What the research prioritization scores mean, and how we might build a real welfare model AI-generated narration, prepared September twelfth, twenty twenty-six. This is an audio adaptation of the written score audit, expanded to include a possible next stage: learning from The Unjournal's past human prioritization decisions. It describes a proposal, not an adopted scoring system. The live ratings have not been recalculated using the welfare model discussed here. Let me begin with the most important correction. The current impact numbers, such as seven point four or six point eight out of ten, are judgments made by a language model against a verbal rubric. They were not produced by a numerical calculation of policy influence, changed funding, health effects, income gains, animal welfare, or catastrophic risk. We can reproduce the arithmetic that combines the existing criteria into an overall evaluation-priority score. But we cannot reconstruct a quantitative welfare calculation that yielded seven point four rather than six point eight, because no such calculation was performed or recorded. That distinction explains some of the confusing results on the impact page. Consider the paper called Reducing Prescription Errors Through Information Intervention. Its revised impact score is seven point four, and it appears second in the impact ranking. The stored explanation says that its global-welfare and value-of-information case is strong, but not exceptional. Those statements are not literally inconsistent. The rubric describes scores from seven to eight as strong, and nine to ten as exceptional. The problem is what the ranking appears to imply. The paper is second only among the forty-one records that currently have revised impact assessments, not second among all one thousand two hundred and nineteen papers in the collection, and certainly not second among research opportunities in general. Different papers in that group were also assessed through different model routes and with different levels of reasoning effort. So the defensible interpretation is modest: among the currently reassessed subset, this model gave only one paper a score above seven point four. The rank is not evidence that the prescription paper has the second-highest expected global welfare impact. The decimal point also suggests more precision than the process supports. The main dashboard performs a separate, real piece of arithmetic. It assigns thirty percent weight to impact, thirty percent to how neglected an additional evaluation would be, twenty percent to timing, ten percent to methodology, and ten percent to likely influence. Each input is on a zero-to-ten scale, and the result is multiplied by ten. For the prescription paper, the inputs are seven point four for impact, nine for evaluation neglectedness, nine point five for timing, eight point four for methodology, and seven point three for influence. The weighted result is eighty-three point nine out of one hundred. That calculation explains the displayed evaluation-priority score. It does not explain where the underlying impact score came from. It also shows why a paper with only a moderately strong impact assessment can rank highly for evaluation: high timing and neglectedness scores pull the total upward. Now consider what the model saw in three concrete cases. The prescription study was conducted in India. The abstract reports two point eight one million prescriptions and seventeen hundred physicians. It finds an eight point six percent reduction in drug-interaction errors. It also extrapolates four point eight million dollars in annual hospitalization savings and one hundred and thirty-four potentially saved lives. This is a relatively concrete path to impact. There is a specific intervention, a plausible implementer, and a measurable change in prescribing behavior. That deserves attention. But the mortality number is an extrapolation, not a count of deaths experimentally shown to have been prevented. The study does not, from the abstract alone, establish that the intervention is better than the best alternative alert system. And India as a country label does not tell us the consumption or health burden of the actual beneficiaries. To estimate welfare impact, we would still need to know how much additional adoption the research causes, what would otherwise happen, what deployment costs, how severe the prevented errors are, how durable the effect is, and who receives the benefits. Those missing quantities make the exact score of seven point four hard to defend. The second example is The Multigenerational Effects of Legal Access to the Pill on Infant Health. This is about historical policy in the United States. The paper reports birth-weight gains and a reduction in low birth weight, concentrated among Black families. It estimates roughly one hundred and fourteen to one hundred and fifteen million dollars in medical savings for one cohort. The model credited a potentially long-lasting health channel and relevance to access rules. But a finding about a policy's historical effects is not the same as the present welfare contribution of publishing one more paper. We need to estimate how much this evidence changes a current or future decision beyond what was already known. Calling the beneficiaries disadvantaged within the United States does not settle the global distributional issue. We do not know their relevant lifetime consumption distribution. And the medical-savings estimate is not equivalent to cash delivered to low-income households. The savings might accrue to households, insurers, providers, or taxpayers, and some reported expenditure may not correspond to real resources released for another use. The third example is Tracking Inequality: Teachers and the Allocation of Educational Opportunities. A factual correction matters here: the study is in Italy, not the United States. Italy is still a high-income country, so the underlying distributional concern remains. The intervention gives teachers information intended to correct beliefs about high-achieving students from disadvantaged backgrounds. It changes recommendations and enrollment in more demanding academic tracks for some students, especially boys, without detectable short-run academic harm. The model credited a feasible intervention and a clear mechanism. But the path from a changed recommendation to greater global welfare runs through several further steps: enrollment, completion, skills, employment, consumption, health, subjective well-being, and possible effects on other students. Several of these links were projected rather than observed in the supplied evidence. A more demanding school track is not automatically a net social gain if places are fixed or if much of the earnings advantage is redistributed from someone else. Again, the paper may be useful. The current score does not quantify its value. So what would a more meaningful calculation look like? The basic causal chain is this. Research changes beliefs. Changed beliefs alter funding, policy, implementation, or the tools used to make later decisions. Those choices change outcomes for humans and animals. Finally, a stated welfare framework assigns value to those outcomes. For a funding decision, a simple first approximation is: the relevant budget, multiplied by the research-caused change in the allocation share, multiplied by the additional outcome per dollar relative to the displaced alternative, multiplied by welfare per unit of outcome. Then subtract implementation costs, displacement, and harms. For a policy or technology where a budget is not the natural measure, use the eligible population, multiplied by the research-caused change in coverage, multiplied by the incremental outcome for each additionally covered person or animal, multiplied by duration and a welfare weight. Again, subtract costs and harms. The phrase research-caused is doing a great deal of work. A large program and an influential-sounding topic do not establish that this particular paper changes a decision. The comparison must be against the world in which policymakers, funders, and researchers have all the other available evidence but not this paper, or not this evaluation of the paper. It also matters whether the work contributes indirectly by building a better research tool, measurement method, model, or body of understanding. That route fits the same framework, but with additional links. The new tool improves later evidence. Better evidence changes later beliefs. Those beliefs improve later choices. Those choices change welfare. An indirect pathway may be large, but each extra link adds uncertainty and a chance of double-counting contributions made by other work. The value of commissioning an evaluation is another distinct counterfactual. It is the difference between publication and use of the research with the evaluation versus without it. An evaluation might catch an important error, reduce misplaced confidence, improve implementation, or accelerate justified adoption. A paper with large possible impact is not automatically the paper for which one more evaluation has the highest value. Here is a deliberately hypothetical funding example. Suppose a paper causes a two-percentage-point shift in a ten-million-dollar budget. Suppose the destination produces point zero zero five more healthy life-years per dollar than the activity it displaces. The calculation gives one thousand additional healthy life-years. If implementation costs displace benefits worth two hundred healthy life-years, the net is eight hundred. But if the paper changes the allocation by only zero point two percentage points, the answer falls by a factor of ten. This illustrates why a named policymaker or a large budget does not, by itself, establish high value of information. For the prescription study, imagine using the paper's extrapolation of one hundred and thirty-four potentially prevented deaths merely as a source anchor. Suppose further that the research causes uptake equal to twenty percent of the scale behind that estimate, that half of the extrapolated effect survives closer scrutiny and delivery constraints, and that each prevented death yields thirty healthy life-years. The illustrative result is four hundred and two healthy life-years. At two percent additional uptake, it is forty point two. Neither number is an estimate from the paper. The uptake, adjustment, and life-years are invented sensitivity inputs. Their purpose is to reveal which assumptions determine the answer, not to create a new precise rating. Distributional weighting is especially important for income and consumption outcomes. A reasonable first pass is to use real consumption per person on a common purchasing-power basis. Income is an imperfect proxy, and a dollar in a government budget is not automatically a dollar of household consumption. With log utility, a one-percent consumption gain has the same welfare value for a rich person as a one-percent gain for a poor person. An equal dollar gain is worth more to the poorer person, because it represents a larger proportional change. David's concern is that log utility may still give too much relative weight to proportional gains among people near the top of the global income distribution. We can represent that concern using a constant-relative-risk-aversion utility function with more curvature. For an intuitive comparison, take annual consumption of one thousand dollars and fifty thousand dollars. Under log utility, the same one-percent increase has equal value at both levels. With a curvature parameter of one point five, the poorer person's one-percent gain is worth about seven times as much. With curvature of two, it is worth about fifty times as much. For the same small dollar gain, log utility favors the poorer person by roughly fifty to one. Curvature of one point five raises that ratio to about three hundred and fifty-four to one. Curvature of two raises it to about two thousand five hundred to one. These are illustrations, not recommended final moral weights. A useful pilot would display results for curvature values of one, one point five, and two. If a central exploratory case is needed, one point five captures the stated preference for somewhat more curvature than log utility while keeping the alternatives visible. Health, suffering, animal welfare, and catastrophic risk require additional assumptions. The safest practice is to keep a ledger in natural units first: consumption changes, healthy life-years, disability severity and duration, animal experiences, and changes in catastrophe risk. Only then convert them under named moral assumptions. For a rough consumption-to-health bridge, GiveWell currently assigns one unit to doubling one person's consumption for one year and two point three units to averting one year lived at a disability weight of one. Borrowing that ratio makes a reference consumption doubling worth about point four three five health-year equivalents. This is a pragmatic pilot convention, not an established scientific equivalence. For comparisons involving nonhuman animals, Rethink Priorities' Moral Weight Project is a useful starting point. It asks about welfare capacity under explicit assumptions, rather than treating species weights as a poll of human sympathy. Its published chicken estimates span a very wide range. That uncertainty should be carried into the result, not hidden in a single score. The relevant animal calculation would multiply the number of animals by the research-caused change in coverage, the years affected, the improvement as a fraction of the relevant welfare range, and an adjusted cross-species welfare-range estimate. One must specify the neutral point, distinguish the negative portion from the full welfare range, and avoid applying sentience adjustments twice. For catastrophic and existential risk, the simplest form is the research-attributable change in risk multiplied by the welfare loss avoided. Both terms must use the same time horizon. Results should separate present generations from longer-future scenarios and state assumptions about population, well-being, survival, population ethics, and discounting. In practice, the hardest quantity may be whether the paper changes the risk by even one part in ten million. A huge possible loss does not make that causal link known. Rethink Priorities' Portfolio Builder provides a useful architecture for handling diminishing returns to funding, uncertain payoffs, failure, and possible backfire. But it is an allocation tool. It does not estimate how much an individual paper changes a policy or funding decision. Our model needs that attribution layer before the portfolio machinery becomes useful. There is also a second, complementary project: use The Unjournal's historical prioritization record to learn what human decision-makers have actually treated as important. The available evidence may include explicit human priority ratings, papers placed on shortlists, papers actually commissioned for evaluation, structured criteria, and comments made during prioritization. These sources are not interchangeable. A paper may fail to be commissioned because of evaluator availability, timing, duplication, or operational constraints rather than low expected welfare value. Comments may contain the clearest reasoning, while numerical ratings may reflect older criteria or inconsistent use of the scale. The first stage should therefore reconstruct the decision history before fitting a model. For each paper and decision point, record what information was available at the time, which human judgments were expressed, whether the paper advanced, and any operational reason for the outcome. Separate individual ratings from team decisions, and separate a prioritization decision from later evidence about whether the paper was commissioned. Next, convert the source material into candidate explanatory features. These can include beneficiary geography and consumption, scale, policy or funding pathway, neglectedness, tractability, time to impact, methodological credibility, observed behavior versus hypothetical preferences, representativeness, external validity, relevance to catastrophic risk, animal populations affected, decision-maker engagement, and the value of another evaluation. Some features can be measured directly. Others should be carefully coded from comments, with the original text retained for audit. For example, we can test whether studies of hypothetical preferences, especially in non-representative samples, received lower ratings than otherwise similar studies based on observed choices. We can test whether health-policy studies in wealthy countries usually fall into a lower range, and identify the circumstances in which they do not: pandemic relevance, implications for cheap and scalable interventions elsewhere, unusually neglected policy margins, or unusually strong evidence that a decision will change. Then fit several deliberately interpretable models. Start with descriptive tables and matched comparisons. Follow with a regularized ordinal model for human ratings and a separate model for advancement or commissioning. A small decision tree or rule list can surface thresholds and exceptions. More flexible prediction can be used as a diagnostic, but should not become the explanation. Any inferred rule should come with sample size, uncertainty, time period, and counterexamples. Cross-validation should split by time, and ideally by related paper family, so near-duplicates do not inflate apparent performance. Comments used to create a feature must not leak the final decision label back into that feature. Most importantly, this exercise learns a descriptive model of past human prioritization. It does not automatically reveal the correct moral weights. Historical choices combine values, empirical beliefs, institutional remit, evaluator supply, and path dependence. The most useful result may be a set of discrepancies: places where stated welfare principles predict one ordering but past choices consistently imply another. Those discrepancies can inform a revised AI prompt. Some stable human heuristics may become explicit rubric rules. Other patterns may reveal biases we want to correct rather than reproduce. The resulting AI system should show both layers: a forecast of how The Unjournal has historically prioritized similar work, and a welfare analysis built from explicit causal and moral assumptions. A practical implementation can proceed as a small pilot. First, audit six contrasting papers in full: the three discussed here, a direct lower-income consumption intervention, an animal-welfare intervention, and a concrete catastrophic-risk paper. For each one, record every welfare input with units and label it as a source estimate, an analyst assumption, or unknown. Second, reconstruct a clean historical dataset of human ratings, advancement, commissioning, and comments. Document missingness and operational constraints. Third, fit interpretable descriptive models and extract candidate rules with counterexamples. Fourth, compare those historical rules with the explicit welfare calculations. Review where they agree, where they diverge, and whether the divergence reflects legitimate institutional considerations or something we want to change. Fifth, revise the AI prioritization process only after the pilot makes the intended calculation and its uncertainty visible. Until then, the zero-to-ten impact ratings should be presented as ordinal model assessments. Bands may be more honest than decimal points. A later interface could let viewers adjust assumptions about consumption curvature, human and animal moral weights, sentience, future welfare, and risk. That extension should keep empirical inputs separate from moral choices and should never overwrite shared evidence or other viewers' settings. It remains a note for later, not part of the present implementation. The goal is ambitious but coherent. We want to estimate the value of research through the beliefs it changes, the decisions those beliefs improve, and the resulting effects on well-being, suffering, flourishing, and survival. We should also recognize contributions that operate through better tools and deeper understanding. The current ratings do not yet do that calculation. The next step is not to pretend they do. It is to build a small, auditable model in which the causal chain, the welfare assumptions, the missing evidence, and the sources of disagreement are all visible. End of AI-generated narration.