The Unjournal · Method audit · September 12, 2026

From research scores to welfare estimates

What the current impact scores mean, where their apparent precision comes from, and how we could move toward explicit estimates of research value through policy, funding, tools, and downstream welfare.

Status: This explains and critiques the current system. A six-paper calculation pilot and 18-paper screening expansion are now available. They do not replace the current 0–10 scores: none of the 24 cases yet supports a defensible total expected-welfare estimate for the research product.
View the current impact leaderboard Read the current scoring prompt Explore the welfare and research-value pilot See the proposed next step Historical human calibration

Audio briefing and transcript

The listening-oriented transcript is ready. The narrated audio is still being prepared.

1. The key finding

No numerical welfare calculation produced the current 7.4 or 6.8 scores. A model read paper material and assigned each number against a verbal rubric. The site can reproduce the ordering and the separate weighted arithmetic used on the main evaluation dashboard, but it cannot reconstruct 7.4 from a chain of policy influence, outcome changes, and welfare weights because that calculation was never performed or recorded.

41 of 1,219public records had revised impact assessments in the audited snapshot
Ordinal judgmentsthe 0–10 values are rankings under a rubric, not welfare units or probabilities
Separate questionspotential research impact and the value of commissioning an evaluation need different counterfactuals

This explains how a paper described as having a “strong but not exceptional” welfare case could rank second: it was second within an incomplete 41-paper subset, and only one paper in that subset scored above 7.4. That rank does not establish global superiority over the other 1,178 records or over research outside the collection.

2. How the displayed numbers were obtained

PaperImpact scoreRank among 41AI weighted evaluationEarlier holistic score
Reducing Prescription Errors7.4/10283.9/10078/100
Multigenerational Effects of Legal Access to the Pill6.8/10Joint 678.6/10077/100
Tracking Inequality6.8/10Joint 678.9/10079/100
A Job I Like or a Job I Can Get6.0/10Joint 1575.7/10056/100

The columns are different quantities. Rank is an ordering within the assessed subset. The weighted evaluation score helps choose which papers might benefit from evaluation. The earlier holistic score is another model judgment. None of the four papers had a published human impact-category rating contributing to its impact rank.

Show the evaluation-score formula and arithmetic

The dashboard’s default AI evaluation lens uses:

10 × (0.30 × impact + 0.30 × evaluation neglectedness + 0.20 × timing + 0.10 × methodology + 0.10 × influence)

Each criterion is on 0–10. Missing criteria cause the remaining weights to be renormalized.

  • Prescription errors: 10 × (0.30×7.4 + 0.30×9.0 + 0.20×9.5 + 0.10×8.4 + 0.10×7.3) = 83.9
  • Pill access: 10 × (0.30×6.8 + 0.30×8.4 + 0.20×9.3 + 0.10×7.6 + 0.10×6.8) = 78.6
  • Tracking Inequality: 10 × (0.30×6.8 + 0.30×8.5 + 0.20×9.0 + 0.10×8.0 + 0.10×7.0) = 78.9

This arithmetic explains the evaluation scores. It does not explain the underlying impact values. High timing and evaluation-neglectedness judgments can sustain a high evaluation priority even when impact is more moderate.

3. What the examples show

The stored explanations identify plausible pathways, but key causal and distributional quantities remain unknown. Open each example for the evidence credited and the reason its numerical placement remains uncertain.

Reducing Prescription Errors Through Information Intervention

The study is in India. Its abstract describes 2.81 million prescriptions, 1,700 physicians, an 8.6% reduction in drug-interaction errors, and extrapolations of $4.8 million in annual hospitalization savings and 134 potentially saved lives. The experimental outcome is errors; the mortality figure is an extrapolation rather than experimentally observed deaths prevented. The abstract’s comparator receives no information, which does not establish superiority over the best alternative alert system. Read the abstract.

Why it merits attention: a specific implementer, concrete design choice, large operational setting, and potentially serious health stakes.

What is missing: the fraction of rollout caused or accelerated by the research, delivery costs, beneficiary circumstances, the validity and time basis of the mortality extrapolation, and gains over competing designs. These unknowns do not establish 7.4.

The Multigenerational Effects of Legal Access to the Pill on Infant Health

This concerns historical US policy. The abstract reports 12–16 grams higher birth weight and a 2.7–3.4% relative reduction in low birth weight among affected births, with gains concentrated among Black families. It estimates roughly $114–115 million in medical savings for one cohort. These are estimates of a policy’s historical effects, not the welfare created by publishing this paper. Read the NBER page.

Why it may contribute: evidence on a multigenerational health channel could affect how access rules are understood.

What is missing: how much this evidence changes present decisions beyond an established literature, the beneficiaries’ consumption distribution, the incidence and real opportunity cost of medical savings, and checks against double-counting health, earnings, and financial benefits.

Tracking Inequality: Teachers and the Allocation of Educational Opportunities

This study is in Italy, not the US. It studies feedback to teachers that changes recommendations and demanding-track enrollment for some high-achieving disadvantaged students, particularly boys, without detectable short-run academic harm. It does not establish lifetime earnings or welfare effects. Read the trial registration or the NBER page.

Why it may contribute: a feasible intervention addresses a specific allocation mechanism affecting disadvantaged students.

What is missing: beneficiary consumption, durable gains, additional adoption caused by the research, costs, effects on other students, and whether selective-track gains are net gains or partly redistributed advantages. Italy is a high-income setting, so the distributional concern remains.

4. A proposed welfare calculation

The calculation should follow the causal chain from research to changed choices, then to outcomes, then to welfare. For research that changes a funding allocation:

Researchchanges beliefs or tools
Decisionsfunding or policy changes
Outcomespeople or animals are affected
Welfarebenefits, harms, and survival
Research contribution ≈ relevant budget × research-caused change in allocation share × incremental outcome per dollar × welfare per outcome − net costs and harms

For a rule or technology without a useful budget measure:

Research contribution ≈ eligible population × research-caused change in coverage × incremental outcome per covered person × duration × welfare weight − net costs and harms

The counterfactual matters at every step. The research-caused change is relative to what would happen without the paper. For reallocations, the relevant return is the difference between the funded option and the displaced alternative. For large reallocations, integrate over a funding-response curve rather than assuming a constant return. Under uncertainty, estimate the expectation of the whole product because uptake, reach, effectiveness, and cost can be correlated.

Worked funding example — entirely hypothetical

Suppose a paper causes a two-percentage-point shift in a $10 million budget toward an option delivering 0.005 more healthy life-years per dollar than the displaced option:

$10,000,000 × 0.02 × 0.005 = 1,000 healthy life-years

If implementation costs displace benefits worth 200 healthy life-years, net value is 800. The answer changes tenfold if the attributable shift is 0.2 rather than two percentage points. A named policymaker or large program budget alone tells us little about that parameter.

Illustration using the prescription paper’s reported extrapolation

Take 134 potential deaths averted only as a source-reported extrapolation at the paper’s reference scale. Suppose, solely for sensitivity analysis, the research causes uptake equal to 20% of that scale, half the extrapolated effect survives scrutiny and delivery constraints, and each death prevented yields 30 healthy life-years:

134 × 0.20 × 0.50 × 30 = 402 healthy life-years

At 2% additional uptake, the answer is 40.2. Neither result is a paper estimate, confidence interval, or reconstruction of 7.4. The uptake shares, adjustment, and remaining healthy years are invented inputs that show where the uncertainty lies.

Research impact and evaluation value need separate counterfactuals

A paper’s research contribution compares decisions with its findings or tools against decisions without them. The value of commissioning an evaluation compares publication and use of that paper with an evaluation against publication and use without one. An evaluation may correct errors, reduce misplaced confidence, improve methods, or accelerate warranted uptake. Neither the current impact score nor the dashboard’s weighted evaluation formula computes that value.

5. Distributional and moral assumptions

Income and consumption

Use real consumption per person on a common purchasing-power and price-year basis when possible, and apply the calculation over the beneficiary distribution rather than a country average. For curvature η, compare log utility (η = 1) with more curved alternatives. The table uses hypothetical annual consumption of $1,000 and $50,000.

CurvatureValue ratio: same 1% gain, poorer vs richerValue ratio: same small dollar gain
η = 1 · log utility50×
η = 1.5 · more curvature7.07×353.6×
η = 2 · greater curvature50×2,500×

Under log utility, the same proportional gain has equal value at either consumption level; an equal dollar gain is already valued more for the poorer person. Greater curvature favors the poorer recipient in both comparisons. A sensible pilot would show η = 1, 1.5, and 2 side by side. Using 1.5 as a central scenario would be an exploratory assumption, not an established empirical coefficient or a settled moral preference.

Health, income, animals, and catastrophic risk

Keep an outcome ledger in natural units first: consumption changes, life-years, health severity and duration, animal experiences, and changes in catastrophe risk. Convert only under named assumptions.

Human health: as a pragmatic first bridge, GiveWell currently assigns one unit to doubling a person’s consumption for a year and 2.3 to averting a year lived with disability of weight one. This would make one reference consumption doubling about 0.435 health-year equivalents. It is a modeling convention, not a scientific equivalence. Comparable human health improvements should initially receive equal weight across countries; the incidence of costs and consumption consequences can still vary.

Animals: Rethink Priorities’ Moral Weight Project offers a starting point for cross-species comparisons. Its published chicken welfare-range estimates have a median of 0.332 and 5th and 95th percentiles of 0.002 and 0.869 relative to the stated human reference. These figures already adjust for sentience probability and subjective-experience rate. A median is not a mean, and the range figures must use definitions compatible with the modeled welfare change.

Catastrophic and existential risk: use the research-attributable change in risk multiplied by the welfare loss avoided, with both defined over the same horizon. Show present-generation and longer-future scenarios separately, with future population, survival paths, welfare assumptions, population ethics, and time preference stated. The paper-to-risk causal link is likely to dominate uncertainty.

What the Rethink Priorities Portfolio Builder contributes

The Portfolio Builder separates funding-response curves, uncertain payoffs including failure and backfire, and different decision procedures. That is useful architecture for this project. It is a donor-allocation tool, however; it does not estimate the extra policy or funding influence caused by an individual research paper. That attribution layer still has to be modeled.

Consumption curvature, diminishing returns to cause funding, and aversion to risky social payoffs are three different concepts and should use separate parameters. Published Portfolio Builder cases describe their inputs as illustrative; those allocations should not become default paper rankings.

6. What the pilot found

We have now run the proposed six-paper pilot and applied the same protocol at screening depth to 18 more papers. None supports a defensible total expected-welfare estimate with the evidence currently assembled. Four pilot cases do support useful conditional calculations, and the wider screen makes the missing research-to-decision links explicit.

  1. Identify the actual decision and the without-paper alternative.
  2. Record reachable scale, the change attributable to the research, outcome effects, beneficiary circumstances, duration, costs, displacement, and plausible harms.
  3. Tag each input as source estimate, analyst assumption, or unknown, with units and provenance.
  4. Report natural outcomes first, then welfare results under several named scenarios.
  5. Show which uncertain parameters reverse the ranking. Preserve “needs evidence” where a numerical total is not defensible.
  6. Compare those outputs with historical human priorities and commissioned evaluations, looking for interpretable patterns by beneficiaries, field, method, outcomes, observed versus hypothetical behavior, neglectedness, and decision pathway.

Explore the six detailed cases and 18-paper expansion, including the input provenance, low/central/high scenarios, natural outcomes, separate evaluation-value judgments, and downloadable JSON.

Historical choices can help reveal useful heuristics, but they should not be treated as ground-truth welfare labels. Selection reflects changing strategy, availability, evaluator capacity, and the value of evaluation as well as expected research impact. Comments may clarify those mechanisms better than a score alone.

The public 0–10 numbers should remain labeled as ordinal model assessments until a transparent mapping to quantified outcomes is specified and tested. Bands or whole numbers would better match the present precision.

Deferred viewer controls

A later version could let viewers adjust consumption curvature, consumption-to-health conversion, species welfare ranges and sentience beliefs, future-welfare assumptions, and risk attitudes. Personal settings should change that viewer’s comparison without overwriting shared evidence or ratings. Empirical disagreement, welfare-capacity uncertainty, and moral preference should remain visibly distinct.

7. Learning from past human prioritization

A separate, privacy-safe analysis asks what The Unjournal’s recorded priorities and evaluation choices reveal about its past practice. The source snapshot contains 325 Coda prioritization records, including 238 with at least one human rating and 57 that became published evaluations. A terminal-outcome model uses 194 rows whose current status can conservatively be treated as either selected or deprioritized/blocked; pending cases are excluded.

325 recordshistorical prioritization entries in the Coda snapshot
238 ratedrecords with at least one human prioritization rating
57 publishedrecords marked as published Unjournal evaluations

The initial descriptive patterns are suggestive rather than rules. For example, 4 of 17 terminal records coded as wealthy-country work were selected, compared with 32 of 73 coded as lower- or middle-income-country work. Papers coded as using observed choices were selected in 9 of 19 cases, compared with 2 of 7 for hypothetical preferences. The cells are small, the coding is preliminary, and topic, methods, publication status, timing, evaluator availability, and other constraints are confounded.

A five-fold model using the prior human rating alone distinguished the terminal outcomes reasonably well in this sample (AUC 0.80; Brier score 0.174). Adding the currently available metadata did not improve out-of-sample performance (AUC 0.769; Brier score 0.190). This is a useful warning against turning a few descriptive correlations into scoring rules.

This exercise does not recover moral weights or a welfare function. Historical ratings bundle research impact, value of an additional evaluation, credibility, timing, open-access and author constraints, evaluator supply, and changing organizational priorities. The model can describe choices and generate hypotheses for review; income curvature, cross-species weights, future-life assumptions, and other moral choices must remain explicit and separate.

Read the modeling and validation plan · Download the privacy-safe aggregate results

Sources and reproducibility

Scores and rank refer to the verified September 12, 2026 snapshot. All worked welfare calculations on this page are hypothetical sensitivity examples. They do not replace or retroactively generate the current ratings.