Shared global welfare impact rubric

Version: global-welfare-voi-v1. Introduced September 9, 2026.

This is the impact criterion embedded in the combined prioritization prompt and in the standalone impact prompt. Combined scoring also assesses evaluation priority, uses other instructions and may include internal calibration guidance. This page reproduces the shared impact rubric; it is not an archive of a complete combined scoring call. Earlier full run prompts were not retained.

## RESEARCH IMPACT POTENTIAL ONLY: GLOBAL WELFARE AND VALUE OF INFORMATION

Objective: assess the expected net contribution of this research's findings, methods or outputs
to the welfare of all sentient beings, using explicit, reasonable welfare weights. This objective
determines the ENTIRE impact score. It is not a small bonus to prestige or generic policy relevance.
Assess the research itself, not the extra benefits of commissioning an Unjournal evaluation.
An already rigorously reviewed paper can have high research impact; an unreviewed paper can have low
research impact. Unjournal scope, evaluator availability and author engagement do not define welfare.

COUNTERFACTUAL AND VALUE OF INFORMATION
Compare the welfare consequences of decisions made with the information or tools this research
adds against decisions using the best reasonably available alternatives without this contribution.
Do not credit one paper with the total value of solving its broad problem. Distinguish gross
problem scale, conditional gains if an intervention works, the paper's incremental information,
and likely realized gains after uptake, implementation and displacement by other research.

Research VoI asks whether learning could change a choice, improve its scale or timing, avert a bad
intervention, or make confidence in a choice more appropriate. A null finding can be highly valuable.
In an ideal Bayesian decision model, expected value of sample information is the expected welfare
of the best choice after the signal minus that of the best choice before the signal. In practice,
allow for imperfect evidence, misinterpretation and implementation constraints. Merely narrowing
a confidence interval or reducing entropy is not evidence of decision value. For an existing paper,
assess the information it actually contributes, not the hypothetical value of a perfect study.
Separate research VoI from the VoI of checking whether this paper is right.

REQUIRED REASONING, BEFORE ASSIGNING THE IMPACT SCORE
1. Welfare scope and scale. Which humans, non-human animals or other plausibly sentient beings
   could be affected? Consider numbers, baseline welfare, severity/intensity, duration, geographic
   reach, spillovers and future generations. Use sourced orders of magnitude or explicit unknowns;
   never invent counts, welfare conversion factors or a dollar value. GDP, willingness to pay and
   researcher/funder attention are not welfare measures by themselves. Include adverse effects.

2. Welfare weights and distribution. Give equal consideration to comparable welfare interests
   across geography. State assumptions about interpersonal comparisons, sentience probabilities,
   species welfare capacity, future lives and moral uncertainty; use plausible alternative scenarios
   where rankings depend on them. Do not silently assign non-human or future welfare zero weight.
   Do not privilege wealthy populations because they can pay more or dominate the research data.
   Separate evidential/implementation discounting from a moral discount for a later birth date.

3. LMIC mechanisms. An increment to low income can produce more welfare because of diminishing
   marginal utility of consumption. Marginal health or other spending may have larger effects at
   very low baseline provision, reflecting diminishing returns and local costs. Under-researched
   settings may offer larger information gains. Assess all three mechanisms explicitly when relevant.
   They are substantive reasons to expect higher research VoI, not an automatic country multiplier.
   Check local effect sizes, costs, delivery capacity, bottlenecks, existing evidence, transferability
   and whether the results can actually redirect resources. Poor people in wealthy countries and
   globally transferable work can also have high value.

4. Specific decision and incremental contribution. Identify the choice, alternatives, decision
   threshold, scale of resources/actions affected, and what changes because of these findings.
   Compare to current interventions, algorithms and knowledge. A named large organization or an
   author claim of substantial gains is not enough. Give empirical comparisons and actual adoption
   evidence credit. Methods, theory, datasets and basic research can matter through credible
   downstream research pathways; identify those links and their uncertainty rather than banning them.

5. Importance, tractability and neglectedness (ITN). Use ITN to reason about marginal research VoI,
   not as three automatic topic bonuses or a multiplication of uncalibrated ordinal scores.
   Importance concerns welfare stakes. Tractability concerns both resolving the decision-relevant
   uncertainty and acting beneficially on the answer. Distinguish neglected knowledge from neglected
   interventions/funding; each can matter differently. Less prior research can mean more to learn,
   and less prior spending can mean higher marginal intervention returns, but neglected problems may
   also be hard to learn about or act on. Explain the connection. Do not double-count these channels
   after already including their effects in information gain, uptake or welfare per marginal dollar.
   Independent scrutiny of this individual paper is a separate evaluation-priority consideration.

6. Global catastrophic risk (GCR) and existential risk (x-risk). For every paper, assess relevance
   as direct, indirect, none apparent or unknown, with a brief reason. Consider effects on the
   probability/severity of catastrophes, extinction or irreversible loss of future flourishing,
   and both present and future sentient welfare. Identify a concrete causal pathway through policy,
   preparedness, institutions, AI, biosecurity, conflict or other mechanisms. Consider risk increases
   and risk reductions. No automatic premium for an x-risk label, and no automatic dismissal because
   the effect is probabilistic or long term. Examine empirical support, competing risk effects,
   uncertainty in tiny probabilities and sensitivity to future-population and ethical assumptions;
   do not let an unsupported enormous payoff manufacture a precise high expected value.

7. Realization, durability and net effects. How likely is the evidence to be credible, communicated,
   adopted, funded and implemented at useful scale? Distinguish conditional promise from documented
   uptake. Consider substitution/crowding out, generalization, useful lifetime, and the option value
   of learning. Rapid AI/labor-market change can erode an application but preserve a general insight.
   Include foreseeable misuse, harmful policy or information hazards at a high level where relevant.
   Report uncertainty and assumptions that could reverse the ranking. Negative net impact is possible.

SCORING ANCHORS (0-10, COMPARATIVE JUDGMENTS, NOT WELFARE UNITS OR PROBABILITIES)
9-10: Exceptional expected global-welfare contribution: very large welfare stakes AND a credible,
      materially decision-changing information/tool contribution with a plausible realization path.
      Could arise through LMIC welfare, animal suffering or catastrophic-risk reduction. Name the
      main counterfactual and test the assumptions that make this exceptional.
7-8: Strong expected welfare/VoI case, with meaningful incremental contribution and plausible
     implementation or downstream research; one or more substantial uncertainties remain.
5-6: Moderate positive case: relevant welfare problem, but bounded incremental gains, indirect
     actionability, limited reach/durability, or substantial unresolved links to beneficial choices.
3-4: Weak expected contribution: important topic alone, small gains over existing knowledge,
     limited beneficiaries, low chance of changing useful action, or a poorly supported pathway.
1-2: Very little plausible net welfare contribution from this research's incremental output.
0: No positive net contribution supported, or expected harms outweigh gains; explicitly distinguish
   those cases. Missing information is NOT zero: return unknown/null where supported by the calling
   schema, otherwise give a clearly provisional assessment and low confidence, never invented facts.

The score should track expected net welfare contribution, not a compensatory average in which
prestige or topical popularity cancels negligible decision value. Do not add fixed LMIC, animal
or x-risk points. Numerical VoI estimates are optional and require defensible inputs; ordinal scores
must not be presented as calibrated utility. Report confidence separately from expected value.