Scoring methodology

Scores estimate evaluation priority and global impact potential. They are not ratings of research quality. They are provisional and AI-generated unless a human rating is identified.

The dashboard combines several scoring layers. Its controls let you choose a display lens, rating source, or custom weights. For some papers, a preliminary Opus re-evaluation changes the score and screening label while its toggle is on. The global welfare impact rubric defines the revised impact assessments.

Discovery sources include NBER, arXiv, CEPR, EA Forum research links, Semantic Scholar, OpenAlex, SSRN, RePEc, and selected AI-research organizations. Discovery and scoring normally refresh twice weekly; feedback aggregates can refresh up to hourly.

Titles, authors, and source-provided abstracts come from discovery sources. Curated summaries and discovery notes have separate labels. Evaluation rationales, suggested claims, and relevance matches are AI analyses; check their factual claims against the linked research.

Core principle

Prioritization = expected value of commissioning an evaluation, not quality endorsement. A prominent but flawed paper may score higher than a rigorous but obscure one, because independent evaluation adds more value there.

Two display lenses (controls on the dashboard)

The list can be scored and sorted through either of two transparent weightings of the six sub-scores — toggle Evaluation-relevance weighting in the controls, and unfold Score weights to see the live weights:

CriterionResearch relevance (readers)Evaluation priority (UJ, default)
Potential decision relevance40%30%
Methodological potential25%10%
Real-world influence20%10%
Prominence / attention15%0%
Timing value0%20%
Neglectedness / value of evaluation0%30%

Human and synthesis ratings: Each rater's latest rating for a paper counts once. Direct 0–100 ratings are used as entered; quick votes map from Strong No through Strong Yes to 10, 30, 50, 70, and 90. Human-only is a weighted mean: a quick rating has weight 1; a full rating without detailed written discussion has weight 2; a public full rating with detailed discussion has weight 3; and a team full rating with detailed discussion has weight 4. “Detailed” currently means at least 120 characters or 20 words. The AI score has weight 3 in the synthesis, measured in the same quick-rating equivalents. These weights are admittedly ad hoc and should be revisited as evidence accumulates. Public submissions are not identity-verified, so treat small human samples as indicative rather than a team decision.

Research relevance ranks work by importance, rigor, and relevance for readers looking for useful research. Evaluation priority (the default) adds timing and value-of-evaluation (neglectedness) — the expected value of commissioning an independent Unjournal evaluation. If a sub-score is missing for a paper, its weight is dropped and the rest renormalized; out-of-scope papers are floored.

Model two-track scoring

Underneath the lenses, the AI model treats prominent and less-prominent work differently. Prominence is not a quality bonus for obscure work; it mainly tells us whether independent evaluation could matter because the work is already influential, or because a neglected but solid piece might otherwise be missed.

CriterionProminent workLess-prominent work
Potential decision relevance40%30%
Timing value25%15%
Real-world influence20%20%
Methodological potential10%25%
Prominence / existing attention5%Context only

For prominent work (NBER, CEPR, World Bank, top journals), decision-relevance dominates; methodological weakness can make public evaluation more valuable rather than less. For less-prominent work, the data, identification strategy, clarity, and evaluability matter more, because serious flaws make evaluation less likely to change decisions.

Scoring rubrics (0–10 each)

1. Global Decision-Relevance (most important)

9–10: Directly informs active decisions by major funders or policymakers (GiveWell cost-effectiveness, WHO policy, climate treaty design). Specific organizations can be named.

7–8: Addresses a recognized global priority with clear policy implications, but the link to specific decisions is less direct.

5–6: Relevant to global welfare in a general sense. Interesting for the field but specific decision-relevance is moderate.

3–4: Tangentially related to global priorities. Primarily academic interest.

1–2: No clear connection to decisions affecting global welfare.

Field-specific: Development and global-health work should connect to intervention choice, delivery, or scale; AI-governance work should bear on a specific governance or institutional choice rather than merely discuss AI; animal-welfare work should identify the affected population, welfare margin, and plausible decision-maker; climate work should change an emissions, damages, adaptation, or policy-design margin; forecasting and meta-science work should connect accuracy or research practice to a concrete institutional decision.

2. Prominence

9–10: NBER working paper, top-5 journal, Nobel/Clark laureate, >500 citations, major media coverage.

7–8: Well-known department, strong journal, recognized researcher, >100 citations.

5–6: Decent institution, field journal, established researcher.

3–4: Less-known institution, newer researcher, workshop paper.

1–2: Unknown author, self-published, no institutional backing.

Note: NBER/CEPR/World Bank/IMF sources are treated as high-attention signals, not proof of quality or publication stage. Some NBER papers are later published in journals, and this should be checked when timing matters. Prominent flawed work can be valuable to evaluate because people may already be using it.

3. Real-World Influence

9–10: Already cited in policy documents, GiveWell/Open Phil analyses, government reports. Named organizations are using this.

7–8: Likely to influence decisions soon. In an active policy debate. Authors have policy connections.

5–6: Could influence decisions if findings hold up. Relevant to active debates but not yet cited.

3–4: Academic contribution with indirect policy relevance.

1–2: Purely academic exercise with no clear path to influence.

4. Timing Value

9–10: Working paper/preprint released in last 6 months. No peer review yet. Authors actively seeking feedback.

7–8: Working paper 6–18 months old. Under review but not yet published.

5–6: Recently published (1–2 years) in a venue where more review would add value. R&R at journal.

3–4: Published 2+ years ago but still influential. Adds transparency but less urgency.

1–2: Old published work with established peer review. Feedback largely moot.

By methodology: RCTs & field experiments benefit most from early feedback (pre-registration, pre-analysis). Policy reports have narrow windows. Theoretical work is less time-sensitive.

5. Methodological Potential

For prominent work: This is a secondary consideration. If it’s prominent and decision-relevant, score 7+ and move on. Quality assessment is for the evaluation stage.

For less-prominent work (the tie-breaker):

9–10: Innovative methodology, strong identification strategy, credible real-world outcome data where possible, reproducible analysis with shared code/data.

7–8: Solid methods appropriate for the research question.

5–6: Acceptable methods, nothing particularly noteworthy.

3–4: Methodological concerns that would make evaluation difficult.

1–2: Not really quantitative. Literature review, opinion piece, or purely conceptual.

Field-appropriate standards are illustrative, not mechanical. We should not penalize fields where RCTs are impossible, but observed choices and measured outcomes usually deserve more weight than recall or hypothetical stated-preference evidence when both are available.
Development/health: RCTs, DiD, regression discontinuity, IV
Environmental/climate: Integrated assessment models, panel data, natural experiments
AI governance: Mixed methods, surveys, formal models
Animal welfare: Revealed-preference or behavioral evidence where available; stated preference, DCEs, and welfare calculations where direct evidence is limited
Political science: Quasi-experimental, panel data, surveys
Macro/trade: DSGE, gravity equations, synthetic control

AI screening recommendations

The labels below come from the original AI screening score. The large displayed score uses the selected weighting and rating source, so it can differ from the screening label. Neither is a team decision.

Score rangeRecommended actionWhat it means
75–100AI shortlistThe model recommends human assessment next. This does not mean the paper has a human rating or that Unjournal has decided to evaluate it.
50–74WatchlistPotentially relevant, but not currently a high-priority evaluation candidate.
25–49DeprioritizeThe AI screening recommendation is to give other candidates attention first.
<25Out of scopeNot quantitative social science, or fundamentally outside UJ coverage.

Calibration

The scoring prompt uses examples and guidance from past Unjournal prioritization judgments. This does not establish accuracy on new papers: the pilot comparison with 12 previously rated papers found a mean absolute AI–human difference of 17.3 points. Those historical ratings are not a blinded holdout test. You can read more about UJ’s prioritization process.

Prompt and run stability

Calibration and stability answer different questions. Calibration asks whether AI ratings line up with human judgments. Stability asks whether the same papers keep similar ratings when the model is run again or a substantively equivalent instruction is reworded. A system can be stable but wrong, so stability is a reproducibility diagnostic rather than an accuracy claim. Our pilot covers 24 selected papers with 192 repeated ratings and remains exploratory; it is not a full-dashboard reliability estimate. Read the report and methods note.

Four-stage pipeline

  1. Suggesting — A paper is suggested (by AI or human) with a 0–100 rating and discussion of relevance
  2. Assessing — A second team member gives an independent rating (without seeing the first)
  3. Voting — If avg rating ≥ 65%, the field group votes (Strong Yes to Strong No)
  4. Evaluation — An evaluation manager commissions 2+ public evaluations via PubPub