# Learning from The Unjournal's past prioritization decisions Prepared September 12, 2026. This is an empirical modeling plan. It does not change the scorer, ratings, or public dashboard. ## What this exercise can and cannot recover The historical data can help us describe how The Unjournal has prioritized research in practice. It can show patterns such as: - whether work on health policy in wealthy countries tended to receive lower ratings after accounting for methods, likely influence, and evaluation value; - whether papers using observed choices tended to outrank papers using hypothetical preferences, especially with convenience samples; - which combinations of cause area, method, publication status, policy relevance, and evaluation tractability were associated with advancing to evaluation; - where AI scores systematically depart from human ratings and choices. It cannot, by itself, identify The Unjournal's welfare function or the moral weights its team members would endorse. The recorded ratings combine several judgments: potential impact of the research, value of an additional evaluation, methodological credibility, timing, open-access constraints, author engagement, evaluator availability, and changing organizational priorities. A paper may have failed to advance because no evaluator was available, rather than because its expected welfare contribution was low. The right description is therefore **historical decision calibration**, not revealed welfare preferences. The welfare model proposed in the accompanying report should remain a separate, explicit layer. Historical patterns can help test whether its outputs resemble past practice and identify unexplained departures. They should not silently determine income curvature, cross-species weights, the value assigned to future lives, or other moral assumptions. ## Current evidence base These counts come from local snapshots and should be regenerated before modeling. No names or comment text are reported here. | Source | Snapshot and usable coverage | What it can contribute | |---|---|---| | Coda `rsx_prioritization` export | September 8; 325 rows, 324 with a usable title, 238 with at least one prioritization rating | Main historical ratings, workflow status, cause and publication categories, some dates, votes, and extensive private discussion fields | | Rating roles in that export | 210 suggester ratings and 156 assessor ratings | Separate first-pass and second-pass judgments; averaging them would hide systematic disagreement by role | | Current Coda workflow outcomes | 57 rows marked as published Unjournal evaluations; 2 awaiting evaluations; 4 seeking evaluators; 1 selecting an evaluation manager; 6 awaiting author consent or an update | Descriptive progression through the evaluation pipeline, subject to stale and right-censored status data | | Unjournal evaluation database | July 28; 340 research records; 65 papers with 1,024 criterion-level evaluator ratings | Independent confirmation that a substantial set was actually evaluated; 59 papers have a `gp_relevance` evaluation criterion and 64 have an overall evaluation rating | | Link between the two historical sources | 323 of the prioritization rows match a research record; 58 matched papers have downstream evaluator-rating records | A useful selected-paper follow-up sample, though downstream evaluation ratings are outcomes after selection rather than pre-selection features | | Current public dashboard | September 10; 1,219 papers, of which 897 are non-historical AI-scored records | Modern AI scores and structured paper metadata | | Direct current AI–Coda overlap | The existing comparator finds only 2 human-rated papers with a current non-historical AI score | Too small for a direct observational validation without rescoring historical papers | | Existing stratified rescore | A September 12 private output covers 30 human-rated papers. It reports a 15.2-point mean absolute error, AI scores 4.3 points lower on average, and only 36.7% agreement under its action-band mapping | A warning that present AI scores are not yet well calibrated to historical ratings; this is not a welfare validation | | Public dashboard feedback aggregate | September 10; 17 current ratings across 13 papers; no populated category-rating fields in the published summary | Useful as a future prospective validation stream, currently too small to train on | | Calibration-anchor page | 22 anchors across six areas; no saved review responses in the current response sidecar | A review interface and candidate anchor set, not yet an independently reviewed calibration dataset | The historical Coda table has broad but uneven topical coverage. In the September snapshot, the largest labeled groups include 35 development economics/governance papers in LMICs, 34 global-health papers in LMICs, 28 animal-welfare/markets papers, 23 innovation/meta-science papers, 23 social-impact-of-AI papers, and 22 catastrophic-risk/forecasting papers. Eighty-three rows lack the compact cause label. Small groups cannot support credible area-specific coefficients without pooling. Private comments are much more prevalent than public feedback: 296 of 325 Coda rows contain text in at least one discussion field. This is potentially valuable evidence about reasons, but it also creates the largest privacy and leakage risks. The row-level text should remain private, and no comment should be quoted publicly without its existing explicit publication permission. ## What the existing calibration script does `pipeline/calibrate_from_human.py` currently performs a limited score-comparison exercise: 1. It loads suggester and assessor percentiles and averages the available values. 2. It tries to match those papers to current AI-scored dashboard papers by URL, normalized title, or a loose 40-character title prefix. 3. It calculates mean AI-minus-human differences overall, by cause area, publication status, and human-score band. 4. It can rescore a 30-paper sample split across low, middle, and high human ratings. 5. With `--apply`, it writes verbal score adjustments into a snippet that the production prompt loads. This is a useful early diagnostic, but it should not be treated as a calibrated model. The natural overlap is only two papers. The optional rescore supplies the model with title, URL, cause area, and publication status, but no abstract. It stratifies by score band rather than cause area despite saying otherwise in its docstring. Matching does not use DOI and the loose prefix match can create false positives. The saved output retains only summary results and the five largest positive and negative errors, rather than the full auditable sample. The action comparison also applies different cutoffs to human and AI scores: a human score of 65 is called “prioritize now,” while the AI threshold is 75. The resulting 36.7% agreement is therefore informative as an alarm, but not a clean accuracy estimate. Cause-area corrections with only three or four papers are highly unstable. Finally, deriving a correction and appending it to the prompt on the same sample provides no held-out test of whether it improves future judgments. ## Proposed targets Use several targets rather than compressing the history into one label. ### 1. Human priority ratings Model the suggester and assessor ratings separately, then model their difference. These are ordinal percentile judgments about the value of pursuing an evaluation, not cardinal welfare estimates. A hierarchical model can partially pool the two roles while allowing systematic role differences. Keep the explicit average rating as a reporting outcome, but do not use it as the only target. Report disagreement and uncertainty when only one rating exists. ### 2. Advancement through the workflow Treat progression as a sequence of conditional decisions: - moved from initial assessment to final consideration; - selected or advanced to author contact/evaluator search; - evaluation started; - evaluation completed and published. These transitions answer different questions. Completion depends heavily on author permission, evaluator availability, management capacity, and elapsed time. It should not be used as a simple “positive” label for research value. The current export mainly records each paper's latest status. Before fitting a transition model, reconstruct status history and dates if Coda revision data are available. Without that history, use conservative snapshot outcomes and explicitly treat recent papers as right-censored. ### 3. Evaluation-stage relevance The downstream evaluation database contains `gp_relevance` ratings for 59 papers. This can test whether papers selected as globally relevant were later judged that way by evaluators. It remains post-selection evidence from a narrow sample, and it may reflect the quality or findings of the completed work rather than the information available during prioritization. ### 4. Prospective team ratings The dashboard's blind-first rating form should become the cleanest future target. Preserve the rating made before AI scores are revealed, the subsequent revision, and the stated reasoning. The current public aggregate has only 17 ratings across 13 papers, so this is a validation stream for now rather than training data. ## Features to construct Build two feature sets and keep them separate. ### Paper-only features available before the human decision - **Decision pathway:** type of decision, plausible funder or policymaker, whether the decision margin is concrete, and evidence of actual use. - **Affected population:** country or region, income setting, estimated beneficiary consumption where available, disadvantaged subgroup, non-human species, and future generations. - **Reach and transfer:** study population, potential reachable population, geographic portability, institutional requirements, and likely duration of relevance. - **Outcome type:** income/consumption, health, mortality, subjective well-being, suffering, animal welfare, catastrophic-risk probability, knowledge/tool contribution, or an intermediate proxy. - **Evidence type:** observed behavior or administrative outcomes; incentivized choice; stated or hypothetical preferences; expert elicitation; model/simulation; theory; review or meta-analysis. - **Sample quality:** representative population, relevant policy population, convenience/online/student sample, expert sample, sample size, attrition, and missingness. - **Method:** randomized experiment, natural experiment, panel/event study, cross-sectional observational analysis, forecasting exercise, structural model, theory, or qualitative synthesis. - **Evaluation value:** publication stage, existing independent scrutiny, replication/data/code status, evaluability, author engagement evidence, and time sensitivity. - **Welfare-model inputs:** the parts of the multiplicative chain that are directly supported: decision scale, attributable influence, incremental outcome, costs, and outcome recipient. Missing quantities remain missing. Country income group is only a proxy. A study in a wealthy country can affect poor people, non-human animals, pandemic risk, or a neglected global policy. A study in an LMIC does not automatically reach low-consumption beneficiaries or change policy. Include interactions and pathway evidence instead of imposing a country bonus or penalty. ### Reason tags from human comments Code the comments into a compact, preregistered set of reason tags: - high or low potential welfare stakes; - direct versus speculative policy or funding pathway; - reach, transferability, and beneficiary distribution; - observed behavior versus hypothetical preferences; - representative or policy-relevant sample versus convenience sample; - credible identification and measurement; - prior scrutiny, replication, and publication status; - evaluation tractability and likely value added; - author/evaluator/management feasibility; - scope, open-access, or non-paper exclusion; - uncertainty or missing evidence. Preserve direction and target. “Hypothetical preferences are adequate for this forecasting question” is different from “hypothetical preferences are a major limitation.” A generic sentiment score would lose the useful content. The first coding pass should be done on a private, stratified sample of about 50 papers, with two independent coders for at least 20–30 of them. Refine the codebook using disagreement, then measure tag-level agreement. Any model-assisted coding of private comments should run only in an approved private environment. External model calls require the applicable per-paper permission. Public outputs should contain aggregate tag frequencies and model coefficients only, with suppression for small cells; no names, private text, or paper-level reconstructed comments. ## Modeling approach The sample is too small for a high-dimensional black-box model. Start with interpretable, regularized models and treat apparent rules as hypotheses. 1. **Descriptive tables.** Show rating and advancement distributions by broad cause area, region/income setting, evidence type, method, outcome type, and publication stage. Report sample sizes and intervals. Do not publish tiny cells. 2. **Rating model.** Fit an ordinal or bounded-outcome hierarchical regression for suggester and assessor ratings, with shrinkage across cause areas. Include year/cohort and role effects. A simple elastic-net regression is a useful benchmark. 3. **Sequential workflow model.** Fit separate regularized logistic models for each stage transition. Include paper age or censoring time for incomplete cases. A discrete-time hazard model is preferable once status histories are available. 4. **Interaction tests.** Pre-specify a short set tied to actual hypotheses: wealthy-country health × direct global transfer; hypothetical preferences × convenience sample; top-journal publication × remaining scrutiny gap; methods/tool paper × clear downstream decision use; and catastrophic-risk relevance × concrete decision pathway. Use hierarchical shrinkage so one or two papers do not create a “rule.” 5. **Flexible check.** Fit a shallow boosted-tree or generalized additive model as a diagnostic. Compare its out-of-sample performance with the simpler model, then use partial-dependence and local explanations to look for missed nonlinearity. 6. **Rule distillation.** Translate only stable patterns into candidate heuristics. A rule should survive bootstrap resampling, an out-of-time test, and removal of any single cause area. Phrase it descriptively: “In the 2023–26 sample, papers with X tended to receive Y-point lower ratings, conditional on Z.” It should not become an automatic prompt instruction until a team member reviews the examples and plausible alternative explanations. The model should estimate associations, not claim causal effects of paper characteristics. For example, a negative coefficient on “hypothetical preferences” may partly reflect the topics for which those measures are used. Matching or weighting on adjacent features can reduce obvious imbalance but will not solve unobserved confounding. ## Validation Use paper-level, time-aware splits. Near-duplicate versions, related working papers, and the same research project must stay in the same fold. A random row split would leak project and period-specific information. Report: - mean absolute error and rank correlation for ratings; - calibration plots and Brier score for workflow transitions; - precision and recall among the top 10, 20, and 50 candidates, reflecting actual evaluation capacity; - uncertainty intervals from paper-level bootstrap resampling; - error and calibration by broad cause, income setting, evidence type, and method where sample sizes allow; - leave-one-cause-area-out performance; - performance using paper-only features versus paper-plus-comment tags. The paper-only model is the test of whether the heuristic can help score new research. The comment-augmented model is mainly an explanation of historical reasoning. If comments were written after a decision, they must not enter a predictive validation set. The difference between the two models can show how much the recorded reasons explain, without pretending that post-decision explanations were available prospectively. Compare the proposed model with three baselines: the overall historical mean, a simple cause/publication-status model, and the current AI scorer applied blind to the historical rating and status. Do not tune the prompt on the test set. Reserve the newest sufficiently mature cohort as a final test and report all specification changes made after looking at earlier results. ## Main limits The historical records are a selected set of research that someone already surfaced. They do not represent all potentially evaluable work. Ratings are missing for 87 of the 325 rows, and missingness is unlikely to be random. Cause categories and organizational strategy changed over time. Several topical cells are small or blank. Workflow outcomes also mix judgments with constraints. Author refusal, open-access status, evaluator supply, funding, and management capacity affect whether an evaluation happens. Publication of an evaluation is delayed and right-censored. A paper evaluated years ago had more time to complete than a recent candidate. Raters may differ in scale use and may have seen prior votes, discussions, or AI output. Suggester and assessor ratings are therefore not independent replications. Comments may be written or edited after outcomes are known. Downstream evaluator scores apply only to selected papers and cannot identify how rejected papers would have been evaluated. Most fundamentally, several different objectives are bundled in the ratings. A model could predict historical decisions well while learning an outdated workflow constraint or a preference that the current team would reject. This is why the result should be presented as a description, with counterexamples and sensitivity checks, and reviewed before it changes scoring. ## Implementable phased pipeline ### Phase 0: freeze a private analysis snapshot Create a gitignored analysis directory containing source hashes, snapshot dates, schemas, and record counts. Deduplicate by DOI and canonical URL first, then use title matching only for unresolved cases. Produce a match-review queue; never accept loose partial-title matches automatically. Define the status-to-transition map with the team before fitting models. Record which fields were available at the date of each decision. Keep row-level ratings, status histories, comments, and identifiers private. ### Phase 1: descriptive audit Build a paper-level table with separate suggester/assessor ratings, current workflow status, decision dates, basic paper features, and privacy-safe IDs. Publish only aggregate tables with minimum cell sizes. This phase should answer whether the user's proposed patterns are even visible before adjusting for other factors. Deliverables: - private data-quality and linkage report; - public aggregate coverage table; - distributions and transition counts; - a list of missing fields that prevent a welfare-pathway calculation. ### Phase 2: structured feature and comment coding pilot Code 50 stratified papers across score ranges, workflow outcomes, cause areas, and time periods. Include deliberate pairs, such as hypothetical-preference versus observed-choice studies addressing similar questions. Double-code a subset and revise the codebook before scaling. Deliverables: - private row-level coded dataset; - public codebook; - agreement results and aggregate tag counts; - adjudicated examples cleared for public use, if any. ### Phase 3: fit and validate descriptive models Fit the rating and sequential workflow models with paper-only features. Add comment tags in a separate explanatory specification. Run time-based, cause-held-out, and bootstrap validation. Compare against the current AI system on exactly the same historical paper packets. Deliverables: - model card with targets, exclusions, coefficients, uncertainty, and validation; - a small set of stable candidate heuristics with counterexamples; - an error audit showing where the AI and historical judgments disagree; - no automatic production changes. ### Phase 4: connect the descriptive model to the welfare framework For a small, diverse paper set, create explicit welfare-pathway records using the accompanying multiplicative model: decision scale, research-attributable influence, incremental outcomes, costs, affected groups, and named moral scenarios. Compare these calculations with historical ratings and the descriptive model. This phase asks useful questions without forcing agreement: Did the team historically give more weight to tractability than the explicit welfare model does? Are rich-country papers favored because of clearer decision pathways? Are animal and catastrophic-risk papers penalized because effects are harder to quantify? Which differences reflect evidence and which reflect moral assumptions? Deliverables: - six to twelve auditable welfare-pathway case studies; - sensitivity analysis for income curvature and cross-cause weights; - a decomposition of differences into empirical assumptions, evaluation-value considerations, and moral assumptions. ### Phase 5: prospective test before prompt integration Run the candidate heuristics and welfare-pathway template alongside the existing process for a new cohort. Collect blind human ratings before revealing either AI result. Track rank changes, decision time, reasons, and eventual workflow progression. Only after this test should reviewed rules enter the AI prompt or post-processing. Keep the historical model, explicit welfare scenarios, and final team decision visible as separate outputs. Do not let public feedback, unreviewed comment coding, or a small cause-area coefficient automatically steer production ratings. ## Recommended immediate next step Begin with Phase 0 and the 50-paper coding pilot. The dataset is large enough to learn whether the proposed patterns recur, but not large enough to skip careful coding, pooling, and validation. In parallel, use six to twelve papers for the explicit welfare-pathway exercise. That pairing should tell us whether the main problem is missing empirical inputs, poorly calibrated scoring heuristics, unresolved moral assumptions, or some combination of the three.