The Unjournal research prioritization prototype
Back to the paper dashboard
Methods note and live pilot

How stable are the AI priority scores?

A useful score should not swing sharply because the same model was run again, a harmless instruction was rephrased, or a comparable model was used. We are adding explicit tests for those forms of fragility.

Stability is not accuracy. A model can repeat the same mistaken judgment every time. Prompt stability is therefore an early reproducibility check; comparison with existing human judgments remains a separate, descriptive test.

The idea in three checks

The approach adapts inter-rater reliability: repeated runs, prompt variants, or models play the role of different raters scoring the same frozen set of papers.

Check 1

Measure ordinary rerun variation

Run each wording several times with the paper, rubric, model, settings, and output schema fixed. Differences within a wording estimate ordinary run-to-run variation. This pilot used GPT-5.5 through Codex headless at medium reasoning effort for every run.

Check 2

Separate wording sensitivity

First average the four baseline runs and the four cautious-wording runs. Then compare those two averages. This reduces the chance that a random high or low execution is mistaken for a wording effect.

See the two exact wordings

Baseline: “Be calibrated, conservative, and honest about scope fit.”

Cautious plain language: “Keep the assessment well calibrated. Avoid optimistic score inflation and describe the paper's fit with Unjournal's scope candidly.”

Separate validation

Compare with people and models

Compare each model's stability profile and then compare its paper-level ratings with prior Unjournal prioritization ratings. Here “calibration” means correspondence with that human reference, not whether either side has identified the true research value.

Variation within each wording
Difference between wording means
Validation against human judgments

Live pilot status

The page emphasizes aggregate diagnostics. It identifies the public papers and the largest disagreement, but does not publish the full set of paper-level rationales.

What is being tested: every stability number below describes AI model scores.

“Previously human-rated papers” names a paper cohort selected because earlier Unjournal prioritization ratings are available. Those paper-level human ratings were not shown to the model and do not enter the stability calculations. AI and human scores are compared separately in the AI–human reference panel.

Balanced crossed design: two wordings × four executions

Each paper receives eight AI scores. Only the final closing instruction changes; all other scoring inputs remain fixed. We initially planned three executions, then added a fourth to every cell when the pre-specified convergence check was not fully met.

Closing instructionExecution 1Execution 2Execution 3Execution 4
Baseline wordingB1B2B3B4
Cautious wordingC1C2C3C4
  • Ordinary rerun variation: compare all six execution pairs within B1–B4 and within C1–C4.
  • Wording sensitivity: compare mean(B1…B4) with mean(C1…C4). This balanced comparison is our adaptation; the paper's framework treats prompt variants as rater-equivalents and can use multiple repetitions.
  • Convergence check: recompute the same outcomes after two, three, and four executions per wording. This shows which findings settle as runs are added and which remain boundary-sensitive.

Four repeats give a better estimate of ordinary execution variation, but they do not estimate the full range of possible prompts or future backend changes. Treating the two instructions as substantively equivalent is a reviewed design judgment.

LoadingReading pilot data...
Which papers are in the pilot?

Reading the frozen pilot sample…

Prior recorded scores are prioritization and expected-impact ratings, not research-quality grades.

Preliminary report

Expanded pilot, 1 September 2026

Reading the expanded pilot results…

What this does and does not show

The crossed design separates ordinary execution variation from one reviewed wording contrast. It does not estimate sensitivity to every reasonable prompt, model, or future backend version.

Twelve selected papers per cohort are still too few for precise population claims. The two cohorts also differ in their source text and score distribution. Their contrast is descriptive; prior human rating did not cause the observed difference in AI stability.

The paper-level human ratings were withheld from the prompts. Some papers and evaluations are public, so prior model exposure cannot be ruled out. Human-score comparisons are descriptive and should not be read as blinded validation.

Earlier 12-paper pilot (archived)

The first round used six papers per cohort, two executions per wording, and a separate third baseline check. It found the same broad pattern: ordinary rerun variation was larger than the tested wording effect, with more instability in the historical comparison cohort. The expanded results above supersede its numerical estimates.

What the metrics measure

Agreement on numeric scores

Interval Krippendorff alpha for the overall 0–100 rating and each 0–10 component. It compares squared score gaps for ratings of the same paper with squared gaps between arbitrary ratings from the whole cohort. The paper treats about 0.8 as a useful, context-dependent reference point, not a universal cutoff.

Operationally important movement

The mean, median, and maximum within-paper score range, plus the number of papers that cross 25, 50, 65, or 75 across runs. A modest wobble can matter if it changes a workflow decision.

Ranking and action stability

Mean pairwise rank correlation and nominal agreement on recommended action. A high alpha can coexist with a few important shortlist reversals, so both views matter.

How is Krippendorff alpha calculated here?

Observed disagreement is built from pairs of ratings assigned to the same paper. For numeric scores, each pair contributes its squared point difference. Expected disagreement pools every rating in the selected cohort and calculates the same quantity for arbitrary pairs, ignoring which paper each score belongs to.

alpha = 1 − observed disagreement / expected disagreement. The practical question is: are two ratings of the same paper substantially closer than two ratings drawn from the cohort at random?

For example, suppose Paper A receives 60 and 70, while Paper B receives 70 and 80. The within-paper squared gap is 100. Across all six possible pairs from the pooled scores, the average squared gap is 133.3. Alpha is therefore 1 − 100 / 133.3 = 0.25. This does not mean 25% correct; it means same-paper disagreement is 25% smaller than the pooled benchmark.

A wider spread between papers increases expected disagreement and can raise alpha without reducing rerun movement. This is why the page also reports score movement, boundary changes, and typical separation between paper means.

What is normal? There is no task-independent normal value. The PSS paper's experiments found mostly high fixed-prompt stability, but weaker results for some low-signal tasks; several prompt-variation tests fell below 0.8 even under relatively mild paraphrasing. Those experiments mostly classify categorical outcomes, use up to 500 items and 30 runs, and include one 10-point outcome. Our 24-paper, 0–100 prioritization task is too different for its alpha values to serve as a direct benchmark. The useful comparison is the pattern: whether the estimate settles as runs are added and whether mild wording changes are less consequential than ordinary rerun noise.

Why not display the model's own confidence as the answer?

A self-reported confidence number is another model output. It may be useful, but it does not show how the pipeline actually behaves under reruns or prompt changes. Observed stability measures that behavior directly.

Metric glossary: what the statistics mean
Prompt Stability Score (PSS)
A family of observed-agreement tests. It asks whether AI ratings change when we rerun the same prompt or make a meaning-preserving wording change. It does not test whether the ratings are correct.
Intra-prompt PSS
Agreement among repeated AI runs with wording held fixed. Here it is estimated separately for the baseline and cautious wordings, rather than assuming one wording represents all run variation.
Inter-prompt PSS
Agreement between reviewed, equivalent wordings. Here each wording score is the mean of four executions, so the comparison is less affected by ordinary run-to-run variation. The long scoring rubric stays fixed.
Crossed design
Every paper is scored under every wording × execution combination. This lets us compare the observed wording shift with movement that occurs even when wording is unchanged.
Mean absolute wording effect
For each paper, take the absolute difference between its cautious-wording mean and baseline-wording mean, then average across papers. It measures the size, not the direction, of this wording contrast.
Mean signed wording effect
The average cautious-minus-baseline difference. A negative value means the cautious wording lowered scores on average; an interval spanning zero does not show a clear directional shift.
Krippendorff alpha
For numeric scores, observed disagreement is the average squared gap among ratings of the same paper. Expected disagreement is the corresponding average squared gap among ratings pooled across the cohort, ignoring paper identity. Alpha is 1 minus their ratio. Thus 0.50 means same-paper disagreement is half the pooled benchmark; 0.25 means it is three-quarters; 0 means it is equal; and a negative value means it is larger. These are not percentages correct. The paper uses about 0.8 as a context-dependent diagnostic reference, not a universal pass/fail rule.
Typical between-paper separation
First average all eight AI scores for each paper. Then take the absolute difference between every pair of paper means and average those differences. This provides a like-for-like point-scale benchmark for ordinary rerun movement within the selected cohort.
95% paper-bootstrap interval
An uncertainty interval made by repeatedly resampling whole papers and recomputing alpha. Keeping each paper's reruns together respects the repeated-rating structure. Very small cohorts produce wide, sometimes negative intervals.
Within-paper range
The highest AI priority score minus the lowest for the same paper under the selected test. We report the mean and median across papers, plus the maximum fragile case.
Rank correlation
Mean pairwise Spearman correlation between AI run rankings. 1 preserves the paper order exactly; 0 indicates no monotonic relationship; negative values reverse the order.
Recommended-action alpha
Chance-adjusted agreement on categorical AI recommendations, such as shortlist or do not prioritize. This uses nominal rather than numeric alpha because the action labels are categories.
Same broad action tier
The share that stayed in the same broad action tier (below 25, 25–49, 50–74, or 75+). Ordinary-rerun summaries use paper–wording pairs; wording-sensitivity summaries use papers. It is an operational summary, not a statistical reliability coefficient.
Boundary flip
The count and share of papers with AI scores on both sides of a dashboard boundary (25, 50, 65, or 75) across reruns. Even modest score movement can matter near a workflow threshold.

How this is adapted to Unjournal prioritization

  1. Freeze the inputs. Use the same public title, abstract, authors, source, publication status, field hint, and cause-area hint for every condition.
  2. Use two deliberate cohorts. One contains 12 newer AI human-impact/governance papers. The comparison cohort contains 12 previously human-prioritized papers.
  3. Cross wording with execution. Run each of two reviewed closing instructions four times. The long prioritization rubric, aggregate human-calibration guidance, and JSON schema remain unchanged.
  4. Use one model first. Establish whether the current production setup is internally repeatable before spending effort comparing providers or model families.
  5. Do not mix fallbacks. A failed run stays failed; it is not silently replaced with a different model. Otherwise a supposed prompt test would partly become a model test.
  6. Publish compact aggregates, with selected cases when useful. The public page shows reliability and boundary diagnostics, identifies the public paper list, and explains the largest disagreement. The full set of rationales stays out of the main report.

This first pilot is intentionally narrower than paraphrasing the entire prompt. The stable core contains policy choices and calibration anchors; changing those would test a different construct, not harmless wording.

Exact closing instructions used in this pilot

These are the two closing instructions used in the balanced crossed test. Each was executed four times per paper. The paper text, long Unjournal scoring rubric, calibration guidance, component definitions, and JSON output schema stayed fixed.

Baseline (baseline) Be calibrated, conservative, and honest about scope fit.
Cautious plain language (cautious_plain) Keep the assessment well calibrated. Avoid optimistic score inflation and describe the paper's fit with Unjournal's scope candidly.

How results should change the project

FindingLikely diagnosisPractical response
Low repeatability within a wordingRun-level nondeterminism, ambiguous papers, underspecified criteria, or inadequate model capability.Inspect the largest same-wording ranges, tighten edge-case rules, and add repetitions before drawing prompt conclusions.
Runs disagree about existing public scrutinyThe scorer is reconstructing a consequential factual input differently on each run.Supply a deterministic scrutiny assessment or verified evidence bundle, and require the scorer to mark missing evidence as unclear rather than infer it afresh.
Stable within wordings but different wording meansThe closing-instruction contrast is changing the judgment beyond ordinary rerun movement.Review the wording-sensitive cases and revise or standardize the fragile instruction before relying on sharp thresholds.
High alpha but many 65-point flipsOverall agreement looks good, but operational decisions near a boundary remain fragile.Show a stability flag near the threshold, request a human second look, or avoid treating 65 as a mechanically sharp boundary.
Stable but poorly matched to humansThe system is reproducibly wrong or systematically miscalibrated.Change calibration anchors or rubric interpretation; prompt stability alone offers no reassurance.
Models disagreeCould reflect capability, training, or interpretation differences.Compare each model's own stability and its error against blinded human anchors; do not choose from agreement alone.

Sources and implementation

Implementation note: the dashboard adaptation uses paper-level bootstrap resampling so the repeated ratings for one paper stay together. This is a deliberate variation from the package implementation described in the paper's appendix, which resamples long-format annotation records.