Agreement on numeric scores
Interval Krippendorff alpha for the overall 0–100 rating and each 0–10 component. It compares squared score gaps for ratings of the same paper with squared gaps between arbitrary ratings from the whole cohort. The paper treats about 0.8 as a useful, context-dependent reference point, not a universal cutoff.
Operationally important movement
The mean, median, and maximum within-paper score range, plus the number of papers that cross 25, 50, 65, or 75 across runs. A modest wobble can matter if it changes a workflow decision.
Ranking and action stability
Mean pairwise rank correlation and nominal agreement on recommended action. A high alpha can coexist with a few important shortlist reversals, so both views matter.
How is Krippendorff alpha calculated here?
Observed disagreement is built from pairs of ratings assigned to the same paper. For numeric scores, each pair contributes its squared point difference. Expected disagreement pools every rating in the selected cohort and calculates the same quantity for arbitrary pairs, ignoring which paper each score belongs to.
alpha = 1 − observed disagreement / expected disagreement. The practical question is: are two ratings of the same paper substantially closer than two ratings drawn from the cohort at random?
For example, suppose Paper A receives 60 and 70, while Paper B receives 70 and 80. The within-paper squared gap is 100. Across all six possible pairs from the pooled scores, the average squared gap is 133.3. Alpha is therefore 1 − 100 / 133.3 = 0.25. This does not mean 25% correct; it means same-paper disagreement is 25% smaller than the pooled benchmark.
A wider spread between papers increases expected disagreement and can raise alpha without reducing rerun movement. This is why the page also reports score movement, boundary changes, and typical separation between paper means.
What is normal? There is no task-independent normal value. The PSS paper's experiments found mostly high fixed-prompt stability, but weaker results for some low-signal tasks; several prompt-variation tests fell below 0.8 even under relatively mild paraphrasing. Those experiments mostly classify categorical outcomes, use up to 500 items and 30 runs, and include one 10-point outcome. Our 24-paper, 0–100 prioritization task is too different for its alpha values to serve as a direct benchmark. The useful comparison is the pattern: whether the estimate settles as runs are added and whether mild wording changes are less consequential than ordinary rerun noise.