Opus 5.5 re-evaluation and score adjustment

On 1 October 2026 we re-evaluated 26 papers with a second model, Claude Opus 5.5, and compared it with the GPT-5.5 scores that the dashboard normally uses. Opus scores about 15 points lower. This note explains how we handled that, and how much weight the result can bear.

Experimental. The adjusted scores come from a small calibration sample (n=44 papers), are not validated against human judgment, and the two models disagree about the order of papers within the top group. Treat small differences as noise. You can switch the Opus scores off (see the end of this page).

What was done

Calibration sample

We planned five papers in each GPT-5.5 score band (70-79, 60-69, 50-59, 40-49 and below 40), chosen at random within the band from papers outside the top group. Opus reached its usage limit before the run finished, so we have 19 of the planned 25 papers: 5, 5, 5 and 4 in the four upper bands, and none below 40. Pooling these 19 with the 25 top papers gives 44 paired scores. (METR is not in the calibration: its earlier GPT-5.5 result was an out-of-scope judgment, a different question from score level.)

Result: a roughly constant gap of 15 points

Group (by GPT-5.5 score)PapersMean GPT-5.5Mean OpusMean differenceMedian difference
Top group (frozen top 25)2577.965.1-12.8-14
70-79575.053.8-21.2-20
60-69566.446.6-19.8-19
50-59554.039.0-15.0-14
40-49446.228.2-18.0-17
Below 40Not scored (Opus hit its usage limit)

Opus is lower in every group, by 13 to 21 points, and the gap does not shrink at lower scores. That pattern is what a scale difference looks like. If the top-group drop were mainly regression to the mean (extreme scores moving back towards the middle), the gap would be near zero further down; it is not.

The fit. Regressing Opus on GPT-5.5 across all 44 papers gives Opus = -22.8 + 1.11 x GPT (95% bootstrap interval for the slope 0.99 to 1.24; for the intercept -31.7 to -15.3; residual standard deviation 6.6 points; R-squared 0.79). A constant shift of about 15 points (mean difference -15.3) fits almost as well (residual standard deviation 6.7), and the slope is consistent with 1. We therefore use the simplest adjustment: adjusted Opus score = Opus score + 15, clipped to 0-100. Any more elaborate map would be false precision with this sample.

Effect on labels. Of the 34 papers with a GPT-5.5 score of 65 or more, 23 fall below 65 under Opus, and 30 of the 44 papers get a different recommended action (mostly shortlist to watchlist, and watchlist to lower priority). The adjusted score is on the GPT-5.5 scale, so the dashboard derives the shortlist/watchlist label for these papers from the adjusted score using the usual thresholds (75 and above, 50 to 74, below 50) and shows Opus's own recommendation separately in each paper's details.

How it enters the ranking. A re-evaluated paper’s priority is computed exactly as for every other paper (same sub-scores, impact component, weighting, human ratings and lens) and then shifted by the difference between its adjusted Opus score and its GPT-5.5 overall score, capped to 0-100. This keeps the comparison like-for-like: the Opus result moves a paper up or down relative to its GPT-5.5 position instead of replacing the usual composite.

Papers whose earlier score was not from GPT-5.5. The calibration compared Opus 5.5 with GPT-5.5 only, so the +15 is applied only where the earlier score is a GPT-5.5 score. Four of the 26 papers instead carry an earlier score from an older Opus run; for those the shift is the Opus 5.5 score minus that earlier score, with no offset, which assumes the two Opus versions score on the same scale. A paper whose earlier score came from any other model would get no shift at all, though its Opus result would still be shown. Each paper’s details say which case applies.

Caveats

Turning it off

On the main dashboard, untick “Use Opus re-evaluations (experimental)” next to the other ranking controls, or add ?opus=0 to the address. With it off, every paper is ranked and labelled exactly as before using GPT-5.5 scores. Each re-evaluated paper’s details still show the Opus result for reference.

Papers in the calibration sample

Scores are on the 0-100 scale. “Opus” is the raw Opus score, before the +15 adjustment. The 25 top papers are not listed here; each carries its own Opus result in its details on the dashboard.

PaperGPT-5.5 bandGPT-5.5OpusDifference
Did the Affordable Care Act Save Lives?70-797656-20
Health Insurance Underwriting and the Heterogeneous Effects of the Affordable Care Act70-797652-24
The Impact of Fiscal Policy to Promote Healthy Diets: Evidence from the Navajo Nation70-797650-26
Is a Dollar a Dollar? How Transfer Design Shapes Household Spending70-797455-19
The Anatomy and Evolution of Survey Error70-797356-17
Applied and Filtered: An End-to-End Algorithmic Fairness Audit of A Public Employment Agency60-696847-21
Did COVID Change the Black Neighborhood Startup Deficit? Evidence from the Startup Cartography Project60-696850-18
Labor Mobility and the Level of Unemployment in a Currency Union60-696850-18
Mapping General-Purpose AI Governance in Twenty AI Middle-Power Jurisdictions60-696643-23
Endogenous Selection and Spillovers: Bayesian Inference for Policy-Relevant Causal Effects60-696243-19
AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality50-595847-11
What does it take to catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring50-595636-20
Reshaping China’s Labour Market: AI’s Dual Impacts on Employee Adaptation and Employer Demand50-595236-16
Toward a Global Regime for Compute Governance: Building the Pause Button50-595238-14
Vegan versus meat-based pet foods: Owner-reported palatability behaviours and implications for canine and feline welfare50-595238-14
From Waste to Wealth: Exploring the Role of Biogas in Circular Economy Transitions Toward Zero-Waste Systems40-494932-17
Hydrogen mobility ecosystem acceleration using system dynamics modeling40-494730-17
Stress-testing university AI governance: A prospective method for locating policy breakpoints40-494724-23
AI and the Economy: An Economic Examination of Production, Distribution, Firms, Labor, and Welfare40-494227-15