Opus 5.5 re-evaluation and score adjustment
On 1 October 2026 we re-evaluated 26 papers with a second model, Claude Opus 5.5, and compared it with the GPT-5.5 scores that the dashboard normally uses. Opus scores about 15 points lower. This note explains how we handled that, and how much weight the result can bear.
Experimental. The adjusted scores come from a small calibration sample (n=44 papers), are not validated against human judgment, and the two models disagree about the order of papers within the top group. Treat small differences as noise. You can switch the Opus scores off (see the end of this page).
What was done
- Opus 5.5 re-scored 25 papers from the top of the dashboard (the top 15 AI-related and top 10 other papers by default priority, frozen on 30 September 2026) and, separately, the METR time-horizons paper, whose scope we were discussing. It used the same thorough scoring prompt and the same public-scrutiny search as the GPT-5.5 refinement step.
- Each paper keeps its GPT-5.5 score, action and rationales. The Opus result is stored alongside them, not in place of them.
- To see whether Opus is simply harsher, or whether the gap depends on score level, Opus also re-scored a stratified random sample of lower-ranked papers (listed below).
Calibration sample
We planned five papers in each GPT-5.5 score band (70-79, 60-69, 50-59, 40-49 and below 40), chosen at random within the band from papers outside the top group. Opus reached its usage limit before the run finished, so we have 19 of the planned 25 papers: 5, 5, 5 and 4 in the four upper bands, and none below 40. Pooling these 19 with the 25 top papers gives 44 paired scores. (METR is not in the calibration: its earlier GPT-5.5 result was an out-of-scope judgment, a different question from score level.)
Result: a roughly constant gap of 15 points
| Group (by GPT-5.5 score) | Papers | Mean GPT-5.5 | Mean Opus | Mean difference | Median difference |
|---|---|---|---|---|---|
| Top group (frozen top 25) | 25 | 77.9 | 65.1 | -12.8 | -14 |
| 70-79 | 5 | 75.0 | 53.8 | -21.2 | -20 |
| 60-69 | 5 | 66.4 | 46.6 | -19.8 | -19 |
| 50-59 | 5 | 54.0 | 39.0 | -15.0 | -14 |
| 40-49 | 4 | 46.2 | 28.2 | -18.0 | -17 |
| Below 40 | Not scored (Opus hit its usage limit) | ||||
Opus is lower in every group, by 13 to 21 points, and the gap does not shrink at lower scores. That pattern is what a scale difference looks like. If the top-group drop were mainly regression to the mean (extreme scores moving back towards the middle), the gap would be near zero further down; it is not.
The fit. Regressing Opus on GPT-5.5 across all 44 papers gives Opus = -22.8 + 1.11 x GPT (95% bootstrap interval for the slope 0.99 to 1.24; for the intercept -31.7 to -15.3; residual standard deviation 6.6 points; R-squared 0.79). A constant shift of about 15 points (mean difference -15.3) fits almost as well (residual standard deviation 6.7), and the slope is consistent with 1. We therefore use the simplest adjustment: adjusted Opus score = Opus score + 15, clipped to 0-100. Any more elaborate map would be false precision with this sample.
Effect on labels. Of the 34 papers with a GPT-5.5 score of 65 or more, 23 fall below 65 under Opus, and 30 of the 44 papers get a different recommended action (mostly shortlist to watchlist, and watchlist to lower priority). The adjusted score is on the GPT-5.5 scale, so the dashboard derives the shortlist/watchlist label for these papers from the adjusted score using the usual thresholds (75 and above, 50 to 74, below 50) and shows Opus's own recommendation separately in each paper's details.
How it enters the ranking. A re-evaluated paper’s priority is computed exactly as for every other paper (same sub-scores, impact component, weighting, human ratings and lens) and then shifted by the difference between its adjusted Opus score and its GPT-5.5 overall score, capped to 0-100. This keeps the comparison like-for-like: the Opus result moves a paper up or down relative to its GPT-5.5 position instead of replacing the usual composite.
Papers whose earlier score was not from GPT-5.5. The calibration compared Opus 5.5 with GPT-5.5 only, so the +15 is applied only where the earlier score is a GPT-5.5 score. Four of the 26 papers instead carry an earlier score from an older Opus run; for those the shift is the Opus 5.5 score minus that earlier score, with no offset, which assumes the two Opus versions score on the same scale. A paper whose earlier score came from any other model would get no shift at all, though its Opus result would still be shown. Each paper’s details say which case applies.
Caveats
- Small sample. With 44 papers the offset is reliable to roughly 3 points and the slope to roughly 0.12. Nothing was tested below a GPT-5.5 score of 40, so the adjustment is an extrapolation there.
- Order within the top group. Ranks agree strongly across the whole calibration sample (Spearman correlation 0.95) but hardly at all within the top 25 (0.17, not distinguishable from zero). The two models do not reliably agree about which top paper is better, and a shift of 15 points cannot fix that. The top group was also narrower in range, which lowers any correlation inside it.
- The top group dropped a little less. Its mean gap was 12.8, so adding 15 puts the adjusted Opus scores about 2 points above the original GPT-5.5 scores on average. Papers with a GPT-5.5 score in the 70s dropped by about 21.
- Effort level. The 70-79 band consists entirely of GPT-5.5 papers that had a second, higher-effort pass. Those papers fell more than medium-effort ones (adjusting for score level, about 7 points more; interval 2 to 11). A possible reason is that GPT-5.5's refinement step inflated them, so some regression to the mean sits inside GPT's own tiers. We did not find a difference between AI-related and other areas (mean gaps of 15.4 and 15.2).
- Not human-validated. We did not compare either model with human ratings here. A scale correction between two models says nothing about which one is closer to a good prioritization.
- Measurement noise. We have not measured how much Opus scores vary on repeated runs. The stability pilot found ordinary rerun variation for GPT-5.5.
Turning it off
On the main dashboard, untick “Use Opus re-evaluations (experimental)” next to the other ranking controls, or add ?opus=0 to the address. With it off, every paper is ranked and labelled exactly as before using GPT-5.5 scores. Each re-evaluated paper’s details still show the Opus result for reference.
Papers in the calibration sample
Scores are on the 0-100 scale. “Opus” is the raw Opus score, before the +15 adjustment. The 25 top papers are not listed here; each carries its own Opus result in its details on the dashboard.
| Paper | GPT-5.5 band | GPT-5.5 | Opus | Difference |
|---|---|---|---|---|
| Did the Affordable Care Act Save Lives? | 70-79 | 76 | 56 | -20 |
| Health Insurance Underwriting and the Heterogeneous Effects of the Affordable Care Act | 70-79 | 76 | 52 | -24 |
| The Impact of Fiscal Policy to Promote Healthy Diets: Evidence from the Navajo Nation | 70-79 | 76 | 50 | -26 |
| Is a Dollar a Dollar? How Transfer Design Shapes Household Spending | 70-79 | 74 | 55 | -19 |
| The Anatomy and Evolution of Survey Error | 70-79 | 73 | 56 | -17 |
| Applied and Filtered: An End-to-End Algorithmic Fairness Audit of A Public Employment Agency | 60-69 | 68 | 47 | -21 |
| Did COVID Change the Black Neighborhood Startup Deficit? Evidence from the Startup Cartography Project | 60-69 | 68 | 50 | -18 |
| Labor Mobility and the Level of Unemployment in a Currency Union | 60-69 | 68 | 50 | -18 |
| Mapping General-Purpose AI Governance in Twenty AI Middle-Power Jurisdictions | 60-69 | 66 | 43 | -23 |
| Endogenous Selection and Spillovers: Bayesian Inference for Policy-Relevant Causal Effects | 60-69 | 62 | 43 | -19 |
| AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality | 50-59 | 58 | 47 | -11 |
| What does it take to catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring | 50-59 | 56 | 36 | -20 |
| Reshaping China’s Labour Market: AI’s Dual Impacts on Employee Adaptation and Employer Demand | 50-59 | 52 | 36 | -16 |
| Toward a Global Regime for Compute Governance: Building the Pause Button | 50-59 | 52 | 38 | -14 |
| Vegan versus meat-based pet foods: Owner-reported palatability behaviours and implications for canine and feline welfare | 50-59 | 52 | 38 | -14 |
| From Waste to Wealth: Exploring the Role of Biogas in Circular Economy Transitions Toward Zero-Waste Systems | 40-49 | 49 | 32 | -17 |
| Hydrogen mobility ecosystem acceleration using system dynamics modeling | 40-49 | 47 | 30 | -17 |
| Stress-testing university AI governance: A prospective method for locating policy breakpoints | 40-49 | 47 | 24 | -23 |
| AI and the Economy: An Economic Examination of Production, Distribution, Firms, Labor, and Welfare | 40-49 | 42 | 27 | -15 |