Internal scoping aid for prioritization, drafted with AI assistance; not an evaluation and not an endorsement. Source: docs/scope/metr-time-horizons.md in the project repository.
Status: internal scoping aid for prioritization, prepared 2026-10-01. Not an evaluation and not an endorsement. Contains no non-public information.
Paper: Kwa, West, Becker, et al. (METR), arXiv 2503.14499 (2025; later versions through v4, July 2026; listed as NeurIPS 2025). Facts below on the abstract were checked against the arXiv page; items marked "verify" come from background knowledge and should be checked against the paper before being relied on.
Why this memo exists
The dashboard's earlier GPT-5.5 pass scored the paper 23 and labelled it out_of_scope (since superseded: the Opus 5.5 thorough re-evaluation of 2026-10-01, run with this memo as context, gave 58 / monitor), even though its components were high (decision relevance 7.1, prominence 7.5, real-world influence 7.0). The paper also underpins crux UJG05 (capability metrics as policy triggers). The label likely reflects a "technical ML benchmark" reading of scope. This memo sets out the competing readings so a person can decide.
What the paper does
- Proposes the 50% time horizon: the length of task, measured by how long skilled humans take, at which an AI system succeeds about half the time.
- Builds a task pool from three sources: HCAST (agentic software and general-skills tasks), RE-Bench (ML research-engineering tasks), and 66 novel shorter tasks (the abstract says "66 novel shorter tasks"; the paper's name for this suite is SWAA, verify).
- Human baselining: domain-skilled people were timed on the tasks; task length is the human completion time.
- Statistical model: for each model, success is regressed on log human time with a logistic fit; the 50% crossing point is the horizon.
- Headline results: Claude 3.7 Sonnet had a horizon of roughly 50 minutes; horizons across frontier models have doubled about every seven months since 2019, possibly faster in 2024.
- Extrapolation: if the trend and its external validity hold, within about five years AI systems could automate many software tasks that take humans a month.
- Stated drivers: reliability, adapting to mistakes, reasoning and tool use. The authors discuss external validity and dangerous-capability implications as limitations.
How it is used
- Widely cited as a forecasting anchor for AI capability timelines and for "automation of AI R&D" scenarios (it feeds directly into cruxes about intelligence-explosion speed and economic scenarios in this dashboard).
- Used by AI-safety institutes, labs' frontier-safety frameworks and think tanks as a shorthand for capability growth, and sometimes proposed as a trigger metric for policy thresholds.
- Frequently over-read in media. MIT Technology Review (5 Feb 2026, "This is the most misunderstood graph in AI") argues that a one-hour horizon does not mean a model can replace an hour of real work, and that the tasks are almost all software and benchmark-style.
Scope question, argued both ways
Out of scope (technical ML capability measurement). The object studied is an AI benchmark; the method is ML evaluation; there is no human-subjects welfare outcome, no economic outcome, and no policy intervention. Unjournal's usual remit is quantitative social science and decision-relevant empirical work. Technical AI/CS papers are excluded in the repository's scope statement. Evaluators from ML might also review it through venues such as NeurIPS.
In scope (quantitative measurement and forecasting used directly in welfare-relevant decisions). The paper is less a model-building paper than a measurement-and-forecasting paper whose central output (a doubling time and a five-year extrapolation) is used in decisions by funders, governments and labs. Unjournal explicitly covers forecasting methodology, measurement validity and catastrophic-risk-relevant evidence. The questions that matter here are social-science questions: does the measure mean what users take it to mean, are the human baselines a valid yardstick, is the statistical inference honest, and is it fit to serve as a policy trigger. Few other papers combine such high influence with such open measurement questions.
Which parts fit. In: construct validity and the mapping to real-world work, the baseline sample and incentives, the statistical model and uncertainty, extrapolation, and fitness as a policy trigger and forecasting input. Out (or only for specialist input): detailed task-suite engineering, agent scaffolding choices, and model-by-model benchmark mechanics.
What evaluators must understand and scrutinize
- Construct validity. Does "human time" index difficulty in a way that is comparable across tasks and between humans and models? Tasks are self-contained and algorithmically scorable, which differs from messy real work.
- Human baselines. Who was timed, how selected, how paid, how much context they had (baseliners were skilled but often unfamiliar with the specific codebase; verify), how failures and long tasks were handled, and whether incentives differ from workplace settings. Later evidence on context-rich experts being much faster would shift horizons.
- Statistical model. Logistic fit on log time, the handling of task clustering and few long tasks, confidence intervals on the horizon, the sensitivity of the 80% horizon (substantially shorter than the 50% horizon), and the weight on a small number of models in the fit.
- Extrapolation risk. An exponential fit over about six years, sparse data at long horizons, possible regime changes (reasoning models, post-2024 acceleration), and the leap from "doubling every seven months" to "a month-long task in five years".
- Contamination and selection. Training-data overlap, task selection that favours automatable tasks, and benchmark saturation.
- Mapping to economic capability. Software task length is not occupational substitution; check what evidence links horizon to labour-market or productivity outcomes (including field-experiment evidence on developer productivity, verify).
- Updates and critiques. Subsequent versions (to v4) and METR's follow-up time-horizon work, independent replications, and critiques such as the MIT Technology Review piece, the Gary Marcus commentary and others. An evaluator should read the current version, not v1.
- Fitness as a policy trigger. Gaming and Goodhart risk, measurement noise at thresholds, task-suite maintenance, and whether a single scalar suits a regulatory trigger.
Evaluator mix (2-3 people, from these profiles)
- ML evaluation and benchmarking: agentic-benchmark design, contamination, scaffolding sensitivity.
- Psychometrics or item-response measurement: construct validity, difficulty scales, human-vs-machine comparability.
- Statistics and forecasting: logistic fits with clustered items, uncertainty, trend extrapolation, forecasting track records.
- Economics of AI and task-based labour: mapping capability to work and productivity.
- AI policy / governance: use as a threshold or trigger.
Suggested pairing: one measurement/statistics person (psychometrics or forecasting), one ML-evaluation person, and, if a third is available, an economics-of-AI or AI-policy person. Avoid evaluators who work at METR or on closely competing time-horizon benchmarks without declaring conflicts.
Scope recommendation
Treat as in scope as a nonstandard or applied evaluation, not a conventional academic paper evaluation: restrict the brief to measurement validity, statistical inference, extrapolation and decision use, and explicitly exclude detailed engineering of the task suite. Because the paper has gone through several versions and a NeurIPS publication, check whether a stable, citable version exists and what an Unjournal evaluation would add beyond that review. The original GPT-5.5 score of 23 / out_of_scope looked too low for a paper with this influence; on the narrower brief, a middling-to-high score is defensible, tempered by the existing scrutiny (substantial public critique already exists, which reduces neglectedness).
Suggested evaluation questions
- Does the 50% time horizon measure a meaningful, comparable quantity across models and tasks, and what does it not measure?
- How sensitive are the horizon and the doubling time to baseline selection, baseliner context, task mix and statistical specification? Are reported intervals adequate?
- How reliable is the extrapolation to month-long tasks, and what would falsify it?
- What evidence links horizon to economically relevant work, and how large is the gap?
- Is the metric fit to act as a policy or safety trigger, and under what safeguards?
- How should decision-makers (funders, regulators, forecasters) use the headline number differently after the later updates and critiques?