The Unjournal · Research prioritization
Paper-specific consideration page · current tranche rank 1

Measuring AI Ability to Complete Long Software Tasks

Heavy public attention and existing public critique
Why this page exists. We are considering whether an independent evaluation of this paper would be useful. The synthesis score uses provisional weights, and the human sample is small. It is one input to the decision.
25AI evaluation-priority (shifted by the Opus 5.5 re-evaluation)
25Original AI lens before the Opus shift (not used for the synthesis)
81Human aggregate · n=2 · effective weight 5.0
60Human–AI synthesis

Why this paper is being considered

The paper proposes the 50% time horizon (the human task length at which an AI agent succeeds half the time) and reports that horizons have doubled roughly every seven months since 2019. Its extrapolation is widely used in capability forecasts and policy discussion. The scope question is whether it belongs in an Unjournal evaluation at all: it is a technical benchmark paper, but its measurement, baselining, statistical and decision-use questions are social-science questions. See the scope discussion below.

What an expert evaluation could add: Substantive public criticism already targets human baselines, benchmark realism and horizon extrapolation. A further evaluation needs to resolve those disputes and their policy implications rather than repeat them. Evidence and commissioning case. AI-assisted judgment, 2026-10-07; separate from ratings and completed evaluations.

AI-generated criterion ratings and reasoning

Scores come from AI prioritization; the accompanying explanations include AI-assisted source checks. These are provisional judgments, separate from human ratings and commissioned evaluations.

Decision relevance

7.1/10

The research has a moderately strong global-welfare/VoI case because it may improve decisions about frontier AI monitoring, pacing, deployment safeguards, and preparedness by translating benchmark progress into a human-time task horizon. Its largest welfare pathway is indirect GCR/x-risk reduction through better capability measurement; current-human labor and software productivity implications matter too, but the paper does not quantify distributional welfare or LMIC effects. The main limitations are external validity, uncertainty about extrapolating the doubling trend, and whether the findings actually change policy thresholds rather than feeding general AI-progress discourse.

  • Paper claim to check Frontier AI agents have a roughly 50-minute 50%-task-completion time horizon on the studied software and research tasks. Source abstract

Value of added scrutiny

4.0/10

The paper is among the most visible in AI forecasting. It is published at NeurIPS 2025, METR has extended it in follow-up work, and MIT Technology Review and independent bloggers have critiqued how the headline chart is read. Substantial scrutiny already exists, which lowers the value of added scrutiny but does not settle the scope question.

Timing

4.5/10

The paper was first posted in March 2025 and revised in July 2026; arXiv now lists a NeurIPS 2025 reference, so it is no longer best treated as an unreviewed preprint. Some independent public commentary exists, including technical blog reviews and summaries, but the supplied scrutiny evidence was empty and I would not treat the scrutiny gap as fully closed. Timing value for UJ is therefore only moderate: feedback may still matter for policy interpretation, but the main academic review window has likely passed.

  • Source record Publication-stage and date evidence should be checked in the linked paper record. TARGETED_CURATED

Methodological potential

6.5/10

A useful evaluation would require technical expertise in AI benchmarking, statistics, software engineering task design, and AI governance. The hardest issues are external validity, benchmark selection, human-time baselining, extrapolation uncertainty, and whether a 50% completion threshold maps to policy-relevant autonomy or risk. This is close to the Unjournal's AI governance interests but outside its usual quantitative social-science core because it is mainly a capabilities benchmark paper.

  • Paper claim to check Frontier AI agents have a roughly 50-minute 50%-task-completion time horizon on the studied software and research tasks. Source abstract

Prominence

7.5/10

The scoring model estimated prominence from the paper's venue, authors, institutional setting, and visibility. The model did not supply a criterion-specific explanation. Current public-attention status: Heavy public attention and existing public critique.

  • Conference publication The paper appears in the NeurIPS 2025 proceedings, so it has passed one round of ML-venue review. NeurIPS

Likely influence

7.0/10

The scoring model estimated how far the findings could shape later research or decisions, without supplying a criterion-specific explanation. Current public-attention status: Heavy public attention and existing public critique. See public-attention evidence below.

  • Conference publication The paper appears in the NeurIPS 2025 proceedings, so it has passed one round of ML-venue review. NeurIPS

What the paper says

Source abstract · Abstract text from the arXiv API record; whitespace normalized.

Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear. To quantify the capabilities of AI systems in terms of human capabilities, we propose a new metric: 50%-task-completion time horizon. This is the time humans typically take to complete tasks that AI models can complete with 50% success rate. We first timed humans with relevant domain expertise on a combination of RE-Bench, HCAST, and 66 novel shorter tasks. On these tasks, current frontier AI models such as Claude 3.7 Sonnet have a 50% time horizon of around 50 minutes. Furthermore, frontier AI time horizon has been doubling approximately every seven months since 2019, though the trend may have accelerated in 2024. The increase in AI models' time horizons seems to be primarily driven by greater reliability and ability to adapt to mistakes, combined with better logical reasoning and tool use capabilities. We discuss the limitations of our results -- including their degree of external validity -- and the implications of increased autonomy for dangerous capabilities. If these results generalize to real-world software tasks, extrapolation of this trend predicts that within 5 years, AI systems will be capable of automating many software tasks that currently take humans a month.

Claims to check

  • Frontier AI agents have a roughly 50-minute 50%-task-completion time horizon on the studied software and research tasks.
  • The estimated frontier time horizon has doubled approximately every seven months since 2019, with possible acceleration in 2024.
  • If the trend generalizes, AI systems may automate many software tasks that currently take humans a month within about five years.
Methodological or theoretical issues flagged for evaluation

A useful evaluation would require technical expertise in AI benchmarking, statistics, software engineering task design, and AI governance. The hardest issues are external validity, benchmark selection, human-time baselining, extrapolation uncertainty, and whether a 50% completion threshold maps to policy-relevant autonomy or risk. This is close to the Unjournal's AI governance interests but outside its usual quantitative social-science core because it is mainly a capabilities benchmark paper.

Opus 5.5 re-evaluation (experimental)

Opus raw score: 58/100
Adjusted (+15): 73/100
Opus own action: watchlist
Label from adjusted score: watchlist
Earlier score (gpt-5.5): 23/100
Shift applied to the AI score: +50

Read with care. This is an experimental re-evaluation by a different model (Claude Opus 5.5, 2026-10-01). Across a calibration sample Opus scored about 15 points lower than GPT-5.5, so 15 points are added to compare it with GPT-5.5 scores. The calibration is a small sample (n=44) and is not human-validated, and the two models disagree about the order of papers within the top group. The AI score at the top of this page is the usual evaluation-priority score shifted by the difference between the adjusted Opus score and the earlier holistic score; the synthesis uses that shifted value, and the unshifted score is shown beside it. On the dashboard you can switch the Opus scores off. Read the methods note.

Opus rationale (AI-generated)

This METR paper introduced the '50% time horizon' and the finding that frontier-AI horizons double roughly every seven months. It has arguably become the most-cited single number in debates about AI capability timelines. It is used by lab frontier-safety frameworks, AI Safety/Security Institutes, forecasters, think tanks and funders such as Open Philanthropy / Coefficient Giving, and it underpins our own crux on capability metrics as policy triggers. On a narrow reading of our scope (technical ML benchmarking) it is out. But the questions that matter for its use are measurement and social-science ones: construct validity of the human-time yardstick, baseline selection and incentives, logistic inference with clustered items, the reliability of an exponential extrapolation from few models, and whether a single scalar suits a regulatory trigger. That makes a scoped, nonstandard evaluation defensible, and I would not label it out_of_scope. Two things temper the score. First, the paper is NeurIPS 2025 and has been revised through v4 (July 2026), so the original working paper has moved on. Second, it has already had substantial public critique and follow-up work, so the marginal value of another review is lower than for an unvetted paper of similar influence. I would put it in 'monitor / consider as an applied evaluation': worth commissioning if we can recruit a psychometrics or forecasting statistician plus an ML-evaluation specialist with no METR conflicts, and if the brief is restricted to measurement validity, inference, extrapolation and decision use. Our value-add would be a careful, citable assessment of how much weight policymakers should put on the headline doubling time, which current critiques (mostly journalism and blog posts) do not provide.

Dashboard details and provenance

Discovery source: TARGETED_CURATED
Publication status: Published other journal
Release date: 2025-03-18
Scoring model: gpt-5.5 (codex headless, medium)
Model holistic score: 23

Full AI dashboard scoring rationale

This is influential AI capability-measurement work from METR, with a plausible path to informing AI governance decisions by labs, regulators, insurers, and governments about autonomy thresholds and dangerous-capability evaluation. However, it is primarily a technical AI benchmarking paper, now apparently accepted at NeurIPS 2025, rather than quantitative social science or policy evaluation in the Unjournal's usual sense. An Unjournal evaluation could be useful if framed as AI-governance scrutiny of policy-relevant measurement claims, but ordinary UJ commissioning would probably add less value than review by technical evals and AI-safety experts.

AI decision-relevance rationale

The research has a moderately strong global-welfare/VoI case because it may improve decisions about frontier AI monitoring, pacing, deployment safeguards, and preparedness by translating benchmark progress into a human-time task horizon. Its largest welfare pathway is indirect GCR/x-risk reduction through better capability measurement; current-human labor and software productivity implications matter too, but the paper does not quantify distributional welfare or LMIC effects. The main limitations are external validity, uncertainty about extrapolating the doubling trend, and whether the findings actually change policy thresholds rather than feeding general AI-progress discourse.

AI timing assessment

The paper was first posted in March 2025 and revised in July 2026; arXiv now lists a NeurIPS 2025 reference, so it is no longer best treated as an unreviewed preprint. Some independent public commentary exists, including technical blog reviews and summaries, but the supplied scrutiny evidence was empty and I would not treat the scrutiny gap as fully closed. Timing value for UJ is therefore only moderate: feedback may still matter for policy interpretation, but the main academic review window has likely passed.

Intake, review, and crux connections

AI governance and the economics of AI policy: priority papers · 2026-09-30

Papers identified in The Unjournal's September 2026 AI-governance scoping (David Reinstein's internal planning). On integration (2026-09-30), titles, authors, dates and abstracts were re-resolved from canonical sources (arXiv, Crossref, NBER, or the publisher's own page), never from the scoping notes; the papers were deduplicated against the dashboard and scored by the standard Codex GPT-5.5 (medium reasoning) subscription path. Five related papers already on the dashboard but still awaiting a genuine model score were re-scored by the same path and are labeled as surfaced existing records

The papers were identified in The Unjournal's September 2026 AI-governance scoping (David Reinstein's internal planning). Several are central to live policy debates but were missing from the dashboard or not yet scored. Inclusion is not an endorsement or a completed Unjournal team decision.

Community crux

This is the most misunderstood graph in AI · 92% match

Directly proposes and estimates METR-style 50% task-completion time horizons for AI software tasks.

Community crux

AI Futures Timelines and Takeoff Model: Dec 2025 Update · 78% match

Estimates time-horizon growth trends, directly bearing on exponential versus faster capability projections.

Community crux

METR's 14h 50% Horizon Impacts The Economy More Than ASI · 72% match

Pins AI coding capability to task-duration thresholds used in downstream economic and timeline forecasts.

Scope discussion

Internal scoping aid for prioritization, drafted with AI assistance; not an evaluation and not an endorsement. Source: docs/scope/metr-time-horizons.md in the project repository.

Status: internal scoping aid for prioritization, prepared 2026-10-01. Not an evaluation and not an endorsement. Contains no non-public information.

Paper: Kwa, West, Becker, et al. (METR), arXiv 2503.14499 (2025; later versions through v4, July 2026; listed as NeurIPS 2025). Facts below on the abstract were checked against the arXiv page; items marked "verify" come from background knowledge and should be checked against the paper before being relied on.

Why this memo exists

The dashboard's earlier GPT-5.5 pass scored the paper 23 and labelled it out_of_scope (since superseded: the Opus 5.5 thorough re-evaluation of 2026-10-01, run with this memo as context, gave 58 / monitor), even though its components were high (decision relevance 7.1, prominence 7.5, real-world influence 7.0). The paper also underpins crux UJG05 (capability metrics as policy triggers). The label likely reflects a "technical ML benchmark" reading of scope. This memo sets out the competing readings so a person can decide.

What the paper does

  • Proposes the 50% time horizon: the length of task, measured by how long skilled humans take, at which an AI system succeeds about half the time.
  • Builds a task pool from three sources: HCAST (agentic software and general-skills tasks), RE-Bench (ML research-engineering tasks), and 66 novel shorter tasks (the abstract says "66 novel shorter tasks"; the paper's name for this suite is SWAA, verify).
  • Human baselining: domain-skilled people were timed on the tasks; task length is the human completion time.
  • Statistical model: for each model, success is regressed on log human time with a logistic fit; the 50% crossing point is the horizon.
  • Headline results: Claude 3.7 Sonnet had a horizon of roughly 50 minutes; horizons across frontier models have doubled about every seven months since 2019, possibly faster in 2024.
  • Extrapolation: if the trend and its external validity hold, within about five years AI systems could automate many software tasks that take humans a month.
  • Stated drivers: reliability, adapting to mistakes, reasoning and tool use. The authors discuss external validity and dangerous-capability implications as limitations.

How it is used

  • Widely cited as a forecasting anchor for AI capability timelines and for "automation of AI R&D" scenarios (it feeds directly into cruxes about intelligence-explosion speed and economic scenarios in this dashboard).
  • Used by AI-safety institutes, labs' frontier-safety frameworks and think tanks as a shorthand for capability growth, and sometimes proposed as a trigger metric for policy thresholds.
  • Frequently over-read in media. MIT Technology Review (5 Feb 2026, "This is the most misunderstood graph in AI") argues that a one-hour horizon does not mean a model can replace an hour of real work, and that the tasks are almost all software and benchmark-style.

Scope question, argued both ways

Out of scope (technical ML capability measurement). The object studied is an AI benchmark; the method is ML evaluation; there is no human-subjects welfare outcome, no economic outcome, and no policy intervention. Unjournal's usual remit is quantitative social science and decision-relevant empirical work. Technical AI/CS papers are excluded in the repository's scope statement. Evaluators from ML might also review it through venues such as NeurIPS.

In scope (quantitative measurement and forecasting used directly in welfare-relevant decisions). The paper is less a model-building paper than a measurement-and-forecasting paper whose central output (a doubling time and a five-year extrapolation) is used in decisions by funders, governments and labs. Unjournal explicitly covers forecasting methodology, measurement validity and catastrophic-risk-relevant evidence. The questions that matter here are social-science questions: does the measure mean what users take it to mean, are the human baselines a valid yardstick, is the statistical inference honest, and is it fit to serve as a policy trigger. Few other papers combine such high influence with such open measurement questions.

Which parts fit. In: construct validity and the mapping to real-world work, the baseline sample and incentives, the statistical model and uncertainty, extrapolation, and fitness as a policy trigger and forecasting input. Out (or only for specialist input): detailed task-suite engineering, agent scaffolding choices, and model-by-model benchmark mechanics.

What evaluators must understand and scrutinize

  1. Construct validity. Does "human time" index difficulty in a way that is comparable across tasks and between humans and models? Tasks are self-contained and algorithmically scorable, which differs from messy real work.
  2. Human baselines. Who was timed, how selected, how paid, how much context they had (baseliners were skilled but often unfamiliar with the specific codebase; verify), how failures and long tasks were handled, and whether incentives differ from workplace settings. Later evidence on context-rich experts being much faster would shift horizons.
  3. Statistical model. Logistic fit on log time, the handling of task clustering and few long tasks, confidence intervals on the horizon, the sensitivity of the 80% horizon (substantially shorter than the 50% horizon), and the weight on a small number of models in the fit.
  4. Extrapolation risk. An exponential fit over about six years, sparse data at long horizons, possible regime changes (reasoning models, post-2024 acceleration), and the leap from "doubling every seven months" to "a month-long task in five years".
  5. Contamination and selection. Training-data overlap, task selection that favours automatable tasks, and benchmark saturation.
  6. Mapping to economic capability. Software task length is not occupational substitution; check what evidence links horizon to labour-market or productivity outcomes (including field-experiment evidence on developer productivity, verify).
  7. Updates and critiques. Subsequent versions (to v4) and METR's follow-up time-horizon work, independent replications, and critiques such as the MIT Technology Review piece, the Gary Marcus commentary and others. An evaluator should read the current version, not v1.
  8. Fitness as a policy trigger. Gaming and Goodhart risk, measurement noise at thresholds, task-suite maintenance, and whether a single scalar suits a regulatory trigger.

Evaluator mix (2-3 people, from these profiles)

  • ML evaluation and benchmarking: agentic-benchmark design, contamination, scaffolding sensitivity.
  • Psychometrics or item-response measurement: construct validity, difficulty scales, human-vs-machine comparability.
  • Statistics and forecasting: logistic fits with clustered items, uncertainty, trend extrapolation, forecasting track records.
  • Economics of AI and task-based labour: mapping capability to work and productivity.
  • AI policy / governance: use as a threshold or trigger.

Suggested pairing: one measurement/statistics person (psychometrics or forecasting), one ML-evaluation person, and, if a third is available, an economics-of-AI or AI-policy person. Avoid evaluators who work at METR or on closely competing time-horizon benchmarks without declaring conflicts.

Scope recommendation

Treat as in scope as a nonstandard or applied evaluation, not a conventional academic paper evaluation: restrict the brief to measurement validity, statistical inference, extrapolation and decision use, and explicitly exclude detailed engineering of the task suite. Because the paper has gone through several versions and a NeurIPS publication, check whether a stable, citable version exists and what an Unjournal evaluation would add beyond that review. The original GPT-5.5 score of 23 / out_of_scope looked too low for a paper with this influence; on the narrower brief, a middling-to-high score is defensible, tempered by the existing scrutiny (substantial public critique already exists, which reduces neglectedness).

Suggested evaluation questions

  1. Does the 50% time horizon measure a meaningful, comparable quantity across models and tasks, and what does it not measure?
  2. How sensitive are the horizon and the doubling time to baseline selection, baseliner context, task mix and statistical specification? Are reported intervals adequate?
  3. How reliable is the extrapolation to month-long tasks, and what would falsify it?
  4. What evidence links horizon to economically relevant work, and how large is the gap?
  5. Is the metric fit to act as a policy or safety trigger, and under what safeguards?
  6. How should decision-makers (funders, regulators, forecasters) use the headline number differently after the later updates and critiques?

Public attention and use

The paper is among the most visible in AI forecasting. It is published at NeurIPS 2025, METR has extended it in follow-up work, and MIT Technology Review and independent bloggers have critiqued how the headline chart is read. Substantial scrutiny already exists, which lowers the value of added scrutiny but does not settle the scope question.

Conference publication listing / discoverability

NeurIPS 2025 proceedings version

The paper appears in the NeurIPS 2025 proceedings, so it has passed one round of ML-venue review.

Source: NeurIPS · Relationship: peer-reviewed venue

What the search did not establish

  • No formal independent replication of the horizon estimates surfaced in the targeted search.
  • We did not search government and safety-institute documents systematically, though the metric is commonly cited there.
  • The most decision-relevant next signal is whether independent re-analysis or later versions change the doubling-time estimate.

Targeted search checked 2026-10-01. Search scope: Targeted exact-title searches across the open web, institutional and author pages, news and public social-media results, EA Forum/LessWrong, and policy/white-paper contexts. Evidence records distinguish commissioning or report use from independent discussion, media attention, indexing, and post-publication policy use. A search miss is reported as uncertainty, not proof of absence. Entries with check_depth "quick" rest on roughly one to two searches and should be read as especially uncertain; entries checked on 2026-10-01 were added for the shortlist of papers without human ratings.

Human feedback so far

The privacy-safe aggregate contains 2 current ratings: 2 team and 0 public. The human mean is 81.2/100. The team has not made a final prioritization decision.

(repeating above: Is this really the most recent work in this area? Has it not been updated since February 2026? -- NICE: on the arxiv site, I see a July 20, 2026 version. Even if not, the methods could be considered carefully, and people could also consider the very question of whether it needs updating and perhaps take steps to update it itself. This paper is very sparsely rich, and evaluators may have to drill down and ask the authors questions for more context and detail. There is likely to be some computer science technical content, but also at least some social science/psychometric content and statistical and modeling considerations that economists and others will be able to consider. )

David Reinstein · 89/100 · 2026-10-02