Paper-specific consideration page · current tranche rank 4
Measuring Mid-2025 LLM-Assistance on Novice Performance in Biology
Shen Zhou Hong, Alex Kleinman, Alyssa Mathiowetz · Catastrophic risks, the long-term future, forecasting
Substantial attention in AI-bio and forecasting communities; decision use not yet established
Why this page exists. We are considering whether an independent evaluation of this paper would be useful. The synthesis score uses provisional weights, and the human sample is small. It is one input to the decision.
90AI evaluation-priority (shifted by the Opus 5.5 re-evaluation)
81Original AI lens before the Opus shift (not used for the synthesis)
84Human aggregate · n=3 · effective weight 6.0
86Human–AI synthesis
Why this paper is being considered
This pre-registered, investigator-blinded randomized trial tests whether mid-2025 LLM access helps novices complete wet-lab biology workflows. It found no significant difference on the primary endpoint but numerically better results on several tasks. It bears directly on how AI-biosecurity risk assessments should weigh in-silico benchmarks against physical-world uplift. An evaluation would focus on power (low completion rates), outcome definitions, the 'internet only' counterfactual, and external validity to later models.
What an expert evaluation could add: A controlled study is evaluable, but its value for current governance depends on model recency and the decisions it informs. Existing paper-specific public scrutiny needs a fuller check. Evidence and commissioning case. AI-assisted judgment, 2026-10-07; separate from ratings and completed evaluations.
AI-generated criterion ratings and reasoning
Scores come from AI prioritization; the accompanying explanations include AI-assisted source checks. These are provisional judgments, separate from human ratings and commissioned evaluations.
Decision relevance
7.8/10
The incremental welfare contribution is potentially strong because the paper informs a concrete, high-stakes choice: whether AI-biosecurity governance should rely on benchmark evidence or require physical-world uplift studies before drawing conclusions about novice operational capability. The study does not solve biosecurity risk and should not be credited with the total value of pandemic prevention, but it adds rare empirical evidence about real novice performance, showing no significant primary completion effect and only modest, uncertain uplift on pooled and ordinal progress measures. Its value is highest if it changes model-evaluation standards, release-policy thresholds, or funder priorities for AI-biosecurity testing. The GCR pathway is direct but uncertain; LMIC welfare enters mainly through unequal pandemic vulnerability rather than through a studied LMIC intervention. The main downside risk is misinterpretation, either as excessive reassurance or as overclaiming from post-hoc positive patterns.
Paper claim to check In a pre-registered, investigator-blinded RCT with 153 novice participants, LLM access did not significantly increase completion of the core reverse genetics workflow: 5.2% in the LLM arm versus 6.6% in the internet-only arm. Source abstract
Value of added scrutiny
8.5/10
The study has drawn substantial independent attention in biosecurity and forecasting circles, including an EA Forum post, a Forecasting Research Institute follow-up, and a METR lessons write-up. Its framing as an informative null is being debated, which makes it a good candidate for careful evaluation.
Search result No government or frontier-lab policy document citing the trial by name surfaced, although it may be used in safety cases that are not public. Unjournal public-attention search
Timing
8.2/10
Submitted to arXiv on February 18, 2026, this is about seven months old as of September 17, 2026: still a working paper/preprint, likely early enough for feedback to affect revisions, follow-up studies, and policy interpretation. The supplied targeted public-scrutiny search found no distinct leads for substantial independent review, replication, critique, author response, or PubPeer discussion; this should not be treated as proof of no scrutiny, especially given possible private review by funders or advisors, but it does suggest a live scrutiny gap. Because the topic is moving fast, timing value is high, though the usefulness of evaluating mid-2025 systems will decay as newer models and lab-assistance interfaces emerge.
Source record Publication-stage and date evidence should be checked in the linked paper record. EAFORUM
Methodological potential
8.4/10
Evaluation would require careful handling of dual-use details, expertise in wet-lab biology, and statistical review of a complex RCT with low event rates, post-hoc pooled Bayesian modeling, ordinal progression endpoints, and possible multiple-comparison concerns. External validity is hard: novice Boston-area participants in a controlled BSL-2 lab, blocked communication tools, provided equipment, and mid-2025 LLMs may not represent future models, expert users, malicious actors, or real acquisition constraints. A useful evaluation should audit the pre-registration and SAP against reported analyses, inspect reproducibility of code, assess whether task design validly proxies dangerous capability uplift, and avoid both over-reassuring and over-alarming interpretations.
Paper claim to check In a pre-registered, investigator-blinded RCT with 153 novice participants, LLM access did not significantly increase completion of the core reverse genetics workflow: 5.2% in the LLM arm versus 6.6% in the internet-only arm. Source abstract
Prominence
6.5/10
The scoring model estimated prominence from the paper's venue, authors, institutional setting, and visibility. The model did not supply a criterion-specific explanation. Current public-attention status: Substantial attention in AI-bio and forecasting communities; decision use not yet established.
Forum discussion A forum post discusses the trial and argues about the right question to ask of it. We did not check the author's relationship to the study team. EA Forum
Likely influence
7.4/10
The scoring model estimated how far the findings could shape later research or decisions, without supplying a criterion-specific explanation. Current public-attention status: Substantial attention in AI-bio and forecasting communities; decision use not yet established. See public-attention evidence below.
Forum discussion A forum post discusses the trial and argues about the right question to ask of it. We did not check the author's relationship to the study team. EA Forum
What the paper says
Source abstract · Source-supplied abstract resolved from the linked bibliographic record.
Large language models (LLMs) perform strongly on biological benchmarks, raising concerns that they may help novice actors acquire dual-use laboratory skills. Yet, whether this translates to improved human performance in the physical laboratory remains unclear. To address this, we conducted a pre-registered, investigator-blinded, randomized controlled trial (June-August 2025; n = 153) evaluating whether LLMs improve novice performance in tasks that collectively model a viral reverse genetics workflow. We observed no significant difference in the primary endpoint of workflow completion (5.2% LLM vs. 6.6% Internet; P = 0.759), nor in the success rate of individual tasks. However, the LLM arm had numerically higher success rates in four of the five tasks, most notably for the cell culture task (68.8% LLM vs. 55.3% Internet; P = 0.059). Post-hoc Bayesian modeling of pooled data estimates an approximate 1.4-fold increase (95% CrI 0.74-2.62) in success for a "typical" reverse genetics task under LLM assistance. Ordinal regression modelling suggests that participants in the LLM arm were more likely to progress through intermediate steps across all tasks (posterior probability of a positive effect: 81%-96%). Overall, mid-2025 LLMs did not substantially increase novice completion of complex laboratory procedures but were associated with a modest performance benefit. These results reveal a gap between in silico benchmarks and real-world utility, underscoring the need for physical-world validation of AI biosecurity assessments as model capabilities and user proficiency evolve.
Claims to check
In a pre-registered, investigator-blinded RCT with 153 novice participants, LLM access did not significantly increase completion of the core reverse genetics workflow: 5.2% in the LLM arm versus 6.6% in the internet-only arm.
LLM access was associated with modest, uncertain performance gains across tasks, including a post-hoc pooled estimate of about 1.4x higher success for a typical reverse-genetics task with a 95% credible interval of 0.74 to 2.62.
Participants using LLMs progressed further through intermediate procedural steps across tasks, with posterior probabilities of a positive ordinal-progression effect reported at 81% to 96%, suggesting binary completion outcomes may miss partial capability uplift.
Methodological or theoretical issues flagged for evaluation
Evaluation would require careful handling of dual-use details, expertise in wet-lab biology, and statistical review of a complex RCT with low event rates, post-hoc pooled Bayesian modeling, ordinal progression endpoints, and possible multiple-comparison concerns. External validity is hard: novice Boston-area participants in a controlled BSL-2 lab, blocked communication tools, provided equipment, and mid-2025 LLMs may not represent future models, expert users, malicious actors, or real acquisition constraints. A useful evaluation should audit the pre-registration and SAP against reported analyses, inspect reproducibility of code, assess whether task design validly proxies dangerous capability uplift, and avoid both over-reassuring and over-alarming interpretations.
Opus 5.5 re-evaluation (experimental)
Opus raw score: 72/100
Adjusted (+15): 87/100
Opus own action: watchlist
Label from adjusted score: shortlist
Earlier score (gpt-5.5): 78/100
Shift applied to the AI score: +9
Read with care. This is an experimental re-evaluation by a different model (Claude Opus 5.5, 2026-10-01). Across a calibration sample Opus scored about 15 points lower than GPT-5.5, so 15 points are added to compare it with GPT-5.5 scores. The calibration is a small sample (n=44) and is not human-validated, and the two models disagree about the order of papers within the top group. The AI score at the top of this page is the usual evaluation-priority score shifted by the difference between the adjusted Opus score and the earlier holistic score; the synthesis uses that shifted value, and the unshifted score is shown beside it. On the dashboard you can switch the Opus scores off. Read the methods note.
Opus rationale (AI-generated)
This is an unusual and, I think, genuinely valuable candidate: a pre-registered, investigator-blinded RCT (n=153) measuring whether LLM assistance actually improves novice performance on physical laboratory tasks modelling a viral reverse-genetics workflow. It sits directly on a decision that is live right now — whether in-silico biology benchmark scores are adequate evidence of real-world misuse uplift. Frontier-lab safety teams, the UK AI Security Institute, US CAISI, NTI | bio, the Johns Hopkins Center for Health Security, RAND and Open Philanthropy's biosecurity programme are all making resource and threshold decisions that this evidence bears on, and the 'benchmarks overstate physical-world capability' finding is exactly the kind of result that gets cited quickly in policy documents. The scrutiny gap is wide: it is an arXiv working paper with no visible independent review, a small author team, no NBER/CEPR-style institutional signal, and a headline result that is quotable in a way its statistical content may not support. That combination — consequential, likely to be cited, methodologically non-trivial, unvetted — is close to the textbook Unjournal case. My honest concerns are three. First, it is not quantitative social science in our usual sense: the subject matter is experimental biosecurity assessment, and the evaluable content is trial design, power, pre-registration fidelity and Bayesian/ordinal inference rather than economics. We would need evaluators comfortable with RCT methodology plus some biosecurity literacy, which is a narrower pool than usual. Second, the primary endpoint has ~5-7% completion in both arms, so the null is weak evidence of absence; an evaluation's most useful contribution would be assessing whether the study was powered for a policy-relevant effect and how much weight the post-hoc Bayesian 1.4-fold estimate (CrI 0.74-2.62) and the 81-96% posterior-positive ordinal results should carry given multiplicity. Third, any evaluation has to respect the information-hazard boundary the authors evidently designed around and not push for operational specifics. On balance I would put this in monitor, high in the band, and would move it to prioritize if we can line up an evaluator with trial-methodology and biosecurity-assessment experience — the marginal value of independent scrutiny here is high precisely because the result is about to be used in both directions of an argument.
Dashboard details and provenance
Discovery source: EAFORUM
Scoring model: gpt-5.5 (codex headless, high)
Model holistic score: 78.0
Full AI dashboard scoring rationale
This seems high value for AI-biosecurity decisions: it directly tests whether LLMs help novices make progress on wet-lab tasks approximating a viral reverse genetics workflow, rather than relying only on in silico benchmark performance. The users would include the Frontier Model Forum, Anthropic/OpenAI/Google DeepMind safety teams, METR, national AI safety institutes, biosecurity funders such as Open Philanthropy and the Packard Foundation, and policy groups advising governments on frontier model release and biological capability evaluations. It is an arXiv working paper, apparently without journal peer review or substantial public expert scrutiny from the supplied search, and the combination of high policy salience, pre-registered RCT methods, post-hoc Bayesian claims, and information-hazard-sensitive interpretation makes independent Unjournal evaluation especially useful. The main concern is scope and evaluability: this is closer to AI governance/biosecurity evaluation than ordinary social science, and a good evaluation would need both statistical expertise and domain knowledge in synthetic biology and AI safety.
AI decision-relevance rationale
The incremental welfare contribution is potentially strong because the paper informs a concrete, high-stakes choice: whether AI-biosecurity governance should rely on benchmark evidence or require physical-world uplift studies before drawing conclusions about novice operational capability. The study does not solve biosecurity risk and should not be credited with the total value of pandemic prevention, but it adds rare empirical evidence about real novice performance, showing no significant primary completion effect and only modest, uncertain uplift on pooled and ordinal progress measures. Its value is highest if it changes model-evaluation standards, release-policy thresholds, or funder priorities for AI-biosecurity testing. The GCR pathway is direct but uncertain; LMIC welfare enters mainly through unequal pandemic vulnerability rather than through a studied LMIC intervention. The main downside risk is misinterpretation, either as excessive reassurance or as overclaiming from post-hoc positive patterns.
AI timing assessment
Submitted to arXiv on February 18, 2026, this is about seven months old as of September 17, 2026: still a working paper/preprint, likely early enough for feedback to affect revisions, follow-up studies, and policy interpretation. The supplied targeted public-scrutiny search found no distinct leads for substantial independent review, replication, critique, author response, or PubPeer discussion; this should not be treated as proof of no scrutiny, especially given possible private review by funders or advisors, but it does suggest a live scrutiny gap. Because the topic is moving fast, timing value is high, though the usefulness of evaluating mid-2025 systems will decay as newer models and lab-assistance interfaces emerge.
Earlier public-scrutiny search
The supplied OpenAlex/Bing/PubPeer targeted search returned no distinct leads for substantial independent expert reviews, replications, critiques, author responses, or extended public discussions of this specific paper. This is not proof that no scrutiny exists; the study may have received private advisory-board, funder, or community feedback, but no material public scrutiny was provided in the search evidence.
No supporting source links are stored.
Public attention and use
The study has drawn substantial independent attention in biosecurity and forecasting circles, including an EA Forum post, a Forecasting Research Institute follow-up, and a METR lessons write-up. Its framing as an informative null is being debated, which makes it a good candidate for careful evaluation.
The Forecasting Research Institute compared earlier expert and superforecaster predictions with this trial's results. It treats the trial as a reference result.
Source: Forecasting Research Institute · Relationship: external commentary
The organisation that ran the trial published a plain-language summary.
Source: Active Site · Relationship: author-written
What the search did not establish
No government or frontier-lab policy document citing the trial by name surfaced, although it may be used in safety cases that are not public.
No independent statistical re-analysis surfaced.
The most decision-relevant next signal is a follow-up trial with newer models and whether safety frameworks change thresholds in response.
Targeted search checked 2026-10-01. Search scope: Targeted exact-title searches across the open web, institutional and author pages, news and public social-media results, EA Forum/LessWrong, and policy/white-paper contexts. Evidence records distinguish commissioning or report use from independent discussion, media attention, indexing, and post-publication policy use. A search miss is reported as uncertainty, not proof of absence. Entries with check_depth "quick" rest on roughly one to two searches and should be read as especially uncertain; entries checked on 2026-10-01 were added for the shortlist of papers without human ratings.
Human feedback so far
The privacy-safe aggregate contains 3 current ratings: 1 team and 2 public. The human mean is 83.5/100. The team has not made a final prioritization decision.
Worth prioritizing because it provides rare pre-registered randomized evidence on an important and rapidly evolving question. The null primary result is informative, while the unexpectedly low completion rate and more suggestive post-hoc findings make independent methodological scrutiny particularly valuable. Evaluators should focus on statistical power, outcome selection, and the interpretation of pre-specified versus post-hoc findings.