AI-assisted prioritization assessment, 7 October 2026. These are provisional commissioning judgments; they are not completed Unjournal evaluations or team selections.
A commissioned evaluation is most useful when a qualified independent reader can check consequential unresolved claims and make the results usable in a real decision. Public attention can increase that value by giving a correction or clarification a larger audience. Existing scrutiny reduces the value of repeating the checks it has already done.
For this collection, the case remains promising for focused evaluations. The public discussion is uneven: there is a substantial computational reproduction of Economic Scenarios for Transformative AI, some methodological criticism of METR’s task horizons, and an active debate about liability and insurance. Many other candidates have publicity, citations or downstream use without a public assessment of their central claims. The right response is to specify what an expert evaluation would add for each paper before choosing it.
Implications for the themes
Economic scenarios and forecasts. Reproducing a calculation is valuable but leaves assumptions, parameter calibration, labor adjustment, demand, distribution and policy interpretation open. A scenario is conditional on its inputs. A forecast can be evaluated now through its elicitation, aggregation and calibration methods; waiting for its horizon to arrive is unnecessary. FRI’s forecast elicitation and Anthropic’s structural scenarios could usefully be compared, while preserving their different questions and time horizons.
Recursive self-improvement and AI races. Separate a feedback mechanism from a claim about its size, persistence or imminence. Check model assumptions, empirical indicators and the steps from theory to a policy recommendation. A paired evaluation of the RSI model and the intelligence-explosion synthesis could assess those steps together. For race models, compare which assumptions generate racing or pacing and whether proposed policies survive plausible extensions.
Liability, insurance and compute governance. Formal incentives need to connect to feasible institutions: insurability, information, capital requirements, enforcement, evasion and jurisdiction. Related skeptical scholarship helps define the questions but does not constitute a review of every new model. A thematic comparison may be more useful than several isolated reports repeating the same objections.
Labor markets and development. Prioritize measurement, identification, adoption and exposure assumptions, observation periods and what the evidence implies for adjustment policy. An early null result does not establish an enduring absence of effects. Reviews should distinguish observed employment changes from projected exposure, and account for differences between richer and poorer countries.
Capability benchmarks and governance maps. An expert can check measurement, baselines, coding reliability and external validity. The commissioning case also needs a clear governance decision that depends on the result. A technical benchmark or a descriptive inventory alone may have a weaker fit for this evaluation stream.
What the commissioned output should contain
For a selected paper, ask for a small set of explicit claims with supporting passages, an assessment of how well each is supported, the assumptions and sensitivity checks that matter, and the implications for a named policy or funding choice. Where useful, compare nearby papers, respond to the strongest existing criticism and invite an author response. This can turn scattered discussion into a public, source-linked assessment without reproducing work already done.
The opportunity cost matters. A narrower expert review or paired assessment may be better than a full report when only one uncertainty remains. Defer when the central claims are already well covered, the proposed checks cannot resolve them, the policy window has passed, or the work falls outside scope. No commissioning decision or numerical ranking change follows automatically from these notes.
Search method and limits
The 7 October check searched public discussion of the leading candidates and common source/version URLs, and reviewed public Hypothes.is annotations. The earlier personal-annotation search helped locate comments; the existence or absence of someone’s annotations does not establish research quality or evaluation value. Reading notes, informal reactions, author publicity, related scholarship, computational reproduction and independent methodological critique are labeled separately below.
The search was strongest for the working shortlist, papers getting attention, and METR. Other records are marked partial or pending rather than treated as unreviewed. Some sources are about older versions or related work. Search failures and sparse results leave uncertainty. Reports of institutional use indicate a possible audience; they do not validate the claims. Source classifications describe the work found, not verified reviewer credentials or a complete inventory of criticism.
The scheduled prioritization pass receives these dated public notes and fresh search leads. It revisits high-priority scrutiny on a bounded monthly cycle, preserving prior scores and keeping editorial judgments separate from human ratings. Public notes remain curated and their dates do not advance merely because a build ran.
Working shortlist · Targeted public-scrutiny check · assessment dated 2026-10-07
Institutional coverage raises visibility; it leaves the model’s policy interpretation largely unchecked. A comparative theory evaluation with Cooperating against Catastrophe could test whether racing results survive realistic extensions.
Remaining gap: Firm heterogeneity, endogenous resources, safety spillovers, dynamic learning and how the critical threshold maps to real-world quantities.
Case for commissioning: Promising if an industrial-organization/game-theory expert can identify assumptions that change the policy conclusion. Pairing the models may add more than two separate summaries.
Working shortlist · Targeted public-scrutiny check · assessment dated 2026-10-07
The search found sparse public discussion rather than a substantive independent review. A paired theory assessment could clarify when transparency and pacing improve outcomes and how those conclusions compare with the AGI Race model.
Remaining gap: Information assumptions, equilibrium selection, safety progress, commitment and enforcement, and the steps from a two-firm model to disclosure or coordination policy.
Case for commissioning: Promising as a companion theory paper if the evaluation can test policy robustness rather than merely restate the equilibria. Sparse search results leave uncertainty about existing scrutiny.
Evidence and its limits (1)
Hacker News discussion(Brief public reaction). A short supportive discussion and a link to related press commentary; no detailed independent methods review was found there.
Working shortlist · Targeted public-scrutiny check · assessment dated 2026-10-07
No substantive direct review was found in the targeted search. Related skepticism about AI liability and insurance defines useful tests, but does not establish that this paper’s proposed mechanisms fail.
Remaining gap: Judgment-proof defendants, liability allocation, compulsory insurance, correlated catastrophic losses, capitalization and what insurers can observe or enforce.
Case for commissioning: Conditional on a tractable quantitative/formal core. Commission legal-economic mechanism checks and compare the insurance papers, rather than a broad doctrinal overview.
Working shortlist · Targeted public-scrutiny check · assessment dated 2026-10-07
Author explanations and broader insurance debates leave the paper-specific mechanism tests open. A comparative review could assess the conditions under which insurance actually improves frontier-AI safety.
Remaining gap: Risk observability, incentives to reduce legal exposure versus underlying harm, insurer capital and exclusions, compulsory coverage and regulation of the insurance market.
Case for commissioning: Potentially useful paired review with judgment-proofness and staged-access models. Existing skeptical scholarship should shape the questions; avoid treating related criticism as a completed assessment of this draft.
Working shortlist · Targeted public-scrutiny check · assessment dated 2026-10-07
News coverage reports a mismatch between public preferences and regulatory options. A public evaluation would add most by checking survey measurement, sampling and the inference from preferences to regulatory legitimacy.
Remaining gap: Representativeness, question wording, cross-country comparability, statistical uncertainty and which conclusions follow from stated preferences.
Case for commissioning: Conditional on a concrete regulatory decision and access to methods/data. Publicity alone does not close the measurement gap; confirm the decision use before commissioning.
Working shortlist · Targeted public-scrutiny check · assessment dated 2026-10-07
Downstream fiscal modeling and monetary-policy commentary show a potential audience, but do not validate the forecast elicitation. An expert can assess the method now, including uncertainty and conversion from capability to economic outcomes.
Remaining gap: Expert selection, question interpretation, aggregation, calibration, correlated judgment and the distinction between capability-conditional forecasts and calendar-time predictions.
Case for commissioning: Promising methodological review with explicit claim assessment. Compare forecast construction with structural scenarios and identify what can be checked now; avoid promising to verify long-horizon outcomes immediately.
Evidence and its limits (2)
Yale Budget Lab fiscal and economic outlook(Downstream use). Uses the forecasts in fiscal modeling with adjustments to productivity definitions. Use establishes relevance, not accuracy.
Productivity and employment: AI’s promise and pitfalls(Independent policy commentary). Discusses forecast uncertainty and monetary-policy implications; this is not evidence of central-bank adoption or a methodological replication.
Working shortlist · Targeted public-scrutiny check · assessment dated 2026-10-07
Public explanation and a trade-data follow-up raise useful questions about the estimates. Neither is a fully independent validation. An expert assessment could test uncertainty in diversion estimates and the policy conclusions about export controls.
Remaining gap: Double counting, price and chip-equivalent conversions, destination versus final use, extrapolation from partial routes, and the effect of smuggling on the counterfactual effectiveness of controls.
Case for commissioning: Promising if underlying estimates are auditable. Focus on bounds and decisions that change across them; pair with governed-access policy only where the economic mechanisms connect.
Evidence and its limits (2)
How banned AI chips end up in China(Contributor commentary). Explains why substantial evasion need not imply controls are ineffective. The contributor helped develop underlying estimates, so this is not fully independent scrutiny.
Malaysia–China mirror-trade follow-up(Same-organization follow-up). Trade discrepancies are consistent with diversion but are not proof of the full smuggling quantity.
Press coverage establishes interest without independently checking the attribution of release delays. A focused evaluation could test coding and causal interpretation before those results are used in regulatory arguments.
Remaining gap: Sample selection, regional release definitions, treatment of non-release, attribution to regulation rather than business choices, and uncertainty around coded reasons.
Case for commissioning: Promising empirical audit if event-level data and coding can be inspected. Keep GDPR-related findings separate from claims about AI Act effects: the sample ends before its August 2026 enforcement date.
Getting attention now · Targeted check; conference discussion not reviewed · assessment dated 2026-10-07
Brookings conference exposure provides a potential policy audience, but the conference discussion was not reviewed in this search. An evaluation case depends on checking the paper’s own evidence and the distributional/political mechanisms it identifies.
Remaining gap: The link between economic effects, beliefs and political contention, identification or model assumptions, and the policy conclusions that survive alternative explanations.
Case for commissioning: Conditional and worth inspecting the conference discussion first to avoid repeating it. A public political-economy claim assessment could still fill gaps; visibility should not itself trigger a downgrade.
Evidence and its limits (1)
BPEA Fall 2026 conference(Conference exposure; discussion unreviewed). Lists an AI session including the paper. The availability of a recording does not establish what scrutiny it received.
Public reading and author discussion do not yet amount to a comprehensive empirical test. A paired expert review with the intelligence-explosion paper could clarify which feedback claims are supported and which indicators would change the conclusion.
Remaining gap: The mapping from research input to capability progress, bottlenecks, returns to reinvestment, measurement of feedback strength, and the conditions for sustained acceleration.
Case for commissioning: Potentially high value if the framework is influencing decisions and an economics-of-innovation expert can assess its assumptions. Use the current arXiv version, distinguish mechanism from forecast, and commission one paired assessment rather than overlapping reports.
Evidence and its limits (2)
METR preliminary economics note(Author explanation / earlier version). Frames the model and missing empirical inputs. It is not independent scrutiny and predates the September arXiv draft.
RSI reading series, essay 11(Independent reading commentary). Distinguishes positive feedback from acceleration and asks what quantity is accelerating. A useful issue lead, not a complete quantitative validation.
The search found author/institutional material without a substantial independent paper-specific review. A useful evaluation would test the economic and enforcement assumptions behind governed access.
Remaining gap: Adversary responses, delay estimates, substitution and evasion, access controls, costs and the circumstances in which buying time improves outcomes.
Case for commissioning: Conditional on an auditable quantitative model and a decision that turns on its parameters. Compare with smuggling evidence and test realistic enforcement scenarios before commissioning a broad policy review.
Evidence and its limits (1)
RAND governed-access report(Original / institutional source). The report and companion author work explain the proposal; they do not constitute independent scrutiny.
Public reactions and policy discussion leave a large gap between evidence of AI-assisted research and claims about a sustained intelligence explosion. The useful evaluation is a claim-by-claim assessment paired with the RSI economics model.
Remaining gap: Automation measurements and baselines, compounding conditional claims, the timescale and reference path for acceleration, and evidence connecting each proposed policy to welfare outcomes.
Case for commissioning: A targeted synthesis review is promising if it separates evidence, conditional extrapolation and recommendations. Ask an evaluator to assess the strongest quantitative premises and policy links; a conventional replication alone would miss the main contribution.
An independent reproduction covers the published calculations; public debate leaves the economic assumptions open. A focused expert evaluation of calibration, labor adjustment, demand and policy interpretation could still add substantial value.
Remaining gap: How sensitive are wages, employment and growth to task substitution, capital supply, new-task creation and worker reallocation? Which scenario assumptions have empirical support, and how far do the policy conclusions follow?
Case for commissioning: Promising if scoped around assumptions and external validity. Reuse the existing reproduction, compare alternative models, and separate conditional scenarios from predictions. A macro/labor expert assessment could make the strongest objections and the authors’ answers usable together.
Evidence and its limits (6)
Econ-ARK reproduction(Computational reproduction). Reports matching all 226 checkable published numbers after recovering twelve implementation details. This establishes reproducibility of the modeled calculations, not the realism of the assumptions.
AI, redistribution, and the size of the pie(Independent economic commentary). Argues adjustment may be easier and fiscal capacity larger than pessimistic interpretations suggest; uses a longer horizon than the paper’s 2030 scenarios.
Will AI soon lead to double-digit GDP growth?(Related economic argument). References the scenarios while questioning whether sustained double-digit growth is likely soon. It is a broader argument, not a full review of this model.
The mirage of AI-driven GDP growth(Public critique). Challenges capital and demand assumptions. The author discloses AI assistance; the objections need an expert response rather than being treated as established refutations.
Scenario explorer and outside feedback(Author response and disclosure). Describes incorporated economist feedback and unresolved limitations; outside commenters were not asked to endorse the work.
Public annotation on demand effects(Reading note). One of seven public annotations by another account, largely summaries in Chinese. Credentials are unknown; these do not establish independent expert review.
A coverage map can inform where to commission evaluations, but its value depends on reliable classification and what counts as an independent assessment. Paper-specific scrutiny remains to be checked.
Remaining gap: Sampling, coding reliability, omitted assessments and whether the map distinguishes publicity, audits and claim evaluation.
Case for commissioning: Potentially useful enabling work. Audit the map first and use it to select consequential evidence gaps, rather than assuming that more evaluations everywhere are valuable.
No independent-review source is recorded here. This does not establish that none exists.
A comparative governance map may support policy learning. The commissioning case depends on coding reliability and a decision that changes when the map is corrected; independent scrutiny has not been established here.
Remaining gap: Jurisdiction coverage, legal-source dates, comparability and distinctions between announced and implemented rules.
Case for commissioning: Conditional on a specific comparative-policy use. A narrow source/coding audit may add more than a general expert review.
No independent-review source is recorded here. This does not establish that none exists.
An evaluation of safety frameworks needs clear criteria and evidence about actual implementation. Independent public scrutiny of this assessment has not yet been checked in detail.
Case for commissioning: Scope decision first: commission a governance measurement review only if it can test consequential comparisons, rather than repeating provider descriptions.
No independent-review source is recorded here. This does not establish that none exists.
Labour markets and development · Partial check; earlier-version commentary · assessment dated 2026-10-07
Commentary on the earlier version emphasizes small early labor-market effects; it does not assess the current revision. A useful expert review would check identification and the observation window before extrapolating to future disruption.
Remaining gap: Measurement of adoption and task changes, earnings/hours estimates, power, heterogeneous effects and how the revised evidence differs from the earlier paper.
Case for commissioning: Promising empirical labor-market evaluation if data access and revision changes are clear. Explain what early evidence can and cannot tell policymakers; verify current public scrutiny before final selection.
Evidence and its limits (1)
Large language models, small labor-market effects(Commentary on earlier version). A slow-transition interpretation of the older draft. This is not an independent assessment of the current Still Waters revision.
Labour markets and development · Partial search; fuller scrutiny check needed · assessment dated 2026-10-07
The trial may warrant evaluation on its own merits, but applying AI to one matching setting is a weaker fit for this stream’s focus on AI’s primary economic impacts and governance. No substantial independent review was established in the limited check.
Remaining gap: Randomization and estimand, implementation, employment outcomes, heterogeneity and transfer to other settings.
Case for commissioning: Route through the wider research dashboard if promising. Confirm scope before allocating this round’s expert capacity.
No independent-review source is recorded here. This does not establish that none exists.
Labour markets and development · Partial public-scrutiny check · assessment dated 2026-10-07
Public commentary highlights distributional and connectivity differences, without independently validating the exposure estimates. An expert review could check task mapping and whether the development-policy conclusions follow.
Remaining gap: Cross-country task content, exposure versus adoption, data coverage, digital complements and the distribution of productivity gains within countries.
Case for commissioning: Promising if measurement uncertainty can be connected to a concrete development decision. Compare the exposure framework with evidence on adoption constraints rather than treating exposure as realized employment effects.
Evidence and its limits (1)
Disruption without dividend commentary(Independent public commentary). Emphasizes how aggregate exposure measures can hide unequal opportunities; does not replicate or validate the underlying estimates.
Labour markets and development · Partial search; fuller scrutiny check needed · assessment dated 2026-10-07
Lessons from earlier digital adoption may inform AI development policy, but the transfer to AI needs explicit testing. No substantial independent review was established in the limited check.
Remaining gap: Representativeness, causal versus descriptive evidence, infrastructure and organizational complements, and which mechanisms plausibly transfer to AI.
Case for commissioning: Potentially useful development review if a concrete policy choice depends on the analogy. Compare with task-exposure evidence and identify testable adoption constraints.
No independent-review source is recorded here. This does not establish that none exists.
Found on the dashboard · Targeted public-scrutiny check · assessment dated 2026-10-07
Author summaries were found, with little substantive independent criticism in the search. An expert review could test the staged-access and liability mechanism against insurance, enforcement and evasion alternatives.
Remaining gap: Monitoring, thresholds, damages and proof, liability allocation, equilibrium assumptions and comparative effects of staged versus unrestricted access.
Case for commissioning: Promising as a companion in the liability/access theme if its formal results affect a real design choice. Verify the current draft and avoid repeating shared objections across reports.
Found on the dashboard · Paper-specific scrutiny check pending · assessment dated 2026-10-07
Conceptual distinctions matter when they change liability rules, but a public commissioned review needs testable consequences. Independent paper-specific scrutiny remains to be checked.
Remaining gap: Consistency of individuation criteria, boundary cases and their consequences for damages, responsibility and enforcement.
Case for commissioning: Conditional on a clear legal-economic mechanism. Pair with agentic liability work if the shared conceptual issue is the main contribution.
No independent-review source is recorded here. This does not establish that none exists.
Found on the dashboard · Paper-specific scrutiny check pending · assessment dated 2026-10-07
The interaction framework needs a tractable link to outcomes or legal decisions before a full evaluation is commissioned. Existing independent scrutiny is not yet established here.
Remaining gap: How interaction categories map to responsibility, incentives, evidential burdens and difficult multi-agent cases.
Case for commissioning: Scope decision first. A focused legal-economic claim assessment could be useful if the framework changes a concrete rule; otherwise a broad doctrinal review may have lower fit.
No independent-review source is recorded here. This does not establish that none exists.
Found on the dashboard · Paper-specific scrutiny check pending · assessment dated 2026-10-07
Claims that containment changes open-model development could matter for export-control policy. The public-scrutiny check is pending, so attention should not yet be interpreted as validation or neglect.
Remaining gap: Identification of containment effects, alternative explanations, measurement of openness and strategic firm responses.
Case for commissioning: Potentially useful if the causal/model core is auditable and policy counterfactuals differ materially. Compare with the smuggling and governed-access evidence.
No independent-review source is recorded here. This does not establish that none exists.
Found on the dashboard · Paper-specific scrutiny check pending · assessment dated 2026-10-07
A verification taxonomy is most useful if it helps choose mechanisms that work under realistic evasion. Independent scrutiny remains to be checked.
Remaining gap: Coverage, adversarial responses, verification cost, false positives and enforcement feasibility across jurisdictions.
Case for commissioning: Conditional on a quantitative or otherwise checkable comparison. Use a targeted feasibility assessment rather than treating a catalogue as evidence of effectiveness.
No independent-review source is recorded here. This does not establish that none exists.
Found on the dashboard · Paper-specific scrutiny check pending · assessment dated 2026-10-07
A broad account of AI and fragmentation needs specific assessable claims to justify expert commissioning. Independent public scrutiny is still to be checked.
Remaining gap: Empirical support for fragmentation channels, scenario assumptions, distributional effects and policy counterfactuals.
Case for commissioning: Conditional on identifying a consequential quantitative core and a live decision. Otherwise a thematic synthesis may add more than a full paper review.
No independent-review source is recorded here. This does not establish that none exists.
Found on the dashboard · Paper-specific scrutiny check pending · assessment dated 2026-10-07
An empirical judgment study can inform human oversight, but generalization beyond its populations and tasks is central. Independent scrutiny remains to be checked.
Remaining gap: Sampling, task validity, experimental design, uncertainty and transfer from study decisions to operational AI oversight.
Case for commissioning: Potentially useful measurement review if linked to a governance decision. Avoid extrapolating military/general-public comparisons beyond the observed design.
No independent-review source is recorded here. This does not establish that none exists.
Found on the dashboard · Paper-specific scrutiny check pending · assessment dated 2026-10-07
Case studies may inform institutional design, but recommendations need a checkable evidential bridge. Independent paper-specific scrutiny remains to be checked.
Remaining gap: Selection of cases, coding, counterfactual reporting routes, generalization to AI firms and feasibility of proposed protections.
Case for commissioning: Conditional on a concrete office-design decision. A focused assessment of recommendations and evidence may be more useful than a generic qualitative review.
No independent-review source is recorded here. This does not establish that none exists.
Found on the dashboard · Paper-specific scrutiny check pending · assessment dated 2026-10-07
A national-regulation framework may help policy comparison if its criteria are reliable and meaningful. Independent scrutiny is not yet established here.
Remaining gap: Source dates, scoring/coding reliability, implicit welfare weights and evidence connecting framework scores to outcomes.
Case for commissioning: Conditional on a decision use and an auditable comparison. Check the criteria and disputed classifications before commissioning a full report.
No independent-review source is recorded here. This does not establish that none exists.
Substantive public criticism already targets human baselines, benchmark realism and horizon extrapolation. A further evaluation needs to resolve those disputes and their policy implications rather than repeat them.
Remaining gap: Human task-time baselines, reliability thresholds, selection of software tasks, trend estimation and the mapping from benchmark completion to economically useful autonomous work.
Case for commissioning: Potentially valuable targeted measurement review, conditional on this stream’s scope. Require a specific governance or economic decision; technical benchmark assessment alone may fit elsewhere.
Is METR underestimating LLM time horizons?(Independent alternative analysis). A different reliability criterion produces a steeper extrapolation; interpretation is sensitive to human baselines and extrapolation choices.
A controlled study is evaluable, but its value for current governance depends on model recency and the decisions it informs. Existing paper-specific public scrutiny needs a fuller check.
Remaining gap: Statistical power, task/outcome validity, participant population, treatment contrast and transfer to newer models.
Case for commissioning: Scope and recency decision first. If still policy-relevant, commission a bounded methods/claim assessment; avoid assuming a new replication is warranted solely because the topic is important.
No independent-review source is recorded here. This does not establish that none exists.
A framework for evaluation may improve later commissioned work, but its own review needs assessable claims about measurement. Independent scrutiny remains to be checked.
Remaining gap: Validity of proposed evaluation constructs, evidence for recommended practices and generalization across applications.
Case for commissioning: Conditional enabling value. A focused comparison of evaluation methods may improve the whole program more than a conventional full-paper report.
No independent-review source is recorded here. This does not establish that none exists.
A model of learning about harms may clarify adoption policy, but this older paper’s remaining scrutiny gap needs checking. Prominence or citations alone do not settle the commissioning case.
Remaining gap: Learning assumptions, reversibility, welfare accounting and robustness of the optimal adoption rule.
Case for commissioning: Revisit if the mechanism is currently used and a theoretical check could change policy. Prefer a thematic comparison when the same assumptions recur elsewhere.
No independent-review source is recorded here. This does not establish that none exists.
Copyright policy affects AI incentives and welfare, but an additional review needs a live decision and a specific uncovered mechanism. Public scrutiny of this older paper remains to be checked.
Remaining gap: Welfare assumptions, market structure, licensing/enforcement and the policy options’ comparative effects.
Case for commissioning: Conditional on current policy use and the remaining gap. Avoid duplicating well-covered legal debates; commission a focused economic assessment if it changes the choice.
No independent-review source is recorded here. This does not establish that none exists.