AI-assisted working evaluation: Forecasting the Economic Effects of AI
Evaluation date: 2026-10-04
Last attention check recorded: 2026-10-05
Storage reconciliation: 2026-10-07
AI-assistance disclosure. This working assessment supports The Unjournal’s prioritization process. Its judgments and ratings have not been adopted as a commissioned human evaluation or a team decision. A human evaluator should verify the factual claims, methods, and ratings before signing or submitting an official evaluation.
The substantive analysis below preserves the fuller existing draft. The attention date records the earlier public-brief check; storage reconciliation does not imply a new evidence search.
Executive summary
Forecasting the Economic Effects of AI is a strong and unusually useful elicitation study, but it should not be read as a validated macroeconomic forecasting model. Its main contribution is to make several relevant groups' beliefs comparable under standardized AI-capability scenarios. The survey design is careful in several respects: explicit conditional and unconditional forecasts, uncertainty ranges, detailed resolution criteria, coherence checks, economist reweighting, and multiple aggregation approaches.
The central empirical object is therefore what these recruited respondents believed in late 2025 / early 2026 under the survey's common framing. That is decision-relevant evidence. It is not direct evidence that 2.5% GDP growth, roughly 3.5–4% growth under a rapid-AI scenario, or 55% labor-force participation in 2050 will occur. These long-horizon forecasts have not resolved and cannot yet be validated against outcomes.
The biggest methodological limitation is selection. The study contacted 4,866 experts and obtained 147 expert completions before filtering, with a final sample of 69 economists, 27 AI-industry professionals and 25 AI-policy professionals. Reweighting the economist sample for experience and geography is useful, but cannot correct selection on unobserved AI optimism, willingness to spend about eight hours on the survey, forecasting interest, or other attitudes related to the outcomes. Non-economist expert samples are not reweighted.
A second limitation is interpretation of the scenarios. The rapid scenario is a rich joint bundle spanning AI R&D, software work, autonomous agency, broad cognitive work and robotics. Conditional forecasts answer “what would the economy look like in this jointly specified world?” They do not isolate the causal effect of any one capability. The authors acknowledge that slow-versus-rapid comparisons are not clean causal effects, but some downstream summaries can easily read more causally than the design warrants.
The clearest presentation problem is the phrase that part of the rapid-scenario decline in labor-force participation is “equivalent to around 10 million lost jobs.” The underlying arithmetic is a difference in the number of people in the labor force, not a direct estimate of jobs destroyed. Lower labor-force participation can reflect unemployment, retirement, education, caregiving, discouraged-worker exit or voluntary nonparticipation. I would change the headline wording to “roughly 10–11 million fewer people in the labor force relative to the slow scenario,” with an explicit note that this is not a forecast of 10 million job losses.
The paper is already being used outside the forecasting community. NBER records funding from Open Philanthropy, now Coefficient Giving, and feedback from Congressional Budget Office seminar attendees. Yale's Budget Lab used the scenarios in May 2026 and again in a July tax analysis. This is genuine downstream use, but it is not institutional endorsement of the forecasts: Budget Lab explicitly treats them as scenarios and adds its own modeling assumptions.
Bottom line: This is a high-value public-evaluation target. The strongest contribution is the standardized elicitation and transparent decomposition of beliefs. The weakest interpretation is treating the resulting numbers as calibrated predictions or causal effects. The highest-value follow-up is an audit of nonresponse/selection, weighting, coherence interventions, the disagreement decomposition, and replication access.
Claims to assess
- Relevant expert groups expect substantial AI-capability progress by 2030. This is reasonably supported as a description of the recruited samples. Generalization beyond those samples is limited by selection and, outside economists, lack of reweighting.
- The median economist expects about 2.5% annual U.S. GDP growth, above most government/private baselines used in the paper. Descriptively supported; not evidence yet that 2.5% is more accurate.
- Under the rapid scenario, experts forecast GDP growth around 4%. Supported as a conditional survey result. The exact economist medians are about 3.3% for 2025–29 and 3.5% for 2045–49, with superforecasters and AI experts somewhat higher.
- Rapid AI could push LFPR to about 55% by 2050. Supported as an elicited median. This point estimate is sensitive enough to weighting that the robustness appendix matters.
- This implies about 10 million lost jobs. The wording is not supported by the underlying quantity. It is a difference in labor-force participation, not a direct job-loss estimate.
- Forecasters mostly disagree about economic consequences conditional on capabilities rather than about AI capabilities themselves. Suggestive, but the decomposition depends on the coarse scenario taxonomy and common scenario framing. It should be stated narrowly.
- Surveyed experts' policy-effect forecasts are informative. They are informative as expert beliefs, not causal estimates of policy effects.
Methods audit
Recruitment and nonresponse
The survey ran October 2025 to February 2026. It targeted economists, AI-industry professionals, AI-policy professionals, highly accurate forecasters, and the general public. Participants received detailed background material and resolution criteria; average completion time was about eight hours.
Low expert response relative to invitations is the major external-validity concern. The economist weights adjust toward the invited frame on experience and geography, but only on observed variables. They cannot repair unobserved selection. The paper itself reports that weighting moves GDP estimates modestly toward conservatism and can move longer-horizon LFPR estimates by a moderate-to-large amount.
A useful robustness package would report the headline tables: unweighted; current weights; trimmed weights; richer response-propensity weights using all invitation-frame covariates; and a sensitivity analysis in which nonresponse is correlated with an unobserved AI-optimism variable.
Scenario design and anchoring
Common, detailed scenarios are a real strength because they make cross-group answers comparable. But they also induce a shared frame. Cross-group convergence is therefore convergence conditional on common definitions, historical information and anchors, not independent convergence on an unconstrained forecast.
The rapid scenario bundles many dimensions. Respondents are told that scenarios describe capabilities rather than adoption and should themselves account for regulation, integration delays, social norms and complementary capital. That means a large amount of causal structure remains inside each respondent's mental model. Two respondents can accept the same capability scenario while imagining different adoption paths.
A valuable future split-sample design would vary how much adoption information is fixed, or vary historical anchors, to measure how much the common frame drives convergence.
Aggregation
The main results use weighted medians of each reported percentile separately. This is robust and easy to communicate, but medians of different quantiles do not constitute a coherent joint predictive distribution in general.
For disagreement analyses, the authors fit parametric distributions to each participant's 10th/50th/90th percentiles and pool them. This adds modeling choices: distribution family, tail treatment, winsorization and treatment of fitted distributions that fail coherence checks. The variance-decomposition conclusions should be rerun under flexible quantile interpolation or alternative distribution families.
Coherence interventions
Checking probability sums and internal consistency is sensible. Participants with flagged responses were contacted and could revise their answers. The paper reports small aggregate pre/post changes, which is reassuring.
For full auditability, a replication package should include anonymized raw forecasts, flags, messages sent to respondents, revised values, removal rules, weights and analysis code. I did not locate a public microdata/code package from the FRI or NBER landing pages in a targeted search on 4 October; this is a request for clarification, not evidence that no package exists.
The labor-force result needs special care
LFPR is a sensible metric because it captures labor-force exit that unemployment does not. The economist median is 61.0% unconditionally by 2030, 59.3% under rapid AI in 2030, and 55.0% under rapid AI in 2050.
The long-horizon rapid result is also highly uncertain. The pooled distribution is wide, and reweighting matters more for LFPR than for some GDP results. Any use of the 55% point estimate should therefore carry the uncertainty and weighting sensitivity with it.
Most importantly, fewer labor-force participants are not mechanically “lost jobs.” If the authors want a jobs measure, it should be modeled separately from labor-force participation.
Disagreement decomposition
The paper's attempt to distinguish disagreement about capability trajectories from disagreement about economic consequences is valuable. But the result is not invariant to the scenario bins. If slow/moderate/rapid are coarse, two forecasters with meaningfully different capability trajectories may still be placed in the same bin, causing capability disagreement to appear as within-scenario economic disagreement.
The paper also finds substantial within-forecaster uncertainty. Statements that “forecasters disagree mainly about economic consequences” should specify the component being decomposed and avoid implying that between-person disagreement is the largest source of total predictive uncertainty.
Downstream use and current attention
- NBER says the study was funded by Open Philanthropy/Coefficient Giving and thanks CBO seminar attendees for feedback. This shows funder and policy-institution interest, not endorsement.
- Yale Budget Lab's May analysis uses Karger et al. to construct fiscal/macroeconomic scenarios. Budget Lab had to alter the productivity translation to align the paper's measure with its macro model, which is a useful reminder that reuse involves additional assumptions.
- Yale Budget Lab's July analysis again uses the economist subsample in tax counterfactuals; its methodology note layers GDP, factor-share and distributional assumptions onto the survey inputs.
- A fresh check on 4 October did not identify a newer major institutional adoption that changes this picture. The key signal remains continued use in policy analysis.
Robustness work I would prioritize
- Nonresponse sensitivity under multiple weighting schemes.
- Leave-one-economist-subpopulation-out analyses.
- Separate AI-industry and AI-policy results in every headline robustness table.
- Full before/after coherence-intervention tables.
- Scenario-anchor and adoption-assumption split-sample tests in a future wave.
- Alternative distribution fits and treatment of incoherent fitted distributions.
- Alternative forecast aggregation rules.
- A prospective scoring protocol for 2030 outcomes.
- Replace “lost jobs” with the labor-force quantity actually estimated.
- Release or clarify access to the microdata/code needed for independent reproduction.
Questions for the authors
- Is an anonymized microdata and code package public or shareable, including pre-coherence forecasts, revisions, flags and weights?
- What observable characteristics are available for nonresponding invitees beyond experience and geography?
- How many forecasts were flagged, revised, retained or removed by each coherence rule and outcome?
- Can the headline GDP/LFPR tables be shown unweighted, with current weights, with trimmed weights and under richer response-propensity weighting?
- How sensitive is the disagreement decomposition to distribution families, winsorization and exclusion of incoherent fitted distributions?
- Could future waves test framing/anchor effects by varying scenario detail or baseline information?
- Would the authors revise the “10 million lost jobs” wording to distinguish labor-force participation from employment?
- Which outputs do the authors regard as most transportable into structural/fiscal models, and which require substantial additional assumptions?
- Is there a prospective plan for scoring 2030 forecast accuracy and calibration?
- Can AI-industry and AI-policy samples be reported separately more consistently?
Evaluation ratings
- Importance / decision relevance: 4.5 / 5. High-stakes outcomes and demonstrated downstream scenario use.
- Elicitation design: 4 / 5. Detailed, uncertainty-aware and unusually transparent.
- Population representativeness: 2.5 / 5. Strong selection concern; limited reweighting.
- Robustness shown: 3.5 / 5. Good appendices, but nonresponse and distribution-fit sensitivity can go further.
- Predictive validation: 2 / 5, not yet testable. 2030/2050 forecasts have not resolved.
- Transparency / reproducibility: 3 / 5 pending replication access. Methods are detailed; public microdata/code was not located in the sources checked.
- Value of further public evaluation: 4.5 / 5. Several concrete, adjudicable checks could improve interpretation and reuse.
Overall assessment
I put fairly high confidence on the descriptive statement “this is what the recruited samples answered,” moderate confidence that broad patterns would survive reasonable re-analysis, and much lower confidence in the numeric results as predictions of 2030/2050 outcomes. Keeping those levels separate is the main interpretive discipline the paper and its users need.
A good Unjournal evaluation should focus less on whether 3.5% versus 4% GDP growth feels plausible and more on the parts that can actually be audited now: recruitment and nonresponse, weights, coherence processing, distribution fitting, disagreement decomposition, replication materials, and downstream translation assumptions.