# AI-assisted working evaluation: Economic Scenarios for Transformative AI **Evaluation date:** 2026-10-04 **Last attention check recorded:** 2026-10-07 **Storage reconciliation:** 2026-10-07 > **AI-assistance disclosure.** This working assessment supports The Unjournal’s prioritization process. Its judgments and ratings have not been adopted as a commissioned human evaluation or a team decision. A human evaluator should verify the factual claims, methods, and ratings before signing or submitting an official evaluation. ## Executive summary *Economic Scenarios for Transformative AI* is unusually useful because it turns broad claims about AI-driven growth and labor displacement into a small set of explicit parameters: the share of work AI can affect, diffusion/adoption, task-level productivity gains, automation versus augmentation, new-task creation, capital supply, and labor-market adjustment. The authors are clear that the three named scenarios are **not predictions** and receive no probabilities. The paper's computational implementation is now unusually well checked. An independent Econ-ARK reproduction reimplemented the model from the paper, recovered twelve implementation details that were rounded or left unstated, and reports that it reproduces **all 226 checkable published numbers** and the scenario explorer's monthly paths to numerical precision. This materially changes the best Unjournal evaluation task: another arithmetic/code replication has low marginal value. The important questions are now **whether the scenario assumptions and economic structure are good representations of the worlds being discussed**. The largest substantive limitation is that the spectacular outcomes in the extreme scenario are chiefly consequences of **exogenous scenario paths**. By 2030 the extreme case assumes a large affected task mass, rapid diffusion, high task-level productivity, 90% automation among affected instances, and no reinstatement of new human tasks. Those assumptions can be useful for stress testing, but the model does not itself establish how likely they are. The labor-market results also depend strongly on a deliberately coarse structure: workers are divided into cognitive and all-other occupations; AI directly affects only the cognitive group; cross-group mobility is frictional; relative wages are sticky; and capital supply follows an assumed elasticity rather than household saving dynamics. This is a sensible transparent benchmark, but extreme-scenario unemployment and distributional results should be read as **model-implied conditional outcomes**, not direct empirical forecasts. The public survey is a useful elicitation of intuitions, but the headline statement that the typical respondent lands near the substantial scenario is partly model-mediated. Of 10,980 surveyed adults, the scenario-output table uses the 3,259 who answered all five model-input questions. Several economically important parameters not elicited from respondents—new-task reinstatement, posting speed, wage rigidity, and capital-supply elasticity—are fixed at the paper's substantial-scenario values for every respondent. **Bottom line:** This is a high-value evaluation object, but the independent reproduction means the focus should shift almost entirely from code correctness to model validation, scenario calibration, sensitivity, and interpretation. The most valuable comparison is with external evidence and forecasts—including Karger et al.—and with alternative macro structures, not another line-by-line reproduction. ## What the paper contributes The model maps five broad technological/labor assumptions into paths for GDP, wages, factor shares, employment and unemployment through 2030. It combines: - a task-based production model; - automation versus augmentation; - new human-task creation; - an ideas/R&D channel; - two labor-market groups; - frictional matching and occupational mobility; - sticky relative wages; - an upward-sloping capital supply. The three named scenarios span a wide range: - **Modest:** GDP 1.6% above a no-AI path by 2030. - **Substantial:** GDP 8.3% above the no-AI path by 2030; cognitive wages roughly flat relative to the no-AI path; labor share around 56%. - **Extreme:** GDP 32.4% above the no-AI path by 2030; annual growth reaches about 15%; labor share falls to about 45%; cognitive wages are 11.5% below the no-AI path and cognitive unemployment reaches 17.9%. These are conditional outputs, not probability-weighted forecasts. ## Computational reproducibility: now a major strength The external Econ-ARK REMARK by Alan Lujan independently reimplemented the monthly model. We did not rerun that model in this evaluation. The reproduction report states: - all **226 checkable published numbers** reproduce once twelve omitted/rounded implementation details are recovered; - the reimplementation matches the scenario explorer's monthly paths to about (3.5 imes 10^{-14}); - with the paper's printed calibration alone, 163 of 183 model outputs reproduce; printing one more digit in key parameters is enough to recover all 183; - one within-month ordering is ambiguous in the paper, but the explorer code resolves it and the economically relevant differences are tiny. This provides unusually thorough computational scrutiny. It also identifies an actionable transparency improvement: later versions should print the extra precision and state the few implementation details that currently live only in code. Crucially, this reproduction validates the **mapping from assumptions to reported outputs**, not the assumptions themselves. That is now where evaluation effort should go. ## The central substantive issue: scenario assumptions drive the outcomes The paper is admirably explicit that the paths are scenarios. In Table 1, several of the most consequential 2030 values are labeled **scenario assumptions**. For the substantial and extreme scenarios, key inputs include: - affected task mass by 2030; - diffusion/adoption by 2030; - task-level productivity gains; - automation share; - the “reinstatement ratio” for new human tasks. The extreme case assumes, among other things, automation of 90% of affected task instances and zero new human tasks reinstated for automated tasks. The combination is deliberately severe. The paper's value is therefore not that it estimates a 32% GDP increase or 18% cognitive unemployment as the expected future. It says: **if a specified bundle of capability, diffusion, productivity, automation and adjustment assumptions obtains, this model implies those outcomes**. That conditional should stay attached whenever headline figures are cited. ## Production structure ### Gross complementarity across tasks The baseline elasticity of substitution across tasks is 0.5, so tasks are gross complements. This creates weak-link behavior: gains on automated tasks are limited by tasks that remain human-performed. This is defensible and grounded in task-model literature, but the correct elasticity at the macro level under rapid AI transition is highly uncertain. If substitution across tasks or occupations is easier, output and labor-demand responses can look very different. A global sensitivity analysis over the substitution elasticity would add more value than treating 0.5 as a common structural constant. ### Cognitive versus all-other occupations The model divides the labor force into two groups. Direct AI exposure is concentrated in the cognitive group; non-cognitive occupations benefit indirectly from higher demand. This makes the mechanism legible, but it is coarse in several ways: - many “non-cognitive” jobs contain administrative, communication or planning tasks exposed to AI; - cognitive occupations vary enormously in task composition; - physical AI/robotics is excluded; - workers differ in skills, wealth, age, geography, household structure and switching costs, none of which is represented. The extreme result that non-cognitive wages rise 34% while cognitive wages fall is therefore a sharp implication of the two-group equilibrium structure, not something directly observed or established by task-exposure evidence. ## Capital supply is a major distributional lever The model's baseline elasticity of capital supply is 3. The authors are explicit that this is not directly measured for the transition they model. Their robustness section shows that changing capital supply can materially alter who receives the gains from AI. This deserves even more prominence. The model does not currently solve household saving with an Euler equation; the paper notes that the relationship between high growth and returns to capital could change future versions. For near-transformative scenarios, capital accumulation is not a side issue. Data-center construction, energy, chips and complementary physical investment can become constraints. Conversely, rapid saving/investment responses can prevent the rental price of capital from rising as much as the baseline implies. The distributional headline—labor share falling to 45% in the extreme scenario—should therefore be paired with capital-supply sensitivity. ## Labor-market adjustment and wage rigidity The labor-market module is one of the paper's strongest contributions because it does not assume frictionless instantaneous occupational switching. Displaced cognitive workers search across both groups with an occupational-switching disadvantage. But the unemployment result is strongly shaped by: - wage rigidity; - posting/hiring speed in expanding occupations; - the cross-occupation search discount; - the two-group aggregation. The paper already varies wage rigidity and shows large differences. In the extreme scenario, cognitive unemployment ranges from very low when wages adjust flexibly to over 20% when wages are stickier. This is exactly the kind of sensitivity that should accompany any headline unemployment number. The next step is not merely more values of the same parameter. It is validation against historical large reallocations and, where possible, individual-level transition data by occupation. ## New tasks / reinstatement The reinstatement ratio is substantively central. In the extreme scenario it is set to zero: automated tasks are not offset by newly created tasks for cognitive labor. In the substantial case it is 0.25 and in the modest case 0.50. That is a transparent stress-test choice, but evidence on new-task creation is one of the hardest unknowns in automation economics. The difference between “AI does existing tasks” and “AI reorganizes production around many new human-complementary tasks” can dominate the labor-demand result. A high-value evaluation should show full outcome surfaces over automation share and reinstatement ratio rather than only the three bundled scenario combinations. ## Innovation / recursive improvement The paper's endogenous innovation channel is deliberately conservative. Even in the extreme scenario, the modeled R&D contribution is small because research remains constrained by physical tasks and the semi-endogenous growth structure. At the same time, the authors say the extreme scenario may plausibly be associated with recursive self-improvement. This creates an important interpretive point: **the model does not generate its extreme capability path from recursive self-improvement; it takes the rapid path as an exogenous scenario input**. That makes the companion recursive-self-improvement literature highly relevant. The Cunningham et al. framework addresses whether AI-R&D feedback could sustain acceleration; this paper addresses what some economic consequences would be conditional on a fast capability/adoption path. Neither should be used as if it validates the other's uncertain input. ## Survey evidence The Morning Consult survey includes 10,980 US adults, but only 3,259 answered all five items needed to run an individual model simulation. The paper uses survey weights for this complete-case sample. This raises several evaluation questions: - How different are complete responders from the full representative sample? - Are item nonresponses correlated with AI familiarity or extreme expectations? - What happens under imputation rather than complete-case restriction? - How sensitive are the model-implied public scenarios to the four non-survey parameters fixed at the substantial-scenario values? The public is also not necessarily the most informative group for forecasting specialized AI capability or macroeconomic adjustment. The survey is valuable for understanding public expectations, not as validation that the substantial scenario is objectively most likely. A particularly useful cross-paper exercise would feed **Karger et al.'s expert conditional/unconditional forecasts** into this framework, or derive scenario parameters consistent with those forecasts, to see where expert and public expectations truly differ. ## “Resources exist to compensate the losers” In the extreme scenario the model finds aggregate gains large enough that a transfer around 9% of GDP could, arithmetically, restore cognitive workers' aggregate income to its no-AI level. This is a useful accounting exercise, but it is not a welfare result. The model does not track: - household wealth ownership; - taxes and behavioral responses; - unemployment duration by person; - lost occupation-specific human capital; - nonpecuniary costs; - geographic concentration; - implementation constraints. The statement should therefore remain “aggregate resources are sufficient in the model,” not “a feasible compensation system would make everyone whole.” ## Missing channels The authors themselves list many important omissions: - aggregate-demand feedback; - business cycles; - financial-market disruption; - catastrophic risks; - detailed worker heterogeneity; - physical robotics; - some research/automation feedback; - household saving dynamics. These omissions are not reasons to discard the framework. They define where results are more conditional and where extensions matter most. For The Unjournal, the most decision-relevant question is which omitted channels plausibly change the **sign or order of magnitude** of the labor, growth and distribution outcomes rather than merely adding detail. ## October 7 attention check retained from the working draft The paper became CEPR Discussion Paper 21939 on 15 September 2026. The independent Econ-ARK reproduction adds unusually strong computational scrutiny. A fresh 7 October check did not identify clear government, central-bank, or funder operational use comparable to Yale Budget Lab's use of Karger et al.; attention is substantial in economics and AI-policy discussion, but demonstrated operational uptake should not be overstated. Appropriate uses now include stress-testing macro/labor outcomes conditional on explicit transition assumptions and identifying which structural assumptions drive disruptive versus benign adjustment. The paper does not establish probabilities for its scenarios or the likelihood of recursive self-improvement. ## Current attention and reuse Attention is substantial: - the paper became **CEPR Discussion Paper DP21939** on 15 September 2026; - Anthropic released a public interactive scenario explorer; - it received rapid discussion in economics venues such as Marginal Revolution; - the authors list extensive comments from prominent macro/labor economists; - an independent Econ-ARK reproduction has already appeared. This is a much stronger external-review/reproducibility environment than many candidates on the prioritization dashboard. That lowers the value of a generic review and raises the bar for an Unjournal evaluation: it should add substantive model criticism, parameter synthesis or cross-model comparison that is not already provided by the reproduction. ## Highest-value evaluation work 1. Treat the Econ-ARK reproduction as the computational baseline rather than duplicating it. 2. Build global sensitivity surfaces over affected mass, diffusion, productivity, automation and reinstatement rather than only three bundled scenarios. 3. Add sensitivity to task-substitution elasticity and occupational grouping. 4. Validate labor-transition parameters against large historical reallocations and current occupation-switching data. 5. Re-estimate survey-implied outcomes with nonresponse/imputation checks. 6. Compare the public-survey scenario with Karger et al.'s expert forecasts. 7. Separate the effect of capital-supply assumptions on GDP, wages and labor share. 8. Explore at least one alternative model with endogenous saving/capital accumulation. 9. Link extreme capability paths to the recursive-self-improvement model rather than treating “extreme” as a free-standing technological trajectory. 10. Make explicit which conclusions are robust across structures versus artifacts of the two-group/task model. ## Questions for the authors 1. Which scenario parameters do you think have the weakest empirical basis and dominate outcome uncertainty? 2. Can you release a global sensitivity analysis over the five technological/adjustment inputs rather than one-at-a-time robustness checks? 3. How do results change with a higher task-substitution elasticity? 4. How much of the extreme cognitive-unemployment result is due to the zero reinstatement assumption? 5. What historical episodes best validate the occupational-switching and vacancy-posting parameters at the scale of the substantial/extreme scenarios? 6. Can the survey analysis report characteristics and answers of complete versus incomplete responders? 7. What do survey-implied outcomes look like if the four fixed structural parameters are varied instead of held at substantial-scenario values? 8. How would you map Karger et al.'s expert forecasts into this model? 9. How should readers interpret the extreme scenario given that the model does not endogenously generate recursive AI capability growth? 10. Will a future version endogenize household saving/capital accumulation, and which current distributional results do you expect to be most sensitive? ## October 7 structural audit The paper's main residual uncertainty is not arithmetic but **model structure**. The cognitive/non-cognitive two-group labor market compresses heterogeneity in transferability, licensing, geography, age, wage loss, and labor-force exit. Wage rigidity matters greatly for whether adjustment appears as unemployment or lower wages. The extreme scenario is also the regime in which a reduced-form capital-supply elasticity is least trustworthy, because data centers, power, transmission, chips, and construction all face physical lags. The independent Econ-ARK reproduction is especially informative: it reproduces all 226 checkable published numbers and the public explorer paths to numerical precision, while documenting twelve implementation details needed beyond the paper's printed text. Most are higher-precision calibration values or explorer definitions; one concerns within-month ordering. This is a specification-transparency issue rather than evidence that the headline calculations are wrong. A strong follow-up should jointly vary task substitution, automation versus augmentation, new-task creation, wage rigidity, search frictions, and capital adjustment rather than relying on bundled scenarios or one-at-a-time checks. It should also inspect the survey microdata and mapping pipeline if those become available. ## Current Unjournal metrics (provisional) Reference group for percentile metrics: serious economics / AI-policy research encountered in the last three years that aims to inform comparable policy or research decisions. These are provisional AI-assisted estimates retained from the 7 October working draft. The ranges express subjective uncertainty, not fitted statistical intervals or team ratings. | Metric | Midpoint | 90% subjective credible interval | |---|---:|---:| | Overall assessment | 84 | 72–92 | | Claims, strength and characterization of evidence | 80 | 68–89 | | Methods: justification, reasonableness, validity, robustness | 87 | 76–94 | | Advancing knowledge and practice | 89 | 78–95 | | Logic and communication | 91 | 82–97 | | Open, collaborative, replicable science | 96 | 89–99 | | Relevance to global priorities / usefulness for practitioners | 91 | 81–97 | | Journal tier this work should reach (0–5) | 4.3 | 3.6–4.9 | | Journal tier this work will reach (0–5) | 4.2 | 3.3–4.9 | ## Diagnostic ratings (0–5) - **Importance / decision relevance: 4.5 / 5.** - **Model clarity / usefulness as a common framework: 4.5 / 5.** - **Computational reproducibility: 5 / 5 after the Econ-ARK reproduction.** - **Empirical calibration of structural parameters: 3 / 5.** - **Scenario calibration / probability information: 2.5 / 5 by design; scenarios are not probabilistic forecasts.** - **Labor-market realism: 3 / 5.** - **Robustness currently shown: 3.5 / 5; strong on selected structural parameters, less complete on the scenario bundle.** - **Value of additional public evaluation: 4 / 5, provided it focuses on model validity rather than arithmetic replication.** ## Overall assessment The paper succeeds at its stated purpose: it provides a transparent common language for translating different beliefs about AI capability and adoption into economic outcomes. The new independent reproduction strongly supports the correctness of the implementation. What remains uncertain is more economically important: whether the parameter bundles correspond to plausible future paths, whether a two-group task model captures the relevant reallocation, how capital adjusts, and whether new tasks emerge. That makes a substantive Unjournal evaluation worthwhile, but it should explicitly build on rather than duplicate the external reproduction. The strongest contribution would be a cross-model and cross-forecast audit that shows which dramatic outcomes survive alternative assumptions and which are specific to the scenario design. ## Sources - [Anthropic scenario explorer and overview](https://www.anthropic.com/institute/econ-scenarios) - [CEPR DP21939](https://cepr.org/publications/dp21939) - [Econ-ARK independent reproduction](https://econ-ark.github.io/econ-scenarios/) The [public reproduction repository](https://github.com/econ-ark/econ-scenarios) contains the implementation and checks.