AI-assisted working assessment; human adoption pending
Initial evaluation: 2026-10-04; last substantive update: 2026-10-07; last attention check: 2026-10-07.

The October 7 check is retained from the scenarios working branch; this storage reconciliation did not conduct a new broad attention search.

Download Markdown source · Editable repo source (collaborator access)

Human input and prioritization discussion

AI-assisted working evaluation: Economic Scenarios for Transformative AI

Evaluation date: 2026-10-04

Last attention check recorded: 2026-10-07

Storage reconciliation: 2026-10-07

AI-assistance disclosure. This working assessment supports The Unjournal’s prioritization process. Its judgments and ratings have not been adopted as a commissioned human evaluation or a team decision. A human evaluator should verify the factual claims, methods, and ratings before signing or submitting an official evaluation.

Executive summary

Economic Scenarios for Transformative AI is unusually useful because it turns broad claims about AI-driven growth and labor displacement into a small set of explicit parameters: the share of work AI can affect, diffusion/adoption, task-level productivity gains, automation versus augmentation, new-task creation, capital supply, and labor-market adjustment. The authors are clear that the three named scenarios are not predictions and receive no probabilities.

The paper's computational implementation is now unusually well checked. An independent Econ-ARK reproduction reimplemented the model from the paper, recovered twelve implementation details that were rounded or left unstated, and reports that it reproduces all 226 checkable published numbers and the scenario explorer's monthly paths to numerical precision. This materially changes the best Unjournal evaluation task: another arithmetic/code replication has low marginal value. The important questions are now whether the scenario assumptions and economic structure are good representations of the worlds being discussed.

The largest substantive limitation is that the spectacular outcomes in the extreme scenario are chiefly consequences of exogenous scenario paths. By 2030 the extreme case assumes a large affected task mass, rapid diffusion, high task-level productivity, 90% automation among affected instances, and no reinstatement of new human tasks. Those assumptions can be useful for stress testing, but the model does not itself establish how likely they are.

The labor-market results also depend strongly on a deliberately coarse structure: workers are divided into cognitive and all-other occupations; AI directly affects only the cognitive group; cross-group mobility is frictional; relative wages are sticky; and capital supply follows an assumed elasticity rather than household saving dynamics. This is a sensible transparent benchmark, but extreme-scenario unemployment and distributional results should be read as model-implied conditional outcomes, not direct empirical forecasts.

The public survey is a useful elicitation of intuitions, but the headline statement that the typical respondent lands near the substantial scenario is partly model-mediated. Of 10,980 surveyed adults, the scenario-output table uses the 3,259 who answered all five model-input questions. Several economically important parameters not elicited from respondents—new-task reinstatement, posting speed, wage rigidity, and capital-supply elasticity—are fixed at the paper's substantial-scenario values for every respondent.

Bottom line: This is a high-value evaluation object, but the independent reproduction means the focus should shift almost entirely from code correctness to model validation, scenario calibration, sensitivity, and interpretation. The most valuable comparison is with external evidence and forecasts—including Karger et al.—and with alternative macro structures, not another line-by-line reproduction.

What the paper contributes

The model maps five broad technological/labor assumptions into paths for GDP, wages, factor shares, employment and unemployment through 2030. It combines:

The three named scenarios span a wide range:

These are conditional outputs, not probability-weighted forecasts.

Computational reproducibility: now a major strength

The external Econ-ARK REMARK by Alan Lujan independently reimplemented the monthly model. We did not rerun that model in this evaluation. The reproduction report states:

This provides unusually thorough computational scrutiny. It also identifies an actionable transparency improvement: later versions should print the extra precision and state the few implementation details that currently live only in code.

Crucially, this reproduction validates the mapping from assumptions to reported outputs, not the assumptions themselves. That is now where evaluation effort should go.

The central substantive issue: scenario assumptions drive the outcomes

The paper is admirably explicit that the paths are scenarios. In Table 1, several of the most consequential 2030 values are labeled scenario assumptions.

For the substantial and extreme scenarios, key inputs include:

The extreme case assumes, among other things, automation of 90% of affected task instances and zero new human tasks reinstated for automated tasks. The combination is deliberately severe.

The paper's value is therefore not that it estimates a 32% GDP increase or 18% cognitive unemployment as the expected future. It says: if a specified bundle of capability, diffusion, productivity, automation and adjustment assumptions obtains, this model implies those outcomes.

That conditional should stay attached whenever headline figures are cited.

Production structure

Gross complementarity across tasks

The baseline elasticity of substitution across tasks is 0.5, so tasks are gross complements. This creates weak-link behavior: gains on automated tasks are limited by tasks that remain human-performed.

This is defensible and grounded in task-model literature, but the correct elasticity at the macro level under rapid AI transition is highly uncertain. If substitution across tasks or occupations is easier, output and labor-demand responses can look very different.

A global sensitivity analysis over the substitution elasticity would add more value than treating 0.5 as a common structural constant.

Cognitive versus all-other occupations

The model divides the labor force into two groups. Direct AI exposure is concentrated in the cognitive group; non-cognitive occupations benefit indirectly from higher demand.

This makes the mechanism legible, but it is coarse in several ways:

The extreme result that non-cognitive wages rise 34% while cognitive wages fall is therefore a sharp implication of the two-group equilibrium structure, not something directly observed or established by task-exposure evidence.

Capital supply is a major distributional lever

The model's baseline elasticity of capital supply is 3. The authors are explicit that this is not directly measured for the transition they model. Their robustness section shows that changing capital supply can materially alter who receives the gains from AI.

This deserves even more prominence. The model does not currently solve household saving with an Euler equation; the paper notes that the relationship between high growth and returns to capital could change future versions.

For near-transformative scenarios, capital accumulation is not a side issue. Data-center construction, energy, chips and complementary physical investment can become constraints. Conversely, rapid saving/investment responses can prevent the rental price of capital from rising as much as the baseline implies.

The distributional headline—labor share falling to 45% in the extreme scenario—should therefore be paired with capital-supply sensitivity.

Labor-market adjustment and wage rigidity

The labor-market module is one of the paper's strongest contributions because it does not assume frictionless instantaneous occupational switching. Displaced cognitive workers search across both groups with an occupational-switching disadvantage.

But the unemployment result is strongly shaped by:

The paper already varies wage rigidity and shows large differences. In the extreme scenario, cognitive unemployment ranges from very low when wages adjust flexibly to over 20% when wages are stickier. This is exactly the kind of sensitivity that should accompany any headline unemployment number.

The next step is not merely more values of the same parameter. It is validation against historical large reallocations and, where possible, individual-level transition data by occupation.

New tasks / reinstatement

The reinstatement ratio is substantively central. In the extreme scenario it is set to zero: automated tasks are not offset by newly created tasks for cognitive labor. In the substantial case it is 0.25 and in the modest case 0.50.

That is a transparent stress-test choice, but evidence on new-task creation is one of the hardest unknowns in automation economics. The difference between “AI does existing tasks” and “AI reorganizes production around many new human-complementary tasks” can dominate the labor-demand result.

A high-value evaluation should show full outcome surfaces over automation share and reinstatement ratio rather than only the three bundled scenario combinations.

Innovation / recursive improvement

The paper's endogenous innovation channel is deliberately conservative. Even in the extreme scenario, the modeled R&D contribution is small because research remains constrained by physical tasks and the semi-endogenous growth structure.

At the same time, the authors say the extreme scenario may plausibly be associated with recursive self-improvement. This creates an important interpretive point: the model does not generate its extreme capability path from recursive self-improvement; it takes the rapid path as an exogenous scenario input.

That makes the companion recursive-self-improvement literature highly relevant. The Cunningham et al. framework addresses whether AI-R&D feedback could sustain acceleration; this paper addresses what some economic consequences would be conditional on a fast capability/adoption path. Neither should be used as if it validates the other's uncertain input.

Survey evidence

The Morning Consult survey includes 10,980 US adults, but only 3,259 answered all five items needed to run an individual model simulation. The paper uses survey weights for this complete-case sample.

This raises several evaluation questions:

The public is also not necessarily the most informative group for forecasting specialized AI capability or macroeconomic adjustment. The survey is valuable for understanding public expectations, not as validation that the substantial scenario is objectively most likely.

A particularly useful cross-paper exercise would feed Karger et al.'s expert conditional/unconditional forecasts into this framework, or derive scenario parameters consistent with those forecasts, to see where expert and public expectations truly differ.

“Resources exist to compensate the losers”

In the extreme scenario the model finds aggregate gains large enough that a transfer around 9% of GDP could, arithmetically, restore cognitive workers' aggregate income to its no-AI level.

This is a useful accounting exercise, but it is not a welfare result. The model does not track:

The statement should therefore remain “aggregate resources are sufficient in the model,” not “a feasible compensation system would make everyone whole.”

Missing channels

The authors themselves list many important omissions:

These omissions are not reasons to discard the framework. They define where results are more conditional and where extensions matter most.

For The Unjournal, the most decision-relevant question is which omitted channels plausibly change the sign or order of magnitude of the labor, growth and distribution outcomes rather than merely adding detail.

October 7 attention check retained from the working draft

The paper became CEPR Discussion Paper 21939 on 15 September 2026. The independent Econ-ARK reproduction adds unusually strong computational scrutiny. A fresh 7 October check did not identify clear government, central-bank, or funder operational use comparable to Yale Budget Lab's use of Karger et al.; attention is substantial in economics and AI-policy discussion, but demonstrated operational uptake should not be overstated.

Appropriate uses now include stress-testing macro/labor outcomes conditional on explicit transition assumptions and identifying which structural assumptions drive disruptive versus benign adjustment. The paper does not establish probabilities for its scenarios or the likelihood of recursive self-improvement.

Current attention and reuse

Attention is substantial:

This is a much stronger external-review/reproducibility environment than many candidates on the prioritization dashboard. That lowers the value of a generic review and raises the bar for an Unjournal evaluation: it should add substantive model criticism, parameter synthesis or cross-model comparison that is not already provided by the reproduction.

Highest-value evaluation work

  1. Treat the Econ-ARK reproduction as the computational baseline rather than duplicating it.
  2. Build global sensitivity surfaces over affected mass, diffusion, productivity, automation and reinstatement rather than only three bundled scenarios.
  3. Add sensitivity to task-substitution elasticity and occupational grouping.
  4. Validate labor-transition parameters against large historical reallocations and current occupation-switching data.
  5. Re-estimate survey-implied outcomes with nonresponse/imputation checks.
  6. Compare the public-survey scenario with Karger et al.'s expert forecasts.
  7. Separate the effect of capital-supply assumptions on GDP, wages and labor share.
  8. Explore at least one alternative model with endogenous saving/capital accumulation.
  9. Link extreme capability paths to the recursive-self-improvement model rather than treating “extreme” as a free-standing technological trajectory.
  10. Make explicit which conclusions are robust across structures versus artifacts of the two-group/task model.

Questions for the authors

  1. Which scenario parameters do you think have the weakest empirical basis and dominate outcome uncertainty?
  2. Can you release a global sensitivity analysis over the five technological/adjustment inputs rather than one-at-a-time robustness checks?
  3. How do results change with a higher task-substitution elasticity?
  4. How much of the extreme cognitive-unemployment result is due to the zero reinstatement assumption?
  5. What historical episodes best validate the occupational-switching and vacancy-posting parameters at the scale of the substantial/extreme scenarios?
  6. Can the survey analysis report characteristics and answers of complete versus incomplete responders?
  7. What do survey-implied outcomes look like if the four fixed structural parameters are varied instead of held at substantial-scenario values?
  8. How would you map Karger et al.'s expert forecasts into this model?
  9. How should readers interpret the extreme scenario given that the model does not endogenously generate recursive AI capability growth?
  10. Will a future version endogenize household saving/capital accumulation, and which current distributional results do you expect to be most sensitive?

October 7 structural audit

The paper's main residual uncertainty is not arithmetic but model structure. The cognitive/non-cognitive two-group labor market compresses heterogeneity in transferability, licensing, geography, age, wage loss, and labor-force exit. Wage rigidity matters greatly for whether adjustment appears as unemployment or lower wages. The extreme scenario is also the regime in which a reduced-form capital-supply elasticity is least trustworthy, because data centers, power, transmission, chips, and construction all face physical lags.

The independent Econ-ARK reproduction is especially informative: it reproduces all 226 checkable published numbers and the public explorer paths to numerical precision, while documenting twelve implementation details needed beyond the paper's printed text. Most are higher-precision calibration values or explorer definitions; one concerns within-month ordering. This is a specification-transparency issue rather than evidence that the headline calculations are wrong.

A strong follow-up should jointly vary task substitution, automation versus augmentation, new-task creation, wage rigidity, search frictions, and capital adjustment rather than relying on bundled scenarios or one-at-a-time checks. It should also inspect the survey microdata and mapping pipeline if those become available.

Current Unjournal metrics (provisional)

Reference group for percentile metrics: serious economics / AI-policy research encountered in the last three years that aims to inform comparable policy or research decisions. These are provisional AI-assisted estimates retained from the 7 October working draft. The ranges express subjective uncertainty, not fitted statistical intervals or team ratings.

MetricMidpoint90% subjective credible interval
Overall assessment8472–92
Claims, strength and characterization of evidence8068–89
Methods: justification, reasonableness, validity, robustness8776–94
Advancing knowledge and practice8978–95
Logic and communication9182–97
Open, collaborative, replicable science9689–99
Relevance to global priorities / usefulness for practitioners9181–97
Journal tier this work should reach (0–5)4.33.6–4.9
Journal tier this work will reach (0–5)4.23.3–4.9

Diagnostic ratings (0–5)

Overall assessment

The paper succeeds at its stated purpose: it provides a transparent common language for translating different beliefs about AI capability and adoption into economic outcomes. The new independent reproduction strongly supports the correctness of the implementation.

What remains uncertain is more economically important: whether the parameter bundles correspond to plausible future paths, whether a two-group task model captures the relevant reallocation, how capital adjusts, and whether new tasks emerge.

That makes a substantive Unjournal evaluation worthwhile, but it should explicitly build on rather than duplicate the external reproduction. The strongest contribution would be a cross-model and cross-forecast audit that shows which dramatic outcomes survive alternative assumptions and which are specific to the scenario design.

Sources

The public reproduction repository contains the implementation and checks.