AI-assisted working assessment; human adoption pending
Initial evaluation: 2026-10-04; last substantive update: 2026-10-04; last attention check: 2026-10-05.

October 5 is the recorded public-brief attention check. The October 7 commissioning case and storage reconciliation are separate updates.

Download Markdown source · Editable repo source (collaborator access)

AI-assisted working evaluation: The AGI Race and Existential Risk

Evaluation date: 2026-10-04

Last attention check recorded: 2026-10-05

Storage reconciliation: 2026-10-07

AI-assistance disclosure. This working assessment supports The Unjournal’s prioritization process. Its judgments and ratings have not been adopted as a commissioned human evaluation or a team decision. A human evaluator should verify the factual claims, methods, and ratings before signing or submitting an official evaluation.

The substantive analysis below preserves the fuller existing draft. The attention date records the earlier public-brief check; storage reconciliation does not imply a new evidence search.

Executive summary

The AGI Race and Existential Risk is a clean and unusually evaluable theory paper. It formalizes a simple strategic mechanism: firms with scarce resources allocate effort between speed and caution; moving faster raises the chance of arriving first but leaves less for safety; catastrophic failure harms all firms. In the model, competition can therefore generate excessive speed, while the effect of adding resources depends on how crowded the race is.

The formal comparative statics are clear, and the paper is commendably explicit about its scope. Its main limitation is not an internal mathematical problem but external validity. The strongest policy-facing results depend on a deliberately stylized environment: symmetric firms, fixed short-run resource budgets, a winner-take-most race, independent exponential arrival thresholds, one-dimensional speed-versus-caution allocation, and a specific reduced-form relationship between caution and catastrophe risk. There is no empirical calibration that tells us where the frontier-AI sector sits relative to the paper's critical market-size threshold.

The paper itself explicitly cautions readers not to treat it as a general case for concentration, resource restriction, or public provision, because market power, distribution, misuse, geopolitical competition, antitrust enforcement and the quality of public institutions are outside the model. That caveat is essential. The paper is best read as identifying mechanisms and counterexamples to simple claims such as “more resources always increase risk” or “more competition is always safer,” not as establishing the net effect of any real-world intervention.

A fresh Becker Friedman Institute research brief on 30 September 2026 has increased the paper's current visibility. This is still primarily academic/economics attention rather than evidence of broad policy adoption. The marginal value of an Unjournal evaluation is therefore high if it focuses on robustness to alternative strategic structures and empirical mapping, rather than re-proving the algebra.

Main claims

  1. In the baseline symmetric model, more competing firms increase aggregate speed and the probability of catastrophe conditional on arrival. Strong within-model result.
  2. There is a critical number of competitors above which firms continue racing even when the expected value of an AGI arrival is negative for each firm. Strong within-model result; empirical location of the threshold is unknown.
  3. More per-firm resources always accelerate expected arrival, but can raise or lower conditional risk depending on whether the race is above or below the critical threshold. Strong within-model comparative static; external validity depends on the fixed-resource and safety technology assumptions.
  4. A firm can benefit from a credible commitment to move more slowly because rivals respond by slowing as well. A useful strategic mechanism, derived locally around the symmetric equilibrium.
  5. A cautious public entrant can improve welfare in some environments. Valid in the modeled environment, but relies on assumptions about the public entrant's safety, objectives, resources and strategic effects.
  6. A planner controlling only the number of firms prefers monopoly in the model. Correct conditional result, but not a general claim about competition policy; the paper itself explicitly says so.

Strengths of the model

The paper makes the strategic forces unusually transparent. It separates:

This decomposition is genuinely useful. It also highlights that “resources” and “competition” cannot be evaluated independently of how they change the allocation between speed and safety.

The paper discusses several assumptions rather than hiding them. It explains why it uses exponential arrival times, why fixed resources are intended as a short-run constraint, and why the winner-take-all payoff can be relaxed to a winner-take-most setup without overturning the basic incentive.

Main external-validity concerns

Symmetric firms

The main equilibrium analysis assumes identical firms. Real frontier developers differ in compute, model quality, governance, safety effort, organizational structure, access to capital, business models and likely response to catastrophe risk.

Heterogeneity could matter qualitatively. A safer leading firm, a much faster but riskier entrant, or a firm with much lower private exposure to global harm can change the strategic externality. The paper itself lists heterogeneous firms as a major extension.

A high-value robustness exercise would solve a small asymmetric model with two or three firm types and ask which comparative statics survive.

Fixed resources and entry

The paper studies both per-firm resource expansion and fragmentation of a fixed industry resource pool, but the headline “competition increases risk” result can be sensitive to which margin is held fixed.

In reality, entry can both:

The model is most compelling as a short-run strategic benchmark. It is much less clear how to map the number of firms into a fixed total resource pool over a multi-year horizon.

One-dimensional speed-versus-safety tradeoff

The model assumes a budget tradeoff: more resources devoted to speed mean fewer resources devoted to caution. This captures an important possibility, but some safety work can be complementary to capability work, reusable across firms, or itself accelerate development by making experiments more reliable.

Likewise, some “speed” investments can improve safety indirectly. A richer model with public-good safety research, shared standards, or complementary safety/capability investment could change the competition result.

Exponential arrival times

The exponential assumption is made for tractability and gives a memoryless race in which winner identity is independent of calendar time. The authors note that some results extend to proportional-hazards and Weibull structures.

This is a reasonable theory simplification, but real AI development has path dependence, lumpy training runs, learning from rivals, changing bottlenecks, and endogenous revision of strategy. The memoryless setup removes dynamic learning and escalation/de-escalation mechanisms that may be central to a real “race.”

Reduced-form catastrophe risk

The probability of catastrophe depends on a “safety ratio” through a specific exponential functional form. This creates a disciplined tradeoff, but there is no empirical calibration for the slope parameter or for how safety effort translates into reduced catastrophic risk.

The critical market-size cutoff therefore has no current empirical estimate. Without a plausible range for the prize, catastrophe loss, discount rate, arrival technology and safety-effectiveness parameter, the model cannot tell decisionmakers whether a real industry is above or below the cutoff.

Internalization of catastrophic harm

All firms suffer a common loss when catastrophe occurs. This captures shared exposure, but real firms may internalize only a small or heterogeneous fraction of global harm. Owners, workers, decisionmakers and society can have different stakes and horizons.

The welfare wedge may therefore be larger or differently structured than in the baseline model. It would be useful to separate private catastrophe cost from social catastrophe cost explicitly and explore heterogeneous internalization.

Interpreting the policy sections

The paper's policy exercises are best read as mechanism demonstrations.

An Unjournal evaluation should preserve this distinction between comparative statics in a model and net policy effects in the world.

Robustness work that would add the most value

  1. Solve the model with heterogeneous firms in speed technology, safety technology and catastrophe internalization.
  2. Separate private catastrophe losses from social catastrophe losses.
  3. Allow safety effort to have public-good spillovers across firms.
  4. Allow some capability investment to complement rather than substitute for safety.
  5. Make total industry resources endogenous to entry and concentration.
  6. Replace exponential arrival with a staged/dynamic development process with learning and strategy revision.
  7. Explore asymmetric information about safety quality and progress.
  8. Calibrate plausible parameter ranges, even coarsely, to show whether the critical threshold can be located empirically.
  9. Compare the model directly with Fudenberg–Koh and other recent pacing/competition models to identify which results are robust across structures.
  10. Distinguish conditional catastrophe probability from cumulative catastrophe probability over a finite policy horizon.

Questions for the authors

  1. Which main comparative statics do you expect to survive heterogeneity in firm size and safety quality?
  2. Can private and social catastrophe losses be separated explicitly, and does the critical threshold change qualitatively?
  3. How would safety knowledge spillovers across firms affect the result that consolidation improves safety in the baseline model?
  4. What happens if entry expands total industry resources rather than reallocating a fixed pool?
  5. Can you provide a dynamic version with observable progress, learning and strategy revision?
  6. Is there any empirical basis for placing today's frontier sector above or below the critical market-size threshold?
  7. How sensitive are the public-entry and commitment results to asymmetric firm safety?
  8. Which policy conclusions do you regard as mechanism illustrations versus claims ready for empirical policy use?
  9. Would you pair this model with Fudenberg–Koh or other race models in a comparative robustness exercise?
  10. What real-world data would be most informative for calibrating the speed-safety tradeoff?

Current attention

The paper appeared as NBER Working Paper 35276 in May 2026 and received seminar feedback at Chicago, Stanford, Zurich, Yale and UIUC. On 30 September, the Becker Friedman Institute published a dedicated research brief summarizing the model and its policy mechanisms. That is a meaningful increase in current academic visibility, but I do not see comparable evidence yet of broad decisionmaker or media uptake.

Evaluation ratings

Overall assessment

This is a strong theory candidate for The Unjournal because its claims are crisp enough to audit and its policy relevance is obvious, while its limitations are also unusually clear.

The most valuable evaluation would not ask whether the model “proves” that concentration, resource constraints, commitments or public entry are desirable. It does not. It would ask which strategic mechanisms remain after relaxing symmetry, fixed resources, one-dimensional safety investment and memoryless arrivals, and whether available data can locate real frontier development anywhere near the model's critical threshold.

That would turn the paper from an elegant set of conditional comparative statics into a more decision-useful map of which conclusions are robust and which depend on specific modeling choices.

Sources