← AI economics & governance: papers and shortlisting

AI governance briefing: overview

A formatted version of the audio briefing recorded in late September 2026 (about 16 minutes; listen or download, plain transcript). Working notes, not Unjournal decisions or evaluations. Some details have changed since recording, including the shortlist, the paper list and some dates; the briefs page is kept current and takes precedence.

Contents: The bottom line · The live policy debate · The papers

This briefing covers papers on AI governance and the economics of AI policy that The Unjournal is considering for evaluation: what each one says, how much attention it is getting right now, and the policy debates it feeds into. These are working notes, not Unjournal decisions or evaluations. It runs about fifteen minutes, in three parts. First, the bottom line. Second, the live policy debate. Third, the papers themselves.

1. The bottom line

Only a few of our candidate papers are actually shaping the policy debate right now. The clearest case is the Forecasting Research Institute's paper, Forecasting the Economic Effects of AI. It is already being used in fiscal and Federal Reserve policy analysis. The most discussed paper in the whole set is METR's time horizons paper, which measures how long a task an AI agent can complete. And the Active Site biology trial, which tested whether AI helps novices in a wet lab, is being cited in current press on AI and biosecurity.

Most of the rest of the shortlist gets very little outside attention: the liability papers, the insurance papers, the game theory papers, and the public opinion survey. That is not a reason to drop them. It means we would be evaluating them before anyone else has weighed in, which is arguably where The Unjournal adds the most value.

The research also turned up three strong recent papers that were not on the first list. One is GovAI's study of delays to frontier AI releases in the EU and the UK. The second is Beraja and Yuchtman's paper, Why Is AI So Contentious, from the Brookings Papers on Economic Activity. The third is a paper called The Economics of Recursive Self-Improvement, by Tom Cunningham, Phil Trammell, Basil Halperin and colleagues. Each one makes claims an evaluator can check, and each one sits on a live policy fight.

Recency is not everything: if a paper is still being used and discussed, it is relevant. That brings back two older papers, METR's time horizons and the biology trial. The remaining question for those two is whether capability measurement counts as governance evidence, or as technical AI safety work, which is outside The Unjournal's focus.

Here is the current working shortlist of five.

  1. the Forecasting Research Institute's paper on the economic effects of AI.
  2. The Economics of Recursive Self-Improvement, with the AGI Race paper as a companion. This is the pacing debate.
  3. Epoch's estimate of chip smuggling to China, with RAND's paper, Buying Time Against AI Proliferation, as a companion. This is the export-control debate.
  4. GovAI's paper on delays to frontier AI in the EU and UK. This is the EU AI Act debate.
  5. METR's time horizons, read together with its twenty twenty-six update, if capability measurement counts. If it does not, the fifth slot goes to the liability set: Gabriel Weil and Cristian Trout on insurance, and Joshua Gans on staged access.

The reserves are Beraja and Yuchtman, the biology trial, the seven-country public opinion survey by Lundgren and Tallberg, the West Point experiment on automation bias, and the paper Who Evaluates AI's Social Impacts, as a low-cost pilot.

2. The live policy debate

Six policy fights are running right now, and our papers feed into each of them.

The first is pacing, and the idea of an intelligence explosion. On the twenty-eighth of July, more than eleven hundred staff at the major AI labs signed a letter called Pacing the Frontier. It asks Washington to back an international effort to deliberately pace automated AI development. Then, on the twenty-eighth of September, a GovAI-led paper with Geoffrey Hinton and Yoshua Bengio among its authors argued that automating AI research could compress years of progress into months. It called for auditors, mandatory reporting and pacing agreements. The papers that speak to this fight are the recursive self-improvement paper, the AGI Race paper, Fudenberg and Koh, and METR's time horizons.

The second fight is export controls on AI models themselves, not just on chips. On the twelfth of June, the Commerce Department required a licence to export two frontier models, then eased the rule for trusted partners. Critics call the episode ad hoc, and say it may even help China. The relevant papers are Epoch on smuggling, RAND on governed access, a paper on how US containment pushed China toward open models, How to Catch a GPU, and GovAI's study of delays in Europe.

The third fight is federal preemption of state AI laws. A House discussion draft in June would freeze state laws on frontier AI for three years. It has not been introduced as a bill, and more than two hundred state lawmakers oppose it. The liability and insurance papers speak to this fight, because state liability law is exactly what preemption would freeze.

The fourth fight is California's next step. On the eighteenth of September, the governor ordered state agencies to recommend AI safety rules by the sixteenth of November. Those recommendations are to cover outside evaluators, verification of company safety frameworks, and incident reporting. The relevant papers are SaferAI's ratings of lab safety frameworks, Who Evaluates AI's Social Impacts, and the paper on designing an AI whistleblower office.

The fifth fight is enforcement of the EU AI Act. The AI Office's powers over general-purpose AI providers took effect on the second of August. Meanwhile the so-called Digital Omnibus pushed deadlines for high-risk systems back to twenty twenty-seven and twenty twenty-eight. GovAI's delay study is central here.

The sixth fight is the new US-China channel. On the twenty-sixth of September, the two governments agreed to set up a channel for AI incidents, with talks resuming in November. Neither side is committing to slow down. Barry Eichengreen's paper on AI in a fragmenting world, the AGI Race paper, and How to Catch a GPU all bear on it.

3. The papers

I'll take them in four groups: the working shortlist, the strong new finds, the older papers to reconsider, and the rest.

The working shortlist

The Forecasting Research Institute's paper, Forecasting the Economic Effects of AI, came out as an NBER working paper in April. The team collected structured forecasts from economists, AI company staff, policy researchers, superforecasters and the public. The median respondent expects US growth of about two and a half percent a year, above baseline forecasts of about two percent. In a rapid AI scenario, experts expect growth near four percent, with labour force participation falling from sixty-two to fifty-five percent by twenty fifty. This paper has the clearest policy use of any in our set. The Yale Budget Lab built fiscal scenarios directly on it, and the monetary economists Stephen Cecchetti and Kermit Schoenholtz used it to frame advice to the Fed's AI task force. It fits us well, as a follow-up to the forecasting tournament we already evaluated. It needs a forecasting methodologist and a macroeconomist.

Epoch's report on chip smuggling to China came out on the twenty-ninth of April. It combines indictments, investigative reporting, and gaps in trade data. For example, China records three point eight billion dollars of server imports from Malaysia, while Malaysia records only six hundred million dollars of exports. Epoch's median estimate is that six hundred and sixty thousand top-end AI chips were smuggled through twenty twenty-five, roughly a third of China's AI computing power. Tom's Hardware and a widely read AI policy blog repeat the estimate. It is a transparent statistical model built on public evidence, which suits a trade economist and a Bayesian modeller.

The AGI Race and Existential Risk, by Ethan Bueno de Mesquita, Wioletta Dziuda and Mattias Polborn, is an NBER paper from May. Firms split resources between speed and safety. When the field is crowded, effort shifts toward speed and the risk of catastrophe rises. Above a critical number of firms, they race even when AGI has negative expected value. It got University of Chicago press coverage and modest social media attention. It is strong theory with checkable assumptions.

Cooperating against Catastrophe, by Drew Fudenberg and Andrew Koh, is only six days old. In their game, firms either race or pace their capabilities as safety research improves. A key result is that hiding internal capabilities makes pacing harder to sustain, which supports disclosure rules. So far it has only social media shares, though it is already in the top fifth for its age.

Then there are the two insurance papers. Gabriel Weil proposes mandatory liability insurance for frontier developers, scaled to plausible worst cases, so that insurers act as private regulators, plus punitive damages for near misses. Cristian Trout asks when insurance really regulates, and concludes it can only work for frontier AI with a carefully designed mandate. Neither paper has much outside attention; Trout's main exposure is his own op-ed in Lawfare. Note that Trout works at an AI underwriting company, which is a relevant interest.

Last in this group is Lundgren and Tallberg's seven-country survey experiment, with over fourteen thousand respondents. People strongly back AI regulation, and prefer safety-oriented rules by about twelve percentage points. That gap is about four times larger than the gaps on public versus private rules, or national versus international rules. The design is strong, but attention so far is minimal.

The strong new finds

GovAI's paper, Delays to Frontier AI in the EU and UK, built a dataset of three hundred and seventy-five model releases from four major labs. Eleven percent were delayed or never released in the EU, and seven percent in the UK. The main cause was data protection law, not the AI Act. In twenty twenty-six, US national security controls became a bigger barrier than European regulation. MLex and Euronews covered it. It has a clear design, coding an evaluator can check, and a causal claim at the centre of the EU simplification fight. This is a strong candidate.

Beraja and Yuchtman's paper, Why Is AI So Contentious, was presented at the Brookings Papers on Economic Activity on the twenty-fifth of September. They explain the backlash against AI as displacement fears plus a moral grievance: AI was built from people's words and creative work without their consent. They argue this backlash will drive future regulation. It comes from a top venue, and its political economy claims can be tested against survey and legislative data.

The Economics of Recursive Self-Improvement models the feedback loops by which AI speeds up AI research, and calibrates them with existing estimates. The finding is that these loops are not yet strong enough to cause runaway acceleration, but they appear to be strengthening. It has explicit parameters an evaluator can challenge, and it is the key economic question in the pacing debate. It looks like one of the best fits.

RAND's report, Buying Time Against AI Proliferation, models when export limits, licensing and compute controls push users into monitored channels, and when they backfire. The economics are clean, but it has had little press. It pairs naturally with Epoch.

Two more new papers are getting a lot of attention. One is the intelligence explosion paper from GovAI and partners, with its striking statistic that lightly supervised AI work on research rose from one percent to twenty-six percent between March and August. It is closer to a white paper, so it is best evaluated alongside the recursive self-improvement paper. The other is Anthropic's Economic Scenarios for Transformative AI, by Anton Korinek, Chad Jones and colleagues, in which labour's share of income falls from sixty to forty-five percent in the extreme scenario. It overlaps with the forecasting paper and is closer to economics than governance, and the authorship comes from a lab.

The older papers to reconsider

METR's time horizons paper defines an AI system's time horizon: the length of task, measured in human working time, that it completes half the time. Horizons doubled roughly every seven months, and METR's updated tracker suggests faster growth. This is the most discussed paper in the set. It has had coverage in forty-six news outlets, including the Washington Post this month, and appears in four policy documents. There are also serious critiques. MIT Technology Review called it the most misunderstood graph in AI, pointing to wide uncertainty, a task set that is mostly coding, and the fact that human time is not the same as difficulty. If we evaluate it, we should evaluate it together with the twenty twenty-six update and METR's own notes on its limitations.

The Active Site biology trial was pre-registered and blinded, with one hundred and fifty-three participants. Its main result was null: about five percent completed a lab workflow with AI help, versus about seven percent with only the internet. A later analysis suggests modest uplift on individual steps, with wide uncertainty. Vox used it on the nineteenth of September, and it has a strong attention score. The design is checkable. One co-author is at METR, so evaluators should be independent of METR.

The Evaluation Science paper, by Laura Weidinger, Deborah Raji and colleagues, is a position paper with no new data. It argues that AI needs a proper evaluation science, like the ones that developed for cars, aviation and drugs. Brookings cited it last week in the debate over outside evaluators' access to lab models. But an evaluation of it would be mostly conceptual, so it is a weaker candidate.

Athey and Scott Morton's paper on AI, competition and welfare argues that upstream market power in AI can hurt displaced workers twice. It is a reasonable theory candidate if we want an antitrust angle. As for the two older Gans papers, on learning about harms and on copyright, both look like weaker candidates now. A newer paper by Andrew Koh and a co-author, called Technology Speed Limits, carries the harms question further. And Gans's own August paper on staged access and liability is fresher.

The rest

Who Evaluates AI's Social Impacts, accepted at ICML, has real, checkable claims. The authors hand-coded one hundred and eighty-six developer reports and two hundred and forty-eight outside evaluations. On a scale of zero to three, developers' own reporting averages under one, while outside evaluations average over two and a half, and developer reporting is falling. The weak point is sampling: the outside sources were chosen because they evaluate a particular dimension, so part of the gap is built in. It would make a good, low-cost pilot for a methods-minded evaluator.

SaferAI's ratings of lab safety frameworks score companies from thirty-four percent down to eight percent. The rubric is contestable, which is good material for evaluation.

The West Point experiment found that military cadets were better calibrated about AI advice than the public, and less than half as likely to follow a mistaken AI. Defense One covered it.

How to Count AIs, by Arbel, Salib and Goldstein, proposes an algorithmic corporation to make AI agents legally accountable. It is the most cited of the liability papers.

The remaining papers are a weaker fit, mostly because they are descriptive or have little attention. These are the paper on US containment and China's open models, How to Catch a GPU, Eichengreen's paper, the whistleblower case studies, the middle-powers mapping, and a framework for evaluating national AI regulation. Eichengreen's paper does suit the theme of The Unjournal's Geopolitics of AI group.

That's the end of the briefing. The briefs page on the Unjournal prioritization dashboard has every paper, with links, summaries and attention sources. You can rate any of them there, and your ratings count on the dashboard.