AI-assisted working assessment; human adoption pending
Initial evaluation: 2026-10-04; last substantive update: 2026-10-05; last attention check: 2026-10-05.

October 5 is the recorded public-brief attention check. The October 7 commissioning case and storage reconciliation are separate updates.

Download Markdown source · Editable repo source (collaborator access)

AI-assisted working evaluation: Delays to Frontier AI in the EU and UK

Evaluation date: 2026-10-04

Last substantive update: 2026-10-05

Last attention check recorded: 2026-10-05

Storage reconciliation: 2026-10-07

Paper: John Lidiard, Oleksandra Vereschak, Tom Gibbs, and Markus Anderljung, Delays to Frontier AI in the EU and UK Primary source: https://www.governance.ai/research-paper/delays-to-frontier-ai-in-the-eu-and-uk

AI-assistance disclosure. This working assessment supports The Unjournal’s prioritization process. Its judgments and ratings have not been adopted as a commissioned human evaluation or a team decision. A human evaluator should verify the factual claims, methods, and ratings before signing or submitting an official evaluation.

Executive summary

GovAI assembled 375 public LLM releases by Meta, Google, OpenAI, and Anthropic from June 2018 through May 2026 and compared release timing in the United States, European Union, and United Kingdom. The descriptive release-history dataset is the strongest part of the study. The authors find that 11% of releases were delayed or not released in the EU and 7% in the UK.

The causal attribution is weaker. Of 68 delayed or non-released cases, the authors tentatively classify 56 as primarily regulatory, most often related to data protection. The report is commendably explicit that assessing the causes of delay is its weakest part. Companies rarely provide complete explanations, public statements can be strategic, and many classifications require inference from rollout patterns, product differences, and surrounding regulatory events.

The most important currentness correction is now concrete: the dataset ends in May 2026, whereas European Commission enforcement powers for general-purpose-AI obligations under the AI Act began on 2 August 2026. The paper's finding that it saw "no strong evidence" of AI-Act-caused delays is therefore a pre-enforcement result. It should not be presented as evidence about the current post-August enforcement regime.

This timing creates a high-value evaluation design for The Unjournal. Rather than only reviewing a static June report, an evaluator can reproduce the release dates and cause coding and then extend the same protocol across the August enforcement boundary.

Bottom line: high-value target as a replication plus update. The descriptive evidence is fairly auditable; cause coding needs independent recoding and agreement statistics; and a post-August extension now has more decision value than re-litigating the pre-enforcement result alone.

Main claims

  1. EU and UK users sometimes receive frontier-model releases later than US users. Strong descriptively within the authors' release definition and four-company sample.
  2. About 11% of releases were delayed or withheld in the EU and 7% in the UK over June 2018-May 2026. Supported by the dataset, subject to unit-of-analysis sensitivity.
  3. Regulatory barriers explain most observed delays or non-releases. Plausible but substantially less secure, because causal coding is inferential.
  4. Data-protection rules were the most common regulatory barrier. Supported as the authors' classification, not as a clean treatment-effect estimate.
  5. The AI Act had not produced strong evidence of delays in the sample. Correctly scoped only to the pre-enforcement period.
  6. Frontier-access losses from regulation may be smaller than some public discussion suggests. A useful hypothesis, with limited external validity beyond these firms, models, and dates.

Dataset construction

The study covers public releases by four major frontier developers. Releases through 31 October 2025 were identified and reviewed manually. From November 2025 onward, an AI-assisted process was used for initial identification, followed by multiple rounds of human review.

Important scope choices:

These choices are reasonable but mean the headline percentage partly reflects the study's ontology of a "release."

Unit-of-analysis sensitivity

The 375 observations are not 375 independent technological events. One underlying model family can produce several rows through app/API modes, modalities, versions, or context windows. A single legal or product issue can therefore affect multiple observations.

A replication should show results under several units:

This would clarify whether the aggregate delay rate is broadly distributed or driven by a small set of model families and rollout practices.

Firm composition

Meta has a notably higher delay/non-release rate than the other firms, around 26% for the EU and 15% for the UK in the authors' dataset. The aggregate result can therefore be sensitive to firm composition.

Useful robustness tables would include:

Cause coding

This is the central inferential weakness.

Companies rarely publish a complete causal account. They may emphasize regulation when other factors also matter, or avoid mentioning legal concerns while making product changes in response to them. The authors therefore combine public statements, rollout sequencing, regulator actions, product characteristics, and surrounding context.

A strong public evaluation should have at least two independent coders classify every delayed/non-release case while blind to the published causal label. It should report:

The headline "56 of 68 primarily regulatory" result should also be presented under strict and permissive definitions.

Cause classification is not a treatment-effect estimate

Even if a classification that regulation mattered is correct, it does not identify how many days of delay were caused by regulation. A release might have been delayed for product-readiness reasons and then delayed further by legal review.

The study is mainly a cause-classification exercise, not a causal treatment-effect design. A more causal extension might exploit policy timing or cross-jurisdiction comparisons where assumptions are defensible, but anticipation and other contemporaneous changes would still matter.

The AI Act conclusion needs an explicit time boundary

The original report itself notes that the relevant AI Act enforcement period was outside the dataset. As of 5 October 2026, European Commission and AI Act Service Desk materials state that enforcement powers for GPAI obligations have applied since 2 August 2026.

The defensible summary is:

Through May 2026, before the Commission's GPAI enforcement powers began, the study found no strong evidence that the AI Act was a major cause of the observed EU/UK delays.

That is narrower than "GDPR, not the AI Act, caused the delays" and better aligned with what the data can show.

What the descriptive results establish

The study usefully shows that during its observation period European users were not routinely denied access to every leading model for long periods.

The authors report that:

These facts are informative about the pre-enforcement period and should not be mechanically projected forward.

External validity

The four frontier developers are appropriate for a frontier-access question, but they may have unusually mature legal and compliance teams. Smaller developers could face different burdens.

The study also does not fully measure:

Current policy context

The study's decision relevance has increased because the regulatory regime changed after the sample ended. European Commission materials state that the AI Office's enforcement powers for GPAI obligations began on 2 August 2026.

This makes a post-August extension especially valuable. It should ideally pre-register inclusion and attribution rules before coding the new cases.

Highest-value robustness and extension work

  1. Reproduce all 375 release-date rows from public sources.
  2. Obtain or publish the row-level dataset with evidence links.
  3. Independently recode all 68 delayed/non-release cases.
  4. Report inter-rater agreement for cause and confidence.
  5. Show strict and permissive regulatory-attribution bounds.
  6. Reaggregate by model family and company-quarter.
  7. Report leave-one-company-out and equal-firm-weight results.
  8. Extend the sample from June through at least October 2026 using the same rules.
  9. Pre-register post-2-August inclusion and attribution rules.
  10. Separate GDPR/data protection, AI Act, national law, export controls, product readiness, and mixed causes.

Questions for the authors

  1. Is the row-level dataset and evidence file public or available for replication?
  2. Can you provide the exact rubric for assigning primary and secondary causes?
  3. How many cases were independently double-coded before publication, and what was inter-rater agreement?
  4. How do the 11% and 7% rates change when model families rather than release rows are the unit?
  5. How sensitive are results to excluding Meta?
  6. Can you publish a strict lower bound consisting only of cases with explicit company or regulator statements?
  7. Which cases change category under a multi-cause framework?
  8. Are you extending the dataset beyond 2 August 2026?
  9. Would you pre-register the extension's inclusion and attribution rules?
  10. Which outcome besides release delay would best capture regulatory burden or innovation effects?

Unjournal-style ratings

Overall assessment

This is a good empirical governance object because much of it can be checked. The release-history dataset is the main strength; causal attribution from public evidence is the main weakness.

The policy regime has now moved beyond the period the study observes. The most useful evaluation is therefore clear: audit the original release and cause coding, then extend the same protocol across the 2 August 2026 enforcement boundary.

References