The Unjournal · Research prioritization
Paper-specific consideration page · Top AI-related papers, rank 7

What Work Does Generative AI Do?

Policy-audience dissemination; no independent critique found
Why this page exists. This paper is on a shortlist of papers we are considering for independent evaluation. No human ratings have been submitted yet, so the score below comes from the AI scoring pass alone. Treat it as one provisional input; the team has made no prioritization decision.
76AI evaluation-priority (shifted by the Opus 5.5 re-evaluation)
81Original AI lens before the Opus shift (not used for the synthesis)
None yetHuman aggregate · no ratings submitted
Not availableHuman–AI synthesis needs at least one human rating

Why this paper is being considered

Using a nationally representative survey, the paper builds occupation- and task-level indexes of generative-AI use and concludes that adoption is widespread but shallow. It informs forecasts of labour-market effects and how quickly AI diffuses inside jobs. An evaluation would focus on survey measurement of AI use, the task taxonomy, representativeness, and how adoption translates into productivity or employment effects.

AI-generated criterion ratings and reasoning

Scores come from AI prioritization; the accompanying explanations include AI-assisted source checks. These are provisional judgments, separate from human ratings and commissioned evaluations.

Decision relevance

8.0/10

The paper informs how policymakers and funders should interpret early evidence on generative AI's labor-market effects. If current adoption is broad but shallow, and if exposure scores explain only part of actual use, then governments, labor agencies, education providers, AI-governance organizations, and philanthropic funders should be cautious about using exposure indexes alone to forecast displacement, target retraining, regulate workplace AI, or estimate productivity gains. The task-level adoption indexes could be inputs for BLS/OECD measurement work, Department of Labor workforce planning, Federal Reserve and CEA macro-labor analysis, World Bank and ILO digital-labor guidance, and philanthropic AI-impact prioritization.

  • Paper claim to check A nationally representative worker survey can produce task-level generative AI adoption indexes linked to detailed occupations and tasks. Source abstract

Value of added scrutiny

8.5/10

The authors' Federal Reserve Bank of St. Louis published a public summary, and a CEPR VoxEU column on measuring what work generative AI does appears alongside it. These show policy-audience reach. We found no independent critique, and the St. Louis summary is from the authors' own institution.

Timing

8.5/10

This appears to be an early NBER working paper rather than a peer-reviewed publication, so feedback could still affect revisions and downstream use. The supplied public-scrutiny search did not identify substantial independent expert review, replication, or critique of this specific paper; the results are the NBER PDF, related NBER material, and an NBER conference listing. That leaves a meaningful scrutiny gap for a prominent, likely-cited working paper in an active policy area, though the search evidence is only a lead set and cannot prove that no expert discussion exists elsewhere.

  • Source record Publication-stage and date evidence should be checked in the linked paper record. NBER

Methodological potential

7.5/10

The main challenge is that the paper is primarily a measurement and descriptive adoption study, not a causal estimate of productivity, displacement, wages, or welfare effects. Evaluation would need to assess survey representativeness, task taxonomy validity, weighting, measurement error in self-reported AI use, comparisons with chat-log-based classifications, and whether the indexes are stable enough to guide policy in a very fast-moving technology environment.

  • Paper claim to check A nationally representative worker survey can produce task-level generative AI adoption indexes linked to detailed occupations and tasks. Source abstract

Prominence

9.0/10

The scoring model estimated prominence from the paper's venue, authors, institutional setting, and visibility. The model did not supply a criterion-specific explanation. Current public-attention status: Policy-audience dissemination; no independent critique found.

  • Author-institution summary The Federal Reserve Bank of St. Louis published a summary of the findings. Its research department issued the working paper, so this is not independent. St. Louis Fed

Likely influence

7.0/10

The scoring model estimated how far the findings could shape later research or decisions, without supplying a criterion-specific explanation. Current public-attention status: Policy-audience dissemination; no independent critique found. See public-attention evidence below.

  • Author-institution summary The Federal Reserve Bank of St. Louis published a summary of the findings. Its research department issued the working paper, so this is not independent. St. Louis Fed

What the paper says

Source abstract · Source-supplied abstract wording; HTML entities and whitespace normalized. Not independently compared with the paper PDF.

We measure how workers use genAI for their jobs in a nationally representative survey linking genAI adoption to detailed occupations and tasks. Our data provide the first task-level genAI adoption indexes, which we show can inform analyses of genAI’s labor market impact. Exposure scores explain some, but far from all, of the variation in adoption across occupations and tasks. We also distinguish our indexes from measures based on genAI platform chat logs, which differ conceptually and tend to over-classify chats into generic activities spanning many occupations. Finally, we highlight that current adoption is widespread but shallow: genAI is used across many occupations and tasks, yet within most of them, fewer than half of workers adopt. This indicates substantial variation among workers doing very similar work, suggesting that understanding who adopts may matter as much as understanding which tasks genAI assists.

Claims to check

  • A nationally representative worker survey can produce task-level generative AI adoption indexes linked to detailed occupations and tasks.
  • Existing generative-AI exposure scores explain some, but far from all, variation in actual adoption across occupations and tasks.
  • Current generative-AI use is widespread across occupations and tasks but shallow within most of them, with fewer than half of workers adopting within many task categories.
Methodological or theoretical issues flagged for evaluation

The main challenge is that the paper is primarily a measurement and descriptive adoption study, not a causal estimate of productivity, displacement, wages, or welfare effects. Evaluation would need to assess survey representativeness, task taxonomy validity, weighting, measurement error in self-reported AI use, comparisons with chat-log-based classifications, and whether the indexes are stable enough to guide policy in a very fast-moving technology environment.

Opus 5.5 re-evaluation (experimental)

Opus raw score: 58/100
Adjusted (+15): 73/100
Opus own action: watchlist
Label from adjusted score: watchlist
Earlier score (gpt-5.5): 78/100
Shift applied to the AI score: -5

Read with care. This is an experimental re-evaluation by a different model (Claude Opus 5.5, 2026-10-01). Across a calibration sample Opus scored about 15 points lower than GPT-5.5, so 15 points are added to compare it with GPT-5.5 scores. The calibration is a small sample (n=44) and is not human-validated, and the two models disagree about the order of papers within the top group. The AI score at the top of this page is the usual evaluation-priority score shifted by the difference between the adjusted Opus score and the earlier holistic score; the synthesis uses that shifted value, and the unshifted score is shown beside it. On the dashboard you can switch the Opus scores off. Read the methods note.

Opus rationale (AI-generated)

This is an NBER working paper by Bick (St. Louis Fed), Blandin (Vanderbilt) and Deming (Harvard), extending their influential nationally representative survey of generative-AI adoption to the occupation-task level. It is clearly within Unjournal's scope as quantitative social science on the economic impacts of AI. It speaks directly to a live measurement debate: whether theoretical exposure scores or AI-platform chat logs (such as the Anthropic Economic Index and OpenAI usage studies) are good proxies for actual work use. The finding that chat logs over-classify generic activities and that adoption is 'widespread but shallow' would matter to statistical agencies (BLS, Census BTOS team), Fed and CBO forecasters, and AI-labour researchers, and to the AI labs whose usage data are increasingly cited in policy. As a recent working paper with no visible independent review, an evaluation could usefully probe survey design, the validity of self-reported task use, task-cell sample sizes, and the fairness of the chat-log comparison while the authors can still revise. Concerns: it is descriptive measurement of US workers, with limited direct welfare or LMIC pathways. Unjournal assessors have historically rated tech/AI papers lower. The headline numbers will date quickly. I would place it in the monitor band, as a reasonable candidate if Unjournal wants an anchor evaluation in the AI-labour measurement cluster.

Dashboard details and provenance

Discovery source: NBER
Scoring model: gpt-5.5 (codex headless, high)
Model holistic score: 78.0

Full AI dashboard scoring rationale

This is a strong Unjournal candidate: an NBER working paper by prominent labor economists measuring task-level generative AI adoption in a nationally representative worker survey, with direct relevance for decisions about labor-market policy, workforce training, productivity measurement, inequality, and AI governance. Organizations such as the Department of Labor, BLS, OECD, ILO, CEA, Federal Reserve researchers, World Bank digital-economy teams, Open Philanthropy, and AI-policy groups could use this to decide whether AI exposure measures are tracking real adoption and which workers or tasks merit intervention. It is still a working paper and the supplied scrutiny search finds only NBER/author-related material and conference listings, not substantial independent expert review; evaluation would add value by checking survey design, task classification, weighting, interpretation of “shallow” adoption, and the comparison with exposure and chat-log measures. The main limitation is that the evidence is U.S.-focused and descriptive rather than a causal estimate of welfare, productivity, displacement, or inequality effects, so an evaluation would need to focus on measurement validity and decision usefulness rather than a clean policy treatment effect.

AI decision-relevance rationale

The paper informs how policymakers and funders should interpret early evidence on generative AI's labor-market effects. If current adoption is broad but shallow, and if exposure scores explain only part of actual use, then governments, labor agencies, education providers, AI-governance organizations, and philanthropic funders should be cautious about using exposure indexes alone to forecast displacement, target retraining, regulate workplace AI, or estimate productivity gains. The task-level adoption indexes could be inputs for BLS/OECD measurement work, Department of Labor workforce planning, Federal Reserve and CEA macro-labor analysis, World Bank and ILO digital-labor guidance, and philanthropic AI-impact prioritization.

AI timing assessment

This appears to be an early NBER working paper rather than a peer-reviewed publication, so feedback could still affect revisions and downstream use. The supplied public-scrutiny search did not identify substantial independent expert review, replication, or critique of this specific paper; the results are the NBER PDF, related NBER material, and an NBER conference listing. That leaves a meaningful scrutiny gap for a prominent, likely-cited working paper in an active policy area, though the search evidence is only a lead set and cannot prove that no expert discussion exists elsewhere.

Earlier public-scrutiny search

The supplied search results do not show substantial independent expert scrutiny of this specific paper. They point to the NBER working-paper PDF, NBER digest material for a related generative-AI adoption paper, another NBER working paper page, and an NBER conference listing. These are not independent reviews, replications, critiques, or extended expert discussions, so they do not materially reduce the marginal value of Unjournal evaluation; however, the failed OpenAlex search and limited search scope mean remaining uncertainty is nonzero.

No supporting source links are stored.

Intake, review, and crux connections

Community crux

Conference Report: Threshold 2030 - Modeling AI Economic Futures · 68% match

Representative task-level adoption data directly informs whether AI impacts should appear in near-term labor and productivity aggregates.

Community crux

Economic growth under transformative AI (Trammell & Korinek) · 49% match

Worker-task adoption measures help assess whether genAI is substituting for labor or mainly assisting existing work.

Community crux

Past Automation Replaced Jobs. AI Will Replace Workers. · 42% match

Widespread but shallow task use informs whether AI displacement proceeds gradually or waits for higher capability thresholds.

Public attention and use

The authors' Federal Reserve Bank of St. Louis published a public summary, and a CEPR VoxEU column on measuring what work generative AI does appears alongside it. These show policy-audience reach. We found no independent critique, and the St. Louis summary is from the authors' own institution.

Author-institution summary visible circulation

St. Louis Fed On the Economy: What work does generative AI do?

The Federal Reserve Bank of St. Louis published a summary of the findings. Its research department issued the working paper, so this is not independent.

Source: St. Louis Fed · Relationship: author institution
Research dissemination listing / discoverability

Equitable Growth working-paper listing

The paper is listed by the Washington Center for Equitable Growth.

Source: Equitable Growth · Relationship: bibliographic index

What the search did not establish

  • No exact-title EA Forum or LessWrong discussion surfaced in the targeted search.
  • No comparison paper or replication using other survey sources surfaced.
  • The most decision-relevant next signal is whether forecasters and agencies adopt its adoption indexes.

Targeted search checked 2026-10-01. Search scope: Targeted exact-title searches across the open web, institutional and author pages, news and public social-media results, EA Forum/LessWrong, and policy/white-paper contexts. Evidence records distinguish commissioning or report use from independent discussion, media attention, indexing, and post-publication policy use. A search miss is reported as uncertainty, not proof of absence. Entries with check_depth "quick" rest on roughly one to two searches and should be read as especially uncertain; entries checked on 2026-10-01 were added for the shortlist of papers without human ratings.

Human feedback so far

No human ratings have been submitted for this paper yet, so there is no human aggregate or synthesis score. If you know the work, rate or discuss it on the dashboard card; the team reviews feedback before it affects prioritization.