Why this paper is being considered
Using Danish administrative data linked to worker and workplace surveys, the paper finds precise null effects of AI chatbot adoption on earnings and recorded hours, alongside task restructuring and occupational switching. It is among the most-cited early causal studies of generative-AI labour-market effects. An evaluation would focus on the difference-in-differences design, adoption measurement, and external validity beyond Denmark and the early period.
What an expert evaluation could add: Commentary on the earlier version emphasizes small early labor-market effects; it does not assess the current revision. A useful expert review would check identification and the observation window before extrapolating to future disruption. Evidence and commissioning case. AI-assisted judgment, 2026-10-07; separate from ratings and completed evaluations.
AI-generated criterion ratings and reasoning
Scores come from AI prioritization; the accompanying explanations include AI-assisted source checks. These are provisional judgments, separate from human ratings and commissioned evaluations.
Decision relevance
8.0/10
The paper informs decisions about whether governments and funders should prioritize worker retraining, AI diffusion subsidies, labor-market monitoring, employment protection, education-system adaptation, or slower/more targeted AI deployment. It is especially relevant to OECD, ILO, World Bank, EU and national labor ministries, Open Philanthropy and other AI-governance funders, and research teams at major AI labs trying to distinguish causal labor-market evidence from exposure forecasts and product-performance studies.
- Paper claim to check Linked Danish adoption surveys and administrative labor records show widespread chatbot adoption and task reorganization but precise near-term null effects on earnings and recorded hours. Source text (unverified type)
Value of added scrutiny
8.0/10
The paper has circulated under an earlier title and is widely discussed in the economics literature, but our two searches mainly found bibliographic records and the author's page. No independent critique or policy use was documented by what we saw.
Timing
8.0/10
The paper is an NBER working paper first issued in May 2025 and revised in March 2026, so independent evaluation could still matter for interpretation, revisions, and policy uptake. Search evidence indicates media, blog, and informal public discussion, but I do not see clear evidence of substantial independent expert replication or a formal critique; this keeps the marginal value of Unjournal evaluation relatively high, though the paper has already received seminar and conference feedback.
- Source record Publication-stage and date evidence should be checked in the linked paper record. TARGETED_CURATED
Methodological potential
8.0/10
The core empirical setting is Denmark, with strong administrative data but limited external validity for lower-income countries, informal labor markets, and sectors such as agriculture, education, and health. Evaluation should separate causal near-term wage/hour evidence from descriptive adoption patterns and avoid overinterpreting null short-run effects as evidence about long-run displacement or welfare.
- Paper claim to check Linked Danish adoption surveys and administrative labor records show widespread chatbot adoption and task reorganization but precise near-term null effects on earnings and recorded hours. Source text (unverified type)
Prominence
8.5/10
The scoring model estimated prominence from the paper's venue, authors, institutional setting, and visibility. The model did not supply a criterion-specific explanation. Current public-attention status: Widely circulated research; documented decision use not confirmed; quick check.
- Author page The first author lists the paper and its earlier versions. Anders Humlum
Likely influence
7.5/10
The scoring model estimated how far the findings could shape later research or decisions, without supplying a criterion-specific explanation. Current public-attention status: Widely circulated research; documented decision use not confirmed; quick check. See public-attention evidence below.
- Author page The first author lists the paper and its earlier versions. Anders Humlum
What the paper says
Source text (unverified type) · Text supplied by the discovery source; its status as a formal abstract has not been verified.
Linked adoption surveys and Danish administrative labor records show widespread chatbot adoption and task reorganization but precise near-term null effects on earnings and recorded hours. The current title replaces the paper's earlier title, Large Language Models, Small Labor Market Effects.
Claims to check
- Linked Danish adoption surveys and administrative labor records show widespread chatbot adoption and task reorganization but precise near-term null effects on earnings and recorded hours.
- Employer AI initiatives, training, and task reorganization may change the structure of work before effects appear in wages, hours, or aggregate employment.
- Adopters appear to move toward higher-paying occupations where AI chatbots are more relevant, but these transitions are still too small to shift average earnings.
Methodological or theoretical issues flagged for evaluation
The core empirical setting is Denmark, with strong administrative data but limited external validity for lower-income countries, informal labor markets, and sectors such as agriculture, education, and health. Evaluation should separate causal near-term wage/hour evidence from descriptive adoption patterns and avoid overinterpreting null short-run effects as evidence about long-run displacement or welfare.
Opus 5.5 re-evaluation (experimental)
Read with care. This is an experimental re-evaluation by a different model (Claude Opus 5.5, 2026-10-01). Across a calibration sample Opus scored about 15 points lower than GPT-5.5, so 15 points are added to compare it with GPT-5.5 scores. The calibration is a small sample (n=44) and is not human-validated, and the two models disagree about the order of papers within the top group. The AI score at the top of this page is the usual evaluation-priority score shifted by the difference between the adjusted Opus score and the earlier holistic score; the synthesis uses that shifted value, and the unshifted score is shown beside it. On the dashboard you can switch the Opus scores off. Read the methods note.
Opus rationale (AI-generated)
This is an NBER working paper by prominent labour economists (Humlum is well known for his work on automation and labour markets). It is one of the most cited empirical studies of generative AI's real labour-market effects. Linking adoption surveys to Danish administrative records, it reports precise near-term null effects on earnings and hours despite widespread chatbot adoption and task reorganisation. That finding is central to current debates on AI and jobs, and it feeds into how the OECD, the IMF, national labour ministries, central banks and AI-governance funders (e.g., Open Philanthropy) judge how urgently to fund transition programmes and how to forecast productivity pass-through. The work is squarely in Unjournal's wheelhouse, and it is still a working paper, now retitled and reframed ('Still Waters, Rapid Currents'). That suggests an active revision, so evaluator feedback could still shape it. An evaluation would add value on some specific questions: how well the precise nulls stand up to selection in adoption and to the choice of occupations; how much self-reported time savings can be trusted; and whether the shift in framing from 'small effects' to 'rapid currents' is supported by the evidence. Concerns: it is a wealthy-country setting with limited direct LMIC transfer. Many parallel studies have appeared, and the paper may soon go through journal review. The 2023-24 window is also obsolescing fast. Even so, it is a strong candidate at the boundary between monitoring it and prioritising it now.
Dashboard details and provenance
Full AI dashboard scoring rationale
This is a prominent NBER working paper on one of the highest-salience empirical questions in AI policy: whether rapid generative-AI adoption is already changing wages, hours, task content, and worker mobility. It would be useful for labor ministries, OECD/ILO/World Bank analysts, AI-governance funders, and firms designing training or worker-transition policy, and as an NBER working paper revised in March 2026 it still appears to need independent public evaluation rather than only seminar feedback. The main concern is external validity: Danish administrative data are unusually strong, but the welfare and distributional implications for lower-income settings, informal labor markets, agriculture, education, and health are indirect rather than directly estimated here.
AI decision-relevance rationale
The paper informs decisions about whether governments and funders should prioritize worker retraining, AI diffusion subsidies, labor-market monitoring, employment protection, education-system adaptation, or slower/more targeted AI deployment. It is especially relevant to OECD, ILO, World Bank, EU and national labor ministries, Open Philanthropy and other AI-governance funders, and research teams at major AI labs trying to distinguish causal labor-market evidence from exposure forecasts and product-performance studies.
AI timing assessment
The paper is an NBER working paper first issued in May 2025 and revised in March 2026, so independent evaluation could still matter for interpretation, revisions, and policy uptake. Search evidence indicates media, blog, and informal public discussion, but I do not see clear evidence of substantial independent expert replication or a formal critique; this keeps the marginal value of Unjournal evaluation relatively high, though the paper has already received seminar and conference feedback.
Intake, review, and crux connections
AI, global health, and development curation
· 2026-07-02
Manual curation of 25 candidate leads, followed by canonical source-page verification, deduplication, and selective import
This recovers the useful public-paper content of an earlier development-branch curation without merging its obsolete code. Only research records with verified canonical source pages and clear relevance were retained. Papers already present are labeled as surfaced rather than added; blog posts, registry-only entries, unverifiable leads, industry benchmarks, and background papers without a direct AI question were excluded. Inclusion documents a deliberately focused intake, not endorsement or a completed Unjournal team decision.
Public attention and use
The paper has circulated under an earlier title and is widely discussed in the economics literature, but our two searches mainly found bibliographic records and the author's page. No independent critique or policy use was documented by what we saw.
This was a quick check (a few targeted searches). Treat the gaps below as especially uncertain.
Author page listing / discoverability
The first author lists the paper and its earlier versions.
Source: Anders Humlum · Relationship: author-written
Bibliographic index listing / discoverability
The paper is indexed in RePEc.
Source: RePEc · Relationship: bibliographic index
Earlier version listing / discoverability
An earlier version circulated through the Becker Friedman Institute under a different title.
Source: BFI · Relationship: author institution
What the search did not establish
- This was a quick check; coverage under the earlier title was not systematically searched.
- No EA Forum or LessWrong discussion surfaced under the current title.
- The most decision-relevant next signal is whether updated data show effects emerging after the early period.
Quick search checked 2026-10-01. Search scope: Targeted exact-title searches across the open web, institutional and author pages, news and public social-media results, EA Forum/LessWrong, and policy/white-paper contexts. Evidence records distinguish commissioning or report use from independent discussion, media attention, indexing, and post-publication policy use. A search miss is reported as uncertainty, not proof of absence. Entries with check_depth "quick" rest on roughly one to two searches and should be read as especially uncertain; entries checked on 2026-10-01 were added for the shortlist of papers without human ratings.