The Unjournal · Research prioritization
Paper-specific consideration page · Top other papers, rank 10

Targeting, Personalization, and Engagement in an Agricultural Advisory Service

Academic and practitioner context; quick check
Why this page exists. This paper is on a shortlist of papers we are considering for independent evaluation. No human ratings have been submitted yet, so the score below comes from the AI scoring pass alone. Treat it as one provisional input; the team has made no prioritization decision.
73AI evaluation-priority (shifted by the Opus 5.5 re-evaluation)
81Original AI lens before the Opus shift (not used for the synthesis)
None yetHuman aggregate · no ratings submitted
Not availableHuman–AI synthesis needs at least one human rating

Why this paper is being considered

The paper reports iterative experiments on targeting calls in an agricultural advisory service serving over one million farmers in India. Policies designed with current data gave much smaller on-policy gains than off-policy evaluation suggested (up to 8 percent). It bears on how digital advisory services should target and personalise at scale. An evaluation would focus on the on-policy versus off-policy gap, the engagement outcome, and whether engagement translates into farm outcomes.

AI-generated criterion ratings and reasoning

Scores come from AI prioritization; the accompanying explanations include AI-assisted source checks. These are provisional judgments, separate from human ratings and commissioned evaluations.

Decision relevance

8.0/10

Agricultural advisory/extension services (e.g., Precision Agriculture for Development-style digital advisory, Digital Green, government extension programs) face a concrete allocation problem: contact capacity is far smaller than the farmer base, so targeting and personalization rules directly determine cost-effectiveness. Evidence on which targeted-calling policies raise engagement — and ideally downstream adoption — informs how funders and implementers spend limited outreach budgets across millions of farmers. Off-policy evaluation methods that let programs value new policies without running every arm as a fresh RCT are themselves decision-relevant, since they promise cheaper iteration for any org running a large advisory service.

  • Paper claim to check Iterative randomized-call experiments plus off-policy evaluation can estimate the value of alternative targeting/personalization policies for an advisory service serving over one million Indian farmers. Source abstract

Value of added scrutiny

8.0/10

The paper is hosted by Harvard Kennedy School and Harvard Business School pages, and the broader program has generated practitioner writing. We found no independent critique of this paper.

Timing

9.0/10

Newly posted NBER working paper (w34951), not yet peer-reviewed, so this is the maximum-value window: independent evaluation of the identification and off-policy evaluation assumptions could inform revisions before publication, and the methods are directly reusable by implementers now.

  • Source record Publication-stage and date evidence should be checked in the linked paper record. NBER

Methodological potential

8.0/10

No methodological rationale is stored.

  • Paper claim to check Iterative randomized-call experiments plus off-policy evaluation can estimate the value of alternative targeting/personalization policies for an advisory service serving over one million Indian farmers. Source abstract

Prominence

9.0/10

The scoring model estimated prominence from the paper's venue, authors, institutional setting, and visibility. The model did not supply a criterion-specific explanation. Current public-attention status: Academic and practitioner context; quick check.

Likely influence

7.0/10

The scoring model estimated how far the findings could shape later research or decisions, without supplying a criterion-specific explanation. Current public-attention status: Academic and practitioner context; quick check. See public-attention evidence below.

What the paper says

Source abstract · Full source abstract recovered from the NBER paper page; HTML entities and whitespace normalized.

We conduct a series of iterative experiments to evaluate approaches for optimally targeting calls for an agricultural advisory service serving over one million farmers in rural India. We estimate the value of alternative targeted policies using “off-policy” evaluation on data from randomized call assignments. When we evaluate off-policy using held-out data from the same time periods used to design the targeted policies, we find that targeted policies increase engagement by up to 8%. However, when we design a policy using current data and then implement the policy for randomly selected users in subsequent weeks, our “on-policy” evaluation estimates show realized gains that are substantially smaller. We develop tools to diagnose the causes of this underperformance and adopt a transfer-learning approach to policy learning and off-policy evaluation that accounts for temporal changes in farmer behavior, substantially improving off-policy estimates of performance in subsequent weeks. We further develop novel approaches to targeted policy design that respond to organizational objectives and resource constraints, including distributional goals (e.g., reaching women farmers) and prioritizing farmers likely to benefit most from the information on downstream outcomes (e.g., yields). We propose a method for organizations to quantify trade-offs between objectives, and we demonstrate the value of targeting for alleviating these trade-offs.

Claims to check

  • Iterative randomized-call experiments plus off-policy evaluation can estimate the value of alternative targeting/personalization policies for an advisory service serving over one million Indian farmers.
  • Optimally targeted and personalized calling policies increase farmer engagement relative to untargeted or default policies.
  • Off-policy (counterfactual) evaluation on RCT call data yields credible value estimates for policies that were not themselves fully deployed.
Methodological or theoretical issues flagged for evaluation

No structured evaluation-challenges field is stored.

Opus 5.5 re-evaluation (experimental)

Opus raw score: 66/100
Opus own action: watchlist
Label from adjusted score: watchlist
Earlier score (opus): 74/100
Shift applied to the AI score: -8

Read with care. This is an experimental re-evaluation by a different model (Claude Opus 5.5, 2026-10-01). Across a calibration sample Opus scored about 15 points lower than GPT-5.5, so 15 points are added to compare it with GPT-5.5 scores. The calibration is a small sample (n=44) and is not human-validated, and the two models disagree about the order of papers within the top group. The AI score at the top of this page is the usual evaluation-priority score shifted by the difference between the adjusted Opus score and the earlier holistic score; the synthesis uses that shifted value, and the unshifted score is shown beside it. On the dashboard you can switch the Opus scores off. Read the methods note.

Different reference score. The earlier score for this paper came from an older Opus run (opus), not GPT-5.5, so no +15 offset is applied: the shift is the Opus 5.5 score minus that earlier score, which assumes the two Opus versions share a scale.

Opus rationale (AI-generated)

This NBER working paper by Susan Athey, Shawn Cole and Shanjukta Nath is squarely in Unjournal's development-economics wheelhouse. It is a large-scale, iterative field experiment with an Indian agricultural advisory service reaching over a million farmers. Its most decision-relevant finding is a caution: ML targeting that looks good under standard off-policy evaluation (up to +8% engagement) delivered much smaller gains when deployed in later weeks. The authors propose a transfer-learning fix and a framework for trading off engagement, reaching women farmers, and expected yield benefits. The evidence bears on decisions by digital extension providers such as Precision Development and government advisory programmes, and by funders of digital agriculture (Gates Foundation, FCDO, World Bank, J-PAL/ATAI). More broadly it matters to any LMIC programme using ML to target nudges or benefits. As a recent NBER working paper with no apparent independent review, an evaluation could check whether the off-policy vs on-policy comparison is well powered and properly compared, whether the transfer-learning correction is validated out of sample rather than tuned after the fact, and whether engagement gains translate into welfare. Concerns: the main outcome is engagement, not income or yields, and the operational gains are modest. Much of the value is methodological, so an evaluator needs both causal ML/policy-learning expertise and agricultural-extension knowledge. I'd rate this a solid monitor-to-borderline-prioritize candidate rather than a clear top priority.

Dashboard details and provenance

Discovery source: NBER
Release date: March 2026
Scoring model: opus (headless)
Model holistic score: 74

Full AI dashboard scoring rationale

This is a high-prominence NBER working paper by Susan Athey (a leading figure in causal ML and off-policy evaluation), Shawn Cole (a well-known development/finance economist), and Shanjukta Nath, studying how to optimally target and personalize an agricultural advisory service reaching over one million smallholder farmers in rural India. It sits squarely in Unjournal's strongest cause area (development economics/LMICs) and speaks directly to a live operational question for agricultural extension programs: given scarce contact budgets, whom should you call, when, and with what content to maximize engagement and downstream farmer outcomes. The methodological contribution — iterative field experiments combined with off-policy (counterfactual policy) evaluation to value alternative targeting rules — is exactly the kind of ML-for-policy design where an independent evaluation adds value, because off-policy evaluation is notoriously sensitive to assumptions (propensity estimation, overlap/positivity, variance of importance weights) that referees at a general venue may not scrutinize closely. It is a working paper with no peer review yet, so feedback is timely and actionable. The main caution is honest calibration: the abstract emphasizes engagement/uptake of calls rather than final welfare outcomes (yields, income, farmer well-being), so decision-relevance hinges on how well engagement proxies for real impact — worth flagging but not disqualifying.

AI decision-relevance rationale

Agricultural advisory/extension services (e.g., Precision Agriculture for Development-style digital advisory, Digital Green, government extension programs) face a concrete allocation problem: contact capacity is far smaller than the farmer base, so targeting and personalization rules directly determine cost-effectiveness. Evidence on which targeted-calling policies raise engagement — and ideally downstream adoption — informs how funders and implementers spend limited outreach budgets across millions of farmers. Off-policy evaluation methods that let programs value new policies without running every arm as a fresh RCT are themselves decision-relevant, since they promise cheaper iteration for any org running a large advisory service.

AI timing assessment

Newly posted NBER working paper (w34951), not yet peer-reviewed, so this is the maximum-value window: independent evaluation of the identification and off-policy evaluation assumptions could inform revisions before publication, and the methods are directly reusable by implementers now.

Public attention and use

The paper is hosted by Harvard Kennedy School and Harvard Business School pages, and the broader program has generated practitioner writing. We found no independent critique of this paper.

This was a quick check (a few targeted searches). Treat the gaps below as especially uncertain.

Author-institution listing listing / discoverability

Harvard Kennedy School CID publication page

The Center for International Development lists the paper.

Source: Harvard Kennedy School · Relationship: author institution
Related practitioner context related context, not uptake

Precision Development: customized digital advice summary

A practitioner summary of related randomized evidence on digital advisory services in India. It is context, not use of this paper.

Source: Precision Development · Relationship: research network

What the search did not establish

  • This was a quick check; treat the gaps as especially uncertain.
  • No news, forum or direct policy use surfaced.
  • The most decision-relevant next signal is whether advisory providers adopt the targeting methods.

Quick search checked 2026-10-01. Search scope: Targeted exact-title searches across the open web, institutional and author pages, news and public social-media results, EA Forum/LessWrong, and policy/white-paper contexts. Evidence records distinguish commissioning or report use from independent discussion, media attention, indexing, and post-publication policy use. A search miss is reported as uncertainty, not proof of absence. Entries with check_depth "quick" rest on roughly one to two searches and should be read as especially uncertain; entries checked on 2026-10-01 were added for the shortlist of papers without human ratings.

Human feedback so far

No human ratings have been submitted for this paper yet, so there is no human aggregate or synthesis score. If you know the work, rate or discuss it on the dashboard card; the team reviews feedback before it affects prioritization.