EconBase
← All papers

Semiparametric Efficiency in Policy Learning with General Treatments

Yue Fang, Geert Ridder, Haitian Xie

arXiv 22 Dec 2025 · Econometrics

arXiv:2512.19230 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Recent literature on policy learning has primarily focused on regret bounds of the learned policy. We provide a new perspective by developing a unified semiparametric efficiency framework for policy learning, allowing for general treatments that are discrete, continuous, or mixed. We provide a characterization of the failure of pathwise differentiability for parameters arising from deterministic policies. We then establish efficiency bounds for pathwise differentiable parameters in randomized policies, both when the propensity score is known and when it must be estimated. Building on the convolution theorem, we introduce a notion of efficiency for the asymptotic distribution of welfare regret, showing that inefficient policy estimators not only inflate the variance of the asymptotic regret but also shift its mean upward. We derive the asymptotic theory of several common policy estimators, with a key contribution being a policy-learning analogue of the Hirano-Imbens-Ridder (HIR) phenomenon: the inverse propensity weighting estimator with an estimated propensity is efficient, whereas the same estimator using the true propensity is not. We illustrate the theoretical results with an empirically calibrated simulation study based on data from a job training program and an empirical application to a commitment savings program.

Citation extraction

45
references
97
in-text mentions
45
distinct cited
4
self-citations
10,671
main-text words

appendix boundary found by appendix_command · 41% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Kitagawa, Toru and Tetenov, Aleksey (2018) Who should be treated? Empirical welfare maximization methods for treatment choice1.00094100%
2Crippa, Federico (2025) Regret analysis in threshold policy design1.00053100%
3Athey, Susan and Wager, Stefan (2021) Policy learning with observational data0.92843100%
4Mbakop, Eric and Tabord-Meehan, Max (2021) Model selection for treatment choice: Penalized welfare maximization0.92843100%
5Hirano, Keisuke and Imbens, Guido W and Ridder, Geert (2003) Efficient estimation of average treatment effects using the estimated propensity score self0.81142100%
6Ai, Chunrong and Fang, Yue and Xie, Haitian (2024) Data-driven Policy Learning for Continuous Treatments self0.7373367%
7Xiaohong Chen and Han Hong and Alessandro Tarozzi (2008) Semiparametric efficiency in GMM models with auxiliary data0.73732100%
8Ai, Chunrong and Linton, Oliver and Motegi, Kaiji and Zhang, Zheng (2021) A unified framework for efficient estimation of general treatment models0.72116438%
9Ashraf, Nava and Karlan, Dean and Yin, Wesley (2006) Tying Odysseus to the mast: Evidence from a commitment savings product in the Philippines0.64422100%
10Yue Fang and Jin Xi and Haitian Xie (2025) Model Selection for Multivalued-Treatment Policy Learning in Observational Studies self0.64422100%

Showing the top 10 of 45 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
12606.016590.40511