Yue Fang, Geert Ridder, Haitian Xie
arXiv 22 Dec 2025 · Econometrics
arXiv:2512.19230 · PDF · DOI · OpenAlex · Extracted main text
Recent literature on policy learning has primarily focused on regret bounds of the learned policy. We provide a new perspective by developing a unified semiparametric efficiency framework for policy learning, allowing for general treatments that are discrete, continuous, or mixed. We provide a characterization of the failure of pathwise differentiability for parameters arising from deterministic policies. We then establish efficiency bounds for pathwise differentiable parameters in randomized policies, both when the propensity score is known and when it must be estimated. Building on the convolution theorem, we introduce a notion of efficiency for the asymptotic distribution of welfare regret, showing that inefficient policy estimators not only inflate the variance of the asymptotic regret but also shift its mean upward. We derive the asymptotic theory of several common policy estimators, with a key contribution being a policy-learning analogue of the Hirano-Imbens-Ridder (HIR) phenomenon: the inverse propensity weighting estimator with an estimated propensity is efficient, whereas the same estimator using the true propensity is not. We illustrate the theoretical results with an empirically calibrated simulation study based on data from a job training program and an empirical application to a commitment savings program.
appendix boundary found by appendix_command · 41% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Kitagawa, Toru and Tetenov, Aleksey (2018) Who should be treated? Empirical welfare maximization methods for treatment choice | 1.000 | 9 | 4 | 100% |
| 2 | Crippa, Federico (2025) Regret analysis in threshold policy design | 1.000 | 5 | 3 | 100% |
| 3 | Athey, Susan and Wager, Stefan (2021) Policy learning with observational data | 0.928 | 4 | 3 | 100% |
| 4 | Mbakop, Eric and Tabord-Meehan, Max (2021) Model selection for treatment choice: Penalized welfare maximization | 0.928 | 4 | 3 | 100% |
| 5 | Hirano, Keisuke and Imbens, Guido W and Ridder, Geert (2003) Efficient estimation of average treatment effects using the estimated propensity score self | 0.811 | 4 | 2 | 100% |
| 6 | Ai, Chunrong and Fang, Yue and Xie, Haitian (2024) Data-driven Policy Learning for Continuous Treatments self | 0.737 | 3 | 3 | 67% |
| 7 | Xiaohong Chen and Han Hong and Alessandro Tarozzi (2008) Semiparametric efficiency in GMM models with auxiliary data | 0.737 | 3 | 2 | 100% |
| 8 | Ai, Chunrong and Linton, Oliver and Motegi, Kaiji and Zhang, Zheng (2021) A unified framework for efficient estimation of general treatment models | 0.721 | 16 | 4 | 38% |
| 9 | Ashraf, Nava and Karlan, Dean and Yin, Wesley (2006) Tying Odysseus to the mast: Evidence from a commitment savings product in the Philippines | 0.644 | 2 | 2 | 100% |
| 10 | Yue Fang and Jin Xi and Haitian Xie (2025) Model Selection for Multivalued-Treatment Policy Learning in Observational Studies self | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 45 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | 2606.01659 | 0.405 | 1 | 1 |