Chunrong Ai, Zeqi Wu, Zheng Zhang
arXiv 1 Jun 2026 · Econometrics
arXiv:2606.01659 · PDF · DOI · OpenAlex · Extracted main text
This paper explores policy learning from observational data, focusing on a nonlinear welfare criterion in a binary treatment setting. The nonlinear criterion is inspired by scenarios where policymakers prioritize specific population segments. We model this criterion using a utility function that encompasses potential outcomes and intermediate parameters, with the latter capturing higher moments of the outcome distributions. When formulated in the context of observational data, both the intermediate parameters and the welfare criterion depend on the propensity score, which we estimate using machine-learning techniques. To address bias in machine learning estimates, we introduce a novel reweighting-based debiasing approach that offers a promising alternative to traditional orthogonality-based methods. To tackle the complexities of infinite-dimensional policy spaces, we employ sieve approximations and $K$-fold cross-validation for model selection, thereby fully automating the policy-learning process. Despite these complexities, we demonstrate that both the welfare regret and the average welfare regret of our proposed policy learning method satisfy an oracle inequality, thereby providing theoretical guarantees on the performance of the estimated policy relative to the best possible policy. This finding extends the existing results from linear to nonlinear welfare criteria, from finite-dimensional to infinite-dimensional policy spaces, and from a known propensity score to a machine-learned one.
appendix boundary found by appendix_command · 24% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Athey, Susan and Wager, Stefan (2021) Policy Learning With Observational Data | 1.000 | 5 | 3 | 100% |
| 2 | Mbakop, Eric and Tabord-Meehan, Max (2021) Model Selection for Treatment Choice: Penalized Welfare Maximization | 0.937 | 17 | 5 | 82% |
| 3 | Kitagawa, Toru and Tetenov, Aleksey (2018) Who Should Be Treated? Empirical Welfare Maximization Methods for Treatment Choice | 0.928 | 4 | 4 | 100% |
| 4 | Crippa, Federico (2025) Regret Analysis in Threshold Policy Design | 0.644 | 2 | 2 | 100% |
| 5 | Liu, Nan and Liu, Yanbo and Sasaki, Yuya and Wan, Yuanyuan (2025) Nonparametric Uniform Inference in Binary Classification and Policy Values | 0.644 | 2 | 2 | 100% |
| 6 | Boyd, Stephen P. and Vandenberghe, Lieven (2004) Convex Optimization | 0.511 | 2 | 2 | 50% |
| 7 | Fan, Yanqin and Qi, Yuan and Xu, Gaoqian (2025) Policy Learning with $ $-Expected Welfare | 0.511 | 2 | 1 | 100% |
| 8 | Farrell, Max H. and Liang, Tengyuan and Misra, Sanjog (2021) Deep Neural Networks for Estimation and Inference | 0.511 | 2 | 1 | 100% |
| 9 | Jiao, Yuling and Shen, Guohao and Lin, Yuanyuan and Huang, Jian (2023) Deep Nonparametric Regression on Approximate Manifolds: Nonasymptotic Error Bounds with Polynomial Prefactors | 0.511 | 2 | 1 | 100% |
| 10 | Lecué, Guillaume and Mitchell, Charles (2012) Oracle Inequalities for Cross-Validation Type Procedures | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 63 scored citations.