EconBase
← All papers

Inference on Optimal Policy Values and Other Irregular Functionals via Smoothing

Justin Whitehouse, Morgane Austern, Vasilis Syrgkanis

arXiv 15 Jul 2025 · Econometrics

arXiv:2507.11780 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Constructing confidence intervals for the value of an optimal treatment policy is an important problem in causal inference. Insight into the optimal policy value can guide the development of reward-maximizing, individualized treatment regimes. However, because the functional that defines the optimal value is non-differentiable, standard semi-parametric approaches for performing inference fail to be directly applicable. Existing approaches for handling this non-differentiability fall roughly into two camps. In one camp are estimators based on constructing smooth approximations of the optimal value. These approaches are computationally lightweight, but typically place unrealistic parametric assumptions on outcome regressions. In another camp are approaches that directly de-bias the non-smooth objective. These approaches don't place parametric assumptions on nuisance functions, but they either require the computation of intractably-many nuisance estimates, assume unrealistic $L^\infty$ nuisance convergence rates, or make strong margin assumptions that prohibit non-response to a treatment. In this paper, we revisit the problem of constructing smooth approximations of non-differentiable functionals. By carefully controlling first-order bias and second-order remainders, we show that a softmax smoothing-based estimator can be used to estimate parameters that are specified as a maximum of scores involving nuisance components. In particular, this includes the value of the optimal treatment policy as a special case. Our estimator obtains $\sqrt{n}$ convergence rates, avoids parametric restrictions/unrealistic margin assumptions, and is often statistically efficient.

Citation extraction

83
references
152
in-text mentions
83
distinct cited
6
self-citations
16,265
main-text words

appendix boundary found by appendix_command · 32% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Luedtke, Alexander R and Van Der Laan, Mark J (2016) Statistical inference for the mean outcome under a possibly non-unique optimal treatment strategy1.000124100%
2Levis, Alexander W and Bonvini, Matteo and Zeng, Zhenghao and Keele,… (2023) Covariate-assisted bounds on causal effects with instrumental variables1.000103100%
3Shi, Chengchun and Lu, Wenbin and Song, Rui (2020) Breaking the Curse of Nonregularity with Subagging–-Inference of the Mean Outcome under Optimal Treatment Regimes1.00084100%
4Chernozhukov, Victor and Chetverikov, Denis and Demirer, Mert and Du… (2018) Double/debiased machine learning for treatment and structural parameters1.00064100%
5Gupta, Chirag (2022) Post-hoc calibration without distributional assumptions1.00063100%
6Laber, Eric B and Lizotte, Daniel J and Qian, Min and Pelham, Willia… (2014) Dynamic treatment regimes: Technical challenges and applications0.92843100%
7Semenova, Vira (2023) Aggregated Intersection Bounds and Aggregated Minimax Values0.92843100%
8Hirano, Keisuke and Porter, Jack R (2012) Impossibility results for nondifferentiable functionals0.84333100%
9Goldberg, Yair and Song, Rui and Zeng, Donglin and Kosorok, Michael R (2014) Comment on “Dynamic treatment regimes: Technical challenges and applications”0.81142100%
10Van Der Laan, Mark J and Rubin, Daniel (2006) Targeted maximum likelihood learning0.73732100%

Showing the top 10 of 83 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Semiparametric Off-Policy Inference for Optimal Policy Values under Possible Non-Uniqueness0.92853
2Nonparametric Bayesian Policy Learning0.64422
3Inference on Welfare and Value Functionals under Optimal Treatment Assignment0.51121
4On the Lower Confidence Band for the Optimal Welfare in Policy Learning0.40511
5Policy Learning with Abstention0.40511
6Adaptive Estimation of Aggregated Values of Conditional Linear Programs0.40511