EconBase
← All papers

Characterization of Efficient Influence Function for Off-Policy Evaluation Under Optimal Policies

Haoyu Wei

arXiv 20 May 2025 · Mathematics — Statistics Theory

arXiv:2505.13809 · PDF · Extracted main text

Abstract

Off-policy evaluation (OPE) provides a powerful framework for estimating the value of a counterfactual policy using observational data, without the need for additional experimentation. Despite recent progress in robust and efficient OPE across various settings, rigorous efficiency analysis of OPE under an estimated optimal policy remains limited. In this paper, we establish a concise characterization of the efficient influence function (EIF) for the value function under optimal policy within canonical Markov decision process models. Specifically, we provide the sufficient conditions for the existence of the EIF and characterize its expression. We also give the conditions under which the EIF does not exist.

Citation extraction

41
references
78
in-text mentions
41
distinct cited
0
self-citations
10,826
main-text words

appendix boundary found by appendix_command · 30% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Shi, Chengchun and Zhang, Sheng and Lu, Wenbin and Song, Rui (2022) Statistical inference of the value function for reinforcement learning in infinite-horizon settings0.96118889%
2Whitehouse, Justin and Austern, Morgane and Syrgkanis, Vasilis (2025) Inference on Optimal Policy Values and Other Irregular Functionals via Smoothing0.9285380%
3Uehara, Masatoshi and Shi, Chengchun and Kallus, Nathan (2022) A review of off-policy evaluation in reinforcement learning0.92844100%
4Shi, Chengchun (2025) Statistical Inference in Reinforcement Learning: A Selective Survey0.7373367%
5Athey, Susan and Wager, Stefan (2021) Policy learning with observational data0.64422100%
6Kosorok, Michael R and Laber, Eric B (2019) Precision medicine0.64422100%
7Laber, Eric B and Lizotte, Daniel J and Qian, Min and Pelham, Willia… (2014) Dynamic treatment regimes: Technical challenges and applications0.64422100%
8Neu, Gergely and Jonsson, Anders and Gómez, Vicen c (2017) A unified view of entropy-regularized markov decision processes0.64422100%
9Shi, Chengchun and Wan, Runzhe and Chernozhukov, Victor and Song, Rui (2021) Deeply-debiased off-policy interval estimation0.64422100%
10Shi, Chengchun and Zhu, Jin and Shen, Ye and Luo, Shikai and Zhu, Ho… (2024) Off-policy confidence interval estimation with confounded markov decision process0.64422100%

Showing the top 10 of 41 scored citations.