Xiaohong Chen, Zhengling Qi
arXiv 17 Jan 2022 · Mathematics — Statistics Theory · 5 citations (OpenAlex)
arXiv:2201.06169 · PDF · DOI · OpenAlex · Extracted main text
We study the off-policy evaluation (OPE) problem in an infinite-horizon Markov decision process with continuous states and actions. We recast the $Q$-function estimation into a special form of the nonparametric instrumental variables (NPIV) estimation problem. We first show that under one mild condition the NPIV formulation of $Q$-function estimation is well-posed in the sense of $L^2$-measure of ill-posedness with respect to the data generating distribution, bypassing a strong assumption on the discount factor $\gamma$ imposed in the recent literature for obtaining the $L^2$ convergence rates of various $Q$-function estimators. Thanks to this new well-posed property, we derive the first minimax lower bounds for the convergence rates of nonparametric estimation of $Q$-function and its derivatives in both sup-norm and $L^2$-norm, which are shown to be the same as those for the classical nonparametric regression (Stone, 1982). We then propose a sieve two-stage least squares estimator and establish its rate-optimality in both norms under some mild conditions. Our general results on the well-posedness and the minimax lower bounds are of independent interest to study not only other nonparametric estimators for $Q$-function but also efficient estimation on the value of any target policy in off-policy settings.
appendix boundary found by appendix_command · 57% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Blundell, Chen \ Kristensen (2007) `Semi-nonparametric iv estimation of shape-invariant engel curves', Econometrica 75(6), 1613–1669 | 1.000 | 7 | 3 | 100% |
| 2 | Farahmand, Ghavamzadeh, Szepesvári \ Mannor (2016) `Regularized policy iteration with nonparametric function spaces', The Journal of Machine Learning Research 17(1), 4809–4874 | 1.000 | 6 | 3 | 100% |
| 3 | Kallus \ Uehara (2019) `Efficiently breaking the curse of horizon: Double reinforcement learning in infinite-horizon processes', arXiv preprint arXiv:1… | 1.000 | 5 | 3 | 100% |
| 4 | Stone (1982) `Optimal global rates of convergence for nonparametric regression', The Annals of Statistics pp. 1040–1053 | 0.928 | 4 | 4 | 100% |
| 5 | Shi, Zhang, Lu \ Song (2020) `Statistical inference of the value function for reinforcement learning in infinite horizon settings', Journal of the Royal Stat… | 0.874 | 11 | 2 | 100% |
| 6 | Uehara, Imaizumi, Jiang, Kallus, Sun \ Xie (2021) `Finite sample analysis of minimax offline reinforcement learning: Completeness, fast rates and first-order efficiency', arXiv p… | 0.874 | 5 | 2 | 100% |
| 7 | Jin, Yang \ Wang (2021) Is pessimism provably efficient for offline rl?, in `International Conference on Machine Learning', PMLR, pp. 5084–5096 | 0.843 | 3 | 3 | 100% |
| 8 | Chen \ Christensen (2018) `Optimal sup-norm rates and uniform inference on nonlinear functionals of nonparametric IV regression', Quantitative Economics 9… | 0.759 | 16 | 9 | 44% |
| 9 | Chen \ Christensen (2015) `Optimal uniform convergence rates and asymptotic normality for series estimators under weak dependence and weak conditions', Jo… | 0.754 | 7 | 5 | 43% |
| 10 | Huang et al (1998) `Projection estimation in multiple regression with application to functional anova models', Annals of Statistics 26(1), 242–272 | 0.737 | 3 | 3 | 67% |
Showing the top 10 of 63 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Inference on Time Series Nonparametric Conditional Moment Restrictions Using General Sieves | 0.737 | 3 | 2 |
| 2 | Welfare Analysis in Dynamic Models | 0.644 | 2 | 2 |
| 3 | Adaptive Estimation and Uniform Confidence Bands for Nonparametric Structural Functions and Elasticities | 0.405 | 1 | 1 |