Vasilis Syrgkanis, Ruohan Zhan
arXiv 17 Feb 2023 · Statistics — Machine Learning · publishedOperations Research (2025) · 1 citations (OpenAlex)
arXiv:2302.08854 · PDF · DOI · OpenAlex · Extracted main text
We study estimation and inference using data collected by reinforcement learning (RL) algorithms. These algorithms adaptively experiment by interacting with individual units over multiple stages, updating their strategies based on past outcomes. Our goal is to evaluate a counterfactual policy after data collection and estimate structural parameters, such as dynamic treatment effects, that support credit assignment and quantify the impact of early actions on final outcomes. These parameters can often be defined as solutions to moment equations, motivating moment-based estimation methods developed for static data. In RL settings, however, data are often collected adaptively under nonstationary behavior policies. As a result, standard estimators fail to achieve asymptotic normality due to time-varying variance. We propose a weighted generalized method of moments (GMM) approach that uses adaptive weights to stabilize this variance. We characterize weighting schemes that ensure consistency and asymptotic normality of the weighted GMM estimators, enabling valid hypothesis testing and uniform confidence region construction. Key applications include dynamic treatment effect estimation and dynamic off-policy evaluation.
appendix boundary found by appendix_titled_section at “Supplementary Results for High-dimensional Markovian Models” · 74% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Robins JM (2004) Optimal structural nested models for optimal sequential decisions | 1.000 | 9 | 3 | 100% |
| 2 | Lewis G, Syrgkanis V (2020) Double/debiased machine learning for dynamic treatment effects via g-estimation | 1.000 | 7 | 5 | 100% |
| 3 | Hadad V, Hirshberg DA, Zhan R, Wager S, Athey S (2021) Confidence intervals for policy evaluation in adaptive experiments | 1.000 | 7 | 4 | 100% |
| 4 | Zhang K, Janson L, Murphy S (2021) Statistical inference with m-estimators on adaptively collected data | 1.000 | 5 | 3 | 100% |
| 5 | Zhan R, Hadad V, Hirshberg DA, Athey S (2021) Off-policy evaluation via adaptive weighting with data from contextual bandits | 0.928 | 4 | 3 | 100% |
| 6 | Deshpande Y, Mackey L, Syrgkanis V, Taddy M (2018) Accurate inference for adaptive linear models | 0.811 | 4 | 2 | 100% |
| 7 | Bibaut A, Dimakopoulou M, Kallus N, Chambaz A, van Der Laan M (2021) Post-contextual-bandit inference | 0.737 | 3 | 2 | 100% |
| 8 | Cattaneo MD, Masini RP, Underwood WG (2022) Yurinskii's coupling for martingales | 0.737 | 3 | 2 | 100% |
| 9 | Chen M, Beutel A, Covington P, Jain S, Belletti F, Chi EH (2019) Top-k off-policy correction for a reinforce recommender system | 0.737 | 3 | 2 | 100% |
| 10 | Chakraborty B, Moodie EE, Chakraborty B, Moodie EE (2013) Semi-parametric estimation of optimal dtrs by modeling contrasts of conditional mean outcomes | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 39 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Estimating Causal Effects from Data Generated by Stochastic Algorithms | 0.405 | 1 | 1 |