EconBase
← All papers

Post Reinforcement Learning Inference

Vasilis Syrgkanis, Ruohan Zhan

arXiv 17 Feb 2023 · Statistics — Machine Learning · publishedOperations Research (2025) · 1 citations (OpenAlex)

arXiv:2302.08854 · PDF · DOI · OpenAlex · Extracted main text

Abstract

We study estimation and inference using data collected by reinforcement learning (RL) algorithms. These algorithms adaptively experiment by interacting with individual units over multiple stages, updating their strategies based on past outcomes. Our goal is to evaluate a counterfactual policy after data collection and estimate structural parameters, such as dynamic treatment effects, that support credit assignment and quantify the impact of early actions on final outcomes. These parameters can often be defined as solutions to moment equations, motivating moment-based estimation methods developed for static data. In RL settings, however, data are often collected adaptively under nonstationary behavior policies. As a result, standard estimators fail to achieve asymptotic normality due to time-varying variance. We propose a weighted generalized method of moments (GMM) approach that uses adaptive weights to stabilize this variance. We characterize weighting schemes that ensure consistency and asymptotic normality of the weighted GMM estimators, enabling valid hypothesis testing and uniform confidence region construction. Key applications include dynamic treatment effect estimation and dynamic off-policy evaluation.

Citation extraction

39
references
81
in-text mentions
39
distinct cited
0
self-citations
35,410
main-text words

appendix boundary found by appendix_titled_section at “Supplementary Results for High-dimensional Markovian Models” · 74% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Robins JM (2004) Optimal structural nested models for optimal sequential decisions1.00093100%
2Lewis G, Syrgkanis V (2020) Double/debiased machine learning for dynamic treatment effects via g-estimation1.00075100%
3Hadad V, Hirshberg DA, Zhan R, Wager S, Athey S (2021) Confidence intervals for policy evaluation in adaptive experiments1.00074100%
4Zhang K, Janson L, Murphy S (2021) Statistical inference with m-estimators on adaptively collected data1.00053100%
5Zhan R, Hadad V, Hirshberg DA, Athey S (2021) Off-policy evaluation via adaptive weighting with data from contextual bandits0.92843100%
6Deshpande Y, Mackey L, Syrgkanis V, Taddy M (2018) Accurate inference for adaptive linear models0.81142100%
7Bibaut A, Dimakopoulou M, Kallus N, Chambaz A, van Der Laan M (2021) Post-contextual-bandit inference0.73732100%
8Cattaneo MD, Masini RP, Underwood WG (2022) Yurinskii's coupling for martingales0.73732100%
9Chen M, Beutel A, Covington P, Jain S, Belletti F, Chi EH (2019) Top-k off-policy correction for a reinforce recommender system0.73732100%
10Chakraborty B, Moodie EE, Chakraborty B, Moodie EE (2013) Semi-parametric estimation of optimal dtrs by modeling contrasts of conditional mean outcomes0.64422100%

Showing the top 10 of 39 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Estimating Causal Effects from Data Generated by Stochastic Algorithms0.40511