Vivek F. Farias, Andrew A. Li, Tianyi Peng, Andrew Zheng
arXiv 6 Jun 2022 · Machine Learning · 18 citations (OpenAlex)
arXiv:2206.02371 · PDF · DOI · OpenAlex · Extracted main text
We consider experiments in dynamical systems where interventions on some experimental units impact other units through a limiting constraint (such as a limited inventory). Despite outsize practical importance, the best estimators for this `Markovian' interference problem are largely heuristic in nature, and their bias is not well understood. We formalize the problem of inference in such experiments as one of policy evaluation. Off-policy estimators, while unbiased, apparently incur a large penalty in variance relative to state-of-the-art heuristics. We introduce an on-policy estimator: the Differences-In-Q's (DQ) estimator. We show that the DQ estimator can in general have exponentially smaller variance than off-policy evaluation. At the same time, its bias is second order in the impact of the intervention. This yields a striking bias-variance tradeoff so that the DQ estimator effectively dominates state-of-the-art alternatives. From a theoretical perspective, we introduce three separate novel techniques that are of independent interest in the theory of Reinforcement Learning (RL). Our empirical evaluation includes a set of experiments on a city-scale ride-hailing simulator.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | J. L. Doob (1935) The limiting distributions of certain statistics | 0.405 | 1 | 1 | 100% |
| 2 | P. E. Greenwood and W. Wefelmeyer (1995) Efficiency of empirical estimators for markov chains | 0.405 | 1 | 1 | 100% |
| 3 | G. L. Jones (2004) On the markov chain central limit theorem | 0.405 | 1 | 1 | 100% |
| 4 | V. R. Konda (2002) Actor-critic algorithms, 2002 | 0.405 | 1 | 1 | 100% |
| 5 | C. D. Meyer, Jr (1980) The condition of a finite markov chain and perturbation bounds for the limiting probabilities | 0.405 | 1 | 1 | 100% |
Showing the top 5 of 5 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.