Jin Li, Ye Luo, Zigan Wang, Xiaowei Zhang
arXiv 6 Mar 2021 · Statistics — Machine Learning · 1 citations (OpenAlex)
arXiv:2103.04021 · PDF · DOI · OpenAlex · Extracted main text
In the standard data analysis framework, data is collected (once and for all), and then data analysis is carried out. However, with the advancement of digital technology, decision-makers constantly analyze past data and generate new data through their decisions. We model this as a Markov decision process and show that the dynamic interaction between data generation and data analysis leads to a new type of bias -- reinforcement bias -- that exacerbates the endogeneity problem in standard data analysis. We propose a class of instrument variable (IV)-based reinforcement learning (RL) algorithms to correct for the bias and establish their theoretical properties by incorporating them into a stochastic approximation (SA) framework. Our analysis accommodates iterate-dependent Markovian structures and, therefore, can be used to study RL algorithms with policy improvement. We also provide formulas for inference on optimal policies of the IV-RL algorithms. These formulas highlight how intertemporal dependencies of the Markovian environment affect the inference.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Sutton, R. S. and Barto, A. G (2018) Reinforcement Learning: An Introduction | 1.000 | 8 | 5 | 100% |
| 2 | Konda, V. R. and Tsitsiklis, J. N (2003) On actor-critic algorithms | 0.928 | 4 | 3 | 100% |
| 3 | Hadad, V., Hirshberg, D. A., Zhan, R., Wager, S., and Athey, S (2021) Confidence intervals for policy evaluation in adaptive experiments | 0.843 | 3 | 3 | 100% |
| 4 | Zhan, R., Hadad, V., Hirshberg, D. A., and Athey, S (2021) Off-policy evaluation via adaptive weighting with data from contextual bandits | 0.843 | 3 | 3 | 100% |
| 5 | Kushner, H. J. and Yin, G. G (2003) Stochastic Approximation and Recursive Algorithms and Applications | 0.737 | 3 | 2 | 100% |
| 6 | Almeida, H., Fos, V., and Kronlund, M (2016) The real effects of share repurchases | 0.693 | 5 | 1 | 100% |
| 7 | Bottou, L., Peters, J., Quiñonero-Candela, J., Charles, D. X., Chick… (2013) Counterfactual reasoning and learning systems: The example of computational advertising | 0.644 | 2 | 2 | 100% |
| 8 | Calvano, E., Calzolari, G., Denicolò, V., and Pastorello, S (2020) Artificial intelligence, algorithmic pricing, and collusion | 0.644 | 2 | 2 | 100% |
| 9 | Chen, X. and White, H (2002) Asymptotic properties of some projection-based Robbins–Monro procedures in a Hilbert space | 0.644 | 2 | 2 | 100% |
| 10 | Melo, F. S., Meyn, S. P., and Ribeiro, M. I (2008) An analysis of reinforcement learning with function approximation | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 53 scored citations.