arXiv 28 Aug 2021 · Econometrics
arXiv:2108.12547 · PDF · DOI · OpenAlex · Extracted main text
This paper identifies and addresses dynamic selection problems in online learning algorithms with endogenous data. In a contextual multi-armed bandit model, a novel bias (self-fulfilling bias) arises because the endogeneity of the data influences the choices of decisions, affecting the distribution of future data to be collected and analyzed. We propose an instrumental-variable-based algorithm to correct for the bias. It obtains true parameter values and attains low (logarithmic-like) regret levels. We also prove a central limit theorem for statistical inference. To establish the theoretical properties, we develop a general technique that untangles the interdependence between data and actions.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Mila Nambiar, David Simchi-Levi \ He Wang (2019) Dynamic Learning and Pricing with Model Misspecification | 0.737 | 3 | 2 | 100% |
| 2 | Martin J. Wainwright (2019) High-Dimensional Statistics: A Non-Asymptotic Viewpoint. Cambridge University Press | 0.737 | 3 | 2 | 100% |
| 3 | Roman Vershynin (2018) High-Dimensional Probability: An Introduction with Applications in Data Science. Cambridge University Press | 0.644 | 2 | 2 | 100% |
| 4 | Ying Zhong, L. Jeff Hong \ Guangwu Liu (2021) Earning and Learning with Varying Cost | 0.644 | 2 | 2 | 100% |
| 5 | Alexander Goldenshluger \ Assaf Zeevi (2013) A Linear Response Bandit Problem | 0.585 | 3 | 1 | 100% |
| 6 | Hamsa Bastani, Mohsen Bayati \ Khashayar Khosravi (2021) Mostly Exploration-Free Algorithms for Contextual Bandits | 0.511 | 2 | 1 | 100% |
| 7 | Gene H. Golub \ Charles F. Van Loan (2013) Matrix Computations | 0.511 | 2 | 1 | 100% |
| 8 | Jin Li, Ye Luo \ Xiaowei Zhang (2021) Causal Reinforcement Learning: An Instrumental Variable Approach self | 0.511 | 2 | 1 | 100% |
| 9 | Joseph G. Altonji, Todd E. Elder \ Christopher R. Taber (2005) Selection on Observed and Unobserved Variables: Assessing the Effectiveness of Catholic Schools | 0.405 | 1 | 1 | 100% |
| 10 | Isaiah Andrews, James Stock \ Liyang Sun (2019) Weak Instruments in IV Regression: Theory and Practice | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 43 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Dynamic Decision-Making under Model Misspecification | 0.405 | 1 | 1 |