Jingwen Zhang, Yifang Chen, Amandeep Singh
arXiv 16 Nov 2022 · Econometrics · 2 citations (OpenAlex)
arXiv:2211.08649 · PDF · DOI · OpenAlex · Extracted main text
The deployment of Multi-Armed Bandits (MAB) has become commonplace in many economic applications. However, regret guarantees for even state-of-the-art linear bandit algorithms (such as Optimism in the Face of Uncertainty Linear bandit (OFUL)) make strong exogeneity assumptions w.r.t. arm covariates. This assumption is very often violated in many economic contexts and using such algorithms can lead to sub-optimal decisions. Further, in social science analysis, it is also important to understand the asymptotic distribution of estimated parameters. To this end, in this paper, we consider the problem of online learning in linear stochastic contextual bandit problems with endogenous covariates. We propose an algorithm we term $\epsilon$-BanditIV, that uses instrumental variables to correct for this bias, and prove an $\tilde{\mathcal{O}}(k\sqrt{T})$ upper bound for the expected regret of the algorithm. Further, we demonstrate the asymptotic consistency and normality of the $\epsilon$-BanditIV estimator. We carry out extensive Monte Carlo simulations to demonstrate the performance of our algorithms compared to other methods. We show that $\epsilon$-BanditIV significantly outperforms other existing methods in endogenous settings. Finally, we use data from real-time bidding (RTB) system to demonstrate how $\epsilon$-BanditIV can be used to estimate the causal impact of advertising in such settings and compare its performance with other existing methods.
appendix boundary found by appendix_titled_section at “Appendix” · 64% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C (2011) Improved algorithms for linear stochastic bandits | 0.928 | 5 | 4 | 80% |
| 2 | Aramayo, N., Schiappacasse, M., and Goic, M (2022) A multi-armed bandit approach for house ads recommendations | 0.737 | 3 | 2 | 100% |
| 3 | Auer, P (2002) Using confidence bounds for exploitation-exploration trade-offs | 0.644 | 2 | 2 | 100% |
| 4 | Chu, W., Li, L., Reyzin, L., and Schapire, R (2011) Contextual bandits with linear payoff functions | 0.644 | 2 | 2 | 100% |
| 5 | Dani, V., Hayes, T. P., and Kakade, S. M (2008) Stochastic linear optimization under bandit feedback | 0.644 | 2 | 2 | 100% |
| 6 | Johnson, G. A., Lewis, R. A., and Nubbemeyer, E. I (2017) Ghost ads: Improving the economics of measuring online ad effectiveness | 0.644 | 2 | 2 | 100% |
| 7 | Tang, L., Jiang, Y., Li, L., Zeng, C., and Li, T (2015) Personalized recommendation via parameter-free contextual bandits | 0.644 | 2 | 2 | 100% |
| 8 | Chen, H., Lu, W., and Song, R (2021) Statistical inference for online decision making: In a contextual bandit setting | 0.630 | 8 | 3 | 25% |
| 9 | Bastani, H. and Bayati, M (2020) Online decision making with high-dimensional covariates | 0.511 | 2 | 2 | 50% |
| 10 | Bojinov, I., Simchi-Levi, D., and Zhao, J (2020) Design and analysis of switchback experiments | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 52 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Adaptive Principal Component Regression with Applications to Panel Data | 0.405 | 1 | 1 |