EconBase
← All papers

Sequential Decision Problems with Missing Feedback

Filippo Palomba

arXiv 25 Jul 2025 · Econometrics

arXiv:2507.19596 · PDF · Extracted main text

Abstract

This paper investigates the challenges of optimal online policy learning under missing data. State-of-the-art algorithms implicitly assume that rewards are always observable. I show that when rewards are missing at random, the Upper Confidence Bound (UCB) algorithm maintains optimal regret bounds; however, it selects suboptimal policies with high probability as soon as this assumption is relaxed. To overcome this limitation, I introduce a fully nonparametric algorithm-Doubly-Robust Upper Confidence Bound (DR-UCB)-which explicitly models the form of missingness through observable covariates and achieves a nearly-optimal worst-case regret rate of $\widetilde{O}(\sqrt{T})$. To prove this result, I derive high-probability bounds for a class of doubly-robust estimators that hold under broad dependence structures. Simulation results closely match the theoretical predictions, validating the proposed framework.

Citation extraction

30
references
43
in-text mentions
31
distinct cited
0
self-citations
35,588
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Lattimore and Szepesvári (2020)1.00074100%
2Auer, Cesa-Bianchi and Fischer (2002) Finite-Time Analysis of the Multiarmed Bandit Problem0.81142100%
3Horvitz and Thompson (1952) A Generalization of Sampling Without Replacement from a Finite Universe0.64422100%
4Vershynin (2018)0.64422100%
5Wald (1947)0.51121100%
6Adusumilli (2024) Risk and Optimal Policies in Bandit Experiments0.40511100%
7Ahrens, Chernozhukov, Hansen, Kozbur, Schaffer and Wiemann (2025) An Introduction to Double/Debiased Machine Learning0.40511100%
8Athey and Wager (2021) Policy Learning With Observational Data0.40511100%
9Bang and Robins (2005) Doubly Robust Estimation in Missing Data and Causal Inference Models0.40511100%
boyd2004ConvexOptimizationunmatched citation key boyd2004ConvexOptimization0.40511100%

Showing the top 10 of 31 scored citations. 1 of these could not be matched to a bibliography entry, so only the citation key is shown.