EconBase
← All papers

The Adaptive Doubly Robust Estimator for Policy Evaluation in Adaptive Experiments and a Paradox Concerning Logging Policy

Masahiro Kato, Shota Yasui, Kenichiro McAlinn

arXiv 8 Oct 2020 · Machine Learning

arXiv:2010.03792 · PDF · DOI · OpenAlex · Extracted main text

Abstract

The doubly robust (DR) estimator, which consists of two nuisance parameters, the conditional mean outcome and the logging policy (the probability of choosing an action), is crucial in causal inference. This paper proposes a DR estimator for dependent samples obtained from adaptive experiments. To obtain an asymptotically normal semiparametric estimator from dependent samples with non-Donsker nuisance estimators, we propose adaptive-fitting as a variant of sample-splitting. We also report an empirical paradox that our proposed DR estimator tends to show better performances compared to other estimators utilizing the true logging policy. While a similar phenomenon is known for estimators with i.i.d. samples, traditional explanations based on asymptotic efficiency cannot elucidate our case with dependent samples. We confirm this hypothesis through simulation studies.

Citation extraction

53
references
138
in-text mentions
53
distinct cited
3
self-citations
7,076
main-text words

appendix boundary found by appendix_command · 33% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1van der Laan, M. J (2008) The construction and analysis of adaptive group sequential designs, 20081.000145100%
2Hadad, V., Hirshberg, D. A., Zhan, R., Wager, S., and Athey, S (2021) Confidence intervals for policy evaluation in adaptive experiments0.98017594%
3Kato, M., Ishihara, T., Honda, J., and Narita, Y (2020) Adaptive experimental design for efficient treatment effect estimation: Randomized allocation via contextual bandit algorithm self0.96510490%
4Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C… (2018) Double/debiased machine learning for treatment and structural parameters0.9619489%
5Hirano, K., Imbens, G., and Ridder, G (2003) Efficient estimation of average treatment effects using the estimated propensity score0.9568588%
6Dudḱ, M., Langford, J., and Li, L (2011) Doubly Robust Policy Evaluation and Learning0.92844100%
7van der Laan, M. J. and Lendle, S. D (2014) Online targeted learning, 20140.87452100%
8Zhang, K., Janson, L., and Murphy, S (2020) Inference for batched bandits0.8307457%
9Hahn, J (1998) On the role of the propensity score in efficient semiparametric estimation of average treatment effects0.81142100%
10Henmi, M. and Eguchi, S (2004) A paradox concerning nuisance parameters and projected estimating functions0.7373367%

Showing the top 10 of 53 scored citations.