Masahiro Kato, Shota Yasui, Kenichiro McAlinn
arXiv 8 Oct 2020 · Machine Learning
arXiv:2010.03792 · PDF · DOI · OpenAlex · Extracted main text
The doubly robust (DR) estimator, which consists of two nuisance parameters, the conditional mean outcome and the logging policy (the probability of choosing an action), is crucial in causal inference. This paper proposes a DR estimator for dependent samples obtained from adaptive experiments. To obtain an asymptotically normal semiparametric estimator from dependent samples with non-Donsker nuisance estimators, we propose adaptive-fitting as a variant of sample-splitting. We also report an empirical paradox that our proposed DR estimator tends to show better performances compared to other estimators utilizing the true logging policy. While a similar phenomenon is known for estimators with i.i.d. samples, traditional explanations based on asymptotic efficiency cannot elucidate our case with dependent samples. We confirm this hypothesis through simulation studies.
appendix boundary found by appendix_command · 33% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | van der Laan, M. J (2008) The construction and analysis of adaptive group sequential designs, 2008 | 1.000 | 14 | 5 | 100% |
| 2 | Hadad, V., Hirshberg, D. A., Zhan, R., Wager, S., and Athey, S (2021) Confidence intervals for policy evaluation in adaptive experiments | 0.980 | 17 | 5 | 94% |
| 3 | Kato, M., Ishihara, T., Honda, J., and Narita, Y (2020) Adaptive experimental design for efficient treatment effect estimation: Randomized allocation via contextual bandit algorithm self | 0.965 | 10 | 4 | 90% |
| 4 | Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C… (2018) Double/debiased machine learning for treatment and structural parameters | 0.961 | 9 | 4 | 89% |
| 5 | Hirano, K., Imbens, G., and Ridder, G (2003) Efficient estimation of average treatment effects using the estimated propensity score | 0.956 | 8 | 5 | 88% |
| 6 | Dudḱ, M., Langford, J., and Li, L (2011) Doubly Robust Policy Evaluation and Learning | 0.928 | 4 | 4 | 100% |
| 7 | van der Laan, M. J. and Lendle, S. D (2014) Online targeted learning, 2014 | 0.874 | 5 | 2 | 100% |
| 8 | Zhang, K., Janson, L., and Murphy, S (2020) Inference for batched bandits | 0.830 | 7 | 4 | 57% |
| 9 | Hahn, J (1998) On the role of the propensity score in efficient semiparametric estimation of average treatment effects | 0.811 | 4 | 2 | 100% |
| 10 | Henmi, M. and Eguchi, S (2004) A paradox concerning nuisance parameters and projected estimating functions | 0.737 | 3 | 3 | 67% |
Showing the top 10 of 53 scored citations.