arXiv 30 Mar 2024 · Statistics — Methodology · 1 citations (OpenAlex)
arXiv:2404.00221 · PDF · DOI · OpenAlex · Extracted main text
Public policies and medical interventions often involve dynamic treatment assignments, in which individuals receive a sequence of interventions over multiple stages. We study the statistical learning of optimal dynamic treatment regimes (DTRs) that determine the optimal treatment assignment for each individual at each stage based on their evolving history. We propose a novel, doubly robust, classification-based method for learning the optimal DTR from observational data under the sequential ignorability assumption. The method proceeds via backward induction: at each stage, it constructs and maximizes an augmented inverse probability weighting (AIPW) estimator of the policy value function to learn the optimal stage-specific policy. We show that the resulting DTR achieves an optimal convergence rate of $n^{-1/2}$ for welfare regret under mild convergence conditions on estimators of the nuisance components.
appendix boundary found by appendix_command · 33% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Sakaguchi, S (2025) Estimation of Optimal Dynamic Treatment Assignment Rules Under Policy Constraints self | 1.000 | 7 | 5 | 100% |
| 2 | Murphy, S. A (2005) A Generalization Error for Q-learning | 1.000 | 7 | 3 | 100% |
| 3 | Zhang, B., A. A. Tsiatis, E. B. Laber, and M. Davidian (2013) Robust Estimation of Optimal Dynamic Treatment Regimes for Sequential Treatment Decisions | 1.000 | 5 | 4 | 100% |
| 4 | Zhao, Y. Q., D. Zeng, E. B. Laber, and M. R. Kosorok (2015) New Statistical Learning Methods for Estimating Optimal Dynamic Treatment Regimes | 1.000 | 5 | 4 | 100% |
| 5 | Nie, X., E. Brunskill, and S. Wager (2021) Learning When-to-Treat Policies | 0.928 | 4 | 3 | 100% |
| 6 | Athey, S. and S. Wager (2021) Policy Learning with Observational Data | 0.894 | 7 | 4 | 71% |
| 7 | Zhang, Y., E. B. Laber, M. Davidian, and A. A. Tsiatis (2018) Interpretable Dynamic Treatment Regimes | 0.843 | 3 | 3 | 100% |
| 8 | Zhou, Z., S. Athey, and S. Wager (2023) b): Offline Multi-Action Policy Learning: Generalization and Optimization | 0.822 | 18 | 6 | 56% |
| 9 | Jiang, N. and L. Li (2016) Doubly Robust Off-policy Value Evaluation for Reinforcement Learning, in | 0.737 | 3 | 2 | 100% |
| 10 | Robins, J. M (1997) Causal Inference From Complex Longitudinal Data in Latent Variable Modeling and Applications to Causality, in | 0.737 | 3 | 2 | 100% |
Showing the top 10 of 40 scored citations.