EconBase
← All papers

Policy Learning for Optimal Dynamic Treatment Regimes with Observational Data

Shosei Sakaguchi

arXiv 30 Mar 2024 · Statistics — Methodology · 1 citations (OpenAlex)

arXiv:2404.00221 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Public policies and medical interventions often involve dynamic treatment assignments, in which individuals receive a sequence of interventions over multiple stages. We study the statistical learning of optimal dynamic treatment regimes (DTRs) that determine the optimal treatment assignment for each individual at each stage based on their evolving history. We propose a novel, doubly robust, classification-based method for learning the optimal DTR from observational data under the sequential ignorability assumption. The method proceeds via backward induction: at each stage, it constructs and maximizes an augmented inverse probability weighting (AIPW) estimator of the policy value function to learn the optimal stage-specific policy. We show that the resulting DTR achieves an optimal convergence rate of $n^{-1/2}$ for welfare regret under mild convergence conditions on estimators of the nuisance components.

Citation extraction

35
references
106
in-text mentions
40
distinct cited
1
self-citations
10,813
main-text words

appendix boundary found by appendix_command · 33% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Sakaguchi, S (2025) Estimation of Optimal Dynamic Treatment Assignment Rules Under Policy Constraints self1.00075100%
2Murphy, S. A (2005) A Generalization Error for Q-learning1.00073100%
3Zhang, B., A. A. Tsiatis, E. B. Laber, and M. Davidian (2013) Robust Estimation of Optimal Dynamic Treatment Regimes for Sequential Treatment Decisions1.00054100%
4Zhao, Y. Q., D. Zeng, E. B. Laber, and M. R. Kosorok (2015) New Statistical Learning Methods for Estimating Optimal Dynamic Treatment Regimes1.00054100%
5Nie, X., E. Brunskill, and S. Wager (2021) Learning When-to-Treat Policies0.92843100%
6Athey, S. and S. Wager (2021) Policy Learning with Observational Data0.8947471%
7Zhang, Y., E. B. Laber, M. Davidian, and A. A. Tsiatis (2018) Interpretable Dynamic Treatment Regimes0.84333100%
8Zhou, Z., S. Athey, and S. Wager (2023) b): Offline Multi-Action Policy Learning: Generalization and Optimization0.82218656%
9Jiang, N. and L. Li (2016) Doubly Robust Off-policy Value Evaluation for Reinforcement Learning, in0.73732100%
10Robins, J. M (1997) Causal Inference From Complex Longitudinal Data in Latent Variable Modeling and Applications to Causality, in0.73732100%

Showing the top 10 of 40 scored citations.