arXiv 9 Jun 2021 · Econometrics · publishedQuantitative Economics (2025) · 7 citations (OpenAlex)
arXiv:2106.05031 · PDF · DOI · OpenAlex · Extracted main text
Many policies involve dynamics in their treatment assignments, where individuals receive sequential interventions over multiple stages. We study estimation of an optimal dynamic treatment regime that guides the optimal treatment assignment for each individual at each stage based on their history. We propose an empirical welfare maximization approach in this dynamic framework, which estimates the optimal dynamic treatment regime using data from an experimental or quasi-experimental study while satisfying exogenous constraints on policies. The paper proposes two estimation methods: one solves the treatment assignment problem sequentially through backward induction, and the other solves the entire problem simultaneously across all stages. We establish finite-sample upper bounds on worst-case average welfare regrets for these methods and show their optimal $n^{-1/2}$ convergence rates. We also modify the simultaneous estimation method to accommodate intertemporal budget/capacity constraints.
appendix boundary found by appendix_command · 41% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Rodríguez, J., F. Saltiel, and S. Urzúa (2022) Dynamic treatment effects of job training | 0.928 | 4 | 3 | 100% |
| 2 | Athey, S. and S. Wager (2021) Policy learning with observational data | 0.899 | 11 | 4 | 73% |
| 3 | Kitagawa, T. and A. Tetenov (2018) b): Who should be treated? Empirical welfare maximization methods for treatment choice | 0.860 | 11 | 6 | 64% |
| 4 | Nie, X., E. Brunskill, and S. Wager (2021) Learning when-to-treat policies | 0.811 | 4 | 2 | 100% |
| 5 | Weymark, J. A (1981) Generalized Gini inequality indices | 0.737 | 5 | 2 | 60% |
| 6 | Meyer, B. D (1995) Lessons from the U.S | 0.737 | 3 | 3 | 67% |
| 7 | Robins, J. M (1997) Causal inference from complex longitudinal data in latent variable modeling and applications to causality, in | 0.737 | 3 | 2 | 100% |
| 8 | Zhou, Z., S. Athey, and S. Wager (2023) Offline multi-action policy learning: Generalization and optimization | 0.707 | 17 | 4 | 35% |
| 9 | Jiang, N. and L. Li (2016) Doubly robust off-policy value evaluation for reinforcement learning, in | 0.644 | 4 | 2 | 50% |
| 10 | Sakaguchi, S (2024) Robust learning for optimal dynamic treatment regimes with observational data, ArXiv:2404.00221 self | 0.644 | 4 | 2 | 50% |
Showing the top 10 of 68 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Sequential Learning of Optimal Dynamic Treatment Regimes with Observational Data | 1.000 | 7 | 5 |
| 2 | Who Should Get Vaccinated? Individualized Allocation of Vaccines Over SIR Network | 0.405 | 1 | 1 |
| 3 | Constrained Classification and Policy Learning | 0.405 | 1 | 1 |
| 4 | Evidence Aggregation for Treatment Choice | 0.405 | 1 | 1 |
| 5 | Who With Whom? Learning Optimal Matching Policies | 0.405 | 1 | 1 |