Jessy Xinyi Han, Devavrat Shah
arXiv 18 Nov 2025 · Machine Learning
arXiv:2511.14133 · PDF · Extracted main text
Estimating causal effects on time-to-event outcomes from observational data is particularly challenging due to censoring, limited sample sizes, and non-random treatment assignment. The need for answering such "when-if" questions--how the timing of an event would change under a specified intervention--commonly arises in real-world settings with heterogeneous treatment adoption and confounding. To address these challenges, we propose Synthetic Survival Control (SSC) to estimate counterfactual hazard trajectories in a panel data setting where multiple units experience potentially different treatments over multiple periods. In such a setting, SSC estimates the counterfactual hazard trajectory for a unit of interest as a weighted combination of the observed trajectories from other units. To provide formal justification, we introduce a panel framework with a low-rank structure for causal survival analysis. Indeed, such a structure naturally arises under classical parametric survival models. Within this framework, for the causal estimand of interest, we establish identification and finite sample guarantees for SSC. We validate our approach using a multi-country clinical dataset of cancer treatment outcomes, where the staggered introduction of new therapies creates a quasi-experimental setting. Empirically, we find that access to novel treatments is associated with improved survival, as reflected by lower post-intervention hazard trajectories relative to their synthetic counterparts. Given the broad relevance of survival analysis across medicine, economics, and public policy, our framework offers a general and interpretable tool for counterfactual survival inference using observational data.
appendix boundary found by appendix_command · 59% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Robins, James M. and Finkelstein, Dianne M (2000) Correcting for Noncompliance and Dependent Censoring in an AIDS Clinical Trial with Inverse Probability of Censoring Weighted (I… | 0.843 | 3 | 3 | 100% |
| 2 | Anish Agarwal and Devavrat Shah and Dennis Shen (2024) Synthetic Interventions self | 0.763 | 9 | 4 | 44% |
| 3 | Alberto Abadie and Alexis Diamond and Jens Hainmueller and (2010) Synthetic Control Methods for Comparative Case Studies: Estimating the Effect of California's Tobacco Control Program | 0.644 | 2 | 2 | 100% |
| 4 | Cox, D. R (1972) Regression Models and Life‐Tables | 0.644 | 2 | 2 | 100% |
| 5 | Chenyin Gao and Zhiming Zhang and Shu Yang (2024) Causal Customer Churn Analysis with Low-rank Tensor Block Hazard Model | 0.644 | 2 | 2 | 100% |
| 6 | Han, Jessy Xinyi and Koh, Min Jung and Boussi, Leora and Sorial, Mar… (2025) Global outcomes and prognosis for relapsed/refractory mature T-cell and NK-cell lymphomas: results from the PETAL consortium self | 0.511 | 2 | 1 | 100% |
| 7 | Shu Hu and George H. Chen (2024) Fairness in Survival Analysis with Distributionally Robust Optimization | 0.405 | 1 | 1 | 100% |
| 8 | Anastasios A. Tsiatis (1981) A Large Sample Study of Cox's Regression Model | 0.405 | 1 | 1 | 100% |
| 9 | Abadie, Alberto and Gardeazabal, Javier (2003) The Economic Costs of Conflict: A Case Study of the Basque Country | 0.405 | 1 | 1 | 100% |
| 10 | Alberto Abadie and Anish Agarwal and Devavrat Shah (2025) A Causal Inference Framework for Data Rich Environments self | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 30 scored citations.