EconBase
← All papers

Causal Forecasting in Panel Data: A Two-Way Synthetic Forecasting Approach

Dennis Shen

arXiv 16 Jun 2026 · Econometrics

arXiv:2606.18512 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Estimating causal effects in panel data is a central problem in policy evaluation. Existing methods largely address retrospective questions of the form: what would have happened to a target unit under a different intervention during the observed panel? In many applications, however, decision-makers face prospective questions: what will happen to a target unit under an intervention it has not yet experienced, beyond the observed panel? This article develops a framework for answering such causal forecasting questions by integrating the retrospective counterfactual logic of synthetic-controls-based approaches with the extrapolative structure of multivariate time-series forecasting. Building on the latent factor models that justify unit-side regressions in synthetic controls, we impose low-rank temporal structure on the latent time factors to identify prospective causal forecast estimands. We operationalize this strategy through the Two-Way Synthetic Forecasting estimator, or TWSF, which learns cross-unit relationships from pre-treatment outcomes and combines them with a time-series model learned from treated donor trajectories under the intervention of interest. Under suitable conditions, we establish finite-sample forecasting error bounds that imply pointwise consistency and introduce an orthogonalized correction that yields asymptotic normality and thus enables pointwise inference. We extend the framework to fixed multi-step forecasting horizons through both direct and recursive procedures, each of which inherits analogous pointwise guarantees. We corroborate the theory with simulation studies and illustrate the practical utility of TWSF by studying the public-health impact of opening NFL stadiums during the 2020 season.

Citation extraction

69
references
112
in-text mentions
69
distinct cited
7
self-citations
19,968
main-text words

appendix boundary found by appendix_command · 43% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Agarwal, Anish and Shah, Devavrat and Shen, Dennis (2026) Synthetic Interventions: Extending Synthetic Controls to Multiple Treatments self1.000105100%
2Agarwal, Anish and Alomar, Abdullah and Shah, Devavrat (2022) On Multivariate Singular Spectrum Analysis and Its Variants1.00093100%
3Bernardo García Bulle and Dennis Shen and Devavrat Shah and Anette E… (2022) Public health implications of opening National Football League stadiums during the COVID-19 pandemic self0.87482100%
4Eli Ben-Michael and Avi Feller and Jesse Rothstein (2021) The Augmented Synthetic Control Method0.84333100%
5Navonil Deb and Raaz Dwivedi and Sumanta Basu (2026) Counterfactual Forecasting for Panel Data0.84333100%
6A. Abadie and A. Diamond and J. Hainmueller (2010) Synthetic Control Methods for Comparative Case Studies: Estimating the Effect of California's Tobacco Control Program.0.73732100%
7Anish Agarwal and Devavrat Shah and Dennis Shen (2025) On Model Identification and Out-of-Sample Prediction of PCR with Applications to Synthetic Controls self0.64441100%
8Jungjun Choi and Ming Yuan and (2024) Matrix Completion When Missing Is Not at Random and Its Applications in Causal Panel Data Models0.64422100%
9Athey, Susan and Bayati, Mohsen and Doudchenko, Nikolay and Imbens,… (2021) Matrix completion methods for causal panel data models0.64422100%
10Jushan Bai and Serena Ng and (2021) Matrix Completion, Counterfactuals, and Factor Analysis of Missing Data0.64422100%

Showing the top 10 of 69 scored citations.