Jann Spiess, Guido Imbens, Amar Venugopal
arXiv 1 May 2023 · Econometrics · 6 citations (OpenAlex)
arXiv:2305.00700 · PDF · DOI · OpenAlex · Extracted main text
Motivated by a recent literature on the double-descent phenomenon in machine learning, we consider highly over-parameterized models in causal inference, including synthetic control with many control units. In such models, there may be so many free parameters that the model fits the training data perfectly. We first investigate high-dimensional linear regression for imputing wage data and estimating average treatment effects, where we find that models with many more covariates than sample size can outperform simple ones. We then document the performance of high-dimensional synthetic control estimators with many control units. We find that adding control units can help improve imputation performance even beyond the point where the pre-treatment fit is perfect. We provide a unified theoretical perspective on the performance of these high-dimensional models. Specifically, we show that more complex models can be interpreted as model-averaging estimators over simpler ones, which we link to an improvement in average performance. This perspective yields concrete insights into the use of synthetic control when control units are many relative to the number of pre-treatment periods.
appendix boundary found by appendix_command · 62% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Abadie, Alberto, Alexis Diamond, and Jens Hainmueller (2010) Synthetic control methods for comparative case studies: Estimating the effect of california’s tobacco control program | 0.928 | 5 | 3 | 80% |
| 2 | LaLonde, Robert J (1986) Evaluating the econometric evaluations of training programs with experimental data | 0.928 | 5 | 3 | 80% |
| 3 | Dehejia, Rajeev H and Sadek Wahba (1999) Causal effects in nonexperimental studies: Reevaluating the evaluation of training programs | 0.843 | 4 | 3 | 75% |
| 4 | Dehejia, Rajeev H and Sadek Wahba (2002) Propensity score-matching methods for nonexperimental causal studies | 0.843 | 4 | 3 | 75% |
| 5 | Liang, Tengyuan, Alexander Rakhlin, and Xiyu Zhai (2020) On the Multiple Descent of Minimum-Norm Interpolants and Restricted Lower Isometry of Kernels | 0.811 | 4 | 2 | 100% |
| 6 | Bartlett, Peter L, Philip M Long, Gábor Lugosi, and Alexander Tsigler (2020) Benign overfitting in linear regression | 0.737 | 3 | 2 | 100% |
| 7 | Hastie, Trevor, Andrea Montanari, Saharon Rosset, and Ryan J Tibshir… (2022) Surprises in high-dimensional ridgeless least squares interpolation | 0.737 | 3 | 2 | 100% |
| 8 | Shen, Dennis, Peng Ding, Jasjeet Sekhon, and Bin Yu (2022) Same Root Different Leaves: Time Series and Cross-Sectional Methods in Panel Data | 0.644 | 2 | 2 | 100% |
| 9 | Kato, Masahiro and Masaaki Imaizumi (2022) Benign-overfitting in conditional average treatment effect prediction with linear regression | 0.644 | 2 | 2 | 100% |
| 10 | Bruns-Smith, David, Oliver Dukes, Avi Feller, and Elizabeth L Ogburn (2023) Augmented balancing weights as undersmoothed regressions | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 25 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Causal Models for Longitudinal and Panel Data: A Survey | 0.511 | 2 | 1 |
| 2 | Prediction Risk and Estimation Risk of the Ridgeless Least Squares Estimator under General Assumptions on Regression Errors | 0.405 | 1 | 1 |