EconBase
← All papers

Benign-Overfitting in Conditional Average Treatment Effect Prediction with Linear Regression

Masahiro Kato, Masaaki Imaizumi

arXiv 10 Feb 2022 · Econometrics · 3 citations (OpenAlex)

arXiv:2202.05245 · PDF · DOI · OpenAlex · Extracted main text

Abstract

We study the benign overfitting theory in the prediction of the conditional average treatment effect (CATE), with linear regression models. As the development of machine learning for causal inference, a wide range of large-scale models for causality are gaining attention. One problem is that suspicions have been raised that the large-scale models are prone to overfitting to observations with sample selection, hence the large models may not be suitable for causal prediction. In this study, to resolve the suspicious, we investigate on the validity of causal inference methods for overparameterized models, by applying the recent theory of benign overfitting (Bartlett et al., 2020). Specifically, we consider samples whose distribution switches depending on an assignment rule, and study the prediction of CATE with linear models whose dimension diverges to infinity. We focus on two methods: the T-learner, which based on a difference between separately constructed estimators with each treatment group, and the inverse probability weight (IPW)-learner, which solves another regression problem approximated by a propensity score. In both methods, the estimator consists of interpolators that fit the samples perfectly. As a result, we show that the T-learner fails to achieve the consistency except the random assignment, while the IPW-learner converges the risk to zero if the propensity score is known. This difference stems from that the T-learner is unable to preserve eigenspaces of the covariances, which is necessary for benign overfitting in the overparameterized setting. Our result provides new insights into the usage of causal inference methods in the overparameterizated setting, in particular, doubly robust estimators.

Citation extraction

76
references
164
in-text mentions
76
distinct cited
0
self-citations
6,707
main-text words

appendix boundary found by appendix_command · 38% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Bartlett, P. L., Long, P. M., Lugosi, G., and Tsigler, A (2020) Benign overfitting in linear regression0.754631143%
2Kennedy, E. H (2020) Optimal doubly robust estimation of heterogeneous causal effects0.73732100%
3Künzel, S. R., Sekhon, J. S., Bickel, P. J., and Yu, B (2019) Metalearners for estimating heterogeneous treatment effects using machine learning0.64441100%
4Abrevaya, J., Hsu, Y.-C., and Lieli, R. P (2015) Estimating Conditional Average Treatment Effects0.64422100%
5Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C… (2018) Double/debiased machine learning for treatment and structural parameters0.64422100%
6Hahn, J (1998) On the Role of the Propensity Score in Efficient Semiparametric Estimation of Average Treatment Effects0.64422100%
7Imbens, G. W. and Rubin, D. B (2015) Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction0.64422100%
8Koltchinskii, V. and Lounici, K (2017) Concentration inequalities and moment bounds for sample covariance operators0.5853333%
9Nie, X. and Wager, S (2020) Quasi-Oracle Estimation of Heterogeneous Treatment Effects0.58531100%
10Assmann, S., Pocock, S., Enos, L., and Kasten, L (2000) Subgroup analysis and other (mis)uses of baseline data in clinical trials0.51121100%

Showing the top 10 of 76 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Double and Single Descent in Causal Inference with an Application to High-Dimensional Synthetic Control0.64422