EconBase
← All papers

CATE Lasso: Conditional Average Treatment Effect Estimation with High-Dimensional Linear Regression

Masahiro Kato, Masaaki Imaizumi

arXiv 25 Oct 2023 · Econometrics · 3 citations (OpenAlex)

arXiv:2310.16819 · PDF · DOI · OpenAlex · Extracted main text

Abstract

In causal inference about two treatments, Conditional Average Treatment Effects (CATEs) play an important role as a quantity representing an individualized causal effect, defined as a difference between the expected outcomes of the two treatments conditioned on covariates. This study assumes two linear regression models between a potential outcome and covariates of the two treatments and defines CATEs as a difference between the linear regression models. Then, we propose a method for consistently estimating CATEs even under high-dimensional and non-sparse parameters. In our study, we demonstrate that desirable theoretical properties, such as consistency, remain attainable even without assuming sparsity explicitly if we assume a weaker assumption called implicit sparsity originating from the definition of CATEs. In this assumption, we suppose that parameters of linear models in potential outcomes can be divided into treatment-specific and common parameters, where the treatment-specific parameters take difference values between each linear regression model, while the common parameters remain identical. Thus, in a difference between two linear regression models, the common parameters disappear, leaving only differences in the treatment-specific parameters. Consequently, the non-zero parameters in CATEs correspond to the differences in the treatment-specific parameters. Leveraging this assumption, we develop a Lasso regression method specialized for CATE estimation and present that the estimator is consistent. Finally, we confirm the soundness of the proposed method by simulation studies.

Citation extraction

78
references
169
in-text mentions
78
distinct cited
0
self-citations
6,724
main-text words

appendix boundary found by appendix_command · 59% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1X Nie and S Wager (2020) Quasi-oracle estimation of heterogeneous treatment effects0.8435360%
2James J. Heckman, Hidehiko Ichimura, and Petra E. Todd (1997) Matching as an econometric evaluation estimator: Evidence from evaluating a job training programme0.7373367%
3Stefan Wager and Susan Athey (2018) Estimation and inference of heterogeneous treatment effects using random forests0.7373367%
4Sören R. Künzel, Jasjeet S. Sekhon, Peter J. Bickel, and Bin Yu (2019) Metalearners for estimating heterogeneous treatment effects using machine learning0.6938250%
5Jerzy Neyman (1923) Sur les applications de la theorie des probabilites aux experiences agricoles: Essai des principes0.64422100%
6Donald B. Rubin (1974) Estimating causal effects of treatments in randomized and nonrandomized studies0.64422100%
7R. Tibshirani (1996) Regression shrinkage and selection via the lasso0.64422100%
8Sara van de Geer, Peter Bühlmann, Ya’acov Ritov, and Ruben Dezeure (2014) On asymptotically optimal confidence regions and tests for high-dimensional models0.6069322%
9Jennifer L. Hill (2011) Bayesian nonparametric modeling for causal inference0.5506317%
10Peter Bühlmann and Sara van de Geer (2011) Statistics for high-dimensional data0.5299222%

Showing the top 10 of 78 scored citations.