arXiv 5 Mar 2024 · Statistics — Methodology
arXiv:2403.03240 · PDF · DOI · OpenAlex · Extracted main text
This study investigates the estimation and the statistical inference about Conditional Average Treatment Effects (CATEs), which have garnered attention as a metric representing individualized causal effects. In our data-generating process, we assume linear models for the outcomes associated with binary treatments and define the CATE as a difference between the expected outcomes of these linear models. This study allows the linear models to be high-dimensional, and our interest lies in consistent estimation and statistical inference for the CATE. In high-dimensional linear regression, one typical approach is to assume sparsity. However, in our study, we do not assume sparsity directly. Instead, we consider sparsity only in the difference of the linear models. We first use a doubly robust estimator to approximate this difference and then regress the difference on covariates with Lasso regularization. Although this regression estimator is consistent for the CATE, we further reduce the bias using the techniques in double/debiased machine learning (DML) and debiased Lasso, leading to $\sqrt{n}$-consistency and confidence intervals. We refer to the debiased estimator as the triple/debiased Lasso (TDL), applying both DML and debiased Lasso techniques. We confirm the soundness of our proposed method through simulation studies.
appendix boundary found by appendix_command · 44% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Qingliang Fan, Yu-Chin Hsu, Robert P. Lieli, and Yichong Zhang (2022) Estimation of conditional average treatment effects with high-dimensional data | 0.874 | 8 | 2 | 100% |
| 2 | Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo,… (2018) Double/debiased machine learning for treatment and structural parameters | 0.843 | 4 | 4 | 75% |
| 3 | Sara van de Geer, Peter Bühlmann, Ya’acov Ritov, and Ruben Dezeure (2014) On asymptotically optimal confidence regions and tests for high-dimensional models | 0.843 | 20 | 6 | 60% |
| 4 | T. Tony Cai and Zijian Guo (2017) Confidence intervals for high-dimensional linear regression: Minimax rates and adaptivity | 0.737 | 3 | 3 | 67% |
| 5 | Adel Javanmard and Andrea Montanari (2014) Confidence intervals and hypothesis testing for high-dimensional regression | 0.737 | 3 | 3 | 67% |
| 6 | Peter L. Bartlett, Philip M. Long, Gábor Lugosi, and Alexander Tsigler (2020) Benign overfitting in linear regression | 0.737 | 3 | 2 | 100% |
| 7 | Jason Abrevaya, Yu-Chin Hsu, and Robert P. Lieli (2015) Estimating conditional average treatment effects | 0.644 | 2 | 2 | 100% |
| 8 | James J. Heckman, Hidehiko Ichimura, and Petra E. Todd (1997) Matching as an econometric evaluation estimator: Evidence from evaluating a job training programme | 0.644 | 2 | 2 | 100% |
| 9 | Sören R. Künzel, Jasjeet S. Sekhon, Peter J. Bickel, and Bin Yu (2019) Metalearners for estimating heterogeneous treatment effects using machine learning | 0.644 | 2 | 2 | 100% |
| 10 | Wenjing Zheng and Mark J van der Laan (2011) Cross-validated targeted minimum-loss-based estimation | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 57 scored citations.