Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
44,420 characters · 19 sections · 45 citation commands
CATE Lasso: Conditional Average Treatment Effect Estimation with High-Dimensional Linear Regression
\affil[1]{Department of Basic Science, The University of Tokyo} \affil[2]{RIKEN Center for Advanced Intelligence Project}
\thispagestyle{empty}
Estimating the causal effects of binary treatment from observations is a central task in various fields, such as economics Wager2018, medicine Assmann2000, and online advertisement Bottou13. Specifically, we consider a case where there is a binary treatment and investigate conditional average treatment effects hahn1998role,Heckman1997,Abrevaya2015, defined as a difference of the expected values of binary treatments' scalar outcomes conditioned on covariates. CATEs have garnered attention as a quantity that captures the heterogeneity among individuals' treatment effects in the population.
Estimating CATEs requires regression models that approximate the conditional expectation of outcomes. This study assumes two linear regression models between a potential outcome and covariates of each treatment. Then, CATEs are defined as a difference of the two linear regression models. Under this setting, our interest lies in consistently estimating CATEs when the two linear regression models have high-dimensional and non-sparse parameters.
Estimating parameters in high-dimensional regression models is usually challenging, especially when the dimension is larger than the sample size. The solutions of least squares may lack desirable properties such as consistency, typically held in low-dimensional models under mild conditions. Various estimation and inference approaches have been proposed for high-dimensional linear regression models. This study focuses on linear regression with sparsity using the Lasso tibshirani96regression, zhao06a, vandeGeer2008.
We first formulate our problem using the Neyman-Rubin causal models Neyman1923, Rubin1974, which define a potential outcome for each treatment. We observe one of the potential outcomes corresponding to our actual treatment. For each treatment, between the potential outcome and covariates, we assume linear regression models; that is, there are two linear models corresponding to binary treatments. Then, we define CATEs as a difference between the two linear models.
For the linear models of each treatment, we assume that parameters are separable into treatment-specific and common parameters. While the treatment-specific parameters take different values between linear models in each treatment, the common parameters are the same. Therefore, when taking the difference between two linear regression models for each binary treatment, the common parameters disappear, and only the treatment-specific parameters remain. Therefore, even under high-dimensional and non-sparse linear regression models, the total dimension of CATE linear models depends only on that of the treatment-specific parameters. Hence, if the dimension of the treatment-specific parameters is low, we can employ properties similar to ones obtained under sparsity. Because we do not explicitly assume sparsity for linear regression models in potential outcome level, we refer to this sparsity as implicit sparsity.
By utilizing this implicit sparsity, we propose the Lasso-based regression specialized for CATE estimation, referred to as the CATE Lasso. The CATE Lasso regularizes the parameters by adding an $\ell_1$-norm for differences of the parameters, while the standard Lasso regularizes the parameters themselves. Surprisingly, even if linear models are high-dimension and non-sparse in a potential outcome level, we can still show consistency by utilizing the assumption. Furthermore, this assumption includes a case where linear models in potential outcomes are sparse. Thus, we develop a regression method for CATE estimation with high-dimensional and non-sparse liner models in a potential outcome level under the implicit sparsity assumption.
In summary, our contribution lies in the proposal of the Lasso method specialized for CATE estimation. Our method allows us to estimate the CATE without assuming the sparsity for linear model in a potential outcome level. We also show the consistency of our proposed estimator. Surprisingly, although we cannot show consistency for each linear model of each treatment, we can show consistency for the CATE defined as a difference of the linear models. Our method does not require nuisance estimators such as the propensity score and conditional mean function as in the inverse probability weighting (IPW) and doubly robust (DR) estimators. For that points, our method has advantages compared to the IPW and DR-based methods.
Organization. The structure of this study is organized as follows. Initially, we formulate our problem in Section (ref) and summarize the notations in Section (ref). Section (ref) is dedicated to defining our high-dimensional linear regression models with implicit sparsity. For these linear regression models, we introduce Lasso-based estimators in Section (ref), referred to as the CATE Lasso estimator. Subsequently, Section (ref) provides theoretical results for the CATE Lasso estimator. Experimental results that confirm the validity of our proposed estimators are showcased in Section (ref). In Section (ref), we introduce related work.
In this section, we define our problem setting. Our formulation is based on the Neyman-Rubin potential outcome framework Neyman1923,Rubin1974, which defines potential outcomes and observations, separately.
Suppose that there are two treatments $\mathcal{D} := \{1, 0\}$. In a typical situation, treatment $d=1$ corresponds to the active treatment, and treatment $d=0$ corresponds to the control treatment. For example, in a clinical trial, $d=1$ corresponds to a new drug, and $d=0$ corresponds to the placebo. We then posit the existence of potential outcome random variables $Y^1, Y^0 \in\mathbb{R}$ corresponding respectively to the treatment $d=1$ and $d=0$. Additionally, we assume that there are $p$-dimensional covariates $X\in\mathcal{X}\subset\mathbb{R}^p$, where $\mathcal{X}$ is the covariate space.
Let $P := (P^1, P^0)$ denote a set of distributions of $P^1$ and $P^0$, which are joint distributions of $(Y^1, X)$ and $(Y^0, X)$. Let $\mathbb{E}_P[\cdot]$ be an expectation under $P$.
In this study, our interest lies in the CATE at $X = x \in \mathcal{X}$ defined as
This quantity has been widely used in empirical studies of various fields, such as epidemiology, economics, and political science, because it captures heterogeneous treatment effects for each individual represented by a characteristic $x\in\mathcal{X}$.
Although there exist $Y^1$ and $Y^0$ as potential outcomes, we can only observe one of the outcomes, corresponding to an actual treatment $D\in\mathcal{D}$. Let $D \in \mathcal{D}$ denote an actual treatment indicator. By using $Y^1$, $Y^0$, and $D$, we define an observed outcome $Y\in\mathbb{R}$ as
Here, if $D = 1$, we observe $Y = Y^1$; if $D = 0$, we observe $Y = Y^0$.
Then, we define our observations. Let $n$ be the sample size. For each $i \in [n] := \{1,2,\dots,n\}$, let $(Y_i, D_i, X_i)$ be an independent and identically distributed (i.i.d.) copy of $(Y, D, X)$. Then, we suppose that the following $n$ samples are observable:
Recall that we defined $P$ as a set of $P^1$ and $P^0$, which are joint distributions of potential outcomes and covariates, $(Y^1, X)$ and $(Y^0, X)$. Similarly, we denote a joint distribution of observations $(Y, D, X)$ by $Q$. For the data-generating process, let $P_0$ and $Q_0$ be sets of the true distributions that generates the observations $\left\{(Y_i, D_i, X_i)\right\}^n_{i=1}$. Let $f_{P_0}$ be $f_0$.
Our goal is to estimate the CATE $f_0(x)$ by using $\left\{(Y_i, D_i, X_i)\right\}^n_{i=1}$. In particular, we aim to obtain a consistent estimator of the CATE $f_0(x)$ at a point $X = x$.
For identification of $\tau(x)$, we assume the unconfoundedness.
Assumption (ref) expresses a setting wherein the assignment is independent of the output conditioned on the covariates. This is the standard approach in treatment effect estimation Rosenbaum1983.
Assumption (ref) allows us to avoid overlap in treatment assignments. By this assumption, we can guarantee the identifiability of the assignment and the parameters imbens_rubin_2015.
Let $\mathcal{P} := \{1,\dots,p\}$. Let us define a ($n\times p$)-matrix $ \bm{X} :=
=
$, and a $n$-dimensional vector $ \mathbb{D} := (D_1,D_2,...,D_n)^\top $. Let $m^1 := \sum^n_{i=1}D_i$ and $m^0 := \sum^n_{i=1}( 1- D_i) = n - m_1$. For each $d\in\mathcal{D}$ and $k\in\mathcal{P}$, let us define $m^d$-dimensional column vectors $\widetilde{\mathbb{Y}}^d$ and $\widetilde{\mathbb{X}}^d_k$ as $ \widetilde{\mathbb{Y}}^d = (Y_i)^\top_{i\in\{1,2,\dots, n\}| D_i=d}$, $ \widetilde{\mathbb{X}}^d_k = (X_{i,k})^\top_{i\in\{1,2,\dots, n\}| D_i=d} $. Let us also define a ($m^d \times p$)-matrix $\widetilde{\bm{X}}^d$ as $ \widetilde{\bm{X}}^d =
$. We also denote the ($i,j$)-element of $\widetilde{\bm{X}}^d$ as $\widetilde{X}^d_{i, j}$. For a vector $v = (v_1,\dots, v_p)\in\mathbb{R}^p$, we denote its $\ell_q$ norm by $\|v\|_q := (\sum^p_{i=1}|v_i|^q)$ for $q \geq 1$.
This section introduces high-dimensional linear regression models with implicit sparsity.
In this study, we consider a linear relationship between a potential outcome $Y^d$ ($d\in\mathcal{D}$) and covariates $X$. For each $d \in \mathcal{D}$ and $P$, we posit the following linear regression model:
where $\bm{\beta}^d \in \mathbb{R}^p$ is a $p$-dimensional parameter, while $\varepsilon^d$ is an independent noise variable with a zero mean and finite variance. In our analysis, we make the following assumptions about $\varepsilon^d$.
Let $\varepsilon^d_i$ be an i.i.d. copy of $\varepsilon^d$, and $\bm{\varepsilon}^d$ be $(\varepsilon^d_1\ \cdots\ \varepsilon^d_n)^\top$.
We also allow for our linear regression models to be high-dimensional; that is, $p > n$. In such a situation, a common approach to obtaining a consistent estimator is to assume that $\bm{\beta}^d$ has the sparsity, that is, most of the elements of $\bm{\beta}^d$ is zero. However, this study does not assume sparsity directly to potential outcomes; instead, we leverage the property of CATE that emerges from the differentiation between $Y^1$ and $Y^0$.
Our key assumption for the CATE estimation is that the data generating models for $Y^1$ and $Y^0$ have several parameters in common. These common parameters plays an important role for a sparsity-like property, without the explicit sparsity assumption.
Specifically, we suppose that the parameter $\bm{\beta}^d_0$ can be separated into two parameters $\bm{\alpha}^d_0$ and $\bm{\gamma}_0$.
We refer to $\bm{\alpha}^d_0 \in\mathbb{R}^{s_0}$ as an individual parameter and $\bm{\gamma}_0$ as a common parameter. We also give a covariate form $X=(W^\top, Z^\top)^\top$ with subvectors $W \in \mathbb{R}^{s_0}$ and $Z \in \mathbb{R}^{p - s_0}$, that correspond to $\bm{\alpha}^d_0$ and $\bm{\gamma}_0$, respectively. Then, the linear regression model (ref) is rewritten as
under $P_0$. We assume that we do not know which elements in covariates $X$ correspond to $W$ and $Z$.
We firstly consider the following unified linear model derived from the original linear model (ref) for $d \in \mathcal{D}$:
where
Here, $\overline{Y}$ is an unobservable random variable defined using potential outcomes, and $\overline{\varepsilon}$ is an error term that is a composite of $\varepsilon^1$, $\varepsilon^0$, and $D$. It should also be noted that $\mathbb{E}[\overline{\varepsilon}|X] = 0$ almost surely.
This model (ref) demonstrates the implicit sparsity. By considering the difference between $Y^1$ and $Y^0$ and using (ref), we can eliminate the term $Z^\top\bm{\gamma}$ from the regression model. Consequently, we obtain the following:
Note that the linear regression model in (ref) only has a $s_0$-dimensional parameter $\bm{\alpha}^1_0 - \bm{\alpha}^0_0$. This situation is regarded that the $p$-dimensional parameter $\bm{\beta}_0$ in the model (ref) has only $s_0$ non-zero elements, and the other elements are zero which is ignored. In other words, we can regard $\bm{\beta}_0$ in the model (ref) has the sparsity in spite that we do not assume the sparsity on $\bm{\beta}^1_0$ and $\bm{\beta}^0_0$. We refer to this property as the implicit sparsity.
The Lasso tibshirani96regression is an estimation method with a regularization that adds the $\ell_1$-norm penalty on the parameters. In this section, we propose the CATE Lasso estimator, which utilizes the implicit sparsity introduced in Section (ref). Unlike the standard Lasso penalizing the parameter itself, the CATE Lasso penalizes a difference between two parameters.
We focus on the fact that for each $d\in\mathcal{D}$, the following linear regression model holds under $P$:
where $\widetilde{\bm{\varepsilon}}^d$ is an $m^d$-dimensional vector whose $i$-th element $\widetilde{\bm{\varepsilon}}^d_i$ is equal to $\varepsilon^d_j$ such that $j = \min\{k\in[N]| \sum^k_{s=1} \mathbbm{1}[D_s = d] = i\}$.
Let $\bm{\beta}_0$, $\bm{\alpha}_0$ and $\bm{\gamma}_0$ be the parameter of distributions that generate observations; that is, they are the true parameters of $\bm{\beta}$, $\bm{\alpha}$ and $\bm{\gamma}$ under $P_0$. Then, under the linear regression models, for $x = \left(w, z\right)$, where $w\in\mathcal{W}$ and $z\in \mathcal{Z}$, the CATE $f_0(x)$ can be rewritten as
We estimate $\bm{\beta}_0$ by using the least squares with $\ell_1$-penalty. Because we cannot observe $Y^1_i - Y^0_i$, we instead minimize two squared losses defined between $Y^1_i$ and $\widetilde{X}^\top_i \bm{\beta}^1$ and between $Y_i^0$ and $\widetilde{X}^\top_i \bm{\beta}^0$. Unlike the standard Lasso, we penalize $\bm{\beta}^1 - \bm{\beta}^0$, not $\bm{\beta}^1$ and $\bm{\beta}^0$, separately. We refer to our Lasso as the CATE Lasso.
First, we consider a estimation strategy for CATEs under the implicit sparsity. In other words, we consider the regularization for the difference between $\bm{\beta}^1$ and $\bm{\beta}^0$ to utilize the implicit sparsity. Let $\widehat{\bm{\beta}}^1$ and $\widehat{\bm{\beta}}^0$ be estimators of $\bm{\beta}_0^1$ and $\bm{\beta}_0^0$, which are defined as
where $\lambda > 0$ is a penalty coefficient and $\delta\in(0, 1)$ is a balancing weight.
In this strategy, we regularize a difference $\bm{\beta}^1 - \bm{\beta}^0$, unlike the standard Lasso. This regularization allows us to employ the sparsity in $\bm{\beta}^1 - \bm{\beta}^0$, while we can assume that $\bm{\beta}^d_0$ is not sparse as well as the standard Lasso. Similar estimators have been proposed in the literature of Lasso, such as the fused Lasso Tibshirani2005 and $\ell_1$-trend filtering Kim2009.
In the optimization problem (ref), there can be multiple solutions unlike the standard Lasso because we only impose a regularization for $\bm{\beta}^1 - \bm{\beta}^0$. Therefore, this section develops a uniquely determined estimator of $\bm{\beta}^0$. We refer to this estimator as the CATE Lasso estimator, a Lasso specialized for CATE estimation.
In preparation, we consider the following optimization problem, which is equivalent to the problem (ref):
where $\bm{\beta}$ is a parameter which represents $\bm{\beta}_0 = \bm{\beta}^1_0 - \bm{\beta}^0_0$. This problem is mathematically equivalent to our previous definition of $(\widehat{\bm{\beta}}^1, \widehat{\bm{\beta}}^0)$ in (ref), since we can achieve $\widehat{\bm{\beta}}^1 = \widehat{\bm{\beta}} + \widehat{\bm{\beta}}^0$ by the estimators $(\widehat{\bm{\beta}}, \widehat{\bm{\beta}}^0)$ in (ref). Then, we estimate $\bm{\beta}_0$ with the following procedure.
\paragraph{(I) Interpolating estimator of $\bm{\beta}^0_0$.} First, we estimate $\bm{\beta}^0_0$ as
The estimator $\widehat{\bm{\beta}}^0$ has the following analytical solution:
where $ \{(\widetilde{\bm{X}}^0)^\top\widetilde{\bm{X}}^0\}^{\dagger}$ is a pseudoinverse of $(\widetilde{\bm{X}}^0)^\top\widetilde{\bm{X}}^0$. Such an estimator is referred to as an interpolating estimator or a minimum-norm estimator Bartlett2020 because $\widehat{\bm{\beta}}^0$ is a parameter with the smallest $L^2$ norm $\bm{\beta}^0$ among ($\|\bm{\beta}^0\|_2$) that perfectly fit $\mathbb{Y}^0$ with $p \geq n$. Note that we cannot identify $\bm{\beta}^0_0$ itself even if using this estimator; however, as shown in Section (ref), we can use it for estimating $\bm{\beta}_0$.
\paragraph{(II) Lasso for estimating $\bm{\beta}_0$.} Using the interpolating estimator $\widehat{\bm{\beta}}^0$ and the Lasso, we estimate $\bm{\beta}_0$ as
for some $\lambda > 0$. This optimization can be solved by the standard algorithm (e.g. the coordinate descent method) used for the Lasso. Thus, we obtain the CATE Lasso estimator $\widehat{\bm{\beta}}$ for $\bm{\beta}_0$.
Note that the choice of $\delta$ does not affect the estimation (we can include $\delta$ into $\lambda$). However, in practice, we may stabilize our method by a suitable choice of $\delta$.
\paragraph{(III) Estimator for CATE $f_0(x)$.} Using the CATE Lasso estimator $\widehat{\bm{\beta}}$, we estimate the CATE $f_0(x)$ for each $x\in\mathcal{X}$ as
To the best of our knowledge, such a type of estimator is novel in the literature of CATE estimation.
This section provides several theoretical properties of our CATE Lasso estimator. Our analysis is inspired by vandeGeer2014.
First, we consider a case where $\bm{X}$ is non-random (fixed-design), $\varepsilon^d$ follows a sub-Gaussian distribution, and $p$ is fixed. We show more detailed arguments about the analysis for a fixed design in Appendix (ref). Based on the results of the fixed design, we show consistency of $\widehat{\bm{\beta}}$ under a random design, where $\bm{X}$ is random, $\varepsilon^d$ is non-Gaussian, and $p = p_n \to \infty$ as $n\to\infty$.
To derive a tight upper bound for the estimation error, we follow Buhlmann2011 in asserting that the compatibility condition plays a crucial role in identifiability. To define the compatibility condition, for a $p\times 1$ vector $\bm{\beta}$ and a subset $\mathcal{S}_0 \subseteq \mathcal{P}$, we introduce a notation that denotes the sparsity of $\bm{\beta}_0$. Let $\mathcal{S}_0 \subseteq \mathcal{P}$ be a set such that
where $\beta_{0, j}$ is the $j$-th element of $\bm{\beta}_0$. For an index set $\mathcal{S}_0 \subset \mathcal{P}$, let $\bm{\beta}_{\mathcal{S}_0}$ and $\bm{\beta}_{\mathcal{S}^c_0}$ be vectors whose $j$-th elements are defined as follows:
respectively, where $\beta_j$ is the $j$-th element of $\bm{\beta}$. Therefore, we obtain $\bm{\beta} = \bm{\beta}_{\mathcal{S}_0} + \bm{\beta}_{\mathcal{S}^c_0}$.
Note that from the definition of the individual and common parameters, it holds that
that is, $j (\in \mathcal{S}_0)$-th element of $\bm{\beta}^d$ corresponds to $\bm{\alpha}^d$ and $j (\in \mathcal{S}^c_0)$-th element of $\bm{\beta}^d$ corresponds to $\bm{\gamma}$. Here, note that $s_0 = |\mathcal{S}_0|$, where recall that $s_0$ is defined as the dimension of $\bm{\alpha}$ in Section (ref).
Let $\widehat{\Sigma}^1 := (\widetilde{\bm{X}}^1)^\top \widetilde{\bm{X}}^1 / n$. Here, we define the compatibility condition.
Let us denote the diagonal elements of $\widehat{\Sigma}^1$, by $\left(\widehat{\sigma}^1_j\right)^2 := \widehat{\Sigma}^1_{j, j}$, for $j\in\mathcal{P}$. Using the compatibility condition, we make the following assumption.
Furthermore, we make the following assumption.
The matrix $\{(\widetilde{\bm{X}}^0)^\top \widetilde{\bm{X}}^0\}^\dagger(\widetilde{\bm{X}}^1)^\top \widetilde{\bm{X}}^1$ represents a discrepancy between $\widetilde{\bm{X}}^0$ and $\widetilde{\bm{X}}^1$. When $\widetilde{\bm{X}}^0$ and $\widetilde{\bm{X}}^1$ are identical or having common eigenvectors with similar eigenvalues, the matrix is approximately an identity, hence Assumption (ref) is satisfied. It represents the correlation between treatments and covariates, thus explains the structure specific to CATE.
Then, we derive an oracle inequality and consistency result for the CATE Lasso estimator. The proof is shown in Appendix (ref).
This theorem implies that under a proper $t$ and penalty $\lambda$ (oracle), we can show the convergence of the estimator $\widehat{\bm{\beta}}$ to the true value $\bm{\beta}_0$ with high probability.
Let $\Sigma^1$ be the population of $\widehat{\Sigma}^1$. Based on the results of a fixed-design, we show the consistency of $\widehat{\bm{\beta}}$ under a random design where the covariates $X_i$ and treatment assignments $D_i$ are random and $p = p_n \to \infty$ as $n\to \infty$. By assuming $\widetilde{X}^d$ is sub-Gaussian, we can obtain the following theorem.
This result implies consistency of $\widehat{\bm{\beta}}$; that is, $\widehat{\bm{\beta}} \xrightarrow{\mathrm{p}}\bm{\beta}_0$.
Note that although these results indicate consistency of $\widehat{\bm{\beta}}$, it does not ensure the convergence of $\widehat{\bm{\beta}}^1$ and $\widehat{\bm{\beta}}^0$ to $\bm{\beta}^1_0$ and $\bm{\beta}^0_0$, respectively. In fact, since we only regularize $\bm{\beta}^1 - \bm{\beta}^0$, the minimizers $\widehat{\bm{\beta}}^1$ and $\widehat{\bm{\beta}}^0$ of the objective function might not be unique. This finding suggests that even if we can not estimate $\bm{\beta}^1_0$ and $\bm{\beta}^0_0$ consistently, it is still possible to consistently estimate $\bm{\beta}_0$ under the implicit sparsity assumption.
From Theorem (ref), we obtain the following corollary.
We used the Cauchy-Schwarz inequality as $ \big|x^\top\widehat{\bm{\beta}} - f_P(x) \big| \leq \|x\|_2\|\widehat{\bm{\beta}} - \bm{\beta}_0\|_2$, and $\|\widehat{\bm{\beta}} - \bm{\beta}_0\|_2 \leq \|\widehat{\bm{\beta}} - \bm{\beta}_0\|_1$.
This section briefly introduces related work. More details and open issues, including an extension to the debiased Lasso, are discussed in Appendix (ref).
Early work on CATEs are Heckman1997 and Heckman2005. With the rise of machine learning algorithms, various methods for CATE estimation have been proposed. Some of them are summarized as meta-learners by Kunzel2019 (see the following section). There is a stream of work that employs neural networks Johansson2016,Shalit2017,Shi2019,Hassanpour2020Learning,curth2021nonparametric,curth2021inductive, utilizing methods and properties of neural networks, such as representation learning bengio2014representation and multi-task learning Caruana1997. yoon2018ganite applies the generative adversarial nets for CATE estimation. Furthermore, methods that utilize Gaussian processes Alaa2017,Alaa2018, deep kernel learning Zhang2020, boosting, tree-based methods Zeileis2008,Su2009,ImaiSt2011,Kang2012,Lipkovich2011,loh2012,Wager2018,Athey2019,Chatla2020, nearest neighbor matching, series estimation, and Bayesian additive regression trees have been developed Hill2011. Numerous methods employing machine learning algorithms have also been proposed Li2017,Kallus2017,Powers2017,Subbaswamy2018,Zhao2019,Nie2020,Hahn2020.
Certain CATE estimators can be categorized into meta-learners. Representatives are listed below:
Many methods for CATE estimation can be categorized into a meta-learner. For instance, our CATE Lasso is an instance of the T-learner.
The IPW-learner and DR-learner extend the IPW estimator Horvitz1952 and DR estimator bang2005drestimation, originally proposed for ATE estimation, to CATE estimation.
In the context of model selection, the IPW-learner and the DR-learner have garnered attention as in schuler2018comparison and Saito2020. ninomiya2022information and ninomiya2021selective combine the IPW-learner with the Lasso. As existing studies have pointed out, the advantages of the IPW-learner stem from the unbiasedness to the risk for $Y^1 - Y^0$, while the T-learner combines two risks for $Y^1$ and $Y^0$ separately. However, the IPW-learner requires the true value for the propensity score $p(D = 1 | X)$. When it is unknown and replace it with an estimate, we cannot enjoy the unbiasedness. Furthermore, under high-dimensional models, the estimation of $p(D = 1 | X)$ itself becomes problematic because we need to assume something like sparsity for estimating $p(D = 1 | X)$ with high-dimensional $X$. The DR-learner further requires an estimate of $\mathbb{E}[Y^d|X]$ to estimate $\mathbb{E}[Y^1- Y^0|X]$, in addition to $p(D = 1 | X)$. In contrast, our method does not suffer from the problem.
To verify the soundness of the CATE Lasso, we exhibit simulation studies in this section. We compare our method with the T-Learners using the OLS and Lasso.
First, for regression models' parameters $\bm{\beta}^d$ for $d\in\mathcal{D}$, we generate each element of its first $s_0$ elements $\bm{\beta}^d_1, \dots, \bm{\beta}^d_{s_0}$ from the uniform distribution whose support is $[-10, 10]$, setting the other elements are zero, $\bm{\beta}^d_{s_0 + 1}= \cdots = \bm{\beta}^d_{p} = 0$. We also generate $\bm{\theta}$ from the uniform distribution with a support $[-1, 1]$, which is used to construct $p(D = 1| X)$.
Let $n$, $p$, and $s_0$ be the sample size, dimension of $X$, and sparsity parameter, defined later. Then, we generate $\{(X_i, D_i, Y_i)\}^n_{i=1}$ as follows. We generate $V_i$ from the $p-1$-dimensional standard normal distribution and define $X_i = (1\ \ V^\top_i)^\top$. For $d\in\mathcal{D}$, let us generate $\epsilon^d_i$ from the standard distribution. Then, we obtain $Y^d_i$ as $Y_i = X^\top_i\bm{\beta}^d$. We also construct $p(D = 1 | X = X_i)$ as $p(D = 1| X = X_i) = \frac{1}{1 + \exp(-X^\top_i \bm{\theta} + \eta_i)}$, where $\eta_i$ is generated from the standard distribution. From a Bernoulli distribution with a parameter $p(D = 1 | X = X_i)$, we obtain $D_i \in \mathcal{D}$. Thus, we obtain $(X_i, D_i, Y_i)$ and obtain $\{(X_i, D_i, Y_i)\}^n_{i=1}$ by generating $(X_i, D_i, Y_i)$ $n$ times independently.
By using $\{(X_i, D_i, Y_i)\}^n_{i=1}$, we predict $Y^1_i - Y^0_i$ given $X_i$. We conduct experiments for $n = 500$, $p = 300, 1000$, and $s_0 = 10, 25, 50$. When $p = 300$, $p > n$ does not hold, but we use this setting for confirming the effectiveness of the CATE Lasso.
We conduct each experiment $100$ times and show the root mean squared errors of the CATE Lasso, the OLS, and the Lasso in Figure (ref) and (ref).
We can confirm that the CATE Lasso shows preferable performances in all cases. Because we do not assume sparsity for each potential outcome, the Lasso does not perform well. As $s_0$ increases, the performance of the CATE Lasso approaches to the OLS, which is an expected behavior because as $s_0$ increases, we cannot exploit the implicit sparsity.
We also show the additional results, such as experiments with the IPW-Learner, and a semi-synthetic datasets in Appendix (ref).
We examined CATE estimation using a high-dimensional linear regression model. By assuming implicit sparsity, which arises from the difference between two potential outcome linear regression models, we proposed the CATE Lasso estimators. For these estimators, we demonstrated several theoretical properties, such as consistency. Subsequently, we presented experimental results to validate our proposed estimators. Our proposed estimators represent novel, simple, and practical approaches for CATE estimation. An open issue for future work is confidence intervals and the semiparametric efficiency of these estimators.
\onecolumn