Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
43,369 characters · 7 sections · 50 citation commands
Sparsity Double Robust Inference of Average Treatment Effects
Average treatment effect estimation is a core problem in causal inference, and has been the topic of a considerable amount of recent literature imbens2015causal. In this paper, we focus on the task average treatment effect estimation with high-dimensional confounders: We have access to $n$ i.i.d. samples $(X_i, \, Y_i, \, W_i) \in \xx \times \mathbb{R} \times \cb{0, \, 1}$, where $X_i$ denotes high-dimensional pre-treatment features ($\xx \subset \mathbb{R}^p$ with $p \gg n$), $W_i$ is the treatment assignment, and $Y_i$ is our outcome of interest. Causal effects are defined via potential outcomes $\cb{Y_i(0), \, Y_i(1)}$, such that we observe $Y_i = Y_i(W_i)$ and the average treatment effect is defined as $\tau = \mathbb{E}{Y_i(1) - Y_i(0)}$ neyman1923applications,rubin1974estimating. Finally, we assume that there are no unmeasured confounders, i.e., the treatment assignment $W_i$ may not be randomized, but can be treated as such once we control for $X_i$, i.e., $\cb{Y_i(0), \, Y_i(1)} \indep W_i \cond X_i$ rosenbaum1983central. Throughout, we also assume overlap, such that $\eta \leq \mathbb{P} \p{W_i \cond X_i = x} \leq 1 - \eta$ for all $x$ and some $\eta > 0$.
In the low-dimensional case, one of the most prominent approaches to average treatment effect estimation is via augmented inverse-propensity weighting robins1994estimation,
where $e(x) = \mathbb{P} \p{W_i \cond X_i = x}$ is the propensity score, $\mu_{(w)}(x) = \mathbb{E} \p{Y_i(w) \cond X_i = x}$ are conditional response surfaces, and the quantities above with hats are estimates thereof. A celebrated property of this estimator is that it is double robust, meaning that it is consistent whenever either \smash{$\he(x)$} or the \smash{$\hat{\mu}_{(w)}(x)$} are consistent scharfstein1999adjusting. Moreover, \smash{$\hat{\tau}$} is $\sqrt{n}$-consistent and semiparametrically efficient whenever the following risk bounds hold farrell2015robust
This statement is not sensitive to the structure of the estimators \smash{$\he(x)$} or the \smash{$\hat{\mu}_{(w)}(x)$} provided we use an appropriate type of sample splitting chernozhukov2016double,zheng2011cross, and thus allows for considerable methodological flexibility. For example, farrell2018deep establish conditions under which (ref) holds when \smash{$\he(x)$} or the \smash{$\hat{\mu}_{(w)}(x)$} are fit using neural networks. These results on augmented inverse-propensity weighting can also be applied when $X_i$ is high dimensional; however, in this case, the required risk bound can be difficult to satisfy. In particular, except in extreme cases, the condition (ref) effectively requires both $\mu_{(w)}(x)$ and $e(x)$ to admit very sparse representations.
In this paper, we study a doubly robust construction that is specifically designed for the high-dimensional case, and can be used for valid inference of $\tau$ under substantially weaker sparsity assumptions than standard augmented inverse-propensity weighting. We focus on the case where $\mu_{(w)}(x)$ and $e(x)$ have a high dimensional linear-logistic specification (we omit intercepts for conciseness of presentation),
and consider an estimator that is $\sqrt{n}$-consistent for $\tau$ under the condition that either $\theta$ or the $\beta_{(w)}$ (but not necessarily both) satisfy the type of sparsity condition that is usually required for high-dimensional inference javanmard2014confidence,van2014asymptotically,zhang2014confidence. We refer to this property as sparsity double robustness.
The issue of sparsity doubly robustness has been an open question since the recent development of high-dimensional inference. This literature requires sparsity level $o(\sqrt{n}/\log p)$ for inference, a condition stronger than $o(n/\log p)$ needed for consistent estimation. Such a gap has only been addressed very recently in javanmard2015biasing, who found that the sparsity level of only one parameter needs to satisfy $o(\sqrt{n}/\log p)$, not both. However, their work only addresses the linear models and heavily relies on the Gaussianity assumption of the design. In this paper, we show that such sparsity doubly robustness result holds true for nonlinear models without Gaussian designs.
Our method starts with a functional form that closely resembles (ref). However, we choose our estimators of $\mu_{(w)}(x)$ and $e(x)$ in ways that carefully exploit the geometry of sparseness in (ref) and are thus able to improve on its performance. A closely related estimator has been independently studied by tan2018model, who considered potentially misspecified models but did not provide results on sparsity doubly robustness. Our main construction is as follows, modulo some algorithmic tweaks (including a type of sample splitting):
As discussed in Section (ref), we can study this estimator from two different perspectives. If $\beta_{(w)}$ is very sparse, then the solution to (ref) converges at a fast rate, while the solution to the propensity model (ref) effectively debiases \smash{$\hat{\beta}_{(w)}$} even if \smash{$\hat{\theta}_{(w)}$} is not particularly accurate. Meanwhile, if $\theta_{(w)}$ is very sparse, then the converse holds. Our proof exploits this idea to establish sparsity double robustness.
The idea of fitting a propensity model that can also leverage the shape of the conditional response surface has generated considerable interest in recent years. The key observation here is that, in addition to being a consistent estimator when $\theta$ is very sparse, (ref) also “balances” the inverse-propensity weighted features among the treated and control samples in finite samples chan2015globally,hainmueller,imai2014covariate,tan2017regularized,zhao2016covariate
The advantage of balancing is that, if the linear model for $Y$ is well specified, then balancing as in (ref) is sufficient for eliminating confounding, even when \smash{$\hat{\theta}_{(w)}$} itself may be inconsistent or misspecified athey2016approximate,hirshberg2017balancing,kallus2018balanced,zhao2017entropy,zubizarreta2015stable. Note that, here, we estimate separate models for \smash{$\mathbb{P} \p{W_i = 0 \cond X_i = x}$} and \smash{$\mathbb{P} \p{W_i = 1 \cond X_i = x}$}, parametrized by \smash{$\theta_{(0)}$} and \smash{$\theta_{(1)}$} respectively. This parametrization is based on ((ref)) and reads $$\mathbb{P} \p{W_i=w\cond X_i=x}=1/(1+\exp(-x'\theta_{(w)})) \qquad \text{for}\qquad w\in\{0,1 \}. $$ Notice that by ((ref)), we have that $\theta_{(1)}=\theta $ and $ \theta_{(0)}=-\theta $. Asymptotically, we expect both parameter vectors to be consistent, \smash{$-\hat{\theta}_{(0)}, \, \hat{\theta}_{(1)} \approx \theta$}, but finite-sample differences between \smash{$\hat{\theta}_{(0)}$} and \smash{$\hat{\theta}_{(1)}$} play a key role in enabling the balance imai2014covariate.
Our main finding is that an estimator constructed via the above “balancing” principle achieves sparsity double robustness, meaning that it attains $\sqrt{n}$-consistency given strong enough sparsity assumptions $o(\sqrt{n}/\log p)$ on either $\theta$ or the $\beta_{(w)}$, but not necessarily both. As discussed further below, this property is considerably stronger than the standard double robustness property (ref) in the high-dimensional setup (ref).
Double robust and/or semiparametrically efficient estimation has a long tradition in the literature on causal inference chernozhukov2016double,farrell2015robust,hahn1998role,hirano2003efficient,newey2018cross, robins1,robins1994estimation,scharfstein1999adjusting,tan2010bounded,van2006targeted. More recently, it has been shown that with high dimensional confounders, we can improve the behavior of double-robust-type estimators by having them directly exploit the geometry of sparsity.
As one of the first result in this direction, athey2016approximate showed, given sufficient sparsity on the outcome function in (ref), $\lVert \beta_{(w)} \rVert_0 \ll \sqrt{n} / \log(p)$, we can achieve $\sqrt{n}$-consistency without any assumptions on the propensity score beyond overlap by simply using weights that balance moments as follows (the \smash{$\hat{\beta}_{(w)}$} are estimated via the lasso):
Conceptually, this approach is related to several papers that stress the important of covariate balance for accurate estimation of treatment effects chan2015globally,imai2014covariate,kallus2018balanced,zhao2016covariate,zubizarreta2015stable. hirshberg2017balancing establish conditions under which this estimator is efficient.
The main downside of the approximate residual balancing estimator (ref) is that it always requires sparsity of the outcome model, and cannot use a well specified and sparse propensity model to compensate for a complex outcome model. Our sparsity double robustness result, which only requires strong sparsity of either $\theta$ or the $\beta_{(w)}$ in (ref) directly addresses this limitation; and, as shown in our experiments, yields substantial gains in accuracy when $\theta$ is in fact sparse.
Our result is most closely related to a recent proposal by chernozhukov2018double, who studied any linear functional whose Riesz representer admits an (approximate) linear representation. In another paper, chernozhukov2016double considers theoretical results for estimators based on learning conditional mean function and the propensity score. In both papers, the key condition is that the product of $\ell_2$-loss for learning the two nuisance parameters is $o(n^{-1/2})$, a condition referred to as rate double robustness; see Definition 2 in Smucler2019unifying. Sufficient conditions for rate double robustness have been provided in these works in terms of sparsity levels. For example, Remark 5.2 of chernozhukov2016double shows that rate double robustness is guaranteed when the product of two sparsity levels is $o(n)$, while Remark 7 of chernozhukov2018double points out that under the assumption of bounded $\ell_1$-norm of both parameters, rate double robustness holds whenever one of the sparsity levels is $o(\sqrt{n}/\log p)$.
The sparsity doubly robustness in this paper contributes to the literature by providing a different perspective. We show that efficient estimation is also possible in certain cases in which rate double robustness might not hold. One such example is when the logistic parameter has bounded $\ell_1$-norm and has sparsity level $o(\sqrt{n}/\log p)$ and the conditional parameter has sparsity level $o(n^{3/4}/\log p)$ with potentially large $\ell_1$-norm. In this example, we can still derive $1/\sqrt{n}$-consistency although we are not aware of any results that can guarantee rate double robustness.
In addition, our work is also different from chernozhukov2018double in terms of specification. In the context of average treatment effect estimation, the formulation in chernozhukov2018double means that we need there to exist (potentially sparse) vectors $\xi_{(0)}$ and $\xi_{(1)}$ whose $\ell_1$-norms are bounded (see Definition 3 or 4 therein) as well as such that $|1/(1 - e(x)) - x' \xi_{(0)}| \approx 0$ and $| 1/e(x) -x' \xi_{(1)}|\approx 0$ uniformly across $x$. This may be a reasonable assumption if $x$ was in fact constructed as a basis expansion of some simpler measured features; however, it appears to be difficult to justify more generally. One contribution of this paper relative to chernozhukov2018double is that we achieve sparsity double robustness using the natural linear-logistic specification (ref).
We also note two recent papers that consider estimators that resemble ours. ning2017high consider an estimator that, in the spirit of belloni2014inference, first fit a penalized covariate-balancing propensity model, and then re-fit without penalty those coefficients that correspond to features that are relevant to outcome modeling. Meanwhile, tan2018model augments a penalized covariate-balancing propensity model in an outcome regression; it turns out that his covariate-balancing mechanism designed to address the issue of misspecification is also helpful for relaxing sparsity requirements. Neither paper, however, achieves sparsity double robustness as discussed here; rather, they require both the outcome parameter vector $\beta$ and the propensity parameter vector $\theta$ to be ultra-sparse---or, if there is misspecification they require the population minimizers of both the outcome and propensity loss functions to be ultra-sparse. Under the framework of Smucler2019unifying,Rotnizky2019mixed, ning2017high,tan2018model are classified as examples of model double robustness, which means that one of the models (either conditional mean or propensity score) is misspecified. Rate double robustness requires that the product of the $\ell_2$-norms of the estimation errors in two models is of the order $o(n^{-1/2})$.
Whenever a parameter is identified through a moment condition, like (ref), a direct loss minimization that does not take into account this moment condition may not guarantee desirable properties. Controlling inferential features of high-dimensional estimates is extremely difficult; most, if not all, require strict sparsity conditions. We aim to control optimality at estimation by directly embedding the leading term of the bias into a constraint of newly designed estimators.
The main idea behind our construction is that we use estimators \smash{$\hat{\beta}_{(0)}$}, etc., of $\beta_{(0)}$, etc., that have two complementary properties. When the underlying parameter $\beta_{(0)}$ is ultra-sparse, then \smash{$\hat{\beta}_{(0)}$} converges to $\beta_{(0)}$ in $\ell_1$-norm. Furthermore, even when $\beta_{(0)}$ is not ultra-sparse, \smash{$\hat{\beta}_{(0)}$} still has a useful covariate-balancing property implied by its Karush-Kuhn-Tucker (KKT) conditions that can be put to good use (and similar guarantees hold for \smash{$\hat{\beta}_{(1)}$}, \smash{$\hat{\theta}_{(0)}$} and \smash{$\hat{\theta}_{(1)}$}). We then provide two separate consistency and asymptotic normality proofs for our estimator: One that assumes that $\beta_{(w)}$ is ultra-sparse and relies on KKT conditions for the \smash{$\hat{\theta}_{(w)}$} estimator to debias a very accurate \smash{$\hat{\beta}_{(w)}$} estimator, and a second that assumes that $\theta$ is ultra-sparse and relies KKT conditions for the \smash{$\hat{\beta}_{(w)}$} estimator to debias a very accurate \smash{$\hat{\theta}_{(w)}$} estimator. Of course, only one of these arguments needs to hold for us to achieve asymptotic normality, and thus our estimator is sparsity double robust. This argument was inspired by the one used by chernozhukov2018double; however, as discussed in the related works section, chernozhukov2018double make the somewhat unusual assumption that $1/e(x)$ can be approximated by a sparse linear model (rather than the assumption we make here, i.e., a sparse logistic model for $e(x)$).
In this section, we briefly sketch the argument behind our main formal result, and use it to motivate the form of our estimator. One of the main ingredients is moment targeting: We design estimators such that they satisfy certain moment conditions that help reduce the bias at estimation. The construction for estimators for $\beta_{(w)}$ and $\theta_{(w)} $ is based on the structure of the bias in the final estimator for $\mathbb{E} \mu_{(w)}(X_i) $. We emphasize that the argument here is only heuristic; formal arguments are given in the appendix.
Given these preliminaries, observe that the treatment effect estimator under consideration can be written in a familiar form \[ \hat \tau = \hat \mu_{(1)} - \hat \mu_{(0)}, \] where $\hat\mu_{(w)}$ is an estimate of $\mu_{(w)} = \mathbb{E}{Y_i(w)}$. We use \[ \hat{\mu}_{(w)}=\frac{1}{n}\sum_{i=1}^{n} X_{i}' \hat{\beta}_{(1)} + \hat \gamma_ i (w) \mathds{1}\{W_i=w\} (Y_i - X_{i}' \hat{\beta}_{(w)}), \ \ \ \ \hat \gamma_i (w) = 1 + \exp(- X_i'\hat{\theta}_{(w)}), \] and construct \smash{$\hat{\mu}_{(0)}$} analogously. Our goal is to choose $\hat{\theta}_{(w)}$ (and hence $\hat{\gamma}_{i}(w)$) such as to control the errors \smash{$\hat{\mu}_{(w)} - \mu_{(w)}$} under flexible sparsity conditions.
To motivate our choice of $\hat{\theta}_{(1)}$, let us first consider the case where $\beta_{(1)}$ is very sparse, i.e., $\|\beta_{(1)}\|_0\ll \sqrt{n}/\log p $. Notice that
where $\varepsilon_{i,(w)}=Y_i(w)-X_i'\beta_{(w)} $. The first two terms on the right hand side are asymptotically normal with mean zero under weak consistency conditions on $\hat{\theta}_{(1)}$ that only require a moderate amount of sparsity on $\theta$. Meanwhile, the last term can be bounded using Holder's inequality,
Under sparsity assumption $\|\beta_{(1)}\|_{0}=o(\sqrt{n}/\log p)$, we can typically obtain $\|\hat{\beta}_{(1)}-\beta_{(1)}\|_{1}=o_P(1/\sqrt{\log p})$ via sparse methods negahban2012unified. Meanwhile, the KKT conditions for the estimator in (ref) with $w=1$ automatically yields tan2017regularized $$ \left\Vert n^{-1}\sum_{i=1}^{n}\left[1-W_{i}(1+\exp(-X_{i}'\hat{\theta}_{(1)}))\right]X_{i} \right\Vert_{\infty}=O_P(\sqrt{n^{-1}\log p}), $$ thus bounding the bias to the order of $o_P(n^{-1/2})$. This is the first example of moment targeting. The KKT condition of the estimator provides a convenient moment condition for the purpose of bias reduction.
The above argument closely mirrors the argument used by athey2016approximate to obtain $\sqrt{n}$-consistent estimates of $\tau$ when $\|\beta_{(w)}\|_0\ll \sqrt{n}/\log p$. The main difference with our approach is that athey2016approximate do not fit a model for $\theta$, but instead directly optimize the weights \smash{$\hat{\gamma}$} via quadratic programming as in javanmard2014confidence and zubizarreta2015stable. That in turn, leads to somewhat loss of flexibility whenever the outcome model is not sparse.
Here, the fact that we also model $\theta$ enables us to alternatively exploit sparsity in $\theta$ and correspondingly relax assumptions on $\beta_{(1)}$. To do so, note that
Again, the sum of the first three terms above is asymptotically Gaussian on the $\sqrt{n}$-scale under only weak assumptions on $\hat{\beta}_{(1)}$. To handle the last term, we can use Taylor expansion to argue that (we will make this rigorous in the proof of our main result)
Now, given sufficient sparsity on $\theta$, i.e., $\|\theta\|_0\ll \sqrt{n}/\log p $ we can verify that $\|\hat{\theta}_{(1)}-\theta \|_1=o_P(1/\sqrt{\log p})$. Meanwhile, the the KKT condition for the estimator $\hat{\beta}_{(1)}$ in (ref) automatically yields that the first component above is $O_P(\sqrt{n^{-1}\log p})$. Thus, we also expect $\hat{\mu}_{(0)}$ to be accurate when $\theta$ is very sparse, even if $\beta$ is not. This is another example of moment targeting in that the KKT condition for $\hat{\beta}_{(1)} $ again provides a convenient bound for bounding the bias.
The above discussion provides some helpful conceptual guidance on how to pick good estimators of the unknown $\beta_{(w)}$ and $\theta_{(w)}$. To achieve optimality in most general terms, we invoke a special scheme of sample-splitting similar to cross-fitting. Under the usual cross-fitting scheme, the influence function is evaluated on observations that are not used to estimate the nuisance parameters (in our case $\beta_{(w)}, \theta_{(w)} $). Cross-fitting has been used to reduce bias terms in many semiparametric and high-dimensional models chernozhukov2016double,newey2018cross,schick1986asymptotically,zheng2011cross. Here, our approach requires us to only cross-fit $\hat \beta_{(w)}$, but not $\hat \theta_{(w)}$.
The entire sample is divided into two parts $\mathcal{J}$ and $\mathcal{J}^c$. For $F\in\{\mathcal{J},\mathcal{J}^c,\}$, estimators trained using the sample $F$ are denoted with $\hat \beta_{(w),F}$ and $\hat \theta_{(w),F}$, respectively. For expositional simplicity, we assume that $|\mathcal{J}|=|\mathcal{J}^c|=n/2 $. Then, for $ (w,F)\in \{0,1\}\times \{\mathcal{J},\mathcal{J}^c \}$, we define the estimator of the mean
where the weight function is defined in-sample \[ \hat \gamma_i (w, F ) = 1 + \exp(- X_i'\hat{\theta}_{(w),F}) , \] $\hat \theta_{(w), F}$ is defined in Algorithm (ref), and $\hat \beta_{(w),F}$ is given by
Algorithm (ref) presents details of the propensity estimation. The loss functions in (ref) and (ref) were recently utilized in tan2018model but the proposed average treatment effects estimator therein does not achieve sparsity double robustness.
The method presented here splits the sample into two subsamples. Although one can easily follow the same principle and split the sample into multiple subsamples, we do not pursue this option here for notational simplicity. We now define \[ \hat \mu _{(1)} =(\hat{\mu}_{(1),\mathcal{J} }+\hat{\mu}_{(1),\mathcal{J}^c })/2. \] Similarly, we can define $\hat \mu _{(0)} $. Then, the average treatment effect estimator is defined as
We now turn to a formal characterization of the average treatment effect estimator in ((ref)), with the aim of providing asymptotic Gaussianity whenever one, but not both, of $\theta$, $\beta$ is estimated consistently. We begin by listing some theoretical assumptions necessary for the development of the theoretical guarantees.
First, we assume that the covariate space and the parameter space are both subsets of Euclidean space; specifically, we assume that $X \in [a,b]^p$ and $\theta \in \mathcal{B}_1 (r) \subset \mathbb{R}^p$ for some bounded $r>0$, where $\mathcal{B}_1(r)$ is an $\ell_1$ ball with radius $r$.
The results discussed below hold whenever, the tuning parameters (in Algorithm (ref)) are chosen to be proportional to $\sqrt{\log(p)/n}$. Moreover, we also assume that $\kappa$ is chosen to be larger than $\| \theta_{(w)}\|_1$. Our procedure is not particularly sensitive to the choice of $\kappa$; in practice it suffices to choose a large enough number.
Our next assumption controls the regularity properties of the errors within both models (ref). Let $\varepsilon_{i,(w)} = Y_i - X_i'\beta_{(w)}$ and $v_{i,(w)} = \mathds{1}\{W_i=w\} - e_{(w)}(X_i)$. Note that in the context of models (ref) the unconfoundedness assumption implies $\varepsilon_{i,(w)} \perp v_{i,(w)} | X_i$ and from now on we will work with this slightly weaker assumption.
Now, observe that Assumption (ref) is very weak and in particular it is not implying consistent estimation in the outcome model. The boundedness of $\|X_i\|_\infty$ and $\|\theta\|_1$ guarantees the overlap condition.
Finally, in the context of the average treatment effects, in order to provide confidence intervals an estimate of the asymptotic variance of $\hat \tau$ is needed. We show an asymptotic variance of $\hat \tau$ takes the form of
Observe that $\Omega$ is the variability induced primarily from the variability of the design $X$. The other two terms can be viewed as properly normalized unexplained variance of the models (ref).
To define variance estimates, we define estimates of $\Omega$, $V_{(1)}$ and $V_{(0)}$ separately. We set
as well as
In the above display, $\hat \beta_{(w)}$ and $\hat \theta_{(w)}$ could be the ones computed on one sample, $\mathcal{J}$ or could be different ones; for example to borrow strength across samples we consider \[ \hat \beta_{(w)} = (\hat \beta_{(w),\mathcal{J}} +\hat \beta_{(1),\mathcal{J}^c} )/2, \qquad \hat \theta_{(w)} = (\hat \theta_{(w),\mathcal{J}} +\hat \theta_{(w),\mathcal{J}^c})/2 . \] Now, we define the variance estimate as
for $\hat \Omega$ and $\hat V_{(w)}$ defined in (ref) and (ref), respectively. We show that the construction above is appropriate for such circumstances and leads to asymptotically optimal confidence sets.
By well known results hahn1998role,newey1994asymptotic,robins1, $\psi$ in ((ref)) is the efficient influence function and $V_*$ is the semiparametric efficiency lower bound. Therefore, our estimator $\hat{\tau}$ in ((ref)) is a semiparametrically efficient estimator. As discussed before, the sparsity requirement in Theorem (ref) is considerably weaker that needed by existing estimators in the high-dimensional linear-logistic model (ref), including the methods discussed in athey2016approximate, belloni2014inference, farrell2015robust, ning2017high and tan2018model, in that we only need either the outcome model or the propensity model to be ultra sparse (but not both).
In this section we present numerical work where we contrast the behavior of the introduced method with the existing approaches. We consider the following design setting, $X_{i}\sim N(0,\Sigma)$ with $\Sigma_{i,j}=\rho^{|i-j|}$ and $\rho=0.6$. We set the sample size and the number of covariates to be $n=500$ and $p=600$, respectively. The following structure for the parameters is used. We consider the propensity model where $$\theta=a_{\theta}(1,0,1,0,1,0,...,0,1,0,0...,0)'$$ with $\|\theta\|_{0}=s_{\theta}$. We vary the value of $s_{\theta}$ and set $a_{\theta}$ such that $\sqrt{\theta'\Sigma_{X}\theta}=1$.
Similarly, for the outcome model we consider $$\beta_{(1)}=a_{\beta}(1,0,1,0,1,0,...,0,1,0,0...,0)'$$ and set $\beta_{(0)}=-\beta_{(1)}$. In other words, non-zero entries appear only on indices with odd numbers. For $a_\beta$, we consider two cases. In the first, we have homoskedastic Erros: The error term is generated from a centered $\chi^{2}(1)$ distribution (since it is light tailed and asymmetric). We set $\sqrt{\beta_{(1)}'\Sigma_{X}\beta_{(1)}}=\sqrt{2}$. We use $\sqrt{2}$ to get an $R^{2}$ of 50% (because the error has variance 2). The second has heteroskedastic errors: $\varepsilon_{i,(1)} $ is generated according to $$\p{4\times\mathbf{1}\{e(X_{i})\leq0.5\}+\mathbf{1}\{e(X_{i})>0.5\}}\xi_i$$ with $\xi_i$ being a centered $\chi^2(1)$ variable independent of $X_i$. Observe that $\varepsilon_{i,(0)}$ is still a centered $\chi^2(1)$ variable. We also consider other values of $a_\beta$ such that the $R^2$ in the homoskedastic case is 10%. We report the mean squared error (MSE) and coverage probability of 95% confidence interval (CP). We compare our methods with two popular alternatives:
The results are reported in Tables (ref) and (ref). Table (ref) indicated that in the baseline case with extremely sparse $\beta_{(w)}$ and $\theta$, AIPW, approximate residual balancing method and SDR perform very similarly. Note that this is as expected; all of the methods should be achieving the same asymptotic variance. However, when either the propensity score model or the conditional mean function or both are not extremely sparse, the SDR method delivers smaller MSE. This confirms our theoretical results, which state that our method is guaranteed to provide efficient estimation even if there is lack of extreme sparsity.
Perhaps a more direct way of formalizing this intuition is through analyzing the rate of the remainder similar to the discussion in newey2018cross; one way is to express the rate of the remainder for the asymptotic expansion in Theorem (ref) in terms of $\|\beta_{(w)}\|_0$ and $\|\theta\|_0$. One can use our technical arguments to show that, compared to AIPW, the remainder of the SDR estimator has the same order or smaller order of magnitude. In Table (ref), we also report the results by setting $a_\beta$ such that in the homoscedasticity case we have an R-squared of 10%. The pattern is quite similar; we observe comparable performance when we have extreme sparsity in both models and the proposed method has lower MSE in the absence of such sparsity in either model.