Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
38,255 characters · 10 sections · 41 citation commands
A Simple and Efficient Estimation of the Average Treatment Effect in the Presence of Unmeasured Confounders
\setcounter{section}{0}
\baselineskip0.6cm
{ Keywords: Average treatment effect; Unmeasured confounders; Semiparametric efficiency; Endogeneity.}
A common approach to account for individual heterogeneity in the treatment effect literature on observational data is to assume that there exist confounders, and conditional on these confounders, there is no systematic selection into the treatment (i.e., the so-called Unconfounded Treatment Assignment condition suggested in rosenbaum1983central, rosenbaum1984reducing). Under this assumption, several procedures for estimating the average treament effect (hereafter ATE) have been proposed, including the weighting procedure (rosenbaum1987model, hirano2003efficient, tan2010bounded, imai2014covariate, chan2016globally, yiu2018covariate); the matching procedure (rosenbaum2002observational, rosenbaum2002covariance, dehejia1999causal); and the regression procedure (heckman1997matching, heckman1998matching, imbens2006mean, chen2008semiparametric). For example survey, see imbens2009recent and imbens2015causal. A critical requirement in this literature is that all confounders are observed and available to researchers. In applications, however, it is often the case that some confounders are either not observed or not availale. In this case, the average treatment effect is only partially identified even with the aid of some insrumental variables (see Imbens1994Identification, Joshua1996Identification, Abadie2003Semiparametric,Abadie2002Instrumental,Tan2006Regression,Cheng2009Efficient,Ogburn2015Doubly for examples).\newline
Recently wang2016bounded suggested a noval identification condition of ATE when some confounders are not available. Under their condition, they showed that the semiparametric efficient influence function of ATE depends on five unknown functionals. They proposed to parameterize all five functionals, estimate those functionals with appropriate parametric approaches, plug the estimated functionals into the influence function, and then estimate the ATE from the estimated influence function. They established that their estimator is consistent if certain functionals are correctly parameterized and attains the semiparametric efficiency bound if all functionals are correctly specified. In applications, it is quite possible\ that some or all of the five functionals are misspecified and consequently their estimator could be inefficient or worse, inconsistent. This paper proposes an alternative, intuitive and easy to compute estimation that does not require parameterization of any of the five unknown functionals. We estbalish that under some sufficient conditions the proposed estimator is consistent, asymptotically normally distributed and attains the semiparametric efficiency bound. Moreover, the proposed procedure provides a natural and convenient estimate of the asymptotic variance. \newline
The paper is organized as follows. Section (ref) describes the basic framework. Section (ref) describes the proposed estimation and derives the large sample properties of the proposed estimator. Section (ref) presents a consistent variance estimator. Since the proposed procedure depends on smoothing parameters, Section (ref) presents a data driven method for selecting the smoothing paprameters. Section (ref) reports a small scale simulation study to evaluate the finite sample performance of the proposed estimator. Some concluding remarks are in Section (ref). All technical proofs are relegated to the Appendix and the supplementary material.
Let $D\in \{0,1\}$ denote the binary treatment indicator, and let $Y(1)$ and $Y(0)$ denote the potential outcomes when an individual is assigned to the treatment and control group respectively. The parameter of interest is the population average treatment effect $\tau = \mathbb{E}[Y(1)-Y(0)]$. Estimation of $\tau $ is complicated by the presence of confounders and the fact that $Y(1)$ and $Y(0)$ cannot be observed simultaneously. To distinguish observed confounders from unobserved confounders, we shall use $X$ to denote the observed confounders and use $U$ to denote the unmeasured confounders. It is well established in the literature that, when all confounders are observed, the following Unconfounded Treatment Assignment condition is sufficient to identify $\tau $ :
When $U$ is unmeasured, we have the classical omitted variable problem, causing the treament indicator $D$ to be endogenous. To tackle the endogeneity problem, instrumental variable is often the preferred choice. Let $Z\in \{0,1\}$ denote the variable satisfying the following classical instrumental variable conditions:
wang2016bounded showed that Asssumptions (ref)- (ref) alone do not identify $\tau $, but if in addition one of the following conditions holds:
then ATE is identified and can be expressed as
where
\newline Furthermore, wang2016bounded derived the efficient influence function for $\tau $: {
} where $f_{Z|X}(Z|X)$ is the conditional probability mass function of $Z$ given $X$. Clearly, the efficient influence function depends on five unknown functionals: $\delta (X)$, $\delta ^{D}(X)$, $f_{Z|X}$, $p_{0}^{Y}(X)= \mathbb{E}[Y|Z=0,X]$ and $p_{0}^{D}(X)=\mathbb{E}[D|Z=0,X]$. They proposed to parameterize all five functionals, estimate the functionals with appropriate parametric approaches, and plug the estimated functionals into the efficient influence function to estimate $\tau $. They established that their estimator of $\tau $ is consistent and asymptotically normally distributed if
and their estimator attains the semiparametric efficiency bound only when all five functionals are correctly specified. The main goal of this paper is to present an alternative, intuitive and easy approach to compute estimator that does not require parameterization of any of the functionals and is always consistent and asymptotically normal and attains the semiparametric efficiency bound.
To motivate our estimation procedure, we rewrite the treatment effect coefficient. Applying the tower law of conditional expectation, we obtain:
The above expression suggests a natural and intuitive plugin estimation, with $f_{Z|X}(Z|X)$ and $\delta ^{D}(X)$ replaced by some consistent estimates. There are many approaches to estimate these functionals including parametric and nonparametric approaches, but as noted by hirano2003efficient, not all estimates can lead to efficient estimation of $\tau $. In this paper, we present an intuitive and easy way to compute estimates of functionals that ensure efficiency of the plugin estimation of $ \tau $. To illustrate our procedure, we notice that the following conditions hold for any integrable functions $u_{1}(X)$ and $u_{2}(X)$:
and (ref) and (ref) uniquely determine $f_{Z|X}(Z|X)$ and $\delta ^{D}(X)$. These conditions impose restrictions on the unknown functionals and they must be taken into account when estimating those functionals. One difficulty with these conditions is that they must be imposed in an infinite dimmensional functional space. To overcome this difficulty, we propose to impose the conditions on a smaller sieve space. Specifically, let $u_{K}(X)=(u_{K,1}(X),\ldots ,u_{K,K}(X))^{\top }$ denote a known basis functions that can approximate any suitable function $u(X)$ arbitrarily well (see chen2007large or Appendix (ref) for further dicussion). Conditions (ref) and (ref) imply for any integers $K_{1}$ and $K_{2}$:
and
We shall construct estimates of the functionals by imposing the above conditions. To ensure consistency, we shall allow $K_{1}$ and $K_{2}$ to increase with sample size at appropriate rates.
Consider estimation of $f_{Z|X}(Z|X)^{-1}$. An obvious approach is to solve $ \left\{ w_{i},i=1,2,...,N\right\} $ from the sample analogue of (ref):
But there are many solutions and all solutions are consistent estimates of $ f_{Z|X}(Z|X)^{-1}$. The question is which solution is the best estimate of $ f_{Z|X}(Z|X)^{-1}$ in the sense of ensuring efficient estimation of $\tau $. Let $\rho (v)$ denote a strictly increasing and concave function and let $ \rho ^{\prime }(v)$ denote its first derivative. Denote
with $\hat{\lambda}_{K_{1}}\in \mathbb{R}^{K}$ maximizing the following objective function
It is easy to show that $N\hat{p}(X)$ satisfies (ref). Moreover, $N \hat{p}(X)$ can be interpreted as a generalized empirical likelihood estimator of $f_{Z|X}(1|X)^{-1}$ (see Appendix (ref)) and hence is the best estimate. The fact that $\hat{G}(\lambda )$ is globally concave implies that its maximand is easy to compute. \newline
Applying the same idea to (ref), we have
with $\hat{\beta}_{K_{1}}\in \mathbb{R}^{K_{1}}$ maximizing the following globally concave objective function
Again, $N\hat{q}(X)$ satisfies (ref) and can be interpreted as a generalized empirical likelihood estimatior of $f_{Z|X}(0|X)^{-1}$. \newline
The $\rho (v)$ function can be any increasing and strictly concave function. Some examples include $\rho (v)=-\exp (-v)$ for the exponential tilting kitamura1997information, imbens1998information, $\rho (v)=\log (1+v)$ for the empirical likelihood owen1988empirical, qin1994empirical, $ \rho (v)=-(1-v)^{2}/2$ for the continuous updating of the generalized method of moments hansen1982large, hansen1996finite and $\rho (v)=v-\exp (-v)$ for the inverse logistic.
Having estimated $f_{Z|X}(Z|X)^{-1}$, we now apply the same principle to estimate $\delta _{D}(X)$. But there is one difference. Here $\delta _{D}(X)\in \lbrack -1,1]$ and the $\rho (v)$ function is not suitable. We shall use the following strictly convex function
whose derivative is the tanh function $f^{\prime }(x)=\frac{e^{x}-e^{-x}}{ e^{x}+e^{-x}}$ with range $[-1,1]$. We estimate $\delta ^{D}(X)$ by
with $\hat{\gamma}_{K_{2}}\in \mathbb{R}^{K_{2}}$ maximizing the following globally concave function
Again, $\hat{\delta}^{D}(X)$ can be interpreted as a generalized empirial likelihood estimator and hence is the best estimate. \newline
Finally, the plugin estimator of $\tau $ is given by
To establish the large sample properties of $\widehat{\tau }$, we shall impose the following assumptions:
Assumption (ref) ensures the asymptotic variance to be bounded. Assumption (ref) restricts the covariates to be bounded. This condition, though restrictive, is commonly imposed in the nonparametric regression literature. Assumption (ref) requires the probability function to be bounded away from 0 and 1. Condition of this sort is familiar in the literature. Assumption (ref) is needed to control for the approximation bias, and they are commonly imposed in the nonparametric literature. Assumption (ref) imposes restrictions on the smoothing parameter so that the proposed estimator of ATE is root-N consistent. This condition, however, is practically unhelpful. We shall present a data driven approach to determine $ K_{1}$ and $K_{2}$. Assumption (ref) is a mild restriction on $\rho $ and is satisfied by all important special cases considered in the literature.\\
Under the above assumptions, the following theorem establishes the consistency, asymptotic normality and the semiparmetric efficiency of $\hat{ \tau}$.
Sketched proof can be found in Appendix (ref) and detailed proofs are provided in the supplementary material.
To conduct the statistical inference on $\tau $, we need a consistent estimator of the asymptotic variance of $\widehat{\tau }$. Note that the asymptotic varance of $\widehat{\tau }$, { \
} depends on five unknown functionals. Direct estimation of the variance requires replacing the five unknown functionals with consistent estimates. In this section, we present an alternative estimation that does not require estimation of those functionals.\newline
To illustrate the idea, we denote:{
} and
with $\theta \triangleq (\lambda ,\beta ,\gamma ,\tau )^{\top }$. Let $\hat{ \theta}\triangleq (\hat{\lambda}_{K_{1}},\hat{\beta}_{K_{1}},\hat{\gamma} _{K_{2}},\hat{\tau})^{\top }$ and ${\theta }^{\ast }\triangleq ({\lambda } _{K_{1}}^{\ast },{\beta }_{K_{1}}^{\ast },{\gamma }_{K_{2}}^{\ast },{\tau } )^{\top }$. Then $\hat{\theta}$ is the moment estimator solving the following moment condition:
Applying Mean Value Theorem, we obtain
where $\tilde{\theta}=(\tilde{\lambda}_{K_{1}},\tilde{\beta}_{K_{1}},\tilde{ \gamma}_{K_{2}},\tilde{\tau})^{\top}$ lies on the line joining $\hat{\theta}$ and $\theta ^{\ast }$. We show in the supplemental material that
Note that
where $\mathbf{e}_{2K_{1}+K_{2}+1}$ is a $(2K_{1}+K_{2}+1)$-dimensional column vector whose last element is $1$ and other components are all of $0$ 's. \newline
Combining (ref), (ref) and (ref), we obtain
which in turn implies
where
Therefore, we can define the sandwich estimator for the efficient variance $ V_{eff}$ by
where
The large sample properties of the proposed estimator permit a wide range of values of $K_{1}$ and $K_{2}$. This presents a dilemma for applied researchers who have only one finite sample and would like to have some guidance on the selection of smoothing parameters. In this section, we present a data-driven approach to select $K_{1}$ and $K_{2}$. Notice that $ f_{Z|X}(1|X)^{-1}$, $f_{Z|X}(0|X)^{-1}$ and $\delta ^{D}(X)$ satisfy the following regression equations:
Since $N\hat{p}(X)$, $N\hat{q}(X)$ and $\hat{\delta}^{D}(X)$ are consistent estimators of $f_{Z|X}(1|X)^{-1}$, $f_{Z|X}(0|X)^{-1}$ and $\delta ^{D}(X)$ respectively, the mean-squared-error (MSE) of the nuisance parameters $(\hat{ \lambda}_{K_{1}},\hat{\beta}_{K_{1}})$ and $\hat{\gamma}_{K_{2}}$ are defined by
The smoothing parameters $K_{1}$ and $K_{2}$ shall be chosen to minimize $ MSE_{1}$ and $MSE_{2}$. Specifically, denote the upper bounds of $K_{1}$ and $K_{2}$ by $\bar{K}_{1}$ and $\bar{K}_{2}$ (e.g. $\bar{K}_{1}=\bar{K}_{2}=5$ in our simulation studies). The data-driven $K_{1}$ and $K_{2}$ are given by
In this section, we conduct a small scale simulation study to evaluate the finite sample performance of the proposed estimator. To evaluate the performance of our estimator against the existing alternatives, particularly the estimators proposed by wang2016bounded , we adopt the exact same design (i.e., the same data generating processes (DGP)). In each Monte Carlo run, we generate sample of data from DGP for two sizes: $N=500$ and $N=1000$ respectively, and from each sample we compute our estimator and other existing estimators. We then repeat the Monte Carlo runs for $500$ times. \newline
The observed baseline covariates are $X=(1,X_{2})$, where $X$ include an intercept term and a continuous random variable $X_{2}$ uniformly distributed on the interval $(-1,-0.5)\cup (0.5,1)$. The unmeasured confounder $U$ is a Bernoulli random variable with mean 0.5. The instrumental variable $Z$, treatment variable $D$ and outcomes variable $ Y\in \{0,1\}$ are generated according to the simulation design of wang2016bounded. The true value of the average treatment effect is $\tau =0.087$.\newline
We compute the proposed estimator (cbe), the naive estimator, the multiply robust estimator (mr) and the bounded multiply robust estimator (b-mr) proposed by wang2016bounded. Details of calculations are given below.
The multiply robust estimator (mr) and the bounded multiply robust estimator (b-mr) proposed by wang2016bounded depend on parameterization of five unknown functionals. In their paper they considered several models, denoted by $\mathcal{M}_{1}$, $\mathcal{M}_{2}$ and $\mathcal{M}_{3}$ (see wang2016bounded for a detailed discussion of the model specification). Following wang2016bounded, we consider scenarios where some or all functionals are misspecified. \newline \baselineskip0.6cm
Table 1 reports the bias, standard deviation (Stdev), and the root mean square error (RMSE) of $\widehat{\tau }$ from the 500 Monte Carlo runs. In each Monte Carlo run, we use the data driven approach to select $K_{1}$ and $ K_{2},$ and their histograms are depicted in Figure 1. The estimated asymptotic variances are reported in Table 2. \newline
Glancing at these tables, we have the following observations:
Overall, the simulation results show that the proposed estimator out-performs the existing estimators.
Most of the existing treatment effect literature on observational data assume that all confounders are observed and available to researchers. In applications, it is often the case that some confounders are not observed or not available. wang2016bounded studied identification and estimation of the average treament effect when some confounders are not observed. They propose to parameterize five unknown functionals and show that their estimation is consistent when certain functionals are correctly specified and is efficient when all functionals are correctly specified. This paper proposes an alternative estimation. Unlike wang2016bounded , the proposed estimation does not parameterize any of the functionals and is always consistent. Moreover, the proposed estimator attains the semiparametric efficiency bound. A simple asymptotic variance estimator is presented, and a small scale simulation study suggests the practicality of the proposed procedure.\newline
Our procedure only applies to the binary treatment with unmeasured confounders. However, other forms of treatment, such as multiple valued or continuous treatment, may arise in applications. Extension of the proposed methodology to those forms of treatment with unmeasured confounders is certainly of great interest. This extension shall be pursued in a future project.
\singlespacing