Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
105,564 characters · 21 sections · 89 citation commands
Improving Estimation Efficiency via Regression-Adjustment in Covariate-Adaptive Randomizations with Imperfect Compliance
Randomized experiments have become increasingly popular in economic research. One commonly used randomization method employed by economists to ensure balance between treatment and control is covariate-adaptive randomization (CAR) (B09, B09), in which subjects are randomly assigned to treatment and control within strata formed by a few key pretreatment variables. However, subject compliance with the random assignment is usually imperfect. We survey all publications using randomized experiments in eight leading economics journals from 2015 to 2022 and identify eleven papers that used CARs with imperfect compliance.\footnote{See Section (ref) for more details.}
When subjects do not comply with the assignment in CARs, researchers usually estimate the local average treatment effects (LATEs) for the compliers using the two-stage least squares (TSLS) method with treatment assignment as an instrumental variable and covariates and strata fixed effects as exogenous controls. Actually, all eleven papers mentioned above estimate the LATE in this way. We simply denote this estimator as TSLS. Recently, anseletal2018 proposed an S estimator (denoted as S) which aggregates IV estimators for each stratum. BG21 proposed a fully saturated estimator with strata dummies, which we call the unadjusted estimator (NA) as it does not use covariates. The standard theory for the consistency of TSLS requires both correct specification of the conditional mean model and homogeneous treatment effect. In contrast, both S and NA estimators are consistent under CARs without requiring correct specifications, homogeneous treatment effect, or identical treatment assignment probability across strata. anseletal2018 further show the S estimator is the most efficient among all the estimators discussed in their paper (Proposition 7).
The existing literature lacks a systematic study and comparison of various LATE estimators under CARs. TSLS and S estimators impose different linear conditional mean models, which can be viewed as different types of linear regression adjustments. Then, under what conditions the TSLS estimator, like the S estimator, is consistent even when the regression adjustments are misspecified? How is the efficiency comparison among TSLS, S, and NA estimators when all of them are consistent? Is the S estimator the most efficient among all linearly adjusted LATE estimators? Can other potentially misspecified nonlinear regression adjustments lead to more efficient LATE estimators? Last, what is the semiparametric efficiency bound (SEB) for LATE estimation under (CARs) and how can we achieve it?
In this paper, we provide answers to all these questions. Specifically, we follow the framework that was recently established by BCS17 to study causal inference under CARs, which allows for heterogeneous assignment probabilities and treatment effects. We first show that (1) TSLS with both strata dummies and covariates as exogenous controls is inconsistent if both the assignment probabilities and treatment effects are heterogeneous across strata; (2) even when TSLS is consistent (especially when the treatment assignment probabilities are homogeneous), its usual heteroskedasticity robust standard error is conservative due to the cross-sectional dependence introduced by CARs;\footnote{This point is consistent with the result in anseletal2018 for their estimator $\hat \beta_2$. However, $\hat \beta_2$ is computed by TSLS with only strata dummies under the assumption of homogeneous assignment probabilities, but no covariates as exogenous control variables.} (3) the correct asymptotic variance of the TSLS estimator may be greater than that of the NA estimator, which defeats the purpose of using covariates in the regression.
We then propose a general adjusted estimator using the doubly robust moment for LATE with a consistent estimator of the assignment probability and potentially misspecified regression adjustments based on covariates. The doubly robust moment for LATE has been derived by F07late and used for estimating LATE by SUW22 and H22. But we are the first to apply it under CARs and investigate the potential efficiency improvements when the regression adjustments are misspecified. We show that our inference method (1) achieves the exact asymptotic size under the null despite the cross-sectional dependence introduced by CARs, (2) is robust to adjustment misspecification, and (3) achieves the SEB when the adjustments are correctly specified. The SEB for LATE under CARs is also new to the literature and complements those bounds derived by F07late and A22.\footnote{F07late derived the SEB for LATE assuming i.i.d. data. However, CARs can introduce cross-sectional dependence, and thus, violate the independence assumption. A22 derived the SEB for average treatment effect under CARs but without covariates. The SEB for LATE under CARs but without covariates is a byproduct of our result by letting our covariates be an empty set.}
Finally, we compare the efficiency of our LATE estimators with three specific forms of regression adjustments: (1) the optimal linear adjustment (denoted as L), which yields the most efficient estimator among all linearly adjusted estimators, (2) the nonlinear logistic adjustment (denoted as NL), and (3) a combination of linear and nonlinear adjustments (denoted as F) which is more efficient than both linear and nonlinear adjustments and new to the literature. We also extend anseletal2018 by showing that their S estimator is asymptotically equivalent to our estimator L, thus is optimal among the linearly adjusted estimators but less efficient than estimator F. We further give conditions under which estimators with nonparametric (denoted as NP) and regularized (denoted as R) regression adjustments achieve the SEB. Figure (ref) visualizes the partial order of efficiency of these estimators.
Our paper is related to several lines of research. HH12,MHZ15,MQLH18,O21,SY13,ZZ20,Y18,YS20 studied inference of either the average treatment effect (ATE) or quantile treatment effect (QTE) under CARs without considering covariates. BCS17,BCS18,BL16,F18,L13,L16,LD20,LiD20,LTM20,LY20,NW20,SYZ10,YYS20,ZD20 studied the estimation and inference of ATEs using a variety of regression methods under various randomization schemes. jiang2021b examine regression-adjusted estimation and inference of QTEs under CARs. Based on pilot experiments, T18 and B19 devise optimal randomization designs that may produce an ATE estimator with the lowest variance. BG21 further examine the optimal design with imperfect compliance. All the above works, except BG21, assume perfect compliance, while we contribute to the literature by studying the LATE estimators in the context of CARs and regression adjustment, which allows imperfect compliance. renliu2021 study the regression-adjusted LATE estimator in completely randomized experiments for a binary outcome using finite population asymptotics. We differ from their work by considering the regression-adjusted estimator in covariate-adaptive randomizations for a general outcome using the superpopulation asymptotics. Finally, our paper also connects to a vast literature on estimation and inference in randomized experiments, including HHK11, athey2017, abadie2018, T18, BRS19, B19, JL20, among many others.
Acronyms. In this paper, we refer to the optimally linearly adjusted, nonlinearly (logistic) adjusted, and nonparametrically adjusted estimators as L, NL, and NP, respectively. We also use NA and S to denote the fully saturated and S estimators proposed by BG21 and anseletal2018, respectively. F denotes the estimator with adjustments that improve upon both optimal linear and nonlinear adjustments, while R denotes the estimator with regularized adjustments. We will provide more details about these estimators below.
Let $Y_{i}$ denote the observed outcome of interest for individual $i$; write $Y_{i} = Y_{i}(1)D_{i} + Y_{i}(0)(1-D_{i})$, where $Y_{i}(1)$ and $Y_{i}(0)$ are the potential treated and untreated outcomes for the individual $i$, respectively, and $D_{i}$ is a binary random variable indicating whether the individual $i$ received treatment ($D_{i}=1$) or not ($D_{i}=0$) in the actual study. One could link $D_{i}$ to the treatment assignment $A_{i}$ in the following way: $D_{i} = D_{i}(1)A_{i} + D_{i}(0)(1-A_{i})$, where $D_{i}(a)$ is the individual $i$'s treatment outcome upon receiving treatment status $A_{i}=a$ for $a=0,1$; $D_{i}(a)$ is a binary random variable. Define $Y_{i}(D_{i}(a)): = Y_{i}(1)D_{i}(a) + Y_{i}(0)(1-D_{i}(a))$, so we can write $Y_{i}=Y_{i}(D_{i}(1))A_{i}+Y_{i}(D_{i}(0))(1-A_{i})$. Individual $i$ belongs to stratum $S_{i}$ and possesses covariate vector $X_{i}$, where $X_{i}$ does not include the constant term. The support of the vectors $\{X_{i}\}_{i=1}^{n}$ is denoted $\text{Supp}(X),$ while the support of $\{S_{i}\}_{i=1}^{n}$ is $\mathcal{S}$, which is a finite set.
A researcher can observe the data $\{Y_{i},D_{i},A_{i},S_{i},X_{i}\}_{i=1}^{n}$. Define $[n]:=\{1,2,...n\}$, $p(s):=\mathbb{P}(S_{i}=s)$, $n(s):=\sum _{i\in\lbrack n]}1\{S_{i}=s\}$, $n_{1}(s):=\sum_{i\in\lbrack n]}A_{i} 1\{S_{i}=s\}$, $n_{0}(s):=n(s)-n_{1}(s)$, $S^{(n)}:=(S_{1},\ldots,S_{n})$, $X^{(n)}:=(X_{1},\ldots,X_{n})$, and $A^{(n)}:=(A_{1},\ldots,A_{n})$. We make the following assumptions on the data generating process (DGP) and the treatment assignment rule.
Several remarks are in order. First, Assumption (ref)(i) allows for the treatment assignment $A^{(n)}$, and thus, the observed outcome $\{Y_i\}_{i \in [n]}$ to be cross-sectionally dependent, which is usually the case for CARs. Second, Assumption (ref)(ii) implies that the treatment assignment $A^{(n)}$ are generated only based on strata indicators. Third, Assumption (ref)(iii) imposes that the strata sizes are roughly balanced. Fourth, BCS17 show that Assumption (ref)(iv) holds under several covariate-adaptive treatment assignment rules such as simple random sampling (SRS), biased-coin design (BCD), adaptive biased-coin design (WEI) and stratified block randomization (SBR).\footnote{For completeness, we briefly repeat their descriptions in Appendix (ref).} Note that we only require $B_{n}(s)/n(s)=o_{p}(1)$, which is weaker than the assumption imposed by BCS17 but the same as that imposed by BCS18 and ZZ20. Fifth, Assumption (ref)(v) implies there are no defiers. Last, Assumption (ref)(vi) is a standard moment condition.
Throughout the paper, we are interested in estimating the local average treatment effect (LATE), which is denoted by $\tau$ and defined as \[ \tau:=\mathbb{E}\sbr[1]{Y(1) - Y(0)|D(1)>D(0)}; \] that is, we are interested in the ATE for the compliers (angristimbens1994, angristimbens1994).
To motivate our work, we give three examples of prominent economic datasets that use CARs and have imperfect compliance.
\defcitealias{royer2015}{Royer et al. (2015)} \defcitealias{himmler2019}{Himmler et al. (2019)} \defcitealias{bolhaar2019}{Bolhaar et al. (2019)} \defcitealias{angrist2021}{Angrist et al. (2021)}
We survey the common practice for analyzing experiments in the empirical economics literature. Our survey is limited to articles that contain the term “experiment" in their title or abstract and are published between January 2015 and December 2022 in eight journals: the American Economic Journal: Applied Economics (AEJ: Applied), American Economic Journal: Economic Policy (AEJ: Policy), American Economic Review, Econometrica, Journal of Political Economy, Quarterly Journal of Economics (QJE), \textit{Review of Economics and Statistics} (ReStat), and \textit{Review of Economic Studies}. We then manually select the articles that use CARs and report imperfect compliance. Table (ref) tabulates the articles found in our survey. It shows that all the papers in our sample use TSLS with covariates and strata fixed effects to estimate the LATE. This finding motivates us to study the statistical properties of this commonly used TSLS estimator in Section (ref) before proposing our new estimator.
Our survey shows that empirical researchers using CARs usually estimate LATE via TSLS regressions with strata dummies and covariates. The first and second stages of the TSLS regression can be formed as
where $\{a_s\}_{s \in \mathcal S}$ and $\{\alpha_{s}\}_{s\in\mathcal{S}}$ are the strata fixed effects.
Denote the TSLS estimator of $\tau$ by $\hat{\tau}_{TSLS}$. To study the asymptotic properties of $\hat{\tau}_{TSLS}$, we follow BCS17 and anseletal2018 and make the following additional assumption on the treatment assignment mechanism.
Three remarks are in order. First, Assumption (ref) is used to analyze the TSLS estimator only and is not needed for all the analyses in later sections in the paper. Second, it implies Assumption (ref)(iv). Third, we have $\gamma(s) = \pi(s)(1-\pi(s))$ for SRS and $\gamma(s) < \pi(s)(1-\pi(s))$ for the other three randomization designs mentioned after Assumption (ref). Specifically, for BCD and SBR, we have $\gamma(s) = 0$, which means the assignment rules achieve the strong balance.
Following empirical researchers, we also consider the usual IV heteroskedasticity-robust standard error estimator for TSLS estimator $\hat{\tau}_{TSLS}$, which is denoted as $\hat{\sigma}_{TSLS,naive}/\sqrt{n}$.\footnote{The detailed definition of $\hat{\sigma}_{TSLS,naive}$ can be found in the proof of Theorem (ref).} We compare $\hat \tau_{TSLS}$ with BG21's (BG21) fully saturated estimator (denoted as $\hat \tau_{NA}$) for $\tau$ under CAR, which does not use any covariates $X_i$. The asymptotic variance of $\hat \tau_{NA}$ is then denoted as $\sigma_{NA}^2$, which is given in BG21. In Section (ref), we further show that $\hat \tau_{NA}$ is a special case of our general estimator whose asymptotic variance is derived in the proof of Theorem (ref).
Theorem (ref) highlights one advantage and three limitations of the commonly used TSLS estimator under CARs. The advantage is that the TSLS estimator can consistently estimate the LATE under certain conditions without assuming the linear regression in (ref) being correctly specified. So the reason for incorporating covariates in the regression is to improve estimation efficiency. The first limitation is that the TSLS estimator is inconsistent when both the treatment effect and the probabilities of treatment assignment vary across strata. To ensure its consistency, economists should thus keep the target assignment probability ($\pi(s)$) equal across all strata in the experimental design stage, which may not be satisfied in reality (see, for example, the first dataset in Section (ref)). The second limitation is that the heteroskedasticity-robust standard error reported by standard software such as STATA is conservative and inconsistent unless $\gamma(s) = \pi(1-\pi)$. However, this condition is violated when treatment is not assigned independently, such as BCD and SBR, which are widely used in RCTs. With the cross-sectional dependence among treatment assignments, it is expected that the usual heteroskedasticity-robust standard error is inconsistent. The third limitation is that the asymptotic variance $\sigma^2_{TSLS}$ may not be smaller than that of the unadjusted estimator, which goes against the purpose of using covariates in the regression. In this paper, we develop estimators that have the same advantage but avoid all these limitations. Specifically, our proposed LATE estimators are (1) consistent even under misspecification of regression models, (2) consistent even when the probabilities of treatment assignment are heterogeneous across strata, and (3) guaranteed to be weakly more efficient than the unadjusted estimator. We also provide consistent estimators of the asymptotic variances for our LATE estimators.
In this section, we propose a general regression-adjusted LATE estimator for $\tau$. Define $\mu^{D}(a,s,x) := \mathbb{E}\sbr[1]{D(a)|S = s, X=x}$ and $\mu ^{Y}(a,s,x) := \mathbb{E}\sbr[1]{Y(D(a))|S=s, X=x}$ for $a=0,1$ as the true specifications. In practice, these are unknown and empirical researchers employ working models $\overline{\mu}^{D}(a,s,x)$ and $\overline{\mu}^{Y}(a,s,x)$, which may differ from the true specifications. We then proceed to estimate the working models with estimators $\hat{\mu}^{D}(a,s,x)$ and $\hat{\mu}^{Y}(a,s,x)$. As the working models are potentially misspecified, their estimators are potentially inconsistent for the true specifications.
To further differentiate $\mu^b(\cdot)$, $\overline{\mu}^b(\cdot)$, and $\hat \mu^b(\cdot)$ for $b \in \{D,Y\}$, we consider an example that $\mu^D(a,s,x)$ follows a probit model, i.e., $\mu^D(a,s,x) = F_N(\tilde \alpha_{a,s}+x^\top\tilde \beta_{a,s})$, where $F_N(\cdot)$ is the standard normal CDF, and $\tilde \alpha_{a,s}$ and $\tilde \beta_{a,s}$ are the regression coefficients which are allowed to depend on assignment $a$ and stratum $s$. However, the researcher does not know the correct specification and instead uses a logit model $\overline{\mu}^D(a,s,x) = \lambda(\alpha_{a,s}+x^\top \beta_{a,s})$ as the working model, where $\lambda(\cdot)$ is the logistic CDF. Then $(\alpha_{a,s},\beta_{a,s})$ are the pseudo true values that depend on how they are estimated and can be defined as the probability limits of the chosen estimator $(\hat \alpha_{a,s},\hat \beta_{a,s})$. For instance, we can estimate the regression coefficients in the logistic model via logistic quasi MLE or nonlinear least squares. As the logistic model is misspecified, the two estimation methods lead to two different pseudo true values. Suppose we estimate $(\alpha_{a,s},\beta_{a,s})$ by quasi MLE and denote their estimators as $(\hat \alpha_{a,s},\hat \beta_{a,s})$. The estimator of the working model is then $\hat \mu^D(a,s,x) = \lambda(\hat \alpha_{a,s}+x^\top\hat \beta_{a,s})$.
In CAR, the targeted assignment probability for stratum $s$, $\pi(s)$, is usually known or can be consistently estimated by $\hat{\pi}(s) := \frac{n_{1}(s)}{n(s)}$. Then our proposed estimator of LATE based on the doubly robust moments\footnote{For reference of doubly robust moments, see RobinsRotnitzkyZhao1994, RR95, ScharfsteinRotnitzkyRobins1999, RobinsRotnitzkyvanderLaan2000, hiranoimbens2001, F07late, wooldridge2007, rothefirpo2019 etc; see sloczynskiwooldridge2018 and seamanvansteelandt2018 for recent reviews.} is
Given the double robustness and the consistency of $\hat \pi(s)$, our estimator $\hat \tau$ is consistent even when the working models $(\hat \mu^D(\cdot),\hat \mu^Y(\cdot))$ are misspecified. Our analysis also takes into account the cross-sectional dependence of the treatment statuses caused by the randomization and is therefore different from the double robustness literature that mostly focuses on the observational data with independent treatment statuses. Furthermore, our general adjusted estimator is numerically invariant to the stratum-specific location shift because
Therefore, using adjustments $\hat \mu^b(a,S_i,X_i)$ and $\hat \mu^b(a,S_i,X_i) - \mathbb{E}(\mu^b(a,S_i,X_i)|S_i)$ for $b \in \{D,Y\}$ are numerically equivalent.
Assumption (ref) requires $\hat \mu^b(\cdot)$ to be a consistent estimator of $\overline{\mu}^b(\cdot)$ for $b=D,Y$. For instance, we can consider a linear working model $\overline{\mu}^{Y}(a,s,X_{i})=X_{i}^{\top}\beta_{a,s}$, where the pseudo true value $\beta_{a,s}$ is defined as the probability limit of the OLS estimator $\hat \beta_{a,s}$ from regressing $Y_i$ on $X_i$ using observations in $I_a(s)$. Then, the estimator $\hat{\mu}^{Y}(a,s,X_{i})$ can be written as $X_{i}^{\top}\hat{\beta}_{a,s},$ and Assumption (ref)(i) reduces to
which holds automatically because by definition, $\hat{\beta}_{a,s}\overset{p}{\longrightarrow}\beta _{a,s}$, and we will assume $\mathbb EX_i^2 <\infty$. This example shows that we do not need to assume the working model $\overline{\mu}^{Y}(a,s,X_{i})=X_{i}^{\top}\beta_{a,s}$ is correctly specified. A similar remark applies to Assumption (ref)(ii) and nonlinear working models such as the logistic regression mentioned earlier. We verify Assumption (ref) for general parametric adjustments in Section (ref) below.
To state our first main result below, we need to introduce extra notation. Let $\mathcal{D}_{i}: = \{Y_{i}(1), Y_{i}(0), D_{i}(1), D_{i}(0), X_{i}\}$, $W_{i} := Y_{i}(D_{i}(1)),$ $Z_{i} :=Y_{i}(D_{i}(0))$, $\tilde{W}_{i} :=W_{i}-\mathbb{E}[W_{i}|S_{i}],$ $\tilde{Z}_{i}:=Z_{i}-\mathbb{E}[Z_{i}|S_{i}],$ $\tilde{X}_{i}:=X_{i}-\mathbb{E}[X_{i}|S_{i}]$, $\tilde{D}_{i}(a) :=D_{i}(a)-\mathbb{E}[D_{i}(a)|S_{i}]$ for $a=0,1,$ and
Several remarks are in order. First, Theorem (ref)(i) establishes the limiting distribution of our adjusted LATE estimator, which also implies its consistency. Our estimator inherits the advantage of the TSLS estimator because it remains consistent even when the adjustment $\overline{\mu}^b(\cdot)$ is misspecified, but avoids its limitation because our estimator remains consistent when $\pi(s)$ varies across strata. Additionally, the terms $\sigma_0^2$, $\sigma_1^2$, and $\sigma_2^2$ in the asymptotic variance of our regression-adjusted LATE estimator represent the sampling variations from the control units within each stratum, the treatment units within each stratum, and the strata itself, respectively.
Second, Theorem (ref)(ii) gives a consistent estimator of this asymptotic variance, which depends on the working model $\overline{\mu}^{b}(a,s,x)$ for $(a,b)\in\{0,1\}\times\{D,Y\}$. Different working models lead to different estimation efficiencies.
Third, Theorem (ref)(iii) further shows that our general regression-adjusted estimator achieves the semiparametric efficiency bound $\underline{\sigma}^{2}$ derived in Theorem (ref) below when the working models are correctly specified.
Fourth, when there are no adjustments so that $\overline{\mu}^{Y}(\cdot)$ and $\overline{\mu}^{D}(\cdot)$ are zero, we obtain
In this case, our estimator coincides numerically with BG21's (BG21) fully saturated estimator (i.e., NA). Indeed, we can verify that $\sigma^{2}$ defined above is the same as the asymptotic variance of the fully saturated estimator derived by BG21.
Several remarks are in order. First, Theorem (ref) suggests that the asymptotic variance of any regular root-$n$ consistent and asymptotically normal semiparametric estimator of LATE is bounded from below by $\underline{\sigma}^{2}$. Second, the proof of Theorem (ref) follows the arguments of A22, who accounted for the cross-sectional dependence of $\{A_{i}\}_{i\in\lbrack n]}$. Third, the efficiency bound here differs slightly from the one derived by F07late under unconfoundedness for observational data because here the covariates $X_{i}$ only affect the conditional mean models (i.e., $\mu^{b}(a,s,x)$ for $a=0,1$, $b=\{D,Y\}$) but not the \textquotedblleft propensity score" $\pi(\cdot)$. Fourth, Theorem (ref) implies that various CARs (with or without achieving strong balance) lead to the same SEB for LATE estimation. Such a result is consistent with what A22 found for ATE under general randomization schemes.
In this section, we consider estimating $\overline{\mu} ^{b}(a,s,x)$ for $a=0,1$, $s \in\mathcal{S}$, and $b = D,Y$ via parametric regressions. Note that we do not require $\overline{\mu}^{b}(a,s,x)$ to be correctly specified. Suppose that
where $\Lambda_{a,s}^{b}(\cdot)$ for $(a,b,s) \in\{0,1\} \times\{D,Y\}\times \mathcal{S}$ is a known function of $X_{i}$ up to some finite-dimensional parameter (i.e., $\theta_{a,s}$ and $\beta_{a,s}$). The researchers have the freedom to choose the functional forms of $\Lambda_{a,s}^{b}(\cdot)$, the parameter values of $(\theta_{a,s},\beta_{a,s})$, and the methods of estimation. As mentioned above, because the parametric models are potentially misspecified, different estimation methods of the same model can lead to distinctive pseudo true values. We will discuss several detailed examples in Sections (ref), (ref), and (ref) below. Here, we first focus on the general setup.
Define the estimators of $(\theta_{a,s},\beta_{a,s})$ as $(\hat{\theta} _{a,s},\hat{\beta}_{a,s})$, and hence the corresponding feasible parametric regression adjustments as
Assumption (ref)(i) means that $(\hat{\theta} _{a,s},\hat{\beta}_{a,s})$ are consistent estimators for $(\theta_{a,s},\beta_{a,s})$. Assumption (ref)(ii) means that the parametric models are smooth in their parameters, which is true for many widely used regression models such as linear, logit, and probit regressions. This restriction can be further relaxed to allow for non-smoothness under less intuitive entropy conditions.
Theorem (ref) generalizes the intuition in (ref) and shows that Assumption (ref) holds for general parametric models as long as the parameters are consistently estimated.
In this section, we consider working models that are linear in $\Psi_{i,s}$ where $\Psi_{i,s}=\Psi_{s}(X_{i})$ is a function of $X_{i}$ and its functional form can vary across $s\in\mathcal{S}$. Specifically, suppose, for $a=0,1$ and $s\in\mathcal{S}$, that $\overline{\mu}^{Y}(a,s,X)=\Psi_{i,s}^{\top}t_{a,s}$ and $\overline{\mu} ^{D}(a,s,X)=\Psi_{i,s}^{\top}b_{a,s}$, where $t_{a,s}$ and $b_{a,s}$ are the regression coefficients whose values are freely chosen by the researchers. The restriction that the function $\Psi_{s}(\cdot)$ does not depend on $a=0,1$ is innocuous as, if it does, we can stack them up and denote $\Psi_{i,s} =(\Psi_{1,s}^{\top}(X_{i}),\Psi_{0,s}^{\top}(X_{i}))^{\top}$. Similarly, it is also innocuous to impose that the function $\Psi_{s}(\cdot)$ is the same for modeling $\overline{\mu}^{Y}(a,s,X)$ and $\overline{\mu}^{D}(a,s,X)$.
Given that all values of $t_{a,s}$ and $b_{a,s}$ lead to consistent estimators of LATE, a natural question to ask is what values give the most precise estimator. Let the asymptotic variance of the adjusted LATE estimator $\hat{\tau}$ be as $\sigma^{2}$, which depends on $(\overline{\mu}^{Y}(a,s,X) ,\overline{\mu }^{D}(a,s,X) )$, and thus, $(t_{a,s},b_{a,s})$. Let $\Theta^*$ be the collection of optimal linear coefficients that minimize the asymptotic variance of $\hat{\tau}$ over all possible $(t_{a,s},b_{a,s})$, i.e.,
Assumption (ref) requires that the regressor $\Psi_{i,s}$ does not contain a constant term. In fact, (ref) and (ref) imply that our estimator is numerically invariant to a stratum-specific location shift. The following theorem characterizes the set of optimal linear coefficients.
The optimality result in Theorem (ref) relies on two key restrictions: (1) the regressor $\Psi_{i,s}$ is the same for treated and control units and (2) both the adjustments $\overline{\mu}^{Y}(a,s,X)$ and $\overline{\mu}^{D}(a,s,X)$ are linear. It is possible to have nonlinear adjustments that are more efficient. We will come back to this point in Sections (ref), (ref), and (ref).
In view of Theorem (ref), the optimal linear coefficients are not unique. In order to achieve the optimality, we only need to consistently estimate one point in $\Theta^{*}$. For the rest of the section, we choose $(\theta_{a,s}^{L},\beta_{a,s}^{L})$ with the corresponding optimal linear adjustments
We estimate $(\theta_{a,s}^{L},\beta_{a,s}^{L})$ by $(\hat{\theta} _{a,s}^{L},\hat{\beta}_{a,s}^{L})$, where
Then, the feasible linear adjustments can be defined as
Suppose that $\mathcal{S}=\{1,\ldots,S\}$ for some integer $S>0$. It is clear that $\hat{\theta}_{a,s}^{L}$ and $\hat{\beta}_{a,s}^{L}$ are the OLS-estimated slopes of the following two linear regressions using observations in $I_{a}(s)$:
The asymptotic variance of the LATE estimator with the optimal linear adjustments ($\hat{\tau}_{L}$) takes the form of (ref) with $\{\overline{\mu}^{b}(a,s,X_{i})\}_{b = D,Y, a=0,1, s \in\mathcal{S}}$ in (ref)--(ref) defined in (ref). It is also guaranteed to be weakly smaller than that of both $\hat \tau_{NA}$ and $\hat \tau_{TSLS}$, which addresses the Freedman's critique (F08b, F08b, F081). When $\Psi_{i,s} = X_{i}$, this asymptotic variance is the same as that of anseletal2018's S estimator, derived in Section (ref) of the Online Supplement. This implies the S estimator is the most efficient LATE estimator adjusted by linear functions of $X_{i}$, and thus, more efficient than $\hat{\tau}_{TSLS}$ and $\hat \tau_{NA}$.
It is also common to consider a linear model for $\overline{\mu} ^{Y}(a,s,X_{i})$ and a logistic model for $\overline{\mu}^{D}(a,s,X_{i})$, i.e., \[ \overline{\mu}^{Y}(a,s,X_{i})=\mathring{\Psi}_{i,s}^{\top}t_{a,s} \quad\text{and}\quad\overline{\mu}^{D}(a,s,X_{i})=\lambda(\mathring{\Psi }_{i,s}^{\top}b_{a,s}), \] where $\mathring{\Psi}_{i,s}=(1,\Psi_{i,s}^{\top})^{\top}$, $\Psi_{i,s}=\Psi_{s}(X_{i})$ and $\lambda(u)=\exp(u)/(1+\exp(u))$ is the logistic CDF. As the model for $\overline{\mu}^{D}(a,s,X_{i})$ is non-linear, the optimality result established in the previous section does not apply. We can consider fitting the linear and logistic models by OLS and (quasi) MLE, respectively, and call this method the nonlinear (logistic) adjustment. Specifically, define
where
It is clear that $\hat{\theta}_{a,s}^{OLS}$ and $\hat{\beta}_{a,s}^{MLE}$ are the OLS and ML estimates of the following two stratum-specific (logistic) regressions using observations in $I_{a}(s)$:
In the logistic regression, we do allow the regressor $\mathring{\Psi}_{i,s}$ to contain the constant term. Suppose $\hat{\theta}_{a,s}^{OLS}=(\hat{h} _{a,s}^{OLS},\hat{\underline{\theta}}_{a,s}^{OLS,\top})^{\top}$, where $\hat{h}_{a,s}^{OLS}$ is the intercept. Then, because our adjusted LATE estimator is invariant to the stratum-specific location shift of the adjustment term, using $\hat{\mu}^{Y}(a,s,X_{i})=\mathring{\Psi}_{i,s}^{\top}\hat{\theta}_{a,s} ^{OLS}=\hat{h}_{a,s}^{OLS}+\Psi_{i,s}^{\top}\hat{\underline{\theta}} _{a,s}^{OLS}$ and $\hat{\mu}^{Y}(a,s,X_{i})=\Psi_{i,s}^{\top}\hat {\underline{\theta}}_{a,s}^{OLS}$ produce the exact same LATE estimator. In addition, we have $\hat {\underline{\theta}}_{a,s}^{OLS}=\hat{\theta}_{a,s}^{L}$ by construction. This means $\hat{\mu}^{Y}(a,s,X_{i})$ used here is the same as that for the optimal linear adjustment. In contrast, because the logistic regression is nonlinear, the non-intercept part of $\hat{\beta}_{a,s}^{MLE}$ does not equal $\hat{\beta}_{a,s}^{L}$. The limits of $\hat{\theta}_{a,s}^{OLS}$ and $\hat{\beta}_{a,s}^{MLE}$ are defined as
which imply that the working models are
Several remarks are in order. First, the nonlinear (logistic) adjustment is not optimal in the sense that it does not necessarily minimize the asymptotic variance of the corresponding LATE estimator over the class of linear/logistic adjustments. Second, the nonlinear (logistic) adjustment is not necessarily less efficient than the optimal linear adjustment studied in Section (ref) as $\mu^{D}(a,s,X_{i})$ could be nonlinear. In fact, as Theorem (ref) shows, if the adjustments are correctly specified, then $\hat{\tau}_{NL}$ can achieve the semiparametric efficiency bound. Compared with the linear probability model considered in Section (ref), the logistic model is expected to be less misspecified, especially when the regressor $\Psi_{i,s}$ contains nonlinear transformations of $X_{i}$ such as interactions and quadratic terms. Third, we will further justify the intuition above in Section (ref), in which we let $\Psi_{i,s}$ be the sieve basis functions with an increasing dimension and show that the nonlinear (logistic) method can consistently estimate the correct specification under some regularity conditions. Fourth, one theoretical shortcoming of the nonlinear (logistic) adjustment is that, unlike the optimal linear adjustment, it is not guaranteed to be more efficient than no adjustment. We address this issue in Section (ref) below.
Following the lead of CF21, we can treat the nonlinear (logistic) adjustments as regressors and obtain the optimal linear coefficients as proposed in Section (ref). Let $\theta_{a,s}^{OLS} = (h_{a,s}^{OLS},\underline{\theta }_{a,s}^{OLS})$ be the probability limit of $\hat \theta_{a,s}^{OLS}$ defined in (ref). If $\beta_{a,s} ^{MLE}$ were known, the nonlinear (logistic) adjustment can be viewed as a linear adjustment. Specifically, denote
where $d_{\Psi}$ is the dimension of $\Psi_{i,s}$. Then, the nonlinear (logistic) adjustment can be written as
Similarly, we can replicate no adjustments and the optimal linear adjustments with $\Phi_{i,s}$ defined in (ref) as regressors by letting
with $(t_{a,s},b_{a,s}) = 0$ and $(t_{a,s},b_{a,s}) = (t_{a,s}^{L} ,b_{a,s}^{L})$, respectively, where
Based on Theorem (ref), we can further improve all three types of adjustments by setting the linear coefficients of $\Phi_{i,s}$ as
where $\tilde{\Phi}_{i,s} = \Phi_{i,s} - \mathbb{E}(\Phi_{i,s}|S_{i}=s)$. The final linear adjustments with $\theta_{a,s}^{F}$ and $\beta_{a,s}^{F}$ are
Because $\beta_{a,s}^{MLE}$ is unknown, we can replace it by its estimate proposed in Section (ref), i.e., define
Then, we define the estimators of $\theta_{a,s}^{F}$ and $\beta_{a,s}^{F}$ as
The corresponding feasible adjustments are
Theorem (ref) shows that by refitting nonlinear (logistic) adjustment in a linear regression with optimal linear coefficients, we can further improve the efficiency of the adjusted LATE estimator. As a by-product, $\hat{\tau}_{F}$ is guaranteed to be weakly more efficient than the LATE estimator without any adjustments ($\hat{\tau}_{NA}$).
In this section, we consider the nonparametric regression as the adjustments for our LATE estimator. Specifically, we use linear and logistic sieve regressions to estimate the true specifications $\mu^{Y}(a,s,X_{i})$ and $\mu^{D}(a,s,X_{i})$, respectively. For implementation, the nonparametric adjustment is exactly the same as nonlinear (logistic) adjustment studied in Section (ref). Theoretically, we will let the regressors $\mathring{\Psi }_{i,s}$ in (ref) be sieve basis functions whose dimensions will diverge to infinity as the sample size increases. For notational simplicity, we suppress the subscript $s$ and denote the sieve regressors as $\mathring{\Psi}_{i,n}\in\Re^{h_{n}}$, where the dimension $h_{n}$ can diverge with the sample size. The corresponding feasible regression adjustments are
where
We finally denote the corresponding adjusted LATE estimator as $\hat{\tau}_{NP}$.
Assumption (ref) is standard for linear and logistic sieve regressions. We refer to HIR03 and c07 for more discussions. The quantity $\zeta(h_{n})$ in Assumption (ref)(iv) depends on the choice of basis functions. For example, $\zeta(h_{n}) = O(h_{n}^{1/2})$ for splines and $\zeta(h_{n}) = O(h_{n})$ for power series.
The nonlinear (logistic) and nonparametric adjustments are numerically identical if the same set of regressors are used. Theorem (ref) then shows that the nonlinear (logistic) adjustment with technical regressors performs well because it can closely approximate the correct specification. Under the asymptotic framework that the dimension of the regressors diverges to infinity and the approximation error converges to zero, the nonlinear (logistic) adjustment can be viewed as the nonparametric adjustment, which achieves the SEB. In fact, if we estimate both $\mu^Y(a,s,X)$ and $\mu^D(a,s,X)$ by linear sieve regressions, under similar conditions to Assumption (ref), we can show that such an adjusted estimator also achieves the SEB. So does anseletal2018's (anseletal2018) S estimator when their $X_i$ is replaced by sieve bases of $X_i$ because it is asymptotically equivalent to our estimator L with optimal linear adjustment.
In this section, we consider the case where the regressor $\mathring{\Psi}_{i,n}\in\Re^{p_{n}}$ has dimension $p_{n}$ that can be much higher than $n$. In this case, we can no longer use the nonlinear (logistic) (nonparametric) adjustment method. Instead, we need to regularize the least squares and logistic regressions. Specifically, let
and the corresponding adjusted LATE estimator is denoted as $\hat{\tau}_{R}$, where
where $\{\varrho_{n,a}(s)\}_{a=0,1,s\in\mathcal{S}}$ are tuning parameters, $\hat{\Omega}^{b}=\text{diag}(\hat{\omega}_{1}^{b},\cdots,\hat{\omega }_{p_{n}}^{b})$ is a diagonal matrix of data-dependent penalty loadings for $b=D,Y$, and $\|\cdot\|_1$ is the $\ell_1$ norm.\footnote{ We provide more details about $\hat{\Omega}^{b}$ in Section (ref) of the Online Supplement.}
We maintain the following assumptions for Lasso and logistic Lasso regressions.
Assumption (ref) is standard in the literature and we refer interested readers to BCFH13 for more discussion.
Due to the approximate sparsity, the Lasso method consistently estimate the correct specification, which explains why the corresponding estimator can achieve the SEB.
Three data generating processes (DGPs) are used to assess the finite sample performance of the estimation and inference methods introduced in the paper. Suppose that
where $\{X_{i}, Z_{i}\}_{i\in[n]}, \alpha(\cdot, \cdot)$, $\{a_{i}, b_{i}, c_{i}\}_{i=0,1}$ and $\{\varepsilon_{j,i}\}_{j\in[4], i\in[n]}$ are specified as follows.
For each data generating process, we consider the four randomization schemes (SRS, WEI, BCD, SBR) defined as in Examples (ref)--(ref) in Appendix (ref), respectively. Specifically, for WEI and BCD, we set $f(x) = (1-x)/2$ and $\lambda = 0.75$, respectively.
We compute the true LATE effect $\tau_{0}$ using Monte Carlo simulations, with sample size being 10,000 and the number of Monte Carlo simulations being 1,000. We gauge the size and power of various tests by testing the hypotheses $H_{0}: \tau=\tau_{0}$ and $H_{0}: \tau=\tau_{0}+1$, respectively. All the tests are carried out at 5% level of significance, and with the number of Monte Carlo simulations being 10,000.
For DGPs(i)-(ii), we consider the following estimators.
For DGP(iii), we consider the estimator with no adjustments (NA), and the lasso estimators $\hat{\theta}_{a,s}^{R}$ and $\hat{\beta}_{a,s}^{R}$ defined in ((ref)) with $\mathring{\Psi}_{i,n}=(1,\Psi_{i,n}^{\top})^{\top}=(1,X_{i}^{\top})^{\top}$. The tuning parameters are choosing as: $\varrho_{n,a}(s) = 1.1 \sqrt{n_{a}(s)}F_{N}^{-1}( 1-1/(p_{n}\log(n_{a}(s))))$.
Tables (ref)-(ref) present the empirical sizes and powers of the true null $H_{0}: \tau=\tau_{0}$ and false null $H_{0}: \tau=\tau_{0}+1$ under DGPs (i)-(iii), respectively. We also report the ratio of the median length of the confidence intervals of a particular estimator to that of the NA estimator is in the corresponding bracket. Note that none of the working models in DGPs (i)-(iii) is correctly specified. Consider DGP (i). When $n=200$, both the NA and TSLS estimators are slightly under-sized. Both the NP and SNP estimators are oversized because the numbers of sieve regressors are relatively large compared to the sample size, while the R estimator has the correct size thanks to the Lasso selection of the sieve regressors. The L estimator performs the same as the S estimator. All other estimators have sizes close to the nominal level of 5%. This confirms that our estimation and inference procedures are robust to misspecification.
In terms of power, the NA estimator has the lowest power, corroborating the belief that one should carry out the regression adjustment whenever covariates correlate with the potential outcomes. The powers of the other estimators are much higher. In particular, the power of the F estimator is higher than those of the NA, TSLS, L, and NL estimators, which is consistent with our theory that the F estimator is weakly more efficient than those estimators. The NP, SNP, and R estimators enjoy the highest powers as a nonparametric model could approximate the true specification very well. The NP and SNP estimators have more size distortions than the R estimator when the sample size is 200. When the sample size is increased to 400, virtually all the sizes and powers of the estimators improve, and all the observations continue to hold.
We also report the ratio of the median length of the confidence intervals of a particular estimator to that of the NA estimator in the corresponding parentheses. Generally speaking, the confidence intervals of the TSLS and adjusted estimators (L, NL, F, NP, and R) are 20%-30% shorter, in terms of the median, than that of the NA estimator.
Most observations uncovered in DGP (i) carry forward to DGP (ii). Two new patterns emerge. First, the powers of the L, S, NL, F, NP, SNP, and R estimators are much higher than those of the NA and TSLS estimators. Second, the ratio of the median length of the confidence intervals of the TSLS estimator is as wide as that of the NA estimator, whereas the confidence intervals of the adjusted estimators (L, NL, F, NP, and R) become 25%-40% shorter, in terms of the median, than that of the NA estimator. This is probably because the true specifications for $Y_{i}(a)$ become more nonlinear.
We now consider DGP (iii). In this setting, only the NA and R estimators are feasible. When $n=200$, both estimators have the correct sizes but the R estimator has considerably higher power. When $n=400$, the sizes of these two estimators remain relatively unchanged, while their powers improve with a diverging gap. The confidence intervals of the R estimator are 60%-65% shorter, in terms of the median, than that of the NA estimator.
In Section (ref) of the Online Supplement, we simulate data with heterogeneous $\pi(s)$. We find that all estimators except TSLS have their empirical rejection rates close to the nominal size of 5% under the null. TSLS, on the other hand, has around 15% rejection rate when $n=1200$. This indicates the TSLS estimator is inconsistent when $\pi(s)$ is heterogeneous, in line with Theorem (ref).
If researchers want to use parametric adjustments without tuning parameters, we recommend the F estimator, which is guaranteed to be weakly more efficient than TSLS, L, and NL estimators. Regressors $\Psi_{i,s}$ can include linear, quadratic and interaction terms of the original covariates. If researchers want to achieve the SEB by using sieve bases and/or the dimension of covariates is high relative to the sample size, we recommend the R estimator.
Banking the unbanked is considered to be the first step toward broader financial inclusion -- the focus of the World Bank's Universal Financial Access 2020 initiative.\footnote{https://www.worldbank.org/en/topic/financialinclusion/brief/achieving-universal-financial-access-by-2020} In a field experiment with a CAR design, dupasetal2018 examined the impact of expanding access to basic saving accounts for rural households living in three countries: Uganda, Malawi, and Chile. In particular, apart from the intent-to-treat effects for the whole sample, they also studied the local average treatment effects for the households who actively used the accounts. This section presents an application of our regression adjusted estimators to the same dataset to examine the LATEs of opening bank accounts on savings balance-- a central outcome of interest in their study.
We focus on the experiment conducted in Uganda. The sample consists of 2,160 households who were randomized with a CAR design. Specifically, within each of 41 strata formed by gender, occupation, and bank branch, half of households were randomly allocated to the treatment group, the other half to the control one. Households in the treatment group were then offered a voucher to open bank accounts with no financial costs. However, not every treated household ever opened and used the saving accounts for deposit. In fact, among those households with treatment assignment, only 41.87% of them opened the accounts and made at least one deposit within 2 years. Subject compliance is therefore imperfect in this experiment.
The randomization design apparently satisfies statements (i), (ii) and (iii) of Assumption (ref). The target fraction of treatment assignment is 1/2. Because $\max_{s\in\mathcal{S}}|\frac{B_{n}(s)}{n(s)}|\approx0.056$, it is plausible to claim that Assumption (ref)(iv) is also satisfied. Since households in the control group need to pay for the fees of opening accounts while the treated ones bear no financial costs, no-defiers statement in Assumption (ref)(v) holds plausibly in this case.
One of the key analyses in dupasetal2018 is to estimate the treatment effects on savings for active users -- households who actually opened the accounts and made at least one deposit within 2 years. We follow their footprints to estimate the same LATEs at savings balance.\footnote{Savings balance includes savings in formal financial intuitions, mobile money, cash at home or in secret place, savings in ROSCA/VSLA, savings with friends/family, other cash savings, total formal savings, total informal savings, and total savings (See dupasetal2018 for details). We use data from the first follow-up survey and exclude other cash savings because only 2% of the households in the sample reported having it.} To maintain comparability, for each outcome variable, we also keep $X_{i}$ similar to those used in dupasetal2018 for our adjusted estimators.\footnote{The description of these estimators is similar to that in Section (ref). Except for savings in formal financial institutions, mobile money, and total formal savings, $X_{i}$ includes baseline value for the outcome of interest, baseline value of total income, and a dummy for missing observations. For savings in formal financial institutions, mobile money, and total formal savings, since their baseline values are all zero, we set $X_{i}$ as the baseline value of total savings, baseline value of total income, and a dummy for missing observations.} Due to the low dimension of covariates used in the regression adjustments, we focus on the performance of the methods “NA", “TSLS", “L", “NL", and “F".
Table (ref) presents the LATE estimates and their standard errors (in parentheses) estimated by these methods.\footnote{For each outcome variable, we filter out the observations with missing values of outcome variables or the strata with less than 10 observations. The total trimmed observations are less than 10% of the whole sample in dupasetal2018).} These results lead to four observations. First, consistent with the theoretical and simulation results, the standard errors for the LATE estimates with regression adjustments are lower than those without adjustments. This observation holds for all the outcome variables and all the regression adjustment methods. Over the eight outcome variables, the standard errors estimated by regression adjustments are on average around 8% lower than those without adjustment. In particular, when the outcome variable is total informal savings, the standard errors obtained via the further improvement adjustment -- “F" method is about 18% lower than those without adjustment. This means that regression adjustments, with the similar covariates used in dupasetal2018, can achieve sizable efficiency gains in estimating the LATEs.
\newcolumntype{L}{>{\arraybackslash}X} \newcolumntype{C}{>{\arraybackslash}X}
Second, the standard errors for the regression-adjusted LATE estimates are mostly lower than those obtained by the usual TSLS procedure. Especially, when the outcome variables are mobile money and total informal savings, the standard errors obtained via “F" method are about 7.1% and 5%, respectively, lower than those by TSLS. When the outcome variable is savings in friends/family, the standard error estimated by the optimal linear adjustment -- “L" method is around 6.7% lower than that obtained by TSLS. This means that, compared with our regression-adjusted methods, TSLS is generally less efficient to estimate the LATEs under CAR.
Third, the standard errors for the LATE estimates with regression adjustments are similar in terms of magnitude. This implies that all the regression adjustments achieve similar efficiency gain in this case.
Finally, as in dupasetal2018, for the households who actively use bank accounts, we find that reducing the cost of opening a bank account can significantly increase their savings in formal institutions. We also observe the evidence of crowd-out -- mainly moving cash from saving at home to saving in bank.