Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
57,824 characters · 5 sections · 0 citation commands
A Consistent Variance Estimator for 2SLS When Instruments Identify Different LATEs
Since the series of seminal papers by Imbens and Angrist (1994), Angrist and Imbens (1995), and Angrist, Imbens, and Rubin (1996), the local average treatment effect (LATE) has played an important role in providing useful guidance to many policy questions. The key underlying assumption is treatment effect heterogeneity, i.e. each individual has a different causal effect of treatment on outcome. Assume a binary treatment, $D_{i}$, and an outcome variable $Y_{i}$. Let $Y_{1i}$ and $Y_{0i}$ denote the potential outcomes of individual $i$ with and without the treatment, respectively. The heterogeneous individual treatment effect is $Y_{1i}-Y_{0i}$, but this cannot be identified because $Y_{1i}$ and $Y_{0i}$ are never observed at the same time. Therefore, researchers typically focus on the average treatment effect (ATE), $E[Y_{1i}-Y_{0i}]$. However, unless the treatment is randomly assigned, a naive estimate of ATE is biased because of selection into treatment.
Instrumental variables are used to overcome this endogeneity problem. If an instrument $Z_{i}$ which is randomly assigned, independent of $Y_{1i}$ and $Y_{0i}$, and correlated with the treatment $D_{i}$ is available, then ATE of those whose treatment status can be changed by the instrument, thus the local ATE, is identified. Assume $Z_{i}$ is binary and define $D_{1i}$ and $D_{0i}$ be $i$'s treatment status when $Z_{i}=1$ and $Z_{i}=0$, respectively. The LATE theorem of Imbens and Angrist (1994) shows that
That is, the instrumental vaiables (IV) estimand (or the Wald estimand) is equal to the ATE of those with $D_{1i}>D_{0i}$, who are called compliers. Those who take the treatment regardless of the instrument value, $D_{1i}=D_{0i}=1$, are always-takers, and those who do not take the treatment anyway, $D_{1i}=D_{0i}=0$, are never-takers. We cannot identify ATE of always-takers and never-takers in general. By the monotonicity assumption of Imbens and Angrist, we exclude defiers who behave in the opposite way with compliers, $D_{1i}<D_{0i}$. Since the compliers are specific to the instrument $Z_{i}$, LATE is instrument-specific.
The above setting can be generalized to situations where the number of (excluded) instruments is greater than the number of endogenous variables. The two-stage least squares (2SLS) estimator is commonly used to estimate the causal effect in such cases. Without loss of generality, consider mutually exclusive binary instruments, $Z_{i}^{j}$ for $j=1,...,q$. Let $D_{zi}^{j}$ be $i$'s potential treatment status when $Z_{i}^{j}=z$ where $z=0,1$. Each instrument identifies a version of LATE because compliers may differ for each $Z_{i}^{j}$. The 2SLS estimand is a weighted average of treatment effects for the instrument-specific compliers:
where $0\leq\xi_{j}\leq1$ and $\sum_{j}\xi_{j}=1$. This is first shown by Imbens and Angrist (1994) for a discrete instrument which can be written as mutually exclusive binary instruments. Angrist and Imbens (1995) and Heckman and Vytlacil (2005) extend this result to multi-valued treatment with covariates, and a continuous instrument with covariates, respectively. These works provided theoretical foundations to interpret 2SLS point estimates as a weighted average of LATEs, and empirical researchers have done so, either explicitly or implicitly.
If the 2SLS estimand is a weighted average of more than one LATE, then the over-identifying restrictions test (the J test, hereinafter) which is typically conducted along with 2SLS would reject the null hypothesis that the moment condition is correctly specified.\footnote{The J test can also reject the null due to invalid instruments. Kitagawa (2015) proposes a specification test for instrument validity under treatment effect heterogeneity.} What is less well known and often overlooked in the literature is that the conventional standard errors are no longer correct under misspecification of the moment condition. This fact has been neglected and the standard errors have been routinely calculated assuming the LATEs are identical. I derive the asymptotic distribution of 2SLS when the estimand is a weighted average of LATEs and propose a consistent estimator of the asymptotic variance robust to multiple LATEs. The correct standard error based on the proposed variance estimator (the multiple-LATEs-robust standard error, hereinafter) can be substantially different from the conventional heteroskedasticity-robust one even for a large sample size, or even for p-values above any usual significance level.
Two recent papers cover similar topics. Koles\'{a}r (2013) shows that under treatment effect heterogeneity the 2SLS estimand is a convex combination of LATEs while the limited information maximum likelihood (LIML) estimand may not. Angrist and Fernandez-Val (2013) propose an estimand for new subpopulations by reweighting covariate-specific LATEs. However, neither of the two papers considers correct variance estimation of 2SLS.
In the next Section, I show that the postulated moment condition of 2SLS is misspecified when there is more than one LATE. The asymptotic distribution of 2SLS estimators in such a case is derived, and a consistent variance estimator is proposed. In Section (ref), I discuss practical implications of using 2SLS with multiple instruments. In particular, Angrist and Krueger (1991), Angrist and Evans (1998), and Thornton (2008) are replicated and the originally reported conventional standard errors are compared to the multiple-LATE-robust standard errors. Section (ref) presents simulation results that show (i) the multiple-LATEs-robust standard error is closer to the standard deviation of the 2SLS estimator than the conventional standard error, (ii) the bias in the conventional standard error tends to be large when the first stage F statistic or the p-value of the J statistic are small, and (iii) the conclusion of $t$ tests may change if the multiple-LATEs-robust standard error is used regardless of the magnitude of the bias in the conventional standard error. Section (ref) concludes. The proofs of propositions are collected in the appendix.
I first define moment condition models for IV and 2SLS estimators to derive the asymptotic distribution. I maintain the assumption that the treatment variable and instruments are binary for simplicity of exposition, but this will be relaxed later in this section. The observed outcome can be written as
where $\rho_{i}=Y_{1i}-Y_{0i}$ and $\eta_{i}=Y_{0i}-E[Y_{0i}]$. Since the individual treatment effect $\rho_{i}$ cannot be identified, a version of ATE becomes the parameter of interest. Let $\rho$ be the parameter, and $\alpha$ be a nuisance parameter for the intercept. Rewriting (ref) in a familiar regression form using $\alpha$ and $\rho$, we get
It is straightforward to see that regressing $Y_{i}$ on a constant and $D_{i}$ yields
which are the solutions to the moment condition
Although the moment condition (ref) is satisfied at $(\alpha_{ols},\rho_{ols})$, $\rho_{ols}$ does not identify an interesting population parameter (the ATE on the treated, in this case) because $D_{i}$ is endogenous, i.e. $Y_{0i}\not\perp D_{i}$, and
Now suppose that there is a binary instrument $Z_{i}^{1}$ which satisfies Conditions 1-3 of Imbens and Angrist (1994). The IV estimator is based on the moment condition
where the unique solution $(\alpha_{IV}^{1},\rho_{IV}^{1})$ is given by
The second equality of (ref) holds by Theorem 2 of Imbens and Angrist (1994). For OLS or IV moment conditions, the model is just-identified because the number of unknown parameters is equal to the number of equations in the moment condition. This implies that there always exists a solution that makes the moment condition equal to zero.\footnote{This means that $\rho_{ols}$ should be interpreted as the projection coefficient. For IV, (ref) holds even if $Z_{i}^{1}$ is not independent of $Y_{0i}$, and thus $E[Z_{i}^{1}\eta_{i}]\neq0$. In this case, the second equality of (ref) does not hold. In other words, $0=E[Z_{i}^{1}e_{i}^{1}]$ is not the instrument validity condition under treatment effect heterogeneity.} The asymptotic distributions of the OLS and IV estimators are derived by expanding the first-order conditions (FOC) around their estimands and using the fact that the respective moment condition holds. Thus, the conventional standard errors are correct under treatment effect heterogeneity.
This positive conclusion does not hold if there are more instruments than the endogenous parameters. In this case, the moment condition is over-identified and the assumption that there exists a unique solution to the moment condition may be violated. Suppose that there are two valid instruments, $Z_{i}^{1}$ and $Z_{i}^{2}$. If we use each instrument one at a time, we would get $(\alpha_{IV}^{1},\rho_{IV}^{1})$ and $(\alpha_{IV}^{2},\rho_{IV}^{2})$, where each corresponds to a different LATE. Assume $\rho_{IV}^{1}\neq\rho_{IV}^{2}$. I emphasize that they are the unique solutions to the corresponding moment conditions. The 2SLS estimator that uses both instruments at the same time is calculated by estimating $D_{i} = \delta + Z_{i}^{1}\pi_{1} + Z_{i}^{2}\pi_{2} + u_{i}$ by OLS in the first stage to produce the fitted value $\hat{D}_{i}$, and then by estimating $Y_{i} = \alpha + \hat{D}_{i}\rho + e_{i}$ by OLS in the second stage. This estimator is equivalent to the linear GMM using the sample mean of outer-products of the instrument vector as a weight matrix, based on the moment condition
for a unique parameter vector $(\alpha^{*},\rho^{*})$. To see if there exists such a solution, solve the first equation for $\alpha^{*}$, substitute it into the second and third equations, and solve them for $\rho^{*}$ to have
By the LATE theorem, (ref) implies that $\rho^{*}=\rho_{IV}^{1}=\rho_{IV}^{2}$, but this contradicts the assumption that $\rho_{IV}^{1}\neq\rho_{IV}^{2}$. In other words, if $(\alpha_{IV}^{1},\rho_{IV}^{1})$ were the unique solution to the first two equations of (ref), then it should satisfy the last equation, but this implies that the two LATEs are the same. Thus, there does not exist a unique parameter that satisfies (ref) and the moment condition is misspecified. It is stressed that misspecification of the moment condition does not imply invalidity of the instruments under treatment effect heterogeneity.
This definition is not directly related to functional form or data-generating process misspecification. For instance, the FOC of the quasi-maximum likelihood forms a just-identified moment condition with a unique solution. Therefore, it is a correctly specified moment condition, although the distribution is misspecified. When the moment condition is misspecified, the GMM estimator is consistent for a value that minimizes the population criterion function. This value is called the pseudo-true value, and it is the 2SLS estimand in this context. Imbens and Angrist (1994) show that the 2SLS estimand is
Evaluated at $(\alpha_{0},\rho_{0})$, the moment function (ref) does not hold. This conclusion continues to hold with more than two instruments and covariates.
Misspecified moment conditions under treatment effect heterogeneity have important implications. First, the J test will reject the null hypothesis (ref) asymptotically. It is not surprising that researchers often face a significant J test statistic when multiple instruments are used. If we can rule out the possibility of invalid instruments either by a statistical test such as Kitagawa (2015) or by an economic reasoning, the rejection is due to treatment effect heterogeneity. Thus, the J test has little relevance once heterogeneity is already assumed. Second, the asymptotic variance of 2SLS will be different from the standard one, and the conventional heteroskedasticity-robust variance estimator would be inconsistent. It is surprising that this has been overlooked in the literature. In the following propositions, I derive the asymptotic distribution of 2SLS and propose a consistent variance estimator. This variance estimator should always be used to calculate the standard error of 2SLS, regardless of the J test results.
To formally derive the asymptotic distribution, I consider the model (ref) with covariates where the treatment variable and instruments can take multiple values or even be continuous as in Angrist and Imbens (1995) and Heckman and Vytlacil (2005). Although I assume a single endogenous variable, extensions to multiple endogenous variables are straightforward. Assume that there are valid instruments $Z_{i}^{1},Z_{i}^{2},...,Z_{i}^{q}$. Let $(Y_{i},\mathbf{X}_{i},\mathbf{Z}_{i})_{i=1}^{n}$ be an iid sample, where $\mathbf{X}_{i}=(\mathbf{W}_{i}',D_{i})'$, $\mathbf{Z}_{i}=(\mathbf{W}_{i}',Z_{i}^{1},\cdots,Z_{i}^{q})'$, and $\mathbf{W}_{i}$ be an $l\times1$ vector of covariates including a constant. The first and second stages are
where $\boldsymbol{\beta} = (\boldsymbol{\gamma}',\rho)'$, and $\boldsymbol{\gamma}$, $\boldsymbol{\delta}$, and $\boldsymbol{\pi}$ are conformable parameters. The 2SLS estimator is
where $\mathbf{X}\equiv (\mathbf{X}_{1},\cdots,\mathbf{X}_{n})'$ is an $n\times (l+1)$ matrix, $\mathbf{Z}\equiv (\mathbf{Z}_{1},\cdots,\mathbf{Z}_{n})'$ is an $n\times (l+q)$ matrix, and $\mathbf{Y} \equiv (Y_{1},\cdots,Y_{n})'$ is an $n\times 1$ vector. The following proposition establishes the asymptotic distribution of 2SLS estimators when there is more than one LATE in the general setting.
The next proposition proposes a consistent estimator of the asymptotic variance matrix of 2SLS robust to multiple-LATEs.
The formula of $\boldsymbol{\hat{\Sigma}_{MR}}$ is different from that of the conventional heteroskedasticity-robust variance estimator:
Under constant treatment effects, both $\boldsymbol{\hat{\Sigma}_{MR}}$ and $\boldsymbol{\hat{\Sigma}_{C}}$ have the same probability limit, but they are generally different in finite samples. $\boldsymbol{\hat{\Sigma}_{MR}}$ is consistent for the true asymptotic variance matrix even when the postulated moment condition is misspecified, and thus can be used regardless of whether there is one or more than one LATE. In contrast, $\boldsymbol{\hat{\Sigma}_{C}}$ is consistent only if the underlying LATEs are identical. This is also true for the standard errors based on $\boldsymbol{\hat{\Sigma}_{MR}}$ and $\boldsymbol{\hat{\Sigma}_{C}}$.
When there is a single endogenous variable without covariates, Proposition 1 coincides with the result in the proof of Theorem 3 of Imbens and Angrist (1994) when the first stage is known but needs to be estimated.\footnote{There are typos in the proof of Theorem 3 of Imbens and Angrist (1994). Their matrix $\Delta$ shoud read $\Delta = \left(
\right)$.} Their derivation is based on the stacked moment condition that consists of FOCs of the first and second stages, which is a special case of two-step estimators of Newey and McFadden (1994). They use the condition that the population fitted value of the endogenous variable is uncorrelated with $e_{i}$, where $e_{i}$ is defined in Proposition 1. In other words, the estimated first stage is used as an instrument. For example, the condition is $E[(\delta+\pi_{1}Z_{i}^{1}+\pi_{2}Z_{i}^{2})e_{i}]=0$ for a two instruments case, which does not necessarily imply $E[Z_{i}^{1}e_{i}]=E[Z_{i}^{2}e_{i}]=0$. This makes their asymptotic variance and its estimator robust to violations of the underlying moment condition of 2SLS, $E[Z_{i}^{1}e_{i}]=E[Z_{i}^{2}e_{i}]=0$. Thus, they coincide with $\boldsymbol{\Sigma_{MR}}$ and $\boldsymbol{\hat{\Sigma}_{MR}}$. Even in such cases, however, their formula has not been used in practice. Econometric software packages such as Stata do not estimate their asymptotic variance but estimate the standard GMM one assuming correct specification, which leads to wrong standard errors. The main contribution of this paper is to make a novel observation that 2SLS using multiple instruments under treatment effect heterogeneity is a special case of misspecified GMM of Hall and Inoue (2003). Specifically, Proposition 1 is a special case of their Theorem 2.
Since the 2SLS estimator is the linear GMM using $\left(\mathbf{Z}'\mathbf{Z}\right)^{-1}$ as a weight matrix, we may consider an alternative GMM estimator based on another weight matrix. This will lead to a different weighted average of LATEs, which may be more appealing than the conventional 2SLS estimand. Let $E[\mathbf{L}_{i}\mathbf{L}_{i}']$ be an alternative symmetric positive definite matrix where $\mathbf{L}_{i}$ is an $(l+q)\times1$ vector, and let $\left(\mathbf{L}'\mathbf{L}\right)^{-1}$ be the sample weight matrix, where $\mathbf{L}$ is an $n\times(l+q)$ matrix. The alternative GMM estimator based on the same moment condition but a different weight matrix is given by
Let $\boldsymbol{\beta_{a}}$ be the probability limit of $\boldsymbol{\tilde{\beta}}$. The asymptotic distribution of $\sqrt{n}(\boldsymbol{\tilde{\beta}}-\boldsymbol{\beta_{a}})$ and a consistent variance estimator can be obtained by replacing $\mathbf{Z}_{i}\mathbf{Z}_{i}'$ with $\mathbf{L}_{i}\mathbf{L}_{i}'$, $\mathbf{Z}'\mathbf{Z}$ with $\mathbf{L}'\mathbf{L}$, and $\hat{e}_{i}$ with $\tilde{e}_{i}=Y_{i}-\mathbf{X}'_{i}\boldsymbol{\tilde{\beta}}$, whenever they appear in Propositions 1 and 2.
How often are researchers interested in a weighted average of LATEs? Probably more common than one might think. First of all, a discrete instrument with a binary treatment identifies a weighted average of LATEs where each LATE corresponds to a pair of two adjacent values of the instrument. This is the parameter defined in Theorem 2 of Imbens and Angrist (1994), and it is one way to average different LATEs because a discrete instrument can be written as a set of mutually exclusive binary instruments. Second, a binary instrument identifies the average causal response (ACR, Angrist and Imbens, 1995) when the treatment variable takes multiple values, e.g. years of schooling. ACR is a weighted average of LATEs for each value of the treatment variable, and is a widely accepted concept in the literature, e.g. Bleakley and Chin (2004), Lochner and Moretti (2004), Elder and Lubotsky (2009).
On the other hand, researchers typically use 2SLS with multiple instruments to increase efficiency and because it is often unclear which instrument gives the strongest identification, assuming that the treatment effect heterogeneity is minimal, if not constant. Examples include Angrist and Chen (2011), Angrist and Evans (1998), Angrist and Krueger (1991, 1994), Angrist, Lavy, Schlosser (2010), Evans and Garthwaite (2012), Evans and Lien (2005), Siminski (2013), Stephens Jr. and Yang (2014), and Thornton (2008), among others. Even in these cases, the multiple-LATEs-robust standard errors can provide safeguards against potential violations of minimal or constant treatment effect heterogeneity.
For the rest of this section, I replicate three well-known studies and show the multiple-LATEs-robust standard errors can be substantially different from the reported ones. The first example is Angrist and Krueger (1991), who study the returns to education. The authors avoid endogeneity of education by instrumenting it with quarter of birth (QOB). Individuals who were born at the end of the year enter school at a younger age compared with their classmates. As a result, they are required to take more compulsory schooling before they reach a legal dropout age. Angrist and Krueger estimate the following 2SLS model:
where $E_{i}$ is education, $X_{i}$ is a vector of covariates including a constant, $Y_{ic}$ is year of birth (YOB), $Q_{ij}$ is QOB, and $W_{i}$ is weekly wage. If we assume that $X_{i}$ only contains a constant, then the first stage equation (ref) is saturated. In this case, the 2SLS estimand is a weighted average of returns to education where averaging takes place on three different levels. First, for each level of education, it is ATE for those who would have additional schooling due to their QOB and YOB. Second, it is the ACR averaged over different levels of education which takes values from 0 to 20. Lastly, the ACR is averaged over different years of birth. By using interaction terms between QOB and YOB dummies as instruments in the first stage, it is assumed that the level of education varies with the QOB-YOB interactions, and thus is expected to increase the model fit. If the authors were interested in the returns to education for each YOB, interaction terms between YOB and education as well as YOB dummies could be included in the second stage. Instead, what is estimated in the second stage is an average of returns to education across different years of birth while controlling for the level effect of YOB, because there is no reason that a particular year, e.g. men born in 1930, is more interesting than the cohort of those born in 1930-1939. Therefore it is important to correctly calculate the standard error of the point estimates in this example.
Table (ref) shows replication results of Tables IV-VI in Angrist and Krueger along with the multiple-LATEs-robust standard errors (Column MR, in bold). The results for covariates are suppressed. There are a few interesting findings. First, even with large sample sizes, the two standard errors are substantially different. The conventional ones (Column C) are underestimated in all specifications. Second, large p-values do not necessarily mean that the two standard errors are similar. Third, the first stage F statistics are below the rule of thumb, 10, except for the Column (0), indicating that the instruments may be weak. One may wonder if the difference between the two standard errors are driven by the weak instruments. In the next section, I provide a simulation result that the two standard errors can be substantially different even when the F statistics are very large, although there is a weak negative relationship between the difference of the two standard errors and the magnitude of the F statistic. Finally, point estimates averaged over a large set of instruments are more robust. This case is illustrated by Table 3 Column (0) where three QOB dummies are used as only instruments with YOB dummies as covariates. Surprisingly, the return to education is estimated to be negative and significant. Further inspection reveals that it is a linear combination of three IV estimates, -0.0191 (0.0272) using only the first quarter as an instrument, -1.3167 (5.2517) using the second quarter, 0.2858 (0.1932) using the third quarter, where the numbers in parentheses are conventional IV standard errors. Apparently, an imprecisely estimated point estimate with the second QOB is the main reason for the negative point estimate in Column (0). In practice, a researcher might get around a similar problem by using a different instrument, but a better alternative is to get the 2SLS estimate based on a larger set of instruments.
The second example is Angrist and Evans (1998) who use the sex of mother's first two children as instruments to estimate the effect of family size on mother's labor supply. The instruments two-boys and two-girls are based on the fact that American parents tended to have a third child when their first two children were of the same sex. Each of the instruments identifies LATE of those whose fertility was affected by their children's sex mix, and the two LATEs are not necessarily the same. For instance, a subpopulation with certain cultural background may have relatively large number of two-girls compliers, and their ATE may be lower than that of two-boys compliers in the population. The 2SLS model used by Angrist and Evans is
where $M_{i}$ is an indicator for more than 2 children, $TB_{i}$ and $TG_{i}$ are indicators for first two boys and first two girls, $X_{i}$ is a vector of covariates including mother's age, age at first birth, race and Hispanic indicators, a firstborn boy indicator, and a constant, and $Y_{i}$ is an indicator for whether the respondent worked for pay in the Census year. The OLS estimate of $\rho$ is -.167, but it is argued to exaggerate the causal effect of fertility on female labor supply due to selection bias. Using the instruments one at a time, we get the IV estimates -.201 (.045) for two-boys and -.059 (.035) for two-girls instrument, where the numbers in parentheses are conventional IV standard errors. The estimates are quite different, and it is also difficult to compare them with the OLS estimate because the latter is for the whole population while the IV estimates are for complier subpopulations. Since the ultimate goal is to estimate the overall effect of having more than two children on mother's labor supply, one way to proceed is to calculate an average of the two IV estimates. 2SLS estimand is a particular weighted average where the weights are calculated based on the relative strength of each instrument.\footnote{In this example, 2SLS estimand is not exactly equal to a weighted average of covariate-specific LATEs, because the first stage is not fully saturated. This may weaken its causal interpretation. Nevertheless, Angrist (2001) shows that 2SLS estimates are very similar to those based on a semiparametric procedure of Abadie (2003) which allows robust causal interpretations, and Angrist and Pischke (2009) argue that this is likely to hold for other cases.}
Table (ref) shows replication results of Table 7 in Angrist and Evans (1998). First of all, the 2SLS estimates are smaller in magnitude than the OLS estimates, providing evidence that the OLS overestimates the true effect. Second, unlike the replication of Angrist and Krueger (1991), the two standard errors are almost the same, even when the p-values are quite small. Since they are similar, there is a sizable gain in precision by combining the two instruments compared with using a single instrument even when we use the multiple-LATEs-robust standard errors. Note that standard errors converge in probability to zero. Thus, the similarity of two standard errors in Table (ref) should not be misinterpreted as the similarity of the asymptotic variances. Lastly, the 2SLS point estimates are weighted averages of the two IV estimates, where the weight for two-boys instrument is .38. Since the weight is completely determined by the first stage, the same weight is used across different dependent variables in the second stage. In this example, two-boys instrument receives less weight because the first-stage coefficient is smaller, which implies that the absolute size of the compliers is smaller than that of two-girls instrument. The multiple-LATES-robust standard error can be computed not only for 2SLS but also for any other weighted averages of LATEs, as long as they can be written as a GMM estimator as in (ref).
The last example is Thornton (2008), who studies the impact of learning HIV status on purchasing condoms in rural Malawi using randomly assigned monetary incentives and distance from results centers as instruments. The dependent variable $Y_{ij}$ is an indicator for condom purchase, reported buying condoms, or reported having sex at the follow up survey, or the number of condoms bought for person $i$ in village $j$. The endogenous variables are $G_{ij}$ and $G_{ij}\times HIV_{ij}$, where $G_{ij}$ indicates knowledge of HIV status and $HIV_{ij}$ indicates an HIV-positive diagnosis. In particular, the interaction term $G_{ij}\times HIV_{ij}$ is included in the main equation to investigate the differential effect of receiving a positive HIV diagnosis. Both endogenous variables are instrumented with the same set of variables. The 2SLS model used by Thornton is
where $Z_{ijc}$ for $c=1,...,5$ are being offered any incentive, the amount of the incentive, the amount of the incentive squared, the distance from the HIV result center, and distance-squared, and $X_{i}$ is a vector of covariates including an indicator for male, $Male_{ij}$, as well as age, age-squared, a district dummy, $HIV_{ij}$, and a constant. The first stage for $G_{ij}\times HIV_{ij}$ has the same set of regressors. To account for the differential effects of gender and HIV-positive diagnosis on getting the test results, the interaction terms are also used as instruments. The author notes that the resulting estimates of $\rho_{1}$ and $\rho_{2}$ are weighted averages of LATEs, but argues that differences in LATEs across instruments may be minimal, which justifies the use of the conventional standard errors.
Table (ref) show replication results of Table 7 in Thornton (2008). The two standard errors are quite different, especially for $\hat{\rho}_{2}$. The point estimate of $\rho_{2}$ with $Y_{ij}=$Number of condoms bought is significantly different from zero at 5% level using the conventional standard error, but it is no longer significant even at 10% level when the multiple-LATEs-robust standard error is used. Thus, even though the treatment effects heterogeneity is assumed to be minimal, one should report the multiple-LATEs-robust standard errors to make correct inferences. The J test is conducted using the cluster-robust variance matrix estimator and the p-values are under any reasonable nominal size. The column under $CD$ shows the Cragg-Donald statistic for multiple endogenous variables (Cragg and Donald, 1993), which can be used to measure the weak instrument problem using the Stock and Yogo (2005) weak instrument critical values under iid and homoskedasticity assumptions.
This section provides simulation evidence to answer three questions: (i) whether the multiple-LATEs-robust standard error correctly approximates the true standard deviation of the 2SLS estimator, (ii) whether the difference between the two standard errors is driven by the weak instruments, and (iii) whether there is a relationship between the difference in the magnitude of the two standard errors and the p-value of the J test.
I generate a random subsample without replacement from the 1980 Census Public Use Micro Samples dataset of Angrist and Evans (1998) that contains information on $254,654$ married women. The 2SLS model is the same as (ref) and (ref), except that an indicator for multiple second birth is also added as an instrument. Angrist and Evans show evidence that this instrument identifies LATE different from that of two-boys or two-girls instruments, and give further analysis on the issue. For simulation, I simply use three instruments together without introducing additional parameters that partly adjust for the difference in LATEs, so that the underlying moment condition model is over-identified and potentially misspecified.
Table (ref) shows the mean and standard deviation of the 2SLS estimator, and the means of the multiple-LATEs-robust (MR) and the conventional standard errors (C) for different dependent variables based on 10,000 replications for the sample sizes 1,000, 2,000, and 5,000. The proportion of the heteroskedasticity-robust first-stage F statistics smaller than the rule of thumb 10 is 0.25% for $n=1,000$, and zero for the other sample sizes. Thus, there is little concern for weak instruments. In an unreported simulation result, the multiple-LATEs-robust standard error tends to approximate the standard deviation of the 2SLS estimator well with marginal F statistics under 10. Across specifications, the conventional standard error based on $\boldsymbol{\hat{\Sigma}_{C}}$ underestimates the standard deviation. Therefore, inferences based on the conventional standard errors can be misleading. In contrast, the multiple-LATEs-robust standard error estimates the standard deviation more accurately. The last column shows the actual rejection probability of the J test at 5% significance level. Since the p-values for the full sample ($n=254,654$) are .0345, .1599, .0971, and .9158 for each dependent variable, respectively, the result shows that the J test exhibits very low power, at least for $Y_{i}=\textit{Worked for pay}$ and $Y_{i}=\textit{Hours per week}$.
Figure (ref) shows the relationships between the percentage difference between the two standard errors (the mutiple-LATEs robust standard error minus the conventional standard error divided by the average of the two) and the p-value of the J test, and the F statistic, respectively, for different sample sizes when $Y_{i}=\textit{Worked for pay}$. The results are similar for other dependent variables, and thus not reported. The multiple-LATEs-robust standard error is likely to differ much from the conventional one when the p-value of the J test or the F statistic is relatively small, though the negative correlation is quite weak for the F statistic. Points with a $+$ marker indicate that the conclusion of the $t$ tests changes with the correct standard error where the null hypothesis is that the coefficient equals to zero, i.e. statistically significant results may not be significant anymore (or the other way around). In particular, this happens regardless of the p-value of the J test or the magnitude of the F statistic.
In sum, the simulation result shows that the multiple-LATEs-robust standard error can be very different from the conventional one even with large F statistics and p-values of the J test. Furthermore, the conclusion of the empirical study may change if the multiple-LATEs-robust standard error is used.
Two-stage least squares (2SLS) estimators are widely used in practice. When heterogeneity is present in treatment effects, 2SLS point estimates can be interpreted as weighted averages of the local average treatment effects (LATE). I show that the conventional standard errors, typically generated by econometric software packages such as Stata, are incorrect in this case. The over-identifying restrictions test is often used to test the presence of heterogeneity, but it is not useful in this context because it can also reject due to invalid instruments. I provide a simple standard error formula for 2SLS which is correct regardless of whether there are multiple LATEs or not. In addition, this standard error is robust to invalid instruments, and can be used for bootstrapping to achieve asymptotic refinements under treatment effect heterogeneity. I recommend practitioners to always use the proposed multiple-LATEs-robust standard error for 2SLS.