EconBase
← Back to paper

Weak Instrumental Variables: Limitations of Traditional 2SLS and Exploring Alternative Instrumental Variable Estimators

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

49,337 characters · 14 sections · 54 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
titlepage\begin{center} Rheinische Friedrich-Wilhelms-Universität Bonn Master Programme in Economics { \onehalfspacing Term Paper for Research Module in Econometrics & Statistics} { \onehalfspacing Weak Instrumental Variables: Limitations of Traditional 2SLS and Exploring Alternative Instrumental Variable Estimators} { Submitted by:} { Aiwei Huang Madhurima Chandra Laura Malkhasyan} { \linespread{2} Supervised by: JProf. Dr. Dominik Liebl} \begin{center} { February 05, 2020} \end{center} \end{center}

\setcounter{page}{1} \chapter{Introduction} Instrumental variables estimation has gained considerable traction in recent decades as a tool for causal inference, particularly amongst empirical researchers. However, this has also highlighted the importance of taking a deeper look at the theoretical properties underlying one's estimator of choice, whether it is instrumental variable estimator or any other estimator. A case-in-point is the well-known 1991 quarter-of-birth study on returns to schooling estimates by angrist1991does and the equally famous rebuttal of their results published by bound1995problems.

The bound1995problems critique of the quarter-of-birth study in essence formed the starting point of the weak instruments literature. They provided simulation evidence that the instruments used by angrist1991does were very weak and this in turn led to misleading results. Interestingly, their study was not the first time that the poor finite-sample behavior of the instrumental variables estimator (two-stage least squares estimator in particular) was discussed. Previously, numerous authors had presented results on the magnitude of the finite-sample bias of the 2SLS estimator towards OLS, such as nagar1959bias, basmann1960asymptotic, richardson1968exact and sawa1969exact. However, their critique was the first to suggest some practices for the detection of possible presence of weak instruments (reporting first-stage F-statistic and the $R^2$ of the first stage regression), which have since become standard practice for using instrumental variables in empirical studies.

The literature on weak instruments have since developed considerably and continues to do so. staiger1997stock showed that conventional asymptotic approaches fail to provide good approximations for weak instruments, and formalized the problem of weak instruments by introducing `weak-instrument asymptotics' which mimics the situation of weak instruments better. stock2002survey provide a comprehensive survey of the weak instruments literature, wherein they emphasize that the problem of bias of the two-stage least-squares estimator is not solely a small sample problem. It affects estimates carried out on large sample sizes as well, as is the case with the quarter-of-birth study which has a dataset of around 300,000 observations.

Further, given the poor finite-sample properties of the two-stage least squares estimator, numerous alternate estimators have also been proposed. These include the USSIV estimator suggested by angrist1995split, Fuller-k estimator by fuller1977some, bias-adjusted 2SLS estimator by donald2001choosing, JIVE estimators by angrist1999jackknife and blomquist1999small. Another estimator, the LIML was formalized by anderson1949estimation. In our paper, we take a deeper look at two of these estimators - JIVE and LIML estimators.

We introduce the two conditions for instrumental variables estimation, the first of which will be a recurring theme in our paper, since it connects to the problem of weak instruments. Consider the standard population regression model:

align*[align* omitted — 83 chars of source]

where $\varepsilon_i$ is the error term representing omitted factors that determine $Y_i$. Variables correlated with the error term are termed endogenous variables and those uncorrelated with the error term are exogenous. A valid instrument $Z_i$ must satisfy two conditions:

itemize• Instrument Relevance: corr ($Z_i$, $X_i$) $\neq 0.$\\ It is the implications of this condition, specifically the strength (or weakness) of this correlation between the instrument and the endogenous regressor, that we are interested in. Instruments that explain little of the variation in X are called weak instruments. A precise definition is introduced in section 3.1. • Instrument Exogeneity: corr ($Z_i$, $\varepsilon_i$) = 0\\ Note that this condition cannot be statistically tested, since it involves the covariance between $Z_i$ and the unobserved $\varepsilon_i$.

If number of instruments (K) equals the number of endogenous regressors (L) we have the just or exactly identified case. If the number of instruments exceeds the number of endogenous regressors, K $\geq$ L, then we have the over-identified case.

This paper makes three contributions. First, we provide a detailed theoretical discussion on the properties of the standard two-stage least squares estimator in the presence of weak instruments, and introduce and derive two alternative estimators. Second, we conduct Monte-Carlo simulations to compare the finite-sample behavior of the different estimators, particularly in the weak-instruments case. Third, we apply the estimators to a real-world context; we employ the different estimators to calculate returns to schooling.

The rest of the paper is sectioned as follows. Chapter 2 presents a derivation of the 2SLS estimator and its limitations in the weak instruments case. Chapter 3 discusses methods to test for weak instruments. Chapter 4 presents the alternative estimators (JIVE and LIML). Chapter 5 presents the simulation study. Chapter 6 presents the application to returns to schooling. Chapter 7 concludes.

\chapter{Two-Stage Least Squares and its Limitations}

Basic Model

Throughout the paper, unless stated otherwise, we consider the following basic model with two equations, following the notation of angrist1999jackknife:

align*[align* omitted — 71 chars of source]

$Y_i$ is scalar. $X_i$ is $1\times L$ row vector of potentially endogenous regressors. $Z_i$ is $1\times K$ row vector, with $K\geq L$. $K-L$ is the number of over-identifying restrictions. In matrix notation:\\

align[align omitted — 237 chars of source]

where (2.1) denotes the structural equation, (2.2) is the first stage. $\mathbf{Y}$ and $\varepsilon$ are $N \times 1$ vectors. $\mathbf{X}$ and $\eta$ are $N \times L$ matrices, and $\mathbf{Z}$ is $N \times K$ matrix. If the vectors of regressors and that of instruments have $M$ common elements, then $M$ columns of the $N \times L$ matrix $\eta$ have their elements as zero. The following assumptions hold for our model:

enumerate• Conditional on instruments $Z_i$, the error term $\varepsilon_i$ has expectation zero ($E[\varepsilon_i|Z_i]=0$) and variance $\sigma^2$. • $E[\eta|\mathbf Z]=0$ and $E[\eta_i^\prime\eta_i|\mathbf Z]=\Sigma_\eta$, with rank $L-M$. • $E[\varepsilon_i\eta_i^\prime|\mathbf Z]=\sigma_{\varepsilon\eta}$, where $\sigma_{\varepsilon\eta}$ is an L-dimensional column vector. • All observations of ($Y_i, X_i, Z_i$) are independent and identically distributed.

Concentration Parameter

The Concentration Parameter $\mu^2$, given by:

equation[equation omitted — 71 chars of source]

measures the strength of the instruments. It is unitless. An important question to be addressed at this point is that since $\pi$, and thus $\mu^2$, is unknown, how is the researcher supposed to know whether $\mu^2$ value is low enough for the instrument under study to be weak? This is addressed by the fact that the concentration parameter can be interpreted in terms of the first stage F statistic, discussed in Section 3. If the sample size is large, $E(F) \cong \mu^2/K + 1$, and thus $F-1$ can be treated as an estimator of $\mu^2 /K$, which gives us a convenient way to test for weak instruments.

The concept of $\mu^2$ is essentially the starting point for understanding the weak instruments literature. rothenberg1984approximating showed that the concentration parameter plays the role we commonly associate with the sample size or number of observations: as $\mu^2$ becomes large, the normal distribution becomes a good approximation for the 2SLS distribution. Likewise a small concentration parameter leads to a non-normal distribution of 2SLS, and the estimator becomes biased. Thus $\mu^2$ can be interpreted as the effective sample size.

Two-Stage Least Squares Estimator

It is helpful to characterize the 2SLS and (later on the jackknife estimator as well), as a feasible version of the ideal but infeasible instrumental variables estimator. Adhering to the notation defined in section 2.1, we derive the formula for the 2SLS estimator using this characterization. The derivation is adopted, with modifications to the notation, from brucehansen2019 and angrist1999jackknife. First, we estimate the ideal estimator $\hat\beta_{OPT}$ using the optimal instrument $Z\pi$, by ordinary least-squares:

equation[equation omitted — 94 chars of source]

$\pi$ is unknown however, and so the optimal estimator $\hat\beta_{OPT}$ is infeasible.

We use an estimate of $\pi$ instead, and denote it by $\hat\pi$. $\hat\pi$ is estimated from the reduced form regression which yields $\hat\pi = (\mathbf{Z}'\mathbf{Z})^{-1}(\mathbf{Z}'\mathbf{X})$. Substituting $\hat\pi$ into (2.4),

equation[equation omitted — 112 chars of source]
equation[equation omitted — 243 chars of source]
equation[equation omitted — 173 chars of source]

This is the two-stage least squares estimator. $\hat\beta_{2SLS}$ can be thought of as an estimator with the constructed instrument $(\mathbf{Z}\hat\pi)$ where $\hat\pi = (\mathbf{Z}'\mathbf{Z})^{-1}(\mathbf{Z}'\mathbf{X})$. Representing (2.7) in terms of the projection matrix $P_{Z}=\mathbf Z(\mathbf Z'\mathbf Z)^{-1}\mathbf Z'$:

equation[equation omitted — 93 chars of source]

The two-stage least squares is by far the most widely used estimator for instrumental variables estimation, however, as we briefly discussed in chapter 1, it has certain concerning limitations, which we expand upon in the next section.

Limitations of 2SLS: Bias

In the weak instruments case, the 2SLS estimator fails to provide unbiased, reliable estimates. As shown in Section 2.3, we arrived at the 2SLS estimator using an estimate of $\pi$ since the true value of the first-stage coefficients is unknown, which implies a certain amount of over-fitting of the first-stage equation, leading to bias in the direction of the expectation of the OLS estimator of $\beta$.

There is considerable theoretical research noting the poor finite-sample properties of instrumental variables estimator. Results on the magnitude of this bias have been provided by nagar1959bias, richardson1968exact, sawa1969exact and buse1992bias. Further, as shown by noteworthy results of bound1995problems, with weak instruments particularly, the finite sample behavior worsens and the estimator can become biased as well as inconsistent.

Bias in Single Instrument case

To visualize the problem intuitively, first let us consider the simple case of a single (weak) instrument and endogenous regressor:

$$Y_i = X_i\beta + \varepsilon_i$$ $$X_i = Z_i\pi_1 + \eta_i$$

In this case, the 2SLS estimator is simply given by the ratio of the two covariances:

equation[equation omitted — 76 chars of source]

However, given the fact that our instrument is very weak, $\sigma_{X_iZ_i}=0.$ Hence, $\hat\beta_{2SLS}$ does not exist. In the simplest case, the 2SLS estimator is just the ratio of two covariances, and with weak instruments, the 2SLS or general instrumental variables estimator does not exist.

Approximate Expression of Bias in Multiple-Instruments case

In this section we provide two different expressions approximating the bias of the two-stage least squares estimator in the general case and show how it centers around ordinary least squares. Defining the `relative bias' of 2SLS to be its bias relative to the inconsistency of OLS, buse1992bias derived an expression for the approximate bias of $\hat\beta_{2SLS}$ using power series approximations. This result holds even when the errors are not normally distributed:

equation[equation omitted — 83 chars of source]

where N is the sample size and K is the number of excluded instruments. This expression is approximately inversely proportional to $\mu^2/(K-2)$, as shown:

equation[equation omitted — 173 chars of source]

where $\mu^2$ denotes the concentration parameter (measure of the strength of the instrument as discussed in Section 1.). Note that in the above equation, the term $\frac{\sigma_{\varepsilon\eta }}{\sigma_{\eta }^{2}}$ approximately equals the asymptotic bias of the OLS estimator when the instrument explains little of the variation of $\mathbf X$. Hence, from equation (2.10) it is clearly seen that for $K>2$, the bias of the 2SLS estimator relative to OLS is inversely proportional to the concentration parameter, hence, weaker the instrument(s), lower is the concentration parameter, and higher is the bias of 2SLS towards the OLS estimate.

It can be seen from the above formulation that increasing the number of instruments with the explanatory power remaining constant, causes the relative bias of the 2SLS to only increase.

Before moving on to the second expression, it is important to briefly introduce alternate asymptotic representations that are generally used for the weak instruments case. As discussed by stock2002survey, for weak instruments, conventional asymptotic approximations to finite-sample distributions are quite poor. Two alternate asymptotic methods commonly employed are weak-instrument asymptotics (involving a sequence of models chosen to keep $\mu^2/K$ constant as sample size $N \rightarrow \infty$) pioneered by staiger1997stock and many-instrument or group asymptotics (involving sequence of models with fixed instruments and normal errors, where $K$ is proportional to $N$ and $\mu^2/K$ converges to a constant finite limit), first proposed by bekker1994alternative. In the 1995 working paper version of angrist1999jackknife, it was referred to as group asymptotics.

We present an expression for the approximate bias of the 2SLS estimator using group asymptotics, wherein we let the number of instruments grow proportional to rate of the sample size. This keeps the instruments weak. The detailed derivation of this expression is provided in the appendix A.1.

equation[equation omitted — 121 chars of source]

where F is the population analog of the F-statistic for the joint significance of the instruments in the first-stage regression. $\frac{\sigma_{\varepsilon\eta }}{\sigma_{\eta }^{2}}$ approximately equals the asymptotic bias of the OLS estimator. Given a weak first-stage (weak instruments case), $F\rightarrow0$, and we can see that the bias approaches $ \frac{\sigma_{\varepsilon\eta }}{\sigma_{\eta }^{2}}$. With a strong first-stage, $F\rightarrow \infty$ and then the 2SLS bias goes to 0.

Inconsistency of the 2SLS estimator

In this section we relate the two conditions for instrumental variables estimation to the consistency property of the 2SLS estimator. If the weak correlation (between the instrument(s) and the endogenous variable) is coupled together with even a small violation of the second condition of instrumental variables estimation (which is the instrument exogeneity condition) then we have an inconsistent estimator. This insight was first discussed by bound1995problems.

To represent the problem discussed above, let us consider the probability limit of the 2SLS estimator:

equation[equation omitted — 127 chars of source]

where $\mathbf {\hat X}$ is the projection of $\mathbf X$ onto $\mathbf Z$, and ${\sigma _{\mathbf {\hat X},\varepsilon }}$ is the covariance between $\mathbf {\hat X}$ and $\varepsilon$.

From equation (2.13) we can intuitively understand that if $\sigma _{\mathbf {\hat X}}^{2}$ is small, which implies that we have a weak instrument, then as long as ${\sigma _{\mathbf {\hat X},\varepsilon }}$ is zero the estimator will be consistent. However, suppose we have a weak instrument and also, ${\sigma _{\mathbf {\hat X},\varepsilon }}$ is small but non-zero, then even that small correlation between $\mathbf Z$ (and thus, $\mathbf {\hat X}$) and the structural error term $\varepsilon$ can lead to a large inconsistency.

Hence, with weak instruments, even moderate correlation between instrument and structural error term can magnify the inconsistency of IV estimator.

\chapter{Testing for Weak Instruments} Testing for presence of weak instruments is, at the time of writing, an active field of research. For a detailed overview, see stock2002survey. For the purpose of our study, we limit our attention to two tests - the widely-used first-stage F-statistic and the Anderson-Rubin Test, which has gained resurgence in recent years in light of new developments in instrumental variables research.

Defining the `Weakness' precisely

stock2002testing posit that the definition of weak instruments depends on the inferential task to be carried out, and cannot be resolved in the abstract. One approach is to define a set of instruments to be weak if $\mu^2/K$ is small enough that inferences based on conventional normal approximating distributions are misleading. For instance, if a researcher wants their 2SLS estimate bias to be small, one measure of whether an instrument(s) is strong is whether $\mu^2/K$ is large enough such that the 2SLS relative bias (relative to the bias of ordinary least squares) is below a certain threshold, for example the relative bias is below 10%. Hence, to be deemed a `weak' instrument, the 2SLS estimate using that instrument should have relative bias above 10%. The definition we discussed (and use for our simulation) is based on relative bias, another definition (for instance on size of test) may result in a different cut-off value.

First Stage F-statistic

The first-stage F-statistic is the F-statistic testing the hypothesis that the coefficients on the instruments equal zero ($\pi=0$) in the first stage of two stage least squares. stock2002testing show that the definition of weak instruments discussed above implies a threshold value for $\mu^2/K$, under weak asymptotics. A weak instrument will have a $\mu^2/K$ value (and hence, an F-statistic value, since F$-$1 can be treated as an estimator of $\mu^2/K$ as discussed in Section 2.2) lower than the threshold. For the case of a single endogenous regressor, staiger1997stock provide a rule-of thumb threshold of 10: a value less than 10 indicates that the instruments are weak, in which case the 2SLS estimator is biased and 2SLS t-statistics and confidence intervals are unreliable.

stock2002survey provide a table listing critical values of the first-stage F-statistic such that the relative bias of 2SLS estimates is greater than 10%, for different numbers of instruments. The authors arrived at those critical values based on weak-instrument asymptotic approximations. We include a subset of this table (which is relevant for our simulations) as Table B.1 in appendix B for reference.

Anderson-Rubin Test

The AR test is a hypothesis test that has the property of being valid whether instruments are strong, weak or even irrelevant ($\pi=0$). It tests the null hypothesis $\beta$ = $\beta_0$ using the statistic. It was proposed by anderson1949estimation.

equation[equation omitted — 180 chars of source]

One definition of the LIML estimator is that it minimizes $AR(\beta)$. With fixed instruments and normal errors, the quadratic forms in the numerator and denominator of (3.1) are independent chi-squared random variables under the null hypothesis, and $AR(\beta_0)$ has an exact $F_{K,T - K}$ null distribution. Under the more general conditions of weak-instrument asymptotics, $AR(\beta_0)$ $\xrightarrow{\text{d}}$ ${\chi_k}^2/K$ under the null hypothesis, regardless of the value of $\mu^2/K$. Thus the AR statistic provides a fully robust test of the hypothesis $\beta$ = $\beta_0$. The set of values of $\beta$ that are not rejected by a 5% Anderson–Rubin test will constitute a 95% confidence set for $\beta$. The logic behind the Anderson–Rubin statistic is that it never assumes instrument relevance, and the AR confidence set will have a coverage probability of 95% in large samples, regardless of the strength or weakness of instruments. In light of the importance given to the problem of weak instruments in recent years, this test has gained traction among econometricians, who increasingly advocate for its use for robust inference with weak instruments (See staiger1997stock). Particularly, recent research has shown the AR confidence set to be optimal in the single-endogenous-regressor just-identified setting. \chapter{Alternative Estimators} To tackle the problem of finite-sample bias of IV, which as we have shown, is particularly problematic in the presence of weak instruments. In this section we present a detailed overview of two alternative estimators which in theory exhibit better finite sample properties. Both of these estimators, the Jackknife IV estimator and the LIML estimator, fall under the broader class of k-class estimators. These estimators partially robust ie less sensitive to weak instruments; they are more reliable in comparison to 2SLS estimates.

Jackknife Instrumental Variables Estimator

\textbf For an intuitive sense of the functioning of the jackknife instrumental variables estimator, consider the term `jackknife' as presented in statistical literature. The jackknife estimator of a parameter is found by leaving out each observation from a dataset, calculating the estimate and computing the average of these calculations. Given a sample of size $N$, the jackknife estimate is found by aggregating the estimates of each sub-sample of size $(N-1)$. In a similar vein, the estimator we describe in this section replaces the usual fitted values from the reduced form regression of the two-stage least squares by `omit-one' fitted values. The key feature of the JIVE estimator is that it eliminates the correlation between the fitted values and the structural equation errors, as explain in detail below. The fitted value of the standard 2SLS estimator is only asymptotically independent of the structural error ($\varepsilon_i$), but the JIVE is independent even in finite samples. The derivation we present in this section is adopted from angrist1999jackknife. To derive the expression formally, consider again the expressions for $\hat\beta_{OPT}$ and $\hat\beta_{2SLS}$ from section 2.2, equations (2.4) and (2.5). $\hat\beta_{OPT}$ employed the optimal instrument $\mathbf Z\pi$ and $\hat\beta_{2SLS}$ used an estimate of $\mathbf Z\pi$, denoted by $\mathbf Z\hat \pi$. Now we introduce the new estimator for $\beta$ by using a different estimate of the optimal instrument, $Z_{i}\tilde{\pi}$.

From the formulation of approximate bias of the 2SLS estimator we presented in section 2.4.2 (equation 2.11), it is clear that increasing the number of instruments while keeping the explanatory/predictive power of the instruments constant, leads to an increase in the bias of $\hat\beta_{2SLS}$. However, we can see that for the optimal estimator, increasing the number of instruments while keeping $Z_i\pi$ fixed will have no effect on the properties of $\hat\beta_{OPT}$. Thus, for finite samples in the presence of many instruments, $\hat\beta_{2SLS}$ fares much worse than $\hat\beta_{OPT}$. Recall the representation of the 2SLS estimator using the projection matrix $P_Z$ from Section 2 (equation 2.8). We can see that the first-stage fitted values $\mathbf Z\hat{\pi}$ can be written as:

equation[equation omitted — 75 chars of source]

The term $P_Z\eta$ in the above equation is correlated with the error term of the first-stage $\eta$ (see equation 1.2) and hence with the structural error term $\varepsilon$. Put differently, since $\hat \pi$ is estimated on the full sample which includes the $i$th observation, it is correlated with $\eta_i$, which is correlated with $\varepsilon_i$. This correlation is given by:

equation[equation omitted — 310 chars of source]

Due to this correlation between $\hat \pi$ and $\varepsilon_i$, $\hat\beta_{2SLS}$ is biased for $\beta$. While the correlation disappears asymptotically, it holds implications in finite samples.

In devising the new instrument, we attempt to keep the correlation in (4.2) equal to zero. The problem stems from $\hat \pi$ being estimated on the full sample which includes the $i$th observation. For the new estimator which has the constructed instrument $Z_{i}\tilde{\pi}$, $\tilde\pi$ is estimated not on the full sample but on the sample with the $i$th observation removed. Therefore, the new estimated instrument will be independent of $\varepsilon_i$ even in finite samples.

The $i$th row of the estimated instrument for 2SLS, $\mathbf Z\hat{\pi}$, where $\hat{\pi}=(\mathbf Z'\mathbf Z)^{-1}(\mathbf Z'\mathbf X)$, is given by

equation[equation omitted — 86 chars of source]

Remove the $i$th row from the matrices of regressors and instruments and denote them by $\mathbf X(i)$ and $\mathbf Z(i)$. Thus JIVE removes the dependence between $Z_{i}\hat{\pi}$ and the regressor $X_{i}$.

The corresponding estimate of $\pi$ and the constructed instrument will be:

equation[equation omitted — 90 chars of source]
equation[equation omitted — 101 chars of source]

$\varepsilon_{i}$ and $X_{j}$ are independent when $i\neq j$, so it follows that

equation[equation omitted — 151 chars of source]

Hence, now we have $E[\varepsilon_{i}Z_{i}\tilde{\pi}(i)]=0$, so we have removed the correlation between the fitted values and the structural error presented in equation (4.2). The JIVE estimator is equal to:

equation[equation omitted — 112 chars of source]

where $\hat{X}_{JIVE}$ is $N\times L$ dimensional matrix with the $i$ th row $Z_{i}\tilde{\pi}(i)$. \\

blomquist1999small summarize the construction of the JIVE estimator by providing the following algorithm:

algorithm[algorithm omitted — 510 chars of source]

The estimator for $Z_i\pi$, $Z_{i}\tilde{\pi}(i)$ is consistent. $\hat{\beta}_{JIVE}$ has the same probability limit and first-order asymptotic distribution as $\hat{\beta}_{OPT}$ and $\hat{\beta}_{2SLS}$, under conventional fixed model asymptotics (stock2002survey).

Limited Information Maximum-Likelihood Estimator

Limited Information Maximum Likelihood (LIML) is an alternative method to estimate the parameters of the structural equation. It was formalized by anderson1949estimation. The derivation of the LIML estimator is shown in the appendix, from equations (A.13) to (A.19).\\ We follow the same model as described in Section 2, and extend it as shown below:

equation[equation omitted — 117 chars of source]
equation[equation omitted — 99 chars of source]

where $\mathbf{Z_0} = \mathbf{X_0}$.\\ Here, $\mathbf{X_0}$ is the matrix of exogenous variables and $\mathbf{X_1}$ of endogenous variables. $\mathbf{Z_0} = \mathbf{X_0}$ is the matrix of included exogenous variables, $\mathbf{Z_1}$ is the matrix of excluded exogenous variables. Because the LIML estimator is based on the structural equation for $\mathbf{Y}$ combined with the first-stage equation for $\mathbf{X_1}$, it is called `limited information'. LIML estimator is given by:

equation[equation omitted — 146 chars of source]

The LIML estimator has some excellent properties when the number of excluded instruments (number of columns in matrix $\mathbf X_1$) in the first-stage equation and the sample size are large. Although the LIML estimator and the 2SLS estimator are asymptotically equivalent in the standard large sample theory, they are quite different in case of many instruments or many weak instruments. It has no finite moments, which implies that its density tends to have very thick tails. anderson2010asymptotic show that the LIML estimator shows asymptotic optimality with many weak instruments . The LIML estimator is asmptotically efficient in higher order, while the 2SLS estimator is inconsistent, shown by kunitomo1987third.

\chapter{Simulation Study} In this chapter, we firstly explore the finite sample behavior of different IV estimators as well as OLS estimator by Monte Carlo simulation. Then, we discuss the performance of all these estimators in 4 models, namely Just-identification with strong IV, Just-identification with weak IV, Over-identification with strong IVs, Over-identification with weak IVs. The experiment presented here follows from the simulations presented in angrist1999jackknife, davidson2006case.

Finite Sample Properties

To maintain simplicity in the design of our experiment, we let the error terms of both reduced-form and structural equation to be homoscedastic, all the regressions to be linear.\\ The structural equation is defined as follow,

equation[equation omitted — 61 chars of source]

where $\iota$ is a vector of ones and $X$ is the endogenous variable, which is a one-dimensional column vector generated by the reduced-form equation,

equation[equation omitted — 96 chars of source]

here, the first column of matrix $Z$ is also $\iota$, the remaining columns $z_i$ are IID variables with mean zero and unit variance.\\ Without loss of generality, we set $\beta_0 = \beta _1 = 1$, $\sigma_{\varepsilon}^2 = 1$, $\pi_0 = 0$, $\pi_1,..., \pi_{K-1}$ to be equal. Under the exogeneity condition, $Z$ is uncorrelated with $\varepsilon$, therefore we know that the correlation between $X$ and $\varepsilon$ is only through $\eta$. The correlation coefficient between $\varepsilon$ and $\eta$ is denoted by $\rho$ (the covariance between the two is denoted by $\sigma_{\varepsilon\eta}$). We carry out 3 experiments - we explore how variation in $\rho$, the strength of instruments and also the number of instruments impacts the finite sample behavior of the estimators under study. The limiting $R^2$ of the reduced-form equation, denoted by $R_{\infty}^2$, is given by $R_{\infty}^2 = 1/(1+ \sigma_\eta^2/\parallel \pi \parallel^2)$. We normalized $\sigma_\eta^2 + \parallel \pi \parallel^2$ to 1 so as to restrict $R_{\infty}^2 \in [0, 1]$. $R_{\infty}^2$ is a monotonically increasing function of the `Concentration Parameter', discussed in Section 2. Matrix $Z$ has $K$ columns and $K-2$ over-identifying restrictions. We vary $\rho$ and $R_{\infty}^2$ at interval of 0.01. Every experiment is performed with sample size 25, 50, 100, 200, 400, 800. All the experiments are replicated 1000 times. Since LIML and JIVE do not have first and second moments, we employ Median Bias, 0.5 quantile of the estimation minus the true value, to evaluate the central tendency of different estimators.\\ In the first experiment, we vary $\rho$, keep $K=7$, $R_{\infty}^2=0.1$ (implies $\parallel \pi \parallel^2=0.1, \sigma_\eta^2=0.9$) constant. Result is shown in Figure B.1. The median bias of OLS, 2SLS and JIVE proportionally increases to $\rho$ for all sample sizes. For LIML, as the sample size reaches 200 and above, its median bias is negligible. In the second experiment, we vary $R_{\infty}^2$ and keep $K=7$, $\rho=0.9$ fixed. Figure B.2 clearly shows that the growth of $R_{\infty}^2$ leads to a decline in the median bias for all four estimators. It is noteworthy that at sample sizes of 100 and above, LIML shows much lower median bias (compared to other estimators) even at small values of $R_{\infty}^2$. JIVE shows an interesting trend, it fluctuates around OLS, while maintaining the same overall trend as OLS - decrease in median bias as $R_{\infty}^2$ decreases. We change the value of $K$ in the third experiment and keep $R_{\infty}^2=0.1$, $\rho=0.9$ fixed. In Figure B.3, except for OLS and LIML, increase in number of instruments $K$ leads to growth of median bias. At small sample size (50 or lower), the median bias of LIML increases with $K$, however, at sizes 100 and above, the median bias appears to be stable around zero.

Performance of the Estimators

We use 4 sets of parameters corresponding to 4 models. The parameter-setting is determined through two hypothesis tests. We employ the first-stage F statistic to detect the strength of instruments and the Anderson-Rubin statistic as a fully robust test for $\beta_1$ in the two cases with weak instruments. We apply the same model in the previous section but fix sample size to be equal to 200 and replicate 5000 times. To compare estimators, median bias and coverage probability for $95\%$ confidence interval (estimated value plus or minus 1.96 times the asymptotic standard error) are employed. Coverage probability is the proportion of the time that the interval contains the true value, in our case $\beta_1 = 1$. Unlike OLS, 2SLS and LIML, the asymptotic standard error of JIVE is calculated based on $\hat{X}_{JIVE}.$

itemize• Model 1: Just-identification with strong instrument\\ Let $\beta_0=\beta_1=1$, $\pi_0=0$, $\pi_1=0.3$. Here, $K=2$ and \begin{align*} \begin{pmatrix} \varepsilon\\ \eta \end{pmatrix} &\sim N \begin{pmatrix} \begin{pmatrix} 0\\ 0 \end{pmatrix}\!\!,& \begin{pmatrix} 0.25 & 0.20\\ 0.20 & 0.25 \end{pmatrix} \end{pmatrix} \end{align*} • Model 2: Just-identification with weak instrument\\ Let $\beta_0=\beta_1=1$, $\pi_0=0$, $\pi_1=0.2$. Here, $K=2$ and \begin{align*} \begin{pmatrix} \varepsilon\\ \eta \end{pmatrix} &\sim N \begin{pmatrix} \begin{pmatrix} 0\\ 0 \end{pmatrix}\!\!,& \begin{pmatrix} 1.0 & 0.9\\ 0.9 & 1.0 \end{pmatrix} \end{pmatrix} \end{align*} • Model 3: Over-identification with strong instruments\\ Let $\beta_0=\beta_1=1$, $\pi_0=0$, $\pi_1=\pi_2=...=\pi_{15}=0.3$. Here, $K=16$ and \begin{align*} \begin{pmatrix} \varepsilon\\ \eta \end{pmatrix} &\sim N \begin{pmatrix} \begin{pmatrix} 0\\ 0 \end{pmatrix}\!\!,& \begin{pmatrix} 0.25 & 0.10\\ 0.10 & 0.25 \end{pmatrix} \end{pmatrix} \end{align*} • Model 4: Over-identification with weak instruments\\ Let $\beta_0=\beta_1=1$, $\pi_0=0$, $\pi_1=\pi_2=...=\pi_{15}=0.1$. Here, $K=16$ and \begin{align*} \begin{pmatrix} \varepsilon\\ \eta \end{pmatrix} &\sim N \begin{pmatrix} \begin{pmatrix} 0\\ 0 \end{pmatrix}\!\!,& \begin{pmatrix} 0.25 & 0.20\\ 0.20 & 0.25 \end{pmatrix} \end{pmatrix} \end{align*}

Figures B.4 to B.7 present the distribution of $\beta_0$ and $\beta_1$ in Models 1-4 respectively. In Model 1, the OLS estimator of $\beta_1$ is notably biased. In Model 2, the OLS estimator of $\beta_1$ is even more biased. Besides, the 2SLS, LIML and JIVE display a very wide range of estimates for both $\beta_0$ and $\beta_1$. Since we are more interested in the slope coefficient ($\beta_1$), we generate Figures B.8 to B.11 to look at $\beta_1$ in more detail. After we exclude some outliers and focus on the interval $[-3, 5]$, in Figure B.9 we observe notable negative skew in 2SLS, LIML and JIVE estimators. In figure B.6, corresponding to Model 3, we see that $\beta_1$ of the OLS estimator shows a remarkable reduction of bias compared to Model 2. In addition, for all estimators, the distribution of $\beta_1$ is concentrated on a small range which contains the true value. However, in Model 4, the bias of OLS estimate of increases again and the distribution of LIML and JIVE estimates disperses slightly. We include a series of quantiles around $\beta_1$ in Table B.2 to display the dispersion. The JIVE estimator returns a surprisingly large number of outliers on both tails in the just-identified with weak IV model. If we had a smaller number of replications, then in spite of the high coverage probability, we might have seen some extreme JIVE estimates, quite far from the true value.

According to the results of our MC simulations, we find that the LIML estimator performs well (in terms of median bias) in most of models, despite the fact that JIVE has the highest coverage probability in all the models. We observe that JIVE does not dominate LIML in any case/model. Between 2SLS and JIVE, we do not observe JIVE performing uniformly better than 2SLS in the models we considered. While we do observe JIVE do well in overidentified case as discussed in angrist1999jackknife, however, our findings are more in line with results of davidson2006case.

\chapter{Application to Returns to Schooling}

In this section, we revisit the famous paper and its (`provocative' as termed by bound1995problems) results that led to the beginning of the weak instruments literature: angrist1991does, who use quarter of birth as an instrument for estimating the impact of educational attainment on earnings. On the same data, we confirm the weakness of the instruments used in the paper using the first-stage F-statistic and proceed to apply estimators more robust to weak instruments.

Endogeneity of education is a well-known problem that economists face while estimating the effect of education on earnings. The reason for the endogeneity is omitted variables, such as the `ability', which can be correlated with both educational attainment and earnings of an individual. As a result, OLS will give biased estimates for the return to education. To correct this bias quarter of birth is included in the regression as an instrumental variable. Association between the quarter of birth and schooling is explained by compulsory schooling requirements in the United States. According to school start age policy, children are required to enter school in the fall of the calendar year in which they turn 6. While compulsory schooling laws allow students to leave school after they turned 16. As a result, students who were born earlier in the calendar year tend to attend school for a shorter period of time than those born at the end of the year. So interaction of the two requirements of schooling laws generates variation in educational attainment for the students who graduate right after their 16th birthday. Quarter of birth must satisfy the instrument relevance and exogeneity conditions to be a valid instrument for educational attainment. The relevance condition implies that the instrument must be correlated with the endogenous variable, in this case with the years of schooling. Higher the correlation between these two variables, stronger is the first stage in the two-stage least squares estimation. To satisfy the second condition quarter of birth must effect earnings only via its effect on schooling years. In general, it is not possible to test this condition statistically. Angrist and Krueger argue that student's birthday can not be correlated with other personal features which may affect earnings, thus the variation in education due to the individual's birthday is exogenous (angrist1991does). For our estimations, we used the dataset from angrist1991does which is taken from 1980 US Census. The sample consists of 329,509 men, who were born in 1930-1939. The dataset includes information on the quarter of birth, year of birth, state of birth, years of schooling and earnings for this sample. All figures and tables are presented in Appendix B.\\ Figure B.12 shows the relation between quarter of birth and schooling (first-stage). The graph indicates that there is an upward trend in average years of schooling for men born in 1930-1939. There is also a persistent seasonal pattern in education. Men born at the beginning of the year tend to have less schooling on average than those who were born later in the year. Figure B.13 illustrates the reduced form, which is the relationship between the quarter of birth and wages. Here we also notice a pattern where the 3rd and 4th quarter of births correspond to a higher log weekly wages. Our general model is given by:

equation[equation omitted — 102 chars of source]
equation[equation omitted — 66 chars of source]

Here $E_{i}$ is the schooling of ith individual, $Y_{ic}$ is a dummy variable, denoting if the ith individual was born in cth year, $Q_{ij}$ is a dummy variable showing quarter of birth of the ith individual and $W_{i}$ is the weakly wage. In Table B.3 we present results of the first stage estimations. The first and third columns of the table indicate that individuals born in the last quarter of the year had about 0.10 year more schooling compared to the men born in the first to third quarters. The second, fourth and fifth columns show the estimates of each quarter of birth in comparison with the first quarter. Naturally, the largest difference is between the last and the first quarter, which is around 0.15 year, independent of including year of birth and state of birth as control variables in the regression. Table B.4 attempts to replicate the main results of Table 2 from angrist1999jackknife. We compare 2SLS, LIML and JIVE estimators for three different specifications of the model based on angrist1991does data. Additionally, we report the first-stage F-statistic and $R^2$ values. Column 1 shows the results of our estimation wherein we regress wages on schooling, including three quarter of birth dummies as instruments and nine year of birth dummies as control variables. Here, our first stage F-statistic is greater than the critical value (for the case of K = 3 where K is the number of instruments) of 9.08, according to Table B.1 provided in Appendix B. Hence in the first model the excluded instruments are strong. Moving to column 2, to control for the age-related trends we include interactions of the year of birth dummies with quarter of birth dummies as instruments in the second model. This leads to a slightly lower 2SLS estimate as well as a slight decrease in standard errors. However, the value of the first-stage F-statistic reduces dramatically to 4.91. The small value of the F-statistic indicates that the excluded instruments and educational attainment are only weakly correlated, which as we discussed in section 2.4.2 can lead to finite-sample bias in the 2SLS estimate. Looking at the $R^{2}$ in columns (1) and (2), we notice that compared to the case of the first column, the explanatory power of the instruments does not increase much in the second specification. Finally, in the last specification (column 3), we increase the number of instruments to 180, taking year of birth $\times$ quarter of birth, and state of birth $\times$ quarter of birth interactions as instruments. The reasoning behind increasing the number of instruments by including interaction terms is to ensure that seasonal differences do not vary by state and birth year. While this results in about 40 percent reduction in standard errors compared to the column (2), however, the value of the first-stage F-statistic again decreases, compared to the preceding model. This indicates that while attempting to increase the precision of the 2SLS estimates, the weakness of the instruments, however, leads to the 2SLS estimates becoming biased.

In Table B.4 we show the estimates of the alternative estimators discussed previously (LIML and JIVE) which are considered to have better finite sample properties when the instruments are weak. As we can see, JIVE estimates vary considerably from the other two. Given the encouraging performance of the LIML estimator in our simulation study, and considering the closeness of the LIML and 2SLS estimates in our application to returns to schooling, we suspect that in this specific case, LIML and 2SLS probably give more reliable results than JIVE. However, we do not at all recommend using 2SLS estimator in case of weak instruments in general, considering the theoretical discussion and simulation results presented previously.

\chapter{Conclusion} In our study, we hope to have provided the reader a concise but comprehensive overview of selected aspects of the weak instruments literature.\\ When we face the issue of weak instruments, the best solution is of course to find better, stronger instruments. However, in empirical practice this is easier said than done. Hence, we think research into alternative estimators which can give more reliable estimates than the two-stage least squares estimator in the weak instruments case, is very relevant. From our small study, we find the LIML estimator to perform the best when the correlation between the instrument and the endogenous explanatory variable is low. We posit the LIML be a possible solution in the (fairly common) case that a researcher has a weak instrument(s) and is not in a position to find other (stronger) instruments. Development of methods for robust inference of weak instruments such as the Anderson-Rubin statistic is a very promising area of further research.