EconBase
← Back to paper

Weak Instruments, First-Stage Heteroskedasticity, the Robust F-Test and a GMM Estimator with the Weight Matrix Based on First-Stage Residuals

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

47,261 characters · 9 sections · 33 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Weak Instruments, First-Stage Heteroskedasticity, the Robust F-Test and a GMM Estimator with the Weight Matrix Based on First-Stage Residuals

abstract\baselineskip=13pt This paper is concerned with the findings related to the robust first-stage F-statistic in the Monte Carlo analysis of IAndrewsREStat2018, who found in a heteroskedastic grouped-data design that even for very large values of the robust F-statistic, the standard 2SLS confidence intervals had large coverage distortions. This finding appears to discredit the robust F-statistic as a test for underidentification. However, it is shown here that large values of the robust F-statistic do imply that there is first-stage information, but this may not be utilized well by the 2SLS estimator, or the standard GMM estimator. An estimator that corrects for this is a robust GMM estimator, denoted GMMf, with the robust weight matrix not based on the structural residuals, but on the first-stage residuals. For the grouped-data setting of Andrews (2018), this GMMf estimator gives the weights to the group specific estimators according to the group specific concentration parameters in the same way as 2SLS does under homoskedasticity, which is formally shown using weak instrument asymptotics. The GMMf estimator is much better behaved than the 2SLS estimator in the IAndrewsREStat2018 design, behaving well in terms of relative bias and Wald-test size distortion at more standard values of the robust F-statistic. We show that the same patterns can occur in a dynamic panel data model when the error variance is heteroskedastic over time. We further derive the conditions under which the Stock and Yogo (2005) weak instruments critical values apply to the robust F-statistic in relation to the behaviour of the GMMf estimator.

JEL Classification: C12, C36

Keywords: Instrumental Variables, Weak Instruments, Heteroskedasticity, F-Test, GMM, Grouped-Data IV, Dynamic Panel Data

\thispagestyle{empty}

\baselineskip=20pt

\pagenumbering{arabic} \setcounter{page}{1}

Introduction

It is commonplace to report the first-stage F-statistic as a test for underidentification in linear single endogenous variable models estimated by two-stage least squares (2SLS). This could either be a non-robust or robust version of the test, with robustness to for example heteroskedasticity, serial correlation and/or clustering. Under maintained assumptions, these are valid tests for the null $H_{0}:\pi=0$ in the first-stage linear specification $x=Z\pi+v$, where $x$ is the endogenous explanatory variable in the model of interest $y=x\beta+u$, and $Z$ are the instruments. If the null is not rejected, then this is an indication that the relevance condition of the instruments does not hold and that the 2SLS estimator does not provide a meaningful estimate of the parameter of interest $\beta$. A rejection of the null does, however, not necessarily imply that the 2SLS estimator is well behaved. This follows the work of StaigerStock1997 and StockYogo2005, with the latter providing critical values for the first-stage non-robust F-statistic for null hypotheses of weak instruments in terms of bias of the 2SLS estimator relative to that of the OLS estimator and Wald test size distortion. These non-robust weak instruments F-tests are valid only under conditional homoskedasticity, no serial correlation and no clustering of both the first-stage errors $v$ and the structural errors $u$, and do not apply to the robust F-test in general designs, see BunHaan2010, OleaPflueger2013 and IAndrewsREStat2018. For general designs OleaPflueger2013 proposed the effective first-stage F-statistic and critical values linked to the Nagar bias of the 2SLS estimator, whereas IAndrewsREStat2018 obtained valid two-step identification robust confidence sets.

This paper is concerned with the findings related to the robust F-statistic in the Monte Carlo analysis of IAndrewsREStat2018. In a cross sectional heteroskedastic design he found that even for very large values of the robust F-statistic, the standard 2SLS confidence intervals had large coverage distortions. For example, for a high endogeneity design, “the 2SLS confidence set has a 15% coverage distortion even when the mean of the first-stage robust F-statistic is 100,000”, IAndrewsREStat2018. This is a striking finding and appears to discredit the robust F-statistic as a test for underidentification. However, it is shown here that large values of the robust F-statistic do imply that there is first-stage information, but this may not be utilized well by the 2SLS estimator, or GMM estimators that incorporate heteroskedasticity in the structural error $u$ only.

The IAndrewsREStat2018 design is the same as a grouped data one, see AngristJoE1991 and the discussion in AngristPischke2009, where the instruments are mutually exclusive group membership indicators. Denoting the groups by $s=1,...,S$, the group specific concentration parameter values are determined by the ratios $\pi_{s}^{2}/\sigma_{v,s}^{2}$, where $\sigma_{v,s}^{2}$ is the group specific variance of the first-stage error $v$. The 2SLS estimator is a weighted average of the group specific estimators of $\beta$, giving more weight to large concentration parameter groups if $v$ is homoskedastic. However, as shown in Section (ref), this may not happen under heteroskedasticity, where 2SLS gives more weight to high variance $\sigma_{v,s}^{2}$ groups, everything else constant. In the design of IAndrewsREStat2018 we consider here, there is one informative group, leading to the large value of the robust F-statistic, but this group has a small variance $\sigma_{v,s}^{2}$, and therefore gets only a relatively small weight in the 2SLS estimator.

An estimator that correctly gives larger weights to more informative groups is a robust GMM estimator, not using the structural residuals $\widehat{u}$, but the first-stage residuals $\widehat{v}$ in the robust weight matrix. This estimator, called GMMf, is introduced in Section (ref) and gives the weights to the group-specific estimators according to the group-specific concentration parameters in the same way as 2SLS does under homoskedasticity. This is further formally shown using weak instrument asymptotics in Section (ref). Section (ref) discusses the potential problems of the standard GMM estimator that uses a robust weight matrix based on the conditional variances of the structural errors $u$. Monte Carlo results in Section (ref) show that the GMMf estimator exploits the available information well, with much better relative bias and Wald test size properties than the 2SLS estimator for values of the robust F-statistic in line with those of the non-robust F-statistic and behaviour of the 2SLS estimator in the homoskedastic case.

Section (ref) shows that similar patterns can occur when considering a simple AR(1) dynamic panel data model with heteroskedasticity of the idiosyncratic shocks in the time dimension, with estimation based on the forward orthogonal deviations transformation.

For a general setting, we report in Section (ref) the conditions under which the Stock and Yogo (2005) critical values can be applied to the robust F-statistic in relation to the behaviour of the GMMf estimator. These conditions are derived in Appendix (ref). Whilst these have limited applicability, the fully homoskedastic design is a special case.

Grouped-Data IV Model, First-Stage F-Statistic and 2SLS Weights

We consider the model as in IAndrewsREStat2018, which is the same as a grouped-data IV setup,

eqnarray*[eqnarray* omitted — 84 chars of source]

for $i=1,...,n$, where the $S$-vector $z_{i}\in\left\{ e_{1},...,e_{S}\right\} $, with $e_{s}$ an $S$-vector with $s$th entry equal to $1$ and zeros everywhere else, for $s=1,\ldots,S$. Assumptions for standard asymptotic normality results for the IV estimator hold and the variance of the limiting distribution of the parameters can be estimated consistently.

The variance-covariance structure for the errors is modeled fully flexibly by group, and specified as \[ \left(\left(

array[array omitted — 29 chars of source]

\right)|z_{i}=e_{s}\right)\sim\left(0,\Sigma_{s}\right). \]

equation[equation omitted — 155 chars of source]

At the group level, we therefore have for group member $j$ in group $s$

eqnarray[eqnarray omitted — 105 chars of source]

\[ \left(

array[array omitted — 31 chars of source]

\right)\sim\left(0,\Sigma_{s}\right) \] for $j=1,...,n_{s}$ and $s=1,...,S$, with $n_{s}$ the number of observations in group $s$, $\sum_{s=1}^{S}n_{s}=n$, see also BekkervdPloeg2005. We assume that $\lim_{n\rightarrow\infty}\frac{n_{s}}{n}=f_{s}$, with $0<f_{s}<1$.

The OLS estimator of $\pi_{s}$ is given by $\widehat{\pi}_{s}=\overline{x}_{s}=\frac{1}{n_{s}}\sum_{j=1}^{n_{s}}x_{js}$ and $Var\left(\widehat{\pi}_{s}\right)=\sigma_{v,s}^{2}/n_{s}$. The OLS residual is $\widehat{v}_{js}=x_{js}-\overline{x}_{s}$ and the estimator for the variance is given by $V\widehat{a}r\left(\widehat{\pi}_{s}\right)=\widehat{\sigma}_{v,s}^{2}/n_{s}$, where $\widehat{\sigma}_{v,s}^{2}=\frac{1}{n_{s}}\sum_{j=1}^{n_{s}}\widehat{v}_{js}^{2}$. Let $Z$ be the $n\times S$ matrix of instruments. For the vector $\pi$ the OLS estimator is given by \[ \widehat{\pi}=\left(Z^{\prime}Z\right)^{-1}Z^{\prime}x=\left(\overline{x}_{1},\overline{x}_{2},...,\overline{x}_{S}\right)^{\prime}. \] Let

eqnarray[eqnarray omitted — 182 chars of source]

where $\text{diag}\left(q_{s}\right)$ is a diagonal matrix with $s$th diagonal element $q_{s}$. Then the robust estimator of $Var\left(\widehat{\pi}\right)$ is given by

eqnarray*[eqnarray* omitted — 212 chars of source]

The non-robust variance estimator is

eqnarray*[eqnarray* omitted — 267 chars of source]

The group (or instrument) specific IV estimators for $\beta$ are given by

equation[equation omitted — 134 chars of source]

with $\overline{y}_{s}=\frac{1}{n_{s}}\sum_{j=1}^{n_{s}}y_{js}$, and the 2SLS estimator for $\beta$ is, with $P_{Z}=Z\left(Z^{\prime}Z\right)^{-1}Z^{\prime}$,

eqnarray*[eqnarray* omitted — 400 chars of source]

the standard result that $\widehat{\beta}_{2sls}$ is a linear combination of the instrument specific IV estimators, (see e.g.\ Windmeijer2019). The weights are given by

equation[equation omitted — 115 chars of source]

and hence the 2SLS estimator is here a weighted average of the group specific estimators.

For the group specific estimates, the first-stage F-statistics are equal to the Wald statistics for testing the null hypotheses $H_{0}:\pi_{s}=0$, and are given by

equation[equation omitted — 173 chars of source]

for $s=1,...,S$. For each group specific IV estimator $\widehat{\beta}_{s}$ the standard weak instruments results of StaigerStock1997 and StockYogo2005 apply. As these are just-identified models, we can relate the values of the F-statistics to Wald-test size distortions.

The robust first-stage F-statistic for testing $H_{0}:\pi=0$ is given by

eqnarray*[eqnarray* omitted — 269 chars of source]

It is therefore clear, that if $F_{r}$ is large, then at least one of the $F_{\pi_{s}}$ is large.

The non-robust F-statistic is given by

eqnarray*[eqnarray* omitted — 406 chars of source]

From ((ref)) and ((ref)) it follows that the weights for the 2SLS estimator are related to the individual F-statistics as follows \[ w_{2sls,s}=\frac{n_{s}\overline{x}_{s}^{2}}{\sum_{l=1}^{S}n_{l}\overline{x}_{l}^{2}}=\frac{\widehat{\sigma}_{v,s}^{2}F_{\pi_{s}}}{\sum_{l=1}^{S}\widehat{\sigma}_{v,l}^{2}F_{\pi_{l}}}. \] Under first-stage homoskedasticity, $\sigma_{v,s}^{2}=\sigma_{v,l}^{2}$, for $s,l=1,\ldots,S$. Then $\widehat{\sigma}_{v,s}^{2}\approx\widehat{\sigma}_{v,l}^{2}$ for all $s,l$, and hence $F\approx\frac{1}{S}\sum_{s=1}^{S}F_{\pi_{s}}$. Then the weights are given by $w_{2sls,s}\approx\frac{F_{\pi_{s}}}{\sum_{l=1}^{S}F_{\pi_{l}}}\approx\frac{F_{\pi_{s}}}{SF}$, so we see that then the groups with the larger individual F-statistics get the larger weights in the 2SLS estimator under homoskedasticity.

This is not necessarily the case under heteroskedasticity. For two groups with equal value of the F-statistic, the group with the larger variance gets the larger weight, and indeed, a large variance weakly identified group could dominate the 2SLS estimator. As shown in the Monte Carlo exercises below, this is exactly what happens in the design of IAndrewsREStat2018. The robust F-statistic is large because one of the groups has a large value of the individual F-statistic. However, this group has a very small variance $\sigma_{v,s}^{2}$ and hence gets a small weight in the 2SLS estimator, resulting in a poor performance of the estimator in terms of (relative) bias and Wald-test size.

An Alternative GMM Estimator

Clearly, one would like to use an estimator that gives larger weights to more strongly identified groups, independent of the value of $\sigma_{v,s}^{2}$, mimicking the weights of the 2SLS estimator under homoskedasticity of the first-stage errors. This is achieved by the following GMM estimator, denoted GMMf, with the extension f for first stage,

eqnarray*[eqnarray* omitted — 338 chars of source]

with $\widehat{\Omega}_{v}=$ $\sum_{i=1}^{n}\widehat{v}_{i}^{2}z_{i}z_{i}^{\prime}$ as defined in ((ref)). This looks like the usual GMM estimator, but instead of the structural residuals $\widehat{u}$, the first-stage residuals $\widehat{v}$ are used in the weight matrix. It clearly links directly to the robust F-statistic, as the denominator is equal to $SF_{r}$.

It follows that

eqnarray[eqnarray omitted — 447 chars of source]

with \[ w_{gmmf,s}=\frac{F_{\pi_{s}}}{\sum_{l=1}^{S}F_{\pi_{l}}}=\frac{F_{\pi_{s}}}{SF_{r}}, \] and hence the groups with the larger F-statistics get the larger weights, independent of the values of $\sigma_{v,s}^{2}$, mimicking the 2SLS weights under homoskedasticity of the first-stage errors.

Weak Instrument Asymptotics

We can formalize the results obtained above further using weak instruments asymptotics (WIA). For each group $s=1,...,S$ define \[ \pi_{s}=\frac{c_{s}}{\sqrt{n_{s}}}. \] The limit for $n_{s}\rightarrow\infty$, $s=1,\ldots,S$, of the group specific concentration parameters are then given by

equation[equation omitted — 75 chars of source]

Then \[ \widehat{\pi}_{s}=\overline{x}_{s}=\frac{1}{n_{s}}\sum_{j=1}^{n_{s}}\left(\frac{c_{s}}{\sqrt{n_{s}}}+v_{js}\right)=\frac{c_{s}}{\sqrt{n_{s}}}+\overline{v}_{s}, \] and

eqnarray*[eqnarray* omitted — 367 chars of source]

where $\mu_{s}=c_{s}/\sigma_{v,s}$ and $\mathcal{T}_{s}\sim N\left(0,1\right)$. We get the standard WIA result that \[ F_{\pi_{s}}=\frac{n_{s}\overline{x}_{s}^{2}}{\widehat{\sigma}_{v,s}^{2}}\overset{d}{\rightarrow}\left(\mu_{s}+\mathcal{T}_{s}\right)^{2}\sim\chi_{1,\mu_{s}^{2}}^{2}, \] where $\chi_{1,\mu_{s}^{2}}^{2}$ is the non-central chi-squared distribution with 1 degree of freedom and non-centrality parameter $\mu_{s}^{2}$.

From ((ref)) it then follows that \[ w_{2sls,s}\overset{d}{\rightarrow}\frac{\sigma_{v,s}^{2}\left(\mu_{s}+\mathcal{T}_{s}\right)^{2}}{\sum_{l=1}^{S}\sigma_{v,l}^{2}\left(\mu_{l}+\mathcal{T}_{l}\right)^{2}}, \] with the $\mathcal{T}_{l}$ independent $N\left(0,1\right)$ variables, for $l=1,...,S$.

For the weights $w_{gmmf,s}$, \[ w_{gmmf,s}\overset{d}{\rightarrow}\frac{\left(\mu_{s}+\mathcal{T}_{s}\right)^{2}}{\sum_{l=1}^{S}\left(\mu_{l}+\mathcal{T}_{l}\right)^{2}}. \]

Consider for illustration the case where there are two groups. Table (ref) presents some results for the average values of $w_{2sls,1}$ and $w_{gmmf,1}$ after randomly drawing 100,000 values of $\mathcal{T}_{1}$ and $\mathcal{T}_{2}$. In the first row, there is homoskedasticity, $\sigma_{v,1}^{2}=\sigma_{v,2}^{2}=5$, and both groups have equal concentration parameters, $\mu_{1}^{2}=\mu_{2}^{2}=5.76$, which is the value of the concentration parameter for the group-specific Wald tests to have a maximal rejection frequency of 10% at the 5% level. Then $E\left(w_{2sls,1}\right)=E\left(w_{gmmf,1}\right)=0.5$ and both estimators will give on average equal weight to the group specific estimators.

table[table omitted — 574 chars of source]

The second row considers the case where there is a large difference in the variances, $\sigma_{v,1}^{2}=5$, and $\sigma_{v,2}^{2}=0.1$, but $\mu_{1}^{2}=\mu_{2}^{2}=5.76$ as before. We find for this case that $E\left(w_{2sls,1}\right)=0.95$, i.e.\ almost all weight will on average be given to the high variance group 1. The expected weight for the GMMf estimator is in this case not affected by the relative values of the $\sigma_{v,s}^{2}$ and remains at $E\left(w_{gmmf,1}\right)=0.5$. If we subsequently reduce the value of $c_{1}$ such that $\mu_{1}^{2}=1.96$, then $E\left(w_{2sls,1}\right)=0.84$, i.e.\ the 2SLS estimator will give more weight to $\widehat{\beta}_{1}$, the estimator in the group with the smaller concentration parameter, but larger variance. In contrast, $E\left(w_{gmmf,1}\right)=0.32$ for this case, giving less weight to the less informative group.

Variance of $u$

So far, focus has been on first-stage heteroskedasticity, with the robust GMMf estimator exploiting the first-stage information by assigning larger weights to the groups with larger group specific concentration parameters independent of the values of $\sigma_{v,s}^{2}$. Consider next the infeasible robust GMM group IV estimator, given by

eqnarray*[eqnarray* omitted — 385 chars of source]

Whereas $\widehat{\beta}_{gmm}$ is the best, normal, consistent and efficient estimator under standard asymptotics, from the analysis above it is clear that the weights may not be optimal under WIA. We have under WIA that \[ \frac{n_{s}\overline{x}_{s}^{2}}{\sigma_{u,s}^{2}}\overset{d}{\rightarrow}\frac{\sigma_{v,s}^{2}}{\sigma_{u,s}^{2}}\left(\mu_{s}+\mathcal{T}_{s}\right)^{2} \] and so \[ w_{gmm,s}\overset{d}{\rightarrow}\frac{\frac{\sigma_{v,s}^{2}}{\sigma_{u,s}^{2}}\left(\mu_{s}+\mathcal{T}_{s}\right)^{2}}{\sum_{l=1}^{S}\frac{\sigma_{v,l}^{2}}{\sigma_{u,l}^{2}}\left(\mu_{s}+\mathcal{T}_{s}\right)^{2}}. \] Clearly, if $u$ is homoskedastic, $\sigma_{u,s}^{2}=\sigma_{u,l}^{2}$ for $\forall$$s,l$, then the infeasible GMM estimator has the same WIA limiting distribution as the 2SLS estimator and suffers from the same problems as described above for 2SLS. If $\sigma_{u,s}^{2}=\kappa\sigma_{v,s}^{2}$ for all $s$ then $\widehat{\beta}_{gmm}$ behaves like the GMMf estimator, the latter in that case also the efficient estimator under standard asymptotics. For other cases the behaviour of $\widehat{\beta}_{gmm}$ depends on whether $\sigma_{v,s}^{2}/\sigma_{u,s}^{2}$ assigns relatively larger or smaller weights to the more informative groups.

An alternative is to weight by $\sigma_{u,s}^{2}\sigma_{v,s}^{2}$, such that

eqnarray*[eqnarray* omitted — 458 chars of source]

The resulting weights are then as for the standard GMM estimator under first-stage homoskedasticity. This would clearly improve efficiency if $\sigma_{u,s}^{2}$ is relatively small for the more informative groups, but can assign again less weight to more informative groups if their values of $\sigma_{u,s}^{2}$ are relatively large.

Some Monte Carlo Results

We consider here the heteroskedastic design of IAndrewsREStat2018 with $S=10$ groups, $\beta=0$ and moderate endogeneity. Table 9 in the Supplementary Appendix C.3 of IAndrewsREStat2018 presents the values of the conditional group specific variance matrices $\Sigma_{s}$ as defined in ((ref)) and the first-stage parameters, denoted $\pi_{0s}$, for $s=1,\ldots,10$. Results for the high endogeneity case are given in Appendix (ref). We multiply the first-stage parameters $\pi_{0}$ by $0.04$, such that the value of the robust $F_{r}$ is just over $80$ on average for $10,000$ replications and sample size $n=10,000$. The group sizes are equal in expectation with $P\left(z_{i}=e_{s}\right)=0.1$ for all $s$. The first two rows of Table (ref) present the values of $\pi_{s}$ and $\sigma_{v,s}^{2}$ for $s=1,\ldots,10$.

Table (ref) presents the estimation results. The non-robust F-statistic is small, $F=1.41$ and the effective F-statistic of OleaPflueger2013, denoted $F_{eff}$, is equal to the non-robust F in this grouped-data IV design. Although the robust F-statistic is large, $F_{r}=80.23$, the 2SLS estimator $\widehat{\beta}_{2sls}$ is poorly behaved. Its relative bias equal to $0.699$ and the Wald test rejection frequency for $H_{0}:\beta=0$ is equal to $0.534$ at the 5% level. In contrast, the GMMf estimator is unbiased and its Wald-test rejection frequency equal to $0.049$ at the 5% level.

table[table omitted — 965 chars of source]

The details as given in Table (ref) below make clear what is happening. It reports the fixed values of $\pi_{s}$, $\sigma_{v,s}^{2}$, $\mu_{n,s}^{2}=1000\pi_{s}^{2}/\sigma_{v,s}^{2}$ and the mean values of $F_{\pi_{s}}$, $w_{2sls,s}$ and $w_{gmmf,s}=F_{\pi_{s}}/\sum_{l=1}^{S}F_{\pi_{l}}$. Identification in the first group is strong, with an average value of $F_{\pi_{1}}=789.5$. Identification in all other 9 groups is very weak, with the largest average value for $F_{\pi_{5}}=2.23$. But the variance in group 1 is very small, and some of the variances in the other groups are quite large. This leads to the low average value of $w_{2sls,1}=0.127$, showing that the 2SLS estimator does not utilize the identification strength of the first group, with larger weight given to higher variance, but lower concentration-parameter groups.

table[table omitted — 2,502 chars of source]

Table (ref) further shows that for the GMMf estimator almost all weight is given to the first group, with the average of $w_{gmmf,1}$ equal to $0.984$, resulting in the good behaviour of the GMMf estimator in terms of bias and Wald test size. In this case the standard deviation of the GMMf estimator is quite large relative to that of the 2SLS estimator. This is driven by the value of $\sigma_{u,1}^{2}$, which in this design is equal to $1.10$, much larger than $\sigma_{v,1}^{2}$. Reducing the value of $\sigma_{u,1}^{2}$ (and the value for $\sigma_{uv,1}$ accordingly to keep the same correlation structure within group 1), will reduce the standard deviation of the GMMf estimator.

Figure (ref) displays the rejection frequencies of the robust Wald tests for testing $H_{0}:\beta=0$ for varying values of the robust F-statistic $F_{r}$ for the 2SLS and GMMf estimators. Different values of $F_{r}$ are obtained by different values of $d$ when setting the first-stage parameters $\pi=d\pi_{0}$. It is clear that the Wald test based on the GMMf estimator is much better behaved in terms of size than the test based on the 2SLS estimator, with hardly any size distortion for mean values of $F_{r}$ larger than 5. The right panel of Figure (ref) shows that the bias of the GMMf estimator, relative to that of the OLS estimator, is also substantially smaller than that of the 2SLS estimator, with the relative bias smaller than $0.10$ for mean values of $F_{r}$ larger than 9.

figure[figure omitted — 259 chars of source]

Dynamic Panel Data Model

Next, consider the dynamic AR(1) panel data specification

equation[equation omitted — 70 chars of source]

for $i=1,...,n$, and $t=2,..,T$, and for $\left|\gamma\right|<1$. Let $y_{it}^{\ast}$ and $y_{i,t-1}^{\ast}$ be the forward orthogonal deviations transformed variables, see ArellanoBover1995, and $y_{i}^{\ast}$ and $y_{i,-1}^{\ast}$ the associated $\left(T-2\right)$ vectors. Let the $\left(T-2\right)\times\left(T-1\right)\left(T-2\right)/2$ matrix of instruments $Z_{i}$ be defined as \[ Z_{i}=\left[

array[array omitted — 172 chars of source]

\right]. \] Further, let the $n\left(T-2\right)$ vectors $y^{\ast}=\left(y_{1}^{\ast\prime},y_{2}^{\ast\prime},...,y_{n}^{\ast\prime}\right)^{\prime}$ and $y_{-1}^{\ast}=\left(y_{1,-1}^{\ast\prime},y_{2,-1}^{\ast\prime},...,y_{n,-1}^{\ast\prime}\right)^{\prime}$, and the $n\left(T-2\right)\times\left(T-1\right)\left(T-2\right)/2$ instrument matrix $Z=\left[Z_{1}^{\prime},Z_{2}^{\prime},...,Z_{n}^{\prime}\right]^{\prime}$. The 2SLS estimator is then given by \[ \widehat{\gamma}_{2sls}=\left(y_{-1}^{\ast\prime}Z\left(Z^{\prime}Z\right)^{-1}Z^{\prime}y_{-1}^{\ast}\right)^{-1}y_{-1}^{\ast\prime}Z\left(Z^{\prime}Z\right)^{-1}Z^{\prime}y^{\ast} \] and is consistent and asymptotic normal under standard asymptotics, regularity assumptions and the assumption of no serial correlation in $u_{it}$, $E\left(u_{it}u_{is}\right)=0$ for $t\neq s$. It is efficient under conditional homoskedasticity, $E\left(u_{i}u_{i}^{\prime}|Z_{i}\right)=\sigma_{u}^{2}I_{T-1}$.

The first stage for the 2SLS estimator is here given by \[ y_{i,-1}^{\ast}=Z_{i}^{\prime}\pi+v_{i}^{\ast}. \] The difference here compared to the grouped data IV example is that the $v_{it}^{\ast}$ are not drawn separately, but the processes are driven by the $u_{it}$ only. Let $\widehat{\pi}$ be the OLS estimator of $\pi$, $\widehat{v}_{i}^{\ast}=y_{i,-1}^{\ast}-Z_{i}^{\prime}\widehat{\pi}$ and \[ \widehat{\Omega}_{v}^{\ast}=\sum_{i=1}^{n}Z_{i}^{\prime}\widehat{v}_{i}^{\ast}\widehat{v}_{i}^{\ast\prime}Z_{i}. \] Then the robust F-statistic for $H_{0}:\pi=0$ is given by \[ F_{r}=y_{-1}^{\ast\prime}Z\left(\widehat{\Omega}_{v}^{\ast}\right)^{-1}Z^{\prime}y_{-1}^{\ast}/k_{z}, \] where $k_{z}=\left(T-1\right)\left(T-2\right)/2$. Accordingly, the GMMf estimator is here given by \[ \widehat{\gamma}_{gmmf}=\left(y_{-1}^{\ast\prime}Z\left(\widehat{\Omega}_{v}^{\ast}\right)^{-1}Z^{\prime}y_{-1}^{\ast}\right)^{-1}y_{-1}^{\ast\prime}Z\left(\widehat{\Omega}_{v}^{\ast}\right)^{-1}Z^{\prime}y^{\ast}. \]

Under conditional heteroskedasticity the standard two-step GMM estimator is efficient under standard asymptotics. Denote $\widehat{u}_{i}^{\ast}=y_{i}^{\ast}-\widehat{\gamma}_{2sls}y_{i,-1}^{\ast}$, and let \[ \widehat{\Omega}_{u}^{\ast}=\sum_{i=1}^{n}Z_{i}^{\prime}\widehat{u}_{i}^{\ast}\widehat{u}_{i}^{\ast\prime}Z_{i}. \] The two-step GMM estimator is then given by \[ \widehat{\gamma}_{gmm2}=\left(y_{-1}^{\ast\prime}Z\left(\widehat{\Omega}_{u}^{\ast}\right)^{-1}Z^{\prime}y_{-1}^{\ast}\right)^{-1}y_{-1}^{\ast\prime}Z\left(\widehat{\Omega}_{u}^{\ast}\right)^{-1}Z^{\prime}y^{\ast}. \]

Following Arellano2003, the 2SLS estimator can also be obtained as a weighted average of cross-sectional 2SLS estimators.

eqnarray*[eqnarray* omitted — 335 chars of source]

where

eqnarray*[eqnarray* omitted — 486 chars of source]

with the $n$-vectors $y_{t}^{\ast}=\left(y_{1t}^{\ast},...,y_{nt}^{\ast}\right)^{\prime}$, $y_{t,-1}^{\ast}=$ $\left(y_{1,t-1}^{\ast},...,y_{n,t-1}^{\ast}\right)^{\prime}$, and the $n\times\left(t-1\right)$ matrix $Z_{t}=\left[

array[array omitted — 41 chars of source]

\right]$, with $y_{t}=\left(y_{1t},...,y_{nt}\right)^{\prime}$. We therefore see that we are in a similar setup as the grouped-data IV example, with the group-specific IV estimators here the cross-section specific ones. The 2SLS estimator may therefore again give too much weight to less informative groups if the associated cross-sectional variance of the first-stage error in the forward orthogonal deviations transformed model is large.

We illustrate this with the following design. We set $n=200$, $T=5$, $\gamma=0.9$, and draw $\eta_{i}\sim N\left(0,1\right)$ and $y_{i0}\sim N\left(\frac{\eta_{i}}{1-\gamma},\frac{1}{1-\gamma^{2}}\right)$. The data are then generated according to model ((ref)) for $t=1,..,5$. The $u_{it}$ are independently drawn, $u_{it}\sim N\left(0,\sigma_{u,t}^{2}\right)$, and so are iid at the cross sectional level. Table 4 presents the estimation results for the design with $\sigma_{u,t}=1$ for $t=1,2,4,5$, whereas $\sigma_{u,3}=4$.

The OLS estimator of $\gamma$ in the transformed model $y^{\ast}=\gamma y_{-1}^{\ast}+u^{\ast}$ is denoted $\widehat{\gamma}_{ols}$. Whereas the GMMf estimator takes fully account of the clustering of the first-stage errors, an alternative estimator denoted $\widehat{\gamma}_{\widehat{\sigma}_{v}^{2}}$ takes account of the period specific variances only and is defined as \[ \widehat{\gamma}_{\widehat{\sigma}_{v}^{2}}=\sum_{t=2}^{T-1}w_{\widehat{\sigma}_{v}^{2},t}\widehat{\gamma}_{t}, \] with \[ w_{\widehat{\sigma}_{v}^{2},t}=\frac{\left(y_{t,-1}^{\ast\prime}Z_{t}\left(Z_{t}^{\prime}Z_{t}\right)^{-1}Z_{t}^{\prime}y_{t,-1}^{\ast}\right)/\widehat{\sigma}_{v,t}^{2}}{\sum_{l=2}^{T-1}y_{l,-1}^{\ast\prime}Z_{l}\left(Z_{l}^{\prime}Z_{l}\right)^{-1}Z_{l}^{\prime}y_{l,-1}^{\ast}/\widehat{\sigma}_{v,l}^{2}}, \] where $\widehat{\sigma}_{v,t}^{2}=\widehat{v}_{t}^{\prime}\widehat{v}_{t}/n$, with $\widehat{v}_{t}=y_{t,-1}^{\ast}-Z_{t}\widehat{\pi}_{t}$.

Estimation results for this design are presented in Table (ref). Due to the time heteroskedasticity, the robust F-statistic has a larger mean than the non-robust F-statistic, 6.74 vs 1.44. The OLS estimator is severely downward biased. The 2SLS estimator also has a large downward bias and its relative bias is 0.70. The two-step GMM estimator is less biased, but still has a large relative bias of 0.45. The GMMf and $\widehat{\gamma}_{\widehat{\sigma}_{v}^{2}}$ estimators perform better in terms of bias, and have relative biases of 0.25 and 0.16 respectively.

table[table omitted — 733 chars of source]

Table (ref) displays the time specific information and paints a similar picture as that given in Table (ref) for the grouped-data IV case. The F-statistics for $t=2$ and $t=3$ are small and the first-stage variances are relatively large. The average F-statistic for $t=4$ is relatively large, but the first-stage variance is small. The 2SLS estimator gives a relatively large weight to the uninformative periods, whereas the first-stage variance weighted estimator gives a large weight to the informative $t=4$ period.

table[table omitted — 628 chars of source]

Figure (ref) displays the relative bias of the estimators for varying values of the robust F-statistic $F_{r}$ by varying the value of $\sigma_{u,3}=\left\{ 1,1.3,1.6...,6.1\right\} $. We see again that the GMMf and $\widehat{\gamma}_{\widehat{\sigma}_{v}^{2}}$ (GMMf$_{diag}$) estimators utilize the information as conveyed by the robust F-statistic well. The 2SLS estimator has a large relative bias for all values of the mean of $F_{r}$. Whereas the two-step GMM estimator's performance improves with increasing value of $F_{r}$ in this setting, its relative bias remains high at 0.385 at the largest mean value of $F_{r}$ considered, $14.38$. The relative biases of the GMMf and $\widehat{\gamma}_{\widehat{\sigma}_{v}^{2}}$ estimators are respectively 0.119 and 0.069 at that value of the mean of $F_{r}$.

figure[figure omitted — 144 chars of source]

Testing for Weak Instruments

Using the GMMf estimator as a generalization of the 2SLS estimator to deal with general forms of first-stage heteroskedasticity, we derive in the Appendix under what conditions the weak-instruments StockYogo2005 critical values derived for the non-robust F-test and the properties of the 2SLS estimator under full homoskedasticity apply to the robust F-test and the properties of the GMMf estimator. We focus here on standard cross-sectional heteroskedasticity, but results apply to cluster and/or serially correlated designs.

Consider again the standard linear model

eqnarray*[eqnarray* omitted — 85 chars of source]

where $z_{i}$ is a $k_{z}$-vector of instruments, and where other exogenous variables, including the constant have been partialled out. General conditional heteroskedasticity is specified as

eqnarray*[eqnarray* omitted — 218 chars of source]

Further, let \[ \Omega_{u}=E\left[\sigma_{u}^{2}\left(z_{i}\right)z_{i}z_{i}^{\prime}\right];\ \ \Omega_{v}=E\left[\sigma_{v}^{2}\left(z_{i}\right)z_{i}z_{i}^{\prime}\right];\ \ \Omega_{uv}=E\left[\sigma_{uv}\left(z_{i}\right)z_{i}z_{i}^{\prime}\right], \] and the unconditional variances and covariance \[ \sigma_{u}^{2}=E_{z}\left[\sigma_{u}^{2}\left(z_{i}\right)\right];\ \ \sigma_{v}^{2}=E_{z}\left[\sigma_{v}^{2}\left(z_{i}\right)\right];\ \ \sigma_{uv}=E_{z}\left[\sigma_{uv}\left(z_{i}\right)\right]. \]

The robust F-statistic and GMMf estimator are given by

eqnarray*[eqnarray* omitted — 226 chars of source]

StockYogo2005 derived critical values for the non-robust F-statistic under homoskedasticity for the weak-instruments hypothesis on the relative bias of the 2SLS estimator, relative to that of the OLS estimator. In Appendix (ref) we show that these critical values apply to the robust F-statistic for relative bias of the GMMf estimator, relative to that of the OLS estimator if $\Omega_{uv}=\delta\Omega_{v}$ and $\sigma_{uv}=\delta\sigma_{v}^{2}$, where $\delta$ is some arbitrary constant.

For the Wald test size distortion, we show in Appendix (ref) that the StockYogo2005 critical values apply to the GMMf based Wald test if $\Omega_{uv}=\delta\Omega_{v}$ and $\Omega_{u}=\kappa\Omega_{v}$, with $\delta$ and $\kappa$ some arbitrary constants. The condition $\Omega_{u}=\kappa\Omega_{v}$ implies that the GMMf estimator is also the efficient estimator under standard asymptotics.

Whilst these conditions imply a limited applicability of the StockYogo2005 critical values for the robust F-statistic in relation to the behaviour of the GMMf estimator, it is a generalization of, and includes, the homoskedastic case. It also encompasses the illustrative example of OleaPflueger2013, where they considered a design with $E\left[\left(u_{i}~v_{i}\right)^{\prime}\left(u_{i}~v_{i}\right)\right]=\Sigma$ and $E\left[\left(\left(u_{i}~v_{i}\right)^{\prime}\left(u_{i}~v_{i}\right)\right)\otimes z_{i}z_{i}^{\prime}\right]=a^{2}\Sigma\otimes I_{k_{z}}$, and where the non-robust F-statistic gives an overestimate of the information content for the 2SLS estimator when $a>1$.

Conclusions

This paper has shown why large values of the first-stage robust F-statistic may not translate in good behaviour of the 2SLS estimator. In the heteroskedastic grouped-data design of IAndrewsREStat2018, this is the case because a highly informative group had a relatively small first-stage variance, and the 2SLS estimator gives more weight to groups with small concentration parameters but large first-stage variances. A robust GMM estimator, called GMMf, with the robust weight matrix estimated using the first-stage residuals, remedies this problem and gives larger weights to more informative groups. This is independent of the values of the first-stage variances and is a generalization of the 2SLS estimator in that it mimics what the 2SLS estimator does under first-stage homoskedasticity. A large value of the robust F-statistic indicates that there is first-stage information resulting in a well behaved GMMf estimator, also confirmed in an AR(1) dynamic panel data model. We have provided the conditions under which the StockYogo2005 weak instruments critical values developed for the non-robust F-statistic and relative bias and Wald test size distortion of the 2SLS estimator apply to the robust F-statistic and the behaviour of the GMMf estimator.

\global\long\global\long\setcounter{table}{0}\setcounter{figure}{0}\setcounter{equation}{0}

\global\long\global\long