EconBase
← Back to paper

Robust Conditional Wald Inference for Over-Identified IV

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

41,486 characters · 8 sections · 48 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

\defcitealias{lmmp22}{LMMP (2022)} \defcitealias{StockYogo05}{SY (2005)}

\doparttoc \faketableofcontents

\pagestyle{plain}

\thispagestyle{empty}

\setcounter{footnote}{0}

titlepage\newgeometry{margin=1.2in} \begin{center} \\ \\ {Robust Conditional Wald Inference for Over-Identified IV}\let\thefootnote\relax\footnotetext{* Lee: Princeton University and NBER (email: [email removed]); McCrary: Columbia University and NBER (email: [email removed]); Moreira: FGV EPGE (email: [email removed]); Porter: University of Wisconsin (email: [email removed]); Yap: Princeton University (email: [email removed]). }\\ \end{center} \begin{center} \ \\ \end{center} \singlespacing \begin{center} { By David S. Lee, Justin McCrary, Marcelo J. Moreira, \\ Jack Porter, Luther Yap*} \end{center} \begin{center} {November 2023} \end{center} \begin{abstract} \begin{singlespace} For the over-identified linear instrumental variables model, researchers commonly report the 2SLS estimate along with the robust standard error and seek to conduct inference with these quantities. If errors are homoskedastic, one can control the degree of inferential distortion using the first-stage $F$ critical values from StockYogo05, or use the robust-to-weak instruments Conditional Wald critical values of Moreira03. If errors are non-homoskedastic, these methods do not apply. We derive the generalization of Conditional Wald critical values that is robust to non-homoskedastic errors (e.g., heteroskedasticity or clustered variance structures), which can also be applied to nonlinear weakly-identified models (e.g. weakly-identified GMM). Keywords: Instrumental Variables, Weak Instruments, $t$-ratio, First-stage $F$-statistic, Conditional Wald \end{singlespace} \end{abstract} \restoregeometry

\setcounter{footnote}{0}

\onehalfspacing

Introduction

This paper considers inference in the over-identified linear instrumental variables model and its generalization to weakly-identified models and GMM more generally. The core problem of inference in the weak IV literature is that when instruments are weak, conventional asymptotic approximations are poor, causing standard inference procedures (like Wald or $t$-ratio-based inference) to over-reject, even under the null. Indeed, Dufour97 pointed out that any confidence set that is bounded with probability 1 (like the usual $\hat{\beta}\pm1.96\cdot\hat{se}(\hat{\beta})$) could have the potential to cover the true parameter 0 percent of the time (i.e., zero percent confidence level).

Moreira03 provided a generalized approach to constructing inference procedures that addressed this weak instrument problem via data-dependent critical values; this approach was first demonstrated in the case when errors are homoskedastic. Moreira03 presented, for the hypothesis that the parameter is equal to a particular value, conditional versions of the "trinity" of test procedures: Likelihood Ratio (LR), Lagrange Multiplier (LM), and Wald tests. The critical values for each of these test statistics were functions of the observed data and the null hypothesis.

Since then, a number of efforts have generalized these tests to accommodate non-homoskedastic settings, given the widespread preference of applied researchers to remain somewhat agnostic about the properties of the errors in the linear model.\footnote{ See, for example, AndrewsMoreiraStock04, Kleibergen05, Andrews16.} Curiously, while there have been efforts to generalize the Conditional LR (CLR) and LM tests to general non-homoskedastic errors, to the best of our knowledge, the extension of the Conditional Wald in such a way has been neglected.

In this paper, we derive the extension of Conditional Wald to non-homoskedastic settings. This effort delivers a Wald-based inference procedure that is valid, similar, robust to arbitrarily weak instruments, robust to HAC error structures, and applicable to more general weakly-identified settings like GMM.

There are a number of practical reasons to revisit a testing procedure rooted in a Wald approach. First, for the linear instrumental variables model, applied research has revealed a preference for Wald-based inference. Most typically, researchers compute and report the 2SLS estimator and robust standard errors, regardless of concerns about instrument weakness. In homoskedastic settings, researchers could pursue two different options for conducting Wald-based inference using the computed $t$-ratio. Researchers can either use the critical value function for Conditional Wald in Moreira03, or they can use the first-stage $F$-statistic along with the critical value tables in StockYogo05 and the Bonferroni arguments used in \hbox{StaigerStock97}. When the errors are non-homoskedastic -- as is typically allowed in modern empirical work -- the values in StockYogo05 tables no longer reliably control size distortions, as pointed out in AndrewsStockSun19. The contribution of the current paper is to provide a method for computing critical values for the $t$-ratio that will deliver valid, robust inference, even in non-homoskedastic settings.\footnote{In this paper, acceptance/rejection of the null is the result of comparing a single statistic with a valid (in this case, data-dependent) critical value. A different, "two-step" inference approach where two different procedures (one "robust to weak instruments" and the other non-robust) are combined to form an overall valid procedure (which can also accommodate non-homoskedastic settings) is proposed by Andrews18. }

A second reason to consider a robust-to-HAC version of Conditional Wald of Moreira03 is that there are two recent studies pointing to power advantages of Conditional Wald in the homoskedastic, over-identified setting and in the non-homoskedastic just-identified setting. vandesijpewindmeijer2023 analyze power for the over-identified, homoskedastic case, and provide simulation evidence that Conditional Wald using 2SLS tends to produce shorter confidence set lengths, compared to CLR ( Moreira03). This is a particularly striking finding, in light of papers that point to the near-optimality, in terms of power, of CLR. Furthermore, lmmpy23, in the context of the just-identified (robust to HAC errors) IV model, show that two different Wald-based -- $VtF$ and Conditional Wald -- confidence intervals appear to be almost always shorter than that of AndersonRubin49, a recommended benchmark in the literature. Thus, developing the Conditional Wald robust to HAC errors is not only already aligned with practitioner practice, but these recent studies suggest that it may even have power advantages in the form of shorter confidence intervals.

Our motivation for deriving CW critical values is entirely practical and stems from taking as given practitioners' apparent preference for computing the 2SLS point estimate and robust standard error (presuming non-homoskedasticity), and finding critical values that lead to valid inference. Our approach is thus different from identifying the optimal test after having defined a class of procedures and an objective function. Nevertheless, the two studies mentioned above do suggest the possibility that in terms of power and confidence interval length, CW could fare well compared to existing alternatives for the over-identified model.

Section (ref) establishes the notation we use for the standard linear IV model with non-homoskedastic errors, Section (ref) derives the critical values for Robust Conditional Wald tests based on 2SLS, LIML, two-step, and CUE GMM estimators, Section (ref) extends the test to nonlinear weakly-identified models (e.g. GMM), and Section (ref) concludes.

The Linear IV Model

The standard linear IV model is represented by

eqnarray*[eqnarray* omitted — 63 chars of source]

where $y_{1}$ $\left( n\times 1\right) $ is the dependent variable, $Y_{2}$ $ \left( n\times p\right) $ are the endogenous variable(s) of interest, and $Z$ $\left( n\times k\right) $ are the excluded instruments, while $u$ $\left( n\times 1\right) $ and $V_{2}$ $\left( n\times p\right) $ are the unobserved structural-form errors. The single endogenous regressor case simply corresponds to $p=1$. We will always take $k\geq p$ with $k=p$ corresponding to the just-identified model and $k>p$ corresponding to the over-identified model. The parameter of interest is $\beta $. It is straightforward to accommodate additional covariates (including a constant), but we omit their inclusion in the exposition below.

The reduced-form model is:

eqnarray*[eqnarray* omitted — 68 chars of source]

where $u\equiv v_{1}-V_{2}\beta $. It will be convenient to write the model in a matrix form:

equation*[equation* omitted — 28 chars of source]

where $Y=\left[ y_{1}:Y_{2}\right] $, $V=\left[ v_{1}:V_{2}\right] $, and $A= \left[ \beta :I_{p}\right] $. We will use the notation $Y_{i}$, $V_{i}$, $ Z_{i}$, etc, to denote the $i$-th row of the corresponding matrix.

In Moreira03, the rows of $V$ were assumed to be i.i.d. This paper relaxes this assumption for the derivation of the Conditional Wald test robust to different DGPs. We are motivated by the observation that applied researchers typically prefer not to make the assumption of homoskedasticity, and often they are interested in a clustered error structure, for example.

The Robust Conditional Wald Tests

In this section, we derive the Conditional Wald (CW) tests robust to HAC errors for the linear model given in section (ref).

The Wald statistic is formed by three elements: a null value $\beta _{0}$, an estimator $\widehat{\beta }_{n}$, and a robust asymptotic variance estimator $\widehat{A.Var}$ for $\sqrt{n}\left( \widehat{\beta }_{n}-\beta _{0}\right) $:

equation*[equation* omitted — 183 chars of source]

In section (ref), we review the class of linear GMM estimators for $\beta $, which includes common estimators like 2SLS, LIML, efficient two-step GMM, and the CU (continuously updating) GMM estimator, all of which can be used for constructing a Robust Conditional Wald test. In section (ref), we review the robust variance estimators based on the asymptotic distribution of $\sqrt{n}\left( \widehat{\beta }_{n}-\beta _{0}\right) $. With these components in hand, we can form robust versions of the Wald statistic for various estimators, $\widehat{\beta }_{n}$. Note that there are no new results in Sections (ref) and (ref), and there are many references that detail these standard results (as one example, see NeweyMcFadden94). We review a selected set of well-established facts about GMM to highlight that the Robust Conditional Wald test we derive in (ref) is not specific to the leading case in applied work -- 2SLS -- and can be applied to tests based on other estimators of the parameter of interest. Note that the multitude of different estimators that could be employed arises in the over-identified case; in contrast, for example, in the single instrument case, all of the estimators we discuss below collapse to the standard IV estimator.

With that as context, an interested reader can skip to section (ref), in which we show how to apply the conditional argument of Moreira03 to obtain a critical value function for the Wald statistic that is robust to instrument weakness. The critical value function is then used to form a robust similar test, which can be inverted to generate confidence intervals for $\beta$.

Estimators

Estimators like 2SLS or LIML can be viewed as particular GMM estimators based on the linear moment condition:

equation*[equation* omitted — 152 chars of source]

where $b=\left( 1,-\beta ^{\prime }\right) ^{\prime }$. A GMM estimator for $ \beta $ is the minimizer of the criterion

equation[equation omitted — 182 chars of source]

where the weighting matrix may or may not depend on the unknown coefficient $ \beta $. Different choices of $W_{n}\left( \beta \right) $ will lead to different estimators $\widehat{\beta }_{n}$.

For the 2SLS estimator, the weighting matrix does not depend on $ \beta $:

equation*[equation* omitted — 137 chars of source]

where $\widehat{V}_{u}$ is an estimator of $V_{u}$, which is the variance of $u$. Because $ \widehat{Q}_{n}\left( \beta \right) $ is quadratic in $\beta $, it is straightforward to show that the resulting estimator is 2SLS:

equation[equation omitted — 210 chars of source]

We can re-express the estimator as

equation*[equation* omitted — 127 chars of source]

where $\widehat{Y}_{2}=NY_{2}$ and $N=Z\left( Z^{\prime }Z\right) ^{-1}Z^{\prime }$ is the usual projection matrix. That is, in the first stage, we first regress $Y_{2}$ on $Z$ to obtain the fitted values $\widehat{ Y}_{2}$. In the second stage, we regress $y_{1}$ on the fitted values $ \widehat{Y}_{2}$.

For the LIML estimator, the weight matrix is formed by using $\beta $ along with residuals from OLS regressions of $y_{1}$ and $Y_{2}$ on $Z$. Let $b=\left( 1,-\beta ^{\prime }\right) ^{\prime }$, $ N=Z\left( Z^{\prime }Z\right) ^{-1}Z^{\prime }$, and $M=I-N$. Then

equation*[equation* omitted — 184 chars of source]

where $\widehat{V}=MY=MV$. The GMM\ criterion is no longer quadratic in $ \beta $ once we use

equation*[equation* omitted — 197 chars of source]

Instead it is a ratio of quadratic forms:

equation*[equation* omitted — 381 chars of source]

The minimum of

equation*[equation* omitted — 114 chars of source]

is well-known to lead to the LIML\ estimator (see DavidsonMacKinnon21 among others). This estimator is proportional to the eigenvector associated to the smallest eigenvalue $\underline{\lambda } _{n}$of the characteristic polynomial $\left\vert Y^{\prime }NY-\lambda .Y^{\prime }MY\right\vert =0$.

Once we consider different weighting functions $W_{n}\left( \beta \right) $, we can wonder if there is the \textquotedblleft best\textquotedblright\ possible choice. The answer depends if errors are heteroskedastic, clustered, etc. Only in special cases, such as with homoskedastic errors, are the 2SLS and LIML\ estimators \textquotedblleft best.\textquotedblright\ To obtain the weighting function that optimally accounts for heteroskedasticity, clustering, serial correlation, and other departures from homoskedastic errors with no serial correlation, one first considers the (infeasibly estimated) variance of

equation*[equation* omitted — 315 chars of source]

Under general conditions for the DGPs, the limiting variance exists and is given by

equation*[equation* omitted — 138 chars of source]

Since we do not observe the errors $V_{i}$, we can make this feasible by replacing them with, as an example, the OLS\ residuals $ \widehat{V}_{i}$. There are different estimators for the variances and covariances above, each one of them suited to different assumptions on the DGPs. For example, typically, one uses the variance estimate of White80 for heteroskedastic errors that are serially uncorrelated:

equation*[equation* omitted — 164 chars of source]

Henceforth, we will employ the broader notation $\widehat{\Omega }_{n}$ without explicitly specifying its formulae for different departures from homoskedasticity. Examples of robust variance estimators include White80 for heteroskedasticity, NeweyWest87 and Andrews91 for both heteroskedasticity and autocorrelation (HAC), and CameronGelbachMiller11 for clustered errors. See AndrewsMoreiraStock04 for the IV model.

The GMM criterion is then

eqnarray*[eqnarray* omitted — 497 chars of source]

with $\widetilde{\beta }_{n}$ being a preliminary consistent estimator of $ \beta $. Again, this criterion

equation*[equation* omitted — 95 chars of source]

is quadratic in $\beta $ and we can easily find its closed-form solution:

equation[equation omitted — 164 chars of source]

which is the two-step GMM\ estimator. It simplifies to the 2SLS estimator if $W_{n}^{-1}$ is proportional to $n^{-1}Z^{\prime }Z$. When the weight matrix depends on $\beta $, we obtain

eqnarray*[eqnarray* omitted — 382 chars of source]

which is the Continuously Updating (CU) GMM estimator, proposed by HansenSingleton82.

Wald Test Statistics

Finally, the usual Wald test statistics are based on the standard asymptotic approximation to the distribution of the GMM estimators. Under the true parameter $\beta _{0}$,

equation*[equation* omitted — 207 chars of source]

for $b_{0}=\left( 1,-\beta _{0}^{\prime }\right) $. For convenience, we derive the asymptotic distribution where we use the parameter $\beta _{0}$ in the criterion function:

equation*[equation* omitted — 170 chars of source]

This setup allows for the possibility that the limiting behavior of $ W_{n}\left( \beta _{0}\right) $ is not necessarily proportional to $ V_{0}^{-1}$. Hence, $\widehat{\beta }_{n}$ is not necessarily optimal. Under the usual asymptotics, the distribution of estimators which minimize $ \overline{Q}_{n}\left( \beta \right) $ is the same as if we had used $ \widehat{Q}_{n}\left( \beta \right) $ instead, where the weight uses a preliminary estimator or uses $\beta $ itself (derivations of these results are standard and can be found, e.g. in NeweyMcFadden94). As a result, the 2SLS and LIML estimators are asymptotically equivalent, while the two-step GMM\ and CUE estimators are asymptotically equivalent as well. We derive the asymptotic distribution for the 2SLS and two-step GMM estimators from equations ((ref)) and ((ref)) and, so, for the LIML\ and continuously updating estimators as well.

For the 2SLS estimator, we can write

equation*[equation* omitted — 255 chars of source]

Assuming that the following probability limits exist,

equation*[equation* omitted — 454 chars of source]

we then have $n^{1/2}\left( \widehat{\beta }_{n}-\beta _{0}\right) \rightarrow _{d}N\left( 0,B_{0}^{-1}A_{0}B_{0}^{-1}\right) $, where

eqnarray*[eqnarray* omitted — 305 chars of source]

We can find some consistent estimators for $A_{0}$ and $B_{0}$, and derive a Wald statistic for the 2SLS and LIML\ estimators:

eqnarray*[eqnarray* omitted — 595 chars of source]

with $\widehat{b}_{n}=\left( 1,-\widehat{\beta }_{n}^{\prime }\right) ^{\prime }$ based on the respective 2SLS/LIML\ estimator.\footnote{ We typically use $\left( \widehat{b}_{n}\otimes I_{k}\right) ^{\prime } \widehat{\Omega }_{n}\left( \widehat{b}_{n}\otimes I_{k}\right) $ as a consistent estimator for $V_{0}$. However, other estimators are possible, including $\left( b_{0}\otimes I_{k}\right) ^{\prime }\widehat{\Omega } _{n}\left( b_{0}\otimes I_{k}\right) $.}

Likewise, for the two-step GMM\ estimator, we find that $n^{1/2}\left( \widehat{\beta }_{n}-\beta _{0}\right) \rightarrow _{d}N\left( 0,B_{0}^{-1}\right) $, where

equation*[equation* omitted — 167 chars of source]

We can find a consistent estimator for $B_{0}$ and derive a Wald statistic for the two-step GMM\ and continuously updating estimator:

eqnarray*[eqnarray* omitted — 389 chars of source]

with $\widehat{b}_{n}=\left( 1,-\widehat{\beta }_{n}^{\prime }\right) ^{\prime }$ based on the GMM/CU estimators.\footnote{ We can also use here either the null value $\beta _{0}$ or the preliminary estimator $\widetilde{\beta }_{n}$ for the variance estimator.}

Valid Critical Value Functions

As emphasized in Dufour97, since the nuisance parameter representing the strength of the first stage may be arbitrarily close to zero, then the usual constant critical values cannot be valid; indeed, Dufour97 points out that any valid confidence set in this context must be unbounded with positive probability, which clearly cannot be the case with a constant critical value for any of the Wald statistics mentioned above. To derive valid critical values, using the conditioning strategy of Moreira03, we begin by defining the quantity

equation*[equation* omitted — 93 chars of source]

where the $k$-dimensional vector $R_{1}$ is the first column of $R$ and the $ k\times p$-matrix $R_{2}$ is the last $p$ columns of $R$. The standardization avoids multiplication by the sample size $n$. The asymptotic variance of $vec(R)$ is

equation*[equation* omitted — 275 chars of source]

where the matrix $\Sigma $ is being partitioned by submatrices of columns/rows of dimensions $1$ and $p$. Analogously, we can use the estimator

equation*[equation* omitted — 354 chars of source]

Up to a scale of the sample size $n$, the GMM criterion is

equation*[equation* omitted — 246 chars of source]

where $\overline{W}_{n}\left( \beta \right) =\left( n^{-1}Z^{\prime }Z\right) ^{1/2}W_{n}\left( \beta \right) \left( n^{-1}Z^{\prime }Z\right) ^{1/2}$. It is clear that only $R$ and the weight function $\overline{W} _{n}\left( \beta \right) $ fully determine the estimator $\widehat{\beta } _{n}$. To illustrate this connection, recall that the 2SLS estimator results if $W_{n}^{-1}$ is proportional to $n^{-1}Z^{\prime }Z$. For such a weight, we have $\overline{W}_{n}\left( \beta \right) =I_{k}$, and we trivially have the 2SLS being dependent only on $R$. Indeed, the 2SLS estimator can be written as

equation*[equation* omitted — 95 chars of source]

The same holds for the other estimators as well. For example, we take the LIML\ estimator. When $W_{n}\left( \beta \right) ^{-1}=\widehat{V}_{u}\left( \beta \right) \cdot n^{-1}Z^{\prime }Z$, we have $\overline{W}_{n}\left( \beta \right) =\widehat{V}_{u}\left( \beta \right) ^{-1}I_{k}$. The LIML\ estimator solves

equation*[equation* omitted — 175 chars of source]

Having found that the estimators are completely determined by the standardized reduced-form coefficients $R$ and the function $\overline{W} _{n}\left( \beta \right) $, we can turn our attention to the Wald statistics.

The Wald statistic for the 2SLS/LIML\ estimators has the form

equation*[equation* omitted — 393 chars of source]

Hence, it is a function of $R$ and $\widehat{\Sigma }_{n}$ (or, for LIML, $ \widehat{\Phi }_{n}$). Likewise, the Wald statistic for the two-step GMM\ and CU\ estimators can be written as

equation*[equation* omitted — 322 chars of source]

which again depends only on $R$ and $\widehat{\Sigma }_{n}$ (as long as the preliminary estimator $\widetilde{\beta }_{n}$ depends only on $R$ and $ \widehat{\Sigma }_{n}$ as well, such as the 2SLS estimator). In short, the Wald statistics associated with any of the estimators we have discussed above are functions of $R$, $\widehat{\Sigma }_{n}$, and $\widehat{\Phi }_{n}$ as shown above.

We now apply the conditioning approach of Moreira03, beginning by finding a useful transformation of $R$:

equation*[equation* omitted — 164 chars of source]

That is, $R_{u}=R_{1}-R_{2}\beta _{0}$. Note that the asymptotic variance of $R_{0}$ is given by

equation*[equation* omitted — 218 chars of source]

This quantity can of course be consistently estimated as well (regardless of identification of $\beta $):

equation*[equation* omitted — 293 chars of source]

Consider a transformation of $R$.

equation*[equation* omitted — 113 chars of source]

Given $\widehat{\Sigma }_{n}$, there is a one-to-one transformation between the pair $R$ and $R_{2}$ and the pair $R_{u}$ and $\widehat{D}$. Since we have established that all of the Wald statistics above can be written as functions of $R, \widehat{\Sigma}_n,$ and $\widehat{\Phi}_n$, this means that they can also be written as functions of $R_u, \widehat{D}, \widehat{\Sigma}_n, $ and $\widehat{\Phi}_n$. Importantly, adopting the appropriate assumptions relevant for HAC (e.g. see Kleibergen05 or Andrews16), it can be shown that $\left( R_{u},\widehat{D}\right) \rightarrow _{d}\left( \mathcal{R}_{u},\mathcal{D}\right) $, where $\mathcal{R}_{u}$ and $\mathcal{D}$ are asymptotically normal and independent, with $\mathcal{R}_{u}$ being mean zero with a variance matrix that can be consistently estimated under the null -- that is, $\mathcal{R}_{u}\sim N\left( 0,\Sigma _{uu}\right) $ under the null. As Moreira03 shows, this allows one to establish the distribution of test statistics even in the presence of the unknown nuisance parameter (the mean of $R_2$), since the distribution of $\mathcal{R}_{u}$ conditional on $\mathcal{D}$ is the same as the marginal distribution.\footnote{ We will not standardize here the $R_{u}$ and $D$ statistics. However, we could have worked with their respective standardized versions, $S=\left[ \left( b_{0}^{\prime }\otimes I_{k}\right) \Sigma \left( b_{0}\otimes I_{k}\right) \right] ^{-1/2}Rb_{0}$ and $T=\left[ \left( A_{0}^{\prime }\otimes I_{k}\right) \Sigma ^{-1}\left( A_{0}\otimes I_{k}\right) \right] ^{-1/2}\left( A_{0}^{\prime }\otimes I_{k}\right) \Sigma ^{-1}vec\left( R\right) $, as in MoreiraMoreira19.}

We can write all Wald statistics as

equation*[equation* omitted — 122 chars of source]

(where we explicitly state the distribution of $R_{u}$ depends on the sample size $n$). Its asymptotic behavior is given by

equation*[equation* omitted — 95 chars of source]

where $\Phi =plim$ $\widehat{\Phi }_{n}=plim$ $n^{-1}V^{\prime }MV$ (if the process is ergodic, $\Phi $ is just the variance of the reduced-form errors $ V$). We then find the $1-\alpha $ quantile, say, $c_{\alpha }\left( d,\Sigma ,\Phi \right) $ of the null asymptotic distribution of

equation*[equation* omitted — 135 chars of source]

The final conditional test rejects the null when

equation*[equation* omitted — 205 chars of source]

Generalization to weakly-identified models (including GMM)

Summarizing the setup in Andrews16, it is assumed that there is a sequence of models $F_{n}\left( \theta ,\gamma \right) $, which is indexed by the sample size $n$. To illustrate the extension, we focus on a parameter of interest $\theta \in \Theta $, and presume there is an $l\times 1$ consistently estimable nuisance parameter $\gamma \in \Gamma $. The objective is to test the null hypothesis $\theta =\theta _{0}$, presuming the availability of three quantities: 1) a standardized sample moment vector (or distance function) evaluated at the null, $h_{n}\left( \theta _{0}\right) $\footnote{ For the linear model, we can take either the (standardized) moment condition $h_{n}\left( \beta \right) =\left( Z^{\prime }Z\right) ^{-1/2}Z^{\prime }\left( y_{1}-Y_{2}\beta \right) $ or the distance function $h\left( \Pi ,\beta \right) =\left( Z^{\prime }Z\right) ^{-1}Z^{\prime }Y-\left[ \Pi \beta :\Pi \right] $.}; 2) a sample gradient of $h_{n}\left( \theta \right) $ with respect to $\theta $ evaluated at the null, $\Delta h_{n}\left( \theta _{0}\right) $; and 3) the consistent estimate $\hat{\gamma}$ for $\gamma $.

The main assumptions in Andrews16 are that for any true value $ \left( \theta ,\gamma \right) \in \Theta \times \Gamma $:

equation*[equation* omitted — 267 chars of source]

and

equation*[equation* omitted — 428 chars of source]

and $\hat{\gamma}\overset{p}{\rightarrow }\gamma $. It is further assumed that $\Sigma _{h\theta }$ and $\Sigma _{hh}$ are continuous in $\gamma $, and hence consistently estimable.

The mean $m\left( \theta _{0}\right) $ belongs to a set $M\left( \mu ,\gamma \right) \subseteq R^{k}$, with $\mu \in \mathcal{M}$, and is defined so that when $\theta =\theta _{0}$, $m\left( \theta _{0}\right) =0$. The goal is to test the null hypothesis $\left( m\left( \theta _{0}\right) ,\mu \right) =\left( 0,\mu \right) $ against the alternative $\left( m\left( \theta _{0}\right) ,\mu \right) =\left( \mathcal{M}\backslash \left\{ 0\right\} ,\mu \right) $ , for any unknown value of $\mu $.

We can once again consider the $k\times 1$ quantity

equation*[equation* omitted — 143 chars of source]

which, by construction is independent of $h\left( \theta _{0}\right) $. The Wald statistic based on the 2SLS estimator for the linear model simplifies to

eqnarray*[eqnarray* omitted — 372 chars of source]

We can thus define a nonlinear analog as

eqnarray*[eqnarray* omitted — 419 chars of source]

(where we have suppressed the dependence on $\theta _{0}$), which will converge in distribution to

equation*[equation* omitted — 223 chars of source]

After substituting in $\Delta h=\mathcal{D}+\Sigma _{\theta h}\Sigma _{hh}^{-1}h$, then it is easy to compute the $\left( 1-\alpha \right) $th conditional quantile defined by

equation*[equation* omitted — 102 chars of source]

The test is straightforward to implement as follows: reject the hypothesis if and only if

equation*[equation* omitted — 281 chars of source]

This test will have the property, under the null, that

equation*[equation* omitted — 154 chars of source]

for all values of $d$ and hence

equation*[equation* omitted — 136 chars of source]

as desired.

Conclusion

We are motivated by providing an inference method for researchers interested in the over-identified linear instrumental variables model, and who have a preference for using the 2SLS estimator $\hat{\beta}$ for inference, and who do not wish to rely on the assumption of homoskedasticity. If errors are assumed to be homoskedastic, one can use the results of StaigerStock97 and StockYogo05 to control the amount of distortion in inference. As noted in AndrewsStockSun19, the tables in StockYogo05 do not apply to non-homoskedastic settings. Andrews18 provides a conservative two-step procedure that builds on StockYogo05 for more general DGPs.

To accommodate practitioners' preference for using the 2SLS estimator and conventional robust standard errors, we present the robust Conditional Wald (data-dependent) critical values for the Wald statistics robust to heteroskedastic, autocorrelated, and/or clustered errors, which turns out to be a relatively straightforward extension of the Conditional Wald test of Moreira03; its derivation has been neglected in the weak-IV literature, which has provided a number of other non-Wald procedures that are both robust to non-homoskedastic errors and to arbitrarily weak instruments.

Using existing results from the weak IV literature, we also generalize the procedure to apply to the more general nonlinear models that are typically estimated via minimum distance or GMM. We can explore several Wald statistics within the nonlinear setup as well, contingent on the weights employed in the criterion function. The final conditional test would substitute the conventional critical value with a conditional quantile.

{

singlespace

}