EconBase
← Back to paper

Weak Identification in Discrete Choice Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

141,998 characters · 21 sections · 152 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Weak Identification in Discrete Choice Models

abstractWe study the impact of weak identification in discrete choice models, and provide insights into the determinants of identification strength in these models. Using these insights, we propose a novel test that can consistently detect weak identification in commonly applied discrete choice models, such as probit, logit, and many of their extensions. Furthermore, we demonstrate that when the null hypothesis of weak identification is rejected, Wald-based inference can be carried out using standard formulas and critical values. A Monte Carlo study compares our proposed testing approach against commonly applied weak identification tests. The results simultaneously demonstrate the good performance of our approach and the fundamental failure of using conventional weak identification tests for linear models in the discrete {choice} model context. Furthermore, we compare our approach against those commonly applied in the literature in two empirical examples: married women labor force participation, and US food aid and civil conflicts.

\noindentKeywords:{ Discrete Choice Models; Weak Instruments; Weak identification; Identification Testing}

Introduction

A prevalent aspect of econometric research concerns estimating the causal impact of some policy relevant treatment variable $y_{2}$ on an outcome variable of interest $y_{1}$. The outcome $y_{1}$ is often qualitative in nature, and the treatment $y_{2}$ is often endogenous when using observational data in empirical studies. For {example}, there is a growing body of research that studies the causal effect of {certain} economic conditions on the incidence of civil conflict in developing countries. {In this context,} economic conditions may be summarized by a state variable such as “economic growth" (see e.g. miguel2004economic) or by a policy tool {such as} US Food Aid (see nunn2014us). In such settings, the most common modelling strategy is to characterize the qualitative outcome variable $y_{1}$ {as a known function of a latent quantitative variable $y_{1}^{\ast }$, and with $y_{1}^{\ast }$ driven by a regression equation:}

equation[equation omitted — 105 chars of source]

{where $y_{2i}$ denotes the (scalar) variable whose causal impact is of interest}, $x_{i}$\ denotes a vector of $k_{x}$ exogenous variables and, for the sake of expositional simplicity, {we consider $i=1,2,...,n$ as indicating independent and identically distributed (i.i.d.) cross-sectional realizations of the respective random variables.}\footnote{In the introduction, we use the terminology “exogenous” to refer to the explanatory variables $x_i$ and to the instrumental variables $z_i$. In Section (ref) we define, following newey1999nonparametric, a precise concept of control variables that is related, but not equivalent, to the common concept of exogeneity.} {We restrict our attention to settings where the relationship between the unobservable $y_{1i}^{\ast}$ and the observable $y_{1i}$ is given by a threshold crossing mechanism and conventionally specified as follows:}\footnote{At the cost of more involved notations, the methodology developed in this paper can easily be extended to a wide variety of multinomial models, such as, for instance, ordered probit models. To some extent, the binary case considered here is the most extreme case of information loss with respect to the observability of the latent variable.}

equation*[equation* omitted — 45 chars of source]

{The causal analysis of interest is conducted through} statistical inference on the true unknown value of the causal parameter $\alpha $ that must be carefully defined in order to account {for the (possible) presence of simultaneity. {However, more often than not, the treatment variable $y_{2}$ is not exogenous, which means that the structural model ((ref)) {can} not be interpreted as a model for the conditional expectation of $y_{1}^{\ast }$\ given $y_{2}$ and $x$. For this reason, identification of the structural parameters in (ref), {and in particular the causal effect $\alpha$,} requires a set of valid (i.e., exogenous) instrumental variables (hereafter, IVs) denoted throughout by $z$.}

Critically, identification of the causal effect relies on the relevance of the underlying instruments to the treatment variable, i.e., the “strength" of the IVs. {The consequences and detection of weak IVs has been extensively studied in} linear models, but it is currently unclear how the instrument strength in binary models affects identification of $\alpha$ and therefore any causal interpretation we may obtain in a given analysis.

To illustrate this point, consider the concrete example given by nunn2014us for estimating the impact of US food aid on the incidence of civil conflicts. Let $y_{2i}$ denote the amount of US food aid to country $i$, and assume we are interested in analyzing if $y_{2i}$ has a causal impact on the probability of civil conflict, with the incidence of conflict denoted by a binary variable $y_{1i}$. In this setting, one must be concerned about the existence of reverse causality (“Do countries receive US aid precisely because they are not doing well?") or common cause (“Could US strategic objectives be a common cause for both conflict and food aid receipts?”) regarding these two variables, {which leads nunn2014us to use lagged US wheat production as an IV to identify the causal impact of US food aid.} Whilst nunn2014us consider various versions of linear probability and hazard models involving different definitions for the binary outcome of civil conflict, {across the various specifications the only measures of identification strength used by nunn2014us to assess the validity of their conclusions are those explicitly designed for linear models, such as the Kleibergen-Paap F-statistics from the first stage regression (kleibergen2006generalized), which are not statistically valid in either binary or hazard models.}

The goal of this paper is to understand, characterize, and quantify the concept of identification strength as it pertains to discrete choice models. We make three primary contributions. First, we give a novel characterization of identification strength in endogenous discrete choice models which demonstrates that identification can be significantly impacted by factors other than the linear correlation between the instruments {and the endogenous variables.} Our second contribution is to use this characterization of identification strength to propose a consistent test for the null hypothesis that “identification is so weak that point estimators are inconsistent," while under the alternative consistent estimation is warranted. Our final contribution is to demonstrate that, under the alternative, we can carry out Wald-based inference in the standard manner.

We now discuss these contributions in more detail, and place them into the broader literature on weak identification.

Testing Identification Strength: Existing Literature and Contributions

Since the analysis of staiger1997instrumental, practitioners have used the well-regarded “rule-of-thumb" to measure instrument strength in the case where $y_{1i}^\ast$ is observed. The magnitude of the $F$-statistic from the reduced form regression equation is arguably the most common measure for determining instrument strength in the linear regression model. Subsequent to the development of the rule-of-thumb, several influential refinements of this measure, and indeed the very concept of weak instruments in the linear model, have been put forward. stock2005testing provide a quantitative definition of weak instruments in the linear model, and use this definition to propose a formal test for instrument weakness. While the approach of stock2005testing relies on conditionally homoskedastic and serially uncorrelated regression errors, an extension of the stock2005testing testing strategy to heteroskedastic and serially correlated errors is devised in olea2013robust.

However, when one moves to general nonlinear models, the impact of instrument weakness on the resulting estimates is more difficult to ascertain. As presented in antoine2009efficient,antoine2012efficient, and following the work of hahn2002discontinuities and caner2009testing, there can exist a range of identification strengths in nonlinear models, between the extreme cases of weak identification (when estimators are not consistent) and strong identification (when estimators are consistent and asymptotically normal at the $n^{1/2}$ rate). Indeed, these authors have shown that generalized method of moments (GMM) estimators can be consistent at a rate slower than the canonical rate of $n^{1/2}$, but only in the case of a convergence rate strictly larger than $n^{1/4}$ is standard inference based on the normal distribution approximation warranted. The key issue is that, when convergence is too slow and the model is nonlinear, second-order terms in Taylor expansions, which govern the behavior of the estimator, may not be negligible relative to the first-order terms, so that standard asymptotic inference may no longer be valid. Such slow rates of convergence have also been documented in the case of many weak instruments (see newey2009generalized and references therein) while a general study of nearly strong instruments is available in andrews2012estimation.

Using this characterization of varying identification strength, antoine2017testing have devised a testing strategy that is capable of detecting (certain levels of) instrument strength in nonlinear models estimated by GMM. The proposed test, dubbed the distorted J-test (DJ test), is based on computing the GMM J-test statistic at a perturbed value of the continuously updated GMM (CUGMM) estimator. The logic behind the test is that, if identification is truly weak, a small perturbation of the moments within the J-statistic will not significantly alter its value, while if identification is not weak this perturbation will result in a significant increase in the value of the J-statistic. Similar to other inference strategies robust to weak identification, the approach explicitly relies on the nature of the CUGMM objective function, which, as originally pointed out by stock2000gmm, automatically controls the behaviour of the GMM objective function under weak identification.

Interestingly, antoine2017testing have demonstrated that their DJ test is akin to the standard rule-of-thumb when the model is linear and homoskedastic. In contrast, they stress (see also windmeijer2019two for related work in the context of clustering) that this DJ test differs from standard “robustified” versions of the rule-of-thumb in case of a heteroskedastic linear model. We note, in particular, that when using linear probability models, one is faced (besides the well-known criticisms of this approach) with a severely heteroskedastic linear model.

Herein, we adapt the general testing strategy of antoine2017testing to the case of discrete choice models and construct a consistent test for the null hypothesis that the instruments are too weak to allow consistent point estimation. Following the nomenclature of antoine2017testing, we also refer to this test as a distorted J-test (DJ test) in this binary model context. Similar to antoine2017testing, we demonstrate that our DJ test can be interpreted as a natural “generalized rule-of-thumb" in the context of discrete choice models, in the sense that {this test appropriately modifies the standard approach to account for both heteroskedasticity and non-linearity.

We compare the performance of this test with the aforementioned existing} approaches both through Monte Carlo experiments and the analyses of two empirical examples. Monte Carlo results show that our DJ test, albeit conservative, {has respectable power. However, the crucial feature of this approach is its ability to discern} that the underlying estimator may not be reliable, while in contrast, the standard rule-of-thumb, because it overlooks information lost due to the nonlinearity of the model, will severely over-reject the null of weak identification. When applied to the two examples with real data, our DJ test is able to unambiguously determine when the null of weak identification should be rejected (as in the textbook example of the causal effect of education of married women on their labor force participation, with strong instruments like parents education), while it rightly questions the use of {standard inference approaches} when identification appears weak, as in the second empirical example. In particular, contrary to the naive rule-of-thumb and Stock and Yogo test results, the DJ test casts some doubt on the strength of the IV used in nunn2014us and thus on the consistency of the estimated negative effect on war offset and the conclusion that food aid may prolong the duration of conflict.

In addition to the development of our DJ test, this paper also reinforces the asymptotic theory developed in antoine2009efficient,antoine2012efficient regarding inference with nearly-strong instruments. By characterizing the strength of instruments in terms of a drifting data generating process, a la staiger1997instrumental and stock2000gmm, we demonstrate that once the null hypothesis of estimator inconsistency has been rejected, Wald-based inference can be performed as normal, up to the effect of pretesting.\footnote{For simplicity, and following antoine2017testing, we choose to overlook the effect of pretesting on the resulting inferences in this current work.} This result is in stark contrast to the existing results for general nonlinear models under weak identification, where it has been shown that standard inference is only warranted once the rate of convergence is strictly larger than $n^{1/4}$. {The ability to perform standard inference in this setting stems from the fact that discrete choice models, while nonlinear, are built from latent linear models, which ensures that they are close enough to linear models to permit} standard inference once the underlying estimator is consistent. {While the convergence rate of the resulting estimator may be very slow, the studentization performed in computing Wald test statistics make their behavior consistent with the standard critical values.} In short, if our DJ test rejects the null of estimator inconsistency (which will be accomplished asymptotically with probability one under the alternative), the practitioner can safely apply standard inference procedures.

In this respect, our recommendation remains true to the widespread practice of a two-stage decision rule: a pretest for weak IV followed by standard inference when the null of weak identification is rejected. Of course, an alternative would be to use more computationally demanding inference strategies that are robust to weak identification. The robust approach proposed by kleibergen2005testing has been extended by magnusson2010inference to the context of limited dependent variable models. More generally, while the existence of weak IV is a common phenomena, there is little theoretical evidence regarding the properties of GMM estimators in endogenous discrete choice models. Using Monte Carlo simulations, dufour2013weak demonstrate the poor behavior of Wald and Likelihood Ratio tests in the presence of weak instruments. finlay2009implementing analyze the Wald test in probit models with weak instruments, and find that the test can significantly over-reject the null hypothesis.

We note that the development of a consistent test for weak instrument in discrete choice models is particularly important since the similarity between linear models and common discrete choice models has led researcher to apply tests that are appropriate for linear models in this nonlinear context. In particular, it is relatively common to see researchers apply the rule-of-thumb developed for the linear model to detect the presence of weak instruments in discrete choice models: see, e.g., miguel2004economic, arendt2005does, mckenzie2011can, cawley2012medical, block2013education and goto2016cartel. However, the above studies do not question the validity of the rule-or-thumb when it is applied in discrete choice models. Other researchers prefer to abandon the discrete choice framework in favor of the linear probability model; see, e.g., lochner2004effect, powell2005importance, kinda2010investment, ruseski2014sport. Besides the fact that they are heavily heteroskedastic, linear probability models are by definition misspecified. Since our DJ test is based on a distortion of the standard J-test statistic for misspecification, it should not be used in the context of misspecified moment models.

The remainder of the paper is organized as follows.

Section 2 introduces our model setup and assumptions. The key maintained assumption is the existence of a control function, in which the conditional probability distribution of the structural error term, given all the variables in the reduced form regression, coincides with the conditional distribution of the structural error term conditional on the reduced form error term. The control function approach for probit with endogeneity has been pioneered by rivers1988limited and led them to put forward a {two-stage conditional maximum likelihood (2SCML) approach, and blundell2004endogeneity propose a nonparametric extension that does not require certain of the parametric assumptions underlying the 2SCML approach.} In this section, we note that a GMM framework allows us to obtain asymptotically equivalent estimators for the structural parameters without necessarily resorting to a two-stage approach. Moreover, we show that our GMM approach is also versatile enough to encompass the Quasi-LIML approach of wooldridge2014quasi.

In Section 3, we present our DJ test and prove its asymptotic properties: size control (under the null of weak identification) and consistency (under the alternative). We further demonstrate that as long as the estimators are consistent (i.e., under the alternative to the null hypothesis of weak identification), standard Wald-style inference can be applied. This stands in contrast to the general case of identification strength for nonlinear models considered in antoine2009efficient,antoine2012efficient and andrews2014gmm, where it is shown that in nonlinear models standard inference approaches are warranted only when the rate of convergence is faster than $n^{1/4}$. Lastly, we demonstrate that, in the context of a discrete choice model, the DJ test can be interpreted as a generalized rule-of-thumb that accounts for the nonlinear nature of the probit model.

Monte Carlo experiments in Section 4 compare the finite-sample properties of our proposed test as well as the performance of other weak IV tests. Section 5 applies our test in two empirical examples: married women labor force participation (wooldridge2010econometric), and US food aid and civil conflicts (nunn2014us). Section 6 concludes.

General Framework

{blundell2004endogeneity propose a control function (hereafter, CF) approach to conduct inference on the structural parameters of endogenous binary choice models. In this and the next section, we examine the impact of weak instruments on such a CF approach to inference. However, we first demonstrate the general point that a CF approach allows us to see} both the 2SCML of rivers1988limited and the Quasi-LIML approach of wooldridge2014quasi as particular cases of a class of GMM estimators, which we discuss in Section (ref). While these GMM estimators can always be characterized by a one-step minimization problem, {using similar arguments to those in Section 6 of newey1994large,} we can also interpret the estimator of the structural parameters as a two-step estimator, {whereby a preliminary plug-in estimator (obtained from a reduced form regression equation) is used within the moments. After establishing the general framework, in Section (ref) we then sketch the weak IV issue in the context of probit models.}

Model and Control Function Approach

newey1999nonparametric suggest that the key for a CF approach is to start from a triangular simultaneous equations model. In the context of endogenous binary choice models, this entails specifying structural and reduced form regression equations, and the mechanism generating the binary responses.

{The structural equation characterizes the response of an unobservable} endogenous variable $y_{1i}^{\ast }$, {conditional on a scalar-valued endogenous variable $y_{2i}$ and a $k_x$-dimensional vector of explanatory variables $x_{i}$}, as the sum of an {unknown structural} function $g\left( y_{2i},x_{i}\right) $ and a structural error term $u_{i}$:\footnote{While imbens2009identification propose an even more general structural model where the error term $u_{i}$ may not be additively separable at the cost of more restrictive independence assumptions, such an extension is beyond the scope of this paper.}

equation[equation omitted — 116 chars of source]

{For sake of expositional simplicity, we will maintain the following linear specification for the structural function }

equation*[equation* omitted — 82 chars of source]

but we note that the analysis remains applicable to any situation where $g\left( y_{2i},x_{i}\right) $ is a parametric function of $\left( y_{2i},x_{i}\right) $; the case of nonparametric $g(\cdot)$ is beyond the scope of this current paper, and is left for future research. {Our primary focus of interest is the case where only the sign of the quantitative structural variable $y_{1i}^{\ast }$ is observable, which yields the structural equation defining the observed binary outcome $y_{1i}$:}\footnote{The binary choice model allows us to address the issue of weak identification in the case of maximum information loss going from the quantitative latent variable $y_{1i}^{\ast } $ to the observed variable $y_{1i}$. However, we note that the general methodology developed in this paper would be similarly relevant for any observation scheme that would define $y_{1i}$ as a known function of $y_{1i}^{\ast }$ and $x_{i}$\ (see e.g Tobit model, Gompit model, disequilibrium model, etc.).}

equation*[equation* omitted — 45 chars of source]

A reduced form, or first stage, regression equation relates the endogenous explanatory variable $y_{2i}$\ {to a $k_z$-dimensional vector of valid instrumental variables, $z_i$, and the explanatory variables $x_{i}$}:

equation[equation omitted — 115 chars of source]
remark\normalfont While we have chosen {to view the reduced form} regression equation ((ref)) as the specification of a conditional expectation, we could alternatively follow the quasi-LIML estimation approach of wooldridge2014quasi. In his approach, the reduced form regression equation is only required to be a linear projection of $y_{2i}$ onto $x_{i}$ and $z_{i}$. We will always assume that $ x_{i}$ includes a constant, so that the reduced form error term $v_{i}$ has a zero mean. That is, instead of ((ref)), we could have assumed \begin{equation} y_{2i}=x_{i}^{\prime }\pi +z_{i}^{\prime }\xi +v_{i},\;\mathbb{E}[ v_{i}] =0 , with Cov\left( \left[ \begin{array}{c} x_{i} \\ z_{i} \end{array} \right] ,v_{i}\right) =0 . \end{equation}
remark\normalfont As noted by blundell2004endogeneity, the reduced form error term $v_{i}$ often appears to be conditionally heteroskedastic. Taking this possibility into account will allow us to devise more efficient estimators when the reduced form error term is deduced from a conditional expectation rather then from only a linear projection. We will actually combine the advantages of both approaches ((ref)) and ((ref)) by assuming that: \begin{equation} y_{2i}=x_{i}^{\prime }\pi +z_{i}^{\prime }\xi +v_{i},\;\mathbb{E}[v_{i}\left\vert x_{i},z_{i}\right] =0 \end{equation} {However, it must be acknowledged that the linearity assumption for the conditional expectation is restrictive, and prevents us from considering cases} where the endogenous explanatory variable $y_{2i}$ is itself qualitative.\footnote{{We also note that, while blundell2004endogeneity propose a nonparametric estimator of the possibly nonlinear regression function $\pi \left( x_{i},z_{i}\right) $, a given nonlinear parametric form of this regression function would not result either in a significant change in our proposed methodology.}}

{As stressed by newey1999nonparametric, the CF approach does not assume that $ x_{i} $ and $z_{i}$ are valid instruments, in that the approach does not require}

equation[equation omitted — 78 chars of source]

but instead only that

equation[equation omitted — 125 chars of source]

Moreover, it is worth realizing that {neither equation ((ref)) or equation ( (ref)) implies the other.} While we will {eventually maintain a stronger version of equation ((ref))}, i.e., $u_i$ conditionally independent of $x_i,z_i$ given $v_i$, there is no reason to believe that $v_{i}$ is itself independent of $x_{i},z_{i}$ , which jointly with the former conditional independence would be tantamount to joint independence of $(u_{i},v_{i})$ and $(x_{i},z_{i})$, and would in turn imply ((ref)). {In particular, such independence would rule out the possibility of conditional heteroskedasticity for the error term $v_{i}$ in the reduced form regression equation ((ref)).}

As clearly defined by wooldridge2015control, “a control function is a variable that, when added to a regression, renders a policy variable appropriately exogenous." {Typically, the restriction in ((ref))} allows us to rewrite equation ((ref)) as

equation[equation omitted — 137 chars of source]

where \[ \varepsilon_{i}=y_{1i}^{\ast }-\mathbb{E}\left[y_{1i}^{\ast }\mid v_{i},x_{i},z_{i} \right] =u_{i}-\mathbb{E}[u_{i}\left\vert v_{i}\right], \] {which ensures, by definition, that the policy variable is appropriately exogenous; i.e.,} \[ \mathbb{E}[\varepsilon_{i}\left\vert y_{2i},x_{i},v_{i}\right] =0 . \]

In their seminal work, {rivers1988limited note that the only assumption needed to obtain valid inference in the probit model} is that the conditional distribution of $u_{i}$ given $v_{i}$ is normal with a mean that is linear in $v_{i}$ and with a fixed variance. While this condition is satisfied if $(u_{i},v_{i})$ is jointly normal, joint normality is not required in general. Similarly, for general discrete choice models, a CF approach can be constructed by assuming that $\mathbb{E}\left[u_{i}\mid v_{i} \right] $ is linear in $v_{i}$\ and that $\varepsilon_{i}=u_{i}-\mathbb{E}\left[u_{i}\mid v_{i} \right] $ \ is independent of $v_{i}$, along with an assumption that $\varepsilon_i$ has a known continuous cumulative distribution function denoted by $\Phi $. We assume that this probability distribution is symmetric, i.e., $\Phi(\varepsilon)=1-\Phi(-\varepsilon)$, which, together with (ref), allows us to write

flalign*\Pr \left[y_{1i}=1 \mid v_{i},x_{i},z_{i}\right] &=\Pr \left\{\varepsilon _{i}>-g\left( y_{2i},x_{i}\right) -\mathbb{E}\left[u_{i}\mid v_{i}\right] \mid v_{i},x_{i},z_{i}\right\} \\ &=\Phi \left\{ g\left( y_{2i},x_{i}\right) +\mathbb{E}[u_{i}\left\vert v_{i}\right] \right\}.

We now collect the maintained assumptions on the general model in (ref)-(ref).

\noindentAssumption 1: The following conditions are satisfied.

(A.1) (Observation scheme) The observed data $\left\{ s_{i}\right\} _{i=1}^{n}=\left\{ (y_{1i},y_{2i},x_{i}',z_{i}')'\right\} _{i=1}^{n}$ are from an i.i.d. sample and for some $\kappa >0,\mathbb{E}\left[ \left\Vert s_{i}\right\Vert ^{2+\kappa }\right] <\infty .$

(A.2) (Reduced form regression): $ y_{2i}=\pi \left( x_{i},z_{i}\right) +v_{i},$ and $\mathbb{E}[v_{i}\mid x_{i},z_{i} ] =0.$

(A.3) (Structural equation): (i) $\mathbb{E}[u_{i}\left\vert v_{i},x_{i},z_{i}\right] =\mathbb{E}[u_{i}\left\vert v_{i}\right]$; (ii) $\Phi$ is a known cumulative distribution function, twice continuously differentiable and strictly increasing, such that $\Phi (\varepsilon)=1-\Phi (-\varepsilon)$; and (iii) for some unknown parameter $\tilde{\rho}\in\mathbb{R}$, $$ \Pr [y_{1i}=1\mid v_{i},x_{i},z_{i}] =\Phi [ g\left( y_{2i},x_{i}\right) +\tilde{\rho}v_{i}].$$

(A.4) (Linearity): The unknown functions $g(\cdot,\cdot)$ and $\pi(\cdot,\cdot)$ are linear:

itemize• For unknown parameters $\alpha\in\mathbb{R}$ and $\beta\in\mathbb{R}^{k_x}$, $g\left( y_{2i},x_{i}\right) =\alpha y_{2i}+x_{i}^{\prime }\beta$; • For unknown parameters $\pi\in\mathbb{R}^{k_x}$ and $\xi\in\mathbb{R}^{k_z}$, $\pi \left( x_{i},z_{i}\right)=x_{i}^{\prime }\pi +z_{i}^{\prime }\xi $.

(A.5)(Parameters) The unknown parameters $\theta=(\theta_1',\theta_2')'$, where $\theta_1:=(\tilde{\rho},\alpha,\beta')'$ and $\theta_2:=(\pi',\xi')'$, are of dimension $p=2+2k_x+k_z$. We have $\theta_1\in\Theta_1\subset\mathbb{R}^{k_x+2}$, $\theta_2\in\Theta_2\subset\mathbb{R}^{k_x+k_z}$, $\Theta:=\Theta_1\times\Theta_2$ and $\Theta$ is compact. For $\theta^0$ denoting the unknown true value of $\theta$, we have $\theta^0\in\text{Int}(\Theta)$.

{As already mentioned, the linearity in Assumption (A.4) is innocuous and what follows can be extended to settings where $g\left( y_{2i},x_{i}\right) $ has any parametric single-index structure and to cases where $\pi \left( x_{i},z_{i}\right) $ has any parametric form.} In the more general nonparametric setting, newey1999nonparametric demonstrate that identification by {CF of the structural model is tantamount} to assuming that there is no functional relationship between the random variables $y_{2i},x_{i}$ and $v_{i}$ (see newey1999nonparametric for a precise definition of this concept). With a linear structural function $g\left( y_{2i},x_{i}\right) $, identification of the {structural parameter $\alpha $\ is equivalent to assuming that $y_{2i}$ is not a linear combination of $x_{i}$ and $ v_{i}$, meaning that the reduced form regression depends on $z_{i}$, i.e., $\xi \neq 0$.}

To give a more concise treatment, throughout the remainder we restrict our analysis to the case where $\Phi $ is the CDF of the standard normal distribution and refer to the model:

equation*[equation* omitted — 139 chars of source]

as a probit model. Since only the sign of the latent variable $y_{1i}^{\ast }$ is observed, the probit model generally requires the normalization condition $\text{Var}(u_{i})=1$. {However, it is without loss of generality to instead consider the normalization condition}

equation*[equation* omitted — 100 chars of source]

If $\rho $ denotes the linear correlation coefficient between $ u_{i}$ and $v_{i}$, the above normalization ensures that \[ \text{Var}\left( u_{i}\right) =\tilde{\rho}^{2}\text{Var}\left( v_{i}\right) +1=\rho ^{2}\text{Var}\left( u_{i}\right) +1, \]where $\sigma _{v}=\sqrt{\text{Var}\left( v_{i}\right) }$, \[ \text{Var}\left( u_{i}\right) =\frac{1}{1-\rho ^{2}},\;\tilde{\rho}=\frac{\rho }{ \sigma _{v}\sqrt{1-\rho ^{2}}} , \]and where we have that $\tilde{\rho}$\ is monotonic in $\rho $. Of course, the {simultaneity/endogeneity} problem is in evidence if and only if $\rho \neq 0$ or equivalently $\tilde{\rho}\neq 0$.

Estimating Equations

Throughout the remainder, we partition the parameter vector as $\theta=(\theta_1',\theta_2')'$, where

equation*[equation* omitted — 165 chars of source]

The vector $\theta _{1}$ (resp., $\theta _{2}$) represents the vector of structural (resp., reduced-from) parameters. Following Assumption 1, the {true value of the reduced form parameters $\theta _{2}$ is defined by} the conditional moment restrictions

equation[equation omitted — 199 chars of source]

{For fixed $\theta _{2}$, the true value of the structural parameters $\theta _{1}$\ is} defined by the conditional moment restrictions

equation[equation omitted — 291 chars of source]

and where $$v_{i}\left( \theta _{2}\right) =r_{2i}\left( \theta _{2}\right) =y_{2i}-x_{i}^{\prime }\pi -z_{i}^{\prime }\xi.$$ As usual, we will handle conditional moment restrictions by choosing vectors of instrumental functions, denoted respectively as $\tilde{b}\left( x_{i},z_{i}\right) $\ for ((ref)) and $\tilde{a}\left( y_{2i},x_{i},z_{i}\right) $ for ((ref)), where it is assumed that the moments $\mathbb{E}[\|\tilde{a}( y_{2i}, x_{i},z_{i})\|^{2+\kappa}]$ and $\mathbb{E}[\|\tilde{b}( x_{i},z_{i})\|^{2+\kappa}]$ are finite for some $\kappa>0$. For a given choice of instrumental functions $\tilde{a}\left( .,.,.\right) $\ and $\tilde{b}\left( .,.\right) $, we maintain the following identification assumption.

\noindentAssumption 2 (Identification): The true unknown value $\theta ^{0}=(\theta_1^{0'},\theta_2^{0'})'\in\text{Int}(\Theta)$ is the unique solution $\theta \in \Theta $ to the following moment restrictions:

alignat*{4} Reduced form:\quad&\quad \mathbb{E}[ \tilde{b}( x_{i},z_{i}) r_{2i}( \theta _{2}) ] =0&\iff&\quad\theta_2=\theta^0_2,\\ Structural:\quad&\quad \mathbb{E}[ \tilde{a}( y_{2i},x_{i},z_{i}) r_{1i}( \theta _{1},\theta _{2}^0) ] =0\quad&\iff&\quad\theta_1=\theta^0_1 .

We can summarize the unconditional moment conditions in Assumption 2 as follows: for $H\ge p$, and $H$-dimensional vectors $a_{i}$ and $b_{i}$ of the same dimension, define

equation*[equation* omitted — 321 chars of source]

then Assumption 2 implies that the moment function $g_i(\theta)$ satisfies

eqnarray*[eqnarray* omitted — 98 chars of source]

A GMM estimator of $\theta^0$ can then be constructed using the moment function

equation[equation omitted — 263 chars of source]

In particular, for $W_n$ a sequence of positive-definite $H\times H$ weighting matrix, we can estimate $\theta^0$ using the GMM estimator $$ \hat\theta_n=\arg\min_{\theta\in\Theta}\bar{g}_n(\theta)'W_{n}\bar{g}_n(\theta), where \bar{g}_{n}(\theta )=\frac{1}{n}\sum_{i=1}^{n}g_i(\theta)\equiv

pmatrix[pmatrix omitted — 58 chars of source]

'. $$

remark\normalfont {In general, imposing that some components of the vectors $ a_{i}$ and $b_{i}$ are zero prevents us from choosing optimal instruments, and ultimately results in $\hat\theta_n$ being an inefficient estimator of $\theta^0$.} The characterization of optimal instrumental functions for the joint set ((ref)) and ((ref)) of conditional moment restrictions is non-standard because they correspond to different conditioning variables. The optimal instrumental functions in this case have been characterized by kawaguchi2017moment (see also ai2003efficient for a general study). Their result implies that in case of overidentification and simultaneity ($\tilde{\rho}\neq 0$), the first set $r_{1i}(\theta)$\ of moment conditions is also informative about $\theta _{2}$, so that a more efficient estimator of $\theta _{2}$ (and in turn $ \theta _{1}$) is obtained by an appropriate choice of $a_{i}$ in which all of its components are non-zero.

{While the specific choice of instrumental functions $a_{i}$ and $b_{i}$ may be sub-optimal, this choice allows us to demonstrate the equivalence between a GMM-based approach and the 2SCML approach of rivers1988limited. In particular, for $g_{1i}(\theta)$ and $g_{2i}(\theta)$ defined as in equation (ref)}, we have that

flalign*Cov\left[ g_{1i}(\theta ^{0}),g_{2i}(\theta ^{0})\right] &=\mathbb{E}\left[ \tilde{a}\left( y_{2i},x_{i},z_{i}\right) \tilde{b}^{\prime }\left( x_{i},z_{i}\right) r_{1i}\left( \theta^0 \right) r_{2i}\left( \theta^0_{2}\right) \right]\&=\mathbb{E}\left\{ \tilde{a}\left( y_{2i},x_{i},z_{i}\right) \tilde{b}^{\prime }\left( x_{i},z_{i}\right) r_{2i}\left( \theta _{2}^{0}\right) \mathbb{E}[r_{1i}\left( \theta ^{0}\right) \left\vert y_{2i},x_{i},z_{i}\right] \right\} =0.

Thus, an efficient GMM estimator based on the moment functions in (ref) can be defined as

eqnarray*[eqnarray* omitted — 370 chars of source]

for an appropriate choice of the weighting matrices $W_{1n}$ and $W_{2n}$. Consequently, the components of the first-order conditions for the structural parameters $\theta_1$ are given by

equation[equation omitted — 150 chars of source]

Equation ((ref)) allows us to see the estimator $\hat{\theta}_{1n}$ as a two-step estimator based on the moment conditions

equation[equation omitted — 134 chars of source]

where the nuisance parameter $\theta _{2}^{0}$\ is replaced by a consistent first-step estimator $\hat{\theta}_{2n}$. From (ref), we can see that the estimator $\hat{\theta}_{1n}$ is the solution in $\theta _{1}=\left( \tilde{\rho},\alpha ,\beta ^{\prime }\right) ^{\prime }$\ to the $\left( 2+k_{x}\right) $ orthogonality conditions

equation[equation omitted — 351 chars of source]

The optimal instruments associated with estimation of $\theta_1^0$ in equation (ref) (i.e., where $\theta^0_2$ is known) are given by any consistent estimator of:

eqnarray[eqnarray omitted — 539 chars of source]

where

eqnarray*[eqnarray* omitted — 331 chars of source]

and $\phi \left( x\right) =d\Phi\left(x\right)/dx $ is the probability density function associated to $\Phi $.

Therefore, if one were to choose a consistent estimator of $\gamma_i^*$ as instruments, the estimator $ \hat{\theta}_{1n}$ can be seen as the solution in $\theta _{1}=\left( \tilde{ \rho},\alpha ,\beta ^{\prime }\right) ^{\prime }$\ to the equations:

equation[equation omitted — 463 chars of source]

Equation ((ref)) shows that, for any choice of a consistent first-step estimator $ \hat{\theta}_{2n}$, the estimator $\hat{\theta}_{1n}$\ is a 2SCML estimator a la rivers1988limited.

The Weak IV Issue in the Probit Model

The representation in equation (ref) demonstrates that the general class of GMM estimators for $\theta_1$ defined in equation ((ref)) contains both 2SCML and Quasi-LIML estimators as particular cases. Therefore, we can ascertain the impact of instrument weakness, on these and related methods, by studying instrument weakness in this general class of GMM estimators.

However, before moving to a general study, we give some intuition on the potential impacts of instrument weakness in the case of probit model. These implications are most easily elucidated in the infeasible case where we replace the optimal instruments in equation (ref) with their infeasible counterpart $\gamma_i^\ast$, and where we replace the estimator $\hat\theta_{2n}$ by the true value $\theta^0_2$.

Under these simplification, and under the one-to-one transformation of $\theta_1$ defined by $$ \eta _{1}=\tilde{\rho},\quad\eta _{2}=\alpha +\tilde{\rho},\quad\eta _{3}^{}=\beta - \tilde{\rho}{\pi}^0, $$the infeasible estimator $\tilde\eta_n$ of $\eta^0$ (and thus $\theta_1^0$) can be defined as the solution to $$ \sum_{i=1}^{n}\gamma_i^* \left\{ y_{1i}-\Phi \left[ \eta _{1}\left( -z_{i}^{\prime }{\xi}^0\right) +\eta _{2}y_{2i}+x_{i}^{\prime }\eta _{3}^{}\right] \right\} =\sum_{i=1}^{n}w_i D_i\left\{ y_{1i}-\Phi \left[ \eta _{1}\left( -z_{i}^{\prime }{\xi}^0\right) +\eta _{2}y_{2i}+x_{i}^{\prime }\eta _{3}^{}\right] \right\}=0, $$where $\gamma_i^*=w_i D_i$, $w_i={1}/{\Phi_i(\theta^0)[1-\Phi_i(\theta^0)]}$ and $D_i=\phi_i(\theta^0)(-z_i'\xi^0,y_{2i},x_i')'$.\footnote{The simplification made in the term $D_i$, i.e., replacing $v_i(\theta^0_2)$ by $-z_i'\xi^0$, follows from the row operation on $\gamma^*_i$ which does not affect the solution of the linear equations in (ref) asymptotically.} A Taylor expansion allows us to heuristically write

flalign*y_{1i}&-\Phi \left[ \eta _{1}\left( -z_{i}^{\prime }{\xi}^{0}\right) +\eta _{2}y_{2i}+x_{i}^{\prime }\eta _{3}^\right] \nonumber \\ &\approx y_{1i}-\Phi _{i}\left( \theta ^{0}\right) -\phi _{i}\left( \theta ^{0}\right) \left[ \left( -z_{i}^{\prime }\xi ^{0}\right) \left( \eta _{1}-\eta _{1}^{0}\right) +y_{2i}\left( \eta _{2}-\eta _{2}^{0}\right) +x_{i}^{\prime }\left( \eta _{3}-\eta _{3}^{0}\right)\right].

Using this expansion within the infeasible estimating equations, $\tilde{\eta}_n$ can be seen to solve $$ \sum_{i=1}^{n}w_i D_i\left( \tilde{y}_{1i}-D_i'\eta \right)=0, \text{ where }\tilde{y}_{1i} =y_{1i}-\Phi(\theta^0)+D_i'\eta^0. $$

Consequently, $\tilde{\eta}_n$ is obtained from a weighted least squares regression of $\tilde{y}_{1i}$ on the explanatory variables $D_i= \phi_i(\theta^0)(-z_i'\xi^0,y_{2i},x_i')'$. While the above estimating equations are not identical to those in equation (ref), it is clear from comparing the two that they are of a similar form, and therefore whatever implications are drawn about the later will be sustained by the former.

This regression-based viewpoint yields two important, and interrelated, implications for inference in endogenous binary choice models. First, the linear regression that is considered is not the one suggested by a linear probability model, which would be based on explanatory variables $z_{i}^{\prime }\xi ^{0},y_{2i}, x_{i}$, and not the weighted versions in $D_i$. Second, since the explanatory variables in the regression are weighted by $\phi_i(\theta^0)$, it is inappropriate to focus solely on the contribution of $z_i'\xi^0$ in the reduced form regression as a measure of instrument strength.

remark\normalfont Before moving on, we note that the above type of estimation approach has been dubbed “two-stage residual inclusion" (2SRI) estimation by terza2008two. In particular, using the first stage consistent estimators $\hat{\theta}_{2n}=(\hat{\pi}_{n}^{\prime },\hat{\xi}_{n}^{\prime })^{\prime },$ the estimated first stage residual \[ \hat{v}_{i}=y_{2i}-x_{i}^{\prime }\hat{\pi}_{n}-z_{i}^{\prime }\hat{\xi} _{n} \] is included in the computation of the generalized residual \[ r_{1i}\left( \theta _{1},\theta _{2}\right) =y_{1i}-\Phi \left[ \alpha y_{2i}+x_{i}^{\prime }\beta +\tilde{\rho}v_i(\theta_2)\right]. \] We know from hausman1978specification that, in a fully linear model and as far as estimation of structural parameters $\alpha $\ and $\beta $ is concerned, 2SRI is equivalent to 2SLS. The inclusion of the residual $\hat{v}_{i}$\ in the regression equation ensures that naive OLS would coincide with 2SLS. In addition, terza2008two dub “Two-stage predictor substitution" (2SPS) the direct generalization of 2SLS to our nonlinear context, meaning that in the structural equation, the endogenous variable is simply replaced by its first stage adjusted value, leading to the generalized residual: \begin{eqnarray*} \hat{u}_{i} &=&y_{1i}-\Phi \left[ \alpha \hat{y}_{2i}+x_{i}^{\prime }\beta \right] \\ \hat{y}_{2i} &=&x_{i}^{\prime }\hat{\pi}_{n}+z_{i}^{\prime }\hat{\xi}_{n} \end{eqnarray*} Not surprisingly, terza2008two show that in a nonlinear model, 2SPS is not equivalent anymore to 2SRI and only the latter provides a consistent estimator of structural parameters. The intuition is quite clear. Due to the non-linearity of the function $\Phi \left( .\right) $, plugging in $\hat{y}_{2i}$ to instrument $y_{2i}$ does not fix satisfactorily the endogeneity bias problem.

As alluded to above, it can be misleading to set the focus on the contribution of $z_{i}^{\prime }\xi ^{0}$\ in the reduced form regression to gauge the instrument strength, as is done when using the standard rule-of-thumb. Doing so is akin to overlooking the impact of nonlinearity\ in the same way as that it is wrong to confuse the correct 2SRI and the flawed 2SPS. Indeed, as the above arguments clarify, the relevant variable for capturing instrument strength is not $z_i$, as in the standard linear case, but $ \phi _{i}(\theta ^{0})z_{i}$. Thus, the assessment of identification strength should rather be based on the variability of $\phi _{i}(\theta ^{0})z_{i}^{\prime }\xi ^{0}$.

We can easily illustrate the impact of moving from $z_i'\xi^0$ to $\phi(\theta^0)z_i'\xi^0$ in terms of instrument strength in the probit model, so that $\phi (\cdot)$ is the probability density function of the Gaussian distribution.\footnote{The conclusions given below will remain valid for any other probability distribution with thin tails, such that the variability of the $\phi _{i}(\theta ^{0})z_{i}$ is drastically different from the one of $z_{i}$.} First we recall that that for a real valued variable $\nu $ and any given number $c$, the absolute value of the function $h(\nu )=\nu \phi (c+\nu )$ is decreasing in $\left\vert \nu \right\vert $\ when the latter value is larger than the absolute value of the roots of the polynomial $\left[ 1-c\nu -\nu ^{2}\right] $. Moreover, the rate of this decrease is sharp (converging swiftly to zero) due to the thin tails of the Gaussian distribution.

Using this argument, one may realize that the multiplication of $z_{i}^{\prime }\xi ^{0}$ by

equation*[equation* omitted — 188 chars of source]

erases the variability of $z_{i}^{\prime }\xi ^{0}$, by pruning all its large values. For $Z\sim \mathcal{N}\left( 0,\sigma _z^{2}\right) $, it is useful to illustrate the above point by comparing the variance of $Z\phi (1+Z)$ as a percentage of the variance of $Z$. For various values of $\sigma _z^{2}$ , we collect these ratios in Table (ref) below.

table[table omitted — 740 chars of source]

The results in Table (ref) constitute compelling evidence on the likely flaws of the standard rule-of-thumb in the probit context. It is also worth stressing that, while Table (ref) only displays results with the normalized function $\phi (1+Z),$ the pruning impact of large values of $z_{i}^{\prime }\xi ^{0}$\ within the function $ \phi (\cdot)$ may actually be magnified in finite sample by a large value of the parameter $\tilde{\rho}^{0}$. We may then expect that the pruning effect documented in Table (ref) will be even more detrimental for small values of $\sigma _{v}$\ and/or a large degree of endogeneity $\rho ,$ with both cases corresponding to a large value of $\tilde{\rho}$. These possible perverse effects for the naive rule-of-thumb will be confirmed by the Monte Carlo experiments in Section (ref). These experiments will show that the standard rule-of-thumb will be more prone to over-reject the null of weak instruments in the case of strong simultaneity ($\rho $ close to one) and/or a large signal to noise ratio $\sigma _{z}/\sigma _{v}$\ in the reduced form regression.

A Test for Instruments Weakness

Intuition

Several authors, {such as} kleibergen2005testing, caner2009testing, chaudhuri2017score, stock2000gmm, and antoine2017testing, have discussed the advantages of a continuously updated GMM (CUGMM) approach to efficient GMM estimation in case of possible weak identification. Following the latter two authors, in our context the advantage of the CUGMM approach is that, irrespective of identification weakness, the asymptotic behavior of the CUGMM criterion is always controlled. This feature of the CUGMM criterion will ultimately allow us to obtain a test for instrument weakness that is size controlled and consistent.

To see that this key feature remains true in our setting, recall the specific moment conditions underlying this analysis given by equation (ref); namely, for $\theta_1=(\tilde\rho,\alpha,\beta')'$ and $\theta_2=(\pi',\xi')'$, and $g_{1i}(\theta)=\tilde{a}(y_{2i},x_i,z_i)r_{1i}(\theta_1,\theta_2)$, $g_{2i}(\theta)=\tilde{b}(x_i,z_i)r_{2i}(\theta_2)$, \[ g_{i}(\theta )=r_{1i}(\theta )a\left( y_{2i},x_{i},z_{i}\right) +r_{2i}(\theta_2 )b\left( x_{i},z_{i}\right)=

pmatrix[pmatrix omitted — 46 chars of source]

'. \] Defining the weighting matrix \[ S_{n}(\theta )=

bmatrix[bmatrix omitted — 51 chars of source]

,\;\;S_{jj,n}(\theta)=\frac{1}{n}\sum_{i=1}^{n}\left[g_{j,i}(\theta )-\bar{g}_{j,n}(\theta )\right]\left[g_{j,i}(\theta )-\bar{g}_{j,n}(\theta )\right]^{\prime }, \;(j=1,2), \]we consider a CUGMM estimator (hereafter, CUE) that takes into account the block diagonal structure of the population variance matrix. Then our CUE of $\theta^0$ based on $\bar{g}_n(\theta)=(\bar{g}_{1n}(\theta)',\bar{g}_{2n}(\theta)')'$ is defined as

eqnarray*[eqnarray* omitted — 208 chars of source]

where the notation $J_n(\theta,\tilde\theta)$ differentiates the occurrences of $\theta$ in the moments, $\bar{g}_n(\theta)$, from those in the weighting matrix, $S^{-1}_n(\tilde\theta)$.

{The critical feature of the criterion $J_{n}(\theta,\tilde\theta ) $ is that, by definition,}

equation[equation omitted — 104 chars of source]

while, since $\text{Cov}\left[g_{1i}(\theta^0),g_{2i}(\theta^0)\right]=0$, it follows that $J_{n}\left( \theta ^{0},\theta^0\right) $ {converges in {distribution to a chi-square random variable with $H$ degrees of freedom, denoted throughout as $\chi ^{2}(H)$}}.\footnote{We note that a similar bound remains valid for a general CUGMM setup that does not make use of the block diagonal structure. For the reasons given previously, we focus on this more particular case.}

The general validity of this upper bound, {regardless of the instrument strength, and, hence consistency of $\hat{\theta}_n$,} is the reason why we resort to CUGMM. This upper bound will allow us to control the size of our test for weak identification.\footnote{The upper bound ((ref)) is generally invalid if a first-step estimator of $\theta ^{0}${ is used to estimate the optimal instrumental functions.} The only way to incorporate optimal instrumental functions for $a(y_{2i},x_i,z_i)$ and $b(x_i,z_i)$ would be to use them with a free value of $\theta $\ like in the weighting matrix of CUGMM. The discussion of this alternative approach is left for future research. Also, we note that in the just identified case, the minimum $J_{n}\left(\hat{\theta}_{n},\hat\theta_n\right) $ of $J_{n}(\theta )$\ is asymptotically, with probability one, equal to zero and $S_{n}^{-1}(\theta ) $\ is immaterial. In particular, when using the first-order conditions of some M-estimator, including two-stage conditional maximum likelihood or quasi-LIML, the weighting matrix is irrelevant.}

The key intuition for our test of weak identification is the following observation. {Under weak identification, there are certain directions of the parameter space where the CUGMM objective function $J_{n}( \cdot,\hat\theta_n) $ is flat in the neighbourhood of $\hat{\theta}_{n}$. In these directions, if we distort $\hat{\theta}_{n}$ by some “small” value, say $\Delta_n\in\mathbb{R}^{p}$, and evaluate $J_n(\cdot,\hat\theta_n)$ at $\hat{\theta}^\delta_n=\hat{\theta}_n+\Delta_n$, then the value of $J_n(\hat\theta_n^\delta,\hat\theta_n)$ should not differ “significantly” from that of $J_n(\hat\theta_n,\hat\theta_n)$. Herein, the concept of} “significance" means that $J_{n}( \hat{\theta}_{n}^{\delta },\hat\theta_n) $ exceeds some {pre-specified quantile of the $\chi ^{2}(H)$ distribution.}

{Critically, however, since the objective function scales the squared norm of the sample mean $\bar{g} _{n}(\theta )$, by the factor $n$, when identification is not weak the distortion introduces a wedge between $\bar{g}_n(\hat{\theta}_{n}^{\delta })$\ and $\bar{g}_n(\hat{\theta}_{n})$. Therefore, if identification is not weak, so long as the distortion goes to zero sufficiently slowly with $n$, the criterion $J_{n}( \hat{\theta}_{n}^{\delta },\hat\theta_n) $ diverges asymptotically and thus exceeds (with probability going to one) the chosen quantile of the $\chi ^{2}(H)$ distribution. Throughout the remainder, we refer to this testing procedure as a distorted J-test.}\footnote{It is worth noting that this test is dubbed the “distorted J-test" because it uses the J statistic proposed by hansen1982large in the overidentified case to test for the validity of a set of moments. The terminology is a bit misleading since our test may work even in the just identified case ($H=p$). There are actually two possible points of view: either one chooses to perform the distorted J-test test in a just identified setting ($H=p$), or in the overidentified setting ($H>p$).}

The null hypothesis of weak identification

As already discussed in Section (ref), {weak instruments impact estimation of the structural parameters through the structural moment function} \[ g_{1i}\left( \theta \right) =\tilde{a}\left( y_{2i},x_{i},z_{i}\right) r_{1i}\left( \theta _{1},\theta _{2}\right), \text{ where } r_{1i}\left( \theta _{1},\theta _{2}\right)=y_{1i}-\Phi \left[ (\tilde{\rho}+\alpha) y_{2i}+x_{i}^{\prime }(\beta-\tilde{\rho}\pi) -\tilde{\rho}z_{i}^{\prime }\xi\right]. \] {The impact of weak instruments can be most easily disentangled under the parameterization}

equation[equation omitted — 202 chars of source]

which allows us to restate the moment function as $$ g_{1i}(\eta,\theta_2)=\tilde{a}(y_{2i},x_i,z_i)\tilde{r}_{1i}(\eta,\theta_2), \text{ where }\tilde{r}_{1i}\left( \eta ,\theta _{2}\right)=y_{1i}-\Phi \left[ -\eta _{1}z_{i}^{\prime }\xi +\eta _{2}y_{2i}+x_{i}^{\prime }\eta _{3}\right]. $$

Following staiger1997instrumental and stock2000gmm, we use a drifting data generating process (DGP) to capture instrument weakness, so that population expectations are viewed as being $n$-dependent. However, to paraphrase lewbel2019identification, we do not actually believe that the DGP is changing as $n$ changes, but use the drifting DGP concept in order to obtain more reliable asymptotic approximations in the context of weak identification. To this end, we consider that the population expectation of $\bar{g}_{1n}(\eta,\theta_2)$ is defined as \[ m_{1n}\left( \eta, \theta _{2}\right) =\mathbb{E}_{n}\left[\sum_{i=1}^{n} \tilde{a}\left( y_{2i},x_{i},z_{i}\right) \tilde{r}_{1i}\left( \eta, \theta _{2}\right) \right]/n. \] {Under this drifting DGP, we are obliged to see $\theta^0_2$, and hence $\eta^0$, as $n$-dependent, so that the maintained identification assumption should technically be recast as $$m_{1n}\left( \eta,\theta _{2}\right) =0\iff (\eta,\theta_2)=(\eta^0_n,\theta_{2n}^0).$$ However, to keep the notational burned to a minimum, we only make the true-values dependence on $n$ explicit when absolutely necessary.}

{Following the approach of stock2000gmm (see their Section 2.3), the following decomposition of $m_{1n}\left( \eta, \theta _{2}\right)$ will ultimately allow us to isolate the impact of instrument weakness

eqnarray*[eqnarray* omitted — 374 chars of source]

In particular, since $m_{1n}\left( \eta^{0}, \theta _{2}^{0}\right) =0$, we have}

equation[equation omitted — 274 chars of source]

As explained in Section (ref), instrument weakness is encapsulated by the explanatory variable $\phi_i \left( \theta ^{0}\right) z_{i}^{\prime }\xi^{0}$. {The impact of this explanatory variable on instrument strength can be directly obtained by linearising $m_{1n}\left( \eta, \theta _{2}^{0}\right) $ around $\eta_1^0$ to obtain}

eqnarray[eqnarray omitted — 540 chars of source]

where $\eta _{1n}^{\ast }$ denotes a component-by-component intermediate value, which can vary according to the components of the function $\tilde{a}(.)$.

Equation (ref) allows us to write the decomposition in equation (ref) in the following semi-separable form, which clearly partitions the directions of weakness in the parameter space: for some real, positive, and deterministic sequence $\varsigma_{n}\rightarrow\infty$ as $n\rightarrow\infty$, with $\varsigma_{n}=O(\sqrt{n})$, possibly $o(\sqrt{n})$,

equation[equation omitted — 159 chars of source]

where

flalign*&q_{11,n}\left( \eta \right) =\varsigma_n\left[ m_{1n}\left( \eta, \theta _{2}^{0}\right) -m_{1n}\left( \eta _{1}^{0},\eta _{2},\eta_{3}, \theta _{2}^{0}\right) \right],\\ &q_{12,n}\left( \eta _{2},\eta _{3}\right) =m_{1n}\left( \eta _{1}^{0},\eta _{2},\eta_{3}, \theta _{2}^{0}\right) =\sum_{i=1}^{n}\mathbb{E}_n\left[\tilde{a}\left( y_{2i},x_{i},z_{i}\right) \tilde{r}_{1i}\left( \eta _{1}^{0},\eta _{2},\eta_{3}, \theta _{2}^{0}\right) \right]/n.

Given this decomposition of $m_{1n}\left( \eta, \theta _{2}^{0}\right)$, the identification strength of $\eta_1$ is entirely determined by equation (ref) and therefore $q_{11,n}(\eta)/\varsigma_n$. In particular, the rate $\varsigma_n$ can be thought of as encapsulating the speed with which the curvature of the moments approaches zero in the $\eta_1$ direction, and thus $\varsigma_n$ determines the degree of identification weakness. If $\varsigma_n$ diverges like $\sqrt{n}$, the speed at which this curvature vanishes is matched by the rate at which information accumulates in the sample, i.e., $\sqrt{n}$, and there is no hope that $\eta^0_1$ can be identified from sample information; i.e., $\eta^0_1$ is weakly identified. In contrast, the identification of $\eta_2,\eta_3$ is determined by $q_{12,n}(\eta_2,\eta_3)$ and is not afflicted by identification weakness. That is, in this rotated parameter space of $\eta$, identification weakness only occurs in the $\eta_1$ direction and does not permeate the remaining directions in the parameter space. The representation in equation (ref) is conformable, but not equivalent, to the decomposition employed by stock2000gmm to study the behavior of GMM under weak identification (see Remark (ref) for details). We maintain the following conditions on $m_{1n}(\eta,\theta^0_2)$, which has the same form as Assumption C in stock2000gmm.

\noindentAssumption 3: For $\varsigma_n=O(\sqrt{n})$, possibly $o(\sqrt{n})$, $m_{1n}\left( \eta, \theta _{2}^{0}\right) ={q_{11,n}\left( \eta \right)/\varsigma_{n}}+q_{12,n}\left( \eta _{2},\eta _{3}\right)$:

(i) $q_{11,n}\left( \eta \right)\rightarrow q_{11}\left( \eta \right)$ as $n\rightarrow\infty$ uniformly in $\eta $, where $q_{11}\left( \eta ^{0}\right) =0$, and $q_{11}(\cdot)$ is uniformly continuous (and hence bounded) in $\eta$.

(ii) $q_{12,n}\left( \eta_2, \eta_3 \right)\rightarrow q_{12}\left( \eta_2,\eta_3 \right)$ as $n\rightarrow\infty$ uniformly in $\eta_2,\eta_3 $. For all $n\ge1$, $q_{12,n}\left( \eta _{2},\eta _{3}\right) $ satisfies $q_{12,n}\left( \eta _{2},\eta _{3}\right) =0\Longleftrightarrow \left( \eta _{2},\eta _{3}\right) =\left( \eta _{2}^{0},\eta _{3}^{0}\right)$, and is continuously differentiable, with ${\partial q_{12,n}\left( \eta _{2},\eta _{3}\right)}/\partial (\eta _{2}, \eta _{3}^{\prime })'$ full column rank at $(\eta^0_2,\eta_3^{0'})'$.

remark{ \normalfont Assumption 3(i) is justified by the decomposition in equation ((ref)) and Assumptions 1 and 2. Secondly, we note that Assumption 3 is natural in our context. Assumption 3(ii) enforces that, for $q_{12,n}\left( \eta _{2},\eta _{3}\right)=\sum_{i=1}^{n}\mathbb{E}_n\left\{ \tilde{a}\left( y_{2i},x_{i},z_{i}\right) \left[ y_{1i}-\Phi \left(-\eta^0 _{1}z_{i}^{\prime }\xi^0 +\eta _{2}y_{2i}+x_{i}^{\prime }\eta _{3}\right) \right] \right\}/n$, \begin{eqnarray*} -\frac{\partial q_{12,n}\left( \eta _{2},\eta _{3}\right) }{\partial (\eta _{2}, \eta _{3}^{\prime })'}=\frac{1}{n}\mathbb{E}_n\left\{ \sum_{i=1}^{n}\tilde{a}\left( y_{2i},x_{i},z_{i}\right) \phi _{i}\left( \eta_1^0,\eta_2,\eta_3, \theta _{2}^{0}\right) ( y_{2i}\;\vdots\; x_{i}^{\prime }) \right\} \end{eqnarray*}has full column rank at $(\eta_2^0,\eta_3^{0'})'$. This is tightly related to the requirement that the components of $( y_{2i}\;\vdots\; x_{i}^{\prime })$\ be linearly independent, since they coincide with the explanatory variables of the latent structural equation.}

For the set, $$\Upsilon(\theta^0_2):=\left\{\eta\in\mathbb{R}^{k_x+2}\;:\;\eta=(\tilde\rho,\alpha+\tilde\rho,\beta'-\tilde\rho\pi^{0'})',\text{ for some }\theta_1=(\tilde\rho,\alpha,\beta')'\in\Theta_1\right\},$$ we state the null hypothesis of weak identification as follows.

\noindentNull Hypothesis of Weak Identification:

equation[equation omitted — 343 chars of source]

The set $\Upsilon(\theta^0_2)$ denotes the set of structural parameters under the parametrization in (ref), and with $\theta_2=\theta^0_2$, so that the supremum over $\eta$ in ((ref)) is akin to a supremum over the structural parameters $\theta _{1}$, given the true value $\theta _{2}^{0}$\ of the reduced form parameters. Both sets of structural parameters, the initial one $\Theta _{1}$ and the reparameterized one $ \Upsilon \left( \theta _{2}^{0}\right) $ are compact subsets of $ \mathbb{R}^{k_{x}+2}$. Based on the decomposition of (ref), the identification strength of $\eta_1$ is determined by the rate $\varsigma_n$, and $\varsigma_n=O(\sqrt{n})$ implies that even asymptotically, the population objective function is nearly flat in $\eta_1$. Such asymptotic behavior of the objective function will lead to inconsistent estimation of $\eta^0_1$ in the rotated parameter space and for the structural parameter $\theta_1$ in the original parameter space $\Theta_1$.

remark\normalfont{ It is worth noting that this definition of weak identification is a generalization of stock2000gmm since it is considered at the true value $\theta _{2}^{0}$\ of the parameters of the reduced form regression equation. This must be seen as the relevant extension of the concept of weak instruments for the context of control variables. As explained in Section 2.3, the relevant explanatory variables for the structural equation are ${\phi _{i}\left( \eta, \theta _{2}^{0}\right) ( z_{i}^{\prime }\xi ^{0},y_{2i},x_{i}^{\prime })^{\prime } }$. In particular, it is the impact $z_{i}^{\prime }\xi ^{0}$, at the true value $\xi ^{0}$, that matters for identification and the pruning effect of $\phi _{i}\left( \eta, \theta _{2}^{0}\right) $, also at the true value $\theta _{2}^{0}=\left( \pi ^{0\prime },\xi ^{0\prime }\right) ^{\prime }$. This extension is made possible by the reinforced identification condition in Assumption 2 (identification of $\theta^0 _{2}$ by the second set of moment conditions in isolation) and the choice of block-diagonal weighting matrix.}

A Distorted J-test (DJ test) for the Null of Weak Identification

The decomposition in equation ((ref)), along with Assumption 3, clarifies and confines the weak identification issue, under the parametrization $\zeta=(\eta',\theta_2')'$, to the $\eta_1$ direction. Therefore, to construct a distorted testing approach for weak identification along the lines proposed in Section (ref), it is precisely this direction, and only this direction, that should be distorted.

To this end, let $\mathcal{Z}$ denote the parameter space of $\zeta$ and define the infeasible CUE $$ \hat\zeta_n=\operatorname*{argmin}_{\zeta\in\mathcal{Z}}\bar{g}_n(\zeta)'S_n^{-1}(\zeta)\bar{g}_n(\zeta), $$ and consider distorting the first component of $\hat\zeta_n$ as \[ \hat{\zeta}_{n}^{\delta }=\hat{\zeta}_{n}+

bmatrix[bmatrix omitted — 39 chars of source]

'=[\hat{\tilde{\rho}}_n,\hat{\tilde\rho}_n+\hat\alpha_n,\hat\beta'_n-\hat{\tilde\rho}_n\pi^{0'},\hat\theta_{2n}^{\prime}]'+

bmatrix[bmatrix omitted — 41 chars of source]

'. \] Under the change of basis, this is equivalent to distorting the CUE $\hat\theta_n$ as \[ \left[

array[array omitted — 54 chars of source]

\right] +\left[

array[array omitted — 44 chars of source]

\right] , where \Delta^0 _{1n}=

bmatrix[bmatrix omitted — 64 chars of source]

' , \]which distorts the entire vector of structural parameters $\theta_1$. However, the above perturbation of $\hat\theta_n$ is infeasible as it depends on the unknown $\pi^0$. A feasible perturbation can be produced by replacing $\pi^0$ with its estimated value $\hat\pi_n$, which yields

equation[equation omitted — 323 chars of source]

As explained in Section (ref), under weak identification, if we distort the CUE $\hat{\theta}_{n}$ by some small value in the directions of weak identification, i.e., $\eta_1$, the value of the GMM criterion at $\hat{ \theta}_{n}^{\delta }$\ should not differ significantly from the criterion evaluated at $\hat{\theta}_{n}$. More precisely, recalling the definitions of $\bar{g}_n(\theta)$ and $S_n(\theta)$ given in Section (ref), \[ J_{n}( \theta,\tilde\theta ) =n\bar{g}_{n}( \theta ) ^{\prime }S_{n}^{-1}( \tilde\theta ) \bar{g}_{n}( \theta ), \quad J_{n}( \hat{\theta}_{n},\hat\theta_n) =\min_{\theta\in\Theta }J_{n}( \theta ,\theta), \]we introduce the distorted J-test statistic: \[ J_{n}^{\delta }=n\bar{g}_{n}( \hat\theta^\delta_n ) ^{\prime }S_{n}^{-1}( \hat\theta_n ) \bar{g}_{n}( \hat\theta^\delta_n ). \]

To deduce the behavior of $J^\delta_n$ under the null of weak identification, we must maintain a regularity condition on the Jacobian of the moments. However, given that our null of weak identification is local about $\eta_1$, at the fixed value of $\theta_2^0$, we are only required to maintain the following assumption.\footnote{We note that Assumption 4 is guaranteed under Assumption 1 and a functional central limit theorem. See the proof of Lemma (ref) in the Appendix for details. We state this result as an assumption to ease the comparison with standard results.}

\noindentAssumption 4: Uniformly over $\Upsilon \left( \theta_{2}^{0}\right) $, $\sqrt{n}\left\{\partial \bar{g}_n(\eta,\theta^0_2)/\partial\eta_1-\mathbb{E}_n\left[\partial \bar{g}_n(\eta,\theta^0_2)/\partial\eta_1\right]\right\}\Rightarrow\Psi(\eta,\theta_2^0)$, for $\Psi(\eta,\theta_2^0)$ a mean-zero Gaussian process, and where $\Rightarrow$ denotes weak convergence in the sup-norm.

proposition[Lack of Consistency] If Assumptions 1-4 are satisfied, and if ${\mathbb{E}[\|\tilde{a}(y_{2i},x_i,z_i)z_i'\|^2]<\infty}$, then under the null of weak identification, for any $\delta _{n}=o(1)$, \[ \operatorname*{plim}_{n\rightarrow \infty }\sqrt{n}\left[ \bar{g}_{n}( \hat{\theta} _{n}^{\delta }) -\bar{g}_{n}( \hat{\theta}_{n}) \right] =0 . \]In addition, if $\sup_{\theta \in \Theta }\left\Vert S_{n}^{-1}( \theta ) \right\Vert =O_{p}(1)$, then \[ \operatorname*{plim}_{n\rightarrow \infty }\left[ J_{n}^{\delta }-J_{n}( \hat{\theta} _{n},\hat\theta_n) \right] =0. \]

Proposition (ref) demonstrates that under the null of weak identification, the curvature of the objective function is insensitive to a small departure from the CUE, indicating the lack of consistency of $\hat{\theta}_{n}$. By adapting the general testing approach of antoine2017testing, Proposition (ref) paves the way for a testing strategy for weak instruments in discrete choice models. Recall that the number of model parameters is $p=2+2k_x+k_z$, and $H$ denotes the number of moments.

theorem[Distorted J-test: Under the Null] Under Assumptions 1-4 and the null of weak identification, for any deterministic sequence $\delta _{n}=o(1)$, define the distorted J-test by the rejection region: \[ W_{n}^{\delta }=\left\{ J_{n}^{\delta }>\chi _{1-\alpha }^{2}\left( H+1-p\right) \right\}, \] where $\chi _{1-\alpha }^{2}\left( H+1-p\right) $ is the $ \left( 1-\alpha \right) $\ quantile of the Chi-square distribution with $(H+1-p)$ degrees of freedom. Under the null hypothesis of weak identification, $W_{n}^{\delta }$ {has asymptotic size of at most} $\alpha $.

{As discussed in Section (ref), the CUGMM framework allows us to control the size of our test by ensuring that we can obtain a convenient upper bound for $J_n^\delta$ under the null of weak identification. Since there is only a single direction of weakness in the rotated parameter space, this bound can be based on the $\chi^2(H+1-p)$ distribution; please see the proof of Theorem (ref) for details. While the test statistic $J_n^\delta$ coincides with the one given in Section (ref), we have improved the asymptotic power of the test $W_n^\delta$ by using a critical value calculated from $\chi ^{2}\left( H+1-p\right) $ instead of $\chi ^{2}\left( H\right) $. This power gain is obviously important since we may be afraid that our test would be overly conservative.}

Estimation and Testing Under the Alternative

In this section, we prove that $W_n^\delta$, the distorted J-test based on $J_n^\delta$, is consistent under the alternative. Before presenting this result, we first discuss the asymptotic behavior of the CUE under the alternative.

Estimation Under the Alternative

We first deduce the properties of the infeasible CUE for the rotated parameter $\zeta=(\eta',\theta_2')'$. The vector $\zeta$ represents the following change of basis in the parameter space:

equation[equation omitted — 273 chars of source]

For $\mathcal{Z}$ denoting the parameter space of $\zeta$, the CUE of $\zeta^0$ is given by

equation*[equation* omitted — 161 chars of source]

Once the asymptotic properties of $\hat\zeta_n$ have been deduced, the asymptotic behavior of $\hat\theta_n$ can be ascertained by applying the change of basis $\hat\theta_n=R\hat\zeta_n$ in equation (ref).

To deduce the properties of $\hat\zeta_n$ under the alternative, we first recall that the null of weak identification, defined by (ref), implies that $$ \sup_{\eta\in\Upsilon(\theta^0_2)}\left\|\frac{1}{n}\mathbb{E}_{n}\left\{ \sum_{i=1}^{n}\frac{\partial g_i(\eta,\theta^0_2)}{\partial\eta_1}\right\}\right\|=\sup_{\eta\in\Upsilon(\theta^0_2)}\left\|\frac{1}{n}\mathbb{E}_{n}\left\{ \sum_{i=1}^{n}\tilde{a} \left( y_{2i},x_{i},z_{i}\right) \phi _{i}\left( \eta ^{},\theta _{2}^{0}\right) z_{i}^{\prime }\xi ^{0}\right\}\right\|=O(1/\sqrt{n}). $$The alternative hypothesis to this null implies the existence of a deterministic sequence $\varsigma_{n}=o(\sqrt{n})$ such that $$ \limsup_{n\rightarrow\infty}\sup_{\eta\in\Upsilon(\theta^0_2)}\left\|\frac{1}{n}\mathbb{E}_{n}\left\{ \sum_{i=1}^{n}\tilde{a} \left( y_{2i},x_{i},z_{i}\right) \phi _{i}\left( \eta ^{},\theta _{2}^{0}\right) z_{i}^{\prime }\xi ^{0}\right\}\varsigma_{n}\right\|>0. $$ To deduce the behavior of the CUE $\hat\zeta_n$ under the alternative, we slightly reinforce this condition as follows.

\noindentAssumption 5: Under the alternative hypothesis, there exists a deterministic sequence $\varsigma_{n}=o(\sqrt{n})$ and a continuous, and deterministic vector function $V^0(\eta)$ such that, ${\inf_{\eta\in\Upsilon(\theta^0_2)}\|V^0(\eta)\|>0}$, and\footnote{We note here that $V^0(\eta)$ technically depends on the drifting pseudo-true value $\theta^0_2$, but subsume this dependence in the definition to simply notations.} \[ \lim_{n\rightarrow\infty}\sup_{\eta\in\Upsilon(\theta^0_2)}\left\|\frac{1}{n}\mathbb{E}_{n}\left\{ \sum_{i=1}^{n}\tilde{a} \left( y_{2i},x_{i},z_{i}\right) \phi _{i}\left( \eta ^{},\theta _{2}^{0}\right) z_{i}^{\prime }\xi ^{0}\right\} \varsigma _{n}-V^0(\eta)\right\|=0. \]

remark\normalfont Even though Assumption 5 arguably limits the scope of the alternative hypothesis, it is more general than if we were to follow the approach of staiger1997instrumental and characterize identification strength only through the reduced form regression equation. In the latter case, one would consider that the reduced form regression evolves according to the drifting DGP \[ \mathbb{E}_{n}[y_{2i}\left\vert x_{i},z_{i}\right] =x_{i}^{\prime }\pi^0 +z_{i}^{\prime }\xi^0 _{n}. \]Under the null of weak identification, we have that $\xi^0 _{n}=O(1/\sqrt{n}). $ In contrast, Assumption 5 would require that, for some $\gamma^0\in\mathbb{R}^{k_z} $ with $\|\gamma^0\|>0$ and some $\varsigma_{n}=o(\sqrt{n})$, $$\lim_{n\rightarrow \infty }\varsigma_{n}\xi ^0_{n}=\gamma^0,\;\text{ and } V^0(\eta)=\mathbb{E}_n\left[\tilde{a}(y_{2i},x_i,z_i)\phi_i(\eta,\theta^0_2)z_i'\right]\gamma^0\neq0. $$However, as explained in Section (ref), this approach to characterize identification strength is not sufficient in our opinion, since it only accounts for the instrument strength in the reduced form regression, $\xi^0_n$, and does not account for the interactions between the instrumental function $\tilde{a}(y_{2i},x_i,z_i)$ and $\phi _{i}\left( \eta,\theta _{2n}^{0}\right) z_{i}^{\prime }\xi _{n}^{0}$, which may result in the pruning of large realizations of the instruments via the behavior of $\phi_{i}\left( \eta,\theta _{2n}^{0}\right) $.

By defining the alternative hypothesis using Assumption 5, we clearly partition the two possible cases for estimation of $\zeta^0$: (i) if identification is weak, $\hat{\zeta}_{n}$\ is not consistent (as implied by Proposition (ref)), nor are other commonly applied estimators such as 2SCML or Quasi-LIML estimators; (ii) when identification is not weak, $\hat{\zeta}_{n}$ is consistent.

proposition[Consistency]If Assumptions 1-5 are satisfied, and if ${\sup_{\zeta\in \mathcal{Z}}\|S_n^{-1}(\zeta)\|=O_p(1)}$, then $\|\hat\zeta_n-\zeta^0\|=o_p(1)$.

The asymptotic distribution of $\hat\zeta_n$ depends on the behavior of the Jacobian for the moments. Under Assumption 3 and 5, the scaled Jacobian of the moment functions, as defined below in Lemma (ref), is full rank under the following mild assumption, which, if we take $\tilde{b}(x_i,z_i)=\left(x_i'\;:\;z_i'\right)'$ is nothing but the standard rank condition on the reduced form regression.

\noindentAssumption 6: For all $n\ge1$, $\mathbb{E}_n[\tilde{b}(x_i,z_i)(x_i'\;:\;z_i')]$ has column rank $(k_x+k_z)=\text{dim}(\theta_2)$.

lemmaUnder Assumption 1-6, for a given sequence $\varsigma_{n}=o(\sqrt{n})$, the matrix $$M=\operatorname*{plim}_{n\rightarrow\infty}\left\{\frac{\partial\bar{g}_n(\zeta^0)}{\partial\zeta'}\right\}\Lambda_n,\text{ where }\Lambda_n=\begin{bmatrix}\varsigma_{n}&\textbf{O}_{p-1}\\\textbf{O}_{p-1}&\textbf{I}_{p-1} \end{bmatrix}, $$exists and is full column rank.

Given the full-rank nature of the scaled Jacobian, we would expect the CUE to be asymptotically normal. In particular, under the alternative (as defined by Assumptions 3 and 5), we can then deduce the following result.

theorem[Asymptotic Normality] If Assumptions 1-6 are satisfied then $$\sqrt{n}\Lambda_n^{-1}(\hat{\zeta}_n-\zeta^0)\overset{d}{\rightarrow}\mathcal{N}\left(0,[M'S^{-1}M]^{-1}\right),\text{ where }S:=\operatorname*{plim}_{n\rightarrow \infty }S_n(\zeta^0).$$

As expected, all entries of $\zeta$, save for $\eta _{1}$, are $\sqrt{n}$-consistent and asymptotically normal GMM estimators. In contrast, the direction $\eta _{1}$ converges at the $\{\sqrt{n}/\varsigma_{n}\}$-rate, which is possibly slower than $\sqrt{n}$. Of course, our goal is not to conduct inference on $\zeta^0$, but on $\theta^0$. By the change of basis in (ref), $\theta=R\zeta$, and Theorem (ref) implies that the feasible CUGMM estimator $\hat{\theta}_{n}$ satisfies

equation[equation omitted — 168 chars of source]

Importantly, since the matrix $R$ is not diagonal, the slower rate of $\{\sqrt{n}/\varsigma_{n}\}$ pollutes the entire vector of structural parameters $\theta_{1}=(\tilde{\rho},\alpha,\beta')'$, which follows from the change of basis $\theta=R\zeta$. Therefore, all structural parameter estimates in the probit model converge at the slower $\{\sqrt{n}/\varsigma_n\}$-rate.

Equation (ref) itself does not directly provide a feasible inference strategy since the matrix $R$ depends on the unknown $\pi^0$. Of course the matrix $R$ may be consistently estimated. However, as explained by antoine2012efficient (see the discussion of their Theorem 4.5), a sufficient condition to ensure that the estimation of $R$ does not pollute the asymptotic distribution in (ref) is that the matrix $R$ is estimable at a rate faster than $n^{1/4}$. In the case of the probit model, the matrix $R$ only depends on the unknown true reduced form parameter $\pi^0 $, which is strongly identified and consistently estimable at the $\sqrt{n}$-rate. Therefore, if $\hat{R}_{n}$ denotes the matrix $R$ where $\pi^0 $ is replaced by $\hat{\pi}_{n}$, we can conclude that, following Theorem 4.5 in antoine2012efficient,

equation[equation omitted — 180 chars of source]
remark{\normalfont The result in equation ((ref)) implies that $\sqrt{n}\Lambda _{n}^{-1}\hat{R}_{n}^{-1}(\hat{\theta}_{n}-\theta _{}^{0})$ behaves like a mean-zero Gaussian random variable, whose variance can be consistently estimated by $$[\Lambda_n\hat R_n^\prime\{\partial\bar{g}_n(\hat\theta_n)/\partial\theta'\}'S^{-1}_n(\hat\theta_n)\{\partial\bar{g}_n(\hat\theta_n)/\partial\theta'\}\hat R_n\Lambda_n]^{-1} .$$ {However, Theorem (ref) does not say that the common estimator of the variance-matrix of $\sqrt{n}(\hat\theta_n-\theta^0)$, obtained using the standard formula $$[\{\partial\bar{g}_n(\hat\theta_n)/\partial\theta'\}'S^{-1}_n(\hat\theta_n)\{\partial\bar{g}_n(\hat\theta_n)/\partial\theta'\}]^{-1},$$ is well-behaved, which follows by noting that the matrix $\frac{\partial\bar{g}_n(\theta^0)}{\partial\theta}S^{-1}_n(\theta^0)\frac{\partial\bar{g}_n(\theta^0)}{\partial\theta'}$ is asymptotically singular unless $\varsigma_n=O(1)$.} Fortunately, Theorem 5.1 in antoine2012efficient allows us to conclude that standard formulas for Wald inference based on the GMM estimator $ \hat{\theta}_{n}$ are asymptotically valid. The main intuition is that the Studentization implied by Wald inference cancels out the required rescaling terms. This is all the more important given that the rescaling factor $ \varsigma _{n}$ is unknown in practice. We stress that this result is in contrast to the general nonlinear case where the asymptotic normality requires faster than $n^{1/4}$ convergence rate, and it is only due to the specificities of the probit model that we are able to conduct valid Wald inference as soon as identification is not genuinely weak. That is, any near weakness, even as severe as $ \varsigma _{n}$ being arbitrarily close to $\sqrt{n}$, will still allow us to compute a consistent GMM estimator and apply standard formulas for Wald inference based on this estimator. }

The Power of the Distorted J-Test

The key to ensuring that the size of $W_{n}^{\delta}$ is asymptotically controlled is the equivalence between $J_n$, the usual $J$-statistic, and $J_n^\delta$, the distorted $J$-statistic, that obtains under the null of weak identification. However, as demonstrated by Proposition (ref) and Theorem (ref), under the alternative hypothesis the CUE is consistent and asymptotically normal. Therefore, there is no reason to suspect that $J_n$ and $J_{n}^{\delta}$ will be asymptotically equivalent under the alternative, at least under reasonable choices for the tuning parameter $\delta_n$.

The following result demonstrates that under the alternative, the distorted J-test, $W_n^\delta$, is a consistent test for the null of weak instruments across a wide range of choices for the perturbation sequence $\delta_n$.

theorem[Distorted J-test: Under the Alternative] If Assumptions 1-6 are satisfied, then $W_n^\delta$ is consistent under the alternative so long as $\{\sqrt{n}/\varsigma_{n}\}\delta_n\rightarrow\infty$ as $n\rightarrow\infty$.
remark\normalfont Theorem (ref) implies that our choice of $\delta_{n}$ has important consequences for the power of the distorted $J$-test. All else equal, the test is more powerful the slower $\delta_n$ goes to zero. However, it is also helpful to understand how fast $\delta_n$ can converge to zero before the result of Theorem (ref) is invalidated. To this end, consider the rate requirement on $\delta_n$ that results from parametrizing $\varsigma_{n}$ as $\varsigma_{n}=n^{\lambda}$ for some $0<\lambda<1/2$. Using this parametrization, we see that the distorted $J$-test is consistent so long as $\delta_nn^{1/2-\lambda}\rightarrow\infty$, and clarifies that if $\delta_n$ goes to zero too fast, i.e., if $\delta_n\ll n^{\lambda-1/2}$, the test can not be consistent.
remark\normalfont It is worth keeping in mind that Assumption 2 maintains that both the structural and reduced form moments are correctly specified. Thus, when the observed data lead to a rejection of $W_n^\delta$, we immediately conclude that it is not due to misspecification of the moment conditions but due to their identification power. However, if the model is misspecified, but we reject the null of weak identification, then we can actually consistently test for model misspecification. Indeed, under the alternative, the standard overidentification test $$\{J_n(\hat\theta_n,\hat\theta_n)>\chi^2_{1-\alpha}(H-p)\},$$ remains a consistent test for model misspecification. As such, if we reject the null of weak identification, we can compare the value of $J_n(\hat\theta_n,\hat\theta_n)$ against $\chi^2_{1-\alpha}(H-p)$ to deduce a consistent test for model misspecification.

Testing Procedure

We now explain one approach to implement our distorted J-test in practice. The key step in the testing procedure is to choose the perturbation (tuning parameter) $\delta_{n}$. To this end, we take $\delta_n=\delta/r_n$, and fix $r_n=\log \{\log(n)\}$. It is then possible to choose $\delta$ using a data-driven approach.

To present our approach to choosing $\delta$, first recall that the perturbation $\delta_n=\delta/\log\{\log(n)\}$ can be thought of as only being applied to the single direction of weakness in the rotated parameter space; namely, the parameter $\eta_1$, which by equation (ref) is nothing but $\tilde\rho$. Therefore, it is with respect to the magnitude of ${\hat{\tilde{\rho}}}_n$ that the perturbation $\delta_n$ should be chosen.

To ensure the value of $\delta_n$ is sufficiently close to the magnitude of $\hat{\tilde{\rho}}_n$, we design a grid of $m$ candidate points for $\delta$ by dissecting the standard confidence interval of $\hat{\tilde{\rho}}_n$ into $m$ equal regions. For the $i$-th region, we set $\delta_i$, $i=1,\dots,m$, to be to the midpoint of the $i$-th region. This produces a grid of $m$ perturbations with $i$-th value, $i=1,\dots,m$ given by $\delta_{i,n}=\delta_i/r_n$.

Whilst it is possible to use any given $\delta_{n,i}$ to conduct the test, we suggest carrying out the test across the entire grid of $\delta_{n,i}$ values and then appropriately modify the critical value via a Bonferroni correction. In particular, let $J_{n,i}^{\delta}$ denote the test statistic $J_{n}^\delta$ calculated under the perturbation $\delta_{n,i}$. This approach would lead us to reject the null of weak identification if $$ \max_{i\in\{1,\dots,m\}}J_{n,i}^\delta>\chi^2_{1-\alpha/m}(H+1-p). $$

Using the above decision rule, our approach can be implemented using the following four steps.

enumerate• Compute $\hat{\theta}_{n}=\operatorname*{argmin}_{\theta\in\Theta}J_n(\theta,\theta)$; • For a given choice of $m$, choose the sequence of tuning parameter $\delta_n=\delta/r_n$, as described above; • For each $i=1,\dots,m$, compute the test statistic $J_{n,i}^{\delta}$, as defined in Section (ref); • Rejection rule: reject if $ \max_{i\in\{1,\dots,m\}}J_{n,i}^\delta>\chi^2_{1-\alpha/m}(H+1-p). $

Under the null hypothesis, the testing procedure is size controlled for any choice of $\delta_{n,i}=o(1)$, while under the alternative the choice of $\delta_{n,i}$ only has implications for the power of the test. Moreover, since the values of $\delta_i$ are chosen from some compact set, dividing by $\log\{\log(n)\}$ ensures that $\delta_{n,i}=o(1)$ under both the null and alternative.

Generalizing the Rule-of-Thumb to Probit Models

{We begin our discussion on the so-called “rule-of-thumb", initially inspired by the work of staiger1997instrumental, in the infeasible situation where the latent endogenous variable $y_{1i}^{\ast }$\ is observable,} meaning that we would consider a bivariate linear model. {For sake of expositional simplicity, let us consider a simplification of this model whereby} the vector $x_{i}$ only contains a constant, so that the model becomes

eqnarray[eqnarray omitted — 132 chars of source]

The rule-of-thumb starts from the reduced form regression and its OLS estimator for $\xi$, \[ \hat{\xi}_{n}=(\tilde{Z}^{\prime }\tilde{Z})^{-1}\tilde{Z}^{\prime }\tilde{Y}_{2} , \] where for $\mathbf{1}_n$ a $(n\times 1)$-vector of ones

align*[align* omitted — 147 chars of source]

and where $\bar{y}_{2n}=\frac{1}{n}\sum_{i=1}^{n}y_{2i} $ and $\bar{Z}_{n}$\ denotes the $\left( n\times k_{z}\right) $ matrix whose $j^{th}$-column has all its entries equal to \[ \bar{z}_{j,n}=\frac{1}{n}\sum_{i=1}^{n}z_{ij} . \] Let $F_{n}$\ denote the F-test statistic for testing the null hypothesis that the vector $\xi $ of coefficients {for the variables} $z_{i}$ in the reduced form regression {are zero}. Under the assumption of conditional homoskedasticity {for the error term $v_{i}$,} the F-test statistic can be written as

equation*[equation* omitted — 172 chars of source]

with $\hat{\sigma}_{v,n}^{2}$ a consistent estimator of variance of $v_i$, $\sigma^2_v$. The rule-of-thumb amounts to conclude that instruments are {strong (i.e., consistent estimation is feasible)} if $F_{n}$ {exceeds a pre-specified threshold value, which differs from the standard critical value used to test the null hypothesis $\text{H}_{0}:\xi =0,$ and which has} been extensively documented by stock2005testing. The rationale for this rule can be understood from the drifting DGP considered in Remark 7. Under the alternative hypothesis to the null of weak identification, for $n$ large,

equation[equation omitted — 224 chars of source]

{Therefore, under the null of weak identification ($\varsigma _{n}= \sqrt{n}$), $F_{n}$ in equation ((ref)) has a finite limit, whilst under} the alternative ($\varsigma _{n}=o\left( \sqrt{n}\right) $) {the statistic $F_n$ diverges} to infinity with a slope defined by the squared norm of $\gamma ^{0}$ and a weighting matrix that is proportional to $\text{Var}(z_{i})/\text{Var}(v_{i})$. This sounds like a natural criterion to measure {instrument strength in the infeasible model (ref),} since the reduced form regression will lead to the control variable $v_{i}=y_{2i}-\pi -z_{i}^{\prime }\xi $\ and endogeneity in the structural equation will be controlled thanks to the two-stage residual inclusion (2SRI):

equation[equation omitted — 142 chars of source]

{Since identification of $\eta _{1}=\tilde{\rho}$ in equation ((ref)) depends on the variation of} \[ z_{i}^{\prime }\xi _{n}^{0}\sim \frac{z_{i}^{\prime }\gamma ^{0}}{\varsigma _{n}} , \] it may sound natural to assess the magnitude of $\gamma ^{0}$\ after normalization by the variance of $z_{i}$. As noted by 16324, “IVs can be weak and the F statistic small, either because $\gamma$ is close to zero or because the variability of $z_i$ is low relative to the variability of $v_i$.” However, the F-test statistic follows a Fisher distribution (and asymptotically a distribution $\chi ^{2}\left( k_{z}\right) /k_{z}$) under the null $\text{H}_{0}:\xi =0$\ only when {the reduced form error term $v_{i}$} is conditionally homoskedastic. When one is {concerned with the} presence of conditional heteroskedasticity in this equation (i.e., non-constant $\text{Var}[v_{i}\mid z_{i}] $), one may consider the heteroskedasticty corrected Fisher test statistic \[ F_{n}^{\ast }=\frac{n-k_{z}}{k_{z}}\left[ \hat{\xi}_{n}^{\prime }\hat{\Sigma} _{n}^{-1}\hat{\xi}_{n}\right], \] where $\hat{\Sigma}_{n}$ is a consistent estimator of the asymptotic variance of $\sqrt{n}( \hat{\xi}_{n}-\xi _{n}^{0}) $. While stock2005testing propose to extend the use of the rule-of-thumb by using instead $F_{n}^{\ast }$\ in case of conditional heteroskedasticity, several authors, including andrews2018valid and olea2013robust, have documented the disappointing performance of the heteroskedasticity corrected rule-of-thumb. One may help to clarify this issue by noting that, denoting $\tilde{z}_{i}$ to be the $i$-th column vector of the matrix $\tilde{Z}'$, {for $n$ large and for $\sigma _{v}^{2}(z_{i})=\text{Var}[v_{i}\mid z_{i}]$},

equation[equation omitted — 341 chars of source]

Equation ((ref)) is a straightforward extension of a result {provided by antoine2017testing, and makes explicit {how robustifying the} test statistic for heteroskedasticity modifies the rule-of-thumb.} This modification is arguably puzzling since what really matters{ for identification power, namely the residual inclusion of $v_{i}$ in the structural equation ((ref)), is not fully captured by $\sigma _{v}^{2}(z_{i})$.} More precisely, the conditional heteroskedasticity that intuitively matters in the structural equation {is instead} \[ \sigma _{u}^{2}(z_{i})=\text{Var}[u_{i}\left\vert z_{i}\right] =\tilde{\rho} ^{2}\text{Var}[v_{i}\left\vert z_{i}\right] +\text{Var}[\varepsilon _{i}\left\vert z_{i} \right] . \]

This intuition is confirmed by antoine2017testing who show that, when nesting the IV estimation procedure in a GMM framework, the distorted J-test leads to a decision rule {based on the following weighted norm of $\gamma ^{0}$}: \[ \frac{n}{\varsigma _{n}^{2}}\gamma ^{0\prime }\text{Var}\left( z_{i}\right) \left[ \mathbb{E}\left( \tilde{z}_{i}\tilde{z}_{i}^{\prime }\sigma _{u}^{2}(z_{i})\right) \right] ^{-1}\text{Var}\left( z_{i}\right) \gamma ^{0} . \] {In the context of the probit model,} where only the sign $y_{1i}$ of $y_{1i}^{\ast }$\ is observed, the 2SRI equation becomes \[ y_{1i}=\Phi \left[ \alpha y_{2i}+\beta +\tilde{\rho}\left( y_{2i}-\pi -z_{i}^{\prime }\xi \right) \right] +\varepsilon _{i}, \] for some error term $\varepsilon_i$, {and the conditional heteroskedasticity in the structural equation takes the form}

equation*[equation* omitted — 316 chars of source]

{One may then expect that any generalized rule-of-thumb for probit models must account not only for} this conditional heteroskedasticity but also the impact of the non-linearity in the structural equation. In the simple context of Remark 7, we may then expect that the {key element to obtain a decision rule about weak instruments in the probit model is the magnitude of the vector} \[ V^{0}\left( \eta \right) =\mathbb{E}_n\left[ \tilde{a}\left( y_{2i},z_{i}\right) \phi _{i}\left( \eta,\theta _{2}^{0}\right) z_{i}^{\prime }\right] \gamma ^{0}, \text{ where }\|\gamma^0\|>0. \]

More generally, since the alternative to weak identification, defined by Assumption 5, is tantamount to the non-nullity of the vector $V^{0}\left( \eta \right) $, the generalized rule-of-thumb should be based on a norm of $V^{0}\left( \eta \right) $. We argue that we do have a well-suited generalization for the standard rule-of-thumb when applying a decision rule that rejects the null of weak identification if the norm $\left\Vert U\right\Vert $, of a certain vector $U$, exceeds a specified threshold with the following definition for $U$.

enumerate$U=\sqrt{n}\text{Var}\left( z_{i}\right)^{1/2} /\sigma_v\xi^{0}$ for a linear model with conditional homoskedasticity (i.e. the standard rule-of-thumb); • $U=\sqrt{n}\left[ \mathbb{E}\left( \tilde{z}_{i}\tilde{z}_{i}^{\prime }\sigma _{u}^{2}(z_{i})\right) \right] ^{-1/2}\text{Var}\left( z_{i}\right) \xi^{0}$ for a linear model with conditional heteroskedasticity (i.e. the generalization of the standard rule-of-thumb proposed by antoine2017testing); • $U=\sqrt{n}S_{11,n}^{-1/2}\left( \theta ^{0}\right) \mathbb{E}_n\left[ \tilde{a} \left( y_{2i},z_{i}\right) \phi _{i}\left( \eta,\theta _{2}^{0}\right) z_{i}^{\prime }\right] \xi^{0}\delta_n$ for the probit model ((ref)) (in the context of Remark 7) and more generally $U=\sqrt{n}S_{11,n}^{-1/2}\left( \theta ^{0}\right) V^{0}\left( \eta \right)\delta_n/\varsigma_n$, where the perturbation term $\delta_n$ is introduced by the design of the distorted J-test.

It is worth realizing that this generalized rule-of-thumb is, for $n$ large, precisely what is performed by our test for the null of weak identification based on the distorted J-test statistic.\footnote{ It is worth realizing that our comparison between the so-called “rules-of-thumb" is based only on the definition of the test statistic. We do not enter into debates regarding alternative definitions of the null of weak identification based either on Assumption 3, or the 2SLS relative bias, Wald test size distortion, Nagar bias, etc..} To see this, we extend the argument of antoine2017testing by noting that under the alternative, our distorted J-test statistic sets the focus on the norm of \[ U=S_{n}^{-1/2}\left( \theta ^{0}\right) \sqrt{n}\bar{g}_{n}( \hat{\theta }_{n}^{\delta }) , \] where \[ \bar{g}_{n}( \hat{\theta}_{n}^{\delta }) =\bar{g}_{n}( \hat{ \theta}_{n}) +\left[

array[array omitted — 105 chars of source]

\right]. \] {Noting that,} \[ \sqrt{n}\left[ \bar{g}_{1n}( \hat{\theta}_{n}^{\delta }) -\bar{g} _{1n}( \hat{\theta}_{n}) \right] =\sqrt{n}\frac{\partial \bar{g} _{1n}}{\partial \eta _{1}}( \eta _{1n}^{\ast },\hat{\eta}_{2n},\hat{\eta} _{3n},\hat{\theta}_{2n}) \delta _{n}, \] where $\eta _{1n}^{\ast }$ denotes a component-by-component intermediate value between the first coefficient of $\hat{\theta}_{n}$\ and $\hat{\theta} _{n}^{\delta }$, under the alternative hypothesis to the null of weak identification

eqnarray*[eqnarray* omitted — 524 chars of source]

and where\footnote{The $O_p(1/\sqrt{n})$ term in the expansion of ${\partial \bar{g}_{1n}( \eta _{1n}^{\ast },\hat{\eta}_{2n},\hat{\eta}_{3n},\hat{\theta}_{2n}) }/{\partial \eta _{1}}$ can be deduced via a Taylor series expansion, re-arranging terms, and noting that the derivative of the Jacobian, in the $\eta_1$ direction, is also degenerate at the $\varsigma_{n}$-rate.} \[ \frac{1}{n}\mathbb{E}_{n}\left\{ \sum_{i=1}^{n}\tilde{a}\left( y_{2i},z_{i}\right) \phi _{i}\left( \eta ,\theta _{2}^{0}\right) z_{i}^{\prime }\xi ^{0}\right\} \sim \frac{V^{0}\left( \eta \right) }{\varsigma _{n}} \] is the dominant term since $\varsigma _{n}=o\left( \sqrt{n}\right) .$ To summarize, under the alternative hypothesis to the null of weak identification, and for a $\delta_n$ such that $\{\sqrt{n}/\varsigma_{n}\}\delta_n\rightarrow \infty$,

eqnarray*[eqnarray* omitted — 306 chars of source]

{which diverges as $n\rightarrow\infty$ and yields a natural generalization of the rule-of-thumb to probit models.}

Monte Carlo: Conventional Weak IV Tests v.s. Distorted $J$-test

In this section, we verify the properties of the distorted $J$-test (hereafter, DJ test) and compare this test against three commonly used weak IV tests, which, even though they are not designed for discrete choice models, have been widely applied in the literature on discrete choice modelling: (i) the staiger1997instrumental standard rule-of-thumb (SS); (ii) stock2005testing (SY); and (iii) the robust weak IV test of olea2013robust (Robust).

We generate observed data according to

align[align omitted — 98 chars of source]

where \(z_i\sim \mathcal{N}(0,\sigma_z^2)\) is i.i.d. univariate, \((u_i,v_i)'\) is i.i.d. homoskedastic and normally distributed, and \((u_i,v_i)'\) is independent of \(z_i\). We set \(\beta=0.5\), \(\alpha=1\) and \(\pi=0.3\). In addition, we take \(\rho=corr(u_i,v_i)\in\{0.5,~0.95\}\), and \(\sigma_u=1/\sqrt{1-\rho^2}\) (to ensure the normalization of \(\text{Var}[u_i|y_{2i},z_i]=1\)). To characterize the potential instrument weakness, we adjust the value of $\xi$ to restrict the correlation between the endogenous regressor $y_{2i}$ and the instrument $z_i$ to be $corr(y_{2i},z_i)=\gamma/n^\lambda$, with \(\gamma=1.5\), and we consider a grid of values for $\lambda\in\{0.5,~0.4,~0.3,~0.2,~0.1\}$.

Since the performance of the DJ test and the standard weak IV tests may depend on $\sigma_{z}$ and $\sigma_{v}$, we simulate data using the following grids: $\sigma_z\in\{0.2,~0.5,~1,~5,~10\}$ and $\sigma_v\in\{0.2,~0.5,~1,~5,~10\}$. For each Monte Carlo trial, we take the sample size to be one of \(n=500,~5000,~10000\) and consider $N=1000$ Monte Carlo replications.

Across each Monte Carlo design, \(\theta=(\widetilde{\rho},\alpha,\beta,\pi,\xi)'\) is estimated by CUGMM with a single degree of over-identification. We choose the instrument functions $a_i=a(y_{2i},z_i)=(1,y_{2i},z_i,z_i^2,0,0)'$ and $b_i=b(z_i)=(0,0,0,0,1,z_i)'$. The DJ test is implemented following the procedure presented in Section (ref).\footnote{For computational simplicity, in the Monte Carlo simulations, we adopt the perturbation $\delta_n=\hat{\tilde{\rho}}/\log(\log(n))$, where $\hat{\tilde{\rho}}$ is the CUGMM estimate of $\tilde{\rho}$ in each Monte Carlo replication. This procedures is a simplified version of the data-driven approach developed in Section (ref).} Using a 5$\%$ significant level, we reject the null hypothesis of weak instruments in accordance to Theorem (ref); i.e., we reject the null if $J^\delta_n>\chi^2_{0.95}(H+1-p)$, where in this case $H=6$, $p=5$ and $\chi^2_{0.95}(H+1-p)=5.99$. Theoretically, the hypotheses of the DJ test correspond to $H_0:\lambda=0.5$, and the alternative to $\lambda<0.5$.\footnote{We note that the null hypothesis of each test are slightly different: DJ- $\text{H}_{0}:\lambda=0.5$; SS- $F_n<10$ as an informal null hypothesis; SY- the triple $\{\xi,\sigma_v^2,\sigma_z^2\}$ is such that 2SLS relative bias or Wald test size distortion is larger than a given tolerance using the Cragg-Donald statistic; the Robust test regards that the Nagar bias exceeds a fraction of the benchmark as null. Although the definitions of the weak instrument are different for each test, their null hypothesis are consistent in the sense to capture situations under which the instrument is weak.} However, we note that in finite samples, it is hardly the case that $\lambda$ alone determines the behavior of the CUEs.

Given this, to compare the behavior of the DJ test with the conventional linear tests, we introduce two sets of criteria to assess the potential impact of instrument weakness in finite samples: the behavior of the CUE and the size distortions of the associated Wald statistic. Specifically, we compute the bias, standard deviation (s.d.) and relative root mean square error (rrmse) as below (taking $\alpha$ as an example) to measure the estimation performance under different designs:

align[align omitted — 281 chars of source]

where $\bar{\widehat{\alpha}}=1/N\sum_{l=1}^N\widehat{\alpha}_{l}$, $\widehat{\alpha}_{l}$ stands for the $l$-th Monte Carlo CUGMM estimate and $\alpha^0$ is the true value. As proven in Sections (ref) and (ref), under the null the CUE is consistent, while under the alternative, the estimator will be consistent and asymptotically normal, albeit with non-standard rates. Unlike stock2005testing, who choose the relative bias of 2SLS to OLS as one criterion to detect weak instruments, here we consider the bias, the s.d. and the rrmse defined in (ref) instead, for the following reasons. For the IV probit model (ref), the CUE (and other commonly adopted estimation methods) does not have a closed-from expression. Therefore, the usual notion of `bias towards OLS' under potential IV weakness in linear models is not valid in this nonlinear context, with the potential impact of the IV weakness now being complicated by the nonlinear features of the model. In this case, there is no guarantee that the positive and negative biases will not offset each other and lead to a spuriously small overall bias. Therefore, to capture the instrument strength and the resulting performance of the CUGMM estimation procedure, we rely on the bias, standard deviation and rrmse of the estimator.

In addition, to better understand weakness in this discrete choice model, we conduct a Wald test of $\text{H}_0:\alpha=\alpha^0$ and compute its size distortion, relative to the $5\%$ significant level, across all the Monte Carlo designs. We carry out this Wald test for two different estimation methods: the CUE considered in this paper and the 2SCML estimator proposed by rivers1988limited. The size distortion of the Wald test is widely used to capture instrument weakness; see e.g. staiger1997instrumental and stock2005testing. This measure reflects not only the performance of the hypothesis test, but also the coverage rate of confidence intervals associated with the two estimation methods.

Under the null hypothesis of $\lambda=0.5$, the performance of the CUE and the rejection probabilities for the different testing procedures are collected in Table (ref) ($\rho=0.5$) and Table (ref) ($\rho=0.95$). For brevity, we only report the estimation results for the structural parameter of interest, $\alpha$, and Wald test size distortions under five designs: $(\sigma_z,\sigma_v)\in\{(1,0.2),(1,10),(1,1),(0.2,1),(10,1)\}$. Additional results for all designs can be obtained from the authors.

Simulation results in Tables (ref) and (ref) confirm our asymptotic results. When $\lambda=0.5$, CUGMM estimation of $\alpha^0$ is inconsistent and behaves poorly in general. More specifically, the biases are unstable, and the s.d. and rrmse do not decrease (in any noticeable way) as the sample size increases, especially when the endogeneity degree is high ($\rho=0.95$). However, under the alternative, $\lambda<0.5$, the s.d. and rrmse drop dramatically as $n$ increases. In addition, the asymptotic normality of the CUE under $\lambda<0.5$ is verified by viewing the standardized sampling distributions of the estimators across the Monte Carlo replications, which is given in Figures (ref) and (ref). The sampling distributions exhibit easily detectable bi-modality when $\lambda$ is 0.5, or close to 0.5, especially when $\sigma_v$ is small and $\rho$ is large, indicating that a standard inference approach, relying on the normal approximation, is likely to perform poorly in those cases.

The results in Tables (ref) and (ref) also show that the behavior of the Wald test varies across the different designs even when $\lambda=0.5$. For a moderate level of endogeneity ($\rho=0.5$), we see relative small size distortions, less than $5\%$, in most cases for the Wald tests based on both 2SCML and CUEs. However, for a high degree of endogeneity ($\rho=0.95$), the Wald tests are significantly over-sized, with the size distortions for the Wald test based on 2SCML being much larger than those based on CUGMM. One exception, however, is the case of $(\sigma_z,\sigma_v)=(1,10)$ and $\rho=0.95$, where the Wald size distortions based on both estimation methods are less than $5\%$. {For $(\sigma_z,\sigma_v)=(1,10)$ and $\rho=0.95$ case, the size distortion based on the CUGMM is 0.008 when $n=10000$, indicating that the $95\%$ confidence interval coverage rate is quite accurate even though $\lambda=0.5$ ($corr(y_{2i},z_i)=0.015$). As such, this design constitutes additional evidence that the value of $\lambda$ is not the only key in determining inference performance in weakly identified discrete choice models.}

The false rejection rates of SS, SY, Robust and DJ under $\lambda=0.5$ are displayed in Tables (ref) and (ref). Firstly, as expected, the DJ test is asymptotically conservative, i.e., the size is less than the significance level of $5\%$.\footnote{Results not reported here due to space limitations also confirm that DJ test is conservative.} The size of the DJ test varies between 1.0% and 1.9% under $\rho=0.5$, and between 1.3% and 3.1% when $\rho=0.95$. However, we note that the DJ test is much less conservative than the Robust approach of olea2013robust, which is extremely conservative, and gives virtually zero rejections across all designs where identification is weak. Therefore, while the DJ test is conservative, we can conclude that it is much less conservative than the Robust approach, and can be relatively close to the nominal level (5%) when the degree of endogeneity is large.

In addition, we see that blindly applying conventional weak instrument tests can lead to poor outcomes. For example, for the design with $(\sigma_z,\sigma_v)=(1,0.2)$ and a high level of endogeneity ($\rho=0.95$ in Table (ref)), the rejection rates of SS and SY (10%)\footnote{The rejection rate of SY (10%) is computed based on the critical value of a maximal 10$\%$ size distortion of a 5$\%$ Wald test, provided by stock2005testing.} are all larger than 10$\%$ across different sample sizes, and are 13.8$\%$ and 18.5$\%$ respectively when $n=10000$. However, the rrmse in this case does not decrease as $n$ increases, and the rrmese for the estimated $\alpha$ is between 910% and 1060$\%$ of the true value. Moreover, both of the Wald size distortions exceed their nominal size by at least 10$\%$. In particular, the 2SCML size distortion is between 17$\%$ and 27$\%$, while the CUGMM size distortion is between 11$\%$ to 17$\%$. Therefore, the identification is weak, but the SS and SY approaches can suggest the opposite, and hence fail to control size. In addition, false rejection rates for other designs, not reported here for brevity, demonstrate a similar pattern of over-rejection for SS and SY tests. Hence, in line with the analysis in Section (ref), when assessing identification strength in discrete choice models, the conventional weak IV tests of SS and SY may fail to provide reliable conclusions regarding identification strength, especially if the degree of endogeneity is high.

Figure (ref) ($\rho=0.5$) and Figure (ref) ($\rho=0.95$) display the power of the four tests. Due to the conservativeness of DJ test, size adjusted power of DJ and of the conventional tests are also computed and compared in Figures (ref) and (ref).\footnote{Size adjusted power is computed as follows: obtain the 95$\%$ quantile of the test statistic from the 1000 Monte Carlo replications when $\lambda=0.5$ and use it as the critical value for cases when $\lambda<0.5$.} The resulting power curves show that the DJ test is consistent as the sample size diverges, and as identification strength increases. Moreover, in cases with high endogeneity (Figure (ref)), the unadjusted power of the DJ is higher than that of the Robust test across most designs. Furthermore, Figures (ref) and (ref) demonstrate that the DJ-test displays non-negligible power even when identification is close to being weak, i.e, when $\lambda=0.4$ or $\lambda=0.3$, which gives convincing numerical evidence of the results in Theorem (ref).

Empirical Illustrations

In this section, we apply our distorted $J$-test in two well-known empirical examples to test for the presence of weak instruments. We then contrast the results of our tests with those obtained from conventional weak IV tests for linear models, namely the SS, the SY, and the Robust tests.

Labor Force Participation of Married Women

We first study the impact of education on married women's labor force participation (hereafter LFP), when education, measured as the women's years of schooling, is treated as an endogenous treatment. We use data from the University of Michigan Panel Study of Income Dynamics (PSID) for the year 1975\footnote{The data is publicly available at wooldridge2010econometric Supplemental Content.}, which have been used in several studies. mroz1987sensitivity provides an extensive analysis of the women's hours of labor supply, and considers a range of specifications including potential endogeneity of several regressors, the use of different instrumental variables and controls for self-selection into labor force participation. As a text book example, wooldridge2010econometric used the same dataset to study women's LFP decisions, and the potential endogeneity of education is tested after estimating an IV probit model using rivers1988limited two-step conditional maximum likelihood estimator (2SCML). In what follows we use similar specification as in wooldridge2010econometric.

The PSID consists of data on 753 married, Caucasian women who are between 30 and 60 years of age at the time the sampling. The dependent variable LFP is a binary response that equals unity if the respondent worked at some time during the year, and zero otherwise. Exogenous regressors include spousal income, the individual's work experience and its square, age, the number of children less than six years old, and the number of children older than six years old. The individual's education, measured as years of schooling, is considered to be endogenous. Following the strategy in wooldridge2010econometric, the individual's family education, which are recorded as the years of schooling for both the individual's father and mother, are used as instruments for education.

Estimated coefficients and the average partial effects on the probability of LFP for all regressors are presented in Table (ref) using two estimation methods: 2SCML as used in wooldridge2010econometric\footnote{For the LFP example, the 2SCML estimation allows for heteroskedastic standard errors.} and CUGMM. More specifically, for the 2SCML, the first step is to regress the endogenous regressor on the instruments and all other exogenous regressors to obtain the reduced form residual. The second step is to run a probit maximum likelihood estimation of the binary response on the endogenous and the exogenous regressors, and the reduced form residual. The CUGMM estimation with over-identification degree one is conducted using $a_i=(1,y_{2i},x_i',z_{i}',\textbf{0}_{k+2}')'$ and $b_i=(\textbf{0}_{k+3}',1,x_i',z_i')'$, where $y_{2i}$, $x_i$ and $z_{i}$ denote the standardized variables corresponding to the women's education, exogenous regressors and two instruments, and $k$ is the number of exogenous regressors and the intercept. The first step estimation of the 2SCML and the reduced form of the CUGMM are listed in the first and fourth columns of Table (ref) respectively. Both the two IVs are highly significant based on both estimation methods. The CUGMM estimation results are reported in columns four through six. Broadly speaking, the CUGMM and 2SCML results are similar, with both methods providing evidence that education has a significant positive effect: one extra year of education increasing the probability of LFP by 5.87 percentage points for both the 2SCML and the CUGMM. Hansen's $J$-statistic is 0.122 which is less than $\chi^2_{0.95}(1)=3.84$, therefore we fail to reject the null that all the moments are valid.

The weak IV test results are collected in Table (ref) for all four tests, SS, SY, Robust and DJ. The Kleibergen-Paap $F$-statistic kleibergen2006generalized is 81.89, based on which the SS rule-of-thumb and the SY test both reject the null that IVs are weak.\footnote{The Kleibergen-Paap $F$-statistic is utilized when allowing for heteroskedastic standard error. The reduced form regression $F$-statistic and the Cragg-Donald statistic are 95.70, when assuming homoskedastic standard error. SY rejects its null according to the critical values of the maximal desired size distortion $5\%$ and 10% of a 5% Wald test.} For the Robust test, the effective $F$-statistic is 91.44, the critical values for the tolerance thresholds $\{5\%,10\%\}$ are 11.59 and 8.58, respectively.\footnote{The estimated effective degrees of freedom of the Robust test for the tolerance thresholds $\{5\%,10\%\}$ are both about 1.8. See olea2013robust for the definitions of the effective $F$-statistic, the tolerance threshold and the effective degrees of freedom. The Robust test statistic and the critical values are obtained using the Stata command "weakivtest" (pflueger2014robust) under heteroskedastic-robust estimation.} Comparing the effective $F$-statistic 91.44 to the critical values, the Robust test also rejects the null of weak IV.

Finally, for the DJ test, the perturbation $\delta_n$ is computed as in Section (ref), using $m=20$ candidate grid points. This choice of $m$ leads us to use the critical value $\chi^2_{1-0.05/20}(H+1-p)=11.98$, where we note that we have used $H=19$ moments and estimated $p=18$ parameters. Of the candidate grid points, three lead to a value of the DJ statistic larger than 11.98, leading us to soundly reject the null of weak identification. The rejection conclusion of the DJ test is quite straightforward: when perturbing the CUE $\hat{\theta}_n$ by $\delta_n$, the value of the $J$-statistic increases dramatically from 0.122 to a maximum of 17.44, implying that the CUGMM criterion is sensitive to even small departures. Overall, results reported in Table (ref) suggest that the DJ test and the three conventional tests for linear models all agree in this example.

US Food Aid and Civil Conflicts

In the second example we examine the impact of US food aid on the incidence of civil conflicts in recipient countries. The research in nunn2014us was motivated by concerns that humanitarian food aid may be ineffective and may even promote civil conflicts. The main challenge of this study is the potential endogeneity of US food aid due to reverse causality and joint determination. Their identification strategy relies on using the product of the lagged US wheat production and the average probability of receiving any US food aid for each country as the instrumental variable for wheat aid. nunn2014us estimate many variations of binary and duration models and consider different kinds of wars, different controls and alternative specifications.

Herein, we focus on the cases of onset and offset of civil wars as considered in nunn2014us. More specifically, we estimate the impact of US wheat aid on the probability of civil war onset after a period of peace, or on the probability of civil war offset after a period of war (columns (3)-(9), Table 7, nunn2014us), using precisely the same datasets and model controls as in nunn2014us.\footnote{Data sets used to construct the incidence of conflict, US food aid, US wheat production and other variables include the UCDP/PRIO Armed Conflict Dataset Version 4-2010, the Food and Agriculture Organization's FAOSTAT database and the data from the United States Department of Agriculture. See nunn2014us for more detailed information.} We examine the IV strength by applying our DJ test to the model, as well as the three conventional weak IV tests for linear models. The dataset in this analysis involves observations on 78 non-OECD countries from 1971 to 2006.

For the onset analysis, the data used are those country-year observations that have no intra-state civil conflict in the previous period (columns (3)-(6), Table 7, nunn2014us). The event indicator for civil war onset {is defined as one} if it is the first period of a intra-state conflict episode, and zero otherwise. nunn2014us estimate a logistic discrete time hazard model for the onset of war, controlling for the previous duration of peace {using a third degree polynomial}. The US wheat aid in year $t$ is instrumented by the product of US wheat production in year $t-1$ and the probability of receiving any US food aid between 1971 and 2006 for each country. For the purpose of the paper, we estimate a binary probit model for the onset of war. The summary statistics for the data used in the onset analysis are given in part (a) of Table (ref).

Using the specification of controls in columns (3) of Table 7 in nunn2014us, in Part (a) of Table (ref), we present the estimated coefficients and average partial effects from both 2SCML probit\footnote{The 2SCML in this example allows intragroup correlation for standard errors, clustered by countries.} and CUGMM with the degree of overidentification equal to unity. For comparison purposes, column (1) of Table (ref)-(a) gives the estimated average partial effect of US wheat aid on the onset of war as reported by nunn2014us using a 2SCML logit approach. For CUGMM, we use $a_i=(x_i',z_i,z_i^2,z_i^3,z_ix_{1i},\textbf{0}_{k_{x}+1}')'$ and $b_i=(\textbf{0}_{k_{x}+3}',1,z_i,x_i')'$ to construct moments. The variables $x_i$ and $z_i$ denote the standardized variables of exogenous regressors and the instrument, $x_{1i}$ is the non-standardized onset duration, and $k_{x}$ is the number of exogenous regressors {(including an intercept)}. Columns (2) and (5) of Table (ref)-(a) demonstrate that the IV is significantly related to the endogenous regressor of wheat aid at the 1% significant level by both estimation methods. However, the estimates of interest, the treatment effects of the US wheat aid on onset are statistically insignificant from both estimation methods, same as the result from nunn2014us in Column (1). \footnote{The insignificance of US food aid on onset of civil conflict is also pointed out by nunn2014us. However, without reliable evidence on instrument strength, we should be cautious when drawing any conclusions based on standard inference procedures.} Estimates for other coefficients are quite stable and similar across the three sets of results. Finally, Hansen's $J$-statistic is 0.553, less than the critical value $\chi^2_{0.95}(1)=3.84$, thus we cannot reject the null that moments are all valid. This evidence leads to the suspicion that the potential weakness of the IV could be one of the possible reasons for the unstable estimates of the US wheat aid coefficient.

This suspicion is verified by the DJ test. The perturbation for the onset analysis is chosen as described in Section (ref), again using $m=20$ candidate grid points. Panel (a) of Table (ref) demonstrates that the DJ test cannot reject the null of weak identification. In contrast to the earlier example in Section (ref), in this example perturbing the estimators by $\delta_n$ does not lead to a significant change in the corresponding J-statistic, which indicates a lack of curvature and thus identification weakness. Across the entire grid of candidate $\delta_n$ values, the maximum of the DJ statistics is 7.5, which is less than the corresponding 5% critical value given by $\chi^2_{1-0.05/20}(H+1-p)=11.98$, and which is based on using $H=12$ moments to estimate $p=11$ parameters. However, when we apply the conventional SS, SY and Robust tests to the onset of the civil conflict case, the SS and SY tests all return a rejection of the weak IV hypothesis and the Robust test also rejects the null if the tolerance threshold is greater than $10\%$. As shown in Table (ref)-(a), the reduced form regression Kleibergen-Paap $F$-statistic for SS and SY is 26.07, much larger than 10 and the critical values 16.38 and 8.96 of SY.\footnote{To be consistent with nunn2014us, standard errors (s.e.) are computed using clustered s.e. by countries. Kleibergen-Paap $F$-statistic kleibergen2006generalized is utilized when allowing for intragroup correlation s.e. The critical values 16.38 and 8.96 of SY test are based on the desired maximal size distortion 5% and 10$\%$ of a 5$\%$ Wald test, respectively.} The Robust test effective $F$-statistic 26.39 is also larger than its 10$\%$ tolerance critical value 23.11.\footnote{The effective $F$-statistic and critical values are computed using the Stata command "weakivtest" (pflueger2014robust). The critical value 23.11 is for the case of effective degrees of freedom one and the tolerance threshold $10\%$. Robust test fails to reject the weak instrument based on the critical value of 5% tolerance.} In summary, for this onset example, the conventional weak IV tests reject the weak IV hypothesis, while our DJ test suggests the opposite. This serves as a reminder that applying conventional weak IV tests for linear models to binary outcome models should be cautioned.

Subsequently, we have repeated this analysis for the other 5 specifications considered in nunn2014us (columns (4)-(8) of Table 7, nunn2014us), which include different controls and the study for conflict offset after a period of war. Our DJ test fails to reject the null of weak instrument for all cases except for column (7), whilst the SS and SY tests all result in a rejection of the weak instrument hypothesis.\footnote{Not all results are reported due to space limitation. SS and SY tests are based on Kleibergen-Paap $F$-statistic}. The DJ test is implemented using the same $a_i$ and $b_i$ as those used in Table (ref). The perturbation is again chosen as in Section (ref) with $m=20$.} The Robust test also rejects the null in some cases.\footnote{Based on the critical value 23.11 ($\tau=10\%$), the Robust test rejects weak IV of the analysis in columns (4) and (8), but fails to reject in columns (5), (6) and (7). Results are obtained by using the Stata command "weakivtest" pflueger2014robust and clustered s.e.} In part (b) of Table (ref) and Table (ref), we report the estimation results for the probability of offset of civil war after a period of war for the specification in column (6) of Table 7, nunn2014us, as well as the test results for weak instrument.

One important result to note is that nunn2014us estimate a significant and negative effect for offset of war, indicating that aid prolongs civil wars with 1,000 MT extra of US wheat aid reducing the probability of civil war offset by 0.04 percentage point. The causal effects estimated by the Probit model in Panel (b) of Table (ref) are also both negative, with statistical significance for the 2SCML results but not for the CU-GMM results. However, as shown in part (b) of Table (ref), if one applies the DJ test using the same methodology as above, the DJ statistic varies between 1.50 and 9.46, which is again less than the corresponding critical value of 11.98, indicating that identification may be weak in this example. If identification is indeed weak, as the DJ test suggests, conducting standard inference on the estimated treatment effect is no longer valid. Therefore, the conclusion that US food aid prolongs civil conflict should be viewed with caution.

Conclusion

Estimating the causal effects of policy relevant treatment variables is the key goal of many empirical studies in economics and other diverse fields. Instrumental variables play a crucial role in the identification and estimation of treatment effects when the treatment is endogenous, but weak instruments have been identified as a potentially serious problem, with consequences including inconsistent estimation {and, consequently, invalid statistical inference.} Consequences and detection of weak identification due to instrument weakness have been extensively studied for linear models, but similar issues have not been thoroughly studied for discrete outcome models. In search for a suitable weak identification test, empirical researchers have often resorted to the inappropriate use of linear model weak IV tests for discrete outcome models, or the use of a linear probability model with a 2SLS estimator treating the discrete outcomes as continuous. The suitability of these linear tests in this nonlinear setting is not usually questioned in many empirical studies (see dufour2013weak and li2019bivariate for additional analysis on the performance of the stock2005testing testing approach in binary models).

This paper proposes a much needed weak identification test in endogenous discrete choice models. The proposed test has desirable asymptotic properties including size control under the null of weak identification, and consistency under the alternative. Moreover, we demonstrate that once the null of weak identification is rejected, standard Wald-based inference can be applied as usual. Our Monte Carlo results demonstrate that, whilst the conventional stock2005testing and staiger1997instrumental tests are often over-sized, and thus fail to reliably detect weakness, our test always controls size and has reasonable power. We apply this testing approach to two empirical examples in the literature, and demonstrate that there are importance instances where our approach produces contradictory conclusions to the commonly applied linear testing approaches. Analyzing the causal impact of U.S. food aid on civil conflict, our approach fails to reject the null of weak identification, however, several commonly applied linear testing approaches all conclude that identification is not weak.

Another key contribution of the paper is the construction of comprehensive concept of weak identification in discrete choice models, based not only on the convergence rate of drifting moments, but also on the respective weight of the key parameters, including variances of error terms and the level of simultaneity. This allows us to provide a unified GMM estimation framework for examining both linear and nonlinear models, and for comparing the asymptotic properties of GMM estimators against other conventional two-step estimators for endogenous discrete choice models. While building on the general testing strategy of antoine2017testing, the test proposed in this paper is based on a null hypothesis of genuine identification weakness, and not the nearly-strong identification null hypothesis analyzed in, e.g., andrews2012estimation and antoine2017testing.

The conclusion {our research gives} to empirical researchers wishing to evaluate identification weakness in discrete choice models {is clear: the} canonical tests developed for linear models are not suitable for nonlinear models, are likely to be overly optimistic, and can fail to detect genuinely weak identification. Our recommendation is a two-step approach. Conduct our testing approach in a first-step, then, if the null is rejected, one can be very confident that identification is not weak, and conventional inference can proceed as usual. If the null of weak identification cannot be rejected, identification robust inference methods (as proposed in stock2000gmm or magnusson2010inference) would be more suitable to assert the significance of any estimated causal effects.

Furthermore, our asymptotic theory is conformable with the point of view on weak identification defended by 16324: “weak instruments should not be thought of as merely a small-sample problem, and the difficulties associated with weak instruments can arise even if the sample size is very large.” We do see weak identification as a population problem (i.e. independent of the sample size): either the GMM estimator is not consistent (under the null of weak identification) or it is consistent (under the alternative). In this respect, the device of using a drifting DGP, as contemplated in the weak identification literature, can be seen as a way to disentangle point identification (a maintained hypothesis in the framework of weak identification) and existence of a consistent estimator. This point of view may look at odds with the one put forward by lewbel2019identification where it is stated that: “a parameter that is weakly identified (meaning that standard asymptotics provide a poor finite sample approximation to the actual distribution of the estimator) when $n$ = 100 may be strongly identified when $n$ = 1000.” However, for all practical purpose, the methodological recommendation may not be so different: in our case, it is only for a large enough sample size that our test may allow us to reject the null of weak identification. In these circumstances, the researcher can trust the consistency of the estimator and confidently use {Wald-based} inference.

Acknowledgements

Frazier was supported by the Australian Research Council's Discovery Early Career Researcher Award funding scheme (DE200101070), and the Australian Centre for Excellence in Mathematics and Statistics. Zhang acknowledges travel support provided by the International Association for Applied Econometrics (IAAE). The authors thank the seminar participants at University of Cambridge, CEMFI, University of Oxford, University College of London, University of Surrey, University of Warwick, Boston University, University of Sydney, the 10th French Econometrics Conference PSE, the 2019 IAAE conference, and the 2019 North American Summer Meeting of the Econometric society for helpful comments.