Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
95,579 characters · 17 sections · 90 citation commands
Detecting Identification Failure in Moment Condition Models
JEL Classification: C11, C12, C13, C32, C36.\newline Keywords: Asset Pricing, Uniform Inference, Global Identification, Indirect Inference.
\baselineskip=18.0pt \thispagestyle{empty} \setcounter{page}{0}
The Generalized Method of Moments (GMM) of Hansen1982 is a powerful estimation framework which does not require the model to be fully specified parametrically. Under regularity conditions, the estimates are consistent and asymptotically Gaussian. In particular, the moments should uniquely identify the finite-dimensional parameters. This is very difficult to verify in practice and, as noted in Newey1994a, is often assumed. Yet, when identification fails or nearly fails, the Central Limit Theorem provides a poor finite sample approximation for the distribution of the estimates. This has motivated a vast amount of research on tests which are robust to identification failure. An empirically relevant problem, which remains less explored, is of determining, for a given set of estimating moments, whether local and global identification actually hold.
The contribution of this paper is two-fold: first, it introduces a quasi-Jacobian matrix which is singular under both local (first-order) and global identification failure and is informative about the coefficients involved in the identification failure. This is the main contribution of the paper as it provides an approach similar to Cragg1993 and Stock2005 but in a non-linear setting. Second, the information is used to construct an identification robust subvector test which does not require a priori knowledge of the identification structure. The test is asymptotically non-conservative under strong identification. It is asymptotically efficient for strongly just-identified models.
The quasi-Jacobian matrix is the best linear approximation of the sample moment function over a region of the parameters where these moments are close to zero. To find the best linear approximation, a sup-norm (or $\ell_\infty$-norm) loss is used to minimize the largest deviation from the linear approximation. This is known as a Chebyshev approximation problem which can be solved fairly quickly using convex optimization software. In the population, the quasi-Jacobian has full rank if, and only if, the parameters are both globally and locally identified. When either global or local identification fails, it is singular in all directions associated with the identification failure. (Non)-singularity of the quasi-Jacobian can be used to check whether identification holds numerically when it is not feasible analytically.
The asymptotic behaviour of the quasi-Jacobian matrix is studied under three identification regimes: including strong, semi-strong, and weak (or set) identification. Under strong identification, the moment conditions are informative, have a unique solution, under semi-strong identification are less informative but sufficiently so that for estimates to be consistent and asymptotically Gaussian. Antoine2009, Andrews2012 showed that: under (semi)-strong identification, standard inference methods such as the t-test with standard normal critical values are asymptotically valid.\footnote{The term (semi)-strong will refer to cases where identification can be either strong or semi-strong. Antoine2009 further distinguish between nearly-strong and nearly-weak identification. Under the latter, the limiting distribution may be non-Gaussian. Here, when this is the case, it will be referred to as higher-order local identification.} Under weak and set identification, the moments are insufficiently informative compared to sampling uncertainty and multiple distant solutions to the moment conditions appear plausible, even in large samples, so that the parameters cannot be consistently estimated and standard inference methods are not asymptotically valid. The Supplement also considers higher-order local identification, where the solution is unique but not locally identified; it can be consistently estimated but with non-Gaussian limiting distribution. Under (semi)-strong identification, the quasi-Jacobian is shown to be asymptotically equivalent to the usual Jacobian: after re-scaling, it is asymptotically non-singular. Under higher-order and weak identification the quasi-Jacobian is asymptotically singular with eigenvalues vanishing in directions where identification fails. It is thus informative about the presence of identification failures and which directions are not identified.
Building on these results, this paper constructs a simple test procedure for subvector hypotheses on the parameters $\theta=(\theta_1^\prime,\theta_2^\prime)^\prime \in \mathbb{R}^{d_\theta}$ of the form:
Subvector inference as described in ((ref)) is quite prevalent in empirical work where only a few structural parameters $\theta_1$ are typically of interest. The remaining $\theta_2$ nuisance parameters describe other features of the data generating process needed for estimation. For instance, in the empirical application only $2$ preference parameters are of interest while the remaining $10$ coefficients parameterize the law of motion for consumption and dividends which is not of immediate interest. The paper relies on the Anderson1949 test statistic for simplicity. The critical values take the form $\chi^2_{d_g-d}$ where $d_g$ is the number of moments and $d$ is determined using an Identification Category Selection (ICS) procedure based on the singular values of the quasi-Jacobian matrix. This is a projection inference procedure where the ICS step estimates the number of (semi)-strongly identified nuisance parameters to reduce the degrees of freedom.
Monte-Carlo simulations illustrate the results for a simple consumption-based asset pricing model. In the empirical application, the procedure is used to conduct joint inference on risk-aversion and the inverse elasticity of substitution in the long-run risks model of bansal2004. The results suggest that several nuisance parameters are weakly identified but not all; some are (semi)-strongly identified. This implies that standard inferences based on t or Wald statistics are not asymptotically valid and full projection inference is valid, but conservative. Given the number of parameters in the application, the standard approach of performing test inversion using a grid search is very computationally demanding. Instead, an adaptive sampling procedure based on the Population Monte Carlo (PMC) principle draws uniformly on level sets of the objective function. This makes it possible to conduct robust inference on more complex models like the empirical application: the quasi-Jacobian and 5,000 uniform draws on the confidence set are computed in about 4 hours on a desktop computer.
After a review of the literature and an overview of the notation, Section (ref) introduces the setting, the procedure and provides more details about the quasi-Jacobian, the test, and the identification regimes. Section (ref) derives the asymptotic behaviour of the quasi-Jacobian matrix, and Section (ref) results for the test. Section (ref) gives Monte-Carlo evidence for the results, and Section (ref) the empirical application. Appendices (ref), (ref) provide proofs for the main results. The Supplement includes sample R code to compute the quasi-Jacobian and for inference, a description of the PMC algorithm used to generate draws, and additional results for higher-order identification.
The literature on the identification of economic models is quite vast, and an extensive review is given in Lewbel2018. Within this literature, this paper mainly relates to three topics: local and global identification of finite-dimensional parameters in the population, detection of identification failure in finite samples, and identification robust inference.
Koopmans1950 provide one of the earliest general formulations of the identification problem at the population level. To paraphrase the authors, the main problem is to determine whether the distribution of the data, assumed to be generated from a given class of models, is consistent with a unique set of structural parameters. In the likelihood setting, Fisher1967, Rothenberg1971 introduced sufficient conditions for local and global identification. Komunjer2012 provides weaker global identification conditions for GMM.
In linear models, global identification amounts to a rank condition on the slope of the moments. This insight was used in pre-testing linear IV models for identification failure using a first-stage F-statistic or rank tests, Cragg1993, Stock2005, kleibergen2006. Pre-tests based on the null of strong identification appear in Hahn2002 for linear IV and Inoue2011, Bravo2012 for non-linear models. Pre-testing for strong identification could make size control difficult when the pre-test has low power. For non-linear models, Wright2003 uses a rank test and antoine2020 a distorted J-statistic to detect local identification failure. Arellano2012 develop a test for underidentification of a single coefficient.
Given the impact of (near) identification failure on standard inferences, a large body of literature has developed identification robust tests. Much of the literature is concerned with inference on the full parameter vector, e.g. Anderson1949, Stock2000, Kleibergen2005, Andrews2016. Projection inference can be used to conduct subvector inference from these tests dufour1997. Alternatively, Bonferroni methods combined with a $C(\alpha)$ test can be used, Chaudhuri2011, andrews2017. For homoskedastic linear IV models, guggenberger2012a propose critical values for a subset Anderson-Rubin test which improve power over full projection inference. In the same setting, guggenberger2019 propose a data-driven choice of critical values based on a measure of identification strength of the nuisance parameters, and kleibergen2021 considers subvector conditional Likelihood-Ratio inference. This paper relies on the Anderson-Rubin statistic for inference, which is the simplest to implement. More powerful test statistics exist such as the conditional quasi-Likelihood Ratio. The main challenge there is in computing the critical values by simulation, which requires to repeatedly minimize non-linear and potentially multi-modal objective functions.\footnote{This is difficult for non-convex problems, see e.g. Nemirovsky1983 for the complexity of the minimization problem and Nesterov2018 for the practical implications and software limitations.}
Given knowledge about the source of a potential identification failure, and a specific structure in the underlying model Andrews2012,Andrews2013,Andrews2014, Cheng2015, han2019, Cox2020 propose identification robust tests which are asymptotically non-conservative and powerful under strong identification. These papers rely on a data-driven choice of critical value; it is determined by an ICS statistic built from model-specific knowledge about the source and form of the identification failure. This paper proposes and studies an ICS statistic which does not rely on model-specific information to determine identification status. The choice of robust critical values can coincide with Andrews2012's least-favorable critical value, see Appendix (ref) for an example. andrews2017 proposes an ICS based on the singular values of sample Jacobian which measures local but not global identification strength. His test applies to GMM and likelihood problems.
Under higher-order identification, estimates are consistent but the delta-method is not valid. The limiting distribution is non-standard Rotnitzky2000, Dovonon2018. This issue is known but much less studied than weak and set identifications. Dovonon2019 study identification robust tests under second-order identification, and Lee2018 conduct standard inference under known second-order identification structure.
For any matrix (or vector) $A$, $\|A\| = \sqrt{\sum_{i,j} A_{i,j}^2} = \sqrt{\text{trace}(AA^\prime)}$ is the Frobenius (Euclidian) norm of $A$. For any square matrix $A$, $\lambda_j(A)$ refers to the j-th eigenvalues of $A$, in increasing order if $A$ is symmetric positive semi-definite; $\lambda_{\max}(A)$ and $\lambda_{\min}(A)$ refer to its largest and smallest eigenvalue, respectively, $\lambda_{1}(A),\dots,\lambda_{d}(A)$ are the first d eigenvalues of $A$ in increasing order. For a weighting matrix $W_n(\theta)$, the norm $\|\bar g_n(\theta)\|^2_{W_n}$ is computed as $\bar g_n(\theta)^\prime W_n(\theta) \bar g_n(\theta)$. The abbreviation wpa 1 will be used to abreviate “with probability approaching 1.” For $\varepsilon >0$, $B_{\varepsilon}(\theta)$ is a closed $\varepsilon$-ball around $\theta$.
Following Hansen1982, the econometrician wants to estimate the solution vector $\theta_0$ to the system of unconditional moment equations:
where $\theta_0 = (\theta_{10}^\prime,\theta_{20}^\prime)^\prime \in \overline{\Theta} = \overline{\Theta}_1 \times \overline{\Theta}_2$, a compact subset of $\mathbb{R}^{d_\theta}$, $\text{dim}(\bar{g}_n) = d_g \geq d_\theta$. $\bar g_n(\theta)= 1/n\sum_{i=1}^n g(z_i,\theta)$ is the sample vector of moment conditions, $(z_i)_{i=1,\dots,n}$ is a sample of iid or stationary random variables. The parameter $\gamma_0 \in \Gamma$ indexes the true distribution of the data $(z_i)$, including the true $\theta_0$. It has the form $\Gamma = \{ \gamma=(\theta,\omega), \theta \in \overline{\Theta}, \omega\in\Omega\}$. $\Omega$ indexes features of the data generating process beyond $\theta$ that are relevant to identification and weak convergence. $\Gamma = \overline{\Theta} \times \Omega$ is a compact subset of a metric space with a metric $\|\theta-\tilde \theta\| + d(\omega,\tilde\omega)$ between $\gamma=(\theta,\omega)$ and $\tilde\gamma=(\tilde\theta,\tilde\omega)$ that induces weak convergence for $(z_i,z_{i+m})$ for any $i,m \geq 1$.\footnote{For reduce the number of coefficients involved in the notation below, this distance will be written as $\|\theta-\tilde\theta\| + d(\gamma,\tilde \gamma)$. See Andrews2012 for a discussion of these conditions.} The operator $\mathbb{E}_{\gamma_0}$ denotes the expectation under $\gamma=\gamma_0$. $g(\theta,\gamma_0) = \mathbb{E}_{\gamma_0}(\bar g_n(\theta))$ is then the population vector of moment conditions evaluated at the true $\gamma_0 \in \Gamma$ and a coefficient $\theta$. Throughout, it is assumed that $\theta_0$ is such that $g(\theta_0,\gamma_0)=0$. The function $g(\cdot,\gamma)$ is assumed to be continuously differentiable on $\Theta$ for all $\gamma$.
Given the sample moments $\bar{g}_n$ and a sequence of positive definite weighting matrices $W_n(\theta)$ converging to $W(\theta)$, the GMM estimator $\hat \theta_n$ solves the sample minimization problem:
where $\Theta = \Theta_1 \times \Theta_2$ is the optimization space.
Assumption (ref) i. implies that $\Theta$ strictly contains $\overline{\Theta}$ so that issues arising when a parameter is on the boundary are not considered here.\footnote{See Cox2020 for results on identification and boundary robust inference.} The connected neighborhood condition plays the role of Assumption ACP iv. in Andrews2012. It implies that we can find sequences $\gamma_n$ along a continuous path in $\Gamma$ leading to $\gamma_0$ such that $0<\|\gamma_n-\gamma_0\| \to 0$. Together with a continuity condition in Assumption (ref) below, it allows to interpolate converging subsequences into converging sequences of parameters in one of the desired identification categories. This is similar to Assumption B2 in andrews2020 and Assumption 14 in Cox2020. Condition ii. is a uniform convergence condition, implied by a uniform CLT. Condition iii. ensures that $\|\cdot\|_{W_n}$ is equivalent to $\|\cdot\|$ so that the choice of $W_n$ does not alter the identifiability of the parameters.
The following steps provide a general overview of the computation of the quasi-Jacobian matrix, the ICS, and test procedure used in the paper. In the following, the matrix $P_{\theta_1}^\perp$ is an orthogonal projection matrix, projecting on the space orthogonal to $\theta_1$. It can be written as $P_{\theta_1}^\perp = \text{diag}(0_{d_{\theta_1}},1_{d_{\theta_2}})$ so that it only selects elements associated with $\theta_2$. The matrix $\overline{V}_n$ is a weighted average of estimates of $\text{var}[\sqrt{n}\overline{g}_n(\theta_b)]$, with weights proportional to $\hat{K}_n(\theta_b)$, described in more details below.\footnote{For iid data, $\text{var}[\sqrt{n}\overline{g}_n(\theta_b)]$ is approximated using $\frac{1}{n} \sum_{i=1}^n g(z_i,\theta_b)g(z_i,\theta_b)^\prime - \overline{g}_n(\theta_b) \overline{g}_n(\theta_b)^\prime$; for dependent data a HAC estimator is used.}
In the procedure, $\chi^2_{d_g - \hat{d}_n}(1-\alpha)$ is the $1-\alpha$ quantile of a $\chi^2$ distribution with $d_g - \hat{d}_n$ degrees of freedom, $d_g$ is the number of moment conditions. In the following, the number of draws $B$ is assumed to be sufficiently large for the finite-$B$ approximation error to be negligible. The $\ell_\infty$ regression ((ref)) is known as a Chebyshev (or minimax) approximation problem and can be cast as a linear programming problem boyd2004. It can be solved with a few lines of code using the cvx convex optimization toolkit.\footnote{See Supplemental Appendix (ref) for sample R code which implements the method.} ((ref)) is also solved using cvx. Finally, note that, in the procedure, the intercept $A_{n,\infty}$ and the mean $\mu_n$ are nuisance parameters, only $B_{n,\infty}$ and $\Sigma_n$ are used in steps 3-4. On the computation side: Appendix (ref) outlines a sequential Algorithm to sample on the level set (step 2i.), the quasi-Jacobian is only computed once; it is defined whether the sample moments are differentiable, or not. The standard Jacobian requires differentiability and needs to be evaluated at every grid point. Instead of the $\ell_\infty$ loss, one could use the $\ell_2$-norm which yields least-squares solutions $(A_{n,LS},B_{n,LS})$. Some technical difficulties arise because the identified set typically has measure zero, and stronger assumptions are required to derive the properties of $B_{n,LS}$ compared to $B_{n,\infty}$. The re-scaling in step 3 is discussed below. The following provides further details about the steps outlined above.
The quasi-Jacobian matrix $B_{n,\infty}$ is defined as the slope of a local linear approximation for $\bar g_n (\cdot)$ over an estimate of the identified set.
In practice, the minimization problem ((ref)) is solved over a finite grid as in ((ref)). The grid can be generated using Monte-Carlo or quasi-Monte-Carlo methods Robert2004,Lemieux2009. In the simulations, the Sobol sequence was used. In the empirical application, $d_\theta=12$ is relatively large, and the set of $\theta$ where $\hat{K}_n(\theta) >0$ is fairly narrow; the acceptance rate is very low. A very large number of draws would be needed to find sufficiently many $\theta_b$ with non-zero weight, i.e. $\hat{K}_n(\theta_b)>0$. The empirical application relies on a sequential sampling principle called Population Monte Carlo cappe2004. It constructs a sequence of proposal distributions that approximate the target distribution with increasing accuracy, see Appendix (ref) for details. These proposals can be re-purposed to compute confidence sets, reducing the additional time required for test inversion. It can also be used to compute $B_{n,\infty}$ for different values of $\kappa_n$ as a sensitivity analysis.
The kernel is assumed to have compact support. The uniform kernel, $K(x) = \mathbbm{1}_{x \in [-1,1]}$, was used in the simulations and empirical results.\footnote{The estimated $B_{n,\infty}$ is nearly numerically identical using the cosine or Epanechnikov kernels.} The first condition ensures that $\hat{K}_n(\cdot)$ selects the identified set with wpa 1 under weak identification. The second ensures that $B_{n,\infty}$ only captures the first-order Jacobian term in local expansions under (semi)-strong identification. Otherwise, it would also capture nonlinear terms from the remainder.
To illustrate the usefulness of detecting identification failure, consider the following simple data-driven test procedure. It is based on the Anderson-Rubin statistic for non-linear GMM models as described in Stock2000. To test null hypotheses of the form $H_0: \theta_1 = \theta_{10}$, compute the sample statistic: \[ \text{AR}_n(\theta_{10}) = \inf_{\theta_2 \in \Theta_2} n \left( \bar{g}_n(\theta_{10},\theta_2)^\prime \hat{V}^{-1}_n(\theta_{10},\theta_2) \bar{g}_n(\theta_{10},\theta_2) \right), \] where $\hat V_n(\theta)$ consistently estimates the asymptotic variance $\lim_{n\to\infty} \text{var}(\sqrt{n}\bar{g}_n(\theta))$. The test rejects at a nominal level $\alpha \in (0,1)$ if $\text{AR}_n(\theta_{10}) > \chi^2_{d_g - \hat d_n}(1-\alpha)$ where $\chi^2_{d_g - \hat d_n}(1-\alpha)$ is the $1-\alpha$ quantile of a chi-square distribution with $d_g - \hat d_n$ degrees of freedom. $\hat d_n \in \{0,\dots,d_{\theta_2}\}$ is computed using an identification category selection (ICS) procedure based on the quasi-Jacobian and its singular values. The procedure, described below, evaluates the number of nuisance parameters in $\theta_2$ which are potentially weakly/set identified. Using $\hat d_n = 0$ yields the largest critical value and amounts to full projection inference Dufour2005. Using $\hat d_n = d_{\theta_2}$ yields the smallest critical value which provides valid, non-conservative inferences when all of the nuisance parameters are strongly identified. Intermediate values of $\hat d_n$ improve power compared to full projection while ensuring robustness if a subset of the nuisance parameters is weakly identified. A confidence set for $\theta_1$ collects all values of $\theta_1$ for which $\text{AR}_n(\theta_{1}) \leq \chi^2_{1-\alpha}(d_g - \hat d_n)$ using the same $\hat d_n$.
The choice of $\hat d_n$ should be invariant to rescaling the sample moments $\bar{g}_n$ and/or the parameters $\theta$. To this end, the procedure relies on two normalization matrices: $\bar{V}_n = \int_\Theta \hat{V}_n(\theta) \hat \pi_n(\theta)d\theta$ and $\Sigma_n$, where $\hat \pi_n(\theta) = \hat{K}_n(\theta)/\int_{\Theta} \hat{K}_n(\theta) d\theta$. $\bar{V}_n$ an average of asymptotic variance estimators for $\lim_{n\to\infty} n \text{var}(\bar{g}_n(\theta))$. It is used to ensure the procedure is invariant to re-scaling and rotating the sample moments. $\Sigma_n$ is the $\ell_\infty$-covariance matrix minimizing $\sup_{\theta \in \Theta} \left( \log|\Sigma| + \|\theta-\mu\|_{\Sigma^{-1}}^2 \right) \hat{K}_n(\theta)$ over $\mu$ and $\Sigma$. These quantities are readily available from the steps required to compute $B_{n,\infty}$. It is important to use an estimate of the variance $\Sigma_n$ of $\theta$ on $\Theta_n$ rather than the variance of $B_{n,\infty}$ or of the sample Jacobian. When the model is set or weakly identified, the variance $\Sigma_n$ - which measures the size of the set $\Theta_n$ - does not go to zero in directions where identification fails.\footnote{Lemma (ref) shows that $\Sigma_n^{-1/2}$ is bounded above under weak identification in directions where identification fails.} The variance of $B_{n,\infty}$ or the Jacobian could be arbitrarily small, however.\footnote{Take $g(z_i;\theta) = 0$, for all $\theta$, $z_i$ a.s.. The variances of both $B_{n,\infty}$ and the Jacobian are zero; yet, $\Sigma_n \neq 0$.} Hence, $B_{n,\infty} \Sigma_n^{-1/2}$ is vanishing in directions where identification fails and is invariant to rescaling the coefficients $\theta$.
Let $P_{\theta_1}^\perp$ be the projection matrix on the orthogonal of the span of $\theta_1$, compute the singular values of the normalized $\bar{V}_n^{-1/2} B_{n,\infty} P_{\theta_1}^\perp \Sigma_n^{-1/2}$: \[ \lambda_{jn} = \lambda_j( P_{\theta_1}^\perp \Sigma_n^{-1/2} P_{\theta_1}^\perp B_{n,\infty}^\prime \bar{V}_n^{-1} B_{n,\infty} P_{\theta_1}^\perp \Sigma_n^{-1/2} P_{\theta_1}^\perp )^{1/2}, \] where $\lambda_{j}$ denotes the j-th eigenvalue in increasing order so that $0 \leq \lambda_{1n} \leq \dots \leq \lambda_{d_{\theta}n}$. By projection, the smallest $d_{\theta_1}$ singular values are equal to zero. Take $\underline{\lambda}_n \to 0$, a decreasing sequence such that $\kappa_n = o(\underline{\lambda}_n)$, and compute: \[ \hat d_n = \#\{ j \in \{ d_{\theta_1}+1,\dots,d_\theta\}, \lambda_{jn} > \underline{\lambda}_n \}, \] where $\#$ counts the number of singular values $\lambda_{jn}$ which are greater than the threshold $\underline{\lambda}_n$.
\paragraph{Choice of Tuning Parameters:} A default choice is the uniform kernel $K(x) = \mathbbm{1}_{x \in [-1,1]}$. Then, the role of the pair $(\kappa_n, W_n)$ is to estimate the solution set of parameter(s) $\theta_0$ such that $g(\theta_0,\gamma_0)=0$. For this choice of kernel, if a law of the iterated logarithm applies, then, pointwise, $\text{liminf}_{n\to\infty} \hat{K}_n(\theta_0) =1$ almost surely using $W_n(\theta_0) = \text{var}[\sqrt{n}\overline{g}_n(\theta_0)]^{-1}$ and $\kappa_n = \sqrt{2 \log[\log(n)]/n}$.\footnote{A law of the iterated logarithm implies $\text{limsup}_{n\to\infty} \sqrt{n/(2\log[\log(n)])}\|\overline{g}_n(\theta_0)\|_{W_n} = 1$ almost surely, also $K(x) = 1$ for all $x \in [-1,1]$, see e.g. petrov1995; and kosorok2008, VanderVaart1996 for references applying to empirical processes, which are not pointwise.} In that sense, efficient weighting and $\kappa_n = \sqrt{2\log\log(n)/n}$ are asymptotically optimal and makes $\hat{K}_n(\theta)$ invariant to linear transformations of the moments.
The role of the normalized $B_{n,\infty}$ and the threshold $\underline{\lambda}_n$ is analogous to the ICS procedure in Andrews2012, and the subsequent literature. Here, it is shown that if $d$ nuisance parameters are weakly identified then at least $d$ singular values are $O_p(\kappa_n)$. Hence, wpa 1, they are smaller than $\underline{\lambda}_n$, if $\underline{\lambda}_n = o(\kappa_n)$. As a result, $\hat d_n$ is no greater than the number (semi)-strongly identified nuisance parameters wpa 1, which leads to valid inferences under weak identification. Typically, using larger values of $\underline{\lambda}_n$ in an ICS procedure is desirable for robust inference since it correctly detects identification failures with greater probability in finite samples. However, it also makes the test more conservative under semi-strong identification since it incorrectly detects identification failure with greater probability. This implies a trade-off between power for semi-strongly identified models with robustness for weakly identified models. The normalization $\Sigma_n^{-1/2}$ in the procedure improves on this by making the behaviour of the ICS statistic more distinct between these two regimes. The normalization $B_{n,\infty} P_{\theta_1}^\perp\Sigma_n^{-1/2} P_{\theta_1}^\perp$ preserves the asymptotic singularity under weak identification but the normalized matrix diverges at a $\kappa_n^{-1}-$rate when identification is strong.
The main component of the procedure is the quasi-Jacobian. To better understand the main differences with the Jacobian, the following derives its properties for $n = \infty$, using a positive definite $W(\theta)$ and the uniform kernel $K(x) = \mathbbm{1}_{x \in [-1,1]}$. Take \[ (A_\infty,B_\infty) = \lim_{\kappa \to 0} \left( \text{argmin}_{A,B} (\sup_{\|g(\theta,\gamma_0)\|_W \leq \kappa} \|g(\theta,\gamma_0)-A-B \theta\|) \right),\] where the $\sup$ is taken over $\theta \in \Theta$ with $\kappa>0$. To compute the Jacobian, $\partial_\theta g(\theta_0,\gamma_0)$, one would use the set $\|\theta-\theta_0\| \leq \kappa$; the main difference is the choice of neighborhood.
This difference suggests that, unlike the Jacobian, the properties of the quasi-Jacobian depend on the set $\Theta_0 = \{ \theta \in \Theta, g(\theta,\gamma_0)=0 \}$, which collects all solutions to the moment condition. For any given value $\gamma_0 \in \Gamma$, there are three possibilities, either: i. $\Theta_0$ is non-singleton, ii. $\Theta_0 = \{\theta_0\}$ is singleton and $\partial_\theta g(\theta_0,\gamma_0)$ is singular, or iii. $\Theta_0 = \{\theta_0\}$ is singleton and $\partial_\theta g(\theta_0,\gamma_0)$ has full rank. Under i., $\theta_0$ is not globally identified. Under ii. and iii. $\theta_0$ is globally identified but only locally identified under iii. Consistency and asymptotic normality require iii., i.e. strong identification, and standard inference need not be asymptotically valid under i. or ii. The following Theorem relates the rank of $B_\infty$ to identifications i., ii., and iii.
The dependence of $\Theta_0$, $B_\infty$ on $\gamma_0$ is omitted to simplify notation. The condition for globally identified models holds with $\alpha=2$ if $g(\cdot,\gamma_0)$ is twice continuously differentiable with bounded second derivative around $\theta_0$. Theorem (ref) shows that $B_\infty$ is singular as soon as $\gamma_0$ is such that global or local identification fails (1). An immediate implication of Theorem (ref) is that for all $v \neq 0$ such that $B_\infty v \neq 0$, $P_v \Theta_0 = P_v \{\theta_0\}$; i.e. the parameter is point identified in direction $v$. This contrasts with the Jacobian which can have full rank without global identification. $B_\infty$ is singular in all directions in which global identification fails (3), or local identification fails (4); these directions may vary depending on $\gamma_0$. $B_\infty$ has full rank only if $\gamma_0$ is such that both global and local identification hold (2). Theorem (ref) holds for both just and over-identified models. The identified set $\Theta_0$ can be arbitrary, e.g. discrete. The results do require correct specification, $\Theta_0$ non-empty, for the quasi-Jacobian $B_\infty$ to be well defined. Even though Theorem (ref) is fairly general, the main results will be restricted to settings where either i. $\Theta_0$ is non-singleton, or iii. $\Theta_0$ is singleton and $\partial_\theta g(\theta_0,\gamma_0)$ has full rank. Additional results for ii. are given in Appendix (ref).
The Jacobian generally does not have property (1) or (3). When using projection methods for subvector inference, one can concentrate out nuisance parameters that are both globally and locally identified. The Jacobian can only determine the latter which is not sufficient for consistency. The following illustrates (3) and gives a sketch of the proof using a simple non-linear model where $\Theta_0$ is non-singleton but the Jacobian has full rank for all $\theta \in \Theta_0$.
\paragraph{Intuition for linear models.} For linear models, the sup-norm approximation is exact with $B_{\infty} = \mathbb{E}(x_i x_i^\prime)$ and $\mathbb{E}(z_i x_i^\prime)$ for OLS and IV, respectively. The quasi-Jacobian coincides with the Jacobian and it is singular when the regressors are multicollinear or the instruments are not relevant. Both are singular in directions where the rank condition fails.
\paragraph{Non-linear models: a pen and pencil example.} Consider a simple MA(1) process: \[ y_t = \sigma( e_t + \vartheta e_{t-1}), \quad e_t \overset{iid}{\sim} (0,1), \] where $\theta = (\vartheta,\sigma^2) \in \mathbb{R} \times \mathbb{R}_{+}$ are the parameters of interest. The model is estimated using the following set of moment conditions (the dependence on $\gamma$ is omitted in this example): \[ g(\theta) := \mathbb{E}\left(
\right)^\prime = 0. \] Whenever $\vartheta_0 \not\in \{-1,0,1\}$ and $\sigma_0^2 >0$, this system of equations has two distinct solutions: $\theta_0^1 = (\vartheta_0,\sigma_0^2)$ and $\theta_0^2 = (1/\vartheta_0,\vartheta_0^2\sigma_0^2)$. Imposing invertibility (i.e. $|\vartheta_0|<1$), or non-invertibility (i.e. $|\vartheta_0|>1$) restores identification so that, intuitively, only one dimension is unidentified. Both solutions are locally identified: the Jacobian $\partial_\theta g (\theta)$ has full rank at both values; it is uninformative about the global identification failure in this example. The goal of this example is to show that $B_{\infty}$ is informative about the lack of global identification and the direction in which identification fails. Without the quasi-Jacobian, one would need to check with pen and pencil whether $g(\theta)=0$ has multiple solutions, or not.
The first step is to find a one-to-one linear reparameterization $\beta = (\beta_1^\prime,\beta_2^\prime)^\prime$ such that $\beta_1$ is uniquely identified but $\beta_2$ is not. Let $v_2 = (\theta_0^1-\theta_0^2)/\|\theta_0^1-\theta_0^2\|$ and pick any orthogonal $v_1 \perp v_2$ such that $\|v_1\|=1$. By construction: $v_1^\prime (\theta_0^1-\theta_0^2) = 0$ and $v_2^\prime (\theta_0^1-\theta_0^2) = \|\theta_0^1-\theta_0^2\|^2 >0$. This implies that $\theta_0^1$ and $\theta_0^2$ are equal in direction $v_1$ but distinct in direction $v_2$. Pick $\beta_1 = v_1^\prime \theta$, $\beta_2 = v_2^\prime \theta$. As desired: the mapping is one-to-one, with $\beta_1$ uniquely and $\beta_2$ set identified. Property (3) in Theorem (ref) implies that directions in which $B_\infty$ is non-singular must be associated with a unique value for $\theta_0$. This first step illustrates how these directions can be constructed from the set $\Theta_0$. Importantly, the linear reparametrization need not be computed explicitly in practice, as explained below.
The second step is to show that $B_\infty$ is informative about the identification failure and contains information about the reparametrization above. In the MA(1) model, the set $\Theta_0 = \{\theta_0^1,\theta_0^2\}$ has two points. Take $\kappa >0$ and compute the intercept and slope $A_{\kappa,\infty},B_{\kappa,\infty}$: \[ (A_{\kappa,\infty},B_{\kappa,\infty}) = \text{argmin}_{A,B} \left( \sup_{ \theta \in \Theta, \|g(\theta)\| \leq \kappa} \|g(\theta) - A - B\theta\| \right), \] here using the uniform kernel $K(x) = \mathbbm{1}_{x \in [-1,1]}$, and $W = I$ for analytical simplicity. Notice that for $(A,B) = 0$, $\sup_{ \theta \in \Theta, \|g(\theta)\| \leq \kappa} \|g(\theta)\| \leq \kappa$. Also, because $\|g(\theta_0^1)\| = \|g(\theta_0^2)\| = 0 \leq \kappa$, the solution $A_{\kappa,\infty},B_{\kappa,\infty}$ is such that $\|g(\theta) - A_{\kappa,\infty} - B_{\kappa,\infty}\theta\| \leq \kappa$ for $\theta \in \{\theta_0^1,\theta_0^2\}$. Using the triangular inequality and its reverse, this implies $\|B_{\kappa,\infty}(\theta_0^1-\theta_0^2)\| \leq 2\kappa + \|g(\theta_0^1)\| + \|g(\theta_0^2)\| = 2\kappa$. Now, express this in terms of the direction vector $v_2$ constructed above: \[ \|B_{\kappa,\infty} v_2\| \leq \frac{2\kappa}{\|\theta_0^1-\theta_0^2\|} \to 0, \text{ as } \kappa \searrow 0. \]
In the limit, the quasi-Jacobian $B_{\infty}$ is singular in the direction $v_2$ where identification fails. This implies that $v_2$ is a right-singular vector associated with the singular value $0$. The singular value decomposition of $B_{\infty}$ is informative about the directions of identification failure and the linear reparametrization from the first step. While the linear reparamerization requires knowledge of $\Theta_0$ and computing all possible $\theta_0^1 - \theta_0^2$ with $\{\theta_0^1,\theta_0^2\} \subseteq \Theta_0$, Theorem (ref) implies that the right-singular vectors of $B_\infty$ associated with the singular value $0$ span all directions of identification failure $\theta_0^1 - \theta_0^2$.
In large samples and under Assumptions (ref)-(ref), $\|B_{n,\infty} v_2\| \leq 2 \kappa_n / \|\theta_0^1-\theta_0^2\|$ wpa $1$ (using the same $K$, $W$). As a result, for any sequence $\underline{\lambda}_n$ such that $\kappa_n = o(\underline{\lambda}_n)$, $\sigma_{\min}(B_{n,\infty}) \leq \underline{\lambda}_n$ wpa $1$ which signals the identification failure, as desired. To illustrate, Figure (ref) compares the distribution of the largest and smallest singular values of the Jacobian $\partial_\theta \overline{g}_n(\hat\theta_n)$, quasi-Jacobian $B_{n,\infty}$, and scaled quasi-Jacobian $\overline{V}_n^{-1/2} B_{n,\infty} \Sigma_n^{-1/2}$ with the same cutoff $\underline{\lambda}_n$. The scaling makes the singular values scale invariant. The Jacobian fails to detect the lack of identification, even for large $n$ (left panel) and also with the scaling $\overline{V}_n^{-1/2} \partial_\theta \overline{g}_n(\hat\theta_n) \Sigma_n^{-1/2}$. The quasi-Jacobian detects the identification failure since the smallest singular value is below the cutoff. However, the largest singular value is also close to the cutoff. With the scaling, the largest singular value diverges while the smallest one shrinks to zero (right panel).
The test procedure described above is said to be robust to identification failure if it has asymptotic null rejection probability bounded above by the nominal size, i.e.: \[ \limsup_{n\to\infty} \sup_{\gamma \in \Gamma, \theta = (\theta_{10}^\prime,\theta_2^\prime)^\prime \in\overline{\Theta}} \mathbb{P}_{\gamma}\left(\text{AR}_n(\theta_{10}) > \chi_{1-\alpha}^2(d_g - \hat d_n) \right) \leq \alpha. \] In the limit, the worst-case rejection rate should be no greater than the nominal size $\alpha$. Following Andrews2012, this can be determined from the asymptotic properties of the test for specific sequences of parameters $(\theta_n,\gamma_n) \in \overline{\Theta} \times \Gamma$.
The function $\delta$ indicates whether the solution $\theta_0$ to the moment condition $g(\theta,\gamma_0)=0$ is unique for a given $\gamma = \gamma_0$. The second part of the assumption implies that when $\delta(\gamma_0)=0$, there is at least one $\theta \neq \theta_0$ such that $g(\theta,\gamma_0)=0$. Sequences such that $\gamma_n \to \gamma_0$, $\delta(\gamma_0)=0$, satisfy $\delta(\gamma_n)\to 0$ since $\delta$ is continuous. The properties of $\hat\theta_n$ depend on the rate at which $\delta(\gamma_n)$ converges to zero. Under Assumptions (ref) and (ref), Lemma (ref) shows that $\|\hat\theta_n-\theta_n\|=o_p(1)$, the estimator is consistent, if $\sqrt{n}\delta(\gamma_n)\to \infty$. When $\sqrt{n}\delta(\gamma_n) = O(1)$, the estimator is generally not consistent, see e.g. Stock2000.
In the MA(1) example, pick $\overline{\Theta} = \{\vartheta_0,\sigma_0^2\} \in (\mathbb{R}/\{-1,0,1\}) \times (\mathbb{R}_+/\{0\})$, then $\delta(\gamma)=0$ regardless of $\gamma$ as long as the two distinct solutions $\theta_0^1,\theta_0^2 \in \Theta$. The second inequality only holds for $\varepsilon < \|\theta_0^1-\theta_0^2\|$, with $C=1$, which implies $\overline{\varepsilon} \in (0,\|\theta_0^1-\theta_0^2\|)$.
To give another example, consider a linear IV regression: $g(\theta,\gamma) = \mathbb{E}_{\gamma}[z_i(y_i - x_i^\prime\theta)] = \mathbb{E}_{\gamma}[z_ix_i^\prime](\theta_n - \theta)$ using $y_i = x_i^\prime \theta_n + u_i$ and $\mathbb{E}_{\gamma}(u_i z_i)=0$. Here $\|g(\theta,\gamma)\| \geq \sigma_{\min}(\mathbb{E}_{\gamma}[z_ix_i^\prime])\|\theta_n-\theta\|$ so that $\delta(\gamma) = \sigma_{\min}(\mathbb{E}_{\gamma}[z_ix_i^\prime])$ and $h(\varepsilon) = \varepsilon$. The inequality holds with equality, i.e. $C=1$, when $\theta_n-\theta$ is the right singular vector of $\mathbb{E}_{\gamma}[z_ix_i^\prime]$ associated with the smallest singular value. Here $\delta(\gamma_0)=0$ implies $\mathbb{E}_{\gamma_0}[z_ix_i^\prime]$ singular, and the model is underidentified.\footnote{Additional derivations for a non-linear regression model are given in Appendix (ref).}
The dichotomy between $\delta$ and $h$ in Assumption (ref) allows to construct a measure of global identification strength used to categorize the sequences $\gamma_n$.\footnote{A similar decomposition can be found in Chen2007 to isolate the effect of the sieve dimension $k$ on the shape of the objective in nonparametric estimation.} Let $\Gamma_0 = \{ \gamma \in \Gamma, \delta(\gamma)=0\}$ and $\Gamma_1 = \Gamma / \Gamma_0$. $\Gamma_0$ collects all DGPs such that $\theta_0$ is not uniquely identified, and in $\Gamma_1$ those that are point identified. Let $\Gamma_0(\infty) = \{ \gamma_n \in \Gamma, \gamma_n \to \gamma_0 \in \Gamma_0, \sqrt{n}\delta(\gamma_n)\to\infty\}$, $\Gamma_0(b) = \{ \gamma_n \in \Gamma, \gamma_n \to \gamma_0 \in \Gamma_0, \lim_{n\to\infty}\sqrt{n}\delta(\gamma_n) = b <\infty\}$. In the following, any converging sequence $\gamma_n$ will be assumed to belong to one of $\Gamma_0(b)$ for some $b \geq 0$, $\Gamma_0(\infty)$, or converges in $\Gamma_1$. These will be referred to as weak, semi-strong, and strong sequences.
Assumption (ref) provides sufficient conditions to establish asymptotic normality of $\hat\theta_n-\theta_n$ at a potentially slower than $\sqrt{n}$-rate. Condition i. is standard and ensures the model is locally identified. Condition ii. allows the Jacobian to be vanishing at a slower than $\sqrt{n}$-rate in some directions. Conditions iii. is a stochastic equicontinuity condition. Condition iv. implies that the Taylor remainder is quadratic under the weaker norm $\|\partial_\theta g(\theta_n,\gamma_n)(\cdot)\|$, which is the relevant norm for convergence when $\gamma_n\in \Gamma_0(\infty)$. Indeed, Lemma (ref) establishes that $\sqrt{n}\|\partial_\theta g(\theta_n,\gamma_n)(\hat\theta_n-\theta_n)\|=O_p(1)$. Condition iv. excludes settings where the non-linear remainder dominates the first-order term.\footnote{These second or higher-order identification issues are not considered in the main text, additional results for the quasi-Jacobian under higher-order identification are given in the Supplement.} Condition v. is analogous to Assumption 3iv in Antoine2012. It requires a rescaling for which the Jacobian is non-singular in the limit. For instance, under a singular value decomposition of the form $\partial_\theta g(\theta_n,\gamma_n) = U D_n V^\prime$, we have $\partial_\theta g(\theta_n,\gamma_n) H_n = UV^\prime = R_0$. The rescaling corrects for the possibly vanishing, but non-zero, terms in the diagonal $D_n$. antoine2021 discuss conditions relating to Assumption (ref) in more detail.
Proposition (ref) implies that the test is asymptotically valid for any choice of $\hat{d}_n \in \{0,\dots,d_{\theta_2}\}$ and asymptotically non-conservative if $\hat{d}_{n}=d_{\theta_2}$ wpa 1. Furthermore, for just-identified models $\text{QLR}_n(\theta_1) = \text{AR}_n(\theta_1)$, and the test is asymptotically efficient if $\hat{d}_{n}=d_{\theta_2}$ wpa 1.
\paragraph{Linear reparameterization.} As in the MA(1) example, the derivations rely on a one-to-one linear reparameterization $\beta = M\theta = (\beta_1^\prime,\beta_2^\prime)^\prime$ with $\beta_1$ uniquely and $\beta_2$ set identified. The following steps construct the reparameterization, which is not implemented in practice: the span of right-singular vectors associated with singular values below $\underline{\lambda}_n$ consistently estimates the span of identification failure. The following applies to just and over-identified models.
First, take $\gamma_0 \in \Gamma$, collect all solutions to the moment conditions $\Theta_0 = \{ \theta \in \Theta, g(\theta,\gamma_0) = 0 \}$. Let $V_2 = \text{span}( \{ v_2 = \theta_0^1 - \theta_0^2, (\theta_0^1,\theta_0^2) \in \Theta_0 \times \Theta_0 \})$ and $V_1 = V_2^\perp$. If $V_2 = \{0\}$, then $V_1 = \mathbb{R}^{d_\theta}$ which implies that $\Theta_0$ is a singleton; i.e. the parameters are uniquely identified. This is the case when $\gamma_0 \in \Gamma_1$. If $\{0\} \subset V_2$ strictly, then $V_1 \subset \mathbb{R}^{d_\theta}$ strictly; i.e. the parameters are set identified. This is the case when $\gamma_0 \in \Gamma_0$. As in the MA(1) example, by projection $P_{V_1}(\theta_0^1 - \theta_0^2) = 0$ for any two $\theta_0^1,\theta_0^2 \in \Theta_0$; i.e. the solution is unique on $V_1$. In contrast, for any non-zero $v_2 \in V_2$, there exists two distinct $\theta_0^1,\theta_0^2 \in \Theta_0 \times \Theta_0$ s.t. $v_2^\prime(\theta_0^1-\theta_0^2) \neq 0$, by construction. Define $\beta_1$ as the projection of $\theta$ on $V_1$ and $\beta_2$ the projection on $V_2$. The matrix $M$ combines the bases of $V_1$ and $V_2$. As illustrated by the MA(1) example, it may not be possible to improve on this linear reparameterization with a non-linear one without some further structure on the moments or the model. The reparameterization is defined up to a rotation on $V_1$ and $V_2$, respectively.
For testing $H_0: \theta_1 = \theta_{10}$, the identification status of the nuisance parameters $\theta_2$ matters. Consider a further sub-decomposition $(\beta_1,\beta_{21},\beta_{22})$ where only $\beta_{22}$ is unidentified under the restriction $\theta_1 = \theta_{10}$. To find it, take $V_{22} = \text{span}( \{ \theta_0^1-\theta_0^2, (\theta_0^1,\theta_0^2) \in \Theta_0 \times \Theta_0, P_{\theta_1}\theta_0^1=P_{\theta_1}\theta_0^2 = \theta_{10} \} )$ and follow the same steps as above. By construction, $V_{22}$ is a subset of $V_2$, also $\theta_1$ is in $V_{22}^\perp$ and $\beta_{22}$ is the subset of $\theta_2$ which is unidentified under $H_0$.\footnote{Note that, by linearity and by construction, $\text{span}(P_{V_{22}}) = \text{span}(P_{V_{22}} P_{\theta_1}^\perp) \subseteq \text{span}(P_{V_{2}} P_{\theta_1}^\perp)$.}
Now consider sequences $(\theta_n,\gamma_n) \to (\theta_0,\gamma_0)$ with $\gamma \in \Gamma_0$. Combine the linear reparameterization with the continuity of $g$ with respect to $\theta$ and $\gamma$ to find, using the Maximum Theorem, that for all $(\theta_n,\gamma_n) \to (\theta_0,\gamma_0)$, any $\varepsilon >0$, and letting $\beta_n = M\theta_n$:\footnote{To apply the Maximum Theorem, note that by continuity of $g(\cdot,\gamma_0)$ and compactness of $\Theta$, both $\Theta_0$ and $\mathcal{B}_2^0$ are compact subsets of $\mathbb{R}^{d_\theta}$ and $\mathbb{R}^{d_{\beta_2}}$, respectively. Similar equations can be derived for $(\beta_1,\beta_{21},\beta_{22})$ with the added constraint $\theta_1 = \theta_{1n}$.}
where $\mathcal{B}_2^0 = P_{V_2}\Theta_0$ is the identified set for $\beta_2$ when $(\theta,\gamma) = (\theta_0,\gamma_0)$. The first limit implies $\beta_1$ is consistently estimable, while the second and third imply that the population objective function becomes flat (only) on $\mathcal{B}_2^0$. The decomposition so far separates $\beta_1$ point identified from $\beta_2$ set unidentified when $\gamma=\gamma_0$.\footnote{Note that for the class of models considered in Andrews2012, their parameter $\beta$ which is point identified and determines identification strength is included in the vector $\beta_1$ constructed here.}
If there is a single source of identification failure, then $\sup_{\beta_2 \in \mathcal{B}_2^0} \sqrt{n}\|g(\beta_{1n},\beta_2,\gamma_n)\|$ is determined by a scalar subset of $\gamma_n$, and is bounded above for weak sequences. To illustrate, consider the linear IV example again with a single endogenous regressor $x_i$ and one instrument $z_i$. In this case $\sqrt{n}\|g(\beta_{1n},\beta_2,\gamma_n)\| = \sqrt{n}|\text{cov}(x_i,z_i)| \times |\beta_2-\beta_{2n}|$ depends on the scalar $\sqrt{n}|\text{cov}(x_i,z_i)|$ being bounded which characterizes weak sequences. In the case with multiple sources of identification failure, there may be mixed identification strength, and some components of $\beta_2$ may be (semi)-strongly identified, so the reparameterization needs to be further refined. This is deferred to Appendix (ref).
Assumption (ref) adds this additional structure to ((ref))-((ref)), where $\beta_1$ are assumed semi-strongly and $\beta_2$ weakly identified. The first part Assumption (ref)i. implies $\beta_1$ is consistently estimable, allowing for some components to be semi-strongly identified. The second and third part imply the objective function is flat with respect to $\beta_2$ but only on the identified set $\mathcal{B}_2^0$. For the quasi-Jacobian, the $\limsup_{n\to\infty} \sup_{\beta_2 \in \mathcal{B}_{2}^0} \sqrt{n}\|g(\beta_{1n},\beta_2)\| < \infty$ implies that $\|\overline{g}_n(\beta_{1n},\beta_2)\|_{W_n} \leq \kappa_n$ uniformly in $\beta_2 \in \mathcal{B}_2^0$ with increasing probability so that Step 2.i of the procedure consistently estimates the identified set and all directions of identification failure. Similarly, condition ii. repeats the conditions under the restriction that $\theta_1 = \theta_{1n}.$ The parameters $(\beta_1,\beta_{21})$ correspond to the directions that are consistently estimable. To simplify notation, the Proposition below denotes as $\phi$ these $d_\theta - d_{\theta_1} - d_{\beta_{22}} = d_\phi$ coefficients that are consistently estimable and semi-strongly identified under $H_0: \theta_1=\theta_{1n}$.
Proposition (ref) implies that the test procedure has limiting null rejection probability bounded by the nominal size for weak sequences as long as $\hat d_n \leq d_{\beta_1}+d_{\beta_{21}} - d_{\theta_1} = d_\phi$ wpa 1, since $\mathbb{P}_{\gamma_n}\left(\text{AR}_n(\theta_{1n}) \geq \chi^2_{1-\alpha}(d_g-\hat d_n)\right) \leq \mathbb{P}_{\gamma_n}\left( \inf_{\phi \in \Phi}\|\bar{g}_n(\theta_{1n},\phi,\beta_{22n})\|_{V_n^{-1}} \geq \chi^2_{1-\alpha}(d_g-d_\phi)\right)+o(1) \to \alpha$. Note that Assumption (ref) with respect to $\phi$ is implied by Assumption (ref) ii.
As discussed above, the properties of the ICS and test procedure are tied to those of the quasi-Jacobian under different identification regimes. The following derives the large sample behaviour of the sup-norm and least-squares quasi-Jacobian matrices $B_{n,\infty}$ under strong, semi-strong, and weak identification.
The proof is given in Appendix (ref). Theorem (ref) implies that, for (semi)-strong sequences, the quasi-Jacobian, and the Jacobian are asymptotically equivalent after re-scaling to a non-singular limit. For non-smooth moments, where the sample Jacobian is not defined as in quantile-IV regression or SMM estimation of discrete choice models, $B_{n,\infty}$ can be used in the sandwich formula to compute standard errors for $\hat\theta_n$. Assumption (ref) v. implies $\lambda_{\min}(H_n \partial_\theta g(\theta_n,\gamma_n)^\prime \partial_\theta g(\theta_n,\gamma_n) H_n) = \lambda_{\min}(R_0^\prime R_0) + o(1) \to 1$, hence: \[ \lambda_{\min}( B_{n,\infty}^\prime B_{n,\infty} ) = \lambda_{\min}\Big(\partial_\theta g(\theta_n,\gamma_n)^\prime \partial_\theta g(\theta_n,\gamma_n)\Big) \left( 1 + o_p(1) \right). \] For sufficiently strong sequences such that $\underline{\lambda}_n^2 = o\left(\lambda_{\min}(\partial_\theta g(\theta_n,\gamma_n)^\prime \partial_\theta g(\theta_n,\gamma_n))\right)$, where $\underline{\lambda}_n$ is the cutoff in Section (ref), this implies that $\hat d_n = d_{\theta_2}$ wpa 1.
Theorem (ref) shows that when $\theta$ is not uniquely identified, the quasi-Jacobian vanishes at a $\kappa_n$ rate in all directions associated with the identification failure. The span of these directions has dimension $d_{\beta_2}$ so that $B_{n,\infty}$ vanishes on a subspace of dimension $d_{\beta_2}$. Hence, small singular values are indicative of an identification failure, and the number of weakly identified coefficients. The constants involved in the $O_p$ terms are made explicit in the proof. The following Proposition extends these results to $B_{n,\infty}P_{\theta_1}^\perp$, which focuses on the identification status of the nuisance parameters only. For both results, the proof is similar to the derivations used for the MA(1) example.
As discussed above, the ICS procedure used to compute $\hat d_n$ relies on two normalizations that ensure invariance to rescaling of the sample moments and/or the parameters. The first normalizing matrix is $\Sigma_n$ computed in the procedure outlined above. $\Sigma_n^{-1/2}$ is shown to be bounded above in directions associated with the identification failure in Lemma (ref), so that Proposition (ref) extends to the normalized $B_{n,\infty}P_{\theta_1}^\perp \Sigma_n^{-1/2} P_{\theta_1}^\perp$. Under strong identification, Lemma (ref) implies that $\Sigma_n^{-1/2} = O(\kappa_n^{-1})$ so that $B_{n,\infty}P_{\theta_1}^\perp \Sigma_n^{-1/2} P_{\theta_1}^\perp$ diverges at a $\kappa_n^{-1}$ rate in $d_{\theta}-d_{\theta_1}$ directions. As a result, $B_{n,\infty}P_{\theta_1}^\perp \Sigma_n^{-1/2} P_{\theta_1}^\perp$ vanishes at a $\kappa_n$-rate in directions where identification fails, and diverges at a $\kappa_n^{-1}$-rate when all parameters are strongly identified.
The second normalizing matrix is $\overline{V}_n = \int_\Theta \hat V_n(\theta) \hat\pi_n(\theta)d\theta$, where $\hat V_n(\theta)$ is an estimator of the asymptotic variance $\lim_{n \to \infty} \text{var}_{\gamma_n}(\sqrt{n}\overline{g}_n(\theta))$. The Assumption below requires $\hat V_n(\theta)$ consistent and asymptotically non-singular so that the normalization does not alter the asymptotic properties of $B_{n,\infty}P_{\theta_1}^\perp \Sigma_n^{-1/2}$.
Theorem (ref) establishes the uniform validity of the test procedure described in Section (ref) under strong, semi-strong, and weak sequences. First, it is shown that the normalizations do not affect the predictions of Theorems (ref), (ref), and Proposition (ref). Then, since $\overline{\Theta}$ and $\Gamma$ are compact, the worst-case rejection probability is attained by a converging subsequence which, using the stated assumptions, can be interpolated into a converging sequence in either $\Gamma_0(b)$ for some $b \in [0,\infty)$, $\Gamma_0(\infty)$, or converging in $\Gamma_1$. The result then relies on two properties. The first is that $\hat d_n \geq d_{\beta_{22}}$ under weak identification, and the second is that $\text{AR}_n(\theta_{1n}) = \inf_{\theta_2 \in \Theta_2} \text{AR}_n(\theta_{1n},\theta_2) \leq \text{AR}_n(\theta_{1n},\hat\phi_n,\beta_{22n})$ which has a standard chi-squared limiting distribution with degrees of freedom that only depend on the dimension of $\bar{g}_n$, and the number of identified nuisance parameters. For just-identified models, the resulting procedure is efficient under strong identification since it uses the smallest valid critical value, and is equivalent to a quasi-Likelihood ratio test. For over-identified model, the test uses the smallest valid critical value for the projected AR test so it is non-conservative within that class. The results above can be extended to some other existing robust test statistics. For instance, the K-statistic of Kleibergen2005 is such that, under additional regularity conditions, $\text{K}_n(\theta_{1n}) = \inf_{\theta_2 \in \Theta_2} \text{K}_n(\theta_{1n},\theta_2) \leq \text{K}_n(\theta_{1n},\hat\phi_n,\beta_{22n})$ which also has a chi-squared limiting distribution with reduced degrees of freedom.
The finite-sample properties of the quasi-Jacobian matrix and the test procedure are illustrated using a consumption capital asset pricing model (CAPM) as in Wright2003.
Let $\delta$, $\gamma$ measure time preference and relative risk aversion. $C_t,D_t,R_t$ are real consumption, dividends, and the gross asset return at time $t$. The Euler equation is: $\mathbb{E}_t[ \delta R_{t+1}(C_{t+1}/C_t)^{-\gamma}-1 ] = 0$, where $C_{t+1}/C_t$ measures consumption growth. $R_{t}$ depends endogenously on $y_{t+1} = ( c_{t+1},d_{t+1} )^\prime$, where $c_{t+1} = \log( C_{t+1}/C_t )$ and $d_{t+1} = \log( D_{t+1}/D_t )$, which follows a first-order vector autoregressive (VAR) process: $y_{t+1} = \mu + \Phi y_t + u_{t+1}$, where $u_{t+1} \overset{iid}{\sim} \mathcal{N}(0,\Lambda)$. The sample moments are: \[ \overline{g}_n(\theta) = \frac{1}{n} \sum_{t=1}^n [\delta R_{t+1}(C_{t+1}/C_t)^{-\gamma}-1]Z_t,\] where $Z_{t} = (1,R_t,C_t/C_{t-1})^\prime$. tauchen1986 illustrates how $(\mu,\Phi,\Lambda)$ affects the finite-sample properties of $\hat\theta_n = (\hat\delta_n,\hat\gamma_n)$. The following considers three DGPs: Rank Failure (RF), Near Rank Failure (NRF) and Full Rank (FR).\footnote{RF, NRF and FR correspond to RF1, NRF1 and FR in Wright2003.} Wright2003 explains that they correspond to $\theta = (\delta,\gamma)$ being set, weakly, and strongly identified. NRF is calibrated to match annual U.S. data kocherlakota1990.
Table (ref) reports rejection rates for the method in Section (ref) (Proj$_1$), full projection inference using $\chi^2_{3}$ (Proj$_2$) and $\chi^2_{2}$ (Proj$_3$) critical values as well a t-test with standard normal critical value ($t_n$). The empirically relevant sample sizes are $n=100,250$. $n=500,1000$ illustrate large sample properties. The parameter space is $\Theta = [0.7,1.1] \times [0,10]$. The t-test does not control size in RF and NRF. It is closer to nominal size for FR. However, as Figure (ref) in Appendix (ref) shows, another global solution $\hat\theta_n \simeq (0.7,10)$ is estimated in about 1% and 0.05% of the replications for $n=100,250$. Here, the parameters are locally strongly identified, but not globally. The sample Jacobian would not detect this issue which leads to some over-rejection for the t-test. In comparison, the proposed procedure (Proj$_1$) has null rejection rates below nominal size across sample sizes and DGPs.
Wright2003, Antoine2009 explain that one coefficient is always strongly identified. Table (ref) and Figure (ref) confirm this. The procedure finds $\gamma$ to be weakly identified in nearly all replications for RF, NRF, and $\delta$ strongly identified. For FR, the procedure finds $\gamma$ weakly identified in $14\%$ and $0.5\%$ of replications when $n=100,250$.
\newcolumntype{g}{>{\columncolor{gray}}c}
Figure (ref) compares the power of the proposed procedure (AR$_1$) with full projection inference (AR$_3$), and projection inference with the nuisance parameter concentrated out (AR$_2$) as well as the t-test when appropriate (FR with $n = 250,500,1000$). The results show power improvement over full projection inference when the nuisance parameter is strongly identified, i.e. when testing hypotheses about $\gamma$. When the model is strongly identified (FR), the procedure is less powerful than the t-test because of over-identification. Result for a just-identified specification with $Z_t = (1,R_t)^\prime$ and a larger $\kappa_n$ are given in Appendix (ref). Another example in that Appendix compares the procedure with Andrews2012 for a non-linear regression.
To illustrate the empirical content which can be gained from the quasi-Jacobian for inference, consider a simulated method of moments estimation of the long-run risks (LRR) model bansal2004. There are two latent variables representing a persistent component to the level of consumption growth $x_{1,t}$ and stochastic volatility $x_{2,t}$:
where $f(x) =\sqrt{x}$ if $x \geq \sigma^2$ and $f(x)=\sigma^2/\sqrt{2\sigma^2 - x}$ as in Calvet2015. Consumption and dividend growth $g_t,d_{d,t}$ are then given by:
where $(e_t,w_t,\eta_t,u_t)\sim \mathcal{N}(0,I)$ iid. Given an Epstein-Zin utility function, equilibrium conditions imply that financial variables, log-price dividend ratio $z_{m,t}$, market return $r_{m,t}$ and the risk-free rate $r_{a,t}$ can be written as:
where the coefficients $(A_{0,m},A_{1,m},A_{2,m},\kappa_{0,m},\kappa_{1,m},A_{0,r},A_{1,r},A_{2,r})$ are computed numerically as a solution of a non-linear system of equations involving the full vector of 12 parameters $\theta = (\rho,\phi_e,\sigma,\nu,\sigma_w,\mu,\mu_d,\phi,\phi_d,\delta,\gamma,\psi^{-1})$ where $\delta$ is the discount factor, $\gamma$ risk-aversion, and $\psi^{-1}$ the inverse intertemporal elasticity of sustitution (IES). See bansal2004 for details. The variables above need to be further time-aggregated from the monthly decision interval to match the quarterly frequency of the data. There are a number of estimations of this model using one of SMM and Indirect Inference,\footnote{See bansal2007,hasseltoft2012,Calvet2015,grammig2018.} GMM,\footnote{See constantinides2011,bansal2012,bansal2016.}, or Bayesian estimation\footnote{See schorfheide2018.} There are, however, several concerns for the identifiability of the parameters. Calvet2015 show that the latent variables $(x_{1,t},x_{2,t})$ cannot be recovered from the data for uncountably many values of $\theta$, resulting in highly irregular GMM and likelihood objective functions. grammig2018 find that the stochastic volatility component is poorly identified and calibrate $\nu = \sigma_w =0$. However, stochastic volatility in long-term consumption growth has important implications for asset prices schorfheide2018. Several papers report estimates with very small standard errors grammig2018, but estimates can vary a lot across estimations. This suggests that some parameters are likely not globally identified but might be locally identified.
The following considers joint inference for the two preference parameters $\theta_1 = (\gamma,\psi^{-1})$. The remaining coefficients are $\theta_2 = (\rho,\phi_e,\sigma,\nu,\sigma_w,\mu,\mu_d,\phi,\phi_d,\delta)$. Amongst these nuisance parameters, it seems reasonable to think that several are (semi)-strongly identified. However, the asset pricing coefficients $(A_{0,m},\dots)$ are highly non-linear functions of $\theta$ so it is arguably more difficult to pin down exactly how many and which ones are well identified. Nevertheless, the results in this paper imply that $B_{n,\infty}P_{\theta_1}^\perp$ can determine how many nuisance parameters are weakly identified with high probability.
The moment conditions used for inference are based on matching the following sample with simulated moments: means of all variables, variances of $g_{t},g_{d,t},z_{m,t}$, AR$(2)$ coefficients of $g_{t}$, and autocorrelation of $g_{t}^2$.\footnote{A quasi-difference $z_{m,t} - 0.95 z_{m,t-1}$ is applied beforehand because $z_{m,t}$ is very persistent making $\hat V_n$ nearly singular, the quasi-differencing solves this issue and makes the estimation below more stable.} These just-identified moments match quantities of interest that are commonly reported in calibrations or post-estimation, see e.g. beeler2012. The estimation is conducted using U.S. data shared by grammig2018 for $(g_t,g_{d,t},z_{m,t},r_{m,t},r_{a,t})$ over 1947Q2-2014Q4, totalling in $n=271$ observations. The simulated moments are computed over $S=2$ samples. The bounds for the optimization space $\Theta$ are $\rho \in [0.9,0.995], \phi_e \in [0,0.1], \sigma \in [10^{-4},0.1], \nu \in [0,0.995], 10^5 \times \sigma_w \in [0,2],\mu \in [-0.035,0.035],\mu_d \in [-0.035,0.035],\phi \in [0,10],\phi_d \in [0,10],\delta \in [0.93,1.2],\gamma \in [0.05,25],\psi^{-1} \in [0.01,3]$. Computations are conducted in R and C++ using Rcpp.
Table (ref) compares the spectrum of the normalized Jacobian and quasi-Jacobian.\footnote{The estimate $\hat\theta_n$ used for the Jacobian is computed by using the calibration in bansal2004 as starting value, and alternating between the Nelder-Mead and bobyqa optimizers until convergence. Note that different seeds for the simulated samples yield very different estimates but similar fitted moments. Also, the sample gradient is not available analytically; it is computed by finite differences which here is quite sensitive to the choice of step size.} Using the threshold $\underline{\lambda}_n = \sqrt{2\log(n_S)/n} = 0.21$ implies that $B_{n,\infty}$ detects $5$ directions of identification failure, with an additional singular value just above the threshold. In comparison, the gradient is small in $3$ directions. After projecting out $\theta_1 = (\gamma,\psi^{-1})$, there are $7$ singular values above the threshold, indicating $7$ (semi)-strongly identified parameters. Hence, inference for $(\gamma,\psi^{-1})$ relies on a $\chi^2_{5}(0.95)=11.1$ critical value. In comparison, full projection relies on $\chi^2_{12}(0.95)=21$, and standard inference $\chi^2_{2}(0.95)=6$.
Figure (ref) reports 5000 draws of $\theta_1$ such that $\text{AR}_n(\theta_1) \leq \chi^2_{5}(0.95)$ using the Population Monte Carlo algorithm in Appendix (ref), plus their convex hull in blue. Values for $\gamma$ are contained in $[5.41,25]$ and $\psi^{-1}\in [0.01,0.90]$. This excludes several regions of interest. First, we can reject $H_0: \psi \leq 1$ at the 95% confidence level, i.e. the IES is strictly greater than unity. Second, we can reject $H_0:\gamma = \psi^{-1}$ and conclude that the utility function is not CRRA. Finally, the confidence set favours $H_1: \gamma > \psi^{-1}$ over $H_0: \gamma \leq \psi^{-1}$. Under $H_1$, households prefer an early resolution of uncertainty; their preference for consumption smoothing is less than their relative risk aversion. Although not reported here, note that full projection inference cannot reject some of these null hypotheses. As a robustness check with respect to tuning parameters, Appendix (ref) finds the same results using $\chi^2_6$ critical values (Figure (ref)) and using a larger value for $\kappa_n$ (Table (ref)).
This paper introduces a quasi-Jacobian matrix which is asymptotically equivalent to the usual Jacobian matrix under strong and semi-strong identification but is asymptotically singular when global identification fails. This can be useful because the Jacobian is not always informative about global identification failures. While the inference procedure relies on the AR statistic, extending the results to the robust score test is straightforward, as discussed earlier. For overidentified models, it could be interesting to extend the theory to more powerful test statistics such as the CQLR/AR test in andrews2017. Another concern could be that a given choice of moments does not identify the parameters but another set of moments might. This is a moment selection problem. In that case, it could be interesting to extend the quasi-Jacobian to a continuum of moment conditions which can be used for conditional GMM estimation Carrasco2000; allowing the use of all available information rather than selecting finite dimensional moments.