Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
75,430 characters · 12 sections · 59 citation commands
Overidentification testing with weak instruments and heteroskedasticity
\thispagestyle{empty}
Exogeneity is a key identifying assumption for validity in instrumental variables (IV) regression. Overidentification tests assess this assumption with the hypothesis specification $\mathbb{H}_0:\mathbb{E}[z_iu_i]=0$ v.s. $\mathbb{H}_1:\mathbb{E}[z_iu_i]\neq0$, where $z_i$ and $u_i$ are the instrument vector and structural error for observation $i$, respectively. These tests can be performed whenever $k_z>k_x$, where $k_z$ and $k_x$ are the number of instruments and endogenous regressors respectively. The Sargan test sargan1958 is an early example of such a test, testing the orthogonality of instruments in the standard linear setting with homoskedasticity. The Hansen test hansen1982, denoted by $J$, generalises the Sargan test to the generalised method of moments (GMM) framework, allowing for overidentification testing on a far greater range of econometric models. The properties of these test statistics are well-known under standard asymptotics (e.g. greene2003).
In this paper we discuss the use of the Kleibergen-Paap ($KP$ hereafter) rank test kleibergen2006 as a heteroskedasticity-robust overidentification test and compare its suitability to the ubiquitous $J$-test. This is the first assessment of the test in an overidentification context, following from windmeijer2018 who presents a framework linking overidentification tests, underidentification tests, the score test, and rank tests. Underidentification tests assess rank hypotheses such as $\mathbb{H}_0:\textup{rank}(\Pi)=k_x-1$ v.s. $\mathbb{H}_1:\textup{rank}(\Pi)=k_x$, where $\Pi$ is the $k_z \times k_x$ first-stage parameter matrix (the $KP$-test is a standard underidentification test, commonly reported in statistical packages such as Stata). By formulating overidentification tests as rank tests on the reduced-form $k_z \times (k_x+1)$ parameter matrix $\bar{\Pi}$ with $\mathbb{H}_0:\textup{rank}(\bar{\Pi})=k_x$ v.s. $\mathbb{H}_1:\textup{rank}(\bar{\Pi})=k_x+1$ instead of the usual $\mathbb{H}_0:\mathbb{E}[z_iu_i]=0$, windmeijer2018 shows that any underidentification test can be used as an overidentification test if performed on an appropriate auxiliary regression, with an analogous result for using overidentification tests as underidentification tests. The paper further shows that the 2SLS-based and LIML-based robust score tests are equivalent to the $J$- and $KP$-tests respectively. The $J$- and $KP$-tests have the same limiting $\chi^2(k_z-k_x)$ distributions under the null when instruments are strong and provide numerically similar values in finite samples. However, their behaviour in a weak instrument setting is unlikely to be similar, and the relative performance of IV methods when instruments are potentially weak is often a key motivator for choosing between them.
staiger1997 derive the limiting distribution of the Sargan test under weak instruments with homoskedasticity, providing a general recommendation to use LIML over 2SLS for overidentification tests based on numerical evidence. We extend these results by deriving the limiting distribution of the robust score test (which nests $J$ and $KP$) under heteroskedastic weak instruments. Due to the highly non-standard nature of the limiting distributions, a direct comparison is difficult. We therefore conduct an extensive Monte Carlo study, and usually find $KP$ preferable to $J$. The $J$-test performs extremely poorly in models with high endogeneity and/or strong heteroskedasticity and in general $KP$ exhibits better size properties and is much less likely than $J$ to be severely size distorted, especially when the number of overidentifying restrictions is small. Test size is dependent on model parameters that are not consistently estimable with weak instruments such as the strength of endogeneity. Consequently, it is difficult for researchers to gauge which test statistic would work better based on parameters estimated from data, and recommend a conservative approach that avoids severe size distortions where possible. We therefore find that our guidance to use the $KP$-test (equivalent to a LIML robust score test) under heteroskedastic weak instruments generalises the guidance from staiger1997 for LIML-based testing under homoskedasticity.
Given these theoretical results, we revisit the classic macroeconomic problem of estimating the elasticity of intertemporal substitution (EIS), the degree to which consumers adjust their consumption path in response to changes in the expected real interest rate. Due to the importance of this parameter for both theory and policy, it is unsurprising that the EIS has been given considerable attention in the empirical literature (e.g. hansen1983 and hall1988). However, there is large variation in values of the EIS that constitute credible estimates, with substantial variation depending on the country, asset of choice or how the data are aggregated. Common concerns regard the strength and validity of the instruments used in estimation, often formed from lagged macroeconomic indicators. Such instruments should be naturally exogenous, but unfortunately are often poorly correlated with endogenous variables at time $t$, particularly as the number of lags increases, and this can lead to substantial weak instrument problems yogo2004. The model also likely suffers from heteroskedasticity, which represents precautionary saving in this model yogo2004, gomes2013.
Within the literature, there are a number of papers where both weak instruments and heteroskedasticity are present, and that reject the overidentifying restrictions using the $J$-test e.g. epstein1991, dacy2011, gomes2011, and gomes2013. This leads to doubt over the validity of the moment conditions and instruments that identify and estimate the model and could potentially, at least partially, explain the variation in estimates of the EIS found in the literature. However, our simulation results suggest that $J$ performs poorly under heteroskedastic weak instruments, with often large over-rejections of the null hypothesis, whereas the $KP$-test does not, which may cast doubt over the rejections of the overidentifying restrictions seen in the literature.
In light of our theoretical and numerical evidence, we compare the performance of the $J$-test and the $KP$-test using the datasets of yogo2004 and pozzi2022. These two applications provide different insights; the yogo2004 application allows us to clearly see the impact of weak instruments via estimating two specifications, by switching between consumption and the real interest rate to be the independent variable. The instrument set with the second normalisation is substantially weaker. On the other hand, the pozzi2022 application allows us to assess test performance with two different sets of instruments, where one set of instruments is more plausibly valid but weaker, and the other is less plausibly valid but stronger.
We provide empirical evidence that the $J$-test over-rejects the overidentifying restrictions, whereas the $KP$-test rejects at a frequency consistent with a nominal 5% level. We find similar patterns of results using datasets from the two papers and argue that the $KP$-test is more reliable than the $J$-test. Both tests provide very similar numerical values when instruments are strong and both do not reject the overidentifying restrictions. However, when instruments are possibly weak, $KP$ continues to not reject the overidentifying restrictions, whereas the $J$-test rejects frequently. We suggest that the substantial number of rejections seen in the literature can be accounted for by the poor behaviour of the $J$-test, rather than there being issues with the instruments or model specification, with the key implication of this being that instrument invalidity is unlikely the cause of variation in EIS obtained in the literature.
In the following section, we present the linear model and assumptions. We also introduce the relevant estimators and test statistics, and describe the framework linking overidentification, underidentification and rank tests. Section (ref) focuses on a theoretical analysis of the estimators and test statistics under heteroskedastic weak instruments. Section (ref) assesses the behaviour of the test statistics in Monte Carlo simulations, and provides numerical evidence that the $KP$-test typically performs better than the $J$-test under heteroskedastic weak instruments. In Section (ref) we provide an overview of the basic lifecycle consumption model, and the moment restrictions to be tested. In Section (ref), we compare the performance of the $J$- and $KP$-tests using the dataset of yogo2004 and suggest that $KP$ has the superior performance. We defer the application using the pozzi2022 dataset to the Appendix purely for space considerations, but emphasise that this application is as valuable and insightful as the yogo2004 application. All proofs and additional simulations are likewise deferred to the Appendix.
Consider the linear IV model
where $y$ is the $n\times 1$ dependent variable, $X$ is the $n\times k_x$ endogenous regressor matrix, $Z$ is the $n\times k_z$ instrument matrix (with $i^{th}$ row $z_i'$), $u$ is the $n\times 1$ structural error term and $V$ is the $n\times k_x$ vector of first-stage errors. $\beta$ is the structural parameter of interest, $\Pi$ is the $k_z\times k_x$ matrix of first-stage parameters, and define $W=[ y \ \ X ]$ (with $i^{th}$ row $w_i'$). Additional exogenous regressors $\tilde{X}$ can be included ((ref)) and ((ref)) and then easily partialled out. The reduced form equation is obtained by substituting ((ref)) into ((ref)) to yield
where $\bar{\Pi} = [ \Pi_y \ \ \Pi ]$ and $\bar{V} = [ V_y \ \ V ]$, with $\Pi_y=\Pi\beta$ and $V_y=u+V\beta$. Equation ((ref)) collects all endogenous variables on the left-hand side and regresses them on the instruments. We make the following assumptions:
Assumption (ref) assumes overidentification to allow for specification testing, with the number of instruments fixed; this rules out many-instrument asymptotics. Assumption (ref) ensures that $Z'Z/n$ converges to a constant matrix with full rank. Assumption (ref) assumes the instrument is valid (and this is the assumption we wish to test). Further, it is general and no specific relationship between the second moments of the errors and instruments is assumed, except that they are finite. Because of the conditional heteroskedasticity, $\Omega_Z$ will in general lack the Kronecker structure seen under homoskedasticity. The assumption states that $\Psi_{Zu}^*$ is a multivariate normal variable with distribution $N(0,\mathbb{E}[u_i^2 z_i z_i'])$, and $\Psi_{ZV}^*$ is a matrix normal variable such that $\textup{vec}(\Psi_{ZV}^*)\sim N(0,\mathbb{E}[v_iv_i'\otimes z_iz_i'])$. These assumptions are standard.
A wide range of estimators are available for IV models. For simplicity, we focus on 2SLS and LIML (denoted by $\hat{\beta}_{2SLS}$ and $\hat{\beta}_{L}$ respectively), defined as
where $P_Z = Z(Z'Z)^{-1}Z'$ and $\hat{\alpha}_L$ is the smallest root of the characteristic polynomial $|W'P_ZW - \alpha W'W| = 0$. 2SLS is the most common IV estimator seen in the applied literature and has been extensively studied within the econometrics literature. The LIML estimator solves the maximum likelihood problem for a single equation within a system of endogenous simultaneous equations anderson1949. LIML has also been studied extensively within the theoretical literature, but has not seen the widespread use that 2SLS has enjoyed in the applied literature, despite exhibiting favourable finite-sample properties. LIML often outperforms 2SLS in finite samples with strong instruments and performs better in homoskedastic weak-instrument settings. Both estimators are part of the wider class of estimators of the form
where clearly $\alpha=0$ for 2SLS and $\alpha = \hat{\alpha}_L$ for LIML.\footnote{Since 2SLS and LIML relate to the $J$ and $KP$ tests respectively, it is simple to study these two estimators and subsequently study test statistics that applied researchers are already familiar with and for which implementation in statistical software already exists. However, generalising our results to other estimators could be an interesting topic for future research.}
All standard IV estimators require instrument exogeneity for consistency. While this assumption is impossible to test directly, when instruments outnumber the endogenous regressors, we can test for the validity of the overidentifying restrictions. Through this, we may be able to find evidence that the restrictions do not hold, and it is recommended to conduct such a test whenever possible. The null and alternative hypotheses are usually specified as
so rejection provides evidence that exogeneity is not satisfied and/or the model is misspecified. For this, it is natural to consider two-step estimators and test statistics when heteroskedasticity may be present. For the general linear GMM estimator $\hat{\beta}_{GMM}=(X'Z\mathcal{G}^{-1} Z'X)^{-1}X'Z\mathcal{G}^{-1} Z'y$, where $\mathcal{G}$ is some positive-definite weighting matrix, asymptotic efficiency is achieved if $\mathcal{G}\overset{p}{\to}\Omega_{Zu}$. Consistency is easily achieved with typical plug-in estimators. The two-step GMM estimator is given by
where $\hat{\Pi}_1$ is the appropriate first-stage estimator for $\hat{\beta}_1$ (e.g. if $\hat{\beta}_1 = \hat{\beta}_{2SLS}$, then $\hat{\Pi}_{2SLS}=(Z'Z)^{-1}Z'X$ and if $\hat{\beta}_1 = \hat{\beta}_{L}$, then $\hat{\Pi}_{L}=(Z'M_{\hat{u}_L}Z)^{-1}Z'M_{\hat{u}_L}X$, with $M_A = I_n - P_A$ and $P_A = A(A'A)^{-1}A'$ for some generic $n \times L$ matrix $A$) and $Z'H_{\hat{u}_1}Z/n\overset{p}{\to}\Omega_{Zu}$. Under standard conditions, both $\hat{\beta}_{2,2SLS}$ and $\hat{\beta}_{2,L}$ are asymptotically efficient and equivalent to the infeasible estimator $\hat{\beta}_{opt}$, given by $\sqrt{n}(\hat{\beta}_{opt}-\beta)\overset{d}{\to}N(0,\Pi'Q_{ZZ}\Omega_{Zu}^{-1}Q_{ZZ}\Pi)$, computable with complete information on $\Pi$ and $\Omega_{Zu}$ windmeijer2018.
The most common heteroskedasticity-robust two-step test statistic is the $J$-test hansen1982, given as
where $\hat{u}_2=y-X\hat{\beta}_2$. Standard implementation sets $\hat{\beta}_1 = \hat{\beta}_{2SLS}$, and under the null $J(\hat{\beta}_1,\hat{\beta}_2)\overset{d}{\to}\chi^2(k_z-k_x)$ with general heteroskedasticity and/or autocorrelation. Two-step estimation can also be used to provide a heteroskedasticity-robust score test. Partition the instrument matrix into $Z = [ Z_1 \ \ Z_2 ]$, where $Z_1$ is an $n\times k_x$ matrix of just-identifying instruments and $Z_2$ is an $n\times(k_z-k_x)$ matrix of remaining overidentifying instruments. Then the robust score statistic is given by
where $\hat{X}=Z\hat{\Pi}$, with $\hat{\Pi}$ the appropriate estimator for $\Pi$ depending on choice of $\hat{\beta}_1$. Again, $S_r(\hat{\beta}_1,\hat{\beta}_2)\overset{d}{\to}\chi^2(k_z-k_x)$ under usual asymptotics. windmeijer2018 proves that ((ref)) and ((ref)) are equivalent. However, the paper also shows that the two-step approach is actually redundant, since $Z_2'M_{\hat{X}}\hat{u}_j = Z_2'M_{\hat{X}}(y - x\hat{\beta}_j) = Z_2'M_{\hat{X}}y$ for $j\in\{1,2\}$. It therefore suffices to use a one-step robust score test given by
with $S_r(\hat{\beta}_1,\hat{\beta}_2) = S_r(\hat{\beta}_1)$ and consequently $J(\hat{\beta}_1,\hat{\beta}_2)=S_r(\hat{\beta}_1)$. Since the tests are numerically equivalent, we will ignore two-step estimators and testing and consider only $\hat{\beta}_1$ and $S_r(\hat{\beta}_1)$ for simplicity.
To demontstrate the link between overidentification and rank tests, it is useful to first consider general underidentification tests. The null and alternative hypotheses for underidentification tests are usually specified as
Identification of $\beta$ requires that $\mathbb{E}[Z'X]$ has full column rank, such that $\Pi=\mathbb{E}[Z'Z]^{-1}\mathbb{E}[Z'X]$ has full column rank; $\Pi$ is reduced-rank under the null. Therefore, there exists some non-zero vector $\zeta\in\mathbb{R}^{k_x}$ such that $\mathbb{E}[Z'X]\zeta=0$.\footnote{$A\in\mathbb{R}^{n\times m}$ is rank deficient iff there exists some non-zero vector $\zeta\in \mathbb{R}^{m}$ that satisfies $A\zeta=0$. The equivalence of underidentification testing and the standard $F$-test for $\mathbb{H}_0:\Pi = 0$ in the $k_x=1$ case is clear.}
To link this with the standard overidentification null of $\mathbb{H}_0:\mathbb{E}[z_iu_i]=0$ v.s. $\mathbb{H}_1:\mathbb{E}[z_iu_i]\neq 0$, consider $u_i = y_i - x_i'\beta$ = $w_i'\psi$, where $\psi=( 1 \ \ -\beta' )'$. It follows that $\mathbb{E}[z_iu_i]=0$ is equivalent to $\mathbb{E}[z_iw_i'\psi]=0$. Since $\mathbb{E}[z_iw_i']$ is an $n\times(k_x+1)$ matrix and $\psi$ is a non-zero $(k_x+1)$-vector with first element normalised to 1, under the null hypothesis of $\mathbb{E}[z_iw_i'\psi]=0$, $\mathbb{E}[z_iw_i']$ must be reduced-rank as this is the only way the homogeneous system of equations can be solved. This in turn implies that the null hypothesis $\mathbb{E}[z_iu_i]=0$ is satisfied if and only if $\mathbb{E}[z_iw_i']$ is reduced-rank. Just as the rank of $\mathbb{E}[z_ix_i']$ determines the rank of $\Pi$ given Assumptions (ref) and (ref), the rank of $\mathbb{E}[z_iw_i']$ determines the rank of the reduced-form parameter matrix $\bar{\Pi}$. Therefore, the familiar overidentification null of the orthogonality of the instruments and structural errors can be restated as a test of rank on the reduced-form parameter matrix as
It follows from this result that rank tests such as the $KP$-test can be used to assess instrument validity.
To link rank testing with the robust score test in ((ref)), create the partition $Z=[ Z_1 \ \ Z_2 ]$, where $Z_1$ is an $n\times k_x$ matrix of just-identifying instruments and $Z_2$ is an $n\times(k_z-k_x)$ matrix of overidentifying instruments. With conformable partitions $\Pi_y = [\Pi_{y,1}' \ \ \Pi_{y,2}' ]'$ and $\Pi = [\Pi_1' \ \ \Pi_2' ]'$, the reduced-form equation $y=Z\Pi_y+V_y$ can be expressed as
where $\tilde{\beta}=\Pi_1^{-1}\Pi_{y,1}$, $\theta = \Pi_{y,2}-\Pi_2\Pi_1^{-1}\Pi_{y,1}$, with $\Pi_1$ nonsingular, and $\eta = V_y - V\Pi_1^{-1}\Pi_{y,1}$. The first line expands $Z\Pi_y = [ Z_1 \ \ Z_2 ][\Pi_{y,1}' \ \ \Pi_{y,2}' ]'$ to $Z_1 \Pi_{y,1} + Z_2 \Pi_{y,2}$, the second line multiplies by $\Pi_1$ and $\Pi_1^{-1}$ and the third line substitutes in the expression for $Z_1\Pi_1$. Under correct model specification, $\Pi_y=\Pi\beta$ follows from ((ref)), which implies that $[\Pi_{y,1}' \ \ \Pi_{y,2}' ]' = [\Pi_1\beta \ \ \Pi_2\beta ]'$. Therefore, we can substitute $\Pi_{y,1}=\Pi_1\beta$ into $\tilde{\beta}$ to yield $\tilde{\beta} = \beta$ and similarly $\theta = \Pi_2\beta - \Pi_2\Pi_1^{-1}(\Pi_1\beta)= \Pi_2\beta -\Pi_2\beta=0$. So, testing the hypothesis $\mathbb{H}_0:\theta=0$ is equivalent to testing for correct specification. This null can be tested using the robust score test in ((ref)). In particular, windmeijer2018 shows that the $KP$-test for $\mathbb{H}_0:\textup{rank}(\bar{\Pi})=k_x$ v.s. $\mathbb{H}_0:\textup{rank}(\bar{\Pi})=k_x+1$ is equivalent to the robust score test estimated via LIML for $\mathbb{H}_0:\theta=0$ v.s. $\mathbb{H}_1:\theta\neq 0$ in ((ref)), given by
where $\hat{X}_L = Z\hat{\Pi}_L = Z(Z'M_{\hat{u}_L}Z)^{-1}Z'M_{\hat{u}_L}X$.
Under strong instruments, $J$ and $KP$ perform similarly. Figure (ref) shows the finite-sample power of the $J$- and $KP$-tests against local-to-zero alternatives across different strengths of endogeneity for a heteroskedastic strong-instrument model with a single overidentifying restriction. Increasing $\alpha$ increases the strength of the heteroskedasticity. Both tests clearly perform similarly, but the $KP$-test has higher power in most of the set-ups shown; the only seeming exception is the high endogeneity design for $\omega>0$, but the $J$-test does not have correct size in this setup. The improvements in power of $KP$ over $J$ are larger in the design with the stronger heteroskedasticity. This demonstrates clear potential for the $KP$-test to be a strong alternative to the usual $J$-test typically reported in applied research when instruments are strong.
However, the more interesting discussion is how $J$- and $KP$ behave under heteroskedastic weak instruments, as performance when identification strength is questionable is often an important motivation behind selecting methods for estimation and inference with instrumental variables. LIML often has favourable properties relative to 2SLS under homoskedastic weak instruments, so we aim to see how this translates into the heteroskedastic case. Further, staiger1997 recommend in general to use a LIML-based Sargan test for overidentification testing, due to the sensitivity of 2SLS in models with high endogeneity or a large number of overidentifying restrictions.
To consider estimation and testing with heteroskedastic weak instruments, we require the following assumption:
This assumption states that the first-stage parameter matrix is local-to-zero at rate $\sqrt{n}$, following staiger1997. We define the concentration matrix as
where $\mathcal{V}_{ZV} = (I_{k_x} \otimes \textup{vec}(I_{k_z}))' (\Omega_{ZV} \otimes Q_{ZZ}^{-1}) (I_{k_x} \otimes \textup{vec}(I_{k_z}))$.\footnote{ When the instruments are orthonormalised, the minimum eigenvalue of $\mu^2/k_z$ gives the weak-instrument test statistic of lewis2022, which with $k_x = 1$ collapses to the test statistic of montiel2013.} It is simple to show that under homoskedasticity, where $\Omega_{ZV} = \Sigma_{V} \otimes Q_{ZZ}$ for $\mathbb{E}[v_iv_i'|z_i] = \Sigma_V$, then $\mathcal{V}_{ZV} = k_z\Sigma_V$, and ((ref)) collapses down to the usual homoskedastic concentration matrix $\mu^2 = \Sigma_V^{-1/2} \Pi'Z'Z\Pi \Sigma_V^{-1/2}\overset{p}{\to} \Sigma_V^{-1/2} C'Q_{ZZ}C\Sigma_V^{-1/2}$.
For heteroskedastic limiting distributions, the following lemma is required.
The interpretations of these results are clear e.g. (a) implies that $Z'X/\sqrt{n}$ converges to a normal distribution, with asymptotic variance accounting for the structure of the heteroskedasticity. Parts (b) and (c) show that $X'P_ZX$ converges to a general noncentral correlated Wishart matrix and $X'P_Zu$ converges to a product of two matrix normals. Assume $Z'\bar{V}/\sqrt{n}\overset{d}{\to}\Psi_{Z\bar{V}}^*$ and let $W'W/n\overset{p}{\to}\Sigma_{\bar{V}}^*$, where $\Sigma_{\bar{V}}^*$ is the reduced-form variance matrix. Throughout the main body of the paper, we leave our limiting distributions in the random matrix form of Lemma (ref) for notational simplicity and clarity, although in Appendix (ref), we restate the results of the paper in vectorised form.
Expression ((ref)) shows that $\tilde{\alpha}^*_L$ is asymptotically random, unlike under standard asymptotics where $n\hat{\alpha}_L \overset{p}{\to} 0$. $\tilde{\alpha}^*_L$ is the minimum eigenvalue of a general noncentral correlated Wishart matrix. By definition, $\tilde{\alpha}_L^*$ is defined as the limit of
i.e. the smallest eigenvalue of $(\frac{1}{n}W'W)^{-1/2}(W'P_ZW)(\frac{1}{n}W'W)^{-1/2}$. Then
where $\xi_1 = Q_{ZZ}^{-1/2}(Q_{ZZ}C + \Psi^*_{ZV})$ and $\xi_2 = Q_{ZZ}^{-1/2}\Psi^*_{Zu}$. By Assumption (ref), the $k_z \times (k_x+1)$ matrix $\Xi = [(\xi_1\beta + \xi_2) \ \ \xi_1 ]$ is distributed $\Xi \sim N(\mathcal{M}_{\xi}, \mathcal{V}_{\xi})$ for $\mathcal{M}_{\xi} = Q_{ZZ}^{1/2}CB$ and $\mathcal{V}_{\xi}$ dependent on the specific structure of the heteroskedasticity. It therefore follows that
where $W_q(p,V,\Theta)$ denotes the non-central Wishart distribution with dimension $q$, $p$ degrees of freedom, scaling matrix $V$ and non-centrality matrix $\Theta$, and $\Lambda=\mathcal{M}_{\xi}'\mathcal{M}_{\xi}$. We can also see that
Therefore, $\tilde{\alpha}^*_L$ is the distribution of the minimum eigenvalue of the limit of the random matrix $(\frac{1}{n}W'W)^{-1/2}(W'P_ZW)(\frac{1}{n}W'W)^{-1/2}$ with distribution
Clearly, the structure of the heteroskedasticity plays an important role in this distribution, although general statements are more difficult here than in staiger1997 due to the loss of the Kronecker variance structure. Further, there are no results for the distribution of extreme eigenvalues of general non-central, correlated Wishart matrices to the best of our knowledge, and deriving results on the distribution of extreme eigenvalues of such matrices is beyond the scope of this paper. This makes precise statements about the behaviour of LIML impossible. We can however see from inspection that $\tilde{\alpha}^*_L$ will be dependent on instrument strength, instrument numerosity and the endogeneity and heteroskedasticity of the model.
This lemma gives a minor adaptation of Lemma 1 from montiel2013, who give the limiting distributions of 2SLS and LIML in a heteroskedastic model with orthonormalised instruments i.e. $Z'Z/n = I_{k_z}$. Although 2SLS and LIML are invariant to this transformation, ((ref)) and ((ref)) do not require the orthonomalising transformation and so the result can be applied to any data satisfying the assumptions specified in Lemma (ref). The expressions are complicated mixtures and ratios of normals, similar to the limiting distributions given by staiger1997. Both the spread and location of the 2SLS and LIML estimators are affected by the structure of the heteroskedasticity, which cannot be consistently estimated under weak-instrument asymptotics. Before moving to the limiting distribution of the $J$ and $KP$-tests, we need the limiting distributions of the first-stage parameter estimators, as these feature in the limits of the variance terms in the test statistics.
The 2SLS estimator for $\Pi$ is asymptotically matrix-normal and with location matrix $C$. Therefore, while the estimator is consistent, it is asymptotically biased. $\tilde{\Pi}^*_L$ is difficult to analyse due to the second term in the expression subtracted from $\tilde{\Pi}^*_{2SLS}$, given that this term is a complicated ratio of random variables with the nonstandard distribution $\tilde{\beta}^*_L$ entering in a nonlinear manner (note that $\hat{\beta}_L$ and $\hat{\Pi}_L$ are estimated simultaneously, whereas of course $\hat{\Pi}_{2SLS}$ is estimated in the first stage and then $\hat{\beta}_{2SLS}$ is estimated in the second stage).
Now we derive the limiting distribution of the robust score test under heteroskedastic weak instruments, which will nest the weak-instrument limits for the $J$- and $KP$-tests. For the limiting distribution of the robust score test, let $\tilde{Z}_2=M_{\hat{X}}Z_2$ (with $i^{th}$ row $\tilde{z}_{2i}'$).
The first part of Assumption (ref) gives general limits for terms that will appear in the decomposition of the denominator of the robust score test. The second part of Assumption (ref) defines two other convergence results that will be important for the limiting distributions in Theorem (ref).
All three terms $\Omega_{\tilde{Z}_{2},u}$, $\Omega_{\tilde{Z}_{2},uV\tilde{\beta}}$ and $\Omega_{\tilde{Z}_{2},\tilde{\beta} V\tilde{\beta}}$ are random matrices by Lemma (ref) and will result in our test statistics becoming ratios of distributions asymptotically, analogous to the staiger1997 results for the homoskedastic Sargan test. The matrix $\tilde{Z}_2'\tilde{Z}_2/n$ is asymptotically non-degenerate with weak instruments. This follows since
and so therefore
which depends nonlinearly on the random matrix normal variable $\tilde{\Pi}^*$. If the model is estimated via LIML, which causes $\tilde{\Pi}^*$ to have a particularly nonstandard distribution, the resulting expression in ((ref)) will be highly nonstandard. The specific structure of the heteroskedasticity will dictate how these nonlinearities in $\tilde{\Pi}^*$ interact, and general statements are difficult. If homoskedasticity was imposed, then e.g. $\Omega_{\tilde{Z}_{2},u} = \sigma_u^2[Q_{22} - Q_2'\tilde{\Pi}^*(\tilde{\Pi}^*Q_{ZZ}\tilde{\Pi}^*)^{-1}\tilde{\Pi}^*Q_2]$. With these components, we now state the main theoretical result of the paper:
All components of the expression except $Q_2'C$ are random under weak-instrument asymptotics and ((ref)) is a highly non-standard distribution. In light of our previous discussion, the above limiting distribution nests both the behaviour of the $J$- and $KP$-tests. Therefore, we have the following corollary:
Whether the $J$- or $KP$-test performs better is not obvious from inspection of the limiting distributions and will be dependent on model attributes such as the strength of endogeneity and heteroskedasticity. Assessing the performance of these tests will be the focus of the next section, where we discuss Monte Carlo results.
The baseline model is given by
We assume $y_i,x_i,z_{1,i},...,z_{k_z,i},u_i,v_i\in\mathbb{R}$ are scalars for $i \in\{1,...,n\}$, with $n=120$ for all simulations. For simplicity, we set $\beta_0 = \beta_1 = \pi_0 =0$. The instruments are independent standard normals and are assumed to be equally informative, such that $\pi_j = \pi = c_0/\sqrt{n}$ for all $j\in\{1,...,k_z\}$. The errors are generated according to $u_i = z_{1,i}^{\alpha} u_i^*$ and $v_i = z_{1,i}^{\alpha} v_i^*$ for $(u_i^*,v_i^*)$ i.i.d. jointly normal with unit variances and correlation coefficient $\rho$. The strength of the heteroskedasticity is determined by $\alpha\in\mathbb{R}_{+}$ and is strictly increasing in the parameter ($\alpha=0$ represents homoskedasticity). We vary $\alpha\in\{0.5,1,1.5\}$ to represent weak, medium and strong heteroskedasticities. Different degrees of endogeneity are considered with $\rho\in\{0.2,0.5,0.95\}$. By varying $\mu^2$ across the interval $[0,32]$, we can see the evolution of test statistic performance as instrument strength increases from extremely weak to strengths commonly seen in applied practice. The number of instruments (not including the constant) is set to $k_z\in\{2,4\}$, leading to 1 and 3 degrees of overidentification. These degrees of overidentification coincide with those found in the two empirical applications in Sections (ref) and (ref).\footnote{Simulations with $k_z=6$ and $k_z=11$ were also conducted for this design; the increases in the size distortions when increasing the degrees of overidentification (particularly for $J$ but also for $KP$ to some extent) happen in a predictable manner. The qualitative comparison of the tests is unchanged. We therefore do not report these results.} Nominal size is calculated as the proportion of experiments for which each test exceeds the 5% critical value for the $\chi^2(k_z-k_x)$ distribution i.e. the correct asymptotic critical value under strong identification. In Appendix (ref), we further report the median bias and 90:10 percentile ranges for the 2SLS and LIML estimates of $\beta_1$, as well as additional size results at the 10% and 1% level for $J$ and $KP$ across a range of parameter configurations.
Figure (ref) presents rejection frequencies of the $J$- and $KP$-tests under the null at the nominal 5% level. In the low endogeneity design with $k_z=2$, test performance is similar under the weak and medium strength heteroskedasticity designs. The $J$-test size is closer to the nominal level for very weak instruments, although both tests still under-reject to significant degrees. It appears as though both tests get closer to nominal size for smaller $\mu^2$ when $\rho = 0.5$ as opposed to $\rho = 0.2$. $KP$ perhaps performs marginally better when $\rho=0.2$ and $J$ performs marginally better when $\rho=0.5$, although the difference is small. The high heteroskedasticity design causes $J$ to under-reject quite severely, whereas for $KP$, the test slightly over-rejecting as $\mu^2$ increases. Performance is largely unaffected by increasing the number of instruments to $4$. In the medium endogeneity design, test performance is again fairly similar across both instrument numbers considered and across different heteroskedasticity strengths, with $J$ slightly over-rejecting and $KP$ slightly under-rejecting. All tests are undersized for $\mu^2<15$ in the three lower endogeneity designs. When $k_z=2$, the two test statistics have similar rejection frequencies for the weak and medium heteroskedasticity designs, differing only significantly in the strong heteroskedasticity design; $J$ somewhat under-rejects even at $\mu^2=32$, whereas $KP$ slightly over-rejects.
The difference between test sizes is stark in the high endogeneity case. In the $k_z=2$ case, $KP$ performs well for $\mu^2>4$, with correct size obtained in the weak and medium heteroskedasticity designs. The test over-rejects slightly in the strongly heteroskedastic case but size distortions are relatively small. However, the $J$-test massively over-rejects, with the problem worsening with increasing strength of heteroskedasticity. In the weakest heteroskedasticity design, the size of the $J$-test peaks at $\approx0.2$ for $\mu^2=2$, before falling to $\approx 0.11$ for $\mu^2=10$. This problem gets increasingly worse as the strength of the heteroskedasticity increases; in the strong design, size peaks at $\approx 0.25$ for $\mu^2=3$, and even at $\mu^2=32$ (typically considered as relatively strong instruments), size is still at $\approx 0.09$. The problem also increases in the degrees of overidentification. For the $KP$-test, performance remains roughly consistent over the different degrees of overidentification, with the size remaining similar for weak and medium heteroskedasticity, whereas for $k_z= 4$, the over-rejection for $J$ in Panel (f) is obvious on inspection.
Figure (ref) shows the rejection frequencies of the test statistics for different values of $\mu^2$ for $\rho\in[-1,1]$, with $\alpha = 1$. When instruments are extremely weak for $k_z=2$, the $J$-test is closer to nominal size than $KP$ for smaller values of $|\rho|$, but for high endogeneity designs size increases sharply, tending toward 0.35, whereas $KP$ does not suffer from over-rejection. For $\mu^2=1$, performance of the two tests is similar for approximately $|\rho|\leq 0.6$, but for larger values of $|\rho|$, the $J$-test rejection frequency starts increasing sharply (at a rate increasing in $|\rho|$). The $KP$-test is clearly superior when we consider either $\mu^2 = 8$ or $\mu^2 = 16$; in the latter case, $KP$ exhibits almost perfectly correct size across all $\rho$ considered, whereas $J$ under-rejects for small $|\rho|$, and over-rejects for large $|\rho|$. A similar pattern is seen in Panel (b) for $k_z=4$, but the increase in size for the $J$-test is much more dramatic. Although the $KP$-test under-rejects more in this model than with $k_z=2$, the difference is not nearly as pronounced as the sensitivity of $J$ to the number of overidentifying restrictions. $KP$ does however have an under-rejection problem for very weak instruments. On balance, Figure (ref) provides evidence in favour of $KP$ when considering the whole parameter space for $\rho$. In particular, it appears that the costs of using $J$ tend to be greater when it is chosen for a model that it is ill-suited to relative to $KP$; $J$ tends to offer only a minor improvement over $KP$ in the models where it is more favourable, whereas $KP$ is sometimes dramatically superior to $J$.
The costs of using $J$ tend to be greater when it is chosen for a model that it is ill-suited to relative to $KP$; $J$ tends to offer only a minor improvement over $KP$ in the models where it is more favourable, whereas $KP$ is sometimes dramatically superior to $J$. From e.g. Figure (ref), $KP$ has an edge over $J$ in a less weakly-identified setting when instruments are perhaps somewhat weak but not extremely weak (as is often the case in empirical work).\footnote{hansen2008 find the median concentration parameter to be 23.6 in their study of applied microeconometric papers, and at this level of instrument strength, the $KP$-test typically outperforms the $J$-test (with this difference often large).} The structure of the heteroskedasticity and degree of endogeneity are important factors when comparing $J$ and $KP$, but these cannot be consistently estimated in a weak instrument case. This leads to the recommendation based on simulation evidence that the $KP$-test should be employed by applied researchers when testing overidentifying restrictions.
In the lifecycle consumption model, an agent maximises their expected lifetime utility by smoothing consumption across an infinite time horizon. epstein1989 and weil1989 propose a generalised class of utility functions $U_t$, defined recursively as
where $\mathbb{E}_t[ \, \cdot \, ]$ denotes the conditional expectation operator w.r.t. the information set at time $t$, $C_t$ is consumption at time $t$, $\delta$ is an intertemporal discount factor and $\varphi = (1 - \gamma)/(1-1/\psi)$. Here, the parameter $\psi$ represents the elasticity of intertemporal substitution and $\gamma$ is the coefficient of relative risk aversion (CRRA).\footnote{When $\varphi = 1$, ((ref)) collapses to the power utility function, leading to the well-known inverse relationship that the EIS is the reciprocal of the CRRA e.g. $\psi = 1/\gamma$.} Following campbell2003 and yogo2004, the individual then maximises ((ref)) w.r.t. the budget constraint
where $W_{t+1}$ is total household wealth at time $t+1$ and $R_{w,t+1}$ is the gross real return on total asset investments at time $t+1$. Solving this optimisation problem yields the Euler equation
where $R_{t+1}$ is the return on the asset in question. By imposing homoskedasticity and joint log-normality on consumption and asset returns (conditional on information available at time $t$), and log-linearising ((ref)), we can derive the return on a riskless asset as
where for the generic variable $A_{t+1}$ we denote $a_{t+1} = \ln A_{t+1}$, $\Delta a_{t+1} = a_{t+1} - a_t$ and $\sigma^2_a$ is the unconditional variance of $\{a_t\}$. Similarly, the risk-premium on risky assets is
Now define the constant terms $\mu_f = [-2\ln \delta + (\varphi-1)\sigma_w^2 - \varphi\sigma_c^2/\psi^2]/2$ and $\mu_r = \mu_f - \sigma_r^2/2 + \varphi\sigma_{rc}/\psi + (1-\varphi) \sigma_{rw}$. Then, ((ref)) can be re-written as
From ((ref)), we can derive estimable regression equations. Define the error $\eta_{t+1} = r_{t+1} - \mathbb{E}_t[r_{t+1}] - \frac{1}{\psi}(\Delta c_{t+1} - \mathbb{E}_t[\Delta c_{t+1}])$. By adding $\eta_{t+1}$ to both sides, ((ref)) becomes
so substituting in the explicit formula for $\eta_{t+1}$ on the L.H.S. of ((ref)) yields
which defines a regression equation in terms of the observed $r_{t+1}$ and $\Delta c_{t+1}$, the unobserved $\eta_{t+1}$ and the unknown constant and slope parameters $\mu_r$ and $1/\psi$ to be estimated. ((ref)) can also be re-normalised in terms of $\Delta c_{t+1}$, yielding the regression equation
where $\mu_c = -\psi \mu_r$ and $u_{t+1} = -\psi\eta_{t+1} =\Delta c_{t+1} - \mathbb{E}_t[\Delta c_{t+1}] - \psi(r_{t+1} - \mathbb{E}_t[r_{t+1}])$. We consider ((ref)) to be the primary regression of interest, as the EIS is usually considered the main parameter of interest. Estimation of this parameter has received substantial attention in the literature hansen1983,hall1988, although a large variety of proposed plausible values have been given.
OLS is inappropriate for both regression equations, as clearly the explanatory variables will be endogenous; the errors are defined as explicit functions of the regressor. However, $u_{t+1}$ and $\eta_{t+1}$ are both errors in expectations of future variables, meaning that they are orthogonal to variables known at time $t$ nakamura2021. Given this, a wide range of lagged macroeconomic variables have been used as instruments in the literature. Suppose we have a $k_z$-vector $Z_{t+1}'$ of lagged macroeconomic indicators $(\, \Tilde{z}_{1,t-1} \,\,\, \Tilde{z}_{2,t-1} \,\,\, ... \,\,\, \Tilde{z}_{k_z,t-1} \, )$ as instruments (e.g. the $k^{th}$ instrument $z_{k,t+1}$ is formed by twice-lagging the $k^{th}$ macroeconomic indicator $\Tilde{z}_{k,t+1}$), stacked into the $T\times k_z$ matrix $Z$. The IV model then implies the moment restrictions $\mathbb{E}_t[Z_{t+1}\eta_{t+1}] = 0$ and $\mathbb{E}_t[Z_{t+1}u_{t+1}] = 0$. From the structure of the error terms $\eta_{t+1}$ and $u_{t+1}$, it is clear that the restrictions $\mathbb{E}_t[Z_{t+1}\eta_{t+1}] = 0$ and $\mathbb{E}_t[Z_{t+1}u_{t+1}] = 0$ are equivalent up to linear transformations yogo2004.
Heteroskedasticity is an important feature of the model, as it has the interpretation of representing precautionary savings. Under heteroskedasticity, yogo2004 shows that the covariance and variance terms that enter the constants in the regressions simply need be replaced by their conditional counterparts. By appropriately redefining the error terms $u_{t+1}$ and $\eta_{t+1}$, we can arrive at the moment restrictions to identify $\psi$ and $1/\psi$. To see this, consider the risk-free log-linearised Euler equation $r_{f,t+1} = \mu_f + \frac{1}{\psi}\mathbb{E}_t[\Delta c_{t+1}]$ for simplicity, but allow for time-varying conditional variances such that
Then, yogo2004 shows that as long as $\mathbb{E}_t[Z_t(\mu_{f,t}-\mu_f)] = 0$, then $1/\psi$ is still identified by the moment condition $\mathbb{E}_t[Z_{t+1}\eta_{t+1}]=0$, with an equivalent result available for the moment condition $\mathbb{E}_t[Z_{t+1}u_{t+1}]=0$ identifying $\psi$.
To test the validity of the moment restrictions implied by ((ref)), we want to test
and for ((ref)), we want to test
which is possible as long as $k_z > 1$. Whilst standard practice for testing ((ref)) and ((ref)) under strong identification is well-established in the literature, we note that weak instruments are a well-known problem in the estimation of the EIS yogo2004,neely2001, as are questions about the validity of the instruments, which stem from rejections of the overidentifying restrictions from $J$-tests epstein1991,gomes2011,gomes2013,dacy2011,pakos2011. Given this and the empirical relevance of heteroskedasticity in our model, we suggest that standard practices regarding overidentification testing may not be suitable for this application.
The dataset used for this application comes from yogo2004, and is commonly used as an empirical application in the weak instrument literature montiel2013,andrews2016. This dataset consists of quarterly data on stock markets at the aggregate level, as well as macroeconomic variables from 11 countries: Australia (AUS), Canada (CAN), France (FRA), Germany (GER), Italy (ITA), Japan (JAP), Netherlands (NTH), Sweden (SWD), Switzerland (SWT), the United Kingdom (UK) and the United States of America (USA). The stock market data come from Morgan Stanley Capital International, and the consumption and interest rate data come from the International Financial Statistics of the International Monetary Fund. For the USA, consumption of nondurables and services is measured in the dataset, but for every other country, only data on total consumption are available. Due to data constraints, the time periods vary across countries, with most time series beginning in 1970, although the data for the USA stretches back to 1947. On top of the consumption and interest rate data, the dataset contains variables for a vector of instruments, consisting of nominal interest rate, inflation, consumption growth and log dividend-price ratio, all of which are twice-lagged. See yogo2004 for a more complete description of the dataset.
Table (ref) reports estimation and test statistic results from the 11 countries. Panel (a) reports results using model ((ref)) i.e. where the EIS $\psi$ is the parameter of interest, and Panel (b) reports results using model ((ref)), where the reciprocal of the EIS $1/\psi$ is the parameter of interest. Column 3 presents the $F_{eff}$-statistic from montiel2013, the current recommended statistic for testing weak instruments under non-homoskedastic errors. $\kappa$ is the simplified, conservative 95% critical value - given the effects that weak instruments have on estimation and inference, we choose the strongest critical value from montiel2013 as the benchmark to rule out weak instruments. Columns 5-6 reports estimates of $\psi$ in Panel (a) and $1/\psi$ in Panel (b), estimated via 2SLS and LIML respectively. Columns 7-8 report the $J$ and $KP$ test statistic values for hypotheses ((ref)) and ((ref)). For convenience, we highlight in bold $F_{eff}$-statistics that fail to reject the null of weak instruments, and values of $J$ and $KP$ that are higher than the critical value $\chi^2_{0.95}(3) = 7.815$.
{6pt}
In Panel (a), we immediately see that both $J$ and $KP$ reject the null of exogeneity for Australia, and fail to reject the null for all other countries. There is some evidence of weak instruments, with 7 of the 11 countries having an $F_{eff}$-statistic lower than the respective conservative 95% critical values, with many of the other specifications attaining $F_{eff}$-statistics only slightly above the critical value. Only the data from France give an $F_{eff}$-statistic substantially over the critical value. It is worth however noting that we are requiring the strongest evidence by using the conservative threshold; if we use the LIML critical values, only 4 specifications would fail to reject the null of weak instruments at the conventional level. However, 2SLS and LIML provide similar point estimates, and the test statistics give similar values in each case too, which typically suggests that identification is not particularly weak, in spite of the $F_{eff}$ values. It appears that both the $J$- and $KP$-tests still perform reasonably well on the whole; as an informal back-of-the-envelope calculation, assuming correctly-sized tests and the null holding in all 11 specifications (with independence between specifications), the probability of receiving at least one false rejection at the 5% level is 43%. This suggests that the single rejection for each test can be reasonably explained as a false rejection. We conclude that Panel (a) therefore fails to provide strong evidence against instrument validity and/or correct specification.
Panel (b) gives a different picture. Every country fails to provide evidence against a large weak-instrument problem (the largest $F_{eff}$-statistic is less than 3). The $J$-test rejects in 6 of the 11 specifications, despite the fact that we have good reason to believe that the instruments are exogenous by construction. Perhaps more importantly, we know that when identification is not weak, $J$ and $KP$ are typically numerically similar and have good size properties. Therefore, given the results from Panel (a) and knowing that that the moment restriction holding in ((ref)) implies it holds in ((ref)), then Panel (b) gives strong evidence of the over-rejection properties found under weak instruments in the Monte Carlo simulations of Section (ref). This is also suggestive of the fact that the degree of endogeneity is likely higher when $\Delta c_{t+1}$ is treated as the endogenous regressor as opposed to $r_{t+1}$.\footnote{ Given the small sample size combined with the weak instrument problem, it is impossible to obtain estimates of the degree of endogeneity that could be considered reliable.} When we look at $KP$, we immediately see the numerical equivalence between the values in Panel (a) and Panel (b) for each country, which follows from the invariance to normalisation of LIML and $KP$.\footnote{ Also note that the LIML estimates are exact reciprocals, whereas this is not the case for 2SLS.} From this, it is clear that in both specifications, the $KP$-test will give the same conclusion for a particular country regardless of whether we have $\psi$ or $1/\psi$ as our unknown parameter, which is potentially a useful property when we consider that the null in ((ref)) implies the null in ((ref)) and vice versa. Under the assumption of power utility functions, this invariance could be advantageous in helping us obtain more reliable estimates of the CRRA (as under power utility it is the inverse of the EIS), despite the significant weak-instrument problem present. We find that $KP$ does not suffer from the same over-rejection problem in Panel (b), and can be considered more reliable for inference than the $J$ counterpart. Given the results of Panel (a), Panel (b) suggests that $J$ is over-rejecting rather than $KP$ falsely failing to reject.
Our results provide a different explanation to the small-scale replication of gomes2011, who replicate yogo2004 but report homoskedastic Sargan tests. They reject the null hypothesis at the 5% level for 4 countries and 10% level for a further two. Given this, they raise doubts about the validity of the instruments and therefore the model specification. However, they do not consider the effects of non-homoskedastic errors or discuss the effects of weak instruments on the size of 2SLS-based overidentification tests. Our results suggest that when appropriately accounting for these two factors, we fail to find evidence against the overidentifying restrictions. The recommendation from this empirical exercise is to $KP$ for overidentification testing, and that instrument validity and/or misspecification of the moment conditions are unlikely to be the cause of the variation of proposed estimates for $\psi$ seen in the literature.
In this paper we have studied the $KP$-test as a test for overidentifying restrictions and provided evidence for its usefulness relative to the standard $J$-test commonly employed in econometric packages. We have derived the limiting distribution of the robust score test, nesting the $J$- and $KP$-tests, under heteroskedastic weak instruments, and find in general that the $KP$-test is favourable, especially when the degree of overidentification is small. This conclusion follows for multiple reasons: firstly, the size distortions for $KP$ tend to be less severe than for $J$ in the models where each performs poorly relative to other e.g. the extreme over-rejection of $J$ in high endogeneity models. Although $J$ seems to perform slightly better under extremely weak identification, both tests have highly nonstandard distributions and neither can be considered reliable here. Under moderately weak identification at levels commonly seen in the applied literature, $KP$ outperforms $J$ across a range of models. Secondly, $KP$ seems to be much less dependent on the strength of the heteroskedasticity; especially when $\rho$ is high, $J$ varies considerably as the strength of the heteroskedasticity changes (implying that this is important information for a researcher to consider about their model), and since the heteroskedastic error structure cannot be consistently estimated under weak instruments, this is problematic. The $KP$-test can therefore be seen as more “robust" to variation in endogeneity or heteroskedasticity strength under weak identification. We therefore recommend its use for applied researchers. This recommendation also provides a generalisation of the guidance from staiger1997, who recommend using LIML-based overidentification tests under homoskedasticity.
Further, we have empirically assessed the $J$- and $KP$-tests and re-examined the validity of the over-identifying restrictions in the estimation of the EIS in the lifecycle consumption model, where weak instruments and heteroskedasticity are common issues. As expected, we provide evidence that the $J$- and $KP$-tests are extremely similar in performance when instruments are strong, with both tests failing to find evidence to reject the overidentifying restrictions. However, when the instruments are possibly weak, then $J$ is empirically much more likely to reject these restrictions than $KP$. With this in mind, we suggest that a large number of papers are possibly rejecting the overidentifying restrictions erroneously due to the poor properties of the $J$-test under weak instruments, rather than finding evidence that the instruments are invalid. Given the previous theoretical results and simulations, as well as the natural exogeneity expected of instruments formed from lagged macroeconomic variables in this model, we argue that this is evidence of the $J$-test having a substantially inflated size under weak instruments, a problem we find not present in the $KP$-test (with this conclusion supported further by the additional application in the Appendix). We therefore suggest two key implications: from an econometric perspective, we provide empirical evidence of the suitability of the $KP$-test as an overidentification test over $J$-test in situations with weak identification and heteroskedastic errors. From a macroeconomic perspective, we suggest that previous evidence found in the literature of possible instrument invalidity stems from issues with the $J$-test, rather than due to the instruments and specification of the model commonly used by empirical macroeconomists. Therefore, instrument validity does not seem to be a plausible factor in the large range of estimates of the EIS presented in the literature.
\printbibliography