EconBase
← Back to paper

Inference in clustered IV models with many and weak instruments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

62,351 characters · 15 sections · 89 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inference in clustered IV models with many and weak instruments

abstractData clustering reduces the effective sample size from the number of observations towards the number of clusters. For instrumental variable models this reduced effective sample size makes the instruments more likely to be weak, in the sense that they contain little information about the endogenous regressor, and many, in the sense that their number is large compared to the sample size. Consequently, weak and many instrument problems for estimators and tests in instrumental variable models are also more likely. None of the previously developed many and weak instrument robust tests, however, can be applied to clustered data as they all require independent observations. Therefore, I adapt the many and weak instrument robust jackknife Anderson--Rubin and jackknife score tests to clustered data by removing clusters rather than individual observations from the statistics. Simulations and a revisitation of a study on the effect of queenly reign on war show the empirical relevance of the new tests. \\[1ex] Keywords: weak instruments, many instruments, clustered data, jackknife.\\ JEL codes: C12, C26, N43.

Introduction

The wide spread use of clustered standard errors in instrumental variable (IV) models shows that studies that use IVs to identify a coefficient on an endogenous regressor often use data that are clustered, as opposed to independent data. In this paper I show that clustered data makes it more likely that the instruments are weak, in the sense that they contain little information about the endogenous regressor, or that they are many, in the sense that their number is large compared to the sample size. This is because the dependence between observations within the same cluster decreases the information in the sample, or put differently reduces the effective sample size. Consequently, weak and many instrument problems, such as unreliability of tests based on two stage least squares (2SLS), are also more likely when the data are clustered and makes the need for tests that are reliable in the presence of many and weak instruments more pressing.

Recently several tests that are robust against many and weak instruments have been proposed crudu2021inference,mikusheva2021inference,matsushita2020jackknife,dovi2022ridge,lim2024conditional,lim2024valid. These are tests for parameters on endogenous regressors in a linear IV model and allow for (i) many instruments in the sense that the number of instruments is a non-negligible fraction of the sample size bekker1994alternative, (ii) weak or irrelevant instruments as measured by a small or zero correlation between the endogenous regressor and the instruments, and (iii) heteroskedastic data. Although the tests can handle heteroskedastic data, they are not applicable to clustered data, as they require independent observations. An exception is the many and weak instrument robust test based on invariance by boot2023identification. However, this test, which build upon an earlier version of this paper, uses an invariance assumption on the data which limits its applicability.

In this paper I therefore extend the jackknife Anderson--Rubin anderson1949estimation test by crudu2021inference and mikusheva2021inference and the jackknife score test by matsushita2020jackknife to clustered data. I start with the jackknife AR test, which is itself an adaptation of the original AR test. The original AR test, although robust against weak instruments, has low power and can be oversized in case the instruments are numerous kleibergen2002pivotal,anatolyev2011specification. crudu2021inference and mikusheva2021inference resolve this issue by removing the terms in the AR statistic that have non-zero expectation. For independent data this can be done by jackknifing the AR statistic, hence yielding the jackknife AR statistic. If the data are clustered, then there are additional terms that have non-zero expectation, which I remove using a new cluster jackknife. This then yields the cluster jackknife AR statistic. Following an earlier version of this paper, the same jackknife also featured in frandsen2023cluster. I then show that the cluster jackknife can also be used to extend the jackknife score test matsushita2020jackknife to a cluster jackknife score test, which in some cases has better power than the cluster jackknife AR test.

To obtain the limiting distribution of the cluster jackknife AR and cluster jackknife score statistics I derive a central limit theorem (CLT) for bilinear forms of clustered random variables with respect to a projection matrix of growing rank. This CLT is a direct generalisation of Lemma A2 in chao2012asymptotic, and, given the widespread use of the lemma, is of independent interest. An alternative way to obtain the limiting distribution is the recently proposed CLT by mikusheva2025estimation, which generalises the CLT by solvsten2020robust to clustered data. Under stricter assumptions, which for example rule out conditional heteroskedasticity, a third option is the CLT by feng2025robust.

Next, I propose four extensions and improvements for the cluster jackknife AR and score tests. Firstly, the cluster jackknife AR and score statistics converge jointly, which allows them to be combined in a conditional linear combination test iandrews2016conditional,lim2024conditional to enhance the power.

Secondly, by leaving out terms from the AR and score statistics, information that can help to conduct inference on the parameters of interest is lost. bekker2015jackknife therefore propose an alternative jackknife procedure, called the symmetric jackknife, which aims at retaining part of this information. crudu2021inference use this alternative jackknife to omit the terms with non-zero expectation from the AR statistic and in that way obtain the symmetric jackknife AR statistic. I extend the symmetric jackknife procedure to clustered data and incorporate it in the cluster symmetric jackknife AR and score statistics. These statistics have a larger signal component about the parameters of interest than the cluster jackknife AR and score statistics.

Thirdly, chao2023jackknife argue that recentring statistics using the jackknife does not work well in case there are many exogenous control variables in the model, since removing these controls introduces dependence between the terms in the statistic, which causes more terms to have non-zero expectation. They therefore propose a different type of jackknife, which can handle many control variables. I also extend this type of jackknife to clustered data.

Fourthly, to enhance the power of the different cluster jackknife tests, I propose alternative variance estimators which transform the implied errors in the estimator by cross-fitting them newey2018cross,kline2020leave. This improves power as it removes the part that biases the variance estimator when the hypothesised value of the parameter of interest is far from its true value mikusheva2021inference.

I explore how the cluster robust tests perform in a Monte Carlo experiment in which the data are dependent within clusters. The cluster jackknife tests attain close to nominal size in a range of settings. When there are many instruments the two tests also have superior power relative to the cluster adaptations of the AR and score kleibergen2005testing tests.

To illustrate the practical relevance of the new tests, I revisit dube2020queens's (dube2020queens) study on the effect of having a queen on the likelihood for a state to be at war. The effect is identified through two instruments based on the family composition of European monarchs and estimated with a data set of 18 polities observed some time between 1480 and 1913, resulting in $3\,586$ polity-year observations. The data are clustered in reigns, which yields 176 clusters of varying size. In the appendix dube2020queens propose three additional instruments based on interactions of the two instruments with exogenous regressors. The authors note however that these interactions are relatively weak instruments. Moreover, as mentioned earlier in this introduction, because the data are clustered the effective sample size is reduced, which makes the number of instruments relatively large compared to the sample size. Therefore, many and weak instrument robust tests are required to conduct inference in the models that include the extra instruments, but since the data are not independent none of the previously developed tests can be applied. I continue the analysis with the new cluster robust tests and find that having a queen significantly increases the likelihood of an armed conflict.

Recently a discussion has emerged on the need to account for clustering in the data. See for example barrios2012clustering, chetverikov2023standard and abadie2023should. In this paper I abstract from this question and instead assume that clustering at the specified level is necessary and adequate.

The remainder of the paper is structured as follows. In (ref) I present the linear IV model for clustered data and argue why many and weak instruments are more likely in these types of models. (ref) discuss the cluster jackknife AR and the cluster jackknife score tests. The four extensions of these tests are covered in (ref). (ref) proceed with the simulation results and the empirical application. Finally, (ref) concludes.

Throughout I use the following notation. $\bs\iota$ is a vector of ones and $\bs e_i$ is $i^{\text{th}}$ unit vector. $\bs A_{(i)}=\bs A\bs e_i$ denotes the $i^{\text{th}}$ column of $\bs A$. I use the projection matrices $\bs P_{A}=\bs A(\bs A'\bs A)^{-1}\bs A'$ and $\bs M_{A}=\bs I-\bs P_{A}$. For $\odot$ the Hadamard product, I write $\bs D_{A}=\bs I\odot\bs A$ for $\bs A$ a square matrix. $\dot{\bs A}=\bs A-\bs D_{A}$ is the matrix with the diagonal elements set to zero. For $\bs A$ a square matrix $\operatorname{\lambda_{\min}}(\bs A)$ and $\operatorname{\lambda_{\max}}(\bs A)$ are the minimum and maximum eigenvalues of $\bs A$. For $\bs A$ not necessarily square $\|\bs A\|_2=\sqrt{\operatorname{\lambda_{\max}}(\bs A'\bs A)}$ is the spectral norm. Let $\otimes$ be the Kronecker product and $\operatorname{vec}(\bs A)$ be the column vectorisation of a matrix. I denote convergence in distribution by $\rightsquigarrow$, convergence in probability by $\overset{p}{\rightarrow}$ and almost sure (a.s.) convergence by $\overset{a.s.}{\rightarrow}$. $\sum_{g\neq h}$ can denote the double sum $\sum_{g=1}^G\sum_{h=1,h\neq g}^H$, and similarly for other types of sums. Whether this is the case and the maxima of the sums depend implicitly on the context. Finally, $C$ denotes a finite positive scalar that is not necessarily the same in each appearance.

Model

Consider the linear IV model

equation[equation omitted — 167 chars of source]

Here $y_i$ is a scalar outcome, $\bs X_i\in\mathbb{R}^{p}$ a vector of endogenous variables, $\bs Z_i\in\mathbb{R}^k$ a vector of IVs and $\varepsilon_i\in\mathbb{R}$ and $\bs\eta_i\in\mathbb{R}^p$ are the second and first stage errors, $i=1,\dots,n$. Assume for now that the exogenous covariates are small in number and have been partialled out. The inclusion of a large number of exogenous covariates is deferred to (ref). Stack the observations in $\bs y=(y_1,\dots,y_n)'$, and similarly for the other variables. The model in (ref) then becomes

equation[equation omitted — 157 chars of source]

The data are divided over $G$ clusters with sizes $n_g$, $g=1,\dots, G$ that are not necessarily equal. Assume that the observations have been ordered per cluster. Denote the indices corresponding to cluster $g$ by $[g]$ and for any $n\times m$ matrix or vector $\bs A$ let $\bs A_{[g]}$ be the $n_g\times m$ matrix with only the rows of $\bs A$ indexed by $[g]$ selected. Also if $m=n$, let $\bs A_{[g,h]}$ be the $n_g\times n_h$ matrix with only the columns of $\bs A_{[g]}$ indexed by $[h]$ selected. I make the following assumptions on the data.

assumptionConditional on $\bs Z$, $\{\bs\varepsilon_{[g]},\bs\eta_{[g]}\}_{g=1}^G$ is independent with mean zero, $\operatorname{E}(\bs\varepsilon_{[g]}\bs\varepsilon_{[h]}'|\bs Z)\allowbreak=\bs\Sigma_{g}$, $\operatorname{E}(\bs\eta_{(i),[g]}\bs\eta_{(j),[h]}'|\bs Z)=\bs\Omega^{(i,j)}_g$, and $\operatorname{E}(\bs\eta_{(i),[h]}\bs\varepsilon_{[g]}'|\bs Z)=\bs\Xi^{(i)}_g$ if $g=h$ and zero otherwise, $g,h=1,\dots, G$, $i,j=1,\dots p$.

This assumption formalises the cluster structure. It requires the observations between clusters to be independent, but allows for dependence of observations within clusters. As a consequence of this dependence, a clustered sample generally contains less information than a same sized sample of independent data. Thus clustering decreases the effective sample size down from the number of observations towards the number of clusters. It is this lower effective sample size that makes the problems with many and weak instruments more likely for clustered data than for independent data.

Many instruments are modelled as the number of instruments being large relative to the sample size. The reduced effective sample size under clustering makes the number of instruments more quickly large compared to the effective sample size. Hence, clustering makes the problems associated with many instruments, such as biased estimators and oversized tests, more likely than when the data are independent.

Similarly, for independent data weak instruments are modelled as the first stage coefficient $\bs\Pi$ in (ref) shrinking to zero at a rate of square root the sample size. Again, the reduced effective sample size under clustering makes that $\bs\Pi$ can shrink at a lower rate for clustered data, making weak instrument problems of biased estimators and oversized tests more likely to occur.

The argument that many and weak instrument problems are more likely with clustered data is more formally worked out in (ref).

In the next sections, I propose two tests that are robust against many and weak instruments and that can handle clustered data. To be robust against weak or even irrelevant instruments, I consider identification robust tests that test the null hypothesis $H_0\colon\bs\beta_0=\bs\beta$ against the alternative $H_1\colon\bs\beta_0\neq\bs\beta$. To be robust against many instruments, I study their behaviour under many instrument asymptotics, in which the number of instruments is proportional to the sample size.

Cluster jackknife AR

The first test robust against many and weak instruments and clustered data that I propose is the cluster jackknife AR test. To understand its functioning it is best to start with the original AR statistic.

For a given value of $\bs\beta$, that is not necessarily equal to $\bs\beta_0$, let $\bs\varepsilon(\bs\beta)=\bs y-\bs X\bs\beta$. Then, under certain conditions, the original AR statistic for the linear IV model converges as

equation[equation omitted — 200 chars of source]

where $\bs V_{AR}(\bs\beta)=\operatorname{var}(\bs Z'\bs\varepsilon(\bs\beta)/\sqrt{n}|\bs Z)$ captures the variance of the moment conditions.

The original AR statistic is robust against weak instruments, but it has some problems in case the instruments are numerous. For homoskedastic data and if the number of instruments is a non-negligible fraction of the sample size, the test based on the feasible version of this statistic has low power kleibergen2002pivotal and can be oversized anatolyev2011specification. Furthermore, if in addition to many instruments, the data are heteroskedastic, then the estimation of the covariance matrix, $\bs V_{AR}(\bs\beta)$, is complex due to the large number of variances and covariances. These estimation problems can add to the size and power problems.

crudu2021inference and mikusheva2021inference therefore adapt the AR statistic in two ways to make it suited for many instruments and heteroskedasticity when the data are independent. First, they substitute $\bs V_{AR}(\bs\beta)$ by the homoskedasticity inspired weighting matrix $\bs Z'\bs Z$. This matrix does not depend on the variances and covariances of the errors and is therefore easier to handle in the asymptotic approximations.

Second, crudu2021inference and mikusheva2021inference note that the $\chi^2_k$ distribution has many degrees of freedom in case there are many instruments. This makes the distribution spread out and causes low power. The $\chi^2_k$ distribution, however, can be changed to a normal distribution if the statistic is properly centred at zero. To do so the terms in the AR statistic that involve squares of the errors and therefore have non-zero expectation need to be removed. This is done efficiently by jackknifing, and which boils down to setting the diagonal elements of the matrix that weighs the errors, $\bs P_{Z}=\bs Z(\bs Z'\bs Z)^{-1}\bs Z'$, to zero.

After rescaling and under certain assumptions, crudu2021inference and mikusheva2021inference then show that chao2012asymptotic's (chao2012asymptotic) central limit theorem can be applied to their jackknife AR statistic, such that it converges as

equation[equation omitted — 238 chars of source]

where $V_{J}(\bs\beta)=\operatorname{var}(\bs\varepsilon(\bs\beta)'\dot{\bs P}_{Z}\bs \varepsilon(\bs\beta)/\sqrt{k}|\bs Z)$.

To extend to jackknife AR to clustered data, note that when the data are clustered not only the terms involving squares of the errors have non-zero expectation, but any product of errors within the same cluster. Hence, to centre the AR statistic it is no longer sufficient to remove only the squares of the errors. Rather, all products of errors within a cluster need to be removed.

I remove all these terms by a cluster jackknife, which sets blocks on the diagonal of $\bs P_{Z}$ to zero, and that can be written more succinctly by introducing the following notation. For given cluster structure and any $n\times n$ matrix $\bs A$, let $\bs B_{A}$ be the $n\times n$ block diagonal matrix with on its diagonal the blocks $\bs A_{[g,g]}$, $g=1,\dots,G$. Also, denote $\ddot{\bs A}=\bs A-\bs B_{A}$ for the matrix with the blocks on its diagonal set to zero.

Using this notation I have that $\bs\varepsilon(\bs\beta_0)'\ddot{\bs P}_{Z}\bs\varepsilon(\bs\beta_0)/k$ is centred at zero. I need the following assumption to derive its distribution.

assumptionConditional on $\bs Z$ and with probability one for all $n$ large enough, it holds that \begin{assenumerate*} • $\operatorname{rank}(\bs P_{Z})=k$; • $\|\bs P_{Z,[g,g]}\|_2^2\leq C<1$; • $k\to\infty$ as $G\to\infty$; • $n_{\max}^6/k\overset{a.s.}{\rightarrow} 0$ for $n_{\max}=\max_{g=1,\dots,G}n_g$; • $0<\allowbreak 1/C\leq\allowbreak\operatorname{\lambda_{\min}}(\bs\Sigma_{g})\leq\allowbreak\operatorname{\lambda_{\max}}(\bs\Sigma_{g})\leq\allowbreak n_gC<\allowbreak\infty$ a.s. for all $g$; • $\operatorname{E}(\varepsilon_i^4|\bs Z)\leq C<\infty$ a.s. for all $i$. \end{assenumerate*}

(ref) implies the full rank assumption on $\bs Z$ and means that there are no redundant instruments. (ref) generalises crudu2021inference's (crudu2021inference) and mikusheva2021inference's (mikusheva2021inference) requirement that the diagonal elements of $\bs P_Z$ are less than one, which is a common assumption in the many instrument literature. It is also related to the measure of cluster leverage in the first stage regression by mackinnon2022leverage and bounds how different the clusters can be. (ref) specifies the asymptotic sampling scheme. It allows for many instruments in the sense that the number of IVs increases proportionally to the sample size and it states that the number of clusters tends to infinity. The fourth item of (ref) bounds the relative cluster sizes and ensures that there is not one cluster that dominates the sample. Such a requirement is common in the literature on clustered data and is, for example, also imposed by djogbenou2019asymptotic and hansen2019asymptotic. Their assumptions on the maximum cluster size are more general however. djogbenou2019asymptotic's (djogbenou2019asymptotic) requirement provides a trade-off between the number of moments of the data that need to exist and the degree of cluster heterogeneity that is allowed. With fourth moments, as I assume, their condition is met when $n_{\max}^3/n\to 0$. hansen2019asymptotic relate the maximum cluster size to the variance of the regressors and show that a central limit theorem holds when $n_{\max}^2/n\to0$. My condition is not conditional on the distribution of the data and stronger than these assumptions. (ref) are regularity conditions that require a positive, but bounded variance and a finite fourth moment. Under these assumptions I have the following result.

theoremUnder (ref) the cluster jackknife AR converges as \begin{equation} \begin{split} \frac{AR_{CLJ}(\bs\beta_0)}{\sqrt{V^{AR}_{CLJ}}}=\frac{1}{\sqrt{V_{CLJ}^{AR}k}}\bs\varepsilon(\bs\beta_0)'\ddot{\bs P}_{Z}\bs \varepsilon(\bs\beta_0)\rightsquigarrow N(0,1), \end{split} \end{equation} with conditional variance \begin{equation} \begin{split} V_{CLJ}^{AR}=\operatorname{var}(\frac{1}{\sqrt{k}}\bs\varepsilon'\ddot{\bs P}_{Z}\bs\varepsilon|\bs Z)=\frac{2}{k}\sum_{g\neq h}\operatorname{tr}(\bs\Sigma_{g}\bs P_{Z,[g,h]}\bs\Sigma_{h}\bs P_{Z,[h,g]}). \end{split} \end{equation}
proofThe result can be shown by adapting the proof of (ref) below.

Cluster jackknife score

The original AR statistic can have low power in overidentified IV models even when the number of IVs remains fixed. kleibergen2002pivotal therefore proposes a score test, which, similarly to the AR statistic is robust against weak instruments, but is more powerful. matsushita2020jackknife show that jackknifing can be used to also make the score test robust against many instruments and heteroskedasticity.

Similarly, the cluster jackknife from the previous section can be used to derive a cluster jackknife score test. For this I require the following assumption.

assumptionConditional on $\bs Z$ and with probability one for all $n$ large enough, \begin{assenumerate} • For any $\bs v\in\mathbb{R}^p$ such that $\bs v'\bs v=1$ and \begin{equation} \begin{split} \bs M_g(\bs v)=\operatorname{E}(\begin{bmatrix} \bs\varepsilon_{[g]}\bs\varepsilon_{[g]}' & \bs\varepsilon_{[g]}\bs v'\bs\eta_{[g]}' \\ \bs\eta_{[g]}\bs v\bs\varepsilon_{[g]}' & \bs\eta_{[g]}\bs v\bs v'\bs\eta_{[g]}' \end{bmatrix}|\bs Z), \end{split} \end{equation} it holds that $0<1/C\leq \operatorname{\lambda_{\min}}(\bs M_g(\bs v))\leq\operatorname{\lambda_{\max}}(\bs M_g(\bs v))\leq n_gC<\infty$; • $\operatorname{E}(\|\bs\eta_i\|^4|\bs Z)\leq C<\infty$; • $\operatorname{\lambda_{\max}}(\bs\Pi'\bs Z_{[g]}'\bs Z_{[g]}\bs\Pi/n_g)\leq C<\infty$. \end{assenumerate}

(ref) is a stronger version of (ref). It requires positive but finite variances and that the second stage errors are not perfectly correlated with a linear combination of the first stage errors. Taking a linear combination of the first stage errors facilitates the derivation of the distribution. (ref) requires finite fourth moments of the first stage errors. (ref) ensures convergence of the scaled instruments.

Then, define the cluster jackknife score statistic as

equation[equation omitted — 145 chars of source]

Write its conditional variance under the null hypothesis as

equation[equation omitted — 651 chars of source]

where $\bs\Sigma=\operatorname{E}(\bs\varepsilon\bs\varepsilon'|\bs Z)$. Also, write the covariance between the cluster jackknife AR and cluster jackknife score as

equation[equation omitted — 250 chars of source]

with $i^{\text{th}}$ element $C_{CLJ,i}=2\sum_{g\neq h}\operatorname{tr}(\bs\Xi_{g}^{(i)\prime}\bs P_{Z,[g,h]}\bs\Sigma_{h}\bs P_{Z,[h,g]})/\sqrt{nk}$ and gather the variances and covariances in

equation[equation omitted — 200 chars of source]

Then the joint limiting distribution of the cluster jackknife AR and cluster jackknife score is given by the following theorem.

theoremUnder (ref) and if the cluster jackknife AR and cluster jackknife score are not perfectly correlated, then \begin{equation} \begin{split} \bs V_{CLJ}^{-1/2}\begin{bmatrix} AR_{CLJ}(\bs\beta_0)\\ \bs S_{CLJ}(\bs\beta_0) \end{bmatrix}\rightsquigarrow N(\bs 0,\bs I). \end{split} \end{equation}
proofThe proof is given in (ref)

To use this result for testing I propose the following conditionally unbiased and consistent estimators for $V^{AR}_{CLJ}$, $\bs V^{S}_{CLJ}$ and $\bs C_{CLJ}$.

equation[equation omitted — 767 chars of source]
theoremUnder (ref) it holds that $\operatorname{E}(\hat{\bs V}_{CLJ}(\bs\beta_0)|\bs Z)=\bs V_{CLJ}$ and, conditional on $\bs Z$, $\hat{\bs V}_{CLJ}(\bs\beta_0)\overset{p}{\rightarrow} \bs V_{CLJ}$.
proofThe proof is given in (ref).

To conclude this subsection, I note that the normal distribution in (ref) obtains only when then number of instruments goes to infinity. For smaller $k$ on the other hand, it is expected that $AR_{CLJ}(\bs\beta_0)/\sqrt{V^{AR}_{CLJ}}$ is closer to a shifted and scaled $\chi^2_k$ distribution. Therefore, if one decides to use the cluster jackknife test in isolation, I suggest to follow mikusheva2021inference and reject the null hypothesis $H_0\colon\bs\beta_0=\bs\beta$ whenever $AR_{CLJ}(\bs\beta)/\sqrt{\hat{V}^{AR}_{CLJ}(\bs\beta)}$ is larger than $(\chi^2_{k,1-\alpha}-k)/\sqrt{2k}$, where $\chi^2_{k,1-\alpha}$ is the $1-\alpha$ quantile of the $\chi^2_k$ distribution. This way the cluster jackknife AR statistic is compared with a shifted and scaled $\chi^2_k$ critical values for small $k$ and standard normal critical values for large $k$.

Extensions

The cluster jackknife AR and score tests can be improved or extended in several ways. Here I provide four of such improvements and extensions.

Combination of AR and score

AR and score statistics can be combined to form a more powerful test moreira2003conditional,kleibergen2005testing,iandrews2016conditional. lim2024conditional combine mikusheva2021inference's (mikusheva2021inference) jackknife AR statistic and matsushita2020jackknife's (matsushita2020jackknife) jackknife score statistic in a conditional linear combination test iandrews2016conditional. The joint convergence of the two cluster jackknife statistics in (ref) suggests that these can be similarly combined. Moreover, (ref) gives formal support for the unproven high level assumption about the joint convergence of the jackknife AR and jackknife score statistics in lim2024conditional.

The combination of the cluster jackknife AR and cluster jackknife score statistic through lim2024conditional's (lim2024conditional) framework, requires only an adaptation of the variance estimators and a rescaling of the score statistic. I give the details in (ref). Since this combination is computationally intensive, however, I will in the following only consider the cluster jackknife AR and cluster jackknife score statistics separately.

Symmetric cluster jackknife

bekker2015jackknife observe that when jackknifing a statistic, the information in the terms that are removed terms is lost. Therefore they propose a different jackknife procedure that retains part of the information in these deleted observations, while still correctly centering the statistic. In particular, bekker2015jackknife propose an alternative to the jackknife HLIM and HFUL estimators from hausman2012instrumental, which were developed for a heteroskedastic many weak IV model. The numerator of bekker2015jackknife's (bekker2015jackknife) estimator can be written as

equation[equation omitted — 203 chars of source]

where $\tilde{\bs y}=\tilde{\bs P}\bs y$ and $\tilde{\bs X}=\tilde{\bs P}\bs X$. Here $\tilde{\bs P}=(\bs I-\bs D_{P_Z})\dot{\bs P}_Z$ denotes the jackknife projection matrix as used for example in ackerberg2009improved. This numerator thus treats all endogenous variables symmetrically, because it jackknifes both the $\bs y$ and the $\bs X$, which leads bekker2015jackknife to call it the symmetric jackknife estimator. crudu2021inference base their jackknife AR statistic on this numerator, with the goal to retain part of the information in the deleted observations.

As the cluster jackknife AR removes entire clusters, rather than individual observations, the potential loss of information is even larger. I therefore generalise the symmetric jackknife to clustered data. In (ref) I show that pre-multiplying a vector by the matrix $\tilde{\bs P}_{CL}=\bs P_{Z}-\bs B_{P_Z}\bs B_{M_Z}^{-1}\bs M_{Z}$, cluster jackknifes the observations in the vector. Then substituting $\ddot{\bs P}_{Z}$ by $(\tilde{\bs P}_{CL}+\tilde{\bs P}_{CL}')/2$ in the cluster jackknife AR and score statistics and their variance estimators yields symmetric cluster jackknife AR and score statistics.

Finally, as in bekker2015jackknife, I conclude that the cluster symmetric jackknife maintains a larger signal component than the cluster jackknife, because the quadratic form on which both cluster symmetric jackknife statistics are built, is

equation[equation omitted — 437 chars of source]

Many controls cluster jackknife

In (ref) I assumed that there are either no exogenous controls, or that they are small in number and have been partialled out of the model. Partialling out, however, introduces dependence between the observations. In case the number of control variables is small compared to the sample size, this dependence is usually asymptotically negligible. If, on the other hand, the number of controls is large compared to the sample size, the dependence might not go away asymptotically.

chao2023jackknife note that the additional dependence between the observations can have as consequence that jackknifing does not remove all terms with non-zero expectation from statistics. They therefore introduce a new type of jackknife, which I call the many controls jackknife, that correctly projects out the control variables, while also centring the statistic.

One particularly relevant type of control that this many controls jackknife can handle is cluster fixed effects to account for within cluster dependence. chao2023jackknife require however that once these effects and other controls are accounted for, the observations are independent, which therefore is different from the model considered in this paper that allows for within cluster dependence after controlling for observables.

Two remarks on the inclusion of many controls in a clustered model are in order. First, if the controls are only cluster specific, partialling out these controls introduces dependence within clusters, but not between clusters. Therefore all tests in this paper can readily be applied to these models. This is for example the case with cluster fixed effects.

Second, the many controls jackknife can be extended to clustered data for the case in which partialling out the controls introduces dependence between the clusters and that after controlling for these controls there remains within cluster dependence. An example of this type of data is when the data are clustered at two levels, where one level subsumes the other, and it is assumed that all dependence on the higher level of clustering can be modelled through a cluster fixed effect, but the dependence on the lower level of clustering cannot.

To extend the many controls jackknife to this type of data, write the model in (ref) as

equation[equation omitted — 236 chars of source]

where $\bs W\in\mathbb{R}^{n\times l}$ are the exogenous control variables. These are partialled out by premultiplying both equations with $\bs M_{W}$, which yields

equation[equation omitted — 223 chars of source]

The cluster jackknife AR and cluster jackknife score statistics then are $\bs V'\bs M_{W}\ddot{\bs P}_{M_{W}\bar{Z}}\bs M_{W}\bar{\bs\varepsilon}(\bs\beta)=\bs V'(\bs P_{M_{W}\bar{Z}}-\bs M_{W}\bs B_{M_{W}\bar{Z}}\bs M_{W})\bar{\bs\varepsilon}(\bs\beta)$, with $\bs V$ either $\bar{\bs\varepsilon}(\bs\beta)=\bar{\bs y}-\bar{\bs X}\bs\beta$ or $\bar{\bs X}$. The partialling out makes that deducting $\bs M_{W}\bs B_{M_{W}\bar{Z}}\bs M_{W}$ from $\bs P_{M_{W}\bar{Z}}$ does not yield a zero block diagonal matrix in general, which makes the cluster jackknife AR and cluster jackknife score statistics not properly centred. I therefore need to find an alternative matrix, say $\bs H$, such that $\bs P_{M_{W}\bar{Z}}-\bs M_{W}\bs H\bs M_{W}$ has zero block diagonal.

To find such a matrix, introduce the following notation. Let $\operatorname{vecb}(\bs A)$ be the column vectorization of the blockdiagonal elements of the $n\times n$ matrix $\bs A$, where we leave the dependence on the clustering structure implicit. Furthermore let $\bs A*\bs B$ be the Khatri-Rao product between $m\times n$ matrices $\bs A$ and $\bs B$, defined as the blockkwise Kronecker product. That is, for certain partitions,

equation[equation omitted — 607 chars of source]

define

equation[equation omitted — 484 chars of source]

In what follows I will take the cluster structure for the partition.

The matrix $\bs H$ then needs to be such that

equation[equation omitted — 145 chars of source]

Start by considering the $h^{\text{th}}$ block and assume $\bs H$ is a blockdiagonal matrix. Then, write the right hand side as

equation[equation omitted — 544 chars of source]

such that $\operatorname{vecb}(\bs P_{M_{W}Z})=(\bs M_{W}*\bs M_{W})\operatorname{vecb}(\bs H)$. Now let $\operatorname{vecb}^{-1}$ be the operator that constructs a blockdiagonal matrix out of a vector, in such a way that for any $n\times n$ matrix $\bs A$ it holds that $\operatorname{vecb}^{-1}(\operatorname{vecb}(\bs A))=\bs B_{A}$. Then, using $\bs H=\operatorname{vecb}^{-1}[(\bs M_{W}*\bs M_{W})^{-1}\operatorname{vecb}(\bs P_{M_{W}Z})]$ in the many controls cluster jackknife AR and score statistics, $\bs V'(\bs P_{M_{W}\bar{Z}}-\bs M_{W}\bs H\bs M_{W})\bar{\bs\varepsilon}(\bs\beta)$, makes them centred, while also correctly handling the many control variables. Substituting $\ddot{\bs P}_{Z}$ by $\bs P_{M_{W}\bar{Z}}-\bs M_{W}\bs H\bs M_{W}$ in (ref) gives the corresponding variance estimators.

Cross-fit variance estimators

The variance estimators in (ref) are not the only possible ones. mikusheva2021inference note that a jackknife AR that uses the equivalent of $\hat{V}_{CLJ}^{AR}(\bs\beta)$ by crudu2021inference has low power against alternatives far from $\bs\beta_0$, due to a bias in the residuals that makes the variance estimate unnecessarily large. Therefore they propose an alternative estimator that removes this bias by using cross-fitted values of the residuals, to improve the power newey2018cross,kline2020leave.

Cross-fit variance estimators for clustered data can be obtained by defining $\tilde{\bs P}(g,h)=\bs Z(\bs Z'_{-([g],[h])}\bs Z_{-([g],[h])})^{-1}\bs Z'_{-([g],[h])}$, where a matrix indexed by $-([g],[h])$ contains all rows of that matrix but those in $[g]$ and $[h]$. Pre-multiplying $\bs \varepsilon(\bs\beta)_{-([g],[h])}$ with $\tilde{\bs P}(g,h)$ therefore gives the leave-cluster-$(g,h)$-out fitted values of $\bs\varepsilon(\bs\beta)$ on $\bs Z$. Denote this by $\tilde{\bs\varepsilon}(\bs\beta;g,h)=\tilde{\bs P}(g,h)\bs\varepsilon(\bs\beta)_{-([g],[h])}$. The cross-fit variance estimators can then be written as

equation[equation omitted — 1,058 chars of source]

Since $g\neq h$ and $\tilde{\bs\varepsilon}(\bs\beta;g,h)$ neither depends on observations in cluster $g$, nor on observations in cluster $h$, the terms with $\tilde{\bs\varepsilon}(\bs\beta;g,h)$ in (ref) have expectation zero. Consequently, the cross-fit variance estimators are unbiased.

Finally, note that substituting $\bs P_{Z}$ with $(\tilde{\bs P}_{CL}+\tilde{\bs P}_{CL}')/2$ in (ref) yields cross-fit variance estimators for the cluster symmetric jackknife statistics.

Simulation results

In this section I explore the finite sample performance of the different cluster robust tests relative to each other and to those for independent data. I generate from a linear IV model with a single endogenous regressor and clustered data as in boot2023identification, which is itself an adaptation of the data generating process (DGP) from hausman2012instrumental to clustered data. To focus this section, I only consider the cluster jackknife AR and cluster jackknife score tests from (ref) separately, hence without the extensions from (ref).

I generate $n=800$ observations as $y_i=\alpha+\beta x_i+\varepsilon_i$ for $i=1,\dots,n$ and $\alpha=0$. Throughout I test $H_0\colon\beta=0$, but the $\beta$ in the DGP can have a different value when investigating the power. The endogenous regressor $x_i$ relates to a single instrumental variable $\tilde{z}_i$, $x_i=\pi\tilde{z}_{i}+\eta_i$. There are $k$ additional instrumental variables that can be used for inference $\bs Z_i=(1,\tilde{z}_i, \tilde{z}_i^2, \tilde{z}_i^3, \tilde{z}_i^4, \tilde{z}_i D_{i1},\dots, \tilde{z}_iD_{ik-4})'$, where the $D_{ij}$, $j=1,\dots,k-4$, are independently distributed Bernoulli(1/2) random variables. $\varepsilon_i$ is a function of random variables, $\varepsilon_i=\rho\eta_i+\sqrt{(1-\rho^2)/(\phi^2+0.86^4)}(\phi v_{1i}+0.86v_{2i})$, where $\rho=0.3$ is the degree of endogeneity and $\phi=1.38072$ as in bekker2015jackknife. The distribution of the other random variables is detailed below.

The $n$ observations are divided over $G=100$ clusters. The clusters can be unbalanced and determined by first setting $n_g=\max\{1, n\exp(\gamma g/G)/[\sum_{g=1}^{G-1}\exp(\gamma g/G)+1]\}$ for $g=1,\dots G-1$ and $n_G=\max\{1,n-\sum_{g=1}^{G-1}n_g\}$. Next, all cluster sizes are ensured to be integer by rounding them down. Finally, to get the correct sample size, the first $n-\sum_{g=1}^Gn_g$ clusters are increased by one. Note that $\gamma$ determines the degree of unbalancedness and $\gamma=0$ implies balanced clusters.

The random variables $\tilde{z}_i$, $\eta_i$, $v_{1i}$ and $v_{2i}$ consist of an idiosyncratic and a cluster common component, weighted by $\lambda=1/2$. To be precise, $\tilde{\bs z}_{[g]}=\sqrt{\lambda}\tilde{\bs z}_{[g]}^\text{ind}+\sqrt{1-\lambda}\tilde{z}_{g}^\text{cl}$, $\bs\eta_{[g]}=\sqrt{\lambda}\bs\eta_{[g]}^\text{ind}+\sqrt{1-\lambda}\eta_{g}^\text{cl}$, $\bs v_{1,[g]}=\sqrt{\lambda}\bs v_{1,[g]}^\text{ind}+\sqrt{1-\lambda}v_{1,g}^\text{cl}$ and $\bs v_{2,[g]}=\sqrt{\lambda}\bs v_{2,[g]}^\text{ind}+\sqrt{1-\lambda}v_{2,g}^\text{cl}$. The cluster common and idiosyncratic components of $z_{i}$ and $\eta_i$ follow standard normal distributions. I draw $v_{1i}^\text{ind}\sim N(0,(\tilde{z}_{i}^\text{ind})^2)$ and $v_{1h}^\text{cl}\sim N(0,(\tilde{z}_{h}^\text{cl})^2)$. The cluster common and idiosyncratic components of $v_{2i}$ come from a normal distribution with mean zero and variance $0.86^2$.

Size

To show the importance to account for clustering, consider the size of the AR, score, jackknife AR and jackknife score tests for independent data and a $t$-test based on 2SLS, standard errors for independent data and normal critical values in the DGP above with $\pi=0.1$, $\gamma=1$ such that the smallest cluster contains $4$ observations and the largest $12$, and $k$ ranging from $2$ to $90$.

The left panel of (ref) shows the rejection over $10\,000$ draws when testing $H_0\colon\beta=0$ at a $5\%$ significance level. Clearly, as these test do not take the clustered dependence into account the tests are size distorted. For smaller values of $k$ all tests are oversized. When $k$ increases, the identification robust tests become closer to size correct. For the AR and score tests this can partly be explained by them being conservative for large $k$. A closer inspection of how $\bs P_{Z}$ changes with $k$, showed that the fall in the rejection rates of the jackknife AR and the jackknife score is due to a reduction in the true variance not captured by the variance estimator. Such a reduction is observed in the current DGP, but there are no guarantees that this holds more generally. Moreover, although the jackknife tests become closer to size correct, they are still oversized for large values of $k$.

Next, in the right panel of the same figure, I apply the cluster robust versions of the AR, score, jackknife AR, and jackknife score tests and the $t$-test based on 2SLS with clustered standard errors to the same data. Observe that for a small $k$ all tests are size correct. When $k$ increases, 2SLS becomes oversized, while the cluster AR becomes conservative. The cluster score, the cluster jackknife AR and the cluster jackknife score remain size correct.

(ref) shows additional simulation results on the size of the cluster robust tests. In particular, by varying the number of clusters it confirms the hypothesis that with clustering in the data, the number of instruments relative to the number of clusters, rather than the number of observations matters for many instrument problems. Furthermore, the appendix gives results on the robustness of the cluster jackknife AR and cluster jackknife score tests to the arguably most stringent assumption, (ref), that limits the size of the largest cluster. It shows that with a dominant cluster the cluster jackknife AR breaks down and becomes oversized, whereas the cluster jackknife score is surprisingly robust.

figure[figure omitted — 885 chars of source]

Power

Given that the cluster AR, cluster score, cluster jackknife AR and cluster jackknife score tests are not oversized, I investigate their power. (ref) shows the rejection rates of the tests when testing $H_0\colon\beta=0$ over $10\,000$ draws when the value of $\beta$ in the DGP varies from $-3$ to $3$. $k$, furthermore, varies between 10 and 30 to show the effect of the number of instruments on the power. Similarly, different instrument strengths are considered by setting $\pi\in\{0.2,0.4\}$. The other parameters in the DGP are as for (ref).

Observe that in each of the four panels the cluster jackknife AR and cluster jackknife score tests have higher power than their counter parts not adapted for many instruments. Moreover, whereas the cluster AR and cluster score tests show a clear drop in power when $k$ increases from 10 to 30, the power the cluster jackknife AR and cluster jackknife score test are almost unaffected.

Furthermore, note that in most panels the cluster score test has higher power than the cluster AR test when the true $\beta$ is close to the tested value of $0$. As is often observed for score tests, the power of the cluster score test drops below that of the cluster AR for $\beta$s further away from $0$. The advantage of the cluster score test over the cluster AR test is mainly observed for a moderate number of strong instruments.

For the cluster jackknife tests the ordering in power is clearer. In the four DGPs of (ref) the cluster jackknife AR test outperforms the cluster jackknife score test when the instruments are weak as shown in the top panels. In the bottom panels, where the instruments are stronger, the difference in power is marginal.

However, it need not always be the case that the cluster jackknife AR test has power better or equal to the cluster jackknife score test. Consider for example an adaptation to the DGP also featured in boot2023identification, in which there is high heteroskedasticity. In this DGP there is an additional heteroskedasticity parameter, $\kappa$, used to scale $v_{1i}^\text{ind}\sim N(0,(\tilde{z}_{i}^\text{ind})^\kappa)$ and $v_{1h}^\text{cl}\sim N(0,(\tilde{z}_{h}^\text{cl})^\kappa)$. Previously $\kappa=2$. (ref) shows the rejection rates of the different tests over $10\,000$ draws when $\kappa=6$. In this case the cluster jackknife score test clearly outperforms the cluster jackknife AR test.

figure[figure omitted — 619 chars of source]
figure[figure omitted — 362 chars of source]

Empirical application

To further illustrate the cluster robust tests I revisit the study by dube2020queens on the effect of female leadership on the likelihood for country to be at war. As the probability for a woman to ascend a throne and come into power might be higher or lower depending on whether the polity is at war, the gender of the leader might be endogenous. dube2020queens therefore propose instruments for it based on historic succession laws for monarchs in Europe. The authors note that a vacant throne was less likely to be taken by a woman if the previous ruler had a male first born child and more likely to be taken by a woman if the previous ruler had a sister. Neither the firstborn's gender, nor whether the previous ruler had a sister seems to directly affect the involvement in armed conflict, which justifies the identification strategy.

To estimate the effect of queenly reign on war, the authors use a data set on 18 European polities observed some time between 1480 and 1913 yielding an unbalanced panel with 3\,586 polity-year observations in total. Throughout the data are clustered on a broad indicator of a reign, resulting in 176 clusters that vary in size. The largest cluster counts 66 observations and the smallest 1.

The estimates from the main model, which uses the two instruments described above, imply that polities lead by a woman are 39 percentage points more likely to be in war in a given year compared to polities lead by a man. As a robustness check, models with additional instruments based on interactions between the original instruments and exogenous regressors are proposed. Table A5 in dube2020queens's (dube2020queens) appendix shows the 2SLS estimates from these models. The smallest and largest estimates imply that polities lead by women are 29 and 50 percentage points more likely to be at war. However, the new instruments are relatively weak, which makes the estimates and their standard errors unreliable and therefore the larger models are not investigated further.

Another problem with these models with extra instruments is that since the data are clustered, the effective sample size will be reduced from the number of observations towards to the number of clusters, making the total number of IVs non-negligible compared to the effective sample size, which exacerbates the problems from which the 2SLS estimates and standard errors suffer. To reliably continue the analysis of the larger models thus requires weak and many instrument robust tests suited for clustered data.

In (ref) I draw the 95% confidence intervals for $\beta$, which is the effect of queenly reign on the likelihood of war relative to polities lead by kings, that I obtained by inverting the cluster AR test, the cluster score test, the cluster jackknife AR test and the cluster jackknife score test, both without cross-fit variance. I consider $\beta$s between -0.1 and 1, as values below -0.1 did not add any information and values above 1 have no sensible interpretation. The groups correspond to the models in columns 1, 2, 4 and 5 of Table A5 in dube2020queens and the full model where I pooled all the instruments and regressors of the other models. The exogenous control variables have been partialled out without the many controls cluster jackknife.

From the figure I conclude, that the cluster score yields confidence intervals that are unbounded from above, and for the third model also from below. (ref) suggests that this can happen when the instruments are relatively weak, in which case it is important to use robust tests. The suspicion of weak instruments is in line with the results from dube2020queens and further confirmed by the relatively wide confidence intervals for the cluster jackknife score test.

Furthermore, the tests other than the cluster score tests yield for most specifications significantly positive effects of female rule on the likelihood of being at war. Only the confidence intervals obtained by inverting the cluster jackknife score for the third model and the cluster AR test for the full model include zero. One can see this more clearly in (ref) in (ref) which reports the exact bounds of the confidence intervals. For the latter model, the cluster jackknife type tests do yield significant results, which therefore shows the benefit of correcting for many instruments.

That there can be a benefit from using tests that are more powerful with many instruments more generally, can be seen from the relative lengths of the cluster jackknife AR and cluster AR confidence intervals. These are $0.932$, $1.025$, $1.218$, $0.796$ and $0.996$ for the five different models respectively. Although cluster jackknife AR does not always yield a narrower confidence interval, the reduction can, with more than 20% in the fourth model, be sizeable.

figure[figure omitted — 1,156 chars of source]

Conclusion

In this paper I argued that the problems associated with many and weak instruments are more likely in case the data are clustered than when they are independent. This makes it particularly important to use tests robust against many and weak instruments when using clustered data, which none of the previously developed robust tests is able of.

I therefore showed that the many and weak instrument robust jackknife AR and jackknife score tests can be extended by removing clusters of observations, rather than individual observations when jackknifing. I furthermore combined the cluster jackknife AR and score in a conditional linear combination test, showed how more information can be retained using the symmetric cluster jackknife, allowed for many control variables and improved the variance estimation via cross-fit.

Monte Carlo simulations showed that the cluster jackknife tests are size correct, whereas the tests for independent data overreject, and that the new tests have good power. In the last section I applied the newly developed tests to the data from dube2020queens and showed that queenly reign has a significant positive effect on the probability for a state to be at war.

Acknowledgements

I thank Stanislav Anatolyev, Tom Boot, Mikkel S\o lvsten, Tom Wansbeek, Tiemen Woutersen and participants at conferences and seminars for insightful comments and discussions. This paper previously circulated under the title Inference in IV models with clustered dependence, many instruments and weak identification.