EconBase
← Back to paper

Identification- and many moment-robust inference via invariant moment conditions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

69,567 characters · 13 sections · 97 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification- and Many Moment-Robust Inference via Invariant Moment Conditions

abstractIdentification-robust hypothesis tests are commonly based on the continuous updating GMM objective function. When the number of moment conditions grows proportionally with the sample size, the large-dimensional weighting matrix prohibits the use of conventional asymptotic approximations and the behavior of these tests remains unknown. We show that the structure of the weighting matrix opens up an alternative route to asymptotic results when, under the null hypothesis, the distribution of the moment conditions satisfies a symmetry condition known as reflection invariance. We provide several examples in which the invariance follows from standard assumptions. Our results show that existing tests will be asymptotically conservative, and we propose an adjustment to attain nominal size in large samples. We illustrate our findings through simulations for various linear and nonlinear models, and an empirical application on the effect of the concentration of financial activities in banks on systemic risk. \\ Keywords: identification-robust inference, many moment conditions, continuous updating.\\ JEL codes: C12, C26.

Introduction

Identification-robust inference in models defined by moment conditions is commonly based on the continuous updating (CU) objective function of \citet*{hansen1996finite}, either directly through anderson1949estimation's (anderson1949estimation) AR statistic, via its score as in kleibergen2005testing's (kleibergen2005testing) K statistic, or via a combination thereof moreira2003conditional,kleibergen2005testing,iandrews2016conditional. While conventional tools yield the limiting distribution of these statistics when the number of moment conditions is small, their asymptotic behavior is unknown when the number of moments increases proportionally with the sample size.

The main obstacle in establishing the limiting distribution of statistics based on the CU objective function is that when the number of moment conditions increases proportionally with the sample size, the weighting matrix for the moment conditions is large-dimensional and does not converge to a well-defined object. As newey2009generalized write for the instrumental variables (IV) model “If the number of instruments grows as fast as the sample size, the number of elements of the weight matrix grows as fast as the square of the sample size. It seems difficult to simultaneously control the estimation error for all these elements." In the linear IV model, a recent stream of the literature circumvents this problem by changing the weighting matrix so that it depends only on the instruments, and not on the second stage errors as in the CU objective function hausman2012instrumental,crudu2021inference,mikusheva2021inference,matsushita2020jackknife,dovi2022ridge,lim2022conditional. However, for nonlinear models, no such results are available, and existing theory requires that the number of moment conditions grows at a slower rate than the sample size, e.g.\ han2006gmm and newey2009generalized.

In this paper, we focus on the CU-GMM objective function when the number of moment conditions increases proportionally to the sample size, similar to the many instrument sequences in linear IV models studied by bekker1994alternative. We show how the difficulty posed by the weighting matrix can be circumvented when the distribution of the moment conditions evaluated at the true parameters is symmetric around zero, referred to as orthant symmetry by efron1969student and as reflection invariance by bekker2008symmetry. We discuss various examples where reflection invariance is implied by existing assumptions: a censored panel data model, median IV regression, linear IV with reflection invariant instruments or regression errors, and GMM-based tests for symmetry.

We show that under reflection invariance, the limiting distribution of the suitably recentered and rescaled CU objective function, and hence that of the AR statistic, follows from known results on the limiting behavior of bilinear forms by \citet*{chao2012asymptotic}. These results indicate that the usual AR test will be conservative under many moments. We propose adjusted critical values that make the procedure asymptotically exact regardless of whether the number of moments is fixed or increasing. The results naturally extend to settings with clustered observations that in a linear IV setting were considered by ligtenberg2023inference. For the heteroskedastic linear IV model, we also study the asymptotic behavior of the score of the CU objective function kleibergen2002pivotal,iandrews2016conditional,matsushita2020jackknife. We derive a new central limit theorem for bilinear forms with an additional third-order term. We provide a set of assumptions under which the additional variance terms that enter under many instruments are negative, indicating that kleibergen2005testing's (kleibergen2005testing) score test will yield conservative inference. We also provide a variance estimator that consistently estimates the additional variance terms.

We assess the finite sample performance of the tests through simulations in two nonlinear settings and one linear setting. Specifically, for the nonlinear settings we consider the moment conditions by honore1992trimmed for censored panel data models and the quantile IV moments developed by chernozhukov2009finite. The simulations highlight the effect of (i) the identification strength, (ii) the number of instruments, and (iii) the effect of deviations from the invariance assumption. The simulation shows that, unlike conventional asymptotic approximations, the many moment-robust tests have close to nominal size control regardless of the identification strength and regardless of the number of moment conditions. Consequently, we document an increase in power when using the many-moment robust tests over the classic AR statistic. In the quantile IV application, the power appears nearly identical to that of the finite sample inference procedure of chernozhukov2009finite. The quantile IV setting also provides a natural setting to test the sensitivity to violations of reflection invariance, as the invariance only holds at the median. We find that the proposed test is reasonably robust, although the rejection rates exceed the nominal size for quantiles distant from the median.

Finally, we consider the heteroskedastic linear IV model under independent and clustered errors. For the case with independent errors, the many instrument correction to the AR and score statistic yields higher power relative to the standard AR and score test and the jackknife AR and score test mikusheva2021inference,matsushita2020jackknife in settings with high heteroskedasticity, while differences are small in settings with low heteroskedasticity. Under clustering, the difference between the standard AR statistic and the many instrument robust version is larger compared to the setting with independent errors. With low heteroskedasticity, a clustered version of the jackknife AR appears more powerful. When we increase the degree of heteroskedasticity, this difference disappears.

We consider an empirical illustration by revisiting the study by langfield2016bank on the effect of concentration of financial activity in the banking sector on systemic risk. As the data are left-censored, the authors employ the reflection invariant moment conditions from honore1992trimmed. These moment conditions are clustered at the bank level, and the number of moment conditions is not negligible relative to the number of banks. The confidence regions from the AR statistic and the many instrument-robust AR statistic are wide, but, similar to langfield2016bank, provide evidence that countries with a relatively large banking sector experience lower financial stability during financial crises than countries with a relatively small banking sector. As expected, the many moment-robust AR statistic offers smaller confidence regions compared to the standard AR statistic.

\paragraph*{Structure} In Section (ref) we discuss the reflection invariance assumption, and several examples where the assumption holds, in Section (ref). Theoretical results under independent moment conditions are given in Section (ref) and under clustered moment conditions in Section (ref). Score-based tests in the linear IV model are discussed in Section (ref). Section (ref) contains the Monte Carlo results for a censored tobit model, quantile IV and linear IV. The empirical application is presented in Section (ref). Section (ref) concludes.

\paragraph*{Notation} For a vector $\bs v$, $\bs D_{v}$ denotes the diagonal matrix with $\bs v$ on its diagonal. For a square matrix $\bs A$, let $\bs D_{A}=\bs A\odot \bs I$, where $\odot$ is the Hadamard product. We use $\dot{\bs A}= \bs A-\bs D_{A}$ for a matrix with all diagonal elements equal to zero. $\bs\iota$ indicates a vector of ones and $\bs e_{k}$ a vector with its $k^{\text{th}}$ entry equal to one and the remaining entries equal to zero. Observations are indexed by $i=1,\ldots,n$ and $j=1,\ldots,n$, parameters of interest by $l=1,\ldots,p$, clusters by $h=1,\ldots, H$, time by $t=1,\ldots, T$. If we require additional indices, we subscript the relevant index, for example $i_{1},i_{2}\ldots$. Let $\bs a_{i}=\bs A'\bs e_{i}$ and $\bs a_{(l)}=\bs A\bs e_l$ denote the $i^\text{th}$ row and $l^\text{th}$ column of a matrix $\bs A$. For random variables $A$ and $B$, $A\overset{(d)}{=}B$ means that $A$ is distributionally equivalent to $B$. $A\overset{(\operatorname{E})}{=}B$ means that $\operatorname{E}[A] = \operatorname{E}[B]$. $\operatorname{E}_{A}[\cdot]$ is the expectation over the distribution of the random variable $A$. $\rightarrow_{d}$ denotes convergence in distribution, $\rightarrow_{p}$ convergence in probability and $\rightarrow_{a.s.}$ almost sure convergence. $a.s.n.$ is short for with probability 1 for all $n$ sufficiently large. For a symmetric $n\times n$ matrix $\bs A$, $\operatorname{\lambda_{\mathrm{min}}}(\bs A)=\lambda_1(\bs A)\leq\ldots\leq\lambda_n(\bs A)=\operatorname{\lambda_{\mathrm{max}}}(\bs A)$ denote its eigenvalues. $C$ denotes a finite positive constant that can differ between appearances. $\mathbbm{1}\{\cdot\}$ is the indicator function.

Continuous updating and invariant moment conditions

We start with a GMM set-up where we have a vector of independent observations $\bs W_{i}$ for $i=1,\ldots,n$. Define the $p\times 1$ vector of parameters $\bs\beta$ and the vector of functions $\bs g_{i}(\bs\beta) = (g_{1}(\bs W_i,\bs\beta), \ldots,g_{k}(\bs W_i,\bs\beta))'$ for $k\geq p$. The true parameter $\bs \beta_0$ satisfies the moment conditions $\operatorname{E}[\bs g(\bs W_{i},\bs\beta_{0})]=\bs 0$. We stack the moment conditions in the $n\times k$ matrix $\bs G(\bs\beta) = [\bs g_{1}(\bs\beta),\ldots,\bs g_{n}(\bs\beta)]'$, with $\operatorname{rank}(\bs G(\bs\beta_{0}))=k$. Define the projector $\bs P(\bs\beta) = \bs G(\bs\beta)(\bs G(\bs\beta)'\bs G(\bs\beta))^{-1}\bs G(\bs\beta)'$.

The continuous updating (CU) objective function introduced by hansen1996finite can be written as

equation[equation omitted — 157 chars of source]

This objective function is closely related to the identification robust Anderson--Rubin (AR) GMM statistic, defined as

equation[equation omitted — 76 chars of source]

For a fixed number of moment conditions $k$, the AR statistic is asymptotically $\chi^2(k)$ distributed when evaluated at $\bs\beta_{0}$. Extending this result to the case where the number of moment conditions grows proportionally with the sample size is challenging as the weighting matrix $\bs G(\bs\beta)'\bs G(\bs\beta)/n$ does not converge to a non-stochastic object.

Invariance assumptions

To derive the asymptotic behavior of (ref) when the number of moment conditions $k$ grows proportional with $n$, we make the following invariance assumption.

assumption[Reflection invariance] Let $\{r_{i}\}_{i=1}^{n}$ be a sequence of independent Rademacher random variables, so that $\operatorname{P}(r_{i}=1)=\operatorname{P}(r_{i}=-1)=1/2$, and define $\bs r=(r_{1},\ldots,r_{n})'$. Then, $\bs G(\bs\beta_{0})\overset{(d)}{=}\bs D_{r}\bs G(\bs\beta_{0})$.

Assumption (ref) is equivalent to stating that the distribution of $\bs g_{i}(\bs\beta_{0})$ is symmetric around 0. Note that this does not preclude the distribution of the moment conditions to differ across observations. We now discuss several examples under which Assumption (ref) holds.

example[Panel tobit model] honore1992trimmed considers a (censored or truncated) two period panel data model \begin{equation} \begin{split} y_{i1}&=\alpha_i+\bs x_{i1}'\bs\beta+\varepsilon_{i1},\\ y_{i2}&=\alpha_i+\bs x_{i2}'\bs\beta+\varepsilon_{i2}. \end{split} \end{equation} If the $\varepsilon_{it}$ are exchangeable conditional on $\bs x_{i1}$, $\bs x_{i2}$ and $\alpha_i$, then $\varepsilon_{i1}-\varepsilon_{i2}$ is symmetrically distributed. Consequently, $y_{i1}-y_{i2}$ is symmetric around $(\bs x_{i1}-\bs x_{i2})'\bs\beta$. honore1992trimmed exploits this symmetry to derive valid moment conditions for $\bs\beta$, which are by the same mechanism reflection invariant. Any two time periods can be used to obtain moment conditions. By restricting the set so that we use each time period only once, we obtain reflection invariant moment conditions. The number of moment conditions then increases linearly with the product of the number of time periods and the dimension of the regressors. See Section (ref) for details on the moment conditions.
example[Median IV regression] chernozhukov2009finite estimate a quantile regression function $q$ that relates a dependent variable $y$ to possibly endogenous variables $\bs x$ and a uniformly distributed rank variable $\varepsilon$ as $y=q(\bs x,\varepsilon)$. Their Assumption 1 implies that for the median $\mathbbm{1}\{y\leq q(\bs x,1/2)\}$ is Bernoulli(1/2) distributed and hence reflection invariant once shifted. If $q$ can be parameterized with parameters $\bs\theta$ and there are exogenous instrumental variables $\bs z$, then we can estimate $\bs\theta$ via GMM with the invariant moment conditions $g_i(\bs\theta)=(1/2-\mathbbm{1} \{y_i\leq q(\bs x_i,1/2;\bs\theta)\})\bs\Psi(\bs z_i)$, where $\bs\Psi(\bs z_i)$ is some function of the instruments. The number of instrumental variables can be large in itself, or after using power series or splines as approximating functions for $\bs\Psi(\cdot)$ newey1990efficient.
example[Heteroskedastic IV with reflection invariant instruments or second stage errors] Consider the standard linear IV model \begin{align} y_{i} &= \bs x_{i}'\bs\beta_{0}+ \varepsilon_{i},\quad \bs x_{i} =\bs \Pi'\bs z_{i} + \bs\eta_{i}, \end{align} with $\bs x_{i}$ the endogenous regressors, $\bs\beta_{0}$ a $p\times 1$ vector, $\bs z_{i}$ the instruments, $\bs \Pi$ a $k\times p$ matrix. We assume an intercept has been projected out and write the model without exogenous control variables. Supplementary Appendix (ref) shows that this does not affect the results if the number of the controls increases slower than the sample size. The moment conditions are $\bs g_{i}(\bs\beta) = \bs z_{i}(y_{i}-\bs x_{i}'\bs\beta)$ and are reflection invariant under $H_{0}\colon\bs\beta=\bs\beta_0$ if $\bs z_{i}$ and $\varepsilon_{i}$ are independent and at least one of them is reflection invariant. Examples of invariant instruments include sex-at-birth instruments angrist1998children and randomized experiments with equal assignment probability to treatment and control groups. This leads to many instruments when these are interacted with control variables dupas2018banking,dube2020queens.
example[GMM based tests for symmetry] bontemps2005testing use Stein's equation and GMM to test for normality of a variable. Let $H_{i}$ be the Hermite polynomial of order $i$ recursively defined as $H_0(x)=1$, $H_1(x)=x$ and $H_i(x)=[xH_{i-1}(x)-\sqrt{i-1}H_{i-2}(x)]/\sqrt{i}$, $i>1$. bontemps2005testing show that for $X\sim N(0,1)$ the moment conditions $\operatorname{E}(H_i(X))=0$ hold. If $X$ is symmetrically distributed but not necessarily normal, then these moment conditions still hold for $i$ odd and are reflection invariant. The number of moment conditions is large if many orders are used.

There are also settings where invariance cannot be argued based on standard assumptions and one wishes to test this assumption. As the reflection invariance is part of the specification, an empty confidence set could point to a failure of this assumption. When the moment conditions are of a product form, such as in the linear IV $\boldsymbol{g}_{i}(\boldsymbol{\beta_{0}}) = \boldsymbol{z}_{i}(y_{i}-\boldsymbol{x}_{i}'\boldsymbol{\beta}_{0}) = \boldsymbol{z}_{i}\varepsilon_{i}$, reflection invariance is testable if it comes from the observable component, e.g.\ the instrument vector $\boldsymbol{z}_{i}$. Whether testing the unobserved component, e.g.\ $\varepsilon_{i}$ in the linear IV, is possible in combination with identification-robust inference is an important question for further research.

Continuous updating under reflection invariance

Under Assumption (ref) we can relate the distribution of the CU objective function with a similar function written in terms of Rademacher random variables. Define, $Q_{r}(\bs\beta_{0})= \bs r'\bs P(\bs\beta_{0})\bs r/n$. Then,

equation*[equation* omitted — 74 chars of source]

As $Q(\bs\beta_{0})$ and $Q_{r}(\bs\beta_{0})$ are distributionally equivalent, it suffices to analyze the asymptotic distribution of $Q_{r}(\bs\beta_{0})$ to obtain the asymptotic distribution of $Q(\bs\beta_{0})$. Likewise, we analyze $\operatorname{AR}_{r}(\bs\beta_{0})=nQ_{r}(\bs\beta_{0})$ to establish the asymptotic distribution of the AR statistic defined in (ref).

Conditioning on $\mathcal{J}=\{\bs g_{i}(\bs\beta_{0})\}_{i=1}^{n}$, the only randomness in $\operatorname{AR}_r(\bs\beta_{0})$ comes from the Rademacher random variables. Under the following assumptions, we can directly apply the CLT for bilinear forms by chao2012asymptotic to obtain the asymptotic distribution of the AR statistic under many instrument sequences.

assumptionWith probability 1 for all sufficiently large $n$: (a) $\operatorname{rank}[\bs P(\bs\beta_{0})]=k$, (b) $\sigma_{n}^2(\bs\beta_{0})>1/C$ where \begin{equation} \sigma_{n}^2(\bs\beta_{0}) = \frac{2}{k}\sum_{i\neq j} P_{ij}(\bs\beta_{0})^2. \end{equation}

Part (a) excludes any redundant moment conditions. Part (b) bounds the variance of the scaled AR statistic away from zero and is required to apply the central limit theorem provided in Lemma A2 by chao2012asymptotic. A stronger alternative assumption that can be found in the many instrument literature is that $P_{ii}(\bs\beta_{0})\leq C<1$ for $i=1,\ldots,n$, see e.g.\ hausman2012instrumental, bekker2015jackknife and anatolyev2019many. The following result follows from Lemma A2 in chao2012asymptotic and its proof.

theoremUnder Assumptions (ref) and (ref), when $k\rightarrow\infty$ as $n\rightarrow\infty$, $(k\sigma_{n}^2(\bs\beta_{0}))^{-1/2}(\operatorname{AR}_{r}(\bm \beta_{0})-k)\rightarrow_{d} N(0,1)$ $a.s.$ As a consequence $(k\sigma_{n}^2(\bs\beta_{0}))^{-1/2}(\operatorname{AR}(\bs \beta_{0})-k) \rightarrow_{d} N(0,1).$

Theorem (ref) shows that the AR statistic needs to be shifted and scaled to have a well-defined asymptotic distribution. A similar result is obtained by anatolyev2011specification for the AR statistic in a homoskedastic IV model with many instruments. While Theorem (ref) requires $k\rightarrow\infty$, we can achieve uniform inference across $k$ by testing based on the quantiles of the distribution of $Z=(2k)^{-1/2}(Z_{1}-k)$ where $Z_{1}\sim \chi^2(k)$. When $k$ is fixed, $\sigma_{n}^2(\bs\beta_{0})\rightarrow_{p} 2$, and hence, we compare $\operatorname{AR}(\bs\beta_{0})$ against the quantiles of a $\chi^2(k)$ distribution. When $k$ increases, the quantiles of $Z$ approach that of the standard normal distribution and Theorem (ref) applies.

Theorem (ref) implies that the usual AR test will be conservative at conventional significance levels. The proof implies that we reject a value of $\bs\beta$ if

equation*[equation* omitted — 173 chars of source]

If $\alpha< 0.3$, we have that $\chi^2(k)_{1-\alpha}>k$ for all $k$. We then see that the usual AR-test will reject less often relative to the many moment robust test. This result does not contradict the fact that both procedures are asymptotically size correct under a fixed number of moments, as in that case $k^{-1}\sum_{i=1}^{n}P_{ii}(\bs\beta)^2\rightarrow_{p}0$.

Clustered moment conditions

Assumption (ref) implies that under the null, randomly flipping the sign of the moment conditions of one observation, does not alter the distribution of the moment conditions of another observation. Many studies however, allow for clustered dependence between observations. If instead of Assumption (ref), we assume that the moment conditions of all observations within a cluster are jointly reflection invariant, we can generalize the AR statistic from above to allow for clustered data by taking similar steps as in the linear IV model analyzed in ligtenberg2023inference.

Assume that the $n$ observations can be grouped in $H$ clusters with $n_h$ observations in cluster $h=1,\dots,H$. Let $[h]$ denote all observations from cluster $h$ and for any $n\times m$ matrix $\bs A$ write $\bs A_{[h]}$ as the $n_h\times m$ submatrix of $\bs A$ with only the rows in $[h]$ selected. Then we adapt Assumption (ref) as follows.

assumptionLet $\{r_{h}\}_{h=1}^{H}$ be a sequence of independent Rademacher random variables. For all $h=1,\dots,H$ it holds that $\bs G(\bs\beta_{0})_{[h]}\overset{(d)}{=}r_h\bs G(\bs\beta_{0})_{[h]}$.

The CU objective function (ref) can be easily adapted to clustered dependence by summing the moment conditions of all observations within a cluster. This allows to write Assumption (ref) in a form that resembles Assumption (ref). Let $\tilde{\bs G}(\bs\beta)_{h}=\bs\iota_{n_h}'\bs G(\bs\beta_0)_{[h]}$ be the summed moment conditions, which we stack in the $H\times k$ matrix $\tilde{\bs G}(\bs\beta_0)$. Then, under Assumption (ref), $\tilde{\bs G}(\bs\beta_0)\overset{(d)}{=}\bs D_{\tilde{r}}\tilde{\bs G}(\bs\beta_0)$, where $\tilde{\bs r}=(r_1,\dots,r_H)'$. We then define $\tilde{\bs P}(\bs\beta_0)=\tilde{\bs G}(\bs\beta_0)(\tilde{\bs G}(\bs\beta_0)'\tilde{\bs G}(\bs\beta_0))^{-1}\tilde{\bs G}(\bs\beta_0)'$. The AR statistic is now given by $\tilde{\text{AR}}( \bs\beta) = \bs\iota_{H}'\tilde{\bs P}(\bs\beta)\bs\iota_{H}$. To derive its distribution, make the following assumption similar to Assumption (ref).

assumptionWith probability 1 for all sufficiently large $n$: (a) $\operatorname{rank}[\tilde{\bs P}(\bs\beta_{0})]=k$, (b) $\tilde{\sigma}_{n}^2(\bs\beta_{0})>1/C$ where \begin{equation} \tilde{\sigma}_{n}^2(\bs\beta_{0}) = \frac{2}{k}\sum_{h_1 \neq h_2} \tilde{P}_{h_1 h_2}(\bs\beta_{0})^2. \end{equation}

Part (a) is ensured by full column rank of $\bs G(\bs\beta_0)$, which also implies that $\operatorname{rank}[\bs P(\bs\beta_0)]=k$. Assumptions (ref) and (ref) then yield the following extension of Theorem (ref) to clustered data.

corollaryUnder Assumptions (ref) and (ref), when $k\rightarrow\infty$ as $H\rightarrow\infty$, $(k\tilde{\sigma}_{n}^2(\bs\beta_{0}))^{-1/2}(\tilde{\operatorname{AR}}(\bs \beta_{0})-k) \rightarrow_{d} N(0,1)$.

We note that under clustering, many moment sequences are defined as a number of moments that grows proportional to the number of clusters. As such, the corrections can become relevant for a relatively small number of moments even in large data sets.

Extension: score-based tests in linear IV

The fact that the standard AR test is conservative under many moment sequences and reflection invariance as shown in Section (ref), raises the question whether the same is true for a statistic based on the score of the CU objective function given in (ref). To gain insight into this question, we specialize to the linear IV model with heteroskedasticity from Example (ref). We then consider the application of Assumption (ref) to analyze the score statistic and show that it is indeed conservative under many moment sequences and reflection invariance.

The following notation will be convenient below: for some $\bs\beta$, not necessarily equal to $\bs\beta_{0}$, $\varepsilon_{i}(\bs\beta)=y_{i}-\bs x_{i}'\bs\beta$ and $\bs\varepsilon(\bs\beta) = (\varepsilon_{1}(\bs\beta),\ldots,\varepsilon_{n}(\bs\beta))'$. The model (ref) is accompanied by the following assumptions.

assumption(a) Conditional on $\bs Z$, $\{\varepsilon_{i},\boldsymbol{\eta}_{i}'\}_{i=1}^{n}$ is independent, with mean zero and $\operatorname{E}[(\varepsilon_{i},\bs\eta_{i}')'(\varepsilon_{i},\bs\eta_{i}')|\bs Z] = \bs \Sigma_{i}=\begin{array}{cccc} (\sigma_{i}^2 & \bs\sigma_{12i}'; & \bs\sigma_{12i}&\bs\Sigma_{22i})\end{array}$ (b) $0<C^{-1}\leq \lambda_{\min}(\bs\Sigma_{i})\leq\lambda_{\max}(\bs\Sigma_{i})\allowbreak\leq C<\infty $ a.s., (c) For all $i$, $\operatorname{E}[\varepsilon_{i}^4|\bs Z]\leq C<\infty$ a.s.\ and $\operatorname{E}[\Vert\bs\eta_{i}\Vert^4|\bs Z]\leq C<\infty$ a.s. (d) Conditional on $\bs Z$, $\bs\varepsilon\overset{(d)}{=}\bs D_{r}\bs\varepsilon$ with $\bs D_{r}$ as in Assumption (ref).

Part (a)--(c) are similar to those made by crudu2021inference, mikusheva2021inference and matsushita2020jackknife for their jackknife tests. Part (d) ensures that the moment conditions are reflection invariant.

To obtain the limiting distribution of the first order conditions of the CU objective function, we make the following assumption on the IV model in (ref).

assumptionConsider $\bs\eta_{i}$ and $\varepsilon_{i}$ as in (ref). Then, $ \bs\eta_{i} = \varepsilon_{i}\bs a_{i} + \bs u_{i},$ where $\bs a_{i}=\bs\sigma_{21i}/\sigma_{i}^2$, and, conditional on $\bs Z$, $\{\bs u_{i},\varepsilon_{i}\}$ are mutually independent.

This assumption can also be found in bekker2003finite, dokotchatoka2020exogeneity, and fraizier2024weak, but it is nevertheless a strong assumption. It is for example satisfied if $(\varepsilon_{i},\bs\eta_{i}')$ is multivariate normal. We use Assumption (ref) to write

equation[equation omitted — 119 chars of source]

for $\bar{\bs z}_{i}=\bs z_{i}'\bs\Pi$ and stack these in the matrix $\bar{\bs Z}$. Define $\boldsymbol{x}_{(l)}$ as the column vector $(x_{1h},\ldots,x_{nh})'$. The score of a half times the CU objective function is

equation*[equation* omitted — 246 chars of source]

Under Assumption (ref), and using the decomposition of $\bs\eta_i$ in Assumption (ref) we find that, conditional on $\bs Z$, $S_{l}(\bs\beta_{0})\overset{(d)}{=}S_{l,r}(\bs\beta_{0})$, with

equation[equation omitted — 446 chars of source]

where $\boldsymbol{a}_{(l)} = (a_{1l},\ldots, a_{nl})'$ and $\bs P = \bs P(\bs\beta_{0})$.

Note that the limiting distribution of the score is not covered by the results in chao2012asymptotic as it contains products of Rademacher variables up to order three. To describe the joint limiting distribution of the AR statistic and the score, we need the following assumptions.

assumption(a) $\sum_{i=1}^{n}\|\bar{\bs z}_{i}\|^2/n\leq C<\infty$ a.s.n., (b) $\displaystyle{\max_{i=1,\ldots,n}}\|\bar{\bs z}_{i}\|^2/n\rightarrow_{a.s.}0$, (c) $\displaystyle{\max_{i=1,\ldots,n}}\|\bar{\bs Z}'\bs V\bs D_{\varepsilon}\bs e_{i}\|^2/n\rightarrow_{a.s.}0,$ (d) $0<C^{-1}\leq\operatorname{\lambda_{\mathrm{min}}}(\frac{1}{n}\bs Z'\bs Z)\leq\operatorname{\lambda_{\mathrm{max}}}(\bs Z'\bs Z/n)\leq C<\infty$ a.s.n., $0<C^{-1}\leq\operatorname{\lambda_{\mathrm{min}}}(\bs Z'\bs D_{\varepsilon}^2\bs Z/n)\leq\operatorname{\lambda_{\mathrm{max}}}(\bs Z'\bs D_{\varepsilon}^2\bs Z/n)\leq C<\infty$ a.s.n., (e) $P_{ii}\leq C<1$ a.s.n.

Part (a) and (b) are standard assumptions under many instruments. Part (a) also appears in chao2012asymptotic and hausman2012instrumental, who instead of (b) require $n^{-2}\sum_{i=1}^{n}\|\bar{\bs z}_{i}\|^4\rightarrow_{a.s.}0$. We see that this condition is implied by Assumption (ref) parts (a) and (b). In particular, (b) is a Lyapunov condition needed for the CLT we employ. Part (c) is another Lyapunov condition needed for the CLT under heteroskedasticity. Part (d) ensures that $\bs Z(\bs Z'\bs D_{\varepsilon}^2\bs Z)^{-1}\bs Z'$ has bounded eigenvalues $a.s.n.$ Part (e) is a `balanced design' assumption, e.g. chao2012asymptotic. Finally, note that we make no assumption on $\bs\Pi$, allowing for weak and even irrelevant instruments.

The joint limiting distribution of the AR statistic and the score evaluated at the true parameter $\bs\beta_{0}$ is given in the following theorem.

theoremLet $\bs T(\bs\beta_{0}) = (k^{-1/2}(\operatorname{AR}(\bs\beta_{0})-k), \sqrt{n} \bs S(\bs\beta_{0}))'$. Under Assumptions (ref) to (ref), when $n\rightarrow\infty$ and $k/n\rightarrow\lambda\in (0,1)$, $\hat{\bs\Sigma}_{n}(\bs\beta_{0})^{-1/2}\bs T(\bs\beta_{0}) \rightarrow_{d} N(\bs 0,\bs I_{p+1})$, with \begin{equation} \hat{\bs\Sigma}_n(\bs\beta) = \begin{pmatrix} \hat\sigma_{n}^2(\bs\beta) & \big[\hat{\bs\Sigma}_n(\bs\beta)\big]_{2:p+1,1}' \\ \big[\hat{\bs\Sigma}_n(\bs\beta)\big]_{2:p+1,1}&\hat{\bs\Omega}(\bs\beta) \end{pmatrix}, \end{equation} an unbiased and consistent estimator of \begin{equation} \bs\Sigma_{n}(\bs\beta_0)=\operatorname{var}(\bs T(\bs\beta_{0})|\mathcal{J})=\begin{pmatrix} \sigma_{n}^2(\bs\beta_0) & \big[\bs\Sigma_n(\bs\beta_0)\big]_{2:p+1,1}' \\ \big[\bs\Sigma_n(\bs\beta_0)\big]_{2:p+1,1}&\bs\Omega(\bs\beta_0) \end{pmatrix}. \end{equation} Both the variance and the estimator are detailed in Appendix (ref).
proofSee Appendix (ref).

In Appendix (ref) we show that the variance of the score, $\bs\Omega(\bs\beta_0)$, can be decomposed into terms that are present in the case with a limited number of moments, $\bs\Omega^{L}(\bs\beta_0)$, and additional terms labeled $\bs\Omega^{H}(\bs\beta_{0})$ that appear due to the presence of many moments. That is, $\bs\Omega(\bs\beta_0)=\bs\Omega^{L}(\bs\beta_0)+\bs\Omega^{H}(\bs\beta_0)$. The most important property of the additional variance terms in $\bs \Omega^{H}(\bs\beta_{0})$ is given by the following result.

lemmaSuppose that Assumptions (ref) to (ref) hold, and furthermore, $\max_{i=1,\ldots,n}P_{ii}\leq 0.9$ and $ n^{-1}\sum_{i=1}^{n}V_{ii}P_{ii}>0$. Then, $\bs\Omega^{H}(\bs\beta_{0})$ from Theorem (ref) is negative definite.
proofSee Appendix (ref).

Lemma (ref) implies that the use of the conventional inference procedures based on the score will be conservative. The condition that $n^{-1}\sum_{i=1}^{n}V_{ii}P_{ii}>0$ makes it clear that this is more likely to occur, and asymptotically only occurs, when the number of instruments is a nonnegligible fraction of the sample size. For instance, if $\{\varepsilon_{i}\}_{i=1}^{n}$ is itself a sequence of Rademacher variables and $\bs Z$ has independent elements with finite fourth moment. \citet*{bai2007asymptotics} show that $\frac{1}{n}\sum_{i=1}^{n}V_{ii}P_{ii} - k^2/n^2 \rightarrow_{a.s.} 0$, and hence (c) requires that $k/n\rightarrow\lambda>0$.

Simulation results

In this section we assess the finite sample properties of the proposed tests through a simulation study on the panel tobit model from honore1992trimmed discussed in Example (ref), the quantile IV model from chernozhukov2006finite discussed in Example (ref), and the linear IV model discussed in Example (ref). To contrast the tests developed under many moment sequences with those developed under fixed moment conditions, we label the former as many instrument (MI) and the latter as fixed-$k$ tests.

Panel tobit

Consider the censored panel data model as in honore1992trimmed

equation[equation omitted — 172 chars of source]

where $y_{it}$ is the censored version of a latent variable $y_{it}^*$ that relates to regressors $\bs x_{it}\in\mathbb{R}^{p}$, unobserved individual fixed effects $\alpha_{i}$ and an error $\varepsilon_{it}$.

honore1992trimmed assumes that for each individual $i$ the $\varepsilon_{it}$ are independent and identically distributed over time $t$ conditional on $\bs x_{i1}$, $\bs x_{i2}$ and $\alpha_i$, allowing for heteroskedasticity across individuals. Under this assumption, or conditional exchangeability more generally, the difference between $y_{it_1}^*$ and $y_{it_2}^*$ for some $t_1$ and $t_2$, as denoted by $\Delta_{t_1,t_2}y_{i}^*$, conditional on $\bs x_{it_1}$ and $\bs x_{it_2}$ is symmetrically distributed around $\Delta_{t_1,t_2}\bs x_{i}'\bs\beta=(\bs x_{it_1}-\bs x_{it_2})'\bs\beta$. Furthermore, if $y_{it_1}^*$ and $y_{it_2}^*$ are both greater than zero, neither is affected by the censoring and therefore also the observed values are unaffected by the censoring. honore1992trimmed combines these two observations to define two sets of outcomes for ($y_{it_1}$,$y_{it_2}$)

equation*[equation* omitted — 889 chars of source]

Conditional on $(x_{it_1},x_{it_2})$, ($y_{it_1}$,$y_{it_2}$) falls with equal probability in $A_{1i,t_1,t_2}$ and $B_{1i,t_1,t_2}$, which leads to the following reflection invariant moment condition

equation[equation omitted — 198 chars of source]

Similarly, honore1992trimmed notes that, conditional on $(\bs x_{it_1},\bs x_{it_2})$, the expected vertical distance from $(y_{it_1},y_{it_2})$ given that $(y_{it_1},y_{it_2})$ lies in $A_{1i,t_1,t_2}$, to the $45^\circ$ line through $(\Delta_{t_1,t_2}\bs x_{i}'\bs\beta,0)$ if $\Delta_{t_1,t_2}\bs x_{i}'\bs\beta\geq0$ or through $(0,-\Delta_{t_1,t_2}\bs x_{i}'\bs\beta)$ if $\Delta_{t_1,t_2}\bs x_{i}'\bs\beta<0$, equals the expected horizontal distance from $(y_{it_1},y_{it_2})$ given that $(y_{it_1},y_{it_2})$ lies in $B_{1i,t_1,t_2}$, to the same line. These distances are given by $-(y_{it_1}-y_{it_2}-\Delta_{t_1,t_2}\bs x_{i}'\bs\beta)$ and $(y_{it_1}-y_{it_2}-\Delta_{t_1,t_2}\bs x_{i}'\bs\beta)$, which leads to the second reflection invariant moment condition

equation[equation omitted — 636 chars of source]

The sets in which the observations are unaffected by censoring can be slightly enlarged, to yield two additional moment conditions. These moment conditions are also the difference between either indicators of events with equal probability or of equal expected distances, and hence are also reflection invariant. To be precise, define

equation*[equation* omitted — 752 chars of source]

Then in addition to (ref) and (ref), the following reflection invariant moment conditions hold

equation*[equation* omitted — 803 chars of source]

Any two periods $t_1\neq t_2$ can be used to obtain moment conditions, which brings the potential number of moment conditions per individual to $4\binom{T}{2}p$. In what follows, however, we will not use the same time period twice for the MI AR, since although the moment conditions are all reflection invariant on their own, they are not necessarily jointly invariant when the same period is used twice. To illustrate this, note that the invariance of the moment conditions stems from differences in exchangeable errors. Then take $\varepsilon_{it}$ discrete with $\operatorname{P}(\varepsilon_{it}=1)=2/3$ and $\operatorname{P}(\varepsilon_{it}=-2)=1/3$, and independent over $t=1,2,3$. In this case, $(\varepsilon_{it}-\varepsilon_{it'})$ is reflection invariant, but $(\varepsilon_{i1}-\varepsilon_{i2},\varepsilon_{i2}-\varepsilon_{i3})$ is not, as for example $\operatorname{P}((\varepsilon_{i1}-\varepsilon_{i2},\varepsilon_{i2}-\varepsilon_{i3})=(0,-3))=4/27$, whereas $\operatorname{P}((\varepsilon_{i1}-\varepsilon_{i2},\varepsilon_{i2}-\varepsilon_{i3})=(0,3))=2/27$. Hence, we will use $4\lfloor T/2\rfloor p$ moment conditions.

figure[figure omitted — 716 chars of source]
figure[figure omitted — 728 chars of source]

We use the simulation design in honore1992trimmed to study the size and power of the MI AR test compared to benchmarks that we explain below. We make two adaptations to the DGP, such that $T$ and $p$ can be flexible and to include heteroskedasticity. We generate $n=200$ observations from (ref), where $x_{1,it}=\alpha_i+\eta_{it}$, where $\alpha_{i},\eta_{it}\sim N(0,1)$. The other regressors $x_{l,it}$, $l=2,\dots, p$ are also standard normally distributed and $p$ varies over $\{2,\ldots,6\}$. Finally, $\varepsilon_{it}\sim N(0,\alpha_{i}^2)$. The results are averaged over $10,000$ draws from the DGP.

As benchmarks we include the fixed-$k$ AR test that uses the same moment conditions as the MI AR test, the AR test that uses all possible moment conditions, and the test based on honore1992trimmed's (honore1992trimmed) panel tobit estimator. For the latter we used the STATA package as on the author's website. To the best of our knowledge, there are few other tests for coefficients in panel tobit models that are as generally applicable as honore1992trimmed's (honore1992trimmed) panel tobit estimator.

Figure (ref) shows the rejection rates of the MI AR test and the benchmarks when testing the true null hypothesis $H_0\colon\bs\beta=\bs\iota$ at a $5\%$ significance level. The number of moment conditions increases with $T$ and $p$. For $T=2$ all moment conditions are jointly reflection invariant, hence the two fixed-$k$ AR tests have identical rejection rates. These rejection rates are well below the desired $5\%$ rate and decreasing with the number of moment conditions. For $T=6$ the fixed-$k$ AR test that uses all moment conditions never rejects. Also the rejection rates of the fixed-$k$ AR statistic that uses only the invariant moment conditions are too low and again decreasing with the number of moments. honore1992trimmed's (honore1992trimmed) panel tobit estimator on the other hand is oversized and more so with many moment conditions. Only the MI AR statistic is close to size correct uniformly over the number of moment conditions.

Next, we investigate the power of the MI AR test relative to the fixed-$k$ AR that uses all moment conditions and the fixed-$k$ AR test that uses only the invariant moment conditions. We do not include test based on honore1992trimmed's (honore1992trimmed) panel tobit estimator as it was not size correct in Figure (ref). In Figure (ref) we show the rejection rates when testing the marginal hypothesis $H_{0}\colon\beta_{1}=1$ over $1,000$ data sets when we vary $\beta_{1}$ in the DGP from $-0.5$ to $1.5$. The other coefficients in $\bs\beta$ remain fixed at $1$. We reject the null hypothesis only if the statistic for $\beta_{1}=1$ exceeds the critical value for all values of $\beta_{2}$ to $\beta_{p}$. One can alternatively view this as, for a given value of $\beta_1$, minimizing the AR and MI AR statistics over $\beta_2$ to $\beta_p$ and checking whether the minimum exceeds the critical value chernozhukov2009finite. In this case the minimization is by means of a grid search, but other algorithms can be used as well. From the figure we conclude that the identification robust tests have lower rejection rates than the nonrobust test based on the panel tobit estimator. More importantly, the MI AR test has higher power than the fixed-$k$ AR tests and the power difference increases with the number of moment conditions.

figure[figure omitted — 945 chars of source]

Median IV regression

We now consider the median IV regression from Example (ref). This provides a natural setting to study the effect of violations of Assumption (ref) as the moment conditions in the quantile IV regression are symmetrically distributed for the median, but not for other quantiles. We slightly adapt the simulation set-up from chernozhukov2006finite to increase the number of instrumental variables and to have nonreflection invariant moment conditions. To be precise, there is a dependent variable $y_i$ that relates to a single endogenous regressor $x_i$ as $y_i=-1+x_i+\varepsilon_i$. The endogenous regressor relates to $k$ instruments $\bs z_i$ as $x_i=\bs z_i'\bs\iota_k\pi+\nu_i$ for $i=1,\dots,n$. Since the moment conditions consist of a product of $\tau-\mathbbm{1}\{y_i\leq\theta_{1}(\tau)+\theta_{2}(\tau)x_i\}$ with the instruments $\bs z_{i}$, we draw the instruments from a skewed distribution for the moment conditions to no longer be reflection invariant at $\tau\neq 1/2$. We draw from a Gamma distribution with shape parameter $1/\zeta$, rate parameter $\sqrt{1/(\sigma^2\zeta)}$ and we subtract $\sqrt{\sigma^2/\zeta}$, such that the distribution has mean zero, variance $\sigma^2$ and skewness governed by $\zeta$. Figure (ref) in Appendix (ref) shows how the skewness increases with $\zeta$. We set $\zeta=3/2$ and $\sigma^2=1$. The errors $\varepsilon_i$ and $\nu_i$ are standard normally distributed with correlation $\rho=0.8$. The first stage coefficient $\pi=0.5$. The number of observations is set to $n=100$. The results are averaged over $1,000$ draws from the DGP.

We fit the conditional quantile model $q(x_i,\tau;\bs\theta)=\theta_{1}(\tau)+\theta_{2}(\tau)x_i$, such that the true values for the intercept and slope parameters are $\theta_{1}(\tau)=-1+\Phi^{-1}(\tau)$ for $\Phi$ the CDF of the standard normal distribution, and $\theta_{2}(\tau)=1$. We use the moment conditions as suggested in Example (ref) with $\bs\Psi(\bs z_{i})=\bs z_{i}$ to test values of $\bs\theta$ with a fixed-$k$ AR and the MI AR test. As benchmarks, we include the exact test by chernozhukov2009finite with $1,000$ draws to approximate the critical values, the weak identification robust score test by jun2008weak, a GMM based test with the limiting distribution derived as in andrews1994empirical and arellano2009gmm, the inverse quantile regressor (IQR) by chernozhukov2006instrumental and the smoothed IVQR by kaplan2017smoothed. See Appendix (ref) for details of these benchmarks.

The left panels of Figure (ref) show the size of the fixed-$k$ AR, MI AR and the test by chernozhukov2009finite for $k=10$ and $k=30$ against the chosen quantile. The MI AR test appears relatively robust to small deviations from the invariance assumption. Only for $k=30$ and quantiles higher than $0.8$ there is a sudden increase in the rejection rate, but size distortions stay limited. The fixed-$k$ AR is conservative due to the many instruments, while the test by chernozhukov2009finite attains the nominal size for all quantiles.

The right panels of the same figure show the size of the other benchmarks. It strikes that, except for the test based on the GMM estimator when $k=10$, all tests are oversized.

Further results on size and power at $\tau=1/2$ can be found in Appendix (ref) showing that the MI AR test offers size and power close to the exact test.

Linear IV model

For the linear IV, we base our simulations on hausman2012instrumental. We adapt the DGP slightly to allow for clustering in the data and vary the degree of heteroskedasticity. We generate $n=800$ observations from $y_i=\alpha+\beta x_{i}+\varepsilon_i$, $i=1,\dots,n$, where $\alpha=0$. We test $H_{0}\colon\beta = 0$. The regressor is generated as $x_{i}=\pi \tilde{z}_{i}+\eta_i$, whose distribution we detail below. In addition to the instrumental variable $\tilde{z}_i$ with coefficient $\pi$, there are $k$ other instrumental variables stacked in $\bs z_i=(1,\tilde{z}_i, \tilde{z}_i^2, \tilde{z}_i^3, \tilde{z}_i^4, \tilde{z}_i D_{i1},\dots, \tilde{z}_iD_{ik-4})'$, where the $D_{ij}$, $j=1,\dots,k-4$, are independent Bernoulli(1/2) random variables.

To allow the data to be clustered, we divide the observations over $H=100$ clusters containing between 4 and 12 observations. The random variables $\tilde{z}_i$, $\eta_i$, $v_{1i}$ and $v_{2i}$ consist of a cluster common component and an idiosyncratic component weighted by $\lambda$. That is, $\tilde{\bs z}_{[h]}=\sqrt{\lambda}\tilde{\bs z}_{[h]}^\text{ind}+\sqrt{1-\lambda}\tilde{z}_{h}^\text{cl}$, $\bs\eta_{[h]}=\sqrt{\lambda}\bs\eta_{[h]}^\text{ind}+\sqrt{1-\lambda}\eta_{h}^\text{cl}$, $\bs v_{1,[h]}=\sqrt{\lambda}\bs v_{1,[h]}^\text{ind}+\sqrt{1-\lambda}v_{1,h}^\text{cl}$ and $\bs v_{2,[h]}=\sqrt{\lambda}\bs v_{2,[h]}^\text{ind}+\sqrt{1-\lambda}v_{2,h}^\text{cl}$ for $h=1,\dots, H$. For the independent setting, we set $\lambda=1$, for the clustered setting $\lambda=0.5$. We draw the cluster common and idiosyncratic components of $z_{i}$ and $\eta_i$ from standard normal distributions. We generate $v_{1i}^\text{ind}\sim N(0,(\tilde{z}_{i}^\text{ind})^\kappa)$ and $v_{1h}^\text{cl}\sim N(0,(\tilde{z}_{h}^\text{cl})^\kappa)$, where $\kappa$ controls the degree of heteroskedasticity. We start with $\kappa=2$ as in hausman2012instrumental and increase it to $\kappa = 6$. The cluster common and idiosyncratic components of $v_{2i}$ come from a normal distribution with mean zero and variance $0.86^2$. Finally, the second stage error generated as $\varepsilon_i=\rho\eta_i+\sqrt{(1-\rho^2)/(\phi^2+0.86^4)}(\phi v_{1i}+0.86v_{2i})$, where $\rho=0.3$ is the degree of endogeneity and $\phi=1.38072$ as in the implementation by bekker2015jackknife. Appendix (ref) considers a DGP with skewed moment conditions where the AR test rejects around 12% in the most extreme scenario, while the score test rejects 8% of the cases.

figure[figure omitted — 951 chars of source]

We investigate the size of the MI AR and MI score, and compare this to the fixed-$k$ AR, the fixed-$k$ score, the jackknife AR, the jackknife score test and the HFUL when the data are independent. See Appendix (ref) for details on the implementation of the different tests. The rejection rates are calculated over $10,000$ data sets.

The left panel of Figure (ref) shows the rejection rate of the different tests when testing $H_0\colon\beta=0$ at a $5\%$ level for $k$ ranging from $2$ to $50$ and $\pi=\sqrt{8/n}$ as in hausman2012instrumental. For weak instruments all identification robust tests are close to size correct as long as the number of instruments is not too high. In line with the theory from Sections (ref) and (ref), the rejection rates of the fixed-$k$ AR and score tests drop when the number of instruments increases. The MI and jackknife AR and score tests remain size correct. HFUL is conservative for few weak instruments, but becomes oversized for many instruments.

figure[figure omitted — 1,006 chars of source]

In the right panel of Figure (ref), we consider a DGP in which the data are clustered. The figure shows the rejection rates of the cluster fixed-$k$ AR, the cluster MI AR and the cluster jackknife AR for the same hypothesis and 10,000 data sets. Appendix (ref) details the exact implementation. This figure is similar to the figure for independent data in the sense that the cluster fixed-$k$ AR test is conservative, whereas the cluster MI AR and cluster jackknife AR tests are size correct.

figure[figure omitted — 920 chars of source]

We next study the power against $H_0\colon\beta=0$ of the tests from the previous subsection when varying the true $\beta$ between $-2$ and $2$. We fix $\pi=\sqrt{128/n}$ to attain reasonable power, set the number of instruments $k=30$ and vary the degree of heteroskedasticity by choosing $\kappa\in\{2,6\}$. The values of the other parameters in the DGP are unaltered.

Figure (ref) shows the power for independent data. We focus first on the AR statistics, which we have marked with a circle. In the left panel, with mild heteroskedasticity, we see that although the power of the three different types of AR tests are comparable, the MI AR has a slight edge over the other two. In the right panel, with stronger heteroskedasticity, the power differences are bigger. Especially the jackknife AR has low rejection rates.

The score tests are marked with a square. In both panels, they outperform the AR tests, which might be due to the relatively strong instruments. In the left panel we furthermore see that the score tests have comparable power to HFUL. In the right panel we again have higher power for the fixed-$k$ and MI tests have higher power than for the jackknife test. In this case also HFUL does not perform as well as the fixed-$k$ and MI score.

For clustered data, as shown in Figure (ref), the relative performance of the MI AR test is more nuanced. It still outperforms the fixed-$k$ test and by a larger margin than for independent data. The jackknife variant is clearly more powerful when the degree of heteroskedasticity is low, although the difference disappears if we increase the degree of heteroskedasticity in the right panel.

Empirical application: the effect of bank bias on stability

We revisit the study by langfield2016bank on the effect of the relative concentration of finance in the banking sector in Europe on financial stability. The effects of this concentration, also called “bank bias", are ambiguous: some argue that a larger banking sector increases systemic risk, whereas others argue that larger banking mitigates systemic risk langfield2016bank. To study the question empirically, langfield2016bank construct a panel with 3981 observations on 467 banks in 20 countries observed over maximally 12 years.

The outcome variable is the bank-specific systemic risk intensity defined as the bank's SRIKS, a measure of systemic risk provided by the New York University's \href{https://vlab.stern.nyu.edu/}{Volatility Laboratory}, over the bank's total assets. The variable of interest is the country-specific bank-market ratio, defined as the amount of bank assets over the stock and private bond market capitalization. langfield2016bank estimate the effect of the bank-market ratio on the bank systemic risk intensity controlling for bank and time fixed effects, a crisis indicator and an interaction term between this indicator and the bank-market ratio.

Since a negative systemic risk intensity does not contribute to systemic risk, the outcome variable is censored at zero. langfield2016bank therefore use honore1992trimmed's (honore1992trimmed) panel tobit specification in their Table 2. As shown in Section (ref), the panel tobit uses many symmetric moment conditions to estimate the effect of interest. In this section we use the same moment conditions in identification robust tests. We the model,

equation[equation omitted — 159 chars of source]

where $y_{it}$ is the censored systemic risk intensity of bank $i$ in year $t$, $\alpha_i$ and $\nu_{t}$ are bank and year fixed effects, $x_{it}$ is the bank-market ratio of the country that bank $i$ is in at time $t$, $d_{it}$ is a crisis indicator, and $\varepsilon_{it}$ is an error term that is conditionally exchangeable over time.

We take differences of the observations per bank to get rid of the bank fixed effects. We then calculate the symmetrically distributed events as in Section (ref), which we multiply with the corresponding difference of the bank-market ratio, $x_{it}$, the crisis dummy, $d_{it}$, and their interaction. Since there may be dependence between the moment conditions per bank, we cluster the moment conditions on the bank level. This leaves us with a total of $12$ moment conditions, which is large relative to the 467 banks in the sample.

The AR tests jointly test hypothesized values for all parameters in the model. As we are mainly interested in $\beta_{1}$ and $\beta_{3}$, the coefficients on the bank market ratio and its interaction with the crisis dummy, we conduct a grid search only over $\beta_{1}$ and $\beta_{3}$, and minimize the AR statistics over all other parameters, which is computationally more efficient chernozhukov2009finite.

figure[figure omitted — 846 chars of source]

Figure (ref) shows the $95\%$ confidence regions for the effect of bank-market ratio and its interaction with a crisis dummy on systemic risk intensity, based on inverting the AR and MI AR statistics and honore1992trimmed's (honore1992trimmed) panel tobit estimator. In the left panel a crisis is defined as a year in which a country's real house prices drop by at least $10\%$, wheres in the right panel a crisis is defined as a year in which a country's real stock prices drop by at least $20\%$. Given that in the simulations in Section (ref), the AR statistic that uses all moment conditions has lower power compared to the AR statistic that only uses the reflection invariant moment conditions, we limit ourselves to the latter. To mitigate the effect of possible inaccuracies in the numerical optimization, we exclude a value $(\beta_1,\beta_3)$ from the confidence region only when the minimum of all statistics in a $3\times 3$ grid around the hypothesized point is above the critical value. This enlarges the confidence regions, but makes them more reliable. The confidence region for the panel tobit estimator is not smoothed and matches the results of columns I and III in Table 2 of langfield2016bank.

The figure shows that the confidence regions from the AR and MI AR statistics are considerably larger than those obtained with the panel tobit estimator. When the crisis dummy captures a decrease in house prices, the robust confidence regions both contain $(0,0)$, whereas the confidence region for the panel tobit estimator excludes a zero effect for the interaction term. In contrast, when the crisis dummy captures a stock crisis, the robust confidence regions only contain zero for the bank-market ratio. The interaction term is significantly different from zero, indicating that countries with a high bank-market ratio have lower financial stability during crises, than countries with a low bank-market ratio.

Importantly, for the grids considered here, the confidence regions of the AR are larger than those of the MI AR. For the housing and stock market crises, the AR confidence regions are 10% and 15% larger than the MI AR confidence regions.

Conclusion

We develop a new approach for identification-robust inference under many moment conditions based on the continuous updating AR GMM statistic. Since the presence of a large-dimensional weighting matrix prohibits the use of conventional asymptotic approximations, we show how reflection invariance in the moment conditions can be used to establish asymptotic results. The many moment condition correction to the variance of the AR statistic negative, indicating that conventional approximations lead to overly wide confidence intervals. We also provide an analysis of the linear IV model that provides conditions under which the same is true for the score statistic. Monte Carlo simulations further show close to nominal size of the developed procedures regardless of the identification strength and the number of moment conditions, as well as an increase in power under many moment conditions. Empirically, we use the corrected AR test to construct confidence regions for effect of concentration of financial activities in banks on systemic risk intensity and find evidence for a negative effect of bank concentration on systemic risk during stock market crises.

Acknowledgements

We thank the editor Michael Jansson, the associate editor and two referees for their detailed comments. We also thank Federico Crudu, Paul Elhorst, Jinyong Hahn, Bruce Hansen, Art\={u}ras Juodis, Lingwei Kong, Nick Koning, Zhipeng Liao, Xinwei Ma, Sophocles Mavroeidis, Jack Porter, Andres Santos, Xiaoxia Shi, Yixao Sun, Aico van Vuuren, and Tom Wansbeek for many insightful comments. Tom Boot's work was supported by the Dutch Research Council (NWO) under grant VI\allowbreak.Veni\allowbreak.201E\allowbreak.11.