EconBase
← Back to paper

Robust Permutation Tests in Linear Instrumental Variables Regression

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

74,245 characters · 11 sections · 100 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Robust Permutation Tests in Linear Instrumental Variables Regression

\def\spacingset#1{ {1.7}} \spacingset{0.5}

\if11 \fi

\if01 \fi

abstractThis paper develops permutation versions of identification-robust tests in linear instrumental variables regression. Unlike the existing randomization and rank-based tests in which independence between the instruments and the error terms is assumed, the permutation Anderson-Rubin (AR), Lagrange Multiplier (LM) and Conditional Likelihood Ratio (CLR) tests are asymptotically similar and robust to conditional heteroskedasticity under standard exclusion restriction i.e. the orthogonality between the instruments and the error terms. Moreover, when the instruments are independent of the structural error term, the permutation AR tests are exact, hence robust to heavy tails. As such, these tests share the strengths of the rank-based tests and the wild bootstrap AR tests. Numerical illustrations corroborate the theoretical results. {\it Keywords:} Anderson-Rubin statistic, Asymptotic and exact tests, Conditional likelihood ratio statistic, Heteroskedasticity, Identification, Lagrange Multiplier statistic

\spacingset{1.9}

Introduction

The use of instrumental variables (IVs) is widespread across many disciplines. One of the main inferential issues that require taking careful account in the IV regression is to protect against a possible weak correlation between instruments and endogenous regressors e.g. treatment variable. To this end, it is desirable to use identification-robust tests that are similar, at least asymptotically, when the instruments have an arbitrary degree of explanatory power (identification strength) for the endogenous regressor, and have good power when the instruments are informative.\footnote{For a comprehensive treatment of similar tests, see Lehmann-Romano(2005) and Linnik(2008).} This paper proposes permutation (randomization) test versions of heteroskedasticity-robust Anderson-Rubin(1949)'s AR, Kleibergen(2002)'s LM and Andrews-Guggenberger(2019)'s CLR tests in IV regression within a super-population framework.\footnote{The LM test is a score test that uses an outer-product-of-the-gradient information matrix estimator, and the CLR test is a heteroskedasticity-robust version of the CLR test proposed by Moreira(2003) which is a LR test with Neyman structure, see Lehmann-Romano(2005) for the latter.} We demonstrate the uniform asymptotic similarity of these tests under standard conditions and characterize the conditions under which the permutation AR tests are exact. Permutation inference is attractive because of its exactness when relevant assumptions hold. It also bears a natural connection to IV methods since IVs provide an exogenous variation independent of the unobserved confounders for the endogenous regressor in the IV regression and the permutation inference typically seeks to exploit independence of two sets of variables by permuting the elements of one variable holding the other fixed to replicate the distribution of a test statistic at hand.

Despite the link, the literature on randomization inference in IV models is scarce. Imbens-Rosenbaum(2005) develop exact permutation tests in an IV setting under a finite population framework, where the instruments (or transformations thereof, such as ranks) are permuted while the quantities that do not depend on the IVs are held fixed. Given the advantages of the permutation method in an IV setting as exemplified by Imbens-Rosenbaum(2005), we broaden its scope by proposing the permutation versions of the trinity of aforementioned identification-robust test statistics in a super-population framework. The contributions of the paper are as follows. First, we propose two permutation AR (PAR) statistics for the IV setting, denoted as \(\mathrm{PAR}_1\) and \(\mathrm{PAR}_2\). These directly extend the work of Diciccio-Romano(2017), who develop permutation tests for regression coefficients in heteroskedastic linear regression, but do not study the IV regression. Specifically, the \(\mathrm{PAR}_1\) statistic permutes the rows of the instrument matrix affecting the endogenous regressor, while the \(\mathrm{PAR}_2\) statistic permutes the null-restricted residuals in the structural equation. We establish the finite sample validity of the PAR tests under an independence assumption similar to that in Andrews-Marmer(2008). We then show their asymptotic validity under the usual exogeneity condition, allowing for conditional heteroskedasticity. As such, our result differs from the finite-population results of Imbens-Rosenbaum(2005). Second, we propose a permutation LM (PLM) test. Constructing its permutation version is nontrivial because the LM statistic is a nonlinear score statistic. To achieve this, we permute the residuals from the first-stage estimation of the reduced-form equation and generate a permutation endogenous regressor.\footnote{See Freedman-Lane(1983) and Diciccio-Romano(2017) for residual permutation tests in linear models.} By combining this regressor with the permuted null-restricted residuals in the structural equation, we construct the PLM statistic.\footnote{This approach is distinct from the score statistic construction proposed by Hemeriketal2020 based on sign-flipping in generalized linear models.} Third, we propose a permutation CLR (PCLR) test. This test generates the conditional permutation distribution by permuting the null-restricted residuals of the structural equation while keeping the nuisance parameter estimates fixed. Although the PCLR test shares some similarities with the conditional permutation tests proposed by Rosenbaum1984 and Hennessyetal2016 in different contexts, such a permutation test has not been considered in the IV literature. Fourth, as a main technical contribution, we show the uniform asymptotic similarity of the proposed permutation tests and derive their asymptotic power under local alternatives and strong identification. To the best of our knowledge, these results have not been shown before in the context of permutation/randomization inference. To establish uniform asymptotic similarity, we adapt the asymptotic results from Andrews-Guggenberger(2019), Andrews-Guggenberger(2019s) and Andrews-Cheng-Guggenberger(2020) to the permutation inference setting.

Fifth, we provide a rigorous analysis of the PAR tests with varying numbers of IVs, identifying the conditions under which these tests maintain asymptotic validity. To substantiate this, we interpret the PAR statistics as double-indexed permutation statistics and make a novel use of the CLT established by Pham-Mocks-Sroka(1989). Such an application appears to be new in the literature.

Finally, we compare the robust permutation tests to both the identification-robust rank-based tests of Andrews-Marmer(2008) and Andrews-Soares(2007), and the wild bootstrap tests of Davidson-MacKinnon(2012) through simulations. In terms of controlling Type I error, we find that the robust permutation tests outperform the rank-based tests under the standard exclusion restriction, and perform on par with or better than the wild bootstrap tests in the heteroskedastic designs considered. {\flushleft{\bf Discussion of the related literature.}} The literature on linear IV model is vast. Earlier surveys on this topic are provided by Stock-Wright-Yogo(2002), Dufour(2003), and Andrews-Stock(2007), and recent contributions include Moreira-Moreira(2019), Young(2020) and the survey of Andrews-Stock-Sun(2019). Andrews-Marmer(2008) develop rank-based AR tests under the independence assumption between the instruments and the structural error term. Under their assumption, the $\mathrm{PAR}_1$ test proposed here is exact while the $\mathrm{PAR}_2$ test is so when there are no included non-constant exogenous variables. Andrews-Soares(2007) consider a rank-version of the CLR statistic of Moreira(2003). Their test is robust to heavy-tailed errors but not to heteroskedasticity. Rosenbaum1984 and Hennessyetal2016 propose conditional permutation tests in the causal inference context, which is different from our setting. However, the strategies used for the PCLR test validity may prove useful for showing the asymptotic validity of their tests under weak conditions. Also, the PCLR test is quite distinct from the conditional test based on rerandomization by LiDingRubin2018, where the conditioning statistic is a permutation quantity as opposed to the sample quantity considered here. Hemeriketal2020 propose a score statistic in generalized linear models based on sign-flipping. The residual permutation construction of the PLM test provides an alternative method for constructing a permutation score statistic in their framework. This paper complements the literature on bootstrap inference in the linear IV model, including works such as Davidson-MacKinnon(2008), Moreira-Porter-Suarez(2009), and Davidson-MacKinnon(2012), which document the improved performance of bootstrap inference over asymptotic inference. The key to our results is that the identification-robust PAR, PLM, and PCLR statistics are studentized. The use of studentization to obtain an improved test is not new; a well-known example is the higher-order accuracy of the bootstrap-$t$ confidence interval in the one-sample problem Lehmann-Romano(2005). Other examples include So-Shin(1999), who develop a persistence-robust test for autoregressive models, and Neuhaus(1993), who propose a permutation test for a two-sample problem with randomly censored data. For further extensions and applications of permutation tests based on studentized statistics, see Janssen(1997), Neubert-Brunner(2007), Pauly(2011) and Chung-Romano(2013), Chung-Romano(2016). Several recent papers that consider randomization inference in different contexts include Chernozhukov-Hansen-Jansson(2009), Canay-Romano-Shaikh(2017), Bugni-Canay-Shaikh(2018), Canay-Kamat(2018), Ganong-Jager(2018) and Dufour-Flachaire-Khalaf(2019). {\flushleft{\bf Organization of the paper.}} The paper is organized as follows. Section (ref) introduces the model and the identification-robust test statistics and develops the permutation tests. Section (ref) provides simulation results comparing the performance of alternative test procedures. Section (ref) presents an empirical application. We briefly conclude in Section (ref). Supplemental Appendix contains all the proofs of our theoretical results and additional simulation evidence. We also provide the replication R codes of the empirical applications.

Robust permutation tests

Model setup

We develop permutation-based tests for a restriction on a scalar structural coefficient

equation[equation omitted — 53 chars of source]

in the linear IV regression model:

align[align omitted — 109 chars of source]

where $y=[y_{1},\dots, y_{n}]^{\prime}$ and $Y=[Y_1,\dots, Y_n]^{\prime}$ are $n$-vectors of endogenous variables, $X=[X_1,\dots, X_n]^{\prime}$ is an $n\times p$ matrix of exogenous covariates whose first column is the $n$-vector of ones $\iota=[1,\dots, 1]^{\prime}$, $W=[W_1,\dots, W_n]^{\prime}$ is an $n\times k\ (k\geq 1)$ matrix of IVs that does not include $\iota$ and the variables in $X$, $u=[u_1, \dots, u_n]^{\prime}$ is an $n$-vector of structural error terms, $V=[V_1,\dots, V_n]^{\prime}$ is an $n$-vector of reduced-form error terms; $\gamma$ is a $p$-parameter vector, and $\mathit{\Gamma}$ and $\mathit{\Psi}$ are $k$- and $p$-vectors of reduced-form coefficients, respectively. Let $Z=[Z_1,\dots, Z_n]^{\prime}\equiv M_{X}W$, where $M_A\equiv I_n-P_A\equiv I_n-A(A^{\prime}A)^{-1}A^{\prime}$ for a matrix $A$ of full column rank. For later use, it will be convenient to express the equation (ref) as

align[align omitted — 61 chars of source]

where $\xi\equiv(X^{\prime}X)^{-1}X^{\prime}W\mathit{\Gamma}+\mathit{\Psi}$. We also define $\tilde{u}_i(\theta)\equiv y_i-Y_i^{\prime}\theta-X_i^{\prime}(X^{\prime}X)^{-1}X^{\prime}(y-Y\theta)$, $\tilde{u}(\theta)\equiv [\tilde{u}_1(\theta),\dots, \tilde{u}_n(\theta)]^{\prime}$, $\tilde{Y}=[\tilde{Y}_1,\dots, \tilde{Y}_n]^{\prime}\equiv M_{X} Y$, and

align[align omitted — 422 chars of source]

The first heteroskedasticity and identification-robust test statistic is the AR statistic defined as

equation*[equation* omitted — 117 chars of source]

The null asymptotic distribution of the AR statistic is chi-square regardless of the values of $\mathit{\Gamma}$ because $n^{1/2}\hat{m}$ is asymptotically normal under (ref). Another asymptotically pivotal test statistic is the following heteroskedasticity-robust version of Kleibergen(2002)'s LM statistic

equation[equation omitted — 169 chars of source]

As defined, $\hat{J}$ in (ref) is a vector of residuals from the regression of the sample Jacobian $\hat{G}$ on the sample moment function in (ref). When $\mathit{\Gamma}$ has full rank i.e. the identifying power of $W$ is strong, $\hat{J}$ converges to a matrix of full rank under mild conditions, and the LM statistic is asymptotically chi-squared under (ref). However, even when $\mathit{\Gamma}$ is “small” so that the noise and signal parts of $Y$ are of a similar magnitude or the former dominates the latter, $\hat{J}$ and $\hat{m}$ are asymptotically independent after suitable normalizations, and consequently, the LM statistic is still chi-squared in the limit.

The CLR-type test statistic considered next has an approximate Neyman structure, such that its conditional asymptotic distribution, given a sufficient statistic for the vector of nuisance parameters $\mathit{\Gamma}$ is independent of $\mathit{\Gamma}$. Let

align[align omitted — 274 chars of source]

where $\Omega^\epsilon(\cdot, \cdot)$ is an eigenvalue-adjusted version of the $2\times 2$ matrix defined as

align[align omitted — 389 chars of source]

and the $k\times k$ matrix ${K}_{ij}(\hat{\mathcal{V}})$ is the $(i, j)$ submatrix of ${K}(\tilde{\mathcal{V}})$ given by

align[align omitted — 402 chars of source]

With an appropriate choice of $\hat{\mathcal{V}}$ specified below, the transformation $\hat{\mathcal{T}}$ in (ref) provides an asymptotically sufficient statistic for $\mathit{\Gamma}$ that is independent of the statistic $\hat{\mathcal{S}}$ in (ref). Andrews-Guggenberger(2019), Andrews-Guggenberger(2019s) propose two different versions of $\hat{\mathcal{V}}$, and for the linear IV model, they recommend the following choice of $\hat{\mathcal{V}}$:

align[align omitted — 527 chars of source]

The CLR-type statistic is then defined as

align[align omitted — 262 chars of source]

where $\lambda_{\min}(\cdot)$ denotes the smallest eigenvalue of a matrix. In Supplemental Appendix, we consider another CLR-type statistic of Andrews-Guggenberger(2019) based on a different specification for $\hat{\mathcal{V}}$. The $\text{CLR}$ in (ref) is denoted as $\text{QLR}_P$ in Andrews-Guggenberger(2019s), and it is shown to be asymptotically equivalent to Moreira(2003)'s CLR statistic in the homoskedastic linear IV regression with fixed instruments and multiple endogenous variables under all strengths of identification. \footnotetext{The eigenvalue-adjustment of Andrews-Guggenberger(2019s) is as follows. Let $A$ be a nonzero positive semi-definite matrix of dimension $p\times p$ that has a spectral decomposition $A=N\Delta N^{\prime}$, where $\Delta=\mathrm{diag}(\lambda_1,\dots,\lambda_p)$, $\lambda_1\geq\dots\geq\lambda_{p}\geq 0,$ is the diagonal matrix that consists of the eigenvalues of $A$, and $N$ is an orthogonal matrix of the corresponding eigenvectors. Given a constant $\epsilon>0$, the eigenvalue adjusted matrix is defined as $A^{\epsilon}\equiv N \Delta^{\epsilon}N^{\prime}$, where $\Delta^{\epsilon}\equiv\mathrm{diag}(\max\{\lambda_1,\lambda_{1}\epsilon\},\dots, \max\{\lambda_p,\lambda_{1}\epsilon\})$. They recommend $\epsilon=0.01$. We refer to Andrews-Guggenberger(2019) for further properties of the eigenvalue-adjustment procedure.}

Main results

We begin by recalling the basic notion of randomization tests from Chapter 15.2 of Lehmann-Romano(2005). Let $\mathbb{G}_n$ be the set of all permutations $\pi=[\pi(1), \dots, \pi(n)]$ of $\{1,\dots, n\}$, and denote a permutation version of a generic test statistic $R$ by $\text{PR}=R^\pi$. The distribution function corresponding to permutations in $\mathbb{G}_n$ is

equation[equation omitted — 126 chars of source]

where $1(\cdot)$ is the indicator function. The computation of (ref) and related quantities (at a particular value of $x$) for permutation tests entails computing $n!$ test statistics which is impractical even for $n$ relatively small. So instead, a stochastic approximation based on uniform random draws from $\mathbb{G}_n$ is often used. To this end, let us fix $N\in \{1,\dots,|n!|\}$ and consider a subset of permutations $\mathbb{G}_n'=\{\pi_1,\dots, \pi_N\}\subseteq \mathbb{G}_n$, where $\pi_1$ is the identity permutation, and $\pi_2,\dots, \pi_N$ are i.i.d. uniformly distributed on $\mathbb{G}_n$. The permutation distribution function corresponding to $\mathbb{G}_n'$ is

equation[equation omitted — 129 chars of source]

Let $R^{\pi}_{(1)}\leq \dots\leq R^{\pi}_{(N)}$ be the order statistics of $\{R^\pi: \pi\in\mathbb{G}_n'\}$, and for a nominal significance level $\alpha\in(0,1)$, define

align[align omitted — 274 chars of source]

where $I(\cdot) $ denotes the integer part of a number. Clearly, $N^{+}$ and $N^{0}$ are the number of values $R_{(j)}^{{\pi}}=\mathrm{PR}_{(j)}, j = 1,\dots, N$, that are greater than $R_{(r)}^\pi$ and equal to $R_{(r)}^\pi$, respectively. A randomization test is then defined as

equation[equation omitted — 190 chars of source]

The finite sample results presented below hold for any $N$ but we will assume that $N\to\infty$ as $n\to \infty$ for the asymptotic results. We next describe the robust permutation test statistics. Let $W_\pi=[W_{\pi(1)},\dots, W_{\pi(n)}]^{\prime}$ and $\tilde{u}_{\pi}(\theta)=[\tilde{u}_{\pi(1)}(\theta),\dots, \tilde{u}_{\pi(n)}(\theta)]^{\prime}$ be an instrument matrix and a residual vector obtained by permuting the rows of $W$ and the elements of $\tilde{u}(\theta)$, respectively, for a permutation $\pi\in\mathbb{G}_n'$. Set

equation*[equation* omitted — 269 chars of source]

We consider two heteroskedasticity-robust permutation AR statistics defined as:

align[align omitted — 347 chars of source]

The $\text{PAR}_1$ statistic is based on the permutation of the rows of the instrument matrix $W$, and the $\text{PAR}_2$ statistic uses the permutation of the null-restricted residuals $\tilde{u}(\theta_0)$. The finite sample validity of these tests are shown under the following condition.

assumption$\{(W_i^{\prime}, X_i^{\prime}, u_i)^{\prime}\}_{i=1}^n$ are i.i.d., and $W_i$ and $[X_i^{\prime}, u_i]^{\prime}$ are independently distributed.

Andrews-Marmer(2008) develop exact rank-based AR tests under assumptions similar to Assumption (ref). When $W$ is independent of $X$ and $u$, $W^{\prime}M_X\tilde{u}(\theta_0)$ and $W_\pi^{\prime}M_X\tilde{u}(\theta_0)$ have the same distribution under $H_0:\theta=\theta_0$, so the $\text{PAR}_1$ test is exact. On the other hand, the $\text{PAR}_2$ test is not, in general, exact because $W^{\prime}M_X\tilde{u}(\theta_0)$ and $W^{\prime}M_X\tilde{u}_\pi(\theta_0)$ may not have identical distributions. However, when $X=\iota$ in Assumption (ref), $Z$ and $u_\pi$ are independent, $Z^{\prime}\tilde{u}(\theta_0)$ and $Z^{\prime}\tilde{u}_\pi(\theta_0)$ are identically distributed, and consequently, the $\text{PAR}_2$ test is exact. The following result summarizes the finite sample validity of the PAR tests.

proposition[Finite sample validity] Under Assumption (ref) and $H_0:\theta=\theta_0$, $\operatorname{E}[\phi_n^{\mathrm{PAR}_1}]=\alpha$ for $\alpha\in(0,1)$. If $X=\iota$ in Assumption (ref), then $\operatorname{E}[\phi_n^{\mathrm{PAR}_2}]=\alpha$.

Even if $[X_i^{\prime}, u_i]^{\prime}$ and $W_i$ are not distributed independently but satisfy the orthogonality condition $\operatorname{E}[(W_i^{\prime}, X_i^{\prime})^{\prime}u_i]=0$, we show that the permutation AR tests are asymptotically similar and heteroskedasticity-robust. The argument for permutation LM statistic is more involved because it is difficult to construct an estimator of the reduced-form coefficients $\mathit{\Gamma}$ whose (asymptotic) distribution remains invariant to a permutation of the data. The OLS estimators of the reduced-form coefficients and the corresponding residuals are

equation*[equation* omitted — 208 chars of source]

Now permute the residuals $\hat{V}_{\pi}=[\hat{V}_{\pi(1)}^{\prime},\dots, \hat{V}_{\pi(n)}^{\prime}]^{\prime}$, and let

equation*[equation* omitted — 122 chars of source]

The idea of permuting the residuals to obtain asymptotically valid randomization tests appears in Freedman-Lane(1983) and Diciccio-Romano(2017). The Jacobian estimator used in the permutation LM statistic is

align[align omitted — 225 chars of source]

It is not difficult to see that in strongly identified models where $\mathit{\Gamma}$ is a fixed vector of full rank, $\hat{J}^{\pi}-n^{-1}Z^{\prime}Z\mathit{\Gamma}\stackrel{p^{\pi}}{\longrightarrow} 0$ in probability, where $\stackrel{p^{\pi}}{\longrightarrow}$ denotes the convergence in probability with respect to the probability measure of $\pi$ uniformly distributed over $\mathbb{G}_n$ conditional on the data. The permutation LM statistic is then defined as

equation[equation omitted — 213 chars of source]

Finally, we define the permutation CLR statistic as follows:

align[align omitted — 297 chars of source]

where $\hat{\mathcal{T}}$ is as defined in (ref), and

align[align omitted — 215 chars of source]

Analogously to (ref) and (ref), define the permutation distributions of the statistic $\mathrm{PCLR}=\mathrm{CLR}(\hat{\mathcal{S}}^\pi, \hat{\mathcal{T}})$ given $\hat{\mathcal{T}}$, as $\hat{F}_n^{\mathrm{PCLR}}(x, \hat{\mathcal{T}}) \equiv(n!)^{-1}\sum_{\pi\in\mathbb{G}_n}1(\mathrm{CLR}(\hat{\mathcal{S}}^\pi, \hat{\mathcal{T}})\leq x)$ and

equation[equation omitted — 170 chars of source]

Let $\mathrm{PCLR}_{(r)}(\hat{\mathcal{T}})$ be the $r$-th order statistic of $\{\mathrm{CLR}(\hat{\mathcal{S}}^\pi, \hat{\mathcal{T}}): \pi\in\mathbb{G}_n'\}$ for $r$ defined in ((ref)), and $a^{\mathrm{PCLR}}$ be as in (ref). The nominal level $\alpha$ PCLR test $\phi_n^{\mathrm{PCLR}}$ is 1 if $\mathrm{CLR}(\hat{\mathcal{S}}, \hat{\mathcal{T}})>\mathrm{PCLR}_{(r)}(\hat{\mathcal{T}}),$ $a^{\mathrm{PCLR}}$ if $\mathrm{CLR}(\hat{\mathcal{S}}, \hat{\mathcal{T}})=\mathrm{PCLR}_{(r)}(\hat{\mathcal{T}})$ and $0$ otherwise. Let $P$ denote the distribution of the vector $(W_i^{\prime}, X_i^{\prime}, u_i, V_i)^{\prime}$ and we shall index the relevant quantities by $P$ in the sequel. Define

equation[equation omitted — 225 chars of source]

$Z_i^{*}\equiv W_i-\operatorname{E}_P[W_iX_i^{\prime}](\operatorname{E}_P[X_iX_i^{\prime}])^{-1}X_i$, and

equation[equation omitted — 192 chars of source]

We maintain the following assumptions to develop the asymptotic results.

assumption\begin{enumerate}[label=(\alph*)] For some constants $\delta, \delta_1>0$ and $M_0<\infty$, • $\{(W_i^{\prime}, X_i^{\prime}, u_i, V_i)^{\prime}\}_{i=1}^n$ are i.i.d. with distribution $P$, • $\operatorname{E}_P[u_i(W_i^{\prime}, X_i^{\prime})]=0$, • $\operatorname{E}_P[\Vert(W_i^{\prime}, X_i^{\prime}, u_i)^{\prime}\Vert^{4+\delta}]<M_0$, • $\lambda_{\min}(A)\geq \delta_1$ for $A\in\big\{\operatorname{E}_P[(W_i^{\prime}, X_i^{\prime})^{\prime}(W_i^{\prime}, X_i^{\prime})], \operatorname{E}_P[{Z}_i^{*}{Z}_i^{*\prime}], \Sigma_P, \sigma_P^2\}$, • $\operatorname{E}_P[V_i(W_i^{\prime}, X_i^{\prime})]=0$, • $\operatorname{E}_P[\vert V_i\vert^{4+\delta}]<M_0$, • $\lambda_{\min}(\Sigma_P^V-\Sigma_P^{Vu}(\sigma_P^2)^{-1}\Sigma_P^{uV})\geq \delta_1$. \end{enumerate}

Next we define the parameter space for the permutation tests.

align[align omitted — 353 chars of source]

Assumption (ref) does not impose virtually any restriction on the degree of association between $Y$ and $W$ (after controlling for the effect of the exogenous covariates), thus allowing for an arbitrary identification strength on the part of the instruments. The distribution $P$ in Assumption (ref)(ref) is allowed to vary with the sample size i.e. $P=P_n$. For simplicity, we suppress the dependence on $n$. Assumption (ref)(ref) and (ref) impose finite $4+\delta$ moment on the error terms and the exogenous variables, and are slightly stronger than the cross moment restrictions used in Andrews-Guggenberger(2019). For pointwise asymptotic results or when $P$ does not vary with $n$, weaker restrictions would be sufficient. The independence assumption between the instruments and the error terms, which is typically made in the randomization test literature (see, for example, Imbens-Rosenbaum(2005)), is not maintained here; the instruments are only assumed to satisfy the standard exogeneity conditions in Assumption (ref)(ref) and (ref). If independence assumptions are maintained, as shown in Proposition (ref), the PAR tests are exact and do not require the finite $4+\delta$ moment assumptions. Assumption (ref)(ref) and (ref) require that the second moment matrices of the instruments and exogenous covariates, and covariance matrices are (uniformly) nonsingular, and are similar to the assumptions employed by Andrews-Guggenberger(2017), Andrews-Guggenberger(2019) and Andrews-Cheng-Guggenberger(2020) when developing identification-robust (but not singularity-robust) AR, LM and CLR tests. These conditions can be restrictive but can potentially be relaxed as in Andrews-Guggenberger(2019) (see also Dufour-Taamouti(2007)) using generalized inverses of the relevant matrices. However, doing so would entail a considerable complication in the proof, and so, for simplicity, we maintain Assumption (ref)(ref) and (ref). The condition $\lambda_{\min}(\Sigma_P^V-\Sigma_P^{Vu}(\sigma_P^2)^{-1}\Sigma_P^{uV})\geq \delta_1$ is needed for the asymptotic similarity of the PLM test. The main result of this paper is given in the following theorem which establishes the uniform asymptotic validity of the robust permutation tests.

theorem[Asymptotic similarity] Suppose that $H_0:\theta=\theta_0$ holds, and $N\to\infty$ and $n\to\infty$. Then, for $\alpha\in (0,1)$ and the permutation statistics \begin{equation*} \mathrm{PR}\in\{\mathrm{PAR}_1, \mathrm{PAR}_2, \mathrm{PLM}, \mathrm{PCLR}\} \end{equation*} with the corresponding parameter spaces $\mathcal{P}_0\in\{\mathcal{P}^{\mathrm{PAR}_1}_0, \mathcal{P}^{\mathrm{PAR}_2}_0, \mathcal{P}^{\mathrm{PLM}}_0, \mathcal{P}^{\mathrm{PCLR}}_0\}$, respectively, \begin{equation*} \operatornamewithlimits{\lim\sup \ }_{n\rightarrow \infty}\sup_{P\in\mathcal{P}_0} \operatorname{E}_{P}[\phi_n^{\mathrm{PR}}] = \operatornamewithlimits{\lim\inf \ }_{n\rightarrow \infty}\inf_{P\in\mathcal{P}_0} \operatorname{E}_{P}[\phi_n^{\mathrm{PR}}] =\alpha. \end{equation*}

Theorem (ref) shows that the permutation tests are asymptotically similar (hence asymptotically of correct size) i.e. the asymptotic null rejection probability is equal to the nominal significance over a large class of data generating processes. Thus, the advantage of the $\mathrm{PAR}_1$ and $\mathrm{PAR}_2$ tests relative to the existing tests is that they inherit the desirable properties of both the exact and asymptotic AR tests (including other simulation-based tests that are asymptotically valid); the PAR tests are exact under the same assumptions as the rank-based tests are, and asymptotically pivotal and heteroskedasticity-robust under the assumptions used to show the validity of the asymptotic tests. The main technical tool for the asymptotics of the permutation statistics are the combinatorial central limit theorem based on a Lindeberg type condition Motoo(1956) for the sums

equation[equation omitted — 139 chars of source]

The fact that the intercept term $\iota$ is included in $X$ in (ref)-(ref) implies that the two quantities in (ref) have mean zero with respect to the distribution of $\pi$, which, in turn, plays an important role for the asymptotic validity of the permutation tests (see also Remark 3.1 of Diciccio-Romano(2017)). The proof of asymptotic similarity of the permutation tests uses the generic method of Andrews-Cheng-Guggenberger(2020) for establishing the asymptotic size of tests. Next, we shall derive the asymptotic distribution of the permutation statistics under the local alternatives $\theta_n=\theta_0+h_\theta n^{-1/2}$ with $h_\theta \in\mathbb{R}$ fixed, assuming strong identification. Let $\chi_l^2(\eta^2)$ denote noncentral chi-square random variable with degrees of freedom $l$ and noncentrality parameter $\eta^2$.

proposition[Asymptotic local power under strong identification] Let Assumption (ref) hold, and $h_\theta \in\mathbb{R}$ be fixed. Assume that the model is strongly identified in the sense that $n^{1/2}\Vert\mathit{\Gamma}\Vert\to\infty$ as $n\to\infty$, and for $G_P\equiv\operatorname{E}_P[{Z}_i^{*}{Z}_i^{*\prime}]\mathit{\Gamma}$, $\eta^2\equiv \lim_{n\to\infty}G_P'\Sigma_P^{-1}G_Ph_\theta^2$ exists. Then, under $H_1:\theta_n=\theta_0+h_\theta n^{-1/2}$, $\mathrm{AR}\displaystyle \stackrel{d}{\longrightarrow} \chi^2_k(\eta^2)$, $\mathrm{LM}\displaystyle \stackrel{d}{\longrightarrow} \chi^2_1(\eta^2)$ and $\mathrm{CLR}\displaystyle \stackrel{d}{\longrightarrow} \chi^2_1(\eta^2)$ as $n\to\infty$, and \begin{align} &\sup_{x\in\mathbb{R}}\vert \tilde{F}_N^{\mathrm{PAR}_1}(x)-P[\chi^2_k\leq x]\vert \stackrel{p}{\longrightarrow} 0,\\ &\sup_{x\in\mathbb{R}}\vert \tilde{F}_N^{\mathrm{PAR}_2}(x)-P[\chi^2_k\leq x]\vert \stackrel{p}{\longrightarrow} 0,\\ &\sup_{x\in\mathbb{R}}\vert \tilde{F}_N^{\mathrm{PLM}}(x)-P[\chi^2_1\leq x]\vert \stackrel{p}{\longrightarrow} 0,\\ &\sup_{x\in\mathbb{R}}\vert \tilde{F}_N^{\mathrm{PCLR}}(x, \hat{\mathcal{T}})-P[\chi^2_1\leq x]\vert\stackrel{p}{\longrightarrow} 0,\\ \end{align} as $N\to\infty$ and $n\to\infty$.

It follows that the asymptotic local power of the PAR tests is equal to that of the asymptotic AR test given by $P[\chi_k^2(\eta^2)>r_{1-\alpha}(\chi^2_k)]$, where $r_{1-\alpha}(\chi^2_k)$ is the $1-\alpha$ quantile of $\chi^2_k$ random variable. Moreover, under strong identification, the PLM and PCLR tests, and the asymptotic LM and CLR tests have the same asymptotic local power equal to $P[\chi_1^2(\eta^2)>r_{1-\alpha}(\chi^2_1)]$.

PAR tests with a diverging number of IVs

This section considers an extension of the PAR tests to cases where the number of IVs, $k$, may grow with the sample size $n$. This extension comprises two main parts. First, the derivation of the asymptotic distribution of the PAR statistics requires a proof strategy different from the fixed-$k$ case. This is accomplished by interpreting a dominant term in the PAR statistic (after centering and scaling) as a double-indexed permutation statistic. Second, we derive the asymptotic distribution of the AR statistic under the condition that $k^3/n\to 0$ as $n\to\infty$, which is the same rate used by Andrews-Stock(2007) who establish the asymptotic validity of the AR, LM and CLR tests in homoskedastic linear IV models.

For simplicity, we assume that $X=\iota$, although the proof can be extended to incorporate a fixed number exogenous covariates in a straightforward manner. In the present case, the $\mathrm{PAR}_1$ and $\mathrm{PAR}_2$ tests are (distributionally) equivalent because $W_\pi' M_{\iota}u(\theta_0)=\sum_{i=1}^n(W_{\pi(i)}-\bar{W})u_i(\theta_0)=\sum_{i=1}^n(W_{i}-\bar{W})u_{\pi^{-1}(i)}(\theta_0)=Z'u_{\pi^{-1}}(\theta_0)$ and $W_\pi^{\prime}M_\iota\tilde{\Sigma}_u(\theta_0) M_\iota W_\pi=\sum_{i=1}^n(W_{\pi(i)}-\bar{W})(W_{\pi(i)}-\bar{W})'\tilde{u}_i(\theta_0)^2=\sum_{i=1}^n(W_{i}-\bar{W})(W_{i}-\bar{W})'\tilde{u}_{\pi^{-1}(i)}(\theta_0)^2$, where $\bar{W}\equiv n^{-1}W'\iota$ and $\pi^{-1}$ is the inverse of $\pi$ i.e. $\pi\circ\pi^{-1}$ is identity permutation, and is uniformly distributed over $\mathbb{G}_n$. When $Z$ and $u(\theta_0)$ are independent under $H_0:\theta=\theta_0$, the PAR tests are still exact since Proposition (ref) holds for any $k$ provided the statistics are well-defined.

Let $P_{ij}\equiv Z_i'(Z'Z)^{-1}Z_j$ denote the $(i,j)$ element of $P_Z$, $W_{ij}, j=1,\dots, k,$ be the $j$-th element of $W_i$, $\sum_{i\neq j}$ denote the summation over distinct values of the indices $i, j=1,\dots, n$, and $O_{a.s.}(\cdot)$ abbreviate $O(\cdot)$ almost surely. The asymptotic validity of the PAR tests are established under the following conditions in addition to Assumption (ref)(ref),(ref) and (ref).

assumptionAs $k,n\to\infty$, \begin{enumerate}[label=(\alph*)] • $\max_{1\leq i\leq n}\left(\sum_{j=1, j\neq i}^n\vert P_{ij}\vert \right)=O_{a.s.}\left(\frac{\sum_{i\neq j}P_{ij}^2}{n\max_{i,j, i\neq j} |P_{ij}|}\right)$, • $\frac{\sum_{i\neq j}P_{ij}^2}{n^2\max_{i,j, i\neq j} P_{ij}^2}\stackrel{a.s.}{\longrightarrow} 0$, • $k^{-1}\sum_{i=1}^nP_{ii}^2\stackrel{a.s.}{\longrightarrow} 0$, • $\sup_{j\leq k}(\operatorname{E}[W_{1j}^4u_1^4]+\operatorname{E}[W_{1j}^4]+\operatorname{E}[u_1^4])<M_0<\infty$. \end{enumerate}

Assumptions (ref) is key for deriving the asymptotic distribution of the PAR statistics under the diverging-$k$ asymptotics. Assumptions (ref)(ref) states that the row sums of the off-diagonal elements of $P_Z$ are roughly equal, Assumptions (ref)(ref) requires that the off-diagonal elements of $P_Z$ are small relative to its largest element. By standard arguments, $\sum_{i=1}^nP_{ii}^2 \leq \sum_{i=1}^nP_{ii}=k$, hence $k^{-1}\sum_{i=1}^nP_{ii}^2\leq 1$. Assumption (ref)(ref) makes a stronger requirement that the left-hand side of the latter inequality actually converges to $0$.

To gain insight into Assumption (ref)(ref)-(ref), let $W$ be a matrix of group dummy variables i.e. the $j$-th column of $W$ has entries $W_{ij}=1, i=n_{j-1}+1,\dots, n_{j-1}+n_j, j=1,\dots, k,$ with $n_0=0, n=\sum_{j=1}^kn_j$, and $0$ otherwise. Let $n_{\min}\equiv \min_{1\leq j\leq k}n_j$ and $n_{\max}\equiv \max_{1\leq j\leq k}n_j$. By direct calculations, one can see that Assumption (ref)(ref) is equivalent to $(1-n_{\max}^{-1})n n_{\min}^{-1}/(k-\sum_{j=1}^kn_j^{-1})=O_{a.s.}(1)$, which holds if $n_{\max}/n_{\min}$ is bounded. Assumption (ref)(ref) is equivalent to $(k-\sum_{j=1}^k n_j^{-1})/(n^2/n_{\min})$ being bounded. The latter holds because $(k-\sum_{j=1}^k n_j^{-1})/(n^2/n_{\min})\leq k^{-2}(k-\sum_{j=1}^k n_j^{-1})\leq k^{-2}(k-n^{-1}k^2)\leq 1$ using $(n/n_{\min})\leq k$ and $\sum_{j=1}^k n_j^{-1}\geq k^2/(\sum_{j=1}^kn_j)$ which follows from the convexity of $x\mapsto x^{-1}$. Finally, (ref)(ref) holds if $n_{\min}\stackrel{a.s.}{\longrightarrow} \infty$, since $k^{-1}\sum_{i=1}^nP_{ii}^2=k^{-1}\sum_{j=1}^kn_j^{-2}\leq n_{\min}^{-2}\stackrel{a.s.}{\longrightarrow} 0$.

Assumption (ref)(ref) replaces Assumption (ref)(ref). The boundedness of the cross moment $\operatorname{E}[W_{1j}^4u_1^4]$ is needed because we allow for a growing $k$ which, in turn, imposes a restriction on the consistency of the covariance matrix estimator. Furthermore, the marginal moment conditions in Assumption (ref)(ref) are slightly weaker than Assumption (ref)(ref) because we only show the tests have asymptotically correct level, as opposed to size (uniform validity). The main results of this section is as follows.

propositionLet Assumptions (ref)(ref),(ref), (ref), and (ref) hold with $P$ independent of $n$. Assume that $k, N\to\infty$ and $k^3/n\to 0$ as $n\to\infty$. If $H_0:\theta=\theta_0$ is true, then, for $\alpha\in (0,1)$ and the permutation statistics $\mathrm{PAR}\in\{\mathrm{PAR}_1, \mathrm{PAR}_2\}$, it holds that \begin{equation*} \lim_{n\to \infty}\operatorname{E}_{P}[\phi_n^{\mathrm{PAR}}]=\alpha. \end{equation*}

The asymptotic distribution of the PAR statistics is obtained as follows. Letting $\tilde{\sigma}^2\equiv n^{-1}\tilde{u}(\theta_0)'\tilde{u}(\theta_0)$, we can rewrite

align[align omitted — 400 chars of source]

It can be shown that the first summand in the above expression is negligible while the second summand converges to $k$. The key insight for establishing the asymptotic distribution of the third summand is to view it as a double-indexed permutation statistic of the form $\sum_{i\neq j}a_{ij}b_{\pi(i)\pi(j)}$ to which one can apply a CLT established by Pham-Mocks-Sroka(1989).

Recently, Mikusheva-Sun(2022) propose a heteroskedasticity-robust Jacknife AR (JAR) test that is based on $\sum_{i\neq j}P_{ij}\tilde{u}_{i}(\theta_0)\tilde{u}_{j}(\theta_0)$ and is asymptotically valid under the condition $k/n\to \tau\in(0,1)$ among others. Our proof strategy is likely to be useful for deriving the asymptotic distribution of a permutation version of their statistic. We leave this extension as well as those of the PLM and PCLR statistics under conditions less stringent than those in Assumption (ref) to future research.

Simulations

This section presents a simulation evidence on the performance of the proposed permutation tests. The data were generated according to

align*[align* omitted — 170 chars of source]

where $\theta=0$, $d=1$, $\mathit{\Gamma}=(1,\dots, 1)^{\prime}\sqrt{\lambda/(nk)}\,(k\times 1)$ and $\lambda\in\{0.1, 4, 20\}$, $X_{2i}$ is $(p-1)\times 1$ vector of non-constant included exogenous variables with $p\in\{1, 5\}$, and $V_i=\rho u_i+\sqrt{1-\rho^2}\epsilon_i$ with $\rho=0.5$. The parameters $\gamma_1, \gamma_2, \psi_1$, and $\mathit{\Psi}_2$ are set equal to $0$. As specified below, $W_i$ has zero mean and unit covariance matrix except for the first part of simulations in Section (ref), hence $\lambda=n\mathit{\Gamma}^{\prime}\operatorname{E}[W_iW_i^{\prime}]\mathit{\Gamma}\approx\mathit{\Gamma}^{\prime}W^{\prime}W\mathit{\Gamma}$. We consider the values $\lambda\in\{0.1, 4, 20\}$ which correspond to very weak, weak and strong identification. The tested restriction is $H_0:\theta=\theta_0=0$. We implement the heteroskedasticity-robust asymptotic tests denoted as AR, LM, CLR, their permutation versions PAR$_1$, PAR$_2$, PLM, PCLR, the normal score and Wilcoxon score rank-based AR tests of Andrews-Marmer(2008) denoted as RARn and RARw, respectively, the normal score and Wilcoxon score rank-based CLR statistics of Andrews-Soares(2007) denoted as RCLRn and RCLRw, respectively, and the wild bootstrap AR and LM tests of Davidson-MacKinnon(2012) denoted as WAR and WLM, respectively. For both the asymptotic and permutation CLR-type tests, we do not make the eigenvalue-adjustment (i.e. $\epsilon=0$ in footnote (ref)). In all simulations below, the null rejection probabilities of the tests (including CLR, PCLR, the rank-based and the wild bootstrap tests considered below) are computed using 2000 replications with 999 simulated samples for each replication while we display some power curves associated with the tests in Supplemental Appendix. We consider the following three cases in turn.

Heavy tails

To examine the finite sample validity, we consider heavy-tailed observations where each element of the random vector $(W_i^{\prime}, X_{2i}^{\prime}, u_i, \epsilon_i)^{\prime}$ is drawn independently from standard Cauchy distribution, and set $k\in\{5, 10\}$, $n\in\{50, 100\}$, and $\lambda=4$. Table (ref) presents the null rejection probabilities. The asymptotic tests all underreject. The PAR$_1$ test has nearly correct levels as predicted by Theorem (ref) in all cases, and the $\mathrm{PAR}_2$ test, which is exact only when $p=1$, has slightly more accurate rejection rates in the case $p=1$ than in the case $p=5$. The RARn, RARw, RCLRn and RCLRw tests are all robust against heavy-tailed errors when the IVs, the exogenous covariates and the error terms are independent as stated in Assumption (ref). The RARn and RARw tests are exact and this is borne out in the simulation results. The RCLRn and RCLRw tests slightly underreject which may be attributed to their asymptotic nature. Note, however, that, under the independence assumption and heavy-tailed errors, the rank-based tests have better power properties than the asymptotic tests, see Andrews-Soares(2007) and Andrews-Marmer(2008). In such cases, the permutation tests are likely to be dominated by the rank-based tests in terms of power given the local asymptotic equivalence between the permutation tests and the asymptotic tests under strong identification established in Proposition (ref).

Independence vs. standard exclusion restriction

The next set of simulations highlights the difference between the independence and the standard exclusion restriction assumption in homoskedastic setting. Two cases are considered:

enumerate$(W_i^{\prime}, X_{2i}^{\prime}, u_i, \epsilon_i)^{\prime}\sim t_{5}[0, I_{k+p+1}]$, where $t_{5}[0, I_{k+p+1}]$ stands for $(k+p+1)$-variate $t$-distribution with degrees of freedom $5$ and covariance matrix $I_{k+p+1}$. Under this setting, these random variables have finite fourth moments, and $(W_i^{\prime}, X_{2i}^{\prime})^{\prime}$ satisfies the standard exclusion restriction in Assumption (ref): $\operatorname{E}[(u_i, V_i^{\prime})^{\prime}(W_i^{\prime}, X_{i}^{\prime})]=0$ but $(u_i, \epsilon_i)^{\prime}$ and $(W_i^{\prime}, X_{i}^{\prime})^{\prime}$ are dependent;\footnote{Because the distribution is fixed i.e. does not vary with the sample size, the permutation tests remain asymptotically valid under the finite fourth moment assumption.} • $(W_i^{\prime}, X_{2i}^{\prime}, u_i, \epsilon_i)^{\prime}\sim N[0, I_{k+p+1}]$. Clearly, in this case the instruments and the error terms are independent.

The result for the overidentified models with $n=100$ and $k=5$ is reported in Table (ref). The permutation tests perform better than the asymptotic tests regardless of whether the instruments and the error terms are independent or not. The rank-based tests, displayed in the bottom part of Table (ref), have nearly correct level in the independent case but overreject in the dependent case. The result for the just-identified design is $k=1$ are similar to the over-identified case, thus is relegated to Supplemental Appendix.

Conditional heteroskedasticity

The next part of the simulations considers designs with conditional heteroskedasticity. The data are generated according to ((ref)) and ((ref)) with

align*[align* omitted — 94 chars of source]

where $W_{i1}$ is the first element of $W_i$, and similarly to the previous case

enumerate$(W_i^{\prime}, X_{2i}^{\prime}, \upsilon_i, \epsilon_i)^{\prime}\sim t_{5}[0, I_{k+p+1}]$; • $(W_i^{\prime}, X_{2i}^{\prime}, \upsilon_i, \epsilon_i)^{\prime}\sim N[0, I_{k+p+1}]$.

The sample size and the number of instruments are $n=100$ and $k\in\{2, 5, 10\}$. The results for $\lambda=4$ are displayed in Tables (ref). The asymptotic AR test tends to underreject and performs poorly compared to the LM statistic. The permutation tests perform better than their asymptotic counterparts in most cases, and on par with the wild bootstrap in the independent case. However, when the instruments satisfy the standard exogeneity condition, and there are included exogenous variables, the permutation tests appear to have an edge over the wild bootstrap AR test which overrejects. Qualitatively, the same observations made in the homoskedastic case for the rank-based tests are also observed in the heteroskedastic case as the rank-based tests reject by a substantial margin in the dependent case. See Supplemental Appendix for additional simulation evidence for the cases $\lambda\in\{0.1, 20\}$.

table[table omitted — 2,060 chars of source]
table[table omitted — 3,509 chars of source]
table[table omitted — 3,249 chars of source]

Empirical application

Bazzi-Clemens(2013) re-examine cross-sectional IV regression results about the effect of foreign aid on economic growth from several published studies. The variables in the “aid-growth" IV regressions estimated by Rajan-Subramanian(2008) and Bazzi-Clemens(2013) are

itemize$y_i:$ the average annual growth of per capita GDP; • $Y_i:$ the foreign-aid receipts to GDP ratio (Aid/GDP); • $W_{i1}:$ a variable constructed from aid-recipient population size, aid-donor population size, colonial relationship, and language traits (see Appendix A of Bazzi-Clemens(2013)); • $X_i:$ constant, initial per capita GDP, initial level of policy, initial level of life expectancy, geography, institutional quality, initial inflation, initial M2/GDP, initial budget balance/GDP, Revolutions, Ethnic fractionalization, Sub-Saharan Africa dummy and East Asia dummy.

The data cover the period 1970-2000 and the sample size is $n=78$. Bazzi-Clemens(2013) consider the following four specifications:

itemize• Specification 1: The baseline specification of Rajan-Subramanian(2008) with the variables as above; • Specification 2: log population is included in the second stage of Specification 1; • Specification 3: $W_{i2}=$ log population replaces the instrument $W_{i1}$ in Specification 1; • Specification 4: An instrument, $W_{i3}$, constructed from the colonial ties indicators only replaces $W_{i1}$ in Specification 1.

Specifications 1-4 are just-identified as there is a single instrument for the single endogenous regressor, hence we only consider the AR and PAR test results. In addition, we consider Specifications 5-7 with IVs $(W_{i1}, W_{i2})'$, $(W_{i2}, W_{i3})'$ and $(W_{i1}, W_{i2}, W_{i3})'$, respectively, to examine whether the over-identified specifications could have any instrumentation power. $W_{i2}$ is included in the latter three cases because it is the strongest instrument as reported in Bazzi-Clemens(2013). We also include the confidence interval based on the two-stage least squares $t$-test, denoted as $t_{\mathrm{2sls}}$. Tables (ref) and (ref) report the $95\%$ confidence intervals for the coefficient on Aid/GDP. In all specifications, the Breusch-Pagan test $p$-values from the two separate reduced-form regressions of the endogenous variables on the exogenous variables show an evidence of heteroskedasticity.

The results for Specifications 1-4 in Table (ref) qualitatively agree with the homoskedastic CLR confidence intervals reported in Bazzi-Clemens(2013). The PAR$_2$ confidence intervals are nearly identical but slightly shorter than the asymptotic and wild bootstrap confidence intervals in Specifications 1 and 3. The very wide confidence intervals in Specifications 2 and 4 that do not use log population or the instruments based on it indicate that the population size instrument used in the other specifications has indeed the most identifying power. The above results also extend to Table (ref). Some noteworthy findings are as follows. The WAR confidence intervals are wider than the AR and PAR confidence intervals. The PLM confidence intervals are shorter than the LM and WLM confidence intervals, but still wider than the other confidence intervals. Somewhat surprisingly, in Specification 6, the AR, PAR and $\mathrm{PCLR}$ confidence intervals exclude $0$ thereby pointing towards borderline significant effect of the foreign aid on growth despite the rather small value of the first-stage $F$-statistic 17.90. Also, in Specification 7, the $\mathrm{PAR}$ confidence intervals produce significant results which underscore the utility of the proposed tests. However, these results should be interpreted with caution because the population size might affect the growth through multiple channels as argued by Bazzi-Clemens(2013). Overall, the uncertainty regarding the effect of foreign aid in the sample data may be expressed more accurately by the permutation and the wild bootstrap confidence intervals as the unobserved confounders in the “aid-growth" regression are more likely to be orthogonal to, rather than independent of the population size.

table[table omitted — 2,851 chars of source]
table[table omitted — 3,276 chars of source]

Conclusion

The PAR$_1$ test shares the strengths of the rank/permutation tests and the wild bootstrap tests because it is exact under independence and heteroskedasticity-robust under standard assumptions. This test is recommended when the IVs are randomly assigned, hence independent of the exogenous covariates and the error term. The PAR$_2$ test also shows a satisfactory performance in the simulations. Because the PAR tests do not require the correct specification of the first-stage equation as they directly uses the IVs in the equation $\hat{m}(\theta_0)=n^{-1}Z'\tilde{u}(\theta_0)$ Dufour-Taamouti(2007), they have robustness advantages relative to the other permutation tests. Moreover, the PCLR test is found to perform reasonably well in the simulations. Assuming that (ref)-(ref) are correctly specified, as we have shown, the PCLR tests are similar and equivalent to the CLR tests of Andrews-Guggenberger(2019) asymptotically. In homoskedastic IV models, the latter tests are asymptotically equivalent or reduce to the CLR test of Moreira(2003) which is known to be optimal when there is single endogenous regressor Andrews-Moreira-Stock(2006). However, despite being equivalent to the PCLR test under strong identification, the PLM test shows an inferior performance in terms of power. Based on the properties mentioned and the simulation evidence, we recommend reporting both the PAR and PCLR test results in practice. Looking ahead, we plan to address the important issue of cluster and identification-robust inference in IV models using the current approach in another work.

\includepdf[pages=-]{SA.pdf}