Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
56,603 characters · 14 sections · 79 citation commands
Pivotal and identification-robust nonparametric inference in linear IV models
\def\spacingset#1{ {#1}} \spacingset{1}
\if11 \fi
\if01 {
} \fi
{\it Keywords:} Weak Identification, Specification Testing, Subvector Inference
\spacingset{1.8}
We consider cross-section data observations and the linear model popular from micro-econometrics
where $Y_{2}$ are endogenous variables, $X_{1}$ are exogenous control variables, and $X_{2}$ are exogenous instrumental variables. Over the last 30 years, it has become clear that standard asymptotic approximations may reflect poorly what is observed even for large samples when there is weak correlation between instrumental variables and endogenous explanatory variables. Alternative asymptotic frameworks have been developed to account for potentially weak identification, along with tests that deliver reliable inference about parameters of interest, see e.g., SS, SW, Moreira, kleibergen_pivotal_2002,Kleib, Andrews2012, andrews_identification_2019, Andrews2016, and andrews_conditional_2016,andrews_geometric_2016. Surveys on weak identification issues include SWY, Dufour, HH, AS, and Andrews2019. These inference procedures are robust to identification strength and uniformly control size. They have good power properties when the first stage is linear and well specified, in particular, the conditional likelihood ratio test of Moreira is nearly optimal, see AMS2006 and andrews_optimal_2019. However, reliance on a parametric first stage may be misleading. From an empirical perspective, Dieterle2016 have documented significant nonlinearities in first-stage regression in several applied microeconomics papers. Since practitioners typically have little prior information on the form of the relation between endogenous variables and instruments, one may consider estimating the reduced form nonparametrically, e.g., using an increasing number of approximating series. However, nonparametrically estimated instruments cannot be relied upon under weak identification, see JP11 and mikusheva_inference. Indeed, if identification is not strong enough, the statistical variability of a nonparametric estimator will dominate the signal we aim to estimate.
In a recent work, building on the Integrated Conditional Moment (ICM) principle proposed by Bierens82, antoine_identification-robust_2023 develop two inference procedures that are robust to any identification pattern and unknown heteroskedasticity, and do not rely on a linear projection in the first-stage equation. In particular, they study an ICM test that tests at the same time for the value of the parameter and the specification of the model. However, critical values should be simulated for each value of the parameter.
Our present work builds on and improves upon the ICM test. We propose a new heteroskedastic version of their ICM statistic, labeled HICM. Its key feature is that it is asymptotically pivotal, i.e., under the null hypothesis $H_0: \beta=\beta_{0}$, its asymptotic distribution is independent of $\beta_0$. Since building a confidence interval for $\beta$ involves inverting a test statistic for $H_0$ a large number of times, the key advantage of our test is computational, as using critical values that are independent of parameter values significantly speeds up the process. In our implementation, we found that computation time can be reduced by a factor close to 70. In practice, this means that when ICM-based inference can take weeks in some instances, it would reduce to a few hours with our new procedure.
We also develop two additional inference procedures that build on this particular feature and cannot be entertained using the ICM test of antoine_identification-robust_2023. First, we investigate subvector inference and obtain valid identification-robust inference on some parameters of interest, regardless of the underlying identification of the other parameters not under test. Second, we devise a {\em pure} specification test that is powerful independently of identification strength and of the particular functional form of the reduced form that links instruments and endogenous variables. In both cases, inference is conservative but easily implemented. In addition, the different tests uniformly control size irrespective of identification strength.
We contribute to the recent literature concerned with joint issues of specification and weak identification. With a focus on inference under potential misspecification, KleibergenZhan propose a doubly-robust Lagrange multiplier statistic to deliver identification-robust and valid inference on the pseudo-true value in a moment-based model. With a focus on specification testing, AntoineFrazierRenault2026WP develop procedures in GMM and minimum distance settings that are robust to weak identification. Dovonon2026 propose a modified J-test with an increasing number of instruments that is asymptotically pivotal and agnostic on identification, but requires to choose the number and identity of nonlinear transformations of the instrumental variables. The procedures proposed here complement these approaches, while allowing the first-stage equation to be nonparametric.
We illustrate the finite sample properties of our procedures in a series of simulations. When the model is correctly specified, we find that the level of our HICM-based inference is well controlled, that it has significant power advantages compared to existing procedures such as the one of SW when the reduced form equation is nonlinear, and that it is also competitive for a linear reduced form. Similarly, when testing whether the model is correctly specified, our HICM-based specification test demonstrates excellent size and power properties. We also demonstrate the relevance and power of our procedure in an empirical application that studies the effects of population decline in Mexico on land concentration in the sixteenth century, using the data and framework of Mexico.
Our paper is organized as follows. In Section (ref), we introduce our framework and our new HICM test statistic. We discuss how it can be used for powerful and identification-robust inference that is compatible with subvector inference, as well as for specification testing. The asymptotic properties of our inference procedures and specification test based on HICM are studied in Section (ref). Their finite sample properties are investigated in a series of Monte Carlo experiments and in an empirical application in Sections (ref) and (ref). We conclude in Section (ref).
The influence of exogenous control variables $X_{1}$ can be projected out through orthogonal projection in ((ref)), which does not influence our reasoning but simplifies exposition. Hence, in what follows, we consider a structural equation of the form
augmented by a first-stage reduced form equation for $Y_2$,
The variables $Z$ include the instruments $X_{2}$, but may also include the exogenous $X_{1}$ to account for potential nonlinearities in $X_{1}$ in the function $\Pi(\cdot)$. Combining (ref) and (ref) yields \[ y_i - Y_{2i}'\beta_{0} = \Pi'(Z_i) \left(\beta-\beta_0\right) + \varepsilon_i, \qquad \text{where} \quad \varepsilon_i = u_i + V_{2i}' \left(\beta - \beta_{0}\right) \quad \text{and} \quad \mathbb{E}\,\left(\varepsilon_i|Z_i\right)=0 \, . \] We propose to test the null hypothesis
which considers at the same time the null hypothesis $\beta = \beta_0$ and the correct specification of the model. In what follows, we introduce our proposed HICM test, then we discuss subvector inference and specification testing.
antoine_identification-robust_2023 introduced the ICM test statistic. Building on Bierens82, ${H}_{0}$ is shown to be equivalent to
Bierens' Integrated Conditional Moment (ICM) statistic is
where $\mu$ is some symmetric measure with support $\mathbb{R}^{k}$. The ICM statistic of antoine_identification-robust_2023 is a standardized version of the latter that can be written as the following quadratic form
$W$ is the matrix of generic element $n^{-1}w\left(Z_{j}-Z_{m}\right)$, with $w(z) = \int_{\mathbb{R}^{k}} \exp(i s'z) \, d\mu(s)$, and $\widehat{\omega}$ is a (semiparametric) estimator of $\omega = \mathbb{E}\, (\operatorname{Var}(Y|Z))$. Hence, the ICM statistic (ref) resembles the AR statistic, with $W$ replacing $P_{Z}$, the orthogonal projection on $Z$. As apparent from its construction, it is designed to test {the correct specification of the model} together with the parameter value, as does the AR test under the assumption of a linear reduced form. While under homoskedasticity, the statistic has an asymptotic distribution independent of $\beta_0$, under heteroskedasticity, critical values depend on the considered $\beta_0$. A confidence set is generally obtained by inverting the $\operatorname{ICM}$ test as $\left\{ \beta_{0} : \operatorname{ICM}(\beta_{0}) < c_{1-\alpha}(Z, \beta_0) \right\}$. In particular, this means that, under heteroskedasticity, critical values $c_{1-\alpha}(Z, \beta_0)$ should be simulated for each candidate $\beta_0$, see antoine_identification-robust_2023 for details. Unless $\beta_0$ is unidimensional, the associated computational cost can be prohibitive.
We now construct a version of $\operatorname{ICM}$ that directly accounts for unknown heteroskedasticity in its construction. Let $\Omega(Z) = \operatorname{Var}(Y|Z)$, the conditional variance of $Y$ given $Z$. We rewrite our null hypothesis of interest in the equivalent form
From Bierens82's results, this is also equivalent to
Following a rationale similar to the one above, we can build an ICM statistic for this hypothesis. Given a consistent estimator $\widehat{\Omega}(\cdot)$ of $\Omega(\cdot)$, we thus define
where $\widehat{\Omega} = \operatorname{diag}\left( \widehat{\Omega}(Z_i), i = 1, \ldots n\right) $, and $e$ is the $n$-vector of ones. If the model is correctly specified with parameter $\beta_0$, twe show that $\operatorname{HICM}(\beta_0)$ asymptotically follows the same distribution as $G' W G$, where $G\sim N (\mbox{\bf 0}, \mbox{\bf I})$. This distribution is independent of the particular value of $\beta_0$. We can then easily simulate the distribution of our statistic under $H_{0}$ and recover a critical value as the $(1-\alpha)$ quantile of the distribution of $G' W G$, denoted as $c_{1-\alpha}(Z)$. A confidence set is obtained by inverting the $\operatorname{HICM}$ test, that is $\left\{ \beta_{0} : \operatorname{HICM}(\beta_{0}) < c_{1-\alpha}(Z) \right\}$. As we will see, the use of critical values that are independent of $\beta_0$ is computationally very advantageous.
Practically, we need to detail a number of elements of our statistic. First, one needs to estimate the conditional variance $\Omega(\cdot)$. One should note that weak identification does not preclude consistent estimation of this quantity. The conditional variance can be estimated parametrically if one is ready to make an assumption on its functional form. Otherwise, we can resort to nonparametric conditional variance estimation. Several consistent estimators have been developed for a univariate $Y$, and generalize easily. To make things concrete, we focus on kernel smoothing, which is used in our simulations and application. Let \[ \overline{Y}(z) = \left(nb_n\right)^{-1} \sum_{i=1}^n Y_i K \left( ( Z_i - z)/ b_n \right) \] based on the $n$ iid observations $(Y_{i},Z_{i})$, a kernel $K(\cdot)$, and a bandwidth $b_n$. With $e = (1, \ldots , 1)'$, let $\widehat{f} (z) = \overline{e} (z)$ and $\widehat{Y} (z) = \overline{Y}(z)/\widehat{f} (z)$. The conditional variance estimator is defined as
This estimator, studied by NCM, is a generalization of the kernel conditional variance and is positive definite whenever $K(\cdot)$ is positive. We can even consider \[ \widehat{\Omega} (z) = \left(nb_n\right)^{-1} \frac{\sum_{i=1}^n Y_{i} Y'_{i} K \left( ( Z_i - z)/ b_n \right)}{\widehat{f}(z)} \, , \] which estimates the uncentered moments $\mathbb{E}\, \left( Y'Y\right)$, and thus avoids preliminary estimation of $\mathbb{E}\, (Y|Z)$. Indeed, we only need $\widehat{\Omega} (z)$ to estimate $\operatorname{Var}(y - Y_2'\beta_0|Z)$, which equals $\mathbb{E}\,[(y - Y_2'\beta_0)^2|Z]$ under $H_0$, so estimation of uncentered cross moments is sufficient for this purpose.
Another element of our statistic is the function $w(\cdot)$, or equivalently the measure $\mu$. The role of $w(\cdot)$ resembles the one of the kernel in nonparametric estimation, but, in contrast, it is a {\em fixed user-chosen function that does not vary with the sample size}. To make this explicit, we will impose that the squared integral of $w(\cdot)$ equals one.\footnote{A more involved restriction would be to impose a similar condition on the Frobenius norm of $W$.} The condition for $\mu$ having support $\mathbb{R}^{k}$ translates into the restriction that $w(\cdot)$ should have a strictly positive Fourier transform almost everywhere. Examples include products of triangular, normal, logistic, see JKB95, Student, including Cauchy, see DK2002, or Laplace densities. To achieve scale invariance, we recommend, as in Bierens82, to scale the exogenous instruments by a measure of dispersion, such as their empirical standard deviation.
If $Z$ has bounded support, results from Bierens82 yield that ${H}_{0}$ holds if and only if $ \mathbb{E}\, \left[ (b_0' \Omega(Z) b_0)^{-1/2} \left( y - Y_{2}'\beta_0 \right) \exp(s'Z) \right] = 0 $ for all $s$ in an arbitrary neighborhood of $0$ in $\mathbb{R}^q$. Hence $\mu$ can be taken as any symmetric probability measure that contains $0$ in the interior of its support. For instance, we can consider the product of uniform distributions on $[-\pi,\pi]$, so that $w(\cdot)$ is the product of sinc functions. As noted by Bierens82, there is no loss of generality assuming a bounded support, as his above-mentioned equivalence result equally applies to a one-to-one transformation of $Z$, which can be chosen with a bounded image. Hence, we will adopt this assumption in our theoretical analysis.
On a more general level, Bierens82's principle replaces conditional moment restrictions by a continuum of unconditional moments involving the complex exponential function. Other functions have been used beyond this choice, see Bierens90 and BP97. stinchcombe_white_1998 characterize a large class of functions that could generate an equivalent set of unconditional moments. As detailed by LP2013, this yields a full collection of potential estimators under strong (or semi-strong) identification, such as the ones developed by dominguez_consistent_2004, AntoineLavergne, and escanciano_simple_2018 among others. This would also yield a collection of test statistics that could be used under weak identification, see Chen2025 for a recent instance. Here, we focus on a particular application of the ICM, which is suitable for theoretical investigation and practical implementation, and we leave for future work the investigation of the relative merits of these different ICM-type tests.
Let us partition our model (ref) as
Our statistic writes \[ \operatorname{HICM}(\beta_{0,1},\beta_{0,2}) = b_{0}' Y' \left[ (e \otimes b_0 )' \widehat{\Omega} (e \otimes b_0 ) \right]^{-1/2} W \left[ (e \otimes b_0 )' \widehat{\Omega} (e \otimes b_0 ) \right]^{-1/2} Y b_{0} \, , \] with $b_0=(1, -\beta_{0,1}',-\beta_{0,2}')$. Suppose we aim at identification-robust inference on $\beta_{0,1}$ only, so $\beta_{0,2}$ can be seen as a nuisance parameter that cannot be easily partialed out or reliably estimated.
We resort to a {\em conservative projection approach}, see Dufour1997. Our test statistic is \[ \operatorname{HICM}^*(\beta_{0,1}) \equiv \min_{\beta_{2}} \operatorname{HICM}(\beta_{0,1},\beta_2) \, . \] Since $\operatorname{HICM}^*(\beta_{0,1}) \leq \operatorname{HICM}(\beta_{0,1},\beta_{0,2})$, we can rely on the critical value $c_{1-\alpha}(Z)$ from the null asymptotic distribution of $\operatorname{HICM}$, which can be easily simulated as explained above. The approach is conservative but allows to control the level of the test. Inverting the test yields the confidence set $\left\{ \beta_{0,1} : \operatorname{HICM}^*(\beta_{0,1}) < c_{1-\alpha}(Z) \right\}$ for $\beta_1$. Valid subvector inference on $\beta_1$ thus obtains without any assumptions on identification of $\beta_2$.
In line with the previous subvector inference procedure, our specification test is based on \[ \operatorname{HICM}^* = \min_{\beta} \operatorname{HICM}(\beta) \, . \] Under correct specification, there is a $\beta_0$ such that ((ref)) holds. Since $\operatorname{HICM}^* \leq \operatorname{HICM}(\beta_0)$, we can rely on the simulated null distribution of $\operatorname{HICM}(\beta_0)$, which is independent of $\beta_0$, to bound the distribution of $\operatorname{HICM}^*$. Our asymptotic test rejects the correct specification of the model whenever $\operatorname{HICM}^* > c_{1-\alpha}(Z)$. From the above inequality, we have control of asymptotic size, but the test is conservative. If the model is incorrectly specified, then there is no $\beta_0$ such that \[ \mathbb{E}\, \left[ (b_0' \Omega(Z) b_0)^{-1/2} \left( y - Y_{2}'\beta_0 \right) \exp( i s'Z) \right] = 0 \quad \forall s\in\mathbb{R}^k \, . \] When the above quantity is uniformly bounded away from zero, we expect $\operatorname{HICM}^*$ to diverge and our test to be consistent. In the next section, we study the behavior of our specification test under some local alternatives and show that it has non-trivial power.
We consider the following assumptions.
For the sake of simplicity, we do not use a double index for observations, and we denote by $\left\{Y_{1}, \ldots , Y_{n}\right\}$ the independent copies from $Y$ for a sample size $n$. The assumption of a constant distribution for $Z$ could be weakened but is made to formalize that identification strength is related to the conditional distribution of $Y$ given $Z$ only. The boundedness of $Z$ could be relaxed, at the cost of more technicalities, but, as previously noted, a one-to-one transformation of $Z$ can be entertained to ensure it holds.
We denote by ${\cal P}$ the class of distributions on which our observations lie. Let ${\cal E}$ be a class of vector-valued functions $\Pi(\cdot)$ and let $N \left( \varepsilon,{\cal E}, L^{2} (Q)\right)$ be the covering number of ${\cal E}$, that is, the minimum number of $L^{2} (Q)$ $\varepsilon$-balls needed to cover ${\cal E}$, where a $L_{2} (Q)$ $\varepsilon$-ball around $\Pi(\cdot)$ is the set of vector functions $\left\{ h \in L^{2} (Q) \ : \ \int_{}^{}{ \|h-\Pi\|^{2} \, dQ } < \varepsilon \right\}$.
AndrewsHB and van_der_vaart_bracketing_1994, among others, exhibit classes of smooth functions that fulfill the above conditions.
Let ${\cal O}$ be a class of matrix-valued functions, and let $N \left( \varepsilon,{\cal O}, L_{2} (Q)\right)$ be the covering number of ${\cal O}$, defined similarly as above.
This assumption entails, in particular, that conditional variance estimation does not affect the asymptotic behavior of our statistics. There is a tension between the generality of the class of functions ${\cal O}$ and the class of possible distributions ${\cal P}$. When $\Omega (\cdot)$ is of parametric form, Assumption (ref) will be satisfied for a large class of distributions. When $\Omega(\cdot)$ is considered nonparametric and estimated accordingly, one typically assumes that its components are smooth functions, and to prove (iii), one has to show that $\widehat{\Omega}(\cdot)$ also satisfies the same smoothness conditions with probability converging to 1. Such results have been derived, see e.g., andrews_nonparametric_1995 for kernel estimators or cattaneo_optimal_2013 for partitioning estimators. Uniform convergence of nonparametric regression estimators (and their derivatives) generally requires the domain of the functions to be bounded and the absolutely continuous components of the distributions of the conditioning variables to have densities bounded away from zero on their support. When they are not, andrews_nonparametric_1995 discusses the use of a vanishing trimming that is compatible with the stochastic equicontinuity results of AndrewsHB. Condition (iv) is dealt with in the literature on honest confidence intervals using $L^{2}$ norm, see robins_adaptive_2006 and the references therein.
We denote by $c_{1-\alpha} (Z)$ the critical value of $\operatorname{HICM}$ obtained by the simulation-based method detailed above.\footnote{We neglect the approximation error due to a finite number of simulations by assuming the number of simulations is infinite so that the critical values are exact.} Let ${\cal P}_{\beta_{0}}$ be the subset of distributions in ${\cal P}$ such that $\beta=\beta_{0}$ in ((ref)). The following result establishes that our tests control size uniformly over a large class of probability distributions.
Our theorem readily implies that our test is asymptotically valid whatever the identification strength. Indeed, for any sequence $\Pi_{n}(\cdot), n \geq 1$, of functions in ${\cal E}$, that can decrease in norm to zero arbitrarily fast, our result yields asymptotic validity under this sequence, see vaart_weak_2000. The result also readily implies the uniform asymptotic validity of our subvector inference and specification tests.
We adopt here a large local alternatives setup similar to BP97.
Condition (i) allows studying the power of our tests against weak and semi-strong identification when considering a test of $H_{0}: \ \beta = \beta_{1}$ where $\beta_{1} \neq \beta_{0}$, the true parameter value. Condition (ii) is the strong identification case, and we consider local alternatives $H_{1n}: \ \beta_{1n} = \beta_{0} + \tilde{c}_{n} \frac{\delta}{\sqrt{n}}$, where $\delta \neq 0$ is fixed. In both cases, the object of interest is the asymptotic power of our two tests when $\tilde{c}_{n}$ becomes large.
Result (i) shows that, under weak identification, power is non-trivial for a large enough $\tilde{c}_{n}$. Result (ii) implies that, under strong identification, power is non-trivial under a sequence of Pitman local alternatives for $\tilde{c}_{n}$ large enough. A similar statement can be derived for our subvector inference procedure.
We consider a sequence of local alternatives \[ H_{1,n}: \ \min_{b = (1, - \beta')' : \beta \in {\cal B}} \int_{}^{}{ \left| \mathbb{E}\, \left[ (b' \Omega(Z) b)^{-1/2} (y - Y'_{2}\beta) \exp(is'Z) \right] \right|^2 \, d\mu(s) } = \tilde{d}_n^2 / {n} \, , \] where $ \tilde{d}_n, n = 1, \ldots$ is a positive real sequence uniformly bounded away from zero. We remain agnostic on the source of misspecification and on identification strength. We denote by ${\cal P}$ the class of distributions on which our observations lie under Assumptions (ref)-(i), (ref) and (ref).
To understand what our result implies, assume that \[ \mathbb{E}\,( y_i | Z_i) = \Pi(Z_i)'\beta_0 + \delta_n(Z_i), \quad i = 1, \ldots n \, . \] This encompasses situations where $Z$ has a direct effect on $y$, and is thus endogenous, or situations where the relationship between $y$ and $Y_2$ is nonlinear, and thus not appropriately modeled. If $\delta_n(Z) = \tilde{d}_n \delta(Z)/\sqrt{n}$ for a non-zero function $\delta(Z)$, then our specification test has some non-trivial power for $\tilde{d}_n$ large enough, provided $\delta(\cdot)$ is not linearly related to the components of $\Pi(\cdot)$, which is implicit in the formulation of $H_{1,n}$. This holds whether $\Pi(\cdot)$ is fixed as in Assumption (ref)-(ii), or whether $\Pi(Z) = \frac{C(Z)}{\sqrt{n}}$ as in Assumption (ref)-(i). In that sense, our specification test is really robust to identification strength.
We generate data following the model
where $c$ is a constant that controls the strength of the identification and $Y_{2i}$ is univariate. The joint distribution of $(u_{i},v_{2i})$ is a bivariate normal with mean $\mathbf{0}$, unit unconditional variances, and unconditional correlation $\rho$. We set $\alpha_0=\beta_0=\gamma_0=0$ and $\rho=0.8$. We consider three different specifications for the function $f(\cdot)$: (i) a polynomial function of degree 3 proportional to $z- 2z^3/5$, (ii) a linear function, and (iii) a function compatible with first-stage group heterogeneity, see Abadie2024, proportional to $\left( 2 z_2 - 1\right) \left( z_1-2z_1^3/5 \right)$. Here $Z$ (or $Z_1$) is deterministic with values evenly spread between -2 and 2 and $Z_{2}$ follows a Bernoulli with probability 1/2. Also, $f(Z)$ is centered and scaled to have variance one to ensure that the different cases are comparable. We consider heteroskedasticity depending on the first component of $Z$ of the form \[ \sigma(z)= \sqrt{\frac{3(1+z^2)}{7}} \, . \] Finally, $\delta$ controls the degree of misspecification. When $\delta = 0$, the model that excludes $Z$ from the structural equation is well-specified and similar to the one considered by antoine_identification-robust_2023, while it is misspecified when $\delta \neq 0$. In all our experiments, $w(\cdot) = \operatorname{sinc}(\pi \cdot)$, which corresponds to a uniform density $\mu$, while conditional covariances are estimated through kernel smoothing with a Gaussian kernel and rule-of-thumb bandwidth. We consider 5,000 replications.
For this subsection, we set $\delta=0$ to ensure that the model is always correctly specified, and we focus on delivering inference on $\beta_0$. We compare the performance of the HICM test introduced in this paper and the ICM one. Simulations results reported in antoine_identification-robust_2023 show that among competing procedures, such as the heteroskedasticity-robust conditional likelihood ratio test, the S test proposed by SW is the only one that controls size well when increasing the number of instruments, so we focus on this test in our comparison.
We consider three versions of S, with 1, 3, or 7 instruments in the polynomial and linear models. These instruments were obtained by fitting piecewise linear functions on intervals defined by the quartiles of $Z$, e.g., the three considered instruments are $1(z\leq 0)$, $z \times 1(z\leq0)$, and $z$. For the group heterogeneity model, we implement S based on a reduced form with 3 instruments, namely the continuous variable $Z_1$, the discrete one $Z_2$, and an interaction term. We then increase the number of instruments to 7 and 15. We construct these instruments as piecewise linear and interaction terms on intervals defined by the quartiles of $Z_1$, e.g., the seven considered instruments are $1(z_1\leq0)$, $z_1 \times 1(z_1 \leq 0)$, $z_1$, $z_2 \times 1(z_1 \leq 0)$, $z_2$, $z_2\times z_1 \times 1(z_1\leq 0)$, $z_2\times z_1$. The critical value for the S statistic is the quantile from the chi-square distribution with degrees of freedom equal to the number of instruments. For ICM, we compute p-values based on 499 simulated values of the statistic for each considered hypothesis $H_0:\beta=\bar\beta$. For our new test based on HICM, we compute 1999 simulated values of the statistic, which are independent of the value of $\beta$ under $H_0$.
Our benchmark is the heteroskedastic version of the polynomial model with a degree of weakness $c = 3$ and a sample size $n = 201$. We then increase $c$ to 7, then we consider a linear model with $c=3$, and a polynomial group model with $c=3$ and $n=401$. The empirical sizes are reported in Table (ref), for one variable $Z$ with $n=201$, or two variables $Z_1$ and $Z_2$ with $n=401$ (the empirical levels under the null hypothesis $H_0: \beta = 0$ do not depend upon other simulation details). Size is controlled by the three procedures, but HICM tends to be more conservative than S, and ICM is the most conservative.
In Figure (ref), we represent the power curves of the different tests when testing the null hypothesis that $H_0:\beta=\bar\beta$ with $\bar\beta\in[-1.5,1.5] \times \sqrt{n} /c$. In the benchmark polynomial model, S with 1 IV does not have more than trivial power. The power of S with 3 instruments is comparable to the one of HICM. With 7 instruments, it decreases and is comparable to the power of ICM. Increasing identification strength does not improve the power of S with only one IV. With a linear specification, S with one IV is the most powerful procedure as expected. Interestingly, HICM and ICM have power comparable to S with 3 IV, and dominate S with 7 IV. In a model with group heterogeneity, S with 3 IV does not have any power, while S with 7 IV is slightly dominated by HICM. Increasing to 15 IV lowers the power of S, which becomes comparable to that of ICM. Overall, HICM performs extremely well and is at the same time less undersized and more powerful than ICM. By contrast, S lacks power when too few instruments are used, while its power worsens if too many instruments are used. As it is never clear how to select the (number of) instruments, our HICM inference procedure appears particularly attractive.
We investigate data-generating processes similar to the ones above, but with $\delta$ varying away from zero. There are not many competitors that can be compared to our specification test. We consider two procedures relying on unconditional moments obtained from a given set of instruments. First, the J-test of overidentification based on the continuously updated GMM with conservative critical values, hereafter JCUE, as recently discussed in AntoineFrazierRenault2026WP\footnote{AntoineFrazierRenault2026WP also propose a conditional approach that exploits information about the existence of strongly identified directions in the parameter space to define data-dependent critical values that are less conservative than the robust ones. We do not explicitly consider this approach here.}. Regardless of identification strength, the CUE objective function is asymptotically upper-bounded by a random variable distributed as a chi-square with degrees of freedom equal to the number of instruments, so one obtains simple conservative critical values. We also compute the jackknife T-specification test proposed by Chao2014, hereafter JackT. The test is based on (i) estimating \(\beta_0\) by HFULL, a heteroskedasticity-robust version of the Fuller estimator proposed by hausman_instrumental_2012, and (ii) considering a jackknife version of the overidentification statistic, based on the objective function of the JIVE2 estimator proposed by AngristImbensKrueger1999. The associated statistic is asymptotically distributed as a chi-square with degrees of freedom equal to the degree of overidentification, when the number of instruments increases with the sample size, or when it is fixed under maintained homoskedasticity. Hence, while the test is generally valid under the framework of {\em many weak instruments}, it may not be in the setup considered here.
We consider different versions of JCUE and JackT, respectively with 1, 3, and 7 instruments for one variable $Z$, and 3, 7, or 15 instruments for two variables $Z_1$ and $Z_2$, see Section (ref) for details (note that JackT cannot be used with one instrument only, as there is exact identification). The empirical sizes are reported in Table (ref). Our specification test and the tests based on JCUE are undersized as expected. The test based on JackT with 3 instruments is oversized for the polynomial model, but size controls improve with 7 instruments, as well as for the linear and polynomial group models. In Figure (ref), we present power curves. For our benchmark polynomial model, the JCUE test has trivial power with one instrument. Power improves with 3 instruments, then decreases with 7 instruments. Our HICM$^*$ test has similar power to the best of the JCUE tests. The JackT tests are more powerful than our test for small values of $\delta$, but their power curves flatten for higher values. As for JCUE, increasing the number of instruments negatively affects power. In the stronger identification setup, all tests are powerful, but JCUE with only one instrument. In the linear model and the polynomial group model, we found the same main features as in the benchmark model. However, JackT with 3 instruments has trivial power for the polynomial group model. Throughout the different designs, the small sample properties of both JackT and JCUE are affected by the number of instruments. By constrast, our specification test consistently controls size and is powerful regardless of identification strength and the particular functional form that links instruments and endogenous variables.
We revisit antoine_identification-robust_2023's empirical application which extends some of the results presented in Mexico. They study the impact of a large population collapse in 16th-century Mexico on land institutions by adopting an instrumental-variables empirical strategy based on the characteristics of a massive epidemic in the mid-1570s, which is believed to have been caused by a rodent-transmitted pathogen that emerged after several years of drought were followed by a period of above-average rainfall. As explained by these authors, the sharp decline in population lowered the costs and increased the benefits of acquiring land from indigenous villages in many areas. As their three excluded instrumental variables, they use proxies for climate conditions : (i) drought, the sum of the 2 lowest consecutive PDSI values between 1570 and 1575 (more negative numbers indicate severe and prolonged drought), (ii) rainfall, the maximum PDSI between 1576 and 1580 (as a measure of excess rainfall), and (iii) gap, the difference between the minimum PDSI between 1570 and 1575 and the maximum between 1576 and 1580.\footnote{The Palmer Drought Severity Index (PDSI) is a normalized measure of soil moisture that captures deviations from typical conditions at a given location.}
Using the data constructed by Mexico, we estimate the short-term effects of the population collapse in the model
where $y_i$ is the inverse hyperbolic sine of the percent rural population living in hacienda communities in 1900, $Y_{2i}$ is the population decline in municipality measured as the log ratio of 1650 and 1570 density, $X_{2i}$ represents the vector of the climate instruments, and $X_{1i}$ is a vector of control variables of geographic and economic features, namely the standard deviation of PDSI, a measure of maize productivity, various measures of elevation and slope, as well as the log of tributary density in 1570 and governorship-level fixed effects, see Column 6 in Table 2 in Mexico.
In Table (ref), we report confidence regions obtained with three inference procedures that simultaneously test the null and the validity of the model: (i) HICM introduced in this paper, (ii) ICM introduced in antoine_identification-robust_2023, and (iii) S introduced in SW assuming a linear reduced form. Following the analysis in antoine_identification-robust_2023, we use as instruments the two most reliable variables gap and drought, as well as regional dummies as additional controls.\footnote{When using the three climate instruments with regional dummies, the model is rejected.} All three procedures yield a significant and negative effect of the population decline on the percent rural population living in hacienda as expected: a decrease in the ratio of 1650 to 1570 density increases the likelihood of having more large estates per area in 1900. The confidence set obtained with HICM is the narrowest and is fully enclosed in the ones obtained with S or ICM. Computationally, the advantage of HICM over ICM is clear. Since its critical value does not depend on the parameter values, it only needs to be computed once for the entire grid of candidates. Accordingly, and similarly to S, the associated confidence set can be obtained much faster than by ICM. With two instruments and a grid of 2,500 candidate points, it takes under 50 seconds for HICM, but over 52 minutes for ICM!
While Mexico and antoine_identification-robust_2023 assume throughout that the effect of the population collapse $Y_{2}$ on land tenure $y$ is linear, we relax this assumption by considering instead nonlinear specifications with either 2 or 3 endogeneous variables, specifically
Both specifications explore whether more negative values of the population collapse - that is, more severe population declines - are driving the previous results.\footnote{We thank I. Andrews for this suggestion.} Specification S2 introduces an indicator which targets observations below/above the 20% quantile of $Y_2$ (equal to -1.89), whereas Specification S3 includes an additional indicator to target observations above the 80% quantile of $Y_2$ (equal to -0.89).
Multivariate confidence regions are obtained using a two-dimensional grid of evenly spread points over $[-1.5,-0.5]\times[-8,1]$ (with respective steps of 0.01 and 0.05) for Specification S2 and a three-dimensional grid of evenly spread points over $[-2.7, 0.2]\times[-7.5,0]\times[-18, 4.5]$ (with respective steps 0.05, 0.05 and 0.1) for Specification S3. Other details are similar to above. We only consider the procedures based on HICM and S: based on our previous results, it would take weeks to obtain a multivariate confidence region with ICM for Specification S3. For both specifications, HICM provides finite and bounded confidence regions, while S provides large and possibly unbounded ones. In Figure (ref), we report the 95% confidence regions for the parameters of Specification S2. The HICM confidence region is small and bounded, while that obtained with S is much larger and possibly unbounded.\footnote{As a robustness check, we consider much larger grids of candidate points. Figure (ref) in our Appendix (ref) strongly suggests that the S confidence region is indeed unbounded.} In Table (ref), we report the 95% confidence intervals obtained with HICM by projecting the multivariate confidence regions. With Specification S2, the effect of the population collapse below the quantile at 20% is negative, significant, and in line with our results for Specification S1. The effect above the 20% quantile is also negative and significant, potentially larger in magnitude, but also less precisely estimated. Given the width of confidence intervals, the previously estimated linear Specification S1 cannot be ruled out. With Specification S3, the confidence intervals are in line with previous specifications, but only $c_2$ is significantly different from 0 at 95% confidence level.
To put these results into perspective, we revisit part of the counterfactual analysis in Mexico. Specifically, to obtain what landholdings would have been in the absence of a population collapse, we subtract off the predicted marginal effect of the population change in each municipality from the actual 1900 outcome. We find that the distribution of hacienda population changes substantially under our counterfactual. Our results are reported in Table (ref) for the median and the third quartile of the percentage of 1900 population living in haciendas, respectively 16.7% and 44.4%. When we remove the effect of the population collapse given by our estimation strategy, we find that the percentage of 1900 population living in haciendas drops by 1.5% to 4.6% with Specification S1, and by 1.1% to 9.1% with Specification S2. The drop is potentially larger at the third quartile with a lower bound of 2.7% for the linear specification and 2.4% for Specification S2. Overall, such an impact is economically meaningful and practically relevant.
We have developed and studied a new set of inference procedures in a linear IV model based on HICM, an heteroskedastic version of the ICM statistic of antoine_identification-robust_2023. The tests are easy to implement, robust to any identification pattern and unknown heteroskedasticity, and nonparametric with respect to the first-stage equation. The key innovative feature is the independence of critical values from parameter values. As a result, inference is greatly facilitated. In addition, this also allows for straightforward subvector inference and pure specification testing. In simulations, our procedure is competitive with the ICM test of antoine_identification-robust_2023. While subvector inference is conservative, we have shown that in an application that it can yield bounded and informative confidence regions for parameters of interest.