EconBase
← Back to paper

Weak-instrument-robust subvector inference in instrumental variables regression: A subvector Lagrange multiplier test and properties of subvector Anderson-Rubin confidence sets

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

91,652 characters · 15 sections · 176 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Weak-instrument-robust subvector inference in instrumental variables regression: A subvector Lagrange multiplier test and properties of subvector Anderson-Rubin confidence sets

abstractWe propose a weak-instrument-robust subvector Lagrange multiplier test for instrumental variables regression. We show that it is asymptotically size-correct under a technical condition or as the number of instruments grows to infinity. This is the first weak-instrument-robust subvector test for instrumental variables regression to recover the degrees of freedom of the commonly used non-weak-instrument-robust Wald test. Additionally, we provide a closed-form solution for subvector confidence sets obtained by inverting the subvector Anderson-Rubin test. We show that they are centered around a k-class estimator. We show that the subvector confidence sets for single coefficients of the causal parameter are jointly bounded if and only if Anderson's likelihood-ratio {test} rejects the null hypothesis that the first-stage regression parameter is of reduced rank, that is, that the causal parameter is not identified. Finally, we show that if a confidence set obtained by inverting the Anderson-Rubin test is bounded and nonempty, it is equal to a Wald-based confidence set with a data-dependent confidence level. We explicitly compute this Wald-based confidence {set and its confidence level}.

\noindentKeywords causal inference, instrumental variables regression, weak instruments, test inversion

{Accepted for publication in the Journal of Econometrics.}

Introduction

Instrumental variables regression is essential for empirical economics. Identifying the causal effect requires at least as many instruments as endogenous (confounded) variables. Finding valid instruments is difficult in practice, so historically much research considers the causal effect of a single endogenous variable, instrumented by another. However, the trend is towards the availability of larger and more detailed data, allowing modeling settings with more endogenous variables.

For causal effect estimation to influence policy, uncertainty quantification is essential. Typically, tests based on the estimator's asymptotic normality are used to construct confidence sets and $p$-values. staiger1997instrumental note that in many applications, the signal-to-noise ratio in the first stage is low, leading to size distortion of such tests. They prove that, for a single instrument and endogenous variable, if the expected first-stage $F$-statistic exceeds 10, asymptotically, Wald-based 95% confidence sets have at least 85% coverage. This resulted in the following widely adopted heuristic: If the first-stage $F$-statistic is greater than 10, the instrument is considered strong, and Wald-based confidence sets are considered valid.

We believe such pre-testing is flawed. Even though the heuristic is applied more generally, staiger1997instrumental's staiger1997instrumental results for the Wald test's size do not apply for multiple instruments or multiple endogenous variables stock2002testing. In a recent paper, lee2022valid show that that the first stage $F$-statistic rule of 10 can still lead to severe size distortion of the TSLS Wald test and that a larger cutoff is needed for size control. On the other hand, pre-testing might be too conservative, preventing inference given weak instruments but a strong causal effect.

An alternative is using weak-instrument-robust tests. These include the Anderson-Rubin test anderson1949estimation, kleibergen2002pivotal's kleibergen2002pivotal Lagrange multiplier test, and moreira2003conditional's moreira2003conditional conditional likelihood-ratio test. They have the correct size asymptotically, even if the signal-to-noise ratio in the first stage decreases such that the first-stage $F$-statistic is of constant order (weak instrument asymptotics).

For multiple endogenous variables, multivariate confidence sets for the causal parameter are difficult to interpret. Subvector tests and confidence sets for individual coefficients of the causal parameter, similar to $t$-tests in linear regression, are more useful. A naive approach to construct such subvector confidence sets is by projecting the multivariate confidence set onto the coefficient of interest dufour2005projection. However, this results in a loss of power.

Subvector variants of the Anderson-Rubin and conditional likelihood-ratio tests have been proposed by guggenberger2012asymptotic and kleibergen2021efficient. However, with more instruments than endogenous variables, the limiting distributions have higher degrees of freedom than the corresponding Wald test. Consequently, additional instruments, for example by modeling non-linear relations with interactions and non-linear basis expansions, possibly increase the size of the confidence sets. This is undesirable, as the additional instruments should improve identification.

We propose a subvector variant of the Lagrange multiplier test. This recovers the degrees of freedom of the Wald test, independently of the number of instruments. While the test is not necessarily more powerful than the subvector conditional likelihood-ratio test, to our knowledge, this is the first weak-instrument-robust subvector test in instrumental variables regression to achieve this. Our test fills a gap in the weak-instrument-robust testing literature for instrumental variables regression, as shown in (ref).

table[table omitted — 497 chars of source]

dufour2005projection note that the confidence sets obtained by inverting the Anderson-Rubin test are quadrics, as are their projections. We provide a closed-form solution for the inverse Anderson-Rubin test confidence sets using the more powerful critical values by guggenberger2012asymptotic and show that they are centered around a k-class estimator. Additionally, we prove that the subvector confidence sets for a single coefficient of the causal parameter are jointly bounded if and only if anderson1951estimating's anderson1951estimating likelihood-ratio test rejects the hypothesis that the first-stage regression parameter is of reduced rank, that is, that the causal parameter is not identified. Finally, we show that if the confidence set obtained by inverting the Anderson-Rubin test is bounded and nonempty, it is equal to a Wald-based confidence set around a k-class estimator. We explicitly compute this k-class estimator's $\kappa$ parameter and the Wald confidence set's confidence level.

Implementations for k-class estimators, the (subvector) Anderson-Rubin, (conditional) like\-li\-hood-ratio, Lagrange multiplier, and Wald tests, and the confidence sets obtained by inversion can be found in our open-source software package \href{https://github.com/mlondschien/ivmodels/}{ivmodels} for Python. We hope that this simplifies and thus possibly popularizes the use of weak-instrument-robust tests in empirical research. The ivmodels software package is available on \href{https://pypi.org/project/ivmodels/}{PyPI} and \href{https://anaconda.org/conda-forge/ivmodels}{conda-forge}. See the GitHub repository at \href{https://github.com/mlondschien/ivmodels/}{github.com/mlondschien/ivmodels} and the documentation at \href{https://ivmodels.readthedocs.io/}{ivmodels.readthedocs.io} for more details. Code and instructions to reproduce the tables and figures in this paper are available at \href{https://github.com/mlondschien/ivmodels-simulations}{\texttt{github.com/mlondschien/ivmodels-simulations}}.

Main results

We present our main contributions in (ref). See londschien2025overview for an overview of the literature on weak-instrument-robust inference for instrumental variables regression.

Consider a linear instrumental variables regression model with multiple endogenous covariates. In a typical application, one wishes to make (subvector) inferences for each component of the causal parameter. To model this, we split the endogenous covariates and the components of the causal parameter into two parts, formalized in (ref). Our parameter of interest is $\beta_0$ and we treat $\gamma_0, \Pi_X,$ and $\Pi_W$ as nuisance parameters.

theoremEnd[malte,restate command=modelone]{model} Let $y_i = X_i^T \beta_0 + W_i^T \gamma_0 + \varepsilon_i \in {\mathbb{R}}$ with $X_i = Z_i^T \Pi_X + V_{X, i} \in {\mathbb{R}}^{{m_x}}$ and $W_i = Z_i^T \Pi_W + V_{W, i} \in {\mathbb{R}}^{{m_w}}$ for random vectors $Z_i \in {\mathbb{R}}^k, V_{X, i}\in {\mathbb{R}}^{{m_x}}, V_{W, i}\in{\mathbb{R}}^{{m_w}}$, and $\varepsilon_i \in {\mathbb{R}}$ for $i=1\ldots, n$ and parameters $\Pi_X \in {\mathbb{R}}^{k \times {{m_x}}}$, $\Pi_W \in {\mathbb{R}}^{k \times {{m_w}}}$, $\beta_0 \in {\mathbb{R}}^{{m_x}}$, and $\gamma_0 \in {\mathbb{R}}^{{m_w}}$. We call the $Z_i$ instruments, the $X_i$ endogenous covariates of interest, the $W_i$ endogenous covariates not of interest, and the $y_i$ outcomes. The $V_{X, i}, V_{W, i}$, and $\varepsilon_i$ are errors. These need not be independent across observations. Let $Z \in {\mathbb{R}}^{n \times k}, X \in {\mathbb{R}}^{n \times {{m_x}}}, W \in {\mathbb{R}}^{n \times {{m_w}}}$, and $y \in {\mathbb{R}}^n$ be the matrices of stacked observations. In strong instrument asymptotics, we assume that $\Pi := \begin{pmatrix} \Pi_X & \Pi_W \end{pmatrix}$ is fixed and of full column rank $m := {{m_x}} + {{m_w}}$. In \emph{weak instrument asymptotics} staiger1997instrumental, we assume that $\sqrt{n} \, \Pi = \sqrt{n} \begin{pmatrix} \Pi_X & \Pi_W \end{pmatrix}$ is fixed and of full column rank $m$. Thus, $\Pi = {\cal O}(\frac{1}{\sqrt{n}})$. Both asymptotics imply that $k \geq m$.

In practice, it is common that additional exogenous covariates $C$ enter the model equation. If these are not of interest, one can reduce to (ref) by replacing $y$, $X$, $W$, and $Z$ with their residuals after regressing out $C$. This also allows for the inclusion of an intercept by centering the data. If the exogenous covariates $C$ are of interest, that is, one would like to make inference for their causal effect on $y$, one can include them into the model as both instruments and additional endogenous covariates. We treat this explicitly in Appendix (ref).

In standard asymptotics, the expected first-stage $F$-statistic (or its multivariate extension, see anderson1951estimating, anderson1951estimating) grows with the sample size. staiger1997instrumental note that in many applications, even though the sample size $n$ is large, the $F$-statistic for regressing $X$ on $Z$ is small. In these settings, standard asymptotics no longer apply, motivating weak instrument asymptotics, where the first-stage $F$-statistic is of constant order.

We assume that a central limit theorem applies to the sums $Z^T \varepsilon$, $Z^T V_X$, and $Z^T V_W$.

theoremEnd[malte,restate command=assumptionone]{assumption} Let $$ \Psi := \begin{pmatrix} \Psi_{\varepsilon} & \Psi_{V_X} & \Psi_{V_W} \end{pmatrix} := (Z^T Z)^{-1/2} Z^T \begin{pmatrix} \varepsilon & V_X & V_W \end{pmatrix} \in {\mathbb{R}}^{k \times (1 + m)}. $$ Assume there exist $\Omega \in {\mathbb{R}}^{(1 + m) \times (1 + m)}$ and $Q \in {\mathbb{R}}^{k \times k}$ positive definite such that, as $n \to \infty$, \begin{align*} &\mathrm{(a)} \ \ \frac{1}{n} \begin{pmatrix}\varepsilon & V_X & V_W\end{pmatrix}^T \begin{pmatrix}\varepsilon & V_X & V_W \end{pmatrix} \overset{{\mathbb{P}}}{\to} \Omega, \\ &\mathrm{(b)} \ \ \operatorname{vec}(\Psi) \overset{d}{\to} {\cal N}(0, \Omega \otimes \mathrm{Id}_k), and \\ &\mathrm{(c)} \ \ \frac{1}{n} Z^T Z \overset{{\mathbb{P}}}{\to} Q, \end{align*} where $\operatorname{Cov}(\operatorname{vec}(\Psi)) = \Omega \otimes \mathrm{Id}_k$ means $\operatorname{Cov}(\Psi_{i, j}, \Psi_{i', j'}) = 1_{i = i'} \cdot \Omega_{j, j'}$.

(ref) is standard in the literature. It is equivalent to the assumption made by kleibergen2002pivotal and staiger1997instrumental and is implied by assumptions made by guggenberger2012asymptotic, guggenberger2019more, and kleibergen2021efficient. If the $(Z_i, \varepsilon_i, V_{X, i}, V_{W, i})$ are i.i.d.\ with finite second moments and if the noise terms $\varepsilon, V_{X, i}$, and $V_{W, i}$ are homoscedastic, centered, and uncorrelated with the $Z_i$, then (ref) holds londschien2025overview. {Crucially, our results do not allow for general forms of heteroskedasticity. For inference that is robust to heteroskedasticity, see guggenberger2024powerful.}

A subvector Lagrange multiplier test

We propose a subvector extension of the Lagrange multiplier test statistic from kleibergen2002pivotal. For any $A \in {\mathbb{R}}^{p \times q}$, let $P_A := A (A^T A)^\dagger A^T$, where $\dagger$ denotes the Moore–Penrose inverse, be the projection matrix onto the column span of $A$. Write $M_A := \mathrm{Id}_{p} - P_A$ for the projection onto its orthogonal complement.

theoremEnd[malte,restate command=defsubvectorklmteststatistic]{definition} Let $$ \tilde S(\beta, \gamma) := \begin{pmatrix} X & W \end{pmatrix} - (y - X \beta - W \gamma) \frac{(y - X \beta - W \gamma)^T M_Z \begin{pmatrix} X & W \end{pmatrix}}{(y - X \beta - W \gamma)^T M_Z (y - X \beta - W \gamma)}. $$ The subvector Lagrange multiplier test statistic is \begin{equation} \operatorname{LM}(\beta) := (n - k) \min_{\gamma \in {\mathbb{R}}^{{m_w}}} \frac{(y - X \beta - W \gamma)^T P_{P_Z \tilde S(\beta, \gamma)} (y - X \beta - W \gamma)}{(y - X \beta - W \gamma)^T M_Z (y - X \beta - W \gamma)}. \end{equation}

This reduces to kleibergen2002pivotal's kleibergen2002pivotal Lagrange multiplier test if ${{m_w}}=0$.

guggenberger2012asymptotic analyse the subvector Lagrange multiplier test statistic obtained by plugging in the {limited information maximum likelihood (LIML)} estimator {(see, e.g., Definition 6 of londschien2025overview, {londschien2025overview})} $\hat\gamma_\mathrm{LIML}$ using outcomes $y - X \beta$, covariates $W$, and instruments $Z$ for $\gamma$ into Equation (ref). They propose a data-generating process for which this subvector Lagrange multiplier test with $\chi^2({{m_x}})$ critical values is size distorted and thus show that it is not weak-instrument-robust. {The LIML minimizes the Anderson-Rubin test statistic, that is, $\hat\gamma_\mathrm{LIML} = \operatorname*{argmin}_\gamma \operatorname{AR}(\beta, \gamma)$ (see, e.g., Corollary 9 of londschien2025overview, londschien2025overview) and guggenberger2012asymptotic show that plugging it into the Anderson-Rubin test statistic yields a size-correct test with $\chi^2(k - {{m_w}})$ critical values.} {However, the minimization over $\gamma$ in (ref)} is not equivalent to plugging in the LIML. We show that our proposed subvector Lagrange multiplier test has the correct size under a technical condition.

theoremEnd{technical_condition} Assume there exists a $\gamma^\star \in {\mathbb{R}}^{{m_w}}$ such that $$ \gamma^\star = \gamma_0 + (\Pi_W^T Z^T P_{P_Z \tilde S(\beta_0, \gamma^\star)} W)^{-1} \Pi_W^T Z^T P_{P_Z \tilde S(\beta_0, \gamma^\star)} \varepsilon, $$ or, equivalently, $$ \Pi_W^T Z^T P_{P_Z \tilde S(\beta_0, \gamma^\star)} (\varepsilon + W(\gamma_0 - \gamma^\star)) = 0. $$
theoremEnd{theorem} Consider (ref) and assume that (ref) and (ref) hold. Under the null $\beta = \beta_0$, under both strong and weak instrument asymptotics, the subvector Lagrange multiplier test statistic is bounded from above by a random variable that is asymptotically $\chi^2({{m_x}})$ distributed.

The proof of (ref) uses ideas from kleibergen2002pivotal and guggenberger2012asymptotic. We provide some intuition below. For the full proof see the proof of (ref) in Appendix (ref). This extends (ref) to the case with included exogenous variables.

Under weak instrument asymptotics, kleibergen2002pivotal's kleibergen2002pivotal Lagrange multiplier test statistic can be shown to be equal to

equation*[equation* omitted — 175 chars of source]

where $\Psi_\varepsilon \sim {\cal N}(0, \sigma^2_\varepsilon \cdot \mathrm{Id}_k), \ \Psi_X \sim {\cal N}(Q^{1/2} \Pi, \Omega_{V} \otimes \mathrm{Id}_k),$ and $\Psi_{\tilde X(\beta)} \sim {\cal N}(\mu(\beta), \tilde\Omega(\beta) \otimes \mathrm{Id}_k) $ {for some $\tilde \Omega(\beta)$,} are asymptotically jointly Gaussian. Importantly, $\tilde X(\beta)$, {equal to $\tilde S(\beta)$ from (ref) when ${{m_w}}=0$,} is constructed such that $\Psi_\varepsilon + \Psi_{X} (\beta_0 - \beta)$ and {$\Psi_{\tilde X(\beta)} := (Z^T Z)^{-1/2} Z^T \tilde X(\beta)$} are asymptotically independent. Thus, under the null $\beta = \beta_0$, conditionally on $\Psi_{\tilde X(\beta_0)}$, we have that $\| P_{\Psi_{\tilde X(\beta_0)}} \Psi_\varepsilon \|^2 \to_{\mathbb{P}} \sigma_\varepsilon^2 \chi^2(m)$, where $m = \rank(P_{\Psi_{\tilde X(\beta_0)}})$. Integrating over $\Psi_{\tilde X(\beta_0)}$, this holds also unconditionally and $\operatorname{LM}(\beta_0) \to_d \chi^2(m)$ follows from $\hat\sigma_\varepsilon^2(\beta_0) \to_{\mathbb{P}} \sigma_\varepsilon^2$.

In contrast, guggenberger2012asymptotic construct $\gamma^\dagger$ in such a manner that

align*[align* omitted — 363 chars of source]

where $\Psi_\varepsilon \sim {\cal N}(0, \sigma^2_\varepsilon \cdot \mathrm{Id}_k)$ { and $\Psi_{\tilde W(\beta_0, \gamma_0)} := (Z^T Z)^{-1/2} Z^T \tilde W(\beta_0, \gamma_0) \sim {\cal N}(\mu, \tilde\Omega \otimes \mathrm{Id}_k)$} are asymptotically jointly Gaussian and independently distributed, {for $(\tilde X(\beta, \gamma), \tilde W(\beta, \gamma) ) := \tilde S(\beta, \gamma)$ and some invertible $\tilde \Omega$}. Thus, under the null $\beta = \beta_0$, conditionally on $\Psi_{\tilde W(\beta_0, \gamma_0)}$ we have that $\| M_{\Psi_{\tilde W(\beta_0, \gamma_0)}} \Psi_\varepsilon \|^2 \to_{\mathbb{P}} \sigma_\varepsilon^2 \chi^2(k - {{m_w}})$, where $k - {{m_w}} = \rank(M_{\Psi_{\tilde W(\beta_0, \gamma_0)}}) = \rank(\mathrm{Id}_k - P_{\Psi_{\tilde W(\beta_0, \gamma_0)}})$. This then also holds unconditionally. The subvector Anderson-Rubin test statistic is then bounded from above by a $\chi^2(k - {{m_w}}) / (k - {{m_w}})$ random variable, as $\operatorname*{plim} \hat\sigma_\varepsilon^2(\beta, \gamma^\dagger) / \mathrm{const(\gamma^\dagger)} \geq \sigma_\varepsilon^2$ and $\min_\gamma \operatorname{AR}(\beta_0, \gamma) \leq \operatorname{AR}(\beta_0, \gamma^\dagger)$.

Note that the $\gamma^\dagger$ is constructed such that $\Psi_\varepsilon + \Psi_W (\gamma_0 - \gamma^\dagger) = M_{\Psi_{\tilde W(\beta_0, \gamma_0)}} \Psi_\varepsilon \neq M_{\Psi_{\tilde W(\beta_0, \gamma^\dagger)}} \Psi_\varepsilon$. The random variable $\Psi_{\tilde W(\beta_0, \gamma^\dagger)}$ is asymptotically independent of $\Psi_\varepsilon + \Psi_{W} (\gamma_0 - \gamma^\dagger)$, but not of $\Psi_\varepsilon$, and so, {for the subvector Lagrange multiplier test}, the argument based on conditioning on $\Psi_{\tilde W(\beta_0, \gamma^\dagger)}$ does not work.

The {subvector} Lagrange multiplier test statistic is equal to $\min_\gamma \operatorname{LM}(\beta, \gamma)$, where $$\operatorname{LM}(\beta, \gamma) = \frac{1}{\hat\sigma_\varepsilon^2(\beta, \gamma)} \| P_{[\Psi_{\tilde X(\beta, \gamma)}, \Psi_{\tilde W(\beta, \gamma)}]} (\Psi_\varepsilon + \Psi_{X} (\beta - \beta_0) + \Psi_W (\gamma_0 - \gamma)) \|^2.$$ If we plug in $\gamma^\dagger$ as above, we obtain $\operatorname{LM}(\beta_0, \gamma^\dagger) = \frac{1}{\hat\sigma_\varepsilon^2(\beta_0, \gamma^\dagger)} \| P_{[\Psi_{\tilde X(\beta_0, \gamma^\dagger)}, \Psi_{\tilde W(\beta_0, \gamma^\dagger)}]} M_{\Psi_{\tilde W(\beta_0, \gamma_0)}} \Psi_\varepsilon \|^2$. The first projection is onto the column span of $[\Psi_{\tilde X(\beta_0, \gamma^\dagger)}, \Psi_{\tilde X(\beta_0, \gamma^\dagger)}]$, which is not independent of $\Psi_\varepsilon$. (ref) ensures the existence of a $\gamma^\star$ such that $$\operatorname{LM}(\beta_0, \gamma^\star) \overset{{\mathbb{P}}}{\to} \frac{1}{\hat\sigma_\varepsilon^2(\beta_0, \gamma^\star)} \| P_{M_{Q^{1/2} \Pi_W} [\Psi_{\tilde X(\beta_0, \gamma^\star)}, \Psi_{\tilde W(\beta_0, \gamma^\star)}]} (\Psi_\varepsilon + \Psi_{V_W} (\gamma_0 - \gamma^\star)) \|^2,$$ where $\Psi_\varepsilon^\star := \Psi_\varepsilon + \Psi_{V_W} (\gamma_0 - \gamma^\star) \sim {\cal N}(0, \sigma_{\varepsilon^\star}^2 \cdot \mathrm{Id}_k)$ is independent of $[\Psi_{\tilde X(\beta_0, \gamma^\star)}, \Psi_{\tilde W(\beta_0, \gamma^\star)}]$. We can thus apply the conditioning argument as above to show that $$\operatorname{LM}(\beta_0, \gamma^\star) \overset{d}{\to} \chi^2({{m_x}}), \text{ where }{{m_x}} = m - {{m_w}} = \rank(P_{M_{Q^{1/2} \Pi_W} [\Psi_{\tilde X(\beta_0, \gamma^\star)}, \Psi_{\tilde W(\beta_0, \gamma^\star)}]}).$$ The result then follows from $\hat\sigma_\varepsilon^2(\beta_0, \gamma^\star) \to_{\mathbb{P}} \sigma_{\varepsilon^\star}^2$ and $\min_\gamma \operatorname{LM}(\beta_0, \gamma) \leq \operatorname{LM}(\beta_0, \gamma^\star)$.

We were unable to prove that the (ref) holds in general with high probability and leave this as future work. However, we empirically observe that the statement of (ref) holds in great generality for finite sample sizes and fixed number of instruments $k$, see also the numerical analyses in (ref). We thus conjecture the following:

theoremEnd[malte,restate command=conjsubvectorklmteststatistic]{conjecture} (ref) holds with probability tending to 1 as $n \to \infty$.

Recall the issue with using $\gamma^\dagger$ as in guggenberger2012asymptotic: $\Psi_{\tilde W(\beta_0, \gamma^\dagger)}$ is independent of $\Psi_\varepsilon + \Psi_W (\gamma_0 - \gamma^\dagger)$, but not of $\Psi_\varepsilon$. However, we can show that as $k \to \infty$, we have that $\Psi_{\tilde W(\beta_0, \gamma^\dagger)} \to_{\mathbb{P}} \Psi_{\tilde W(\beta_0, \gamma_0)}$. This yields the following result.

theoremEnd{theorem} Consider (ref) and assume that (ref) holds as $k = k(n) \to \infty$ as $n\to\infty$. Here, one needs to reformulate (ref) (b) with convergence of the difference between $\Psi$ and a Gaussian random variable to zero for all probabilities of Borel-measurable sets (see also bentkus2003dependence, bentkus2003dependence and chernozhukov2017central, chernozhukov2017central). Under the null $\beta = \beta_0$, under both strong and weak instrument asymptotics, the subvector Lagrange multiplier test statistic is bounded from above by a random variable that is asymptotically $\chi^2({{m_x}})$ distributed.

For the proof of (ref), see the proof of (ref) in Appendix (ref). (ref) extends (ref) to the case with included exogenous variables.

Note that (ref) as $k = k(n) \to \infty$ is a stronger assumption than (ref). Lemma 1 of londschien2025overview can be extended to allow for a $k = k(n)$ that slowly grows with $n$, for example $k(n) = \log(n)$, by invoking a central limit theorem in slowly growing dimension, see also bentkus2003dependence. We note that an arbitrarily slow growth of $k = k(n)$ in (ref) is sufficient for (ref).

(ref) directly imply the following corollary:

theoremEnd[category=index,malte_intro,restate command=corsubvectorlagrangemultipliertest]{corollary} Consider (ref) and assume the conditions of (ref) or (ref). Let $F_{\chi^2({{m_x}})}$ be the cumulative distribution function of a chi-squared random variable with ${{m_x}}$ degrees of freedom. Under both strong and weak instrument asymptotics, a test that rejects the null $H_0: \beta = \beta_0$ whenever $\operatorname{LM}(\beta) > F^{-1}_{\chi^2({{m_x}})}(1 - \alpha)$ has asymptotic size at most $\alpha$.

Thus, under (ref) or if $k = k(n) \to \infty$, the weak-instrument-robust subvector Lagrange multiplier test achieves the same degrees of freedom as the commonly used Wald test that is not robust to weak instruments. To our knowledge, this is the first weak-instrument-robust subvector test for the causal parameter in instrumental variables regression to achieve this. The subvector Lagrange multiplier test fills a gap in the literature on weak-instrument-robust inference for instrumental variables regression, as shown in (ref).

We numerically analyze the size (and power) of various tests under (a variant of) guggenberger2012asymptotic's (guggenberger2012asymptotic) data-generating process in (ref). Our proposed subvector Lagrange multiplier test is size-correct and substantially more powerful than their subvector Anderson-Rubin test, which uses $\chi^2(k - {{m_w}})$ critical values. We compare properties of various tests in instrumental variables regression in (ref).

table[table omitted — 706 chars of source]

In practice, finding valid instruments is challenging and applications where the number of instruments is much larger than the number of endogenous covariates are rare. However, we encourage improving identification by modeling non-linear relations between instruments and endogenous covariates with interactions and non-linear basis expansions. If using the non-weak-instrument-robust Wald test, the addition of such instruments possibly decreases the concentration parameter $\mu^2 / k := n \cdot {\lambda_\mathrm{min}\mleft(\Omega_V^{-1} \Pi^T Q \Pi\mright)} / k$ staiger1997instrumental,stock2002testing, degrading the the test's approximation quality and possibly leading to size distortion. If using the weak-instrument-robust Anderson-Rubin test, the critical values increase with the number of instruments and additional instruments possibly increase the size of confidence intervals. These issues do not affect the subvector Lagrange multiplier test, which is robust to weak instruments and whose critical values do not depend on the number of instruments, as described by (ref).

In (ref), we apply various tests to data from card1993using and tanaka2010risk. The addition of instrument interactions generally leads to smaller $p$-values for the Lagrange multiplier test, which is not the case for the Anderson-Rubin and conditional likelihood-ratio tests. {As a score-type test, the Lagrange multiplier test statistic is zero at all stationary points of the likelihood or equivalently of the Anderson-Rubin test statistic. As pointed out by a reviewer and established by andrews2006optimal and kleibergen2007generalizing, there are $m + 1$ such stationary points. In practice, this results in reduced power around such stationary points, see also (ref). Confidence sets obtained by inversion of the subvector Lagrange multiplier test are possibly disjoint, with the number of components equal to the number of small eigenvalues of the concentration matrix (the number of weakly identified endogenous variables) plus one. We observe this phenomenon in (ref). }

Even if there is only one endogenous covariate, subvector inference is needed to construct weak-instrument-robust confidence sets for included exogenous covariates. In a typical instrumental variables specification, inference is needed for the causal effect of endogenous covariates after accounting for included exogenous covariates. However, there are also use cases where the focus is on the causal effect of an exogenous covariate after accounting for an endogenous covariate, and instruments are required to identify the nuisance parameter thams2022identifying. For example, tanaka2010risk discuss the effect of the effect of (exogenous) gender on risk preferences, after accounting for (endogenous) income.

kleibergen2021efficient suggest to test a hypothesis on the causal effect of an included exogenous covariate with their subvector conditional likelihood-ratio test by including it both as an endogenous covariate and an instrument. However, both their proof and that of the full vector conditional likelihood-ratio test moreira2003conditional depend on invertibility of the covariance estimate $\hat\Omega = \frac{1}{n - k} X^T M_Z X$. This is not the case if an included exogenous covariate is included into both $X$ and $Z$. londschien2025overview explicitly computes the limiting distribution of the full vector conditional likelihood-ratio test statistic with included exogenous covariates. Their limiting distribution is strictly smaller than the one that would be obtained by simply including the included exogenous covariates as both endogenous covariates and instruments. That is, testing hypotheses on the causal effect of included exogenous covariates by simply including them into both instruments and endogenous variables and applying moreira2003conditional's (moreira2003conditional) conditional likelihood-ratio test is conservative and possibly leads to unnecessarily large confidence sets. It is not clear whether londschien2025overview's (londschien2025overview) result on the distribution of the conditional likelihood-ratio test statistic with included exogenous covariates generalizes to kleibergen2021efficient's kleibergen2021efficient subvector test.

In contrast, we treat the setting with included exogenous covariates explicitly and prove versions of (ref) and (ref) that allow for included exogenous variables under test in (ref) and (ref) in Appendix (ref).

The optimization problem in (ref) is non-convex and thus difficult to solve. In our ivmodels software package, which was used for the simulations in (ref), we use the quasi-Newton method of Broyden–Fletcher–Goldfarb–Shannon (BFGS). This is the default option of scipy.optimize.minimize in Python. While this is not guaranteed to find the global optimum, in practice we observe that it suffices to find a local optimum next to the LIML for the test to be size-correct. See Appendix (ref) for details.

Closed-form subvector confidence sets obtained by inverting the Anderson-Rubin test

The Anderson-Rubin test is an important test for instrumental variables regression. anderson1949estimation showed that the statistic is asymptotically $\chi^2(k) / k$ distributed, independently of instrument strength. guggenberger2012asymptotic proposed the following subvector variant.

theoremEnd[malte, restate command=subvectorandersonrubinteststatistic]{definition}[guggenberger2012asymptotic, guggenberger2012asymptotic] The subvector Anderson-Rubin test statistic is \begin{align*} \operatorname{AR}(\beta) &:= \min_{\gamma \in {\mathbb{R}}^{{m_w}}} \frac{n - k}{k - {{m_w}}} \frac{(y - X \beta - W \gamma)^T P_Z (y - X \beta - W \gamma)}{(y - X \beta - W \gamma)^T M_Z (y - X \beta - W \gamma)} \\ &= \frac{n - k}{k - {{m_w}}} \frac{(y - X \beta - W \hat\gamma_\mathrm{LIML})^T P_Z (y - X \beta - W \hat\gamma_\mathrm{LIML})}{(y - X \beta - W \hat\gamma_\mathrm{LIML})^T M_Z (y - X \beta - W \hat\gamma_\mathrm{LIML})}, \end{align*} where $\hat\gamma_\mathrm{LIML}$ is the LIML estimator using outcomes $y - X \beta$, covariates $W$, and instruments $Z$.

guggenberger2012asymptotic showed that under (ref) and (ref), the subvector Anderson-Rubin test statistic is asymptotically bounded from above by a ${\chi^2(k - {{m_w}}) / (k - {{m_w}})}$ distributed random variable, also independently of instrument strength. This yields weak-instrument-robust confidence sets for subvectors of the causal parameter via test inversion: $$ \operatorname{CI}_{\operatorname{AR}}(1 - \alpha) := \{ \beta \in {\mathbb{R}}^{{{m_x}}} \mid (k - {{m_w}}) \cdot \operatorname{AR}(\beta) \leq F^{-1}_{\chi^2(k - {{m_w}})}(1 - \alpha) \}. $$ We give an explicit form for these confidence sets.

theoremEnd[normal]{proposition} The inverse Anderson-Rubin test confidence sets describe quadrics in ${\mathbb{R}}^{{{m_x}}}$. Let $S := \begin{pmatrix} X & W \end{pmatrix}$ and $B \in {\mathbb{R}}^{{{m_x}} \times m}$ have ones on the diagonal and zeros elsewhere such that $BS = X$. Let $F_{\chi^2(k - {{m_w}})}$ denote the cumulative distribution function of a $\chi^2(k - {{m_w}})$ random variable and let $$ \hat\beta_\mathrm{k}(\kappa) := \left( S^T (\mathrm{Id}_n - \kappa M_Z) S \right)^{\dagger} S^T (\mathrm{Id}_n - \kappa M_Z) y, $$ where $\dagger$ denotes the Moore-Penrose pseudoinverse, be the k-class estimator for $(\beta_0^T, \gamma_0^T)^T$ using covariates $S$, instruments $Z$, and outcome $y$. Define \begin{align*} \kappa_{\operatorname{AR}}(\alpha) &= 1 + F^{-1}_{\chi^2(k - {{m_w}})}(1 - \alpha) / (n - k) , \ \ \ \hat\sigma^2(\kappa) := \frac{1}{n-k} \| M_Z (y - S \hat\beta_\mathrm{k}(\kappa)) \|^2, \ and \\ A(\kappa) &:= \left( B \left(S^T (\mathrm{Id}_n - \kappa M_Z) S \right)^{-1} B^T \right)^{-1} = X^T (\mathrm{Id}_n - \kappa M_Z) X \\ & - X^T (\mathrm{Id}_n - \kappa M_Z) W \left(W^T (\mathrm{Id}_n - \kappa M_Z) W \right)^{-1} W^T (\mathrm{Id}_n - \kappa M_Z) X. \end{align*} If ${{m_w}} > 0$, let $\kappa_\mathrm{max} := {\lambda_\mathrm{min}\mleft((W^T M_Z W)^{-1} W^T P_Z W\mright)} + 1$, where $\lambda_\mathrm{min}$ denotes the minimal eigenvalue of a matrix. Else $\kappa_{\max} := \infty$. If $\kappa_{\operatorname{AR}}(\alpha) \leq \kappa_\mathrm{max}$, then \begin{multline*} \operatorname{CI}_{\operatorname{AR}}(1 - \alpha) = \Big\{ \beta \in {\mathbb{R}}^{{{m_x}}} \mid \left(\beta - B \hat\beta_\mathrm{k}(\kappa_{\operatorname{AR}}(\alpha))\right)^T A(\kappa_{\operatorname{AR}}(\alpha)) \left(\beta - B\hat\beta_\mathrm{k}(\kappa_{\operatorname{AR}}(\alpha)) \right) \\ \leq \hat\sigma^2(\kappa_{\operatorname{AR}}(\alpha)) \cdot \big(F^{-1}_{\chi^2(k - {{m_w}})}(1 - \alpha) - k \operatorname{AR}(\hat\beta_\mathrm{k}(\kappa_{\operatorname{AR}}(\alpha)))\big)\Big\}. \end{multline*} Else, $\operatorname{CI}_{\operatorname{AR}}(1 - \alpha) = {\mathbb{R}}^{{{m_x}}}$.

For the proof of (ref), see the proof of (ref) in Appendix (ref). In (ref), we also compute the closed-form solutions of the subvector inverse Wald and likelihood-ratio tests. To derive conditions for the subvector inverse Anderson-Rubin test confidence sets to be bounded, we introduce the following technical condition.

theoremEnd[malte, restate command=tctwo]{technical_condition} Let $S = \begin{pmatrix} X & W \end{pmatrix}$. \begin{enumerate}[(a)] • If ${{m_w}} > 0$, then $$ \lambda := {\lambda_\mathrm{min}\mleft( (S^T M_Z S)^{-1} S^T P_Z S\mright)} < {\lambda_\mathrm{min}\mleft( (X^T M_Z X)^{-1} X^T P_Z X\mright)}, $$ or, equivalently, $$ {\lambda_\mathrm{min}\mleft( X^T (P_Z + (1 - \lambda) M_Z) X \mright)} > 0. $$ • Let $S^{(-i)}$ be $S$ with the $i$-th column removed. Then, for all $i = 1, \ldots, m$, $$ \lambda < {\lambda_\mathrm{min}\mleft( (S^{(-i)T} M_Z S^{(-i)})^{-1} S^{(-i)T} P_Z S^{(-i)} \mright)} $$ or, equivalently, $$ {\lambda_\mathrm{min}\mleft( S^{(-i)T} (P_Z + (1 - \lambda) M_Z) S^{(-i)} \mright)} > 0. $$ \end{enumerate}

If the noise in (ref) is absolutely continuous with respect to the Lebesgue measure, then (ref) holds with probability one.

theoremEnd[malte_intro]{proposition} Assume (ref) (ref) {holds}. Let $\alpha > 0$. Let \begin{align*} J_\mathrm{LIML} &:= (k - {{m_w}}) \cdot \min_b \operatorname{AR}(b) \\ &= (n - k) \ {\lambda_\mathrm{min}\mleft( \left(\begin{pmatrix} X & W & y \end{pmatrix}^T M_Z \begin{pmatrix} X & W & y \end{pmatrix} \right)^{-1}\begin{pmatrix} X & W & y \end{pmatrix}^T P_Z \begin{pmatrix} X & W & y \end{pmatrix}\mright)} \end{align*} be the LIML variant of the J-statistic guggenberger2012asymptotic and let $$ \lambda = (n - k) \ {\lambda_\mathrm{min}\mleft( \left(\begin{pmatrix} X & W \end{pmatrix}^T M_Z \begin{pmatrix} X & W \end{pmatrix} \right)^{-1}\begin{pmatrix} X & W \end{pmatrix}^T P_Z \begin{pmatrix} X & W \end{pmatrix}\mright)}. $$ be anderson1951estimating's anderson1951estimating likelihood-ratio test statistic of reduced rank. Then the (subvector) inverse Anderson-Rubin test confidence set $\operatorname{CI}_{\operatorname{AR}}(1 - \alpha)$ is bounded and nonempty if and only if \begin{align*} J_\mathrm{LIML} \leq F^{-1}_{\chi^2(k - {{m_w}})}(1 - \alpha) < \lambda \ \Leftrightarrow \ F_{\chi^2(k - {{m_w}})} \left(J_\mathrm{LIML} \right) \leq 1 - \alpha < F_{\chi^2(k - {{m_w}})} \left( \lambda \right). \end{align*}

For the proof of (ref), see the proof of (ref) in Appendix (ref). (ref) additionally provides a condition for the (subvector) inverse likelihood-ratio test confidence set to be bounded. If the (ref) (ref) does not hold, then the subvector inverse Anderson-Rubin test confidence set for $1 - \alpha = F^{-1}_{\chi^2(k - {{m_w}})}(\lambda)$ may be bounded. See also (ref) in {Appendix}.

One is typically interested in subvector inference for a single component of the causal parameter, that is, ${{m_x}} = 1$, as this is easier to interpret. We characterize when the subvector inverse Anderson-Rubin test confidence sets for individual components are (jointly) bounded and nonempty (thus, forming confidence intervals).

theoremEnd[malte_intro]{corollary} Consider structural equations $y_i = S_i^T \beta_0 + \varepsilon_i$ and $S_i = Z_i^T \Pi + V_i$. We are interested in inference for each component of the causal parameter $\beta_0$. Thus, for each covariate $i = 1, \ldots, m$, we separate $X \leftarrow S^{(i)}$ and $W \leftarrow S^{(-i)}$ and construct confidence intervals $\operatorname{CI}_{\operatorname{AR}}^{(i)}(1 - \alpha)$ for the $i$-th component of $\beta_0$. Assume (ref) (ref) {holds.} Then \begin{enumerate}[(a)] • The subvector inverse Anderson-Rubin test confidence sets are jointly (un)bounded. That is, if $\operatorname{CI}_{\operatorname{AR}}^{(i)}(1 - \alpha)$ is (un)bounded for any $i$, then $\operatorname{CI}_{\operatorname{AR}}^{(i)}(1 - \alpha)$ is (un)bounded for all $i$. • The subvector inverse Anderson-Rubin test confidence sets at level $\alpha$ are bounded (and thus confidence intervals) if and only if anderson1951estimating's anderson1951estimating likelihood-ratio test rejects the null hypothesis that $\Pi$ is of reduced rank at level $\alpha$. • The subvector inverse Anderson-Rubin test confidence sets are jointly (non)empty. That is, if $\operatorname{CI}_{\operatorname{AR}}^{(i)}(1 - \alpha)$ is (non)empty for any $i$, then $\operatorname{CI}_{\operatorname{AR}}^{(i)}(1 - \alpha)$ is (non)empty for all $i$. • The subvector inverse Anderson-Rubin test confidence sets at level $\alpha$ are jointly empty if and only if the LIML variant of the J-statistic guggenberger2012asymptotic $J_\mathrm{LIML} > F^{-1}_{\chi^2(k - m + 1)}(1 - \alpha)$. \end{enumerate}

For the proof of (ref), see the proof of (ref) in Appendix (ref). (ref) additionally gives a condition for the (subvector) inverse likelihood-ratio test confidence sets to be bounded.

{ (ref) summarizes properties of the confidence sets obtained by inverting the Anderson-Rubin test that follow directly from (ref). We hope that this characterization assists in the interpretation of subvector inference with the Anderson-Rubin test. As pointed out by a reviewer, these properties are implicit in the geometric construction of subvector statistics in guggenberger2012asymptotic and kleibergen2021efficient. If $m = {{m_x}} = 1$, anderson1951estimating's anderson1951estimating likelihood-ratio test of reduced rank reduces to the $F$-test of the first-stage regression of $X$ on $Z$ using $\chi^2(k)$ critical values. For this special case, kleibergen2007generalizing notes that the inverse Anderson-Rubin test confidence sets are unbounded if and only if the first-stage $F$-test rejects the null hypothesis of underidentification. For the setting with multiple endogenous covariates (${{m_w}} > 0$), kleibergen2021efficient shows in their Theorem 12a that for values of $\beta$ far from $\beta_0$, the subset Anderson-Rubin statistic equals anderson1951estimating's anderson1951estimating likelihood-ratio test for reduced rank, implying (ref) (b). Property (d) is a direct consequence of the definition of the subvector Anderson-Rubin test and the critical values first derived by guggenberger2012asymptotic.}

(ref) (ref) only holds if ${{m_x}} = 1$, that is, if we are testing a single component of the causal parameter. Otherwise, the condition for boundedness of the inverse Anderson-Rubin test and anderson1951estimating's anderson1951estimating likelihood-ratio test for reduced rank use different critical values. See also (ref) (middle panel), where the 2-dimensional 90% confidence set for the causal parameter is unbounded, even though the $p$-value of the anderson1951estimating's anderson1951estimating likelihood-ratio test for reduced rank is 0.07. As predicted by (ref), the 1-dimensional 90% subvector confidence sets are bounded.

figure[figure omitted — 1,096 chars of source]

On the other hand, the critical values $F^{-1}_{\chi^2(k-m + 1)}(1 - \alpha)$ in (ref) (ref) do not match those of the LIML variant of the J-statistic $F^{-1}_{\chi^2(k-m)}(1 - \alpha)$. In particular, rejection of the LIML J-test does not imply that the inverse Anderson-Rubin confidence sets are empty.

dufour2005projection note that if ${{m_w}} = 0$, confidence sets obtained by inverting the Anderson-Rubin test are quadrics. They propose constructing subvector confidence sets by projection but do not give an explicit closed form for the confidence sets. Nor do they show that the confidence sets are centered around a k-class estimator. Additionally, they do not use the more powerful $\chi^2(k - {{m_w}})$ critical values by guggenberger2012asymptotic, resulting in a substantial loss of power, see (ref).

See (ref) in Appendix (ref) for the definitions of the Wald and likelihood-ratio test statistics. Let $\operatorname{CI}_{\operatorname{Wald}, \hat\beta_k(\kappa)}(1 - \alpha)$ be the $1-\alpha$ confidence set obtained by inverting the subvector Wald test centered around the k-class estimator $\hat\beta_k(\kappa)$ with parameter $\kappa$ and let $\operatorname{CI}_{\operatorname{LR}}(1 - \alpha)$ be the $1 - \alpha$ confidence set obtained by inverting the subvector likelihood-ratio test, As for the subvector Anderson-Rubin test, these are quadrics (see (ref) in Appendix (ref)). Notably, they differ only in the critical value and the $\kappa$ parameter of the k-class estimator they are centered around. We derive the following result:

theoremEnd[malte_intro,restate,restate command=propinversearequaltowald,category=propinversearequaltowald]{proposition} Let $\alpha > 0$. Assume that $J_\mathrm{LIML} \leq F^{-1}_{\chi^2(k-{{m_w}})}(1 - \alpha) < \lambda$ such that $\operatorname{CI}_{\operatorname{AR}}$ is bounded and nonempty by (ref). Let \begin{multline*} s(\kappa) := y^T (\mathrm{Id}_n - \kappa M_Z ) y - y^T (\mathrm{Id}_n - \kappa M_Z ) X (X^T (\mathrm{Id}_n - \kappa M_Z ) X)^{-1} X^T (\mathrm{Id}_n - \kappa M_Z ) y. \end{multline*} Recall that $\kappa_{\operatorname{AR}}(\alpha) = 1 + F^{-1}_{\chi^2(k - {{m_w}})}(1 - \alpha) / (n-k)$ and define \begin{align*} \alpha_{\operatorname{Wald} \mid \operatorname{AR}}(\alpha) &:= 1 - F_{\chi^2({{m_x}})}(-s(\kappa_{\operatorname{AR}}(\alpha)) / \hat\sigma^2_{\operatorname{Wald}}(\kappa_{\operatorname{AR}}(\alpha))) and \\\alpha_{\operatorname{LR} \mid \operatorname{AR}}(\alpha) &:= 1 - F^{-1}_{\chi^2({{m_x}})}( F^{-1}_{\chi^2(k - {{m_w}})}(1 - \alpha) - J_\mathrm{LIML}), \end{align*} where $\hat\sigma^2_{\operatorname{Wald}}(\kappa) = \frac{1}{n-m}\| y - X \hat\beta_\mathrm{k}(\kappa) \|^2$. Then $$ \operatorname{CI}_{\operatorname{AR}}(1 - \alpha) = \operatorname{CI}_{\operatorname{LR}}(1 - \alpha_{\operatorname{LR} \mid \operatorname{AR}}(\alpha)) = \operatorname{CI}_{\operatorname{Wald}_{\hat\beta_\mathrm{k}(\kappa_{\operatorname{AR}}(\alpha))}}(1 - \alpha_{\operatorname{Wald} \mid \operatorname{AR}}(\alpha)). $$
proofEndWe apply (ref). Write $s := s(\kappa_{\operatorname{AR}}(\alpha))$. \paragraph*{Step 1: $J_\mathrm{LIML} \leq F^{-1}_{\chi^2(k - {{m_w}})}(1 - \alpha) < \lambda$ implies $s \leq 0$.} From \begin{multline*} J_\mathrm{LIML} \leq F^{-1}_{\chi^2(k - {{m_w}})}(1 - \alpha) < \lambda \\ \Leftrightarrow {\lambda_\mathrm{min}\mleft( (S^T M_Z S)^{-1} S^T P_Z S \mright)} \leq \kappa_{\operatorname{AR}}(\alpha) - 1 < {\lambda_\mathrm{min}\mleft( (\begin{pmatrix} S & y \end{pmatrix}^T M_Z \begin{pmatrix} S & y \end{pmatrix})^{-1} \begin{pmatrix} S & y \end{pmatrix}^T P_Z \begin{pmatrix} S & y \end{pmatrix} \mright)} \\ \overset{(ref)}{\Leftrightarrow} {\lambda_\mathrm{min}\mleft( S^T (P_Z - \kappa_{\operatorname{AR}}(\alpha) M_Z) S \mright)} \geq 0 > {\lambda_\mathrm{min}\mleft( \begin{pmatrix} S & y \end{pmatrix}^T (P_Z - \kappa_{\operatorname{AR}}(\alpha) M_Z) \begin{pmatrix} S & y \end{pmatrix} \mright)}. \end{multline*} The eigenvalues of $\begin{pmatrix} S & y \end{pmatrix}^T (P_Z - \kappa_{\operatorname{AR}}(\alpha) M_Z) \begin{pmatrix} S & y \end{pmatrix}$ and $ S^T (P_Z - \kappa_{\operatorname{AR}}(\alpha) M_Z) S $ interleave and thus \begin{multline*} \lambda_2 \left(\begin{pmatrix} S & y \end{pmatrix}^T (P_Z - \kappa_{\operatorname{AR}}(\alpha) M_Z) \begin{pmatrix} S & y \end{pmatrix} \right) \geq {\lambda_\mathrm{min}\mleft( S^T (P_Z - \kappa_{\operatorname{AR}}(\alpha) M_Z) S \mright)} \geq 0 \\ \Rightarrow \det( \begin{pmatrix} S & y \end{pmatrix}^T (P_Z - \kappa_{\operatorname{AR}}(\alpha) M_Z) \begin{pmatrix} S & y \end{pmatrix} ) \leq 0. \end{multline*} Applying the formula for the determinant of a block matrix, we get \begin{multline*} \det( S^T (P_Z - \kappa_{\operatorname{AR}}(\alpha) M_Z) S ) \cdot s = \det( \begin{pmatrix} S & y \end{pmatrix}^T (P_Z - \kappa_{\operatorname{AR}}(\alpha) M_Z) \begin{pmatrix} S & y \end{pmatrix} ) \Rightarrow s \leq 0. \end{multline*} \paragraph*{Step 2: $\operatorname{CI}_{\operatorname{AR}}(1 - \alpha) = \operatorname{CI}_{\operatorname{Wald}_{\hat\beta_\mathrm{k}(\kappa_{\operatorname{AR}}(\alpha))}}(1 - \alpha_{\operatorname{Wald} \mid \operatorname{AR}}(\alpha))$} We calculate \begin{multline*} \sigma_{\operatorname{AR}}^2(\kappa_{\operatorname{AR}}(\alpha)) \cdot ( F^{-1}_{\chi^2(k - {{m_w}})}(1 - \alpha) - k \operatorname{AR}(\hat\beta_\mathrm{k}(\kappa_{\operatorname{AR}}(\alpha)))) \\ = (\kappa_{\operatorname{AR}}(\alpha) - 1) \cdot \| M_Z (y - S \hat\beta_\mathrm{k}(\kappa_{\operatorname{AR}}(\alpha)) ) \|^2 - \| P_Z (y - S \hat\beta_\mathrm{k}(\kappa_{\operatorname{AR}}(\alpha)) ) \|^2 \\ = \left( y^T - y^T \left(P_Z - \kappa_{\operatorname{AR}}(\alpha) M_Z \right) S \left( S^T \left( P_Z - \kappa_{\operatorname{AR}}(\alpha) M_Z \right) S \right)^{-1} S^T \right) \left(P_Z - \kappa_{\operatorname{AR}}(\alpha) M_Z \right) \\ \left( y - S \left( S^T \left( P_Z - \kappa_{\operatorname{AR}}(\alpha) M_Z \right) S \right)^{-1} S^T \left( P_Z - \kappa_{\operatorname{AR}}(\alpha) M_Z \right) y \right) = - s. \end{multline*} Clearly $\hat\sigma^2_{\operatorname{Wald}}(\kappa_{\operatorname{AR}}(\alpha)) \geq 0$ and thus \begin{multline*} \hat\sigma^2_{\operatorname{Wald}}(\kappa_{\operatorname{AR}}(\alpha)) \cdot F^{-1}_{\chi^2({{m_x}})}(1 - \alpha_{\operatorname{Wald} \mid \operatorname{AR}}(\alpha)) = - s \\ = \sigma_{\operatorname{AR}}^2(\kappa_{\operatorname{AR}}(\alpha)) \cdot ( F^{-1}_{\chi^2(k - {{m_w}})}(1 - \alpha) - k \operatorname{AR}(\hat\beta_\mathrm{k}(\kappa_{\operatorname{AR}}(\alpha)))). \end{multline*} \paragraph*{Step 3: $\operatorname{CI}_{\operatorname{AR}}(1 - \alpha) = \operatorname{CI}_{\operatorname{LR}}(1 - \alpha_{\operatorname{LR} \mid \operatorname{AR}}(\alpha))$} From (ref), the two confidence sets are equal if $$ \kappa_{\operatorname{AR}}(\alpha) = \kappa_{\operatorname{LR}}(\alpha_{\operatorname{LR} \mid \operatorname{AR}}(\alpha)) \Leftrightarrow F^{-1}_{\chi^2(k - {{m_w}})}(1 - \alpha) = J_\mathrm{LIML} + F^{-1}_{\chi^2({{m_x}})}(1 - \alpha_{\operatorname{LR} \mid \operatorname{AR}}(\alpha)). $$ By assumption $J_\mathrm{LIML} \leq F^{-1}_{\chi^2(k - {{m_w}})}(1 - \alpha)$, so we can solve for $\alpha_{\operatorname{LR} \mid \operatorname{AR}}(\alpha)$.

Let $\hat\kappa_\mathrm{LIML}$ be the $\kappa$ parameter of the LIML's k-class parametrization ${\hat\beta_\mathrm{LIML} = \hat\beta_\mathrm{k}(\hat\kappa_\mathrm{LIML})}$ londschien2025overview. As ${J_\mathrm{LIML} = (n - k) (\hat\kappa_\mathrm{LIML} - 1)}$, if ${J_\mathrm{LIML} \leq F^{-1}_{\chi^2(k-{{m_w}})}(1 - \alpha)}$, then ${\kappa_{\operatorname{AR}}(\alpha) \geq \hat\kappa_\mathrm{LIML}}$.

Numerical analysis of the subvector Lagrange multiplier test

In this section, we compare the empirical size and power of the proposed subvector Lagrange multiplier test with other subvector tests.

We start by comparing the empirical sizes of the different tests for the data-generating process proposed by guggenberger2012asymptotic in (ref). As noted by guggenberger2012asymptotic, the subvector variant of the Lagrange multiplier test obtained by plugging in the LIML estimator is size-inflated. The Lagrange multiplier test proposed by us is size-correct, {meaning that it does not reject the null hypothesis more often than the nominal level $\alpha$.}

Next, we compare the power of the different tests using a variation of guggenberger2012asymptotic's (guggenberger2012asymptotic) data-generating process in (ref). The subvector Lagrange multiplier test proposed by us has power comparable to the subvector conditional likelihood-ratio test and among the highest power out of the tests that are size-correct according to (ref).

Finally, we compare the empirical sizes of the different tests for the data-generating process proposed by kleibergen2021efficient in (ref). Again, the subvector Lagrange multiplier test we propose is size-correct.

Empirical sizes of subvector tests for guggenberger2012asymptotic's (guggenberger2012asymptotic) data-generating process

guggenberger2012asymptotic suggest a data-generating process under which the subvector variant of the Lagrange multiplier test obtained by plugging in the LIML instead of direct minimization as in (ref) is size inflated. In this data-generating process

equation*[equation* omitted — 522 chars of source]

We simulate according to this model (see Appendix (ref) for details) 10'000 times, each time drawing $n=1000$ observations, and report the empirical sizes of the different subvector tests in (ref).

table[table omitted — 888 chars of source]

We replicate the size-distortion found by guggenberger2012asymptotic for the subvector Lagrange multiplier test obtained by {plugging} in the LIML estimator. However, our proposed subvector Lagrange multiplier test is size-correct. (ref) in Appendix (ref) shows quantile-quantile plots comparing the $p$-values of the different tests to the uniform distribution. Empirical sizes of the different subvector tests for $n=50$ and $100$ are in Appendix (ref). The subvector Anderson-Rubin, conditional likelihood-ratio, and Lagrange multiplier tests are size-correct as long as $n/k > 2$.

Power of subvector tests for a variation of guggenberger2012asymptotic's (guggenberger2012asymptotic) data-generating process

Next, we compare the power of different tests using a variation of guggenberger2012asymptotic's (guggenberger2012asymptotic) data-generating process. In the original process, identification is so weak that all size-correct tests have no power above the significance level. To address this, we increase identification by scaling $\Pi_W$ such that $\sqrt{n} \| \Pi_W \| = 10$. We then plot the rejection frequencies of weak-instrument-robust tests in (ref).

figure[figure omitted — 519 chars of source]

The subvector Anderson-Rubin test with guggenberger2019more's guggenberger2019more critical values is indistinguishable from the test using $\chi^2(k-{{m_w}})$ critical values. Both are size-correct but have substantially less power than the (also size-correct) subvector conditional likelihood-ratio and subvector Lagrange multiplier tests. The power curves of the subvector Lagrange multiplier and the subvector conditional likelihood-ratio tests overlap, except for a small region around $\beta = 0.92$. Here, the subvector Lagrange multiplier incurs a loss in power due to the test being a {quadratic} function of the {Anderson-Rubin test statistic's derivative}.

{In (ref) in Appendix, we show power curves for the same data-generating process but with correlations $\rho := \langle \Pi_W, \Pi_X \rangle / ( \| \Pi_W \| \| \Pi_X \| ) = 0.8$ and $0.99$. As $\rho$ grows towards 1, identification becomes weaker and the power of all tests decreases. Additionally, the subvector Lagrange multiplier test's power loss around $\beta = 0.92$ becomes more pronounced. This power loss occurs because the derivative of the Anderson-Rubin test statistic is often close to zero around $\beta = 0.92$ when $\rho$ is close to 1, see also the discussion in (ref). }

In (ref), we compare the power of subvector tests for the causal effect of an included exogenous regressor. We use the same data-generating process as in (ref), but test for the causal effect $\delta$ for the first components of the $k=10$ instruments $Z$. That is, we set ${{m_x}} = 0, W \leftarrow

pmatrix[pmatrix omitted — 20 chars of source]

$, $D = Z_1$, and $Z =

pmatrix[pmatrix omitted — 42 chars of source]

$ (see also Appendix \ref{sec:exogenous_variables}). In our variation of \citeauthor{guggenberger2012asymptotic}'s (\citeyear{guggenberger2012asymptotic}) data-generating process, the instruments are valid and thus the true causal effect $\delta_0 = 0$. {We implement the subvector CLR as suggested by \citet{kleibergen2021efficient} by including $D$ as both an endogenous regressor and an instrument.} See also the discussion at the end of (ref).

figure[figure omitted — 680 chars of source]

The power of the subvector Anderson-Rubin test with guggenberger2019more's guggenberger2019more critical values is indistinguishable from the test using $\chi^2(8)$ critical values. The power of the subvector conditional likelihood-ratio and subvector Lagrange multiplier tests are also very similar, with the subvector Lagrange multiplier test being minimally more powerful for $|\delta| \leq 0.15$.

Empirical sizes of subvector tests for kleibergen2021efficient's (kleibergen2021efficient) data-generating process

Let $\Omega_{VV \cdot \varepsilon} := \operatorname{Cov}(

pmatrix[pmatrix omitted — 34 chars of source]

) - \operatorname{Cov}(

pmatrix[pmatrix omitted — 34 chars of source]

, \varepsilon_i ) \operatorname{Var}(\varepsilon_i)^{-1} \operatorname{Cov}( \varepsilon_i,

pmatrix[pmatrix omitted — 34 chars of source]

)$ be the covariance of $

pmatrix[pmatrix omitted — 34 chars of source]

$ conditional on $\varepsilon_i$. \citet{kleibergen2021efficient} prove that under \Cref{ass:1}, the asymptotic distribution of the subvector (conditional) likelihood-ratio test statistic is fully captured by $$ R \Lambda^T \Lambda R^T := n \Omega_{VV \cdot \varepsilon}^{-1/2} \Pi^T Q \Pi \Omega_{VV \cdot \varepsilon}^{-1/2}, $$ where $R$ is orthogonal and $\Lambda$ is diagonal. They parametrize $$ \Lambda =

pmatrix[pmatrix omitted — 59 chars of source]

for 0 \leq \lambda_1, \lambda_2 \leq 100 and R =

pmatrix[pmatrix omitted — 66 chars of source]

for 0 \leq \tau < \pi. $$ We use the same parametrization (see Appendix \ref{sec:simulation_details} for details) and present the empirical sizes of the different tests for $k=100$ and $\Omega = \operatorname{Cov}(

pmatrix[pmatrix omitted — 38 chars of source]

) = \mathrm{Id}_3$ in \Cref{fig:kleibergen19_identity_k100}. Unlike the subvector Anderson-Rubin and (conditional) likelihood-ratio statistics, the distribution of the subvector Lagrange multiplier statistic is not fully captured by $R \Lambda^T \Lambda R^T$. We include additional plots for $\Omega=\mathrm{Id}_3$ and $k=5, 20$ and $\Omega$ as in \citet{guggenberger2012asymptotic} and $k=5, 20, 100$ in \Cref{fig:kleibergen19_guggenberger12_k5_20,fig:kleibergen19_identity_k5_20,fig:kleibergen19_guggenberger12_k100} in Appendix \ref{sec:additional_figures}. \citeauthor{guggenberger2012asymptotic}'s (\citeyear{guggenberger2012asymptotic}) data-generating process corresponds to $\lambda_1 \approx 1, \lambda_2 \approx 120'000$, and $\tau \approx 0$. In all settings, the subvector Lagrange multiplier test we propose is size-correct.

figure[figure omitted — 262 chars of source]

Applications

We apply the Wald, subvector Anderson-Rubin, subvector conditional likelihood-ratio, and our proposed subvector Lagrange multiplier test to two real-world datasets. All estimates, test statistics, and $p$-values were generated using our ivmodels software package for Python. We provide code and instructions to reproduce results at \href{https://github.com/mlondschien/ivmodels-simulations}{github.com/mlondschien/ivmodels-simulations}.

card1993using card1993using: Using geographic variation in college proximity to estimate the return to schooling

card1993using card1993using estimates the causal effect of length of education on hourly wages. Their dataset is based on a 1976 subsample of the National Longitudinal Survey of Young Men and contains 3010 observations. It includes years of education obtained by 1976, hourly wages in 1976, age in 1976, and indicators of whether the individual lived close to a public four-year college, a private four-year college, or a two-year college in 1966. The dataset also contains variables or indicators on race, metropolitan area, region, and family background. card1993using defines (potential) experience as age - education - 6.

Like card1993using and kleibergen2021efficient, we set log-wages as the outcome, include education, experience, and experience squared as endogenous regressors, and race, metropolitan area, region, and family background indicators or variables as included exogenous regressors.

Inference for the causal effect of education of log-wages

We compare three specifications, varying by the number of instruments: (i) Using proximity to a four-year college (public or private), age, and age squared as instruments ($k=3$). This is similar to column 6 B of Table 3 of card1993using. (ii) Using the three college proximity indicators, age, and age squared as instruments ($k=5$). This is equal to kleibergen2021efficient's kleibergen2021efficient specification. (iii) Using age, age squared, and the interaction of the college proximity indicators with an indicator of whether both parents obtained less than 12 years of education as instruments ($k=8$). card1993using argue that this is a valid set of instruments, conditionally on family background variables, and report results using these instruments in column 3 of their Table 5. We present estimates, test statistics, $p$-values, and confidence sets for the causal effect of education on log-wages in (ref) (top).

table[table omitted — 5,156 chars of source]

For all specifications identification is weak, but not too weak to prohibit {informative} inference. In particular, for all specifications and tests, the causal effect of education on log-wages is significant at level $\alpha = 0.01$. Additional instruments improve identification, as measured by anderson1951estimating' anderson1951estimating likelihood-ratio test statistic of reduced rank, which increases from specification (i) to (ii) to (iii). However, as the critical value of the Anderson-Rubin and conditional likelihood-ratio tests increase with the number of instruments, the smallest $p$-values and confidence sets are achieved at specification (ii). This is in contrast to the Lagrange multiplier test, whose $p$-value is smallest for specification (iii). Still, the $p$-values of the conditional likelihood-ratio tests are smaller than those of the Lagrange multiplier test. Also, the confidence sets obtained by inversion of the Lagrange multiplier test are the only ones that are not intervals and contain negative numbers. This is due to the effect of the score being zero both at the minimum and maximum of the likelihood. {The existence of two disjoint components in the confidence sets, given education as the single weakly identified endogenous variable, is consistent with the discussion in (ref).} See (ref) in Appendix (ref) for the $p$-values of the tests for varying values of $\beta$.

Inference of the causal effect of education of log-wages for non-blacks

A natural extension of the above specifications investigates how the causal effect of education on log-wages differs by race. For this, we keep specifications (i), (ii), and (iii), but add the interaction between education and a race indicator to the set of endogenous variables and interact the instruments other than age and age squared with race. We present estimates, test statistics, $p$-values, and confidence sets for the causal effect of education on log-wages for non-blacks in (ref) (bottom).

Again, identification is weak, but not too weak to prohibit {informative} inference. anderson1951estimating's anderson1951estimating likelihood-ratio test statistic of reduced rank increases from specification (i) to (ii) to (iii). Again, the smallest $p$-values and confidence sets for the Anderson-Rubin and conditional likelihood-ratio test are achieved at specification (ii), whereas the Lagrange multiplier test achieves the smallest $p$-value at specification (iii) with $k=14$. The $p$-value of the Lagrange multiplier test for specification (iii) is the smallest among all specifications and weak-instrument-robust tests. Still, the confidence sets obtained by inversion of the Lagrange multiplier test are disconnected and include negative numbers. If one discards the negative part of the confidence set, for example as such values are unlikely according to economic theory, the remaining confidence interval is smallest among those obtained by inverting weak-instrument-robust tests.

The causal effect of education on wages for blacks was not significant at level $\alpha=0.2$ for any of the tests of interest. We present estimates, test statistics, $p$-values, and confidence sets for the causal effect of education on log-wages for blacks in (ref) in Appendix (ref).

tanaka2010risk tanaka2010risk: Risk and time preferences: Linking experimental and household survey data from vietnam

tanaka2010risk study causes of risk preferences in Vietnam. Individuals from 25 households were interviewed for each of 289 villages. tanaka2010risk work with a subsample of in total 181 households from 9 villages. From the interviews, they estimate measures of risk preferences, including the curvature of the utility function. This is the dependent variable in the specifications of our interest. Also measured is household income, which tanaka2010risk split into the village means and the household's income relative to the village mean. These regressors are assumed to be endogenous. Included exogenous regressors are gender, age, education, distance to market, and whether the individual is Chinese or not. Finally, tanaka2010risk use rainfall in the village and an indicator of whether the head of the household can work as instruments for the household income variables.

We consider two specifications: (i) tanaka2010risk's tanaka2010risk original specification as described above (Table 4 column 2 and Table 5 column 2 bottom in their paper) and (ii) a specification where we include the interaction of the head of the household being able to work with rainfall as an additional instrument. tanaka2010risk discuss the estimated conditional causal effect for village mean income, income relative to the village mean, and gender. We report estimates, confidence sets, and $p$-values for the causal effect of these variables in (ref). To make inference on gender, assumed to be exogenous, with the Anderson-Rubin and conditional likelihood-ratio test, we include gender as both an endogenous variable and an instrument.

table[table omitted — 4,183 chars of source]

In specification (i), where the number of instruments is equal to the number of endogenous regressors, the LIML estimate is equal to the TSLS estimate and the Lagrange multiplier and conditional likelihood-ratio test are equal to the Anderson-Rubin test. Adding the interaction of rainfall and the head of the household being able to work as an instrument does not substantially improve identification: anderson1951estimating's anderson1951estimating test statistic of reduced rank increases from 6.051 to 6.071, but the $p$-value and the concentration parameter (which depend on the number of instruments) decrease.

tanaka2010risk note that using the Wald test, the conditional causal effect of mean village income is significant at the 10% level and that the conditional causal effects of income relative to the village mean and gender are not significant. These findings still hold in specification (i) if using weak-instrument-robust tests. After adding the interaction of rainfall and the head of the household being able to work as an additional instrument (specification ii), the $p$-value from the Lagrange multiplier test for the conditional causal effect of village mean income decreases from 0.061 to 0.060, whereas the $p$-values of the Anderson-Rubin and conditional likelihood-ratio tests increases from 0.061 to 0.150 and 0.105. That is, the Lagrange multiplier test is the only weak-instrument-robust test that rejects the null of no conditional causal effect of village mean income on risk preferences at level 10% for specification (ii).

We also report 90% confidence sets. The subvector inverse Lagrange multiplier test confidence set for the causal effect of the village mean income is the only weak-instrument-robust confidence set that does not include zero for specification (ii). {It is composed of three disjoint components, consistent with two weakly identified endogenous variables and the discussion in (ref).} See (ref) in Appendix (ref) for the $p$-values of the tests for varying values of $\beta$.

Conclusion

We introduced a novel weak-instrument-robust subvector Lagrange multiplier test for instrumental variables regression. This test recovers the degrees of freedom of the standard Wald test, a property not achieved by previous weak-instrument-robust subvector tests. We show that this test is asymptotically size-correct, either under a technical condition or as the number of instruments grows to infinity. Numerical simulations confirm that the subvector Lagrange multiplier test maintains correct size and exhibits substantial power, comparable to the subvector conditional likelihood-ratio test. Its ability to handle a large number of instruments without an increase in critical values makes it particularly suitable for modern applications where non-linear relationships are modeled through basis expansions.

In addition to the new test, we provide an analysis of confidence sets obtained by inverting the subvector Anderson-Rubin test. We offer a closed-form solution for these confidence sets, revealing that they are centered around a k-class estimator. A significant direct consequence is that for single coefficients, these subvector confidence sets are jointly bounded if and only if anderson1951estimating's anderson1951estimating likelihood-ratio test rejects the null hypothesis of underidentification. We also establish a direct relationship between bounded inverse Anderson-Rubin test confidence sets and Wald-based confidence sets.

We apply various subvector tests to data from card1993using and tanaka2010risk. There, the subvector Lagrange multiplier test can yield smaller weak-instrument-robust $p$-values when additional instruments through interactions are included.

Future work could focus on proving that the technical condition required for the general size-correctness of the subvector Lagrange multiplier test holds with high probability, as our empirical results suggest. Overall, the subvector Lagrange multiplier test and the properties of subvector Anderson-Rubin confidence sets presented in this paper fill a gap in the literature and provide practitioners with additional powerful and reliable methods for subvector inference in the presence of weak instruments.

\FloatBarrier

Acknowledgements

Malte Londschien is supported by the ETH Foundations of Data Science. We would like to thank Christoph Schultheiss, Cyrill Scheidegger, Fabio Sigrist, Felix Kuchelmeister, Frank Kleibergen, Gianna Wolfisberg, Jonas Peters, Juan Gamella, Leonard Henckel, Markus Ulmer, Maybritt Schillinger, Michael Law, Yuansi Chen, and Zijian Guo for helpful discussions and comments. We also thank the co-editor, associate editor, and two anonymous reviewers for their constructive comments.