EconBase
← Back to paper

Resistant Inference in Instrumental Variable Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

76,873 characters · 17 sections · 104 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Resistant Inference in Instrumental Variable Models

{1.8\baselineskip}

abstractThe classical tests in the instrumental variable model can behave arbitrarily if the data is contaminated. For instance, one outlying observation can be enough to change the outcome of a test. We develop a framework to construct testing procedures that are robust to weak instruments, outliers and heavy-tailed errors in the instrumental variable model. The framework is constructed upon M-estimators. By deriving the influence functions of the classical weak instrument robust tests, such as the Anderson-Rubin test, K-test and the conditional likelihood ratio (CLR) test, we prove their unbounded sensitivity to infinitesimal contamination. Therefore, we construct contamination resistant/robust alternatives. In particular, we show how to construct a robust CLR statistic based on Mallows type M-estimators and show that its asymptotic distribution is the same as that of the (classical) CLR statistic. The theoretical results are corroborated by a simulation study. Finally, we revisit three empirical studies affected by outliers and demonstrate how the new robust tests can be used in practice.

{\it Keywords: } Influence function; Local misspecification; Robust inference; Outlier; Robust test; Weak instrument.

Introduction

The linear instrumental variable (IV) model is recognized as an important tool that can be used to draw causal inferences in nonexperimental data AngristImbensRubin1996. Its applicability spans over different fields and allows, for example, to infer the causal effects of education on labor market earnings angrist1991does, chemotherapy for advanced lung cancer in the elderly earle2001effectiveness and foreign media on authoritarian regimes kern2009opium. In each of these examples, an endogeneity problem occurs where the explanatory variable is correlated with the error term, causing the Least Squares (LS) estimator to be biased. The researchers introduce instrumental variables to resolve this problem. The instrumental variables should be correlated with the (endogenous) explanatory variable and uncorrelated with the error term. When instrumental variables are available, then reliable estimation (and inference) is possible using IV estimators, such as the Two-Stage Least Squares (2SLS) estimator.

In practice, it is difficult to find instruments that satisfy both conditions. In particular, the instruments that researchers propose are oftentimes only weakly correlated with the endogenous explanatory variable andrews2019weak. When the instruments are weak, then classical IV estimators, such as the 2SLS estimator, are biased and $t$-tests based on IV estimators can fail to control the size of the tests and their associated confidence intervals are invalid Nelson1990robust,BoundJaegerBaker1995.

commentIn particular, it is typically advised to use the CLR test to construct confidence sets in the (homoskedastic) linear IV model as it has good power properties irrespective of the strength of the instruments andrews2006optimal, andrews2019weak.

Therefore, it is common to draw inference in the IV model using a two-step procedure. In the first step, the strength of the instruments is tested by means of an $F$-test. When the first-stage $F$ statistic is above a certain threshold (motivated by staiger1997instrumental a cutoff value of 10 is common), the instruments are considered to be strong. When the instruments are strong, then estimation is done with a 2SLS estimator and inference with a $t$-test. When the instruments are weak, then weak instrument robust tests are used anderson1949estimation, kleibergen2002pivotal, moreira2003conditional.

The implementation of the two-step procedure described above by practitioners has typically been imperfect and it is advised to directly rely on weak instrument robust tests, or, smoothly adjust 2SLS $t$-test inference based on the first-stage $F$ andrews2019weak, lee2022valid, keane2023instrument. In particular, in the just-identified model, i.e., with one endogenous variable and one instrumental variable, the anderson1949estimation (AR) test is the uniformly most powerful among the class of unbiased tests moreira2009tests. When there is one endogenous variable and multiple instrumental variables, then the conditional likelihood ratio (CLR) test introduced by moreira2003conditional enjoys good power properties in the linear homoskedastic IV model andrews2006optimal. As the CLR test is equal to the AR test in the just-identified model it is generally advised to use the CLR test.

Recently, young2022consistency pointed out that many empirical IV studies in economics are affected by a few outliers or by small clusters of deviating observations (see also lal2023how in political science). In the two-step procedure, the $F$-test is not robust against outliers ronchetti1982robust. Hence, even one outlier is enough to inflate the first-stage $F$ leading to a false impression that the instrument is strong, while it is weak klooster2021outlier, which eventually results in incorrect inference in the second stage. On the other hand, an outlier could also deflate the $F$-statistic, i.e., the instrument is strong, but due to an outlier it seems weak. In that case, the researcher might unnecessarily dismiss the instrument.

A possible solution could be to clean the data using outlier detection tools and then apply the (classical) procedure. However, this data-analytic strategy has been criticized in the literature, e.g., welsh2002journey, maronna2019robust, heritier2009robust. If the data is cleaned, then one has to correct the inference for the data cleaning step, otherwise the classical asymptotic theory becomes invalid chen2020valid. A typical consequence is underestimation of the (asymptotic) variance which leads to overrejection of the null hypothesis, see klooster2021outlier for an example in the IV model. Moreover, in modern applications involving complex structural models outlier detection and its justification are themselves not trivial. Hence, a more viable strategy is to directly use (outlier) robust methods for estimation and inference. In the context of IV, several studies showed that classical IV estimators, such as the 2SLS and Limited Information Maximum Likelihood (LIML) estimators, are not robust to outliers and, consequently, introduced robust IV estimators freue2013natural, jiao2022robust, solvsten2020robust, ZGR2012. klooster2021outlier developed a robust version of the AR test. However, a general treatment of the problem of (outlier) robust inference in combination with weak instruments is missing.

In this article, we formally show that weak instrument robust tests such as the AR, K kleibergen2002pivotal and CLR tests are not robust to outliers. For this reason, we develop a general framework that can be used to construct robust versions of the AR, K and CLR tests. In particular, we show how to construct an outlier robust CLR test that allows for reliable inference (without having to rely on a first-stage pre-test) when the data might contain some outliers. The robust CLR test provides reliable inference irrespective of the strength of the instrumental variables and only loses a little bit of power compared to the (classical) CLR test when the data does not contain any outliers. Moreover, if the errors follow a heavy-tailed distribution, then our tests are more powerful than the classical ones.

We formalize the outlier contamination using the huber1964 gross-error model $F_{t} = (1-t)F + t G$, where $t$ is a (typically small) contamination proportion. We are interested in drawing inference for the central model $F$, but we assume that it only holds approximately and the data that we observe comes from $F_{t}$. The model $G$ is assumed to be completely unknown and it is the source of the contamination. The goal is to draw inference that is valid at the central model $F$, but also remains stable and reliable when the data is generated according to $F_{t}$. The (classical) AR, K and CLR tests are valid when all the data is generated according to $F$, but, as we show both analytically and numerically, when the data that we observe comes from $F_{t}$, then they can be distorted. This setup is related to the literature on sensitivity analysis and local misspecification, see AndrewsGentzkowShapiro2017, AndrewsGentzkowShapiro2020, AtheyImbens2015, BonhommeWeidner2021, kitamuraOtsuEvdokimov2013 and references therein. The model $F_t$ can be viewed as locally misspecified. If one assumes that the model $F$ holds exactly, then a sensitivity analysis should complement the reported results. In Section (ref), we revisit studies by alesina2011segregation, ananat2011wrong and angrist1991does and show how our tests can be used to study the sensitivity of the results to data contamination. Our approach allows to benefit from the parametric structure, e.g.,\ computational simplicity and interpretability, while being resistant to small but harmful deviations from the assumed model $F$.

When the assumed model $F$ does not even hold approximately, then the use of nonparametric weak instrument robust inference andrews2007rank, andrews2008exactly is advised. Note, however, that nonparametric procedures are typically not designed to be robust to outliers. For example, the sample mean is a nonparametric estimator of the expectation, but it is not robust as one outlier can make it arbitrarily biased (see huber2009 for further discussion). Similarly, weak instrument robust quantile methods introduced by chernozhukov2008instrumental and jun2008weak are also not designed to be robust against outliers.

The setup of the article is as follows. In Section (ref), we introduce the model and our notation. In Section (ref), we introduce the general (robust) framework that allows construction of weak instrument testing procedures that are robust to outliers. In particular, in Section (ref) we show how to construct a robust version of the CLR test that can be readily used in practice. Then, in Section (ref), we study the small sample properties of the robust CLR test and compare its performance to the classical CLR test. In Section (ref), we revisit three empirical studies and show how the robust CLR test can be used in practice. Finally, in Section (ref) we conclude.

Instrumental Variables Model

We assume that the data is generated according to a linear instrumental variable regression model. The model consists of a structural equation ((ref)) and a first-stage equation ((ref)):

align[align omitted — 154 chars of source]

where $y$ and $x$ are endogenous random variables, $z$ is a $k \times 1$ random vector of instrumental variables and $w$ is a $p \times 1$ random vector of control variables. We assume that $\gamma_1, \gamma_2 \in \mathbb{R}^p$ are parameter vectors that both, if necessary, include an intercept, $\pi \in \mathbb{R}^k$ and $\beta \in \mathbb{R}$. We assume that the errors $(u, v)$ are mean zero, with covariance matrix $\Sigma =

pmatrix[pmatrix omitted — 74 chars of source]

$. We assume that the instruments are uncorrelated with the error terms, conditional on the control variables $w$.

We are interested in testing the hypothesis

align[align omitted — 118 chars of source]

in the linear IV model ((ref)) - ((ref)). We do so by constructing weak instrument robust tests based on estimators in the reduced form model. The reduced form equation can be obtained by substituting ((ref)) into ((ref)):

align[align omitted — 93 chars of source]

with $\gamma = (\gamma_1 +\gamma_2 \beta)$, $\delta = \pi \beta$ and $\epsilon = v\beta + u$. We refer to ((ref)) - ((ref)) as the reduced form model.

The tests we introduce in Section (ref) are constructed upon estimators of the parameters $\delta$ and $\pi$ in the reduced form model ((ref)) - ((ref)). To simplify the presentation of our results, we assume that $\gamma_1 = \gamma_2 = 0$ so that $w$ drops out of the model. The results we present extend to the more general case where $\gamma_1 \neq 0$ and $\gamma_2 \neq 0$. In particular, when the estimators are regression equivariant (see rousseeuw1987robust), then assuming $\gamma_1 = \gamma_2 = 0$ is without loss of generality. Thus, we focus on the simplified reduced form model

align[align omitted — 159 chars of source]

Define $\theta = ( \delta^{\top} , \pi^{\top} )^{\top}$, then we assume that the model ((ref)) - ((ref)) is governed by $F_{\theta}$. In practice, we assume that we observe i.i.d.\ data $d_i = (y_i , x_i , z_i^{\top} )^{\top}$, which are random samples from $d = (y , x , z^{\top} )^{\top} \sim F_{\theta}$. We define $Y,X \in \mathbb{R}^n$, which are vectors with entries $y_i$ and $x_i$ respectively, and $Z \in \mathbb{R}^{n \times k}$, the matrix with rows $z_i^{\top}$. Denote the partition of a vector $a \in \mathbb{R}^{2k}$ into $k$ and $k$ components by $a^{\top} = ( a_{(1)}^{\top} , a_{(2)}^{\top} )$ and denote the corresponding partition of $2k \times 2k$ matrices by $A_{(ij)}$, $i,j = 1,2$.

Robust Inference

In this section, we introduce a general (robust) framework to construct weak instrument robust testing procedures based on M-estimators huber1964. We show that classical weak instrument robust tests, such as the AR, K and CLR tests can be obtained by specifying the M-estimators to be the LS estimators. To evaluate whether a statistic is robust, we use the influence function hampel1974influence, hampel1986robust, which we discuss in the next section.

Influence Function

To formally determine whether a statistic is robust, we analyze its influence function. Let $T$ denote a (Fisher consistent) statistical functional, then the influence function is defined as

align[align omitted — 139 chars of source]

where $\Delta_d$ is a point mass at $d$. The heuristic interpretation of the influence function is that it describes the effect of an infinitesimal contamination at the point $d$ on the estimate, standardized by the mass of the contamination hampel1986robust. If the statistical functional $T$ is sufficiently regular, then a von Mises expansion mises1947asymptotic, hampel1974influence yields

align[align omitted — 104 chars of source]

When we consider the approximation ((ref)) over a neighborhood of the model $\mathcal{F}_t = \{F_t = (1-t)F + tG \ | \ G \text{ an arbitrary distribution} \}$, we see that the influence function can be used to linearize the asymptotic bias in a neighborhood of the ideal model $F$. In particular, when we consider the maximum asymptotic bias, we obtain the following relationship

align[align omitted — 99 chars of source]

Hence, when the influence function of a statistical functional $T$ is not bounded, then the maximum bias in a neighborhood of $F$ can be infinite, even when the contamination proportion $t$ is small. Therefore, for a statistical functional to be (locally) robust, a bounded influence function is required.

commentwhere we used $F_t = (1-t)F + tG$ and $\int \operatorname*{IF}(d; T, F) dF = 0$. Taking the supremum on both sides yields \begin{align} \sup_{G} || T(F_t) - T(F) || \approx t \sup_{d} || \operatorname*{IF}(d; T, F) ||. \end{align} Hence, when the influence function of a statistical functional $T$ is not bounded, then the maximum bias in a neighborhood of $F$ can be infinite, even when the contamination proportion $t$ is small. Therefore, for a statistical functional to be (locally) robust, a bounded influence function is required.

M-estimators

We construct test statistics based on the M-estimator $T(F_n) = \left\{ \delta(F_n)^{\top} , \pi(F_n)^{\top} \right\}^{\top}$ of $\theta$ defined by

equation[equation omitted — 290 chars of source]

where $\psi$ is a (general) score function and $F_n$ denotes the empirical distribution function. Note, letting $\psi\{(a,b),c\} = (a - b^{\top}c)b$, results in the classical LS estimating equations, which are worked out in detail in Example (ref). Under general regularity conditions stated in Assumption (ref) in the Appendix, if $(y_i , x_i , z_i^{\top} )^{\top}$ are i.i.d.\ draws from the distribution $F_{\theta}$, then M-estimators are consistent and asymptotically normally distributed clarke1983uniqueness, clarke1986nonsmooth, i.e., as $n \rightarrow \infty$, then

align*[align* omitted — 468 chars of source]

where

align*[align* omitted — 767 chars of source]

The influence functions of these estimators can be calculated using the definition ((ref)) (see hampel1986robust), and are given by

align[align omitted — 492 chars of source]

with $M(F_\theta) = -\int \frac{\partial \Psi}{\partial \theta}(d, \theta) d F_{\theta}$.

comment\begin{align} \operatorname*{IF}\left\{d_i; \delta(\cdot), F_{\theta}\right\} &:= \left[-\int \frac{\partial \psi}{\partial \delta}\{(y,z), \delta\} d F_{\theta}\right]^{-1} \psi\left\{(y_i,z_i), \delta\right\}, \\ \operatorname*{IF}\left\{(x_i,z_i); \pi(\cdot), F_{\theta}\right\} &= \left[-\int \frac{\partial \psi}{\partial \pi}\{(x,z), \pi\} d F_{\theta}\right]^{-1} \psi\left\{(x_i,z_i), \pi\right\}. \end{align}

When we analyze the influence functions ((ref)) - ((ref)) we see that they are only bounded when the function $\psi$ is bounded. Thus, the estimators $\delta(F_n)$ and $\pi(F_n)$ are (locally) robust when $\psi$ is bounded.

As the LS estimators play an important role for the construction of the AR, K and CLR statistics, we explicitly show how can they be considered as a special case of the M-estimators in Example (ref).

example[Least Squares] When we use the following score function: \begin{align*} \psi \colon \left(\mathbb{R} \times \mathbb{R}^k \right) \times \mathbb{R}^k \mapsto \mathbb{R}^{k} \colon \psi\{(a,b),c\} = (a - b^{\top}c)b, \end{align*} then $\delta(F_n) = (Z^{\top}Z)^{-1}Z^{\top}Y$ and $\pi(F_n)= (Z^{\top}Z)^{-1}Z^{\top}X$ are the LS estimators that solve estimating equations \begin{align*} \frac{1}{n} \sum_{i=1}^n \begin{bmatrix} \left\{y_i - z_i^{\top}\delta(F_n)\right\}z_i \\ \left\{x_i - z_i^{\top}\pi(F_n)\right\}z_i \end{bmatrix} = 0. \end{align*} Their influence functions are \begin{align} \operatorname*{IF}\left\{d_i; \delta(\cdot), F_{\theta}\right\} &= \left(\int z z^{\top} dF_{\theta} \right)^{-1}\left(y_i - z_i^{\top}\delta \right)z_i, \\ \operatorname*{IF}\left\{d_i; \pi(\cdot), F_{\theta}\right\} &= \left(\int z z^{\top} dF_{\theta} \right)^{-1}\left(x_i - z_i^{\top}\pi \right)z_i. \end{align} The influence functions of the LS estimators are not bounded, so that one outlying observation $d_i$ can arbitrarily bias the estimators $\delta(F_n)$ and $\pi(F_n)$. Using the influence function, we can now calculate the (co)variance matrices. For instance, assuming homoskedasticity, we have \begin{align*} \Sigma_{\delta \delta}(F_{\theta}) &= \int \operatorname*{IF}\left\{d; \delta(\cdot), F_{\theta}\right\} \operatorname*{IF}\left\{d; \delta(\cdot), F_{\theta}\right\}^{\top} dF_{\theta}\\ &= \int \left(\int z z^{\top} dF_{\theta} \right)^{-1}\left(y - z^{\top}\delta \right)z z^{\top} \left(y - z^{\top}\delta \right)^{\top} \left(\int z z^{\top} dF_{\theta} \right)^{-\top} d F_{\theta} \\ &= \left(\int z z^{\top} dF_{\theta} \right)^{-1} \int \left(y - z^{\top}\delta \right)^2 dF_{\theta}, \end{align*} which can be estimated by \begin{equation*} {\Sigma}_{\delta \delta}(F_n) = \left(\int z z^{\top} dF_{n} \right)^{-1} \int \left\{y - z^{\top}\delta(F_n) \right\}^2 dF_{n} = \left( \frac{1}{n}Z^{\top} Z \right)^{-1}\sigma_{\epsilon}^2(F_n), \end{equation*} where $\sigma_{\epsilon}^2(F_n) = \frac{1}{n}\sum_{i=1}^n\left\{y_i - z_i^{\top}\delta(F_n) \right\}^2$.

The (Robust) Conditional Likelihood Ratio Test

moreira2003conditional shows that in the case of a scalar $\beta$, then the classical CLR statistic can be decomposed into the classical AR, K, and a Wald statistic. Therefore, to introduce a general (robust) version of the CLR statistic, we start by introducing general (robust) versions of the AR, K and Wald statistics. Then, we plug in the general (robust) AR, K and Wald statistics into the decomposition to introduce a general (robust) CLR statistic.

Let $\left\{ \delta(F_n)^{\top} , \pi(F_n)^{\top} \right\}^{\top}$ denote M-estimators of the parameters $( \delta^{\top} , \pi^{\top})^{\top}$ as defined in ((ref)). Then, given a hypothesized $\beta_0$ value, we define (see also andrews2019weak),

align*[align* omitted — 237 chars of source]

where $\Omega(F_{\theta},\beta_0) = \Sigma_{\delta \delta}(F_{\theta}) - \beta_0\left\{\Sigma_{\delta \pi}(F_{\theta}) + \Sigma_{\pi \delta}(F_{\theta})\right\} + \beta_0^2 \Sigma_{\pi \pi}(F_{\theta})$. The general (robust) AR, K and Wald statistics are defined as:

align[align omitted — 449 chars of source]

where $\Lambda(F_{\theta},\beta_0) = \Sigma_{\pi \pi}(F_{\theta}) - \left\{\Sigma_{\pi \delta}(F_{\theta}) - \Sigma_{\pi \pi}(F_{\theta}) \beta_0\right\}\Omega(F_{\theta},\beta_0)^{-1}\left\{\Sigma_{\delta \pi}(F_{\theta}) - \Sigma_{\pi \pi}(F_{\theta}) \beta_0\right\}.$ The general (robust) CLR statistic can then be written as:

equation[equation omitted — 276 chars of source]

In Example (ref), we show how plugging in the LS estimators from Example (ref) results in the classical AR, K, Wald and CLR statistics.

example[Least Squares (continued)] When $\delta(F_n)$ and $\pi(F_n)$ are the LS estimators, and using the estimators $\Omega(F_n, \beta_0)$ and $\Lambda(F_n, \beta_0)$, then we have \begin{align*} g(F_n, \beta_0) &= (Z^{\top} Z)^{-1}Z^{\top}(Y - \beta_0 X), \\ D(F_n, \beta_0) &= (Z^{\top}Z)^{-1}Z^{\top}\left\{X - \frac{\sigma_{u,v}(F_n)}{\sigma_u^2(F_n)}\left(Y - X\beta_0\right)\right\}. \end{align*} Plugging these into the statistics ((ref)), ((ref)) and ((ref)) yields the classical AR, K and Wald statistics (up to a factor $n$): \begin{align*} n AR(F_n, \beta_0) &= \frac{(Y - \beta_0 X)^{\top} Z (Z^{\top}Z)^{-1} Z^{\top}(Y - \beta_0 X)}{\sigma_u^2(F_n)},\\ nK(F_n, \beta_0) &= \frac{1}{\sigma_u^2(F_n)}(Y-X\beta_0)^{\top} P_{ZD(F_n, \beta_0)}(Y-X\beta_0),\\ nW(F_n, \beta_0) &= \left\{X - \frac{\sigma_{u,v}(F_n)}{\sigma_u^2(F_n)}\left(Y - X\beta_0\right) \right\}^{\top}\left[\left\{\sigma_v^2(F_n) -\frac{\sigma_{u,v}(F_n)}{\sigma_u^2(F_n)}\right\}Z^{\top}Z\right]^{-1}\times \\ & \qquad \left\{X - \frac{\sigma_{u,v}(F_n)}{\sigma_u^2(F_n)}\left(Y - X\beta_0\right)\right\}, \end{align*} with $P_{ZD(F_n, \beta_0)} = ZD(F_n, \beta_0)\left\{D(F_n, \beta_0)^{\top}Z^{\top}ZD(F_n, \beta_0)\right\}^{-1}D(F_n, \beta_0)^{\top}Z^{\top}$ and $\sigma_u^2(F_n)= \sigma_{\epsilon}^2(F_n) - 2\beta_0 \sigma_{\epsilon, v}(F_n) + \beta_0^2 \sigma_v^2(F_n)$. Plugging these statistics into the decomposition ((ref)) yields the CLR statistic introduced by moreira2003conditional.

{\it Remark} \ \ We are not the first to use the decomposition ((ref)) to construct a generalization of the CLR statistic. For example, kleibergen2005testing uses the decomposition to introduce a generalized method of moments (GMM) version of the CLR statistic based on GMM versions of the AR stock2000gmm and K statistic. magnusson2010inference uses the decomposition to introduce a CLR test for inference in weakly identified limited dependent variable models. Their goal is to provide tests that can be used for inference in weakly identified models beyond the linear IV model. In this article, we restrict our focus to the linear IV model and leave connections to more general versions open for future research. Note, ronchetti2001robust study the robustness properties of GMM estimators, but they do not consider weak identification.

Robustness Properties

To determine the robustness properties of the general (robust) CLR statistic we derive its influence function. It turns out that directly applying the influence function defined in ((ref)) to the statistical functional $RCLR(\cdot, \beta_0)$ is not meaningful as the influence function is always zero. This happens, because the statistical functional $RCLR(\cdot, \beta_0)$ is not Fisher consistent. A functional $T$ is Fisher consistent when $T(F_{\theta}) = \theta$ for all $\theta$, while $RCLR(F_{\theta}, \beta) = 0$. Therefore, to still obtain a meaningful result, similar as in hampel1986robust, we derive the influence function of the square root of the statistic. Note, generalizations of the influence function towards functionals that are not Fisher consistent have been introduced as well rousseeuw1981influence.

In Proposition (ref), we derive the influence function of the general (robust) CLR statistic, conditional on $D(F_n, \beta_0) = \tilde{D}$, denoted by $RCLR(\cdot, \beta_0 | \tilde{D})$.

propositionUnder the null hypothesis $\beta = \beta_0$, and the regularity conditions in Assumption (ref), the influence function of the general (robust) CLR statistic, conditional on $D(F_n, \beta_0) = \tilde{D}$, is \begin{align*} \operatorname*{IF}\left\{d_i; \sqrt{RCLR(\cdot, \beta_0 | \tilde{D})}, F_{\theta} \right\} = \begin{cases} \operatorname*{IF}\left\{d_i; \sqrt{RAR(\cdot, \beta_0)}, F_{\theta} \right\}, &if \tilde{D} = 0, \\ \operatorname*{IF}\left\{d_i; \sqrt{RK(\cdot, \beta_0 | \tilde{D})}, F_{\theta}\right\}, &if \tilde{D} \neq 0. \end{cases} \end{align*}
proofSee Appendix.

Proposition (ref) shows that the influence function of the RCLR statistic, conditional on $D(F_n, \beta_0) = \tilde{D}$, depends on the influence functions of the RAR statistic and RK statistic. This corresponds to the theory, as $D(F_n, \beta_0) = 0$ suggests that true value of $\pi = 0$ as $D(F_{\theta}, \beta) = \pi$. When $D(F_n, \beta_0) = 0$, it suggests that the instruments are completely irrelevant. In this case, the RCLR statistic is the same as the RAR statistic and therefore has the same influence function as the RAR statistic. When $\tilde{D} \neq 0$, it suggests that $\pi \neq 0$. moreira2003conditional shows that when $\pi \neq 0$, then the CLR statistic converges towards the $K$ statistic as the amount of data goes towards infinity. For this reason, we see the same behavior in the IF of the RCLR statistic as it is equal to the IF of the RK statistic when $\tilde{D} \neq 0$.

We continue with deriving the influence functions of the RAR, RW and RK statistics.

propositionUnder the null hypothesis $\beta = \beta_0$, and the regularity conditions in Assumption (ref), the influence functions of the general (robust) AR, W and K statistics are \begin{align} \operatorname*{IF}\{d_i; \sqrt{RAR(\cdot, \beta_0)}, F_{\theta}\} &= \sqrt{\operatorname*{IF}\{d_i; g(\cdot, \beta_0), F_{\theta}\}^{\top} \Omega(F_{\theta}, \beta_0)^{-1}\operatorname*{IF}\{d_i; g(\cdot, \beta_0), F_{\theta}\}}, \\ \operatorname*{IF}\{d_i; \sqrt{RK(\cdot, \beta_0)}, F_{\theta}\} &= \sqrt{\operatorname*{IF}\{d_i; g(\cdot, \beta_0), F_{\theta}\}^{\top}\pi \left(\pi^{\top}\Omega(F_{\theta},\beta_0)\pi\right)^{-1}\pi^{\top}\operatorname*{IF}\{d_i; g(\cdot, \beta_0), F_{\theta}\}},\\ \operatorname*{IF}\{d_i; \sqrt{RW(\cdot, \beta_0)}, F_{\theta}\} &= \sqrt{\operatorname*{IF}\{d_i; D(\cdot, \beta_0), F_{\theta}\}^{\top} \Lambda(F_{\theta},\beta_0)^{-1}\operatorname*{IF}\{d_i; D(\cdot, \beta_0), F_{\theta}\}}, \end{align} where \begin{align*} \operatorname*{IF}\{d_i; g(\cdot, \beta_0), F_{\theta}\} &= \operatorname*{IF}\{d_i; \delta(\cdot), F_{\theta}\} - \beta_0 \operatorname*{IF}\{d_i; \pi(\cdot), F_{\theta}\} \\ \operatorname*{IF}\{d_i; D(\cdot, \beta_0), F_{\theta}\} &= \operatorname*{IF}\{d_i; \pi(\cdot), F_{\theta}\} - \left\{\Sigma_{\pi \delta}(F_{\theta}) - \Sigma_{\pi \pi}(F_{\theta}) \beta_0\right\}\Omega(F_{\theta}, \beta_0)^{-1}\operatorname*{IF}\{d_i; g(\cdot, \beta_0), F_{\theta}\}. \end{align*} The influence function of the general (robust) K statistic, conditional on $D(F_n, \beta_0) = \tilde{D}$, is: \begin{align*} \operatorname*{IF}\{d_i;& \sqrt{RK(\cdot, \beta_0 | \tilde{D})} , F_{\theta} \} = \sqrt{\operatorname*{IF}\{d_i; g(\cdot, \beta_0), F_{\theta}\}^{\top}\tilde{D} \left\{\tilde{D}^{\top}\Omega(F_{\theta},\beta_0)\tilde{D}\right\}^{-1}\tilde{D}^{\top}\operatorname*{IF}\{d_i; g(\cdot, \beta_0), F_{\theta}\}}. \end{align*}
proofSee Appendix.

Proposition (ref) shows that the influence functions of the RAR and RK statistics depend directly on the influence function of the functional $g(\cdot, \beta_0)$. The influence function of the functional $g(\cdot, \beta_0)$ depends on the influence functions of the functionals $\delta(\cdot)$ and $\pi(\cdot)$.

Robustness Issues With the Conditional Likelihood Ratio Test

In this section, we study the robustness properties of the classical CLR statistic by analyzing its influence function. Example (ref) showed us how the classical CLR statistic can be viewed as a special case where LS estimators are used. Therefore, we can use Propositions (ref) and (ref) to analyze the influence function of the classical CLR statistic. As in Example (ref), we denote the classical AR, K, Wald and CLR statistics as $AR(F_n,\beta_0), K(F_n, \beta_0), W(F_n, \beta_0)$ and $CLR(F_n, \beta_0)$, respectively.

The previous section showed that the influence function of the general (robust) CLR statistic ultimately depends on the influence functions of the estimators $\delta(F_n)$ and $\pi(F_n)$ it is constructed upon. The classical CLR statistic is constructed upon LS estimators. The influence functions of the LS estimators were derived in Example (ref), see ((ref)) and ((ref)). These influence functions are not bounded, so that one outlying observation $d_i = (y_i , x_i , z_i^{\top} )^{\top}$ can arbitrarly bias the estimators $\delta(F_n)$ and $\pi(F_n)$. As the influence functions of $\delta(\cdot)$ and $\pi(\cdot)$ are not bounded, the influence function of $g(\cdot, \beta_0)$ is not bounded, so that the influence functions $\operatorname*{IF}\{d_i; \sqrt{AR(\cdot, \beta_0)}, F_{\theta}\}$ and $\operatorname*{IF}\{d_i;\sqrt{K(\cdot, \beta_0 | \tilde{D})}, F_{\theta}\}$ are not bounded, and finally the influence function $\operatorname*{IF}\{d_i;\sqrt{CLR(\cdot, \beta_0 | \tilde{D})}, F_{\theta}\}$ is not bounded. Therefore, we can conclude that the classical CLR statistic is not (locally) robust.

Intuitively, the classical CLR statistic inherits the robustness properties of the estimators it is constructed upon. As the LS estimators are not robust, the classical CLR statistic is also not robust. However, from Propositions (ref) and (ref) it follows that if the influence function of the estimators would be bounded, then the influence function of the general (robust) CLR statistic would be bounded. From Section (ref), it follows that, when the function $\psi$ is bounded the influence function of the M-estimator is bounded. Therefore, the general (robust) CLR statistic introduced in ((ref)) is robust whenever the M-estimator it is constructed upon has a bounded (general) score function $\psi$. An example of an M-estimator with a bounded influence function is the M-estimator of Mallows type. In Section (ref), we show how these types of estimators can be used to construct a robust CLR statistic that can be used in practice.

comment\subsubsection{\textcolor{red}{Optional Section: }Connection to the Level Influence Function} In this section we derive the level influence function of the CLR test and show how it relates to the results from the previous section. As in heritier1994robust, we denote the level of the test when the underlying distribution is $F$ by $\alpha(F)$, and denote the nominal level $\alpha(F_{\theta})$ by $\alpha$. We define $\mu_q = -(\partial/\partial \zeta) H_q(\eta_{1-\alpha}; \zeta)|_{\zeta = 0}$, where $H_q(\cdot ; \zeta)$ is the cumulative distribution function of a $\chi_q^2(\zeta)$ distribution, and $\eta_{1-\alpha_0}$ is the $1-\alpha_0$ quantile of the central $\chi_q^2$ distribution. Let \begin{align*} F_{\epsilon, n}^L = \left(1 - \frac{\epsilon}{\sqrt{n}} \right)F_{\theta} + \frac{\epsilon}{\sqrt{n}} \Delta_d, \end{align*} then a direct application of Proposition 5 in heritier1994robust shows that the level of the AR test under contamination is given by \begin{align*} \lim_{n \rightarrow \infty} \alpha_{AR}(F_{\epsilon, n}^L) = \alpha + \epsilon^2 \mu_k \cdot \operatorname*{IF}(d; \sqrt{AR(\cdot, \beta_0)}, F_{\theta} )^2 + o(\epsilon^2). \end{align*} Here we see how the influence function of the square root of the AR statistic also shows up when deriving the level influence function. As the influence function of the square root of the AR statistic is not bounded, the test will not be robust, as one outlying observation can have an arbitrary effect on the asymptotic level of the AR test. \textcolor{red}{THE FOLLOWING IS A CONJECTURE, STILL TO BE PROVEN:} The level of the K test under contamination, conditional on $D(F_n, \beta_0) = \tilde{D}$, is given by \begin{align*} \lim_{n \rightarrow \infty} \alpha_{K}(F_{\epsilon, n}^L | D(F_n, \beta_0) = \tilde{D}) = \alpha + \epsilon^2 \mu_1 \cdot \operatorname*{IF}(d; \sqrt{K(\cdot, \beta_0)}, F_{\theta} \big| \pi = \tilde{D})^2 + o(\epsilon^2). \end{align*} The level of the CLR test under contamination, conditional on $D(F_n, \beta_0) = \tilde{D}$, is given by \begin{align*} \lim_{n \rightarrow \infty} \alpha_{CLR}(F_{\epsilon, n}^L | D(F_n, \beta_0) = \tilde{D}) = \begin{cases} \lim_{n \rightarrow \infty} \alpha_{AR}(F_{\epsilon, n}^L) &if \tilde{D} = 0 \\ \lim_{n \rightarrow \infty} \alpha_{K}(F_{\epsilon, n}^L | D(F_n, \beta_0) = \tilde{D}) &if \tilde{D} \neq 0 \end{cases} \end{align*}

Asymptotic Distributions

In this section, we derive the asymptotic distribution of the general (robust) AR, K and CLR statistics defined in ((ref)), ((ref)) and ((ref)), and show how to construct tests based on these statistics. We start with a lemma that plays a key role in deriving all of the asymptotic distributions of the statistics.

lemmaUnder the null hypothesis $\beta = \beta_0$, and the regularity conditions of Assumption (ref), then as $n \rightarrow \infty$, \begin{align*} \sqrt{n} \left\{\begin{matrix} g(F_n, \beta_0) \\ D(F_n, \beta_0) \end{matrix} \right\} \overset{d}{\to} \mathcal{N}\left[\begin{pmatrix} 0 \\ \pi \end{pmatrix}, \left\{\begin{matrix} \Omega(F_{\theta}, \beta_0) & 0 \\ 0 & \Lambda(F_{\theta}, \beta_0) \end{matrix} \right\} \right]. \end{align*}
proofSee Appendix.

From Lemma (ref), we see that under the null hypothesis $\beta = \beta_0$, as $n \rightarrow \infty$,

align[align omitted — 84 chars of source]

Moreover, under the null hypothesis $\beta = \beta_0$ and $\pi = 0$, as $n \rightarrow \infty$,

align*[align* omitted — 64 chars of source]

Based on ((ref)), we can construct the general (robust) AR test as:

align*[align* omitted — 88 chars of source]

where $\chi_{k, 1-\alpha}^2$ denotes the $1-\alpha$ quantile of a $\chi_{k}^2$ distribution.

Next, in Proposition (ref), we derive the asymptotic distribution of the RK statistic.

propositionUnder the null hypothesis $\beta = \beta_0$, and the regularity conditions in Assumption (ref), then as $n \rightarrow \infty$, \begin{align} nRK(F_n, \beta_0) \overset{d}{\to} \chi_1^2. \end{align}
proofSee Appendix.

Based on Proposition (ref), we can introduce the general (robust) K test as:

align*[align* omitted — 86 chars of source]

Finally, we derive the asymptotic distribution of the RCLR statistic.

propositionUnder the null hypothesis $\beta = \beta_0$, and the regularity conditions in Assumption (ref), it holds that, conditional on $D(F_n, \beta_0) = \tilde{D}$ and with $\tilde{W} = \tilde{D}^{\top}\Lambda(F_{\theta}, \beta_0)^{-1} \tilde{D}$, then as $n \rightarrow \infty$, \begin{align} nRCLR(F_n, \beta_0) \overset{d}{\to} \frac{1}{2}\left\{\chi_{k - 1}^2 + \chi_{1}^2 - \tilde{W} + \sqrt{(\chi_{k - 1}^2 + \chi_{1}^2 + \tilde{W})^2 - 4 \tilde{W}\chi_{k - 1}^2 }\right\}, \end{align} where $\chi_{k - 1}^2$ and $\chi_{1}^2$ are independent chi-squared distributed random variables with $k - 1$ and $1$ degrees of freedom.
proofSee Appendix.

Using Proposition (ref) we can introduce the robust CLR test as follows:

align*[align* omitted — 105 chars of source]

where $c_{\alpha}\{\sqrt{n}D(F_n, \beta_0)\}$ denotes the conditional $1-\alpha$ quantile of the asymptotic distribution given in ((ref)). The critical values and confidence sets can be computed using simulation and test inversion moreira2003conditional, andrews2019weak. That is, we find all the values of $\beta_0$ for which the data does not reject the null hypothesis. The confidence set is then $\left\{\beta_0 \ |\ \beta_0 \in \mathbb{R} \text{ and } \phi_{RCLR}(\beta_0) = 0 \right\}$.

A Robust Conditional Likelihood Ratio Test of Mallows Type

In this section, we show how to construct a robust CLR statistic based on Mallows type M-estimators. These estimators solve the following estimating equations:

align*[align* omitted — 270 chars of source]

where we use $\omega(z_i) = \sqrt{1 - z_i^{\top}(Z^{\top}Z)^{-1}z_i}$ as a weight function, and $\rho \colon \mathbb{R} \mapsto \mathbb{R}$ is the Huber function:

align*[align* omitted — 215 chars of source]

The estimates $\sigma_{\epsilon}(F_n)$ and $\sigma_{v}(F_n)$ are robust estimates of the scales. In our case, for the practical implementation in Sections (ref) and (ref), we use the function rlm from the R package MASS mass2002. We use the default robust scale estimator provided, which is based on a re-scaled median absolute deviation of the residuals. In Sections (ref) and (ref), for simplicity, we do the estimation equation per equation and not simulteneously.

Now we only need to obtain estimates for the (co)variances belonging to these estimators. Using Equations ((ref)) and ((ref)), we derive the following influence functions for the Mallows type M-estimators:

align*[align* omitted — 576 chars of source]

where

align*[align* omitted — 192 chars of source]

Using these influence functions, we can calculate the (co)variance matrices. For instance, $\Sigma_{\delta \delta}(F_{\theta}) = \int \operatorname*{IF}\{d; \delta(\cdot), F_{\theta}\} \operatorname*{IF}\{d; \delta(\cdot), F_{\theta}\}^{\top} dF_{\theta} = M_{\delta \delta}(F_{\theta})^{-1} Q_{\delta \delta}(F_{\theta}) M_{\delta \delta}(F_{\theta})^{-\top}$, where

align*[align* omitted — 317 chars of source]

Evaluating $M_{\delta \delta}(\cdot)$ and $Q_{\delta \delta}(\cdot)$ at the empirical distribution $F_n$ yields empirical estimate $\Sigma_{\delta \delta}(F_n)$ of the variance matrix $\Sigma_{\delta \delta}(F_{\theta})$.

With these estimates, we can construct the robust CLR test (and also the robust AR and K tests). We study the performance of this test in different contaminated scenarios in Section (ref). Note, this is just one possible way to construct a robust CLR test. Other promising estimators that can be used to construct a robust CLR statistic, are MM-estimators yohai1987high and robust SUR estimators based on MM-estimators saraceno2021robust.

comment\subsection{An alternative approach} In Section (ref), we showed how to construct general (robust) AR, K and a Wald statistic directly based on reduced form (M-)estimates motivated by andrews2019weak. In this section, we introduce an alternative approach to robustify these statistics and use them to robustify the CLR statistic. As the AR and K statistic are $F$ and score type statistics the idea is to robustify the AR, K and Wald statistics by directly using (bounded-influence) robust $F$, score and Wald tests introduced by heritier1994robust. This can be achieved by first rewriting the model ((ref)) - ((ref)) to a $\beta_0$-restricted model as follows. Given a hypothesized $\beta_0$, we introduce $u(\beta_0) = \epsilon - v\beta_0$ and decompose $v$ into two parts as follows: $v = \rho(\beta_0) u(\beta_0) + v(\beta_0)^*$, where $\rho(\beta_0) = \sigma_{u(\beta_0), v}/\sigma_{u(\beta_0)}^2$. This yields equations \begin{align} y - x \beta_0 &= z^{\top}\kappa(\beta_0) + u(\beta_0) , \\ x - \rho(\beta_0) u(\beta_0) &= z^{\top} \pi(\beta_0) + v(\beta_0)^*, \end{align} where $\kappa(\beta_0) = \delta - \pi \beta_0 = \pi(\beta - \beta_0)$. Under the null hypothesis $\beta = \beta_0$, it holds that $u(\beta_0) = u, \gamma_1(\beta_0) = \gamma_1$ and $\kappa(\beta_0) = 0$. Hence, we can introduce a robust AR test by applying a robust $F$-test to test whether $\kappa(\beta_0) = 0$ in ((ref)), which is done in klooster2021outlier. Similarly, we can robustify the Wald statistic by using a robust Wald statistic heritier1994robust that tests whether $\pi = 0$ in ((ref)). The K statistic can be robustified by first estimating $\pi$ in ((ref)) using an M-estimator, and using this estimate to construct a robust test. This also yields robust AR, K and Wald statistics with a bounded influence function that we can use to robustify the CLR test. The (mathematical) details of this approach can be found in the supplementary material. In practice, we advise to use the robust CLR test based on the reduced form estimators for two reasons. First, in extra simulation studies we found that the performance in terms of robustness of the two different approaches is almost the same, while the alternative approach is more difficult to construct as it also requires an initial estimate of $\rho(\beta_0)$. Second, to construct confidence sets for the (robust) CLR tests, we need to rely on test inversion. For the alternative approach this means we need to obtain new M-estimates of $\kappa(\beta_0)$ and $\pi(\beta_0)$ for each $\beta_0$ value. In contrast, the method proposed based on reduced form estimators do not directly depend on $\beta_0$. This is preferable as the M-estimators we use rely on numerical optimization and do not have closed form solutions. Therefore, they can be vulnerable to numerical optimization errors and can take longer to compute when there is a lot of data (for example when applying it to the angrist1991does data in Section (ref)). Hence, from a computational perspective, the method proposed in this article is preferable.

Simulation Study

In this section, we numerically evaluate the robustness properties of the robust CLR test introduced in Section (ref) compared to the classical CLR test. We consider four different scenarios: a baseline scenario without contamination, two scenarios with contamination by a single outlier and scenario with “distributional" contamination in the error terms. These contaminated scenarios are generated by slightly altering the baseline scenario. The exact details of these different contaminations are given below. We start by introducing the scenario without contamination.

Let $\iota_3$ denote a 3-dimensional vector of ones. We generate data from the following model:

align[align omitted — 76 chars of source]

We generate three exogeneous instruments $z = (z_{1}, z_{2}, z_{3})^{\top}$ and one exogeneous control variable $w$. We draw $z_1, z_2, z_3$ and $w$ from independent standard normal distributions. The errors $u$ and $v$ are drawn from a bivariate normal distribution with variances equal to $1$ and correlation $\rho = 0.5$. We consider two different cases for the parameter $\pi \in \{0.1, 1\}$. When $\pi = 0.1$ we are in the weak instrument case and when $\pi = 1$ we are in the strong instrument case. The sample size is $n = 250$ and we repeat the study $10000$ times. In the simulation, we test $H_0 \colon \beta = 0$ at a 5% significance level.

Results

In Figure (ref), we show the power curves of the RCLR and CLR tests in a scenario without contamination. In this case, we see that both the RCLR and CLR tests are size correct in both the weak and strong instrument cases. Furthermore, we see that the CLR test is more powerful than the RCLR test in both the weak and strong instrument cases. However, we see that the loss of power of the RCLR test is not very large implying that the extra robustness comes at a low cost.

figure[figure omitted — 311 chars of source]

In Figure (ref), we show the power curves of the RCLR and CLR tests in a scenario with contamination by a large outlier in the $y$ variable. The data is first generated according to the baseline scenario and then an outlier is introduced. The outlier is generated by changing the first observation to $(y_{1}, x_{1}, z_{11}, z_{12}, z_{13}, w_{1}) = (20, x_{1}, z_{11}, z_{12}, z_{13}, w_{1})$, which means we only change the data point $y_{1}$ to $20$. In this case, we see that the RCLR test remains almost as powerful as in the case without contamination in both the weak and strong instrument cases. This happens, because the outlier is successfully downweighted by the robust M-estimators. In contrast, the CLR test loses power compared to the uncontaminated case and is now less powerful compared to the RCLR test. The loss of power happens, because the outlier increases the variance estimates of the classical test. This increased variance results in fewer rejections of the null, lowering the power. Overall, the CLR test remains reliable as we can see that it is still size correct. This happens because the (vertical) outlier does not bias the LS estimators $\delta(F_n)$ and $\pi(F_n)$. We would see a similar pattern if the outlier was in any other variable, for example the $x$ or $w$ variable, where the CLR test loses power but remains reliable. Therefore, we can ask the question whether it is necessary to use a robust CLR test. The simple answer is yes, as we demonstrate in the following contamination scenario.

figure[figure omitted — 329 chars of source]

In Figure (ref), we show the power curves of the RCLR and CLR tests in a scenario with contamination by a large outlier in the $y$ and $z$ variable. This outlier was generated by replacing the first observation with $(y_{1}, x_{1}, z_{11}, z_{12}, z_{13}, w_{1}) = (20, x_{1}, 5, z_{12}, z_{13}, w_{1})$. We see that the power curves of the robust CLR test remain size correct in both weak and strong instrument cases and only lose a little bit of power compared to the scenario without contamination. This shows that the robust CLR test effectively downweights the outlier. However, this is not the case for the CLR test. We see that in both the strong and weak instrument cases the CLR test is not size correct anymore and the minimum is shifted to the left. To understand how this happens, for simplicity, we consider the case with only one instrument. In this case the CLR statistic behaves the same as the AR statistic. The AR statistic is minimized when $g(F_n) = \delta(F_n) - \pi(F_n) \beta_0 = 0$. The outlier we constructed can positively bias the estimator $\delta(F_n)$ when a nonrobust estimator is used. Hence, under the null, we obtain $\delta(F_n) - \pi(F_n) \beta_0 = \delta(F_n) \approx \delta + b = \pi \beta + b$, where $b > 0$ denotes a positive bias term. Solving $\pi \beta + b = 0$ for $\beta$ gives $\beta = -b/\pi$. This formula shows that if there is a positive bias $b$, then the minimum will not be at $\beta = 0$ implying there must an overrejection at $\beta = 0$. Furthermore, as in our case the bias is positive and $\pi > 0$, we obtain a negative value for $\beta$ so that the power curves shift to the left. At last, when the instrument is weak and $\pi \approx 0$, then the $\beta = -b/\pi$ term will be large showing that a (small) bias caused by outliers can be more problematic when the instrument is weak compared to the strong instrument case. This last finding also holds more generally for biases that are not necessarily due to outliers small2008war.

figure[figure omitted — 375 chars of source]

In Figure (ref), we show the power curves of the RCLR and CLR tests in a scenario with contamination in the error terms. In this case, the first $50$ error terms $(u_i, v_i), i = 1, \dots, 50$, are drawn from a mean zero bivariate $t(3)$-distribution with correlation $\rho = 0.5$ instead of a mean zero bivariate normal distribution. In this case, we see that the RCLR test is strictly more powerful than the CLR test in both the weak and strong instrument settings. This happens, because the error terms drawn from the $t(3)$-distribution sometimes introduce larger outlying values that increase the variance estimates the CLR statistic relies on. This results in fewer rejections of the null hypothesis when $\beta \neq \beta_0$ lowering the power. The RCLR effectively downweights these outlying values, resulting in better variance estimates, which results in a higher power.

figure[figure omitted — 341 chars of source]

Empirical Examples

In this section, we show how the robust CLR test introduced in Section (ref) can be used in practice by revisiting three empirical studies. First, we consider the data and several specifications in alesina2011segregation who examine the effect of segregation on the quality of government. Second, we revisit the main specifications considered in ananat2011wrong where the effect of (racial) segregation on urban poverty and inequality is studied. Finally, we revisit the staiger1997instrumental specifications for the angrist1991does data where the effect of education on labor market earnings is studied.

Alesina and Zhuravskaya (2011)

alesina2011segregation study the effect of segregation on the quality of government in a cross section of countries. They find that ethnically and linguistically segregated countries have a lower quality of government. Furthermore, they find that there is no relationship between religious segregation and governance. To address endogeneity concerns caused by mobility and endogeneous internal borders, an instrument is constructed for segregation. For more information, data and the construction of the instrument we refer to alesina2011segregation.

In Section 5D alesina2011segregation mention that they carefully examined whether a handful of influential observations drive their results. By exluding influential observations and recalculating their statistics they conclude that this is not the case. However, alesina2011segregation mention that in the specifications of Panel D in Table 7 removing two influential observations leads to the first-stage $F$-statistic dropping from 17.22 to 7.82 making inference based on 2SLS estimator unreliable staiger1997instrumental. alesina2011segregation solve this by also removing the most influential observation in the first stage from the data so that the instrument becomes “strong enough". The practice of manually removing outliers from the data and then relying on classical statistical methods is not recommended for several reasons. The asymptotic distribution of the estimators and test statistics are unknown as they become conditional on the outlier detection procedure(s). The variability may be underestimated, which can lead to an overrejection of the null hypothesis. Lastly, some outliers might be difficult to detect due to masking effects. We refer to Section 4.3 in maronna2019robust and Section 1.3 in heritier2009robust for a more detailed discussion on these issues. Therefore, it is advisable to rely on robust estimation and testing procedures from the start. It seems that in case of alesina2011segregation that the use of the 2SLS estimator might be questionable due to the outliers and the weak instrument. Therefore, we believe that it is desirable to re-evaluate the robustness of the results and the stability of the conclusions. We apply our test to six specifications (“Voice", “Political stability", “Government effectiveness", “Regulatory Quality", “Rule of law" and “Control of corruption") of Panel D in Table 7 of alesina2011segregation. As there is only one instrument, we calculate the $95\%$ confidence sets of the robust AR statistic and the (classical) AR statistic. The results are reported in Table (ref).

table[table omitted — 2,064 chars of source]

From Table (ref), we can see that in specifications IV and V that the confidence set of the robust AR test is shifted compared to the AR confidence set. This suggests that outliers did have an effect as we would expect the confidence set of the robust AR test to be wider than the confidence set of the AR test when there are no outliers, but not shifted. The shift of the confidence set suggests that the LS estimators the AR test is constructed upon are biased due to the outlier(s). Therefore, the confidence sets based on the robust AR are more reliable. In specifications I, II, III and VI, we see that the robust AR confidence sets are wider, but not (much) shifted, compared to the AR confidence set. This suggests that outliers did not have a large effect. Overall, the outliers do not seem to be very problematic as in all specifications the final decision whether to reject or not reject the null hypothesis $H_0 \colon \beta = 0$ remains the same.

commentFrom Table (ref) we can see that in specifications II, III, IV and V that the confidence set of the robust AR confidence set is shifted compared to the AR confidence set. This suggests that outliers did have an effect in these regressions as we would expect the confidence set of the robust AR test to be wider than the confidence set of the AR test when there are no outliers, but not shifted. The shift of the confidence set suggests that the LS estimators the AR test is constructed upon are biased due to the outlier(s). Therefore, the confidence sets based on the robust AR are more reliable. Overall, the outliers do not seem to be very problematic as in all specifications, except specification II, the final decision whether to reject or not reject the null hypothesis $H_0 \colon \beta = 0$ remains the same. When we analyze specification II, we see that the robust confidence set would not reject the null hypothesis, while the classical confidence set would reject the null hypothesis. In this case, it does seem the outliers are the main drivers of the significant result.

Ananat (2011)

ananat2011wrong studies the effect of racial segregation on urban poverty and inequality. To overcome endogeneity issues, a railroad division index is used to instrument for racial segregation. Using this instrumental variable, ananat2011wrong shows that segregation increases metropolitan rates of black poverty and overall black-white income disparities, while decreasing rates of white poverty and inequality within the white population. For more information, data and the construction of the instrument we refer to ananat2011wrong.

klooster2021outlier show that an outlier in the control variable used in the main results of ananat2011wrong inflates the first-stage $F$-statistic from 1.83 to 19.32. As the outlier was not taken into account in the original study, it was assumed that the instrument was strong. Consequently, estimation was done with a 2SLS estimator and inference with a $t$-test. Due to the outlier (and the weak instrument) inference based on the 2SLS estimator might be unreliable. Therefore, to re-evaluate the robustness of the results, we apply our test to the four main specifications (“Gini index whites", “Gini index blacks", “Poverty rate whites", “Poverty rate blacks"), which can be found in columns (3) and (4) of Table 2 in ananat2011wrong. As there is only one instrument, we calculate the $95\%$ confidence sets of the robust AR statistic and the (classical) AR statistic. The results are reported in Table (ref).

table[table omitted — 1,578 chars of source]

When we analyze Table (ref), we find large differences between the robust and classical confidence sets. When there are no outliers in the data, we would expect that the robust confidence sets are only a bit wider than the classical confidence sets. However, in this case, the classical confidence sets are bounded convex sets, while the robust confidence sets are unbounded sets. The shape of the robust confidence sets do correctly suggest that the instrument is weak. In this example, the outlier does have a large effect and we recommend using the robust confidence sets for reliable inference.

Angrist and Krueger (1991)

angrist1991does study the effect of education on labor market earnings. To adress endogeneity issues, quarter of birth instruments are constructed for education. Using these instrumental variables they find a positive relationship between years of education and labor market earnings. We revisit four specifications presented in Table 2 of staiger1997instrumental based on the 1930 - 1939 cohort. For more information, data and the construction of the instrument, we refer to staiger1997instrumental and angrist1991does.

BoundJaegerBaker1995 showed that the relationship between the instruments and the endogeneous regressor is quite weak in certain specifications of angrist1991does. Furthermore, more recently, solvsten2020robust shows that the LIML residuals of the structural equation of a certain specification in angrist1991does is distributed roughly like a normal distribution at the center with outlying errors that closely follow a $t(3)$-distribution (reminiscent of the “distributional" contamination scenario in Section (ref)). Due to the weak instruments (and possible outliers), inference based on the 2SLS estimator used by angrist1991does might be unreliable. Moreover, due to outliers, weak instrument robust tests might be corrupted and/or inefficient. Therefore, to re-evaluate the robustness of the results we replicate the results reported in Panel A of Table 2 in staiger1997instrumental. For each specification, we give the $95\%$ confidence set of the CLR and robust CLR, and the first-stage $F$-statistic. The results are reported in Table (ref).

When we compare the $95\%$ confidence sets of the CLR and RCLR statistics, we note that the RCLR confidence sets are smaller than the CLR confidence sets in every specification. This happens because the RCLR statistic effectively downweights outlying values in the residuals. Similar as in the “distributional" contamination scenario in Section (ref) this results in better variance estimates and hence tighter confidence sets.

table[table omitted — 3,095 chars of source]

Conclusion

In this article, we proposed a general framework to construct weak instrument robust testing procedures that are also robust to outliers in the linear instrumental variable model. The framework is constructed upon M-estimators and we showed that the classical weak instrument robust tests, such as the AR, K and CLR tests, can be obtained by specifying the M-estimators to be the LS estimators. We formally showed that influence functions of the classical test statistics are not bounded and hence not (locally) robust. As the classical test statistics are not robust, we showed how to construct a robust CLR test based on a Mallows type M-estimator. By means of a simulation study, we documented good performance of our robust CLR test in different contamination scenarios. Finally, we illustrated how the robust CLR test can be used in practice by revisiting three different empirical studies.