EconBase
← Back to paper

A Simple Robust Procedure in Instrumental Variables Regression

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

174,438 characters · 20 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Simple Robust Procedure in Instrumental Variables Regression

abstractA common concern in empirical modelling centres around whether estimated regression coefficients are affected by a small set of outlying observations. To conduct outlier robustness checks in practical applications of instrumental variables regressions, the common practice is to run ordinary two stage least squares (2SLS) and remove observations with standardised residuals beyond a chosen cut-off value. Subsequently, the trimmed 2SLS is computed and compared to the original full-sample 2SLS. This paper aims to understand and improve the above heuristic procedure by establishing an asymptotic theory. Specifically, there are three main contributions of the paper. First, the trimmed 2SLS has a positive probability of removing observations even under the null hypothesis where the model contains no outliers. Under this situation, we derive a limiting Normal distribution of the trimmed 2SLS with the asymptotic variance as the ordinary one multiplied by a relative efficiency inflator. Furthermore, a bias correction factor is introduced for the variance estimator of structural errors, which otherwise would be downward biased. Second, a Hausman-type test is constructed to formalize the heuristic procedure of comparing between the two 2SLS estimators. Third, the trimmed 2SLS is a two-step procedure, which can be iterated until a fixed point is reached. The fixed point is shown to have the same first order asymptotics as the Huber-skip M-estimator. Our analysis involves a new class of empirical processes, whose theory would be of independent interest in applied probability. Simulation studies lend support to the asymptotic theory. An empirical illustration to Acemoglu et al. (2019) shows the utility of the proposed method. JEL Classification: C26, C36. \\ Keywords: Outlier robustness; Instrumental variables; Trimmed 2SLS; Fixed point; Huber-skip estimator; Empirical processes.

Introduction

A frequent concern in applied research centres around whether key empirical findings are driven by a tiny set of outliers. In practice, the most common approach to conduct outlier analysis is a trimming method based on the size of regression residuals, see for example Auerbach et al. (1994), De Long and Summers (1994), Fabrizio et al. (2007), Albouy (2012), Acemoglu et al. (2012), Toda and Walsh (2015), and Acemoglu et al. (2019). Those papers in instrumental variables regressions identify observations whose standardised structural (second stage) residual is above or below a chosen cut-off value (like $\pm 1.96$) and re-estimate the 2SLS model without these observations. The trimmed 2SLS is then compared to the original full-sample estimate to heuristically assess outlier sensitivity of estimated regression coefficients. This paper aims to understand and improve the above described heuristic procedure by establishing an asymptotic theory.

In the empirical literature, there is a recurrent critique that questions whether crucial results are robust after removing outlying observations. De Long and Summers (1991) found a positive effect of equipment investment on economic growth, which was followed by discussion of whether outliers drove such a positive connection, see Auerbach et al. (1994) and De Long and Summers (1994). When estimating an institutional impact on economic growth, Acemoglu et al. (2001) attracted extensive discussion of whether outliers undermined the validity of mortality rates as instruments, see Albouy (2012) and Acemoglu et al. (2012). Guthrie et al. (2012) criticised that a result in Chhaochharia and Grinstein (2009) was invalidated by the exclusion of a few observations. Herndon et al. (2014) also commented on the outlier sensitivity of the findings in Reinhart and Rogoff (2010). Toda and Walsh (2015) reported that trimming a small sub-sample of observations drastically changes the estimate of the risk aversion coefficient, making otherwise insignificant inference now significant. After conducting a comprehensive survey study of re-examining around 1400 IVs regressions in 32 papers published in journals of the American Economic Association, Young (2022) reported that most of the empirical findings obtained from the 2SLS models are sensitive to outliers. Broderick et al. (2020) developed a metric related to influence functions named “Approximate Maximum Inference Perturbation" to find small subsets of the data that strongly affect regression estimates. When re-exploring several applied papers using the robustness metric, they reported that empirical results can often be reversed by removing less than one percent of the sample.

All valid empirical results should therefore share one essential feature that the analysis are not affected by a small portion of outliers to a non-negligible degree, see Young (2018). In practical applications of IVs regressions, the common practice is to conduct outlier robustness checks by re-running 2SLS on clean data and comparing the results from the original full-sample ones. The clean data is obtained by identifying and removing observations whose standardised structural (second stage) residuals are beyond a chosen cut-off value. However, such a heuristic procedure in the empirical literature lacks formal justification. The leading purpose of this paper is thus to develop its asymptotic theory under the null hypothesis of no outliers. Specifically, there are three main contributions of the paper.

First, the trimmed 2SLS has a positive probability of removing observations even under the null hypothesis where the model contains no outliers. Under this situation, we derive a limiting Normal distribution of the trimmed 2SLS estimator with the asymptotic variance as the ordinary one multiplied by a relative efficiency inflator. The trimmed estimator is less efficient under the null of no outliers in the sense that its asymptotic variance is larger than the full-sample estimator such that the relative efficiency inflator is strictly greater than one. Furthermore, a bias correction factor is introduced for the variance estimator of structural errors, which otherwise would be inconsistent and downward biased. Simulation studies lend strong support to our established asymptotic theory. Subsequently, an adjusted inference on structural parameters is suggested by taking into account of both the relative efficiency and bias correction factors.

Second, our established asymptotic theory of the trimmed 2SLS estimator is necessary to formalize the outlier robustness checks as statistical tests. A Hausman-type test is thus constructed to assess the outlier sensitivity by comparing between the ordinary and trimmed 2SLS estimators. With our results, we can test whether the parameter of interest changes its value significantly before and after outlier removal, enabling formal statistical investigation of outlier robustness analysis. When outliers are present, we should rely on the trimmed 2SLS estimator that is more robust. When no outliers are present, the full-sample 2SLS is more efficient. Unfortunately, in practice the presence of outliers and any resulting distortion in coefficients is unknown. The proposed test therefore also helps assess whether the gain in robustness offsets the loss of efficiency when using the robust trimmed estimator.

Third, the trimmed 2SLS is actually a two-step procedure, and most empirical studies stop after the initial iteration. In contrast, we suggest iterating the two-step procedure until a fixed point is reached. The paper demonstrates that the fixed point of the iterated procedure has the same first order asymptotics as the Huber (1964) skip regression. Thus, repeated application of the procedure would provide an estimate that is more resistant to outliers. In practice, empirical researchers commonly choose the full-sample 2SLS as a starting point to identify and remove observations. We could use a more robust initial estimator, since the full-sample estimator is not robust to outliers itself especially in presence of high leverage points. Iterating the procedure would also produce a more stable estimate, which is less influenced by the choice of the initial estimator.

An empirical illustration to Acemoglu et al. (2019) shows the utility of the proposed method. The paper tries to estimate the causal effect of democracy on economic growth measured by log GDP per capita. Due to the endogeneity problem, they apply IVs regressions by constructing a spatial instrument named regional waves of democratization. To confirm the validity of their positive estimates on democracy, they choose the cut-off value $1.96$ and perform (and compare to) the trimmed 2SLS. This paper re-explores the outlier robustness analysis conducted in Acemoglu et al. (2019) by using the improved robust procedure that is guided by our asymptotic theory. We find that the Hausman-type tests are rejected at the $5 \%$ significance level for most coefficients. This indicates that the outlier removal estimator is distinct from the full-sample one at least in magnitude. The adjusted inference based on the robust trimmed estimator cannot reject the hypothesis that the coefficient on democracy equals zero at the $5 \%$ significance level. Thus, the result of outlier analysis does not robustly support that democracy has a strictly positive effect causally on economic growth, though there is little evidence for the view that democracy is a constraint for economic development.

Analysis of robust methods in the iterated one-step framework is by no means new. One-step estimators have been investigated by Bickel (1975), Ruppert and Carroll (1980), and Welsh and Ronchetti (2002). The idea of iterating one-step estimators can be found in Dollinger and Staudte (1991), who applied an influence function argument to demonstrate convergence of iteratively re-weighted least squares with smooth weights. Notwithstanding we are interested in binary weights, the spirit of their argument still provides a guidance for our work. Other related work includes for example Cavaliere and Georgiev (2013), who explored the first order autoregression with infinite variance. Huber-skip M and L-estimators have been recently studied by Johansen and Nielsen (2009, 2013, 2016a, 2016b, 2018). Berenguer-Rico and Nielsen (2018) and Berenguer-Rico and Wilms (2018) explored and proposed diagnostic tests on residuals for normality and heteroskedasticity after outlier removal. However, all these works are restricted in ordinary regressions.

There is a closely related paper that tries to solve exactly the same problem in IVs regressions. Kaji (2018) studied the trimmed 2SLS estimator by developing a theory of empirical processes and a functional delta method for quantile processes. His machinery can be used to investigate L-statistics represented outlier-robust estimators (integrals of empirical quantile functions with respect to corresponding random sample selection measures). Compared to Kaji (2018), we characterise the trimmed 2SLS statistics as a different class of empirical processes, for which we also need to develop a theory. Our argument is constructed specifically for the trimmed 2SLS, which thus allow us to explore the procedure in depth. For example, we can study the asymptotics of the variance estimator, iteration of the algorithms, variations of the robustification parameter (cut-off), and the procedure starting with a different initial estimator, like the least trimmed squares. In addition, Kaji (2018) applied the non-parametric bootstrap to implement the test of comparing the full-sample 2SLS to the trimmed one. Although the bootstrapping version of the test is valid under the null, it has the lower power under alternatives where there are significant outliers in the data. This is because the bootstrapping distribution is distorted by those bootstrap re-samples which include outliers, and thus it is distinct from what it should be under the null of no outliers.

Our analysis involves a new class of empirical processes. We derive first order asymptotic expansions for the required new class of weighted and marked empirical processes of residuals using a martingale decomposition argument. This is proceeded in a classical way by adding and subtracting a compensator term to construct a martingale. There are three key proofs to establish our empirical process theory. First, we linearise the compensator by using the first order Taylor expansion. Second, a chaining argument and concentration inequalities are applied to prove that the constructed martingale converges to zero in probability. Third, an asymptotic equi-continuity is demonstrated for the empirical processes by performing a dyadic argument. Technical difficulty here is to show convergence uniformly in all dimensions. The established theory of empirical processes would be of independent interest in applied probability.

We define a filtration $\mathcal{F}_{i - 1} = \sigma (z_{1}, \ldots, z_{i}, r_{1}, \ldots, r_{i - 1}, u_{1}, \ldots, u_{i - 1})$ to construct a martingale. Note that $\{ (z_{i}, r_{i}, u_{i}) \}_{i = 1}^{n}$ is the set of instruments, projection (first stage) errors, structural (second stage) errors. The empirical processes can then be defined from the generalized empirical distribution function of structural residuals

equation*[equation* omitted — 195 chars of source]

where the weights $w_{in}$ are combinations of the normalized $\mathcal{F}_{i - 1}$ measurable instruments $z_{in}$ and $u_{i}^{p}$ are the $\mathcal{F}_{i}$ adapted marks, while $a$, $b$ represent the normalized estimation errors for $\sigma$ (variance of structural errors), $\beta$ (structural parameters). Also note that $\Pi$ is the first stage regression coefficient matrix. We derive the theory that are uniform in $a$, $b$, $c$ and allow for a near $n^{1/4}$ inefficiency in the estimation uncertainties $a$, $b$.

The empirical process literature dates back to Kolmogorov (1933) and Smirnov (1939) who proposed a type of goodness of fit tests that check whether the distribution under the null is well specified by comparing it with the empirical distribution function. To build asymptotics of Kolmogorov-Smirnov type test statistics, Doob (1949) first demonstrated weak convergence of empirical distribution function, and then Donsker (1952) established the empirical process central limit theorem to close a gap in Doob's proof. Because there exists the measurability issue in $D[0,1]$ space, the classical method applies the Skorokhod metric instead of uniform topology to avoid non-measurability\footnote{see details in Billingsley (1968).}. Meanwhile, many researchers still equip $D[0,1]$ space with uniform metric but use the Hoffmann-J\o rgensen\footnote{a sequence of random elements $X_{n}$ weakly converges to the limiting random element $X$ if and only if $\mathsf{E}^{\ast} f(X_{n}) \to \mathsf{E} f(X)$ as $n \to \infty$ for any continuous and bounded function $f$ where $\mathsf{E}^{\ast}$ denotes the outer expectation.} definition of weak convergence instead. With the new concept of weak convergence, they extend the classical theory to the empirical process indexed by VC class of sets or functions using the entropy and bracketing argument\footnote{see the summary in van der Vaart and Wellner (1996).}. However, their work has not yet been extended to deal with the weights $w_{in}$ and marks $u_{i}^{p}$ appearing in our context.

Following the classical idea, Koul and Ossiander (1994) used the entropy argument to first build up the theory for the weighted empirical process in the autoregressive model\footnote{also see Koul (2002) and Koul and Ling (2006).}. Recently, empirical processes including both weights and marks have been analyzed by a series of papers, see Engler and Nielsen (2009), Johansen and Nielsen (2009, 2016a), Jiao and Nielsen (2017), Jiao (2017), Berenguer-Rico and Nielsen (2018), Berenguer-Rico, Johansen, and Nielsen (2019a, b), and Nielsen and Qian (2022). However, all these works are restricted in ordinary regressions.

The paper proceeds as follows: \S (ref) introduces the models and proposes the improved procedure robust to outliers. \S (ref) presents the main results followed by \S (ref) simulation studies. \S (ref) re-explores outlier robustness analysis in Acemoglu et al. (2019). \S (ref) establishes the theory for a new class of empirical processes with proofs in Appendix \S (ref), (ref). Lastly, proofs of main theorems in \S (ref) are shown in Appendix \S (ref).

Model and a robust procedure

The instrumental variables (IVs) regressions with some notations are described first. It is widely known that two stage least squares estimator is sensitive to outliers. We review an outlier robust procedure in the instrumental variables regressions.

Model

Suppose in the cross-sectional settings we have independently and identically distributed data $\{(y_{i}, x_{i}, z_{i})\}_{i = 1}^{n}$ across individuals, where $y_{i}$ is univariate and $x_{i}$, $z_{i}$ are multivariate with dimension $d_{x}$, $d_{z}$. Assume the data $\{(y_{i}, x_{i})\}_{i = 1}^{n}$ satisfies the structural equation

equation[equation omitted — 108 chars of source]

Structural errors $\{u_{i}\}_{i = 1}^{n}$ are univariate i.i.d. random variables with scale $\sigma$ so that $u_{i}/\sigma$ has the known density $\mathsf{f}_{u}(y)$ and distribution function $\mathsf{F}_{u}(y) = \mathsf{P}(u_{i}/\sigma \le y)$ for $y \in \mathbb{R}$ with mean $0$ and variance $1$. In the structural model some elements of $x_{i}$ are endogenous, i.e. $\mathsf{E} x_{i} u_{i} \neq 0_{d_{x}}$, instruments $z_{i}$ are required for estimating the parameter $\beta \in \mathbb{R}^{d_{x}}$ consistently. The first stage regression holds for $\{(x_{i}, z_{i})\}_{i = 1}^{n}$

equation[equation omitted — 127 chars of source]

Instruments $z_{i}$ are assumed to have the zero mean so $\mathsf{E} z_{i} = 0_{d_{z}}$ and to be orthogonal to errors $r_{i}$ in the first stage equation so $\mathsf{E} z_{i} r_{i}^{\prime} = 0_{d_{z} \times d_{x}}$. Innovations $\{r_{i}\}_{i = 1}^{n}$ are $d_{x}$-variate i.i.d. random vector with symmetric and positive definite dispersion matrix $\Sigma \in \mathbb{R}^{d_{x} \times d_{x}}$ so $\Sigma^{-1/2} r_{i}$ follows the density $\mathsf{f}_{r}(x)$ and distribution function $\mathsf{F}_{r}(x)$ for $x \in \mathbb{R}^{d_{x}}$ with mean $0_{d_{x}}$ and identity variance-covariance matrix $I_{d_{x}}$. The parameter $\Pi$ lies in $\mathbb{R}^{d_{z} \times d_{x}}$ and we suppose $d_{x} \le d_{z}$ meaning the number of regressors is less than or equal to the number of instruments. Furthermore, to identify structural parameters $\beta$, instruments $z_{i}$ need to be valid, i.e. $\mathsf{E} z_{i} u_{i} = 0_{d_{z}}$, and informative, i.e. rank condition for identification ($\operatorname{rank} \Pi = d_{x}$ and $\operatorname{rank} \mathsf{E} z_{i} z_{i}^{\prime} = d_{z}$).

To apply the martingale argument, we need to construct the filtration

equation[equation omitted — 162 chars of source]

such that $u_{i}$, $r_{i}$ are $\mathcal{F}_{i}$ measurable and independent of $\mathcal{F}_{i - 1}$ while $z_{i}$ is adapted to $\mathcal{F}_{i - 1}$. Notice in the IVs setting $u_{i}$ is correlated with $r_{i}$. Assume the scaled errors $(u_{i}/\sigma, \Sigma^{-1/2} r_{i})$ have the joint density $\mathsf{f}_{u, r}(y, x)$ and distribution function $\mathsf{F}_{u, r}(y, x)$ for $y \in \mathbb{R}$, $x \in \mathbb{R}^{d_{x}}$. Note the joint distribution $\mathsf{f}_{u, r}, \mathsf{F}_{u, r}$ does also depend on the covariance $\Omega = \mathsf{Cov}(u_{i}/\sigma, \Sigma^{-1/2} r_{i}) = \mathsf{E} (\Sigma^{-1/2} r_{i} u_{i} / \sigma)$, however for simplicity $\Omega \in \mathbb{R}^{d_{x}}$ is suppressed in the notation of joint density. Further suppose the joint density can be decomposed into the conditional and marginal ones so $\mathsf{f}_{u, r}(y, x) = \mathsf{f}_{u|r}(y|x) \mathsf{f}_{r}(x) = \mathsf{f}_{r|u}(x|y) \mathsf{f}_{u}(y)$. Then denote the conditional expected value

equation[equation omitted — 186 chars of source]

Notice $\xi_{y} \in \mathbb{R}^{d_{x}}$ is related to the covariance $\Omega$ between $u_{i} / \sigma$ and $\Sigma^{-1/2} r_{i}$. In practice $\mathsf{f}_{u, r}$, $\mathsf{F}_{u, r}$ will often be assumed to be $(1 + d_{x})$-variate normal, so the scaled error vector $(u_{i} / \sigma, \Sigma^{-1/2} r_{i})$ follows

equation[equation omitted — 310 chars of source]

As the marginal and conditional distribution of multivariate normal are still normally distributed, then $\Sigma^{-1/2} r_{i} | u_{i} / \sigma \sim \mathsf{f}_{r|u}$ follows normal $\mathsf{N} (\Omega u_{i} / \sigma, I_{d_{x}} - \Omega \Omega^{\prime})$. Thus the conditional expectation in ((ref)) has the form $\xi_{y} = \Omega y$. Denote

equation[equation omitted — 134 chars of source]

and introduce the truncated covariance

equation[equation omitted — 244 chars of source]

Further derive $\omega_{c}$ to attain

equation[equation omitted — 232 chars of source]

Under normality in ((ref)), $\omega_{c} = \tau_{2}^{c} \Omega$ and $\xi_{-y} = - \Omega y = - \xi_{y}$ such that $\zeta_{y}^{+} = 0_{d_{x}}$ and $\zeta_{y}^{-} = 2 \xi_{y} = 2 \Omega y$.

The robust procedure we analyzed detects outliers through checking if absolute standardised residuals in the structural equation ((ref)) are beyond the chosen cut-off value $c$ and then calculating the robust two stage least squares estimator from the non-outlying sample. This implicitly assumes symmetry of $u_{i}/\sigma$, while non-symmetry leads to specific forms of bias. We assume symmetry for $\mathsf{f}_{u}$ when analyzing the robust algorithm in \S (ref), but not for the general empirical process results in \S (ref). Let the absolute error $|u_{i}|/\sigma$ in ((ref)) have a density $\mathsf{g}_{u}(y)$ and distribution function $\mathsf{G}_{u}(y) = \mathsf{P}(|u_{i}|/\sigma \le y)$ for $y > 0$. With a symmetry assumption, $\mathsf{G}_{u}(y) = 2\mathsf{F}_{u}(y) - 1$ and $\mathsf{g}_{u}(y) = 2\mathsf{f}_{u}(y)$ for $y > 0$. Define $\psi = \mathsf{G}_{u}(c)$ so the probability of exceeding the cut-off $c$ is $\gamma = 1 - \psi$. Suppose the $k$-th moment of the density $\mathsf{f}_{u}$ exists, then introduce

equation[equation omitted — 182 chars of source]

Thus $\tau_{0}^{c} = \psi$, $\tau_{2} = 1$ while $\tau_{k} = \tau_{k}^{c} = 0$ for odd $k$ when assuming symmetry for $\mathsf{f}_{u}$. Define the conditional variance of $u_{i}/\sigma$ given $( |u_{i}|/\sigma \le c )$ as

equation[equation omitted — 178 chars of source]

This will be used as a bias correction factor for the variance estimate computed from the selected non-outlying sample. With normality assumption ((ref)), we have $u_{i} / \sigma \sim \mathsf{f}_{u}$ follows standard normal $\mathsf{N}(0, 1)$ then $\tau_{2}^{c} = \psi - 2c\mathsf{f}_{u}(c)$, $\tau_{4} = 3$, $\tau_{4}^{c} = 3 \psi - 2c(c^{2} + 3)\mathsf{f}_{u}(c)$.

For matrices $M$, choose the spectral norm $|M| = \max \{ \mathrm{eigen} (M^{\prime} M) \}^{1/2}$ so that for vectors $x$ then $|x|$ is the Euclidean norm. The spectral norm is compatible with respect to the Euclidean norm so $|Mx| \le |M| |x|$. Notice $(dM)$ represents the exterior product of all elements in the matrix $M$ to describe the multivariate integral, see the chapter 2 of Muirhead (1982). For instance consider a vector $x \in \mathbb{R}^{d_{x}}$, then $(dx)$ is the exterior product of some measures on $\mathbb{R}^{d_{x}}$.

Procedures robust to outliers

We first define an iterated version of robust two stage least squares estimators. Two specific examples with different initial estimators are then described.

To carry out outlier analysis, Acemoglu et al. (2019) first compute the full sample 2SLS and choose it as the initial estimator for the parameter $\beta$, then calculate all the residuals in the structural equation and detect outliers if the absolute value of standardised residuals in ((ref)) are beyond the selected cut-off value $1.96$. Running two stage least squares based on the non-outlying observations gives them the updated robust estimator for $\beta$. Finally, they compare two $\beta$ estimates for outlier robustness checks. In principle, a different initial estimators or cut-off values can be chosen and this robust procedure can also be iterated. We define a formal version of the algorithm with the corrected variance estimator as follows.

algorithm[algorithm omitted — 1,395 chars of source]

Algorithm (ref) is a similar procedure as the iterated one-step Huber-skip M-estimator\footnote{one-step updated estimator mimics the Huber (1964) skip estimator, which has criterion function $\rho (y)= \min(y^2, c^2)/2$ as opposed to the Huber estimator with criterion function $\rho (y)= y^{2}/2$ for $|y| \le c$ and $\rho(y)=c|y|-c^2/2$ otherwise, see Hampel et al. (1986, p. 104).} in the instrumental variables regressions. Note that the bias correction factor $\varsigma_{c}^{2}$ is required in ((ref)) in order to have a consistent estimator of $\sigma^{2}$, otherwise $\sigma^{2}$-estimator is downward biased with the probability limit $\varsigma_{c}^{2} \sigma^{2}$.

In \S (ref), we establish asymptotic theory of this algorithm, which can hold uniformly in the cut-off values $c$. The algorithm could start with a robust estimator, while as the procedure applied in Acemoglu et al. (2019) the so called Robustified 2SLS is initiated using the full sample 2SLS. The following algorithm formally defines Robustified 2SLS.

algorithm[algorithm omitted — 1,102 chars of source]

Algorithm (ref) is not very robust with respect to high leverage points when we have i.i.d. cross-sectional data. Thus, to deal with the high leverage points, another typical example of Algorithm (ref) starting with the split sample 2SLS was first proposed but in the classical regression by Hendry (1999) in his empirical work. This saturated idea is to divide the full sample into two sub-samples and use 2SLS calculated from each sub-sample to detect outliers in the other sub-sample. The following algorithm formally defines Saturated 2SLS.

algorithm[algorithm omitted — 1,831 chars of source]

Algorithm (ref) is possibly more robust than the Robustified Two Stage Least Squares when we have prior knowledge that outliers are located in a particular subset of the whole sample. Since the location of contaminated observations is unknown in most practical situations, the choice of the initial sets $\mathcal{I}_{1}$ and $\mathcal{I}_{2}$ should be iterated.

The main results

We start by listing assumptions, then asymptotics for Iterated 2SLS and weak convergence for Robustified 2SLS and Saturated 2SLS. It is followed by comparing efficiency between robust and non-robust 2SLS under the null of no outliers. Finally, I provide the valid inference procedure on $\beta$ when implementing these trimming methods and propose a new Hausman type test to check the presence of influential outliers.

Assumptions

We list the sufficient assumptions for asymptotic theory of Algorithm (ref), (ref), (ref). To simplify analysis, this section focuses on somewhat stronger conditions than they need to be. In section \S (ref) on the one-sided empirical process, we will introduce some weaker assumptions. For example, we impose the symmetricity assumption on $\mathsf{f}_{u}$ in this section but not in section \S (ref), though the idea of reasoning is the same and the results in \S (ref) can be extended to the asymmetric case using the argument in Johansen and Nielsen (2009).

Since $\{ (y_{i}, x_{i}, z_{i}) \}_{i = 1}^{n}$ are i.i.d., instruments $z_{i}$ can be normalized by the rate $n^{-1/2}$ so denote $z_{in} = n^{-1/2} z_{i}$, then we have $M_{z z, n} = \sum_{i = 1}^{n} z_{in} z_{in}^{\prime} = n^{-1} \sum_{i = 1}^{n} z_{i} z_{i}^{\prime} \overset{\mathsf{P}}{\to} \mathsf{E} z_{i} z_{i}^{\prime} = M_{z z}$ by Law of Large Numbers assuming $\mathsf{E} |z_{i}|^{2} < \infty$. Denote the infeasible ideal fitted value $\tilde{x}_{i} = \Pi^{\prime} z_{i}$ while the feasible counterpart is attained by estimating $\Pi$ consistently. The normalized $\tilde{x}_{i}$ is defined as $\tilde{x}_{in} = n^{-1/2} \tilde{x}_{i} = \Pi^{\prime} z_{in}$ such that

equation*[equation* omitted — 260 chars of source]

If $d_{x} \le d_{z}$ and rank condition for identification hold so that $\operatorname{rank} \Pi = d_{x}$, $M_{z z} > 0$, then we have $M_{\tilde{x} \tilde{x}} > 0$ as well.

To carry out asymptotic analysis, the scaled error vector $(u_{i} / \sigma, \Sigma^{-1/2} r_{i})$, instruments $z_{i}$, and the initial estimator $(\widetilde{\beta}, \widetilde{\sigma}^{2})$ must satisfy the following conditions.

assumptionLet $\mathcal{F}_{i - 1} = \sigma(z_{1}, \ldots, z_{i}, r_{1}, \ldots, r_{i - 1}, u_{1}, \ldots, u_{i - 1})$ be an increasing sequence of $\sigma$-fields so $u_{i-1}$, $r_{i - 1}$, $z_{i}$ are $\mathcal{F}_{i-1}$ measurable while $r_{i}$, $u_{i}$ are independent of $\mathcal{F}_{i - 1}$. Suppose $(u_{i}/\sigma, \Sigma^{-1/2} r_{i})$ have continuously differentiable joint, conditional, and marginal densities $\mathsf{f}_{u, r}(y, x) = \mathsf{f}_{u|r}(y|x) \mathsf{f}_{r}(x) = \mathsf{f}_{r|u}(x|y) \mathsf{f}_{u}(y)$ which are positive on $y \in \mathbb{R}$, $x \in \mathbb{R}^{d_{x}}$. Assume $d_{x} \le d_{z}$ and $\operatorname{rank} \Pi = d_{x}$. For $0 \le \kappa < \eta \le 1/4$, choose an integer $s \ge 2$ such that $2^{s - 1} > 1 + (1/4 - \eta)(1 + d_{x})$. Let $e = 1 + 2^{s + 1}$. Suppose \\ $(i)$ the marginal density $\mathsf{f}_{u}(y)$ is symmetric and satisfies for $y \in \mathbb{R}$ \\ $(a)$ tail monotonicity: $y^{e} \mathsf{f}_{u}(y), |y^{e + 1} \dot{\mathsf{f}}_{u}(y)|$ are decreasing for large $y$; \\ $(ii)$ the marginal density $\mathsf{f}_{r}(x)$ satisfies for $x \in \mathbb{R}^{d_{x}}$ \\ $(a)$ moments: $\int_{x \in \mathbb{R}^{d_{x}}} |x|^{4} \mathsf{f}_{r}(x) (dx) < \infty$; \\ $(iii)$ the joint and conditional densities $\mathsf{f}_{u, r}(y, x), \mathsf{f}_{u|r}(y|x)$ satisfy for $y \in \mathbb{R}, x \in \mathbb{R}^{d_{x}}$ \\ $(a)$ boundedness: $\sup_{y \in \mathbb{R}, x \in \mathbb{R}^{d_{x}}} |(1 + y) y^{e - 2} \mathsf{f}_{u|r}(y|x) + y^{e - 1} \dot{\mathsf{f}}_{u|r, y}(y|x)| < \infty$; \\ $(iv)$ the instruments $z_{i}$ satisfy \\ $(a)$ $M_{z z, n} = \sum_{i = 1}^{n} z_{in} z_{in}^{\prime} \overset{\mathsf{P}}{\to} M_{z z} > 0$; \\ $(b)$ $\max_{1 \le i \le n} |n^{1/2 - \kappa} z_{in}| = \mathrm{O}_{\mathsf{P}}(1)$; \\ $(c)$ $n^{-1} \mathsf{E} \sum_{i = 1}^{n} |n^{1/2} z_{in}|^{e} = \mathsf{E} |n^{1/2} z_{in}|^{e} = \mathrm{O}(1)$; \\ $(v)$ the initial estimators $(\widetilde{\beta}, \widetilde{\sigma}^{2})$ satisfy \\ $(a)$ $n^{1/2} (\widetilde{\beta} - \beta) = \mathrm{O}_{\mathsf{P}}(n^{1/4 - \eta})$; \\ $(b)$ $n^{1/2} (\widetilde{\sigma}^{2} - \sigma^{2}) = \mathrm{O}_{\mathsf{P}}(n^{1/4 - \eta})$.

The conditions $(ia, iia, iiia)$ are satisfied in a range of situations. In particular $(ia)$ is satisfied by the normal and t distribution; see Johansen and Nielsen (2016a, Example 3.1), while $(iia, iiia)$ holds for the normal distribution. Condition $(iv)$ is standard for stationary instruments, see Johansen and Nielsen (2016a, Example 3.2). Condition $(v)$ allows the standardised estimation errors to diverge at a rate of $n^{1/4 - \eta}$ rather than being bounded in probability. In particular, $\eta = 1/4$ can be chosen for estimators with standard convergence rates.

Notice $\kappa$ is not present in the moment condition $2^{s - 1} > 1 + (1/4 - \eta)(1 + d_{x})$, thus there is a trade-off between the convergence rate $\eta$ of initial estimators and the required number $s$ of moments for the density $\mathsf{f}_{u, r}$. In standard situations the normalized estimator is bounded so $\eta = 1/4$, then the moment condition to $s$ reduces to $s = 2$. Otherwise when the initial estimator diverges at the rate $\eta$, then the number of moments $s$ grows linearly with the dimension of the regressors $d_{x}$\footnote{also see discussion in Berenguer-Rico, Johansen, and Nielsen (2019a, Remark 3.1, 3.2).}.

Asymptotics of Algorithm (ref)

This subsection shows asymptotic theory for Algorithms (ref). Before establishing asymptotic properties of the iterated estimators for $\beta$, $\sigma^{2}$, we first demonstrate that the updated estimator for $\Pi$ is consistent uniformly in $c \in [c_{+}, \infty)$ for a finite number $c_{+} > 0$, given tightness of previous estimators for $\beta$, $\sigma^{2}$.

theoremConsider Algorithm (ref). Suppose Assumption (ref)$(ia, iia, iiia, iv)$ holds, and that $n^{1/2} (\widehat{\beta}_{c}^{(m)} - \beta)$, $n^{1/2} (\widehat{\sigma}_{c}^{(m)} - \sigma$) are $\mathrm{O}_{\mathsf{P}}(1)$ for any $m \in [0, \infty)$. Then as $n \to \infty$ \begin{equation*} \sup_{c_{+} \le c < \infty} |\widehat{\Pi}_{c}^{(m + 1)} - \Pi| = \mathrm{o}_{\mathsf{P}}(1). \end{equation*}

The proof of the above theorem uses empirical process theory studied in \S (ref). Then with uniform consistency of location estimator for $\Pi$ in the first stage regression ((ref)) we build a one-step stochastic expansion of the updated estimators for structural parameters $\beta$, $\sigma^{2}$ in ((ref)) in terms of its original estimators, kernels, and small remainder terms.

theoremConsider Algorithm (ref). Suppose Assumption (ref)$(ia, iia, iiia, iv)$ holds, and that $n^{1/2} (\widehat{\beta}_{c}^{(m)} - \beta)$, $n^{1/2} (\widehat{\sigma}_{c}^{(m)} - \sigma$) are $\mathrm{O}_{\mathsf{P}}(1)$ for any $m \in [0, \infty)$. Then as $n \to \infty$ and uniformly in $c \in [c_{+}, \infty)$ \begin{align*} n^{1/2} (\widehat{\beta}_{c}^{(m + 1)} - \beta) & = \frac{2c \mathsf{f}_{u}(c)}{\psi} n^{1/2} (\widehat{\beta}_{c}^{(m)} - \beta) + (M_{\tilde{x} \tilde{x}, n} \psi)^{-1} \sum_{i = 1}^{n} \tilde{x}_{in} u_{i} 1_{(|u_{i}| \le \sigma c)} + \mathrm{o}_{\mathsf{P}}(1), \\ n^{1/2} (\widehat{\sigma}_{c}^{(m + 1)} - \sigma) & = \frac{c (c^{2} - \varsigma_{c}^{2}) \mathsf{f}_{u}(c)}{\tau_{2}^{c}} n^{1/2} (\widehat{\sigma}_{c}^{(m)} - \sigma) + \frac{\sigma}{2 \tau_{2}^{c}} n^{-1/2} \sum_{i = 1}^{n} (\frac{u_{i}^{2}}{\sigma^{2}} - \varsigma_{c}^{2}) 1_{(|u_{i}| \le \sigma c)} \\ & + \frac{\mathsf{f}_{u}(c)}{\tau_{2}^{c}} ( \frac{c^{2} - \varsigma_{c}^{2}}{2} \zeta_{c}^{-} - \frac{2 c}{\psi} \omega_{c} )^{\prime} \Sigma^{1/2} n^{1/2} (\widehat{\beta}_{c}^{(m)} - \beta) \\ & - \frac{1}{\psi \tau_{2}^{c}} \omega_{c}^{\prime} \Sigma^{1/2} M_{\tilde{x} \tilde{x}, n}^{-1} \sum_{i = 1}^{n} \tilde{x}_{in} u_{i} 1_{(|u_{i}| \le \sigma c)} + \mathrm{o}_{\mathsf{P}}(1). \end{align*}
remarkThe above theorem explores expansion of Algorithm (ref) (Iterated 2SLS) in the IVs setting where $\mathsf{E} x_{i} u_{i} \neq 0_{d_{x}}$, which is the generalized result of iterated 1-step Huber-skip M-estimators in the classical regression; see Theorem 1 in Jiao and Nielsen (2017). To be specific, the expansion of $\beta$ estimator only depends on its own estimation error in the previous step; see Proof of Theorem (ref) (Appendix (ref)) why $n^{1/2} (\widehat{\sigma}_{c}^{(m)} - \sigma)$ does not occur, and moreover if $\tilde{x}_{in}$ is replaced by $x_{in}$ then it is the same as that in the classical regression. While the expansion of $\sigma^{2}$ estimator does also depend on the estimation error $n^{1/2} (\widehat{\beta}_{c}^{(m)} - \beta)$, this term disappears in the classical regression where $\mathsf{E} x_{i} u_{i} = 0_{d_{x}}$ and $\mathsf{E} r_{i} u_{i} = 0_{d_{x}}$ such that $\xi_{c} = 0_{d_{x}}$, $\zeta_{c}^{+} = \zeta_{c}^{-} = 0_{d_{x}}$, and $\omega_{c} = 0_{d_{x}}$. Thus, the second expansion degenerates to that in Jiao and Nielsen (2017).
remarkIn the one-step expansions, the autoregressive coefficients are $2c \mathsf{f}_{u}(c) / \psi$, $c (c^{2} - \varsigma_{c}^{2}) \mathsf{f}_{u}(c) / \tau_{2}^{c}$. By Johansen and Nielsen (2013, Theorem 3.6), also see Jiao and Nielsen (2017, Theorem 2), Assumption (ref)$(ia)$ implies that they are strictly bounded by one such that \begin{equation} \sup_{c_{+} \le c < \infty} \max \{ |\frac{2c \mathsf{f}_{u}(c)}{\psi}|, |\frac{c (c^{2} - \varsigma_{c}^{2}) \mathsf{f}_{u}(c)}{\tau_{2}^{c}}| \} < 1. \end{equation} This in fact shows the spectral radius of the autoregressive coefficient matrix is smaller than one in the iterative system in Theorem (ref), which is significant to establish the tightness and fixed point result, see the unit cycle boundedness condition ((ref)) in the one-step stochastic expansion system ((ref)), ((ref)), ((ref)) in Proof of Theorem (ref) (Appendix (ref)).

Assumption (ref)$(v)$ with $\eta = 1/4$ corresponds to a standard convergence rate for the initial estimator. Theorem (ref) provides an iterative equation between the updated and original estimators, while its autoregressive coefficient has a spectral radius strictly bounded by the unit cycle, see ((ref)) in Remark (ref). Thus a geometric argument and mathematical induction are then used to show $\widehat{\beta}_{c}^{(m)}$, $\widehat{\sigma}_{c}^{(m)}$ are tight in iteration $m \in [0, \infty)$ and in the cut-off value $c \in [c_{+}, \infty)$.

theoremConsider Algorithm (ref). Suppose Assumption (ref) holds with $\eta = 1/4$. Then as $n \to \infty$ \begin{equation*} \sup_{0 \le m < \infty} \sup_{c_{+} \le c < \infty} |n^{1/2} (\widehat{\beta}_{c}^{(m)} - \beta)| + |n^{1/2} (\widehat{\sigma}_{c}^{(m)} - \sigma)| = \mathrm{O}_{\mathsf{P}}(1). \end{equation*}

The above tightness result is required for building the weak convergence theory of iterated estimators for $\beta$, $\sigma^{2}$. Further combined with Theorem (ref) a corollary follows, to demonstrate uniform consistency of $\widehat{\Pi}_{c}^{(m)}$ in $m$ and $c$.

corollaryConsider Algorithm (ref). Suppose Assumption (ref) holds with $\eta = 1/4$. Then as $n \to \infty$ \begin{equation*} \sup_{0 \le m < \infty} \sup_{c_{+} \le c < \infty} |\widehat{\Pi}_{c}^{(m)} - \Pi| = \mathrm{o}_{\mathsf{P}}(1). \end{equation*}

Since tightness results for iterated estimators of $\beta$, $\sigma^{2}$ have been established by Theorem (ref), we can apply the one-step expansion in Theorem (ref) recursively. Then for any $m \in [0, \infty)$ the stochastic expansions of $m + 1$ step estimators are explored in terms of the initial estimators, kernels, and small remainder terms.

theoremConsider Algorithm (ref). Suppose Assumption (ref) holds with $\eta = 1/4$. Then as $n \to \infty$ and uniformly in $c \in [c_{+}, \infty)$ we have for any $m \in [0, \infty)$ \begin{align*} n^{1/2} (\widehat{\beta}_{c}^{(m + 1)} - \beta) & = \varrho_{\beta \beta, c}^{(m + 1)} n^{1/2} (\widehat{\beta}_{c}^{(0)} - \beta) + \varrho_{\beta \tilde{x} u, c}^{(m + 1)} M_{\tilde{x} \tilde{x}, n}^{-1} \sum_{i = 1}^{n} \tilde{x}_{in} u_{i} 1_{(|u_{i}| \le \sigma c)} + \mathrm{o}_{\mathsf{P}}(1), \\ n^{1/2}(\widehat{\sigma}_{c}^{(m+1)}-\sigma) & = \varrho_{\sigma\sigma,c}^{(m+1)} n^{1/2}(\widehat{\sigma}_{c}^{(0)}-\sigma) + \frac{\sigma}{2} \varrho_{\sigma uu,c}^{(m+1)} n^{-1/2} \sum_{i=1}^{n} ( \frac{u_i^2}{\sigma^2}-\varsigma_c^2 ) 1_{(|u_i|\leq\sigma c)} \\ & + \frac{\mathsf f_u(c)\varrho_{\sigma\beta,c}^{(m+1)}} {\tau_2^c} ( \frac{c^2-\varsigma_c^2}{2}\zeta_c^- - \frac{2c}{\psi}\omega_c )^{\prime} \Sigma^{1/2} n^{1/2}(\widehat{\beta}_{c}^{(0)}-\beta) \\ & + \{ \frac{ \varrho_{\sigma\tilde{x}u,c}^{(m+1)} (c^2-\varsigma_c^2)}{2} \zeta_c^- - \frac{2c\varrho_{\sigma\tilde{x}u,c}^{(m+1)} +\varrho_{\sigma uu,c}^{(m+1)}}{\psi} \omega_c \}^{\prime} \\ &\qquad\qquad\times \Sigma^{1/2} M_{\tilde{x}\tilde{x},n}^{-1} \sum_{i=1}^{n} \tilde{x}_{in}u_i 1_{(|u_i|\leq\sigma c)} + \mathrm{o}_{\mathsf P}(1), \end{align*} where coefficients have expressions \begin{align*} \varrho_{\beta \beta, c}^{(m + 1)} & = \{ \frac{2c \mathsf{f}_{u}(c)}{\psi} \}^{m + 1}, \qquad \varrho_{\beta \tilde{x} u, c}^{(m + 1)} = \frac{\psi^{m + 1} - \{ 2 c \mathsf{f}_{u}(c) \}^{m + 1}}{\psi^{m + 1} \{ \psi - 2 c \mathsf{f}_{u}(c) \}}, \\ \varrho_{\sigma \sigma, c}^{(m + 1)} & = \{ \frac{c (c^{2} - \varsigma_{c}^{2}) \mathsf{f}_{u}(c)}{\tau_{2}^{c}} \}^{m + 1}, \qquad \varrho_{\sigma u u, c}^{(m + 1)} = \frac{(\tau_{2}^{c})^{m + 1} - \{ c (c^{2} - \varsigma_{c}^{2}) \mathsf{f}_{u}(c) \}^{m + 1}}{ (\tau_{2}^{c})^{m + 1} \{ \tau_{2}^{c} - c (c^{2} - \varsigma_{c}^{2}) \mathsf{f}_{u}(c) \}}, \\ \varrho_{\sigma \beta, c}^{(m+1)} &= \sum_{l=0}^{m} \{ \frac{2c\mathsf f_u(c)}{\psi} \}^{m-l} \{ \frac{ c (c^2-\varsigma_c^2) \mathsf{f}_u(c) }{ \tau_2^c } \}^{l}, \\ \varrho_{\sigma\tilde{x}u,c}^{(m+1)} &= \frac{\mathsf f_u(c)} { \tau_2^c\{\psi-2c\mathsf f_u(c)\} } \Bigg[ \frac{ (\tau_2^c)^{m+1} - \{c (c^2-\varsigma_c^2) \mathsf f_u(c)\}^{m+1} }{ (\tau_2^c)^m \{\tau_2^c-c (c^2-\varsigma_c^2) \mathsf f_u(c)\} } \\ &\qquad\qquad - \sum_{l=0}^{m} \{ \frac{2c\mathsf f_u(c)}{\psi} \}^{m-l} \{ \frac{ c (c^2-\varsigma_c^2) \mathsf f_u(c) }{ \tau_2^c } \}^{l} \Bigg]. \end{align*}

Initially, a tight estimator is assumed to be available. It is then iterated using the one-step stochastic expansion presented in Theorem (ref). The spectral radius of the corresponding autoregressive coefficient matrix is uniformly bounded away from one, as shown in ((ref)) of Remark (ref). Consequently, as the number of iterations becomes sufficiently large, the effects of the initial estimation errors vanish and the iterated estimator converges in probability to a fixed point. To characterize this fixed point, let \(m\to\infty\) in Theorem (ref). This gives

align*[align* omitted — 481 chars of source]

such that the fixed-point follows

equation*[equation* omitted — 204 chars of source]
theoremConsider Algorithm (ref). Suppose Assumption (ref) holds with $\eta = 1/4$. Then for all $\epsilon, \delta > 0$ a pair $n_{0} > 0, m_{0} > 0$ exists, so for $n > n_{0}$ and $m > m_{0}$ \begin{equation*} \mathsf{P} \{ \sup_{c_{+} \le c < \infty} | n^{1/2} (\widehat{\beta}_{c}^{(m)} - \widehat{\beta}_{c}^{(\ast)}) | + | n^{1/2} (\widehat{\sigma}_{c}^{(m)} - \widehat{\sigma}_{c}^{(\ast)}) | > \delta \} < \epsilon, \end{equation*} where \begin{align*} n^{1/2} (\widehat{\beta}_{c}^{(\ast)} - \beta) & = \frac{1}{\psi - 2c \mathsf{f}_{u}(c)} M_{\tilde{x} \tilde{x}, n}^{-1} \sum_{i = 1}^{n} \tilde{x}_{in} u_{i} 1_{(|u_{i}| \le \sigma c)}, \\ n^{1/2} (\widehat{\sigma}_{c}^{(\ast)} - \sigma) & = \frac{\sigma}{2 \{ \tau_{2}^{c} - c(c^{2} - \varsigma_{c}^{2})\mathsf{f}_{u}(c) \}} n^{-1/2} \sum_{i = 1}^{n} (\frac{u_{i}^{2}}{\sigma^{2}} - \varsigma_{c}^{2}) 1_{(|u_{i}| \le \sigma c)} \\ & + \frac{\{\frac{(c^{2} - \varsigma_{c}^{2}) \mathsf{f}_{u}(c)}{2} \zeta_{c}^{-} - \omega_{c} \}^{\prime} }{\{ \psi - 2 c \mathsf{f}_{u}(c) \} \{ \tau_{2}^{c} - c (c^{2} - \varsigma_{c}^{2}) \mathsf{f}_{u}(c) \}} \Sigma^{1/2} M_{\tilde{x} \tilde{x}, n}^{-1} \sum_{i = 1}^{n} \tilde{x}_{in} u_{i} 1_{(|u_{i}| \le \sigma c)}. \end{align*}

Based on Theorem (ref), if the initial estimator is bounded in a large compact set with large probability, then any iterated estimator takes values in the same compact set. The proof of Theorem (ref) is to further argue that the deviation between the $m$-fold iterated estimator and the fixed point is the sum of two terms vanishing exponentially and in probability respectively when $m$ and $n$ tend to infinity.

Algorithm (ref) mimics the Huber (1964) skip estimator in the IVs context. By investigating Theorem (ref) we reason that through infinite iterations the fixed point of the algorithm approximates the Huber-skip IVs estimator in the sense that they have the same asymptotic expansion in terms of kernels and thus follow the same limiting distribution (see Theorem (ref)).

Similar to Remark (ref), Theorems (ref) and (ref) extend the $(m + 1)$-fold expansion and the fixed point result of the iterated 1-step Huber-skip M-estimator in classical regression to the IVs setting where $\mathsf{E} x_{i} u_{i} \neq 0_{d_{x}}$ is allowed\footnote{Replacing $\tilde{x}_{in}$ by $x_{in}$ and setting $\mathsf{E} x_{i} u_{i} = 0_{d_{x}}$ such that $\mathsf{E} r_{i} u_{i} = 0_{d_{x}}$, we then have $\xi_{c} = 0_{d_{x}}$, $\zeta_{c}^{+} = \zeta_{c}^{-} = 0_{d_{x}}$, and $\omega_{c} = 0_{d_{x}}$. Thus all results degenerate to those shown in Jiao and Nielsen (2017).}. Under normality of $(u_{i}/\sigma, \Sigma^{-1/2} r_{i})$, $\xi_{c} = \Omega c$, $\zeta_{c}^{+} = 0_{d_{x}}$, $\zeta_{c}^{-} = 2 \Omega c$, and $\omega_{c} = \tau_{2}^{c} \Omega$, so coefficients appearing in expansions in Theorems (ref) and (ref) can be further simplified.

Weak convergence of Algorithm (ref) and (ref)

Algorithms (ref) and (ref) are special versions of Algorithm (ref) with different starting points. Their initial estimators are either the full sample or split sample two stage least squares which do not depend on the cut-off, and so satisfy the tightness property. Therefore, theorems and corollaries in the previous subsection apply for these two algorithms as well. Moreover, since Algorithms (ref) and (ref) start with two stage least squares estimators whose statistical properties are well known, the asymptotic distribution can then be established for these two robust procedures.

When researchers carry out empirical work using instrumental variables regression, their interest mainly lies in making inference on $\beta$ whereas they are only concerned about the consistency of estimators for $\sigma^{2}$, $\Pi$. Theorem (ref) and Corollary (ref) have already shown that iterated estimators of $\sigma^{2}$, $\Pi$ are consistent, so in order to perform inference on structural parameters in ((ref)) we next build distributional theory for the iterated estimator of $\beta$ in Algorithms (ref) and (ref).

Choosing the distinct cut-off value $c$ in the interval $[c_{+}, \infty)$, we then get a process $\mathbb{G}_{n}^{(m + 1)}(c) = n^{1/2} (\widehat{\beta}_{c}^{(m + 1)} - \beta)$ for any $m \in [0, \infty)$. A weak convergence theory for $\mathbb{G}_{n}^{(m + 1)}$ follows from a finite dimensional convergence derived by the expansion in Theorem (ref) and tightness in Theorem (ref) (see Billingsley, 1968). The iterated estimator of $\beta$ as a process is asymptotically approximated by the Gaussian process\footnote{see the definition of Gaussian processes in Adler and Taylor (2009, p.\ 27).}.

We first analyze Algorithm (ref) by providing asymptotics for the full sample two stage least squares estimator $\widetilde{\beta}$ defined through its initial estimate $\widehat{\beta}_{c}^{(0)}$ in ((ref)).

lemmaConsider the full sample two stage least squares \begin{equation*} \widetilde{\beta} = (\widetilde{\Pi}^{\prime} \sum_{i = 1}^{n} z_{i} z_{i}^{\prime} \widetilde{\Pi})^{-1} (\widetilde{\Pi}^{\prime} \sum_{i = 1}^{n} z_{i} y_{i}), \quad where \quad \widetilde{\Pi} = (\sum_{i = 1}^{n} z_{i} z_{i}^{\prime})^{-1} (\sum_{i = 1}^{n} z_{i} x_{i}^{\prime}). \end{equation*} Suppose Assumption (ref)$(ia, iia, iva, ivc)$ holds. Then as $n \to \infty$ \begin{equation*} n^{1/2} (\widetilde{\beta} - \beta) = M_{\tilde{x} \tilde{x}, n}^{-1} \sum_{i = 1}^{n} \tilde{x}_{in} u_{i} + \mathrm{o}_{\mathsf{P}}(1). \end{equation*} Furthermore, we have \begin{equation*} n^{1/2} (\widetilde{\beta} - \beta) \overset{\mathsf{D}}{\to} \mathsf{N}(0_{d_{x}}, \sigma^{2} M_{\tilde{x} \tilde{x}}^{-1}). \end{equation*}

The next step is to establish the weak convergence theory for the process $\mathbb{G}_{n}^{(m + 1)}$ of the iterated $\beta$ estimator where $m \in [0, \infty)$.

theoremConsider Algorithm (ref). Suppose Assumption (ref)$(ia, iia, iiia, iv)$ holds. Denote the process $\mathbb{G}_{n}^{(m + 1)}(c) = n^{1/2} (\widehat{\beta}_{c}^{(m + 1)} - \beta)$ for $c \in [c_{+}, \infty)$ and $m \in [0, \infty)$. Then as $n \to \infty$ we have \begin{equation*} \mathbb{G}_{n}^{(m + 1)}(c) = \varrho_{\beta \beta, c}^{(m + 1)} M_{\tilde{x} \tilde{x}, n}^{-1} \sum_{i = 1}^{n} \tilde{x}_{in} u_{i} + \varrho_{\beta \tilde{x} u, c}^{(m + 1)} M_{\tilde{x} \tilde{x}, n}^{-1} \sum_{i = 1}^{n} \tilde{x}_{in} u_{i} 1_{(|u_{i}| \le \sigma c)} + \mathrm{o}_{\mathsf{P}}(1), \end{equation*} where $\varrho_{\beta \beta, c}^{(m + 1)}$, $\varrho_{\beta \tilde{x} u, c}^{(m + 1)}$ are defined in Theorem (ref). Furthermore, $\mathbb{G}_{n}^{(m + 1)}$ weakly converges to a zero mean Gaussian process $\mathbb{G}^{(m + 1)}$ with variance given as \begin{equation*} \mathsf{Var} \{ \mathbb{G}^{(m + 1)}(c) \} = \{ (\varrho_{\beta \beta, c}^{(m + 1)})^{2} + 2 \tau_{2}^{c} \varrho_{\beta \beta, c}^{(m + 1)} \varrho_{\beta \tilde{x} u, c}^{(m + 1)} + \tau_{2}^{c} (\varrho_{\beta \tilde{x} u, c}^{(m + 1)})^{2} \} \sigma^{2} M_{\tilde{x} \tilde{x}}^{-1}. \end{equation*}

The one-step updated estimator from the full sample 2SLS is of particular interest in Algorithm (ref). To explore its asymptotics let $m = 0$ in Theorem (ref) such that $\varrho_{\beta \beta, c}^{(1)} = 2 c \mathsf{f}_{u}(c) / \psi$ and $\varrho_{\beta \tilde{x} u, c}^{(1)} = \psi^{-1}$, then the corollary below follows.

corollaryConsider Algorithm (ref). Suppose Assumption (ref)$(ia, iia, iiia, iv)$ holds. Denote the process $\mathbb{G}_{n}^{(1)}(c) = n^{1/2} (\widehat{\beta}_{c}^{(1)} - \beta)$ for $c \in [c_{+}, \infty)$. Then as $n \to \infty$ we have \begin{equation*} \mathbb{G}_{n}^{(1)}(c) = \frac{2 c \mathsf{f}_{u}(c)}{\psi} M_{\tilde{x} \tilde{x}, n}^{-1} \sum_{i = 1}^{n} \tilde{x}_{in} u_{i} + (M_{\tilde{x} \tilde{x}, n} \psi)^{-1} \sum_{i = 1}^{n} \tilde{x}_{in} u_{i} 1_{(|u_{i}| \le \sigma c)} + \mathrm{o}_{\mathsf{P}}(1). \end{equation*} Furthermore, $\mathbb{G}_{n}^{(1)}$ weakly converges to a zero mean Gaussian process $\mathbb{G}^{(1)}$ with variance given as \begin{equation*} \mathsf{Var} \{ \mathbb{G}^{(1)}(c) \} = \frac{4 c^{2} \mathsf{f}_{u}^{2}(c) + 4 \tau_{2}^{c} c \mathsf{f}_{u}(c) + \tau_{2}^{c}}{\psi^{2}} \sigma^{2} M_{\tilde{x} \tilde{x}}^{-1}. \end{equation*}

Then we analyze Algorithm (ref) and find asymptotically it is equivalent to Algorithm (ref) although they start with distinct initial estimates when implementing Algorithm (ref).

theoremConsider Algorithm (ref) where $n_{1} = n_{2} = n / 2$. Suppose Assumption (ref)$(ia, iia, iiia , iv)$ holds for each sub-sample set $\mathcal{I}_{1}$, $\mathcal{I}_{2}$. Then as $n \to \infty$ and for any $m \in [0, \infty)$, the process of the estimator $n^{1/2} (\widehat{\beta}_{c}^{(m + 1)} - \beta)$ for $c \in [c_{+}, \infty)$ has the same asymptotic expansion as for Algorithm (ref) and thus weakly converges to the identical Gaussian process reported in Theorem (ref).

The proof involves checking whether the expansion of the first step updated estimator for Algorithm (ref) is the same as Algorithm (ref) using arguments in Theorems (ref) and (ref). Since the two algorithms are iterated from the split half or full sample 2SLS respectively, they have the same asymptotics if their $\widehat{\beta}_{c}^{(1)}$ perform identically when $n \to \infty$.

Finally, when $n \to \infty$ and $m \to \infty$, the process $\mathbb{G}_{n}^{(m)}$ uniformly converges to the limiting process $\mathbb{G}^{(\infty)} = \mathbb{G}^{(\ast)}$ in probability. Since $\varrho_{\beta \beta, c}^{(\infty)} = 0$, $\varrho_{\beta \tilde{x} u, c}^{(\infty)} = 1 / \{ \psi - 2 c \mathsf{f}_{u}(c) \}$, then Theorems (ref) and (ref) immediately show the weak convergence of the fixed point process $\mathbb{G}^{(\ast)}$ for Algorithms (ref), (ref), and (ref).

theoremConsider Algorithm (ref) with tight initial estimators, such as Algorithm (ref), (ref). Suppose Assumption (ref)$(ia, iia, iiia, iv)$ holds. When $m \to \infty$ let the fixed point process $\mathbb{G}_{n}^{(\ast)}(c) = n^{1/2} (\widehat{\beta}_{c}^{(\ast)} - \beta)$ for $c \in [c_{+}, \infty)$ as in Theorem (ref). Then as $n \to \infty$ the process $\mathbb{G}_{n}^{(\ast)}$ has the asymptotic expansion shown in Theorem (ref) and thus weakly converges to a zero mean Gaussian process $\mathbb{G}^{(\ast)}$ with variance given as \begin{equation*} \mathsf{Var} \{ \mathbb{G}^{(\ast)}(c) \} = \frac{\tau_{2}^{c}}{\{ \psi - 2 c \mathsf{f}_{u}(c) \}^{2}} \sigma^{2} M_{\tilde{x} \tilde{x}}^{-1}. \end{equation*}

Consider Algorithms (ref) and (ref). Weak convergence in Theorems (ref) and (ref) implies pointwise convergence, so for any $c \in [c_{+}, \infty)$, $m \in [0, \infty)$, and as $n \to \infty$ it holds

equation[equation omitted — 420 chars of source]

where $\varrho_{\beta \beta, c}^{(m + 1)}$, $\varrho_{\beta \tilde{x} u, c}^{(m + 1)}$ are defined in Theorem (ref).

Johansen and Nielsen (2009, 2013) built up the asymptotic distributions for the one-step Huber-skip M-estimator and its infinite iteration either starting from full sample or split sample least squares in the classical setting where $\mathsf{E} x_{i} u_{i} = 0_{d_{x}}$\footnote{also see Johansen and Nielsen (2016b) for the asymptotic distribution of the $(m + 1)$-step estimator where $m \in [0, \infty)$.}. The distributional results for Algorithms (ref) and (ref) are suprisingly the same as the iterated one-step Huber-skip M-estimators if $\tilde{x}_{in}$ is replaced by $x_{in}$ even though dependence is allowed between regressors and errors in the IVs regression. This is due to the fact that the stochastic expansion for the updated $\beta$ estimator only depends on its own previous step, not on the $\sigma$ estimator (see Theorem (ref) and Remark (ref)). Furthermore, this result is because $\zeta_{c}^{+}$ defined in ((ref), \S (ref)) does not appear in the $\beta$ expansion since the data is demeaned such that $\mathsf{E} z_{i} = 0_{d_{z}}$\footnote{see the details in Proof of Theorem (ref), Appendix (ref).}.

Efficiency comparison

Outlier analysis is trade-off between efficiency and robustness: Algorithms (ref) and (ref) will lose efficiency under the null of no outliers relative to the non-robust ordinary 2SLS since it is with positive probability to wrongly detect outliers, whereas they are indeed more robust under the alternative when there is data contamination. This subsection mainly investigates efficiency loss of two algorithms in return for gaining robustness.

Distributions of $\widehat{\beta}_{c}^{(m + 1)}$ are given by ((ref)) while Lemma (ref) shows the limit of the ordinary two stage least squares estimator $\widetilde{\beta}$. Thus, we can now compare the efficiency of the robust estimator $\widehat{\beta}_{c}^{(m + 1)}$ with respect to the non-robust estimator $\widetilde{\beta}$ under the null of no outliers. Relative efficiency is defined by

align[align omitted — 536 chars of source]

In the rest of the paper we treat $\vartheta_{c}^{(m + 1)}$ as a scalar factor though it is a scaled identity matrix. We are interested in efficiency comparison for two special cases: $m = 0$ and $m \to \infty$ such that

align[align omitted — 576 chars of source]
figure[figure omitted — 1,048 chars of source]

Using the standard normal density for $\mathsf{f}_{u}$, we plot the relative efficiency ((ref)) against cut-off values $c \in [c_{+}, \infty)$ and iterations $m \in [0, \infty)$ in Figure (ref). To further illustrate the 3D Figure (ref), we also plot two 2D figures, Figure (ref): The efficiency ((ref)), ((ref)) against cut-off values; Figure (ref): The efficiency ((ref)) against iteration steps for the cut-off values $c_{1} = 1.645 (10 \%)$, $c_{2} = 1.96 (5 \%)$, $c_{3} = 2.576 (1 \%)$.

Asymptotics in this paper are derived under the null hypothesis of no outliers where Algorithms (ref) and (ref) have a probability to falsely trim the non-outlying observations given the cut-off $c$. Thus, robust estimators have a higher asymptotic variance relative to the ordinary 2SLS, and so the relative efficiency plots lie in the interval $(0, 1)$ as shown in the three relative efficiency figures.

When the cut-off values $c$ become larger, the chance of wrongly detecting outliers tends to be smaller. Subsequently, the robust estimators have the smaller variance, and the efficiency become larger and closer to one, and vice versa if $c$ decreases. First, observe the efficiency plot of the infinite step estimator $\widehat{\beta}_{c}^{(\ast)}$ in Figure (ref), and find that it is consistent with what we discussed. Then check the solid line of the first step estimator $\widehat{\beta}_{c}^{(1)}$, and find that its efficiency tends to one when $c \to 0$ as well as when $c \to \infty$. It is intuitive that for large $c$ the efficiency increases as $c$ increases. However we also find for small $c$ the efficiency increases as $c$ decreases. This is due to the fact that the updated estimator $\widehat{\beta}_{c}^{(1)}$ is asymptotically equivalent to 2SLS when $c \to 0$ as stated in Remark (ref). This fact is also supported by the efficiency plot (ref) for other iterations.

remarkLet $m = 0$ and $n \to \infty$, Theorem (ref) gives the equation \begin{equation*} n^{1/2} (\widehat{\beta}_{c}^{(1)} - \beta) = \frac{2 c \mathsf{f}_{u}(c)}{\psi} n^{1/2} (\widetilde{\beta} - \beta) + (M_{\tilde{x} \tilde{x}, n} \psi)^{-1} \sum_{i = 1}^{n} \tilde{x}_{in} u_{i} 1_{(|u_{i}| \le \sigma c)} + \mathrm{o}_{\mathsf{P}}(1), \end{equation*} uniformly in $c \in [c_{+}, \infty)$. Note that $\widetilde{\beta}$ is 2SLS so $\widehat{\beta}_{c}^{(1)}$ is its updated estimator. Suppose $\mathsf{f}_{u}$ follows the standard normal then $\psi = 2 \mathsf{F}_{u}(c) - 1$, $\tau_{2}^{c} = \psi - 2 c \mathsf{f}_{u}(c)$, and $\dot{\mathsf{f}}_{u}(c) = - c \mathsf{f}_{u}(c)$. Let $c_{+} > 0$ be sufficiently small so that we can investigate the situation where $c \to 0$. Lemma (ref) (Appendix (ref)) shows the weak limit of the kernel term such that \begin{equation*} (M_{\tilde{x} \tilde{x}, n} \psi)^{-1} \sum_{i = 1}^{n} \tilde{x}_{in} u_{i} 1_{(|u_{i}| \le \sigma c)} \overset{\mathsf{D}}{\to} \mathsf{N}(0_{d_{x}}, \frac{\tau_{2}^{c}}{\psi^{2}} \sigma^{2} M_{\tilde{x} \tilde{x}}^{-1}). \end{equation*} By L'H{\^o}pital's rule, we have $\tau_{2}^{c} / \psi^{2} \to 0$ as $c \to 0$ so the kernel term vanishes. Again when $c \to 0$, L'H{\^o}pital's rule gives $2 c \mathsf{f}_{u}(c) / \psi \to 1$, thus $n^{1/2} (\widehat{\beta}_{c}^{(1)} - \beta)$ is asymptotically equivalent to $n^{1/2} (\widetilde{\beta} - \beta)$ indicating an efficiency of unity. In fact, check the efficiency formula given by ((ref)). When $c \to 0$, use $\tau_{2}^{c} / \psi^{2} \to 0$ and $2 c \mathsf{f}_{u}(c) / \psi \to 1$ to get $\psi^{2} / \{ 4 c^{2} \mathsf{f}_{u}^{2}(c) + 4 \tau_{2}^{c} c \mathsf{f}_{u}(c) + \tau_{2}^{c} \} \to 1$, then find $\mathrm{efficiency} (\widehat{\beta}_{c}^{(1)}, \widetilde{\beta}) \to 1$.

Figure (ref) also indicatess the efficiency of the first iteration estimator dominates the infinite step. This fact is consistent with Figure (ref) showing in general that the efficiency decreases as iteration increases regardless of cut-off values $c$. Figure (ref) reveals that the efficiency drops dramatically within the first few steps and converges to the fixed point upon through large iterations, also shown in Figure (ref). Intuition is that as iteration increases the algorithms throw away more observations and thus loss more efficiency until attaining the fixed point. It is not surprising that the critical value $c_{3} = 2.576 (1 \%)$ has the highest relative efficiency and the $c_{1} = 1.645 (10 \%)$ has the lowest while $c_{2} = 1.96 (5 \%)$ is in between.

Figure (ref) ((ref) and (ref)) imply that Algorithms (ref) and (ref) do not lose too much efficiency compared to the ordinary 2SLS if the cut-off value $c$ and iteration step $m$ are appropriately chosen under the null of no outliers, whereas these two procedures can produce a more robust estimate under alternatives where the data is contaminated.

Adjusted inference on structural parameters $\beta$

Assume that the density $\mathsf{f}_{u}$ is known, say standard normal $\mathsf{N}(0, 1)$, and choose the cut-off value $c$ and iteration step $m$, say $c = 1.96 (5 \%)$ and $m = 0$ or $\infty$, when carrying out Algorithms (ref) and (ref). The standard method of conducting inference based on ordinary 2SLS asymptotics\footnote{see ordinary 2SLS asymptotics in Lemma (ref).} are not reliable for three reasons. Firstly, variance $\sigma^{2}$ cannot be estimated consistently without bias correction factor $\varsigma_{c}^{2}$ introduced in ((ref), \S (ref)) Secondly, the standard method fails to take into account of the relative efficiency factor $\vartheta_{c}^{(m + 1)}$ appearing in the asymptotic variance of $\widehat{\beta}_{c}^{(m + 1)}$, see ((ref)) and ((ref)). Thirdly, to obtain the standard error of $\widehat{\beta}_{c}^{(m + 1)}$ the usual inferential procedure incorrectly divides the estimated asymptotic variance by the sub-sample size $\sum_{i = 1}^{n} v_{i, c}^{(m)}$ of all non-outlying observations retained, while the full sample size $n$ should have been used.

On implementing Algorithms (ref) and (ref) to perform valid inference on the structural parameters $\beta$ given the cut-off value $c$ and iteration step $m$, we have to assume the explicit form of $\mathsf{f}_{u}$ in order to obtain the bias correction factor $\varsigma_{c}^{2}$ and relative efficiency factor $\vartheta_{c}^{(m + 1)}$. In practice, it is statistically difficult to estimate the density $\mathsf{f}_{u}$. This is because, without knowing $\varsigma_{c}^{2}$ as $\mathsf{f}_{u}$ is unknown, the standard estimate of $\sigma^{2}$ will be biased downward, which will subsequently produce upward biased standardised residuals for structural errors $u_{i}$ although residuals can be estimated consistently based on beta estimates. Then the estimated density of $\mathsf{f}_{u}$ using the biased standardised residuals will be inconsistent by whatever methods. Thus, assumption on the known density $\mathsf{f}_{u}$ is not avoidable to conduct valid inference on $\beta$ in this paper. Proving validity of bootsrap Kaji (2018) applied nonparametric bootstrap to obtain standard errors of beta estimates when conducting inference, which however does not require the assumed known form of the density $\mathsf{f}_{u}$.

Weak limits of beta estimates in the step $m + 1$ are provided by ((ref)) such that

equation*[equation* omitted — 190 chars of source]

where $\vartheta_{c}^{(m + 1)}$ is defined by ((ref)) and the cut-off $c$ is given. By assuming on the form of $\mathsf{f}_{u}$, $\vartheta_{c}^{(m + 1)}$ is known as well as $\varsigma_{c}^{2}$ so $\sigma^{2}$ can be estimated consistently by $(\widehat{\sigma}_{c}^{(m + 1)})^{2}$, see Theorem (ref). Corollary (ref) implies we can consistently estimate $M_{\tilde{x} \tilde{x}} = \Pi^{\prime} \mathsf{E} z_{i} z_{i}^{\prime} \Pi$ by

equation[equation omitted — 260 chars of source]

Notice that $\sigma^{2}$ and $M_{\tilde{x} \tilde{x}}$ above are estimated using the robust estimator with the sub-sample after trimming all outliers detected, since the estimator of standard errors should be robust itself, otherwise the inference on $\beta$ might lead to incorrect results under the alternative hypothesis where data is contaminated. Then, the beta estimates are asymptotically approximated by

equation*[equation* omitted — 233 chars of source]

For conducting valid inference, the above estimated asymptotic variance can then be further divided by $n$ to obtain standard errors so that

equation*[equation* omitted — 218 chars of source]

The standard inferential procedure fails to take into consideration of $(\vartheta_{c}^{(m + 1)})^{-1}$, so mistakenly applies usual 2SLS asymptotics to get the asymptotic variance of $\widehat{\beta}_{c}^{(m + 1)}$ by $\mathrm{avar}_{\mathrm{2SLS}} = \sigma^{2} M_{\tilde{x} \tilde{x}}^{-1}$. Thus, to address this problem the relative efficiency factor $\vartheta_{c}^{(m + 1)}$ should be adjusted to $\mathrm{avar}_{\mathrm{2SLS}}$, then the normal weak limit is rearranged as

equation*[equation* omitted — 180 chars of source]

Furthermore, without considering the consistency factor $\varsigma_{c}^{-2}$ by the standard method $\sigma^{2}$ is inconsistently estimated by $\varsigma_{c}^{2} (\widehat{\sigma}_{c}^{(m + 1)})^{2}$. Then asymptotic variance of $\widehat{\beta}_{c}^{(m + 1)}$ is falsely estimated by $\widehat{\mathrm{avar}}_{\mathrm{2SLS}, c}^{(m + 1)} = \varsigma_{c}^{2} (\widehat{\sigma}_{c}^{(m + 1)})^{2} (\widehat{M}_{\tilde{x} \tilde{x}, c}^{(m + 1)})^{-1}$, which should be rectified by $\iota_{c}^{(m + 1)} = (\vartheta_{c}^{(m + 1)} \varsigma_{c}^{2})^{-1}$ so as to fix the inferential problem. Therefore, the correct asymtotic variance is the term $\widehat{\mathrm{avar}}_{\mathrm{2SLS}, c}^{(m + 1)}$ estimated by usual asymptotics but adjusted by $\iota_{c}^{(m + 1)}$, then error-free normal approximation can now be expressed as

equation*[equation* omitted — 182 chars of source]

In addition, the standard method incorrectly estimates the variance of $\widehat{\beta}_{c}^{(m + 1)}$ by $\widehat{\mathrm{var}}_{\mathrm{2SLS}, c}^{(m + 1)}$ through dividing $\widehat{\mathrm{avar}}_{\mathrm{2SLS}, c}^{(m + 1)}$ by the sub-sample size $\sum_{i = 1}^{n} v_{i, c}^{(m)}$ of non-outlying observations rather than the full sample size $n$. Thus, the erroneous standard errors are computed as the square root of $\widehat{\mathrm{var}}_{\mathrm{2SLS}, c}^{(m + 1)} = \varsigma_{c}^{2} (\widehat{\sigma}_{c}^{(m + 1)})^{2} (\widehat{\Pi}_{c}^{(m + 1) \prime} \sum_{i = 1}^{n} z_{i} z_{i}^{\prime} v_{i, c}^{(m)} \widehat{\Pi}_{c}^{(m + 1)})^{-1}$. Therefore, correct standard errors should be based on $\widehat{\mathrm{var}}_{\mathrm{2SLS}, c}^{(m + 1)}$ by usual 2SLS asymptotics, however rectified by $n^{-1} \sum_{i = 1}^{n} v_{i, c}^{(m)}$ and $\iota_{c}^{(m + 1)}$. In summary, the valid inference should rely on the following rearranged normal approximation

equation*[equation* omitted — 202 chars of source]

Consider the standard normal $\mathsf{N}(0, 1)$ for $\mathsf{f}_{u}$. Choose the cut-off value $c$ as $1.645 (10 \%)$, $1.96 (5 \%)$, $2.576 (1 \%)$ and set the iteration step $m$ as $0$, $\infty$. Notice that $n^{-1} \sum_{i = 1}^{n} v_{i, c}^{(m)}$ is the proportion of non-outlying observations retained, which has the probability limit $\psi$ under the null of no outliers due to Lemma (ref) and LLN. List in Table (ref) the probability $\psi$ within the cut-off, the bias correction factor $\varsigma_{c}^{2}$ for estimating $\sigma^{2}$, the relative efficiency factor $\vartheta_{c}^{(m + 1)}$ for asymptotic variance of $\widehat{\beta}_{c}^{(m + 1)}$, the adjusting factor $\iota_{c}^{(m + 1)}$ for estimating asymptotic variance, and the adjusting factor $\psi \iota_{c}^{(m + 1)}$ for estimating variance of $\widehat{\beta}_{c}^{(m + 1)}$ under the null of no outliers for large samples.

table[table omitted — 1,407 chars of source]

It is fundamental to note that $\psi$, $\varsigma_{c}^{2}$, $\vartheta_{c}^{(1)}$, $\vartheta_{c}^{(\ast)}$ are between $0$ and $1$, so $\iota_{c}^{(1)} = (\vartheta_{c}^{(1)} \varsigma_{c}^{2})^{-1}$, $\iota_{c}^{(\ast)} = (\vartheta_{c}^{(\ast)} \varsigma_{c}^{2})^{-1}$ and $\psi \iota_{c}^{(1)}$, $\psi \iota_{c}^{(\ast)}$ are greater than $1$, since the trimming method underestimates the variance $\sigma^{2}$ and has the lower efficiency (larger asymptotic variance) than ordinary 2SLS such that the adjusted (estimated) asymptotic variance and standard errors of beta estimates are larger than those provided by usual asymptotics.

Consistent with Figures (ref) and (ref), as the cut-off $c$ increases then $\psi$, $\varsigma_{c}^{2}$, $\vartheta_{c}^{(1)}$, $\vartheta_{c}^{(\ast)}$ increase and approach to $1$ so $\iota_{c}^{(1)}$, $\iota_{c}^{(\ast)}$ decrease and approach to $1$ along with the fact that $\psi \iota_{c}^{(1)}$, $\psi \iota_{c}^{(\ast)}$ approach to $\iota_{c}^{(1)}$, $\iota_{c}^{(\ast)}$. Moreover, $\vartheta_{c}^{(1)}$ in the first step dominates the infinite step iteration $\vartheta_{c}^{(\ast)}$ so that magnitudes $\iota_{c}^{(1)}$, $\psi \iota_{c}^{(1)}$ in the first step is smaller and closer to $1$ than that of the infinite step $\iota_{c}^{(\ast)}$, $\psi \iota_{c}^{(\ast)}$.

Nuisance parameters $\xi_{c}$, $\zeta_{c}^{+}$, $\zeta_{c}^{-}$\footnote{see $\xi_{c}$ in ((ref), \S (ref)) and $\zeta_{c}^{+}$, $\zeta_{c}^{-}$ in ((ref), \S (ref)).} measure the dependence between $x_{i}$ and $u_{i}$ (between $r_{i}$ and $u_{i}$) in the IVs setup. They are significantly important to evaluate the degree of endogeneity in the structural equation, but difficult to estimate in practice. However, performing valid inference on $\beta$ does not require estimating these nuisance parameters, since they all vanish asymptotically and thus do not contribute to standard errors. Therefore, Algorithms (ref) and (ref) can be easily implemented not only to identify outliers and to obtain a robust estimator, but also to carry out valid inference.

A new Hausman type test of outlier robustness

A frequent concern in empirical economics is whether a tiny set of outliers may have invalidated empirical results. The common practice is to carry out robustness checks by redoing the analyses with the sample after trimming all outliers detected by Algorithm (ref) and comparing results from the original ones with the full sample. Instead of heuristically checking the difference between robustified and ordinary 2SLS without knowing whether it is statistically significant, this paper formalizes the outlier robustness check as a new type of Durbin (1954)-Hausman (1978)-Wu (1973) test.

The test is based on trade-off between robustness and efficiency and enables to judge whether the least squares estimation is appropriate or the robust method should be preferred. As discussed in \S (ref), the robust estimator produced by Algorithm (ref) is consistent both under the null and the alternative (although less efficient under the null), whereas ordinary 2SLS is efficient (and consistent) under the null, but inconsistent otherwise. The test statistics is looking on statistically significant difference between robustified and ordinary 2SLS. If the model is correctly specified, the Hausman type test statistics should be rather small under the null of no outliers, since two consistent methods should produce estimates that are very close. On the contrary when outliers have large influence on least squares estimation, the robust trimming method should be very different from the ordinary estimate, so 2SLS should be rejected and the robustified 2SLS should be preferred even if less efficient. The proposed test thus will evaluate whether the gain in robustness is more valuable than the corresponding loss in efficiency.

In practice, empirical researchers run the full sample 2SLS $\widetilde{\beta}$ and compare to the robustified 2SLS $\widehat{\beta}_{c}^{(m + 1)}$ on the sub-sample with all outliers removed by Algorithm (ref) or (ref). The new type of Hausman test can detect whether two estimates are significantly distinct by looking on the L2 norm of difference between $\widetilde{\beta}$ and $\widehat{\beta}_{c}^{(m + 1)}$. Thus, it is essential first to derive the stochastic expansion and weak limit of a sequence of processes $\mathbb{H}_{n}^{(m + 1)}(c) = n^{1/2} (\widehat{\beta}_{c}^{(m + 1)} - \widetilde{\beta})$ for $c \in [c_{+}, \infty)$ and $m \in [0, \infty)$.

theoremConsider Algorithm (ref) or (ref). Suppose Assumption (ref)$(ia, iia, iiia, iv)$ holds. For $c \in [c_{+}, \infty)$ and $m \in [0, \infty)$ denote the process $\mathbb{H}_{n}^{(m + 1)}(c) = n^{1/2} (\widehat{\beta}_{c}^{(m + 1)} - \widetilde{\beta})$. Then as $n \to \infty$ we have \begin{equation*} \mathbb{H}_{n}^{(m + 1)}(c) = (\varrho_{\beta \beta, c}^{(m + 1)} - 1) M_{\tilde{x} \tilde{x}, n}^{-1} \sum_{i = 1}^{n} \tilde{x}_{in} u_{i} + \varrho_{\beta \tilde{x} u, c}^{(m + 1)} M_{\tilde{x} \tilde{x}, n}^{-1} \sum_{i = 1}^{n} \tilde{x}_{in} u_{i} 1_{(|u_{i}| \le \sigma c)} + \mathrm{o}_{\mathsf{P}}(1), \end{equation*} where $\varrho_{\beta \beta, c}^{(m + 1)}$, $\varrho_{\beta \tilde{x} u, c}^{(m + 1)}$ are defined in Theorem (ref). Furthermore, $\mathbb{H}_{n}^{(m + 1)}$ weakly converges to a zero mean Gaussian process $\mathbb{H}^{(m + 1)}$ with variance given as \begin{equation*} \mathsf{Var} \{ \mathbb{H}^{(m + 1)}(c) \} = \{ (\varrho_{\beta \beta, c}^{(m + 1)} - 1)^{2} + 2 \tau_{2}^{c} (\varrho_{\beta \beta, c}^{(m + 1)} - 1) \varrho_{\beta \tilde{x} u, c}^{(m + 1)} + \tau_{2}^{c} (\varrho_{\beta \tilde{x} u, c}^{(m + 1)})^{2} \} \sigma^{2} M_{\tilde{x} \tilde{x}}^{-1}. \end{equation*}

The proof of the above theorem follows from stochastic expansions of $\widehat{\beta}_{c}^{(m + 1)}$ in Theorem (ref), (ref) and $\widetilde{\beta}$ in Lemma (ref). The next corollary establishes a new type of Hausman test for outlier robustness checks, which is based on pointwise convergence of $\mathbb{H}_{n}^{(m + 1)}(c) = n^{1/2} (\widehat{\beta}_{c}^{(m + 1)} - \widetilde{\beta})$ implied by its weak limit in Theorem (ref).

corollaryConsider Algorithm (ref) or (ref). Suppose Assumption (ref)$(ia, iia, iiia, iv)$ holds. For $c \in [c_{+}, \infty)$ and $m \in [0, \infty)$ then as $n \to \infty$ we have \begin{equation*} n^{1/2} (\widehat{\beta}_{c}^{(m + 1)} - \widetilde{\beta}) \overset{\mathsf{D}}{\to} \mathsf{N}\{ 0_{d_{x}}, \mathsf{avar}(\widehat{\beta}_{c}^{(m + 1)} - \widetilde{\beta}) \}, \end{equation*} where $\mathsf{avar}(\widehat{\beta}_{c}^{(m + 1)} - \widetilde{\beta}) = \{ (\varrho_{\beta \beta, c}^{(m + 1)} - 1)^{2} + 2 \tau_{2}^{c} (\varrho_{\beta \beta, c}^{(m + 1)} - 1) \varrho_{\beta \tilde{x} u, c}^{(m + 1)} + \tau_{2}^{c} (\varrho_{\beta \tilde{x} u, c}^{(m + 1)})^{2} \} \sigma^{2} M_{\tilde{x} \tilde{x}}^{-1}$. Then the new type of Hausman test statistics has the weak limit \begin{equation*} H_{n, c}^{(m + 1)} = n (\widehat{\beta}_{c}^{(m + 1)} - \widetilde{\beta})^{\prime} \mathsf{avar}(\widehat{\beta}_{c}^{(m + 1)} - \widetilde{\beta})^{-1} (\widehat{\beta}_{c}^{(m + 1)} - \widetilde{\beta}) \overset{\mathsf{D}}{\to} \chi^{2}_{d_{x}}. \end{equation*}

Hausman (1978) argued that, when two estimators are correlated (one is always consistent but inefficient under the null, the other efficient but not consistent under the alternative), the asymptotic variance of their difference is given by the difference of their respective asymptotic variances. However, his argument requires additional regularity conditions that might not be satisfied in our context. The next remark is to show that the argument does work so $\mathsf{avar}(\widehat{\beta}_{c}^{(m + 1)} - \widetilde{\beta}) = \mathsf{avar}(\widehat{\beta}_{c}^{(m + 1)}) - \mathsf{avar}(\widetilde{\beta})$ when $\mathsf{f}_{u} \sim \mathsf{N}(0, 1)$.

remarkRearrange the expression of $\mathsf{avar}(\widehat{\beta}_{c}^{(m + 1)} - \widetilde{\beta})$ in Corollary (ref) to achieve \begin{equation*} \{ (\varrho_{\beta \beta, c}^{(m + 1)})^{2} - 2 \varrho_{\beta \beta, c}^{(m + 1)} + 1 + 2 \tau_{2}^{c} \varrho_{\beta \beta, c}^{(m + 1)} \varrho_{\beta \tilde{x} u, c}^{(m + 1)} - 2 \tau_{2}^{c} \varrho_{\beta \tilde{x} u, c}^{(m + 1)} + \tau_{2}^{c} (\varrho_{\beta \tilde{x} u, c}^{(m + 1)})^{2} \} \sigma^{2} M_{\tilde{x} \tilde{x}}^{-1}. \end{equation*} Assume $\mathsf{f}_{u} \sim \mathsf{N}(0, 1)$, then $\tau_{2}^{c} = \psi - 2 c \mathsf{f}_{u}(c)$. Recall from Theorem (ref) that \begin{equation*} \varrho_{\beta \beta, c}^{(m + 1)} = \{ \frac{2c \mathsf{f}_{u}(c)}{\psi} \}^{m + 1}, \qquad \varrho_{\beta \tilde{x} u, c}^{(m + 1)} = \frac{\psi^{m + 1} - \{ 2 c \mathsf{f}_{u}(c) \}^{m + 1}}{\psi^{m + 1} \{ \psi - 2 c \mathsf{f}_{u}(c) \}}. \end{equation*} Apply these terms to attain $- 2 \varrho_{\beta \beta, c}^{(m + 1)} + 1 - 2 \tau_{2}^{c} \varrho_{\beta \tilde{x} u, c}^{(m + 1)} = -1$. Further notice that $\mathsf{avar}(\widehat{\beta}_{c}^{(m + 1)}) = \{ (\varrho_{\beta \beta, c}^{(m + 1)})^{2} + 2 \tau_{2}^{c} \varrho_{\beta \beta, c}^{(m + 1)} \varrho_{\beta \tilde{x} u, c}^{(m + 1)} + \tau_{2}^{c} (\varrho_{\beta \tilde{x} u, c}^{(m + 1)})^{2} \} \sigma^{2} M_{\tilde{x} \tilde{x}}^{-1}$ and $\mathsf{avar}(\widetilde{\beta}) = \sigma^{2} M_{\tilde{x} \tilde{x}}^{-1}$, see ((ref)) and Lemma (ref). Thus, we finally shows \begin{align*} \mathsf{avar}(\widehat{\beta}_{c}^{(m + 1)} - \widetilde{\beta}) & = \{ (\varrho_{\beta \beta, c}^{(m + 1)} - 1)^{2} + 2 \tau_{2}^{c} (\varrho_{\beta \beta, c}^{(m + 1)} - 1) \varrho_{\beta \tilde{x} u, c}^{(m + 1)} + \tau_{2}^{c} (\varrho_{\beta \tilde{x} u, c}^{(m + 1)})^{2} \} \sigma^{2} M_{\tilde{x} \tilde{x}}^{-1} \\ & = \{ (\varrho_{\beta \beta, c}^{(m + 1)})^{2} + 2 \tau_{2}^{c} \varrho_{\beta \beta, c}^{(m + 1)} \varrho_{\beta \tilde{x} u, c}^{(m + 1)} + \tau_{2}^{c} (\varrho_{\beta \tilde{x} u, c}^{(m + 1)})^{2} \} \sigma^{2} M_{\tilde{x} \tilde{x}}^{-1} - \sigma^{2} M_{\tilde{x} \tilde{x}}^{-1} \\ & = \mathsf{avar}(\widehat{\beta}_{c}^{(m + 1)}) - \mathsf{avar}(\widetilde{\beta}). \end{align*}

Remark (ref) demonstrates that, in practical applications, it is equivalent either to estimate $\mathsf{avar}(\widehat{\beta}_{c}^{(m + 1)} - \widetilde{\beta})$ or $\mathsf{avar}(\widehat{\beta}_{c}^{(m + 1)}) - \mathsf{avar}(\widetilde{\beta})$ so as to obtain test statistics of outlier robustness developed in Corollary (ref). If $u_{i} / \sigma \sim \mathsf{f}_{u}$ follows $\mathsf{N}(0, 1)$ according to which the cut-off value $c$ is chosen, terms $\tau_{2}^{c}$, $\varrho_{\beta \beta, c}^{(m + 1)}$, $\varrho_{\beta \tilde{x} u, c}^{(m + 1)}$ are known for any iteration step $m$, so the only item left to estimate in the asymptotic variance is $\sigma^{2} M_{\tilde{x} \tilde{x}}^{-1}$.

As discussed in \S (ref), the estimator of asymptotic variance should be robust itself, so $\sigma^{2} M_{\tilde{x} \tilde{x}}^{-1}$ should be estimated by $(\widehat{\sigma}_{c}^{(m + 1)})^{2}$, $\widehat{M}_{\tilde{x} \tilde{x}, c}^{(m + 1)}$, see ((ref)), ((ref)), based on the sub-sample with all outliers removed instead of using the full sample. Otherwise the test might lose power and lead to incorrect results under the alternative, where estimates using the full sample are inconsistent though they are consistent (efficient) under the null of no outliers. Notice that to obtain test statistics many empirical researchers wrongly estimate $\mathsf{avar}(\widetilde{\beta})$ by using the full sample estimates of $\sigma^{2}$, $M_{\tilde{x} \tilde{x}}$, which might produce $\widehat{H}_{n, c}^{(m + 1)} \le 0$ as $\widehat{\mathsf{avar}}(\widehat{\beta}_{c}^{(m + 1)}) - \widehat{\mathsf{avar}}(\widetilde{\beta}) \le 0$ and make the test meaningless under the alternative, see discussion in detail in our empirical application to Acemoglu et al. (2019) in \S (ref). Thus, in practice it is important to avoid this false procedure by applying the robust estimator of asymptotic variance of difference between ordinary and robustified 2SLS.

Once achieving the estimator $\widehat{\mathsf{avar}}(\widehat{\beta}_{c}^{(m + 1)} - \widetilde{\beta})$ of asymptotic variance by

equation[equation omitted — 357 chars of source]

which is consistent under the null and robust under the alternative, then we have

equation*[equation* omitted — 197 chars of source]

and

equation*[equation* omitted — 278 chars of source]

Thus, a new type of Hausman test for outlier robustness has now been established, and it can either be performed as the one-sided test with the chi-squared limit using estimates of the whole parameter vector or as the two-sided test with the normal limit using the corresponding estimates of individual scalar parameters.

Note that the asymptotic variance in ((ref)) is invertible, since the first scalar term in the expression $(\varrho_{\beta \beta, c}^{(m + 1)} - 1)^{2} + 2 \tau_{2}^{c} (\varrho_{\beta \beta, c}^{(m + 1)} - 1) \varrho_{\beta \tilde{x} u, c}^{(m + 1)} + \tau_{2}^{c} (\varrho_{\beta \tilde{x} u, c}^{(m + 1)})^{2} \neq 0$ and its inverse is

equation*[equation* omitted — 319 chars of source]

We therefore avoid the rank deficiency problem described and addressed through generalized inverses in Hausman and Taylor (1981) and Holly (1982).

On implementing Algorithm (ref) or (ref), we are particularly interested in the trimming estimator just updated from the original 2SLS and in the fixed point estimator iterated upon through infinite steps. The next corollary provides two special cases of the test for outlier robustness frequently used in empirics, checking whether $\widetilde{\beta}$ is distinct from $\widehat{\beta}_{c}^{(1)}$ when $m = 0$ or from $\widehat{\beta}_{c}^{(\ast)}$ when $m = \infty$.

corollaryConsider Algorithm (ref) or (ref). Suppose Assumption (ref)$(ia, iia, iiia, iv)$ holds. For $c \in [c_{+}, \infty)$ and a large $n$ then we have for $m = 0$ \begin{equation*} n^{1/2} (\widehat{\beta}_{c}^{(1)} - \widetilde{\beta}) \overset{a}{\sim} \mathsf{N}\{ 0_{d_{x}}, \widehat{\mathsf{avar}}(\widehat{\beta}_{c}^{(1)} - \widetilde{\beta}) \}, \end{equation*} and \begin{equation*} \widehat{H}_{n, c}^{(1)} = n (\widehat{\beta}_{c}^{(1)} - \widetilde{\beta})^{\prime} \widehat{\mathsf{avar}}(\widehat{\beta}_{c}^{(1)} - \widetilde{\beta})^{-1} (\widehat{\beta}_{c}^{(1)} - \widetilde{\beta}) \overset{a}{\sim} \chi^{2}_{d_{x}}, \end{equation*} where \begin{equation*} \widehat{\mathsf{avar}}(\widehat{\beta}_{c}^{(1)} - \widetilde{\beta}) = \frac{\{2 c \mathsf{f}_{u}(c) - \psi \}^{2} + 2 \tau_{2}^{c} \{2 c \mathsf{f}_{u}(c) - \psi \} + \tau_{2}^{c}}{\psi^{2}} (\widehat{\sigma}_{c}^{(1)})^{2} (\widehat{M}_{\tilde{x} \tilde{x}, c}^{(1)})^{-1}. \end{equation*} In addition, we have for $m = \infty$ \begin{equation*} n^{1/2} (\widehat{\beta}_{c}^{(\ast)} - \widetilde{\beta}) \overset{a}{\sim} \mathsf{N}\{ 0_{d_{x}}, \widehat{\mathsf{avar}}(\widehat{\beta}_{c}^{(\ast)} - \widetilde{\beta}) \}, \end{equation*} and \begin{equation*} \widehat{H}_{n, c}^{(\ast)} = n (\widehat{\beta}_{c}^{(\ast)} - \widetilde{\beta})^{\prime} \widehat{\mathsf{avar}}(\widehat{\beta}_{c}^{(\ast)} - \widetilde{\beta})^{-1} (\widehat{\beta}_{c}^{(\ast)} - \widetilde{\beta}) \overset{a}{\sim} \chi^{2}_{d_{x}}, \end{equation*} where \begin{equation*} \widehat{\mathsf{avar}}(\widehat{\beta}_{c}^{(\ast)} - \widetilde{\beta})\footnote{If $\mathsf{f}_{u} \sim \mathsf{N}(0, 1)$, then $\tau_{2}^{c} = \psi - 2 c \mathsf{f}_{u}(c)$ so $\widehat{\mathsf{avar}}(\widehat{\beta}_{c}^{(1)} - \widetilde{\beta})$ and $\widehat{\mathsf{avar}}(\widehat{\beta}_{c}^{(\ast)} - \widetilde{\beta})$ can be further simplified.} = \frac{ \{ 2 c \mathsf{f}_{u}(c) - \psi \}^{2} + 2 \tau_{2}^{c} \{ 2 c \mathsf{f}_{u}(c) - \psi \} + \tau_{2}^{c} }{\{ 2 c \mathsf{f}_{u}(c) - \psi \}^{2}} (\widehat{\sigma}_{c}^{(\ast)})^{2} (\widehat{M}_{\tilde{x} \tilde{x}, c}^{(\ast)})^{-1}. \end{equation*}

Without knowing the asymptotic theory derived in this section, some empirical researchers simply use the similar form of the Hausman test statistics but incorrectly replace the standard error of the difference of two estimates $\widetilde{\beta}$ and $\widehat{\beta}_{c}^{(m + 1)}$ by the standard error of the marginal distribution of the baseline estimate $\widetilde{\beta}$. They subsequently compare this test statistics with the critical value either drawn from the standard normal for the two-sided test or from the chi-square with $d_{x}$ degree of freedom for the one-sided test. The above test procedure is commonly referred to as the heuristic method. Since the heuristic test statistics is constructed in a mistaken way, its size does not converge to the nominal level under the null and it has low statistical power under the alternative. Thus, even the heuristic method is informative for preliminary analysis, it is not consistent in theory and not reliable in practice.

Kaji (2018) applies the non-parametric bootstrap by randomly sampling observations $i$ from the data $\{(y_{i}, x_{i}, z_{i})\}_{i = 1}^{n}$ with replacement to draw the distribution of the $L_{1}$ norm of the difference between $\widehat{\beta}_{c}^{(m + 1)}$ and $\widetilde{\beta}$. To prove the validity of the bootstrap, he first characterises the weak convergence of a class of integrable empirical processes and then build the functional delta method for the corresponding quantile processes. Unlike the asymptotic theory in this paper, the bootstrap in Kaji (2018) does not require assuming the known form of the density $\mathsf{f}_{u}$. However, if the specification assumed on $\mathsf{f}_{u}$ is correct, our Hausman type test would have the higher power than the bootstrapping method under the alternative where outliers in fact present. This could be because under the alternative the bootstrap has a certain probability to sample outlying observations that have large influence on $\beta$ estimation. For these bootstrapping iterations having sampled large outliers from the original data, the test statistics on difference between the robust estimate $\widehat{\beta}_{c}^{(m + 1)}$ and the baseline estimate $\widetilde{\beta}$ would then become rather large, so the distribution of the test statistics drawn from the bootstrap sample would be highly distorted and significantly distinct from what it should be under the null of no outliers.

Simulation studies

In this section, we verify the established asymptotic theory and check their finite sample performance using Monte Carlo simulations. This paper uses a DGP where the structural equation contains an intercept and one endogenous regressor, which is instrumented by one excluded and informative instrument, such that we are in the just-identified case. The structural equation is given as

equation*[equation* omitted — 97 chars of source]

The first stage projection is

equation*[equation* omitted — 267 chars of source]

Note $x_{1i} = 1$ for all $i$, and we choose $\beta_{1} = 2$, $\beta_{2} = 4$ and $\pi_{1} = 0$, $\pi_{2} = 1$. Simulated data $\{ (y_{i}, x_{i}, z_{i}) \}_{i = 1}^{n}$ can then be generated by drawing random samples from the Normal

equation*[equation* omitted — 374 chars of source]

We choose the variances of $u_{i}$, $r_{2i}$, $z_{2i}$ as $\sigma^{2} = 1$, $\Sigma_{2} = 1$, $\sigma_{z_{2}}^{2} = 1$ and covariance between $u_{i}$ and $r_{2i}$ as $\Omega_{2}$ = 0.75, which is the source of endogeneity. The sample size is varied from $50$ to $5000$ to assess the finite sample performance and to check the correctness of the asymptotic theory, i.e. $n \in \{ 50, 100, 200, 500, 1000, 5000 \}$. We set $W = 10000$ replications for our Monte Carlo experiments.

We run Algorithm (ref) with the choices of the initial estimator as the full-sample 2SLS, the cut-off value $1.96$, and the iterations $m \in \{ 1, 5, \textrm{Fixed point} \}$. As a result, for each Monte Carlo replication $w = 1, 2, \ldots, W$, the algorithm produces the $\beta$-estimators $\widehat{\beta}_{1, c, w}^{(m)}$, $\widehat{\beta}_{2, c, w}^{(m)}$ and $\sigma^{2}$-estimators $(\widehat{\sigma}_{c, w}^{(m)})^{2}$ on the simulated data $\{ (y_{i}, x_{i}, z_{i}) \}_{i = 1, w}^{n}$.

figure[figure omitted — 354 chars of source]
figure[figure omitted — 321 chars of source]

Evaluation of asymptotic distribution of $\beta$-estimators

Our asymptotic theory shows

equation*[equation* omitted — 343 chars of source]

The asymptotic variance is given by $(\vartheta_{c}^{(m)})^{-1} \sigma^{2} M_{\tilde{x} \tilde{x}}^{-1}$, where in our simulation specification $\sigma^{2} = 1$ and $M_{\tilde{x} \tilde{x}} = \mathsf{E} \tilde{x}_{i} \tilde{x}_{i}^{\prime} = \Pi^{\prime} \mathsf{E} z_{i} z_{i}^{\prime} \Pi = I_{2}$ as $z_{i}^{\prime} = (1, z_{2i})$ and $\Pi = I_{2}$. This asymptotic distribution is then evaluated by comparing the simulated distribution of $\{ n^{1/2} (\widehat{\beta}_{j, c, w}^{(m)} - \beta_{j}) / \sqrt{(\vartheta_{c}^{(m)})^{-1}} \}_{w = 1}^{W}$ for $j = 1, 2$ to $\mathsf{N}(0, 1)$ when the sample size $n$ increases. Figure (ref) plots kernel density estimates of the simulated distributions and the probability density function of the standard normal distribution as the asymptotic distribution. The graphs show that the standardised $\beta$-estimators converges to a standard normal distribution and that the approximation is very close for the sample size $n = 50$ and larger regardless of the iterations steps $m$. This indicates that the derived relative efficiency factor $\vartheta_{c}^{(m)}$ is consistent with the simulation studies.

Evaluation of consistency of $\sigma^{2}$-estimators

Our asymptotic theory shows $(\widehat{\sigma}_{c}^{(m)})^{2} \overset{\mathsf{P}}{\to} \sigma^{2}$. We assess the consistency by checking the difference between the simulated estimates $\{ (\widehat{\sigma}_{c, w}^{(m)})^{2} \}_{w = 1}^{W}$ and the population one $\sigma^{2} = 1$ when the sample size $n$ increases. Figure (ref) plots means, $97.5 \%$ quantiles, $2.5 \%$ quantiles of the simulated estimates and the true variance of structural errors. The graphs show that $\sigma^{2}$-estimators converges to the population one and that the deviation is already very small for the sample size $n = 50$ regardless of the iterations steps $m$. This indicates that the simulations lend support to the correctness of the derived bias correction factor $\varsigma_{c}^{2}$.

Empirical application to Acemoglu et al. (2019)

This paper now applies the robust procedure proposed in \S (ref) and its theory to recarry out outlier analysis in Acemoglu et al. (2019). Their paper tries to address the long-standing debate of whether democracy does cause economic growth through an empirical study. Acemoglu et al. (2019) find evidence that democracy has a significant and robust positive effect on GDP per capita and further suggest that democracy increases GDP by encouraging investment, increasing schooling, inducing economic reforms and openness, improving the provision of public goods, and reducing social unrest.

To investigate the causal effect of democracy on GDP per capita, Acemoglu et al. (2019) collect an annual panel that comprises 175 countries around the world from the year 1960 to 2010. The panel is unbalanced in the sense that not all observations of GDP and democracy are available for the entire sample. In order to reduce measurement error, they create a consolidated and dichotomous measure of democracy from several datasets and document robustness of their results to other measures of GDP and democracy.

Acemoglu et al. (2019) conduct three specifications that control for the rich dynamics of GDP and endogenous selection into democracy, which otherwise confound the effect of democracy on economic growth. Three specifications consist a dynamic panel model for log GDP per capita, a semi-parametric treatment effects model which does not rely on a parametric assumption for GDP dynamics, the same dynamic panel model but instrumenting democracy by regional waves of democratizations and reversals. We focus on the first and third specification (a dynamic panel model) in Acemoglu et al. (2019) where they carry out outlier analysis for OLS and 2SLS fixed effect estimates.

Their baseline preferred specification constructs a dynamic panel that models log GDP per capita as outcome variable, democracy as treatment variable, and four lags of log GDP per capita as other covariates to control GDP dynamics. They start with running the fixed effect regression implemented as OLS estimator by adding dummy variable regressors for each country and year. Although the fixed effect estimator suffers the Nickell (1981) bias of order one over time horizon resulting from the failure of strict exogenity in dynamic panel model, this bias should be small in their setup where the time dimension of panel is fairly large (each country is on average observed 38.8 times). This motivates their use of the OLS fixed effect estimator as a natural starting point even in dynamic panel models. Acemoglu et al. (2019) then apply Arellano and Bond (1991) GMM estimator that deal with Nickell bias, and further use an even more refined Hahn, Hausman, and Kuersteiner (2002) GMM estimator that address “too many instruments” problem faced by Arellano and Bond estimator. Consistent with their expectation, both type of GMM estimates are similar to the OLS fixed effec estimate such that the coefficient on democracy variable is estimated to be positive and highly significant.

Their third specification is identical to the dynamic panel model above, but they treat democracy as endogenous and exploit IVs strategy by using regional waves of democratization and transitions to nondemocracy as source of exogenous variation in democracy. The instrument regional waves of democratization is constructed as the jack-knifed average of democracy for countries in the same region with the same initial regime cell, which leaves out the own-country observation. Although there is no consensus on the factors that create such waves, the existing evidence suggests that they are not explained by regional economic trends. Thus, Acemoglu et al. (2019) argue that regional waves are valid as instruments such that their exclusion restriction assumption is plausible. They further show from the first stage regression that regional waves are informative such that rank conditions hold for identification. By applying 2SLS fixed effect and Hahn, Hausman, and Kuersteiner (2002) estimator, they reassure that the causal effect of democracy on log GDP per capita is significant and positive.

Three empirical strategies tackle different possibilities and find very similar results, which bolsters their confidence to demonstrate the positive causal effect of democracy. To defend against the critique that their empirical findings are driven by outliers, Acemoglu et al. (2019) explores the sensitivity of their results to outliers for their baseline OLS and 2SLS fixed effect regression.

For each observation in the model ((ref)) and ((ref)), we now have the dependent variable, regressors, and instruments

align*[align* omitted — 501 chars of source]

In regressors, only $\mathrm{Democracy}_{i, t}$ is treated as endogenous, while the rest of $x_{i, t}$ are assumed to be orthogonal to structural errors. Then the structural and first stage regression parameters are

align*[align* omitted — 1,113 chars of source]

In the structural parameter vector $\beta$, we are particulary interested in $\beta_{0}$ on the effect of democracy as well as $\beta_{1}, \ldots, \beta_{4}$ indicating the dynamics of the GDP process\footnote{Aside from structural parameters $\beta_{0}, \beta_{1}, \ldots, \beta_{4}$, Acemoglu et al. (2019) report three more parameters: $(1)$ $\beta_{5} := \beta_{0} / (1 - \beta_{1} - \beta_{2} - \beta_{3} - \beta_{4})$ (long-run equilibrium effect of democracy); $(2)$ $\beta_{6} := e_{25}$, where $e_{j} = \beta_{0} + \beta_{1} e_{j - 1} + \beta_{2} e_{j - 2} + \beta_{3} e_{j - 3} + \beta_{4} e_{j - 4}$ and $e_{0} = e_{-1} = e_{-2} = e_{-3} = 0$ (effect of transition to democracy after 25 years); $(3)$ $\beta_{7} := \beta_{1} + \beta_{2} + \beta_{3} + \beta_{4}$ (persistence of the GDP process).}.

Implicitly assuming normal distribution for structural errors, i.e. $\mathsf{f}_{u} \overset{\mathsf{D}}{=} \mathsf{N}(0, 1)$, Acemoglu et al. (2019) choose the cut-off value $c = 1.96$ for $5 \%$ significance level such that outliers are identified among observations whose absolute values of standardized residuals of structural errors are beyond $1.96$. They then perform the same OLS and 2SLS fixed effect regression but exclude these outlying observations\footnote{This is the first robust method that Acemoglu et al. (2019) apply, and its results are reported in column $(2)$, Table A$8$ for OLS fixed effect regression; and in column $(2)$, Table A$13$ for 2SLS fixed effect regression in their paper. They also use other robust methods, such as Cook's distance, Li's (1985) procedure, and Huber M-estimator; and take into account the influence of outliers in both the first and second stage; see their reported results in other columns in Table A8, A13.}. Investigating the difference between the ordinary (full sample) and robust (sub-sample without outliers) estimates, they argue robustness of their results to outliers.

The above robust method is in fact Algorithm (ref) (Robustified 2SLS), which select the cut-off $c = 1.96$ $(5 \%)$ by assuming $\mathsf{f}_{u} \overset{\mathsf{D}}{=} \mathsf{N}(0, 1)$ and iterates only once by choosing $m = 0$. However, Acemoglu et al. (2019) do not take into consideration the case that observations are classified as outliers and removed just by chance, so consistency does not hold for the updated estimator of the structural error variance if without adjusting the bias correction factor ((ref)), although estimation of structural parameters is still consistent. In addition, by not adjusting standard errors calculated from the usual asymptotoics that does not have the correct asymptotic variance of robust estimates, inference on regression coefficients is invalid. At last, without a formal statistical test we do not know whether the robust estimate is significantly distinct from the ordinary one, thus even when we end up with the robustness conclusion, our argument is not rooted in a correct statistical reasoning.

To solve all these problems, this paper applies Algorithm (ref) to re-explore outlier robustness analysis in Acemoglu et al. (2019). As in their paper, we first compute the full sample 2SLS $\widetilde{\beta}$, then choose the cut-off $c = 1.96$ to run a single iteration $m = 0$ of the algorithm to obtain $\widehat{\beta}_{1.96}^{(1)}$. Besides selecting $m = 0$, we continue to run the algorithm upon through infinite iterations $m = \infty$ until reaching the fixed point $\widehat{\beta}_{1.96}^{(\ast)}$. In adjusting the standard errors of $\widehat{\beta}_{1.96}^{(1)}$ and $\widehat{\beta}_{1.96}^{(\ast)}$ by the factors listed in Table (ref), this paper performs valid inference on structural parameters $\beta$ and conducts the new Hausman type tests to check whether $\widetilde{\beta}$ is different from $\widehat{\beta}_{1.96}^{(1)}$ or $\widehat{\beta}_{1.96}^{(\ast)}$. Our re-examination of Acemoglu et al. (2019) finds that the effect of democracy estimated with the full sample $\widetilde{\beta}_{0}$ is distinct in magnitude from the robust estimates $\widehat{\beta}_{0, 1.96}^{(1)}$ or $\widehat{\beta}_{0, 1.96}^{(\ast)}$ produced by Algorithm (ref), although we find little evidence for the view that democracy is a constraint on economic growth ($\beta_{0} < 0$).

Tables (ref) and (ref) list results of the OLS and 2SLS fixed effect regressions together with formal statistical tests for the hypothesis that outliers have no effect on parameters. The first three columns show the OLS (2SLS) fixed effect and robust estimates in Table (ref) (Table (ref)). Select the cut-off value $c = 1.96$ and set the iteration step $m = -1$, $m = 0$, $m = \infty$, then run Algorithm (ref) to produce three estimates $\widetilde{\beta}$, $\widehat{\beta}_{1.96}^{(1)}$, $\widehat{\beta}_{1.96}^{(\ast)}$. In Column $1$, $\widetilde{\beta}$ represents the full sample OLS (2SLS) estimate in Table (ref) (Table (ref)). In Column $2$, we remove observations whose absolute values of standardized residuals are beyond the cut-off $1.96$ (in the second stage) and re-compute OLS (2SLS) to obtain the robust estimate $\widehat{\beta}_{1.96}^{(1)}$ in Table (ref) (Table (ref)). In Column $3$, this procedure is iterated until the fixed point $\widehat{\beta}_{1.96}^{(\ast)}$ is attained. Values reported below the estimates in parentheses $()$ are standard errors\footnote{Acemoglu et al. (2019) also report standard errors robust against heteroskedasticity and serial correlation at the country level in their Table A$8$, A$13$.}. Note that in Columns $2$ and $3$ in Tables (ref) and (ref) we report the adjusted standard errors\footnote{Acemoglu et al. (2019) do not take the adjustments to the standard errors of $\widehat{\beta}_{1.96}^{(1)}$ reported in Column $(2)$ in their Table A$8$, A$13$.} by applying the theory in \S (ref). In Table (ref) (Table (ref)), ordinary standard errors are multiplied by $\sqrt{6044 / 6336}$ ($\sqrt{6019 / 6309}$) and $\sqrt{\iota_{1.96}^{(1)}}$ in Column $2$; and by $\sqrt{5213 / 6336}$ ($\sqrt{5202 / 6309}$) and $\sqrt{\iota_{1.96}^{(\ast)}}$ in Column $3$. Table (ref) in \S (ref) shows $\iota_{1.96}^{(1)} = 1.612$ and $\iota_{1.96}^{(\ast)} = 1.828$, so in order to produce the reported standard errors in Table (ref) (Table (ref)) we apply the adjusting factor $1.240$ ($1.240$) in Column $2$ and $1.226$ ($1.228$) in Column $3$.

In Tables (ref) and (ref), Columns $4 - 7$ present formal statistical tests for outlier robustness analysis. Columns $4 - 5$ give test results for the null hypothesis that the two parameters estimated by $\widetilde{\beta}$ and $\widehat{\beta}_{1.96}^{(1)}$ are identical. Columns $6 - 7$ present test results for the same hypothesis but comparing $\widetilde{\beta}$ with $\widehat{\beta}_{1.96}^{(\ast)}$. Columns $5$ and $7$ show the new Hausman type test statistics $\widehat{H}_{1.96}^{(1)}$ and $\widehat{H}_{1.96}^{(\ast)}$ that use the standard errors of the difference of two estimators $\widehat{\mathsf{avar}} (\widehat{\beta}_{1.96}^{(1)} - \widetilde{\beta})$ and $\widehat{\mathsf{avar}} (\widehat{\beta}_{1.96}^{(\ast)} - \widetilde{\beta})$\footnote{Consider $c = 1.96$ and $m = 0, \infty$ and note $\mathsf{avar}(\widehat{\beta}_{c}^{(m + 1)} - \widetilde{\beta}) = \mathsf{avar}(\widehat{\beta}_{c}^{(m + 1)}) - \mathsf{avar}(\widetilde{\beta})$. To construct the standard errors of the difference of two estimators, some empirical researchers estimate $\mathsf{avar}(\widetilde{\beta})$ using the full sample and estimate $\mathsf{avar}(\widehat{\beta}_{c}^{(m + 1)})$ using the sub-sample with all outliers removed. This produces the negative estimate of $\mathsf{avar}(\widehat{\beta}_{c}^{(m + 1)} - \widetilde{\beta})$ in this application as $\widehat{\mathsf{avar}}(\widehat{\beta}_{c}^{(m + 1)}) - \widehat{\mathsf{avar}}(\widetilde{\beta}) \le 0$ clearly shown in Tables (ref) and (ref), making the Hausman test meaningless. As suggested in \S (ref), the estimator of $\mathsf{avar}(\widehat{\beta}_{c}^{(m + 1)} - \widetilde{\beta})$ should be robust itself to ensure the power performance under the alternative. In other words, we need to robustly estimate $\mathsf{avar}(\widetilde{\beta})$ using the sub-sample of all non-outlying observations.}, see Corollary (ref) in \S (ref) for construction of the Hausman test statistics. We first apply the two-sided type Hausman test with the normal limit for each individual scalar parameter $\beta_{0}$, $\beta_{1}$, $\ldots$, and $\beta_{4}$, then perform the one-sided test with the chi-squared limit for the whole parameter vector $(\beta_{0}, \beta_{1}, \ldots, \beta_{4})^{\prime}$. Columns $4$ and $6$ show the heuristic test statistics $\widehat{\mathrm{Heu}}_{1.96}^{(1)}$ and $\widehat{\mathrm{Heu}}_{1.96}^{(\ast)}$ for the same hypothesis, which however use the standard error of the marginal distribution of the baseline full sample estimate $\widetilde{\beta}$. P-values are reported in brackets $[]$ below test statistics.

table[table omitted — 3,955 chars of source]
table[table omitted — 3,984 chars of source]

We see from Tables (ref) and (ref) that $\beta_{0}$ on democracy is estimated by 2SLS to have the larger value than OLS fixed effect regressions while the effect of GDP dynamics remains roughly unchanged, which is consistent with our expectation that the IVs strategy affects mostly the coefficient of the endogenous variable democracy. The standard errors of the estimates $\widetilde{\beta}$, $\widehat{\beta}_{1.96}^{(1)}$, $\widehat{\beta}_{1.96}^{(\ast)}$ decrease with iteration, suggesting that it is under the alternative hypothesis where the data is contaminated by outlying observations driving up the standard error of the full sample estimate $\widetilde{\beta}$. Otherwise under the null of no outliers standard errors of $\widehat{\beta}_{1.96}^{(1)}$, $\widehat{\beta}_{1.96}^{(\ast)}$ should become larger than $\widetilde{\beta}$, since by gaining robustness the estimates $\widehat{\beta}_{1.96}^{(1)}$, $\widehat{\beta}_{1.96}^{(\ast)}$ lose efficiency relative to $\widetilde{\beta}$, leading to the larger asymptotic variance. The number of observations used in the OLS and 2SLS estimates are $6336$ and $6309$. Even when the model contains no outliers, Algorithm (ref) throws away around $5 \%$ of the whole observations with a large probability for the cut-off value $1.96$ to produce $\widehat{\beta}_{1.96}^{(1)}$ and $\widehat{\beta}_{1.96}^{(\ast)}$. Thus, under the null of no outliers the expected number of observations retained by Algorithm (ref) should be around $6019$ and $5994$ that are sub-sample sizes of $95 \%$ of the full OLS and 2SLS sample. We notice that the observed number of observations retained by $\widehat{\beta}_{1.96}^{(\ast)}$ are much less than the expected number for both of the OLS and 2SLS fixed effect regressions, although the sub-sample size of $\widehat{\beta}_{1.96}^{(1)}$ for OLS and 2SLS are not far away from the expected number. This fact again indicates that the model does contain outliers.

figure[figure omitted — 600 chars of source]

By only computing point estimates $\widetilde{\beta}$, $\widehat{\beta}_{1.96}^{(1)}$, $\widehat{\beta}_{1.96}^{(\ast)}$ for the OLS or 2SLS fixed effect regressions without applying formal statistical tests of outlier robustness, we cannot draw any conclusion on whether or not the empirical findings are driven by outliers, since the information contained in the standard errors also needs to be considered in order to evaluate the distance between ordinary and robust estimates. For either the simple regression in Table (ref) or IVs specification in Table (ref), both of the heuristic and Hausman tests reject the null hypothesis that the parameter vector $\beta$ estimated by $\widetilde{\beta}$ is identical to the robust estimate $\widehat{\beta}_{1.96}^{(1)}$ or $\widehat{\beta}_{1.96}^{(\ast)}$. It is obvious to note from Columns $4 - 7$ in Tables (ref) and (ref) that the Hausman tests yield strictly the lower p-values than the heuristic tests, which is concordant with our theory in \S (ref) that the Hausman test has the better power performance than the heuristics. In the simple regressions, the Hausman tests reject the null of outlier robustness at the $5 \%$ significance level for most coefficients except $\beta_{2}$, $\beta_{3}$ when compare $\widetilde{\beta}$ with $\widehat{\beta}_{1.96}^{(1)}$; except $\beta_{3}$ when compare $\widetilde{\beta}$ with $\widehat{\beta}_{1.96}^{(\ast)}$. In the IVs regressions, the Hausman tests mostly give the same result but the identity of $\beta_{0}$ is not rejected by comparing $\widetilde{\beta}$ with $\widehat{\beta}_{1.96}^{(1)}$.

table[table omitted — 2,262 chars of source]

Kaji (2018) also re-examines outlier robustness of the results found in Acemoglu et al. (2019). Without the weak limit of the Hausman type test statistics derived in our paper, he carries out the non-parametric bootstrap to draw the distribution of the $L_{1}$ norm of the difference between the baseline full sample estimate and the robust estimate. The bootstrap is implemented by randomly sampling countries $i$ with replacement. Each draw of country $i$ adds its time dimensional number of observations to the bootstrap sample. The bootstrap consists of $10000$ iterations. By comparing between the baseline OLS estimate $\widetilde{\beta}$ and outlier-removed estimate $\widehat{\beta}_{1.96}^{(1)}$, the bootstrap does not reject the null that two estimates are identical, indicating robustness to outliers of the empirical findings in Acemoglu et al. (2019). The reason why the null of outlier robustness is not rejected by Kaji (2018) could be the low power issue faced by the bootstrap under alternatives; see discussion in the last paragraph in \S (ref). Note that Kaji (2018) only applies the tests for checking the difference for each coefficient $\beta_{0}, \beta_{1}, \ldots, \beta_{4}$, but not for the whole vector $\beta$.

Figure (ref) plots the estimated IV regression results on democracy over iterations (with the cut-off value $1.96$). The upper panel reports the point estimates, the ordinary and adjusted $95 \%$ confidence intervals. The lower panel reports the p-values of the Hausman-type tests of comparing between the full-sample and outlier removal estimates. The red line shows that the iterated procedure converges at the iteration step $16$. We can find that the p-values of the tests become smaller (lower than the conventional level $ 0.05$) with increasing iterations. It indicates that we should reply on the trimmed estimates to conduct inference. The outlier removal estimates become closer to zero and the adjusted $95 \%$ confidence intervals contain zero for large iteration numbers.

Table (ref) focuses on $\beta_{0}$ the effect of democracy on log GDP per capita by summarizing regression results from Tables (ref) and (ref). We present estimates of $\beta_{0}$ and compare standard errors with and without the adjustments together with their corresponding inference including $95 \%$ confidence intervals and p-values of the $t$-test for the null hypothesis that $\beta_{0} = 0$. The adjusted standard errors and inference results are reported in parentheses $()$ after the ordinary ones.

We see from Table (ref) that the modified standard errors are uniformly larger than those without taking the adjustments, so the adjusted confidence bands are uniformly wider and p-values are uniformly larger; see the theory in \S (ref). Recall that the baseline full sample and robust estimates of $\beta_{0}$ are statistically different, indicated by the Hausman tests. In the OLS regressions, inference based on $\widehat{\beta}_{0, 1.96}^{(\ast)}$ shows that $\beta_{0}$ is not significantly distinct from $0$. In the 2SLS regressions, we cannot reject $\beta_{0} = 0$ at the $5 \%$ significance level by performing the adjusted inference either based on the baseline estimate $\widetilde{\beta}$ or the robust estimates $\widehat{\beta}_{0, 1.96}^{(1)}$ and $\widehat{\beta}_{0, 1.96}^{(\ast)}$. Thus, we cannot find the evidence robustly supporting the result $\beta_{0} > 0$ that the effect of democracy is strictly positive on economic growth.

Weighted and marked empirical process

Consider the weighted and marked empirical distribution function

equation[equation omitted — 264 chars of source]

with $\mathcal{F}_{i - 1}$ adapted weights $w_{in}$ and $\mathcal{F}_{i}$ measurable marks $u_{i}^{p}$. The filtration $\mathcal{F}_{i - 1}$ is generated by $(z_{1}, \ldots, z_{i}, r_{1}, \ldots, r_{i - 1}, u_{1}, \ldots, u_{i - 1})$ so $z_{in} \in \mathcal{F}_{i - 1}$ while $u_{i}, r_{i} \in \mathcal{F}_{i}$ are independent of $\mathcal{F}_{i - 1}$. The observable data $\{(y_{i}, x_{i}, z_{i})\}_{i = 1}^{n}$ has i.i.d. structure in this paper, so $a \in\mathbb{R}$, $b \in \mathbb{R}^{d_{x}}$ represent normalized estimation errors $\widetilde{a} = n^{1/2} (\widetilde{\sigma} - \sigma)$, $\widetilde{b} = n^{1/2} (\widetilde{\beta} - \beta)$, and $c \in \mathbb{R}$ is the quantile, while $\sigma$, $\Pi$, $\Sigma$ are true parameters for the variance of the structural error, for the location coefficient in the first stage regression, and for the variance of the first stage error respectively. Let normalized instruments $z_{in} = n^{-1/2} z_{i}$ such that $M_{z z, n} = \sum_{i=1}^{n} z_{in} z_{in}^{\prime} = n^{-1} \sum_{i=1}^{n} z_{i} z_{i}^{\prime}$ converges in probability to $\mathsf{E} z_{i} z_{i}^{\prime} = M_{z z}$ by Law of Large Numbers when $\mathsf{E} |z_{i}|^{2} < \infty$. Our interest focuses on weights $w_{in}$ given as either of $1$, $n^{1/2} z_{in} = z_{i}$, $n z_{in}z_{in}^{\prime} = z_{i} z_{i}^{\prime}$ and $p$ as either of $0$, $1$, $2$. To form the empirical process, introduce the compensator

equation[equation omitted — 262 chars of source]

where $\mathsf{E}_{i-1}(\cdot) = \mathsf{E}( \cdot | \mathcal{F}_{i-1})$. Note that $\overline{\mathsf{F}}_{u, n}^{1, 0}(0,0,c) = \mathsf{F}_{u}(c) = \mathsf{P}(u_{i} \le \sigma c)$.

We embed these processes into the space $D[0,1]$ that are processes continuous from the right and with limits of left, where the space is endowed with the Skorokhod metric. We do this as follows. The indicator $1_{(u_i \le \sigma c)}$ and the distribution function $\mathsf{F}_{u}(c)$ can be defined as 0 or 1 when $c$ takes the values $-\infty$ and $\infty$ respectively. We can then define quantiles $c_\psi=\mathsf{F}_{u}^{-1}(\psi)$ for $0 \le \psi \le 1$. Correspondingly we can continously extend the definition of the weighted and marked empirical distribution function and its compensator by chosing $\widehat{\mathsf{F}}_{u, n}^{w, p}(a, b, -\infty) =\overline{\mathsf{F}}_{u, n}^{w, p}(a, b, -\infty) = 0$ while we then have $\widehat{\mathsf{F}}_{u, n}^{w, p}(a, b, \infty) = n^{-1}\sum_{i=1}^{n} w_{in} u_{i}^{p}$ and $\overline{\mathsf{F}}_{u, n}^{w, p}(a, b, \infty) = n^{-1} \sum_{i=1}^{n} w_{in}\mathsf{E}_{i-1} u_{i}^{p}$. We now define the empirical process for $0 \le \psi \le 1$

equation[equation omitted — 228 chars of source]

In the following we first present assumptions, and then results follow for the one-sided process and finally extend to the absolute case.

Assumptions

In this section, the density $\mathsf{f}_{u}$ is not necessarily symmetric. The conditions listed below are weaker than Assumption (ref).

assumptionLet $\mathcal{F}_{i - 1} = \sigma(z_{1}, \ldots, z_{i}, r_{1}, \ldots, r_{i - 1}, u_{1}, \ldots, u_{i - 1})$ be an increasing sequence of $\sigma$-fields so $u_{i-1}$, $r_{i - 1}$, $z_{i}$, $w_{in}$ are $\mathcal{F}_{i-1}$ measurable while $r_{i}$, $u_{i}$ are independent of $\mathcal{F}_{i - 1}$. Suppose $(u_{i}/\sigma, \Sigma^{-1/2} r_{i})$ have continuously differentiable joint, conditional, and marginal densities $\mathsf{f}_{u, r}(y, x) = \mathsf{f}_{u|r}(y|x) \mathsf{f}_{r}(x) = \mathsf{f}_{r|u}(x|y) \mathsf{f}_{u}(y)$ which are positive on $y \in \mathbb{R}$, $x \in \mathbb{R}^{d_{x}}$. Let $p, \eta, \kappa$ be given so $p \in \mathbb{N}_{0}$, $0 \le \kappa < \eta \le 1/4$. Choose $0 < \nu \le 1$, $s \in \mathbb{N}_{0}$ such that \begin{equation} 2^{s - 1} > 1 + (1/4 - \eta)(1 + d_{x}). \end{equation} Suppose \\ $(i)$ the marginal density $\mathsf{f}_{u}(y)$ satisfies for $y \in \mathbb{R}$ \\ $(a)$ moments: $\int_{-\infty}^{\infty} |y|^{2^{s}p / \nu} \mathsf{f}_{u}(y) dy < \infty$; \\ $(b)$ boundedness: $\sup_{y \in \mathbb{R}} |y^{2^{s} p + 1} \mathsf{f}_{u}(y) + y^{2^{s} p + 2} \dot{\mathsf{f}}_{u}(y)| < \infty$; \\ $(c)$ smoothness: a $C_{\mathrm{H}} \in \mathbb{N}$ exists so that for all $\epsilon > 0$ \begin{equation*} \frac{\sup_{y \ge \epsilon} (1 + y^{2^{s}p})\mathsf{f}_{u}(y)}{\inf_{0 \le y \le \epsilon}(1 + y^{2^{s}p})\mathsf{f}_{u}(y)} \le C_{\mathrm{H}}, \qquad \frac{\sup_{y \le -\epsilon} (1 + y^{2^{s}p})\mathsf{f}_{u}(y)}{\inf_{-\epsilon \le y \le 0}(1 + y^{2^{s}p})\mathsf{f}_{u}(y)} \le C_{\mathrm{H}}; \end{equation*} $(ii)$ the marginal density $\mathsf{f}_{r}(x)$ satisfies for $x \in \mathbb{R}^{d_{x}}$ \\ $(a)$ moments: $\int_{x \in \mathbb{R}^{d_{x}}} |x|^{4} \mathsf{f}_{r}(x) (dx) < \infty$; \\ $(iii)$ the joint and conditional densities $\mathsf{f}_{u, r}(y, x), \mathsf{f}_{u|r}(y|x)$ satisfy for $y \in \mathbb{R}, x \in \mathbb{R}^{d_{x}}$ \\ $(a)$ boundedness: $\sup_{y \in \mathbb{R}, x \in \mathbb{R}^{d_{x}}} |(1 + y) y^{2^{s} p - 1} \mathsf{f}_{u|r}(y|x) + y^{2^{s} p} \dot{\mathsf{f}}_{u|r, y}(y|x)| < \infty$; \\ $(iv)$ the instruments $z_{i}$ satisfy \\ $(a)$ $\max_{1 \le i \le n} |n^{1/2 - \kappa} z_{in}| = \mathrm{O}_{\mathsf{P}}(1)$; \\ $(v)$ the weights $w_{in}$ satisfy \\ $(a)$ $n^{-1} \mathsf{E} \sum_{i = 1}^{n} |w_{in}|^{2^{s}} (1 + |n^{1/2} z_{in}|) = \mathrm{O}(1)$; \\ $(b)$ $n^{-1} \sum_{i = 1}^{n} |w_{in}| (1 + |n^{1/2} z_{in}|^{2}) = \mathrm{O}_{\mathsf{P}}(1)$.
remarkAssumption (ref)$(ia, iia, iiia, ivb, ivc)$ implies Assumption (ref) with $s \ge 2$ satisfying ((ref)) when $w_{in}$ is either of $1$, $n^{1/2} z_{in} = z_{i}$, $n z_{in} z_{in}^{\prime} = z_{i} z_{i}^{\prime}$ and $p$ is either of $0$, $1$, $2$. Details are given in Lemma (ref) in the appendix.

Asymptotic results for empirical process

We present three asymptotic results. The first theorem shows that the estimation error for the scale and regression parameters is negligible uniformly in the quantile.

theoremLet $c_{\psi} = \mathsf{F}_{u}^{-1}(\psi)$. Suppose Assumption (ref) holds with $\nu = 1$ and $s \ge 2$ satisfying ((ref)). Then for any $B > 0$ and as $n \to \infty$ \begin{equation*} \sup_{0 \le \psi \le 1} \sup_{|a|, |b| \le n^{1/4 - \eta}B} |\mathbb{F}_{u, n}^{w, p}(a, b, c_{\psi}) - \mathbb{F}_{u, n}^{w, p}(0, 0, c_{\psi})| = \mathrm{o}_{\mathsf{P}}(1). \end{equation*}

The proof involves a chaining argument. For this, we apply an iterated martingale inequality, see Lemma (ref), to explore the tail behaviour of the maximum of a family of martingales. We can, however, split this argument into two parts. First, we keep $b$ fixed and consider variation in $a, c_{\psi}$. Then we keep $a$ fixed and consider variation in $b, c_{\psi}$. Combining these two arguments will finally finish the proof.

The second result provides a linearization of the compensator.

theoremLet $c_{\psi} = \mathsf{F}_{u}^{-1}(\psi)$. Suppose Assumption (ref)$(ia, ib, iia, iiia, vb)$ holds with $\nu = 1$, $s = 0$. Then for any $B > 0$ and as $n \to \infty$ \begin{equation*} \sup_{0 \le \psi \le 1} \sup_{|a|, |b| \le n^{1/4 - \eta}B} |n^{1/2} \{ \overline{\mathsf{F}}_{u, n}^{w, p}(a, b, c_{\psi}) - \overline{\mathsf{F}}_{u, n}^{w, p}(0, 0, c_{\psi}) \} - \mathcal{B}_{\mathsf{F}_{u}, n}(a, b, c_{\psi})| = \mathrm{O}_{\mathsf{P}}(n^{-2\eta}), \end{equation*} where the bias term, with $\xi_{c_{\psi}} = \mathsf{E} (\Sigma^{-1/2} r_{i} | u_{i}/\sigma = c_{\psi})$, is defined as \begin{equation*} \mathcal{B}_{\mathsf{F}_{u}, n}(a, b, c_{\psi}) = \sigma^{p - 1} c_{\psi}^{p} \mathsf{f}_{u}(c_{\psi}) n^{-1/2} \sum_{i=1}^{n} w_{in} ( n^{-1/2} a c_{\psi} + n^{-1/2} \xi_{c_{\psi}}^{\prime} \Sigma^{1/2} b + z_{in}^{\prime} \Pi b ). \end{equation*}

Since $x_{i}$ and $u_{i}$ are correlated so for $u_{i}$ and $r_{i}$, this dependent structure requires a more intricate analysis for the compensator. Similar as the proof for uniform convergence in Theorem (ref), we proceed by first considering variation in $a, c_{\psi}$ while $b = 0$. We then approximate the compensator by combining with the argument for variation in $b, c_{\psi}$ while $a = 0$.

Finally, we argue the empirical process $\mathbb{F}_{u, n}^{w, p}(0, 0, c_{\psi})$ is tight when viewed as a sequence in $n$ of processes on $D[0, 1]$. By Billingsley (1968, Theorem 13.2), two conditions need to be checked. First, it holds by construction that $\mathbb{F}_{u, n}^{w, p}(0, 0, c_{0}) = 0$ where $c_{0} = \mathsf{F}_{u}^{-1}(0) = -\infty$. Second, the next theorem shows that the modulus of continuity is small. The proof uses a dyadic argument and then apply an iterated martingale exponential inequality, see Lemma (ref), to explore the tail probability of the maximum of a family of martingales.

theoremLet $c_{\psi} = \mathsf{F}_{u}^{-1}(\psi)$. Suppose Assumption (ref)$(ia, va)$ holds with $0 < \nu < 1$, $s = 2$. Then for all $\epsilon > 0$ \begin{equation*} \lim_{\phi \downarrow 0} \limsup_{n \to \infty} \mathsf{P} \{ \sup_{0 \le \psi \le \psi^{\dagger} \le 1 : \psi^{\dagger} - \psi \le \phi} |\mathbb{F}_{u, n}^{w, p}(0, 0, c_{\psi^{\dagger}}) - \mathbb{F}_{u, n}^{w, p}(0, 0, c_{\psi})| > \epsilon \} \to 0. \end{equation*}

The empirical distribution functions $\widehat{\mathsf{F}}_{u, n}^{w, p}(a, b, c)$ and $\widehat{\Pi}_{c}^{(m + 1)}$ involve in the updated estimators $\widehat{\beta}_{c}^{(m + 1)}$, $(\widehat{\sigma}_{c}^{(m + 1)})^{2}$, see ((ref)), ((ref)). While $n^{1/2}$-convergence results have been established for the empirical process $n^{1/2} \widehat{\mathsf{F}}_{u, n}^{w, p}(a, b, c)$ by Theorem (ref), (ref), (ref), consistency for the estimator of $\Pi$ is then required to build up the asymptotic distribution theory for the iterated estimators of $\beta$, $\sigma^{2}$. A product moment appearing in $\widehat{\Pi}_{c}^{(m + 1)}$ but not included in $\widehat{\mathsf{F}}_{u, n}^{w, p}(a, b, c)$ is $n^{-1} \sum_{i = 1}^{n} z_{i} r_{i}^{\prime} v_{i, c}^{(m)}$, see ((ref)), ((ref)). Thus, proving consistency of the $\Pi$ estimator, we need the following $n$-convergence theorem for a new class of empirical processes with the weight $n^{1/2} z_{in} = z_{i}$ and mark $r_{i}^{\prime}$ instead of mark $u_{i}$ in the previous case.

theoremLet $c_{\psi} = \mathsf{F}_{u}^{-1}(\psi)$. Suppose Assumption (ref) holds with $\nu = 1$ and $s \ge 2$ satisfying ((ref)). Then for any $B > 0$ and as $n \to \infty$ \begin{equation*} \sup_{0 \le \psi \le 1} \sup_{|a|, |b| \le n^{1/4 - \eta}B} |n^{-1/2} \sum_{i = 1}^{n} z_{in} r_{i}^{\prime} \{ 1_{(u_{i} \le \sigma c_{\psi} +n^{-1/2}a c_{\psi} + z_{in}^{\prime} \Pi b + n^{-1/2}r_{i}^{\prime}b)} - 1_{(u_{i} \le \sigma c_{\psi})}\}| = \mathrm{o}_{\mathsf{P}}(1). \end{equation*}

The proof first considers uniform convergence for the empirical process without weights $n^{1/2} z_{in} = z_{i}$ and marks $r_{i}$. Since we establish $n$-convergence result instead of with the normalization $n^{1/2}$, the H{\"o}lder inequality is then used to prove Theorem (ref) by adding the weights and marks into the process with only indicators.

A result for the two-sided empirical process

Estimators in Algorithm (ref), (ref), (ref) involves indicators depending on the absolute value of residuals of structural errors. We therefore present some results for a class of two-sided weighted and marked empirical processes.

Define the weighted and marked absolute empirical distribution function

equation[equation omitted — 240 chars of source]

We suppose $a$ such that $\sigma + n^{-1/2}a > 0$, in which case it suffices to consider $c \ge 0$. This restriction on $a$ is satisfied when choosing $a$ as $\widetilde{a} = n^{1/2} (\widetilde{\sigma} - \sigma)$ so that $\sigma + n^{-1/2} \widetilde{a} = \widetilde{\sigma} > 0$. Introduce the compensator of $\widehat{\mathsf{G}}_{u, n}^{w, p}(a, b, c)$

equation[equation omitted — 244 chars of source]

Note that $\overline{\mathsf{G}}_{u, n}^{1, 0}(0, 0, c) = \mathsf{G}_{u}(c) = \mathsf{P}(|u_{i}| \le \sigma c)$. Similar as defining the one-sided empirical process $\mathbb{F}_{u, n}^{w, p}(a, b, c_{\psi})$ on the space $D[0, 1]$ equipped with the Skorokhod metric, let $c_{\psi} = \mathsf{G}_{u}^{-1}(\psi)$ for $0 \le \psi \le 1$ so $c_{\psi} \ge 0$ and the absolute empirical process is

equation[equation omitted — 213 chars of source]

We can now derive asymptotic theory for the absolute empirical process from Theorems (ref), (ref), (ref), (ref). These results are presented under more restrictive Assumption (ref), where the distribution $\mathsf{f}_{u}$ of structural error is symmetric, see Remark (ref). In this section, we only consider $w_{in}$ chosen as $1$, $n^{1/2} z_{in} = z_{i}$, $nz_{in} z_{in}^{\prime} = z_{i} z_{i}^{\prime}$ and $p$ as $0$, $1$, $2$.

The first theorem states the absolute processes with estimation errors converge to the one without estimation errors uniformly in the quantile.

theoremLet $c_{\psi} = \mathsf{G}_{u}^{-1}(\psi)$. Suppose Assumption (ref)$(ia, iia, iiia, ivb, ivc)$ holds. Then for any $B > 0$ and as $n \to \infty$ \begin{equation*} \sup_{0 \le \psi \le 1} \sup_{|a|, |b| \le n^{1/4 - \eta} B} |\mathbb{G}_{u, n}^{w, p} ( a, b, c_{\psi}) - \mathbb{G}_{u, n}^{w, p} ( 0, 0, c_{\psi})| = \mathrm{o}_{\mathsf{P}}(1). \end{equation*}

Then we provide the first-order approximation to the absolute compensator.

theoremLet $c_{\psi} = \mathsf{G}_{u}^{-1}(\psi)$. Suppose Assumption (ref)$(ia, iia, iiia, ivc)$ holds. Then for any $B > 0$ and as $n \to \infty$ \begin{equation*} \sup_{0 \le \psi \le 1} \sup_{|a|, |b| \le n^{1/4 - \eta}B} |n^{1/2} \{ \overline{\mathsf{G}}_{u, n}^{w, p}(a, b, c_{\psi}) - \overline{\mathsf{G}}_{u, n}^{w, p}(0, 0, c_{\psi}) \} - \mathcal{B}_{\mathsf{G}_{u}, n}(a, b, c_{\psi})| = \mathrm{O}_{\mathsf{P}}(n^{-2\eta}), \end{equation*} where the bias term, with $\xi_{c_{\psi}} = \mathsf{E} (\Sigma^{-1/2} r_{i} | u_{i}/\sigma = c_{\psi})$, is defined as \begin{align*} \mathcal{B}_{\mathsf{G}_{u}, n}(a, b, c_{\psi}) & = \sigma^{p - 1} c_{\psi}^{p} \mathsf{f}_{u}(c_{\psi}) n^{-1/2} \sum_{i=1}^{n} w_{in} [ \{1 + (-1)^{p} \} n^{-1/2} a c_{\psi} \\ & \quad \,\, + n^{-1/2} \{ \xi_{c_{\psi}} - (-1)^{p} \xi_{-c_{\psi}} \}^{\prime} \Sigma^{1/2} b + \{ 1 - (-1)^{p} \} z_{in}^{\prime} \Pi b ]. \end{align*}

The next theorem extends the tightness result to the case of empirical processes with absolute indicators.

theoremLet $c_{\psi} = \mathsf{G}_{u}^{-1}(\psi)$. Suppose Assumption (ref)$(ia, ivc)$ holds. Then for all $\epsilon > 0$ \begin{equation*} \lim_{\phi \downarrow 0} \limsup_{n \to \infty} \mathsf{P} \{ \sup_{0 \le \psi \le \psi^{\dagger} \le 1 : \psi^{\dagger} - \psi \le \phi} |\mathbb{G}_{u, n}^{w, p}(0, 0, c_{\psi^{\dagger}}) - \mathbb{G}_{u, n}^{w, p}(0, 0, c_{\psi})| > \epsilon \} \to 0. \end{equation*}

While the above theorems deal with $n^{1/2}$-convergence results, the final result considers $n$-convergence for a new type of absolute processes used to prove consistency of the estimator for $\Pi$ in the first stage regression ((ref)).

theoremLet $c_{\psi} = \mathsf{G}_{u}^{-1}(\psi)$. Suppose Assumption (ref)$(ia, iia, iiia, ivb, ivc)$ holds. Then for any $B > 0$ and as $n \to \infty$ \begin{equation*} \sup_{0 \le \psi \le 1} \sup_{|a|, |b| \le n^{1/4 - \eta}B} |n^{-1/2} \sum_{i = 1}^{n} z_{in} r_{i}^{\prime} \{ 1_{(|u_{i} - z_{in}^{\prime} \Pi b - n^{-1/2}r_{i}^{\prime}b| \le \sigma c_{\psi} + n^{-1/2}a c_{\psi})} - 1_{(|u_{i}| \le \sigma c_{\psi})}\}| = \mathrm{o}_{\mathsf{P}}(1). \end{equation*}

The proofs of empirical processes are given in Appendix (ref), but first we discuss a metric on $\mathbb{R}$ and provide some useful inequalities applied repeatedly in the chaining argument in Appendix (ref). Finally, the proofs of the main results follow in Appendix (ref).

Conclusion and discussion

Since most empirical analyses are affected by atypical observations, when many applied economists apply IVs regressions they first run ordinary 2SLS and use the resulting residuals to find non-outlying observations. They then re-run 2SLS on this subset, subsequently iterating this procedure until they obtain robust results.

In this paper we analyze this simple robust algorithm asymptotically, and provide the consistent estimation and valid inferential procedures given the cut-off value. We concentrate on cross-sectional i.i.d. data in this paper, though our analysis can be extended to time series, whether stationary, deterministic trend, or unit root. The results are all derived under the null hypothesis that there are no outliers in the model. Since in this case there is still a positive probability to find outliers using such algorithms, we then propose the concept of gauge - the expected retention rate of falsely discovered outliers - to evaluate their performance.

Future work should focus on building asymptotic theory for the gauge, to give empirical researchers an indirect way of choosing the cut-off. Furthermore, some specification tests could be devised to check whether outliers truly exist with respect to the reference model ((ref), (ref)), as in Jiao and Pretis (2018).

It is well known that the first order asymptotic approximation is fragile in some small finite sample situations. Therefore, it would be of interest to carry out simulation studies to evaluate the finite sample performance of the results in this paper. Likewise it would be of interest to extend these results to situations where outliers are actually present in the data generating process. Possible scenarios include single outliers, clusters of outliers, level shifts, symmetric or non-symmetric outliers. In such situations, we would analyze the potency, which is the retention rate for relevant outliers.