EconBase
← Back to paper

Occasionally Misspecified

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

85,610 characters · 13 sections · 79 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Occasionally Misspecified

abstractWhen fitting a particular Economic model on a sample of data, the model may turn out to be heavily misspecified for some observations. This can happen because of unmodelled idiosyncratic events, such as an abrupt but short-lived change in policy. These outliers can significantly alter estimates and inferences. A robust estimation is desirable to limit their influence. For skewed data, this induces another bias which can also invalidate the estimation and inferences. This paper proposes a robust GMM estimator with a simple bias correction that does not degrade robustness significantly. The paper provides finite-sample robustness bounds, and asymptotic uniform equivalence with an oracle that discards all outliers. Consistency and asymptotic normality ensue from that result. An application to the “Price-Puzzle,” which finds inflation increases when monetary policy tightens, illustrates the concerns and the method. The proposed estimator finds the intuitive result: tighter monetary policy leads to a decline in inflation.

JEL Classification: C11, C12, C13, C32, C36.\newline Keywords: Leveraged outliers, Structural Vector-Autoregression, Instrumental Variables.

\baselineskip=18.0pt \thispagestyle{empty} \setcounter{page}{0}

Introduction

Empirical data is routinely used to fit and test Economic models or predictions. Although the model may explain much of the variation in the data, it may also turn out to be particularly misspecified for some observations. This can result from sudden, yet temporary, changes in policy. To illustrate: monetary policy is often measured via changes in interest rates. Between 1979 and 1982, the Federal Reserve no longer fixed the Federal Funds Rate as a policy tool, targetting monetary aggregates instead coibion2012. Sharp changes in interest rates during that period generate significant identifying power on the effects of monetary policy. Yet, misspecification threatens the validity of the resulting estimates and inferences. Other factors that can cause occasional misspecification include imperfect data matching, or some rare - but significant - prediction errors when generating regressors.

A robust estimator is desirable in these scenarios: being less sensitive to influential outliers. However, robust estimates can be biased and inconsistent when the underlying data is asymmetric. To illustrate: the sample median is more robust than the mean; however, it estimates a different quantity when the data is skewed. This is relevant as many economic variables -- income, prices, and quantity, to name a few -- tend to be skewed. When symmetric data is contaminated asymmetrically, both the mean and median are biased. Further, in a linear regression context, hamilton1992 stresses that robust M-estimators are “designed for protection against wild errors or y-outliers. x-outliers are its Achilles' heel.” Leverage characterizes x-outliers, which is bounded for ordinary least-squares. In the example above: sharp changes in interest rates imply high leverage around 1979-1982. The issue is even more pronounced in non-linear regressions where leverage is not necessarily bounded st1992. This superleverage can further exacerbate the influence of outliers.

This paper proposes a robust Generalized Method of Moments (GMM) estimator with a simple bias-correction step. Building on ronchetti2001, the sample moments are estimated robustly; here using a penalized student log-likelihood criterion. The particular choice of criterion makes the asymptotic asymmetry bias tractable. A linear combination, known as Richardson extrapolation, of two robust moment estimates is asymptotically unbiased. The bias, which depends on higher-order moments, is not estimated. The correction does not degrade robustness significantly. Also, in linear regressions, robust GMM estimates are robust against x-outlier, unlike M-estimates which only screen for large residuals. Given these moment estimates, the model is estimated in the same fashion as a standard GMM.

Finite and large sample results describe the properties of the method against adversarial contamination. First, uniform finite-sample exponential bounds, for cross-sections and mixing time-series, measure how robust moment estimates deviate from their biased target. This provides a worst-case global robustness guarantee for a given level of data contamination. The combination of the student likelihood, which is neither convex nor bounded but has a bounded influence function, with the particular choice of penalty is key for this result.

The large-sample results require the number of outliers to increase more slowly than the sample size. Their influence can grow rapidly: non-robust estimates may be inconsistent, or diverge. This captures the finite-sample setting where a few observations overwhelm the estimation. The bias-corrected robust moment and parameter estimates are shown to be first-order equivalent to an oracle which discards all outliers. Asymptotic normality follows from standard regularity conditions on the oracle. For linear models, the robust GMM estimates can be expressed as weighted least-squares or weighted two-stage least-squares. The weights are easy to compute and report, highlighting which observations were downweighted in the process. This should reduce concerns about black-box results.

Simulations illustrate the small sample properties of the proposed estimator in the presence of x-outliers, which have high leverage. OLS is very sensitive. A robust M-estimator packaged in R is biased and sensitive. Without correction, the procedure is more robust but biased. Bias correction reduces estimation error and improves coverage of t-tests. As the proportion of outliers increases, its performance degrades but remains better than the benchmarks. Undersmoothing, sometimes suggested in the literature, is also less robust than bias correction. Three empirical applications illustrate the relevance of the procedure.

The first estimates the effect of a monetary policy shock on inflation using a structural Vector Autoregressive (VAR) model as in stock2001. OLS estimates a “Price-Puzzle:” predicting an inflation increase when monetary policy tightens. Two historical sub-periods of unusual monetary policy -- including 1979-1982 -- significantly influence this result. The proposed estimates find the intuitive result: a negative impact on inflation. The weights reveal that the two historical subperiods are downweighted to get this result. Robust estimates overweight some observations. Bias correction re-adjusts towards equal weighting.

Recently, young2022 found that many instrumental variable (IV) results involve highly leveraged regressions, and are very sensitive to outliers. Two applications illustrate the methodology in this setting. The first considers the relationship between trade openness and inflation romer1993. The second is about the effect of segregation on the quality of government alesina2011. Both regressions are highly influenced by a few observations. Robust estimates have significantly smaller standard errors, producing more precise inferences. Bias correction reveals non-negligible bias in robust estimates.

\paragraph{Structure of the paper.} Section (ref) motivates the paper with the Price Puzzle example. Section (ref) surveys the existing literature. Section (ref) introduces the setting, sampling assumptions, and the estimator. Derivations for a simplified estimator give insights for the finite and large sample results. Section (ref) provides finite-sample bounds and asymptotic results. Simulated and empirical applications are in Section (ref). Appendices (ref), (ref) give the proofs for the main results and preliminary ones. Supplemental Appendices (ref), (ref), (ref), (ref), (ref), (ref) provide proofs for the preliminary results, simple derivations with leveraged outliers, derivations for influence and leverage in IV regressions, additional simulation and empirical results, and detailed numerical Algorithms to perform the estimation.

Motivating Example: the Price Puzzle

To illustrate the issues considered in this paper, consider estimating the impact of monetary policy with a recursive vector autoregressive (VAR) model as in stock2001. There are three variables: inflation ($\pi_t$), unemployment rate $(u_t)$, and the federal funds rate ($R_t$). The VAR is estimated by OLS with four lags on U.S. data from 1960Q1 to 2000Q4.

Panel a) in Figure (ref) plots the estimated response of inflation to a unit increase in $R_t$. It shows a positive and significant increase in inflation for nearly four consecutive quarters. This was first observed by sims1992 and immediately coined as a `Price Puzzle' by Eichenbaum1992. It has since been studied extensively. rusnak2013 performed a meta-analysis of $1000$ estimates and put forward several potential forms of model misspecification to explain the puzzle. The number of specifications they explore is several times greater than the sample size so there should be some concerns about overfitting, however.

The following presents some simple diagnostics that indicate two time periods strongly influence the estimates. The puzzle begins with a positive and significant initial impact. It is measured by $\beta_1$ in the regression:

align[align omitted — 189 chars of source]

Figure (ref) investigates this regression more closely. Panel a) plots the residuals $\hat{e}_{\pi,t}$ over time. Besides some increased volatility between 1970-1982, there are no obvious outliers in the series. In fact, the skewness and kurtosis are $0.36$ and $3.78$, respectively, not far from a normal distribution. Panel c) approximates the contribution of each $t$ to $\hat{\beta}_1$. Since $\hat{\beta}_n = \sum_{t=1}^n (X^\prime X /n )^{-1} x_t y_t /n$ is a sample mean, $(X^\prime X /n )^{-1} x_t y_t$ approximates the contribution of each $t$ to the mean. Some observations stand out: for instance, 1981Q1 alone positively contributes $\approx 75/n = 0.47$ to $\hat{\beta}_1 = 0.21$, about $3.5$ standard errors.\footnote{Most coefficients in ((ref)) are strongly influenced by a few observations as shown in Table (ref). The contribution reported here is related to Cook's distance which measures changes in predicted values $\hat{y}_t$ when observation $t$ is excluded in the estimation cook1977. Here, the effect of observation $t$ on the estimated regression coefficients is the object of interest -- this will be referred to as contribution.}

figure[figure omitted — 557 chars of source]

Panels b,c) show that, although none of the residuals $\hat{e}_{\pi,t}$ are particularly large, two time periods, around 1974-1975 and 1979-1982, have a disproportionate influence on the results. The latter has historical significance: the Federal Reserve changed to non-borrowed reserves targeting where the interest rate $R_t$ was no longer a fixed policy instrument, as discussed in the introduction. Richmond FED President, Robert P. Black, summarized the tactical change during the October 1979 FOMC meeting as follows:

quote“I often think of our position as being analogous to that of a monopolist in the sense that we control the money supply. A monopolist has a choice of controlling either price or quantity but he can never control both. I believe we’ve been trying to control the quantity of money by setting the price and we have misjudged. We’ve jiggled the price, in terms of the federal funds rate, one way or the other, and we‘ve usually met with less than complete success in judging what quantity of money will be forthcoming from that.” FOMC100679

This has several implications for the VAR estimates. First, $R_t$ was no longer a direct measure of monetary policy: the recursive VAR may not correctly identify monetary shocks during that time period. Importantly, this goes beyond parameter instability. Time-varying parameters, regime-switching, or structural break models would still require $R_t$ to provide a measure of monetary policy shocks. As emphasized by Robert Black, monetary policy was conducted on monetary aggregates at that time, not interest rates. Second, interest rates were significantly more volatile with the policy change;\footnote{This was anticipated and monitored by board members as shown by FOMC Transcripts of 1979-1982.} producing significant regression leverage. This, as highlighted in Figure (ref), gives excess influence to these observations.

Misspecification arises because the central bank relies on multiple policy instruments, the VAR only uses $R_t$. friedman1963 argued that well-known historical events clearly identify large monetary shocks. This narrative approach was popularized by romer1989, romer2004. Narrative and VAR estimates can differ when the central bank relies on different instruments throughout the sample coibion2012,monnet2014. Narrative estimates, however, aggregate multiple types of monetary policies; results cannot be interpreted as e.g. an interest rate shock.

To identify the effect of an interest rate shock, a robust estimation is desirable. However, as noted in the introduction robust M-estimates may be biased and may not be robust to these x-outliers. Because residuals are small, robust M-estimates with Huber loss and high-breakdown MM estimates (rlm, lmRob in R) are nearly identical to Figure (ref) (not reported).

Diagnostics, as presented above, are useful to assess whether the estimation might present some irregularities. A robust estimation, presented below, is meant to reduce the influence of abnormal observations. The two are complementary, see huber2011 for further discussion.

Figure (ref) re-estimates the effect on the same data, with the same model specification: using OLS (panel a), the proposed robust estimator without bias correction (panel b), with bias correction (panel c), with bias correction and a small sample correction (panel d). Without bias correction, the price puzzle remains -- but does not last 4 quarters anymore. With bias correction, the price puzzle disappears; the initial effect is not significant. With the additional adjustment, the effect is qualitatively larger and negative.

As discussed above, the estimates can be seen as weighted least-squares. Figure (ref) compares the weights, for each time period, used by each method on a regular and a log-scale (resp. top, bottom). OLS uses equal weighting (black/dashed). Without bias correction, robust estimates downweigh the leveraged outliers, especially 1979-1982, but overweigh other periods (black/solid). Bias correction re-adjusts towards equal weighting (blue/dot). The small sample adjustment further re-adjusts in that direction (purple/triangle).

figure[figure omitted — 617 chars of source]
figure[figure omitted — 720 chars of source]

Related Literature

The paper is mainly related to the literature on robust estimation, mostly developed in statistics. Textbook references such as huber2011 and maronna2019 survey a wide range of estimators and their properties. To focus the discussion, consider a linear regression: $y_t = x_t^\prime \theta + e_t$. Robust M-estimators minimize the loss $\sum_{t=1}^n\psi( y_t - x_t^\prime \theta )$ over $\theta$. While OLS uses a quadratic $\psi$, least-absolute deviation (LAD), and the Huber1964 loss are non-quadratic. They increase linearly with large residuals $|y_t - x_t^\prime \theta|$. This reduces the influence of y-outliers. Winsorizing and trimming are popular alternatives. Huber1964 notes that trimming can be sensitive around the cutoffs. The first-order condition implies the solution $\hat{\theta}_n$ satifies $\sum_{t=1}^n x_t \psi^\prime(y_t - x_t^\prime \hat{\theta}_n)$. Large residuals $\hat{e}_t = y_t - x_t^\prime \hat{\theta}_n$ are handled by $\psi^\prime$. However, x-outliers with a large $x_t$, are not screened by $\psi^\prime$.\footnote{Mallows type estimators separately screen for leverage, see e.g. carroll1988.} When the distribution of $e_t$ is symmetric and the sample is contaminated symmetrically, robust estimates are consistent and asymptotically normal under regularity conditions. Symmetry is critical. jaeckel1971 derived, for estimating a location parameter, with asymmetric contamination of symmetric data, an asymptotic bias of order $n^{-1/2}$ when the proportion of outliers is $O(n^{-1/2})$ -- i.e. $n_o = O(n^{1/2})$. $n_o$ is the number of outliers in the sample of size $n$. Recently, dalalyan2022 proposed an attractive robust location estimator for multivariate Gaussian or sub-Gaussian data with a high-breakdown point - i.e. robust to a large fraction of outliers in the sample. Here, the finite-sample results are derived using only finite second moment conditions. Also, the focus here is on settings where data is asymmetric, contaminated by a small number of highly influential outliers.

For asymmetric data, the estimator may not be consistent, see carroll1988 for linear regressions. Quasi-Maximum Likelihood estimation, with a student distribution for the errors, is commonly used to estimate volatility models. newey1997 show that the estimates may not be consistent without symmetry conditions. In a parametric setup, cantoni2001 provide analytical bias formulas for generalized linear models, used to correct the first-order condition of the M-estimation. Here, parametric assumptions are not required. zhou2018 derive bias bounds and exponential inequalities for linear regressions with the Huber loss when $e_t$ has finite variance. They do not consider sample contamination and require sub-gaussian regressors - i.e. no x-outliers. These two issues are particularly relevant for the Price Puzzle. Another approach to robustness is to bound the asymptotic bias in a local neighborhood of the model using the influence curve (IC) of hampel1974, see e.g. huber2011. andrews1986 relates the IC to the stability of estimators. Recently, several papers have used the IC to study and bound local misspecification bias for GMM, e.g. andrews2017, armstrong2021, bonhomme2022. Under these local asymptotics, the estimator remains consistent and asymptotically normal with a bias proportional to sampling uncertainty. In this paper, the model is grossly misspecified, but only for $1 \leq n_o \ll n$ outliers. Non-robust estimates can be inconsistent, or diverge: a robust estimation is required. christensen2023 propose global sensitivity analyses on distributional assumptions, the model is otherwise correctly specified. It is common in Economics to apply more robust testing to non-robust estimates, assuming consistency, asymptotic normality -- unlike here. One can adjust standard errors mackinnon2012, critical values muller2020,potscher2023, or both. sasaki2023 propose a test for finite moments at a point, as required for consistency and central limit theory. cowell1996 and cowell2007 consider the robustness properties of inequality measures, e.g. Gini coefficient. Surveying a large number of empirical results, young2022 finds that many IV regressions are highly leveraged and sensitive to a few observations, or clusters of observations.

For GMM estimation, ronchetti2001 proposed a robust estimator that is locally asymptotically robust, using the IC criteria. hill2010, vcivzek2016 consider trimming in GMM estimation. rohatgi2022 use a filter algorithm to screen out outliers in GMM estimation. The median-of-means is popular in prediction problems, which could also be considered here: the dataset is split into $K \geq 2$ subsamples of $m = n/K$ observations. $K$ sample means are computed. The median of the $K$ means is the estimator. The estimate is robust for up to $n_o \leq K/2-1$ outliers, see e.g. lecue2020, laforgue2021. To accommodate an increasing $n_o$, having $K \to \infty$ as $n \to \infty$ is necessary. This introduces a bias, bounded above by $\sigma/\sqrt{m} = \sigma \sqrt{K/n}$.\footnote{For any distribution, the median and the mean differ by at most: $|\text{median}(X) - \mathbb{E}(X)| \leq \sigma(X)$.} Even for $K$ fixed, an asymptotic bias can arise. Without a tractable expression for the bias, it is not clear how one would correct the asymptotic bias. Here, the choice of loss function makes the asymptotic bias tractable. An alternative is undersmoothing where the tuning parameter diverges fast enough that the bias is asymptotically negligible. It only requires to bound the asymptotic bias. Section (ref) illustrates that it is less robust than bias-correction.

Models, Sample, Estimator

This paper considers estimations from unconditional moment restrictions:

align[align omitted — 122 chars of source]

where $z_t \overset{d}{\sim} P$ and the solution $\theta_0 \in \Theta$, a compact subset of $\mathbb{R}^k$. OLS regressions correspond to $g(z_t;\theta) = x_t(y_t - x_t^\prime \theta)$ where $z_t = (y_t,x_t)$ collects the dependent variable and the regressors. For instrumental variable regressions, take $g(z_t;\theta) = w_t(y_t - x_t^\prime \theta)$ where $z_t = (y_t,x_t,w_t)$ collects the dependent variable, the regressors and the instruments. Non-linear estimations also fit into this framework. Concave Likelihood maximization, such as Probit or Logit, would set ((ref)) to be the first-order condition. The main examples are linear.

The dataset consists of $n$ observations but $z_t \sim P$ may not hold for all $t = 1,\dots,n$. This is presented in the following Assumption.

assumption[Sample] There are $n = n_P + n_o$ observations such that \begin{itemize} • for $t \in \{1,\dots,n_P\}$, $z_t \sim P$ for which ((ref)) holds, are either iid or strictly stationary, $\beta$-mixing with rate $\beta_{m} \leq a \exp(- b m)$ for $0 < a,b < \infty$; • for $t \in \{n_P+1,\dots,n \}$ and $0 < A,\alpha < \infty$: \begin{align} z_t \in \mathcal{O}_n := \{ z s.t. \sup_{\theta \in \Theta} \|g(z;\theta)\|^2 \leq A n^{\alpha} \}. \end{align} \end{itemize}

The first $n_P$ observations are such that ((ref)) holds. However, the last $n_o$ observations, or outliers, can be arbitrary in $\mathcal{O}_n$. The ordering between observations simplifies notation and, for time-series, preserves the dependence structure of the good $n_P$ observations. The mixing condition typically holds for stationary VAR models, as in the motivating example. In practice, the user does not know which observations are drawn from $P$ and those that are not. The $n_o$ outliers could be allocated anywhere within the sample. The outliers will be chosen in an adversarial fashion, looking at the least-favorable collection $(z_{n_{P}+1},\dots,z_n) \in \mathcal{O}_n$ for each $\theta$, without restrictions on dependence.

The goal here is to derive finite-sample robustness properties against the worst-case realization of the $n_o$ outliers. Ex-ante, if the $n_o$ outliers are randomly distributed, such that $\mathbb{P}(\sup_{\theta \in \Theta}\|g(z_t;\theta)\| > t) \leq t^{-\varepsilon}$ for some $\varepsilon > 0$. Then $\mathbb{P}( z_t \in \mathcal{O}_n \text{ for each }t=n_{P}+1,\dots,n ) \geq 1-A^{-\varepsilon} n_o n^{-\alpha \varepsilon}$, can be made arbitrarily close to $1$ setting $\alpha$ large enough. In practice, the user does not specify $(A,\alpha)$. For random data contamination, Assumption (ref) can be interpreted as conditioning on a realization with $n_o$ outliers in the set $\mathcal{O}_n$ which has arbitrarily high-probability given an appropriate choice of $n_o$, $A$ and $\alpha$.\footnote{See also Remark 1 in laforgue2021.}

Outliers can take many forms in ((ref)). Figure (ref) illustrates that residuals $y_t - x_t^\prime \theta$ are not the only source of influence, captured here by $x_t(y_t - x_t^\prime \theta)$. High leverage observations are only influential if $|y_t - x_t^\prime\theta| \gg 0$. Likewise, $x_t(y_t - x_t^\prime \theta)$ can be large when neither $y_t - x_t^\prime \theta$ nor $x_t$ are individually large but their product is non-negligible. This implies that screening residuals and regressors separately, as suggested in hamilton1992, can be insufficient. The influence of a single observation can also vary depending on the model specification: a regression that is linear in $x_t$ is typically less leveraged than in a quadratic specification with $(x_t,x_t^2)$ as regressors. Collinearity also plays a role on influence, as $(X^\prime X/n)^{-1} x_t y_t$, reported in Figure (ref), can be greatly inflated by the collinearity factor $(X^\prime X/n)^{-1}$. In the motivating example, the regressors are lagged variables which are autocorrelated, i.e. collinear. A rotation invariance property is important to ensure robustness when there are multiple regressors. For instrumental variable regressions, the relevant quantity $w_t(y_t - x_t^\prime \theta)$ involves the instruments $w_t$ and the residual. In the context of time-series, one concern would be innovation outliers associated with a large shock $y_t - x_t^\prime \theta$. Another, similar to the description in the motivating example, would be additive outliers. Here the effect is isolated, as in a different regime that occurs only once within the sample.

The main concern here is that the sample mean $\overline{g}_n(\theta) = 1/n \sum_{t=1}^n g(z_t;\theta)$ is not a consistent estimator for $\mathbb{E}_P[g(z_t;\theta)]$ when $(n_o n^{\alpha})/n \not \to 0$. This allows to capture the concern that a minority of observations has significant influence, even as the sample size $n$ increases. For $n_o =1$, $\alpha = 1/2$ the estimates are consistent but asymptotically biased, standard error estimates are also affected.\footnote{This is illustrated in Appendix (ref).} For $n_o = 1$, $\alpha = 1$ estimates are inconsistent. They diverge when $\alpha >1$. Mild outliers are also problematic: for $n_o = n^{1/4}$ and $\alpha = 1/4$ estimates are asymptotically biased.

To handle contaminated samples, ronchetti2001 showed that a robust estimate of $\mathbb{E}_P[g(z_t;\theta)]$ is required. The following first computes a robust estimate of $\mu(\theta) = \mathbb{E}_P[g(z_t;\theta)]$, then corrects the first-order asymptotic bias, and finally solves for $\mu(\theta)=0$.

\paragraph{Step 1.} For each $\theta \in \Theta$, find $\hat{\psi}_n(\theta;\nu)$ which minimizes the sample criterion:

align[align omitted — 257 chars of source]

where $\psi = (\mu,\Sigma)$ and $p = \text{dim}(g(z_t;\theta))$. The location and scale parameters are estimated jointly to ensure the first is invariant to rotation and less sensitive to re-scaling. The loss $Q_n$ consists of a student quasi-likelihood plus two penalization terms. The tuning parameter $\nu > 0$ controls the robustness of the estimates. Here, it is not estimated and acts as a critical value. For observations such that $\|g(z_t;\theta)-\mu\|^2_{\Sigma} \ll \nu$, the loss is approximately quadratic and approximates the Gaussian log-likelihood. In contrast, for observations such that $\|g(z_t;\theta)-\mu\|^2_{\Sigma^{-1}} \gg \nu$ the loss is approximately logarithmic. Large values for $g(z_t;\theta)$ have a lesser impact compared to the Gaussian likelihood.

To fully capture robustness, the parameter space $\Psi$ for $\psi$ is unbounded: \[ \Psi = \{ (\mu,\Sigma), \, \mu \in \mathbb{R}^p, 0 < s_0 \leq \lambda_{\min}(\Sigma) \leq \lambda_{\max}(\Sigma) \leq +\infty \},\] where $s_0$ is such that $s_0 \leq \lambda_{\min}(\text{var}_P[g(z_t;\theta)]) < +\infty$ for all $\theta \in \Theta$. In the presence of outliers, the main concern is in estimating a large $\hat{\mu}_n$ and/or $\hat{\Sigma}_n$. Here, setting $s_0 >0$ simplifies some derivations to focus on finite-sample upper bounds.

The robustness of the student log-likelihood has some downsides numerically. Without regularization ($\kappa_1=\kappa_2 = 0$), the derivative $\partial_{\mu} Q_n(\psi;\theta) = 0$ for $\|\mu\| = +\infty$ and any $\Sigma$. The student likelihood becomes flat for larger values of $\|\mu\|$. With non-zero penalties, i.e. $\kappa_1$ and $\kappa_2 \neq 0$, $\partial_{\mu} Q_n(\psi;\theta) \to \infty$ when $\|\mu\| \to \infty$. The combination of the student log-likelihood, which has bounded influence, with this choice of penalty implies the estimates $\hat\psi_n(\theta;\nu)$ are bounded, as shown in the next Section. The self-normalization $\|\mu\|_{\Sigma^{-1}}$ is invariant to rotations of the moments and less sensitive to scale. At the solution $\theta = \theta_0$, $\mu(\theta_0)=\mathbb{E}_P[g(z_t;\theta_0)] = 0$ holds. This motivates penalizing towards zero in this particular setting.

Simultaneously estimating the location and scale parameters can seem problematic. A large $\hat{\Sigma}_n$ is effectively similar to using a large $\nu$, leading to less robust location estimates $\hat{\mu}_n$. The second penalty $\text{trace}(\Sigma)$ is important in that regard, as it ensures $\hat{\Sigma}_n$ cannot be too large in finite samples. This is shown in the next Section.

\paragraph{Step 2.} For each $\theta \in \Theta$, compute:

align[align omitted — 107 chars of source]

This type of adjustment is known as Richardson extrapolation in numerical analysis. Unlike the sample mean, the estimator $\hat{\mu}_n(\theta;\nu)$ is typically biased for $\nu < +\infty$. Taking $\nu \to \infty$ with $n \to\infty$ at an appropriate rate, the adjustment $2 \hat{\mu}_n(\theta;\nu) - \hat{\mu}_n(\theta;\nu/2)$ corrects the first-order asymptotic bias. The bias depends on higher-order moments (see below). Estimating this bias is not straightforward: robustly estimating the first moment is already a challenge in this setting. The correction ((ref)) is simple to implement and widely applicable.

\paragraph{Step 3.} Find $\tilde\theta_{n}$ such that:

align[align omitted — 143 chars of source]

The estimated $\tilde\theta_{n}$ inherits the asymptotic bias properties of the bias corrected moments $\tilde{\mu}_n$.

Step 1. continuously updates both $\mu$ and $\Sigma$ with $\theta$. The scaling $\hat{\Sigma}_n(\theta;\nu)$ used to normalize the estimation of $\hat{\mu}_n(\theta;\nu)$ adapts to the value of $\theta$. Appendix (ref) gives generic Algorithms (ref), (ref) used to compute $\hat{\psi}_n$, $\tilde{\theta}_n$ in the applications. $\tilde{\mu}_n(\theta;\nu)$ is as smooth as $g(z_t;\theta)$ -- cf. implicit function Theorem. Gradient-based optimizers, e.g. gradient-descent or Gauss-Newton, can be used. They are globally convergent under rank conditions forneron2023. Unlike trimmed moments, the estimated $\tilde{\mu}_n(\theta;\nu)$ varies continuously with $\nu$. This implies that the estimates $\tilde{\theta}_n$ can be less sensitive to small changes in tuning parameters. Figures (ref)-(ref) reproduce Figure (ref) with larger values of $\nu$, illustrating that the estimated impulse response function changes continuously with $\nu$.

Numerical software typically proceeds iteratively, see e.g. huber2011. Fix a tuning parameter and fit an initial regression $\hat{\theta}_n^1$. Then, update the scale parameter - here $\hat{\Sigma}_n^1$, re-estimate the regression $\hat{\theta}_n^2$, re-estimate the scale parameter, and repeat until convergence. The same scaling is applied for all $\theta$ at each stage. For least-squares, rreg in Stata and rlm in R proceed this way. Stata's rreg is initialized with a non-robust OLS estimate. The properties of the estimates after many iterations are not easy to derive, especially as scale estimates are less robust than those of location. Here, uniform-in-$\theta$ non-asymptotic concentration inequalities for the joint parameter $\hat\psi_n(\theta;\nu)$ are derived. This gives some finite-sample guarantees for step 1. above.

\paragraph{Intuition for the results.} To better understand the role of the tuning parameter $\nu$ and the bias-correction step, consider estimating a scalar parameter $\theta_0 = \mathbb{E}_P(z_t)$ using: \[ \hat{\mu}_n(\nu) = \frac{1}{n} \sum_{t=1}^n \frac{ z_t }{ 1+|z_t|^2/\nu }, \] which simplifies the first-order condition of $Q_n$ with respect to $\mu$.\footnote{The first-order condition $\partial_\mu Q_n = 0$ reads $\frac{\nu + p}{\nu n} \sum_{t=1}^n \frac{ z_t -\mu }{1+\|z_t - \mu\|^2_{\Sigma^{-1}}/\nu} + \frac{\kappa_1 \mu}{\nu} = 0$.} For any $z$, $\frac{|z|}{1+|z|^2/\nu} \leq \frac{\sqrt{\nu}}{2}$ bounds the influence of a single observation. Let $\mu(\nu) = \mathbb{E}_P \left( z_t/(1+|z_t|^2/\nu) \right)$. If $z_t$ are iid for $t \in \{1,\dots,n_P\}$, regardless of the remaining $n_o$ observations:\footnote{This inequality implies $\mathbb{P}( \sup_{z_t \in \mathcal{O}_n,t>n_P}|\hat{\mu}_n(\nu)-\mu(\nu)| \geq \frac{\sqrt{\nu} n_o}{n} + C\frac{n_P}{n}[\sqrt{\frac{x}{n_P}} + \frac{x}{n_P}] ) \leq 2 \exp ( - x )$ for some constant $C$. This is the form used in a later Theorem. } \[ \mathbb{P}\left( \sup_{z_t \in \mathcal{O}_n,t>n_P}|\hat{\mu}_n(\nu)-\mu(\nu)| \geq \frac{\sqrt{\nu} n_o}{n} + \frac{n_P}{n}\frac{x}{\sqrt{n_P}} \right) \leq 2 \exp \left( - \frac{x^2}{2\sigma^2_\nu + \frac{2}{3}\sqrt{\frac{\nu}{n_P}}x} \right), \] using Bernstein's inequality, with $\sigma^2_\nu = \text{var}_P\left(\frac{z_t}{1+|z_t|^2/\nu}\right) \to \text{var}_P(z_t)$ as $\nu \to \infty$. The right-hand-side is approximately sub-Gaussian for $x \ll \sqrt{n_P/\nu}$ and sub-exponential for $x \gg \sqrt{n_P/\nu}$. The factor $\sqrt{\nu/n_P}$ indicates the rate at which the estimator becomes sub-Gaussian.

As expected, outliers introduce a bias. The worst-case bias is at most $\sqrt{\nu} n_o/n$. Consistency of $\hat{\mu}_n$ requires $(\sqrt{\nu}/n)n_o = o(1)$ and asymptotic normality $(\sqrt{\nu/n})n_o = o(1)$. More contamination $n_o$ requires a smaller $\nu$ to compensate. The same $\nu$ introduces another bias: \[ \mu(\nu) = \theta_0 - \frac{1}{\nu}\mathbb{E}_P\left( \frac{z_t^3}{1+z_t^2/\nu} \right), \] as measured by the last term. It is typically non-zero when the distribution is not symmetric around $0$. The bias is at most $\mathbb{E}_P(|z_t|^3)/\nu$ or $\mathbb{E}_P(|z_t|^2)/(2 \sqrt{\nu})$ if, respectively, the third or second moment is finite. Consistency requires $\nu \to \infty$ and asymptotic normality $\sqrt{n}/\nu = o(1)$. There is some tradeoff between the outlier bias $\sqrt{\nu}n_o/n$, which mandates a smaller $\nu$, and this robustness bias, which compels using a larger $\nu$. A bias reduction that does not significantly degrade robustness can be achieved using $\tilde{\mu}_n = 2 \hat{\mu}_n(\nu) - \hat{\mu}_n(\nu/2)$, since: \[ \tilde{\mu}(\nu) = 2\mu(\nu) - \mu(\nu/2) = \theta_0 - \frac{1}{\nu^2}\mathbb{E}_P\left( \frac{z_t^5}{(1+z_t^2/\nu)(1+2z_t^2/\nu)} \right). \] Now the bias is at most $\mathbb{E}_P(|z_t|^5)/\nu^2$ or $\mathbb{E}_P(|z_t|^4)/\nu^{3/2}$ if, respectively, the fifth or fourth moment is finite. For the former, asymptotic normality only requires $\sqrt{n}/\nu^2 = o(1)$. The effect of a single observation on the estimate $\tilde{\mu}_n$ is no more than $\sqrt{2 \nu} + \sqrt{\nu}/2$, compared to $\sqrt{\nu/2}$ for the non-corrected $\hat{\mu}_n$. The bias correction does require more regularity from the uncontaminated data in terms of moments - 5 instead of 3 finite ones.

Higher-order Richardson extrapolation could further reduce the order of the asymptotic bias. Simulations suggest the following can give better results in small samples. Applying the correction once more using $\tilde{\raisebox{0pt}[0.85\height]{$\tilde{\mu}$}}_n = 2\tilde{\mu}_n(\nu) - \tilde{\mu}_n(\nu/2)$ flips the sign of the asymptotic bias and can have some small sample effects: \[ \tilde{\raisebox{0pt}[0.85\height]{$\tilde{\mu}$}}(\nu) = \theta_0 + \frac{2}{\nu^2} \mathbb{E}_P \left( \frac{z_t^5(1 - 4z_t^4/\nu^2)}{(1+z_t^2/\nu)(1+2z_t^2/\nu)(1+2z_t^2/\nu)(1+4z_t^2/\nu)} \right). \] To illustrate, take $z_t = \theta_0$ constant. Then $\tilde{\mu}(\nu) = \theta_0$ if, and only if, $\theta_0 = 0$ whereas $\tilde{\raisebox{0pt}[0.85\height]{$\tilde{\mu}$}}(\nu) = \theta_0$ if $\theta_0 \in \{\theta_0,-\sqrt{\nu/2},\sqrt{\nu/2}\}$. For finite $\nu$, the bias of $\tilde{\raisebox{0pt}[0.85\height]{$\tilde{\mu}$}}$ has two additional roots. Simulations in Section (ref) indicate small-sample improvements for estimation and inference.\footnote{Note that averaging $2/3 \tilde{\mu}(\nu) + 1/3 \tilde{\raisebox{0pt}[0.85\height]{$\tilde{\mu}$}}(\nu) = \theta_0 + o(\nu^{-2})$ can reduce the asymptotic bias, by the dominated convergence Theorem. This is not pursued here.}

Properties of the Estimator

Finite Sample Bounds

The following Lemma shows the importance of the penalization $\kappa_1,\kappa_2$ in ((ref)) which effectively bounds the parameter space $\Psi$.

lemmaFor any $\theta \in \Theta$ and $\nu >0$, the minimizer $\hat\psi_n = (\hat{\mu}_n,\hat{\Sigma}_n)$ of ((ref)) over $\Psi$ satisfies: \begin{align} \|\hat{\Sigma}_n^{-1/2}\hat{\mu}_n\| \leq \frac{\nu^{3/2}(1+p/\nu)}{2\kappa_1}, \quad trace(\hat{\Sigma}_n) \leq \frac{\nu^2 (1+p/\nu)}{\kappa_2} + \frac{\nu^4(1+p/\nu)^2}{4 \kappa_1\kappa_2} + \frac{p \nu}{\kappa_2}. \end{align}

The dependence of $\hat\psi_n$ on $\theta,\nu$ is omitted to simplify notation. Lemma (ref) implies $\|\hat{\mu}_n\| \leq \nu^{7/2}$ and $\hat{\Sigma}_n \leq \nu^4$, up to constants. Although $\Psi$ is unbounded, the estimates are bounded with probability $1$. In the following, $\Psi$ will be replaced with: \[ \Psi_n = \{ (\mu,\Sigma) \in \Psi \text{ s.t. } (\ref{eq:bounds}) \text{ holds} \},\] without loss of generality. The upper bounds increase rapidly. With Lemma (ref), they imply an envelope function of size $\nu^{17}$ which diverges too quickly to directly apply standard empirical process results, e.g. van1996. Instead, the results directly rely on the functional form of ((ref)) and the following assumption to derive exponential inequalities under cross-sectional and time-series dependence (Lemma (ref)).

assumption$z_t \sim P$, a distribution such that for two $0 \leq M_2,M_4 < \infty$:\\ i. $\sup_{\theta \in \Theta}\mathbb{E}_{P}(\|g(z_t;\theta)\|^2) \leq M_2$, ii. for all $(\theta_1,\theta_2) \in \Theta$, $\|g(z_t;\theta_1) - g(z_t;\theta_2)\| \leq G_t \|\theta_1-\theta_2\|$ with $\mathbb{E}_P(\|G_t\|^2) \leq M_2$, iii. $\sup_{\theta \in \Theta}\mathbb{E}_{P}(\|g(z_t;\theta)\|^4) \leq M_4$. In ii. $G_t = G(z_t)$ is either iid or strictly stationary and mixing with rate $\beta_m$ found in Assumption (ref) i.

Let $Q_\nu = \mathbb{E}_P(Q_{n})$ be the population analog of $Q_n$ without any contamination: \[Q_\nu(\psi;\theta) = \mathbb{E}_P\left[(\nu+p)\log\left(1+\|g(z_t;\theta)-\mu\|^2_{\Sigma^{-1}}/\nu\right)\right] + \log|\Sigma| + \frac{\kappa_1}{\nu}\|\mu\|^2_{\Sigma^{-1}} + \frac{\kappa_2}{\nu} \text{trace}(\Sigma).\]

propositionTake $x \geq 0$ and $1 \leq \nu \leq n$, suppose Assumptions (ref) and (ref) i-ii hold with $z_t$ iid for $t \in \{1,\dots,n_P\}$. For each $\theta \in \Theta$, let $\hat{\psi}_n(\theta;\nu)$ be the minimizer of ((ref)) and $\psi(\theta;\nu)$ the minimizer of $Q_\nu$ on $\Psi$. Set $C_n = 1 + (k + 2p^2 )[\log(p) + \log(\nu) + \log(n_P)],$ with $p = \text{dim}(g)$ and $k = \text{dim}(\theta)$ then: \begin{align*} \mathbb{P} \Bigg( \sup_{\theta\in\Theta} &\sup_{z_t \in \mathcal{O}_n, t>n_P} \left\{ Q_{\nu}(\hat\psi_n(\theta;\nu);\theta)-Q_{\nu}(\psi(\theta;\nu);\theta) \right\} \geq C_{\mathcal{O}} \frac{n_o(\nu+p)}{n}[1 + \log(n) ] \\ &+ L\frac{n_P}{n} (\nu+p)\log(1+\nu p)\left[ \sqrt{\frac{x}{n_P}} + \frac{x}{n_P} + \sqrt{\frac{C_n}{n_P}} + \frac{C_n}{n_P} \right] \Bigg) \leq 4 \exp(-x), \end{align*} for a constant $L$ which depends on $s_0,\kappa_1,\kappa_2,M_2$ and $C_{\mathcal{O}}$ depends on $s_0,M_2,\kappa_1,A,\alpha$. If $z_t$ is strictly stationary and $\beta$-mixing for $t \in \{1,\dots,n_P\}$, then: \begin{align*} \mathbb{P} \Bigg( \sup_{\theta\in\Theta} &\sup_{z_t \in \mathcal{O}_n, t>n_P}\left\{ Q_{\nu}(\hat\psi_n(\theta;\nu);\theta)-Q_{\nu}(\psi(\theta;\nu);\theta) \right\} \geq C_{\mathcal{O}} \frac{n_o(\nu+p)}{n}[1 + \log(n) ] \\ &+ \tilde{L}\frac{n_P}{n} (\nu+p)\log(1+\nu p)\left[ \sqrt{\frac{(x+C_n)x}{n_P}} + \frac{(x+C_n)x}{n_P} + \sqrt{\frac{C_n}{n_P}} + \frac{C_n}{n_P} \right] \Bigg) \leq 12 \exp(-x), \end{align*} for $\tilde{L}$ which additionally depends on the mixing coefficients $a,b$.

Because $\psi(\theta;\nu)$ is a minimizer, $Q_{\nu}(\hat\psi_n(\theta;\nu);\theta)-Q_{\nu}(\psi(\theta;\nu);\theta) \geq 0$ always holds. Proposition (ref) gives exponential inequalities for deviations from the biased solutions $\psi(\theta;\nu)$, uniformly in both parameters $\theta$ and outliers $z_t \in \mathcal{O}_n$, with respect to the loss $Q_\nu$. The bounds only require finite second moments, allowing for heavy tails under $P$. This is important in macroeconomic and financial applications since $P$ typically does not have sub-exponential, or Gaussian, tails.\footnote{Heavy-tailed distributions, unlike the exponential and Gaussian distributions, may not have all finite moments. Student and Pareto are both heavy-tailed distributions.} The worst-case contamination bias is of order $n_0(\nu+p)/n[1+\log(n)]$ which depends on the proportion of outliers $n_o/n$ and the tuning parameter $\nu$. It differs from the $(\sqrt{\nu}/n)n_o$ term for the simple estimator above. The proofs indicate that $(\nu/n)n_o$ corresponds to the influence of outliers when estimating $\Sigma$.

For iid data, similar to Bernstein's inequality, the tails are thin: approximately sub-Gaussian for small $x \ll \sqrt{n_P}$ and sub-exponential for large $x \gg \sqrt{n_P}$.\footnote{The inequality $\mathbb{P}(Z \geq \sqrt{x/n} + x/n + a_n) \leq 4\exp(-x)$ implies $\mathbb{P}(Z \geq u/\sqrt{n} + a_n) \leq 4\min[\exp(-u^2),\exp(-\sqrt{n}u)]$ which is sub-Gaussian for $u \ll \sqrt{n}$ and sub-exponential for $u \gg \sqrt{n}$.} For time-series data, the tails are thicker: approximately sub-Gaussian for $x \ll C_n$, sub-exponential for $C_n \ll x \ll \sqrt{n_P}$ and sub-Weibull for $x \gg \sqrt{n_P}$ with tail parameter $1/2$ vladimirova2020. This is comparable to Bernstein inequalities for sample means of bounded $\beta$-mixing processes in doukhan1994.

Estimating both $\mu$ and $\Sigma$ consistently requires $(\nu/n)\log(n) n_o \to 0$. This is more restrictive than $(\sqrt{\nu}/n)n_o \to 0$ which appears under local asymptotics for $\mu$. This is related to the discussion above on iterative procedures and joint estimation of $\psi$. The dependence on the number of moment conditions $p$ is made explicit to show how it affects the bounds. The $p^2$ term in $C_n$ comes from estimating $p(p+1)/2$ coefficients in $\Sigma$. For the large sample results below, the number of parameters $k$ and moments $p$ will be assumed to be fixed and finite.

Asymptotic Properties

The following builds on Proposition (ref) to derive uniform consistency and then oracle equivalence results which involve the amount of contamination $n_o$ and the bias. The large-sample results can be used to compute standard errors and compute confidence intervals the usual way (i.e. reporting $\tilde{\theta}_n \pm 1.96 \text{se}(\tilde{\theta}_n)$).

corollarySuppose the conditions for Proposition (ref), Assumption (ref) iii hold, and: \[ n_o = o\left( \frac{n}{\nu\log(n)}\right), \, \nu\log(\nu) = o\left( \sqrt{\frac{n}{\log(n)}} \right). \] Let $\psi(\theta;\infty)$ denote the pair $\mu(\theta;\infty) = \mathbb{E}_P[g(z_t;\theta)]$, $\Sigma(\theta;\infty) = \text{var}_P[g(z_t;\theta)]$, then: \[ \sup_{\theta \in \Theta} \left( \sup_{ z_t \in \mathcal{O}_n, t > n_P } \|\hat{\psi}_n(\theta;\nu) - \psi(\theta;\infty)\| \right) = o_p(1). \]

Proposition (ref) and the following two bounds: $|Q_\nu(\psi;\theta) - Q_\infty(\psi;\theta)| \leq O(\nu^{-1})$ and $\|\psi(\theta;\nu)-\psi(\theta;\infty)\| \leq O(\nu^{-1})$, uniformly in $\theta$, imply the uniform consistency result above. Taking the supremum over $\mathcal{O}_n$ ensures the result is robust against the least favorable outliers.

propositionSuppose the conditions of Corollary (ref) hold. Let $\max[\mathbb{E}_P(\|g(z_t;\theta_0)\|^{r + \delta}),\mathbb{E}_P(|G_t|^{r + \delta})] := M_{r,\delta}$ for $r \geq 1$ and $\delta >0$. Let $\overline{g}_{n_P}(\theta) = \frac{1}{n_P} \sum_{t=1}^{n_P} g(z_t;\theta)$, if $M_{3,\delta}$ is finite for some $\delta>0$: \begin{align*} \sup_{\theta \in \Theta} \left( \sup_{ z_t \in \mathcal{O}_n, t > n_P } \| \hat{\mu}_{n}(\theta;\nu) - \overline{g}_{n_P}(\theta) \| \right) = O_p\left(\max\left[\frac{1}{\nu},\frac{\sqrt{\nu}n_o}{n}\right] \right) \end{align*} If, in addition $M_{5,\delta}$ is finite for some $\delta>0$: \begin{align*} \sup_{\theta \in \Theta} \left( \sup_{ z_t \in \mathcal{O}_n, t > n_P } \| \tilde{\mu}_{n}(\theta;\nu) - \overline{g}_{n_P}(\theta) \| \right) = O_p\left(\max\left[\frac{1}{\nu^2},\frac{\sqrt{\nu}n_o}{n}\right]\right). \end{align*}

Using the same two inequalities, and a bound on the score, Proposition (ref) shows that the robust and bias-corrected estimates are uniformly close to an oracle that computes the sample mean using only the good $n_P$ observations. An empirical researcher might want to trim out outliers without altering, as much as possible, the rest of the sample. This oracle result precisely states this property. In that sense, it gives a more desirable characterization than limit theorems for $\hat{\mu}_n(\theta;\nu) - \mu(\theta;\infty)$ and $\tilde{\mu}_n(\theta;\nu) - \mu(\theta;\infty)$.

Similar to non-parametric regressions which derive bias from smoothness, stronger moment conditions are needed to derive faster rates of convergence. Without outliers, OLS estimates are asymptotically normal for iid data when $\mathbb{E}_P(\|x_t y_t\|^2),\mathbb{E}_P(\|x_t \|^4) < \infty$. Here the condition is more restrictive, it reads $\mathbb{E}_P(\|x_t y_t\|^{5+\delta}),\mathbb{E}_P(\|x_t \|^{10 + 2\delta}) < \infty$.

The worst-case impact of outliers is of order $(\sqrt{\nu}/n)n_o$, with and without bias correction. Note that the estimator $\hat{\mu}_n$ is “redescending.” The maximal influence of a single observation $z$ given by $(\sqrt{\nu/2})/n$, is attained at $\|g(z;\theta)-\mu\|_{\Sigma^{-1}} = \sqrt{\nu}$ and then monotonically declines to zero as $\|g(z;\theta)-\mu\|_{\Sigma^{-1}}$ increases.\footnote{This is also discussed in mcdonald1988, huber2011.} The result requires $\hat{\Sigma}_n(\theta;\nu)$ uniformly convergent. Importantly, the influence function is not redescending for $\Sigma$: it is strictly increasing and bounded above by $\nu > \sqrt{\nu}$. Hence, consistency of $\hat{\Sigma}_n$ is more restrictive: $(\nu/n) n_o \to 0$.

assumptioni. $\mathbb{E}_P[g(z_t;\cdot)]$ is continuously differentiable in $\theta \in \Theta$, ii. $\mathbb{E}_P[g(z_t;\theta)]=0$ if, and only if, $\theta = \theta_0 \in \text{int}(\Theta)$, iii. $G(\theta_0) := \partial_\theta \mathbb{E}_P[g(z_t;\theta_0)]$ has full rank, iv. for any $\delta_{n_P} \to 0$, $\sup_{\|\theta - \theta_0\|\leq \delta_{n_P}}\sqrt{n_P}\|\overline{g}_{n_P}(\theta)-\overline{g}_{n_P}(\theta_0) - \partial_\theta \mathbb{E}_P[g(z_t;\theta_0)](\theta-\theta_0)\|/[1+\sqrt{n_P}\|\theta-\theta_0\|] = o_p(1)$, v. $\sqrt{n_P}\overline{g}_{n_P}(\theta_0) \overset{d}{\to} \mathcal{N}(0,\Sigma_0)$, vi. $W_n \overset{p}{\to} W$ positive definite.

Assumption (ref) repeats conditions from newey1994, only for the good $n_P$ observations. They imply consistency and asymptotic normality of $\hat{\theta}_{n_P}$, an oracle estimator which uses only the good $n_P$ datapoints.

theoremSuppose Assumption (ref) and the conditions of Proposition (ref) hold with $M_{5,\delta}$ finite for some $\delta >0$. Suppose $n_o$ and $\nu$ are such that: \[ \frac{\sqrt{n}}{\nu^{2}} = o(1), \text{ and } \sqrt{\frac{\nu}{n}}n_o = o(1).\] Let $\hat{\theta}_{n_P} = \text{argmin}_{\theta \in \Theta} \|\overline{g}_{n_P}( \theta)\|_{W_n}$, the estimator $\tilde{\theta}_n$ satisfies: \begin{align*} &\sup_{z_t \in \mathcal{O}_n, t > n_P}\|\sqrt{n_P}(\tilde{\theta}_n - \hat{\theta}_{n_P})\| = o_p(1), \quad and \quad \sqrt{n_P}(\tilde{\theta}_n - \theta_0) \overset{d}{\to} \mathcal{N}(0,V) \end{align*} for any sequence $z_t \in \mathcal{O}_n$, $t=n_P+1,\dots,n$, where $V = (G^\prime W G)^{-1} G^\prime W \Sigma_0 W G (G^\prime W G)^{-1}$, $G = \partial_\theta \mathbb{E}_P[g(z_t;\theta_0)]$.

Theorem (ref) presents the main result: the bias-corrected estimates are asymptotically equivalent to the oracle $\hat{\theta}_{n_P}$. They inherit its asymptotic properties. The supremum over $\mathcal{O}_n$ ensures robustness against least favorable outliers. The bias is asymptotically negligible if $\nu^2 = o(\sqrt{n})$. If $n_o$ were known, setting $\nu \asymp (n/n_o)^{2/5}$ would achieve the optimal rate in Proposition (ref). For this choice of $\nu$, the condition $\sqrt{n}/\nu^2 = o(1)$ reads $n_o = o(n^{3/8})$. Setting $\nu = O(n^{1/4} \log(n))$ is nearly optimal when $n_o$ becomes arbitrarily close to this bound as it requires implies $n_o = o( n^{3/8}/\sqrt{\log(n)} )$. A data-driven rule is given below to select $\nu$ in practice while enforcing this rate. The large sample properties of $\tilde{\raisebox{0pt}[0.85\height]{$\tilde{\mu}$}}_n(\theta;\nu) = 2 \tilde{\mu}_n(\theta;\nu) - \tilde{\mu}_n(\theta;\nu/2)$, from Section (ref), and the resulting $\tilde{\raisebox{0pt}[0.85\height]{$\tilde{\theta}$}}_n$ follow from those of $\tilde{\mu}(\theta;\nu)$, $\tilde{\theta}_n$.

Assumption (ref) requires $\theta$ to be strongly identified using the good $n_P$ observations. Given that Proposition (ref) does not restrict identification status, one should compute an identification-robust test statistic - e.g. Anderson-Rubin (AR) - from the robust bias-corrected moment estimates. This is related to klooster2023 who consider AR test statistics with bounded influence curve.

With the oracle result (Proposition (ref)) and the regularity conditions (Assumption (ref)), further results could be derived. One could consider two-step GMM with robust weighting $W_n = \hat{\Sigma}_n(\tilde{\theta}_n;\nu)^{-1}$ in the second step, robust overidentifying restrictions, quasi-Likelihood Ratio and Lagrange multiplier tests, etc. This is not pursued here.

propositionSuppose the assumptions for Theorem (ref) hold. For each $\theta \in \Theta$ and $\nu >0$, the estimates $\hat{\mu}_n(\theta;\nu)$, $\tilde{\mu}_n(\theta;\nu)$ satisfy: \[ \hat{\mu}_n(\theta;\nu) = \sum_{t=1}^n \omega_t(\theta;\nu) g(z_t;\theta), \quad \tilde{\mu}_n(\theta;\nu) = \sum_{t=1}^n \tilde{\omega}_t(\theta;\nu) g(z_t;\theta) \] where the weights are given by $\omega_t(\theta;\nu) = \frac{ (1+p/\nu)/n[1+q_t(\theta;\nu)/\nu]^{-1}}{(1+p/\nu)/n\sum_{t=1}^n [1+q_t(\theta;\nu)/\nu]^{-1} + \kappa_1/\nu }$ and $\tilde{\omega}_t(\theta;\nu) = 2\omega_t(\theta;\nu) - \omega_t(\theta;\nu/2)$ using $q_t(\theta;\nu) = \|g(z_t;\theta)-\hat{\mu}_n\|^2_{\hat{\Sigma}_n^{-1}}$. Let $\hat{\varepsilon}_t(\theta) = g(z_t;\theta) - \hat{\mu}_n(\theta;\nu)$, $\tilde{\varepsilon}_t(\theta) = g(z_t;\theta) - \tilde{\mu}_n(\theta;\nu)$. The following weighted variance estimators are consistent: \begin{align*} &\hat{\Sigma}_{n,\omega}(\theta) = \sum_{t=1}^n \omega_t(\theta;\nu) \hat{\varepsilon}_t(\theta) \hat{\varepsilon}_t(\theta)^\prime \overset{p}{\to} \Sigma(\theta), \quad \tilde{\Sigma}_{n,\omega}(\theta) = \sum_{t=1}^n \tilde{\omega}_t(\theta;\nu) \tilde{\varepsilon}_t(\theta) \tilde{\varepsilon}_t(\theta)^\prime \overset{p}{\to} \Sigma(\theta),\end{align*} where $\Sigma(\theta) = \text{var}_P( g(z_t;\theta) )$ here denotes the short-run variance under $P$.

For cross-sections and serially uncorrelated moments, $\hat{\Sigma}_n(\hat{\theta}_n,\nu)$ can be used to estimate $\Sigma_0$ in Theorem (ref). Because of the penalty $\kappa_2 >0$ on $\Sigma$, it tends to be downward biased. An alternative is to use the same weights as $\tilde{\mu}_n$ to match the properties of the estimator more closely. Proposition (ref) above shows that such an estimator is also consistent for the short-run variance. Long-run variance estimates, required for serially correlated moments, are not considered here.

The weighted average representation further implies, for linear models, that $\tilde{\theta}_n$ are weighted least-squares estimates since $\tilde{\mu}_n(\tilde{\theta}_n;\nu) = \sum_{t=1}^n \tilde{\omega}(\tilde{\theta}_n;\nu) x_t( y_t - x_t^\prime \tilde{\theta}_n ) = 0 \Rightarrow \tilde{\theta}_n = ( \sum_{t=1}^n \tilde{\omega}(\tilde{\theta}_n;\nu) x_t x_t^\prime )^{-1} \sum_{t=1}^n \tilde{\omega}(\tilde{\theta}_n;\nu) x_t y_t$. The weighting can be used to interpret the results.

\paragraph{Data-driven choice of tuning parameter $\nu$.} The following describes a data-driven procedure to select the tuning parameter $\nu$. Take $0<a_0<a_1<\dots<a_J$ and $\nu_j = a_j n^{s}$ with $1/4 < s < 1/2$ or $\nu_j = a_j n^{s} \log(n)$ with $1/4 \leq s < 1/2$. The simulated and empirical examples use $s = 1/4$ and $0.5 = \log(a_0) < \dots < \log(a_J) = 35$ so that each $\nu_j = O(n^{1/4}\log(n))$ satisfies the requirements for Theorem (ref).

Using $\nu = \nu_0$ as a baseline, compute a preliminary estimate $\hat{\theta}_n$ and the corresponding moment estimates $\hat{\psi}_n(\hat{\theta}_n;\nu_0)$. In the absence of outliers, it can be shown that $|Q_n(\hat{\psi}_n(\hat{\theta}_n;\nu_0);\nu_j) - Q_n(\hat{\psi}_n(\hat{\theta}_n;\nu_0);\infty)| = O_p(\nu_j^{-1})$. This implies that, in the absence of outliers, the fit should be comparable accross different values of $\nu$: $|Q_n(\hat{\psi}_n(\hat{\theta}_n;\nu_0);\nu_j) - Q_n(\hat{\psi}_n(\hat{\theta}_n;\nu_0);\nu_0)| \leq O_p(\nu_0^{-1})$. The selection rule picks the largest value of $\nu_j$ such that the fit remains comparable: \[ \hat{\nu}_n = \text{max} \left\{ \nu_j, \text{ s.t. } |Q_n(\hat{\psi}_n(\hat{\theta}_n;\nu_0);\nu_j) - Q_n(\hat{\psi}_n(\hat{\theta}_n;\nu_0);\nu_0)| \leq \frac{1+\log(n)}{\nu_0} \right\}. \] The parameters and the moments are only estimated once, to reduce computation, at the smallest $\nu_0$ which produces the most robust estimate of the grid $\nu_0,\dots,\nu_J$.

By design, $\hat{\nu}_n$ has rate $O(n^{s})$ or $O(n^{s} \log(n))$ which satisfies the conditions of Theorem (ref) given restrictions on $n_o$. The following heuristic motivates the choice of criteria. As discussed above, the outliers have an asymptotic impact on non-robust estimates if $n_o n^{\alpha}/n \not \to 0$, and the estimator is robust as long as $n_o = o(\sqrt{n/\nu})$. Set $n_o = c \sqrt{n/\nu_0}$ then the sum over outliers in $Q_n(\cdot;\nu)$ increases proportionally to $\nu c \sqrt{\nu_0/n}\log(1+n^{2\alpha}/\nu) \sim c \nu \sqrt{\nu_0/n}\log(n)$. The change over the $n_P$ terms is a $O_p(\nu_0^{-1})$. For $\nu_0$ relatively small, the upper bound in the criteria above conservatively minors the sum of these two bounds.

Simulated and Empirical Applications

All the estimations below use the same $\kappa_1 = \kappa_2 =10^{-2}$, giving wide bounds in Lemma (ref). With the data-driven choice of $\hat{\nu}_n$, the results are not too sensitive to this choice of penalty.

Simulated Example

To illustrate the finite sample properties of the procedure, consider a linear regression $y_t = x_t^\prime \theta_0 + e_t$. There are three regressors $x_t = (1,x_{1t},x_{2t},x_{3t})$, each $x_{jt}$ and $e_t$ is drawn from $(\chi^2_5 - 5)/\sqrt{10}$ has mean zero and unit variance, $\theta_0 = (0,1,1,1)$. Sample size is $n=150$, several $n_o = 0,1,5,10$ are reported where each outlier has $x_{jt} = \sqrt{n}$ and $y_t = x_t^\prime \theta_\dagger$, $\theta_\dagger = (0,1/2,1/2,1/2)$. In this example, outliers are leveraged to mimic the motivating example.

The simulations compares full sample $\hat{\theta}_{n}^{ols}$, an oracle which discards outliers $\hat{\theta}_{n_P}^{ols}$, R's robust regression estimates $\hat{\theta}_{n}^{rlm}$ with $\hat{\theta}_{n}$, $\tilde{\theta}_{n}$, $\tilde{\raisebox{0pt}[0.85\height]{$\tilde{\theta}$}}_{n}$ computed using $\hat{\nu}_n$ as described above. A further $\hat{\theta}^{un}_{n}$ is computed using $\hat{\nu}_n^2$ to illustrate undersmoothing as opposed to bias correction used in this paper. $\tilde{\raisebox{0pt}[0.85\height]{$\tilde{\theta}$}}_{n}$ applies the correction step twice as discussed at the end of Section (ref).

table[table omitted — 5,601 chars of source]

Table (ref) shows that without outliers ($n_o=0$) the performance of bias-corrected and undersmoothed estimates is comparable to full sample OLS. The robust M-estimates of the intercept $\theta_0$ are biased, because the errors are skewed. The performance of OLS degrades as soon as $n_o = 1$, as expected. The undersmoothed and rlm estimates are also less accurate. The non-corrected estimates $\hat{\theta}_n$ are more robust but biased. Bias correction, $\tilde{\theta}_n$ and $\tilde{\raisebox{0pt}[0.85\height]{$\tilde{\theta}$}}_n$, improves accuracy and rejection rates. The estimators still perform well for $n_o = 5$. Performance degrades for $n_o = 10$. This is perhaps not too surprising since $\log(n_o)/\log(n) \simeq 0.45 > 3/8$ for $n_o = 10$. Additional results for $n = 500$ are reported in Table (ref), Appendix (ref). Tables (ref), (ref) has results with $\nu =O(n^{1/3})$ in the same Appendix.

Empirical Applications

Two empirical applications further illustrate robust-GMM estimator in instrumental variable regression settings.

Trade Openness and Inflation

The second empirical application is also inflation-related. romer1993 estimates the relationship between trade openness and inflation using country time averages between 1973 and 1993. Trade openness, measured by the share of imports to GDP, can be considered as endogenous given that monetary policy affects both inflation and exchange rates. He considers the following specification: \[ \pi_t = \theta_0 + \theta_1 \text{open}_t + \theta_2 \log(\text{pcinc})_t + e_t, \] where $\pi$ measures inflation, $\text{pcinc}$ is per-capita income in 1980, assumed exogenous. romer1993 further adds dummies in some specifications, these are not included here. The instrument for openness is $\log(\text{land})$ measuring the $\log$ of the square-mile surface of the country. The idea is that smaller land area economics should be more open to imports. romer1993 notes that “A few countries in the sample have extremely high average inflation rates.” and is concerned that “the parameter estimates from a linear regression would be determined almost entirely by a handful of observations.” As a remedy, he estimates the regression using the log of average inflation $\log(\pi/100)$. The influence of outliers in linear IV regressions is not intuitive because leverage can be either positive or negative (Lemma (ref)). As a result, unlike OLS, the influence may not have the same sign as the residual: the impact of an outlier is less predictable than with OLS.

table[table omitted — 1,907 chars of source]

Similar to the motivating example, Table (ref) provides diagnostics for both specifications. For $y = \pi/100$, the greatest contributors tend to be severely indebted countries that were particularly affected by the 1980s debt crisis. terra1998 argues that these countries overborrowed in the 1980s and had “less pre-commitment in monetary policy” resulting in higher inflation during the debt crisis.\footnote{terra1998 classifies Argentina, Bolivia, Brazil, Peru, Mexico, Zaire as severely indebted.} In contrast, for $y = \log(\pi/100)$, the greatest contributors are less indebted and other countries terra1998 which have low average inflation.\footnote{Singapore is the country with the lowest average inflation in the sample.} The log increases the influence of low-inflation countries, as one might expect.

table[table omitted — 2,331 chars of source]

The kurtosis indicates the $\log$-transformed regression is less prone to outliers, the standard deviation suggests the estimates will be significantly less accurate. This reflects the larger volatility of log-inflation compared to inflation. Also, the log transformation changes the interpretation of the coefficient $\theta_1$ which may not be desirable. The following replicates the original results and estimates the regression in levels, as in wooldridge2002, to get the desired coefficient interpretation.

Table (ref) confirms that the log-transformed regression is less prone to outliers as the IV and robust estimates are very similar after bias-correction.\footnote{Estimates using a smaller $\nu = 12$ are nearly identical for the log regression (not reported here).} The non-transformed regression is, as romer1993 suspected, sensitive to some datapoints. Robust and bias-corrected estimates indicate IV overestimates the relationship between trade openness and inflation. Standard errors indicate the bias-corrected estimates are more accurate than the IV ones. The estimated effect is about one-third of the non-robust one. The bias correction adjusts the estimates by half to a full standard error. The full dataset of weights used to compute the estimates when $y = \pi/100$ are reported in Tables (ref), (ref), Appendix (ref).

Segregation and the Quality of Government

The third application considers the relationship between racial and religious discrimination and the quality of government. alesina2011 constructed a new dataset on ethnic, linguistic, and religious segregation and fractionalization for a large number of countries. Mobility within a country, which determines segregation, can be endogenous to government quality. To address this particular issue, the authors predict segregation from neighboring country data. The main idea is that when a sub-population is at the border of the neighboring country, the same sub-group is more likely to be located near that border alesina2011. They illustrate using Switzerland as an example: most French speakers live near the French border, and Protestants are more commonly found near the German border. This is one of the papers surveyed in young2022, which finds that published IV regressions tend to be highly leveraged and sensitive to a few observations. The following revisits some of the main results in the original paper. The regression specification is given by: \[ \text{Rule of law}_i = \theta_0 + \theta_1 \text{Segregation}_i + \theta_2 \text{Fractionalization}_i + \text{Controls} + u_i, \] where Segregation and Fractionalization are measured with respect to one of Ethnicity, Language, or Religion leading to three separate IV regressions. The controls are the same as in Table 6, Column 2 of alesina2011. Fractionalization controls for group heterogeneity in each dimension (ethnicity, language, and religion) as measured by a Herfindahl index. If there is only one group in the population, the index is zero. If there are many equal-sized groups, the measure is closer to $1$. See alesina2011 for further details.

table[table omitted — 2,015 chars of source]

Table (ref) shows the 10 highest contributors for the coefficient $\theta_1$ in each regression as well as the sample moments of coefficient contribution. The regressions for ethnicity and language display somewhat heavy tails, as measured by the kurtosis. This indicates that the baseline results -- estimated coefficients, standard errors, or both -- may be sensitive to a few observations. Table (ref) reports standard IV and robust estimates with(out) bias correction. Robust estimates tend to produce more precise inferences, as measured by standard errors. The baseline results indicate that both ethnic and language segregation have a significant, negative impact on the rule of law in a given country.

Robust results indicate that ethnic segregation if the only significant determinant of the rule of law. Unlike the previous example, the estimate implies a larger effect than standard IV. Notice that because the controls are correlated with the instrument, the direction of the change from standard to robust estimates does not necessarily coincide with the contribution to $\theta_1$ in Table (ref).\footnote{If an outlier affects a coefficient on the controls and there is collinearity with the instrument, then robust estimates of $\theta_1$ will change with the coefficients on the controls, as they are correlated. Diagnostics may not fully reflect the multivariate effect of the outliers. In addition, when there are multiple outliers, the direction of change depends on the combined effect of the outliers.} Diagnostics can inform if the results are sensitive to some observations but, as explained in huber2011, are not a substitute for robust estimation. Although non-significant, the coefficient $\theta_2$ for fractionalization does change from positive to negative in the first two regressions. Again, the full dataset of weights used in the three regressions is reported in Tables (ref), (ref), (ref) of Appendix (ref).

table[table omitted — 2,581 chars of source]

Conclusion

It is important to assess the robustness of empirical findings. Without symmetry restrictions, large differences between robust and non-robust estimates could be attributed to 1) improved resilience, or 2) significant asymmetry bias (or a combination of the two). This paper proposes a procedure with a simple asymptotic bias correction so that 2) is less likely. Reporting the implicit estimation weights makes the final results transparent and interpretable. This is illustrated in three empirical applications.