EconBase
← Back to paper

Subsample-Based Estimation under Dynamic Contamination

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

86,981 characters · 14 sections · 43 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Subsample-Based Estimation under Dynamic Contamination

\pagenumbering{arabic}

abstractThis paper studies a structural failure of subsample-based estimation in dynamic time series models. Even under oracle knowledge of contamination locations, removing contaminated observations does not restore the uncontaminated objective. In such settings, contamination propagates through the residual filter and distorts the estimation criterion. As a result, subsample-based estimators are generically inconsistent for the clean-data parameter. We characterise this failure as a structural incompatibility between pointwise subsampling and residual propagation. More generally, the failure arises whenever contamination propagates through transformations that enter the estimation criterion, with dynamic time series models as a leading example. To address it, we propose a propagation-compatible transformation of index sets via a patch removal operator. Under general high-level conditions, this transformation leaves the estimator asymptotically unchanged under the uncontaminated model while restoring consistency under contamination. The results apply to a broad class of residual-based estimators and do not rely on modelling the contamination process.

Keywords: Outliers; Residual propagation; Robustness; Consistency; Asymptotic theory.

Introduction

Outliers are a pervasive feature of economic and financial time series, and more generally of data in many statistical applications. Their primary statistical effect is not merely to generate a few anomalous observations, but to distort inference about the underlying data generating mechanism. Such contamination may induce spurious statistical significance Huber1964, Tsay1986, ChenLiu1993, DaviesGather1993 and create false evidence of structural features such as volatility clustering FransesGhijsels1999, CarneroPenaRuiz2007, structural breaks, regime switching Andrews1993, vDFL99, or persistence and unit-root-like behaviour, even when the clean process exhibits no such features. From an econometric perspective, the object of interest is typically the latent clean process rather than the contaminated observations themselves. Outliers act as external perturbations that may mimic or obscure level shifts and variance changes ChenLiu1993, BoxTiao1975, Tsay1988, FransesHaldrup1994. We refer to such properties of the underlying data-generating mechanism as structural features. If contamination causes the observed series to appear to exhibit a structural feature that is absent from the clean process, inference based on the observed data is misleading. Reliable inference therefore requires consistent estimation of the clean model under contamination, since asymptotic approximations, hypothesis testing, model selection, and specification searches all depend on an accurate characterisation of the underlying dynamics JohansenNielsen2009, DoornikHendry2016.

A common and widely used approach in this setting is to construct estimators based on subsets of the data intended to exclude contaminated observations. Such subsample-based procedures operate without explicit modelling of the contamination mechanism and avoid imposing assumptions on the contamination process itself. They form a standard paradigm for robust estimation when the objective is to recover the clean-data mechanism Huber1964, Ro84, Maronna2006, JohansenNielsen2016. This objective differs fundamentally from approaches that incorporate contamination directly into the data-generating mechanism, where contamination is treated as part of the model and therefore requires additional distributional and structural assumptions for specification and identification. The present paper adopts the former perspective and treats contamination as an external nuisance rather than as part of the data-generating mechanism. Their use in time series, however, implicitly rests on the idea that removing contaminated observations restores the uncontaminated optimisation problem. As we demonstrate below, contamination propagates through the residual filter that defines the estimation criterion, inducing non-local distortions in the residual sequence and hence in the criterion itself.

We show that subsample-based procedures do not generically provide consistent estimation of the clean-data mechanism in dynamic time series models, even under oracle knowledge of contamination locations. In cross-sectional settings, contamination is local, so removing contaminated observations restores the clean-data objective. In dynamic models, however, this equivalence fails. As a result, robustness cannot be achieved by pointwise removal of contaminated observations, but requires control of propagated distortions in the residual-based criterion.

We establish a general inconsistency result for subsample-based procedures under dynamic contamination, showing that this failure is structural and arises from residual propagation through the filtering mechanism rather than from imperfect identification of contaminated observations. We characterise this mechanism by deriving an exact representation of additive and innovative contamination after passage through the residual filter, which induces non-local distortions in the residual-based criterion. We show that this implies a propagation-compatible transformation of index sets that removes the residual footprint of contamination, yielding a patch removal operator governed by the model-implied filter rather than by an ad hoc trimming rule and applicable to a broad class of residual-based procedures without modifying their internal structure. We establish asymptotic invariance under the transformation in the uncontaminated case and asymptotic equivalence between contaminated and uncontaminated estimators once residual propagation is controlled, thereby restoring consistency for the clean-data parameter under contamination.

Existing work on outliers in time series has focused on their effects on residuals, diagnostics, and detection procedures. Early work, beginning with fox_outliers_1972 and later surveyed by MARTIN1979147 and MaYo86, studies additive and innovative outliers and their effects on residuals, diagnostics, and influence functions in dependent data. Classical model-based approaches include chang_estimation_1988 and ChenLiu1993, who analyse outlier effects and develop procedures for parameter estimation and joint estimation of model and outlier effects in ARMA-type settings. More recent contributions, particularly JohansenNielsen2016, develop a unified asymptotic theory for outlier detection procedures, formalising their behaviour through the notion of gauge and establishing calibration results for the fraction of observations flagged as outliers.

These detection-based procedures are closely related to residual-based screening, down-weighting, and trimming methods, including M-type estimators Huber1964, WelshRonchetti2002, least trimmed squares (LTS) and related high-breakdown estimators Ro84, Visek2006a, and forward search methods HadiSimonoff1993, AtkinsonRiani2000. While this provides a rigorous basis for detection and tuning, these procedures operate at the level of individual observations or residual magnitudes. They proceed through identification and removal, down-weighting, or trimming of aberrant observations based on residual diagnostics, a well-known difficulty in time series settings Ljung93. Correct identification or asymptotic calibration of contaminated observations is not sufficient to restore the uncontaminated optimisation problem. The difficulty is structural rather than algorithmic. Even when contaminated observations are correctly excluded, residual propagation continues to distort the residual-based criterion through the dynamic filter. The resulting optimisation problem therefore no longer targets the clean-data parameter asymptotically. This limitation also applies to robust estimation for multivariate dynamic models, such as VARMA settings GarciaBenMartinezYohai1999, and to approaches based on modified residual constructions that limit the effect of individual outliers MulerPenaYohai2009. Despite decades of work on outlier detection and robust estimation in time series, the question of whether consistency for the clean-data parameter can be restored under dynamic contamination without modelling the contamination process has remained open. The central issue is therefore not whether contaminated observations can be identified, but whether removing them restores the uncontaminated estimation problem.

JohansenNielsen2016 emphasise clustered contamination and $\varepsilon$-contamination as important directions for future research. We show that even sparse contamination, when combined with temporal dependence, can undermine consistency through residual propagation, and we provide a general correction that restores consistency for the clean-data parameter under high-level conditions that apply broadly across residual-based estimators.

The remainder of the paper is organised as follows. Section (ref) introduces the contaminated VARMA framework and characterises residual propagation. Section (ref) develops the general estimation framework and formalises the propagation-compatible transformation of index sets. Section (ref) discusses the verification of the assumptions and establishes sufficient conditions under which they hold in standard settings. Section (ref) reports simulation evidence. Section (ref) presents a feasible implementation. Section (ref) provides an empirical illustration. Section (ref) concludes.

Contamination and Its Residual Propagation

Let $\{x_t\}_{t\in\mathbb Z}$ be a $d$-dimensional stochastic process generated by a vector autoregressive moving average (VARMA) model

equation[equation omitted — 75 chars of source]

where $\phi_0(L)$ and $\theta_0(L)$ are lag polynomials with orders $p$ and $q$ and $\varepsilon_t$ is a $d$-dimensional zero mean innovation vector.

In this paper, we follow Definitions 3.1.3 and 3.1.4 in brockwell_time_1991. The VARMA model (ref) is said to be causal if and only if it admits the representation

equation[equation omitted — 66 chars of source]

where $\psi_0(L) = \phi_0^{-1}(L)\theta_0(L)$. It is said to be invertible if and only if it admits the representation

equation[equation omitted — 65 chars of source]

where $\pi_0(L) = \theta_0^{-1}(L)\phi_0(L)$.

Let $\mathcal{V}$ denote the parameter space and let $\varphi_0=(\phi_0,\theta_0)\in\mathcal{V}$ denote the true parameter vector of the data generating process. We restrict the parameter space to values for which the VARMA model is causal and invertible. Under these conditions, the associated filters $\psi(L)$ and $\pi(L)$ exist and are absolutely summable.

To focus on the stochastic dynamic structure, we consider a mean-zero VARMA process without deterministic components. In practice, deterministic terms such as intercepts, trends, seasonal components, or exogenous regressors can be removed by standard preprocessing and are therefore not included in the present formulation.

Outliers are commonly modelled as either additive outliers (AO) or innovative outliers (IO) fox_outliers_1972, AbrahamBox1979, chang_estimation_1988. We begin with the replacement contamination model denby_robust_1979, bustos_robust_1986, MaYo86, which takes the form

equation[equation omitted — 82 chars of source]

where $\delta_t$ is a binary indicator process taking values in $\{0,1\}$, and $\xi_t$ denotes an outlying observation whose distribution does not conform to that of $x_t$. Under this model, the observed series is $y_t$, while $x_t$ remains latent. Given the potentially large magnitude of $\xi_t$, it is convenient to represent contamination in an additive form, leading to

equation[equation omitted — 74 chars of source]

where $\zeta_t = \xi_t - x_t$ captures the deviation of the outlier from the latent process. This specification is referred to as the AO model.

In contrast, under the IO model, the outlier enters through the innovation process and propagates to future observations via the lag structure fox_outliers_1972, chang_estimation_1988,

equation[equation omitted — 92 chars of source]

In this case, the effect of an outlier is transmitted through the AR dynamics, resulting in persistent deviations over time. Unlike the replacement model (ref), the IO model therefore allows for time-lagged propagation of contamination.

A key feature of time series models is that contamination at time $t$ affects not only $x_t$, but also the residuals at subsequent times through the dynamic structure of the model. Let $e_t(\varphi)$ denote the residual computed from the latent uncontaminated process $\{x_t\}$ under a candidate parameter $\varphi \in \mathcal{V}$, so that $e_t(\varphi_0)=\varepsilon_t$. Let $\tilde e_t(\varphi)$ denote the residual computed from the contaminated observations $\{y_t\}$. Even when contamination occurs at a single time point, the resulting distortion in the residual sequence is generically not local in time. Because residuals are obtained by filtering the observed series through the model dynamics, contamination entering the data propagates through this filter and affects subsequent residuals.

The following proposition provides an exact characterisation of how contamination propagates through the residual filter.

propositionUnder AO contamination, the residual satisfies \begin{equation} \tilde e_t(\varphi) = e_t(\varphi) + \pi(L)\,\delta_t \zeta_t . \end{equation} Under IO contamination, the residual satisfies \begin{equation} \tilde e_t(\varphi) = e_t(\varphi) + \pi(L)\phi_0^{-1}(L)\,\delta_t \zeta_t . \end{equation}

The proposition characterises how contamination propagates in the residual process. This propagation mechanism is distinct from the more commonly studied propagation in the observed series, which we refer to as level propagation. Level propagation refers to how contamination appears in the observed series itself, while residual propagation refers to how the same contamination appears after the residual filter is applied. The latter is the directly relevant object for any estimator that can be expressed as a function of the residual sequence. Consequently, the observed series alone does not reveal the full effect of contamination on estimation.

sidewaysfigure\caption{Uncontaminated and contaminated series in AR(1), MA(1), and ARMA(1,1) models under AO contamination (left panel) and IO contamination (right panel). Within each panel, rows correspond to the different models, the left column displays the observed series, and the right column displays the corresponding residuals computed using the true model parameters. The shaded band marks the time of the outlier.}

Figure (ref) illustrates these propagation mechanisms for AR(1), MA(1), and ARMA(1,1) processes under AO and IO contamination. Evaluating the residuals at the true parameter value $\varphi_0$ isolates the contamination mechanism by removing the effect of parameter estimation and provides a benchmark in which parameter uncertainty is absent. If residual sequences differ under this benchmark, consistency under contamination cannot be taken for granted.

The propagation patterns described above depend critically on the dynamic structure of the model. Under the causality and invertibility conditions imposed earlier, the filters $\psi(L)$ and $\pi(L)$ are absolutely summable, so that the effect of a single contamination decays over time in magnitude, even when the associated patch is of infinite length. This decay property is fundamental for the subsequent analysis, since it implies that residual distortions induced by contamination become asymptotically negligible beyond a sufficiently large horizon. Without such attenuation, the effect of contamination would persist indefinitely and consistency could not be restored through trimming or patch removal. Thus, causality and invertibility are not merely technical assumptions, but ensure that contamination propagation through the residual filter remains asymptotically controllable and thereby make consistency restoration possible.

We introduce the notion of patches to formalise the propagation effects described above. An outlier patch corresponds to the consecutive set of time indices affected by the propagation of a single outlier through the time series dynamics. If the propagation affects the observed series itself, the resulting set of indices is called an aberrant level patch, or simply a level patch, corresponding to level propagation. If the propagation affects the residual sequence, we obtain an aberrant residual patch, or residual patch, corresponding to residual propagation, defined as a consecutive set of indices for which the residuals do not coincide with the structural errors $e_t(\varphi_0)$ and therefore do not conform to the underlying data-generating mechanism. In both cases, the patch is generated by dynamic propagation from a single outlier, even though no further outliers are present within the patch itself.

Under AO contamination, the observed series differs from the uncontaminated series only at the outlying observation, so the level effect is local. However, the residual filter produces a residual patch in the residual sequence, that is, a set of indices over which the residuals do not coincide with the structural errors $e_t(\varphi_0)$. In an AR($p$) model this residual patch has finite length determined by the AR order $p$. In contrast, when a moving-average component is present, the inverse moving-average filter generates a residual patch of infinite length.

Under IO contamination the propagation pattern differs. In AR models the level effect propagates through the AR dynamics and produces a level patch of infinite length, whereas the residual effect remains local. In MA models the level effect is local, whereas the residual filter produces a residual patch of infinite length. In ARMA models both AR and MA propagation mechanisms operate, so that both level and residual patches are infinite. These propagation patterns are summarised in Table (ref).

table[table omitted — 394 chars of source]

Standard robust estimation procedures rely on trimming or down-weighting observations associated with large residuals, see, e.g., Huber1964, Ro84, Maronna2006. In cross-sectional settings, contamination is local. In dynamic time series models, however, Proposition (ref) shows that contamination propagates through the residual filter and induces non-local distortions in the residual sequence. Removing only directly contaminated observations therefore does not eliminate these distortions.

Propagation therefore affects the optimisation problem directly. Even at the true parameter, contaminated residuals differ from the structural errors underlying the clean criterion, so the usual consistency argument fails. The figures are evaluated at $\varphi_0$ and therefore represent the most favourable benchmark. In practice, $\varphi_0$ is unknown and must be estimated from contaminated data, so the effect of contamination may be more severe. This motivates the need to control residual propagation in the estimation procedure.

The preceding analysis reveals a structural feature of dynamic contamination. Residual propagation introduces a non-negligible perturbation into the criterion. As a consequence, consistency is not a default property and requires additional structure to control the propagated component. Such control may arise from decay of the underlying filter, sparsity of contamination, or an explicit correction of the retained subset, as developed in the next section.

Subsample-based Estimation under Residual Propagation

In this section, we develop a general framework for subsample-based estimation under dynamic contamination. The framework applies to a broad class of residual-based subsample estimators and operates at the level of the propagation mechanism. The analysis is conducted within a causal and invertible VARMA representation, used as a convenient operator framework in which residual propagation can be characterised explicitly.

As shown in the previous section, contamination propagates through the residual filter and induces non-local distortions in the residual-based criterion. This implies a structural requirement for subsample-based estimation in dynamic models. The retained index set must be compatible with the propagation mechanism, since excluding only directly contaminated observations does not eliminate the resulting residual distortions.

We formalise this requirement through a criterion-based framework defined on retained subsets of indices. Within this framework, we introduce a transformation of index sets that removes the residual footprint of contamination and study the conditions under which it preserves the asymptotic behaviour of the estimator under clean data while restoring consistency under contamination.

Residual-based criterion functions

We consider subsample-based estimators defined through residual-based sample criteria. Let $H_T \subset \{1,\dots,T\}$ be a subset of indices with cardinality $h$, and let $\mathcal{H}_T^h$ denote the collection of all such subsets. For any $H_T \in \mathcal{H}_T^h$, define

equation[equation omitted — 106 chars of source]

and let $\hat{\varphi}_T(H_T)$ denote a minimiser or maximiser of $f_T(\varphi,H_T)$. The criterion depends on $\varphi$ only through the residuals $\{e_t(\varphi):t\in H_T\}$.

For simplicity, $\hat{\varphi}_T(H_T)$ is taken to denote an exact optimiser of the sample criterion. The same framework can be extended to approximate optimisers that achieve the optimum up to an $o_p(1)$ error, as in the classical theory of extremum estimators; see, for example, vdV1998. The subset $H_T$ may be fixed, data-dependent, or generated by an algorithm. The present framework is formulated conditional on a given retained subset $H_T$ and does not depend on how that subset is obtained. In particular, the construction of $H_T$ is treated as external to the subsequent estimation problem considered here. The representation is formulated in terms of the structural residuals $e_t(\varphi)$ associated with the latent process, and thus represents the uncontaminated criterion.

The criterion (ref) encompasses a broad class of estimators used in regression and robust statistics, including least squares Amemiya1985, least absolute deviations and quantile regression KoenkerBassett1978, maximum likelihood and general M-estimators Huber1964, NeweyMcFadden1994, as well as trimming-based procedures such as LTS Ro84; see also Maronna2006. Related residual-based subset selection methods studied by JohansenNielsen2009,JohansenNielsen2013,JohansenNielsen2016 also fall within this framework.

The analysis is conducted directly at the level of sample criteria and does not rely on the existence of a population objective function or on uniform convergence. The key idea is that if the criteria evaluated on two subsets are uniformly close, then, under local stability conditions, the corresponding optimisers are also close. This allows us to study subsample transformations without invoking a limiting objective. In contrast to standard extremum arguments, which rely on convergence to a deterministic limit with a unique optimiser Cizek2008, the present approach operates entirely at the level of the sample criterion.

Under contamination, Proposition (ref) shows that the residual-based criterion constructed from contaminated observations differs from its uncontaminated counterpart through the propagated residual component. Unless this propagation effect becomes asymptotically negligible, the contaminated and uncontaminated criteria are not asymptotically equivalent. Consequently, the corresponding optimisation problems share the same optimiser only under specific conditions for which the propagated residual distortion becomes asymptotically negligible. Thus, even when the retained subset is asymptotically free of directly contaminated observations, standard subsample-based estimation does not generically yield consistency for the clean-data parameter under dynamic contamination.

A propagation-compatible transformation of index sets

The structural requirement above leads to the following propagation-compatible transformation of retained index sets.

definition[Patch removal operator] For a given integer $\kappa \geq 0$, define the subset transformation \begin{equation} S^\kappa:\, H_T \mapsto H_T \setminus \bigcup_{t \in H_T^c} \{t+1, \dots, \min(t+\kappa, T)\}. \end{equation} This transformation removes, for each excluded index, a forward block of length at most $\kappa$, thereby accounting for the propagation of contamination in the residual sequence. The transformation $S^\kappa$ induces an operator acting on subsample-based estimators: \begin{equation} K^\kappa:\, \hat{\varphi}_T\, \mapsto\, \hat{\varphi}_T \circ S^\kappa, \end{equation} that is, \begin{equation} (K^\kappa \hat{\varphi}_T)(H_T) = \hat{\varphi}_T(S^\kappa H_T). \end{equation} We refer to $K^\kappa$ as the patch removal operator.

The transformation $S^\kappa$ is not an additional trimming rule applied for robustness in the usual pointwise sense. It is the index-set operation induced by the residual propagation mechanism. Once an index is excluded because it may carry contamination, the residual filter implies that later residuals may also be affected. A propagation-compatible retained set must therefore exclude the corresponding residual footprint. The operator $K^\kappa$ implements this requirement for any subsample-based estimator by composing the original estimator with the transformation $S^\kappa$.

Thus, patch removal operates through the empirical criterion induced by the transformed index set. In the presence of contamination, it prevents residual patches from entering the criterion. Under clean data, it only changes the retained index set and may reduce the effective sample size. The invariance result below shows that this cost is asymptotically negligible under suitable stability conditions.

Let $\alpha \in (0,1)$ denote the trimming proportion associated with $H_T$, and let $\kappa$ control the extent of propagation removal. Then the transformed subset $S^\kappa H_T$ has cardinality at least $(1 - \alpha(\kappa + 1))T$. In the worst case, when residual patches do not overlap, at most $\alpha(\kappa+1)T$ observations are removed. In practice, residual patches may overlap, particularly in the presence of consecutive outliers, so that the effective number of removed observations can be smaller.

Local stability and asymptotic invariance

In this subsection, we analyse the optimisation problem defined by the sample criterion $f_T(\varphi, H_T)$ in terms of the structural residuals $e_t(\varphi)$, corresponding to the uncontaminated formulation. The analysis is conducted directly at the level of sample criteria and does not rely on convergence to a limiting objective function or differentiability.

The first set of conditions imposes local regularity of the sample criterion around its optimiser and captures properties that arise asymptotically for subsample-based estimators. In particular, it ensures separation of the optimiser and local sharpness of the criterion in its neighbourhood.

assumption[Local optimisation stability] There exists an open set $\mathcal N \subset \mathcal V$ such that, as $T \to \infty$, with probability tending to one, the following hold. (i) (Local separation) There exists $\delta>0$ such that, in the case of minimisation, \begin{equation} \inf_{\varphi\notin\mathcal N} f_T(\varphi,H_T) \ge \inf_{\varphi\in\mathcal N} f_T(\varphi,H_T) + \delta, \end{equation} or, in the case of maximisation, \begin{equation} \sup_{\varphi\notin\mathcal N} f_T(\varphi,H_T) \le \sup_{\varphi\in\mathcal N} f_T(\varphi,H_T) - \delta. \end{equation} (ii) (Local sharpness) For every $\varepsilon > 0$, there exists $\eta_\varepsilon > 0$ such that, in the case of minimisation, \begin{equation} \inf_{\substack{\varphi \in \mathcal N:\\ \|\varphi-\hat{\varphi}_T(H_T)\|\ge \varepsilon}} f_T(\varphi,H_T) \ge f_T(\hat{\varphi}_T(H_T),H_T) + \eta_\varepsilon, \end{equation} or, in the case of maximisation, \begin{equation} \sup_{\substack{\varphi \in \mathcal N:\\ \|\varphi-\hat{\varphi}_T(H_T)\|\ge \varepsilon}} f_T(\varphi,H_T) \le f_T(\hat{\varphi}_T(H_T),H_T) - \eta_\varepsilon. \end{equation}

Assumption (ref) formalises local identification and curvature properties of the sample criterion. It is analogous to standard conditions used in extremum estimation, and can be verified under conventional sufficient conditions, such as uniform convergence and uniqueness of a limiting objective. However, the present framework does not rely on the existence of such a limiting objective and instead operates directly at the level of the sample criterion.

assumption[Stability under patch removal] The criterion function is uniformly stable under patch removal, that is, \begin{equation} \sup_{\varphi\in\mathcal V} \bigl| f_T(\varphi,H_T)-f_T(\varphi,S^\kappa H_T) \bigr| \overset{p}{\longrightarrow}0. \end{equation}

Assumption (ref) depends on how the criterion aggregates residuals. Proposition (ref) provides sufficient conditions illustrating that the high-level condition holds for a wide class of residual-based criteria, in particular when the fraction of removed observations is asymptotically negligible and the criterion contributions are suitably controlled.

Together, Assumptions (ref) and (ref) ensure that patch removal does not alter the optimisation problem locally. Local separation restricts attention to a neighbourhood of the optimiser, stability controls the perturbation induced by the transformation, and sharpness ensures that such perturbations do not shift the optimiser.

The following theorem establishes that, under the stated assumptions, local optimisation stability is preserved under patch removal. In particular, the criterion based on $S^\kappa H_T$ satisfies the same local separation and sharpness properties as the original criterion. Furthermore, the patch removal operator is asymptotically invariant, in the sense that the resulting estimator is asymptotically equivalent to the original one.

theorem[Stability and invariance under patch removal] Suppose that the data are generated from the uncontaminated model (ref), and that Assumptions (ref) and (ref) hold. Then, with probability tending to one, the criterion $f_T(\varphi,S^\kappa H_T)$ satisfies the following on the same open set $\mathcal N$. (i) (Transferred local separation) There exists $\delta'>0$ such that, in the case of minimisation, \begin{equation} \inf_{\varphi\notin\mathcal N} f_T(\varphi,S^\kappa H_T) \ge \inf_{\varphi\in\mathcal N} f_T(\varphi,S^\kappa H_T)+\delta', \end{equation} or, in the case of maximisation, \begin{equation} \sup_{\varphi\notin\mathcal N} f_T(\varphi,S^\kappa H_T) \le \sup_{\varphi\in\mathcal N} f_T(\varphi,S^\kappa H_T)-\delta'. \end{equation} (ii) (Transferred local sharpness) For every $\varepsilon>0$, there exists $\eta'_\varepsilon>0$ such that, in the case of minimisation, \begin{equation} \inf_{\substack{\varphi\in\mathcal N:\\ \|\varphi-\hat{\varphi}_T(S^\kappa H_T)\|\ge\varepsilon}} f_T(\varphi,S^\kappa H_T) \ge f_T(\hat{\varphi}_T(S^\kappa H_T),S^\kappa H_T)+\eta'_\varepsilon, \end{equation} or, in the case of maximisation, \begin{equation} \sup_{\substack{\varphi\in\mathcal N:\\ \|\varphi-\hat{\varphi}_T(S^\kappa H_T)\|\ge\varepsilon}} f_T(\varphi,S^\kappa H_T) \le f_T(\hat{\varphi}_T(S^\kappa H_T),S^\kappa H_T)-\eta'_\varepsilon. \end{equation} Furthermore, we have \begin{equation} (K^\kappa \hat{\varphi}_T)(H_T) - \hat{\varphi}_T(H_T) \overset{p}{\longrightarrow}0. \end{equation}

An immediate consequence is that consistency is preserved under patch removal.

corollary[Consistency under patch removal] Under the conditions of Theorem (ref), if $\hat{\varphi}_T(H_T)\overset{p}{\longrightarrow}\varphi_0$, then \begin{equation} (K^\kappa \hat{\varphi}_T)(H_T)\overset{p}{\longrightarrow}\varphi_0. \end{equation}

Contamination and propagation control

We now turn to the contaminated setting and study how contamination affects residual-based criteria. We begin with assumptions on the contamination process and then introduce additional regularity conditions on the criterion function.

The purpose of this subsection is not to treat contamination as a structural feature of the data-generating mechanism, but to isolate the conditions under which its effect on the residual-based criterion can be controlled. Structural features of the clean stochastic mechanism should not be confused with contamination. In particular, even when the clean process has no unit root, contamination may generate spurious persistence or unit-root-like behaviour in the observed data. Such behaviour does not reflect genuine propagation through the clean model, but rather distortion induced by contamination in the residual sequence.

For this reason, the relevant issue here is not stationarity in a purely descriptive sense, but whether the propagation of contamination through the residual filter can be sufficiently attenuated by patch removal. In the present VARMA framework, this is achieved through conditions that yield decay of the propagation effect. The key requirement is therefore not whiteness of the errors per se, but the existence of a representation under which contamination enters through a propagation mechanism whose effect can be controlled. The analysis relies on structural and regularity conditions that ensure that the residual representation is well defined and that this propagation mechanism is suitably attenuated.

assumption[Contamination rate] There exists $\bar{\alpha} \in [0,1/2)$ such that \begin{equation} \limsup_{T\to\infty} T^{-1}\sum_{t=1}^T \delta_t \le \bar{\alpha}. \end{equation}

To ensure that the trimming step is capable of removing contaminated observations asymptotically, the trimming proportion $\alpha$ is assumed to satisfy $\alpha \ge \bar{\alpha}$. The restriction $\bar{\alpha}<1/2$ ensures that contaminated observations form a minority of the sample and is standard in robust statistics; see, for example, Ro84.

To isolate the effect of residual propagation from the separate problem of outlier detection, we impose the following high-level condition. In many applications, extreme observations are readily identifiable in the observed series. Contamination propagates through the dynamic filter and distorts the residual-based criterion.

assumption[Asymptotic absence of contamination] \begin{equation} \Pr\big( \delta_t = 0 for all t \in H_T \big) \longrightarrow 1, \end{equation} as $T \to \infty$.

A sufficient condition, which arises in trimming-based procedures, is that, for sufficiently large $T$, $\min_{t \in \boldsymbol{\tau}} |\tilde{e}_t(\varphi_0)| > \max_{t \notin \boldsymbol{\tau}} |\tilde{e}_t(\varphi_0)|$, where $\boldsymbol{\tau} = \{t : \delta_t = 1\}$ denotes the set of contaminated indices. This enforces strict separation between contaminated and uncontaminated residuals and implies exact classification.

Assumption (ref) is substantially weaker. It does not require separation of residual magnitudes or exact identification of individual observations. Instead, it imposes a one-sided requirement that contaminated observations are asymptotically excluded from the selected subset $H_T$, while uncontaminated observations may be discarded. Thus false positives are allowed, but false negatives are ruled out with probability tending to one. This reflects a conservative selection principle, under which it is preferable to discard additional clean observations in order to eliminate the influence of contamination.

The present paper treats the retained subset $H_T$ as given. Assumption (ref) is imposed as a high-level condition so that the analysis can focus on contamination and residual propagation, rather than on the separate problem of identifying contaminated observations. Importantly, this condition represents a best-case benchmark. Even under this asymptotically uncontaminated subset, standard subsample-based estimation fails in dynamic models, as shown below.

The choice of $\kappa$ is central to the effectiveness of patch removal. A common trimming length $\kappa$ provides a uniform upper bound on the propagation of contamination across observations. Although the effective propagation length may vary across outliers, the analysis requires only such a bound and does not depend on the exact propagation profile of individual contamination events.

For a causal and invertible VARMA model (ref), the coefficients in the representation $\pi(L)$ decay exponentially fast. That is, there exist $M>0$ and $R\in(0,1)$ such that $|\pi_j|\le M R^j$ for all $j$; see brockwell_time_1991. This implies geometric decay of the contamination effect in the residual sequence. Consequently, by choosing $\kappa$ sufficiently large, the residual distortion becomes asymptotically negligible on the retained subset.

If the propagation effect does not decay, as in unit-root-type settings, finite patch removal cannot eliminate the residual footprint of contamination.

assumption[Contamination decay condition] The contamination magnitude, rate, and trimming level satisfy \begin{equation} \alpha T \,\|\delta_T \zeta_T\|_\infty \, R^\kappa \longrightarrow 0 \end{equation} as $T \to \infty$.

Assumption (ref) can be satisfied by an appropriate choice of $\kappa$. Since $R\in(0,1)$, the factor $R^\kappa$ decays exponentially in $\kappa$, whereas $\alpha T\|\delta_T\zeta_T\|_\infty$ may grow with $T$. Under standard breakdown-type formulations, $\|\delta_T \zeta_T\|_\infty$ is not assumed to be bounded. It therefore suffices for $\kappa$ to increase at a logarithmic rate in $\alpha T\|\delta_T\zeta_T\|_\infty$, so that the exponential decay offsets the combined effect of contamination frequency and magnitude. Assumption (ref) thus ensures that the residual effect of contamination becomes asymptotically negligible after patch removal.

At the same time, $\kappa$ is constrained by the feasibility condition

equation[equation omitted — 39 chars of source]

for some constant $c\in(0,1)$, which ensures that the retained subset has cardinality of order $T$. If $\kappa$ increases with $T$, the trimming proportion $\alpha$ may need to decrease accordingly, corresponding to settings with a small, possibly vanishing, contamination proportion.

In the VAR($p$) case, propagation has finite memory and is fully eliminated by choosing $\kappa=p$. Consequently, once the initial subset is asymptotically free of contamination, no further distortion remains in the residual sequence on the retained subset, and Assumption (ref) is no longer required. This simplification is formalised in Theorem (ref) below.

The assumptions introduced so far concern the contamination mechanism and its interaction with the retained subset. We now impose additional regularity conditions on the criterion function itself. These conditions are conceptually distinct from the contamination assumptions above. They are imposed on the clean residual-based representation and are applied after the effect of contamination has been attenuated on the retained subset. Their role is to control the effect of residual perturbations on the criterion. They are not required for Theorem (ref) and Corollary (ref), which are established under the uncontaminated setting.

assumption[Criterion regularity] For every subset $H_T \in \mathcal H_T^h$, the following conditions hold: (a) The criterion admits the representation \begin{equation} f_T(\varphi,H_T) = \frac{1}{|H_T|} \sum_{t\in H_T} m_t\bigl(e_t(\varphi)\bigr). \end{equation} (b) The functions $m_t:\mathbb R \to \mathbb R$ satisfy \begin{equation} |m_t(y)-m_t(x)| \le C\bigl(1+|x|+|y-x|\bigr)|y-x|, \end{equation} for all $x,y\in\mathbb R$, uniformly in $t$, for some constant $C<\infty$. (c) \begin{equation} \sup_{\varphi\in\mathcal V} \frac{1}{|S^\kappa H_T|} \sum_{t\in S^\kappa H_T}|e_t(\varphi)|^2 = O_p(1). \end{equation}

Assumption (ref) imposes a regularity condition on the criterion function conditional on a given retained subset $H_T$. Part (a) restricts attention to criteria that can be written as averages of observation-wise contributions over $H_T$. This is standard in subsample-based estimation and is a special case of (ref), since the criterion depends on $\varphi$ only through the residuals $\{e_t(\varphi): t \in H_T\}$. Part (b) imposes a uniform increment condition on the functions $m_t$, requiring that differences in the criterion contributions are controlled linearly in the size of the perturbation, with a coefficient growing at most linearly in the magnitude of the argument. This accommodates a broad class of smooth and non-smooth loss functions, including general convex and non-convex M-estimation criteria. Part (c) requires that, after patch removal, the average squared residuals over $S^\kappa H_T$ are stochastically bounded uniformly over $\varphi \in \mathcal V$. This does not impose bounded residuals, nor does it require the same property to hold over the full sample or the original subset $H_T$. It allows for occasional large values and is compatible with heteroskedasticity.

Taken together, these conditions are weak enough to cover most commonly used criterion functions in time series and econometrics, while ensuring that perturbations in the residuals induce controlled perturbations of the criterion. The assumption operates directly at the level of criterion increments, rather than imposing unnecessary smoothness conditions.

Let $\tilde{\varphi}_T(H_T)$ denote an estimator obtained by optimising a sample criterion function of the form

equation[equation omitted — 123 chars of source]

where, as in (ref), the criterion depends on $\varphi$ only through the contaminated residuals $\{\tilde e_t(\varphi):t\in H_T\}$.

theorem[Invariance under contamination] Suppose that Assumptions (ref)--(ref) hold. Then, under AO or IO contamination in model (ref), \begin{equation} (K^\kappa \tilde{\varphi}_T)(H_T) - (K^\kappa \hat{\varphi}_T)(H_T)\overset{p}{\longrightarrow}0. \end{equation} Moreover, in the VAR($p$) case with $\kappa = p < \infty$, the result holds without Assumption (ref).

An immediate consequence is that consistency is preserved under contamination.

corollary[Consistency under contamination] Under the conditions of Theorem (ref), if $\hat{\varphi}_T(H_T)\overset{p}{\longrightarrow}\varphi_0$, then \begin{equation} (K^\kappa \tilde{\varphi}_T)(H_T)\overset{p}{\longrightarrow}\varphi_0. \end{equation}

The preceding results do not require $\kappa$ to be fixed and remain valid for sequences $\kappa=\kappa_T$, provided that the assumptions continue to hold. This yields a useful separation in the analysis. The optimisation results are independent of the contamination mechanism, while the roles of $\alpha$ and $\kappa$ arise only through model-specific conditions, such as Assumption (ref), that ensure stability.

The analysis is conducted within a weakly stationary VARMA framework, which provides a convenient operator representation for studying residual propagation. Under weak stationarity, Wold's decomposition theorem wold1938study ensures that the process admits an infinite-order VMA representation. Thus, the propagation mechanism is not specific to finite-order VARMA models, but arises more generally in dynamic systems admitting linear or approximately linear filtering representations. Under standard conditions, such representations can be approximated by causal and invertible VARMA models, making the propagation structure explicit through the associated filters. For this reason, the VARMA framework is sufficiently general to capture a broad class of weakly stationary linear dynamic systems relevant for the present analysis.

Stationarity itself, however, is not essential. What matters is whether the propagation effect induced by contamination decays sufficiently fast on the retained subset. In the present framework, this is ensured through Assumption (ref). When propagation does not decay, finite patch removal is insufficient.

Similar propagation effects also arise in many weakly stationary nonlinear models, including heteroskedastic and regime-switching processes. In such settings, the clean-data dynamics themselves may be nonlinear, while contamination remains an external perturbation acting through the residual filter.

The framework admits the boundary case $\kappa=0$, in which $S^\kappa H_T = H_T$, so that patch removal reduces to the identity operator. In this case Assumption (ref) becomes

equation[equation omitted — 90 chars of source]

This condition is highly restrictive and excludes settings in which residual propagation induces a non-negligible distortion of the criterion. Thus the case $\kappa=0$ does not contradict the generic inconsistency of subsample-based estimation under dynamic contamination. Rather, it corresponds to an exceptional boundary case in which propagation-compatible patch removal is asymptotically unnecessary.

Verification of the assumptions

In this section, we clarify the roles of the assumptions and indicate which of them require verification in applications. Notably, the analysis does not require structural assumptions on the contamination process, such as independence, distributional forms, or temporal dependence. Instead, contamination enters only through its magnitude, its rate, and its propagation through the residual filter.

The propositions below focus primarily on Assumptions (ref) and (ref), which are high-level conditions imposed directly on the optimisation problem and whose generality may not be immediate from their formulation. By contrast, Assumptions (ref)--(ref) depend on the contamination and screening mechanism, while Assumption (ref) is a standard regularity condition satisfied by a broad class of residual-based criteria. The propositions provide sufficient conditions under which Assumptions (ref) and (ref) hold, thereby connecting the present framework to familiar extremum estimation arguments and existing subsample-based robust procedures. The conditions imposed throughout are sufficient rather than necessary and are formulated at a high level to capture the structural features of the problem across a broad class of settings, rather than to characterise minimal assumptions in each component.

We begin by connecting the local optimisation stability condition to the classical extremum estimator framework. The result is stated for minimisation, and the maximisation case is analogous.

propositionConsider the uncontaminated VARMA model (ref). Fix a subset $H_T \in \mathcal H_T^h$, and let \begin{equation} \hat{\varphi}_T(H_T) = \arg\min_{\varphi\in\mathcal V} f_T(\varphi,H_T), \end{equation} where $\mathcal V \subset \mathbb R^m$ is compact and $\varphi_0 \in \mathcal V$. Assume that (i) the mapping $\varphi \mapsto f_T(\varphi,H_T)$ is continuous; (ii) there exists a deterministic function $f:\mathcal V\to\mathbb R$ such that \begin{equation} \sup_{\varphi\in\mathcal V} \bigl|f_T(\varphi,H_T)-f(\varphi)\bigr| \overset{p}{\longrightarrow}0; \end{equation} (iii) the function $f$ is continuous and has a unique minimiser at $\varphi_0$. Then (a) $\hat{\varphi}_T(H_T)\overset{p}{\longrightarrow}\varphi_0$; (b) Assumption (ref) holds.

Proposition (ref) shows that Assumption (ref) is compatible with the standard extremum-estimator framework. In particular, whenever the sample criterion converges uniformly to a deterministic limit with a uniquely identified minimiser, local separation and local sharpness follow asymptotically. Thus, the present framework includes the classical consistency setting as a special case, even though the main results of this paper are formulated directly at the level of sample criteria and do not rely on the existence of a limiting objective function.

propositionSuppose that Assumption (ref)(a) holds for $H_T$ and the corresponding transformed subsets $S^\kappa H_T$ for all sufficiently large $T$. Assume that (i) the number of removed indices is asymptotically negligible, that is, \begin{equation} \frac{|H_T\setminus S^\kappa H_T|}{|H_T|} \overset{p}{\longrightarrow}0; \end{equation} (ii) the criterion contributions are uniformly bounded in probability, \begin{equation} \sup_{\varphi\in\mathcal V}\sup_{t\in H_T} \bigl|m_t(e_t(\varphi))\bigr| = O_p(1). \end{equation} Then Assumption (ref) holds.

Proposition (ref) provides a transparent route to verify Assumption (ref) for criteria constructed as averages of residual contributions. The key requirement is that patch removal affects only an asymptotically negligible fraction of the retained indices, so that the induced perturbation of the empirical criterion vanishes uniformly in the parameter. Formally, this corresponds to condition (i).

To interpret this condition, note that $S^\kappa$ removes, for each initially excluded observation, a patch of length at most $\kappa$. Hence, the total number of additionally removed observations is at most of order $\alpha \kappa T$, while $|H_T| \approx (1-\alpha)T$, yielding the bound

equation[equation omitted — 98 chars of source]

A sufficient condition is therefore that $\alpha \kappa \to 0$. Combined with the contamination decay requirement in Assumption (ref), which typically requires $\kappa$ to grow at a logarithmic rate, this leads to conditions of the form

equation[equation omitted — 72 chars of source]

and, in typical settings, to $\alpha \log T \to 0$. This corresponds to a setting in which contamination is asymptotically sparse, even though the total number of contaminated observations may diverge.

Condition (i) is not the only possible route to stability. An alternative is to impose conditions under which the criteria based on $H_T$ and $S^\kappa H_T$ have asymptotically equivalent behaviour under clean data. However, such an approach typically requires stronger assumptions on the data-generating process and dependence structure, since the thinning induced by $S^\kappa$ must preserve the limiting behaviour of the criterion. The advantage of condition (i) is that it achieves stability by directly controlling the fraction of affected observations, making it particularly suitable in settings with general dependence structures.

For pure VAR$(p)$ models, the effective patch length is finite, so that $\kappa$ can be taken as a fixed constant (e.g.\ $\kappa=p$), rather than growing with $T$. In this case, the condition $\alpha \kappa \to 0$ reduces to $\alpha \to 0$, and the compatibility with Assumption (ref) becomes immediate.

Condition (ii) is stated as a convenient sufficient condition. For unbounded criteria, such as quadratic or Gaussian likelihood contributions, the same conclusion can be obtained under suitable moment and maximal control conditions on the residual sequence.

The preceding propositions show that Assumptions (ref) and (ref) can be verified under familiar conditions. This clarifies the scope of the present framework and its connection to existing subsample-based robust estimators. In particular, the framework is compatible with the robust procedures studied by JohansenNielsen2009,JohansenNielsen2013,JohansenNielsen2016, insofar as these procedures produce a retained subset $H_T$ on which subsequent estimation is based. Procedures such as Huber-skip and LTS differ in how the subset $H_T$ is constructed, but this distinction is immaterial here, as the framework is formulated conditional on $H_T$ and does not depend on how it is obtained. Different procedures may employ one-step rules, such as Huber-skip and LTS, or iterative and sequential methods such as the forward search HadiSimonoff1993, AtkinsonRiani2000. Once a subset $H_T$ is obtained, the patch removal operator can be applied to construct $S^\kappa H_T$, after which estimation proceeds on the transformed subset.

Under the regularity conditions imposed in that literature, including tightness of initial estimators and conditions on the innovation density and regressors, the sample criteria underlying these estimators satisfy the sufficient conditions in Propositions (ref) and (ref). These results therefore provide concrete examples under which Assumptions (ref) and (ref) hold jointly.

The high-level assumptions are thus weaker than the sufficient conditions in Propositions (ref) and (ref), which are themselves mild and hold under familiar regularity conditions. Together with Theorem (ref), this implies that, once residual propagation is properly controlled, the proposed transformation ensures consistent estimation for a broad class of procedures based on subset selection in dynamic models under sufficiently sparse contamination.

Simulation

This section studies the finite-sample implications of the propagation-compatible correction under controlled contamination. The design isolates the estimation problem from the subset selection problem. In standard robust procedures, methods such as Huber-type estimators, LTS, or related approaches produce a retained subset of observations to be used for subsequent estimation. While these procedures differ in how the subset $H_T$ is constructed, the present analysis treats that subset as given and studies the effect of residual propagation and patch removal conditional on it.

In the present design, we therefore work directly with a retained subset $H_T$ containing only uncontaminated observations. This isolates residual propagation and patch removal from the separate problem of outlier detection and represents a best-case benchmark for subsample-based procedures.

In each Monte Carlo replication, the model parameters are re-estimated from the simulated sample using a residual-based subsample estimator defined on the retained set. Since all such procedures share the same second-stage estimation on $H_T$, the specific choice of estimator is not central to the purpose of the experiment. Estimation is carried out by numerically minimising the determinant of the residual covariance matrix over the selected observations after patch removal. The optimisation is initialised at the true parameter value to isolate the effect of contamination and propagation correction from numerical issues. The model order is assumed to be known and correctly specified.

We consider three standard models VAR(1), VMA(1), and VARMA(1,1). Let $\varepsilon_t \sim N_2(0, \Sigma)$ and

equation[equation omitted — 136 chars of source]

The data-generating processes are

equation[equation omitted — 255 chars of source]

Contamination is introduced as AO or IO. A proportion $\alpha \in \{1\%, 5\%, 10\%\}$ of indices is contaminated with magnitude $\zeta \in \{5, 10, 50, 100\}$. Sample sizes are $T \in \{500, 1000\}$.

The trimming parameter $\kappa$ controls the length of residual patches removed after each contaminated observation. The case $\kappa=0$ corresponds to oracle subsampling without accounting for residual propagation, and serves as a benchmark for the failure of standard subsample-based procedures.

Estimator performance is evaluated using total bias and root mean squared error (RMSE)

equation[equation omitted — 274 chars of source]

The simulation results are reported in Tables (ref)--(ref). Across all experiments, a clear failure-correction contrast emerges. When $\kappa=0$, substantial bias arises under contamination despite oracle subsampling, confirming that subsample-based estimation fails even when the retained subset contains only uncontaminated observations. This shows that the failure is not due to imperfect identification of contaminated observations, but arises from residual propagation in the estimation stage itself. The distortion increases sharply with the outlier magnitude $\zeta$, reflecting amplification through the residual-based criterion. In contrast, increasing $\kappa$ restores estimation accuracy, while inducing only a mild efficiency loss under clean data.

Table (ref) reports the VAR(1) results. Under AO contamination, estimation deteriorates sharply at $\kappa=0$, reflecting finite residual propagation that is not removed by standard trimming. Setting $\kappa=1$ restores performance to the clean-data benchmark, consistent with the finite propagation length. IO contamination has no effect in this setting and is omitted, in line with Theorem 2.

Table (ref) reports the VMA(1) results. Under contamination, severe distortions arise at $\kappa=0$, reflecting the accumulation of propagated residual effects. As $\kappa$ increases, performance improves markedly, with particularly strong gains for large $\zeta$. AO and IO yield identical results, as both generate non-local residual effects.

Tables (ref) and (ref) report the VARMA(1,1) results. Under contamination, both AO and IO induce substantial distortions at $\kappa=0$, with AO having a larger initial effect. Increasing $\kappa$ reduces the distortions markedly and brings both cases close to the clean benchmark. The difference between AO and IO diminishes for sufficiently large $\kappa$, reflecting the joint propagation structure.

Taken together, the simulation results confirm the theoretical findings and clarify their implications. First, even under oracle subsampling, substantial bias arises when residual propagation is not controlled, demonstrating that subsample-based estimation fails in dynamic models. Second, the distortion increases with the magnitude of contamination, reflecting its amplification through the residual-based criterion. Third, patch removal restores estimation accuracy across all settings, while incurring only a modest efficiency loss under clean data. These findings highlight that reliable estimation requires controlling residual propagation, rather than relying solely on pointwise trimming.

A Feasible Implementation of the Patch Removal Operator

This section describes a feasible implementation of the patch removal operator for empirical use and illustrates how it can be incorporated into existing robust procedures. Subsample-based robust procedures for contaminated time series produce a retained subset for estimation; see, for example, JohansenNielsen2016. Suppose that such a procedure produces an estimate and a retained subset $H_T$. Patch removal then refines $H_T$ to account for residual propagation.

Given an initial subset $H_T$, define the patch-removed subset $S^\kappa H_T$ according to (ref), where $\kappa \ge 0$ is a user-specified patch length. The estimator is then re-applied to the reduced subset, yielding

equation[equation omitted — 113 chars of source]

This step implements the operator $K^\kappa$ and removes not only directly contaminated observations but also the subsequent residual patches induced by dynamic propagation. The screening step serves only to provide an initial conservative subset.

In finite samples, the initial screening step may misclassify observations, particularly when it relies on level-based diagnostics. For this reason, the procedure may be iterated. Let $\tilde{\varphi}_T^{(\kappa,0)}=\tilde{\varphi}_T(H_T)$ and $H_T^{(0)}=H_T$. The recursive scheme is given in Algorithm (ref). Under the conditions of Theorem (ref), a single application of $K^\kappa$ is sufficient for asymptotic validity, so iteration serves only as a finite-sample refinement. When the initial screening step is accurate, the algorithm typically stabilises quickly. When classification is imperfect, the refinement step may improve the approximation to the uncontaminated subset.

algorithm[algorithm omitted — 968 chars of source]

The choice of $\kappa$ is governed by the propagation structure of the underlying dynamic model. In particular, the effective propagation length depends on the AR and moving-average orders $(p,q)$ and on the contamination mechanism. This connection is formalised in Assumption (ref). For pure VAR($p$) models, residual patches have finite length, so that $\kappa$ can be taken equal to $p$. For models with a moving-average component, residual patches are of infinite length but decay geometrically, so that $\kappa$ must typically increase with the sample size at a logarithmic rate. Thus, the choice of $\kappa$ is linked to the underlying VARMA specification.

In practice, however, neither the model orders $(p,q)$ nor the contamination mechanism are known in advance, and both may be difficult to infer reliably in contaminated data. This creates a circular problem in that the choice of $\kappa$ depends on the model specification, while the model specification itself is obscured by contamination. We therefore adopt a pragmatic joint-selection approach. A collection of candidate specifications indexed by $(\kappa,p,q)$ is considered. For each candidate, $\kappa$ is chosen large enough to remove the dominant part of the residual patch, but not so large that the retained subset becomes too small. The estimation procedure in Algorithm (ref) is then applied to each candidate.

Candidate specifications may be screened using the observed propagation patterns in the contaminated series. Although uncontaminated residuals are not observable, the behaviour of observed aberrant level and residual patches still provides useful diagnostic information. In particular, Table (ref) shows that AO contamination does not generate persistent level patches, but may induce persistent residual patches, whereas IO contamination generates persistent level patches together with model-dependent residual patch behaviour. These qualitative features can be used to rule out implausible combinations of contamination type, patch length, and dynamic specification.

For estimation itself, explicit classification of contamination as AO or IO is not essential. The role of $\kappa$ is to remove the dominant part of residual propagation, and a sufficiently conservative choice typically yields robustness regardless of the underlying mechanism. The AO--IO distinction is more relevant for interpretation and prediction, since different contamination mechanisms imply different propagation patterns and hence potentially different out-of-sample implications. For this reason, the distinction is used here mainly as a diagnostic device and is revisited in the empirical illustration.

After this preliminary reduction of the candidate set, we require a simple rule to compare the remaining models. The robust model selection literature contains substantially more elaborate proposals. For example, Müller01122005 combine a robust penalised in-sample criterion with a robust estimate of conditional prediction loss obtained via a stratified bootstrap, and note that the associated computational burden can be substantial. Model selection, however, is not the primary objective of the present paper. Our purpose is only to adopt a simple and operational comparison rule for illustration.

We therefore employ a normalised information criterion based on the average log-likelihood, that is, AIC on a per-observation scale:

equation[equation omitted — 130 chars of source]

where $|S^\kappa H_T|$ is the effective sample size after subsampling and patch removal, and $k$ is the number of estimated parameters. This normalisation reduces the scale effect induced by differing effective sample sizes and facilitates heuristic comparison across candidates. The criterion is used here solely as a practical device rather than as a fully developed contribution to robust model selection.

Empirical Illustration

We consider two daily Icelandic river flow series covering the period from 1 January 1972 to 31 December 1974, yielding a total of $T=1096$ observations, in which the flow is measured in cubic metres per second. The data originate from the Hydrological Survey of the National Energy Authority of Iceland and were first analysed by ToThGu85. The dataset has subsequently been used in the time series literature, including Tsa98 and TeYa14, where bivariate river flow systems are analysed using threshold and smooth transition models, respectively. The two series correspond to the rivers Jökulsá eystra and Vatndalsá. The former is the larger river, with a drainage basin that includes a glacier, whereas the latter has a smaller basin with a substantial contribution from groundwater. This leads to smoother dynamics in the former and more pronounced variation in the latter.

In the present analysis, we focus on the observed bivariate river flow series and denote them by $y_{1t}$ and $y_{2t}$. The series exhibit strong serial correlation and a nontrivial temporal structure, making them suitable for illustrating the framework developed in this paper. Unlike earlier studies, we do not incorporate exogenous variables such as precipitation or temperature, as our aim is not to construct a structural model of river flow dynamics.

We do not impose any assumption on the presence of outliers in the data. The empirical illustration is designed to examine the behaviour of the proposed procedure under potential contamination, irrespective of whether such contamination is present in the observed series. The analysis is therefore conducted in a robustness framework rather than as a diagnostic exercise. While the dataset has often been used to study nonlinear and regime-dependent dynamics, our objective here is different in that it is used to provide a realistic temporal structure within which the impact of contamination on residual-based estimation can be examined.

Accordingly, the dataset serves as a representative example of a multivariate time series with nontrivial autocorrelation. This setting allows us to study how contamination, if present, propagates through temporal dependence and affects residual-based estimation, and to assess how the proposed procedure mitigates such effects in practice. For reference, Figure (ref) displays the series together with observations identified by the iterative procedure in Algorithm (ref).

Model comparison is carried out over a range of VARMA$(p,q)$ specifications. For each candidate model, estimation is conducted using the Huber-skip procedure with trimming proportion $\alpha=0.1$, followed by patch removal with $\kappa=5$. Model comparison is based on a Gaussian log-likelihood computed from the residuals, and we report the corresponding $\mathrm{AIC}_{\mathrm{avg}}$ (ref) in Table (ref).

table[table omitted — 808 chars of source]

We restrict attention to a single subsample-based procedure for illustration. While alternative approaches, such as LTS, could also be employed, the focus is not on comparing robust estimation algorithms, but on isolating the effect of residual propagation and patch removal on the estimation criterion. The Huber-skip procedure is therefore used as a representative method that produces a data-dependent subset and permits repeated estimation across candidate models. For each specification, we report the full sample estimator, the Huber-skip estimator without patch removal, and the corresponding estimator with patch removal. These estimators differ in how contamination is handled. The full sample estimator uses the observed series directly, Huber-skip applies pointwise trimming, and patch removal additionally accounts for residual propagation. The comparison between Huber-skip and patch removal therefore isolates the effect of residual propagation, as any remaining discrepancy after trimming reflects contamination effects that are not removed by standard robust procedures.

The trimming proportion $\alpha=0.1$ is chosen to allow for a moderate level of contamination, so that the retained subset is expected to be dominated by observations consistent with the underlying data-generating mechanism. For each $(p,q)$, we report two versions of the information criterion, one based on the retained subset after patch removal and one based on the full sample. The use of $\mathrm{AIC}_{\mathrm{avg}}$ is heuristic, serving as a per-observation measure for comparing candidate models rather than as a formal robust model selection procedure.

Several features emerge from the results. First, the patch removal procedure generally yields lower values of $\mathrm{AIC}_{\mathrm{avg}}$, reflecting an improved Gaussian fit once extreme observations and their propagation effects are removed. Second, for certain specifications, such as $(p,q)=(1,2)$ and $(2,3)$, the information criterion becomes extremely large under both approaches, indicating numerical instability in the estimated residual covariance matrix. This behaviour is not driven by subsampling or patch removal, but reflects the well-known sensitivity of VARMA models with MA components, which require nonlinear optimisation and may yield ill-conditioned estimates in finite samples.

A related feature arises for $(p,q)=(2,3)$, where $|S^\kappa H_T| = |H_T|$, indicating that identified extreme observations are concentrated at the end of the sample, so that patch removal becomes inactive. Taken together, these patterns suggest that irregular values for higher-order specifications are driven primarily by numerical difficulties associated with the MA component, rather than by the subsampling or patch removal procedure. This distinction highlights that the proposed method primarily affects the statistical properties of the estimation criterion, whereas numerical stability is governed by the underlying VARMA structure.

Finally, the model minimising $\mathrm{AIC}_{\mathrm{avg}}$ differs across the two approaches. The VARMA$(2,0)$ specification minimises $\mathrm{AIC}_{\mathrm{avg}}$ under patch removal, whereas the VARMA$(2,1)$ model achieves the lowest value under the full sample criterion. This discrepancy illustrates that departures from Gaussianity, including the presence of extreme observations, can materially affect likelihood-based model comparison.

We report all model specifications without selection or filtering to provide a transparent assessment across configurations, including cases exhibiting numerical instability. The grid is restricted to $p \geq 1$, and no improvement is observed from higher-order specifications. Overall, the results should be interpreted as illustrating how the proposed procedure modifies likelihood-based model comparison, rather than as definitive evidence for a particular model choice.

figure[figure omitted — 397 chars of source]

For clarity, the full sample estimates are expressed in terms of the observed process $y_t$, whereas the robust estimates approximate the latent uncontaminated process $x_t$.

Based on Table (ref), we adopt the VARMA$(2,0)$ specification. For this model, we set $\kappa=2$ in the iterative patch removal procedure and report estimates from the full sample, Huber-skip without patch removal, and patch removal. The three estimators differ in whether they are based directly on the observed series $y_t$ or aim to recover the latent process $x_t$.

description\begin{equation} \begin{aligned} z_{1, t} &= 1.233 z_{1, t-1} + 0.064 z_{2, t-1} - 0.258 z_{1, t-2} - 0.072 z_{2, t-2} + \varepsilon_{1, t}, \\ z_{2, t} &= 0.164 z_{1, t-1} + 0.993 z_{2, t-1} - 0.173 z_{1, t-2} - 0.064 z_{2, t-2} + \varepsilon_{2, t}, \end{aligned} \end{equation} where $z_{1, t} = y_{1, t} - 3.239$ and $z_{2, t} = y_{2, t} - 2.094$. • \begin{equation} \begin{aligned} z_{1, t} &= 1.206 z_{1, t-1} + 0.103 z_{2, t-1} - 0.222 z_{1, t-2} - 0.107 z_{2, t-2} + \varepsilon_{1, t}, \\ z_{2, t} &= -0.017 z_{1, t-1} + 1.210 z_{2, t-1} + 0.015 z_{1, t-2} - 0.258 z_{2, t-2} + \varepsilon_{2, t}, \end{aligned} \end{equation} where $z_{1, t} = x_{1, t} - 3.077$ and $z_{2, t} = x_{2, t} - 1.929$. • \begin{equation} \begin{aligned} z_{1, t} &= 1.291 z_{1, t-1} + 0.072 z_{2, t-1} - 0.310 z_{1, t-2} - 0.070 z_{2, t-2} + \varepsilon_{1, t}, \\ z_{2, t} &= -0.008 z_{1, t-1} + 1.270 z_{2, t-1} + 0.010 z_{1, t-2} - 0.312 z_{2, t-2} + \varepsilon_{2, t}, \end{aligned} \end{equation} where $z_{1, t} = x_{1, t} - 3.544$ and $z_{2, t} = x_{2, t} - 1.933$.

The estimated coefficients differ non-negligibly across the three approaches. The Huber-skip estimator removes extreme observations through pointwise trimming, but residual propagation still distorts the estimated dependence structure. By contrast, the patch removal estimator accounts for this propagation mechanism and yields estimates more consistent with the latent uncontaminated process. The resulting differences are systematic rather than incidental and are consistent with the theoretical results on criterion distortion under dynamic contamination.

The observations identified by the iterative procedure are illustrated in Figure (ref). A comparison with Figure (ref) suggests that the flagged observations are more consistent with IO-type contamination, although no formal classification is attempted. In particular, extreme observations are often followed by gradual adjustments or short-lived reversals, reflecting propagation through the AR structure.

Conclusion

This paper establishes that subsample-based estimation, as commonly used in robust time series analysis, is generically inconsistent in dynamic settings when contamination propagates through the residual filter. The failure is structural and arises because residual propagation distorts the criterion defining the estimator itself, rather than merely contaminating a small number of observations. Consequently, removing contaminated observations does not in general recover the clean-data objective.

We introduce a propagation-compatible transformation of retained index sets, formalised through the patch removal operator. The operator provides a general correction for a broad class of residual-based subsample estimators without modifying their internal structure. Under suitable conditions, it leaves the uncontaminated estimator asymptotically unchanged while restoring consistency under contamination. The analysis also clarifies an important distinction across model classes. For pure VAR models the required correction is finite, whereas moving-average components may generate propagation effects with infinite memory.

These results have direct implications for empirical work. When outliers are suspected, inference should not be based on standard estimators applied to the contaminated sample, or on subsample procedures that remove only the visibly aberrant observations. Instead, inference should be based on an estimator that remains consistent for the clean-data mechanism under contamination. This is particularly important when the aim is to assess structural features of the clean data generating mechanism or model adequacy more generally.

The paper also points to several natural extensions, which fall into four directions. First, at the level of estimation, an important extension is the estimation of the contamination magnitude $\zeta_t$. This problem is inherently conditional on the one studied here, as meaningful estimation of contamination effects requires consistent estimation of the clean-model parameters. Further extensions concern the patch removal mechanism, including the use of heterogeneous trimming lengths across contamination events or across observations. While the present analysis shows that a uniform bound on $\kappa$ suffices for consistency, relaxing this restriction may improve flexibility in practice.

Second, forecasting is a natural extension. In many applications the primary objective is to predict the clean process $x_t$. Since contamination is typically sparse and does not reflect the underlying dynamics, forecasts should be based on a consistent estimate of the clean model, so as to capture the behaviour of the system in the absence of contamination.

Third, an important direction concerns inference. Outliers and residual propagation induce a structural distortion in estimation, under which standard hypothesis testing is conducted on estimators that are not consistent under contamination. Consequently, hypothesis tests may produce spurious statistical evidence, including misleading conclusions on model parameters and on structural features of the underlying data-generating mechanism, thereby distorting the interpretation of fitted models. By restoring consistency at the estimation stage, the present framework provides a basis for developing inference procedures that remain valid under contamination, including tests that account for residual propagation.

Finally, there are extensions related to model evaluation and selection. In particular, the development of subsample-based information criteria tailored to this framework remains an open problem. The average AIC used in the present paper serves only as a practical device, and a principled criterion for contaminated dynamic models with patch removal requires separate analysis.

These directions are important but raise distinct issues, and are therefore best studied separately. The present paper provides a foundation for such developments.

There are also clear limitations. First, the analysis is developed under settings in which contamination is sufficiently sparse. This is not merely a technical restriction. When contamination is frequent or persistent, it becomes difficult to distinguish external outliers from features that belong to the underlying structural mechanism. Second, patch removal necessarily reduces the effective sample size. Consequently, finite-sample convergence cannot be faster than in uncontaminated full sample estimation, and is slower when contamination is substantial or when the required patch length is large. This cost is unavoidable if consistency for the clean parameter is to be preserved.

For these reasons, the main practical recommendation of the paper is simple. If there are good reasons to suspect outliers in a dynamic time series, empirical conclusions should be based on an estimation procedure that is consistent for the clean model under contamination. The framework developed here provides a principled way to achieve this by correcting for residual propagation, thereby restoring the validity of subsample-based estimation in dynamic models.

\ifanonymous

\else

Acknowledgements

We acknowledge financial support from the Jan Wallander and Tom Hedelius Foundation, Grant No.\ P2016-0293:1. Sandberg also acknowledges support from the same foundation, Grant No.\ P22-0264. The computations were enabled by resources provided by the Swedish National Infrastructure for Computing (SNIC) at UPPMAX, partially funded by the Swedish Research Council under grant agreement No.\ 2018-05973.

Data Availability Statement

The empirical dataset consists of daily Icelandic river flows originally analysed by ToThGu85 and widely used in the time series literature. The data are publicly available from the original source but cannot be redistributed due to licensing restrictions. All code required to reproduce the simulation and empirical results is available at \href{https://github.com/yukai-yang/Robust_Experiments}{github.com/yukai-yang/Robust_Experiments} under the MIT license. \fi