Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
107,232 characters · 16 sections · 71 citation commands
Information-Enriched Selection of Stationary and Non-Stationary Autoregressions using the Adaptive LassoWe thank Christoph Hanck for valuable comments that significantly improved the paper. We are indebted to Anders B. Kock for his supportive feedback on several parts of this paper. We also thank Yannick Hoga, Matei Demetrescu, and the Winter 2022/23 Ruhrmetrics Research Seminar participants for their helpful suggestions.
\setcounter{footnote}{0}
\setcounter{footnote}{0} Over the last two decades, shrinkage estimators have become increasingly popular in the econometric literature and relevant for economic applications. A substantial body of work considers Lasso-type estimators (Tibshirani1996) and establishes conditions for convergence to their unpenalised analogue in the correct model: such \enquote{oracle efficient} procedures select the set of relevant variables with probability converging to 1 and ensure that the coefficient estimators have the identical asymptotic distributions as the unpenalised estimators if the relevant variables were known from the outset (cf. FanLi2001).
Recent contributions apply shrinkage estimators to dependent and non-identic\-ally distributed data, which often arise in macroeconomic modelling. For example, CanerHan2014 use group bridge estimation to select the correct factor number in approximate factor models. MedeirosMendes2012 establish the oracle property for the adaptive Lasso of zou2006adaptive in sparse high-dimensional models for stationary time series. KockCallot2015 present oracle inequalities for high-dimensional stationary vector autoregression (VAR) models.
Lasso-type estimators have also been suggested for models with non-stationary variables. Leeetal2022 propose the twin adaptive Lasso for oracle-efficient estimation of predictive regressions with mixed root predictors. LiaoPhillips2015 devise a consistent adaptive shrinkage procedure that selects the cointegration rank and the lag order in cointegrated VARs. Wilms2016 discuss forecasting high-dimensional time series based on $\ell_1$-penalised maximum likelihood estimators for sparse cointegration. CanerKnight2013 show that bridge estimators achieve conservative model selection in stationary and non-stationary autoregressions. kock2016consistent proves that the (BIC-tuned) adaptive Lasso is an oracle-efficient estimator in augmented Dickey-Fuller (ADF) regressions. Similarly to the multivariate procedure in LiaoPhillips2015, adaptive Lasso may thus distinguish between stationary and non-stationary data and simultaneously select the correct sparsity pattern in the lags of the dependent variable.
The key feature of adaptive Lasso is an adaptive penalty applied to the individual model coefficients. To implement the adaptive Lasso for ADF models, kock2016consistent suggests using weights based on ordinary least squares (OLS) estimates from an auxiliary ADF regression. This approach aligns with zou2006adaptive who shows that using weights based on consistent coefficient estimates is sufficient for the oracle property. To our knowledge, no contributions in the literature address tailored weight specification for autoregressive (AR) models beyond using OLS estimates. In a high-dimensional cross-section framework, Huangetal2008 prove that using marginal regression to obtain weights that need not be consistent for the true coefficients ensures the oracle property. Their simulation results indicate an improved selection performance under liberal tuning using the AIC. However, their asymptotic analysis relies on critical assumptions like partial orthogonality and a variation of the irrepresentable condition for the Lasso, which is not generally satisfied in time series regressions.\footnote{See meinshausen2006high for the original irrepresentable condition of the Lasso.}
If information indicative of a regressor's relevance other than its OLS coefficient estimate is available, incorporating this information into the adaptive weight may yield a Lasso estimator with improved properties. In ADF models, the potentially non-stationary regressor $y_{t-1}$ is of particular interest since its coefficient determines the order of integration. The unit root test literature suggests a variety of identification principles for the latter that could be informative for the weight. Combining two identification principles, we propose an enhanced weight for $y_{t-1}$ in estimating ADF models using the adaptive Lasso. This information-enriched weight implies several properties of the Lasso solution path that are beneficial for (consistent) tuning, especially under stationarity. The resulting estimator is shown to share the oracle property of the adaptive Lasso estimator in kock2016consistent and can be consistently tuned using BIC. Moreover, we prove that the estimator resulting from the modified penalty weight enjoys favourable properties like a relaxed shrinkage when required. Simulation evidence demonstrates that our modification enhances the correct selection rate of stationary and non-stationary models in finite samples, further motivating the use of adaptive Lasso for unit root testing. Our method, therefore, improves upon the estimator in kock2016consistent, particularly in this respect.
The remainder of this article is organised as follows. In (ref), we discuss the statistical framework and model selection of ADF regressions using the adaptive Lasso. (ref) introduces the modified adaptive weight and presents several theoretical results on the estimator's asymptotic properties, emphasising the detection of a stationary stochastic component. (ref) elaborates on the merit of information enrichment to mitigate spurious selection of an irrelevant variable using the example of a zero-mean model. Simulation studies highlighting the benefit of the modified weight are presented in (ref), where we also discuss handling deterministic trends. In (ref), we apply the procedure to model selection for German energy and headline inflation rates. (ref) concludes and gives an outlook to potential further applications of the information enrichment principle. (ref) presents additional simulation results. (ref) contains proofs for the theoretical results of (ref).
We will use the following conventions throughout the manuscript. $\xrightarrow{p}$ and $\xrightarrow{d}$ denote convergence in probability and distribution, respectively. Coefficients of the true ADF model have a superscript $\star$. We write $\lambda=\Theta(T^\kappa)$ for $0<\liminf_{T\rightarrow\infty}\lambda/T^\kappa\leq\limsup_{T\rightarrow\infty}\lambda/T^\kappa<\infty$, $\kappa\in\mathbb{R}$. Several (estimated) quantities often denoted by Greek symbols depend on the $\ell_1$ penalty and the sample size, which we express by sub-indexing with $\lambda$ and $T$ when important for the exposition. Coefficients $\beta_i^\star$ of the relevant variables identify the data-generating model $\mathcal{M}:=\left\{i:\beta_i^\star\neq0\right\}$. We refer to an estimator $\widehat{\mathcal{M}}$ as model selection consistent if $\lim_{T\to\infty}\ensuremath{\mathrm{P}}\left(\widehat{\mathcal{M}} = \mathcal{M}\right) = 1$ and conservative if $\lim_{T\to\infty}\ensuremath{\mathrm{P}}\left(\widehat{\mathcal{M}} \supseteq \mathcal{M}\right) = 1$ with $\lim_{T\to\infty}\ensuremath{\mathrm{P}}\left(\widehat{\mathcal{M}}\supset\mathcal{M}\right) > 0$, where $\widehat{\mathcal{M}}$ denotes the estimated model.
We consider time series $y_t$ obtained from an autoregressive data-generating process (DGP)
where $\bm z_t = (1,t,\dots,t^q)$ gathers deterministic regressors with coefficients $\bm \psi$. Leading cases are intercept-only models where $\bm z_t = 1$ and models with a linear time trend, $\bm z_t = (1, t)'$. The stochastic component $x_t$ is driven by the error process $u_t$, introducing a unit root into the observable variable $y_t$ if $\varrho = 1$ and is weakly stationary if $\lvert\varrho\rvert<1$. We make the following assumptions about $u_t$.
(ref) states that $u_t$ has a linear process representation with \enquote{long-run} variance (LRV) $\omega^2 := \lim_{T\rightarrow\infty} \operatorname{E}[T^{-1}(\sum_{t=1}^T u_t)^2]<\infty$. This is common in the time series literature and covers various serially dependent processes, including stationary ARMA models. The lag polynomial $\phi(L)$ is invertible, meaning that $u_t$ has an AR representation with potentially indefinitely many lags. This ensures the existence of the ADF($\infty$) representation of $y_t$,
with degree-$q$ deterministic time polynomial $d_t := \bm z_t'\bm \psi$, $\rho^\star\in(-2,0]$ and $\bm \delta^\star := (\delta^\star_1,\dots,\delta^\star_p)'$, where $\sum_{j=1}^p\delta_j^\star<1$. The stochastic component of $y_t$ is stationary if and only if all solutions to the characteristic polynomial
are outside the unit circle, viz. $\rho^\star\in(-2,0)$. If $\rho^\star = 0$, then (ref) has a unit root. Selecting a model that includes (omits) $y_{t-1}$ thus is equivalent to classifying $y_t$ as stationary (non-stationary). We thus refer to $y_{t-1}$ as the inference regressor.
In this paper, we deal with the $\ell_1$-penalised estimation of ADF($p$) models
which either incorporate or approximate the true model (ref), dependent on the choice of $p$ and the DGP's lag order. If the DGP is not sparse, the coefficients $\rho_p$ and $\bm \delta_p$ are flawed with an approximation error, with $\varepsilon_{p,\,t}$ subject to a truncation error inducing autocorrelation. ChangPark2002 establish that $p=o(T^{1/3})$ as $T\to\infty$ ensures consistent estimation with the approximation and truncation errors vanishing asymptotically.
This setting is a generalisation of the sparse processes considered in kock2016consistent and aligns with the concept of \enquote{approximate sparsity} in the high-dimensional regression literature, cf. Chernozhukov2013. Most of the results presented in this paper hold under (ref). We indicate exceptions by invoking the following assumption, which ensures sparsity of the DGP (ref) and hence an ADF representation with $p<\infty$.
The adaptive Lasso optimisation problem to (ref) with $d_t=0$ and given the hyperparameter $\lambda\geq0$ governing the $\ell_1$ penalty obtains as
with its solution denoted by
All adaptive weights $w_1 := \lvert\widehat{\rho}\rvert^{-1}$, $w_{2,\,j} := \lvert\widehat{\delta}_j\rvert^{-1}$ are determined by the initial (OLS) estimates $\widehat{\rho}$ and $\widehat{\delta}_j$ of (ref) and the parameters $\gamma_1>0,\gamma_2>0$ are chosen a priori. The loss function (ref) is as in kock2016consistent except that we add a factor of 2 in front of the $\ell_1$-penalty term for mathematical convenience.
Assuming sparsity, kock2016consistent proves oracle efficiency of $\widehat\bm \beta_\lambda$ for a suitable sequence of the tuning parameter $\lambda\in\lambda_{\mathcal{O}}$, with $\lambda_{\mathcal{O}}$ denoting a continuum of all sequences satisfying certain growth rate conditions. He further establishes for $\rho^\star\in(-2,0]$, i.e., irrespective of whether the model is stationary or not, that
with the non-zero entries of $\bm \beta^\star := (\rho^\star,\bm \delta^{\star'})'$ defining the true model $\mathcal{M}$ and $\widehat{\bm \varepsilon}_\lambda = \Delta \bm y_t - \widehat\rho_\lambda \bm y_{t-1} - \sum_{j=1}^p \widehat\delta_{j,\,\lambda} \Delta \bm y_{t-j}$.\footnote{$\lVert\bm x\rVert_0 = \sum_{i=1}^p \mathbb{I}(\lvert x_i\rvert>0)$ denotes the Hamming distance of a $p$-vector $\bm x$ and $\mathbb{I}$ the indicator function.} Hence, the BIC-tuned adaptive Lasso estimator $\widehat\bm \beta_{\lambda_\textup{BIC}}$ is model selection consistent. In particular, $\widehat\bm \beta_{\lambda_\textup{BIC}}$ consistently classifies stationary and non-stationary models. kock2016consistent cautions that consistent selection via the BIC does not imply the estimation error $\widehat{\bm \delta}_{\lambda_\textup{BIC}} - \bm \delta^\star$ to converge to zero, but rather to be bounded.\footnote{We thank Anders B. Kock for drawing our attention to this detail.} Nevertheless, the oracle property establishes the existence of an estimator that is asymptotically equivalent to the unpenalised OLS estimator of the true model $\mathcal{M}$. It thus appears logical that $\lambda_\mathcal{O}$ asymptotically provides the global minimum of the BIC's objective function.
Due to its oracle properties and the availability of fast algorithms for the optimisation problem underlying the computation like LARS (Efronetal2004), the adaptive Lasso is an attractive tool for model selection in (ref). However, it requires initial estimators to elicit weights for all coefficients. While OLS weights perform reasonably in finite samples, a targeted weight specification may enhance the small-sample performance in discriminating stationary and non-stationary models.\footnote{We thank Anders B. Kock for kindly providing his supplementary simulation results to kock2016consistent upon which we base our Monte Carlo setup.}
It is an open research question of how weights should be chosen in a potentially non-stationary model where it is often of prime interest to make a well-substantiated choice concerning inclusion or omission of $y_{t-1}$, especially for small $T$. In the following, we suggest an alternative weight for $\rho^\star$ that enhances the performance of the adaptive Lasso in distinguishing between stationary and non-stationary autoregressions.
To improve the performance of the adaptive Lasso in discriminating between stationary and non-stationary models, we modify the adaptive weight to exhibit more attractive properties than for $w_1 = \lvert\widehat\rho\rvert^{-1}$. Instead of using only the OLS coefficient estimate, we suggest enriching $w_1$ using a statistic that exploits distinct stochastic orders of stationary and non-stationary $y_t$ to obtain a more favourable adaptive penalty. To this end, we suggest simulation-based computation of a scaling factor to $\widehat{\rho}$ that behaves antithetically to the OLS estimator regarding the limits attained in stationary and non-stationary models. We use the quantile range statistic $J_\alpha$ by HerwartzSiedenburg2010, which is based on distinct convergence rates of OLS estimators in balanced and unbalanced time series regressions.
Phillips1986 shows that for $y_t$ and $q_t$ satisfying Assumption (ref),
where $\widehat\zeta$ is the OLS estimator in the regression $y_t = \zeta q_t + \epsilon_t$.\footnote{Implications of this result are known as spurious regression.} $W_w(a)$ and $W_w(a)$ are independent standard Brownian motion. Hence, $\widehat\zeta$ is inconsistent for the true parameter $\zeta = 0$ but converges to a r.v. with a non-standard distribution. If $y_t$ is weakly-stationary, then $\widehat{\zeta}\xrightarrow[]{p}0$ where $\widehat\zeta= O_p(T^{-1})$. Given these characteristics, we suggest the modified adaptive weight $\Breve{w}_1$.
We refer to $\Breve w_1$ as an information-enriched weight and the resulting Lasso estimator as adaptive Lasso with information enrichment (ALIE). For intuition as to why (ref) is useful, note that $J_\alpha= O_p(1)$ for $\rho^\star = 0$. While the asymptotic distribution of $J_\alpha$ has no closed form, simulation results indicate that it is well approximated for small $T$, cf. the left panel in Figure (ref). The distribution has much probability mass on the interval $(1, \infty)$ and so (ref) has the tendency to upscale the penalty $w_1$ when $y_{t-1}$ is non-stationary. Under stationarity, $J_\alpha$ approaches a degenerate limit at zero at rate $T$, and so (ref) may downscale $w_1$. (ref) states the asymptotic behaviour of $\Breve{w}_1$.
Note that $\Breve w_1$ grows unlimited as $T\to\infty$ in a non-stationary model by part (ref) of (ref), just as $w_1$. However, small-sample properties of the weights differ as is illustrated in the right panel of (ref), which shows samples of $\Breve w_1$ and $w_1$ in $\log$-$\log$ space for DGP (ref) with $u_t\sim i.i.d.\,N(0,1)$ and $T=100$. We deduce that the joint distribution for $\rho^\star=0$ (orange dots) has the most probability mass above the dashed bisector line, i.e., $\Breve w_1 > w_1$ on average, implying a higher penalty from setting $\dot\rho\neq0$.
Under stationarity, $w_1-\lvert\rho^\star\rvert^{-1}=O_p(T^{-1/2})$ whereas $\Breve{w}_1\xrightarrow[]{p}0$ at rate $O_p(T^{-1})$ by part (ref) of (ref). The effect is observed in the right panel of (ref) where samples (green dots) indicate that $\Breve{w}_1<w_1$ for $\rho^\star=-.1$, on average. The faster convergence rate and the zero limit of $\Breve w_1$ are expected to benefit the finite sample performance since a small penalty from setting $\dot\rho\neq0$ renders ALIE less likely to exclude $y_{t-1}$ from the model, misclassifying $y_t$ as non-stationary. We have several remarks.
We next present results on consistency and consistent tuning of ALIE.
(ref) establishes the asymptotic (oracle-)equivalence of ALIE and the adaptive Lasso with OLS-based weights of kock2016consistent, henceforth AL, for sparse processes.\footnote{Different lower bounds on $\gamma_1$ and $\gamma_2$ are due to distinct convergence rates of the initial estimates in a non-stationary model where $\widehat{\rho}= O_p(T^{-1})$ and $\widehat{\delta}_j-\delta_j^\star = O_p(T^{-1/2})$.} This implies consistency of ALIE for sequences $\lambda \in \lambda_\mathcal{O}$ satisfying the same asymptotic conditions as required for AL: Given $\lambda \in \lambda_\mathcal{O}$, ALIE discriminates consistently between the sets of relevant and irrelevant regressors and the estimated coefficients are asymptotically equivalent to OLS estimates in the true model. Furthermore, ALIE can be tuned using BIC to achieve consistent model selection, giving practical guidance on choosing $\lambda$ in applications.\footnote{The Lasso literature also considers plug-in estimators for $\lambda_\mathcal{O}$, but to our knowledge, no method performs as well as the BIC in our application.}
In the following, we present several asymptotic results for ALIE regarding consistent classification of $y_{t-1}$, conditional on the regularisation parameter $\lambda$. For ease of exposition and comparability with the results for AL in kock2016consistent, we consider an ADF(0) model and set $\gamma_1 = 1$.
(ref) is along the lines of Theorem 3 in kock2016consistent and considers properties in non-stationary models for $T\to\infty$. Part (ref) states that $\lambda\to0$ yields an inconsistent classifier. By part (ref), $\lambda\to\infty$ is required for ALIE to correctly classify a non-stationary model in the limit. Both results also apply to AL. Part (ref) states that if $\lambda$ converges to some constant $c>0$, then $\widehat{\rho}_\lambda = 0$ with non-zero probability $F_{\Breve{\mathcal{H}}}(c)$.
We next discuss the asymptotic behaviour of ALIE in a stationary model conditional on $\lambda$.
(ref) is our main result and describes the asymptotic behaviour of $\ensuremath{\mathrm{P}}\left(\widehat{\rho}_\lambda=0\right)$ for ALIE if $\rho^\star\in(-2,0)$, conditional on $\lambda$. Part (ref) states that $\lambda=\Theta(T^{1+\kappa})$ with $\kappa<1$ is required for $\ensuremath{\mathrm{P}}\left(\widehat{\rho}_\lambda = 0\right) = 0$ asymptotically, implying 100% power. By part (ref), ALIE will incorrectly classify the model as non-stationary with probability one as $T\to\infty$ if $\kappa>1$. Part (ref) considers $\lambda=\Theta(T^2)$, the knife edge between parts (ref) and (ref). The classification result then is random. This finding generalises an analogous result in kock2016consistent, who imposes an additional assumption on the mode of convergence of a specific quantity. We prove in (ref) in (ref) that this assumption is not necessary, resulting in a different distribution of the activation probabilities.
We conclude that ALIE performs consistent classification for $\rho^\star\in(-2,0]$ if $\lambda=\Theta(T^{1+\kappa})$, $\kappa<1$, satisfying part (ref) of (ref) ($\lambda\to\infty$ required if $\rho^\star = 0$) and part (ref) of (ref) ($\kappa<1$ required if $\rho^\star\in(-2,0))$. By Theorem 4 in kock2016consistent, AL only has power for $\lambda = O(T)$. (ref) relaxes this constraint to $\lambda = O(T^2)$ for ALIE, which is a direct consequence of the properties of $\Breve{w}_1$ stated in part (ref) of (ref).
While the $\lambda$-conditions of (ref) are weaker than $\lambda_\mathcal{O} = o(\sqrt{T})$, as required for the oracle property, cf. (ref), they suffice for ALIE to consistently select stationary models. If the goal is distinguishing stationary and non-stationary data by classification, the lack of oracle efficiency is negligible. Importantly, (ref) foreshadows a power advantage of ALIE over AL in stationary ADF($p$) models under consistent tuning of $\lambda$, and thus also for oracle-efficient tuning using BIC. We will discuss this in the next section.
The asymptotic analysis for the ADF(0) model presented in the previous section builds on the activation probability of $y_{t-1}$, which is analysed based on the first-order condition (FOC) for $\dot\rho_\lambda = 0$. We next transfer the results of the previous section to general ADF models in terms of the penalty parameter $\lambda$ by examining the Lasso solution path. For this, we use the activation threshold $\lambda_0$.
For every $\lambda \in \Lambda_{0,\, \beta_i^\star}$, the FOC for $\dot{\beta}_i(\lambda) = 0$ is just binding, i.e., it holds with weak inequality whilst the coefficient $\dot{\beta}_i(\lambda)$ remains set to zero. Following the interpretation in Efronetal2004 closely, we declare regressor $x_i$ active as soon as the FOC is binding, so that $x_i \in \widehat{\mathcal{M}}_{\lambda_{0,\,\beta_i^\star}}$ even though $\dot{\beta}_{i}(\lambda_{0,\,\beta_i^\star}) = 0$. Note that the FOC for $\dot{\beta}_i(\lambda) = 0$ can also be just binding with a zero coefficient if the regressor is being removed from the active set or is on the verge of a sign reversal.\footnote{A projected sign reversal is the sole reason for a variable to be deactivated by the LARS algorithm with the Lasso modification, see Section 3.1 in Efronetal2004.} By imposing the right-hand derivative to be zero while the left-hand derivative is non-zero, (ref) only considers knots on the solution path where variable $i$ is added to the active set. Note that the solution path for the ADF(0) model in the previous section features a single activation knot $\Lambda_{0,\,\rho^\star} = \lambda_{0,\,\rho^\star}$ since $y_{t-1}$ is the only regressor.
In least angle regression (LAR), the $\Lambda_{0,\, \beta_i^\star}$ are singletons because active variables cannot be deactivated. However, Lasso and its variants allow for repeated activation and deactivation of variables along the solution path.\footnote{See Efronetal2004 for details on the LARS algorithm in Lasso mode.} Still, the earliest activation knot of variable $i$, $\lambda_{0,\, \beta_i^\star}$, is the most informative because it is mainly driven by the shrinkage applied to the associated coefficient. (ref) describes the asymptotic behaviour of the activation thresholds for the relevant and the irrelevant variables in stationary and non-stationary ADF models.
Part (ref) of (ref) corroborates an asymptotically equivalent effect of $w_1$ and $\Breve w_1$ when $\rho^\star=0$, cf. (ref). Denote by $\lambda_{0,\,\rho^\star\in(-2,0)}$ the activation threshold for a relevant stationary $y_{t-1}$ for brevity. Part (ref) implies that $$\lambda_{0,\,\rho^\star\in(-2,0)} = o\left(\Breve\lambda_{0,\,\rho^\star\in(-2,0)}\right)\quad\forall\quad\gamma_1>0.$$ This is one main feature of information enrichment: $\Breve w_1$ implies $\Breve\lambda_{0,\,\rho^\star\in(-2,0)} =\Theta(T^{1+\gamma_1})$, whereas AL's $w_1$ yields $\lambda_{0,\,\rho^\star\in(-2,0)} =\Theta(T)$ -- just as for the $\lambda_{0,\,\delta_j^\star\neq0}$ of the relevant $\Delta y_{t-j}$ with $\sqrt{T}$-consistent weights $w_{2,\,j}$. Part (ref) asserts the stochastic orders of the $\lambda_{0,\,\delta_j^\star}$ associated with the relevant and irrelevant $\Delta y_{t-j}$. These orders are the same for ALIE and AL, irrespective of $\rho^\star$.
We conclude from (ref) that distinct orders of $w_1$ and $\Breve{w}_1$ influence the selection ability of the adaptive Lasso for $\rho^\star \in (-2,0)$. However, asymptotic performance differences cannot be explained by growth rate considerations for $\rho^\star=0$, where benefits of information enrichment are different, as discussed in the next section. We focus on $\rho^\star \in (-2,0)$ in the remainder.
As for the ADF(0) model considered in (ref), the asymptotic shrinkage applied to $\widehat{\rho}_\lambda$ in an ADF($p$) model depends on the growth rate of the $\lambda$-sequence. ALIE has the virtue that below $\Breve{\lambda}_{0,\, \rho^\star \in (-2,0)}$ the shrinkage on $\widehat{\rho}_\lambda$ vanishes asymptotically, as we assert in (ref).
Importantly, zero asymptotic shrinkage is not synonymous with unbiasedness or consistency because $\widehat{\rho}_\lambda$ may be subject to omitted variable bias or an approximation error. Unbiasedness is only guaranteed at $\lambda_\mathcal{O}$ where the shrinkage on the coefficients of the relevant variables vanishes for $T\to\infty$, as required for the oracle property. (ref) is useful nonetheless: it transfers the findings for the stationary ADF(0) model in (ref) to ADF($p$) models. (ref) further implies an asymptotic ordering in every possible Lasso solution path across all permissible stationary linear models, as the subsequent corollary details.
Part (ref) of (ref) states that ALIE activates $y_{t-1}$ first w.p. 1 in any stationary linear model as $T\to\infty$. This feature is unique to ALIE and has not arisen in any other Lasso variant, to our knowledge. Part (ref) asserts that $\widehat\rho_\lambda$ escapes penalisation as some $\Delta y_{t-j}$ enters the active set (with zero coefficient) at $\lambda^{(2)}$, rendering $\widehat\rho_{\lambda^{(2)}}$ equivalent to the unpenalised OLS estimator of $\rho$ in the (possibly underspecified) model $\Delta y_t = \rho y_{t-1} + u_t$, with the non-zero probability limit $\rho^{\star\star}$. Remarkably, although $\rho^{\star \star}$ does not necessarily equal $\rho^\star$, detection of stationarity is asymptotically guaranteed at $\lambda^{(2)}$.\footnote{(ref) hence motivates another principle to test for a unit root: checking whether $y_{t-1}$ is activated first. This test is consistent but has low power, as simulation evidence suggests.} The results of (ref) are not granted to hold for AL since $\lambda_{0,\,\rho^\star}$ has the same stochastic order as the $\lambda_{0,\,\delta^\star_j}$ of the relevant $\Delta y_{t-j}$. Therefore, the asymptotic activation knot of $y_{t-1}$ is subject to the DGP and thus unknown.
Combining (ref), (ref) and (ref), we obtain a further result on ALIE's solution path.
(ref) states that once $y_{t-1}$ has been activated moving down the solution path, the probability to encounter a deactivation knot converges asymptotically to zero. This renders ALIE robust to restrictive tuning, a common finite-sample feature of consistent tuning criteria that is also observed for the BIC, cf. Efronetal2004.
In applications, $\lambda$ is often selected by data-driven methods, e.g., cross-validation or information criteria, usually targeting some $\lambda = O(\lambda_\mathcal{O})$. Even though the asymptotic properties of these tuning methods have been established under various conditions, their stochastic behaviour in small samples needs to be better understood. Efronetal2004 remark that consistent tuning criteria like the BIC tend to choose too large values for $\lambda$ in small samples, thereby neglecting relevant variables. To emulate this behaviour, we assume the growth rate of a tuned parameter sequence $\lambda_\textup{IC}$ to vary stochastically in finite samples. For comparison of $\lambda$-sequences with varying growth rates, we apply the $\log_T$-transformation. Occasions of restrictive tuning as addressed above translate to $\log_T(\lambda_\textup{IC}) > \log_T(\lambda_\mathcal{O})$. Note that in $\log_T$-space, all $\lambda$-sequences are asymptotically bounded and can be considered to be drawn from a continuous distribution function $F_{T,\,\lambda_\textup{IC}}(\log_T \lambda \vert \bm \varepsilon_T, \bm \beta^\star)$, shorthand $F_{T,\,\lambda_\textup{IC}}$. We account for the dependence of $F_{T,\,\lambda_\textup{IC}}$ on the DGP by conditioning on the innovations $\bm \varepsilon_T$ and the vector of true coefficients $\bm \beta^\star$. The distribution function $F_{T,\,\lambda_\textup{IC}}$ has support on $\mathbb{R}^+$ and converges to its limit $F_{\lambda_\textup{IC}}$, inheriting the asymptotic properties of the tuning criterion applied.\footnote{E.g., $F_{\lambda_\textup{BIC}}$ has all probability mass on $\lambda\in\lambda_\mathcal{O}$ asymptotically while $F_{\lambda_\textup{AIC}}$ distributes some probability mass below $\lambda_\mathcal{O}$, see kock2016consistent.} This convergence also acknowledges the fact that $\lambda_\mathcal{O}$ is only guaranteed to exist as $T \to \infty$. Under consistent tuning using the BIC, deviations of $\lambda_\textup{BIC}$ from $\lambda_\mathcal{O}$ can happen as $\log_T \lambda_\textup{BIC} \to \log_T \lambda_\mathcal{O}$. This is reminiscent of the local alternatives regarding coefficients widely used in the unit root literature.
The activation probability of $y_{t-1}$ for Lasso-type estimators is bounded by $F_{T,\,\lambda_\textup{IC}}$,
with generic first activation threshold $\widetilde{\lambda}_{0,\, \rho^\star}$. (ref), (ref) and (ref) point to an advantage for ALIE over AL concerning the conditional activation probability $\ensuremath{\mathrm{P}}\left(\widehat{\rho}_{\lambda_\text{IC}} \neq 0 \vert \lambda_\text{IC} < \widetilde{\lambda}_{0,\,\rho^\star}\right)$. The same applies for $F_{T,\,\lambda_\text{IC}}\left(\log_T \widetilde{\lambda}_{0,\,\rho^\star}\right)$, which is ultimately driven by the different growth rates of $\lambda_{0,\,\rho^\star}$ and $\Breve\lambda_{0,\,\rho^\star}$ (cf. (ref)). Maximising $\ensuremath{\mathrm{P}}\left(\widehat{\rho}_{\lambda_\text{IC}} \neq 0\right)$ requires maximising $F_{T, \lambda_\text{IC}}\left(\log_T \widetilde{\lambda}_{0,\,\rho^\star} \right)$ which can either be achieved using another tuning method or by lifting $\widetilde{\lambda}_{0,\,\rho^\star}$. However, a different tuning method may deteriorate classification performance for $\rho^\star=0$. Information enrichment via $\Breve{w}_1$ does not share this disadvantage as it manipulates $\widetilde{\lambda}_{0,\,\rho^\star}$, conditional on $\rho^\star$. Since the CDF $F_{T,\,\lambda_\text{IC}}$ is monotonically increasing in $\lambda$, raising the growth rate of $\Breve\lambda_{0,\,\rho^\star \in (-2,0)}$ adds robustness against the finite sample characteristics of $\lambda_\textup{IC}$, a prominent feature in our simulation studies.
(ref) illustrates the finite sample impact of the results in (ref) and (ref). Here, we compare simulated $\lambda_0$ distributions for the regressors in an over-specified ADF model to the BIC-tuned penalty parameter $\lambda_\text{BIC}$ for ALIE and AL. We set $\rho^\star=0$ or $\rho^\star=-.05$, let $\bm \delta^\star=(.275,\,.25,\,.2,\,.15,\,.1)'$ and estimate model (ref) with $p=10$. For ease of interpretation, we adopt the popular choice $\gamma_1 = \gamma_2 = 1$ which translates to $\lambda_0 = O_p(1)$ for all irrelevant regressors (dashed lines). Activation thresholds of the relevant regressors (solid lines), i.e., $\lambda_{0,\,\rho^\star\in(-2,0)}$ and $\lambda_{0,\,\delta_j^\star\neq0}$, are $O_p(T)$ and $\Breve\lambda_{0,\,\rho^\star\in(-2,0)} = O_p(T^2)$, by (ref). Note that $F_{T,\,\lambda_\textup{BIC}}$ converges to a stable asymptotic distribution $F_{\lambda_\textup{BIC}}$ with support on $\mathbb{R}^+$.
For $\rho^\star=-.05$ (top panel), the asymptotic implications of (ref), part (ref) already emerge for $T=75$. The empirical distribution of $\lambda_{0,\,\rho^\star \in (-2,0)}$ is shifted towards larger values, reducing the overlap with the density of $\lambda_\textup{BIC}$. Following (ref), this feature suggests an improved activation probability from information enrichment. Consistent with the theory, these effects become more pronounced for $T=200$. We also find high congruence between distributions of the remaining variables for ALIE and AL for $T=75$ and $T=200$, corroborating part (ref) of (ref).
The bottom panel of (ref) shows results for $\rho^\star = 0$. Unlike for $\rho^\star=-.05$, the $\Delta y_{t-j}$ are weakly correlated with $y_{t-1}$, meaning that estimates are expected to be less sensitive to the inclusion or omission of $y_{t-1}$. We hence find the distributions of the $\lambda_{0,\,\delta_j^\star}$ to be virtually identical for ALIE and AL.\footnote{The distributions of $\lambda_{0,\,\delta_j^\star}$ are identical for the irrelevant lags since $\Delta y_{t-j}$ is always stationary.} However, $\Breve\lambda_{0,\,\rho^\star}$ displays a shift of probability mass towards small values of $\Breve\lambda_{0,\,\rho^\star}$, reducing overlap with the density of $ \lambda_\textup{BIC}$. This is a desirable property that, however, does not stem from the asymptotic results detailed in this section. The shape of this distribution matters most in applications because the oracle properties impose zero overlaps between $F_{T,\, \lambda_\textup{BIC}}$ and $\Breve\lambda_{0,\,\rho^\star = 0}$ only as $T \to \infty$. The following section provides further insights by discussing information enrichment in a zero-mean model.
Comparing the medians (vertical lines) of $\lambda_{0,\, \rho^\star}$ for $\rho^\star = -.05$ and $\rho^\star = 0$ reveals how combining two imperfectly correlated statistics translates into an improved performance via projection on the activation thresholds: information enrichment drives an additional wedge between $\Breve\lambda_{0,\,\rho^\star = 0}$ and $\Breve\lambda_{0,\,\rho^\star \in (-2,0)}$, which is captured in the following remark.
As $\lambda_{0, \rho^\star = 0}$ and $\Breve\lambda_{0, \rho^\star = 0}$ have the same stochastic order, ALIE's size advantage shown in (ref) and our simulation results must root from differing shapes of the respective distributions $F_{\lambda_0}$, $F_{\Breve{\lambda}_0}$ for AL and ALIE. Therefore, the relationship between information enrichment and the shape of $F_{\Breve{\lambda}_0}$ will be critical in our considerations below.
Suppose the setup
with $\sigma_\epsilon,\sigma_{\nu,\,j}<\infty$, $\mu=0$, and statistics $z_j$, $ j=1,\dots,q$, providing additional information on $\mu$. Also, let $\{\epsilon,\nu_1,\cdots,\nu_q\}$ be a set of independent noise terms. The measurements $x$ and $z_j$ can represent single observations or sample averages. By defining $\mu := \rho^\star y_{t-1}$, inference on $\rho^\star$ is tantamount to inference on $\mu$ with the null hypothesis being $\mu = 0$.\\ An adaptive Lasso estimator for $\mu$ in (ref) is the minimiser
with the FOC for $\widehat\mu_w = 0$ resulting in
The adaptive Lasso estimator suggested by zou2006adaptive uses exclusively $x$ to determine the penalty weight, $w:=\lvert1/x\rvert^\gamma$. Setting $\dot{\mu} = 0$ and $\mu = 0$ in (ref), one obtains $w=\lvert1/\epsilon\rvert^\gamma$ and $\lambda_0:=\lvert\epsilon\rvert^\gamma\lvert\epsilon\rvert$. We suppose the penalty parameter $\lambda_\textup{IC}$ to be drawn from a tuning distribution $F_{\lambda_\textup{IC}}$ independent of $x$ and $z_j$. Because the model (ref) has a single regressor, $\ensuremath{\mathrm{P}}\left(\widehat\mu_w\neq0\bigr\vert\lambda_0>\lambda_\textup{IC}\right) = 1$ is ensured. The probability to misclassify $\mu$ simplifies to
An information-enriched adaptive Lasso estimator $\widehat\mu_{\Breve w}$ adds the $z_{j}$, yielding
as $z_j = \nu_j$ since $\mu = 0$. The enriched estimator has the activation probability
Accordingly, reducing the classification error of $\mu$ requires shrinking $\Breve\lambda_0$, i.e., shifting $F_{\Breve\lambda_0}$ closer to zero. In (ref), we examine how the (erroneous) activation probabilities compare for $\widehat\mu_w$ and $\widehat\mu_{\Breve w}$.
Part (ref) of (ref) states how the probability for a general adaptive Lasso estimator $\widehat\mu_w$ to activate $\mu$ in model (ref) depends on the data $x$ through $\lambda_0$ and the employed tuning procedure, viz. $\ensuremath{\mathrm{P}}\left(\widehat\mu_w\neq0\right)=\ensuremath{\mathrm{P}}\left(\lambda_0>\lambda_\textup{IC}\right)$. Part (ref) asserts that $\ensuremath{\mathrm{P}}\left(\widehat\mu_{\Breve{w}}\neq0\right)\to0$ as $q\to\infty$ for an information-enriched adaptive Lasso estimator $\widehat\mu_{\Breve w}$, i.e., enriching the adaptive weight $\Breve w$ with additional information on $\mu$ obtains an estimator that dominates any $\widehat\mu_w$ in terms of activation probability.
Part (ref) of (ref) gives a condition for $\ensuremath{\mathrm{P}}\left(\widehat\mu_{\Breve w}\neq0\right)$ to be bounded from above by the activation probability of any adaptive Lasso estimator for some $q\in(0,\infty)$: we require that $\ensuremath{\mathrm{P}}\left(\widehat\mu_w\neq0\right)=a$ corresponds to an upper tail event of $\lambda_0$ that is less likely to occur for $\Breve\lambda_0$. While this statement is independent of $q$, note that $\lim_{q \to \infty} A_q = (0,1)$ is a consequence of part (ref). Hence, enriching the information on $\mu$ enlarges the potential of $\widehat\mu_{\Breve w}$ to improve upon $\widehat\mu_w$ as the set of dominating activation probabilities $a$ converges to the entire unit interval.
We next examine the information-enriched adaptive Lasso properties in a Gaussian zero mean model.
Result (ref) of (ref) states that in a zero-mean model, $\prod_{j=1}^q \operatorname{E}(\lvert \nu_j\rvert^\gamma) = o(1)$ is sufficient for the probability of $\widehat\mu\neq 0$ to converge to zero as the number of informative statistics $q$ diverges. Since the $\lvert\nu_j\rvert$ are half-Gaussians, this condition is equivalent to restricting the signal-to-noise ratio of the enriching statistics $z_j$ in $\Breve{w}=(1/\prod_{j=1}^q z_j)^\gamma$. Result (ref) of (ref) provides a tangible condition for $\lim_{q\rightarrow\infty}\ensuremath{\mathrm{P}}\left(\widehat\mu_{\Breve w}\neq0\right) = 0$, requiring an upper bound on the noise added to $\Breve w$ by the enriching statistics. Note that $\prod_{j=1}^q \sigma_{\nu,\,j} = o\left( (\pi/2)^{q/2}\right)$ is trivially satisfied if $\sigma_{\nu,\,j} = 1 \, \forall j$, i.e., it is permitted to employ a (possibly infinite) set of independent Gaussian statistics to enrich $\Breve{w}$.
The advantage of $\widehat\mu_{\Breve w}$ over $\widehat\mu_w$ can be seen from the left panel of (ref), which compares activation probabilities for $\gamma = 1$ and Gaussian $z_j$. Information enrichment may yield significant improvements, even for $q=1$ (red curve). Including additional information results in even superior estimators. A comparison of the activation probabilities for $n = 50$ and a single correlated enriching statistic is shown in the right panel of (ref). Here, $w=\lvert\overline{\boldsymbol{x}}\rvert^{-1}$ and we set $\Breve{w}=\lvert t_{z_1}\cdot\overline{\boldsymbol{x}}\rvert^{-1}$ where $t_{z_1}$ is the $t$-statistic for $H_0:\mu=0$. Again, we find the significant potential of $\Breve{w}_1$ to mitigate the risk of spurious activations.
In the following, we investigate the empirical properties of ALIE in more detail and compare its performance to AL using simulation studies. For implementing $\breve{w}_1$, it is crucial to scale $y_t$ by a consistent estimator of the long-run standard deviation $\omega$ such that the limit distribution of $J_\alpha$ given $\rho^\star = 0$ is nuisance-parameter-free. We follow HerwartzSiedenburg2010 and estimate $\omega^2$ using the AR spectral density estimator at frequency zero,
suggested by PerronNg1998. The $\widehat{\varepsilon}_{k,\,t}$ are residuals from estimating (ref) by OLS with lag order $p=k$. To compute (ref), we estimate $k$ as $\widehat{k}:= \operatorname*{arg\,min}_{0\leq k\leq k_{\max}} \textup{IC}(k)$ where $\textup{IC}(k)$ is an information criterion of the form
$k_{\max}$ is a maximum lag order and $\tilde\sigma^2_k := (T-k_{\max})^{-1}\sum_{t=k_{\max}+1}^T \widehat{\varepsilon}_{k,\,t}^2$ a deviance estimate. $C_T$ is a penalty function and $\tau_T(k)$ a stochastic adjustment term. Popular choices are $C_T = \log(T)$ and $\tau_T(k) = 0$ (BIC), $C_T = 2$ and $\tau_T(k) = 0$ (AIC). NgPerron2001 have suggested modified versions for truncating long autoregressions. These are MBIC and MAIC where $\tau_T(k) = \left(\tilde{\sigma}_k^2\right)^{-1}\widehat{\rho}^2\sum_{k_{\max}+1}^T y_{t-1}^2$, respectively.
To estimate the LRV by $\widehat{\omega}^2_{\textup{AR}}(k)$, we use BIC to select $k$ if not indicated otherwise for sparse processes covered by (ref) and MAIC for processes that do not allow consistent model selection when $T<\infty$. The adaptive Lasso estimator (ref) is computed in a model with $p=k_{\max} = \lfloor12 (T/100)^{1/4}\rfloor$, cf. Schwert1989.
All simulations were performed using the statistical programming language R R. We used the LARS algorithm implemented in the lars package pkg-lars for computing the Lasso solution paths.
As outlined in (ref), $\Breve{w}_1$ can be calibrated via $\alpha$ and $\sigma_v$. The rationale is that larger values of $\alpha$ narrow the distribution of $J_\alpha$ and thus impact the behaviour of $\Breve{w}_1$. By definition of $\Breve{w}_1$, we expect a positive association between $\alpha$ and $\ensuremath{\mathrm{P}}\left(\widehat{\rho}_\lambda\neq0\right)$ for $\rho^\star\in(-2,0]$. A similar argument can be made for $\sigma_v$, which governs the variance of $J_\alpha$.
To illustrate the effect of both parameters on the activation rate of $y_{t-1}$, we simulate $y_t$ using DGP (ref) for $T=100$ and consider $\alpha\in(0, .5)$ for $\sigma_v=1$ and $\alpha=.1$ for $\sigma_v\in(0,2]$. The results are shown in Figure (ref). Activation rates increase with $\alpha$ so that power gains are accompanied by higher misclassification rates for non-stationary $y_t$ (left panel). Similarly, activation rates increase with $\sigma_v$ (right panel).
Notice that the comparison with AL indicates room for tuning. For example, ALIE could be calibrated to have an activation rate close to that of AL for non-stationary $y_t$ but to exhibit more attractive selection rates under stationarity. Hereafter, we follow HerwartzSiedenburg2010 and set $\sigma_v=1$ and $\alpha=.1$ which seem reasonable. However, we note that different choices may be more suitable when the expected loss from a misclassification is not identical for $\rho^\star=0$ or $\rho^\star\in(-2,0)$. This may be relevant in forecasting where the loss from falsely classifying a model as stationary rather than excluding the inference regressor can be quite different than vice-versa.
To illustrate the effect of $\Breve{w}_1$, we first consider BIC-tuned adaptive Lasso estimates of (ref) with lag order $p=\lfloor12(T/100)^{.25}\rfloor$ for simple processes. For this we simulate $y_t$ according to DGP (ref) without a deterministic component with coefficient $\varrho=\rho^\star+1\in\{.95, 1\}$, and errors $u_t \sim i.i.d.\,N(0,1)$. We consider $T\in\{25,50,100,150,250,500,1000\}$. $\breve{w}_1$ is computed using $J_{.1}$ with the LRV estimated by $\widehat{\omega}_{\text{AR}}^2(0) = \widehat{\sigma}^2_0$.
The results are presented in Table (ref). We find that in the case of a random walk $(\rho^\star=0)$, there is improvement in the rate of correct classification of ALIE over AL for $T<150$, where ALIE never exceeds the probability of activating $y_{t-1}$ of AL. This benefit diminishes for large $T$. A comparison of the inference regressor weights shows that $\Breve{w}_1$ implies a higher average penalty from including $y_{t-1}$ in the model. We also find that the average activation threshold on the Lasso solution path decreases faster for ALIE than for AL as $T$ grows.
For stationary autoregressions ($\rho^\star = -.05$), we also find improvement in the discriminatory power of the procedure. For larger $T$, $\Breve{w}_1$ yields a pronounced advantage. E.g., for $T=250$, ALIE has approximately 16% higher probability of activating $y_{t-1}$ than AL. Finite sample implications of our theoretical results in (ref) are mirrored by the average weights and activation thresholds: while $w_1$ converges at rate $\sqrt{T}$ towards its probability limit $\lvert\rho^\star\rvert^{-1} = 20$, the modified $\Breve{w}_1$ yields a lower adaptive penalty for $T>50$. We also find that $\Breve{\lambda}_{0,\,\rho^\star}$ grows at a faster rate than $\lambda_{0,\,\rho^\star}$.
To our knowledge, no literature has examined the properties of adaptive Lasso in the ADF regression (ref) with $d_t\neq0$ to date. For trend filtering, tibshirani2011solution give a comprehensive overview of past contributions to fitting polynomial trends or structural breaks components using the Lasso and its variants. While these approaches to removing deterministic components appear promising, a theoretical analysis is beyond the scope of this article. We instead investigate the performance of the estimators on detrended data by simulation.\footnote{kock2016consistent does not discuss properties of AL when $d_t\neq0$ and notes that oracle properties carry over to models with detrended data.}
Computation of $\Breve{w}_1$ requires special care in the presence of deterministic terms. For the asymptotic distribution of $J_\alpha$ to be invariant to $\bm \psi\neq0$ one needs to adapt (ref) and the LRV estimator (ref), for which OLS or quasi-difference detrending has been suggested, cf. PerronQu2007. Detrending affects the distribution of the OLS estimator $\widehat\zeta$ in (ref) and hence alters the null distribution of $J_\alpha$. To see this for OLS detrending, note that $W_y$ in (ref) is replaced by
with $\bm D(a) := \lim_{n\rightarrow\infty}\bm S_n^{-1}\bm z_{[na]}$, $a\in[0,1]$ for a suitable scaling matrix $\bm S_n$. $\bm D(a)=1$ under demeaning where $\bm z_t = 1$ and $\bm D(a)=(1,a)'$ if $\bm z_t = (1,t)'$. Simulation results for $\bm z_t = (1,t)'$ indeed reveal higher kurtosis, accompanied by a shift of the distribution of $J_\alpha$ towards zero. The simulated null distribution of $J_{\alpha}^{\tau}$ under OLS detrending has more probability mass below the threshold of $1$ than the distribution of $J_\alpha$, implying an undesired deflation of $\Breve{w}_1^\tau$ relative to $\Breve{w}_1$ in the (no-)intercept model, see the left panel of Figure (ref).\footnote{This shift is even stronger pronounced for QD detrending where $W_y^d$ is replaced by the corresponding projection of the QD-detrended data, see Phillips1998.} This tendency to yield smaller weights for non-stationary $y_t$ is also seen from a comparison of the cumulative distribution functions (CDFs) of $\Breve{w}_1^\tau$ and $w_{1}$ in the right panel of (ref).
We thus expect OLS detrending to inflate activations for $\rho^\star\in(-2,0]$.
We propose the trend-adjusted simulated weight $\Breve{w}_{1}^{\tau^\star}$ based on a modified trend-adjusted statistic $J_{\alpha,\,\sigma_v}^{\tau^\star}$.
The procedure is similar to (ref) with the difference that we incorporate a linear time trend in the regressions underlying the LRV estimation and the simulation. Following the results in section (ref), we suggest simulating random walks with $\sigma_v\leq1$ in step 3 to attenuate spurious activations of $y_t$. Simulation results for $J_{\alpha,\,\sigma_v}^{\tau^\star}$ and the associated $\Breve{w}_{1}^{\tau^\star}$ in Figure (ref) indeed reveal null distributions with more favourable properties than for the OLS-detrended variants $J^\tau_\alpha$ and $w^\tau_1$.\footnote{We use the ad-hoc choice $\sigma_v = .75$ for better size control of $J_{\alpha,\,\sigma_v}^{\tau^\star}$ for small $T$.}
We next investigate finite sample properties of ALIE and AL for sparse processes and also evaluate the ability of the (inconsistent) plain Lasso (PL) to distinguish between stationary and non-stationary models. Data are simulated with the ADF-DGP
with
and $k^\star+1$ zero initial values where the lag order is $k^\star := \dim(\bm \delta^\star)$. To investigate the model selection performance when there is sparsity in the coefficients of the $\Delta y_{t-j}$, we set $\bm \delta_A^\star := (.4, .3, .2)'$ and $\bm \delta_B^\star := (.4, .3, .2, 0, 0, 0, -.2, 0, 0, .2)'$. $\bm \delta_C^\star := .7$ yields a short AR process. $\bm \delta_D^\star:= (-.4, 0, .7)'$ is taken from CanerKnight2013 and is a low power setting for unit root tests. We compute Lasso estimates of model (ref) with lag order $p=\lfloor12\cdot(T/100)^{.25}\rfloor$ and tune $\lambda$ using BIC. We consider $T\in\{25,50,100,150,200,250\}$, let $\rho^\star=0$ for non-stationarity and $\rho^\star = -.05$ to investigate the stationary case. For $\Breve w_1$ we estimate the LRV by $\widehat\omega^2_\text{AR}(k)$ and compute $J_\alpha$ as outlined in Sections (ref) and (ref).
Accommodating linear trends is non-trivial. While model selection consistency should be possible for detrended data, unreported simulations indicate that OLS detrending leads to prohibitive activation rates of $y_{t-1}$ for the DGPs under consideration when $\rho^\star=0$. To provide some guidance for empirical application, we present results for first-difference (FD) detrending (SchmidtPhillips1992), which seems more reliable for sparse processes.\FloatBarrier
(ref) shows activation rates of $y_{t-1}$ without deterministic components (top panel), for FD demeaning (middle panel) and for FD detrending (bottom panel). Spurious activation rates for $\rho^\star = 0$ (first sub-panel) indicate deficiencies for AL in small samples and across most settings for $\bm \delta^\star$. For instance, at $T=25$ observations, the activation probability of $y_{t-1}$ in settings $\bm \delta_A^\star$ and $\bm \delta_B^\star$ exceeds 40%, whereas ALIE achieves considerably lower rates. Similar outcomes are found for demeaned and detrended data. Although the discrepancy between ALIE and AL diminishes in $T$, these results indicate an advantage of $\Breve{w}_1$ over $w_1$ in finite samples.
Results for $\rho^\star=-.05$ (second sub-panel) corroborate favourable finite sample implications of our theoretical results in (ref) under stationary. ALIE's activation rates for $y_{t-1}$ are above AL's across most settings. For example, ALIE has power advantages of up to 20% for sample sizes between $T=100$ and $T=150$ in setting $\bm \delta_B^\star$. Outcomes are qualitatively similar when adjusting for deterministic components. However, all methods suffer from reduced power under detrending, where gains from information enrichment are smaller. This is most pronounced in scenario $\bm \delta_D^\star$ where none of the methods dominates.
(ref) also indicates the mediocre performance of the BIC-tuned plain Lasso in most settings: PL often cannot compete with AL, let alone ALIE in power, despite comparable activation rates under non-stationarity. At worst, the inconsistency of the plain Lasso entails erratic behaviour, cf. setting $\bm \delta_B^\star$ for $\rho^\star = 0$ where rejection rates display upward surges or plateau as $T$ increases.
Further insight into the methods' suitability as classifiers for stationary and non-stationary models is given in (ref). Here, we report the positive predictive value (PPV) and the negative predicted value (NPV),
based on pooled simulation results $\mathcal{S}$ for $\rho^\star\in(-2,0]$ for $T=100$.\footnote{We consider $\widehat{\rho}_\lambda<0$ a \enquote{positive}.} The comparison indicates that ALIE is the superior classifier. Especially for NPV, ALIE dominates the other methods. Regarding PPV, ALIE scores at least as high as AL and is always ahead of PL.
Extended results on model selection capabilities of AL and ALIE are presented in Tables (ref)--(ref). These report correct ($\widehat{\mathcal{J}} = \mathcal{J}$) and conservative ($\mathcal{J}\in\widehat{\mathcal{J}}$) selection of the lag pattern, correct model selection ($\widehat{\mathcal{M}} = \mathcal{M}$), as well as log-scale medians of $\lambda_{0,\,\rho^\star}$ and $\Breve{\lambda}_{0,\,\rho^\star}$. The outcomes largely align with our theoretical results in (ref). Discrepancies in medians of the $\lambda_{0\,\rho^\star}$ indicate potential for improved selection of $y_{t-1}$ by ALIE. The finite sample impact of $\Breve{w}_1$ on the selection of the $\Delta y_{t-j}$ is minor under non-stationarity, since $\mathrm{P}(\widehat{\mathcal{J}} = \mathcal{J})$ and $\mathrm{P}(\mathcal{J}\in\widehat{\mathcal{J}})$ are often similar for both procedures. However, ALIE is more likely to select the correct lag structure in stationary regressions in larger samples. Moreover, we find evidence that BIC tuning of ALIE achieves correct model selection and that ALIE often has higher $\mathrm{P}(\widehat{\mathcal{M}} = \mathcal{M})$ than AL.\footnote{Notice that $\lvert\bm \delta_B^\star\rvert>p = k_{\max}$ if $T\in(0, 50)$ due to Schwert's rule so that selection of the true model is infeasible in this setting.}
To investigate the effect of information enrichment when $u_t$ satisfies (ref) but not the stronger condition of (ref), we consider a small simulation study for moving average errors. Although selecting the correct model is infeasible for $T<\infty$ in an ADF($p$) framework, it is interesting to study the implications of using $\Breve{w}_1$ given the empirical relevance of such processes.
We generate $y_t$ as in (ref) with $\varrho=1$ or $\varrho=.95$ and the errors $u_t$ follow a first-order moving average,
We consider $\theta\in\{-.8, -.4, .4, .8\}$ and sample sizes $T\in\{25,50,100,150,250, 500\}$.
The results are shown in (ref). Error processes with large negative MA coefficients are particularly challenging for all methods in non-stationary models, which is reminiscent of the large size distortions reported for standard unit root tests under these conditions, cf. NgPerron2001. Overall, we find the discrepancy in AL and ALIE activation rates for $\rho^\star=0$ to be small. ALIE has some advantages in small samples occasionally, but activation rates of AL improve quickly with $T$ and are slightly superior to ALIE in large samples. Depending on the DGP, the plain Lasso may exhibit a strong tendency to activate $y_{t-1}$ spuriously. Especially for MA coefficients with large magnitude, this issue does not ease and may even exacerbate with $T$.
For $\rho^\star = -.05$, we find that ALIE can be considerably more powerful than AL. An exception is $\theta=-.8$, where both methods perform equally well. Again, we find PL to be inferior due to its inconsistency. High detection rates of stationary models are associated with a tendency to misclassify non-stationary data. PL thus is significantly less reliable than its adaptive competitors.
The recent turmoil in the energy sector due to the Russian invasion of Ukraine and prevailing market frictions from the COVID pandemic have thrust price inflation into the forefront of public and scientific discussions in a way not witnessed in decades. It is broadly recognised that accurate projections of key variables are essential for the success of monetary policy-making. For monetary targeting, inflation forecasts rank among the most critical measures. Workhorse tools for medium-term inflation forecasting include structural macroeconomic models built on Phillips curves. While these models typically entail a forward-looking component summarising survey-based inflation expectations, they also incorporate a stochastic part for sufficiently capturing the dynamics over time. Model selection for this time series component usually includes testing for a unit root and lag truncation.
We apply the penalised estimators considered in this article to German seasonally adjusted consumer price index (CPI) data, performing model selection for energy commodities and all-product inflation rates (CPIenrg2022, CPIall2022). The dataset encompasses 81 quarterly CPI observations from 2001-Q4 to 2021-Q4, beginning with the introduction of the Euro and ending before the markets were shocked by the war in Ukraine.\footnote{We restrict ourselves to this period to mitigate the effects of stylised features of long macroeconomic time series, such as structural change in the first and second moments.} We consider variable selection in the ADF regression
(ref) includes a constant since the ECB monetary policy strives to maintain the inflation rate $\pi_t$ in a corridor around a positive target, implying a stationary process with a non-zero mean.
Besides model selection using the shrinkage estimators, we apply information criteria for lag selection in conjunction with the classical ADF and quasi-difference demeaned ADF tests (DFQD) of Elliottetal1996. For the OLS-estimated models, we consider an exhaustive search for a sparsity pattern using BIC (denoted $\text{BIC}^\star$) as well as an estimation of the truncation lag (by AIC/MAIC). As before, $\lambda$ for PL, AL, and ALIE is tuned using BIC. For ALIE, we use $\Breve{w}_1$ with $J_\alpha$ computed based on $\widehat\omega^2_\textup{AR}(k)$ for $k$ selected by BIC, MAIC, or $k=11$ according to the Schwert1989 rule, which is the maximum lag order ($p_{\max}$) considered across all procedures.
The outcomes are shown in Table (ref). For both inflation rates, the consistent $\text{BIC}^\star$ procedure estimates the lag pattern $\widehat{\mathcal{J}}=\{1, 4, 8\}$ which indicates short-run dependence on the seasonal lags $j=4$ and $j=8$. Reassuringly, the Lasso estimators select the same lag pattern as $\text{BIC}^\star$, except for PL for energy commodities. None of the classical tests reject a non-stationary model at the 5% level for either series. ALIE classifies the all-products inflation rate as stationary for all three specifications of $\widehat\omega_\textup{AR}^2(k)$. It also includes $\pi_{t-1}$ in the energy inflation rate model for LRV estimates based on BIC and $p_\textup{max}$. PL and AL classify both series as non-stationary.
Previous research on penalised estimation has considered single identification principles to elicit data-dependent penalty weights. A prominent example is the adaptive Lasso, for which zou2006adaptive recommends using consistent coefficient estimates. While oracle properties grant asymptotic equivalence between such adaptive procedures and their non-penalised counterparts under appropriate conditions, simulation studies often point to mediocre finite sample performance of ad-hoc implementations.
In this article, we proposed combining two identification principles in computing penalty weights for consistent and oracle-efficient estimation using the adaptive Lasso. With ADF regressions being our use case, we propose to enrich a consistent coefficient estimate with additional information from a (simulation-based) statistic that exploits different stochastic orders of stationary and non-stationary data. We established that the enhanced weight promotes a zero shrinkage property and ensures perpetual activation of the corresponding regressor in stationary models as $T\to\infty$. In particular, we highlighted that beneficial consequences of these features for inference by consistent selection arise from an improved adaptive sorting of the activation knots on the Lasso solution path. We have further extended the theoretical analysis of the asymptotic properties of the adaptive Lasso for the AR(1) models examined in kock2016consistent to general ADF($p$) and ARMA models, even allowing $p\to\infty$ beyond the (approximate) sparsity assumption.
Simulation evidence supports our theoretical results, showing that the information-enriched ALIE estimator dominates the adaptive Lasso and is significantly more reliable than the (possibly inconsistent) plain Lasso regarding both the selection of $y_{t-1}$ and recovering the entire model. Additionally, ALIE improves the classification of stationary and non-stationary models in finite samples under detrending, which appears to be a weakness common to all Lasso variants compared to unpenalised regression-based tests. Abstracting from ADF regressions and as another novelty, we showcase the ability of information enrichment to combine hypothesis tests without type I error accumulation or, under consistent tuning, even with shrinking size. In contrast to most popular combination methods, ALIE does not require standardised evidence, e.g., p-values. Hence, ALIE can be considered a testbed for combining different estimators without establishing asymptotic distributions.
We apply the information-enriched estimator in a model selection exercise for German consumer price data from 2001-Q4 to 2021-Q4. The results indicate that a stationary ADF model is better suited to model the headline inflation rate than a non-stationary autoregression. We also find evidence supporting a stationary model for the German energy commodity inflation rate.
There are various avenues for further research: we consider it relevant to investigate whether the performance can be further improved by alternative specifications of the adaptive weight $\Breve{w}$. A first idea would be to swap the simulation-based statistic $J_\alpha$ with a different statistic. Our theory suggests benefits from adding further weakly correlated estimators to the computation of $\Breve{w}$. Second, since our method focuses exclusively on the weight for $y_{t-1}$, another path would be to elicit improved weights for the lags of $\Delta y_t$ to enhance further the classification rates concerning model selection. In this regard, it would also be worthwhile to investigate adaptive Lasso estimators that include deterministic components, perhaps in a flexible fashion as suggested in tibshirani2011solution and compare them with detrending approaches.
Another aspect yet to be investigated is the benefit of using ALIE for forecast averaging. Due to its superior selection power under consistent tuning, ALIE could help sort out bad predictors better than adaptive Lasso or comparable procedures can by drawing on other forecasts of the same predictor. In more general terms, it would also be instructive to analyse ALIE's performance under conservative model selection and to benchmark against pre-test or averaging estimators as suggested by Hansen2007, Hansen2010. We consider it also promising to extend our approach to multivariate models, e.g., for penalised estimation of the cointegration rank in VAR models as in the framework of LiaoPhillips2015.
Certain violations of our assumptions, e.g., smooth transitions in the (unconditional) variance, have received much attention in the unit root test literature during the last two decades as features of economic time series. Studying the behaviour of our approach in such a framework and making it adaptive would be of interest to a broad range of applications in empirical macroeconomic research. One opportunity is to devise a procedure robust to non-stationary volatility by employing scaled estimators based on non-parametric estimation of the series' variance profile as in Beare2018.
Finally, tuning the adaptive Lasso to detect local alternatives while exhibiting reliable model selection properties is an unresolved problem that is particularly challenging in the context of correlated regressors. We are currently investigating this issue.\\