The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
107,232 characters
Information-Enriched Selection of Stationary and Non-Stationary Autoregressions using the Adaptive LassoWe thank Christoph Hanck for valuable comments that significantly improved the paper. We are indebted to Anders B. Kock for his supportive feedback on several parts of this paper. We also thank Yannick Hoga, Matei Demetrescu, and the Winter 2022/23 Ruhrmetrics Research Seminar participants for their helpful suggestions.
\maketitle
\vspace*{-1em}
\begin{abstract}
\noindent\textit{\textbf{Abstract}}\\[.5ex]
We propose a novel approach to elicit the weight of a potentially non-stationary regressor in the consistent and oracle-efficient estimation of autoregressive models using the adaptive Lasso. The enhanced weight builds on a statistic that exploits distinct orders in probability of the OLS estimator in time series regressions when the degree of integration differs. We develop a new theoretical framework focusing on activation knots to establish results on the benefit of our approach for detecting stationarity when a tuning criterion selects the $\ell_1$ penalty parameter. Monte Carlo evidence shows that our proposal is superior to using OLS-based weights, as suggested by Kock [Econom. Theory, 32, 2016, 243--259]. We apply the modified estimator to model selection for German inflation rates after the introduction of the Euro. The results indicate that energy commodity price inflation and headline inflation are best described by stationary autoregressions.
\\[2ex]
\noindent\textit{\textbf{Keywords:}} Adaptive Lasso, Lasso path, Autoregression, Variable selection, Time series, Unit root testing\newline
\noindent\textit{\textbf{JEL classifications:}} C22, C52, E31
\end{abstract}
\setcounter{footnote}{0}
\newpage
\section{Introduction}
\setcounter{footnote}{0}
Over the last two decades, shrinkage estimators have become increasingly popular in the econometric literature and relevant for economic applications. A substantial body of work considers Lasso-type estimators (\cite{Tibshirani1996}) and establishes conditions for convergence to their unpenalised analogue in the correct model: such \enquote{oracle efficient} procedures select the set of relevant variables with probability converging to 1 and ensure that the coefficient estimators have the identical asymptotic distributions as the unpenalised estimators if the relevant variables were known from the outset (cf. \cite{FanLi2001}).
Recent contributions apply shrinkage estimators to dependent and non-identic\-ally distributed data, which often arise in macroeconomic modelling. For example, \textcite{CanerHan2014} use group bridge estimation to select the correct factor number in approximate factor models. \textcite{MedeirosMendes2012} establish the oracle property for the adaptive Lasso of \textcite{zou2006adaptive} in sparse high-dimensional models for stationary time series. \textcite{KockCallot2015} present oracle inequalities for high-dimensional stationary vector autoregression (VAR) models.
Lasso-type estimators have also been suggested for models with non-stationary variables. \textcite{Leeetal2022} propose the twin adaptive Lasso for oracle-efficient estimation of predictive regressions with mixed root predictors. \textcite{LiaoPhillips2015} devise a consistent adaptive shrinkage procedure that selects the cointegration rank and the lag order in cointegrated VARs. \textcite{Wilms2016} discuss forecasting high-dimensional time series based on $\ell_1$-penalised maximum likelihood estimators for sparse cointegration. \textcite{CanerKnight2013} show that bridge estimators achieve conservative model selection in stationary and non-stationary autoregressions. \textcite{kock2016consistent} proves that the (BIC-tuned) adaptive Lasso is an oracle-efficient estimator in augmented Dickey-Fuller (ADF) regressions. Similarly to the multivariate procedure in \textcite{LiaoPhillips2015}, adaptive Lasso may thus distinguish between stationary and non-stationary data and simultaneously select the correct sparsity pattern in the lags of the dependent variable.
The key feature of adaptive Lasso is an adaptive penalty applied to the individual model coefficients. To implement the adaptive Lasso for ADF models, \textcite{kock2016consistent} suggests using weights based on ordinary least squares (OLS) estimates from an auxiliary ADF regression. This approach aligns with \textcite{zou2006adaptive} who shows that using weights based on consistent coefficient estimates is sufficient for the oracle property. To our knowledge, no contributions in the literature address tailored weight specification for autoregressive (AR) models beyond using OLS estimates. In a high-dimensional cross-section framework, \textcite{Huangetal2008} prove that using marginal regression to obtain weights that need not be consistent for the true coefficients ensures the oracle property. Their simulation results indicate an improved selection performance under liberal tuning using the AIC. However, their asymptotic analysis relies on critical assumptions like partial orthogonality and a variation of the irrepresentable condition for the Lasso, which is not generally satisfied in time series regressions.\footnote{See \textcite{meinshausen2006high} for the original irrepresentable condition of the Lasso.}
If information indicative of a regressor's relevance other than its OLS coefficient estimate is available, incorporating this information into the adaptive weight may yield a Lasso estimator with improved properties. In ADF models, the potentially non-stationary regressor $y_{t-1}$ is of particular interest since its coefficient determines the order of integration. The unit root test literature suggests a variety of identification principles for the latter that could be informative for the weight. Combining two identification principles, we propose an enhanced weight for $y_{t-1}$ in estimating ADF models using the adaptive Lasso. This information-enriched weight implies several properties of the Lasso solution path that are beneficial for (consistent) tuning, especially under stationarity. The resulting estimator is shown to share the oracle property of the adaptive Lasso estimator in \textcite{kock2016consistent} and can be consistently tuned using BIC. Moreover, we prove that the estimator resulting from the modified penalty weight enjoys favourable properties like a relaxed shrinkage when required. Simulation evidence demonstrates that our modification enhances the correct selection rate of stationary and non-stationary models in finite samples, further motivating the use of adaptive Lasso for unit root testing. Our method, therefore, improves upon the estimator in \textcite{kock2016consistent}, particularly in this respect.
The remainder of this article is organised as follows. In \Cref{sec:dnsapr}, we discuss the statistical framework and model selection of ADF regressions using the adaptive Lasso. \Cref{sec:alieirw} introduces the modified adaptive weight and presents several theoretical results on the estimator's asymptotic properties, emphasising the detection of a stationary stochastic component. \Cref{sec:nmm} elaborates on the merit of information enrichment to mitigate spurious selection of an irrelevant variable using the example of a zero-mean model. Simulation studies highlighting the benefit of the modified weight are presented in \Cref{sec:simev}, where we also discuss handling deterministic trends. In \Cref{sec:empapp}, we apply the procedure to model selection for German energy and headline inflation rates. \Cref{sec:conc} concludes and gives an outlook to potential further applications of the information enrichment principle. \Cref{sec:amcr} presents additional simulation results. \Cref{sec:proofs} contains proofs for the theoretical results of \Cref{sec:alieirw,sec:nmm}.
We will use the following conventions throughout the manuscript. $\xrightarrow{p}$ and $\xrightarrow{d}$ denote convergence in probability and distribution, respectively. Coefficients of the true ADF model have a superscript $\star$. We write $\lambda=\Theta(T^\kappa)$ for $0<\liminf_{T\rightarrow\infty}\lambda/T^\kappa\leq\limsup_{T\rightarrow\infty}\lambda/T^\kappa<\infty$, $\kappa\in\mathbb{R}$. Several (estimated) quantities often denoted by Greek symbols depend on the $\ell_1$ penalty and the sample size, which we express by sub-indexing with $\lambda$ and $T$ when important for the exposition. Coefficients $\beta_i^\star$ of the relevant variables identify the data-generating model $\mathcal{M}:=\left\{i:\beta_i^\star\neq0\right\}$. We refer to an estimator $\widehat{\mathcal{M}}$ as model selection consistent if $\lim_{T\to\infty}\ensuremath{\mathrm{P}}\left(\widehat{\mathcal{M}} = \mathcal{M}\right) = 1$ and conservative if $\lim_{T\to\infty}\ensuremath{\mathrm{P}}\left(\widehat{\mathcal{M}} \supseteq \mathcal{M}\right) = 1$ with $\lim_{T\to\infty}\ensuremath{\mathrm{P}}\left(\widehat{\mathcal{M}}\supset\mathcal{M}\right) > 0$, where $\widehat{\mathcal{M}}$ denotes the
estimated model.
\section{Model selection of autoregressions using the adaptive Lasso}
\label{sec:dnsapr}
We consider time series $y_t$ obtained from an autoregressive data-generating process (DGP)
\begin{align}
\label{eq:thedgp}
y_t = \bm z_t'\bm \psi + x_t, \qquad x_t = \varrho x_{t-1} + u_t, \qquad t=1,\dots,T,
\end{align}
where $\bm z_t = (1,t,\dots,t^q)$ gathers deterministic regressors with coefficients $\bm \psi$. Leading cases are intercept-only models where $\bm z_t = 1$ and models with a linear time trend, $\bm z_t = (1, t)'$. The stochastic component $x_t$ is driven by the error process $u_t$, introducing a unit root into the observable variable $y_t$ if $\varrho = 1$ and is weakly stationary if $\lvert\varrho\rvert<1$. We make the following assumptions about $u_t$.
\begin{restatable}[Linear process errors]{assum}{lperrors}
\label{assum:lperrors}
$u_t=\phi(L)\varepsilon_t= \varepsilon_t + \sum_{j=1}^\infty\phi_j\varepsilon_{t-j}$, where the lag polynomial satisfies $\phi(z)\neq0$ for all $\lvert z\rvert\leq1$ and $\sum_{j=1}^\infty j\lvert\phi_j\rvert<\infty$. The innovations $\varepsilon_t$ satisfy $\varepsilon_t\sim i.i.d.\, (0,\sigma^2)$, $\sigma^2<\infty$, and $\operatorname{E}(\varepsilon_t^4)<\infty$.
\end{restatable}
\Cref{assum:lperrors} states that $u_t$ has a linear process representation with \enquote{long-run} variance (LRV) $\omega^2 := \lim_{T\rightarrow\infty} \operatorname{E}[T^{-1}(\sum_{t=1}^T u_t)^2]<\infty$. This is common in the time series literature and covers various serially dependent processes, including stationary ARMA models. The lag polynomial $\phi(L)$ is invertible, meaning that $u_t$ has an AR representation with potentially indefinitely many lags. This ensures the existence of the ADF($\infty$) representation of $y_t$,
\begin{equation}
\Delta y_t = \, d_t + \rho^\star y_{t-1} + \sum_{j=1}^\infty \delta_j^\star \Delta y_{t-j} + \varepsilon_t,
\label{eq:adf_dgp}
\end{equation}
with degree-$q$ deterministic time polynomial $d_t := \bm z_t'\bm \psi$, $\rho^\star\in(-2,0]$ and $\bm \delta^\star := (\delta^\star_1,\dots,\delta^\star_p)'$, where $\sum_{j=1}^p\delta_j^\star<1$. The stochastic component of $y_t$ is stationary if and only if \textit{all} solutions to the characteristic polynomial
\begin{align}
(1-z) - \rho^\star z - \sum_{j=1}^p (1-z)z^j\delta_j^\star = 0 \label{eq:charpolyadf}
\end{align}
are outside the unit circle, viz. $\rho^\star\in(-2,0)$. If $\rho^\star = 0$, then \eqref{eq:charpolyadf} has a unit root. Selecting a model that includes (omits) $y_{t-1}$ thus is equivalent to classifying $y_t$ as stationary (non-stationary). We thus refer to $y_{t-1}$ as the \textit{inference regressor}.
In this paper, we deal with the $\ell_1$-penalised estimation of ADF($p$) models
\begin{align}
\Delta y_t =& \, d_t + \rho_p y_{t-1} + \sum_{j=1}^p \delta_{p,\,j} \Delta y_{t-j} + \varepsilon_{p,\,t},
\label{eq:adfreg}
\end{align}
which either incorporate or approximate the true model \eqref{eq:adf_dgp}, dependent on the choice of $p$ and the DGP's lag order. If the DGP is not sparse, the coefficients $\rho_p$ and $\bm \delta_p$ are flawed with an approximation error, with $\varepsilon_{p,\,t}$ subject to a truncation error inducing autocorrelation. \textcite{ChangPark2002} establish that $p=o(T^{1/3})$ as $T\to\infty$ ensures consistent estimation with the approximation and truncation errors vanishing asymptotically.
This setting is a generalisation of the sparse processes considered in \textcite{kock2016consistent} and aligns with the concept of \enquote{approximate sparsity} in the high-dimensional regression literature, cf. \textcite{Chernozhukov2013}. Most of the results presented in this paper hold under \Cref{assum:lperrors}. We indicate exceptions by invoking the following assumption, which ensures sparsity of the DGP \eqref{eq:thedgp} and hence an ADF representation with $p<\infty$.
\begin{restatable}[Sparsity]{assum}{sparsity}
\label{assum:sparsity}
$u_t = \phi(L)\varepsilon_t$ such that $\phi(L)^{-1}$ exists and is a finite-order AR lag polynomial. The innovations $\varepsilon_t$ satisfy the conditions of \Cref{assum:lperrors}.
\end{restatable}
The adaptive Lasso optimisation problem to \eqref{eq:adfreg} with $d_t=0$ and given the hyperparameter $\lambda\geq0$ governing the $\ell_1$ penalty obtains as
\begin{align}
\Psi_T(\dot\rho,\dot\bm \delta\vert\lambda) := \sum_{t=1}^T \left(\Delta y_t - \dot\rho y_{t-1} - \sum_{j=1}^p \dot\delta_j \Delta y_{t-j} \right)^2 +&\, 2\lambda\left( w_1^{\gamma_1} \lvert\dot\rho\rvert + \sum_{j=1}^p w_{2,\,j}^{\gamma_2} \left\lvert\dot\delta_j\right\rvert\right),
\label{eq:adaptive lassooptim}
\end{align}
with its solution denoted by
\begin{align}
\widehat{\bm \beta}_\lambda := (\widehat{\rho}_\lambda, \widehat\bm \delta_\lambda')' = \operatorname*{arg\,min}_{\dot\rho,\, \dot\bm \delta} \Psi_T(\dot\rho,\dot\bm \delta\vert\lambda).\label{eq:ALestimator}
\end{align}
All adaptive weights $w_1 := \lvert\widehat{\rho}\rvert^{-1}$, $w_{2,\,j} := \lvert\widehat{\delta}_j\rvert^{-1}$ are determined by the initial (OLS) estimates $\widehat{\rho}$ and $\widehat{\delta}_j$ of \eqref{eq:adfreg} and the parameters $\gamma_1>0,\gamma_2>0$ are chosen a priori. The loss function \eqref{eq:adaptive lassooptim} is as in \textcite{kock2016consistent} except that we add a factor of 2 in front of the $\ell_1$-penalty term for mathematical convenience.
Assuming sparsity, \textcite{kock2016consistent} proves oracle efficiency of $\widehat\bm \beta_\lambda$ for a suitable sequence of the tuning parameter $\lambda\in\lambda_{\mathcal{O}}$, with $\lambda_{\mathcal{O}}$ denoting a continuum of all sequences satisfying certain growth rate conditions. He further establishes for $\rho^\star\in(-2,0]$, i.e., irrespective of whether the model is stationary or not, that
\begin{align*}
\lim_{T\to\infty}\ensuremath{\mathrm{P}}\left(\widehat{\mathcal{M}}_{\lambda_\text{BIC}} = \mathcal{M}\right) = 1, \quad\text{with}\quad
\lambda_\text{BIC} := \operatorname*{arg\,min}_\lambda\ \log\left(\frac{\widehat{\bm \varepsilon}_\lambda'\widehat{\bm \varepsilon}_\lambda}{T}\right) + \left\lVert\widehat{\bm \beta}_\lambda\right\rVert_0\frac{\log(T)}{T}
\end{align*}
with the non-zero entries of $\bm \beta^\star := (\rho^\star,\bm \delta^{\star'})'$ defining the true model $\mathcal{M}$ and $\widehat{\bm \varepsilon}_\lambda = \Delta \bm y_t - \widehat\rho_\lambda \bm y_{t-1} - \sum_{j=1}^p \widehat\delta_{j,\,\lambda} \Delta \bm y_{t-j}$.\footnote{$\lVert\bm x\rVert_0 = \sum_{i=1}^p \mathbb{I}(\lvert x_i\rvert>0)$ denotes the Hamming distance of a $p$-vector $\bm x$ and $\mathbb{I}$ the indicator function.} Hence, the BIC-tuned adaptive Lasso estimator $\widehat\bm \beta_{\lambda_\textup{BIC}}$ is model selection consistent. In particular, $\widehat\bm \beta_{\lambda_\textup{BIC}}$ consistently classifies stationary and non-stationary models.
\textcite{kock2016consistent} cautions that consistent selection via the BIC does not imply the estimation error $\widehat{\bm \delta}_{\lambda_\textup{BIC}} - \bm \delta^\star$ to converge to zero, but rather to be bounded.\footnote{We thank Anders B. Kock for drawing our attention to this detail.} Nevertheless, the oracle property establishes the existence of an estimator that is asymptotically equivalent to the unpenalised OLS estimator of the true model $\mathcal{M}$. It thus appears logical that $\lambda_\mathcal{O}$ asymptotically provides the global minimum of the BIC's objective function.
\section{An information-enriched weight}
\label{sec:alieirw}
Due to its oracle properties and the availability of fast algorithms for the optimisation problem underlying the computation like LARS (\cite{Efronetal2004}), the adaptive Lasso is an attractive tool for model selection in \eqref{eq:adfreg}. However, it requires initial estimators to elicit weights for all coefficients. While OLS weights perform reasonably in finite samples, a targeted weight specification may enhance the small-sample performance in discriminating stationary and non-stationary models.\footnote{We thank Anders B. Kock for kindly providing his supplementary simulation results to \textcite{kock2016consistent} upon which we base our Monte Carlo setup.}
It is an open research question of how weights should be chosen in a potentially non-stationary model where it is often of prime interest to make a well-substantiated choice concerning inclusion or omission of $y_{t-1}$, especially for small $T$. In the following, we suggest an alternative weight for $\rho^\star$ that enhances the performance of the adaptive Lasso in distinguishing between stationary and non-stationary autoregressions.
\subsection{Simulation-based weight estimation}
\label{sec:sbwe}
To improve the performance of the adaptive Lasso in discriminating between stationary and non-stationary models, we modify the adaptive weight to exhibit more attractive properties than for $w_1 = \lvert\widehat\rho\rvert^{-1}$. Instead of using only the OLS coefficient estimate, we suggest enriching $w_1$ using a statistic that exploits distinct stochastic orders of stationary and non-stationary $y_t$ to obtain a more favourable adaptive penalty. To this end, we suggest simulation-based computation of a scaling factor to $\widehat{\rho}$ that behaves antithetically to the OLS estimator regarding the limits attained in stationary and non-stationary models. We use the quantile range statistic $J_\alpha$ by \textcite{HerwartzSiedenburg2010}, which is based on distinct convergence rates of OLS estimators in balanced and unbalanced time series regressions.
\begin{restatable}[]{assum}{simrw}
\label{assum:simrw}
$y_t$ has DGP \eqref{eq:thedgp} with $d_t=0$ and $u_t$ satisfies \Cref{assum:lperrors}. $q_t$ is generated by $q_t = q_{t-1} + \nu_t$, $\nu_t\sim i.i.d.\,N(0,\sigma_\nu^2)$, and $q_0=0$. $\{u_t\}_{t=1}^T$ and $\{\nu_t\}_{t=1}^T$ are independent.
\end{restatable}
\textcite{Phillips1986} shows that for $y_t$ and $q_t$ satisfying Assumption \ref{assum:simrw},
\begin{align}
\widehat{\zeta} := \frac{\sum_{t=1}^T q_t y_t}{\sum_{t=1}^T q_t^2} \, \xrightarrow{d}\,\frac{\omega}{\sigma_\nu}\int_0^1 W_w(a)W_y(a)\textup{d}a \biggr/ \int_0^1 W_w(a)^2\textup{d}a,
\label{eq:phillips1986}
\end{align}
where $\widehat\zeta$ is the OLS estimator in the regression $y_t = \zeta q_t + \epsilon_t$.\footnote{Implications of this result are known as \emph{spurious regression}.} $W_w(a)$ and $W_w(a)$ are independent standard Brownian motion. Hence, $\widehat\zeta$ is inconsistent for the true parameter $\zeta = 0$ but converges to a r.v. with a non-standard distribution. If $y_t$ is weakly-stationary, then $\widehat{\zeta}\xrightarrow[]{p}0$ where $\widehat\zeta= O_p(T^{-1})$. Given these characteristics, we suggest the modified adaptive weight $\Breve{w}_1$.
\begin{restatable}[$\Breve{w}_1$]{algo}{wbrewe}
\noindent
\begin{enumerate}
\item Scale $y_t$ using a consistent estimate of the long-run standard deviation $\widehat{\omega}$.
\item Obtain Monte Carlo OLS estimates $\widehat\zeta^{(r)}$, $r=1,\dots,R$ in the regression
\begin{align}
y_t = \zeta^{(r)} q_t^{(r)} + \epsilon_t^{(r)}, \ t=1,\dots,T, \label{eq:simres}
\end{align} with $q_t^{(r)}$ simulated as stated in \Cref{assum:simrw}.
\item Compute the inter-quantile range $J_\alpha := \left\lvert\widehat{\zeta}_{1-\alpha/2}-\widehat{\zeta}_{\alpha/2}\right\rvert$ for some $\alpha\in(0,1)$, where $\widehat{\zeta}_{\alpha}$ denotes the empirical $\alpha$ quantile of $\{\widehat\zeta_r\}_{r=1}^R$.\footnote{One could use some discrete variant of $\bar{J}:= \int_0^1 J_\alpha\mathrm{d}\alpha$. However, this is less favourable for our purpose than choosing a small value for $\alpha$, as we further discuss in Section \ref{sec:alphatuning}.}
\item Compute the modified adaptive weight
\begin{align}
\Breve w_1:=w_1\cdot J_\alpha = \biggr\lvert\frac{\widehat{\rho}}{J_\alpha}\biggr\rvert^{-1}.
\label{eq:mintweights}
\end{align}
\end{enumerate}\qed
\end{restatable}
We refer to $\Breve w_1$ as an information-enriched weight and the resulting Lasso estimator as \emph{adaptive Lasso with information enrichment} (ALIE). For intuition as to why \eqref{eq:mintweights} is useful, note that $J_\alpha= O_p(1)$ for $\rho^\star = 0$. While the asymptotic distribution of $J_\alpha$ has no closed form, simulation results indicate that it is well approximated for small $T$, cf. the left panel in Figure \ref{fig:jvsrho}. The distribution has much probability mass on the interval $(1, \infty)$ and so \eqref{eq:mintweights} has the tendency to upscale the penalty $w_1$ when $y_{t-1}$ is non-stationary. Under stationarity, $J_\alpha$ approaches a degenerate limit at zero at rate $T$, and so \eqref{eq:mintweights} may downscale $w_1$. \Cref{lem:crates} states the asymptotic behaviour of $\Breve{w}_1$.
\begin{restatable}[Stochastic order of $\Breve{w}_1$]{lem}{crates}
\label{lem:crates}
Let $\Breve w_1 := \lvert\widehat{\rho}/J_\alpha\rvert^{-1}$.
\begin{enumerate*}
\item\label{item:lem:crates:H0} If $\rho^\star=0$, the modified weight $\Breve{w}_1$ diverges as $T\rightarrow\infty$ and $\Breve{w}_1^{-1} = O_p(T^{-1})$.
\item\label{item:lem:crates:H1} If $\rho^\star\in(-2,0)$, we have $\Breve{w}_1= O_p(T^{-1})$.
\end{enumerate*}
\end{restatable}
\begin{proof}
See \Cref{sec:proofs}.
\end{proof}
Note that $\Breve w_1$ grows unlimited as $T\to\infty$ in a non-stationary model by part \ref{item:lem:crates:H0} of \cref{lem:crates}, just as $w_1$. However, small-sample properties of the weights differ as is illustrated in the right panel of \Cref{fig:jvsrho}, which shows samples of $\Breve w_1$ and $w_1$ in $\log$-$\log$ space for DGP \eqref{eq:thedgp} with $u_t\sim i.i.d.\,N(0,1)$ and $T=100$. We deduce that the joint distribution for $\rho^\star=0$ (orange dots) has the most probability mass above the dashed bisector line, i.e., $\Breve w_1 > w_1$ on average, implying a higher penalty from setting $\dot\rho\neq0$.
Under stationarity, $w_1-\lvert\rho^\star\rvert^{-1}=O_p(T^{-1/2})$ whereas $\Breve{w}_1\xrightarrow[]{p}0$ at rate $O_p(T^{-1})$ by part \ref{item:lem:crates:H1} of \cref{lem:crates}. The effect is observed in the right panel of \Cref{fig:jvsrho} where samples (green dots) indicate that $\Breve{w}_1<w_1$ for $\rho^\star=-.1$, on average. The faster convergence rate and the zero limit of $\Breve w_1$ are expected to benefit the finite sample performance since a small penalty from setting $\dot\rho\neq0$ renders ALIE less likely to exclude $y_{t-1}$ from the model, misclassifying $y_t$ as non-stationary. We have several remarks.
\begin{figure}[t]
\centering
\caption{Simulated distribution of $J_\alpha$ and joint distributions of regressor weights}
\vspace{.25cm}
\label{fig:jvsrho}
\resizebox{\textwidth}{!}{
\includegraphics{j_vs_rhostar.png}
}
\begin{minipage}{\textwidth}
\vspace{.2cm}
\scriptsize\textit{Notes:} Left: Gaussian Kernel density estimates of simulated null distributions of $J_\alpha$ with $\alpha = .1$ and $R = 150$, bandwidth selected using the \textcite{Silverman1986} rule. Right: Samples of OLS and modified inference regressor weights for $T=100$. The dashed bisector line marks the set of identical weights. 50000 replications based on Gaussian random walks.
\end{minipage}
\end{figure}
\begin{rem}
\label{rem:whyJ}
Other enriching statistics could be used for $\Breve{w}_1$. The class of statistics discussed by \textcite{Stock1999} seems promising. Similarly to $J_\alpha$, these exploit different orders of probability of stationary and non-stationary $y_t$ and obtain by transforming $v_T(\nu) := (\widehat{\omega}^2T)^{-1/2}y_{\lfloor T\nu\rfloor}$, $0\leq\nu\leq1$ by a continuous functional $g(\cdot)$, with $\widehat{\omega}^2$ a LRV estimate.\footnote{Generally, $g: \mathcal{D}[0,1]\rightarrow\mathbb{R}$ with $\mathcal{D}[0,1]$ the Skorokhod space such that the continuous mapping theorem (CMT) can be applied.} For $\rho^\star = 0$ and under suitable conditions, $v_T \xrightarrow[]{d} W_y$ and so $g(v_T)\xrightarrow[]{d} g(W_y)$, while for $\rho^\star\in(-2,0)$, $v_T\xrightarrow[]{p}0$ and thus $g(v_T)\xrightarrow[]{p}g(0)$. Generally, one could work backwards from some asymptotic null distribution suitable for weighting $y_{t-1}$ to a corresponding sample analogue of $g(\cdot)$. However, this is outside the scope of this paper.
\end{rem}
\begin{rem}
\label{rem:Jtuning}
$J_\alpha$ can be modified: $\alpha$ governs the estimated quantile range, thereby determining the distribution's scale and shape. Also, the scale and the location can be controlled via $\sigma_v$. Both parameters impact $\Breve w_1$ and hence the performance of ALIE. This flexibility is beneficial when accommodating a deterministic trend: for $\rho^\star=0$ the limit of $J_\alpha$ involves projections that affect $\Breve{w}_1$. We elaborate on this in \Cref{sec:hdt,sec:alphatuning}.
\end{rem}
We next present results on consistency and consistent tuning of ALIE.
\begin{restatable}[Oracle properties and consistent tuning]{prop}{plugandplim}
\label{prop:plugandplim}
Consider DGP \eqref{eq:thedgp} under \Cref{assum:lperrors}. The oracle properties of the adaptive Lasso in model \eqref{eq:adfreg} as presented in \textcite{kock2016consistent} are retained by ALIE for a sequence of tuning parameters $\lambda_\mathcal{O}$ satisfying conditions (i) $\lambda_\mathcal{O} \big/ \sqrt{T} \to 0$, (ii) $\lambda_\mathcal{O} / T^{1-\gamma_1} \to \infty$ and (iii) $\lambda_\mathcal{O} \big/ T^{(1 - \gamma_2)/2} \to \infty$, where $\gamma_1>1/2$, $\gamma_2>0$. Selecting $\lambda$ by BIC renders ALIE model selection consistent under \Cref{assum:sparsity}.
\end{restatable}
\begin{proof}
See \Cref{sec:proofs}.
\end{proof}
\Cref{prop:plugandplim} establishes the asymptotic (oracle-)equivalence of ALIE and the adaptive Lasso with OLS-based weights of \textcite{kock2016consistent}, henceforth AL, for sparse processes.\footnote{Different lower bounds on $\gamma_1$ and $\gamma_2$ are due to distinct convergence rates of the initial estimates in a non-stationary model where $\widehat{\rho}= O_p(T^{-1})$ and $\widehat{\delta}_j-\delta_j^\star = O_p(T^{-1/2})$.} This implies consistency of ALIE for sequences $\lambda \in \lambda_\mathcal{O}$ satisfying the \textit{same} asymptotic conditions as required for AL: Given $\lambda \in \lambda_\mathcal{O}$, ALIE discriminates consistently between the sets of relevant and irrelevant regressors and the estimated coefficients are asymptotically equivalent to OLS estimates in the true model. Furthermore, ALIE can be tuned using BIC to achieve consistent model selection, giving practical guidance on choosing $\lambda$ in applications.\footnote{The Lasso literature also considers plug-in estimators for $\lambda_\mathcal{O}$, but to our knowledge, no method performs as well as the BIC in our application.}
\subsection{Asymptotics of the adaptive Lasso estimators}
\label{sec:aotale}
In the following, we present several asymptotic results for ALIE regarding consistent classification of $y_{t-1}$, conditional on the regularisation parameter $\lambda$. For ease of exposition and comparability with the results for AL in \textcite{kock2016consistent}, we consider an ADF(0) model and set $\gamma_1 = 1$.
\begin{restatable}[]{thm}{aronelpnull}
\label{thm:aronelpnull}
Let $\Delta y_t = \rho^\star y_{t-1} + u_t$ with $\rho^\star=0$ and $u_t$ satisfies \Cref{assum:lperrors}. Denote $\widehat{\rho}_\lambda$ the minimiser of $\Psi(\dot\rho)=\sum_{t=1}^T(\Delta y_t - \dot\rho y_{t-1})^2 + 2\lambda\lvert\dot\rho\rvert\breve{w}$ with $\breve{w}=\lvert \widehat{\rho}/J_\alpha\rvert^{-1}$, for some $\alpha\in(0,1)$ and denote $\widehat{\rho}$ the OLS estimator of $\rho^\star$.
\begin{enumerate}
\item If $\lambda \rightarrow 0$, then $\lim_{T\rightarrow\infty}\ensuremath{\mathrm{P}}\left(\widehat\rho_\lambda=0\right) = 0$.\label{item:thm:aronelpnull:P1}
\item If $\lambda \rightarrow c\in(0,\infty)$, then $\lim_{T\rightarrow\infty}\ensuremath{\mathrm{P}}\left(\widehat\rho_\lambda=0\right) = F_{\Breve{\mathcal{H}}}(c)>0$ with $F_{\Breve{\mathcal{H}}}$ the CDF of the r.v. $\Breve{\mathcal{H}}$ as defined in the proof of this theorem.\label{item:thm:aronelpnull:P2}
\item If $\lambda \rightarrow \infty$, then $\lim_{T\rightarrow\infty}\ensuremath{\mathrm{P}}\left(\widehat\rho_\lambda=0\right) = 1$.\label{item:thm:aronelpnull:P3}
\end{enumerate}
\end{restatable}
\begin{proof}
See \Cref{sec:proofs}.
\end{proof}
\Cref{thm:aronelpnull} is along the lines of Theorem 3 in \textcite{kock2016consistent} and considers properties in non-stationary models for $T\to\infty$. Part \ref{item:thm:aronelpnull:P1} states that $\lambda\to0$ yields an inconsistent classifier. By part \ref{item:thm:aronelpnull:P3}, $\lambda\to\infty$ is required for ALIE to correctly classify a non-stationary model in the limit. Both results also apply to AL. Part \ref{item:thm:aronelpnull:P2} states that if $\lambda$ converges to some constant $c>0$, then $\widehat{\rho}_\lambda = 0$ with non-zero probability $F_{\Breve{\mathcal{H}}}(c)$.
We next discuss the asymptotic behaviour of ALIE in a stationary model conditional on $\lambda$.
\begin{restatable}[]{thm}{aronelpalt}
\label{thm:aronelpalt}
Let $\Delta y_t = \rho^\star y_{t-1} + u_t$ with $\rho^\star\in(-2,0)$ under \Cref{assum:lperrors}. Denote $\widehat{\rho}_\lambda$ the minimiser of $\Psi(\dot\rho)=\sum_{t=1}^T(\Delta y_t - \dot\rho y_{t-1})^2 + 2\lambda\lvert\dot\rho\rvert\breve{w}_1$ with $\lambda=\Theta(T^{1+\kappa})$, $\kappa\in\mathbb{R}$, $\breve{w}_1=\lvert \widehat{\rho}/J_\alpha\rvert^{-1}$ with $\alpha\in(0,1)$ and $\widehat{\rho}$ the OLS estimator of $\rho^\star$ with $\widehat{\rho}\xrightarrow[]{p}\rho^{\star\star}$.
\begin{enumerate}
\item\label{item:thm:aronelpalt:kappa<1} If $\lambda=\Theta(T^{1+\kappa})$ with $\kappa<1$, then $\lim_{T\rightarrow\infty}\ensuremath{\mathrm{P}}\left(\widehat{\rho}_\lambda=0\right) = 0$.
\item\label{item:thm:aronelpalt:kappa=1} For $\lambda=\Theta(T^{1+\kappa})$ with $\kappa=1$ so that $J\lambda/T\xrightarrow{d}\mathcal{C}$, denote by $c$ a realisation of the r.v. $\mathcal{C}$.
\begin{enumerate}
\item\label{item:thm:aronelpalt:kappa=1lim0} If $c\in\left[0,\,{\rho^{\star\star}}^2\operatorname{E}\left(y_{t-1}^2\right)\right)$, then $\lim_{T\rightarrow\infty}\ensuremath{\mathrm{P}}\left(\widehat{\rho}_\lambda = 0\vert\mathcal{C}=c\right) = 0$.
\item\label{item:thm:aronelpalt:kappa=1limp} If $c = {\rho^{\star\star}}^2\operatorname{E}\left(y_{t-1}^2\right)$, then $\lim_{T\rightarrow\infty}\ensuremath{\mathrm{P}}\left(\widehat{\rho}_\lambda=0\vert\mathcal{C}=c\right) = F(0)$ with $F$ the CDF of the sum of two zero-mean Gaussians as defined in the proof of this theorem.
\item\label{item:thm:aronelpalt:kappa=1lim1} If $c\in\left({\rho^{\star\star}}^2\operatorname{E}\left(y_{t-1}^2\right),\,\infty\right)$, then $\lim_{T\rightarrow\infty}\ensuremath{\mathrm{P}}\left(\widehat{\rho}_\lambda = 0\vert\mathcal{C}=c\right) = 1$.
\end{enumerate}
\item\label{item:thm:aronelpalt:kappa>1} If $\lambda=\Theta(T^{1+\kappa})$ with $\kappa>1$, then $\lim_{T\rightarrow\infty}\ensuremath{\mathrm{P}}\left(\widehat{\rho}_\lambda=0\right)= 1$.
\end{enumerate}
\end{restatable}
\begin{proof}
See \Cref{sec:proofs}.
\end{proof}
\Cref{thm:aronelpalt} is our main result and describes the asymptotic behaviour of $\ensuremath{\mathrm{P}}\left(\widehat{\rho}_\lambda=0\right)$ for ALIE if $\rho^\star\in(-2,0)$, conditional on $\lambda$. Part \ref{item:thm:aronelpalt:kappa<1} states that $\lambda=\Theta(T^{1+\kappa})$ with $\kappa<1$ is required for $\ensuremath{\mathrm{P}}\left(\widehat{\rho}_\lambda = 0\right) = 0$ asymptotically, implying 100\% power. By part \ref{item:thm:aronelpalt:kappa>1}, ALIE will incorrectly classify the model as non-stationary with probability one as $T\to\infty$ if $\kappa>1$. Part \ref{item:thm:aronelpalt:kappa=1} considers $\lambda=\Theta(T^2)$, the knife edge between parts \ref{item:thm:aronelpalt:kappa<1} and \ref{item:thm:aronelpalt:kappa>1}. The classification result then is random. This finding generalises an analogous result in \textcite{kock2016consistent}, who imposes an additional assumption on the mode of convergence of a specific quantity. We prove in \cref{lem:rhostarstar,lem:knifeCLT} in \Cref{sec:proofs} that this assumption is not necessary, resulting in a different distribution of the activation probabilities.
We conclude that ALIE performs consistent classification for $\rho^\star\in(-2,0]$ if $\lambda=\Theta(T^{1+\kappa})$, $\kappa<1$, satisfying part \ref{item:thm:aronelpnull:P3} of \Cref{thm:aronelpnull} ($\lambda\to\infty$ required if $\rho^\star = 0$) and part \ref{item:thm:aronelpalt:kappa<1} of \Cref{thm:aronelpalt} ($\kappa<1$ required if $\rho^\star\in(-2,0))$.
By Theorem 4 in \textcite{kock2016consistent}, AL only has power for $\lambda = O(T)$. \Cref{thm:aronelpalt} relaxes this constraint to $\lambda = O(T^2)$ for ALIE, which is a direct consequence of the properties of $\Breve{w}_1$ stated in part \ref{item:lem:crates:H1} of \cref{lem:crates}.
While the $\lambda$-conditions of \Cref{thm:aronelpalt} are weaker than $\lambda_\mathcal{O} = o(\sqrt{T})$, as required for the oracle property, cf. \Cref{prop:plugandplim}, they suffice for ALIE to consistently select stationary models. If the goal is distinguishing stationary and non-stationary data by classification, the lack of oracle efficiency is negligible. Importantly, \Cref{thm:aronelpalt} foreshadows a power advantage of ALIE over AL in stationary ADF($p$) models under consistent tuning of $\lambda$, and thus also for oracle-efficient tuning using BIC. We will discuss this in the next section.
\subsection{Solution path properties}
\label{sec:ctuning}
The asymptotic analysis for the ADF(0) model presented in the previous section builds on the activation probability of $y_{t-1}$, which is analysed based on the first-order condition (FOC) for $\dot\rho_\lambda = 0$. We next transfer the results of the previous section to general ADF models in terms of the penalty parameter $\lambda$ by examining the Lasso solution path. For this, we use the activation threshold $\lambda_0$.
\begin{restatable}[Activation knots]{defn}{lambda0_definition}
\label{def:lambda0_definition}
Be $x_i\in(y_{t-1}, \Delta y_{t-1}, \dots, \Delta y_{t-p})'$, $\dot\bm \beta := (\dot\rho,\,\dot\bm \delta')'$ and consider the coefficient functions $\dot{\beta}_i(\lambda) \in \mathbb{R}$, $i=1,\dots,p+1$ along the adaptive Lasso solution path to \eqref{eq:adaptive lassooptim}. Define $\lambda_{0,\,\beta_i^\star}$ the earliest activation threshold of $x_i$, i.e.,
\begin{align*}
\lambda_{0,\, \beta_i^\star} :=& \max \Lambda_{0,\, \beta_i^\star}, \\
\Lambda_{0,\, \beta_i^\star} :=& \left\{\lambda : \ \ \partial_{+} \dot\beta_i(\lambda) = 0 \quad \land \quad \partial_{-} \dot\beta_i(\lambda) \neq 0 \quad \land \quad \dot\beta_i(\lambda) = 0\right\},
\end{align*}
where $\partial_{+}\dot\beta_i(\lambda)$ and $\partial_{-}\dot\beta_i(\lambda)$ denote derivatives of $\dot\beta_i(\lambda)$ with respect to $\lambda$ from above and from below, respectively.
\qed
\end{restatable}
For every $\lambda \in \Lambda_{0,\, \beta_i^\star}$, the FOC for $\dot{\beta}_i(\lambda) = 0$ is just binding, i.e., it holds with weak inequality whilst the coefficient $\dot{\beta}_i(\lambda)$ remains set to zero. Following the interpretation in \textcite{Efronetal2004} closely, we declare regressor $x_i$ \textit{active} as soon as the FOC is binding, so that $x_i \in \widehat{\mathcal{M}}_{\lambda_{0,\,\beta_i^\star}}$ even though $\dot{\beta}_{i}(\lambda_{0,\,\beta_i^\star}) = 0$. Note that the FOC for $\dot{\beta}_i(\lambda) = 0$ can also be just binding with a zero coefficient if the regressor is being removed from the active set or is on the verge of a sign reversal.\footnote{A projected sign reversal is the sole reason for a variable to be deactivated by the LARS algorithm with the Lasso modification, see Section 3.1 in \textcite{Efronetal2004}.} By imposing the right-hand derivative to be zero while the left-hand derivative is non-zero, \Cref{def:lambda0_definition} only considers knots on the solution path where variable $i$ is \textit{added} to the active set. Note that the solution path for the ADF(0) model in the previous section features a single \emph{activation} knot $\Lambda_{0,\,\rho^\star} = \lambda_{0,\,\rho^\star}$ since $y_{t-1}$ is the only regressor.
In least angle regression (LAR), the $\Lambda_{0,\, \beta_i^\star}$ are singletons because active variables cannot be deactivated. However, Lasso and its variants allow for repeated activation and deactivation of variables along the solution path.\footnote{See \textcite{Efronetal2004} for details on the LARS algorithm in Lasso mode.} Still, the earliest activation knot of variable $i$, $\lambda_{0,\, \beta_i^\star}$, is the most informative because it is mainly driven by the shrinkage applied to the associated coefficient. \Cref{prop:lambdanaught} describes the asymptotic behaviour of the activation thresholds for the relevant and the irrelevant variables in stationary and non-stationary ADF models.
\begin{restatable}[]{prop}{lambdanaught}
\label{prop:lambdanaught}
Consider model \eqref{eq:adfreg} for $y_t$ generated by \eqref{eq:thedgp} under \Cref{assum:lperrors}. Denote $\lambda_{0,\,\rho^\star}$, $\Breve\lambda_{0,\,\rho^\star}$ and $\lambda_{0,\,\delta_j^\star}$ activation thresholds as defined in \Cref{def:lambda0_definition} for $w_1$, $\Breve{w}_1$, and $w_{2,\,j}$, respectively.
\begin{enumerate*}
\item\label{item:prop:lambdanaught:rho=0} If $\rho^\star=0$, then $\lambda_{0,\,\rho^\star}= O_p(T^{1-\gamma_1})$ and $\Breve\lambda_{0,\, \rho^\star}= O_p(T^{1-\gamma_1})$.
\item\label{item:prop:lambdanaught:rho<0} If $\rho^\star\in(-2,0)$, then $\lambda_{0,\,\rho^\star}=\Theta(T)$ and $\Breve\lambda_{0,\,\rho^\star}=\Theta(T^{1+\gamma_1})$.
\item\label{item:prop:lambdanaught:delta=0} If $\delta_j^\star=0$, then $\lambda_{0,\,\delta_j^\star} = O_p\left(T^{\frac{1-\gamma_2}{2}}\right)$. If $\delta_j^\star\neq0$, then $\lambda_{0,\,\delta_j^\star} =\Theta(T)$.
\end{enumerate*}
\end{restatable}
\begin{proof}
See \Cref{sec:proofs}.
\end{proof}
Part \ref{item:prop:lambdanaught:rho=0} of \Cref{prop:lambdanaught} corroborates an asymptotically equivalent effect of $w_1$ and $\Breve w_1$ when $\rho^\star=0$, cf. \Cref{thm:aronelpnull}. Denote by $\lambda_{0,\,\rho^\star\in(-2,0)}$ the activation threshold for a relevant stationary $y_{t-1}$ for brevity. Part \ref{item:prop:lambdanaught:rho<0} implies that $$\lambda_{0,\,\rho^\star\in(-2,0)} = o\left(\Breve\lambda_{0,\,\rho^\star\in(-2,0)}\right)\quad\forall\quad\gamma_1>0.$$ This is one main feature of information enrichment: $\Breve w_1$ implies $\Breve\lambda_{0,\,\rho^\star\in(-2,0)} =\Theta(T^{1+\gamma_1})$, whereas AL's $w_1$ yields $\lambda_{0,\,\rho^\star\in(-2,0)} =\Theta(T)$ -- just as for the $\lambda_{0,\,\delta_j^\star\neq0}$ of the \emph{relevant} $\Delta y_{t-j}$ with $\sqrt{T}$-consistent weights $w_{2,\,j}$. Part \ref{item:prop:lambdanaught:delta=0} asserts the stochastic orders of the $\lambda_{0,\,\delta_j^\star}$ associated with the relevant and irrelevant $\Delta y_{t-j}$. These orders are the same for ALIE and AL, irrespective of $\rho^\star$.
We conclude from \Cref{prop:lambdanaught} that distinct orders of $w_1$ and $\Breve{w}_1$ influence the selection ability of the adaptive Lasso for $\rho^\star \in (-2,0)$. However, asymptotic performance differences cannot be explained by growth rate considerations for $\rho^\star=0$, where benefits of information enrichment are different, as discussed in the next section. We focus on $\rho^\star \in (-2,0)$ in the remainder.
As for the ADF(0) model considered in \Cref{thm:aronelpalt}, the asymptotic shrinkage applied to $\widehat{\rho}_\lambda$ in an ADF($p$) model depends on the growth rate of the $\lambda$-sequence. ALIE has the virtue that below $\Breve{\lambda}_{0,\, \rho^\star \in (-2,0)}$ the shrinkage on $\widehat{\rho}_\lambda$ vanishes asymptotically, as we assert in \Cref{lem:ZeroShrinkage}.
\begin{restatable}[Zero asymptotic shrinkage]{lem}{ZeroShrinkage}
\label{lem:ZeroShrinkage}
Denote $\widehat{\varepsilon}_{t,\,p,\,\lambda}$ the period-$t$ residual associated with the ALIE estimators $\widehat{\rho}_{\lambda}$ and $\widehat{\delta}_{j,\,\lambda}$ under the conditions of \Cref{prop:lambdanaught} and consider $\rho^\star \in (-2,0)$. For any sequence $\lambda = o\left(\Breve{\lambda}_{0,\, \rho^\star \in (-2,0)}\right)$, the FOC for $\widehat{\rho}_\lambda = 0$ is asymptotically always binding with
\begin{equation*}
\lim_{T \to \infty} \frac{1}{T} \sum_{t=1}^{T} y_{t-1} \widehat{\varepsilon}_{t,\,p,\,\lambda} = 0,
\end{equation*}
amounting to a vanishing shrinkage applied to $\widehat{\rho}_{\lambda}$ as $T\to\infty$.
\end{restatable}
\begin{proof}
See \Cref{sec:proofs}.
\end{proof}
Importantly, zero asymptotic shrinkage is not synonymous with unbiasedness or consistency because $\widehat{\rho}_\lambda$ may be subject to omitted variable bias or an approximation error. Unbiasedness is only guaranteed at $\lambda_\mathcal{O}$ where the shrinkage on the coefficients of the relevant variables vanishes for $T\to\infty$, as required for the oracle property. \Cref{lem:ZeroShrinkage} is useful nonetheless: it transfers the findings for the stationary ADF(0) model in \Cref{thm:aronelpalt} to ADF($p$) models. \Cref{lem:ZeroShrinkage} further implies an asymptotic ordering in every possible Lasso solution path across all permissible stationary linear models, as the subsequent corollary details.
\begin{restatable}[Knot asymptotics]{cor}{KnotOneAsymptotics}
\label{cor:knot1asymptotics}
Be $\lambda^{(m)}$ the $m^{th}$ knot on ALIE's solution path characterised by the knots $\lambda^{(1)} > \lambda^{(2)} > ... > 0$ under the conditions of \Cref{lem:ZeroShrinkage} and $\gamma_1,\gamma_2>0$. If $\rho^\star \in (-2,0)$, then
\begin{enumerate*}
\item\label{item:cor:knot1asymptotics:P1} $\lambda^{(1)} \xrightarrow[]{p} \Breve\lambda_{0,\, \rho^\star}$ and
\item\label{item:cor:knot1asymptotics:P2} $\widehat{\rho}_{\lambda^{(2)}} \xrightarrow[]{p} \rho^{\star\star} \in (-2,0)$.
\end{enumerate*}
\end{restatable}
\begin{proof}
See \Cref{sec:proofs}.
\end{proof}
Part \ref{item:cor:knot1asymptotics:P1} of \Cref{cor:knot1asymptotics} states that ALIE activates $y_{t-1}$ first w.p.~1 in any stationary linear model as $T\to\infty$. This feature is unique to ALIE and has not arisen in any other Lasso variant, to our knowledge. Part \ref{item:cor:knot1asymptotics:P2} asserts that $\widehat\rho_\lambda$ escapes penalisation as some $\Delta y_{t-j}$ enters the active set (with zero coefficient) at $\lambda^{(2)}$, rendering $\widehat\rho_{\lambda^{(2)}}$ equivalent to the unpenalised OLS estimator of $\rho$ in the (possibly underspecified) model $\Delta y_t = \rho y_{t-1} + u_t$, with the non-zero probability limit $\rho^{\star\star}$. Remarkably, although $\rho^{\star \star}$ does not necessarily equal $\rho^\star$, detection of stationarity is asymptotically guaranteed at $\lambda^{(2)}$.\footnote{\Cref{cor:knot1asymptotics} hence motivates another principle to test for a unit root: checking whether $y_{t-1}$ is activated first. This test is consistent but has low power, as simulation evidence suggests.} The results of \cref{cor:knot1asymptotics} are not granted to hold for AL since $\lambda_{0,\,\rho^\star}$ has the same stochastic order as the $\lambda_{0,\,\delta^\star_j}$ of the relevant $\Delta y_{t-j}$. Therefore, the asymptotic activation knot of $y_{t-1}$ is subject to the DGP and thus unknown.
Combining \Cref{prop:plugandplim}, \Cref{lem:ZeroShrinkage} and \Cref{cor:knot1asymptotics}, we obtain a further result on ALIE's solution path.
\begin{restatable}[Perpetual activation]{thm}{PermanentActivation}
\label{thm:permactive}
Under the conditions of \Cref{lem:ZeroShrinkage} and for any sequence $\lambda \in \left[0, \lambda^{(1)}\right]$,
\begin{equation}
\lim_{T \to \infty} \ensuremath{\mathrm{P}}\left(y_{t-1} \in \widehat{\mathcal{M}}_\lambda\right) = 1.
\end{equation}
\end{restatable}
\begin{proof}
See \Cref{sec:proofs}.
\end{proof}
\Cref{thm:permactive} states that once $y_{t-1}$ has been activated moving down the solution path, the probability to encounter a deactivation knot converges asymptotically to zero. This renders ALIE robust to restrictive tuning, a common finite-sample feature of consistent tuning criteria that is also observed for the BIC, cf. \textcite{Efronetal2004}.
\begin{rem}
\label{rem:LARsimilarity}
The property that $y_{t-1}$ is asymptotically never removed from the active set once activated is reminiscent of the LAR solution path to \eqref{eq:adaptive lassooptim}. Still, the solution paths of ALIE and LAR are not necessarily equal since \Cref{thm:permactive} only applies to $y_{t-1}$ in the case of ALIE but to \emph{every} regressor in LAR, since there are no deactivations. Moreover, coefficient sign changes do not generate knots on the LAR solution path, as is the case with the Lasso solution path. For ALIE, \Cref{thm:permactive} requires any possible deactivation knot of $y_{t-1}$ to be simultaneously an activation knot of $y_{t-1}$, turning it into the functional equivalent of a sign change.
\end{rem}
\begin{rem}
\label{rem:countzerorho}
Due to the piece-wise linearity of the solution path $\dot{\rho}_\lambda$ in $\lambda$, sign changes can occur only once between two neighbouring knots that do not affect the activation status of $y_{t-1}$. Consequently, for an asymptotic solution path with $m$ knots $\lambda^{(m)} < \lambda^{(m-1)} < \cdots < \lambda^{(2)}$ unrelated to $y_{t-1}$, there can only be a maximum of $m-1 \geq p$ instances where $\dot{\rho}_\lambda = 0$.\footnote{Referring to an asymptotic solution path, we do not suggest a certain ordering of knots, but rather asymptotics to have kicked in.} However, these spots constitute a countable set of singletons and thus are negligible with respect to the Lebesgue measure.
\end{rem}
\subsection{Properties of ALIE under tuning}
In applications, $\lambda$ is often selected by data-driven methods, e.g., cross-validation or information criteria, usually targeting some $\lambda = O(\lambda_\mathcal{O})$. Even though the asymptotic properties of these tuning methods have been established under various conditions, their stochastic behaviour in small samples needs to be better understood. \textcite{Efronetal2004} remark that consistent tuning criteria like the BIC tend to choose too large values for $\lambda$ in small samples, thereby neglecting relevant variables. To emulate this behaviour, we assume the growth rate of a tuned parameter sequence $\lambda_\textup{IC}$ to vary stochastically in finite samples.
For comparison of $\lambda$-sequences with varying growth rates, we apply the $\log_T$-transformation. Occasions of restrictive tuning as addressed above translate to $\log_T(\lambda_\textup{IC}) > \log_T(\lambda_\mathcal{O})$. Note that in $\log_T$-space, all $\lambda$-sequences are asymptotically bounded and can be considered to be drawn from a continuous distribution function $F_{T,\,\lambda_\textup{IC}}(\log_T \lambda \vert \bm \varepsilon_T, \bm \beta^\star)$, shorthand $F_{T,\,\lambda_\textup{IC}}$. We account for the dependence of $F_{T,\,\lambda_\textup{IC}}$ on the DGP by conditioning on the innovations $\bm \varepsilon_T$ and the vector of true coefficients $\bm \beta^\star$. The distribution function $F_{T,\,\lambda_\textup{IC}}$ has support on $\mathbb{R}^+$ and converges to its limit $F_{\lambda_\textup{IC}}$, inheriting the asymptotic properties of the tuning criterion applied.\footnote{E.g., $F_{\lambda_\textup{BIC}}$ has all probability mass on $\lambda\in\lambda_\mathcal{O}$ asymptotically while $F_{\lambda_\textup{AIC}}$ distributes some probability mass below $\lambda_\mathcal{O}$, see \textcite{kock2016consistent}.} This convergence also acknowledges the fact that $\lambda_\mathcal{O}$ is only guaranteed to exist as $T \to \infty$. Under consistent tuning using the BIC, deviations of $\lambda_\textup{BIC}$ from $\lambda_\mathcal{O}$ can happen as $\log_T \lambda_\textup{BIC} \to \log_T \lambda_\mathcal{O}$. This is reminiscent of the local alternatives regarding coefficients widely used in the unit root literature.
The activation probability of $y_{t-1}$ for Lasso-type estimators is bounded by $F_{T,\,\lambda_\textup{IC}}$,
\begin{align}
\ensuremath{\mathrm{P}}\left(\widehat{\rho}_{\lambda_\text{IC}} \neq 0\right)
=&\, \ensuremath{\mathrm{P}}\left(\widehat{\rho}_{\lambda_\text{IC}} \neq 0 \vert \lambda_\text{IC} < \widetilde{\lambda}_{0,\,\rho^\star}\right) \cdot \ensuremath{\mathrm{P}}\left(\lambda_\text{IC} < \widetilde{\lambda}_{0,\,\rho^\star}\right)\notag\\
=&\, \ensuremath{\mathrm{P}}\left(\widehat{\rho}_{\lambda_\text{IC}} \neq 0 \vert \lambda_\text{IC} < \widetilde{\lambda}_{0,\,\rho^\star}\right) \cdot F_{T,\,\lambda_\text{IC}}\left(\log_T \widetilde{\lambda}_{0,\,\rho^\star}\right), \label{eq:rho_activ_prob}
\end{align}
with generic first activation threshold $\widetilde{\lambda}_{0,\, \rho^\star}$. \Cref{lem:ZeroShrinkage}, \Cref{cor:knot1asymptotics} and \Cref{thm:permactive} point to an advantage for ALIE over AL concerning the conditional activation probability $\ensuremath{\mathrm{P}}\left(\widehat{\rho}_{\lambda_\text{IC}} \neq 0 \vert \lambda_\text{IC} < \widetilde{\lambda}_{0,\,\rho^\star}\right)$. The same applies for $F_{T,\,\lambda_\text{IC}}\left(\log_T \widetilde{\lambda}_{0,\,\rho^\star}\right)$, which is ultimately driven by the different growth rates of $\lambda_{0,\,\rho^\star}$ and $\Breve\lambda_{0,\,\rho^\star}$ (cf. \cref{prop:lambdanaught}).
Maximising $\ensuremath{\mathrm{P}}\left(\widehat{\rho}_{\lambda_\text{IC}} \neq 0\right)$ requires maximising $F_{T, \lambda_\text{IC}}\left(\log_T \widetilde{\lambda}_{0,\,\rho^\star} \right)$ which can either be achieved using another tuning method or by lifting $\widetilde{\lambda}_{0,\,\rho^\star}$. However, a different tuning method may deteriorate classification performance for $\rho^\star=0$. Information enrichment via $\Breve{w}_1$ does not share this disadvantage as it manipulates $\widetilde{\lambda}_{0,\,\rho^\star}$, \textit{conditional on} $\rho^\star$. Since the CDF $F_{T,\,\lambda_\text{IC}}$ is monotonically increasing in $\lambda$, raising the growth rate of $\Breve\lambda_{0,\,\rho^\star \in (-2,0)}$ adds robustness against the finite sample characteristics of $\lambda_\textup{IC}$, a prominent feature in our simulation studies.
\begin{figure}[htbp]
\centering
\caption{Comparison of $\lambda_0$-distributions in ADF models}
\label{fig:lambda0dist}
\vspace{.25cm}
\begin{minipage}{\textwidth}
\resizebox{\textwidth}{!}{
\includegraphics{lambda_0_ridges.png}
}
\end{minipage}
\begin{minipage}{\textwidth}
\resizebox{\textwidth}{!}{
\includegraphics{H0_lambda_0_ridges.png}
}
\begin{minipage}{\textwidth}
\scriptsize\textit{Notes:} DGP \eqref{eq:adfdgp}. $\Breve{w}_1$ is computed using the AR spectral density estimator of \textcite{PerronNg1998} with $k=10$. Ridges show Gaussian kernel density estimates of activation thresholds as in \Cref{def:lambda0_definition} with joint bandwidth estimated according to the \textcite{Silverman1986} rule. Vertical lines show medians. Dashed lines indicate \emph{irrelevant} variables. 5000 replications.
\end{minipage}
\end{minipage}
\end{figure}
\Cref{fig:lambda0dist} illustrates the finite sample impact of the results in \Cref{prop:lambdanaught} and \Cref{cor:knot1asymptotics}. Here, we compare simulated $\lambda_0$ distributions for the regressors in an over-specified ADF model to the BIC-tuned penalty parameter $\lambda_\text{BIC}$ for ALIE and AL. We set $\rho^\star=0$ or $\rho^\star=-.05$, let $\bm \delta^\star=(.275,\,.25,\,.2,\,.15,\,.1)'$ and estimate model \eqref{eq:adfreg} with $p=10$. For ease of interpretation, we adopt the popular choice $\gamma_1 = \gamma_2 = 1$ which translates to $\lambda_0 = O_p(1)$ for \textit{all} irrelevant regressors (dashed lines). Activation thresholds of the relevant regressors (solid lines), i.e., $\lambda_{0,\,\rho^\star\in(-2,0)}$ and $\lambda_{0,\,\delta_j^\star\neq0}$, are $O_p(T)$ and $\Breve\lambda_{0,\,\rho^\star\in(-2,0)} = O_p(T^2)$, by \Cref{prop:lambdanaught}. Note that $F_{T,\,\lambda_\textup{BIC}}$ converges to a stable asymptotic distribution $F_{\lambda_\textup{BIC}}$ with support on $\mathbb{R}^+$.
For $\rho^\star=-.05$ (top panel), the asymptotic implications of \cref{prop:lambdanaught}, part \ref{item:prop:lambdanaught:rho<0} already emerge for $T=75$. The empirical distribution of $\lambda_{0,\,\rho^\star \in (-2,0)}$ is shifted towards larger values, reducing the overlap with the density of $\lambda_\textup{BIC}$. Following \cref{eq:rho_activ_prob}, this feature suggests an improved activation probability from information enrichment. Consistent with the theory, these effects become more pronounced for $T=200$. We also find high congruence between distributions of the remaining variables for ALIE and AL for $T=75$ and $T=200$, corroborating part \ref{item:prop:lambdanaught:delta=0} of \Cref{prop:lambdanaught}.
The bottom panel of \Cref{fig:lambda0dist} shows results for $\rho^\star = 0$. Unlike for $\rho^\star=-.05$, the $\Delta y_{t-j}$ are weakly correlated with $y_{t-1}$, meaning that estimates are expected to be less sensitive to the inclusion or omission of $y_{t-1}$. We hence find the distributions of the $\lambda_{0,\,\delta_j^\star}$ to be virtually identical for ALIE and AL.\footnote{The distributions of $\lambda_{0,\,\delta_j^\star}$ are identical for the irrelevant lags since $\Delta y_{t-j}$ is always stationary.} However, $\Breve\lambda_{0,\,\rho^\star}$ displays a shift of probability mass towards small values of $\Breve\lambda_{0,\,\rho^\star}$, reducing overlap with the density of $ \lambda_\textup{BIC}$. This is a desirable property that, however, does not stem from the asymptotic results detailed in this section. The shape of this distribution matters most in applications because the oracle properties impose zero overlaps between $F_{T,\, \lambda_\textup{BIC}}$ and $\Breve\lambda_{0,\,\rho^\star = 0}$ only as $T \to \infty$. The following section provides further insights by discussing information enrichment in a zero-mean model.
Comparing the medians (vertical lines) of $\lambda_{0,\, \rho^\star}$ for $\rho^\star = -.05$ and $\rho^\star = 0$ reveals how combining two imperfectly correlated statistics translates into an improved performance via projection on the activation thresholds: information enrichment drives an additional wedge between $\Breve\lambda_{0,\,\rho^\star = 0}$ and $\Breve\lambda_{0,\,\rho^\star \in (-2,0)}$, which is captured in the following remark.
\begin{rem}
\label{rem:lambda0_score}
By parts \ref{item:prop:lambdanaught:rho=0} and \ref{item:prop:lambdanaught:rho<0} of \Cref{prop:lambdanaught},
\begin{equation*}
\frac{\Breve\lambda_{0,\,\rho^\star\in(-2,0)}}{\Breve\lambda_{0,\,\rho^\star=0}} = O_p(T^{2\gamma_1})
\quad \textup{and} \quad
\frac{\lambda_{0,\,\rho^\star\in(-2,0)}}{\lambda_{0,\,\rho^\star=0}} = O_p(T^{\gamma_1}),
\end{equation*}
i.e., information enrichment accelerates the separation between the activation thresholds of relevant and irrelevant $y_{t-1}$. Of course, AL could be set up to yield the same discriminatory power as ALIE by doubling $\gamma_1$, thereby putting twice the weight on the OLS estimate $\widehat\rho$ used in $w_1$. However, this is not sensible because the rationale to favour $\Breve{w}_1$ over $w_1$ is that $\widehat\rho$ is deemed insufficient in the first place.\qed
\end{rem}
\subsection{Information enrichment in a zero-mean model}
\label{sec:nmm}
As $\lambda_{0, \rho^\star = 0}$ and $\Breve\lambda_{0, \rho^\star = 0}$ have the same stochastic order, ALIE's size advantage shown in \Cref{fig:lambda0dist} and our simulation results must root from differing shapes of the respective distributions $F_{\lambda_0}$, $F_{\Breve{\lambda}_0}$ for AL and ALIE. Therefore, the relationship between information enrichment and the shape of $F_{\Breve{\lambda}_0}$ will be critical in our considerations below.
Suppose the setup
\begin{align}
\begin{split}
x =&\,\mu + \epsilon, \qquad \epsilon \sim (0, \sigma_\epsilon), \label{eq:nmm} \\
z_{j} =&\, \mu + \nu_{j}, \quad \nu_{j}\sim (0,\sigma_{\nu,\,j}),
\end{split}
\end{align}
with $\sigma_\epsilon,\sigma_{\nu,\,j}<\infty$, $\mu=0$, and statistics $z_j$, $ j=1,\dots,q$, providing additional information on $\mu$. Also, let $\{\epsilon,\nu_1,\cdots,\nu_q\}$ be a set of independent noise terms. The measurements $x$ and $z_j$ can represent single observations or sample averages. By defining $\mu := \rho^\star y_{t-1}$, inference on $\rho^\star$ is tantamount to inference on $\mu$ with the null hypothesis being $\mu = 0$.\\
An adaptive Lasso estimator for $\mu$ in \eqref{eq:nmm} is the minimiser
\begin{equation}
\widehat{\mu}_w := \operatorname*{arg\,min}_{\dot{\mu}} \left( x - \dot{\mu} \right)^2 + 2 \lambda w \lvert \dot{\mu} \rvert, \label{eq:nmmadalasso}
\end{equation}
with the FOC for $\widehat\mu_w = 0$ resulting in
\begin{align}
\lambda w^\gamma \geq \lvert 1 \cdot \left( x - \dot{\mu} \right)\rvert
\Leftrightarrow\lambda \geq w^{-\gamma} \lvert \mu - \dot{\mu} + \epsilon \rvert. \label{eq:nmmadlfoc}
\end{align}
The adaptive Lasso estimator suggested by \textcite{zou2006adaptive} uses exclusively $x$ to determine the penalty weight, $w:=\lvert1/x\rvert^\gamma$. Setting $\dot{\mu} = 0$ and $\mu = 0$ in \eqref{eq:nmmadlfoc}, one obtains $w=\lvert1/\epsilon\rvert^\gamma$ and $\lambda_0:=\lvert\epsilon\rvert^\gamma\lvert\epsilon\rvert$. We suppose the penalty parameter $\lambda_\textup{IC}$ to be drawn from a tuning distribution $F_{\lambda_\textup{IC}}$ independent of $x$ and $z_j$. Because the model \eqref{eq:nmm} has a single regressor, $\ensuremath{\mathrm{P}}\left(\widehat\mu_w\neq0\bigr\vert\lambda_0>\lambda_\textup{IC}\right) = 1$ is ensured. The probability to misclassify $\mu$ simplifies to
\begin{equation*}
\ensuremath{\mathrm{P}}\left(\widehat\mu_{w}\neq0\right) = \ensuremath{\mathrm{P}}\left(\lambda_0>\lambda_\textup{IC}\right).
\end{equation*}
An information-enriched adaptive Lasso estimator $\widehat\mu_{\Breve w}$ adds the $z_{j}$, yielding
\begin{align*}
\Breve{w} := \left\lvert \frac{1}{\epsilon\cdot\prod_{j=1}^q z_{j}} \right\rvert^\gamma, \quad \Breve{\lambda}_0 := \left\lvert\epsilon\cdot\prod_{j=1}^q\nu_{j}\right\rvert^\gamma\lvert\epsilon\rvert,
\end{align*}
as $z_j = \nu_j$ since $\mu = 0$. The enriched estimator has the activation probability
\begin{equation}
\ensuremath{\mathrm{P}}\left(\widehat\mu_{\Breve w}\neq0\right) = \ensuremath{\mathrm{P}}\left(\Breve\lambda_0>\lambda_\textup{IC}\right) = 1 - \ensuremath{\mathrm{P}}\left(\Breve\lambda_0 \leq \lambda_\textup{IC}\right).
\end{equation}
Accordingly, reducing the classification error of $\mu$ requires shrinking $\Breve\lambda_0$, i.e., shifting $F_{\Breve\lambda_0}$ closer to zero. In \Cref{thm:nmmlmtwo}, we examine how the (erroneous) activation probabilities compare for $\widehat\mu_w$ and $\widehat\mu_{\Breve w}$.
\begin{restatable}[Activation probabilities]{thm}{nmmlmtwo}
\label{thm:nmmlmtwo}
Let $\mu=0$ in \eqref{eq:nmm} and consider the adaptive Lasso estimators $\widehat\mu_w$ and $\widehat\mu_{\Breve w}$ with penalty weights $w$ and $\Breve w$, respectively. Denote $\lambda_0\sim F_{\lambda_0}$ and $\Breve{\lambda}_0\sim F_{\Breve{\lambda}_0}$ the thresholds for which the FOC \eqref{eq:nmmadlfoc} is just binding and be $\lambda_\textup{IC}\sim F_{\lambda_\textup{IC}}$ the tuned penalty parameter in \eqref{eq:nmmadalasso}. Let $F_{\lambda_\textup{IC}}$, $F_{\Breve{\lambda}_0}$, and $F_{\lambda_0}$ be continuous distributions that are tight on $\mathbb{R}^+$ and be $F_{\lambda_\textup{IC}}$ exogenously given.
\begin{enumerate}
\item\label{item:thm:nmmlmtwo:P1}
\begin{enumerate}
\item\label{item:thm:nmmlmtwo:P1:a} $\ensuremath{\mathrm{P}}\left(\widehat\mu_w\neq0\right)=1-\int_{0}^\infty F_{\lambda_0}(\lambda_\textup{IC})\,\mathrm{d}F_{\lambda_\textup{IC}}\in[0,1]$.
\item\label{item:thm:nmmlmtwo:P1:b} If $\lim_{q\to\infty}\operatorname{E}(\Breve{\lambda}_0)=0$, then $\lim_{q\to\infty}\ensuremath{\mathrm{P}}\left(\widehat\mu_{\Breve{w}}\neq0\right)=0$.
\end{enumerate}
\item\label{item:thm:nmmlmtwo:P2} Let $A_q := \left\{a\in(0,1): F^{-1}_{\Breve{\lambda}_0}(1-a) < F^{-1}_{\lambda_0}(1-a)\right\}$ for $q\in(0,\infty)$. If $\ensuremath{\mathrm{P}}\left(\widehat\mu_w\neq0\right)=a\in A_q$, then $\ensuremath{\mathrm{P}}\left(\widehat\mu_{\Breve{w}}\neq0\right)\leq \ensuremath{\mathrm{P}}\left(\widehat\mu_w\neq0\right)$.
\end{enumerate}
\end{restatable}
\begin{proof}
See \Cref{sec:proofs}.
\end{proof}
Part \ref{item:thm:nmmlmtwo:P1:a} of \Cref{thm:nmmlmtwo} states how the probability for a \textit{general} adaptive Lasso estimator $\widehat\mu_w$ to activate $\mu$ in model \eqref{eq:nmm} depends on the data $x$ through $\lambda_0$ and the employed tuning procedure, viz. $\ensuremath{\mathrm{P}}\left(\widehat\mu_w\neq0\right)=\ensuremath{\mathrm{P}}\left(\lambda_0>\lambda_\textup{IC}\right)$. Part \ref{item:thm:nmmlmtwo:P1:b} asserts that $\ensuremath{\mathrm{P}}\left(\widehat\mu_{\Breve{w}}\neq0\right)\to0$ as $q\to\infty$ for an information-enriched adaptive Lasso estimator $\widehat\mu_{\Breve w}$, i.e., enriching the adaptive weight $\Breve w$ with additional information on $\mu$ obtains an estimator that dominates any $\widehat\mu_w$ in terms of activation probability.
Part \ref{item:thm:nmmlmtwo:P2} of \Cref{thm:nmmlmtwo} gives a condition for $\ensuremath{\mathrm{P}}\left(\widehat\mu_{\Breve w}\neq0\right)$ to be bounded from above by the activation probability of any adaptive Lasso estimator for some $q\in(0,\infty)$: we require that $\ensuremath{\mathrm{P}}\left(\widehat\mu_w\neq0\right)=a$ corresponds to an upper tail event of $\lambda_0$ that is less likely to occur for $\Breve\lambda_0$. While this statement is independent of $q$, note that $\lim_{q \to \infty} A_q = (0,1)$ is a consequence of part \ref{item:thm:nmmlmtwo:P1:b}. Hence, enriching the information on $\mu$ enlarges the potential of $\widehat\mu_{\Breve w}$ to improve upon $\widehat\mu_w$ as the set of dominating activation probabilities $a$ converges to the entire unit interval.
We next examine the information-enriched adaptive Lasso properties in a \emph{Gaussian} zero mean model.
\begin{restatable}[Gaussian zero means]{cor}{qlimnmm}\label{cor:qlimnmm}
Let the assumptions of \Cref{thm:nmmlmtwo} hold and be $\epsilon\sim N(0,\sigma_\epsilon)$ and $\nu_j\sim N(0,\sigma_{\nu,\,j})$. Consider the information-enriched adaptive Lasso estimator $\widehat\mu_{\breve w}$ with penalty weight $\Breve w$.
\begin{enumerate}
\item\label{item:cor:qlimnmm:P1} If $\prod_{j=1}^q \operatorname{E}(\lvert \nu_j\rvert^\gamma) = o(1)$, then $\lim_{q\rightarrow\infty} \ensuremath{\mathrm{P}}\left(\widehat\mu_{\Breve w}\neq0\right) = 0.$
\item\label{item:cor:qlimnmm:P2} Let $\gamma=1$. If $\prod_{j=1}^q \sigma_{\nu,\,j} = o\left( (\pi/2)^{q/2}\right)$, then $\lim_{q\rightarrow\infty} \ensuremath{\mathrm{P}}\left(\widehat\mu_{\Breve w}\neq0\right) = 0.$
\end{enumerate}
\end{restatable}
\begin{proof}
See \Cref{sec:proofs}.
\end{proof}
Result \ref{item:cor:qlimnmm:P1} of \Cref{cor:qlimnmm} states that in a zero-mean model, $\prod_{j=1}^q \operatorname{E}(\lvert \nu_j\rvert^\gamma) = o(1)$ is sufficient for the probability of $\widehat\mu\neq 0$ to converge to zero as the number of informative statistics $q$ diverges. Since the $\lvert\nu_j\rvert$ are half-Gaussians, this condition is equivalent to restricting the signal-to-noise ratio of the enriching statistics $z_j$ in $\Breve{w}=(1/\prod_{j=1}^q z_j)^\gamma$. Result \ref{item:cor:qlimnmm:P2} of \Cref{cor:qlimnmm} provides a tangible condition for $\lim_{q\rightarrow\infty}\ensuremath{\mathrm{P}}\left(\widehat\mu_{\Breve w}\neq0\right) = 0$, requiring an upper bound on the noise added to $\Breve w$ by the enriching statistics. Note that $\prod_{j=1}^q \sigma_{\nu,\,j} = o\left( (\pi/2)^{q/2}\right)$ is trivially satisfied if $\sigma_{\nu,\,j} = 1 \, \forall j$, i.e., it is permitted to employ a (possibly infinite) set of independent Gaussian statistics to enrich $\Breve{w}$.
\begin{figure}[t]
\centering
\caption{Activation probabilities in univariate Gaussian zero-mean models}
\vspace{.25cm}
\label{fig:actratenmm}
\resizebox{\textwidth}{!}{
\includegraphics{activations.png}
}
\begin{minipage}{\textwidth}
\scriptsize\textit{Notes}: $\gamma=1$. Left: Simulated activation probabilities for $n=1$. Standard Gaussian enriching statistics used in $\widehat\mu_{\Breve{w}}$, $\sigma_\epsilon=\sigma_{\nu,j}=1$. Right: Simulated activation probabilities for $n=50$ with $w=\lvert\overline{\boldsymbol{x}}\rvert^{-1}$ and $\Breve{w}=\lvert t_{z_1}\cdot\overline{\boldsymbol{x}}\rvert^{-1}$ where $t_{z_1}$ is the $t$-statistic for $H_0:\mu=0$ with $\sigma_\epsilon=1,\ \sigma_{\nu}=2,\ \operatorname{corr}(\epsilon, \nu) = .2$. $5\cdot10^4$ replications.
\end{minipage}
\end{figure}
The advantage of $\widehat\mu_{\Breve w}$ over $\widehat\mu_w$ can be seen from the left panel of \Cref{fig:actratenmm}, which compares activation probabilities for $\gamma = 1$ and Gaussian $z_j$. Information enrichment may yield significant improvements, even for $q=1$ (red curve). Including additional information results in even superior estimators. A comparison of the activation probabilities for $n = 50$ and a single correlated enriching statistic is shown in the right panel of \Cref{fig:actratenmm}. Here, $w=\lvert\overline{\boldsymbol{x}}\rvert^{-1}$ and we set $\Breve{w}=\lvert t_{z_1}\cdot\overline{\boldsymbol{x}}\rvert^{-1}$ where $t_{z_1}$ is the $t$-statistic for $H_0:\mu=0$. Again, we find the significant potential of $\Breve{w}_1$ to mitigate the risk of spurious activations.
\section{Implementation and Monte Carlo evidence}
\label{sec:simev}
In the following, we investigate the empirical properties of ALIE in more detail and compare its performance to AL using simulation studies. For implementing $\breve{w}_1$, it is crucial to scale $y_t$ by a consistent estimator of the long-run standard deviation $\omega$ such that the limit distribution of $J_\alpha$ given $\rho^\star = 0$ is nuisance-parameter-free. We follow \textcite{HerwartzSiedenburg2010} and estimate $\omega^2$ using the AR spectral density estimator at frequency zero,
\begin{align}
\widehat{\omega}^2_{\textup{AR}}(k) := \frac{\widehat{\sigma}_k^2}{(1 - \sum_{j=1}^k \widehat{\delta}_j)^2}, \qquad \widehat{\sigma}^2_k := (T-k)^{-1} \sum_{t=k+1}^T \widehat{\varepsilon}_{k,\,t}^2, \label{eq:s2ar}
\end{align}
suggested by \textcite{PerronNg1998}. The $\widehat{\varepsilon}_{k,\,t}$ are residuals from estimating \eqref{eq:adfreg} by OLS with lag order $p=k$. To compute \eqref{eq:s2ar}, we estimate $k$ as $\widehat{k}:= \operatorname*{arg\,min}_{0\leq k\leq k_{\max}} \textup{IC}(k)$ where $\textup{IC}(k)$ is an information criterion of the form
\begin{align}
\textup{IC}(k) := \log(\tilde{\sigma}^2_k) - C_T \frac{\tau_T(k) + k}{T-k_{\max}}.\label{eq:icadf}
\end{align}
$k_{\max}$ is a maximum lag order and $\tilde\sigma^2_k := (T-k_{\max})^{-1}\sum_{t=k_{\max}+1}^T \widehat{\varepsilon}_{k,\,t}^2$ a deviance estimate. $C_T$ is a penalty function and $\tau_T(k)$ a stochastic adjustment term. Popular choices are $C_T = \log(T)$ and $\tau_T(k) = 0$ (BIC), $C_T = 2$ and $\tau_T(k) = 0$ (AIC). \textcite{NgPerron2001} have suggested modified versions for truncating long autoregressions. These are MBIC and MAIC where $\tau_T(k) = \left(\tilde{\sigma}_k^2\right)^{-1}\widehat{\rho}^2\sum_{k_{\max}+1}^T y_{t-1}^2$, respectively.
To estimate the LRV by $\widehat{\omega}^2_{\textup{AR}}(k)$, we use BIC to select $k$ if not indicated otherwise for sparse processes covered by \Cref{assum:sparsity} and MAIC for processes that do not allow consistent model selection when $T<\infty$. The adaptive Lasso estimator \eqref{eq:ALestimator} is computed in a model with $p=k_{\max} = \lfloor12 (T/100)^{1/4}\rfloor$, cf. \textcite{Schwert1989}.
All simulations were performed using the statistical programming language R \parencite{R}. We used the LARS algorithm implemented in the \texttt{lars} package \parencite{pkg-lars} for computing the Lasso solution paths.
\subsection{Rationale for treating \texorpdfstring{$\alpha$}{\textalpha} and \texorpdfstring{$\sigma_v$}{\textsigma} as tuning parameters}
\label{sec:alphatuning}
As outlined in \Cref{rem:whyJ}, $\Breve{w}_1$ can be calibrated via $\alpha$ and $\sigma_v$. The rationale is that larger values of $\alpha$ narrow the distribution of $J_\alpha$ and thus impact the behaviour of $\Breve{w}_1$. By definition of $\Breve{w}_1$, we expect a positive association between $\alpha$ and $\ensuremath{\mathrm{P}}\left(\widehat{\rho}_\lambda\neq0\right)$ for $\rho^\star\in(-2,0]$. A similar argument can be made for $\sigma_v$, which governs the variance of $J_\alpha$.
\begin{figure}[tbp]
\centering
\caption{Information enriched weights --- $\alpha$ and $\sigma_v$ as tuning parameters}
\label{fig:alphatuning}
\vspace{.25cm}
\resizebox{\textwidth}{!}{
\includegraphics{J_alpha_sigma.png}
}
\begin{minipage}{\textwidth}
\vspace{.2cm}
\scriptsize\textit{Notes:} Model \eqref{eq:adfreg} with $d_t = 0$ and $p=12$. DGP \eqref{eq:thedgp} with $u_t \sim i.i.d.\,N(0,1)$ and $T=100$. $\Breve{w}_{1}$ computed with $J_\alpha$ with LRV estimate $\widehat{\omega}^2_{\textup{AR}}(k=12)$. Activation rates for $\sigma_v = 1$ (left) and $\alpha = .1$ (right). 25000 replications.
\end{minipage}
\end{figure}
To illustrate the effect of both parameters on the activation rate of $y_{t-1}$, we simulate $y_t$ using DGP \eqref{eq:thedgp} for $T=100$ and consider $\alpha\in(0, .5)$ for $\sigma_v=1$ and $\alpha=.1$ for $\sigma_v\in(0,2]$. The results are shown in Figure \ref{fig:alphatuning}. Activation rates increase with $\alpha$ so that power gains are accompanied by higher misclassification rates for non-stationary $y_t$ (left panel). Similarly, activation rates increase with $\sigma_v$ (right panel).
Notice that the comparison with AL indicates room for tuning. For example, ALIE could be calibrated to have an activation rate close to that of AL for non-stationary $y_t$ but to exhibit more attractive selection rates under stationarity. Hereafter, we follow \textcite{HerwartzSiedenburg2010} and set $\sigma_v=1$ and $\alpha=.1$ which seem reasonable. However, we note that different choices may be more suitable when the expected loss from a misclassification is not identical for $\rho^\star=0$ or $\rho^\star\in(-2,0)$. This may be relevant in forecasting where the loss from falsely classifying a model as stationary rather than excluding the inference regressor can be quite different than vice-versa.
\subsection{Properties for simple autoregressions}
\label{sec:simpleproc}
To illustrate the effect of $\Breve{w}_1$, we first consider BIC-tuned adaptive Lasso estimates of \eqref{eq:adfreg} with lag order $p=\lfloor12(T/100)^{.25}\rfloor$ for simple processes. For this we simulate $y_t$ according to DGP \eqref{eq:thedgp} without a deterministic component with coefficient $\varrho=\rho^\star+1\in\{.95, 1\}$, and errors $u_t \sim i.i.d.\,N(0,1)$. We consider $T\in\{25,50,100,150,250,500,1000\}$. $\breve{w}_1$ is computed using $J_{.1}$ with the LRV estimated by $\widehat{\omega}_{\text{AR}}^2(0) = \widehat{\sigma}^2_0$.
\begin{table}[tbp]
\setlength{\tabcolsep}{10pt}
\centering
\caption{Properties of adaptive Lasso estimators for AR(1) processes}
\label{tab:simplear}
\vspace{.3cm}
\resizebox{\textwidth}{!}{
\begin{tabular}{lccccccccc}
\toprule
\multicolumn{3}{c}{ } & \multicolumn{3}{c}{AL} & \multicolumn{1}{c}{ } & \multicolumn{3}{c}{ALIE} \\
\cmidrule(l{3pt}r{3pt}){4-6} \cmidrule(l{3pt}r{3pt}){8-10}
$T$ & $p$ & $\rho^\star$ & $\mathrm{P}(\widehat\rho_\lambda\neq0)$ & $\log(w_1)$ & $\log(\lambda_{0,\,\rho^\star})$ & & $\mathrm{P}(\widehat\rho_\lambda\neq0)$ & $\log(\breve{w}_1)$ & $\log(\breve{\lambda}_{0,\,\rho^\star})$\\
\midrule
\addlinespace[-.2cm]
\multicolumn{10}{l}{\textbf{}}\\
25 & $8$ & & $.17$ & $2.14$ & $-.09$ & & $.1$ & $3.78$ & $-1.7$\\
50 & $10$ & & $.06$ & $3.29$ & $-.38$ & & $.05$ & $4.43$ & $-1.41$\\
100 & $12$ & & $.02$ & $4.18$ & $-.46$ & & $.03$ & $5.13$ & $-1.39$\\
150 & $13$ & & $.02$ & $4.6$ & $-.38$ & & $.02$ & $5.56$ & $-1.31$\\
250 & $15$ & & $.01$ & $5.19$ & $-.45$ & & $.02$ & $6.06$ & $-1.29$\\
500 & $22$ & & $.01$ & $5.94$ & $-.48$ & & $.01$ & $6.82$ & $-1.36$\\
1000 & $31$ & \multirow{-7}{*}{ $0$} & $0$ & $6.65$ & $-.51$ & & $.01$ & $7.49$ & $-1.32$\\
\addlinespace[-.2cm]
\multicolumn{10}{l}{\textbf{}}\\
25 & $8$ & & $.18$ & $1.89$ & $.04$ & & $.15$ & $3.01$ & $-1.11$\\
50 & $10$ & & $.1$ & $2.59$ & $.29$ & & $.12$ & $2.76$ & $.09$\\
100 & $12$ & & $.12$ & $2.85$ & $.91$ & & $.17$ & $2.61$ & $1.16$\\
150 & $13$ & & $.21$ & $2.87$ & $1.35$ & & $.28$ & $2.34$ & $1.87$\\
250 & $15$ & & $.41$ & $2.91$ & $1.87$ & & $.56$ & $2.01$ & $2.76$\\
500 & $22$ & & $.89$ & $2.95$ & $2.54$ & & $.96$ & $1.48$ & $4.01$\\
1000 & $31$ & \multirow{-7}{*}{\centering\arraybackslash $-.05$} & $1$ & $2.98$ & $3.23$ & & $1$ & $.86$ & $5.34$\\
\bottomrule
\end{tabular}
}
\begin{minipage}{\textwidth}
\vspace{.25cm}
\scriptsize\textit{Notes:}
Model \eqref{eq:adfreg} with $p=\lfloor12\cdot(T/100)^{.25}\rfloor$ and $d_t=0$. DGP \eqref{eq:adfdgp} with $\varepsilon_t \sim i.i.d.\,N(0,1)$. $\Breve{w}_1$ computed with $J_{.1}$. LRV estimated with $\widehat{\omega}^2_\textup{AR}(0)$. $\ensuremath{\mathrm{P}}\left(\widehat{\rho}_\lambda\neq0\right)$ is the average activation rate of $y_{t-1}$. $\log(\breve w_1)$ and $\log(w_1)$ are log scale median penalty weights. $\log(\lambda_{0,\,\rho^\star})$ and $\log(\Breve{\lambda}_{0,\,\rho^\star})$ are log scale median activation thresholds for $\lambda$. 5000 replications.
\end{minipage}
\end{table}
The results are presented in Table \ref{tab:simplear}. We find that in the case of a random walk $(\rho^\star=0)$, there is improvement in the rate of correct classification of ALIE over AL for $T<150$, where ALIE never exceeds the probability of activating $y_{t-1}$ of AL. This benefit diminishes for large $T$. A comparison of the inference regressor weights shows that $\Breve{w}_1$ implies a higher average penalty from including $y_{t-1}$ in the model. We also find that the average activation threshold on the Lasso solution path decreases faster for ALIE than for AL as $T$ grows.
For stationary autoregressions ($\rho^\star = -.05$), we also find improvement in the discriminatory power of the procedure. For larger $T$, $\Breve{w}_1$ yields a pronounced advantage. E.g., for $T=250$, ALIE has approximately 16\% higher probability of activating $y_{t-1}$ than AL. Finite sample implications of our theoretical results in \Cref{sec:sbwe,sec:dnsapr} are mirrored by the average weights and activation thresholds: while $w_1$ converges at rate $\sqrt{T}$ towards its probability limit $\lvert\rho^\star\rvert^{-1} = 20$, the modified $\Breve{w}_1$ yields a lower adaptive penalty for $T>50$. We also find that $\Breve{\lambda}_{0,\,\rho^\star}$ grows at a faster rate than $\lambda_{0,\,\rho^\star}$.
\subsection{Handling deterministic terms}
\label{sec:hdt}
To our knowledge, no literature has examined the properties of adaptive Lasso in the ADF regression \eqref{eq:adfreg} with $d_t\neq0$ to date. For trend filtering, \textcite{tibshirani2011solution} give a comprehensive overview of past contributions to fitting polynomial trends or structural breaks components using the Lasso and its variants. While these approaches to removing deterministic components appear promising, a theoretical analysis is beyond the scope of this article. We instead investigate the performance of the estimators on detrended data by simulation.\footnote{\textcite{kock2016consistent} does not discuss properties of AL when $d_t\neq0$ and notes that oracle properties carry over to models with detrended data.}
Computation of $\Breve{w}_1$ requires special care in the presence of deterministic terms. For the asymptotic distribution of $J_\alpha$ to be invariant to $\bm \psi\neq0$ one needs to adapt \eqref{eq:simres} and the LRV estimator \eqref{eq:s2ar}, for which OLS or quasi-difference detrending has been suggested, cf. \textcite{PerronQu2007}. Detrending affects the distribution of the OLS estimator $\widehat\zeta$ in \eqref{eq:phillips1986} and hence alters the null distribution of $J_\alpha$. To see this for OLS detrending, note that $W_y$ in \eqref{eq:phillips1986} is replaced by
\begin{align}
W_y^D(a) := W_y(a) - \left(\int_0^1 \bm D(s)\bm D(s)'\text{d}s\right)^{-1}\left(\int_0^1 \bm D(s)W_y(s)\text{d}s\right) \bm D(a),
\end{align}
with $\bm D(a) := \lim_{n\rightarrow\infty}\bm S_n^{-1}\bm z_{[na]}$, $a\in[0,1]$ for a suitable scaling matrix $\bm S_n$. $\bm D(a)=1$ under demeaning where $\bm z_t = 1$ and $\bm D(a)=(1,a)'$ if $\bm z_t = (1,t)'$. Simulation results for $\bm z_t = (1,t)'$ indeed reveal higher kurtosis, accompanied by a shift of the distribution of $J_\alpha$ towards zero. The simulated null distribution of $J_{\alpha}^{\tau}$ under OLS detrending has more probability mass below the threshold of $1$ than the distribution of $J_\alpha$, implying an undesired deflation of $\Breve{w}_1^\tau$ relative to $\Breve{w}_1$ in the (no-)intercept model, see the left panel of Figure \ref{fig:jdetr}.\footnote{This shift is even stronger pronounced for QD detrending where $W_y^d$ is replaced by the corresponding projection of the QD-detrended data, see \textcite{Phillips1998}.} This tendency to yield smaller weights for non-stationary $y_t$ is also seen from a comparison of the cumulative distribution functions (CDFs) of $\Breve{w}_1^\tau$ and $w_{1}$ in the right panel of \ref{fig:jdetr}.
\begin{figure}[t]
\centering
\caption{Distribution of $J_\alpha$ and CDFs of weights for $y_{t-1}$ under trend adjustment}
\vspace{.25cm}
\label{fig:jdetr}
\resizebox{.95\textwidth}{!}{
\includegraphics{J_detrending.png}
}
\begin{minipage}{\textwidth}
\vspace{.2cm}
\scriptsize\textit{Notes:} DGP \eqref{eq:thedgp} with $u_t\sim i.i.d.\,N(0,1)$. $T = 100$. Left: Gaussian kernel density estimates. $J$ computed without deterministic components ($J_\alpha$), for OLS-detrended data ($J_\alpha^\tau$), and with simulation-based trend adjustment ($J_{\alpha,\,\sigma_v}^{\tau^\star}$) with $\sigma_v = .75$. $\alpha = .1$ and LRV estimated by $\widehat{\omega}^2_{\textup{AR}}(k = 0) = \widehat{\sigma}_0^2$ throughout. Bandwidth selected using the \textcite{Silverman1986} rule. Right: empirical CDFs of inference regressor weights. $10^5$ replications based on Gaussian random walks.
\end{minipage}
\end{figure}
We thus expect OLS detrending to inflate activations for $\rho^\star\in(-2,0]$.
We propose the trend-adjusted simulated weight $\Breve{w}_{1}^{\tau^\star}$ based on a modified trend-adjusted statistic $J_{\alpha,\,\sigma_v}^{\tau^\star}$.
\begin{restatable}[$\Breve{w}_1^{\tau^\star}$]{algo}{wbrewedet}
\label{algo:wbrewedet}
\noindent
\begin{enumerate}
\item Estimate the LRV using model \eqref{eq:adfreg} with $\bm z_t = (1,\,t)'$ by
\begin{align*}
\widehat{\omega}^2_{\textup{AR}}(k) =&\, \frac{\widehat{\sigma}_k^2}{(1 - \sum_{j=1}^k \widehat{\delta}_j)^2}, \quad \widehat{\sigma}^2_k = (T-k-2)^{-1} \sum_{t=k+1}^T \widehat{\varepsilon}_{k,\,t}^2
\end{align*}
for suitable $k$.
\item Scale $y_t$ by $\widehat{\omega}_{\textup{AR}}(k)^{-1}$.
\item Obtain the range statistic $J_{\alpha,\,\sigma_v}^{\tau^\star} := \left\lvert\widehat{\zeta}^{(r)}_{1-\alpha/2}-\widehat{\zeta}^{(r)}_{\alpha/2}\right\rvert$ for simulated OLS estimates $\widehat\zeta^{(r)}$, in the model $y_t = \mu^{(r)} + \varsigma^{(r)} t + \zeta^{(r)} x_t^{(r)} + \nu_t^{(r)}$ with $x_t^{(r)}$ satisfying \Cref{assum:simrw} and $r=1,\dots,R$.
\item Compute $\Breve{w}_{1}^{\tau^\star} := \left\lvert\widehat{\rho} / J_{\alpha,\,\sigma_v}^{\tau^\star}\right\rvert$ with $\widehat{\rho}$ the OLS estimator of $\rho^\star$ in model \eqref{eq:adfreg} with $d_t$ as in step 1.\qed
\end{enumerate}
\end{restatable}
The procedure is similar to \ref{sec:sbwe} with the difference that we incorporate a linear time trend in the regressions underlying the LRV estimation and the simulation. Following the results in section \ref{sec:alphatuning}, we suggest simulating random walks with $\sigma_v\leq1$ in step 3 to attenuate spurious activations of $y_t$. Simulation results for $J_{\alpha,\,\sigma_v}^{\tau^\star}$ and the associated $\Breve{w}_{1}^{\tau^\star}$ in Figure \ref{fig:jdetr} indeed reveal null distributions with more favourable properties than for the OLS-detrended variants $J^\tau_\alpha$ and $w^\tau_1$.\footnote{We use the ad-hoc choice $\sigma_v = .75$ for better size control of $J_{\alpha,\,\sigma_v}^{\tau^\star}$ for small $T$.}
\subsection{Sparse processes}
\label{sec:adfdgp}
We next investigate finite sample properties of ALIE and AL for sparse processes and also evaluate the ability of the (inconsistent) plain Lasso (PL) to distinguish between stationary and non-stationary models. Data are simulated with the ADF-DGP
\begin{equation}
y_t(1-B(L)) = v_t, \qquad v_t\overset{i.i.d.}{\sim}N(0,1), \qquad t=1,\dots,T, \label{eq:adfdgp}
\end{equation}
with
\begin{align*}
B(L)=\,(\rho^\star + \delta_1^\star) L+ (\delta_2^\star - \delta_1^\star) L^2 + \dots + (\delta^\star_{k^\star-1}-\delta^\star_{k^\star}) L^{k^\star}- \delta^\star_{k^\star} L^{k^\star+1},
\end{align*}
and $k^\star+1$ zero initial values where the lag order is $k^\star := \dim(\bm \delta^\star)$. To investigate the model selection performance when there is sparsity in the coefficients of the $\Delta y_{t-j}$, we set $\bm \delta_A^\star := (.4, .3, .2)'$ and $\bm \delta_B^\star := (.4, .3, .2, 0, 0, 0, -.2, 0, 0, .2)'$. $\bm \delta_C^\star := .7$ yields a short AR process. $\bm \delta_D^\star:= (-.4, 0, .7)'$ is taken from \textcite{CanerKnight2013} and is a low power setting for unit root tests. We compute Lasso estimates of model \eqref{eq:adfreg} with lag order $p=\lfloor12\cdot(T/100)^{.25}\rfloor$ and tune $\lambda$ using BIC. We consider $T\in\{25,50,100,150,200,250\}$, let $\rho^\star=0$ for non-stationarity and $\rho^\star = -.05$ to investigate the stationary case. For $\Breve w_1$ we estimate the LRV by $\widehat\omega^2_\text{AR}(k)$ and compute $J_\alpha$ as outlined in Sections \ref{sec:sbwe} and \ref{sec:hdt}.
Accommodating linear trends is non-trivial. While model selection consistency should be possible for detrended data, unreported simulations indicate that OLS detrending leads to prohibitive activation rates of $y_{t-1}$ for the DGPs under consideration when $\rho^\star=0$. To provide some guidance for empirical application, we present results for first-difference (FD) detrending (\cite{SchmidtPhillips1992}), which seems more reliable for sparse processes.\FloatBarrier
\begin{figure}
\centering
\caption{Activation rates of Lasso estimators under sparsity}
\label{fig:sp_rates}
\vspace{.25cm}
\resizebox{.97\textwidth}{!}{
\includegraphics{sp_fig.png}
}
\begin{minipage}{\textwidth}
\vspace{.2cm}
\scriptsize\textit{Notes:} DGP \eqref{eq:adfdgp}. $\bm \delta_A^\star := (.4, .3, .2)'$, $\bm \delta_B^\star := (.4, .3, .2, 0, 0, 0, -.2, 0, 0, .2)'$, $\bm \delta_C^\star := .7$, $\bm \delta_D^\star:= (-.4, 0, .7)'$. Model \eqref{eq:adfreg} with $p=\lfloor 12 \cdot (T/100)^{.25}\rfloor$. AL and ALIE are adaptive Lasso estimators based on $w_1$ and $\Breve{w}_1$, respectively. PL is the plain Lasso. Top panel: no adjustment. Middle panel: FD-demeaned data. Bottom panel: FD-detrended data. $\Breve{w_1}$ computed as discussed in \Cref{sec:sbwe,sec:hdt} with LRV estimate $\widehat\omega^2_{\textup{AR}}(k)$. $k$ selected by BIC with $k_{\max} = p$. 5000 Monte Carlo replications.
\end{minipage}
\end{figure}
\begin{table}[t]
\centering
\caption{Classification metrics of Lasso estimators for sparse processes}
\label{tab:cmetrics}
\vspace{.25cm}
\setlength{\tabcolsep}{12pt}
\resizebox{\textwidth}{!}{
\begin{tabular}{lcd{3.2}d{3.2}cd{3.2}d{3.2}cd{3.2}d{3.2}}
\toprule
\multicolumn{1}{c}{ } & \multicolumn{1}{c}{ } & \multicolumn{2}{c}{PL} & \multicolumn{1}{c}{ } & \multicolumn{2}{c}{AL} & \multicolumn{1}{c}{ } & \multicolumn{2}{c}{ALIE} \\
\cmidrule(l{3pt}r{3pt}){3-4} \cmidrule(l{3pt}r{3pt}){6-7} \cmidrule(l{3pt}r{3pt}){9-10}
& $\bm \delta^\star$ & \multicolumn{1}{c}{PPV} & \multicolumn{1}{c}{NPV} & & \multicolumn{1}{c}{PPV} & \multicolumn{1}{c}{NPV} & & \multicolumn{1}{c}{PPV} & \multicolumn{1}{c}{NPV}\\
\midrule
\addlinespace[.2cm]
\multicolumn{10}{l}{no adjustment}\\
\addlinespace[.2cm] & $\bm \delta_A^\star$ & .79 & .74 & \ & .90 & .82 & \ & .94 & .93\\
& $\bm \delta_B^\star$ & .73 & .67 & \ & .86 & .73 & \ & .94 & .87\\
& $\bm \delta_C^\star$ & .90 & .72 & \ & .95 & .71 & \ & .95 & .84\\
& $\bm \delta_D^\star$ & .79 & .78 & \ & .85 & .56 & \ & .88 & .62\\
\addlinespace[.2cm]
\multicolumn{10}{l}{constant}\\
\addlinespace[.2cm] & $\bm \delta_A^\star$ & .79 & .73 & \ & .89 & .81 & \ & .93 & .92\\
& $\bm \delta_B^\star$ & .73 & .66 & \ & .87 & .72 & \ & .94 & .83\\
& $\bm \delta_C^\star$ & .90 & .71 & \ & .94 & .70 & \ & .96 & .80\\
& $\bm \delta_D^\star$ & .79 & .76 & \ & .85 & .56 & \ & .88 & .60\\
\addlinespace[.2cm]
\multicolumn{10}{l}{linear trend}\\
\addlinespace[.2cm] & $\bm \delta_A^\star$ & .67 & .63 & \ & .78 & .66 & \ & .83 & .78\\
& $\bm \delta_B^\star$ & .68 & .62 & \ & .76 & .64 & \ & .83 & .70\\
& $\bm \delta_C^\star$ & .76 & .66 & \ & .83 & .64 & \ & .85 & .68\\
& $\bm \delta_D^\star$ & .59 & .70 & \ & .67 & .55 & \ & .68 & .56\\
\bottomrule
\end{tabular}
}
\begin{minipage}{\textwidth}
\vspace{.2cm}
\scriptsize\textit{Notes:} See \Cref{fig:sp_rates} for the DGP and the model. $T=100$. PPV and NPV denote the positive and negative predictive values. Adjustments are done using FD-detrending.
\end{minipage}
\end{table}
\Cref{fig:sp_rates} shows activation rates of $y_{t-1}$ without deterministic components (top panel), for FD demeaning (middle panel) and for FD detrending (bottom panel). Spurious activation rates for $\rho^\star = 0$ (first sub-panel) indicate deficiencies for AL in small samples and across most settings for $\bm \delta^\star$. For instance, at $T=25$ observations, the activation probability of $y_{t-1}$ in settings $\bm \delta_A^\star$ and $\bm \delta_B^\star$ exceeds 40\%, whereas ALIE achieves considerably lower rates. Similar outcomes are found for demeaned and detrended data. Although the discrepancy between ALIE and AL diminishes in $T$, these results indicate an advantage of $\Breve{w}_1$ over $w_1$ in finite samples.
Results for $\rho^\star=-.05$ (second sub-panel) corroborate favourable finite sample implications of our theoretical results in \Cref{sec:ctuning} under stationary. ALIE's activation rates for $y_{t-1}$ are above AL's across most settings. For example, ALIE has power advantages of up to 20\% for sample sizes between $T=100$ and $T=150$ in setting $\bm \delta_B^\star$. Outcomes are qualitatively similar when adjusting for deterministic components. However, all methods suffer from reduced power under detrending, where gains from information enrichment are smaller. This is most pronounced in scenario $\bm \delta_D^\star$ where none of the methods dominates.
\Cref{fig:sp_rates} also indicates the mediocre performance of the BIC-tuned plain Lasso in most settings: PL often cannot compete with AL, let alone ALIE in power, despite comparable activation rates under non-stationarity. At worst, the inconsistency of the plain Lasso entails erratic behaviour, cf. setting $\bm \delta_B^\star$ for $\rho^\star = 0$ where rejection rates display upward surges or plateau as $T$ increases.
Further insight into the methods' suitability as classifiers for stationary and non-stationary models is given in \Cref{tab:cmetrics}. Here, we report the positive predictive value (PPV) and the negative predicted value (NPV),
\begin{align*}
\textup{PPV} = \frac{\sum_{s\in\mathcal{S}} \mathbb{I}\left(\widehat{\rho}^{(s)}_\lambda\neq0 \big\vert \rho^\star\in(-2,0)\right)}{\sum_{s\in\mathcal{S}} \mathbb{I}\left(\widehat{\rho}^{(s)}_\lambda\neq0 \big\vert \rho^\star\in(-2,0]\right)}, \quad \textup{NPV} = \frac{\sum_{s\in\mathcal{S}} \mathbb{I}\left(\widehat{\rho}^{(s)}_\lambda=0 \big\vert \rho^\star=0\right)}{\sum_{s\in\mathcal{S}} \mathbb{I}\left(\widehat{\rho}^{(s)}_\lambda=0 \big\vert \rho^\star\in(-2,0]\right)},
\end{align*}
based on pooled simulation results $\mathcal{S}$ for $\rho^\star\in(-2,0]$ for $T=100$.\footnote{We consider $\widehat{\rho}_\lambda<0$ a \enquote{positive}.} The comparison indicates that ALIE is the superior classifier. Especially for NPV, ALIE dominates the other methods. Regarding PPV, ALIE scores at least as high as AL and is always ahead of PL.
Extended results on model selection capabilities of AL and ALIE are presented in Tables \ref{tab:sp_c_extended_size}--\ref{tab:sp_ct_extended_power}. These report correct ($\widehat{\mathcal{J}} = \mathcal{J}$) and conservative ($\mathcal{J}\in\widehat{\mathcal{J}}$) selection of the lag pattern, correct model selection ($\widehat{\mathcal{M}} = \mathcal{M}$), as well as log-scale medians of $\lambda_{0,\,\rho^\star}$ and $\Breve{\lambda}_{0,\,\rho^\star}$. The outcomes largely align with our theoretical results in \Cref{sec:ctuning}. Discrepancies in medians of the $\lambda_{0\,\rho^\star}$ indicate potential for improved selection of $y_{t-1}$ by ALIE. The finite sample impact of $\Breve{w}_1$ on the selection of the $\Delta y_{t-j}$ is minor under non-stationarity, since $\mathrm{P}(\widehat{\mathcal{J}} = \mathcal{J})$ and $\mathrm{P}(\mathcal{J}\in\widehat{\mathcal{J}})$ are often similar for both procedures. However, ALIE is more likely to select the correct lag structure in stationary regressions in larger samples. Moreover, we find evidence that BIC tuning of ALIE achieves correct model selection and that ALIE often has higher $\mathrm{P}(\widehat{\mathcal{M}} = \mathcal{M})$ than AL.\footnote{Notice that $\lvert\bm \delta_B^\star\rvert>p = k_{\max}$ if $T\in(0, 50)$ due to Schwert's rule so that selection of the true model is infeasible in this setting.}
\subsection{ARMA processes}
\label{sec:asparsity}
\begin{figure}[htbp]
\centering
\caption{Activation rates of Lasso estimators for MA errors}
\label{fig:arimarates}
\vspace{.25cm}
\resizebox{\textwidth}{!}{
\includegraphics{ARMA_fig.png}
}
\begin{minipage}{\textwidth}
\vspace{.2cm}
\scriptsize\textit{Notes:} DGP \eqref{eq:thedgp} with $\bm \psi=0$ and $u_t$ as in \eqref{eq:maerrors}. Model \eqref{eq:adfreg} with $d_t=0$ and $p=\lfloor 12 \cdot (T/100)^{.25}\rfloor$. AL and ALIE are adaptive Lasso estimators based on $w_1$ and $\Breve{w}_1$, respectively. PL is the plain Lasso. $J_{\alpha}$ with $\alpha=.1$ computed with LRV estimate $\widehat\omega^2_{\textup{AR}}(k)$, $k$ selected by MAIC with $k_{\max} = p$. 5000 Monte Carlo replications.
\end{minipage}
\end{figure}
To investigate the effect of information enrichment when $u_t$ satisfies \Cref{assum:lperrors} but not the stronger condition of \Cref{assum:sparsity}, we consider a small simulation study for moving average errors. Although selecting the correct model is infeasible for $T<\infty$ in an ADF($p$) framework, it is interesting to study the implications of using $\Breve{w}_1$ given the empirical relevance of such processes.
We generate $y_t$ as in \eqref{eq:thedgp} with $\varrho=1$ or $\varrho=.95$ and the errors $u_t$ follow a first-order moving average,
\begin{align}
u_t = \varepsilon_t + \theta \varepsilon_{t-1}, \qquad \varepsilon_t\overset{i.i.d.}{\sim}N(0,1).
\label{eq:maerrors}
\end{align}
We consider $\theta\in\{-.8, -.4, .4, .8\}$ and sample sizes $T\in\{25,50,100,150,250, 500\}$.
The results are shown in \Cref{fig:arimarates}. Error processes with large negative MA coefficients are particularly challenging for all methods in non-stationary models, which is reminiscent of the large size distortions reported for standard unit root tests under these conditions, cf. \textcite{NgPerron2001}. Overall, we find the discrepancy in AL and ALIE activation rates for $\rho^\star=0$ to be small. ALIE has some advantages in small samples occasionally, but activation rates of AL improve quickly with $T$ and are slightly superior to ALIE in large samples. Depending on the DGP, the plain Lasso may exhibit a strong tendency to activate $y_{t-1}$ spuriously. Especially for MA coefficients with large magnitude, this issue does not ease and may even exacerbate with $T$.
For $\rho^\star = -.05$, we find that ALIE can be considerably more powerful than AL. An exception is $\theta=-.8$, where both methods perform equally well. Again, we find PL to be inferior due to its inconsistency. High detection rates of stationary models are associated with a tendency to misclassify non-stationary data. PL thus is significantly less reliable than its adaptive competitors.
\section{Empirical application}
\label{sec:empapp}
The recent turmoil in the energy sector due to the Russian invasion of Ukraine and prevailing market frictions from the COVID pandemic have thrust price inflation into the forefront of public and scientific discussions in a way not witnessed in decades. It is broadly recognised that accurate projections of key variables are essential for the success of monetary policy-making. For monetary targeting, inflation forecasts rank among the most critical measures. Workhorse tools for medium-term inflation forecasting include structural macroeconomic models built on Phillips curves. While these models typically entail a forward-looking component summarising survey-based inflation expectations, they also incorporate a stochastic part for sufficiently capturing the dynamics over time. Model selection for this time series component usually includes testing for a unit root and lag truncation.
\begin{figure}[htbp]
\centering
\caption{Quarterly German consumer price indices and inflation rates}
\vspace{.25cm}
\label{fig:cpiinf}
\resizebox{\textwidth}{!}{
\includegraphics{ea_plot.png}
}
\begin{minipage}{\textwidth}
\vspace{.25cm}
\scriptsize\textit{Notes:} Left: seasonally adjusted quarterly German consumer price indices for all products (\cite{CPIall2022}) and energy commodities (\cite{CPIenrg2022}) including fuel, electricity, and gasoline from 2001-Q4 to 2021-Q4. The base year is 2015. Right: quarterly year-on-year inflation rates.
\end{minipage}
\end{figure}
\begin{table}[htbp]
\centering
\caption{Model selection outcomes for German price inflation rates}
\label{tab:infoutcomes}
\vspace{.25cm}
\setlength{\tabcolsep}{6pt}
\resizebox{\textwidth}{!}{
\begin{tabular}{lcc d{3.2} cccc d{3.2} cc}
\toprule
& \multicolumn{4}{c}{all products} & & \multicolumn{4}{c}{energy commodities} &\\
\cmidrule(l{0pt}r{0pt}){2-5} \cmidrule(l{0pt}r{0pt}){7-10}
procedure & IC & lags & $t$ & $\pi_{t-1}$ & & IC & lags & $t$ & $\pi_{t-1}$ &\\
\hline
\multirow{2}{*}{ADF} & $\text{BIC}^\star$ & 1, 4, 8 & -.47 & exclude & & $\text{BIC}^\star$ & 1, 4, 8 & -.82 & exclude &\\
& AIC & 1--8 & -1.78 & exclude & & AIC & 1--11 & -2.61 & exclude &\\
\\
\multirow{2}{*}{DFQD} & $\text{BIC}^\star$ & 1, 4, 8 & -.66 & exclude & & $\text{BIC}^\star$ & 1, 4, 8 & -.66 & exclude &\\
& MAIC & -- & -1.41 & exclude & & MAIC & 1--8 & -1.40 & exclude &\\
\\
PL & & 1, 4, 8 & & exclude & & & 1, 4, 6, 8 & & exclude &\\
AL & & 1, 4, 8 & & exclude & & & 1, 4, 8 & & exclude &\\
ALIE & BIC ($J_\alpha$) & 1, 4, 8 & & \textit{\textbf{include}} & & BIC ($J_\alpha$) & 1, 4, 8 & & \textit{\textbf{include}} &\\
ALIE & $p_\textup{max}$ ($J_\alpha$) & 1, 4, 8 & & \textit{\textbf{include}} & & $p_\textup{max}$ ($J_\alpha$) & 1, 4, 8 & & \textit{\textbf{include}} &\\
ALIE & MAIC ($J_\alpha$) & 1, 4, 8 & & \textit{\textbf{include}} & & MAIC ($J_\alpha$) & 1, 4, 8 & & exclude &\\
\bottomrule
\addlinespace[.3cm]
\end{tabular}
}
\begin{minipage}{\textwidth}
\vspace{.25cm}
\scriptsize\textit{Notes:} Model selection results for \eqref{eq:infreg}. Lasso estimators are computed using FD-demeaned data and for $p=p_\textup{max}=\lfloor12\cdot(100/T)^{.25}\rfloor=11$ with $\lambda$ tuned by BIC. IC reports the criterion for selecting the lag pattern/truncation lag of the unpenalised procedures. $\text{BIC}^\star$ denotes exhaustive lag pattern selection with BIC. For ALIE, we report the criterion used in the LRV estimation for $J_\alpha$. $5\%$ critical values used for ADF and DFQD are $-2.89$ and $-1.95$, respectively.
\end{minipage}
\end{table}
We apply the penalised estimators considered in this article to German seasonally adjusted consumer price index (CPI) data, performing model selection for energy commodities and all-product inflation rates (\cite{CPIenrg2022, CPIall2022}). The dataset encompasses 81 quarterly CPI observations from 2001-Q4 to 2021-Q4, beginning with the introduction of the Euro and ending before the markets were shocked by the war in Ukraine.\footnote{We restrict ourselves to this period to mitigate the effects of stylised features of long macroeconomic time series, such as structural change in the first and second moments.} We consider variable selection in the ADF regression
\begin{align}
\Delta\pi_t = \mu + \rho\pi_{t-1} + \sum_{j=1}^p \delta_j \Delta\pi_{t-j} + e_t. \label{eq:infreg}
\end{align}
\Cref{eq:infreg} includes a constant since the ECB monetary policy strives to maintain the inflation rate $\pi_t$ in a corridor around a positive target, implying a stationary process with a non-zero mean.
Besides model selection using the shrinkage estimators, we apply information criteria for lag selection in conjunction with the classical ADF and quasi-difference demeaned ADF tests (DFQD) of \textcite{Elliottetal1996}. For the OLS-estimated models, we consider an exhaustive search for a sparsity pattern using BIC (denoted $\text{BIC}^\star$) as well as an estimation of the truncation lag (by AIC/MAIC). As before, $\lambda$ for PL, AL, and ALIE is tuned using BIC. For ALIE, we use $\Breve{w}_1$ with $J_\alpha$ computed based on $\widehat\omega^2_\textup{AR}(k)$ for $k$ selected by BIC, MAIC, or $k=11$ according to the \textcite{Schwert1989} rule, which is the maximum lag order ($p_{\max}$) considered across all procedures.
The outcomes are shown in Table~\ref{tab:infoutcomes}. For both inflation rates, the consistent $\text{BIC}^\star$ procedure estimates the lag pattern $\widehat{\mathcal{J}}=\{1, 4, 8\}$ which indicates short-run dependence on the seasonal lags $j=4$ and $j=8$. Reassuringly, the Lasso estimators select the same lag pattern as $\text{BIC}^\star$, except for PL for energy commodities. None of the classical tests reject a non-stationary model at the 5\% level for either series. ALIE classifies the all-products inflation rate as stationary for all three specifications of $\widehat\omega_\textup{AR}^2(k)$. It also includes $\pi_{t-1}$ in the energy inflation rate model for LRV estimates based on BIC and $p_\textup{max}$. PL and AL classify both series as non-stationary.
\section{Conclusion}
\label{sec:conc}
Previous research on penalised estimation has considered single identification principles to elicit data-dependent penalty weights. A prominent example is the adaptive Lasso, for which \textcite{zou2006adaptive} recommends using consistent coefficient estimates. While oracle properties grant asymptotic equivalence between such adaptive procedures and their non-penalised counterparts under appropriate conditions, simulation studies often point to mediocre finite sample performance of ad-hoc implementations.
In this article, we proposed combining two identification principles in computing penalty weights for consistent and oracle-efficient estimation using the adaptive Lasso. With ADF regressions being our use case, we propose to enrich a consistent coefficient estimate with additional information from a (simulation-based) statistic that exploits different stochastic orders of stationary and non-stationary data. We established that the enhanced weight promotes a \emph{zero shrinkage} property and ensures \emph{perpetual activation} of the corresponding regressor in stationary models as $T\to\infty$. In particular, we highlighted that beneficial consequences of these features for inference by consistent selection arise from an improved adaptive sorting of the activation knots on the Lasso solution path. We have further extended the theoretical analysis of the asymptotic properties of the adaptive Lasso for the AR(1) models examined in \textcite{kock2016consistent} to general ADF($p$) and ARMA models, even allowing $p\to\infty$ beyond the (approximate) sparsity assumption.
Simulation evidence supports our theoretical results, showing that the information-enriched ALIE estimator dominates the adaptive Lasso and is significantly more reliable than the (possibly inconsistent) plain Lasso regarding both the selection of $y_{t-1}$ and recovering the entire model. Additionally, ALIE improves the classification of stationary and non-stationary models in finite samples under detrending, which appears to be a weakness common to all Lasso variants compared to unpenalised regression-based tests. Abstracting from ADF regressions and as another novelty, we showcase the ability of information enrichment to combine hypothesis tests without type I error accumulation or, under consistent tuning, even with shrinking size. In contrast to most popular combination methods, ALIE does not require standardised evidence, e.g., p-values. Hence, ALIE can be considered a testbed for combining different estimators without establishing asymptotic distributions.
We apply the information-enriched estimator in a model selection exercise for German consumer price data from 2001-Q4 to 2021-Q4. The results indicate that a stationary ADF model is better suited to model the headline inflation rate than a non-stationary autoregression. We also find evidence supporting a stationary model for the German energy commodity inflation rate.
There are various avenues for further research: we consider it relevant to investigate whether the performance can be further improved by alternative specifications of the adaptive weight $\Breve{w}$. A first idea would be to swap the simulation-based statistic $J_\alpha$ with a different statistic. Our theory suggests benefits from adding further weakly correlated estimators to the computation of $\Breve{w}$. Second, since our method focuses exclusively on the weight for $y_{t-1}$, another path would be to elicit improved weights for the lags of $\Delta y_t$ to enhance further the classification rates concerning model selection. In this regard, it would also be worthwhile to investigate adaptive Lasso estimators that include deterministic components, perhaps in a flexible fashion as suggested in \textcite{tibshirani2011solution} and compare them with detrending approaches.
Another aspect yet to be investigated is the benefit of using ALIE for forecast averaging. Due to its superior selection power under consistent tuning, ALIE could help sort out bad predictors better than adaptive Lasso or comparable procedures can by drawing on other forecasts of the same predictor. In more general terms, it would also be instructive to analyse ALIE's performance under conservative model selection and to benchmark against pre-test or averaging estimators as suggested by \textcite{Hansen2007, Hansen2010}. We consider it also promising to extend our approach to multivariate models, e.g., for penalised estimation of the cointegration rank in VAR models as in the framework of \textcite{LiaoPhillips2015}.
Certain violations of our assumptions, e.g., smooth transitions in the (unconditional) variance, have received much attention in the unit root test literature during the last two decades as features of economic time series. Studying the behaviour of our approach in such a framework and making it adaptive would be of interest to a broad range of applications in empirical macroeconomic research. One opportunity is to devise a procedure robust to non-stationary volatility by employing scaled estimators based on non-parametric estimation of the series' variance profile as in \textcite{Beare2018}.
Finally, tuning the adaptive Lasso to detect local alternatives while exhibiting reliable model selection properties is an unresolved problem that is particularly challenging in the context of correlated regressors. We are currently investigating this issue.\\
\newpage