EconBase
← Back to paper

Beyond the Oracle Property: Adaptive LASSO in Cointegrating Regressions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

143,923 characters · 13 sections · 64 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Beyond the Oracle Property: Adaptive LASSO in Cointegrating Regressions with Local-to-Unity Regressors

abstractThis paper derives new asymptotic results for the adaptive LASSO estimator in cointegrating regressions, allowing for uncertainty about whether the regressors are exact unit root processes. We study model selection probabilities, estimator consistency, and limiting distributions under standard and moving-parameter asymptotics. We further derive uniform convergence rates and the fastest local-to-zero rates detectable by the estimator under conservative and consistent tuning. For consistent tuning, we construct confidence regions that are easy to implement, uniformly valid over the parameter space, and achieve sure asymptotic coverage without requiring knowledge or estimation of local-to-unity or long-run covariance parameters. Simulation results reveal that the finite-sample distribution of the adaptive LASSO estimator can deviate substantially from the oracle property, whereas moving-parameter asymptotics provide much more accurate approximations. Consequently, in addition to being infeasible in applications due to their dependence on non-estimable nuisance parameters, oracle-based confidence regions are often too small to achieve adequate coverage in empirically relevant scenarios with small but non-zero coefficients. In contrast, the proposed confidence regions are always feasible and deliver reliable coverage across the parameter space. An empirical application to predicting the U.S. unemployment rate illustrates their practical usefulness for quantifying uncertainty around adaptive LASSO estimates. \noindentKeywords: Adaptive LASSO, Confidence regions, Local-to-unity regressors, Moving-parameter asymptotics, Shrinkage estimation, Variable selection \noindentJEL classification: C22, C51, C52, C61

Introduction

In recent years, the availability of large datasets comprising numerous economic and financial variables has become the rule rather than the exception. Consequently, practitioners using traditional methods are frequently confronted with the challenge of selecting a small number of relevant variables from an extensive pool of potential covariates. In this context, statistical methods that simultaneously perform estimation and variable selection, such as variants of the least absolute shrinkage and selection operator (LASSO) introduced in Tibshirani96, are becoming increasingly popular in econometrics.

While contributions such as WangEtAl07b, RenZhang10, MedeirosMendes16, AdamekEtAl23, and ChenEtAl25 examine the use of LASSO-type estimators in models with (locally) stationary time series, a growing body of recent research considers models with highly persistent and endogenous regressors exhibiting unit root or local-to-unit root behavior. For example, LiaoPhillips15 propose adaptive shrinkage methods to estimate vector error correction models. KooEtAl20 and MeiShi24 consider so-called predictive regressions with high-dimensional stationary and unit root regressors and derive certain asymptotic properties of LASSO-type procedures in this context. SmeekesWijler21 consider asymptotic properties of a LASSO-type estimator in a high-dimensional error correction model. Schweikert22 proposes an adaptive group LASSO method to estimate structural breaks in cointegrating regressions and TuXie23 follow a similar agenda in predictive regressions with a fixed number of highly persistent regressors. LeeEtAl22 derive asymptotic properties of LASSO-type estimators in regressions with a fixed number of stationary and (potentially cointegrated) local-to-unity regressors. GonzaloPitarakis25 propose a test for cointegration in a regression model with a fixed number of regressors based on the residuals of the adaptive LASSO estimator.

Typically, these papers consider regression models where the regressors can be split into a set of relevant regressors, i.e., those with a non-zero regression coefficient, and a set of irrelevant regressors, i.e., those with a regression coefficient being exactly equal to zero. Then these articles, among other things, often derive model selection probabilities and sometimes also the limiting distribution of the LASSO-type estimator for the set of non-zero coefficients. If the procedure identifies zero coefficients with probability approaching one and if the limiting distribution for the non-zero coefficients coincides with that of the ordinary least squares (OLS) estimator applied to the true model, the estimator is said to possess the “oracle property” FanLi01. While the oracle property certainly appears convenient, it has to be interpreted with extreme caution. The oracle property primarily characterizes the asymptotic behavior of the penalized estimation method under the assumption that coefficients are either exactly zero or sufficiently large (in absolute value, relative to the sample size). It thus offers very limited guidance in empirically relevant situations where some coefficients are small (in absolute value, relative to sample size), but not exactly zero. To study the asymptotic properties when coefficients are allowed to be small but unequal to zero, one has to let the true coefficients depend on sample size.

The contributions of LeeEtAl22 and TuXie23 take a step forward by providing such moving-parameter asymptotic properties of the LASSO-type procedures under consideration. Their results, however, are restricted to specific sequences, as they focus on particular rates at which the true coefficients go to zero, and also couple these rates to the choice of the tuning parameter. We discuss the implications of these restrictions at several points in this paper.

As a first illustration that these restrictions are not innocuous, we refer to Kock16, who analyzes the asymptotic properties of the adaptive LASSO estimator in stationary and non-stationary autoregressions. For example, in the AR(1) case $$ \Delta y_t = \rho_T y_{t-1} + \ensuremath{\varepsilon}_t, $$ with i.i.d. errors $\ensuremath{\varepsilon}_t$, the estimator's behavior critically depends on both the value of $\rho_T$ as well as the choice of the tuning parameter. In particular, when $\rho_T$ approaches zero at rate $1/T$, the procedure fails to detect the coefficient as non-zero if the tuning parameter diverges, whereas it succeeds with positive probability if the tuning parameter remains bounded. Moreover, if the tuning parameter converges to zero as well, the procedure detects the coefficient as non-zero with probability approaching one. What remains unresolved, however, is the cut-off rate for the local-to-zero coefficient that can still be detected by the estimator when the tuning parameter diverges.

These so-called local-to-zero rates are motivated not only by theoretical considerations but also by their relevance in empirical applications. For example, they provide a natural framework for modeling weak signal-to-noise ratios, which play an important role in, e.g., the analysis of stock return predictability CampbellYogo06,Campbell08,Phillips15,DemetrescuEtAl22. In this context, a detailed analysis of the properties of LASSO-type estimators under general sequences at which the true coefficients are allowed to go to zero, reveals the fastest local-to-zero rates that can still be detected by the estimator.

In this paper, we provide a comprehensive analysis of the asymptotic properties of the adaptive LASSO estimator Zou06 applied to cointegrating regressions with potentially local-to-unity regressors. Allowing for deviations from exact unit roots is important, as macroeconomic variables often display high persistence without being unit-root processes, see, e.g., Jensen09 for evidence on inflation, and HwangValdes24 for further discussion. In particular, we derive model selection probabilities, estimator consistency, and limiting distributions, while allowing the true coefficients to freely move through the parameter space along arbitrary sequences. In addition, we derive uniform convergence rates and the local-to-zero rates of the true coefficients that can still be detected by the estimator. These findings will shed light on, e.g., the signal-to-noise ratios practitioners can accept when employing the adaptive LASSO estimator in empirical applications. We complete our theoretical analysis with providing uniformly valid confidence regions in the typical regime when the tuning parameter diverges.

Before we discuss our results in more detail, note that in the context of classical linear regression models with non-stochastic regressors, PoetscherSchneider09 provide a comprehensive analysis of the adaptive LASSO estimator, examining both its asymptotic behavior and finite-sample properties. In particular, they derive uniform convergence rates and highlight how the choice of tuning parameters influences the performance of the estimation procedure. In contrast to the analysis of PoetscherSchneider09, we have to overcome several difficulties to derive the results of this paper. First, we are dealing with different convergence rates of the estimators and second, we have to account for the stochastic and non-stationary nature of the regressors, which leads to the occurrence of stochastic integrals in the limit. Moreover, the second-order bias terms in the limiting distribution of the OLS estimator also affect the limiting distribution of the adaptive LASSO estimator. Finally, we also provide results for the multivariate case which is not addressed in the aforementioned article.

Based on the asymptotic study of model selection probabilities, we distinguish between two regimes determined by the large-sample behavior of the tuning parameter: consistent model selection (or “consistent tuning”), where zero coefficients are found with asymptotic probability equal to one, and conservative model selection (“conservative tuning”), where zero coefficients are detected with asymptotic probability less than one. The asymptotic properties of the adaptive LASSO estimator differ substantially between these two cases, with the main massages as follows: In the conservatively tuned case, the estimator is uniformly $T$-consistent for parameter estimation and the cut-off rate for local-to-zero coefficients that can be detected by the procedure is $1/T$. In the consistently tuned case, the uniform convergence rate depends on the tuning parameter and is slower than $1/T$. Deviations of the true parameter from zero of rate $1/T$ cannot be discovered by the estimator. The fastest local-to-zero rate that is still detectable with positive probability again depends on the tuning parameter and is slower than $1/T$. Moreover, in the consistently tuned case, the detailed theoretical analysis of the adaptive LASSO estimator allows us to construct confidence regions that have coverage probability approaching one uniformly over the parameter space, without requiring any knowledge or estimation of local-to-unity or long-run covariance parameters. Although inspired by the non-stochastic regressor case considered in AmannSchneider23, extending the construction of such regions to the stochastic regressors present in the unit-root or local-to-unity setting substantially changes the nature of the problem.

The theoretical analysis is complemented by an extensive simulation study. The results show that the finite-sample distribution of the adaptive LASSO estimator often deviates substantially from what is suggested by the oracle property, whereas the limiting distributions derived under moving-parameter asymptotics capture the finite-sample properties of the procedure more closely. Moreover, the poor approximation quality of the oracle property to the finite-sample distribution of the adaptive LASSO estimator is also reflected in the performance of the confidence regions based on the oracle property. In particular, we find that the oracle-based confidence regions are often too small to achieve adequate coverage in empirically relevant scenarios with small but non-zero coefficients, whereas the confidence regions proposed in this paper achieve adequate coverage across the entire parameter space.

Taken together, the theoretical and simulation results indicate that the oracle property provides an incomplete characterization of both the asymptotic and finite-sample properties of the adaptive LASSO estimator. A full understanding instead requires a moving-parameter framework of the type considered in this paper. Moving beyond the oracle property in this way enables the construction of uniformly valid confidence regions under consistent tuning.

Finally, an empirical application to predicting the U.S. unemployment rate illustrates the usefulness of the proposed confidence regions for quantifying uncertainty around adaptive LASSO estimates.

The paper is organized as follows. Section (ref) introduces the model and states the assumptions. Section (ref) contains our theoretical contributions: Section (ref) derives the large-sample properties of the adaptive LASSO estimator in a fixed-parameter asymptotic framework in the univariate regressor case, Section (ref) considers a moving-parameter asymptotic framework in the univariate regressor case, and Section (ref) extends the results to the multivariate case. Section (ref) then derives the uniform confidence region based on the adaptive LASSO estimator. Section (ref) presents the simulation results and Section (ref) contains the empirical application. Section (ref) summarizes and concludes. All proofs are provided in the appendix, which also contains additional simulation and empirical results.

We use the following notation: $\lfloorx\rfloor$ denotes the integer part of $x \in \ensuremath{{\mathbb R}}$, $L$ is the backward-shift operator, $\text{diag}(\cdot)$ denotes a diagonal matrix with elements specified throughout, and $\ensuremath{\overline{\ensuremath{{\mathbb R}}}} \coloneqq \ensuremath{{\mathbb R}} \cup \{-\infty,\infty\}$. With $\ensuremath{\Rightarrow}$ and $\ensuremath{\overset{p}{\longrightarrow}}$ we denote weak convergence and convergence in probability, respectively, and all limits apply as the sample size $T$ tends to infinity. We denote a normal distribution with mean $\mu$ and covariance matrix $\Sigma$ as $\mathcal{N}\left(\mu,\Sigma\right)$. The symbol $I_k$ denotes the $k$-dimensional identity matrix. For any event $E$, the indicator function $\ensuremath{\mathbbm{1}}\{E\}$ equals one if $E$ occurs and zero otherwise. If a sequence $a_T$ is identical to $a\in\ensuremath{{\mathbb R}}$ for all $T$, we write $a_T \equiv a$. By $\omega$ we denote an element of the sample space of the underlying probability space and $(\omega)$ attached to a random variable denotes its realization for this particular $\omega$.

Setting and Assumptions

As motivated in the introduction, we consider a cointegrating regression model with local-to-unity regressors of the form

align[align omitted — 114 chars of source]

for $t = 1,\dots,T$, where $c \coloneqq \text{\rm diag}(c_1,\ldots,c_k)$ with $c_j\geq0$, and $x_0 = O_{\ensuremath{{\mathbb P}}}(1)$. For $c=0$, the model encompasses classical cointegrating regressions with unit root regressors. Following LeeEtAl22, we treat the number of regressors $k$ as fixed. For notational brevity, we exclude deterministic components from (ref). For $\{w_t\}_{t \in \ensuremath{{\mathbb Z}}} \coloneqq \{[u_t,v_t']'\}_{t \in \ensuremath{{\mathbb Z}}}$ we impose the following assumption.

assumptionLet $w_t = \Psi(L) \ensuremath{\varepsilon}_t = \sum_{j=0}^\infty \Psi_j \ensuremath{\varepsilon}_{t-j}$, with $\sum_{j=0}^\infty j \Vert \Psi_j\Vert < \infty$ and $\det(\Psi(1)) \neq 0$, where $\{\ensuremath{\varepsilon}_t\}_{t \in \ensuremath{{\mathbb Z}}}$ is a $(1+k)$-dimensional strictly stationary ergodic martingale difference sequence with natural filtration $\mathcal{F}_t \coloneqq \sigma\left(\{\ensuremath{\varepsilon}_s\}_{-\infty}^t\right)$, conditional covariance matrix $\Sigma \coloneqq \ensuremath{{\mathbb E}}(\ensuremath{\varepsilon}_t\ensuremath{\varepsilon}_t'| \mathcal{F}_{t-1}) > 0$, and $\sup_{t \geq 1}\ensuremath{{\mathbb E}}(\Vert\ensuremath{\varepsilon}_t\Vert^r|\mathcal{F}_{t-1}) < \infty$ a.s.\ for some $r > 4$.

Conditions similar to Assumption (ref) are common in the cointegrating regression literature, see, e.g., WagnerHong16 for a detailed discussion. In particular, Assumption (ref) allows for regressor endogeneity and error serial correlation, but excludes cointegration among the elements of $x_t$.\footnote{As shown by LeeEtAl22, allowing for cointegration among the regressors requires a different estimation strategy, termed the twin adaptive LASSO. Extending our analysis to this estimator in the more general setting is left for future research.} Under Assumption (ref), the process $\{w_t\}_{t \in \ensuremath{{\mathbb Z}}}$ fulfills a functional central limit theorem of the form

align[align omitted — 181 chars of source]

where $W(r) = [W_{u \cdot v}(r),W_v(r)']'$ is a $(1 + k)$-dimensional vector of independent standard Brownian motions and $\Omega \coloneqq \sum_{h = -\infty}^\infty \ensuremath{{\mathbb E}}(w_0 w_h')>0$ denotes the long-run covariance matrix of $\{w_t\}_{t \in \ensuremath{{\mathbb Z}}}$. Note that the results in this paper also hold under alternative sets of assumptions as long as they imply the functional central limit theorem for $\{w_t\}_{t \in \ensuremath{{\mathbb Z}}}$ given in (ref), see, e.g., IbragimovPhillips08 and deJong03UP for possible other conditions.

The target of our investigation, the adaptive LASSO estimator Zou06 of $\beta_T$ in (ref), is defined as

align[align omitted — 265 chars of source]

where $\lambda_T > 0$ and $\gamma \geq 1$ are tuning parameters. In contrast to the classical LASSO estimator of Tibshirani96, the penalty term for the $j$-th coefficient in (ref) contains the reciprocal of the absolute value of a preliminary estimator $\hat\beta_j^0$ of $\beta_{T,j}$, where $\beta_{T,j}$ denotes the $j$-th element of $\beta_T$. Its aim is to increase the penalty term if $\beta_{T,j}$ seems small to encourage shrinking, and to penalize less if $\beta_{T,j}$ appears to be large in order to reduce the bias. In practice, $\gamma$ is often chosen as 1 or 2 and $\lambda_T$ is typically selected based on cross-validation or information criteria. In line with the recommendations in LeeEtAl22, we set $\gamma = 1$ and $\hat\beta^0 = \hat\beta$, where $\hat\beta$ denotes the OLS estimator of $\beta_T$ in (ref). Under Assumption (ref), the limiting distribution of $\hat\beta$ is given by

align[align omitted — 195 chars of source]

where $J_v^c(r) \coloneqq \int_0^r e^{(r-s)c}dB_v(s)$ and $\Delta_{vu} \coloneqq \sum_{h=0}^\infty \ensuremath{{\mathbb E}}(v_0u_h)$, see, Phillips88. To simplify notation, we define $\zeta_{vv}^c \coloneqq \int_0^1 J_v^c(r)J_v^c(r)'dr$.

If at least one regressor is endogenous, the limiting distribution of the OLS estimator is contaminated by second-order bias terms. In contrast to the classical cointegrating regressions with unit root regressors, where such bias terms can be addressed, for example, using the fully modified approach of PhillipsHansen90, in the local-to-unity regressor case, these bias terms additionally depend on the unknown local-to-unity parameters $c_j$. Since these parameters are not consistently estimable, the resulting bias is difficult to correct, see, e.g., Phillips23 and the references therein for a detailed discussion. Consequently, constructing asymptotically valid confidence intervals or hypothesis tests in cointegrating regressions with local-to-unity regressors is non-trivial, see, e.g., MagdalinosPhillips09 for an instrumental variables approach and HwangValdes24 for a modified low-frequency transformed and augmented OLS method. In Section (ref), we show that asymptotically valid uniform confidence regions can nevertheless be constructed from the consistently tuned adaptive LASSO estimator which are straightforward to implement, and do not require any knowledge of local-to-unity parameters.

remarkOur results also extend to the predictive regression setting KooEtAl20,MeiShi24, where the regressor at time $t$ is $x_{t-1}$ rather than $x_t$. In this setting, the limiting distribution of the OLS estimator coincides with $\mathcal{Z}^c$, except that $\Delta_{vu}$ needs to be replaced by $\sum_{h=1}^\infty \ensuremath{{\mathbb E}}(v_0u_h)$.

Before deriving the asymptotic properties of the adaptive LASSO estimator in the multivariate case, where no closed-form solution of the minimization problem in (ref) is available, we consider the univariate case, where the minimization problem has an explicit solution of the form

equation[equation omitted — 236 chars of source]

with $\tilde\lambda_T \coloneqq 0.5\lambda_T (\sum_{t=1}^T x_t^2)^{-1}$ and $\sum_{t=1}^T x_t^2 = O_\ensuremath{{\mathbb P}}(T^{-2})$, compare PoetscherSchneider09. Analyzing the univariate case in detail facilitates a transparent derivation of the estimator’s asymptotic properties and provides insights into the underlying mechanisms, helping to clarify the multivariate results.

Equation (ref) reveals that in the univariate case the adaptive LASSO estimator can be represented solely in terms of the OLS estimator and that the tuning parameter $\lambda_T$ only affects the estimator in its “standardized” version $\tilde\lambda_T$, where the term $\sum_{t=1}^T x_t^2$ can be viewed as a measure of variation in the regressor. This explains why for the asymptotic study in the subsequent section, the large-sample behavior of $\lambda_T$ relative to $T^{-2}$ is important -- which is the rate at which $\sum_{t=1}^T x_t^2$ stabilizes.

Figure (ref) illustrates the relationship between $\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}}$ and $\hat\beta$ for different values of $\tilde\lambda_T$.

figure[figure omitted — 319 chars of source]

Asymptotic Theory

In this section, we investigate the large-sample behavior of the adaptive LASSO estimator under two different asymptotic regimes regarding the model selection properties of the procedure. We speak of consistent model selection (or “consistent tuning”) if all zero coefficients are detected with asymptotic probability equal to one, whereas the case where at least one zero coefficient is set to zero by the estimator with limiting probability strictly less than one is referred to as conservative model selection (“conservative tuning”). Formally, this definition only relates to zero coefficients and poses no requirement on the non-zero coefficients in the model. Which regime applies depends on the limiting behavior of the tuning parameter sequence $\lambda_T$, as will be clarified below.

We first consider the univariate regressor case. In Section (ref), we set $\beta_T \equiv \beta$ and study how the behavior of the tuning parameter sequence $\lambda_T$ affects both model selection and parameter estimation. In addition, we derive the asymptotic distribution of the estimator when $\beta_T \equiv \beta$ is fixed under both model selection regimes. The insights from Section (ref) into which limiting behavior of $\lambda_T$ leads to what type of model selection regime then serve as a starting point for the detailed analysis of the large-sample behavior of the adaptive LASSO estimator in Section (ref). In that section, we adopt a moving-parameter framework in which the true parameter $\beta_T$ may vary with the sample size $T$. This allows to determine local-to-zero and uniform convergence rates, and to derive asymptotic distributions that, as will become apparent in the simulation study in Section (ref), more accurately capture the estimator’s finite-sample properties. The analysis is again conducted under both conservative and consistent model selection.

After utilizing the explicit expression of the adaptive LASSO estimator in the univariate case, we turn to the multivariate case in Section (ref). In this case, the absence of a closed-form solution necessitates different techniques for deriving the asymptotic properties. Unlike in the univariate case, we do not separately present results under fixed-parameter asymptotics, since these are encompassed by the moving-parameter framework and do not provide additional insights beyond those already established in the univariate analysis. As before, we study the asymptotic behavior of the estimator under both conservative and consistent model selection.

Finally, in Section (ref), we use the insights from Section (ref) to construct asymptotically valid uniform confidence regions for the consistently tuned adaptive LASSO estimator.

Fixed-Parameter Asymptotics in the Univariate Case

As outlined above, we begin by deriving asymptotic results for the adaptive LASSO estimator in the univariate regressor case under a fixed-parameter framework, i.e., by setting $\beta_T \equiv \beta$ in (ref) with $\beta \in \ensuremath{{\mathbb R}}$ fixed. We first examine the large-sample properties of the estimator with respect to model selection.

proposition[Model selection] Let $\{y_t\}_{t\in\ensuremath{{\mathbb Z}}}$ and $\{x_t\}_{t\in\ensuremath{{\mathbb Z}}}$ be generated by (ref) and (ref) with $k=1$ and $\beta_T\equiv\beta$, and let $\{w_t\}_{t\in\ensuremath{{\mathbb Z}}}$ satisfy Assumption (ref). \begin{enumerate} • Let $\beta\neq 0$. If $T^{-2}\lambda_T \to 0$, then $\ensuremath{{\mathbb P}}\left(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} = 0\right) \to 0$. • Let $\beta=0$. \begin{enumerate} • If $\lambda_T \to \lambda_0$, $0 \leq \lambda_0 < \infty$, then $$ \ensuremath{{\mathbb P}}\left(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} = 0\right) \to \ensuremath{{\mathbb P}}\left((\zeta_{vv}^c)^{1/2}\left|\mathcal{Z}^c\right| \leq \sqrt{\frac{\lambda_0}{2}}\right) < 1. $$ • If $\lambda_T \to \infty$, then $\ensuremath{{\mathbb P}}\left(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} = 0\right) \to 1$. \end{enumerate} \end{enumerate}

Proposition (ref) reveals the role of the tuning parameter sequence $\lambda_T$ for model selection of the adaptive LASSO: In case $\lambda_T \to \lambda_0$ with $0 \leq \lambda_0 < \infty$, the estimator detects zero coefficients with probability smaller than one asymptotically, resulting in conservative model selection. In contrast, when $\lambda_T \to \infty$, the estimator sets zero coefficients equal to zero with probability approaching one and consequently leads to consistent model selection. In the following, we therefore refer to the case $\lambda_T \to \lambda_0$, $0 \leq \lambda_0 < \infty$, as conservative tuning, whereas the case $\lambda_T\to\infty$ is termed consistent tuning.\footnote{Under conservative tuning with $\lambda_0 = 0$, the adaptive LASSO estimator is equivalent to OLS in the sense that $T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \hat\beta) = o_\ensuremath{{\mathbb P}}(1)$. This follows directly from the proof of Theorem (ref)(ref) in Section (ref). The result also holds in a moving-parameter framework and extends to the multivariate case.} Moreover, the condition $T^{-2}\lambda_T \to 0$ is a basic requirement for the tuning parameter as it ensures that the probability of the adaptive LASSO estimator incorrectly setting the coefficient to zero vanishes asymptotically. While this condition is automatically fulfilled under conservative tuning, it controls the rate at which $\lambda_T$ may diverge under consistent tuning. We will assume this condition in all subsequent statements in this section.

We now derive the asymptotic properties of the adaptive LASSO estimator with respect to parameter estimation.

proposition[Parameter estimation] Let $\{y_t\}_{t \in \ensuremath{{\mathbb Z}}}$ and $\{x_t\}_{t \in \ensuremath{{\mathbb Z}}}$ be generated by (ref) and (ref) with $k=1$ and $\beta_T \equiv \beta$, let $\{w_t\}_{t \in \ensuremath{{\mathbb Z}}}$ satisfy Assumption (ref), and assume $T^{-2}\lambda_T \to 0$. \begin{enumerate} • $\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta = o_\ensuremath{{\mathbb P}}(1)$. • If $T^{-1}\lambda_T \to \tilde\lambda_0$, $0 \leq \tilde\lambda_0 < \infty$, then $T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta) = O_\ensuremath{{\mathbb P}}(1)$. • If $T^{-1}\lambda_T \to \infty$ then $\lambda_T^{-1}T^2(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta) = O_\ensuremath{{\mathbb P}}(1)$. \end{enumerate}

From Proposition (ref)(ref), we learn that the basic condition $T^{-2}\lambda_T \to 0$, which ensures that non-zero coefficients are not falsely put to zero as shown in Proposition (ref)(ref), also guarantees that the procedure is consistent for $\beta$ with respect to parameter estimation. Proposition (ref)(ref) shows that in case of conservative tuning or for consistent tuning with a slowly diverging tuning parameter sequence (in the sense that $T^{-1}\lambda_T$ stays bounded), the convergence rate of the adaptive LASSO estimator is $T^{-1}$ and coincides with the rate of OLS. However, when the estimator is tuned consistently and $\lambda_T$ tends to infinity fast enough so that $T^{-1}\lambda_T$ diverges also, Proposition (ref)(c) reveals that the convergence rate of the adaptive LASSO estimator is only $T^{-2}\lambda_T$, which is slower than $T^{-1}$ in this case.

We now derive the limiting distribution of the adaptive LASSO estimator.

proposition[Limiting distribution] Let $\{y_t\}_{t \in \ensuremath{{\mathbb Z}}}$ and $\{x_t\}_{t \in \ensuremath{{\mathbb Z}}}$ be generated by (ref) and (ref) with $k=1$ and $\beta_T \equiv \beta$, let $\{w_t\}_{t \in \ensuremath{{\mathbb Z}}}$ satisfy Assumption (ref), and assume $T^{-2}\lambda_T \to 0$. \begin{enumerate} • Let $\lambda_T \to \lambda_0$, $0 \leq \lambda_0 < \infty$. \begin{itemize} • If $\beta \neq 0$, then $T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta) \Rightarrow \mathcal{Z}^c$. • If $\beta = 0$, then \begin{align*} T(\ensuremath{\hat\beta_{\rm\tiny AL}} - \beta) \Rightarrow \ensuremath{\mathbbm{1}}\left\{(\zeta_{vv}^c)^{1/2}\left|\mathcal{Z}^c\right| > \sqrt{\frac{\lambda_0}{2}}\right\}\left(\mathcal{Z}^c - \frac{\lambda_0}{2\zeta_{vv}^c}(\mathcal{Z}^c)^{-1}\right). \end{align*} \end{itemize} • Let $\lambda_T \to \infty$. \begin{itemize} • If $T^{-1} \lambda_T \to \tilde\lambda_0$, $0 \leq \tilde\lambda_0 < \infty$, then \begin{align*} T(\ensuremath{\hat\beta_{\rm\tiny AL}} - \beta) \Rightarrow \ensuremath{\mathbbm{1}}\left\{\beta \neq 0\right\} \left(\mathcal{Z}^c - (\zeta_{vv}^c)^{-1}\frac{\tilde\lambda_0}{2\beta}\right). \end{align*} • If $T^{-1} \lambda_T \to \infty$, then \begin{align*} \lambda_T^{-1}T^2(\ensuremath{\hat\beta_{\rm\tiny AL}} - \beta) \Rightarrow - \ensuremath{\mathbbm{1}}\left\{\beta \neq 0\right\} (\zeta_{vv}^c)^{-1}\frac{1}{2\beta}. \end{align*} \end{itemize} \end{enumerate}

Under conservative tuning, Proposition (ref)(ref) reveals that the asymptotic distribution of the adaptive LASSO estimator coincides with the one of OLS if $\beta\neq 0$. In case $\beta = 0$, however, the limiting distribution consists of an atomic part at zero incurred by the positive probability of the estimator being equal to zero, and an absolutely continuous part for when the estimator is not equal to zero.

Proposition (ref)(ref) further shows that under consistent tuning, the limiting distribution of the adaptive LASSO estimator fully collapses to pointmass at zero whenever $\beta = 0$.\footnote{For completeness, note that the proofs of (b1) and (b2) in Proposition (ref) reveal that $\delta_T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta) \Rightarrow 0$ for any sequence $\delta_T \to \infty$ if $\beta = 0$ and $\lambda_T \to \infty$.} In case $\beta \neq 0$, both the convergence rate of the adaptive LASSO estimator and its limiting distribution depend on how $\lambda_T$ diverges in relation to $T$. If $\lambda_T$ tends to infinity slower than $T$, the limiting distribution of $T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta)$ coincides with the one of OLS. If $\lambda_T$ diverges at rate $T$, rate-$T$ consistency is still maintained as is already shown in Proposition (ref)(b1), but the asymptotic distribution now deviates from the one of OLS by a random shift that is inversely proportional to and has the opposite sign of $\beta$. This random shift in the limiting distribution of the adaptive LASSO estimator is not detected in LeeEtAl22, as their assumptions imply that $\tilde\lambda_0 = 0$. Moreover, in contrast to the second order bias term in the limiting distribution of the OLS estimator, the random shift in the limiting distribution of the adaptive LASSO estimator does not vanish if the regressor is exogenous. Lastly, if $\lambda_T$ diverges faster than $T$, the convergence rate becomes slower than $T$, as shown in Proposition (ref)(b2). In this case, for the appropriately scaled estimator $\lambda_T^{-1}T^2(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta)$, the term “corresponding to OLS” no longer appears in the limit. Finally, Proposition (ref) confirms that the rates established in Proposition (ref) are indeed sharp.

The results in this section can be used to explicitly formulate a condition for the so-called “oracle property” of the adaptive LASSO estimator, a term coined by FanLi01, established for the adaptive LASSO estimator in a classical linear regression model in Zou06.

corollary[“Oracle property”] Let $\{y_t\}_{t \in \ensuremath{{\mathbb Z}}}$ and $\{x_t\}_{t \in \ensuremath{{\mathbb Z}}}$ be generated by (ref) and (ref) with $k=1$ and $\beta_T \equiv \beta$, let $\{w_t\}_{t \in \ensuremath{{\mathbb Z}}}$ satisfy Assumption (ref), and assume $T^{-1} \lambda_T + \lambda_T^{-1} \to 0$. Then $\ensuremath{{\mathbb P}}\left(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} = 0\right) \to \ensuremath{\mathbbm{1}}\{\beta=0\}$ and $T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} -\beta) \Rightarrow \ensuremath{\mathbbm{1}}\left\{\beta \neq 0\right\}\mathcal{Z}^c$.

Corollary (ref) states that, under consistent tuning with $\lambda_T$ diverging more slowly than $T$, the adaptive LASSO estimator identifies non-zero coefficients and sets null coefficients to zero with probability approaching one. Moreover, its limiting distribution coincides with that of the OLS estimator whenever $\beta\neq 0$.

Results similar to Corollary (ref) often constitute the main asymptotic findings in the literature on LASSO-type estimators across various contexts, see, e.g., MedeirosMendes16, SmeekesWijler21, Schweikert22, LeeEtAl22, TuXie23, and ChenEtAl25. However, although such results may seem convenient, they have to be interpreted with extreme caution. While the “oracle property” represents the large-sample performance of the estimator in situations where regression coefficients are either equal to zero or “relatively large” (in absolute value and in relation to sample size), it does not shed light on the empirically relevant case where some coefficients are “relatively small” rather than being exactly equal to zero. In Section (ref), we analyze the large-sample properties of the adaptive LASSO estimator within an asymptotic framework that also accommodates this case.

Moving-Parameter Asymptotics in the Univariate Case

In this section, we study the asymptotic behavior of the adaptive LASSO estimator in the univariate regressor case within a moving-parameter framework, where the unknown coefficient $\beta_T$ may vary with $T$. This framework overcomes the limitations of the fixed-parameter setting, in which the true coefficient is restricted to being either exactly zero or asymptotically large relative to sample size. Such a dichotomy is unsatisfactory, since in finite samples the coefficient may be non-zero yet small, especially when the signal-to-noise ratio is low. The smaller the non-zero coefficient that an estimator can still reliably detect, the better its performance. A key advantage of the moving-parameter framework is that it reveals the local-to-zero rate at which the estimator can detect non-zero coefficients.

In the following theorem, we derive the model selection probabilities of the adaptive LASSO under both conservative and consistent tuning.

theorem[Model selection] Let $\{y_t\}_{t \in \ensuremath{{\mathbb Z}}}$ and $\{x_t\}_{t \in \ensuremath{{\mathbb Z}}}$ be generated by (ref) and (ref) with $k=1$, and let $\{w_t\}_{t \in \ensuremath{{\mathbb Z}}}$ satisfy Assumption (ref). \begin{enumerate} • If $\lambda_T \to \lambda_0$, $0 \leq \lambda_0< \infty$, and $T\beta_T \to \beta_0 \in \ensuremath{\overline{\ensuremath{{\mathbb R}}}}$, then $$ \ensuremath{{\mathbb P}}\left(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} = 0\right) \to \ensuremath{{\mathbb P}}\left((\zeta_{vv}^c)^{1/2}\left|\mathcal{Z}^c + \beta_0\right| \leq \sqrt{\frac{\lambda_0}{2}}\right) < 1. $$ • If $\lambda_T \to \infty$ and $\lambda_T^{-1/2}T\beta_T \to \tilde\beta_0 \in \ensuremath{\overline{\ensuremath{{\mathbb R}}}}$, then $$ \ensuremath{{\mathbb P}}\left(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} = 0\right) \to \ensuremath{{\mathbb P}}\left((\zeta_{vv}^c)^{1/2} \leq \frac{1}{\sqrt{2}}\left| \tilde\beta_0 \right|^{-1}\right). $$ \end{enumerate}
remarkTheorem (ref) describes the asymptotic behavior of the model selection probabilities for arbitrary sequences of $\beta_T$ in the sense that all accumulation points of the selection probabilities can be obtained in the following way: apply the result to subsequences and observe that for every such subsequence, we can select a further subsequence such that relevant quantities, i.e., $T\beta_T$ or $\lambda_T^{-1/2}T\beta_T$, converge to a limit in $\ensuremath{\overline{\ensuremath{{\mathbb R}}}}$. A similar comment also applies to Theorems (ref)--(ref) below.

Part (a) of Theorem (ref) shows that under conservative tuning, if $\beta_T$ is bounded away from zero or converges to zero at rate slower than $T^{-1}$, i.e., $|\beta_0| = \infty$, the estimator can detect the coefficient as non-zero with asymptotic probability equal to one. If $\beta_T \equiv 0$ or $\beta_T$ converges to zero at rate $T^{-1}$ or faster, i.e., $\beta_0 \in \ensuremath{{\mathbb R}}$, the estimator will set the coefficient equal to zero with positive probability less than one even asymptotically.

To interpret the results in (b) of Theorem (ref) in a meaningful way, we suppose that the basic condition $T^{-2}\lambda_T \to 0$ from Section (ref) holds, which we also assume to hold in all subsequent statements in this section. Part (b) of Theorem (ref) then reveals that if $\beta_T$ is bounded away from zero or converges to zero at rate slower than $T^{-1}\lambda_T^{1/2}$, i.e., $|\tilde\beta_0| = \infty$, the estimator can detect the coefficient as non-zero with asymptotic probability equal to one. If $\beta_T$ converges to zero exactly at rate $T^{-1}\lambda_T^{1/2}$, i.e., $\tilde\beta_0 \in \ensuremath{{\mathbb R}}$, $\tilde\beta_0 \neq 0$, the estimator will set the coefficient equal to zero with positive probability less than one asymptotically. Finally, if $\beta_T \equiv 0$ or $\beta_T$ converges to zero with rate faster than $T^{-1}\lambda_T^{1/2}$, the estimator will set the coefficient equal to zero with asymptotic probability equal to one.

remarkAs discussed above, Theorem (ref) shows that in the consistently tuned case, the relevant local-to-zero rate is $T^{-1}\lambda_T^{1/2}$ in the sense that coefficients that converge to zero slower than that will be detected as non-zero with asymptotic probability equal to one and coefficients that converge to zero faster will be detected as non-zero with asymptotic probability equal to zero. In LeeEtAl22, it appears that coefficients converging to zero at rate $T^{-\delta}$ for any $\delta \in (0,1)$ pose no difficulty for the consistently tuned adaptive Lasso in the sense that they will be detected as non-zero with asymptotic probability equal to one. This, however, is made possible by an assumption that links the true coefficient to the tuning parameter (through the parameter $\delta$), thereby masking the dependence of the local-to-zero rate on the tuning parameter.

Next, we analyze estimation consistency of the adaptive LASSO estimator for the parameter $\beta_T$.

theorem[Parameter estimation] Let $\{y_t\}_{t \in \ensuremath{{\mathbb Z}}}$ and $\{x_t\}_{t \in \ensuremath{{\mathbb Z}}}$ be generated by (ref) and (ref) with $k=1$, let $\{w_t\}_{t \in \ensuremath{{\mathbb Z}}}$ satisfy Assumption (ref), and assume $T^{-2}\lambda_T \to 0$. \begin{enumerate} • It holds that $\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta_T = o_\ensuremath{{\mathbb P}}(1)$. • If $\lambda_T \to \lambda_0$, $0 \leq \lambda_0 < \infty$, then $T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta_T) = O_\ensuremath{{\mathbb P}}(1)$. • If $\lambda_T \to \infty$, then $\lambda_T^{-1/2}T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta_T) = O_\ensuremath{{\mathbb P}}(1)$. \end{enumerate}

Theorem (ref)(a) shows that if $T^{-2}\lambda_T \to 0$, the adaptive LASSO estimator is not only consistent (cf.\ Proposition (ref)), but also uniformly consistent. Parts (b) and (c) reveal that the uniform convergence rate depends on the tuning regime. Under conservative tuning, the estimator is rate-$T$ consistent, whereas under consistent tuning, it is only rate-$T\lambda_T^{-1/2}$ consistent.

We now derive the limiting distribution of the adaptive LASSO estimator under arbitrary sequences of $\beta_T$.

theorem[Limiting distribution] Let $\{y_t\}_{t \in \ensuremath{{\mathbb Z}}}$ and $\{x_t\}_{t \in \ensuremath{{\mathbb Z}}}$ be generated by (ref) and (ref) with $k=1$, let $\{w_t\}_{t \in \ensuremath{{\mathbb Z}}}$ satisfy Assumption (ref), and assume $T^{-2}\lambda_T \to 0$. \begin{enumerate} • If $\lambda_T \to \lambda_0$, $0 \leq \lambda_0 < \infty$, and $T\beta_T \to \beta_0 \in \ensuremath{\overline{\ensuremath{{\mathbb R}}}}$, then \begin{align*} T(\ensuremath{\hat\beta_{\rm\tiny AL}} - \beta_T) \Rightarrow & \ensuremath{\mathbbm{1}}\left\{(\zeta_{vv}^c)^{1/2}\left|\mathcal{Z}^c + \beta_0\right| > \sqrt{\frac{\lambda_0}{2}}\right\} \left(\mathcal{Z}^c - \frac{\lambda_0}{2\zeta_{vv}^c}(\mathcal{Z}^c + \beta_0)^{-1}\right) \\ & - \ensuremath{\mathbbm{1}}\left\{(\zeta_{vv}^c)^{1/2}\left|\mathcal{Z}^c + \beta_0\right| \leq \sqrt{\frac{\lambda_0}{2}}\right\}\beta_0. \end{align*} • If $\lambda_T \to \infty$ and $\lambda_T^{-1/2}T\beta_T \to \tilde\beta_0 \in \ensuremath{\overline{\ensuremath{{\mathbb R}}}}$, then $\lambda_T^{-1/2}T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta_T) \Rightarrow$ \begin{align*} \begin{cases} -\ensuremath{\mathbbm{1}}\left\{(\zeta_{vv}^c)^{1/2} > a_0\right\}(2\tilde\beta_0\zeta_{vv}^c)^{-1} -\ensuremath{\mathbbm{1}}\left\{(\zeta_{vv}^c)^{1/2} \leq a_0\right\}\tilde\beta_0 & if 0 < |\tilde\beta_0| < \infty\\ 0 & otherwise, \end{cases} \end{align*} where $a_0 \coloneqq 1/(\sqrt{2}|\tilde\beta_0|)$. \end{enumerate}

From Theorem (ref)(ref), we learn that under conservative tuning, if $\beta_T$ is bounded away from zero or converges to zero at rate slower than $T^{-1}$, i.e., $|\beta_0| = \infty$, the limiting distribution of $T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta_T)$ is $\mathcal{Z}^c$ and thus coincides with the limiting distribution of OLS. If $\beta_T \equiv 0$ or $\beta_T$ converges to zero at rate $T^{-1}$ or faster, i.e., $\beta_0 \in \ensuremath{{\mathbb R}}$, the limiting distribution consists of an atomic as well as an absolutely continuous part.

Part (ref) of the above theorem shows that under consistent tuning and using the correct scaling factor, the limit of $\lambda_T^{-1/2}T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta_T)$ collapses to zero if $\beta_T \equiv 0$ or $\beta_T$ converges to zero faster than $T\lambda_T^{-1/2}$, i.e. $\tilde\beta_0 = 0$, or if $\beta_T$ is bounded away from zero or converges to zero slower than $T\lambda_T^{-1/2}$, i.e., $|\tilde\beta_0| = \infty$. However, if $\beta_T$ converges to zero exactly at rate $T^{-1}\lambda_T^{1/2}$, i.e., $0 < |\tilde\beta_0| < \infty$, the limit is random and contains an atomic as well as an absolutely continuous part. Interestingly, all remaining randomness originates from the regressor $x_t$, but not from the errors $u_t$. The dependence on $u_t$ disappears because the influence of the rate-$T$ consistent OLS estimator vanishes asymptotically if $\hat \beta - \beta_T$ is scaled by $\lambda_T^{-1/2}T$ rather than $T$, whereas the dependence on $x_t$ appears in the limit through $\tilde\lambda_T$.\footnote{For more details, we refer to Lemma (ref)(c) in Appendix (ref).}

remarkTheorem (ref)(b) shows that the uniform convergence rate for the adaptive Lasso estimator under consistent tuning is, indeed, $T^{-1}\lambda_T^{1/2}$ and that scaling the estimation error with the larger factor $T$ will result in a stochastically unbounded sequence if $\beta_T$ converges to zero at rate $T^{-1}\lambda_T^{1/2}$. For completeness, we also list the limiting distribution of $T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta_T)$ for arbitrary sequences of $\beta_T$: If $\lambda_T \to \infty$, such that $T^{-2}\lambda_T \to 0$, and $\lambda_T^{-1/2}T\beta_T \to \tilde\beta_0 \in \ensuremath{\overline{\ensuremath{{\mathbb R}}}}$, then \begin{align*} T(\ensuremath{\hat\beta_{\rm\tiny AL}} - \beta_T) \Rightarrow \begin{cases} -\beta_0 & if \tilde\beta_0 = 0 \\ -\rm sign(\tilde\beta_0)\infty & if 0 < |\tilde\beta_0| < \infty \\ \mathcal{Z}^c - 0.5(\zeta_{vv}^c\bar\beta_0)^{-1} & if |\tilde\beta_0| = \infty, \end{cases} \end{align*} where $T\beta_T \to \beta_0 \in \ensuremath{\overline{\ensuremath{{\mathbb R}}}}$ and $\lambda_T^{-1}T\beta_T \to \bar\beta_0 \in \ensuremath{\overline{\ensuremath{{\mathbb R}}}}$. Hence, $T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta_T)$ collapses to pointmass at $-\beta_0$ whenever $\beta_T \equiv 0$ or $\beta_T$ converges to zero at rate $T^{-1}$ or faster, i.e., $\beta_0 \in \ensuremath{{\mathbb R}}$ and $\tilde\beta_0 = 0$. The limiting distribution of $T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta_T)$ is random if $\beta_T$ is bounded away from zero or converges to zero at rate $T^{-1}\lambda_T$ or slower, i.e., $\bar\beta_0 \neq 0$ and $|\tilde\beta_0| = \infty$. In this case, the limiting distribution coincides with the one of OLS if $\beta_T$ converges to zero slower than $T^{-1}\lambda_T$, i.e., $|\bar\beta_0| = |\tilde\beta_0| = \infty$. However, if $|\bar\beta_0| < \infty$, the limiting distribution of the adaptive LASSO estimator deviates from the one of OLS by a random shift that is inversely proportional to and has the opposite sign of $\bar\beta_0$, analogously to what we have seen under fixed-parameter asymptotics in Proposition (ref)(b1). Again, the random shift is not detected in LeeEtAl22, and, in contrast to the second order bias term in the limiting distribution of the OLS estimator, it does not vanish if the regressor is exogeneous. In all other cases, the total mass of $T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta_T)$ escapes to $-\infty$ or $\infty$.
remarkRemark (ref) illustrates that in the consistently tuned case, while the adaptive Lasso estimator can detect local-to-zero rates of any order greater than $T^{-1}\lambda_T^{1/2}$, in order to obtain the same limiting distribution as OLS, the true coefficient must be of even larger order of magnitude, i.e., greater than $T^{-1}\lambda_T$. In the setting of LeeEtAl22 with $\beta_T = \beta T^{-\delta}$ for $\delta \in (0,1)$, $\beta \neq 0$, and $T^{-(1-\delta)}\lambda_T \to 0$, it automatically holds that $\lambda_T^{-1}T\beta_T \to \infty$.

The Multivariate Case

We now turn to the multivariate regressor case and investigate the asymptotic properties of the adaptive LASSO estimator within a moving-parameter framework. Since detailed results for the fixed-parameter framework have already been presented and discussed in the univariate case, our focus here remains on the more general setting, where the true coefficients are allowed to vary with sample size, see also the discussion at the beginning of Section (ref). With respect to notation, please note that the subscript $j$ continues to denote the $j$-th element of the vector to which it is attached, e.g., $\ensuremath{\hat\beta_{{\scriptstyle\text{\rm\tiny AL}},j}}$ denotes the $j$-th component of $\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}}$.

We start by deriving model selection probabilities for the adaptive LASSO estimator under conservative as well as consistent tuning for certain relevant sequences of $\beta_T$.

theorem[Model selection] Let $\{y_t\}_{t \in \ensuremath{{\mathbb Z}}}$ and $\{x_t\}_{t \in \ensuremath{{\mathbb Z}}}$ be generated by (ref) and (ref), and let $\{w_t\}_{t \in \ensuremath{{\mathbb Z}}}$ satisfy Assumption (ref). \begin{enumerate} • If $\lambda_T \to \lambda_0$, $0 \leq \lambda_0 < \infty$, and $T\beta_T \to \beta_0 \in \ensuremath{\overline{\ensuremath{{\mathbb R}}}}^k$, then $$ \ensuremath{{\mathbb P}}\left(\ensuremath{\hat\beta_{{\scriptstyle\text{\rm\tiny AL}},j}} = 0\right) \to 0, $$ if $|\tilde\beta_{0,j}| = \infty$. • If $\lambda_T \to \infty$ and $\lambda_T^{-1/2}T\beta_T \to \tilde\beta_0 \in \ensuremath{\overline{\ensuremath{{\mathbb R}}}}^k$ then $$ \ensuremath{{\mathbb P}}\left(\ensuremath{\hat\beta_{{\scriptstyle\text{\rm\tiny AL}},j}} = 0\right) \to \begin{cases} 1 & \text{ if } \tilde\beta_{0,j} = 0 \\ 0 & \text{ if } |\tilde\beta_{0,j}| = \infty. \end{cases} $$ \end{enumerate}

Before we discuss the results in Theorem (ref) in detail, we extend the statement in Theorem (ref)(ref) in the following remark to obtain a more comprehensive picture for the model selection properties in the conservatively tuned case.

remarkWe point out two additional special cases for the model selection properties in the conservatively tuned case. Let $\lambda_T \to \lambda_0$, $0 \leq \lambda_0 < \infty$. Then: \begin{enumerate} • For $T\beta_T \to \beta_0 \in \ensuremath{\overline{\ensuremath{{\mathbb R}}}}^k$ with $\beta_{0,j} = 0$ we have $$ \liminf_{T \rightarrow \infty} \ensuremath{{\mathbb P}}\left(\ensuremath{\hat\beta_{{\scriptstyle\text{\rm\tiny AL}},j}} = 0\right) > 0. $$ • For $\beta_T \equiv \beta \in \ensuremath{{\mathbb R}}^k$, $\mathcal{A} \coloneqq \{j: \beta_j \neq 0\}$, and $\hat\mathcal{A} \coloneqq \{j: \ensuremath{\hat\beta_{{\scriptstyle\text{\rm\tiny AL}},j}} \neq 0\}$ we have $$ \limsup_{T \rightarrow \infty} \ensuremath{{\mathbb P}}\left(\hat\mathcal{A} = \mathcal{A} \right) < 1. $$ \end{enumerate}

In line with the results from the univariate case, Theorem (ref)(ref) shows that under conservative tuning, the estimator can detect coefficients with local-to-zero rates of order greater than $T^{-1}$ as non-zero with asymptotic probability equal to one. Importantly, in the multivariate case, this property depends solely on the rate of the coefficient under consideration and is unaffected by the behavior of the other components of $\beta_T$. Moreover, coefficients converging to zero with rate faster than $T^{-1}$ will be set to zero with positive asymptotic probability as can be seen in Remark (ref)(ref). The smallest detectable local-to-zero rate under conservative tuning therefore remains $T^{-1}$. Remark (ref)(ref) illustrates that this tuning regime is indeed conservative, also in the multivariate case.

For the consistently tuned case, a meaningful interpretation again requires the basic condition $T^{-2}\lambda_T \to 0$, which we assume to hold throughout this section. As in the univariate case, Theorem (ref)(ref) then reveals that the estimator detects coefficients with local-to-zero rates of order greater than $T^{-1}\lambda_T^{1/2}$ as non-zero with asymptotic probability equal to one, while coefficients converging to zero with rate faster than $T^{-1}\lambda_T^{1/2}$ will always be set to zero with asymptotic probability equal to one. As under conservative tuning, these properties only depend on the rate of the coefficient under consideration and are not affected by the behavior of the remaining components of $\beta_T$. Hence, in the multivariate setting the smallest detectable local-to-zero rate under consistent tuning continues to be $T^{-1}\lambda_T^{1/2}$.

We turn to parameter estimation consistency in the following theorem.

theorem[Parameter estimation] Let $\{y_t\}_{t \in \ensuremath{{\mathbb Z}}}$ and $\{x_t\}_{t \in \ensuremath{{\mathbb Z}}}$ be generated by (ref) and (ref), let $\{w_t\}_{t \in \ensuremath{{\mathbb Z}}}$ satisfy Assumption (ref), and assume $T^{-2}\lambda_T \to 0$. \begin{enumerate} • It holds that $\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta_T = o_\ensuremath{{\mathbb P}}(1)$. • If $\lambda_T \to \lambda_0$, $0 \leq \lambda_0< \infty$, then $T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta_T) = O_\ensuremath{{\mathbb P}}(1)$. • If $\lambda_T \to \infty$, then $\lambda_T^{-1/2}T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta_T) = O_\ensuremath{{\mathbb P}}(1)$. \end{enumerate}

Theorem (ref) shows that if $T^{-2}\lambda_T \to 0$, the adaptive LASSO estimator is uniformly consistent also in the multivariate case and its uniform convergence rate depends on the tuning regime. Under conservative tuning, the estimator is rate-$T$ consistent, whereas under consistent tuning, the rate decreases to $T\lambda_T^{-1/2}$, just as in the univariate case.

We now derive the limiting distribution of the adaptive LASSO estimator under arbitrary sequences of $\beta_T$.

theorem[Limiting distribution] Let $\{y_t\}_{t \in \ensuremath{{\mathbb Z}}}$ and $\{x_t\}_{t \in \ensuremath{{\mathbb Z}}}$ be generated by (ref) and (ref), let $\{w_t\}_{t \in \ensuremath{{\mathbb Z}}}$ satisfy Assumption (ref), and assume $T^{-2}\lambda_T \to 0$. \begin{enumerate} • If $\lambda_T \to \lambda_0$, $0 \leq \lambda_0 < \infty$ and $T\beta_T \to \beta_0 \in \ensuremath{\overline{\ensuremath{{\mathbb R}}}}^k$, then $T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta_T) \Rightarrow \operatorname{argmin}_{z \in \ensuremath{{\mathbb R}}^k} V_{\beta_0}^c(z)$, where $$ V_{\beta_0}^c(z) \coloneqq z'\zeta_{vv}^cz - 2z'\left(\int_0^1 J_v^c(r)dB_u(r) + \Delta_{vu}\right) + \lambda_0 \sum_{j=1}^k A_j(z_j,\beta_{0,j}) $$ and $$ A_j(z_j,\beta_{0,j}) \coloneqq \begin{cases} 0 & \text{ if } |\beta_{0,j}| = \infty \text{ or } z_j = 0 \\ \frac{|z_j|}{|\mathcal{Z}^c_j|} & \text{ if } \beta_{0,j} = 0 \text{ and } z_j\neq 0 \\ \frac{|\beta_{0,j} + z_j| - |\beta_{0,j}|}{|\beta_{0,j} + \mathcal{Z}^c_j|} & \text{ otherwise.} \end{cases} $$ • If $\lambda_T \to \infty$ and $\lambda_T^{-1/2}T\beta_T \to \tilde\beta_0 \in \ensuremath{\overline{\ensuremath{{\mathbb R}}}}^k$, then $\lambda_T^{-1/2}T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta_T) \Rightarrow \operatorname{argmin}_{z \in \ensuremath{{\mathbb R}}^k} \tilde{V}_{\tilde\beta_0}^c(z)$, where \begin{align*} \tilde{V}_{\tilde\beta_0}^c(z) \coloneqq z'\zeta_{vv}^cz + \sum_{j=1}^k \tilde{A}_j(z_j,\tilde\beta_{0,j}) \end{align*} and $$ \tilde{A}_j(z_j,\tilde\beta_{0,j}) \coloneqq \begin{cases} 0 & \text{ if } |\tilde\beta_{0,j}| = \infty \text{ or } z_j = 0 \\ \infty & \text{ if }\tilde\beta_{0,j} = 0 \text{ and } z_j \neq 0 \\ \frac{|z_j + \tilde\beta_{0,j}|}{|\tilde\beta_{0,j}|} - 1 & \text{ otherwise.} \end{cases} $$ • If $\lambda_T \to \infty$, $T\beta_T \to \beta_0 \in \ensuremath{\overline{\ensuremath{{\mathbb R}}}}^k$, and $\lambda_T^{-1}T\beta_T \to \bar\beta_0 \in \ensuremath{\overline{\ensuremath{{\mathbb R}}}}^k$, then $T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta_T) \Rightarrow \operatorname{argmin}_{z \in \ensuremath{{\mathbb R}}^k} \bar{V}_{\bar\beta_0}^c(z)$, where $$ \bar{V}_{\bar\beta_0}^c(z) \coloneqq z'\zeta_{vv}^cz - 2z'\left(\int_0^1 J_v^c(r)dB_u(r) + \Delta_{vu}\right) + \sum_{j=1}^k \bar{A}_j(z_j,\beta_{0,j},\bar\beta_{0,j}) $$ and $$ \bar{A}_j(z_j,\beta_{0,j},\bar{\beta}_{0,j}) \coloneqq \begin{cases} 0 & \text{ if } |\bar\beta_{0,j}| = \infty \text{ or } z_j = 0 \\ \infty & \text{ if } \bar\beta_{0,j} = \beta_{0,j} = 0 \text{ and } z_j \neq 0 \\ \text{\rm sign}(z_j + 2\beta_{0,j})\text{\rm sign}(z_j)\infty & \text{ if } \bar\beta_{0,j} = 0, \, 0 < |\beta_{0,j}| < \infty, \text{ and } z_j \neq 0 \\ \text{\rm sign}(z_j)\text{\rm sign}(\beta_{0,j})\infty & \text{ if } \bar\beta_{0,j} = 0, \, |\beta_{0,j}| = \infty, \text{ and } z_j \neq 0 \\ \frac{-\text{\rm sign}(\bar\beta_{0,j})z_j}{|\bar\beta_{0,j}|} & \text{ otherwise.} \end{cases} $$ \end{enumerate}

The limiting distributions presented in Theorem (ref) are defined implicitly. While we cannot explicitly minimize $V_{\beta_0}^c(z)$, $\tilde{V}_{\tilde\beta_0}^c(z)$, and $\bar{V}_{\bar\beta_0}^c(z)$ for fixed $\beta_0$, $\tilde\beta_0$, and $\bar\beta_0$ in general, there are a number of special cases worth pointing out. First, in line with Theorem (ref)(ref), Theorem (ref)(ref) shows that under conservative tuning with either $\lambda_0 = 0$ or $|\beta_{0,j}| = \infty$ for all $j = 1,\ldots,k$, the limiting distribution of $T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta_T)$ is $\mathcal{Z}^c$ and thus coincides with the limiting distribution of OLS. Second, in line with Theorem (ref)(ref), part (ref) shows that under consistent tuning, the limit of $\lambda_T^{-1/2}T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta_T)$ collapses to zero whenever $\tilde\beta_{0,j} = 0$ or $|\tilde\beta_{0,j}| = \infty$ for all $j = 1,\ldots,k$. Moreover, whenever the limit is stochastic, all randomness originates from the regressors $x_t$, but not from the errors $u_t$ (see also the discussion in the univariate case). Finally, part (ref) reveals that under consistent tuning, the limiting distribution of $T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta_T)$ coincides with the limiting distribution of OLS if $|\bar{\beta}_{0,j}| = \infty$ for all $j = 1,\ldots,k$. Conversely, the limit of $T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta_T)$ collapses to zero whenever $|\bar\beta_{0,j}| = 0$ for all $j = 1,\ldots,k$. Both results are consistent with Remark (ref).

remarkA similar comment as in Remark (ref) also applies to Theorems (ref)--(ref).

The following proposition offers additional insights into the limiting distribution of $\lambda_T^{-1/2}T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta_T)$ under consistent tuning, as derived in Theorem (ref)(ref).

propositionFor a fixed $\omega$ in the sample space of the underlying probability space, the point $m = m(\omega) \in \ensuremath{{\mathbb R}}^k$ is a minimizer of $\tilde{V}_{\tilde\beta_0}^c(z)(\omega)$ if and only if \begin{align*} \begin{cases} m_j = 0 & if \tilde\beta_{0,j} = 0 \\ (\zeta_{vv}^c(\omega)m)_j = 0 & if |\tilde\beta_{0,j}| = \infty \\ (\zeta_{vv}^c(\omega)m)_j = - \frac{\rm sign(m_j+\tilde\beta_{0,j})}{2|\tilde\beta_{0,j}|} & if 0 < |\tilde\beta_{0,j}| < \infty and m_j\neq - \tilde\beta_{0,j} \\ |(\zeta_{vv}^c(\omega)m)_j| \leq \frac{1}{2|\tilde\beta_{0,j}|} & if 0 < |\tilde\beta_{0,j}| < \infty \text{ and } m_j = - \tilde\beta_{0,j}, \end{cases} \end{align*} where $\zeta_{vv}^c$ is the same random matrix as in the definition of $\tilde{V}_{\tilde\beta_0}^c(z)$ in Theorem (ref)(ref).

Building on Proposition (ref), the following theorem shows that the set of minimizers of $\tilde{V}_{\tilde\beta_0}^c(z)$ taken over over all $\tilde\beta_0 \in \ensuremath{\overline{\ensuremath{{\mathbb R}}}}^k$ is contained in a random set that does not depend on $\tilde\beta_0$.

theoremDefine the random set $$ \mathcal{M}^c \coloneqq \left\{m\in\ensuremath{{\mathbb R}}^k:m_j(\zeta_{vv}^cm)_j\leq \frac{1}{2},\,j=1,\ldots,k\right\}, $$ where $\zeta_{vv}^c$ is the same as in the definition of $\tilde{V}_{\tilde\beta_0}^c(z)$ in Theorem (ref)(ref). Then, $$ \underset{\tilde\beta_0 \in \ensuremath{\overline{\ensuremath{{\mathbb R}}}}^k}{\bigcup} \operatorname{argmin}_{z \in \ensuremath{{\mathbb R}}^k}\tilde V_{\tilde\beta_0}^c(z) \subseteq \mathcal{M}^c, $$ where the set inclusion holds surely, i.e., for all $\omega$ in the sample space of the underlying probability space.
remarkFor later use, we will assume that the random matrix $\zeta_{vv}^c$ in Theorem (ref)(ref) satisfies $\lim_{T \to \infty} T^{-2} \sum_{t=1}^Tx_t(\omega)x_t'(\omega) = \zeta_{vv}^c(\omega)$ for all $\omega$. This can be achieved using Skorohod's representation theorem.

Theorem (ref) together with Remark (ref) shows that the union of limits of $\lambda_T^{-1/2} T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta_T)$ over all possible parameter sequences is contained in the set $\mathcal{M}^c$, which is a compact set for each realization of $\zeta_{vv}^c$. This observation allows us to construct uniformly valid confidence regions centered at the adaptive LASSO estimator, as developed in the following subsection. Before doing so, let us examine the set $\mathcal{M}^c$ in more detail.

Clearly, all randomness in $\mathcal{M}^c$ stems from the regressors $x_t$, i.e., no randomness arises from the regression errors $u_t$. For expositional convenience, we focus on the pure unit root case ($c=0$): There, $\zeta_{vv}^c$ has expectation $0.5\,\Omega_{vv}$, where $\Omega_{vv}$, given by the $k \times k$ bottom-right block of $\Omega$, the long-run covariance matrix of $\{v_t\}_{t \in \mathbb{Z}}$. Hence, on average, $\mathcal{M}^c$ is given by $\{m\in \ensuremath{{\mathbb R}}^k : m_j (\Omega_{vv} m)_j \leq 1,\,j=1,\ldots,k \}$. Thus, on average, the set $\mathcal{M}^c$ becomes smaller as the variability of $\{v_t\}_{t \in \mathbb{Z}}$ increases. In the univariate regressor case, $\Omega_{vv}$ reduces to the long-run variance of $v_t$. Normalizing this variance to one implies that, on average, $\mathcal{M}^c$ coincides with the interval $[-1,1]$. Consequently, in one dimension and on average, we recover the same interval as PoetscherSchneider09 and AmannSchneider23. Note, however, that the corresponding sets in these two papers are non-random, as the regressors are treated as deterministic.

A Universal Confidence Region Under Consistent Tuning

We now use the observation from Theorem (ref) to construct a confidence region that has asymptotic coverage probability equal to one. To this end, we define a “slightly larger” finite-sample analogue $\widehat{\mathcal{M}}_T(\ensuremath{\varepsilon}) \coloneqq \{m\in\ensuremath{{\mathbb R}}^k : m_j((T^{-2}\sum_{t=1}^Tx_tx_t')m)_j \leq \frac{1}{2} + \ensuremath{\varepsilon},\,j=1,\ldots,k\}$ of $\mathcal{M}^c$, where $\ensuremath{\varepsilon} > 0$ but arbitrarily small. The following theorem shows that this set can be used to construct a confidence region based on the adaptive LASSO estimator that asymptotically holds any prescribed coverage level.

theorem[Confidence regions] Let $\{y_t\}_{t \in \ensuremath{{\mathbb Z}}}$ and $\{x_t\}_{t \in \ensuremath{{\mathbb Z}}}$ be generated by (ref) and (ref), let $\{w_t\}_{t \in \ensuremath{{\mathbb Z}}}$ satisfy Assumption (ref), and let $T^{-2}\lambda_T \to 0$ and $\lambda_T \to \infty$. Then $$ \lim_{T \to \infty} \inf_{\beta \in \ensuremath{{\mathbb R}}^k} \ensuremath{{\mathbb P}}_\beta\left(\beta \in \ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - T^{-1}\lambda_T^{1/2}\widehat{\mathcal{M}}_T(\ensuremath{\varepsilon})\right) = 1 $$ for any $\ensuremath{\varepsilon} > 0$.

Theorem (ref) delivers valid confidence regions, as the coverage probability holds uniformly over the parameter space. Technically, this is achieved by taking the infimum over the parameter space prior to letting $T$ tend to infinity. The key underlying idea is that $\beta \in \{\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \lambda_T^{1/2}T^{-1}m : m\in \widehat{\mathcal{M}}_T(\ensuremath{\varepsilon})\}$ holds if and only if $\lambda_T^{-1/2}T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta) \in \widehat{\mathcal{M}}_T(\ensuremath{\varepsilon})$. The latter event can then be approximated by $\operatorname{argmin}_{z \in \ensuremath{{\mathbb R}}^k}\tilde V_{\tilde\beta_0}^c(z) \in \mathcal{M}^c$ for some $\tilde\beta_0 \in \ensuremath{\overline{\ensuremath{{\mathbb R}}}}^k$, and this inclusion surely holds for any $\tilde\beta_0$ by Theorem (ref).

In practice, we propose to construct the confidence region for $\beta$ based on $\widehat{\mathcal{M}}_T(0)$. Since no closed-form solution is available, $\widehat{\mathcal{M}}_T(0)$ can be computed numerically following the the description below, and the confidence region is then given by $\{\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \lambda_T^{1/2}T^{-1}m : m\in \widehat{\mathcal{M}}_T(0)\}$.

If confidence intervals for individual components $\beta_j$ are of interest, these can be obtained by $$ [\ensuremath{\hat\beta_{{\scriptstyle\text{\rm\tiny AL}},j}} - \lambda_T^{1/2}T^{-1}\overline{m}_j, \ensuremath{\hat\beta_{{\scriptstyle\text{\rm\tiny AL}},j}} - \lambda_T^{1/2}T^{-1}\underline{m}_j], $$ where $\overline{m}_j\coloneqq \max\{m_j : m\in \widehat{\mathcal{M}}_T(0)\}$ and $\underline{m}_j\coloneqq \min\{m_j : m\in \widehat{\mathcal{M}}_T(0)\}=-\overline{m}_j$. The quantity $\overline{m}_j$ can be computed numerically using sequential quadratic programming without explicitly constructing the set $\widehat{\mathcal{M}}_T(0)$. Specifically, we solve the constrained maximization problem defining $\overline{m}_j$ directly. To reduce the risk of convergence to local optima, the algorithm can be initialized from multiple random starting values, retaining the largest value of $m_j$ obtained across these runs.

The coordinate-wise intervals can also be used to construct the set $\widehat{\mathcal{M}}_T(0)$. Since $\widehat{\mathcal{M}}_T(0)$ is contained in the Cartesian product of these intervals, i.e., $\widehat{\mathcal{M}}_T(0) \subseteq \mathcal{B} \coloneqq \prod_{j=1}^{k}[\underline{m}_j, \overline{m}_j]$, one can either construct a grid within $\mathcal{B}$ or, to speed up computation for large $k$, randomly sample vectors from $\mathcal{B}$, retaining only those that satisfy the constraints defining $\widehat{\mathcal{M}}_T(0)$.\footnote{MATLAB code for computing $\overline{m}_j$ and $\underline{m}_j$ for all $j=1,\ldots,k$, as well as $\widehat{\mathcal{M}}_T(0)$ for general $k \times k$ symmetric positive definite matrices, is available on the first author's personal website.}

It may appear curious to consider confidence regions whose asymptotic coverage probability equals one. To explain this, recall that the consistently tuned adaptive LASSO estimator exhibits the behavior that only randomness stemming from the regressors $x_t$ can persist asymptotically. This occurs because the scaling induced by the uniform convergence rate is not sufficiently large for stochastic variation from the error terms $u_t$ to survive in the limit. When the regressors are treated as non-random, this even manifests in entirely non-random limits which are contained in a compact set. AmannSchneider23 show that this non-random set can be utilized in a similar way to construct confidence regions with uniform asymptotic coverage probability equal to one, but that any slightly smaller region has asymptotic coverage probability equal to zero.

When the regressors $x_t$ are random, a slightly different picture arises. All possible limits of $\lambda_T^{-1/2}T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta)$ are contained in a set that is compact for a fixed realization of the limiting regressor matrix $\zeta_{vv}^c$. As a result, confidence regions can be constructed that still achieve uniform asymptotic coverage equal to one. However, it cannot be shown anymore that this probability will drop to zero when the regions are made smaller, an effect of $x_t$ being stochastic.

It may be possible to exploit the remaining randomness to construct confidence regions with asymptotic coverage strictly less than one, but we leave a formal investigation of this question for future research. Nevertheless, the regions proposed here possess a key advantage: they do not rely on the asymptotic distribution of either the adaptive LASSO estimator or the rescaled regressors. As a consequence, their construction avoids the need for any knowledge or estimation of local-to-unity and long-run covariance parameters, as well as for accounting for second-order bias terms in the limiting distribution of the adaptive LASSO estimator. This universality distinguishes our approach from existing methods and, to the best of our knowledge, represents the first construction of LASSO-based uniformly valid confidence regions in regressions with unit root or local-to-unity regressors.

Simulation Results

This section presents simulation results. Section (ref) analyzes the approximation quality of the asymptotic results to the finite-sample distribution of the adaptive LASSO estimator, whereas Section (ref) focuses on the empirical coverage probabilities of the uniform confidence regions.

Finite-Sample Distributions

We investigate the approximation quality of our theoretical results to the finite-sample distribution of the adaptive LASSO estimator under both conservative and consistent tuning for various sequences $\beta_T$ and different sample sizes. We also compare the finite-sample distribution of the adaptive LASSO to that of OLS to analyze how much it deviates from what is suggested by the oracle property in empirically relevant scenarios where some coefficients are small rather than exactly equal to zero.

We generate data according to (ref) and (ref) for the univariate unit root case with $[u_t,v_t]'\sim\mathcal{N}\left(0,I_2\right)$ i.i.d. across $t$, and present results for $\beta_T\in\{0.1\beta, \beta/T^{1/2}, \beta/T, \lambda_T^{1/2}\beta/T\}$, with $\beta=1$, $\lambda_T\in\{1,T^{1/4},T^{1/2},T\}$, and $T\in\{25,50,100,250,1000\}$. All results are based on $10{,}000$ Monte Carlo replications.\footnote{To focus on the main effects, we omit error serial correlation and regressor endogeneity from the model. While our empirical findings remain qualitatively similar when these features are included -- as well as under changes in the variances of $u_t$ and $v_t$ or deviations from normality -- the approximation quality of the limiting distributions derived in Theorem (ref) and Remark (ref) may be reduced, particularly in small- to medium-sized samples.} The choices for $\beta_T$ cover the cases where $\beta_T$ is bounded away from zero, converges to zero at rate $T^{1/2}$ LeeEtAl22, converges to zero at rate $T^{-1}$ (the cut-off rate under conservative tuning), and converges to zero at the slower rate $T^{-1}\lambda_T^{1/2}$ (the cut-off rate under consistent tuning). With respect to the tuning parameter $\lambda_T$, the choice $\lambda_T\equiv1$ leads to conservative tuning, while the other three choices lead to consistent tuning. Importantly, in case $\beta_T=\beta/T^{1/2}$, only $\lambda_T=T^{1/4}$ fulfills the condition in LeeEtAl22 that $T^{-1/2}\lambda_T + \lambda_T^{-1} \to 0$.

Separately for the four choices of $\beta_T$, Figures (ref)--(ref) display the finite-sample distributions of $T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta_T)$ (under conservative tuning) and $\lambda_T^{-1/2}T(\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - \beta_T)$ (under consistent tuning). The distributions consist of an atomic mass, drawn at the height corresponding to the relative frequency $p$ of the event $\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} = 0$, and a continuous component (rescaled to integrate to $1 - p$), representing the density of the non-zero estimates.\footnote{All displayed densities are smoothed by using the Gaussian kernel.} The figures also display the finite-sample distribution of the OLS estimator, $T(\hat\beta - \beta_T)$, as well as the case-specific limiting distribution of the adaptive LASSO estimator from Theorem (ref) evaluated at $\beta_{0,T} \coloneqq T\beta_T$ (under conservative tuning) and $\tilde\beta_{0,T} \coloneqq \lambda_T^{-1/2}T\beta_T$ (under consistent tuning).\footnote{Densities of limiting distributions are obtained by simulation, where Brownian motions are approximated by normalized sums of $10{,}000$ i.i.d. standard normal random variables and stochastic integrals are approximated accordingly.} Replacing the limiting parameters $\beta_0$ and $\tilde \beta_0$ with their finite-sample counterparts allows us to account for the size of $\beta_T$ relative to the sample size when evaluating the approximation quality of the limiting distributions derived in Theorem (ref) for the finite-sample distributions.

Figure (ref) presents the results for $\beta_T \equiv 0.1\beta$. In general, the adaptive LASSO estimator identifies the true coefficient as non-zero with probability approaching one as sample size $T$ increases. However, by construction, the empirical probability of incorrectly setting the non-zero coefficient to zero increases with the order at which $\lambda_T$ diverges. Under conservative tuning, the finite-sample distribution of the estimator approaches that of the OLS estimator, with the two distributions becoming virtually indistinguishable already for $T = 100$. Notably, the limiting distribution derived in Theorem (ref)(a) evaluated at $\beta_{0,T}$ already provides a good approximation to the finite-sample distribution of the adaptive LASSO estimator for small $T$, e.g., $T = 25$. In the context of Lemma (ref) in Appendix (ref), this indicates that $\mathcal{Z}^c_T$ and $\zeta_{vv,T}$ converge quickly to their asymptotic counterparts, such that the finite-sample distribution of the procedure is effectively governed by $\beta_{0,T}$. Under consistent tuning, the figure shows that the scaling factor implied by the uniform rate causes all mass of the distribution of the adaptive LASSO estimator to collapse at zero as $T$ increases. Moreover, the limiting distribution derived in Theorem (ref)(ref), evaluated at $\tilde\beta_{0,T}$, approximates the finite-sample distribution more accurately the larger the order of $\lambda_T$. For $\lambda_T = T$, the approximation is already quite accurate for small $T$.

We now turn to Figure (ref), which presents the results for $\beta_T=\beta/T^{1/2}$. Under conservative tuning, the adaptive LASSO estimator still identifies the smaller coefficient as non-zero with probability approaching one as $T$ increases and its distribution quickly approaches that of the OLS estimator. Under consistent tuning, however, its properties now depend on the order of $\lambda_T$. As before, for $\lambda_T\in\{T^{1/4}, T^{1/2}\}$, the procedure correctly identifies the true coefficient as non-zero with probability approaching one and the scaling factor implied by the uniform rate causes all mass to collapse at zero as $T$ increases. For $\lambda_T = T$, however, the estimator incorrectly sets the true coefficient to zero with probability approaching $0.68$. As a result, even for large $T$, the distribution of the estimator consists of both an atomic and a continuous part. Nevertheless, the limiting distribution derived in Theorem (ref)(ref), evaluated at $\tilde\beta_{0,T}$, provides a good approximation to the finite-sample distribution even for small $T$.

Figure (ref) presents the results for $\beta_T = \beta/T$, which is the cut-off rate for detection of non-zero coefficients under conservative tuning. As a result, under conservative tuning, the adaptive LASSO estimator incorrectly sets the true coefficient to zero with probability approaching $0.43$ and its distribution -- which is approximated very well by the limiting distribution derived in Theorem (ref)(ref), evaluated at $\beta_{0,T}$, already for small $T$ -- consists of both an atomic and a continuous part. Under consistent tuning, the estimator incorrectly sets the true coefficient to zero with probability approaching one, such that its distribution collapses to an atomic part at $-\tilde\beta_{0,T}$.

Finally, Figure (ref) presents the results for $\beta_T = \sqrt{\lambda_T}\beta/T$, which is the cut-off rate for the detection of non-zero coefficients under consistent tuning.\footnote{Since $\lambda_T \equiv 1$ implies $\beta_T = \beta/T$, the results under conservative tuning coincide with those shown in Figure (ref).} As a result, under consistent tuning, the adaptive LASSO estimator incorrectly sets the true coefficient to zero with probability approaching $0.68$ and its distribution consists of both an atomic and a continuous part. The limiting distribution derived in Theorem (ref)(ref), evaluated at $\tilde\beta_{0,T}$, approximates the finite-sample distribution of the estimator more accurately the larger the order of $\lambda_T$. For $\lambda_T = T$, the approximation is already quite accurate for small $T$.

Now, we analyze the finite-sample distribution of $T(\hat\beta_{{\scriptstyle\text{\rm\tiny AL}}} - \beta_T)$ under consistent tuning, which is presented in Figures (ref)--(ref) in Appendix (ref) separately for the four choices of $\beta_T$ alongside the corresponding limits from Remark (ref) and the finite-sample distribution of the OLS estimator.

Figure (ref) presents the results for $\beta_T \equiv 0.1\beta$. As seen previously in Figure (ref), the adaptive LASSO estimator identifies the true coefficient as non-zero with probability approaching one as $T$ increases. However, the order of $\lambda_T$ has a substantial impact on the estimator’s distribution. For $\lambda_T \in \{T^{1/4}, T^{1/2}\}$, its distribution approaches that of the OLS estimator. In contrast, for $\lambda_T = T$, the distribution approaches to that of the OLS estimator shifted to the left. This stochastically bounded shift significantly distorts the distribution even for small $T$. As a result, when $\lambda_T = T$, the adaptive LASSO estimator fails to exhibit the oracle property, despite correctly identifying the true coefficient as non-zero with probability approaching one.

We now turn to Figure (ref), which presents the results for $\beta_T = \beta/T^{1/2}$. For $\lambda_T \in \{T^{1/4}, T^{1/2}\}$, the adaptive LASSO estimator identifies the true coefficient as non-zero with probability approaching one as $T$ increases (c.f. Figure (ref)), but only for $\lambda_T = T^{1/4}$ does its distribution approach that of OLS. For $\lambda_T = T^{1/2}$, on the other hand, the distribution of the procedure approaches the one of OLS shifted to the left and it thus loses its oracle property. For $\lambda_T = T$, the adaptive LASSO estimator incorrectly sets the true coefficient to zero with probability approaching $0.68$ and its distribution consists of both an atomic part and a continuous part. As $T$ increases, both the location of the atomic part and the region where the continuous part of the distribution has mass shift toward $-\infty$. While the behavior of the estimator in case $\lambda_T = T^{1/4}$ is already described in LeeEtAl22, the cases $\lambda_T \in \{T^{1/2}, T\}$ are not covered by their results, but are in line with our asymptotic results in Remark (ref).

Figure (ref) presents the results for $\beta_T=\beta/T$. In this case, the adaptive LASSO estimator incorrectly sets the true coefficient to zero with probability approaching one (see also Figure (ref)), causing its distribution to collapse to an atomic part at $-\beta$. The collapse occurs more rapidly the larger the order at which $\lambda_T$ diverges.

Finally, Figure (ref) presents the results for $\beta_T = \sqrt{\lambda_T}\beta/T$. As this is the cut-off rate for the detection of non-zero coefficients under consistent tuning, the adaptive LASSO estimator incorrectly sets the true coefficient to zero with probability approaching $0.68$ (see also Figure (ref)), and its distribution consists of both an atomic and a continuous part. As $T$ increases, both the location of the atomic part and the region where the continuous part of the distribution has mass shift toward $-\infty$.

Overall, the simulation results reveal that the finite-sample distribution of the adaptive LASSO estimator can deviate substantially from what is suggested by the oracle property, especially under consistent tuning. By contrast, the limiting distributions derived in under moving-parameter asymptotics capture the finite-sample properties much more accurately. They not only provide reasonable approximations for the absolutely continuous part of the estimator's distribution but also convey information on the relative frequency with which the coefficient is set to zero.

Coverage Probabilities

We now study the coverage probabilities of the confidence regions introduced in Section (ref) and benchmark them against confidence regions based on the oracle property.

Data are generated as in the previous subsection. In the univariate case, the confidence region from Theorem (ref) simplifies to the interval $\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} \mp \sqrt{\lambda_T}/\sqrt{2\sum_{t=1}^T x_t^2}$, which we label the “Uniform CI”. The oracle-based interval (“Oracle CI”) is given by $[\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - q_{1-\alpha/2}/T, \ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} - q_{\alpha/2}/T]$, where $q_{\alpha}$ denotes the $\alpha$-quantile of $\mathcal{Z}^c$ defined in Equation (ref) in Section (ref). The quantiles are obtained by simulation as described in the previous subsection. We report results for $\alpha=0.05$ and $\alpha=0.01$, corresponding to nominal $95\%$ and $99\%$ oracle confidence intervals, respectively. In addition, we consider the interval $\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} \mp \sqrt{\lambda_T}/T$, labeled “asymptotic Uniform CI”, which replaces $T^{-2}\sum_{t=1}^T x_t^2$ in the definition of $\widehat{\mathcal{M}}_T(0)$ by the expectation of its limit $\zeta_{vv}^c$, which is equal to 1 in this case.

For each interval, we compute the empirical coverage probability for the true parameter value $\beta \in [-0.6,0.6]$ using an equidistant grid with step size $0.01$. As the results are symmetric in $\beta$, we report them only for $|\beta|$. Practitioners following the oracle property are typically interested in confidence intervals only when $\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}}\neq 0$, as they assume $\beta=0$ otherwise. Accordingly, when $\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}} = 0$ we collapse the Oracle CI to the singleton $\{0\}$, which covers the true parameter $\beta$ only if $\beta = 0$.\footnote{Results are qualitatively similar without this restriction.} The Uniform CI is constructed the same way for all values of $\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}}$. As before, all results reported in this subsection are based on $10{,}000$ Monte Carlo replications.

Figure (ref) reports coverage probabilities for those choices of tuning parameters $\lambda_T$ from the previous subsection that lead to consistent tuning, i.e., $\lambda_T\in\{T^{1/4},T^{1/2},T\}$. Since the oracle property requires $T^{-1}\lambda_T+\lambda_T^{-1}\to 0$ (see, e.g., Corollary (ref)), it does not hold for $\lambda_T=T$. Consequently, the Oracle CI is asymptotically invalid in this case, whereas the Uniform CI remains asymptotically valid. The figure shows that the coverage probabilities of all confidence intervals are close to one when $\beta=0$. However, for small deviations of $\beta$ away from zero, the coverage probability of the Oracle CI drops sharply, often to values below $0.5$, while the coverage probability of the Uniform CI remains relatively stable. As $\beta$ moves further away from zero, the coverage probability of the Oracle CI eventually recovers. Nevertheless, it performs poorly in the region of primary interest, namely when $\beta$ is small but non-zero. More generally, we find that the higher the rate at which $\lambda_T$ diverges, or the larger the sample size $T$, the more stable the coverage probability of the Uniform CI across values of $\beta$. When $\lambda_T=T$, the asymptotically invalid Oracle CI can exhibit coverage probabilities close to zero even in large samples, whereas the Uniform CI remains asymptotically valid and performs well even in small samples. Finally, the coverage probabilities of the asymptotic Uniform CI are generally much lower than those of the Uniform CI. This reflects the point already emphasized in Section (ref) that the confidence sets must be constructed using the realized regressors $x_t$, rather than quantities based on (limiting) distributional properties only.

We next repeat the analysis after scaling $\lambda_T$ by a factor of four in each case, leaving its rate of divergence unchanged. Figure (ref) presents the results. Scaling $\lambda_T$ in this way further improves the coverage probability of the Uniform CI across all sample sizes and divergence rates, while the coverage probability of the Oracle CI deteriorates further. This outcome reflects two opposing effects of increasing $\lambda_T$. First, more estimates $\ensuremath{\hat\beta_{\scriptstyle\text{\rm\tiny AL}}}$ are shrunk to zero, which worsens the performance of the Oracle CI. Second, the width of the Uniform CI increases, which improves its coverage.

Finally, we relax the assumption of i.i.d.\ standard normal regression errors and regressor innovations and examine the performance of the confidence intervals under error serial correlation and regressor endogeneity. Specifically, we generate $u_t=\rho_1 u_{t-1}+e_t+\rho_2\nu_t$ and $v_t=\nu_t+0.5\nu_{t-1}$, where $[e_t,\nu_t]'\sim\mathcal{N}(0,(4/9)I_2)$ i.i.d.\ across $t$.\footnote{Rescaling the variance ensures that $\Omega_{vv}=1$ and thus simplifies the comparison with the previous results.} The parameters $\rho_1$ and $\rho_2$ govern the degree of error serial correlation and regressor endogeneity, respectively. Figure (ref) reports results for $\rho_1=\rho_2=0.6$ and the scaled tuning parameters.\footnote{When simulating the quantiles of $\mathcal{Z}^c$ for the oracle intervals, we use the true long-run covariance parameters to capture the dependence structure in the data. In applications, these quantities are unknown and must be estimated. In the presence of local-to-unity regressors this approach becomes infeasible as $\mathcal{Z}^c$ also depends on the local-to-unity parameters which are not consistently estimable. In contrast, the uniform confidence intervals are unaffected by these issues.} The figure shows that the previous findings remain intact in the presence of serial correlation and regressor endogeneity.

To complete the analysis, Tables (ref) and (ref) in Appendix (ref) report the lengths of the confidence intervals underlying Figures (ref)--(ref). As expected, the Uniform CI is typically longer than the Oracle CI, but the difference is usually moderate and diminishes further as the sample size increases. In some instances, however, the Uniform CI can be substantially longer than the Oracle CI, but this occurs either for unfavorable realizations of $x_t$ or in settings where the Oracle CI is asymptotically invalid and exhibits coverage probabilities close to zero. We therefore conclude that the uniform confidence region proposed in Section (ref) constitutes a useful tool for quantifying uncertainty around adaptive LASSO estimates under consistent tuning in empirical applications.

figure[figure omitted — 4,947 chars of source]
figure[figure omitted — 4,790 chars of source]
figure[figure omitted — 4,784 chars of source]
figure[figure omitted — 4,800 chars of source]
figure[figure omitted — 3,070 chars of source]
figure[figure omitted — 3,133 chars of source]
figure[figure omitted — 3,283 chars of source]

Empirical Illustration

We apply the adaptive LASSO estimator within a predictive regression framework to forecast the U.S. monthly unemployment rate (UNRATE). As potential predictors, we include the variables considered by BuckmannJoseph23: the 3-month Treasury bill (TB3MS), real personal income (RPI), industrial production (INDPRO), consumption (DPCERA3M086SBEA), the S&P 500 price index (S&P 500), business loans (\texttt{BUSLOANS}), the consumer price index (\texttt{CPIAUCSL}), the oil price (\texttt{OILPRICEx}), and the M2 money stock (\texttt{M2SL}). In addition, we include the four variables most frequently selected by the standardized LASSO in MeiShi24 when predicting the U.S. unemployment rate one month ahead using a 20-year rolling window: initial jobless claims (\texttt{CLAIMSx}), the number of unemployed less than 5 weeks (\texttt{UEMPLT5}), the number of unemployed 5 to 14 weeks (\texttt{UEMP5TO14}), and the number of unemployed 15 weeks and over (\texttt{UEMP15OV}). All series are obtained from the FRED-MD macroeconomic database McCrackenNg16, with their respective FRED-MD codes indicated in parentheses. Throughout, we use the raw data without applying any transformations.

MeiShi24 demonstrate the usefulness of LASSO-type methods for predicting the U.S. unemployment rate using the full set of FRED-MD variables, considering multiple horizons (1, 2, and 3 months) and rolling windows of 10, 20, and 30 years, and benchmarking their results against a random walk with drift and an autoregressive model. In contrast, our focus is not on relative predictive performance, but on quantifying uncertainty around adaptive LASSO estimates using the uniformly valid confidence intervals proposed in Section (ref). Although we compare the magnitudes of adaptive LASSO coefficients to their OLS counterparts, oracle-based confidence intervals for OLS are infeasible because the limiting distribution of the OLS estimator is distorted by nuisance parameters arising from endogeneity, serial correlation, and local-to-unity predictors.

We report results for one-month-ahead out-of-sample forecasts based on a 20-year rolling window using data from January 1959 to December 2024. The sample from January 1959 to December 1979 is used for initial estimation, while forecasts are evaluated over the period January 1980 to December 2024.

Within each rolling window, the penalization parameter $\lambda_T$ is selected via time-series cross-validation following HyndmanAthanasopoulos18. For each candidate value of $\lambda_T$, the adaptive LASSO is calculated using the first 60% of observations in the window and then used to generate a one-month-ahead forecast. The estimation sample is then expanded recursively by one observation at a time until the end of the window, producing a sequence of forecasts. The value of $\lambda_T$ is chosen to minimize the resulting root mean squared forecasting error across those forecasts. The candidate set for $\lambda_T$ consists of a grid ranging from zero to the smallest value that shrinks all coefficients to zero when estimated on the full window.

The left panel of Figure (ref) in Appendix (ref) shows the adaptive LASSO and OLS forecasts alongside the observed unemployment rate, while the right panel reports the corresponding forecast errors. Over the full evaluation period, OLS attains a root mean squared forecasting error of $0.82$, whereas the adaptive LASSO reduces this by $11\%$ to $0.73$. This improvement largely reflects the adaptive LASSO’s ability to accommodate structural changes during and in the aftermath of the COVID-19 pandemic. The spike in forecast errors both for OLS and adaptive LASSO in the beginning of COVID-19 is driven by the abrupt surge in CLAIMSx. While the adaptive LASSO adjusts rapidly in subsequent months, the OLS estimator fails to do so, as illustrated in Figure (ref).

In line with the motivation of this paper, we find that the coefficients for the four labor-market variables (CLAIMSx, UEMPLT5, UEMP5TO14, UEMP15OV) are frequently estimated to be small but non-zero, whereas the remaining coefficients are often shrunk to zero by the adaptive LASSO. Figure (ref) reports the rolling-window adaptive LASSO estimates for the coefficients corresponding to the labor-market variables together with their uniformly valid confidence intervals, benchmarked against the corresponding OLS estimates. The resulting confidence intervals appear plausible, often widening during and in the aftermath of crisis episodes. This behavior can be partly attributed to increases in the penalization parameter $\lambda_T$ (see Figure (ref) in Appendix (ref)) and partly to changes in the underlying variables. Importantly, a larger $\lambda_T$ does not necessarily imply wider confidence intervals. For example, during and in the aftermath of the COVID-19 crisis, $\lambda_T$ is elevated, yet the confidence intervals for the coefficients on CLAIMSx and UEMP5TO14 become noticeably narrower.

Figure (ref) in Appendix (ref) shows the results for the remaining variables. Although the adaptive LASSO often sets their coefficients to zero, the associated uncertainty can still be relatively large, particularly during crises.

Overall, the application highlights the usefulness of the adaptive LASSO for estimating relationships among economic variables in the presence of structural changes or shocks. The confidence intervals proposed in this paper are plausible and allow to quantify uncertainty around adaptive LASSO estimates. Although the confidence intervals can occasionally be wide, they are robust to endogeneity, serial correlation, and local-to-unity parameters, making them a valuable tool for empirical applications.

figure[figure omitted — 890 chars of source]

Summary and Conclusions

This paper analyzes the asymptotic behavior of the adaptive LASSO estimator in cointegrating regressions with local-to-unity regressors under moving-parameter asymptotics. We establish model selection probabilities, estimation consistency, limiting distributions, and uniform convergence rates, as well as the fastest local-to-zero rates that remain detectable by the estimator. As these rates depend critically on the tuning regime, the results characterize the smallest signal-to-noise ratios that can reliably be detected under both conservative and consistent tuning. In addition, under consistent tuning, we construct uniformly valid confidence regions for the regression coefficients that are straightforward to compute and do not require any knowledge or estimation of nuisance parameters associated with long-run covariance matrices or local-to-unity parameters.

Our simulation study demonstrates that the finite-sample distribution of the adaptive LASSO estimator often differs substantially from what is implied by the oracle property. In contrast, the limiting distributions derived under moving-parameter asymptotics provide accurate approximations and also successfully capture the empirical frequency with which coefficients are set to zero. As a result, the proposed uniform confidence regions exhibit stable coverage probabilities across the parameter space, whereas confidence regions based on the oracle property perform poorly when the true coefficients are close to zero. The empirical application complements these findings by illustrating the usefulness of the proposed confidence regions for quantifying uncertainty around adaptive LASSO estimates in practice.

Several promising directions for future research emerge. First, an extension of the analysis to the twin adaptive LASSO proposed in LeeEtAl22 for settings that allow also for stationary regressors and cointegration among local-to-unity regressors would extend the framework to a broader class of econometric models. Second, further work should focus on the properties of the proposed confidence regions, including their potential usefulness for hypothesis testing. Third, theoretical guidance on choosing the penalization parameter that balances the performance of the adaptive LASSO estimator and the size of the confidence regions would further enhance the empirical applicability of the proposed methods. Finally, extending the results to high-dimensional regressions represents a natural next step toward data-rich applications.

Acknowledgements

We thank the participants of a research seminar at the Vienna University of Economics and Business in 2025 and of the 2025 Econometrics Workshop at TU Dortmund University for their valuable comments and suggestions.

\paragraph*{Declaration of Interest} The authors have no conflicts of interest to declare.

thebibliography{38} \expandafter\ifx\csname natexlab\endcsname\relax\def\natexlab#1{#1}\fi \bibitem[{Adamek et al.(2023)Adamek, Smeeks and Wilms}]{AdamekEtAl23} Adamek, R., Smeeks, S. and Wilms, I. (2023). \newblock Lasso inference for high-dimensional time series. \newblock Journal of Econometrics 235, 1114--1143. \bibitem[{Amann and Schneider(2023)}]{AmannSchneider23} Amann, N. and \textsc{Schneider, U.} (2023). \newblock Uniform asymptotics and confidence regions based on the adaptive lasso with partially consistent tuning. \newblock \textit{Econometric Theory} \textbf{39}, 1097--1122. \bibitem[{Buckmann and Joseph(2023)}]{BuckmannJoseph23} \textsc{Buckmann, M.} and \textsc{Joseph, A.} (2023). \newblock An interpretable machine learning workflow with an application to economic forecasting. \newblock \textit{International Journal of Central Banking} \textbf{19}, 449--522. \bibitem[{Campbell(2008)}]{Campbell08} \textsc{Campbell, J. Y.} (2008). \newblock Viewpoint: Estimating the equity premium. \newblock \textit{Canadian Journal of Economics} \textbf{41}, 1--21. \bibitem[{Campbell and Yogo(2006)}]{CampbellYogo06} \textsc{Campbell, J. Y.} and \textsc{Yogo, M.} (2006). \newblock Efficient tests of stock return predictability. \newblock \textit{Journal of Financial Economics} \textbf{81}, 27--60. \bibitem[{Chen et al.(2025)Chen, Li, Li and Linton}]{ChenEtAl25} \textsc{Chen, J.}, \textsc{Li, D.}, \textsc{Li, Y.-N.} and \textsc{Linton, O.} (2025). \newblock Estimating time-varying networks for high-dimensional time series. \newblock \textit{Journal of Econometrics} \textbf{249}, 105941. \bibitem[{de Jong(2003)}]{deJong03UP} \textsc{de Jong, R.} (2003). \newblock Nonlinear estimators with integrated regressors but without exogeneity. \newblock Mimeo. \bibitem[{Demetrescu et al.(2022)Demetrescu, Georgiev, Rodrigues and Taylor}]{DemetrescuEtAl22} \textsc{Demetrescu, M.}, \textsc{Georgiev, I.}, \textsc{Rodrigues, P. M. M.} and \textsc{Taylor, A. M. R.} (2022). \newblock Testing for episodic predictability in stock returns. \newblock \textit{Journal of Econometrics} \textbf{227}, 85--113. \bibitem[{Fan and Li(2001)}]{FanLi01} \textsc{Fan, J.} and \textsc{Li, R.} (2001). \newblock Variable selection via nonconcave penalized likelihood and its oracle properties. \newblock \textit{Journal of the American Statistical Association} \textbf{96}, 1348--1360. \bibitem[{Geyer(1996)}]{Geyer96TR} \textsc{Geyer, C.} (1996). \newblock On the asymptotics of convex stochastic optimization. \newblock Unpublished manuscript. \bibitem[{Gonzalo and Pitarakis(2025)}]{GonzaloPitarakis25} \textsc{Gonzalo, J.} and \textsc{Pitarakis, J.-Y.} (2025). \newblock Detecting sparse cointegration. \newblock Preprint 2501.13839, arxiv. \bibitem[{Hwang and Vald\'{e}s(2024)}]{HwangValdes24} \textsc{Hwang, J.} and \textsc{Vald\'{e}s, G.} (2024). \newblock Low frequency cointegrating regression with local to unity regressors and unknown form of serial dependence. \newblock \textit{Journal of Business and Economic Statistics} \textbf{42}, 160--173. \bibitem[{Hyndman and Athanasopoulos(2018)}]{HyndmanAthanasopoulos18} \textsc{Hyndman, R. J.} and \textsc{Athanasopoulos, G.} (2018). \newblock \textit{Forecasting: Principles and Practice}. \newblock OTexts. \bibitem[{Ibragimov and Phillips(2008)}]{IbragimovPhillips08} \textsc{Ibragimov, R.} and \textsc{Phillips, P. C. B.} (2008). \newblock Regression asymptotics using martingale convergence methods. \newblock \textit{Econometric Theory} \textbf{24}, 888--947. \bibitem[{Jensen(2009)}]{Jensen09} \textsc{Jensen, M. J.} (2009). \newblock The long-run fisher effect: Can it be tested? \newblock \textit{Journal of Money, Credit and Banking} \textbf{41}, 221--231. \bibitem[{Kock(2016)}]{Kock16} \textsc{Kock, A. B.} (2016). \newblock Consistent and conservative model selection with the adaptive {LASSO} in stationary and nonstationary autoregressions. \newblock \textit{Econometric Theory} \textbf{32}, 243--259. \bibitem[{Koo et al.(2020)Koo, Anderson, Seo and Yao}]{KooEtAl20} \textsc{Koo, B.}, \textsc{Anderson, H. M.}, \textsc{Seo, M. H.} and \textsc{Yao, W.} (2020). \newblock High-dimensional predictive regression in the presence of cointegration. \newblock \textit{Journal of Econometrics} \textbf{219}, 456--477. \bibitem[{Lee et al.(2022)Lee, Shi and Gao}]{LeeEtAl22} \textsc{Lee, J. H.}, \textsc{Shi, Z.} and \textsc{Gao, Z.} (2022). \newblock On {LASSO} for predictive regression. \newblock \textit{Journal of Econometrics} \textbf{229}, 322--349. \bibitem[{Liao and Phillips(2015)}]{LiaoPhillips15} \textsc{Liao, Z.} and \textsc{Phillips, P. C. B.} (2015). \newblock Automated estimation of vector error correction models. \newblock \textit{Econometric Theory} \textbf{31}, 581--646. \bibitem[{Magdalinos and Phillips(2009)}]{MagdalinosPhillips09} \textsc{Magdalinos, T.} and \textsc{Phillips, P. C. B.} (2009). \newblock Econometric inference in the vicinity of unity. \newblock Preprint, CoFie Working Paper 7, Singapore Management University. \bibitem[{McCracken and Ng(2016)}]{McCrackenNg16} \textsc{McCracken, M. W.} and \textsc{Ng, S.} (2016). \newblock {FRED-MD}: A monthly data-base for macroeconomic research. \newblock \textit{Journal of Business and Economic Statistics} \textbf{34}, 574--589. \bibitem[{Medeiros and Mendes(2017)}]{MedeirosMendes16} \textsc{Medeiros, M. C.} and \textsc{Mendes, E. F.} (2017). \newblock $\ell_1$-regularization of high-dimensional time-series models with non-gaussian and heteroskedastic errors. \newblock \textit{Journal of Econometrics} \textbf{191}, 255--271. \bibitem[{Mei and Shi(2024)}]{MeiShi24} \textsc{Mei, Z.} and \textsc{Shi, Z.} (2024). \newblock On lasso for high dimensional predictive regression. \newblock \textit{Journal of Econometrics} \textbf{242}, 105809. \bibitem[{Park and Phillips(1988)}]{ParkPhillips88} \textsc{Park, J. Y.} and \textsc{Phillips, P. C. B.} (1988). \newblock Statistical inference in regressions with integrated processes: Part 1. \newblock \textit{Econometric Theory} \textbf{4}, 468--497. \bibitem[{Phillips(1988)}]{Phillips88} \textsc{Phillips, P. C. B.} (1988). \newblock Regression theory for near-integrated time series. \newblock \textit{Econometrica} \textbf{56}, 1021--1043. \bibitem[{Phillips(2015)}]{Phillips15} \textsc{Phillips, P. C. B.} (2015). \newblock Pitfalls and possibilities in predictive regression. \newblock \textit{Journal of Financial Econometrics} \textbf{13}, 521--555. \bibitem[{Phillips(2023)}]{Phillips23} \textsc{Phillips, P. C. B.} (2023). \newblock Estimation and inference with near unit roots. \newblock \textit{Econometric Theory} \textbf{39}, 221--263. \bibitem[{Phillips and Hansen(1990)}]{PhillipsHansen90} \textsc{Phillips, P. C. B.} and \textsc{Hansen, B. E.} (1990). \newblock Statistical inference in instrumental variables regression with i(1) processes. \newblock \textit{Review of Economic Studies} \textbf{57}, 99--125. \bibitem[{P\"otscher and Schneider(2009)}]{PoetscherSchneider09} \textsc{P\"otscher, B. M.} and \textsc{Schneider, U.} (2009). \newblock On the distribution of the adaptive {LASSO} estimator. \newblock \textit{Journal of Statistical Planning and Inference} \textbf{139}, 2775--2790. \bibitem[{Ren and Zhang(2010)}]{RenZhang10} \textsc{Ren, Y.} and \textsc{Zhang, X.} (2010). \newblock Subset selection for vector autoregressive processes via adaptive {L}asso. \newblock \textit{Statistics and Probability Letters} \textbf{80}, 1705--1712. \bibitem[{Saikkonen and Choi(2004)}]{SaikkonenChoi04} \textsc{Saikkonen, P.} and \textsc{Choi, I.} (2004). \newblock Cointegrating smooth transition regressions. \newblock \textit{Econometric Theory} \textbf{20}, 301--340. \bibitem[{Schweikert(2022)}]{Schweikert22} \textsc{Schweikert, K.} (2022). \newblock Oracle efficient estimation of structural breaks in cointegrating regressions. \newblock \textit{Journal of Time Series Analysis} \textbf{43}, 83--104. \bibitem[{Smeekes and Wijler(2021)}]{SmeekesWijler21} \textsc{Smeekes, S.} and \textsc{Wijler, E.} (2021). \newblock An automated approach towards sparse single-equation cointegration modelling. \newblock \textit{Journal of Econometrics} \textbf{221}, 247--276. \bibitem[{Tibshirani(1996)}]{Tibshirani96} \textsc{Tibshirani, R.} (1996). \newblock Regression shrinkage and selection via the {L}asso. \newblock \textit{Journal of the Royal Statistical Society Series B} \textbf{58}, 267--288. \bibitem[{Tu and Xie(2023)}]{TuXie23} \textsc{Tu, Y.} and \textsc{Xie, X.} (2023). \newblock Penetrating sporadic return predictability. \newblock \textit{Journal of Econometrics} \textbf{237}, 105509. \bibitem[{Wagner and Hong(2016)}]{WagnerHong16} \textsc{Wagner, M.} and \textsc{Hong, S. H.} (2016). \newblock Cointegrating polynomial regressions: Fully modified {OLS} estimation and inference. \newblock \textit{Econometric Theory} \textbf{32}, 1289--1315. \bibitem[{Wang et al.(2007)Wang, Li and Tsai}]{WangEtAl07b} \textsc{Wang, H.}, \textsc{Li, G.} and \textsc{Tsai, L.} (2007). \newblock Regression coefficient and autoregressive order shrinkage and selection via the {L}asso. \newblock \textit{Journal of the Royal Statistical Society Series B} \textbf{69}, 63--78. \bibitem[{Zou(2006)}]{Zou06} \textsc{Zou, H.} (2006). \newblock The adaptive {L}asso and its oracle properties. \newblock \textit{Journal of the American Statistical Association} \textbf{101}, 1418--1429.