EconBase
← Back to paper

Variable Selection in High Dimensional Linear Regressions with Parameter Instability

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

254,299 characters · 23 sections · 69 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Variable Selection in High Dimensional Linear Regressions with Parameter Instability

{\setstretch{1.1}

abstractThis paper considers the problem of variable selection allowing for parameter instability. It distinguishes between signal and pseudo-signal variables that are correlated with the target variable, and noise variables that are not, and investigate the asymptotic properties of the One Covariate at a Time Multiple Testing (OCMT) method proposed by chudik2018one under parameter insatiability. It is established that OCMT continues to asymptotically select an approximating model that includes all the signals and none of the noise variables. Properties of post selection regressions are also investigated, and in-sample fit of the selected regression is shown to have the oracle property. The theoretical results support the use of unweighted observations at the selection stage of OCMT, whilst applying down-weighting of observations only at the forecasting stage. Monte Carlo and empirical applications show that OCMT without down-weighting at the selection stage yields smaller mean squared forecast errors compared to Lasso, Adaptive Lasso, and boosting.

{ Keywords: Lasso, One Covariate at a time Multiple Testing, OCMT, Parameter instability, Variable selection, Forecasting}

{ JEL Classifications: C22, C52, C53, C55} } \thispagestyle{empty} \pagenumbering{arabic}

Introduction

Models fitted to statistical relationships could be subject to parameter instabilities. In an extensive early study, StockWatson1996 find that a large number of time series regressions in economics are subject to breaks. ClementsHendry1998 consider parameter instability to be one of the main sources of forecast failure. This problem has been addressed at the estimation/forecasting stage for a given set of selected regressors. Typical solutions are either to use rolling windows or exponential down-weighting. For instance, pesaran2007selection, pesaran2011forecast and inoue2017rolling consider the choice of an observation window, and hyndman2008forecasting and pesaran2013optimal, respectively consider exponential and non-exponential down-weighting of the observations. There are also Bayesian approaches to prediction that allow for the possibility of breaks over the forecast horizon, such as chib1998estimation, koop2004forecasting, and pesaran2006forecasting. rossi2013advances provides a review of the literature on forecasting under instability. There are also related time varying parameter and regime switching models that are used for forecasting. See, for example, hamilton1988rational and dangl2012predictive . This literature does not address the problem of variable selection and takes the model specification as given.

The theory of variable selection in the presence of parameter instability is still largely underdeveloped. The application of penalized regression methods to variable selection is often theoretically justified under two key parameter stability assumptions: the stability of the coefficients in the data generating process and the stability of the correlation matrix of the covariates in the active set. Under these assumptions, the penalized regression methods can proceed using the full sample without down-weighting or separating the variable selection from the estimation stage. However, in the presence of parameter instability penalized regression methods must be adapted to simultaneously deal with selection and parameter change. There are a number of recent studies that use machine learning techniques to allow for parameter instability, in particular penalized regression, especially the Least Absolute Shrinkage and Selection Operator (Lasso) initially proposed by tibshirani1996regression. For example, qian2016shrinkage consider a linear regression model with a finite number of covariates but allow for an unknown number of breaks and use group fused Lasso by 10.1007/978-3-642-40728-4_9 to consistently estimate the number of breaks and their locations. lee2016lasso have proposed a Lasso procedure that allows for threshold effects. kapetanios2018time have proposed a time-varying Lasso procedure, where all the parameters of the model vary locally. fan2014nonparametric suggest an extension of the screening procedure initially proposed by fan2008sure to the case where the regression coefficients vary smoothly with an observable exposure variable. Also recently, yousuf2019boosting propose an interesting boosting procedure for the estimation of high-dimensional models with locally time varying parameters. These studies focus on specific forms of discrete or continuous time varying parameter models, and often carry out variable selection and estimation simultaneously using the penalized regression or boosting procedures.

This paper proposes the use of One Covariate at a Time Multiple Testing (OCMT) procedure proposed by chudik2018one which is readily adapted to the task of variable selection under parameter instability. The key insight comes from the fact that coefficients of the noise variables that do not enter the data generating process are zero at all times. Consequently, using unweighted observations at the variable selection stage will be most effective in removing noise variables, while using weighted observations at the estimation stage can provide gains in terms of mean squared forecast errors. In this study, we allow the marginal effects of signals on the target variable, as well as the correlation of the covariates under consideration, to vary over time, assuming time variations in the marginal effects are not correlated with the signals. We establish the conditions required for OCMT with unweighted observations to select a model that contains all the signal variables and none of the noise variables with probability approaching one as the sample size, $T$, and the number of covariates under consideration, $N$, tend to infinity.

Clearly, it is also possible to use penalized regression methods with unweighted observations for the purpose of variable selection, and then estimate the selected model by the least squares method using weighted observations. However, as far as we know, there are no studies that consider the choice of the penalty term to achieve variable selection consistency under parameter instability. It is hoped that the present paper provides an impetus for further theoretical analysis of penalized regression techniques under parameter instability. Although at this stage a comparison of the assumptions required for variable selection consistency of OCMT and Lasso under parameter instability is not possible, in Section (ref) we provide a discussion of the assumptions required for the variable selection consistency of Lasso under parameter stability that are comparable with the those required for the validity of the OCMT procedure.

The OCMT procedure selects variables based on the statistical significance of the net effect of the covariates in the active set on the target variable, one-at-a-time subject to the multiple testing nature of the inferential problem involved. The idea of using one-at-a-time regressions is not unique to OCMT and has been used in boosting as well as in screening approaches. See, for example, buhlmann2006boosting and fan2018sure as prominent examples of these approaches. What is unique about the OCMT procedure is its inferentially motivated stopping rule without resorting to the use of information criteria, or penalized regression after the initial stage.\ In the case of models with stable parameters, chudik2018one establish that OCMT asymptotically selects an approximating model that includes all the signals and none of the noise variables. This model can contain covariates that do not enter the data generating process for the target variable but exhibit non-zero correlation with at least one signal, known as pseudo-signals.

Lasso and OCMT exploit different aspects of the low-dimensional structure assumed for the underlying data generating process. Lasso restricts the magnitude of the correlations within signals as well as the correlations between signals and the remaining covariates in the active set. OCMT limits the rate at which the number of pseudo-signals, $k_{T}^{\ast }$, rises with the sample size, $T$. Under parameter stability, the variable selection consistency of Lasso has been investigated by zhao2006model, meinshausen2006high and more recently by lahiri2021necessary. These conditions, and how they compare with the conditions that underlie OCMT, are discussed in Section (ref) of the paper. Although Lasso does not directly impose any restrictions on $k_{T}^{\ast }$, its Irrepresentable Condition (IRC), by restricting the magnitude of correlations within and between the signals and pseudo-signals, does have implications for the number of pseudo-signals that Lasso selects. OCMT requires $k_{T}^{\ast }$\ not to rise faster than $\sqrt{T}$. When this condition is violated, then the true signals must end up as common factors for the pseudo-signals, and what matters is the number of residuals (from the regressions of pseudo-signals on the common factors) that are correlated with the residuals of the true signals from the same set of common factors. sharifvaghefi2023variable shows that such common factors can be estimated from the principal components of the covariates in the active set and the OCMT\ condition on the number pseudo-signals, now defined in terms of the correlation of the residuals, is no longer restrictive.\footnote{ Another extension of OCMT\ is provided by su2023one who allow for unknown potentially non-linear relationship between the signals and the target variable.} Once the model is selected, Theorem (ref) establishes how the convergence rate of estimated coefficients of the selected variables depends on $k_{T}^{\ast }$. The regular convergence rate of $\sqrt{T}$ is achieved only if $k_{T}^{\ast }$ is fixed in $T$. A similar issue also arises for Lasso, as shown by lahiri2021necessary who establishes that the Lasso procedure cannot achieve both variable selection consistency and $\sqrt{T}$-consistency in coefficient estimation. As noted above, the focus of the present paper is on the application of OCMT to variable selection in the presence of parameter instability, broadly defined. To the best of our knowledge, there are no studies that investigate the variable selection properties of Lasso under parameter instability.

To take account of the time variations in the coefficients of the signals, we consider their time averages and distinguish between strong signals whose average marginal effects go to a non-zero value, semi-strong signals whose average marginal effects converge to zero, but sufficiently slow, and weak signals whose average marginal effects approach to zero quite fast. In this way we allow for variety of time variations that could arise in practice. Strong signals tend to have non-zero effects at all times, semi-strong signals could have zero effects during some periods, with weak signals enter the model relatively rarely. Weak signals are often indistinguishable from noise variables. In our theoretical analysis we will focus on selection of strong and semi-strong signals.

We provide three main theorems in support of our proposed variable selection method. Under certain fairly general regularity conditions we show that the probability of OCMT selecting the approximating model that contains all the signals (strong and semi-strong) and none of the noise variables approaches to one as $T $ goes to infinity. Our results apply both when $N$ is fixed as well as when $N$ goes to infinity jointly with $T$, covering the case where $ N \gg T$. We also establish conditions under which (a) least squares estimates of the coefficients of selected covariates converge to zero unless they are signals, and (b) the average squared residuals of the selected model achieves the oracle rate for regression models with time-varying coefficients. These theoretical findings provide a formal justification for application of statistical techniques from the time-varying parameters literature to the post OCMT selected model. Our Monte Carlo experiments show that the OCMT procedure with weighted observations only at the estimation stage outperforms, in terms of mean squared forecast errors, Lasso and Adaptive Lasso (A-Lasso by zou2006adaptive), as well as boosting by buhlmann2006boosting, under many different settings.

Finally, we provide three empirical applications, forecasting monthly rates of price changes of 28 stocks in Dow Jones using large number of financial, economic and technical indicators, forecasting output growths across 33 countries using a large number of macroeconomic indicators, and forecasting euro area output growth using ECB surveys of 25 professional forecasters. To save space the third application is included in the online supplement. We generate a large number of forecasts using OCMT with and without down-weighting of the observations at the selection stage and compare the results with the forecasts obtained using Lasso, A-Lasso and boosting. The empirical results are in line with our theoretical and MC findings and suggest that using down-weighted observations at the selection stage of the OCMT procedure worsens forecast performance in terms of mean squared forecast errors and mean directional forecast accuracy. The empirical results also provide that OCMT with no down-weighting at the selection stage outperforms, in terms of mean squared forecast errors, boosting, Lasso and A-Lasso.

The rest of the paper is organized as follows: Section (ref) sets out the model specification. Section (ref) explains the basic idea behind the OCMT procedure for variable selection without down-weighting in the presence of parameter instability. {Section (ref) provides a discussion of key assumptions of Lasso and OCMT under parameter stability.} Section (ref) discusses the technical assumptions and the asymptotic properties of the OCMT procedure under parameter instability. Section (ref) provides the details of the Monte Carlo experiments and a summary of the main results. Section (ref) presents the empirical applications, and Section (ref) concludes. The paper is also accompanied with three online supplements. A theory supplement contains the mathematical proofs of the theorems and related lemmas. A Monte Carlo supplement provides additional summary tables, the full set of Monte Carlo results, as well as the description of the algorithms used for Lasso, A-Lasso and boosting. Further details of the empirical applications are given in an empirical supplement.

Notations: Generic finite positive constants are denoted by $C_{i}$ for $i=1,2,\cdots $. $\lVert \mathbf{A}\rVert _{2}$ and $\lVert \mathbf{A} \rVert _{F}$ denote the spectral and Frobenius norms of matrix $\mathbf{A}$, respectively. $\text{tr}(\mathbf{A})$ and $\lambda _{i}(\mathbf{A})$ denote the trace and the $i^{th}$ eigenvalue of a square matrix $\mathbf{A}$, respectively. $\left\Vert \mathbf{x}\right\Vert $ denotes the $\ell _{2}$ norm of vector $\mathbf{x}$. {\ If $\{f_{n}\}_{n=1}^{\infty }$ and $ \{g_{n}\}_{n=1}^{\infty }$ are both positive sequences of real numbers, then we say $f_{n}=\ominus (g_{n})$ if there exist $n_{0}\geq 1$ and positive constants $C_{0}$ and $C_{1}$, such that $\inf_{n\geq n_{0}}\left( f_{n}/g_{n}\right) \geq C_{0}$ and $\sup_{n\geq n_{0}}\left( f_{n}/g_{n}\right) \leq C_{1}$. }Similarly, if $f_{iT}$\ and $g_{iT}$\ are positive double sequences of real numbers for $i=1,2,3,\cdots $; and $ T=1,2,3,\cdots $, then $f_{iT}=\ominus (g_{iT})$\ if there exist $T_{0}\geq 1 $\ and positive constants $C_{0}$\ and $C_{1}$, such that $\inf_{T\geq T_{0}}\left( f_{iT}/g_{iT}\right) \geq C_{0}$\ and $\sup_{T\geq T_{0}}\left( f_{iT}/g_{iT}\right) \leq C_{1}$.

Model specification under parameter instability

We consider the following data generating process (DGP) for the target variable, $y_{t}$, in terms of the signal variables $(x_{it}$, for $ i=1,2,...,k$)

equation[equation omitted — 116 chars of source]

with time-varying parameters, $\left\{ \beta _{it}\text{, } i=1,2,...,k\right\} $, and an error term, $u_{t}$. Intercepts and other pre-selected variables can also be included.\footnote{ See the working paper version of the paper available at \url{https://doi.org/10.24149/gwp394r2}.} Since the parameters are time-varying we refer to the covariate $i$ as \textquotedblleft signal" if its average marginal effect, $\bar{\beta}_{i,T}=T^{-1} \sum_{t=1}^{T}\mathbb{E}(\beta _{it})$, is not equal to zero. The strength of the signal can be captured by the exponent coefficient $\vartheta _{i}$ in $\bar{\beta}_{i,T}=\ominus (T^{-\vartheta _{i}})$. For $\vartheta _{i}=0$ , the signal is strong and the average marginal effect, $\bar{\beta}_{i,T}$, does not converge to zero. For $0<\vartheta _{i}<1/2$, the signal is semi-strong and the average marginal effect converges to zero, but not too fast. For $\vartheta _{i}\geq 1/2$, the average marginal effect tends to zero very fast, making it infeasible for the OCMT procedure to distinguish such weak signals from noise, unless weak signals are sufficiently correlated with at least one strong or semi-strong signal. In this paper, we do not impose any restrictions on the correlations among signals, and we focus only on the covariates with strong and semi-strong signals, where $ 0\leq \vartheta _{i}<1/2$. For simplicity of exposition, unless specified otherwise, we will refer to both strong and semi-strong signals simply as signals.

The identity of the $k$ signals are unknown, and the task facing the investigator is to select the signals from a set of covariates under consideration, $\mathcal{S}_{Nt}=\{x_{1t},x_{2t},\cdots ,x_{Nt}\}$, known as the active set, with $N$, the number of covariates in the active set, possibly much larger than $T$, the number of data points available for estimation prior to forecasting. The time variations in $\beta _{it}$, for $ i=1,2,...,k$, are assumed to be exogenous, in the sense that $\beta _{it}$ are distributed independently of the covariates in the active set $\mathcal{S }_{Nt}$. This assumption rules out correlated time variations that can arise in non-linear regressions where $y_{t}$ is a non-linear function of the signals. One important example is given by the bilinear model

equation*[equation* omitted — 82 chars of source]

where it is assumed that $\beta _{it}$ systematically varies with $x_{it}$. Nevertheless, in the context of linear regressions, our assumptions about parameter instability includes many models of parameter instability studied in the literature. Specifically, our analysis accommodates cases where the coefficients vary continuously following a stochastic process as in the standard random coefficient model,

equation*[equation* omitted — 75 chars of source]

or could change at discrete time intervals, as

equation*[equation* omitted — 139 chars of source]

where $T_{0}=1$ and $T_{S}=T$.

In this paper we follow chudik2018one and consider the application of the OCMT procedure for variable selection even when the parameters are time-varying, and provide theoretical arguments in favour of using the full sample of data available without down-weighting. We first recall that OCMT's variable selection is based on the net effect of $x_{it}$ on $y_{t}$. However, when the regression coefficients and/or the correlations across the covariates in the active set are time-varying, the net effects will also be time-varying and we need to base our selection on average net effects. The average net effect of the covariate $x_{it}$ on $y_{t}$ can be defined as

equation*[equation* omitted — 90 chars of source]

By substituting $y_{t}$ from ((ref)) we can further write $\bar{ \theta}_{i,T}$ as (noting that $\beta _{jt}$ and $x_{it}$ are assumed to be independently distributed)

equation*[equation* omitted — 143 chars of source]

where $\sigma _{ij,t}=\mathbb{E}(x_{it}x_{jt})$, and $\bar{\sigma} _{iu,T}=T^{-1}\sum_{t=1}^{T}\mathbb{E}(x_{it}u_{t})$. In what follows we allow for a mild degree of correlation between $x_{it},$ and $u_{t},$ by assuming that $\bar{\sigma}_{iu,T}=O(T^{-\epsilon _{i}}),$ for some $ \epsilon _{i} \geq 1/2$. In this case the average net effect of the $i^{th}$ covariate simplifies to

equation*[equation* omitted — 145 chars of source]

In line with our assumption about the average marginal effects, namely that $ \bar{\beta}_{i,T}=\ominus (T^{-\vartheta _{i}}),$ for some $0\leq \vartheta _{i}<1/2$, we distinguish between covariates with strong and semi-strong net effects, and the noise variables whose net effects, averaged over time, tend to zero sufficiently fast. Specifically, for covariates with strong or semi-strong net effects we set $\bar{\theta}_{i,T}=\ominus (T^{-\vartheta _{i}}),$ for some $0\leq \vartheta _{i}<1/2$, and for the noise variables we shall assume that $\bar{\theta}_{i,T}=\ominus (T^{-\epsilon _{i}}),$ for some $\epsilon _{i}\geq 1/2$.

In what follows, we first describe the OCMT procedure and then discuss the conditions under which the approximating model (that includes all the signals and none of the noise variables) is selected with probability approaching one by OCMT.

Parameter instability and OCMT

The OCMT procedure begins with $N$ separate regressions, for each of the $N$ covariates in the active set $\mathcal{S}_{Nt}$. Specifically, the focus is on the statistical significance of $\phi _{i,T}$ in the following simple regressions:

equation[equation omitted — 121 chars of source]

where

equation[equation omitted — 239 chars of source]

with $\bar{\sigma}_{ii,T}=$ $T^{-1}\sum_{t=1}^{T}\sigma _{ii,t}$. Due to non-zero cross-covariate correlations, knowing whether $\phi _{i,T}$ $\ ($or equivalently $\bar{\theta}_{i,T}$) is zero does not necessarily allow us to establish whether $\bar{\beta}_{i,T}$ is sufficiently close to zero or not. There are four possibilities:

{

tabular[tabular omitted — 550 chars of source]

}

{ {$\dagger $ The signals are assumed to be (semi) strong such that $0\leq \vartheta _{i}<1/2$.}}

for some $0\leq \vartheta _{i}<1/2$, and $\epsilon _{i}\geq 1/2$. { To simplify the exposition, we consider the covariates $x_{it}$, for $ i=1,2,\cdots ,k$, as signals, and for $i=k+1,k+2,\cdots ,k+k_{T}^{\ast }$, as pseudo-signals. The remaining covariates in the active set, }$\left\{ { x_{it},}\text{ for }{i=k+k_{T}^{\ast }+1,k+k_{T}^{\ast }+2,\cdots ,N} \right\} ${, are classified as (pure) noise variables. We assume that the number of signals, $k$, is a finite fixed integer but we allow the number of pseudo-signals, denoted by $k_{T}^{\ast }$, to grow with $N$ and $T$.} Notice, if the covariate $x_{it}$ is a noise variable, then $\bar{\theta} _{i,T}$ converges to zero very fast. Therefore, down-weighting of observations at the variable selection stage is likely to be inefficient for eliminating the noise variables. Moreover, for a signal to remain hidden, we need the terms of higher order, $\ominus (T^{-\vartheta _{j}})$ with $0\leq \vartheta _{i}<1/2$, to exactly cancel out such that $\theta _{i,T}$ becomes a lower order, i.e. $\ominus (T^{-\varepsilon _{i}})$, that tends to zero at a sufficiently fast rate (with $\epsilon _{i}\geq 1/2)$. This combination of events seem quite unlikely, and to simplify the theoretical derivations in what follows we abstract from such a possibility and assume that there are no hidden signals and consider a single stage version of the OCMT procedure for variable selection. To allow for hidden signals, chudik2018one extend the OCMT method to have multiple stages.

flushleftThe OCMT procedure
enumerate• For $i=1,2,\cdots ,N$, regress $y_{t}$ on $x_{it}$; $y_{t}=\phi _{i,T}x_{it}+\eta _{it}$; and compute the $t$-ratio of $\phi _{i,T}$, given by \begin{equation} t_{i,T}=\frac{\hat{\phi}_{i,T}}{s.e.\left( \hat{\phi}_{i,T}\right) }=\frac{ \sum_{t=1}^{T}x_{it}y_{t}}{\hat{\sigma}_{i}\sqrt{\sum_{t=1}^{T}x_{it}^{2}}}, \end{equation} where $\hat{\phi}_{i,T}=\left( \sum_{t=1}^{T}x_{it}^{2}\right) ^{-1}\left( \sum_{t=1}^{T}x_{it}y_{t}\right) $ is the least squares estimator of $\phi _{i,T}$, $\hat{\sigma}_{i}^{2}=T^{-1}\sum_{t=1}^{T}\hat{\eta}_{it}^{2}$, and $\hat{\eta}_{it}=y_{t}-\hat{\phi}_{i,T}x_{it}$, is the regression residual. • Consider the critical value function, $c_{p}(N,\delta )$, defined by \begin{equation} c_{p}(N,\delta )=\Phi ^{-1}\left( 1-\frac{p}{2N^{\delta }}\right) , \end{equation} where $\Phi ^{-1}(.)$ is the inverse of a standard normal distribution function, $\delta $ is a finite positive constant, and $p$ is the nominal size of the tests to be set by the investigator. • Given $c_{p}(N,\delta )$, the selection indicator is given by \begin{equation} \hat{\mathcal{J}}_{i}=\mathbf{\mathit{I}}\left[ \lvert t_{i,T}\rvert >c_{p}(N,\delta )\right] ,\ for\ i=1,2,\cdots ,N. \end{equation} The covariate $x_{it}$ is selected if $\hat{\mathcal{J}}_{i}=1$.

{OCMT uses the t-ratio of $\phi _{i,T}$, defined by ((ref)), to select the signals (strong as well as semi-strong), $\{x_{it}:i=1,2,\cdots ,k\}$, and none of the noise variables, $\{x_{it}:k+k_{T}^{\ast }+1,k+k_{T}^{\ast }+2,\cdots ,N\}$. The selected model is referred to as an approximating model since it can include pseudo-signals, $ \{x_{it}:k+1,k+2,\cdots ,k+k_{T}^{\ast }\}$, that proxy for the true signals. To deal with the multiple testing nature of the problem, the critical value $c_{p}(N,\delta )$ used for the separate-induced tests is chosen to be an appropriately increasing function of $N$, by setting $\delta >0$. The choice of $\delta $ is guided by our theoretical derivations, to be discussed below in Section (ref). }

Before presenting our technical assumptions and theoretical results under parameter instability, it is instructive to discuss and compare the key conditions under which Lasso and OCMT lead to consistent model selection under parameter stability.

Lasso and OCMT under parameter stability

{As formally established by zhao2006model and meinshausen2006high, three main conditions are required for the Lasso variable selection to be consistent. Here we follow lahiri2021necessary who also considers the convergence of Lasso estimated coefficients to their true values. The key condition is the \textquotedblleft Irrepresentable Condition" (IRC) that places restrictions on the magnitudes of the sample correlations across the signals, }${\mathbf{x }_{1t}=(x_{1t},x_{2t},\cdots ,x_{kt})^{\prime }}$, and the rest of the { covariates in the active set, namely $\mathbf{x}_{2t}=(x_{k+1,t},x_{k+2,t}, \cdots ,x_{Nt})^{\prime }$. Let

equation*[equation* omitted — 129 chars of source]

be the $N\times N$ matrix of sample correlations of the covariates in the active set, partitioned conformably to $\mathbf{x}_{t}=(\mathbf{x} _{1t}^{\prime },\mathbf{x}_{2t}^{\prime })^{\prime }$. The IRC can be written as

equation[equation omitted — 160 chars of source]

where $\left\Vert .\right\Vert _{\infty }$ is the $\ell _{\infty }$ norm of a vector, $\text{sign}(.)$ is the sign function, and $\boldsymbol{\beta } _{0}=\left( \beta _{01},\beta _{02},\cdots ,\beta _{0k}\right) ^{\prime }$ is the $k\times 1$ vector of the coefficients of the signals. The following example provides more intuition on how IRC imposes restrictions on the magnitudes of the sample correlations between the covariates in the active set. }

exampleSuppose the DGP for $y_{t}$ contains only two signals, $x_{1t}$ and $x_{2t}$ . Denote the sample correlation coefficient between $x_{1t}$ and $x_{2t}$ by $\hat{\rho}$, and the sample correlation coefficients of $x_{1t}$ and $ x_{2t} $ with the rest of the covariates in the active set, $ x_{3,t},x_{4,t},\cdots ,x_{Nt}$, by $\hat{\rho}_{i1}$ and $\hat{\rho}_{i2}$, for $i=3,4,\cdots ,N$, respectively. Then, after some algebra, the IRC given by ((ref)) simplifies to \begin{equation*} \max_{i\in \{3,4,\cdots ,N\}}\left\vert (\hat{\rho}_{i1}-\hat{\rho}\hat{\rho} _{i2})sign(\beta _{01})+(\hat{\rho}_{i2}-\hat{\rho}\hat{\rho}_{i1}) sign(\beta _{02})\right\vert \leq 1-\hat{\rho}^{2}. \end{equation*} There are two cases: (A) $\text{sign}\left( \beta _{01}\right) =\text{sign} \left( \beta _{02}\right) $ and (B) $\text{sign}\left( \beta _{01}\right) \neq \text{sign}\left( \beta _{02}\right) $. Under case (A) it follows that the IRC condition is met if \begin{equation*} \max_{i\in \{3,4,\cdots ,N\}}\left\vert \hat{\rho}_{i1}+\hat{\rho} _{i2}\right\vert \leq 1+\hat{\rho}. \end{equation*} Similarly under case (B) it is required that \begin{equation*} \max_{i\in \{3,4,\cdots ,N\}}\left\vert \hat{\rho}_{i1}-\hat{\rho} _{i2}\right\vert \leq 1-\hat{\rho}. \end{equation*}

From the above example, it is clear that IRC places restrictions on the magnitude of sample correlation among signals ($\hat{\rho}$ in the above example), as well as the magnitude of sample correlation between signals and pseudo-signals ($\hat{\rho}_{i1}$ and $\hat{\rho}_{i2}$). Notably, the IRC is met for noise variables but need not hold for pseudo-signals. OCMT also has no difficulty in dealing with noise variables, and is very effective at eliminating them. However, for consistent estimation of the approximate model, post OCMT selection, it is necessary to restrict the number of selected covariates relative to the sample size, $T$. To this end, OCMT assumes that the number of pseudo-signals, $k_{T}^{\ast }$, could grow at an order less than the square root of the number of observations, namely

equation*[equation* omitted — 83 chars of source]

It is important to note that OCMT does not place any restrictions on the magnitude of correlations of signals and pseudo-signals. Instead, it limits the number of covariates that are correlated with the signals ($k_{T}^{\ast } $). Clearly, the IRC could be violated even when the number of pseudo-signals grows at an order less than $\sqrt{T}$. Hence the OCMT's requirement on the number of pseudo-signals allows for cases where the IRC does not hold, and vice versa.

The condition on the number of pseudo-signals ($k_{T}^{\ast }$) in the OCMT framework has been recently relaxed by sharifvaghefi2023variable. To illustrate how this is done, suppose there are no noise variables and hence the signals, $\mathbf{x}_{1t}=\left( x_{1t},x_{2t},\cdots ,x_{kt}\right) ^{\prime }$, are correlated with all the remaining covariates in the active set. In this case if $N \gg \sqrt{T} $, a straightforward application of OCMT will not be valid. But, we can model the correlation between the signals, $\mathbf{x}_{1t}$, and the remaining covariates, {$\mathbf{x}_{2t}$, } as

equation*[equation* omitted — 156 chars of source]

The signals thus act as strong factors for the pseudo-signals. Given that the identity of signals and pseudo-signals are unknown and the number of pseudo-signals is large, it is reasonable to propose the existence of latent factors, $\mathbf{f}_{t}$, that are common across the covariates in the active set. This idea can be formally expressed as:

equation*[equation* omitted — 121 chars of source]

where $\boldsymbol{\psi }_{i}$ is vector of factor loadings, and $ \varepsilon _{it}$ refers to the idiosyncratic components that are weakly cross-correlated such that

equation[equation omitted — 125 chars of source]

Substituting $x_{it}$ into the DGP for $y_{t}$, given by ((ref)), we obtain:

equation*[equation* omitted — 121 chars of source]

with $\boldsymbol{\delta }_{0}=\sum_{i=1}^{k}\beta _{i0}\boldsymbol{\psi } _{i}$. When the common factors, $\mathbf{f}_{t}$, and idiosyncratic components, $\varepsilon _{it}$ , are known, this model would correspond to that presented in working paper version of our work, where common factors $ \mathbf{f}_{t}$ can be used as preselected variables. Since $\mathbf{f}_{t}$ and $\varepsilon _{it}$ are not known, sharifvaghefi2023variable shows that when both $N$ and $T$ are large the OCMT selection can be carried out using the principal component estimators of $\mathbf{f}_{t}$ and $ \varepsilon _{it}$, denoted by $\hat{\mathbf{f}}_{t}$ and $\hat{\varepsilon} _{it}$, using all the covariates in the active set. The large $N$ is required for consistent estimation of the common factors. As a result, the OCMT\ condition on the number of pseudo-signals now relates to the number of $\varepsilon _{it}$ for $i=k+1,k+2,...,N$ that are correlated with $ \varepsilon _{it}$ for $i=1,2,...,k$, which is bounded under condition ((ref)).

For variable selection consistency of Lasso under parameter stability, the literature further requires the penalty term, $\lambda _{T}$, to grow at an order greater than $\sqrt{T}$ such that:

equation*[equation* omitted — 192 chars of source]

where $\mathbf{x}_{2t}^{\perp} $ is the part of variation in $\mathbf{x} _{2t} $ that is orthogonal to $\mathbf{x}_{1t} $ and $u_{t}$ is the error term in the data generating process. The exact choice of $\lambda _{T}$ in practice is often unclear, with practitioners typically relying on cross-validation methods.

A third condition required by Lasso for variable selection consistency is the beta-min condition:

equation*[equation* omitted — 192 chars of source]

where $\left\vert .\right\vert _{j}$ denotes the absolute value of the $ j^{th}$ element of a vector. Given that $\lambda _{T}$ must grow at an order greater than $\sqrt{T}$, we can conclude from the beta-min condition that $ \beta _{i0}\gg \frac{1}{\sqrt{T}}$ for $i=1,2,\cdots ,k$. For example, lahiri2021necessary assumes that $\beta _{i0}\gg \sqrt{\frac{k\log (T)}{T}} $. The OCMT's requirement on the strength of signals (under parameter stability) is given by $\beta _{i0}=\ominus (T^{-\vartheta _{i}})$, for some $0\leq \vartheta _{i}<1/2$. This condition is essentially very similar to the Lasso's beta-min condition.

Asymptotic properties of OCMT under parameter instability

We establish the asymptotic properties of the OCMT procedure for variable selection assuming the time variations in $\beta _{it}$ for $i=1,2,...,k$ are distributed independently of the regressors in the active set. We also make additional assumptions that bound the degree of time variations in $ \beta _{it}$ and $x_{it}$, in addition to assuming the exponentially decaying tail probabilities for $\beta _{it}$ and $x_{it}$. Our assumptions on $x_{it}$, $i=1,2,...,k$ and their correlations with the other variables in the active set are in line with those assumed in the literature. A formal statement of these assumptions are set out in Section (ref). Theorem (ref) establishes that OCMT continues to asymptotically select an approximating model that includes all the signals and none of the noise variables. Additional assumptions are required for investigating the asymptotic properties of the least squares estimates of the post OCMT selected model. These assumptions and the related theorems are provided in Section (ref). Theorem (ref) establishes the rate at which the least squares estimates of the coefficients of the selected model converge to their true time averages. It is shown that the regular convergence rate of $\sqrt{T}$ is achieved only if $k_{T}^{\ast }$ (the number of selected covariates) is fixed in $T$. Irregular convergence rates result when $k_{T}^{\ast }$ rises in $T$. Theorem (ref) shows that the sum of squared residuals of the estimated model converges in probability to its limiting value at the oracle rate of $\sqrt{T}$. The limiting value consists of two components: the first is the unavoidable uncertainty due to the unobserved error term, $u_{t}$, and the second is the cost (in terms of fit) of ignoring the time variations in the coefficients of the signals.

Suppose the target variable, $y_{t}$, is generated by ((ref)) in terms of $x_{it}$ for $i=1,2,...,k$, and $\mathbf{x}_{t}=(x_{1t},x_{2t}, \cdots ,x_{kt},x_{k+1,t},....,x_{Nt})^{\prime }$ is the $N\times 1$ vector of covariates in the active set ($N \gg k$). Let $\bar{\beta}_{i,T}\equiv T^{-1}\sum_{t=1}^{T}\mathbb{E}(\beta _{it})$, for $i=1,2,...,k$, and $\bar{ \theta}_{i,T}=\sum_{j=1}^{k}\left( T^{-1}\sum_{t=1}^{T}\mathbb{E}(\beta _{jt})\sigma _{ij,t}\right) +\bar{\sigma}_{iu,T},$ for $i=1,2,...,N$, where $ \sigma _{ij,t}=\mathbb{E}(x_{it}x_{jt})$, and $\bar{\sigma} _{iu,T}=T^{-1}\sum_{t=1}^{T}\mathbb{E}(x_{it}u_{t})$. Define the filtrations $\mathcal{F}_{t}^{u}=\sigma (u_{t},u_{t-1},\cdots )$, $\mathcal{F} _{t}^{x}=\sigma (\mathbf{x}_{t},\mathbf{x}_{t-1},\cdots )$, and $\mathcal{F} _{jt}^{\beta }=\sigma (\beta _{jt},\beta _{j,t-1},\cdots )$, for $ j=1,2,\cdots ,k$. Set $\mathcal{F}_{t}^{\beta }=\cup _{j=1}^{k}\mathcal{F} _{jt}^{\beta }$ and $\mathcal{F}_{t}=\mathcal{F}_{t}^{q}\cup \mathcal{F} _{t}^{\mathrm{a}}\cup \mathcal{F}_{t}^{\beta }\cup \mathcal{F}_{t}^{u}$, and consider the following assumptions:

Assumptions

assumption[Coefficients of signals] \textcolor{white}{enter} \newline (a) The number of signals, $k$, is a finite fixed integer. (b) $\beta _{jt}$ , $j=1,2,\cdots ,k$, are distributed independently of $x_{it^{\prime }}$, $ i=1,2,\cdots ,N$, and $u_{t^{\prime }}$ for all $t$ and $t^{\prime }$. (c) The signals are (semi) strong in the sense that $\bar{\beta}_{j,T}=\ominus (T^{-\vartheta _{j}})$ for $0\leq \vartheta _{j}<1/2$, $j=1,2,...,k.$ (d) There are no hidden signals in the sense that $\bar{\theta}_{j,T}=\ominus (T^{-\vartheta _{j}})$, for $0\leq \vartheta _{j}<1/2$, $j=1,2,...,k$.
assumption[Martingale difference processes] \textcolor{white}{enter} \newline For $i, i^{\prime} =1,2,\cdots ,N$, $j=1,2,\cdots ,k$, and $t=1,2,\cdots ,T$ , (a) $\mathbb{E}\left[ x_{it}x_{ i^{\prime} t}-\mathbb{E} (x_{it}x_{i^{\prime} t})|\mathcal{F}_{t-1}\right] =0$, (b) $\mathbb{E}\left[ u_{t}^{2}-\mathbb{E}\left( u_{t}^{2}\right) |\mathcal{F}_{t-1}\right] =0$, (c) $\mathbb{E}\left[ x_{it}u_{t}-\mathbb{E}(x_{it}u_{t})|\mathcal{F}_{t-1} \right] =0$, where $T^{-1}\sum_{t=1}^{T}\mathbb{E}(x_{it}u_{t})=O(T^{- \epsilon _{i}}),$ with $\epsilon _{i} \geq 1/2$, and (d) $\mathbb{E}\left[ \beta _{j t}-\mathbb{E}(\beta _{j t})|\mathcal{F}_{t-1}\right] =0$.
assumption[Exponential decaying probability tails] \textcolor{white}{enter} \newline There exist sufficiently large positive constants $C_{0}$ and $C_{1}$, and $ s>0$ such that for all $\alpha >0$, (a) $\sup_{i,t}\Pr (|x_{it}|>\alpha )\leq C_{0}\exp (-C_{1}\alpha ^{s})$, (b) $\sup_{i,t}\Pr (|\beta _{it}|>\alpha )\leq C_{0}\exp (-C_{1}\alpha ^{s})$, and (c) $\sup_{t}\Pr (|u_{t}|>\alpha )\leq C_{0}\exp (-C_{1}\alpha ^{s})$.

{Before presenting the theoretical results, we briefly discuss the rationale behind our assumptions and compare them with the assumptions typically made in the high-dimensional linear regressions and the parameter instability literature. }

Assumption (ref)(a) posits that the number of signals is a fixed integer. This is crucial to ensure that the random variable $y_{t}$ has a distribution with an exponentially decaying probability tail. Under the premise that the covariates $x_{it}$ for all $i$ and $t$ are non-random and fixed, which is a common assumption in the penalized regression setting, it becomes permissible for the number of signals to grow with the sample size at an order slower that $\sqrt{T}$. Assumption (ref)(b) is common in the literature under parameter instability and restrict the distribution of time-varying parameters to be independent of the covariates. Assumption (ref)(c) is an identification assumption needed to distinguish signals from noise variables and is similar to the beta-min condition already discussed in Section (ref). Finally, Assumption (ref)(d) ensures that there are no hidden signals. As discussed in Section (ref), we make this assumption to simplify the theoretical derivations, and one can use the multi-stage OCMT procedure suggested by chudik2018one to allow for hidden signals.

To establish that the OCMT procedure with the critical value function $ c_{p}(N,\delta )=\Phi ^{-1}\left( 1-\frac{p}{2N^{\delta }}\right) $ does not select any of the noise variables with a probability approaching one as $N$ and $T$ go to infinity, we need to show that the t-statistic given by ((ref)) follows a distribution with exponentially decaying tails. We utilize the concentration inequality of an exponential decaying rate to accomplish this goal. Assumptions (ref) and (ref) place constraints on the sequence of random variables, $x_{it}$ for $i=1,2,\ldots ,N$, $\beta _{jt}$ for $j=1,2,\ldots ,k$, and $u_{t}$ such that they adhere to a martingale difference process and exhibit exponential decaying probability tails. These assumptions are sufficient to establish the exponential decaying concentration inequality, as provided in Lemma (ref) in the online theory supplement. Notably, these assumptions could be relaxed provided that the exponential decaying concentration inequality holds. For example, Theorem 1 of merlevede2011bernstein and Lemma D1 of the online theory supplement for chudik2018one establishes that this inequality can be achieved while allowing for weak time-series dependence. In penalized regression literature, a commonly held assumption is that the covariates are non-random and fixed. Moreover, error terms $\{u_{t}\}_{t=1}^{T}$ are typically assumed to be serially independent. See, for example, see zhao2006model, javanmard2013model, lee2015model, belloni2014inference, javanmard2020flexible, and lahiri2021necessary. Additionally, in the Lasso literature it is often assumed that $u_{t}$ possesses an exponentially decaying probability tail. See, for example, javanmard2018debiasing, hansen2019factor, fan2020factor, and javanmard2020flexible.

Variable selection consistency

As mentioned in Section (ref), the purpose of this paper is to provide the theoretical argument for applying the OCMT procedure with no down-weighting at the variable selection stage in linear high-dimensional settings subject to parameter instability. We now show that under the assumptions set out in Section (ref), the OCMT procedure selects the approximating model that contains all the signals; $ \{x_{it}:i=1,2,\cdots ,k\}$; and none of the noise variables; $ \{x_{it}:k+k_{T}^{\ast }+1,k+k_{T}^{\ast }+2,\cdots ,N\}$. The event of choosing the approximating model is defined by

equation[equation omitted — 204 chars of source]

Note that the approximating model can contain pseudo-signals. In what follows, we show that $\Pr (\mathcal{A}_{0})\rightarrow 1$, as $ N,T\rightarrow \infty $.

theoremConsider the DGP for $y_{t}$, $t=1,2,\cdots ,T$, given by ((ref)), and the set $\mathcal{S}_{Nt}= \{x_{1t},x_{2t},\cdots ,x_{Nt}\}$ that contains $k$ signals, $k_{T}^{\ast }$ pseudo-signals, and $N-k-k_{T}^{\ast }$ noise variables. Suppose that Assumptions (ref)-(ref) hold and $N=\ominus (T^{\kappa })$ with $ \kappa >0$. Then, there exist finite positive constants $C_{0}$ and $C_{1}$ such that, for any $0<\pi <1$ and any null sequence $d_{T}>0$, the probability of selecting the approximating model $\mathcal{A}_{0}$, as defined by ((ref)), by the OCMT procedure with the critical value function $c_{p}(N,\delta )$ given by ((ref)), for some $\delta >0$, is \begin{equation} \Pr (\mathcal{A}_{0})=1-O\left[ T^{\kappa \left( 1-\mathcal{X}_{NT}\left( \frac{1-\pi }{1+d_{T}}\right) ^{2}\delta \right) }\right] -O\left[ T^{\kappa }\exp \left( -C_{0}T^{C_{1}}\right) \right] , \end{equation} where, \begin{equation*} \textstyle\mathcal{X}_{NT}=\inf_{i\in \{k+k^{\ast }+1,\cdots ,N\}}\frac{\bar{ \sigma}_{\eta _{i},T}^{2}\bar{\sigma}_{x_{i},T}^{2}}{\bar{\omega}_{iy,T}^{2}} , \end{equation*} $\bar{\sigma}_{x_{i},T}^{2}=T^{-1}\sum_{t=1}^{T}\mathbb{E}(x_{it}^{2})$, $ \bar{\omega}_{iy,T}^{2}=T^{-1}\sum_{t=1}^{T}\mathbb{E}(x_{it}^{2}y_{t}^{2}| \mathcal{F}_{t-1})$, $\bar{\sigma}_{\eta _{i},T}^{2}=T^{-1}\sum_{t=1}^{T} \mathbb{E}(\eta _{it}^{2})$, $\eta _{it}=y_{t}-\phi _{i,T}x_{it}$, and $\phi _{i,T}$ is defined by ({(ref)}).

This theorem shows that the probability of selecting the approximating model is unaffected by parameter instability, so long as the average net effects of the signals are non-zero or converge to zero sufficiently slowly in $T$, as defined formally by Assumption (ref). The theorem also highlights the importance of an appropriate choice of $\delta $ for model selection consistency. Corollary (ref) in the online theory supplement shows that if the covariates in the active set are generated by a stationary process and the noise variables are independent of $y_{t}$ then $ \mathcal{X}_{NT}=1$. As a result, for any $\delta >1$, OCMT consistently selects the approximating model, $\mathcal{A}_{0}$. Notably, $c_{p}(N,\delta )$ is reasonably stable with respect to small increases in $\delta $ in the neighborhood of $\delta =1$ and the extensive Monte Carlo studies in chudik2018one also suggest that setting $\delta =1$ performs well in practice.\footnote{ One could also use the heteroscedasticity and/or autocorrelation robust standard errors in computation of t-statistics given by ((ref)) to ensure the consistent selection of the approximating model for any $\delta >1 $ in a more general setup.}

Properties of the post OCMT selected model

To investigate the asymptotic properties of the least squares estimates of the selected model (post OCMT) we require the following additional assumption:

assumption[Eigenvalues] The eigenvalue condition \begin{equation*} \lambda _{\min }\left[ T^{-1}\sum\limits_{t=1}^{T}\mathbb{E}(\mathbf{x}_{ \tilde{k}_{T},t}\mathbf{x}_{\tilde{k}_{T},t}^{\prime })\right] >c>0, \end{equation*} holds, where $\mathbf{x}_{\tilde{k}_{T},t}$, for $t=1,2,...,T$ are the $ \tilde{k}_{T}\times 1$ vector of observations on signals ($k$) and pseudo-signals ($k_{T}^{\ast }$) with $\tilde{k}_{T}=k+k_{T}^{\ast }$.

This assumption ensures that the post OCMT selected model can be consistently estimated subject to certain regularity conditions to be discussed below. The post OCMT selected model can be written as

equation*[equation* omitted — 90 chars of source]

where $\hat{\mathcal{J}}_{i}=\mathbf{\mathit{I}}\left[ \lvert t_{i,T}\rvert >c_{p}(N,\delta )\right] ,\ $defined by ((ref)). Also $ \sum_{i=1}^{N}\hat{\mathcal{J}}_{i}=\hat{k}_{T}$, where $\hat{k}_{T}$ is the number of covariates selected by OCMT. By Theorem (ref) the probability that the selected model contains the signals tends to unity as $T\rightarrow \infty $. We can further write

equation[equation omitted — 171 chars of source]

where $\mathbf{w}_{t}=\left( w_{1t},w_{2t},\cdots ,w_{\hat{k}_{T}t}\right) ^{\prime }$. The least squares (LS) estimator of selected coefficients, $ \boldsymbol{\gamma }_{T}=\left( \gamma _{1},\gamma _{2},\cdots ,\gamma _{ \hat{k}_{T}}\right) ^{\prime }$, is given by

equation[equation omitted — 208 chars of source]

In establishing the rate of convergence of $\hat{\mathbf{\gamma }}_{T}$ we distinguish between two cases: when the vector of signals, $\mathbf{x} _{k,t}=\left( x_{1t},x_{2t},\cdots ,x_{kt}\right) ^{\prime }$ is included in $\mathbf{w}_{t}$ as a subset, and when this is not the case. But we know from Theorem (ref) that the probability of the latter tends to zero at a sufficiently fast rate. The following theorem provides the conditions under which the estimates of the coefficients of the selected signals and pseudo-signals of the approximating model tend to their true mean values, defined formally below.

theoremLet the DGP for $y_{t}$, $t=1,2,\cdots ,T$, be given by ((ref)) and write down the regression model selected by the OCMT procedure as ((ref)). Suppose that Assumptions (ref) -(ref) hold and the number of pseudo-signals, $k_{T}^{\ast }$, grow with $T$ such that $ k_{T}^{\ast}=\ominus (T^{d})$ with $0\leq d<\frac{1}{2}$. Consider the least squares (LS) estimator of $\boldsymbol{\gamma }_{T}=\left( \gamma _{1},\gamma_{2},\cdots ,\gamma _{\hat{k}_{T}}\right) ^{\prime }$, given by ( (ref)). \begin{enumerate} • If $\mathbb{E}(\beta _{it})=\beta _{i}$ for all $t$, then, \begin{equation*} \left\| \hat{\boldsymbol{\gamma}}_{T} - \boldsymbol{\gamma}_{T}^{*} \right\| = O_p \left(T^{\frac{d-1}{2}} \right), \end{equation*} where $\boldsymbol{\gamma}_{T}^{*} = (\gamma _{1}^{*}, \gamma _{2}^{*}, \cdots, \gamma_{\hat{k}_{T}}^{*})^{\prime} $, and \begin{equation*} \left\{ \begin{matrix} \gamma _{\ell}^{*} \in \boldsymbol{\beta }=(\beta _{1},\beta _{2},\cdots,\beta _{k})^{\prime }, & if w_{\ell t}\in \mathbf{x}_{kt} \\ \gamma _{\ell}^{*} = 0, & otherwise. \end{matrix} \right. \end{equation*} • If $\mathbb{E}\left( \mathbf{x}_{\tilde{k}_{T},t} \mathbf{x}_{ \tilde{k}_{T},t}^{\prime}\right) $ is a fixed time-invariant matrix, where $ \tilde{k}_{T} = k + k^{\ast}_{T}$, then, \begin{equation*} \left\Vert \hat{\boldsymbol{\gamma }}_{T}-\boldsymbol{\gamma }_{T}^{\diamond }\right\Vert =O_{p}\left( T^{\frac{d-1}{2}}\right) , \end{equation*} where $\boldsymbol{\gamma }_{T}^{\diamond }=(\gamma _{1T}^{\diamond },\gamma _{2T}^{\diamond },\cdots ,\gamma _{\hat{k} _{T},T}^{\diamond })^{\prime }$ , and \begin{equation*} \left\{ \begin{matrix} \gamma _{\ell ,T}^{\diamond }\in \boldsymbol{\bar{\beta}}_{T}=(\bar{\beta} _{1T},\bar{\beta}_{2T},\cdots ,\bar{\beta}_{kT})^{\prime }, & if w_{\ell t}\in \mathbf{x}_{kt} \\ \gamma _{\ell ,T}^{\diamond }=0, & otherwise, \end{matrix} \right. \end{equation*} and $\bar{\beta}_{iT}=T^{-1}\sum_{t=1}^{T}\mathbb{E}(\beta _{it})$, $ i=1,2,\cdots ,k$. \end{enumerate}
remark{\ The above theorem builds on Theorem (ref) and establishes that in the post OCMT selected model estimated by LS only signals will end up having non-zero limiting values, as $N$ and $ T\rightarrow \infty $. This theorem also shows that the convergence rate of the LS estimators depends on $d$, defined by $k_{T}^{\ast }=\ominus (T^{d})$ , and the regular $\sqrt{T}$ rate of convergence is achieved only if $d=0$. Similarly, lahiri2021necessary establishes that the Lasso procedure cannot achieve both variable selection consistency and $\sqrt{T}$ -consistency in coefficient estimation. }
remarkThe conditions of Theorem (ref) are met in the case of random coefficient models where $\beta _{it}=\beta _{i}+\sigma _{it}\xi _{it}$, and $\xi _{it}$ are distributed independently of the signals, and the LS estimator of $\boldsymbol{\gamma }_{T}^{\ast }$ is consistent, so long as $0\leq d<1/2$. Interestingly, if signal and pseudo-signal variables are generated by a stationary process, and hence they satisfy condition (ii) of Theorem (ref), then we can extend the random coefficient model to have time-varying means, and still estimate $ \boldsymbol{\gamma }_{T}^{\ast }$ consistently by LS.

Lastly, we consider the fit of the post OCMT selected regression in terms of its residuals given by

equation[equation omitted — 155 chars of source]

It is worth noting that even when all the signal variables are correctly selected, the forecasts based on the selected model will be biased due to parameter instability. The implications of parameter instability for the in-sample fit of the selected regression is derived in Proposition (ref) of the online theory supplement, abstracting from variable selection uncertainty. In what follows we derive the asymptotic properties of the sum of squared residuals (SSR) of the selected model, namely $\sum_{t=1}^{T}\hat{\eta}_{t}^{2}$, taking account of the costs associated with variable selection uncertainty and parameter instability. To this end we need the following assumption on the cross correlation of parameter heterogeneity.

assumption[Weak time dependence] $h_{ij,t}=x_{it}x_{jt}(\beta _{it}-\bar{\beta} _{iT})(\beta _{jt}-\bar{\beta}_{jT})$ is weakly correlated over time such that \begin{equation*} \sum_{t=1}^{T}\sum_{t^{\prime }=1}^{T}cov(h_{ij,t},h_{ij,t^{\prime }})=O(T), for i,j=1,2,..,k, \end{equation*} where $cov(.,.)$ is the covariance operator.
remarkAssumption (ref) is a high-level assumption. Here is an example of conditions under which this assumption holds. Suppose, Assumptions (ref) and (ref) hold, and the cross products of coefficients of the signals follow martingale difference processes such that \begin{equation*} \mathbb{E}\left[ \beta _{it}\beta _{jt}-\mathbb{E}(\beta _{it}\beta _{jt})| \mathcal{F}_{t-1}\right] =0, for i=1,2,\cdots ,k,\ j=1,2,\cdots ,k, and t=1,2,\cdots ,T. \end{equation*} Then, $\sum_{t=1}^{T}\sum_{t^{\prime }=1}^{T}\text{cov}(h_{ij,t},h_{ij,t^{ \prime }})=O(T)$. See Lemma (ref) in the online theory supplement for a proof. \color{black}

The following theorem establishes the limiting property of SSR of the post OCMT selected model.

theoremLet the DGP for $y_{t}$, $t=1,2,\cdots ,T$ be given by ((ref)) and write down the regression model selected by the OCMT procedure as ((ref)). Suppose that Assumptions (ref)- (ref) hold and the number of pseudo-signals, $ k_{T}^{\ast }$, grow with $T$ such that $k_{T}^{\ast }=\ominus (T^{d})$ with $0 \leq d < \frac{1}{2} $. Consider the residuals of the selected model, estimated by LS and given by ((ref)). \begin{enumerate} • If $\mathbb{E}(\beta _{it})=\beta _{i}$ for all $t$, then \begin{equation} \textstyle T^{-1}SSR=\bar{\sigma}_{u,T}^{2}+\bar{\Delta}_{\beta ,T}+O_{p}\left( T^{-\frac{1}{2}}\right) +O_{p}\left( T^{d-1}\right) , \end{equation} where $\bar{\sigma}_{u,T}^{2}=T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( u_{t}^{2}\right) $, and $\bar{\Delta}_{\beta ,T}=T^{-1}\sum_{t=1}^{T}\text{tr }\left( \boldsymbol{\Sigma }_{\mathbf{x}_{k},t}\boldsymbol{\Omega }_{\beta ,t}\right) $ are non-negative, with $\boldsymbol{\Sigma }_{\mathbf{x} _{k},t}\equiv \left( \sigma _{ijt,x}\right) $, $\boldsymbol{\Omega }_{\beta ,t}\equiv \left( \sigma _{ijt,\beta }\right) $ for $i,j=1,2,\cdots ,k$, and $ \sigma _{ijt,x}=\mathbb{E}\left( x_{it}x_{jt}\right) $, $\sigma _{ijt,\beta }=\mathbb{E}\left[ (\beta _{it}-\beta _{i})(\beta _{jt}-\beta _{j})\right] $. • $\ $Let $\tilde{k}_{T}=k+k_{T}^{\ast }$ and suppose that $\mathbb{E}\left( \mathbf{x}_{\tilde{k}_{T},t}\mathbf{x}_{ \tilde{k}_{T},t}^{\prime }\right) $ is time-invariant (fixed). Then, \begin{equation} \textstyle T^{-1}SSR=\bar{\sigma}_{u,T}^{2}+\bar{\Delta}_{\beta ,T}^{\ast }+O_{p}\left( T^{-\frac{1}{2}}\right) +O_{p}\left( T^{d-1}\right) , \end{equation} where $\bar{\Delta}_{\beta ,T}^{\ast }=T^{-1}\sum_{t=1}^{T}\text{tr}\left( \boldsymbol{\Sigma }_{\mathbf{x}_{k},t}\boldsymbol{\Omega }_{\beta ,t}^{\ast }\right) $ is non-negative, with $\boldsymbol{\Omega }_{\beta ,t}^{\ast }\equiv \left( \sigma _{ijt,\beta }^{\ast }\right) $ for $i,j=1,2,\cdots ,k$ , and $\sigma _{ijt,\beta }^{\ast }=\mathbb{E}\left[ (\beta _{it}-\bar{\beta} _{i,T})(\beta _{jt}-\bar{\beta}_{j,T})\right] $. \end{enumerate}
remarkThe condition $d<\frac{1}{2}$ in Theorem (ref) ensures that the number of pseudo-signals grows sufficiently slowly in $T$, which in turn ensures that $T^{1-d}<T^{-\frac{1}{ 2}}$ and hence from equations ((ref)) and ((ref)), we can conclude that the average of squared residuals ($T^{-1}$SSR) of the Post OCMT selected model convergences at the same rate of $T^{-\frac{1}{2}}$ under both scenarios (i) and (ii).

Results ((ref)) and ((ref)) in Theorem (ref) show that the SSR of the selected model depends on ($i$) the unavoidable uncertainty due to the unobserved error term, $u_{t}$, given by the term $ \bar{\sigma}_{u,T}^{2}$, ($ii$) the cost (in terms of fit) of ignoring the time variation in the coefficients of the signals, $\boldsymbol{\beta }_{it}$ , $i=1,2,\cdots ,k$, as given by the term $\bar{\Delta}_{\beta ,T}$ and $ \bar{\Delta}_{\beta ,T}^{\ast }$, respectively, and ($iii$) the $O_{p}\left( T^{-1/2}\right) $ term due to sampling uncertainty (which will be present even in the absence of variable selection uncertainty), and $(iv)$ the $ O_{p}\left( T^{d-1}\right) $ term which is due to variable selection uncertainty, and will be dominated by $O_{p}\left( T^{-1/2}\right) $ when $ d<1/2$. Therefore, the cost of variable selection can be controlled when using OCMT if the number of pseudo-signals, $k_{T}^{\ast }$, do not rise faster than $\sqrt{T}$. However, to reduce the cost associated with parameter instability more information about the nature of time variations in $\boldsymbol{\beta }_{it}\boldsymbol{\ }$and $\sigma _{ijt,x}$ are required. For example, $\bar{\Delta}_{\beta ,T}$ (or $\bar{\Delta}_{\beta ,T}^{\ast }$) could be lower if $\boldsymbol{\Omega }_{\beta ,t}$ is close to zero in some periods, or if there are cancelling effects from negative $ \sigma _{ijt,x}$ ($\sigma _{ijt,x}^{\ast }$) when $\sigma _{ijt,\beta }$ is positive, namely $\sigma _{ijt,x}\sigma _{ijt,\beta }<0$ ($\sigma _{ijt,x}^{\ast }\sigma _{ijt,\beta }<0$), for some $i\neq j$ and some $t$. This finding for the in-sample fit is similar to the results for mean squared forecast errors in the presence of breaks in the literature, such as Proposition 2 of pesaran2007selection or equation (20) of pesaran2013optimal, where the main focus is to minimize the MSFE by mitigating the cost of parameter instability at the expense of increased sampling uncertainty by appropriate weighting of the observations.

Monte Carlo evidence

We use Monte Carlo (MC) techniques to compare finite sample performance of OCMT with and without down-weighting at the selection stage, as well as comparing the OCMT\ results with those of Lasso, A-Lasso, and boosting. In these comparisons we consider the number of selected covariates ($\hat{k} _{T} $), the true positive rate (TPR), the false positive rate (FPR), and the one-step-ahead mean squared forecast error (MSFE) of the selected models. Sub-section (ref) outlines the MC designs, sub-section (ref) provides a summary of how the OCMT, Lasso, A-Lasso, and boosting procedures are implemented, and finally sub-section (ref) presents the main MC findings. {Details of Lasso, A-Lasso, and boosting procures and how they are implemented are provided in Section (ref) of the online Monte Carlo supplement.}

Simulation design

We consider the following data generating process (DGP):

equation*[equation* omitted — 110 chars of source]

where the four signals $\tilde{x}_{jt},$ $j=1,2,3,4$ have non-zero, time-varying means $\mu _{jt}=\mathbb{E}\left( \tilde{x}_{jt}\right) $. To simplify the exposition of the DGP we consider the demeaned covariates, $ x_{jt}=\tilde{x}_{jt}-\mu _{jt}$ (so that $\mathbb{E}\left( x_{jt}\right) =0$ ), and write the DGP equivalently as

equation[equation omitted — 119 chars of source]

where

equation[equation omitted — 75 chars of source]

Since $c_{t}$ is a free parameter, without loss of generality we also treat $ \left\{ d_{t}\,,\text{ }t=1,2,...,T\right\} $ as free parameters.

For each MC replication, $r=1,2,...,R$, the target variable, $y_{t}$, is generated as random draws using ((ref)). The signal variables $ x_{jt} $, $j=1,2,3,4$, are unknown and belong to a set $\mathcal{S} _{Nt}=\left\{ x_{1t},x_{2t},\cdots ,x_{Nt}\right\} $. The vector of covariates $\mathbf{x}_{t}=\left( x_{1t},x_{2t},\cdots ,x_{Nt}\right) ^{\prime }$ is generated as $\mathbf{x}_{t}=\boldsymbol{R}_{t}^{1/2} \boldsymbol{\varepsilon }_{t}$, where $\boldsymbol{\varepsilon }_{t}=\left( \varepsilon _{1t},\varepsilon _{2t},\cdots ,\varepsilon _{Nt}\right) ^{\prime }$. $\left\{ \varepsilon _{it}\right\} $ are generated as AR(1) processes with GARCH(1,1) innovations

equation*[equation* omitted — 199 chars of source]

using the starting values $\varepsilon _{i,0}\sim IIDN\left( 0,1\right) $. The parameters were generated heterogeneously as independent draws, $ \rho_{i\varepsilon }\sim IIDU\left( 0,0.95\right) $. $e_{\varepsilon _{i}t}\sim IIDN\left( 0,\sigma_{\varepsilon _{i},t}^{2}\right) $, with $ \sigma_{\varepsilon_{i},t}^{2}$ given by

equation*[equation* omitted — 227 chars of source]

where $\alpha _{1\varepsilon _{i}}\sim IIDU(0,0.2)$, and $\alpha _{2\varepsilon _{i}}\sim IIDU(0.6,0.75)$. The error terms, $\left\{ u_{t}\right\} _{t=1}^{T}$, in ((ref)) are generated as $ IIDN(0,\sigma _{ut}^{2})$ with $\sigma _{ut}^{2}$ following the GARCH(1,1) specification

equation*[equation* omitted — 127 chars of source]

using $u_{0}\sim \mathcal{N}(0,1)$, $\alpha _{1u}=0.2$ and $\alpha _{2u}=0.75 $.

As our baseline DGP we consider a model with stable parameters, and set $ \beta _{jt}=1$ for $j=1,2,3,4$. We also set $c_{t}=0$ and $\mu _{jt}=1$ in ( (ref)), which yields $d_{t}=4$. In addition, we set $\rho _{y,t}=0$ when the baseline model is static and $\rho _{y,t}=0.3$ when the baseline model is dynamic. In the dynamic case we set $y_{0}=(1-\rho _{y,1})^{-1}d_{1}$. In the case of models with parameter instability we consider a mixed deterministic-stochastic model and generate $\beta _{jt}$ as

equation*[equation* omitted — 86 chars of source]

where $b_{jt}$ are deterministic and $\eta _{jt}$ are AR(1) processes with GARCH(1,1) innovations,

equation*[equation* omitted — 120 chars of source]

using the starting values $\eta _{j,0}\sim IID\mathcal{N}\left( 0,1\right) $ , and $\rho _{\eta j}=0.5$, for all $j$. $\left\{ e_{\eta _{j}t}\right\} $ follows a normal distribution with mean zero, and variance $\sigma _{\eta _{j}t}^{{2}}$ given by

equation*[equation* omitted — 206 chars of source]

where $\alpha _{1\eta _{j}}=0.2$ and $\alpha _{2\eta _{j}}=0.75$. We set $ \tau _{\eta _{j}}$ such that deterministic variations in $\beta _{jt}$ are quite large relative to the stochastic variations. To this end we set $\tau _{\eta _{j}}$ (using simulations) so that

equation*[equation* omitted — 179 chars of source]

For the deterministic components of the slope coefficients ($b_{jt}$, for $ j=1,2,3,4$), we consider the following specifications

equation[equation omitted — 223 chars of source]

and

equation[equation omitted — 167 chars of source]

where $[.]$ is the nearest integer function.

We \ also set $c_{t}=0$ in ((ref)) and generate the intercept as $ d_{t}=\sum_{j=1}^{k}\beta _{jt}\mu _{jt},$ where

equation[equation omitted — 235 chars of source]

and

equation[equation omitted — 173 chars of source]

In this design, the jumps in $b_{jt}$ and $\mu _{jt}$, for $j=1,2$, have opposite signs and the jumps in $b_{jt}$ and $\mu _{jt}$, for $j=3,4$, have the same sign.

The $N\times N$ correlation matrix of the covariates, $\mathbf{R}_{t}\equiv (r_{ij,t})$, are set as $r_{ij,t}=r_{t}^{|i-j|}$, for all $i,j=1,2,\cdots ,N$ . We allow for a break in the correlation matrix and set $r_{t}$ equal to 0.9 in the first half of the sample and 0.4 in the second half of the sample. Also, we consider two possibilities for $\rho _{y,t}$. In the static scenario we set $\rho _{y,t}=0$ for all $t$. In the dynamic scenario we allow for a switch in $r_{y,t}$ and set it as

equation[equation omitted — 167 chars of source]

For the static and dynamic models with parameter instabilities, the parameter $\tau _{u}$ is calibrated by simulations to ensure that the R-squared of the linear regression of $y_{t}$ on a constant term, the signal variables $\left\{ x_{1t},x_{2t},x_{3t},x_{4t}\right\} $, and (in experiments with $\rho _{y,t}\neq 0$) the lagged dependent variable is equal to $30\%$ (low fit) and $50\%$ (high fit). The same value of $\tau _{u}$ is used for the corresponding static and dynamic models without parameter instabilities.

We base the MC results on $R=2,000$ replications, and consider $N\in \left\{ 20,40,100\right\} $ and $T\in \left\{ 100,200,500\right\} $, combinations. These choices of $\left( N,T\right) $ cover our empirical applications. For each pair of $(N,T)$, there are four experiments in case of the models with no parameter instabilities, and four experiments in the case of models with parameter instabilities, corresponding to the two choices of $\tau _{u}$ (low and high fit), $\rho _{yt}$ (static to dynamic). In total, we carry out eight different experiments.

Selection and estimation methods using weighted and unweighted observations

Let $\mathbf{w}_{t}=(\mathbf{x}_{t}^{\prime },y_{t})^{\prime }$, $ t=1,2,\cdots ,T$ be the (unweighted) set of available observations, and denote the corresponding set of down-weighted observations by $\hat{\mathbf{w }}_{t}(\lambda )=\lambda ^{T-t}\mathbf{w}_{t}$ where $0<\lambda \leq 1$ is the down-weighting coefficient. We are not arguing for the use of exponential down-weighting -- but use it as an example. There are also non-exponential type down-weighting schemes that one can use, e.g. pesaran2013optimal. We will consider the following selection/estimation methods: (1) OCMT with down-weighted observations $\{\hat{\mathbf{w}} _{t}(\lambda )\}_{t=1}^{T}$ used at both selection and estimation stages; (2) OCMT with the unweighted observations, $\{\mathbf{w}_{t}\}_{t=1}^{T}$, used at the selection stage and down-weighted observations, $\{\hat{\mathbf{w }}_{t}(\lambda )\}_{t=1}^{T}$, used at the estimation stage; (3) OCMT using unweighted observations, $\{\mathbf{w}_{t}\}_{t=1}^{T}$, at both selection and estimation stages; (4,5 & 6) Lasso, A-Lasso, and boosting also using unweighted observations, $\{\mathbf{w}_{t}\}_{t=1}^{T}$; and (7,8 & 9) Lasso, A-Lasso, and boosting with down-weighted observations, $\{\hat{ \mathbf{w}}_{t}(\lambda )\}_{t=1}^{T}$ used as inputs.

We also implement a two-step procedures based on Lasso, A-Lasso and boosting. In the first step, we apply Lasso, A-Lasso and boosting to the original (unweighted) observations and select the variables with non-zero coefficients. In the second step, we estimate the corresponding post-selected model by LS using the weighted observations. Overall, the MSFEs of these procedures were higher than that of direct application of Lasso, A-Lasso and boosting to the weighted observations. The results are available in Section S-2 of the online MC supplement.

We consider two sets of values for the down-weighting coefficient, $\lambda $ : (1) Light down-weighting with $\lambda =\left\{ 0.975,0.98,0.985,0.99,0.995,1\right\} $, and (2) Heavy down-weighting with $ \lambda =\left\{ 0.95,0.96,0.97,0.98,0.99,1\right\} $. For each of the above two sets of exponential down-weighting schemes (light/heavy) we focus on simple average forecasts computed over the individual forecasts obtained for each value of $\lambda $ in the set under consideration.

Simulation results

A summary of the main results are provided in Tables (ref) to (ref), with additional summary tables highlighting the effects of down-weighting at the selection stage, and the differences between static versus dynamic models provided in the online MC supplement. Table (ref) give the number of selected covariates ($\hat{k}_{T}$), TPR and FPR of OCMT, Lasso, A-Lasso and boosting without down-weighting. Panel A of this table reports the results for different $N$ and $T$ combinations, averaged across the four experiments without parameter instabilities, and panel B of the table gives the corresponding results for the four experiments with parameter instabilities. The results show that all the methods under consideration have higher average TPR for models with stable parameters compared to the ones with parameter instabilities. This is to be expected, as the models with parameter instabilities are subject to an additional source of uncertainty.

We further observe that the lower average TPR of OCMT in the models with parameter instabilities is associated with a lower average number of selected covariates, and hence a lower average FPR. On the other hand, the other procedures tend, on average, to select more covariates in the models with parameter instabilities and hence have a higher average FPR relative to the models without parameter instabilities. Lastly, OCMT most of the times selects fewer covariates relative to Lasso, A-Lasso, and boosting , while maintaining the TPR at a similar level. As a result, OCMT has mostly the lowest average FPR among the selection methods under consideration. Summary Tables (ref) and (ref) in the online MC supplement provide further results on the effects of down-weighting on TPR and FPR. The results consistently show that down-weighting of observations provides no gains for OCMT in terms of average TPR and FPR. This is also true for other methods in majority but not all cases.

Table (ref) focusses on the one-step-ahead MSFEs and provides comparative results on the effects of down-weighting across the methods (OCMT, Lasso, A-Lasso and boosting). As in Table (ref), Panel A of Table (ref) gives average MSFEs for the four experiments without parameter instabilities, and Panel B gives the corresponding results for the experiments with parameter instabilities. As expected, in the absence of parameter instabilities, using unweighted observations gives the lowest MSFE across all the methods. Moreover, for all $N$ and $T$ combinations and different down-weighting scenarios, the average MSFE of each method is lower in the case of models with stable parameters as compared to those with parameter instabilities. This observation is consistent with our finding in Theorem (ref) about the cost of time-variation in the coefficients on the in-sample fit of the estimated model. As can be seen, for models with parameter instabilities, down-weighting does improve the forecasting performance of OCMT (with and without down-weighting in the selection stage), Lasso, and A-Lasso. However, by comparing the MSFEs of OCMT with and without down-weighting at the selection stage, we see that the down-weighting at the selection stage always results in deterioration of the forecast accuracy of OCMT, which is in line with our main theoretical result. Last but not least, the results in Table (ref) show that OCMT with down-weighting only at the estimation stage almost always has the lowest average MSFE among all the methods for all choices of $N$, $T$, and different down-weighting scenarios. In fact, in the case of experiments with parameter instabilities OCMT with down-weighting (light or heavy) at the estimation stage only, always beats Lasso, A-Lasso and boosting with light or heavy down-weighting in terms of the one-step-ahead MSFE.

Table (ref) compares the performance of OCMT with the down-weighting option at the estimation stage to that of the other procedures, using the same set of down-weighting parameter ($\lambda $). Specifically, we report the MSFE of Lasso, A-Lasso, and boosting relative to that of OCMT. Since the relative MSFE ranking of OCMT, Lasso, A-Lasso, and boosting does not appear to be affected by no/light/heavy down-weighting options, as a summary measure, we simply average relative MSFE values across individual experiments and the three (no/light/heavy) down-weighting options. However, we provide the relative MSFE results for the models without and with parameter instabilities separately, on left and right panels of Table (ref). Two observations stand out from this table. First, the reported average relative MSFEs are almost always greater than one for all the $N$ and $T$ choices, indicating that OCMT outperforms Lasso, A-Lasso, and boosting. Second, the degree to which OCMT outperforms Lasso and A-Lasso tends to increase with the degree of parameter instability. This is less so if we compare OCMT\ with boosting.

Tables (ref), (ref), and (ref) in the online MC supplement provide further details about the performance of the methods under consideration in static and dynamic experiments. In Table (ref), we compare the number of selected covariates, the TPR, and the FPR of each method without down-weighting across static and dynamic models. For various $N$ and $T$ combinations the reported results are averaged across four experiments (with/without parameter instabilities and with/without high-fit). The results show that all the methods tend to select fewer covariates in the dynamic models relative to the static ones, and hence have a lower TPR and FPR. This is expected, as in the dynamic models, part of the variation in the target variable is explained by its own lag rather than the signal variables. Consequently, in Tables (ref) and (ref), which are about the MSFE in static and dynamic models, respectively, we see that all the methods have a higher MSFE in dynamic models relative to the static ones. Additionally, the results in Tables (ref) and (ref) show that the MSFE for models with stable parameters is always lower than the ones with parameter instabilities, regardless of whether the model is static or not.

Overall, the results of our MC studies suggest that the OCMT procedure without down-weighting at the selection stage is a useful method to deal with variable selection in linear regression settings with parameter instability.

Empirical applications

The rest of the paper considers empirical applications whereby the forecast performance of the proposed OCMT approach with no down-weighting at the selection stage is compared with those of Lasso and A-Lasso. In particular, we consider the following two applications:\footnote{ We also consider forecasting euro area quarterly output growth using the European Central Bank (ECB) survey of professional forecasters as our third application. The results of this application can be found in Section (ref) of the online empirical supplement.}

itemize• Forecasting monthly rate of price changes for 28 (out of 30) stocks in Dow Jones using a relatively large number of financial, economic, as well as technical indicators. • Forecasting quarterly output growth rates across 33 countries using macro and financial variables.

In each application, we first compare the performance of OCMT with and without down-weighted observations at the selection stage. We then consider the comparative performance of OCMT (with variable selection carried out without down-weighting) relative to Lasso and A-Lasso, with and without down-weighting. For down-weighting we make use of exponentially down-weighted observations, namely $\hat{x}_{it}(\lambda )=\lambda ^{T-t}x_{it}$, and $\hat{y}_{t}(\lambda )=\lambda ^{T-t}y_{t}$, where $y_{t}$ is the target variable to be forecasted, $x_{it}$, for $i=1,2,...,N$ are the covariates in the active set, and $\lambda $ is the exponential decay coefficient. We consider the same two sets of values for the degree of exponential decay, $\lambda $, as in the MC section: (1) Light down-weighting with $\lambda =\left\{ 0.975,0.98,0.985,0.99,0.995,1\right\} $ , and (2) Heavy down-weighting with $\lambda =\left\{ 0.95,0.96,0.97,0.98,0.99,1\right\} $. For each of the above two sets of exponential down-weighting schemes we again focus on simple average forecasts computed over the individual forecasts obtained for each value of $ \lambda $ in the set under consideration.

For forecast evaluation we consider Mean Squared Forecasting Error (MSFE) and Mean Directional Forecast Accuracy (MDFA), together with related pooled versions of\ Diebold-Mariano (DM), and Pesaran-Timmermann (PT) test statistics. A panel version of diebold2002comparing test is proposed by pesaran2009forecasting. Let $q_{lt}\equiv e_{ltA}^{2}-e_{ltB}^{2}$ be the difference in the squared forecasting errors of procedures $A$ and $B$ , for the target variable $y_{lt}$ ($l=1,2,...,L)$ and $t=1,2,...,T_{l}^{f}$ , where $T_{l}^{f}$ is the number of forecasts for target variable $l$ (could be one or multiple step ahead) under consideration. Suppose $ q_{lt}=\alpha _{l}+\varepsilon _{lt}$ with $\varepsilon _{lt}\sim \mathcal{N} (0,\sigma _{l}^{2})$. Then under the null hypothesis of $H_{0}:\alpha _{l}=0$ for all $l$ we have

equation*[equation* omitted — 266 chars of source]
equation*[equation* omitted — 267 chars of source]

Note that $V(\bar{q})$ needs to be modified in the case of multiple-step ahead forecast errors, due to the serial correlation that results in the forecast errors from the use of over-lapping observations. There is no adjustment needed for one-step ahead forecasting, since it is reasonable to assume that in this case the loss differentials are serially uncorrelated. However, to handle possible serial correlation for $h$-step ahead forecasting with $h>1$, we can modify the panel DM test by using the Newey-West type estimator of $\sigma _{l}^{2}$.

The $MDFA$ statistic compares the accuracy of forecasts in predicting the direction (sign) of the target variable, and is computed as

equation*[equation* omitted — 138 chars of source]

where $\mathbf{1}(w>0)$ is the indicator function takes the value of $1$ when $w>0$ and zero otherwise, $\text{sgn}(w)$ is the sign function, $y_{lt}$ is the actual value of dependent variable at time $t$ and $y_{lt}^{f}$ is its corresponding predicted value. To evaluate statistical significance of the directional forecasts for each method, we also report a pooled version of the test suggested by pesaran1992simple:

equation*[equation* omitted — 110 chars of source]

where $\hat{P}$ is the estimator of the probability of correctly predicting the sign of $y_{lt}$, computed by

equation*[equation* omitted — 226 chars of source]
equation*[equation* omitted — 233 chars of source]

Finally,$\ \hat{V}(\hat{P})=T_{Lf}^{-1}\hat{P}^{\ast }(1-\hat{P}^{\ast }),$ and

equation*[equation* omitted — 264 chars of source]

The last term of $\hat{V}(\hat{P^{\ast }})$ is negligible and can be ignored. Under the null hypothesis, that prediction and realization are independently distributed, PT is asymptotically distributed as a standard normal distribution.

Forecasting monthly returns of stocks in Dow Jones

In this application the focus is on forecasting one-month ahead stock returns, defined as monthly change in natural logarithm of stock prices. We consider stocks that were part of the Dow Jones index in 2017m12, and have non-zero prices for at least 120 consecutive data points (10 years) over the period 1980m1 and 2017m12. We ended up forecasting 28 blue chip stocks. \footnote{ Visa and DowDuPont are excluded since they have less than 10 years of historical price data.} Daily close prices for all the stocks are obtained from Data Stream. For stock $i$, the price at the last trading day of each month is used to construct the corresponding monthly stock prices, $P_{it}$. Finally, monthly returns are computed by $r_{i,t+1}=100\ln(P_{i,t+1}/P_{it}) $, for $i=1,2,...,28$. For all 28 stocks we use an expanding window starting with the observations for the first 10 years ($T=120$). The active set for predicting $r_{i,t+1}$ consists of 40 financial, economic, and technical variables.\footnote{ All regressions include the intercept as the only conditioning (pre-selected) variable.} The full list and the description of the indicators considered can be found in Section (ref) of online empirical supplement.

Overall we computed 8,659 monthly forecasts for the 28 target stocks. The results are summarized as average forecast performances across the different variable selection procedures. Table (ref) reports the effects of down-weighting at the selection stage of the OCMT procedure. It is clear that down-weighting worsens the predictive accuracy of OCMT. From the Panel DM tests, we can also see that down-weighting at the selection stage worsens the forecasts significantly. Panel DM test statistics is -5.606 (-11.352) for light (heavy) versus no down-weighing at the selection stage. Moreover, Table (ref) shows that the OCMT procedure with no down-weighting at the selection stage dominates Lasso, A-Lasso and boosting in terms of MSFE and the differences are statistically highly significant.

Further, OCMT outperforms Lasso, A-Lasso and boosting in terms of Mean Directional Forecast Accuracy (MDFA), measured as the percent number of correctly signed one-month ahead forecasts across all the 28 stocks over the period 1990m2-2017m12. See Table (ref). As can be seen from this table, OCMT\ with no down-weighting performs the best; correctly predicting the direction of 56.057% of 8,659 forecasts, as compared to 55.769%, which we obtain for Lasso, A-Lasso and boosting forecast, at best. This difference is highly significant considering the very large number of forecasts involved. It is also of interest that the better of performance of OCMT is achieved with a much fewer number of selected covariates as compared to Lasso, A-Lasso and boosting. As can be seen from the last column of Table (ref), Lasso, A-Lasso and boosting on average select many more covariates than OCMT (1-15 variables as compared to 0.072 for OCMT).

So far we have focused on average performance across all the 28 stocks. Table (ref) provides the summary results for individual stocks, showing the relative performance of OCMT in terms of the number of stocks, using MSFE and MDFA criteria. The results show that OCMT performs better than Lasso, A-Lasso and boosting in the majority of the stocks in terms of MSFE and MDFA. OCMT outperforms Lasso, A-Lasso and boosting in at least 22 out of 28 stocks in terms of MSFE, under no down-weighting, and almost universally when Lasso, A-Lasso and boosting are implemented with down-weighting. Similar results are obtained when we consider MDFA criteria, although the differences in performance are somewhat less pronounced. Overall, we can conclude that the better average performance of OCMT (documented in Tables (ref) and (ref)) is not driven by a few stocks and holds more generally.

Forecasting quarterly output growth rates across 33 countries

We consider one and two years ahead predictions of output growth for 33 countries (20 advanced and 13 emerging). We use quarterly data from $1979Q2$ to $2016Q4$ taken from the GVAR dataset.\footnote{ The GVAR dataset is available at \url{https://sites.google.com/site/gvarmodelling/data}.} We predict $ \Delta_{4}y_{it}=y_{it}-y_{i,t-4}$, and $\Delta _{8}y_{it}=y_{it}-y_{i,t-8},$ where $y_{it}$, is the log of real output for country $i$. We adopt the following direct forecasting equations:

equation*[equation* omitted — 157 chars of source]

where we consider $h=4$ (one-year-ahead forecasts) and $h=8$ (two-years-ahead forecasts). Given the known persistence in output growth, in addition to the intercept in the present application we also condition on the most recent lagged output growth, denoted by $ \Delta_{1}y_{it}=y_{it}-y_{i,t-1}$, and confine the variable selection to list of variables set out in Table (ref) in the online empirical supplement. Overall, we consider a maximum of 15 covariates in the active set covering quarterly changes in domestic variables such as real output growth, real short term interest rate, and long-short interest rate spread and quarterly change in the corresponding foreign variables.

We use expanding samples, starting with the observations on the first 15 years (60 data points), and evaluate the forecasting performance of the three methods over the period 1997Q2 to 2016Q4.

Tables (ref) and (ref), respectively, report the MSFE of OCMT for one-year and two-year ahead forecasts of output growth, with and without down-weighting at the selection stage. Consistent with the previous application, down-weighting at the selection stage worsens the forecasting accuracy. Moreover, in Tables (ref) and (ref), we can see that OCMT (without down-weighting at the selection stage) outperforms Lasso, A-Lasso and boosting in two-year ahead forecasting. In the case of one-year ahead forecasts, OCMT and Lasso are very close to each other and both outperform A-Lasso and boosting. Table (ref) summarizes country-specific MSFE and DM findings for OCMT relative to Lasso, A-Lasso and boosting. The results show OCMT under-performs Lasso in more than half of the countries for one-year ahead horizon, but outperforms Lasso, A-Lasso and boosting in more than 70 percent of the countries in the case of two-year ahead forecasts. It is worth noting that while Lasso generally outperforms OCMT in the case of one-year ahead forecasts, overall its performance is not statistically significantly better. See Panel DM test of Table (ref). On the other hand we can see from Table (ref) that overall OCMT significantly outperforms Lasso in the case of the two-year ahead forecasts.

Finally in Tables (ref) and (ref) we reports MDFA and PT test statistics for OCMT, Lasso, A-Lasso and boosting. Overall, OCMT has a slightly higher MDFA and hence predicts the direction of real output growth better than Lasso, A-Lasso and boosting in most cases. The PT test statistics suggest that while all the methods perform well in forecasting the direction of one-year ahead real output growth, none of the methods considered are successful at predicting the direction of two-year ahead output growth.

It is also worth noting that as with the previous applications, OCMT selects very few variables from the active set (0.1 on average for both horizons, with the maximum number of selected variables being 2 for $h=4$ and $8$). On the other hand, Lasso on average selects 2.7 variables from the active set for $h=4$, and $1$ variable on average for $h=8$. Maximum number of variables selected by Lasso is $9$ and $13$ for $h=4$, $8$, respectively (out of possible $15$). Again as to be expected, A-Lasso selects a fewer number of variables as compared to Lasso (2.3 and 0.8 on average for $h=4,8$ , respectively), but this does not lead to a better forecast performance in comparison with Lasso. Boosting on average selects $2.7$ variables from the active set for $h=4$, and $1.4$ variables on average for $h=8$.

In conclusion, down-weighting at both selection and forecasting stages deteriorates OCMT's MSFE for both one--year and two-years ahead forecast horizons, as compared to down-weighting only at the forecasting stage. Moreover, light down-weighting at the forecasting stage improves forecasting performance for both horizons. Statistically significant evidence of forecasting skill is found for OCMT relative to Lasso only in the case of two-years ahead forecasts. However, it is interesting that none of the big data methods can significantly beat the simple (light down-weighted) AR(1) baseline model.

Concluding remarks

The penalized regression approach has become the de facto benchmark in the literature on variable selection in the context of linear regression models. But, barring a few exceptions (such as kapetanios2018time ), these studies focus on models with stable parameters, and do not consider the implications of parameter instabilities for variable selection. Recently, chudik2018one proposed OCMT as an alternative procedure to penalized regression. One feature of the OCMT procedure is the fact that the problem of variable selection is separated from the forecasting stage, in contrast to the penalized regression techniques where the variable selection and estimation are carried out simultaneously. Using OCMT one can decide whether to use the weighted observations at the variable selection stage or not, without preempting whether to down-weight and how to down-weight the observations at the forecasting stage.

We have provided theoretical arguments for using the unweighted observations at the selection stage of OCMT, and down-weighted observations at the forecasting stage of OCMT. Our MC results as well as empirical applications uniformly suggest that OCMT without down-weighting at the selection stage outperforms, in terms of mean squared forecast errors, Lasso, Adaptive Lasso, boosting, as well as when OCMT is applied with down-weighted observations.

table[table omitted — 3,665 chars of source]
table[table omitted — 5,688 chars of source]
table[table omitted — 2,089 chars of source]
table[table omitted — 2,060 chars of source]
table[table omitted — 3,776 chars of source]
table[table omitted — 4,319 chars of source]
table[table omitted — 5,679 chars of source]
table[table omitted — 7,215 chars of source]
table[table omitted — 3,858 chars of source]

{\ { \setstretch{1.15} }}

\setcounter{section}{0} \setcounter{table}{0} \setcounter{page}{1} \setcounter{equation}{0} \setcounter{proposition}{0} \setcounter{corollary}{0} \setcounter{proofcorollary}{0} \setcounter{proofproposition}{0}

\quad

center[center omitted — 436 chars of source]

This online theory supplement has three sections. Section (ref) provides the proofs of Theorems 1 to 3, and additional propositions and corollaries. Section (ref) establishes the main lemmas needed for the proof of the theorems in Section (ref). Section (ref) contains the complementary lemmas needed for the proofs of the main lemmas in the previous section.

Notations: Generic finite positive constants are denoted by $C_{i}$ for $i=1,2,\cdots $ and $c $. They can take different values in different instances. $\lVert \mathbf{A}\rVert _{2}$, $\lVert \mathbf{A}\rVert _{F}$, $ \lVert \mathbf{A}\rVert _{\infty}$ and $\lVert \mathbf{A}\rVert _{1}$ denote the spectral, Frobenius, row, and column norms of matrix $\mathbf{A}$, respectively. $\lambda_{i}(\mathbf{A}) $ denotes the $i^{th} $ eigenvalue of a square matrix $\mathbf{A} $. $\text{tr}(\mathbf{A})$ and $\text{det}( \mathbf{A})$ are the trace and determinant of a square matrix $\mathbf{A}$, respectively. $\left\| \mathbf{x} \right\| $ denotes the $\ell_{2} $ norm of vector $\mathbf{x} $. If $\{f_{n}\}_{n=1}^{\infty }$ is any real sequence and $\{g_{n}\}_{n=1}^{\infty }$ is a sequence of positive real numbers, then $f_{n}=O(g_{n})$, if there exists a positive constant $C_{0}$ and $n_0 $ such that $\lvert f_{n}\rvert /g_{n}\leq C_{0}$ for all $n > n_0$. $ f_{n}=o(g_{n})$ if $f_{n}/g_{n}\rightarrow 0$ as $n\rightarrow \infty $. If $ \{f_{n}\}_{n=1}^{\infty }$ and $\{g_{n}\}_{n=1}^{\infty }$ are both positive sequences of real numbers, then $f_{n}=\ominus (g_{n})$ if there exist $ n_{0}\geq 1$ and positive constants $C_{0}$ and $C_{1}$, such that $ \inf_{n\geq n_{0}}\left( f_{n}/g_{n}\right) \geq C_{0}$ and $\sup_{n\geq n_{0}}\left( f_{n}/g_{n}\right) \leq C_{1}$. If $\{f_{n}\}_{n=1}^{\infty }$ is a sequence of random variables and $\{g_{n}\}_{n=1}^{\infty }$ is a sequence of positive real numbers, then $f_{n}=O_p(g_{n})$, if for any $ \varepsilon > 0 $, there exists a positive constant $B_{\varepsilon}$ and $ n_{\varepsilon} $ such that $\Pr \left(\lvert f_{n}\rvert > g_{n} B_{\varepsilon} \right) < \varepsilon $ for all $n > n_{\varepsilon}$.

Proof of the Theorems

This section provides the proofs of Theorems 1 to 3. The proofs are based on lemmas presented in Section (ref). Among these, Lemmas (ref) and (ref) are key. For each covariate $i=1,2,\cdots ,N$, Lemmas (ref) establishes exponential probability inequalities for the t-ratio multiple tests conditional on the average net effect, $\bar{\theta}_{i,T}$, being either of the order $\ominus (T^{-\varepsilon _{i}})$ for some $\varepsilon _{i}>1/2$, or of the order $ \ominus (T^{-\vartheta _{i}})$, for some $0\leq \vartheta _{i}<1/2$. For DGP given by ((ref)), Lemma (ref) provides asymptotic properties of LS estimator of coefficients and SSR of a regression model that includes all the signals and pseudo-signals. This lemma establishes that the coefficients of pseudo-signals estimated by LS converges to zero so long as $k_{T}^{\ast }=\ominus (T^{d})$ grows at a slow rate relative to $T$ , i.e. $0\leq d<1/2$. This lemma also shows that the SSR of the regression model converges to that of the oracle model, which includes only the signals.

flushleftAdditional notations and definitions: Throughout this section we consider the following events: \begin{equation} \textstyle\mathcal{A}_{0}=\mathcal{H}\cap \mathcal{G},\ where\ \mathcal{H}=\left\{ \sum_{i=1}^{k}\hat{\mathcal{J}}_{i}=k\right\} \ and \ \mathcal{G}=\left\{ \sum_{i=k+k_{T}^{\ast }+1}^{N}\hat{\mathcal{J}} _{i}=0\right\} , \end{equation} where $\{\hat{\mathcal{J}}_{i}\text{ for }i=1,2,\cdots ,N\}$ are the selection indicators defined by ((ref)). $\mathcal{A} _{0} $ is the event of selecting the approximating model, defined by ((ref)). $\mathcal{H}$ is the event that all signals are selected, and $\mathcal{G}$ is the event that no noise variable is selected. To simplify the exposition, with slight abuse of notation, we denote the probability of an event $\mathcal{E}$ conditional on $\bar{\theta}_{i,T}$ being of order $\ominus (T^{-a})$ by $\Pr [\mathcal{E}|\bar{\theta} _{i,T}=\ominus (T^{-a})],$where $a $ is a nonnegative constant.

Proof of Theorem (ref)

{\ To establish result ((ref)), first note that $ \mathcal{A}_{0}^{c}=\mathcal{H}^{c}\cup \mathcal{G}^{c}$ and hence ($ \mathcal{H}^{c}$ denotes the complement of $\mathcal{H}$)

equation[equation omitted — 192 chars of source]

where $\mathcal{H}$ and $\mathcal{G}$ are given by ((ref)). We also have $\mathcal{H}^{c}=\{\sum_{i=1}^{k}\hat{\mathcal{J}}_{i}<k\}$ and $ \mathcal{G}^{c}=\{\sum_{i=k+k_{T}^{\ast }+1}^{N}\hat{\mathcal{J}}_{i}>0\}$. Let's consider $\Pr \left( \mathcal{H}^{c}\right) $ and $\Pr \left( \mathcal{ \ G}^{c}\right) $ in turn. We have $\textstyle\Pr (\mathcal{H}^{c})\leq \sum_{i=1}^{k}\Pr (\hat{\mathcal{J}} _{i}=0)$. But for any signal

equation*[equation* omitted — 274 chars of source]

where $0\leq \vartheta _{i}<1/2$ and hence by Lemma (ref), we can conclude that there exist sufficiently large finite positive constants $ C_{0}$ and $C_{1}$ such that $\textstyle\Pr (\hat{\mathcal{J}}_{i}=0)=O\left[ \exp (-C_{0}T^{C_{1}}) \right] $. Since by Assumption (ref), the number of signals is finite we can further conclude that

equation[equation omitted — 98 chars of source]

for some finite positive constants $C_{0}$ and $C_{1}$. In the next step note that

equation*[equation* omitted — 200 chars of source]

But for any noise variable $\textstyle\Pr (\hat{\mathcal{J}}_{i}=1)=\Pr \left[ \lvert t_{i,T}\rvert >c_{p}(N,\delta )|\bar{\theta}_{i,T}=\ominus (T^{-\epsilon _{i}})\right] ,$ where $\epsilon _{i}\geq 1/2$ and hence by Lemma (ref), we can conclude that there exist sufficiently large finite positive constants $C_{0}$ and $C_{1}$ such that for any $0<\pi <1$, $\textstyle\Pr (\hat{\mathcal{J}}_{i}=1)\leq \exp \left[ -\frac{(1-\pi )^{2}\bar{\sigma}_{\eta _{i},T}^{2}\bar{\sigma}_{x_{i},T}^{2}c_{p}^{2}(N, \delta )}{2\bar{\omega}_{iy,T}^{2}(1+d_{T})^{2}}\right] +\exp (-C_{0}T^{C_{1}})$, in which $\bar{\sigma}_{x_{i},T}^{2}=T^{-1}\sum_{t=1}^{T} \mathbb{E}(x_{it}^{2})$, $\bar{\omega}_{iy,T}^{2}=T^{-1}\sum_{t=1}^{T} \mathbb{E}(x_{it}^{2}y_{t}^{2}|\mathcal{F}_{t-1})$, $\bar{\sigma}_{\eta _{i},T}^{2}=T^{-1}\sum_{t=1}^{T}\mathbb{E}(\eta _{it}^{2})$, }$\eta _{it}=y_{t}-\phi _{i,T}x_{it}$, and $\phi _{i,T}$ is defined in ((ref)). {Therefore,

equation*[equation* omitted — 172 chars of source]

where $\mathcal{X}_{NT}=\inf_{i\in \{k+k^{\ast }+1,k+k^{\ast }+2,\cdots ,N\}} \frac{\bar{\sigma}_{\eta _{i},T}^{2}\bar{\sigma}_{x_{i},T}^{2}}{\bar{\omega} _{iy,T}^{2}}$. By result (II) of Lemma (ref) in online theory supplement we can further conclude that for any $0<\pi <1$,

equation[equation omitted — 189 chars of source]

Using ((ref)) and ((ref)) in ((ref)), we obtain $\Pr (\mathcal{A} _{0}^{c})=O\left( N^{1-\mathcal{X}_{NT}\left( \frac{1-\pi }{1+d_{T}}\right) ^{2}\delta }\right) +O\left[ N\exp (-C_{0}T^{C_{1}})\right] $ and $\Pr ( \mathcal{A}_{0})=1-O\left( N^{1-\mathcal{X}_{NT}\left( \frac{1-\pi }{1+d_{T}} \right) ^{2}\delta }\right) -O\left[ N\exp (-C_{0}T^{C_{1}})\right] $, which completes the proof. }

Proof of Theorem (ref)

For any $B>0$,

align*[align* omitted — 510 chars of source]

Since $\Pr \left( T^{\frac{1-d}{2}}\left\Vert \hat{\boldsymbol{\gamma }}_{T}- \boldsymbol{\gamma }_{T}^{\ast }\right\Vert >B|\mathcal{A}_{0}^{c}\right) $ and $\Pr \left( \mathcal{A}_{0}\right) $ are less than or equal to one, we can further write,

equation*[equation* omitted — 330 chars of source]

By conditioning on $\mathcal{A}_{0}$ the dimension of vector $\hat{ \boldsymbol{\gamma }}_{T}$ is at most equal to $k+k_{T}^{\ast}$ and by assumption $k_{T}^{\ast}=\ominus (T^{d})$ where $0\leq d<1/2$. Therefore, by Lemma (ref) in online theory supplement, conditional on $\mathcal{A} _{0} $, $\left\Vert \hat{\boldsymbol{\gamma }}_{T}-\boldsymbol{\gamma } _{T}^{\ast }\right\Vert $ is $O_{p}\left( T^{\frac{d-1}{2}}\right) $. By Theorem 1, we also have $\lim_{T\rightarrow \infty }\Pr \left( \mathcal{A} _{0}^{c}\right) =0$. Hence, for any $\varepsilon >0$, there exists $ B_{\varepsilon }>0$ and $T_{\varepsilon }>0$ such that

equation*[equation* omitted — 261 chars of source]

Therefore, $\Pr \left( T^{\frac{1-d}{2}}\left\Vert \hat{\boldsymbol{\gamma }} _{T}-\boldsymbol{\gamma }_{T}^{\ast }\right\Vert >B_{\varepsilon }\right) <\varepsilon \text{ for all }T>T_{\varepsilon }, $ and we conclude that

equation*[equation* omitted — 159 chars of source]

as required. Similar lines of arguments can be used to show that if $\mathbb{ E}\left( \mathbf{x}_{\tilde{k}_{T},t}\mathbf{x}_{\tilde{k}_{T},t}^{\prime }\right) $ is a fixed time-invariant matrix, then $\left\Vert \hat{\boldsymbol{\gamma }}_{T}-\boldsymbol{\gamma } _{T}^{\diamond }\right\Vert =O_{P}\left( T^{\frac{d-1}{2}}\right)$, which completes the proof.

Proof of Theorem (ref)

Let $D_{T}=T^{-1}\sum_{t=1}^{T}\hat{\eta}_{t}^{2}-\left(\bar{\Delta} _{\beta,T} +\bar{\sigma}_{u,T}^{2}\right) $. For any $B>0$,

align*[align* omitted — 321 chars of source]

Since $\Pr \left( T^{\frac{1}{2}}\left\vert D_{T}\right\vert >B|\mathcal{A} _{0}^{c}\right) $ and $\Pr \left( \mathcal{A}_{0}\right) $ are less than or equal to one, we can further write,

equation*[equation* omitted — 206 chars of source]

By conditioning on $\mathcal{A}_{0}$, the number of selected covariates is at most equal to $k+k_{T}^{\ast }$ and by assumption $k_{T}^{\ast }=\ominus (T^{d})$, where $0\leq d<1/2$. Therefore, by Lemma (ref) in online theory supplement, conditional on $\mathcal{A}_{0}$, $D_{T}$ is $O_{p}\left( T^{- \frac{1}{2}}\right) $. By Theorem 1, we also have $\lim_{T\rightarrow \infty }\Pr \left( \mathcal{A}_{0}^{c}\right) =0$. Hence, for any $ \varepsilon >0$, there exists $B_{\varepsilon }>0$ and $T_{\varepsilon }>0$ such that $\Pr \left( T^{\frac{1}{2}}\left\vert D_{T}\right\vert >B_{\varepsilon }| \mathcal{A}_{0}\right) +\Pr \left( \mathcal{A}_{0}^{c}\right) <\varepsilon, \text{ for all }T>T_{\varepsilon }$. Therefore, $\Pr \left( T^{\frac{1}{2}}\left\vert D_{T}\right\vert >B_{\varepsilon }\right) <\varepsilon \text{ for all }T>T_{\varepsilon }, $ and we conclude that

equation*[equation* omitted — 154 chars of source]

Furthermore, by Lemma (ref), $\bar{\Delta}_{\beta,T} $ is non-negative. Following similar lines of argument we get that if $\mathbb{E} \left( \mathbf{x}_{\tilde{k}_{T},t}\mathbf{x}_{\tilde{k}_{T},t}^{\prime }\right) $ is a fixed time-invariant matrix, then,

equation*[equation* omitted — 161 chars of source]

with $\bar{\Delta}_{\beta,T}^{\ast} \geq 0 $ which completes the proof.

Propositions and corollaries

propositionSuppose the target variable $y_{t}$ is generated according to ((ref)), and Assumptions (ref)-(ref) hold. Consider the following regression equation: \begin{equation} y_{t}=\sum_{i=1}^{k}x_{it}\gamma _{iT}+\eta _{t}=\mathbf{x}_{k,t}^{\prime } \boldsymbol{\gamma }_{T}+\eta _{t}, t=1,2....,T \end{equation} where $\boldsymbol{\gamma }_{T}$ is defined by \begin{equation} \boldsymbol{\gamma }_{T}=\arg \min_{\mathbf{b}}T^{-1}\sum_{t=1}^{T}\mathbb{E} \left( y_{t}-\mathbf{x}_{k,t}^{\prime }\mathbf{b}\right) ^{2}. \end{equation} Then there exists a positive constnt $\epsilon \geq 1/2$, such that \begin{equation*} \boldsymbol{\gamma }_{T}=\left[ T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( \mathbf{ x}_{k,t}\mathbf{x}_{k,t}^{\prime }\right) \right] ^{-1}T^{-1}\sum_{t=1}^{T} \sum_{i=1}^{k}\mathbb{E}\left( \mathbf{x}_{k,t}x_{it}\right) \mathbb{E} \left( \beta _{it}\right) +d_{T}\boldsymbol{\tau }_{k}, \end{equation*} where $d_{T}=O\left( T^{-\epsilon }\right) $, and $\boldsymbol{\tau }_{k}$ is the $k\times 1$ vector of ones. Also, if the expected value of $\beta _{it}$ for $i=1,2,\cdots ,k$ are time-invariant, i.e., $\mathbb{E}\left( \beta _{it}\right) =\beta _{i}$, then $\gamma _{iT}=\beta _{i}+d_{T}$ for $ i=1,2,\cdots ,k$ and there exists $\varrho \geq 1$ such that \begin{equation*} \mathbb{E}\left( \eta _{t}^{2}\right) =\Delta _{\beta ,t}+\mathbb{E} (u_{t}^{2})+e_{T}, \end{equation*} where $e_{T}=O\left( T^{-\varrho }\right) $ \begin{equation} \Delta _{\beta ,t}=\sum_{i=1}^{k}\sum_{j=1}^{k}\sigma _{ijt,x}\sigma _{ij,\beta }=tr\left( \boldsymbol{\Sigma }_{\mathbf{x}_{k},t} \boldsymbol{\Omega }_{\beta ,t}\right) \geq 0, \end{equation} $\boldsymbol{\Sigma }_{\mathbf{x}_{k},t}\equiv \left( \sigma _{ijt,x}\right) $, $\boldsymbol{\Omega }_{\beta ,t}\equiv \left( \sigma _{ijt,\beta }\right) $, for $i,j=1,2,\cdots ,k$, $\sigma _{ijt,x}=\mathbb{E}\left( x_{it}x_{jt}\right) $, and $\sigma _{ijt,\beta }=\mathbb{E}\left[ (\beta _{it}-\beta _{i})(\beta _{jt}-\beta _{j})\right] $. Alternatively, if the covariance matrix of the signals are time-invariant, i.e., $\mathbb{E}\left( \mathbf{x}_{k,t}\mathbf{x} _{k,t}^{\prime }\right) =\boldsymbol{\Sigma }_{x_{k}}$, then $\gamma _{iT}= \bar{\beta}_{iT}+d_{T}$ for $i=1,2,\cdots ,k$, where $\bar{\beta} _{iT}=T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( \beta _{it}\right) $, and there exists $\varrho \geq 1$ such that \begin{equation*} \mathbb{E}\left( \eta _{t}^{2}\right) =\Delta _{\beta ,t}^{\ast }+\mathbb{E} (u_{t}^{2})+e_{T} \end{equation*} where $e_{T}=O\left( T^{-\varrho }\right) $ \begin{equation} \Delta _{\beta ,t}^{\ast }=\sum_{i=1}^{k}\sum_{j=1}^{k}\sigma _{ij,x}\sigma _{ijt,\beta }^{\ast }=tr\left( \boldsymbol{\Omega }_{\beta ,t}^{\ast } \boldsymbol{\Sigma }_{\mathbf{x}_{k}}\right) \geq 0, \end{equation} $\boldsymbol{\Omega }_{\beta ,t}^{\ast }\equiv \left( \sigma _{ijt,\beta }^{\ast }\right) ,$ for $i,j=1,2,\cdots ,k$, and $\sigma _{ijt,\beta }^{\ast }=\mathbb{E}\left[ (\beta _{it}-\bar{\beta}_{i,T})(\beta _{jt}-\bar{\beta} _{j,T})\right] $.
remarkProposition (ref) shows that in a linear regression model that does not consider parameter instability, the deviation of each coefficient from the simple time-average of the corresponding coefficients in the DGP approaches zero. Moreover, $\Delta _{\beta ,t}\geq 0$ and $\Delta _{\beta ,t}^{\ast }\geq 0$ represent the costs, in mean squared error sense, of neglecting parameter instability.
corollaryLet $y_{t}$ for $t=1,2,\cdots ,T$ be generated by ( (ref)), and consider the active set $\mathcal{S}_{Nt}= \{x_{1t},x_{2t},\cdots ,x_{Nt}\}$ which contains $k$ signals, $k_{T}^{\ast }$ pseudo-signals, and $N-k-k_{T}^{\ast }$ noise variables. Suppose Assumptions (ref)-(ref) hold and the noise variables, $x_{it}$ $ i=k+k_{T}^{\ast }+1,k+k_{T}^{\ast }+2,\cdots ,N$, are independent of the target, $y_{t}$, and have time-invaraint unconditional variances, $\mathbb{V} \left( x_{it}^{2}\right) =\sigma _{i}^{2}$ for $i=k+k_{T}^{\ast }+1,k+k_{T}^{\ast }+2,\cdots ,N$, and $N=\ominus (T^{\kappa })$ with $\kappa >0$. Then, there exist finite positive constants $C_{0}$ and $C_{1}$ such that, for any $\pi $ in $(0,1)$ and any null sequence $d_{T}>0$, the probability of selecting the approximating model $\mathcal{A}_{0}$, defined by ((ref)), by the OCMT procedure with the critical value function $c_{p}(N,\delta )$ given by ((ref)), for some $ \delta >0$, is given by \begin{equation} \Pr (\mathcal{A}_{0})=1-O\left[ T^{\kappa \left( 1-\left( \frac{1-\pi }{ 1+d_{T}}\right) ^{2}\delta \right) }\right] -O\left[ T^{\kappa }\exp \left( -C_{0}T^{C_{1}}\right) \right] . \end{equation}
remarkCorollary (ref) shows that if we further assume that the noise variables are independent of $y_{t} $ and their variance does not change over time, then for any $\delta > 1 $, the OCMT procedure consistently selects the approximating model.

Proof of propositions and corollaries

proofpropositionSince the objective function for $\boldsymbol{\gamma }_{T}$ is convex and, by Assumption (ref), $ T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( \mathbf{x}_{k,t}\mathbf{x} _{k,t}^{\prime }\right) $ is invertible, then by the first-order condition of the minimization in ((ref)) we have \begin{equation*} \boldsymbol{\gamma }_{T}=\left[ T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( \mathbf{ x}_{k,t}\mathbf{x}_{k,t}^{\prime }\right) \right] ^{-1}T^{-1}\sum_{t=1}^{T} \mathbb{E}\left( \mathbf{x}_{k,t}y_{t}\right) . \end{equation*} Substituting $y_{t}$ from ((ref)), now yields \begin{equation*} \boldsymbol{\gamma }_{T}=\left[ T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( \mathbf{ x}_{k,t}\mathbf{x}_{k,t}^{\prime }\right) \right] ^{-1}T^{-1}\sum_{t=1}^{T} \sum_{i=1}^{k}\mathbb{E}\left( \mathbf{x}_{k,t}x_{it}\beta _{it}\right) + \left[ T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( \mathbf{x}_{k,t}\mathbf{x} _{k,t}^{\prime }\right) \right] ^{-1}T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( \mathbf{x}_{k,t}u_{it}\right) . \end{equation*} By part (c) of Assumptions (ref), all the elements of the $k\times 1$ vector $T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( \mathbf{x}_{k,t}u_{it}\right) $ are $O\left( T^{-\epsilon }\right) $ for some $\epsilon \geq 1/2$. Moreover, by Assumptions (ref) and (ref), all the element of $k\times k$ matrix $\left[ T^{-1}\sum_{t=1}^{T}\mathbb{E} \left( \mathbf{x}_{k,t}\mathbf{x}_{k,t}^{\prime }\right) \right] ^{-1}$ are finite fixed numbers. Since, by Assumption (ref), the number of signals, $k$, is a finite fixed number, we can further conclude that all the elements of $k\times 1$ vector, \begin{equation*} \left[ T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( \mathbf{x}_{k,t}\mathbf{x} _{k,t}^{\prime }\right) \right] ^{-1}T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( \mathbf{x}_{k,t}u_{it}\right) , \end{equation*} are $O\left( T^{-\epsilon }\right) $ for some $\epsilon \geq 1/2$ and consequently we can write \begin{equation*} \boldsymbol{\gamma }_{T}=\left[ T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( \mathbf{ x}_{k,t}\mathbf{x}_{k,t}^{\prime }\right) \right] ^{-1}T^{-1}\sum_{t=1}^{T} \sum_{i=1}^{k}\mathbb{E}\left( \mathbf{x}_{k,t}x_{it}\beta _{it}\right) +d_{T}\boldsymbol{\tau }_{k}, \end{equation*} where $d_{T}=O\left( T^{-\epsilon }\right) $ for some $\epsilon \geq 1/2$ and $\boldsymbol{\tau }_{k}$ is the $k\times 1$ vector of ones. By Assumption (ref), $\beta _{it}$ is independent of $x_{jt}$ for all $ i,j=1,2,\cdots ,k $, therefore, \begin{equation*} \boldsymbol{\gamma }_{T}=\left[ T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( \mathbf{ \ x}_{k,t}\mathbf{x}_{k,t}^{\prime }\right) \right] ^{-1}T^{-1} \sum_{t=1}^{T}\sum_{i=1}^{k}\mathbb{E}\left( \mathbf{x}_{k,t}x_{it}\right) \mathbb{E}\left( \beta _{it}\right) +d_{T}\boldsymbol{\tau }_{k}. \end{equation*} Consider first the case where $\mathbb{E}\left( \beta _{it}\right) $ is time-invariant and set $\mathbb{E}\left( \beta _{it}\right) =\beta _{i}$. Then \begin{align*} \boldsymbol{\gamma }_{T}=& \left[ T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( \mathbf{x}_{k,t}\mathbf{x}_{k,t}^{\prime }\right) \right] ^{-1}T^{-1} \sum_{t=1}^{T}\sum_{i=1}^{k}\mathbb{E}\left( \mathbf{x}_{k,t}x_{it}\right) \beta _{i}+d_{T}\boldsymbol{\tau }_{k} \\ =& \left[ T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( \mathbf{x}_{k,t}\mathbf{x} _{k,t}^{\prime }\right) \right] ^{-1}T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( \mathbf{x}_{k,t}\sum_{i=1}^{k}x_{it}\beta _{i}\right) +d_{T}\boldsymbol{\tau }_{k} \\ =& \left[ T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( \mathbf{x}_{k,t}\mathbf{x} _{k,t}^{\prime }\right) \right] ^{-1}T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( \mathbf{x}_{k,t}\mathbf{x}_{k,t}^{\prime }\boldsymbol{\beta }\right) +d_{T} \boldsymbol{\tau }_{k} \\ =& \left[ T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( \mathbf{x}_{k,t}\mathbf{x} _{k,t}^{\prime }\right) \right] ^{-1}\left[ T^{-1}\sum_{t=1}^{T}\mathbb{E} \left( \mathbf{x}_{k,t}\mathbf{x}_{k,t}^{\prime }\right) \right] \boldsymbol{ \ \beta }+d_{T}\boldsymbol{\tau }_{k}=\boldsymbol{\beta }+d_{T}\boldsymbol{\ \tau }_{k}, \end{align*} where $\boldsymbol{\beta }=\left( \beta _{1},\beta _{2},\cdots ,\beta _{k}\right) ^{\prime }$. So, in this case $\gamma _{iT}$ would converge to the expected value of $\beta _{it}$ at $T\rightarrow \infty $. Moreover, \begin{equation*} \eta _{t}=y_{t}-\sum_{i=1}^{k}x_{it}(\beta _{i}+d_{T}). \end{equation*} By substituting for $y_{t}$ from ((ref)), we have \begin{equation*} \eta _{t}=\sum_{i=1}^{k}x_{it}\left( \beta _{it}-\beta _{i}\right) +u_{t}+d_{T}\sum_{i=1}^{k}x_{it}. \end{equation*} Therefore, by Assumptions (ref) and (ref), \begin{equation*} \mathbb{E}\left( \eta _{t}^{2}\right) =\sum_{i=1}^{k}\sum_{j=1}^{k}\sigma _{ijt,x}\sigma _{ijt,\beta }+\mathbb{E}(u_{t}^{2})+e_{T}. \end{equation*} where $e_{T}=O\left( T^{-\varrho }\right) $ for some $\varrho \geq 1$, $ \sigma _{ijt,x}=\mathbb{E}\left( x_{it}x_{jt}\right) $, $\sigma _{ijt,\beta }=\mathbb{E}\left[ (\beta _{it}-\beta _{i})(\beta _{jt}-\beta _{j})\right] $ . We further have \begin{equation*} \Delta _{\beta ,t}=\sum_{i=1}^{k}\sum_{j=1}^{k}\sigma _{ijt,x}\sigma _{ijt,\beta }=tr\left( \boldsymbol{\Omega }_{\beta ,t}\boldsymbol{\ \Sigma }_{\mathbf{x}_{k},t}\right) , \end{equation*} where $\boldsymbol{\Omega }_{\beta ,t}\equiv \left( \sigma _{ijt,\beta }\right) $ and $\boldsymbol{\Sigma }_{\mathbf{x}_{k},t}\equiv \left( \sigma _{ijt,x}\right) $ for $i,j=1,2,\cdots ,k$. By result 9(b) on page 44 of Lutkepohl1996handbook, we can further write \begin{equation*} tr\left( \boldsymbol{\Omega }_{\beta ,t}\boldsymbol{\Sigma }_{\mathbf{ \ x}_{k},t}\right) \geq k\left[ det\left( \boldsymbol{\Omega }_{\beta ,t}\right) \right] ^{1/k}\left[ det\left( \boldsymbol{\Sigma }_{ \mathbf{x}_{k},t}\right) \right] ^{1/k}. \end{equation*} But $k$ is a finite fixed integer. Furthermore, $\text{det}\left( \boldsymbol{\Omega }_{\beta ,t}\right) \geq 0$ and $\text{det}\left( \boldsymbol{\Sigma }_{\mathbf{x}_{k},t}\right) >0$, since $\boldsymbol{\ \Omega }_{\beta ,t}$ and $\boldsymbol{\Sigma }_{\mathbf{x}_{k},t}$ are positive semi-definite and positive definite matrices, respectively. So, we can conclude that $\Delta _{\beta ,t}\geq 0$. Consider now a second case where $\mathbb{E}\left( \mathbf{x}_{k,t}\mathbf{x} _{k,t}^{\prime }\right) $ is time-invariant and set $\mathbb{E}\left( \mathbf{x}_{k,t}\mathbf{x}_{k,t}^{\prime }\right) =\boldsymbol{\Sigma }_{{x} _{k}}$. Then \begin{align*} \boldsymbol{\gamma }_{T}=& \left[ T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( \mathbf{x}_{k,t}\mathbf{x}_{k,t}^{\prime }\right) \right] ^{-1} \sum_{i=1}^{k} \mathbb{E}\left( \mathbf{x}_{k,t}x_{it}\right) \left[ T^{-1}\sum_{t=1}^{T} \mathbb{E}\left( \beta _{it}\right) \right] +d_{T} \boldsymbol{\tau }_{k} \\ =& \left[ T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( \mathbf{x}_{k,t}\mathbf{x} _{k,t}^{\prime }\right) \right] ^{-1}\sum_{i=1}^{k}\mathbb{E}\left( \mathbf{ x }_{k,t}x_{it}\bar{\beta}_{iT}\right) +d_{T}\boldsymbol{\tau }_{k} \\ =& \left[ T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( \mathbf{x}_{k,t}\mathbf{x} _{k,t}^{\prime }\right) \right] ^{-1}\mathbb{E}\left( \mathbf{x}_{k,t} \mathbf{x}_{k,t}^{\prime }\bar{\boldsymbol{\beta }}_{T}\right) +d_{T} \boldsymbol{\tau }_{k} \\ =& \left[ T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( \mathbf{x}_{k,t}\mathbf{x} _{k,t}^{\prime }\right) \right] ^{-1}\left[ T^{-1}\sum_{t=1}^{T}\mathbb{E} \left( \mathbf{x}_{k,t}\mathbf{x}_{k,t}^{\prime }\right) \right] \bar{ \boldsymbol{\beta }}_{T}+d_{T}\boldsymbol{\tau }_{k}=\bar{\boldsymbol{\beta } }_{T}+d_{T}\boldsymbol{\tau }_{k} \end{align*} where $\bar{\boldsymbol{\beta }}_{T}=\left( \bar{\beta}_{1T},\bar{\beta} _{2T},\cdots ,\bar{\beta}_{kT}\right) ^{\prime }$ and $\bar{\beta} _{iT}=T^{-1}\sum_{t=1}^{k}\mathbb{E}\left( \beta _{it}\right) $. So, in this case $\gamma _{iT}$ would converge to the simple average of expected value of $\beta _{it}$ across time. Moreover, \begin{equation} \eta _{t}=y_{t}-\sum_{i=1}^{k}x_{it}(\bar{\beta}_{iT}+d_{T}) \end{equation} By substituting for $y_{t}$ from ((ref)), we have \begin{equation} \eta _{t}=\sum_{i=1}^{k}x_{it}\left( \beta _{it}-\bar{\beta}_{iT}\right) +u_{t}+d_{T}\sum_{i=1}^{k}x_{it}. \end{equation} Therefore, by Assumptions (ref) and (ref), \begin{equation} \mathbb{E}\left( \eta _{t}\right) ^{2}=\sum_{i=1}^{k}\sum_{j=1}^{k}\sigma _{ij,x}\sigma _{ijt,\beta }^{\ast }+\mathbb{E}(u_{t}^{2})+e_{T}, \end{equation} where $e_{T}=O\left( T^{-\varrho }\right) $ for some $\varrho \geq 1$, $ \sigma _{ij,x}=\mathbb{E}\left( x_{it}x_{jt}\right) $ and $\sigma _{ijt,\beta }^{\ast }=\mathbb{E}\left[ (\beta _{it}-\bar{\beta}_{i,T})(\beta _{jt}-\bar{\beta}_{j,T})\right] $. We further have \begin{equation*} \Delta _{\beta ,t}^{\ast }=\sum_{i=1}^{k}\sum_{j=1}^{k}\sigma _{ij,x}\sigma _{ijt,\beta }^{\ast }=tr\left( \boldsymbol{\Omega }_{\beta ,t}^{\ast } \boldsymbol{\Sigma }_{\mathbf{x}_{k}}\right) \end{equation*} where $\boldsymbol{\Omega }_{\beta ,t}^{\ast }\equiv \left( \sigma _{ijt,\beta }^{\ast }\right) $ and $\boldsymbol{\Sigma }_{\mathbf{x} _{k}}\equiv \left( \sigma _{ij,x}\right) $ for $i,j=1,2,\cdots ,k$. By result 9(b) on page 44 of Lutkepohl1996handbook, we can further write \begin{equation*} tr\left( \boldsymbol{\Omega }_{\beta ,t}^{\ast }\boldsymbol{\Sigma } _{ \mathbf{x}_{k}}\right) \geq k\left[ \text{det}\left( \boldsymbol{\Omega } _{\beta ,t}^{\ast }\right) \right] ^{1/k}\left[ \text{det}\left( \boldsymbol{ \ \Sigma }_{\mathbf{x}_{k}}\right) \right] ^{1/k}. \end{equation*} But $k$ is a finite fixed integer. Furthermore, $\text{det}\left( \boldsymbol{\Omega }_{\beta ,t}^{\ast }\right) \geq 0$ and $\text{det}\left( \boldsymbol{\Sigma }_{\mathbf{x}_{k},t}\right) >0$, since $\boldsymbol{\ \Omega }_{\beta ,t}^{\ast }$ and $\boldsymbol{\Sigma }_{\mathbf{x}_{k},t}$ are positive semi-definite and positive definite matrices, respectively. So, we can conclude that $\Delta _{\beta ,t}^{\ast }\geq 0$.
proofcorollaryBy Theorem (ref), we have that under Assumptions (ref)-(ref), there exist finite positive constants $C_{0}$ and $ C_{1} $ such that, for any $0<\pi <1$, the probability of selecting the approximating model $\mathcal{A}_{0}$, as defined by ((ref)), is given by \begin{equation} \Pr (\mathcal{A}_{0})=1-O\left[ T^{\kappa \left( 1-\mathcal{X}_{NT}\left( \frac{1-\pi }{1+d_{T}}\right) ^{2}\delta \right) }\right] -O\left[ T^{\kappa }\exp \left( -C_{0}T^{C_{1}}\right) \right] , \end{equation} where \begin{equation*} \textstyle\mathcal{X}_{NT}=\inf_{i\in \{k+k^{\ast }+1,\cdots ,N\}}\frac{\bar{ \sigma}_{\eta _{i},T}^{2}\bar{\sigma}_{x_{i},T}^{2}}{\bar{\omega}_{iy,T}^{2}} . \end{equation*} with $\bar{\sigma}_{x_{i},T}^{2}=T^{-1}\sum_{t=1}^{T}\mathbb{E}(x_{it}^{2})$ , $\bar{\omega}_{iy,T}^{2}=T^{-1}\sum_{t=1}^{T}\mathbb{E} (x_{it}^{2}y_{t}^{2}|\mathcal{F}_{t-1})$, $\bar{\sigma}_{\eta _{i},T}^{2}=T^{-1}\sum_{t=1}^{T}\mathbb{E}(\eta _{it}^{2})$, $\eta _{it}=y_{t}-\phi _{i,T}x_{it}$, and $\phi _{i,T}$ is defined in ((ref)). Note that, \begin{align*} \bar{\sigma}_{\eta _{i},T}^{2}& =T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( y_{t}^{2}\right) +\phi _{i,T}^{2}\left[ T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( x_{it}^{2}\right) \right] -2\phi _{i,T}\left[ T^{-1}\sum_{t=1}^{T}\mathbb{E} \left( x_{it}y_{t}\right) \right] \\ & =\bar{\sigma}_{y,T}^{2}+\phi _{i,T}^{2}\bar{\sigma}_{x_{i},T}^{2}-2\phi _{i,T}\bar{\theta}_{i,T}=\bar{\sigma}_{y,T}^{2}-\phi _{i,T}^{2}\bar{\sigma} _{x_{i},T}^{2}. \end{align*} But, $x_{it}$ for all $i\in \{k+k_{T}^{\ast }+1,k+k_{T}^{\ast }+2,\cdots ,N_{T}\}$ are independent of $y_{t}$ and hence $\phi _{i,T}=0$. Consequently, $\bar{\sigma}_{\eta _{i},T}^{2}=\bar{\sigma}_{y,T}^{2}$ for $ i\in \{k+k_{T}^{\ast }+1,k+k_{T}^{\ast }+2,\cdots ,N_{T}\}$. Moreover, since $x_{it}$ for all $i\in \{k+k_{T}^{\ast }+1,k+k_{T}^{\ast }+2,\cdots ,N_{T}\}$ are independent of $y_{t}$, we can write \begin{equation*} \bar{\omega}_{iy,T}^{2}=T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( x_{it}^{2}| \mathcal{F}_{t-1}\right) \mathbb{E}\left( y_{t}^{2}|\mathcal{F}_{t-1}\right) =T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( x_{it}^{2}\right) \mathbb{E}\left( y_{t}^{2}\right) , \end{equation*} for $i\in \{k+k_{T}^{\ast }+1,k+k_{T}^{\ast }+2,\cdots ,N_{T}\}$. Therefore, \begin{equation*} \textstyle\mathcal{X}_{NT}=\inf_{i\in \{k+k^{\ast }+1,\cdots ,N\}}\frac{\bar{ \sigma}_{y,T}^{2}\bar{\sigma}_{x_{i},T}^{2}}{T^{-1}\sum_{t=1}^{T}\mathbb{E} \left( x_{it}^{2}\right) \mathbb{E}\left( y_{t}^{2}\right) }, \end{equation*} In cases where $\mathbb{E}(x_{it}^{2})$ for $i\in \{k+k_{T}^{\ast }+1,k+k_{T}^{\ast }+2,\cdots ,N_{T}\}$ are time-invariant, we can conclude that $\mathcal{X}_{NT}=1$ and hence the probability of selecting the approximating model is given by \begin{equation*} \Pr (\mathcal{A}_{0})=1-O\left[ T^{\kappa \left( 1-\left( \frac{1-\pi }{ 1+d_{T}}\right) ^{2}\delta \right) }\right] -O\left[ T^{\kappa }\exp \left( -C_{0}T^{C_{1}}\right) \right] , \end{equation*} as required. Note that $d_{T}\rightarrow 0$ and $T\rightarrow \infty $ and $ \pi $ is an arbitrary constant between zero and one.

Since $\pi $ is an arbitrary constant between zero and one, result ((ref)) of Corollary (ref) implies that for any $\delta >1$, we can select an approximating model with probability approaching one as $N$ and $T$ grows to infinity.

Main lemmas

lemmaLet $y_{t}$ be a target variable generated by equation ((ref)), and $x_{it}$ be a covariate in the active set $\mathcal{S} _{Nt}=\{x_{1t},x_{2t}, \cdots ,x_{Nt}\}$. Under Assumptions (ref), and (ref), we have \begin{equation*} \mathbb{E}\left[ y_{t}x_{it}-\mathbb{E}(y_{t}x_{it})|\mathcal{F}_{t-1}\right] =0, for i=1,2,\cdots ,N, \end{equation*} and \begin{equation*} \mathbb{E}\left[ y_{t}^{2}-\mathbb{E}(y_{t}^{2})|\mathcal{F}_{t-1}\right] =0. \end{equation*}
proofFor $i=1,2,\cdots ,N$, we have \begin{equation*} \mathbb{E}(y_{t}x_{it}|\mathcal{F}_{t-1})\textstyle= \sum_{j=1}^{k}\mathbb{E} (\beta_{jt}|\mathcal{F}_{t-1}) \mathbb{E}(x_{jt}x_{it}|\mathcal{F}_{t-1})+ \mathbb{E}(u_{t}x_{it}| \mathcal{F}_{t-1}). \end{equation*} By Assumption (ref), $\mathbb{E}(\beta _{jt}| \mathcal{F}_{t-1})=\mathbb{E }(\beta _{jt})$, $\mathbb{E}(x_{jt}x_{it}| \mathcal{F}_{t-1})=\mathbb{E} (x_{jt}x_{it})$, and $\mathbb{E}(u_{t}x_{it}| \mathcal{F}_{t-1})=\mathbb{E} (u_{t}x_{it})$. Therefore, \begin{equation*} \textstyle\mathbb{E}(y_{t}x_{it}|\mathcal{F}_{t-1})= \sum_{j=1}^{k} \mathbb{E }(\beta _{jt})\mathbb{E}(x_{jt}x_{it})+\mathbb{E}(u_{t}x_{it})= \mathbb{E} (y_{t}x_{it}). \end{equation*} Also to establish the last result, note that $y_{t}$ can be written as \begin{equation*} \textstyle y_{t}= \sum_{j=1}^{k}\beta _{jt}x_{jt}+u_{t} = \mathbf{x} _{k,t}^{\prime }\boldsymbol{\beta}_{t} + u_{t}, \end{equation*} where $\mathbf{x}_{k,t}=(x_{1t},x_{2t},\cdots ,x_{kt})^{\prime }$, and $ \boldsymbol{\beta }_{t}=(\beta _{1t},\beta _{2t},\cdots ,\beta _{kt})^{\prime }$. Hence, \begin{equation*} \begin{split} \mathbb{E}(y_{t}^{2}|\mathcal{F}_{t-1})& \textstyle=\mathbb{E}(\boldsymbol{ \beta}_{t}^{\prime }|\mathcal{F}_{t-1})\mathbb{E}(\mathbf{x}_{t}\mathbf{x} _{t}^{\prime }|\mathcal{F}_{t-1})\mathbb{E}(\boldsymbol{\beta}_{t}| \mathcal{ F}_{t-1})+\mathbb{E}(u_{t}^{2}|\mathcal{F}_{t-1})+2\mathbb{E}( \boldsymbol{ \beta}_{t}^{\prime }|\mathcal{F}_{t-1})\mathbb{E}(\mathbf{x}_{t}u_{t}| \mathcal{F}_{t-1}) \\ & \textstyle=\mathbb{E}(\boldsymbol{\beta}_{t}^{\prime })\mathbb{E}(\mathbf{x }_{t}\mathbf{x}_{t}^{\prime })\mathbb{E}(\boldsymbol{\beta}_{t})+ \mathbb{E} (u_{t}^{2})+2\mathbb{E}(\boldsymbol{\beta}_{t}^{\prime })\mathbb{E}(\mathbf{x }_{t}u_{t})=\mathbb{E}(y_{t}^{2}). \end{split} \end{equation*}
lemmaLet $y_{t}$ be a target variable generated by equation ((ref)). Under Assumptions (ref)-(ref), for any value of $ \alpha >0$, there exist some positive constants $C_{0}$ and $C_{1}$ such that \begin{equation*} \sup_{t}\Pr (\lvert y_{t}\rvert >\alpha )\leq C_{0}\exp (-C_{1}\alpha ^{s/2}). \end{equation*}
proofNote that \begin{equation*} \textstyle\lvert y_{t}\rvert \leq \sum_{j=1}^{k}\lvert \beta _{jt}x_{jt}\rvert +\lvert u_{t}\rvert . \end{equation*} Therefore, \begin{equation*} \textstyle\Pr (\lvert y_{t}\rvert >\alpha )\leq \sum_{j=1}^{k}\lvert \beta _{jt}x_{jt}\rvert +\lvert u_{t}\rvert >\alpha ), \end{equation*} and by Lemma (ref) for any $0<\pi _{i}<1$, $i=1,2,\cdots ,k +1$, with $\sum_{i=1}^{k+1}\pi_{j}=1$, we can further write \begin{equation*} \Pr (\lvert y_{t}\rvert >\alpha )\textstyle\leq \sum_{j=1}^{k}\Pr (\lvert \beta _{jt}x_{jt}\rvert >\pi _{j}\alpha )+\Pr (\lvert u_{t}\rvert >\pi_{k+1}\alpha ). \end{equation*} Moreover, by Lemma (ref), we have \begin{equation*} \begin{split} \Pr (\lvert \beta _{jt}x_{jt}\rvert >\pi _{j}\alpha )& \leq \Pr [\lvert x_{jt}\rvert >(\pi _{j}\alpha )^{1/2}]+\Pr [\lvert \beta _{jt}\rvert >(\pi _{i}\alpha )^{1/2}], \end{split} \end{equation*} and hence \begin{equation*} \begin{split} \Pr (\lvert y_{t}\rvert >\alpha )& \textstyle\leq \sum_{j=1}^{k}\Pr [\lvert x_{jt}\rvert >(\pi _{j}\alpha )^{1/2}]+\sum_{j=1}^{k}\Pr [\lvert \beta _{jt}\rvert >(\pi _{j}\alpha )^{1/2}]+\Pr (\lvert u_{t}\rvert >\pi _{k+1}\alpha ), \end{split} \end{equation*} Therefore, under Assumptions (ref)-(ref), we can conclude that for any value of $\alpha >0$, there exist some positive constants $C_{0}$ and $C_{1}$ such that \begin{equation*} \sup_{t}\Pr (\lvert y_{t}\rvert >\alpha )\leq C_{0}\exp (-C_{1}\alpha ^{s/2}) . \end{equation*}
lemmaLet $x_{it}$ be a covariate in the active set, $\mathcal{S}_{Nt}=\{x_{1t},x_{2t},\cdots ,x_{Nt}\}$. Suppose Assumptions (ref)-(ref) hold and $\zeta_T = \ominus(T^{\lambda})$ for some $\lambda > 0$. Then, if $0 < \lambda \leq (s+2)/(s+4)$, for any $0 <\pi < 1$, \begin{equation*} \Pr (|\mathbf{x}_{i}^{\prime}\mathbf{x}_{j}-\mathbb{E}(\mathbf{x} _{i}^{\prime }\mathbf{x}_{j})|>\zeta _{T})\leq \exp\left(- \frac{(1-\pi)^{2} \zeta_{T}^2}{2 T \bar{\omega}_{i j,T}^2 } \right), \end{equation*} where, $\mathbf{x}_{i}=(x_{i1},x_{i2},\cdots ,x_{iT})^{\prime }$ and $\bar{ \omega}_{i j,T}^2 = T^{-1}\sum_{t=1}^{T} \mathbb{E}\left(x_{it}^2 x_{jt}^2 | \mathcal{F}_{t-1}\right) $. Also, if $\lambda >(s+2)/(s+4)$, there exists a finite positive constant $C_{1}$, \begin{equation*} \Pr (|\mathbf{x}_{i}^{\prime }\mathbf{x}_{j}-\mathbb{E}(\mathbf{x} _{i}^{\prime }\mathbf{x}_{j})|>\zeta_{T}) \leq \exp \left( -C_{1}\zeta_{T}^{s/(s+1)}\right), \end{equation*} for all $i,j = 1, 2, \cdots, N $. .
proofNote that $[\mathbf{x}_{i}^{\prime}\mathbf{x}_{j} -\mathbb{E}(\mathbf{x} _{i}^{\prime} \mathbf{x}_{j})] = \sum_{t=1}^{T}[x_{it} x_{jt} - \mathbb{E} (x_{it} x_{jt})]$, for all $i$ and $j$. By Assumption (ref) we have \begin{equation*} \mathbb{E}\left[ x_{it}x_{jt}-\mathbb{E}(x_{it}x_{jt})|\mathcal{F}_{t-1} \right] =0, \end{equation*} for $i,j=1,2,\cdots ,N$. Moreover, by Assumption (ref), for all $i = 1,2, \cdots, N $ and $\alpha > 0$, there exist some finite positive constants $C_0$ and $C_1 $ such that, \begin{equation*} \sup_{t}\Pr (|x_{i t}|>\alpha )\leq C_{0}\exp (-C_{1}\alpha^{s}). \end{equation*} Therefore, by Lemma (ref), for all $i,j = 1,2, \cdots, N $ and $\alpha > 0$, \begin{equation*} \sup_{t}\Pr (|x_{it}x_{jt}|>\alpha )\leq C_{0}\exp (-C_{1}\alpha ^{s/2}). \end{equation*} Hence, by Lemma (ref), if $0<\lambda \leq (s+2)/(s+4) $, for any $0 < \pi < 1$, \begin{equation*} \textstyle \Pr (|\mathbf{x}_{i}^{\prime}\mathbf{x}_{j} -\mathbb{E}(\mathbf{x} _{i}^{\prime} \mathbf{x}_{j})|>\zeta_{T}) \leq \exp\left(- \frac{(1-\pi)^{2} \zeta_{T}^2}{2 T \bar{\omega}_{i j,T}^2 } \right). \end{equation*} Moreover, if $\lambda >(s+2)/(s+4)$, then there exists a finite positive constant $C_{1}$, such that \begin{equation*} \textstyle\Pr (|\mathbf{x}_{i}^{\prime}\mathbf{x}_{j} -\mathbb{E}(\mathbf{x} _{i}^{\prime} \mathbf{x}_{j})|>\zeta_{T}) \leq \exp \left( -C_{1}\zeta _{T}^{s/(s+1)}\right). \end{equation*}
lemmaLet $y_{t}$ be a target variable generated by the DGP given by ((ref)) and $x_{it}$ be a covariate in the active set, $\{x_{1t},x_{2t},\cdots ,x_{Nt}\}$. Suppose Assumptions (ref)- (ref) hold and $\zeta_T =\ominus(T^{\lambda})$ for some $\lambda > 0$. Then, if $0<\lambda \leq (s+4)/(s+8)$, for any $0 < \pi < 1$, \begin{equation*} \Pr(|\mathbf{x}_{i}^{\prime }\mathbf{y}-\theta_{i,T}| > \zeta _{T})\leq \exp\left(- \frac{(1-\pi)^{2} \zeta_{T}^2}{2 T \bar{\omega}_{iy,T}^2 } \right), \end{equation*} where $\mathbf{x}_{i}=(x_{i1},x_{i2},\cdots,x_{iT})^{\prime }$, $\mathbf{y} =(y_{1},y_{2},\cdots,y_{T})^{\prime }$, $\theta _{i,T}=T\bar{\theta}_{i,T}= \mathbb{E}(\mathbf{x}_{i}^{\prime }\mathbf{y})$ and $\bar{\omega}_{i y,T}^2 = T^{-1} \sum_{t=1}^{T} \mathbb{E}\left(x_{it}^2 y_{t}^2 | \mathcal{F}_{t-1} \right) $. Also, if $\lambda >(s+4)/(s+8)$, there exists a finite positive constant $C_1$ such that \begin{equation*} \Pr (|\mathbf{x}_{i}^{\prime }\mathbf{y}-\theta_{i,T} | > \zeta _{T}) \leq \exp \left(-C_{1}\zeta_{T}^{s/(s+1)}\right), \end{equation*} for all $i=1,2,\cdots ,N$.
proofNote that $[\mathbf{x}_{i}^{\prime}\mathbf{y} - \theta_{i,T}] = \sum_{t=1}^{T}[x_{it} y_{t} - \mathbb{E}(x_{it} y_{t})]$, for all $i$. By Lemma (ref) \begin{equation*} \mathbb{E}\left[ x_{it}y_{t}-\mathbb{E}(x_{it}y_{t})|\mathcal{F}_{t-1} \right] =0, \end{equation*} for $i=1,2,\cdots ,N$. Moreover, by Assumption (ref), for all $i = 1,2, \cdots, N $ and $\alpha > 0$, there exist some finite positive constants $ C_0 $ and $C_1 $ such that, \begin{equation*} \sup_{t}\Pr (|x_{i t}|>\alpha )\leq C_{0}\exp (-C_{1}\alpha^{s}). \end{equation*} Also, by Lemma (ref), there exist some finite positive constants $C_0$ and $C_1 $ such that, \begin{equation*} \sup_{t}\Pr (\lvert y_{t}\rvert >\alpha )\leq C_{0}\exp (-C_{1}\alpha ^{s/2}) . \end{equation*} Therefore, by Lemma (ref), for all $i = 1,2, \cdots, N $ and $ \alpha > 0$, \begin{equation*} \sup_{t}\Pr (|x_{it}y_{t}|>\alpha )\leq C_{0}\exp (-C_{1}\alpha ^{s/4}). \end{equation*} Hence, by Lemma (ref), if $0<\lambda \leq (s+4)/(s+8) $, for any $0 < \pi < 1$, \begin{equation*} \textstyle \Pr (|\mathbf{x}_{i}^{\prime}\mathbf{y} -\theta_{i,T}|>\zeta_{T}) \leq \exp\left(- \frac{(1-\pi)^{2} \zeta_{T}^2}{2 T \bar{\omega}_{i y,T}^2 } \right). \end{equation*} Moreover, if $\lambda >(s+4)/(s+8)$, there exists a finite positive constant $C_{1}$, \begin{equation*} \textstyle\Pr (|\mathbf{x}_{i}^{\prime}\mathbf{y}_{j} -\theta_{i,T}|>\zeta_{T}) \leq \exp \left( -C_{1}\zeta _{T}^{s/(s+1)}\right). \end{equation*}
lemmaLet $y_{t}$ be a target variable generated by equation ( (ref)) and $x_{it}$ be a covariate in the active set, $\mathcal{S} _{Nt}=\{x_{1t},x_{2t},\cdots ,x_{Nt}\}$. Suppose Assumptions (ref)- (ref) hold and $\zeta_T =\ominus(T^{\lambda})$ for some $\lambda > 0$. Consider the projection regression of $y_{t}$ on $x_{it}$ as \begin{equation*} y_{t}=\phi _{i,T}x_{it}+\eta _{it}, \end{equation*} where the projection coefficient $\phi _{i,T}$ is given by ((ref)). Then, if $0<\lambda \leq(s+4)/(s+8)$, there exist sufficiently large positive constants $C_{0}$, $C_{1}$ and $C_{2}$ such that \begin{equation*} \Pr \left[ \lvert \boldsymbol{\eta }_{i}^{\prime }\mathbf{M}_{x_{i}} \boldsymbol{\eta }_{i}-\mathbb{E}(\boldsymbol{\eta }_{i}^{\prime } \boldsymbol{\eta }_{i})\rvert >\zeta _{T}\right] \leq \exp (-C_{0}T^{-1}\zeta _{T}^{2})+\exp (-C_{1}T^{C_{2}}), \end{equation*} where $\boldsymbol{\eta }_{i}=(\eta _{i1},\eta _{i2},\cdots ,\eta _{iT})^{\prime }$ and $\mathbf{M}_{x_{i}}=\mathbf{I}-T^{-1}\mathbf{x} _{i}(T^{-1}\mathbf{x}_{i}^{\prime }\mathbf{x}_{i})^{-1}\mathbf{x} _{i}^{\prime }$ with $\mathbf{x}_{i}=(x_{i1},x_{i2},\cdots ,x_{iT})^{\prime } $. Also, if $\lambda >(s+4)/(s+8)$, there exist sufficiently large positive constants $C_{0}$, $C_{1}$ and $C_{2}$ such that \begin{equation*} \Pr \left[ \lvert \boldsymbol{\eta }_{i}^{\prime }\mathbf{M}_{x_{i}} \boldsymbol{\eta }_{i}-\mathbb{E}(\boldsymbol{\eta }_{i}^{\prime } \boldsymbol{\eta }_{i})\rvert >\zeta _{T}\right] \leq \exp (-C_{0}\zeta _{T}^{s/(s+1)})+\exp (-C_{1}T^{C_{2}}), \end{equation*} for all $i=1,2,\cdots ,N$.
proofNote that $\boldsymbol{\eta }_{i}^{\prime }\mathbf{M}_{x_{i}}\boldsymbol{ \eta }_{i}=\mathbf{y}^{\prime }\mathbf{M}_{x_{i}}\mathbf{y}$, where $\mathbf{ y}=(y_{1},y_{2},\cdots ,y_{T})^{\prime }$. By Assumption (ref), we have \begin{equation*} \mathbb{E}\left[ x_{it}^2-\mathbb{E}(x_{it}^2)|\mathcal{F}_{t-1} \right] =0, \end{equation*} for $i=1,2,\cdots ,N$. By Lemma (ref), we also have \begin{equation*} \mathbb{E}\left[ y_{t}x_{it}-\mathbb{E}(y_{t}x_{it})|\mathcal{F}_{t-1}\right] =0, \end{equation*} for $i=1,2,\cdots ,N$, and \begin{equation*} \mathbb{E}\left[ y_{t}^{2}-\mathbb{E}(y_{t}^{2})|\mathcal{F}_{t-1}\right] =0. \end{equation*} Moreover, by Assumption (ref), for all $i = 1,2, \cdots, N $ and $ \alpha > 0$, there exist some finite positive constants $C_0$ and $C_1 $ such that, \begin{equation*} \sup_{t}\Pr (|x_{i t}|>\alpha )\leq C_{0}\exp (-C_{1}\alpha^{s}). \end{equation*} Also, by Lemma (ref), there exist some finite positive constants $C_0$ and $C_1 $ such that, \begin{equation*} \sup_{t}\Pr (\lvert y_{t}\rvert >\alpha )\leq C_{0}\exp (-C_{1}\alpha ^{s/2}) . \end{equation*} Therefore by Lemma (ref), we can conclude that there exist sufficiently large positive constants $C_{0}$, $C_{1}$, and $C_{2}$ such that if $0<\lambda \leq (s+4)/(s+8)$, then \begin{equation*} \Pr \left[ \lvert \boldsymbol{\eta }_{i}^{\prime }\mathbf{M}_{x_{i}} \boldsymbol{\eta }_{i}-\mathbb{E}(\boldsymbol{\eta }_{i}^{\prime } \boldsymbol{\eta }_{i})\rvert >\zeta _{T}\right] \leq \exp (-C_{0}T^{-1}\zeta _{T}^{2})+\exp (-C_{1}T^{C_{2}}), \end{equation*} and if $\lambda >(s+4)/(s+8)$, then \begin{equation*} \Pr \left[ \lvert \boldsymbol{\eta }_{i}^{\prime }\mathbf{M}_{x_{i}} \boldsymbol{\eta }_{i}-\mathbb{E}(\boldsymbol{\eta }_{i}^{\prime } \boldsymbol{\eta }_{i})\rvert >\zeta _{T}\right] \leq \exp (-C_{0}\zeta _{T}^{s/(s+1)})+\exp (-C_{1}T^{C_{2}}), \end{equation*} for all $i=1,2,\cdots ,N$.
lemmaLet $y_{t}$ be a target variable generated by equation ( (ref)) and $x_{it}$ be a covariate in the active set, $\mathcal{S} _{Nt}=\{x_{1t},x_{2t},\cdots ,x_{Nt}\}$. Suppose Assumptions (ref)- (ref) hold and consider the projection regression of $y_{t}$ on $x_{it}$ as \begin{equation} y_{t}=\phi _{i,T}x_{it}+\eta _{it}, \end{equation} where $\phi _{i,T}$ is given in ((ref)). Define, \begin{equation*} t_{i,T}=\frac{T^{-1/2}\mathbf{x}_{i}^{\prime }\mathbf{y}}{\sqrt{T^{-1} \boldsymbol{\eta }_{i}^{\prime }\mathbf{M}_{x_{i}}\boldsymbol{\eta }_{i}} \sqrt{T^{-1}\mathbf{x}_{i}^{\prime }\mathbf{x}_{i}}}, \end{equation*} where $\mathbf{x}_{i}=(x_{i1},x_{i2},\cdots ,x_{iT})^{\prime }$, $\mathbf{y} =(y_{1},y_{2}\allowbreak ,\cdots ,y_{T})^{\prime }$, $\boldsymbol{\eta } _{i}=(\eta _{i1},\eta _{i2},\cdots ,\eta _{iT})^{\prime }$ and $\mathbf{M} _{x_{i}}=\mathbf{I}-T^{-1}\mathbf{x}_{i}(T^{-1}\mathbf{x}_{i}^{\prime } \mathbf{x}_{i})^{-1}\mathbf{x}_{i}^{\prime }$. Then, there exist sufficiently large finite positive constants $C_{0}$ and $C_{1}$ such that for any $0<\pi <1$, any null sequence $d_{T} >0 $, and $\epsilon _{i}\geq \frac{1}{2}$ \begin{equation*} \Pr \left[ \lvert t_{i,T}\rvert >c_{p}(N,\delta )|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right] \leq \exp \left[ -\frac{(1-\pi )^{2}\bar{\sigma} _{\eta _{i},T}^{2}\bar{\sigma}_{x_{i},T}^{2}c_{p}^{2}(N,\delta )}{2\bar{ \omega}_{iy,T}^{2}(1+d_{T})^{2}}\right] +\exp (-C_{0}T^{C_{1}}), \end{equation*} where $c_{p}(N,\delta )$ is defined by ((ref)), $\theta _{i,T}=T \bar{\theta}_{i,T}=\mathbb{E}(\mathbf{x}_{i}^{\prime }\mathbf{y})$, $\bar{ \sigma}_{\eta _{i},T}^{2}=T^{-1}\sum_{t=1}^{T}\mathbb{E}\left( \eta _{it}^{2}\right) $, $\bar{\sigma}_{x_{i},T}^{2}=T^{-1}\sum_{t=1}^{T}\mathbb{E }\left( x_{it}^{2}\right) $ and $\bar{\omega}_{iy,T}^{2}=T^{-1}\sum_{t=1}^{T} \mathbb{E}\left( x_{it}^{2}y_{t}^{2}|\mathcal{F}_{t-1}\right) $. Also, if $ c_{p}(N,\delta )=o(T^{1/2-\vartheta -c})$ for any $0\leq \vartheta_{i} <1/2$ and a finite positive constant $c$, then, there exist some finite positive constants $C_{0}$ and $C_{1}$ such that \begin{equation*} \textstyle\Pr \left[ \lvert t_{i,T}\rvert >c_{p}(N,\delta )|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right] \geq 1-\exp (-C_{0}T^{C_{1}}). \end{equation*}
proofWe have $\lvert t_{i,T}\rvert =\mathcal{A}_{iT}\mathcal{B}_{iT}$, where, \begin{equation*} \mathcal{A}_{iT}=\frac{|T^{-1/2}\mathbf{x}_{i}^{\prime }\mathbf{y}|}{\bar{ \sigma} _{\eta _{i,T}} \bar{\sigma}_{x_{i,T}}}, \end{equation*} and \begin{equation*} \mathcal{B}_{iT}=\frac{\bar{\sigma}_{\eta _{i},T}\bar{\sigma}_{x_{i},T}}{ \sqrt{T^{-1}\boldsymbol{\eta }_{i}^{\prime }\mathbf{M}_{x_{i}}\boldsymbol{ \eta }_{i}}\sqrt{T^{-1}\mathbf{x}_{i}^{\prime }\mathbf{x}_{i}}}. \end{equation*} In the first case where $\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})$ for some $\epsilon _{i}\geq 1/2$, by using Lemma (ref) we have \begin{equation*} \begin{split} \Pr \left[ \lvert t_{i,T}\rvert >c_{p}(n,\delta )|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right] & \leq \Pr \left[ \mathcal{A} _{iT}>c_{p}(N,\delta )/(1+d_{T})|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}}) \right] + \\ & \quad \ \Pr \left[ \mathcal{B}_{iT}>1+d_{T}|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right] , \end{split} \end{equation*} where $d_{T}\rightarrow 0$ as $T\rightarrow \infty $. By using Lemma (ref), \begin{equation*} \begin{split} & \Pr \left[ \mathcal{B}_{iT}>1+d_{T}|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right] \\ & \qquad \leq \Pr \left( \lvert \frac{\bar{\sigma}_{\eta _{i},T}\bar{\sigma} _{x_{i},T}}{\sqrt{T^{-1}\boldsymbol{\eta }_{i}^{\prime }\mathbf{M}_{x_{i}} \boldsymbol{\eta }_{i}}\sqrt{T^{-1}\mathbf{x}_{i}^{\prime }\mathbf{x}_{i}}} -1\rvert >d_{T}|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right) \\ & \qquad \leq \Pr \left( \lvert \frac{(T^{-1}\boldsymbol{\eta }_{i}^{\prime } \mathbf{M}_{x_{i}}\boldsymbol{\eta }_{i})(T^{-1}\mathbf{x}_{i}^{\prime } \mathbf{x}_{i})}{\bar{\sigma}_{\eta _{i},T}^{2}\bar{\sigma}_{x_{i},T}^{2}} -1\rvert >d_{T}|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right) \\ & \textstyle\qquad =\Pr \left[ \mathcal{M}_{iT}+\mathcal{R}_{iT}+\mathcal{M} _{iT}\mathcal{R}_{iT}>d_{T}|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}}) \right] \end{split} \end{equation*} where $\mathcal{M}_{iT}=|(T^{-1}\mathbf{x}_{i}^{\prime }\mathbf{x}_{i})/\bar{ \sigma}_{x_{i},T}^{2}-1|$ and $\mathcal{R}_{iT}=|(T^{-1}\boldsymbol{\eta } _{i}^{\prime }\mathbf{M}_{x_{i}}\boldsymbol{\eta }_{i})/\bar{\sigma}_{\eta _{i},T}^{2}-1|$. By using Lemmas (ref) and (ref) , for any values of $0<\pi _{i}<1$ with $\sum_{i=1}^{3}\pi _{i}=1$ and a strictly positive constant, $c$, we have \begin{equation*} \begin{split} & \Pr \left[ \mathcal{B}_{iT}>1+d_{T}|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right] \\ & \textstyle\qquad \leq \Pr \left[ \mathcal{M}_{iT}>\pi _{1}d_{T}|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right] +\Pr \left[ \mathcal{R}_{iT}>\pi _{2}d_{T}|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right] + \\ & \textstyle\qquad \quad \ \Pr \left[ \mathcal{M}_{iT}>\frac{\pi _{3}}{c} d_{T}|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right] +\Pr \left[ \mathcal{R}_{iT}>c|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right] . \end{split} \end{equation*} First, consider $\Pr \left[ \mathcal{M}_{iT}>\pi _{1}d_{T}|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right] $, and note that \begin{equation*} \Pr \left[ \mathcal{M}_{iT}>\pi _{1}d_{T}|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right] =\Pr \left[ |\mathbf{x}_{i}^{\prime }\mathbf{x} _{i}-\mathbb{E}(\mathbf{x}_{i}^{\prime }\mathbf{x}_{i})|>\pi _{1}\bar{\sigma} _{x_{i},T}^{2}Td_{T}|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right] . \end{equation*} Therefore, by Lemma (ref), there exist some constants $C_{0}$ and $C_{1}$ such that, \begin{equation*} \Pr \left[ \mathcal{M}_{iT}>\pi _{1}d_{T}|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right] \leq \exp (-C_{0}T^{C_{1}}). \end{equation*} Similarly, \begin{equation*} \textstyle\Pr \left[ \mathcal{M}_{iT}>\frac{\pi _{3}}{c}d_{T}|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right] \leq \exp (-C_{0}T^{C_{1}}). \end{equation*} Also note that \begin{equation*} \Pr \left[ \mathcal{R}_{iT}>\pi _{2}d_{T}|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right] =\Pr \left[ |\boldsymbol{\eta }_{i}^{\prime } \mathbf{M}_{x_{i}}\boldsymbol{\eta }_{i}-\mathbb{E}(\boldsymbol{\eta } _{i}^{\prime }\boldsymbol{\eta }_{i})|>\pi _{2}\bar{\sigma}_{\eta _{i},T}^{2}Td_{T}|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right] . \end{equation*} Therefore, by Lemma (ref), there exist some constants $C_{0}$ and $C_{1}$ such that, \begin{equation*} \Pr \left[ \mathcal{R}_{iT}>\pi _{2}d_{T}|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right] \leq \exp (-C_{0}T^{C_{1}}). \end{equation*} Similarly, \begin{equation*} \textstyle\Pr \left[ \mathcal{R}_{iT}>c|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right] \leq \exp (-C_{0}T^{C_{1}}). \end{equation*} Therefore, we can conclude that there exist some constants $C_{0}$ and $ C_{1} $ such that, \begin{equation*} \Pr \left[ \mathcal{B}_{iT}>1+d_{T}|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right] \leq \exp (-C_{0}T^{C_{1}}) \end{equation*} Now consider $\Pr \left[ \mathcal{A}_{iT}>c_{p}(N,\delta )/(1+d_{T})|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right] $, which is equal to \begin{eqnarray*} &&\Pr \left( \frac{\left\vert \mathbf{x}_{i}^{\prime }\mathbf{y}-\theta _{i,T}+\theta _{i,T}\right\vert }{\bar{\sigma}_{\eta _{i},T}\bar{\sigma} _{x_{i},T}}>T^{1/2}\frac{c_{p}(N,\delta )}{1+d_{T}}|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right) \\ &&\qquad \ \leq \Pr \left( |\mathbf{x}_{i}^{\prime }\mathbf{y}-\theta _{i,T}|>\frac{\bar{\sigma}_{\eta _{i},T}\bar{\sigma}_{x_{i},T}}{1+d_{T}} T^{1/2}c_{p}(N,\delta )-\lvert \theta _{i,T}\rvert \ |\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right) . \end{eqnarray*} Note that since $\epsilon _{i}\geq 1/2$ and $c_{p}(N,\delta )\rightarrow \infty $ as N and consequently T goes to infinity, the first term on the right hand side of the inequality dominate the second one. Moreover, Since $ c_{p}(N,\delta )=o(T^{\lambda })$ for all values of $\lambda >0$, by Lemma (ref), for any $0<\pi <1$, \begin{equation*} \textstyle\Pr \left[ |\mathbf{x}_{i}^{\prime }\mathbf{y}|>\frac{\bar{\sigma} _{\eta _{i},T}\bar{\sigma}_{x_{i},T}}{1+d_{T}}T^{1/2}c_{p}(N,\delta )|\theta _{i,T}=\ominus (T^{1-\epsilon _{i}})\right] \leq \exp \left[ -\frac{(1-\pi )^{2}\bar{\sigma}_{\eta _{i},T}^{2}\bar{\sigma}_{x_{i},T}^{2}c_{p}^{2}(N, \delta )}{2\bar{\omega}_{iy,T}^{2}(1+d_{T})^{2}}\right] . \end{equation*} Given the probability upper bound for $\mathcal{A}_{iT}$ and $\mathcal{B} _{iT}$, we can conclude that there exist some finite positive constants $ C_{0}$ and $C_{1}$ such that for any $0 < \pi < 1$ \begin{equation*} \Pr \left[ \lvert t_{i,T}\rvert >c_{p}(N,\delta )|\theta _{i,T}=\ominus(T^{1-\epsilon _{i}})\right] \leq \exp \left[-\frac{ (1-\pi)^{2} \bar{\sigma}_{\eta_{i},T}^{2} \bar{\sigma}_{x_{i},T}^{2} c_{p}^{2}(N,\delta)}{2 \bar{\omega}_{iy, T}^{2} (1+d_{T})^{2}} \right] + \exp (-C_{0}T^{C_{1}}). \end{equation*} Let's consider the next case where $\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})$ for some $0\leq \vartheta _{i}<1/2$. We know that \begin{equation*} \Pr \left[ \lvert t_{i,T}\rvert >c_{p}(N,\delta )|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right] =1-\Pr \left[ \lvert t_{i,T} \rvert <c_{p}(N,\delta )|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right] . \end{equation*} By Lemma (ref), \begin{equation*} \begin{split} & \Pr \left[ \lvert t_{i,T}\rvert <c_{p}(N,\delta )|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right] \leq \Pr \left[ \mathcal{A}_{iT}<\sqrt{1+d_{T}} c_{p}(N,\delta )|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right] + \\ & \qquad \qquad \qquad \qquad \qquad \qquad \ \Pr \left[ \mathcal{B}_{iT}<1/ \sqrt{1+d_{T}}|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right] . \end{split} \end{equation*} Since $\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})$, for some $0\leq \vartheta _{i}<1/2$ and $c_{p}(N,\delta )=o(T^{1/2-\vartheta -c})$, for any $ 0 \leq \vartheta < 1/2 $, $|\theta _{i,T}| - \bar{\sigma}_{\eta_{i},T} \bar{ \sigma}_{x_{i},T}[(1+d_{T})T]^{1/2}c_{p}(N,\delta ) = \ominus (T^{1-\vartheta _{i}})>0$ and by Lemma (ref), we have \begin{equation*} \begin{split} & \Pr \left[ \mathcal{A}_{iT}<\sqrt{1+d_{T}}c_{p}(N,\delta )|\theta_{i,T} = \ominus (T^{1-\vartheta_{i}})\right] \\ & \qquad = \Pr \left[ \frac{|T^{-1/2}\mathbf{x}_{i}^{\prime}\mathbf{y} -T^{-1/2}\theta_{i,T} + T^{-1/2} \theta_{i,T}|}{\bar{\sigma}_{\eta_{i},T} \bar{\sigma}_{x_{i},T}} < \sqrt{1+d_{T}} c_{p}(N,\delta) | \theta_{i,T} = \ominus (T^{1-\vartheta _{i}}) \right] \\ & \qquad \leq \Pr \left[ |\mathbf{x}_{i}^{\prime }\mathbf{y}-\theta _{i,T}|>|\theta_{i,T}|- \bar{\sigma}_{\eta_{i},T}\bar{\sigma} _{x_{i},T}[(1+d_{T})T]^{1/2}c_{p}(N,\delta )|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right] . \end{split} \end{equation*} Therefore, by Lemma (ref), there exist some finite positive constants $C_{0}$ and $C_{1}$ such that, \begin{equation*} \Pr \left[ |\mathbf{x}_{i}^{\prime }\mathbf{y}-\theta _{i,T}|>|\theta _{i,T}|-\bar{\sigma}_{\eta_{i},T} \bar{\sigma} _{x_{i},T}[(1+d_{T})T]^{1/2}c_{p}(N,\delta )|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right] \leq \exp (-C_{0}T^{C_{1}}), \end{equation*} and therefore \begin{equation*} \Pr \left[ \mathcal{A}_{iT}<\sqrt{1+d_{T}}c_{p}(N,\delta )|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right] \leq \exp (-C_{0}T^{C_{1}}). \end{equation*} Now let consider the probability of $\mathcal{B}_{iT}$, \begin{equation*} \begin{split} & \Pr \left( \mathcal{B}_{iT}<1/\sqrt{1+d_{T}}|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right) \\ & =\Pr \left( \frac{\bar{\sigma}_{\eta_{i},T} \bar{\sigma}_{x_{i},T}}{\sqrt{ T^{-1} \boldsymbol{\eta }_{i}^{\prime }\mathbf{M}_{x_{i}}\boldsymbol{\eta } _{i}}\sqrt{T^{-1}\mathbf{x}_{i}^{\prime }\mathbf{x}_{i}}}<\frac{1}{\sqrt{ 1+d_{T}}}|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right) \\ & \qquad =\Pr \left( \frac{(T^{-1}\boldsymbol{\eta }_{i}^{\prime }\mathbf{M} _{x_{i}}\boldsymbol{\eta }_{i})(T^{-1}\mathbf{x}_{i}^{\prime }\mathbf{x}_{i}) }{\bar{\sigma}_{\eta_{i},T}^{2}\bar{\sigma}_{x_{i},T}^{2}}>1+d_{T}|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right) \\ & \textstyle\qquad \leq \Pr (\mathcal{M}_{iT}+\mathcal{R}_{iT}+\mathcal{M} _{iT}\mathcal{R}_{iT}>d_{T}|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})) , \end{split} \end{equation*} where $\mathcal{M}_{iT}=|(T^{-1}\mathbf{x}_{i}^{\prime}\mathbf{x}_{i})/\bar{ \sigma}_{x_{i},T}^{2}-1|$ and $\mathcal{R}_{iT}=|(T^{-1}\boldsymbol{\eta} _{i}^{\prime }\mathbf{M}_{x_{i}}\boldsymbol{\eta }_{i})/\bar{\sigma}_{\eta _{i},T}^{2}-1|$. By using Lemmas (ref) and (ref) , for any values of $0<\pi _{i}<1$ with $\sum_{i=1}^{3}\pi _{i}=1$ and a positive constant, $c$, we have \begin{equation*} \begin{split} & \Pr \left[ \mathcal{B}_{iT}<1/\sqrt{1+d_{T}}|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right] \\ & \textstyle\qquad \leq \Pr \left[ \mathcal{M}_{iT}>\pi _{1}d_{T}|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right] +\Pr \left[ \mathcal{R} _{iT}>\pi _{2}d_{T}|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right] + \\ & \textstyle\qquad \quad \ \Pr \left[ \mathcal{M}_{iT}>\frac{\pi _{3}}{c} d_{T}|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right] +\Pr \left[ \mathcal{R}_{iT}>c|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right] . \end{split} \end{equation*} Let's first consider the $\Pr \left[ \mathcal{M}_{iT}>\pi _{1}d_{T}|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right] $. Note that \begin{equation*} \Pr \left[ \mathcal{M}_{iT}>\pi _{1}d_{T}|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right] =\Pr \left[ |\mathbf{x}_{i}^{\prime }\mathbf{x} _{i}-\mathbb{E}(\boldsymbol{x}_{i}^{\prime }\boldsymbol{x}_{i})|>\pi _{1} \bar{\sigma}_{x_{i},T}^{2}Td_{T}|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}}) \right] . \end{equation*} So, by Lemma (ref), we know that there exist some constants $C_{0}$ and $C_{1}$ such that, \begin{equation*} \Pr \left[ \mathcal{M}_{iT}>\pi _{1}d_{T}|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right] \leq \exp (-C_{0}T^{C_{1}}). \end{equation*} Similarly, \begin{equation*} \textstyle\Pr \left[ \mathcal{M}_{iT}>\frac{\pi _{3}}{c}d_{T}|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right] \leq \exp (-C_{0}T^{C_{1}}). \end{equation*} Also note that \begin{equation*} \Pr \left[ \mathcal{R}_{iT}>\pi _{2}d_{T}|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right] =\Pr \left[ |\boldsymbol{\eta }_{i}^{\prime } \mathbf{M}_{x_{i}}\boldsymbol{\eta }_{i}-\mathbb{E}(\boldsymbol{\eta } _{i}^{\prime }\boldsymbol{\eta }_{i})|>\pi _{2}\bar{\sigma} _{\eta_{i},T}^{2}Td_{T}|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right] . \end{equation*} Therefore, by Lemma (ref), there exist some constants $C_{0}$ and $C_{1}$ such that, \begin{equation*} \Pr (\mathcal{R}_{iT}>\pi _{2}d_{T}|\theta _{i,T} = \ominus (T^{1-\vartheta _{i}}))\leq \exp (-C_{0}T^{C_{1}}). \end{equation*} Similarly, \begin{equation*} \textstyle\Pr (\mathcal{R}_{iT}>c|\theta _{i,T} =\ominus (T^{1-\vartheta _{i}}))\leq \exp (-C_{0}T^{C_{1}}). \end{equation*} Therefore, we can conclude that there exist some constants $C_{0}$ and $ C_{1} $ such that, \begin{equation*} \Pr \left[ \mathcal{B}_{iT}<1/\sqrt{1+d_{T}}|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right] \leq \exp (-C_{0}T^{C_{1}}). \end{equation*} So, overall we conclude that \begin{equation*} \begin{split} & \Pr \left[ \lvert t_{i,T}\rvert >c_{p}(N,\delta )|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right] \\ & \qquad \qquad =1-\Pr \left[ \lvert t_{i,T} \rvert <c_{p}(N,\delta )|\theta _{i,T}=\ominus (T^{1-\vartheta _{i}})\right] \geq 1-\exp (-C_{0}T^{C_{1}}). \end{split} \end{equation*}
lemmaSuppose $y_{t}$ are generated by \begin{equation} y_{t}=\sum_{i=1}^{k}x_{it}\beta _{it}+u_{t} for t=1,2,\cdots ,T, \end{equation} and consider the LS estimator of the following regression augmented with the additional $l_{T}$ regressors from the active set: \begin{equation*} y_{t}=\mathbf{x}_{kt}^{\prime }\boldsymbol{\phi }+\mathbf{s}_{t}^{\prime } \boldsymbol{\delta }_{T}+\eta _{t}, \end{equation*} where $\mathbf{x}_{kt}=(x_{1t},x_{2t},\cdots ,x_{kt})^{\prime }$, is the $ k\times 1$ vector of signals, $\mathbf{s}_{t}$ is the $l_{T}\times 1$ vector of additional regressors, $\boldsymbol{\phi }=(\phi _{1},\phi _{2},\cdots ,\phi _{k})^{\prime }$ and $\boldsymbol{\delta }=(\delta _{1},\delta _{2},\cdots ,\delta _{l_{T}})^{\prime }$ are the associated coefficients. The LS estimator of $\boldsymbol{\ \gamma }_{T}=(\boldsymbol{\phi }^{\prime },\boldsymbol{\delta }_{T}^{\prime })^{\prime }$ is \begin{equation} \hat{\boldsymbol{\gamma }}_{T}=\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W} \right) ^{-1}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{y}\right), \end{equation} where $\mathbf{W}=(\mathbf{w}_{1},\mathbf{w}_{2},\cdots ,\mathbf{w} _{T})^{\prime }$, $\mathbf{w}_{t}=\left( \mathbf{x}_{kt}^{\prime },\mathbf{s} _{t}^{\prime }\right) ^{\prime }$ and $\mathbf{y}=(y_{1},y_{2},\cdots ,y_{T})^{\prime }$. The model error is \begin{equation} \boldsymbol{\hat{\eta}}=\mathbf{y}-\mathbf{W}\boldsymbol{\hat{\gamma}}_{T}. \end{equation} Suppose that $\lambda _{\min }\left[ T^{-1}\mathbb{E}(\mathbf{W^{\prime }} \mathbf{W})\right] >c>0$, and $l_{T}=\ominus (T^{d})$, where $0\leq d<\frac{1 }{2}$. Moreover suppose that Assumptions (ref)-(ref) holds. Now, \begin{enumerate} • If $\mathbb{E}(\beta _{it})=\beta _{i}$ for all $t$, then \begin{equation*} \left\Vert \hat{\boldsymbol{\gamma }}_{T}-\boldsymbol{\gamma }_{T}^{\ast }\right\Vert =O_{p}\left( T^{-\frac{1- d}{2}}\right) , \end{equation*} where $\boldsymbol{\gamma }_{T}^{\ast }=(\boldsymbol{\beta }^{\prime }, \mathbf{\ 0}_{l_{T}}^{\prime })^{\prime }$ and $\boldsymbol{\beta }=(\beta _{1},\beta _{2},\cdots ,\beta _{k})^{\prime }$. Under Assumption (ref) we also have \begin{equation*} T^{-1}\boldsymbol{\hat{\eta}}^{\prime }\boldsymbol{\hat{\eta}}=\bar{\sigma} _{u,T}^{2}+\bar{\Delta}_{\beta,T}+O_{p}\left( \frac{1}{\sqrt{T}}\right) +O_{p}\left( T^{-(1-d)}\right) , \end{equation*} where $\bar{\sigma}_{u,T}^{2}=T^{-1}\sum_{t=1}^{T}\mathbb{E} \left(u_t^2\right) $, and $\bar{\Delta}_{\beta,T} = T^{-1} \sum_{t=1}^{T} \text{tr} \left( \boldsymbol{\Sigma}_{\mathbf{x}_k,t} \boldsymbol{\Omega} _{\beta,t} \right)$ are non-negative, with $\boldsymbol{\Sigma}_{\mathbf{x} _k,t} \equiv \left(\sigma_{ijt,x}\right)$, $\boldsymbol{\Omega}_{\beta,t} \equiv \left(\sigma_{ijt,\beta}\right)$ for $i,j = 1,2, \cdots, k$, and $ \sigma_{ijt,x} = \mathbb{E}\left( x_{it}x_{jt}\right)$, $\sigma_{ijt,\beta }= \mathbb{E}\left[(\beta_{it} - \beta_{i})(\beta_{jt} - \beta_{j})\right]$. • If $\mathbb{E}\left( \mathbf{w}_{t}\mathbf{w}_{t}^{\prime }\right) $ is time-invariant, then \begin{equation*} \left\Vert \hat{\boldsymbol{\gamma }}_{T}-\boldsymbol{\gamma }_{T}^{\diamond }\right\Vert =O_{p}\left( T^{-\frac{1-d}{2}}\right) , \end{equation*} where $\boldsymbol{\gamma }_{T}^{\diamond }=(\boldsymbol{\bar{\beta}} _{T}^{\prime },\mathbf{0}_{l_{T}}^{\prime })^{\prime }$, $\boldsymbol{\bar{ \beta}}_{T}=(\bar{\beta}_{1T},\bar{\beta}_{2T},\cdots ,\bar{\beta} _{kT})^{\prime }$, and $\bar{\beta}_{iT}=T^{-1}\sum_{t=1}^{T}\mathbb{E} (\beta _{it})$. If Assumption (ref) also holds, then \begin{equation*} T^{-1}\boldsymbol{\hat{\eta}}^{\prime }\boldsymbol{\hat{\eta}}= \bar{\sigma} _{u,T}^{2}+\bar{\Delta}_{\beta,T}^{\ast}+ O_{p}\left( \frac{1}{\sqrt{T}} \right) +O_{p}\left( T^{-(1-d)} \right) , \end{equation*} where $\bar{\Delta}_{\beta,T}^{\ast} = T^{-1} \sum_{t=1}^{T} \text{tr} \left( \boldsymbol{\Sigma}_{\mathbf{x}_k,t} \boldsymbol{\Omega} _{\beta,t}^{\ast} \right)$ is non-negative, with $\boldsymbol{\Omega} _{\beta,t}^{\ast} \equiv \left(\sigma_{ijt,\beta}^{\ast}\right)$ for $i,j = 1,2, \cdots, k$, and $\sigma_{ijt,\beta }^{\ast}=\mathbb{E}\left[ (\beta _{it}-\bar{\beta}_{i,T})(\beta _{jt}-\bar{\beta}_{j,T})\right]$. \end{enumerate}
proofIn the first scenario, where $\mathbb{E}(\beta _{it})=\beta _{i}$ for all $t$ , we can write ((ref)) as \begin{equation*} y_{t}=\sum_{i=1}^{k}x_{it}\beta _{i}+\sum_{i=1}^{k}x_{it}\left( \beta _{it}-\beta _{i}\right) +u_{t}=\sum_{i=1}^{k}x_{it}\beta _{i}+\sum_{i=1}^{k}r_{it}+u_{t}=\mathbf{x}_{kt}^{\prime }\boldsymbol{\beta }+ \mathbf{r}_{t}^{\prime }\boldsymbol{\tau }+u_{t}, \end{equation*} where $r_{it}=x_{it}\left( \beta _{it}-\beta _{i}\right) $, $\mathbf{r} _{t}=(r_{1t},r_{2t},\cdots ,r_{kt})^{\prime }$, and $\boldsymbol{\tau }$ is a $k\times 1$ vector of ones. We can further write the DGP in a following matrix format, \begin{equation} \mathbf{y}=\mathbf{X}_{k}\boldsymbol{\beta }+\mathbf{R}\boldsymbol{\tau}+ \mathbf{u}, \end{equation} where $\mathbf{X}_{k}=(\mathbf{x}_{k1},\mathbf{x}_{k2},\cdots ,\mathbf{x} _{kT})^{\prime }$, $\mathbf{R}=(\mathbf{r}_{1},\mathbf{r}_{2},\cdots , \mathbf{\ \ r}_{T})^{\prime }$ and $\mathbf{u}=(u_{1},u_{2},\cdots ,u_{T})^{\prime }$. By substituting ((ref)) into ((ref)), we obtain \begin{equation*} \hat{\boldsymbol{\gamma }}_{T}=\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W} \right) ^{-1}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{X}_{k}\boldsymbol{ \beta }\right) +\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) ^{-1}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{R}\boldsymbol{\tau }\right) +\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) ^{-1}\left( T^{-1} \mathbf{W}^{\prime }\mathbf{u}\right) , \end{equation*} where $\mathbf{W=(X}_{k}\mathbf{,S)}$, and $\mathbf{S}=(\mathbf{s}_{1}, \mathbf{s}_{2},\cdots ,\mathbf{s}_{T})^{\prime }$. Since $\boldsymbol{\gamma }_{T}^{\ast }=(\boldsymbol{\beta }^{\prime },\mathbf{0}_{l_{T}}^{\prime })^{\prime }$, $\mathbf{X}_{k}\boldsymbol{\beta }=\mathbf{X}_{k}\boldsymbol{ \beta }+\mathbf{S}\mathbf{0}_{l_{T}}=\mathbf{W}\boldsymbol{\gamma } _{T}^{\ast }$, which in turn allows us to write the above result as: \begin{equation*} \hat{\boldsymbol{\gamma }}_{T}=\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W} \right) ^{-1}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) \boldsymbol{ \gamma }_{T}^{\ast }+\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) ^{-1}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{R}\boldsymbol{\tau }\right) +\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) ^{-1}\left( T^{-1} \mathbf{W}^{\prime }\mathbf{u}\right) , \end{equation*} and hence \begin{equation} \hat{\boldsymbol{\gamma }}_{T}-\boldsymbol{\gamma }_{T}^{\ast }=\left( T^{-1} \mathbf{W}^{\prime }\mathbf{W}\right) ^{-1}\left( T^{-1}\mathbf{W}^{\prime } \mathbf{R}\boldsymbol{\tau }\right) +\left( T^{-1}\mathbf{W}^{\prime } \mathbf{W}\right) ^{-1}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{u}\right) . \end{equation} We can further write \begin{align*} \hat{\boldsymbol{\gamma }}_{T}-\boldsymbol{\gamma }_{T}^{\ast }=& \left\{ \left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) ^{-1}-\left[ \mathbb{E} \left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) \right] ^{-1}\right\} \left( T^{-1} \mathbf{W}^{\prime }\mathbf{R}\boldsymbol{\tau } \right) + \\ & \left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) \right] ^{-1}\left(T^{-1} \mathbf{W}^{\prime }\mathbf{R}\boldsymbol{\ \tau }\right) + \\ & \left\{ \left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) ^{-1}-\left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) \right] ^{-1}\right\} \left\{ T^{-1}\left[ \left( \mathbf{W}^{\prime }\mathbf{u} \right) -\mathbb{E}\left( \mathbf{W}^{\prime }\mathbf{u}\right) \right] \right\} + \\ & \left\{ \left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) ^{-1}-\left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) \right] ^{-1}\right\} \left[ T^{-1}\mathbb{E}\left( \mathbf{W}^{\prime }\mathbf{u} \right) \right] + \\ & \left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) \right] ^{-1}\left\{ T^{-1}\left[ \left( \mathbf{W}^{\prime }\mathbf{u}\right) - \mathbb{E}\left( \mathbf{W}^{\prime }\mathbf{u}\right) \right] \right\} + \\ & \left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) \right] ^{-1}\left[ T^{-1} \mathbb{E}\left( \mathbf{W}^{\prime }\mathbf{u}\right) \right]. \\ \end{align*} Hence, by the sub-additive property of norms and Lemma (ref), we have \begin{align*} \left\Vert \hat{\boldsymbol{\gamma }}_{T}-\boldsymbol{\gamma }_{T}^{\ast }\right\Vert \leq & \left\Vert \left( T^{-1}\mathbf{W}^{\prime }\mathbf{W} \right) ^{-1}-\left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W} \right) \right] ^{-1}\right\Vert _{F}\left\Vert T^{-1} \mathbf{W}^{\prime } \mathbf{R}\boldsymbol{\tau } \right\Vert + \\ & \left\Vert \left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W} \right) \right] ^{-1}\right\Vert _{2}\left\Vert T^{-1}\mathbf{W}^{\prime } \mathbf{R}\boldsymbol{\tau } \right\Vert + \\ & \left\Vert \left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) ^{-1}-\left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) \right] ^{-1}\right\Vert _{F}\left\Vert T^{-1}\left[ \left( \mathbf{W}^{\prime } \mathbf{u}\right) -\mathbb{E}\left( \mathbf{W}^{\prime }\mathbf{u}\right) \right] \right\Vert _+ \\ & \left\Vert \left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) ^{-1}-\left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) \right] ^{-1}\right\Vert _{F}\left\Vert T^{-1}\mathbb{E}\left( \mathbf{W}^{\prime } \mathbf{u}\right) \right\Vert + \\ & \left\Vert \left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W} \right) \right] ^{-1}\right\Vert _{2}\left\Vert T^{-1}\left[ \left( \mathbf{ \ W}^{\prime }\mathbf{u}\right) -\mathbb{E}\left( \mathbf{W}^{\prime } \mathbf{\ u}\right) \right] \right\Vert + \\ & \left\Vert \left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W} \right) \right] ^{-1}\right\Vert_{2} \left\Vert T^{-1}\mathbb{E}\left( \mathbf{W}^{\prime }\mathbf{\ u}\right) \right\Vert \end{align*} By Assumption (ref) \begin{equation*} \left \Vert T^{-1}\mathbb{E}\left( \mathbf{W}^{\prime }\mathbf{u}\right) \right\Vert = \left \Vert T^{-1}\sum_{t=1}^{T}\mathbb{E}(\mathbf{w} _{t}u_{t}) \right \Vert = O\left(T^{-\frac{2\epsilon-d}{2} }\right), \end{equation*} where $\epsilon \geq 1/2 $. Assumptions (ref) and (ref) imply that $\mathbf{W}$ and $\mathbf{u}$ satisfy condition (i) and (ii) of Lemma (ref), and by Lemmas (ref) and (ref), \begin{equation*} \left\Vert T^{-1}\left[\mathbf{W}^{\prime }\mathbf{u} -\mathbb{E}\left( \mathbf{W}^{\prime }\mathbf{\ u}\right)\right] \right\Vert =O_{p}\left( T^{- \frac{1-d}{2}}\right). \end{equation*} Similarly, \begin{equation*} \left\Vert T^{-1}\left[ \left( \mathbf{W}^{\prime }\mathbf{W}\right) - \mathbb{E}\left( \mathbf{W}^{\prime }\mathbf{W}\right) \right] \right\Vert _{F}=O_{p}\left(T^{-(1/2-d)}\right) , \end{equation*} and since $l_{T}=\ominus (T^{d})$ with $0\leq d<1/2$, by Lemma (ref), \begin{equation*} \left\Vert \left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) ^{-1}-\left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) \right] ^{-1}\right\Vert _{F}=O_{p}\left(T^{-(1/2-d)}\right). \end{equation*} Now consider $\left\Vert T^{-1}\mathbf{W}^{\prime }\mathbf{R}\boldsymbol{\ \tau }\right\Vert $. Note that the row $j$ and column $i$ of $l_{T}\times p$ matrix $T^{-1}\mathbf{W}^{\prime }\mathbf{R}$ is equal to $ T^{-1}\sum_{t=1}^{T}w_{jt}r_{it}$. Hence the $j^{\text{th}}$ element of $ l_{T}\times 1$ vector $T^{-1}\mathbf{W}^{\prime }\mathbf{R}\boldsymbol{\tau } $ is equal $T^{-1}\sum_{i=1}^{k}\sum_{t=1}^{T}w_{jt}r_{it}$. In other words, $T^{-1}\mathbf{W}^{\prime }\mathbf{R}\boldsymbol{\tau }=T^{-1}\sum_{i=1}^{k} \sum_{t=1}^{T}\mathbf{w}_{t}r_{it}$. Therefore, (recalling that $ r_{it}=x_{it}\left( \beta _{it}-\beta _{i}\right) $) \begin{align*} \left\Vert T^{-1}\mathbf{W}^{\prime }\mathbf{R}\boldsymbol{\tau }\right\Vert ^{2}=& \left\Vert T^{-1}\sum_{i=1}^{k}\sum_{t=1}^{T}(\mathbf{w} _{t}r_{it})\right\Vert ^{2}\leq \sum_{i=1}^{k}\left\Vert T^{-1}\sum_{t=1}^{T} \mathbf{w}_{t}x_{it}\left( \beta _{it}-\beta _{i}\right) \right\Vert ^{2} \\ =& T^{-2}\sum_{i=1}^{k}\sum_{t=1}^{T}\sum_{t^{\prime }=1}^{T}\mathbf{w} _{t}^{\prime }\mathbf{w}_{t^{\prime }}x_{it}x_{it^{\prime }}\left( \beta _{it}-\beta _{i}\right) \left( \beta _{it^{\prime }}-\beta _{i}\right) \\ =& T^{-2}\sum_{i=1}^{k}\sum_{t=1}^{T}\sum_{t^{\prime }=1}^{T}\sum_{\ell =1}^{k+l_{T}}w_{\ell t}w_{\ell t^{\prime }}x_{it}x_{it^{\prime }}\left( \beta _{it}-\beta _{i}\right) \left( \beta _{it^{\prime }}-\beta _{i}\right) . \\ & \end{align*} Since, by Assumption (ref), $\beta _{it}$ for $i=1,2,\cdots ,k$ are distributed independently of $\mathbf{w}_{t}$ for $t=1,2,\cdots ,T$, we can further write, \begin{align*} & \mathbb{E}\left\Vert T^{-1}\mathbf{W}^{\prime }\mathbf{R}\boldsymbol{\tau } \right\Vert ^{2}\leq T^{-2}\sum_{i=1}^{k}\sum_{t=1}^{T}\sum_{t^{\prime }=1}^{T}\sum_{\ell =1}^{k+\ell _{T}}\mathbb{E}\left( w_{\ell t}w_{\ell t^{\prime }}x_{it}x_{it^{\prime }}\right) \mathbb{E}\left[ \left( \beta _{it}-\beta _{i}\right) \left( \beta _{it^{\prime }}-\beta _{i}\right) \right] \\ & \qquad \leq T^{-2}\sum_{i=1}^{k}\sum_{t=1}^{T}\sum_{t^{\prime }=1}^{T}\sum_{\ell =1}^{k+\ell _{T}}\left\vert \mathbb{E}\left( w_{\ell t}w_{\ell t^{\prime }}x_{it}x_{it^{\prime }}\right) \right\vert \times \left\vert \mathbb{E}\left[ \left( \beta _{it}-\beta _{i}\right) \left( \beta _{it^{\prime }}-\beta _{i}\right) \right] \right\vert \\ & \qquad \leq T^{-2}\left( k+\ell _{T}\right) sup_{i,\ell ,t,t^{\prime }}\left\vert \mathbb{E}\left( w_{\ell t}w_{\ell t^{\prime }}x_{it}x_{it^{\prime }}\right) \right\vert \sum_{i=1}^{k}\sum_{t=1}^{T}\sum_{t^{\prime }=1}^{T}\left\vert \mathbb{E} \left[ \left( \beta _{it}-\beta _{i}\right) \left( \beta _{it^{\prime }}-\beta _{i}\right) \right] \right\vert \end{align*} Since $\mathbf{W}$ satisfy condition (i) of Lemma (ref), we have $\sup_{i,\ell ,t,t^{\prime }}\left\vert \mathbb{E} (w_{\ell t}w_{\ell t^{\prime }}x_{it}x_{it^{\prime }})\right\vert <C<\infty $ . Also, note that for any $t^{\prime }<t$, \begin{equation*} \mathbb{E}\left[ \left( \beta _{it}-\beta _{i}\right) \left( \beta _{it^{\prime }}-\beta _{i}\right) \right] =\mathbb{E}\left[ \left( \beta _{it^{\prime }}-\beta _{i}\right) \mathbb{E}\left( \beta _{it}-\beta _{i}| \mathcal{F}_{t-1}\right) \right] , \end{equation*} and by Assumption (ref), $\mathbb{E}\left( \beta _{it}-\beta _{i}| \mathcal{F}_{t-1}\right) =0$. Therefore, \begin{align*} \sum_{t=1}^{T}\sum_{t^{\prime }=1}^{T}\left\vert \mathbb{E}\left[ \left( \beta _{it}-\beta _{i}\right) \left( \beta _{it^{\prime }}-\beta _{i}\right) \right] \right\vert & =\sum_{t=1}^{T}\left\vert \mathbb{E}\left[ \left( \beta _{it}-\beta _{i}\right) ^{2}\right] \right\vert +2\sum_{t=2}^{T}\sum_{t^{\prime }=1}^{t}\left\vert \mathbb{E}\left[ \left( \beta _{it}-\beta _{i}\right) \left( \beta _{it^{\prime }}-\beta _{i}\right) \right] \right\vert \\ & =\sum_{t=1}^{T}\left\vert \mathbb{E}\left[ \left( \beta _{it}-\beta _{i}\right) ^{2}\right] \right\vert =O(T). \end{align*} Since, by Assumption (ref), $k$ is also a finite fixed integer, we conclude that \begin{equation*} \mathbb{E}\left\Vert T^{-1}\mathbf{W}^{\prime }\mathbf{R}\boldsymbol{\tau } \right\Vert ^{2}=O\left( T^{-(1-d)}\right) , \end{equation*} and hence, by Lemma (ref), \begin{equation*} \left\Vert T^{-1}\mathbf{W}^{\prime }\mathbf{R}\boldsymbol{\tau }\right\Vert =O_{p}\left( T^{-\frac{1-d}{2}}\right) . \end{equation*} So, we can conclude that \begin{equation*} \left\Vert \hat{\boldsymbol{\gamma }}_{T}-\boldsymbol{\gamma }_{T}^{\ast }\right\Vert =O_{p}\left( T^{-\frac{1-d}{2}}\right) , \end{equation*} as required. In the next step, consider the mean squared errors of the model, $T^{-1} \boldsymbol{\hat{\eta}}_{T}^{\prime }\boldsymbol{\hat{\eta}}_{T}$. By substituting $y$ from ((ref)) into equation ((ref)) for the model error, we have \begin{equation*} \boldsymbol{\hat{\eta}}=\mathbf{y}-\mathbf{W}\boldsymbol{\hat{\gamma}}_{T}= \mathbf{X}_{k}\boldsymbol{\beta }+\mathbf{R}\boldsymbol{\tau }+\mathbf{u}- \mathbf{W}\boldsymbol{\hat{\gamma}}_{T}. \end{equation*} Since $\mathbf{X}_{k}\boldsymbol{\beta }=\mathbf{W}\boldsymbol{\gamma } _{T}^{\ast }$, where $\boldsymbol{\gamma }_{T}^{\ast }=(\boldsymbol{\beta } ^{\prime },\mathbf{0}_{l_{T}}^{\prime })^{\prime }$, we can further write, \begin{equation*} \boldsymbol{\hat{\eta}}=\mathbf{R}\boldsymbol{\tau }+\mathbf{u}-\mathbf{\ W} \left( \boldsymbol{\hat{\gamma}}_{T}-\boldsymbol{\gamma }_{T}^{\ast }\right) . \end{equation*} Therefore, \begin{align*} T^{-1}\boldsymbol{\hat{\eta}}^{\prime }\boldsymbol{\hat{\eta}}& =T^{-1}\left[ \mathbf{R}\boldsymbol{\tau }+\mathbf{u}-\mathbf{W}\left( \boldsymbol{\ \ \hat{\gamma}}_{T}-\boldsymbol{\gamma }_{T}^{\ast }\right) \right] ^{\prime } \left[ \mathbf{R}\boldsymbol{\tau }+\mathbf{u}-\mathbf{W}\left( \boldsymbol{ \hat{\gamma}}_{T}-\boldsymbol{\gamma }_{T}^{\ast }\right) \right] \\ & =T^{-1}\left( \mathbf{R}\boldsymbol{\tau }+\mathbf{u}\right) ^{\prime }\left( \mathbf{R}\boldsymbol{\tau }+\mathbf{u}\right) +T^{-1}\left[ \mathbf{ W}\left( \boldsymbol{\hat{\gamma}}_{T}-\boldsymbol{\gamma }_{T}^{\ast }\right) \right] ^{\prime }\left[ \mathbf{W}\left( \boldsymbol{\ \hat{\gamma} }_{T}-\boldsymbol{\gamma }_{T}^{\ast }\right) \right] - \\ & \qquad 2T^{-1}\left[ \mathbf{W}\left( \boldsymbol{\hat{\gamma}}_{T}- \boldsymbol{\gamma }_{T}^{\ast }\right) \right] ^{\prime }\left( \mathbf{R} \boldsymbol{\tau }+\mathbf{u}\right) \\ & =T^{-1}\left( \boldsymbol{\tau }^{\prime }\mathbf{R}^{\prime }\mathbf{R} \boldsymbol{\tau }+\mathbf{u}^{\prime }\mathbf{u}\right) +2T^{-1}\boldsymbol{ \tau }^{\prime }\mathbf{R}^{\prime }\mathbf{u}+\left( \boldsymbol{\hat{\gamma }}_{T}-\boldsymbol{\gamma }_{T}^{\ast }\right) ^{\prime }\left( T^{-1} \mathbf{W}^{\prime }\mathbf{W}\right) \left( \boldsymbol{\hat{\gamma}}_{T}- \boldsymbol{\gamma }_{T}^{\ast }\right) - \\ & \qquad 2\left( \boldsymbol{\hat{\gamma}}_{T}-\boldsymbol{\gamma } _{T}^{\ast }\right) ^{\prime }\left[ T^{-1}\left( \mathbf{W}^{\prime } \mathbf{\ R}\boldsymbol{\tau }+\mathbf{W}^{\prime }\mathbf{u}\right) \right] . \end{align*} By substituting for $\boldsymbol{\hat{\gamma}}_{T}-\boldsymbol{\gamma } _{T}^{\ast }$ from ((ref)), we get \begin{align*} T^{-1}\boldsymbol{\hat{\eta}}^{\prime }\boldsymbol{\hat{\eta}}& =T^{-1}\left( \boldsymbol{\tau }^{\prime }\mathbf{R}^{\prime }\mathbf{R} \boldsymbol{\ \tau }+\mathbf{u}^{\prime }\mathbf{u}\right) +2T^{-1} \boldsymbol{\tau }^{\prime }\mathbf{R}^{\prime }\mathbf{u}+ \\ & \qquad \left[ T^{-1}\left( \mathbf{W}^{\prime }\mathbf{R}\boldsymbol{\tau } +\mathbf{W}^{\prime }\mathbf{u}\right) \right] ^{\prime }\left( T^{-1} \mathbf{W}^{\prime }\mathbf{W}\right) ^{-1}\left[ T^{-1}\left( \mathbf{W} ^{\prime }\mathbf{R}\boldsymbol{\tau }+\mathbf{W}^{\prime }\mathbf{u}\right) \right] - \\ & \qquad 2\left[ T^{-1}\left( \mathbf{W}^{\prime }\mathbf{R}\boldsymbol{\ \tau }+\mathbf{W}^{\prime }\mathbf{u}\right) \right] ^{\prime }\left( T^{-1} \mathbf{W}^{\prime }\mathbf{W}\right) ^{-1}\left[ T^{-1}\left( \mathbf{W} ^{\prime }\mathbf{R}\boldsymbol{\tau }+\mathbf{W}^{\prime }\mathbf{u}\right) \right] \\ & =T^{-1}\left( \boldsymbol{\tau }^{\prime }\mathbf{R}^{\prime }\mathbf{R} \boldsymbol{\tau }+\mathbf{u}^{\prime }\mathbf{u}\right) +2T^{-1}\boldsymbol{ \tau }^{\prime }\mathbf{R}^{\prime }\mathbf{u}- \\ & \qquad \left[ T^{-1}\left( \mathbf{W}^{\prime }\mathbf{R}\boldsymbol{\tau } +\mathbf{W}^{\prime }\mathbf{u}\right) \right] ^{\prime }\left( T^{-1} \mathbf{W}^{\prime }\mathbf{W}\right) ^{-1}\left[ T^{-1}\left( \mathbf{W} ^{\prime }\mathbf{R}\boldsymbol{\tau }+\mathbf{W}^{\prime }\mathbf{u} \right) \right] . \end{align*} we can further write \begin{align*} & T^{-1}\boldsymbol{\hat{\eta}}^{\prime }\boldsymbol{\hat{\eta}}=T^{-1} \mathbb{E}\left( \boldsymbol{\tau }^{\prime }\mathbf{R}^{\prime }\mathbf{R} \boldsymbol{\tau }+\mathbf{u}^{\prime }\mathbf{u}\right) +T^{-1}\left\{ \left[ \boldsymbol{\tau }^{\prime }\mathbf{R}^{\prime }\mathbf{R}\boldsymbol{ \ \tau }-\mathbb{E}\left( \boldsymbol{\tau }^{\prime }\mathbf{R}^{\prime } \mathbf{R}\boldsymbol{\tau }\right) \right] +\left[ \mathbf{u}^{\prime } \mathbf{u}-\mathbb{E}\left( \mathbf{u}^{\prime }\mathbf{u}\right) \right] \right\} + \\ & \quad 2T^{-1}\boldsymbol{\tau }^{\prime }\mathbf{R}^{\prime }\mathbf{u}- \left[ T^{-1}\left( \mathbf{W}^{\prime }\mathbf{R}\boldsymbol{\tau }+\mathbf{ W}^{\prime }\mathbf{u}\right) \right] ^{\prime }\left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) \right] ^{-1}\left[ T^{-1}\left( \mathbf{W}^{\prime }\mathbf{R}\boldsymbol{\tau }+\mathbf{W}^{\prime }\mathbf{ u}\right) \right] - \\ & \quad \left[ T^{-1}\left( \mathbf{W}^{\prime }\mathbf{R}\boldsymbol{\tau }+ \mathbf{W}^{\prime }\mathbf{u}\right) \right] ^{\prime }\left\{ \left( T^{-1} \mathbf{W}^{\prime }\mathbf{W}\right) ^{-1}-\left[ \mathbb{E}\left( T^{-1} \mathbf{W}^{\prime }\mathbf{W}\right) \right] ^{-1}\right\} \left[ T^{-1}\left( \mathbf{W}^{\prime }\mathbf{R}\boldsymbol{\tau }+\mathbf{W} ^{\prime }\mathbf{u}\right) \right] . \end{align*} Therefore, \begin{equation} \begin{split} & T^{-1}\boldsymbol{\hat{\eta}}^{\prime }\boldsymbol{\hat{\eta}}-T^{-1} \mathbb{E}\left( \boldsymbol{\tau }^{\prime }\mathbf{R}^{\prime }\mathbf{R} \boldsymbol{\tau }+\mathbf{u}^{\prime }\mathbf{u}\right) \leq \\ & \qquad T^{-1}\left[ \boldsymbol{\tau }^{\prime }\mathbf{R}^{\prime } \mathbf{\ \ R}\boldsymbol{\tau }-\mathbb{E}\left( \boldsymbol{\tau }^{\prime }\mathbf{R}^{\prime }\mathbf{R}\boldsymbol{\tau }\right) \right] +T^{-1} \left[ \mathbf{u}^{\prime }\mathbf{u}-\mathbb{E}\left( \mathbf{u}^{\prime } \mathbf{u}\right) \right] + 2T^{-1}\boldsymbol{\tau }^{\prime }\mathbf{R} ^{\prime }\mathbf{u} + \\ & \qquad \left\Vert T^{-1}\left[ \mathbf{W}^{\prime } \mathbf{R}\boldsymbol{ \tau }+ \mathbf{W}^{\prime } \mathbf{u} - \mathbb{E}\left(\mathbf{W} ^{\prime} \mathbf{u}\right) \right] \right\Vert^{2} \left\Vert \left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) \right] ^{-1}\right\Vert _{2} + \left\Vert T^{-1} \mathbb{E}\left(\mathbf{W}^{\prime } \mathbf{u}\right) \right\Vert ^{2} \left\Vert \left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) \right] ^{-1}\right\Vert _{2} + \\ & \qquad \left\Vert T^{-1} \left[ \mathbf{W}^{\prime } \mathbf{R}\boldsymbol{ \tau }+\mathbf{W}^{\prime }\mathbf{u} - \mathbb{E} \left(\mathbf{W}^{\prime } \mathbf{u} \right)\right] \right\Vert ^{2}\left\Vert \left( T^{-1}\mathbf{W }^{\prime }\mathbf{W}\right) ^{-1}-\left[ \mathbb{E}\left( T^{-1}\mathbf{W} ^{\prime }\mathbf{W}\right) \right] ^{-1}\right\Vert _{F} + \\ & \qquad \left\Vert T^{-1} \mathbb{E} \left(\mathbf{W}^{\prime } \mathbf{u} \right) \right\Vert ^{2}\left\Vert \left( T^{-1}\mathbf{W}^{\prime }\mathbf{W }\right) ^{-1}-\left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W} \right) \right] ^{-1}\right\Vert _{F}. \end{split} \end{equation} First, consider $T^{-1}\left[ \boldsymbol{\tau }^{\prime }\mathbf{R}^{\prime }\mathbf{R}\boldsymbol{\tau }-\mathbb{E}\left( \boldsymbol{\tau }^{\prime } \mathbf{R}^{\prime }\mathbf{R}\boldsymbol{\tau }\right) \right] $. Note that \begin{equation*} \boldsymbol{\tau }^{\prime }\mathbf{R}^{\prime }\mathbf{R}\boldsymbol{\tau }= \boldsymbol{\tau }^{\prime }\left( \sum_{t=1}^{T}\mathbf{r}_{t}\mathbf{r} _{t}^{\prime }\right) \boldsymbol{\tau }=\sum_{t=1}^{T}\left( \boldsymbol{\ \tau }^{\prime }\mathbf{r}_{t}\right) \left( \mathbf{r}_{t}^{\prime } \boldsymbol{\tau }\right) =\sum_{t=1}^{T}\left( \sum_{i=1}^{k}r_{it}\right) \left( \sum_{j=1}^{k}r_{jt}\right) =\sum_{i=1}^{k}\sum_{j=1}^{k}\sum_{t=1}^{T}r_{it}r_{jt}. \end{equation*} Recalling that $r_{it}=x_{it}(\beta _{it}-\beta _{i})$, and hence, \begin{equation*} T^{-1}\left[ \boldsymbol{\tau }^{\prime }\mathbf{R}^{\prime }\mathbf{R} \boldsymbol{\tau }-\mathbb{E}\left( \boldsymbol{\tau }^{\prime }\mathbf{R} ^{\prime }\mathbf{R}\boldsymbol{\tau }\right) \right] =\sum_{i=1}^{k} \sum_{j=1}^{k}\left( T^{-1}\sum_{t=1}^{T} \tilde{r}_{ij,t}\right) , \end{equation*} where \begin{equation*} \tilde{r}_{ij,t}= r_{it}r_{jt} - \mathbb{E}(r_{it} r_{jt}) \end{equation*} Now consider $\mathbb{E}\left( T^{-1}\sum_{t=1}^{T}\tilde{r}_{ij,t}\right) ^{2}$ and note that \begin{align*} \mathbb{E}\left( T^{-1}\sum_{t=1}^{T}\tilde{r}_{ij,t}\right) ^{2}& = T^{-2}\sum_{t=1}^{T}\sum_{t^{\prime }=1}^{T}\mathbb{E}\left( \tilde{r}_{ij,t} \tilde{r}_{ij,t^{\prime }}\right). \end{align*} By Assumption (ref), $T^{-2}\sum_{t=1}^{T}\sum_{t^{ \prime }=1}^{T}\mathbb{E}\left( \tilde{r}_{ij,t}\tilde{r}_{ij,t^{\prime }}\right)=O\left( T^{-1}\right) ,$ and hence, by Lemma (ref), it follows that \begin{equation*} \left\vert T^{-1}\sum_{t=1}^{T}\tilde{r}_{ij,t}\right\vert =O_{p}\left( \frac{1}{\sqrt{T}}\right) . \end{equation*} Since by Assumption (ref), $k$ is a finite fixed integer, we can further conclude that \begin{equation} T^{-1}\left[ \boldsymbol{\tau }^{\prime }\mathbf{R}^{\prime }\mathbf{R} \boldsymbol{\tau }-\mathbb{E}\left( \boldsymbol{\tau }^{\prime }\mathbf{R} ^{\prime }\mathbf{R}\boldsymbol{\tau }\right) \right] =\sum_{i=1}^{k} \sum_{j=1}^{k}\left( T^{-1}\sum_{t=1}^{T}\tilde{r}_{ij,t}\right) =O_{p}\left( \frac{1}{\sqrt{T}}\right) . \end{equation} Now, consider, $T^{-1}\boldsymbol{\tau }^{\prime }\mathbf{R}^{\prime } \mathbf{u}$. Note that \begin{equation*} T^{-1}\boldsymbol{\tau }^{\prime }\mathbf{R}^{\prime }\mathbf{u}=T^{-1} \boldsymbol{\tau }^{\prime }\left( \sum_{t=1}^{T}\mathbf{r}_{t}u_{t}\right) =T^{-1}\sum_{t=1}^{T}\boldsymbol{\tau }^{\prime }\mathbf{r} _{t}u_{t}=T^{-1}\sum_{t=1}^{T}\sum_{i=1}^{k}r_{it}u_{t}=\sum_{i=1}^{k}\left( T^{-1}\sum_{t=1}^{T}r_{it}u_{t}\right) . \end{equation*} We have \begin{equation*} \mathbb{E}\left( T^{-1}\sum_{t=1}^{T}r_{it}u_{t}\right) ^{2}=T^{-2}\sum_{t=1}^{T}\mathbb{E}\left( r_{it}^{2}u_{t}^{2}\right) +2T^{-2}\sum_{t=2}^{T}\sum_{t^{\prime }=1}^{t}\mathbb{E}\left( r_{it}r_{it^{\prime }}u_{t}u_{t^{\prime }}\right) . \end{equation*} Since $r_{it}=x_{it}(\beta _{it}-\beta _{i})$, and $\beta _{it}$ for $ i=1,2,\cdots ,k$ are distributed independently of $x_{js}$, $j=1,2,\cdots ,N$ , and $u_{s}$ for all $t$ and $s$, we can further write for any $t^{\prime }<t$ \begin{align*} \mathbb{E}\left( r_{it}r_{it^{\prime }}u_{t}u_{t^{\prime }}\right) & = \mathbb{E}\left( x_{it}u_{t}x_{it^{\prime }}u_{t^{\prime }}\right) \mathbb{E} \left[ (\beta _{it}-\beta _{i})(\beta _{it^{\prime }}-\beta _{i})\right] \\ & =\mathbb{E}\left( x_{it}u_{t}x_{it^{\prime }}u_{t^{\prime }}\right) \mathbb{E}\left\{ (\beta _{it^{\prime }}-\beta _{i})\mathbb{E}\left[ (\beta _{it}-\beta _{i})|\mathcal{F}_{t-1}\right] \right\} . \end{align*} But, by Assumption (ref), $\mathbb{E}\left[ (\beta _{it}-\beta _{i})| \mathcal{F}_{t-1}\right] =0$ and thus $\mathbb{E}\left( r_{it}r_{it^{\prime }}u_{t}u_{t^{\prime }}\right) =0$ for any $t^{\prime }<t$. Therefore, \begin{equation*} \mathbb{E}\left( T^{-1}\sum_{t=1}^{T}r_{it}u_{t}\right) ^{2}=T^{-2}\sum_{t=1}^{T}\mathbb{E}\left( r_{it}^{2}u_{t}^{2}\right) =O\left( \frac{1}{T}\right) . \end{equation*} Hence, by Lemma (ref), $\left\vert T^{-1}\sum_{t=1}^{T}r_{it}u_{t}\right\vert =O_{p}\left( \frac{1}{\sqrt{T}} \right) $. Since, by Assumption (ref), $k$ is a finite fixed integer, we conclude that \begin{equation} T^{-1}\boldsymbol{\tau }^{\prime }\mathbf{R}^{\prime }\mathbf{u} =\sum_{i=1}^{k}\left( T^{-1}\sum_{t=1}^{T}r_{it}u_{t}\right) =O_{p}\left( \frac{1}{\sqrt{T}}\right) . \end{equation} By substituting ((ref)) and ((ref)) into ((ref)), and noting that $\left \Vert T^{-1}\mathbb{E}\left( \mathbf{W}^{\prime }\mathbf{u}\right) \right\Vert^2 = O\left(T^{-(2\epsilon-d) }\right)$, for some $\epsilon \geq 1/2 $, \begin{equation*} \left\Vert T^{-1}\left[ \mathbf{W}^{\prime }\mathbf{R}\boldsymbol{\tau }+ \mathbf{W}^{\prime }\mathbf{u} - \mathbb{E}\left(\mathbf{W}^{\prime }\mathbf{ u} \right) \right] \right\Vert ^{2} = O_{p}(T^{-(1-d)}), \end{equation*} \begin{equation*} \left\Vert \left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) ^{-1}-\left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) \right] ^{-1}\right\Vert _{F}=O_{p}(T^{-(1/2-d)}), \end{equation*} and \begin{equation*} T^{-1}\left[ \mathbf{u}^{\prime }\mathbf{u}-\mathbb{E}\left(\mathbf{\ u} ^{\prime }\mathbf{u}\right) \right] =O_{p}(1/\sqrt{T}), \end{equation*} we conclude that \begin{equation*} T^{-1}\boldsymbol{\hat{\eta}}^{\prime }\boldsymbol{\hat{\eta}} =\sum_{i=1}^{k}\sum_{j=1}^{k}\left( T^{-1}\sum_{t=1}^{T}\sigma _{ijt,x}\sigma _{ijt,\beta }\right) +\bar{\sigma}_{u,T}^{2}+O_{p}\left( \frac{1}{\sqrt{T}}\right) +O_{p}\left( T^{-(1-d)}\right) , \end{equation*} where $\sigma _{ijt,x}=\mathbb{E}\left( x_{it}x_{jt}\right) $, $\sigma _{ijt,\beta }=\mathbb{E}\left[ (\beta _{it}-\beta _{i})(\beta _{jt}-\beta _{j})\right] $, and $\bar{\sigma}_{u,T}^{2}=T^{-1}\mathbb{E}\left( \mathbf{u} ^{\prime }\mathbf{u}\right) $. We further have \begin{equation*} \bar{\Delta}_{\beta,T} = \sum_{i=1}^{k}\sum_{j=1}^{k}\left( T^{-1}\sum_{t=1}^{T}\sigma_{ijt,x}\sigma _{ijt,\beta }\right) = T^{-1}\sum_{t=1}^{T} \left( \sum_{i=1}^{k}\sum_{j=1}^{k}\sigma_{ijt,x}\sigma _{ijt,\beta }\right) = \frac{1}{T} \sum_{t=1}^{T} tr\left(\boldsymbol{ \Omega}_{\beta,t} \boldsymbol{\Sigma}_{\mathbf{x}_k,t}\right), \end{equation*} where $\boldsymbol{\Omega}_{\beta,t} \equiv \left(\sigma_{ijt,\beta}\right)$ and $\boldsymbol{\Sigma}_{\mathbf{x}_k,t} \equiv \left(\sigma_{ijt,x}\right)$ for $i,j = 1,2, \cdots, k$. By result 9(b) on page 44 of Lutkepohl1996handbook, we can further write \begin{equation*} tr\left(\boldsymbol{\Omega}_{\beta,t} \boldsymbol{\Sigma}_{\mathbf{x} _k,t}\right) \geq k \left[det\left(\boldsymbol{\Omega}_{\beta,t} \right) \right]^{1/k}\left[det\left(\boldsymbol{\Sigma}_{\mathbf{x} _k,t}\right)\right]^{1/k}. \end{equation*} But $k $ is a finite fixed integer. Furthermore, $\text{det}\left( \boldsymbol{\Omega}_{\beta,t} \right) \geq 0 $ and $\text{det}\left( \boldsymbol{\Sigma}_{\mathbf{x}_k,t}\right) > 0 $, since $\boldsymbol{\Omega} _{\beta,t}$ and $\boldsymbol{\Sigma}_{\mathbf{x}_k,t} $ are positive semi-definite and positive definite matrices, respectively. So, we can conclude that $\bar{\Delta}_{\beta,T} \geq 0 $ as required. In the second scenario, where $\mathbb{E}\left( \mathbf{w}_{t}\mathbf{w} _{t}^{\prime }\right) $ is time-invariant, we can write ((ref)) as \begin{equation*} y_{t}=\sum_{i=1}^{k}x_{it}\bar{\beta}_{iT}+\sum_{i=1}^{k}x_{it}\left( \beta _{it}-\bar{\beta}_{iT}\right) +u_{t}=\sum_{i=1}^{k}x_{it}\bar{\beta} _{iT}+\sum_{i=1}^{k}h_{it}+u_{t}=\mathbf{x}_{kt}^{\prime }\bar{\boldsymbol{ \beta }}+\mathbf{h}_{t}^{\prime }\boldsymbol{\tau }+u_{t}, \end{equation*} where $h_{it}=x_{it}\left( \beta _{it}-\bar{\beta}_{iT}\right) $, and $ \mathbf{h}_{t}=(h_{1t},h_{2t},\cdots ,h_{kt})^{\prime }$. We can further write the DGP in ((ref)) in matrix format as \begin{equation*} \mathbf{y}=\mathbf{X}_{k}\bar{\boldsymbol{\beta }}+\mathbf{H}\boldsymbol{ \tau }+\mathbf{u}, \end{equation*} where $\mathbf{H}=(\mathbf{h}_{1},\mathbf{h}_{2},\cdots ,\mathbf{h} _{T})^{\prime }$. Now, by using the similar lines of arguments as in the first scenario, we obtain \begin{equation*} \hat{\boldsymbol{\gamma }}_{T}-\boldsymbol{\gamma }_{T}^{\diamond }=\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) ^{-1}\left( T^{-1}\mathbf{W} ^{\prime }\mathbf{H}\boldsymbol{\tau }\right) +\left( T^{-1}\mathbf{W} ^{\prime }\mathbf{W}\right) ^{-1}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{u} \right) . \end{equation*} We can further use the similar lines of arguments as in the first scenario and write \begin{align*} \left\Vert \hat{\boldsymbol{\gamma }}_{T}-\boldsymbol{\gamma }_{T}^{\diamond }\right\Vert \leq & \left\Vert \left( T^{-1}\mathbf{W}^{\prime }\mathbf{W} \right) ^{-1}-\left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W} \right) \right] ^{-1}\right\Vert _{F}\left\Vert T^{-1}\mathbf{W}^{\prime } \mathbf{H}\boldsymbol{\tau }\right\Vert + \\ & \left\Vert \left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W} \right) \right] ^{-1}\right\Vert _{2}\left\Vert T^{-1}\mathbf{W}^{\prime } \mathbf{H}\boldsymbol{\tau }\right\Vert + \\ & \left\Vert \left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) ^{-1}-\left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) \right] ^{-1}\right\Vert _{F}\left\Vert T^{-1}\left[ \left( \mathbf{W}^{\prime } \mathbf{u}\right) -\mathbb{E}\left( \mathbf{W}^{\prime }\mathbf{u}\right) \right] \right\Vert _+ \\ & \left\Vert \left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) ^{-1}-\left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) \right] ^{-1}\right\Vert _{F}\left\Vert T^{-1}\mathbb{E}\left( \mathbf{W}^{\prime } \mathbf{u}\right) \right\Vert + \\ & \left\Vert \left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W} \right) \right] ^{-1}\right\Vert _{2}\left\Vert T^{-1}\left[ \left( \mathbf{W }^{\prime }\mathbf{u}\right) -\mathbb{E}\left( \mathbf{W}^{\prime }\mathbf{\ u}\right) \right] \right\Vert + \\ & \left\Vert \left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W} \right) \right] ^{-1}\right\Vert_{2} \left\Vert T^{-1}\mathbb{E}\left( \mathbf{W}^{\prime }\mathbf{\ u}\right) \right\Vert \end{align*} We know that $\left\Vert T^{-1}\mathbb{E}\left( \mathbf{W}^{\prime }\mathbf{ \ u}\right) \right\Vert = O\left(T^{-\frac{2 \epsilon - d}{2}}\right) $ for some $\epsilon \geq 1/2 $. Also, \begin{equation*} \left\Vert T^{-1}\left[ \left( \mathbf{W}^{\prime }\mathbf{u}\right) - \mathbb{E}\left( \mathbf{W}^{\prime }\mathbf{u}\right) \right] \right\Vert =O_{p}\left( T^{-\frac{1 - d}{2}}\right) , \end{equation*} and \begin{equation*} \left\Vert \left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) ^{-1}-\left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) \right] ^{-1}\right\Vert _{F}=O_{p}\left( T^{-(1/2 - d)}\right) . \end{equation*} Now consider $\left\Vert T^{-1}\mathbf{W}^{\prime }\mathbf{H}\boldsymbol{ \tau }\right\Vert $. By using the similar lines of arguments as in the first scenario, we have \begin{equation*} \left\Vert T^{-1}\mathbf{W}^{\prime }\mathbf{H}\boldsymbol{\tau }\right\Vert ^{2}\leq T^{-2}\sum_{i=1}^{k}\sum_{\ell =1}^{k+l_{T}}\sum_{t=1}^{T}\sum_{t^{\prime }=1}^{T}w_{\ell t}w_{\ell t^{\prime }}x_{it}x_{it^{\prime }}\left( \beta _{it}-\bar{\beta}_{i}\right) \left( \beta _{it^{\prime }}-\bar{\beta}_{i}\right) . \end{equation*} Since, by Assumption (ref), $\beta _{it}$ for $i=1,2,\cdots ,k$ are distributed independently of $\mathbf{w}_{t}$ for $t=1,2,\cdots ,T$, we can further write, \begin{align*} \mathbb{E}\left\Vert T^{-1}\mathbf{W}^{\prime }\mathbf{H}\boldsymbol{\tau } \right\Vert ^{2}& \leq T^{-2}\sum_{i=1}^{k}\sum_{\ell =1}^{k+l_{T}}\sum_{t=1}^{T}\sum_{t^{\prime }=1}^{T}\mathbb{E}\left( w_{\ell t}w_{\ell t^{\prime }}x_{it}x_{it^{\prime }}\right) \mathbb{E}\left[ \left( \beta _{it}-\bar{\beta}_{i}\right) \left( \beta _{it^{\prime }}-\bar{\beta} _{i}\right) \right] \\ & =T^{-2}\sum_{i=1}^{k}\sum_{\ell =1}^{k+l_{T}}\sum_{t=1}^{T}\mathbb{E} \left( w_{\ell t}^{2}x_{it}^{2}\right) \mathbb{E}\left[ \left( \beta _{it}- \bar{\beta}_{i}\right) ^{2}\right] + \\ & \quad T^{-2}\sum_{i=1}^{k}\sum_{\ell =1}^{k+l_{T}}\sum_{t=1}^{T}\sum_{t^{\prime }\neq t}\mathbb{E}\left( w_{\ell t}w_{\ell t^{\prime }}x_{it}x_{it^{\prime }}\right) \mathbb{E}\left[ \left( \beta _{it}-\bar{\beta}_{i}\right) \left( \beta _{it^{\prime }}-\bar{\beta} _{i}\right) \right] . \end{align*} Since, by Assumption (ref), $\mathbb{E}\left[ w_{\ell t}w_{\ell ^{\prime }t}-\mathbb{E}(w_{\ell t}w_{\ell ^{\prime }t})|\mathcal{F}_{t-1}\right] =0$ for all $\ell $, $\ell ^{\prime }$ and $t=1,2,\cdots ,T$, we have for any $ t^{\prime }\neq t$ \begin{equation*} \mathbb{E}\left( w_{\ell t}w_{\ell t^{\prime }}x_{it}x_{it^{\prime }}\right) =\mathbb{E}\left( w_{\ell t}x_{it}\right) \mathbb{E}\left( w_{\ell t^{\prime }}x_{it^{\prime }}\right) . \end{equation*} Therefore, \begin{align*} & \sum_{t=1}^{T}\sum_{t^{\prime }\neq t}\mathbb{E}\left( w_{\ell t}w_{\ell t^{\prime }}x_{it}x_{it^{\prime }}\right) \mathbb{E}\left[ \left( \beta _{it}-\bar{\beta}_{i}\right) \left( \beta _{it^{\prime }}-\bar{\beta} _{i}\right) \right] \\ & \qquad =\sum_{t=1}^{T}\sum_{t^{\prime }\neq t}\mathbb{E}\left( w_{\ell t}x_{it}\right) \mathbb{E}\left( w_{\ell t^{\prime }}x_{it^{\prime }}\right) \mathbb{E}\left[ \left( \beta _{it}-\bar{\beta}_{i}\right) \left( \beta _{it^{\prime }}-\bar{\beta}_{i}\right) \right] . \end{align*} Since $\mathbb{E}\left( \mathbf{w}_{t}\mathbf{w}_{t}^{\prime }\right) $ is time-invariant, we can further write \begin{align*} & \sum_{t=1}^{T}\sum_{t^{\prime }\neq t}\mathbb{E}\left( w_{\ell t}w_{\ell t^{\prime }}x_{it}x_{it^{\prime }}\right) \mathbb{E}\left[ \left( \beta _{it}-\bar{\beta}_{i}\right) \left( \beta _{it^{\prime }}-\bar{\beta} _{i}\right) \right] \\ & \qquad =\mathbb{E}\left( w_{\ell t}x_{it}\right) ^{2}\sum_{t=1}^{T}\sum_{t^{\prime }\neq t}\mathbb{E}\left[ \left( \beta _{it}-\bar{\beta}_{i}\right) \left( \beta _{it^{\prime }}-\bar{\beta} _{i}\right) \right] . \end{align*} Note that, by Assumption (ref), for any $t^{\prime }\neq t$, $\mathbb{E} \left[ \left( \beta _{it}-\bar{\beta}_{i}\right) \left( \beta _{it^{\prime }}-\bar{\beta}_{i}\right) \right] =\left[ \mathbb{E}\left( \beta _{it}\right) -\bar{\beta}_{i}\right] \left[ \mathbb{E}\left( \beta _{it^{\prime }}\right) -\bar{\beta}_{i}\right] $. Therefore \begin{align*} & \sum_{t=1}^{T}\sum_{t^{\prime }\neq t}\mathbb{E}\left( w_{\ell t}w_{\ell t^{\prime }}x_{it}x_{it^{\prime }}\right) \mathbb{E}\left[ \left( \beta _{it}-\bar{\beta}_{i}\right) \left( \beta _{it^{\prime }}-\bar{\beta} _{i}\right) \right] \\ & \qquad =\left[ \mathbb{E}\left( w_{\ell t}x_{it}\right) \right] ^{2}\sum_{t=1}^{T}\sum_{t^{\prime }\neq t}\left[ \mathbb{E}\left( \beta _{it}\right) -\bar{\beta}_{i}\right] \left[ \mathbb{E}\left( \beta _{it^{\prime }}\right) -\bar{\beta}_{i}\right] . \end{align*} We can further write, \begin{align*} & \sum_{t=1}^{T}\sum_{t^{\prime }\neq t}\mathbb{E}\left( w_{\ell t}w_{\ell t^{\prime }}x_{it}x_{it^{\prime }}\right) \mathbb{E}\left[ \left( \beta _{it}-\bar{\beta}_{i}\right) \left( \beta _{it^{\prime }}-\bar{\beta} _{i}\right) \right] \\ & \qquad =\left[ \mathbb{E}\left( w_{\ell t}x_{it}\right) \right] ^{2}\left\{ \sum_{t=1}^{T}\sum_{t^{\prime }=1}^{T}\left[ \mathbb{E}\left( \beta _{it}\right) -\bar{\beta}_{i}\right] \left[ \mathbb{E}\left( \beta _{it^{\prime }}\right) -\bar{\beta}_{i}\right] -\sum_{t=1}^{T}\left[ \mathbb{ E}\left( \beta _{it}\right) -\bar{\beta}_{i}\right] ^{2}\right\} \\ & \qquad =\left[ \mathbb{E}\left( w_{\ell t}x_{it}\right) \right] ^{2}\left\{ \sum_{t=1}^{T}\left[ \mathbb{E}\left( \beta _{it}\right) -\bar{ \beta}_{i}\right] \right\} \left\{ \sum_{t^{\prime }=1}^{T}\left[ \mathbb{E} \left( \beta _{it^{\prime }}\right) -\bar{\beta}_{i}\right] \right\} - \\ & \qquad \quad \left[ \mathbb{E}\left( w_{\ell t}x_{it}\right) \right] ^{2}\sum_{t=1}^{T}\left[ \mathbb{E}\left( \beta _{it}\right) -\bar{\beta}_{i} \right] ^{2}. \end{align*} But, $\sum_{t=1}^{T}\left[ \mathbb{E}\left( \beta _{it}\right) -\bar{\beta} _{i}\right] =0$, and therefore, \begin{equation*} \sum_{t=1}^{T}\sum_{t^{\prime }\neq t}\mathbb{E}\left( w_{\ell t}w_{\ell t^{\prime }}x_{it}x_{it^{\prime }}\right) \mathbb{E}\left[ \left( \beta _{it}-\bar{\beta}_{i}\right) \left( \beta _{it^{\prime }}-\bar{\beta} _{i}\right) \right] =-\left[ \mathbb{E}\left( w_{\ell t}x_{it}\right) \right] ^{2}\sum_{t=1}^{T}\left[ \mathbb{E}\left( \beta _{it}\right) -\bar{\beta}_{i} \right] ^{2}. \end{equation*} So, \begin{align*} & \mathbb{E}\left\Vert T^{-1}\mathbf{W}^{\prime }\mathbf{H}\boldsymbol{\tau } \right\Vert ^{2} \\ & \qquad \leq T^{-2}\sum_{i=1}^{p}\sum_{\ell =1}^{p+l_{T}}\sum_{t=1}^{T}\left\{ \mathbb{E}\left( w_{\ell t}^{2}x_{it}^{2}\right) \mathbb{E}\left[ \left( \beta _{it}-\bar{\beta} _{i}\right) ^{2}\right] -\left[ \mathbb{E}\left( w_{\ell t}x_{it}\right) \right] ^{2}\left[ \mathbb{E}\left( \beta _{it}\right) -\bar{\beta}_{i} \right] ^{2}\right\} \\ & \qquad =O\left( T^{-(1 - d)}\right) , \end{align*} and hence, by Lemma (ref), \begin{equation*} \left\Vert T^{-1}\mathbf{W}^{\prime }\mathbf{H}\boldsymbol{\tau }\right\Vert =O_{p}\left( T^{-\frac{1 - d}{2}}\right) . \end{equation*} So, we conclude that \begin{equation*} \left\Vert \hat{\boldsymbol{\gamma }}_{T}-\boldsymbol{\gamma }_{T}^{\diamond }\right\Vert =O_{p}\left( T^{-\frac{1 - d}{2}}\right) . \end{equation*} Lastly, consider the model mean squared errors for the second scenario. Following the same lines of argument as in the first scenario, we can write, \begin{equation} \begin{split} & T^{-1}\boldsymbol{\hat{\eta}}^{\prime }\boldsymbol{\hat{\eta}}-T^{-1} \mathbb{E}\left( \boldsymbol{\tau }^{\prime }\mathbf{H}^{\prime }\mathbf{H} \boldsymbol{\tau }+\mathbf{u}^{\prime }\mathbf{u}\right) \leq \\ & \qquad T^{-1}\left[ \boldsymbol{\tau }^{\prime }\mathbf{H}^{\prime } \mathbf{\ \ H}\boldsymbol{\tau }-\mathbb{E}\left( \boldsymbol{\tau }^{\prime }\mathbf{H}^{\prime }\mathbf{H}\boldsymbol{\tau }\right) \right] +T^{-1} \left[ \mathbf{u}^{\prime }\mathbf{u}-\mathbb{E}\left( \mathbf{u}^{\prime } \mathbf{u}\right) \right] + 2T^{-1}\boldsymbol{\tau }^{\prime }\mathbf{H} ^{\prime }\mathbf{u} + \\ & \qquad \left\Vert T^{-1} \left[\mathbf{W}^{\prime}\mathbf{H}\boldsymbol{ \tau }+ \mathbf{W}^{\prime} \mathbf{u} - \mathbb{E}\left(\mathbf{W}^{\prime} \mathbf{u}\right)\right] \right\Vert ^{2}\left\Vert \left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) \right] ^{-1}\right\Vert _{2}+ \left\Vert T^{-1} \mathbb{E}\left(\mathbf{W}^{\prime} \mathbf{u}\right) \right\Vert ^{2}\left\Vert \left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime } \mathbf{W}\right) \right] ^{-1}\right\Vert _{2} + \\ & \qquad \left\Vert T^{-1} \left[\mathbf{W}^{\prime}\mathbf{H}\boldsymbol{ \tau }+ \mathbf{W}^{\prime} \mathbf{u} - \mathbb{E}\left(\mathbf{W}^{\prime} \mathbf{u}\right)\right] \right\Vert ^{2} \left\Vert \left( T^{-1}\mathbf{W} ^{\prime }\mathbf{W}\right) ^{-1}-\left[ \mathbb{E}\left( T^{-1}\mathbf{W} ^{\prime }\mathbf{W}\right) \right] ^{-1}\right\Vert _{F} + \\ & \qquad \left\Vert T^{-1} \mathbb{E}\left(\mathbf{W}^{\prime} \mathbf{u} \right) \right\Vert ^{2} \left\Vert \left( T^{-1}\mathbf{W}^{\prime }\mathbf{ W}\right) ^{-1}-\left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W} \right) \right] ^{-1}\right\Vert _{F} \\ \end{split} \end{equation} First, consider $T^{-1}\left[ \boldsymbol{\tau }^{\prime }\mathbf{H}^{\prime }\mathbf{H}\boldsymbol{\tau }-\mathbb{E}\left( \boldsymbol{\tau }^{\prime } \mathbf{H}^{\prime }\mathbf{H}\boldsymbol{\tau }\right) \right] $. Note that \begin{equation*} \boldsymbol{\tau }^{\prime }\mathbf{H}^{\prime }\mathbf{H}\boldsymbol{\tau }= \boldsymbol{\tau }^{\prime }\left( \sum_{t=1}^{T}\mathbf{h}_{t}\mathbf{h} _{t}^{\prime }\right) \boldsymbol{\tau }=\sum_{t=1}^{T}\left( \boldsymbol{\ \tau }^{\prime }\mathbf{r}_{t}\right) \left( \mathbf{r}_{t}^{\prime } \boldsymbol{\tau }\right) =\sum_{t=1}^{T}\left( \sum_{i=1}^{k}h_{it}\right) \left( \sum_{j=1}^{k}h_{jt}\right) =\sum_{i=1}^{k}\sum_{j=1}^{k}\sum_{t=1}^{T}h_{it}h_{jt}. \end{equation*} Recalling that $h_{it}=x_{it}(\beta _{it}-\bar{\beta}_{iT})$, and hence, \begin{equation*} T^{-1}\left[ \boldsymbol{\tau }^{\prime }\mathbf{H}^{\prime }\mathbf{H} \boldsymbol{\tau }-\mathbb{E}\left( \boldsymbol{\tau }^{\prime }\mathbf{H} ^{\prime }\mathbf{H}\boldsymbol{\tau }\right) \right] =\sum_{i=1}^{k} \sum_{j=1}^{k}\left( T^{-1}\sum_{t=1}^{T}\tilde{h}_{ij,t}\right) , \end{equation*} where \begin{equation*} \tilde{h}_{ij,t}=h_{it}h_{jt}-\mathbb{E}(h_{it}h_{jt}). \end{equation*} Now consider $\mathbb{E}\left( T^{-1}\sum_{t=1}^{T}\tilde{h}_{ij,t}\right) ^{2}$ and note that \begin{equation*} \mathbb{E}\left( T^{-1}\sum_{t=1}^{T}\tilde{h}_{ij,t}\right) ^{2}=T^{-2}\sum_{t=1}^{T}\sum_{t^{\prime }=1}^{T}\mathbb{E}\left( \tilde{h} _{ij,t}\tilde{h}_{ij,t^{\prime }}\right) . \end{equation*} By Assumption (ref), $T^{-2}\sum_{t=1}^{T}\sum_{t^{ \prime }=1}^{T}\mathbb{E}\left( \tilde{h}_{ij,t}\tilde{h}_{ij,t^{\prime }}\right) =O\left( T^{-1}\right) $, and hence, by Lemma (ref), it follows that \begin{equation*} \left\vert T^{-1}\sum_{t=1}^{T}\tilde{h}_{ij,t}\right\vert =O_{p}\left( \frac{1}{\sqrt{T}}\right) . \end{equation*} Since by Assumption (ref), $k$ is a finite fixed integer, we can further conclude that \begin{equation} T^{-1}\left[ \boldsymbol{\tau }^{\prime }\mathbf{H}^{\prime }\mathbf{H} \boldsymbol{\tau }-\mathbb{E}\left( \boldsymbol{\tau }^{\prime }\mathbf{H} ^{\prime }\mathbf{H}\boldsymbol{\tau }\right) \right] =\sum_{i=1}^{k} \sum_{j=1}^{k}\left( T^{-1}\sum_{t=1}^{T}\tilde{h}_{ij,t}\right) =O_{p}\left( \frac{1}{\sqrt{T}}\right) . \end{equation} Now, consider, $T^{-1}\boldsymbol{\tau }^{\prime }\mathbf{H}^{\prime } \mathbf{u}$. Note that \begin{equation*} T^{-1}\boldsymbol{\tau }^{\prime }\mathbf{H}^{\prime }\mathbf{u}=T^{-1} \boldsymbol{\tau }^{\prime }\left( \sum_{t=1}^{T}\mathbf{h}_{t}u_{t}\right) =T^{-1}\sum_{t=1}^{T}\boldsymbol{\tau }^{\prime }\mathbf{h} _{t}u_{t}=T^{-1}\sum_{t=1}^{T}\sum_{i=1}^{k}h_{it}u_{t}=\sum_{i=1}^{k}\left( T^{-1}\sum_{t=1}^{T}h_{it}u_{t}\right) . \end{equation*} We have \begin{equation*} \mathbb{E}\left( T^{-1}\sum_{t=1}^{T}h_{it}u_{t}\right) ^{2}=T^{-2}\sum_{t=1}^{T}\mathbb{E}\left[ \left( h_{it}u_{t}\right) ^{2} \right] +T^{-2}\sum_{t=1}^{T}\sum_{t^{\prime }\neq t}\mathbb{E}\left( h_{it}h_{it^{\prime }}u_{t}u_{t^{\prime }}\right) . \end{equation*} Since $h_{it}=x_{it}(\beta _{it}-\bar{\beta}_{iT})$, and $\beta _{it}$ for $ i=1,2,\cdots ,k$ are distributed independently of $x_{js}$, $j=1,2,\cdots ,N$ , and $u_{s}$ for all $t$ and $s$, we can further write for any $t^{\prime }\neq t$ \begin{equation*} \mathbb{E}\left( h_{it}h_{it^{\prime }}u_{t}u_{t^{\prime }}\right) =\mathbb{E }\left( x_{it}u_{t}x_{it^{\prime }}u_{t^{\prime }}\right) \mathbb{E}\left[ (\beta _{it}-\bar{\beta}_{iT})(\beta _{it^{\prime }}-\bar{\beta}_{iT})\right] . \end{equation*} But, by Assumption (ref), $\mathbb{E}\left[ x_{it}u_{t}-\mathbb{E} (x_{it}u_{t})|\mathcal{F}_{t-1}\right] =0$ and we also have $\mathbb{E} (x_{it}u_{t})=0$ for $i=1,2,\cdots ,k$ and thus for any $t^{\prime }\neq t$ we have \begin{equation*} \mathbb{E}\left( x_{it}u_{t}x_{it^{\prime }}u_{t^{\prime }}\right) =\mathbb{ \ E}\left( x_{it}u_{t}\right) \mathbb{E}\left( x_{it^{\prime }}u_{t^{\prime }}\right) =0. \end{equation*} Therefore, \begin{equation*} \mathbb{E}\left( T^{-1}\sum_{t=1}^{T}h_{it}u_{t}\right) ^{2}=T^{-2}\sum_{t=1}^{T}\mathbb{E}\left[ \left( h_{it}u_{t}\right) ^{2} \right] =O\left( \frac{1}{T}\right) . \end{equation*} Hence, by Lemma (ref), $\left\vert T^{-1}\sum_{t=1}^{T}h_{it}u_{t}\right\vert =O_{p}\left( \frac{1}{\sqrt{T}} \right) $. Since, by Assumption (ref), $k$ is a finite fixed integer, we conclude that \begin{equation} T^{-1}\boldsymbol{\tau }^{\prime }\mathbf{H}^{\prime }\mathbf{u} =\sum_{i=1}^{k}\left( T^{-1}\sum_{t=1}^{T}h_{it}u_{t}\right) =O_{p}\left( \frac{1}{\sqrt{T}}\right) . \end{equation} By substituting ((ref)) and ((ref)) into ((ref)), and noting that $ \left \Vert T^{-1}\mathbb{E}\left( \mathbf{W}^{\prime }\mathbf{u}\right) \right\Vert^2 = O\left(T^{-(2\epsilon-d) }\right)$, for some $\epsilon \geq 1/2 $, \begin{equation*} \left\Vert T^{-1}\left[ \mathbf{W}^{\prime }\mathbf{R}\boldsymbol{\tau }+ \mathbf{W}^{\prime }\mathbf{u} - \mathbb{E}\left(\mathbf{W}^{\prime }\mathbf{ u} \right) \right] \right\Vert ^{2} = O_{p}(T^{-(1-d)}), \end{equation*} \begin{equation*} \left\Vert \left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) ^{-1}-\left[ \mathbb{E}\left( T^{-1}\mathbf{W}^{\prime }\mathbf{W}\right) \right] ^{-1}\right\Vert _{F}=O_{p}(T^{-(1/2-d)}), \end{equation*} and \begin{equation*} T^{-1}\left[ \mathbf{u}^{\prime }\mathbf{u}-\mathbb{E}\left(\mathbf{\ u} ^{\prime }\mathbf{u}\right) \right] =O_{p}(1/\sqrt{T}), \end{equation*} we conclude that \begin{equation*} T^{-1}\boldsymbol{\hat{\eta}}^{\prime }\boldsymbol{\hat{\eta}} =\sum_{i=1}^{k}\sum_{j=1}^{k}\left( T^{-1}\sum_{t=1}^{T}\sigma _{ijt,x}\sigma _{ijt,\beta }^{\ast }\right) +\bar{\sigma}_{u,T}^{2}+O_{p} \left( \frac{1}{\sqrt{T}}\right) +O_{p}\left(T^{-(1-d)}\right), \end{equation*} where $\sigma _{ijt,\beta }^{\ast }=\mathbb{E}\left[ (\beta _{it}-\bar{\beta} _{i,T})(\beta _{jt}-\bar{\beta}_{j,T})\right] $, $\bar{\beta} _{iT}=T^{-1}\sum_{t=1}^{T}\mathbb{E}(\beta _{it})$, and $\bar{\sigma} _{u,T}^{2}=T^{-1}\mathbb{E}\left( \mathbf{u}^{\prime }\mathbf{u}\right) $. We further have \begin{equation*} \bar{\Delta}_{\beta ,T}^{\ast }=\sum_{i=1}^{k}\sum_{j=1}^{k}\left( T^{-1}\sum_{t=1}^{T}\sigma _{ijt,x}\sigma _{ijt,\beta }^{\ast }\right) =T^{-1}\sum_{t=1}^{T}\left( \sum_{i=1}^{k}\sum_{j=1}^{k}\sigma _{ijt,x}\sigma _{ijt,\beta }^{\ast }\right) =\frac{1}{T}\sum_{t=1}^{T} tr\left( \boldsymbol{\Omega }_{\beta ,t}^{\ast }\boldsymbol{\Sigma }_{ \mathbf{x}_{k},t}\right) , \end{equation*} where $\boldsymbol{\Omega }_{\beta ,t}^{\ast }\equiv \left( \sigma _{ijt,\beta }^{\ast }\right) $ and $\boldsymbol{\Sigma }_{\mathbf{x} _{k},t}\equiv \left( \sigma _{ijt,x}\right) $ for $i,j=1,2,\cdots ,k$. By result 9(b) on page 44 of Lutkepohl1996handbook, we can further write \begin{equation*} tr\left( \boldsymbol{\Omega }_{\beta ,t}^{\ast }\boldsymbol{\Sigma }_{ \mathbf{x}_{k},t}\right) \geq k\left[ \text{det}\left( \boldsymbol{\Omega } _{\beta ,t}^{\ast }\right) \right] ^{1/k}\left[ \text{det}\left( \boldsymbol{ \Sigma }_{\mathbf{x}_{k},t}\right) \right] ^{1/k}. \end{equation*} But $k$ is a finite fixed integer. Furthermore, $\text{det}\left( \boldsymbol{\Omega }_{\beta ,t}^{\ast }\right) \geq 0$ and $\text{det}\left( \boldsymbol{\Sigma }_{\mathbf{x}_{k},t}\right) >0$, since $\boldsymbol{ \Omega }_{\beta ,t}^{\ast }$ and $\boldsymbol{\Sigma }_{\mathbf{x}_{k},t}$ are positive semi-definite and positive definite matrices, respectively. So, we can conclude that $\bar{\Delta}_{\beta ,T}^{\ast }\geq 0$ as required.
lemmaLet $y_{t} $ $t = 1, 2, \cdots, T $ be generated by ((ref)). Suppose Assumption (ref) and (ref) hold, and the cross products of coefficients of the signals in DGP for $y_t$ follow martingale difference processes such that \begin{equation*} \mathbb{E}\left[ \beta _{it}\beta _{jt}-\mathbb{E}(\beta _{it}\beta _{jt})| \mathcal{F}_{t-1}\right] =0, for i=1,2,\cdots ,k,\ j=1,2,\cdots ,k, and t=1,2,\cdots ,T. \end{equation*} Then, $\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\text{cov}(h_{ij,t},h_{ij,t^{ \prime }})=O(T)$ where $h_{ij,t}=x_{it}x_{jt}(\beta _{it}-\bar{\beta} _{iT})(\beta _{jt}-\bar{\beta}_{jT})$.
proofTo show this, let $\tilde{h}_{ij,t}=h_{ij,t}-\mathbb{E}\left( h_{ij,t}\right) $. We have \begin{align*} \sum_{t=1}^{T}\sum_{t^{\prime }=1}^{T}cov(h_{ij,t},h_{ij,t^{\prime }})& =\sum_{t=1}^{T}\mathbb{E}\left( \tilde{h}_{ij,t}^{2}\right) +2\sum_{t=2}^{T}\sum_{t^{\prime }=1}^{t}\mathbb{E}\left( \tilde{h}_{ij,t} \tilde{h}_{ij,t^{\prime }}\right) \\ & =\sum_{t=1}^{T}\mathbb{E}\left( \tilde{h}_{ij,t}^{2}\right) +2\sum_{t=2}^{T}\sum_{t^{\prime }=1}^{t}\mathbb{E}\left[ \tilde{h} _{ij,t^{\prime }}\mathbb{E}\left( \tilde{h}_{ij,t}|\mathcal{F}_{t-1}\right) \right] . \end{align*} But, $\mathbb{E}\left( \tilde{h}_{ij,t}|\mathcal{F}_{t-1}\right) =\mathbb{E} \left( h_{ij,t}|\mathcal{F}_{t-1}\right) -\mathbb{E}\left( h_{ij,t}\right) $ and under the conditions mentioned in this Lemma, \begin{align*} \mathbb{E}\left( h_{ij,t}|\mathcal{F}_{t-1}\right) & =\mathbb{E}\left( x_{it}x_{jt}|\mathcal{F}_{t-1}\right) \mathbb{E}\left[ (\beta _{it}-\bar{ \beta}_{iT})(\beta _{jt}-\bar{\beta}_{jT})|\mathcal{F}_{t-1}\right] \\ & =\mathbb{E}\left( x_{it}x_{jt}\right) \left\{ \mathbb{E}(\beta _{it}\beta _{jt}|\mathcal{F}_{t-1})-\bar{\beta}_{jT}\mathbb{E}(\beta _{it}|\mathcal{F} _{t-1})-\bar{\beta}_{iT}\mathbb{E}(\beta _{jt}|\mathcal{F}_{t-1})+\bar{\beta} _{iT}\bar{\beta}_{jT}\right\} \\ & =\mathbb{E}\left( x_{it}x_{jt}\right) \left\{ \mathbb{E}(\beta _{it}\beta _{jt})-\bar{\beta}_{jT}\mathbb{E}(\beta _{it})-\bar{\beta}_{iT}\mathbb{E} (\beta _{jt})+\bar{\beta}_{iT}\bar{\beta}_{jT}\right\} \\ & =\mathbb{E}\left( x_{it}x_{jt}\right) \mathbb{E}\left[ (\beta _{it}-\bar{ \beta}_{iT})(\beta _{jt}-\bar{\beta}_{jT})\right] =\mathbb{E}\left( h_{ij,t}\right) . \end{align*} Therefore, $\mathbb{E}\left( \tilde{h}_{ij,t}|\mathcal{F}_{t-1}\right) =0$. Hence, $\sum_{t=1}^{T}\sum_{t^{\prime }=1}^{T}\text{cov}(h_{ij,t},h_{ij,t^{ \prime }})=\sum_{t=1}^{T}\mathbb{E}\left( \tilde{h}_{ij,t}^{2}\right) =O\left( T\right) $.