Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Bridging factor and sparse models
frontmatter\runtitle{Bridging factor and sparse models}
\begin{aug}
,
\and
\address[A]{Department of Operations Research and Financial Engineering, Princeton University
\printead{e1}}
\address[B]{Center for Statistics and Machine Learning, Princeton University
\printead{e2}}
\address[C]{Department of Economics, Pontifical Catholic University of Rio de Janeiro
\printead{e3}}
\end{aug}
\begin{abstract}
Factor and sparse models are two widely used methods to impose a low-dimensional structure in high-dimensions. However, they are seemingly mutually exclusive. We propose a lifting method that combines the merits of these two models in a supervised learning methodology that allows for efficiently exploring all the information in high-dimensional datasets. The method is based on a flexible model for high-dimensional panel data, called factor-augmented regression model with observable and/or latent common factors, as well as idiosyncratic components. This model not only includes both principal component regression and sparse regression as specific models but also significantly weakens the cross-sectional dependence and facilitates model selection and interpretability. The method consists of several steps and a novel test for (partial) covariance structure in high dimensions to infer the remaining cross-section dependence at each step. We develop the theory for the model and demonstrate the validity of the multiplier bootstrap for testing a high-dimensional (partial) covariance structure. The theory is supported by a simulation study and applications.
\end{abstract}
\begin{keyword}[class=MSC]
\kwd[Primary ]{62H25}
\kwd{62F03}
\end{keyword}
\begin{keyword}
\kwd{factor models}
\kwd{penalized least-squares}
\kwd{high-dimensional inference}
\kwd{(partial) covariance structure}
\kwd{prediction}
\kwd{machine learning}
\kwd{supervised learning}
\end{keyword}
Introduction
With the emergence of new and large datasets in almost all disciplines, the correct characterization of the dependence among variables is of substantial importance. Usually, to achieve this goal, the literature has followed two seemingly orthogonal tracks over the last two decades. On the one hand, factor models have become an essential tool to summarize information in large datasets under the assumption that the remaining dependence structure is negligible. For instance, panel factor models are now applied to various problems, ranging from forecasting to causal inference and network analysis. However, on the other hand, there have been significant advances in parameter estimation in ultra high-dimensions under the assumption of sparsity or weak sparsity. That is, a variable depends only on a (very) small subset of the other variables. For an overview of these two topics and their exciting developments, see FLZZ20.
In this paper, we take an alternative route and combine the best of the two worlds described above to better characterize the dependence structure in high-dimensional datasets. More specifically, we consider that the covariance structure of a large set of variables, organized in a panel data format, is characterized as a combination of a factor structure, where factors can be either observed, unobserved, or both, and a weakly-sparse idiosyncratic component. This formulation is general enough to accommodate many data-generating processes of interest in economics, finance, epidemiology, and energy/engineering, for example. The proposed methodology has two ingredients: a multi-step estimation procedure and a new test for structure in high dimensional (partial) covariance matrices. The steps of the estimation procedure are as follows. In the first one, we take the original data and remove the effects of any observed factors. These factors can be deterministic terms such as seasonal dummies and/or trends or other observed covariates. The first step can be parametric or nonparametric, low or high dimensional. A latent factor model is then estimated using the residuals from the first stage. In a third step, we model the dependence among idiosyncratic terms as a weakly sparse regression estimated by the Least Absolute Shrinkage and Selection Operator (LASSO). A final additional fourth step can be used to build models for out-of-sample forecasting, taking into account each of the components in the previous steps. At each step, the null hypothesis of no remaining cross-section dependence can be tested by the proposed test for the (partial) covariance structure in high dimensions.
Our approach has many statistical applications. It can enhance high-dimensional prediction, select more interpretable variables, construct counterfactuals for policy evaluations, and depict (partial) correlation networks, among many others.
Motivation
Let $\boldsymbol Y_t:=(Y_{1,t},\ldots, Y_{n,t})'$ be a random vector generated as $Y_{i,t}=\boldsymbol\lambda_i'\boldsymbol F_t+ U_{i,t}$, for $i\in[n]$, $t\in[T]$, where $\boldsymbol\Sigma:=\boldsymbol\mathbb{E}(\boldsymbol U_t\boldsymbol U_t')$, with $\boldsymbol U_t:=(U_{1,t},\ldots, U_{n,t})'$, is not necessarily diagonal. Fix one component of interest $i\in[n]$, which serves as a response variable. Consider the following predictive models:
equation[equation omitted — 272 chars of source]
where $\boldsymbol Y_{-i,t}$ and $\boldsymbol U_{-i,t}$ are, respectively, vectors with the elements of $\boldsymbol Y_t$ and $\boldsymbol U_t$ without the $i$-th entry. $\mathcal{M}_3$ is a factor augmented regression since it is the same as $\mathbb{E}(Y_{i,t}| \boldsymbol F_t,\boldsymbol Y_{-i,t})$.
Suppose that we observe both $\boldsymbol F_t$ and $\boldsymbol U_{-i,t}$. Which one of the three models above is best in terms of mean square error ($\mathsf{MSE}$) for prediction? Comparison between $\mathcal{M}_1$ and $\mathcal{M}_2$ is not clear since it depends, among other things, on the magnitude of $\boldsymbol \Sigma$ relative to $\boldsymbol\Lambda'\boldsymbol\Lambda$, where $\boldsymbol\Lambda :=(\boldsymbol\lambda_1,\ldots,\boldsymbol\lambda_n)'$.
However, since the $\sigma$-algebras generated by $\boldsymbol Y_{-i,t}$ and $\boldsymbol F_t$ are both included in the $\sigma$-algebra generated by $(\boldsymbol F_t,\boldsymbol U_{-i,t})$, it is not surprising that $\mathsf{MSE}(\mathcal{M}_3)\leq\min[\mathsf{MSE}(\mathcal{M}_1),\mathsf{MSE}(\mathcal{M}_2)]$.
The same will hold true if we replace the models in (ref) by their best linear projections, which we denote by $\mathcal{\widetilde{M}}_j$ for $j\in\{1,2,3\}$. Therefore, we can write the “gains” of $\widetilde{\mathcal{M}}_3$ when compared to $\widetilde{\mathcal{M}}_1$ and $\widetilde{\mathcal{M}}_2$:
align*[align* omitted — 381 chars of source]
where $\boldsymbol\vartheta_i$ is coefficients of the projection of $U_{i,t}$ onto $\boldsymbol U_{-i,t}$; $\boldsymbol\Sigma_{-i,-i}$ is $\boldsymbol\Sigma$ excluding the $i$-th row and column; $\boldsymbol\Delta_{1,i} := \boldsymbol\Lambda_i -\boldsymbol\beta_i'\boldsymbol\Lambda_{-i}$; $\boldsymbol\Delta_{2,i} := \boldsymbol\beta_i-\boldsymbol\vartheta_i$; and $\boldsymbol\beta_i$ is the coefficient of the projection of $Y_{i,t}$ onto $\boldsymbol Y_{-i,t}$. From the previous expressions, it becomes evident that both $\widetilde{\mathcal{M}}_1$ and $\widetilde{\mathcal{M}}_2$ are restrictions on $\widetilde{\mathcal{M}}_3$. Broadly speaking, whenever one does not expect to have an exact factor model, there are potential gains of taking into account the contribution of the idiosyncratic components $\boldsymbol U_{-i,t}$. Therefore, we use $\widetilde{\mathcal{M}}_3$ as the base model for the estimation methodology described in Section (ref).
Main Contributions and Comparison with the Literature
The contributions of this paper are multi-fold. First, our methodology bridges the gap between two apparently competing methods for high-dimensional modeling; see, for example, the discussion in illusion2017 and jFyKkW2020. This yields a vast number of potential applications and spin-offs. For instance, in jFrMmM2020a, we apply the methods developed here to evaluate the effects of interventions and contribute to the literature on synthetic controls and related methods by combining the approaches of lGtM2016, and cCrMmM2018. Therefore, in our setup, both a common strong factor structure and weak sparsity can coexist.
Second, the methodology proposed here contributes to the forecasting literature. For instance, in the second application considered in this paper, we build forecasting models based on a large cross-section of macroeconomic variables. We call this method the FarmPredict. We show that combining factors and a sparse regression strongly outperforms the traditional principal component regression as in jSmW2002b. Therefore, FarmPredict can be an additional contribution to the forecasting and machine learning toolkit. The method can be easily extended to a multivariate setting combining factor-augmented vector autoregressions (FAVAR) as in bBjBpE2005 and sparse vector models as in aKlC2015 and rMmMeM2019. Our methodology can also be applied in areas beyond economics. For example, it can be useful to construct forecasting models for the spread of infectious diseases, such as COVID-19, by taking into account co-movements and spillovers among different locations.
Third, we show the consistency of factor estimation based on the residuals of a first-step regression. Our results hold for both parametric (linear or nonlinear) and nonparametric first stage. A high-dimensional first stage is also allowed. Note that current results in the literature consider that factors are estimated based on observed data, and our derivations favor a much more flexible and general setup Bai2003,jBsN2002,jBsN2006. More specifically, our methodology allows settings where there are both observed and latent factors and trend-stationary data. In the latter, the trend can be first removed by (nonparametric) first-stage regression. Whenever the unobserved factors and the observed covariates are correlated, the method proposed in mhP2006 can be used, and all results follow directly.
Fourth, we contribute to the LASSO literature. LASSO can not be model selection consistent for highly correlated variables. By decomposing covariates into factors and idiosyncratic components, namely the idea of lifting, we decorrelate the variables and make the model selection condition much easier to hold; see, for example, jFyKkW2020. We show the consistency of the estimates based on the residuals of the previous steps. Our results are derived under restrictions on the population covariance matrix of the data and not on the estimated one, as is usual in many papers. See, for example, vandegeer2009. Furthermore, we derive our results under much mild conditions than the ones considered in jFyKkW2020.
Fifth, we extend the results in vCdCkK2013,CCK2018 to strong-mixing data in order to construct hypothesis tests for covariance and partial covariance structure in high dimensions.\footnote{alex2020 also extended vCdCkK2013. However, their setup differs from ours as they only consider the case of independent and identically distributed data.} We also establish the consistency of a new estimator of the partial covariance matrix in high-dimensions and strong-mixing data. Our proposed tests can be used to infer if the (partial) covariance matrix of a high-dimensional random vector is diagonal or block-diagonal. More generally, we can test any pre-defined structure. Furthermore, we show that the test remains valid when we use the residuals from a previous step estimation to compute the covariance matrix. This result allows us to apply the test to the multi-stage estimation procedure proposed here. These are important results for many applications in economics, finance, epidemiology, and many other areas. For instance, our inference procedure can serve as a diagnostic and misspecification tool. For panel data models with interactive fixed effects as in mhP2006, jB2009, rMmW2015 and jByL2017, our test can be directly applied to uncover the dependence structure among cross-sectional units before and after accounting for common factor components. If the factor structure is informative enough, we expect the idiosyncratic covariance matrix to be almost sparse. If this is not the case, we may have possibly underestimated the number of factors. One popular application is in asset pricing, as discussed in pGeOoS2019 and the empirical section of this paper. There are a huge number of proposed factors as described in gFsGdX2020. We can apply our methodology to test for omitted factors and estimate network connections among firms, as in fxDkY2014 and cBgGgL2020. Finally, as a diagnostic tool, our paper tackles the same problem as pGeOoS2019. However, we take an alternative solution strategy that relies on a much different set of hypotheses; see also pGeOoS2020.
Although our results are derived under the assumption that the number of factors is known, simulation results presented in the Supplementary Material provide evidence that the test has good finite-sample properties even when the number of factors is determined by data-driven methods commonly found in the literature. In addition, due to factor augmentation, our method is robust to overestimating the number of factors. Over the past years, a vast number of papers proposed different methods to test for covariance structure in high dimensions. See, for example, cai2017 and the references therein. To the best of our knowledge, we complement all the previous papers by simultaneously considering high-dimensions, strong-mixing data with mild distributional assumptions, and pre-estimation when constructing tests for both covariance and partial covariance structure.
Finally, it is essential to highlight the theoretical challenges that we tackle in this paper. First, to derive the properties of our proposed test for (partial) covariance structure, we prove a new high dimensional Central Limit Theorem (CLT) for strong mixing sequences. To our knowledge, this CLT is new in the literature and allows us to apply Gaussian approximation results in a much more general framework. Second, as a side result, to develop the test, we first show the consistency of kernel-based estimation of a high-dimensional long-run covariance matrix of strong-mixing processes. This is also a new result with significant consequences for the theory of high-dimensional regression with dependent errors. Second, all the non-asymptotic bounds for our multi-step estimators are derived under the assumption that the distributions of the random variables in the model have polynomial tails, and the estimation errors of previous steps are also taken into account. This introduces several difficulties in proving the results but makes our method well-suited for applications with fat-tailed data, such as those observed in financial applications. Finally, in the derived bounds, the strong mixing coefficient appears explicit and can be allowed to grow with the sample size. This not only introduces technical challenges but makes our results very general.
Summarizing, our approach provides:
enumerate• A systematic way to unify factor and sparse models to construct statistical specifications which use all available information and that can be applied, for example, to
\begin{enumerate}
• Forecasting in a high-dimensional setting;
• Construction of counterfactuals to aggregate data; or
• Estimation of partial correlation networks;
\end{enumerate}
• An inferential procedure to test for general structures in covariance and partial covariance matrices which can be useful, among other things, for:
\begin{enumerate}
• Test for misspecification in factor models; or
• Test for nontrivial links among idiosyncratic units.
\end{enumerate}
Organization of the Paper
In addition to this Introduction, the paper is organized as follows. We present the model setup and assumptions in Section (ref). The theoretical results are presented in Section (ref). We discuss the empirical application in Section (ref). Section (ref) concludes. All proofs, additional discussion, simulation, and empirical results are deferred to the Supplementary Material due to space constraints. Tables and figures in the Supplementary Material are referenced with an “S” before the number.
Notation
All random quantities (real-valued, vectors and matrices) are defined in a common probability space $(\Omega,\mathscr{F},\P)$. We denote random variables by an upper case letter, $X$, and its realization by a lower case letter, $X=x$. The expected value operator is with respect to the $\P$ law, such that $\mathbb{E}(X):=\int_\Omega X(\omega) \mathrm{d}\P(\omega)$. Matrices and vectors are written in bold letters $\boldsymbol X$. Except for the number of factors, $r$, and the number of covariates, $k$, defined below, all other dimensions are allowed to depend on the sample size ($T$). However, we omit this dependency throughout the paper to avoid clustering the notation prematurely. Also, we write $[n]:=\{1,\dots,n\}$ for $n\in\N$ and denote the cardinality of a set $\mathcal{S}$ by $|\mathcal{S}|$.
We use $\|\,\cdot\,\|_p$ to denote the $\ell^p$ norm for $p\in[1,\infty]$, such that for a $d$-dimensional (possibly random) vector $\boldsymbol{X}=(X_1,\ldots, X_d)'$, we have $\|\boldsymbol{X}\|_p:=(\sum_{i=1}^d|X_i|^p)^{1/p}$ for $p\in[1,\infty)$ and $\|\boldsymbol{X}\|_\infty:=\sup_{i\leq d} |X_i|$. If $\boldsymbol X$ is a $(m\times n)$ possibly random matrix, then $\|\boldsymbol{X}\|_p$ denotes the matrix $\ell^p$-induced norm and $\|\boldsymbol X\|_{\max}$ denotes the maximum entry in absolute terms of the matrix $\boldsymbol X$. Note that whenever $\boldsymbol X$ is random, then $\|\boldsymbol{X}\|_p$ for $p\in[1,\infty]$ and $\|\boldsymbol{X}\|_{\max}$ are random variables. We also reserve the symbol $\|\,\cdot\,\|$ without subscript for the Euclidean norm $\|\cdot\| :=\|\cdot\|_2$ for both vectors and matrices. We denote the $\ell^p(\P)$ norm of $X$ by ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_p$ for $p\in[0,\infty]$, i.e., ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_p:=(\mathbb{E} |X|^p)^{1/p}$ for $p\in[1,\infty)$ and ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_\infty$ is the essential supremum of $X$ (in respect to $\P$). Also, when $\boldsymbol X$ is a random vector, we define ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \boldsymbol X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_p:=\sup_{\|\boldsymbol{u}\|\leq 1}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \boldsymbol{u}'\boldsymbol X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_p$. Note that $\|\boldsymbol X\|_p$ is random variable while ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \boldsymbol X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_p$ is non-random scalar for $p\in[0,\infty]$.
For any vector $\boldsymbol{X}$, $\mathsf{diag}\,(\boldsymbol{X})$ is a diagonal matrix whose diagonal is the elements of $\boldsymbol{X}$. $\1(A)$ is an indicator function ion the event $A$, i.e, $\1(A) = 1$ if $A$ is true and $0$ otherwise. For any matrix (possibly random) $\boldsymbol M\in \mathbb{R}^{n\times T}$, $M_{i,t}\in\mathbb{R}$ represents the entry for row $i$ and column $t$, $\boldsymbol M_{t}\in\mathbb{R}^n$ is the column-vector with all rows of column $t$ and $\boldsymbol M_{i,\cdot}\in \mathbb{R}^T$ is a column-vector with the transpose of all elements of row $i$. We decide not to write $\boldsymbol M_{t}\in\mathbb{R}^n$ as $\boldsymbol M_{\cdot,t}$ to avoid a cumbersome notation. Finally, for non-negative sequences $x_m$ and $y_m$, we write $x\lesssim y$ if there is a constant $C$ independent of $m$ such that $x_m\leq C y_m$ for all $m$. Also, we write $x\asymp y$ if both $x\lesssim y$ and $y \lesssim x$. Similarly, for non-negative random sequences $X_m$ and $Y_m$, we write $X_m\lesssim_\P Y_m$ if for every $\epsilon>0$ there is a constant $C$ independent of $m$ such that $\P(X_m\leq C Y_m)\leq \epsilon$ for all $m$. Also, $X_m\asymp_\P Y_m$, if both $X_m\lesssim_\P Y_m$ and $Y_m \lesssim_\P X_m$.
Setup and Method
Data Generating Process
We consider a very general panel data model rich enough to nest several important cases in economics, finance, and related areas. We define the following Data Generating Process (DGP).
assumption[DGP] For $T\geq 4$ and $n\geq 2$, the process $\{Y_{i,t}:i\in[n], t\in [T]\}$ is generated by covariate-adjusted factor model
\begin{equation}
Y_{i,t} =\boldsymbol\gamma_{i}' \boldsymbol X_{i,t}+\underbrace{\boldsymbol\lambda_i'\boldsymbol F_t, +U_{i,t}}_{=:R_{i,t}}
\end{equation}
where $\boldsymbol X_{i,t}$ is a $k$-dimensional observable (random) vector which may also include a constant term and is typically used for adjustments of heterogeneity, seasonality, and covariate dependence, $\boldsymbol F_t$ is a $r$-dimensional vector of common latent factors, and $U_{i,t}$ is a zero mean idiosyncratic component.\footnote{For simplicity, we assume that all the units $i$ have the same number of covariates ($k$). The framework can certainly accommodate situations where $k_i$ depends on $i$. It also includes cross-sectional regression as a specific example.} The unknown parameters are $\boldsymbol\gamma_i\in\mathbb{R}^{k}$, the factor loadings $\boldsymbol\lambda_i$, and the covariance matrix of the idiosyncratic components. Finally, we assume that $\boldsymbol X_{i,t}$, $\boldsymbol F_t$ and $U_{i,t}$ are mutually uncorrelated, but can be serially autocorrelated.
RemIn Assumption (ref) we consider that $k$, the dimension of $\boldsymbol X_{i,t}$, is finite and fixed. Furthermore, the relation between $Y_{i,t}$ and $\boldsymbol X_{i,t}$ is linear. This is for the sake of exposition. As our theoretical results are written in terms of the consistency rate of the first-step estimation, the DGP can be made much more general by just changing the rates.
RemThe assumption that $\boldsymbol X_{i,t}$, $\boldsymbol F_t$ and $U_{i,t}$ are mutually uncorrelated can be relaxed. Whenever $\boldsymbol X_{i,t}$ is correlated with $\boldsymbol F_t$ and $U_{i,t}$, and the interest lies of the estimation of the parameters $\boldsymbol\gamma_i$, $i\in[n]$, the method proposed by mhP2006 can applied in the first-stage of the procedure considered in this paper and our theoretical results will follow. Nevertheless, we provide several examples where the assumption that $\boldsymbol X_{i,t}$, $\boldsymbol F_t$ and $U_{i,t}$ are mutually uncorrelated is reasonable.
Our modeling strategy does not stop at (ref). Instead, we attempt to further explain $U_{i,t}$ and impose dynamics on $\boldsymbol F_t$. The former allows us to use other idiosyncratic components to further explain $Y_{i,t}$ and hence, increase the information set. The latter builds a dynamic model $\boldsymbol F_t$ to facilitate out-of-sample prediction, for instance.
For each $i\in[n]$, let $\boldsymbol W_{i,t}$ be a vector whose elements form a (non-empty) subset of $\left(\boldsymbol U_{-i,t}', \boldsymbol U_{t-1}',\ldots,\boldsymbol U_{t-l_U}'\right)'$, where $l_U$ is a non-negative integer (much) smaller than $T$. This subset of variables attempts to explain further $U_{i,t}$ by the following population regression model:
equation[equation omitted — 91 chars of source]
where $\mathbb{E}\left(\boldsymbol W_{i,t}'V_{i,t}\right) = \boldsymbol 0$, for $i\in[n]$ and $t\in[T]$.
For simplicity of exposition, we assume that the dimension of $\boldsymbol W_{i,t}$ is the same for all $i\in[n]$, which we denote by $d_W$. Clearly, $d_W\leq n (l_U+1)$.
Similarly, for a non-negative integer $l_F$ (much) smaller than $T$, let $\boldsymbol G_{j,t}$ be a vector whose elements form a (non-empty) subset of $\left(\boldsymbol F_{t}', \boldsymbol F_{t-1}',\ldots,\boldsymbol F_{t-l_F}'\right)'$, for $j\in[r]$, which attempts to explain further the $j^{th}$ latent factor. Consider the following population regression model:
equation*[equation* omitted — 77 chars of source]
where $\mathbb{E}\left(\boldsymbol G_{j,t}'{V_{j,t}^F}\right) = \boldsymbol 0$, for $j\in[r]$ and $t\in[T]$. Once again, we assume that the dimension of $\boldsymbol G_{j,t}$ is the same for all $j\in[r]$, which we denote by $d_G$, where $d_G\leq r(l_F+1)$. Note that we do not necessarily exclude $\boldsymbol F_t$ from $\boldsymbol G_{t}$ to contemplate cases when the contemporaneous factor is available in a prediction exercise. For future reference we write (ref) as
equation[equation omitted — 108 chars of source]
where $\boldsymbol G_t :=(\boldsymbol G_{1,t}',\dots, \boldsymbol G_{r,t}')'$, $\boldsymbol V_{t}^F:=( V_{1,t}^F,\dots, V_{r,t}^F)'$, and $\boldsymbol P:=(\boldsymbol\rho_1:\dots:\boldsymbol\rho_r)'$.
Ex[Asset Pricing Models]
Suppose $Y_{i,t}$ is the return of an asset $i$ at time $t$ and let $\boldsymbol X_{i,t}:=\boldsymbol X_t$ be a set of $k$ observable common risk factors, such as the market returns or Fama-French factors eFkF1993,eFkF2015. $\boldsymbol F_t$ can be a set of additional, non-observable risk factors. Several asset pricing models, such as the Capital Asset Pricing Model (CAPM) or the Arbitrage Pricing Theory (APT) model, are nested into this general framework. The idiosyncratic terms, $U_{i,t}$, can be non-trivially correlated across assets, representing links among firms that are not captured by the common factor structure. In this case, we may be interested in estimating a regression where $\boldsymbol W_{i,t}=\boldsymbol U_{-i,t}$ such that:
\[
U_{i,t} =\boldsymbol \theta_i'\boldsymbol U_{-i,t} + V_{i,t}.
\]
The coefficients $\boldsymbol \theta_i$ represent the links among firms after controlling for the factors. The structures of the covariance and partial covariance matrix of $\boldsymbol U_t=\left(U_{1,t},\ldots,U_{n,t}\right)'$ also shed light on these potential links.
Ex[Networks]
Model (ref) also complements the network specifications discussed in Barigozzi and Hallin (2016,2017b)\nocite{mBmH2016,mBmH2017b} and mBcB2019. Furthermore, as discussed in the previous example, the test proposed here can be used to detect network links as in fxDkY2014, and cBgGgL2020. As another example, $Y_{i,t}$ can be the (realized) volatility of financial assets and $\boldsymbol X_{i,t}:=\boldsymbol X_t$ can be volatility factors as in eAeG2021.
Ex[FAVAR]
In the case where the index $i$ represents a different dependent (endogenous) variable and $U_{i,t}$ is a dependent process, model (ref) turns out to be equivalent to the Factor Augmented Vector Autoregressive (FAVAR) model of bBjBpE2005. In this case, $\boldsymbol{X}_{i,t}$ may include a constant and seasonal dummies, and $\boldsymbol W_{i,t}=\left(\boldsymbol U_{t-1}',\ldots,\boldsymbol U_{t-l_U}'\right)'$. Furthermore, in order to construct $h$-step-ahead out-of-sample forecasts, we can set $\boldsymbol G_{j,t}=(\boldsymbol F_{t-h}',\ldots,\boldsymbol F_{t-l_F}')'$.
Ex[Panel Data Models]
Model (ref) is the panel model with iterative fixed-effects considered in lGtM2016, where the authors propose an alternative to the Synthetic Control method of aAjG2003 to evaluate the effects of regional policies. Model (ref) is also in the heart of the FarmTreat method of jFrMmM2020a, where the authors set $\boldsymbol W_{i,t}=\boldsymbol U_{-i,t}$.
The Method
The method proposed here for estimation, inference, and prediction consists of multiple stages, where the residuals' covariance structure can be tested at the end of each stage.
enumerate• For each $i\in[n]$ run the regression:
\[
Y_{i,t}=\boldsymbol\gamma_i'\boldsymbol{X}_{i,t}+R_{i,t},\quad \,t\in[T],
\]
and compute $\widehat{R}_{i,t}:= Y_{i,t} - \widehat{\boldsymbol\gamma}_i'\boldsymbol{X}_{i,t}$. The first stage may consist of a regression on a constant, a deterministic time trend, and seasonal dummies, for instance, or, as in Example (ref), a regression on observed factors. After removing the contribution from the observables, we can use the test for the null hypothesis of no remaining (partial) covariance structure to check if the (partial) covariance of $R_{i,t}$ is dense or sparse. If it is dense, we move to Step 2. Otherwise, we jump directly to Step 3. This first parametric, low-dimensional step can be replaced by a nonlinear/nonparametric regression or by a high-dimensional model when, for example, the number of observed factors is large. As pointed out in Remark (ref), the Pesaran's (2006) estimator can be also used whenever correlation between $\boldsymbol{X}_{i,t}$ and $R_{i,t}$ is allowed. This will be discussed more in the subsequent sections.
• Write $\boldsymbol{R}_t:=(R_{1,t},\ldots,R_{n,t})'$ and $\boldsymbol{R}_t=\boldsymbol\Lambda\boldsymbol{F}_t+\boldsymbol{U}_t$.
This step consists of estimating $\boldsymbol\Lambda$ and $\boldsymbol{F}_t$, for $t\in[T]$, through principal component analysis (PCA) \footnote{Another approach is to use the joint estimation as in agarwal2012noisy.
The key difference between the two approaches is the optimization-based approach (fully-iterated) and the one-step approach as well as the different assumptions behind the two approaches. The PCA approach is based on the strong factor assumption with a large eigengap but does not impose the sparse structure on the idiosyncratic component covariance matrix. See FLM2013 for additional discussion. }
of $\widehat{\boldsymbol R}_t$ and compute
\[
\widehat{\boldsymbol U}_t=\widehat{\boldsymbol R}_t - \widehat{\boldsymbol\Lambda}\widehat{\boldsymbol F}_t.
\]
After estimating the factors and loadings, we apply our testing procedure to check for the remaining covariance structure in $\boldsymbol U_t$. The second-step estimation can be carried out also by dynamic factor models. In Section (ref) we discuss the selection of the number of factors.
• Now, define $\widehat{\boldsymbol W}_{i,t}$ be a non-empty subset of $(\widehat{\boldsymbol U}_{-i,t}', \widehat{\boldsymbol U}_{t-1}',\ldots,\widehat{\boldsymbol U}_{t-l_U}')'$ where $l_U$ is a non-negative integer (lag). This includes contemporary regression for association studies and lagged regression for prediction of $U_{i,t}$ as two specific examples. The third step consists of a sparse regression to estimate the following model for each $i\in[n]$:
\[
\widehat{U}_{i,t}=\boldsymbol\theta_i'\widehat{\boldsymbol{W}}_{i,t} +V_{i,t};\qquad t\in[T].
\]
The regression in Step 3 provides useful augmentation for reducing the error in explaining $Y_{i,t}$ in (ref) further from $U_{i,t}$ to $V_{i,t}$ and hence the prediction error for $Y_{i,t}$; see (ref).
• (For partial covariance analysis only) Estimate the following sparse regression model for each $i,j\in[n]$:
\[
\widehat{U}_{i,t}=\boldsymbol\theta_{i,j}'\widehat{\boldsymbol{U}}_{-ij,t} +V_{i,j,t};\qquad t\in[T],
\]
where $\widehat{\boldsymbol{U}}_{-ij,t}$ is the vector $\widehat{\boldsymbol{U}}_{t}$ without $i^{th}$ and $j^{th}$ elements. Let $\{\hat V_{i,j,t}\}$ be the residuals. Then compute the partial correlation as the sample correlation of $\{(\hat V_{i,j,t}, \hat V_{j,i,t})\}_{t=1}^T$.
• (For forecasting only) Define $\widehat{\boldsymbol G}_{t}:=(\widehat{\boldsymbol G}_{1,t}',\dots, \widehat{\boldsymbol G}_{r,t}')'$ where $\widehat{\boldsymbol G}_{j,t}$ is a non-empty subset of $(\widehat{\boldsymbol F}_{t}', \widehat{\boldsymbol F}_{t-1}',\ldots,\widehat{\boldsymbol F}_{t-l_F}')'$ for each $j\in[r]$;
and $l_F$ is a non-negative integer (lag) \footnote{We write at this generality to accommodate the cross-sectional applications in which no latent factor needs to be predicted as in the principal component regression. In this case, $\widehat{\boldsymbol{G}}_{j,t} = \widehat{F}_{j,t}$ and this step is not needed.}. This step is multiple linear regression to estimate the following model for each $j\in[r]$:
\[
\widehat{F}_{j,t}=\boldsymbol\rho_j'\widehat{\boldsymbol{G}}_{j,t} +V_{j,t}^F,\qquad t\in[T].
\]
This step aims at establishing a predictive model for latent factors.
The estimator $\widehat{\boldsymbol P}:=(\widehat{\boldsymbol\rho}_1:\dots:\widehat{\boldsymbol\rho}_r)'$ will be used in the predictive model defined in (ref).
Predictive Models
In a pure prediction exercise, one is usually interested in the linear projection of $Y_{i,t}$ onto $(\boldsymbol X_{i,t}',\boldsymbol G_t', \boldsymbol W_{i,t}')'$ motivated by the discussion in Section (ref). This results in the factor-augmented regression model (FARM)
equation[equation omitted — 227 chars of source]
in which $\boldsymbol P {\boldsymbol G}_t$ predicts $\boldsymbol F_t$; see (ref). This
can used for prediction, following the steps described in Section (ref), by
equation[equation omitted — 270 chars of source]
We call the prediction model {FarmPredict}.
Note that model (ref) is equivalent to using the predictors $ \boldsymbol X_{i,t}, {\boldsymbol Y}_{-i,t}$ and ${\boldsymbol F}_t$, which augment predictors $ \boldsymbol X_{i,t}, {\boldsymbol Y}_{-i,t}$ by using the common factors ${\boldsymbol F}_t$. The form in (ref) mitigates the collinearity issues in high dimensions. Model (ref) also bridges factor regression ($\boldsymbol \theta_i = \boldsymbol 0$) on one end and (sparse) regression on the other end with $\boldsymbol \lambda_i = \boldsymbol \Lambda_{-i}' \boldsymbol \theta_i$, where $\boldsymbol \Lambda_{-i}$ is the loading matrix without the $i$th row. In the latter, if we set $\boldsymbol W_{i,t}=\boldsymbol U_{-i,t}$ and $\boldsymbol G_t = \boldsymbol F_t$ (hence, $\boldsymbol P = \boldsymbol I_r$), model (ref) becomes a (sparse) regression model:
equation*[equation* omitted — 162 chars of source]
In this case, (ref) decorrelates the predictor $\boldsymbol R_{-i,t}$, which makes the model selection consistency much easier to satisfy and forms the basis of FarmSelect in jFyKkW2020. Our contribution in this specific task is to allow heterogeneity adjustments, resulting in the estimated data $\boldsymbol R_t$. In general, for model (ref) with sparsity, FarmPredict chooses additional idiosyncratic components to enhance the prediction of the factor regression.
Ex[First-Order FAVAR]
Consider the case of model (ref) where $\boldsymbol X_{i,t}=1$ and $l_U=l_F=1$, and set $\boldsymbol W_{i,t}=\boldsymbol U_{t-1}$ and $\boldsymbol G_t=\boldsymbol F_{t-1}$. Therefore, we can write
\begin{equation}
\begin{split}
\boldsymbol Y_t & = (\boldsymbol I - \boldsymbol \Theta)\boldsymbol \gamma + (\boldsymbol\Lambda\boldsymbol P - \boldsymbol\Theta\boldsymbol\Lambda)\boldsymbol F_{t-1} + \boldsymbol\Theta\boldsymbol Y_{t-1} + \boldsymbol V_{t}^Y,\\
& = \boldsymbol \Theta_{0} + \boldsymbol \Theta_{F}\boldsymbol F_{t-1} + \boldsymbol\Theta \boldsymbol Y_{t-1} + \boldsymbol V_{t}^Y, \end{split}
\end{equation}
where $\boldsymbol Y_{t}=(Y_{1t},\ldots,Y_{nt})'$, $\boldsymbol\gamma=(\gamma_1,\ldots,\gamma_n)$, and $\boldsymbol\Theta=(\boldsymbol\theta_1:\cdots:\boldsymbol\theta_n)'$. Model (ref) is a first-order FAVAR model.
Covariance Structure and Inference
In several applications, the structure of the idiosyncratic components $\boldsymbol U_t=(U_{1,t},\ldots, U_{n,t})'$ is the objective of interest. An
estimator for $\boldsymbol\Sigma := \mathbb{E}(\boldsymbol U_t \boldsymbol U_t')$ could be simply given by
equation[equation omitted — 138 chars of source]
The task of estimating $\boldsymbol\Sigma$ is well documented in literature even in high-dimensions; see, for example, FLZZ20, oLmW2021b, and the references therein. Nevertheless, we show that (ref) can be used within our testing framework.
In order to proper understand the (linear) relation between a pair $(U_{i,t},U_{j,t})$ of $\boldsymbol U_t$, a simple covariance estimate sometimes is not enough. In applications, it is often desirable to directly measure how $U_{i,t}$ and $U_{j,t}$ are connected. By direct connection, we mean the relation between those units removing the contribution of other variables of $\boldsymbol U_t$. For this purpose, we use the partial covariance between $U_{i,t}$ and $U_{j,t}$, defined for any pair $i,j\in[n]$ as $\pi_{i,j} := \mathbb{E}(V_{i,j,t}V_{j,i,t})$, where $V_{i,j,t}:=U_{i,t} - \mathsf{Proj}(U_{i,t}|\boldsymbol U_{-ij,t})$ and $\mathsf{Proj}(U_{i,t}|\boldsymbol U_{-ij,t})$ denotes the linear projection of $U_{i,t}$ onto the space spanned by all the units except $i$ and $j$, which we denote by $\boldsymbol U_{-ij,t}$. As in jPpWnZjZ2009, we suggest to estimate the partial covariance matrix $\boldsymbol\Pi := (\pi_{i,j})$ by
equation[equation omitted — 191 chars of source]
where $\widehat{V}_{i,j,t}:=\widehat{U}_{i,t}-\widehat{\boldsymbol\theta}_{i,j}'\widehat{\boldsymbol U}_{-ij,t}$ is the residual of the LASSO regression of $\widehat{U}_{i,t}$ onto $\widehat{\boldsymbol U}_{-ij,t}$ obtained in step 4 for $i,j\in[n]$.
We also would like to conduct formal tests on the population structure of $\boldsymbol U_t$. Specifically, we propose a test for the following null hypothesis:
equation[equation omitted — 158 chars of source]
for a given subset $\mathcal{D}$, where, for a given $(n\times n)$ matrix $\boldsymbol M$, the notation $\boldsymbol M_\mathcal{D}$ denotes the $d:=|\mathcal{D}|$-dimensional vector of $\mathsf{vec}\,(\boldsymbol M)$ indexed by the elements in $\mathcal{D}\subseteq[n]^2$. Note we allow $d$ to diverge as $n,T\to\infty$. For testing the structure on the partial covariance matrix, consider:
equation[equation omitted — 164 chars of source]
The null hypotheses (ref) and (ref) nest several cases of interest. The most common would be to test for a diagonal or a block diagonal structure in $\boldsymbol\Sigma$ and/or $\boldsymbol\Pi$. But it also accommodates other structures.\footnote{With minor changes, the proposed test can also be used to test the null $\boldsymbol M\mathsf{vec}\,(\boldsymbol\Sigma)=\boldsymbol m$ for some $d\times n^2$ matrix $\boldsymbol M$ and $d$-dimensional vector $\boldsymbol m$ where $d:=d_T$ is also a function of $T$.} The challenges for testing (ref) and (ref) can be summarized as follows:
enumerate• As we allow for both $n$ and $d$ to diverge as $T$ grows, sometimes at a faster rate, we have a high-dimensional test where some sort of Gaussian approximation result for dependent data must be deployed as we also allow the number covariances to be tested $d$ to diverge. In this case, a high-dimensional long-run covariance matrix must be estimated if one expects to get the (asymptotic) correct test size.
• We do not observe $\{\boldsymbol U_t\}$ or $\{V_{i,j,t}\}$. Instead, we have an estimate of both from a postulated model on observable random variables. Therefore, the estimation error must be considered to claim some sort of asymptotic properties of the test. In fact, it is not uncommon to obtain estimates of both $\{\boldsymbol U_t\}$ and $\{V_{i,j,t}\}$ from a multi-stage estimation procedure as we illustrate later.
We propose to test (ref) using the statistic
equation[equation omitted — 164 chars of source]
The quantiles of $S_\mathcal{D}^\Sigma$ are approximated by a Gaussian bootstrap. To describe the procedure, let $\boldsymbol\Upsilon_{\Sigma}$ denote the $(d\times d)$ covariance matrix for the vectorized submatrix $(\widetilde{\sigma}_{i,j})_{(i,j)\in\mathcal{D}}$, where $\widetilde{\sigma}_{i,j}:=\frac{1}{T}\sum_{t=1}^T U_{i,t}U_{j,t}$. Since the process $\{\boldsymbol U_{t}\}$ might present some form of temporal dependence (refer to Assumption (ref)(c)), we estimate $\boldsymbol\Upsilon_{\Sigma}$ using a Newey-West-type estimator. For a given $\mathcal{K}\in\K$, where $\K$ is a class of kernel functions described below in (ref) and bandwidth $h>0$ ,
$\boldsymbol\Upsilon_\Sigma$ is estimated by
equation[equation omitted — 303 chars of source]
where $\widehat{\boldsymbol D}_{\Sigma,t}$ is a $d$-dimensional vector with entries given by $\widehat{U}_{i,t}\widehat{U}_{j,t}-\widehat{\sigma}_{i,j}$ for $(i,j)\in\mathcal{D}$, where $\widehat{\sigma}_{i,j}$ is the $(i,j)$ element of $\widehat {\boldsymbol \Sigma}$ defined in (ref). Finally, let $c^*_\Sigma(\tau)$ be the conditional $\tau$-quantile of the Gaussian bootstrap $S^*_\mathcal{D}:=\|\boldsymbol Z^*_\Sigma \|_\infty$ where $\boldsymbol Z^*_{\Sigma}|\boldsymbol X,\boldsymbol Y \sim \mathcal{N}(\boldsymbol 0,\widehat{\boldsymbol\Upsilon}_\Sigma)$, i.e.
\[c^*_\Sigma(\tau):=\inf\{q\in \mathbb{R}:\P(S^*_\mathcal{D}\leq q|\boldsymbol X,\boldsymbol Y)\geq \tau\}\]
Theorem (ref) demonstrates the validity of Gaussian bootstrap procedure described above, i.e., it states conditions under which the $\tau$-quantile of the test statistic (ref) can be approximated by $c^*_\Sigma(\tau)$ in the appropriate sense.
Similarly, the test statistic for (ref) is given by
equation[equation omitted — 163 chars of source]
Let $\boldsymbol\Upsilon_{\Pi}$ denote the $(d\times d)$ covariance matrix of $(\widetilde{\pi}_{i,j})_{(i,j)\in\mathcal{D}}$ where $\widetilde{\pi}_{i,j}:=\frac{1}{T}\sum_{t=1}^TV_{i,j,t}V_{j,i,t}$. $\boldsymbol\Upsilon_\Pi$ is estimated by
equation[equation omitted — 272 chars of source]
where $\widehat{\boldsymbol D}_{\Pi,t}$ is a $d$-dimensional vector with entries given by $\widehat{V}_{i,j,t}\widehat{V}_{j,i,t}-\widehat{\pi}_{i,j}$ for $(i,j)\in\mathcal{D}$. Also, let $c^*_\Pi(\tau)$ be the conditional $\tau$-quantile of the Gaussian bootstrap $S^*_\mathcal{D}:=\|\boldsymbol Z^*_\Pi \|_\infty$ where $\boldsymbol Z^*_\Pi|\boldsymbol X,\boldsymbol Y \sim \mathcal{N}(\boldsymbol 0,\widehat{\boldsymbol\Upsilon}_\Pi)$. Theorem (ref) establish conditions for the validity of Gaussian bootstrap to approximate the quantiles of (ref).
Theoretical Results
In this section, we collect all the theoretical guarantees for estimating the model (ref) by using the proposed multi-stage method described above. Specifically, Subsection (ref) presents non-asymptotic bounds for the (parametric) estimation, Subsection (ref) deals similar results concerning forecasting, and Subsection (ref) deals with inference on the (partial) covariance structure of $\boldsymbol U_t$.
To present the results, it is convenient to use a compact notation. For each $i\in[n]$, we define the $T$-dimensional vectors $\boldsymbol Y_{i,\cdot}:=(Y_{i,1},\ldots, Y_{i,T})'$ and $\boldsymbol U_{i,\cdot}:=(U_{i,1},\ldots, U_{i,T})'$. We also define the $(T\times k)$ matrix of covariates $\boldsymbol X_{i,\cdot}:=(\boldsymbol X_{i,1},\ldots, \boldsymbol X_{i,T})'$, for each $i\in[n]$ and the $(T\times r)$ matrix of factors $\boldsymbol F:=(\boldsymbol F_1,\dots, \boldsymbol F_T)'$, such that (ref) can be represented as
equation[equation omitted — 294 chars of source]
for each cross-sectional unit $i$, where $\boldsymbol R_{i,\cdot}:=\boldsymbol F\boldsymbol\lambda_i+ \boldsymbol U_{i,\cdot}$.
We define for each $t\in[T]$, the $n$-dimensional vectors $\boldsymbol Y_t:=(Y_{1,t},\ldots, Y_{n,t})'$ and $\boldsymbol U_t:=(U_{1,t},\ldots, U_{n,t})'$; and the $nk$-dimensional vector $\boldsymbol X_t:=(\boldsymbol X_{1,t}',\ldots, \boldsymbol X_{n,t}')'$. Also, set the $(n\times nk)$ block diagonal matrix $\boldsymbol \Gamma$ whose block diagonal is given by $(\boldsymbol\gamma_1',\dots, \boldsymbol\gamma_n')$ and the $(n\times r)$ loading matrix $\boldsymbol\Lambda:=(\boldsymbol\lambda_1,\dots,\boldsymbol\lambda_n)'$. Then, (ref) can also be represented as panel time series
equation[equation omitted — 251 chars of source]
where $\boldsymbol R_t:=\boldsymbol\Lambda \boldsymbol F_t+ \boldsymbol U_t$.
Estimation
We start by stating the following assumption.
assumption[Moments and Dependency] Consider the following:
\begin{enumerate}[(a)]
• The stochastic process $\{\boldsymbol Z_t:= (\boldsymbol X_{S,t}',\boldsymbol F_t',\boldsymbol U_{t}')':t\in [T]\}$ is weakly stationary for each $T\in\N$, where $\boldsymbol X_{S,t}$ denotes the vector $\boldsymbol X_{t}$ after excluding all deterministic (non-random) components. Furthermore, the strong mixing coefficient of $\boldsymbol Z_t$ is denoted by $\alpha_m$.
• Define $\boldsymbol{\mathcal{U}}_t:=(\boldsymbol U_{t}',\boldsymbol U_{t-1}',\dots, \boldsymbol U_{t-l}')'$ for some integer $l\geq 0$ and let $b>0$ be a finite constant such that $\lambda_{\min}\left[\mathbb{E} \left(\boldsymbol{\mathcal{U}}_t\boldsymbol{\mathcal{U}}_t'\right)\right]\geq b^2$, where $\lambda_{\min}(\cdot)$ is the minimum eigenvalue of $(\cdot)$.
• Assume there exists an universal constant $C>0$ such for all $t,s\in [T]$, $T\geq 2$ and $i\in[n]$:
\begin{enumerate}
• ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \boldsymbol Z_{t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p+\epsilon}\leq C$ for some constants $p\geq 8$ and $\epsilon>0$
• ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert n^{-1/2}\left[\boldsymbol U_s'\boldsymbol U_t - \mathbb{E}(\boldsymbol U_s'\boldsymbol U_t)\right]
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p}\leq C$
• ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert n^{-1/2}\sum_{i=1}^n\lambda_{j,i}U_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p}\leq C$
• ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \|(\boldsymbol X_i'\boldsymbol X_i/T)^{-1}\|
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p}\leq C$.
\end{enumerate}
\end{enumerate}
Assumption ((ref).a) excludes the deterministic components of $\boldsymbol X_t$ to accommodate possibly non-random non-stationary, but uniformly bounded covariates as in Assumption ((ref).c). Assumption ((ref).b) ensures that the parameter in (ref) is well defined since, by the Cauchy interlacing theorem, we have $\inf_i \lambda_{\min}\left[\mathbb{E}\left(\boldsymbol W_{i,t} \boldsymbol W_{i,t}\right)\right]\geq \lambda_{\min}\left[\mathbb{E}\left( \boldsymbol{\mathcal{U}}_t\boldsymbol{\mathcal{U}}_t'\right)\right]$. Finally, Assumptions ((ref).a) and ((ref).c) allow us to apply a Marcinkiewicz-Zygmund type inequality for partial sums to deal with polynomial tails (see Lemma (ref) in the Supplementary Material).
For each $i\in[n]$, let $\boldsymbol R_i :=\boldsymbol F\boldsymbol\lambda_i+ \boldsymbol U_i$ denote the unobservable error term in (ref), $\widehat{\boldsymbol\gamma}_i$ the least-squares estimator of $\boldsymbol \gamma_i$ and $\widehat{\boldsymbol R}_{i}:=\boldsymbol Y_t-\boldsymbol X_t \widehat{\boldsymbol\gamma}_i$ the vector of residuals. Also, set $\widehat{\boldsymbol R}:=(\widehat{\boldsymbol R}_1,\dots, \widehat{\boldsymbol R}_n)'$ and $\boldsymbol R:=(\boldsymbol R_1,\ldots, \boldsymbol R_n)'$. We must control for estimation error in the first step of the proposed methodology. The next result gives a bound for the maximum entry of the $(n\times T)$ matrix $\widehat{\boldsymbol R}-\boldsymbol R$ when the first stage is conducted by OLS in a linear setup. Note that in this case we assume that $\boldsymbol X_{i,t}$, $\boldsymbol F_t$ and $U_{i,t}$ are mutually uncorrelated.
We state the results below in terms of the strong mixing coefficient sequence, whose definition is presented here for convenience. For $m\in\{0,\ldots,T-1\}$, define
equation[equation omitted — 149 chars of source]
where $\mathcal{Z}_s^t$ is the $\sigma$-algebra generated by $(\boldsymbol Z_{s},\ldots, \boldsymbol Z_t)$ for $1\leq s\leq t\leq T$.
Note that $\alpha_m$ might depend on both $T$ and $n$.
theoremUnder Assumptions (ref) and (ref):
\[
\|\widehat{\boldsymbol R}-\boldsymbol R\|_{\max}\lesssim_\P \frac{\mathscr{R}_\alpha\sqrt{r}k^{1 + \tfrac{3}{p}} n^{4/p}}{T^{1/2-1/p}},
\]
where $\mathscr{R}_\alpha:=\mathscr{R}_\alpha(T,n):=\left[\sum_{m=0}^{T-1} (m+1)^{(p/2)-2}\alpha_m^{1-p/(p+\epsilon)}\right]^{\tfrac{2}{p}}$ and $\alpha_m$ is defined by (ref).
RemEven though we will treat the number of factors, $r$, and the number of covariates, $k$, fixed (not depending on $T$ or $n$), we kept them explicit in the result above. Furthermore, if we assume that $\alpha_{m}\leq K (m+1)^{-c}$ for all $m$, where $K$ is a constant that might depend on $n$ and $c\geq 0$ is an universal constant, then
\[\mathscr{R}_{\alpha}(T,n)\lesssim K^{\frac{2\epsilon}{p(p+\epsilon)}}\times
\begin{cases}
1 &; c>\frac{(p-2)(p+\epsilon)}{2\epsilon}\\
(\log T)^{\frac{2}{p}} &; c=\frac{(p-2)(p+\epsilon)}{2\epsilon}\\
T^{\frac{2}{p}\left[\frac{(p-2)(p+\epsilon)}{2\epsilon}+1\right]} &; c<\frac{(p-2)(p+\epsilon)}{2\epsilon}.
\end{cases}
\]
RemWhen the first step of the method involves more complicated estimation, such as Pesaran's (2006) method, instrumental variables, or LASSO, we write $\|\widehat{\boldsymbol R}-\boldsymbol R\|_{\max} \lesssim_\P \varrho_R$, where $\varrho_R$ is a non-negative sequence of $n$ and $T$. This approach is adopted systematically in the following theorems.
Define the $(n\times T)$ matrices $\boldsymbol Y := (\boldsymbol Y_1,\ldots,\boldsymbol Y_T)$ and $\boldsymbol U := (\boldsymbol U_1,\ldots,\boldsymbol U_T)$; and the $(nk\times T)$ matrix $\boldsymbol X:=(\boldsymbol X_{1},\ldots, \boldsymbol X_T)$. We can write (ref) in the matrix form as
equation[equation omitted — 137 chars of source]
Notice that $\widehat{\boldsymbol R}= \boldsymbol\Lambda\boldsymbol F' + \widetilde{\boldsymbol U}$ where $\widetilde{\boldsymbol U}:=\boldsymbol U + \widehat{\boldsymbol R}-\boldsymbol R$ and $(\boldsymbol\Lambda,\boldsymbol F)$ can be estimated by Principal Component Analysis (PCA), which minimizes $\|\widehat{\boldsymbol R} -\boldsymbol\Lambda\boldsymbol F'\|_F^2$ with respect to $\boldsymbol\Lambda$ and $\boldsymbol F$, subject to $\boldsymbol F'\boldsymbol F/T=\boldsymbol I_r$. The solution $\widehat{\boldsymbol F}$ is the matrix whose columns are $\sqrt{T}$ times $r$ eigenvectors of the top $r$ eigenvalues of $\widehat{\boldsymbol R}'\widehat{\boldsymbol R}$ and $\widehat{\boldsymbol\Lambda} = \widehat{\boldsymbol R}\widehat{\boldsymbol F}/T$.
Since we do not observe $\boldsymbol U$, in the third step of the method we use $\widehat{\boldsymbol U}:=\widehat{\boldsymbol R}-\widehat{\boldsymbol\Lambda}\widehat{\boldsymbol F}'$ instead. Therefore, we must control for the estimation error in the previous steps: $\widehat{\boldsymbol U} - \boldsymbol U$. Also, the loading matrix $\boldsymbol\Lambda$ and the factors $\boldsymbol F$ are not separably identified since $\boldsymbol\Lambda\boldsymbol F_t = \boldsymbol\Lambda\boldsymbol H'\boldsymbol H\boldsymbol F_t$ for any matrix $\boldsymbol H$ such that $\boldsymbol H'\boldsymbol H=\boldsymbol I_r$. If we let $\boldsymbol H:=T^{-1}\boldsymbol V^{-1}\widehat{\boldsymbol F}'\boldsymbol F\boldsymbol\Lambda'\boldsymbol\Lambda$, where $\boldsymbol V$ is the $(r\times r)$ diagonal matrix containing the $r$ largest eigenvalues of $\widehat{\boldsymbol R}'\widehat{\boldsymbol R}/T$ in decreasing order, we have that $\boldsymbol H\boldsymbol F_t$ is identified as $\boldsymbol\Lambda \boldsymbol F_t$ is identified.
The result below first appeared in Bai2003 for the case of $(n,T)$ diverging and was further extended to hold uniformly in $(i\leq n,t\leq T)$ by FLM2013.
However, both consider the case when the factor model is estimated using the actual data instead of the “estimated” ones (residuals) as in our case. Therefore, the next result is a generalization that takes into account the pre-estimation error term and quantifies how the error impacts the precision of factor analysis.
assumption[Factor Model] Assume:
\begin{enumerate}[(a)]
• $\mathbb{E}(\boldsymbol F_t) = \boldsymbol 0$, $\mathbb{E}(\boldsymbol F_t \boldsymbol F_t') = \boldsymbol I_r$, and $\boldsymbol\Lambda'\boldsymbol\Lambda$ is a diagonal matrix;
• All eigenvalues of $\boldsymbol\Lambda'\boldsymbol\Lambda/n$ are bounded away from zero and infinity as $n\to\infty$;
• $\|\boldsymbol\Sigma - \boldsymbol\Lambda\boldsymbol\Lambda'\|\lesssim 1$; and
• $\|\boldsymbol\Lambda\|_{\max}\lesssim 1$.
\end{enumerate}
RemAssumption (ref) is standard in the literature. Assumption $\mathbb{E}(\boldsymbol F_t)=\boldsymbol 0$ is not restrictive as our approach considers a first-step estimation which may include a constant in the set of regressors. Assumption ((ref).a) is also needed for identifiability of the factor structure. Assumption ((ref).c) imposes a strong factor structure.
theoremUnder Assumptions (ref) --(ref) , let $\varrho_R$ be a non-negative sequence of $n$ and $T$ such that $\|\widehat{\boldsymbol R}-\boldsymbol R\|_{\max} \lesssim_\P \varrho_R$. Then,
\begin{enumerate}
• $\max_{t\leq T}\|\widehat{\boldsymbol F}_t - \boldsymbol H \boldsymbol F_t\|_2\lesssim_\P \frac{1}{\sqrt{T}} + \frac{T^{1/p}}{\sqrt{n}} +\varrho_R(nT)^{1/p}$,
• $\max_{i\leq n}\|\widehat{\boldsymbol\lambda}_i - \boldsymbol H \boldsymbol\lambda_i\|_2\lesssim_\P \frac{\mathscr{R}_\alpha n^{2/p}}{\sqrt{T}} + \frac{1}{\sqrt{n}} +\varrho_R$, and
• $
\|\widehat{\boldsymbol U} - \boldsymbol U\|_{\max}\lesssim_\P \frac{\mathscr{R}_\alpha n^{2/p}}{T^{1/2-1/p}}+\frac{T^{1/p}}{\sqrt{n}} + \varrho_R(nT)^{1/p}$,
\end{enumerate}
provided that $\mathscr{R}_\alpha n^{4/p}/\sqrt{T} + (nT)^{1/p}\varrho_R \lesssim 1$, where $\mathscr{R}_\alpha$ is defined in Theorem (ref).
By setting $\varrho_R=0$ we have the case of no estimation error in the first step. Note that in order to
have the error $\|\widehat{\boldsymbol U} - \boldsymbol U\|_{\max}$ vanishing in probability we must have the pre-estimation error $\|\widehat{\boldsymbol R}-\boldsymbol R\|_{\max}$ of order (in probability) smaller than $(nT)^{-1/p}$.
We have decided not to replace $\varrho_R$ in Theorem (ref) with the rate obtained in Theorem (ref) as the latter only applies to the least-squares estimator. In some applications, however, the first step of the procedure could be done using a different type of estimator. For instance, a penalized adaptive Huber regression as in fan2017estimation if the number of features $k$ is comparable or even larger than $T$ and the tail of the distribution of $\boldsymbol{X}_t$ is heavy. By stating Theorem (ref) in terms of a generic rate, it is easier to account for the effect of different estimators. By combining Theorems (ref) and (ref) we have the following corollary.
corollaryUnder the same assumptions of Theorems (ref) and (ref), when OLS is used in the first-stage to obtain $\widehat{\boldsymbol R}$, we have
\[
\|\widehat{\boldsymbol U} - \boldsymbol U\|_{\max}\lesssim_\P \frac{\mathscr{R}_\alpha n^{5/p}}{T^{1/2-2/p}}+\frac{T^{1/p}}{\sqrt{n}}=:\varpi_U.
\]
We propose to estimate (ref) by LASSO using $\widehat{\boldsymbol U}$ in place of $\boldsymbol U$. Specifically, for a regularization parameter $\xi> 0$, we denote by $\widehat{\boldsymbol\theta}_i$ a minimizer of $\mathcal{Q}_i$ given by
equation[equation omitted — 203 chars of source]
The next result presents non-asymptotic bounds for the (in sample) prediction error and the $\ell_1$-estimation error for the LASSO estimator (ref) based on the “estimated data” and quantifies how the estimation errors impact on the choice of regularization parameter and rates of convergence.
theoremLet $\varrho_U$ be a non-negative sequence of $n$ and $T$ such that $\|\widehat{\boldsymbol U}- \boldsymbol U\|_{\max} \lesssim_\P \varrho_U$ and assume that Assumption (ref) holds. For every $\epsilon>0$ there is a constant $0<C_\epsilon<\infty$ such that if the penalty parameter is set $\xi\geq C_\epsilon\xi_0$, then for any minimizer $\widehat{\boldsymbol\theta_i}$ of (ref), with probability at least $1-\epsilon$:
\[
\max_{i\in[n]} \left[(\widehat{\boldsymbol\theta}_i - \boldsymbol\theta_i)'\mathbb{E}\left(\boldsymbol W_{i,t} \boldsymbol W_{i,t}'\right) (\widehat{\boldsymbol\theta}_{i} - \boldsymbol\theta_i)
+ \xi\|\widehat{\boldsymbol\theta}_i - \boldsymbol\theta_i\|_1\right]\leq 8 \frac{\xi^2s_0}{b^2},
\]
provided that $\frac{\xi_1 s_0}{b}\leq K_\epsilon$ where $s_k:=\max_{i\in[n]}\|\boldsymbol\theta_i\|_k$ for $k\in\{0,1,2\}$, $K_\epsilon$ is a positive constant only depending on $\epsilon$, and
\begin{align*}
\xi_0 &:=\left(1+s_2\right) \mathscr{L}_\alpha\left[n(l+1)\right]^{2/p}T^{-1/2}+\left(1+s_1\right)\left[(nT)^{1/p}\varrho_U + \varrho_U^2\right], \\
\xi_1 &:= \mathscr{L}_\alpha\left[n(l+1)\right]^{4/p}T^{-1/2}+\left[(nT)^{1/p}\varrho_U + \varrho_U^2\right]
\end{align*}
with $\mathscr{L}_\alpha:= \mathscr{L}_\alpha(T,n,l):=\left[\sum_{t=0}^{T-1}(m+1)^{(p/2)-2}\alpha_{(m-l)\lor 0}^{1-p/(p+\epsilon)}\right]^{2/p}$.
Once again, we purposely avoided replacing $\varrho_U$ in Theorem (ref) with the rate of Corollary (ref) to make it readily applicable to the case when a different type of factor model was used or, as a matter of fact, any other pre-estimation procedure. By setting $\varrho_U$ equal to $\varpi_U$, the rate of Corollary (ref), we have the next result.
corollaryIf $\varrho_U$ defined in Theorem (ref) is taken to be rate given by Corollary (ref) and $l< T$ then under the conditions of the Theorem (ref):
\[
\max_{i\in[n]}\|\widehat{\boldsymbol\theta}_i - \boldsymbol\theta_i\|_1\lesssim_\P \frac{s_0\left(1+s_1\right)}{b^2}\left[\frac{\mathscr{L}_\alpha n^{6/p}}{T^{(1/2)-(3/p)}} +\frac{T^{2/p}}{n^{(1/2)-(1/p)}}\right]=:\varpi_\theta .
\]
Forecasting
Recall that in the context of out-of-sample forecasting, our object of interest is $\widetilde{Y}_{i,t}:={\boldsymbol\gamma_i}'\boldsymbol X_{i,t}+ {\boldsymbol\lambda_i}'\boldsymbol P {\boldsymbol G}_t +{\boldsymbol\theta_i}'{\boldsymbol W}_{i,t}$ for $i\in[n]$ and $t\geq T$, which we estimate using $\widehat{Y}_{i,t}:=\widehat{\boldsymbol\gamma_i}'\boldsymbol X_{i,t}+\widehat{\boldsymbol\lambda_i}'\widehat{\boldsymbol P}\widehat{\boldsymbol G}_t +\widehat{\boldsymbol\theta_i}'\widehat{\boldsymbol W}_{i,t}$. The next result bounds (in probability) the prediction error bound in terms of all estimation errors from previous steps.
theoremUnder Assumptions (ref) and (ref) , let $\varrho_\gamma, \varrho_U, \varrho_\theta, \varrho_P,\varrho_\lambda$ and $\varrho_F$ be non-negative sequence on $n$ and $T$ such that, uniformly in $i\in[n]$, $\|\widehat{\boldsymbol\gamma}_i-\boldsymbol\gamma_i\|_{1} \lesssim_\P \varrho_\gamma$, $\|\widehat{\boldsymbol U}-\boldsymbol U\|_{\max} \lesssim_\P \varrho_U$, $\|\widehat{\boldsymbol \theta}_i-\boldsymbol\theta_i\|_{1} \lesssim_\P \varrho_\theta$, $\|\widehat{\boldsymbol P}-\boldsymbol P\|_{2} \lesssim_\P \varrho_P$, $\|\widehat{\boldsymbol \lambda}_i-\boldsymbol\lambda_i\|_{2} \lesssim_\P \varrho_\lambda$, and $\|\widehat{\boldsymbol F}_t-\boldsymbol F_t\|_{2} \lesssim_\P \varrho_F$, respectively. Then, for every $t\geq 1$,
\[
\max_{i\in[n]}\left|\widehat{Y}_{i,t}-\widetilde{Y}_{i,t}\right|\lesssim_\P (\varrho_\gamma+\varrho_\theta) n^{1/p} + \varrho_U s_1+ \varrho_P + \varrho_\lambda +\varrho_F,\]
where $s_1$ is defined in Theorem (ref).
RemIf $\varrho_\gamma, \varrho_U, \varrho_\theta,\varrho_\lambda$ and $\varrho_F$ are taken to be rates appearing in Theorems (ref)--(ref), we have $\varrho_\gamma\lor \varrho_\lambda\lor \varrho_F\lesssim \varrho_U$ and $\varrho_U\lesssim \varrho_\theta$. If further, $\boldsymbol P$ is estimated via OLS (for each $j\in[r])$ using $\{\widehat{\boldsymbol F}_t,t\in[T]\}$, we have $\varrho_P \lesssim T^{-1/2}\lor \varrho_F\lesssim \varrho_F$. Therefore, Theorem (ref) reduces to
\[
\max_{i\in[n]}\left|\widehat{Y}_{i,t}-\widetilde{Y}_{i,t}\right|\lesssim_\P \varpi_\theta\left(n^{1/p} +s_1\right),
\]
where $\varpi_\theta$ is the rate in Corollary (ref).
Inference on Covariance and Partial Covariance Matrices
We now obtain the null distributions of our test statistics for the structures of the covariance and the partial covariance. Recall the setup and notation of Section (ref). Further define $\widetilde{\boldsymbol \Sigma}:=T^{-1}\sum_{t=1}^T\boldsymbol U_t \boldsymbol U_t'$ and $\widetilde{\boldsymbol \Sigma}_\mathcal{D}$ as the element of $\widetilde{\boldsymbol \Sigma}$ indexed by $\mathcal{D}\subseteq[n]^2$.
Also, for any pair of random vectors $\boldsymbol X,\boldsymbol Y$ of the same dimensional, $d$, say; define the distance in distribution
\[
\rho(\boldsymbol X, \boldsymbol Y):=\sup_{A\in\mathcal{R}}|\P( \boldsymbol X\in A) - \P(\boldsymbol Y \in A)|,
\]
where $\mathcal{R}$ is the class of all rectangles in the from $\bigtimes_{j=1}^d (a_j,b_j]$ for some $-\infty\leq a_j\leq b_j\leq \infty$ and $j\in[d]$.
We assume that the kernel $\mathcal{K}(\cdot)$ appearing in the covariance estimator defined by (ref) belongs to the class defined in andrews91 which we reproduce below for convenience
equation[equation omitted — 196 chars of source]
This includes most of the well-known kernels used in the literature. To avoid confusion, it is worth pointing out that our tuning parameter $h$, also called bandwidth parameter by andrews91, is supposed to diverge, as opposed to the bandwidth in the density kernel estimation setup, which is expected to shrink to zero.
The following result shows how accurately the covariance matrix elements are estimated and validate the bootstrap method.
theoremFor $\mathcal{D}\subseteq [n]^2$, let $\widetilde{\boldsymbol J}:=\sqrt{T}(\widetilde{\boldsymbol\Sigma}_\mathcal{D} - \boldsymbol\Sigma_\mathcal{D})$ and $\boldsymbol{\mathcal{G}}$ be a zero-mean Gaussian vector with the same covariance matrix of $\widetilde{\boldsymbol J}$, i.e., $\boldsymbol{\mathcal{G}}\sim N(\boldsymbol{0},\boldsymbol\Upsilon_\Sigma)$.
Under Assumptions (ref)--(ref), if further
\begin{enumerate}[(a)]
• $\{\boldsymbol U_t:t\in[T]\}$ is fourth-order stationary process for each $T$;
• The strong mixing coefficient of $\{\boldsymbol U_t:t\in[T]\}$, $\alpha_m$, obeys $\alpha_m\leq K m^{-r}$ for $r>\left(\frac{p+\epsilon}{\epsilon}\right)\left(\frac{p}{4}-1\right)\lor \frac{2p}{p-4}$ where $p\geq 8$ and $\epsilon>0$ are defined in Assumption (ref); and a constant $K$ that might depend on $d$;
• the minimum eigenvalue of $\boldsymbol\Upsilon_\Sigma$ is greater or equal to $\underline{c}$, for some $\underline{c}>0$, then,
\end{enumerate}
\begin{align*}
\rho(\widetilde{\boldsymbol J},\boldsymbol{\mathcal{G}})&\lesssim \frac{\log T}{c^2}\left[ \frac{\log d}{T^{(1/2)-\kappa}} +\frac{K^{1-4/p}\log d}{T^{(r/2)(1-4/p) -1}} + \frac{(\log d)^{3/2}}{T^{1/4}} + \frac{d^{4/p}(\log d)^2}{T^{1/2-2/p}}\right] \\
&\qquad +\frac{\left[ d(\log d)^{(3/4)p-4}\log T \log (dT) \right]}{T^{1/4}c^{p/(p-4)}}^{\frac{2}{p-4}}+ \frac{d^{2/(p+2)}\sqrt{\log (dT)}+ K}{T^{r\kappa -1/2}},
\end{align*}
where $d:=|\mathcal{D}|$, $\kappa:=\kappa(p,r):=\frac{1 + D_p/2}{2(r+D_p/2)}\land 1/2$ and $D_p := \frac{p}{p+2}$.
Let $\widehat{\boldsymbol J}:=\sqrt{T}(\widehat{\boldsymbol\Sigma}_\mathcal{D} - \boldsymbol\Sigma_\mathcal{D})$, then
\[
\rho(\widehat{\boldsymbol J},\boldsymbol{\mathcal{G}})\lesssim\rho(\widetilde{\boldsymbol J},\boldsymbol{\mathcal{G}}) + \inf_{\delta>0}\left[\delta_1\sqrt{1\lor \log (d/\delta)} + \P(\|\widehat{\boldsymbol J}-\widetilde{\boldsymbol J}\|_\infty>\delta)\right]
.\]
Let $\widetilde{\boldsymbol \Upsilon}_\Sigma$ be any positive semi definite estimator of $\boldsymbol\Upsilon_\Sigma$ and $\boldsymbol{\mathcal{G}}^*|\boldsymbol X,\boldsymbol Y\sim N(\boldsymbol 0,\widetilde{\boldsymbol \Upsilon}_\Sigma)$, then
\[
\rho(\widehat{\boldsymbol J},\boldsymbol{\mathcal{G}}^*)\lesssim \rho(\widehat{\boldsymbol J},\boldsymbol{\mathcal{G}}) + \inf_{\delta>0}\left[\delta \log d(1\lor |\log d|) + \P(\|\widetilde{\boldsymbol\Upsilon}_\Sigma - \boldsymbol\Upsilon_\Sigma\|_{\max}>\delta)\right].
\]
RemThe first result in Theorem (ref) bounds the Komolgorov distance between the (unobservable) process $\left\{\frac{1}{\sqrt{T}}\sum_{t=1}^T \boldsymbol U_t \boldsymbol U_t '-\mathbb{E}\left(\boldsymbol U_t \boldsymbol U_t'\right)\right\}_{T\geq 1}$ and a Gaussian process with the same covariance structure. It is a direct consequence of a more general Central Limit Theorem result for high-dimensional alpha mixing sequences (see Theorem (ref) in the Supplementary Material). The second one is similar but controls for the difference between $\boldsymbol U_t-\widehat{\boldsymbol U}_t$ and, therefore, takes into account the estimation error. Finally, the last result ensures a bootstrap validity provided we can estimate the covariance matrix in an appropriate sense.
RemTheorem (ref) seem complicated. However, they only depend on $d$, $T$, $p$, $r$, and the “quality” of the estimators $\widehat{\boldsymbol U}$ and $\widetilde{\boldsymbol \Upsilon}$. The latter allows a different selection of estimators for any of the stages and the bootstrap procedure. If we were to specialized Theorem (ref) to incorporate the rates obtained in Theorem (ref) and (ref), and set $\widetilde{\boldsymbol\Upsilon}=\widehat{\boldsymbol \Upsilon}$ defined by (ref), we obtain a sufficient condition to ensure the bootstrap validity depending only on $n,T,r$ and $p$.
corollaryUnder the same conditions of Theorem (ref), if $\widetilde{\boldsymbol\Upsilon}_\Sigma=\widehat{\boldsymbol \Upsilon}_\Sigma$ defined by (ref) with $\mathcal{K}\in\K$, where $\mathcal{K}$ is defined by (ref), then, uniformly in $\mathcal{D}\subseteq[n]^2$,
\[
\|\widehat{\boldsymbol\Upsilon}_\Sigma - \boldsymbol\Upsilon_\Sigma\|_{\max}\lesssim_\P h\left[\varpi_U(nT)^{3/p} +n^{8/p}/\sqrt{T}\right],
\]
where $h>0$ is the bandwidth parameter of the covariance estimator and $\varpi_U$ is the rate appearing in Corollary (ref). If further, as $h, n,T\to\infty$:
\begin{enumerate}[(a)]
• $\varrho_\Sigma=o(1)$, where $\varrho_\Sigma$ is the rate appearing in the first result of Theorem (ref) with $d$ replaced by $n^2$;
• $(\log n)^{3/2}\left(\sqrt{T}\varpi_U^2+ \frac{ n^{9/p}}{\sqrt{T}}+\frac{n^{6/p}}{T^{1/2-1/p}} +\frac{1}{n^{1/2-9/p}}\right) = o(1);$
• $(\log n)^3h\left[\varpi_U(nT)^{3/p} +d^{8/p}/\sqrt{T}\right] = o(1)$, then,
\end{enumerate}
\[
\sup_{\mathcal{D}}\sup_{\tau\in(0,1)}\left|\P\left[S_\mathcal{D}^\Sigma\leq c^*_\Sigma(\tau)\right] -\tau\right|=o(1),\]
where the first supremum is over all null hypotheses of the form (ref) indexed by $\mathcal{D}\subseteq[n]^2$.
The next theorem shows how well the elements of the partial autocovariance matrix are estimated and gives the conditions under which the bootstrap test is valid. In the calculation the partial covariance (ref), we use the residual $\widehat{V}_{i,j,t}$ of the LASSO regression of $\widehat{U}_{i,t}$ onto $\widehat{\boldsymbol U}_{-ij,t}$. Define $\widetilde{\boldsymbol\Pi}_\mathcal{D}:=(\widetilde{\pi}_{i,j})_{(i,j)\in\mathcal{D}}$ for $\mathcal{D}\subseteq[n]^2$ and recall from Section (ref) that $\widetilde{\pi}_{i,j}:=\frac{1}{T}\sum_{t=1}^TV_{i,j,t}V_{j,i,t}$ for $i,j\in[n]$, $\boldsymbol\Pi_\mathcal{D}:=\mathbb{E}\left(\widetilde{\boldsymbol \Pi}_\mathcal{D}\right)$, and $\boldsymbol\Upsilon_\Pi$ denotes the covariance of $\boldsymbol\Pi_\mathcal{D}$.
theoremFor $\mathcal{D}\subseteq [n]^2$, let $\widetilde{\boldsymbol Q}:=\sqrt{T}(\widetilde{\boldsymbol\Pi}_\mathcal{D} - \boldsymbol\Pi_\mathcal{D})$ and $\boldsymbol{\mathcal{H}}$ be a zero-mean Gaussian vector with the same covariance matrix of $\widetilde{\boldsymbol Q}$, i.e., $\boldsymbol{\mathcal{H}}\sim N(\boldsymbol 0,\boldsymbol\Upsilon_\Pi)$.
Under the same assumptions and notation of Theorem (ref), with $\boldsymbol\Upsilon_\Sigma$ and $\underline{c}$ replaced by $\boldsymbol\Upsilon_\Pi$ and $\underline{b}$, respectively in condition (c), we have
\begin{align*}
\rho(\widetilde{\boldsymbol Q},\boldsymbol{\mathcal{H}})&\lesssim \frac{\log T
}{b^2}\left[ \frac{\log d}{T^{(1/2)-\kappa}} +\frac{K^{1-4/p}\log d}{T^{(r/2)(1-4/p) -1}} +\frac{(\log d)^{3/2}}{T^{1/4}} + \frac{ d^{4/p}(\log d)^2}{T^{1/2-2/p}}\right]\\
&\qquad +\frac{\left[ d(\log d)^{(3/4)p-4}\log T \log (dT) \right]}{T^{1/4}b^{p/(p-4)}}^{\frac{2}{p-4}}+ \frac{d^{1/(p/2+1)}\sqrt{\log d}+ K}{T^{r\kappa -1/2}},
\end{align*}
where $d,\kappa:=\kappa(p,r)$ are defined in Theorem (ref).
Let $\widehat{\boldsymbol Q}:=\sqrt{T}(\widehat{\boldsymbol\Pi}_\mathcal{D} - \boldsymbol\Pi_\mathcal{D})$, then
\[
\rho(\widehat{\boldsymbol Q},\boldsymbol{\mathcal{H}})\lesssim\rho(\widetilde{\boldsymbol Q},\boldsymbol{\mathcal{H}}) + \inf_{\delta>0}\left[\delta_1\sqrt{1\lor \log (d/\delta)} + \P(\|\widehat{\boldsymbol Q}-\widetilde{\boldsymbol Q}\|_\infty>\delta)\right]
.\]
Let $\widetilde{\boldsymbol \Upsilon}_\Pi$ be any positive semi definite estimator of $\boldsymbol\Upsilon_\Pi$ and $\boldsymbol{\mathcal{H}}^*|\boldsymbol X,\boldsymbol Y\sim N(\boldsymbol 0,\widetilde{\boldsymbol \Upsilon}_\Pi)$, then
\[
\rho(\widehat{\boldsymbol Q},\boldsymbol{\mathcal{H}}^*)\lesssim \rho(\widehat{\boldsymbol Q},\boldsymbol{\mathcal{H}}) + \inf_{\delta>0}\left[\delta \log d(1\lor |\log d|) + \P(\|\widetilde{\boldsymbol\Upsilon}_\Pi - \boldsymbol\Upsilon_\Pi\|_{\max}>\delta)\right].
\]
Similar comments as those appearing in Remarks (ref) apply to Theorem (ref) as well, which results in the following corollary.
corollaryUnder the same conditions of Theorem (ref), if $\widetilde{\boldsymbol\Upsilon}_\Pi=\widehat{\boldsymbol \Upsilon}_\Pi$ defined by (ref) with $\mathcal{K}\in\K$, where $\mathcal{K}$ is defined by (ref), then, uniformly in $\mathcal{D}\subset[n]^2$,
\[
\|\widehat{\boldsymbol\Upsilon}_\Pi - \boldsymbol\Upsilon_\Pi\|_{\max}\lesssim_\P h\left[((1 + \widetilde{s}_1) \varpi_U + \varrho_\chi n^{1/p})(1+\widetilde{s}_1)^3(nT)^{3/p} + \frac{(1+\widetilde{s}_2)^4n^{8/p}}{\sqrt{T}}\right] ,
\]
where $h>0$ is the bandwidth of the covariance estimator, $\widetilde{s}_k:=\max_{(i,j)\in\mathcal{D}}\|\boldsymbol\chi_{i,j}\|_k$ for $k\in\{0,1,2\}$, $\varpi_U$ is the rate appearing in Corollary (ref), and $\varrho_\chi$ is the rate appearing in Corollary (ref) with $s_0$ and $s_1$ replaced by $\widetilde{s}_0$ and $\widetilde{s}_1$, respectively, and $l=0$. If further, as $h, n,T\to\infty$:
\begin{enumerate}[(a)]
• $\varrho_\Pi=o(1)$, where $\varrho_\Pi$ is the rate in the first result of Theorem (ref) with $d$ replaced by $n^2$;
• $(\log n)^{3/2}\left[(1+\widetilde{s}_1+\varrho_\chi)^2 \left(\sqrt{T}\varpi_U^2+ \frac{ n^{9/p}}{\sqrt{T}}+\frac{n^{6/p}}{T^{1/2-1/p}} +\frac{1}{n^{1/2-9/p}}\right) + \varrho^2_\chi n^{4/p}\sqrt{T}\right] = o(1);$
• $(\log n)^3h\left\{[(1 + \widetilde{s}_1) \varpi_U + \varrho_\chi n^{1/p}](1+\widetilde{s}_1)^3(nT)^{3/p} + \frac{(1+M_2)^4n^{8/p}}{\sqrt{T}}\right\} = o(1)$, then,
\end{enumerate}
\[
\sup_{\mathcal{D}}\sup_{\tau\in(0,1)}|\P\left[S_\mathcal{D}^\Pi\leq c^*_\Pi(\tau)\right] -\tau|=o(1),\]
where the first supremum is over all null hypotheses of the form (ref) indexed by $\mathcal{D}\subseteq[n]^2$.
RemAs opposed to the case of testing covariance, when testing partial covariance in high dimensions, the sparse structure plays a role in terms of $\widetilde{s}_0$ appearing in the conditions (b) and (c). Therefore, these assumptions restrict the cases where the proposed partial covariance test has the correct asymptotic size. For instance, in the case of a complete dense partial covariance structure, we are likely to have $\widetilde{s}_0$ of the order of $n$ and, therefore, conditions $(b)$ and $(c)$ are not expected to hold.
Applications
Factor Models and Network Structure in Asset Returns
We illustrate our methodology by studying the factor structure of asset returns. We consider monthly close-to-close excess returns from a cross-section of 9,456 firms traded on the New York Stock Exchange. The data starts on November 1991 and runs until December 2018. There are 326 monthly observations in total. In addition to the returns, we also consider 16 monthly factors: Market ($\mathsf{MKT}$), Small-minus-Big ($\mathsf{SMB}$), High-minus-Low ($\mathsf{HML}$), Conservative-minus-Aggressive ($\mathsf{CMA}$), Robust-minus-Weak ($\mathsf{RMW}$), earning/price ratio, cash-flow/price ratio, dividend/price ratio, accruals, market beta, net share issues, daily variance, daily idiosyncratic variance, 1-month momentum, and 36-month momentum. The firms are grouped according to 20 industry sectors as in tMmG1999. The following sectors are considered:\footnote{The number between parenthesis indicates the number of firms in our sample that belong to each sector.} Mining (602), Food (208), Apparel (161), Paper (81), Chemical (513), Petroleum (48), Construction (68), Primary Metals (133), Fabricated Metals (186), Machinery (710), Electrical Equipment (782), Transportation Equipment (166), Manufacturing (690), Railroads (25), Other transportation (157), Utilities (411), Department Stores (67), Retail (1018), Financial (3419), and Other (11).
We start the analysis by looking at the correlation matrix for monthly returns of a sample of nine sectors: Mining, Food, Petroleum, Construction, Manufacturing, Utilities, Department Stores, Retail, and Financial. Figure (ref) plots the correlations that are larger than 0.15 in absolute value. We test for the null of a diagonal covariance matrix. The null hypothesis is strongly rejected with $p$-value much lower than 1% for all sectors. To conduct the test of the covariance matrix we use the simple sample estimator as described in the paper. However, the correlations plotted in Figure (ref) and in the subsequent ones are based on the nonlinear shrinkage estimator proposed by oLmW2020.
figure[figure omitted — 390 chars of source]
We proceed by regressing the daily returns on the observed 16 factors. Figure (ref) presents the estimated correlations for the first-stage residuals, namely factor-adjusted returns. We focus on the nine sectors as before. The first-stage regression is efficient in removing the correlation within specific sectors in some cases. The most notable ones are Financial and Retail, followed by Construction, Petroleum, and Manufacturing. On the other hand, Utilities, Department Stores, Mining, and Food still display a dense covariance matrix.
figure[figure omitted — 457 chars of source]
The second step is to conduct a principal component analysis on the residuals of the first stage. The eigenvalue ratio procedure selects two factors, while all four information criteria points to a single factor. We proceed with two factors. Note that, by construction, the principal component factors are orthogonal to all the 16 risk factors considered in the first stage. Figure (ref) shows the estimated correlations for the residuals (idiosyncratic component) of the second-stage. The latent factor is not able to reduce the correlations within each sector. However, when we consider the partial correlations the conclusions are much different. As can be seen from Figure (ref) that the partial correlation matrices are (almost) diagonal. In addition, we are not able to reject the null of a diagonal covariance matrix at a 5% significance level.
figure[figure omitted — 450 chars of source]
figure[figure omitted — 453 chars of source]
To shed some light on the links among different sectors, we report how often variables from sector $i$ are selected in the third-stage LASSO regression for firms in sector $j$. The numbers are normalized by the total number of firms in each sector and are presented in Figure (ref). The most interesting fact is that covariates from the financial sector are the ones most frequently selected for all the other sectors. Other sectors, such as Mining, Chemical, Machinery, Electrical Equipment, Manufacturing, and Retail are also frequently selected. This may indicate that there are industry factors, specifically a “financial factor”, that is unmodeled in the first two stages. However, if we augment the set of regressors in the first stage by the value-weighted portfolio from the financial sector, although the remaining dependence among firms is attenuated, particularly for Department Stores, we do not get close to an exact factor model. This finding suggests that there are hidden links among firms.
figure[figure omitted — 454 chars of source]
Forecasting US Industrial Production
The second application consists of forecasting monthly US industrial production using a large set of monthly macroeconomic variables. We compare four different models: (1) Autoregressive model; (2) Sparse LASSO Regression (SR); (3) Principal Component Regression (PCR); and (4) FarmPredict.
We use variables from the August 2022 vintage of the FRED-MD database, which is a large monthly macroeconomic dataset designed for empirical analysis in data-rich macroeconomic environments. The dataset is updated in real-time through the FRED database and is available from Michael McCraken's webpage.\footnote{https://research.stlouisfed.org/econ/mccracken/fred-databases/. For further details, we refer to mMsN2016.} .
Our sample extends from January 1960 to December 2019 (719 observations), and only variables with all observations in the period are used (122 variables). The dataset is divided into eight groups: (i) output and income; (ii) labor market; (iii) housing; (iv) consumption, orders, and inventories; (v) money and credit; (vi) interest and exchange rates; (vii) prices; and (viii) stock market. Finally, all series are transformed in order to become stationary.
In order to highlight the gains of exploring all relevant information in the dataset, we construct one-step ahead forecasts for the first-order difference of the logarithm of the monthly industrial production index (IP, growth rate): $Y_{IP,t}$.
We compare the following models:
enumerate• Autoregressive model (AR):
\[
\widehat{Y}_{IP, t+1|t}^{(\texttt{AR})} = \widehat{\phi}_{0} + \widehat{\phi}_{1}\widehat{Y}_{IP,t} + \ldots + \widehat{\phi}_{p}\widehat{Y}_{IP,t-p+1},
\]
where $\widehat{\phi}_{0},\widehat{\phi}_{1},\ldots,\widehat{\phi}_{p}$ are OLS estimates. The value of $p$ is selected by BIC.
• Sparse regression (SR):
\[
\widehat{Y}_{IP, t+1|t}^{(\texttt{SR})} = \widehat{\beta}_{0} + \widehat{\boldsymbol\beta}_{1}'\boldsymbol Y_{t} + \ldots + \widehat{\boldsymbol\beta}_{p}'\boldsymbol Y_{t-p+1},
\]
$\widehat{\beta}_{0},
\widehat{\boldsymbol\beta}_{1}\ldots,
\widehat{\boldsymbol\beta}_{p}$ are LASSO estimates and $\boldsymbol Y_t=(Y_{1t},\ldots,Y_{nt})$ with $n=122$. The penalty parameter is selected by modified BIC as in hWbLcL2009.
• Principal Component Regression (PCR):
\[
\widehat{Y}_{IP,t+1|t}^{(\texttt{PCR})} = \widehat{\pi}_{0} + \widehat{\boldsymbol\pi}_{1}'\widehat{\boldsymbol{F}}_t +\cdots+\widehat{\boldsymbol\pi}_{q}'\widehat{\boldsymbol{F}}_{t-q+1},
\]
where $\widehat{\boldsymbol F}_{t}$ is the estimate of the $(k \times 1)$ vector of factors $\boldsymbol F_t$ given by the first $k$ principal components of $\boldsymbol Y_t-\widehat{\boldsymbol\mu}$ with $\widehat{\boldsymbol\mu}$ being the sample average of $\boldsymbol Y_t$. The parameters of the model are computed by OLS regression of $Y_{j,t}$ on a constant and lags of $\widehat{\boldsymbol F}_t$. The lag $q$ is selected by BIC.
• \textbf{AR - Principal Component Regression (\texttt{AR-PCR}):}
\[
\widehat{Y}_{IP,t+1|t}^{(\texttt{AR-PCR})} = \widehat{\alpha}_{0} + \widehat{\boldsymbol\alpha}_{1}'Y_{IP,t} + \ldots + \widehat{\alpha}_{p}Y_{IP,t-p+1} + \widehat{\boldsymbol\varrho}_{1}'\widehat{\boldsymbol{F}}_t +\cdots+\widehat{\boldsymbol\varrho}_{p}'\widehat{\boldsymbol{F}}_{t-p+1},
\]
where $\widehat{\boldsymbol F}_{t}$ is the estimate of the $(k \times 1)$ vector of factors $\boldsymbol F_t$ given by the first $k$ principal components of $\boldsymbol Y_t-\widehat{\boldsymbol\mu}$ with $\widehat{\boldsymbol\mu}$ being the sample average of $\boldsymbol Y_t$. The parameters of the model are computed by OLS regression of $Y_{j,t}$ on a constant, its own lags and lags of $\widehat{\boldsymbol F}_t$. The lag orders $p$ and $q$ are selected by BIC.
• \textbf{\texttt{FarmPredict}:}
\[
\widehat{Y}_{i,t+1|t}^{(\texttt{FarmPredict})} =\widehat{\boldsymbol\mu}_{j}+
\widehat{\boldsymbol\lambda}_i'\widehat{\boldsymbol P}_{j1}'\widehat{\boldsymbol F}_t+\cdots+\widehat{\boldsymbol\lambda}_i'\widehat{\boldsymbol P}_{jp}'\widehat{\boldsymbol F}_{t-p+1}+\widehat{\boldsymbol\theta}_{1i}'\widehat{\boldsymbol U}_{t} + \ldots + \widehat{\boldsymbol\theta}_{pi}'\widehat{\boldsymbol U}_{t-p+1},
\]
where
$\widehat{\boldsymbol U}_t=\left(\widehat{U}_{1,t},\ldots,\widehat{U}_{n,t}\right)'$ and $\widehat{U}_{i,t}=Y_{i,t}-\widehat{\boldsymbol\lambda}_i'\widehat{\boldsymbol F}_t$, $i\in[n]$. The estimates $\widehat{\theta}_{0i},
\widehat{\boldsymbol\theta}_{1i}\ldots,
\widehat{\boldsymbol\theta}_{pi}$, $i\in[n]$, are given by LASSO. The penalty parameter is selected by the modified BIC and the value of $p$ is set to 24.
The forecasts are based on a rolling-window framework of a fixed length of 480 observations, starting in January 1960. Therefore, the forecasts start on January 1990. The last forecasts are for December 2019. Note that the AR model only considers information concerning the own past of the variable of interest. SR and PCR/AR-PCR/ expand the information by two opposing routes. While $\texttt{SR}$ uses a sparse combination of the set of variables, $\texttt{PCR}$ and $\texttt{AR-PCR}$ consider a factor structure (dense model). In the case of $\texttt{AR-PCR}$, lags of the dependent variable are also included. FarmPredict combines these two approaches and uses the full information available. The number of factors is set to 1.
Figure (ref) reports the ratios of cumulative MSE of the FarmPredict model against the cumulative MSE of the other benchmarks over the forecasting period. Several conclusions emerge from the plot. First, FarmPredict outperforms the PCR model over the entire out-of-sample period. It is also, in general, superior to the AR, SR, and AR-PCR models, apart from 2004 and 2008. During this period, the economy experienced housing and financial crises. Furthermore, the number of out-of-sample forecasting periods was also quite small. It is clear that the performance of FarmPredict improves drastically after 2008, and over the entire sample, the MSE ratio of the FarmPredict model over the AR benchmark is 0.9080, while the SR, PCR, and AR-PCR models have the following ratios, respectively: 0.9217, 1.0249, and 0.9215.
figure[figure omitted — 399 chars of source]
Conclusions
We propose a new methodology that bridges the gap between sparse regressions and factor models and evaluates the gains of increasing the information set via factor augmentation. Our proposal consists of several steps. In the first one, we filter the data for known factors (trends, seasonal adjustments, covariates). In the second step, we estimate a latent factor structure. Finally, in the last part of the procedure, we estimate a sparse regression for the idiosyncratic components. We also propose a new test for remaining structures in both high-dimensional covariance and partial covariance matrices. Our test can be used to evaluate the benefits of adding more structure to the model. Our paper has also a number of important side results. First, we proved the consistency of kernel estimation of long-run covariance matrices in high dimensions where both the number of observations and variables grows. Second, we derive the theoretical properties of factor estimation on the residuals of a first-step process. Third, the proposed test can be used as a diagnostic tool for factor models.
We evaluate our methodology with simulations and real data. The simulations show the test has good size and power properties even when the true number of factors is unknown and must be determined from the data. If the number of factors is underestimated, we observe size distortions. In particular, this is the case when the eigenvalue ratio test is used to determine the number of latent factors. The simulations also show that there are major informational gains when combining factor models and sparse regressions in a forecasting exercise. Two applications are considered.
acks[Acknowledgments]
Medeiros gratefully acknowledges the partial financial support from CNPq and CAPES. Fan's research is partially supported by ONR grant N00014-22-1-2340 and NSF grants DMS-2210833, DMS-2053832, and DMS-2052926. We are grateful to Caio Almeida, Matteo Barigozzi, Gilberto Boareto, Gustavo Bulhões, Giuseppe Cavaliere, Frank Diebold, Bruno Ferman, Marcelo Fernandes, Claudio Flores, Conrado Garcia, Eric Ghysels, Alexander Giessing, Nathalie Gimenes, Marcelo J. Moreira, Henrique Pires, Yuri Saporito, and Rodrigo Targino for helpful comments. We also thank seminar participants at the SofiE online seminar series, Princeton University, the University of Amsterdam, the University of Pennsylvania, the University of Illinois at Urbana-Champaign, the Federal University of São Carlos, Rutgers University, the University of North Carolina at Chapel Hill, the University of Chicago, the University of California at Riverside, and Columbia University for a number of valuable comments. Finally, we are deeply grateful to Michele Lenza, Eduardo F. Mendes, and Michael Wolf for the careful reading of the paper and the many insightful discussions which led to a much-improved version of this manuscript. This manuscript has been also presented at a number of conferences, and we thank all the participants for their very useful comments.
supplement\setcounter{section}{0}
\setcounter{table}{0}
\setcounter{figure}{0}
\setcounter{page}{1}
\stitle
\startcontents[sections]
\begin{center}
CONTENTS
\end{center}
\printcontents[sections]{l}{1}{\setcounter{tocdepth}{3}}
\section{Introduction}
The goal of this supplement is to provide additional results as well as the proofs of all theoretical results in the main body of the paper.
This Supplementary Material is organized as follows. We start by discussing guidelines for practical implementation of the method in Section (ref). Section (ref) provides some additional simulation results. Section (ref) presents results for the case with geometric mixing and Exponential Tails. Section (ref) contains all the proofs of the results on paper. Finally, Section (ref) collects auxiliary lemmas.
\section{Guide to Practice}
The methodology in this paper involves several steps. The first step consists of identifying known covariates that we may want to control for. It may involve the removal of deterministic trends and seasonal effects, for instance. This can be done either by parametric or nonparametric regressions. It is important to notice, however, that the convergence rates of the estimations in the subsequent steps will be influenced by the convergence rate of the estimation in the first part of the procedure.
After the data are filtered in the first step, one can test for the remaining covariance structure. If the covariance matrix of the filtered data is (almost) diagonal, there is no need to estimate a latent factor structure, and the practitioner may jump directly to the third step.
On the other hand, if the covariance of the filtered data is dense, a latent factor model should be considered, and the number of factors must be determined. To determine the number of factors, we consider either the eigenvalue ratio test of sAaH2013 or the information criteria put forward in jBsN2002. The factors can be estimated by the usual methods.
The next step involves a sparse regression in order to estimate any remaining links between idiosyncratic components. Before running the last step, we may test for a diagonal covariance matrix of the idiosyncratic terms. If the null is not rejected, there is no need for additional estimation. In case of rejection, we can proceed with a LASSO regression. We recommend that the penalty term is selected by some Information Criterion (IC) as advocated by hWbLcL2009 and mMeM2016. If out-of-sample forecasting is the goal, an additional final step is necessary. In this case, lags of factors and idiosyncratic terms must be determined. The practitioner may rely on the usual information criteria available.
Finally, concerning the estimation of the long-run matrices, the usual methods discussed in the literature can be used here to select the kernel and the bandwidth. We use the simple Bartlett kernel with bandwidth given as $\lfloor T/3 \rfloor$.
\section{Simulation}
In this section, we report simulation results divided into two parts. In the first one, we evaluate the finite-sample properties of the test for the remaining covariance structure. In the second part, we highlight the informational gains when considering both the common factors and the idiosyncratic component. We simulate 1,000 replications of the following model for various combinations of sample size ($T$) and number of variables ($n$):
\begin{align}
Y_{i,t}&=\boldsymbol\lambda_i'\boldsymbol{F}_t+U_{i,t},\,i=[n],\,t=[T],\\
\boldsymbol{F}_t&=0.8 \boldsymbol{F}_{t-1} + \boldsymbol E_{t},\\
U_{i,t}&= I(i=1)(\theta_{12} U_{2t} + \theta_{13} U_{3t} + \theta_{14} U_{4t} + \theta_{15} U_{5t})+ V_{i,t}
\\
V_{i,t} &=\phi V_{it-1} + \epsilon_{i,t}
\end{align}
where $I(\cdot)$ is the indicator function, $\{\epsilon_{i,t}\}$ is a sequence of independent t-distributed random variables with 10 degrees of freedom, and $\{\boldsymbol E_{t}\}$ is a sequence of $r$-dimensional mutually independent random vectors t-distributed with 10 degrees of freedom. Furthermore, $\{\epsilon_{i,t}\}$ and $\{\boldsymbol E_{t}\}$ are mutually independent for all time periods, factors and variables. For each Monte Carlo replication, the vector of loadings is sampled from a Gaussian distribution with mean -6 and standard deviation 0.2 for $i=1$ and mean 2 and unit variance for $i=2,\ldots,n$. The coefficients $\theta_{12}$, $\theta_{13}$, $\theta_{14}$, and $\theta_{15}$ are equal to zero or 0.8, 0.9, -0.7, and 0.5, respectively. We set the number of factors, $r$, equal to 3. $\phi$ can be either 0 or 0.5. Note that equation (ref) is a special case of equation (ref) with $\boldsymbol{W}_{i,t}=\boldsymbol{U}_{-i,t}$.
\subsection{Test for Remaining Covariance Structure}
We start by reporting results for the test of no remaining structure on the covariance matrix of $\boldsymbol{U}_t=(U_{1t},\ldots,U_{nt})'$. The null hypothesis considered is that all the covariances between the first variable ($i=1$) and the remaining ones are all zero. For size simulations we set $\theta_{12}=\theta_{13}=\theta_{14}=\theta_{15}=0$ in the DGP. To evaluate the effects of factor estimation as well as the methods in selecting the number of factors, we consider the following scenarios: (1) factors are known, and there is no estimation involved; (2) factors are estimated by principal components, but the number of factors is known; (3) the number of factors is determined by the eigenvalue ratio procedure of sAaH2013; (4)-(7) the number of factors is determined by one of the four information criteria proposed by jBsN2002 as defined by
\[
\begin{matrix*}[l]
\textnormal{IC}_1=\log[S(r)]+r\frac{n+T}{nT}\log\left(\frac{nT}{n+T}\right) &
\textnormal{IC}_2=\log[S(r)]+r\frac{n+T}{nT}\log C_{nT}^2\\
\textnormal{IC}_3=\log[S(r)]+r\frac{\log C_{nT}^2}{C_{nT}^2} &
\textnormal{IC}_4=\log[S(r)]+r\frac{(n+T-k)\log(nT)}{nT}.
\end{matrix*}
\]
where $S(r)=\frac{1}{nT}\|\boldsymbol R - \widehat{\boldsymbol\Lambda}_r\widehat{\boldsymbol F}_r\|_2^2$ and $C_{nT}:=\sqrt{\min(n,T)}$.
Table (ref) and (ref) reports the results of the empirical size of test for different significance levels. We consider the case of $\phi=0$ in Table (ref) and $\phi=0.5$ in (ref). The tables present the results when the factors are known in panel (a), the factors are unknown but the number of factors is known in panel (b), or the number of factors is estimated either by the information criterion $\textnormal{IC}_1$ in panel (c) or the eigenvalue ratio procedure in panel (d).
A number of facts emerge from the inspection of the results in Table (ref). First, size distortions are small when the factors are known. In this case, the test is undersized when the pair $(n,T)$ is small. When the factor is not known but the true number of factors is available, the size distortions are high only when $T=100$ and $n=50$ due to inaccurate estimation of factors. However, the distortions disappear when the pair $(T,n)$ grows. In this case, the empirical size is similar to the situation reported in Panel (a). The finite performance of the test in the case where the number of factors is selected by information criterion $\textnormal{IC}_1$ is almost indistinguishable from the case reported in Panel (b). However, the results with the eigenvalue ratio procedure are much worse when $T=100$ and $n=50$. In this case, the procedure selects fewer factors than the true number $r=3$. For instance, the procedure selects 2 or fewer factors in 36% of the replications. Just as comparison, for $T=100$ and $n=50$, $\textnormal{IC}_1$ underdetermines the number of factors only in 3.10% of the cases. The latter also confirms that overestimation of the number of factors will not have a big adversarial effect. For all the other combinations of $T$ and $n$ all the data-driven methods select the correct number of factors in almost all replications.
When the idiosyncratic components are autocorrelated the size distortions are higher, as reported in Table (ref). This is mainly caused by the well-known difficulties in the estimation of the long-run covariance matrix.
Table (ref) report the results of the empirical power with $\phi=0$, $\beta_{12}=0.8$, $\beta_{13}=0.9$, $\beta_{14}=-0.7$, and $\beta_{15}=-0.5$ in the DGP. When the factors are known, the test always rejects the null and the empirical power is one for any significance level. On the other hand, when factors must be estimated but the number of factors is known, the power decreases, as depicted in panel (b) in the table. Nevertheless, for $T=500,700$ the power is reasonably high, especially when the test is conducted at a $10\%$ significance level. For $T=100$, the performance deteriorates as $n$ grows. The results are similar when data-driven procedures are used to determine the number of factors, and the conclusions are mostly the same.
Table (ref) report power results in a similar setting as above but with $\phi=0.5$. The above conclusions are mostly the same if $\phi=0$ or $\phi=0.5$.
The main message of the simulation exercise is that the finite-sample performance of the proposed tests depends on the correct selection of factors. Nevertheless, for the DGP considered here, the usual data-driven methods available in the literature to determine the true number of factors seem to work reasonably well.
\subsection{Informational Gains}
The goal of this simulation is to compare, in a prediction environment, the three-stage method developed in the paper by evaluating the information gains in predicting $Y_{1t}$ by three different methods. First, the predictions are computed from a LASSO regression of $Y_{1t}$ on all the other $n-1$ variables. This is the Sparse Regression (SR) approach. Second, we consider a principal component regression (PCR), i.e., an ordinary least squares (OLS) regression of the variable of interest on factors computed from the pool of other variables. Finally, we consider predictions constructed from the method proposed here, the FarmPredict methodology. Table (ref) presents the results. The table presents the average mean squared error (MSE) over 5-fold cross-validation (CV) subsamples. As in the size and power simulations, we consider different combinations of $T$ and $n$. We report results for the case where $\theta_{12}=0.8$, $\theta_{13}=0.9$, $\theta_{14}=-0.7$, and $\theta_{15}=-0.5$ in the DGP.
According to the DGP, the theoretical MSE is 0.25 when all the information is used. When just a factor is used, the MSE is 2.21. From the table is clear that there are significant informational gains when we consider both factors and the cross-dependence between idiosyncratic components. Several conclusions emerge from the table. First, it is clear that when the sample size increases the MSE reduces. This is expected. Second, the PCR's MSE and FarmPredict's MSE are close to their theoretical values of 2.21 and 0.25 when the sample increases. The performance of the FarmPredict is quite remarkable when $T=500$ or $T=700$ and is always superior to Sparse Regression and PCR.
\begin{table}[tb]
\caption{Simulation Results: Size with $\phi=0$.}
\begin{minipage}{0.9\linewidth}
\begin{footnotesize}
The table reports the empirical size of the test of the remaining covariance structure. Panel (a) reports the case where the factors are known, whereas Panel (b) considers that the factors are unknown but the number of factors is known. Panels (c) and (d) present the results when the number of factors is determined, respectively, by the eigenvalue ratio test and the information criterion $IC_1$. Factors are estimated by the usual principal component algorithm. Three nominal significance levels are considered: 0.01, 0.05, and 0.10. The table reports the results for the case where $\phi=0$ in (ref).
\end{footnotesize}
\end{minipage}
\resizebox{0.9\linewidth}{!}{
\begin{threeparttable}
\begin{tabular}{lllllllllllll}
\hline
&&\multicolumn{11}{c}{Panel(a): Known factors}\\
&& \multicolumn{3}{c}{\underline{$T=100$}} && \multicolumn{3}{c}{\underline{$T=500$}} && \multicolumn{3}{c}{\underline{$T=700$}} \\
&& 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 \\
\hline
$n=0.5\times T$ && 0.08 & 0.03 & 0.01 && 0.10 & 0.05 & 0.01 && 0.09 & 0.04 & 0.01\\
$n=1\times T$ && 0.06 & 0.02 & 0.00 && 0.07 & 0.03 & 0.01 && 0.10 & 0.05 & 0.01\\
$n=2\times T$ && 0.07 & 0.02 & 0.00 && 0.07 & 0.02 & 0.00 && 0.08 & 0.04 & 0.00\\
$n=3\times T$ && 0.05 & 0.01 & 0.00 && 0.08 & 0.04 & 0.01 && 0.07 & 0.04 & 0.01\\
\\
&&\multicolumn{11}{c}{\underline{\textbf{Panel(b): Known number of factors}}}\\
&& \multicolumn{3}{c}{\underline{$T=100$}} && \multicolumn{3}{c}{\underline{$T=500$}} && \multicolumn{3}{c}{\underline{$T=700$}} \\
&& 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 \\
\hline
$n=0.5\times T$ && 0.23 & 0.13 & 0.02 && 0.14 & 0.06 & 0.02 && 0.11 & 0.05 & 0.01\\
$n=1\times T$ && 0.13 & 0.06 & 0.01 && 0.09 & 0.04 & 0.01 && 0.12 & 0.05 & 0.01\\
$n=2\times T$ && 0.09 & 0.04 & 0.01 && 0.07 & 0.04 & 0.01 && 0.09 & 0.04 & 0.00\\
$n=3\times T$ && 0.06 & 0.02 & 0.00 && 0.07 & 0.04 & 0.01 && 0.07 & 0.03 & 0.01\\
\\
&&\multicolumn{11}{c}{\underline{\textbf{Panel(c): Information criterion ($\textnormal{IC}_1$)}}}\\
&& \multicolumn{3}{c}{\underline{$T=100$}} && \multicolumn{3}{c}{\underline{$T=500$}} && \multicolumn{3}{c}{\underline{$T=700$}} \\
&& 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 \\
\hline
$n=0.5\times T$ && 0.23 & 0.12 & 0.02 && 0.13 & 0.06 & 0.02 && 0.10 & 0.06 & 0.01 \\
$n=1\times T$ && 0.13 & 0.08 & 0.02 && 0.10 & 0.03 & 0.01 && 0.11 & 0.05 & 0.01 \\
$n=2\times T$ && 0.11 & 0.05 & 0.01 && 0.07 & 0.04 & 0.01 && 0.10 & 0.05 & 0.01 \\
$n=3\times T$ && 0.08 & 0.03 & 0.01 && 0.07 & 0.03 & 0.01 && 0.07 & 0.03 & 0.01 \\
\\
&&\multicolumn{11}{c}{\underline{\textbf{Panel(d): Eigenvalue ratio}}}\\
&& \multicolumn{3}{c}{\underline{$T=100$}} && \multicolumn{3}{c}{\underline{$T=500$}} && \multicolumn{3}{c}{\underline{$T=700$}} \\
&& 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 \\
\hline
$n=0.5\times T$ && 0.46 & 0.35 & 0.23 && 0.13 & 0.06 & 0.02 && 0.10 & 0.05 & 0.01 \\
$n=1\times T$ && 0.12 & 0.06 & 0.02 && 0.08 & 0.04 & 0.01 && 0.12 & 0.05 & 0.01 \\
$n=2\times T$ && 0.09 & 0.04 & 0.01 && 0.08 & 0.04 & 0.01 && 0.09 & 0.04 & 0.00 \\
$n=3\times T$ && 0.06 & 0.02 & 0.00 && 0.08 & 0.04 & 0.01 && 0.06 & 0.03 & 0.01 \\
\hline
\end{tabular}
\end{threeparttable}}
\end{table}
\begin{table}[htbp]
\caption{\textbf{Simulation Results: Size with $\phi=0.5$.}}
\begin{minipage}{0.9\linewidth}
\begin{footnotesize}
The table reports the empirical size of the test of the remaining covariance structure. Panel (a) reports the case where the factors are known, whereas Panel (b) considers that the factors are unknown but the number of factors is known. Panels (c) and (d) present the results when the number of factors is determined, respectively, by the eigenvalue ratio test and the information criterion $IC_1$. Factors are estimated by the usual principal component algorithm. Three nominal significance levels are considered: 0.01, 0.05, and 0.10. The table reports the results for the case where $\phi=0.5$ in (ref).
\end{footnotesize}
\end{minipage}
\resizebox{0.9\linewidth}{!}{
\begin{threeparttable}
\begin{tabular}{lllllllllllll}
\hline
&&\multicolumn{11}{c}{\underline{\textbf{Panel(a): Known factors}}}\\
&& \multicolumn{3}{c}{\underline{$T=100$}} && \multicolumn{3}{c}{\underline{$T=500$}} && \multicolumn{3}{c}{\underline{$T=700$}} \\
&& 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 \\
\hline
$n=0.5\times T$ && 0.09 &0.04 &0.01 &&0.12 &0.07 &0.01 &&0.10 &0.05 &0.01\\
$n=1\times T$ && 0.07 &0.03 &0.00 &&0.07 &0.03 &0.01 &&0.11 &0.06 &0.01\\
$n=2\times T$ && 0.08 &0.02 &0.00 &&0.08 &0.03 &0.00 &&0.09 &0.05 &0.00\\
$n=3\times T$ && 0.05 &0.02 &0.00 &&0.09 &0.04 &0.01 &&0.07 &0.04 &0.01\\
\\
&&\multicolumn{11}{c}{\underline{\textbf{Panel(b): Known number of factors}}}\\
&& \multicolumn{3}{c}{\underline{$T=100$}} && \multicolumn{3}{c}{\underline{$T=500$}} && \multicolumn{3}{c}{\underline{$T=700$}} \\
&& 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 \\
\hline
$n=0.5\times T$ && 0.25 &0.15 &0.3 &&0.15 &0.07 &0.02 &&0.12 &0.06 &0.01\\
$n=1\times T$ && 0.13 &0.07 &0.01 &&0.09 &0.04 &0.01 &&0.14 &0.06 &0.02\\
$n=2\times T$ && 0.09 &0.04 &0.01 &&0.08 &0.04 &0.01 &&0.09 &0.05 &0.00\\
$n=3\times T$ && 0.08 &0.02 &0.00 &&0.08 &0.04 &0.01 &&0.08 &0.03 &0.01\\
\\
&&\multicolumn{11}{c}{\underline{\textbf{Panel(c): Information criterion ($\textnormal{IC}_1$)}}}\\
&& \multicolumn{3}{c}{\underline{$T=100$}} && \multicolumn{3}{c}{\underline{$T=500$}} && \multicolumn{3}{c}{\underline{$T=700$}} \\
&& 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 \\
\hline
$n=0.5\times T$ &&0.45 &0.41 &0.29 &&0.15 &0.07 &0.02 &&0.11 &0.06 &0.01\\
$n=1\times T$ &&0.15 &0.09 &0.02 &&0.10 &0.04 &0.01 &&0.14 &0.06 &0.01\\
$n=2\times T$ &&0.09 &0.04 &0.01 &&0.09 &0.04 &0.01 &&0.09 &0.05 &0.00\\
$n=3\times T$ &&0.07 &0.03 &0.00 &&0.10 &0.04 &0.01 &&0.08 &0.03 &0.01\\
\\
&&\multicolumn{11}{c}{\underline{\textbf{Panel(d): Eigenvalue ratio}}}\\
&& \multicolumn{3}{c}{\underline{$T=100$}} && \multicolumn{3}{c}{\underline{$T=500$}} && \multicolumn{3}{c}{\underline{$T=700$}} \\
&& 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 \\
\hline
$n=0.5\times T$ && 0.25 &0.14 &0.04 &&0.14 &0.07 &0.02 &&0.13 &0.06 &0.01\\
$n=1\times T$ && 0.15 &0.07 &0.02 &&0.10 &0.04 &0.01 &&0.13 &0.06 &0.02\\
$n=2\times T$ && 0.11 &0.05 &0.01 &&0.08 &0.05 &0.01 &&0.10 &0.05 &0.00\\
$n=3\times T$ && 0.08 &0.03 &0.01 &&0.09 &0.04 &0.01 &&0.08 &0.03 &0.01\\
\hline
\end{tabular}
\end{threeparttable}}
\end{table}
\begin{table}[tb]
\caption{\textbf{Simulation Results: Power ($\phi=0$).}}
\begin{minipage}{0.9\linewidth}
\begin{footnotesize}
The table reports the empirical power of the test of the remaining covariance structure. Panel (a) reports the case where the factors are known, whereas Panel (b) considers that the factors are unknown but the number of factors is known. Factors are estimated by the usual principal component algorithm. Three nominal significance levels are considered: 0.01, 0.05, and 0.10.
\end{footnotesize}
\end{minipage}
\resizebox{0.9\linewidth}{!}{
\begin{threeparttable}
\begin{tabular}{lllllllllllll}
\hline
&&\multicolumn{11}{c}{\underline{\textbf{Panel(a): Known factors}}}\\
&& \multicolumn{3}{c}{\underline{$T=100$}} && \multicolumn{3}{c}{\underline{$T=500$}} && \multicolumn{3}{c}{\underline{$T=700$}} \\
&& 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 \\
\hline
$n=0.5\times T$ && 1 & 1 & 1 && 1 & 1 & 1 && 1 & 1 & 1\\
$n=1\times T$ && 1 & 1 & 1 && 1 & 1 & 1 && 1 & 1 & 1\\
$n=2\times T$ && 1 & 1 & 1 && 1 & 1 & 1 && 1 & 1 & 1\\
$n=3\times T$ && 1 & 1 & 1 && 1 & 1 & 1 && 1 & 1 & 1\\
\\
&&\multicolumn{11}{c}{\underline{\textbf{Panel(b): Known number of factors}}}\\
&& \multicolumn{3}{c}{\underline{$T=100$}} && \multicolumn{3}{c}{\underline{$T=500$}} && \multicolumn{3}{c}{\underline{$T=700$}} \\
&& 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 \\
\hline
$n=0.5\times T$ && 0.35 & 0.19 & 0.03 && 0.99 & 0.98 & 0.83 & 0.99 & 0.99 & 0.95\\
$n=1\times T$ && 0.20 & 0.08 & 0.01 && 0.82 & 0.60 & 0.11 & 0.95 & 0.81 & 0.32\\
$n=2\times T$ && 0.15 & 0.07 & 0.01 && 0.82 & 0.55 & 0.11 & 0.94 & 0.82 & 0.34\\
$n=3\times T$ && 0.09 & 0.03 & 0.00 && 0.79 & 0.52 & 0.10 & 0.94 & 0.80 & 0.31\\
&&\multicolumn{11}{c}{\underline{\textbf{Panel(c): Eigenvalue ratio}}}\\
&& \multicolumn{3}{c}{\underline{$T=100$}} && \multicolumn{3}{c}{\underline{$T=500$}} && \multicolumn{3}{c}{\underline{$T=700$}} \\
&& 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 \\
\hline
$n=0.5\times T$ && 0.14 &0.09 &0.02 &&0.99 &0.97 &0.83&& 0.99 &0.99 &0.94 \\
$n=1\times T$ && 0.18 &0.07 &0.01 &&0.84 &0.59 &0.13&& 0.95 &0.81 &0.33 \\
$n=2\times T$ && 0.17 &0.07 &0.01 &&0.83 &0.54 &0.11&& 0.94 &0.82 &0.34 \\
$n=3\times T$ && 0.08 &0.03 &0.00 &&0.82 &0.53 &0.10&& 0.95 &0.81 &0.34 \\
\\
&&\multicolumn{11}{c}{\underline{\textbf{Panel(d): Information criterion ($\textnormal{IC}_1$)}}}\\
&& \multicolumn{3}{c}{\underline{$T=100$}} && \multicolumn{3}{c}{\underline{$T=500$}} && \multicolumn{3}{c}{\underline{$T=700$}} \\
&& 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 \\
\hline
$n=0.5\times T$ && 0.15 &0.09 &0.02 &&0.99 &0.97 &0.83&& 0.99 &0.99 &0.94 \\
$n=1\times T$ && 0.20 &0.07 &0.01 &&0.84 &0.60 &0.13&& 0.95 &0.81 &0.33 \\
$n=2\times T$ && 0.17 &0.07 &0.01 &&0.83 &0.55 &0.11&& 0.94 &0.82 &0.34 \\
$n=3\times T$ && 0.09 &0.03 &0.00 &&0.82 &0.56 &0.10&& 0.95 &0.81 &0.34 \\
\hline
\end{tabular}
\end{threeparttable}}
\end{table}
\begin{table}[htbp]
\caption{\textbf{Simulation Results: Power ($\phi=0.5$).}}
\begin{minipage}{0.9\linewidth}
\begin{footnotesize}
The table reports the empirical power of the test of the remaining covariance structure. Panel (a) reports the case where the factors are known, whereas Panel (b) considers that the factors are unknown but the number of factors is known. Factors are estimated by the usual principal component algorithm. Three nominal significance levels are considered: 0.01, 0.05, and 0.10.
\end{footnotesize}
\end{minipage}
\resizebox{0.9\linewidth}{!}{
\begin{threeparttable}
\begin{tabular}{lllllllllllll}
\hline
&&\multicolumn{11}{c}{\underline{\textbf{Panel(a): Known factors}}}\\
&& \multicolumn{3}{c}{\underline{$T=100$}} && \multicolumn{3}{c}{\underline{$T=500$}} && \multicolumn{3}{c}{\underline{$T=700$}} \\
&& 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 \\
\hline
$n=0.5\times T$ && 1 & 1 & 1 && 1 & 1 & 1 && 1 & 1 & 1\\
$n=1\times T$ && 1 & 1 & 1 && 1 & 1 & 1 && 1 & 1 & 1\\
$n=2\times T$ && 1 & 1 & 1 && 1 & 1 & 1 && 1 & 1 & 1\\
$n=3\times T$ && 1 & 1 & 1 && 1 & 1 & 1 && 1 & 1 & 1\\
\\
&&\multicolumn{11}{c}{\underline{\textbf{Panel(b): Known number of factors}}}\\
&& \multicolumn{3}{c}{\underline{$T=100$}} && \multicolumn{3}{c}{\underline{$T=500$}} && \multicolumn{3}{c}{\underline{$T=700$}} \\
&& 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 \\
\hline
$n=0.5\times T$ && 0.38 & 0.18 & 0.03 && 1.00 & 1.00 & 0.91 && 1.00 & 1.00 & 1.00\\
$n=1\times T$ && 0.21 & 0.09 & 0.02 && 0.89 & 0.69 & 0.13 && 1.00 & 0.92 & 0.39\\
$n=2\times T$ && 0.18 & 0.07 & 0.01 && 0.98 & 0.59 & 0.13 && 1.00 & 0.96 & 0.36\\
$n=3\times T$ && 0.10 & 0.03 & 0.00 && 0.91 & 0.66 & 0.11 && 1.00 & 0.92 & 0.40\\
&&\multicolumn{11}{c}{\underline{\textbf{Panel(c): Eigenvalue ratio}}}\\
&& \multicolumn{3}{c}{\underline{$T=100$}} && \multicolumn{3}{c}{\underline{$T=500$}} && \multicolumn{3}{c}{\underline{$T=700$}} \\
&& 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 \\
\hline
$n=0.5\times T$ && 0.15 & 0.11 &0.03 &&1.00 &1.00 &0.96 &&1.00 &1.00 &1.00\\
$n=1\times T$ && 0.22 & 0.08 &0.01 &&0.99 &0.70 &0.16 &&1.00 &0.95 &0.38\\
$n=2\times T$ && 0.17 & 0.07 &0.01 &&0.94 &0.62 &0.12 &&0.98 &0.88 &0.37\\
$n=3\times T$ && 0.09 & 0.03 &0.00 &&0.89 &0.59 &0.11 &&1.00 &0.87 &0.40\\
&&\multicolumn{11}{c}{\underline{\textbf{Panel(d): Information criterion ($\textnormal{IC}_1$)}}}\\
&& \multicolumn{3}{c}{\underline{$T=100$}} && \multicolumn{3}{c}{\underline{$T=500$}} && \multicolumn{3}{c}{\underline{$T=700$}} \\
&& 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 && 0.10 & 0.05 & 0.01 \\
\hline
$n=0.5\times T$ && 0.15 & 0.11 &0.03 &&1.00 &1.00 &0.96 &&1.00 &1.00 &1.00\\
$n=1\times T$ && 0.22 & 0.08 &0.01 &&0.99 &0.70 &0.16 &&1.00 &0.95 &0.38\\
$n=2\times T$ && 0.17 & 0.07 &0.01 &&0.94 &0.62 &0.12 &&0.98 &0.88 &0.37\\
$n=3\times T$ && 0.09 & 0.03 &0.00 &&0.89 &0.59 &0.11 &&1.00 &0.87 &0.40\\
\hline
\end{tabular}
\end{threeparttable}}
\end{table}
\begin{table}[tb]
\caption{\textbf{Simulation Results: Informational Gains}}
\begin{minipage}{0.9\linewidth}
\begin{footnotesize}
The table reports the average mean squared error (MSE) of three different prediction models over 5-fold cross-validation subsamples. The goal is to predict the first variable using information from the remaining $n-1$. Panel (a) considers the case of Sparse Regression (SR) where $Y_{1t}$ is LASSO-regressed on all the other variables. Panel (b) shows the results of Principal Component Regression (PCR). Finally, Panel (c) presents the results of \texttt{FarmPredict}. “N/A” means “not available”. Note that there is no factor selection for Sparse Regression. “Known Number” means that the number of factors is known.
\end{footnotesize}
\end{minipage}
\resizebox{0.9\linewidth}{!}{
\begin{threeparttable}
\begin{tabular}{lllllllllllll}
\hline
&&\multicolumn{11}{c}{\underline{\textbf{Panel(a): Sparse Regression (SR)}}}\\
&& \multicolumn{3}{c}{\underline{Known Number}} && \multicolumn{3}{c}{\underline{Eigenvalue Ratio}} && \multicolumn{3}{c}{\underline{Information Criterion ($\textnormal{IC}_1$)}} \\
&& $T=100$ & $T=500$ & $T=700$ && $T=100$ & $T=500$ & $T=700$ && $T=100$ & $T=500$ & $T=700$ \\
\hline
$n=0.5\times T$ && 0.60 & 0.35 & 0.34 && N/A & N/A& N/A&& N/A& N/A& N/A\\
$n=1\times T$ && 0.42 & 0.38 & 0.32 && N/A & N/A& N/A&& N/A& N/A& N/A\\
$n=2\times T$ && 0.40 & 0.35 & 0.31 && N/A & N/A& N/A&& N/A& N/A& N/A\\
$n=3\times T$ && 0.40 & 0.35 & 0.30 && N/A & N/A& N/A&& N/A& N/A& N/A\\
\\
&&\multicolumn{11}{c}{\underline{\textbf{Panel(b): Principal Component Regression (PCR)}}}\\
&& \multicolumn{3}{c}{\underline{Known Number}} && \multicolumn{3}{c}{\underline{Eigenvalue Ratio}} && \multicolumn{3}{c}{\underline{Information Criterion ($\textnormal{IC}_1$)}} \\
&& $T=100$ & $T=500$ & $T=700$ && $T=100$ & $T=500$ & $T=700$ && $T=100$ & $T=500$ & $T=700$ \\
\hline
$n=0.5\times T$ &&3.82 &3.12 &3.01 &&4.69 &3.12 &3.01 &&3.26 &3.04 &2.34\\
$n=1\times T$ &&3.09 &2.35 &2.34 &&4.05 &3.35 &3.34 &&3.22 &3.02 &2.32\\
$n=2\times T$ &&3.14 &2.97 &2.21 &&4.13 &3.97 &2.21 &&3.29 &3.21 &2.27\\
$n=3\times T$ &&3.83 &3.00 &2.33 &&3.83 &3.00 &2.33 &&3.12 &3.00 &2.28\\
&&\multicolumn{11}{c}{\underline{\textbf{Panel(c): FarmPredict}}}\\
&& \multicolumn{3}{c}{\underline{Known Number}} && \multicolumn{3}{c}{\underline{Eigenvalue Ratio}} && \multicolumn{3}{c}{\underline{Information Criterion ($\textnormal{IC}_1$)}} \\
&& $T=100$ & $T=500$ & $T=700$ && $T=100$ & $T=500$ & $T=700$ && $T=100$ & $T=500$ & $T=700$ \\
\hline
$n=0.5\times T$ && 0.50 &0.33 &0.31 &&0.52 &0.33 &0.31 &&0.50& 0.34 &0.30 \\
$n=1\times T$ && 0.32 &0.29 &0.28 &&0.37 &0.29 &0.28 &&0.53& 0.28 &0.27 \\
$n=2\times T$ && 0.27 &0.27 &0.26 &&0.28 &0.27 &0.26 &&0.32& 0.28 &0.28 \\
$n=3\times T$ && 0.22 &0.21 &0.21 &&0.22 &0.21 &0.21 &&0.34& 0.27 &0.27 \\
\hline
\end{tabular}
\end{threeparttable}}
\end{table}
\section{Results for the case with Geometric Mixing and Exponential Tails}
This section extends the results from the main text (Theorems (ref)-(ref)) to the case where the random quantities admit an exponential tail and a strong mixing coefficient with exponential decay (see Assumption (ref)(a)-(c) for a precise statement).
Before we begin, we need some additional pieces of notation. For any convex function $\psi:\mathbb{R}^+\to\mathbb{R}^+ $ such that $\psi(0)=0$ and $\psi(x)\to\infty$ as $x\to\infty$ and (real-valued) random variable $X$, we denote its Orlicz-norm by ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert x
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_\psi$, which is defined by ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_\psi:=\inf \left\{C>0:\mathbb{E}\left[\psi\left(\frac{|X|}{C}\right)\right]\leq 1\right\}$. In particular, we have the $\ell^p$ Orlicz-norm of $X$ by ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_p$ for $p\in[0,\infty)$ by setting $\psi(x)=x^p$ and ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}$ the exponential Orlicz-norm for $\gamma>0$ by setting $\psi(x)=\exp(x^\gamma)-1$ for $\gamma\geq 1$ and $\psi(x)$ is the convex hull of $x\mapsto \exp(x^\gamma)-1$ for $\gamma \in(0,1)$ (to ensure convexity). Also, when $\boldsymbol X$ is a random vector, we define its Orlicz-norm by ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \boldsymbol X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_\psi:=\sup_{\|\boldsymbol{u}\|\leq 1}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \boldsymbol{u}'\boldsymbol X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_\psi$.
We strengthen Assumption ((ref).c) by adding the following condition. Recall that $\{\alpha_{m}\}_m$ denotes the strong mixing coefficients $\{\boldsymbol Z_t\}_t$ as defined in (ref).
\begin{assumption}[\textbf{Moments and Dependency: Exponential Case}] . There are universal constants $C, K_1,\gamma_1,\gamma_2>0$ such for all $t,s\in [T]$, $T\geq 2$ and $i\in[n]$,
\begin{enumerate}[(a)]
• ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \boldsymbol Z_{t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma_2}}\leq C$
• $\alpha_{m}\leq \exp(-K_1m^{\gamma_1})$ for $1\leq m<T$.
• $\gamma < 1$ where $\gamma$ is defined by $1/\gamma=1/\gamma_1+2/\gamma_2$
• ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert n^{-1/2}\left[\boldsymbol U_s'\boldsymbol U_t - \mathbb{E}(\boldsymbol U_s'\boldsymbol U_t)\right]
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma_2}}\leq C$
• ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert n^{-1/2}\sum_{i=1}^n\lambda_{j,i}U_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma_2}}\leq C$
• ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \|(\boldsymbol X_i'\boldsymbol X_i/T)^{-1}\|
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma_2}}\leq C$
• $\frac{(\log n)^{(2/\gamma) - 1}}{T}\leq \frac{1}{C_1}$ where $C_1$ is positive constant only depending on $C$, $K_1$, $\gamma$ and $\gamma_1$.
\end{enumerate}
\end{assumption}
\begin{theorem} Under Assumptions (ref),(ref), and (ref):
\[
\|\widehat{\boldsymbol R}-\boldsymbol R\|_{\max}\lesssim_\P \frac{\sqrt{r}k[\log(n)\log(nTk)]^{1/\gamma_2} \sqrt{\log(nk)}}{\sqrt{T}}.
\]
\end{theorem}
\begin{theorem} Under Assumptions (ref) --(ref) , let $\varrho_R$ be a non-negative sequence of $n$ and $T$ such that $\|\widehat{\boldsymbol R}-\boldsymbol R\|_{\max} \lesssim_\P \varrho_R$. Then
\begin{enumerate}[(a)]
• $\max_{t\leq T}\|\widehat{\boldsymbol F}_t - \boldsymbol H \boldsymbol F_t\|_2\lesssim_\P \frac{1}{\sqrt{T}} + \frac{[\log T]^{1/\gamma_2}}{\sqrt{n}} + \varrho_R[\log (nT)]^{1/\gamma_2}$
• $\max_{i\leq n}\|\widehat{\boldsymbol\lambda}_i - \boldsymbol H \boldsymbol\lambda_i\|_2\lesssim_\P \sqrt{\frac{\log n}{T}} + \frac{1}{\sqrt{n}} + \varrho_R$
• $\|\widehat{\boldsymbol U} - \boldsymbol U\|_{\max}\lesssim_\P (\log T)^{1/\gamma_2}\sqrt{\frac{\log n}{T}}+\frac{(\log T)^{1/\gamma_2}}{\sqrt{n}} + \varrho_R[\log (nT)]^{1/\gamma_2}$,
\end{enumerate}
provided that $\sqrt{\log n/T} + [\log(nT) \varrho_R]^{1/\gamma_2} \lesssim 1$.
\end{theorem}
\begin{theorem} Let $\varrho_U$ be a non-negative sequence of $n$ and $T$ such that $\|\widehat{\boldsymbol U}- \boldsymbol U\|_{\max} \lesssim_\P \varrho_U$ and assume that Assumptions (ref) and (ref) hold. For every $\epsilon>0$ there is a constant $0<C_\epsilon<\infty$ such that if the penalty parameter is set $\xi\geq C_\epsilon\xi_0$, then for any minimizer $\widehat{\boldsymbol\theta_i}$ of (ref), with probability at least $1-\epsilon$:
\[
\max_{i\in[n]} \left[(\widehat{\boldsymbol\theta}_i - \boldsymbol\theta_i)'\mathbb{E}\left(\boldsymbol W_{i,t} \boldsymbol W_{i,t}'\right) (\widehat{\boldsymbol\theta}_{i} - \boldsymbol\theta_i)
+ \xi\|\widehat{\boldsymbol\theta}_i - \boldsymbol\theta_i\|_1\right]\leq 8 \frac{\xi^2s_0}{b^2},
\]
provided that $\frac{\xi_1 s_0}{b}\leq K_\epsilon$ where $s_k:=\max_{i\in[n]}\|\boldsymbol\theta_i\|_k$ for $k\in\{0,1,2\}$, $K_\epsilon$ is a positive constant only depending on $\epsilon$, and
\begin{align*}
\xi_0 &:=\left(1+s_2\right)\sqrt{\frac{\log[n(l+1)]}{T}}+\left(1+s_1\right)\left[[\log(nT)]^{1/\gamma_2}\varrho_U + \varrho_U^2\right], \\
\xi_1 &:= \sqrt{\frac{\log [n(l+1)]}{T}}+\left[[\log(nT)]^{1/\gamma_2}\varrho_U + \varrho_U^2\right].
\end{align*}
\end{theorem}
\begin{theorem} Under Assumptions (ref), (ref), and (ref) , let $\varrho_\gamma, \varrho_U, \varrho_\theta, \varrho_P,\varrho_\lambda$ and $\varrho_F$ be non-negative sequence on $n$ and $T$ such that, uniformly in $i\in[n]$, $\|\widehat{\boldsymbol\gamma}_i-\boldsymbol\gamma_i\|_{1} \lesssim_\P \varrho_\gamma$, $\|\widehat{\boldsymbol U}-\boldsymbol U\|_{\max} \lesssim_\P \varrho_U$, $\|\widehat{\boldsymbol \theta}_i-\boldsymbol\theta_i\|_{1} \lesssim_\P \varrho_\theta$, $\|\widehat{\boldsymbol P}-\boldsymbol P\|_{2} \lesssim_\P \varrho_P$, $\|\widehat{\boldsymbol \lambda}_i-\boldsymbol\lambda_i\|_{2} \lesssim_\P \varrho_\lambda$, and $\|\widehat{\boldsymbol F}_t-\boldsymbol F_t\|_{2} \lesssim_\P \varrho_F$, respectively. Then, for every $t\geq 1$,
\[
\max_{i\in[n]}\left|\widehat{Y}_{i,t}-\widetilde{Y}_{i,t}\right|\lesssim_\P (\varrho_\gamma+\varrho_\theta) (\log n)^{1/\gamma_2} + \varrho_U s_1+ \varrho_P + \varrho_\lambda +\varrho_F,\]
where $s_1$ is defined in Theorem (ref).
\end{theorem}
\begin{theorem}
For $\mathcal{D}\subseteq [n]^2$, let $\widetilde{\boldsymbol J}:=\sqrt{T}(\widetilde{\boldsymbol\Sigma}_\mathcal{D} - \boldsymbol\Sigma_\mathcal{D})$ and $\boldsymbol{\mathcal{G}}$ be a zero-mean Gaussian vector with the same covariance matrix of $\widetilde{\boldsymbol J}$, i.e., $\boldsymbol{\mathcal{G}}\sim N(\boldsymbol{0},\boldsymbol\Upsilon_\Sigma)$.
Under Assumptions (ref)--(ref), if further
\begin{enumerate}[(a)]
• $\{\boldsymbol U_t:t\in[T]\}$ is fourth-order stationary process for each $T$;
• the minimum eigenvalue of $\boldsymbol\Upsilon_\Sigma$ is greater or equal to $\underline{c}$, for some $\underline{c}>0$,
\end{enumerate}
then,
\begin{align*}
\rho(\widetilde{\boldsymbol J},\boldsymbol{\mathcal{G}})&\lesssim \frac{(\log T)^{\gamma_1 + 1} \log d + \big[\log (dT)\big]^{2/\gamma}(\log d)^2\log T}{\sqrt{T}\underline{c}^2} \\
&\qquad + \frac{(\log d)^{2} +(\log d)^{3/2}\log T+ \log d (\log T)^{\gamma_1+1}\log (dT)}{T^{1/4}\underline{c}^2},
\end{align*}
where $d:=|\mathcal{D}|$.
Let $\widehat{\boldsymbol J}:=\sqrt{T}(\widehat{\boldsymbol\Sigma}_\mathcal{D} - \boldsymbol\Sigma_\mathcal{D})$, then
\[
\rho(\widehat{\boldsymbol J},\boldsymbol{\mathcal{G}})\lesssim\rho(\widetilde{\boldsymbol J},\boldsymbol{\mathcal{G}}) + \inf_{\delta>0}\left[\delta_1\sqrt{1\lor \log (d/\delta)} + \P(\|\widehat{\boldsymbol J}-\widetilde{\boldsymbol J}\|_\infty>\delta)\right]
.\]
Let $\widetilde{\boldsymbol \Upsilon}_\Sigma$ be any positive semi definite estimator of $\boldsymbol\Upsilon_\Sigma$ and $\boldsymbol{\mathcal{G}}^*|\boldsymbol X,\boldsymbol Y\sim N(\boldsymbol 0,\widetilde{\boldsymbol \Upsilon}_\Sigma)$, then
\[
\rho(\widehat{\boldsymbol J},\boldsymbol{\mathcal{G}}^*)\lesssim \rho(\widehat{\boldsymbol J},\boldsymbol{\mathcal{G}}) + \inf_{\delta>0}\left[\delta \log d(1\lor |\log d|) + \P(\|\widetilde{\boldsymbol\Upsilon}_\Sigma - \boldsymbol\Upsilon_\Sigma\|_{\max}>\delta)\right].
\]
\end{theorem}
\begin{theorem}
For $\mathcal{D}\subseteq [n]^2$, let $\widetilde{\boldsymbol Q}:=\sqrt{T}(\widetilde{\boldsymbol\Pi}_\mathcal{D} - \boldsymbol\Pi_\mathcal{D})$ and $\boldsymbol{\mathcal{H}}$ be a zero-mean Gaussian vector with the same covariance matrix of $\widetilde{\boldsymbol Q}$, i.e., $\boldsymbol{\mathcal{H}}\sim N(\boldsymbol 0,\boldsymbol\Upsilon_\Pi)$.
Under the same assumptions and notation of Theorem (ref) with $\boldsymbol\Upsilon_\Sigma$ and $\underline{c}$ replaced by $\boldsymbol\Upsilon_\Pi$ and $\underline{b}$, respectively in condition (c), we have
\begin{align*}
\rho(\widetilde{\boldsymbol Q},\boldsymbol{\mathcal{H}})&\lesssim \frac{(\log T)^{\gamma_1 + 1} \log d + \big[\log (dT)\big]^{2/\gamma}(\log d)^2\log T}{\sqrt{T}\underline{b}^2} \\
&\qquad + \frac{(\log d)^{2} +(\log d)^{3/2}\log T+ \log d (\log T)^{\gamma_1+1}\log (dT)}{T^{1/4}\underline{b}^2},
\end{align*}
where $d:=|\mathcal{D}|$.
Let $\widehat{\boldsymbol Q}:=\sqrt{T}(\widehat{\boldsymbol\Pi}_\mathcal{D} - \boldsymbol\Pi_\mathcal{D})$, then
\[
\rho(\widehat{\boldsymbol Q},\boldsymbol{\mathcal{H}})\lesssim\rho(\widetilde{\boldsymbol Q},\boldsymbol{\mathcal{H}}) + \inf_{\delta>0}\left[\delta_1\sqrt{1\lor \log (d/\delta)} + \P(\|\widehat{\boldsymbol Q}-\widetilde{\boldsymbol Q}\|_\infty>\delta)\right]
.\]
Let $\widetilde{\boldsymbol \Upsilon}_\Pi$ be any positive semi definite estimator of $\boldsymbol\Upsilon_\Pi$ and $\boldsymbol{\mathcal{H}}^*|\boldsymbol X,\boldsymbol Y\sim N(\boldsymbol 0,\widetilde{\boldsymbol \Upsilon}_\Pi)$, then
\[
\rho(\widehat{\boldsymbol Q},\boldsymbol{\mathcal{H}}^*)\lesssim \rho(\widehat{\boldsymbol Q},\boldsymbol{\mathcal{H}}) + \inf_{\delta>0}\left[\delta \log d(1\lor |\log d|) + \P(\|\widetilde{\boldsymbol\Upsilon}_\Pi - \boldsymbol\Upsilon_\Pi\|_{\max}>\delta)\right].
\]
\end{theorem}
\section{Proof of the Theorems}
\subsection{Proof of Theorems (ref) and (ref)}
We first upper bound $|\widehat{R}_{i,t}-R_{i,t}|$. Recall that $\widehat{R}_{i,t}-R_{i,t}=(\widehat{\boldsymbol\gamma}_i-\boldsymbol\gamma_i)'\boldsymbol X_{i,t}$. Then, by subsequent application of Hölder's inequality, we have
\begin{align}
|\widehat{R}_{i,t}-R_{i,t}|
\leq \|\widehat{\boldsymbol\gamma}_i-\boldsymbol\gamma_i\|\|\boldsymbol X_{i,t}\|\leq\|\widehat{\boldsymbol\Sigma}_i^{-1}\| \|\widehat{\boldsymbol v}_i\|\|\boldsymbol X_{i,t}\|\leq k\|\widehat{\boldsymbol\Sigma}_i^{-1}\| \|\widehat{\boldsymbol v}_i\|_\infty\|\boldsymbol X_{i,t}\|_\infty,
\end{align}
where $\widehat{\boldsymbol\Sigma}_i:=\boldsymbol X_i'\boldsymbol X_i/T$ and $\widehat{\boldsymbol v}_i:= \boldsymbol X_i'\boldsymbol R_i/T$. Therefore,
\begin{equation}
\|\widehat{\boldsymbol R}- \boldsymbol R\|_{\max}\leq k\left(\max_{i}\left\|\widehat{\boldsymbol\Sigma}_i^{-1}\right\|\right)\left(\max_{i,j}\left|\widehat{v}^{(j)}_i\right|\right)\left(\max_{i,t,j}\left|X^{(j)}_{i,t}\right|\right),
\end{equation}
where $\widehat{v}_i^{(j)}$ and $X_{i,t}^{(j)}$ denote the $j$-th component of $\widehat{\boldsymbol v}_i$ and $\boldsymbol X_{i,t}$ respectively.
Under Assumption (ref), ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \|\widehat{\boldsymbol\Sigma}_i^{-1}\|
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_\psi\leq C$. By definition, $R_{i,t} = \boldsymbol \lambda_i'\boldsymbol F_t + U_{i,t}$, then
\[
{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert R_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_\psi\leq \|\boldsymbol \lambda_i\| {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \boldsymbol F_t
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_\psi + {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert U_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_\psi\leq (1+ \sqrt{r}\|\Lambda\|_{\max}){\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \boldsymbol Z_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert} \leq (1+C\sqrt{r})C
\]
under Assumptions (ref) and (ref). Also, $\boldsymbol X_{i,t}R_{i,t}$ is a function of $\boldsymbol Z_{t}$. Then, $\{\boldsymbol X_{i,t}R_{i,t}\}_t$ is a zero-mean sequence with strong mixing coefficient upper bounded by the strong mixing coefficient of $\{\boldsymbol Z_t\}_t$.
For the polynomial case, applying the union bound followed by Markov's inequality we conclude that
\[
\max_{i}\left\|\widehat{\boldsymbol\Sigma}_i^{-1}\right\|\lesssim_\P n^{1/p}
\quad\textnormal{and}\quad
\max_{i,t,j}\left|X^{(j)}_{i,t}\right|\lesssim_\P (nkT)^{1/p}.
\]
Also, by the Cauchy-Schwartz inequality
\[
{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X_{i,t}^{(j)}R_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{(p+\epsilon)/2}\leq {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X_{i,t}^{(j)}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p+\epsilon}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert R_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p+\epsilon}\leq (1+C\sqrt{r})C^2.
\]
Then, by Lemma (ref), $\max_{i,j}|\widehat{v}^{(j)}_i|\lesssim_\P\frac{\mathscr{R}_\alpha\sqrt{r}(nk)^{2/p}}{\sqrt{T}}$. Plugging the last three probability bounds back on (ref) yields the result for the polynomial case.
Similarly, for the exponential case, the union bound followed by Lemma (ref) gives us
\[
\max_{i}\|\widehat{\boldsymbol\Sigma}_i^{-1}\|\lesssim_\P (\log n)^{1/\gamma_2}\quad
\textnormal{and}\quad \max_{i,t,j}|X^{(j)}_{i,t}|\lesssim_\P [\log(nkT)]^{1/\gamma_2}.
\]
Also, by Lemma (ref),
\[
{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X_{i,t}^{(j)}R_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma_2/2}}\lesssim {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X_{i,t}^{(j)}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma_2}}^2\lor {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert R_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma_2}}^2\lesssim 1.
\]
Hence, $\max_{i,j}|\widehat{v}^{(j)}_i|\lesssim_\P\sqrt{\frac{\log (nk)}{T}}$ by Lemma (ref). Plugging the last three probability bounds back on (ref) yields the result for the exponential case,
\subsection{Proof of Theorems (ref) and (ref)}
The proof is an adaption of the proof of Theorem 4 and Corollary 1 in FLM2013, henceforth FLM, to accommodate (i) the serial dependency (strong mixing sequences), (ii) polynomial tails, and (iii) the estimation error in the sample covariance matrix. For part (a), we use expression (A.1) in Bai2003 to obtain the following identity
\begin{equation}
\widehat{\boldsymbol F}_t - \boldsymbol H \boldsymbol F_t = \left(\frac{\boldsymbol V}{n}\right)^{-1}\left[\frac{1}{T}\sum_{s=1}^T\widehat{\boldsymbol F}_s\frac{\mathbb{E}(\boldsymbol U_s'\boldsymbol U_t)}{n} + \frac{1}{T}\sum_{s=1}^T\left(\widehat{\boldsymbol F}_s\widetilde{\zeta}_{st} + \widehat{\boldsymbol F}_s\widetilde{\eta}_{st} +\widehat{\boldsymbol F}_s\widetilde{\xi}_{st}\right)\right],
\end{equation}
where $\boldsymbol V$ is a $(r\times r)$ diagonal matrix whose diagonal is given by the eigenvalues of $\widehat{\boldsymbol\Lambda}'\widehat{\boldsymbol\Lambda}$ in decreasing fashion with $\widehat{\boldsymbol\Lambda}:=\widehat{\boldsymbol R}\widehat{\boldsymbol F}/T$ ; and $\widetilde{\zeta}_{st},\widetilde{\eta}_{st}$ and $\widetilde{\xi}_{st}$ are defined before Lemma (ref).
By Assumptions (ref)(d) and (ref) we have $\|\boldsymbol R\|_{\max}\leq r\|\boldsymbol\Lambda\|_{\max}\|\boldsymbol F\|_{\max} +\|\boldsymbol U\|_{\max}\lesssim_\P g(nT)$. Applying Lemma (ref) we conclude that $\|\widehat{\boldsymbol\Sigma} - \widetilde{\boldsymbol\Sigma}\|_{\max} \lesssim_\P \varrho_R[g(nT) + \varrho_R]\lesssim_\P 1$. Finally, $\psi
g_\alpha(n)/\sqrt{T}\lesssim 1$ also by assumption. Then, $\|(\frac{\boldsymbol V}{n})^{-1}\| \lesssim_\P 1$ by Lemma (ref). Using the results (a)-(d) of Lemma (ref) we can bound in probability each of the terms in brackets of (ref) in $\ell_2$ norm, uniformly in $t\leq T$. Result (a) follows since
\[
\max_{t\in[T]}\|\widehat{\boldsymbol F}_t - \boldsymbol H \boldsymbol F_t\|\lesssim_\P \left[\tfrac{1}{\sqrt{T}} + \frac{g(T)}{\sqrt{n}} + g(nT)\varrho_R\right],
\]
where $g(x)= x^{1/p}$ under Assumption ((ref).c) and $g(x)=[\log (x)]^{1/\gamma_2}$ under Assumption ((ref).d).
For part (b) we use the fact that $\widehat{\boldsymbol\Lambda}:=\widehat{\boldsymbol R}\widehat{\boldsymbol F}/T$ and set $\widehat{\boldsymbol F}'\widehat{\boldsymbol F} = \boldsymbol I_r$ to write
\begin{equation}
\widehat{\boldsymbol\lambda}_i - \boldsymbol H \boldsymbol\lambda_i = \frac{1}{T}\sum_{t=1}^T\boldsymbol H \boldsymbol F_t \widetilde{U}_{i,t} + \frac{1}{T}\sum_{t=1}^T\widehat{R}_{i,t}(\widehat{\boldsymbol F}_t - \boldsymbol H\boldsymbol F_t) + \boldsymbol H\left(\frac{1}{T}\sum_{t=1}^T\boldsymbol F_t\boldsymbol F_t' - \boldsymbol I_r\right)\boldsymbol\lambda_i.
\end{equation}
The first term can be upper bounded in $\ell_2$ norm, uniformly in $i\leq n $, by
\[
\sqrt{r}\|\boldsymbol H\|\max_{i\leq n}\max_{j\leq r}\left|\frac{1}{T}\sum_{t=1}^T F_{jt} \widetilde{U}_{i,t}\right|\lesssim_\P g_1(n)/\sqrt{T} + \varrho_R,
\]
where the equality follows from Lemma (ref)(b) and (e), and $g_1(x)= \mathscr{R}_\alpha x^{2/p}$ under Assumption ((ref).c) and $g_1(x)=\sqrt{\log (x)}$ under Assumption ((ref).d). The $\ell_2$-norm of the second term is upper bounded uniformly in $i\leq n $ by
\[
\left(\max_{i\leq n}\frac{1}{T}\sum_{t=1}^T\widehat{R}_{i,t}^2 \frac{1}{T}\sum_{t=1}^T\| \widehat{\boldsymbol F}_t - \boldsymbol H\boldsymbol F_t\|^2\right )^{1/2} \lesssim_\P \left[\frac{1}{T} +(1/\sqrt{n}+)^2\right]^{1/2},
\]
where the first term after the equality follows from Lemma (ref)(d) together with the theorem's assumption and the second term from Lemma (ref)(e). Finally, the last term of (ref) is upper bounded by
\[
\|\boldsymbol H\|\|\max_{i\leq n}\boldsymbol\lambda_i\|\left\|\frac{1}{T}\sum_{t=1}^T\boldsymbol F_t\boldsymbol F_t' - \boldsymbol I_r\right\|\lesssim_\P 1/\sqrt{T},
\]
where the last term is $\lesssim_\P 1/\sqrt{T}$ by the maximum inequality and Assumption (ref).
Plugging the last three displays back into (ref) yields result (b).
For part (c) we have $\|\widehat{\boldsymbol U}-\boldsymbol U\|_{\max} = \| \boldsymbol\Lambda\boldsymbol F' - \widehat{\boldsymbol\Lambda}\widehat{\boldsymbol F}' +\widehat{\boldsymbol R} - \boldsymbol R\|_{\max} \leq \| \widehat{\boldsymbol\Lambda}\widehat{\boldsymbol F}' - \boldsymbol\Lambda\boldsymbol F' \|_{\max} +\|\widehat{\boldsymbol R} - \boldsymbol R\|_{\max}$. The last term is $\lesssim_\P \varrho_R$ by assumption. For the first term we use the decomposition
\begin{align}
\widehat{\boldsymbol\lambda}_i'\widehat{\boldsymbol F}_t - \boldsymbol\lambda_i'\boldsymbol F_t&=(\widehat{\boldsymbol\lambda}_i - \boldsymbol H\boldsymbol\lambda_i)'(\widehat{\boldsymbol F}_t - \boldsymbol H\boldsymbol F_t) + (\boldsymbol H\boldsymbol\lambda_i)'(\widehat{\boldsymbol F}_t - \boldsymbol H\boldsymbol F_t) \nonumber\\
&\qquad+ (\widehat{\boldsymbol\lambda}_i - \boldsymbol H\boldsymbol\lambda_i)'\boldsymbol H\boldsymbol F_t +\boldsymbol\lambda_i'(\boldsymbol H'\boldsymbol H - \boldsymbol I_r)\boldsymbol F _t.
\end{align}
Therefore, we can upper bound the left hand side as
\begin{align*}
|\widehat{\boldsymbol\lambda}_i'\widehat{\boldsymbol F}_t - \boldsymbol\lambda_i'\boldsymbol F_t|&\leq\|\widehat{\boldsymbol\lambda}_i - \boldsymbol H\boldsymbol\lambda_i\|\|\widehat{\boldsymbol F}_t - \boldsymbol H\boldsymbol F_t\| + \|\boldsymbol H\boldsymbol\lambda_i\|\|\widehat{\boldsymbol F}_t - \boldsymbol H\boldsymbol F_t\| \\
&\qquad+ \|\widehat{\boldsymbol\lambda}_i - \boldsymbol H\boldsymbol\lambda_i\|\|\boldsymbol H\boldsymbol F_t\| +\|\boldsymbol\lambda_i\|\|\boldsymbol F _t\|\|\boldsymbol H'\boldsymbol H - \boldsymbol I_r\|.
\end{align*}
Now, we bound in probability, uniformly in $i\leq n$ and $t\leq T$, each of the four terms above. The first one is given by parts (a) and (b). $\max_{i\leq n}\|\boldsymbol H\boldsymbol\lambda_i\|\leq \|\boldsymbol H\| \max_{i\leq n}||\boldsymbol \lambda_i\|\lesssim_\P r\|\boldsymbol\Lambda\|_{\max} \lesssim_\P 1$ by Lemma (ref)(b) and Assumption (ref)(d). Thus, the second term is bounded by part (a). For the third term, $\max_{t\leq T}\|\boldsymbol H\boldsymbol F_t\|\leq \|\boldsymbol H\|\max_{t\leq T}||\boldsymbol F_t\|\lesssim_\P g(T)$ by Lemma (ref)(b) and Assumption (ref). Finally, $\|\boldsymbol H'\boldsymbol H - \boldsymbol I_r\| \lesssim_\P 1/\sqrt{T} + 1/\sqrt{n} + \varrho_R$ by Lemma (ref)(c). Hence, the last term is $\lesssim_\P g(T)(1/\sqrt{T} + 1/\sqrt{n} + \varrho_R)$ by Assumptions (ref)(d) and (ref).
\subsection{Proof of Theorems (ref) and (ref)}
For $i\in[n]$, let $\mathcal{Q}_i(\boldsymbol a):=\|\widehat{\boldsymbol U}_{i,\cdot}-\boldsymbol a' \widehat{\boldsymbol W}_{i,\cdot}\|_2^2/T$ for $\boldsymbol a\in\mathbb{R}^d$. We have that $L(\widehat{\boldsymbol\theta}_i)+\xi\|\widehat{\boldsymbol\theta}_i\|_1\leq \mathcal{Q}(\boldsymbol \theta_i) +\xi\|\boldsymbol \theta_i\|_1$ by definition of $\widehat{\boldsymbol\theta}_i$. Also, since $\mathcal{Q}_i(\boldsymbol\theta)$ is a quadratic function, it implies that $(\widehat{\boldsymbol\theta}_i -\boldsymbol\theta_i)'\widehat{\boldsymbol\Sigma}_i(\widehat{\boldsymbol\theta}_i -\boldsymbol\theta_i)\leq 2\widehat{\boldsymbol E}_i'(\widehat{\boldsymbol\theta}_i -\boldsymbol\theta_i)+\xi(\|\boldsymbol\theta_i\|_1 - \|\widehat{\boldsymbol\theta}_i\|_1)$ where $\widehat{\boldsymbol\Sigma}_i:= \widehat{\boldsymbol W}_{i,\cdot} \widehat{\boldsymbol W}_{i,\cdot}'/T$ and $\widehat{\boldsymbol E}_i :=(\widehat{\boldsymbol U}_{i,\cdot}-\boldsymbol \theta_i' \widehat{\boldsymbol W}_{i,\cdot})'\widehat{\boldsymbol W}_{i,\cdot}/T$. By Holder's inequality, we have $|\widehat{\boldsymbol E}_i'(\widehat{\boldsymbol\theta}_i -\boldsymbol\theta_i)|\leq \|\widehat{\boldsymbol E}_i\|_\infty\|\widehat{\boldsymbol\theta}_i -\boldsymbol\theta_i\|_1$. Then, for $\xi\geq 4\|\widehat{\boldsymbol E}_i\|_\infty$ we have
\begin{equation}
(\widehat{\boldsymbol\theta}_i -\boldsymbol\theta_i)'\widehat{\boldsymbol\Sigma}_i(\widehat{\boldsymbol\theta}_i -\boldsymbol\theta_i)\leq \xi/2\|\widehat{\boldsymbol\theta}_i -\boldsymbol\theta_i\|_1 +\xi(\|\boldsymbol\theta_i\|_1 - \|\widehat{\boldsymbol\theta}_i\|_1).
\end{equation}
For any index set $\mathcal{S}\subseteq [n]$, by the decomposability of the $\ell_1$ norm (refer to Definition 1 in Negahban2012) followed by the triangle inequality, we have $\|\widehat{\boldsymbol\theta_i}\|_1= \|\widehat{\boldsymbol\theta}_{i,\mathcal{S}}\|_1 +\|\widehat{\boldsymbol\theta}_{i,\mathcal{S}^c}\|_1 \geq \|\boldsymbol\theta_{i,\mathcal{S}}\|_1 - \|\widehat{\boldsymbol\theta}_{i,\mathcal{S}} -\boldsymbol\theta_{i,\mathcal{S}}\|_1+\|\widehat{\boldsymbol\theta}_{i,\mathcal{S}^c}\|_1 $ and $\|\widehat{\boldsymbol\theta}_i -\boldsymbol\theta_i\|_1 = \|\widehat{\boldsymbol\theta}_{i,\mathcal{S}} -\boldsymbol\theta_{i,\mathcal{S}}\|_1 + \|\widehat{\boldsymbol\theta}_{i,\mathcal{S}^c} -\boldsymbol\theta_{i,\mathcal{S}^c}\|_1\leq \|\widehat{\boldsymbol\theta}_{i,\mathcal{S}} -\boldsymbol\theta_{i,\mathcal{S}}\|_1 + \|\widehat{\boldsymbol\theta}_{i,\mathcal{S}^c} -\boldsymbol\theta_{i,\mathcal{S}^c}\|_1 $. Plugging it back in (ref) yields
\begin{equation}
2(\widehat{\boldsymbol\theta}_i-\boldsymbol\theta_i)'\widehat{\boldsymbol\Sigma}_i(\widehat{\boldsymbol\theta}_i -\boldsymbol\theta_i) + \xi\|\widehat{\boldsymbol\theta}_{i,\mathcal{S}^c} -\boldsymbol\theta_{i,\mathcal{S}^c}\|_1\leq
3\xi\|\widehat{\boldsymbol\theta}_{i,\mathcal{S}} -\boldsymbol\theta_{i,\mathcal{S}}\|_1 +4\xi\|\boldsymbol\theta_{i,\mathcal{S}^c}\|_1.
\end{equation}
We then conclude that any minimizer of (ref) obeys $\widehat{\boldsymbol\theta}_i -\boldsymbol\theta_i \in \mathbb{C}(\mathcal{S}_i,3):=\{\boldsymbol x\in\mathbb{R}^d:\|\boldsymbol x_{\mathcal{S}^c_i}\|_1\leq 3\|\boldsymbol x_{\mathcal{S}_i}\|_1\|_1\}$ where $\mathcal{S}_i:=\{j:\theta_{i,j}\neq 0\}$ is the support of $\boldsymbol\theta_i$.
Recall the definition of the compatibility constant appearing in vandegeer2009, reproduced below for convenience.
\begin{Def}
For an $n\times n$ matrix $\boldsymbol{M}$, a set $\mathcal{S}\subseteq\{1,\ldots,n\}$ and a scalar $\zeta \geq 0$, the compatibility constant is given by
\begin{equation}
\kappa(\boldsymbol{M},\mathcal{S},\zeta):=\inf\left\{\frac{\|\boldsymbol{x}\|_{\boldsymbol{M}}\sqrt{|\mathcal{S}|}}{\|\boldsymbol{x}_\mathcal{S}\|_1}:\boldsymbol{x}\in\mathbb{R}^n:\|\boldsymbol{x}_{\mathcal{S}^c}\|_1\leq\xi\|\boldsymbol{x}_{\mathcal{S}}\|_1 \right\},
\end{equation}
where $\|\boldsymbol{x}\|_{\boldsymbol{M}} = \sqrt{\boldsymbol{x}'{\boldsymbol{M}} \boldsymbol{x}}$.
Moreover, we say that $(\boldsymbol{M},\mathcal{S},\zeta)$ satisfies the compatibility condition if $\kappa(\boldsymbol M,\mathcal{S},\zeta)>0$.
\end{Def}
Notice that the square of the compatibility constant is close related to the minimum of the $\ell_1$-norm of the eigenvalues of $\boldsymbol M$, restricted to a cone in $\mathbb{R}^n$. By definition of the compatibility constant $\widehat{\kappa}_i:=\kappa( \widehat{\boldsymbol\Sigma}_i,\mathcal{S}_0,3)$, we have that $\|\widehat{\boldsymbol\theta}_{i,\mathcal{S}} -\boldsymbol\theta_{i,\mathcal{S}}\|_1\leq \sqrt{(\widehat{\boldsymbol\theta}_i -\boldsymbol\theta_i)'\widehat{\boldsymbol\Sigma}_i(\widehat{\boldsymbol\theta}_i -\boldsymbol\theta_i)} \sqrt{|\mathcal{S}_i|}/\widehat{\kappa}_i$. Apply this inequality to (ref) and use the fact that $4ab<a^2 + 4b^2$ for non-negative $a,b\in\mathbb{R}$, to obtain
\begin{equation}
(\widehat{\boldsymbol\theta}_i -\boldsymbol\theta_i)'\widehat{\boldsymbol\Sigma}_i(\widehat{\boldsymbol\theta}_i -\boldsymbol\theta_i) + \xi\|\widehat{\boldsymbol\theta}_{i} -\boldsymbol\theta_i\|_1\leq 4\xi^2|\mathcal{S}_i|/\widehat{\kappa}_i^2,
\end{equation}
provided that $\xi\geq 4\|\widehat{\boldsymbol E}_i\|_\infty$.
Let $\boldsymbol\Sigma_i:=\mathbb{E}(\boldsymbol W_{i,t} \boldsymbol W_{i,t}')$ and $\kappa_i:=\kappa( \boldsymbol\Sigma_i,\mathcal{S}_0,3)$ . Note that $\min_i\kappa_i^2\geq b^2>0$ under Assumption ((ref).g) because
\[\lambda_{\min}\left[\mathbb{E}\left(\boldsymbol{\mathcal{U}}_t\boldsymbol{\mathcal{U}}_t'\right)\right]\leq\min_i \lambda_{\min}(\boldsymbol\Sigma_i)\leq \min_i \min_{\boldsymbol x\in\mathbb{C}(\mathcal{S}_i,3)}\left\{\frac{\boldsymbol x'\boldsymbol\Sigma_i \boldsymbol x}{\boldsymbol x'\boldsymbol x}\right\}\leq \min_i \kappa_i^2,\]
where in the first inequality we use Cauchy interlacing theorem and in the last on we use the fact that $\boldsymbol x' \boldsymbol x \geq \boldsymbol x_{\mathcal{S}_i}' \boldsymbol x_{\mathcal{S}_i}$ and the bound $\|\boldsymbol x_{\mathcal{S}_i}\|_1\leq \sqrt{|\mathcal{S}_i}\|\boldsymbol x_{\mathcal{S}_i}\|_2$.
Also, we may use Lemma (ref) with $\zeta =3$ and $\alpha =1/2$ to assert that if $\|\widehat{\boldsymbol\Sigma}_i-\boldsymbol\Sigma_i\|_{\max}\leq \xi_1$ for some $\xi_1>0$, such that $32\xi_1 |\mathcal{S}_i|/\kappa_i^2\leq 1$ then, $\widehat\kappa_i^2 \geq \kappa_i^2/2$ and $\boldsymbol x'\boldsymbol\Sigma_i \boldsymbol x/2\leq \boldsymbol x'\widehat{\boldsymbol \Sigma}_i \boldsymbol x$ for $\boldsymbol x\in\mathbb{C}(\mathcal{S}_i,3)$. Use the last three inequalities in (ref) to conclude that, for $i\in[n]:$
\begin{equation}
\tfrac{1}{2}(\widehat{\boldsymbol\theta}_i -\boldsymbol\theta_i)'\boldsymbol\Sigma_i(\widehat{\boldsymbol\theta}_i -\boldsymbol\theta_i) + \xi\|\widehat{\boldsymbol\theta}_{i} -\boldsymbol\theta_i\|_1\leq 8\xi^2|\mathcal{S}_i|/b^2,
\end{equation}
provided that $\xi\geq 4\|\widehat{\boldsymbol E}_i\|_\infty$ and $\|\widehat{\boldsymbol\Sigma}_i-\boldsymbol\Sigma_i\|_{\max}\leq \xi_1$ such that $32\xi_1 |\mathcal{S}_i|/b\leq 1$.
\subsubsection{Probabilistic Bounds}
We now bound in probability the events $\{\xi\geq 4\|\widehat{\boldsymbol E}_i\|_\infty\}$ and $\{\|\widehat{\boldsymbol\Sigma}_i-\boldsymbol\Sigma_i\|_{\max}\leq \xi_1\}$, uniformly in $i\in[n]$. For the former, by definition of $\boldsymbol W_{i,t}$, we have $\max_{i,t}\|\boldsymbol W_{i,t}\|_\infty\leq \|\boldsymbol U\|_{\max}$ and $\max_{i,t}\|\widehat{\boldsymbol W}_{i,t} - \boldsymbol W_{i,t}\|_{\infty}\leq \|\widehat{\boldsymbol U} - \boldsymbol U\|_{\max}$. Also, by the definition of $V_{i,t}$, the triangle inequality followed by Hölder's inequality, we have $\max_{i,t}|V_{i,t}|\leq \max_{i,t}|U_{i,t}| +\max_{i,t}\|\boldsymbol\theta_i\|_1\boldsymbol|\boldsymbol W_{i,t}\|_\infty\leq (1+\max_i\|\boldsymbol\theta_i\|_1)\|\boldsymbol U\|_{\max}$. Similarly, we conclude $\max_{i,t}|\widehat{V}_{i,t} - V_{i,t}| \leq (1+\max_i\|\boldsymbol\theta_i\|_1)\|\widehat{\boldsymbol U} - \boldsymbol U\|_{\max}$.
First, we decompose
\[T\widehat{\boldsymbol E}_i=\boldsymbol W_{i,\cdot} \boldsymbol V_{i,\cdot} + \boldsymbol W_{i,\cdot} (\widehat{\boldsymbol V}_{i,\cdot}-\boldsymbol V_{i,\cdot})+(\widehat{\boldsymbol W}_{i,\cdot} -\boldsymbol W_{i,\cdot} )\boldsymbol V_{i,\cdot}+(\widehat{\boldsymbol W}_{i,\cdot} -\boldsymbol W_{i,\cdot} )(\widehat{\boldsymbol V}_{i,\cdot}-\boldsymbol V_{i,\cdot}),\]
and bound each term individually as
\begin{align*}
\max_i\|\boldsymbol W_{i,\cdot} (\widehat{V}_{i,\cdot}-V_{i,\cdot})/T\|_\infty&\leq\max_{i,t}\|\boldsymbol W_{i,t}\|_\infty\max_{i,t}|\widehat{V}_{i,t} - V_{i,t}| \\ &\leq (1+\max_{i} \|\boldsymbol\theta_i\|_1)\|\boldsymbol U\|_{\max}\|\widehat{\boldsymbol U} - \boldsymbol U\|_{\max},
\end{align*}
\begin{align*}
\max_i\|(\widehat{\boldsymbol W}_{i,\cdot} -\boldsymbol W_{i,\cdot} )V_{i,\cdot}/T\|_\infty&\leq\max_{i,t}\|\widehat{\boldsymbol W}_{i,t} -\boldsymbol W_{i,t} \|_\infty\max_{i,t}| V_{i,t}| \\ &\leq (1+\max_{i} \|\boldsymbol\theta_i\|_1)\|\boldsymbol U\|_{\max}\|\widehat{\boldsymbol U} - \boldsymbol U\|_{\max},
\end{align*}
and
\begin{align*}
\max_i\|(\widehat{\boldsymbol W}_{i,\cdot} -\boldsymbol W_{i,\cdot} )(\widehat{V}_{i,\cdot}-V_{i,\cdot})/T\|_\infty&\leq\max_{i,t}\|\widehat{\boldsymbol W}_{i,t} -\boldsymbol W_{i,t}\|_\infty\max_{i,t}|\widehat{V}_{i,t} - V_{i,t}| \\ &\leq (1+\max_{i} \|\boldsymbol\theta_i\|_1)\|\widehat{\boldsymbol U} - \boldsymbol U\|_{\max}^2.
\end{align*}
Then,
\begin{align}
\max_{i}\|\widehat{\boldsymbol E}_i\|_\infty&\leq \max_{i}\|\boldsymbol W_{i,\cdot} \boldsymbol V_{i,\cdot} /T\|_\infty\\
&\qquad +(1+\max_{i} \|\boldsymbol\theta_i\|_1)(2\|\boldsymbol U\|_{\max}\|\widehat{\boldsymbol U} - \boldsymbol U\|_{\max} +\|\widehat{\boldsymbol U} - \boldsymbol U\|_{\max}^2 )\nonumber.
\end{align}
We now bound ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \boldsymbol W_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_\psi $ and ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert V_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_\psi$. For the former, ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \boldsymbol W_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_\psi \leq \sup_t {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \boldsymbol U_{\cdot,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_\psi$ and for the latter
\begin{align}
{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert V_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_\psi &= {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert U_{i,t}-\boldsymbol\theta_i' \boldsymbol W_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_\psi\nonumber\\
&\leq {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert U_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}+{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \boldsymbol\theta_i' \boldsymbol W_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_\psi\nonumber\\
&\leq {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert U_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}+\|\boldsymbol \theta_i\|_2{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \boldsymbol W_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_\psi\nonumber\\
&\leq (1+\|\boldsymbol \theta_i\|_2)\sup_t{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \boldsymbol U_{\cdot,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_\psi.
\end{align}
By definition, we have that $\boldsymbol W_{i,t}V_{i,t} = f_i(\boldsymbol U_{\cdot,t},\dots, \boldsymbol U_{\cdot, t-l})$ for some mapping $f_i$, for each $i\in[n]$ and $t\in[T]$. Then, by Lemma (ref), we have that the strong mixing coefficients of the sequence $\{W_{i,t}^{(j)}V_{i,t}\}_t$ are such that $\widetilde{\alpha}_{m}\leq \alpha_{(m-l)\lor 0}$, where $W_{i,t}^{(j)}$ denote the $j$-th component of $\boldsymbol W_{i,t}$ for $j\in[d]$.
Under Assumption (ref)(c), by the Cauchy-Schwartz inequality and the last two bounds, we have
\begin{align*}
{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert W_{i,t}^{(j)}V_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/2+\epsilon/2}&\leq {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert W_{i,t}^{(j)}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p+\epsilon}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert V_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p+\epsilon}\\
&\leq (1+\|\boldsymbol \theta_i\|_2)\sup_t{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \boldsymbol U_{\cdot,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p+\epsilon}^2\leq C^2\left(1+\max_i\|\boldsymbol \theta_i\|_2\right).
\end{align*}
Therefore, by Lemma (ref) followed by Lemma (ref), we conclude that for $p\in[4,\infty)$
\[
\max_{i,j}|\sum_{t=1}^T W_{i,t}^{(j)}V_{i,t}| \lesssim_\P\left(1+\max_i\|\boldsymbol \theta_i\|_2\right) \mathscr{L}_\alpha\left[n(l+1)\right]^{2/p}\sqrt{T},
\]
where $\mathscr{L}_\alpha:=\big(\sum_{t=0}^{T-1}(m+1)^{(p/2)-2}\alpha_{(m-l)\lor 0}^{1-p/(p+\epsilon)}\big)^{2/p}$. By assumption $\|\widehat{\boldsymbol U}- \boldsymbol U\|_{\max} \lesssim_\P\varrho_U$, and the union bound followed by Markov's inequality give us $\|\boldsymbol U\|_{\max} \lesssim_\P nT^{1/p}$. Then, from (ref), we obtain
\begin{align}
\max_{i}\|\widehat{\boldsymbol E}_i\|_\infty&\lesssim_\P (1+\max_i\|\boldsymbol \theta_i\|_2) \mathscr{L}_\alpha\big[n(l+1)\big]^{2/p}T^{-1/2}\\
&\qquad +(1+\max_{i} \|\boldsymbol\theta_i\|_1)\left[(nT)^{1/p}\varrho_U + \varrho_U^2\right]\nonumber=:\xi_0.
\end{align}
Similarly, under Assumption (ref)(d) and Lemma (ref) we have
\begin{align*}
{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert W_{i,t}^{(j)}V_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma_2/2}}&\lesssim{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert W_{i,t}^{(j)}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma_2}}^2\lor {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert V_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma_2}}^2\\
&\lesssim (1+\|\boldsymbol \theta_i\|_2)\sup_t{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \boldsymbol U_{\cdot,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma_2}}
\lesssim(1+\max_i\|\boldsymbol \theta_i\|_2).
\end{align*}
Therefore, by Lemma (ref) we conclude that
\[
\max_{i,j}\left|\sum_{t=1}^T W_{i,t}^{(j)}V_{i,t}\right| \lesssim_\P(1+\max_i\|\boldsymbol \theta_i\|_2) \sqrt{T\log [n(l+1)]}.
\]
The union bound followed by Lemma (ref) give us $\|\boldsymbol U\|_{\max} \lesssim_\P [\log (nT)]^{1/\gamma_2}$. Finally, we use (ref) once again to obtain
\begin{align}
\max_{i}\|\widehat{\boldsymbol E}_i\|_\infty&\lesssim_\P (1+\max_i\|\boldsymbol \theta_i\|_2) \sqrt{\log[n(l+1)]}T^{-1/2}\\
&\qquad +(1+\max_{i} \|\boldsymbol\theta_i\|_1)\left\{[\log(nT)]^{1/\gamma_2}\varrho_U + \varrho_U^2\right\}\nonumber:=\xi_0.
\end{align}
We now bound $\max_i\|\widehat{\boldsymbol\Sigma}_i-\boldsymbol\Sigma_i\|_{\max}$. Let $\widetilde{\boldsymbol \Sigma}_i:=\boldsymbol W_{i,\cdot}\boldsymbol W_{i,\cdot}'/T$. From the definition of $\boldsymbol W_{i,t}$ and Lemma (ref) we have that
\[\max_i\|\widehat{\boldsymbol\Sigma}_i-\widetilde{\boldsymbol\Sigma}_i\|_{\max}\leq \|\widehat{\boldsymbol U} - \boldsymbol U\|_{\max}(2\|\boldsymbol U\|_{\max} + \|\widehat{\boldsymbol U} - \boldsymbol U\|_{\max}).\]
Also by definition, we have that $\boldsymbol W_{i,t}\boldsymbol W_{i,t}' = g_i(\boldsymbol U_{\cdot,t},\dots, \boldsymbol U_{\cdot, t-l})$ for some function $g_i$ for each $i\in[n]$ and $t\in[T]$. Then, by Lemma (ref) we have that the strong mixing coefficients of the sequence $\{\boldsymbol W_{i,t}\boldsymbol W_{i,t}'\}_t$, $\check{\alpha}_{m}\leq \alpha_{(m-l)\lor 0}$. Let $W_{i,t}^{(j)}$ denote the $j$-th component of $\boldsymbol W_{i,t}$ for $j\in[d]$.
Under Assumption (ref)(c) and the Cauchy-Schwartz inequality we have
\[
{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert W_{i,t}^{(j_1)}W_{i,t}^{(j_2)}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/2+\epsilon/2}\leq {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert W_{i,t}^{(j_1)}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p+\epsilon}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert W_{i,t}^{(j_2)}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p+\epsilon}\leq \sup_t{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \boldsymbol U_{\cdot,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p+\epsilon}^2\leq C^2;\qquad j_1,j_2\in[d].
\]
Therefore, by Lemma (ref) followed by Lemma (ref), we conclude that for $p\in[4, \infty)$
\[\max_{i,j_1,j_2}\left|\sum_{t=1}^T W_{i,t}^{(j)} W_{i,t}^{(j_2)} -\mathbb{E}\left(W_{i,t}^{(j)} W_{i,t}^{(j_2)}\right)\right| \lesssim_\P \mathscr{L}_\alpha\left[n(l+1)\right]^{4/p}\sqrt{T},\]
where $\mathscr{L}_\alpha$ is defined above. Hence, by the triangle inequality
\begin{align}
\max_i\|\widehat{\boldsymbol\Sigma}_i-\boldsymbol\Sigma_i\|_{\max}&\leq \max_i\|\widetilde{\boldsymbol\Sigma}_i-\boldsymbol\Sigma_i\|_{\max} + \max_i\|\widehat{\boldsymbol\Sigma}_i-\widetilde{\boldsymbol\Sigma}_i\|_{\max}\\
&\lesssim_\P \frac{\mathscr{L}_\alpha\big[n(l+1)\big]^{4/p}}{\sqrt{T}} + (nT)^{1/p}\varrho_U + \varrho_U^2=:\xi_1\nonumber.
\end{align}
Similarly, under Assumption (ref)(d) and Lemma (ref) we have, for $j_1,j_2\in[d]$,
\[
{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert W_{i,t}^{(j_1)}W_{i,t}^{(j_2)}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma_2/2}}\lesssim {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert W_{i,t}^{(j_1)}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma_2}}\lor {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert W_{i,t}^{(j_2)}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma_2}}\leq C_{\gamma_2}\sup_t{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \boldsymbol U_{\cdot,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma_2}}^2\leq C_{\gamma_2}C.
\]
Thus, by Lemma (ref) we conclude that
\[\max_{i,j_1,j_2}\left|\sum_{t=1}^T W_{i,t}^{(j)} W_{i,t}^{(j_2)} -\mathbb{E}\left(W_{i,t}^{(j)} W_{i,t}^{(j_2)}\right)\right|\lesssim_\P \sqrt{T\log [n(l+1)]}.\]
Therefore, by the triangle inequality
\begin{align}
\max_i\|\widehat{\boldsymbol\Sigma}_i-\boldsymbol\Sigma_i\|_{\max}&\leq \max_i\|\widetilde{\boldsymbol\Sigma}_i-\boldsymbol\Sigma_i\|_{\max} + \max_i\|\widehat{\boldsymbol\Sigma}_i-\widetilde{\boldsymbol\Sigma}_i\|_{\max}\\
&\lesssim_\P \sqrt{\frac{\log [n(l+1)]}{T}} + [\log (nT)]^{1/\gamma_2}\varrho_U + \varrho_U^2=:\xi_1\nonumber.
\end{align}
\subsection{Proof of Theorems (ref) and (ref)}
Decompose the prediction error as
\begin{align*}
\widehat{Y}_{i,t}-\widetilde{Y}_{i,t}&=(\widehat{\boldsymbol\gamma}_i-\boldsymbol\gamma_i)'\boldsymbol X_{i,t} +\widehat{\boldsymbol\lambda}_i'\widehat{\boldsymbol P}\widehat{\boldsymbol G}_t - \boldsymbol\lambda_i'\boldsymbol P\boldsymbol G_t +\widehat{\boldsymbol\theta}_i' \widehat{\boldsymbol W}_{i,t} - \boldsymbol\theta_i'\boldsymbol W_{i,t} \\
&= (\widehat{\boldsymbol\gamma}_i-\boldsymbol\gamma_i)'\boldsymbol X_{i,t}+\big[\boldsymbol\lambda_i'\boldsymbol P + (\widehat{\boldsymbol\lambda}_i'\widehat{\boldsymbol P}-\boldsymbol\lambda_i'\boldsymbol P)\big](\widehat{\boldsymbol G}_t-\boldsymbol G_t)- (\widehat{\boldsymbol\lambda}_i'\widehat{\boldsymbol P}-\boldsymbol\lambda_i'\boldsymbol P)\boldsymbol G_t\\
&\qquad +\big[\boldsymbol\theta_i +(\widehat{\boldsymbol\theta}_i- \boldsymbol\theta_i) \big]' (\widehat{\boldsymbol W}_{i,t} - \boldsymbol W_{i,t}) + (\widehat{\boldsymbol\theta}_i- \boldsymbol\theta_i)'\boldsymbol W_{i,t},
\end{align*}
and
\begin{align*}
\widehat{\boldsymbol\lambda}_i'\widehat{\boldsymbol P}-\boldsymbol\lambda_i'\boldsymbol P &= \big[\boldsymbol\lambda_i +(\widehat{\boldsymbol\lambda}_i - \boldsymbol\lambda_i)\big](\widehat{\boldsymbol P} - \boldsymbol P) + (\widehat{\boldsymbol\lambda}_i - \boldsymbol\lambda_i)\boldsymbol P.
\end{align*}
For the first term we have
\begin{align*}
\max_{i\in[n]}\left|(\widehat{\boldsymbol\gamma}_i-\boldsymbol\gamma_i)'\boldsymbol X_{i,t}\right|\leq \max_{i\in[n]} \|\widehat{\boldsymbol\gamma}_i-\boldsymbol\gamma_i\|_1\max_{i\in[n]}\|\boldsymbol X_{i,t}\|_\infty \lesssim_\P \varrho_\gamma\max_{i\in[n]}\|\boldsymbol X_{i,t}\|_\infty).
\end{align*}
For the second term,
\[
\max_{i\in[n]}|\widehat{\boldsymbol\lambda}_i'\widehat{\boldsymbol P}-\boldsymbol\lambda_i'\boldsymbol P |\lesssim_\P (\varrho_ P + \varrho_\lambda)
\]
and, therefore,
\[
\max_{i\in[n]}|\widehat{\boldsymbol\lambda}_i'\widehat{\boldsymbol P}\widehat{\boldsymbol G}_t - \boldsymbol\lambda_i'\boldsymbol P\boldsymbol G_t|\lesssim_\P (\varrho_ P + \varrho_\lambda +\varrho_F).
\]
Finally, for the last term, Hölder's inequality yields
\begin{align*}
|\widehat{\boldsymbol\theta}_i' \widehat{\boldsymbol W}_{i,t} - \boldsymbol\theta_i'\boldsymbol W_{i,t}|
&\leq(\|\boldsymbol\theta_i\|_1 +\|\widehat{\boldsymbol\theta}_i- \boldsymbol\theta_i\|_1 ) \|\widehat{\boldsymbol W}_{i,t} - \boldsymbol W_{i,t}\|_\infty + \|\widehat{\boldsymbol\theta}_i- \boldsymbol\theta_i\|_1\|\boldsymbol W_{i,t}\|_\infty.
\end{align*}
As a consequence,
\[\max_{i\in[n]}|\widehat{\boldsymbol\theta}_i' \widehat{\boldsymbol W}_{i,t} - \boldsymbol\theta_i'\boldsymbol W_{i,t}|\lesssim_\P \max_{i\in[n]} \|\boldsymbol \theta_i\|_1\varrho_U + \varrho_\theta\max_{i\in[n]}\|\boldsymbol W_{i,t}\|_\infty
\]
and
\begin{align*}
\max_{i\in[n]}|\widehat{Y}_{i,t}-\widetilde{Y}_{i,t}|&\lesssim_\P \varrho_\gamma\max_{i\in[n]}\|\boldsymbol X_{i,t}\|_\infty + \varrho_U\max_{i\in[n]} \|\boldsymbol \theta_i\|_1 + \varrho_\theta\max_{i\in[n]}\|\boldsymbol W_{i,t}\|_\infty + \varrho_ P + \varrho_\lambda +\varrho_F.
\end{align*}
Finally, $\max_{i\in[n]}\|\boldsymbol X_{i,t}\|_\infty\lesssim_\P n^{1/p}$ in the polynomial case and $\max_{i\in[n]}\|\boldsymbol X_{i,t}\|_\infty\lesssim_\P (\log n)^{1/\gamma_2}$ in the exponential one. Similarly, $\max_{i\in[n]}\|\boldsymbol W_{i,t}\|_\infty\lesssim_\P d_W^{1/p}$ in the polynomial case and $\max_{i\in[n]}\|\boldsymbol W_{i,t}\|_\infty\lesssim_\P(\log d_W)^{1/\gamma_2}$ in the exponential one.
\subsection{Proof of Theorems (ref) and (ref)}
For $\mathcal{D}\in[n]^2$, let $\boldsymbol N_t:=(N_{1,t},\dots,N_{d,t})' :=\left[U_{i,t}U_{j,t} - \mathbb{E} \left(U_{i,t}U_{j,t}\right):(i,j)\in\mathcal{D}\right]$ and $\widehat{\boldsymbol N}_t:=\left[\widehat{U}_{i,t}\widehat{U}_{j,t} - \mathbb{E}\left(U_{i,t}U_{j,t}\right):(i,j)\in\mathcal{D}\right]$ where $d:=|\mathcal{D}|$. Then,
\[
\widetilde{\boldsymbol J} = \frac{1}{\sqrt{T}}\sum_{t\in[T]}\boldsymbol N_t\quad\textnormal{and}\quad \widehat{\boldsymbol J} = \frac{1}{\sqrt{T}}\sum_{t\in[T]}\widehat{\boldsymbol N}_t.
\]
For the polynomial case (Theorems (ref)), by Cauchy Schwartz inequality and Assumption (ref), \[
{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert U_{i,t}U_{j,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{(p+\epsilon)/2}\leq{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert U_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p+\epsilon}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert U_{j,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p+\epsilon}\leq C^2.
\]
Then, ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert N_{j,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{(p+\epsilon)/2}\leq 2C^2$ for $j\in[d]$; for the exponential case (Theorem (ref)). By Lemma (ref) and Assumption (ref), ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert U_{i,t}U_{j,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma_2/2}}\lesssim 1$. Then, ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert N_{j,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma_2/2}}\lesssim 1$ for $j\in[d]$. Also, the mixing coefficient of $\{\boldsymbol N_t,t\in[T]\}$ are upper bounded by those of $\{\boldsymbol U_t,t\in[T]\}$ by Lemma (ref).
Therefore, we can apply Theorem (ref)(a) to bound $\rho(\widetilde{\boldsymbol J},\boldsymbol{\mathcal{G}})$ with “$p=p/2$” and “$\epsilon= \epsilon/2$” for the polynomial case and the first result of Theorem (ref) follows. Similarly, we can apply Theorem (ref)(b) to bound $\rho(\widetilde{\boldsymbol J},\boldsymbol{\mathcal{G}})$ with “$\gamma_1=\gamma_1$” and “$\gamma_2= \gamma_2/2$” for the exponential case and the first result of Theorem (ref) follows.
By triangle inequality followed Lemma (ref) (applied twice) and Corollary 1 in cck_anti-concentration we have
\begin{align*}
\rho(\widehat{\boldsymbol J},\boldsymbol{\mathcal{G}})&\leq \rho(\widehat{\boldsymbol J}, \widetilde{\boldsymbol J}) + \rho(\widetilde{\boldsymbol J},\boldsymbol{\mathcal{G}})\\
&\lesssim \rho(\widetilde{\boldsymbol J},\boldsymbol{\mathcal{G}}) + \inf_{\delta_1>0}\left\{\Delta(\boldsymbol{\mathcal{G}},\delta_1) + \P(\|\widehat{\boldsymbol J} - \widetilde{\boldsymbol J}\|_\infty>\delta_1)\right\}\\
&\lesssim\rho(\widetilde{\boldsymbol J},\boldsymbol{\mathcal{G}}) + \inf_{\delta_1>0}\left\{\delta_1\sqrt{1\lor \log d} + \P(\|\widehat{\boldsymbol J} - \widetilde{\boldsymbol J}\|_\infty>\delta_1)\right\},
\end{align*}
where $\Delta(\cdot,\cdot)$ is defined by (ref). Recall that $\boldsymbol\Upsilon$ is the variance of $\boldsymbol N_t$ and $ \boldsymbol{\mathcal{G}}^*|\boldsymbol X,\boldsymbol Y\sim N(\boldsymbol 0,\widehat{\boldsymbol\Upsilon})$. Then, on the event $\{\|\widehat{\boldsymbol\Upsilon} - \boldsymbol\Upsilon\|_{\max}\leq \delta_2\}$ for $\delta_2>0$ by Theorem 1.1 in fang_koike (see also Lemma 2.1 in clt_hd for this particular application)
\[\sup_{t\in\mathbb{R}}|\P(\|\boldsymbol{\mathcal{G}}\|_\infty\leq t)-\P(\| \boldsymbol{\mathcal{G}}^*\|_\infty\leq t|\boldsymbol X,\boldsymbol Y)|\lesssim\delta_2 \log d(1\lor |\log d|),\]
and
\[\rho(\boldsymbol{\mathcal{G}}, \boldsymbol{\mathcal{G}}^*)\lesssim \inf_{\delta_2>0}\left\{\delta_2 \log d(1\lor |\log d|) + \P(\|\widehat{\boldsymbol\Upsilon} - \boldsymbol\Upsilon\|_{\max}>\delta_2)\right\}.\]
\subsection{Proof of Corollary (ref)}
Since $S^*|\boldsymbol X,\boldsymbol Y$ has no point-mass, we have that $\tau = \P(S^*\leq c^*(\tau)|\boldsymbol X,\boldsymbol Y)$ almost surely for every $\tau\in(0,1)$ and $\mathcal{D}\in[n]^2$. Then, under the Null, using the notation as in the proof of Theorem (ref), we write for $\delta_1,\delta_2>0$:
\begin{align*}
\sup_\mathcal{D}\sup_\tau\left|\P\left[S\leq c^*(\tau)\right] -\tau\right| &= \sup_\mathcal{D}\sup_\tau\left|\P\left[S\leq c^*(\tau)\right] -\P(S^*\leq c^*(\tau)\right|\\
&\leq \sup_\mathcal{D}\rho(\widehat{\boldsymbol J}, \boldsymbol Z^*)\\
&\lesssim \varrho_\Sigma + \delta_1\sqrt{1\lor \log n} + \sup_\mathcal{D}\P(\|\widehat{\boldsymbol J} - \widetilde{\boldsymbol J}\|_\infty>\delta_1)\\
&\quad + \delta_2 \log n(1\lor |\log n|) + \sup_\mathcal{D}\P(\|\widehat{\boldsymbol\Upsilon} - \boldsymbol\Upsilon\|_{\max}>\delta_2).
\end{align*}
Let $\gamma_1$ and $\gamma_2$ denote the rate appearing on Lemmas (ref) and (ref), respectively, with $\varrho_U$ equals to the rate of Collorary (ref) and $\mathcal{D}=[n]^2$ . By assumption, $\varrho_\Sigma\lor (\log n)^{3/2}\gamma_1\lor(\log n)^3\gamma_2 = o(1)$ as $T,n\to\infty$ when we set $\delta_1=\gamma_1\log n$ and $\delta_2=\gamma_2\log n$. Therefore, all the terms in the last expression vanish in probability, and the result follows.
\subsection{Proof of Theorems (ref) and (ref)}
The proof is parallel to the proof of Theorem (ref). We outline the main differences. For $\mathcal{D}\in[n]^2$, let $\boldsymbol K_t:=(K_{1,t},\dots,K_{d,t})' :=\left[V_{i,j,t}V_{j,i,t} - \mathbb{E} (V_{i,j,t}V_{j,i,t}):(i,j)\in\mathcal{D}\right]$ and $\widehat{\boldsymbol K}_t:=\left[\widehat{V}_{i,j,t}\widehat{V}_{j,i,t} - \mathbb{E} (V_{i,j,t}V_{j,i,t}):(i,j)\in\mathcal{D}\right]$ where $d:=|\mathcal{D}|$. Then,
\[
\widetilde{\boldsymbol Q} = \tfrac{1}{\sqrt{T}}\sum_{t\in[T]}\boldsymbol K_t\quad\textnormal{and}\quad \widehat{\boldsymbol Q} = \tfrac{1}{\sqrt{T}}\sum_{t\in[T]}\widehat{\boldsymbol K}_t.\]
For the polynomial case (Theorem (ref)), by Cauchy Schwartz inequality, the bound in (ref) with $\boldsymbol W_{i,t} = \boldsymbol U_{-ij,t}$, and Assumption (ref), we write \[
{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert V_{i,j,t}V_{j,i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{(p+\epsilon)/2}\leq{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert V_{i,j,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p+\epsilon}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert V_{j,i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p+\epsilon}\leq C^2_\chi,
\]
where $C_\chi:=(1+\max_{i,j\in\mathcal{D}}\|\boldsymbol \chi_{i,j}\|_2)C$. Then, ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert K_{j,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{(p+\epsilon)/2}\leq 2C^2_\chi$ for $j\in[d]$. Also, the mixing coefficient of $\{\boldsymbol K_t,t\in[T]\}$ are upper bounded by those of $\{\boldsymbol U_t,t\in[T]\}$ by Lemma (ref).
Therefore, we can apply Theorem (ref)(a) to bound $\rho(\widetilde{\boldsymbol Q},\boldsymbol{\mathcal{H}})$ with “$p=p/2$” and “$\epsilon= \epsilon/2$” and the first result of Theorem (ref) follows. The other two results can be obtained as in the proof do Theorem (ref), with $\widehat{\boldsymbol J}$, $\boldsymbol{\mathcal{G}}$, and $\boldsymbol{\mathcal{G}}^*$ replaced by $\widehat{\boldsymbol Q}$, $\boldsymbol{\mathcal{H}}$, and $\boldsymbol{\mathcal{H}}^*$, respectively.
For the exponential case (Theorem (ref)), we replace the Cauchy Schwartz inequality and Assumption (ref) used above by Lemma (ref) and Assumption (ref), respectively, and conclude that
\[
{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert V_{i,j,t}V_{j,i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma_2/2}}\lesssim{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert V_{i,j,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma_2}}^2\lor {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert V_{j,i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma_2}}^2\leq C^2_\chi.
\]
Finally, we apply Theorem (ref)(b) to bound $\rho(\widetilde{\boldsymbol Q},\boldsymbol{\mathcal{H}})$ with “$\gamma_1 = \gamma_1$” and “$\gamma_2= \gamma_2/2$” and the first of Theorem (ref) follows.
\subsection{Proof of Corollary (ref)}
The proof is identical to the proof of Corollary (ref) with Lemmas (ref) and (ref) replaced by Lemmas (ref) and (ref), respectively.
\section{High-Dimension Central Limit Theorem for Strong Mixing sequences}
The setup for this section is the following. Let $\{\boldsymbol X_1,\dots, \boldsymbol X_T\}$ be a sequence of zero-mean random vectors taking value in $\mathbb{R}^d$, whose strong mixing coefficient we denote by $\{\alpha_n:n\in\N\}$. Define $\boldsymbol S=T^{-1/2}\sum_{t=1}^T\boldsymbol X_t$ and let $\boldsymbol Z\sim N(0,\boldsymbol\Sigma)$ where $\boldsymbol\Sigma = \mathbb{E}(\boldsymbol S \boldsymbol S')$. The next result upper bounds
\[
\rho(\boldsymbol S, \boldsymbol Z):=\sup_{A\in\mathcal{R}}|\P( \boldsymbol S\in A) - \P(\boldsymbol Z \in A)|,
\]
where $\mathcal{R}$ is the class of all rectangles in the from $\bigtimes_{j=1}^d (a_j,b_j]$ for some $-\infty\leq a_j\leq b_j\leq \infty$ and $j\in[d]$. We also define for $s>0$
\begin{equation}
\Delta(\boldsymbol Z,s):= \sup_{z\in\mathbb{R}^d}\{\P(\boldsymbol Z\leq z+s) - \P(\boldsymbol Z\leq z)\}.
\end{equation}
Note that we may assume, without loss of generality, that all the elements on the diagonal of $\boldsymbol\Sigma$ are positive. Otherwise, we could exclude the entries of $\boldsymbol S$ (and $\boldsymbol Z)$ that are $0$ almost surely. Also, since $\rho(\boldsymbol D\boldsymbol S, \boldsymbol D\boldsymbol Z)=
\rho(\boldsymbol S, \boldsymbol Z)$ for any diagonal matrix $\boldsymbol D$, we may assume that all the elements in the diagonal of $\boldsymbol\Sigma$ are $1$, i.e., $\boldsymbol\Sigma$ is a correlation matrix. Finally, we assume that $\boldsymbol \Sigma$ is positive definite and denote by $\sigma_*^2\in(0,1)$ its smallest eigenvalue.
\begin{theorem}[High-Dimension Central Limit Theorem for Strong Mixing sequences]
\begin{enumerate}[(a)]
• \underline{Polynomial Case}: If for some $p\in[4,\infty)$ and $\epsilon>0$, ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p+\epsilon}\leq B$ for $i\in[d]$ and $t\in [T]$ for some constant $B$ that might dependent on $T$ and $p$; and $\alpha_n\leq K n^{-r}$ for $r\geq 0$ and a constant $K$ that might depend on $d$, then
\begin{align*}
\rho\left(\boldsymbol S,\boldsymbol Z\right)&\lesssim \frac{B^2}{\sigma_*^2}\left[\frac{\mathscr{A}_2}{\sqrt{T}} + \frac{\mathscr{A}_1}{T^{(1/2)-\kappa}} +\frac{K^{1-2/p}}{T^{(r/2)(1-2/p) -1}}\right]\log T\log d\\
&\qquad +\frac{(B\mathscr{A}_{\sqrt{T},4})^2(\log d)^{3/2}\log T}{T^{1/4}\sigma_*^2} + \frac{(B\mathscr{A}_{\sqrt{T},p})^2d^{2/p}(\log d)^2\log T}{T^{1/2-1/p}\sigma_*^2}\nonumber\\
&\qquad +\frac{\left[ B\mathscr{A}_{\sqrt{T},p} d(\log d)^{(3/2)p-4}\log T \log (dT) \right]}{T^{1/4}\sigma_*^{p/(p-2)}}^{1/(p-2)},\nonumber\\
&\qquad+ \frac{1}{T^{r\kappa -1/2}}\left[d^{1/(p+1)}(\mathscr{A}_{T^{\kappa},p}B)^{D_p}\sqrt{1\lor \log d}+ K \right]\nonumber,
\end{align*}
where $\kappa:=\frac{1/2 + D_p/4}{r+D_p/2}\land 1/2$, $D_p := \frac{p}{p+1}$, $\mathscr{A}_{k,p}:=\left[\sum_{n=0}^{k-1} (n+1)^{p/2-2}\alpha_n^{\epsilon/(p+\epsilon)}\right]^{\tfrac{2}{p}}$ for $k\in\N$, $\mathscr{A}_1:=\sum_{n=0}^{\lfloor T^\kappa \rfloor}\alpha_n^{1-2/p}$, and $\mathscr{A}_2:=\sum_{n=1}^{\lfloor \sqrt{T} \rfloor}n \alpha_n^{1-2/p}$. In particular, if $r>(\tfrac{p+\epsilon}{\epsilon})(p/2-1)\lor \tfrac{2}{1-2/p}$ then $\mathscr{A}_{k,p},\mathscr{A}_{1}$ and $\mathscr{A}_{2}$ can be bounded by constant that does not depend on $T$ or $d$.
• \underline{Exponential Case}: If $\{X_{i,t}: t\in[T]\}$ fulfills the conditions of Lemma (ref) for $i\in [d]$, then for $T\geq 16\lor (C\log d)^{1/\gamma -1/2}$, $d\geq 2$, and $d\sqrt{T}\geq \exp(1/\gamma) - 1 $,
\begin{align*}
\rho\left(\boldsymbol S_{X},\boldsymbol Z\right)&\lesssim \frac{(\log T)^{\gamma_1 + 1} \log d + \big[\log (dT)\big]^{2/\gamma}(\log d)^2\log T}{\sqrt{T}\sigma_*^2} \\
&\qquad + \frac{(\log d)^{2} +(\log d)^{3/2}\log T+ \log d (\log T)^{\gamma_1+1}\log (dT)}{T^{1/4}\sigma_*^2}.
\end{align*}
\end{enumerate}
\end{theorem}
\begin{proof}
We adapt the classical “big block-small block” technique commonly used to prove central limit theorems with dependence. Recall we assume $T\geq 4$ and consider two sequences of positive integers $a:=a_T$ and $b:=b_T$ such that $b\leq a$ and $a+b\leq T/2$. Let $m:=\lceil T/(a+b)\rceil$ and define for $j\in[m-1]$ consecutive blocks of size $a$ and $b$
with index set $\mathcal{A}_j:=\{((j-1)(a+b)+1,\dots, (j-1)(a+b)+a\}$ and $\mathcal{B}_j:=\{(j-1)(a+b)+a+1,\dots j(a+b)\}$.
Finally, set $\mathcal{A}_{m}:=\{m(a+b) +1, \dots, T\}$, which might be empty. Define,
\[
\boldsymbol A_j := \sum_{t\in \mathcal{A}_j}\boldsymbol X_t,\quad j\in[m]\qquad\textnormal{and}\qquad\boldsymbol B_j = \sum_{t\in\mathcal{B}_j} \boldsymbol X_t,\qquad j\in[m-1]
\]
such that
\[
\boldsymbol S:= \boldsymbol S_X:=\tfrac{1}{\sqrt{T}}\sum_{t=1}^T \boldsymbol X_t =\tfrac{1}{\sqrt{T}}\sum_{j=1}^{m} \boldsymbol A_j + \tfrac{1}{\sqrt{T}}\sum_{j=1}^{m-1} \boldsymbol B_j=:\boldsymbol S_A + \boldsymbol S_B.
\]
Also, let $\{\widetilde{\boldsymbol A}_j:j\in[m]\}$ be an independent sequence such that $\boldsymbol A_j$ and $\widetilde{\boldsymbol A}_j$ have the same distribution for $j\in[m]$, and define $\boldsymbol S_{\widetilde{A}} := T^{-1/2}\sum_{j=1}^{m} \widetilde{\boldsymbol A}_j$.
We start by applying Lemma (ref) with $X=\boldsymbol S_X$, $Y=\boldsymbol S_A$ and $d(x,y)=\|x-y\|_\infty$ to obtain, for $A\in\mathcal{R}$ and $s>0$,
\begin{align*}
|\P(\boldsymbol S_X\in A)- \P(\boldsymbol S_A\in A)|\leq \P(\|\boldsymbol S_B\|_\infty>s) + \P(\boldsymbol S_A\in A^s\setminus A)\lor \P(\boldsymbol S_A\in A\setminus A^{-s}).
\end{align*}
The right hand side of the last expression can be upper bounded as
\begin{align*}
\P(\boldsymbol S_A\in A^s\setminus A) &= \P(\boldsymbol Z\in A^s\setminus A) + \big[\P(\boldsymbol S_A\in A^s\setminus A)-\P(\boldsymbol Z\in A^s\setminus A)\big]\\
&\leq \P(\boldsymbol Z\in A^s\setminus A)+2\rho(\boldsymbol S_A,\boldsymbol Z)\\
&\leq \P(\boldsymbol Z\in A^s\setminus A)+2\rho(\boldsymbol S_A,\boldsymbol S_{\widetilde{A}}) + 2\rho(\boldsymbol S_{\widetilde{A}},\boldsymbol Z).
\end{align*}
We can proceed similarly to conclude that $\P(\boldsymbol S_A\in A\setminus A^{-s})\leq \P(\boldsymbol Z\in A\setminus A^{-s})+2\rho(\boldsymbol S_A,\boldsymbol S_{\widetilde{A}}) + 2\rho(\boldsymbol S_{\widetilde{A}},\boldsymbol Z)$. Since $\P(\boldsymbol Z\in A^s\setminus A)\lor\P(\boldsymbol Z\in A\setminus A^{-s})\leq 2\Delta(\boldsymbol Z, s)$, we can take the supremum over $A\in\mathcal{R}$ to write, for every $s>0$,
\begin{align*}
\rho(\boldsymbol S_X,\boldsymbol S_A)\leq \P(\|\boldsymbol S_B\|_\infty>s)+2\Delta(\boldsymbol Z, s) +2\rho(\boldsymbol S_A,\boldsymbol S_{\widetilde{A}}) + 2\rho(\boldsymbol S_{\widetilde{A}},\boldsymbol Z).
\end{align*}
By the triangle inequality,
\begin{align*}
\rho\left(\boldsymbol S_X,\boldsymbol Z\right)&\leq \rho\left(\boldsymbol S_X,\boldsymbol S_A\right) + \rho\left(\boldsymbol S_A,\boldsymbol S_{\widetilde{A}}\right) +\rho\left(\boldsymbol S_{\widetilde{A}},\boldsymbol Z\right).
\end{align*}
Finally, combine the last two displays we obtain the inequality
\begin{align}
\rho\left(\boldsymbol S_X,\boldsymbol Z\right)&\leq 3\rho\left(\boldsymbol S_{\widetilde{A}},\boldsymbol Z\right) +2\Delta(\boldsymbol Z, s) + \P\left(\|\boldsymbol S_B \|_\infty>s\right) + 3 \rho\left(\boldsymbol S_A,\boldsymbol S_{\widetilde{A}}\right).
\end{align}
We now proceed to bound each of the terms on the right-hand side.
\textbf{Step 1.} Recall that $\{\widetilde{\boldsymbol A}_j:j\in[m]\}$ is a zero-mean independent sequence, so we deal with the first term in (ref) applying Theorem 2.2 in clt_hd in the form
\[
\boldsymbol S_{\widetilde{A}}:=\frac{1}{\sqrt{T}}\sum_{j=1}^m\widetilde{\boldsymbol A}_j = \frac{1}{\sqrt{m}}\sum_{j=1}^m\boldsymbol W_j\quad\textnormal{and}\quad \boldsymbol W_j:=\sqrt{\frac{m}{T}}\widetilde{\boldsymbol A}_j,
\]
which give us
\begin{align}
\rho\left(\boldsymbol S_{\widetilde{A}},\boldsymbol Z\right)&\lesssim \log m\left[\frac{(\log d) M_\Sigma}{\sigma_*^2} +\frac{(\log d)^{3/2}\sqrt{\mu_4}}{m\sigma_*^2} + \frac{(\mathcal{M}\log d)^2}{m\sigma_*^2}\right]\\
&\qquad + \inf_{x>0}\left[\frac{\log d\sqrt{\log m \log (dm)H(x)}}{\sqrt{m}\sigma_*^2} + \frac{x(\log d)^{3/2}}{\sqrt{m}\sigma_*}\right],\nonumber
\end{align}
where $M_\Sigma:=\|\boldsymbol\Sigma_{\boldsymbol S_{X}} -\boldsymbol\Sigma_{\boldsymbol S_{\widetilde{A}}}\|_{\max}$, $\mu_4 := \max_{i\in[d]}\sum_{j=1}^m{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert W_{j,i}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_4^4$, $\mathcal{M}:={\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \max_{j,i} |W_{j,i}|
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_4$ and $H(x):=\max_{j\in[m]}\mathbb{E}\|\boldsymbol W_j\|_\infty^4\1\{\|\boldsymbol W_j\|_\infty>x\}$ for $x>0$.
From Lemma (ref) we have, for $p\in[2,\infty)$ and $j\in[m]$, ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \widetilde{ A}_{j,i}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p}={\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert A_{j,i}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p}={\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \sum_{t\in\mathcal{A}_j}X_{t,i}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p}\lesssim\sqrt{a+b}\mathscr{A}_{a,p} L_{p}\lesssim\sqrt{a}\mathscr{A}_{a,p} L_{p}$ where $L_p:=\max_{i\in[d]}\max_{t\in [T]}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X_{t,i}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p+\epsilon}$, $W_{j,i}$ and $X_{t,i}$ denotes the $i$-th element of $\boldsymbol W_j$ and $\boldsymbol X_t$ respectively, for $i\in[d]$. Thus, for $i\in[d]$, $j\in[m]$ and $p\in[2,\infty)$.
\[{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert W_{j,i}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_p\lesssim L_p\mathscr{A}_{a,p}\sqrt{\frac{ma}{T}}\leq L_p\mathscr{A}_{a,p}=:R_p.\]
Expression (1,12a) and (1.12b) in rio2017asymptotic give us, for $t,s\in[T]$, $i,k\in[d]$ and $p\in[1,\infty]$,
\[
\left|\mathbb{E}(X_{t,i}X_{s,k})\right| \leq 2\alpha_{|t-s|}^{1-2/(p+\epsilon)}\|X_{t,i}\|_{p+\epsilon}\|X_{t,k}\|_{p+\epsilon}\leq 2L_p^2\alpha_{|t-s|}^{1-2/p}.
\]
Recall that $\boldsymbol S_{\widetilde{A}}^G\sim N(\boldsymbol 0,\boldsymbol\Sigma_{\boldsymbol S_{\widetilde{A}}})$ and $\boldsymbol Z\sim N(\boldsymbol 0,\boldsymbol\Sigma_{\boldsymbol S_{X}})$, where
\[
\boldsymbol\Sigma_{\boldsymbol S_{\widetilde{A}}}=\frac{1}{T}\sum_{j=1}^m \sum_{t\in\mathcal{A}_j} \sum_{s\in\mathcal{A}_j} \mathbb{E}(\boldsymbol X_t \boldsymbol X_s')\quad\textnormal{and}\quad \boldsymbol\Sigma_{\boldsymbol S_{X}}=\frac{1}{T} \sum_{t=1}^T \sum_{s=1}^T \mathbb{E}( \boldsymbol X_t \boldsymbol X_s').
\]
Define $\boldsymbol\Delta_\ell:= \sum_{i=1}^{a-1}\sum_{j=1}^{a-1} \mathbb{E} (X_{i+\ell} X'_{a+j+\ell})$ for $\ell\geq 1$, then using the bound above we have
\[\|\boldsymbol\Delta_\ell\|_{\max}\leq 2L_p^2\sum_{m=1}^{a-1} m \alpha_m^{1-2/p},\quad \]
Then,
\begin{align*}
M_\Sigma & \leq \tfrac{4 m}{T} \max_\ell\|\boldsymbol\Delta_\ell\|_{\max} + \frac{m}{T}\max_{j\in[m]}\|\mathbb{E} (\boldsymbol B_j\boldsymbol B_j')\|_{\max} + \frac{1}{T}\sum_{|t-s|>a} \mathbb{E} (\boldsymbol X_t \boldsymbol X_s') \\
& \lesssim \frac{L_p^2}{T}\left[m\mathscr{A}_2 + mb\mathscr{A}_1(b) + T^2\alpha_a^{1-2/p}\right],
\end{align*}
where $\mathscr{A}_1(m):=\sum_{n=0}^{m-1}\alpha_n^{1-2/p}$ for $m\geq 1$ and $\mathscr{A}_2:=\sum_{n=1}^{a-1}n \alpha_n^{1-2/p}$.
Also,
\begin{align*}
\mu_4 &\leq m\max_{i\in[d]}\max_{j\in[m]}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert W_{j,i}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_4^4\lesssim m R_4^4\\
\mathcal{M}&\leq {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \max_{j,i} |W_{j,i}|
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_p\leq \left(md\max_{j,i} |W_{j,i}|^p\right)^{1/p}\lesssim R_p(md)^{1/p}\\
H(x) &\leq \max_{j\in[m]}\mathbb{E}\|\boldsymbol W_j\|_\infty^p/x^{p-4}\lesssim R_p^p d/x^{p-4}
\end{align*}
Plug these bounds back into (ref) and set $x=\left[\frac{\log m \log (dm) R_p d}{\sigma_*^2\log d}\right]^{1/(p-2)}$ to equate the terms inside the infimum. As a consequence, we are left with
\begin{align}
\rho\left(\boldsymbol S_{\widetilde{A}},\boldsymbol Z\right)&\lesssim \frac{L_p^2\left[m\mathscr{A}_2 + mb\mathscr{A}_1(b) + T^2\alpha_a^{1-2/p}\right]\log m \log d}{T\sigma_*^2}\\
&\qquad +\frac{R_4^2(\log d)^{3/2}\log m}{\sqrt{m}\sigma_*^2} + \frac{R_p^2(md)^{2/p}(\log d)^2\log m}{m\sigma_*^2}\nonumber\\
&\qquad +\frac{\left[ R_p d(\log d)^{(3/2)p-4}\log m \log (dm) \right]}{\sqrt{m}\sigma_*^{p/(p-2)}}^{1/(p-2)}.\nonumber
\end{align}
\textbf{Step 2.} By Nazarov's inequality for Gaussian random vectors (Theorem 1 in nazarov) we can bound the second term in (ref) as
\[
\Delta\left(\boldsymbol Z,s\right)\lesssim s\sqrt{ 1\lor \log d}.
\]
Similarly to Step 1, Lemma (ref) give us for $j\in[m]$, $i\in[d]$ and $p\in[2,\infty)$
\[{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \widetilde{ B}_{j,i}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p}={\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert B_{j,i}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p}={\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \sum_{t\in\mathcal{B}_j}X_{t,k}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p}\lesssim\sqrt{b}\mathscr{A}_{b,p} L_p,\]
where $B_{j,k}$ and $N_{t,k}$ denotes the $i$-th element of $\boldsymbol B_j$ and $\boldsymbol X_t$ respectively, for $i\in[d]$; and $\mathscr{A}_{b}:=\left[\sum_{0\leq m <b} (m+1)^{p/2-2}\alpha_m^{\epsilon/(p+\epsilon)}\right]^{\tfrac{2}{p}}$. We can then bound the third term in (ref) by Markov's inequality as
\[\P\left(\|\boldsymbol S_B\|_\infty>s\right) \leq d
\left(\frac{\sqrt{mb/T}\mathscr{A}_{b,p} L_{p}}{s}\right)^{p}.\]
We can set $s=s^*:=\big[d^{1/p}\sqrt{mb/T}\mathscr{A}_{b,p} L_{p}\big]^{p/(p+1)}$ to equate (up to log terms) the two terms above containing $s$ and obtain
\begin{equation*}
\Delta\left(\boldsymbol Z,s\right) + \P\left(\|\boldsymbol S_B\|_\infty>s\right) \lesssim \left(d^{1/p}\sqrt{\tfrac{mb}{T}}\mathscr{A}_{b,p} L_{p}\right)^{\tfrac{p}{p+1}}\sqrt{1\lor \log d}.
\end{equation*}
\textbf{Step 3.} Notice that any measurable $A\subseteq\mathbb{R}^2$ we have $|\P[(\boldsymbol A_1, \boldsymbol A_2)\in A] - \P[\widetilde{\boldsymbol A}_1,\widetilde{\boldsymbol A}_2\in A]|\leq \alpha_b$ where $\{\alpha_n,n\in \N\}$. Then, the last term in (ref) can be upper bounded by $(m-1)\alpha_b$ by induction, i.e,
\begin{equation*}
\rho\left(\boldsymbol S_A,\boldsymbol S_{\widetilde{A}}\right) \leq (m-1)\alpha_b.
\end{equation*}
\textbf{Step 4.} Applying
the bound derived in the steps above back into (ref) we are left with
\begin{align}
\rho\left(\boldsymbol S_{X},\boldsymbol Z\right)&\lesssim \frac{L_p^2\left[m\mathscr{A}_2 + mb\mathscr{A}_1(b) + T^2\alpha_a^{1-2/p}\right]\log m \log d}{T\sigma_*^2}\\
&\qquad +\frac{R_4^2(\log d)^{3/2}\log m}{\sqrt{m}\sigma_*^2} + \frac{R_p^2(md)^{2/p}(\log d)^2\log m}{m\sigma_*^2}\nonumber\\
&\qquad +\frac{\left[ R_p d(\log d)^{(3/2)p-4}\log m \log (dm) \right]}{\sqrt{m}\sigma_*^{p/(p-2)}}^{1/(p-2)},\nonumber\\
&\qquad+ \left(d^{1/p}\sqrt{\tfrac{mb}{T}}\mathscr{A}_{b,p} L_{p}\right)^{\tfrac{p}{p+1}}\sqrt{1\lor \log d} + m\alpha_b\nonumber\\
&=:(I) + (II) + (III) + (IV) + (V) + (VI)\nonumber.
\end{align}
We now must choose sequences $a$ and $b$ to balance all the terms appearing on the right-hand side. Let $a =1\lor \lfloor T^{\gamma_a}\rfloor$ and $b =1\lor \lfloor T^{\gamma_b}\rfloor$ for $0\leq \gamma_b\leq \gamma_a <1$ . Then $m\asymp T^{1-\gamma_a}$, $ma/T\asymp 1$, $mb/T\asymp T^{\gamma_b-\gamma_a}$. Comparing $(I)$, $(II)$ and $(IV)$, it seems we cannot do much better than setting $\gamma_a=\gamma_a^*:=1/2$.
Since $\alpha_n \leq K n^{-r}$ then, $m\alpha_b\lesssim K T^{1-\gamma_{a}^*-r\gamma_b} = KT^{1/2-r\gamma_b}$. To equate $(V)$ with $(VI)$ (up to $\log$ terms) and respect the fact that $0\leq \gamma_b\leq \gamma_a^*$ we set
\[
\gamma_b = \gamma_b^*:= \left(\frac{1/2+D_p/4}{r+D_p/2}\right)\land 1/2;\qquad D_p:=\frac{p}{p+1}.
\]
Then,
\begin{equation*}
(V) + (VI) \lesssim \frac{1}{T^{\phi(p,r)}}\left[d^{1/(p+1)}(\mathscr{A}_{b,p} L_{p})^{D_p}\sqrt{1\lor \log d}+ K\right],
\end{equation*}
where
\[\phi(r,p) := \gamma_a^* + r\gamma_b^* - 1=\begin{cases}
\tfrac{D_p(r-1)}{4(r+D_p/2)} &; r\geq 1\\
\frac{r-1}{2} &; 0\leq r< 1.
\end{cases} \]
Therefore,
\begin{align}
\rho\left(\boldsymbol S_{X},\boldsymbol Z\right)&\lesssim \frac{L_p^2}{\sigma_*^2}\left[\frac{\mathscr{A}_2}{\sqrt{T}} + \frac{\mathscr{A}_1(T^{\gamma_b^*})}{T^{(1/2)-\gamma_b^*}} +\frac{K_d^{1-2/p}}{T^{(r/2)(1-2/p) -1}}\right]\log T\log d\\
&\qquad +\frac{(L_4\mathscr{A}_{\sqrt{T},4})^2(\log d)^{3/2}\log T}{T^{1/4}\sigma_*^2} + \frac{(L_p\mathscr{A}_{\sqrt{T},p})^2d^{2/p}(\log d)^2\log T}{T^{1/2-1/p}\sigma_*^2}\nonumber\\
&\qquad +\frac{\left[ L_p\mathscr{A}_{\sqrt{T},p} d(\log d)^{(3/2)p-4}\log T \log (dT) \right]}{T^{1/4}\sigma_*^{p/(p-2)}}^{1/(p-2)},\nonumber\\
&\qquad+ \frac{1}{T^{\phi(p,r)}}\left[d^{1/(p+1)}(\mathscr{A}_{T^{\gamma_b^*},p} L_{p})^{D_p}\sqrt{1\lor \log d}+ K \right]\nonumber,
\end{align}
which concludes the proof of part (a).
The proof for part (b) runs parallel to the proof of part (a). In step 1, we replace Lemma (ref) by Lemma (ref) to conclude that ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert A_{i,j}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}\lesssim \sqrt{a}$ and thus ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert W_{i,j}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}\lesssim 1$, which in turn allow us to obtain
\begin{align*}
M_\Sigma & \lesssim \frac{m}{T}+ \frac{mb}{T} + T\exp\left[-K_1(1-2/p)a^{\gamma_1}\right];\qquad p\in[2,\infty)\\
\mu_4 &\leq m\max_{i\in[d]}\max_{j\in[m]}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert W_{j,i}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_4^4\lesssim m.
\end{align*}
We apply Lemma (ref) (b) followed (d) to write
\begin{align*}
\mathcal{M}&\lesssim {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \max_{j,i} |W_{j,i}|
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}\lesssim \psi_{e^\gamma}^{-1}(md)\max_{j,i} {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert W_{j,i}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}\lesssim \psi_{e^\gamma}^{-1}(md).
\end{align*}
Similarly, we obtain $\mathbb{E}\|\boldsymbol W_j\|_\infty^p = {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \max_{i\in[d]}|W_{j,i}|
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_p^p\lesssim {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \max_{i\in[d]}|W_{j,i}|
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma}}^p\lesssim \big[\psi_{e^\gamma}^{-1}(d)\big]^p$. Set $x=K_x\sqrt{\log d}$ for some $K_x>0$ large enough such that, by Lemma (ref), we have for $d\geq 2$ and $a\geq 4\lor K_x(\log d)^{2/\gamma-1}$,
\[
\P(\|\boldsymbol W_j\|_\infty>x)=\P\left(\left\|\sum_{t\in\mathcal{A}_j} \boldsymbol X_t\right\|_\infty>K_x\sqrt{\frac{T}{ma}}\sqrt{a\log d}\right)\leq (da)^{1-K_x^\gamma} + 2d^{1-K^2_x}.
\]
Then, by Cauchy-Schwartz inequality,
\[
H(x)\leq \max_{j\in[m]}\sqrt{\mathbb{E}\|\boldsymbol W_j\|_\infty^8\P(\|\boldsymbol W_j\|_\infty>x)}\lesssim \left[\psi_{e^\gamma}^{-1}(d)\right]^4\left[(da)^{1-K_x^\gamma} + 2d^{1-K_x^2}\right]^{1/2}
\]
Plug these bounds back into (ref), we are left with
\begin{align}
\rho\left(\boldsymbol S_{\widetilde{A}},\boldsymbol Z\right)&\lesssim \left\{\frac{m}{T} + \frac{mb}{T} + T\exp\left[-K_1(1-2/p)a^{\gamma_1}\right]\right\}\frac{\log m \log d}{\sigma_*^2}\\
&\qquad +\frac{(\log d)^{3/2}\log m}{\sqrt{m}\sigma_*^2} + \frac{\big[\psi_{e^\gamma}^{-1}(md)\big]^{2}(\log d)^2\log m}{m\sigma_*^2}\nonumber\\
&\qquad +\frac{\log d\sqrt{\log m \log (dm)}\big[\psi_{e^\gamma}^{-1}(d)\big]^2\big[(da)^{1-K^\gamma} + 2d^{1-K^2}\big]^{1/4}}{\sqrt{m}\sigma_*^2}\nonumber\\
&\qquad + \frac{(\log d)^{2}}{\sqrt{m}\sigma_*}.\nonumber
\end{align}
For step 2, we can then bound the third term in (ref) by Lemma (ref). Set $s=K_s\sqrt{\frac{mb}{T} \log d \log T}$ for some $K_s>0$ large enough so that, for $d\geq 2$, and $mb\geq 4\lor K_s(\log d)^{2/\gamma-1} $, we have
\begin{align*}
\P\left(\|\boldsymbol S_B\|_\infty>s\right)&=\P\left(\left\|\sum_{j=1}^{m-1}\sum_{t\in\mathcal{B}_j} \boldsymbol X_t\right\|_\infty>K_s\sqrt{mb\log T\log d}\right)\\
&\leq (dmb)^{1-K_s^\gamma(\log T)^{\gamma/2}} + 2d^{1-K_s^2\log T}.
\end{align*}
By plugging the bounds above back into (ref) we obtain
\begin{align}
\rho\left(\boldsymbol S_{X},\boldsymbol Z\right)&\lesssim \left\{\frac{m}{T} + \frac{mb}{T} + T\exp\left[-K_1(1-2/p)a^{\gamma_1}\right]\right\}\frac{\log m \log d}{\sigma_*^2}\\
&\qquad +\frac{(\log d)^{3/2}\log m}{\sqrt{m}\sigma_*^2} + \frac{\big[\psi_{e^\gamma}^{-1}(md)\big]^{2}(\log d)^2\log m}{m\sigma_*^2}\nonumber\\
&\qquad +\frac{\log d\sqrt{\log m \log (dm)}\big[\psi_{e^\gamma}^{-1}(d)\big]^2\big[(da)^{1-K^\gamma_x} + 2d^{1-K_x^2}\big]^{1/4}}{\sqrt{m}\sigma_*^2}\nonumber\\
&\qquad + \frac{(\log d)^{2}}{\sqrt{m}\sigma_*} + (dmb)^{1-K_s^\gamma(\log T)^{\gamma/2}} + 2d^{1-K_s^2\log T}.\nonumber\\
&\qquad + \log d\sqrt{\tfrac{mb}{T} \log T} + m\exp(-K_1b^{\gamma_1})\nonumber.
\end{align}
For the same reason discussed in the polynomial case, we cannot do much better than setting $a \asymp T^{1/2}$. Then, $m\asymp T^{1/2}$. Set $b\asymp (\log T)^{\gamma_1}$ so that the last term is at most $T^{-1/4}$. Note that the numerator of the fourth term above can be bounded uniformly in $d$ and $T$ by a constant only depending on $\gamma$ and the choice of $K_x>1$. Also, $\left(\frac{1}{d}\right)^{K^2_s\log T-1}\lesssim T^{-1}$ provided that $K_s>1$ as $d\geq 2$, and $\left[\frac{1}{d\sqrt{T}(\log T)^{\gamma_1}}\right]^{K^\gamma(\log T)^{\gamma/2}-1}\lesssim T^{-1/2}$ by taking $K_s\geq 2^{1/\gamma}$. Finally, $\psi_{e^\gamma}^{-1}(x) \lesssim (\log x )^{1/\gamma}$, for $x\geq \exp(1/\gamma)-1$ by Lemma (ref)(a). Result (b) follows.
\begin{align}
\rho\left(\boldsymbol S_{X},\boldsymbol Z\right)&\lesssim \frac{(\log T)^{\gamma_1 + 1} \log d}{\sqrt{T}\sigma_*^2}\\
&\qquad +\frac{(\log d)^{3/2}\log T}{T^{1/4}\sigma_*^2} + \frac{\big[\psi_{e^\gamma}^{-1}(d\sqrt{T})\big]^{2}(\log d)^2\log T}{\sqrt{T}\sigma_*^2}\nonumber\\
&\qquad +\frac{\log d\sqrt{\log T \log (d\sqrt{T})}\big[\psi_{e^\gamma}^{-1}(d)\big]^2\big[(d\sqrt{T})^{1-K^\gamma} + 2d^{1-K^2}\big]^{1/4}}{T^{1/4}\sigma_*^2}\nonumber\\
&\qquad + \frac{(\log d)^{2}}{T^{1/4}\sigma_*} + \left(\frac{1}{d\sqrt{T}(\log T)^{\gamma_1}}\right)^{K^\gamma(\log T)^{\gamma/2}-1} + \left(\frac{1}{d}\right)^{K^2\log T-1}.\nonumber\\
&\qquad + \frac{\log d (\log T)^{\gamma_1+1}}{T^{1/4}} \log\left(1\lor d\right) +\frac{1}{T^{1/4}}\nonumber.
\end{align}
Note that the numerator of the fourth term above can be bounded uniformly in $d$ and $T$ by a constant only depending on $\gamma$ and the choice of $K>1$. Also, $\left(\frac{1}{d}\right)^{K^2\log T-1}\lesssim T^{-1}$ as $d\geq 2$ and $K>1$, and $\left[\frac{1}{d\sqrt{T}(\log T)^{\gamma_1}}\right]^{K^\gamma(\log T)^{\gamma/2}-1}\lesssim T^{-1/2}$ by taking $K\geq 2^{1/\gamma}$. Finally, $\psi_{e^\gamma}^{-1}(x) \lesssim (\log x )^{1/\gamma}$ for $x\geq \exp(1/\gamma)-1$ by Lemma (ref)(a). Result (b) follows.
\end{proof}
\section{Auxiliary Lemmas}
\subsection{Concentration Inequalities for strong mixing sequences}
\begin{lemma} Let $\{\boldsymbol X_t: t\in\Z\}$ be a sequence of $d$-dimensional random vector with strong mixing coefficient $\{\alpha^{\boldsymbol X}_n:n\in\N\}$. For a non-negative integer $k$, define $\{\boldsymbol Y_t:=f(\boldsymbol X_t,\dots, \boldsymbol X_{t-k}): t\in\Z\}$ for some measurable $f:\mathbb{R}^{d(k+1)}\to \mathbb{R}^q$ and denote its strong mixing coefficient by $\{\alpha^{\boldsymbol Y}_n:n\in\N\}$. Then $\alpha^{\boldsymbol Y}_n\leq \alpha^{\boldsymbol X}_{(n-k)\lor 0}$ for $n\in\N$.
\end{lemma}
\begin{proof}
Let $\mathcal{X}_s^t$ and $\mathcal{Y}_s^t$ be the $\sigma$-algebra generated by $(X_s,\dots, X_t)$ and $(Y_s,\dots, Y_t)$, respectively where $-\infty\leq s\leq t\leq \infty$. Since $\mathcal{Y}_{-\infty}^t\subseteq \mathcal{X}_{-\infty}^t$ and $\mathcal{Y}_{t+n}^\infty\subseteq \mathcal{X}_{t+n-k}^\infty$, we have $\alpha^{\boldsymbol Y}_n\leq \alpha^{\boldsymbol X}_{n-k}$ for $n\geq k$ and $\alpha^{\boldsymbol Y}_n\leq \alpha^{\boldsymbol X}_0$ for $0\leq n <k$.
\end{proof}
\begin{lemma}[based on Corollary 1.1 and Theorem 6.3 in rio2017asymptotic] Let $S_T = \sum_{t=1}^T X_t$ where $\{X_t:t\in [T]\}$ is a sequence of zero mean real-valued random variables such that ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X_t
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_p<\infty$ for $t\in[T]$, and the mixing coefficients given by $\{\alpha_m:0\leq m <T\}$. Then, for $p\in[2,\infty)$ and $T\in\N$,
\[{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert S_T
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_p\leq a_p\sqrt{T}L_{2,\alpha} + b_p T^{1/p}L_{p,\alpha},\]
where $L_{s,\alpha}:=\left\{\int_0^1[\alpha^{-1}(u)\land T]^{s-1} Q^{s}(u){ du}\right\}^{1/s}$ for $s>0$, $a_p$ and $b_p$ are positive constants only depending on $p$, $\alpha^{-1}(u):=\sum_{0\leq m<T}\1\{u\leq \alpha_m\}$, $Q:=\max_{t\in[T]} Q_t$ and $Q_t(u):=\sup\{x\in\mathbb{R}:\P(|X_t|>x)<u\}$.
If further, $X_t\in\mathcal{L}^q$ for $t\in[T]$ for some $q\in(p, \infty]$, then
\[{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert S_T
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_p\leq c_{p,q} \mathscr{A}_{p,q}\max_{t\in[T]}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X_t
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_q\sqrt{T},\]
where $\mathscr{A}_{p,q}:=\left[\sum_{0\leq m <T} (m+1)^{p-2}\alpha_m^{1-p/q}\right]^{\tfrac{1}{p}}$ and $c_{p,q}$ is a positive constant only depending on $p$ and $q$.
If further, $\alpha_{m}\leq (m+1)^{-r}$ for some constant $r\geq 0$ and all $m\in\N$ then
\[{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert S_T
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_p\leq d_{p,q,r}\mathscr{R}_{p,q,r}(T)\max_{t\in[T]}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X_t
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_q;\qquad \mathscr{R}_{p,q,r}(T):=
\begin{cases}
\sqrt{T} &; r>\nu\\
(\log T+ \lambda)^{\tfrac{1}{p}}\sqrt{T} &; r=\nu\\
T^{[\frac{1}{2}+\frac{p-1 -r(1-p/q)}{p}]\land 1} &; r<\nu,
\end{cases}
\]
where $\nu:=\frac{(p-1)}{1-p/q}$ and $d_{p,q,r}$ is a positive constant only depending on $p$, $q$ and $r$ and $\lambda\approx 1.1$.
\end{lemma}
\begin{proof}
The first result, for $p=2$, follows from Corollary 1.1 in rio2017asymptotic by setting $a_2=b_2 = 1$ and bounding
\[
\sum_{t=1}^T\int_0^1[\alpha^{-1}(u)\land T] Q^2_t(u){ du}\leq T\int_0^1[\alpha^{-1}(u)\land T] \left(\max_{t\in[T]}Q_t\right)^2(u){ du}.
\]
For $p>2$ we rely on Theorem 6.3 in rio2017asymptotic, the concavity of $x\mapsto x^{1/p}$ and the result for $p=2$ to write ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert S_T
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_p\leq a_p\sqrt{T}L_{2,\alpha} + b_p T^{1/p}L_{p,\alpha}$ with $a_p = 2^{1+2(p+1)/p}p^{1/p}\sqrt{p+1}$ and $b_p=\left[\frac{p}{p-1}4^{p+1}(p+1)^{p-1}\right]^{1/p}$.
For the second result, Markov's inequality give us $\P(|X_t|\geq x)\leq\left({\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X_t
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_q/x\right)^q$ for $x>0$. Then, $Q_t(u)\leq \frac{{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X_t
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_q}{u^{1/p}}$ and $Q(u):=\max_{t\in[t]}Q(u)\leq \frac{\mu_q}{u^{1/p}}$, where $\mu_q:=\max_{t\in[T]}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X_t
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_q$. Then, for $s\in[2,p]$, we have
\begin{align*}
L_{s,q}\leq \mu_q\left\{\int_0^1[\alpha^{-1}(u)\land T]^{s-1} u^{-s/q}\mathsf{d} u\right\}^{1/s}\leq \mu_q\left\{\int_0^1[\alpha^{-1}(u)\land T]^{p-1} u^{-p/q}\mathsf{d} u\right\}^{1/s}.
\end{align*}
Combine equation (C.10) in in rio2017asymptotic, the bound $Q(u)\leq \frac{\mu_q}{u^{1/p}}$ and the last expression to obtain
\begin{align*}
L_{s,q}\leq \mu_q\left(\frac{p-1}{1-p/q}\right)^{1/s}\left[\sum_{m=0}^{T-1} (m+1)^{p-2} \alpha_m^{1-s/q}\right]^{1/s},\qquad s\in[2,p].
\end{align*}
We apply this bound to the first result, such that the second result follows with $c_{p,q}=2(a_p\lor b_p)\left(\frac{p-1}{1-p/q}\right)^{1/s}$.
For the last result, if $\alpha_{m}\leq (m+1)^{-r}$ then $\mathscr{A}_{p,q}^p\leq \sum_{m=1}^T m^{p-2-r(1-p/q)}$. For $r>\nu:=\frac{p-1}{1-p/q}$, the sum is convergent as $T\to\infty$, let $C_1:=\lim _{T\to\infty}\mathscr{A}_{p,q}(T)$. For $r=\nu$ we have the harmonic series which is known to diverge at rate slower than $\log T + \lambda$ where $\lambda\approx 1.1$. Finally, for $r<\nu$ we have $T^{-1}\sum_{m=1}^T (m/T)^{p-2-r(1-p/q)}\leq \int_0^1 x^{p-2-r(1-p/q)}\mathsf{d} x =\frac{1}{p-1-r(1-p/q)}=:C_2^p$. Thus,
\[
\mathscr{A}_{p,q}\leq
\begin{cases}
C_1 &; r>\nu\\
(\log T+ \lambda)^{1/p} &; r=\nu;\\
C_2 T^{\frac{p-1 -r(1-p/q)}{p}} &; r<\nu.
\end{cases}
\]
Apply this last bound to the second result yields the last result with $d_{p,q,r}=c_{p,q}(1\lor C_1\lor C_2)$.
\end{proof}
\begin{lemma}[based on Theorem 1 in MPR2011 ] Let $S_T = \sum_{t=1}^T X_t$ where $\{X_t:t\in [T]\}$ is a sequence of zero mean real-valued random variables such that
\begin{enumerate}[(a)]
• There exist two positive constants $\gamma_1$ and $K_1$
such that the strong mixing coefficients of the sequence satisfy $\alpha(m)\leq \exp(-K_1m^{\gamma_1})$ for any $1\leq m<T$ and $T\geq 2$,
• There exist two positive constants $\gamma_2$ and $K_2$
such that $\sup_{1\leq t\leq T,T\in \N}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X_t
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma_2}}\leq K_2$,
• $\gamma < 1$ where $\gamma$ is defined by $1/\gamma=1/\gamma_1+1/\gamma_2$.
\end{enumerate}
Then, there exist positive constants $C_1$, $C_2$, $C_3$ and $C_4$ depending only on $ K_2, K_1, \gamma$ and $\gamma_1$ such that, for $x>0$ and $T\geq 4$,
\[\P(|S_T|\geq x)\leq T\exp\left(-\frac{x^\gamma}{C_1}\right) + \exp\left[-\frac{x^2}{C_2(1+TV)}\right] + \exp\left\{-\frac{x^2}{C_3T}\exp\left[\frac{x^{\gamma(1-\gamma)}}{C_4(\log x)^\gamma}\right]\right\},\]
where $V$ is a finite constant.
In particular, there a constant $C_5$ only depending on $C_3, C_4$ and $\gamma$ such that, for $x>1$,
\begin{align*}
\P(|S_T|\geq x)\leq T\exp\left(-\frac{x^\gamma}{C_1}\right) + \exp\left[-\frac{x^2}{C_2(1+TV)}\right]+\exp\left(-\frac{x^2}{C_5 T}\right).
\end{align*}
Furthermore, for some constant $C_6$ only depending on $C_i$ for $i\in[4]$, $\gamma_1$, $\gamma_2$ and $V$,
\[{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert S_T
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{\psi_\gamma} \leq C_6\sqrt{T},\quad\textnormal{and}\quad \psi_\gamma(x):=K_\gamma x\1_{0\leq x< a_\gamma} + (\exp x^\gamma -1)\1_{x\geq a_\gamma}\]
where $a_\gamma:=(1/\gamma)^{1/\gamma}$ and $K_\gamma:=(\exp a_\gamma^\gamma -1)/a_\gamma$.
\end{lemma}
\begin{proof}
We verify the conditions (2.6), (2.7), and (2.8) of Theorem 1 in MPR2011, henceforth MPR. Assumption (b) ensures that $\mathbb{E}\left[\exp(|X_i /K_2|^{\gamma_2} )\right]\leq 2$ then following Remark 2 in MPR. Condition (2.7) is satisfied with $b = K_2$. From expression (2.5) in MPR we have $\tau(m)\leq 2\int_0^{2\alpha(m)} Q(u) du$, where $Q:=\sup_{1\leq t\leq T, T\in\N} Q_{|X_t|}$ with $Q_{|X_t|}$ denoting the quantile function of $|X_t|$. Cauchy-Schwartz inequality gives us
$\tau(m)\leq 2 \sqrt{2\alpha(m)\int_0^1 Q^2(u) du}$. Given Assumption (b), the integral is finite. Let denote it by $M^2$. By Assumption (a) we have that $\tau(m)\leq 2 \sqrt{2} M\exp\left(-\frac{1}{2} K_1m^{\gamma_1}\right)$, then condition (2.6) is satisfied with $a=2 \sqrt{2} M$ and $c=\frac{1}{2} K_1$. The first result follows.
For the second result, note that the function $x\mapsto x^{(1-\gamma)}/\log x$ is continuously differentiable and coercive for $x>1$, hence it attains its minimum over $x\in(1,\infty)$ which is given by $C_\gamma:=e(1-\gamma)>0$. Define $C_5:=C_3/\exp\left(\frac{C_\gamma^\gamma}{C_4}\right)$, then for $x>1$ the last term of the first result can be upper bounded by $\exp\left(-\frac{x^2}{C_5 T}\right)$, and the second result follows.
For the last result we have, by Fubini's Theorem and $C>0$,
\begin{align*}
\mathbb{E}\left[\psi_\gamma\left(\frac{S_T}{\sqrt{T}C}\right)\right]
&= \int_0^\infty\P\left[|S_T|> C\psi^{-1}(x)\sqrt{T}\right]d x\\
&=\int_0^{a_\gamma}\P\left[|S_T|> \frac{C}{K_\gamma} x\sqrt{T}\right]d x + \int_{a_\gamma}^\infty \P\left\{|S_T|> C \left[\log (1+x)\right]^{1/\gamma} \sqrt{T}\right\}d x\\
&=: I_1 + I_2.
\end{align*}
Lemma (ref) give us ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert T^{-1/2}S_T
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_p\leq K_p$ for any $p\in[2,\infty)$. Then, for the first integral we have
\[
I_1\leq \int_0^{\infty}\P\left[|T^{-1/2}S_T|> \frac{C}{K_\gamma} x\right]d x= \frac{K_\gamma}{C}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert T^{-1/2}S_T
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_1\leq \frac{K_\gamma}{C}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert T^{-1/2}S_T
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_2\leq \frac{K_\gamma K_2}{C}.
\]
For the second one, note that $a_\gamma >1$, then
\begin{align*}
I_2 &= \int_{a_\gamma}^\infty T\left(1+x\right)^{-\tfrac{C^\gamma T^{\gamma/2}}{C_1}} dx +\int_{a_\gamma}^\infty \exp\left(-\frac{C^2[\log (1+x)]^{2/\gamma}}{C_2(T^{-1}+V)}\right)+\exp\left(-\frac{x^2}{C_5 T}\right) d x\\
&\leq \int_{a_\gamma}^\infty T\left(1+x\right)^{-\tfrac{C^\gamma T^{\gamma/2}}{C_1}} dx +\int_{a_\gamma}^\infty \exp\left(-\frac{C^2[\log (1+x)]^{2/\gamma}}{C_7}\right) dx:=I_3 + I_4,
\end{align*}
where $C_7:=C_2(1+V)\lor C_5$. For $C^\gamma/C_1> 1$, we can bound $I_3$ as
\[
I_3 = \frac{T(1+a_\gamma)^{1-\tfrac{C^\gamma T^{\gamma/2}}{C_1}}}{\tfrac{C^\gamma T^{\gamma/2}}{C_1}-1}\leq \frac{T^{1-\gamma/2}(1+a_\gamma)^{1-T^{\gamma/2}}}{\tfrac{C^\gamma }{C_1}-1}\leq \frac{M_\gamma}{\tfrac{C^\gamma }{C_1}-1},
\]
where $M_\gamma:=\sup_{T\geq 4}\big[T^{1-\gamma/2}(1+a_\gamma)^{1-T^{\gamma/2}}\big]$ which is finite because $a_\gamma>0$. For the $C^2/C_7>1$, $\exp\left(-\frac{C^2[\log (1+x)]^{2/\gamma}}{C_7}\right)\leq \exp\left(-[\log (1+x)]^{2/\gamma}\right)$ and
\begin{align*}
\int_{a_\gamma}^\infty\exp\left(-[\log (1+x)]^{2/\gamma}\right) d x &\leq ((a_\gamma\lor 2) - a_\gamma) + \int_{a_\gamma\lor 2}^\infty\exp\left(-[\log (1+x)]^{2/\gamma}\right) d x\\
&\leq 2+ \int_{0}^\infty\exp\left(-[\log (1+x)]^{2}\right) d x\leq 4,
\end{align*}
where we use $\log 3>1$ and the last integral equals $e^{1/4}\sqrt{\pi}(\mathsf{erf}(1/2) +1)/2\leq 2$ with $\mathsf{erf}$ is Gauss error function. Then, by the dominated convergence theorem, we conclude that $\lim_{C\to\infty} I_4 = 0$, hence there exist $C_8>0$ such $I_4\leq 1/3$ for $C\geq C_8$.
Therefore, $\mathbb{E}\psi_\gamma(\tfrac{S_T}{\sqrt{T}C})\leq I_1 + I_3+ I_4\leq 1$ for $C \geq C_6:=3K_\gamma K_2\lor (2C_1)^{1/\gamma} \lor (2C_7)^{1/2}\lor \big[C_1(1+3M)\big]^{1/\gamma}\lor C_8$.
\end{proof}
\begin{lemma} Let $\boldsymbol S_T = \sum_{t=1}^T \boldsymbol X_t$ where $\{\boldsymbol X_t:=(X_{1,t},\dots, X_{n,t})': t\in[T]\}$ is a sequence of zero mean $n$-dimensional random vector such that ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p+\epsilon}<\infty $ for $i\in[n]$ and $t\in[T]$ for $p\in[2,\infty)$ and $\epsilon>0$. Then, for $T\in \N$ and $x>0$,
\[\P(\|\boldsymbol S_T\|_\infty\geq x)\leq \sum_{i=1}^n\left(\frac{c_{p,\epsilon}\sqrt{T} \mathscr{A}_i(T)\max_{t\in[T]}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p+\epsilon}}{x}\right)^p,\]
where $\mathscr{A}_i(T):=\left[\sum_{0\leq m <T} (m+1)^{p-2}\alpha_{i,m}^{1-p/(p+\epsilon)}\right]^{\tfrac{1}{p}}$, $\alpha_{i,m}$ is the strong mixing coefficient of $\{X_{i,t}\}_t$ for $1\leq i\leq n$, and $c_p,\epsilon$ is a constant depending only on $p$ and $\epsilon$.
\end{lemma}
\begin{proof}
The result follows from the union bound, Markov inequality, and Lemma (ref).
\end{proof}
\begin{lemma} Let $\boldsymbol S_T = \sum_{t=1}^T \boldsymbol X_t$ where $\{\boldsymbol X_t:=(X_{1,t},\dots, X_{n,t})':1\leq t\leq T\}$ be a sequence of zero mean $n$-dimensional random variables and write $\boldsymbol S_T$ for its sum. Assume:
\begin{enumerate}[(a)]
• There exist two positive constants $\gamma_1$ and $K_1$
such that the strong mixing coefficients $\alpha_{i,m}\leq \exp(-K_1m^{\gamma_1})$ for $1\leq m<T$, $1\leq i\leq n$, and $T\geq 2$;
• There exist two positive constants $\gamma_2$ and $K_2$
such that $\sup_{1\leq t\leq T,T\in \N}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma_2}}\leq K_2$,
• $\gamma < 1$ where $\gamma$ is defined by $1/\gamma=1/\gamma_1+1/\gamma_2$.
\end{enumerate}
Then there exist positive constants $C_1$ and $C_2$ depending only on $K_2, K_1, \gamma, \gamma_1$, and the covariance structure such that, for $n\geq 2$, $T\geq 4\lor C_1(\log n)^{(2/\gamma) - 1}$, and $K\geq 1/\sqrt{C_1C_2(\log 2)^{2/\gamma}}$,
\begin{align*}
\P(\|\boldsymbol S_T\|_\infty\geq K\sqrt{C_2 T\log n})\leq (nT)^{1-K^\gamma} +2n^{1-K^2}.
\end{align*}
In particular, $\|\boldsymbol S_T\|_\infty\lesssim_\P \big(\sqrt{T\log n}\big)$ whenever $\frac{(\log n)^{(2/\gamma) - 1}}{T}=o(1)$.
\end{lemma}
\begin{proof} Write $\boldsymbol S_T = (S_{1,T},\dots, S_{n,T})'$ where $S_{i,T}=\sum_{t=1}^T X_{i,t}$ for $i\in[n]$. The union bound followed by the second result in Lemma (ref) yields, for every $x>1$ and $T\geq 4$,
\begin{align*}
\P(\|\boldsymbol S_T\|_\infty\geq x)&\leq n\max_{i}\P(\|\boldsymbol S_{i,T}\|_\infty\geq x)\\
&\leq n\left[T\exp\left(-\frac{x^\gamma}{C_1}\right) + \exp\left(-\frac{x^2}{C_2(1+TV)}\right) +\exp\left(-\frac{x^2}{C_5T}\right)\right].
\end{align*}
Set $x = K\left[(C_1\log(nT))^{1/\gamma}\lor \sqrt{C_6T\log (n)}\right]$ for some $K>0$ where $C_6:= (C_2 +V)\lor C_5$. For large enough $T$, we have that $x>1$ and, therefore, $\P(\|\boldsymbol S_T\|_\infty\geq x)\leq (nT)^{1-K^\gamma} +2n^{1-K^2}$. Notice that that the first in brackets is no larger than the second one provided that $(C_1\log(nT))^{2/\gamma}\leq C_6 T\log (n)$, which in turn is implied by $T\geq C_7(\log n)^{(2/\gamma) - 1}$ where $C_7 = C_6^{-1}(2C_1)^{2/\gamma}$. Therefore, for $T\geq C_7(\log n)^{(2/\gamma) - 1}$ we have that $x = K\sqrt{C_6 T\log n}$ and $x>1$ whenever $K\geq 1/\sqrt{C_6C_7(\log 2)^{2/\gamma}}$.
\end{proof}
\subsection{Factor Model Estimation}
\begin{lemma} Let $a_j$ and $b_j$ denote the $j$-th eigenvalue in decreasing order of $\boldsymbol\Sigma$ and $\boldsymbol \Lambda\boldsymbol \Lambda'$ respectively. Then, under Assumption (ref)(b) and $(c)$: (a) $b_j\asymp n$ for $1\leq j\leq r$; (b) $\max_{j\leq n}|a_j-b_j\lesssim 1$; and (c) $a_j\asymp n$ for $1\leq j\leq r$.
\end{lemma}
\begin{proof}
Result $(a)$ follows from the fact that the $r$ eigenvalues of $\boldsymbol\Lambda'\boldsymbol\Lambda$ are also (the only $r$ non-zero) eigenvalues of $\boldsymbol\Lambda\boldsymbol\Lambda'$ and Assumption (ref)(b). Part $(b)$ follows from Weyl's inequality that implies $\max_{j\leq n}|a_j-b_j|\leq \|\boldsymbol\Sigma - \boldsymbol\Lambda\boldsymbol\Lambda'\|\lesssim 1$, where the last equality follows from Assumption (ref)(c). Finally, result $(c)$ follows from part $(a)$ and $(b)$ and the (reverse) triangle inequality.
\end{proof}
Recall that $\boldsymbol\Sigma$ is the $(n\times n)$ covariance matrix of $\boldsymbol U_t = \boldsymbol Z_t -\boldsymbol\Gamma \boldsymbol W_t$. Let $\widetilde{\boldsymbol\Sigma}:=\frac{1}{T}\sum_{t=1}^T \boldsymbol U_t \boldsymbol U_t'$ and $\widehat{\boldsymbol\Sigma}$ the same as $\widetilde{\boldsymbol\Sigma}$ but with $\boldsymbol \Gamma$ replaced by the estimator $\widehat{\boldsymbol\Gamma}$. Let $\widehat{a}_j$ denote the $j$-th eigenvalue in decreasing order of $\widehat{\boldsymbol\Sigma}$.
\begin{lemma} Under the Assumptions (ref) and (ref), let $\varrho_1$ be a non-negative sequence of $n$ and $T$ such that $\| \widehat{\boldsymbol\Sigma}-\widetilde{\boldsymbol\Sigma}\|_{\max}\lesssim_\P \varrho_1$, then:
\begin{enumerate}[(a)]
• $\| \widehat{\boldsymbol\Sigma}-\boldsymbol\Sigma\|_{\max} \lesssim_\P \varrho_1 + g_\alpha(n)/\sqrt{T}$
• $\max_{j\leq n}|\widehat{a}_j -a_j| \lesssim_\P n(\varrho_1 + g_\alpha(n)/\sqrt{T})$
• $\widehat{a}_j\asymp_P n$ for $j\leq r$ provided that $\varrho_1 + g_\alpha(n)/\sqrt{T}\lesssim 1$,
\end{enumerate}
where $g_\alpha(n)= \mathcal{A}_\alpha n^{4/p}$ under Assumptions ((ref).c) and $g_\alpha(n)= \sqrt{\log n}$ under Assumptions ((ref).d).
\end{lemma}
\begin{proof}
Part (a) follows by the triangle and Lemma (ref) or Lemma (ref) , since $ \|\widehat{\boldsymbol\Sigma}-\boldsymbol\Sigma\|_{\max}\leq\| \widehat{\boldsymbol\Sigma}-\widetilde{\boldsymbol\Sigma}\|_{\max} +\| \widetilde{\boldsymbol\Sigma}-\boldsymbol\Sigma\|_{\max} \lesssim_\P \varrho_1 + g(n)/\sqrt{T}$. Part (b) follows from Weyl's inequality, the fact that $\| \widehat{\boldsymbol\Sigma}-\boldsymbol\Sigma\|\leq n\| \widehat{\boldsymbol\Sigma}-\boldsymbol\Sigma\|_{\max}$ and part $(a)$. Part $(c)$ follows from the triangle inequality combined with part $(b)$ and Lemma (ref)(c).
\end{proof}
The Lemmas (ref)-(ref) below are an adaption of Lemmas 8--10 in FLM2013, henceforth FLM, to include the estimation error in the sample covariance matrix. To avoid confusion and make it easier for the reader to follow through with the changes we use the same notation adopted in FLM. In particular, if $\delta_{i,t}$ denotes the $(i,t)$ element of $\boldsymbol \Delta:=\widehat{\boldsymbol R} - \boldsymbol R$ then $\widetilde{U}_{i,t} = U_{i,t} + \delta_{i,t}$ for $i\in[n]$ and $t\in[T]$. We consider that $\|\boldsymbol\Delta\|_{\max}\lesssim_\P \varrho_R$ for some non-negative sequence $\varrho_R$ depending on $n$ and $T$.
Define:
\begin{align*}
\widetilde{\zeta}_{st}&:=\frac{\widetilde{\boldsymbol U}_s' \widetilde{\boldsymbol U}_t }{n}-\frac{\mathbb{E}(\boldsymbol U_s'\boldsymbol U_t)}{n} = \left(\frac{\boldsymbol U_s' \boldsymbol U_t }{n}-\frac{\mathbb{E}(\boldsymbol U_s'\boldsymbol U_t)}{n}\right) +\left(\frac{\boldsymbol U_s' \boldsymbol \delta_t }{n} + \frac{\boldsymbol\delta_s' \boldsymbol U_t}{n} + \frac{\boldsymbol\delta_s' \boldsymbol\delta_t }{n} \right)=: \zeta_{st}+\zeta_{st}^*\\
\widetilde{\eta}_{st}&:=\frac{\boldsymbol F_s'\sum_{i=1}^n \boldsymbol\lambda_i \widetilde{U}_{i,t} }{n} = \frac{\boldsymbol F_s'\sum_{i=1}^n \boldsymbol\lambda_i U_{i,t} }{n} + \frac{\boldsymbol F_s'\sum_{i=1}^n \boldsymbol\lambda_i \delta_{i,t} }{n} =:\eta_{st}+\eta_{st}^*\\
\widetilde{\xi}_{st}&:=\frac{\boldsymbol F_t'\sum_{i=1}^n \boldsymbol\lambda_i \widetilde{U}_{is} }{n} = \frac{\boldsymbol F_t'\sum_{i=1}^n \boldsymbol\lambda_i U_{is} }{n} + \frac{\boldsymbol F_t'\sum_{i=1}^n \boldsymbol\lambda_i \delta_{is}}{n} = \xi_{st}+\xi_{st}^*.
\end{align*}
\begin{lemma} Under Assumption (ref):
\begin{enumerate}[(a)]
• $\zeta_{st}\lesssim_\P 1/\sqrt{n}$
• $\eta_{st}\lesssim_\P 1/\sqrt{n}$
• $\xi_{st}\lesssim_\P 1/\sqrt{n}$
• $\zeta_{st}^*\lesssim_\P \varrho_R + \varrho_R^2$
• $\max_{s,t\leq T}\zeta_{st}^*\lesssim_\P g(nT)\varrho_R + \varrho_R^2$
• $\eta_{st}^*\lesssim_\P \varrho_R$
• $\xi_{st}^*\lesssim_\P \varrho_R$,
\end{enumerate}
where $g(x)= x^{1/p}$ under Assumptions ((ref).c) and $g(x)= [\log x]^{1/\gamma_2}$ under Assumptions ((ref).d);
\end{lemma}
\begin{proof} Parts $(a)$-$(c)$ are straightforward. For $(d)$ we have that $\frac{1}{n}\boldsymbol U_s' \boldsymbol U_t\lesssim_\P 1$ and $\frac{1}{n}\boldsymbol \delta_s' \boldsymbol \delta_t\leq\|\boldsymbol\Delta\|_{\max}^2\lesssim_\P \varrho_R^2$. Then, the other two terms in parentheses in the definition of $ \zeta_{st}^*$ are $\lesssim_\P \varrho_R$ by the Cauchy-Schwartz inequality. Part $(e)$ and $(f)$ follows by similar arguments. Therefore,
\[
\max_{t\leq T}\frac{1}{T}\sum_{s=1}^T\left(\frac{1}{n}\boldsymbol\delta_s'\boldsymbol U_t\right)^2= \max_{t\leq T}\frac{1}{n^2}\boldsymbol U_t'\left(\frac{1}{T}\sum_{s=1}^T\boldsymbol\delta_s\boldsymbol\delta_s'\right)\boldsymbol U_t\leq \|\boldsymbol\Delta\|_{\max}^2\left(\max_{t\leq T}\|\boldsymbol U_t\|_1/n\right)^2
\]
and
\[
\zeta_{st}^*\leq \|\boldsymbol U_s\|_\infty\|\boldsymbol \delta_t\|_\infty + \|\boldsymbol U_t\|_\infty\|\boldsymbol \delta_s\|_\infty + \|\boldsymbol\delta_t\|_\infty\|\boldsymbol\delta_s\|_\infty\leq 2\|\boldsymbol U\|_{\max}\|\boldsymbol\Delta\|_{\max} + \|\boldsymbol\Delta\|_{\max}^2.
\]
\end{proof}
\begin{lemma} Under Assumption (ref):
\begin{enumerate}[(a)]
• $\frac{1}{T}\sum_{t=1}^T\left[\frac{1}{nT}\sum_{s=1}^T\widehat{F}_{js}\mathbb{E}(\boldsymbol U_s'\boldsymbol U_t)\right]^2 \lesssim_\P 1/T$
• $\frac{1}{T}\sum_{t=1}^T\frac{1}{T}\sum_{s=1}^T\widehat{F}_{js}\widetilde{\zeta}_{st}^2\lesssim_\P \left(1/\sqrt{n}+\varrho_R +\varrho_R^2\right)^2$
• $ \frac{1}{T}\sum_{t=1}^T\left[\frac{1}{T}\sum_{s=1}^T\widehat{F}_{js}\widetilde{\eta}_{st}\right]^2\lesssim_\P \left(1/\sqrt{n}+\varrho_R\right)^2$
• $\frac{1}{T}\sum_{t=1}^T\left[\frac{1}{T}\sum_{s=1}^T\widehat{F}_{js}\xi_{st}\right]^2\lesssim_\P \left(1/\sqrt{n}+\varrho_R\right)^2$
• $\frac{1}{T}\sum_{t=1}^T\left\|\widehat{\boldsymbol F}_t -\boldsymbol H \boldsymbol F_t\right\|^2\lesssim_\P 1/T+(1/\sqrt{n}+\varrho_R +\varrho_R^2)^2$.
\end{enumerate}
\end{lemma}
\begin{proof}
Part (a) is unaltered by the presence of a pre-estimation, so it follows directly from Lemma 8(a) in FLM. For part (b), we have that for $s,l\in[n]$ and $j\in[r]$ by Cauchy-Schwartz inequality
\[
\frac{1}{T}\sum_{t=1}^T\left[\frac{1}{T}\sum_{s=1}^T\widehat{ F}_{js}\widetilde{\zeta}_{st}\right]^2\leq\left[\frac{1}{T^2}\sum_{s,l=1}^T\left(\frac{1}{T}\sum_{t=1}^T\widetilde{\zeta}_{st}\widetilde{\zeta}_{lt} \right)^2\right]^{1/2}.
\]
Since $\widetilde{\zeta}_{st}=\zeta_{st}+\zeta_{st}^*\lesssim_\P 1/\sqrt{n} +\varrho_R +\varrho_R^2)$ by Lemma (ref), the term in parentheses is $\lesssim_\P (1/\sqrt{n}+\varrho_R+\varrho_R^2)^2$. Result $(b)$ follows. For (c), by the triangle inequality and Lemma 8(c) in FLM, we have that $\|\sum_{i=1}^n\lambda_{ji}\widetilde{u}_{i,t}\|\leq \|\sum_{i=1}^n\lambda_{ji} U_{i,t}\| +\|\sum_{i=1}^n\lambda_{ji}\delta_{i,t}\|\lesssim_\P \sqrt{n} + n\varrho_R$. Then, we conclude
\[\frac{1}{T}\sum_{t=1}^T[\frac{1}{T}\sum_{s=1}^T\widehat{\boldsymbol F}_{s}\widetilde{\eta}_{st}]^2\leq\frac{1}{Tn^2}\sum_{t=1}^T\|\sum_{i=1}^n\boldsymbol U_{i,t}\lambda_i\|^2\lesssim_\P 1/n+\varrho_R/\sqrt{n}+ \varrho_R^2 .\]
The proof of part (d) is analogous to (c) and is omitted.
For (e), let $[\widehat{\boldsymbol F}_t -\boldsymbol H \boldsymbol F_t]_j$ denote the $j$-th entry of $\widehat{\boldsymbol F}_t -\boldsymbol H \boldsymbol F_t$. Since $\boldsymbol V/n$ is bounded away from zero by Lemma (ref)(c), the fact that $(a+b+c+d)^2\leq 4(a^2 + b^2 + c^2 + d^2)$ and using (ref), we have that $\max_{j\leq r}T^{-1}\sum_t[\widehat{\boldsymbol F}_t -\boldsymbol H \boldsymbol F_t]_j$ is upper bounded by some constant $C<\infty$ times
\begin{align*}
&\left\{\max_{j\leq r}\frac{1}{T}\sum_{t=1}^T \left[\frac{1}{T}\sum_{s=1}^T\widehat{F}_{js}\frac{\mathbb{E}(\boldsymbol U_s'\boldsymbol U_t)}{n}\right]^2 +
\max_{j\leq r}\frac{1}{T}\sum_{t=1}^T \left(\frac{1}{T}\sum_{s=1}^T\widehat{ f}_{js}\widetilde{\zeta}_{st}\right)^2\right.\\
&\left.\qquad +\max_{j\leq r}\frac{1}{T}\sum_{t=1}^T \left(\frac{1}{T}\sum_{s=1}^T\widehat{F}_{js}\widetilde{\eta}_{st} \right)^2 + \max_{j\leq r}\frac{1}{T}\sum_{t=1}^T\left(\frac{1}{T}\sum_{s=1}^T\widehat{ f}_{js}\widetilde{\xi}_{st}\right)^2\right\}.
\end{align*}
The result follows by applying the bounds from (a)--(d) to each of the terms above.
\end{proof}
\begin{lemma} Under Assumptions (ref) and (ref):
\begin{enumerate}[(a)]
• $\max_{t\leq T}\|\frac{1}{nT}\sum_{s=1}^T\widehat{\boldsymbol F}_s\mathbb{E}(\boldsymbol U_s'\boldsymbol U_t)\| \lesssim_\P 1/\sqrt{T})$
• $\max_{t\leq T}\|\frac{1}{T}\sum_{s=1}^T\widehat{\boldsymbol F}_s\widetilde{\zeta}_{st}\|\lesssim_\P \left[g(T)/\sqrt{n} +g(nT)\varrho_R + \varrho_R^2\right]$
• $\max_{t\leq T}\|\frac{1}{T}\sum_{s=1}^T\widehat{\boldsymbol F}_{s}\widetilde{\eta}_{st}\|\lesssim_\P \left[g(T)/\sqrt{n} +\varrho_R\right]$
• $\max_{t\leq T}\|\frac{1}{T}\sum_{s=1}^T\widehat{\boldsymbol F}_{s}\xi_{st}\|\lesssim_\P \left[g(T)(1/\sqrt{n} + \varrho_R)\right]$,
\end{enumerate}
where $g(x)= x^{1/p}$ under Assumption ((ref).c) and $g(x)=[\log x]^{1/\gamma_2}$ under Assumption ((ref).d).
\end{lemma}
\begin{proof}
Part (a) is unaltered by pre-estimation, so it follows directly from Lemma 9(a) in FLM. For part (b), from the Cauchy-Schwartz inequality, we have
\[\max_{t\leq T}\left\|\frac{1}{T}\sum_{s=1}^T\widehat{\boldsymbol F}_s\widetilde{\zeta}_{st}\right\|\leq\left(\frac{1}{T}\sum_{s=1}^T\|\widehat{\boldsymbol F}_s\|^2 \max_{t\leq T}\frac{1}{T}\sum_{s=1}^T\widetilde{\zeta}_{st}^2\right)^{1/2}.\]
The first summation inside the parentheses equals $r$ due to the normalization. For the second summation, by the triangle inequality, we have $\max_{t\leq T}\frac{1}{T}\sum_{s=1}^T\widetilde{\zeta}_{st}^2\leq \max_{t\leq T}\frac{1}{T}\sum_{s=1}^T\zeta_{st}^2 + 2\max_{t\leq T}\frac{1}{T}\sum_{s=1}^T\zeta_{st}\zeta_{st}^* + \max_{t\leq T}\frac{1}{T}\sum_{s=1}^T{\zeta_{st}^*}^2$. For the first term, the maximum inequality followed by Assumption (ref)(e) yields
\[\max_{t\leq T}\frac{1}{T}\sum_{s=1}^T\zeta_{st}^2 \lesssim_\P \left[\psi_{p/2}^{-1}(T)\max_{s,t}\|\zeta^2\|_{\psi_{p/2}}\right] \lesssim_\P \left[\psi_{p/2}^{-1}(T)\max_{s,t}\|\zeta\|_{\psi}^2\right]\lesssim_\P \left[\frac{g_1(T)}{n}\right].\]
where $g_1(x)= x^{2/p}$ or $g_1(x)=[\log x]^{2/\gamma_2}$.
The last one is $\lesssim_\P (g(nT)\varrho_R+\varrho_R^2)^2$ by Lemma (ref)(d). Then, by Cauchy-Schwartz, we have that
$\max_{t\leq T}\frac{1}{T}\sum_{s=1}^T\widetilde{\zeta}_{st}^2\lesssim_\P (\sqrt{g_1(T)/n} +g(nT)\varrho_R+\varrho_R^2)^2$ and result (b) follows.
For (c), by the triangle inequality we have that $\max_{t\leq T}\|\frac{1}{n}\sum_{i=1}^n\boldsymbol\lambda_i\widetilde{U}_{i,t}\|\leq \max_{t\leq T}\|\frac{1}{n}\sum_{i=1}^n\boldsymbol\lambda_iU_{i,t}\| + \max_{t\leq T}\|\frac{1}{n}\sum_{i=1}^n\boldsymbol\lambda_i\delta_{i,t}\|$. For the first term, the maximum inequality followed by Assumption (ref)(f) yields
\[
\max_{t\leq T}\left\|\frac{1}{n}\boldsymbol\Lambda'\boldsymbol U_t\right\| \lesssim_\P \left[\frac{g(T)}{\sqrt{n}}\max_t\left\|\frac{1}{\sqrt{n}}\boldsymbol\Lambda'\boldsymbol U_t\right\|\right] \lesssim_\P \left[g(T)/\sqrt{n}\right].
\]
The second term is upper bounded by $r\|\boldsymbol\Lambda\|_{\max}\|\boldsymbol\Delta\|_{\max} \lesssim_\P \varrho_R)$ by Assumption (ref)(d). We obtain the result since
\begin{equation}
\max_{t\leq T}\left\|\frac{1}{T}\sum_{s=1}^T\widehat{\boldsymbol F}_{s}\widetilde{\eta}_{st}\right\|\leq \left\|\frac{1}{T}\sum_{s=1}^T \widehat{\boldsymbol F}_s \boldsymbol F_s' \right\| \max_{t\leq T}\left\|\frac{1}{n}\sum_{i=1}^n\boldsymbol\lambda_i\widetilde{U}_{i,t}\right\| \lesssim_\P \left[\frac{g(T)}{\sqrt{n}} +\varrho_R\right].
\end{equation}
For (d), by the triangle inequality, $\|\frac{1}{nT}\sum_s\sum_i \boldsymbol\lambda_i\widetilde{U}_{is} \widehat{\boldsymbol F}_{s}\|\leq \|\frac{1}{nT}\sum_s\sum_i \boldsymbol\lambda_iU_{is} \widehat{\boldsymbol F}_{s}\| + \|\frac{1}{nT}\sum_s\sum_i \boldsymbol\lambda_i\delta_{is} \widehat{F}_{is}\|$. Lemma 9(d) of FLM shows that the first term is $\lesssim_\P 1/\sqrt{n}$. For the second term, for each $j\in[r]$:
\[
\left\|\frac{1}{nT}\sum_s\sum_i \boldsymbol\lambda_i\delta_{is} \widehat{F}_{js}\right\|^2\leq \left(\frac{1}{T}\sum_{s=1}^n\left\|\frac{1}{n}\sum_{i=1}^n \boldsymbol\lambda_i\delta_{is}\right\|^2 \widehat{F}_{js}\right)\left(\frac{1}{T}\sum_{s=1}^n \widehat{F}_{js}^2\right)\lesssim_\P \varrho_R^2).
\]
Thus, $\|\frac{1}{nT}\sum_s\sum_i \boldsymbol\lambda_i\widetilde{U}_{i,t} \widehat{\boldsymbol F}_{s}\|\lesssim_\P 1/\sqrt{n} + \varrho_R)$ and by Cauchy-Schwartz inequality:
\begin{equation}
\max_{t\leq T}\left\|\frac{1}{T}\sum_{s=1}^T\widehat{\boldsymbol F}_{s}\xi _{st}\right\|\leq \max_{t\leq T}\left\|\boldsymbol F_t\right\|\left\|\frac{1}{nT}\sum_s\sum_i \boldsymbol\lambda_i\widetilde{U}_{i,t} \widehat{\boldsymbol F}_{s}\right\|\lesssim_\P \left[g(T)(1/\sqrt{n} + \varrho_R)\right].
\end{equation}
\end{proof}
\begin{lemma} Let $\varrho_1 + \psi^{-1}(n^2)/\sqrt{T} \lesssim 1$ where $\varrho_1$ is defined in Lemma (ref). Then, under Assumption (ref) we have:
\begin{enumerate}[(a)]
• $\|\boldsymbol V^{-1}\| \lesssim_\P 1/n$
• $\|\boldsymbol H\|\lesssim_\P 1$
• $\|\boldsymbol H'\boldsymbol H - \boldsymbol I_r\|_F \lesssim_\P 1/\sqrt{T} + 1/\sqrt{n} + \varrho_R$
• $\max_{i\leq n}\frac{1}{T}\sum_{t=1}^T \widehat{R}_{i,t}^2\lesssim_\P \varrho_R(h(nT) + \varrho_R) +g_\alpha(n)/\sqrt{T} + 1$
• $\max_{i\leq n}\max_{j\leq r}\frac{1}{T}\sum_{t=1}^T F_{jt} \widetilde{U}_{i,t} \lesssim_\P g_\alpha(n)/\sqrt{T} + \varrho_R$,
\end{enumerate}
where $g_\alpha(x) = \mathscr{A}_\alpha x^{2/p}$ or $g(x) = \sqrt{\log x}$, and $h(x) = x^{1/p}$ or $h(x) = [\log x]^{1/\gamma_2}$.
\end{lemma}
\begin{proof} We have that $\boldsymbol V^{-1} =\mathsf{diag}\,(1/\widehat{a}_1,\dots, 1/\widehat{a}_r)$ and $1/\widehat{a}_j \asymp_P 1/n$ for $j\leq r$ by Lemma (ref)(c). The result (a) then follows. The normalization tell us $\|\widehat{\boldsymbol F}\|=\sqrt{T}$, Lemma 11(a) in FLM give us $\|\boldsymbol F\|\lesssim_\P \sqrt{T})$, $\|\boldsymbol\Lambda'\boldsymbol\Lambda\|=\widetilde{a}_1 \asymp n$ by Lemma (ref)(a), and from part (a) we have $\|\boldsymbol V^{-1}\|\lesssim_\P 1/n)$. Result (b) then follows since by definition $\boldsymbol H:=T^{-1}\boldsymbol V^{-1}\widehat{\boldsymbol F}'\boldsymbol F\boldsymbol\Lambda'\boldsymbol\Lambda$.
For (c), we have by the triangle inequality,
\[\|\boldsymbol H'\boldsymbol H - \boldsymbol I_r\|_F\leq \|\boldsymbol H'\boldsymbol H - \boldsymbol H' \boldsymbol F'\boldsymbol F/T \boldsymbol H\|_F + \|\boldsymbol H'\boldsymbol F'\boldsymbol F/T\boldsymbol H - \boldsymbol I_r\|_F\]
For the first term, we have
\[\|\boldsymbol H'(\boldsymbol I_r - \boldsymbol F'\boldsymbol F/T) \boldsymbol H\|_F\leq\|\boldsymbol H\|^2\|\boldsymbol I_r - \boldsymbol F'\boldsymbol F/T\|_F \lesssim_\P 1/\sqrt{T}. \]
The second term is equal to $\|\boldsymbol H'\boldsymbol F'\boldsymbol F/T\boldsymbol H - \widehat{\boldsymbol F}' \widehat{\boldsymbol F}/T\|_F$.
For (d) we have
\begin{align*}
\max_{i\leq n}\frac{1}{T}\sum_{t=1}^T \widehat{R}_{i,t}^2 &\leq \max_{i\leq n}\frac{1}{T}\sum_{t=1}^T (\widehat{R}_{i,t}^2 -R_{i,t}^2) + \max_{i\leq n}\frac{1}{T}\sum_{t=1}^T R_{i,t}^2 -\mathbb{E}(R_{i,t}^2) +\max_{i\leq n}\frac{1}{T}\sum_{t=1}^T \mathbb{E}(R_{i,t}^2)\\
&\leq \max_{i,t} |\widehat{R}_{i,t}^2 -R_{i,t}^2| + \max_{i\leq n}\frac{1}{T}\sum_{t=1}^T R_{i,t}^2 -\mathbb{E}(R_{i,t}^2) +\max_{i,t}\mathbb{E}(R_{i,t}^2).
\end{align*}
The last term is $\lesssim 1$ by Assumption (ref)(c) or (d), the middle term $\lesssim_\P g(n)/\sqrt{T}$. The first term is no larger then $\|\boldsymbol\Delta\|_{\max}(2\|\boldsymbol R\|_{\max} + \|\boldsymbol\Delta\|_{\max}) \lesssim_\P \varrho_R(h(nT) + \varrho_R))$. The result (d) then follows.
For (e) we have for each $j\leq r$:
\begin{align*}
\max_{i\leq n}\max_{j\leq r}\left|T^{-1}\sum_t F_{jt}\widetilde{U}_{i,t}\right|&\leq \max_{i,j}\left|T^{-1}\sum_t F_{jt}U_{i,t}\right| + \max_{i,j}\left|T^{-1}\sum_t F_{jt}\delta_{i,t}\right|\\
&\leq \max_{i,j}\left|T^{-1}\sum_t F_{jt}U_{i,t}\right| + \max_{i,j}\left(T^{-1}\sum_t F_{jt}^2 T^{-1}\sum_t\delta_{i,t}^2\right)^{1/2}
\end{align*}
The first term is $\lesssim_\P g(n)/\sqrt{T}$ by the maximum inequality and Assumption (ref) and the second is $\lesssim_\P \varrho_R$.
\end{proof}
\subsection{Penalized Regression Results}
\begin{lemma} Let $\boldsymbol\Sigma_{\boldsymbol U}:=\boldsymbol U'\boldsymbol U/T$ and $\boldsymbol\Sigma_{\boldsymbol V}:=\boldsymbol V'\boldsymbol V/T$ where $\boldsymbol U$ and $\boldsymbol V$ are $(T\times n)$ matrices, then
\[\|\boldsymbol\Sigma_{\boldsymbol U}-\boldsymbol\Sigma_{\boldsymbol V}\|_{\max}\leq \|\boldsymbol U-\boldsymbol V\|_{\max}(2\|\boldsymbol V\|_{\max}+\|\boldsymbol U-\boldsymbol V\|_{\max}).\]
Furthermore, if $\|\boldsymbol\Sigma_{\boldsymbol U}-\boldsymbol\Sigma_{\boldsymbol V}\|_{\max}\leq \alpha\kappa(\boldsymbol\Sigma_{\boldsymbol V},\mathcal{S},3)/(|\mathcal{S}|(1+\zeta)^2)$ for $\mathcal{S}\subseteq [n]$, $\zeta>0$ and $\alpha\in[0,1]$, then for $\|\boldsymbol x_{\mathcal{S}^c}\|_1\leq \zeta \|\boldsymbol x_{\mathcal{S}}\|_1$ we have
\[(1-\alpha)\boldsymbol x'\boldsymbol\Sigma_{\boldsymbol V}\boldsymbol x\leq \boldsymbol x'\boldsymbol\Sigma_{\boldsymbol U} \boldsymbol x\leq (1+\alpha)\boldsymbol x'\boldsymbol\Sigma_{\boldsymbol V}\boldsymbol x.\]
Take the infimum of the expression above over $\{\boldsymbol x\in\mathbb{R}^n:\|\boldsymbol x_{\mathcal{S}^c}\|_1\leq \zeta \|\boldsymbol x_{\mathcal{S}}\|_1\}$ to conclude
\[(1-\alpha)\kappa^2(\boldsymbol\Sigma_{\boldsymbol V},\mathcal{S},\zeta)\leq \kappa^2(\boldsymbol\Sigma_{\boldsymbol U},\mathcal{S},\zeta)\leq(1+\alpha)\kappa^2(\boldsymbol\Sigma_{\boldsymbol V},\mathcal{S},\zeta).\]
\end{lemma}
\begin{proof} By the (reverse) triangle inequality we have $\|\boldsymbol U\|_{\max} - \|\boldsymbol V\|_{\max}\leq \|\boldsymbol U-\boldsymbol V\|_{\max}$. Also, $\|\boldsymbol\Sigma_{\boldsymbol U}-\boldsymbol\Sigma_{\boldsymbol V}\|_{\max}=\max_{1\leq i,j\leq n }|T^{-1}\sum_{t=1}^T U_{i,t}U_{jt} - V_{i,t} V_{ijt}|\leq \max_{i,j,t}|U_{i,t}U_{jt} - V_{i,t} V_{jt}|$ and $
|U_{i,t}U_{jt} - V_{i,t} V_{jt}|\leq |(U_{i,t}-V_{i,t})U_{jt} + (U_{jt}-V_{jt})V_{i,t}|\leq \|\boldsymbol U-\boldsymbol V\|_{\max}(\|\boldsymbol U\|_{\max}+\|\boldsymbol V\|_{\max}).
$. Combine the 3 bounds to obtain the first result.
For the second part of the lemma notice that for any $\boldsymbol x\in\mathbb{R}^n$ we have $|\boldsymbol x'\boldsymbol\Sigma_{\boldsymbol U} \boldsymbol x - \boldsymbol x'\boldsymbol\Sigma_{\boldsymbol V} \boldsymbol x| = |\boldsymbol x'(\boldsymbol\Sigma_{\boldsymbol U} -\boldsymbol\Sigma_{\boldsymbol V} )\boldsymbol x|\leq\|\boldsymbol\Sigma_{\boldsymbol U}-\boldsymbol\Sigma_{\boldsymbol V}\|_{\max}\|\boldsymbol x\|_1^2$. Also, if $\|\boldsymbol x_{\mathcal{S}^c}\|_1\leq \zeta \|\boldsymbol x_{\mathcal{S}}\|_1$ we have that $\|\boldsymbol x\|_1=\|\boldsymbol x_{\mathcal{S}}\|_1 + |\boldsymbol x_{\mathcal{S}^c}\|_1\leq (1+\zeta) \|\boldsymbol x_{\mathcal{S}}\|_1\leq (1+\zeta)\sqrt{\boldsymbol x'\boldsymbol\Sigma_{\boldsymbol V}\boldsymbol x |\mathcal{S}|}/\kappa(\boldsymbol\Sigma_{\boldsymbol V},\mathcal{S},\zeta)$ where the last inequality follows from the definition of compatibility condition. Thus $|\boldsymbol x'\boldsymbol\Sigma_{\boldsymbol U} \boldsymbol x - \boldsymbol x'\boldsymbol\Sigma_{\boldsymbol V} \boldsymbol x| \leq \|\boldsymbol\Sigma_{\boldsymbol U}-\boldsymbol\Sigma_{\boldsymbol V}\|_{\max}(1+\zeta)^2\boldsymbol x'\boldsymbol\Sigma_{\boldsymbol V}\boldsymbol x |\mathcal{S}|/\kappa(\boldsymbol\Sigma_{\boldsymbol V},\mathcal{S},\zeta)\leq\boldsymbol x'\boldsymbol\Sigma_{\boldsymbol V}\boldsymbol x/2 $, where the last inequality follows from the definition of compatibility condition. Therefore, we have that $(1-\alpha)\boldsymbol x'\boldsymbol\Sigma_{\boldsymbol V}\boldsymbol x\leq \boldsymbol x'\boldsymbol\Sigma_{\boldsymbol U} \boldsymbol x\leq (1+\alpha)\boldsymbol x'\boldsymbol\Sigma_{\boldsymbol V}\boldsymbol x$ whenever $\|\boldsymbol x_{\mathcal{S}^c}\|_1\leq \zeta \|\boldsymbol x_{\mathcal{S}}\|_1$.
\end{proof}
\begin{lemma} Let $\boldsymbol W:= (\boldsymbol U,\boldsymbol V)$ and $\boldsymbol Z:=(\boldsymbol X,\boldsymbol Y)$ be $T\times (n+1)$ matrices such that $\|\boldsymbol W - \boldsymbol Z\|_{\max}\leq C_1$ and $\|\boldsymbol Z\|_{\max}\leq C_2$, then for any $\boldsymbol \delta\in\mathbb{R}^n$ we have
\[\|\boldsymbol U'(\boldsymbol V - \boldsymbol U\boldsymbol\delta)/T-\boldsymbol X'(\boldsymbol Y-\boldsymbol X\boldsymbol\delta)/T\|_\infty\leq (1+\|\boldsymbol\delta\|_1)C_1(2C_2+C_1) \]
\end{lemma}
\begin{proof} For convenience let $q:=V-U\delta\in\mathbb{R}^T$ and $r:=Y-X\delta\in\mathbb{R}^T$, then Hölder's inequality gives us $\|r\|_\infty\leq (1+\|\delta\|_1)\|Z\|_{\max}\leq (1+\|\delta\|_1)C_2$ and $\|q-r\|_\infty\leq (1+\|\delta\|_1)\|W-Z\|_{\max}\leq (1+\|\delta\|_1)C_1$. From the (reverse) triangle inequality we obtain $\|q\|_\infty\leq \|q-r\|_\infty + \|r\|_\infty\leq (1+\|\delta\|_1)(C_1 + C_2)$. Now, following the same steps in the proof of the previous Lemma, we can upper bound the right-hand side of the display by $\|U-X\|_{\max}\|q\|_\infty + \|q-r\|_\infty\|X\|_{\max}$, which in turn can be upper bounded by the left-hand side of the display.
\end{proof}
\subsection{Inference Procedure Results}
\begin{lemma} Let $X$ and $Y$ be random elements defined in the same probability space $(\Omega,\mathcal{F},\P)$ taking values in the metric space $(S,d)$. Then for measurable $A$ and $\delta\geq 0$
\begin{align*}
-\P(Y\in A\setminus A^{-\delta}) - \P(d(X,Y)>\delta) &\leq \P(X\in A) - \P(Y\in A)\\
&\leq \P(Y\in A^\delta\setminus A) + \P(d(X,Y)>\delta),
\end{align*}
where $A^\delta:=\{x\in S:d(x,A)\leq \delta\}$, $A^{-\delta}:=S\setminus(S\setminus A)^\delta$ and $d(x,A):=\inf_{y\in A} d(x,y)$.
Let $\mathcal{A}$ be a class of measurable subsets of $S$ then
\[
\rho_\mathcal{A}(X,Y)\leq \inf_{\delta>0}\big[ \P(d(X,Y)>\delta)+\Delta_\mathcal{A}(Y,\delta) \big],
\]
where $\rho_\mathcal{A}(X,Y):=\sup_{A\in \mathcal{A}}|\P(X\in A) -\P(Y\in A)|$ and $\Delta_\mathcal{A}(Y,\delta):=\sup_{A\in \mathcal{A}}\P(Y\in A^\delta\setminus A^{-\delta})$.
In particular, if $d$ is induced by a norm $\|\cdot\|$ we have for all $t\in\mathbb{R}$ and $\delta>0$
\begin{align*}
-\P(t-\delta\leq\|Y\|\leq t) - \P(\|X-Y\|>\delta) &\leq \P(\|X\|\leq t) - \P(\|Y\|\leq t)\\
&\leq \P(t<\|Y\|\leq t+\delta) + \P(\|X-Y\|>\delta).
\end{align*}
\begin{proof}
For the right hand side, we use $\{X\in A\}\cap\{d(X,Y)\leq \delta\}\subseteq\{Y\in A^\delta\}$ to write
\begin{align*}
\P(X\in A) &= \P(X\in A, d(X,Y)\leq \delta) + \P(X\in A,d(X,Y)> \delta)\\
&\leq \P(Y\in A^\delta) + \P(d(X,Y)> \delta)\\
&= \P(Y\in A) + \P(Y\in A^\delta\setminus A) + \P(d(X,Y)> \delta).
\end{align*}
For the left hand side, we use $\{Y\in A^{-\delta}\}\cap\{d(X,Y)\leq \delta\}\subseteq\{X\in A\}$ to write
\begin{align*}
\P(X\in A) &\geq \P(Y\in A^{-\delta}, d(X,Y)\leq \delta) \\
&\geq \P(Y\in A^{-\delta}) - \P(d(X,Y)> \delta)\\
&= \P(Y\in A) - \P(Y\in A\setminus A^{-\delta}) - \P(d(X,Y)> \delta).
\end{align*}
The first result follows. For the second one we have
\begin{align*}
|\P(X\in A) - \P(Y\in A)|&\leq \P(d(X,Y)>\delta) +\P(Y\in A^\delta\setminus A)\lor \P(Y\in A\setminus A^{-\delta})\\
&\leq \P(d(X,Y)>\delta) +\P(Y\in A^\delta\setminus A^{-\delta}).
\end{align*}
Take the supremum over $A\in\mathcal{A}$ we have $\rho_\mathcal{A}(X,Y)\leq \P(d(X,Y)>\delta)+\Delta_\mathcal{A}(Y,\delta)$. Take the infimum over $\delta>0$
to obtain the second result.
For the last one, take $A=B_t$ where $B_t$ is $\|\cdot\|$-closed ball of radius $t$ centered at the origin for $t\geq 0$ and the empty set otherwise, then $\{X\in A\}=\{\|X\|\leq t\}$ and $\{Y\in A\}=\{\|Y\|\leq t\}$. Also $\{Y\in A^\delta\setminus A\} = \{t<\|Y\|\leq t+\delta\}$ since $A^\delta = B_{t+\delta}$ and $\{Y\in A\setminus A^{-\delta}\} = \{t-\delta\leq\|Y\|\leq t\}$ since $A^{-\delta}$ is the open ball of radius $t-\delta$
\end{proof}
\end{lemma}
\begin{lemma} Let $\|\widehat{\boldsymbol U}- \boldsymbol U\|_{\max}\lesssim_\P \varrho_U$ for some positive sequence of $n$ and $T$, then
\[
\left\| \frac{1}{\sqrt{T}} (\widehat{\boldsymbol U}\widehat{\boldsymbol U}' - \boldsymbol U\boldsymbol U') \right\|_{\max}\lesssim_\P \sqrt{T}\varrho_U^2+ \mathscr{R}_\alpha\left(\frac{ n^{9/p}}{\sqrt{T}}+\frac{n^{6/p}}{T^{1/2-1/p}}\right) +\frac{1}{n^{1/2-9/p}}.
\]
\end{lemma}
\begin{proof} By the triangle inequality we have
\begin{align*}
\left\| \frac{1}{\sqrt{T}} (\widehat{\boldsymbol U}\widehat{\boldsymbol U}' - \boldsymbol U\boldsymbol U') \right\|_{\max}\leq \left\| \frac{1}{\sqrt{T}} (\widehat{\boldsymbol U} - \boldsymbol U)(\widehat{\boldsymbol U} - \boldsymbol U)' \right\|_{\max} + 2 \left\| \frac{1}{\sqrt{T}} \boldsymbol U(\widehat{\boldsymbol U} - \boldsymbol U)' \right\|_{\max}.
\end{align*}
For the first term, we have
\begin{align*}
\left\| \frac{1}{\sqrt{T}} (\widehat{\boldsymbol U} - \boldsymbol U)(\widehat{\boldsymbol U} - \boldsymbol U)' \right\|_{\max}\leq\sqrt{T}\|\widehat{\boldsymbol U} - \boldsymbol U\|_{\max}^2 \lesssim_\P \sqrt{T}\varrho_U^2).
\end{align*}
For the second term we use decomposition (ref) to write
\begin{align*}
\frac{1}{\sqrt{T}}\sum_{t=1}^T U_{i,t}(\widehat{U}_{jt} - U_{jt}) &= \frac{1}{\sqrt{T}}\sum_{t=1}^T U_{i,t}(\widehat{\boldsymbol\lambda}_{j}'\widehat{\boldsymbol F}_t - \boldsymbol\lambda_{j}'\boldsymbol F_t + \widehat{R}_{jt} - R_{jt})\\
&= \left[(\widehat{\boldsymbol\lambda}_j - \boldsymbol H \boldsymbol\lambda_j) +\boldsymbol H\boldsymbol\lambda_j\right]'\frac{1}{\sqrt{T}}\sum_{t=1}^T U_{i,t}(\widehat{\boldsymbol F}_t- \boldsymbol H\boldsymbol F_t)\\
&\quad + \left[(\widehat{\boldsymbol\lambda}_j - \boldsymbol H \boldsymbol\lambda_j) +(\boldsymbol H'\boldsymbol H - I_r)\boldsymbol\lambda_j\right]'\frac{1}{\sqrt{T}}\sum_{t=1}^T U_{i,t}\boldsymbol F_t\\
&\quad + (\widehat{\boldsymbol\gamma}_j - \boldsymbol\gamma_j)'\frac{1}{\sqrt{T}}\sum_{t=1}^T U_{i,t}\boldsymbol X_{jt}.
\end{align*}
Apply Cauchy-Schwartz inequality in each term followed by the triangle inequality we obtain
\begin{align}
\left\| \frac{1}{\sqrt{T}} \boldsymbol U(\widehat{\boldsymbol U} - \boldsymbol U)' \right\|_{\max}
&\leq \left[\max_{j\leq n}\|\widehat{\boldsymbol\lambda}_j - \boldsymbol H \boldsymbol\lambda_j\| +\sqrt{r}\|\boldsymbol H\|\|\boldsymbol\Lambda\|_{\max}\right]\max_{i\leq n}\left\|\frac{1}{\sqrt{T}}\sum_{t=1}^T U_{i,t}(\widehat{\boldsymbol F}_t- \boldsymbol H\boldsymbol F_t)\right\|\nonumber\\
&\quad + \left[\max_{j\leq n}\|\widehat{\boldsymbol\lambda}_j - \boldsymbol H \boldsymbol\lambda_j\| +\sqrt{r}\|\boldsymbol H'\boldsymbol H - I_r\|\|\boldsymbol\Lambda\|_{\max}\right]\max_{i\leq n}\left\|\frac{1}{\sqrt{T}}\sum_{t=1}^T U_{i,t}\boldsymbol F_t\right\|\nonumber\\
&\quad + \max_{j\leq n}\|\widehat{\boldsymbol\gamma}_j - \boldsymbol\gamma_j\|\max_{i,j\leq n}\left\|\frac{1}{\sqrt{T}}\sum_{t=1}^T U_{i,t}\boldsymbol X_{jt}\right\|.
\end{align}
Recall $\delta_{i,t}:=\widehat{R}_{i,t}-R_{i,t}$. From (ref) and Holder's inequality, we conclude that
\[
\max_{i,t}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \delta_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/4}\lesssim {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \|\widehat{\boldsymbol\Sigma}_i^{-1}\|
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_p {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \|\widehat{\boldsymbol v}_i\|_\infty
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/2}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \|\boldsymbol X_{i,t}\|_\infty
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_p\lesssim \mathscr{R}_\alpha/\sqrt{T},
\]
where $\mathscr{R}_\alpha$ is defined in Theorem (ref). We now use this last result to show that
\begin{equation}
{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \|(\boldsymbol V/n)(\boldsymbol F_t - \boldsymbol H \boldsymbol F_t)\|_2
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/8} = O\left(\frac{\mathscr{R}_\alpha}{\sqrt{T}} +\frac{1}{\sqrt{n}}\right),
\end{equation}
which in turns uses identity (ref) and the fact that for any random variable $A_{st}$, by Cauchy-Schwartz inequality and the normalization $\widehat{\boldsymbol F}\widehat{\boldsymbol F}/T =\boldsymbol I_r$, we have $\|\frac{1}{T}\sum_{s=1}^T \widehat{\boldsymbol F}_s A_{st}\|\leq \sqrt{r}\left(\frac{1}{T}\sum_{s=1}^T A_{st}^2\right)^{1/2}$. Thus
\[g(A_{st}):={\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \left\|\frac{1}{T}\sum_{s=1}^T \widehat{\boldsymbol F}_s A_{st}\right\|
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/8} \leq \sqrt{r} \left[{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \left(\frac{1}{T}\sum_{s=1}^T A_{st}^2\right)^{1/2}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/8}\right].\]
\begin{enumerate}[(a)]
• Set $A_{st} = \mathbb{E}(\boldsymbol U_s'\boldsymbol U_t)/n$, then $g(A_{st})\lesssim 1/\sqrt{T}$.
• Set $A_{st}=\widetilde{\zeta}_{st}:=(\widetilde{\boldsymbol U}_s'\widetilde{\boldsymbol U}_t - \mathbb{E}(\boldsymbol U_s'\boldsymbol U_t))/n$. By the triangle inequality, ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \left(\frac{1}{T}\sum_{s=1}^T A_{st}^2\right)^{1/2}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/8}\leq \max_{s\in[T]}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \widetilde{\zeta}_{st}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/8}$, and ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \widetilde{\zeta}_{st}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/8}\leq {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \zeta_{st}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/8} + {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \zeta_{st}^*
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/8}$. The first term is $\lesssim 1/\sqrt{n}$ by Assumption (ref)(c.2). The second can be upper bounded by ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \boldsymbol U_s'\boldsymbol\delta_t/n
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/8} + {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \boldsymbol \delta_s'\boldsymbol U_t/n
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/8} +{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \boldsymbol \delta_s'\boldsymbol\delta_t/n
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/8}\lesssim \max_{i\in[n]}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert U_{i,s}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/4}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \delta_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/4} + \max_{i\in[n]}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \delta_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/4}^2$. Thus $g(A_{st}) \lesssim 1/\sqrt{n} + \mathscr{R}_\alpha/\sqrt{T} $.
• Set $A_{st} =\widetilde{\eta}_{st}:=\boldsymbol F_s'\sum_{i=1}^n \boldsymbol\lambda_i (U_{i,t}+\delta_{i,t})/n$, then apply Cauchy-Schwartz twice to obtain
\begin{align*}
g(A_{st})&\leq\left({\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \left(\frac{1}{T}\sum_{s=1}^T \|\boldsymbol F_s\|^2\right)^{1/2}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/4}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \sum_{i=1}^n\boldsymbol\lambda_{i}\frac{U_{i,t}+\delta_{i,t}}{n}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/4}\right)\\
&\lesssim{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \sum_{i=1}^n \boldsymbol\lambda_i \frac{U_{i,t}}{n}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/4} + {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \sum_{i=1}^n \boldsymbol\lambda_i\frac{\delta_{i,t}}{n}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/4}.
\end{align*}
The first term is $\lesssim 1/\sqrt{n}$ by Assumption (ref)(d) and (ref)(c.1); the second is $\lesssim \mathscr{R}_\alpha/\sqrt{T}$. Hence $g(A_{st})\lesssim \frac{1}{\sqrt{n}} + \mathscr{R}_\alpha/\sqrt{T}$.
• Set $A_{st} =\widetilde{\xi}_{st}:=\boldsymbol F_t'\sum_{i=1}^n \boldsymbol\lambda_i (U_{is}+\delta_{is})/n$, then apply Cauchy-Schwartz twice followed by the maximal inequality to obtain
\begin{align*}
g(A_{st})&\leq{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \|\boldsymbol F_t\|^{1/2}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/4}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \left(\frac{1}{T}\sum_{s=1}^T\left\|\sum_{i=1}^n\boldsymbol\lambda_{i}\frac{U_{is}+\delta_{is}}{n}\right\|^2\right)^{1/2}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/4}\\
&\lesssim \max_{s\in[T]}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \sum_{i=1}^n \boldsymbol\lambda_i \frac{U_{is}}{n}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/4} + \max_{s\in[T]}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \sum_{i=1}^n \boldsymbol\lambda_i\frac{\delta_{is}}{n}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/4}.
\end{align*}
The first term in square brackets is $\lesssim 1/\sqrt{n}$ by Assumption (ref)(d) and (ref)(e); the second is $\lesssim \mathscr{R}_\alpha/\sqrt{T}$. Hence $g(A_{st}) \lesssim \frac{1}{\sqrt{n}} + \mathscr{R}_\alpha/\sqrt{T}$.
\end{enumerate}
Finally, use the identity (ref), the triangle inequality twice and the bounds $(a)$-$(d)$ to obtain (ref). Also, by Holder's inequality
\begin{align*}
{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \|U_{i,t}(\boldsymbol V/n)(\boldsymbol F_t - \boldsymbol H \boldsymbol F_t)\|_2
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/9} &\leq {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert U_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p} {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \|U_{i,t}(\boldsymbol V/n)(\boldsymbol F_t - \boldsymbol H \boldsymbol F_t)\|_2
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/8}\\
&\lesssim \frac{ \mathscr{R}_\alpha}{\sqrt{T}} +\frac{1}{\sqrt{n}}.
\end{align*}
The first term of (ref) is $\lesssim_\P n^{9/p}\left(\frac{\mathscr{R}_\alpha}{\sqrt{T}} +\frac{1}{\sqrt{n}}\right)$ due to Lemma (ref)(a), the result above and the maximal inequality; the second term is $\lesssim_\P(\frac{\mathscr{R}_\alpha n^{4/p}}{T^{1/2-1/p}} + \frac{1}{\sqrt{n}})n^{2/p}$ since, from Theorem (ref), we might take $\varrho_R = \frac{\mathscr{R}_\alpha n^{4/p}}{T^{1/2-1/p}}$ in Theorem (ref)(b). The last term is $\lesssim_\P (\mathscr{R}_\alpha n^{3/p}/\sqrt{T})n^{4/p}$. Thus
\begin{equation}
\left\| \frac{1}{\sqrt{T}} \boldsymbol U(\widehat{\boldsymbol U} - \boldsymbol U)' \right\|_{\max} \lesssim_\P \mathscr{R}_\alpha\left(\frac{ n^{9/p}}{\sqrt{T}}+\frac{n^{6/p}}{T^{1/2-1/p}}\right) +\frac{1}{n^{1/2-9/p}}.
\end{equation}
The result then follows.
\end{proof}
\begin{lemma} Let $\|\widehat{\boldsymbol U}- \boldsymbol U\|_{\max}\lesssim_\P \varrho_U$ for some non-negative sequence $\varrho_U$ depending on $n$ and $T$, then, for all $ h >0$ and $\mathcal{D}\subseteq[n]^2$,
\[
\|\widehat{\boldsymbol\Upsilon}_\Sigma - \boldsymbol\Upsilon_\Sigma\|_{\max}\lesssim_\P h[\varrho_U(nT)^{3/p} +d^{8/p}/\sqrt{T}]
\]
in the polynomial case; and
\[
\|\widehat{\boldsymbol\Upsilon}_\Sigma - \boldsymbol\Upsilon_\Sigma\|_{\max}\lesssim_\P h[\varrho_U(\psi^{-1}_p(nT))^3 +\psi_{p/4}^{-1}(n^4)/\sqrt{T}]
\]
in the exponential case, where $d:=|\mathcal{D}|$.
\end{lemma}
\begin{proof} Let $\boldsymbol i:=(i_1,i_2,i_3,i_4)\in\mathcal{D}^2$ be a multi-index where $(i_1,i_2)\in\mathcal{D}$ and $(i_3,i_4)\in\mathcal{D}$. Define for $\boldsymbol i$ and $|\ell|<T$:
\[\widetilde{\gamma}_{\boldsymbol i}^\ell :=\frac{1}{T}\sum_{t=|\ell| +1}^T U_{i_1,t}U_{i_2,t}U_{i_3,t-|\ell|}U_{i_4,t-|\ell|};\qquad \gamma_{\boldsymbol i}^\ell := \mathbb{E}\widetilde{\gamma}_{\boldsymbol i},\]
and $\widehat{\gamma}_{\boldsymbol i}^\ell$ as $ \widetilde{\gamma}_{\boldsymbol i}^\ell$ with $U$'s replaced by $\widehat{U}$'s. Also, define
\[\widetilde{\upsilon}_{\boldsymbol i} := \sum_{|\ell|<T} k(\ell/h)\widetilde{\gamma}_{\boldsymbol i}^\ell\qquad \upsilon_{\boldsymbol i} := \sum_{|\ell|<T}\gamma_{\boldsymbol i}^\ell,\]
and $\widehat{\upsilon}_{\boldsymbol i}$ as $ \widetilde{\upsilon}_{\boldsymbol i}$ with $U$'s replaced by $\widehat{U}$'s. Then we write
\begin{align}
\widetilde{\upsilon}_{\boldsymbol i}-\upsilon_{\boldsymbol i} = \sum_{|\ell|<T} k(\ell/h)(\widetilde{\gamma}_{\boldsymbol i}^\ell - \gamma_{\boldsymbol i}^\ell) + \sum_{|\ell|<T} (k(\ell/h)-1)\gamma_{\boldsymbol i}^\ell.
\end{align}
By Cauchy-Schwartz inequality we have that ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert U_{i_1,t}U_{i_2,t}U_{i_3,t-|\ell|}U_{i_4,t-|\ell|}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{(p+\epsilon)/4}\leq C^4$, then by Lemma (ref) we obtain ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \widetilde{\gamma}_{\boldsymbol i}^\ell - \gamma_{\boldsymbol i}^\ell
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/4} = \lesssim \sqrt{T-|\ell|}/T = \lesssim 1/\sqrt{T}$,
the $\mathcal{L}_{p/4}$-norm of the first term is bounded by
\[h\sum_{|\ell|<T} |h^{-1} k(\ell/h)|{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \widetilde{\gamma}_{\boldsymbol i}^\ell - \gamma_{\boldsymbol i}^\ell
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/4} \lesssim \left( \frac{h}{\sqrt{T}}\int |k(u)|du \right)= \lesssim h/\sqrt{T},\]
whereas the second term is deterministic and is shown to be $\lesssim h/\sqrt{T}$ by andrews91. Thus ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \widetilde{\upsilon}_{\boldsymbol i}-\upsilon_{\boldsymbol i}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/4} = \lesssim h/\sqrt{T})$ uniformly in $ \boldsymbol i\in[n]^4$. Thus, union bound followed by Markov's inequality give us, for $x>0$,
\begin{equation}
\P(\max_{\boldsymbol i}|\widetilde{\upsilon}_{\boldsymbol i}-\upsilon_{\boldsymbol i}|\geq x)\leq d^2\max_{\boldsymbol i}\P(|\widetilde{\upsilon}_{\boldsymbol i}-\upsilon_{\boldsymbol i}|\geq x)\lesssim\left(\frac{d^{8/p}h}{x\sqrt{T}}\right)^{p/4}.
\end{equation}
We now use the fact that for any $\boldsymbol x:=(x_1,\dots,x_4)'\in\mathbb{R}^4$ and $\boldsymbol y:=(y_1,\dots, y_4)'\in\mathbb{R}^4$ we have $|\prod_{i=1}^4 x_i - \prod_{i=1}^4 y_i|\lesssim \sum_{i=0}^{3}\|\boldsymbol x-\boldsymbol y\|_\infty^{3-i}\|\boldsymbol y\|_\infty^i$ combined with the fact that $\|\widehat{\boldsymbol U} - \boldsymbol U\|_{\max}\lesssim 1$ to obtain
\begin{align*}
\max_{\boldsymbol i, \ell} |\widehat{\gamma}_{\boldsymbol i}^\ell - \widetilde{\gamma}_{\boldsymbol i}^\ell|&\leq \max_{\boldsymbol i,t,\ell}|\widehat{U}_{i_1,t}\widehat{U}_{i_2,t}\widehat{U}_{i_3,t-|\ell|}\widehat{U}_{i_4,t-|\ell|} - U_{i_1,t}U_{i_2,t}U_{i_3,t-|\ell|}U_{i_4,t-|\ell|}|\\
&\lesssim (\|\widehat{\boldsymbol U} - \boldsymbol U\|_{\max}\|\boldsymbol U\|_{\max}^3)\\
&\lesssim_\P \{\varrho_U[\psi^{-1}(nT)]^3\}
\end{align*}
Therefore we conclude
\begin{align}
\max_{\boldsymbol i}|\widehat{\upsilon}_{\boldsymbol i} - \widetilde{\upsilon}_{\boldsymbol i}|&\leq \max_{\boldsymbol i, \ell} |\widehat{\gamma}_{\boldsymbol i}^\ell - \widetilde{\gamma}_{\boldsymbol i}^\ell|\sum_{|\ell|<T} |k(\ell/h)|\\
&\lesssim_\P \left\{ h\varrho_U[\psi^{-1}(nT)]^3\int |k(u)|du\right\} \lesssim_\P h\varrho_U[\psi^{-1}(nT)]^3.\nonumber
\end{align}
The result then follows from the triangle inequality $\|\widehat{\boldsymbol\Upsilon} - \boldsymbol\Upsilon\|_{\max}\leq \max_{\boldsymbol i}|\widehat{\upsilon}_{\boldsymbol i} - \widetilde{\upsilon}_{\boldsymbol i}| +\max_{\boldsymbol i}|\widetilde{\upsilon}_{\boldsymbol i}-\upsilon_{\boldsymbol i}|$, expression (ref) and (ref).
\end{proof}
\begin{lemma} Let $\|\widehat{\boldsymbol U}- \boldsymbol U\|_{\max}\lesssim_\P \varrho_U$ and $\max_{i,j\in[n]}\|\widehat{\boldsymbol\chi}_{i,j} - \boldsymbol\chi_{i,j} \|_1\lesssim_\P \varrho_\chi$ for some non-negative sequences $\varrho_U$ and $\varrho_\chi$ depending on $n$ and $T$, then
\begin{enumerate}[(a)]
• $\max\limits_{(i,j)\in\mathcal{D},t\in [T]} |\widehat{V}_{i,j,t}-V_{i,j,t}|\lesssim_\P (1 + \widetilde{s}_1) \varrho_U + \varrho_\chi n^{1/p}$;
• $\max\limits_{(i,j)\in\mathcal{D}} \left|\frac{1}{\sqrt{T}}\sum_{t=1}^T(\widehat{V}_{i,j,t}\widehat{V}_{j,i,t} - V_{i,j,t}V_{i,j,t}) \right|\lesssim_\P (1+\widetilde{s}_1+\varrho_\chi)^2 \varrho_{UU} + \varrho^2_\chi n^{4/p}\sqrt{T}$,
\end{enumerate}
where $\widetilde{s}_1:=\max_{(i,j)\in\mathcal{D}}\|\boldsymbol\chi_{i,j}\|_1$ and $\varrho_{UU}$ is the rare appearing in Lemma (ref).
\end{lemma}
\begin{proof} By the triangle inequality we have, for $i,j\in[n]$ and $t\in[T]$,
\begin{align*}
\widehat{V}_{i,j,t}-V_{i,j,t}
&= \widehat{U}_{i,t} -U_{i,t} -\widehat{\boldsymbol \chi}_{i,j}' \widehat{\boldsymbol U}_{-ij,t} + \boldsymbol \chi_{i,j}'\boldsymbol U_{-ij,t}\\
&= \widehat{U}_{i,t} -U_{i,t} -\widehat{\boldsymbol \chi}_{i,j}' (\widehat{\boldsymbol U}_{-ij,t} -\boldsymbol U_{-ij,t}) - (\widehat{\boldsymbol \chi}_{i,j} -\boldsymbol \chi_{i,j})'\boldsymbol U_{-ij,t}\\
&= \widehat{U}_{i,t} -U_{i,t}-\big[\boldsymbol \chi_{i,j} + (\widehat{\boldsymbol \chi}_{i,j} -\widehat{\boldsymbol \chi}_{i,j})\big]' (\widehat{\boldsymbol U}_{-ij,t} -\boldsymbol U_{-ij,t}) -(\widehat{\boldsymbol \chi}_{i,j} -\boldsymbol \chi_{i,j})'\boldsymbol U_{-ij,t},
\end{align*}
Therefore, result (a) follows by Hölder's and the triangle inequality since
\begin{align*}
\max_{i,j,t} |\widehat{V}_{i,j,t}-V_{i,j,t}|
&\leq (\|\boldsymbol \chi_{i,j}\|_1 +\| \widehat{\boldsymbol \chi}_{i,j} -\boldsymbol \chi_{i,j}\|_1 )\|\widehat{\boldsymbol U}_{-ij,t} - \boldsymbol U_{-ij,t}\|_\infty \\
&\qquad +|\widehat{U}_{i,t}- U_{i,t}| + \|\widehat{\boldsymbol \chi}_{i,j} -\boldsymbol \chi_{i,j}\|_1\|\boldsymbol U_{-ij,t}\|_\infty \\
&\lesssim (1+\max_{(i,j)\in\mathcal{D}}\|\boldsymbol\chi_{i,j}\|_1) \|\widehat{\boldsymbol U}-\boldsymbol U\|_{\max} + \max_{(i,j)\in\mathcal{D}}\|\widehat{\boldsymbol\chi}_{i,j}-\boldsymbol\chi_{i,j}\|_1\|\boldsymbol U\|_{\max}\\
&\lesssim_\P \big[(1 +\max_{(i,j)\in\mathcal{D}}\|\boldsymbol\chi_{i,j}\|_1) \varrho_U + \varrho_\chi n^{1/p}\big].
\end{align*}
For (b), we write for $i,j\in[n]$ and $t\in[T]$
\begin{align*}
\widehat{V}_{i,j,t}\widehat{V}_{j,i,t} - V_{i,j,t}V_{i,j,t}&=\widehat{U}_{i,t}^2 - U_{i,t}^2 +\widehat{\boldsymbol\chi}_{i,j}'(\widehat{\boldsymbol U}_{-ij,t}\widehat{\boldsymbol U}_{-ij,t}' - \boldsymbol U_{-ij,t}\boldsymbol U_{-ij,t}')\widehat{\boldsymbol\chi}_{i,j}\\
&\qquad +(\widehat{\boldsymbol\chi}_{i,j} - \boldsymbol\chi_{i,j})' \boldsymbol U_{-ij,t}\boldsymbol U_{-ij,t}'(\widehat{\boldsymbol\chi}_{i,j} - \boldsymbol\chi_{i,j}).
\end{align*}
Then, by Holder's and the triangle inequality, we obtain
\begin{align*}
\max_{i,j}\left|\tfrac{1}{\sqrt{T}}\sum_{t\in[T]}(\widehat{V}_{i,j,t}\widehat{V}_{j,i,t} - V_{i,j,t}V_{i,j,t})\right| &\leq (1+\max_{i,j\in[n]}\|\widehat{\boldsymbol\chi}_{i,j}\|_1)^2\|\tfrac{1}{\sqrt{T}}(\widehat{\boldsymbol U}\widehat{\boldsymbol U}'-\boldsymbol U\boldsymbol U')\|_{\max} \\
&\quad +\max_{i,j\in[n]}\|\widehat{\boldsymbol\chi}_{i,j} - \boldsymbol\chi_{i,j}\|_1^2\|\tfrac{1}{\sqrt{T}}\boldsymbol U\boldsymbol U')\|_{\max}.
\end{align*}
Now, $\max_{i,j\in[n]}\|\widehat{\boldsymbol\chi}_{i,j}\|_1\leq \max_{i,j\in[n]}\|\boldsymbol\chi_{i,j}\|_1 +\max_{i,j\in[n]}\|\widehat{\boldsymbol\chi}_{i,j} - \boldsymbol\chi_{i,j}\|_1 \lesssim_\P \widetilde{s}_1+\varrho_\chi$. The order in probability of $\|\tfrac{1}{\sqrt{T}}(\widehat{\boldsymbol U}\widehat{\boldsymbol U}'-\boldsymbol U\boldsymbol U')\|_{\max}$ is given by Lemma (ref). By Cauchy-Schwartz inequality and Assumption (ref) we have ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert U_{i,t} U_{j,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/2}\leq {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert U_{i,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_p{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert U_{j,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_p\leq C^2$, then by the triangle inequality, ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \sum_{t\in[T]}U_{i,t}U_{j,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/2}\leq \sum_{t\in[T]}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert U_{i,t} U_{j,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p/2}\lesssim T$. Finally, by the union bound followed by Markov's inequality, $\|\tfrac{1}{\sqrt{T}}\boldsymbol U\boldsymbol U')\|_{\max}\lesssim_\P n^{4/p}\sqrt{T}$. The result follows.
\end{proof}
\begin{lemma}
Let $\|\widehat{\boldsymbol U}- \boldsymbol U\|_{\max}\lesssim_\P \varrho_U$ and $\max_{i,j\in[n]}\|\widehat{\boldsymbol\chi}_{i,j} - \boldsymbol\chi_{i,j} \|_1\lesssim_\P \varrho_\chi$ for some non-negative sequences $\varrho_U$ and $\varrho_\chi$ depending on $n$ and $T$ then
\[
\|\widehat{\boldsymbol\Upsilon}_\Pi - \boldsymbol\Upsilon_\Pi\|_{\max}\lesssim_\P h\left((1 + \widetilde{s}_1) \varrho_U + \varrho_\chi n^{1/p}(1+\widetilde{s}_1)^3(nT)^{3/p} + \frac{(1+\widetilde{s}_2)^4n^{8/p}}{\sqrt{T}}\right),
\]
where $\widetilde{s}_k:=\max_{(i,j)\in\mathcal{D}}\|\boldsymbol\chi_{i,j}\|_k$ for $k\in\{1,2\}$.
\end{lemma}
\begin{proof} The proof is parallel to the proof of Lemma (ref), refer to it for details. It suffices to bound in $\max_{i,j,t}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert V_{i,j,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p+\epsilon}$, $\max_{i,j,t}|\widehat{V}_{i,j,t}-V_{i,j,T}|$, and $\max_{i,j,t} |V_{i,j,t}|$, where the maximum is over $(i,j)\in\mathcal{D}\subseteq[n]^2$ and $t\in[T]$. For the first term we have, $\max_{i,j,t}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert V_{i,j,t}
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p+\epsilon}\leq (1+\max_{i,j}\|\boldsymbol\chi_{i,j}\|_2)\sup_{t\in[T]}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \boldsymbol U_t
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{p+\epsilon}\leq (1+M_2)C$ by Assumption (ref). Lemma (ref)(a) bounds the second term. The last one is upper bounded by $(1+\max_{(i,j)\in\mathcal{D}}\|\boldsymbol\chi_{ij}\|_1)\|\boldsymbol U \|_{\max} \lesssim_\P \left[(1+M)(nT)^{1/p}\right]$.
\end{proof}
\subsection{Orlicz Norm Results}
In this subsection, we show some useful results for Orlicz norms. All these results can be found in VW1996 for the case of $\gamma\geq 1$. Recall that for $\gamma\in (0,1)$ we define
\[
\psi_{e^\gamma}(x) := \mathsf{co} (x\mapsto \exp(x^\gamma)-1),
\]
where $\mathsf{co}(f)$ denote the convex hull of $f$.
\begin{lemma} We have
\begin{enumerate}[(a)]
• $\psi_{e^\gamma}(x) = K_\gamma x\1_{0\leq x<a_\gamma} +[\exp(x^\gamma)-1]\1_{x\geq a_\gamma}$ where $K_\gamma :=(\exp a_\gamma^\gamma-1)/a_\gamma$ and $a_\gamma:=\inf\{x\in\mathbb{R}_+: x\geq ((1-\gamma)/\gamma)^{1/\gamma}, K_\gamma\leq \gamma\exp(x^\gamma)/x^{1-\gamma}$. Also $((1-\gamma)/\gamma)^{1/\gamma}\leq a_\gamma\leq (1/\gamma)^{1/\gamma}$.
• For $p\in[1,\infty)$ and $\gamma>0$, ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_p\leq C{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}$ for some constant $C$ depending only on $p$ and $\gamma$.
• For $0<\gamma_1\leq \gamma_2$, ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma_1}}\leq C{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma_2}}$ for constant $C$ depending only on $\gamma_1$ and $\gamma_2$.
• For random variables $X_1,\dots ,X_n$, ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \max_{j\in[n]}X_j
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{\psi_{e^\gamma}}\leq C\psi^{-1}_{e^\gamma}(n)\max_{j\in[n]}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X_j
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma}} $ for some constant $C$ only depending on $\gamma$.
\end{enumerate}
\end{lemma}
\begin{proof}
For part (a), it is straightforward to verify that $x\mapsto \exp(x^\gamma)-1$ is convex on the interval $[((1-\gamma)/\gamma)^{1/\gamma}\lor 0,\infty)$. Also, $\psi_{e^\gamma}$ is continuous at $z_\gamma$ (hence continuous on $\mathbb{R}_+)$. Therefore, convexity follows as the left derivative of $\psi_{e^\gamma}$ is no larger than the right derivative at $a_\gamma$ since, by definition, $K_\gamma\leq \gamma a_\gamma^{\gamma-1} \exp(a^\gamma_\gamma)$.
For part (b), we have $x^p = x^p 1 = f(x^p) + f^*(1)$ for any $x\in\mathbb{R}_+$ and $f:\mathbb{R}_+\to \mathbb{R}\cup \{\pm\infty\}$, where $f^*$ denote the convex conjugate of $f$. Take $f(x)=\psi_{e^\gamma}(x^{1/p})$, $x = X(\omega)/{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}$ (where we assume ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}>0$, otherwise the result is trivial because $X=0$ a.s) and take expectation with respect to the law of $X$ to obtain
\[\frac{\mathbb{E}|X|^p}{{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}^p}\leq \mathbb{E}\psi_{e^\gamma}\left(\frac{X}{{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}}\right) + f^*(1)\leq 1 + f^*(1),\]
where $\psi^*(1)=\sup_{x\geq 0}\{x - \psi_{e^\gamma}(x^{1/p})\}<\infty$ and only depends on $p$ and $\gamma$, hence we might take $C= (1+\psi^*(1))^{1/p}$ to conclude.
For (d), we have that $\psi_e^\gamma$ is convex (by part (a)), nondecreasing, nonzero function vanishing at the origin. Note that $x^\gamma + y^\gamma\leq 2 (xy)^\gamma$ for $x,y\geq 1$, then $\frac{\psi_{e^\gamma}(x)\psi_{e^\gamma}(y)}{\psi_{e^\gamma}(2^{1/\gamma} xy)} \leq \frac{\exp(x^\gamma + y^\gamma)}{\exp(2(xy)^\gamma)}\leq 1$ for $x,y\geq a_\gamma\geq 1$. The result then follows from Lemma 2.2.2 in VW1996.
\end{proof}
\begin{lemma} If ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}<\infty$ for $\gamma>0$ , there are constants $C_1>0$ and $C_2>0$ such that
\[ \P(|X|>x)\leq C_1\exp[-(x/C_2)^\gamma]\qquad x>0 .\]
In particular, if $0<{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}<\infty$, we might take $C_1 = 2 + \exp \big[(a_\gamma/{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma})^\gamma\big]\1_{0<\gamma<1}$ and $C_2 = {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}$, where $a_\gamma$ is defined in Lemma (ref)(a).
Conversely, if there are constants $C_1>0$ and $C_2>0$ such that $\P(|X|>x)\leq C_1\exp[-(x/C_2)^\gamma]$ for $x>0$, then
\[{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}\leq
\begin{cases}
C_2\big[(1+2C_1)^{1/\gamma}\lor 2K_\gamma C_1 \Gamma(1/\gamma)/\gamma\big];&0< \gamma <1\\
C_2(1+C_1)^{1/\gamma}; & \gamma\geq 1,
\end{cases}
\]
where $\Gamma(\cdot)$ denotes the Gamma function.
\end{lemma}
\begin{proof} If ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}=0$ then $X=0$ a.s and the inequality holds for any choice of $C_1,C_2>0$. For the case when $0<{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}<\infty$ we have, by Markov inequality and the fact that $x\mapsto \exp \big[(x/{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma})^\gamma\big]$ is non-decreasing
\begin{align*}
\P(|X|\geq x) &=\P\left(\exp \big[(|X|/{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma})^\gamma\big]\geq\exp \big[(x/{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma})^\gamma\big]\right)\\
&\leq \exp \big[-(x/{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma})^\gamma\big]\mathbb{E} \exp \big[(|X|/{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma})^\gamma\big].
\end{align*}
Using Lemma (ref)(a), we have
\begin{align*}
\mathbb{E} \exp \big[(|X|/{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma})^\gamma\big] &=\mathbb{E} \exp \big[(|X|/{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma})^\gamma\big]\1_{|X|<a_\gamma}
+ \mathbb{E} \exp \big[(|X|/{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma})^\gamma\big]\1_{|X|\geq a_\gamma}\\
&\leq \exp \big[(a_\gamma/{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma})^\gamma\big]\1_{0<\gamma<1}+ \mathbb{E}\psi_{e^\gamma}(|X|/{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}) + 1\\
&\leq \exp \big[(a_\gamma/{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma})^\gamma\big]\1_{0<\gamma<1}+ 2.
\end{align*}
Combine the last two displays to obtain the first result.
For the converse, we have for $c>0$, by Fubini's Theorem,
\begin{align*}
\mathbb{E} \exp(|X/c|^\gamma) -1 &= \int\int_0^{|x|^\gamma}c^{-\gamma}\exp (c^{-\gamma} y) dy \P(dx)\\
&=c^{-\gamma} \int_0^\infty\P(|X|\geq x^{1/\gamma})\exp(c^{-\gamma} x) dx.
\end{align*}
Since $\P(|X|>x)\leq C_1\exp[-(x/C_2)^\gamma]$, for $c>C_2$, we have
\begin{align*}
\mathbb{E} \exp[(|X/c|)^\gamma] -1 &\leq c^{-\gamma} C_1\int_0^\infty\exp\left[-x(C_2^{-\gamma} - c^{-\gamma})\right]dx\leq \frac{c^{-\gamma} C_1}{C_2^{-\gamma}-c^{-\gamma}} = \frac{C_1}{(c/C_2)^{\gamma}-1}.
\end{align*}
Also
\begin{align*}
\mathbb{E}|X| = \int_0^\infty\P(|X|>x)dx\leq C_1\int_0^\infty\exp[-(x/C_2)^\gamma]dx = \tfrac{C_1C_2}{\gamma}\Gamma(1/\gamma).
\end{align*}
Therefore using the last two displays and Lemma (ref)(a), we have, for $c>C_2$,
\begin{align*}
\mathbb{E}\psi_{e^\gamma}(|X|/c) &\leq \frac{K_\gamma}{c}\mathbb{E}|X|\1_{0<\gamma<1}+ \mathbb{E} \exp[(|X|/x)^\gamma] -1\\
&\leq \frac{K_\gamma C_1C_2}{c\gamma}\Gamma(1/\gamma)\1_{0<\gamma<1} + \frac{C_1}{(c/C_2)^{\gamma}-1}.
\end{align*}
For $\gamma\geq 1$ the right hand side is less or equal than $1$ for $c\geq C_2(1+C_1)^{1/\gamma}$ hence ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}\leq C_2(1+C_1)^{1/\gamma}$. For $0<\gamma<1$, the right hand side is less or equal than $1$ for $c\geq C_2(1+2C_1)^{1/\gamma}\lor 2K_\gamma C_1 C_2 \Gamma(1/\gamma)/\gamma$ then ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}\leq C_2(1+2C_1)^{1/\gamma}\lor 2K_\gamma C_1 C_2 \Gamma(1/\gamma)/\gamma$.
\end{proof}
\begin{lemma} If ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}<\infty$ and ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert Y
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}<\infty$ for $\gamma>0$ then ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert XY
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma/2}}<\infty$ for $\left({\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}^2\lor {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert Y
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}^2\right)$. In particular, ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert XY
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^{\gamma/2}}\leq C\left({\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}^4\lor {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert Y
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}^2\right)$ where $C=5^{2/\gamma}$ for $\gamma\geq 1$; and $C=(1+2C_1)^{2/\gamma}\lor 2K_{\gamma/2} C_0 \Gamma(2/\gamma)/\gamma$ with $C_0 = 2(2 + \exp \big[(a_\gamma/({\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}\land{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert Y
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma} ))^\gamma\big] )$ for $\gamma\in(0,1)$ provided that ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert Y
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}\land {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert Y
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}>0$.
\end{lemma}
\begin{proof}
If ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}=0$ or ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert Y
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}=0$ then $XY=0$ a.s and the inequality holds trivially. Assume that $0<{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}<\infty$ and $0<{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert Y
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}<\infty$. From Lemma (ref) we have for $x>0$
\begin{align*}
\P(|X|>x)\leq C_X\exp[-(x/{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma})^\gamma]\\
\P(|Y|>x)\leq C_Y\exp[-(x/{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert Y
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma})^\gamma],
\end{align*}
for positive constants $C_X$ and $C_X$. Then, by the union bound,
\begin{align*}
\P(|XY|\geq x)&\leq \P(|X|\geq \sqrt{x})+ \P(|Y|\geq \sqrt{x})
\\
&\leq C_X \exp(-x^{\gamma/2}/{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}^\gamma)+ C_Y \exp\big[-x^{\gamma/2}/{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert Y
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}^\gamma)\\
&\leq 2(C_X\lor C_Y)\exp\left[-\left(\frac{x}{{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert X
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}^2\lor {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert Y
\right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{e^\gamma}^2}\right)^{\gamma/2}\right].
\end{align*}
Apply once again Lemma (ref) in the other direction to conclude.
\end{proof}