Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
{8pt}
{8pt}
titlepage\begin{center}
\begin{spacing}{1.5}
{ Panel Data Estimation and Inference: \\Homogeneity versus Heterogeneity}
\end{spacing}
$^{\ast}${\sc Jiti Gao}, $^{\dag}${\sc Fei Liu}, $^{\ast}${\sc Bin Peng} and $^{\ddag}${\sc Yayi Yan}
\begingroup
\footnote{Gao and Peng would like to acknowledge the Australian Research Council Discovery Projects Program for its financial support under Grant Numbers: DP250100063. Liu's research was financially supported by National Natural Science Foundation of China under Grant Number 72203114. Yan acknowledges the financial support by the NSFC under the grant number 72303142 and the Fundamental Research Funds for the Central Universities under grant numbers 2022110877 and 2023110099. The authors contributed equally to this paper and are credited in alphabetical order.}
\addtocounter{footnote}{-1}
\endgroup
$^{\ast}$Monash University
$^{\dag}$Nankai University
$^\ddag$Shanghai University of Finance and Economics
\today
\end{center}
\begin{abstract}
In this paper, we define an underlying data generating process that allows for different magnitudes of cross-sectional dependence, along with time series autocorrelation. This is achieved via high-dimensional moving average processes of infinite order (HDMA($\infty$)). Our setup and investigation integrates and enhances homogenous and heterogeneous panel data estimation and testing in a unified way. To study HDMA($\infty$), we extend the Beveridge-Nelson decomposition to a high-dimensional time series setting, and derive a complete toolkit set. We exam homogeneity versus heterogeneity using Gaussian approximation, a prevalent technique for establishing uniform inference. For post-testing inference, we derive central limit theorems through Edgeworth expansions for both homogenous and heterogeneous settings. Additionally, we showcase the practical relevance of the established asymptotic theory by (1). connecting our results with the literature on grouping structure analysis, (2). examining a nonstationary panel data generating process, and (3). revisiting the common correlated effects (CCE) estimators. Finally, we verify our theoretical findings via extensive numerical studies using both simulated and real datasets.
{\it Keywords}: homogeneity, heterogeneity, weak and strong cross-sectional dependence, Gaussian approximation, (non)stationary panel
{\it JEL Classification:} C12, C18, C23, C55
\end{abstract}
Introduction
Panel data analysis has seen its popularity in the past thirty years or so. Comprehensive reviews have been conducted at different stages, while the literature evolves. See, for example, ARELLANO20013229, Petersen2009, CP2015, hsiao2022analysis, etc. Among all challenges raised in different surveys, this article aims to offer a unified framework and a set of toolkit to
enumerate[leftmargin=24pt, parsep=2pt, topsep=2pt]
• simultaneously test homogeneity (i.e., $\mathbb{H}_0$) vs. heterogeneity (i.e., $\mathbb{H}_1$);
• develop valid inference under either $\mathbb{H}_0$ or $\mathbb{H}_1$, and account for both weak and strong cross-sectional dependence (WCD and SCD), as well as the dependence along the time dimension.
In what follows, we review the relevant literature, point out the challenges, and then highlight our contributions.
We start with the hypothesis testing about homogeneity against heterogeneity, which has always been a central topic in empirical studies. Without loss of generality, we consider a simple setup as follows:
eqnarray[eqnarray omitted — 65 chars of source]
where $x_{it}$, $\mu_i$, and $\epsilon_{it}$ are all scalars, $(i,t)\in [N]\times [T]$, $[L]\coloneqq \{1,\ldots,L \}$ for any given positive integer $L$, $N$ stands for the number of individuals, and $T$ stands for the total number of periods.
More often than not, the key hypotheses are
eqnarray[eqnarray omitted — 141 chars of source]
Sometimes, the hypotheses are even more straightforward, e.g.,
eqnarray*[eqnarray* omitted — 123 chars of source]
The model (ref) extends the location model of Lazarus to panel data settings. To settle (ref), a large literature (PY2008,GAO2020329; and references therein) adopts the quadratic test statistics. However, as explained in fan2015power, the tests based on quadratic forms often suffer from low powers, and fail to detect the sparse alternatives. Therefore, fan2015power and YU2024105458 provide power enhanced test statistics to tackle this issue.
To the best of our knowledge, the above literature largely (if not all) ignores the dependence along both dimensions. As time series autocorrelation (TSA) has been well discussed in the literature (see FanYao; Gao2007), we justify the necessity of accounting for the cross-sectional dependence (CD) here. As surveyed by CP2015, CD is likely to be the rule rather than the exception, and it sometimes goes beyond WCD due to omitted variables as documented in GX2021. It is then reasonable to call for a complete toolkit set that is robust to the presence of dependence.
To better present our motivations, we start with four datasets, of which $N$ and $T$ stand for the number of individuals and the number of time periods respectively.
itemize[leftmargin=12pt, parsep=2pt, topsep=2pt]
• Dataset 1: the U.S. macroeconomic dataset assembled by MN2016.
• Dataset 2: the climate data of 37 stations from the U.K. Meteorological Office.
• Dataset 3: the bank equity return data constructed by Baron2021.
• Dataset 4: the realized volatility data of 16 international stock markets.
We shall provide numerical evidences demonstrating that the four datasets exemplify both WCD and SCD, highlighting the need for a unified framework to model these varying dependencies.
For each dataset, we observe $\{x_{it}\}$ as defined in (ref), which yields the pairwise correlation $r_{ij}$ for $\forall i,j\in [N]$. For $\forall\tau \in[0,1]$, we can calculate:
eqnarray*[eqnarray* omitted — 100 chars of source]
where $I(\cdot)$ stands for the indicator function. This allows us to plot $p(\tau)$ against $\tau$, as shown in Figure (ref) below. Here, $\tau$ represents a specific correlation threshold, while $p(\tau)$ measures the percentage of absolute correlation values that exceed the threshold. We also calculate the following measure:
eqnarray[eqnarray omitted — 97 chars of source]
which is adopted from Assumption C of BN2002, and represents the magnitude of cross-sectional dependence of a panel dataset. Having presented these measures, we proceed.
Dataset 1 --- We examine a time period spanning from October 2003 to September 2023, resulting in a total of $T = 240$ observations along the time dimension. After removing variables with missing values, we are left with $N = 127$ macro variables. In the first sub-figure of Figure (ref), we observe that around 20% of the absolute correlations $\{|r_{ij}|\mid i>j\}$ are greater than 0.8, and roughly more than 50% of these correlations have absolute values exceeding 0.5. Additionally, $\overline{\rho}=66$ is approximately $N/2$. Thus, we find a high degree of correlation among these macro variables. An intuitive thought is that many of these macro variables are generated by the same set of unobservable shocks.
Dataset 2 --- We analyze temperature and sunshine data from the U.K. Meteorological office. There are 37 stations in total, widely distributed across the U.K. Each station reports both temperature and sunshine monthly, resulting in $N = 74$ individual time series. After handling missing values, we focus on the period from January 1950 to February 2023, resulting in $T = 878$ observations. In the second sub-figure of Figure (ref), over 80% of the absolute correlations $\{|r_{ij}|\mid i>j\}$ are greater than 0.5, and approximately 30% of these correlations exceed 0.8. We have $\overline{\rho}=52$, which is almost the same as the sample size of the individual dimension. Thus, it is evident that these climate data exhibit a high degree of correlation.
Dataset 3 --- We study a dataset of real bank equity returns for 46 advanced and emerging economies ($N=46$) in the period from 1870 to 2016 ($T=147$). This dataset, constructed by Baron2021, illustrates the influence of banking crises on subsequent output gaps and credit contractions. The estimated values for $p(\tau)$ are presented in the third sub-figure of Figure (ref). Obviously, this dataset has the weakest cross-sectional dependence among the four datasets, and most of $r_{ij}$'s are less than 0.6, which might be a signal of WCD. The value $\overline{\rho}$ of $\eqref{def.rho}$ is 13 which is relatively small.
Dataset 4 --- A final example is a realized volatility dataset, which exhibits strong connections across different equity markets. This dataset consists of realized volatility data for 16 international stock market indices ($N=16$) and is computed using tick-by-tick stock index data from Refinitiv DataScope Select. The dataset covers 3809 common trading days ($T=3809$) over a 16-year period from January 4, 2005, to February 26, 2021. For this dataset, we compute the pairwise correlations between realized volatilities and present the estimated values of $p(\tau)$ in the fourth sub-figure of Figure (ref). More than 50% of the correlations are above 0.5, indicating a strong connectedness in realized volatility across markets. The value $\overline{\rho}$ of $\eqref{def.rho}$ is 8.
figure[figure omitted — 127 chars of source]
Four datasets offer examples of CD with different magnitude. The issue at hand is certainly a cause for concern, as inferring $\mu_i$ of (ref) is a cornerstone of data analysis (e.g., Hamilton1994; Lazarus; CP2024). Without a grasp of the dependence magnitude, the existing literature offers scant guidance on how to infer the homogenous/heterogeneous means, let alone more intricate scenarios including homogenous/heterogeneous trends (ROBINSON20124,WSX2023). See Example (ref) and Example (ref) of Section (ref) for the purpose of demonstration. To our knowledge, although CP2015 formalize the definitions of WCD and SCD, only Assumption 3 of goncalves_2011 addresses this issue using a set of high-level conditions. Yet, the underlying data generating mechanism remains underexplored. In related research, ROBINSON2011, ROBINSON2012 and LEE2016 impose a linear system structure to represent cross-sectional dependence. This approach, while convenient in theory, is infeasible in practice due to the fact that there is no natural ordering in the cross-sectional dimension. When solely WCD is of interest, Assumption C of BN2002, which regulates dependence across cross-sections and time using various moments, is often cited. Nonetheless, the underlying data generating process remains somewhat obscure.
Considering the aforementioned points, our contributions in this paper are as follows:
enumerate[leftmargin=24pt, parsep=2pt, topsep=2pt]
• First, we define an underlying data generating process that allows for different magnitude of CD, along with TSA. This is achieved via high-dimensional moving average processes of infinite order (HDMA($\infty$)), which automatically generalizes the spatial structure introduced by Robinson and his co-authors in recent years. The framework is important in the sense that as noted by brockwell1991time and FanYao, the Wold decomposition theorem ensures a formal linear representation exists for any stationary time series with no deterministic components, and HDMA($\infty$) naturally incorporates this result into a panel data framework.
• To the best of our knowledge, HDMA($\infty$) has not been carefully explored in the literature of panel data analysis. Our setup and investigation significantly integrates and enhances both homogenous and heterogeneous panel data modelling and testing (such as Pesaran2006, PY2008, fan2015power, YU2024105458). To study HDMA($\infty$), we extend the BN decomposition (e.g., BN1981,PS1992) to a high-dimensional time series setting, and derive a complete set of toolkit.
Our development offers theoretical justification for some high level assumptions of the literature (e.g., goncalves_2011, BN2002), and also generalizes Assumption 2 of Pesaran2006 and many extensions since then. Additionally, our investigation complements the work of fan2015power, who specifically study cases where $\frac{T}{\sqrt{N}}\to 0$, by considering a broader range of scenarios and relaxing the independence assumptions employed in PY2008 and YU2024105458.
• We exam homogeneity against heterogeneity using Gaussian approximation, a prevalent technique for establishing uniform inference (e.g., chernozhuokov2022improved, and references therein). For post-testing inference, we derive Central Limit theorems through Edgeworth expansions for both homogenous and heterogeneous settings. Notably, the demand for Gaussian approximation in panel data analysis has been increasing recently, as exemplified in Section 4 of SJW_2024 and Section 4 of LLS2024. Our study also contributes to this research direction by providing a set of foundational conditions and deriving a set of useful basic results.
• We showcase the practical relevance of the established asymptotic theory by (1). connecting our results with the literature on grouping structure analysis such as those surveyed in BM2015 and SSP2016, (2). examining a nonstationary panel data generating process presented in PM1999, and (3). revisiting the common correlated effects (CCE) estimators of Pesaran2006. Typically, when investigating nonstationary panel data, one has to impose cross-sectional independence such as PM1999, DGP2021 and HUANG2021198 due to technical constraints. Our study offers a set of complete toolkit to account for the dependence of unit root precesses.
• Finally, we evaluate our theoretical findings via extensive numerical studies using both simulated and real datasets.
The remainder of this paper is structured as follows. In Section (ref), we present the underlying data generating process in detail, and show its practical relevance. The corresponding asymptotic properties under homogenous (i.e., $\mathbb{H}_0$ of (ref)) and heterogeneous (i.e., $\mathbb{H}_1$ of (ref)) settings are given in Sections (ref) and (ref) respectively. Building on Sections (ref) and (ref), we provide the test statistic to exam (ref) in Section (ref). In Section (ref), we revisit the CCE estimators of Pesaran2006, and some results presented in PM1999 to showcase the practical relevance of the results of Section (ref). Section (ref) conducts extensive simulation studies to exam our theoretical results. An empirical study is given in Section (ref) to exam whether the rational expectations of financial markets are homogeneous or heterogeneous. Section (ref) concludes with a few remarks. The preliminary lemmas and proofs are regulated to the online appendices.
Notations --- Before proceeding, we introduce some notations and present a few useful facts to facilitate development. Throughout, vectors and matrices are always in bold font. We let $\mathsf{i}$ be the imaginary unit; for a matrix $\mathbf{A}=\{a_{ij}\}_{m\times n}$, let $\mathbf{A}^+$ define Moore-Penrose inverse, and let
eqnarray*[eqnarray* omitted — 326 chars of source]
define respectively its Spectral norm, column norm, row norm, and entry wise norm; moreover, we always use $^\dag$ and $^\sharp$ to represent the column and row of a matrix, e.g.,
\[
\mathbf{A} =(\mathbf{a}_1^\dag,\ldots,\mathbf{a}_n^\dag)=(\mathbf{a}_1^\sharp,\ldots,\mathbf{a}_m^\sharp)^\top .
\]
Given two conformable matrices $\mathbf{A}$ and $\mathbf{B}$, we let $\mathbf{A}\circ \mathbf{B}$ denote its Hadamard product. For a vector $\mathbf{v} =(v_1,\ldots, v_p)^\top$, we let $|\mathbf{v}|_\infty\coloneqq \max_{i}v_i$. $\mathbf{I}_p$ stands for a $p\times p$ identify matrix, and when no misunderstanding arises, we write $\mathbf{I}$. $\mathbf{1}_p$ stands for a $p\times 1$ vector of ones. $\mathbf{e}_j$ always stands for a selection column vector with the $j^{th}$ individual being 1 and others being 0. For two positive constants $a$ and $b$, $a\asymp b$ stands for $a=O(b)$ and $b=O(a)$; for $a,b\in \mathbb{R}$, $a\wedge b=\min \{a,b\}$ and $a\vee b=\max \{a,b\}$. We always let $\Phi(x)$ and $\phi(x)$ be the CDF and PDF of the standard normal distribution, and let $\widetilde{\phi}(x)\coloneqq \exp(-x^2/2)$ for notational simplicity. Thus, $\sqrt{2\pi}\phi(x)=\widetilde{\phi}(x)$. The cumulant generating function of a random variable $x$ is defined by $C(u)=\log E[\exp(u x)],$ and we have
eqnarray*[eqnarray* omitted — 109 chars of source]
where $\psi(u)$ defines the characteristic function of $x$. Finally, $E^*(\cdot)$ and $\text{Pr}^*(\cdot)$ always refer to the operations induced by the sample space.
The Setup and Asymptotic Properties
We firstly present the main results achieved in this paper, and then outline the establishment of some key results before presenting any asymptotic results.
figure[figure omitted — 1,371 chars of source]
enumerate• Edgeworth expansion for both homogeneous and heterogeneous cases. Our derivation relies on the high-dimensional Beveridge-Nelson (BN) decomposition, as panel data can be viewed as a high-dimensional (HD) time series when stacked along the cross-sectional dimension. The main structure of the development for Theorems (ref) and (ref) is established by equation (ref) and Lemma (ref), while some essential technical details are provided in Lemmas (ref) and (ref).
The strategy is the same for both homogeneous and heterogeneous cases overall, although in the heterogeneous case we only take summation over the time dimension for each individual. Our development offers theoretical justification for some high level assumptions of the literature (e.g., goncalves_2011, BN2002), and also generalize Assumption 2 of Pesaran2006 and many extensions since then.
• Bootstrap statistics for both homogeneous and heterogeneous cases. The key technique for investigating our bootstrap statistics relies on the properties of the \(m\)-dependent time series. In deriving Edgeworth expansions for the homogeneous and heterogeneous bootstrap statistics, we adopt the proof strategy of tik1981, which decomposes the \(m\)-dependent series into a sequence of segments. Elements within each segment may remain dependent, while different segments are 1-dependent conditional on the observed data.
By applying the asymptotic results for each segment (established in Lemmas (ref) and (ref)) and exploiting the independence of nonadjacent segments, we extend the Edgeworth expansion results from independent data to the \(m\)-dependent bootstrap series in Theorems (ref) and (ref) for the homogeneous and heterogeneous statistics, respectively.
• Gaussian approximation for uniform inference. Applying the martingale decomposition technique to HDMA($\infty$), we establish that its partial sum processes can be uniformly approximated by the summation of independent vectors with negligible approximation errors. Building on this result, the Gaussian approximation theorem for high-dimensional independent random vectors developed in chernozhuokov2022improved is used to justify the validity of Gaussian approximation for HDMA($\infty$) in Theorem (ref).
Based on Theorem (ref), the continuity of the maximum of the Gaussian distribution and the robust inference of high-dimensional covariance matrix estimation proposed in gao2024robust, we derive high-dimensional Gaussian multiplier bootstrap approximations in Theorem (ref) for inference purpose.
To proceed, we recall the notations defined in the end of Section (ref), and present a result which is independent of the assumptions to be adopted, and facilitates the Edgeworth expansion for both homogeneous and heterogeneous cases.
lemmaLet $\{ H_n(x)\mid n\ge 0\}$ be Probabilist's Hermite polynomials. The Fourier transformation of $\phi(x) H_n(x)$ is $ \widetilde{\phi}(x)(\mathsf{i}x)^n$.
The result is built on the generating function of the Probabilist's Hermite polynomials, and, together with (ref), naturally connects cumulants of a random variable with the density function of a standard normal distribution.
We are now ready to formulate our ideas. To exam (ref), we need to have a good understanding about both homogenous (i.e., $\mathbb{H}_0$) and heterogeneous (i.e., $\mathbb{H}_1$) cases. The investigation does not only offer post-testing inference under either $\mathbb{H}_0$ or $\mathbb{H}_1$, but also helps to establish the test statistic. Having said that, we respectively investigate (ref) under the null $\mathbb{H}_0$ in Section (ref), and under the alternative $\mathbb{H}_1$ in Section (ref). In Section (ref), we assemble the results of both sections to finalize the test statistic.
Firstly, we explain the necessity of accounting for WCD and SCD in a unified framework. For simplicity, suppose that $\mathbb{H}_0$ holds, and define the following HDMA($\infty$) process for $\{x_{it}\}$ of (ref):
eqnarray[eqnarray omitted — 102 chars of source]
where $\mathbf{x}_t\coloneqq (x_{1t},\ldots, x_{Nt})^\top$, $\mathbf{B}(L)\coloneqq \sum_{\ell=0}^{\infty} \mathbf{B}_\ell L^\ell$ with $L$ being the lag operator, $\{\mathbf{B}_\ell\mid \ell\ge 0\}$ is a set of $N\times N$ matrices, $\pmb{\varepsilon}_t =(\varepsilon_{1t},\ldots, \varepsilon_{Nt})^\top$, and $\{\varepsilon_{it}\}$ are independent and identically distributed (i.i.d.) over both $(i,t)$ with mean 0. Each $\mathbf{B}_\ell$ admits the following representation:
eqnarray[eqnarray omitted — 183 chars of source]
The HDMA($\infty$) of (ref) offers the flexibility to account for different types of dependence. For example, simple algebra (Appendix (ref)) shows three types of dependence as follows:
itemize[leftmargin=24pt, parsep=2pt, topsep=2pt]
• CD: $\text{Cov}(x_{i1},x_{j1}) = \sum_{\ell =0}^{\infty} \mathbf{b}_{\ell i}^{\sharp\top} \mathbf{b}_{\ell j}^\sharp$;
• TSA: $\text{Cov}(x_{it},x_{is}) = \sum_{\ell=0}^{\infty} \mathbf{b}_{\ell+t-s, i}^{\sharp\top} \mathbf{b}_{\ell i}^\sharp$ for $t>s$;
• $\text{CD}+\text{TSA}$: $\text{Cov}(x_{it},x_{js}) =\sum_{\ell=0}^{\infty} \mathbf{b}_{\ell+t-s, i}^{\sharp\top} \mathbf{b}_{\ell j}^\sharp$ for $t>s$.
Loosely speaking, we require the following conditions to hold: (i) CD does not necessarily shrink as $|i-j|$ increases, which is evident in view of Example (ref) below; (ii) TSA shrinks as $t-s$ increases, and may vary with respect to $i$; (iii) $\text{CD}+\text{TSA}$ inherits the properties of (i) and (ii).
To see the difference between WCD and SCD, we provide the following examples.
exampleWhen $\mathbf{B}_0=\mathbf{I}$ and $\mathbf{B}_\ell=\mathbf{0}$ for $\ell\ge 1$, we have $\mathbf{x}_t=\mu\cdot \mathbf{1}_N+ \pmb{\varepsilon}_t$, which gives a set of i.i.d. panel data over both dimensions. This is an extreme case of WCD, and
\begin{eqnarray*}
\frac{1}{\sqrt{NT}}\sum_{i=1}^N\sum_{t=1}^T(x_{it}-\mu) =\frac{1}{\sqrt{NT}}\sum_{t=1}^T\sum_{i=1}^N\varepsilon_{it},
\end{eqnarray*}
which will help us to drive the asymptotic distribution. In this case $\|\mathbf{B}_0\|_2=1<\infty$.
exampleSuppose that $\mathbf{x}_t$ is generated as follows.
\begin{eqnarray*}
\mathbf{x}_t =\mu\cdot \mathbf{1}_N + \begin{pmatrix}
\frac{1}{\sqrt{N}} & \cdots & \frac{1}{\sqrt{N}} \\
\vdots & \ddots & \vdots \\
\frac{1}{\sqrt{N}} & \cdots & \frac{1}{\sqrt{N}} \\
\end{pmatrix} \pmb{\varepsilon}_t,
\end{eqnarray*}
which infers $\mathbf{B}_0 =\mathbf{1}_N (\mathbf{1}_N^\top/\sqrt{N}) $ and $\mathbf{B}_\ell=\mathbf{0}$ for $\ell\ge 1$. Apparently, each time series $\{x_{it}\mid t\in[T]\}$ is perfectly correlated with the others, and simple algebra gives that
\begin{eqnarray*}
\frac{1}{N\sqrt{T}}\sum_{i=1}^N\sum_{t=1}^T(x_{it} - \mu ) =\frac{1}{\sqrt{NT}}\sum_{t=1}^T\sum_{i=1}^N \varepsilon_{it}.
\end{eqnarray*}
Therefore, SCD requires a different normalizer to derive the asymptotic distribution, which directly affects the construction of the confidence interval in practice. Notably, $\|\mathbf{B}_0\|_2=\sqrt{\lambda_{\max} (\frac{1}{N}\mathbf{1}_N \mathbf{1}_N^\top\mathbf{1}_N \mathbf{1}_N^\top )}=\sqrt{N}$, which diverges as $N\to \infty$.
A few findings emerge in view of both examples. First, (ref) extends ROBINSON2011 and offers detailed data generating process for goncalves_2011. Second, inferring $\mu$ under WCD and SCD respectively requires different normalizers to construct the standard deviations. A challenge arises naturally, as one needs to decide the magnitude of CD prior to analysis. Third, the magnitude of cross-sectional dependence will mathematically influence the Spectral norm of $\mathbf{B}_\ell$'s from the modelling perspective. In what follows, we shall account for these findings in our investigation.
Inference under $\mathbb{H}_0$
Firstly, we infer the homogenous mean via the following statistic:
eqnarray[eqnarray omitted — 175 chars of source]
where $L_N$ is generic notation, and varies with respect to the magnitude of CD. Obviously, $L_N=N$ for Example (ref), and $L_N=N^2$ for Example (ref). Here, $\sigma_x^2 \coloneqq \lim_{N,T}\frac{1}{L_N T}E[S_{NT}^2]$.
To facilitate development, we present the first assumption.
assumption• \begin{enumerate}[leftmargin=24pt, parsep=2pt, topsep=2pt]
• $\{\varepsilon_{it}\}$ are i.i.d. over both $i$ and $t$, and satisfy that $E[\varepsilon_{it}]=0$ and $E[\varepsilon_{it}^2]=1$. In addition, $\varepsilon_{11}$ has the characteristic function $\psi(u)\coloneqq E[\exp(\mathsf{i}u\varepsilon_{11})]$ with $u\in \mathbb{R}$, and has cumulants $\kappa_r$ for $r\in [J]$ with a fixed $J\ (\ge 4)$.
• Suppose that for $L_N\in [N, N^2]$,
\begin{enumerate}[leftmargin=24pt, parsep=2pt, topsep=2pt]
• $\limsup_{N}\sum_{\ell =0}^{\infty}\ell \cdot C_{N\ell}<\infty$, where $C_{N\ell}\coloneqq \sqrt{\frac{N}{L_N}}\|\mathbf{B}_\ell\|_2$;
• $\limsup_{N}\sqrt{\frac{N}{L_N}}\|\mathbf{B} \|_1<\infty$, where $\mathbf{B}\coloneqq \mathbf{B}(1)$.
\end{enumerate}
\end{enumerate}
Assumption (ref).1 is standard. See saulis for detailed definition of cumulants. In general, $\mathbf{B}_\ell\ne \mathbf{0}$ for $\ell \ge 1$, so we need to regulate the elements of $\mathbf{B}_\ell$ using Assumption (ref).2, which ensures that after suitable normalization, the Spectral norms of $\mathbf{B}_\ell$'s are summable. Obviously, Assumption (ref).2 is satisfied for both Examples (ref) and (ref). While $L_N\in (N, N^2)$, the magnitude of cross-sectional dependence is in between both cases. By Lemma (ref) of the appendix, Assumption (ref) entails that
eqnarray*[eqnarray* omitted — 108 chars of source]
which does not only incorporate long run covariance along the time dimension, but also accounts for WCD and SCD automatically.
Having these conditions in hand, we present the first main result of this paper.
theoremUnder Assumption (ref), as $(N,T)\to (\infty,\infty)$,
\begin{eqnarray*}
\sup_{u\in\mathbb{R}}\left|F_{NT}(u)-\Phi(u)-\frac{\beta_3}{6} (1-u^2)\phi(u) \right|=O\left(NT\left(\frac{\|\mathbf{B} \|_1}{\sqrt{L_NT}}\right)^4 \vee \frac{1}{T^2}\right),
\end{eqnarray*}
where $F_{NT}(u)\coloneqq \Pr(\widetilde{S}_{NT}\le u) $, $\beta_3$ is defined in (ref) for the sake of presentation, and
\[
|\beta_3|=O\left(NT\left(\frac{\|\mathbf{B} \|_1}{\sqrt{L_NT}}\right)^3 \vee \frac{1}{T^{3/2}}\right).
\]If $\kappa_3= 0$, then $\beta_3= 0$.
Theorem (ref) gives an Edgeworth expansion for $\widetilde{S}_{NT}$, and a considerably simplified form will be presented in Corollary (ref) under a slightly more restrictive condition. The detailed definition of $\beta_3$ is omitted here due to its cumbersome notation. It is worth mentioning that if the skewness of $\varepsilon_{it}$ is 0 (i.e., $\kappa_3= 0$), the term $\frac{\beta_3}{6} (1-u^2)\phi(u)$ automatically vanishes. Then $F_{NT}(u)$ converges to $\Phi(u)$ in a much faster rate. In general, we do not have $\kappa_3= 0$, so we keep the statement of Theorem (ref) as it is.
It is noteworthy that for either Example (ref) or Example (ref), the result reduces to
eqnarray*[eqnarray* omitted — 149 chars of source]
where $|\beta_3|=O(\frac{1}{\sqrt{NT}} \vee \frac{1}{T^{3/2}})$. Sequentially, for both examples, Theorem (ref) gives the following Berry-Esseen bound:
eqnarray*[eqnarray* omitted — 121 chars of source]
which says that, for the panel data with TSA and WCD/SCD, the Berry-Esseen bound is usually at the order of $\frac{1}{\sqrt{NT}}$ unless $N\gg T$ such that $\frac{T}{\sqrt{N}}\to 0$. The presence of the term $\frac{1}{T^{3/2}}$ is due to the BN decomposition which introduces a truncation residual along the time dimension. In a typical panel data setting $N\asymp T$, the truncation residual is negligible. Our result complements the work of fan2015power, who specifically study cases where $\frac{T}{\sqrt{N}}\to 0$, by considering a broader range of scenarios and relaxing the independence assumptions employed in PY2008 and YU2024105458.
With a minor additional restriction, the following corollary holds.
corollaryUnder Assumption (ref), as $(N,T)\to (\infty,\infty)$ and $\frac{N}{T^2}\to 0$,
\begin{eqnarray*}
\sup_{u\in\mathbb{R}}\left|F_{NT}(u)-\Phi(u)-\frac{\beta_3^*}{6} (1-u^2)\phi(u) \right|=O_P\left(\frac{\sqrt{N}\|\mathbf{B}\|_1^2}{L_NT}\right),
\end{eqnarray*}
where
\[
\beta_3^*\coloneqq \frac{\kappa_3}{\sigma_x^{3/2}L_N^{3/2}T^{1/2}} \sum_{j=1}^N (\mathbf{1}_N^\top \mathbf{b}_j^\dag )^3
\] with $\mathbf{b}_j^\dag $ being the $j^{th}$ column of $\mathbf{B}$, and $\kappa_3$ is defined in Assumption (ref).1.
The condition $\frac{N}{T^2}\to 0$ is rather common in the literature of panel data analysis, and is apparently fulfilled given $N\asymp T$. Using Assumption (ref).2, it is straightforward to see that $\frac{\sqrt{N}\|\mathbf{B}\|_1^2}{L_NT} =\frac{N\|\mathbf{B}\|_1^2}{L_N}\cdot \frac{1}{\sqrt{N}T}\to 0.$
Bootstrap Inference --- Below, we provide a bootstrap procedure to conduct inference. Given the nature of panel data, an intuitive thought is to draw a set of random variables, say, $\{\zeta_{it}\mid i\in [N], t\in [T] \}$ under certain restrictions, and construct the bootstrap counterpart of $S_{NT}$ as follows:
eqnarray*[eqnarray* omitted — 123 chars of source]
It is then straightforward to obtain that
eqnarray[eqnarray omitted — 388 chars of source]
where $\pmb{\zeta}_t=(\zeta_{1t},\ldots, \zeta_{Nt})^\top$, $\mathbf{W}_{ts} =E[\pmb{\zeta}_t\pmb{\zeta}_s^\top]$ for $t\ne s$, and $\mathbf{W}_{0}\coloneqq \mathbf{W}_{tt}$ for simplicity. Correspondingly, we can write $\widetilde{S}_{NT}^2 $ as follows:
eqnarray[eqnarray omitted — 404 chars of source]
Comparing the right hand sides of $E^*(\widetilde{S}_{NT}^{*2})$ and $\widetilde{S}_{NT}^2$, it is obvious that we need (ref) and (ref) to mimic (ref) and (ref) respectively. As $\mathbf{1}_N \mathbf{1}_N^\top $ has rank one, by Lemma (ref) the most obvious form of $\mathbf{W}_{ts}$ should be
eqnarray*[eqnarray* omitted — 77 chars of source]
where $\alpha_{ts}$ is a scalar varying with respect to the distance between $t$ and $s$. Therefore, we conclude that, to have a valid bootstrap procedure, one should replace $\zeta_{it}$ with $\zeta_t$ to ensure there is no cross-sectional variation. The finding nicely fits our study, as we assume no prior information about the magnitude of CD. By doing so, we preserve the dependence along the individual dimension in the bootstrap draws, and also avoid imposing certain order on individuals implicitly.
Formally, the bootstrap procedure is as follows.
\hrule
enumerate[leftmargin=24pt, parsep=2pt, topsep=2pt]
• Draw the bootstrap version of $S_{NT}$ by
\begin{eqnarray*}
\widetilde{S}_{NT}^*\coloneqq\frac{1}{\sigma_x\sqrt{L_NT}}\sum_{i=1}^N\sum_{t=1}^T(x_{it}-\mu)\zeta_{t},
\end{eqnarray*}
where $\zeta_{t}$ is an $m$-dependent time series.
• Repeat the above procedure $R$ times to obtain the sampling distribution of $\widetilde{S}_{NT}^*$.
\hrule
Accordingly, we impose the following assumption.
assumption• \begin{enumerate}[leftmargin=24pt, parsep=2pt, topsep=2pt]
• Let $a(\cdot)$ be a symmetric Lipschitz continuous kernel defined on $[-1, 1]$ such that $a(0)=1$ and $\lim_{|x|\to 0 } \frac{1-a(x)}{|x|^{q_a}} = C_{q_a}$ for $q_a \in \{1,2\}$ and $0 < C_{q_a} < \infty$. Assume $\limsup_{N}\sum_{\ell =0}^{\infty}\ell^{q_a} \cdot C_{N\ell}<\infty$, and $\int_{-1}^1a(u) \exp(-\mathsf{i}u x) \mathrm{d}u\ge 0$ for $x\in \mathbb{R}$.
• Let $E[\zeta_{t}]=0$, $E[\zeta^2_{t}]=1$, $E|\zeta_{t}|^4<\infty$, and $E[\zeta_{t}\zeta_{s}]=a(\frac{t-s}{m})$ for $\forall t, s\in [T]$, where $m$ satisfies $m\rightarrow\infty $ and $\frac{m}{\sqrt{T}}\rightarrow 0$ as $T\rightarrow\infty$.
\end{enumerate}
Assumption (ref) is a typical assumption in the literature of dependent wild bootstrap. We refer interested readers to Shao2015 for a comprehensive review of this line of research. Several conventional kernel functions satisfy such conditions. For example, for the Bartlett kernel, $q_a= 1$ and $C_1 = 1$; for the Parzen, Tukey-Hanning, QS kernels, and the trapezoidal functions, $q_a= 2$ and the values of $C_2$ vary and all satisfy $C_2<\infty$. See Andrews1991 for comments on different kernel functions. Notably, when the Bartlett kernel is adopted, the condition $\limsup_{N}\sum_{\ell =0}^{\infty}\ell^{q_a} \cdot C_{N\ell}<\infty$ reduces to Assumption (ref).2.a.
Using Assumptions (ref)-(ref), the following lemma and theorem hold for the bootstrap procedure.
lemmaUnder Assumptions (ref)-(ref), as $(N,T)\to (\infty,\infty)$, the following results hold:
\begin{enumerate}[leftmargin=24pt, parsep=2pt, topsep=2pt]
• $E^*[(\widetilde{S}_{NT}^*)^2]= 1+o_P(1)$;
• $E^\ast[(\sum_{t=(s-1)m+1}^{sm} \mathbf{1}_N^\top \mathbf{x}_t\zeta_t)^2]=O_P(mL_N)$, for $s=1,\ldots,\lfloor \frac{T}{m} \rfloor$;
• $E^\ast[(\sum_{t=\lfloor\frac{T}{m}\rfloor m+1}^{T} \mathbf{1}_N^\top \mathbf{x}_t\zeta_t)^2]=O_P(mL_N)$.
\end{enumerate}
Lemma (ref) summarizes some basic statistical properties to facilitate the investigation on the bootstrap statistic for the homogeneous cases. With them in hand, we present the following theorem.
theoremUnder Assumptions (ref)-(ref), as $(N,T)\to (\infty,\infty)$,
\begin{enumerate}[leftmargin=24pt, parsep=2pt, topsep=2pt]
• $\sup_{u\in \mathbb{R}}\left|\text{\normalfont Pr}^*(\widetilde{S}_{NT}^* \le u ) -\Phi(u)-\frac{1}{6\widetilde{\sigma}^{\ast3}}(1-u^2)E^\ast [\widetilde{S}_{NT}^{*3}]\phi(x)\right|=O_P\left(\frac{m}{T}\right)$;
• $\sup_{u\in \mathbb{R}}\left|\text{\normalfont Pr}^*(\widetilde{S}_{NT}^* \le u) -\Pr(\widetilde{S}_{NT}\le u)\right|=O_P\left(\sqrt{\frac{m}{T}}\right)$;
• $\text{\normalfont MSE}(\widetilde{\sigma}^{\ast 2})=\frac{2m}{T}\int_{-1}^1 a^2(u)du+\frac{C_{q_a}^2}{ \sigma_x^4m^{2q_a}}\Delta_{q_a}^2+o_P(m^{-2q_a})+o\left(\frac{m}{T}\right)$;
\end{enumerate}
where $\widetilde{\sigma}^{\ast2}= E^\ast[\widetilde{S}_{NT}^{*2}]$ and $\Delta_{q_a}=L_N^{-1}\sum_{s=-\infty}^{\infty}\sum_{\ell=0}^{\infty} |s|^{q_a}\mathbf{1}_N^\top \mathbf{B}_\ell \mathbf{B}_{\ell+|s|}^\top \mathbf{1}_N $.
Theorem (ref).1 establishes the first-order Edgeworth expansion for the bootstrap statistic $\widetilde{S}_{NT}^*$, extending the results of tik1981 to the panel data framework. Conditional on the sample, the rates of the first two results of Theorem (ref) are optimal, as the bootstrap draws $\{\zeta_t\}$ are $m$-dependent time series data only. Theorem (ref).3 demonstrates that the bootstrap covariance estimator can consistently estimate the true covariance, and infers that MSE is minimized at
$$m_{\text{opt}}= \big(\frac{C_{q_a}\Delta_{q_a}}{2\sigma_x^4\int_{-1}^1 a^2(u)du}\big)^{2/(2q_a+1)} T^{1/(2q_a+1)}.$$
Inference under $\mathbb{H}_1$
Under $\mathbb{H}_1$, the model (ref) admits the following vector form:
eqnarray[eqnarray omitted — 89 chars of source]
where $\pmb{\mu}=(\mu_1,\ldots, \mu_N)^\top$, and the rest settings are identical to those in (ref).
Inferring heterogeneity is slightly more complicated, as there are two options:
enumerate[leftmargin=24pt, parsep=2pt, topsep=2pt]
• Infer a specific individual $\mu_i$;
• Infer $\pmb{\mu}$ as a whole.
It is noteworthy that, when studing each $\mu_i$, the conditional expectations and probabilities involved in Lemma (ref) and Theorem (ref) are still random variables, we therefore do not take $\max$ over $i$ in these results. In order to study the uniform inference over all individuals (i.e., inferring $\pmb{\mu}$ as a whole), we provide Theorems (ref) and (ref), which will help us establish the corresponding test statistic in Section (ref).
We start with the first choice, and consider the following quantify:
eqnarray*[eqnarray* omitted — 59 chars of source]
where $\widetilde{p}_i\coloneqq \frac{1}{\sigma_{p,i}\sqrt{T}}p_i\coloneqq \frac{1}{\sigma_{p,i}\sqrt{T}} \sum_{t=1}^T(\mathbf{e}_i^\top \mathbf{x}_t-\mu_i)$, $\sigma_{p,i}^2 \coloneqq \lim_N\mathbf{e}_i^\top \mathbf{B}\mathbf{B}^\top \mathbf{e}_i>0$, and $\mathbf{e}_i$ is a selection vector as defined in Section (ref).
To proceed, we need more structures to investigate heterogeneity.
assumptionSuppose that $\max_i\sum_{\ell=1}^{\infty}\ell^2 \|\mathbf{b}_{\ell i}^{\sharp}\|_2^2<\infty$ and $\max_i\|\mathbf{b}_{i}^{\sharp}\|_2<\infty$, where $\mathbf{b}_{\ell i}^{\sharp}$ is defined in (ref), and $\mathbf{B}=(\mathbf{b}_{1}^{\sharp},\ldots, \mathbf{b}_{N}^{\sharp})^\top$.
Assumption (ref) regulates the rows of $\mathbf{B}_\ell$'s and $\mathbf{B}$, and is rather minor in view of Examples (ref) and (ref). Under this condition, we are able to present the following theorem.
theoremUnder Assumptions (ref) and (ref), as $(N,T)\to (\infty, \infty)$,
\begin{eqnarray*}
\max_{i}\sup_{u\in\mathbb{R}}\left|F_i(u)-\Phi(u)-\frac{\beta_{i3}}{6} (1-u^2)\phi(u) \right|=O\left( \frac{1}{T}\right),
\end{eqnarray*}
where $\max_{i}|\beta_{i3}|=O( \frac{1}{T^{1/2}})$. If $\kappa_3= 0$, then $\beta_{i3}= 0$.
Again, if the skewness of $\varepsilon_{it}$ is 0, the term $\frac{\beta_{i3}}{6} (1-u^2)\phi(u) $ vanishes. Then $F_i(u)$ converges to $\Phi(u)$ at a faster rate. In general, we do not have $\kappa_3= 0$, so we keep the statement of Theorem (ref) as it is.
To infer the distribution of $\widetilde{p}_i$ in practice, we provide a heterogeneous version of the dependent wild bootstrap procedure.
\hrule
enumerate[leftmargin=24pt, parsep=2pt, topsep=2pt]
• Draw the bootstrap version of $\widetilde{p}_{i}$ by
\begin{eqnarray*}
\widetilde{p}^\ast_{i}\coloneqq \frac{1}{\sigma_{p,i}\sqrt{T}}\sum_{t=1}^T(\mathbf{e}_i^\top \mathbf{x}_t-\mu_i)\zeta_{t},
\end{eqnarray*}
where $\zeta_{t}$ is an $m$-dependent time series satisfying Assumption (ref).
• Repeat the above procedure $R$ times to obtain the sampling distribution of $\widetilde{p}^\ast_{i}$.
\hrule
For the heterogeneous case, we present the following lemma to facilitate the development of the asymptotic distribution of $\widetilde{p}^\ast_{i}$ conditional on the sample.
lemmaUnder Assumptions (ref)-(ref), as $(N,T)\to (\infty,\infty)$, for each $i$,
\begin{enumerate}[leftmargin=24pt, parsep=2pt, topsep=2pt]
• $ E^*[(\widetilde{p}_{i}^*)^2]= 1+o_P(1)$;
• $ E^\ast[(\sum_{t=(s-1)m+1}^{sm} \mathbf{e}_i^\top \mathbf{x}_t\zeta_t)^2]=O_P(mL_N)$, for $s=1,\ldots,\lfloor \frac{T}{m} \rfloor$;
• $ E^\ast[(\sum_{t=\lfloor\frac{T}{m}\rfloor m+1}^{T} \mathbf{e}_i^\top \mathbf{x}_t\zeta_t)^2]=O_P(mL_N)$.
\end{enumerate}
Using Lemma (ref), we are able to establish the asymptotic distribution of $\widetilde{p}^\ast_{i}$ conditional on the sample, which is an extension of Theorem (ref) for the heterogeneous bootstrap statistics.
theoremUnder Assumptions (ref)-(ref), as $(N,T)\to (\infty,\infty)$, for each $i$,
\begin{enumerate}[leftmargin=24pt, parsep=2pt, topsep=2pt]
• $\sup_{u\in \mathbb{R}}\left|\text{\normalfont Pr}^*(\widetilde{p}^\ast_{i} \le u ) -\Phi(u)-\frac{1}{6\widetilde{\sigma}_i^{\ast3}}(1-u^2)E^\ast [\widetilde{p}_{i}^{*3}]\phi(x)\right|=O_P\left(\frac{m}{T}\right)$;
• $\sup_{u\in \mathbb{R}}\left|\text{\normalfont Pr}^*(\widetilde{p}^\ast_{i} \le u) -F_i(u)\right|=O_P\left(\sqrt{\frac{m}{T}}\right)$;
• $\text{\normalfont MSE}(\widetilde{\sigma}_i^{\ast 2})=\frac{2m}{T}\int_{-1}^1 a^2(u)du+\frac{C_{q_a}^2}{ \sigma_{p,i}^4m^{2q_a}}\Delta_{q_a,i}^2+o_P(m^{-2q_a})+o\left(\frac{m}{T}\right)$;
\end{enumerate}
where $\widetilde{\sigma}_i^{\ast 2}= E^\ast[\widetilde{p}_{i}^{*2}]$ and $\Delta_{q_a,i}=\sum_{s=-\infty}^{\infty}\sum_{\ell=0}^{\infty} |s|^{q_a}\mathbf{e}_i^\top \mathbf{B}_\ell \mathbf{B}_{\ell+|s|}^\top \mathbf{e}_i$.
Theorems (ref) and (ref) jointly imply that we are able to infer every single $\mu_i$. The discussion under Theorem (ref) still applies here.
We then explore the second option under (ref) when inferring heterogeneity. Mathematically, it means that one is concerned with $\max_i \widetilde{p}_i$ rather than any individual $\widetilde{p}_i$. This is useful, as there is an increasing literature concerning about the uniform inference in the panel data setting (e.g., LLS2024, SJW_2024). The following result contributes to this line of research.
theoremLet $\{\mathbf{z}_t\mid t\in [T]\}$ be independent Gaussian random vectors in $\mathbb{R}^{N}$ such that $E[\mathbf{z}_t]=\mathbf{0}$ and $\mathrm{Var}(\mathbf{z}_t) = \mathbf{B}\mathbf{B}^\top$. Let Assumption (ref).1 hold and $\max_{i}\sum_{\ell=1}^{\infty}\ell \|\mathbf{b}_{\ell i}^{\sharp}\|_2<\infty$. As $(N,T)\to (\infty,\infty)$,
\begin{eqnarray*}
&&\sup_{u\in \mathbb{R}}\left|\Pr\left(\left|\frac{1}{\sqrt{T}}\sum_{t=1}^{T}(\mathbf{x} _t-\pmb{\mu})\right|_{\infty} \leq u\right) - \Pr\left(\left|\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\mathbf{z} _t\right|_{\infty} \leq u\right)\right| \notag \\
&=&O\left( \left(\frac{ N^{2/J}(\log N)^5}{T}\right)^{1/4}+\sqrt{\frac{N^{2/J}(\log N)^{3-2/J}}{T^{1-2/J}}}\right).
\end{eqnarray*}
Theorem (ref) establishes a Gaussian approximation within the panel data framework. Notably, this approximation remains valid irrespective of whether the cross-sectional dependence is WCD or SCD. Also, it is worth mentioning that under the sub-Gaussian condition of $\varepsilon_{it}$, the term $N^{2/J}$ in the above theorem can be replaced by $\log N$. Thus if we impose an exponential tail assumption, $N$ can even diverge at an exponential rate of $T$.
To have a practically feasible version of Theorem (ref), we need to know $\mathrm{Var}(\mathbf{z}_t) = \mathbf{B}\mathbf{B}^\top \eqqcolon \pmb{\Omega}$. Thus, define the high-dimensional long-run covariance matrix estimator by
eqnarray*[eqnarray* omitted — 278 chars of source]
and accordingly, define the Gaussian multiplier bootstrap approximate by
eqnarray*[eqnarray* omitted — 118 chars of source]
where $\{\mathbf{z}_t^* \mid t\in [T]\}$ is a vector of i.i.d. $N$-dimensional Gaussian random variables with $\mathbf{z}_t^*\sim N(\mathbf{0},\mathbf{I}_N)$.
theoremLet Assumption (ref).1 hold with $J>4$, and let Assumption (ref).1 hold. Suppose that (1) $\max_{i}\sum_{\ell=1}^{\infty}\ell^{q_\alpha}\|\mathbf{b}_{\ell i}^{\sharp}\|_2<\infty$, (2) $\frac{N^2T\log T}{(T\widetilde{m}\log N)^{J/4}}\to 0$, and (3) $\widetilde{m} \asymp T^{1/(2q_{\alpha}+1)}$. Then
{
\begin{eqnarray*}
&&\sup_{u\in \mathbb{R}}\left|\mathrm{Pr}^*\left(\left|\widehat{\bm{\Omega}}^{1/2}\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\mathbf{z}_t^* \right|_{\infty} \leq u\right) - \Pr\left(\left|\frac{1}{\sqrt{T}}\sum_{t=1}^{T} (\mathbf{x} _t-\pmb{\mu}) \right|_{\infty} \leq u\right)\right| \notag \\
&=&O_P\left((\log N)^{5/4} (\widetilde{m}/T)^{1/4}\right) + O\left( \left(\frac{ N^{2/J}(\log N)^5}{T}\right)^{1/4}+\sqrt{\frac{N^{2/J}(\log N)^{3-2/J}}{T^{1-2/J}}}\right).
\end{eqnarray*}}
Theorem (ref) builds upon the robust inference of high-dimensional covariance matrix estimation presented in gao2024robust and the continuity of the maximum of a Gaussian distribution. Similar to Theorem (ref), Theorem (ref) accommodates various types of CD.
The condition $\max_{i}\sum_{\ell=1}^{\infty}\ell^{q_\alpha}\|\mathbf{b}_{\ell i}^{\sharp}\|_2<\infty$ is slightly more restrictive than those required in Assumption (ref). This is not surprising, as we now need to investigate the uniform inference rather than derive CLT for any individual. Nevertheless, we only require an algebraic decay rate of the temporal dependence. $\widetilde{m}\asymp T^{1/(2q_{\alpha}+1)}$ corresponds to the optimal level of bandwidth in terms of minimizing the asymptotic mean squared error of each element in $\widehat{\bm{\Omega}}$. The condition $\frac{N^2T\log T}{(T\widetilde{m}\log N)^{J/4}}\to 0$ imposes restrictions on $(\widetilde{m}, N, T)$ jointly, which are easy to realize. If $\varepsilon_{it}$ is sub-Gaussian, $J$ can be arbitrarily large. Then the restrictions can be much simplified.
In the following subsection, we show that the above results offer the theoretical framework for the purpose of inference.
Test Statistic
According to the results established in Section (ref) and (ref), we are now ready to consider an $L_\infty$-based test statistic:
$$
Q_{NT} = \max_i \sqrt{T}(\overline{x}_i - \overline{x}),
$$
where $\overline{x}_i = \frac{1}{T}\sum_{t=1}^{T}x_{it}$ and $\overline{x} = \frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T}x_{it}$. Under $\mathbb{H}_0$, assume $N \leq L_N < N^2$ which implies that $\overline{x}-\mu = o_P(1/\sqrt{T})$, then simple algebra yields that
$$
Q_{NT} = \left|\frac{1}{\sqrt{T}}\sum_{t=1}^{T}(\mathbf{x}_t-\mu\cdot \mathbf{1}_N) \right|_{\infty} + o_P(1).
$$
Hence, we can calculate the critical value of $Q_{NT}$ using Theorem (ref). In addition, our test statistic can detect a class of sparse local alternatives at the rate $T^{-1/2}$. Specifically, under the sparse local alternatives such that $\mathbb{H}_1:\ \mu_i = \mu + a_{T}\delta_i$ with $a_T = 1/\sqrt{T}$ and $|\delta_i| > 0$ for some $i$ (our $L_\infty$-based test is still powerful even only one individual violates the null hypothesis), we have
$$
Q_{NT} = \max_i\left(\frac{1}{\sqrt{T}}\sum_{t=1}^{T}(x_{it} - \mu) + \delta_i\right) + o_P(1),
$$
and thus $Q_{NT}$ diverges to infinity if $ \sqrt{T}a_T \to \infty$. We now state the following proposition.
propositionLet Assumptions (ref) and (ref).1 hold with $J>4$. Additionally, suppose that (1) $\max_{i}\sum_{\ell=1}^{\infty}\ell^{q_\alpha}\|\mathbf{b}_{\ell i}^{\sharp}\|_2<\infty$, (2) $\frac{N^2T\log T}{(T\widetilde{m}\log N)^{J/4}}\to 0$, (3) $\widetilde{m} \asymp T^{1/(2q_{\alpha}+1)}$ and (4) $\overline{x} - \mu = o_P(1/\sqrt{T})$. Then under the null hypothesis
{
\begin{eqnarray*}
\sup_{u\in \mathbb{R}}\left|\Pr(Q_{NT}\le u)- \mathrm{Pr}^*\left(\left|\widehat{\bm{\Omega}}^{1/2}\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\mathbf{z}_t^* \right|_{\infty} \leq u\right)\right|=o_P(1),
\end{eqnarray*}
where $\{\mathbf{z}_t^* \mid t\in [T]\}$} is a sequence of i.i.d. $N$-dimensional Gaussian random vectors with $\mathbf{z}_t^*\sim N(\mathbf{0},\mathbf{I}_N)$.
The distributional approximation established in Proposition 1 enables us to explore the known features of the partial sum of the bootstrapped versions of Gaussian samplers for inferential purposes.
Extensions
In this section, we consider three extensions. We firstly connect the above investigation with the literature on grouping structure analysis such as those surveyed in BM2015 and SSP2016. In the second extension, we relax some restrictions imposed on the panel data unit root processes of PM1999. Finally, we revisit the model of Pesaran2006 and the corresponding CCE estimators using the results established above.
Connection with Grouping Analysis
In this subsection, we explain how the above study can be connected with the current literature on grouping structure analysis. The numerical implementation of grouping can be easily done via $K$-mean or the agglomerative hierarchical clustering algorithm (hastie2009elements). In what follows, we explain how the above results are connected with the grouping analysis.
Note that $\mu_i$ can be estimated by $\overline{x}_i$ as in Section (ref), so $\mu_i$ is already known approximately. To proceed, we introduce the grouping structure, and impose the necessary conditions. Denote $J_0$ sets of indices (i.e., $\mathscr{G}_1,\ldots, \mathscr{G}_{J_0}$) such that
eqnarray*[eqnarray* omitted — 210 chars of source]
where $J_0 \ge 1$ is a fixed positive integer, and $\sharp \mathscr{G}_{j}$ stands for the cardinality of $\mathscr{G}_{j}$.
assumption• Suppose that
\begin{enumerate}[leftmargin=24pt, parsep=2pt, topsep=2pt]
• $\displaystyle \max_{j\in [J]}\max_{i\in \mathscr{G}_{j}} | \mu_{i}-\overline{\mu}_j|\le c_{N} $ with $c_{N}\to 0$, where $\overline{\mu}_j$'s are fixed numbers;
• $\displaystyle \min_{j_1\ne j_2} | \overline{\mu}_{j_1}- \overline{\mu}_{j_2}|\ge c_0 >0$, where $c_0$ is a constant.
\end{enumerate}
Typically, one assumes that $\mu_{i}\equiv \overline{\mu}_j$ for all $i\in \mathscr{G}_{j}$, so $c_{N}\equiv 0$. Assumption (ref).1 relaxes this restriction slightly. Assumption (ref).2 requires all groups to be well partitioned. If $J_0$ is known, we can estimate the group structure as follows:
eqnarray[eqnarray omitted — 161 chars of source]
where $S(\mathcal{G}, \pmb{\nu}) \coloneqq \frac{1}{N}\sum_{j=1}^{J_0}\sum_{i\in \mathcal{G}_j}| \overline{x}_i- \nu_j|^2$, $\mathcal{G} \coloneqq (\mathcal{G}_1,\ldots, \mathcal{G}_{J_0})$, $\pmb{\nu} \coloneqq(\nu_1,\ldots, \nu_{J_0})$, $\widehat{\mathcal{G}}\coloneqq (\widehat{\mathcal{G}}_1,\ldots, \widehat{\mathcal{G}}_{J_0})$, and $\widehat{\pmb{\nu}}\coloneqq (\widehat{\nu}_1,\ldots, \widehat{\nu}_{J_0})$. Here $\mathcal{G}_1,\ldots, \mathcal{G}_{J_0}$ are $J_0$ sets of indices such that
eqnarray*[eqnarray* omitted — 148 chars of source]
In (ref), $\widehat{\mathcal{G}}$ and $\widehat{\pmb{\nu}}$ estimate $\mathscr{G}\coloneqq (\mathscr{G}_1,\ldots, \mathscr{G}_N)$ and $\overline{\pmb{\mu}} \coloneqq (\overline{\mu}_1,\ldots, \overline{\mu}_{J_0})$, respectively.
Since $J$ is not known in practice, we estimate $J_0$ using the following information criterion:
eqnarray[eqnarray omitted — 97 chars of source]
where $\text{IC}(J)\coloneqq S(\widehat{\mathcal{G}}_{\mid J}, \widehat{\pmb{\nu}}_{\mid J}) + \rho_{NT}J$, $J^*$ is a user-specified large fixed constant, and $\rho_{NT}$ is a tuning parameter satisfying that
eqnarray*[eqnarray* omitted — 103 chars of source]
In (ref), $\widehat{\pmb{\nu}}_{\mid J}\coloneqq (\widehat{\nu}_{1\mid J},\ldots,\widehat{\nu}_{J\mid J})$ and $\widehat{\mathcal{G}}_{\mid J}\coloneqq (\widehat{\mathcal{G}}_{1\mid J},\ldots, \widehat{\mathcal{G}}_{J\mid J})$ are obtained via (ref) by assuming the number of groups being $J$. A natural choice of $\rho_{NT}$ is $[\log(N+T)]^{-1}$.
propositionLet Assumptions (ref)-(ref) hold. Then the following results hold:
\begin{enumerate}[leftmargin=24pt, parsep=2pt, topsep=2pt]
• $\displaystyle\max_{i\in [N]}|\overline{x}_i -\sum_{j\in [J_0]}\overline{\mu}_j I(i\in \mathscr{G}_j)|=O_P(\sqrt{\log(N)/T} +c_{N})$.
• Given $J_0$, $\Pr(\widehat{\mathcal{G}}_j=\mathscr{G}_j)\to 1$ for all $j\in [J_0]$, where $\widehat{\mathcal{G}}_j$ is obtained via (ref).
• When $J_0$ is unknown, $\Pr(\widehat{J}=J_0)\to 1$, where $\widehat{J}$ is obtained via (ref).
\end{enumerate}
The first result of this proposition is evident building on the investigation in Section (ref) and Assumption (ref).1. After successfully partitioning the individuals, one can further investigate the homogeneity/heterogeneity for the $j^{th}$ group, e.g.,
eqnarray*[eqnarray* omitted — 100 chars of source]
using the results developed in Sections (ref)-(ref).
Nonstationary Panel Data
We now revisit the nonstationary panel data model studied in PM1999. Consider the following data generating process:
eqnarray[eqnarray omitted — 81 chars of source]
where $\mathbf{y}_t=(y_{1t},\ldots, y_{Nt})^\top$, $\mathbf{y}_0=\mathbf{0}$ without loss of generality, and $\mathbf{B}(L)\pmb{\varepsilon}_t$ is the same as that in (ref). As presented in Section 2 of PM1999, when studying nonstationary panel data, a key quantity is
eqnarray[eqnarray omitted — 127 chars of source]
which is also the foundation of some basic results of BAI200982 and DGP2021. Using the results of Section (ref), we are now able to account for the cross-sectional dependence of these unit root processes, and present the following proposition.
propositionSuppose that Assumption (ref) holds with $L_N=N$, and let further that $\lim_{N}\frac{1}{N}\|\mathbf{B} \|^2\to b$. Then as $(N,T)\to (\infty,\infty)$, $\frac{1}{NT^2}\sum_{i=1}^N\sum_{t=1}^T y_{it}^2\to_P\frac{b}{2}$.
Previously, one has to impose cross-sectional independence such as PM1999, DGP2021 and HUANG2021198 due to technical constraints.
CCE Estimators
We now revisit the model of Pesaran2006 and the corresponding CCE estimators using the results established above. Accordingly, we let $L_N\equiv N$ in what follows. The model is as follows:
eqnarray*[eqnarray* omitted — 193 chars of source]
where only $\{(y_{it}, \mathbf{w}_{it})\mid i\in [N], t\in [T] \}$ are observable, $\mathbf{w}_{it}$ is a $k\times 1$ vector, $\mathbf{f}_t$ is an $m\times 1$ vector, and the dimensions of the other variables are defined accordingly. Both $m$ and $k$ are finite. The model admits a vector form:
eqnarray*[eqnarray* omitted — 101 chars of source]
where we have $\mathbf{Y}_i =(y_{i1},\ldots, y_{iT})^\top$, $\mathbf{W}_i =(\mathbf{w}_{i1},\ldots, \mathbf{w}_{iT})^\top$, $\mathbf{F} =(\mathbf{f}_1,\ldots, \mathbf{f}_T)^\top$, and $\pmb{\epsilon}_i =(\epsilon_{i1},\ldots, \epsilon_{iT})^\top$. In the homogenous setting (e.g., Westerlund2018 and references therein), one further assumes that
eqnarray[eqnarray omitted — 92 chars of source]
It is then natural to question whether (ref) holds practically.
To exam (ref), we briefly review the CCE approach. When eliminating the unobservable factor structure, the CCE approach utilizes the following form:
eqnarray[eqnarray omitted — 321 chars of source]
and for notational simplicity, we rewrite (ref) as
eqnarray*[eqnarray* omitted — 81 chars of source]
where the definitions of $\mathbf{z}_{it}$, $\mathbf{C}_i$, and $\mathbf{u}_{it}$ are self-evident. Simple algebra yields that
eqnarray*[eqnarray* omitted — 98 chars of source]
where $\overline{\mathbf{C}}\coloneqq \frac{1}{N}\sum_{i=1}^N \mathbf{C}_i $, $\overline{\mathbf{Z}}\coloneqq \frac{1}{N}\sum_{i=1}^N \mathbf{Z}_i $ with $\mathbf{Z}_i =(\mathbf{z}_{i1},\ldots, \mathbf{z}_{iT})^\top$, and $\overline{\mathbf{U}}\coloneqq \frac{1}{N}\sum_{i=1}^N \mathbf{U}_i $ with $\mathbf{U}_i =(\mathbf{u}_{i1},\ldots, \mathbf{u}_{iT})^\top$. Consequently, the CCE estimators of $\pmb{\theta}_i$ and $\pmb{\theta}$ are respectively defined by
eqnarray[eqnarray omitted — 453 chars of source]
where $\mathbf{M}_{\overline{\mathbf{Z}}} =\mathbf{I}_T-\overline{\mathbf{Z}}(\overline{\mathbf{Z}}^\top \overline{\mathbf{Z}})^+ \overline{\mathbf{Z}}^\top$.
The key part of the CCE approach is that $\mathbf{M}_{\overline{\mathbf{Z}}} $ offers a good approximation of $\mathbf{M}_{\mathbf{F}}$, because
\[
\overline{\mathbf{C}}^+ =\overline{\mathbf{C}}^\top(\overline{\mathbf{C}}\, \overline{\mathbf{C}}^\top)^{-1}
\]
under the conditions: (1) $\overline{\mathbf{C}}$ has full row rank, and (2) $\overline{\mathbf{U}}$ is asymptotically negligible.
Based on (ref), we construct the following quantities for all $j\in [k]$
eqnarray*[eqnarray* omitted — 187 chars of source]
where $\mathbf{e}_j$ is a selection vector as defined in Section (ref). We further let
eqnarray*[eqnarray* omitted — 166 chars of source]
where $(\widehat{\pmb{\xi}}_{j1},\ldots,\widehat{\pmb{\xi}}_{jT}) = (\widehat{\pmb{\mathbf{U}}}_{1,1}\circ \widehat{\pmb{\mathbf{U}}}_{1,1+j},\ldots, \widehat{\pmb{\mathbf{U}}}_{N,1}\circ \widehat{\pmb{\mathbf{U}}}_{N,1+j})^\top$, $\widehat{\pmb{\mathbf{U}}}_i= \mathbf{M}_{\overline{\mathbf{Z}}} \mathbf{Z}_i,$ and $\widehat{\pmb{\mathbf{U}}}_{i,j}$ stands for the $j^{th}$ column of $\widehat{\pmb{\mathbf{U}}}_{i}$.
To facilitate the development, we impose the following conditions.
assumption• \begin{enumerate}[leftmargin=24pt, parsep=2pt, topsep=2pt]
• (1) Suppose that $\mathbf{f}_t$ is covariance stationary with absolute summable auto covariances, and $\frac{1}{T}\mathbf{F}^\top \mathbf{F}\to_P \pmb{\Sigma}_{\mathbf{f}}>0$. (2) $\overline{\mathbf{C}}\to_P \mathbf{C}$, where $\mathbf{C}$ has full row rank $m$. (3) $\frac{1}{T}\overline{\mathbf{U}}^\top \overline{\mathbf{U}}=O_P(\frac{1}{N})$, $\frac{1}{T} \mathbf{F}^\top \overline{\mathbf{U}}=O_P(\frac{1}{\sqrt{NT}})$ and $\frac{1}{T}\mathbf{V}_i^\top \mathbf{M}_{\mathbf{F}} \pmb{\epsilon}_i=\frac{1}{T}\mathbf{V}_i^\top \pmb{\epsilon}_i+O_P(\frac{1}{T})$ for $\forall i$. (4) $\frac{1}{T}\mathbf{V}_i^\top \mathbf{V}_i\to_P\pmb{\Sigma}_{\mathbf{V}_i}>0$ for $\forall i$, and $\frac{1}{NT}\sum_{i=1}^N\mathbf{V}_i^\top \mathbf{V}_i\to_P\pmb{\Sigma}_{\mathbf{V}}>0$.
• Suppose that $\pmb{\epsilon}_t =(\epsilon_{1t},\ldots, \epsilon_{Nt})^\top \coloneqq\sum_{\ell=0}^{\infty} \mathbf{B}_\ell \pmb{\varepsilon}_{t-\ell}$ fulfills Assumption (ref), and the conditions of Theorem (ref). Let $\mathbf{v}_t\coloneqq f(\pmb{\epsilon}_{t-1},\ldots, \pmb{\epsilon}_{t-p})$, where $p$ is fixed, and $f(\pmb{\epsilon}_{t-1},\ldots, \pmb{\epsilon}_{t-p})$ admits a linear combination of $\pmb{\epsilon}_{t-1},\ldots, \pmb{\epsilon}_{t-p}$.
\end{enumerate}
The first two conditions of Assumption (ref).1 are standard. In the third condition, $\frac{1}{T}\overline{\mathbf{U}}^\top \overline{\mathbf{U}}=O_P(\frac{1}{N})$ and $\frac{1}{T} \mathbf{F}^\top \overline{\mathbf{U}}=O_P(\frac{1}{\sqrt{NT}})$ are rather standard in view of the development of Section (ref), and the following expansions:
eqnarray*[eqnarray* omitted — 310 chars of source]
where $\mathbf{U}_t=(\mathbf{u}_{1t},\ldots, \mathbf{u}_{Nt})^\top$. The requirement $\frac{1}{T}\mathbf{V}_i^\top \mathbf{M}_{\mathbf{F}} \pmb{\epsilon}_i=\frac{1}{T}\mathbf{V}_i^\top \pmb{\epsilon}_i+O_P(\frac{1}{T})$ can be verified by the development of Section (ref). Assumption (ref).2 allows $\pmb{\epsilon}_t$ and $\mathbf{v}_t$ to be weakly dependent.
Under these condition, the following proposition holds.
propositionSuppose that Assumption (ref) holds and $N\asymp T$. Then for all $j\in [k]$
\begin{eqnarray*}
\sup_{u\in \mathbb{R}}\left|\Pr(Q_j\le u)- \mathrm{Pr}^*\left(\left|\widehat{\bm{\Omega}}_j^{1/2}\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\mathbf{z}_t^* \right|_{\infty} \leq u\right)\right|=o_P(1),
\end{eqnarray*}
where $\{\mathbf{z}_t^* \mid t\in [T]\}$ is a sequence of i.i.d. $N$-dimensional Gaussian random vectors with $\mathbf{z}_t^*\sim N(\mathbf{0},\mathbf{I}_N)$.
According to Proposition (ref), all $Q_j$'s should not fall in the rejection region yielded by the bootstrap draws under (ref). We further examine this result in the simulation studies.
On top of these extensions, we may also revisit the time trend analyses of gh2006, cgl2012, ROBINSON20124 and WSX2023, which are of course mathematically involved due to the nonparametric nature. We leave them for future study.
Simulation
In this section, we conduct simulation studies to examine the theoretical results about testing and inference of Section (ref), and also verify our argument about the CCE estimators in Section (ref).
Simulation 1 (Testing) --- The data generating process (DGP) is simplified as follows:
eqnarray*[eqnarray* omitted — 106 chars of source]
where $t=-200, \ldots,0, 1,\ldots, T$, $\rho_x=0.3$, $\pmb{\mu}\coloneqq (\mu_1,\ldots, \mu_N)^\top$, $\mathbf{x}_t\coloneqq (x_{1t},\ldots, x_{Nt})^\top$, and $\pmb{\nu}_t\coloneqq (\nu_{1t},\ldots, \nu_{Nt})^\top$. As well understood, the AR(1) process admits an MA($\infty$) representation thus suiting the definition of (ref). To introduce cross-sectional dependence, we let $\pmb{\Sigma}_\nu=\{\rho_\nu^{|i-j|}\}_{N\times N}$. The observations from $t=-200,\ldots, 0$ are burn-in sample in order to eliminate the impact of the initial value. We let $\{\nu_{it} \}$ be i.i.d. over both $i$ and $t$, and be generated in three cases:
enumerate[leftmargin=48pt, parsep=2pt, topsep=2pt]
• $\nu_{it}\sim N(0,1)$;
• $\nu_{it}\sim t_8$, where $t_8$ stands for a $t$-distribution with a degree freedom 8;
• $\nu_{it}\sim \Gamma(2,2)-1$, where $\Gamma(2,2)$ stands for a Gamma distribution with a shape parameter 2 and a scale parameter 0.5. Therefore, $\Gamma(2,2)-1$ has mean 0.
Case 1 is a symmetric distribution with thin tails; Case 2 is a symmetric distribution with heavy tails; and Case 3 is an asymmetric distribution. To exam the size and power of the proposed test in Section (ref), for each case we consider three scenarios for $\pmb{\mu}$:
enumerate[parsep=2pt, topsep=2pt]
• $\mu_i\equiv 0$ for all $i$;
• $\mu_1=\frac{4}{\sqrt{T}}$, and $\mu_i\equiv 0$ for $i\ge 2$;
• $\mu_1=1$, and $\mu_i\equiv 0$ for $i\ge 2$.
Scenarios (a)-(c) are designed to examine size, local power, and global power respectively.
Notably, the above DGP is a special case of PY2008 and YU2024105458. Thus, for the purpose of comparison, we also consider the approaches of these two papers (referred to as PY and YYX respectively). To put everything on equal footing, we consider the null of YU2024105458 (i.e., $\mathbb{H}_0:\mu_i\equiv 0$ for all $i$), and modify the test statistic of PY2008 accordingly in an obvious manner. For the sake of space, we refer interested readers to their papers for detailed implementation. As PY and YYX methods calculate the asymptotic variances neglecting the dependence of the residuals (e.g., PY2008 and YU2024105458), we anticipate some distorted size or power. Additionally, YU2024105458 rely on the Gaussian assumption, so we anticipate the DGPs of Cases 2 and 3 will further distort size or power. For our method (referred to as GLPY), we calculate $Q_{NT}$ under the null for each dataset, and obtain the 95% confidence interval (denoted by $\text{CI}_\infty$) via $ |\widehat{\bm{\Omega}}^{1/2}\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\mathbf{z}_t^* |_{\infty}$ based on 399 bootstrap replications. After $R$ replications, we calculate the rejection rate as follows:
eqnarray*[eqnarray* omitted — 101 chars of source]
where the subindex $j$ stands for the corresponding values obtained in the $j^{th}$ simulation replication. Similarly, we will report the rejection rates for PY and YYX approaches. For our method, we expect that $\Delta_{test}$ is sufficiently close to 0.05 for the scenario (a), is reasonably close to 1 for the scenario (c), and is in between 0 and 1 for the scenario (b). For simplicity, we let $a(\cdot)$ be Bartlett kernel, and take suggestions from gao2024robust to set $\widetilde{m}=\lfloor 1.75 T^{1/3}\rfloor$ for simplicity. Additionally, we let $R=1000$, $N \in \{100,150,200 \}$, $T\in \{ 200, 400 \}$, and $\rho_\nu \in \{0.5, 0.95 \}$.
The results are summarized in Table (ref). Overall, our approach has reasonable size, local power, and global power irrespective to the magnitude of CD (i.e., the value of $\rho_\nu$) as expected. Due to omitting dependence, PY and YYX methods tend to over reject, which is evident in view of the rejection rates of scenario (a) of Cases 1-3.
table[table omitted — 4,028 chars of source]
Simulation 2 (Inference) --- In this simulation, we examine the bootstrap inferences documented in Sections (ref) and (ref). For simplicity, we consider the three cases identical to Simulation 1 with $\pmb{\mu}=\mathbf{0}_{N\times 1}$, and we set $m=\lfloor 1.75 T^{1/3}\rfloor$. For each dataset, we infer a few quantities. For the homogenous case, we calculate $\widetilde{S}_{NT}$ of Section (ref), and simulate its distribution via $\widetilde{S}_{NT}^*$ based on 399 bootstrap replications, where $\widetilde{S}_{NT}^*$ is self-normalized by construction and does not require any prior knowledge about $L_N$. Using the bootstrap draws, we construct the 95% confidence interval of $\widetilde{S}_{NT}$, denoted by $\text{CI}_{S}$. Second, for each $i$, we calculate $\widetilde{p}_i$ of Section (ref), and construct its distribution via $\widetilde{p}_i^*$ based on 399 bootstrap replications. Accordingly, we construct the 95% confidence interval of each $\widetilde{p}_i$, denoted by $\text{CI}_{p_i}$.
After $R$ simulation replications, we calculate the following measures:
eqnarray*[eqnarray* omitted — 400 chars of source]
where again $j$ indexes the $j^{th}$ simulation replication. We anticipate that $\Delta_{\text{HM}}$ and $\Delta_{\text{HE}} $ are close to 0.05, and $\text{Sd}_{\text{HE}}$ is close to 0 indicating the proposed method in Section (ref) is stable for all $i$'s.
table[table omitted — 1,937 chars of source]
Table (ref) shows that most values are as expected. While the DGPs cover a symmetric distribution with thin tails, a symmetric distribution with heavy tails, and an asymmetric distribution, the above results are reasonably good irrespective of the magnitude of CD.
Simulation 3 (CCE) --- We then consider the following panel data model:
eqnarray*[eqnarray* omitted — 117 chars of source]
where all elements are scalar for simplicity. In this case, $m=1$ and $k=1$, so $m\le k+1$ (a typical requirement for CCE estimators) is fulfilled. We examine the size and power of Proposition (ref).
The DGP is as follows. $\{\epsilon_{it}\}$ follows the identical DGP of $\{x_{it}\}$ as in Cases 1-3 of Simulation 1 with $\pmb{\mu}=\mathbf{0}_{N\times 1}$ and $\rho_{\nu}=0.5$. We let $f_t\sim N(0,1.5)$, $\gamma_i \sim N(0.8, 1)$, $\Gamma_i\sim N(-0.2, 2)$, and $\{v_{it}\coloneqq \epsilon_{i,t-1} \}$. For each case of $\{\epsilon_{it}\}$, we further consider three scenarios for $\{\theta_i\}$:
itemize[parsep=2pt, topsep=2pt]
• $\theta_i\equiv 1$;
• $\theta_1=1+\frac{4}{\sqrt{T}}$, and $\theta_i\equiv 1$ for $i\ge 2$;
• $\theta_1=2$, and $\theta_i\equiv 1$ for $i\ge 2$.
Similar to Simulation 1, scenarios (a)-(c) are designed to evaluate size, local power, and global power. For each dataset, we calculate $Q_1$ and generate the corresponding confidence interval (say, $\text{CI}_{Q_1}$) as in Proposition (ref). After $R$ simulation replications, we report the following measure:
eqnarray*[eqnarray* omitted — 101 chars of source]
where $j$ still indexes the $j^{th}$ simulation replication. $\Delta_{\text{Q}}$ should be close to 0.05 and 1 for the scenarios (a) and (c) respectively, and should be in between 0 and 1 for the scenario (b). As shown in Table (ref), the results are as expected. For scenario (a) of Case 3, our approach is slightly under-size with large $N$, and it might be due to the fact that the error component of Case 3 is skewed.
table[table omitted — 874 chars of source]
A Case Study
Heterogeneous expectations, which may arise in financial markets because investors interpret and react to available information differently, are crucial for understanding variations in asset prices, portfolio allocations, and market dynamics. As a result, the study of heterogeneous expectations has become an important focus in recent years, with a number of works highlighting its implications for financial theory and practice ame2020, Brun2021, Giglio2021.
Notably, Giglio2021 document significant and persistent cross-sectional variations in individual investors' return expectations, which can only be partially explained by demographic factors. Similarly, DI2024 reveal substantial heterogeneity in equity return expectations among institutional investors and investment consultants and it can be linked to asset managers' portfolios. These findings challenge traditional finance models based on homogeneous rational expectations and motivate the development of alternative frameworks.
In this study, we revisit some prior work in this area and formally test the heterogeneity in equity return expectations using the newly proposed inference method. Furthermore, we investigate the relationships between equity premium expectations and equity valuations, as explored by DI2024.
Variables and Data
Following DI2024, we examine three types of subjective return expectations held by asset managers for U.S. equities. The equity return expectation ($ere_{it}$) is defined as the (geometric) nominal equity return forecast for large-cap U.S. equities over a 10-year horizon, as reported in public disclosures\footnote{As noted by DI2024, the forecast horizons for return expectation data range from 1 to 50 years. However, most asset managers provide forecasts close to a 10-year horizon. Consequently, this study focuses on 10-year horizon forecasts.}. The equity premium expectation over yield ($epey_{it}$) is calculated by subtracting the horizon-matched log nominal Treasury yield from the nominal equity return expectation. Similarly, the equity premium expectation over cash ($epec_{it}$) is derived by subtracting the expected annualized return on cash (over the next 10 years) from the equity return expectation.
For the equity valuation, we utilize the cyclically adjusted price-to-earnings ratio ($cape_t$), defined as the log ratio of real equity prices to real earnings averaged over the past decade. Introduced by campbell1988, this measure of equity valuation is widely recognized in the finance literature JL2019. Additionally, we include the past 12-month return of the S&P 500 index ($pr_t$) to capture momentum effects and the horizon-matched Treasury yield ($rf_t$) to account for the risk-free rate.
The data used to construct these variables are available on the website maintained by the authors of DI2024. By focusing on a fixed 10-year forecast horizon, we obtain a panel dataset comprising observations from 45 asset managers across 109 months, covering the period from November 1997 to April 2021. Descriptive statistics for these variables are presented in Table (ref).
table[table omitted — 758 chars of source]
Model Specifications and Inference
To analyze the heterogeneity in equity return expectations, we estimate three econometric models, each progressively incorporating more explanatory variables to account for the potential determinants of the expectations:
itemize[leftmargin=60pt, parsep=2pt, topsep=2pt]
• $y_{it} = \mu_i+\epsilon_{it}$;
• $y_{it} = \alpha_i+\beta_icape_t+\epsilon_{it}$;
• $y_{it} = \alpha_i+\beta_icape_t+\gamma_ipr_t+\delta_irf_t+\epsilon_{it}$,
where $y_{it}\in\{epey_{it}, epec_{it}, ere_{it}\}$. Model 1 is the baseline model that captures the individual-specific mean levels of equity return expectations, where $\mu_i$ represents the time-invariant mean for each asset manager $i$. In Model 2, the equity return expectations are modeled as a function of the $cape_t$, with $\beta_i$ representing the sensitivity of each manager's expectations to equity valuations. DI2024 reveal countercyclical expectations for U.S. equity returns, by showing that the expectations are negatively associated with $cape_t$. In this study, we further investigate whether the relationship between return expectations and valuations differs across managers. The inclusion of $pr_t$ and$rf_t$ in Model 3 further enables us to examine whether these additional factors, such as historical market performance, can help explain cross-sectional differences in expectations.
For each specification, we conduct hypothesis tests to evaluate the presence of heterogeneity in key parameters. Specifically:
enumerate[leftmargin=24pt, parsep=2pt, topsep=2pt]
• We test $\mathbb{H}_0^1$: $\mu_i = \mu$ for all $i$ in Model 1;
• We test $\mathbb{H}_0^2$: $\alpha_i = \alpha$ and $\mathbb{H}_0^3$: $\beta_i = \beta$ for all $i$ respectively in Models 2 and 3.
These tests provide evidence regarding both the variation in average return expectations and the diversity in how managers incorporate fundamental and market-driven information into their forecasts. Rejecting these null hypotheses would indicate the heterogeneity of return expectations among the institutional investors and investment consultants.
Estimation and Testing Results
We first estimate three specifications of equity return expectations using both heterogeneous and homogeneous regression models. The estimated coefficients for each model, along with their 95% confidence intervals, are presented in Table (ref). Across all models, the cyclically adjusted price-to-earnings ratio is found to be significantly negatively associated with asset managers' forecasts of equity returns and premiums, indicating countercyclical patterns in equity return expectations: when equity valuations are higher, asset managers expect lower future returns. This result aligns with the findings of DI2024 and contrasts with the procyclical expectations among retail investors revealed by GS2014.
Comparing the outcomes from heterogeneous and homogeneous estimations, we can identify noticeable differences. In particular, most homogeneous coefficient estimates for the intercept and $cape$ are larger (in absolute values) than the corresponding average values from heterogeneous models. This discrepancy highlights the importance of allowing for heterogeneity in the data. Homogeneous models, by assuming no diversity across asset managers, may overestimate the influence of equity valuation on return expectations.
Furthermore, both estimation approaches consistently suggest that past returns exert no significant influence on asset managers' expectations regarding future equity premiums. This finding suggests that asset managers may not rely on recent return trends when forming long-term return expectations, potentially due to the forward-looking nature of their strategies. It also indicates that their expectations are primarily driven by fundamental factors, such as valuations and macroeconomic conditions, rather than past market performance. Importantly, these results hold consistently across all three measures of equity return expectations.
To further investigate the heterogeneity in asset managers' expectations for U.S. equity returns, we perform a sequence of heterogeneity tests as outlined in Section (ref). The results of these tests are summarized in Panel A of Table (ref). For Model 1, the null hypothesis of homogeneous means is rejected at the 0.01 significance level for all three types of expectations: equity returns, equity premium over yield, and equity premium over cash. This indicates substantial variation in the average levels of return expectations across asset managers.
For Models 2 and 3, the heterogeneity in expectations remains significant even after accounting for equity valuations, past returns, and risk-free rates. The results for $\mathbb{H}_0^3$ further reveal that the relationship between equity return expectations and equity valuations varies significantly across institutional investors and investment consultants. This finding indicates that differences among asset managers extend beyond simple average return expectations; they also exhibit diverse sensitivities to fundamental factors such as equity evaluations. Collectively, these findings confirm the presence of heterogeneity in U.S. equity premium expectations, which aligns with recent literature emphasizing the role of diverse beliefs and preferences in shaping market outcomes Brun2021.
In order to illustrate the heterogeneity test results for the intercept in each model, we provide the distributions of bootstrap test statistics in Figure (ref). For comparison, the test statistics and critical values for the homogeneity test of slope coefficients proposed by PY2008 (PY) and the homogeneity test of intercepts developed by YU2024105458 (YYX) are also computed for $\mathbb{H}_0^3$ and $\mathbb{H}_0^2$, respectively, under Model 2. The results are reported in Panel A of Table (ref).
Notably, the YU2024105458's test originally examines heterogeneity in intercepts under the null hypothesis of known homogeneous intercepts. To adapt this framework to our setting, we modify their statistics to test heterogeneity against an unknown constant by specifying the intercept in the null as its homogeneous estimator. This adjustment ensures comparability with our proposed methodology while maintaining the robustness of their test.
table[table omitted — 2,400 chars of source]
As a robustness check, we follow DI2024 and expand the dataset to include observations with forecast horizons close to ten years, rather than limiting the sample to exactly 10-year-horizon. The heterogeneity test results for this expanded sample are reported in Panel B of Table (ref). Most tests continue to reject the null hypothesis of homogeneity at the 0.01 significance level. The only exception is the one for the mean value of equity premiums over cash, which indicates a slightly weaker rejection (at a 0.05 significance level) of the homogeneity. Nevertheless, the overall robustness of the test results confirms that our findings are not sensitive to the inclusion of additional observations with wider forecast horizons.
table[table omitted — 3,763 chars of source]
figure[figure omitted — 1,286 chars of source]
Conclusion
In this paper, we introduce an underlying data generating process that allows for different magnitude of CD, along with TSA. This is achieved via high-dimensional moving average processes of infinite order (HDMA($\infty$)), which automatically generalizes the spatial structure introduced by Robinson and his co-authors in recent years. The framework is important in the sense that as noted by brockwell1991time and FanYao, the Wold decomposition theorem ensures a formal linear representation exists for any stationary time series with no deterministic components, and HDMA($\infty$) naturally incorporates this result into a panel data framework.
Our setup and investigation significantly integrates and enhances both homogenous and heterogeneous panel data modelling and testing (such as Pesaran2006, PY2008, fan2015power, YU2024105458). To study HDMA($\infty$), we extend the BN decomposition (e.g., BN1981,PS1992) to a high-dimensional time series setting, and derive a complete set of toolkit. It is worth mentioning our investigation complements the work of fan2015power, who specifically study cases where $\frac{T}{\sqrt{N}}\to 0$, by considering a broader range of scenarios and relaxing the independence assumptions employed in PY2008 and YU2024105458.
We exam homogeneity against heterogeneity using Gaussian approximation, a prevalent technique for establishing uniform inference (e.g., chernozhuokov2022improved, and references therein). For post-testing inference, we derive Central Limit theorems through Edgeworth expansions for both homogenous and heterogeneous settings. Notably, the demand for Gaussian approximation in panel data analysis has been increasing recently, as exemplified in Section 4 of SJW_2024 and Section 4 of LLS2024. Our study also contributes to this research direction by providing a set of foundational conditions and deriving a set of useful basic results.
We showcase the practical relevance of the established asymptotic theory by (1). connecting our results with the literature on grouping structure analysis such as those surveyed in BM2015 and SSP2016, (2). examining a nonstationary panel data generating process presented in PM1999, and (3). revisiting the common correlated effects (CCE) estimators of Pesaran2006. Typically, when investigating nonstationary panel data, one has to impose cross-sectional independence such as PM1999, DGP2021 and HUANG2021198 due to technical constraints. Our study offers a set of toolkit to account for the dependence of unit root precesses.
Finally, we verify our theoretical findings via extensive numerical studies using both simulated and real datasets.
{
{
{2.pt plus 0ex}
}
\setcounter{page}{1}
center[center omitted — 397 chars of source]
\setcounter{equation}{0}
\setcounter{lemma}{0}
\setcounter{section}{0}
\setcounter{table}{0}
\setcounter{figure}{0}
\setcounter{remark}{0}
\setcounter{corollary}{0}
Throughout the proofs, we suppose that without loss of generality, $\mu \equiv 0$ for the homogenous case, and $\mu_i\equiv 0$ for all $ i\in [N]$ for the heterogeneous case. We shall not mention them again unless misunderstanding may arise.
This document is structured as follows: Appendix (ref) points out that the proposed inference methods can accommodate unbalanced panels after some necessary modifications; Appendix (ref) provides some foundational facts used throughout the proofs; Appendix (ref) contains all preliminary lemmas; Appendix (ref) presents the theoretical proofs of the main results; and Appendix (ref) details the proofs of the preliminary lemmas.
Unbalanced Panel Data
More often than not, one encounters unbalanced panel data practically. Given the missing proportion is asymptotically negligible, the proposed inference methods can accommodate unbalanced panels after some necessary modifications.
For instance, we let $\mathbb{T}_i$ denote the sample set for individual $i$ and $T_i\coloneqq \sharp \mathbb{T}_i$. The $L_\infty$-based test statistic is then redefined as
$$
Q_{NT} = \max_i\left|\sqrt{T_i}(\overline{x}_i - \overline{x})\right|,
$$
where $\overline{x}_i = \frac{1}{T_i}\sum_{t=1}^{T_i}x_{it}$ and $\overline{x} = \frac{1}{N}\sum_{i=1}^{N}\frac{1}{T_i}\sum_{t=1}^{T_i}x_{it}$.
Additionally, the high-dimensional long-run covariance matrix estimator for unbalanced panels is defined as $ \widehat{\bm{\Omega}}\coloneqq (\widehat{\Omega}_{in})$, where
eqnarray*[eqnarray* omitted — 211 chars of source]
and the Gaussian multiplier bootstrap approximation is accordingly defined by
eqnarray*[eqnarray* omitted — 200 chars of source]
where $\{z_{it}^* \mid t\in\mathbb{T}_i \}$ is a sequence of i.i.d. $ N(0,1)$.
For the CCE estimation, the estimators and test statistics for unbalanced panels can also be updated. Let $\mathbb{N}_t$ contain the indices for individuals that have valid observations at time $t$ and let $N_t$ be the number of indices in this set. Additionally, for each $i$, define $\overline{\mathbf{Z}}_{\cdot i}\coloneqq (\overline{\mathbf{z}}_{t_1},\cdots, \overline{\mathbf{z}}_{t_{T_i}})^\top$, where $t_1,\cdots,t_{T_i}\in \mathbb{T}_i$ and $\overline{\mathbf{z}}_t\coloneqq \frac{1}{N_t}\sum_{i\in\mathbb{N}_t}\mathbf{z}_{it}$.
The CCE estimators of $\pmb{\theta}_i$ and $\pmb{\theta}$ can be then defined as
eqnarray*[eqnarray* omitted — 473 chars of source]
where $\mathbf{M}_{\overline{\mathbf{Z}}_{\cdot i}} =\mathbf{I}_T-\overline{\mathbf{Z}}_{\cdot i}(\overline{\mathbf{Z}}_{\cdot i}^\top \overline{\mathbf{Z}}_{\cdot i})^{-1} \overline{\mathbf{Z}}_{\cdot i}^\top $.
The heterogeneity test statistic is then given by
eqnarray*[eqnarray* omitted — 197 chars of source]
where $\mathbf{e}_j$ is a selection vector. For constructing the bootstrap statistics, we use $\widehat{\pmb{\Omega}}_{j}=(\widehat{\pmb{\Omega}}_{j,in})$, where
eqnarray*[eqnarray* omitted — 200 chars of source]
where $\widehat{\xi}_{j,it}$ denotes the $t$-th element of $\widehat{\pmb{\mathbf{U}}}_{i,1}\circ \widehat{\pmb{\mathbf{U}}}_{i,1+j}$, with
$\widehat{\pmb{\mathbf{U}}}_i= \mathbf{M}_{\overline{\mathbf{Z}}_{\cdot i}} \mathbf{Z}_i,$ and $\widehat{\pmb{\mathbf{U}}}_{i,j}$ standing for the $j^{th}$ column of $\widehat{\pmb{\mathbf{U}}}_{i}$. The distribution of $Q_j$ is then approximated by
eqnarray*[eqnarray* omitted — 202 chars of source]
where $\{z_{it}^* \mid t\in\mathbb{T}_i \}$ is a sequence of i.i.d. $ N(0,1)$.
Some Facts
We present a few facts in this section, which will be repeatedly used in the proofs.
On Cumulant --- Note that for a generic cumulant, we have for a constant $c$
eqnarray[eqnarray omitted — 61 chars of source]
We refer the interested reader to saulis for more details about cumulants.
On BN Decomposition --- Simple algebra shows the Beveridge and Nelson (BN) decomposition under the high-dimensional (HD) setting is as follows:
eqnarray[eqnarray omitted — 91 chars of source]
where $\widetilde{\mathbf{B}}(L)\coloneqq\sum_{\ell=0}^{\infty}\widetilde{\mathbf{B}}_\ell L^\ell$ with $\widetilde{\mathbf{B}}_\ell\coloneqq \sum_{k=\ell+1}^{\infty}\mathbf{B}_k$.
On Dependence --- To calculate CD, for $\forall i,j\in [N]$ it is easy to obtain that
eqnarray*[eqnarray* omitted — 119 chars of source]
To calculate TSA, for $t> s$ we write
eqnarray*[eqnarray* omitted — 789 chars of source]
To calculate CD + TSA, for $t> s$ we write
eqnarray*[eqnarray* omitted — 788 chars of source]
On Cumulants --- Recall the notation of Section (ref). Let $\widetilde{\phi}(x)$ be the characteristic function of a standard normal distribution. We further define its cumulants by $\gamma_r$ for $r\ge 1$. We now consider a generic distribution function $G(x)$ with a characteristic functions $\chi(u)$ and cumulants $\beta_r$. Then by Taylor expansion, we have
\[
\log \frac{\chi(u)}{\widetilde{\phi}(x)} =\log \chi(u) - \log \widetilde{\phi}(x)=\sum_{r=1}^\infty (\beta_r-\gamma_r)\cdot \frac{(\mathsf{i} u)^r}{r!}
\]
and
eqnarray[eqnarray omitted — 155 chars of source]
On Exponential function --- By Taylor expansion,
eqnarray[eqnarray omitted — 73 chars of source]
Taylor theorem yields that
eqnarray[eqnarray omitted — 111 chars of source]
where $\widetilde{u}$ is between 0 and $u$.
Preliminary Lemmas
In this appendix, we provide some useful preliminary lemmas. The first three lemmas are either obvious or have been carefully proved in the literature, so we do not provide the proofs herewith. We will provide the proofs for Lemmas (ref) to (ref).
lemmaA matrix $\mathbf{A}$ in $\mathbb{R}^{n\times n}$ has rank one if and only if it can be written as the outer product of two nonzero vectors in $\mathbb{R}^{n}$ (i.e., $\mathbf{A} =\mathbf{x}\mathbf{y}^\top$).
lemma[Esseen's smoothing Lemma]
Let $H(\cdot)$ be a distribution with 0 expectation and characteristic function $\chi(\cdot)$. Suppose $H(x)-G(x)$ vanishes at $\pm \infty$ and that $G(\cdot)$ has a derivative $g(\cdot)$ such that $|g|_\infty \le m$. Finally, suppose that $g$ has a continuously differentiable Fourier transform $\xi(\cdot)$ such that $\xi(0)=1$ and $\xi^{(1)}(0)=0$. Then
\begin{eqnarray*}
|H -G|_\infty\le \frac{1}{\pi}\int_{-a}^a \left|\frac{\chi(u)-\xi(u)}{u} \right|\mathrm{d}u+\frac{24m}{\pi \cdot a}
\end{eqnarray*}
where $a >0$.
See Chapter XVI, section 3 of feller1971introduction for details about Esseen's smoothing Lemma.
lemmaLet $\mathsf{i}$ be the imaginary unit. Then for $\forall \theta\in [0,1)$
\begin{eqnarray*}
\sup_x\left|\exp(\mathsf{i}x) -\sum_{j=0}^r \frac{(\mathsf{i}x)^j}{j!} \right| \le \min\left\{\frac{2}{r!}|x|^{r+\theta}, \frac{|x|^{r+1}}{(r+1)!} \right\}.
\end{eqnarray*}
See tik1981 for example.
lemma• Under Assumption (ref), the following results hold:
\begin{enumerate}[leftmargin=24pt, parsep=2pt, topsep=2pt]
• For $t> 1$, $\sum_{s=1}^t \mathbf{x}_{s}$ admits two representations:
\begin{enumerate}[leftmargin=24pt, parsep=2pt, topsep=2pt]
• $\sum_{s=1}^t \mathbf{x}_{s} = \mathbf{B} \sum_{s=1}^t\pmb{\varepsilon}_{s}-\widetilde{\mathbf{B}}(L)\pmb{\varepsilon}_{t}+\widetilde{\mathbf{B}}(L)\pmb{\varepsilon}_{0}$,
• $\sum_{s=1}^t \mathbf{x}_{s} = \sum_{\ell=1}^t (\mathbf{B}-\widetilde{\mathbf{B}}_{t-\ell})\pmb{\varepsilon}_{\ell}-\sum_{\ell=-\infty}^{0}(\widetilde{\mathbf{B}}_{t-\ell} -\widetilde{\mathbf{B}}_{-\ell})\pmb{\varepsilon}_{\ell}\eqqcolon \sum_{\ell=-\infty}^t \pmb{\mathcal{B}}_{t\ell} \pmb{\varepsilon}_\ell$,
\end{enumerate}
where $\sum_{\ell=0}^{\infty}\sqrt{\frac{N}{L_N}}\|\widetilde{\mathbf{B}}_\ell\|_2<\infty$, and $\pmb{\mathcal{B}}_{t\ell} = - \widetilde{\mathbf{B}}_{t-\ell} + \widetilde{\mathbf{B}}_{-\ell}$ for $-\infty\leq \ell\leq 0$; $\pmb{\mathcal{B}}_{t\ell} = \mathbf{B}-\widetilde{\mathbf{B}}_{t-\ell}$ for $1\leq \ell \leq t$;
• $\left|\frac{1}{\sqrt{L_N}} \sum_{\ell=-\infty}^0 \mathbf{1}_N^\top\pmb{\mathcal{B}}_{T\ell} \pmb{\varepsilon}_{\ell}\right|=O_P(1)$;
• $\left|\frac{1}{\sqrt{L_N}}\sum_{t=1}^T \mathbf{1}_N^\top\widetilde{\mathbf{B}}_{T-\ell} \pmb{\varepsilon}_{\ell} \right|=O_P(1)$.
\end{enumerate}
Write
eqnarray[eqnarray omitted — 1,027 chars of source]
where the definition of $\mathbf{B}_v^*(L)$ for $v\ge 0$ should be obvious. In connection with the BN decomposition of (ref), we have
eqnarray[eqnarray omitted — 101 chars of source]
where
eqnarray*[eqnarray* omitted — 430 chars of source]
lemma• Under Assumptions (ref) and (ref), the following results hold:
\begin{enumerate}[leftmargin=24pt, parsep=2pt, topsep=2pt]
• $\| E[(\pmb{\varepsilon}_{t}\otimes \pmb{\varepsilon}_{t})(\pmb{\varepsilon}_{t}^\top \otimes \pmb{\varepsilon}_{t}^\top)]\|_2 =O(N)$;
• $\left|\frac{1}{L_NT}\sum_{t=1}^T \mathbf{B}_0^*(L)\text{\normalfont vec}(\pmb{\varepsilon}_{t} \pmb{\varepsilon}_{t}^\top) - \frac{1}{L_N} \sum_{\ell=0}^{\infty} \mathbf{1}_N^\top \mathbf{B}_\ell \mathbf{B}_\ell^\top \mathbf{1}_N \right|=O_P\left(\frac{1}{\sqrt{T}}\right)$;
• $\left|\frac{1}{L_NT}\sum_{t=1}^T\sum_{v=1}^\infty \mathbf{B}_v^*(L)\text{\normalfont vec}(\pmb{\varepsilon}_{t-v} \pmb{\varepsilon}_{t}^\top ) \right|=O_P\left(\frac{1}{\sqrt{T}}\right)$.
\end{enumerate}
Proofs of the Main Results
proof[Proof of Theorem (ref)]
• First, we denote the following notation to facilitate the development. Let
\begin{eqnarray}
\mathbf{B} =(\mathbf{b}_1^\dag,\ldots, \mathbf{b}_N^\dag)\quad and\quad\widetilde{\mathbf{B}}_{\ell}=(\widetilde{\mathbf{b}}_{\ell, 1}^\dag,\ldots, \widetilde{\mathbf{b}}_{\ell, N}^\dag)
\end{eqnarray}
where $\mathbf{B}$ and $\widetilde{\mathbf{B}}_{\ell}$ have been defined in (ref) already.
By Lemma (ref), we write
\[
\sum_{s=1}^t \mathbf{x}_{s} = \sum_{\ell=-\infty}^t \pmb{\mathcal{B}}_{t\ell}\pmb{\varepsilon}_{\ell},
\]
where the definition of $\pmb{\mathcal{B}}_{t\ell}$ is the same as that in Lemma (ref). Thus, $\widetilde{S}_{NT}$ can also be written as
\[
\widetilde{S}_{NT}= \frac{1}{\sigma_x\sqrt{L_N T}} \sum_{\ell=-\infty}^T\mathbf{1}_N^\top\pmb{\mathcal{B}}_{T\ell} \pmb{\varepsilon}_{\ell}\eqqcolon \sum_{\ell=-\infty}^T\mathbf{1}_N^\top\overline{\pmb{\mathcal{B}}}_{\ell} \pmb{\varepsilon}_{\ell},
\]
where $\frac{1}{\sigma_x\sqrt{L_NT}} \pmb{\mathcal{B}}_{T\ell} \eqqcolon \overline{\pmb{\mathcal{B}}}_{\ell}$, and we have suppressed $N$ and $T$ in $\overline{\pmb{\mathcal{B}}}_{\ell}$ for notational simplicity.
To proceed, denote by $\psi_{NT}(u)$ the characteristic function of $\widetilde{S}_{NT}$. Thus,
\begin{eqnarray}
\psi_{NT}(u) &=& E\left[\exp\left(\mathsf{i}u \sum_{\ell=-\infty}^T\mathbf{1}_N^\top\overline{\pmb{\mathcal{B}}}_{\ell} \pmb{\varepsilon}_{\ell}\right)\right]\notag \\
&=&\prod_{\ell=-\infty}^T \prod_{j=1}^NE [\exp (\mathsf{i}u \mathbf{1}_N^\top\overline{\mathbf{b}}_{\ell ,j}^\dag \varepsilon_{j\ell} ) ]\notag\\
&=& \prod_{\ell=-\infty}^T \prod_{j=1}^N\psi(\widetilde{b}_{\ell j} u) ,
\end{eqnarray}
where $\overline{\mathbf{b}}_{\ell ,j}^\dag$ stands for the $j^{th}$ column of $\overline{\pmb{\mathcal{B}}}_{\ell}$, $\widetilde{b}_{\ell j}\coloneqq \mathbf{1}_N^\top \overline{\mathbf{b}}_{\ell ,j}^\dag$, and the second equality follows from $\{\varepsilon_{it}\}$ being i.i.d. over both dimensions.
Using (ref), we are able to calculate the $r^{th}$ cumulant $\beta_r$ of $\widetilde{S}_{NT}$ for $r\ge 1$. Obviously, we have $\beta_1 =0$ and $\beta_2 =1$. For $r\ge 3$, write
\begin{eqnarray}
\beta_r &=& (-\mathsf{i})^r \frac{\mathrm{d}^r}{\mathrm{d}u^r}\log \psi_{NT}(u) |_{u=0}\notag \\
&=&\sum_{\ell=-\infty}^T\sum_{j=1}^N (-\mathsf{i})^r\frac{\mathrm{d}^r}{\mathrm{d}u^r}\log\psi(\widetilde{b}_{\ell j}u)|_{u=0} \notag \\
&=&\sum_{\ell=-\infty}^T\sum_{j=1}^N\widetilde{b}_{\ell j}^r (-\mathsf{i})^r\frac{\mathrm{d}^r}{\mathrm{d}u^r}\log\psi(u)|_{u=0}\notag \\
&=&\kappa_r\left(\sum_{\ell=1}^T\sum_{j=1}^N\widetilde{b}_{\ell j}^r +\sum_{\ell=-\infty}^0\sum_{j=1}^N\widetilde{b}_{\ell j}^r \right),
\end{eqnarray}
where the second equality follows from (ref), and the third equality follows from (ref). For the second term on the right hand side of (ref), we note that
\begin{eqnarray}
\left|\sum_{\ell=-\infty}^0\sum_{j=1}^N\widetilde{b}_{\ell j}^r\right|&\le &\frac{1}{(\sigma_x^2 L_NT)^{r/2}}\sum_{\ell=-\infty}^0\sum_{j=1}^N|\mathbf{1}_N^\top \mathbf{b}_{T\ell, j}^\dag|^r\notag \\
&\le &\frac{1}{(\sigma_x^2 L_N T)^{r/2}}\left(\sum_{\ell=-\infty}^0\sum_{j=1}^N|\mathbf{1}_N^\top \mathbf{b}_{T\ell, j}^\dag|^2\right)^{r/2}\notag\\
&= &\frac{1}{(\sigma_x^2 T)^{r/2}}\left(\frac{1}{L_N}\sum_{\ell=-\infty}^0 \mathbf{1}_N^\top \pmb{\mathcal{B}}_{T\ell}\pmb{\mathcal{B}}_{T\ell}^\top \mathbf{1}_N\right)^{r/2}\notag \\
&=&O\left(\frac{1}{T^{r/2}}\right),
\end{eqnarray}
where $\mathbf{b}_{T\ell, j}^\dag$ stands for the $j^{th}$ column of $\pmb{\mathcal{B}}_{T\ell}$, the second inequality follows from the fact that for a vector $\mathbf{x}$, $|\mathbf{x}|_{p_1}\le |\mathbf{x} |_{p_2}$ for any $p_1> p_2 \ge 1$, and the last equality follows from the proof of Lemma (ref).2. Thus, (ref) and (ref) together infer that
\begin{eqnarray}
\beta_r =\kappa_r \sum_{\ell=1}^T\sum_{j=1}^N\widetilde{b}_{\ell j}^r +O\left(\frac{1}{T^{r/2}}\right).
\end{eqnarray}
Recall the definitions of (ref), and write further that
\begin{eqnarray}
&&\frac{1}{(\sigma_x^2 L_NT)^{r/2}}\sum_{\ell=1}^T\sum_{j=1}^N|\mathbf{1}_N^\top \mathbf{b}_{T\ell, j}^\dag|^r \notag \\
&\le& \frac{2^{r-1}}{(\sigma_x^2 L_NT)^{r/2}}\sum_{\ell=1}^T\sum_{j=1}^N[|\mathbf{1}_N^\top \mathbf{b}_j^\dag|^r+|\mathbf{1}_N^\top \widetilde{\mathbf{b}}_{T-\ell, j}^\dag|^r]\notag\\
&=& \frac{2^{r-1}}{(\sigma_x^2 L_N)^{r/2}T^{r/2-1}} \sum_{j=1}^N |\mathbf{1}_N^\top \mathbf{b}_j^\dag|^r+O\left(\frac{1}{T^{r/2}}\right),
\end{eqnarray}
where the first inequality follows from $(a+b)^r\le 2^{r-1}(a^r+b^r)$ for $a,b\ge 0$ and $r>1$, and the last step follows from a development similar to (ref). Thus, by (ref) and (ref), we can obtain that
\begin{eqnarray}
|\beta_r|& \le &\frac{2^{r-1}}{(\sigma_x^2 L_N)^{r/2}T^{r/2-1}} \sum_{j=1}^N(\mathbf{1}_N^\top \mathbf{b}_j^\dag )^r +O\left(\frac{1}{T^{r/2}}\right)\notag\\
&\le &\frac{2^{r-1}}{(\sigma_x^2 L_N)^{r/2}T^{r/2-1}} \cdot N\|\mathbf{B} \|_1^r +O\left(\frac{1}{T^{r/2}}\right)\notag\\
&=& O\left(\Delta_{NT}(r) \vee \frac{1}{T^{r/2}} \right),
\end{eqnarray}
where $\Delta_{NT}(r)\coloneqq NT\left(\frac{\|\mathbf{B} \|_1}{\sqrt{L_NT}}\right)^r$.
Note that the right hand side of (ref) reduces to $O\left(\frac{1}{ (NT)^{r/2-1}} \vee \frac{1}{T^{r/2}} \right)$ for both Examples (ref) and (ref) by simple calculation. In addition, using (ref), we write for $r\ge 3$
\begin{eqnarray}
|\beta_r| &=& O\left(\Delta_{NT}(r) \vee \frac{1}{T^{r/2}} \right) =O(1) \frac{1}{\left(\frac{\sqrt{L_NT}}{(NT)^{1/r} \|\mathbf{B}\|_1} \wedge T^{1/2}\right)^r}\notag \\
&\le &O (1) \frac{1}{\left(\frac{\sqrt{L_NT}}{(NT)^{1/3} \|\mathbf{B}\|_1} \wedge T^{1/2}\right)^r} = O\left( \frac{1}{\widetilde{\Delta}_{NT}^r}\right),
\end{eqnarray}
where $\widetilde{\Delta}_{NT} \coloneqq \frac{\sqrt{L_N}T^{1/6}}{N^{1/3} \|\mathbf{B}\|_1} \wedge T^{1/2}$.
By applying (ref) to $\psi_{NT}(u)$ and the characteristic function of $N(0,1)$, we have
\begin{eqnarray}
\psi_{NT}(u) &=&\exp\left\{ \sum_{r=3}^\infty \beta_r\cdot\frac{(\mathsf{i} u)^r}{r!}\right\} \cdot\widetilde{\phi}(u)\notag \\
&=&\left\{1 +\sum_{n=1}^\infty\frac{1}{n!}\left(\sum_{r=3}^\infty \beta_r\cdot\frac{(\mathsf{i} u)^r}{r!}\right)^n \right\}\cdot\widetilde{\phi}(u),
\end{eqnarray}
where $\beta_r$ is bounded by (ref), and the second equality follows from (ref). Let $f_{NT}(w)$ be the density function of $\widetilde{S}_{NT}$, and Fourier inversion of (ref) leads to the following expansion:
\begin{eqnarray*}
f_{NT}(w) &=&\phi(w) +\frac{1}{2\pi}\int_{\mathbb{R}} \exp(-\mathsf{i}uw)\cdot \beta_3\cdot \frac{(\mathsf{i}u)^3}{3!}\cdot\widetilde{\phi}(u)\mathrm{d}u \notag \\
&&+\frac{1}{2\pi}\int_{\mathbb{R}} \exp(-\mathsf{i}uw)\cdot\widetilde{\psi}_{NT}(u)\cdot\widetilde{\phi}(u)\mathrm{d}u\notag \\
&=&\phi(w) + \frac{\beta_3}{6} H_3(w)\phi(w)\notag \\
&&+\frac{1}{2\pi}\int_{\mathbb{R}} \exp(-\mathsf{i}uw)\cdot\widetilde{\psi}_{NT}(u)\cdot\widetilde{\phi}(u)\mathrm{d}u,
\end{eqnarray*}
where $\widetilde{\psi}_{NT}(u)\coloneqq \sum_{r=4}^\infty \beta_r\cdot\frac{(\mathsf{i} u)^r}{r!} + \sum_{n=2}^\infty\frac{1}{n!}\left(\sum_{r=3}^\infty \beta_r\cdot\frac{(\mathsf{i} u)^r}{r!}\right)^n$, and the second equality follows from Lemma (ref).
Next, we bound $\frac{1}{2\pi}\int_{\mathbb{R}} \exp(-\mathsf{i}uw)\cdot \widetilde{\psi}_{NT}(u)\cdot\widetilde{\phi}(u)\mathrm{d}u$. First, note
\begin{eqnarray}
\left|\sum_{r=4}^\infty \beta_r\cdot\frac{(\mathsf{i} u)^r}{r!} \right|\widetilde{\phi}(u) &\le & |\beta_4| \widetilde{\phi}(u) \sum_{r=4}^\infty \frac{|u|^r}{r!} \notag \\
&\le & |\beta_4| \widetilde{\phi}(u)\exp(|u|)\notag \\
&=&|\beta_4| \exp(-u^2/2+|u|),
\end{eqnarray}
where the second inequality follows from (ref). Second, in connection with (ref), assuming that $|u|\le \widetilde{\Delta}_{NT}$ and using Taylor theorem twice as in (ref) we write
\begin{eqnarray}
&&\left|\sum_{n=2}^\infty\frac{1}{n!}\left(\sum_{r=3}^\infty \beta_r\cdot\frac{(\mathsf{i} u)^r}{r!}\right)^n \right|\widetilde{\phi}(u)\notag \\
&\le &O(1)\sum_{n=2}^\infty\frac{1}{n!} \left| \sum_{r=3}^\infty \cdot\frac{(\mathsf{i} u/\widetilde{\Delta}_{NT} )^r}{r! }\right|^n \widetilde{\phi}(u)\notag\\
&\le & O(1)\frac{1}{\widetilde{\Delta}_{NT}^6}\widetilde{\phi}(u) \widetilde{u}^6=O(1)\frac{1}{\widetilde{\Delta}_{NT}^6},
\end{eqnarray}
where $|\widetilde{u}|$ is in between 0 and $\widetilde{\Delta}_{NT}$, and the last equality follows from $\widetilde{\phi}(u) \widetilde{u}^6$ being uniformly bounded. Thus, by (ref) and (ref), we can write
\begin{eqnarray}
&&\sup_{|u|\le \widetilde{\Delta}_{NT}}\left|\frac{1}{2\pi}\int_{\mathbb{R}} \exp(-\mathsf{i}uw)\cdot \widetilde{\psi}_{NT}(u)\cdot\widetilde{\phi}(u)\mathrm{d}u\right|\notag \\
&=&O\left( \Delta_{NT}(4) \vee \frac{1}{T^2} + \frac{1}{\widetilde{\Delta}_{NT}^6}\right) = O\left( \Delta_{NT}(4) \vee \frac{1}{T^2} \right),
\end{eqnarray}
where the second equality follows from Assumption (ref).2.
Thus, according to (ref) and (ref), we write further
\begin{eqnarray}
&&\sup_{w\in\mathbb{R}}\left|f_{NT}(w)-\phi(w)-\frac{\beta_3}{6} H_3(w)\phi(w) \right|\notag \\
&\le &\frac{1}{2\pi}\int_{\mathbb{R}}\left|\psi_{NT}(u) -\widetilde{\phi}(u) - \frac{\beta_3}{6} (\mathsf{i}u)^3 \widetilde{\phi}(u)\right|\mathrm{d}u\notag\\
&= &\frac{1}{2\pi}\int_{|u|\ge \widetilde{\Delta}_{NT}} |\widetilde{\psi}_{NT}(u) |\widetilde{\phi}(u)\mathrm{d}u +\frac{1}{2\pi}\int_{|u|< \widetilde{\Delta}_{NT}} |\widetilde{\psi}_{NT}(u) |\widetilde{\phi}(u)\mathrm{d}u \notag \\
&=&O\left( \Delta_{NT}(4) \vee \frac{1}{T^2}\right),
\end{eqnarray}
where, in the last step, $\int_{|u|\ge \widetilde{\Delta}_{NT}} |\widetilde{\psi}_{NT}(u) |\widetilde{\phi}(u)\mathrm{d}u$ can be arbitrarily small by standard operation, and $\int_{|u|< \widetilde{\Delta}_{NT}} |\widetilde{\psi}_{NT}(u) |\widetilde{\phi}(u)\mathrm{d}u $ is bounded by (ref).
To study the CDF, we let $G(w) =\Phi(w)+\frac{\beta_3}{6}(1-w^2)\phi(w)$, where simple algebra shows that
\begin{eqnarray*}
\frac{\mathrm{d}( (1-w^2)\phi(w) )}{\mathrm{d}w} = H_3(w)\phi(w) .
\end{eqnarray*}
Thus, $G(w)$ has a characteristic function $\xi(w) =(1+\frac{\beta_3}{6}(\mathsf{i}w)^3)\widetilde{\phi}(w)$. We then invoke Esseen's smoothing Lemma. First, we let $a$ of Lemma (ref) be sufficiently large, so the second term on the right hand side of Lemma (ref) becomes negligible. Then, similar to (ref), we study the first term and can obtain that
\begin{eqnarray*}
|F_{NT}(w)-G(w)|_\infty=O\left( \Delta_{NT}(4) \vee \frac{1}{T^2}\right).
\end{eqnarray*}
The proof is now completed.
proof[Proof of Corollary (ref)]
• According to (ref) and Lemma (ref).1, we write
\begin{eqnarray*}
\beta_3&=&\kappa_3 \sum_{\ell=1}^T\sum_{j=1}^N\widetilde{b}_{\ell j}^3+O\left(\frac{1}{T^{3/2}}\right)\notag\\
&=&\frac{\kappa_3}{(L_NT\sigma_x)^{3/2}} \sum_{\ell=1}^T\sum_{j=1}^N (\mathbf{1}_N^\top \mathbf{b}_j^\dag )^3-\frac{3\kappa_3}{(L_NT\sigma_x)^{3/2}} \sum_{\ell=1}^T\sum_{j=1}^N(\mathbf{1}_N^\top \mathbf{b}_j^\dag )^2(\mathbf{1}_N^\top \widetilde{\mathbf{b}}_{T-\ell,j}^\dag ) \notag \\
&&+\frac{3\kappa_3}{(L_NT\sigma_x)^{3/2}} \sum_{\ell=1}^T\sum_{j=1}^N (\mathbf{1}_N^\top \mathbf{b}_j^\dag) (\mathbf{1}_N^\top \widetilde{\mathbf{b}}_{T-\ell,j}^\dag )^2-\frac{\kappa_3}{(L_NT\sigma_x)^{3/2}} \sum_{\ell=1}^T\sum_{j=1}^N( \mathbf{1}_N^\top \widetilde{\mathbf{b}}_{T-\ell,j}^\dag )^3 \notag \\
&&+O\left(\frac{1}{T^{3/2}}\right)\notag \\
&\eqqcolon &\frac{\kappa_3}{\sigma_x^{3/2}}( \beta_{3,1}-3\beta_{3,2}+3\beta_{3,3}-\beta_{3,4})+O\left(\frac{1}{T^{3/2}}\right),
\end{eqnarray*}
where $\mathbf{b}_j^\dag$ and $\widetilde{\mathbf{b}}_{T-\ell,j}^\dag$ are defined in (ref), and the definitions of $\beta_{3,j}$ for $j\in[4]$ are obvious.
For $\beta_{3,2}$, we write
\begin{eqnarray*}
|\beta_{3,2}|&\le&\frac{\|\mathbf{B}\|_1^2}{(L_NT)^{3/2}} \sum_{j=1}^N \sum_{\ell=1}^T |\mathbf{1}_N^\top \widetilde{\mathbf{b}}_{T-\ell,j}^\dag |\notag\\
&\le&\frac{\|\mathbf{B}\|_1^2}{(L_NT)^{3/2}} \sqrt{NT}\left( \sum_{j=1}^N\sum_{\ell=1}^T |\mathbf{1}_N^\top \widetilde{\mathbf{b}}_{T-\ell,j}^\dag|^2\right)^{1/2}\notag\\
&=&\frac{\sqrt{NT}\|\mathbf{B}\|_1^2}{(L_NT)^{3/2}} \left(\sum_{\ell=1}^T\mathbf{1}_N^\top \widetilde{\mathbf{B}}_{T-\ell}\widetilde{\mathbf{B}}_{T-\ell}^\top \mathbf{1}_N\right)^{1/2}\notag\\
&=&\frac{\sqrt{N}\|\mathbf{B}\|_1^2}{L_NT } \left(\frac{N}{L_N}\sum_{\ell=1}^T \|\widetilde{\mathbf{B}}_{T-\ell}\|_2^2\right)^{1/2}\notag \\
&=&O_P\left(\frac{\sqrt{N}\|\mathbf{B}\|_1^2}{L_NT}\right),
\end{eqnarray*}
where the last line follows from the proof of Lemma (ref).2.
For $\beta_{3,3}$, we write
\begin{eqnarray*}
|\beta_{3,3}|&\le&\frac{\|\mathbf{B}\|_1}{(L_NT)^{3/2}} \sum_{j=1}^N \sum_{\ell=1}^T |\mathbf{1}_N^\top \widetilde{\mathbf{b}}_{T-\ell,j}^\dag |^2\notag\\
&=&\frac{\|\mathbf{B}\|_1}{(L_NT)^{3/2}} \sum_{\ell=1}^T\mathbf{1}_N^\top \widetilde{\mathbf{B}}_{T-\ell}\widetilde{\mathbf{B}}_{T-\ell}^\top \mathbf{1}_N\notag\\
&=&O\left(\frac{\|\mathbf{B}\|_1}{L_N^{1/2}T^{3/2}}\right) ,
\end{eqnarray*}
where the second equality follows from the proof of Lemma (ref).2.
For $\beta_{3,4}$, we write
\begin{eqnarray*}
|\beta_{3,4}|&\le&\frac{1}{(L_NT)^{3/2}} \left(\sum_{j=1}^N \sum_{\ell=1}^T |\mathbf{1}_N^\top \widetilde{\mathbf{b}}_{T-\ell,j}^\dag |^2\right)^{3/2}\notag\\
&=&\frac{1}{(L_NT)^{3/2}} \left( \sum_{\ell=1}^T\mathbf{1}_N^\top \widetilde{\mathbf{B}}_{T-\ell}\widetilde{\mathbf{B}}_{T-\ell}^\top \mathbf{1}_N\right)^{3/2}\notag \\
&=&O\left(\frac{1}{L_N T^{3/2}}\right),
\end{eqnarray*}
where the first inequality follows a development similar to (ref).
By Assumption (ref).2, it is easy to know that among $\beta_{3,j}$ with $j=2,3,4$, the term $|\beta_{3,2}|$ offers the lowest rate. In connection with the decomposition of $\beta_3$, the proof is now completed.
proof[Proof of Theorem (ref)]
• (1). Let $\mathcal{S}_{NT}^*= \sum_{t=1}^T\mathbf{1}_N^\top \mathbf{x}_t\zeta_t$ and $\sigma_{\mathcal{S}}^{\ast}=\text{Var}^\ast(\mathcal{S}_{NT}^*)^{\frac{1}{2}}$. We can observe that $\widetilde{S}_{NT}^*/\widetilde{\sigma}^\ast=\mathcal{S}_{NT}^*/\sigma_{\mathcal{S}}^\ast$. We aim to establish the Edgeworth expansion for $\widetilde{\mathcal{S}}_{NT}^*\coloneqq \mathcal{S}_{NT}^*/\sigma_{\mathcal{S}}^\ast$ through investigating its characteristic function:
\begin{equation*}
\phi^\ast_{NT}(u)=E^\ast [\exp ( \mathsf{i} u\widetilde{\mathcal{S}}_{NT}^* )].
\end{equation*}
Specifically, we adopt a proof strategy similar to those used by tik1981. Let $\mathcal{T}_m \coloneqq \lfloor \frac{T}{m} \rfloor$, where we decompose $\widetilde{\mathcal{S}}_{NT}^*$ into a summation of 1-dependent series, conditional on the observations:
\begin{equation*}
\widetilde{\mathcal{S}}_{NT}^*=\sum_{s=1}^{\mathcal{T}_m} A_s+ A_{\mathcal{T}_m+1},
\end{equation*}
where $A_s \coloneqq \frac{1}{\sigma_{\mathcal{S}}^\ast}\sum_{t=(s-1)m+1}^{sm} \mathbf{1}_N^\top \mathbf{x}_t\zeta_t$ and $A_{\mathcal{T}_m+1} \coloneqq \frac{1}{\sigma_{\mathcal{S}}^\ast}\sum_{t=\mathcal{T}_m+1}^{T} \mathbf{1}_N^\top \mathbf{x}_t\zeta_t$. Additionally, define
\begin{eqnarray*}
A_{s,l}^c=\sum_{|t-s|>l} A_t \quad and \quad A_{s,0}^c = \widetilde{\mathcal{S}}_{NT}^*,
\end{eqnarray*}
for $l=1,\ldots, 4$ and $s=1,\ldots, \mathcal{T}_m+1$. By the nature of 1-dependent random variables, we know that $A_{s,l}^c$ and $A_{s}$ are independent conditional on the observations for $l=1,\ldots, 4$. This decomposition enables us to apply established techniques to handle such series.
We are now ready to make the following expansion of $\exp ( \mathsf{i} u\widetilde{\mathcal{S}}_{NT}^* )$ for each given $s$:
\begin{eqnarray}
\exp ( \mathsf{i} u\widetilde{\mathcal{S}}_{NT}^* )&=& \exp ( \mathsf{i} u A_{s,1}^c )+ \{\exp ( \mathsf{i} u (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c ) )-1 \}\exp ( \mathsf{i} u A_{s,1}^c ),
\end{eqnarray}
and this process can be iterated for the second term involving $\exp ( \mathsf{i} u A_{s,1}^c )$ and the subsequent terms. Finally, we obtain
\begin{eqnarray}
\exp ( \mathsf{i} u\widetilde{\mathcal{S}}_{NT}^* )&=&\exp ( \mathsf{i} u A_{s,1}^c )+ \{\exp ( \mathsf{i} u (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c ) )-1 \}\exp ( \mathsf{i} u A_{s,2}^c )
\notag\\
&&+\sum_{l=3}^{4}\prod_{k=1}^{l-1} \{\exp ( \mathsf{i} u (A_{s,k-1}^c - A_{s,k}^c ) )-1 \}\exp ( \mathsf{i} u A_{s,l}^c )\notag\\
&&+\prod_{k=1}^{4} \{\exp ( \mathsf{i} u (A_{s,k-1}^c - A_{s,k}^c ) )-1 \}\exp ( \mathsf{i} u A_{s,4}^c ).
\end{eqnarray}
Accordingly, we can write the first-order derivative of $\phi^\ast_{NT}(u)$ as
\begin{eqnarray}
\frac{\mathrm{d}\phi^\ast_{NT}(u)}{\mathrm{d}u}&=&\mathsf{i}E^\ast [\widetilde{\mathcal{S}}_{NT}^* \exp ( \mathsf{i} u \widetilde{\mathcal{S}}_{NT}^* ) ]
\notag\\
&=& \mathsf{i}\sum_{s=1}^{\mathcal{T}_m+1} E^\ast [A_s\exp ( \mathsf{i} u A_{s,1}^c ) ]\notag \\
&&+\mathsf{i}\sum_{s=1}^{\mathcal{T}_m+1} E^\ast [A_s \{\exp ( \mathsf{i} u (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c ) )-1 \}\exp ( \mathsf{i} u A_{s,2}^c ) ] \notag\\
&&+\mathsf{i}\sum_{s=1}^{\mathcal{T}_m+1}\sum_{l=3}^{4} E^\ast \Big[A_s\prod_{k=1}^{l-1} \{\exp ( \mathsf{i} u (A_{s,k-1}^c - A_{s,k}^c ) )-1 \}\exp ( \mathsf{i} u A_{s,l}^c ) \Big] \notag\\
&&+\mathsf{i}\sum_{s=1}^{\mathcal{T}_m+1} E^\ast\Big[A_s \prod_{k=1}^{4} \{\exp ( \mathsf{i} u (A_{s,k-1}^c - A_{s,k}^c ) )-1 \}\exp ( \mathsf{i} u A_{s,4}^c )\Big]\notag\\
&\eqqcolon &\mathsf{J}_{NT,1}(u)+\cdots+\mathsf{J}_{NT,4}(u).
\end{eqnarray}
We then investigate the four terms on the right hand side one by one.
For the first term, it is clear to see that
\begin{eqnarray}
\mathsf{J}_{NT,1}(u) =\mathsf{i}\sum_{s=1}^{\mathcal{T}_m+1} E^\ast[A_s]E^\ast [\exp ( \mathsf{i} u A_{s,1}^c)]=0.
\end{eqnarray}
To investigate $\mathsf{J}_{NT,2}(u)$, we invoke Lemma (ref) by letting $r=2$,
\begin{eqnarray}
\mathsf{J}_{NT,2}(u)&=&\mathsf{i}\sum_{s=1}^{\mathcal{T}_m+1} E^\ast\Big[A_s\Big\{\mathsf{i}u (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c )-\frac{1}{2}u^2 (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c )^2+R^\ast_{s,1}(u)\Big\}\exp ( \mathsf{i} u A_{s,2}^c )\Big]
\notag\\
&=&-u\sum_{s=1}^{\mathcal{T}_m+1} E^\ast [A_s (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c )\exp ( \mathsf{i} u A_{s,2}^c ) ]\notag \\
&&-\frac{1}{2}\mathsf{i}u^2\sum_{s=1}^{\mathcal{T}_m+1} E^\ast [A_s (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c )^2\exp ( \mathsf{i} u A_{s,2}^c ) ]
\notag\\
&&+\mathsf{i}\sum_{s=1}^{\mathcal{T}_m+1}E^\ast [A_s R^\ast_{s,1} (u)\exp ( \mathsf{i} u A_{s,2}^c ) ]
\notag\\
&\eqqcolon &\mathsf{J}_{NT,2,1}(u)+\mathsf{J}_{NT,2,2}(u)+\mathsf{J}_{NT,2,3}(u),
\end{eqnarray}
where $R^\ast_{s,1}(u)=\sum_{l=3}^{\infty}\frac{\mathsf{(iu)}^l}{l!} (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c )^l$ satisfies $|R^\ast_{s,1}(u)|\leq \frac{\mathsf{u}^3}{3!} |\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c |^3$.
Since $A_s$ and $\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c$ are conditionally independent with $A_{s,2}^c$, we can further write
\begin{eqnarray*}
\mathsf{J}_{NT,2,1}(u)&=&-u\sum_{s=1}^{\mathcal{T}_m+1} E^\ast [A_s (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c ) ]E^\ast [\exp ( \mathsf{i} u A_{s,2}^c ) ] \notag\\
&=&u\sum_{s=1}^{\mathcal{T}_m+1} E^\ast [A_s (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c ) ]E^\ast [\exp ( \mathsf{i} u \widetilde{\mathcal{S}}_{NT}^* )-\exp ( \mathsf{i} u A_{s,2}^c ) ]\notag\\
&&-u\sum_{s=1}^{\mathcal{T}_m+1} E^\ast [A_s (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c ) ] \phi^\ast_{NT}(u),
\end{eqnarray*}
where the second equality follows from the fact that $\phi^\ast_{NT}(u)=E^\ast [\exp ( \mathsf{i} u \widetilde{\mathcal{S}}_{NT}^*)]$.
For the first term on the right hand side of $\mathsf{J}_{NT,2,1}(u)$, we adopt a similar expansion procedure for $\exp ( \mathsf{i} u \widetilde{\mathcal{S}}_{NT}^* )$ as in (ref), and obtain
\begin{eqnarray}
&&\exp ( \mathsf{i} u \widetilde{\mathcal{S}}_{NT}^* )-\exp ( \mathsf{i} u A_{s,2}^c ) \notag \\
&=& \{\exp ( \mathsf{i} u (\widetilde{\mathcal{S}}_{NT}^* - A_{s,2}^c ) )-1 \}\exp ( \mathsf{i} u A_{s,2}^c )\notag\\
&=& \{\exp ( \mathsf{i} u (\widetilde{\mathcal{S}}_{NT}^* - A_{s,2}^c ) )-1 \}\exp ( \mathsf{i} u A_{s,3}^c )\notag\\
&&+ \{\exp ( \mathsf{i} u (\widetilde{\mathcal{S}}_{NT}^* - A_{s,2}^c ) )-1 \} \{\exp (\mathsf{i} u (A_{s,2}^c - A_{s,3}^c ) )-1 \}\exp ( \mathsf{i} u A_{s,3}^c ).
\end{eqnarray}
For the first term in (ref),
\begin{eqnarray}
&&|E^\ast [ \{\exp (\mathsf{i} u (\widetilde{\mathcal{S}}_{NT}^* - A_{s,2}^c ) )-1 \}\exp ( \mathsf{i} u A_{s,3}^c ) ] | \notag \\
&\leq& |E^\ast [ \{\exp\big(\mathsf{i} u\big(\widetilde{\mathcal{S}}_{NT}^* - A_{s,2}^c ) )-1 \} |\notag\\
&\leq& u |E^\ast [\widetilde{\mathcal{S}}_{NT}^* - A_{s,2}^c ] |+\frac{u^2}{2}E^\ast [ (\widetilde{\mathcal{S}}_{NT}^* - A_{s,2}^c )^2 ]\notag\\
&=&\frac{u^2}{2}E^\ast [ (\widetilde{\mathcal{S}}_{NT}^* - A_{s,2}^c )^2 ],
\end{eqnarray}
for any given $u$, where the first inequality holds by the conditional independence between $\widetilde{\mathcal{S}}_{NT}^* - A_{s,2}^c$ and $A_{s,3}^c$ and $|E[\exp( \mathsf{i} u A_{s,3}^c)]|\leq1$, and the second inequality is an application of Lemma (ref) with $r=1$.
Additionally, we can use the inequality $|e^{iu_1}-e^{iu_2}|\leq |u_1-u_2|$ and then Cauchy-Schwarz inequality sequentially to obtain
\begin{eqnarray}
&& |E^\ast [ \{\exp ( \mathsf{i} u(\widetilde{\mathcal{S}}_{NT}^* - A_{s,2}^c))-1\}\{\exp(\mathsf{i} u\big(A_{s,2}^c - A_{s,3}^c))-1\}\exp ( \mathsf{i} u A_{s,3}^c ) ] |\notag\\
&\leq&E^\ast [ |\exp ( \mathsf{i} u (\widetilde{\mathcal{S}}_{NT}^* - A_{s,2}^c ) )-1 |\cdot |\exp (\mathsf{i} u (A_{s,2}^c - A_{s,3}^c ) )-1 | ]\notag\\
&\leq&u^2 E^\ast [ (\widetilde{\mathcal{S}}_{NT}^* - A_{s,2}^c )^2 ]^{\frac{1}{2}}E^\ast [ (A_{s,2}^c - A_{s,3}^c )^2 ]^{\frac{1}{2}}.
\end{eqnarray}
By (ref), (ref), and (ref),
\begin{eqnarray*}
&& |E^\ast [\exp ( \mathsf{i} u \widetilde{\mathcal{S}}_{NT}^* )-\exp ( \mathsf{i} u A_{s,2}^c ) ] |\notag \\
&\le &\frac{u^2}{2}E^\ast [ (\widetilde{\mathcal{S}}_{NT}^* - A_{s,2}^c )^2 ] +u^2 E^\ast [ (\widetilde{\mathcal{S}}_{NT}^* - A_{s,2}^c )^2 ]^{\frac{1}{2}}E^\ast [ (A_{s,2}^c - A_{s,3}^c )^2 ]^{\frac{1}{2}}.
\end{eqnarray*}
Substituting this result into the first term in $\mathsf{J}_{NT,2,1}(u)$, we obtain
\begin{eqnarray}
&&\sum_{s=1}^{\mathcal{T}_m+1} |E^\ast [A_s (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c ) ] |\cdot |E^\ast [\exp ( \mathsf{i} u \widetilde{\mathcal{S}}_{NT}^* )-\exp ( \mathsf{i} u A_{s,2}^c ) ] |\notag\\
&\leq&\frac{u^2}{2}\sum_{s=1}^{\mathcal{T}_m+1} |E^\ast [A_s (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c ) ] |\cdot E^\ast [ (\widetilde{\mathcal{S}}_{NT}^* - A_{s,2}^c )^2 ]\notag\\
&&+u^2\sum_{s=1}^{\mathcal{T}_m+1} |E^\ast [A_s (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c ) ] |\cdot E^\ast [ (\widetilde{\mathcal{S}}_{NT}^* - A_{s,2}^c )^2 ]^{\frac{1}{2}}E^\ast [ (A_{s,2}^c - A_{s,3}^c )^2 ]^{\frac{1}{2}}\notag\\
&=&O_P\left(\left(\frac{m}{T}\right)^2\mathcal{T}_m \right) =O_P\left(\frac{m}{T}\right)
\end{eqnarray}
for any given $u$, where the first equality holds by Lemma (ref) and the second equality holds because $m\mathcal{T}_m=O(T)$.
For the second term in $\mathsf{J}_{NT,2,1}(u)$,
\begin{eqnarray}
\sum_{s=1}^{\mathcal{T}_m+1} E^\ast [A_s (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c ) ] &=& \sum_{s=1}^{\mathcal{T}_m+1} E^\ast [A_s (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c ) ]\notag\\
&=&\sum_{s=1}^{\mathcal{T}_m+1} E^\ast [A_s ]\widetilde{\mathcal{S}}_{NT}^*
\notag\\
&=&E^\ast [\widetilde{\mathcal{S}}_{NT}^{*2} ]=1.
\end{eqnarray}
By (ref) and (ref), we have
\begin{eqnarray}
\mathsf{J}_{NT,2,1}(u)&=&-u\phi^\ast_{NT}(u)+O_P\left(\frac{m}{T}\right),
\end{eqnarray}
for any given $u\in(-\infty,\infty)$.
For $\mathsf{J}_{NT,2,2}(u)$, using arguments similar to those in (ref), we obtain
\begin{eqnarray}
\mathsf{J}_{NT,2,2}(u)&=&-\frac{1}{2}\mathsf{i}u^2\sum_{s=1}^{\mathcal{T}_m+1} E^\ast [A_s (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c )^2 ]\phi^\ast_{NT}(u)+O_P\left(\frac{m}{T}\right).
\end{eqnarray}
For $\mathsf{J}_{NT,2,3}(u)$, since $|R^\ast_{s,1}(u)|\leq \frac{\mathsf{u}^3}{3!} |\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c |^3$, it is straightforward to show
\begin{eqnarray}
\mathsf{J}_{NT,2,3}(u)=O_P\left(\frac{m}{T}\right).
\end{eqnarray}
Combining (ref), (ref), and (ref) gives
\begin{eqnarray}
\mathsf{J}_{NT,2}(u)=-u\phi^\ast_{NT}(u)-\frac{1}{2}\mathsf{i}u^2\sum_{s=1}^{\mathcal{T}_m+1} E^\ast [A_s (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c )^2 ]\phi^\ast_{NT}(u)+O_P\left(\frac{m}{T}\right).
\end{eqnarray}
After finishing the investigation of $\mathsf{J}_{NT,2}(u)$, we proceed to study $\mathsf{J}_{NT,3}(u)$. In a similar way to the development for $\mathsf{J}_{NT,2,1}(u)$, we can easily show that
\begin{eqnarray*}
\mathsf{J}_{NT,3}(u)&=&\mathsf{i}\sum_{s=1}^{\mathcal{T}_m+1}\sum_{l=3}^{4} E^\ast\Big[A_s\prod_{k=1}^{l-1} \{\exp ( \mathsf{i} u (A_{s,k-1}^c - A_{s,k}^c ) )-1 \}\Big] \phi^\ast_{NT}(u)+o_P\left(\frac{m}{T}\right) \notag\\
&\eqqcolon& \mathsf{J}_{NT,3,1}(u)+ \mathsf{J}_{NT,3,2}(u),
\end{eqnarray*}
where $\mathsf{J}_{NT,3,1}(u)$ and $\mathsf{J}_{NT,3,2}(u)$ contain the terms with $l=3$ and $l=4$, respectively.
For $\mathsf{J}_{NT,3,1}(u)$, we invoke the Taylor expansion and the inequality in Lemma (ref) (with $r=1$) to write
\begin{eqnarray*}
&&\mathsf{J}_{NT,3,1}(u)\notag \\
&=&\mathsf{i}\sum_{s=1}^{\mathcal{T}_m+1}E^\ast\Big[A_s\big(\mathsf{i}u\big(\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c\big)+ R^\ast_{s,2} (u)\big)\big(\mathsf{i}u\big(A_{s,1}^c- A_{s,2}^c\big)+ R^\ast_{s,3} (u)\big)\Big]\phi^\ast_{NT}(u),
\end{eqnarray*}
where $R^\ast_{s,2}(u)=\sum_{l=2}^{\infty}\frac{\mathsf{(iu)}^l}{l!} \big(\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c\big)^l$ and $R^\ast_{s,3}(u)=\sum_{l=2}^{\infty}\frac{\mathsf{(iu)}^l}{l!} \big(A_{s,1}^c- A_{s,2}^c\big)^l$ satisfy $|R^\ast_{s,2}(u)|\leq \frac{\mathsf{u}^2}{2}\big(\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c\big)^2$ and $|R^\ast_{s,3}(u)|\leq \frac{\mathsf{u}^2}{2}\big(A_{s,1}^c- A_{s,2}^c\big)^2$.
By Lemma (ref), we have
\begin{eqnarray}
\mathsf{J}_{NT,3,1}(u)&=&-\mathsf{i}u^2\sum_{s=1}^{\mathcal{T}_m+1}E^\ast [A_s (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c ) (A_{s,1}^c- A_{s,2}^c ) ]\phi^\ast_{NT}(u)+O_P\left(\frac{m}{T}\right).
\end{eqnarray}
Analogously to (ref), the inequality $|e^{iu_1}-e^{iu_2}|\leq |u_1-u_2|$ can be applied here to show that
\begin{eqnarray*}
\mathsf{J}_{NT,3,2}(u)&=&O_P\left(\frac{m}{T}\right).
\end{eqnarray*}
Together with (ref), it yields
\begin{eqnarray}
\mathsf{J}_{NT,3}(u)&=&-\mathsf{i}u^2\sum_{s=1}^{\mathcal{T}_m+1}E^\ast [A_s (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c ) (A_{s,1}^c- A_{s,2}^c ) ]\phi^\ast_{NT}(u)+O_P\left(\frac{m}{T}\right).
\end{eqnarray}
For $\mathsf{J}_{NT,3}(u)$, we can use similar arguments to obtain
\begin{eqnarray}
\mathsf{J}_{NT,4}(u)&=&o_P\left(\frac{m}{T}\right).
\end{eqnarray}
In summary of (ref), (ref), (ref), (ref) and (ref),
\begin{eqnarray*}
\frac{\mathrm{d}\phi^\ast_{NT}(u)}{\mathrm{d}u}&=&-u\phi^\ast_{NT}(u)-\frac{1}{2}\mathsf{i}u^2\sum_{s=1}^{\mathcal{T}_m+1} E^\ast [A_s (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c )^2 ]\phi^\ast_{NT}(u)\notag\\
&&-\mathsf{i}u^2\sum_{s=1}^{\mathcal{T}_m+1}E^\ast [A_s (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c ) (A_{s,1}^c- A_{s,2}^c ) ]\phi^\ast_{NT}(u)+O_P\left(\frac{m}{T}\right).
\end{eqnarray*}
It is worth noting that
\begin{eqnarray*}
&&\sum_{s=1}^{\mathcal{T}_m+1} E^\ast [A_s ( (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c )^2+2\big(\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c ) (A_{s,1}^c- A_{s,2}^c ) ) ]\notag\\
&=&\sum_{s=1}^{\mathcal{T}_m+1} E^\ast [A_s (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c ) (\widetilde{\mathcal{S}}_{NT}^* + A_{s,1}^c-2A_{s,2}^c ) ]\notag\\
&=&\sum_{s=1}^{\mathcal{T}_m+1}E^\ast [A_s (\widetilde{\mathcal{S}}_{NT}^{*2} - A_{s,1}^{c2} )]-2\sum_{s=1}^{\mathcal{T}_m+1}E^\ast [A_s (\widetilde{\mathcal{S}}_{NT}^* - A_{s,1}^c )A_{s,2}^c ]\notag\\
&=&E^\ast [\widetilde{\mathcal{S}}_{NT}^{*3} ].
\end{eqnarray*}
Therefore, it follows that
\begin{eqnarray*}
\frac{\mathrm{d}\phi^\ast_{NT}(u)}{\mathrm{d}u}&=&-u\phi^\ast_{NT}(u)-\frac{1}{2}\mathsf{i}u^2 E^\ast [\widetilde{\mathcal{S}}_{NT}^{*3} ]\phi^\ast_{NT}(u)
+O_P\left(\frac{m}{T}\right)
\end{eqnarray*}
for any given $u\in(-\infty, \infty)$. Integration of this relation gives
\begin{eqnarray*}
\phi^\ast_{NT}(u)&=&\exp\big(-\frac{1}{2}u^2-\frac{1}{6}\mathsf{i}u^3 E^\ast [\widetilde{\mathcal{S}}_{NT}^{*3} ]\big) +O_P\left(\frac{m}{T}\right)\notag\\
&=&\exp\big(-\frac{1}{2}u^2\big)\Big\{1-\frac{1}{6}\mathsf{i}u^3 E^\ast [\widetilde{\mathcal{S}}_{NT}^{*3} ]\Big\} +O_P\left(\frac{m}{T}\right),
\end{eqnarray*}
where the second equality holds due to Taylor expansion of $\exp\left(-\frac{1}{6}\mathsf{i}u^3 E^\ast\left[\widetilde{\mathcal{S}}_{NT}^{*3}\right]\right)$, and the higher-order terms in this expansion, beyond the second order, are all bounded by the probability order $O_P\left(\frac{m}{T}\right)$. Moreover, by Esseen smoothing inequality in Lemma (ref), we can finally establish the desired Edgeworth expansion for the CDF of $\widetilde{\mathcal{S}}_{NT}^\ast$:
\begin{eqnarray*}
\sup_{u\in \mathbb{R}}\left|\normalfont Pr^\ast(\widetilde{\mathcal{S}}_{NT}^{*}\le u) -\Phi(u) -\frac{1-u^2}{6}E^\ast\big[\widetilde{\mathcal{S}}_{NT}^{*3}\big]\phi(x)\right|=O_P\left(\frac{m}{T}\right).
\end{eqnarray*}
This completes the proof of Theorem (ref).1.
(2). Since $E^\ast\big[\widetilde{S}_{NT}^{*3}\big]=O_P\big(\sqrt{\frac{m}{T}}\big)$ and $\widetilde{\sigma}^\ast=1+o_P(1)$, it follows from Theorem (ref).1 that
\begin{eqnarray}
\sup_{u\in \mathbb{R}}\left|\normalfont Pr^*(\widetilde{S}_{NT}^* \le u) -\Phi(u)\right|=O_P\left(\sqrt{\frac{m}{T}}\right).
\end{eqnarray}
Combining the results in Theorem (ref) and (ref), we can readily obtain
\begin{eqnarray*}
&&\sup_{u\in \mathbb{R}}\left|\normalfont Pr^*(\widetilde{S}_{NT}^* \le u) -\Pr(\widetilde{S}_{NT}\le u)\right|\notag \\
&\leq&\sup_{u\in \mathbb{R}}\left|\normalfont Pr^*(\widetilde{S}_{NT}^* \le u) -\Phi(u)\right|
+\sup_{u\in \mathbb{R}}\left|\Pr(\widetilde{S}_{NT}\le u)-\Phi(u)\right|
\notag\\
&=&O_P\left(\sqrt{\frac{m}{T}}\right).
\end{eqnarray*}
It leads to the desired result in Theorem (ref).2.
(3). To prove Theorem (ref).3, it suffices to show the following results:
\begin{eqnarray}
&&E[E^\ast[\widetilde{S}_{NT}^{*2}]]-E[\widetilde{S}_{NT}^{2} ]=-\frac{C_{q_a}}{ \sigma_x^2m^{q_a}}\Delta_{q_a}+o_P(m^{-q_a}),
\notag\\
&&Var( E^\ast[\widetilde{S}_{NT}^{*2}])=\frac{2m}{T}\int_{-1}^1 a^2(u)du+o\left(\frac{m}{T}\right).
\end{eqnarray}
where $\Delta_{q_a}=L_N^{-1}\sum_{s=-\infty}^{\infty}\sum_{\ell=0}^{\infty} |s|^{q_a}\mathbf{1}_N^\top \mathbf{B}_\ell \mathbf{B}_{\ell+|s|}^\top \mathbf{1}_N $ and
$C_{q_a}$ is defined in Theorem (ref).
It is clear to see that
\begin{eqnarray*}
\sigma_x^2\big(E[E^\ast[\widetilde{S}_{NT}^{*2}]]-E[\widetilde{S}_{NT}^{2} ]\big)&=&
\frac{1}{L_NT}\sum_{t=1}^T\sum_{s=1}^T \Big\{ a\Big(\frac{t-s}{m}\Big)-1\Big\} \mathbf{1}_N^\top E\big[ \mathbf{x}_{t}\mathbf{x}_{s}^\top\big]\mathbf{1}_N
\notag\\
&=& \frac{1}{L_N}\sum_{s=-m}^{m} \Big\{ a\Big(\frac{s}{m}\Big)-1\Big\} \mathbf{1}_N^\top E\big[ \mathbf{x}_{1}\mathbf{x}_{1+s}^\top\big]\mathbf{1}_N
\notag\\
&&-\frac{2}{L_N}\sum_{s=1}^{m}\frac{s}{T} \Big\{ a\Big(\frac{s}{m}\Big)-1\Big\} \mathbf{1}_N^\top E\big[ \mathbf{x}_{1}\mathbf{x}_{1+s}^\top\big]\mathbf{1}_N
\notag\\
&&-\frac{2}{L_N}\sum_{s=m+1}^{T-1}\frac{T-s}{T} \mathbf{1}_N^\top E\big[ \mathbf{x}_{1}\mathbf{x}_{1+s}^\top\big]\mathbf{1}_N
\notag\\
&:=&\mathcal{I}_1+\mathcal{I}_2+\mathcal{I}_3,
\end{eqnarray*}
where the definitions of $\mathcal{I}_1$, $\mathcal{I}_2$ and $\mathcal{I}_3$ are obvious. Here, the expression for $\mathcal{I}_3$ is derived from the fact that $a(s/m)=0$ for $s \geq m+1$.
We now proceed to investigate each term individually. For the first term, we employ the properties of the kernel function as specified in Assumption (ref) to establish its convergence. For $\forall \epsilon>0$, let $\epsilon^\ast = \frac{1}{2}\epsilon|\Delta_{q_a}|^{-1}$.
By Assumption (ref), there exists a positive constant $ \varsigma_\epsilon $ such that $\Big|\frac{1-a(u)}{|u|^{q_a}}-C_{q_a}\Big|<\epsilon^\ast$ for any $|u|\leq \varsigma_\epsilon$. Let $m^\ast = \lfloor m \varsigma_\epsilon\rfloor $. Without loss of generality, we assume $m^\ast\leq m$, as $ \varsigma_\epsilon $ can be chosen sufficiently small. Then, we can write
\begin{eqnarray*}
\mathcal{I}_1&=&\frac{1}{L_N}\sum_{s=-m^\ast}^{m^\ast} \Big\{ a\Big(\frac{s}{m}\Big)-1\Big\} \mathbf{1}_N^\top E\big[ \mathbf{x}_{1}\mathbf{x}_{1+s}^\top\big]\mathbf{1}_N\notag \\
&&+\frac{2}{L_N}\sum_{s=m^\ast+1}^{m} \Big\{a\Big(\frac{s}{m}\Big)-1\Big\} \mathbf{1}_N^\top E\big[ \mathbf{x}_{1}\mathbf{x}_{1+s}^\top\big]\mathbf{1}_N
\notag\\
&:=&\mathcal{I}_{1,1}+\mathcal{I}_{1,2}.
\end{eqnarray*}
For $ \mathcal{I}_{1,1}$, it is clear to see that
\begin{eqnarray*}
\mathbf{1}_N^\top E\big[ \mathbf{x}_{1}\mathbf{x}_{1+|s|}^\top\big]\mathbf{1}_N&=&\sum_{\ell=0}^{\infty} [(\mathbf{1}_N^\top \mathbf{B}_{\ell+|s|})\otimes (\mathbf{1}_N^\top \mathbf{B}_\ell)]E[vec(\pmb{\varepsilon}_{1-\ell} \pmb{\varepsilon}_{1-\ell}^\top)]
\notag\\
&&+\sum_{\ell=0}^{\infty}\sum_{v=0}^{|s|+\ell-1}[(\mathbf{1}_N^\top \mathbf{B}_v)\otimes (\mathbf{1}_N^\top \mathbf{B}_{\ell})] E[\text{vec}(\pmb{\varepsilon}_{1-\ell} \pmb{\varepsilon}_{1+|s|-v}^\top )]
\notag\\
&&+\sum_{\ell=0}^{\infty}\sum_{v>|s|+\ell}^{\infty} [(\mathbf{1}_N^\top \mathbf{B}_v)\otimes (\mathbf{1}_N^\top \mathbf{B}_{\ell})]E[ \text{vec}(\pmb{\varepsilon}_{1-\ell} \pmb{\varepsilon}_{1+|s|-v}^\top )]
\notag\\
&=&\sum_{\ell=0}^{\infty} \mathbf{1}_N^\top \mathbf{B}_\ell \mathbf{B}_{\ell+|s|}^\top \mathbf{1}_N.
\end{eqnarray*}
Note that \begin{eqnarray*}
&&L_N^{-1}\sum_{k=1}^{\infty}\sum_{\ell=0}^{\infty}k^{q_a} \big|\mathbf{1}_N^\top \mathbf{B}_\ell \mathbf{B}_{\ell+k}^\top \mathbf{1}_N\big|<\infty\notag \\
&\le & \frac{N}{L_N} \sum_{\ell=0}^{\infty}\|\mathbf{B}_\ell \|_2 \sum_{k=1}^{\infty}k^{q_a}\| \mathbf{B}_{\ell+k}\|_2\notag \\
&\le & \sqrt{\frac{N}{L_N}} \sum_{\ell=0}^{\infty}\|\mathbf{B}_\ell \|_2 \sqrt{\frac{N}{L_N}}\sum_{k=1}^{\infty}k^{q_a}\| \mathbf{B}_{k}\|_2<\infty,
\end{eqnarray*}
where the last line follows from Assumption (ref).1. Additionally, since $|s/m|\leq \varsigma_\epsilon$ for $|s|\leq m^\ast$, we have
\begin{eqnarray}
\big|m^{q_a}\mathcal{I}_{1,1}+C_{q_a}\Delta_{q_a}\big|<|\Delta_{q_a}|\epsilon^\ast=\frac{1}{2}\epsilon,
\end{eqnarray}
for sufficiently large $N$ and $T$. For $\mathcal{I}_{1,2}$, it is clear to see that $|s/m|\geq \varsigma_\epsilon$ and for sufficiently large $N$ and $T$,
\begin{eqnarray}
\big|m^{q_a}\mathcal{I}_{1,2}\big|\leq \frac{2(\max_u(a(u))+1)}{\varsigma_\epsilon^{q_a}L_N}\sum_{s=m^\ast+1}^{m}s^{q_a}\big|\mathbf{1}_N^\top E\big[ \mathbf{x}_{1}\mathbf{x}_{1+s}^\top\big]\mathbf{1}_N\big|
&<&\frac{1}{2}\epsilon,
\end{eqnarray}
where the inequality holds due to the fact that $L_N^{-1}\sum_{s=m^\ast+1}^{m}|s|^{q_a}\big|\mathbf{1}_N^\top E\big[ \mathbf{x}_{1}\mathbf{x}_{1+s}^\top\big]\mathbf{1}_N\big|$ converges to zero as both $m,m^\ast\rightarrow\infty$. By (ref) and (ref), we have
\begin{eqnarray*}
\big|m^{q_a}\mathcal{I}_{1}+C_{q_a}\Delta_{q_a}\big|<\epsilon,
\end{eqnarray*}
for sufficiently large $N$ and $T$. Since $\epsilon$ is an arbitrarily small positive number, we obtain
\begin{eqnarray}
\mathcal{I}_{1}=-C_{q_a}m^{-q_a}\Delta_{q_a}+o_P(m^{-q_a}).
\end{eqnarray}
We then proceed to study $\mathcal{I}_{2}$. Specifically, we have
\begin{eqnarray}
|\mathcal{I}_{2}|&\leq&\frac{2(\max_u(a(u))+1)}{L_NT}\sum_{s=1}^{m} s \big|\mathbf{1}_N^\top E\big[ \mathbf{x}_{1}\mathbf{x}_{1+s}^\top\big]\mathbf{1}_N\big|
\notag\\
&=&O\left(\frac{1}{T}\right)=o(m^{-q_a}),
\end{eqnarray}
where the second equality holds because $m^2/T \rightarrow 0$, as required in Assumption (ref).
For $\mathcal{I}_{3}$,
\begin{eqnarray}
|\mathcal{I}_{3}|&\leq& \frac{2}{L_N}\sum_{s=m+1}^{T-1} \big|\mathbf{1}_N^\top E\big[ \mathbf{x}_{1}\mathbf{x}_{1+s}^\top\big]\mathbf{1}_N\big|
\notag\\
&\leq& \frac{2}{L_Nm^2}\sum_{s=m+1}^{T-1}s^2 \big|\mathbf{1}_N^\top E\big[ \mathbf{x}_{1}\mathbf{x}_{1+s}^\top\big]\mathbf{1}_N\big|
\notag\\
&=&o\left(\frac{1}{m^2}\right)=o_P(m^{-q_a}),
\end{eqnarray}
where the first equality follows from
\begin{eqnarray*}
\frac{1}{L_N}\sum_{s=m+1}^{T-1}s^2 \big|\mathbf{1}_N^\top E\big[ \mathbf{x}_{1}\mathbf{x}_{1+s}^\top\big]\mathbf{1}_N\big|&\leq&\frac{1}{L_N}\sum_{s=m+1}^{T-1}\sum_{\ell=0}^{\infty}s^2 \big| \mathbf{1}_N^\top \mathbf{B}_\ell \mathbf{B}_{\ell+s}^\top \mathbf{1}_N\big| = o(1),
\end{eqnarray*}
with the equality holding by Assumption (ref).
Combining (ref), (ref) and (ref), we obtain the first desired result in (ref). Next, we investigate the variance of $E^\ast[\widetilde{S}_{NT}^{*2}]$.
For notational simplicity, let $\mathcal{U}_{ts}=L_N^{-1}\mathbf{1}_N^\top\mathbf{x}_{t}\mathbf{x}_{s}^\top\mathbf{1}_N$. Additionally, define
\begin{eqnarray*}
\varsigma (s)=E[\mathcal{U}_{1,1+s}],\,\upsilon (s_1,s_2,s_3)= E[\mathcal{U}_{1,1+s_1}\mathcal{U}_{1+s_2,1+s_3}],
\end{eqnarray*}
and $\varrho (s_1,s_2,s_3)=\upsilon (s_1,s_2,s_3)- \varsigma(s_1) \varsigma(s_2-s_3)- \varsigma(s_2) \varsigma(s_1-s_3)- \varsigma(s_3) \varsigma(s_1-s_2)$.
To derive the variance of $E^\ast[\widetilde{S}_{NT}^{*2}]$, we write
\begin{eqnarray*}
&& \sigma_x^4 Var( E^\ast[\widetilde{S}_{NT}^{*2}])
\notag\\
&=& \frac{1}{T^2}\sum_{t_1,t_2=1}^T\sum_{s_1=1-t_1}^{T-t_1}\sum_{s_2=1-t_2}^{T-t_2} a\Big(\frac{s_1}{m}\Big)a\Big(\frac{s_2}{m}\Big) Cov\big(\mathcal{U}_{t_1,t_1+s_1}, \mathcal{U}_{t_2,t_2+s_2}\big)
\notag\\
&=& \frac{1}{T^2}\sum_{s_1=-m}^{m}\sum_{s_2=-m}^{m} a\Big(\frac{s_1}{m}\Big)a\Big(\frac{s_2}{m}\Big) Cov\big(\sum_{t_1=\psi_{1,s_1}}^{\psi_{2,s_1}}\mathcal{U}_{t_1,t_1+s_1}, \sum_{t_2=\psi_{1,s_2}}^{\psi_{2,s_2}}\mathcal{U}_{t_2,t_2+s_2}\big)
,
\end{eqnarray*}
where $\psi_{1,s}=\max\{1,1-s\}$, $\psi_{2,s}=\min\{T,T-s\}$.
We initially focus on the case of $0\leq s_1\leq s_2$. Using standard results for the fourth moment of stationary time series Anderson2011, we obtain
\begin{eqnarray*}
\frac{1}{T-s_2}Cov\big(\sum_{t_1=1}^{T-s_1}\mathcal{U}_{t_1,t_1+s_1}, \sum_{t_2=1}^{T-s_2}\mathcal{U}_{t_2,t_2+s_2}\big)&:=&\varpi_{s_1,s_2,1}+ \varpi_{s_1,s_2,2}+\varpi_{s_1,s_2,3},
\end{eqnarray*}
where
\begin{eqnarray*}
\varpi_{s_1,s_2,1}&=& \sum_{r=0}^{s_2-s_1}\{\varsigma(r)\varsigma(r+s_1-s_2)+\varsigma(r-s_2)\varsigma(r+s_1)\}
\notag\\
&&+\sum_{r=s_2-s_1+1}^{T-s_1-1}\Big(1-\frac{r-(s_2-s_1)}{T-s_2}\Big)\{\varsigma(r)\varsigma(r+s_1-s_2)+\varsigma(r-s_2)\varsigma(r+s_1)\}
\notag\\
&&+\sum_{r=-(T-s_2-1)}^{-1}\Big(1-\frac{|r|}{T-s_2}\Big)\{\varsigma(r)\varsigma(r+s_1-s_2)+\varsigma(r-s_2)\varsigma(r+s_1)\}
\notag\\
\varpi_{s_1,s_2,2}&=& \sum_{r=0}^{s_2-s_1}\varrho (s_1,-r,s_2-r)
+\sum_{r=s_2-s_1+1}^{T-s_1-1}\Big(1-\frac{r-(s_2-s_1)}{T-s_2}\Big)\varrho (s_1,-r,s_2-r)
\notag\\
&&+\sum_{r=-(T-s_2-1)}^{-1}\Big(1-\frac{|r|}{T-s_2}\Big)\varrho (s_1,-r,s_2-r).
\end{eqnarray*}
Regarding terms with $\varpi_{s_1,s_2,1}$, applying the same arguments as in the proof of Corollary 8.3.1 of Anderson2011, we find that
\begin{eqnarray}
&&\frac{1}{T^2}\sum_{s_1= 0}^{m}\sum_{s_2=s_1+1}^{m} a\Big(\frac{s_1}{m}\Big)a\Big(\frac{s_2}{m}\Big) (T-s_2)\varpi_{s_1,s_2,1}
\notag\\
&=&\frac{1}{T}\sum_{s_1= 0}^{m}\sum_{s_2=s_1+1}^{m}\sum_{r=-\infty}^{\infty} a\Big(\frac{s_1}{m}\Big)a\Big(\frac{s_2}{m}\Big) \{\varsigma(r+s_2)\varsigma(r+s_1)+\varsigma(r-s_2)\varsigma(r+s_1)\}(1+o(1)).
\notag\\
\end{eqnarray}
Using equation (46) in Section 8.3 of Anderson2011, we can demonstrate that the terms with $\varpi_{s_1,s_2,2}$ are asymptotically negligible:
\begin{eqnarray}
&&\frac{1}{T^2}\sum_{s_1= 0}^{m}\sum_{s_2=s_1+1}^{m} a\Big(\frac{s_1}{m}\Big)a\Big(\frac{s_2}{m}\Big) (T-s_2)|\varpi_{s_1,s_2,2}|
\notag\\
&=&\frac{\kappa_4}{T^2}\sum_{s_1= 0}^{m}\sum_{s_2=s_1+1}^{m} a\Big(\frac{s_1}{m}\Big)a\Big(\frac{s_2}{m}\Big) (T-s_2)|\varsigma(s_1)||\varsigma(s_2)|(1+o(1))
\notag\\
&\leq&\frac{\kappa_4\max_u a(u)^2}{T}\sum_{s_1= 0}^{m}\sum_{s_2=s_1+1}^{m}|\varsigma(s_1)||\varsigma(s_2)|(1+o(1)) = O\left(\frac{1}{T}\right)=o\left(\frac{l}{T}\right),
\end{eqnarray}
where $\kappa_4=\sigma_x^{-4}E[|L_N^{-1/2}\mathbf{1}_N^\top\mathbf{x}_{t}|^4]-3$ and the second equality follows from $\sum_{s= 0}^{\infty}|\varsigma(s)|=O(1)$. Combining (ref) and (ref), we obtain
\begin{eqnarray*}
&&\frac{1}{T}\sum_{s_1= 0}^{m}\sum_{s_2=s_1+1}^{m} a\Big(\frac{s_1}{m}\Big)a\Big(\frac{s_2}{m}\Big) Cov\big(\sum_{t_1=1}^{T-s_1}\mathcal{U}_{t_1,t_1+s_1}, \sum_{t_2=1}^{T-s_2}\mathcal{U}_{t_2,t_2+s_2}\big)
\notag\\
&=&\frac{1}{T}\sum_{s_1= 0}^{m}\sum_{s_2=s_1+1}^{m} \sum_{r=-\infty}^\infty a\Big(\frac{s_1}{m}\Big)a\Big(\frac{s_2}{m}\Big) \{\varsigma(r+s_2)\varsigma(r+s_1)+\varsigma(r-s_2)\varsigma(r+s_1)\}+o\left(\frac{l}{T}\right).
\end{eqnarray*}
In the case of other pairings of $s_1$ and $s_2$, similar arguments yield the leading-order terms. Ultimately, we establish that
\begin{eqnarray*}
&&\sigma_x^4 Var( E^\ast[\widetilde{S}_{NT}^{*2}]) \notag \\
&=&\frac{1}{T}\sum_{s_1= -m}^{m}\sum_{s_2=-m}^{m} \sum_{r=-\infty}^\infty a\Big(\frac{s_1}{m}\Big)a\Big(\frac{s_2}{m}\Big) \{\varsigma(r+s_2)\varsigma(r+s_1)+\varsigma(r-s_2)\varsigma(r+s_1)\}
+o\left(\frac{l}{T}\right)
\notag\\
&:=&\mathcal{I}_4+\mathcal{I}_5+o\left(\frac{l}{T}\right),
\end{eqnarray*}
where the definitions of $\mathcal{I}_4$ and $\mathcal{I}_5$ are obvious.
For $\mathcal{I}_4$, we can further write
\begin{eqnarray}
\frac{T}{m}\mathcal{I}_4&=&\frac{1}{m} \sum_{r=-\infty}^\infty\sum_{s_1= r-m}^{r+m}\sum_{s_2=r-m}^{r+m} a\Big(\frac{s_1-r}{m}\Big)a\Big(\frac{s_2-r}{m}\Big) \varsigma(s_1)\varsigma(s_2)
\notag\\
&=&\frac{1}{m} \sum_{r=-\infty}^\infty\sum_{s_1= r-m}^{r+m}\sum_{s_2=r-m}^{r+m}a^2\Big(\frac{r}{m}\Big) \varsigma(s_1)\varsigma(s_2)+o(1)
\notag\\
&=&\sigma_x^4\int_{-1}^1 a^2(u)du +o(1),
\end{eqnarray}
where the second equality follows from Lipschitz continuity of $a(u)$, such that $|a(\frac{s-r}{m})-a(\frac{-r}{m})|=O(\frac{s}{m})$ and from the fact $\sum_{s=-\infty}^{\infty}|s||E[\mathcal{U}_{1,1+s}]|=O(1)$, both implied by Assumption (ref). By similar arguments, we can obtain
\begin{eqnarray}
\frac{T}{m}\mathcal{I}_5=\sigma_x^4\int_{-1}^1 a^2(u)du +o(1).
\end{eqnarray}
(ref) and (ref) together yield the second desired result in (ref). Therefore, Theorem (ref).3 is established.
proof[Proof of Theorem (ref)]
• Before proceeding, we introduce some notation that will be repeatedly used in this proof. Note that by Lemma (ref) $\widetilde{p}_i$ admits the following decomposition:
\begin{eqnarray*}
p_i &=& \mathbf{e}_i^\top \sum_{t=1}^T\mathbf{x}_t = \mathbf{e}_i^\top \mathbf{B} \sum_{s=1}^T\pmb{\varepsilon}_{s}-\mathbf{e}_i^\top\widetilde{\mathbf{B}}(L)\pmb{\varepsilon}_{T}+\mathbf{e}_i^\top\widetilde{\mathbf{B}}(L)\pmb{\varepsilon}_{0} = \mathbf{e}_i^\top\sum_{\ell =-\infty}^T \pmb{\mathcal{B}}_{T\ell}\pmb{\varepsilon}_\ell.
\end{eqnarray*}
Let $\mathbf{B}_\ell=(\mathbf{b}_{\ell 1}^\sharp,\ldots, \mathbf{b}_{\ell N}^\sharp)^\top$ and $\widetilde{\mathbf{B}}_{\ell}=(\widetilde{\mathbf{b}}_{\ell 1}^\sharp,\ldots, \widetilde{\mathbf{b}}_{\ell N}^\sharp)^\top$,
where $\mathbf{B}_\ell$ and $\widetilde{\mathbf{B}}_{\ell}$ have been defined in (ref) and (ref) respectively. We can then write
\begin{eqnarray}
\mathbf{e}_i^\top\widetilde{\mathbf{B}}(1)&=&\mathbf{e}_i^\top\sum_{\ell=0}^{\infty}\widetilde{\mathbf{B}}_\ell =\sum_{\ell=0}^{\infty}\widetilde{\mathbf{b}}_{\ell i}^{\sharp \top} = \sum_{\ell=0}^{\infty}\sum_{k=\ell+1}^\infty\mathbf{b}_{k i}^{\sharp\top} = \sum_{\ell=1}^{\infty}\ell \mathbf{b}_{\ell i}^{\sharp \top},
\end{eqnarray}
which yields that $E |\mathbf{e}_i^\top\widetilde{\mathbf{B}}(L)\pmb{\varepsilon}_{T} |^2 = \sum_{\ell=1}^{\infty}\ell^2 \|\mathbf{b}_{\ell i}^{\sharp}\|_2^2$. Therefore, $|\mathbf{e}_i^\top\widetilde{\mathbf{B}}(L)\pmb{\varepsilon}_{T}| =O_P(1)$. Similarly, we obtain that $|\mathbf{e}_i^\top\widetilde{\mathbf{B}}(L)\pmb{\varepsilon}_{0}| =O_P(1)$. We can then conclude that
\begin{eqnarray*}
Var\left(\frac{1}{\sqrt{T}}p_i\right) = \mathbf{e}_i^\top \mathbf{B}\mathbf{B}^\top \mathbf{e}_i +o(1)\eqqcolon \sigma_{p,i}^2+o(1).
\end{eqnarray*}
In what follows, we study $\frac{1}{\sigma_{p,i}\sqrt{T}}p_i$, and write
\begin{eqnarray*}
\widetilde{p}_i\coloneqq \frac{1}{\sigma_{p,i}\sqrt{T}}p_i =\frac{1}{\sigma_{p,i}\sqrt{T}} \mathbf{e}_i^\top\sum_{\ell =-\infty}^T \pmb{\mathcal{B}}_{T\ell}\pmb{\varepsilon}_\ell\eqqcolon \mathbf{e}_i^\top\sum_{\ell =-\infty}^T \dot{\pmb{\mathcal{B}}}_{\ell}\pmb{\varepsilon}_\ell,
\end{eqnarray*}
where $\frac{1}{\sigma_{p,i}\sqrt{T}}\pmb{\mathcal{B}}_{T\ell}\eqqcolon \dot{\pmb{\mathcal{B}}}_{\ell}= \{\dot{\mathtt{b}}_{\ell, ij}\}_{N\times N}$, and we have suppressed $T$ in $\dot{\pmb{\mathcal{B}}}_{\ell}$ for notational simplicity.
To proceed, denote by $\psi_{i}(u)$ the characteristic function of $\widetilde{p}_i$. Thus,
\begin{eqnarray}
\psi_i(u) &=& E\left[\exp\left(\mathsf{i}u \sum_{\ell=-\infty}^T\mathbf{e}_i^\top\dot{\pmb{\mathcal{B}}}_{\ell} \pmb{\varepsilon}_{\ell}\right)\right]\notag \\
&=&\prod_{\ell=-\infty}^T \prod_{j=1}^NE [\exp (\mathsf{i}u \dot{\mathtt{b}}_{\ell ,ij}\varepsilon_{j\ell} )] = \prod_{\ell=-\infty}^T \prod_{j=1}^N\psi(\dot{\mathtt{b}}_{\ell ,ij} u),
\end{eqnarray}
where the second equality follows from $\{\varepsilon_{it}\}$ being i.i.d. over both dimensions. We have suppressed $T$ in $\dot{\mathtt{b}}_{\ell ,ij}$ for notational simplicity.
Using (ref), we are able to calculate the $r^{th}$ cumulant $\beta_{ir}$ of $\widetilde{p}_i$ for $r\ge 1$. Obviously, we have $\beta_{i1} =0$ and $\beta_{i2} =1$. For $r\ge 3$, write
\begin{eqnarray}
\beta_{ir} &=& (-\mathsf{i})^r \frac{\mathrm{d}^r}{\mathrm{d}u^r}\log \psi_i(u) |_{u=0}\notag \\
&=&\sum_{\ell=-\infty}^T\sum_{j=1}^N (-\mathsf{i})^r\frac{\mathrm{d}^r}{\mathrm{d}u^r}\log\psi(\dot{\mathtt{b}}_{\ell ,ij} u) |_{u=0} \notag \\
&=&\sum_{\ell=-\infty}^T\sum_{j=1}^N \dot{\mathtt{b}}_{\ell ,ij}^r (-\mathsf{i})^r\frac{\mathrm{d}^r}{\mathrm{d}u^r}\log\psi(u)|_{u=0}\notag \\
&=&\kappa_r\left(\sum_{\ell=1}^T\sum_{j=1}^N \dot{\mathtt{b}}_{\ell ,ij}^r +\sum_{\ell=-\infty}^0\sum_{j=1}^N \dot{\mathtt{b}}_{\ell ,ij}^r \right),
\end{eqnarray}
where the second equality follows from (ref), and the third equality follows from (ref).
For the second term on the right hand side of (ref), we note that
\begin{eqnarray}
\max_i\left|\sum_{\ell=-\infty}^0\sum_{j=1}^N \dot{\mathtt{b}}_{\ell ,ij}^r \right|&\le & \max_i \frac{1}{(\sigma_{p,i}^2 T)^{r/2}}\sum_{\ell=-\infty}^0\sum_{j=1}^N|b_{T\ell, ij}|^r\notag \\
&\le &\max_i\frac{1}{(\sigma_{p,i}^2 T)^{r/2}}\left(\sum_{\ell=-\infty}^0\sum_{j=1}^N |b_{T\ell, ij}|^2\right)^{r/2}\notag\\
&\le &\max_i O(1)\frac{1}{T^{r/2}}\left(\sum_{\ell=-\infty}^0 \| \widetilde{\mathbf{b}}_{\ell i}^\sharp\|^2 \right)^{r/2} = O\left(\frac{1}{T^{r/2}}\right),
\end{eqnarray}
where $\pmb{\mathcal{B}}_{T\ell}\eqqcolon \{b_{T\ell, ij}\}_{N\times N}$, the second inequality follows from the fact that for a vector $\mathbf{x}$, $|\mathbf{x}|_{p_1}\le |\mathbf{x} |_{p_2}$ for any $p_1> p_2 \ge 1$, the third inequality follows from Lemma (ref).1, and the last equality follows from Assumption (ref) and (ref).
Thus, (ref) and (ref) together infer that
\begin{eqnarray*}
\max_i\left|\beta_{ir} -\kappa_r \sum_{\ell=1}^T\sum_{j=1}^N \dot{\mathtt{b}}_{\ell ,ij}^r\right|=O\left(\frac{1}{T^{r/2}}\right).
\end{eqnarray*}
Recall the definitions of $\dot{\mathtt{b}}_{\ell ,ij}^r$ for $\ell\ge 1$ and $j\ge 1$, and write further that
\begin{eqnarray*}
\max_i\left| \sum_{\ell=1}^T\sum_{j=1}^N \dot{\mathtt{b}}_{\ell ,ij}^r \right|&\le& \max_i\frac{2^{r-1}}{(\sigma_{p,i}^2 T)^{r/2}}\sum_{\ell=1}^T\sum_{j=1}^N[|b_{ij}|^r+| \widetilde{b}_{T-\ell, ij}|^r]\notag\\
&=& \max_i\frac{2^{r-1}}{\sigma_{p,i}^rT^{r/2-1}} \sum_{j=1}^N | b_{ij} |^r+O\left(\frac{1}{T^{r/2}}\right)\notag \\
&\le & \max_i\frac{2^{r-1}}{\sigma_{p,i}^rT^{r/2-1}} \left(\sum_{j=1}^N | b_{ij} |^2\right)^{r/2}+O\left(\frac{1}{T^{r/2}}\right)\notag \\
&=&\max_i\frac{2^{r-1}}{\sigma_{p,i}^rT^{r/2-1}} \|\mathbf{b}_{i}^{\sharp}\|_2^r+O\left(\frac{1}{T^{r/2}}\right),
\end{eqnarray*}
where the first inequality follows from $(a+b)^r\le 2^{r-1}(a^r+b^r)$ for $a,b\ge 0$ and $r>1$, the first equality follows from a development similar to (ref), and the second inequality follows from the fact that for a vector $\mathbf{x}$, $|\mathbf{x}|_{p_1}\le |\mathbf{x} |_{p_2}$ for any $p_1> p_2 \ge 1$. Thus, by (ref) and (ref), we can obtain that for $r\ge 3$
\begin{eqnarray}
|\beta_{ir}|& \le &\max_i\frac{2^{r-1}}{\sigma_{p,i}^rT^{r/2-1}} \|\mathbf{b}_{i}^{\sharp}\|_2^r+O\left(\frac{1}{T^{r/2}}\right) \notag \\
&=& O\left( \frac{1}{T^{r/2 -1}} \right) \le O\left( \frac{1}{(T^{1/6})^r} \right),
\end{eqnarray}
where the equality follows from Assumption (ref), and the second inequality follows from $r\ge 3$.
By applying (ref) to $\psi_i(u)$ and the characteristic function of $N(0,1)$, we have
\begin{eqnarray}
\psi_i(u) &=&\exp\left\{ \sum_{r=3}^\infty \beta_{ir}\cdot\frac{(\mathsf{i} u)^r}{r!}\right\} \cdot\widetilde{\phi}(u)\notag \\
&=&\left\{1 +\sum_{n=1}^\infty\frac{1}{n!}\left(\sum_{r=3}^\infty \beta_{ir}\cdot\frac{(\mathsf{i} u)^r}{r!}\right)^n \right\}\cdot\widetilde{\phi}(u),
\end{eqnarray}
where $\beta_{ir}$ is bounded by (ref), and the second equality follows from (ref). Let $f_i(w)$ be the density function of $\widetilde{p}_i$, and the Fourier inversion of (ref) leads to the following expansion:
\begin{eqnarray*}
f_i(w) &=&\phi(w) +\frac{1}{2\pi}\int_{\mathbb{R}} \exp(-\mathsf{i}uw)\cdot \beta_{i3}\cdot \frac{(\mathsf{i}u)^3}{3!}\cdot\widetilde{\phi}(u)\mathrm{d}u \notag \\
&&+\frac{1}{2\pi}\int_{\mathbb{R}} \exp(-\mathsf{i}uw)\cdot\widetilde{\psi}_i(u)\cdot\widetilde{\phi}(u)\mathrm{d}u\notag \\
&=&\phi(w) + \frac{\beta_{i3}}{6} H_3(w)\phi(w)\notag \\
&&+\frac{1}{2\pi}\int_{\mathbb{R}} \exp(-\mathsf{i}uw)\cdot\widetilde{\psi}_i(u)\cdot\widetilde{\phi}(u)\mathrm{d}u
\end{eqnarray*}
where $\widetilde{\psi}_i(u)\coloneqq \sum_{r=4}^\infty \beta_{ir}\cdot\frac{(\mathsf{i} u)^r}{r!} + \sum_{n=2}^\infty\frac{1}{n!}\left(\sum_{r=3}^\infty \beta_{ir}\cdot\frac{(\mathsf{i} u)^r}{r!}\right)^n$, and the second equality follows from Lemma (ref).
We then bound $\frac{1}{2\pi}\int_{\mathbb{R}} \exp(-\mathsf{i}uw)\cdot \widetilde{\psi}_i(u)\cdot\widetilde{\phi}(u)\mathrm{d}u$. First, note
\begin{eqnarray}
\left|\sum_{r=4}^\infty \beta_{ir}\cdot\frac{(\mathsf{i} u)^r}{r!} \right|\widetilde{\phi}(u) &\le & |\beta_{i4}| \widetilde{\phi}(u) \sum_{r=4}^\infty \frac{|u|^r}{r!} \notag \\
&\le & |\beta_{i4}| \widetilde{\phi}(u)\exp(|u|) =|\beta_{i4}| \exp(-u^2/2+|u|),
\end{eqnarray}
where the second inequality follows from (ref). Second, in connection with (ref), assuming that $|u|\le T^{1/6}$ and using Taylor theorem twice as in (ref) we write
\begin{eqnarray}
&&\left|\sum_{n=2}^\infty\frac{1}{n!}\left(\sum_{r=3}^\infty \beta_{ir}\cdot\frac{(\mathsf{i} u)^r}{r!}\right)^n \right|\widetilde{\phi}(u)\notag \\
&\le &O(1)\sum_{n=2}^\infty\frac{1}{n!} \left| \sum_{r=3}^\infty \cdot\frac{(\mathsf{i} u/T^{1/6} )^r}{r! }\right|^n \widetilde{\phi}(u)\notag\\
&\le & O(1)\frac{1}{T}\widetilde{\phi}(u) \widetilde{u}^6=O(1)\frac{1}{T},
\end{eqnarray}
where $|\widetilde{u}|$ is in between 0 and $T^{1/6}$, and the last equality follows from $\widetilde{\phi}(u) \widetilde{u}^6$ being uniformly bounded. Thus, by (ref), (ref) and (ref), we can write
\begin{eqnarray}
\sup_{|u|\le T^{1/6}}\left|\frac{1}{2\pi}\int_{\mathbb{R}} \exp(-\mathsf{i}uw)\cdot \widetilde{\psi}_i(u)\cdot\widetilde{\phi}(u)\mathrm{d}u\right| &=&O\left( \frac{1}{T} \right).
\end{eqnarray}
Thus, according to (ref) and (ref), we write further
\begin{eqnarray}
&&\sup_{w\in\mathbb{R}}\left|f_i(w)-\phi(w)-\frac{\beta_{i3}}{6} H_3(w)\phi(w) \right|\notag \\
&\le &\frac{1}{2\pi}\int_{\mathbb{R}}\left|\psi_i(u) -\widetilde{\phi}(u) - \frac{\beta_{i3}}{6} (\mathsf{i}u)^3 \widetilde{\phi}(u)\right|\mathrm{d}u\notag\\
&= &\frac{1}{2\pi}\int_{|u|\ge T^{1/6}} |\widetilde{\psi}_i(u) |\widetilde{\phi}(u)\mathrm{d}u +\frac{1}{2\pi}\int_{|u|< T^{1/6}} |\widetilde{\psi}_i(u) |\widetilde{\phi}(u)\mathrm{d}u = O\left( \frac{1}{T}\right),
\end{eqnarray}
where, in the last step, $\int_{|u|\ge T^{1/6}} |\widetilde{\psi}_i(u) |\widetilde{\phi}(u)\mathrm{d}u$ can be arbitrarily small by standard operation, and $\frac{1}{2\pi}\int_{|u|< T^{1/6}} |\widetilde{\psi}_i(u) |\widetilde{\phi}(u)\mathrm{d}u$ is bounded by (ref).
To study the CDF, we let $G_i(w) =\Phi(w)+\frac{\beta_{i3}}{6}(1-w^2)\phi(w)$. Simple algebra shows that
\begin{eqnarray*}
\frac{\mathrm{d}( (1-w^2)\phi(w) )}{\mathrm{d}w} = H_3(w)\phi(w) .
\end{eqnarray*}
Thus, $G_i(w)$ has a characteristic function $\xi_i(w) =(1+\frac{\beta_{i3}}{6}(\mathsf{i}w)^3)\widetilde{\phi}(w)$. We then invoke Esseen's smoothing Lemma. First, we let $a$ involved in Lemma (ref) be sufficiently large, so the second term on the right hand side of Lemma (ref) becomes negligible. Then, similar to (ref), we study the first term and can obtain that
\begin{eqnarray*}
|F_i(w)-G_i(w)|_\infty=O\left( \Delta_{NT}(4) \vee \frac{1}{T^2}\right).
\end{eqnarray*}
The proof of is now completed.
proof[Proof of Theorem (ref)]
• (1). For the heterogeneous bootstrap statistics, we can apply a similar approach to what has been used in the proof of Theorem (ref).1 to establish its Edgeworth expansion.
Let $\mathcal{P}_{i}^*= \sum_{t=1}^T\mathbf{e}_i^\top \mathbf{x}_t\zeta_{t}$ and $\sigma_{\mathcal{P},i}^{\ast}=\text{Var}^\ast(\mathcal{P}_{i}^*)^{\frac{1}{2}}$. We need to establish the Edgeworth expansion for $\widetilde{\mathcal{P}}_{i}^*\coloneqq \mathcal{P}_{i}^*/\sigma^\ast_{\mathcal{P},i}$.
First, we make the following decomposition for $\widetilde{\mathcal{P}}_{i}^*$:
\begin{equation*}
\widetilde{\mathcal{P}}_{i}^*=\sum_{s=1}^{\mathcal{T}_m} B_{i,s}+ B_{i,\mathcal{T}_m+1},
\end{equation*}
where $\mathcal{T}_m \coloneqq \lfloor \frac{T}{m} \rfloor$, $B_{i,s} \coloneqq \sigma_{\mathcal{P},i}^{\ast-1}\sum_{t=(s-1)m+1}^{sm} \mathbf{e}_i^\top \mathbf{x}_t\zeta_t$ and $B_{i,\mathcal{T}_m+1} \coloneqq \sigma_{\mathcal{P},i}^{\ast-1}\sum_{t=\mathcal{T}_m+1}^{T} \mathbf{e}_i^\top \mathbf{x}_t\zeta_t$. Additionally, define $B_{i,s,l}^c=\sum_{|t-s|>l} B_{i,t}$ and $B_{i,s,0}^c = \widetilde{\mathcal{P}}_{i}^*$
for $l=1,\ldots, 4$ and $s=1,\ldots, \mathcal{T}_m+1$.
To study the characteristic function $\phi^\ast_{i}(u)=E^\ast [\exp ( \mathsf{i} u\widetilde{\mathcal{P}}_{i}^* ) ]$, we require the following expansion for $\exp ( \mathsf{i} u\widetilde{\mathcal{P}}_{i}^* ) $, which can be derived analogously to (ref):
\begin{eqnarray*}
\exp ( \mathsf{i} u\widetilde{\mathcal{P}}_{i}^* )&=&\exp ( \mathsf{i} u B_{i,s,1}^c )+ \{\exp ( \mathsf{i} u (\widetilde{\mathcal{P}}_{i}^* - B_{i,s,1}^c ) )-1 \}\exp ( \mathsf{i} u B_{i,s,2}^c)\notag\\
&&+\sum_{l=3}^{4}\prod_{k=1}^{l-1} \{\exp ( \mathsf{i} u (B_{i,s,k-1}^c - B_{i,s,k}^c ) )-1 \}\exp ( \mathsf{i} u B_{i,s,l}^c )\notag\\
&&+\prod_{k=1}^{4} \{\exp ( \mathsf{i} u (B_{i,s,k-1}^c - B_{i,s,k}^c ) )-1 \}\exp ( \mathsf{i} u B_{i,s,4}^c ).
\end{eqnarray*}
It further yields the following expansion for the first derivative of $\phi^\ast_{i}(u)$:
\begin{eqnarray}
\frac{\mathrm{d}\phi^\ast_{i}(u)}{\mathrm{d}u}
&=& \mathsf{i}\sum_{s=1}^{\mathcal{T}_m+1} E^\ast [B_{i,s}\exp ( \mathsf{i} u B_{i,s,1}^c ) ]\notag \\
&&+\mathsf{i}\sum_{s=1}^{\mathcal{T}_m+1} E^\ast [B_{i,s} \{\exp ( \mathsf{i} u (\widetilde{\mathcal{P}}_{i}^* - B_{i,s,1}^c ) )-1 \}\exp ( \mathsf{i} u B_{i,s,2}^c ) ]\notag\\
&&+\mathsf{i}\sum_{s=1}^{\mathcal{T}_m+1}\sum_{l=3}^{4} E^\ast\Big[B_{i,s}\prod_{k=1}^{l-1} \{\exp ( \mathsf{i} u (B_{i,s,k-1}^c - B_{i,s,k}^c ) )-1 \}\exp ( \mathsf{i} u B_{i,s,l}^c )\Big] \notag\\
&&+\mathsf{i}\sum_{s=1}^{\mathcal{T}_m+1} E^\ast\Big[B_{i,s} \prod_{k=1}^{4} \{\exp ( \mathsf{i} u (B_{i,s,k-1}^c - B_{i,s,k}^c ) )-1 \}\exp ( \mathsf{i} u B_{i,s,4}^c )\Big]\notag\\
&\eqqcolon &\mathsf{J}_{i,1}(u)+\cdots+\mathsf{J}_{i,4}(u),
\end{eqnarray}
where $\mathsf{J}_{i,1}(u),\ldots,\mathsf{J}_{i,4}(u)$ are defined obviously.
For $\mathsf{J}_{i,1}(u)$, it is straightforward to show that $\mathsf{J}_{i,1}(u)=0$. For $\mathsf{J}_{i,2}(u)$,
\begin{eqnarray*}
\mathsf{J}_{i,2}(u)&=&-u\sum_{s=1}^{\mathcal{T}_m+1} E^\ast [B_{i,s} (\widetilde{\mathcal{P}}_{i}^* - B_{i,s,1}^c )\exp ( \mathsf{i} u B_{i,s,2}^c ) ]\notag \\
&&-\frac{1}{2}\mathsf{i}u^2\sum_{s=1}^{\mathcal{T}_m+1} E^\ast [B_{i,s} (\widetilde{\mathcal{P}}_{i}^* - B_{i,s,1}^c )^2\exp ( \mathsf{i} u B_{i,s,2}^c ) ] \notag\\
&&+\mathsf{i}\sum_{s=1}^{\mathcal{T}_m+1}E^\ast [B_{i,s} R^\ast_{i,s,1} (u)\exp ( \mathsf{i} u B_{i,s,2}^c ) ]\notag\\
&\eqqcolon &\mathsf{J}_{i,2,1}(u)+\mathsf{J}_{i,2,2}(u)+\mathsf{J}_{i,2,3}(u),
\end{eqnarray*}
where $R^\ast_{i,s,1}(u)=\sum_{l=3}^{\infty}\frac{\mathsf{(iu)}^l}{l!} (\widetilde{\mathcal{P}}_{i}^* - B_{i,s,1}^c )^l$ satisfies $|R^\ast_{i,s,1}(u)|\leq \frac{\mathsf{u}^3}{3!} |\widetilde{\mathcal{P}}_{i}^* - B_{i,s,1}^c |^3$, which is a direct application of Lemma (ref).
By adopting similar arguments to those in the proof of (ref) and involking Lemma (ref), for each $i$, we obtain
\begin{eqnarray*}
\big|\mathsf{J}_{i,2,1}(u)+u\phi^\ast_{i}(u)\big|=O_P\left(\frac{m}{T}\right),
\end{eqnarray*}
and $\Big|\mathsf{J}_{i,2,2}(u)+\frac{1}{2}\mathsf{i}u^2\sum_{s=1}^{\mathcal{T}_m+1} E^\ast [B_{i,s} (\widetilde{\mathcal{P}}_{i}^* - B_{i,s,1}^c )^2 ]\phi^\ast_{i}(u)\Big|=O_P\left(\frac{m}{T}\right)$ for any $u\in(-\infty,\infty)$.
In summary of these results, we can readily obtain
\begin{eqnarray}
|\mathsf{J}_{i,2}(u)-\mathsf{J}^\ast_{i,2}(u) |=O_P\left(\frac{m}{T}\right),
\end{eqnarray}
where $\mathsf{J}^\ast_{i,2}(u)=-u\phi^\ast_{i}(u)-\frac{1}{2}\mathsf{i}u^2\sum_{s=1}^{\mathcal{T}_m+1} E^\ast [B_{i,s} (\widetilde{\mathcal{P}}_{i}^* - B_{i,s,1}^c )^2 ]\phi^\ast_{i}(u)$.
Using arguments analogous to those for (ref), we can further show
\begin{eqnarray}
|\mathsf{J}_{i,3}(u)-\mathsf{J}^\ast_{i,3}(u) |=O_P\left(\frac{m}{T}\right)\quadand \quad |\mathsf{J}_{i,4}(u) |=o_P\left(\frac{m}{T}\right),
\end{eqnarray}
where
\begin{eqnarray*}
\mathsf{J}^\ast_{i,3}(u)&=&-\mathsf{i}u^2\sum_{s=1}^{\mathcal{T}_m+1}E^\ast [B_{i,s} (\widetilde{\mathcal{P}}_{i}^* - B_{i,s,1}^c ) (B_{i,s,1}^c- B_{i,s,2}^c ) ]\phi^\ast_{i}(u).
\end{eqnarray*}
By combining (ref), (ref), and (ref), we obtain
\begin{eqnarray}
\big|\frac{\mathrm{d}\phi^\ast_{i}(u)}{\mathrm{d}u}-\mathsf{J}^\ast_{i,2}(u)-\mathsf{J}^\ast_{i,3}(u)\big|=O_P\left(\frac{m}{T}\right),
\end{eqnarray}
where $\mathsf{J}^\ast_{i,2}(u)$ and $\mathsf{J}^\ast_{i,3}(u)$ satisfy
\begin{eqnarray*}
&&\mathsf{J}^\ast_{i,2}(u)+\mathsf{J}^\ast_{i,3}(u) = -u\phi^\ast_{i}(u)-\frac{1}{2}\mathsf{i}u^2\sum_{s=1}^{\mathcal{T}_m+1} E^\ast [B_{i,s} (\widetilde{\mathcal{P}}_{i}^* - B_{i,s,1}^c )^2 ]\phi^\ast_{i}(u)\notag\\
&&-\mathsf{i}u^2\sum_{s=1}^{\mathcal{T}_m+1}E^\ast [B_{i,s} (\widetilde{\mathcal{P}}_{i}^* - B_{i,s,1}^c ) (B_{i,s,1}^c- B_{i,s,2}^c ) ]\phi^\ast_{NT}(u) \notag\\
&=&-u\phi^\ast_{i}(u)-\frac{1}{2}\mathsf{i}u^2 \sum_{s=1}^{\mathcal{T}_m+1} E^\ast [B_{i,s} ( (\widetilde{\mathcal{P}}_{i}^* - B_{i,s,1}^c )^2+2 (\widetilde{\mathcal{P}}_{i}^* - B_{i,s,1}^c ) (B_{i,s,1}^c- B_{i,s,2}^c ) ) ]\phi^\ast_{i}(u)
\notag\\
&=&-u\phi^\ast_{i}(u)-\frac{1}{2}\mathsf{i}u^2 E^\ast [\widetilde{\mathcal{P}}_{i}^{*3} ]\phi^\ast_{i}(u).
\end{eqnarray*}
Finally, in connection with Taylor expansion of $\exp (-\frac{1}{6}\mathsf{i}u^3 E^\ast[\widetilde{\mathcal{S}}_{i}^{*3} ] )$ and Lemma (ref), integration of (ref) leads to the following result for the characteristic function $\phi^\ast_{i}(u)$:
\begin{eqnarray*}
\Big|\phi^\ast_{i}(u)-\exp\big(-\frac{1}{2}u^2\big) \big\{1-\frac{1}{6}\mathsf{i}u^3 E^\ast [\widetilde{\mathcal{P}}_{i}^{*3} ] \big\}\Big| =O_P\left(\frac{m}{T}\right).
\end{eqnarray*}
Using Esseen smoothing inequality, this condition is sufficient to establish Theorem (ref).1.
(2). By Theorem (ref) and Theorem (ref).1, it follows for each $i$ that
\begin{eqnarray*}
&&\sup_{u\in \mathbb{R}}\left|\normalfont Pr^*(\widetilde{p}_{i}^* \le u) -\Pr(\widetilde{p}_{i}\le u)\right|\notag \\
&\leq&\sup_{u\in \mathbb{R}}\left|\normalfont Pr^*(\widetilde{p}_{i}^* \le u) -\Phi(u)\right|
+\sup_{u\in \mathbb{R}}\left|\Pr(\widetilde{p}_{i}\le u)-\Phi(u)\right| = O_P\left(\sqrt{\frac{m}{T}}\right).
\end{eqnarray*}
It completes the proof of Theorem (ref).2.
(3) Similar arguments used in the proof of Theorem (ref).3 can be directly applied here to establish the desired result in Theorem (ref).3. Consequently, the details are omitted for brevity.
proof[Proof of Theorem (ref)]
• By using the BN decomposition in Lemma (ref), we have
\begin{eqnarray*}
\frac{1}{\sqrt{T}}\sum_{t=1}^T(\mathbf{x}_{t} -\bm{\mu}) &=& \frac{1}{\sqrt{T}} \sum_{t=1}^T [\mathbf{B} -(1-L)\widetilde{\mathbf{B}}(L)] \pmb{\varepsilon}_{t} \notag \\
&=&\frac{1}{\sqrt{T}}\sum_{t=1}^T\mathbf{B} \pmb{\varepsilon}_{t} - \frac{1}{\sqrt{T}}\sum_{t=1}^T\widetilde{\mathbf{B}}(L)\pmb{\varepsilon}_{t}+ \frac{1}{\sqrt{T}}\sum_{t=1}^T \widetilde{\mathbf{B}}(L)\pmb{\varepsilon}_{t-1}\notag \\
&=&\frac{1}{\sqrt{T}}\sum_{t=1}^T\mathbf{B} \pmb{\varepsilon}_{t}-\frac{1}{\sqrt{T}}\widetilde{\mathbf{B}}(L)\pmb{\varepsilon}_{T}+\frac{1}{\sqrt{T}}\widetilde{\mathbf{B}}(L)\pmb{\varepsilon}_{0}.
\end{eqnarray*}
Note that for any $\epsilon>0$, we have
\begin{eqnarray*}
&&\Pr\left(\left|\frac{1}{\sqrt{T}}\sum_{t=1}^{T}(\mathbf{x}_{t}-\bm{\mu})\right|_{\infty} \leq u \right) \notag \\
&\leq& \Pr\left(\left|\frac{1}{\sqrt{T}}\widetilde{\mathbf{B}}(L)\pmb{\varepsilon}_{T}-\frac{1}{\sqrt{T}}\widetilde{\mathbf{B}}(L)\pmb{\varepsilon}_{0}\right|_{\infty}\geq \epsilon \right)+\Pr\left(\left|\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\mathbf{B} \pmb{\varepsilon}_{t}\right|_{\infty} \leq u + \epsilon\right)
\end{eqnarray*}
and
\begin{eqnarray*}
&&\Pr\left(\left|\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\mathbf{z} _t\right|_{\infty} \leq u\right) \notag \\
&=&\Pr\left(\left|\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\mathbf{z} _t\right|_{\infty} \leq u+\epsilon\right)-\Pr\left(u<\left|\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\mathbf{z} _t\right|_{\infty} \leq u + \epsilon\right).
\end{eqnarray*}
Hence, we have
\begin{eqnarray*}
&&\sup_{u\in \mathbb{R}}\left|\Pr\left(\left|\frac{1}{\sqrt{T}}\sum_{t=1}^{T}(\mathbf{x} _t-\bm{\mu})\right|_{\infty} \leq u\right) - \Pr\left(\left|\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\mathbf{z} _t\right|_{\infty} \leq u \right)\right| \notag \\
&\leq& \Pr\left(\left|\frac{1}{\sqrt{T}}\widetilde{\mathbf{B}}(L)\pmb{\varepsilon}_{T}-\frac{1}{\sqrt{T}}\widetilde{\mathbf{B}}(L)\pmb{\varepsilon}_{0}\right|_{\infty}\geq \epsilon \right) \notag \\
&&+ \sup_{u\in \mathbb{R}}\left|\Pr\left(\left|\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\mathbf{B} \pmb{\varepsilon}_{t}\right|_{\infty} \leq u\right) - \Pr\left(\left|\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\mathbf{z} _t\right|_{\infty} \leq u \right)\right| \notag\\
&& + \sup_{u\in \mathbb{R}}\left| \Pr\left(\left|\left|\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\mathbf{z} _t\right|_{\infty}-u\right| \leq \epsilon \right)\right| \eqqcolon I_1+I_2+I_3.
\end{eqnarray*}
Consider $I_1$ first. Given $E[\varepsilon_{it}^J] < \infty$ and $\max_{i}\sum_{\ell=1}^{\infty}\ell \|\mathbf{b}_{\ell i}^{\sharp}\|_2<\infty$, we next show that each element in $\widetilde{\mathbf{B}}(L)\pmb{\varepsilon}_{t}$ has the finite $J^{th}$ moment to imply
\[
E [ |\widetilde{\mathbf{B}}(L)\pmb{\varepsilon}_{t} |_\infty^J ] = O(N).
\]
Let $\widetilde{\mathbf{B}}_{\ell} = (\widetilde{b}_{\ell,ij} )_{1\leq i,j\leq N}$. Then, for the $i^{th}$ element in $\widetilde{\mathbf{B}}(L)\pmb{\varepsilon}_{t}$, we have
\begin{eqnarray*}
&&\left(E(\sum_{\ell=0}^{\infty}\sum_{j=1}^{N}\widetilde{b}_{\ell,ij}\varepsilon_{j,t-\ell})^{J}\right)^{1/J} \leq \sum_{\ell=0}^{\infty}\left(E(\sum_{j=1}^{N}\widetilde{b}_{\ell,ij}\varepsilon_{j,t-\ell})^{J}\right)^{1/J} \notag\\
&\leq& \sum_{\ell=0}^{\infty}\frac{14.5J}{\log J}\left[\left(\sum_{j=1}^{N}E(\widetilde{b}_{\ell,ij}\varepsilon_{j,t-\ell})^J \right)^{1/J} + \left(\sum_{j=1}^{N}E(\widetilde{b}_{\ell,ij}\varepsilon_{j,t-\ell})^2 \right)^{1/2} \right] \notag\\
&\leq& \frac{29J(E[\varepsilon_{it}^J])^{1/J}}{\log J} \sum_{\ell=0}^{\infty}\ell\|\mathbf{b}_{\ell i}^{\sharp}\|_2 =O(1),
\end{eqnarray*}
where the first inequality follows from the triangle inequality, and the second inequality follows from the Rosenthal inequality for independent variables (e.g., Johnson).
Then choose $\epsilon = \sqrt{\frac{N^{2/J}}{T^{1-2/J}}}$ and by using the Markov inequality, we have
$$
I_1 = O\left(\frac{N/T^{J/2}}{N/T^{J/2-1}} \right) =O(T^{-1}).
$$
Consider $I_2$. Similarly, we can show that each element in $\mathbf{B}\pmb{\varepsilon}_{t}$ has bounded the $J^{th}$ moment and thus $E[|\mathbf{B}\pmb{\varepsilon}_{t}|_\infty^J] = O(N)$. Then by using high-dimensional Gaussian approximations for independent random vectors (cf., Theorem 2.5 of chernozhuokov2022improved), we have
\begin{eqnarray*}
&&\sup_{u\in \mathbb{R}}\left|\Pr\left(\left|\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\mathbf{B}\pmb{\varepsilon}_{t}\right|_{\infty} \leq u\right) - \Pr\left(\left|\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\mathbf{z} _t\right|_{\infty} \leq u\right)\right| \notag \\
&=&O\left( \left(\frac{ N^{2/J}(\log N)^5}{T}\right)^{1/4}+\sqrt{\frac{N^{2/J}(\log N)^{3-2/J}}{T^{1-2/J}}}\right).
\end{eqnarray*}
Consider $I_3$. Using Lemma A.1 in chernozhukov2017central, we have
$$
\sup_{u\in \mathbb{R}}\left| \Pr\left(\left|\left|\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\mathbf{z} _t\right|_{\infty}-u\right| \leq \epsilon \right)\right| = O(\epsilon \sqrt{\log N}) = o \left(\sqrt{\frac{N^{2/J}(\log N)^{3-2/J}}{T^{1-2/J}}}\right)
$$
when choosing $\epsilon = \sqrt{\frac{N^{2/J}}{T^{1-2/J}}}$. The proof is now completed.
proof[Proof of Theorem (ref)]
• To complete this theorem, it is sufficient to prove
\[
|\widehat{\bm{\Omega}}-\bm{\Omega}|_{\max} = O_P(\sqrt{\widetilde{m}\log N/T}).
\]
Then Theorem (ref) follows from Theorem (ref) and Proposition 2.1 in chernozhuokov2022improved.
Define $\mathbf{y}_t = (y_{1t},\ldots,y_{Nt})^\top = \sum_{\ell=0}^{\infty}\mathbf{B}_{\ell}\bm{\varepsilon}_{t-\ell}$. Let $y_{it}^*$ be the coupled version of $y_{it}$ with $\bm{\varepsilon}_0^*$ replacing $\bm{\varepsilon}_0$ and $\mathbf{B}_\ell = \left(b_{\ell,ij} \right)_{1\leq i,j\leq N}$. By using Rosenthal inequality for independent variables, we have
\begin{eqnarray*}
&&\left(E(y_{it}y_{jt}-y_{it}^*y_{jt}^*)^J\right)^{1/J} \leq O(1)\max_{i} \left(E(y_{it}-y_{it}^*)^J\right)^{1/J}\notag\\
&\leq&O(1)\max_i\left(E(\sum_{j=1}^{N}b_{t,ij}\varepsilon_{j,0})^{J}\right)^{1/J} \notag\\
&\leq& O(1)\max_i\left[\left(\sum_{j=1}^{N}E(b_{t,ij}\varepsilon_{j0})^J \right)^{1/J} + \left(\sum_{j=1}^{N}E(b_{t,ij}\varepsilon_{j0})^2 \right)^{1/2} \right] \notag\\
&\leq& O(1)\max_i\|\mathbf{b}_{t i}^{\sharp}\|_2.
\end{eqnarray*}
Then by using Lemma A.8 (1) of gao2024robust, if $\frac{N^2T\log T}{(T\widetilde{m}\log N)^{J/4}}\to 0$, we have
$$
\max_{1\leq i,j\leq N}\left|\frac{1}{T}\sum_{t,s=1}^{T}a\left(\frac{t-s}{\widetilde{m}}\right)(y_{it}y_{js}-E(y_{it}y_{js})) \right| = O_P\left(\sqrt{\widetilde{m}\log N/T}\right).
$$
Let $\bm{\Omega} = \{\omega_{ij}\}_{1\leq i,j\leq N}$. By using standard arguments for the bias term (e.g., the proof of Theorem 2.2 in gao2024robust), we have
\[
\left|\frac{1}{T}\sum_{t,s=1}^{T}a\left(\frac{t-s}{\widetilde{m}}\right)E(y_{it}y_{js}) - \omega_{ij} \right| = O(\widetilde{m}^{-q_{\alpha}}).
\]
Then, to complete the proof, it is sufficient to show
\begin{eqnarray*}
\max_{1\leq i,j\leq N}\left|\frac{1}{T}\sum_{t,s=1}^{T}a\left(\frac{t-s}{\widetilde{m}}\right)((x_{it}-\mu_i)(x_{js}-\mu_j)-(x_{it}-\overline{x}_i)(x_{js}-\overline{x}_j)) \right| = O_P\left(\sqrt{\widetilde{m}\log N/T}\right).
\end{eqnarray*}
By Lemma 1 (1) in gao2024robust we have
\begin{eqnarray*}
&&\max_{1\leq i,j\leq N}\left|\frac{1}{T}\sum_{t,s=1}^{T}a\left(\frac{t-s}{\widetilde{m}}\right)((x_{it}-\mu_i)(x_{js}-\mu_j)-(x_{it}-\overline{x}_i)(x_{js}-\overline{x}_j))\right|\notag\\
&\leq&2\max_{1\leq j\leq N}\left|\frac{1}{T}\sum_{t,s=1}^{T}a\left(\frac{t-s}{\widetilde{m}}\right)y_{js}\right|\max_{1\leq i\leq N}\left|\mu_i-\overline{x}_i\right| + \max_{1\leq i\leq N}(\mu_i-\overline{x}_i)^2\left|\frac{1}{T}\sum_{t,s=1}^{T}a\left(\frac{t-s}{\widetilde{m}}\right) \right| \notag\\
&=& O_P\left(\widetilde{m}\log N/T\right). \notag
\end{eqnarray*}
The proof is now completed.
proof[Proof of Proposition (ref)]
• (1). Using BN decomposition, we obtain that
\begin{eqnarray*}
\max_i|\overline{x}_i -\mu_i|&=& \frac{1}{T}\left|\mathbf{e}_i^\top \mathbf{B} \sum_{s=1}^T\pmb{\varepsilon}_{s}-\mathbf{e}_i^\top\widetilde{\mathbf{B}}(L)\pmb{\varepsilon}_{T}+\mathbf{e}_i^\top\widetilde{\mathbf{B}}(L)\pmb{\varepsilon}_{0} \right|.
\end{eqnarray*}
By Corollary 2.1 of JSW_2023, we can obtain that
\begin{eqnarray*}
\Pr\left(\max_i\left|\frac{1}{\sqrt{\sum_{s=1}^T E[\zeta_{is}^2]}} \sum_{s=1}^T\zeta_{is}\right| \ge \epsilon\right) &\le &\sum_{i=1}^N \Pr\left(\left|\frac{1}{\sqrt{\sum_{s=1}^T E[\zeta_{is}^2]}} \sum_{s=1}^T\zeta_{is}\right|\ge \epsilon\right)\notag \\
&\le & 2N \exp(-\epsilon^2/4)= 2N \exp(\log(1/N^c))\notag \\
&=&2N/N^c\to 0,
\end{eqnarray*}
where $\zeta_{is}\coloneqq \mathbf{e}_i^\top \mathbf{B} \pmb{\varepsilon}_{s}$, and $\epsilon=2\sqrt{\log(N^c)}$ with $c>1$. After neglecting the residuals from the BN decomposition and using Assumption (ref).1, the first result is proved.
(2). Let $\overline{\pmb{\mu}} =(\mu_1,\ldots, \mu_{J_0})$. Since $J_0$ is known, we immediately obtain that
\begin{eqnarray*}
0\le S(\widehat{\mathcal{G}}, \widehat{\pmb{\nu}})\le S(\mathscr{G}, \overline{\pmb{\mu}})= O_P( \log(N)/T +c_{N}^2),
\end{eqnarray*}
where the last step is due to Assumption (ref).1 and the first result of this lemma. If one individual is allocated in a wrong group, then we can always show that $S(\mathscr{G}, \overline{\pmb{\mu}})$ is a better option. Therefore, we conclude that $\Pr(\widehat{\mathcal{G}}_j=\mathscr{G}_j)\to 1$ for all $j\in [J_0]$.
(3). Without loss of generality, we consider two cases: (i) $J=J_0-1$ and (ii) $J=J_0+1$. For case (i) (i.e.,$J= J_0-1$), we write
\begin{eqnarray*}
&&S(\widehat{\mathcal{G}}_{\mid J}, \widehat{\pmb{\nu}}_{\mid J})+\rho_{NT}J \notag \\
&= &\frac{1}{N}\sum_{j=1}^J\sum_{i\in \widehat{\mathcal{G}}_{j\mid J}}| \overline{x}_i- \sum_{j}\overline{\mu}_j I(i\in \mathscr{G}_j)+\sum_{j}\overline{\mu}_j I(i\in \mathscr{G}_j) - \widehat{\nu}_{j\mid J}|^2+\rho_{NT}J \notag \\
&\ge & \frac{1}{N}\sum_{j=1}^J\sum_{i\in \widehat{\mathcal{G}}_{j\mid J}}| \sum_{j}\overline{\mu}_j I(i\in \mathscr{G}_j) - \widehat{\nu}_{j\mid J}|^2 -o_P (1)\notag \\
&\ge &c^*-o_P (1)\notag \\
&> &S(\mathscr{G}, \overline{\pmb{\mu}})+\rho_{NT}J_0,
\end{eqnarray*}
where the first inequality follows from the first result of this lemma, and we must be able to find a $c^*$ using Assumption (ref).2 and $J=J_0-1$.
For case (ii) (i.e.,$J= J_0+1$), write
\begin{eqnarray*}
&&S(\widehat{\mathcal{G}}_{\mid J}, \widehat{\pmb{\nu}}_{\mid J})+\rho_{NT}(J_0+1)\notag \\
&=& \frac{1}{N}\sum_{j=1}^J\sum_{i\in \widehat{\mathcal{G}}_{j\mid J}}| \sum_{j}\overline{\mu}_j I(i\in \mathscr{G}_j)-\widehat{\nu}_{j\mid J} |^2+\rho_{NT} (J_0+1)+O_P(\sqrt{\log(N)/T} +c_{N})\notag \\
&> &S(\mathscr{G}, \overline{\pmb{\mu}})+\rho_{NT}J_0,
\end{eqnarray*}
where the last step follows in view of the fact that $\rho_{NT}/(\sqrt{\log(N)/T} +c_{N})\to \infty$.
The proof is now completed.
proof[Proof of Proposition (ref)]
• Write
\begin{eqnarray*}
&&\frac{1}{NT^2}\sum_{t=1}^T \mathbf{y}_t ^{\top} \mathbf{y}_t \notag \\
&=&\frac{1}{NT^2}\sum_{t=1}^T \left( \mathbf{B} \sum_{s=1}^t\pmb{\varepsilon}_{s} -\widetilde{\mathbf{B}} (L)\pmb{\varepsilon}_{t} +\widetilde{\mathbf{B}} (L)\pmb{\varepsilon}_{0} \right)^\top \left( \mathbf{B} \sum_{s=1}^t\pmb{\varepsilon}_{s} -\widetilde{\mathbf{B}} (L)\pmb{\varepsilon}_{t} +\widetilde{\mathbf{B}} (L)\pmb{\varepsilon}_{0} \right) ,
\end{eqnarray*}
where the leading term is obviously $\frac{1}{NT^2}\sum_{t=1}^T \sum_{s_1=1}^t \sum_{s_2=1}^t \pmb{\varepsilon}_{s_1} ^{\top}\mathbf{B} ^{\top}\mathbf{B} \pmb{\varepsilon}_{s_2} $.
We now focus on the leading term $\frac{1}{NT^2}\sum_{t=1}^T\sum_{s_1=1}^t\sum_{s_2=1}^t\pmb{\varepsilon}_{s_1} ^{\top}\mathbf{B} ^{\top}\mathbf{B} \pmb{\varepsilon}_{s_2} $ in what follows. Firstly, write
\begin{eqnarray*}
&&\frac{1}{NT^2}\sum_{t=1}^TE \left[\sum_{s_1=1}^t\pmb{\varepsilon}_{s_1} ^{\top}\mathbf{B} ^{\top}\mathbf{B} \sum_{s_2=1}^t\pmb{\varepsilon}_{s_2} \right] = \frac{1}{NT^2}\sum_{t=1}^T\sum_{s=1}^tE \left[\pmb{\varepsilon}_{s} ^{\top}\mathbf{B} ^{\top}\mathbf{B} \pmb{\varepsilon}_{s} \right]\notag \\
&=&\frac{1}{NT^2}\sum_{t=1}^T\sum_{s=1}^t vec(\mathbf{B} ^{\top}\mathbf{B} )^\top E [\pmb{\varepsilon}_{s} \otimes\pmb{\varepsilon}_{s} ] = \frac{1}{NT^2}\sum_{t=1}^T\sum_{s=1}^t vec(\mathbf{B} ^{\top}\mathbf{B} )^\top (\mathbf{e}_1^\top,\ldots, \mathbf{e}_N^\top)^\top \notag \\
&= &\frac{1}{NT^2}\sum_{t=1}^T\sum_{s=1}^t \|\mathbf{B} \|^2 = \left(\int_0^1x \mathrm{d}x+ O\left(\frac{1}{T}\right) \right) \frac{1}{N}\|\mathbf{B} \|^2 \to \frac{b }{2},
\end{eqnarray*}
where the second equality follows from the vectorization operation, the fifth equality follows from the definition of Riemann integral, and the last step follows from the condition $\lim_{N}\frac{1}{N}\|\mathbf{B} \|^2\to b$.
We further note a few facts:
\begin{enumerate}[leftmargin=24pt, parsep=2pt, topsep=2pt]
• $E\big[(\pmb{\varepsilon}_{s_1} ^{\top}\mathbf{B} ^{\top}\mathbf{B} \pmb{\varepsilon}_{s_2} - E[\pmb{\varepsilon}_{s_1} ^{\top}\mathbf{B} ^{\top}\mathbf{B} \pmb{\varepsilon}_{s_2} ]) (\pmb{\varepsilon}_{s_3} ^{\top}\mathbf{B} ^{\top}\mathbf{B} \pmb{\varepsilon}_{s_4} - E[\pmb{\varepsilon}_{s_3} ^{\top}\mathbf{B} ^{\top}\mathbf{B} \pmb{\varepsilon}_{s_4} ])\big]=0$ when
\begin{enumerate}[leftmargin=24pt, parsep=2pt, topsep=2pt]
• three or four of $s_1, s_2, s_3, s_4$ are mutually different;
• $s_1=s_2$, and $s_3\ne s_1$ (or $s_4\ne s_1$);
• $s_1\ne s_2$, and $s_4=s_3=s_1$.
\end{enumerate}
• for $s_1\ne s_2$, $E[\pmb{\varepsilon}_{s_1} ^{\top}\mathbf{B} ^{\top}\mathbf{B} \pmb{\varepsilon}_{s_2} ]^2 =E[\pmb{\varepsilon}_{s_1} ^{\top} \mathbf{B} ^{\top}\mathbf{B} \mathbf{B} ^{\top}\mathbf{B} \pmb{\varepsilon}_{s_1}^{u }] \le \|\mathbf{B} \|_2^4 N =O(N)$.
• Note that
\begin{eqnarray*}
&&E|\pmb{\varepsilon}_{s} ^{\top}\mathbf{B} ^{\top}\mathbf{B} \pmb{\varepsilon}_{s} |^2 = vec(\mathbf{B} ^{\top}\mathbf{B} )^\top E [(\pmb{\varepsilon}_{s} \otimes\pmb{\varepsilon}_{s}) (\pmb{\varepsilon}_{s}^{\top}\otimes\pmb{\varepsilon}_{s}^{\top})] vec(\mathbf{B} ^{\top}\mathbf{B} )\notag \\
&\le & O(N)vec(\mathbf{B} ^{\top}\mathbf{B} )^\top vec(\mathbf{B} ^{\top}\mathbf{B}) = O(N)\operatorname*{\normalfont\textrm{trace}}(\mathbf{B} ^{\top}\mathbf{B} \mathbf{B} ^{\top}\mathbf{B} )\notag \\
&\le &O(N) \text{vec}(\mathbf{B} ^{\top})^\top (\mathbf{B} \otimes \mathbf{B} ^{\top}) \text{vec}(\mathbf{B} ^{\top}) \le O(N)\|\text{vec}(\mathbf{B} ^{\top})\|^2 = O(N^2),
\end{eqnarray*}
where the first inequality follows from Lemma (ref), the second equality follows from the fact that $\operatorname*{\normalfont\textrm{trace}}(\mathbf{C}^\top\mathbf{D}) = \text{vec}(\mathbf{C})^\top \text{vec}(\mathbf{D})$ for $\forall \mathbf{C},\mathbf{D}\in \mathbb{R}^{N\times N}$, the second inequality follows from the fact that $\operatorname*{\normalfont\textrm{trace}}(\mathbf{A}_1\mathbf{A}_2\mathbf{A}_3\mathbf{A}_4) =\text{vec}(\mathbf{A}_1)^\top(\mathbf{A}_2\otimes \mathbf{A}_4^\top )\text{vec}(\mathbf{A}_3^\top)$ for any conformable matrices $\mathbf{A}_1,\mathbf{A}_2,\mathbf{A}_3,\mathbf{A}_4$ (Bernstein), and the last step follows from the condition $\lim_{N}\frac{1}{N}\|\mathbf{B} \|^2\to b$.
\end{enumerate}
We then write
\begin{eqnarray*}
&&E\left[\frac{1}{NT^2}\sum_{t=1}^T\sum_{s_1=1}^t\sum_{s_2=1}^t(\pmb{\varepsilon}_{s_1} ^{\top}\mathbf{B} ^{\top}\mathbf{B} \pmb{\varepsilon}_{s_2} - E[\pmb{\varepsilon}_{s_1} ^{\top}\mathbf{B} ^{\top}\mathbf{B} \pmb{\varepsilon}_{s_2} ])\right]^2\notag \\
&=&\frac{1}{N^2T^4}\sum_{t_1=1}^T\sum_{t_2=1}^T\sum_{s_1=1}^{t_1}\sum_{s_2=1}^{t_1}\sum_{s_3=1}^{t_2}\sum_{s_4=1}^{t_2}E\big[(\pmb{\varepsilon}_{s_1} ^{\top}\mathbf{B} ^{\top}\mathbf{B} \pmb{\varepsilon}_{s_2} - E[\pmb{\varepsilon}_{s_1} ^{\top}\mathbf{B} ^{\top}\mathbf{B} \pmb{\varepsilon}_{s_2} ]) \notag \\
&&\cdot (\pmb{\varepsilon}_{s_3} ^{\top}\mathbf{B} ^{\top}\mathbf{B} \pmb{\varepsilon}_{s_4} - E[\pmb{\varepsilon}_{s_3} ^{\top}\mathbf{B} ^{\top}\mathbf{B} \pmb{\varepsilon}_{s_4} ])\big] \notag \\
&=&\frac{1}{N^2T^4}\sum_{t_1=1}^T\sum_{t_2=1}^T\sum_{s_1=1}^{t_1} E|\pmb{\varepsilon}_{s_1} ^{\top}\mathbf{B} ^{\top}\mathbf{B} \pmb{\varepsilon}_{s_1} - E[\pmb{\varepsilon}_{s_1} ^{\top}\mathbf{B} ^{\top}\mathbf{B} \pmb{\varepsilon}_{s_1} ]|^2 \notag \\
&&+\frac{1}{N^2T^4}\sum_{t_1=1}^T\sum_{t_2=1}^T\sum_{ s_1,s_2=1, s_1\ne s_2}^{t_1} E|\pmb{\varepsilon}_{s_1} ^{\top}\mathbf{B} ^{\top}\mathbf{B} \pmb{\varepsilon}_{s_2} - E[\pmb{\varepsilon}_{s_1} ^{\top}\mathbf{B} ^{\top}\mathbf{B} \pmb{\varepsilon}_{s_2} ]|^2\notag \\
&\le &O(1)\frac{1}{N^2T}(E|\pmb{\varepsilon}_{s_1} ^{\top}\mathbf{B} ^{\top}\mathbf{B} \pmb{\varepsilon}_{s_1} |^2+E|\pmb{\varepsilon}_{s_1} ^{\top}\mathbf{B} ^{\top}\mathbf{B} \pmb{\varepsilon}_{s_2} |^2) = O(1)\frac{1}{T},
\end{eqnarray*}
where the last step follows from the above facts.
Putting everything together, we immediately obtain that $\frac{1}{NT^2}\sum_{t=1}^T \mathbf{y}_t ^{\top} \mathbf{y}_t \to_P\frac{b }{2}.$
proof[Proof of Proposition (ref)]
• In what follows, we focus on $\widehat{\pmb{\theta}}_i$, and write
\begin{eqnarray*}
\widehat{\pmb{\theta}}_i-\pmb{\theta}_i &=& (\mathbf{W}_i^\top \mathbf{M}_{\overline{\mathbf{Z}}} \mathbf{W}_i)^{-1} \mathbf{W}_i^\top \mathbf{M}_{\overline{\mathbf{Z}}}\mathbf{F}\pmb{\gamma}_i + (\mathbf{W}_i^\top \mathbf{M}_{\overline{\mathbf{Z}}} \mathbf{W}_i)^{-1} \mathbf{W}_i^\top \mathbf{M}_{\overline{\mathbf{Z}}} \pmb{\epsilon}_i.
\end{eqnarray*}
We study $\mathbf{M}_{\overline{\mathbf{Z}}}$ first.
\begin{eqnarray*}
\mathbf{M}_{\overline{\mathbf{Z}}} &=&\mathbf{I}_T-(\mathbf{F}\overline{\mathbf{C}}+\overline{\mathbf{U}})[(\mathbf{F}\overline{\mathbf{C}}+\overline{\mathbf{U}})^\top (\mathbf{F}\overline{\mathbf{C}}+\overline{\mathbf{U}})]^+ (\mathbf{F}\overline{\mathbf{C}}+\overline{\mathbf{U}})^\top.
\end{eqnarray*}
We now investigate $(\mathbf{F}\overline{\mathbf{C}}+\overline{\mathbf{U}})^\top (\mathbf{F}\overline{\mathbf{C}}+\overline{\mathbf{U}})$, and write
\begin{eqnarray*}
(\mathbf{F}\overline{\mathbf{C}}+\overline{\mathbf{U}})^\top (\mathbf{F}\overline{\mathbf{C}}+\overline{\mathbf{U}})=\overline{\mathbf{C}}^\top \mathbf{F}^\top \mathbf{F} \overline{\mathbf{C}}+\overline{\mathbf{U}}^\top \overline{\mathbf{U}} + \overline{\mathbf{C}}^\top \mathbf{F}^\top \overline{\mathbf{U}}+ \overline{\mathbf{U}}^\top \mathbf{F} \overline{\mathbf{C}}.
\end{eqnarray*}
By Assumption (ref).1,
\begin{eqnarray*}
\frac{1}{T} (\mathbf{F}\overline{\mathbf{C}}+\overline{\mathbf{U}})^\top (\mathbf{F}\overline{\mathbf{C}}+\overline{\mathbf{U}}) = \frac{1}{T}\overline{\mathbf{C}}^\top \mathbf{F}^\top \mathbf{F} \overline{\mathbf{C}}+O_P\left(\frac{1}{N}+\frac{1}{\sqrt{NT}}\right).
\end{eqnarray*}
Also, note that by Assumptions (ref).1
\begin{eqnarray*}
\left( \overline{\mathbf{C}}^\top \frac{\mathbf{F}^\top \mathbf{F} }{T}\overline{\mathbf{C}}\right)^+ = \overline{\mathbf{C}}^+ \left(\frac{\mathbf{F}^\top \mathbf{F} }{T}\right)^+ (\overline{\mathbf{C}}^\top)^+ ,
\end{eqnarray*}
where $\overline{\mathbf{C}}^+ =\overline{\mathbf{C}}^\top(\overline{\mathbf{C}}\overline{\mathbf{C}}^\top)^{-1}$ and $(\overline{\mathbf{C}}^\top)^+ =(\overline{\mathbf{C}}\overline{\mathbf{C}}^\top)^{-1}\overline{\mathbf{C}}$.
Thus, similar to (40) of Pesaran2006, it is straightforward to obtain that
\begin{eqnarray*}
\frac{1}{T} \mathbf{W}_i^\top \mathbf{M}_{\overline{\mathbf{Z}}}\mathbf{F}\pmb{\gamma}_i =O_P\left(\frac{1}{N}+\frac{1}{\sqrt{NT}}\right).
\end{eqnarray*}
We then focus on $\frac{1}{T}\mathbf{W}_i^\top \mathbf{M}_{\overline{\mathbf{Z}}} \pmb{\epsilon}_i$, and write
\begin{eqnarray*}
\frac{1}{T}\mathbf{W}_i^\top \mathbf{M}_{\overline{\mathbf{Z}}} \pmb{\epsilon}_i &=&\frac{1}{T}\mathbf{W}_i^\top \mathbf{M}_{\mathbf{F}} \pmb{\epsilon}_i + O_P\left(\frac{1}{N}+\frac{1}{\sqrt{NT}}\right)\notag \\
&=&\frac{1}{T}\mathbf{V}_i^\top \pmb{\epsilon}_i + O_P\left(\frac{1}{N}+\frac{1}{T}+\frac{1}{\sqrt{NT}}\right),
\end{eqnarray*}
where the first equality follows from (45) of Pesaran2006, and the second equality follows from Assumption (ref).1.
Note further that Assumption (ref).2 ensures that $\mathbf{v}_{it}\epsilon_{it}$ admits an MA($\infty$) process via the second order BN decomposition. See (ref) and PS1992 for example. Thus, by the proofs of Lemma (ref), it is easy to know that $\frac{1}{T}\mathbf{V}_i^\top \pmb{\epsilon}_i$ can be further decomposed via the second order BN decomposition. Also, we note that $\widehat{\pmb{\theta}}-\pmb{\theta} =O_P(\frac{1}{\sqrt{NT}})$ under the null and the condition $N\asymp T$ (see Westerlund2018 for detailed development). Finally, putting everything together, we invoke Theorem (ref) and Assumption (ref).2, and the result follows immediately.
Proofs of the Preliminary Lemmas
proof[Proof of Lemma (ref)]
• First, we note that the Probabilist's Hermite polynomials
\[
\left\{H_n(x)\mid (-1)^n \exp\left(\frac{x^2}{2}\right) \frac{\mathrm{d}^n}{\mathrm{d}x^n}\widetilde{\phi}(x)\ \text{ for }\ n \ge 0\right\}
\]
have the following generating function:
\[
\exp\left( xu-\frac{u^2}{2}\right) = \sum_{n=0}^\infty H_n(x)\frac{u^n}{n!}.
\]
Thus,
\begin{eqnarray}
\sum_{n=0}^\infty\phi(x) H_n(x)\frac{u^n}{n!}&=&\frac{1}{\sqrt{2\pi}}\exp\left(-\frac{x^2}{2}+xu-\frac{u^2}{2}\right) \notag \\
&=& \frac{1}{\sqrt{2\pi}}\exp\left(-\frac{(x-u)^2}{2}\right) .
\end{eqnarray}
We now apply the Fourier transformation to both sides of (ref), and note that the right hand side can be further written as follows:
\begin{eqnarray*}
&& \int_{\mathbb{R}}\exp(\mathsf{i}w)\cdot\frac{1}{\sqrt{2\pi}}\exp\left(-\frac{(x-u)^2}{2}\right) \mathrm{d}x = \exp\left(\mathsf{i}w u-\frac{1}{2}w^2\right)\notag \\
&& = \exp(\mathsf{i}w u)\widetilde{\phi}(w) = \sum_{n=0}^\infty\widetilde{\phi}(w) \frac{(\mathsf{i}wu)^n}{n!} = \sum_{n=0}^\infty\widetilde{\phi}(w)(\mathsf{i}w)^n\cdot \frac{u^n}{n!},
\end{eqnarray*}
where the first equality is obvious in view of the characteristic function of the normal distribution, and the third equality follows from (ref).
By comparing the Fourier transformation of both sides of (ref), the result follows immediately.
proof[Proof of Lemma (ref)]
• (1).a The expression $\mathbf{B}(L) = \mathbf{B} -(1-L)\widetilde{\mathbf{B}}(L)$ of (ref) is the so-called BN decomposition from PS1992. We then write
\begin{eqnarray*}
\sum_{s=1}^t \mathbf{x}_{s} &=& \sum_{s=1}^t [ \mathbf{B} -(1-L)\widetilde{\mathbf{B}}(L)] \pmb{\varepsilon}_{s} \notag \\
&=&\mathbf{B} \sum_{s=1}^t\pmb{\varepsilon}_{s} - \sum_{s=1}^t\widetilde{\mathbf{B}}(L)\pmb{\varepsilon}_{s}+ \sum_{s=1}^t \widetilde{\mathbf{B}}(L)\pmb{\varepsilon}_{s-1}\notag \\
&=&\mathbf{B} \sum_{s=1}^t\pmb{\varepsilon}_{s}-\widetilde{\mathbf{B}}(L)\pmb{\varepsilon}_{t}+\widetilde{\mathbf{B}}(L)\pmb{\varepsilon}_{0}.
\end{eqnarray*}
By Assumption (ref).2,
\begin{eqnarray*}
\sum_{\ell=0}^{\infty}\sqrt{\frac{N}{L_N}}\|\widetilde{\mathbf{B}}_\ell\|_2\le \sum_{\ell=0}^{\infty}\sqrt{\frac{N}{L_N}}\sum_{k=\ell+1}^{\infty}\| \mathbf{B}_k\|_2=\sum_{\ell=1}^{\infty}\ell \sqrt{\frac{N}{L_N}}\| \mathbf{B}_\ell\|_2=\sum_{\ell=1}^{\infty}\ell\cdot C_{N\ell}<\infty.
\end{eqnarray*}
(1).b According to (1).a, write
\begin{eqnarray*}
\sum_{s=1}^t \mathbf{x}_{s}&=&\mathbf{B} \sum_{s=1}^t\pmb{\varepsilon}_{s}-\sum_{\ell=0}^{t-1}\widetilde{\mathbf{B}}_\ell \pmb{\varepsilon}_{t-\ell}-\sum_{\ell=t}^{\infty}\widetilde{\mathbf{B}}_\ell \pmb{\varepsilon}_{t-\ell}+\sum_{\ell=0}^{\infty}\widetilde{\mathbf{B}}_\ell \pmb{\varepsilon}_{-\ell} \notag \\
&=&\mathbf{B} \sum_{\ell=1}^t\pmb{\varepsilon}_{\ell}-\sum_{\ell=1}^{t}\widetilde{\mathbf{B}}_{t-\ell} \pmb{\varepsilon}_{\ell}-\sum_{\ell=0}^{\infty}\widetilde{\mathbf{B}}_{t+\ell} \pmb{\varepsilon}_{-\ell}+\sum_{\ell=0}^{\infty}\widetilde{\mathbf{B}}_\ell \pmb{\varepsilon}_{-\ell} \notag \\
&=& \sum_{\ell=1}^t (\mathbf{B}-\widetilde{\mathbf{B}}_{t-\ell})\pmb{\varepsilon}_{\ell}-\sum_{\ell=0}^{\infty}(\widetilde{\mathbf{B}}_{t+\ell} -\widetilde{\mathbf{B}}_\ell)\pmb{\varepsilon}_{-\ell}\notag\\
&=& \sum_{\ell=1}^t (\mathbf{B}-\widetilde{\mathbf{B}}_{t-\ell})\pmb{\varepsilon}_{\ell}-\sum_{\ell=-\infty}^{0}(\widetilde{\mathbf{B}}_{t-\ell} -\widetilde{\mathbf{B}}_{-\ell})\pmb{\varepsilon}_{\ell} \eqqcolon \sum_{\ell=-\infty}^t \pmb{\mathcal{B}}_{t\ell} \pmb{\varepsilon}_\ell.
\end{eqnarray*}
(2). Note that for a given vector $\mathbf{v}$, $|\mathbf{v}|_2\le |\mathbf{v}|_1$ in which $|\mathbf{v}|_j$ with $j=1,2$ defines its $L^j$ norm. In connection with the fact that $\sum_{\ell=0}^{\infty}\sqrt{\frac{N}{L_N}}\|\widetilde{\mathbf{B}}_\ell\|_2<\infty $ of the first result, we obtain that $\sum_{\ell=0}^{\infty} \frac{N}{L_N}\|\widetilde{\mathbf{B}}_\ell\|_2^2< \infty.$ We are now able to write
\begin{eqnarray*}
E\left|\frac{1}{\sqrt{L_N}} \sum_{\ell=-\infty}^0 \mathbf{1}_N^\top\pmb{\mathcal{B}}_{T\ell} \pmb{\varepsilon}_{\ell}\right|^2 &=&\frac{1}{L_N}\sum_{\ell=-\infty}^0\mathbf{1}_N^\top\pmb{\mathcal{B}}_{T\ell}\pmb{\mathcal{B}}_{T\ell}^\top \mathbf{1}_N\notag \\
&=&\frac{1}{L_N}\sum_{\ell=-\infty}^0\mathbf{1}_N^\top (\widetilde{\mathbf{B}}_{T-\ell} -\widetilde{\mathbf{B}}_{-\ell})(\widetilde{\mathbf{B}}_{T-\ell} -\widetilde{\mathbf{B}}_{-\ell})^\top \mathbf{1}_N\notag \\
&\le &\frac{2}{L_N}\sum_{\ell=T}^{\infty}\|\mathbf{1}_N\|_2^2 \|\widetilde{\mathbf{B}}_{\ell}\|_2^2+\frac{2}{L_N}\sum_{\ell=0}^\infty\|\mathbf{1}_N\|_2^2 \|\widetilde{\mathbf{B}}_{\ell}\|_2^2\notag \\
&\le &\frac{4N}{L_N}\sum_{\ell=0}^\infty \|\widetilde{\mathbf{B}}_{\ell}\|_2^2 <\infty.
\end{eqnarray*}
(3). By the first result of this lemma, we have
\begin{eqnarray*}
E\left| \frac{1}{\sqrt{L_N}}\sum_{t=1}^T \mathbf{1}_N^\top\widetilde{\mathbf{B}}_{T-\ell} \pmb{\varepsilon}_{\ell} \right|^2 &=&\frac{1}{L_N}\mathbf{1}_N^\top \left(\sum_{t=1}^T \widetilde{\mathbf{B}}_{T-\ell}\right) \left(\sum_{t=1}^T \widetilde{\mathbf{B}}_{T-\ell}\right)^\top\mathbf{1}_N\notag\\
&\le &\left(\sqrt{\frac{N}{L_N}}\sum_{\ell=0}^{T-1}\|\widetilde{\mathbf{B}}_{\ell}\|_2\right)^2 \le \left(\sqrt{\frac{N}{L_N}}\sum_{\ell=0}^{\infty}\|\widetilde{\mathbf{B}}_{\ell}\|_2\right)^2<\infty.
\end{eqnarray*}
The proof is now completed.
proof[Proof of Lemma (ref)]
• (1). For simplicity, we drop index $t$, so write $\pmb{\varepsilon}=(\varepsilon_1,\ldots, \varepsilon_N)^\top$. It is obvious that
\begin{eqnarray*}
&&\| E[(\pmb{\varepsilon} \otimes \pmb{\varepsilon} )(\pmb{\varepsilon} ^\top \otimes \pmb{\varepsilon} ^\top)]\|_2\le \sqrt{\|E[(\pmb{\varepsilon} \otimes \pmb{\varepsilon} )(\pmb{\varepsilon} ^\top \otimes \pmb{\varepsilon} ^\top)]\|_1\|E[(\pmb{\varepsilon} \otimes \pmb{\varepsilon} )(\pmb{\varepsilon} ^\top \otimes \pmb{\varepsilon} ^\top)]\|_\infty}=O(N).
\end{eqnarray*}
(2). Write
\begin{eqnarray*}
\frac{1}{L_NT}\sum_{t=1}^T \mathbf{B}_0^*(L)vec(\pmb{\varepsilon}_{t} \pmb{\varepsilon}_{t}^\top)&=& \frac{1}{L_NT}\sum_{t=1}^T\mathbf{B}_0^* vec(\pmb{\varepsilon}_{t} \pmb{\varepsilon}_{t}^\top)-\frac{1}{L_NT} \widetilde{\mathbf{B}}_0^*(L) vec(\pmb{\varepsilon}_{T} \pmb{\varepsilon}_{T}^\top)\notag \\
&&+ \frac{1}{L_NT} \widetilde{\mathbf{B}}_0^*(L) vec(\pmb{\varepsilon}_{0} \pmb{\varepsilon}_{0}^\top).
\end{eqnarray*}
Note that
\begin{eqnarray*}
&&E[\widetilde{\mathbf{B}}_0(L) vec(\pmb{\varepsilon}_{T} \pmb{\varepsilon}_{T}^\top) vec(\pmb{\varepsilon}_{T} \pmb{\varepsilon}_{T}^\top)^\top \widetilde{\mathbf{B}}_0(L)^\top ]\notag \\
&\le & \lambda_{\max}(E[(\pmb{\varepsilon} \otimes \pmb{\varepsilon} )(\pmb{\varepsilon} ^\top \otimes \pmb{\varepsilon} ^\top)])\sum_{\ell=0}^{\infty}\|\widetilde{\mathbf{B}}_{0\ell}^* \|_2 \notag \\
&\le & O(N)\sum_{\ell=0}^{\infty} \sum_{k=\ell+1}^{\infty}\|(\mathbf{1}_N^\top \mathbf{B}_k)\otimes (\mathbf{1}_N^\top \mathbf{B}_{k})\|_2 \le O(1)N^2 \sum_{\ell=0}^{\infty} \ell \|\mathbf{B}_{\ell}\|_2^2,
\end{eqnarray*}
where the second inequality follows from the first result of this lemma. Thus,
\begin{eqnarray*}
\frac{1}{L_NT} |\widetilde{\mathbf{B}}_0^*(L) \text{vec}(\pmb{\varepsilon}_{T} \pmb{\varepsilon}_{T}^\top)|&=&O_P(1)\frac{1}{L_NT} \sqrt{N^2 \sum_{\ell=0}^{\infty} \ell \|\mathbf{B}_{\ell}\|_2^2} \notag \\
&\le &O_P(1)\frac{1}{\sqrt{L_N}T} \sum_{\ell=0}^{\infty} \frac{\sqrt{N^2 \ell}}{\sqrt{L_N}} \|\mathbf{B}_{\ell}\|_2 = O_P(1)\frac{1}{ T}.
\end{eqnarray*}
Similarly, we have
\begin{eqnarray*}
\frac{1}{L_NT} |\widetilde{\mathbf{B}}_0^*(L) \text{vec}(\pmb{\varepsilon}_0 \pmb{\varepsilon}_0^\top)|= O_P(1)\frac{1}{ T}.
\end{eqnarray*}
Therefore, we need only to consider $\frac{1}{L_NT}\sum_{t=1}^T\mathbf{B}_0^* \text{vec}(\pmb{\varepsilon}_{t} \pmb{\varepsilon}_{t}^\top)$ in what follows. By the proof similar to Proposition (ref), we obtain that
\begin{eqnarray*}
\left|\frac{1}{L_NT}\sum_{t=1}^T \mathbf{B}_0^*(L)\text{vec}(\pmb{\varepsilon}_{t} \pmb{\varepsilon}_{t}^\top) - \frac{1}{L_N} \sum_{\ell=0}^{\infty} \mathbf{1}_N^\top \mathbf{B}_\ell \mathbf{B}_\ell^\top \mathbf{1}_N \right|=O_P\left(\frac{1}{\sqrt{T}}\right).
\end{eqnarray*}
(3). Note that
\begin{eqnarray*}
&&\frac{1}{L_NT}\sum_{t=1}^T\sum_{v=1}^\infty \mathbf{B}_v^*(L)\text{vec}(\pmb{\varepsilon}_{t-v} \pmb{\varepsilon}_{t}^\top ) \notag \\
&=&\sum_{v=1}^\infty \frac{1}{L_NT}\sum_{t=1}^T \mathbf{B}_v^* \text{vec}(\pmb{\varepsilon}_{t-v} \pmb{\varepsilon}_{t}^\top )+\sum_{v=1}^\infty \frac{1}{L_NT} \widetilde{\mathbf{B}}_v^*(L) \text{vec}(\pmb{\varepsilon}_{T-v} \pmb{\varepsilon}_{T}^\top)\notag \\
&&+\sum_{v=1}^\infty \frac{1}{L_NT} \widetilde{\mathbf{B}}_v^*(L) \text{vec}(\pmb{\varepsilon}_{0-v} \pmb{\varepsilon}_{0}^\top).
\end{eqnarray*}
In what follows, we consider the three terms on the right hand side one by one.
Write
\begin{eqnarray*}
&&E\left|\sum_{t=1}^T\sum_{v=1}^\infty \mathbf{B}_v^* \text{vec}(\pmb{\varepsilon}_{t-v} \pmb{\varepsilon}_{t}^\top ) \right|^2 = \sum_{t,s=1}^T\sum_{v,k=1}^\infty\mathbf{B}_v^* E[\text{vec}(\pmb{\varepsilon}_{t-v} \pmb{\varepsilon}_{t}^\top )\text{vec}(\pmb{\varepsilon}_{s-k} \pmb{\varepsilon}_{s}^\top )^\top]\mathbf{B}_k^{*\top}\notag \\
&=&\sum_{t=1}^T\sum_{v,k=1}^\infty\mathbf{B}_v^* E[\text{vec}(\pmb{\varepsilon}_{t-v} \pmb{\varepsilon}_{t}^\top )\text{vec}(\pmb{\varepsilon}_{t-k} \pmb{\varepsilon}_{t}^\top )^\top]\mathbf{B}_k^{*\top}\notag \\
&\le &O(1)T\sum_{v=1}^\infty\|\mathbf{B}_v^*\|_2^2\leq O(1)NT\left(\sum_{v=1}^\infty\sum_{\ell=0}^{\infty} \| \mathbf{B}_\ell\|_2 \|\mathbf{B}_{\ell+v}\|_2\right)^2 \notag \\
&\le &O(1)NT \left(\sum_{\ell=0}^{\infty} \| \mathbf{B}_\ell\|_2 \right)^2.
\end{eqnarray*}
Thus, we have
\begin{eqnarray*}
\left|\frac{1}{L_NT}\sum_{t=1}^T \sum_{t=1}^T\sum_{v=1}^\infty \mathbf{B}_v^* \text{vec}(\pmb{\varepsilon}_{t-v} \pmb{\varepsilon}_{t}^\top )\right|=O_P(1)\frac{\sqrt{N}}{\sqrt{L_N}T}\sum_{\ell=0}^{\infty} \| \mathbf{B}_\ell\|_2 =O_P(1)\frac{1}{\sqrt{T}},
\end{eqnarray*}
where the second equality follows from Assumption (ref). Similar to the proof of the second result, it is easy to know that the term $\sum_{v=1}^\infty \frac{1}{L_NT}\sum_{t=1}^T \mathbf{B}_v^* \text{vec}(\pmb{\varepsilon}_{t-v} \pmb{\varepsilon}_{t}^\top )$ offers the lowest rate. Thus, we have
\begin{eqnarray*}
\left|\frac{1}{L_NT}\sum_{t=1}^T\sum_{v=1}^\infty \mathbf{B}_v^*(L)\text{vec}(\pmb{\varepsilon}_{t-v} \pmb{\varepsilon}_{t}^\top ) \right|= O_P(1)\frac{1}{\sqrt{T}}.
\end{eqnarray*}
The proof is now completed.
proof[Proof of Lemma (ref)]
• (1). We now show that
\begin{eqnarray*}
E^*[S_{NT}^{*2}]= \sigma_x^2+o_P(1),
\end{eqnarray*}
which then infers the desired result. It suffices to show that
\begin{eqnarray}
\frac{1}{L_NT} \sum_{t=1}^T \mathbf{x}_{t}^\top \mathbf{1}_N\mathbf{1}_N^\top \mathbf{x}_{t}= \frac{1}{L_NT} \sum_{t=1}^T E[\mathbf{x}_{t}^\top \mathbf{1}_N\mathbf{1}_N^\top \mathbf{x}_{t}]+o_P(1),
\end{eqnarray}
and
\begin{eqnarray}
\frac{1}{L_NT}\sum_{k=1}^{T-1}\sum_{t=1}^{T-k} \mathbf{x}_{t}^\top \mathbf{W}_{ts}\mathbf{x}_{t+k}=\frac{1}{L_NT}\sum_{k=1}^{T-1}\sum_{t=1}^{T-k} E[\mathbf{x}_{t}^\top \mathbf{1}_N \mathbf{1}_N^\top \mathbf{x}_{t+k}]+o_P(1).
\end{eqnarray}
We start with (ref). By (ref) and Lemma (ref), we immediately obtain that
\begin{eqnarray}
\frac{1}{L_NT} \sum_{t=1}^T \mathbf{x}_{t}^\top \mathbf{1}_N\mathbf{1}_N^\top \mathbf{x}_{t}=\frac{1}{L_N} \sum_{\ell=0}^{\infty} \mathbf{1}_N^\top \mathbf{B}_\ell \mathbf{B}_\ell^\top \mathbf{1}_N+O_P\left(\frac{1}{\sqrt{T}}\right).
\end{eqnarray}
Therefore, the proof of (ref) is completed.
We then consider (ref). Let $a_{k/m}\coloneqq a\left(\frac{k}{m}\right)$ for notational simplicity. Firstly, we consider
\begin{eqnarray}
&&E\left|\frac{1}{L_NT}\sum_{k=1}^{T-1}\sum_{t=1}^{T-k} (\mathbf{x}_{t}^\top \mathbf{W}_{ts}\mathbf{x}_{t+k}-E[\mathbf{x}_{t}^\top \mathbf{W}_{ts}\mathbf{x}_{t+k}])\right|\notag \\
&=&E\left|\frac{1}{L_NT}\sum_{k=1}^{T-1} a_{k/m} \sum_{t=1}^{T-k} (\mathbf{x}_{t}^\top \mathbf{1}_N \mathbf{1}_N^\top\mathbf{x}_{t+k}-E[\mathbf{x}_{t}^\top \mathbf{1}_N \mathbf{1}_N^\top\mathbf{x}_{t+k}])\right|\notag \\
&=&\sum_{k=1}^{T-1} a_{k/m} E\left|\frac{1}{L_NT} \sum_{t=1}^{T-k} (\mathbf{x}_{t}^\top \mathbf{1}_N \mathbf{1}_N^\top\mathbf{x}_{t+k}-E[\mathbf{x}_{t}^\top \mathbf{1}_N \mathbf{1}_N^\top\mathbf{x}_{t+k}])\right|\notag \\
&=&O_P(1)\sum_{k=1}^{T-1} a_{k/m} \cdot \frac{1}{\sqrt{T}} = O_P\left(\frac{m}{\sqrt{T}}\right),
\end{eqnarray}
where the third equality follows from a development similar to that for (ref), and the last equality follows from Assumption (ref). Sequentially, we consider
\begin{eqnarray}
&&\frac{1}{L_NT}\sum_{k=1}^{T-1}\sum_{t=1}^{T-k} E[\mathbf{x}_{t}^\top \mathbf{W}_{ts}\mathbf{x}_{t+k}] = \frac{1}{L_NT}\sum_{k=1}^{T-1}\sum_{t=1}^{T-k} E[\mathbf{x}_{t}^\top \mathbf{1}_N \mathbf{1}_N^\top\mathbf{x}_{t+k}] \notag \\
&&+\frac{1}{L_NT}\sum_{k=1}^{T-1}\sum_{t=1}^{T-k} (a_{k/m}-1) E[\mathbf{x}_{t}^\top \mathbf{1}_N \mathbf{1}_N^\top\mathbf{x}_{t+k}].
\end{eqnarray}
Note that
\begin{eqnarray}
&&\left|\frac{1}{L_NT}\sum_{k=1}^{T-1}\sum_{t=1}^{T-k} (a_{k/m}-1) E[\mathbf{x}_{t}^\top \mathbf{1}_N \mathbf{1}_N^\top\mathbf{x}_{t+k}] \right|\notag \\
&\le & \frac{1}{L_N}\sum_{k=1}^{d_T} |a_{k/m}-1|\cdot |E[\mathbf{x}_{t}^\top \mathbf{1}_N \mathbf{1}_N^\top\mathbf{x}_{t+k}] |+O(1)\frac{1}{L_N}\sum_{k=d_T+1}^{\infty} |E[\mathbf{x}_{t}^\top \mathbf{1}_N \mathbf{1}_N^\top\mathbf{x}_{t+k}] |\notag \\
&\le & \frac{1}{L_N} |E[\mathbf{x}_{t}^\top \mathbf{1}_N \mathbf{1}_N^\top\mathbf{x}_{t+1}] |\sum_{k=1}^{d_T} \frac{k}{m} +O(1)\frac{1}{L_N}\sum_{k=d_T+1}^{\infty} |E[\mathbf{x}_{1}^\top \mathbf{1}_N \mathbf{1}_N^\top\mathbf{x}_{1+k}] | \rightarrow 0
\end{eqnarray}
by letting $d_T^2/m+1/d_T\to 0$, where the second inequality follows from Assumption (ref).
By (ref), (ref) and (ref), we have proved (ref).
Up to this point, we have verified both (ref) and (ref). Thus, the result in Lemma (ref).1 follows.
Using arguments similar to those employed in the proof of Lemma (ref).1, we can also establish the results in Lemmas (ref).2-3. The detailed steps are omitted here to avoid unnecessary repetition.
proof[Proof of Lemma (ref)]
• For the first result, it suffices to show
\begin{eqnarray}
\Big|\frac{1}{T}\sum_{s,t=1}^T\mathbf{e}_i^\top \mathbf{x}_t\mathbf{x}_s^\top\mathbf{e}_i a\Big(\frac{t-s}{m}\Big)-\sigma_{p,i}^2\Big|=o_P(1),
\end{eqnarray}
for each $i$.
In a manner analogous to (ref), we can employ similar arguments as those used in (ref) along with Lemma (ref) and Assumption (ref) to derive the following result:
\begin{eqnarray}
&&\Big|\frac{1}{T}\sum_{t=1}^T(\mathbf{e}_i^\top \mathbf{x}_t\mathbf{x}_t^\top\mathbf{e}_i-E[\mathbf{e}_i^\top \mathbf{x}_t\mathbf{x}_t^\top\mathbf{e}_i])\Big|\notag \\
&=&\Big|\frac{1}{T}\sum_{t=1}^T\mathbf{e}_i^\top \mathbf{B}(L)(\pmb{\varepsilon}_t\pmb{\varepsilon}_t^\top-E[\pmb{\varepsilon}_t\pmb{\varepsilon}_t^\top])\mathbf{B}(L)^\top\mathbf{e}_i\Big|
=O_P\left(\frac{1}{\sqrt{T}}\right).
\end{eqnarray}
Recall that we have defined $a_{k/m}= a\left(\frac{k}{m}\right)$. We can write
\begin{eqnarray}
&&\left|\frac{1}{T}\sum_{k=1}^{T-1}\sum_{t=1}^{T-k} a_{k/m}\big(\mathbf{e}_i^\top \mathbf{x}_t\mathbf{x}_{t+k}^\top\mathbf{e}_i-E[\mathbf{e}_i^\top \mathbf{x}_t\mathbf{x}_{t+k}^\top\mathbf{e}_i] \big)\right|\notag \\
&=&\sum_{k=1}^{T-1} a_{k/m} \left|\frac{1}{T} \sum_{t=1}^{T-k} \big(\mathbf{e}_i^\top \mathbf{x}_t\mathbf{x}_{t+k}^\top\mathbf{e}_i-E[\mathbf{e}_i^\top \mathbf{x}_t\mathbf{x}_{t+k}^\top\mathbf{e}_i] \big)\right|\notag \\
&=&O_P\left(\sum_{k=1}^{T-1} a_{k/m} \cdot \frac{1}{\sqrt{T}}\right) = O_P\left(\frac{m}{\sqrt{T}}\right).
\end{eqnarray}
Additionally, using arguments similar to those in (ref), we can further show that
\begin{eqnarray}
\left|\frac{1}{T}\sum_{k=1}^{T-1}\sum_{t=1}^{T-k}(a_{k/m}-1)E[\mathbf{e}_i^\top \mathbf{x}_t\mathbf{x}_{t+k}^\top\mathbf{e}_i] \right|=o_P(1).
\end{eqnarray}
In light of the results in (ref), (ref) and (ref), we are ready to establish (ref) and conclude that Lemma (ref).1 holds. Using similar arguments, we can obtain the desired results in Lemmas (ref).2-3. The detailed steps are omitted here.