Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Estimation of Panel Data Models with Nonlinear Factor Structure$^$
\dedicatory{\normalfont \today}
abstractPanel data models with unobserved heterogeneity in the form of interactive effects standardly assume that the time effects -- or “common factors” -- enter linearly. This assumption is unnatural in the sense that it pertains to the unobserved component of the model, and there is rarely any reason to believe that this component takes on a particular functional form. This is in stark contrast to the relationship between the observables, which can often be credibly argued to be linear. Linearity in the factors has persevered mainly because it is convenient, and that it is better than standard fixed effects. The present paper relaxes this assumption. It does so by combining the common correlated effects (CCE) approach to standard interactive effects with the method of sieves. The new estimator -- abbreviated “SCCE” -- retains many of the advantages of CCE, including its computational simplicity, and good small-sample and asymptotic properties, but is applicable under a much broader class of factor structures that includes the linear one as a special case. This makes it well-suited for a wide range of empirical applications.
JEL Classification: C13; C14; C33;
Keywords: nonlinear factor models, CCE, sieve estimation, panel data, unobserved heterogeneity.
Introduction
Consider the panel variable $y_{i,t} \in \mathbb{R}$, observable for $i=1, \dots, N$ cross-sectional units and $t=1, \dots, T$ time periods. Following the bulk of the existing literature on interactive effects models for such variables (see chudik_pesaran_2015, and karabiyik2019, for surveys), we assume that $y_{i,t}$ depends linearly on a vector of regressors $\*x_{i,t} \in \mathbb{R}^{d}$. The main difference compared to said literature is that in the present paper the interactive effects are not necessarily linear. The particular model that we will be considering is given by
align[align omitted — 137 chars of source]
where $\varepsilon_{i,t} \in \mathbb{R}$ and $\*v_{i,t} \in \mathbb{R}^{d}$ are idiosyncratic errors that are assumed to be independent of one another. However, this does not mean that $\*x_{i,t}$ is exogenous. The reason is the presence of unobserved heterogeneity in the form of interactive effects, as captured by $g_i(\*f_t)$ and $\*G_i(\*f_t)$. Here $\*f_t\in \mathbb{R}^{m}$ is a vector of unobserved common factors, whose dimension, $m$, is fixed and does not need to be known. The way that $\*f_t$ enters (ref) and (ref) is determined by $g_i(\cdot) \in \mathbb{R}$ and $\*G_i(\cdot) = [G_{1,i}(\cdot),\ldots, G_{d,i}(\cdot)]'\in \mathbb{R}^{d}$, where $g_i(\cdot),G_{1,i}(\cdot),\ldots, G_{d,i}(\cdot) : \mathbb{R}^m \to \mathbb{R}$ are some unknown, real-valued functions. For reasons that will soon become clear, it is convenient to write $g_i(\cdot)$ and $\*G_i(\cdot)$ as functions of $\*f_t$, and to let dependence on $i$ be absorbed by these functions. This is not necessary, though, but we can just as well write $g_i(\*f_t) = g_t(\+\gamma_i)$, where $\+\gamma_i\in \mathbb{R}^{m}$ is a vector of factor loadings, or even $g_i(\*f_t) = g(\*f_t,\+\gamma_i)$. The same is true for $\*G_i(\cdot)$. Either way, because the unobserved interactive effects enter both (ref) and (ref), $\*x_{i,t}$ is potentially endogenous. Our goal is to consistently estimate and conduct inference on $\+\beta$ in this case.
The existing interactive effects literature can be divided into two main strands (see, for example, chudik_pesaran_2015, and karabiyik2019). One is based on the common correlated effects (CCE) approach of pesaran2006. This approach is extremely simple to implement as it has a closed form and it does not require accurate estimation of the number of factors, $m$; however, this simplicity has a price in that $g_i(\cdot)$ and $\*G_i(\cdot)$ must be linear, such that $g_i(\*f_t) =\+\gamma_i' \*f_t $ and $\*G_i(\*f_t) = \+\Gamma_i' \*f_t$, where $\+\Gamma_i\in \mathbb{R}^{m\times d}$.\footnote{As DeVos2019 point out, requiring $\*G_i(\cdot)$ to be linear is restrictive not only in itself but also because it rules out models in which there are regressors that enter (ref) nonlinearly.} The other strand is based on the quasi maximum likelihood (QML) approach of bai2009. This approach is more general than CCE in that it does not require $\*G_i(\cdot)$ to be linear but then it is also computationally much more involved.\footnote{There are some studies based on generalized methods of moments (see, for example, Ahn2013, and Robertson2015). These are similar to those based on QML in that $\*G_i(\cdot)$ is not required to be linear.} Both approaches require that $g_i(\cdot)$ is linear.
Requiring $g_i(\cdot)$ to be linear is very convenient because it means that given a consistent estimate of the space spanned by $\*f_t$, the source of the endogeneity of $\*x_{i,t}$ can be netted out using simple least squares (LS) orthogonal projections.\footnote{As is well known from the classical common-factor literature, factors and loadings cannot be identified separately; at best, one can consistently estimate only the spaces they span.} Of course, just because linearity is simple, does not mean it is correct, a fact that is now being increasingly recognized in the literature. For example, in climate economics, unobserved factors such as common technological adaptation or weather are likely to exert nonlinear effects. Temperature's effect on productivity and agricultural yield, for instance, often exhibit threshold behavior; small variations may have little impact, but larger variations can after a certain point become extremely damaging (see hsiang16). In empirical asset pricing, it is well known that conventional linear factor models cannot price securities whose payoffs are nonlinear functions of the factors, especially in the presence of derivative securities (see bansal). Therefore, nonlinear factor models offer more appropriate and flexible pricing kernels. In macroeconomics, evidence indicates that inflation responds disproportionally to large, common economic shocks, while small shocks have little or no effect (see bobeica). Another example, which we will come back to in our empirical illustration, is from labor economics. voig investigates the effects of intersectoral technology-skill complementarity on skill demand. The study notes that technological changes in one sector of the economy are likely to have nonlinear effects on the skill demand in other sectors. These are a few examples but there are many more. In fact, it is difficult to imagine a scenario of empirical relevance in which it is known that $g_i(\cdot)$ and $\*G_i(\cdot)$ are linear. This scenario is ruled out almost by construction, as $g_i(\cdot)$ and $\*G_i(\cdot)$ are a part of the unobserved part of the model.
Although the functional forms of $g_i(\cdot)$ and $\*G_i(\cdot)$ are generally unknown, the relationship between $y_{i,t}$ and $\*x_{i,t}$ is often better understood. Economic theory frequently implies linearity, and empirical work commonly supports it. A vast majority of studies therefore assumes that the relationship is linear, and so do we. What we have in mind is an empirical researcher interested in estimating the average marginal effect of $y_{i,t}$ and $\*x_{i,t}$. Consistent with previous research in the area, this part of the model is taken to be linear, which means that the average marginal effect is given by $\+\beta$. The model is also parsimonious, which is again consistent with previous research, and this raises concerns about unobserved heterogeneity due to omitted variables. Little is known about the nature of this heterogeneity, except for the fact that it is likely present. The researcher therefore wants to infer $\+\beta$ while controlling for as much unobserved heterogeneity as possible.\footnote{This scenario is inspired in part by cross-sectional studies such as Cattaneo03072018, and GALBRAITH2020609. Here there are “core” regressors that are known to enter the model for the dependent variable linearly, and there are “cofounders” that may enter nonlinearly. In this part of the literature, the cofounders are assumed to be observed and the main challenge is how to control for their -- potentially nonlinear -- effect.} The approach in this paper does exactly that.
Our main contribution is that we allow $g_i(\cdot)$ and $\*G_i(\cdot)$ to be nonlinear without requiring it. And if they are nonlinear, the only assumption we make is that the functions should be smooth. Hence, in our paper, $g_i(\cdot)$ and $\*G_i(\cdot)$ are completely nonparametric, which means that the type of unobserved heterogeneity that can be permitted is very general. We also do not require any knowledge of the functional form of $g_i(\cdot)$ and $\*G_i(\cdot)$, which means that researchers are spared of the usual problem in practice of having to specify a parametric model for the unobserved heterogeneity.
Linear interactive effects represent the state of the art, but most empirical research is still based on additive fixed effects specifications, which makes it necessary to decide on which effects to include. Moreover, since fixed effects are known to be insufficient in many applications, there is also a need for control variables and data on those. If data are available, there is -- at least in principle -- a further need to decide on how the controls should enter the model, although linear specifications are so common that they are hardly ever questioned. Controlling for unobserved heterogeneity therefore involves a lot of choices and assumptions that need not be correct (see LU2014194, for an excellent, general discussion). The proposed nonparametric treatment of $g_i(\cdot)$ and $\*G_i(\cdot)$ enables a completely agnostic approach to unobserved heterogenity, which makes it very attractive from an applied point of view.
The new estimation procedure can be seen as a sieve extension of the CCE approach to linear interactive effects models, and is henceforth referred to as “SCCE”. Our preference to build on CCE is first and foremost its computational simplicity. This is important even under linearity, and it is essential in more general models where other approaches need not be feasible. Another reason for building on CCE is its excellent small-sample properties. Again, while important under linearity, good performance in small samples is not something that can be taken for granted in more general models.
The idea is to use the cross-sectional averages of $y_{i,t}$ and $\*x_{i,t}$ to construct a first-pass “mock” proxy for $\*f_t$. On their own, these averages are unlikely to span the full factor space. SCCE addresses this by expanding the averages via the method of sieves (see chen2007, for a review). This method generates a sequence of nonlinear transformations of the averages that enrich their span, which makes it possible to approximate $g_i(\cdot)$ and $\*G_i(\cdot)$ without specifying their functional form. We then project out the variation explained by this sequence and estimate $\+\beta$ by applying pooled LS to the resulting projection errors. Our asymptotic results establish that the new estimator is consistent at the usual parametric rate and asymptotically normal. Extensive Monte Carlo simulations confirm good small sample properties under a variety of data generating processes.
The present paper is not the first to allow for nonlinearities in panel data. However, existing studies typically do not consider the same type of nonlinearity as we do. The closest study in terms of estimation strategy is that of su12, who also use CCE and sieve estimation. However, they assume that the factors enter linearly and place the nonlinearity on the regressors, which is the complete opposite of what we do.\footnote{An incomplete list of other -- even more distant -- papers that consider nonlinear panel data models include aty, boneva17, fer21, and frey. We refer to free for further references and discussion.} For reasons already given, we believe that our model has greater appeal not only from an empirical point of view but also theoretically because the nonlinearity is placed on the factors, which -- in contrast to the regressors -- are unobserved. The only other paper that we are aware of to consider a model with this type of nonlinearity is free. However, their estimation approach, which is an extension of the QML approach of bai2009, is computationally much more involved, and hence not very user-friendly. Specifically, the nonlinearity of the interactive effects is handled through an approximation that involves an infinite number of factors, and the estimation of slope coefficients, factors and loadings is carried out jointly, which requires solving a high-dimensional, non-convex optimization problem. The procedure is therefore likely slow and highly dependent on the starting values used. By contrast, the proposed SCCE estimator has a closed form, and is therefore both fast and numerically stable. Moreover, unlike free, who for technical reasons only prove consistency of their procedure, SCCE is shown to be asymptotically normal, which means that it supports standard inference.
The rest of the paper proceeds as follows: Section (ref) introduces the SCCE estimator, whose asymptotic and small-sample properties are investigated in Sections (ref) and (ref), respectively. In Section (ref), we apply the new estimator to investigate the drivers of the increasing wage inequality between high- and low-skilled workers in U.S. manufacturing. Section (ref) concludes. All proofs, and some additional empirical and Monte Carlo results are reported in the paper's appendix.
Notation: For a real matrix $\*A$, we use $\mathrm{tr}(\*A)$, $\mathrm{rank}(\*A)$ and $\|\*A\| \equiv \sqrt{\mathrm{tr}\,(\*A'\*A)}$ to denote its trace, rank and Frobenius (Euclidean) norm, respectively, where $\equiv$ signifies definitional equality. $\mathrm{vec}(\*A)$ vectorizes $\*A$ by stacking its columns on top of each other. The maximum and minimum eigenvalues of $\*A$ are denoted $\lambda_{\max}(\*A)$ and $\lambda_{\min}(\*A)$, respectively. If $\*A_1,\dots, \*A_N$ are matrices, $\overline{\*A} \equiv \frac{1}{N} \sum_{i=1}^N \*A_i$ denotes their average. For any $T$-rowed matrix $\mathbf{A}$, we define its projection error matrix $\mathbf{M}_{A} \equiv \mathbf{I}_{T} - \*P_{A}$, where $\*P_{A}\equiv \mathbf{A}(\mathbf{A}'\mathbf{A})^{+}\mathbf{A}'$ with $(\mathbf{A}'\mathbf{A})^+$ being the Moore--Penrose inverse of $\mathbf{A}'\mathbf{A}$ when $\mathbf{A}$ is not of full column rank. Given a vector function $\*g_i = [g_1(\cdot), \dots, g_n(\cdot)]' \in \mathbb{R}^{n}$ of some vector $\mathbf{f} = [f_1, \dots, f_m]' \in \mathbb{R}^{m}$, the gradient matrix is given by $\nabla \*g_i(\mathbf{f}) \equiv \frac{\partial \*g_i(\mathbf{f})}{\partial \mathbf{f}'} \in \mathbb{R}^{n \times m}$, whose $(i,j)$-th element is given by $\frac{\partial g_i(\mathbf{f})}{\partial f_j}$. If we further let $\*a = [a_1, \dots, a_m]' \in \mathbb{R}^{m}$ be a vector of nonnegative integers, and define $|\*a| \equiv \sum_{j=1}^m a_j$ for such vectors, then the $|\*a|$-th partial derivative of $g(\cdot)$ is given by $\nabla^{\*a} g(\*f) \equiv \frac{\partial^{|\*a|}}{\partial f_{1}^{a_1}, \dots, \partial f_{m}^{a_m}} g(\*f)\in \mathbb{R}^n$. The notation $\*x_n = O_p(a_n)$ means that a random vector $\*x_n$ is at most of order $a_n$ in probability, where $a_n$ is some deterministic sequence, while $\*x_n = o_p(a_n)$ means it is of smaller order in probability than $a_n$. We use w.p.1 (w.p.a.1) to denote with probability (approaching) one, $\overset{p}{\to}$ to denote convergence in probability, $\overset{d}{\to}$ to denote convergence in distribution, and $\lfloor x \rfloor $ to denote the largest integer less than $x$. Finally, $C < \infty$ denotes a generic, positive constant whose value may differ from one case to another.
Motivation and Estimation Procedure
It is convenient to write the model presented in Section (ref) on time-stacked vector form. Let us therefore introduce $\*y_i = \left[y_{i,1}, \dots, y_{i,T}\right]' \in \mathbb{R}^{T}$, $\*X_{i}= \left[\*x_{i,1}, \dots ,\*x_{i,T}\right]' \in \mathbb{R}^{T \times d}$, $\+\varepsilon_i = \left[\varepsilon_{i, 1}, \dots, \varepsilon_{i, T}\right]' \in \mathbb{R}^{T}$, $\*V_i = [\*v_{i, 1}, \dots, \*v_{i, T}]' \in \mathbb{R}^{T \times d}$, $\*g_i(\*F) = [g_i(\*f_1), \dots, g_i(\*f_T)]' \in \mathbb{R}^{T}$ and $\*G_i(\*F) = [\*G_i(\*f_1), \dots, \*G_i(\*f_T)]' \in \mathbb{R}^{T \times d}$. In this notation, (ref) and (ref) can be written as
align[align omitted — 116 chars of source]
Let us further denote by $\*z_{i,t} \equiv [y_{i,t}, \*x'_{i,t}]' \in \mathbb{R}^{d+1}$ the vector of observables and by $\*Z_i \equiv [\*y_i, \*X_i] = [\*z_{i,1}, \dots, \*z_{i,T}]' \in \mathbb{R}^{T \times (d+1)}$ its stacked version. Combining (ref) and (ref), the data generating process of this last matrix can be written in the following way:
align[align omitted — 58 chars of source]
where $\mathcal{G}_i(\*F) \equiv [\*g_i(\*F)+\*G_i(\*F)\+\beta, \*G_i(\*F)] = [\mathcal{G}_i(\*f_1), \dots, \mathcal{G}_i(\*f_T)]' \in \mathbb{R}^{T \times (d+1)}$, $\mathcal{G}_i(\cdot):\mathbb{R}^{m}\to\mathcal{G}_i\left(\mathbb{R}^m\right)\subseteq\mathbb{R}^{d+1}$, and $\*U_i \equiv [\+\varepsilon_i + \*V_i \+\beta , \*V_i] = [\*u_{i, 1}, \dots, \*u_{i, T}]' \in \mathbb{R}^{T \times (d+1)}$.
Equation (ref) is a nonlinear factor model for $\*Z_i$. One of the few papers to discuss estimation of such models is Amemiya. However, their QML-style estimation approach requires that the functional form of $\mathcal{G}_i(\cdot)$ is known. The same is true for the principal components-based approach of WANG2022180. These approaches are therefore not of any use to us. feng2023 recently considered the unknown $\mathcal{G}_i(\cdot)$ case. In our context with time as one of the two panel data dimensions, his approach involves time period-wise sample splitting, $K$-nearest neighbors matching of cross-sectional units and finally local-linear principal component estimation. Not only do these steps make for a computationally costly procedure, but the data also have to be independent over time, which is unrealistic. We therefore need a different approach.
One possible way to estimate (ref) is to ignore the functional form of $\mathcal{G}_i(\cdot)$ altogether and to treat its columns as $d+1$ distinct factors. One could then try to approximate these factors using the cross-sectional average $\widehat{\*F} \equiv\overline{\*Z}$ of $\*Z_i$, as prescribed by standard CCE. However, this is not as simple as it may sound. To appreciate the issues involved, it is useful to write out $\widehat{\*F}$ in terms of the elements of (ref) as
align[align omitted — 89 chars of source]
Because $\overline{\*U} \overset{p}{\to}\*0_{T \times (d+1)}$ as $N \to \infty$ for a given $T$ under general conditions, $\widehat{\*F} \overset{p}{\to} \overline{\mathcal{G}}(\*F)$.\footnote{As pointed out in Section (ref), the asymptotic results of this paper are based on letting both $N$ and $T$ to infinity. The fact that $\overline{\*U} \overset{p}{\to}\*0_{T \times (d+1)}$ as $N \to \infty$ for a given $T$ is therefore not enough. It is intuitive, though, and the main purpose of this discussion is to provide intuition. Formal proofs are provided in the appendix.} However, this does not necessarily mean that $\widehat{\*F}$ is consistent for the space spanned by $\*F$. In fact, there are a number of conditions that have to be met for $\widehat{\*F}$ to be consistent. Suppose for a moment that $\*g_i(\*F)$ and $\*G_i(\*F)$ are linear such that $\*g_i(\*F) = \*F \+\gamma_i$ and $\*G_i(\*F) = \*F \+\Gamma_i$. In this case, $\overline{\mathcal{G}}(\*F) = \*F[\overline{\+\gamma}+\overline{\+\Gamma}\+\beta, \overline{\+\Gamma}] \equiv \*F\overline{\*C}$, and hence $\widehat{\*F} \overset{p}{\to} \*F\overline{\*C}$. Not even this is enough to ensure that $\widehat{\*F}$ spans the space of $\*F$ asymptotically, unless we assume that $\operatorname{rank} (\overline{\*C})=m = d+1$. The reason for this last condition is that we need $\overline{\*C}$ to be invertible in order to show that $\*M_{F\overline{C}} = \*M_{F}$.\footnote{If $m < d+1$, there exists a full rank matrix $\overline{\*B}$, say, such that $m$ of the columns of $\widehat{\*F}\overline{\*B}$ converge to $\*F$, while the remaining $d+1-m$ columns converge to zero (see KARABIYIK201760).} Hence, since $\widehat{\*F}$ is consistent for $\*F\overline{\*C}$, it spans the space of $\*F$ asymptotically. However, this is possible only under very special circumstances. In general, there is no reason to expect $\widehat{\*F}$ to converge to anything meaningful, which means that there is also no reason to expect the resulting CCE estimator of $\+\beta$ to be consistent. In fact, as feng2023 points out, the number of primitive factors required to capture the nonlinearity of $\mathcal{G}_i(\cdot)$ is likely large, and much larger than the value of $d$ implied by parsimonious economic models. To take one example, HOLLY2010160 use CCE to study the effect of income ($d=1$) on housing prices in the U.S. They apply the information criteria of bai_ng with which they find evidence of up to six factors.
The present paper can be seen as a reaction to the problem described in the previous paragraph. The proposed solution consists of the following three-step procedure designed to deal with the potential nonliearity of $\mathcal{G}_i(\cdot)$:
step[Initial estimation]\ \normalfont
Compute $\widehat{\*F}$. This is our “mock” proxy of the factor component.
step[Sieve approximation of the factor space]\ \normalfont
As pointed out earlier, $\widehat{\*F}$ is likely inconsistent for the space spanned by $\mathcal{G}_i(\cdot)$. To address this issue, in the second step we employ the method of sieves. The main idea is to approximate $\mathcal{G}_i(\cdot)$ using a set of user-specified basis functions with $\widehat{\*F}$ as their argument. Let us therefore denote by $\*p=[p_1(\cdot), \dots, p_K(\cdot)]' \in \mathbb{R}^{K}$ a vector of $K$ such basis functions. The sought sieve basis matrix is given by
$\widehat{\*P} = \left[\*p(\widehat{\*f}_1), \dots, \*p(\widehat{\*f}_T)\right]' \in \mathbb{R}^{T \times K}$.
step[Estimation of $\+\beta$] \normalfont
Given $\widehat{\*P}$, our proposed estimator of $\boldsymbol{\beta}$, denoted by $\widehat{\boldsymbol{\beta}}_{SCCE}$, is given by
\begin{align}
\widehat{\boldsymbol{\beta}}_{SCCE} \equiv \left(\sum_{i=1}^N \mathbf{X}_{i}'\mathbf{M}_{\widehat{P}} \mathbf{X}_{i}\right)^{-1}\sum_{i=1}^N \mathbf{X}_{i}' \mathbf{M}_{\widehat{P}} \mathbf{y}_{i}.
\end{align}
Some remarks are in order. We start by justifying Step (ref). As already pointed out, the number of factors required to capture the nonlinearity of $\mathcal{G}_i(\cdot)$ is likely large. We also know that $\widehat{\*f}_t \overset{p}{\to} \overline{\mathcal{G}}(\*f_t)$, where $\overline{\mathcal{G}}(\cdot): \mathbb{R}^m \to \overline{\mathcal{G}}\left(\mathbb{R}^m\right) \subseteq \mathbb{R}^{d+1}$. Hence, $\widehat{\*f}_t$ effectively recovers a low-dimensional projection of the space spanned by $\mathcal{G}_i(\*f_t)$. Applying the sieve basis to $\widehat{\*f}_t$ re-expands this projection into a high-dimensional function space, which can be made rich enough to capture all relevant factor transformations. Had $\*f_t$ been observed, we could have formed $\*p(\*f_t)$, and under mild conditions $\mathcal{G}_i(\*f_t)$ would be well approximated by $\+\alpha_{\mathcal{G}_i}' \*p(\*f_t)$ for some $\+\alpha_{\mathcal{G}_i}\in \mathbb{R}^{K \times (d+1)}$. The coefficient matrix $\+\alpha_{\mathcal{G}_i}$ need not be known but will be absorbed in the estimation of the model. In Lemma (ref) of the appendix, we show that the same approximation argument applies also to $\*p(\widehat{\*f}_t)$; that is, $\+\alpha_{\mathcal{G}_i}'\*p(\widehat{\*f}_t)$ can be used as a proxy for $\mathcal{G}_i(\*f_t)$. Hence, $\*p(\widehat{\*f}_t)$ can be consistent in this sense even if $\widehat{\*f}_t$ is not. This last result seems new to the literature and is therefore interesting on its own.
As chen2007 points out, the approximation properties of a sieve typically depend on the support, smoothness, and other functional characteristics of the target function. However, these features are often unknown in practice. In such cases, the specific choice of sieve space is less important, provided it achieves the desired rate of approximation. Commonly used sieve methods with well-understood approximation properties include finite-dimensional linear sieves such as splines, wavelets, power series and Fourier series, all of which achieve comparable approximation rates (see chen2007, and bel15).
An alternative to the method of sieves is kernel estimation. However, the method of sieves has (at least) three advantages that in our opinion makes it relatively more attractive (see chen2007, for a comprehensive review on sieve estimation). First, the method is computationally very simple, which means that it preserves the empirical appeal of CCE. Second, the approach nests the conventional linear factor specification, implying that it can be used to construct formal statistical tests of the presence of factor nonlinearities. In the empirical illustration of Section (ref), we elaborate on this point. Finally, compared to fully nonparametric approaches, the method of sieves enables faster rates of convergence.
The pooled estimator considered in Step (ref) is our version of pesaran2006's pesaran2006 pooled CCE (CCEP) estimator. This estimator is based on within pooling; that is, we sum the data over the cross-section before taking the ratio. pesaran2006 also introduces a mean-group CCE (CCEMG) estimator, which is based on between pooling. Here, one would first compute $N$ estimators obtained by applying CCE to each cross-sectional unit, and then average those. Both estimators -- CCEP and CCEMG -- usually show good finite sample properties. However, within pooling is relatively more efficient; hence, our choice of estimator. In the empirical illustration of Section (ref), we compare SCCE to CCEP and CCEMG.
Assumptions and Asymptotic Results
assumption[Idiosyncratic errors]
\leavevmode
\begin{enumerate}[label=(\alph*)]
• For each $i$, $\{(\varepsilon_{i,t}, \*v_{i,t}): t \geq 1\}$ is a strictly stationary and $\alpha$-mixing process with mixing coefficients $\alpha_{i}(j)$ satisfying $\sum_{j=1}^\infty j^2 \alpha_i(j)^{\frac{\eta}{4+\eta}} \leq C$ for some $\eta > 0$. The process is also independent across $i$ with $\mathbb{E}(\varepsilon_{i,t}) = 0$, $\mathbb{E}(\*v_{i,t}) = \*0_{d \times 1}$, $\max_{t=1,\dots,T} \max_{i=1,\dots,N} \mathbb{E}(\|\*v_{i,t}\|^{4+\eta}) \leq C$ and $\max_{t=1,\dots,T} \max_{i=1,\dots,N}\mathbb{E}(|\varepsilon_{i,t}|^{4 + \eta}) \leq C$.
• Let $\sigma^2_{\varepsilon, i} \equiv \lim_{T \to \infty} \frac{1}{T} \mathbb{E}(\+\varepsilon_i'\+\varepsilon_i) \in \mathbb{R}$ and $\+\Sigma_{v,i} \equiv \lim_{T \to \infty} \frac{1}{T} \mathbb{E}(\*V_i' \*V_i) \in \mathbb{R}^{d \times d}$. The following applies to these variances: $\inf_{i=1, \dots, N} \sigma_{\varepsilon, i}^2 > 0$, $\sup_{i=1, \dots, N} \sigma_{\varepsilon, i}^2 \leq C$, $\sigma^2_{\varepsilon} = \lim_{N \to \infty} \frac{1}{N} \sum_{i=1}^N \sigma_{\varepsilon, i}^2 \leq C$, $\inf_{i=1, \dots, N}\lambda_{\min}(\+\Sigma_{v,i}) > 0$, $\sup_{i=1, \dots, N} \|\+\Sigma_{v,i} \| \leq C$ and $\+\Sigma_v \equiv \lim_{N \to \infty} \frac{1}{N} \sum_{i=1}^N \+\Sigma_{v,i} \in \mathbb{R}^{d \times d}$ is positive definite. Moreover, the covariance matrices $\+\Omega_{\varepsilon, i} \equiv \mathbb{E}(\+\varepsilon_i \+\varepsilon_i')\in \mathbb{R}^{T \times T}$ and $\+\Omega_{v, i} \equiv \mathbb{E}(\*V_{i}\*V'_{i})\in \mathbb{R}^{T \times T}$ have eigenvalues bounded away from zero and infinity.
\end{enumerate}
assumption[Common factors]
$\{\*f_t: t \geq 1\}$ is a strictly stationary and $\alpha$-mixing process with mixing coefficient $\alpha_0(j)$ satisfying $\sum_{j=1}^{\infty} j^2 \alpha_0(j)^{\frac{\eta}{4+ \eta}} \leq C$ with $\eta > 0$. Also, $\sup_{t=1, \dots, T}\mathbb{E}(\|\*f_{t}\|^{4+\eta}) \leq C$ and $\+\Sigma_F \equiv \lim_{T \to \infty} \frac{1}{T} \mathbb{E}(\*F'\*F)\in \mathbb{R}^{m \times m}$ is positive definite.
assumption[Independence] $\*f_t$, $\varepsilon_{i, s}$ and $\*v_{j, l}$ are mutually independent for all $t$, $i$, $j$, $s$ and $l$.
Assumptions (ref)--(ref) are similar to Assumptions 1--2 in pesaran2006. Hence, we will only discuss main differences and restrictions here. Assumptions (ref)(ref) and (ref) impose strict stationarity on $\varepsilon_{i,t}$, $\*v_{i,t}$ and $\*f_t$, and specify $\alpha$-mixing conditions as in su12 with mixing rates vanishing to zero sufficiently fast. This implies that $\alpha_i(j) = o(j^{-(3+ \frac{12}{\eta})})$, assuming that $\alpha_i(j)$ is approximately a $p$-series. The decay of the mixing rates depends inversely on the parameter $\eta$; the smaller is $\eta$, the faster the mixing rates vanish. As pointed out by su12, $\alpha$-mixing is not particularly restrictive. The Monte Carlo results reported in the appendix suggest that Assumptions (ref)(ref) and (ref) can be relaxed to allow for weak cross-sectional correlation in $\varepsilon_{i,t}$ and $\*v_{i,t}$, and non-stationarity in $\*f_t$. Assumption (ref) is necessary and cannot be relaxed.
To ensure accurate approximation of $\mathcal{G}_i(\*f_t)$ using our proposed method, we need to impose smoothness conditions with respect to $\*f_t$. Let $\mathcal{R}_f \subseteq \mathbb{R}^m$ denote the support of $\*f_t$. Classical series methods often require $\mathcal{R}_f$ to be bounded, such that, for instance, $\mathcal{R}_f=[0,1]^m$, which can be restrictive. We therefore follow su12, and chen2005, and allow the support to be unbounded, taking $\mathcal{R}_f = \mathbb{R}^m$ and using the following weighted sup-norm metric:
align[align omitted — 271 chars of source]
where $\omega \geq 0$ and $z(\*f) \equiv [1+\|\*f\|^{2}]^{-\omega}$ is a weight function. When $\omega = 0$, (ref) reduces to the usual sup-norm, which would suffice if $\mathcal{R}_f$ would indeed be a bounded subset of $\mathbb{R}^m$ (see, for example, lee16, su12, chen2005, and chen2007).
A typical smoothness class of functions is the H\"older-smooth -- or “$p$-smooth” -- class of functions. This class is particularly popular in econometrics as functions belonging to this class can be well approximated by the method of sieves (see chen2007, for a discussion). The H\"older space $\Lambda^\lambda(\mathcal{R}_f)$ with smoothness $\lambda > 0$ is a space of functions $g: \mathcal{R}_f \to \mathbb{R}$ in which the first $\lfloor \lambda \rfloor $ derivatives are bounded and the last $\lfloor\lambda \rfloor$ derivatives are H\"older continuous with exponent $r = \lambda - \lfloor \lambda \rfloor \in (0,1]$. The H\"older norm is given by
align[align omitted — 246 chars of source]
We borrow the following definition of H\"older-smooth functions from chen2007, and chen2005.
definition[H\"older-smooth functions]
Let $\Lambda^\lambda(\mathcal{R}_f, \omega)$ denote a weighted H\"older space of functions $g: \mathcal{R}_f \to \mathbb{R}$ such that $g(\cdot)[1 + \|\cdot\|^2]^{-\frac{\omega}{2}}$ is in $\Lambda^\lambda(\mathcal{R}_f)$. A weighted H\"older ball with radius $r$ is given by $\Lambda_r^\lambda(\mathcal{R}_f, \omega) \equiv \{g \in \Lambda^\lambda(\mathcal{R}_f, \omega): \|g(\cdot)[1 + \|\cdot\|^2]^{-\frac{\omega}{2}}\|_{\Lambda^\lambda} \leq r $\}. A function $g(\cdot)$ is said to be $H(\lambda, \omega)$-smooth on $\mathcal{R}_f$ if it belongs to the weighted H\"older ball $\Lambda_r^\lambda(\mathcal{R}_f, \omega)$ for some $\lambda>0$, $r \in (0,\infty)$ and $\omega \geq0$.
If $\omega=0$, the weighted H\"older ball reduces to the standard H\"older ball $\Lambda_r^\lambda(\mathcal{R}_f) \equiv \{g \in \Lambda^\lambda(\mathcal{R}_f): \|g(\cdot)\|_{\Lambda^\lambda} \leq r \}$. As before, if $\mathcal{R}_f$ would be a bounded subset of $\mathbb{R}^m$ this would suffice. Here, however, $\mathcal{R}_f = \mathbb{R}^m$, and a standard H\"older ball would exclude any function whose magnitude grows without bound. Hence, even simple linear interactive effects with deterministic factors, such as $g_i(\*f_t) = \+\gamma_i'\*f_t=\*1_{m}'\*f_t$, would be ruled out, which is clearly not desirable in our set-up.
assumption[Function class]
Define the set $\mathcal{S}\equiv \{g_i(\cdot),G_{j,i}(\cdot): i=1, \dots, N, j=1, \dots, d\}$, where the mappings $g_i(\cdot)$ and $G_{j,i}(\cdot)$ are defined as before.
\begin{enumerate}[label=(\alph*)]
• For any $\widetilde{g}_i(\cdot)\in \mathcal{S}$, $\widetilde{g}_i(\cdot) \in \Lambda_r^{\lambda_i}(\mathcal{R}_f, \omega_i)$ for some $\lambda_i > 0$, $\omega_i \geq 0$.
• $\sup_{t=1,\hdots,T}\int_{\mathbb{R}^m} [1+\|\*f_t\|^{2}]^{\overline{\omega}} d\mu(\*f_t) < \infty$ for some $\overline{\omega} > \sup_{i=1,\dots, N}\left(\omega_i + \lambda_i\right)$, where $d\mu(\*f_t) = w(\*f_t)d\*f_t$ and $w(\*f_t)$ is the probability density function of $\*f_t$.
• For any $\widetilde{g}_i(\cdot) \in \Lambda_r^{\lambda_i}(\mathcal{R}_f, \omega_i)$, there is a function $\Pi_{\infty K} \widetilde{g}_i(\cdot)\equiv \+\alpha_{\widetilde{g}_i}' \*p(\cdot)$ in the sieve space $\mathfrak{G}_K \equiv \{h(\cdot) = \*a'\*p(\cdot)\}$ such that the approximation error satisfies $\| \widetilde{g}_i(\cdot) - \Pi_{\infty K}\widetilde{g}_i(\cdot) \|_{\infty, \overline{\omega}} = O\left(K^{-\frac{\lambda_i}{m}}\right)$.
• $\mathbb{E}\left[\widetilde{g}_i(\*f_t)\right] = 0$.
\end{enumerate}
assumption[Rank condition]
$\overline{\mathcal{G}}(\cdot): \mathbb{R}^m \to \overline{\mathcal{G}}\left(\mathbb{R}^m\right) \subseteq \mathbb{R}^{d+1}$ is $H(\lambda, \omega)$-smooth and injective. Further, $\nabla \overline{\mathcal{G}}\left(\*f_t\right) \in \mathbb{R}^{(d+1) \times m}$ exists and is of full column-rank $m$ w.p.1, such that for all $\*f_t \in \mathbb{R}^{m}$, $\operatorname{rank} \left(\nabla \overline{\mathcal{G}}\left(\*f_t\right)\right) = m \leq d+1$.
Assumption (ref)(ref) specifies a weighted smoothness condition for each function in the set $\mathcal{S}$. Recall that $\mathcal{G}_i(\*f_t)\equiv[g_i(\*f_t) + \+\beta'\*G_i(\*f_t), \, \*G_i(\*f_t)']'\in \mathbb{R}^{d+1}$, which we can write equivalently as
align[align omitted — 222 chars of source]
This decomposition shows that $\mathcal{G}_i(\*f_t)$ consists of a nonlinear component $[g_i(\*f_t), \*G_i(\*f_t)']'$ and a linear transformation of $\*G_i(\*f_t)$, $[\+\beta'\*G_i(\*f_t), \*0_{d\times 1}']'$. Because linear transformations and constant shifts preserve H\"older smoothness and approximation rates, both terms in (ref) belong to the same smoothness class. Consider the function $\widetilde{g}_i(\cdot) \in \mathcal{S}$. Equation (ref) implies that any assumption placed on $\widetilde{g}_i(\cdot)$ applies also to $\mathcal{G}_i(\cdot)$, and vice versa. We will make use of this feature throughout the paper. Assumption (ref)(ref) restricts the tails of the marginal density of $\*f_t$. This assumption is necessary since we allow $\mathcal{R}_f$ to be unbounded. Assumption (ref)(ref) quantifies the error due to the approximation of the $H(\lambda_i, \omega_i)$-functions by the $K$-dimensional linear sieves in $\mathfrak{G}_K$. Assumption (ref)(ref) is a standard identification condition that holds whenever the model in (ref) and (ref) includes an intercept.
Assumption (ref) is our version of condition (21) in pesaran2006. Three things are notable here. First, provided that $\nabla \overline{\mathcal{G}}(\*f_t)$ is of full column rank $m \leq d+1$, the inverse function theorem ensures that $\overline{\mathcal{G}}(\cdot)$ is locally invertible into its image, which is sufficient for our analyses. Second, the smoothness of $\overline{\mathcal{G}}(\cdot)$ follows from the regularity conditions imposed on each function in the set $\mathcal{S}$ in Assumption (ref). Similarly to (ref), $\overline{\mathcal{G}}(\*f_t)$ inherits the differentiability and smoothness of its elements. Third, in the special case of linear interactive effects in which $g_i(\*f_t)\equiv \+\gamma_i'\*f_t$ and $\*G_i(\*f_t)\equiv \+\Gamma_i'\*f_t$, we can write $\overline{\mathcal{G}}(\*f_t)=[\overline{\+\gamma}+\overline{\+\Gamma}\+\beta, \, \overline{\+\Gamma}]'\*f_t\equiv \overline{\*C}'\*f_t$, such that $\nabla \overline{\mathcal{G}}(\*f_t)=\overline{\*C}'$. In this case, Assumption (ref) reduces to condition (21) in pesaran2006.
assumption[Basis functions]
Define the set of sieve matrices $\mathcal{P} \equiv \{\widetilde{\*P}, \widehat{\*P}\}$, where $\widetilde{\mathbf{P}} = \left[\widetilde{\*p}\left(\*f_1\right), \dots, \widetilde{\*p}\left(\*f_T\right)\right]' = \left[\*p\left(\overline{\mathcal{G}}\left(\*f_1\right)\right), \dots, \*p\left(\overline{\mathcal{G}}\left(\*f_T\right)\right) \right]'$ and $\widehat{\*P}$ is as before.
\begin{enumerate}[label=(\alph*)]
• For any $\*P \in \mathcal{P}$, uniformly in $K$ and $T$, there exist constants $0 < \underline{C} < \overline{C} < \infty$, such that $\underline{C} \leq \lambda_{\min}[{\mathbb{E}(\frac{1}{T}\*P' \*P})]\leq \lambda_{\max}[\mathbb{E}(\frac{1}{T}\*P' \*P)] \leq \overline{C}$.
• $\sup_{k=1, \dots, K} \mathbb{E}[|p_k(\cdot)|^{8 + \eta}]\leq C$, where $\eta>0$ is as in Assumption (ref) and $p_k(\cdot)$ is the $k$-th sieve basis function of $\*P \in \mathcal{P}$.
• For large enough $K$, there exists a sequence of constants $\zeta_0(K)$ and $\zeta_1(K)$ satisfying $\sup_{\*f}\|\*p(\*f)\|\leq \zeta_0(K)$, $\sup_{\*f}\|\nabla\*p(\*f)\|\leq \zeta_1(K)$ with $\zeta_0(K)\geq 1, \, \zeta_1(K) \geq 1$. Further, $\sqrt{NT}K^{-\frac{\lambda_i}{m}}\to 0$ with $\lambda_i > 0$, $\frac{\sqrt{T}K^{\frac{1}{4}}}{\sqrt{N}}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)^{2} \to 0$, $\sqrt{\frac{K}{T}}\zeta_{0}(K)^{\frac{3}{2}}\to 0$, as $N,T\to \infty$.
\end{enumerate}
assumption[Gram matrices of sieve coefficients]
Define $\+\Sigma_{\alpha_g} \equiv \frac{1}{N}\sum_{i=1}^N \+\alpha_{g_i}\+\alpha_{g_i}' \in \mathbb{R}^{K\times K}$ and $\+\Sigma_{\alpha_G} \equiv \frac{1}{N}\sum_{i=1}^N \+\alpha_{G_i}\+\alpha_{G_i}' \in \mathbb{R}^{K \times K}$. Then, there exist constants $0 < \underline{C} < \overline{C} < \infty$, such that for all $j\in \{\alpha_g, \alpha_G\}$, $\liminf_{N\to \infty}\lambda_{\min}\left(\+\Sigma_{j}\right) \geq \underline{C}$ and $\limsup_{N\to \infty} \lambda_{\max}\left(\+\Sigma_{j}\right) \leq \overline{C}$.
Assumption (ref)(ref) is standard in the sieve estimation literature (see, for example, newey97, or chen2007). As usual in this literature, $K$ is allowed to grow with the sample size to balance the bias-variance trade-off. Increasing $K$ implies decreasing bias, but increasing variance, and vice versa (see, for example, bel15, su12, or chen2007). Assumption (ref)(ref) is not uncommon (see, for example, su12). Assumption (ref)(ref) restricts how fast $K$ may grow, how large or steep the sieve basis can be, and limits the relative growth of $N$ and $T$. The sequences $\zeta_{0}(K)$ and $\zeta_{1}(K)$ bound the size of the basis vector and its first derivatives, respectively. Bounding $\zeta_0(K)$ avoids overly large series terms, while bounding $\zeta_1(K)$ ensures that small input errors do not explode after inserting $\widehat{\*f}_t$. In practice, many common bases such as splines and Fourier series satisfy $\zeta_0(K) = O(\sqrt{K})$ and $\zeta_1(K)=O\left(K^{\frac{3}{2}}\right)$, while others such as Legendre polynomials or power series can have $\zeta_0(K) = O(K)$ and $\zeta_1(K)=O\left(K^3\right)$ (see, for example, bel15, and newey97).
The condition $\sqrt{NT}K^{-\frac{\lambda_i}{m}}\to 0$ keeps the approximation error introduced by the series estimation asymptotically negligible as $K$ grows and is common in the literature (see, for example, su12). $\frac{\sqrt{T}K^{\frac{1}{4}}}{\sqrt{N}}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)^{2}\to 0$ controls higher-order effects of using $\widehat{\*f}_t$ as input for the basis functions. It is our analogue of the CCE condition $\frac{T}{N^2}\to 0$ when $\*g_i\left(\*F\right)$ and $\*G_i\left(\*F\right)$ are linear (see pesaran2006). $\sqrt{\frac{K}{T}}\zeta_0(K)^{\frac{3}{2}}\to 0$ disciplines the growth of basis terms relative to the time dimension. These conditions can be used to determine an admissible corridor for the sieve dimension, $K$, such that $\widehat{\boldsymbol{\beta}}_{SCCE}$ is consistent. Suppose that $\zeta_0(K) = O\left(K^a\right)$ and $\zeta_1(K) = O\left(K^b\right)$ with $a, b \geq 0$. Then, the condition that $\sqrt{NT}K^{-\frac{\lambda_{i}}{m}}=o(1)$ determines the lower bound, such that $\left(NT\right)^{\frac{m}{2\lambda_{i}}}<<K$, while $\sqrt{\frac{T}{N}}K^{\frac{1}{4}}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)^{2}=o(1)$ and $\sqrt{\frac{K}{T}}\zeta_{0}(K)^{\frac{3}{2}}=o(1)$ determine the upper bound, such that $K<<\min\left\{ T^{\frac{1}{1+3a}},\left(\frac{N}{T}\right)^{\frac{2}{1+6a+8b}}\right\}$. Hence, $K$ should satisfy $\left(NT\right)^{\frac{m}{2\lambda_{i}}}<<K<<\min\left\{ T^{\frac{1}{1+3a}},\left(\frac{N}{T}\right)^{\frac{2}{1+6a+8b}}\right\}$. Importantly, we do not require the optimal rate of convergence as in stone to achieve consistency of $\widehat{\+\beta}_{SCCE}$. Any choice of $K$ in the admissible corridor defined above suffices.
Assumption (ref) imposes uniform lower and upper bounds on the eigenvalues of the Gram matrices of the sieve coefficient vectors $\+\alpha_{g_i}$ and $\+\alpha_{G_i}$. Because in our specification the factor loadings are absorbed into these vectors, $\+\Sigma_j$ acts as the second-moment matrix of the corresponding factor loadings. Assumption (ref) guarantees that this matrix remains well-conditions as $N\to \infty$.
theoremUnder Assumptions (ref)--(ref), as $N,\,T\to\infty$,
\begin{eqnarray}
\sqrt{NT}(\widehat{\+\beta}_{SCCE} - \+\beta) \overset{d}{\to} \mathcal{N}(\*0_{d\times 1},\+\Sigma_v^{-1}\+\Theta \+\Sigma_v^{-1}),
\end{eqnarray} where
\begin{align}
\+\Theta \equiv \lim_{N,T\to\infty}\frac{1}{NT}\sum_{i=1}^{N}\mathbb{E}\left(\mathbf{V}_{i}'\boldsymbol{\Omega}_{\varepsilon,i}\mathbf{V}_{i}\right)
\end{align} with $\boldsymbol{\Omega}_{\varepsilon,i} \equiv \mathbb{E}\left(\+\varepsilon_i\+\varepsilon_i'\right)$.
Theorem (ref), whose proof can be found in the appendix, establishes that our proposed SCCE estimator is consistent, and that the rate of convergence is the usual parametric one. The estimator is also asymptotically normally distributed, which means that it supports standard inference based on Student-$t$ and Wald tests. However, this requires a consistent estimator of the covariance matrix $\+\Sigma_v^{-1}\+\Theta \+\Sigma_v^{-1}$. A naturally consistent estimator of $\+\Sigma_v$ is given by $\widehat{\+\Sigma}_{v}\equiv\frac{1}{NT}\sum_{i=1}^N \widehat{\*V}_i'\widehat{\*V}_i$, where $\widehat{\*V}_i = [\widehat{\*v}_{i,1}, \dots, \widehat{\*v}_{i,T}]' \equiv \*M_{\widehat{P}}\*X_i$. To estimate $\+\Theta$, analogously to pesaran2006, one can use the following heteroskedasticity and autocorrelation consistent (HAC) estimator:
align[align omitted — 160 chars of source]
where $\widehat{\+\Theta}_l\equiv \frac{1}{NT} \sum_{i=1}^N \sum_{t=l+1}^T \widehat{\varepsilon}_{i,t}\widehat{\varepsilon}_{i,t-l}\widehat{\*v}_{i,t}\widehat{\*v}_{i,t-l}'$ with $\widehat{\+\varepsilon}_i = [\widehat{\varepsilon}_{i,1}, \dots, \widehat{\varepsilon}_{i,T}]' \equiv \*M_{\widehat{P}}(\*y_i - \*X_i \widehat{\+\beta}_{SCCE})$ and $L$ being the window size. Consistency of this last estimator requires $L \to \infty$ such that $\frac{L}{T}\to 0$.
As an alternative to the above HAC estimator, one could use bootstrapping, which tend to work better in small samples. We resample $\*Z_i^* \equiv [\*y_i^*,\*X_i^*]$ from $\{\*Z_1,\ldots, \*Z_N\}$ with replacement, as proposed by staus_devos. For each bootstrap sample, we compute $\widehat{\+\beta}_{SCCE}^*$, which is $\widehat{\+\beta}_{SCCE}$ with $[\*y_i,\*X_i]$ replaced by $[\*y_i^*,\*X_i^*]$. This procedure is repeated many times. Confidence intervals can then be constructed from the resulting bootstrapped distribution of $\widehat{\+\beta}_{SCCE}$. In the empirical illustration of Section (ref), we focus on the bootstrapped confidence intervals, although we also report HAC-based results.
Monte Carlo Simulations
A large-scale Monte Carlo study was conducted to evaluate the small-sample properties of the new estimator. The full set of results is too numerous to report here in full. Consequently, in this section, we focus on a representative subset. The full set of results is provided in the appendix. The data generating process is given by a restricted version of (ref) and (ref) with $m=d=2$ and $\+\beta = [1, 1]'$. All elements of the idiosynchratic errors $\varepsilon_{i,t}$ and $\*v_{i,t}$ are drawn independently from $\mathcal{N}(0,1)$.
Denote by $G_{s,i}\left(\*f_t\right)\in \mathbb{R}$ the $s$-th row of $\*G_i\left(\*f_t\right)$. We consider the following two experiments:
E1 (Nonlinear factor structure). We generate
align[align omitted — 392 chars of source]
for $s=1, 2$, where the elements of $\*f_t = [f_{1,t}\, f_{2,t}]'\in \mathbb{R}^2$ and $\gamma_{1,i}, \gamma_{2,i}, \gamma_{3,i}, \Gamma_{1s,i}, \Gamma_{2s,i}$ are drawn independently from $\mathcal{N}(0,1)$, whereas $\Gamma_{3s,i}, \Gamma_{4s,i}$ are drawn from $\mathcal{N}(1,1)$.
E2 (Linear factor structure). In this experiment,
align[align omitted — 131 chars of source]
where $\+\gamma_i = [\gamma_{1,i}, \gamma_{2,i}]' \in \mathbb{R}^2$ and $\+\Gamma_{s,i} = [\Gamma_{1s,i}, \Gamma_{2s,i}]' \in \mathbb{R}^2$ for all $s=1,2$, where $\*f_t$, $\gamma_{1,i}$, $\gamma_{2,i}$, $\Gamma_{s1,i}$ and $\Gamma_{s2,i}$ are as in E1.
A word on implementation: We use univariate cubic spline polynomials as the sieve base. Denote by $\widehat{f}_{r,t} \in \mathbb{R}$ the $r$-th element of $\widehat{\*f}_t$. Define
align[align omitted — 224 chars of source]
where $(\widehat{f}_{r,t} - \theta_{r,j})_+^{3} = \max\{(\widehat{f}_{r,t} - \theta_{r,j})^3, 0\}$ and $\theta_{r,1},\dots, \theta_{r,J}$ are knot values computed as the $\frac{j}{J+1}$-th empirical quantile of $\widehat{f}_{r,1},\dots, \widehat{f}_{r,T}$. In terms of the notation of Section (ref), we have $\*p(\widehat{\*f}_{t}) = [\*p(\widehat{f}_{1,t})',\dots,\*p(\widehat{f}_{d+1,t})' ]'\in \mathbb{R}^{(d+1)(4+J)}$. Since $d=2$ in this section, the sieve basis contains $K = 3(4 + J)$ terms in total. The number of knots is set to $J = C \lfloor T^{\frac{1}{4}} \rfloor$, which makes $K$ increase with $T$ while ensuring that the sieve dimension remains well within the theoretical growth restrictions of Assumption (ref)(ref). The constant $C$ had little effect on the results. We therefore focus on the case when $C=1$ and put the rest of the results in the appendix.\footnote{The appendix also reports results for $J = \lfloor T^{\frac{1}{3}} \rfloor$, $J = \lfloor T^{\frac{1}{5}} \rfloor$ and $J = \lfloor T^{\frac{1}{10}} \rfloor$.}
figure[figure omitted — 1,362 chars of source]
Figure (ref) presents absolute bias and root mean squared error (RMSE) based on 1,000 replications. According to the results, the SCCE estimator performs well even in samples as small as $N=T=20$, which is desirable given that such sample sizes are not uncommon in practice. Performance improves as both $N$ and $T$ increase, which corroborates Theorem (ref) and the consistency of the estimator. Performance is good in both experiments, although it is generally better in E2 than in E1, which is to be expected given the relatively simple, linear factor structure in E2. The fact that SCCE works well regardless of the factor structure being considered is reassuring because it means that it can be applied without prior knowledge in this regard.
Empirical Illustration
We apply our method to explore possible reasons for the rise in wage inequality between high-skilled and low-skilled workers in U.S. manufacturing over recent decades. Following the approach of voig, we investigate whether these persistent differences in skill premium can be attributed to so-called “intersectoral technology skill complementarity (ITSC)”. In the labor economics literature, ITSC refers to the idea that skilled labor in intermediate production stages complements skilled labor in later stages, such as final processing or product integration. This complementarity is intersectoral because it takes place between different industries connected through the supply chain. For example, skilled workers in the software industry can increase the productivity of downstream sectors like electronics, which then raises the demand for skilled workers in those sectors. In this way, skill requirements are transmitted across sectors, creating a multiplier effect; when one sector uses more skilled labor, it increases the demand for skilled labor in related sectors. As a result, ITSC can raise the wage premium for skilled workers and contribute to a widening wage gap.
voig employs unbalanced data on 358 U.S. manufacturing sectors covering the years 1958-2005. Recently, yin21, and juo22 applied versions of the CCEP and CCEMG estimators to a balanced subset of the same dataset that covers 313 sectors. Both explore the impact of ITSC on wage inequality. Our interest in these studies stems from the fact that they assume the unobserved factors to enter the model linearly, which might not be appropriate in this context. As voig points out, technological changes in one sector can indirectly influence skill demand in related industries through intersectoral linkages, often in nonlinear ways. One example is the effect of automation on labor demand; initially, new technologies may raise demand for high-skilled workers, but as automation expands, mid-skilled jobs decline at an accelerating rate, leading to wage polarization across sectors. In fact, ace22 show that automation-driven task displacement is one of the central drivers of U.S. wage inequality, operating in ways that standard linear or factor-augmented models fail to capture. Ignoring such nonlinear adjustment mechanisms can lead to omitted variables bias and hence misleading conclusions.
Our proposed SCCE approach allows us to capture nonlinearities in the unobserved factors and is therefore able to handle concerns like those just mentioned. It should therefore be well-suited for studying the impact of ITSC on wage inequality. We use the same balanced subsample of voig's voig dataset as yin21, and juo22.\footnote{The dataset is available from the website of Journal of Business & Economic Statistics at: https://www.tandfonline.com/doi/full/10.1080/07350015.2019.1623044\#d1e18402.} Hence, in this illustration, $N = 313$ and $T = 48$, which according to the Monte Carlo results reported in Section (ref) should be more than enough to ensure good performance. The model is also the same, except that we do not presume a linear factor structure. It is given by
align[align omitted — 228 chars of source]
where $\frac{w_{L,i,t}}{w_{H,i,t}}$ is the relative wage of low-skilled to high-skilled workers. In our analysis, we focus on two key regressors. The first is $\ln(\sigma_{i,t})$, which captures input skill intensity and serves as a proxy for ITSC. This measure -- constructed by voig -- is defined as the weighted average share of white-collar workers involved in the production of industry $i$-s intermediate manufacturing inputs.\footnote{The weights are derived from input-output expenditure data.} Accordingly, one of our primary coefficients of interest is $\beta_{1,i}$. The second regressor, $\ln\left(\frac{H_{i,t}}{L_{i,t}}\right)$, captures the relative demand for high- versus low-skilled labor, with $\beta_{2,i}$ being the corresponding coefficient. The vector $\*z_{i,t} \in \mathbb{R}^6$ contains some control variables that the previous literature has deemed relevant (see voig), and is given by
align[align omitted — 652 chars of source]
where $k_{i,t}^{\text{equip}}$ is real capital equipment per worker, $(OCAM/K)_{i,t}$ is the sectoral share of office, computing and accounting equipment, $(HT/K)_{i,t}$ is the sectoral share of high-technology capital, $R\&D_{\text{lag},i,t}$ is the lagged research and development (R&D) intensity, and $OS_{i,t}^{\text{broad}}$ and $OS_{i,t}^{\text{narr}}$ are broad and narrow measures of outsourcing, respectively. Sector-specific, time-invariant effects, $\alpha_i$, are accounted for by including a vector of ones in the sieve base. As before, $g_i(\*f_t)$ captures unobserved heterogeneity where the functional form of $g_i(\cdot)$ is essentially unrestricted. This is the main difference when compared to yin21, and juo22, who -- as pointed out earlier -- assume that $g_i(\*f_t) = \+\gamma_i'\*f_t$ even though there are reasons to suspect this restriction to be violated.
As for the number of knots, $J$, we use the same specification as in Section (ref) with $C=1, \dots, 5$. We focus on the results for the specification with $C=1$, and just briefly discuss the results for the other values as a robustness check. Before we come to the estimation results, we investigate the properties of the estimated factors. In Figure (ref), we plot $\widehat{\*f}_t = [\overline{\ln (\sigma_t)}, \ \overline{\ln \left(\frac{H_t}{L_t}\right)}, \ \overline{\*z}_t]'$ over time. All factor estimates appear to be trending, suggesting that they are not stationary. This observation is supported by a formal augmented Dickey-Fuller test, the results of which are reported in the appendix. The factors are therefore likely unit root non-stationary, which is not permitted under our assumptions. We therefore proceed to transform our data by taking first differences before applying our estimation procedure.
figure[figure omitted — 1,015 chars of source]
Similarly to yin21, we distinguish between a “narrow” and a “full” model. The full model is the one in (ref). The narrow model is the same, except that the vector of controls, $\*z_{i,t}$, is excluded. We estimate these two models by using three estimators; our own, SCCE, as well as the CCEP and CCEMG estimators of pesaran2006. The reason for including the last two is twofold. First, we are interested to see to what extent we can replicate the results of yin21, and juo22, which are based on the same two estimators.\footnote{yin21, and juo22 do not transform their data by taking first differences. We use the same approach for implementing the CCEP and CCEMG estimators to ensure comparability with their results. Unreported results suggest that the results are not affected much by taking differences, which is partly expected given that CCE has been shown to be robust to non-stationarity (see kap2011).} Second, because the linear factor specification employed by the CCEP and CCEMG estimators is nested within our more general nonlinear model, we can assess whether linearity is in fact met.
The estimation results are presented in Table (ref). The first thing to note is that the CCEP and CCEMG results do in fact replicate those reported by yin21, and juo22 for both models, which is reassuring. Next we compare these results based on assuming a linear factor structure with those obtained by using SCCE. The overall picture is that failing to account for nonlinearities leads to an overestimation of the effect of ITSC on wage inequality in absolute terms. Consider $\ln(\sigma_{i,t})$. In the narrow model, the estimated effect of this regressor decreases in absolute value from -0.59 (-0.61) for CCEMG (CCEP) to -0.34 for SCCE, which means that the estimates with linearity imposed are about 40% larger in absolute value than that without. The conclusions for the full model are basically the same. Because the estimated sign is always negative, our results support the conventional wisdom in the literature that ITSC causes increased wage inequality. The results for $\ln \left(\frac{H_{i,t}}{L_{i,t}}\right)$ are qualitatively the same but go in the other direction. The estimated sign is positive instead of negative with CCEP and CCEMG underestimating the effect of relative demand on wage inequality when compared to SCCE. The difference relative to SCCE is about the same in magnitude compared to before. However, it is statistically more pronounced as now the CCEMG and CCEP estimates are no longer covered by the SCCE confidence intervals. Nevertheless, since the estimated sign is always positive, the conclusion is in line with the corresponding literature; namely, that high-skilled and low-skilled labor are substitutes.
landscape\begin{table}[h]
\begin{threeparttable}
\caption{Estimation results.}
\begin{tabular}{lc c c c c c c c c}
\toprule
& \multicolumn{4}{c}{Narrow model} & \multicolumn{4}{c}{Full model} \\
\cmidrule(lr){2-5} \cmidrule(lr){6-9}
Regressor & SCCE & CCEP & CCEMG & & SCCE & CCEP & CCEMG & \\
\midrule
$\ln(\sigma_{i,t})$ & -0.34 & -0.61 & -0.59 & & -0.33 & -0.52 & -0.71 & \\
& (-0.67; -0.03) & (-0.97; -0.25) & (-1.01; -0.15) & & (-0.70; 0.00) & (-0.88; -0.17) & (-1.17; -0.24) & \\
$\ln\left(\frac{H_{i,t}}{L_{i,t}}\right)$ & 0.56 & 0.40 & 0.36 & & 0.61 & 0.48 & 0.45 & \\
& (0.50; 0.61) & (0.33; 0.46) & (0.33; 0.38) & & (0.55; 0.66) & (0.42; 0.55) & (0.43; 0.48) &\\
$z_{1,i,t}$ & & & & & -0.03 & -0.19 & -1.64 &\\
& & & & & (-0.47; 0.51) & (-0.44; 0.06) & (-2.74; -0.63) &\\
$z_{2,i,t}$ & & & & & 3.15 & 1.23 & -3.29 & \\
& & & & & (0.40; 6.02) & (-0.37; 2.78) & (-6.66; 0.10) & \\
$z_{3,i,t}$ & & & & & -1.03 & 0.25 & 2.28 & \\
& & & & & (-2.57; 0.42) & (-0.43; 1.01) & (0.19; 4.28) & \\
$z_{4,i,t}$ & & & & & -0.20 & 0.20 & 1.44 & \\
& & & & & (-0.74; 0.35) & (-0.29; 0.64) & (-0.23; 3.04) & \\
$z_{5,i,t}$ & & & & & -0.17 & -0.08 & -1.20 & \\
& & & & & (-0.46; 0.08) & (-0.19; 0.06) & (-2.28; -0.22) & \\
$z_{6,i,t}$ & & & & & 0.07 & -0.09 & 0.20 & \\
& & & & & (-0.23; 0.38) & (-0.26; 0.08) & (-0.50; 0.98) & \\
\bottomrule
\end{tabular}
\begin{tablenotes}
• \scriptsize Notes: The table reports point estimates of the coefficients of the model in (ref) with the associated 95% bootstrap confidence intervals appearing within parentheses. Results are reported for three estimators, SCCE, CCEP and CCEMG. Results with 95% confidence intervals based on HAC standard errors for our SCCE estimator can be found in appendix, Table (ref). The regressors of interest are $\ln\left(\sigma_{i,t}\right)$, which is a measure of the input skill intensity, and $\ln\left(\frac{H_{i,t}}{L_{i,t}}\right)$, which is the ratio of high-skilled to low-skilled labor. The six control variables are: real capital equipment per worker ($z_{1,i,t}$), the sectoral share of office, computing and accounting equipment ($z_{2,i,t}$), the difference between high-tech and computer capital share ($z_{3,i,t}$), R&D intensity ($z_{4,i,t}$), narrow outsourcing ($z_{5,i,t}$), and the difference between broad and narrow outsourcing ($z_{6,i,t}$). The dependent variable is the log of the relative wage of low-skilled to high-skilled workers.
\end{tablenotes}
\end{threeparttable}
\end{table}
To ensure the above findings are not driven by the chosen number of knots in the sieve base, we re-estimate the model for $C=2, \dots, 5$, which implies that $J\in \{2, 4, 6, 8, 10\}$. Figure (ref) shows that the estimated impacts of $\ln(\sigma_{i,t})$ and $\ln \left(\frac{H_{i,t}}{L_{i,t}}\right)$ are very stable across the values of $J$. The results are therefore quite robust in this regard.\footnote{In the appendix, we also re-estimate the specification in (ref) using HAC standard errors. According to the results in Table (ref) and Figure (ref), the HAC-based confidence intervals are slightly tighter when compared to bootstrapped ones reported in Figure (ref).}
figure[figure omitted — 1,296 chars of source]
We have seen that the SCCE estimates differ quite markedly from those of CCEP and CCEMP. It is tempting to conclude that these differences in the results are due to nonlinearity in the factor component. In order to shed some light on this issue, we test if the factor component is in fact linear. Because CCEP can be seen as a restricted version of our SCCE estimator, we can apply ney's ney $C(\alpha)$ statistic to test the significance of the nonlinear factor transformations included in the sieve base. Under the null hypothesis that the factors enter linearly, the nonlinear transformations should be insignificant, in which case $C(\alpha)$ should be asymptotically chi-squared distributed. The observed test value in the narrow model is given by $94.6$, which is highly significant even at the 1% level. The test results for the full model are almost identical. Linearity is therefore rejected in both models, suggesting that the observed differences in the estimation results are likely due to omitted nonlinearity on behalf of CCEP and CCEMG.
Concluding Remarks
The existing econometric literature on interactive effects panel data models supposes that the common factors enter in a linear fashion. However, since the factors are unobserved, so is their functional form, and there is typically little or no empirical or theoretical guidance. In fact, in most scenarios of empirical relevance, it is not possible to rule out factor nonlinearities. As a response to this, the present paper proposes a new approach -- called “SCCE” -- in which the factors are permitted to enter the model through an unknown functional form. The factors can enter linearly but they are not required to, and if they enter nonlinearly the new approach does not require any knowledge thereof. It is therefore very general in this regard. Interestingly, despite this generality, SCCE is not only computationally simple, easy to implement, and fast, but it also has excellent small-sample and asymptotic properties.
appendix\setcounter{lemma}{0}
\setcounter{theorem}{0}
\setcounter{table}{0}
\setcounter{figure}{0}
\section{Auxiliary Results and Proof of Theorem (ref)}
\@startsection{subsection}{2}
\z@{-.5\linespacing\@plus-.7\linespacing}{.5\linespacing}
{\normalfont}{Rate Results}
Before we start our asymptotic analysis, it is useful to establish several auxiliary rate results. These results will be invoked repeatedly throughout this appendix. We begin by noting that the basis functions satisfy the following standard growth rates (see, for example, newey97):
\begin{align}
\zeta_0(K)=\begin{cases}
O\left(\sqrt{K}\right) & for splines\\
O\left(K\right) & for power series
\end{cases}, \qquad\zeta_{1}(K)=\begin{cases}
O\left(K^{\frac{3}{2}}\right) & for splines\\
O\left(K^{3}\right) & for power series
\end{cases}.
\end{align}
We also note that under the conditions of Assumption (ref)(ref),
\begin{align}
\sqrt{\frac{T}{N}}K^{\frac{1}{4}}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)^{2}=o(1)&\iff\frac{1}{\sqrt{N}}=o\left(\frac{1}{\sqrt{T}K^{\frac{1}{4}}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)^{2}}\right)\notag \\
&\iff\frac{1}{N}=o\left(\frac{1}{T\sqrt{K}\zeta_{0}(K)^{3}\zeta_{1}(K)^{4}}\right),\\
\sqrt{\frac{K}{T}}\zeta_{0}(K)^{\frac{3}{2}}=o(1)&\iff\frac{1}{\sqrt{T}}=o\left(\frac{1}{\sqrt{K}\zeta_{0}(K)^{\frac{3}{2}}}\right)\notag \\
&\iff\frac{1}{T}=o\left(\frac{1}{K\zeta_{0}(K)^{3}}\right),\\
\sqrt{NT}K^{-\frac{\lambda_{i}}{m}}=o(1)&\iff K^{-\frac{\lambda_{i}}{m}}=o\left(\frac{1}{\sqrt{NT}}\right).
\end{align} By using these results and the rates in (ref), we can show the following:
\begin{align}
\frac{\sqrt{KT}\zeta_{1}(K)^{2}}{N\zeta_{0}(K)}&=o\left(\frac{\sqrt{KT}\zeta_{1}(K)^{2}}{T\sqrt{K}\zeta_{0}(K)^{4}\zeta_{1}(K)^{4}}\right)=o\left(\frac{1}{\sqrt{T}\zeta_{0}(K)^{4}\zeta_{1}(K)^{2}}\right)=o(1),\\
\frac{K^{\frac{1}{4}}\zeta_{1}(K)}{\sqrt{N\zeta_{0}(K)}}&=o\left(\frac{K^{\frac{1}{4}}\zeta_{1}(K)}{\sqrt{T}K^{\frac{1}{4}}\zeta_{0}(K)^{2}\zeta_{1}(K)^{2}}\right)=o\left(\frac{1}{\sqrt{T}\zeta_{0}(K)^{2}\zeta_{1}(K)}\right)=o(1),\\
\frac{K}{NT}\zeta_{1}(K)^{2}&=o\left(\frac{K}{NK\zeta_{0}(K)^{3}\zeta_{1}(K)^{2}}\right)=o\left(\frac{1}{N\zeta_{0}(K)^{3}\zeta_{1}(K)^{2}}\right)=o(1),\\
\frac{K^{\frac{1}{4}}}{\sqrt{T}}\zeta_{0}(K)^{\frac{3}{2}}&=o\left(\frac{K^{\frac{1}{4}}\zeta_{0}(K)^{\frac{3}{2}}}{\sqrt{K}\zeta_{0}(K)^{\frac{3}{2}}}\right)=o\left(\frac{1}{K^{\frac{1}{4}}}\right)=o(1),\\
\frac{K^{\frac{5}{2}}}{N^{\frac{3}{2}}}\zeta_{1}(K)^{3}&=o\left(\frac{K^{\frac{5}{2}}\zeta_{1}(K)^{3}}{T^{\frac{3}{2}}K^{\frac{3}{4}}\zeta_{0}(K)^{\frac{9}{2}}\zeta_{1}(K)^{6}}\right)=o\left(\frac{K^{\frac{7}{4}}}{K^{\frac{3}{2}}\zeta_{0}(K)^{\frac{9}{2}}\zeta_{0}(K)^{\frac{9}{2}}\zeta_{1}(K)^{3}}\right)\notag \\
&=o\left(\frac{K^{\frac{1}{4}}}{\zeta_{0}(K)^{9}\zeta_{1}(K)^{3}}\right)=o(1),\\
\frac{K^{\frac{4m-4\lambda_{i}}{4m}}}{N}\zeta_{1}(K)^{2}&=o\left(\frac{KK^{-\frac{\lambda_{i}}{m}}\zeta_{1}(K)^{2}}{T\sqrt{K}\zeta_{0}(K)^{3}\zeta_{1}(K)^{4}}\right)=o\left(\frac{\sqrt{K}K^{-\frac{\lambda_{i}}{m}}}{K\zeta_{0}(K)^{3}\zeta_{0}(K)^{3}\zeta_{1}(K)^{2}}\right)\notag \\
&=o\left(\frac{K^{-\frac{\lambda_{i}}{m}}}{\sqrt{K}\zeta_{0}(K)^{6}\zeta_{1}(K)^{2}}\right)=o\left(\frac{1}{\sqrt{NTK}\zeta_{0}(K)^{6}\zeta_{1}(K)^{2}}\right)\notag \\
&=o(1),\\
\frac{K^{\frac{13m-8\lambda_{i}}{8m}}}{\sqrt{N}}\zeta_{0}(K)^{\frac{3}{4}}\zeta_{1}(K)&=o\left(\frac{K^{\frac{13}{8}}K^{-\frac{\lambda_{i}}{m}}\zeta_{0}(K)^{\frac{3}{4}}\zeta_{1}(K)}{\sqrt{T}K^{\frac{1}{4}}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)^{2}}\right)=o\left(\frac{K^{\frac{11}{8}}K^{-\frac{\lambda_{i}}{m}}}{\sqrt{K}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{0}(K)^{\frac{3}{4}}\zeta_{1}(K)}\right)\notag \\
&=o\left(\frac{K^{\frac{7}{8}}K^{-\frac{\lambda_{i}}{m}}}{\zeta_{0}(K)^{\frac{9}{4}}\zeta_{1}(K)}\right)=o(1),\\
\frac{K^{\frac{3}{4}}}{\sqrt{N}}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)&=o\left(\frac{K^{\frac{3}{4}}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)}{\sqrt{T}K^{\frac{1}{4}}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)^{2}}\right)=o\left(\frac{\sqrt{K}}{\sqrt{K}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)}\right)\notag \\
&=o\left(\frac{1}{\zeta_{0}(K)^{\frac{5}{2}}}\right)=o(1),\\
K^{\frac{m-4\lambda_{i}}{4m}}\zeta_{0}(K)^{\frac{3}{2}}&=o\left(\frac{K^{\frac{1}{4}}\zeta_{0}(K)^{\frac{3}{2}}}{\sqrt{NT}}\right)=o\left(\frac{K^{\frac{1}{4}}\zeta_{0}(K)^{\frac{3}{2}}}{\sqrt{NK}\zeta_{0}(K)^{\frac{3}{2}}}\right)\notag \\
&=o\left(\frac{1}{\sqrt{N}K^{\frac{1}{4}}}\right)=o(1),\\
\sqrt{\frac{T}{N}}\zeta_{1}(K)^{2}&=o\left(\frac{\sqrt{T}\zeta_{1}(K)^{2}}{\sqrt{T}K^{\frac{1}{4}}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)^{2}}\right)=o\left(\frac{1}{K^{\frac{1}{4}}\zeta_{0}(K)^{\frac{3}{2}}}\right)=o(1),\\
\frac{\sqrt{T}K}{N^{\frac{3}{2}}}\zeta_{1}(K)^{4}&=o\left(\frac{\sqrt{T}K\zeta_{1}(K)^{4}}{T^{\frac{3}{2}}K^{\frac{3}{4}}\zeta_{0}(K)^{\frac{9}{2}}\zeta_{1}(K)^{6}}\right)=o\left(\frac{K^{\frac{1}{4}}}{\sqrt{T}\sqrt{K}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{0}(K)^{\frac{9}{2}}\zeta_{1}(K)^{2}}\right)\notag\\
&=o\left(\frac{1}{\sqrt{T}K^{\frac{1}{4}}\zeta_{0}(K)^{\frac{11}{2}}\zeta_{1}(K)^{2}}\right)=o(1),\\
\frac{\sqrt{T}K^{\frac{m-2\lambda_{i}}{m}}}{\sqrt{N}}\zeta_{1}(K)^{2}&=\left(\frac{K^{\frac{m-2\lambda_{i}}{m}}}{K^{\frac{1}{4}}\zeta_{0}(K)^{\frac{3}{2}}}\right)=o\left(\frac{K^{\frac{3}{4}}}{NT\zeta_{0}(K)^{\frac{3}{2}}}\right)\notag \\
&=o\left(\frac{K^{\frac{3}{4}}}{NK\zeta_{0}(K)^{3}\zeta_{0}(K)^{\frac{3}{2}}}\right)=o\left(\frac{1}{NK^{\frac{1}{4}}\zeta_{0}(K)^{\frac{9}{2}}}\right)=o(1),\\
\sqrt{NT}K^{\frac{m-8\lambda_{i}}{4m}}\zeta_{0}(K)^{\frac{3}{2}}&=o\left(\frac{\sqrt{NT}K^{\frac{1}{4}}\zeta_{0}(K)^{\frac{3}{2}}}{NT}\right)=o\left(\frac{K^{\frac{1}{4}}\zeta_{0}(K)^{\frac{3}{2}}}{\sqrt{NT}}\right)\notag \\
&=o\left(\frac{1}{\sqrt{N}K^{\frac{1}{4}}}\right)=o(1),\\
\sqrt{\frac{T}{N}}K\zeta_{1}(K)^{2}&=o\left(\frac{K}{K^{\frac{1}{4}}\zeta_{0}(K)^{\frac{3}{2}}}\right)=o\left(\frac{K^{\frac{3}{4}}}{\zeta_{0}(K)^{\frac{3}{2}}}\right)=o(1),\\
TK^{-\frac{4\lambda_{i}}{m}}&=o\left(\frac{T}{N^{4}T^{4}}\right)=o\left(\frac{1}{N^{4}T^{3}}\right)=o(1),\\
\frac{K^{\frac{3}{2}}}{T^{\frac{3}{2}}}\zeta_{0}(K)&=o\left(\frac{K^{\frac{3}{2}}\zeta_{0}(K)}{K^{\frac{3}{2}}\zeta_{0}(K)^{\frac{9}{2}}}\right)=o\left(\frac{1}{\zeta_{0}(K)^{\frac{7}{2}}}\right)=o(1), \\
\frac{K^{\frac{17}{8}}}{N}\zeta_0(K)^{\frac{3}{4}}\zeta_1(K)^2 &= o\left(\frac{K^{\frac{17}{8}}\zeta_{0}(K)^{\frac{3}{4}}\zeta_{1}(K)^{2}}{T\sqrt{K}\zeta_{0}(K)^{3}\zeta_{1}(K)^{4}}\right)=o\left(\frac{K^{\frac{13}{8}}}{K\zeta_{0}(K)^{3}\zeta_{0}(K)^{\frac{9}{4}}\zeta_{1}(K)^{2}}\right)\notag \\
&=o\left(\frac{K^{\frac{5}{8}}}{\zeta_{0}(K)^{\frac{21}{4}}\zeta_{1}(K)^{2}}\right) = o(1), \\
\sqrt{T}K^{\frac{1}{2}-\frac{\lambda_{i}}{m}}\zeta_{1}(K)&=o\left(\frac{\sqrt{TK}}{\sqrt{NT}}\zeta_{1}(K)\right)=o\left(\frac{\sqrt{K}\zeta_{1}(K)}{\sqrt{T}K^{\frac{1}{4}}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)^{2}}\right)\notag \\
&=o\left(\frac{K^{\frac{1}{4}}}{\sqrt{K}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)}\right)=o\left(\frac{1}{K^{\frac{1}{4}}\zeta_{0}(K)^{3}\zeta_{1}(K)}\right)\notag \\
&=o(1).
\end{align}
\@startsection{subsection}{2}
\z@{-.5\linespacing\@plus-.7\linespacing}{.5\linespacing}
{\normalfont}{A Useful Lemma}
\begin{lemma}
Let $\*f_t \in \mathbb{R}^m$, $\widehat{\*f}_t \in \mathbb{R}^{d+1}$, $\mathcal{G}_i(\cdot): \mathbb{R}^m \to \mathcal{G}_i\left(\mathbb{R}^m\right) \subseteq\mathbb{R}^{d+1}$, and $\overline{\mathcal{G}}(\cdot): \mathbb{R}^m \to \overline{\mathcal{G}}\left(\mathbb{R}^m\right) \subseteq\mathbb{R}^{d+1}$. Under our Assumptions (ref)--(ref), as $N,T \to \infty$,
\begin{enumerate}[label=(\roman*)]
• $\left\|\mathcal{G}_i\left(\*f_t\right) - \+\alpha_{\mathcal{G}_i}' \*p\left(\overline{\mathcal{G}}\left(\*f_t\right)\right)\right\| = o_p(1)$,
• $\left\|\*p (\widehat{\*f}_t) - \*p\left(\overline{\mathcal{G}}\left(\*f_t\right)\right)\right\| = O_p\left(\frac{\zeta_1(K)}{\sqrt{N}}\right)=o_p(1)$.
\end{enumerate}
\end{lemma}
Proof of Lemma (ref). Consider (ref). Let us define a mapping $h_i(\cdot): \mathbb{R}^{d+1} \supseteq \mathcal{R}_{\overline{\mathcal{G}}} \to \mathbb{R}^{d+1}$ and $h_i(\*u) \equiv \mathcal{G}_i\left(\overline{\mathcal{G}}^{-1}\left(\*u\right)\right)$ for some vector $\*u \in \mathcal{R}_{\overline{\mathcal{G}}}$, with $\*u \equiv \overline{\mathcal{G}}\left(\*f_t\right)$ and $\mathcal{R}_{\overline{\mathcal{G}}} \equiv \overline{\mathcal{G}}\left(\mathbb{R}^m\right)$. The mappings $\overline{\mathcal{G}}(\cdot)$ and $\mathcal{G}_i(\cdot)$ are defined as before. By Assumption (ref), $\overline{\mathcal{G}}(\cdot)$ is injective, $H(\lambda, \omega)$-smooth, and has full column rank. Hence, by inverse function theorem, for each $\*u \in \mathcal{R}_{\overline{\mathcal{G}}}$, there exists a neighborhood in which the inverse $\overline{\mathcal{G}}^{-1}(\cdot): \mathbb{R}^{d+1}\supseteq \mathcal{R}_{\overline{\mathcal{G}}} \to \mathbb{R}^m$ is well defined and $H(\lambda, \omega)$-smooth. Consequently, $h_i\left(\overline{\mathcal{G}}\left(\*f_t\right)\right) = \mathcal{G}_i\left(\overline{\mathcal{G}}^{-1}\left(\overline{\mathcal{G}}\left(\*f_t\right)\right)\right) = \mathcal{G}_i\left(\*f_t\right)$. Since a composition of smooth functions is smooth, $h_i(\cdot)$ itself is H\"older-smooth on $\mathcal{R}_{\overline{\mathcal{G}}}$. This means that by Assumption (ref)(ref) there exists a function $\Pi_{\infty K}h_i \equiv \+\alpha_{h_i}'\*p(\cdot)$, such that
\begin{align}
\left\|h_i(\*u) - \Pi_{\infty K}h_i(\*u) \right\|_{\infty, \widetilde{\omega}} = \sup_{\*u} \left\|h_i(\*u) - \+\alpha_{h_i}'\*p(\*u)\right\| \left[1 + \left\|\*u\right\|^2\right]^{-\frac{\tilde{\omega}}{2}} = O\left(K^{-\frac{\lambda_i}{m}}\right).
\end{align}
Before we proceed, let us first state the following theorem (also known as the "Changing Variables Theorem"), which is necessary in order to prove Lemma (ref)(ref) (see for example, evans2015measure, Theorem 3.9):
\begin{theorem}
Let $f: \mathbb{R}^n \to \mathbb{R}^m$ be Lipschitz continuous, $n \leq m$. Then, for each $\mathcal{L}^n$-summable function $g: \mathbb{R}^n \to \mathbb{R}$,
\begin{align}
\int_{\mathbb{R}^n} g(x) J(f(x))dx = \int_{\mathbb{R}^m} \left[\sum_{x \in f^{-1}\{y\}} g(x)\right] d \mathcal{H}^n(y),
\end{align} where a function $g$ is called $\mathcal{L}^n$-summable if $\left|g\right|$ has a finite integral. Here, $J(f(x))$ is a Jacobian of $f(x)$ defined as $J(f(x)) \equiv \sqrt{\det \nabla f(x)'\nabla f(x)}$, and $\mathcal{H}^n(y)$ is the Hausdorff measure.\footnote{The Hausdorff measure generalizes length/area/volume to possibly lower (or even non-integer) dimensions, which allows a measurement in lower-dimensional sets within a space $\mathbb{R}^n$. In contrast, Lebesgue measure can only measure $n$-dimensional volume in this case. For a formal definition of the Hausdorff measure, see evans2015measure.}
\end{theorem}
We would like to point out two things regarding Theorem (ref). First, note that if the function $f$ is injective, then every $y$ in $f\left(\mathbb{R}^n\right)$ can be mapped to at most one point in $\mathbb{R}^n$, such that
\begin{align}
\sum_{x \in f^{-1}\{y\}}g(x) = \begin{cases} g\left(f^{-1}\left(y\right)\right) &\text{if} \ y \in f\left(\mathbb{R}^n\right) \\ 0 \ &\text{otherwise}\end{cases}.
\end{align} Hence, equation (ref) reduces to
\begin{align}
\int_{\mathbb{R}^n} g(x) J(f(x))dx = \int_{f\left(\mathbb{R}^n\right)} g\left(f^{-1}\left(y\right)\right) d\mathcal{H}^n(y).
\end{align} Second, by Assumption (ref), $\overline{\mathcal{G}}(\cdot)$ is H\"older-smooth, and $\nabla \overline{\mathcal{G}}(f_t)$ exists and has full column rank $m$.
Together, these conditions imply that $\overline{\mathcal{G}}(\cdot)$ is locally Lipschitz on $\mathbb R^m$, which is sufficient for the application of Theorem (ref).
Now consider
\begin{align}
\mathbb{E}\left(\left\|h_i (\*u) - \+\alpha_{h_i}'\*p(\*u)\right\|^2\right) = \int_{\mathcal{R}_{\overline{\mathcal{G}}}} \left\|h_i (\*u) - \+\alpha_{h_i}'\*p(\*u)\right\|^2 w_u(\*u) d \mathcal{H}^m(\*u),
\end{align} where $h_i(\*u) = \left[h_{i,1}(\*u), \dots, h_{i,d+1}(\*u)\right]' \in \mathbb{R}^{d+1}$, $\+\alpha_{h_i} \in \mathbb{R}^{K \times (d+1)}$ is a sieve coefficient matrix, $\*p(\cdot) \in \mathbb{R}^{K}$ is a vector of basis functions, $w_u(\*u)$ is a probability density function of the vector $\*u$, and $\mathcal{H}^m(\*u)$ is a Hausdorff measure since $m \leq d+1$. By multiplying and dividing (ref) by $\left(1 + \left\|\*u\right\|^2\right)^{\widetilde{\omega}}$, we obtain
\begin{align}
&\mathbb{E}\left(\left\|h_i (\*u) - \+\alpha_{h_i}'\*p(\*u)\right\|^2\right) \notag\\
& = \int_{\mathcal{R}_{\overline{\mathcal{G}}}} \left\|h_i (\*u) - \+\alpha_{h_i}'\*p(\*u)\right\|^2 w_u(\*u) d \mathcal{H}^m(\*u)\notag \\
&= \int_{\mathcal{R}_{\overline{\mathcal{G}}}} \left[ \left\|h_i (\*u) - \+\alpha_{h_i}'\*p(\*u)\right\|\left(1 + \left\|\*u\right\|^2\right)^{-\frac{\tilde{\omega}}{2}}\right]^2 \left(1 + \left\|\*u\right\|^2\right)^{\widetilde{\omega}}w_u(\*u) d \mathcal{H}^m(\*u) \notag\\
& \leq \left\|h_i (\*u) - \+\alpha_{h_i}'\*p(\*u)\right\|^2_{\infty, \widetilde{\omega}} \int_{\mathcal{R}_{\overline{\mathcal{G}}}}\left(1 + \left\|\*u\right\|^2\right)^{\widetilde{\omega}} w_u(\*u) d\mathcal{H}^m(\*u) \notag\\
& = O\left(K^{-\frac{2\lambda_i}{m}}\right)\int_{\mathcal{R}_{\overline{\mathcal{G}}}}\left(1 + \left\|\*u\right\|^2\right)^{\widetilde{\omega}} w_u(\*u) d\mathcal{H}^m(\*u),
\end{align} where the last equality holds by Assumption (ref)(ref) for a weight $\widetilde{\omega} \geq 0$. By Assumption $\ref{A5}$, $\overline{\mathcal{G}}(\cdot) \in \Lambda_r^{\lambda}(\mathcal{R}_f, \omega)$ and hence, for some generic $\*f_t$,
\begin{align}
\sup_{\*f_t} \left\|\overline{\mathcal{G}}(\*f_t) \left(1+ \left\|\*f_t\right\|^2\right)^{-{\frac{\omega}{2}}}\right\|_{\Lambda^\lambda} &\leq C, \notag\\
\Longleftrightarrow \left\|\overline{\mathcal{G}}(\*f_t)\right\| &\leq C \left(1+ \left\|\*f_t\right\|^2\right)^{{\frac{\omega}{2}}} \notag\\
\Longleftrightarrow \left\|\overline{\mathcal{G}}(\*f_t)\right\|^2 &\leq C \left(1+ \left\|\*f_t\right\|^2\right)^{\omega} \notag\\
\Longleftrightarrow \left(1 + \left\|\overline{\mathcal{G}}(\*f_t)\right\|^2\right)^{\widetilde{\omega}} &\leq C' \left(1+ \left\|\*f_t\right\|^2\right)^{\omega \widetilde{\omega}} = C' \left(1+ \left\|\*f_t\right\|^2\right)^{\overline{\omega}},
\end{align} for some constants $C, C' \in (0, \infty)$ and a weight $\overline{\omega} \geq 0$. This means that $\overline{\mathcal{G}}(\cdot)$ is bounded by the polynomial growth of $\*f_t$. Thus,
\begin{align}
\left(1 + \left\|\*u\right\|^2\right)^{\widetilde{\omega}} \equiv \left(1 + \left\|\overline{\mathcal{G}}\left(\*f_t\right)\right\|^2\right)^{\widetilde{\omega}} \leq C' \left(1+ \left\|\*f_t\right\|^2\right)^{\overline{\omega}}.
\end{align} As discussed before, the inverse $\overline{\mathcal{G}}^{-1}(\cdot)$ exists and is well defined under Assumption (ref), such that
\begin{align} \left(1+\left\|\overline{\mathcal{G}}^{-1}\left(\*u\right)\right\|^2\right)^{\overline{\omega}} &= \left(1+\left\|\overline{\mathcal{G}}^{-1}\left(\overline{\mathcal{G}}\left(\*f_t\right)\right)\right\|^2\right)^{\overline{\omega}} \notag \\
&=\left(1+\left\|\left(\overline{\mathcal{G}}^{-1}\circ\overline{\mathcal{G}}\right)\left(\*f_t\right)\right\|^2\right)^{\overline{\omega}} = \left(1+\left\|\*f_t\right\|^2\right)^{\overline{\omega}}.
\end{align} We can make use of the results in (ref) and (ref) to arrive at the following expression for (ref):
\begin{align}
\mathbb{E}\left(\left\|h_i (\*u) - \+\alpha_{h_i}'\*p(\*u)\right\|^2\right) &= O\left(K^{-\frac{2\lambda_i}{m}}\right)\int_{\mathcal{R}_{\overline{\mathcal{G}}}}\left(1 + \left\|\*u\right\|^2\right)^{\widetilde{\omega}} w_u(\*u) d\mathcal{H}^m(\*u) \notag \\
&\leq O\left(K^{-\frac{2\lambda_i}{m}}\right)\int_{\mathcal{R}_{\overline{\mathcal{G}}}}\left(1 + \left\|\overline{\mathcal{G}}^{-1}\left(\*u\right)\right\|^2\right)^{\overline{\omega}} w_u(\*u) d\mathcal{H}^m(\*u) \notag \\
&= O\left(K^{-\frac{2\lambda_i}{m}}\right) \int_{\mathbb{R}^m} \left(1 + \left\|\*f_t\right\|^2\right)^{\overline{\omega}}w_u(\overline{\mathcal{G}}(\*f_t)) J\left( \overline{\mathcal{G}}(\*f_t)\right) d\*f_t \\
&= O\left(K^{-\frac{2\lambda_i}{m}}\right) \int_{\mathbb{R}^m} \left(1 + \left\|\*f_t\right\|^2\right)^{\overline{\omega}} w_f(\*f_t)d\*f_t,
\end{align} where we use Theorem (ref) for (ref) and define $w_f(\*f_t) \equiv w_u(\overline{\mathcal{G}}(\*f_t))J\left( \overline{\mathcal{G}}(\*f_t)\right) $, with $J\left(\overline{\mathcal{G}}(\*f_t)\right) \equiv \sqrt{\det \left[\nabla \overline{\mathcal{G}}(\*f_t)'\nabla \overline{\mathcal{G}}(\*f_t)\right]}$ being the Jacobian of $\overline{\mathcal{G}}(\*f_t)$. By Assumption (ref)(ref), $\sup_{t=1, \dots, T}\int_{\mathcal{R}_f}\left(1+\left\|\*f_t\right\|^2\right)^{\overline{\omega}}w_f(\*f_t)d\*f_t < \infty$, which allows us to conclude that
\begin{align}
\mathbb{E}\left(\left\|h_i (\*u) - \+\alpha_{h_i}'\*p(\*u)\right\|^2\right)= O\left(K^{-\frac{2\lambda_i}{m}}\right) O(1) = O\left(K^{-\frac{2\lambda_i}{m}}\right) .
\end{align} Hence,
\begin{align}
\left\|h_i (\*u) - \+\alpha_{h_i}'\*p(\*u)\right\| &= \left\|h_i\left(\overline{\mathcal{G}}\left(\*f_t\right)\right) - \+\alpha_{h_i}'\*p\left(\overline{\mathcal{G}}\left(\*f_t\right)\right)\right\| \notag \\
&= \left\|\mathcal{G}_i\left(\*f_t\right) - \+\alpha_{\mathcal{G}_i}'\*p\left(\overline{\mathcal{G}}\left(\*f_t\right)\right)\right\| \notag\\
&=O_p\left(K^{-\frac{\lambda_i}{m}}\right)=o_p(1),
\end{align} where the last equality follows from the fact that $K^{-\frac{\lambda_i}{m}}=o(1)$. This established part (ref) of the lemma.
We move on to part (ref). By the mean-value theorem, there exists a vector $\*f_t^* \in \mathbb{R}^{d+1}$ which lies element-wise between $\overline{\mathcal{G}}(\*f_t)$ and $\widehat{\*f}_t$, such that
\begin{align}
\left\|\*p(\widehat{\*f}_t) - \*p\left(\overline{\mathcal{G}}\left(\*f_t\right)\right)\right\| &= \left\|\nabla \*p\left(\*f_t^*\right)\left(\widehat{\*f}_t - \overline{\mathcal{G}}\left(\*f_t\right)\right)\right\| \notag\\
& \leq \left\|\nabla \*p\left(\*f_t^*\right) \right\| \left\|\left(\widehat{\*f}_t - \overline{\mathcal{G}}\left(\*f_t\right)\right)\right\| \notag\\
& \leq \zeta_1\left(K\right) \left\|\overline{\*u}_t\right\|\notag \\
&= O_p\left(\frac{\zeta_1\left(K\right)}{\sqrt{N}}\right) \notag \\
&=O_p\left(o(1)\right) = o_p(1),
\end{align} where we have used that $\overline{\*u}_t = O_p\left(\frac{1}{\sqrt{N}}\right)$ and the result in (ref). This establishes (ref), and hence the proof of Lemma (ref) is complete. $\blacksquare$
\@startsection{subsection}{2}
\z@{-.5\linespacing\@plus-.7\linespacing}{.5\linespacing}
{\normalfont}{Proof of Theorem (ref)}
Recall from Assumption (ref) that $\widetilde{\mathbf{P}} = \left[\widetilde{\*p}\left(\*f_1\right), \dots, \widetilde{\*p}\left(\*f_T\right)\right]' = \left[\*p\left(\overline{\mathcal{G}}\left(\*f_1\right)\right), \dots, \*p\left(\overline{\mathcal{G}}\left(\*f_T\right)\right) \right]'$, where the mapping $\overline{\mathcal{G}}(\cdot)$ is defined in Assumption (ref). Further, we define $\+\alpha_{g_i}\in \mathbb{R}^{K \times 1}$ and $\+\alpha_{G_i} \in \mathbb{R}^{K \times d}$ as the sieve coefficient vector and matrix, when approximating $\*g_i(\*F)$ and $\*G_i(\*F)$, respectively. We continue by normalizing $\widehat{\mathbf{P}}$. As a convenient way to partly deal with the growing dimension of $\widetilde{\mathbf{P}}$, similarly to newey97, we replace $\widehat{\mathbf{P}}$ by $\widehat{\mathbf{P}}\left(\mathbb{E}\left[\frac{1}{T}\widetilde{\mathbf{P}}' \widetilde{\mathbf{P}}\right]\right)^{-\frac{1}{2}}$, where $\*A^{\frac{1}{2}}$ is such that $\left(\*A^{\frac{1}{2}}\right)'\*A^{\frac{1}{2}}=\*A$ for any positive semidefinite matrix $\*A$. This is possible because, by Assumption (ref)(ref), $\lambda_{\min}\left[\mathbb{E}\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)\right] > 0$. Moreover, this normalization preserves projection operators such that $\*M_{\widehat{P}} = \*M_{\widehat{P}\left(\mathbb{E}[\frac{1}{T} \widetilde{P}'\widetilde{P}]\right)^{-\frac{1}{2}}}$. Because of this normalization, we can assume without loss of generality that $\mathbb{E}\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right) = \*I_K$.
We begin by re-writing equation (ref) for $\*y_{i}$ in the following way:
\begin{align}
\*y_i &= \*X_i\+\beta + \*g_i\left(\*F\right) + \+\varepsilon_i \notag \\
&= \*X_i \+\beta + \widehat{\*P}\+\alpha_{g_i} + \+\varepsilon_i - \left(\widehat{\*P}\+\alpha_{g_i} - \*g_i\left(\*F\right)\right),
\end{align} where
\begin{align}
\widehat{\*P}\+\alpha_{g_i} - \*g_i\left(\*F\right) &= \left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right) \+\alpha_{g_i} + \left(\widetilde{\mathbf{P}} \+\alpha_{g_i} - \*g_i\left(\*F\right)\right).
\end{align}
Insertion into $\widehat{\+\beta}_{SCCE}$ yields
\begin{align}
\sqrt{NT}\left(\widehat{\+\beta}_{SCCE} - \+\beta\right) = \left(\frac{1}{NT}\sum_{i=1}^N \*X_i'\*M_{\widehat{P}}\*X_i \right)^{-1} \frac{1}{\sqrt{NT}}\sum_{i=1}^N\*X_i'\*M_{\widehat{P}}\left[\+\varepsilon_i - \left(\widehat{\mathbf{P}}\+\alpha_{g_i} - \*g_i\left(\*F\right)\right)\right].
\end{align}
A key first step in establishing the asymptotic distribution of $\sqrt{NT}\left(\widehat{\+\beta}_{SCCE} - \+\beta\right)$ is to derive a cleaned up asymptotic representation that is free of all errors coming from the functional approximation. We start by considering $\*M_{\widehat P} \*X_i$. Equation (ref) for $\*X_i$ can be manipulated in the same way as (ref) for $\*y_i$ to obtain
\begin{align}
\*X_i = \*G_i\left(\*F\right) + \*V_i = \widehat{\mathbf{P}} \+\alpha_{G_i} + \*V_i - \left(\widehat{\mathbf{P}} \+\alpha_{G_i} - \*G_i\left(\*F\right)\right),
\end{align} where analogously to $\widehat{\mathbf{P}} \+\alpha_{g_i} - \*g_i\left(\*F\right)$ we can write
\begin{align}
\widehat{\mathbf{P}} \+\alpha_{G_i} - \*G_i\left(\*F\right) &= \left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right) \+\alpha_{G_i} + \left(\widetilde{\mathbf{P}} \+\alpha_{G_i} - \*G_i\left(\*F\right)\right).
\end{align}
It follows that
\begin{align}
\*M_{\widehat{P}}\*X_i = \*M_{\widehat{P}}\left[\*V_i - \left(\widehat{\mathbf{P}} \+\alpha_{G_i} - \*G_i\left(\*F\right)\right)\right].
\end{align}
We now make use of this last expression for $\*M_{\widehat{P}}\*X_i$ in order to evaluate the numerator of (ref). The following expansion will be used:
\begin{align}
&\frac{1}{\sqrt{NT}}\sum_{i=1}^N\*X_i'\*M_{\widehat{P}}\left[\+\varepsilon_i - \left(\widehat{\mathbf{P}}\+\alpha_{g_i} - \*g_i\left(\*F\right)\right) \right] \notag \\
&=\frac{1}{\sqrt{NT}}\sum_{i=1}^N \left[\*V_i - \left(\widehat{\mathbf{P}} \+\alpha_{G_i} - \*G_i\left(\*F\right)\right)\right]'\*M_{\widehat{P}}\left[\+\varepsilon_i - \left(\widehat{\mathbf{P}}\+\alpha_{g_i} - \*g_i\left(\*F\right)\right) \right] \notag \\
&= \frac{1}{\sqrt{NT}}\sum_{i=1}^N \left[\*V_i - \left(\widehat{\mathbf{P}} \+\alpha_{G_i} - \*G_i\left(\*F\right)\right)\right]'\*M_{\widetilde{P}}\left[\+\varepsilon_i - \left(\widehat{\mathbf{P}}\+\alpha_{g_i} - \*g_i\left(\*F\right)\right) \right] \notag \\
&- \frac{1}{\sqrt{NT}}\sum_{i=1}^N \left[\*V_i - \left(\widehat{\mathbf{P}} \+\alpha_{G_i} - \*G_i\left(\*F\right)\right)\right]'\left(\*M_{\widetilde{P}} - \*M_{\widehat{P}}\right)\left[\+\varepsilon_i - \left(\widehat{\mathbf{P}}\+\alpha_{g_i} - \*g_i\left(\*F\right)\right) \right] \notag \\
&= \sum_{j=1}^8 \*D_j,
\end{align} where
\begin{align}
\*D_1 &= \frac{1}{\sqrt{NT}}\sum_{i=1}^N \*V_i' \*M_{\widetilde{P}} \+\varepsilon_i, \notag \\
\*D_2 &= - \frac{1}{\sqrt{NT}}\sum_{i=1}^N \*V_i' \*M_{\widetilde{P}}\left(\widehat{\mathbf{P}} \+\alpha_{g_i} - \*g_i\left(\*F\right)\right),\notag \\
\*D_3 &= - \frac{1}{\sqrt{NT}}\sum_{i=1}^N \left(\widehat{\mathbf{P}} \+\alpha_{G_i} - \*G_i\left(\*F\right)\right)'\*M_{\widetilde{P}}\+\varepsilon_i, \notag\\
\*D_4 &= \frac{1}{\sqrt{NT}}\sum_{i=1}^N \left(\widehat{\mathbf{P}} \+\alpha_{G_i} - \*G_i\left(\*F\right)\right)'\*M_{\widetilde{P}}\left(\widehat{\mathbf{P}} \+\alpha_{g_i} - \*g_i\left(\*F\right)\right),\notag \\
\*D_5 &= - \frac{1}{\sqrt{NT}}\sum_{i=1}^N \*V_i'\left(\*M_{\widetilde{P}} - \*M_{\widehat{P}}\right)\+\varepsilon_i, \notag\\
\*D_6 &= \frac{1}{\sqrt{NT}}\sum_{i=1}^N \*V_i' \left(\*M_{\widetilde{P}} - \*M_{\widehat{P}}\right)\left(\widehat{\mathbf{P}} \+\alpha_{g_i} - \*g_i\left(\*F\right)\right),\notag \\
\*D_7 &= \frac{1}{\sqrt{NT}}\sum_{i=1}^N \left(\widehat{\mathbf{P}}\+\alpha_{G_i} - \*G_i\left(\*F\right)\right)'\left(\*M_{\widetilde{P}} - \*M_{\widehat{P}}\right)\+\varepsilon_i,\notag \\
\*D_8 &= - \frac{1}{\sqrt{NT}}\sum_{i=1}^N \left(\widehat{\mathbf{P}} \+\alpha_{G_i} - \*G_i\left(\*F\right)\right)'\left(\*M_{\widetilde{P}} - \*M_{\widehat{P}}\right)\left(\widehat{\mathbf{P}} \+\alpha_{g_i} - \*g_i\left(\*F\right)\right).
\end{align}
We now proceed to evaluate each of the terms appearing above, starting with $\*D_5$. Analogous to (A.12) in westerlundpetrova2019, we have
\begin{align}
\*M_{\widetilde P} - \*M_{\widehat P} &= \left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)' + \left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ \widetilde{\mathbf{P}}' \notag \\
& \quad + \widetilde{\mathbf{P}}\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)' + \widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ - \left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^+\right]\widetilde{\mathbf{P}}',
\end{align} from which it follows that
\begin{align}
-\*D_5 & \equiv \frac{1}{\sqrt{NT}}\sum_{i=1}^N \*V_i'\left(\*M_{\widetilde P} - \*M_{\widehat{P}}\right)\+\varepsilon_i \notag \\
&= \frac{1}{\sqrt{NT}}\sum_{i=1}^N \*V_i'\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}' \widehat{\mathbf{P}}\right)^+\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)' \+\varepsilon_i + \frac{1}{\sqrt{NT}}\sum_{i=1}^N \*V_i'\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}' \widehat{\mathbf{P}}\right)^+ \widetilde{\mathbf{P}}'\+\varepsilon_i \notag \\
&\quad + \frac{1}{\sqrt{NT}}\sum_{i=1}^N \*V_i'\widetilde{\mathbf{P}}\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\+\varepsilon_i + \frac{1}{\sqrt{NT}}\sum_{i=1}^N \*V_i'\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ - \left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^+\right]\widetilde{\mathbf{P}}'\+\varepsilon_i \notag \\
&\equiv \*d_{5,1} + \*d_{5,2} + \*d_{5,3} + \*d_{5,4} ,
\end{align}
with an implicit definition of $\*d_{5,1}$, $\*d_{5,2}$, $\*d_{5,3}$ and $\*d_{5,4}$. In what remains of this appendix, implicit definitions like these will be used repeatedly to save space. Before we begin evaluating each of the four terms in (ref), let us state the following lemma, Lemma (ref), which will be used throughout this proof. The proof of Lemma (ref) is provided after the proof of Theorem (ref).
\begin{lemma}
Under the conditions of Theorem (ref),
\begin{enumerate}[label=(\roman*)]
• $
\left\|\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}-\*I_K\right\| = O_p\left(\sqrt{\frac{K}{T}}\zeta_0(K)\right)$,
• $\lambda_{\min}\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)=1+o_p(1)$, $\lambda_{\max}\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)=1+o_p(1)$,
• $\lambda_{\min}\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)=1+o_p(1)$, $\lambda_{\max}\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)=1+o_p(1)$.
\end{enumerate}
\end{lemma}
We start by taking the second term of $\*D_5$, $\*d_{5,2} \equiv \frac{1}{\sqrt{NT}}\sum_{i=1}^N \*V_i'\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}' \widehat{\mathbf{P}}\right)^+ \widetilde{\mathbf{P}}'\+\varepsilon_i$, in detail. The evaluation of the other terms is similar. Consider some matrices $\*A$ and $\*B$. Then, $\mathrm{tr}(\*B\*A\*B')\leq \lambda_{\max}(\*A)\mathrm{tr}(\*B\*B')$ whenever $\*A$ is positive semidefinite (see lut, result (12) on page 44). Under Assumptions (ref) and (ref) this allows us to obtain
\begin{align}
& \mathbb{E}\left(\left\|\*d_{5,2}\right\|^2\right) \notag\\
&=\mathbb{E}\left(\left\|\frac{1}{\sqrt{NT}}\sum_{i=1}^N \*V_i' \left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}' \widehat{\mathbf{P}}\right)^+ \widetilde{\mathbf{P}}' \+\varepsilon_i \right\|^2\right) \notag\\
& \leq \mathbb{E}\left[\mathrm{tr}\left(\frac{1}{NT}\sum_{i=1}^N\sum_{j=1}^N\+\varepsilon_i'\widetilde{\mathbf{P}}\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\mathbb{E}\left(\*V_i\*V_j'\right) \left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}' \widehat{\mathbf{P}}\right)^+ \widetilde{\mathbf{P}}' \+\varepsilon_j\right)\right] \notag \\
&\leq \sum_{i=1}^N \lambda_{\max}\left(\+\Omega_{v,i}\right)\mathbb{E}\left[\mathrm{tr}\left(\frac{1}{NT}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}' \widehat{\mathbf{P}}\right)^+ \widetilde{\mathbf{P}}' \mathbb{E}\left(\+\varepsilon_i\+\varepsilon_i'\right)\widetilde{\mathbf{P}}\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+\left(\widetilde{\mathbf{P}} - \widehat{\mathbf{P}}\right)'\right)\right] \notag \\
&\leq \sum_{i=1}^N\lambda_{\max}\left(\+\Omega_{v,i}\right)\lambda_{\max}\left(\+\Omega_{\varepsilon,i}\right)\mathbb{E}\left[\frac{1}{NT}\left\|\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}' \widehat{\mathbf{P}}\right)^+ \widetilde{\mathbf{P}}'\right\|^2\right] \notag \\
&\leq \sup_{i=1, \dots, N}\lambda_{\max}\left(\+\Omega_{v,i}\right)\sup_{i=1, \dots, N}\lambda_{\max}\left(\+\Omega_{\varepsilon,i}\right)\mathbb{E}\left[\frac{1}{NT}\sum_{i=1}^N\left\|\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}' \widehat{\mathbf{P}}\right)^+ \widetilde{\mathbf{P}}'\right\|^2\right] \notag \\
&= O(1)\mathbb{E}\left[\frac{1}{NT}\sum_{i=1}^N\left\|\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}' \widehat{\mathbf{P}}\right)^+ \widetilde{\mathbf{P}}'\right\|^2\right] \notag \\
&= O\left(\frac{1}{T}\right)\mathbb{E}\left[\left\|\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}' \widehat{\mathbf{P}}\right)^+ \widetilde{\mathbf{P}}'\right\|^2\right],
\end{align} where the last equality holds by Assumption (ref)(ref), since the covariance matrices $\+\Omega_{\varepsilon,i}$ and $\+\Omega_{v,i}$ have eigenvalues bounded away from zero and infinity.
We are left to analyze $\mathbb{E}\left[\left\|\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}' \widehat{\mathbf{P}}\right)^+ \widetilde{\mathbf{P}}'\right\|^2\right]$ in (ref). For that, add and subtract $\frac{1}{T}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\widetilde{\mathbf{P}}'$ from $\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}' \widehat{\mathbf{P}}\right)^+ \widetilde{\mathbf{P}}'$ and obtain
\begin{align}
\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ \widetilde{\mathbf{P}}' & = \frac{1}{T}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}} \right)\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ \widetilde{\mathbf{P}}' \notag\\
& = \frac{1}{T}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\widetilde{\mathbf{P}}' + \frac{1}{T}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left[\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+} -\*I_K\right] \widetilde{\mathbf{P}}'.
\end{align} Hence,
\begin{align}
\mathbb{E}\left[\left\|\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ \widetilde{\mathbf{P}}' \right\|^2\right] &\leq 2\mathbb{E}\left[\left\|\frac{1}{T}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right) \widetilde{\mathbf{P}}' \right\|^2\right] \notag \\
&\quad + 2\mathbb{E}\left[\left \|\frac{1}{T}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left[\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ - \*I_K \right] \widetilde{\mathbf{P}}' \right\|^2\right],
\end{align} since $(a+b)^2 \leq 2(a^2+b^2)$.
We begin by evaluating $\mathbb{E}\left[\left\|\frac{1}{T}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right) \widetilde{\mathbf{P}}' \right\|^2\right]$. First, note how
\begin{align}
&\mathbb{E}\left[\left\|\frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\right\|^4\right] \notag\\
& = \mathbb{E}\left[\left(\left\|\frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\right\|^2\right)^2\right] \notag \\
&=\mathbb{E}\left[\left(\frac{1}{T}\sum_{t=1}^T \left\|\*p(\widehat{\*f}_t)-\widetilde{\*p}(\*f_t)\right\|^2\right)^2\right] \notag \\
&=\mathbb{E}\left[\frac{1}{T^2}\sum_{t=1}^T\left\|\*p(\widehat{\*f}_t)-\widetilde{\*p}(\*f_t)\right\|^4+\frac{2}{T^2}\sum_{t
<s}\left\|\*p(\widehat{\*f}_t)-\widetilde{\*p}(\*f_t)\right\|^2\left\|\*p(\widehat{\*f}_s)-\widetilde{\*p}(\*f_s)\right\|^2\right] \notag \\
&=\mathbb{E}\left[\frac{1}{T^2}\sum_{t=1}^T\left\|\*p(\widehat{\*f}_t)-\widetilde{\*p}(\*f_t)\right\|^4\right]+\mathbb{E}\left[\frac{2}{T^2}\sum_{t< s}\left\|\*p(\widehat{\*f}_t)-\widetilde{\*p}(\*f_t)\right\|^2\left\|\*p(\widehat{\*f}_s)-\widetilde{\*p}(\*f_s)\right\|^2\right] \notag \\
&= \frac{1}{T^2}\sum_{t=1}^T \mathbb{E}\left(\left\|\*p(\widehat{\*f}_t)-\widetilde{\*p}(\*f_t)\right\|^4\right) + \frac{2}{T^2}\sum_{t< s} \mathbb{E}\left[\left\|\*p(\widehat{\*f}_t)-\widetilde{\*p}(\*f_t)\right\|^2\left\|\*p(\widehat{\*f}_s)-\widetilde{\*p}(\*f_s)\right\|^2\right],
\end{align} where we have used $\left(\sum_{i}a_i\right)^2 = \sum_ia_i^2+2\sum_{i< j}a_ia_j$. Using this fact again and the Cauchy-Schwarz inequality, the first term in (ref) can be expressed as
\begin{align}
&\frac{1}{T^{2}}\sum_{t=1}^{T}\mathbb{E}\left[\left\Vert \mathbf{p}(\widehat{\mathbf{f}}_{t})-\widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert ^{4}\right] \notag\\
& =\frac{1}{T^{2}}\sum_{t=1}^{T}\mathbb{E}\left[\left(\left\Vert \mathbf{p}(\widehat{\mathbf{f}}_{t})-\widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert ^{2}\right)^2\right]\notag \\
&=\frac{1}{T^{2}}\sum_{t=1}^{T}\mathbb{E}\left[\left(\sum_{k=1}^{K}\left|p_{k}(\widehat{\mathbf{f}}_{t})-\widetilde{p}_{k}(\mathbf{f}_{t})\right|^{2}\right)^{2}\right] \notag \\
&=\frac{1}{T^{2}}\sum_{t=1}^{T}\mathbb{E}\left[\sum_{k=1}^{K}\left|p_{k}(\widehat{\mathbf{f}}_{t})-\widetilde{p}_{k}(\mathbf{f}_{t})\right|^{4}+2\sum_{k< l}\left|p_{k}(\widehat{\mathbf{f}}_{t})-\widetilde{p}_{k}(\mathbf{f}_{t})\right|^{2}\left|p_{l}(\widehat{\mathbf{f}}_{t})-\widetilde{p}_{l}(\mathbf{f}_{t})\right|^{2}\right]\notag \\
&=\frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{k=1}^{K}\mathbb{E}\left[\left|p_{k}(\widehat{\mathbf{f}}_{t})-\widetilde{p}_{k}(\mathbf{f}_{t})\right|^{4}\right]+\frac{2}{T^{2}}\sum_{t=1}^{T}\sum_{k< l}\mathbb{E}\left[\left|p_{k}(\widehat{\mathbf{f}}_{t})-\widetilde{p}_{k}(\mathbf{f}_{t})\right|^{2}\left|p_{l}(\widehat{\mathbf{f}}_{t})-\widetilde{p}_{l}(\mathbf{f}_{t})\right|^{2}\right]\notag \\
&\leq \frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{k=1}^{K}\sup_{k}\mathbb{E}\left[\left|p_{k}(\widehat{\mathbf{f}}_{t})-\widetilde{p}_{k}(\mathbf{f}_{t})\right|^{4}\right]+\frac{2}{T^{2}}\sum_{t=1}^{T}\sum_{k< l}\sup_{k,l}\mathbb{E}\left[\left|p_{k}(\widehat{\mathbf{f}}_{t})-\widetilde{p}_{k}(\mathbf{f}_{t})\right|^{2}\left|p_{l}(\widehat{\mathbf{f}}_{t})-\widetilde{p}_{l}(\mathbf{f}_{t})\right|^{2}\right]\notag \\
&\leq\frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{k=1}^{K}\sup_{k}\mathbb{E}\left[\left|p_{k}(\widehat{\mathbf{f}}_{t})-\widetilde{p}_{k}(\mathbf{f}_{t})\right|^{4}\right]\notag \\
&\quad +\frac{2}{T^{2}}\sum_{t=1}^{T}\sum_{k< l}\sup_{k,l}\sqrt{\mathbb{E}\left[\left|p_{k}(\widehat{\mathbf{f}}_{t})-\widetilde{p}_{k}(\mathbf{f}_{t})\right|^{4}\right]}\sqrt{\mathbb{E}\left[\left|p_{l}(\widehat{\mathbf{f}}_{t})-\widetilde{p}_{l}(\mathbf{f}_{t})\right|^{4}\right]}\notag \\
&\leq\frac{K}{T^{2}}\sum_{t=1}^{T}\sup_{k}\mathbb{E}\left[\left|p_{k}(\widehat{\mathbf{f}}_{t})-\widetilde{p}_{k}(\mathbf{f}_{t})\right|^{4}\right]\notag \\
&\quad +\frac{2K^{2}}{T^{2}}\sum_{t=1}^{T}\sqrt{\sup_{k}\mathbb{E}\left[\left|p_{k}(\widehat{\mathbf{f}}_{t})-\widetilde{p}_{k}(\mathbf{f}_{t})\right|^{4}\right]}\sqrt{\sup_{l}\mathbb{E}\left[\left|p_{l}(\widehat{\mathbf{f}}_{t})-\widetilde{p}_{l}(\mathbf{f}_{t})\right|^{4}\right]}\notag \\
&=\frac{K}{T}O\left(\sup_{k}\mathbb{E}\left[\left|p_{k}(\widehat{\mathbf{f}}_{t})-\widetilde{p}_{k}(\mathbf{f}_{t})\right|^{4}\right]\right)\notag \\
&\quad +\frac{2K^{2}}{T}O\left(\sqrt{\sup_{k}\mathbb{E}\left[\left|p_{k}(\widehat{\mathbf{f}}_{t})-\widetilde{p}_{k}(\mathbf{f}_{t})\right|^{4}\right]}\sqrt{\sup_{l}\mathbb{E}\left[\left|p_{l}(\widehat{\mathbf{f}}_{t})-\widetilde{p}_{l}(\mathbf{f}_{t})\right|^{4}\right]}\right)\notag \\
&=\frac{K}{T}O\left(\frac{\zeta_{1}(K)^{4}}{N^{2}}\right)+\frac{K^{2}}{T}O\left(\frac{\zeta_{1}(K)^{4}}{N^{2}}\right)\notag \\
&=O\left(\frac{K^{2}}{TN^{2}}\zeta_{1}(K)^{4}\right),
\end{align} which holds because under Assumption (ref)(ref), we have $\sup_k \mathbb{E}\left(\left|p_k(\cdot)\right|^{8+\eta}\right)\leq C$ and $\sup_k \mathbb{E}\left(\left|\widetilde{p}_k(\cdot)\right|^{8+\eta}\right)\leq C$ for some $\eta > 0$. This implies that the $(8+\eta)$-th moments of $\left|p_k(\cdot)-\widetilde{p}(\cdot)\right|$ are uniformly bounded as well because by Minkowski inequality, $\sup_k\mathbb{E}\left(\left|p_k(\cdot)-\widetilde{p}(\cdot)\right|^{8+\eta}\right)^{\frac{1}{8+\eta}}\leq \sup_k\left[\mathbb{E}\left(\left|p_k(\cdot)\right|^{8+\eta}\right)\right]^{\frac{1}{8+\eta}}+\sup_k\left[\mathbb{E}\left(\left|\widetilde{p}_k(\cdot)\right|^{8+\eta}\right)\right]^{\frac{1}{8+\eta}}$. Consequently, the random variable $\left\{\left\|\*p(\widehat{\*f}_t)-\widetilde{\*p}(\*f_t)\right\|\right\}$ is uniformly integrable (see for example Theorem 6.1 in dasgupta). This is important because uniform integrability combined with convergence in probability implies convergence in moments. This in turn implies that we can use the bound in Lemma (ref)(ref) to bound $\mathbb{E}\left(\left|p_k(\widehat{\*f}_t)-\widetilde{p}_k(\*f_t)\right|^4\right)$. The argument goes as follows: Since by Lemma (ref)(ref) $\left\|\*p(\widehat{\*f}_t)-\widetilde{\*p}(\*f_t)\right\|=O_p\left(\frac{\zeta_1(K)}{\sqrt{N}}\right)$, we have $\left\|\*p(\widehat{\*f}_t)-\widetilde{\*p}(\*f_t)\right\|^4=O_p\left(\frac{\zeta_1(K)^4}{N^2}\right)$. Further, it is easy to see that for any $k\in \left\{1, \dots, K\right\}$,
\begin{align}
\left|p_k(\widehat{\*f}_t) -\widetilde{p}_k(\*f_t)\right|^4 \leq \sum_{k=1}^K\left|p_k(\widehat{\*f}_t) -\widetilde{p}_k(\*f_t)\right|^4 = \left\|\*p(\widehat{\*f}_t)-\widetilde{\*p}(\*f_t)\right\|^4,
\end{align} implying component-wise convergence in probability of the vector $\*p(\widehat{\*f}_t)-\widetilde{\*p}(\*f_t)$, such that $\left|p_k(\widehat{\*f}_t) -\widetilde{p}_k(\*f_t)\right|^4 =O_p\left(\frac{\zeta_1(K)^4}{N^2}\right)$. We can therefore conclude that $\mathbb{E}\left(\left|p_k(\widehat{\*f}_t)-\widetilde{p}_k(\*f_t)\right|^4\right)=O\left(\frac{\zeta_1(K)^4}{N^2}\right)$ for all $k=1, \dots, K$.
For the second term in (ref), we use the Cauchy-Schwarz inequality and the result we obtained in (ref). We arrive at
\begin{align}
&\frac{2}{T^2}\sum_{t\leq s} \mathbb{E}\left[\left\|\*p(\widehat{\*f}_t)-\widetilde{\*p}(\*f_t)\right\|^2\left\|\*p(\widehat{\*f}_s)-\widetilde{\*p}(\*f_s)\right\|^2\right] \notag \\
&\leq \frac{2}{T^2}\sum_{t\leq s}\sqrt{\mathbb{E}\left[\left\|\*p(\widehat{\*f}_t)-\widetilde{\*p}(\*f_t)\right\|^4\right]}\sqrt{\mathbb{E}\left[\left\|\*p(\widehat{\*f}_s)-\widetilde{\*p}(\*f_s)\right\|^4\right]} \notag \\
&= \frac{2}{T^2}\frac{T(T-1)}{2}O\left(\frac{K^2}{N^2}\zeta_1(K)^4\right) \notag \\
&= O\left(\frac{K^2}{N^2}\zeta_1(K)^4\right).
\end{align} Consequently,
\begin{align}
\mathbb{E}\left[\left\|\frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\|^4\right] &= O\left(\frac{K^{2}}{TN^{2}}\zeta_{1}(K)^{4}\right) + O\left(\frac{K^2}{N^2}\zeta_1(K)^4\right) \notag \\
&=O\left(\frac{K^{2}}{N^{2}}\zeta_{1}(K)^{4}\right).
\end{align}
We now move on to $\mathbb{E}\left[\|\frac{1}{T}(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}})\widetilde{\mathbf{P}}'\|^2\right]$. To bound this term, we can use (ref) and the Cauchy-Schwarz inequality to obtain
\begin{align}
\mathbb{E}\left[\left\|\frac{1}{T}(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}})\widetilde{\mathbf{P}}'\right\|^2\right] & \le\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{2}\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\right\Vert ^{2}\right]\notag \\
&\leq\sqrt{\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{4}\right]}\sqrt{\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\right\Vert ^{4}\right]} \notag \\
&= \sqrt{O\left(\frac{K^2}{N^2}\zeta_1(K)^4\right)}\sqrt{\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\right\Vert ^{4}\right]} \notag \\
&=O\left(\frac{K}{N}\zeta_1(K)^2\right)\sqrt{\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\right\Vert ^{4}\right]},
\end{align} where by using Jensen's inequality and the normalization condition $\mathbb{E}(\widetilde{\*p}(\*f_t)\widetilde{\*p}(\*f_t)')=\*I_K$, we can show that
\begin{align}
\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\right\Vert ^{4}\right] &=\mathbb{E}\left[\left(\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\right\Vert ^{2}\right)^{2}\right]\notag\\
&=\mathbb{E}\left[\left(\frac{1}{T}\sum_{t=1}^{T}\left\Vert \widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert ^{2}\right)^{2}\right]\notag \\
&=\mathbb{E}\left[\frac{1}{T^{2}}\sum_{t=1}^{T}\left\Vert \widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert ^{4}+\frac{2}{T^{2}}\sum_{t<s}\left\Vert \widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert ^{2}\left\Vert \widetilde{\mathbf{p}}(\mathbf{f}_{s})\right\Vert ^{2}\right]\notag \\
&=\mathbb{E}\left[\frac{1}{T^{2}}\sum_{t=1}^{T}\left\Vert \widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert ^{4}\right]+\mathbb{E}\left[\frac{2}{T^{2}}\sum_{t<s}\left\Vert \widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert ^{2}\left\Vert \widetilde{\mathbf{p}}(\mathbf{f}_{s})\right\Vert ^{2}\right]\notag \\
&\leq\sup_{\mathbf{f}_{t}}\mathbb{E}\left[\frac{1}{T^{2}}\sum_{t=1}^{T}\left\Vert \widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert ^{4}\right]+\sup_{\mathbf{f}_{t}}\mathbb{E}\left[\frac{2}{T^{2}}\sum_{t<s}\left\Vert \widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert ^{2}\left\Vert \widetilde{\mathbf{p}}(\mathbf{f}_{s})\right\Vert ^{2}\right]\notag \\
&\leq\frac{1}{T^{2}}\sum_{t=1}^{T}\sup_{\mathbf{f}_{t}}\mathbb{E}\left[\left\Vert \widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert ^{4}\right]+\frac{2}{T^{2}}\sum_{t<s}\sup_{\mathbf{f}_{t}}\mathbb{E}\left[\left\Vert \widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert ^{2}\left\Vert \widetilde{\mathbf{p}}(\mathbf{f}_{s})\right\Vert ^{2}\right]\notag \\
&\leq\frac{1}{T^2}\sum_{t=1}^T \mathrm{tr}\left[\mathbb{E}\left(\widetilde{\mathbf{p}}(\mathbf{f}_{t})\widetilde{\mathbf{p}}(\mathbf{f}_{t})'\right)\right]\zeta_0(K)^2 \notag\\
& \quad+\frac{2}{T^{2}}\sum_{t<s}\sqrt{\sup_{\*f_t}\mathbb{E}\left(\left\|\widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\|^4\right)}\sqrt{\sup_{\*f_s}\mathbb{E}\left(\left\|\widetilde{\mathbf{p}}(\mathbf{f}_{s})\right\|^4\right)}\notag \\
&\leq\frac{K\zeta_{0}(K)^{2}}{T}+\frac{2}{T^{2}}\frac{T(T-1)}{2}K\zeta_{0}(K)^{2}\notag \\
&=\frac{K\zeta_{0}(K)^{2}}{T}+\frac{T-1}{T}K\zeta_{0}(K)^{2}\notag \\
&=\frac{K\zeta_{0}(K)^{2}}{T}+K\zeta_{0}(K)^{2}-\frac{K\zeta_{0}(K)^{2}}{T}\notag \\
&=K\zeta_{0}(K)^{2},
\end{align} and hence,
\begin{align}
\mathbb{E}\left[\left\|\frac{1}{T}(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}})\widetilde{\mathbf{P}}'\right\|^2\right] &\leq O\left(\frac{K}{N}\zeta_1(K)^2\right)\sqrt{\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\right\Vert ^{4}\right]} \notag \\
&= O\left(\frac{K}{N}\zeta_1(K)^2\right)\sqrt{O\left(K\zeta_{0}(K)^{2}\right)}\notag \\
&=O\left(\frac{K^{\frac{3}{2}}}{N}\zeta_1(K)^2\zeta_0(K)\right).
\end{align}
We now analyze the second term in (ref) in detail. First, applying H\"older's inequality, we get
\begin{align}
&\mathbb{E}\left\{ \left\Vert \frac{1}{T}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left[\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right]\widetilde{\mathbf{P}}'\right\Vert ^{2}\right\} \notag \\
&\leq\mathbb{E}\left\{ \left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{2}\left\Vert \left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right\Vert ^{2}\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\right\Vert ^{2}\right\} \notag \\
&\leq\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{6}\right]^{\frac{1}{3}}\mathbb{E}\left[\left\Vert \left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right\Vert ^{6}\right]^{\frac{1}{3}}\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\right\Vert ^{6}\right]^{\frac{1}{3}}.
\end{align}
The first term in (ref) can be written as
\begin{align}
&\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{6}\right] \notag\\
&=\mathbb{E}\left[\left(\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{2}\right)^{3}\right]\notag \\
&=\mathbb{E}\left[\left(\frac{1}{T}\sum_{t=1}^{T}\left\|\mathbf{p}(\widehat{\mathbf{f}}_{t})-\widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\|^{2}\right)^{3}\right]\notag \\
&=\mathbb{E}\left[\frac{1}{T^{3}}\sum_{t=1}^{T}\left\Vert \mathbf{p}(\widehat{\mathbf{f}}_{t})-\widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert ^{6}+\frac{3}{T^{3}}\sum_{t\neq s}\left\Vert \mathbf{p}(\widehat{\mathbf{f}}_{t})-\widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert ^{4}\left\Vert \mathbf{p}(\widehat{\mathbf{f}}_{s})-\widetilde{\mathbf{p}}(\mathbf{f}_{s})\right\Vert^2\right. \notag \\
&\quad \left.+\frac{6}{T^{3}}\sum_{t<s<r}\left\Vert \mathbf{p}(\widehat{\mathbf{f}}_{t})-\widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert^2 \left\Vert \mathbf{p}(\widehat{\mathbf{f}}_{s})-\widetilde{\mathbf{p}}(\mathbf{f}_{s})\right\Vert^2 \left\Vert \mathbf{p}(\widehat{\mathbf{f}}_{r})-\widetilde{\mathbf{p}}(\mathbf{f}_{r})\right\Vert^2\right] \notag \\
&=\mathbb{E}\left[\frac{1}{T^{3}}\sum_{t=1}^{T}\left\Vert \mathbf{p}(\widehat{\mathbf{f}}_{t})-\widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert ^{6}\right]\notag \\
&\quad+\mathbb{E}\left[\frac{3}{T^{3}}\sum_{t\neq s}\left\Vert \mathbf{p}(\widehat{\mathbf{f}}_{t})-\widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert ^{4}\left\Vert \mathbf{p}(\widehat{\mathbf{f}}_{s})-\widetilde{\mathbf{p}}(\mathbf{f}_{s})\right\Vert^2 \right]\notag \\
&\quad+\mathbb{E}\left[\frac{6}{T^{3}}\sum_{t<s<r}\left\Vert \mathbf{p}(\widehat{\mathbf{f}}_{t})-\widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert^2 \left\Vert \mathbf{p}(\widehat{\mathbf{f}}_{s})-\widetilde{\mathbf{p}}(\mathbf{f}_{s})\right\Vert^2 \left\Vert \mathbf{p}(\widehat{\mathbf{f}}_{r})-\widetilde{\mathbf{p}}(\mathbf{f}_{r})\right\Vert^2 \right],
\end{align} where we have used $\left(\sum_i a_i\right)^3=\sum_i a_i^3+3\sum_{i\neq l}a_i^2a_l + 6\sum_{i<l<m}a_ia_la_m$. While the first term in (ref) can be bounded using the same arguments as in (ref), the second and third terms can be bounded by arguments analogous to those used to evaluate (ref). It follows that
\begin{align}
\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{6}\right]&=O\left(\frac{K^{3}}{N^{3}T^{2}}\zeta_{1}(K)^{6}\right) + O\left(\frac{K^{3}}{N^{3}T}\zeta_{1}(K)^{6}\right) + O\left(\frac{K^{3}}{N^{3}}\zeta_{1}(K)^{6}\right)\notag \\
&= O\left(\frac{K^{3}}{N^{3}}\zeta_{1}(K)^{6}\right).
\end{align}
Terms of the form in (ref) and (ref) will reappear with different exponents throughout the proof. The key point is that their asymptotic behavior follows the same pattern; after pre-multiplying by $\frac{1}{T^\delta}$ with $\delta>0$, only the terms involving the maximal number of distinct indices survive. These sums scale with $O(T^\delta)$ while all other components grow more slowly, typically $O(T^{\delta-\eta})$ with $0<\eta\leq \delta$. After dividing by $T^\delta$ those lower-order terms vanish. From now on we will make use of this observation and Assumption (ref)(ref) to bound higher-order moments, instead of computing them explicitly.
We have evaluated the first term on the right-hand side of (ref). For the second term in the same equation, we use
\begin{align}
\widehat{\mathbf{P}}'\widehat{\mathbf{P}} = \widetilde{\mathbf{P}}'\widetilde{\mathbf{P}} + \left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\widetilde{\mathbf{P}} + \widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right) + \left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right).
\end{align}This allows us to write
\begin{align}
\left\|\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}-\*I_K\right\|
&\le \left\|\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}-\*I_K\right\| +2\left\|\frac{1}{T}(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}})'\widetilde{\mathbf{P}}\right\| +\left\|\frac{1}{\sqrt{T}}(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}})\right\|^2.
\end{align}By Lemma (ref)(ref), and the known orders of the variances of $\left\|\frac{1}{T}(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}})'\widetilde{\mathbf{P}}\right\|$ from (ref) and $\left\|\frac{1}{\sqrt{T}}(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}})\right\|^2$ from (ref), (ref) can be bounded as
\begin{align}
\left\|\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}-\*I_K\right\|
&= O_p\left(\sqrt{\frac{K}{T}}\zeta_0(K)\right)
+ O_p\left(\frac{K^{\frac{3}{4}}}{\sqrt{N}}\zeta_1(K)\sqrt{\zeta_0(K)}\right)
+ O_p\left(\frac{K}{N}\zeta_1(K)^2\right) \notag \\
&= O_p\left(\sqrt{\frac{K}{T}}\zeta_0(K)\right)
+ O_p\left(\frac{K^{\frac{3}{4}}}{\sqrt{N}}\zeta_1(K)\sqrt{\zeta_0(K)}\right),
\end{align} since $\frac{\frac{K}{N}\zeta_1(K)^2}{\sqrt{\frac{K}{T}}\zeta_0(K)} = \frac{\sqrt{KT}}{N}\frac{\zeta_1(K)^2}{\zeta_0(K)}=o(1)$ by (ref) and $\frac{\frac{K}{N}\zeta_1(K)^2}{\frac{K^{\frac{3}{4}}}{\sqrt{N}}\zeta_1(K)\sqrt{\zeta_0(K)}} = \frac{K^{\frac{1}{4}}\zeta_1(K)}{\sqrt{N}\sqrt{\zeta_0(K)}}=o(1)$ by (ref). Hence, the third term above is dominated by the first two. The two terms that remain will appear again and again. It is therefore convenient to introduce
\begin{align}
\delta_{NTK} \equiv \max\left\{\sqrt{\frac{K}{T}}\zeta_0(K), \, \frac{K^{\frac{3}{4}}}{\sqrt{N}}\zeta_1(K)\sqrt{\zeta_0(K)}\right\}.
\end{align} By (ref) and (ref) we have
\begin{align}
\delta_{NTK}=o(1).
\end{align}
This result will be used repeatedly in the sequel. We continue by noting that
\begin{align}
\left\|\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}-\*I_K\right\| =O_p\left(\delta_{NTK}\right).
\end{align}
By Lemma (ref)(ref), asymptotically $\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}$ will have full rank $K$, as $N \to \infty$, and therefore the Moore-Penrose inverse of this matrix reduces to the standard inverse. This implies that
\begin{align}
\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ - \*I_K = \left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{-1} - \*I_K = -\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{-1}\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}} - \*I_K\right).
\end{align}
If $\lambda$ is an eigenvalue of $\*A$, then $\lambda^2$ is an eigenvalue of $\*A^2$ and $\lambda^{-1}$ is an eigenvalue of $\*A^{-1}$. Moreover, because $\lambda_{\max}(\*A) \geq \lambda_{\min}(\*A)$, we have $\lambda_{\max}(\*A)^{-1} \leq \lambda_{\min}(\*A)^{-1}$. Together with $\mathrm{tr}\,(\*B\*A\*B') \leq \lambda_{\max}(\*A) \mathrm{tr}\,(\*B\*B')$, these results yield
\begin{align}
\left\|\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ - \*I_K \right\|^2 &= \mathrm{tr}\left[\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}} - \*I_K\right)\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{-1}\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{-1}\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}} - \*I_K\right)\right] \notag\\
&= \mathrm{tr}\left[\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}} - \*I_K\right)\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{-2}\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}} - \*I_K\right)\right] \notag \\
&\leq \lambda_{\max}\left[\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{-1}\right]^2\mathrm{tr}\left[\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}} - \*I_K\right)\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}} - \*I_K\right)\right] \notag \\
&\leq \lambda_{\max}\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{-2}\mathrm{tr}\left[\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}} - \*I_K\right)\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}} - \*I_K\right)\right] \notag \\
& \leq \lambda_{\min}\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{-2}\left\|\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}} - \*I_K\right)\right\|^2.
\end{align} Hence, by Lemma (ref)(ref) and the result in (ref)
\begin{align}
\left\|\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ - \*I_K \right\|& \leq \lambda_{\min}\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{-1}\left\|\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}} - \*I_K\right)\right\|\notag \\
&= O_p(1) O_p\left(\delta_{NTK}\right) \notag \\
&=O_p\left(\delta_{NTK}\right).
\end{align} For convenience, let us define $\left\|\*E_K\right\|\equiv \left\|\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ - \*I_K \right\|=\left\|\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{-1} - \*I_K \right\|=O_p(\delta_{NTK})$. To bound the second term in (ref), we use that $\left\|\*E_K\right\|^6$ is non-negative. It follows that
\begin{align}
\mathbb{E}\left[\left\Vert \mathbf{E}_{K}\right\Vert ^{6}\right] =\int_0^\infty \mathbb{P}\left(\left\Vert \mathbf{E}_{K}\right\Vert ^{6}\geq x\right)dx,
\end{align} where $\mathbb{P}(A)$ is the probability of the event $A$. Performing change of variables, such that $x=t^6$ and consequently $dx=6t^5 dt$ yields
\begin{align}
\mathbb{E}\left[\left\Vert \mathbf{E}_{K}\right\Vert ^{6}\right]& = \int_0^\infty \mathbb{P}\left(\left\Vert \mathbf{E}_{K}\right\Vert ^{6}\geq x\right)dx = 6\int_0^\infty t^5\mathbb{P}\left(\left\Vert \mathbf{E}_{K}\right\Vert ^{6}\geq t^6\right)dt\notag \\
&=6\int_{0}^{\infty}t^{5}\mathbb{P}\left(\left\Vert \mathbf{E}_{K}\right\Vert \geq t\right)dt\notag \\
&=6\int_{0}^{C_1\delta_{NTK}}t^{5}\mathbb{P}\left(\left\Vert \mathbf{E}_{K}\right\Vert \geq t\right)dt+6\int_{C_1\delta_{NTK}}^{\infty}t^{5}\mathbb{P}\left(\left\Vert \mathbf{E}_{K}\right\Vert \geq t\right)dt\notag\\
&\equiv I_{1}+I_{2},
\end{align} with implicit definitions of $I_1$ and $I_2$. We can bound $I_1$ as
\begin{align}
I_{1}&\equiv6\int_{0}^{C_{1}\delta_{NTK}}t^{5}\mathbb{P}\left(\left\Vert \mathbf{E}_{K}\right\Vert \geq t\right)dt \leq6\int_{0}^{C_{1}\delta_{NTK}}t^{5}\cdot1dt =\left[t^{5}\right]_{0}^{C_{1}\delta_{NTK}}+C_{2}\notag \\
&=C_{1}^{5}\delta_{NTK}^{5}+C_{2},
\end{align} since $\mathbb{P}\left(\left\Vert \mathbf{E}_{K}\right\Vert \geq t\right)\leq 1$. Therefore, $I_{1}=O\left(\delta_{NTK}^{5}\right)$. Let us now consider $I_2$;
\begin{align}
I_{2}&\equiv6\int_{C_{2}\delta_{NTK}}^{\infty}t^{5}\mathbb{P}\left(\left\Vert \mathbf{E}_{K}\right\Vert \geq t\right)dt\le6\int_{C_{2}\delta_{NTK}}^{\infty}t^{5}C_{3}\exp(-C_{4}t^{2})dt\notag \\
&=6C_{3}\int_{C_{2}\delta_{NTK}}^{\infty}t^{4}\exp(-C_{4}t^{2})tdt,
\end{align} where we have used the fact that bounded random variables belong to the class of sub-gaussian variables, such that one can bound $\mathbb{P}\left(\left\Vert \mathbf{E}_{K}\right\Vert \geq t\right)\leq C_{3}\exp(-C_{4}t^{2})$ (see, for example, rudel08). Let $u=t^2$ implying $du=2tdt$ and in turn $tdt=\frac{1}{2}du$. Integration by substitution and by parts now yields
\begin{align}
I_{2}&\leq6C_{3}\int_{C_{2}\delta_{NTK}}^{\infty}t^{4}\exp(-C_{4}t^{2})tdt\notag \\
&=3C_{3}\lim_{r\to\infty}\int_{C_{2}^{2}\delta_{NTK}^{2}}^{r}u^{2}\exp(-C_{4}u)du\notag \\
&=-\frac{3C_3}{C_{4}}\lim_{r\to\infty}\left[u^{2}\exp(-C_{4}u)\right]_{C_{2}^{2}\delta_{NTK}^{2}}^{r}+\frac{3C_3}{C_{4}}\lim_{r\to\infty}\int_{C_{2}^{2}\delta_{NTK}^{2}}^{r}2u\exp(-C_{4}u)du\notag \\
&=-\frac{3C_3}{C_{4}}\lim_{r\to\infty}\left[u^{2}\exp(-C_{4}u)\right]_{C_{2}^{2}\delta_{NTK}^{2}}^{r}-\frac{6C_3}{C_{4}^2}\lim_{r\to\infty}\left[u\exp(-C_{4}u)\right]_{C_{2}^{2}\delta_{NTK}^{2}}^{r}\notag \\
&\quad +\frac{6C_3}{C_{4}^2}\lim_{r\to\infty}\int\exp(-C_{4}u)du\notag \\
&=-\frac{3C_3}{C_{4}}\lim_{r\to\infty}\left[u^{2}\exp(-C_{4}u)\right]_{C_{2}^{2}\delta_{NTK}^{2}}^{r}-\frac{6C_3}{C_{4}^2}\lim_{r\to\infty}\left[u\exp(-C_{4}u)\right]_{C_{2}^{2}\delta_{NTK}^{2}}^{r}\notag \\
&\quad -\frac{6C_3}{C_{4}^{3}}\lim_{r\to\infty}\left[\exp(-C_{4}u)\right]_{C_{2}^{2}\delta_{NTK}^{2}}^{r}\notag \\
&=\lim_{r\to\infty}\left[-\left(\frac{3C_3}{C_4}u^2+\frac{6C_3}{C_4^2}u+\frac{6C_3}{C_4^2}\right)\exp\left(-C_4 u\right)\right]^r_{C_2^2\delta_{NTK}^2}\notag \\
&=\left(\frac{3C_3C_2^2}{C_4}\delta_{NTK}^4 + \frac{6C_3C_2^2}{C_4^2}\delta_{NTK}^2+\frac{6C_3}{C_4^2}\right)\exp\left(-C_4C_2^2\delta_{NTK}^2\right),
\end{align} where $C_1,\dots, C_4$ are nonnegative constants. Since $\delta_{NTK}\to0$ and by the continuous mapping theorem, $\exp(-\delta_{NTK}^{2})\to\exp(0)=1$, we can show that $I_2=O(1)$.
Thus,
\begin{align}
\mathbb{E}\left[\left\Vert \mathbf{E}_{K}\right\Vert ^{6}\right] \equiv \mathbb{E}\left[\left\|\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ - \*I_K \right\|^6\right]=I_1 + I_2 = O\left(\delta_{NTK}^5\right) + O(1) = O(1),
\end{align} since $\delta_{NTK}\to 0$.
We are left to bound the third term in (ref). Here, we can rely on the same steps used to derive (ref) to arrive at
\begin{align}
\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\right\Vert ^{6}\right] &= \mathbb{E}\left[\left(\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\right\Vert ^{2}\right)^3\right] = \mathbb{E}\left[\left(\frac{1}{T}\sum_{t=1}^{T}\left\Vert \widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert ^{2}\right)^{3}\right]\notag \\
&\leq K\zeta_0(K)^4.
\end{align}
By combining the results, we obtain the following bound for (ref):
\begin{align}
&\mathbb{E}\left\{ \left\Vert \frac{1}{T}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left[\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right]\widetilde{\mathbf{P}}'\right\Vert ^{2}\right\} \notag \\
&\leq\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{6}\right]^{\frac{1}{3}}\mathbb{E}\left[\left\Vert \left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right\Vert ^{6}\right]^{\frac{1}{3}}\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\right\Vert ^{6}\right]^{\frac{1}{3}} \notag \\
&= \left[O\left(\frac{K^3}{N^3}\zeta_1(K)^6\right)\right]^{\frac{1}{3}}\left[O\left(1\right)\right]^{\frac{1}{3}}\left[O\left(K\zeta_0(K)^4\right)\right]^{\frac{1}{3}} \notag \\
&=O\left(\frac{K^{\frac{4}{3}}}{N}\zeta_1(K)^2\zeta_0(K)^{\frac{4}{3}}\right),
\end{align} and consequently,
\begin{align}
\mathbb{E}\left[\left\|\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ \widetilde{\mathbf{P}}' \right\|^2\right] &\leq 2\mathbb{E}\left[\left\|\frac{1}{T}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right) \widetilde{\mathbf{P}}' \right\|^2\right]\notag \\
&\quad + 2\mathbb{E}\left[\left \|\frac{1}{T}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left[\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ - \*I_K \right] \widetilde{\mathbf{P}}' \right\|^2\right] \notag\\
&=O\left(\frac{K^{\frac{3}{2}}}{N}\zeta_1(K)^2\zeta_0(K)\right) + O\left(\frac{K^{\frac{4}{3}}}{N}\zeta_1(K)^2\zeta_0(K)^{\frac{4}{3}}\right).
\end{align} Therefore,
\begin{align}
\mathbb{E}\left(\left\|\*d_{5,2}\right\|^2\right) &\equiv \mathbb{E}\left[\left\|\frac{1}{\sqrt{NT}}\sum_{i=1}^N \*V_i'\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}' \widehat{\mathbf{P}}\right)^+ \widetilde{\mathbf{P}}'\+\varepsilon_i\right\|^2\right] \notag \\
&\leq O\left(\frac{1}{T}\right)\mathbb{E}\left[\left\|\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}' \widehat{\mathbf{P}}\right)^+ \widetilde{\mathbf{P}}'\right\|^2\right]\notag \\
&= O\left(\frac{K^{\frac{3}{2}}}{NT}\zeta_1(K)^2\zeta_0(K)\right) + O\left(\frac{K^{\frac{4}{3}}}{NT}\zeta_1(K)^2\zeta_0(K)^{\frac{4}{3}}\right).
\end{align} This allows us to conclude that
\begin{align}
\left\|\*d_{5,2}\right\| = O_p\left(\frac{K^{\frac{3}{4}}}{\sqrt{NT}}\sqrt{\zeta_0(K)}\zeta_1(K)\right) + O_p\left(\frac{K^{\frac{2}{3}}}{\sqrt{NT}}\zeta_0(K)^{\frac{2}{3}}\zeta_1(K)\right).
\end{align}
The above arguments can be applied also to the other terms in (ref). The third term, $\*d_{5,3} \equiv \frac{1}{\sqrt{NT}}\sum_{i=1}^N \*V_i'\widetilde{\mathbf{P}}\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\+\varepsilon_i$, is of the same order as $\left\|\*d_{5,2}\right\|$;
\begin{align}
\left\|\*d_{5,3}\right\| = O_p\left(\frac{K^{\frac{3}{4}}}{\sqrt{NT}}\sqrt{\zeta_0(K)}\zeta_1(K)\right) + O_p\left(\frac{K^{\frac{2}{3}}}{\sqrt{NT}}\zeta_0(K)^{\frac{2}{3}}\zeta_1(K)\right).
\end{align}
For $\*d_{5,4}\|$, we can use the same steps as for analyzing $\mathbb{E}\left(\left\|\*d_{5,2}\right\|^2\right)$ to obtain
\begin{align}
&\mathbb{E}\left(\left\|\*d_{5,4}\right\|^2\right) \notag\\
& = \mathbb{E} \left[\left\|\frac{1}{\sqrt{NT}}\sum_{i=1}^N \*V_i'\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ - \left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^+\right]\widetilde{\mathbf{P}}'\+\varepsilon_i\right\|^2\right] \notag \\
&\leq \mathbb{E}\left[\mathrm{tr}\left(\frac{1}{NT}\sum_{i=1}^N\sum_{j=1}^N\+\varepsilon_i'\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ - \left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^+\right]\widetilde{\mathbf{P}}'\mathbb{E}\left(\*V_i \*V_j'\right)\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ - \left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^+\right]\widetilde{\mathbf{P}}'\+\varepsilon_j\right)\right] \notag \\
&= \mathbb{E}\left[\mathrm{tr}\left(\frac{1}{NT}\sum_{i=1}^N\+\varepsilon_i'\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ - \left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^+\right]\widetilde{\mathbf{P}}'\+\Omega_{v,i}\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ - \left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^+\right]\widetilde{\mathbf{P}}'\+\varepsilon_i\right)\right] \notag \\
&\leq \frac{1}{NT} \sum_{i=1}^N \lambda_{\max}\left(\+\Omega_{v,i}\right)\mathbb{E}\left[\mathrm{tr}\left(\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ - \left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^+\right]\widetilde{\mathbf{P}}'\mathbb{E}\left(\+\varepsilon_i\+\varepsilon_i'\right)\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ - \left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^+\right]\widetilde{\mathbf{P}}'\right)\right] \notag \\
&\leq \frac{1}{NT} \sum_{i=1}^N \lambda_{\max}\left(\+\Omega_{v,i}\right)\lambda_{\max}\left(\+\Omega_{\varepsilon,i}\right)\notag\\
&\quad \times \mathbb{E}\left[\mathrm{tr}\left(\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ - \left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^+\right]\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ - \left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^+\right]\widetilde{\mathbf{P}}'\right)\right] \notag \\
&= \frac{1}{NT} \sum_{i=1}^N \lambda_{\max}\left(\+\Omega_{v,i}\right)\lambda_{\max}\left(\+\Omega_{\varepsilon,i}\right)\mathbb{E}\left(\left\|\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ - \left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^+\right]\widetilde{\mathbf{P}}'\right\|^2\right) \notag \\
& = \sup_{i=1, \dots, N}\lambda_{\max}\left(\+\Omega_{v,i}\right)\sup_{i=1, \dots, N}\lambda_{\max}\left(\+\Omega_{\varepsilon,i}\right) \frac{1}{NT}\sum_{i=1}^N\mathbb{E}\left(\left\|\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ - \left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^+\right]\widetilde{\mathbf{P}}'\right\|^2\right) \notag \\
&= \sup_{i=1, \dots, N}\lambda_{\max}\left(\+\Omega_{v,i}\right)\sup_{i=1, \dots, N}\lambda_{\max}\left(\+\Omega_{\varepsilon,i}\right) \frac{1}{T}\mathbb{E}\left(\left\|\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ - \left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^+\right]\widetilde{\mathbf{P}}'\right\|^2\right) \notag \\
&= O\left(\frac{1}{T}\right)\mathbb{E}\left(\left\|\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ - \left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^+\right]\widetilde{\mathbf{P}}'\right\|^2\right).
\end{align} Further use of
\begin{align}
\left\|\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right\|&=\left\|\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}+\mathbf{I}_{K}-\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right\|\notag\\
&=\left\|\left[\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right]-\left[\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right]\right\|\notag\\
&\leq\left\|\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right\|+\left\|\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right\|,
\end{align}
and the Cauchy-Schwarz inequality yields
\begin{align}
&\mathbb{E}\left(\left\Vert \widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\widetilde{\mathbf{P}}'\right\Vert ^{2}\right) \notag \\
&=\mathbb{E}\left\{ \mathrm{tr}\left[\frac{1}{T}\widetilde{\mathbf{P}}\left[\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\left[\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\widetilde{\mathbf{P}}'\right]\right\}\notag \\
&=\mathbb{E}\left\{ \mathrm{tr}\left[\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\left[\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\left[\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\right]\right\}\notag \\
&=\mathbb{E}\left\{ \left\|\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\left[\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\right\|^{2}\right\}\notag \\
&\leq\mathbb{E}\left[\left\|\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right\|^{2}\left\|\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right\|^{2}\right]\notag\\
&\leq\sqrt{\mathbb{E}\left[\left\Vert \frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right\Vert ^{4}\right]}\sqrt{\mathbb{E}\left[\left\Vert \left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right\Vert ^{4}\right]}\notag\\
&\leq \sqrt{\mathbb{E}\left[\left\Vert \frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right\Vert ^{4}\right]}\sqrt{\mathbb{E}\left[\left\Vert \left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right\Vert ^{4}\right]+\mathbb{E}\left[\left\|\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right\|^{4}\right]}\notag \\
&\leq \sqrt{\mathbb{E} \left[\left\|\frac{1}{\sqrt{T}} \widetilde{\mathbf{P}}\right\|^4 \left\|\frac{1}{\sqrt{T}} \widetilde{\mathbf{P}}\right\|^4\right]}\sqrt{\mathbb{E}\left[\left\Vert \left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right\Vert ^{4}\right]+\mathbb{E}\left[\left\|\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right\|^{4}\right]} \notag\\
&\leq \sqrt{\mathbb{E}\left[\left\|\frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\right\|^8\right]}\sqrt{\mathbb{E}\left[\left\Vert \left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right\Vert ^{4}\right]+\mathbb{E}\left[\left\|\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right\|^{4}\right]}.
\end{align} From our previous derivations in (ref) and (ref), we have
\begin{align}
\mathbb{E}\left[\left\|\frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\right\|^8\right]=K\zeta_0(K)^6,
\end{align} and
\begin{align}
\mathbb{E}\left[\left\|\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+-\*I_K\right\|^4\right] = O(1).
\end{align} Thus, we are left to bound $\mathbb{E}\left[\left\|\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^+ - \*I_K\right\|^4\right]$. Here, we can employ exactly the same strategy we used to bound (ref). By Lemma (ref)(ref), asymptotically $\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}$ will have full rank $K$, as $N\to \infty$ and therefore, the Moore-Penrose inverse of this matrix reduces to the standard inverse also in this case. Thus, similarly to (ref)--(ref), we obtain
\begin{align}
\left\|\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^+ - \*I_K\right\| &\leq \lambda_{\min} \left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{-1}\left\|\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right\| \notag\\
&=O_p(1) O_p\left(\sqrt{\frac{K}{T}}\zeta_0(K)\right)\notag\\
&=O_p\left(\sqrt{\frac{K}{T}}\zeta_0(K)\right),
\end{align} which holds by Lemma (ref)(ref) and (ref)(ref). Moreover, we can use exactly the same arguments as in (ref)--(ref) to show that
\begin{align}
\mathbb{E}\left[\left\|\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^+ - \*I_K\right\|^4\right] = O\left(\frac{K^2}{T^2}\zeta_0(K)^4\right)+O(1)=O(1),
\end{align} since $\frac{K^2}{T^2}\zeta_0(K)^4=o\left(\frac{K^2\zeta_0(K)^4}{K^2\zeta_0(K)^6}\right)=o\left(\frac{1}{\zeta_0(K)^2}\right)=o(1)$ under Assumption (ref)(ref).
Therefore, the bound for the expression in (ref) is given by
\begin{align}
&\mathbb{E}\left(\left\Vert \widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\widetilde{\mathbf{P}}'\right\Vert ^{2}\right)\notag\\
&\leq \sqrt{\mathbb{E}\left[\left\|\frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\right\|^8\right]}\sqrt{\mathbb{E}\left[\left\Vert \left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right\Vert ^{4}\right]+\mathbb{E}\left[\left\|\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right\|^{4}\right]}\notag \\
&=\sqrt{O\left(K\zeta_0(K)^6\right)}\sqrt{O(1)}\notag\\
&=O\left(\sqrt{K}\zeta_0(K)^3\right),
\end{align} yielding
\begin{align}
\mathbb{E}\left(\left\|\*d_{5,4}\right\|^2\right) &\leq O\left(\frac{1}{T}\right)\mathbb{E}\left(\left\|\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ - \left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^+\right]\widetilde{\mathbf{P}}'\right\|^2\right)\notag\\
&=O\left(\frac{\sqrt{K}}{T}\zeta_0(K)^3\right),
\end{align} and thus,
\begin{align}
\left\|\*d_{5,4}\right\| = O_p\left(\frac{K^{\frac{1}{4}}}{\sqrt{T}}\zeta_0(K)^{\frac{3}{2}}\right).
\end{align}
The first term in (ref), $\*d_{5,1}\equiv \frac{1}{\sqrt{NT}}\sum_{i=1}^N \*V_i'\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}' \widehat{\mathbf{P}}\right)^+\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\+\varepsilon_i$, can be analyzed very similarly to (ref). In particular,
\begin{align}
\mathbb{E}\left(\left\|\mathbf{d}_{5,1}\right\|^{2}\right) &=\mathbb{E}\left[\left\|\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\boldsymbol{\varepsilon}_{i}\right\|^{2}\right] \notag\\
&\leq O\left(\frac{1}{T}\right)\mathbb{E}\left[\left\|\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\right\|^{2}\right],
\end{align} where
\begin{align}
&\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\right\Vert ^{2}\right]\notag\\
&\leq2\mathbb{E}\left[\left\|\frac{1}{T}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\right\|^{2}\right]+2\mathbb{E}\left[\left\|\frac{1}{T}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left[\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right]\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\right\|^{2}\right]\notag\\
&\leq2\mathbb{E}\left[\left\|\frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\|^{4}\right]+2\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{2}\left\Vert \left[\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right]\right\Vert ^{2}\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{2}\right]\notag\\
&\leq2\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{4}\right]+2\left\{ \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{6}\right]\right\} ^{\frac{1}{3}}\left\{ \mathbb{E}\left[\left\Vert \left[\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right]\right\Vert ^{6}\right]\right\} ^{\frac{1}{3}}\notag\\
& \quad \times \left\{ \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{6}\right]\right\} ^{\frac{1}{3}}\notag\\
&=O\left(\frac{K^{2}}{N^{2}}\zeta_{1}(K)^{4}\right)+O\left(\frac{K^{3}}{N^{3}}\zeta_{1}(K)^{6}\right)^{\frac{2}{3}}O(1)\notag\\
&=O\left(\frac{K^{2}}{N^{2}}\zeta_{1}(K)^{4}\right)+O\left(\frac{K^{2}}{N^{2}}\zeta_{1}(K)^{4}\right)\notag\\
&=O\left(\frac{K^{2}}{N^{2}}\zeta_{1}(K)^{4}\right).
\end{align} Here we have used the results in (ref), (ref) and (ref). Thus,
\begin{align}
\left\|\*d_{5,1}\right\| = O_p\left(\frac{K}{NT}\zeta_1(K)^2\right).
\end{align}
By adding the results,
\begin{align}
\left\|\*D_5\right\| &= \left\|\frac{1}{\sqrt{NT}}\sum_{i=1}^N \*V_{i}' (\*M_{\widetilde{P}} - \*M_{\widehat P}) \+\varepsilon_{i} \right\| \notag \\
&\leq \left\|\*d_{5,1}\right\| + \left\|\*d_{5,2}\right\| + \left\|\*d_{5,3}\right\| + \left\|\*d_{5,4}\right\| \notag \\
&=O_p\left(\frac{K}{NT}\zeta_1(K)^2\right) + O_p\left(\frac{K^{\frac{3}{4}}}{\sqrt{NT}}\sqrt{\zeta_0(K)}\zeta_1(K)\right) \notag \\
&\quad + O_p\left(\frac{K^{\frac{2}{3}}}{\sqrt{NT}}\zeta_0(K)^{\frac{2}{3}}\zeta_1(K)\right) + O_p\left(\frac{K^{\frac{1}{4}}}{\sqrt{T}}\zeta_0(K)^{\frac{3}{2}}\right)\notag \\
&= o_p(1),
\end{align} where we have used (ref), (ref) and (ref) to show the last equality.
We now move to $\mathbf{D}_{6},$ which we
again expand using (ref);
\begin{align}
\mathbf{D}_{6} & =\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'(\mathbf{M}_{\widetilde{P}}-\mathbf{M}_{\widehat{P}})\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\nonumber \\
& =\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\nonumber \\
& \quad +\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\nonumber \\
& \quad +\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\widetilde{\mathbf{P}}\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\nonumber \\
& \quad +\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\nonumber \\
& \equiv \*d_{6,1}+ \*d_{6,2} + \*d_{6,3} + \*d_{6,4}.
\end{align}
As before, we proceed by evaluating each of the four terms on the
right-hand side of (ref). In light of the decomposition in
(ref), it is important to note that each of
these four terms is actually comprised of two terms, as each component
in $\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)$
must be evaluated separately. We start by considering the second term in (ref):
\begin{align}
\*d_{6,2} & \equiv\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\nonumber \\
& =\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\+\alpha_{g_i}\nonumber \\
& \quad +\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right) \notag \\
&\equiv \*d^{(1)}_{6,2} + \*d^{(2)}_{6,2}.
\end{align}
Both terms on the right can be treated the same way as $\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\boldsymbol{\varepsilon}_{i}.$
The treatment of the second term is actually almost identical. The
reason is that $\mathbf{V}_{i}$ is mean zero and independent across
$i$. It is also independent of $\mathbf{F}$, and thus of any function
thereof, including $\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})$.
Hence, using the same arguments as for $\*d_{5,2}$, we can consider the following variance expression for $\*d_{6,2}^{(2)}$:
\begin{align}
& \mathbb{E}\left(\|\*d_{6,2}^{(2)}\|^2 \right)\\
& = \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right\Vert^2\right] \notag \\
& \leq \mathbb{E}\left[\mathrm{tr}\left(\frac{1}{NT}\sum_{i=1}^N\sum_{j=1}^N\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)'\widetilde{\mathbf{P}}\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\mathbb{E}\left(\*V_i\mathbf{V}_{j}'\right) \right.\right. \notag\\
& \quad \times \left. \left. \left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\left(\widetilde{\mathbf{P}}\+\alpha_{g_j}-\*g_j(\mathbf{F})\right)\right)\right] \notag \\
&\leq \frac{1}{NT}\sum_{i=1}^N\lambda_{\max}\left(\+\Omega_{v,i}\right) \notag\\
& \quad \times \mathbb{E}\left[\mathrm{tr}\left(\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)'\widetilde{\mathbf{P}}\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right)\right] \notag \\
&=\sup_{i=1,\dots,N}\lambda_{\max}(\boldsymbol{\Omega}_{v,i})\mathbb{E}\left[\frac{1}{NT}\sum_{i=1}^{N}\left\|\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\left(\widetilde{\mathbf{P}}\boldsymbol{\alpha}_{g_{i}}-\mathbf{g}_{i}(\mathbf{F})\right)\right\|^{2}\right]\notag\\
&\leq\sup_{i=1,\dots,N}\lambda_{\max}(\boldsymbol{\Omega}_{v,i})\notag\\
&\quad\times\mathbb{E}\left[\left\|\frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\|^{2}\left\Vert \left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\right\Vert ^{2}\left\|\frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\right\|^{2}\frac{1}{N}\sum_{i=1}^{N}\left\|\frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\boldsymbol{\alpha}_{g_{i}}-\mathbf{g}_{i}(\mathbf{F})\right)\right\|^{2}\right]\notag\\
&\leq\sup_{i=1,\dots,N}\lambda_{\max}(\boldsymbol{\Omega}_{v,i})\left\{ \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}\left\{ \mathbb{E}\left[\left\Vert \left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}\notag\\
&\quad\times\left\{ \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}\frac{1}{N}\sum_{i=1}^{N}\left\{ \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\boldsymbol{\alpha}_{g_{i}}-\mathbf{g}_{i}(\mathbf{F})\right)\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}\notag\\
&\leq\sup_{i=1,\dots,N}\lambda_{\max}(\boldsymbol{\Omega}_{v,i})\left\{ \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}\left\{ \mathbb{E}\left[\left\Vert \left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}+\mathbf{I}_{K}\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}\notag\\
&\quad\times\left\{ \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}\frac{1}{N}\sum_{i=1}^{N}\left\{ \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\boldsymbol{\alpha}_{g_{i}}-\mathbf{g}_{i}(\mathbf{F})\right)\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}\notag\\
&\leq\sup_{i=1,\dots,N}\lambda_{\max}(\boldsymbol{\Omega}_{v,i})\left\{ \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}\notag\\
&\quad\times\left\{ \mathbb{E}\left[\left\Vert \left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right\Vert ^{8}\right]+\mathbb{E}\left(\left\Vert \mathbf{I}_{K}\right\Vert ^{8}\right)\right\} ^{\frac{1}{4}}\left\{ \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}\notag\\
&\quad\times\frac{1}{N}\sum_{i=1}^{N}\left\{ \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\boldsymbol{\alpha}_{g_{i}}-\mathbf{g}_{i}(\mathbf{F})\right)\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}\notag\\
&=O(1)\left[O\left(\frac{K^{4}}{N^{4}}\zeta_{1}(K)^{8}\right)\right]^{\frac{1}{4}}\left[O(1)+O\left(K^{8}\right)\right]^{\frac{1}{4}}\left[K\zeta_{0}(K)^{6}\right]^{\frac{1}{4}}\notag\\
&\quad\times\frac{1}{N}\sum_{i=1}^{N}\left\{ \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\boldsymbol{\alpha}_{g_{i}}-\mathbf{g}_{i}(\mathbf{F})\right)\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}\notag\\
&=O\left(\frac{K^{\frac{13}{4}}}{N}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)^{2}\right)\frac{1}{N}\sum_{i=1}^{N}\left\{ \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\boldsymbol{\alpha}_{g_{i}}-\mathbf{g}_{i}(\mathbf{F})\right)\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}.
\end{align} Most of the terms in (ref) are either known from previous derivations or can be extrapolated using those. The only new term is
\begin{align}
&\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\boldsymbol{\alpha}_{g_{i}}-\mathbf{g}_{i}(\mathbf{F})\right)\right\Vert ^{8}\right]\notag\\
&=\mathbb{E}\left[\left(\frac{1}{T}\sum_{t=1}^{T}\left\|g_{i}(\mathbf{f}_{t})-\boldsymbol{\alpha}'_{g_{i}}\widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\|^{2}\right)^{4}\right]\notag\\
&=\frac{1}{T^{4}}\sum_{t=1}^{T}\mathbb{E}\left[\left\Vert g_{i}(\mathbf{f}_{t})-\boldsymbol{\alpha}'_{g_{i}}\widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert ^{8}\right]\notag\\
&\quad+\frac{4}{T^{4}}\sum_{t\neq s}\mathbb{E}\left[\left\Vert g_{i}(\mathbf{f}_{t})-\boldsymbol{\alpha}'_{g_{i}}\widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert ^{6}\left\Vert g_{i}(\mathbf{f}_{s})-\boldsymbol{\alpha}'_{g_{i}}\widetilde{\mathbf{p}}(\mathbf{f}_{s})\right\Vert ^{2}\right]\notag\\
&\quad+\frac{6}{T^{4}}\sum_{t<s}\mathbb{E}\left[\left\Vert g_{i}(\mathbf{f}_{t})-\boldsymbol{\alpha}'_{g_{i}}\widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert ^{4}\left\Vert g_{i}(\mathbf{f}_{s})-\boldsymbol{\alpha}'_{g_{i}}\widetilde{\mathbf{p}}(\mathbf{f}_{s})\right\Vert ^{4}\right]\notag\\
&\quad+\frac{12}{T^{4}}\sum_{t\neq s\neq r}\mathbb{E}\left[\left\Vert g_{i}(\mathbf{f}_{t})-\boldsymbol{\alpha}'_{g_{i}}\widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert ^{4}\left\Vert g_{i}(\mathbf{f}_{s})-\boldsymbol{\alpha}'_{g_{i}}\widetilde{\mathbf{p}}(\mathbf{f}_{s})\right\Vert ^{2}\left\Vert g_{i}(\mathbf{f}_{r})-\boldsymbol{\alpha}'_{g_{i}}\widetilde{\mathbf{p}}(\mathbf{f}_{r})\right\Vert ^{2}\right]\notag\\
&\quad+\frac{24}{T^{4}}\sum_{t<s<r<q}\mathbb{E}\left[\left\Vert g_{i}(\mathbf{f}_{t})-\boldsymbol{\alpha}'_{g_{i}}\widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert ^{2}\left\Vert g_{i}(\mathbf{f}_{s})-\boldsymbol{\alpha}'_{g_{i}}\widetilde{\mathbf{p}}(\mathbf{f}_{s})\right\Vert ^{2}\left\Vert g_{i}(\mathbf{f}_{r})-\boldsymbol{\alpha}'_{g_{i}}\widetilde{\mathbf{p}}(\mathbf{f}_{r})\right\Vert ^{2} \right.\notag \\
&\quad \left. \times \left\Vert g_{i}(\mathbf{f}_{q})-\boldsymbol{\alpha}'_{g_{i}}\widetilde{\mathbf{p}}(\mathbf{f}_{q})\right\Vert ^{2}\right].
\end{align} By Assumption (ref)(ref) and (ref), for $g_{i}(\cdot)\in \Lambda_r^{\lambda_i}(\mathcal{R}_f, \omega_i)$ there exists $\Pi_{\infty K}g_{i}(\cdot)=\+\alpha'_{g_i}\*p(\cdot)\in \mathfrak{G}_K$, such that for any $\overline{\omega} > \sup_{i=1,\dots, N}(\omega_i + \lambda_i)$ we can write
\begin{align}
\left\|g_i(\cdot)-\Pi_{\infty K}g_i(\cdot)\right\|_{\infty, \overline{\omega}} = \sup_{\*f \in \mathcal{R}_f} \left|\left[g_i(\*f)-\+\alpha'_{g_i}\*p(\*f)\right]\left[1+\left\|\*f\right\|^2\right]^{-\frac{\overline{\omega}}{2}}\right| \leq CK^{-\frac{\lambda_i}{m}},
\end{align} where we have used the definition of the weighted sup-norm from (ref). Then, by Assumption (ref),
\begin{align}
&\frac{1}{T^{4}}\sum_{t=1}^{T}\mathbb{E}\left[\left\Vert g_{i}(\mathbf{f}_{t})-\boldsymbol{\alpha}'_{g_{i}}\widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert ^{8}\right] = \frac{1}{T^{4}}\sum_{t=1}^{T}\mathbb{E}\left[\left(\left\Vert g_{i}(\mathbf{f}_{t})-\boldsymbol{\alpha}'_{g_{i}}\widetilde{\mathbf{p}}(\mathbf{f}_{t})\right\Vert ^{2}\right)^{4}\right]\notag\\
&=\frac{1}{T^{3}}\int_{\mathcal{R}_f}\left[\left|g_{i}(\mathbf{f}_{t})-\boldsymbol{\alpha}'_{g_{i}}\widetilde{\mathbf{p}}(\mathbf{f}_{t})\right|^{2}\right]^{4}w_{f}(\mathbf{f}_{t})d\mathbf{f}_{t}\notag\\
&=\frac{1}{T^{3}}\int_{\mathcal{R}_f}\left\{ \left[\left(g_{i}(\mathbf{f}_{t})-\boldsymbol{\alpha}'_{g_{i}}\widetilde{\mathbf{p}}(\mathbf{f}_{t})\right)\left(1+\left\Vert \mathbf{f}_{t}\right\Vert ^{2}\right)^{-\frac{\overline{\omega}}{2}}\right]^{2}\left(1+\left\Vert \mathbf{f}_{t}\right\Vert ^{2}\right)^{\overline{\omega}}\right\} ^{4}w_{f}(\mathbf{f}_{t})d\mathbf{f}_{t}\notag\\
&\leq\frac{1}{T^{3}}\left(\left\|g_{i}(\cdot)-\Pi_{\infty K}g_{i}(\cdot)\right\|_{\infty,\overline{\omega}}\right)^{8}\int_{\mathcal{R}_f}\left(1+\left\Vert \mathbf{f}_{t}\right\Vert ^{2}\right)^{4\overline{\omega}}w_{f}(\mathbf{f}_{t})d\mathbf{f}_{t}\notag\\
&=O\left(\frac{K^{-\frac{8\lambda_{i}}{m}}}{T^{3}}\right)O(1)\notag\\
&=O\left(\frac{K^{-\frac{8\lambda_{i}}{m}}}{T^{3}}\right),
\end{align} with $w_f(\*f_t)$ denoting the density of $\*f_t$. The steps above can be applied to every component in (ref). Using exactly the same arguments as for (ref), only the sum with the maximum number of cross-terms will survive. We can thereforee conclude that
\begin{align}
\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\boldsymbol{\alpha}_{g_{i}}-\mathbf{g}_{i}(\mathbf{F})\right)\right\Vert ^{8}\right]&=O\left(K^{-\frac{8\lambda_{i}}{m}}\right),
\end{align} and consequently,
\begin{align}
\mathbb{E}\left(\left\|\*d_{6,2}^{(2)}\right\|^2\right)&\leq O\left(\frac{K^{\frac{13}{4}}}{N}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)^{2}\right)\frac{1}{N}\sum_{i=1}^{N}\left\{ \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\boldsymbol{\alpha}_{g_{i}}-\mathbf{g}_{i}(\mathbf{F})\right)\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}} \\
&=O\left(\frac{K^{\frac{13}{4}}}{N}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)^{2}\right)\left[O\left(K^{-\frac{8\lambda_{i}}{m}}\right)\right]^{\frac{1}{4}}\\
&=O\left(\frac{K^{\frac{13m-8\lambda_i}{4m}}}{N}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)^{2}\right),
\end{align} which implies
\begin{align}
\left\|\*d_{6,2}^{(2)}\right\|=O\left(\frac{K^{\frac{13m-8\lambda_i}{8m}}}{\sqrt{N}}\zeta_{0}(K)^{\frac{3}{4}}\zeta_{1}(K)\right).
\end{align}
The analysis of the first term in (ref), $\*d_{6,2}^{(1)}$, is similar. We can
follow the same steps and assumptions as in (ref) to obtain
\begin{align}
& \mathbb{E} \left(\|\*d_{6,2}^{(1)}\|^2\right) \notag\\
& \equiv \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\+\alpha_{g_i}\right\Vert^2\right] \nonumber \\
&= \mathbb{E}\Bigg[\mathrm{tr}\Bigg(\frac{1}{NT}\sum_{i=1}^N \+\alpha_{g_i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\widetilde{\mathbf{P}} \left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\mathbb{E}\left(\*V_i\mathbf{V}_{i}'\right)\notag\\
& \quad \times \left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\+\alpha_{g_i}\Bigg)\Bigg] \notag \\
&\leq \frac{1}{NT}\sum_{i=1}^N\lambda_{\max}\left(\+\Omega_{v,i}\right)\notag\\
& \quad \times \mathbb{E}\left[\mathrm{tr}\left(\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\+\alpha_{g_i}\+\alpha_{g_i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\widetilde{\mathbf{P}} \left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\right)\right] \notag \\
&\leq \sup_{i=1, \dots, N}\lambda_{\max}\left(\+\Omega_{v,i}\right) \notag \\
& \quad \times \frac{1}{T}\mathbb{E}\left[\mathrm{tr}\left(\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\frac{1}{N}\sum_{i=1}^N\+\alpha_{g_i}\+\alpha_{g_i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\widetilde{\mathbf{P}} \left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\right)\right] \notag \\
&= \sup_{i=1, \dots, N}\lambda_{\max}\left(\+\Omega_{v,i}\right) \notag \\
&\quad \times \frac{1}{T}\mathbb{E} \left[\mathrm{tr}\left(\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\+\Sigma_{\alpha_g}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\widetilde{\mathbf{P}} \left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\right)\right] \notag \\
&\leq \sup_{i=1, \dots, N}\lambda_{\max}\left(\+\Omega_{v,i}\right) \lambda_{\max}\left(\+\Sigma_{\alpha_g}\right) \frac{1}{T}\mathbb{E}\left(\left\|\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\|^2\right) \notag \\
& =O\left(\frac{1}{T}\right)\mathbb{E}\left(\left\|\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\|^2\right),
\end{align}
where $\+\Omega_{v,i}$ and $\+\Sigma_{\alpha_g}$ are defined as in Assumption (ref)(ref) and (ref), respectively.
For the term in expectations in (ref), we use the H\"older and Cauchy-Schwarz inequalities repeatedly to arrive at
\begin{align}
&\mathbb{E}\left[\left\|\frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\|^{2}\right]\notag\\
&\leq\mathbb{E}\left[\left\|\frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\right\|^{2}\left\|\frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\|^{2}\right]\notag\\
&\leq\sqrt{\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\right\Vert ^{4}\right]}\sqrt{\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{4}\right]}\notag\\
&\leq\sqrt{\mathbb{E}\left[\left\|\frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\|^{4}\left\|\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\right\|^{4}\right]}\sqrt{\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{4}\right]}\notag\\
&\leq\left\{ \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}\left\{ \mathbb{E}\left[\left\Vert \left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}\sqrt{\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{4}\right]}\notag\\
&\leq\left\{ \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}\left\{ \mathbb{E}\left[\left\Vert \left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right\Vert ^{8}\right]+\mathbb{E}\left(\left\|\mathbf{I}_{K}\right\|^{8}\right)\right\} ^{\frac{1}{4}}\notag\\
&\quad\times\sqrt{\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{4}\right]}\notag\\
&\leq\left\{ O\left(\frac{K^{4}}{N^{4}}\zeta_{1}(K)^{8}\right)\right\} ^{\frac{1}{4}}\left\{ O(1)+O\left(K^{8}\right)\right\} ^{\frac{1}{4}}\notag\\
&\quad\times\sqrt{\mathbb{E}\left[\left\|\frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}'\right\|^{4}\left\|\frac{\sqrt{T}}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\|^{4}\right]}\notag\\
&\leq O\left(\frac{K}{N}\zeta_{1}(K)^{2}\right)O\left(K^{2}\right)\left\{ \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}'\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}\left\{ \mathbb{E}\left[\left\Vert \frac{\sqrt{T}}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}\notag\\
&=O\left(\frac{K^{3}}{N}\zeta_{1}(K)^{2}\right)\left[O\left(K\zeta_{0}(K)^{6}\right)\right]^{\frac{1}{4}}\left[T^{4}O\left(\frac{K^{4}}{N^{4}}\zeta_{1}(K)^{8}\right)\right]^{\frac{1}{4}}\notag\\
&=O\left(\frac{K^{3}}{N}\zeta_{1}(K)^{2}\right)O\left(K^{\frac{1}{4}}\zeta_{0}(K)^{\frac{3}{2}}\right)O\left(\frac{TK}{N}\zeta_{1}(K)^{2}\right)\notag\\
&=O\left(\frac{TK^{\frac{17}{4}}}{N^{2}}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)^{4}\right),
\end{align} leading to
\begin{align}
\mathbb{E}\left(\left\|\*d_{6,2}^{(1)}\right\|^2\right)&\leq O\left(\frac{1}{T}\right)\mathbb{E}\left(\left\|\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\|^2\right)\notag\\
&=O\left(\frac{K^{\frac{17}{4}}}{N^{2}}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)^{4}\right).
\end{align} Consequently,
\begin{align}
\left\|\*d_{6,2}^{(1)}\right\|=O_p\left(\frac{K^{\frac{17}{8}}}N\zeta_{0}(K)^{\frac{3}{4}}\zeta_{1}(K)^{2}\right).
\end{align}
Insertion into (ref) yields the following bound for $\*d_{6,2}$:
\begin{align}
\left\|\*d_{6,2}\right\| &\equiv \left\Vert \frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right\Vert \nonumber \\
& \leq\left\|\*d_{6,2}^{(1)}\right\|
+\left\|\*d_{6,2}^{(2)}\right\|\nonumber \\
& =O_p\left(\frac{K^{\frac{17}{8}}}N\zeta_{0}(K)^{\frac{3}{4}}\zeta_{1}(K)^{2}\right) + O\left(\frac{K^{\frac{13m-8\lambda_i}{8m}}}{\sqrt{N}}\zeta_{0}(K)^{\frac{3}{4}}\zeta_{1}(K)\right).
\end{align}
The order of $\*d_{6,3}$ is the same as that of $\*d_{6,2}$;
\begin{align}
\left\Vert \*d_{6,3}\right\Vert & \equiv\left\|\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\widetilde{\mathbf{P}}\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right\|\nonumber \\
& =O_p\left(\frac{K^{\frac{17}{8}}}N\zeta_{0}(K)^{\frac{3}{4}}\zeta_{1}(K)^{2}\right) + O\left(\frac{K^{\frac{13m-8\lambda_i}{8m}}}{\sqrt{N}}\zeta_{0}(K)^{\frac{3}{4}}\zeta_{1}(K)\right).
\end{align}
To bound $\*d_{6,1}$, let us again expand using (ref);
\begin{align}
\*d_{6,1} & \equiv \frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right) \nonumber \\
& =\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\+\alpha_{g_i}\nonumber \\
&\quad +\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right) \notag \\
&\equiv \*d_{6,1}^{(1)} + \*d_{6,1}^{(2)}.
\end{align}
We start with $\*d_{6,1}^{(1)}$. By using exactly the same steps as for $\*d_{6,2}^{(1)}$ in (ref), we can show that
\begin{align}
\mathbb{E}\left(\|\*d_{6,1}^{(1)}\|^2\right) &=\mathbb{E}\left[\left\|\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\+\alpha_{g_i}\right\|^2\right] \notag \\
&\leq \sup_{i=1, \dots, N} \lambda_{\max}\left(\+\Omega_{v,i}\right) \lambda_{\max}\left(\+\Sigma_{\alpha_g}\right) \notag \\
&\quad \times \frac{1}{T}\mathbb{E}\left(\left\|\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\|^2\right) \notag \\
&\leq O\left(\frac{1}{T}\right) \mathbb{E}\left(\left\|\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\|^2\right) \notag \\
&= O\left(\frac{1}{T}\right)\mathbb{E}\left(\left\|\frac{1}{T}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\|^{2}\right)\notag \\
&\leq O\left(\frac{1}{T}\right)\mathbb{E}\Bigg(\left\|\frac{\sqrt{T}}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\|^{2}\left[\left\|\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}+\mathbf{I}_{K}\right\|^{2}\right]\notag\\
& \quad \times\left\|\frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\|^2\left\|\frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\|^{2}\Bigg)\notag \\
&\leq O\left(\frac{1}{T}\right)\left\{ \mathbb{E}\left[\left\Vert \frac{\sqrt{T}}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}\left\{ \mathbb{E}\left[\left\Vert \left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right\Vert ^{8}\right]+\mathbb{E}\left[\left\Vert \mathbf{I}_{K}\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}\notag \\
&\quad \times \left\{ \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{8}\right]\right\} ^{\frac{1}{2}}\notag \\
&=O\left(\frac{1}{T}\right)\left[T^{4}O\left(\frac{K^{4}}{N^{4}}\zeta_{1}(K)^{8}\right)\right]^{\frac{1}{4}}\left[O(1)+O\left(K^{8}\right)\right]^{\frac{1}{4}}\left[O\left(\frac{K^{4}}{N^{4}}\zeta_{1}(K)^{8}\right)\right]^{\frac{1}{2}}\notag \\
&=O\left(\frac{1}{T}\right)O\left(\frac{TK}{N}\zeta_{1}(K)^{2}\right)O\left(K^{2}\right)O\left(\frac{K^2}{N^2}\zeta_{1}(K)^4\right)\notag \\
&=O\left(\frac{K^5}{N^3}\zeta_{1}(K)^6\right).
\end{align}
It follows that
\begin{align}
\left\|\*d_{6,1}^{(1)}\right\| &\leq \left\|\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\+\alpha_{g_i}\right\| \notag \\
&= O\left(\frac{K^{\frac{5}{2}}}{N^{\frac{3}{2}}}\zeta_{1}(K)^3\right).
\end{align}
$\*d_{6,1}^{(2)}$ can be analyzed in exactly the same fashion as $\*d_{6,2}^{(2)}$, resulting in
\begin{align}
&\mathbb{E}\left(\left\Vert \mathbf{d}_{6,1}^{(2)}\right\Vert ^{2}\right)\notag\\
&=\mathbb{E}\left[\left\|\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\left(\widetilde{\mathbf{P}}\boldsymbol{\alpha}_{g_{i}}-\mathbf{g}_{i}(\mathbf{F})\right)\right\|^{2}\right]\notag\\
&\leq\sup_{i=1,\dots,N}\lambda_{\max}\left(\boldsymbol{\Omega}_{v,i}\right)\frac{1}{T}\mathbb{E}\left[\left\|\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\left(\widetilde{\mathbf{P}}\boldsymbol{\alpha}_{g_{i}}-\mathbf{g}_{i}(\mathbf{F})\right)\right\|^{2}\right]\notag\\
&=O(1)\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\frac{1}{\sqrt{N}}\sum_{i=1}^{N}\frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\boldsymbol{\alpha}_{g_{i}}-\mathbf{g}_{i}(\mathbf{F})\right)\right\Vert ^{2}\right]\notag\\
&\leq\mathbb{E}\left[\left\|\frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\|^{2}\left\Vert \left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\right\Vert ^{2}\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{2}\frac{1}{N}\sum_{i=1}^{N}\left\Vert \frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\boldsymbol{\alpha}_{g_{i}}-\mathbf{g}_{i}(\mathbf{F})\right)\right\Vert ^{2}\right]\notag\\
&\leq\left\{ \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{8}\right]\right\} ^{\frac{2}{4}}\left\{ \mathbb{E}\left[\left\Vert \left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}} \notag\\
& \quad \times \left\{ \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\boldsymbol{\alpha}_{g_{i}}-\mathbf{g}_{i}(\mathbf{F})\right)\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}\notag\\
&=\sqrt{O\left(\frac{K^{4}}{N^{4}}\zeta_{1}(K)^{8}\right)}\left\{ O\left(K^{8}\right)\right\} ^{\frac{1}{4}}\left\{ O\left(K^{-\frac{8\lambda_{i}}{m}}\right)\right\} ^{\frac{1}{4}}\notag\\
&=O\left(\frac{K^{2}}{N^{2}}\zeta_{1}(K)^{4}\right)O\left(K^{2}\right)O\left(K^{-\frac{2\lambda_{i}}{m}}\right)\notag\\
&=O\left(\frac{K^{\frac{4m-4\lambda_{i}}{2m}}}{N^{2}}\zeta_{1}(K)^{4}\right),
\end{align} with $\lambda_{\max}\left(\+\Omega_{v,i}\right)$ defined as in Assumption (ref)(ref). Consequently,
\begin{align}
\left\|\*d_{6,1}^{(2)}\right\| = O_p\left(\frac{K^{\frac{4m-4\lambda_i}{4m}}}{N}\zeta_1(K)^2\right).
\end{align}
Since the bounds for all terms in (ref) are now known, we
can conclude that
\begin{align}
\left\Vert \*d_{6,1}\right\Vert & \equiv\left\Vert \frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right\Vert \notag \\
& \leq \left\|\*d_{6,1}^{(1)}\right\| + \left\|\*d_{6,1}^{(2)}\right\| \notag \\
&= O\left(\frac{K^{\frac{5}{2}}}{N^{\frac{3}{2}}}\zeta_{1}(K)^3\right) + O_p\left(\frac{K^{\frac{4m-4\lambda_i}{4m}}}{N}\zeta_1(K)^2\right).
\end{align}
$\*d_{6,4}$ can be expanded in the following fashion:
\begin{align}
\left\|\*d_{6,4}\right\| &\equiv \left\|\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right\| \notag \\
&\leq \left\|\frac{1}{\sqrt{NT}}\sum_{i=1}^N \*V_i'\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\widetilde{\mathbf{P}}'(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}})\+\alpha_{g_i}\right\| \notag \\
&\quad +\left\|\frac{1}{\sqrt{NT}}\sum_{i=1}^N \*V_i'\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\widetilde{\mathbf{P}}'(\widetilde{\mathbf{P}}\+\alpha_{g_i} -\*g_i(\*F))\right\| \notag \\
& \equiv \left\|\*d_{6,4}^{(1)}\right\| + \left\|\*d_{6,4}^{(2)}\right\|.
\end{align}
Repeated use of the same steps as before gives
\begin{align}
&\mathbb{E}\left(\left\|\*d_{6,4}^{(1)}\right\|^2\right)\notag\\
& = \mathbb{E} \left[\left\|\frac{1}{\sqrt{NT}}\sum_{i=1}^N \*V_i'\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\widetilde{\mathbf{P}}'(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}})\+\alpha_{g_i}\right\|^2\right] \notag \\
&\leq \sup_{i=1, \dots, N}\lambda_{\max}\left(\+\Omega_{v,i}\right)\lambda_{\max}\left(\+\Sigma_{\alpha_g}\right)\notag\\
& \quad \times \mathbb{E}\left[\left\|\frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\left[\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}'\frac{1}{\sqrt{T}}(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}})\right\|^2\right] \notag \\
& \leq O(1)\mathbb{E}\left[\left\|\frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\right\|^{2}\left\|\left[\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\right\|^{2}\left\|\frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}'\right\|^{2}\left\|\frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\|^{2}\right]\notag \\
&\leq\left\{ \mathbb{E}\left(\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\right\Vert ^{8}\right)\right\} ^{\frac{2}{4}}\left\{ \mathbb{E}\left[\left\Vert \left[\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}\left\{ \left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert ^{8}\right\} ^{\frac{1}{4}}\notag \\
&\leq\sqrt{O\left(K\zeta_{0}(K)^{6}\right)}\left\{ \mathbb{E}\left[\left\|\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right\|^{8}\right]+\mathbb{E}\left[\left\Vert \left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}\notag \\
&\quad\times\left[O\left(\frac{K^{4}}{N^{4}}\zeta_{1}(K)^{8}\right)\right]^{\frac{1}{4}}\notag \\
&=O\left(\sqrt{K}\zeta_{0}(K)^{3}\right)O(1)O\left(\frac{K}{N}\zeta_{1}(K)^{2}\right)\notag \\
&=O\left(\frac{K^{\frac{3}{2}}}{N}\zeta_{0}(K)^{3}\zeta_{1}(K)^{2}\right),
\end{align}
and
\begin{align}
&\mathbb{E}\left(\left\|\*d_{6,4}^{(2)}\right\|^2\right) \notag\\
& = \mathbb{E}\left[\left\|\frac{1}{\sqrt{NT}}\sum_{i=1}^N \*V_i'\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\widetilde{\mathbf{P}}'(\widetilde{\mathbf{P}}\+\alpha_{g_i} -\*g_i(\*F))\right\|^2\right] \notag \\
&\leq \sup_{i=1, \dots, N}\lambda_{\max}\left(\+\Omega_{v,i}\right)\frac{1}{N}\sum_{i=1}^N\mathbb{E}\left[\left\|\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^+ - \left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^+\right]\widetilde{\mathbf{P}}'\frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}} \+\alpha_{g_i} - \*g_i\left(\*F\right)\right)\right\|^2\right] \notag \\
&\leq O(1)\frac{1}{N}\sum_{i=1}^{N}\mathbb{E}\left[\left\|\frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\left[\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}'\frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\boldsymbol{\alpha}_{g_{i}}-\mathbf{g}_{i}(\mathbf{F})\right)\right\|^{2}\right]\notag\\
&\leq\mathbb{E}\left[\left\|\frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\right\|^{2}\left\Vert \left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right\Vert ^{2}\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\right\Vert ^{2}\left\Vert \frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\boldsymbol{\alpha}_{g_{i}}-\mathbf{g}_{i}(\mathbf{F})\right)\right\Vert ^{2}\right]\notag \\
&\leq\left\{ \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}\right\Vert ^{8}\right]\right\} ^{\frac{2}{4}}\left\{ \mathbb{E}\left[\left\Vert \left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}\notag\\
& \quad\times \left\{ \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\boldsymbol{\alpha}_{g_{i}}-\mathbf{g}_{i}(\mathbf{F})\right)\right\Vert ^{8}\right]\right\} ^{\frac{1}{4}}\notag\\
&=\sqrt{O\left(K\zeta_{0}(K)^{6}\right)}O(1)\left[O\left(K^{-\frac{8\lambda_{i}}{m}}\right)\right]^{\frac{1}{4}}\notag\\
&=O\left(\sqrt{K}\zeta_{0}(K)^{3}\right)O\left(K^{-\frac{2\lambda_{i}}{m}}\right)\notag\\
&=O\left(K^{\frac{m-4\lambda_{i}}{2m}}\zeta_{0}(K)^{3}\right).
\end{align} Consequently,
\begin{align}
\left\|\*d_{6,4}^{(1)}\right\| = O_p\left(\frac{K^{\frac{3}{4}}}{\sqrt{N}}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)\right),
\end{align} and
\begin{align}
\left\|\*d_{6,4}^{(2)}\right\| = O_p\left(K^{\frac{m-4\lambda_{i}}{4m}}\zeta_{0}(K)^{\frac{3}{2}}\right),
\end{align} which establishes the following bound for $\*d_{6,4}$:
\begin{align}
\left\|\*d_{6,4}\right\| &\leq \left\|\*d_{6,4}^{(1)}\right\| + \left\|\*d_{6,4}^{(2)}\right\| \notag \\
&= O_p\left(\frac{K^{\frac{3}{4}}}{\sqrt{N}}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)\right) + O_p\left(K^{\frac{m-4\lambda_{i}}{4m}}\zeta_{0}(K)^{\frac{3}{2}}\right).
\end{align}
Therefore,
\begin{align}
\left\|\*D_6\right\| &\equiv \left\|\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'(\mathbf{M}_{\widetilde{P}}-\mathbf{M}_{\widehat{P}})\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right\| \notag \\
&\leq \left\|\*d_{6,1}\right\| + \left\|\*d_{6,2}\right\| + \left\|\*d_{6,3}\right\| + \left\|\*d_{6,4}\right\| \notag \\
&= O\left(\frac{K^{\frac{5}{2}}}{N^{\frac{3}{2}}}\zeta_{1}(K)^3\right) + O_p\left(\frac{K^{\frac{4m-4\lambda_i}{4m}}}{N}\zeta_1(K)^2\right) \notag \\
&\quad +O_p\left(\frac{K^{\frac{17}{8}}}N\zeta_{0}(K)^{\frac{3}{4}}\zeta_{1}(K)^{2}\right) + O\left(\frac{K^{\frac{13m-8\lambda_i}{8m}}}{\sqrt{N}}\zeta_{0}(K)^{\frac{3}{4}}\zeta_{1}(K)\right)\notag \\
&\quad +O_p\left(\frac{K^{\frac{3}{4}}}{\sqrt{N}}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)\right) + O_p\left(K^{\frac{m-4\lambda_{i}}{4m}}\zeta_{0}(K)^{\frac{3}{2}}\right)\notag \\
&=o_p(1),
\end{align} where we have made use of (ref), (ref), (ref), (ref), (ref) and (ref) to establish the last equality.
$\mathbf{D}_{7}$ has the same basic structure as $\mathbf{D}_{6}$;
\begin{align}
\mathbf{D}_{7} & \equiv\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\left(\widehat{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)'\left(\mathbf{M}_{\widetilde{P}}-\mathbf{M}_{\widehat{P}}\right)\boldsymbol{\varepsilon}_{i}\nonumber \\
& =\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\left(\widehat{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\boldsymbol{\varepsilon}_{i}\nonumber \\
& \quad +\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\left(\widehat{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\boldsymbol{\varepsilon}_{i}\nonumber \\
& \quad +\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\left(\widehat{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)'\widetilde{\mathbf{P}}\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\boldsymbol{\varepsilon}_{i}\nonumber \\
& \quad +\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\left(\widehat{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)'\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\widetilde{\mathbf{P}}'\boldsymbol{\varepsilon}_{i}\nonumber \\
& \equiv \*d_{7,1} + \*d_{7,2} + \*d_{7,3} + \*d_{7,4},
\end{align}
where it is straightforward to show using the same arguments as before that
\begin{align}
\left\|\*d_{7,1}\right\| & =O_{p}\left(\left\Vert \*d_{6,1}\right\Vert \right)=O\left(\frac{K^{\frac{5}{2}}}{N^{\frac{3}{2}}}\zeta_{1}(K)^3\right) + O_p\left(\frac{K^{\frac{4m-4\lambda_i}{4m}}}{N}\zeta_1(K)^2\right),\nonumber \\
\left\Vert \*d_{7,2}\right\Vert & =O_{p}\left(\left\Vert \*d_{6,2}\right\Vert \right) =O_p\left(\frac{K^{\frac{17}{8}}}N\zeta_{0}(K)^{\frac{3}{4}}\zeta_{1}(K)^{2}\right) + O\left(\frac{K^{\frac{13m-8\lambda_i}{8m}}}{\sqrt{N}}\zeta_{0}(K)^{\frac{3}{4}}\zeta_{1}(K)\right),\nonumber \\
\left\Vert \*d_{7,3}\right\Vert & =O_{p}\left(\left\Vert \*d_{6,3}\right\Vert \right)= O_p\left(\frac{K^{\frac{17}{8}}}N\zeta_{0}(K)^{\frac{3}{4}}\zeta_{1}(K)^{2}\right) + O\left(\frac{K^{\frac{13m-8\lambda_i}{8m}}}{\sqrt{N}}\zeta_{0}(K)^{\frac{3}{4}}\zeta_{1}(K)\right),\nonumber \\
\left\Vert \*d_{7,4}\right\Vert & =O_{p}\left(\left\Vert \*d_{6,4}\right\Vert \right)= O_p\left(\frac{K^{\frac{3}{4}}}{\sqrt{N}}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)\right) + O_p\left(K^{\frac{m-4\lambda_{i}}{4m}}\zeta_{0}(K)^{\frac{3}{2}}\right),
\end{align} which under our conditions implies
\begin{align}
\left\|\mathbf{D}_{7}\right\| & \leq\left\Vert \*d_{7,1}\right\Vert +\left\Vert \*d_{7,2}\right\Vert +\left\Vert \*d_{7,3}\right\Vert +\left\Vert \*d_{7,4}\right\Vert = o_p(1).
\end{align}
We now continue with $\mathbf{D}_{8}.$ Application of (ref) yields
\begin{align}
-\mathbf{D}_{8} & \equiv\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\left(\widehat{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)'\left(\mathbf{M}_{\widetilde{P}}-\mathbf{M}_{\widehat{P}}\right)\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\nonumber \\
& =\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\left(\widehat{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\nonumber \\
&\quad +\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\left(\widehat{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\nonumber \\
&\quad +\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\left(\widehat{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)'\widetilde{\mathbf{P}}\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\nonumber \\
&\quad +\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\left(\widehat{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)'\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\nonumber \\
& \equiv \*d_{8,1}+\*d_{8,2}+ \*d_{8,3}+\*d_{8,4}.
\end{align}
Consider the second term, $\*d_{8,2}$. The evaluation of this term is similar to that of $\*d_{6,2}$, except that we cannot rely on the same law of large numbers argument to take care of the cross-sectional sum. We instead use the following approach:
\begin{align}
& \left\Vert \*d_{8,2}\right\Vert \notag\\
& \equiv\left\Vert \frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\left(\widehat{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right\Vert \nonumber \\
& \leq\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\left\Vert \left(\widehat{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right\Vert \nonumber \\
& \leq\sqrt{NT}\left\Vert \left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\right\Vert \frac{1}{N}\sum_{i=1}^{N}\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)\right\Vert \left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right\Vert \nonumber \\
& \leq\sqrt{NT}\left\Vert \left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\right\Vert \sqrt{\frac{1}{N}\sum_{i=1}^{N}\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)\right\Vert ^{2}} \notag \\
&\quad \times\sqrt{\frac{1}{N}\sum_{i=1}^{N}\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right\Vert ^{2}}\nonumber \\
& =\sqrt{NT}\left[O_{p}\left(\frac{K^{\frac{3}{4}}}{\sqrt{N}}\sqrt{\zeta_0(K)}\zeta_1(K)\right)+ O_p\left(\frac{K^{\frac{2}{3}}}{\sqrt{N}}\zeta_0(K)^{\frac{2}{3}}\zeta_1(K)\right)\right]\notag \\
&\quad \times \left[O_p\left(\frac{\zeta_{1}(K)^{2}}{N}\right) + O_p\left(K^{-\frac{2\lambda_i}{m}}\right)\right] \notag \\
&= \sqrt{NT}o_p(1)\left[O_p\left(\frac{\zeta_{1}(K)^{2}}{N}\right) + O_p\left(K^{-\frac{2\lambda_i}{m}}\right)\right]\notag \\
&= o_p\left(\frac{\sqrt{T}\zeta_1(K)^2}{\sqrt{N}}\right)+o_p\left(\sqrt{NT}K^{-\frac{2\lambda_i}{m}}\right),
\end{align} where we have used the results in (ref) and (ref), as well as the Cauchy-Schwarz inequality followed by
\begin{align}
&\frac{1}{N}\sum_{i=1}^{N}\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right\Vert ^{2} \notag \\
& \leq\frac{2}{N}\sum_{i=1}^{N}\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\+\alpha_{g_i}\right\Vert ^{2}+\frac{2}{N}\sum_{i=1}^{N}\left\Vert \frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right\Vert ^{2}\nonumber \\
&= O_p\left(\frac{\zeta_1(K)^2}{N}\right) + O_p\left(K^{-\frac{2\lambda_i}{m}}\right).
\end{align} This last result holds because by Assumption (ref) and Lemma (ref)(ref) we have
\begin{align}
\frac{1}{N}\sum_{i=1}^N\left\|\frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\+\alpha_{g_i}\right\|^2 &= \mathrm{tr}\left[\frac{1}{T}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\frac{1}{N}\sum_{i=1}^N\+\alpha_{g_i}\+\alpha_{g_i}'\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\right] \notag \\
&= \mathrm{tr}\left[\frac{1}{T}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\+\Sigma_{\alpha_g}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\right] \notag \\
&\leq \lambda_{\max}\left(\+\Sigma_{\alpha_g}\right) \mathrm{tr}\left[\frac{1}{T}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\right] \notag \\
&=\lambda_{\max}\left(\+\Sigma_{\alpha_g}\right)\left\|\frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\right\|^2 \notag \\
&= O_p(1) O_p\left(\frac{\zeta_1(K)^2}{N}\right) \notag \\
&= O_p\left(\frac{\zeta_1(K)^2}{N}\right).
\end{align}
Exactly the same arguments can be used to show that $\frac{1}{N}\sum_{i=1}^{N}\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)\right\Vert ^{2}$
is of the same order as $\frac{1}{N}\sum_{i=1}^{N}\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right\Vert ^{2}$.
$\*d_{8,3}$ is of the same order as $\*d_{8,2}$; that is,
\begin{align}
\left\|\*d_{8,3}\right\|=o_p\left(\frac{\sqrt{T}\zeta_1(K)^2}{\sqrt{N}}\right)+o_p\left(\sqrt{NT}K^{-\frac{2\lambda_i}{m}}\right).
\end{align}
As for $\*d_{8,1}$, by the usual arguments together with $\left\|\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+} \left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)' \right\| = O_p\left(\frac{K}{N}\zeta_1(K)^2\right)$ from (ref), we can show that
\begin{align}
\left\Vert \*d_{8,1}\right\Vert & =\sqrt{NT}O_{p}\left(\frac{K}{N}\zeta_1(K)^2\right) \left[O_{p}\left(\frac{\zeta_{1}(K)^{2}}{N}\right) + O_p\left(K^{-\frac{2\lambda_i}{m}}\right)\right]\nonumber \\
& = O_{p}\left(\sqrt{NT}\frac{K}{N^{2}}\zeta_{1}(K)^{4}\right)+O_{p}\left(\sqrt{NT}\frac{K^{\frac{m-2\lambda_{i}}{m}}}{N}\zeta_{1}(K)^{2}\right)\notag \\
&=O_{p}\left(\frac{\sqrt{T}K}{N^{\frac{3}{2}}}\zeta_{1}(K)^{4}\right)+O_{P}\left(\frac{\sqrt{T}K^{\frac{m-2\lambda_{i}}{m}}}{\sqrt{N}}\zeta_{1}(K)^{2}\right).
\end{align}
Similarly, by using the known order of $\left\|\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\widetilde{\mathbf{P}}'\right\| = O_p\left(K^{\frac{1}{4}}\zeta_0(K)^{\frac{3}{2}}\right)$ from (ref), we obtain
\begin{align}
\left\|\*d_{8,4}\right\| & =\sqrt{NT}O_p\left(K^{\frac{1}{4}}\zeta_0(K)^{\frac{3}{2}}\right)\left[O_{p}\left(\frac{\zeta_{1}(K)^{2}}{N}\right) + O_p\left(K^{-\frac{2\lambda_i}{m}}\right)\right] \notag \\
&= O_{p}\left(\sqrt{NT}\frac{K^{\frac{1}{4}}}{N}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)^{2}\right)+O_{P}\left(\sqrt{NT}K^{\frac{m-8\lambda_{i}}{4m}}\zeta_{0}(K)^{\frac{3}{2}}\right)\notag\\
&=O_{p}\left(\frac{\sqrt{T}K^{\frac{1}{4}}}{\sqrt{N}}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)^{2}\right)+O_{P}\left(\sqrt{NT}K^{\frac{m-8\lambda_{i}}{4m}}\zeta_{0}(K)^{\frac{3}{2}}\right).
\end{align}
This allows us to arrive at the order of $\mathbf{D}_{8}$;
\begin{align}
\left\|\mathbf{D}_{8}\right\| & \leq\left\|\*d_{8,1}\right\|+\left\Vert \*d_{8,2}\right\Vert +\left\Vert \*d_{8,3}\right\Vert +\left\Vert \*d_{8,4}\right\Vert \nonumber \\
& = O_{p}\left(\frac{\sqrt{T}K}{N^{\frac{3}{2}}}\zeta_{1}(K)^{4}\right)+O_{P}\left(\frac{\sqrt{T}K^{\frac{m-2\lambda_{i}}{m}}}{\sqrt{N}}\zeta_{1}(K)^{2}\right) \notag \\
&\quad + o_p\left(\frac{\sqrt{T}\zeta_1(K)^2}{\sqrt{N}}\right)+o_p\left(\sqrt{NT}K^{-\frac{2\lambda_i}{m}}\right) \notag \\
& \quad +O_{p}\left(\frac{\sqrt{T}K^{\frac{1}{4}}}{\sqrt{N}}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)^{2}\right)+O_{P}\left(\sqrt{NT}K^{\frac{m-8\lambda_{i}}{4m}}\zeta_{0}(K)^{\frac{3}{2}}\right) \notag \\
&=o_p(1),
\end{align} where the last equality follows from (ref), (ref), (ref), (ref), and conditions $\sqrt{NT}K^{-\frac{\lambda_i}{m}}=o(1)$ and $\frac{\sqrt{T}K^{\frac{1}{4}}}{\sqrt{N}}\zeta_{0}(K)^{\frac{3}{2}}\zeta_{1}(K)^{2}=o(1)$ in Assumption (ref)(ref).
Let us now consider $\mathbf{D}_{2}$, which we can expand in the following fashion using (ref):
\begin{align}
-\mathbf{D}_{2} & \equiv\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\mathbf{M}_{\widetilde{P}}\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\nonumber \\
& =\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\mathbf{M}_{\widetilde{P}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\+\alpha_{g_i}+\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\mathbf{M}_{\widetilde{P}}\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\nonumber \\
& \equiv \*d_{2,1}+\*d_{2,2}.
\end{align}
We start with the second term on the right, $\*d_{2,2}$.
Note that $\mathbf{M}_{\widetilde{P}}$ is an idempotent matrix
whose eigenvalues are either zero or one (see abadir2005,
Exercise 8.56). Therefore, $\lambda_{\max}(\mathbf{M}_{\widetilde{P}})=1$.
Then, using Assumptions (ref)(ref), (ref), and the definitions of the variances in Assumption (ref)(ref),
\begin{align}
&\mathbb{E}\left[\left\|\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\mathbf{M}_{\widetilde{P}}\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right\|^{2}\right] \notag \\
& =\mathbb{E}\left[\frac{1}{NT}\sum_{i=1}^{N}\sum_{j=1}^{N}\mathrm{tr}\left(\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)'\mathbf{M}_{\widetilde{P}}\mathbb{E}\left(\mathbf{V}_{i}\mathbf{V}_{j}'\right)\mathbf{M}_{\widetilde{P}}\left(\widetilde{\mathbf{P}}\+\alpha_{g_j}-\*g_j(\mathbf{F})\right)\right)\right]\nonumber \\
& =\frac{1}{N}\sum_{i=1}^{N}\frac{1}{T}\mathbb{E}\left[\mathrm{tr}\left(\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)'\mathbf{M}_{\widetilde{P}}\boldsymbol{\Omega}_{v,i}\mathbf{M}_{\widetilde{P}}\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right)\right]\nonumber \\
& \leq\frac{1}{N}\sum_{i=1}^{N}\lambda_{\max}\left(\boldsymbol{\Omega}_{v,i}\right)\frac{1}{T}\mathbb{E}\left[\mathrm{tr}\left(\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)'\mathbf{M}_{\widetilde{P}}\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right)\right]\nonumber \\
& \leq\frac{1}{N}\sum_{i=1}^{N}\lambda_{\max}\left(\boldsymbol{\Omega}_{v,i}\right)\frac{1}{T}\mathbb{E}\left[\lambda_{\max}(\mathbf{M}_{\widetilde{P}})\mathrm{tr}\left(\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)'\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right)\right]\nonumber \\
& =\frac{1}{N}\sum_{i=1}^{N}\lambda_{\max}\left(\boldsymbol{\Omega}_{v,i}\right)\frac{1}{T}\mathbb{E}\left[\mathrm{tr}\left(\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)'\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right)\right]\nonumber \\
& \leq\frac{1}{N}\sum_{i=1}^{N}\lambda_{\max}\left(\boldsymbol{\Omega}_{v,i}\right)\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right\Vert ^{2}\right]\nonumber \\
& \leq\sup_{i=1,\dots,N}\lambda_{\max}\left(\boldsymbol{\Omega}_{v,i}\right)\frac{1}{N}\sum_{i=1}^{N}\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right\Vert ^{2}\right]\nonumber \\
& =O(1)O\left(K^{-\frac{2\lambda_i}{m}}\right)\nonumber \\
& =O\left(K^{-\frac{2\lambda_i}{m}}\right),
\end{align}
implying
\begin{equation}
\left\Vert \*d_{2,2}\right\Vert =O_{p}\left(K^{-\frac{\lambda_i}{m}}\right).
\end{equation}
For $\*d_{2,1}$, we use
\begin{align}
\mathbb{E} \left(\|\*d_{2,1}\|^2\right) & = \mathbb{E}\left[\left\Vert \frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\mathbf{M}_{\widetilde{P}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\+\alpha_{g_i}\right\Vert ^{2}\right] \notag \\
& =\mathbb{E}\left[\frac{1}{NT}\sum_{i=1}^{N}\sum_{j=1}^{N}\mathrm{tr}\left(\+\alpha_{g_i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\mathbf{M}_{\widetilde{P}}\mathbb{E}\left(\mathbf{V}_{i}\mathbf{V}_{j}'\right)\mathbf{M}_{\widetilde{P}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\+\alpha_{g_j}\right)\right]\nonumber \\
& =\frac{1}{N}\sum_{i=1}^{N}\frac{1}{T}\mathbb{E}\left[\mathrm{tr}\left(\+\alpha_{g_i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\mathbf{M}_{\widetilde{P}}\boldsymbol{\Omega}_{v,i}\mathbf{M}_{\widetilde{P}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\+\alpha_{g_i}\right)\right]\nonumber \\
& \leq \sup_{i=1, \dots, N} \lambda_{\max} \left(\+\Omega_{v,i}\right)\mathbb{E} \left[\mathrm{tr}\left(\frac{1}{T} \left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right) \frac{1}{N}\sum_{i=1}^N\+\alpha_{g_i}\+\alpha_{g_i}'\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\right)\right] \notag \\
&= \sup_{i=1, \dots, N} \lambda_{\max}\left(\+\Omega_{v,i}\right)\mathbb{E} \left[\mathrm{tr}\left(\frac{1}{T} \left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right) \+\Sigma_{\alpha_g}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\right)\right]\notag \\
&\leq \sup_{i=1, \dots, N}\lambda_{\max}\left(\+\Omega_{v,i}\right)\lambda_{\max}\left(\+\Sigma_{\alpha_g}\right) \mathbb{E} \left[\left\|\frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\right\|^2\right]\notag \\
&=O(1)O(1)O\left(\frac{\zeta_1(K)^2}{N}\right)=O\left(\frac{\zeta_1(K)^2}{N}\right).
\end{align}
Hence,
\begin{equation}
\left\|\*d_{2,1}\right\|=O_{p}\left(\frac{\zeta_{1}(K)}{\sqrt{N}}\right).
\end{equation}
Insertion into (ref) leads to the following:
\begin{align}
\left\|\mathbf{D}_{2}\right\| & \leq\left\|\*d_{2,1}\right\|+\left\Vert \*d_{2,2}\right\Vert \nonumber \\
& =O_{p}\left(\frac{\zeta_{1}(K)}{\sqrt{N}}\right)+O_{p}\left(K^{-\frac{\lambda_i}{m}}\right)\nonumber \\
& = o_p(1),
\end{align} which holds by Assumption (ref)(ref) and the result in (ref).
$\mathbf{D}_{3}$ has the same basic structure as $\mathbf{D}_{2}$. Note in particular how
\begin{align}
-\mathbf{D}_{3} & \equiv\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\left(\widehat{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)'\mathbf{M}_{\widetilde{P}}\boldsymbol{\varepsilon}_{i}\nonumber \\
& =\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\+\alpha_{G_i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\mathbf{M}_{\widetilde{P}}\boldsymbol{\varepsilon}_{i}+\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\left(\widetilde{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)'\mathbf{M}_{\widetilde{P}}\boldsymbol{\varepsilon}_{i}\nonumber \\
& \equiv \*d_{3,1} + \*d_{3,2},
\end{align}
where under our assumptions it is straightforward to show that
\begin{align}
\left\|\*d_{3,1}\right\| & =O_{p}\left(\left\Vert \*d_{2,1}\right\Vert \right)=O_{p}\left(\frac{\zeta_{1}(K)}{\sqrt{N}}\right),\\
\left\Vert \*d_{3,2}\right\Vert & =O_{p}\left(\left\Vert \*d_{2,2}\right\Vert \right)=O_{p}\left(K^{-\frac{\lambda_i}{m}}\right),
\end{align}
and thus,
\begin{align}
\left\|\mathbf{D}_{3}\right\| & \leq\left\Vert \*d_{3,1}\right\Vert +\left\Vert \*d_{3,2}\right\Vert \nonumber \\
& =O_{p}\left(\frac{\zeta_{1}(K)}{\sqrt{N}}\right)+O_{p}\left(K^{-\frac{\lambda_i}{m}}\right)\nonumber \\
& = o_p(1),
\end{align} where the last equality holds similarly to the one in $\left\|\*D_2\right\|$.
We continue with $\mathbf{D}_{4}$;
\begin{align}
\mathbf{D}_{4} & \equiv\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\left(\widehat{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)'\mathbf{M}_{\widetilde{P}}\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\nonumber \\
& =\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\left[\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\+\alpha_{G_i}\right]'\mathbf{M}_{\widetilde{P}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\+\alpha_{g_i}\notag \\
& \quad +\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\left[\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\+\alpha_{G_i}\right]'\mathbf{M}_{\widetilde{P}}\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\nonumber \\
& \quad +\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\left(\widetilde{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)'\mathbf{M}_{\widetilde{P}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\+\alpha_{g_i} \notag \\
& \quad +\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\left(\widetilde{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)'\mathbf{M}_{\widetilde{P}}\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\nonumber \\
& \equiv \*d_{4,1}+\*d_{4,2}+\*d_{4,3}+\*d_{4,4}.
\end{align} For the first term on the right, $\*d_{4,1}$, we can use the Cauchy-Schwarz inequality to obtain
\begin{align}
\left\|\*d_{4,1}\right\|&=\left\|\frac{1}{\sqrt{NT}}\sum_{i=1}^N\+\alpha_{G_i}' \left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\*M_{\widetilde{P}}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\+\alpha_{g_i}\right\| \notag \\
& \leq \frac{1}{\sqrt{NT}} \sum_{i=1}^N \left\|\+\alpha_{G_i}' \left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\*M_{\widetilde{P}}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\+\alpha_{g_i}\right\| \notag \\
& \leq \sqrt{NT} \left\|\frac{1}{T} \left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\*M_{\widetilde{P}}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\right\| \frac{1}{NK}\sum_{i=1}^N \left\|\+\alpha_{G_i}\right\|\left\|\+\alpha_{g_i}\right\| \notag \\
& \leq \sqrt{NT} \left\|\frac{1}{T} \left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\*M_{\widetilde{P}}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\right\| \sqrt{\frac{1}{NK}\sum_{i=1}^N\left\|\+\alpha_{G_i}\right\|^2}\sqrt{\frac{1}{NK}\sum_{i=1}^N\left\|\+\alpha_{g_i}\right\|^2} \notag \\
& = \sqrt{NT} K O_p\left(\frac{\zeta_1(K)^2}{N}\right)O_p(1) \notag \\
&= O_p\left(\sqrt{\frac{T}{N}} K \zeta_1(K)^2\right),
\end{align} which holds under Assumption (ref) and the fact that the trace of a square matrix is the sum of its eigenvalues (abadir2005, Exercise 7.27). Note how
\begin{align}
\frac{1}{NK}\sum_{i=1}^N\left\|\+\alpha_{g_i}\right\|^2 &= \frac{1}{NK}\sum_{i=1}^N \mathrm{tr}\left(\+\alpha_{g_i}\+\alpha_{g_i}'\right) = \mathrm{tr}\left(\frac{1}{NK}\sum_{i=1}^N \+\alpha_{g_i}\+\alpha_{g_i}'\right) \notag \\
&= \frac{1}{K}\mathrm{tr}\left(\+\Sigma_{\alpha_g}\right) = \frac{1}{K}\sum_{k=1}^K \lambda_k\left(\+\Sigma_{\alpha_g}\right) \leq \frac{1}{K} K \lambda_{\max}\left(\+\Sigma_{\alpha_g}\right)= O(1),
\end{align} and we can similarly show that $\frac{1}{NK}\sum_{i=1}^N \left\|\+\alpha_{G_i}\right\|^2 = O(1)$. Further,
\begin{align}
& \left\|\frac{1}{T}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\*M_{\widetilde{P}}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\right\| \notag\\
&= \sqrt{\mathrm{tr}\left[\frac{1}{T}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\*M_{\widetilde{P}}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\frac{1}{T}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\*M_{\widetilde{P}}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\right]} \notag \\
&\leq \sqrt{\mathrm{tr}\left[\frac{1}{T}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\*M_{\widetilde{P}}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\right]\mathrm{tr}\left[\frac{1}{T}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\*M_{\widetilde{P}}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\right]}\notag \\
&\leq \lambda_{\max}\left(\*M_{\widetilde{P}}\right)\mathrm{tr}\left[\frac{1}{T}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\right]\notag \\
&=\lambda_{\max}\left(\*M_{\widetilde{P}}\right)\left\|\frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\right\|^2 \notag \\
&= O_p(1) O_p\left(\frac{\zeta_1(K)^2}{N}\right) = O_p\left(\frac{\zeta_1(K)^2}{N}\right),
\end{align} with which we can bound $\*d_{4,1}$.
We can apply exactly the same steps to
\begin{align}
\*d_{4,2}\equiv\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\left[\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\+\alpha_{G_i}\right]'\mathbf{M}_{\widetilde{P}}\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right).
\end{align} The only difference here is that instead of a second $O_{p}\left(\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\right\Vert \right)$,
we need to bound $O_{p}\left(\left\|\frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right\|\right)$ instead. From (ref) we know that this
term is of order $O_{p}\left(K^{-\frac{\lambda_i}{m}}\right)$. Therefore, using the same argumentation as before,
\begin{align}
\left\|\*d_{4,2}\right\| & \equiv\left\|\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\+\alpha_{G_i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\mathbf{M}_{\widetilde{P}}\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right\|\nonumber \\
& \leq \frac{1}{\sqrt{NT}}\sum_{i=1}^N \left\|\+\alpha_{G_i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\mathbf{M}_{\widetilde{P}}\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right\| \notag \\
&\leq \sqrt{NTK}\left\|\frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\*M_{\widetilde{P}}\right\| \frac{1}{N\sqrt{K}}\sum_{i=1}^N \left\|\+\alpha_{G_i}\right\| \left\|\frac{1}{\sqrt{T}} \left(\widetilde{\mathbf{P}}\+\alpha_{g_i} - \*g_i\left(\*F\right)\right)\right\| \notag \\
&\leq \sqrt{NTK}\left\|\frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\*M_{\widetilde{P}}\right\|\sqrt{\frac{1}{NK}\sum_{i=1}^N \left\|\+\alpha_{G_i}\right\|^2}\notag\\
& \quad \times\sqrt{\frac{1}{N}\sum_{i=1}^N \left\|\frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\+\alpha_{g_i} - \*g_i\left(\*F\right)\right)\right\|^2},
\end{align} where
\begin{align}
\left\|\frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\*M_{\widetilde{P}}\right\| &= \sqrt{\mathrm{tr}\left[\frac{1}{T}\*M_{\widetilde{P}}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\*M_{\widetilde{P}}\right]} \notag \\
&= \sqrt{\mathrm{tr}\left[\frac{1}{T}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)' \*M_{\widetilde{P}}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\right]} \notag \\
&\leq \sqrt{\lambda_{\max}\left(\*M_{\widetilde{P}}\right)\mathrm{tr}\left[\frac{1}{T}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\right]} \notag \\
&= \sqrt{\lambda_{\max}\left(\*M_{\widetilde{P}}\right)\left\|\frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)\right\|^2} \notag \\
&= O_p(1) O_p\left(\frac{\zeta_1(K)}{\sqrt{N}}\right) = O_p\left(\frac{\zeta_1(K)}{\sqrt{N}}\right),
\end{align} and therefore,
\begin{align}
\left\|\*d_{4,2}\right\| &\leq \sqrt{NTK}\left\|\frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}} - \widetilde{\mathbf{P}}\right)'\*M_{\widetilde{P}}\right\|\sqrt{\frac{1}{NK}\sum_{i=1}^N \left\|\+\alpha_{G_i}\right\|^2}\sqrt{\frac{1}{N}\sum_{i=1}^N \left\|\frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\+\alpha_{g_i} - \*g_i\left(\*F\right)\right)\right\|^2} \notag \\
&= \sqrt{NTK}O_p\left(\frac{\zeta_1(K)}{\sqrt{N}}\right)O_p(1)O_p\left(K^{-\frac{\lambda_i}{m}}\right) \notag \\
&= O_p\left(\sqrt{T}K^{\frac{1}{2}-\frac{\lambda_i}{m}}\zeta_1(K)\right).
\end{align}
The third term is obviously of the same order as the second and hence,
\begin{equation}
\left\|\*d_{4,3}\right\|\equiv\left\|\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\left(\widetilde{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)'\mathbf{M}_{\widetilde{P}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\+\alpha_{g_i}\right\|=O_p\left(\sqrt{T}K^{\frac{1}{2}-\frac{\lambda_i}{m}}\zeta_1(K)\right).
\end{equation}
It remains to consider $\*d_{4,4}$. Note how
\begin{align}
\left\|\*d_{4,4}\right\|^2 &= \left\|\frac{\sqrt{T}}{\sqrt{N}T}\sum_{i=1}^{N}\left(\widetilde{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)'\mathbf{M}_{\widetilde{P}}\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right\|^2 \notag \\
&=\mathrm{tr}\Bigg[\frac{T}{NT^2}\sum_{i=1}^N\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)'\*M_{\widetilde{P}}\left(\widetilde{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)\notag\\
& \quad \times\left(\widetilde{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)'\mathbf{M}_{\widetilde{P}}\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\Bigg] \notag \\
&\leq \frac{T}{N} \sum_{i=1}^N \mathrm{tr}\left[\frac{1}{T}\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\left(\widetilde{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)'\*M_{\widetilde{P}}\right] \notag \\
&\quad \times \mathrm{tr}\left[\frac{1}{T}\left(\widetilde{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)\left(\widetilde{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)'\*M_{\widetilde{P}}\right] \notag \\
&\leq \frac{T}{N}\sum_{i=1}^N \lambda_{\max}\left(\*M_{\widetilde{P}}\right)^2\left\|\frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\+\alpha_{g_i} - \*g_i\left(\*F\right)\right)\right\|^2\left\|\frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\+\alpha_{G_i} - \*G_i\left(\*F\right)\right)\right\|^2 \notag \\
&\leq T \frac{1}{N}\sum_{i=1}^N\sqrt{\left\|\frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\+\alpha_{g_i} - \*g_i\left(\*F\right)\right)\right\|^4} \sqrt{\left\|\frac{1}{\sqrt{T}}\left(\widetilde{\mathbf{P}}\+\alpha_{G_i} - \*G_i\left(\*F\right)\right)\right\|^4} \notag \\
&= T O_p\left(K^{-\frac{4\lambda_i}{m}}\right) \notag \\
&= O_p\left(T K^{-\frac{4\lambda_i}{m}}\right).
\end{align}
By adding the above terms, we can show that
\begin{align}
\left\|\mathbf{D}_{4}\right\| & \leq\left\Vert \*d_{4,1}\right\Vert +\left\Vert \*d_{4,2}\right\Vert +\left\Vert \*d_{4,3}\right\Vert +\left\Vert \*d_{4,4}\right\Vert \nonumber \\
& =O_p\left(\sqrt{\frac{T}{N}}K\zeta_1(K)^2\right) + O_p\left(\sqrt{T}K^{\frac{1}{2}-\frac{\lambda_i}{m}}\zeta_1(K)\right) + O_p\left(T K^{-\frac{4\lambda_i}{m}}\right) \notag \\
&=o_p(1),
\end{align} where we have used (ref), (ref) and (ref) to establish the last equality.
The only term that remains now is $\mathbf{D}_{1}$, which we write as
\begin{align}
\mathbf{D}_{1} & \equiv\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\mathbf{M}_{\widetilde{P}}\boldsymbol{\varepsilon}_{i}\nonumber \\
& =\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\left[\mathbf{I}_{T}-\widetilde{\mathbf{P}}\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\right]\boldsymbol{\varepsilon}_{i}\nonumber \\
& =\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\boldsymbol{\varepsilon}_{i}-\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\widetilde{\mathbf{P}}\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\boldsymbol{\varepsilon}_{i}\nonumber \\
& \equiv \*d_{1,1}+\*d_{1,2},
\end{align}
where we have used the definition of the projection matrix $\mathbf{M}_{\widetilde{P}}\equiv\mathbf{I}_{T}-\widetilde{\mathbf{P}}\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'$.
Consider the second term on the right-hand side, $\*d_{1,2}$. Here, we proceed as follows:
\begin{align}
\mathbb{E}\left(\|\*d_{1,2}\|^2\right) &= \mathbb{E}\left(\left\Vert \frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\widetilde{\mathbf{P}}\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\boldsymbol{\varepsilon}_{i}\right\Vert ^{2}\right)\nonumber \\
& =\mathbb{E}\left[\frac{1}{NT}\sum_{i=1}^{N}\sum_{j=1}^{N}\mathrm{tr}\left(\boldsymbol{\varepsilon}_{i}'\widetilde{\mathbf{P}}\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\mathbf{V}_{i}\mathbf{V}_{j}'\widetilde{\mathbf{P}}\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\boldsymbol{\varepsilon}_{j}\right)\right]\nonumber \\
& =\mathbb{E}\left[\frac{1}{NT}\sum_{i=1}^{N}\mathrm{tr}\left[\widetilde{\mathbf{P}}\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\mathbb{E}\left(\mathbf{V}_{i}\mathbf{V}_{i}'\right)\widetilde{\mathbf{P}}\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\mathbb{E}\left(\boldsymbol{\varepsilon}_{i}\boldsymbol{\varepsilon}_{i}'\right)\right]\right]\nonumber \\
& \leq\mathbb{E}\left[\frac{1}{NT}\sum_{i=1}^{N}\lambda_{\max}\left(\boldsymbol{\Omega}_{v,i}\right)\mathrm{tr}\left(\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\mathbb{E}\left(\boldsymbol{\varepsilon}_{i}'\boldsymbol{\varepsilon}_{i}\right)\right)\right]\nonumber \\
& \leq\mathbb{E}\left[\frac{1}{NT}\sum_{i=1}^{N}\lambda_{\max}\left(\boldsymbol{\Omega}_{v,i}\right)\mathrm{tr}\left(\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\mathbb{E}\left(\boldsymbol{\varepsilon}_{i}'\boldsymbol{\varepsilon}_{i}\right)\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)\right]\nonumber \\
& \leq\mathbb{E}\left[\frac{1}{NT}\sum_{i=1}^{N}\lambda_{\max}\left(\boldsymbol{\Omega}_{v,i}\right)\lambda_{\max}\left(\boldsymbol{\Omega_{\varepsilon,i}}\right)\mathrm{tr}\left(\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)\right]\nonumber \\
& =\frac{1}{NT}\sum_{i=1}^{N}\lambda_{\max}\left(\boldsymbol{\Omega}_{v,i}\right)\lambda_{\max}\left(\boldsymbol{\Omega}_{\varepsilon,i}\right)\mathrm{tr}\left(\mathbf{I}_{K}\right)\nonumber \\
& =\frac{1}{N}\sum_{i=1}^{N}\lambda_{\max}\left(\boldsymbol{\Omega}_{v,i}\right)\lambda_{\max}\left(\boldsymbol{\Omega}_{\varepsilon,i}\right)\frac{K}{T}\nonumber \\
& =O(1)\frac{K}{T}\nonumber \\
& =O\left(\frac{K}{T}\right).
\end{align}
This allows us to conclude that
\begin{equation}
\left\|\*d_{1,2}\right\|=O_{p}\left(\sqrt{\frac{K}{T}}\right),
\end{equation}
which in turn implies the following:
\begin{align}
\mathbf{D}_{1} & \equiv \*d_{1,1}+\*d_{1,2}\nonumber \\
& \equiv\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\boldsymbol{\varepsilon}_{i}+O_{p}\left(\sqrt{\frac{K}{T}}\right),
\end{align} where the reminder is negligible provided that $\frac{K}{T}=o(1)$, which holds because $\sqrt{\frac{K}{T}}\zeta_0(K)^{\frac{3}{2}}=o(1)$ by Assumption (ref)(ref).
We have shown that $\*D_2 -\*D_8$ are negligible. Therefore, (ref) reduces to
\begin{align}
\frac{1}{\sqrt{NT}}\sum_{i=1}^N\*X_i'\*M_{\widehat{P}}\left[\+\varepsilon_i - \left(\widehat{\mathbf{P}}\+\alpha_{g_i} - \*g_i\left(\*F\right)\right) \+\gamma_i\right] \equiv \sum_{j=1}^8 \*D_j \equiv \frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\boldsymbol{\varepsilon}_{i} + o_p(1).
\end{align} This is the sought cleaned-up asymptotic representation of the numerator of $\sqrt{NT}(\widehat{\+\beta}_{SCCE} - \+\beta)$.
Let us now consider the denominator, which we expand in the usual
fashion:
\begin{align}
\frac{1}{NT}\sum_{i=1}^{N} & \mathbf{X}_{i}'\mathbf{M}_{\widehat{P}}\mathbf{X}_{i}=\frac{1}{NT}\sum_{i=1}^{N}\left[\mathbf{V}_{i}-\left(\widehat{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)\right]'\mathbf{M}_{\widehat{P}}\left[\mathbf{V}_{i}-\left(\widehat{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)\right]\nonumber \\
& =\frac{1}{NT}\sum_{i=1}^{N}\left[\mathbf{V}_{i}-\left(\widehat{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)\right]'\mathbf{M}_{\widetilde{P}}\left[\mathbf{V}_{i}-\left(\widehat{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)\right]\nonumber \\
& \quad -\frac{1}{NT}\sum_{i=1}^{N}\left[\mathbf{V}_{i}-\left(\widehat{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)\right]'\left(\mathbf{M}_{\widetilde{P}}-\mathbf{M}_{\widehat{P}}\right)\left[\mathbf{V}_{i}-\left(\widehat{\mathbf{P}}\+\alpha_{G_i}-\*G_i(\mathbf{F})\right)\right].
\end{align}
Most of the terms in (ref) are either known from before or similar to those considered earlier, and are negligible under our assumptions. The only new terms are
\begin{align}
\*b_1 &\equiv \frac{1}{NT}\sum_{i=1}^{N}\mathbf{V}_{i}'\mathbf{M}_{\widetilde{P}}\mathbf{V}_{i}, \\
\*b_2 &\equiv \frac{1}{NT}\sum_{i=1}^{N}\mathbf{V}_{i}'\left(\mathbf{M}_{\widetilde{P}}-\mathbf{M}_{\widehat{P}}\right)\mathbf{V}_{i}.
\end{align}
Consider the second term, $\*b_2$. From (ref),
\begin{align}
\*b_2 &\equiv \frac{1}{NT}\sum_{i=1}^N\mathbf{V}_{i}'\left(\mathbf{M}_{\widetilde{P}}-\mathbf{M}_{\widehat{P}}\right)\mathbf{V}_{i}\nonumber \\
& =\frac{1}{NT}\sum_{i=1}^N\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\mathbf{V}_{i} +\frac{1}{NT}\sum_{i=1}^N\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\mathbf{V}_{i}\nonumber \\
& \quad +\frac{1}{NT}\sum_{i=1}^N\mathbf{V}_{i}'\widetilde{\mathbf{P}}\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\mathbf{V}_{i} +\frac{1}{NT}\sum_{i=1}^N\mathbf{V}_{i}'\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\widetilde{\mathbf{P}}'\mathbf{V}_{i} \notag \\
&\equiv \*b_{2,1} + \*b_{2,2} + \*b_{2,3} + \*b_{2,4}.
\end{align} We start with the first term, $\*b_{2,1} \equiv \frac{1}{NT}\sum_{i=1}^N\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\mathbf{V}_{i}$, where
\begin{align}
\mathbb{E}\left(\left\|\*b_{2,1}\right\|^2\right) &=\mathbb{E}\left[\left\|\frac{1}{NT}\sum_{i=1}^{N}\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\mathbf{V}_{i}\right\|^{2}\right]\notag\\
&=\mathbb{E}\Bigg[\frac{1}{N^{2}T^{2}}\sum_{i=1}^{N}\sum_{j=1}^{N}\mathrm{tr}\Big(\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\mathbb{E}\left(\mathbf{V}_{i}\mathbf{V}_{j}'\right) \notag\\
& \quad \times \left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\mathbf{V}_{j}\Big)\Bigg]\notag\\
&\leq\frac{1}{N^{2}T^2}\sum_{i=1}^{N}\lambda_{\max}(\boldsymbol{\Omega}_{v,i})^{2}\mathbb{E}\left[\left\|\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\right\|^{2}\right]\notag\\
&\leq\sup_{i=1,\dots,N}\lambda_{\max}(\boldsymbol{\Omega}_{v,i})^{2}\frac{1}{N^{2}T^2}\sum_{i=1}^{N}\mathbb{E}\left[\left\Vert \frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\frac{1}{\sqrt{T}}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\right\Vert ^{2}\right]\notag\\
&=O(1)\frac{1}{NT^2}O\left(\frac{K^{2}}{N^{2}}\zeta_{1}(K)^{4}\right)\notag\\
&=O\left(\frac{K^{2}}{N^{3}T^2}\zeta_{1}(K)^{4}\right),
\end{align} where we have used the result in (ref) and Assumption (ref). Hence,
\begin{align}
\left\|\mathbf{b}_{2,1}\right\|&=\left\Vert \frac{1}{NT}\sum_{i=1}^{N}\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\mathbf{V}_{i}\right\Vert \notag \\
&=O_{p}\left(\frac{K}{N^{\frac{3}{2}}T}\zeta_{1}(K)^{2}\right).
\end{align}
The second term on the right-hand site of (ref), $\*b_{2,2}\equiv \frac{1}{NT}\sum_{i=1}^N\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\mathbf{V}_{i}$, can be
evaluated in a similar way. Hence,
\begin{align}
\mathbb{E}\left(\left\|\mathbf{b}_{2,2}\right\|^{2}\right)&=\mathbb{E}\left[\left\|\frac{1}{NT}\sum_{i=1}^{N}\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\mathbf{V}_{i}\right\|^{2}\right]\notag\\
&=\mathbb{E}\left[\frac{1}{N^{2}T^{2}}\sum_{i=1}^{N}\sum_{j=1}^{N}\mathrm{tr}\left[\mathbf{V}_{i}'\widetilde{\mathbf{P}}\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\mathbb{E}\left(\mathbf{V}_{i}\mathbf{V}_{j}'\right)\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\mathbf{V}_{j}\right]\right]\notag\\
&\leq\sup_{i=1,\dots,N}\lambda_{\max}(\boldsymbol{\Omega}_{v,i})^{2}\frac{1}{N^{2}T^{2}}\sum_{i=1}^{N}\mathbb{E}\left[\left\Vert \left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\right\Vert ^{2}\right]\notag\\
&=O(1)\frac{1}{NT^{2}}\left[O\left(\frac{K^{\frac{3}{2}}}{N}\zeta_1(K)^2\zeta_0(K)\right) + O\left(\frac{K^{\frac{4}{3}}}{N}\zeta_1(K)^2\zeta_0(K)^{\frac{4}{3}}\right)\right]\notag\\
&=O\left(\frac{K^{\frac{3}{2}}}{N^{2}T^{2}}\zeta_{0}(K)\zeta_{1}(K)^{2}\right) + O\left(\frac{K^{\frac{4}{3}}}{N^2T^2}\zeta_1(K)^2\zeta_0(K)^{\frac{4}{3}}\right),
\end{align} where we have used the result from (ref). Hence,
\begin{align}
\left\Vert \mathbf{b}_{2,2}\right\Vert &=\left\Vert \frac{1}{NT}\sum_{i=1}^{N}\mathbf{V}_{i}'\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\mathbf{V}_{i}\right\Vert \notag \\
&=O_{p}\left(\frac{K^{\frac{3}{4}}}{NT}\sqrt{\zeta_{0}(K)}\zeta_{1}(K)\right) + O_p\left(\frac{K^{\frac{2}{3}}}{NT}\zeta_0(K)^{\frac{2}{3}}\zeta_1(K)\right).
\end{align}
The third term on the right-hand site of (ref) is of the same order as the second, implying
\begin{align}
\left\|\*b_{2,3}\right\| &\equiv \left\|\frac{1}{NT}\sum_{i=1}^N\mathbf{V}_{i}'\widetilde{\mathbf{P}}\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}\left(\widehat{\mathbf{P}}-\widetilde{\mathbf{P}}\right)'\mathbf{V}_{i}\right\| \notag \\
& =O_{p}\left(\frac{K^{\frac{3}{4}}}{NT}\sqrt{\zeta_{0}(K)}\zeta_{1}(K)\right).
\end{align}
We are left to analyze the last term. Following the same steps as before yields
\begin{align}
&\mathbb{E}\left(\left\|\mathbf{b}_{2,4}\right\|^{2}\right)=\mathbb{E}\left[\left\|\frac{1}{NT}\sum_{i=1}^{N}\mathbf{V}_{i}'\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\widetilde{\mathbf{P}}'\mathbf{V}_{i}\right\|^{2}\right]\notag\\
&=\frac{1}{N^{2}T^{2}}\sum_{i=1}^{N}\sum_{j=1}^{N}\mathbb{E}\Bigg[\mathrm{tr}\Big(\mathbf{V}_{i}'\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\widetilde{\mathbf{P}}'\mathbb{E}\left(\mathbf{V}_{i}\mathbf{V}_{j}'\right) \notag\\
& \quad\times \widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\widetilde{\mathbf{P}}'\mathbf{V}_{j}\Big)\Bigg]\notag\\
&\leq\sup_{i=1,\dots,N}\lambda_{\max}(\boldsymbol{\Omega}_{v,i})^{2}\frac{1}{N^{2}T^{2}}\sum_{i=1}^{N}\mathbb{E}\left[\left\|\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\widetilde{\mathbf{P}}'\right\|^{2}\right]\notag\\
&=O(1)\frac{1}{NT^{2}}O\left(\sqrt{K}\zeta_{0}(K)^{3}\right)\notag\\
&=O\left(\frac{\sqrt{K}}{NT^{2}}\zeta_{0}(K)^{3}\right),
\end{align} where we have used (ref). Hence,
\begin{align}
\left\Vert \mathbf{b}_{2,4}\right\Vert &=\left\Vert \frac{1}{NT}\sum_{i=1}^{N}\mathbf{V}_{i}'\widetilde{\mathbf{P}}\left[\left(\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)^{+}-\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\right]\widetilde{\mathbf{P}}'\mathbf{V}_{i}\right\Vert \notag\\
&=O_{p}\left(\frac{K^{\frac{1}{4}}}{\sqrt{N}T}\zeta_{0}(K)^{\frac{3}{2}}\right).
\end{align}
Collecting the results,
\begin{align}
\left\|\*b_{2}\right\| &\equiv \left\|\frac{1}{NT}\sum_{i=1}^N\mathbf{V}_{i}'\left(\mathbf{M}_{\widetilde{P}}-\mathbf{M}_{\widehat{P}}\right)\mathbf{V}_{i}\right\| \notag \\
& \leq \left\|\*b_{2,1}\right\| + \left\|\*b_{2,2}\right\| + \left\|\*b_{2,3}\right\| + \left\|\*b_{2,4}\right\| \notag \\
& =O_{p}\left(\frac{K}{N^{\frac{3}{2}}T}\zeta_{1}(K)^{2}\right) + O_{p}\left(\frac{K^{\frac{3}{4}}}{NT}\sqrt{\zeta_{0}(K)}\zeta_{1}(K)\right)\notag \\
&\quad + O_p\left(\frac{K^{\frac{2}{3}}}{NT}\zeta_0(K)^{\frac{2}{3}}\zeta_1(K)\right) + O_{p}\left(\frac{K^{\frac{1}{4}}}{\sqrt{N}T}\zeta_{0}(K)^{\frac{3}{2}}\right) \notag \\
&=o_p(1),
\end{align} where the last equality holds by (ref), (ref) and (ref).
The above results can be used to determine the bound of the first term, $\*b_1$. Note how
\begin{align}
\left\|\*b_1\right\| & \equiv \left\|\frac{1}{NT}\sum_{i=1}^{N}\mathbf{V}_{i}'\mathbf{M}_{\widetilde{P}}\mathbf{V}_{i}\right\| \notag \\
&= \left\Vert \frac{1}{NT}\sum_{i=1}^{N}\mathbf{V}_{i}'\widetilde{\mathbf{P}}\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\mathbf{V}_{i}\right\Vert \nonumber \\
& =\left\|\frac{1}{NT^{2}}\sum_{i=1}^{N}\mathbf{V}_{i}'\widetilde{\mathbf{P}}\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\mathbf{V}_{i}\right\|\nonumber \\
& \leq\frac{1}{NT}\sum_{i=1}^{N}\left\|\frac{1}{T}\mathbf{V}_{i}'\widetilde{\mathbf{P}}\widetilde{\mathbf{P}}'\mathbf{V}_{i}\right\|+\frac{1}{NT}\sum_{i=1}^{N}\left\|\frac{1}{T}\mathbf{V}_{i}'\widetilde{\mathbf{P}}\left[\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right]\widetilde{\mathbf{P}}'\mathbf{V}_{i}\right\|\nonumber \\
& \leq\frac{1}{NT}\sum_{i=1}^{N}\left\|\frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}'\mathbf{V}_{i}\right\|^{2}+\frac{1}{NT}\sum_{i=1}^{N}\left\|\frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}'\mathbf{V}_{i}\right\|^{2}\left\|\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right\|.
\end{align}
The order of $\left\Vert \left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right\Vert $
is known by (ref). Hence, we are only left to investigate the first term in (ref). Using Assumption (ref) and (ref), as well as the normalization condition $\mathbb{E}\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)=\*I_K$, we can show
\begin{align}
\mathbb{E}\left(\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}'\mathbf{V}_{i}\right\Vert ^{2}\right) & =\mathbb{E}\left[\mathrm{tr}\left(\frac{1}{T}\mathbf{V}_{i}'\widetilde{\mathbf{P}}\widetilde{\mathbf{P}}'\mathbf{V}_{i}\right)\right]\nonumber \\
& =\mathbb{E}\left\{\mathrm{tr}\left[\frac{1}{T}\widetilde{\mathbf{P}}'\mathbb{E}\left(\mathbf{V}_{i}\mathbf{V}_{i}'\right)\widetilde{\mathbf{P}}\right]\right\}\nonumber \\
& =\mathbb{E}\left[\mathrm{tr}\left(\frac{1}{T}\widetilde{\mathbf{P}}'\boldsymbol{\Omega}_{v,i}\widetilde{\mathbf{P}}\right)\right]\nonumber \\
& \leq\lambda_{\max}\left(\boldsymbol{\Omega}_{v,i}\right)\mathrm{tr}\left[\mathbb{E}\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)\right]\nonumber \\
&= O(1)O(K) \notag \\
&=O(K).
\end{align}
Consequently,
\begin{align}
\left\|\*b_1\right\| &\equiv \left\Vert \frac{1}{NT}\sum_{i=1}^{N}\mathbf{V}_{i}'\widetilde{\mathbf{P}}\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\mathbf{V}_{i}\right\Vert \notag \\
& \leq\frac{1}{NT}\sum_{i=1}^{N}\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}'\mathbf{V}_{i}\right\Vert ^{2}+\frac{1}{NT}\sum_{i=1}^{N}\left\Vert \frac{1}{\sqrt{T}}\widetilde{\mathbf{P}}'\mathbf{V}_{i}\right\Vert ^{2}\left\Vert \left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}-\mathbf{I}_{K}\right\Vert \nonumber \\
& =\frac{1}{NT}\sum_{i=1}^{N}O_{p}(K)+\frac{1}{NT}\sum_{i=1}^{N}O_{p}(K)O_{p}\left(\sqrt{\frac{K}{T}}\zeta_{0}(K)\right)\nonumber \\
& =O_{p}\left(\frac{K}{T}\right)+O_{p}\left(\frac{K\sqrt{K}}{T\sqrt{T}}\zeta_{0}(K)\right)\nonumber \\
& =O_p\left(\frac{K}{T}\right) + O_p
\left(\frac{K^{\frac{3}{2}}}{T^{\frac{3}{2}}}\zeta_0(K)\right) \notag \\
&=o_p(1),
\end{align} which holds by Assumption (ref)(ref) and the result in (ref).
By using this and applying a law of large numbers to $\frac{1}{NT}\sum_{i=1}^{N}\mathbf{V}_{i}'\mathbf{V}_{i}$,
\begin{align}
\frac{1}{NT}\sum_{i=1}^{N}\mathbf{V}_{i}'\mathbf{M}_{\widetilde{P}}\mathbf{V}_{i} & =\frac{1}{NT}\sum_{i=1}^{N}\mathbf{V}_{i}'\mathbf{V}_{i}-\frac{1}{NT}\sum_{i=1}^{N}\mathbf{V}_{i}'\widetilde{\mathbf{P}}\left(\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)^{+}\widetilde{\mathbf{P}}'\mathbf{V}_{i}\nonumber \\
& =\frac{1}{NT}\sum_{i=1}^{N}\mathbf{V}_{i}'\mathbf{V}_{i}+o_{p}(1)\nonumber \\
& =\boldsymbol{\Sigma}_{v}+o_{p}(1),
\end{align}
where $\boldsymbol{\Sigma}_{v}$ is as defined in Assumption (ref)(ref). Insertion into (ref) yields
\begin{equation}
\frac{1}{NT}\sum_{i=1}^{N}\mathbf{X}_{i}'\mathbf{M}_{\widehat{P}}\mathbf{X}_{i}=\boldsymbol{\Sigma}_{v}+o_{p}(1),
\end{equation}
which, together with (ref) and (ref),
leads to the following cleaned-up asymptotic representation for $\sqrt{NT}\left(\widehat{\boldsymbol{\beta}}_{SCCE}-\boldsymbol{\beta}\right)$:
\begin{align}
\sqrt{NT}\left(\widehat{\boldsymbol{\beta}}_{SCCE}-\boldsymbol{\beta}\right) & =\left(\frac{1}{NT}\sum_{i=1}^{N}\mathbf{X}_{i}'\mathbf{M}_{\widehat{P}}\mathbf{X}_{i}\right)^{-1}\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{X}_{i}'\mathbf{M}_{\widehat{P}}\left[\boldsymbol{\varepsilon}_{i}-\left(\widehat{\mathbf{P}}\+\alpha_{g_i}-\*g_i(\mathbf{F})\right)\right]\nonumber \\
& =\boldsymbol{\Sigma}_{v}^{-1}\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\boldsymbol{\varepsilon}_{i}+o_{p}(1).
\end{align} In order to arrive at the sought asymptotic distribution, we apply
a central limit law for independent processes to $\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\boldsymbol{\varepsilon}_{i}$. This gives
\begin{equation}
\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\boldsymbol{\varepsilon}_{i}\overset{d}{\to}N(\mathbf{0}_{d\times1},\boldsymbol{\Theta})
\end{equation}
as $N,T\to\infty$, where
\begin{align}
\boldsymbol{\Theta} & =\lim_{N,T\to\infty}\mathbb{E}\left[\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\boldsymbol{\varepsilon}_{i}\left(\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\boldsymbol{\varepsilon}_{i}\right)'\right]\nonumber \\
& =\lim_{N,T\to\infty}\frac{1}{NT}\sum_{i=1}^{N}\sum_{j=1}^{N}\mathbb{E}\left[\mathbf{V}_{i}'\mathbb{E}\left(\boldsymbol{\varepsilon}_{i}\boldsymbol{\varepsilon}_{j}\right)\mathbf{V}_{j}\right]\nonumber \\
& =\lim_{N,T\to\infty}\frac{1}{NT}\sum_{i=1}^{N}\mathbb{E}\left[\mathbf{V}_{i}'\mathbb{E}\left(\boldsymbol{\varepsilon}_{i}\boldsymbol{\varepsilon}_{i}'\right)\mathbf{V}_{i}\right]\nonumber \\
& =\lim_{N,T\to\infty}\frac{1}{NT}\sum_{i=1}^{N}\mathbb{E}\left(\mathbf{V}_{i}'\boldsymbol{\Omega}_{\varepsilon,i}\mathbf{V}_{i}\right).
\end{align}
We can therefore show that under Assumptions (ref)--(ref),
\begin{equation}
\sqrt{NT}\left(\widehat{\boldsymbol{\beta}}_{SCCE}-\boldsymbol{\beta}\right)=\boldsymbol{\Sigma}_{v}^{-1}\frac{1}{\sqrt{NT}}\sum_{i=1}^{N}\mathbf{V}_{i}'\boldsymbol{\varepsilon}_{i}+o_{p}(1)\overset{d}{\to}N(\mathbf{0}_{d\times1},\boldsymbol{\Sigma}_{v}^{-1}\boldsymbol{\Omega}\boldsymbol{\Sigma}_{v}^{-1}),
\end{equation} as $N,T\to\infty.$ $\blacksquare$
\@startsection{subsection}{2}
\z@{-.5\linespacing\@plus-.7\linespacing}{.5\linespacing}
{\normalfont}{Proof of Lemma (ref)} Consider (ref). In view of $\mathbb{E}\left(\frac{1}{T}\widetilde{\mathbf{P}}' \widetilde{\mathbf{P}}\right) = \*I_K$, $\mathbb{E}\left(\widetilde{\*p}(\*f_t)\widetilde{\*p}(\*f_t)'\right)= \*I_K$, for all $t=1, \dots, T$. Hence, we can consider the following expression
\begin{align}
&\mathbb{E} \left[\left\|\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}-\*I_K\right\|^2\right]=\mathbb{E}\left\{\mathrm{tr}\left[\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}-\*I_K\right)'\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}-\*I_K\right)\right]\right\} \notag \\
&=\mathbb{E}\left[\frac{1}{T^2}\sum_{t=1}^T\sum_{s=1}^T\mathrm{tr}\left(\widetilde{\*p}(\*f_t)\widetilde{\*p}(\*f_t)'\widetilde{\*p}(\*f_s)\widetilde{\*p}(\*f_s)'\right)\right] -2\mathrm{tr}\left[\mathbb{E}\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)\right]+\mathbb{E}\left[\mathrm{tr}(\*I_K)\right] \notag \\
&=\mathbb{E}\left[\frac{1}{T^2}\sum_{t=1}^T\sum_{s=1}^T\mathrm{tr}\left(\widetilde{\*p}(\*f_t)\widetilde{\*p}(\*f_t)'\widetilde{\*p}(\*f_s)\widetilde{\*p}(\*f_s)'\right)\right] -2\mathrm{tr}\left(\*I_K\right)+\mathrm{tr}(\*I_K) \notag \\
&=\mathbb{E}\left[\frac{1}{T^2}\sum_{t=1}^T\sum_{s=1}^T\mathrm{tr}\left(\widetilde{\*p}(\*f_t)'\widetilde{\*p}(\*f_s)\widetilde{\*p}(\*f_s)'\widetilde{\*p}(\*f_t)\right)\right]-K \notag \\
&=\frac{1}{T^2}\sum_{t=1}^T\mathbb{E}\left(\left\|\widetilde{\*p}(\*f_t)\right\|^4\right)+\frac{1}{T^2}\sum_{t=1}^T\sum_{\substack{s=1 \\ s\neq t}}^T\mathbb{E}\left(\widetilde{\*p}(\*f_t)'\widetilde{\*p}(\*f_s)\widetilde{\*p}(\*f_s)'\widetilde{\*p}(\*f_t)\right)-K \notag \\
&\equiv A_1 + A_2.
\end{align} Starting with $A_1$, let us derive
\begin{align}
A_1 &\equiv \frac{1}{T^2}\sum_{t=1}^T\mathbb{E}\left(\left\|\widetilde{\*p}(\*f_t)\right\|^4\right) \leq \frac{1}{T^2}\sum_{t=1}^T\mathbb{E}\left(\left\|\widetilde{\*p}(\*f_t)\right\|^2\right)\zeta_0(K)^2 \notag \\
&\leq \frac{1}{T^2} \sum_{t=1}^T \mathrm{tr}\left[\mathbb{E}\left(\widetilde{\*p}(\*f_t)\widetilde{\*p}(\*f_t)'\right)\right]\zeta_0(K)^2 =\frac{1}{T^2} \sum_{t=1}^T \mathrm{tr}\left[\*I_K\right]\zeta_0(K)^2 \notag \\
&=O\left(\frac{K \zeta_0(K)^2}{T}\right).
\end{align} Now, let us continue with $A_2$. By Davydov's inequality (see for example bosq, pp. 20-21) for $\alpha$-mixing processes, we have
\begin{align}
A_2 &\equiv \frac{1}{T^2}\sum_{t=1}^T\sum_{\substack{s=1 \\ s\neq t}}^T\mathbb{E}\left(\widetilde{\*p}(\*f_t)'\widetilde{\*p}(\*f_s)\widetilde{\*p}(\*f_s)'\widetilde{\*p}(\*f_t)\right)-K \notag \\
&= \frac{1}{T^2}\sum_{t=1}^T\sum_{\substack{s=1 \\ s\neq t}}^T\sum_{j=1}^K\sum_{k=1}^K\mathbb{E}\left(\widetilde{p}_j(\*f_t)\widetilde{p}_j(\*f_s)\widetilde{p}_k(\*f_t)\widetilde{p}_k(\*f_s)\right)-K \notag \\
&=\frac{1}{T^2}\sum_{t=1}^T\sum_{\substack{s=1 \\ s\neq t}}^T\sum_{j=1}^K\sum_{k=1}^K \text{Cov}\left(\widetilde{p}_j(\*f_t)\widetilde{p}_k(\*f_t), \widetilde{p}_j(\*f_s)\widetilde{p}_k(\*f_s)\right) \notag \\
&\quad + \frac{1}{T^2}\sum_{t=1}^T\sum_{\substack{s=1 \\ s\neq t}}^T\sum_{j=1}^K\sum_{k=1}^K \mathbb{E}\left(\widetilde{p}_j(\*f_t)\widetilde{p}_k(\*f_t)\right)\mathbb{E}\left(\widetilde{p}_j(\*f_s)\widetilde{p}_k(\*f_s)\right)-K\notag \\
&=\frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{\substack{s=1\\
s\neq t
}
}^{T}\sum_{j=1}^{K}\sum_{k=1}^{K}\text{Cov}\left(\widetilde{p}_{j}(\mathbf{f}_{t})\widetilde{p}_{k}(\mathbf{f}_{t}),\widetilde{p}_{j}(\mathbf{f}_{s})\widetilde{p}_{k}(\mathbf{f}_{s})\right)+\frac{T(T-1)}{T^{2}}\sum_{j=1}^{K}1-K\notag \\
&\leq\frac{1}{T^{2}}\frac{8+2\eta}{8+\eta}2^{\frac{\eta}{8+\eta}}\sum_{t=1}^{T}\sum_{\substack{s=1\\
s\neq t
}
}^{T}\sum_{j=1}^{K}\sum_{k=1}^{K}\alpha\left(\left|t-s\right|\right)^{\frac{\eta}{8+\eta}}\mathbb{E}\left(\left|\widetilde{p}_{j}(\mathbf{f}_{t})\widetilde{p}_{k}(\mathbf{f}_{t})\right|^{\frac{8+\eta}{2}}\right)^{\frac{2}{8+\eta}}\notag \\
&\quad \times \mathbb{E}\left(\left|\widetilde{p}_{j}(\mathbf{f}_{s})\widetilde{p}_{k}(\mathbf{f}_{s})\right|^{\frac{8+\eta}{2}}\right)^{\frac{2}{8+\eta}}-\frac{K}{T}\notag \\
&\leq\frac{c_{\eta}}{T^{2}}\sum_{t=1}^{T}\sum_{s\neq t}^{T-1}\sum_{j=1}^{K}\sum_{k=1}^{K}\alpha\left(\left|t-s\right|\right)^{\frac{\eta}{8+\eta}}\left\{ \mathbb{E}\left(\left|\widetilde{p}_{j}(\mathbf{f}_{t})\widetilde{p}_{k}(\mathbf{f}_{t})\right|^{\frac{8+\eta}{2}}\right)\right\} ^{\frac{4}{8+\eta}}-\frac{K}{T}\notag \\
&\leq\frac{c_{\eta}}{T^{2}}\sum_{t=1}^{T}\sum_{s\neq t}^{T-1}\sum_{j=1}^{K}\sum_{k=1}^{K}\alpha\left(\left|t-s\right|\right)^{\frac{\eta}{8+\eta}}\left\{ \mathbb{E}\left(\left|\widetilde{p}_{j}(\mathbf{f}_{t})\right|^{8+\eta}\right)\mathbb{E}\left(\left|\widetilde{p}_{k}(\mathbf{f}_{t})\right|^{8+\eta}\right)\right\} ^{\frac{2}{8+\eta}}-\frac{K}{T}\notag \\
&\leq\frac{c_{\eta}}{T}\Bigg[\sum_{l=1}^{T-1}\alpha(l)^{\frac{\eta}{8+\eta}}\Bigg]\sum_{j=1}^{K}\sum_{k=1}^{K}\left\{ \mathbb{E}\left(\left|\widetilde{p}_{j}(\mathbf{f}_{t})\right|^{8+\eta}\right)\mathbb{E}\left(\left|\widetilde{p}_{k}(\mathbf{f}_{t})\right|^{8+\eta}\right)\right\} ^{\frac{2}{8+\eta}}-\frac{K}{T}\notag \\
&=O\left(\frac{K^{2}}{T}\right),
\end{align} which holds under Assumptions (ref) and (ref)(ref). Here, we define
$c_{\eta}\equiv \frac{8+2\eta}{8+\eta}2^{\frac{\eta}{8+\eta}}$ and $l\equiv \left|t-s\right|$. Taking the results for $A_1$ and $A_2$ together, we can bound $\mathbb{E}\left[\left\|\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}-\*I_K\right\|^2\right]$ as
\begin{align}
\mathbb{E}\left[\left\|\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}-\*I_K\right\|^2\right] &\equiv A_1 + A_2 =O\left(\frac{K\zeta_0(K)^2}{T}\right) + O\left(\frac{K^2}{T}\right) = O\left(\frac{K\zeta_0(K)^2}{T}\right),
\end{align} since typically, $\zeta_0(K)=O(\sqrt{K})$ or $\zeta_0(K)=O(K)$.
Therefore,
\begin{align}
\left\|\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}-\*I_K\right\| O_p\left(\sqrt{\frac{K}{T}}\zeta_0(K)\right) = o_p(1),
\end{align} which holds under the Assumption (ref)(ref) condition that $\sqrt{\frac{K}{T}}\zeta_0(K)^{\frac{3}{2}}=o(1)$.
Now consider (ref). By Lemma (ref)(ref), we have $\mathbb{E}\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)=\*I_K$. Moreover, for any matrix $\*A\in \mathbb{R}^{K\times K}$, every eigenvalue $\lambda_j(A)$ satisfies $\left|\lambda_j(A)\right|\leq \left\|\*A\right\|$, for all $j=1, \dots, K$ (see, for example, horn). Using this we can show that
\begin{align}
\left|\lambda_{\max}\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)-\lambda_{\max}\left[\mathbb{E}\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)\right]\right| &= \left|\lambda_{\max}\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)-\lambda_{\max}\left(\*I_K\right)\right| \notag \\
&\leq \left\|\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}-\*I_K\right\| = o_p(1),
\end{align} which holds by (ref). This implies that
\begin{align}
\lambda_{\max}\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)-\lambda_{\max}\left(\*I_K\right) = \lambda_{\max}\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)-1 = o_p(1),
\end{align} and
\begin{align}
\lambda_{\max}\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right) = 1+o_p(1) = O_p(1).
\end{align} Since exactly the same argumentation applies to $\lambda_{\min}\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)$,
\begin{align}
\lambda_{\min}\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right) = 1+o_p(1) = O_p(1),
\end{align}which establishes (ref).
To prove (ref), we can first consider the following expression
\begin{align}
\left\|\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}-\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right\| &= \left\|\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}-\*I_K-\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}-\*I_K\right)\right\| \notag \\
&\leq \left\|\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}-\*I_K\right\|+\left\|\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}-\*I_K\right\| \notag \\
&= O_p\left(\delta_{NTK}\right)+ o_p(1) = O_p\left(\delta_{NTK}\right) =o_p(1),
\end{align} where the last equality holds under Lemma (ref)(ref), as well as the results in (ref) and (ref). Consequently, by the same arguments used to prove (ref),
\begin{align}
\left|\lambda_{\min}\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)-\lambda_{\min}\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)\right| \leq \left\|\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}-\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right\|= o_p(1).
\end{align} Under Lemma (ref)(ref), and the normalization $\mathbb{E}\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)=\*I_K$,
\begin{align}
\lambda_{\max}\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)=\lambda_{\min}\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)=1+o_p(1),
\end{align} and hence,
\begin{align}
\lambda_{\min}\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)=\lambda_{\min}\left(\frac{1}{T}\widetilde{\mathbf{P}}'\widetilde{\mathbf{P}}\right)+o_p(1)=1+o_p(1)=O_p(1),
\end{align} and also,
\begin{align}
\lambda_{\max}\left(\frac{1}{T}\widehat{\mathbf{P}}'\widehat{\mathbf{P}}\right)=1+o_p(1)=O_p(1),
\end{align} which holds under exactly the same arguments as for the proof of (ref). This establishes (ref). The proof of Lemma (ref) is therefore complete. $\blacksquare$
\section{Additional Results: Empirical Application}
In this section, we provide additional illustrations for our empirical application. Figure (ref) shows the development of wage inequality in U.S. manufacturing over the sample period. Table (ref) presents results from a Dickey-Fuller test for the estimated factors, supporting our hypothesis that they are unit-root non stationary. Figure (ref) plots $\widehat{\*f}_t = [\overline{\ln{\sigma_t}}, \overline{\ln \left(\frac{H_t}{L_t}\right)}, \overline{\*z}_t]'$ over time, after taking first differences, indicating mean reversion. Table (ref) presents SCCE estimates with 95% confidence intervals based on HAC standard errors, and Figure (ref) shows the estimated effects of our main regressors, $\ln(\sigma_{i,t})$ and $\ln\left(\frac{H_{i,t}}{L_{i,t}}\right)$, across different number of knots, $J$, using HAC standard errors.
\setcounter{figure}{0}
\begin{figure}[H]
\caption{Average wage inequality in U.S. manufacturing sector for the time period 1958 -- 2005.}
\begin{minipage}{\textwidth}
\scriptsize \textit{Notes:} We calculate the average wage inequality as $\overline{\ln\left(\frac{w_t^{(L)}}{w_t^{(H)}}\right)}=\frac{1}{N}\sum_{i=1}^N \ln\left(\frac{w_{i,t}^{(L)}}{w_{i,t}^{(H)}}\right)$, for manufacturing sectors $i=1, \dots, N$. Here, $\frac{w_{i,t}^{(L)}}{w_{i,t}^{(H)}}$ is the relative wage of low-skilled workers to high-skilled workers in a U.S. manufacturing sector $i$ at time $t$. This illustration is based on a balanced version of the data from voig.
\end{minipage}
\end{figure}
\begin{table}
\begin{threeparttable}
\caption{Results from Augmented Dickey-Fuller test.}
\begin{tabular*}{0.7\textwidth}{@{\extracolsep{\fill}}lcc}
\toprule
Factor & ADF & $p$-value \\
\midrule
$\overline{\ln(\sigma_{t})}$ & -1.02 & 0.75 \rule{0pt}{3ex}\\
$\overline{\ln\left(\frac{H_{t}}{L_{t}}\right)}$ & -0.57 & 0.88 \rule{0pt}{3ex}\\
$\overline{z}_{1,t}$ &0.98 & 0.99 \rule{0pt}{3ex}\\
$\overline{z}_{2,t}$ & 0.90 & 0.99 \rule{0pt}{3ex}\\
$\overline{z}_{3,t}$ & 0.76 & 0.99 \rule{0pt}{3ex}\\
$\overline{z}_{4,t}$ & -1.49 & 0.54 \rule{0pt}{3ex}\\
$\overline{z}_{5,t}$ & 2.39 & 0.99 \rule{0pt}{3ex}\\
$\overline{z}_{6,t}$ & 2.45 & 0.99 \rule{0pt}{3ex}\\
\bottomrule
\end{tabular*}
\begin{tablenotes}
• \scriptsize \textit{Notes}: The lag order in the ADF test is selected using BIC. $\overline{z}_{1,t}, \dots, \overline{z}_{6,t}$ denote auxiliary regressors used as controls and are further specified in Table (ref).
\end{tablenotes}
\end{threeparttable}
\end{table}
\begin{figure}[p]
\caption{Plotting $\widehat{\*f}_t = [\overline{\ln{\sigma_t}}, \overline{\ln \left(\frac{H_t}{L_t}\right)}, \overline{\*z}_t]'$ over time, after taking first differences.}
\begin{subfigure}{0.49\textwidth}
\caption{$\overline{\ln{\sigma_t}}$ and $\overline{\ln \left(\frac{H_t}{L_t}\right)}$.}
\end{subfigure}
\begin{subfigure}{0.49\textwidth}
\caption{The elements of $\overline{\mathbf{z}}_t$.}
\end{subfigure}
\scriptsize
\begin{minipage}[t]{\textwidth}
\textit{Notes:} The figure plots the estimated factors over the 1958-2005 period. Figure B2.(A) plots the cross-sectional averages of the two key regressors, $\ln\left(\sigma_{i,t}\right)$ and $\ln\left(\frac{H_{i,t}}{L_{i,t}}\right)$, while Figure B2.(B) plots the controls included in $\mathbf{z}_{i,t}$.
\end{minipage}
\end{figure}
\begin{table}[p]
\begin{threeparttable}
\caption{Estimation results for SCCE with HAC standard errors.}
\begin{tabular*}{0.7\textwidth}{@{\extracolsep{\fill}}lcc}
\toprule
& Narrow model & Full model \\
\midrule
$\ln(\sigma_{i,t})$ & -0.34 & -0.33 \\
& (-0.62; -0.06) & (-0.62; -0.05) \\
$\ln\!\left(\frac{H_{i,t}}{L_{i,t}}\right)$ & 0.56 & 0.61 \\
& (0.52; 0.59) & (0.58; 0.64) \\
$z_{1,i,t}$ & & -0.03 \\
& & (-0.37; 0.31) \\
$z_{2,i,t}$ & & 3.15 \\
& & (1.20; 5.11) \\
$z_{3,i,t}$ & & -1.03 \\
& & (-2.09; 0.04) \\
$z_{4,i,t}$ & & -0.20 \\
& & (-0.55; 0.16) \\
$z_{5,i,t}$ & & -0.17 \\
& & (-0.33; -0.00) \\
$z_{6,i,t}$ & & 0.07 \\
& & (-0.14; 0.28) \\
\bottomrule
\end{tabular*}
\begin{tablenotes}
• \scriptsize \textit{Notes}: The table reports point estimates of the coefficients of the model in (ref) with the associated 95% confidence intervals appearing within parentheses, based on HAC standard errors. Results are reported for our SCCE estimator for the narrow and full model. The regressors of interest are $\ln\left(\sigma_{i,t}\right)$, which is a measure of the input skill intensity, and $\ln\left(\frac{H_{i,t}}{L_{i,t}}\right)$, which is the ratio of high-skilled to low-skilled labor. The six control variables are; real capital equipment per worker ($z_{1,i,t}$), the sectoral share of office, computing and accounting equipment ($z_{2,i,t}$), the difference between high-tech and computer capital share ($z_{3,i,t}$), R&D intensity ($z_{4,i,t}$), narrow outsourcing ($z_{5,i,t}$), and the difference between broad and narrow outsourcing ($z_{6,i,t}$). The dependent variable is the log of the relative wage of low-skilled to high-skilled workers.
\end{tablenotes}
\end{threeparttable}
\end{table}
\begin{figure}[p]
\caption{Estimated effects of $\ln \left(\sigma_{i,t}\right)$ and $\ln \left(\frac{H_{i,t}}{L_{i,t}}\right)$ for different number of knots, $J$, using HAC standard errors.}
\begin{subfigure}{0.49\textwidth}
\caption{The effect of $\ln (\sigma_{i,t})$ in the narrow model.}
\end{subfigure}
\begin{subfigure}{0.49\textwidth}
\caption{The effect of $\ln \left(\sigma_{i,t}\right)$ in the full model.}
\end{subfigure}
\begin{subfigure}{0.49\textwidth}
\caption{The effect of $\ln \left(\frac{H_{i,t}}{L_{i,t}}\right)$ in the narrow model.}
\end{subfigure}
\begin{subfigure}{0.49\textwidth}
\caption{The effect of $\ln \left(\frac{H_{i,t}}{L_{i,t}}\right)$ in the full model.}
\end{subfigure}
\scriptsize
\begin{minipage}{\textwidth}
\textit{Notes:} The figure plots point estimates obtained by reestimating (ref) by SCCE for $J\in \{2, 4, 6, 8, 10\}$. The grey area represents 95% confidence intervals based on HAC standard errors.
\end{minipage}
\end{figure}
\section{Additional Results: Monte Carlo Simulations}
This section presents additional results from a small-scale Monte Carlo simulation exercise, with a focus on various extensions that show the flexibility of our SCCE estimator. The specifications are basically Experiment \textbf{E1} in Section (ref) with minor deviations. Hence, we refer the reader there for more details. The following experiments are considered:
\begin{enumerate}
• Here, we report the numerical results of our estimator for the data-generating process as in Experiment \textbf{E1} in Section (ref).
• Here, we report the numerical results of our estimator, assuming interactive effects as in Experiment E2 in section (ref).
• Here, we allow the factors to be unit root non-stationary. More specifically, we generate $\*f_t = \*f_{t-1}+\*w_t$, with $\*w_t \sim \mathcal{N}(\*0_{m \times 1}, 0.05 \cdot \*I_m)$ and $\*f_0=\*0_{m \times 1}$.
• Here, we consider an infeasible estimator, where we use the true factors as input for the sieve basis.
• Here, we consider a DGP exactly as in E1 but use Hermite polynomials instead of splines as a sieve basis. For a definition of Hermite polynomials see, for example, chen2007.
• Here, relax the independence assumption on the error terms in (ref) and (ref), and allow for weak serial and cross-sectional correlations as well as cross-sectional heteroscedasticity. We follow kad25, and generate $\varepsilon_{i,t}=\pi \varepsilon_{i,t-1} + \theta_{i,t}+\sum_{l=1}^L \pi (\theta_{i-l,t} +\theta_{i+l,t})$, where $\varepsilon_{1,0}= \dots = \varepsilon_{N,0}=0$, $L=5$, and $\theta_{i,1} \sim \mathcal{N}(0, \sigma_i^2)$ with $\sigma_i \sim \mathcal{U}(0.5,1)$. Further, $\pi \in \{0.2, 0.5\}$, where $\pi = 0.2$ implies "low", and $\pi = 0.5 $ implies "high" dependence. $\*v_{i,t}$ is generated in the same way.
• Here, we consider different number of knots for the spline basis in Experiment \textbf{E1}. More specifically, we will set $J = C \lfloor T^{\frac{1}{4}} \rfloor$, with $C\in \{2,3,4\}$. Further, we also consider different rates for $T$ i.e., $J=\{T^{\frac{1}{3}}, T^{\frac{1}{5}}, T^{\frac{1}{10}}\}$.
\end{enumerate}
We ran $1{,}000$ simulations with $N,T \in \{20,50,100,200,300\}$. Given the computational load from high-dimensional, nonlinear estimation, the Monte Carlo experiments were executed on LUNARC's COSMOS cluster at Lund University. COSMOS comprises 182 compute nodes, each with two AMD 7413 processors (Milan), resulting in 48 cores per node. All simulations were implemented in Python 3.9.18.
\begin{landscape}
\begin{table}[p]
\begin{threeparttable}
\caption{Simulation results for Experiments E1.A, E1.B, E1.C, E1.D, and E1.E.}
\begin{tabular}{rrrrrrrrrrrr}
\toprule
& & \multicolumn{2}{c}{E1.A} & \multicolumn{2}{c}{E1.B} & \multicolumn{2}{c}{E1.C} & \multicolumn{2}{c}{E1.D} & \multicolumn{2}{c}{E1.E} \\
\cmidrule(lr){3-4}\cmidrule(lr){5-6}\cmidrule(lr){7-8}\cmidrule(lr){9-10}\cmidrule(lr){11-12}
N & T & Bias & RMSE & Bias & RMSE & Bias & RMSE & Bias & RMSE & Bias & RMSE \\
\midrule
20 & 20 & 0.0014 & 0.1420 & 0.0016 & 0.1183 & 0.0017 & 0.1166 & 0.0001 & 0.0660 & 0.0016 & 0.0861 \\
20 & 50 & 0.0011 & 0.0478 & 0.0008 & 0.0420 & 0.0009 & 0.0406 & 0.0003 & 0.0332 & 0.0013 & 0.0449 \\
20 & 100 & 0.0016 & 0.0335 & 0.0022 & 0.0273 & 0.0001 & 0.0260 & 0.0000 & 0.0235 & 0.0017 & 0.0333 \\
20 & 200 & 0.0003 & 0.0258 & 0.0002 & 0.0220 & 0.0005 & 0.0188 & 0.0003 & 0.0177 & 0.0001 & 0.0269 \\
20 & 300 & 0.0008 & 0.0226 & 0.0005 & 0.0201 & 0.0009 & 0.0169 & 0.0003 & 0.0140 & 0.0006 & 0.0228 \\
50 & 20 & 0.0042 & 0.0859 & 0.0011 & 0.0704 & 0.0002 & 0.0714 & 0.0019 & 0.0423 & 0.0001 & 0.0534 \\
50 & 50 & 0.0015 & 0.0317 & 0.0002 & 0.0241 & 0.0008 & 0.0247 & 0.0008 & 0.0224 & 0.0009 & 0.0309 \\
50 & 100 & 0.0009 & 0.0215 & 0.0004 & 0.0159 & 0.0011 & 0.0162 & 0.0001 & 0.0149 & 0.0008 & 0.0226 \\
50 & 200 & 0.0003 & 0.0148 & 0.0002 & 0.0109 & 0.0003 & 0.0121 & 0.0000 & 0.0104 & 0.0001 & 0.0156 \\
50 & 300 & 0.0003 & 0.0131 & 0.0003 & 0.0089 & 0.0001 & 0.0100 & 0.0002 & 0.0086 & 0.0003 & 0.0138 \\
100 & 20 & 0.0054 & 0.0584 & 0.0004 & 0.0484 & 0.0005 & 0.0504 & 0.0007 & 0.0310 & 0.0028 & 0.0383 \\
100 & 50 & 0.0008 & 0.0210 & 0.0000 & 0.0175 & 0.0011 & 0.0169 & 0.0001 & 0.0157 & 0.0009 & 0.0212 \\
100 & 100 & 0.0005 & 0.0143 & 0.0001 & 0.0112 & 0.0001 & 0.0120 & 0.0007 & 0.0107 & 0.0002 & 0.0147 \\
100 & 200 & 0.0003 & 0.0102 & 0.0002 & 0.0075 & 0.0003 & 0.0080 & 0.0002 & 0.0076 & 0.0003 & 0.0108 \\
100 & 300 & 0.0004 & 0.0089 & 0.0002 & 0.0062 & 0.0001 & 0.0066 & 0.0002 & 0.0057 & 0.0005 & 0.0093 \\
200 & 20 & 0.0041 & 0.0568 & 0.0008 & 0.0356 & 0.0006 & 0.0350 & 0.0001 & 0.0206 & 0.0017 & 0.0272 \\
200 & 50 & 0.0006 & 0.0149 & 0.0004 & 0.0125 & 0.0006 & 0.0127 & 0.0002 & 0.0108 & 0.0007 & 0.0145 \\
200 & 100 & 0.0004 & 0.0100 & 0.0003 & 0.0084 & 0.0002 & 0.0081 & 0.0002 & 0.0077 & 0.0002 & 0.0102 \\
200 & 200 & 0.0001 & 0.0069 & 0.0001 & 0.0054 & 0.0002 & 0.0056 & 0.0001 & 0.0050 & 0.0003 & 0.0074 \\
200 & 300 & 0.0000 & 0.0062 & 0.0001 & 0.0043 & 0.0002 & 0.0046 & 0.0002 & 0.0042 & 0.0001 & 0.0067 \\
300 & 20 & 0.0000 & 0.0326 & 0.0024 & 0.0288 & 0.0003 & 0.0291 & 0.0008 & 0.0170 & 0.0000 & 0.0214 \\
300 & 50 & 0.0006 & 0.0120 & 0.0004 & 0.0097 & 0.0001 & 0.0100 & 0.0007 & 0.0088 & 0.0004 & 0.0118 \\
300 & 100 & 0.0002 & 0.0084 & 0.0003 & 0.0067 & 0.0003 & 0.0066 & 0.0003 & 0.0062 & 0.0003 & 0.0087 \\
300 & 200 & 0.0001 & 0.0058 & 0.0002 & 0.0042 & 0.0000 & 0.0046 & 0.0001 & 0.0043 & 0.0002 & 0.0062 \\
300 & 300 & 0.0000 & 0.0051 & 0.0002 & 0.0035 & 0.0000 & 0.0037 & 0.0001 & 0.0036 & 0.0000 & 0.0054 \\
\bottomrule
\end{tabular}
\begin{tablenotes}
• \scriptsize \textit{Notes}: The absolute bias and RMSE are computed as
$\frac{1}{S} \sum_{s=1}^S |\widehat{\beta}_{1,s} - \beta_1 |$ and $\sqrt{\frac{1}{S}\sum_{s=1}^S (\widehat{\beta}_{1,s} - \beta_1)^2}$, respectively, where $S =1,000$ is the number of simulations, and $\widehat{\beta}_{1,s}$ is the estimate of $\beta_1$ from replication $s$. While we focus here on the first coefficient, $\beta_1$, the results for the second coefficient, $\beta_2$, were similar and can be made are available upon request.
\end{tablenotes}
\end{threeparttable}
\end{table}
\end{landscape}
\begin{table}[p]
\begin{threeparttable}
\caption{Simulation results for Experiment E1.F.}
\begin{tabular*}{0.7\textwidth}{@{\extracolsep{\fill}}cccccc}
\toprule
&&\multicolumn{2}{c}{$\pi =0.2$} & \multicolumn{2}{c}{$\pi =0.5$}\\
\cmidrule(lr){3-4} \cmidrule(lr){5-6}
N & T & Bias & RMSE & Bias & RMSE \\
\toprule
20 & 20 & 0.0119 & 0.1911 & 0.0057 & 0.2733 \\
20 & 50 & 0.0003 & 0.0675 & 0.0022 & 0.0971 \\
20 & 100 & 0.0010 & 0.0481 & 0.0018 & 0.0678 \\
20 & 200 & 0.0005 & 0.0341 & 0.0015 & 0.0465 \\
20 & 300 & 0.0008 & 0.0300 & 0.0001 & 0.0402 \\
50 & 20 & 0.0066 & 0.1251 & 0.0049 & 0.1804 \\
50 & 50 & 0.0006 & 0.0433 & 0.0016 & 0.0655 \\
50 & 100 & 0.0003 & 0.0282 & 0.0014 & 0.0421 \\
50 & 200 & 0.0000 & 0.0206 & 0.0003 & 0.0298 \\
50 & 300 & 0.0002 & 0.0174 & 0.0006 & 0.0253 \\
100 & 20 & 0.0018 & 0.1010 & 0.0006 & 0.1303 \\
100 & 50 & 0.0009 & 0.0296 & 0.0005 & 0.0456 \\
100 & 100 & 0.0003 & 0.0187 & 0.0001 & 0.0293 \\
100 & 200 & 0.0004 & 0.0142 & 0.0003 & 0.0201 \\
100 & 300 & 0.0005 & 0.0120 & 0.0008 & 0.0170 \\
200 & 20 & 0.0003 & 0.0856 & 0.0025 & 0.0893 \\
200 & 50 & 0.0004 & 0.0214 & 0.0002 & 0.0329 \\
200 & 100 & 0.0001 & 0.0135 & 0.0002 & 0.0208 \\
200 & 200 & 0.0004 & 0.0091 & 0.0000 & 0.0137 \\
200 & 300 & 0.0003 & 0.0082 & 0.0006 & 0.0120 \\
300 & 20 & 0.0029 & 0.0496 & 0.0047 & 0.0706 \\
300 & 50 & 0.0000 & 0.0172 & 0.0001 & 0.0256 \\
300 & 100 & 0.0003 & 0.0112 & 0.0008 & 0.0169 \\
300 & 200 & 0.0003 & 0.0083 & 0.0004 & 0.0115 \\
300 & 300 & 0.0002 & 0.0067 & 0.0005 & 0.0098 \\
\bottomrule
\end{tabular*}
\begin{tablenotes}
• \scriptsize \textit{Notes}: See Table 1 for an explanation. Here, $\pi$ refers to serial and cross-sectional correlation of the errors in the equations (ref) and (ref).
\end{tablenotes}
\end{threeparttable}
\end{table}
\begin{landscape}
\begin{table}[p]
\begin{threeparttable}
\caption{Simulation results for Experiment E1.G.}
\begin{tabular}{rrrrrrrrrrrrrr}
\toprule
&&\multicolumn{2}{c}{$J = 2\lfloor T^{\frac{1}{4}}\rfloor$} & \multicolumn{2}{c}{$J = 3\lfloor T^{\frac{1}{4}}\rfloor$} & \multicolumn{2}{c}{$J = 4\lfloor T^{\frac{1}{4}}\rfloor$} & \multicolumn{2}{c}{$J=\lfloor T^{\frac{1}{3}}\rfloor$} & \multicolumn{2}{c}{$J=\lfloor T^{\frac{1}{5}}\rfloor$} & \multicolumn{2}{c}{$J=\lfloor T^{\frac{1}{10}}\rfloor$}\\
\cmidrule(lr){3-4} \cmidrule(lr){5-6} \cmidrule(lr){7-8} \cmidrule(lr){9-10} \cmidrule(lr){11-12} \cmidrule(lr){13-14}
N & T & Bias & RMSE & Bias & RMSE & Bias & RMSE & Bias & RMSE & Bias & RMSE & Bias & RMSE \\
\midrule
20 & 20 & 0.2027 & 6.5197 & 0.1158 & 3.3448 & 0.7988 & 28.9344 & 0.0014 & 0.1420 & 0.0004 & 0.1029 & 0.0004 & 0.1029 \\
20 & 50 & 0.0018 & 0.0513 & 0.0022 & 0.0573 & 0.0015 & 0.0648 & 0.0017 & 0.0499 & 0.0011 & 0.0478 & 0.0014 & 0.0459 \\
20 & 100 & 0.0010 & 0.0353 & 0.0011 & 0.0367 & 0.0008 & 0.0377 & 0.0015 & 0.0340 & 0.0018 & 0.0335 & 0.0014 & 0.0328 \\
20 & 200 & 0.0005 & 0.0261 & 0.0003 & 0.0264 & 0.0003 & 0.0266 & 0.0005 & 0.0258 & 0.0004 & 0.0267 & 0.0002 & 0.0264 \\
20 & 300 & 0.0008 & 0.0230 & 0.0007 & 0.0232 & 0.0006 & 0.0234 & 0.0009 & 0.0227 & 0.0007 & 0.0221 & 0.0003 & 0.0224 \\
50 & 20 & 0.2721 & 8.9585 & 0.0924 & 4.8175 & 0.1307 & 2.4790 & 0.0042 & 0.0859 & 0.0001 & 0.0644 & 0.0001 & 0.0644 \\
50 & 50 & 0.0019 & 0.0335 & 0.0019 & 0.0375 & 0.0028 & 0.0427 & 0.0023 & 0.0324 & 0.0015 & 0.0317 & 0.0006 & 0.0312 \\
50 & 100 & 0.0012 & 0.0221 & 0.0011 & 0.0230 & 0.0011 & 0.0241 & 0.0009 & 0.0217 & 0.0010 & 0.0219 & 0.0009 & 0.0221 \\
50 & 200 & 0.0003 & 0.0152 & 0.0003 & 0.0154 & 0.0003 & 0.0156 & 0.0003 & 0.0150 & 0.0002 & 0.0153 & 0.0003 & 0.0155 \\
50 & 300 & 0.0002 & 0.0132 & 0.0001 & 0.0133 & 0.0002 & 0.0134 & 0.0004 & 0.0131 & 0.0002 & 0.0129 & 0.0004 & 0.0135 \\
100 & 20 & 0.0034 & 3.6462 & 0.0082 & 1.2543 & 0.0746 & 3.6349 & 0.0054 & 0.0584 & 0.0031 & 0.0461 & 0.0031 & 0.0461 \\
100 & 50 & 0.0009 & 0.0230 & 0.0017 & 0.0247 & 0.0015 & 0.0269 & 0.0007 & 0.0217 & 0.0008 & 0.0210 & 0.0010 & 0.0212 \\
100 & 100 & 0.0004 & 0.0146 & 0.0004 & 0.0153 & 0.0003 & 0.0157 & 0.0006 & 0.0142 & 0.0004 & 0.0144 & 0.0003 & 0.0146 \\
100 & 200 & 0.0003 & 0.0102 & 0.0003 & 0.0102 & 0.0003 & 0.0103 & 0.0002 & 0.0101 & 0.0002 & 0.0107 & 0.0003 & 0.0106 \\
100 & 300 & 0.0004 & 0.0089 & 0.0004 & 0.0089 & 0.0004 & 0.0090 & 0.0004 & 0.0088 & 0.0004 & 0.0087 & 0.0004 & 0.0092 \\
200 & 20 & 0.0274 & 1.5439 & 0.0200 & 1.2613 & 0.0042 & 2.0946 & 0.0041 & 0.0568 & 0.0020 & 0.0309 & 0.0020 & 0.0309 \\
200 & 50 & 0.0004 & 0.0158 & 0.0009 & 0.0175 & 0.0014 & 0.0205 & 0.0004 & 0.0152 & 0.0006 & 0.0149 & 0.0006 & 0.0145 \\
200 & 100 & 0.0004 & 0.0109 & 0.0004 & 0.0104 & 0.0004 & 0.0107 & 0.0004 & 0.0101 & 0.0004 & 0.0099 & 0.0002 & 0.0099 \\
200 & 200 & 0.0001 & 0.0070 & 0.0000 & 0.0071 & 0.0000 & 0.0072 & 0.0001 & 0.0070 & 0.0001 & 0.0070 & 0.0002 & 0.0071 \\
200 & 300 & 0.0000 & 0.0062 & 0.0001 & 0.0063 & 0.0001 & 0.0064 & 0.0000 & 0.0062 & 0.0000 & 0.0062 & 0.0001 & 0.0064 \\
300 & 20 & 0.0271 & 0.4488 & 0.1172 & 5.2088 & 0.1167 & 2.3462 & 0.0000 & 0.0326 & 0.0002 & 0.0249 & 0.0002 & 0.0249 \\
300 & 50 & 0.0004 & 0.0126 & 0.0003 & 0.0139 & 0.0005 & 0.0155 & 0.0006 & 0.0122 & 0.0006 & 0.0120 & 0.0006 & 0.0117 \\
300 & 100 & 0.0001 & 0.0086 & 0.0000 & 0.0089 & 0.0000 & 0.0092 & 0.0001 & 0.0085 & 0.0002 & 0.0085 & 0.0002 & 0.0085 \\
300 & 200 & 0.0001 & 0.0058 & 0.0002 & 0.0060 & 0.0001 & 0.0060 & 0.0001 & 0.0058 & 0.0001 & 0.0060 & 0.0001 & 0.0061 \\
300 & 300 & 0.0000 & 0.0051 & 0.0000 & 0.0051 & 0.0000 & 0.0051 & 0.0000 & 0.0051 & 0.0000 & 0.0051 & 0.0000 & 0.0053 \\
\bottomrule
\end{tabular}
\begin{tablenotes}
• \scriptsize \textit{Notes}: See Table 1 for an explanation. Here, $J$ is the number of knots and $T$ is the time dimension.
\end{tablenotes}
\end{threeparttable}
\end{table}
\end{landscape}