Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Robust Inference for Multiple Predictive Regressions with an Application on Bond Risk Premia
\def\spacingset#1{
{#1}} \spacingset{1}
\if11
{
} \fi
\if01
{
center[center omitted — 35 chars of source]
} \fi
abstractWe propose a robust hypothesis testing procedure for the predictability of multiple predictors that could be highly persistent. Our method improves the popular extended instrumental variable (IVX) testing Phillips2013PredictiveRU,Kostakisetal2015 in that, besides addressing the two bias effects found in HosseinkouchackDemetrescu2021, we find and deal with the variance-enlargement effect. We show that two types of higher-order terms induce these distortion effects in the test statistic, leading to significant over-rejection for one-sided tests and tests in multiple predictive regressions. Our improved IVX-based test includes three steps to tackle all the issues above regarding finite sample bias and variance terms. Thus, the test statistics perform well in size control, while its power performance is comparable with the original IVX. Monte Carlo simulations and an empirical study on the predictability of bond risk premia are provided to demonstrate the effectiveness of the newly proposed approach.
{\it Keywords:} Highly Persistent Predictors, Sample Splitting, Bias and Variance, Size Control, Bond Return Predictability.
\spacingset{1.9}
Introduction
Testing the predictability of asset returns is the focal point of financial studies for different stakeholders, including academia, practitioners, and policymakers. However, the commonly used statistical inference suffers from size distortions due to the high persistence of (some) predictors and the contemporaneous correlations between the innovations of predictors and dependent variables CampbellYogo2006. The most popular method to solve this problem is the IVX Phillips2013PredictiveRU,Kostakisetal2015, which aims to solve the over-rejection problem of the test statistics based on OLS.
Unfortunately, the IVX-based test Phillips2013PredictiveRU,Kostakisetal2015 still has size distortion, which is magnified in empirically relevant finite samples, one-sided hypotheses, and multiple predictors cases. In Section (ref), we show explicitly that the root of the oversized IVX statistic is not only the {\bf Deformation Effect} (DE) and the {\bf Displacement Effect} (DiE) HosseinkouchackDemetrescu2021,DemetrescuRodrigues2020 but also the {\bf Variance Enlargement Effect} (VEE). In closely related literature, HosseinkouchackDemetrescu2021 provided a structured approach to control the small sample noncentrality of the IVX $t$-statistic for a given instrumental variable. DemetrescuRodrigues2020 combined the residual-augmented regression and the recursive demean method to remove the size distortion of IVX caused by higher-order terms. However, HosseinkouchackDemetrescu2021 did not consider the inference in multiple predictive
regressions, while the recursive demean technique applied by DemetrescuRodrigues2020 would cause the inference to perform worse than our test in terms of power. Additionally, the inference by HosseinkouchackDemetrescu2021 and DemetrescuRodrigues2020 still suffers size distortion, as evidenced by their simulations. XuGuo2022 applied the Lagrange-multipliers principle to the joint test in multiple predictive models. Nevertheless, XuGuo2022 ignored the higher-order terms that generate the three effects mentioned above, and as a result, their test still suffered size distortions in many cases.
To our knowledge, no existing literature deals with the VEE in IVX test statistics in a single or multiple predictive regression. We show that, as the number of predictors increases, the VEE makes the IVX test statistic grow even bigger compared to its theoretical value. Therefore, ignoring the VEE leads to serious size distortion problems in multiple predictive regressions.
Main Results and Contributions
Our new three-step procedure (detailed in Section (ref)) improves the size control of IVX while keeping its good power performance:
enumerate• We entirely eliminate the DE using a novel sample splitting technique without loss of effective sample size. To do so, we use split (evenly from the middle time point) subsamples and obtain three IVX test statistics based on the full sample and two subsamples. The new test statistic is the weighted average of these three test statistics, effectively circumvents the higher-order term causing the DE.
• We considerably mitigate the DiE, which can cause the finite sample mean of the test statistic to deviate from the expected value of its asymptotic distribution. By subtracting the median of this displacement from the test statistic obtained in the first step, the location of the new test statistic and its asymptotic distribution become very close, thus reducing this displacement effect.
• The VEE can cause the sample standard deviation of the test statistic obtained in the second step to be greater than its asymptotic standard deviation. We divide the test statistic in the previous step by the ratio of the two (sample and the asymptotic) standard deviations. The standard deviation of the new test statistic and its asymptotic distribution are very close, thereby alleviating this effect.
Under some mild conditions detailed in Section (ref), the test statistics converge to a standard normal or Chi-square distribution. For more robust inference in the finite sample, we combine the three steps for size control and the Lagrange-multipliers principle by XuGuo2022. The easy-to-implement procedure is summarized in Algorithm (ref). Intensive simulation studies in the main text and Appendix are provided to demonstrate the effectiveness of the newly proposed approach.
In the empirical research, we revisit the predictability of the bond risk premia. Using forward rates, macroeconomic principal components, and their linear combinations as predictors of bond risk premia, we show that some of the predictors (especially the highly persistent ones) proposed by OLS or IVX Phillips2013PredictiveRU,Kostakisetal2015 exhibit no predicting power on bond excess returns based on our test. Additionally, the effective predictors selected by our proposed test have strong economic implications, and these findings are notably different from the IVX test in Kostakisetal2015. The empirical results confirm that our newly developed test statistics are more reliable tools to test the predictability of financial returns. Our test should be helpful to practitioners.
The main contributions of the paper are summarized below.
enumerate• We address the distortion effects induced by the higher-order terms of IVX test statistics in a unified approach integrating univariate and multivariate models. We discuss the newly discovered VEE and show its accumulative nature in the finite sample as the number of predictors increases, which explains the lingering size distortion issue in the financial return predictability test literature. Our new test eliminates the DE and significantly reduces the DiE and VEE.
• The proposed inference procedure performs very well in size control while keeping the good power performance of the original IVX test. Compared to the literature, the proposed method has unique advantages for one-sided tests in the univariate model and joint and marginal tests in the multivariate model.
• From the empirical study perspective, our test sheds light on the significant variables to forecast one-year-ahead excess returns on bonds. Based on our test, the relatively longer-term forward rates are found to have little predicting power for the short-term bond excess return. The evidence also supports the market segmentation theory, which indicates that the long-term forward rates can not forecast the short-term bond excess return.
Brief Literature Review
There is a rich literature on the over-rejection problem in the inference of predictive regression with highly persistent predictors. Early literature such as Bonferroni's method by CampbellYogo2006 and the conditional likelihood method by JanssonMoreira2006 focused on the inference for the univariate model. Choietal2016 introduced the Cauchy type instrumental variable approach for the test in univariate regression. Zhuetal2014 and Yangetal2021 proposed the weighted empirical likelihood approach with a good size control to test predictability in univariate and bivariate predictive models. To conduct inference in multiple predictive models, Elliott2011 and BreitungDemetrescu2015 proposed the variable addition approach, but the power performance is unsatisfactory. HARVEY2021198 used the quasi-GLS demeaned predictor to develop easy-to-implement tests for return predictability in univariate models, displaying attractive finite sample size control and power. The IVX approach is proposed by MagdalinosPhillips2009 and is also studied by Kostakisetal2015, PhillipsLee2016, Yangetal2019, DemetrescuRodrigues2020, HosseinkouchackDemetrescu2021 and Liu2022RobustIW. In the studies of forecasting bond risk premia, Bauer2018 proposed a bootstrap procedure to test the predictability of bond risk premia with highly persistent predictors. However, it still tends to over-reject and does not offer the asymptotic property of the bootstrap procedure. FAN2021269 proposed an augmented factor model and applied their method to forecast the excess return of U.S. government bonds.
The rest of this paper is organized as follows. Section (ref) introduces the higher-order terms causing the distortion effects of IVX test statistics. Section (ref) provides the procedures to construct the test statistics and presents the asymptotic theories for the proposed estimators and the test statistics. Section (ref) reports the Monte Carlo simulation results. Section (ref) presents the empirical application. Section (ref) concludes the paper. The online Appendix contains detailed proofs of the theoretical results, extra simulations, and additional empirical results.
Throughout this paper, SD, WD refer to strong and weak dependence, respectively. The symbols $\Rightarrow$, $\xrightarrow{d}$ and $\xrightarrow{p}$ are used to represent weak convergence and convergence in distribution and in probability, respectively. All limits are for $T\rightarrow \infty$ in the limit theories, and $O_p (1)$ is stochastically asymptotically bounded while $o_p(1)$ is asymptotically negligible. For any matrix $A$, $A^{1/2}$ is a matrix defined as $A^{1/2} = Q_L \operatorname{diag}( \lambda_1^{1/2}, \lambda_2^{1/2},\cdots,\lambda_K^{1/2})Q_L^\top$, in which $Q_L$ is the matrix whose $i$th column is the eigenvectors of $A$, and $\lambda_i$, $i=1,2,\cdots,K$ are the eigenvalues of $A$. It implies that $A^{1/2}A^{1/2}=A$. Also, we define $A^{-1/2}=(A^{1/2})^{-1}$. $0_K$ is a zero vector with dimension K and $\operatorname{I_K}$ is a $K\times K$ identity matrix.
IVX-based Test and the Higher-order Terms
We are interested in testing the predictability of $y_t$, such as bond or stock returns, in a unified framework with univariate or multivariate predictors. Without loss of generality, we discuss the following multiple predictive regression model
align[align omitted — 70 chars of source]
where $\beta=\left(\beta_1,\beta_2,\cdots,\beta_K \right)^\top$. The model is also studied by Phillips2013PredictiveRU,Kostakisetal2015.
The highly persistent variables $x_{t-1}=\left(x_{1,t-1},x_{2,t-1} \cdots,x_{K,t-1}\right)^\top$ are modeled as follows
align[align omitted — 69 chars of source]
where $x_{i,0}=o_p(\sqrt{T})$, $\rho_i=1+c_i/T^\alpha$, $c=\operatorname{diag}(c_1,c_2,\cdots,c_K)$, $v_t=(v_{1,t},v_{2,t},\cdots,v_{K,t})^\top$, $t=1,\cdots,T$, $i=1,\cdots,K$
and $\alpha=0$ or 1, $K$ is the dimension of $x_{t-1}$. We assume no cointegration relationship exists among $x_t$.
Two types of persistence with different values of $c_i$ and $\alpha$ are considered: (1) Strong Dependence [SD]: $\alpha=1$ and $c_i\leq 0$ is a constant; (2) Weak Dependence [WD]: $\alpha=0$ and $|1+c_i|<1$. \footnote{ \textcolor{black}{We also permit $x_{i,t-1}$ with $i=1,2,\cdots,K$ to have mixed types of persistence; that is to say, $\alpha$ could be a function of $i$. In this setting, the theoretical results for the test statistics are almost the same.}}
Define $\mathcal{F}_t=\sigma\{(u_j,v_j^\top)^\top, j \leq t\}$, which is the information set available at time $t$. Here, following Kostakisetal2015, a general weakly dependent innovation structure of the linear process on $(u_t,v_t^\top)^\top$ in equations ((ref)) and ((ref)) is imposed as follows.
assumptionDefine $\epsilon_t=(u_t, \varepsilon_t^\top)^\top$ as martingale difference sequence\footnote{For SD predictors, $u_t$ could be extended from m.d.s. to the linear process.} with respect to the natural filtration $\mathcal{F}_t$, which satisfies
\begin{align}
E\left(\epsilon_t \epsilon_t^\top|\mathcal{F}_t \right) = \Sigma_t,\,\, a.s.,\quad and \quad \sup_{t} E\|\epsilon_t\|^{4+\nu} <\infty
\end{align}
for some $\nu>0$, where $\Sigma_t$ is a positive definite matrix. Assume $v_t$ is a stationary linear process given by
\[
v_{t}=\sum_{j=0}^\infty F_{j} \varepsilon_{t-j},
\]
Here, $F_{0}=I_K$, \textcolor{black}{$K$ is the dimension of $x_t$} \textcolor{black}{($K\geq 1$)}, and $(F_{j})_{j\geq 0}$ is a sequence of constant matrices such that $\sum_{j=0}^\infty F_{j}$ is of full rank and $\sum_{j=0}^\infty j\|F_{j}\|< \infty$. The process $\epsilon_t$ is strictly stationary ergodic satisfying
\begin{align}
\lim_{m \rightarrow \infty} \left\| Cov\left[vec\left( \epsilon_m\epsilon_m^\top\right),vec\left(\epsilon_0\epsilon_0^\top \right) \right]\right\|=0
\end{align}
remarkAs specified by Kostakisetal2015, the sequence $(u_t)_{t=1}^T$ permits the following GARCH(p,q) representation:
\begin{align}
u_t = h_t \eta_t,\quad
h_t^2= \varphi_0+ \sum_{j=1}^p \varphi_j h_{t-j}^2 +\sum_{j=1}^q \bar{ \varphi}_j u_{t-j}^2
\end{align}
where $(\eta_t)$ is an i.i.d. (0,1) sequence. $\varphi_0$ is a positive constant. And $\varphi_j$ and $\bar{\varphi}_j$ are non-negative constants satisfying $\sum_{j=1}^p \varphi_j +\sum_{j=1}^q \bar{ \varphi}_j<1$.
Under Assumption (ref),
the functional central limit theorem (FCLT) for $(u_t,v_t^\top)^\top$ holds by Phillips1992AsymptoticsFL,
align[align omitted — 267 chars of source]
where $[B_u(r), B_v(r)^\top]^\top$ is a vector of Brownian motions and
$\Sigma_{u v}=\sum_{h=-\infty}^{\infty}E\left( u_{t}v_{t+h}\right)= F_x(1) E\left(u_t \varepsilon_{t}\right)$, $\Sigma_{u u}=E\left(u_{t} u_{t}^\top\right)$ and $\Omega_{vv}= \sum_{h=-\infty}^\infty E(v_tv_{t+h}^\top) = F_x(1)E(v_tv_t^\top)F_x(1)^\top= F_x(1)\Sigma_{vv}F_x(1)^\top$.
Under Assumption (ref), the following result holds by equation ((ref)), Phillips1987 and Kostakisetal2015,
align[align omitted — 120 chars of source]
for SD predictors, where $J_x^c(r)=\int_0^r e^{(r-s)c}d B_v(s)$.
Following Phillips2013PredictiveRU, the standard IVX is set to be less persistent than SD predictors to obtain the asymptotic mixture normal distributed estimator $\hat{\beta}_{ivx}$. Specifically, we have
align[align omitted — 161 chars of source]
where $\bar{z}_{t-1}=z_{t-1}- \frac{1}{T}\sum\limits_{t=1}^{T} z_{t-1}$, $z_t=(z_{1,t},z_{2,t}\cdots,z_{K,t})^\top$ and
align[align omitted — 93 chars of source]
where $\rho_z = 1+ c_z/T^\delta $, $c_z<0$ and $1/2<\delta<1$.
In the following, we expand on Proposition 1 of HosseinkouchackDemetrescu2021, which focuses on univariate models, and point out two higher-order terms in multivariate models. We define the random vector ${t_{ivx}}
\equiv \left( \sum_{t=1}^T \bar{z}_{t-1} \bar{z}_{t-1}^\top \hat{u}_t^2 \right)^{-1/2} \sum_{t=1}^T \bar{z}_{t-1} u_t$ to analyze the DE, DiE and VEE of the original IVX Kostakisetal2015,Phillips2013PredictiveRU in a unified (univariate or multivariate) model,\footnote{When $K=1$, under the null hypothesis $H_0:\beta=0$, ${t_{ivx}}$ becomes the t-test statistic defined by HosseinkouchackDemetrescu2021. The three effects exist in univariate models. } since the original IVX test statistic $ Q_{ivx}={t_{ivx}^\top} H_{ivx}^\top (H_{ivx}H_{ivx}^\top )^{-1} H_{ivx} {t_{ivx}}$ is the function of ${t_{ivx}}$ under the null hypothesis $H_0:R\beta =r_J$, where $R$ is a $J\times K$ predetermined matrix with the rank $J$, $r_J$ is a predetermined vector with dimension $J$, and $H_{ivx} = R\left(\sum\limits_{t=1}^{T} \bar{z}_{t-1}x_{t-1}^\top \right)^{-1}
\left(\sum\limits_{t=1}^{T} \bar{z}_{t-1}\bar{z}_{t-1}^\top \hat{u}_t^2\right)^{1/2}$.
To show the following proposition, we define the normalized instrumental variable $z_{t-1}^*=\Omega_{zz}^{-1/2}z_{t-1}$, where $\Omega_{zz}= \Sigma_{uu} \Omega_{vv}/(-2c_z)$ for SD predictors and $\Omega_{zz}= \operatorname{E}(x_{t-1}x_{t-1}^\top u_t^2)$ for WD predictors.
propUnder Assumption (ref), the following results hold for SD predictors.
\begin{align}
{t_{ivx}}
&= \left( \frac{1}{T^{1+\delta}} \sum_{t=1}^T \bar{z}_{t-1} \bar{z}_{t-1}^\top \hat{u}_t^2 \right)^{-1/2}\frac{1}{T^{1 / 2+\delta / 2}} \sum_{t=1}^T \bar{z}_{t-1} u_t \\
&=\left( \frac{1}{T^{1+\delta}} \sum_{t=1}^T \bar{z}_{t-1}^* (\bar{z}_{t-1}^*)^\top \hat{u}_t^2 \right)^{-1/2}\frac{1}{T^{1 / 2+\delta / 2}} \sum_{t=1}^T \bar{z}_{t-1}^* u_t \nonumber
&={Z_T}+{B_T}+{C_T}+ o_p\left[T^{(\delta- 1)/2}\right]. \nonumber
\end{align}
\begin{align}
&{Z_T} = \frac{1}{T^{1 / 2+\delta / 2}} \sum_{t=1}^T z_{t-1}^* u_t \xrightarrow{P} \operatorname{N}(0_K,\operatorname{I_K} ).\nonumber\\
&{B_T}= \varpi_b
\frac{1}{T^{1 / 2+\delta / 2}} \sum_{t=1}^T z_{t-1}^* u_t\xrightarrow{P} 0.\\
& T^{(1-\delta)/2}\operatorname{E}\left({B_T} \right) \rightarrow
- \frac{K+1}{2}{\rho_{u v^*}}/ \sqrt{-2 c_z},
\end{align}
\begin{align}
T^{(1 -\delta) / 2} {C_T}&=\left( \frac{1}{T^{1+\delta}} \sum_{t=1}^T \bar{z}_{t-1} \bar{z}_{t-1}^\top \hat{u}_t^2 \right)^{-1/2}\frac{1}{T^{1 / 2+\delta }} \sum_{t=1}^T z_{t-1} \frac{1}{\sqrt{T}}\sum_{t=1}^T u_t \\
&\Rightarrow -(-c_z/2)^{-1/2} \Sigma_{vv}^{-1/2} \frac{B_u(1)J_x^c(1)}{\sqrt{\Sigma_{uu} }} \nonumber
\end{align}
where ${\rho_{u v^*}} = \Sigma_{vv}^{-1/2} \Sigma_{vu}/ \sqrt{\Sigma_{uu}}$, $\varpi_b = -\frac{1}{2} \left[\frac{1}{T^{1+\delta }} \sum_{t=1}^T z_{t-1}^* (z_{t-1}^*)^\top u_t^2- \operatorname{I}_K \right]$, $0_K$ is a zero vector with dimension K and $\operatorname{I_K}$ is a $K\times K$ identity matrix.
remarkWhen $K=1$, ${\rho_{u v^*}}$ becomes the correlation coefficient between $u_t$ and $v_t$. Hence Proposition (ref) is the generalization of Proposition 1 of HosseinkouchackDemetrescu2021.
The DE, DiE and VEE
Three critical points are drawn from Proposition (ref). First, similar to HosseinkouchackDemetrescu2021, the higher-order term ${B_T}$ comes from the correlation between the terms $ \sum_{t=1}^T \bar{z}_{t-1} \bar{z}_{t-1}^\top \hat{u}_t^2 $ and $ \sum_{t=1}^T \bar{z}_{t-1} u_t$ in the random vector ${t_{ivx}}$. Meanwhile, the higher-order term ${C_T}$ comes from $\sum\nolimits_{t=1}^T z_{t-1} \sum\nolimits_{t=1}^T u_t$ in ${t_{ivx}}$.
\textcolor{black}{Second, with SD predictors, the higher-order term ${B_T}$ leads to not only the DiE shown in equation ((ref)) but also the VEE shown in equation ((ref)).} The DiE causes the center of the finite sample distribution of ${t_{ivx}}$ to deviate from zero to $T^{-(1-\delta)/2}\lim_{T\rightarrow\infty}T^{(1-\delta)/2}\operatorname{E}\left({B_T} \right)$, \textcolor{black}{which makes the one-sided test of the original IVX to suffer from size distortion HosseinkouchackDemetrescu2021.} The VEE causes the covariance matrix of ${t_{ivx}}$ in the finite sample to be enlarged by the higher-order term $B_T$ so that the test statistics $Q_{ivx}$ is enlarged, which is shown in Theorem (ref). Specifically, equation ((ref)) shows that the covariance matrix of ${t_{ivx}}$ in the finite sample is approximately
align[align omitted — 194 chars of source]
Therefore, the covariance matrix of $(\operatorname{I_K} +\hat{\varpi}_b \hat{\varpi}_b^\top )^{-1/2}{t_{ivx}}$ in the finite sample is closer to $\operatorname{I_K}$ than that of ${t_{ivx}}$, where $\hat{\varpi}_b= -\frac{1}{2} \left[ \sum_{t=1}^T z_{t-1}^{**} (z_{t-1}^{**})^\top u_t^2- \operatorname{I}_K \right]$ and $z_{t-1}^{**} = \left(\sum_{t=1}^T z_{t-1}z_{t-1}^\top \hat{u}_t^2 \right)^{-1/2}z_{t-1}$. As a result, the finite sample distribution of the new IVX test statistic $$ \tilde{Q}_{ivx}={t_{ivx}^\top} H_{ivx}^\top \left[H_{ivx} (\operatorname{I_K} +\hat{\varpi}_b \hat{\varpi}_b^\top) H_{ivx}^\top \right]^{-1} H_{ivx} {t_{ivx}}$$ based on $(\operatorname{I_K} +\hat{\varpi}_b \hat{\varpi}_b^\top )^{-1/2}{t_{ivx}}$ is much closer to $\chi_J^2$ than that of the original IVX test statistic $Q_{ivx}$ based on ${t_{ivx}}$.
The following theorem shows that $Q_{ivx} - \tilde{Q}_{ivx} $ is positive, and VEE aggravates as K increases.
thmUnder Assumption (ref) and the null hypothesis $H_0:R\beta =r_J$, it follows that
\begin{align*}
Q_{ivx} - \tilde{Q}_{ivx} & = {t_{ivx}^\top} H_{ivx}^\top \left(H_{ivx} H_{ivx}^\top \right)^{-1}
H_{ivx} \hat{\varpi}_b \hat{\varpi}_b^\top H_{ivx}^\top
\left[H_{ivx} (\operatorname{I_K} +\hat{\varpi}_b \hat{\varpi}_b^\top) H_{ivx}^\top \right]^{-1} H_{ivx} {t_{ivx}} \\
&= {t_{ivx}^\top} H_{ivx}^\top \left(H_{ivx} H_{ivx}^\top \right)^{-1}
H_{ivx} \hat{\varpi}_b \hat{\varpi}_b^\top H_{ivx}^\top
\left(H_{ivx} H_{ivx}^\top \right)^{-1} H_{ivx} {t_{ivx}} + o_p\left[T^{-(1-\delta)/2}\right] \\
&= \sum_{j=1}^K (H_{j}^{ivx})^2 + o_p\left[T^{-(1-\delta)/2}\right]>0,
\end{align*}
where $H^{ivx}=(H_1^{ivx},H_2^{ivx},\cdots,H_K^{ivx})^\top = \hat{\varpi}_b^\top H_{ivx}^\top
\left(H_{ivx} H_{ivx}^\top \right)^{-1} H_{ivx} {t_{ivx}}$.
By Theorem (ref), as the number of predictors $K$ grows, VEE becomes more severe since $Q_{ivx} - \tilde{Q}_{ivx} $ grows bigger. This causes the original IVX to suffer severe size distortion in multiple predictive models.
Third, equation ((ref)) shows that the higher-order term ${C_T}$ leads to the DE, meaning that the finite sample distribution of ${t_{ivx}}$ deviates significantly from the normal distribution.
In summary, the DiE, VEE, and DE induced by the higher-order terms ${B_T}$ and ${C_T}$ in Proposition (ref) explain why the IVX test Kostakisetal2015,Phillips2013PredictiveRU suffers size distortions for one-sided tests and tests in multivariate predictive regression.
A straight-forward idea to eliminate the effects of ${B_T}$ and ${C_T}$ is to subtract their consistent estimators from ${t_{ivx}}$.
However, it is not an easy task to eliminate the effect of ${C_T}$ since it arises from $\sum \nolimits_{t=1}^T z_{t-1}\sum \nolimits_{t=1}^T u_t$ which could not be estimated consistently.
The Improved IVX Test
A Sample Splitting Procedure to Eliminate the DE
To eliminate the size distortion induced by the higher-order term ${C_T}$ in Proposition (ref), we apply a sample splitting procedure to remove the term $\sum\nolimits_{t=1}^{T} z_{t-1} \sum\nolimits_{t=1}^{T} u_t$ in test statistics, which is the source of DE. Intuitively, we split the full sample into two subsamples and construct the IVX estimators based on the full sample and the two subsamples. By setting appropriate weights, the new estimator, which is the weighted sum of the three estimators (full sample, two subsamples), does not contain $\sum\nolimits_{t=1}^{T} z_{t-1} \sum\nolimits_{t=1}^{T} u_t$. As a result, the DE is eliminated by the following steps.
We split the full sample into two subsamples $\{(y_t,x_{t-1})\}_{t=1}^{T_0}$ and $\{(y_t,x_{t-1})\}_{t=T_0}^{T}$, where $T_0=\lfloor \lambda T\rfloor$ and $0<\lambda<1$ is a constant. In practice, we suggest evenly split to subsamples such that $\lambda=0.5$, whose reason is further specified by Remark (ref) in subsection (ref).
First, we define the IVX estimators based on two subsamples as follows.
align[align omitted — 326 chars of source]
where $\bar{z}_{t-1}^a= z_{t-1}- \frac{1}{T_0}\sum\nolimits_{t=1}^{T_0} z_{t-1}$, $\bar{z}_{t-1}^b= z_{t-1}- \frac{1}{T-T_0}\sum\nolimits_{t=T_0+1}^{T} z_{t-1}$.
Second, we apply equations ((ref)), ((ref)), ((ref)) and ((ref)) to obtain the following equations.
small\begin{align}
\left[\sum\limits_{t=1}^{T} \left( z_{t-1}- \frac{1}{T}\sum\limits_{t=1}^{T} z_{t-1} \right) x_{t-1}^\top \right](\hat{\beta}_{ivx}-\beta) & = \sum\limits_{t=1}^{T} z_{t-1}u_t - \frac{1}{T}\sum\limits_{t=1}^{T} z_{t-1} \sum\limits_{t=1}^{T} u_t , \\
\left[\sum\limits_{t=1}^{T_0} \left( z_{t-1}- \frac{1}{T_0}\sum\limits_{t=1}^{T_0} z_{t-1} \right) x_{t-1}^\top \right](\hat{\beta}_a-\beta) &= \sum\limits_{t=1}^{T_0} z_{t-1} u_t- \frac{1}{T_0}\sum\limits_{t=1}^{T_0} z_{t-1} \sum\limits_{t=1}^{T_0} u_t ,\\
\left[\sum\limits_{t=T_0+1}^{T} \left( z_{t-1}- \frac{\sum_{t=T_0+1}^{T} z_{t-1} }{T-T_0} \right) x_{t-1}^\top \right] (\hat{\beta}_b-\beta) & = \sum\limits_{t=T_0+1}^{T} z_{t-1} u_t - \frac{\sum\limits_{t=T_0+1}^{T} z_{t-1}}{T-T_0} \sum\limits_{t=T_0+1}^{T} u_t.
\end{align}
By equations ((ref)) and ((ref)), the following equations hold.
footnotesize\begin{align}
S_a\left[\sum\limits_{t=1}^{T_0} \left( z_{t-1}- \frac{1}{T_0}\sum\limits_{t=1}^{T_0} z_{t-1} \right) {x}_{t-1}^\top \right](\hat{\beta}_a-\beta) &= S_a \sum\limits_{t=1}^{T_0} z_{t-1} u_t- \frac{1}{T}\sum\limits_{t=1}^{T} z_{t-1} \sum\limits_{t=1}^{T_0} u_t ,\\
S_b\left[\sum\limits_{t=T_0+1}^{T} \left( z_{t-1}- \frac{1}{T-T_0}\sum\limits_{t=T_0+1}^{T} z_{t-1} \right) {x}_{t-1}^\top \right] (\hat{\beta}_b-\beta)& = S_b \sum\limits_{t=T_0+1}^{T} z_{t-1} u_t - \frac{1}{T}\sum\limits_{t=1}^{T} z_{t-1} \sum\limits_{t=T_0+1}^{T} u_t,
\end{align}
where
align[align omitted — 502 chars of source]
By subtracting the sum of equations ((ref)) and ((ref)) from equation ((ref)), \textcolor{black}{the term $\sum\nolimits_{t=1}^{T} z_{t-1} \sum\nolimits_{t=1}^{T} u_t$ of $\hat{\beta}_{ivx}$ in equation ((ref)) is removed as follows. }
align[align omitted — 314 chars of source]
where $W_1=\sum\nolimits_{t=1}^{T} \bar{z}_{t-1} x_{t-1}^\top $, $W_2=S_a \sum\nolimits_{t=1}^{T_0} \bar{z}_{t-1}^a x_{t-1}^\top$,
$W_3=S_b \sum\nolimits_{t=T_0+1}^{T} \bar{z}_{t-1}^b x_{t-1}^\top$ and
align[align omitted — 192 chars of source]
Third, we could utilize another instrumental variable estimator with the IV $\tilde{z}_{t-1}$. Define the IV estimator $\hat{\beta}_l $ as follows.
align[align omitted — 132 chars of source]
By equations ((ref)) and ((ref)) and that $W_1-W_2-W_3 = \sum\limits_{t=1}^{T} \tilde{z}_{t-1}x_{t-1}$, it follows that
align[align omitted — 163 chars of source]
The key and desirable property of the new instrumental variable $\tilde{z}_{t-1}$ is
align[align omitted — 74 chars of source]
Thus $\hat{\beta}_l$ could be deemed the instrumental variable estimator using $\tilde{z}_{t-1}$ as IV,
align*[align* omitted — 647 chars of source]
Therefore, the higher-order term ${C_T}$ vanishes,
since the term $\sum\limits_{t=1}^{T} z_{t-1} \sum\limits_{t=1}^{T} u_t$ disappears in $\hat{\beta}_l- \beta$.
thmUnder Assumption (ref), for SD and WD predictors, it follows that
$$
D_T(\hat{\beta}_l- \beta)\xrightarrow{d} \operatorname{MN}\left[0,\Sigma_{zx}^{-1}\Sigma_{zz}\left(\Sigma_{zx}^{-1}\right)^\top \right],
$$
where $D_T=T^{(1+\delta)/2}$ for SD predictors and $D_T=\sqrt{T}$ for WD predictors and
$\Sigma_{zz} = \lambda\left(\operatorname{I_K} - \tilde{S}_a\right)\Omega_{zz}\left(\operatorname{I_K} - \tilde{S}_a\right)^\top +(1-\lambda)\left(1-\tilde{S}_b\right)\Omega_{zz} \left(\operatorname{I_K} - \tilde{S}_b\right)^\top
$, where \quad $\Sigma_{zx}= -c_z^{-1}(1-\tilde{S}_a) \int_0^{\lambda} dJ_x^c(r)\, J_x^c(r)^\top
-c_z^{-1}(1-\tilde{S}_b)\int_{\lambda}^{1} dJ_x^c(r) \, J_x^c(r)^\top -c_z^{-1}\left[ 1- \lambda\tilde{S}_a -(1-\lambda)\tilde{S}_b\right]\operatorname{E}(v_tv_t^\top) $ for SD predictors and
$\Sigma_{zx}=\left[1- \lambda\tilde{S}_a -(1-\lambda)\tilde{S}_b \right] \operatorname{E}\left( x_{t-1}x_{t-1}^\top\right)$ for WD predictors
and
\begin{align}
&S_a \Rightarrow \tilde{S}_a =
\begin{cases}
J_x^c(1) J_x^c(\lambda)^\top \left[ J_x^c(\lambda)^\top J_x^c(\lambda) \right]^{-1},\quad SD; \\
B_v(1) B_v(\lambda)^\top \left[ B_v(\lambda)^\top B_v(\lambda) \right]^{-1},\quad WD;
\end{cases}\\
&S_b \Rightarrow \tilde{S}_b =
\begin{cases}
J_x^c(1) \left[J_x^c(1) - J_x^c(\lambda) \right]^\top \left\{ \left[J_x^c(1) - J_x^c(\lambda) \right]^\top \left[J_x^c(1) - J_x^c(\lambda) \right] \right\}^{-1},\quad SD; \\
B_v(1) \left[B_v(1) - B_v(\lambda) \right]^\top \left\{ \left[B_v(1) - B_v(\lambda) \right]^\top \left[B_v(1) - B_v(\lambda) \right] \right\}^{-1},\quad WD.
\end{cases}\nonumber
\end{align}
Then the test statistic ${Q_l}$ is constructed for the null hypothesis $H_0:R\beta =r_J$.
equation[equation omitted — 201 chars of source]
where $\operatorname{\widehat{Avar}}(\hat{\beta}_l )= H_lH_l^\top$
\footnote{We use the heteroskedasticity consistent estimator similar to that of MacKinnon1985SomeHC.} and
$H_l = \left( \sum_{t=1}^T \tilde{z}_{t-1} x_{t-1}^\top \right)^{-1}\left(\frac{T}{T-2K-1} \sum_{t=1}^T \tilde{z}_{t-1} \tilde{z}_{t-1}^\top \hat{u}_t^2 \right)^{1/2}$.
remarkFollowing the literature, $u_t$ is estimated by OLS, which has a smaller variance than that of the IV estimator for $u_t$. Besides the higher-order terms $B_T$ and $C_T$, XuGuo2022 argue that the estimation error of $u_t$ also causes the size distortion. Moreover, the variance of unconstrained OLS estimator $u_t$ increases with $K$. To avoid size distortion induced by the uncertainty of the estimators of $u_t$, we apply the Lagrange-multipliers principle XuGuo2022 to obtain the constrained OLS estimator $\hat{u}_t$ as follows.
\begin{align*}
(\hat{\mu}_s,\hat{\beta}_s)^\top = \arg\, \min_{\mu,\beta} \left( y_t - \mu - x_{t-1}^\top \beta\right)^2,
s.t. \quad R\beta =r_J,
\end{align*}
and $\hat{u}_t= y_t - \hat{\mu}_s - x_{t-1}^\top \hat{\beta}_s$. Since the procedure to obtain $\hat{u}_t$ uses the information of the null hypothesis $H_0:R\beta =r_J$, the estimator $\hat{u}_t$ is more efficient than the unconstrained OLS estimator. As a result, the size performance of ${Q_l}$ is much better than the case in which the Lagrange-multipliers principle is not applied.
Moreover, we construct the t-test statistic ${Q_l^t}$ when $J=1$ for right side test $H_0:\beta_i=0$ vs $H_a:\beta_i>0$ and left side test $H_0:\beta_i=0$ vs $H_a:\beta_i<0$.
align[align omitted — 131 chars of source]
Note that $Q_l=(Q_l^t)^2$ when $J=1$.
thmUnder Assumption (ref) and the null hypothesis $H_0:R\beta =r_J$, one can show that the limiting distribution of the test statistics ${Q_l^t}$ and ${Q_l}$ are the standard normal and the $\chi^2$-distribution with $J$ degrees of freedom, respectively.
Then, the local power of ${Q_l^t}$ with $J=1$ is shown as follows.
thmUnder Assumption (ref), for SD and WD predictors with $J=1$, it follows that
$$
{Q_l^t} \xrightarrow{d} \operatorname{N}(0,1)+ R\Sigma_{zx}^{-1}\Sigma_{zz}\left(\Sigma_{zx}^{-1}\right)^\top R^\top b,
$$
where $R\beta - r_J=b_\beta/D_T$ and $b_\beta$ is a constant.
Theorem (ref) shows that the convergence rate to the local power of ${Q_l^t}$ is the same as IVX Kostakisetal2015,Phillips2013PredictiveRU and the local power of ${Q_l^t}$ is comparable with that of IVX Kostakisetal2015,Phillips2013PredictiveRU. A similar conclusion could be drawn for the test statistic ${Q_l}$.
Methods to Reduce DiE and VEE
The asymptotically $\chi^2$-distributed test statistic ${Q_l}$ still suffers size distortion due to the DiE and VEE in ${Q_l}$. To analyze these effects, we define the random vector ${t_l}$ as follows.
$${t_l} \equiv \left(\frac{T}{T-2K-1} \frac{1}{T^{1+\delta}} \sum_{t=1}^T \tilde{z}_{t-1} \tilde{z}_{t-1}^\top \hat{u}_t^2 \right)^{-1/2}\frac{1}{T^{1 / 2+\delta / 2}} \sum_{t=1}^T \tilde{z}_{t-1} u_t. $$
Note that
align[align omitted — 234 chars of source]
Thus by equations ((ref)) and ((ref)), $ Q_l={{t_l}^\top} H_l^\top (H_l H_l^\top )^{-1} H_l {t_l}$ is the function of ${t_l}$ under the null hypothesis $H_0:R\beta =r_J$. Therefore, we can analyze and reduce DiE and VEE in ${Q_l}$ by analyzing and reducing those in ${t_l}$.
To analyze the higher-order term of \textcolor{black}{ ${t_l}$} , we define \textcolor{black}{the random vector} based on the first subsample $\{(y_t,x_{t-1})\}_{t=1}^{T_0}$ as
$$
{t_a^p} =\left( \frac{1}{T^{1+\delta}} \sum_{t=1}^{T_0} z_{t-1} z_{t-1}^\top \hat{u}_t^2 \right)^{-1/2}\frac{1}{T^{1 / 2+\delta / 2}} \sum_{t=1}^{T_0} z_{t-1} u_t
$$
and \textcolor{black}{the random vector} based on the second subsample $\{(y_t,x_{t-1})\}_{t=T_0+1}^{T}$ as
$$
{t_b^p} = \left( \frac{1}{T^{1+\delta}} \sum_{t=T_0+1}^{T} z_{t-1} z_{t-1}^\top \hat{u}_t^2 \right)^{-1/2}\frac{1}{T^{1 / 2+\delta / 2}} \sum_{t=T_0+1}^{T} z_{t-1} u_t.
$$
The random vectors ${t_a^p}$ and ${t_b^p}$ are used to show the property of the higher-order terms of ${t_l}$. Since ${t_a^p}$ does not contain $\sum\nolimits_{t=1}^T z_{t-1} \sum\nolimits_{t=1}^T u_t$, it only contains the higher-order terms ${B_T^{p_a}}$ similar to ${B_T}$. Therefore,
align[align omitted — 101 chars of source]
where
$
{Z_T^{p_a}} \xrightarrow{d} \operatorname{N}(0_K,\operatorname{I_K}),\quad \sqrt{T_0(1-\rho_z^2)} \operatorname{E}({B_T^{p_a}})\xrightarrow{P} -{\rho_{u v^*}},\quad {B_T^{p_a}}\xrightarrow{P}0,
$
and
align[align omitted — 101 chars of source]
where
$
{Z_T^{p_b}} \xrightarrow{d} \operatorname{N}(0_K,1),\quad \sqrt{(T-T_0)(1-\rho_z^2)} \operatorname{E}({B_T^{p_b}})\xrightarrow{P} -{\rho_{u v^*}},\quad {B_T^{p_b}}\xrightarrow{P}0
$.
Also, notice that under the null hypothesis $H_0:R\beta =r_J$, it follows that
align[align omitted — 64 chars of source]
where
$
W_a = \left( \frac{1}{T^{1+\delta}} \sum_{t=1}^{T} \tilde{z}_{t-1} \tilde{z}_{t-1}^\top \hat{u}_t^2 \right)^{-1/2}
(\operatorname{I_K}-S_a)
\left( \frac{1}{T^{1+\delta}} \sum_{t=1}^{T_0} z_{t-1} z_{t-1}^\top \hat{u}_t^2 \right)^{1/2}
$
and
$$
W_b = \left( \frac{1}{T^{1+\delta}} \sum_{t=1}^{T} \tilde{z}_{t-1} \tilde{z}_{t-1}^\top \hat{u}_t^2 \right)^{-1/2}
(\operatorname{I_K}-S_b)
\left( \frac{1}{T^{1+\delta}} \sum_{t=T_0+1}^{T} z_{t-1} z_{t-1}^\top \hat{u}_t^2 \right)^{1/2}.
$$
We can apply equation ((ref)) to construct the higher-order terms of ${t_l}$ by the higher-order terms of ${t_a^p}$ and ${t_b^p}$ shown in equations ((ref)) and ((ref)).
Define the normalized instrumental variable $\tilde{z}_{t-1}^*= (\Sigma_{zz})^{-1/2}\tilde{z}_{t-1}$. Then the following proposition holds.
propUnder Assumption (ref), for SD predictors, the following equation holds as $T \rightarrow \infty$,
$$
{t_l}={Z_T^l}+{B_T^l}+o_p\left(T^{\delta / 2-1 / 2}\right),
$$
where
$
{Z_T^l} = (\Sigma_{zz})^{-1/2}\sum\nolimits_{t=1}^T \tilde{z}_{t-1} u_t \stackrel{d}{\rightarrow} \operatorname{N}(0_K,\operatorname{I_K}), \quad \text {and }
$
${B_T^l }\rightarrow 0. \quad$ And
\begin{align}
{B_T^l} &= \varpi_l {Z_T^l}\\
& = W_a {B_T^{p_a}}+ W_b {B_T^{p_b}} + o_p\left[T^{-(1-\delta)/2}\right]
\end{align}
where
$
\varpi_l = -\frac{1}{2}\left\{ \sum\nolimits_{t=1}^T \tilde{z}_{t-1}^* (\tilde{z}_{t-1}^*)^\top \hat{u}_t^2 -\operatorname{I_K}\right\}
$
and
$$
T^{(1 -\delta) / 2} {B_T^l} ={R_T^l} + W_l\, T^{-(1 -\delta) / 2}\operatorname{plim}_{T\rightarrow \infty} \; T^{(1 -\delta) / 2} \operatorname{E}\left({B_T}\right),
$$
where ${R_T^l}= W_a {R_{1,T}^l}+W_b{ R_{2,T}^l}$, $ {R_{1,T}^l}=T^{(1 -\delta) / 2}\left[ {B_T^{p_a}}- \operatorname{E}({B_T^{p_a}}) \right]$ and $ {R_{2,T}^l}=T^{(1 -\delta) / 2}\left[ {B_T^{p_b}}- \operatorname{E}({B_T^{p_b}}) \right]$ satisfying
$
\operatorname{E}\left( {R_{1,T}^l} \right) = \operatorname{E}\left( {R_{2,T}^l} \right) =0
$
and
$
W_l = W_a/\sqrt{\lambda} + W_b/\sqrt{1-\lambda}
$.
remarkThe term ${R_T^l}$ could be regarded as the “residual term" of ${B_T^l}$ since it is the weighted sum of two terms whose expectation is zero. Therefore, $W_l\, T^{-(1 -\delta) / 2} \operatorname{plim}_{T\rightarrow \infty} \; T^{(1 -\delta) / 2} \operatorname{E}\left({B_T}\right)$ is the main source of DiE induced by ${B_T^l}$ rather than the term ${R_T^l}$.
Proposition (ref) shows that the higher-order term ${B_T^l}$ leads to not only the DiE but also the VEE in multivariate models. Due to the DiE, the center of the distribution of ${t_l}$ in the finite sample deviates from zero vector, while the covariance matrix of ${t_l}$ in the finite sample is closer to $\operatorname{I_K}+\varpi_l\varpi_l^\top$ rather than $\operatorname{I_K}$.
First, we try to reduce the DiE.
Although Remark (ref) shows that the higher order term $B_T^l$'s term
$W_l \,T^{-(1 -\delta) / 2} \operatorname{plim}\nolimits_{T\rightarrow \infty} \; T^{(1 -\delta) / 2} \operatorname{E} ({B_T})$ is the main source of size distortion, the operation that subtracting ${t_l}$ by
$W_l\,T^{-(1 -\delta) / 2} \operatorname{plim}\nolimits_{T\rightarrow \infty} \; T^{(1 -\delta) / 2} \operatorname{E}\left({B_T}\right)$
could not eliminate the size distortion caused by the higher order term ${B_T^l}$. The main reason is as follows. $W_l$ converges to a function of Brownian Motion rather than a constant. So the covariance matrix of the random vector $W_l \operatorname{plim}_{T\rightarrow \infty} \; T^{(1 -\delta) / 2} \operatorname{E}\left({B_T}\right)$ does not vanish even when the sample size goes to infinity. As a result, the finite sample covariance matrix of $W_l \, T^{-(1 -\delta) / 2}\operatorname{plim}_{T\rightarrow \infty} \; T^{(1 -\delta) / 2} \operatorname{E}\left({B_T}\right)$ is non-negligible since its rate $T^{-(1 -\delta) / 2}$ converging to zero is very slow. This non-negligible finite sample covariance matrix of $W_l \, T^{-(1 -\delta) / 2}\operatorname{plim}_{T\rightarrow \infty} \; T^{(1 -\delta) / 2} \operatorname{E}\left({B_T}\right)$ leads that the finite sample covariance matrix of ${t_l}-W_l\,T^{(1 -\delta) / 2} \operatorname{plim}_{T\rightarrow \infty} \; T^{(1 -\delta) / 2} \operatorname{E}\left({B_T}\right)$ significantly deviates from its asymptotic covariance matrix $\operatorname{I_K}$. This deviation causes the size distortion of the test statistic constructed by ${t_l}-W_l T^{-(1 -\delta) / 2} \operatorname{plim}_{T\rightarrow \infty} \; T^{(1 -\delta) / 2} \operatorname{E}\left({B_T}\right)$.
To reduce the size distortion arise from $W_l \operatorname{plim}_{T\rightarrow \infty} \; T^{(1 -\delta) / 2} \operatorname{E}\left({B_T}\right)$, we subtract ${t_l}$ by the estimator of $ T^{-(1 -\delta) / 2} \operatorname{plim}_{T\rightarrow \infty} \; T^{(1 -\delta) / 2} \operatorname{E}\left({B_T}\right)$, which is equal to $ - T^{-(1 -\delta) / 2} \frac{K+1}{2}{\hat{\rho}_{u v^*}}/ \sqrt{-2 c_z}$.
We construct the random vector ${\tilde{t}_n}$ as follows.
align[align omitted — 131 chars of source]
where $ {\hat{\rho}_{u v^*}}=\hat{\Sigma}_{vv}^{-1/2} \hat{\Sigma}_{vu} \hat{\Sigma}_{uu}^{-1/2}$ is the consistent estimator of $\rho_{u v^*}$. Note that $ \hat{\Sigma}_{vu}$, $\hat{\Sigma}_{vv}^{-1/2}$ and $\hat{\Sigma}_{uu}$ are the consistent estimators of $\Sigma_{vu}$, $\Sigma_{vv}^{-1/2}$ and $\Sigma_{uu}$, respectively.
In the following, we illustrate why the DiE is reduced in ${\tilde{t}_n}$ in the following aspects. First, by equation ((ref)) and Proposition (ref), it is straightforward that
align[align omitted — 99 chars of source]
where ${B_T^n}= {B_T^l}- T^{-(1 -\delta) / 2} \operatorname{plim}_{T\rightarrow \infty} \; T^{(1 -\delta) / 2} \operatorname{E}\left({B_T}\right) \xrightarrow{P} 0$ and thus
align[align omitted — 196 chars of source]
Second, we show that the distorting effect of ${B_T^n}$ for the size is much smaller than that of ${B_T^l}$ in the following aspects. By Proposition (ref), it follows that $W_l-\left[ \lambda(1-\lambda) \right]^{-1/2}\operatorname{I_K}$ is negative definite matrix and $\left[ \lambda(1-\lambda) \right]^{-1/2}\geq 2$. Therefore, the size distortion induced by $(W_l-\operatorname{I_K}) \operatorname{plim}_{T\rightarrow \infty} \; T^{(1 -\delta) / 2} \operatorname{E}\left({B_T}\right)$ is smaller than that of $W_l \operatorname{plim}_{T\rightarrow \infty} \; T^{(1 -\delta) / 2} \operatorname{E}\left({B_T}\right)$. As a result, the size distortion induced by ${B_T^n}$ is smaller than ${B_T^l}$.
remarkThe intuition of the size distortion induced by ${B_T^n}$ is smaller than that of ${B_T^l}$ can be easier to see using an example of a univariate model. First, $\operatorname{P}(W_l-1>0)\rightarrow 1/2$ when $K=1$, which implies that the median of $W_l-1$ is approximately 0. Moreover, $(W_l-1)^2
\in \left[0,1+ 2\lambda^{-1/2}(1-\lambda)^{-1/2} +\lambda^{-1}(1-\lambda)^{-1}\right]$ such that $W_l-1$ is bounded. It implies that ${B_T^n}$ is much smaller than ${B_T^l}$ generally, which means DiE is reduced significantly in ${\tilde{t}_n}$.
remark\textcolor{black}{Following Remark (ref), $(W_l-1)^2$ reaches the minimum upper bound when $\lambda=0.5$. So we set $\lambda=0.5$ to minimize the DiE of ${\tilde{t}_n}$. Simulation studies in Section (ref) and the appendix support the choice of $\lambda=0.5$ in the finite sample.}
However, VEE still exists in ${\tilde{t}_n}$ since
align[align omitted — 319 chars of source]
Equation ((ref)) holds since $\operatorname{Var}( {B_T^n})= \operatorname{Var}\left[{B_T^l}- T^{-(1 -\delta) / 2} \operatorname{plim}_{T\rightarrow \infty} \; T^{(1 -\delta) / 2} \operatorname{E}\left({B_T}\right) \right]= \operatorname{Var}\left({B_T^l}\right) + o(1)$ induced by equation ((ref)).
To eliminate the VEE, the random vector ${\tilde{t}_m}$ is constructed as follows.
align[align omitted — 205 chars of source]
where $\hat{\varpi}_l = -\frac{1}{2}\left\{ \sum\nolimits_{t=1}^T \tilde{z}_{t-1}^{**} (\tilde{z}_{t-1}^{**} )^\top \hat{u}_t^2 -\operatorname{I_K}\right\}$,
$\tilde{z}_{t-1}^{**} =({\hat{\Sigma}}_{zz})^{-1/2}\tilde{z}_{t-1}$ and
$$
{\hat{\Sigma}}_{zz} = \left(\operatorname{I_K} - S_a\right)\frac{\sum_{t=1}^{T_0}\Delta x_{t-1}\Delta x_{t-1}^\top }{1-\rho_z^2}\left(\operatorname{I_K} - S_a\right)^\top + \left(1-S_b\right)\frac{\sum_{t=T_0+1}^T \Delta x_{t-1}\Delta x_{t-1}^\top }{1-\rho_z^2} \left(\operatorname{I_K} - S_b\right)^\top.
$$
By equations ((ref)), ((ref)), ((ref)) and ((ref)), the estimator for $\beta$ is defined as
align[align omitted — 253 chars of source]
where and
${\tilde{B}_m}\equiv \left( \sum_{t=1}^T \tilde{z}_{t-1} x_{t-1}^\top \right)^{-1}\left( \sum_{t=1}^T \tilde{z}_{t-1} \tilde{z}_{t-1}^\top \hat{u}_t^2 \right)^{1/2}\left(\operatorname{I_K}+ \hat{\varpi}_l \hat{\varpi}_l^\top \right)^{1/2}.$
We apply the following estimator for the asymptotic distribution of $\hat{\beta}_m $ to reduce VEE of IVX.
align[align omitted — 147 chars of source]
Then we define the test statistics for $H_0:R\beta =r_J$.
align[align omitted — 199 chars of source]
propUnder Assumption (ref) and the null hypothesis $H_0:R\beta =r_J$, for SD predictors, it follows that
\begin{align*}
Q_l - {\tilde{Q}_m}
&= {t_l^\top} H_l^\top R^\top \left(R H_l H_l^\top R^\top \right)^{-1}
R H_l \hat{\varpi}_l \hat{\varpi}_l^\top H_l^\top R^\top
\left(R H_l H_l^\top R^\top \right)^{-1} R H_l {t_l} + o_p\left[T^{-(1-\delta)/2}\right]>0.
\end{align*}
Thus far, the DE is eliminated for SD predictors in the finite sample, and the DiE and VEE are reduced significantly in ${\tilde{Q}_m}$.
On the other hand, the asymptotic property of $\tilde{\beta}_m $ and ${\tilde{Q}_m}$ is the same as that of $\tilde{\beta}_l$ and ${\tilde{Q}_l}$ by equation ((ref)), which is shown as follows.
thmUnder Assumption (ref), for SD and WD predictors, it follows that
$$
D_T(\tilde{\beta}_m- \beta)=D_T(\hat{\beta}_l- \beta)+o_p(1)\xrightarrow{d}
\operatorname{MN}\left[0,\Sigma_{zx}^{-1}\Sigma_{zz}\left(\Sigma_{zx}^{-1}\right)^\top \right].
$$
By Theorem (ref), the asymptotic distribution of the test statistic ${\tilde{Q}_m}$ is the same as that of ${Q_l}$, which is shown as follows.
thmUnder Assumption (ref) and the null hypothesis $H_0:R\beta =r_J$, one can show that the limiting distribution of ${\tilde{Q}_m}$ is the $\chi^2$-distribution with $J$ degrees of freedom.
Although ${\tilde{Q}_m}$ performs well in terms of size with SD predictors in the finite sample, it suffers size distortion with WD predictors since the terms to reduce the DiE and VEE for the case with SD predictors are redundant terms with WD predictors. To avoid size distortion with WD predictors, we construct the test statistics by the following steps. First, the weight is introduced ${W_z} =\operatorname{diag}(w_{z_1},w_{z_2},\cdots,w_{z_K})$ such that $w_{z_i} \equiv exp\left[-T(1-\hat{\rho}_i)^2 /K \right]=1+O_p(T^{-1})$ for SD predictors\footnote{We put $K$ in $w_i$ such that the weight $w_{z_i}$ is closer to one in the finite sample with more predictors. Therefore, the size control is improved since the DiE and the VEE are reduced sufficiently.} while $w_{z_i} \equiv exp\left[-T(1-\hat{\rho}_i)^2 /K \right]= o_p(T^{-1})$ for WD predictors,
where $\hat{\rho}_i$ is the consistent estimator of $\rho_i$ for the predictor $x_{i,t}$.
Then the proposed estimators and the test statistics are given as follows.
align[align omitted — 719 chars of source]
where ${B_m}\equiv \left( \sum_{t=1}^T \tilde{z}_{t-1} x_{t-1}^\top \right)^{-1}\left( \sum_{t=1}^T \tilde{z}_{t-1} \tilde{z}_{t-1}^\top \hat{u}_t^2 \right)^{1/2}\left(\operatorname{I_K}+ {W_z}\hat{\varpi}_l\hat{\varpi}_l^\top{W_z}^\top \right)^{1/2}$
and
$\operatorname{\widehat{Avar}}(\hat{\beta}_m ) \equiv H_l \left(\operatorname{I_K}+{W_z}\hat{\varpi}_l\hat{\varpi}_l^\top {W_z}^\top \right)H_l^\top$.
Since $w_{z_i}=1+O_p(T^{-1})$ for SD predictors and $w_{z_i}= o_p(T^{-1})$ for SD predictors, we have ${Q_m}={\tilde{Q}_m}+O_p(T^{-1})$. By this result and Proposition (ref), we have the following proposition.
propUnder Assumption (ref) and the null hypothesis $H_0:R\beta =r_J$, for both SD and WD predictors, it follows that
\begin{align*}
Q_l - {Q_m} =
\begin{cases}
Q_l - {\tilde{Q}_m} +O_p(T^{-1})\\
{t_l^\top} H_l^\top \left(H_l H_l^\top \right)^{-1}
H_l \hat{\varpi}_l \hat{\varpi}_l^\top H_l^\top
\left(H_l H_l^\top \right)^{-1} H_l {t_l}
+ O_p(T^{-1})>0,\; SD;\\
O_p(T^{-1}),\quad WD;
\end{cases}
\end{align*}
By Proposition (ref), the finite sample property of ${Q_m}$ with SD predictors is similar to that of ${\tilde{Q}_m}$ rather than that of ${Q_l}$. Thus ${Q_m}$ with SD predictors still has the nice finite sample property of ${\tilde{Q}_m}$ discussed earlier. Meanwhile, the finite sample property of ${Q_m}$ with WD predictors is similar to that of ${Q_l}$ rather than that of ${\tilde{Q}_m}$. Thus ${Q_m}$ is free of the redundant terms of ${\tilde{Q}_m}$ with WD predictors. Therefore, ${Q_m}$ is free of size distortion for both joint and marginal tests in multiple models with both SD and WD predictors.
Next, we present the one-sided test for $H_0:\beta_i=0$, that is $H_0:R\beta =r_J$ constructed by setting the rank $J$ of matrix $R$ to be one and $r_J=0$. So we construct the t-test statistic ${Q_m^t}$ for right side test $H_0:\beta_i=0$ vs $H_a:\beta_i>0$ and left side test $H_0:\beta_i=0$ vs $H_a:\beta_i<0$.
align[align omitted — 132 chars of source]
Note that $Q_m = (Q_m^t)^2$ when $J=1$. By similar arguments for the nice property of ${Q_m}$, ${Q_m^t}$ performs well in terms of size for one-sided tests in both univariate and multiple models.
By equation ((ref)) and that $\hat{\varpi}_l=o_p(1)$, we have
align[align omitted — 204 chars of source]
Therefore, the limiting distribution of the t-test statistic ${Q_m^t}$ with $J=1$ and Wald type test statistic ${Q_m}$ with $J\geq 1$ under the null hypothesis is stated in the following theorem with its detailed proof given in the Appendix.\footnote{Our procedure works for univariate models as well since ${Q_m^t}$ is equal to the t-test statistic for univariate model when $J=K=1$.}
thmUnder Assumption (ref) and the null hypothesis $H_0:R\beta =r_J$, the limiting distribution of the t-test statistic ${Q_m^t}$ with $J=1$ and those of Wald type test statistic ${Q_m}$ are the standard normal distribution and the $\chi^2$-distribution with $J$ degrees of freedom, respectively.
We summarize our procedure in Algorithm (ref) below.
algorithm[algorithm omitted — 2,777 chars of source]
Monte Carlo Simulations
We demonstrate the effectiveness of the proposed procedure using a multivariate case.
We report the results for 1) the joint test $H_0:\beta=0$ vs $H_a:\beta\neq 0$; 2) the marginal test (two-sided) $H_0:\beta_i=0$ vs $H_a:\beta_i \neq 0$; 3) the marginal test (one-sided) $H_0:\beta_i=0$ vs $H_a:\beta_i > 0$ and $H_0:\beta_i=0$ vs $H_a:\beta_i < 0$.
The DGP of $y_t$ and $x_{t-1}$ is equation ((ref)), ((ref)) and ((ref)), in which $\mu=1$ and $\varphi_0=1$. The DGP of the innovations is as follows: $v_{i,t}=\gamma_i \eta_t + \check{v}_{i,t}$, where $(u_t,\check{v}_{i,t})^\top \;\sim \; i.i.d. \, \operatorname{N}(0_K,\operatorname{I}_{K\times K})$.
Thus the contemporaneous correlation coefficient between $\eta_t$ and $v_{i,t}$ is
$\gamma_i (1+\gamma_i^2)^{-1/2}$, which is the source of the size distortion of the standard IVX test statistics Kostakisetal2015,Phillips2013PredictiveRU and DiE and VEE shown in Proposition (ref).
We report the simulation results for a GARCH model with $\lambda=0.5$, $\varphi_1=0.1$, $\bar{ \varphi}_1=0.85$ and the sample $T=750$ with the nominal size $5\%$. \footnote{To save space, we did not report the simulation results for GARCH model with the sample size $T=250$ and $T=500$ and ARCH and i.i.d. model with the sample size $T=250$, $T=500$ and $T=750$ in the paper, which is similar to the result of GARCH model with the sample size $T=750$. The codes and results are available upon request.} Simulation is repeated $10,000$ times for each setting.
Set $\beta = \frac{b}{(1+K)/2}$ in Table (ref) and $\delta=0.95$ and $c_z=-4-K$ for all Tables. \footnote{\textcolor{black}{This setting (choice of $c_z$) is not exactly the same as the IVX of Phillips2013PredictiveRU, in that it makes the instrumental variable $z_{t-1}$ become less persistent as the number of predictors K grows bigger. Therefore, in this setting, we can see the advantages of our test even when the predictors are less persistent, but the VEE is nonetheless significant (see discussions after Proposition (ref)).}}
First, the results for the size and power performances of the proposed test statistics $Q_l$ and $Q_m$ for joint test $H_0:\beta=0$ are shown in Table (ref) in which $K=2,3,\cdots,10$ and $\beta= \frac{b}{(1+K)/2}$. Second, the size and power performances of the proposed test statistics $Q_l^t$ and $Q_m^t$ for two-sided marginal test $H_0:\beta_{i}=0$ vs $H_a:\beta_{i}\neq 0$ and right side marginal test $H_0:\beta_{i}=0$ vs $H_a:\beta_{i}> 0$ are shown in Panel A and Panel B of Table (ref), in which $K=10$.
The first column of Panel A and Panel B of Table (ref) and the third column of Table (ref) are the size results while the others are power results. In these models, we set $(\rho_1,\rho_2,\cdots,\rho_{K})^\top=(\bm{\rho})_K$, where $\bm{\rho}= (0.996,0.993,1,0.987,0.967,0.95,0.9,0.98,0.92,0.94)^\top$. We show the simulation results for $(\gamma_1,\gamma_2,\cdots,\gamma_K)^\top=(\Gamma)_K$, $\Gamma=(-3,2,1,3,1,0.833,0.667,0.5,0.333,0.167)^\top$ and thus the contemporaneous correlation coefficient $\left[\gamma_1 (1+\gamma_1^2)^{-1/2},\cdots,\gamma_K (1+\gamma_K^2)^{-1/2} \right]^\top$ between $u_t$ and $v_t$ are $(\tilde{\Gamma})_K$ and
$ \tilde{\Gamma}=(-0.949,0.894,0.707,0.949,0.707,0.64,0.555,0.447,0.316,0.164)^\top$.
We have the following findings from Tables (ref) and (ref). First, for the joint test $H_0:\beta=0$, the proposed test statistic ${Q_l}$ still suffers size distortion with multiple predictors ($K\geq 3$) and the size distortion grows bigger as the number of predictors K grows bigger. Meanwhile, the proposed test statistic ${Q_m}$ is almost free of size distortion in different settings. It is because ${Q_l}$ still suffers from DiE and VEE while ${Q_m}$ does not. Second, the size performance of ${Q_m^t}$ is better than that of ${Q_l^t}$ although ${Q_m^t}$ suffers small-scale size distortion for two-sided marginal test $H_0:\beta_{i}=0$ vs $H_a:\beta_{i}\neq 0$ and right side marginal test $H_0:\beta_{i}=0$ vs $H_a:\beta_{i}> 0$. Specifically, the size performances of ${Q_l^t}$ and ${Q_m^t}$ are comparable for right side marginal test $H_0:\beta_{i}=0$, while the size performance of ${Q_m^t}$ is much better than that of ${Q_l^t}$ for two-sided marginal test $H_0:\beta_{i}=0$ vs $H_a:\beta_{i}\neq 0$. Third, the power performances of ${Q_l^t}$ and ${Q_m^t}$ are comparable and quite well.
We also conduct other experiments (a univariate case and other settings of multivariate cases, e.g., for different $\lambda$'s) for the robustness of our conclusions. These results are reported in the Appendix. We obtain a similar conclusion there.
table[table omitted — 3,366 chars of source]
To sum up, the proposed test statistics ${Q_m}$ and ${Q_m^t}$ with $\lambda=0.5$ perform quite well in terms of both size and power in multivariate models. Thus, we recommend that empirical researchers apply our test statistics ${Q_m}$ and ${Q_m^t}$ and set $\lambda=0.5$.
table[table omitted — 1,936 chars of source]
Robust Inference for Bond Risk Premia
Prediction of bond risk premia is crucial for monetary policy and investment decisions. Following the literature 2005Bond, we use the excess log return of n-year U.S. discount bond rx(n), where $n=2,3,4,5$, as the left-hand side variable. We reexamine the predictability of five bond forward rates F1-F5 2005Bond and the first eight macroeconomic principal components M1-M8 (the $\hat f_i$, $i=1,\dots , 8$ in ludvigson2009macro) constructed by 132 macro variables in a multiple predictive regression model. Specifically, the forward rate F1 is the log price of the one-year discount bond; and the forward rate Fn of n-year (n=2,\dots 5) bond is the log price difference between (n-1)-year bond and n-year bond.\footnote{The dataset used in this section is from the Fama-Bliss database available on CRSP, and the macroeconomic factors are from the \href{https://www.sydneyludvigson.com/s/Updated_LN_Macro_Factors_2023FEB.xlsx}{Link of ludvigson2009macro}.} Our main sample period is 1964:01-2003:12, which is the same used in the aforementioned literature. We present the test statistics and the corresponding p-values in parenthesis marked with $\ast$, $\ast\ast$, $\ast\ast\ast$ implying the rejection of the null hypothesis at the 10%, 5%, and 1% levels. We also conduct empirical studies on A1) reexamining the three linear combinations of F1-F5 and M1-M8, which include CP 2005Bond constructed by F1-F5 and {LN1}, {LN2} (the F5, F6 in the original paper of ludvigson2009macro) constructed by M1-M8; and A2) evaluating the predictability during the pandemic period 2020:01-2022:12. Due to space limitations, we put the more details of these two studies in Appendix (ref).
table[table omitted — 1,088 chars of source]
The predictors' persistence parameters $\rho_i$'s are reported in Table (ref). It shows that F1-F5 and CP are likely the SD predictors, while the rest of the variables are likely to be WD predictors. Therefore, applying our robust testing procedure is necessary to avoid the spurious predictability induced by the size distortion discussed earlier. We show the testing results using F1-F5 and M1-M8 as regressors in Table (ref). This is helpful for practitioners to discover which component of CP, the {LN1} and {LN2} plays a vital part in predicting the bond excess returns.
We discuss the main results as follows. First, our method shows less significance (of predictability) than the original IVX Kostakisetal2015,Phillips2013PredictiveRU and ludvigson2009macro. This aligns with the theoretical results in Section (ref) that our method improves the size control performance of the original IVX by reducing the distortion effects of the higher-order terms. Our approach does not detect some predictability of F1-F5 and M1-M8 found by IVX and test statistics based on OLS. At the same time, all predictability of F1-F5 and M1-M8 found by our method are also detected by IVX and test statistics based on OLS except for the predictability of F5 for rx(5). Our test also finds that the linear combinations predictors CP, {LN1} and {LN2} have predicting power for the one-year bond excess returns at all maturities in the main sample period (with higher p-values than the IVX). During the pandemic period of 2020:01-2022:12, our test does not reject the null that CP, {LN1} and {LN2} have no predictability at all maturities, contrary to the OLS and IVX tests, where the former almost rejects all null hypotheses, and the latter finds CP to have predictability. The details are shown in Table (ref) of the Appendix.
table[table omitted — 2,241 chars of source]
From the economic perspective, our empirical result in Table (ref) supports the market segmentation theory
2013Flow,foley2016impact,greenwood2018asset. The market segmentation theory points out that short and long-term securities are not perfectly substitutable, and the short-term bond risk premia could only be predicted by the short-term forward rates and, similarly, the long-term forward rates for long-term bond risk premia. In the main sample, our findings parallel the studies of ludvigson2009macro that CP, {LN1} and {LN2} can predict the expected one-year excess return on bonds (see the Appendix for details). However, during the pandemic, the bond return is more likely to be driven by aggressive monetary policies and epidemic shocks goldstein2021covid, levine2021did, which hamper the predicting power of CP, {LN1} and {LN2}. Our test can detect the spurious predictability of CP, {LN1} and {LN2} in forecasting excess bond returns under different macroeconomic scenarios, while IVX and OLS do not.
Conclusion
This paper improves the popular IVX testing Kostakisetal2015,Phillips2013PredictiveRU by addressing not only DE and DiE, pointed out by HosseinkouchackDemetrescu2021, but also VEE which cannot be ignored, especially in multiple predictive regressions. We introduce a three-step method to deal with all the aforementioned issues regarding finite sample bias and variance-inflating terms. As a result, the size performance of the proposed method for the one-sided test and the test in multiple predictive regressions is improved significantly, while the power performance is comparable with the original IVX. Numerical simulations demonstrate the effectiveness of the newly proposed approach. Moreover, an empirical study of the predictability of bond returns shows that the original IVX rejects the null more often than our method, which supports the theoretical results that our procedure reduces the size distortion induced by the higher-order terms of IVX.
\setcounter{section}{0}
\setcounter{equation}{0}
\setcounter{table}{0}
\setcounter{figure}{0}
center[center omitted — 148 chars of source]
The Appendices include the following parts: Section (ref) contains all technical proofs and additional discussions complementary to the main text. Section (ref) collects extra simulation results. Section (ref) provides further details of the empirical exercise.
\newcounter{counter}[section]
Proofs
To help the readers navigate the testing procedure in the main text, we first explain the names of important test statistics.
description• \textcolor{black}{Here we use $G$ as the general notation of a test statistic. The subscript `ivx' in the random vector (or scalar) $G_{ivx}$ is to show $G_{ivx}$ is constructed by the method of the original IVX Kostakisetal2015. The subscript `a' in the random vector (or scalar) $G_a$ is to show $G_a$ is constructed by time series in the duration of the first subsample, while the subscript `b' in the random vector (or scalar) $G_a$ is to show $G_a$ is constructed by time series in the duration of the second subsample. The subscript `l' in the random vector (or scalar) $G_l$ is to show $G_l$ is constructed after the procedure that DE is eliminated. The subscript `n' in the random vector (or scalar) $G_l$ is to show $G_l$ is constructed after the procedure that DE is eliminated and DiE is reduced. The subscript `m' in the random vector (or scalar) $G_m$ is to show $G_m$ is constructed after the procedure that DE is eliminated and both DiE and VEE are reduced.}
Then, we give the proofs of the theoretical results. Then, we give the proofs of the theoretical results. The equation numbers refer to those in the main text. The new numbered equations in this appendix have the prefix A before the numbers, such as `A1'.
proof[Proof of equation ((ref))]
By equation ((ref)) and the definitions of $\tilde{S}_a$ and $\tilde{S}_b$ in equation ((ref)), it is known that
\begin{align}
\sum\nolimits_{t=1}^{T} \tilde{z}_{t-1}& = (\operatorname{I_K}-S_a) \sum\nolimits_{t=1}^{T_0} z_{t-1}+ (\operatorname{I_K}-S_b) \sum\nolimits_{t=T_0+1}^{T} z_{t-1} \\
& = \sum\nolimits_{t=1}^{T} z_{t-1} -S_a \sum\nolimits_{t=1}^{T_0} z_{t-1} - S_b \sum\nolimits_{t=T_0+1}^{T} z_{t-1} \nonumber\\
& = \sum\nolimits_{t=1}^{T} z_{t-1} - \lambda \sum\nolimits_{t=1}^{T } z_{t-1} - (1-\lambda) \sum\nolimits_{t=1}^{T} z_{t-1}=0 \nonumber
\end{align}
lemmaUnder Assumption (ref), for SD and WD predictors, it follows that
\begin{align}
D_T^{-1} \sum\limits_{t=1}^{T} \tilde{z}_{t-1} u_t \xrightarrow{d} \operatorname{MN}\left(0,\Sigma_{zz} \right)
\end{align}
proof[Proof of Lemma (ref)]
First, we consider the case with SD predictors. To begin with, we first prove that $\left(\sum_{t=1}^{T_0} z_{t-1}^\top u_t, \sum_{t=T_0+1}^{T} z_{t-1}^\top u_t\right)^\top$ follows asymptotically normal distribution. To this aim, we need to prove that
$$
\kappa_1^\top \frac{1}{T^{(1+\delta)/2}}\sum\limits_{t=1}^{T_0} z_{t-1}u_t + \kappa_2^\top \frac{1}{T^{(1+\delta)/2}} \sum\limits_{t=T_0+1}^{T} z_{t-1}u_t \xrightarrow{d}
\operatorname{N}\left[0, \lambda \kappa_1^\top\Omega_{zz}\kappa_1 + (1-\lambda) \kappa_2^\top\Omega_{zz}\kappa_2 \right].
$$
for any real vector $\kappa_1$ and $\kappa_2$ with dimension K and use the Cramer-Wold device. Define
\begin{align}
\breve{z}_{t-1} =
\begin{cases}
\kappa_1^\top z_{t-1},\; 1\leq t\leq T_0\\
\kappa_2^\top z_{t-1},\; T_0+1\leq t\leq T
\end{cases}.
\end{align}
Then
$\kappa_1^\top \frac{1}{T^{(1+\delta)/2}}\sum\limits_{t=1}^{T_0} z_{t-1}u_t + \kappa_2^\top \frac{1}{T^{(1+\delta)/2}} \sum\limits_{t=T_0+1}^{T} z_{t-1}u_t = \frac{1}{T^{(1+\delta)/2}}\sum\limits_{t=1}^{T_0} \breve{z}_{t-1} u_t$. Note that $\{ \frac{1}{T^{(1+\delta)/2}} \breve{z}_{t-1} u_t\}_{t=1}^T$ is a martingale difference sequence. So we need to verify the Lindeberg condition specified by Corollary 3.1 of HallHeyde1980 as follows
\begin{align*}
\sum_{t=1}^T E_{\mathcal{F}_{k-1}}\left(\left\|\frac{1}{T^{(1+\delta)/2}} \breve{z}_{t-1} u_t\right\|^2 {1}\left\{\left\|\frac{1}{T^{(1+\delta)/2}} \breve{z}_{t-1} u_t \right\|>\nu_\kappa \right\}\right) \xrightarrow{P} 0 \quad \forall \nu_\kappa>0
\end{align*}
In the following part, we will prove this equation.
By equation ((ref)) and Cauchy-Schwarz inequality, it follows that
\begin{align}
&\sum_{t=1}^T E_{\mathcal{F}_{k-1}}\left(\left\|\frac{1}{T^{(1+\delta)/2}} \breve{z}_{t-1} u_t\right\|^2 {1}\left\{\left\|\frac{1}{T^{(1+\delta)/2}} \breve{z}_{t-1} u_t \right\|>\nu_\kappa \right\}\right)\\
&\leq \|\kappa_1\|^2 \sum_{t=1}^{T_0} \frac{1}{T^{1+\delta}} \left\|z_{t-1} \right\|^2 E_{\mathcal{F}_{k-1}}\left( u_t^2 {1}\left\{\left\|\frac{1}{T^{(1+\delta)/2}} z_{t-1} u_t \right\|>\nu_\kappa/\|\kappa_1\| \right\}\right) \nonumber\\
&+ \|\kappa_2\|^2 \sum_{t=T_0+1}^{T} \frac{1}{T^{1+\delta}} \left\|z_{t-1} \right\|^2 E_{\mathcal{F}_{k-1}}\left( u_t^2 {1}\left\{\left\|\frac{1}{T^{(1+\delta)/2}} z_{t-1} u_t \right\|>\nu_\kappa/\|\kappa_2\| \right\}\right) \nonumber
\end{align}
It is straightforward that both events $\left\{\left\|\frac{1}{T^{(1+\delta)/2}} z_{t-1} u_t \right\|>\nu_\kappa/\|\kappa_1\| \right\}$ and $\left\{\left\|\frac{1}{T^{(1+\delta)/2}} z_{t-1} u_t \right\|>\nu_\kappa/\|\kappa_2\| \right\}$ implies $\left\{\left\|\frac{1}{T^{(1+\delta)/2}} z_{t-1} u_t \right\|>\nu_\kappa/\max(\|\kappa_1\| ,\|\kappa_2\| )\right\}$. So we have
\begin{align}
{1}\left\{\left\|\frac{1}{T^{(1+\delta)/2}} z_{t-1} u_t \right\|>\nu_\kappa/\|\kappa_2\| \right\}\leq
{1}\left\{\left\|\frac{1}{T^{(1+\delta)/2}} z_{t-1} u_t \right\|>\nu_\kappa/\max(\|\kappa_1\| ,\|\kappa_2\| ) \right\},\\
{1}\left\{\left\|\frac{1}{T^{(1+\delta)/2}} z_{t-1} u_t \right\|>\nu_\kappa/\|\kappa_2\| \right\}\leq
{1}\left\{\left\|\frac{1}{T^{(1+\delta)/2}} z_{t-1} u_t \right\|>\nu_\kappa/\max(\|\kappa_1\| ,\|\kappa_2\| ) \right\}.
\end{align}
And
\begin{align}
\|\kappa_1\|^2 \leq \max(\|\kappa_1\|^2,\|\kappa_2\|^2),\quad \|\kappa_2\|^2\leq \max(\|\kappa_1\|^2,\|\kappa_2\|^2).
\end{align}
By equations ((ref)), ((ref)), ((ref)) and ((ref)), it follows that
\begin{align}
&\sum_{t=1}^T E_{\mathcal{F}_{k-1}}\left(\left\|\frac{1}{T^{(1+\delta)/2}} \breve{z}_{t-1} u_t\right\|^2 {1}\left\{\left\|\frac{1}{T^{(1+\delta)/2}} \breve{z}_{t-1} u_t \right\|>\nu_\kappa \right\}\right)\\
&\leq \max(\|\kappa_1\|^2, \|\kappa_2\|^2) \sum_{t=1}^{T} \frac{1}{T^{1+\delta}} \left\|z_{t-1} \right\|^2 E_{\mathcal{F}_{k-1}}\left( u_t^2 {1}\left\{\left\|\frac{1}{T^{(1+\delta)/2}} z_{t-1} u_t \right\|>\nu_\kappa/\max(\|\kappa_1\| ,\|\kappa_2\| ) \right\}\right) \nonumber
\end{align}
By part (ii) of Lemma 5.2 of the proof of the Online Technical Supplement to Phillips2013PredictiveRU, we have
\begin{align}
\sum_{t=1}^{T} \frac{1}{T^{1+\delta}} \left\|z_{t-1} \right\|^2 E_{\mathcal{F}_{k-1}}\left( u_t^2 {1}\left\{\left\|\frac{1}{T^{(1+\delta)/2}} z_{t-1} u_t \right\|>\nu_\kappa/\max(\|\kappa_1\| ,\|\kappa_2\| ) \right\}\right)
\xrightarrow{P} 0
\end{align}
By equations ((ref)) and ((ref)), the following Lindeberg condition holds.
\begin{align}
\sum_{t=1}^T E_{\mathcal{F}_{k-1}}\left(\left\|\frac{1}{T^{(1+\delta)/2}} \breve{z}_{t-1} u_t\right\|^2 {1}\left\{\left\|\frac{1}{T^{(1+\delta)/2}} \breve{z}_{t-1} u_t \right\|>\nu_\kappa \right\}\right)\xrightarrow{P} 0,\quad
\forall \nu_\kappa>0
\end{align}
Next, we verify the stability condition for $\{ \frac{1}{T^{(1+\delta)/2}} \breve{z}_{t-1} u_t\}_{t=1}^T$. By equation ((ref)), it is known that
\begin{align*}
\frac{1}{T} \sum_{t=1}^{T_0} z_{t-1}z_{t-1}^\top &= \rho_z^2 \frac{1}{T} \sum_{t=2}^{T_0}z_{t-2}z_{t-2}^\top + \rho_z \frac{1}{T} \sum_{t=1}^{T_0}\Delta x_{t-1} \Delta x_{t-1}^\top \\
&+ \rho_z \frac{1}{T} \sum_{t=2}^{T_0} z_{t-2} \Delta x_{t-1}^\top
+ \frac{1}{T} \sum_{t=2}^{T_0} \Delta x_{t-1} z_{t-2}^\top
\end{align*}
Thus
\begin{align}
(1-\rho_z^2)\frac{1}{T} \sum_{t=1}^{T_0} z_{t-1} z_{t-1}^\top &= -\rho_z^2 \frac{1}{T} z_{T_0-1} z_{T_0-1}^\top + \rho_z\frac{1}{T} \sum_{t=1}^{T_0}\Delta x_{t-1} \Delta x_{t-1}^\top \\
& + \rho_z \frac{1}{T} \sum_{t=1}^{T_0} z_{t-2} \Delta x_{t-1}^\top
+ \frac{1}{T} \sum_{t=1}^{T_0} \Delta x_{t-1} z_{t-2}^\top \nonumber\\
&= \lambda \frac{1}{T_0} \sum_{t=1}^{T_0}\Delta x_{t-1} \Delta x_{t-1}^\top + o_p(1). \nonumber
\end{align}
The last step of the above equation holds since $\frac{1}{T^{\delta/2}}z_{T_0-1}=O_p(1)$, $\frac{1}{T^{(1+\delta)/2}} \sum_{t=1}^{T_0} z_{t-2} \Delta x_{t-1}^\top =O_p(1)$ and $\frac{1}{T^{(1+\delta)/2}} \sum_{t=1}^{T_0} \Delta x_{t-1} z_{t-2}^\top =O_p(1)$ by Online Appendix of Kostakisetal2015 and $T_0=\lambda T$.
Similarly,
\begin{align}
(1-\rho_z^2)\frac{1}{T} \sum_{t=T_0+1}^{T} z_{t-1} z_{t-1}^\top =(1-\lambda) \frac{1}{T-T_0} \sum_{t=T_0+1}^{T}\Delta x_{t-1} \Delta x_{t-1}^\top + o_p(1).
\end{align}
By equations ((ref)), ((ref)) and ((ref)), the following stability condition holds.
\begin{align}
\frac{1}{T^{1+\delta}}\sum\limits_{t=1}^{T_0} \breve{z}_{t-1} \breve{z}_{t-1}^\top u_t^2 &=\Sigma_{uu} \frac{1}{T^{1+\delta}}\sum\limits_{t=1}^{T_0} \breve{z}_{t-1} \breve{z}_{t-1}^\top + o_p(1) \\
&\xrightarrow{P} \lambda \kappa_1^\top \Omega_{zz} \kappa_1 + (1-\lambda)\kappa_2^\top \Omega_{zz} \kappa_2 \nonumber
\end{align}
By equations ((ref)) and ((ref)), Corollary 3.1 of HallHeyde1980 yields the following results.
\begin{align}
\frac{1}{T^{(1+\delta)/2}}\sum\limits_{t=1}^{T_0} \breve{z}_{t-1}u_t& =\kappa_1^\top \frac{1}{T^{(1+\delta)/2}}\sum\limits_{t=1}^{T_0} z_{t-1}u_t + \kappa_2^\top \frac{1}{T^{(1+\delta)/2}} \sum\limits_{t=T_0+1}^{T} z_{t-1}u_t \\
&\xrightarrow{d}
\operatorname{N}\left[0, \lambda \kappa_1^\top \Omega_{zz} \kappa_1 + (1-\lambda)\kappa_2^\top \Omega_{zz} \kappa_2\right] \nonumber
\end{align}
for any real vectors $\kappa_1$ and $\kappa_2$. By the Cramer-Wold device and equation ((ref)), for SD predictors, we have
\begin{align}
\left(\frac{1}{T^{(1+\delta)/2}}\sum_{t=1}^{T_0} z_{t-1}u_t, \frac{1}{T^{(1+\delta)/2}}\sum_{t=T_0+1}^{T} z_{t-1}u_t\right)^\top \xrightarrow{d}
\operatorname{N}\left\{0, \operatorname{diag}\left[ \lambda\Omega_{zz} , (1-\lambda)\Omega_{zz}\right]\right\}.
\end{align}
Similarly, for WD predictors, we also have
\begin{align}
\left(\frac{1}{\sqrt{T} }\sum_{t=1}^{T_0} z_{t-1}u_t, \frac{1}{\sqrt{T}}\sum_{t=T_0+1}^{T} z_{t-1}u_t\right)^\top \xrightarrow{d}
\operatorname{N}\left\{ {0}, \operatorname{diag}\left[ \lambda\Omega_{zz} , (1-\lambda)\Omega_{zz}\right]\right\}.
\end{align}
By equation ((ref)), it follows that
\begin{align}
\left(\operatorname{I_K} -S_a,\operatorname{I_K} -S_b \right) \Rightarrow \left(\operatorname{I_K} -\tilde{S}_a,\operatorname{I_K} -\tilde{S}_b \right),
\end{align}
where the joint convergence is guaranteed by that both $\tilde{S}_a$ and $\tilde{S}_b$ are the function of $B_v(\cdot)$. And by the proof of part (i) of Proposition A1 and Lemma 3.2 of PhillipsMagdalinos2009, the joint convergence of $\left(D_T^{-1}\sum_{t=1}^{T_0} z_{t-1}u_t, D_T^{-1} \sum_{t=T_0+1}^{T} z_{t-1}u_t\right)^\top $ and $B_v(\cdot)$ applies. As a result, considering that $\left(\operatorname{I_K} -\tilde{S}_a,\operatorname{I_K} -\tilde{S}_b \right)$ is the function of $B_v(\cdot)$, the joint convergence of $\left(1-S_a,1-S_b \right)$ and $\left(D_T^{-1}\sum_{t=1}^{T_0} z_{t-1}u_t, D_T^{-1} \sum_{t=T_0+1}^{T} z_{t-1}u_t\right)^\top $ holds.
By equations ((ref)) and ((ref)) and the continuous mapping theorem and the joint convergence, the following equation holds for both SD and WD predictors.
\begin{align}
D_T^{-1}\sum_{t=1}^{T} \tilde{z}_{t-1}u_t &= \left(\operatorname{I_K} -S_a,\operatorname{I_K} -S_b \right)\left(D_T^{-1}\sum_{t=1}^{T_0} z_{t-1}u_t, D_T^{-1}\sum_{t=T_0+1}^{T} z_{t-1}u_t\right)^\top \xrightarrow{d} \operatorname{MN}\left(0,\Sigma_{zz} \right). \nonumber
\end{align}
lemmaUnder Assumption (ref), for SD and WD predictors, we have
$$D_T^{-2}\sum_{t=1}^T \tilde{z}_{t-1} \tilde{z}_{t-1}^\top \hat{u}_t^2\Rightarrow \Sigma_{zz}.$$
proof[Proof of Lemma (ref)]
Following the similar procedure for the stability condition shown in equation ((ref)) for $\{ \frac{1}{T^{(1+\delta)/2}} \breve{z}_{t-1} u_t\}_{t=1}^T$ of Proof for Lemma (ref), Lemma (ref) is proved.
lemmaUnder Assumption A.1 and the condition $1/2<\delta<1$, by Taylor expansion of matrix Taylor2016, we have
\begin{align}
&\left( \frac{1}{T^{1+\delta }} \sum_{t=1}^T z_{t-1}^*(z_{t-1}^*)^\top \hat{u}_t^2 \right)^{-1/2} =\left( \frac{1}{T^{1+\delta }} \sum_{t=1}^T z_{t-1}^*(z_{t-1}^*)^\top u_t^2 \right)^{-1/2} +o_p\left[ T^{-(1-\delta)/2} \right] \\
&=
\left[\operatorname{I_K}+\left(\frac{1}{T^{1+\delta }} \sum_{t=1}^T z_{t-1}^*(z_{t-1}^*)^\top u_t^2 -\operatorname{I_K} \right) \right]^{-1/2} +o_p\left[ T^{-(1-\delta)/2} \right] \nonumber\\
& = \operatorname{I_K} -\frac{1}{2} \left(\frac{1}{T^{1+\delta }} \sum_{t=1}^T z_{t-1}^*(z_{t-1}^*)^\top u_t^2 -\operatorname{I_K}\right)+o_p \left(\frac{1}{T^{1+\delta }} \sum_{t=1}^T z_{t-1}^*(z_{t-1}^*)^\top u_t^2 -\operatorname{I_K}\right)+o_p\left[ T^{-(1-\delta)/2} \right] \nonumber \\
& = \operatorname{I_K} -\frac{1}{2} \left(\frac{1}{T^{1+\delta }} \sum_{t=1}^T z_{t-1}^*(z_{t-1}^*)^\top u_t^2 -\operatorname{I_K}\right)+o_p\left[ T^{-(1-\delta)/2} \right].\nonumber
\end{align}
And
\begin{align}
&\left( \frac{1}{T^{1+\delta }} \sum_{t=1}^T \tilde{z}_{t-1}^*(\tilde{z}_{t-1}^*)^\top \hat{u}_t^2 \right)^{-1/2} =\left( \frac{1}{T^{1+\delta }} \sum_{t=1}^T \tilde{z}_{t-1}^*(\tilde{z}_{t-1}^*)^\top u_t^2 \right)^{-1/2} +o_p\left[ T^{-(1-\delta)/2} \right] \\
&=
\left[\operatorname{I_K}+\left(\frac{1}{T^{1+\delta }} \sum_{t=1}^T \tilde{z}_{t-1}^*(\tilde{z}_{t-1}^*)^\top u_t^2 -\operatorname{I_K} \right) \right]^{-1/2} +o_p\left[ T^{-(1-\delta)/2} \right] \nonumber\\
& = \operatorname{I_K} -\frac{1}{2} \left(\frac{1}{T^{1+\delta }} \sum_{t=1}^T \tilde{z}_{t-1}^*(\tilde{z}_{t-1}^*)^\top u_t^2 -\operatorname{I_K}\right)+o_p \left(\frac{1}{T^{1+\delta }} \sum_{t=1}^T \tilde{z}_{t-1}^*(\tilde{z}_{t-1}^*)^\top u_t^2 -\operatorname{I_K}\right)+o_p\left[ T^{-(1-\delta)/2} \right].\nonumber\\
& = \operatorname{I_K} -\frac{1}{2} \left(\frac{1}{T^{1+\delta }} \sum_{t=1}^T \tilde{z}_{t-1}^*(\tilde{z}_{t-1}^*)^\top u_t^2 -\operatorname{I_K}\right)+o_p\left[ T^{-(1-\delta)/2} \right].\nonumber
\end{align}
proof[Proof Lemma (ref)]
By online appendix of Kostakisetal2015, it follows that
\begin{align}
\frac{1}{T^{1+\delta }} \sum_{t=1}^T z_{t-1}^*(z_{t-1}^*)^\top \hat{u}_t^2
& =\Omega_{zz}^{-1/2} \frac{1}{T^{1+\delta }} \sum_{t=1}^T z_{t-1}z_{t-1}^\top \hat{u}_t^2 \Omega_{zz}^{-1/2} \\
&\xrightarrow{P} \Omega_{zz}^{-1/2} \Sigma_{vv}/(-2c_z)\Sigma_{uu} \Omega_{zz}^{-1/2} = \Omega_{zz}^{-1/2} \Omega_{zz}^{1/2}\Omega_{zz}^{1/2} \Omega_{zz}^{-1/2} = \operatorname{I_K}. \nonumber
\end{align}
Thus
\begin{align}
\left(\frac{1}{T^{1+\delta }} \sum_{t=1}^T z_{t-1}^*(z_{t-1}^*)^\top \hat{u}_t^2 - \operatorname{I_K} \right)
\left(\frac{1}{T^{1+\delta }} \sum_{t=1}^T z_{t-1}^*(z_{t-1}^*)^\top \hat{u}_t^2 - \operatorname{I_K} \right)
\xrightarrow{P} 0.
\end{align}
By equation ((ref)) and continuous mapping theorem, the equation about determinant holds as follows.
\begin{align}
\operatorname{det}\left|\left(\frac{1}{T^{1+\delta }} \sum_{t=1}^T z_{t-1}^*(z_{t-1}^*)^\top \hat{u}_t^2 - \operatorname{I_K} \right)
\left(\frac{1}{T^{1+\delta }} \sum_{t=1}^T z_{t-1}^*(z_{t-1}^*)^\top \hat{u}_t^2 - \operatorname{I_K} \right) \right| \xrightarrow{P} 0.
\end{align}
Additionally,
\begin{align}
\operatorname{det}\left|\left(\frac{1}{T^{1+\delta }} \sum_{t=1}^T z_{t-1}^*(z_{t-1}^*)^\top \hat{u}_t^2 - \operatorname{I_K} \right)
\left(\frac{1}{T^{1+\delta }} \sum_{t=1}^T z_{t-1}^*(z_{t-1}^*)^\top \hat{u}_t^2 - \operatorname{I_K} \right) \right| = \lambda_1 \lambda_2 \cdots \lambda_K,
\end{align}
where $\lambda_i>0$, $i=1,2,\cdots,K$ is the singular value of $\frac{1}{T^{1+\delta }} \sum_{t=1}^T z_{t-1}^*(z_{t-1}^*)^\top \hat{u}_t^2 - \operatorname{I_K} $. By equations ((ref)) and ((ref)) and that $\frac{1}{T^{1+\delta }} \sum_{t=1}^T z_{t-1}^*(z_{t-1}^*)^\top \hat{u}_t^2 - \operatorname{I_K}=o_p(1)$, it follows that
\begin{align}
\max_{i=1,2,\cdots,K} \lambda_i \xrightarrow{P} 0.
\end{align}
On the other hand, following the similar procedure of HosseinkouchackDemetrescu2021, we have
\begin{align}
\left( \frac{1}{T^{1+\delta }} \sum_{t=1}^T z_{t-1}^*(z_{t-1}^*)^\top \hat{u}_t^2 \right)^{-1/2} =\left( \frac{1}{T^{1+\delta }} \sum_{t=1}^T z_{t-1}^*(z_{t-1}^*)^\top u_t^2 \right)^{-1/2} +o_p\left[ T^{-(1-\delta)/2} \right].
\end{align}
By Lemma (ref) and ((ref)), equation ((ref)) holds. Also, by the similar procedure, equation ((ref)) is proved.
proof[Proof of Proposition (ref)]
We consider the higher order terms induced by the SD predictors.
By equation ((ref)), it follows that
\begin{align}
z_{t-1}^*=\Omega_{zz}^{-1/2}z_{t-1} =\Omega_{zz}^{-1/2}\left(\rho_z z_{t-1} + \Delta x_{t-1} \right)
= \rho_z z_{t-2}^* + \Delta x_{t-1}^*,
\end{align}
where $\Delta x_{t-1}^*= x_{t-1}^* - {x}_{t-2}^* $. We define $v_t^*= \Omega_{zz}^{-1/2}v_t$ and thus
$\Delta x_{t-1}^* =v_t^* +\frac{c_z}{T^\delta} {x}_{t-2}^* $ and
\begin{align}
\frac{1}{T} \sum_{t=1}^Tv_t^*(v_t^*)^\top \xrightarrow{P} \operatorname{E}\left[ v_t^*(v_t^*)^\top \right]= - 2c_z/\Sigma_{uu}\operatorname{I_K},\\
\frac{1}{T} \sum_{t=1}^T(v_t^*)^\top v_t^* \xrightarrow{P} \operatorname{E}\left[ (v_t^*)^\top v_t^* \right]= - 2c_z/\Sigma_{uu} K. \nonumber
\end{align}
Then by equation ((ref)) in Lemma (ref), we have
\begin{align}
\left[ \frac{1}{T^{1+\delta}} \sum_{t=1}^T z_{t-1}^* (z_{t-1}^*)^\top u_t^2 \right]^{-1/2}\frac{1}{T^{1 / 2+\delta / 2}} \sum_{t=1}^T z_{t-1}^* u_t ={Z_T}+{B_T}+o_p\left[T^{-(1-\delta)/2}\right]
\end{align}
where
$$
{Z_T}= \frac{1}{T^{1 / 2+\delta / 2}} \sum_{t=1}^T z_{t-1}^* u_t
$$
and
$$
\begin{aligned}
{B_T}=- & \frac{1}{2} \left[\frac{1}{T^{1+\delta }} \sum_{t=1}^T z_{t-1}^* (z_{t-1}^*)^\top u_t^2- \operatorname{I_K} \right]
\left(\frac{1}{T^{1 / 2+\delta / 2}} \sum_{t=1}^T z_{t-1}^* u_t\right)
\end{aligned}
$$
By the online appendix of Kostakisetal2015 and the continuous mapping theorem, it is known that ${Z_T} \xrightarrow{d} \operatorname{N}(0_K,\operatorname{I_K})$.
Since $\operatorname{E}
\left(\frac{1}{T^{1 / 2+\delta / 2}} \sum_{t=1}^T z_{t-1}^* u_t\right)=0$, then
\begin{align}
&T^{(1-\delta)/2}\operatorname{E}\left({B_T} \right) \\
&=- \frac{1}{2} T^{(1-\delta)/2}\operatorname{E}\left[\left(\frac{1}{T^{1+\delta }} \sum_{t=1}^T z_{t-1}^* (z_{t-1}^*)^\top u_t^2- \operatorname{I_K} \right)
\left(\frac{1}{T^{1 / 2+\delta / 2}} \sum_{t=1}^T z_{t-1}^* u_t\right)\right]\nonumber\\
&=- \frac{1}{2} T^{(1-\delta)/2} \operatorname{E}\left[\left(\frac{1}{T^{1+\delta }} \sum_{t=1}^T z_{t-1}^* (z_{t-1}^*)^\top u_t^2 \right)
\left(\frac{1}{T^{1 / 2+\delta / 2}} \sum_{t=1}^T z_{t-1}^* u_t\right)\right]\nonumber
\end{align}
And
\begin{align}
&T^{(1-\delta)/2}\operatorname{E}\left[\left(\frac{1}{T^{1+\delta }} \sum_{t=1}^T z_{t-1}^* (z_{t-1}^*)^\top u_t^2 \right)
\left(\frac{1}{T^{1 / 2+\delta / 2}} \sum_{t=1}^T z_{t-1}^* u_t\right)\right]\\
& =\frac{1}{T^{1+2\delta} }\mathrm{E}\left(\sum_{t=1}^T z_{t-1}^* (z_{t-1}^*)^\top z_{t-1}^* u_t^3\right)
+ \frac{1}{T^{1+2\delta} } \mathrm{E}\left(\sum_{t=1}^T \sum_{s=1}^{t-1} z_{s-1}^*(z_{s-1}^*)^\top z_{t-1}^* u_s^2 u_t \right) \nonumber\\
& +\frac{1}{T^{1+2\delta} } \mathrm{E}\left(\sum_{t=1}^{T-1} \sum_{s=t+1}^T z_{s-1}^*(z_{s-1}^*)^\top z_{t-1}^* u_t u_s^2\right) \nonumber\\
& = \frac{1}{T^{1+2\delta} }\mathrm{E}\left(\sum_{t=1}^T z_{t-1}^* (z_{t-1}^*)^\top z_{t-1}^* u_t^3\right)
+\frac{1}{T^{1+2\delta} }\mathrm{E}\left(\sum_{t=1}^{T-1} \sum_{s=t+1}^T z_{s-1}^*(z_{s-1}^*)^\top z_{t-1}^* u_t u_s^2\right). \nonumber
\end{align}
Let $S_0=S_{0,1}+S_{0,2}$, with $S_{0,1}=\mathrm{E}\left(\sum_{t=1}^T z_{t-1}^* (z_{t-1}^*)^\top z_{t-1}^* u_t^3\right)$ and
$$
S_{0,2}=\mathrm{E}\left(\sum_{t=1}^{T-1} \sum_{s=t+1}^T z_{s-1}^*(z_{s-1}^*)^\top z_{t-1}^* u_t u_s^2\right).
$$ We work out $S_{0,2}$ first. By the proof of HosseinkouchackDemetrescu2021,
$$
z_{t-1}^* =\sum_{j=1}^{t-1} c_{j, t-1} v_j^* \text { with } c_{j, t-1}=\frac{\rho_z^{t-1-j}(1-\rho_z)-\rho_j^{t-1-j}(1-\rho_j)}{\rho_j-\rho_z} .
$$
Then for $s\geq t+1$, it follows that
\begin{small}
\begin{align}
&\frac{1}{T^{1+2\delta} }\mathrm{E}\left[ z_{s-1}^*(z_{s-1}^*)^\top z_{t-1}^* u_t u_s^2\right]\\
&=\mathrm{E}(u_s^2)\frac{1}{T^{1+2\delta} } \mathrm{E}\left[\left(\sum_{j=1}^{t-1} c_{j, s-1} (v_j^*)+\sum_{j=t}^{s-1} c_{j, s-1} (v_j^*)\right) \left(\sum_{j=1}^{t-1} c_{j, s-1} (v_j^*)+\sum_{j=t}^{s-1} c_{j, s-1} (v_j^*)\right)^\top \left(\sum_{j=1}^{t-1} c_{j, t-1} (v_j^*)\right) u_t\right] \nonumber\\
& =\Sigma_{uu}\frac{1}{T^{1+2\delta} } \mathrm{E}\left[\left(\sum_{j=1}^{t-1} c_{j, s-1} (v_j^*)\right)\left(\sum_{j=t}^{s-1} c_{j, s-1} (v_j^*)\right)^\top \left(\sum_{j=1}^{t-1} c_{j, t-1} (v_j^*)\right) u_t \right] \nonumber\\
&+\Sigma_{uu}\frac{1}{T^{1+2\delta} }\mathrm{E}\left[\left(\sum_{j=t}^{s-1} c_{j, s-1} (v_j^*)\right) \left(\sum_{j=1}^{t-1} c_{j, s-1} (v_j^*)\right)^\top \left(\sum_{j=1}^{t-1} c_{j, t-1} (v_j^*)\right) u_t \right]. \nonumber
\end{align}
\end{small}
Thus by the independence of innovations and equation ((ref)), it follows that
\begin{align}
&\frac{1}{T^{1+2\delta} } \mathrm{E}\left[\left(\sum_{j=1}^{t-1} c_{j, s-1} (v_j^*)\right)\left( \sum_{j=t}^{s-1} c_{j, s-1} (v_j^*)\right)^\top \left(\sum_{j=1}^{t-1} c_{j, t-1} (v_j^*)\right) u_t \right] \\
&=\frac{1}{T^{1+2\delta} } \mathrm{E}\left[\left(\sum_{j=1}^{t-1} c_{j, s-1} (v_j^*)\right)\left(c_{t, s-1} (v_t^*)+\sum_{j=t+1}^{s-1} c_{j, s-1} (v_j^*)\right)^\top \left(\sum_{j=1}^{t-1} c_{j, t-1} (v_j^*)\right) u_t \right] \nonumber\\
&= \frac{1}{T^{1+2\delta} } \mathrm{E}\left[\left(\sum_{j=1}^{t-1} c_{j, s-1} (v_j^*)\right) c_{t, s-1} (v_t^*)^\top \left(\sum_{j=1}^{t-1} c_{j, t-1} (v_j^*)\right) u_t \right] \nonumber\\
&= \frac{1}{T^{1+2\delta} } \mathrm{E}\left( \sum_{j=1}^{t-1} c_{j, s-1} v_j^* c_{t, s-1} (v_t^*)^\top c_{j, t-1} v_j^* u_t \right) \nonumber\\
&= \frac{1}{T^{1+2\delta} } \mathrm{E}\left( \sum_{j=1}^{t-1} c_{j, s-1} c_{t, s-1} c_{j, t-1} v_j^* (v_t^*)^\top v_j^* u_t \right) \nonumber \\
& = \frac{1}{T^{1+2\delta} } \sum_{j=1}^{t-1} c_{j, s-1} c_{t, s-1} c_{j, t-1} \mathrm{E}\left( v_j^* (v_t^*)^\top v_j^* u_t \right). \nonumber\\
& = \frac{1}{T^{1+2\delta} } \sum_{j=1}^{t-1} c_{j, s-1} c_{t, s-1} c_{j, t-1} \mathrm{E}\left( v_j^* v_t^* (v_j^*)^\top u_t \right) \nonumber\\
& = \frac{1}{T^{1+2\delta} } \sum_{j=1}^{t-1} c_{j, s-1} c_{t, s-1} c_{j, t-1} \mathrm{E}\left(u_t v_t^* (v_j^*)^\top v_j^* \right) \nonumber\\
& = \frac{1}{T^{1+2\delta} } \sum_{j=1}^{t-1} c_{j, s-1} c_{t, s-1} c_{j, t-1} \mathrm{E}\left(u_t v_t^* \right) \mathrm{E}\left[ (v_j^*)^\top v_j^* \right] \nonumber\\
&= \Omega_{zz}^{-1/2} \Sigma_{vu} \left( - 2c_z/\Sigma_{uu}\right) \frac{1}{T^{1+2\delta} } \sum_{j=1}^{t-1} c_{j, s-1} c_{t, s-1} c_{j, t-1}. \nonumber
\end{align}
The last four lines of equation ((ref)) hold since $(v_t^*)^\top v_j^* = v_t^* (v_j^*)^\top $ is a scalar and $j<t$.
And by equations (S.2)-(S.5) of the appendix of HosseinkouchackDemetrescu2021, we have
\begin{align}
\frac{1}{T^{1+2\delta} }\sum_{j=1}^{t-1} c_{j, s-1} c_{t, s-1} c_{j, t-1} \rightarrow \frac{1}{4c_z^2}
\end{align}
Then by equations ((ref)) and ((ref)), it follows that
\begin{align}
&\frac{1}{T^{1+2\delta} } \mathrm{E}\left[\left(\sum_{j=1}^{t-1} c_{j, s-1} (v_j^*)\right)\left( \sum_{j=t}^{s-1} c_{j, s-1} (v_j^*)\right)^\top \left(\sum_{j=1}^{t-1} c_{j, t-1} (v_j^*)\right) u_t \right] \\
&\rightarrow (4c_z^2)^{-1}\Omega_{zz}^{-1/2} \Sigma_{vu} \left( - 2c_z/\Sigma_{uu}\right)= - \Omega_{zz}^{-1/2} \Sigma_{vu} /(2c_z\Sigma_{uu})
\nonumber
\end{align}
On the other hand,
\begin{align}
&\frac{1}{T^{1+2\delta} } \mathrm{E}\left[\left( \sum_{j=t}^{s-1} c_{j, s-1} (v_j^*)\right) \left(\sum_{j=1}^{t-1} c_{j, s-1} (v_j^*)\right)^\top \left(\sum_{j=1}^{t-1} c_{j, t-1} (v_j^*)\right) u_t \right] \\
&=\frac{1}{T^{1+2\delta} } \mathrm{E}\left[\left(c_{t, s-1} (v_t^*)+\sum_{j=t+1}^{s-1} c_{j, s-1} (v_j^*)\right) \left(\sum_{j=1}^{t-1} c_{j, s-1} (v_j^*)\right)^\top \left(\sum_{j=1}^{t-1} c_{j, t-1} (v_j^*)\right) u_t \right] \nonumber\\
&= \frac{1}{T^{1+2\delta} } \mathrm{E}\left[c_{t, s-1} (v_t^*) \left(\sum_{j=1}^{t-1} c_{j, s-1} (v_j^*)^\top\right) \left(\sum_{j=1}^{t-1} c_{j, t-1} (v_j^*)\right) u_t \right] \nonumber\\
&= \frac{1}{T^{1+2\delta} } \mathrm{E}\left( \sum_{j=1}^{t-1} c_{t, s-1} (v_t^*) c_{j, s-1} (v_j^*)^\top c_{j, t-1} (v_j^*) u_t \right) \nonumber\\
&= \frac{1}{T^{1+2\delta} } \mathrm{E}\left( \sum_{j=1}^{t-1} c_{j, s-1} c_{t, s-1} c_{j, t-1} (v_t^*) u_t (v_j^*)^\top (v_j^*) \right)
= \frac{1}{T^{1+2\delta} } \sum_{j=1}^{t-1} c_{j, s-1} c_{t, s-1} c_{j, t-1} \mathrm{E}\left( (v_t^*) u_t (v_j^*)^\top (v_j^*) \right). \nonumber\\
&= \mathrm{E}\left[ v_t^* u_t \right] \mathrm{E}\left[(v_j^*)^\top v_j^* \right] \frac{1}{T^{1+2\delta} } \sum_{j=1}^{t-1} c_{j, s-1} c_{t, s-1} c_{j, t-1} \rightarrow - K\Omega_{zz}^{-1/2} \Sigma_{vu} /(2c_z\Sigma_{uu}) . \nonumber
\end{align}
By equations ((ref)), ((ref)) and ((ref)), it follows that
\begin{align}
\frac{1}{T^{1+2\delta} } \mathrm{E}\left( z_{s-1}^*(z_{s-1}^*)^\top z_{t-1}^* u_t u_s^2\right) \rightarrow - \Omega_{zz}^{-1/2} \Sigma_{vu} /c_z ,\quad s \geq t + 1,
\end{align}
By the same procedure of the appendix of HosseinkouchackDemetrescu2021, we have
\begin{align}
\frac{1}{T^{1+2\delta} }\mathrm{E}\left(\sum_{t=1}^T z_{t-1}^* (z_{t-1}^*)^\top z_{t-1}^* u_t^3\right) \rightarrow 0.
\end{align}
By equations ((ref)), ((ref)), ((ref)) and ((ref))
\begin{align}
T^{(1-\delta)/2}\operatorname{E}\left({B_T} \right) &\rightarrow - \frac{K+1}{2} \Omega_{zz}^{-1/2} \Sigma_{vu} /(-2 c_z).
\end{align}
Since $\Omega_{zz}= \Sigma_{uu} \Omega_{vv}/(-2c_z)$, equation ((ref)) is the same as
\begin{align}
T^{(1-\delta)/2}\operatorname{E}\left({B_T} \right) &\rightarrow - \frac{K+1}{2}{\rho_{uv^*}}/ \sqrt{-2 c_z}.
\end{align}
Thus by equations ((ref)) and ((ref)), Proposition (ref) holds.
proof[Proof of Theorem (ref)]
By the definitions of $Q_{ivx}$ and $\tilde{Q}_{ivx}$, it follows that
\begin{align}
&Q_{ivx}-\tilde{Q}_{ivx}=
\mathbf{t_{ivx}^\top} H_{ivx}^\top \left(H_{ivx} H_{ivx}^\top \right)^{-1} H_{ivx} \mathbf{t_{ivx}}
-\mathbf{t_{ivx}^\top} H_{ivx}^\top \left[H_{ivx} (\operatorname{I_K} +\hat{\varpi}_b \hat{\varpi}_b^\top) H_{ivx}^\top \right]^{-1} H_{ivx} \mathbf{t_{ivx}} \nonumber\\
&=
\mathbf{t_{ivx}^\top} H_{ivx}^\top \left(H_{ivx} H_{ivx}^\top \right)^{-1} \left[H_{ivx} (\operatorname{I_K} +\hat{\varpi}_b \hat{\varpi}_b^\top) H_{ivx}^\top \right]
\left[H_{ivx} (\operatorname{I_K} +\hat{\varpi}_b \hat{\varpi}_b^\top) H_{ivx}^\top \right]^{-1}H_{ivx} \mathbf{t_{ivx}} \nonumber\\
&-\mathbf{t_{ivx}^\top} H_{ivx}^\top \left(H_{ivx} H_{ivx}^\top \right)^{-1} \left(H_{ivx} H_{ivx}^\top \right)
\left[H_{ivx} (\operatorname{I_K} +\hat{\varpi}_b \hat{\varpi}_b^\top) H_{ivx}^\top \right]^{-1} H_{ivx} \mathbf{t_{ivx}} \nonumber\\
&=
\mathbf{t_{ivx}^\top} H_{ivx}^\top \left(H_{ivx} H_{ivx}^\top \right)^{-1} \left[H_{ivx} (\operatorname{I_K} +\hat{\varpi}_b \hat{\varpi}_b^\top) H_{ivx}^\top - H_{ivx} H_{ivx}^\top\right]
\left[H_{ivx} (\operatorname{I_K} +\hat{\varpi}_b \hat{\varpi}_b^\top) H_{ivx}^\top \right]^{-1}H_{ivx} \mathbf{t_{ivx}} \nonumber\\
&=
\mathbf{t_{ivx}^\top} H_{ivx}^\top \left(H_{ivx} H_{ivx}^\top \right)^{-1} \left(H_{ivx} \hat{\varpi}_b \hat{\varpi}_b^\top H_{ivx}^\top \right)
\left[H_{ivx} (\operatorname{I_K} +\hat{\varpi}_b \hat{\varpi}_b^\top) H_{ivx}^\top \right]^{-1}H_{ivx} \mathbf{t_{ivx}}>0.
\end{align}
The last line of equation ((ref)) holds since $\left(H_{ivx} H_{ivx}^\top \right)^{-1} $,
$H_{ivx} \hat{\varpi}_b \hat{\varpi}_b^\top H_{ivx}^\top$ and
$\left[H_{ivx} (\operatorname{I_K} +\hat{\varpi}_b \hat{\varpi}_b^\top) H_{ivx}^\top \right]^{-1}$ are positive definition.
And
\begin{align}
&\mathbf{t_{ivx}^\top} H_{ivx}^\top \left(H_{ivx} H_{ivx}^\top \right)^{-1}
H_{ivx} \hat{\varpi}_b \hat{\varpi}_b^\top H_{ivx}^\top
\left[H_{ivx} (\operatorname{I_K} +\hat{\varpi}_b \hat{\varpi}_b^\top) H_{ivx}^\top \right]^{-1} H_{ivx} \mathbf{t_{ivx}} \\
& - \mathbf{t_{ivx}^\top} H_{ivx}^\top \left(H_{ivx} H_{ivx}^\top \right)^{-1}
H_{ivx} \hat{\varpi}_b \hat{\varpi}_b^\top H_{ivx}^\top
\left(H_{ivx} H_{ivx}^\top \right)^{-1} H_{ivx} \mathbf{t_{ivx}} \nonumber\\
&= \mathbf{t_{ivx}^\top} H_{ivx}^\top \left(H_{ivx} H_{ivx}^\top \right)^{-1}
H_{ivx} \hat{\varpi}_b \hat{\varpi}_b^\top H_{ivx}^\top \left(H_{ivx} H_{ivx}^\top \right)^{-1} \left(H_{ivx} H_{ivx}^\top \right)
\left[H_{ivx} (\operatorname{I_K} +\hat{\varpi}_b \hat{\varpi}_b^\top) H_{ivx}^\top \right]^{-1} H_{ivx} \mathbf{t_{ivx}} \nonumber \\
& - \mathbf{t_{ivx}^\top} H_{ivx}^\top \left(H_{ivx} H_{ivx}^\top \right)^{-1}
H_{ivx} \hat{\varpi}_b \hat{\varpi}_b^\top H_{ivx}^\top
\left(H_{ivx} H_{ivx}^\top \right)^{-1} \left[H_{ivx} (\operatorname{I_K} +\hat{\varpi}_b \hat{\varpi}_b^\top) H_{ivx}^\top \right] \nonumber\\
& \left[H_{ivx} (\operatorname{I_K} +\hat{\varpi}_b \hat{\varpi}_b^\top) H_{ivx}^\top \right]^{-1} H_{ivx} \mathbf{t_{ivx}} \nonumber\\
&= - \mathbf{t_{ivx}^\top} H_{ivx}^\top \left(H_{ivx} H_{ivx}^\top \right)^{-1}
H_{ivx} \hat{\varpi}_b \hat{\varpi}_b^\top H_{ivx}^\top
\left(H_{ivx} H_{ivx}^\top \right)^{-1} \left[H_{ivx} (\operatorname{I_K} +\hat{\varpi}_b \hat{\varpi}_b^\top) H_{ivx}^\top
- H_{ivx} H_{ivx}^\top
\right] \nonumber\\
&\left[H_{ivx} (\operatorname{I_K} +\hat{\varpi}_b \hat{\varpi}_b^\top) H_{ivx}^\top \right]^{-1} H_{ivx} \mathbf{t_{ivx}} \nonumber\\
&= - \mathbf{t_{ivx}^\top} H_{ivx}^\top \left(H_{ivx} H_{ivx}^\top \right)^{-1}
H_{ivx} \hat{\varpi}_b \hat{\varpi}_b^\top H_{ivx}^\top
\left(H_{ivx} H_{ivx}^\top \right)^{-1} \left[H_{ivx} \hat{\varpi}_b \hat{\varpi}_b^\top H_{ivx}^\top\right] \nonumber\\
&\left[H_{ivx} (\operatorname{I_K} +\hat{\varpi}_b \hat{\varpi}_b^\top) H_{ivx}^\top \right]^{-1} H_{ivx} \mathbf{t_{ivx}} \nonumber\\
& = O_p( \hat{\varpi}_b \hat{\varpi}_b^\top ) = o_p\left[T^{-(1-\delta)/2}\right]\nonumber.
\end{align}
The last line holds of equation ((ref)) since $ \hat{\varpi}_b=O_p\left[T^{-(1-\delta)/2}\right]$. By equations ((ref)) and ((ref)), Theorem (ref) holds.
proof[Proof of equation ((ref))]
We first consider the SD case for $S_a$. By the definition of $S_a$ and $S_b$ and part (i) of Lemma B1 of Online Appendix for Kostakisetal2015 , it follows that
\begin{align}
&S_a = \frac{1}{T}\sum\limits_{t=1}^{T} z_{t-1}\left(\frac{1}{T_0}\sum\limits_{t=1}^{T_0} z_{t-1}^\top \right) \left(\frac{1}{T_0}\sum\limits_{t=1}^{T_0} z_{t-1}^\top \frac{1}{T_0}\sum\limits_{t=1}^{T_0} z_{t-1}\right)^{-1}\\
&= \lambda \sum\limits_{t=1}^{T} z_{t-1}\left( \sum\limits_{t=1}^{T_0} z_{t-1}^\top \right) \left( \sum\limits_{t=1}^{T_0} z_{t-1}^\top \sum\limits_{t=1}^{T_0} z_{t-1}\right)^{-1}\nonumber\\
&= \lambda \frac{1}{T^{1/2+\delta}}\sum\limits_{t=1}^{T} z_{t-1}\left( \frac{1}{T^{1/2+\delta}}\sum\limits_{t=1}^{T_0} z_{t-1}^\top \right) \left( \frac{1}{T^{1/2+\delta}}\sum\limits_{t=1}^{T_0} z_{t-1}^\top \frac{1}{T^{1/2+\delta}}\sum\limits_{t=1}^{T_0} z_{t-1}\right)^{-1}\nonumber\\
&= \lambda \frac{1}{T^{1/2+\delta}}\sum\limits_{t=1}^{T} z_{t-1}\left( \frac{1}{T^{1/2+\delta}}\sum\limits_{t=1}^{T_0} z_{t-1}^\top \right) \left( \frac{1}{T^{1/2+\delta}}\sum\limits_{t=1}^{T_0} z_{t-1}^\top \frac{1}{T^{1/2+\delta}}\sum\limits_{t=1}^{T_0} z_{t-1}\right)^{-1}\nonumber\\
&= \lambda \frac{1}{T^{1/2}}x_{T}\frac{1}{T^{1/2}}x_{T_0}^\top \left( \frac{1}{T^{1/2}}x_{T_0}^\top \frac{1}{T^{1/2}}x_{T_0} \right)^{-1}+o_p(1) \nonumber\\
&\Rightarrow \tilde{S}_a= \lambda J_x^c(1) J_x^c(\lambda)^\top \left[ J_x^c(\lambda)^\top J_x^c(\lambda) \right]^{-1}.\nonumber
\end{align}
Similarly,
\begin{align}
&S_b = \frac{1}{T}\sum\limits_{t=1}^{T} z_{t-1}\left(\frac{1}{T-T_0}\sum\limits_{t=T_0+1}^{T} z_{t-1}^\top \right) \left(\frac{1}{T-T_0}\sum\limits_{t=T_0+1}^{T} z_{t-1}^\top \frac{1}{T-T_0}\sum\limits_{t=T_0+1}^{T} z_{t-1}\right)^{-1} \nonumber\\
&= \lambda \sum\limits_{t=1}^{T} z_{t-1}\left( \sum\limits_{t=T_0+1}^{T} z_{t-1}^\top \right) \left( \sum\limits_{t=T_0+1}^{T} z_{t-1}^\top \sum\limits_{t=T_0+1}^{T} z_{t-1}\right)^{-1}\nonumber\\
&= \lambda \frac{1}{T^{1/2+\delta}}\sum\limits_{t=1}^{T} z_{t-1}\left( \frac{1}{T^{1/2+\delta}}\sum\limits_{t=T_0+1}^{T} z_{t-1}^\top \right) \left( \frac{1}{T^{1/2+\delta}}\sum\limits_{t=T_0+1}^{T} z_{t-1}^\top \frac{1}{T^{1/2+\delta}}\sum\limits_{t=T_0+1}^{T} z_{t-1}\right)^{-1}\nonumber\\
&= \lambda \frac{1}{T^{1/2+\delta}}\sum\limits_{t=1}^{T} z_{t-1}\left( \frac{1}{T^{1/2+\delta}}\sum\limits_{t=T_0+1}^{T} z_{t-1}^\top \right) \left( \frac{1}{T^{1/2+\delta}}\sum\limits_{t=T_0+1}^{T} z_{t-1}^\top \frac{1}{T^{1/2+\delta}}\sum\limits_{t=T_0+1}^{T} z_{t-1}\right)^{-1}\nonumber\\
&= \lambda \frac{1}{T^{1/2}}x_{T}\frac{1}{T^{1/2}}(x_{T}-x_{T_0})^\top \left[ \frac{1}{T^{1/2}}(x_{T}-x_{T_0})^\top \frac{1}{T^{1/2}}(x_{T}-x_{T_0}) \right]^{-1} + o_p(1)\nonumber\\
&\Rightarrow \tilde{S}_b= \lambda J_x^c(1) \left[ J_x^c(1) - J_x^c(\lambda) \right]^\top \left\{ \left[ J_x^c(1) - J_x^c(\lambda) \right]^\top \left[ J_x^c(1) - J_x^c(\lambda) \right] \right\}^{-1}.\nonumber
\end{align}
Second, by equation (7) of the supplementary material for Kostakisetal2015, for WD predictors, we have
\begin{align}
\frac{1}{\sqrt{T}}\sum\limits_{t=1}^{T} z_{i,t-1} = \frac{1}{\sqrt{T}}\sum\limits_{t=1}^{T} x_{i,t-1} +o_p(1),\\ \frac{1}{\sqrt{T}}\sum\limits_{t=1}^{T_0} z_{i,t-1} = \frac{1}{\sqrt{T}}\sum\limits_{t=1}^{T_0} x_{i,t-1} +o_p(1)\nonumber
\end{align}
By equation ((ref)), it follows that
\begin{align}
\frac{1}{\sqrt{T}}\sum\limits_{t=1}^{T} x_{i,t-1} = \rho_i \frac{1}{\sqrt{T}}\sum\limits_{t=2}^{T} x_{i,t-2} +\frac{1}{\sqrt{T}}\sum\limits_{t=1}^{T} v_{i,t}.
\end{align}
So
\begin{align}
\frac{1}{\sqrt{T}}\sum\limits_{t=1}^{T} x_{i,t-1} = \frac{1}{1-\rho_i}\frac{1}{\sqrt{T}}\sum\limits_{t=1}^{T} v_{i,t} + o_p(1)\Rightarrow B_{v_i}(1)/(1-\rho_i).\\
\frac{1}{\sqrt{T}}\sum\limits_{t=1}^{T_0} x_{i,t-1} = \frac{1}{1-\rho_i}\frac{1}{\sqrt{T}}\sum\limits_{t=1}^{T_0} v_{i,t}+ o_p(1)\Rightarrow B_{v_i}(\lambda)/(1-\rho_i), \nonumber
\end{align}
where $B_v(\lambda)=\left[B_{v_1}(\lambda),B_{v_2}(\lambda),\cdots,B_{v_K}(\lambda)\right]^\top$.
By equations ((ref)), ((ref)), ((ref)) and ((ref)), for WD predictors, we have
\begin{align*}
S_a \Rightarrow \tilde{S}_a =
B_v(1) B_v(\lambda)^\top \left[ B_v(\lambda)^\top B_v(\lambda) \right]^{-1}.
\end{align*}
Similarly, for WD predictors, by the continuous mapping theorem and the definition of $S_a$ and $S_b$ , it follows that
\begin{align}
S_b \Rightarrow \tilde{S}_b= B_v(1) \left[B_v(1) - B_v(\lambda) \right]^\top \left\{ \left[B_v(1) - B_v(\lambda) \right]^\top \left[B_v(1) - B_v(\lambda) \right] \right\}^{-1}.\nonumber
\end{align}
lemmaUnder Assumption (ref), for SD and WD predictors, it follows that
\begin{align}
D_T^{-2} \sum\limits_{t=1}^{T} \tilde{z}_{t-1} x_{t-1}^\top \Rightarrow \Sigma_{zx}
\end{align}
proof[Proof of Lemma (ref)]
By the definition of $\tilde{z}_{t-1}$ in equation ((ref)), it follows that
\begin{align}
D_T^{-2} \sum\limits_{t=1}^{T} \tilde{z}_{t-1} x_{t-1}^\top = \left(\operatorname{I_K}-S_a \right) D_T^{-2}\sum_{t=1}^{T_0} z_{t-1} x_{t-1}^\top+ \left(\operatorname{I_K}-S_b \right) D_T^{-2}\sum_{t=T_0+1}^{T} z_{t-1} x_{t-1}^\top.
\end{align}
Also
\begin{align}
x_{t-1} = x_{t-2} + \Delta x_{t-1}.
\end{align}
The result of equation ((ref)) times equation ((ref)) is as follows.
\begin{align*}
&\frac{1}{T} \sum_{t=1}^{T_0} z_{t-1}x_{t-1}^\top \\
&= \rho_z \frac{1}{T} \sum_{t=1}^{T_0}z_{t-2} x_{t-2}^\top+ \rho_z \frac{1}{T} \sum_{t=1}^{T_0}z_{t-2} \Delta x_{t-1}^\top + \frac{1}{T} \sum_{t=2}^{T_0} \Delta x_{t-1} x_{t-2}^\top + \frac{1}{T}\sum_{t=1}^{T_0} \Delta x_{t-1} \Delta x_{t-1}^\top \nonumber
\end{align*}
Since $\frac{1}{T^{(1+\delta)/2}} \sum_{t=1}^{T_0}z_{t-2} \Delta x_{t-1}^\top =O_p(1)$, then
\begin{align}
(1-\rho_z)\frac{1}{T} \sum_{t=1}^{T_0} z_{t-1}x_{t-1}^\top
&= \frac{1}{T} \sum_{t=1}^{T_0}\Delta x_{t-1} x_{t-2}^\top + \frac{1}{T}\sum_{t=2}^{T_0} \Delta x_{t-1} \Delta x_{t-1}^\top + o_p(1). \nonumber
\end{align}
As a result,
\begin{align}
&(\operatorname{I_K}-S_a)(1-\rho_z)\frac{1}{T} \sum_{t=1}^{T_0} z_{t-1}x_{t-1}^\top \\
&= (\operatorname{I_K}-S_a)\frac{1}{T} \sum_{t=1}^{T_0} \Delta x_{t-1} x_{t-2}^\top + (\operatorname{I_K}-S_a) \frac{1}{T}\sum_{t=1}^{T_0} \Delta x_{t-1} \Delta x_{t-1}^\top + o_p(1). \nonumber\\
& \Rightarrow (\operatorname{I_K}-\tilde{S}_a)\int_0^{\lambda} d J_x^c(r)J_x^c(r)^\top + \lambda (\operatorname{I_K}-\tilde{S}_a) \operatorname{E}(v_tv_t^\top). \nonumber
\end{align}
Similarly,
\begin{align}
& (\operatorname{I_K}-S_b)(1-\rho_z)\frac{1}{T} \sum_{t=T_0+1}^{T} z_{t-1}x_{t-1}^\top \\
&= (\operatorname{I_K}-S_b)\frac{1}{T} \sum_{t=T_0+1}^{T} \Delta x_{t-1} x_{t-2}^\top + (\operatorname{I_K}-S_b) \frac{1}{T}\sum_{t=T_0+1}^{T} \Delta x_{t-1} \Delta x_{t-1}^\top + o_p(1) \nonumber\\
& \Rightarrow (\operatorname{I_K}-\tilde{S}_a)\int_0^{\lambda} d J_x^c(r) J_x^c(r)^\top + (1 -\lambda) (\operatorname{I_K}-\tilde{S}_b) \nonumber \operatorname{E}(v_tv_t^\top).
\end{align}
By equations ((ref)), ((ref)) and ((ref)), for SD predictors, we have
\begin{align}
D_T^{-2} \sum\limits_{t=1}^{T} \tilde{z}_{t-1} x_{t-1}^\top \xrightarrow{P} \Sigma_{zx}
\end{align}
By the similar procedure, we have the following equation for WD predictors.
\begin{align}
\frac{1}{T} \sum_{t=1}^{T_0} \tilde{z}_{t-1}x_{t-1}^\top & =(\operatorname{I_K}-S_a) \frac{1}{T} \sum_{t=1}^{T_0} z_{t-1}x_{t-1}^\top
= (\operatorname{I_K}-S_a)\frac{1}{T} \sum_{t=1}^{T_0} x_{t-1} x_{t-1}^\top + o_p(1) \\
&\Rightarrow \lambda(\operatorname{I_K}-\tilde{S}_a) \operatorname{E}\left( x_{t-1}x_{t-1}^\top\right) \nonumber
\end{align}
and
\begin{align}
\frac{1}{T} \sum_{t=T_0+1}^{T} \tilde{z}_{t-1}x_{t-1}
&= (\operatorname{I_K}-S_b) \frac{1}{T} \sum_{t=T_0+1}^{T} z_{t-1}x_{t-1}
= (\operatorname{I_K}-S_b)\frac{1}{T} \sum_{t=T_0+1}^{T} x_{t-1} x_{t-1}^\top + o_p(1) \\
&\Rightarrow \lambda(1-\tilde{S}_b) \operatorname{E}\left( x_{t-1}x_{t-1}^\top\right) \nonumber
\end{align}
By equations ((ref)), ((ref)) and ((ref)), for WD predictors, we also have
\begin{align}
D_T^{-2} \sum\limits_{t=1}^{T} \tilde{z}_{t-1} x_{t-1}^\top \xrightarrow{P} \Sigma_{zx}
\end{align}
By equations ((ref)) and ((ref)), Lemma (ref) holds.
proof[Proof of Theorem (ref)]
The illustration of the joint convergence between $D_T^{-2} \sum\limits_{t=1}^{T} \tilde{z}_{t-1} x_{t-1} $ and $D_T^{-1} \sum\limits_{t=1}^{T} \tilde{z}_{t-1} u_t$ is the same as the last part of the proof of Lemma (ref). Therefore, Theorem (ref) holds by Lemma (ref) and Lemma (ref).
The proof of equations ((ref)) and ((ref)) is very similar to the proof of Proposition (ref), so it is omitted here.
lemmaUnder Assumption (ref), for SD and WD predictors, we have
$$D_T^{-2}\sum_{t=1}^T \tilde{z}_{t-1} \tilde{z}_{t-1}^\top \hat{u}_t^2\Rightarrow \Sigma_{zz}.$$
proof[Proof of Lemma (ref)]
Following the similar procedure for the stability condition shown in equation ((ref)) for $\{ \frac{1}{T^{(1+\delta)/2}} \breve{z}_{t-1} u_t\}_{t=1}^T$ of proof of Lemma (ref), Lemma (ref) is proved.
proof[Proof of Proposition (ref)]
This proof is quite similar to the proof of Proposition 1 of HosseinkouchackDemetrescu2021.
By Lemma (ref), it follows that
\begin{align}
\mathbf{t_l} &\equiv \left( \sum_{t=1}^T \tilde{z}_{t-1} \tilde{z}_{t-1}^\top \hat{u}_t^2 \right)^{-1/2} \sum_{t=1}^T \tilde{z}_{t-1} u_t \\
&=\left( \frac{1}{T^{1+\delta}} \sum_{t=1}^T \tilde{z}_{t-1} \tilde{z}_{t-1}^\top \hat{u}_t^2 \right)^{-1/2}\frac{1}{T^{1 / 2+\delta / 2}} \sum_{t=1}^T \tilde{z}_{t-1} u_t \nonumber\\
&=\left[ \frac{1}{T^{1+\delta}} \sum_{t=1}^T \tilde{z}_{t-1}^* (\tilde{z}_{t-1}^*)^\top \hat{u}_t^2 \right]^{-1/2}\frac{1}{T^{1 / 2+\delta / 2}} \sum_{t=1}^T \tilde{z}_{t-1}^* u_t \nonumber\\
&= \left\{\operatorname{I_K} -\frac{1}{2} \left(\frac{1}{T^{1+\delta }} \sum_{t=1}^T \tilde{z}_{t-1}^*(\tilde{z}_{t-1}^*)^\top u_t^2 -\operatorname{I_K}\right)+o_p\left[ T^{-(1-\delta)/2} \right] \right\}
\frac{1}{T^{1 / 2+\delta / 2}} \sum_{t=1}^T \tilde{z}_{t-1}^* u_t \nonumber \\
&= \frac{1}{T^{1 / 2+\delta / 2}} \sum_{t=1}^T \tilde{z}_{t-1}^* u_t -\frac{1}{2} \left(\frac{1}{T^{1+\delta }} \sum_{t=1}^T \tilde{z}_{t-1}^*(\tilde{z}_{t-1}^*)^\top u_t^2 -\operatorname{I_K}\right)\frac{1}{T^{1 / 2+\delta / 2}} \sum_{t=1}^T \tilde{z}_{t-1}^* u_t +o_p\left[ T^{-(1-\delta)/2} \right]\nonumber\\
&=Z_T^l+B_T^l+o_p\left[ T^{-(1-\delta)/2} \right],
\nonumber
\end{align}
where
$Z_T^l = \left[ \Sigma_{zz} \right]^{-1/2} \frac{1}{T^{(1+\delta)/2}}\sum\nolimits_{t=1}^T \tilde{z}_{t-1} u_t$, $B_T^l = \varpi_l Z_T^l$ and $\varpi_l = -\frac{1}{2}\left[\sum\nolimits_{t=1}^T \tilde{z}_{t-1}^* (\tilde{z}_{t-1}^*)^\top \hat{u}_t^2 -1\right]$.
By Lemma (ref) and Lemma (ref), we have
\begin{align}
Z_T^l\xrightarrow{d}\operatorname{N}(0_K,\operatorname{I_K}).
\end{align}
By equations ((ref)), ((ref)) and ((ref)), it follows that
\begin{small}
\begin{align}
&Z_T^l - W_a Z_T^{p_a} - W_b Z_T^{p_b} = \Sigma_{zz}^{-1/2}\frac{1}{T^{(1+\delta)/2}} \sum\nolimits_{t=1}^T \tilde{z}_{t-1} u_t \\
& - \left( \frac{\sum_{t=1}^{T} \tilde{z}_{t-1} \tilde{z}_{t-1}^\top \hat{u}_t^2}{T^{1+\delta}} \right)^{-1/2}
(\operatorname{I_K}-S_a)
\left( \frac{\sum_{t=1}^{T_0} z_{t-1} z_{t-1}^\top \hat{u}_t^2}{T^{1+\delta}} \right)^{1/2}
\left( \frac{\sum_{t=1}^{T_0} z_{t-1} z_{t-1}^\top \hat{u}_t^2 }{T^{1+\delta}} \right)^{-1/2}
\frac{ \sum\nolimits_{t=1}^{T_0} z_{t-1} u_t }{T^{(1+\delta)/2}}
\nonumber\\
&-\left( \frac{\sum_{t=1}^{T} \tilde{z}_{t-1} \tilde{z}_{t-1}^\top \hat{u}_t^2}{T^{1+\delta}} \right)^{-1/2}
(\operatorname{I_K}-S_b)
\left( \frac{\sum_{t=T_0+1}^{T} z_{t-1} z_{t-1}^\top \hat{u}_t^2}{T^{1+\delta}} \right)^{1/2} \nonumber\\
&\cdot \left( \frac{ \sum_{t=T_0+1}^{T} z_{t-1} z_{t-1}^\top \hat{u}_t^2}{T^{1+\delta}} \right)^{-1/2}
\frac{ \sum\nolimits_{t=T_0+1}^T z_{t-1} u_t }{T^{(1+\delta)/2}}
\nonumber\\
&= \Sigma_{zz}^{-1/2}\frac{1}{T^{(1+\delta)/2}} \sum\nolimits_{t=1}^T \tilde{z}_{t-1} u_t - \left( \frac{\sum_{t=1}^{T} \tilde{z}_{t-1} \tilde{z}_{t-1}^\top \hat{u}_t^2}{T^{1+\delta}} \right)^{-1/2}
(\operatorname{I_K}-S_a)
\frac{ \sum\nolimits_{t=1}^{T_0} z_{t-1} u_t }{T^{(1+\delta)/2}}
\nonumber\\
&-\left( \frac{\sum_{t=1}^{T} \tilde{z}_{t-1} \tilde{z}_{t-1}^\top \hat{u}_t^2}{T^{1+\delta}} \right)^{-1/2}
(\operatorname{I_K}-S_b)
\frac{ \sum\nolimits_{t=T_0+1}^T z_{t-1} u_t }{T^{(1+\delta)/2}}
\nonumber
\end{align}
\end{small}
\begin{small}
\begin{align}
&= \Sigma_{zz}^{-1/2}\frac{1}{T^{(1+\delta)/2}} \sum\nolimits_{t=1}^T \tilde{z}_{t-1} u_t - \left( \frac{\sum_{t=1}^{T} \tilde{z}_{t-1} \tilde{z}_{t-1}^\top \hat{u}_t^2}{T^{1+\delta}} \right)^{-1/2}
\frac{ \sum\nolimits_{t=1}^{T_0} \tilde{z}_{t-1} u_t }{T^{(1+\delta)/2}}
\nonumber\\
&-\left( \frac{\sum_{t=1}^{T} \tilde{z}_{t-1} \tilde{z}_{t-1}^\top \hat{u}_t^2}{T^{1+\delta}} \right)^{-1/2}
\frac{ \sum\nolimits_{t=T_0+1}^T z_{t-1} u_t }{T^{(1+\delta)/2}}
\nonumber\\
&= \left[ \Sigma_{zz}^{-1/2} - \left( \frac{\sum_{t=1}^{T} \tilde{z}_{t-1} \tilde{z}_{t-1}^\top \hat{u}_t^2}{T^{1+\delta}} \right)^{-1/2} \right]
\frac{ \sum\nolimits_{t=1}^{T_0} \tilde{z}_{t-1} u_t }{T^{(1+\delta)/2}} - \left[ \Sigma_{zz}^{-1/2} -\left( \frac{\sum_{t=1}^{T} \tilde{z}_{t-1} \tilde{z}_{t-1}^\top \hat{u}_t^2}{T^{1+\delta}} \right)^{-1/2}\right]
\frac{ \sum\nolimits_{t=T_0+1}^T z_{t-1} u_t }{T^{(1+\delta)/2}}
\nonumber\\
&= o_p\left[T^{(1-\delta)/2}\right]O_p(1)+ o_p\left[T^{(1-\delta)/2}\right]O_p(1) = o_p\left[T^{(1-\delta)/2}\right]\nonumber
\end{align}
\end{small}
The last step holds since $\frac{1}{T^{1+\delta}}\sum_{t=1}^{T_0} \tilde{z}_{t-1} \tilde{z}_{t-1}^\top \hat{u}_t^2-\Sigma_{zz}=o_p\left[T^{(1-\delta)/2}\right]$ and $\frac{1}{T^{1+\delta}}\sum_{t=T_0+1}^{T} \tilde{z}_{t-1} \tilde{z}_{t-1}^\top \hat{u}_t^2-\Sigma_{zz}=o_p\left[T^{(1-\delta)/2}\right]$ by the proof of Proposition 1 of HosseinkouchackDemetrescu2021.
Additionally, by equations ((ref)), ((ref)), ((ref)) and ((ref)), it follows that
\begin{align}
B_T^l &= t_l - Z_T^l +o_p\left[ T^{-(1-\delta)/2} \right]
= W_a t_a^p + W_b t_b^p - Z_T^l +o_p\left[ T^{-(1-\delta)/2} \right] \\
& = (W_a Z_T^{p_a} + W_b Z_T^{p_b} -Z_T^l) + W_a B_T^{p_a}+ W_b B_T^{p_b} + o_p\left[T^{-(1-\delta)/2}\right] \nonumber\\
& = W_a B_T^{p_a}+ W_b B_T^{p_b} + o_p\left[T^{-(1-\delta)/2}\right]. \nonumber
\end{align}
The last step holds by equation ((ref)). By equation ((ref))
\begin{align}
R_T^l&= T^{(1 -\delta) / 2} B_T^l - W_l \operatorname{plim}_{T\rightarrow \infty} \; T^{(1 -\delta) / 2} \operatorname{E}\left(B_T\right) \\
& = W_a B_T^{p_a}+ W_b B_T^{p_b}- \left[W_a/\sqrt{\lambda} + W_b/\sqrt{1-\lambda} \right]\operatorname{plim}_{T\rightarrow \infty} \; T^{(1 -\delta) / 2} \operatorname{E}\left(B_T\right) + o_p\left[T^{-(1-\delta)/2}\right] \nonumber\\
& = W_a B_T^{p_a}+ W_b B_T^{p_b}- \left[W_a/\sqrt{\lambda} + W_b/\sqrt{1-\lambda} \right] \left[- \frac{K+1}{2}{\rho_{u v^*}}/ \sqrt{-2 c_z} \right] + o_p\left[T^{-(1-\delta)/2}\right] \nonumber\\
&=W_a R_{1,T}^l+W_b R_{2,T}^l \nonumber
\end{align}
where $ R_{1,T}^l=T^{(1 -\delta) / 2}\left[ B_T^{p_a}- \operatorname{E}(B_T^{p_a}) \right]$ and $ R_{2,T}^l=T^{(1 -\delta) / 2}\left[ B_T^{p_b}- \operatorname{E}(B_T^{p_b}) \right]$.
By equations ((ref)), ((ref)), ((ref)) and ((ref)), Proposition (ref) holds.
proof[Proof of Theorem (ref)]
When $J=K=1$, by equations ((ref)) and ((ref)), we have
\begin{align}
Q_l^t&= \frac{\sum_{t=1}^T \tilde{z}_{t-1}y_t}{\sqrt{\sum_{t=1}^T \tilde{z}_{t-1}\tilde{z}_{t-1}^\top \hat{u}_t^2}} = \frac{\sum_{t=1}^T \tilde{z}_{t-1}(\mu + x_{t-1}^\top \beta + u_t)}{\sqrt{\sum_{t=1}^T \tilde{z}_{t-1}\tilde{z}_{t-1}^\top \hat{u}_t^2}} \\
&= \frac{\sum_{t=1}^T \tilde{z}_{t-1} x_{t-1}^\top \beta }{\sqrt{\sum_{t=1}^T \tilde{z}_{t-1}\tilde{z}_{t-1}^\top \hat{u}_t^2}}
+ \frac{\sum_{t=1}^T \tilde{z}_{t-1} u_t }{\sqrt{\sum_{t=1}^T \tilde{z}_{t-1}\tilde{z}_{t-1}^\top \hat{u}_t^2}} \nonumber \\
&=b \frac{D_T^{-2}\sum_{t=1}^T \tilde{z}_{t-1} x_{t-1}^\top }{\sqrt{D_T^{-2}\sum_{t=1}^T \tilde{z}_{t-1}\tilde{z}_{t-1}^\top \hat{u}_t^2}}
+ \frac{D_T^{-1}\sum_{t=1}^T \tilde{z}_{t-1} u_t }{\sqrt{D_T^{-2}\sum_{t=1}^T \tilde{z}_{t-1}\tilde{z}_{t-1}^\top \hat{u}_t^2}} \nonumber
\end{align}
By equation ((ref)) and Theorem (ref) and Lemma (ref) and (ref), Theorem (ref) holds.
proof[Proof of Theorem (ref)]
By equation ((ref)) and Lemma (ref) and (ref), it follows that
\begin{align}
D_T(\tilde{\beta}_m - \hat{\beta}_{l} ) = O_p\left[ T^{-(1-\delta)/2} \right].
\end{align}
By equation ((ref)) and Theorem (ref), Theorem (ref) holds.
proof[Proof of Proposition (ref)]
By the similar procedure of the proof of Theorem (ref) and the definitions of $ Q_l$ and ${\tilde{Q}_m}$, Proposition (ref) holds.
proof[Proof of Theorem (ref)]
By Theorem (ref) and Lemma (ref) and (ref), Theorem (ref) holds.
proof[Proof of Proposition (ref)]
For SD predictors,
by the definitions of $\tilde{\beta}_m$ and $\hat{\beta}_m$ in equations ((ref)) and ((ref)), $$\tilde{\beta}_m -\hat{\beta}_m = O_p({T^{-1}}). $$
Therefore, by the definitions of $Q_m$ and $\tilde{Q}_m$ , ${Q_m}={\tilde{Q}_m}+O_p(T^{-1})$. So
$$Q_l-Q_m= Q_l-\tilde{Q}_m + \tilde{Q}_m-Q_m= Q_l-\tilde{Q}_m+O_p(T^{-1}).$$
Similarly, for WD predictors, by $ W_z = o_p({T^{-1}})$,
$$Q_l-Q_m= O_p(T^{-1}).$$
By these two results and Proposition (ref), Proposition (ref) holds.
proof[Proof of Theorem (ref)]
By equation ((ref)), for SD and WD predictors, it follows that
\begin{align}
D_T(\tilde{\beta}_m- \beta)=D_T(\hat{\beta}_l- \beta)+o_p(1)\xrightarrow{d}
\operatorname{MN}\left[0,\Sigma_{zx}^{-1}\Sigma_{zz}\left(\Sigma_{zx}^{-1}\right)^\top \right], \quad SD
\end{align}
for SD and WD predictors.
By equation ((ref)) and Lemma (ref) and (ref), Theorem (ref) holds.
Additional Simulation Results
Example 1 (univariate models)
In this example for the univariate model, the DGP is set up by equations ((ref)), ((ref)) and ((ref)), in which $K=1$, $\mu=1$ and $\varphi_0=1$. In the univariate model, we define the notation $Q_m^l$ and $Q_m^t$ as the t-test statistics $t_l$ and $t_m$, respectively. To create the embedded endogeneity among innovations, the innovation processes are generated as $(\eta_t,\varepsilon_t)^\top \;\sim \; i.i.d. \, \operatorname{N}(0_{2\times 1},\Sigma_{2\times 2})$, where $\Sigma=\left(
array[array omitted — 38 chars of source]
\right)$. To save space, we only report the simulation results for GARCH model with the sample $T=750$. \footnote{To save space, we did not report the simulation results for GARCH model with the sample size $T=500$ and $T=250$ and ARCH and i.i.d. model with the sample size $T=750$, $T=500$ and $T=250$ in the paper. Their results are similar to the results of GARCH model with the sample size $T=750$. The codes and results are available upon request.}
First, the results for the comparison of the size performances of the proposed test statistics $t_l$ and $t_m$ for seven cases in GARCH(1,1) model are shown in Tables (ref)-(ref) while the power performances are shown in Figures (ref)-(ref). The size performance of two-sided test $H_0:\beta=0$ vs $H_a:\beta\neq 0$ and the right side test $H_0:\beta=0$ vs $H_a:\beta> 0$ and the left side test $H_0:\beta=0$ vs $H_a:\beta < 0$ are shown in Tables (ref)-(ref), respectively. Meanwhile, the power performances of two-sided test $H_0:\beta=0$ vs $H_a:\beta\neq 0$ and the right side test $H_0:\beta=0$ vs $H_a:\beta> 0$ and left side test $H_0:\beta=0$ vs $H_a:\beta < 0$ are shown in Figures (ref)-(ref), (ref)-(ref) and (ref)-(ref) respectively. In each hypothesis, we show the simulation results for $\phi=0.95$, $0.5$, $-0.1$ and $-0.95$. Additionally, six persistence including two categories are considered in each hypothesis. The first category including case 1 - case 4 is SD with $\alpha=1$, while $c_z=0,-5,-10,-15$ respectively. The second category including case 5 - case 6 is WD with $\alpha=0$, while $c_z=-0.05,-0.1,-0.15$, respectively. And we set $\lambda=0.1,0.25,0.5,0.75,0.9$ in Tables (ref)-(ref) to compare the size performance with different $\lambda$, $\beta= b/T^{(1+\alpha)/2}$ to see the local power and $\lambda=0.5$ in Figures (ref)-(ref), \footnote{Since the following part states that the proposed test statistics perform the best with $\lambda=0.5$, we only show the power performance with $\lambda=0.5$.} and $\varphi_1=0.1$, $\varphi_1=0.1$, $\bar{ \varphi}_1=0.85$ in GARCH(1,1) model and the nominal size to be $5\%$. Simulations are repeated $10,000$ times for each setting.
sidewaystable\caption{Size Performance ($\%$) of $t_l$ and $t_m$ for $H_0:\beta=0$ vs $H_a:\beta\neq 0$}
\resizebox{\textwidth}{!}{
\begin{tabular}{c|c|cccc|cccc|cccc|cccc|cccc}
\hline
& $\lambda$ & \multicolumn{4}{c|}{0.1} & \multicolumn{4}{c|}{0.25} & \multicolumn{4}{c|}{0.5} & \multicolumn{4}{c|}{0.75} & \multicolumn{4}{c}{0.9} \\
\hline
& $\phi$ & 0.95 & 0.5 & -0.1 & -0.95 & 0.95 & 0.5 & -0.1 & -0.95 & 0.95 & 0.5 & -0.1 & -0.95 & 0.95 & 0.5 & -0.1 & -0.95 & 0.95 & 0.5 & -0.1 & -0.95 \\
\hline
\multirow{7}[0]{*}{$t_l$} & Case 1 & 5.8 & 5.6 & 5.1 & 5.7 & 5.9 & 5.2 & 5.0 & 6.2 & 5.8 & 5.3 & 5.1 & 5.9 & 5.7 & 5.4 & 4.5 & 5.6 & 5.7 & 5.6 & 4.7 & 5.5 \\
& Case 2 & 5.5 & 5.4 & 4.4 & 5.5 & 5.3 & 5.1 & 4.8 & 5.4 & 5.7 & 4.9 & 5.0 & 5.4 & 5.6 & 5.3 & 5.2 & 5.5 & 5.3 & 4.9 & 4.9 & 5.6 \\
& Case 3 & 4.8 & 5.1 & 5.0 & 5.4 & 5.3 & 4.9 & 5.0 & 5.0 & 4.5 & 4.9 & 4.6 & 5.0 & 5.1 & 5.3 & 4.8 & 5.1 & 5.6 & 4.5 & 5.0 & 5.1 \\
& Case 4 & 5.0 & 5.0 & 4.5 & 5.4 & 5.1 & 4.9 & 5.1 & 5.0 & 5.1 & 4.8 & 5.0 & 5.2 & 4.8 & 4.7 & 4.9 & 5.0 & 5.4 & 4.8 & 4.7 & 5.7 \\
& Case 5 & 5.0 & 5.0 & 4.8 & 4.6 & 4.8 & 4.7 & 5.2 & 4.5 & 4.6 & 4.7 & 4.8 & 5.1 & 4.8 & 5.0 & 5.3 & 4.8 & 5.0 & 4.9 & 4.8 & 5.1 \\
& Case 6 & 4.5 & 4.5 & 4.9 & 4.5 & 4.3 & 5.2 & 4.9 & 4.8 & 4.8 & 4.9 & 4.6 & 4.5 & 4.7 & 5.0 & 4.6 & 4.6 & 5.2 & 4.7 & 5.0 & 4.3 \\
& Case 7 & 4.6 & 4.6 & 4.5 & 4.5 & 4.6 & 4.9 & 4.7 & 4.9 & 5.0 & 4.8 & 5.0 & 4.6 & 4.9 & 4.8 & 4.7 & 5.1 & 4.3 & 4.5 & 4.3 & 4.4 \\
\hline
\multirow{7}[0]{*}{$t_m$} & Case 1 & 3.8 & 4.3 & 4.0 & 3.6 & 4.3 & 4.2 & 4.3 & 4.3 & 4.2 & 4.3 & 4.4 & 4.2 & 3.9 & 4.4 & 3.9 & 3.6 & 3.6 & 4.4 & 3.8 & 3.8 \\
& Case 2 & 4.6 & 4.2 & 3.6 & 4.3 & 4.5 & 4.3 & 4.2 & 4.6 & 4.8 & 4.3 & 4.2 & 4.5 & 4.5 & 4.7 & 4.6 & 4.5 & 4.1 & 3.8 & 4.1 & 4.5 \\
& Case 3 & 4.3 & 4.3 & 4.2 & 4.7 & 4.8 & 4.2 & 4.3 & 4.4 & 4.2 & 4.4 & 4.0 & 4.3 & 4.6 & 4.5 & 4.1 & 4.3 & 4.6 & 3.9 & 4.3 & 4.3 \\
& Case 4 & 4.9 & 4.3 & 3.9 & 5.1 & 4.8 & 4.4 & 4.4 & 4.7 & 4.7 & 4.4 & 4.3 & 4.9 & 4.3 & 4.1 & 4.3 & 4.5 & 4.9 & 4.2 & 4.2 & 5.1 \\
& Case 5 & 5.2 & 4.9 & 4.7 & 4.8 & 5.0 & 4.7 & 5.1 & 4.7 & 4.8 & 4.6 & 4.7 & 5.2 & 4.8 & 5.0 & 5.1 & 4.9 & 5.1 & 4.8 & 4.7 & 5.2 \\
& Case 6 & 4.5 & 4.5 & 4.9 & 4.5 & 4.3 & 5.2 & 4.9 & 4.8 & 4.8 & 4.9 & 4.6 & 4.5 & 4.7 & 5.0 & 4.6 & 4.6 & 5.2 & 4.7 & 5.0 & 4.3 \\
& Case 7 & 4.6 & 4.6 & 4.5 & 4.5 & 4.6 & 4.9 & 4.7 & 4.9 & 5.0 & 4.8 & 5.0 & 4.6 & 4.9 & 4.8 & 4.7 & 5.1 & 4.3 & 4.5 & 4.3 & 4.4 \\
\hline
\end{tabular}
}
sidewaystable\caption{Size Performance ($\%$) of $t_l$ and $t_m$ for $H_0:\beta=0$ vs $H_a:\beta> 0$}
\resizebox{\textwidth}{!}{
\begin{tabular}{c|c|cccc|cccc|cccc|cccc|cccc}
\hline
& $\lambda$ & \multicolumn{4}{c|}{0.1} & \multicolumn{4}{c|}{0.25} & \multicolumn{4}{c|}{0.5} & \multicolumn{4}{c|}{0.75} & \multicolumn{4}{c}{0.9} \\
\hline
& $\phi$ & 0.95 & 0.5 & -0.1 & -0.95 & 0.95 & 0.5 & -0.1 & -0.95 & 0.95 & 0.5 & -0.1 & -0.95 & 0.95 & 0.5 & -0.1 & -0.95 & 0.95 & 0.5 & -0.1 & -0.95 \\
\hline
\multirow{7}[0]{*}{$t_l$} & Case 1 & 3.6 & 4.1 & 5.7 & 8.4 & 3.1 & 3.8 & 5.7 & 9.1 & 3.1 & 3.6 & 5.6 & 8.7 & 2.9 & 3.6 & 4.9 & 9.0 & 2.8 & 4.2 & 5.5 & 8.4 \\
& Case 2 & 4.0 & 4.4 & 4.6 & 7.1 & 4.0 & 4.1 & 5.0 & 6.8 & 3.6 & 4.1 & 5.2 & 7.2 & 3.4 & 4.3 & 5.3 & 7.6 & 3.6 & 3.9 & 5.2 & 7.5 \\
& Case 3 & 4.3 & 4.4 & 5.5 & 6.5 & 4.0 & 4.1 & 4.9 & 6.1 & 3.9 & 4.2 & 5.3 & 7.1 & 3.5 & 4.0 & 5.2 & 6.8 & 3.4 & 4.0 & 5.3 & 7.2 \\
& Case 4 & 4.3 & 4.5 & 5.3 & 6.3 & 3.9 & 4.3 & 5.3 & 6.3 & 3.8 & 4.2 & 5.3 & 6.7 & 3.5 & 4.1 & 5.2 & 6.8 & 3.6 & 3.9 & 5.0 & 7.1 \\
& Case 5 & 4.3 & 4.2 & 5.4 & 5.8 & 4.5 & 4.4 & 5.5 & 5.7 & 4.4 & 4.5 & 5.1 & 6.1 & 3.9 & 4.6 & 5.2 & 6.1 & 3.7 & 4.3 & 5.2 & 6.5 \\
& Case 6 & 4.6 & 4.3 & 4.9 & 5.4 & 4.5 & 4.9 & 5.4 & 5.4 & 4.2 & 4.7 & 4.9 & 5.3 & 4.5 & 4.6 & 5.0 & 5.5 & 4.2 & 4.3 & 5.2 & 5.6 \\
& Case 7 & 4.3 & 4.8 & 5.0 & 5.0 & 4.3 & 4.7 & 5.1 & 5.4 & 4.6 & 4.6 & 5.1 & 5.6 & 4.3 & 4.8 & 5.1 & 5.7 & 4.4 & 4.5 & 4.7 & 5.4 \\
\hline
\multirow{7}[0]{*}{$t_m$} & Case 1 & 4.3 & 4.3 & 4.6 & 4.6 & 4.1 & 4.3 & 4.9 & 5.1 & 4.2 & 4.0 & 4.8 & 5.1 & 4.0 & 4.3 & 4.1 & 5.0 & 3.6 & 4.5 & 4.6 & 4.4 \\
& Case 2 & 5.3 & 4.7 & 3.9 & 4.3 & 5.4 & 4.7 & 4.4 & 4.1 & 4.9 & 4.6 & 4.4 & 4.4 & 4.8 & 4.9 & 4.5 & 4.6 & 5.0 & 4.4 & 4.4 & 4.5 \\
& Case 3 & 5.5 & 4.8 & 4.7 & 4.4 & 5.4 & 4.4 & 4.2 & 4.0 & 5.0 & 4.6 & 4.6 & 4.6 & 4.6 & 4.4 & 4.5 & 4.6 & 4.5 & 4.3 & 4.6 & 4.8 \\
& Case 4 & 5.5 & 4.7 & 4.6 & 4.8 & 5.1 & 4.5 & 4.7 & 4.7 & 4.9 & 4.5 & 4.6 & 4.9 & 4.6 & 4.4 & 4.6 & 5.1 & 4.6 & 4.1 & 4.4 & 5.4 \\
& Case 5 & 4.7 & 4.4 & 5.3 & 5.6 & 4.9 & 4.5 & 5.5 & 5.5 & 4.7 & 4.7 & 5.0 & 5.9 & 4.3 & 4.8 & 5.1 & 6.0 & 4.0 & 4.5 & 5.1 & 6.3 \\
& Case 6 & 4.6 & 4.4 & 4.9 & 5.4 & 4.5 & 4.9 & 5.4 & 5.4 & 4.2 & 4.7 & 4.9 & 5.3 & 4.5 & 4.6 & 5.0 & 5.5 & 4.2 & 4.3 & 5.2 & 5.6 \\
& Case 7 & 4.3 & 4.8 & 5.0 & 5.0 & 4.3 & 4.7 & 5.1 & 5.4 & 4.6 & 4.6 & 5.1 & 5.6 & 4.3 & 4.8 & 5.1 & 5.7 & 4.4 & 4.5 & 4.7 & 5.4 \\
\hline
\end{tabular}
}
Clearly, the following findings can be observed from Tables (ref)-(ref). First, the proposed test statistics $t_l$ and $t_m$ generally perform well in terms of size for the joint test $H_0:\beta=0$ vs $H_a:\beta\neq 0$. $t_l$ performs well in terms of size for different $\lambda$ while $t_m$ suffers a little bit of size distortion with $\lambda=0.1,0.75,0.9$ for SD predictors. Second, the proposed test statistics $t_l$ suffers size distortion for the right side test $H_0:\beta=0$ vs $H_a:\beta> 0$ and left side test $H_0:\beta=0$ vs $H_a:\beta < 0$ with SD predictors and $\phi=0.95,0.5,-0.95$, while $t_m$ performs well in terms of size, especially with $\lambda=0.5$. In one word, $t_m$ with $\lambda=0.5$ performs well in terms of size for all cases while $t_l$ only performs well in terms of size for the joint test, which is according to the results in Section (ref) in the main text.
Next, it is clearly observed from Figures (ref)-(ref) that the power performance of the proposed test statistics $t_l$ and $t_m$ are quite well and comparable. To sum up, $t_m$ with $\lambda=0.5$ performs well in terms of size and power, which is according to the statements in Section (ref) that we reduce DiE and VEE while keeping its power performance.
figure[figure omitted — 919 chars of source]
figure[figure omitted — 918 chars of source]
figure[figure omitted — 925 chars of source]
figure[figure omitted — 926 chars of source]
figure[figure omitted — 941 chars of source]
figure[figure omitted — 940 chars of source]
figure[figure omitted — 947 chars of source]
figure[figure omitted — 948 chars of source]
figure[figure omitted — 935 chars of source]
figure[figure omitted — 934 chars of source]
figure[figure omitted — 941 chars of source]
figure[figure omitted — 942 chars of source]
Example 2
We observe the following findings from Tables (ref)-(ref), (ref)-(ref), (ref)-(ref) and (ref)-(ref). First, for the joint test $H_0:\beta=0$, the proposed test statistic ${Q_l}$ still suffers from multiple predictors ($K\geq 3$) and the size distortion grows bigger as the number of predictors K grows bigger. Meanwhile, the proposed test statistic ${Q_m}$ is free of size distortion for different settings with $\lambda=0.1,0.25,0.5,0.75,0.9$. With $\lambda=0.1$ and $K=6$, ${Q_m}$ tends to slightly under-reject the null hypothesis $H_0:\beta=0$. It is not surprising since ${Q_l}$ still suffers the DiE and the VEE while ${Q_m}$ does not. Second, the size performance of ${Q_m^t}$ is better than that of ${Q_l^t}$ for both two-sided and one-sided marginal tests although ${Q_m^t}$ suffers a little size distortion. In particular, the size performance of ${Q_l^t}$ and ${Q_m^t}$ are comparable for right side marginal test $H_0:\beta_{i}=0$. And the size performance of ${Q_m^t}$ is much better than that of ${Q_l^t}$ for two-sided marginal test $H_0:\beta_{i}=0$ vs $H_a:\beta_{i}\neq 0$ and left side marginal test $H_0:\beta_{i}=0$ vs $H_a:\beta_{i}< 0$. Third, for two-sided marginal test $H_0:\beta_{i}=0$ vs $H_a:\beta_{i}\neq 0$, right side marginal test $H_0:\beta_{i}=0$ vs $H_a:\beta_{i}> 0$ and left side marginal test $H_0:\beta_{i}=0$ vs $H_a:\beta_{i}< 0$, the size performances of ${Q_l^t}$ and ${Q_m^t}$ become slightly better as $\lambda$ grows bigger while their power performances become worse. Therefore, we recommend the empirical researchers to set $\lambda=0.5$. Fourth, the power performances of ${Q_l^t}$ and ${Q_m^t}$ are comparable and quite well.
sidewaystable\caption{Size Performance ($\%$) of $t_l$ and $t_m$ for $H_0:\beta=0$ vs $H_a:\beta<0$}
\resizebox{\textwidth}{!}{
\begin{tabular}{c|c|cccc|cccc|cccc|cccc|cccc}
\hline
& $\lambda$ & \multicolumn{4}{c|}{0.1} & \multicolumn{4}{c|}{0.25} & \multicolumn{4}{c|}{0.5} & \multicolumn{4}{c|}{0.75} & \multicolumn{4}{c}{0.9} \\
\hline
& $\phi$ & 0.95 & 0.5 & -0.1 & -0.95 & 0.95 & 0.5 & -0.1 & -0.95 & 0.95 & 0.5 & -0.1 & -0.95 & 0.95 & 0.5 & -0.1 & -0.95 & 0.95 & 0.5 & -0.1 & -0.95 \\
\hline
\multirow{7}[0]{*}{$t_l$} & Case 1 & 8.6 & 6.9 & 4.8 & 3.6 & 8.9 & 6.8 & 4.6 & 3.4 & 8.4 & 7.1 & 4.6 & 2.9 & 8.6 & 7.0 & 4.6 & 2.9 & 9.0 & 6.9 & 4.6 & 3.2 \\
& Case 2 & 7.1 & 6.4 & 4.5 & 3.9 & 6.7 & 6.3 & 4.7 & 4.0 & 7.8 & 5.9 & 5.0 & 3.5 & 7.9 & 6.3 & 4.8 & 3.7 & 7.6 & 6.1 & 4.8 & 4.0 \\
& Case 3 & 5.8 & 6.2 & 4.8 & 4.1 & 6.6 & 6.0 & 4.9 & 4.1 & 6.3 & 6.2 & 4.5 & 3.6 & 7.0 & 6.5 & 4.8 & 3.6 & 7.6 & 5.8 & 4.5 & 3.5 \\
& Case 4 & 6.1 & 5.7 & 4.5 & 4.4 & 6.4 & 5.9 & 5.1 & 4.1 & 6.8 & 6.0 & 5.0 & 3.9 & 6.1 & 5.5 & 4.9 & 4.0 & 7.3 & 6.2 & 5.1 & 3.8 \\
& Case 5 & 5.9 & 5.6 & 4.8 & 4.3 & 5.3 & 5.2 & 5.0 & 4.2 & 5.7 & 5.7 & 5.0 & 4.5 & 6.4 & 5.8 & 5.0 & 3.9 & 6.9 & 5.9 & 5.0 & 4.1 \\
& Case 6 & 5.4 & 5.4 & 5.7 & 4.6 & 5.0 & 5.5 & 4.7 & 4.9 & 5.3 & 5.2 & 4.9 & 4.1 & 5.5 & 5.7 & 4.8 & 4.3 & 6.2 & 5.9 & 5.1 & 3.8 \\
& Case 7 & 5.2 & 4.8 & 5.0 & 4.7 & 5.1 & 5.5 & 4.8 & 4.8 & 5.1 & 5.3 & 4.8 & 4.3 & 5.6 & 5.4 & 4.8 & 4.7 & 5.6 & 5.2 & 4.9 & 4.2 \\
\hline
\multirow{7}[0]{*}{$t_m$} & Case 1 & 4.6 & 4.9 & 4.2 & 4.6 & 5.0 & 4.6 & 4.3 & 4.3 & 4.9 & 4.9 & 4.3 & 4.0 & 4.8 & 5.2 & 4.3 & 3.9 & 4.7 & 4.9 & 4.2 & 4.0 \\
& Case 2 & 4.2 & 4.7 & 4.1 & 5.2 & 4.0 & 4.6 & 4.2 & 5.4 & 4.7 & 4.3 & 4.9 & 4.9 & 4.8 & 4.5 & 4.4 & 5.0 & 4.5 & 4.3 & 4.3 & 5.1 \\
& Case 3 & 3.9 & 4.8 & 4.4 & 5.3 & 4.3 & 4.4 & 4.5 & 5.5 & 4.0 & 4.8 & 4.2 & 4.8 & 4.7 & 5.0 & 4.4 & 4.9 & 5.1 & 4.1 & 4.0 & 4.7 \\
& Case 4 & 4.4 & 4.8 & 4.0 & 5.4 & 4.7 & 4.8 & 4.8 & 5.1 & 5.2 & 4.8 & 4.7 & 5.0 & 4.6 & 4.4 & 4.5 & 5.0 & 5.3 & 4.9 & 4.7 & 4.8 \\
& Case 5 & 5.8 & 5.5 & 4.8 & 4.7 & 5.1 & 5.0 & 5.0 & 4.7 & 5.4 & 5.5 & 5.0 & 4.9 & 6.1 & 5.6 & 5.0 & 4.3 & 6.6 & 5.8 & 4.9 & 4.4 \\
& Case 6 & 5.4 & 5.4 & 5.7 & 4.6 & 5.0 & 5.5 & 4.7 & 4.9 & 5.3 & 5.2 & 4.9 & 4.1 & 5.5 & 5.7 & 4.8 & 4.3 & 6.2 & 5.9 & 5.1 & 3.8 \\
& Case 7 & 5.2 & 4.8 & 5.0 & 4.7 & 5.1 & 5.5 & 4.8 & 4.8 & 5.1 & 5.3 & 4.8 & 4.3 & 5.6 & 5.4 & 4.8 & 4.7 & 5.6 & 5.2 & 4.9 & 4.2 \\
\hline
\end{tabular}
}
table[table omitted — 1,741 chars of source]
table[table omitted — 1,744 chars of source]
table[table omitted — 1,745 chars of source]
table[table omitted — 1,744 chars of source]
table[table omitted — 1,870 chars of source]
table[table omitted — 1,768 chars of source]
table[table omitted — 1,870 chars of source]
table[table omitted — 1,869 chars of source]
table[table omitted — 1,864 chars of source]
table[table omitted — 1,865 chars of source]
table[table omitted — 1,865 chars of source]
table[table omitted — 1,864 chars of source]
table[table omitted — 1,871 chars of source]
table[table omitted — 1,873 chars of source]
table[table omitted — 1,869 chars of source]
table[table omitted — 1,871 chars of source]
table[table omitted — 1,871 chars of source]
\setcounter{table}{0}
\setcounter{figure}{0}
Additional Empirical Studies
In this section, we report the results of the two extra empirical studies. Here we give more details of the linear combination variables. The CP factor 2005Bond is the linear combination of the forward rate with $\mathrm{CP}_{t} = \hat{\gamma}_{cp}^{\top} \overrightarrow{\mathrm{\textbf{F}}}_{t}$. Here, $\overrightarrow{\mathrm{\textbf{F}}}_{t} = (\mathrm{F1}_{t},\mathrm{F2}_{t},\mathrm{F3}_{t},\mathrm{F4}_{t},\mathrm{F5}_{t})$ and $\hat{\gamma}_{cp}$ is the OLS estimator of the regression, $\frac{1}{4}\sum_{n=2}^{5}rx(n)_{t+1} = \gamma_{cp}^{\top} \overrightarrow{\mathrm{\textbf{F}}}_{t} + \varepsilon_{cp,t+1}$. The {LN1} factor and {LN2} factor ludvigson2009macro are the linear combination of the macroeconomic factors of $\overrightarrow{\mathrm{\textbf{Ma}}}_{t} = (\mathrm{M1}_{t}, \mathrm{M1}^{3}_{t}, \mathrm{M3}_{t}, \mathrm{M4}_{t}, \mathrm{M8}_{t})^{\top}$ and $\overrightarrow{\mathrm{\textbf{Mb}}}_{t} = (\mathrm{M1}_{t}, \mathrm{M1}^{3}_{t}, \mathrm{M2}_{t}, \mathrm{M3}_{t}, \mathrm{M4}_{t}, \mathrm{M8}_{t})^{\top}$ respectively with $\mathrm{LN1}_{t} = \hat{\vartheta}_{ln1}^{\top}\overrightarrow{\mathrm{\textbf{Ma}}}_{t}$ where $\hat{\vartheta}_{ln1}^{\top}$ is the OLS estimator of the regression $\frac{1}{4}\sum_{n=2}^{5}rx(n)_{t+1} = \vartheta_{ln1}^{\top}\overrightarrow{\mathrm{\textbf{Ma}}}_{t} + \varepsilon_{ln1,t+1}$ and $ \mathrm{LN2}_{t} = \hat{\vartheta}_{ln2}^{\top}\overrightarrow{\mathrm{\textbf{Mb}}}_{t}$ where $\hat{\vartheta}_{ln2}^{\top}$ is the OLS estimator of the regression $\frac{1}{4}\sum_{n=2}^{5}rx(n)_{t+1} = \vartheta_{ln2}^{\top} \overrightarrow{\mathrm{\textbf{Mb}}}_{t} + \varepsilon_{ln2,t+1}$. First, we show the comparison between the proposed test and ludvigson2009macro, whose data sample spanning the period 1964:01-2003:12 and regressors are the linear combinations CP and {LN1}, and CP and {LN2}, respectively. Second, we conduct the empirical study during 2020:01-2022:12 to see whether the predictability of bond risk premia varies during the COVID-19 epidemic period with the aggressive monetary policy. The results are shown in Table (ref).
Different from the original IVX and ludvigson2009macro, Table (ref) shows that the predicting power of CP, {LN1} and {LN2} are found by our method during 1964:01-2003:12 while not found during the COVID-19 epidemic period (2020:01-2022:12). The disappearance of predicting power of CP, {LN1} and {LN2} are most likely caused by the extensively loose monetary policy. Such monetary policy leads to the rapid growth of the debt and deficit of the U.S.; thus, the pricing of the bond is driven by the aggressively eased monetary policy rather than CP, {LN1} and {LN2}. As shown in Table (ref), during the COVID-19 epidemic period, the predicting power of CP, {LN1}, and {LN2} is not found by our method. On the contrary, IVX finds the CP can be useful in predicting bond risk premia, and CP, {LN1} and {LN2} are all found by ludvigson2009macro to be useful predictors except for the capacity of CP to predict $rx(4)$ and $rx(5)$ in Panel B.
table[table omitted — 2,522 chars of source]
To further check the robustness of the main results, we also do the tests of the predictability of F1-F5 and M1-M8 from 1964:01 to 2019:12. Results similar to Table (ref) in Section (ref) are observed in Table (ref). Therefore, the robustness of the main results in Section (ref) is confirmed.
table[table omitted — 2,307 chars of source]