The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
66,045 characters
Identification Robust Inference for the Risk Premium in Term Structure Models
\title{{\Large Identification Robust Inference for the Risk Premium in Term
Structure Models }}
\author{Frank Kleibergen\thanks{
Amsterdam School of Economics, University of Amsterdam, Roetersstraat 11,
1018 WB Amsterdam, The Netherlands. Email: [email removed].} \and
Lingwei Kong\thanks{
Faculty of Economics and Business, University of Groningen, Nettelbosje 2,
9747 AE Groningen, The Netherlands. Email: [email removed]. } }
\date{October 16, 2022}
\maketitle
\begin{abstract}
We propose identification robust statistics for testing hypotheses on the
risk premia in dynamic affine term structure models. We do so using the
moment equation specification proposed for these models in \nocite
{adrian2013pricing}Adrian et al. (2013). We extend the subset (factor)
Anderson-Rubin test from Guggenberger et al. (2012) to models with multiple
dynamic factors and time-varying risk prices. Unlike projection based tests,
it provides a computationally tractable manner to conduct identification
robust tests on a larger number of parameters. We analyze the potential
identification issues arising in empirical studies. Statistical inference
based on the three-stage estimator from Adrian et al. (2013) requires
knowledge of the factors' quality and is misleading without full-rank $\beta
$'s or with sampling errors of comparable size as the loadings. Empirical
applications show that some factors, though potentially weak, may drive the
time variation of risk prices, and weak identification issues are more
prominent in multi-factor models.
\end{abstract}
\thispagestyle{empty}\bigskip
JEL Classification: C12, C58, G10, G12
Keywords: asset pricing; risk premia; robust inference; weak identification\bigskip
\section{Introduction}
A variety of Dynamic affine term structure models (DATSMs) have been
developed since the foundational work by Vasicek (1977)\nocite
{vasicek1977equilibrium} and Cox et al. (1985).\nocite{cox198fiftheory} They
help to understand the movements of bond yields and to analyze financial
markets. DATSMs are empirically appealing for their smooth tractability and
simple characterization of how risks get priced. There are many studies
employing this framework. To list a few: Cochrane and Piazzesi (2005)\nocite
{cochrane2005bond} apply affine term structure models to study time
variation in expected excess bond returns using a single powerful
explanatory factor; Wu and Xia (2016)\nocite{wu2016measuring} use affine
models to summarize the macroeconomic effects of unconventional monetary
policy; Ang and Piazzesi (2003)\nocite{ang2003no} investigate how macro
economic variables affect bond prices and the dynamics of the yield curve,
Buraschi and Jiltsov (2005)\nocite{buraschi2005inflation} study the
properties of the nominal and real risk premia of the term structure of
interest rates and \nocite{golinski2016long}Goli{\'{n}}ski and Zaffaroni
(2016) incorporate long memory state variables into the term structure
model. We adopt the DATSMs setup developed by \nocite{adrian2013pricing}
Adrian et al. (2013) which nests a general class of linear asset pricing
models and can be regarded as a linear asset pricing model with time-varying
risk premia and dynamic factors.
Many recent studies have developed approaches to estimate DATSMs. Most of
them involve a time-consuming numerical optimization procedure which results
from the high non-linearity. The inference concerning the coefficients
suffers similar challenges. Another undesirable feature, as pointed out by,
e.g., Hamilton and Wu (2012)\nocite{hamilton2012identification}, is that the
identification can be problematic. Lack of identification, e.g., due to
unspanned factors (\nocite{adrian2013pricing}Adrian et al. (2013)), and thus
a relatively flat surface of the likelihood also leads to unsatisfying
inference results (e.g., Kan and Zhang (1999),\nocite{kan1999two} \nocite
{gospodinov2017spurious}Gospodinov et al. (2017), \nocite
{kleibergen2009tests}Kleibergen (2009), Hamilton and Wu (2012)\nocite
{hamilton2012identification}, \nocite{dovonon2013testing}Dovonon and Renault
(2013), Beaulieu et al. (2016)\nocite{beaulieu2016less}, Khalaf and Schaller
(2016),\nocite{khalaf2016identification} \nocite{cattaneo2022beta}Cattaneo
et al. (2022)).
Unspanned factors refer to those that only affect the dynamics of bond
prices under the historical measure but not the risk neutral one (see, e.g.,
Joslin et al. (2014)\nocite{joslin2014risk}). The presence of such factors
has been documented in empirical studies (e.g., Ludvigson and Ng (2009)
\nocite{ludvigson2009macro}, \nocite{adrian2013pricing}Adrian et al.
(2013)). They can lead to identification challenges since varying the
parameters of the risk neutral pricing measure - risk premia parameters -
associated with these factors does not strongly influence bond prices.
\nocite{adrian2013pricing}Adrian et al. (2013) allow for the presence of
unspanned factors with the prior knowledge of knowing which factors are
unspanned. Their proposed estimation procedure differs for cases with and
without unspanned factors.
Because of the identification issues, traditional inference methods based on
t-tests and Wald statistics become unreliable for conducting inference on
the risk prices in DATSMs (e.g., \nocite{stock2000gmm}Stock and Wright
(2000) \nocite{kleibergen2005testing}Kleibergen (2005), \nocite
{antoine2020testing}Antoine and Renault (2020), \nocite
{andrews2012estimation}Andrews and Cheng (2012), \nocite
{antoine2022identification}Antoine and Lavergne (2022)). We therefore
propose easy-to-implement identification robust test procedures which are
valid even when the model is not identified due to unspanned factors. The
test procedures we provide can be used to study the time-varying risk-premia
for linear asset pricing models. Our proposed inference procedures use the
framework presented in \nocite{adrian2013pricing}Adrian et al. (2013), where
the risk of bond prices is modeled as a linear functional in observed
factors. The risk of bond prices can then be decomposed into two parts: a
time-constant and a time-varying part. We propose statistics for testing
hypothezes specified on all parameters of the time-varying components and on
just subsets of them.
The paper is organized as follows: Section 2 introduces the DATSM. Section 3
states the three step estimation procedure from Adrian et al. (2013)\nocite
{adrian2013pricing} and shows that it is sensitive to identification issues.
Section 3 also shows the empirical relevance of these identification issues
using the data from Adrian et al. (2013). Section 4 introduces the
identification robust tests for the time-varying risk premia. It conducts a
small simulation experiment and applies then for an one time-varying risk
factor setting. Section 6 introduces the identification robust tests for
testing hypothezes on subsets of the time-varying risk premia. It applies
them to a variety of multi-factor settings using the data from Adrian et al.
(2013). The final sixth section concludes.
We use the following notation throughout the paper: $\otimes \ $and vec$
(\cdot )$ represent respectively the Kronecker product and vectorization
operator; $\Sigma ^{\frac{1}{2}}$ is the lower triangular Cholesky
decomposition of the positive definite symmetric matrix $\Sigma $ such that $
\Sigma =\Sigma ^{\frac{1}{2}}\Sigma ^{\frac{1}{2}^{\prime }}$; for a $
N\times K$ dimensional full rank matrix $A:P_{A}=A(A^{\prime
}A)^{-1}A^{\prime }$ and $M_{A}=I_{N}-P_{A}\text{.}$
\section{Dynamic affine term structure models (DATSMs)}
We briefly discuss the popular class of DATSMs with observed factors.
Instead of working directly with the implied yields on an $n$-period bond as
usually done in the term structure literature, we make use of the excess
holding return of an $n$-period bond as in \nocite{adrian2013pricing}Adrian
et al. (2013).
We first illustrate the model set-up following \nocite{adrian2013pricing}
Adrian et al. (2013) and thereafter consider tests associated with the risk
premia. For $P_{t,n}$ the price at time $t$ of a zero-coupon bond maturing
at time $t+n\text{,}$ the pricing kernel, $M_{t+1},$ is such that
\begin{equation*}
P_{t,n}=E_{t}(M_{t+1}P_{t+1,n-1})\text{.}
\end{equation*}
For $r_{t}$ the one-period short rate and $\lambda _{t}$ the market price of
risk, the pricing kernel is assumed exponential affine in innovation factors
$v_{t}\sim _{i.i.d}N(0,\Sigma _{v}):$
\begin{equation*}
M_{t+1}=\exp \left( -r_{t}-\frac{1}{2}\lambda _{t}^{\prime }\lambda
_{t}-\lambda _{t}^{\prime }\Sigma _{v}^{-\frac{1}{2}}v_{t+1}\right) ,
\end{equation*}
where the market price of risk $\lambda _{t}$ is an affine function of the
observed factors $X_{t}:$
\begin{equation*}
\lambda _{t}=\Sigma _{v}^{-\frac{1}{2}}\left( \lambda _{0}+\Lambda
_{1}X_{t}\right) ,
\end{equation*}
with $\lambda _{0}$ and $\Lambda _{1}$ resp. a $K$-dimensional vector and a $
K\times K$ dimensional matrix, and the $K$-dimensional vector of state
variables $X_{t}$ results from a VAR(1):
\begin{equation*}
X_{t+1}=\mu +\Phi_1 X_{t}+v_{t+1}\text{.}
\end{equation*}
For the one-period (log) excess holding return of a $n$-period bond at $t+1:$
\begin{equation*}
r_{t+1,n}=\ln \left( P_{t+1,n}\right) -\ln \left( P_{t,n+1}\right) -r_{t}
\text{,}
\end{equation*}
with $r_{t}=\ln P_{t,1}$, the structure of the pricing kernel implies that:
\begin{equation*}
E_{t}\left[ \exp \left( -r_{t+1,n}-\frac{1}{2}\lambda _{t}^{\prime }\lambda
_{t}-\lambda _{t}^{\prime }\Sigma _{v}^{-\frac{1}{2}}v_{t+1}\right) \right]
=1.
\end{equation*}
Assuming that $(r_{t+1,n}\text{,}$ $v_{t+1})$ are jointly normal, Adrian et
al. (2013) show that:
\begin{equation*}
E_{t}\left( r_{t+1,n}\right) =\beta ^{(n)\prime }\left( \lambda _{0}+\Lambda
_{1}X_{t}\right) -\frac{1}{2}var(r_{t+1,n})\text{,}
\end{equation*}
with $\beta ^{(n)}=\Sigma _{v}^{-1}cov(v_{t+1},r_{t+1,n})\text{.}$ When
decomposing, $R_{t+1,n}$ into a component correlated with $v_{t+1}$ and an
uncorrelated component/prediction error $e_{t+1,n}:$
\begin{equation*}
\begin{array}{c}
r_{t+1,n}-E_{t}\left( r_{t+1,n}\right) =\beta ^{(n)\prime }v_{t+1}+e_{t+1,n};
\end{array}
\end{equation*}
so
\begin{equation*}
\begin{array}{c}
r_{t+1,n}=\beta ^{(n)\prime }\left( \lambda _{0}+\Lambda _{1}X_{t}\right)
+g^{(n)}(\beta ,\Sigma _{v},\Sigma _{e})+\beta ^{(n)\prime }v_{t+1}+e_{t+1,n}
\text{,}
\end{array}
\end{equation*}
where $g^{(n)}(\beta ,\Sigma _{v},\Sigma _{e})=-\frac{1}{2}var(r_{t+1,n})
\text{,}$ thus, for example, in case $\Sigma _{e}=$var($e_{t+1,n})=\sigma
_{e}^{2}\text{,}$ $g^{(n)}(\beta ,\Sigma _{v},\Sigma _{e})=-\frac{1}{2}
\left( \beta ^{(n)\prime }\Sigma _{v}\beta ^{(n)}+\sigma _{e}^{2}\right)
\text{. }$Additional restrictions are often imposed on the parameters
because of the cross-sectional term structure, but these restrictions are
dropped in \nocite{adrian2013pricing}Adrian et al. (2013)'s approach.
Assumption \ref{assum: model specification} summarizes the model setting.
\begin{assumption}
\label{assum: model specification}~\newline
(a) Consider a $K\times 1$ vector of factors $X_{t},$ $t=0,\cdots ,T$ that
results from a stationary vector autoregression of order 1:
\begin{equation*}
\begin{array}{c}
X_{t+1}=\mu +\Phi_1 X_{t}+v_{t+1},
\end{array}
\end{equation*}
where $v_{t}$ are the innovation shocks (or innovation factors). The log
excess holding return $r_{t+1,n-1}$ satisfies:
\begin{equation*}
{\small
\begin{array}{c}
r_{t+1,n-1}=\beta ^{(n-1)^{\prime }}\left( \lambda _{0}+\Lambda
_{1}X_{t}\right) +g^{(n-1)}\left( \beta ,\Sigma _{v},\Sigma _{e}\right)
+\beta ^{(n-1)^{\prime }}v_{t+1}+e_{t+1,n-1}
\end{array}
}
\end{equation*}
with $g^{(n-1)}(\cdot )$ a parametric function. Furthermore:
\begin{equation*}
(v_{t+1}^{\prime },e_{t+1}^{\prime })^{\prime }|\left\{ X_{s}\right\}
_{s=0}^{t}\sim i.i.d.N\left( 0,\text{diag}(\Sigma _{v},\Sigma _{e})\right) .
\end{equation*}
(b) For $\Sigma _{e}=\sigma _{e}^{2},$ $g^{(n-1)}\left( \beta ,\Sigma
_{v},\Sigma _{e}\right) =\frac{1}{2}\left( \beta ^{(n-1)^{\prime }}\Sigma
_{v}\beta ^{(n-1)}+\sigma _{e}^{2}\right) $.
\end{assumption}
\section{Regression estimator and Wald based inference}
To estimate the price of risk, Adrian et al. (2013) propose a three-step
procedure akin to the two pass procedure from \nocite{fama1973risk}Fama and
MacBeth (1973):
\begin{enumerate}
\item Estimate:
\begin{equation*}
X_{t+1}=\mu +\Phi_1 X_{t}+v_{t+1}\text{,}
\end{equation*}
by least squares to obtain $\hat{\mu}\text{,}$ $\hat{\Phi}\text{,}$ $\hat{v}
_{t}=X_{t}-\hat{\mu}-\hat{\Phi}_1X_{t-1}\text{,}$ $t=1,\ldots ,T$ and $\hat{
\Sigma}_{v}=\frac{1}{T}\sum_{t=1}^{T}\hat{v}_{t}\hat{v}_{t}^{\prime }\text{.}
$
\item Estimate:
\begin{equation*}
r_{t+1,n}=a^{(n)}+d^{(n)\prime }X_{t}+\beta ^{(n)\prime }\hat{v}
_{t+1}+e_{t+1,n}\text{,}
\end{equation*}
by least squares to obtain $\hat{a}^{(n)}\text{,}$ $\hat{d}^{(n)}$ and $\hat{
\beta}^{(n)}\text{,}$ $n=1,\ldots ,N\text{.}$
\item Construct $\hat{a}=(\hat{a}^{(1)}\ldots \hat{a}^{(N)})^{\prime },\text{
}$ $\hat{\beta}=(\hat{\beta}^{(1)}\ldots $ $\hat{\beta}^{(N)})^{\prime },$ $
\hat{d}=(\hat{d}^{(1)}\ldots d^{(n)})^{\prime },$ $\hat{g}=(\hat{g}
^{(1)}\ldots \hat{g}^{(n)})$ for $\hat{g}^{(n)}=g^{(n)}(\hat{\beta},\hat{
\Sigma}_{v},\hat{\Sigma}_{e}),$ $n=1,\ldots ,N\text{,}$ and estimate $
\lambda _{0}$ and $\Lambda _{1}$ using:
\begin{equation*}
\begin{array}{cl}
\hat{\lambda}_{0}= & \left( \hat{\beta}^{\prime }\hat{\beta}\right)
^{^{\prime }-1}\hat{\beta}^{\prime }(\hat{a}+\hat{g}+\hat{\beta}\hat{\mu})
\\
\hat{\Lambda}_{1}= & \left( \hat{\beta}^{\prime }\hat{\beta}\right)
^{^{\prime }-1}\hat{\beta}^{\prime }(\hat{d}+\hat{\beta}\hat{\Phi}).
\end{array}
\end{equation*}
\end{enumerate}
The three-step procedure essentially regresses transformed returns on
estimated $\beta $'s. Two issues can hamper the reliability of the final
stage: the sampling error of estimates resulting from previous stages and
the quality of the $\beta $'s. The first issue is negligible when the
information dominates the asymptotically vanishing sampling errors, and only
slight modifications are necessary for the asymptotic variance estimator
used in test statistics to accommodate for it. However, when modeling the
unspanned factors, \nocite{adrian2013pricing}Adrian et al. (2013) show that
the entries in the $\beta $'s corresponding to unspanned factors are zero
(Assumption \ref{assum:unspanned factors}), which leads to identification
problems (e.g., \nocite{kleibergen2009tests}Kleibergen (2009), Beaulieu et
al. (2013)\nocite{beaulieu2013identification},\nocite
{kleibergen2015unexplained}Kleibergen and Zhan (2015)).
\begin{assumption}
\label{assum:unspanned factors} Potential existence of (nearly) unspanned
factors: $\beta ^{\prime }=(B,C)$, $B$ represents the spanned factors and is
of full rank $K_{B}\leq K$ while $C=O(1/\sqrt{T})$ reflects the (nearly)
unspanned factors. If $K_{B}=K,$ there are no (nearly) unspanned factors.
\end{assumption}
\nocite{adrian2013pricing}Adrian et al. (2013) show that Assumption \ref
{assum:unspanned factors} embeds the unspanned factors that do not affect
the dynamics of bonds under the historical pricing measure. The rows of $
\beta $ comprised of close to zero values correspond to unspanned factors.
\nocite{adrian2013pricing} Adrian et al. (2013) assume that the location of
these rows is known, and their three-step estimation procedure excludes the
unspanned factors in the second step by only including spanned factors. The
third step would otherwise encounter identification issues resulting from a
classical multicollinearity problem. It is straightforward to show that zero
$\beta $'s lead to an identification problem because any value of the $
\lambda $'s associated with the zero elements in $\beta $ goes for the
values of excess returns. When we instead just have small $\beta $'s, which
are comparable in magnitude to the estimation error, we similarly encounter
such an identification problem (e.g., \nocite{kleibergen2009tests}Kleibergen
(2009), \nocite{antoine2009efficient}Antoine and Renault (2009, 2012),
\nocite{antoine2012efficient} \nocite{kleibergen2020robust}Kleibergen and
Zhan (2020), \nocite{kleibergen2019identification}Kleibergen et al. (2022)).
Proposition \ref{prop2} states some well-known results from the weak
identification literature. It shows that the risk premia estimator $\hat{
\Lambda}$ becomes inconsistent in the presence of weak factors because it
converges to a non-normal distribution. This results since varying the value
of the risk premia associated with the weak factors does not change the
excess returns. The asymptotic distribution of the conventional Wald
statistic for testing the null hypothesis $H_{0}:\Lambda _{1}=\Lambda
_{1}^{0}$ then no longer converges to a $\chi ^{2}$ -distribution, and the
same holds for subset testing based on this estimator. Therefore, the
conventional test statistics can be misleading in the presence of unspanned
factors.
\begin{proposition}
\label{prop2}~\newline
Under Assumptions \ref{assum: model specification}.(a), denote $\Lambda
=[\lambda _{0},\Lambda _{1}]$:\newline
(a) If $\beta $ is of full rank:
\begin{equation*}
\begin{array}{c}
\sqrt{T}\left[
\begin{matrix}
\text{vec}\left( \hat{\beta}^{\prime }-\beta ^{\prime }\right) \\
\text{vec}\left( \hat{\Lambda}-\Lambda \right)
\end{matrix}
\right] \rightarrow _{d}N\left( \left[
\begin{matrix}
0 \\
0
\end{matrix}
\right] ,\left[
\begin{matrix}
\mathcal{V}_{\beta } & \mathcal{C}_{\Lambda ,\beta }^{\prime } \\
\mathcal{C}_{\Lambda ,\beta } & \mathcal{V}_{\Lambda }
\end{matrix}
\right] \right) ,
\end{array}
\label{eq:propo1}
\end{equation*}
where $\mathcal{V}_{\beta },\mathcal{C}_{\Lambda ,\beta },\mathcal{V}
_{\Lambda }$ are specified in the proof.\newline
(b) If Assumption \ref{assum:unspanned factors} holds with $K_{B}<K$, then $
\hat{\Lambda}\rightarrow _{d}\Lambda +\epsilon,$ where $\epsilon $ follows a
non-standard distribution, so $\hat{\Lambda}$ no longer converges to the
true value $\Lambda $ at rate $\sqrt{T}$.
\end{proposition}
\begin{proof}
See the Online Supplementary Appendix.
\end{proof}
We focus on inference concerning the time-varying component of risk prices $
\Lambda _{1}$. We therefore demean the one-period (log) excess holding
returns by subtracting its time-series average:
\begin{equation*}
\bar{r}_{t+1,n}=\beta ^{(n)\prime }\left( \Lambda _{1}\bar{X}_{t}\right)
+\beta ^{(n)\prime }\bar{v}_{t+1}+\bar{e}_{t+1,n}\text{,}
\end{equation*}
with $\bar{z}_{t+1,n}=z_{t+1,n}-\bar{z},$ with $\bar{z}=\frac{1}{T}
\sum_{t=1}^{T}z_{t},$ for $z=r,$ $X,$ $v$ and $e$ resp. When stacking the
equations for $N$ different maturities:
\begin{equation*}
R_{t+1}=\beta \left( \Lambda _{1}\bar{X}_{t}\right) +\beta \bar{v}
_{t+1}+e_{t+1}\text{,}
\end{equation*}
where $R_{t+1}=(\bar{r}_{t+1,1}\ldots \bar{r}_{t+1,N})^{\prime }\text{,}$ $
\beta =(\beta ^{(1)}\ldots \beta ^{(n)})^{\prime }\text{,}$ $e_{t}=(\bar{e}
_{t,1}\ldots \bar{e}_{t,N})^{\prime }\text{,}$ the pricing equation closely
resembles the beta-pricing model for the return on (portfolios of) assets:
\begin{equation*}
\text{r}_{t+1}=\beta \lambda +\beta \text{F}_{t+1}+u_{t+1}\text{,}
\end{equation*}
with r$_{t}$ an $N$-dimensional vector with the returns on $N$ assets, $
\beta $ the $N\times K$ dimensional beta matrix and $F_{t}$ a $K$
-dimensional vector of risk factors. A further important similarity that
both models imply is the reduced rank structure that becomes obvious using a
slight re-specification:
\begin{equation*}
\begin{array}{ccc}
R_{t+1}=\beta \left( \Lambda _{1}\ \text{{}}\vdots \ \text{{}}I_{K}\right)
\left(
\begin{array}{c}
\bar{X}_{t} \\
\bar{v}_{t+1}
\end{array}
\right) +e_{t} & &
\end{array}
\end{equation*}
and
\begin{equation*}
\begin{array}{ccc}
\text{r}_{t+1}=\beta \left( \lambda \ \text{{}}\vdots \ \text{{}}
I_{K}\right) \left(
\begin{array}{c}
1 \\
F_{t+1}
\end{array}
\right) +u_{t}, & &
\end{array}
\end{equation*}
where the $N\times 2K$ and $N\times (K+1)$ dimensional matrices $\beta
(\Lambda _{1}\ \text{{}}\vdots \ \text{{}}I_{K})$ and $\beta (\lambda \
\text{{}}\vdots \ \text{{}}I_{K})$ are each at most of rank $K\text{,}$ so
except for the largest $K$ singular values, all, $K$ and 1 resp., singular
values of these matrices are zero. Further for $\Lambda _{1}$ and $\lambda $
to be well defined, $\beta $ should be of full rank. When $\beta $ is near a
reduced rank value, or in other words, if some factors are weak/unspanned,
we encounter an identification issue which is also reflected by more than
just $K$ or 1 resp. singular values of the above matrices being equal or
close to zero.\bigskip
\noindent {{\small
\begin{tabular}{c|c|c|c|c|cc}
\hline\hline
& $\hat{\beta}_{1}$ & $\hat{\beta}_{2}$ & $\hat{\beta}_{3}$ & $\hat{\beta}
_{4}$ & $\hat{\beta}_{5}$ & \\ \hline
{\ (1)} & {\ -0.0094} & {\ 0.0031 } & {\ -0.0008 } & {\ 0.0002 } & {\ 0.0000}
& \\
& (-3.6027) & (2.5058) & (-1.4886) & (0.4344) & (0.0347) & \\
{\ (2)} & {\ -0.0213 } & {\ 0.0057 } & {\ -0.0007 } & {\ -0.0003} & {\
0.0002 } & \\
& (-8.2107) & (4.5416) & (-1.2228) & (-0.7341) & (0.5816) & \\
{\ (3)} & {\ -0.0446} & {\ 0.0070} & {\ 0.0010 } & {\ -0.0005} & {\ -0.0001}
& \\
& (-17.1745) & (5.6040) & (1.8889) & (-1.3535) & (-0.4551) & \\
{\ (4)} & {\ -0.0656} & {\ 0.0048 } & {\ 0.0024 } & {\ 0.0000} & {\ -0.0003}
& \\
& (-25.2648) & (3.8500) & (4.5279) & (0.0386) & (-0.9643) & \\
{\ (5)} & {\ -0.0843} & {\ 0.0003} & {\ 0.0028 } & {\ 0.0007 } & {\ -0.0001}
& \\
& (-32.4670) & (0.2023) & (5.2857) & (1.7192) & (-0.1732) & \\
{\ (6)} & {\ -0.1011 } & {\ -0.0059} & {\ 0.0022} & {\ 0.0010 } & {\ 0.0003}
& \\
& (-38.9474) & (-4.6910) & (4.1584) & (2.6824) & (1.0713) & \\
{\ (7)} & {\ -0.1164 } & {\ -0.0130 } & {\ 0.0008} & {\ 0.0010} & {\ 0.0006}
& \\
& (-44.8511) & (-10.3464) & (1.5602) & (2.5469) & (1.9102) & \\
{\ (8)} & {\ -0.1305 } & {\ -0.0206} & {\ -0.0011} & {\ 0.0005 } & {\ 0.0005}
& \\
& (-50.2742) & (-16.4072) & (-2.0105) & (1.2971) & (1.8297) & \\
{\ (9)} & {\ -0.1435 } & {\ -0.0284} & {\ -0.0033} & {\ -0.0003} & {\ 0.0002}
& \\
& (-55.2826) & (-22.6118) & (-6.1006) & (-0.8978) & (0.6354) & \\
{\ (10)} & {\ -0.1556 } & {\ -0.0361} & {\ -0.0056} & {\ -0.0015} & {\
-0.0005} & \\
& (-59.9307) & (-28.7667) & (-10.3483) & (-3.7935) & (-1.6457) & \\
{\ (11)} & {\ -0.1669} & {\ -0.0436} & {\ -0.0078} & {\ -0.0028} & {\
-0.0014 } & \\
& (-64.2706) & (-34.7284) & (-14.4913) & (-7.1319) & (-4.8558) & \\ \hline
& [0.0000] & [0.0000] & [0.0000] & [0.0000] & [0.0001] & \\ \hline\hline
\end{tabular}
}} ~\bigskip
\noindent {\small \textbf{Table 1}: The least squares estimates of the $
\beta $'s associated with the excess returns for bonds with 11 different
maturities of $6,12,18\cdots ,60$ and 84, 120 months over the sample period
1987:01-2011:12, and the factors are the first five principal components
generated using the cross-section of bond yields for maturities $3,\cdots
,120$ months (same data taken from Adrain et al. (2013)). Based on
Proposition \ref{prop2}.(a), we report LS estimates of $\beta $'s with
t-statistics in round brackets; and$p$-values of $F$-tests in square
brackets, testing the null hypothesis $H_{0}:\beta _{j}=0$ that each column
is jointly zero.. The Kleibergen-Paap rank statistic (see Kleibergen and
Paap (2006)) testing $H_{0}:$ rank($\beta $)=4: 1.6561 [0.9764]. }
\paragraph{Illustrating unspanned and weak factors in an empirical study}
In reality, there is no direct prior knowledge of the number of unspanned or
weak factors. To show their presence and empirical relevance, we use the
data from \nocite{adrian2013pricing}Adrian et al. (2013), i.e., the zero
coupon yield data constructed by G\"{u}rkaynak et al. (2007).\nocite
{gurkaynak2007us} Table 1 shows the factor loading estimates and relevant
tests corresponding to the five principal component (PCA) factors employed
in Adrian et al. (2013). Table 1 shows that many elements of $\beta $ are
small and not statistically significant from zero at the 5\% significance
level. Though the $F$-tests on the columns of $\beta $ are significant, the
rank test of the $\beta $-matrix indicates potential identification
problems, because for these factor loadings we can not reject the reduced
rank null hypothesis at the 5\% significance level.
\begin{figure}[htbp!]
\includegraphics[width=\columnwidth]{fig12-eigenvalue.png}
\caption*{{\small \textbf{Figure 1}: (log) singular values (blue) and the
percentage of variance explained (red) of $\displaystyle\frac{1}{T}
\sum_{t=1}^{T}R_{t}\left(
\begin{array}{c}
\bar{X}_{t} \\
\hat{v}_{t+1}
\end{array}
\right) ^{\prime }\left[ \frac{1}{T}\sum_{t=1}^{T}\left(
\begin{array}{c}
\bar{X}_{t} \\
\hat{v}_{t+1}
\end{array}
\right) \left(
\begin{array}{c}
\bar{X}_{t} \\
\hat{v}_{t+1}
\end{array}
\right) ^{\prime }\right] .$ $R_{t}$ uses the demeaned excess returns of the
same bonds as in Table 1, $\bar{X}_{t}$ involves the five PCA factors in (a)
and uses the five PCA factors with one additional macro factor in (b). }}
\end{figure}
\bigskip
Figure 1.(a) uses the same data set as Table 1 and shows a scree plot of the
(log) singular values of
\begin{equation*}
\frac{1}{T}\sum_{t=1}^{T}R_{t}\left(
\begin{array}{c}
\bar{X}_{t} \\
\hat{v}_{t+1}
\end{array}
\right) ^{\prime }\left[ \frac{1}{T}\sum_{t=1}^{T}\left(
\begin{array}{c}
\bar{X}_{t} \\
\hat{v}_{t+1}
\end{array}
\right) \left(
\begin{array}{c}
\bar{X}_{t} \\
\hat{v}_{t+1}
\end{array}
\right) ^{\prime }\right] ,
\end{equation*}
which estimates $\beta (\Lambda _{1}\ \text{{}}\vdots \ \text{{}}I_{K})$. In
case of strong spanned factors, the smallest $K,$ is five, singular values
should be close to zero while Figure 1.(a) shows that the smallest six
singular values are close to zero which indicates a weak/unspanned factor
problem which is further reflected by the rank test not rejecting rank($
\beta $)=4 in Table 1 at the 5\% significance level.{~}
~\bigskip
\noindent {{\small \
\begin{tabular}{c|c|c|c|c|c|c}
\hline\hline
& $\hat{\beta}_{1}$ & $\hat{\beta}_{2}$ & $\hat{\beta}_{3}$ & $\hat{\beta}
_{4}$ & $\hat{\beta}_{5}$ & $\hat{\beta}_{\mathit{macro}}$ \\ \hline
{\ (1)} & -0.0094 & 0.0032 & -0.0008 & 0.0002 & 0.0000 & 0.0001 \\
& (-3.4930) & (2.4109) & (-1.4749) & (0.4253) & (0.0309) & (0.0516) \\
{\ (2)} & -0.0213 & 0.0057 & -0.0007 & -0.0003 & 0.0002 & 0.0000 \\
& (-7.9353) & (4.3501) & (-1.2115) & (-0.7259) & (0.5717) & (0.0254) \\
{\ (3)} & -0.0446 & 0.0070 & 0.0010 & -0.0005 & -0.0001 & -0.0000 \\
& (-16.5873) & (5.3631) & (1.8715) & (-1.3372) & (-0.4497) & (-0.0037) \\
{\ (4)} & -0.0656 & 0.0049 & 0.0024 & 0.0000 & -0.0003 & 0.0001 \\
& (-24.4149) & (3.6980) & (4.4871) & (0.0346) & (-0.9529) & (0.0655) \\
{\ (5)} & -0.0843 & 0.0003 & 0.0028 & 0.0007 & -0.0001 & 0.0001 \\
& (-31.3724) & (0.2100) & (5.2383) & (1.6937) & (-0.1732) & (0.0801) \\
{\ (6)} & -0.1011 & -0.0059 & 0.0022 & 0.0010 & 0.0003 & 0.0000 \\
& (-37.6206) & (-4.4799) & (4.1208) & (2.6462) & (1.0539) & (0.0333) \\
{\ (7)} & -0.1164 & -0.0130 & 0.0008 & 0.0010 & 0.0006 & -0.0000 \\
& (-43.3100) & (-9.8999) & (1.5456) & (2.5142) & (1.8813) & (-0.0252) \\
{\ (8)} & -0.1305 & -0.0206 & -0.0011 & 0.0005 & 0.0005 & -0.0001 \\
& (-48.5424) & (-15.7013) & (-1.9930) & (1.2811) & (1.8021) & (-0.0512) \\
{\ (9)} & -0.1435 & -0.0284 & -0.0033 & -0.0003 & 0.0002 & -0.0000 \\
& (-53.3870) & (-21.6288) & (-6.0457) & (-0.8872) & (0.6243) & (-0.0204) \\
{\ (10)} & -0.1557 & -0.0361 & -0.0056 & -0.0015 & -0.0005 & 0.0001 \\
& (-57.8990) & (-27.4937) & (-10.2541) & (-3.7506) & (-1.6271) & (0.0745) \\
{\ (11)} & -0.1670 & -0.0435 & -0.0078 & -0.0028 & -0.0014 & 0.0002 \\
& (-62.1305) & (-33.1560) & (-14.3589) & (-7.0557) & (-4.7981) & (0.2297) \\
\hline
& [0.0000] & [0.0000] & [0.0000] & [0.0000] & [0.0002] & [1.0000] \\
\hline\hline
\end{tabular}
}} ~\bigskip
\noindent {\small \textbf{Table 2}: Using the same data as in Table 1 with
one additional macro factor (real activity) that constructed following Ang
and Piazzesi (2003), we report estimates of the $\beta $'s with t-statistics
in round brackets; and $p$-values of $F$-tests in square brackets, testing
the null hypothesis $H_{0}:\beta _{j}=0$ that each column is jointly zero.
Kleibergen-Paap rank statistic testing $H_{0}:$ rank($\beta $ )=5: 0.0025
[1.0000]. \bigskip }
As previously noted in the literature (e.g., \nocite{kleibergen2020robust}
Kleibergen and Zhan (2020), Kleibergen et al. (2022)), weak identification
issues are often present when macro factors are used. Following \nocite
{ang2003no}Ang and Piazzesi (2003), we construct one macro factor, the real
activity measure, which is the first principal component resulting from four
variables that capture real US macro activity: the "Help Wanted Advertising
in Newspapers (HELP)"\footnote{
We use the HELP-Wanted index from \nocite{barnichon2010building}Barnichon
(2010) to match the time periods of the excess returns.} index, unemployment
(UE), the growth rate of employment (EMPLOY), and the growth rate of
industrial production (IP). As shown in Table 2, the macro factor is much
less correlated with returns and thus is more likely to result in an
identification issue. This is further reflected in Figure 1(b) containing
the scree plot which shows that while there are six factors, the smallest
seven singular values are close to zero. Table 2 also shows a tiny value for
the rank test which provides another indication of a weak/unspanned factor
problem. \bigskip
The identification problems revealed in Figures 1-2 and Tables 1-2 not only
affect the validity of the estimators but also the reliability of
traditional inference procedures. \nocite{adrian2013pricing} Adrian et al.
(2013) assume that the positions of the zero rows in $\beta $ are known when
dealing with unspanned factors. We remain agnostic about this and provide
testing procedures concerning $\Lambda _{1}$ which are identification robust
without the need of prior knowledge of the unspanned factors.
\section{Identification robust tests of time-varying risk premia}
The identification robust tests are based on the sample moment vector. The
sample moment vector results from the observed factors having no predictive
power for the prediction error, $\bar{e}_{t+1,n},$ so our sample moment
vector for $\Lambda _{1}$ is:
\begin{equation*}
\begin{array}{c}
f_{T}(\Lambda _{1},X)=\frac{1}{T}\sum_{t=1}^{T}(\bar{X}_{t-1}\otimes (\bar{R}
_{t}-\hat{\beta}\hat{V}_{t}))-\left( \hat{Q}_{XX}\otimes \hat{\beta}\right)
\text{vec(}\Lambda _{1}),
\end{array}
\end{equation*}
with $\hat{Q}_{XX}=\frac{1}{T}\sum_{t=1}^{T}\bar{X}_{t-1}\bar{X}
_{t-1}^{\prime },$ and its derivative with respect to vec($\Lambda _{1})$ is
\begin{equation*}
\begin{array}{c}
q_{T}(X)=-\left( \hat{Q}_{XX}\otimes \hat{\beta}\right) \text{.}
\end{array}
\end{equation*}
We next make an assumption regarding the large sample behavior of the sample
moment vector and its derivative.
\begin{assumption}
\label{assum3}~\newline
Under $H_{0}:\Lambda _{1}=\Lambda _{1}^{0}$,
\begin{equation*}
\sqrt{T}\left(
\begin{array}{c}
f_{T}(\Lambda _{1}^{0},X) \\
\text{vec}(q_{T}(X)-J)
\end{array}
\right) \underset{d}{\rightarrow }\left(
\begin{array}{c}
\psi _{f} \\
\psi _{q}
\end{array}
\right) \text{,}
\end{equation*}
where the Jacobian $J=-(Q_{XX}\otimes \beta ),$ $\hat{Q}_{XX}\underset{p}{
\rightarrow }Q_{XX}\text{,}$ and $\psi _{f}$ and $\psi _{q}$ are $N\times K$
and $NK\times K^{2}$ dimensional random vectors:
\begin{equation*}
\left(
\begin{array}{c}
\psi _{f} \\
\psi _{q}
\end{array}
\right) \sim N\left( 0,V(\Lambda _{1}^{0})\right) \text{,}
\end{equation*}
with
\begin{equation*}
V(\Lambda _{1}^{0})=\left(
\begin{array}{cc}
V_{ff}(\Lambda _{1}^{0}) & V_{qf}(\Lambda _{1}^{0})^{\prime } \\
V_{qf}(\Lambda _{1}^{0}) & V_{qq}(\Lambda _{1}^{0})
\end{array}
\right) ,
\end{equation*}
where $V_{ff}(\Lambda _{1}^{0}),$ $V_{qf}(\Lambda _{1}^{0})$ and $
V_{qq}(\Lambda _{1}^{0})$ are $NK\times NK,$ $NK^{3}\times NK$ and $
NK^{3}\times NK^{3}$ dimensional matrices.
\end{assumption}
Assumption \ref{assum3} is a high-level assumption which resembles
Assumption 1 in Kleibergen (2005) and holds under mild conditions. \nocite
{kleibergen2005testing} Assumption \ref{assum3} holds true irrespective of
Assumption \ref{assum:unspanned factors}. Assumption \ref{assum: model
specification} is sufficient for Assumption \ref{assum3}, but our proposed
test statistics can be applied to more general cases than the model implied
in Assumption \ref{assum: model specification}. For our setting:
\begin{equation*}
\psi _{f}=\psi _{\bar{f}}+\Psi _{q}\text{vec(}\Lambda _{1}^{0})\text{,}
\end{equation*}
with $\Psi _{q}=$vecinv($\psi _{q})$ and
\begin{equation*}
\begin{array}{c}
\sqrt{T}\left( \frac{1}{T}\sum_{t=1}^{T}(\bar{X}_{t-1}\otimes (\bar{R}_{t}-
\hat{\beta}\hat{V}_{t}))-\left( Q_{XX}\otimes \beta \right) \text{vec}
(\Lambda _{1}^{0})\right) \underset{d}{\rightarrow }\psi _{\bar{f}}\text{.}
\end{array}
\end{equation*}
We also have
\begin{equation*}
\psi _{q}=\text{vec}((Q_{XX}\otimes \Psi _{\beta })+(\Psi _{XX}\otimes \beta
))\text{,}
\end{equation*}
where
\begin{equation*}
\begin{array}{cl}
\sqrt{T}\text{vech}(\frac{1}{T}\sum_{t=1}^{T}\bar{X}_{t}\bar{X}_{t}^{\prime
}-Q_{XX})\underset{d}{\rightarrow }\psi _{XX}\text{,} & \Psi _{XX}=\text{
vechinv}(\psi _{XX})\text{,} \\
\sqrt{T}\text{vec}(\hat{\beta}-\beta )\underset{d}{\rightarrow }\psi _{\beta
}\text{ ,} & \Psi _{\beta }=\text{vecinv}(\psi _{\beta })\text{,}
\end{array}
\end{equation*}
with vech$(A)$ containing the unique elements of a symmetric matrix $A$.
Since $\psi _{q}$ has $NK^{3}$ elements while the number of unique elements
in $\Psi _{\beta }$ and $\Psi _{XX}$ equals $NK+\frac{1}{2}K(K+1)\text{,}$
the joint normal distribution of $(\psi _{f}\text{,}$ $\psi _{q})$ is
further allowed to be degenerate. When $v_{t}$ is normal (as assumed in
Assumption \ref{assum: model specification}) or its third moment equals
zero, $\psi _{\beta }$ and $\psi _{XX}$ are also independently distributed.
The identification robust statistics use an estimator of the Jacobian whose
limit behavior under H$_{0}:\Lambda _{1}=\Lambda _{1}^{0}$ is independent of
the limit behavior of the sample moment, see Kleibergen (2005):
\begin{equation*}
\begin{array}{rl}
\hat{D}_{T}(\Lambda _{1},X)= & (\hat{D}_{1,T}(\Lambda _{1},X)\ldots \hat{D}
_{1,T}(\Lambda _{k},X)) \\
\text{vec(}\hat{D}_{T}(\Lambda _{1},X))= & \text{vec}(q_{T}(X))-\hat{V}
_{qf}(\Lambda _{1})\hat{V}_{ff}(\Lambda _{1})^{-1}f_{T}(\Lambda _{1},X) \\
\sqrt{T}\text{vec(}\hat{D}_{T}(\Lambda _{1}^{0},X)-J)\underset{d}{
\rightarrow } & \psi _{q.f}\sim N(0,V_{qq.f}(\Lambda _{1}^{0}))
\end{array}
\end{equation*}
with $V_{qq.f}(\Lambda _{1})=V_{qq}-V_{qf}(\Lambda _{1})V_{ff}(\Lambda
_{1})^{-1}V_{qf}(\Lambda _{1})^{\prime }\text{,}$ $\hat{V}_{qf}(\Lambda
_{1}) $ and $\hat{V}_{ff}(\Lambda _{1})$ consistent estimators of $
V_{qf}(\Lambda _{1}^{0})$ and $V_{ff}(\Lambda _{1}^{0}),$ and $\psi _{q.f}$
independent of $\psi _{f}.$
We can next define the identification robust Factor Anderson-Rubin (FAR),
(Kleibergen) Lagrange multiplier (KLM) and JKLM\ statistics for testing H$
_{0}:\Lambda _{1}=\Lambda _{1}^{0}:$
\begin{equation*}
\begin{array}{cl}
\text{FAR(}\Lambda _{1}^{0})= & T\times f_{T}(\Lambda _{1}^{0},X)^{\prime }
\hat{V}_{ff}(\Lambda _{1}^{0})^{-1}f_{T}(\Lambda _{1}^{0},X)\underset{d}{
\rightarrow }\chi ^{2}(KN) \\
\text{KLM(}\Lambda _{1}^{0})= & T\times f_{T}(\Lambda _{1}^{0},X)^{\prime }
\hat{V}_{ff}(\Lambda _{1}^{0})^{-\frac{1}{2}}P_{\hat{V}_{ff}(\Lambda
_{1}^{0})^{-\frac{1}{2}}\hat{D}_{T}(\Lambda _{1}^{0},X)} \\
\text{\thinspace \thinspace } & \text{\quad \quad \quad \quad \quad \quad }
\hat{V}_{ff}(\Lambda _{1}^{0})^{-\frac{1}{2}}f_{T}(\Lambda _{1}^{0},X)
\underset{d}{\rightarrow }\chi ^{2}(K^{2}) \\
\text{JKLM(}\Lambda _{1}^{0})= & T\times f_{T}(\Lambda _{1}^{0},X)^{\prime }
\hat{V}_{ff}(\Lambda _{1}^{0})^{-\frac{1}{2}}M_{\hat{V}_{ff}(\Lambda
_{1}^{0})^{-\frac{1}{2}}\hat{D}_{T}(\Lambda _{1}^{0},X)} \\
\text{\thinspace \thinspace } & \text{\quad \quad \quad \quad \quad \quad }
\hat{V}_{ff}(\Lambda _{1}^{0})^{-\frac{1}{2}}f_{T}(\Lambda _{1}^{0},X)
\underset{d}{\rightarrow }\chi ^{2}(K(N-K)).
\end{array}
\end{equation*}
The limiting distributions are a direct result of Assumption \ref{assum3}
and do not depend on the rank of the Jacobian or $\beta $, so the limiting
distributions hold regardless of Assumption \ref{assum:unspanned factors}.
\subsection{Illustrative simulation and empirical results}
We conduct a single-factor model simulation study to illustrate the
performance of the proposed robust joint tests. For the data generating
process (DGP), we consider
\begin{equation*}
R_{t}=c+\beta \left( \Lambda _{1}X_{t-1}+v_{t}\right) +e_{t}\text{,}
\end{equation*}
where the parameters are calibrated to data from Adrian et al. (2013). In
particular, we fix the sample size to be $T=300$ and use the eleven excess
returns as in Table 1. We calibrate with the first PCA factor from Adrian et
al. (2013) to mimic the strong identified case and the third one for the
weak identification setting. Figure 2 shows power curves of the conventional
t-statistic and the robust test statistics in both strong and weakly
identified cases. For both settings, FAR, KLM and JKLM tests are all size
correct, while the Wald test is size distorted under weak identification and
size correct but biased for stong identification. For weak identification,
the KLM test has some power loss away from the hypothesized value because of
which it is preferred to combine it in a conditional or unconditional manner
with the J-test to improve power, see Moreira (2003) and Kleibergen (2005).
We use the robust tests to analyze the time-varying component of the risk
premium. A detailed description of the involved excess returns and risk
factors has been discussed previously for Tables 1-2. Figure 3 shows the $p$
-values for testing the risk premium associated with all six factors in a
single factor model with the identification robust tests. A $p$-value larger
than, say, 5\%, implies that we could not reject the null at the 5\%
significance level.\bigskip
\begin{figure}[htbp!]
\includegraphics[width=\columnwidth]{fig3_joint_test_power_curves.jpg}
\caption*{{\small \textbf{Figure 2}: the left panel (a) is the strong
identified case, while the right panel (b) is the weak identified case.
Power curves (rejection frequencies) of Wald (dash-dotted, red), FAR
(solid), KLM (dashed), and JKLM (dotted) testing the $\Lambda _{1}$ (scalar)
with values on the horizontal axis; horizontal dashed line at 5\% and
vertical one at the calibrated value of $\Lambda _{1}$.}}
\end{figure}
\bigskip
Figure 3 shows that for most cases the JKLM test leads to unbounded 95\%
confidence sets even for strong factors, e.g., the first PCA factor, since
the $p$-value curves are above the 5\% line over the whole interval of
analyzed values of the risk premium which is consistent with the smaller
power observed in our simulation exercises and results since the JKLM test
primarily tests misspecification. For all factors, the FAR and KLM tests
provide bounded 95\% confidence sets since only for bounded regions the $p$
-values are above the 5\% line, even for those potentially weak factors such
as the fifth PCA factor and the macro factor. The latter implies that in a
single factor setting, all these risk premia are identified. For the
high-order PCA factors, the robust tests, however, result in 95\% confidence
sets that differ from those resulting from the Wald test. Most striking is
that a zero value for the risk premium is not rejected for strong factors
such as the first and second PCA factors but rejected for potentially weak
factors. For example, the null hypothesis that $\Lambda _{1}=0$ is rejected
by the FAR and KLM test for both the fifth factor and the macro factor. This
is partly in line with Adrian et al. (2013), which highlight the role of the
higher-order principal components as the time variation may be largely
driven by, e.g., the fifth principal component. Therefore, some factors may
be weak but have some importance for interpreting the (time-varying)
expected returns.
\begin{figure}[htbp!]
\includegraphics[height=0.8\textheight,width=\columnwidth]{fig4-Pvalue_curves_joint.png}
\caption*{{\small \textbf{Figure 3}: one-minus-$p$-value curves of Wald (dash-dotted,
red); FAR (solid); KLM (dashed); JKLM (dotted) for testing the $\Lambda _{1}$
(scalar) in a single factor model with values on the horizontal axis and
dotted line at 95\%. This figure uses the same data as in Table 2.}}
\end{figure}
\newpage
\section{Identification robust sub-vector testing}
The identification robust tests introduced in the previous section are for
testing hypotheses specified on all elements of $\Lambda _{1}.$ We are often
interested in testing hypotheses specified on just subsets of the
parameters. When we analyze multi-factor models, testing whether or not a
certain factor risk premium exhibits time variation would require testing a
specific row of $\Lambda _{1}$, while testing whether a factor drives the
time variation would require to test the corresponding column of $\Lambda
_{1}$. Under our current settings, projection-based versions of the
identification robust tests would allow us to test such hypotheses whilst
preserving the size of the test, see \nocite{dufour2005projection}Dufour and
Taamouti (2005). These tests, however, lead to reduced power so we extend
the robust subset FAR test (sFAR) of \nocite{guggenberger2012asymptotic}
Guggenberger et al. (2012) for testing hypotheses on all elements in a row
of $\Lambda _{1}$ which can be similarly extended to test for all elements
in a column of $\Lambda _{1}$.
Without loss of generality, we consider testing the hypothesis that the risk
premia associated with one specific factor, say the first, are all equal to $
\lambda _{1}^{0}$:
\begin{equation*}
\text{H}_{0}:\lambda _{1}=\lambda _{1}^{0},
\end{equation*}
for $\Lambda _{1}=\binom{\lambda _{1}^{0\prime }}{\Lambda _{2}},$ $\lambda
_{1}:K\times 1,$ $\Lambda _{2}:(K-1)\times K.$ Under H$_{0}:\lambda
_{1}=\lambda _{1}^{0},$ the $N\times 2K$ dimensional reduced rank parameter
matrix in the equation for the stacked returns becomes:
\begin{equation*}
\Phi =\left( \beta _{1}\ \text{{}}\vdots \ \text{{}}\beta _{2}\right) \left(
\begin{array}{ccc}
\lambda _{1}^{0\prime } & 1 & 0 \\
\Lambda _{2} & 0 & I_{K-1}
\end{array}
\right) ,
\end{equation*}
so post-multiplying by $\left(\begin{matrix}
{I_{K}} &{-\lambda _{1}^{0}}&{0} \\
{0}&0&{I_{K-1}}
\end{matrix} \right)^{\prime }
$ yields the $N\times (2K-1)$ matrix:
\begin{equation*}
\Phi \left(
\begin{array}{cc}
I_{K} & 0 \\
-\lambda _{1}^{0\prime } & 0 \\
0 & I_{K-1}
\end{array}
\right) =\left( \beta _{1}\ \text{{}}\vdots \ \text{{}}\beta _{2}\right)
\left(
\begin{array}{cc}
0 & 0 \\
\Lambda _{2} & I_{K-1}
\end{array}
\right) =\beta _{2}\left(
\begin{array}{cc}
\Lambda _{2} & I_{K-1}
\end{array}
\right) ,
\end{equation*}
which, since the rank of $\beta _{2}\left(
\begin{array}{cc}
\Lambda _{2} & I_{K-1}
\end{array}
\right) $ equals $K-1,$ shows that H$_{0}$ implies that the smallest $K$
singular values of $\Phi $ times $\left(\begin{matrix}
{I_{K}} &{-\lambda _{1}^{0}}&{0} \\
{0}&0&{I_{K-1}}
\end{matrix} \right) ^{\prime }$ equal zero. The sFAR
statistic for testing H$_{0}:\lambda _{1}=\lambda _{1}^{0}:$
\begin{equation*}
\begin{array}{cc}
\text{sFAR(}\lambda _{1})= & \min_{\Lambda _{2}}\text{FAR}(\Lambda
_{1}(\lambda _{1}^{0},\Lambda _{2})),
\end{array}
\end{equation*}
therefore corresponds with a rank test of H$_{0}:$ rank$\left( \Phi \left(
\begin{array}{cc}
I_{K} & 0 \\
-\lambda _{1}^{0\prime } & 0 \\
0 & I_{K-1}
\end{array}
\right) \right) =K-1.\bigskip $
The bounding distribution of the limiting distribution of the sFAR statistic
relies upon a Kronecker product structure (KPS) asymptotic covariance matrix
of the least squares estimator of the linear model (\nocite
{guggenberger2012asymptotic} see Guggenberger et al. (2012)):
\begin{equation*}
\hat{\Phi}=\frac{1}{T}\sum_{t=1}^{T}R_{t}\left(
\begin{array}{c}
\bar{X}_{t} \\
\hat{v}_{t+1}
\end{array}
\right) ^{\prime }\left[ \frac{1}{T}\sum_{t=1}^{T}\left(
\begin{array}{c}
\bar{X}_{t} \\
\hat{v}_{t+1}
\end{array}
\right) \left(
\begin{array}{c}
\bar{X}_{t} \\
\hat{v}_{t+1}
\end{array}
\right) ^{\prime }\right] ^{-1}=\left( \hat{d}\ \text{{}}\vdots \ \text{{}}
\hat{\beta}\right) .
\end{equation*}
The KPS thus concerns the asymptotic variance of
\begin{equation*}
\begin{array}{cc}
\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\left(
\begin{array}{c}
\bar{X}_{t} \\
\hat{v}_{t+1}
\end{array}
\right) \otimes \tilde{e}_{t}, & \tilde{e}_{t}=e_{t}+\beta (\bar{v}_{t}-\hat{
v}_{t}).
\end{array}
\end{equation*}
We note that $\hat{v}_{t}$ is not directly observed so it adds additional
sampling error when imputing estimates of $\hat{v}_{t}.$ To implement the
sFAR test, we therefore make the following assumption.
\begin{assumption}
\label{assum:KPS}There exists $\Omega \in \mathbb{R}^{2K\times 2K}$ and $
\Sigma \in \mathbb{R}^{N\times N}$ symmetric positive definite matrices such
that $S=\Omega \otimes \Sigma $ and
\begin{equation*}
\begin{array}{c}
\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\left( \left(
\begin{array}{c}
\bar{X}_{t} \\
\hat{v}_{t+1}
\end{array}
\right) \otimes \tilde{e}_{t}\right) \rightarrow _{d}N(0,S).
\end{array}
\end{equation*}
\end{assumption}
The asymptotic normality stated in Assumption \ref{assum:KPS} is a direct
result of Assumption \ref{assum: model specification}. If $\hat{v}_{t}$ is
directly observed, so $\tilde{e}_{t}=e_{t}$, Assumption \ref{assum: model
specification} implies that $\Omega =\mathbb{E}\left( (\bar{X}_{t}^{\prime }
\text{ }\vdots \text{ }\hat{v}_{t+1}^{\prime })^{\prime }(\bar{X}
_{t}^{\prime }\text{ }\vdots \text{ }\hat{v}_{t+1}^{\prime })\right) $ and $
\Sigma =\text{var}(e_{t})$. Because of the additional sampling error due to
the generated regressor $\hat{v}_{t+1}$, Assumption \ref{assum:KPS} does,
however, not provide the exact specifications of $\Omega $ and $\Sigma $.
~\bigskip
\noindent {{\small \
\begin{tabular}{c|c|c|c|cc}
\hline\hline
& (1) & (2) & (3) & (4) & \\ \hline
KPST & 212.0808 & 201.8551 & 197.1613 & 205.8054 & \\
$p$-value & [0.0511] & [ 0.1265] & [ 0.1809] & [0.0910] & \\ \hline\hline
\end{tabular}
}} ~\bigskip
\noindent {\small \textbf{Table 3}: KPS test (KPST) statistics for testing
the null hypothesis that $H_{0}:S=\Omega \otimes \Sigma $ for some $\Omega
\in \mathbb{\ \ R}^{2K\times 2K}$ and $\Sigma \in \mathbb{R}^{N\times N}$
symmetric positive definite matrices. All four cases use excess returns on
bonds with maturities 3, 12, 24, 60, 90, 120 months from Adrian et al.
(2013). For the factors $X_{t}$, (1) uses the macro factor (real activity)
and the level factor (first PCA factor), (2) uses the level factor (first
PCA factor) and the slope factor (second PCA factor), (3) uses the macro
factor (real activity) and the slope factor (second PCA factor) and (4) uses
the macro factor (real activity) and the curvature factor (third PCA
factor). \bigskip }
We use the KPS test (KPST) from \nocite{guggenberger2022test}Guggenberger et
al. (2022) to test for the proximity of a KPS matrix to the covariance
matrix, $S.$ Table 3 reports the KPST results, and shows that the KPS
restriction for $S$ is a realistic assumption since none of these tests
reject the null hypothesis that the covariance matrix has a KPS at the 5\%
significance level. A by-product of the KPS test is the KPS factorization
for $\hat{S},$ see \nocite{guggenberger2022test}Guggenberger et al.(2022).
\begin{proposition}
\label{prop:KPS covariance} Under Assumption \ref{assum:KPS}, and when there
is a consistent estimator for $S$, $\hat{S}$, then in large samples $\hat{S}
\approx (\hat{\Omega}\otimes \hat{\Sigma}),$ where
\begin{equation*}
\hat{\Omega}=\displaystyle \text{vecinv}\left( \left(
\begin{array}{l}
\hat{L}_{11} \\
\hat{L}_{21}
\end{array}
\right) /\hat{L}_{11}\right),\hat{\Sigma}=\text{vecinv}(\hat{L}_{11}\hat{
\sigma}_{1}\hat{N}_{1}^{\prime }),
\end{equation*}
and $\hat{L}_{11},\hat{L}_{21},\hat{\sigma}_{1},\hat{N}_{1}$ are specified
in the proof. The KPS covariance estimator $\hat{\Omega}\otimes \hat{\Sigma}$
provides a consistent estimator for $S$.
\end{proposition}
\begin{proof}
See Appendix.
\end{proof}
Because of the KPS covariance structure, $\hat{\Omega}\otimes \hat{\Sigma},$
we can compute the sFAR statistic using the characteristic polynomial stated
in Proposition \ref{prop:sFAR}.
\begin{proposition}
\label{prop:sFAR} Under Assumptions \ref{assum3} and \ref{assum:KPS}, let $
\hat{V}_{\hat{\Phi}}=(\hat{\Psi}\otimes \hat{\Sigma}),$ $\hat{\Psi}=\hat{W}
^{-1}\hat{\Omega}\hat{W}^{-1}=\left(
\begin{array}{cc}
\hat{\Psi}_{X} & \hat{\Psi}_{XV} \\
\hat{\Psi}_{VX} & \hat{\Psi}_{V}
\end{array}
\right) $, $\hat{W}=\frac{1}{T}\sum_{t=1}^{T}\binom{\bar{X}_{t}}{\hat{v}
_{t+1}}\binom{\bar{X}_{t}}{\hat{v}_{t+1}}^{\prime },$ sFAR$(\lambda
_{1}^{0}),$ for testing $\text{H}_{0}:\lambda _{1}=\lambda _{1}^{0},$ for $
\Lambda _{1}=\binom{\lambda _{1}^{\prime }}{\Lambda _{2}},$ $\lambda
_{1}:K\times 1,$ $\Lambda _{2}:(K-1)\times K,$ equals $T$ times the sum of
the $K$ smallest roots of the characteristic polynomial:
\begin{equation*}
\begin{array}{lc}
\left\vert \mu \left(
\begin{array}{cc}
I_{K} & 0 \\
-\lambda _{1}^{0\prime } & 0 \\
0 & I_{K-1}
\end{array}
\right) ^{\prime }\hat{\Psi}\left(
\begin{array}{cc}
I_{K} & 0 \\
-\lambda _{1}^{0\prime } & 0 \\
0 & I_{K-1}
\end{array}
\right) -\right. & \\
\left. \qquad \qquad \qquad \qquad \left(
\begin{array}{cc}
I_{K} & 0 \\
-\lambda _{1}^{0\prime } & 0 \\
0 & I_{K-1}
\end{array}
\right) ^{\prime }\hat{\Phi}^{\prime }\hat{\Sigma}^{-1}\hat{\Phi}\left(
\begin{array}{cc}
I_{K} & 0 \\
-\lambda _{1}^{0\prime } & 0 \\
0 & I_{K-1}
\end{array}
\right) \right\vert & =0,
\end{array}
\end{equation*}
and $\lim_{T\rightarrow \infty }\text{sFAR(}\lambda _{1})\prec \chi
^{2}(K(N-(K-1)).$ The bound on the limiting distribution holds regardless of
Assumption \ref{assum:unspanned factors}.
\end{proposition}
\begin{proof}
See Appendix.
\end{proof}
\noindent
\bigskip
\begin{figure}[htbp!]
\includegraphics[width=\columnwidth,height=0.3\textheight]{fig5_histgram_compare.jpg}
\caption*{{\small \textbf{Figure 4:} simulated density plots of the sFAR test
statistic (shadowed bins) and the density function of the $\chi ^{2}$
-distribution (dashed black curve). The left panel (a) is the strong
identified case, while the right panel (b) is the weak identified case. }}
\end{figure}
\bigskip
Figure 4 illustrates the $\chi ^{2}$ bound on the limiting distribution of
the sFAR statistic stated in Proposition \ref{prop:sFAR}. The density of the
limiting distribution of the sFAR statistic is simulated for a two-factor
model, which uses the first two PCA factors to mimic the strongly identified
case and the third and the fifth PCA factors to mimic weak identification.
The limiting distribution of the sFAR statistic is $\chi ^{2}$ when the
model is strongly identified. In the weakly identified case, as shown in
Figure 4, the limiting distribution is bounded by the $\chi ^{2}$
distribution. Using $\chi ^{2}$ critical values for the sFAR test thus
controls the size of the test. \bigskip
For projection-based tests on $\lambda _{1},$ the involved test has to be
computed over a grid of points concerning the partialled out parameters.
This becomes computationally burdensome when the number of partialled out
parameters increases because of an increased dimension of $\Lambda _{1}$
resulting from more factors. It makes our proposed approach involving the
sFAR test more empirically appealing because it does not involve an
extensive grid search. In practice, Assumption \ref{assum:KPS} can be
relaxed by using the KPST as a pre-test for conducting robust subset testing
as described in Guggenberger et al. (2022).\bigskip
The value of the sFAR statistic at parameter values distant from zero
provide a diagnostic to indicate if the confidence sets of the hypothesized
parameters are bounded. These tests are therefore indicative of weak
identification.
\begin{proposition}
\label{prop:sFAR far away}\label{prop:min boundary} For tests of $\text{H}
_{0}:\lambda _{1}=c\lambda _{1}^{0}$ with $\lambda _{1}^{0}$ a fixed vector
of length one and $c$ a scalar, $\lim_{c\rightarrow \infty }$ sFAR($c\lambda
_{1}^{0}$), the realized value of the sFAR statistic at a distant value of $
\lambda _{1}$ in the direction of $\lambda _{1}^{0}$, equals $T$ times the
sum of the K smallest roots of
\begin{equation*}
\begin{array}{rl}
\left\vert \left(
\begin{array}{cc}
\bar{\lambda}_{1,\perp }^{0} & 0 \\
0 & I_{K}
\end{array}
\right) ^{\prime }\left[ \mu \hat{\Psi}-\hat{\Phi}^{\prime }\hat{\Sigma}^{-1}
\hat{\Phi}\right] \left(
\begin{array}{cc}
\bar{\lambda}_{1,\perp }^{0} & 0 \\
0 & I_{K}
\end{array}
\right) \right\vert & =0,
\end{array}
\end{equation*}
where $\bar{\lambda}_{1,\perp }^{0}$ is a $K\times (K-1)$ orthonormal matrix
that is orthogonal to $\lambda _{1}^{0}$. The limit sFAR statistic is
uniformly bounded from below by the minimum eigenvalue of $T\hat{\Psi}
_{V}^{-1/2\prime }\hat{\beta}^{\prime }\hat{\Sigma}^{-1}\hat{\beta}\hat{\Psi}
_{V}^{-1/2}$.
\end{proposition}
\begin{proof}
See Appendix.
\end{proof}Proposition \ref{prop:sFAR far away} provides a way of verifying
whether the confidence sets resulting from the sFAR statistic are bounded or
unbounded in specific directions (\nocite{dufour1997some}Dufour (1997),
\nocite{kleibergen2020robust}Kleibergen and Zhan (2020), Kleibergen (2021),
\nocite{khalaf2016identification}Khalaf and Schaller (2016), \nocite
{kleibergen2019identification}Kleibergen et al. (2022)). The minimum
eigenvalue of $T\hat{\Psi}_{V}^{-1/2\prime }\hat{\beta}^{\prime }\hat{\Sigma}
^{-1}\hat{\beta}\hat{\Psi}_{V}^{-1/2}$ is a rank test statistic concerning
the rank of the factor loading matrix $\beta $ (see \nocite
{kleibergen2006generalized}Kleibergen and Paap (2006)), so Proposition \ref
{prop:sFAR far away} shows that the sFAR test evaluated at distant values
relates to the rank of $\beta .$ Proposition \ref{prop:sFAR far away} also
explains that when we encounter weak identification issues with $\beta $'s
close to reduced rank, we have unbounded confidence sets. The lower bound is
sharp when $K=1,$ as indicated in the proof of {Proposition} \ref{prop:min
boundary}, for which case also Theorem 12 in \nocite{kleibergen2021efficient}
Kleibergen (2021) applies. When $K=1$ and the Kleibergen-Paap rank test is
significant at the 5\% significance level, Proposition \ref{prop:min
boundary} implies that the sFAR test leads to bounded 95\% confidence sets
of $\lambda _{1}$. \bigskip
Table 4 reports the Kleibergen-Paap rank test for different factor settings
for the data from Adrian et al. (2013). It shows that, in line with Figure
3, all single-factor model have bounded 95\% confidence sets for the
time-varying risk premia, which are less likely to be bounded when we
include more than three factors. The fifth factor, though identified in a
single-factor setting, suffers from weak identification problems when we
include other factors. Table 4 also shows that the rank test statistic is a
good indicator of unboundedness as small values of the rank test statistics
suggest unbounded confidence sets. \newline
~\bigskip \noindent {{\small
\begin{tabular}{ll|ll|ll|ll|ll}
\hline\hline
(1) & rank & (2) & rank & (3) & rank & (4) & rank & (5) & rank \\
& test & & test & & test & & test & & test \\ \hline
1* & 30075 & 1,2 & 1290 & 1,2,3 & 711.1 & 1,2,3, & 932.9 & 1,2,3, & {9.105}
\\
& {[0.000]} & & {[0.000]} & & {[0.000]} & 4 & {[0.000]} & 4,5$\dagger
\dagger $ & {[0.003]} \\
2* & 4409 & 1,3 & 1487 & 1,2,4 & 1317 & 1,2,3, & {5.103} & & \\
& {[0.000]} & & {[0.000]} & & {[0.000]} & 5$\dagger \dagger $ & {[0.078]}
& & \\
3* & 973.0 & 1,4 & 1073 & 1,2,5 & 26.03 & 1,2,4, & 76.53 & & \\
& {[0.000]} & & {[0.000]} & & {[0.002]} & 5 & {[0.000]} & & \\
4* & 703.2 & 1,5 & 39.82 & 1,3,4 & 603.8 & 1,3,4, & {12.36} & & \\
& {[0.000]} & & {[0.000]} & & {[0.000]} & 5 & {[0.002]} & & \\
5* & 74.62 & 2,3 & 864.7 & 1,3,5 & {13.22} & 2,3,4, & 16.23 & & \\
& {[0.000]} & & {[0.000]} & & {[0.004]} & 5 & {[0.000]} & & \\
& & 2,4 & 990.9 & 1,4,5 & 110.3 & & & & \\
& & & {[0.000]} & & {[0.000]} & & & & \\
& & 2,5 & 63.17 & 2,3,4 & 765.8 & & & & \\
& & & {[0.000]} & & {[0.000]} & & & & \\
& & 3,4 & 825.2 & 2,3,5 & 16.96 & & & & \\
& & & {[0.000]} & & {[0.001]} & & & & \\
& & 3,5 & {11.14} & 2,4,5 & 549.0 & & & & \\
& & & {[0.025]} & & {[0.000]} & & & & \\
& & 4,5 & 445.2 & 3,4,5 & 18.05 & & & & \\
& & & {[0.000]} & & {[0.000]} & & & & \\ \hline\hline
\end{tabular}
}} ~\newline
\noindent {\small \textbf{Table 4}: Kleibergen-Paap rank statistic testing $
H_{0}:\text{rank}(\beta )=K-1$ ($K$ denotes the number of factors) and its
associated [$p$-value] in square brackets, for varying factor combinations.
The colum headed by (i), for i=1,\ldots ,5, states which factor combinations
are used when using i factors. All cases use excess returns on bonds with
maturities of 2, 3, 12, 60, and 120 months and different combinations of the
five PCA factors from Adrian et al. (2013). We mark with one star if the
lower bound of the limit sFAR (see Proposition \ref{prop:min boundary})
indicates bounded 95\% confidence sets in every direction, and mark with
double daggers if the associated 95\% confidence sets of the time-varying
risk premia parameters of one or more factor are unbounded. \bigskip }
\subsection{Power of the sFAR test}
To illlustrate the power of the sFAR test, we compute power curves for two
settings calibrated to the data discussed previously. Figure 6 therefore
shows the two-dimensional power curves that result when jointly testing the
two risk premia parameters associated with a single factor in a two factor
model. The left hand side of Figure 6 shows the power curves for a strongly
identified setting while the right hand side does so for a weakly identified
setting. The power curves on the right hand side show that the sFAR test is
not consistent for weakly identified settings since the rejection
frequencies do not converge to one when we move away from the hypothesized
value.\bigskip
\begin{figure}[htbp!]
\includegraphics[width=\columnwidth]{fig8_sFAR_power.jpg}
\caption*{{\small \textbf{Figure 6}: Simulated power surfaces (rejection frequencies)
of sFAR tests on $\Lambda _{1}$'s w.r.t. the first factor of a two-factor
model: the left panel (a) is a strong identified setting calibrated to the
two-factor model with excess returns of bonds with maturities 3, 60, 120
months using the first and second PCA factors , while the right panel (b) is
a weakly identified setting calibrated to the two-factor model with excess
returns of bonds with maturities 3, 60, 120 months using the third and fifth
PCA factors. Dotted lines $(\gamma _{i},y,0.05)$ are drawn to mark the
positions of the calibrated risk premia values, $\gamma _{i}$'s, at 5\%
level.}}
\end{figure}
\subsection{Identification robust confidence sets for risk premia}
We use the sFAR test to construct confidence sets on the risk premia
resulting from two and three factor models. Figure 7 shows the 90, 95 and
99\% joint confidence sets that result for the two risk premia resulting for
one specific factor in a two factor model using the data from Adrian et al.
(2013) while Figure 8 does so for the three risk premia resulting for one
specific factor in a three factor model. Size correct confidence sets for
the individual risk premia result by projecting the joint confidence sets on
the axes. When using four or more factors, the number of risk premia
concerning one factor is at least four so we have to use projection-based
tests based on the sFAR statistic to be able to visualize these confidence
sets.\ For expository purposes and since Table 4 shows that some of these
confidence sets are unbounded, for example, the one that results when using
all five factors, we therefore refrain from using more than three factors.
\footnote{
The rank tests in Table 4 and Figures 7-8 are not identical. For expository
purposes, we choose a smaller number of test assets in Figures 7-8.}
Figure 7 shows all two dimensional confidence sets for the two risk premia
resulting for one factor for all different specifications using the five PCA
factors discussed previously in a two factor model. The two dimensional
confidence sets in Figure 2 vary a lot. Quite a few are empty so all values
of the parameters are rejected at significance levels which exceed 99\%.
This occurs, for example, when using the first and either the second, third
and fourth PCA factor so the model is misspecified. There are also settings
where the confidence set is bounded and well behaved which occurs, for
example, when using the third and fourth PCA factor. Other confidence sets
are unbounded and/or cover the whole two-dimensional space, which occurs,
for instance, when using the third and fifth PCA. For this combination the
90\% confidence set for the two risk premia on the third factor is unbounded
but excludes an area in the parameter space while the 90\% confidence set of
the two risk premia on the fifth factor covers the whole two-dimensional
space. Table 4 also shows that the combination of the third and fifth PCA
factors leads to a smaller rank test statistic than other factor
combinations when using two-factor models which is in line with the
unbounded confidence set in Figure 7 which relates to the third and fifth
PCA factors. This is all indicative of weak identification when using both
the third and fifth factors.
Figure 8 shows the joint confidence sets for the three risk premia
associated with a single factor in a three factor model. The first column of
Figure 8 does for a factor model containing the first three PCAs as factors
while the second column does so using the first, third and fifth PCAs as
factors. Unlike when using two factors, the first column shows that the
confidence sets are no longer empty but bounded which shows that the risk
premia for the first three PCA factors are well identified and that the
model is no longer misspecified. This is confirmed by the p-value of the
rank test on the $\beta $'s. This is in contrast when using the first, third
and fifth PCAs as factors. The confidence sets in the second column of
Figure 8 are namely all unbounded indicating weak identification of the risk
premia which is further reflected by the $p$-value of the rank test on the $
\beta $'s. Table 4 also shows that the model including the first, third, and
fifth PCA factors has a much smaller rank test statistic than one of the
first three PCA factors within the three-factor model.~\newline
\begin{figure}[htbp!]
\includegraphics[width=\columnwidth]{fig6_2D.jpg}
\caption*{{\small \textbf{Figure 7}: Joint confidence sets from the sFAR
test for the two risk premia of the first of the two listed factors in a two
factor model. (yellow 90\%, light green 95\%, light blue 99\%, dark blue
area contains the remaining values). Excess returns on bonds with maturities
3, 60, 120 months are used. [$p$-value] of Kleibergen-Paap rank statistic
testing $H_0: \text{rank}(\beta)=1$ in square brackets.}}
\end{figure}
\bigskip
\begin{figure}[htbp!]
\includegraphics[width=\columnwidth]{fig7_3D_CS.jpg}
\caption*{{\small \textbf{Figure 8}: Joint confidence sets from the sFAR test for
three-factor models (yellow 90\%, light green 95\%, light blue 99\%, dark
blue area contains the remaining values) with excess returns on bonds with
maturities 3, 60, 120 months, (a.i) testing on $\Lambda _{1}$'s w.r.t. the $
i $-th factor when using first, second and third PCA factors factors
Kleibergen-Paap rank statistic testing $H_{0}$ :rank($\beta $)=2 equals
310.9294 [p-value: 0.0000]); (b.i) testing on $\Lambda _{1}$'s w.r.t. the }$
( ${\small $i+2)$-th factor when using first, third and fifth PCA factors
factors. Kleibergen-Paap rank statistic testing $H_{0}$ :rank($\beta $)=2
equals 0.0654 [p-value 0.7981]).} }
\end{figure}
\bigskip
Figure 8 shows the three dimensional confidence sets that result from the
sFAR test. It results from partialling out the six risk premia associated
with the other factors. Hence, when we compute these confidence sets using
projection with the identification robust tests, we have to specify a
nine-dimensional grid for the risk premia and compute the identification
robust tests for all values on this nine-dimensional grid. This is, or is
close to be, computationally infeasible. Hence, the sFAR test provides a
computationally tractable manner to conduct identification robust tests on
larger number of parameters.
\section{Conclusion}
We propose identification robust test procedures for testing hypotheses on
risk premia in dynamic affine term structure models. The robust subset
factor Anderson-Rubin test extends the sFAR test from the linear asset
pricing model to allow for tests on multiple risk premia and, unlike
projection based testing, provides a computationally tractable manner to
conduct identification robust tests on larger number of parameters. Our
empirical results show that especially in case of multiple factors, weak
identification is pervasive and traditional tests are likely misleading. We
use the empirical settings from the literature on affine term structure
models, see e.g. \nocite{adrian2013pricing} Adrian et al. (2013) and \nocite
{ang2003no}Ang and Piazzesi (2003)), to illustrate our results and the
importance of using weak identification robust test procedures.\newpage
\bibliographystyle{ier}
\bibliography{bibtexrefs}
\newpage