EconBase
← Back to paper

Weak (Proxy) Factors Robust Hansen-Jagannathan Distance For Linear Asset Pricing Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

80,167 characters · 14 sections · 87 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Weak (Proxy) Factors Robust Hansen-Jagannathan Distance For Linear Asset Pricing Models

\linespread{1.3}

\thispagestyle{empty}

abstract{ The Hansen-Jagannathan (HJ) distance statistic is one of the most dominant measures of model misspecification. However, the conventional HJ specification test procedure has poor finite sample performance, and we show that it can be size distorted even in large samples when (proxy) factors exhibit small correlations with asset returns. In other words, applied researchers are likely to falsely reject a model even when it is correctly specified. We provide two alternatives for the HJ statistic and two corresponding novel procedures for model specification tests, which are robust against the presence of weak (proxy) factors, and we also offer a novel robust risk premia estimator. Simulation exercises support our theory. Our empirical application documents the non-reliability of the traditional HJ test since it may produce counter-intuitive results, when comparing nested models, by rejecting a four-factor model but not the reduced three-factor model, while our proposed methods are practically more appealing and show support for a four-factor model for Fama French portfolios. }

Keywords: asset pricing; identification robust statistics; reduced-rank models; model misspecification; rank test

\doublespace

\setcounter{page}{1} \pagenumbering{arabic}

Introduction

Linear factor models have gained tremendous popularity in the empirical asset pricing literature, see e.g., fama1993common, lettau2001consumption, kan2004hansen, kan2008model. The low dimensional factor structure in asset returns is well-documented (e.g., kleibergen2015unexplained), and harvey2016and list hundreds of papers with factors that attempt to explain the cross-section of expected returns. Since so many factors are introduced, the proposed factors are at best proxies for some unobserved common factors, and the asset pricing models are at best approximations. Therefore, it is more appealing to determine whether or not the data reject a model, namely how good a model can approximate the data than to identify important factors (or factors with significant risk premia). The assessment of model performance is where specification tests play a role. To evaluate these factors and diagnose the model specifications, the HJ distance, proposed in hansen1997assessing, has emerged as one of the most dominant measures of model misspecification in the empirical asset pricing literature (e.g., jagannathan1996conditional, kan2004hansen).

However, this paper shows that the conventional HJ statistic can be unreliable. Previous studies have shown that the lack of model identification can lead to spuriously significant risk premia (kleibergen2009tests, 2016, anatolyev2018factor), and a misleading gauge of model fit based on the second pass $R^2$ (kleibergen2015unexplained). This paper demonstrates that when models are weakly identified the HJ statistic does not measure model fit satisfactorily and the specification test via the HJ statistic, which we refer to as the HJ specification test in the sequel, is not reliable. In an empirically relevant setting where proxy factors weakly correlate with the unobserved common factors, the HJ specification test can be size distorted even in large samples. The boundary, determined via the HJ specification test, between correct model specifications and misspecifications begins to blur in these so-called weak identification cases. Another potential issue would be the omitted-strong-factor problem. The resulting strong cross-sectional dependence in the error term can exaggerate all sorts of distortions when some included (proxy)\footnote{Sometimes researchers consider a latent factor structure in asset returns and regard included factors in empirical studies as proxies for priced latent common factors (e.g., kleibergen2015unexplained, giglio2017inference), sometimes factors are assumed to be directly observed and priced weak common factors lead to problems (e.g., kleibergen2009tests, anatolyev2018factor). We mostly adopt the former idea, but our discussions and methods are also valid in the latter case, and we emphasize this by enclose the term proxy in brackets.} factors are weak (kleibergen2015unexplained). One of the reasons for these failures is that sampling errors are no longer negligible asymptotically in the presence of weak (proxy) factors. Therefore, the conventional asymptotic justification may fail in empirically relevant settings, as weak (proxy) factors are commonly observed in many recent studies (e.g., kleibergen2009tests, anatolyev2018factor).

This paper not only shows the potential failure of the HJ test, but also aims to improve the performance of specification tests. This contributes to the literature on providing the identification robust statistical tools. Recent papers have developed different techniques to incorporate some of these aforementioned issues, most of which focus on the identification and inference of risk premia. 2016 provides an estimation approach using shrinkage-based dimension-reduction technique which excludes weak/useless (proxy) factors. anatolyev2018factor propose an estimation procedure based on split-sample instrumental variables regression with proxies for the missing factor structure. giglio2017inference propose a three-pass estimation procedure and bypass the omitted factors bias by projecting risk premia of observed factors on those of strong factors extracted via principal components analysis (PCA). Alongside with these estimation techniques there are identification robust test statistics to correct for the overly optimistic statistical inference of the risk premia (e.g. kleibergen2009tests, kleibergen2019Consumption, kong2018). As for the specification tests of asset pricing models, gospodinov2017spurious discuss the potential power loss of the $\mathcal{J}$ specification test when spurious/useless factors, which are completely uncorrelated with asset returns, are present.

This paper focuses on the robust model specification tests, and provides two easy-to-implement specification test procedures to remedy the size distortion of the HJ test resulting from weak (proxy) factors that are minorly correlated with asset returns. The first proposed test procedure is a two-step Bonferroni-type method, and it is robust against identification issues when the number of asset returns is limited. This method takes into account the identification strength via a first-step confidence set, and we verify that it improves power compared with the $\mathcal{J}$ test. The second approach relies on a novel four-pass estimator, and the test procedure provides valid inference results in an asymptotic framework where the number of assets is comparable to the number of the observation periods. Our proposed four-pass estimator directly leads to a novel risk premia estimator, and thus we also contribute to the literature on estimation of risk premia in the presence of weak (proxy) factors and omitted factors. For linear asset pricing models, the conventional approach for estimating risk premia is known as the Fama-Macbeth (FM) two-pass estimation procedure (fama1973risk), where risk premia estimates result from regressing average asset returns on first-pass estimated risk exposures (factor loadings $\beta$'s). The two-pass procedure is easy to implement but can result in unreliable estimates and inference when some included factors are not strongly correlated with asset returns such that their risk exposures do not dominate corresponding sampling errors (kleibergen2009tests, anatolyev2018factor), which resembles the failure of the 2SLS estimator in instrumental variable regression when instruments are weak. Besides, anatolyev2018factor show that the missing factor structure exacerbates the weak (proxy) factor problem. We show our risk premia estimator is robust to the presence of both weak (proxy) factors and missing factors.

Our empirical application documents the strange behavior of the HJ test. Counter-intuitively, it can reject a four-factor model but not the corresponding three-factor model nested within the four-factor model. We attribute this behavior to the additional fourth factor being a weak proxy factor which leads to a undesirably high rejection rate of the HJ test. Our proposed procedures do not have this problem and reflect the factor structure in asset returns in a more informative way. \\

The paper is organized as follows: Section (ref) reviews the basic model setting and shows the drawbacks of the HJ statistic; Section (ref) and (ref) introduce our proposed model specification test procedures, where Section (ref) discusses our two-step Bonferroni-like method and Section (ref) considers an approach that is valid with a double-asymptotic framework; Section (ref) presents results of our empirical application.

This paper uses the following notation: $P_X$ stands for $X(X'X)^{-1}X'$ for a full column rank matrix $X$, $M_X$ for $I-P_X$, $X^{\frac{1}{2}}$ for the upper triangular matrix from the Cholesky decomposition of the positive definite matrix $X$ such that $X=(X^{\frac{1}{2}})'X^{\frac{1}{2}}$. Besides, in the following discussion, the notation would be more precise with $N, T$ in the sub- or superscripts. For example to model the weak (proxy) factors, it might be reasonable to use notation such as $\beta_{T,N}, d_{g,T,N}, \gamma_{N,T}, \theta_{g,N,T}$ since the parameter values may change according to the sample dimensions in order to model the local to zero behavior. To avoid a more cumbersome notation, we ignore these subscripts when there can be no misunderstanding.

Models and Problems

This section introduces linear asset pricing models and the conventional model selection and specification test procedure based on the HJ distance. We use the term HJ statistic to denote the conventional squared HJ distance estimator, and to distinguish it from the other two estimators, our so-called HJN and HJS statistics, proposed in Sections 3 and 4. We start by introducing our baseline model setting. We next derive asymptotic properties of the HJ statistic in the presence of weak (proxy) factors to clarify the problems we focus on.

Baseline model setting

We work with the linear asset pricing model because of its popularity in empirical studies. It imposes that all asset returns share common risk factors described by a small set of proposed factors. We regard the proposed factors in empirical studies as proxies for latent ones in the form as suggested in (kleibergen2015unexplained). Assumption (ref) summarizes the baseline model setting:

assumptionFor the $N\times 1$ vector of asset gross returns $r_t$, we assume that \begin{align} r_t = c + \beta f_t + u_t, \end{align} with $f_t$ a $K\times 1 $ vector of (possibly) unobserved zero-mean factors, $u_t$ an $N\times 1$ vector of idiosyncratic components, and with a $K\times 1$ vector of proxy factors $g_t$ \begin{align} f_t = & d_g\left( g_t - \mu_g \right) + v_t, \end{align} where $\mu_g=\mathbb{E}g_t$, $g_t$ is uncorrelated with $u_t,v_t$, $d_g$ is of full rank and $g_t,v_t,u_t$ are stationary with finite fourth moments. Furthermore, \begin{align} c= \iota_N \lambda_0 + \beta \lambda_f, \end{align} where $\lambda_0\neq 0 $ is the zero-beta return, $\lambda_f$ is a $K\times 1$ vector of risk premia, and the parameter space of $\lambda_0,\lambda_f$ is compact.

Assumption (ref) describes the beta representation of linear asset pricing models, and the DGP is similar to the one employed in kleibergen2015unexplained. The moment conditions (or in other words the structural assumptions imposed on the constant term), $c= \iota_N \lambda_0 + \beta \lambda_f$, are commonly used in linear asset pricing model (e.g. cochrane2009asset). If $f_t$ is observed, then we would have perfect proxies with $d_g=I_K$, $v_t=0$. Therefore, this model setting also embeds the model specification used in e.g. kleibergen2009tests, anatolyev2018factor where factors $f_t$ are assumed to be observed. Using the observed factors $g_t$, model ((ref)) can be rewritten as

align[align omitted — 75 chars of source]

with $\beta_g = \beta d_g, \bar{g}_t= g_t-\bar{g}, \bar{g}= \sum_{t=1}^T g_t /T, \widetilde{u}_{g,t} = {u}_{g,t}+ \beta_g \left( \bar{g}- \mu_g \right) u_{g,t} = \beta v_t + u_t$. The estimation of the risk premia $\lambda_g$ is usually accomplished by the FM two-pass estimator (fama1973risk, shanken1992ebp). In the first pass, the risk exposures ${\beta}_g$ are estimated by regressing asset returns $r_t$ on a constant and factors $\bar{g}_t$, and in the second pass the FM estimator $\widehat{\lambda}_g$ results from regressing average asset returns $\bar{r}= \sum_{t=1}^T r_t/T$ on an $N\times 1$ unity vector $\iota_N$ and the risk exposure estimates $\widehat{\beta}_g$.

Another well-known representation of asset pricing models is the stochastic discount factor (SDF) representation, based on which the HJ distance is defined. cochrane2009asset shows that for linear asset pricing models, there is a corresponding SDF $m_t$ that is linearly spanned by the latent factors

align[align omitted — 65 chars of source]

with $F_t=(1,f_t')'$, and re-scaled risk premia $\theta= (1/\lambda_0, -V_{ff}^{-1}\lambda_f/ \lambda_0)$. The moment conditions ((ref)) are then equivalent to the following ones

align[align omitted — 89 chars of source]

with $\iota_N$ an $N\times 1 $ unity vector with all entries equal to one. The population pricing errors which are the deviations from the moment conditions ((ref)) are denoted by

align[align omitted — 107 chars of source]

With a linear SDF ((ref)) $e(\theta)=\iota_N -q\theta$, for $q=\mathbb{E}\left(r_tF_t' \right) $ a full column rank $N\times K$ matrix.

hansen1997assessing (HJ) propose the minimum distance between the SDF of an asset pricing model and a set of correct SDFs as a measure of model misspecification. It also serves as a measure of goodness-of-fit. A smaller value of the HJ distance indicates a better model fit, and this is used for model selection. The population squared HJ distance $\delta$ has an explicit expression:

align[align omitted — 67 chars of source]

with $Q_r= \mathbb{E}\left(r_tr_t' \right)$ a full column rank $N\times N$ matrix. With a linear SDF, after some simple algebra, we can write the squared HJ distance explicitly as $\iota_N' \left(Q_r^{-1}-Q_r^{-1} q \left( q'Q_r^{-1} q \right)^{-1} q'Q_r^{-1} \right) \iota_N,$ which is also numerically equal to $\iota_N' (Q_r^{-1}-Q_r^{-1} B (B'Q_r^{-1}B)^{-1} B'Q_r^{-1} ) \iota_N$ with $B=(c,\beta)$, and it is zero if and only if moment conditions ((ref)) hold. Given the observed proxy factors $g_t$, the sample counterpart of the squared HJ distance, the HJ statistic, is

align[align omitted — 116 chars of source]

In a linear asset pricing model, the sample pricing errors $ e_{g,T}(\theta_G) =\sum_{t=1}^{T} e_{g,t}(\theta_G)/T =\iota_N - q_{G,T}\theta_G,~ e_{g,t}(\theta_G)= \iota_N -r_tG_t'\theta_G,~ q_{G,T}= \sum_{t=1}^{T} r_t G_t'/T,~ \widehat{Q}_{r} = \sum_{t=1}^{T} r_tr_t'/T$, $G_t=(1,\bar{g}_t')'$, and the estimator resulting from this quadratic optimization problem is

align[align omitted — 122 chars of source]

jagannathan1996conditional propose the HJ specification test by testing the moment conditions ((ref)) via the HJ statistic. Under the null hypothesis that the moment conditions ((ref)) hold, jagannathan1996conditional show that the asymptotic distribution of the HJ statistic follows a weighted sum of $\chi^2(1)$ random variables, which is because of the weighting matrix used in the HJ statistic. If we weight by the long-run covariance matrix of the sample pricing errors, we would have a regular chi-square-type limiting distribution. However, since we weight the HJ statistic differently with the second moment of asset returns as the weighting matrix, each of these $\chi^2(1)$ random variables has a weight different from one. Therefore, the critical values for the HJ statistic are obtained from the weighted sum of $\chi^2(1)$ random variables

align[align omitted — 72 chars of source]

with $x_i$ being independent $\chi^2(1)$ distributed random variables, and $p_i$ being the positive eigenvalues of the matrix

align*[align* omitted — 205 chars of source]

with $\widehat{S} = S_T(\widehat{\theta}_G)$ and $S_T({\theta}_G)$ a consistent estimator of the long-run variance matrix of the sample pricing errors $e_{g,T}(\theta_G)$ (for example, one may simple choose $S_T({\theta}_G) = \frac{1}{T}\sum_{t=1}^T e_{g,t} ({\theta}_G)e_{g,t} ({\theta}_G)'$, $e_{g,t}({\theta}_G)= \iota_N -r_tG_t'{\theta}_G$ provided that Assumptions (ref)-(ref) hold).

Problems and asymptotic properties

Our interest lies in the performance of the HJ statistic in the presence of weak identification issues, in particular, when observed proxies $g_t$ are only weakly correlated with asset returns. ahn2004small document the poor finite sample performance of the HJ specification test, and they argue that the size distortion is due to the critical value of the test which requires the estimation of the covariance matrix of $e_t(\theta)$ that performs badly with a limited number of observation periods. This is also consistent with the findings in kleibergen2019Consumption, kong2018. In later parts, we show that not only in finite samples but also in large samples, the HJ specification test can be severely size distorted in the presence of weak identification issues. We focus on two issues that can cause the potential deficiency of the HJ statistic.

1. Weak (proxy) factors. The HJ statistic depends on the estimator $\widehat{\theta}_G$. This is a GMM estimator based on the moment conditions ((ref)) with weighting matrix $\widehat{Q}_r$. Similar to the FM risk premia estimator, this estimator can be constructed in two steps. In the first step the regressor in the SDF, $q_g$, is estimated via the estimator $q_{G,T}$, and in the second step $\widehat{\theta}_G$ results from regressing $\widehat{Q}_r^{-\frac{1}{2}}\iota_N $ on the first stage estimates $\widehat{Q}_r^{-\frac{1}{2}}q_{G,T}$. This close link with the FM estimator raises concern for the quality of the $\theta_G$ estimator.

The FM estimator is unreliable under weak identification (e.g., kan1999two, kleibergen2009tests, kleibergen2015unexplained, kleibergen2019Consumption, kong2018, anatolyev2018factor). For linear asset pricing models, the identification strength is reflected by the rank of $B_g=(c,\beta_g)$ (e.g. kleibergen2019Consumption). The weak identification issues result from the empirical observation that the matrix $B_g$ might be of reduced rank or near reduced rank. For example, this can happen when some (proxy) factors used in the estimation are weakly correlated with the asset returns. One way of modeling these weak (proxy) factors is to consider a sequence of models (or a sequence of parameter values) such that along the sequence, factor loadings are smaller and thus less informative for identifying risk premia. For example, suppose the $\beta_g$ matrix is small, modeled by a drifting to zero sequence of order $O(1/\sqrt{T})$, then the sampling errors in the first stage estimator $\widehat{\beta}_g$, which are of the same order, are no longer negligible. These non-negligible sampling errors lead to the asymptotic invalidity of the FM estimator under weak (proxy) factors.

Following the same reasoning, the asymptotic justification for the estimator $\widehat{\theta}_G$ fails when $q_g$ is small and comparable to its sampling error (see Theorem (ref)). Since $q_g=\beta_g V_g$ with $V_g =\mathbb{E}g_tg_t'$, we model weak (proxy) factors using drifting to zero risk exposures (see Assumption (ref)) to mimic the behavior of a small $q_g$, which is in line with the literature on weak factors (e.g. kleibergen2009tests).

2. The missing factor structure. Omitted factors have received attention in recent studies (e.g., kleibergen2015unexplained, giglio2017inference, anatolyev2018factor). When we work with observed factors $g_t$ in a latent factor setting, equation ((ref)) suggests that the omitted factors $v_t$ contribute to the error term $u_{g,t}$, and we also allow that unobserved factors explain most of the cross-sectional dependence in $u_{g,t}$. Similar to the discussion of the FM two-pass risk premia estimator in anatolyev2018factor , the missing factor structure could exacerbate the problem caused by the weak (proxy) factors and enlarge the bias in the estimator $\widehat{\theta}_g$ (see Theorem (ref)), as the presence of an unobserved (missing) factor structure in the error terms creates the classical omitted-variables problem in the second step regression of the $\widehat{\theta}_g$ estimator when some (proxy) factors are weak.

Therefore, the HJ statistic may use an estimate that is potentially far away from the true value, and thus selection and inference based on the HJ statistic can be misleading. Before we continue to verify this, we make two assumptions.

assumptionSuppose $u_t$ can be decomposed into two parts: a missing factor structure with a $K_z\times 1$ ($K_z\geq 0$) vector of unobserved strong factors $z_t$ and weakly cross-sectional correlated noise $e_t$; \begin{align} u_t = \gamma z_t + e_t, \end{align} where (i) $e_t$ (with mean zero and bounded fourth moment $\sup_{i}\mathbb{E} e_{it}^4 < L< \infty$) is independent from $e_s, s\neq t$ and $g_{t'}, \forall t'$; (ii) denote $\Omega_e = \mathbb{E}e_te_t' $, then $\lim_{N,T} \text{tr}\left( \Omega_e \right)/N =a>0$ and $ 0<l<\lim\inf_{N,T} \lambda_{\min} \left( \Omega_e \right) <\lim\sup_{N,T} \lambda_{\max}\left( \Omega_e \right) < L < \infty$ with $\lambda_{\min}(X)$ the smallest eigenvalues of matrix $X$ and $\lambda_{\max}(X)$ the largest eigenvalues of matrix $X$; (iii) $\mathbb{E}\left| \frac{1}{\sqrt{N}} \sum_{i=1}^{N} \left(e_{it}^2- \mathbb{E}e_{it}^2\right) \right|^4 < L < \infty$.
assumptionDenote $Q_r$ by the second moments of $r_t$, $\eta =\left(\eta_{B_g}, \gamma \right) $ with $\eta_{B_g}=\left(c, \eta_{\beta_g} \right), \eta_{\beta_g}=\left(\beta_{g,1} , \sqrt{T}\beta_{g,2} \right), \beta_{g,i} = \beta d_{g,i}, i=1,2$ being of dimension $N\times K_{g,i}, i=1,2$ ($K_{g,1}+K_{g,2}=K, K_{g,2}\geq 0$) and full column rank matrices. (i) for fixed $N$, we assume $ \eta ' \eta $ is a $(1+K+K_z )\times (1+K+K_z )$ positive definite matrix; (ii) as $N,T$ approach to infinity, $ N^{-1}\eta'\eta $ converges to a $(1+K+K_z )\times (1+K+K_z )$ positive definite matrix.

Our framework involves the observed factors $g_t$, and the omitted ones $v_t, z_t$. They have factor loadings $\beta_g, \beta, \gamma$ respectively. Assumption (ref) specifies the strengths of these factors. The loadings, $\beta_{g,2}$, of the $K_{g,2}$ proxy factors $g_{2,t}$ are modeled as drifting to zero sequences, so we call $g_{2,t}$ weak proxy factors. We do not restrict the strength of the priced latent factors $f_t$, and allow for weak priced latent factors. This assumption resembles the factor loading assumption in anatolyev2018factor, but our risk exposure matrix $(\beta_g, \beta, \gamma)$ is of reduced rank.

Assumption (ref), which is similar to assumptions in onatski2012asymptotics and anatolyev2018factor, does not fully rule out the cross-sectional dependence in the idiosyncratic error term $e_t$. This assumption allows the explanatory power of the cross-sectional variation in $e_t$ to be comparable to the weak proxy factors when N,T increase proportionally, and the weak identification issue appears when the explanatory power of the proxy factors are roughly of the same order as $e_t$. When the cross sectional size $N$ is fixed, Assumptions (ref).(ii)-(iii) imposed on the noise term $e_t$ hold naturally as long as Assumption (ref).(i) holds and we can not really distinguish weak and strong factors. In later parts of this paper when $N$ is fixed, we do not make use of the assumptions imposed on $e_t$ but only the independence assumption (Assumption (ref).(i)). We assume that $e_t$ is independent across periods which is consistent with the efficient market hypothesis, and since empirical studies mostly use monthly or even less frequent data, this is not a unrealistic assumption.

lemSuppose Assumption (ref), (ref) hold, let T increase to infinity then \begin{align*} \widehat{B}_g =_d \left(c, \beta_g\right) + \psi_{B_g}/\sqrt{T} \end{align*} where $\widehat{B}_g = q_{G,T}\widehat{Q}_G^{-1}, \widehat{Q}_G=\sum_{t=1}^TG_tG_t'/T, \psi_{B_g}=\left(\psi_c, \psi_{\beta_{g,1}}, \psi_{\beta_{g,2}}\right) $ and $\text{vec}(\psi_{B_g})$ being zero-mean normal random vectors. Proof: See Appendix (ref)}.
theoSuppose Assumptions (ref) and (ref)-(ref) hold, let T increase to infinity with fixed N, then the behavior of the HJ statistic is characterized by: \begin{align*} T \widehat{\delta}^2_g \rightarrow_d \widetilde{\psi}_{B_g} ' M_{Q_r^{-\frac{1}{2}} \left( \eta _{B_g} + (0\vdots \psi_{\beta_{g,2}} ) \right) } \widetilde{\psi}_{B_g} \end{align*} with $\widetilde{\psi}_{B_g}= \psi_{B_g} Q_G \theta_G $.\\ Proof: See Appendix (ref)}.

Theorem (ref) is derived assuming that the linear model is correctly specified, and it suggests that with strong factors, $K_{g,2}=0$, the weighted sum of $\chi^2$'s provides a reasonable approximation (Corollary (ref)).

coroSuppose Assumption (ref) and the assumptions in Theorem (ref) hold with $K_{g,2}=0$, then \begin{align*} T \widehat{\delta}^2_g \sim_d \sum_{i=1}^{N-K-1} {p}_i x_i \end{align*} with $x_i$ being independently $\chi^2(1)$ distributed random variables and $p_i$ being the positive eigenvalues of the matrix $\widehat{S}^{\frac{1}{2}} \left(\widehat{Q}_r^{-1}-\widehat{Q}_r^{-1}q_{G,T} \left(q_{G,T} '\widehat{Q}_r^{-1} q_{G,T} \right)^{-1}q_{G,T} '\widehat{Q}_r^{-1} \right) \widehat{S}^{\frac{1}{2}'}$. \\ Proof: This is a direct result of Theorem (ref).

However, the conventional specification test procedure can be unreliable and suffer from severe size distortion even in large samples due to the irregular distribution of the HJ statistic (Corollary (ref)).

coroSuppose the assumptions in Corollary (ref) hold with $0< K_{g,2} \leq K$, \begin{align} \lim\sup_{N} \lim\sup_{T} \mathbb{P} \left( T\widehat{\delta}_g \geq \widehat{c}_{1-\alpha} \right) =1 \nonumber \end{align} where $\widehat{c}_{1-\alpha} $ is the conventional critical value derived from the distribution ((ref)).\\ Proof: See Appendix (ref)}.

Corollary (ref) shows that under certain conditions the conventional specification test rejects the model specification with probability converging to one even when the moment conditions hold. Thus the conventional specification testing procedure based on the HJ statistic may mistake the "weak identification" resulting from the weak (proxy) factors for model misspecification, and leads to over-rejection when models are correctly specified.

Simulation exercises

We conduct simulation exercises to show that the HJ distance statistic as a model selection criterion might favor the presence of useless factors, and the HJ specification test suffers from severe size distortions.

In the first simulation exercise (Figures (ref), (ref)), we calibrate the data generating process to match the data set of monthly gross asset returns on 25 size and book to market sorted portfolios from 1963 to 1998 and the three Fama French (FF) factors used by lettau2001consumption. The data is simulated in the following way: we simulate three proxy factors $g_t \sim ~i.i.d~ N(0,V_F)$, three omitted factors $v_t~ i.i.d~\sim N(0.99 V_F)$ and three strong factors are then generated by $f_t=0.1*g_t + v_t,$. $V_F$ is calibrated to the sample covariance of the FF factors. We also generate three completely useless factors $w_t\sim i.i.d. N(0,V_F)$. We then generate returns via $r_t=\iota_N + \beta \lambda+\beta f_t +u_t$, $u_t\sim ~i.i.d ~N(0,V_u)$, where we set $\lambda$ to be the sample risk premia estimated via the FM two-pass estimator, $\beta$ is the sample slope parameter between the assets returns and FF factors, and $V_u$ is the sample covariance of the residuals resulting from regressing asset returns on a constant and FF factors from the data.

Figure (ref) compares the density functions of the simulated HJ statistics evaluated with various combinations of the factors $g_t, f_t, w_t$. For example, the black solid curve is drawn using three strong factors $f_t$, the black dashed curve is drawn with two strong factors $f_{1t}, f_{2t}$. Ideally, the black solid curve should be the most left, since the model with three strong factors should be most likely to be selected by the HJ statistic. However, comparing the red solid and black solid curves shows that adding additional useless factors leads to a shift of the distribution to the left, so it reduces the HJ statistic and leads to a "preferred model". The blue sold curve illustrates the density function of the HJ statistic of the model with three weak proxy factors. By construction, the moment conditions ((ref)) are satisfied by the three weak proxy factors, and this model is correctly specified. If we compare the blue solid curve with the blacked dashed one which is constructed with only two strong factors, then the misspecified model with two strong factors is more likely to be selected. These observations imply that the HJ statistic is not a satisfying model selection tool.

figure[figure omitted — 586 chars of source]

The observations in Figure (ref) show that values of the HJ statistic can not properly distinguish between weakly identified models and misspecified models. This is what motivates us to look further into the HJ specification test. Figure (ref) compares two different approaches for approximating the distribution of the HJ statistic: one uses the conventional weighted sum of $\chi^2$'s, from which the critical values of the HJ statistic result, and another one uses the infeasible distribution from Theorem (ref). The left-hand side panels of Figure (ref) use three strong factors, while the right-hand side panels use three week ones.

The upper panels of Figure (ref) show that both approximations for the distribution of the HJ statistic (the conventional weighted sum of $\chi^2(1)$s and the infeasible one from Theorem (ref)) are bad when T is small, and shift to the left compared with the density function of the HJ statistic. This observation is consistent with the one in ahn2004small that the HJ specification test over-rejects correct model specifications in small samples. With a limited number of observation periods, not only sampling errors in the $q_G$ estimators need to be taken into account but also those of other estimators such as the covariance estimator $\widehat{S}$ (e.g. kleibergen2019Consumption, kong2018). The infeasible distribution improves slightly by taking into account the sampling errors in the $q_G$ estimators. When T is large, the randomness in the covariance estimators becomes small, but the sampling errors in $q_{G,T}$ still matter when proxy factors are weak. As shown in the lower panels of Figure (ref), with a larger sample size the conventional approximation works fine when factors are strong but not when weak proxy factors are present. With weak proxy factors, the distribution of the HJ statistic is not properly approximated by the weighted sum of $\chi^2$'s even in large samples, and the HJ specification test is still likely to over-reject models when moment conditions do hold (Corollary (ref)).

figure[figure omitted — 770 chars of source]

Our second simulation exercise considers a simple single factor model in order to further illustrate the size distortion of the HJ specification test. We calibrated parameters to the data set from kroencke2017asset. We simulate the proxy factor $g_t\sim ~i.i.d.~N(0,V_f/4)$, the omitted factor $v_t \sim ~i.i.d.~ N(0, Vff- d_g*(V_f/4)*d_g')$, the latent factor $f_t =d_g g_t + v_t$ and thus the variance of the $f_t$ remain unchanged to different values of the $d_g$. We calibrate $V_f$ to match the sample variance of the consumption growth factor from the data. The factor proposed in kroencke2017asset has been shown to be weak (e.g. kleibergen2019Consumption). Therefore, we choose $\beta$ to be of $10\widehat{\beta}$ with $\widehat{\beta}$ the sample regression parameter from kroencke2017asset, and thus $r_t$ is generated with one single strong factor $f_t$ via $r_t=\iota_N + \beta \lambda + \beta f_t +u_t$, $u_t\sim ~i.i.d ~N(0,V_u)$, where we match $\lambda$ to the estimated risk premium from kroencke2017asset. We arbitrarily choose $d_g=1.9;0.9$ to mimic a strong and weak proxy factor.

Table (ref) shows that the HJ specification tests have poor finite sample performances, size distortions increase with the number of assets and with relative weak proxy factors the distortion is more severe, and these observations support Corollary (ref).

table[table omitted — 712 chars of source]

Specification test with limited N: HJS

As in previous discussions, the HJ specification test can not provide valid inference when weak (proxy) factors are present, and this is because the HJ specification test procedure ignores some non-negligible sampling errors in the estimates of parameters that can not be properly identified in the presence of weak identification issues. In this section, we suggest a numerically simple and identification robust test procedure which replaces the estimates of these parameters with potential identification issues by those lying in a robust confidence set. This approach is related to the widely studied weak instrument problem, where confidence sets with asymptotically correct coverage can be constructed for parameters with potential identification issues (e.g. kleibergen2005testing, mikusheva2010robust).

HJS specification test

Our proposed HJS specification test procedure is conducted in three steps:

Step (1): Construct an identification robust confidence set, $CS_{r, \alpha_1} $, for $\theta_G$ by inverting an Anderson-Rubin (AR) type test statistic (e.g., kleibergen2009tests, gospodinov2017spurious):

align[align omitted — 170 chars of source]

with $c_{1-\alpha_1}$ the $100(1-\alpha_1)\%$ percentile of the $\chi^2(N)$ distribution.

Step (2): Compute the HJS statistic:

align[align omitted — 101 chars of source]

with $\delta_{g,T} = (\theta) e_{g,T}(\theta_G)'\widehat{Q}_{r}^{-1} e_{g,T}(\theta_G) $. To complete the construction of the HJS statistic we set $\widehat{\delta}_g^* = \infty$ when the confidence set $ CS_{r, \alpha_1} $ is empty.

Step (3): This test would then reject the null hypothesis that moment conditions ((ref)) hold if $$T\widehat{\delta}^*_{g}> c^*_{1-\alpha},$$ and the critical value is:

align[align omitted — 95 chars of source]

where $c^*_{1-\alpha}(\theta)$ is the $100(1-\alpha_1)\%$ percentile of the weighted sum of $N$ $\chi^2(1)$ random variables with weights being the non-zero eigenvalues of $ S^{\frac{1}{2}}_T( {\theta}) \widehat{Q}_r^{-1} S^{\frac{1}{2}'}_T( {\theta})$, $\alpha_1,$ $ \alpha_2$ are chosen such that $\alpha_1>0,$ $ \alpha_2>0$ and $\left(1-\alpha_1 \right)\left(1-\alpha_2 \right)=1-\alpha$ with $\alpha$ the overall significance level.

Our HJS specification test procedure combines a less powerful but robust statistic (AR) with a non-robust one (HJ) to incorporate the model identification strength in our testing procedure, and in later discussion we show that this test improves performance in size (compared with the HJ test) and power (compared with the $\mathcal{J}$ test).

Before we proceed to show the size and power performances of the HJS specification test (Theory (ref) and Theory (ref) ), we first discuss the properties of the robust confidence set $CS_{r,\alpha}$. kleibergen2019Consumption study a similar robust risk premia confidence set using the GRS-FAR statistic. They show that this kind of set can be unbounded in certain cases. Therefore, for practical reason, we restrict the parameter space $\Theta$ to be a compact set (Assumption (ref)), of which the robust confidence set is a subset. By construction, when the model is strongly identified we would expect the confidence set to shrink to a point as sample size grows, and when the model is weakly identified or even unidentified the diameter of this set can be arbitrarily large.

lemSuppose the assumptions in Corollary (ref) hold, then \begin{align*} \lim\inf_T \mathbb{P}\left(\theta_G\in CS_{r,\alpha} \right) \geq 1- \alpha \end{align*} Proof: See Appendix (ref).

Lemma (ref) implies that the confidence set covers the true value with the requested probability asymptotically even in the presence of weak (proxy) factors, which is essential for the correct size performance of the HJS test. This result holds under more general cases, for example it holds even when the model is not identified, and this correct coverage probability of the confidence set directly results from the correct size of the identification robust AR test statistic.

theoSuppose the assumptions in Lemma (ref) hold, \begin{align*} \lim\sup_T \mathbb{P}\left(T\widehat{\delta}^*_g \geq c_{1-\alpha}^* \right) \leq \alpha \end{align*} Proof: See Appendix (ref).

Theorem (ref) shows that $c_{1-\alpha}^*$ provides a upper bound for the HJS statistic, which is also a upper bound for the HJ statistic as the HJ statistic is smaller than the HJS statistic by construction, and that the HJS specification test is size correct in the presence of weak (proxy) factors. The proof of it implies that the size property of the HJS specification test is a direct result of Lemma (ref), and given that the lemma holds for more general conditions, we know that the HJS specification test can be extended to more general cases as well. Theorem (ref) also implies that the HJS specification test is conservative, which is understandable as we use the infimum to construct the HJS statistic instead of the supremum. However given the diameter of the robust confidence set can be arbitrarily large, using the supremum can lead to size distortion (see Example (ref)).

Even though it is conservative, the HJS specification test has better power performance compared with another well-know specification test, the $\mathcal{J}$ specification test. The $\mathcal{J}$ specification test statistic is also constructed based on the AR statistic such that

align*[align* omitted — 53 chars of source]

gospodinov2017spurious show that the $\mathcal{J}$ specification test is size correct in the presence of spurious/useless factors, which means $q_G$ is of reduced rank and the model is not identified, but it has a complete power loss in such cases. We extend their results to weakly identified models. Theory (ref) shows that in both unidentified and weakly identified models, the $\mathcal{J}$ specification test suffers from power loss, while our HJS test still maintains proper power performance.

theoSuppose Assumptions (ref), (ref), (ref) hold, but instead of the correct proxy factors $g_t$, proxy factors $\widetilde{g}_t$ are used such that $\widetilde{g}_t$ is a $\widetilde{K}\times 1$ vector, $\left\lVert\iota_N -q_{\widetilde{G}} \theta_{\widetilde{G}}\right\rVert>a > 0, \forall \theta_{\widetilde{G}} \in \Theta $ with $q_{\widetilde{G}} =\mathbb{E}(r_t\widetilde{G}_t'), \widetilde{G}_t=(1,\widetilde{g}_t)$ and the model is misspecified. In addition, assume that the $\psi_{B_{{g}}}$, when we replace ${G}_t$ with $\widetilde{G}_t$, in Lemma (ref) satisfies that $\psi_{B_{{g}}}\sim N(0, Q_{\widetilde{G}} \otimes \Sigma)$ with $Q_{\widetilde{G}} = \mathbb{E}(\widetilde{G}_t\widetilde{G}_t')$ and $\Sigma$ the covariance matrix of $\widetilde{u}_{g,t}$. Let $H=\left(\iota_N, q_{\widetilde{G}} \right)$. \begin{itemize} • (gospodinov2017spurious, Theorem 2, unidentified model under misspecification) Suppose $H$ has a column rank $K+1-k$ for an integer $k\geq 1$, then we have \begin{align*} \mathcal{J} \preceq_d w_{k,i} \end{align*} where $w_{k,i}$ is the smallest eigenvalue of $\mathcal{W}_k\sim \mathcal{W}_k\left(N-K-1+k,I_k \right)$ and $\mathcal{W}_k\left(N-K-1+k,I_k \right)$ denotes the Wishart distribution with $N-K-1+k$ degrees of freedom and a scaling matrix $I_k$. Furthermore, \begin{align*} & \lim\sup_T \mathbb{P}\left( \mathcal{J} \geq c_{\chi_{N-K}^2, 1-\alpha} \right) \leq \alpha, \\ & \lim\inf_T \mathbb{P}\left(T\widehat{\delta}^*_g \geq c_{1-\alpha}^* \right) =1, \end{align*} with $c_{\chi_{N-K}^2, 1-\alpha}$ the $1-\alpha$ quantile of $\chi_{N-K}^2$. • (weakly identified model under misspecification) Suppose $(HQ_{x})'*(HQ_{x})$ with $Q_x=\textit{diag}(I_{K+1-k},\sqrt{T}I_{k})$ converges to a positive definite matrix, then we have \begin{align*} \mathcal{J} \rightarrow_d w_{k,ii} \end{align*} where $w_{k,ii}$ is the smallest eigenvalue of $\mathcal{W}_k\sim \mathcal{W}_k\left(\mu, N-K-1+k,I_k \right)$ and $\mathcal{W}_k\left(\mu,N-K-1+k,I_k \right)$ denotes the non-central Wishart distribution with $N-K-1+k$ degrees of freedom and a scaling matrix $I_k$, a location parameter $\mu$ ($\mu$ is specified in the proof). Furthermore, \begin{align*} & \lim\inf_T \mathbb{P}\left( \mathcal{J} \geq c_{\chi_{N-K}^2, 1-\alpha} \right) <1, \\ & \lim\inf_T \mathbb{P}\left(T\widehat{\delta}^*_g \geq c_{1-\alpha}^* \right) =1, \end{align*} with $c_{\chi_{N-K}^2, 1-\alpha}$ the $1-\alpha$ quantile of $\chi_{N-K}^2$. \end{itemize} Proof: See Appendix (ref).

Simulation exercises

In this section, we conduct a simple simulation exercise with a single-factor model to evaluate the empirical rejection rates (the size and power performance) of our proposed HJS specification test. We calibrate the data generating process in our simulations to match the data set from kroencke2017asset. We simulate the factor $f_t\sim ~i.i.d.~N(0,V_f)$, where we set $V_f$ to match the sample variance of the consumption growth factor. $r_t$ is generated with one factor $f_t$ via $r_t=\iota_N + \beta \lambda+ \beta_\perp d +\beta d_g f_t +u_t$, $u_t\sim ~i.i.d ~N(0,V_u)$, where we match $\lambda$ to the estimated risk premium, $\beta$ is the sample slope parameter between the assets returns and consumption growth factor, and $V_u$ is the sample covariance of the residuals resulting from regressing asset returns on a constant and the consumption growth factor. $\beta_\perp$ is a vector which is orthogonal to $\iota_N, \beta$ and $\left\lVert\sqrt{T}\beta_\perp\right\rVert=1$. We set $T=100$. We use $d_g$ to tune the identification strength of the factors in our simulation exercise where a larger $d_g$ means a stronger factor, and $d$ to tune the model misspecification level where a lager $d$ means a larger deviation from the moment conditions ((ref)) for our simulated data.

For the size performance comparison, we set $d=0$ and thus moment conditions ((ref)) hold for our simulated data. Figure (ref) shows that the HJ specification test is highly size distorted and the distortion only drops down slightly when we increase the identification strength of the factor, while the HJS specification test has a better finite sample behavior and remains size correct.

figure[figure omitted — 280 chars of source]

For the power performance comparison, we set $d_g=0$ which means $f_t$ only serves as a spurious factor. Figure (ref) shows that the rejection frequency of the HJS specification test increases much faster compared with the one of the $\mathcal{J}$ specification test when the level of model misspecification ($d$) increases. The rejection frequency of the $\mathcal{J}$ specification test remains relatively small even when the HJS specification test rejection frequency is close to one, and this implies the HJS specification test has better power performance.

These observations support our theory and show that the HJS specification test has good performance in both size and power.

figure[figure omitted — 413 chars of source]

Specification testing with large N: HJN

In the previous section, we construct the HJS statistic using a robust confidence set of $\theta_G$ since it is only weakly identified with a limited number of asset returns. The HJS specification testing procedure involves optimization steps, which is commonly done in practice through a grid search procedure. In this section, we provide another novel valid specification test statistic, the HJN statistic, which does not involve any time consuming optimization procedure. The construction of the HJN statistic uses a consistent $\theta_G$ estimator, and thus we first introduce our $\theta_G$ estimator and then the HJN statistic.

Four-pass estimator

When we work within a double-asymptotic framework such that both the number of time periods and the number of asset returns grow, weak (proxy) factors do not necessarily lead to a weak identification problem (anatolyev2018factor), which is similar to the case of many weak instruments that information about some parameters though limited, aggregates slowly. Even though $\widehat{\theta}_G$ is not consistent (Theory (ref)), another consistent estimator for $\theta_G$ can be constructed. With an extended number of asset returns, we can estimate $\theta_G$ consistently by removing the missing factor structure via PCA and using an IV-type technique to correct for the remaining issues. The consistent estimator gives another way to construct a statistic for the HJ distance, based on which we propose a novel specification test statistic, our HJN statistic. In the following, we first introduce our four-pass $\theta_G$ estimator with the extended number of asset returns, and thereafter provide the motivation for this four-pass procedure.

We propose the following steps to estimate $\theta_G$ with $N$ base portfolios of gross returns $r_t$:

Step (1): Estimate $\widehat{c}, \widehat{\beta}_g$ in the linear observed-(proxy)-factor model ((ref)) via OLS with $N$ base portfolios of returns.

Step (2): Determine the omitted factor structure using the following two steps:

(2.1) Determine the number of factors, $K_{vz}$, in $\widehat{u}_{g,t}=r_t - \widehat{c}-\widehat{\beta}_g \bar{g}_t$ by

align[align omitted — 183 chars of source]

where $\lambda_j(A)$ is the $j$-th largest eigenvalue of a given matrix $A$, $\widehat{u}_{g}$ is $T\times N$ matrix stacked with the OLS residuals $\widehat{u}_{g,t}$, $K_{vz,\max}$ is an arbitrary upper bound for $K_{vz}$ and $\phi(N,T)$ is a penalty function with the properties $\phi(N,T)\rightarrow 0, \phi(N,T)/(N^{-\frac{1}{2}}+T^{-\frac{1}{2}}) \rightarrow \infty$ (e.g., in later simulation exercise and empirical application we simply choose $\phi(N,T)= N^{-\frac{1}{4}}+T^{-\frac{1}{4}}$);

(2.2): Estimate the $T\times N$ common component matrix ${cc}= {x} {b}'$ stacked with the common components $cc_{t}$, $\widetilde{cc} = \widehat{x} \widehat{b}$, such that $ \widehat{x} $ is equal to $\sqrt{T}$ times the eigenvector associated with the $\widehat{K}_{vz}$ largest eigenvalues of the matrix $\widehat{u}_g\widehat{u}_g' $, and $\widehat{b}= \widehat{x}' \widehat{u}_g/T$ corresponds with the OLS estimator regressing $\widehat{u}_g$ on $\widehat{x}$: $$ \left( \widehat{x}, \widehat{b} \right) = \arg \min_{b_i, x_t ~s.t.~\sum_{t=1}^{T} x_tx_t' /T = I_{\widehat{K}_{vz}} } \sum_{i,t} \left( \widehat{u}_{g,it} - b_{i}' x_{t} \right)^2 $$

Step (3): Split the sample into two non-overlapping subsamples along the time index and remove the missing factor structure from the regressors in the SDF of both subsamples: $$\widetilde{q}_{G,T}^{(i)}= q_{G,T}^{(i)}- \widetilde{cc}_{G,T}^{(i)'} , i=1,2$$ where

align*[align* omitted — 493 chars of source]

Step (4): We then use IV regression to derive two estimators, $\widetilde{\theta}_{G}^{(i)}, i=1,2$, where we use $\widetilde{q}_{G,T}^{(1)}$ as instrument for $\widetilde{q}_{G,T}^{(2)}$ and vice versa. Thereafter, our proposed four-pass estimator is derived by taking the average of both estimators: $$\widetilde{\theta}_G = \sum_{i=1}^2 \widetilde{\theta}_{G}^{(i)}/2, $$ with $ \widetilde{\theta}_{G}^{(1)} = \left(\widetilde{q}^{(1)' }_{G,T} P_{\widetilde{q}^{(2)}_{G,T}} \widetilde{q}^{(1)}_{G,T} \right)^{-1} \widetilde{q}^{(1)'}_{G,T} P_{ \widetilde{q}^{(2)}_{G,T} } \iota_N $ and $ \widetilde{\theta}_{G}^{(2)} = \left(\widetilde{q}^{(2)' }_{G,T} P_{\widetilde{q}^{(1)}_{G,T}} \widetilde{q}^{(2)}_{G,T} \right)^{-1} \widetilde{q}^{(2)'}_{G,T} P_{ \widetilde{q}^{(1)}_{G,T} } \iota_N$.\\

Our estimation approach for $\theta_G$ resolves the problems of the missing factor structure and the weak (proxy) factors simultaneously. We make use of the results from bai2002determining, bai2003inferential and giglio2017inference in step (2) to recover the common components in the error terms using principal component analysis, and we use the instrumental variable idea applied for the factor models, which is used in anatolyev2018factor, in step (4) to solve potential endogeneity issues. Compared with the estimator proposed in anatolyev2018factor, our proposed estimator relaxes the restrictions on the number of omitted factors and the restrictions on the rank of the loadings of all the factors present in the model.

To illustrate why our proposed procedure is robust against weak (proxy) factors and a missing factor structure, we start by comparing it with the conventional $\theta_G$ estimator. To do so, we first rewrite equation ((ref)) ($\iota_N = \mathbb{E}q_{G,T} \theta_G$) as

align*[align* omitted — 79 chars of source]

with $ \epsilon_{q_G} = \left({q}_{G,T} - \mathbb{E}{q}_{G,T} \right)$ which is correlated with $ q_{G,T}$. The term $\epsilon_{q_G}$ vanishes asymptotically, and so it is dominated by $q_{G,T}$ when all proxy factors are strong. The conventional estimator, which results from regressing $\iota_N$ on $ q_{G,T}$, is then valid in large samples, since $\epsilon_{q_G}$ becomes negligible. However, if some (proxy) factors are weak, some columns of $q_{G,t}$ are of the same order as $\epsilon_{q_G}$, then there would be a classic endogeneity problem if we simply regress $\iota_N$ on ${q}_{G,T} $.

To solve the endogeneity problem, a valid instrument can be constructed in our framework with a split-sample technique and this idea is also employed in anatolyev2018factor. Given the independence of the $e_t$ from non-overlapping sub-samples, $q_{G,T}^{(1)}$ can serve as an instrument for $q_{G,T}^{(2)}$ and vice versa when there is no missing factor structure ($K_{vz}=0$) and this is the starting point of our proposed procedure. When there is a missing factor structure with factors that might be correlated across time, $q_{G,T}^{(1)}$ is no longer a valid instrument for $q_{G,T}^{(2)}$. Therefore, we use $\widetilde{q}_{G,T}^{(i)}, i=1,2$ which results from removing the missing factor structure from $q_{G,T}^{(i)},i=1,2$. By doing so, $\widetilde{q}_{G,T}^{(1)}$ is asymptotically uncorrelated with $\widetilde{\epsilon}_{q_G}^{(2)}$, and is a valid instrument.

As shown in Theorem (ref), our estimation procedure provides $\min \{ \sqrt{T}, \sqrt{N} \}$-consistent results for $\theta_G$, of which a non-linear transformation leads to a consistent risk premia estimator (Corollary (ref)).

theoSuppose Assumptions (ref) - (ref), (ref) - (ref) hold, and $N/T \rightarrow c$. \begin{align*} \sqrt{NT} Q^{-1}_{B_g,T}\left(\widetilde{\theta}_{G} - \theta_G\right) \rightarrow O_p(1), \end{align*} with $Q_{B_g,T} = \text{diag}(I_{1+K_{g,1}},\sqrt{T}I_{1+K_{g,1}})$.\footnote{With some additional regular assumptions, we can construct $\widehat{\Sigma}_{\theta_G}$ such that $ \widehat{\Sigma}_{\theta_G}^{1/2}\left(\widetilde{\theta}_{g} - \theta_G\right) \rightarrow_d N(0,I)$. See Appendix (ref). } \\ Proof: See Appendix (ref) .
coroSuppose the assumptions in Theorem (ref) hold, then \begin{align*} \sqrt{NT} Q^{-1}_x\left(\widetilde{\lambda}_{g} - \lambda_g\right) \rightarrow O_p(1), \end{align*} with $\widetilde{\lambda}_{g}=-V_g \text{diag}(0_{K\times 1},I_{K})\widetilde{\theta }_{G}/\widetilde{\theta }_{G,1}$.\\ Proof: This is a direct result of Theorem (ref).

HJN specification test

kleibergen2018identification study risk premia on mimicking portfolios by projecting non-traded factors on traded base portfolios, and then carry out identification robust tests using a set of testing portfolios. We use a similar idea to construct the HJN specification test: \\ Step (1) Estimate $\theta_G$ from a set of N base portfolios $r_t$ of asset returns using our proposed four-pass estimator $\widetilde{\theta}_G$

Step (2) Estimate $q_G$ from a set of testing portfolios $R_t$ of n asset returns and $n$ is fixed such that $\widetilde{q}_G= \frac{1}{T}\sum_{t=1}^{T} R_tG_t'$

Step (3) The HJN statistic is $$ \widetilde{\delta}_{g}^{2} = \widetilde{e}_T' \widehat{Q}_R^{-1} \widetilde{e}_T,$$ with sample pricing errors $\widetilde{e}_T= \iota_n - \widetilde{q}_G\widetilde{\theta}_G$, $\widehat{Q}_R= \frac{1}{T}\sum_{t=1}^{T} R_tR_t'$.\\

Step (4) This test would reject the null hypothesis that moment conditions ((ref)) hold if $$T\widetilde{\delta}_{g}> \widetilde{c}_{1-\alpha},$$ where $\widetilde{c}_{1-\alpha}$ is the $1-\alpha$ quantile of the weighted sum of $N$ $\chi^2(1)$ random variables with weights being the positive eigenvalues of the matrix $\widetilde{S}^{\frac{1}{2}} (\widehat{Q}_R)^{-1}\widetilde{S}^{\frac{1}{2}'}$ with $\widetilde{S}$ a consistent estimator of the long-run variance matrix of the sample pricing errors ${e}_{T,R}(\theta_G)=\iota_n - \widetilde{q}_G {\theta}_G$. \\

Remark: the HJN specification test does not require the base portfolios and testing portfolios to be non-overlapping.

coroSuppose the assumptions in Theorem (ref) hold, then \begin{align*} T \widetilde{\delta}_{g}^{2} \sim_d \sum_{i=1}^{n} {p}_i x_i \end{align*} with $x_i$ being independently $\chi^2(1)$ distributed random variables and $p_i$ being the positive eigenvalues of the matrix $\widetilde{S}^{\frac{1}{2}} (\widehat{Q}_R)^{-1}\widetilde{S}^{\frac{1}{2}'}$ with $\widetilde{S}$ a consistent estimator of the long-run variance matrix of the sample pricing errors ${e}_{T,R}(\theta_G)$ (for example, one may simple choose $\widetilde{S}= \frac{1}{T}\sum_{t=1}^T e_{g,t,R} (\widetilde{\theta}_G)e_{g,t,R} (\widetilde{\theta}_G)', e_{g,t,R}({\theta}_g)= \iota_N -R_tG_t'{\theta}_g$). \\ Proof: See Appendix (ref).

Theorem (ref) shows that our HJN specification test is size correct even with weak (proxy) factors.

theoSuppose the assumptions in Theorem (ref) hold, \begin{align*} \lim\sup_T \mathbb{P}\left(T\widetilde{\delta}_g \geq \widetilde{c}_{1-\alpha} \right) =\alpha \end{align*} with $\widetilde{c}_{1-\alpha}$ the $1-\alpha$ quantile of the weighted sum of $N$ $\chi^2(1)$ random variables with being the positive eigenvalues of the matrix $\widetilde{S}^{\frac{1}{2}} (\widehat{Q}_R)^{-1}\widetilde{S}^{\frac{1}{2}'}$ with $\widetilde{S}$ a consistent estimator of the long-run variance matrix of the sample pricing errors ${e}_{T,R}(\theta_G)$ (for example, one may simple choose $\widetilde{S}= \frac{1}{T}\sum_{t=1}^T e_{g,t,R} (\widetilde{\theta}_G)e_{g,t,R} (\widetilde{\theta}_G)', e_{g,t,R}({\theta}_g)= \iota_N -R_tG_t'{\theta}_g$).\\ Proof: This is a direct result from Corollary (ref).

Simulation exercise and empirical application

Similar to section (ref), we again evaluate the empirical rejection rates of our HJS specification test via simulation exercises.

We calibrate to the data set used in anatolyev2018factor: the monthly returns on 100 Fama-French portfolios sorted by size and book-to-market and three Fama-French factors ($g_t$). From the portfolio returns we obtain the first four principal components (PC), and we regard the first three PCs as priced latent common factors, $f_t$, and the fourth one as the omitted factor, $z_t$. With normalization $\left(f,z\right)'\left(f,z\right)/T =I_{4}$, we set the variance of these factors to be $1$. We regress demeaned returns on $f_t, z_t$ for their risk exposures, and calculate the sample mean $\mu_{\beta\gamma}$ and the sample variance $V_{\beta\gamma}$ of the risk exposures. We compute the sample variance $\sigma_e^2 I_N$ of the residuals after regressing returns on $f_t,z_t$. To maintain the relation between observed factors $g_t$ and PCs $f_t$, we regress $f_t$ on the three Fama French factors to obtain the slope $d_g$ and residual covariance matrix $V_v$, and $d_g$ captures the quality of the proxy factors.

We then simulate our data in the following way. In the first step we simulate observed factors from i.i.d $N(0,I_3)$ and latent factors $f_t$ are generated by $Ad_g g_t + v_t$ with $v_t$ simulated as i.i.d $N\left(0, (I-A)d_gd_g'(I-A)+V_v \right)$. $A$ is a diagonal matrix which we use to adjust the strengths of our (proxy) factors, and we set $A=\text{diag}(I_2,d_{\alpha})$ in our simulations with $d_{\alpha}$ tuning the strength of the simulated factors. As for the corresponding risk exposures, we use $(\beta_i', \gamma_i')'\sim$ i.i.d $N(\mu_{\beta\gamma}, V_{\beta\gamma})$. Then in the end, we generate $r_t=\iota_N+\beta_\perp d+\beta \lambda+\beta f_t +\gamma z_t + e_t$, $e_t\sim ~i.i.d ~N(0,\sigma_e^2 I_N)$, where we match $\lambda$ to the estimated risk premia resulting from the data. $\beta_\perp$ is a vector which is orthogonal to $\iota_N, \beta$ and $\left\lVert\beta_\perp\right\rVert=1$. Similar to the previous simulation setting in Figure (ref) we use $d$ to tune the model misspecification level and $d=0$ when we simulate size curves.

In our simulations, we fix $N=100$. For the HJN specification test we use all the simulate 100 asset gross returns to form the base portfolios and the first 25 to form the testing portfolios, and we use the testing portfolios for the conventional HJ specification test.

Figure (ref) compares the size curves of the HJ specification test and of the HJN specification test. It shows that the HJ specification test is highly size distorted even when $d_{\alpha}$ is large (left hand side panel of Figure (ref)) and the size distortion increases when proxy factors become weaker (smaller $d_{\alpha}$), while the HJN specification test roughly remains size correct. Even with relatively large number of time periods, the conventional HJ specification test still over-reject. Observations from Figure (ref) also seem to imply that our HJN specification test tends to under-reject in finite samples, and to show it performs well when the model is misspecified we also simulate power curves in Figure (ref).

figure[figure omitted — 396 chars of source]

Figure (ref) shows power curves of the HJ specification test and the HJN specification test respectively. The left hand side panel of Figure (ref) uses $d_{\alpha}=0.5$ to mimic one weak proxy factor, while all proxy factors in the right hand side panel are strong with a larger value of $d_{\alpha}$. Figure (ref) shows that the HJN specification test has proper power performance regardless of the presence of weak (proxy) factors, and rejection frequency increases faster when the proxy factor is stronger.

figure[figure omitted — 514 chars of source]

Empirical application

We apply our proposed test procedures on the data set of monthly returns on 100 Fama-French portfolios sorted by size and book-to-market and the three Fama-French factors (market, SmB, HmL) and the momentum factor.

An intuitive measure for the factor structure in asset returns is total variation of the asset returns explained by the principal components\footnote{This corresponds with the nuclear norm of the demeaned asset returns} (kleibergen2015unexplained). We construct the spectral decomposition of the sample covariance matrix of the 100 portfolio returns, and denote $\lambda_1 > \lambda_2 > \cdots $ the characteristic roots (or eigenvalues of the PCs of asset returns) in descending order. We use the characteristic roots ratios (CRRs) ${\lambda_i}/{\left(\sum_{j}\lambda_j\right)}, i=1,2,3,4$, which represent the total variation of the portfolio returns explained by the first four PCs respectively, to check the factor structure of portfolio returns (see Figures (ref) and (ref)).

Figures (ref) and (ref) also report the p-values of specification tests (HJ and HJN) with respect to a three-FF-factor model and a four-factor (adding the momentum) model from 1963-09 to 2019-08 using rolling windows of 240 and 120 months respectively\footnote{We first choose the window size of 240 months in Figure (ref), because our simulations suggest sample size around 300 seems to be enough for carrying out our tests properly.}. For the HJN specification test we use all the 100 asset returns to form the base portfolios and the first 25 to form the testing portfolios, and we use the testing portfolios for the conventional HJ specification test. They also report measures for the presence of a factor structure in the asset returns: the fraction of the total variation of the portfolio returns that is explained by their principal components.

Figures (ref) shows that when comparing nested models, the HJ test can produce counter-intuitive results by rejecting a four-factor model but not the reduced three-factor model (see points near the coordinate '2015-01'). This is an unfortunate outcome since the four-factor mode apparently embeds the three-factor model and if the four-factor model is rejected we would expect the three-factor model to be rejected as well. We attribute this strange behavior to the momentum factor having only weak correlation with the returns and thus inducing a larger rejection rate of the HJ test, while our HJN specification test does not have such problem.

The HJN specification test also captures changes in the factor structure of asset returns in a more sensible way compared with the HJ test. As shown in Figure (ref), when CRRs vary in different time periods (e.g., the total variation of the portfolio returns is mostly explained by the first PC for points near the coordinate '2000-01' while the other PCs only account for a much lower percentage of the variation), the HJN specification tests reflect the changes in the factor structure of asset returns with variations in p-values of tests of a four-factor model, while the HJ specification tests reject both three-factor and four-factor models for most time periods and is not informative for the factor structure of asset returns.

Both of the HJ and the HJN tests in Figure (ref) seem to have larger p values near the coordinate '2015-01' while the patterns of characteristic roots ((a) and (b)) seem to be rather stable. This is because a 240-month window size is a bit too long, and some changes in the factor structure might be averaged out and thus not detected by CRRs. We choose a smaller rolling window size (120 months) in Figure (ref), and it shows the change in the factor structure (Figure (ref).(a)) after the coordinate '2010-01'. Similar to what we observe in Figure (ref), Figure (ref) also shows our HJN specification test respond to the factor structure in the asset returns in a more informative way, while the HJ tests only report small p values for most time periods for both three- and four-factor models.

sidewaysfigure[h!] \caption{ The time series of estimates from 1963-09 to 2019-08 with rolling windows of size T = 240 (20 years) and the number of increments between successive rolling windows being 12 (1 year) (x-axis is labeled with the ending period of each rolling window, for example, the first rolling window uses data from 1963-09 to 1983-08, so x-axis should label the first point with '1983-08'). (a) The fraction of the total variation of the portfolio returns that is explained by their four largest principal components respectively (CRRs) ($ \lambda_i/(\sum_{j=1}^N \lambda_j), i=1,\cdots,4$); (b) The fraction of the total variation of the portfolio returns that is explained by the sum of the four largest principal components ($ \sum_{i=1}^4\lambda_i/(\sum_{j=1}^N \lambda_j)$); (c) p-values of the HJ specification test of the three-FF-factor model (blue), p-values of the HJ specification test of the four-factor (three FF and momentum factors) model (red);(d) p-values of the HJN specification test of the three-FF-factor model (blue), p-values of the HJN specification test of the four-factor model (red).}
sidewaysfigure[h!] \caption{The time series of estimates from 1968-09 to 2019-08 with rolling windows of size T = 240 (20 years) and the number of increments between successive rolling windows being 12 (1 year) (x-axis is labeled with the ending period of each rolling window, for example, the first rolling window uses data from 1968-09 to 1988-08, so x-axis should label the first point with '1988-08'). (a) The fraction of the total variation of the portfolio returns that is explained by their four largest principal components respectively (CRRs) ($ \lambda_i/(\sum_{j=1}^N \lambda_j), i=1,\cdots,4$); (b) (GDP growth) US GDP annual growth rate 1988-2018 (downloaded from the World Bank); (c) p-values of the HJ specification test of the three-FF-factor model (blue), p-values of the HJ specification test of the four-factor (three FF and momentum factors) model (red);(d) p-values of the HJN specification test of the three-FF-factor model (blue), p-values of the HJN specification test of the four-factor model (red).}
sidewaysfigure[h!] \caption{ The time series of estimates from 1963-09 to 2019-08 with rolling windows of size T = 240 (20 years) and the number of increments between successive rolling windows being 12 (1 year) (x-axis is labeled with the ending period of each rolling window). (a) The fraction of the total variation of the portfolio returns that is explained by their four largest principal components respectively (CRRs) ($ \lambda_i/(\sum_{j=1}^N \lambda_j), i=1,\cdots,4$); (b) The fraction of the total variation of the portfolio returns that is explained by the sum of the four largest principal components ($ \sum_{i=1}^4\lambda_i/(\sum_{j=1}^N \lambda_j)$); (c) p-values of the HJ specification test of the three-FF-factor model (blue), p-values of the HJ specification test of the four-factor (three FF and momentum factors) model (red);(d) p-values of the HJN specification test of the three-FF-factor model (blue), p-values of the HJN specification test of the four-factor model (red). }

We observe in Figures (ref) and (ref) that the HJN specification tests in some rolling windows do not reject a four-factor model. To further study this observation, Table (ref) reports results based on the data from 1977-08 to 2019-08. We see in Table (ref) that both the HJ and HJN specification tests reject the three-factor model, and while the HJ specification test rejects the four-factor model, the HJN specification test does not reject it. Our HJS specification test seems to be a bit conservative and does not reject both models in this application. The estimates for the four-factor model using our proposed approach indicate a larger change in values corresponding to the momentum factor, and this might result from the momentum factor being weak. Our specification tests support a four-factor model for Fama French portfolios, and observations show that the momentum factor might only serve as a weak proxy factor which can explain the difference between the HJ and the HJN specification test results and the differences in estimated parameter values. \\

table[table omitted — 541 chars of source]

Conclusions

We show that the HJ statistic is not a valid model selection tool and model specification test statistic when weak (proxy) factors are present. We propose two novel approaches that provide size-correct model specification tests, alongside with which we also propose novel weak (proxy) factors robust risk premia estimators. Our empirical application supports a four factor structure for Fama French portfolios despite that the momentum factor is a weak proxy factor.