EconBase
← Back to paper

One Factor to Bind the Cross-Section of Returns

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

101,311 characters · 19 sections · 49 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

One Factor to Bind the Cross-Section of Returns

abstractWe propose a new non-linear single-factor asset pricing model $r_{it}=h(f_{t}\lambda_{i})+\epsilon_{it}$. Despite its parsimony, this model represents exactly any non-linear model with an arbitrary number of factors and loadings -- a consequence of the Kolmogorov-Arnold representation theorem. It features only one pricing component $h(f_{t}\lambda_{i})$, comprising a nonparametric link function of the time-dependent factor and factor loading that we jointly estimate with sieve-based estimators. Using 171 assets across major classes, our model delivers superior cross-sectional performance with a low-dimensional approximation of the link function. Most known finance and macro factors become insignificant controlling for our single-factor.

\

JEL Codes: G10, G12, C10

Keywords: asset returns, non-linear factor model, Kolmogorov-Arnold, factor zoo

\thispagestyle{empty}

\setcounter{page}{1}

Introduction

Factor models play an important role in finance and economics. The existing literature, however, focuses primarily on linear factor models. In this paper, we propose a new non-linear single factor asset pricing model

equation[equation omitted — 109 chars of source]

where $r_{it}\in\mathbb{R}$ is the return of asset $i$ at time period $t$, $f_{t}>0$ is the single factor, $\lambda_{i}>0$ is the factor loading, $h$ is a continuous function, and $\varepsilon_{it}\in\mathbb{R}$ is an idiosyncratic mean-zero component. Here, the function $h$, the factor values $f_{t}$, and the factor loading values $\lambda_{i}$ are all assumed to be unknown, and both $N$ and $T$ are assumed to be large. We refer to our model as the HFL model.

Our model is different from the typical linear factor model. On the one hand, our model is seemingly more restrictive because it allows for only one time-dependent factor and requires both the time factor and the factor loading to be strictly positive, whereas the linear factor model allows for an arbitrary number of factors and allows all factors and factor loadings to take negative values. On the other hand, our model is more flexible because it includes a nonparametric link function $h$, whereas the linear factor model sets the link function to be identity.

Even though our model has only one factor, the presence of the nonparametric link function $h$ makes our model surprisingly flexible. In particular, using the Kolmogorov-Arnold representation theorem (Kolmogorov1956on,Arnold1957on), we show that any nonparametric factor model with an arbitrary number of factors and with arbitrary interactions between factors and factor loadings can be reduced to model ((ref)). That is, any arbitrary nonparametric model can be exactly represented by our model. It is thus without loss of generality to focus on our model with just one factor.

We apply our model to characterize risk premia in a large cross-section of assets across major asset classes, and propose a sieve-based least squares estimator of the model. This estimator is defined as a solution to an optimization problem minimizing the sum of squared residuals for the model not only over $T$ potential factor values $\{f_{t}\}_{t=1}^{T}$ and $N$ potential factor loading values $\{\lambda_{i}\}_{i=1}^{N}$ but also over potential values of the function $h$ in a sieve space, which we choose to consist of all polynomials of a given order that is slowly increasing in the sample dimensions $(N,T)$.

The Kolmogorov-Arnold representation ensures that the function $h$ must be continuous under very mild regularity conditions, and so can be arbitrarily well approximated by a suitable sieve space for estimation purposes. At the same time, readers familiar with the Kolmogorov-Arnold representation theorem might argue that this theorem implies that even though the function $h$ must be continuous, it may nonetheless be rather irregular, and one may need a high-dimensional sieve space in order to approximate it sufficiently well, making estimation of this model difficult. However, whether the function $h$ can be well approximated by a relatively low-dimensional sieve space for the purposes of a particular application, making estimation feasible, is ultimately an empirical question. Our main empirical result is that, for the cross-section of asset returns across major asset classes, the estimated version of our model based on a relatively low-dimensional sieve space outperforms linear factor models with multiple factors in a number of dimensions.

We test the cross-sectional asset pricing predictions of our model using as test assets U.S. equity portfolios, U.S. and international government bonds, commodities, and foreign currency portfolios, for a total of 171 assets. Specifically, we estimate a cross-sectional regression of the average asset realized excess returns on the average asset excess returns predicted by the model and test several implications of our one factor model, such as that the intercept is zero and the significance of the slope coefficient. Furthermore, we evaluate the fit of the model. Our empirical strategy follows the classic two-step fama1973risk's procedure. The first step of the fama1973risk's procedure, which requires the estimation of the asset specific betas in time-series regressions, is implicit in the algorithm to estimate our single factor model. The second step of the procedure, which requires a cross-sectional regression of average asset excess returns on factor betas, corresponds to our cross-sectional regression.

Our single factor model explains a large fraction of cross-sectional differences in asset returns. We find that a relatively low-dimensional sieve space, with the polynomials of degree as low as four, suffices for the majority of our empirical results. When using all asset classes as test assets, the slope coefficient on the pricing component is significantly different from zero. Furthermore, we cannot reject the null that the slope coefficient is equal to 1, as implied by our model. Moreover, the pricing error, measured by the intercept in the cross-sectional regression, is not significantly different from zero. The HFL model exhibits a cross-sectional regression adjusted R-squared value of 89%.

Next, we compare the HFL model with a number of prominent cross-sectional asset pricing models that have been proposed in the finance and economics literature. First, we consider benchmark equity factor models, such as the CAPM, the fama1993common's three factor model, the fama2015five's five factor model and the five factor model augmented with the momentum factor of jegadeesh1993returns. Second, we consider models using factors based on principal component analysis (PCA). Specifically, we consider three estimators of latent factors based on principal components: the standard estimator based on the covariance matrix of asset returns; the RP-PCA estimator, which is constructed to better account for the cross-section of average asset returns; and the kernel PCA estimator, which is a nonlinear form of PCA. Third, we consider higher-order factor models, which account for non-linearities by including higher order factors, such as the squared equity market return of harvey2000conditional. Specifically, we consider models including the square and cube of equity and PCA factors. Fourth, we consider recent macro-factor models, such as the intermediary capital risk factor model of \citet*{he2017intermediary}, the downside equity market risk factor model of \citet*{lettau2014conditional}, and the liquidity factor model of pastor2003liquidity. Fifth, we consider the macro factors of ludvigson2009macro. Our main result is that, once we include our single factor in cross-sectional regressions, the risk prices associated with the alternative factors are all indistinguishable from zero, with the exception of the size factor in fama1993common.

Furthermore, we compare the HFL model with a large set of factors proposed in the last decades to account for asset risk premia, which is commonly referred to as the factor zoo (see, e.g., \citet*{feng2020taming}). We consider 153 established factors, collected by \citet*{jensen2023there}, and for each of these factors, we first estimate a standard cross-sectional regression of average asset excess returns on the asset factor betas. The betas are measured as slope coefficients in time-series regressions of each asset excess return on the factor. We find that, for a large number of these factors, the risk price, i.e., the slope coefficient in the cross-sectional regression, is statistically significant, although also the estimates for the intercept, a measure of the pricing error of the model, are mostly significantly different from zero. We then repeat the estimation additionally including the predicted values from the HFL model. The results are striking. We find a significant factor risk price estimate for only 3 out of the 153 factors in the factor zoo, while the slope estimates associated with the HFL factor are all statistically significant. Moreover, the pricing error of these models are mostly indistinguishable from zero. Furthermore, we conduct cross-sectional regressions employing both the HFL factor and a diverse set of factors from the factor zoo, employing the double-selection Lasso method. This approach enables simultaneous control over all factors and provides a robust model selection technique for assessing factor contributions in cross-sectional asset pricing regressions (\citet*{feng2020taming}). We find that the HFL component is highly significant, after controlling for the factor zoo, while the majority of the factors in the factor zoo is not.

We then investigate the predictive ability of the HFL model by building portfolios using as signals the predicted asset returns. Specifically, we first estimate the model on all assets and the first 120 months of data. Next, we use the model predicted values to sort assets in five portfolios, and compute the equally-weighted average portfolio returns in the following six months. Finally, we repeat the procedure expanding the estimation window by 6 months each time until the end of the sample. This empirical methodology is designed to capture the strategy of an investor who uses the HFL model to predict future returns rebalancing at a semi-annual frequency. In each period $t$, we sort assets only using information available up to period $t$ and compute excess returns between $t$ and $t+1,$where each period contains 6 months. The first portfolio groups the assets with, each period, the lowest predicted excess returns by the HFL model. The last portfolio groups the assets with, each period, the highest predicted excess returns by the HFL model. Across the portfolios, we document a sizable, and statistically significant, cross-section of monthly average returns, ranging from 0.08% for the first portfolio to 0.77% for the last one. A zero-cost long-short strategy, wherein assets with the highest predicted returns are bought while those with the lowest predicted returns are sold short, yields a monthly average excess return of 0.69%, with a monthly Sharpe ratio of 0.15. We investigate the risk-adjusted performance of the HFL strategy, by regressing the portfolio returns on the \citet*{fama1993common}'s three factors, on the \citet*{fama2015five}'s five factors, and on the five factors augmented by the momentum factor of \citet*{jegadeesh1993returns}, respectively. The estimates of the intercept from these regressions, a measure of risk-adjusted performance, increase from the first to the last portfolio in all three specifications, and are significant and quantitatively large for portfolio 3 to 5, as well as for the long-short portfolio.

Our empirical results extend to alternative samples. First, we show that the HFL component is statistically significant for each individual asset class using the HFL component estimated with all asset classes. Second, we repeat the cross-sectional asset pricing tests using the first and second half of the sample and find results consistent with those obtained using the full sample. Third, we consider a higher frequency time-variation in the asset pricing test using an estimation based on an expanding window. In this case, we find that the slope coefficient associated with our single factor is significantly different from zero for all months, and we can never reject the null hypothesis that the slope coefficient equals to one as implied by our model. Fourth, we augment the set of test assets with 10 equity momentum portfolios and the momentum factor and report that the slope coefficient associated with the HFL component is statistically different from zero.

Our paper relates to a vast literature on factor models in asset pricing. The majority of the literature focuses on linear factor models, such as characteristic-based models (e.g., \citet*{fama1993common,fama1996multifactor,Carhart1997,koijen2017cross}), statistical-based models (e.g., connor1986performance), and macro models (e.g., \citet*{adrian2014financial,he2017intermediary}). There is a literature that studies conditional asset pricing models, where the model is conditionally linear (e.g., \citet*{jagannathan1996conditional,lettau2001resurrecting,kelly2019characteristics}). There is a literature that shows that the same characteristics-based models are present across a large set of asset classes (e.g., \citet*{asness2013value,koijen2018carry,bollerslev2018risk}). A small subset of papers emphasizes non-linearity in explaining asset returns, such as bansal1993no. A growing literature aims at pricing assets with machine learning methods (e.g., \citet*{hutchinson1994nonparametric,gu2020empirical,kelly2023principal,kelly2024virtue}). Apart for using factor models in understanding asset prices, a recent literature highlights the significance of demand forces (e.g., \citet*{koijen2019demand,koijen2020investors}). We contribute to this literature by proposing and testing a parsimonious model that represents exactly any non-linear model with an arbitrary number of factors and factor loadings.

The Kolmogorov-Arnold representation is typically used as a foundation of deep learning and as a benchmark for deriving bounds on neural network approximations; see SH21, hecht1987kolmogorov, K91, and MP99. C04 also used this representation for nonparametric estimation in the regression context. We contribute to this literature by applying this representation to factor models.

Finally, our approach complements the literature on non-linear factor extraction. \citet*{gu2021autoencoder} used an auto-encoder approach. \citet*{SSM98} developed a kernel PCA procedure, which estimates a linear factor model for a non-linear transformation of the vectors $(r_{1t},\dots,r_{Nt})$. In contrast to these alternative methods, our approach is based on the assumption that the factor affects returns via model ((ref)).

The rest of the paper is organized as follows. Section (ref) discusses the motivation and estimation for our one factor model. Section (ref) derives the rate of convergence of our sieve-based least squares estimator. Section (ref) discusses the data and presents summary statistics. Section (ref) presents the cross-sectional asset pricing results and compares the HFL model with alternative models. Section (ref) contains additional results and robustness checks. The Online Appendix contains technical derivations and further empirical results.

Motivation and Estimation

We start our analysis with a very general factor-type model

equation[equation omitted — 140 chars of source]

where $r_{it}$ is the excess return of asset $i$ at time period $t$, $x_{t1},\dots,x_{tk}\in[0,1]$ are time-specific effects/factors, $z_{i1},\dots,z_{im}\in[0,1]$ are asset-specific effects/factor loadings, $g$ is a continuous function, and $\varepsilon_{it}$ is an idiosyncratic mean-zero component. This model is very general as it nests many special cases. For example, it nests any linear factor model \[ r_{it}=\sum_{j=1}^{k}x_{tj}z_{ij}+\varepsilon_{it}, \] any single-index factor model \[ r_{it}=h\left(\sum_{j=1}^{k}x_{tj}z_{ij}\right)+\varepsilon_{it}, \] and any non-linear factor model \[ r_{it}=h(x_{t1}z_{i1},\dots,x_{tk}z_{ik})+\varepsilon_{it}. \] In fact, model ((ref)) does not even require the number of time-specific effects $k$ to coincide with the number of asset-specific effects $m$.

Our first result shows that under very mild regularity conditions, namely continuity of the function $g$, model ((ref)) can be reduced to model ((ref)) with continuous $h$ as a consequence of the Kolmogorov-Arnold theorem (Kolmogorov1956on,Arnold1957on). Specifically, consider the version of Kolmogorov-Arnold theorem in SH21, Theorem 2(ii). That representation implies that there exists a continuous function $q$ and monotone functions $\phi_{1},\dots,\phi_{k},\psi_{1},\dots,\psi_{m}$ such that \[ g(x_{1},\dots,x_{k},z_{1},\dots,z_{m})=q\left(\sum_{j=1}^{k}\phi_{j}(x_{j})+\sum_{j=1}^{m}\psi_{j}(z_{j})\right) \] for all $x_{1},\dots,x_{k},z_{1},\dots,z_{m}\in[0,1]$. Therefore, denoting $\widetilde{f}_{t}=\sum_{j=1}^{k}\phi_{j}(x_{tj})$ and $\widetilde{\lambda}_{i}=\sum_{j=1}^{m}\psi_{j}(z_{ij})$, it follows that \[ g(x_{t1},\dots,x_{tk},z_{i1},\dots,z_{im})=q(\widetilde{f}_{t}+\widetilde{\lambda}_{i}) \] for all $i=1,\dots,N$ and $t=1,\dots,T$. Moreover, by using the identity \[ x+z=\log(\exp(x+z))=\log(\exp(x)\exp(z)),\quad x,z\in\mathbb{R} \] and denoting $h(x)=q(\log(x))$, $f_{t}=\exp(\widetilde{f}_{t})$, and $\lambda_{i}=\exp(\widetilde{\lambda}_{i})$, we have \[ g(x_{t1},\dots,x_{tk},z_{i1},\dots,z_{im})=q(\log(f_{t}\lambda_{i}))=h(f_{t}\lambda_{i}) \] for all $t=1,\dots,T$ and $i=1,\dots,N$. We summarize this result in the following lemma.

lemmaIf random variables $\{y_{it}\}_{i,t}^{N,T}$ satisfy model ((ref)) with continuous $g$, there exist a continuous function $h$, and strictly positive sequences $\{f_{t}\}_{t=1}^{T}$ and $\{\lambda_{i}\}_{i=1}^{N}$ such that \[ r_{it}=h(f_{t}\lambda_{i})+\varepsilon_{it},\quad i=1,\dots,N,\ t=1,\dots,T. \]

This lemma shows that any factor-type model ((ref)) can be reduced to model ((ref)). It is because of this generality, we study the non-linear factor model ((ref)) with just one factor in this paper.

Remark (Alternative Representations). We could consider a more flexible reduction of model ((ref)). Indeed, by applying the Kolmogorov-Arnold representation theorem conditional on $z_{1},\dots,z_{m}$, we can find a continuous function $q$ and monotone functions $\phi_{1},\dots,\phi_{k}$ such that \[ g(x_{1},\dots,x_{k},z_{1},\dots,z_{m})=q\left(\sum_{j=1}^{k}\phi_{j}(x_{j}),z_{1},\dots,z_{m}\right) \] for all $x_{1},\dots,x_{k},z_{1},\dots,z_{m}\in[0,1]$. Further, by applying the Kolmogorov-Arnold representation theorem conditional on $x_{1},\dots,x_{k}$, we can find a continuous function $u$ and monotone functions $\psi_{1},\dots,\psi_{k}$ such that \[ q\left(\sum_{j=1}^{k}\phi_{j}(x_{j}),z_{1},\dots,z_{m}\right)=u\left(\sum_{j=1}^{k}\phi_{j}(x_{j}),\sum_{j=1}^{m}\psi_{j}(z_{j})\right) \] for all $x_{1},\dots,x_{k},z_{1},\dots,z_{m}\in[0,1]$. Thus, denoting $f_{t}=\sum_{j=1}^{k}\phi_{j}(x_{tj})$ and $\lambda_{i}=\sum_{j=1}^{m}\psi_{j}(z_{ij})$, it follows that model ((ref)) implies that

equation[equation omitted — 124 chars of source]

\qed

In order to estimate model ((ref)), we assume that both the factor $f_{t}$ and the factor loading $\lambda_{i}$ are bounded random variables. By rescaling the function $h$, it is then without loss of generality to assume that both $f_{t}$ and $\lambda_{i}$ are taking values in the $(0,1]$ interval. Further, since the function $h$ is continuous, it follows from the Weierstrass theorem that it can be approximated arbitrarily well by a finite-degree polynomial: \[ h(x)=\sum_{j=0}^{K}h_{j}x^{j}+e_{K}(x),\quad x>0, \] where the residual $e_{K}$ converges to zero as the degree of the polynomial $K$ gets large. We therefore propose a sieve-based least squares estimator

equation[equation omitted — 245 chars of source]

where the minimum is taken over all $\{c_{j}\}_{j=0}^{K}\in\mathbb{R}^{K+1}$, $\{\phi_{t}\}_{t=1}^{T}\in(0,1]^{T}$, and $\{l_{i}\}_{i=1}^{N}\in(0,1]^{N}$. The estimators of the factor values $\{f_{t}\}_{t=1}^{T}$ and factor loadings $\{\lambda_{i}\}_{i=1}^{N}$ are then $\{\widehat{f}_{t}\}_{t=1}^{T}$ and $\{\widehat{\lambda}_{i}\}_{i=1}^{N}$, respectively, and the estimator of the function $h$ is the $K$th degree polynomial

equation[equation omitted — 99 chars of source]

Note here that the optimization problem ((ref)) is highly non-convex and may be difficult to solve exactly. To make progress in this direction, we take advantage of the fact that for given values of $\{\phi_{t}\}_{t=1}^{T}$ and $\{l_{i}\}_{i=1}^{N}$, the optimization over $\{c_{j}\}_{j=0}^{K}$ is easy as it can be implemented via OLS. We therefore only need to take care of optimization over $\{\phi_{t}\}_{t=1}^{T}\in(0,1]^{T}$ and $\{l_{i}\}_{i=1}^{N}\in(0,1]^{N}$. In practice, we proceed with optimization over these two sequences as follows. First, we calculate the initial values of $\{\phi_{t}\}_{t=1}^{T}$ and $\{l_{i}\}_{i=1}^{N}$ as the first left and right singular vectors of the $T\times N$ matrix $\{r_{it}\}_{t,i=1}^{T,N}$ shifted and scaled in a way to make sure that their components are taking values in the $(0,1]$ interval. Second, we perform gradient descent starting from the initial values to find a local minimum of the optimization problem. Third, we reoptimize each component of the sequences $\{\phi_{t}\}_{t=1}^{T}$ and $\{l_{i}\}_{i=1}^{N}$ in turn multiple times until the change in the criterion function from reoptimization becomes negligible.

Remark (Other Sieve Spaces). We emphasize that although we focus on the polynomial sieves, there exist many other sieve spaces that could be used to approximate the function $h$ arbitrarily well. For example, one could consider splines, wavelets, or neural networks. Our theory in the next section could be extended to cover these alternative sieve spaces but we have decided to focus on the polynomial sieves for conciseness of the derivations.\qed

Rate of Convergence

In this section, we derive the rate of convergence of our sieve-based least squares estimator. To do so, we need the concept of the sub-Gaussian norm, also known as the $\psi_{2}$ norm, which is defined as follows. Let $\psi_{2}$ be the function defined by $\psi_{2}(x)=\exp(x^{2})$ for all $x\geq0$. Then for any random variable $X$, its sub-Gaussian norm $\|X\|_{\psi_{2}}$ is the smallest constant $C>0$ such that $E[\psi(|X|/C)]\leq1$. In addition, we impose the following assumptions.

assumption[Idiosyncratic Noise] (i) Random variables $\varepsilon_{i,t}$ are mutually indepenent across $i$ and $t$. (ii) $\max_{1\leq i\leq N}\max_{1\leq t\leq T}\|\varepsilon_{it}\|_{\psi_{2}}=O(1).$

Assumption (ref)(i) is fairly standard and can be relaxed to allow weak dependence in exchange for a more complicated analysis. Assumption (ref)(ii) means that the tails of the random variables $\varepsilon_{it}$ are not heavier than those of the Gaussian random variables. This assumption is also rather standard in the literature.

assumption[Approximation Error] There exist constants $\alpha>0$ and $L>0$ such that for each $K$, there exist constants $h_{0}^{K},h_{1}^{K}\dots,h_{K}^{K}\in[-L,L]$ such that the polynomial $h_{K}(x)=\sum_{j=0}^{K}h_{j}^{K}x^{j}$, $x\in[0,1]$, satisfies $\sup_{x\in[0,1]}|h_{K}(x)-h(x)|=O(K^{-\alpha})$.

This assumption means that the function $h$ can be approximated by finite-degree polynomials and specifies how fast the approximation error in the uniform metric converges to zero as we increase the degree of the polynomials.

In light of Assumption (ref), for the theoretical analysis in this section, we assume that optimization in ((ref)) over slope parameters $\{c_{j}\}_{j=0}^{K}$ is performed so that each of these parameters remains bounded: $|c_{j}|\leq L$ for all $j=0,\dots,K$ and the same constant $L$ as that specified in Assumption (ref). In practice, we find that these constraints are not binding if the constant $L$ is chosen large enough.

The following theorem establishes the main result of this section.

theorem[Rate of Convergence] Suppose that Assumptions (ref) and (ref) are satisfied. Then \begin{equation} \frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T}(\widehat{h}(\widehat{f}_{t}\widehat{\lambda}_{i})-h(f_{t}\lambda_{i}))^{2}=O_{P}\left(\frac{(T+N+K)\log(TNK)}{NT}+K^{-2\alpha}\right). \end{equation}

The proof of this theorem is in Appendix (ref). The theorem shows that our sieve-based least-squares estimator is consistent as long $K\to\infty$ together with $N$ and $T$ in such a way that $K\log(TN)=o(NT).$ Following standard terminology of nonparametric estimation, the first and the second terms on the right-hand side of ((ref)) represent the variance and the bias terms, respectively. The variance term is increasing in the degree of the polynomial sieve $K$ and the bias term is decreasing in $K$. Moreover, if $K=o(\max(N,T))$, then it follows from the theorem that \[ \frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T}(\widehat{h}(\widehat{f}_{t}\widehat{\lambda}_{i})-h(f_{t}\lambda_{i}))^{2}=O_{P}\left(\frac{\log(NT)}{\min(N,T)}+K^{-2\alpha}\right). \] Thus, the theorem shows that the variance term converges to zero with a fast rate even if we set $K$ to be rather large.\footnote{Note also that Theorem (ref) derives the convergence rate of the HFL estimator $\widehat{h}(\widehat{f}_{t}\widehat{\lambda}_{i})$ which we use for all of our empirical results. Another interesting question is about convergence rates of the estimators $\widehat{f}_{t}$ and $\widehat{\lambda}_{i}$. It turns out that answering those questions is more difficult. In fact, even deriving conditions under which $f_{t}$ and $\lambda_{i}$ are identified, modulo certain normalizations, is highly non-trivial. In a separate paper (\citet*{Borri2024factor}), we establish identification and consistency of $f_{t}$ and $\lambda_{i}$.}

Data and Summary Statistics

We present a description of the data used for the estimation of the HFL model. In order to estimate the HFL model and perform cross-sectional asset pricing tests, we use data for a range of primary asset classes: equities, bonds, commodities, and currencies.

For equities, we use the fama1993common 25 size and value sorted portfolios and the fama2015five 25 size and operating profitability sorted portfolios. We also include the size, value and operating profitability equity factors. These equity portfolios and factors are from Ken French's data library. For bonds, we consider U.S. bonds and international bonds. For U.S. bonds, we use eleven maturity-sorted government bond portfolios from CRSP's “Fama Bond Portfolios” file with maturities in six month intervals up to ten years. For international bonds, we use the Refinitiv Datastream Government Bond Indices for 8 advanced economies (Australia, Austria, Canada, France, Germany, Japan, Netherlands and the United Kingdom). Specifically, we consider bond indices with maturity 1-3 years, 2 years, 5 years, 7 years and 10 years. These bond indices include coupon payments and bonds denominated in each country local currency. For commodities, we use returns of 21 commodity spot price indices from Datastream. The commodities are aluminum, Brent oil, cocoa, copper, corn, cotton, crude oil, eggs, gasoline, gold, natural gas, oats, platinum, pork bellies, propane, silver, soybean, sugar, tin, and wheat. For foreign exchange, we use 46 portfolios from \citet*{nucera2024currency}. These portfolios are based on popular investment strategies: carry, short-term and long-term momentum, currency value, net foreign assets and liabilities in domestic currencies, term spread, long-term yields, and output gap.

The selection of advanced economies, for the government bond indices, and commodities is guided by the objective of constructing a balanced long sample, which in our case starts in January 1988. The sample ends in December 2017, when the sample of foreign exchange portfolios ends.

Table (ref) reports summary statistics of the test assets and $h\left(f_{t}\lambda_{i}\right)$ estimated from the sieve-based estimator. We report the mean and standard deviations of the excess returns and $h\left(f_{t}\lambda_{i}\right)$ for the whole sample and for each asset class. Returns are in excess of the U.S. risk-free rate and reported in percentage. The “All” sample includes 171 assets, each with 360 monthly return observations. The average excess return is 0.32% per month (3.84% annualized). The mean standard deviation is 4.46% (15.44% annualized). The average excess return predicted by the HFL model across all assets exactly matches the sample counterpart, with a monthly volatility which is lower by one order of magnitude. The lower volatility is explained by the fact that realized returns for assets include an idiosyncratic, and volatile, component, which in our model is captured by the volatility of the residuals. Equities and commodities, on average, have higher and more volatile returns than other assets. Bonds, both U.S. and international, have on average lower, and less volatile, returns. Foreign exchange portfolios, in our sample, have average returns not statistically different from zero. For the individual asset classes, as for the whole sample of assets, the mean excess return predicted by the HFL model is very close to the average asset excess return.

{}

table[table omitted — 1,297 chars of source]

Cross-Sectional Asset Pricing Tests

Baseline Results

In this section, we consider cross-sectional asset pricing tests. These tests assess whether the HFL model can account for cross-sectional differences in average excess returns. The HFL model is

equation[equation omitted — 117 chars of source]

where $r_{it}$ is the time $t$ excess return on asset $i$ and $\varepsilon_{it}$ is an idiosyncratic mean zero disturbance. The term $h(f_{t}\lambda_{i})$ includes the single factor $f_{t}$, common to all assets, and the asset $i$ factor loading $\lambda_{i}$ which are combined by the non-linear link function $h$ by Lemma 2.1. For each asset $i$, Equation ((ref)) implies $E[r_{it}]=E[h(f_{t}\lambda_{i})]$. We empirically test this prediction by estimating the following linear model

equation[equation omitted — 113 chars of source]

where the expectation operator is replaced by the sample mean (i.e., $E[r_{it}]=1/T\sum_{t=1}^{T}r_{it}$), $h(f_{t}\lambda_{i})$'s are replaced by the corresponding estimators $\widehat{h}(\widehat{f}_{t}\widehat{\lambda}_{i})$, $\alpha$ and $\beta$ are the intercept and slope coefficients, and $\epsilon$ is a mean zero residual term. Let $E[r_{it}]$ and $E[h(f_{t}\lambda_{i})]$ denote vectors containing, respectively, the averages of the excess returns and predictions from the model.

In the empirical analysis we test several implications of our one factor model, such as that the intercept is zero and that the slope coefficient is equal to 1. Furthermore, we evaluate the fit of the model with regression R-squared and mean absolute pricing error (MAPE). The latter measures the mean absolute residuals in the cross-sectional regression ((ref)). Our empirical strategy follows the classic two-step fama1973risk's methodology, where the first step (i.e., the asset level time-series regressions to retrieve the asset beta) is implicit in the algorithm to estimate the HFL model. The second step, that is the cross-sectional regression of mean asset returns on their betas, is replaced by the cross-sectional regression ((ref)).

Remark. In the Appendix, we demonstrate that unless the HFL model is correct, the probability limits $\alpha_{0}$ and $\beta_{0}$ of the OLS estimates $\widehat{\alpha}$ and $\widehat{\beta}$ of Equation ((ref)) generally satisfy the restrictions $\alpha_{0}=0$ and $\beta_{0}=1$ only if the factor $f_{t}$ is constant over time. Since the latter does not hold in practice, it follows that the OLS-based tests considered in this section are indeed consistent.\qed

Column (1) of Table ((ref)) reports estimates of ((ref)) using all assets for the period of January 1988 to December 2017. We consider estimates of the HFL model based on four degrees of the polynomial function used to approximate the function $h(.)$.\footnote{We find that using higher order of polynomials generate qualitatively similar results.} The table shows that the slope coefficient ($h\left(f_{t}\lambda_{i}\right)$) is significantly different from zero at conventional confidence levels. The pricing error, captured be the estimate of the intercept $\alpha$, is not significantly different from zero. The model provides a good fit with the adjusted $R^{2}$ of 89% and the MAPE is small at 0.101%.

{}

table[table omitted — 2,821 chars of source]

The good fit of the model is visually confirmed by Figure (ref), which plots the average asset excess returns versus the predicted excess returns using the HFL model. All assets line up along the 45 degree line. Figure (ref) in the Appendix reveals that this result is robust to the exclusion of the intercept from the model used to predict average excess returns.

We compare the HFL model with the standard capital asset pricing model (CAPM) of sharpe1964capital and lintner1965security, the fama1993common's three factor model, the fama2015five's five factors model, and the five factors model augmented with the momentum factor of jegadeesh1993returns.

Columns (2) to (5) of Table ((ref)) present the results of cross-sectional asset pricing tests using all assets. In all models, we include the HFL component ($h\left(f_{t}\lambda_{i}\right)$), and evaluate the effect of including additional factors. First, the slope estimate associated with the HFL component is always significantly different from zero at the 1% confidence level. Second, including the additional factors increases the adjusted R-squareds only marginally, from 89% to a maximum of 95%. In fact, with the exception of the slope coefficient associated with the size factor, which is significant at the 10% level, the slope coefficients associated with all the other factors are never statistically different from zero.

Comparison with PCA factor models

In this section, we compare the HFL model with factor models based on principal component analysis (see, e.g., connor1986performance,connor1988risk). Table (ref) presents the results of cross-sectional asset pricing tests using all assets and three estimators of the latent factors related to PCA. The first estimator is based on the standard PCA of the covariance matrix of asset returns. The second estimator, called RP-PCA, is a generalization of PCA, proposed by lettau2020factors, which includes a penalty term to account for the pricing errors in the cross-sectional regressions based on the RP-PCA factors. RP-PCA is motivated by the poor performance of standard PCA in identifying factors which are relevant to explain the cross-section of average excess returns (see, e.g., onatski2012asymptotics and \citet*{kozak2020shrinking}). The third estimator is the kernel-PCA, proposed by \citet*{scholkopf1998nonlinear}. Kernel-PCA (K-PCA) is a nonlinear form of PCA which allows for the separability of nonlinear data by projecting it onto a higher dimensional space where it becomes linearly separable using kernels.

center[center omitted — 711 chars of source]

In all models, we include the HFL component ($h\left(f_{t}\lambda_{i}\right)$), and evaluate the effect of including, respectively, the first principal component, the first two principal components, and the first three principal components, separately for standard PCA, RP-PCA and K-PCA. The estimates from the cross-sectional asset pricing regressions reveal that the slope coefficient associated with the HFL component is statistically significant at conventional levels in all models, while the slopes associated with the principal components, for all three estimators we consider, are never statically different from zero. Furthermore, including the principal components increases only marginally the cross-sectional adjusted R-squared, from 89% to a maximum of 91%, while the pricing error $\alpha$ is never statistically significant at conventional condidence levels.

Table (ref) in the Appendix reveals similar conclusions for asset pricing estimates of PCA models which include up to 6 factors. These estimates confirm the robustness of the results based on PCA models including the first three components.

Comparison with higher order factor models

In this section, we compare the HFL model with higher order factor models, related to the non-linear version of the CAPM of kraus1976skewness. harvey2000conditional develop a three-moment conditional CAPM where skewness is priced. \citet*{ang2006cross} develop a model where aggregate market volatility is priced. Table (ref) presents the results of cross-sectional asset pricing tests using all assets. In all models, we include the HFL component ($h\left(f_{t}\lambda_{i}\right)$), and evaluate the effect of including higher order factors. Specifically, we consider both a model specification which additionally includes the U.S. equity market excess returns, the squared U.S. equity market excess returns, and the cubed U.S. equity market excess returns, and the second model specification which additionally includes the first principal component, the squared first principal component, and the cubed first principal component. For the principal components, we consider three estimators: the standard PCA based on the covariance matrix of asset returns; the risk premium PCA (RP-PCA) of lettau2020factors, and the kernel PCA (K-PCA) of \citet*{scholkopf1998nonlinear}. In all modes, the slope coefficient associated to the HFL component is statistically significant at conventional confidence levels, while the slopes associated with the higher order factors are never statistically different from zero. Furthermore, including the principal components increases only marginally the cross-sectional R-squared, from 89% to a maximum of 92%. Moreover, the pricing error $\alpha$ is not statistically different from zero in all models, with the exception of the CAPM model augmented with the equity market factor squared and cubed (CAPM$^{3}$) for which the estimate for the intercept is statistically significant at the 10% level.

{}

sidewaystable[H] {\caption{Comparison with PCA Factor Models} } { } {This table presents the results of cross-sectional asset pricing tests using all assets. We estimate the cross-sectional linear model $E[r_{it}]=\alpha+\beta E[h(f_{t}\lambda_{i})]+\epsilon_{i}$ augmented by, respectively, the first principal component (PC1), the first two principal components (PC1+PC2) and the first three principal components extracted by the test asset returns. For the principal components, we consider three estimators: the standard PCA based on the covariance matrix of asset returns; the risk premium PCA (RP-PCA) of lettau2020factors, and the kernel PCA (K-PCA) of \citet*{scholkopf1998nonlinear}. We report the estimates for the slopes and intercept ($\alpha$), and standard errors adjusted with the Fama-MacBeth procedure in parenthesis. The table additionally reports the regression adjusted R-squared, and the mean absolute pricing error (MAPE) in percentage terms. {*}, {*}{*}, and {*}{*}{*} represent significance at the 1%, 5%, and 10% level. We further report the number of assets in the cross-sectional regression and the number of monthly observations for each asset used in the estimation of the averages of the excess returns and predictions from the model. The HFL model is based on a degree of the polynomial used to approximate the function $h(.)$ equal to 4.} { } {} \begin{tabular}{lccccccccccc} \hline & \multicolumn{3}{c}{Standard PCA} & & \multicolumn{3}{c}{RP-PCA} & & \multicolumn{3}{c}{K-PCA}\tabularnewline \hline & PC1 & PC1 + PC2 & PC1--PC3 & & PC1 & PC1 + PC2 & PC1--PC3 & & PC1 & PC1 + PC2 & PC1--PC3\tabularnewline $h\left(f_{t}\lambda_{i}\right)$ & 0.705{*}{*}{*} & 0.696{*}{*}{*} & 0.654{*}{*}{*} & & 0.692{*}{*}{*} & 0.656{*}{*}{*} & 0.635{*}{*}{*} & & 0.807{*}{*}{*} & 0.804{*}{*}{*} & 0.777{*}{*}{*}\tabularnewline & (0.129) & (0.120) & (0.109) & & (0.129) & (0.117) & (0.113) & & (0.174) & (0.176) & (0.169)\tabularnewline PC1 & 0.014 & 0.014 & 0.015 & & 0.014 & 0.016 & 0.016 & & 0.012 & 0.010 & 0.009\tabularnewline & (0.020) & (0.019) & (0.019) & & (0.020) & (0.019) & (0.019) & & (0.018) & (0.017) & (0.017)\tabularnewline PC2 & & 0.001 & 0.001 & & & 0.001 & 0.002 & & & 0.004 & 0.007\tabularnewline & & (0.010) & (0.010) & & & (0.010) & (0.010) & & & (0.013) & (0.014)\tabularnewline PC3 & & & -0.005 & & & & 0.003 & & & & -0.008\tabularnewline & & & (0.010) & & & & (0.010) & & & & (0.014)\tabularnewline $\alpha$ & 0.000 & 0.000 & 0.000 & & 0.000 & 0.000 & 0.001 & & 0.001 & 0.001 & 0.001\tabularnewline & (0.001) & (0.001) & (0.000) & & (0.001) & (0.001) & (0.000) & & (0.001) & (0.001) & (0.001)\tabularnewline Adj $R^{2}$ & 0.907 & 0.907 & 0.912 & & 0.908 & 0.911 & 0.914 & & 0.888 & 0.888 & 0.895\tabularnewline MAPE, % & 0.096 & 0.097 & 0.090 & & 0.095 & 0.094 & 0.089 & & 0.103 & 0.103 & 0.104\tabularnewline Assets & 171 & 171 & 171 & & 171 & 171 & 171 & & 171 & 171 & 171\tabularnewline Months & 360 & 360 & 360 & & 360 & 360 & 360 & & 360 & 360 & 360\tabularnewline & & & & & & & & & & & \tabularnewline \hline \end{tabular}

{\scriptsize}

table[table omitted — 6,004 chars of source]

{\scriptsize}

Comparison with recent macro-factor models

In this section, we compare the HFL model with factors from recent recent macro-factor models. Specifically, we consider \citet*{he2017intermediary}'s intermediary capital risk factor (HKM), the \citet*{lettau2014conditional}'s downside equity market risk factor (LMW), and pastor2003liquidity's liquidity risk factor (PS). Both the HKM and LMW factors are developed to investigate cross-sectional differences in average returns across large cross-section of assets from different asset classes. The PS factor is developed to investigate cross-sectional differences in equity average returns. In all cases, we compare the HFL model with single-index models based on the macro factors.\footnote{For the HKM and PS factors, we rely on the data provided by Assaf Manela and Lubos Pastor, respectively. We construct the LMW factor, for the period of analysis, following \citet*{lettau2014conditional}.}

Table (ref) presents the results of cross-sectional asset pricing tests using all assets. In all models, we include the HFL component ($h\left(f_{t}\lambda_{i}\right)$), and evaluate the effect of including, respectively, the HKM, LMW and PS factors. In all modes, the slope coefficient associated to the HFL component is statistically significant at conventional levels, while the slopes associated with the macro-factor models are never statistically different from zero. Furthermore, including the macro-factors increases only marginally the cross-sectional adjusted R-squared, from 89% to a maximum of 91%. The pricing error $\alpha$ is not statistically different from zero in all specifications.

{}

table[table omitted — 2,428 chars of source]
sidewaystable[H] \caption{Comparison with macro factors of ludvigson2009macro} { } \begin{singlespace} {This table presents the results of cross-sectional asset pricing tests using all assets. We estimate the cross-sectional linear model $E_{i}[r_{it}]=\alpha+\beta E_{i}[h(f_{t}\lambda_{i})]+\epsilon_{i}$ augmented by the 9 macro factors of ludvigson2009macro. These factors, denoted by $F_{LN}$, are eight macro factors extracted from a sample of US government bonds plus the first factor cubed. We report the estimates for the slopes and intercept ($\alpha$), and standard errors adjusted with the Fama-MacBeth procedure in parenthesis. The table additionally reports the regression adjusted R-squared. {*}, {*}{*}, and {*}{*}{*} represent significance at the 1%, 5%, and 10% level. We further report the number of assets in the cross-sectional regression and the number of monthly observations for each asset used in the estimation of the averages of the excess returns and predictions from the model. The HFL model is based on a degree of the polynomial used to approximate the function $h(.)$ equal to 4.} \end{singlespace} {\scriptsize }{\tiny} \begin{tabular}{lcccccccccccccccccc} \hline & \multicolumn{9}{c}{{\tinyWithout HFL}} & \multicolumn{9}{c}{{\tinyWith HFL}}\tabularnewline \hline & {\tiny$F_{LN,1}$} & {\tiny$F_{LN,1}^{3}$} & {\tiny$F_{LN,2}$} & {\tiny$F_{LN,3}$} & {\tiny$F_{LN,4}$} & {\tiny$F_{LN,5}$} & {\tiny$F_{LN6}$} & {\tiny$F_{LN7}$} & {\tiny$F_{LN8}$} & {\tiny$F_{LN,1}$} & {\tiny$F_{LN,1}^{3}$} & {\tiny$F_{LN,2}$} & {\tiny$F_{LN,3}$} & {\tiny$F_{LN,4}$} & {\tiny$F_{LN,5}$} & {\tiny$F_{LN6}$} & {\tiny$F_{LN7}$} & {\tiny$F_{LN8}$}\tabularnewline \hline {\tiny$F_{LN,1}$} & {\tiny-0.125{*}} & & & & & & & & & {\tiny-0.007} & & & & & & & & \tabularnewline & {\tiny(-1.830)} & & & & & & & & & {\tiny(-0.118)} & & & & & & & & \tabularnewline {\tiny$F_{LN,1}^{3}$} & & {\tiny-0.289{*}{*}} & & & & & & & & & {\tiny-0.011} & & & & & & & \tabularnewline & & {\tiny(-2.469)} & & & & & & & & & {\tiny(-0.113)} & & & & & & & \tabularnewline {\tiny$F_{LN,2}$} & & & {\tiny0.049} & & & & & & & & & {\tiny0.023} & & & & & & \tabularnewline & & & {\tiny(1.078)} & & & & & & & & & {\tiny(0.536)} & & & & & & \tabularnewline {\tiny$F_{LN,3}$} & & & & {\tiny0.062} & & & & & & & & & {\tiny-0.002} & & & & & \tabularnewline & & & & {\tiny(1.621)} & & & & & & & & & {\tiny(-0.042)} & & & & & \tabularnewline {\tiny$F_{LN,4}$} & & & & & {\tiny-0.103{*}{*}} & & & & & & & & & {\tiny-0.008} & & & & \tabularnewline & & & & & {\tiny(-2.128)} & & & & & & & & & {\tiny(-0.200)} & & & & \tabularnewline {\tiny$F_{LN,5}$} & & & & & & {\tiny0.037} & & & & & & & & & {\tiny-0.013} & & & \tabularnewline & & & & & & {\tiny(0.989)} & & & & & & & & & {\tiny(-0.324)} & & & \tabularnewline {\tiny$F_{LN6}$} & & & & & & & {\tiny0.166{*}{*}{*}} & & & & & & & & & {\tiny0.056} & & \tabularnewline & & & & & & & {\tiny(3.560)} & & & & & & & & & {\tiny(1.524)} & & \tabularnewline {\tiny$F_{LN7}$} & & & & & & & & {\tiny0.046{*}{*}{*}} & & & & & & & & & {\tiny0.012} & \tabularnewline & & & & & & & & {\tiny(2.727)} & & & & & & & & & {\tiny(0.648)} & \tabularnewline {\tiny$F_{LN8}$} & & & & & & & & & {\tiny-0.047{*}{*}{*}} & & & & & & & & & {\tiny-0.006}\tabularnewline & & & & & & & & & {\tiny(-2.939)} & & & & & & & & & {\tiny(-0.389)}\tabularnewline {\tiny$h\left(f_{t}\lambda_{i}\right)$} & & & & & & & & & & {\tiny0.810{*}{*}{*}} & {\tiny0.808{*}{*}{*}} & {\tiny0.806{*}{*}{*}} & {\tiny0.817{*}{*}{*}} & {\tiny0.801{*}{*}{*}} & {\tiny0.825{*}{*}{*}} & {\tiny0.774{*}{*}{*}} & {\tiny0.701{*}{*}{*}} & {\tiny0.794{*}{*}{*}}\tabularnewline & & & & & & & & & & {\tiny(4.890)} & {\tiny(5.008)} & {\tiny(4.732)} & {\tiny(4.635)} & {\tiny(5.503)} & {\tiny(4.440)} & {\tiny(4.505)} & {\tiny(5.230)} & {\tiny(4.170)}\tabularnewline {\tiny$\alpha$} & {\tiny0.003{*}{*}{*}} & {\tiny0.002{*}{*}{*}} & {\tiny0.002{*}{*}{*}} & {\tiny0.003{*}{*}{*}} & {\tiny0.001{*}} & {\tiny0.004{*}{*}{*}} & {\tiny0.005{*}{*}{*}} & {\tiny0.001{*}{*}} & {\tiny0.005{*}{*}{*}} & {\tiny0.001} & {\tiny0.001} & {\tiny0.000} & {\tiny0.001} & {\tiny0.000} & {\tiny0.000} & {\tiny0.001} & {\tiny0.000} & {\tiny0.001{*}}\tabularnewline & {\tiny(3.473)} & {\tiny(3.246)} & {\tiny(4.026)} & {\tiny(3.388)} & {\tiny(1.775)} & {\tiny(3.991)} & {\tiny(3.865)} & {\tiny(2.032)} & {\tiny(4.555)} & {\tiny(0.854)} & {\tiny(0.859)} & {\tiny(0.364)} & {\tiny(0.886)} & {\tiny(0.686)} & {\tiny(0.851)} & {\tiny(1.562)} & {\tiny(0.700)} & {\tiny(1.770)}\tabularnewline {\tinyAdj $R^{2}$} & {\tiny0.107} & {\tiny0.213} & {\tiny0.036} & {\tiny0.058} & {\tiny0.213} & {\tiny0.016} & {\tiny0.189} & {\tiny0.559} & {\tiny0.192} & {\tiny0.885} & {\tiny0.885} & {\tiny0.894} & {\tiny0.885} & {\tiny0.886} & {\tiny0.887} & {\tiny0.904} & {\tiny0.904} & {\tiny0.888}\tabularnewline {\tinyMonths} & {\tiny360} & {\tiny360} & {\tiny360} & {\tiny360} & {\tiny360} & {\tiny360} & {\tiny360} & {\tiny360} & {\tiny360} & {\tiny360} & {\tiny360} & {\tiny360} & {\tiny360} & {\tiny360} & {\tiny360} & {\tiny360} & {\tiny360} & {\tiny360}\tabularnewline \hline \end{tabular}{\tiny}

Comparison with macro factors of ludvigson2009macro

In this section, we compare the HFL model with nine macro factors proposed by ludvigson2009macro to study bond risk premia.\footnote{The macro factors of ludvigson2009macro are from Sydney Ludvigson's website.} These factors are eight factors extracted using dynamic factor analysis and a panel of US government bonds, and the first factor cubed. For each of these factors (denoted with $F_{LN}$), separately, we first estimate the cross-sectional regression $E[r_{it}]=\alpha+\beta E[F_{LN}]+\epsilon_{i}$. Next, we repeat the estimation additionally including the predicted values obtained with the HFL model (i.e., $E[h(f_{t}\lambda_{i})])$.

Table (ref) summarizes our results. When taken individually, each of the macro factors is significant, except for the second, third and fifth factors. When we additionally include in the regression specification the HFL component, all of the nine macro factors are not statistically different from zero. In contrast, the HFL component is statistically significant in all models.

Comparison with the factor zoo

In this section, we compare the HFL model with a large set of factors proposed in the last decades to account for the cross-section of asset returns (see, e.g., \citet*{feng2020taming}). In particular, we consider factors for 153 characteristics in 13 themes, using data from 93 countries and four regions, constructed by \citet*{jensen2023there}.\footnote{The data for the factor zoo is available through the Global Factor Data website. We use capped value weighted data for all factors and all countries.} For each of these factors (denoted with $f^{zoo}$), separately, we first estimate the cross-sectional regression $E[r_{it}]=\alpha+\beta^{zoo}E[f_{t}^{zoo}]+\epsilon_{i}$. Next, we repeat the estimation additionally including the predicted values obtained with the HFL model (i.e., $E[h(f_{t}\lambda_{i})])$. Figure (ref) summarizes our results. It plots the average of the absolute values of the $t$-statistics in a test of the intercepts and slope coefficients in the cross-sectional regressions, respectively, equal to zero. In the left panel of the figure, we consider the regressions which only include the factor asset betas. In the right panel of the figure, we consider regressions which additionally include the average predicted values from the HFL model. The figure reveals that, for the models which do not include the HFL component, the estimates of the slope coefficients are on average significantly different from zero (average absolute $t$-statistics equal to 2.06). For these models, though, the pricing errors, captured by the estimates of the intercept, are on average statistically significant (average absolute $t$-statistics equal to 2.91). Furthermore, the figure reveals that, in the models which additionally include the average predicted values from the HFL model, the slope coefficients associated with the factor zoo become not statistically different from zero (average absolute $t$-statistics equal to 0.85), while those associated with the HFL component are highly significant (average absolute t-statistics equal to 5.29). Moreover, for the models which includes the HFL component, the estimates of the pricing error also become not statistically different from zero (average $t$-statistics equal to 0.21).

center[center omitted — 1,538 chars of source]

Table (ref) provides further details about the comparison with the factor zoo. If the HFL component is not included (Panel A), the fraction of models with a slope coefficient on the factor zoo which is significant at the 10% confidence level is approximately 78%. When we further include the HFL component, this fraction drops to approximately 2%. Moreover, if the HFL component is included (Panel B), all models have a slope on the HFL component which is significant at the 1% confidence level. In these models, the intercept is never statistically different from zero at a confidence level of 10% or lower. Finally, we find that only three of the 153 factors from the factor zoo are significant with a confidence level of 10%. These factors are the capital turnover of haugen1996commonality; the dollar trading volume of \citet*{brennan1998alternative}; and the share turnover of \citet*{datar1998liquidity}.

The double-selection Lasso of \citet*{belloni2014inference} and \citet*{chernozhukov2018double} allows estimating regressions where the number of right-hand side variables could be large, even larger than the sample size. In practice, double-selection Lasso estimates only one coefficient on the right-hand side at a time by means of a two-step selection method. In the first, factors with a low contribution to the cross-sectional pricing are excluded from the set of controls. In the second step, factors whose covariances with returns are highly correlated in the cross-section with the covariance between returns and a given factor are added to the set of controls. Therefore, the double-selection Lasso achieves dimensionality reduction by eliminating useless and redundant factors. \citet*{feng2020taming} propose double-selection Lasso as model selection method to evaluate the contribution of factors in cross-sectional asset pricing regressions.

The double-selection Lasso is based on two Lasso regressions which contain a regularization parameter. In setting these parameters, we follow a cross-validation, or CV, procedure (see, e.g., \citet*{hastie2009elements}). We first set a broad range of values for each regularization parameter. Next, we split the data into 5 folds, and train the model on 4 folds with one fold held out as validation set, cycling through all folds. This process is repeated for each value in the range of values for the regularization parameters and the model performance is evaluated on the validation set using the mean squared error as the metric. The parameter value with the lowest average mean squared error across all folds is selected as the chosen regularization parameter.

The HFL component using the double-selection Lasso method is statistically significant at the 1% level ($t$-statistics equal to 7.34). In this case, the intercept is not statistically different from zero ($t$-statistics equal to 1.37). Furthermore, Panel C of table (ref) reports summary statistics on the statistical significance of slope coefficients on the factor zoo, after controlling for the HFL component, as well as for the remaining factors in the factor zoo. The table reveals that the fraction of significant factors in the factor zoo, with a confidence of 10% or below, is approximately 10%, or 14 factors (see Table (ref) in the Appendix for the list of significant factors).

Predictability

In this section, we study whether the HFL model can predict returns by investigating the properties of portfolios constructed on the basis of the predictions of the HFL model. We start by estimating the model on all 171 assets and the first 120 months of data.\footnote{The 171 assets contain three equity factors, and the results are qualitatively similar when we exclude them in the predictability exercise.} Next, we use the model predicted returns to sort assets in five portfolios, and compute the equally-weighted average portfolio return in the following six months. Finally, we repeat the procedure expanding the estimation window by 6 months each time until the end of the sample (i.e., we repeat the procedure 40 times in order to reach the end of the sample in December, 2017). In each period, the first four portfolios contain 34 assets each, and the last portfolio contains the remaining 35 assets. This empirical methodology is designed to capture the strategy of an investor who uses the model to predict future returns rebalancing at a semi-annual frequency. In each period $t$, we sort assets only using information available up to period $t$ and compute excess returns between $t$ and $t+1,$where each period contains 6 months. Portfolio 1 groups the assets with, each period, the lowest predicted returns by the HFL model. Portfolio 5 groups the assets with, each period, the highest predicted returns by the HFL model.

Table (ref) describes the properties of the five portfolios, as well as of a long-short strategy long in portfolio 5 and short in portfolio 1, and their risk-adjusted performance. The analysis of Panel A reveals the predictive ability of the HFL model in forecasting future returns. Across the portfolios, there is a sizable cross-section in average monthly returns, ranging from 0.08% for the first portfolio to 0.77% for the last one. Notably, a zero-cost long-short strategy, wherein assets with the highest predicted returns are bought while those with the lowest predicted returns are sold short, yields a monthly average excess return of 0.69% (8.28% annualized), with a monthly Sharpe ratio of 0.15 (0.51 annualized). Despite the relatively short time-series, the cross-sectional return spread is statistically significant. For instance, the standard error for the long-short strategy is 30 basis points. Consequently, the average excess return is more than two standard errors from zero. Similarly, each individual portfolio, except the first two portfolios, offers average returns which are statistically different from zero. Panel B highlights the significant risk-adjusted performance of the strategy based on the HFL model. For each portfolio, we report the estimates of the intercept in a regression of the portfolio return on the \citet*{fama1993common}'s three factors ($\alpha_{FF3})$, on the \citet*{fama2015five}'s five factors ($\alpha_{FF5}$), and on the five factors augmented by the momentum factor of \citet*{jegadeesh1993returns} ($\alpha_{FF5+MOM})$, respectively. These estimates of the intercept are a measure of risk-adjusted performance. First, the table reveals that the $\alpha$ estimates increase from the first to the last portfolio in all three specifications. Second, we find that the alphas for portfolio 3 to 5, as well as for the long-short portfolio, are statistically different from zero at conventional levels. For the long-short portfolio, the monthly alphas are equal, respectively, to 0.60%, 0.78% and 0.73%, and are statically significant at the 5% level.

table[table omitted — 2,553 chars of source]
table[table omitted — 3,382 chars of source]

Additional Results

In this section, we describe four sets of additional results related to our main findings. First, we report the results of asset pricing tests for individual asset classes. Second, we report further asset pricing tests using the first and second half of our sample. Third, we investigate the time-variation in the asset pricing test using an estimation based on an expanding window. Fourth, we consider a larger set of test assets, which further includes the equity momentum portfolios.

Asset pricing tests for individual asset classes

In this section, we present the result of cross-sectional asset pricing tests for individual asset classes. Table (ref) summarizes these results. The slope associated with the HFL component is statistically significant for each asset class. Specifically, it is significant at the 1% level for each asset classes except for the case of commodity, for which the slope is significant at the 10% level. The estimates for the intercept are small in magnitude, and statistically different from zero, at the 5% level, only for US bonds and international bonds. Furthermore, the adjusted R-squareds are above 85% for each asset class, with the exception of equities (47%). The MAPE is the highest for equities (0.13%) and the smallest for US bonds (0.006%). Section (ref) in the Appendix contains scatter plots of average excess returns against predicted excess returns for each individual asset class.

{}

table[table omitted — 2,186 chars of source]

Asset-pricing using alternative samples

In this section, we present results of cross-sectional asset pricing tests by asset class using two alternative samples. The first sample starts in January 1988 and ends in December 2003 (first half). The second sample starts in January 2004 and ends in December 2017 (second half). We report the estimates for these two sub-samples, each including 180 monthly dates, in Table (ref).

Panel A refers to the estimates using the first half of the sample. In the model with all assets, the slope coefficient associated with $h\left(f_{t}\lambda_{i}\right)$ is statistically different from zero. Furthermore, the estimate for the $\alpha$, the model pricing error, is not statistically different from zero and the cross-sectional regression adjusted R-squared is equal to 86%. For the individual asset classes, we reject the null of slope associated with the HFL component equal to zero for all assets, except for commodities (for equities the $p$-value is equal to 9.6%). The cross-sectional regression adjusted R-squared is highest for US bonds (96%) and lowest for equities (60%).

Panel B refers to the, more recent, second half of the sample. Similarly to the results for the first half of the sample, in the model with all assets, the slope coefficient associated with $h\left(f_{t}\lambda_{i}\right)$ is statistically different from zero. The slope point estimate is equal to 0.82, whereas it is equal to 0.85 in the first half of the sample. Furthermore, the estimate for the $\alpha$, the model pricing error, is not statistically different from zero. However, the model fit is lower than in the first half, with a regression adjusted R-squared of 77% and a MAPE of 0.14%. For the individual asset classes, similarly to the estimates for the first half of the sample, we reject the null of slope equal to zero for the HFL component for all asset classes with the exception of commodities. The cross-sectional regression adjusted R-squared is highest for US bonds (98%) and lowest for equities (22%).

Table (ref) in the Appendix presents asset pricing estimates for two samples before and after the global financial crisis of 2007-08, which further confirm the robustness of our results for alternative samples.

Time variation in asset-pricing tests

In this section, we further investigate time-variation in the cross-sectional asset pricing tests. In section (ref), we showed that the estimates for the slope associated with the HFL component, on all assets, are similar and statistically significant in both the first and second half of the sample. In this section, we provide further robustness evidence by plotting, in Figure (ref), time-varying monthly coefficient estimates. The time-varying estimates are based on an expanding window and starting with a minimum window size of 10 years of data at the monthly frequency following eugene1992cross. The figure reveals that the slope associated with $E[h(f_{t}\lambda_{i})]$ is statistically different from zero for all months, as highlighted by the shaded region which denotes a two standard error confidence band. Furthermore, we can never reject the hypothesis that the slope is equal to 1, which is one of the cross-sectional predictions of our model.

{}

table[table omitted — 3,261 chars of source]
figure[figure omitted — 980 chars of source]

Asset-pricing with momentum portfolios

Table (ref) presents estimates of cross-sectional asset pricing tests using a set of test assets which additionally includes 10 momentum portfolios and the momentum factor. Both the momentum portfolios and the momentum factor are from Ken French's website and are based on U.S. stocks monthly returns sorted by the 12-2 past return. The table reveals that the slope coefficient associated with the averages of the predicted values from the HFL component is significantly different from zero for all assets as well as for each individual asset class. The pricing error, captured be the estimate of the intercept, is significantly different from zero only for US bonds (significant at the 5%). The model provides the closest fit for US bonds (adjusted-$R^{2}$ of 99%), although the model fit is similar across the specifications with the exception of equities with a lower value for the adjusted R-squared of 54%. The MAPE ranges from 0.005% (US bonds) to 0.122% (equities).

{}

table[table omitted — 2,121 chars of source]

Conclusion

We propose a parsimonious non-linear single factor asset pricing model $r_{it}=h(f_{t}\lambda_{i})+\epsilon_{it}$. Despite its parsimony, the HFL model represents exactly any non-linear model with an arbitrary number of factors and factor loadings as a consequence of the Kolmogorov-Arnold representation theorem. Empirically, we jointly estimate the link function, the time factor, and the loading with the sieve-based least square estimator.

Using 171 test assets across major asset classes across U.S. equity, U.S. and international bonds, commodities, and currencies, we show that this one factor model delivers superior cross-sectional empirical asset pricing performance with a low-dimensional approximation of the link function. Moreover, controlling for the HFL model, the majority of the established factors for the cross-section of asset returns becomes insignificant. Additionally, we use the HFL model to construct a tradable strategy that generates statistically significant and sizable risk-adjusted long-short excess returns

There are two other broader points we want to make. First, nonlinearity and the higher-order interactions are the key feature of the HFL model. Either as a stand alone tool or incorporating it in other asset pricing settings thus allows a parsimonious, yet general, way to capture the importance of these effects compared to linear asset pricing models. Second, there is significant recent interest in using neural nets and artificial intelligence models in asset-pricing. It is interesting to speculate why these models may deliver superior performance to the existing asset-pricing models. One answer is the Kolmogorov-Arnold representation theorem. This theorem shows that any nonlinear continuous function of many variables can be represented as a continuous outer layer and an inner layer consisting of the sum of functions of single variables. \citet*{hecht1987kolmogorov} argues that this is exactly the structure of a general neural network. Our HFL model thus offers a parsimonious method to capture the performance of neural network models and to evaluate their performance for asset pricing.

References

btSect[econ-aea]{4_Users_nicolaborri_Dropbox_Research_INDEX-EVERYTHING_lyx_bib_ka} \btPrintCited