EconBase
← Back to paper

Dynamic Latent-Factor Model with High-Dimensional Asset Characteristics

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

93,761 characters · 14 sections · 110 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Dynamic Latent-Factor Model with High-Dimensional Asset Characteristics

titlepageWe develop novel estimation procedures with supporting econometric theory for a dynamic latent-factor model with high-dimensional asset characteristics, that is, the number of characteristics is on the order of the sample size. Utilizing the Double Selection Lasso estimator, our procedure employs regularization to eliminate characteristics with low signal-to-noise ratios yet maintains asymptotically valid inference for asset pricing tests. The crypto asset class is well-suited for applying this model given the limited number of tradable assets and years of data as well as the rich set of available asset characteristics. The empirical results present out-of-sample pricing abilities and risk-adjusted returns for our novel estimator as compared to benchmark methods. We provide an inference procedure for measuring the risk premium of an observable nontradable factor, and employ this to find that the inflation-mimicking portfolio in the crypto asset class has positive risk compensation.

Introduction

In this manuscript, we develop novel estimation procedures with supporting econometric theory for a dynamic latent-factor model with high-dimensional asset characteristics. Additionally, we develop econometric methods for characteristic importance measures in this model and, with asymptotically valid inference, an asset pricing test to measure the risk premium of a nontradable factor. We close with provide empirical results to investigate why different crypto assets earn different average returns; conduct inference for crypto's inflation risk premium; and, estimate risk premia of crypto asset excess returns with classic factor models and our new dynamic latent-factor model.

\paragraph{Factor Model} We assume a statistical model where asset excess returns $r_{i,t+1} \in \R$ are a function of common time-varying latent (unobserved) factors, $f_{t+1} \in \R^k$, as dictated by time-varying asset-specific factor loadings $\b_{i,t} \in \R^k,$ that is, for all assets $i \in \{1, \dots, N\}$ and time $t \in \{1, \dots, T\}$

equation[equation omitted — 184 chars of source]

where $\a_{i,t} \in \R$ are average pricing errors of the factor model; $\e^r_{i,t+1} \in \R$ are uncorrelated idiosyncratic errors, i.e., $\E_t [ \e^r_{i,t+1} f_{t+1} ] = 0$; $\G_\b \in \R^{p \times k}$ is a static latent loading parameter; and, $z_{i,t} \in \R^p$ are time-varying asset-specific characteristics where $p$ is high-dimensional, i.e., on the order of $N$ and $T.$ Crucially, we follow the established practice in the literature of assuming the number of factors $k$ is low-dimensional, i.e., $N,T,p >> k \in \{1, 2, 3, \dots\}.$ This model has theoretical underpinnings motivated by a structural model for asset excess returns or by the assumption of no arbitrage, as discussed in our literature review.

\paragraph{Double Selection Lasso Factor Model} The main contribution of this research is to develop new estimation and inference procedures, termed the the Double Selection Lasso Factor Model (DSLFM), to fit the latent factors $f_{t+1}$ and loadings $\G_\b$ in (ref) and to conduct standard asset pricing tests under the novel setting of high-dimensional asset characteristics.

The DSLFM remains consistent with the equilibrium asset pricing principle that risk premia are solely determined by risk exposures and specifies a linear loading mapping $\G_\b$ between characteristics and dynamic factor loadings $\b_{i,t}.$ We have two novel assumptions for $\G_\b.$ First, we develop estimation procedures and large-sample theory that allows $p,T,N\rightarrow\infty.$ Given our focus is studying the cross-section of crypto assets, this assumption is particularly relevant given the numerous asset characteristics available, as previously discussed, as well as the existence of only a small number of tradable assets and years of relevant data such that $p,T,N$ are of similar order. Second, we assume exact row sparsity in $\G_\b$; that is, only a small number of the $p$ asset characteristics determine the content of the factor loading, which matches empirical findings in cross-sectional asset pricing (babiak2021risk and bianchi2022dynamics). These novel assumptions within a dynamic latent-factor model require novel estimation procedures and supporting asymptotic theory.

The DSLFM aims to jointly and consistently estimate the loading matrix $\G_\b$ and latent factors $f_{t+1}.$ If we were to utilize the MSE objective function to minimize over the $p-$dimensional choice vector $\G_\b f_{t+1},$ for each $t,$ the mean-squared error of $\sum_{i} (r_{i,t+1} - z_{i,t}^\top \G_\b f_{t+1})^2$, we will have not only a noisily estimated design matrix when $p \sim N,$ (or, at worst, a nonsingular design matrix when $p>N$) but also a non-convex objective function given the interaction between minimization arguments $f_{t+1}$ and $\G_\b$. The next logical step would be to introduce sparsity in $\G_\b,$ which would amount to adding a regularization parameter to the aforementioned objective function to combat the curse of dimensionality from $z_{i,t}.$ However, although potentially helpful for minimizing MSE by decreasing the variance of the estimator, this regularization introduces a bias in estimation, which would lead to invalid asymptotic inference for asset pricing tests, defeating a goal of this research and, more broadly, the purpose of factor models in this field.

We therefore adapt for our purpose the Double Selection Lasso (DSL) estimator developed by belloni2014inference. The key insight from their work was to introduce an orthogonality wherein, assuming $\G_\b$ is row sparse, the regularization bias from the LASSO first-stage estimation does not pass through to the target parameter of interest when conducting inference. We first estimate for each time period $t$ and each characteristic $j$ the scalar $\G_{\b,j}^\top f_{t+1}$ using DSL; then, stacking these estimates into a $T \times p$ matrix, we use PCA to obtain separate estimates for latent loadings $\wh{\G}_\b$ and factors $\{ \wh{f}_{t+1} \}_{t=1}^T$; and, finally, we soft-threshold $\wh{\G}_\b$ to set numerous rows to zero given the assumption of sparsity for $\G_\b$.

This procedure has several additional benefits. Given the period-by-period cross-sectional regression---mirroring Fama-MacBeth regressions--our estimation procedure accommodates unbalanced panels. This first DSL step does require running $T \times p$ cross-sectional regressions; however, this can be done in parallel and is on the order of minutes in practice as each regression is computationally fast. In the second step, the high dimensionality of the PCA procedure, given we have a $p \times T$ matrix, is adapted from the established theory for $N \times T$ excess return matrices bai2003inferential. The final soft-thresholding step exploits the sparsity in $\G_\b$ to remove noise from characteristics with low signal-to-noise ratios. \footnote{Finally, just as DSL laid the groundwork for the more general Debiased Machine Learning (DML) theory, this work sets up future research to extend the framework with a semi-parametric specification to utilize the rich set of available machine learning estimators that have been shown to handle well the nonlinearities in cross-sectional asset returns gu2020empirical.}

Under standard DSL assumptions adapted to our setting belloni2014inference, high-dimensional PCA assumptions bai2003inferential, and assuming we observe the true number of latent factors bai2002determining, we develop the asymptotic consistency of the latent factors $\wh{f}_{t+1}$ and loading matrix $\Check{\G}_\b$ for the latent factors $f_{t+1}$ and loadings $\G_{\b}$, respectively. Monte Carlo simulations corroborate with finite-sample evidence that the performance of the DSLFM is comparable to or surpasses relevant benchmarks. As is standard in this setting, without further restrictions outlined in bai2013principal, $F^0$ and $\G^0_\b$ are not separately identifiable; hence, the $k \times k$ invertible matrix transformation $H$ appears in each asymptotic result. However, in many cases, knowing $F^0 H$ is equivalent to knowing $F^0;$ for example, using the regressor $F^0$ will give the same predictions as using the regressor $F^0 H$ given they have the same column space. Similarly, the target parameter in the coming inference result is rotation invariant to $H.$

To show the generality of these estimation procedures, we enrich our model---with one of several possible extensions---to address a common question in asset pricing research. We ask whether an observable, nontradable factor $g_{t+1} \in \R$ carries a risk premium: compensation for exposure to the risk factor holding constant exposure to all other sources of risk, i.e., variation with other factors. In the subsequent empirical applications, we investigate a common hypothesis for the crypto asset class: exposure to inflation offers crypto investors a positive risk premium.

Following a recent approach in the literature (giglio2021asset and giglio2021test), we assume the “true” latent factors $f_{t+1}$ can be decomposed into the latent-factor risk premia $\g \in \R^k$ and latent-factor innovations $v_{t+1} \in \R^k,$ that is, $f_{t+1} \coloneqq \g + v_{t+1}.$ Then, we specify the observable factor $g_{t+1}$ as potentially linearly correlated with the latent factors through

equation*[equation* omitted — 51 chars of source]

where $\h \in \R^{k}$ is an unknown parameter mapping the relation between the latent-factor innovations and the observable factors, and $\e^g_{t+1} \in \R$ is measurement error in $g_{t+1}.$

The risk premium of an observable factor---our target parameter in this extension---is defined to be the expected excess return of a portfolio with loading (i.e., beta) of 1 with respect to this factor in $g_{t+1}$ and zero loadings on all other factors; in this model, that parameter is $\g_g \coloneqq \h^\top \g = \h^{0 \top} H H^{-1} \g^0 = \h^{0 \top} \g^0,$ which utilizes the rotation invariant result of giglio2021asset.

We thus extend with our estimation procedure for $\g_g$ in a dynamic latent factor model with high-dimensional characteristics the estimation procedure of giglio2021asset, which is for a static latent factor model. We additionally develop our estimator's large-sample distribution and variance to conduct asset pricing tests on the sign of the observable factor risk premium. We apply this test for the risk premium of the observable inflation factor within the crypto asset panel studied herein.

\paragraph{Empirical Applications} We close with studying three empirical questions: the out-of-sample pricing ability of the DSLFM compared to benchmark methods, the drivers of returns using our bootstrapped inference method, and estimating the inflation risk premium in the crypto asset class with our asset pricing test. We utilize the same novel and rigorous panel of crypto asset returns developed in XYZ.

To begin, we compare the out-of-sample predictive power in the Q3-Q4 2022 data of three models, namely, a three-factor model of size, crypto market, and momentum; a latent three-factor model fit with PCA; and a dynamic latent-factor model fit with IPCA using a subset of the characteristics. We find IPCA outperforms the other models, suggestive of the signal in the characteristics. Its predictive pricing signal outperforms a random walk and it provides economically and statistically significant risk-adjusted returns for the zero-investment portfolio, whereas the other models underperform a random walk yet provide modest long-short risk-adjusted returns.

Utilizing the broader set of asset characteristics, we first establish the comparable out-of-sample predictive ability of the DSLFM compared to the benchmark methods, with supporting bootstrapped characteristic importance measures to elucidate the drivers of returns. Exchanges inflows and outflows were significant characteristics, showing the importance of these onchain measures. While DSLFM achieved a maximum, with one latent factor, out-of-sample Sharpe of $3.3,$ this underperforms ICPA's maximum, with one latent factor, Sharpe ratio of $4.07$.

Additionally, we implement our testing procedure to find that the crypto asset class provides investors a positive inflation risk premium. Early proponents of Bitcoin and other cryptocurrencies framed these as an outside options or hedges against traditional fiat currencies. To study this question, we use our extended model to recover the 10-year expected inflation mimicking portfolio and measure its risk premium. This inflation risk premium was estimated at a statistically significant 1.4 basis points (0.0097 standard error). This translates to a 7.3% annual excess return, suggestive of positive compensation for investors holding an inflation-hedged crypto portfolio, ceteris paribus.

Literature Review

This paper builds the econometric theory literature for high-dimensional factor models. giglio2022factor provide an excellent review of recent machine-learning based factor model applications and relevant econometric theory, including the common asymptotic frameworks of fixed $N$ and $T\rightarrow\infty,$ fixed $T$ and $N\rightarrow \infty,$ and $T,N\rightarrow \infty.$ In short, this paper extends this last asymptotic framework to be high-dimensional on the new dimension of the number of asset characteristics, i.e., $p,T,N \rightarrow \infty$.

Starting from either a structural model for asset excess returns in the style of the Capital Asset Pricing Model sharpe1964capital, or the assumption of no arbitrage, as in Arbitrage Pricing Theory ross1976arbitrage, a stochastic discount factor $m_{t+1} \in \R$ exists and an Euler equation, termed the Law of One Price, holds for asset excess returns $r_{i,t+1} \in \R$ for assets $i \in \{1,2,\dots,N\}$ in time periods $t \in \{1,2,\dots,T\}$

equation*[equation* omitted — 65 chars of source]

which by the definition of the variance and covariance operators,

equation*[equation* omitted — 249 chars of source]

As discussed in Section (ref) we directly assume the statistical model

equation*[equation* omitted — 81 chars of source]

To map this model to its theoretical underpinnings in the Law of One Price, one can assume for all $i$ and $t$: mean zero unobserved idiosyncratic errors $\E_t [\e^r_{i,t+1}] = 0,$ uncorrelated errors $\E_t [ \e^r_{i,t+1} f_{t+1} ] = 0,$ the price of risk associated with the factors to be defined as $\l_t \coloneqq \E_t [ f_{t+1} ],$ and zero average pricing errors $\a_{i,t} = 0$ cochrane2009asset. The factor model posits that asset excess returns $r_{i,t+1}$ signify compensation for asset-specific, time-varying exposure $\b_{i,t} \in \R^k$ to systematic risk factors $f_{t+1} \in \R^k.$

The classic factor model is a static loading observable factor model---in the style of fama1992cross---where pricing errors and static factor loadings $\b_i$ in (ref) using exogenously-defined factors are estimated via the “Fama-MacBeth” two-step procedure fama1973risk, which has rich supporting inferential theory shanken1992estimation. This procedure relies on ex ante declaration of observable factors $f_{t+1}$ formed as a convex combination (i.e., a portfolio) of sorted asset returns $r_{t+1}$ based on asset characteristics $z_{t} \in \R^p.$ This is likely an incomplete model of the relationship between $z_{t}$ and $r_{t+1}$ and prone to overfit.

Recognizing the proliferation of factors, a “Factor Zoo” cochrane2011presidential, explaining the cross section of expected returns, feng2020taming propose the use of Double Selection Lasso belloni2014inference combined with two-pass Fama-MacBeth regressions to evaluate the contribution of a new factor, $g_{t+1},$ explaining asset expected returns above and beyond an existing high-dimensional set of factors. However, as the recent empirical literature has shown (e.g., kelly2019characteristics; chen2020deep), allowing the data to construct the relevant latent factors offers superior explanatory and predictive power as compared to using a set of selected observable factors from the literature.

Factor models with latent factors have been a focal point since the development of APT ross1976arbitrage and early empirical efforts chamberlain1983funds. PCA is the common estimation framework. bai2003inferential develops the core theory for PCA estimation and inference under joint $N,T\rightarrow\infty$ high-dimensional asymptotics, with bai2002determining introducing a novel Information Criterion penalization to ascertain the true number of latent factors. Assumptions for the joint identification of factor loadings and factors are summarized in bai2013principal. This manuscript leverages these contributions yet employs instead a dynamic latent-factor model with $p,N,T\rightarrow\infty$ high-dimensional asymptotics. In the asset pricing context, the static model can still perform well for describing portfolios over time given the dynamics are captured by the dynamic factors; however, this has not performed as well in describing individual asset returns, as noted by ang2009using. Although this model allows the data to statistically inform the factor structure, it fails to incorporate rich asset characteristic data and it assumes a static factor loading that maps systematic risks to excess returns.

A recent and significant methodological advance, instrumented PCA (IPCA) from kelly2019characteristics, utilizes asset characteristics in a linear model of dynamic factor loadings to parameterize the more general semi-parametric method studied in connor2007semiparametric. That is,

equation*[equation* omitted — 64 chars of source]

where $\G_\b \in \R^{p \times k} $ is a loading matrix mapping asset characteristics to the factor loading. IPCA has several benefits, including compressing the $N \times T$ factor loading matrix $\b$ to a lower dimensional $p \times k$ loading matrix $\G_\b$, as well as specifying a time-varying relationship $\b_{i,t}$ between characteristics and returns, which as stated previously appears to be the empirical reality in crypto cross-sectional asset pricing. In a separate theory paper, the authors develop the asymptotic distributional theory---following $N,T\rightarrow \infty$ asymptotics from bai2003inferential---for the factor realizations and loadings under quite general identifying restrictions on loadings and factors kelly2020instrumented.

The IPCA procedure benefits from the following: the efficiency gains from using asset characteristics for estimating the latent factors and their loadings; accommodation of unbalanced panels; maintaining an expected return factor model structure to ascertain the economic relationships among factors and assets via the observable characteristics; and, a parametric model with inference procedures for asset pricing tests. The IPCA estimation procedure, however, is not possible under high-dimensional asset characteristics (i.e., $p > T,N$), or, if regularization is used, produces biased inference for asset pricing tests.

giglio2021asset develop a three-step procedure combining estimation of latent-factor model via PCA with standard two-pass regressions to recover the risk premium of an observable nontradable factor $g_{t+1} \R$, which are potentially correlated with the latent factors:

equation*[equation* omitted — 59 chars of source]

where $v_{t+1}$ are (mean-zero) latent-factor innovations (i.e., $f_{t+1} = \g + v_{t+1}$); $\h \in \R^{k}$ is a linear mapping of the true latent factors to the observed factors; and, $\e^g_{t+1}$ is measurement error. This allows the observable factors to either be some component of the latent factors or just correlated with $v_{t+1}$ and therefore still carry a risk premium. The target parameter is the risk premium associated with the observable factors $\g_g \coloneqq \h \g.$

To demonstrate the generality of the DSLFM, we extend our estimation procedure by adding the procedure of giglio2021asset to address this classic asset pricing test of the recovering the risk premium of an observable factor. The DSLFM theory presented in Section (ref) extends giglio2021asset by incorporating not only dynamic factor loadings, but also high-dimensional asset characteristics.

More recent literature has incorporated a wide array of machine learning-based estimation approaches within the factor model structure. gu2020empirical study a set of machine learning estimation procedures for measuring the equity risk premium to find that deep learning methods outperform out-of-sample. gu2021autoencoder develop a factor model in the structure of IPCA, but allows for non-linear mappings to the factor loadings and the factors through two feed-forward neural networks. Although, in practice they only use linear mappings to the factors, they still obtain out-of-sample predictive $R^2$ and Sharpe ratio gains in relation to benchmark methods. Other notable uses of deep learning within factor model structures are chen2020deep; feng2018deep; guijarro2021deep; among others.

A fundamental difference between common empirical settings for machine learning applications and their use in empirical asset pricing is the uniquely low signal-to-noise DGP. Theory suggests market efficiency prices in the signal, such that, the unforecastable idiosyncratic error dominates. This critical issue significantly compounds the curse-of-dimensionality of using high-dimensional asset characteristics, which further motivates the parsimonious specification of the factor model.

Double Selection Lasso Factor Model

Our goal is to consistently estimate a dynamic latent-factor model and conduct asset pricing tests with valid inference under the novel setting of high-dimensional asset characteristics, which is of particular relevance in the crypto asset class.

Setup

\paragraph{Setting and Observable Random Variables} Assume for time periods $t=1,2,\dots,T$ and assets $i=1,2,\dots,N,$ that we observe realizations of random variables for asset excess returns $r_{i,t+1}\in \R$ and asset characteristics $z_{i,t} \in \R^{p}.$ An asset's excess return is the simple return of asset $i$ from time $t$ to $t+1$ net the assumed simple return of the risk-free rate (e.g., one-month US Treasury Bill). An asset characteristic of asset $i$ is known at time $t:$ for example, the total fees for a crypto protocol between time $t-1$ and $t.$ Note that asset characteristics are information from the previous period to follow the established convention in the literature and to be able to use this model for prediction. Importantly, we will introduce the novel asymptotic assumption for dynamic latent-factor models to let $p$ grow to infinity simultaneously with $N$ and $T.$

\paragraph{Model} Given the highly nonlinear data-generating process observed in empirical asset returns (gu2020empirical; chen2020deep; bianchi2021bond), we specify a semi-parametric factor model---where $\b_{i,t}$ is a function of asset characteristics $z_{i,t}$---studied in recent literature (connor2007semiparametric; connor2012efficient; and fan2016projected). We assume a dynamic latent-factor model where

equation*[equation* omitted — 66 chars of source]

However, given we are interested in conducting inference, we specify a linear model---to build on the foundational work of kelly2019characteristics---for the factor loadings

equation*[equation* omitted — 347 chars of source]

\paragraph{Parameters and Unobserved Random Variables} $f_{t+1} \in \R^k$ are low-dimensional latent factors; $\beta_{i,t} \in \R^k$ are latent-factor loadings; $\G_\b \in \R^{p \times k}$ is an unknown factor loading parameter matrix; and, $\e^r_{i,t+1} \in \R$ and $\e^\b_{i,t} \in \R^k$ are unobserved idiosyncratic errors.

The latent factors $f_{t+1}$ should be interpreted as purely statistical in nature. That is, these risk factors do not necessarily capture fundamental shocks to productive technologies as modeled in canonical theoretical models. Nevertheless, the latent factors capture systemic risk or covariance among asset returns that is non-diversifiable. We follow the literature in restricting $k$ to be a small finite constant (i.e., $k \in \{1,2,3,4,5\}$) that, in our asymptotic theory, does not grow with $p,T,N.$ It should be noted that, although ubiquitous, it is nevertheless a strong assumption: the empirical content of asset returns can be captured by a small number of strictly time-varying systematic risk factors.

The specification of $\b_{i,t}$ provides several benefits. First, we enable the use of a dynamic loading to model a changing relationship (e.g., regime changes) between the cross section of returns and systematic risk. Yet, we reduce the parameter space from a $N \times T$ loading matrix $\beta$ to the $p \times k$ loading matrix $\G_\b.$ Second, we incorporate additional data from the large number of asset characteristics to influence the factor model for returns through the loading matrix $\G_\b$. This addresses a challenge of migrating assets wherein an asset-specific but static $\b_i$ would not be able to capture an asset moving from, for example, a crypto asset earning low fees to one with high protocol fees. The classic way to handle this issue was to sort assets into portfolios of similar characteristics to form test assets, which then compresses the dimensionality of the cross-section. Thus, as discussed in kelly2019characteristics, this model specification skips ad hoc test asset formation to instead accommodate working directly with the high-dimensional system of individual assets. Finally, we assume exact row sparsity in $\G_\b$---precisely stated in the coming Assumption (ref)(ii)---a novel assumption to the literature; that is, only a small number of the $p$ asset characteristics determine the content of the factor loading, which matches empirical findings in cross-sectional asset pricing (kelly2019characteristics; babiak2021risk; bianchi2022dynamics).

\paragraph{Extended Model} In order to show the generality of this approach, we enrich the model---with one of several possible extensions---to address the common question in asset pricing research of whether an observable factor $g_{t+1} \in \R$ carries a risk premium: compensation for exposure to the risk factor holding constant exposure to all other sources of risk, i.e., variation with other factors.

In the context of asset pricing, a factor can be either tradable or nontradable. A tradable factor is a portfolio, that is, a convex combination of tradable asset returns. The risk premium is straightforward to calculate for tradable factors: it is the time series average excess return of the factor. However, many risk factors are nontradable, e.g., inflation expectations, consumption, liquidity, etc. Thus, we must estimate the risk premia of nontradable observable risk factors as the risk premia associated with their tradable portfolio.

Following a recent approach in the literature (giglio2021asset and giglio2021test), first, we assume the aforementioned model for returns is a function of the “true” latent factors $f_{t+1},$ \footnote{To be precise, by true latent factors, we mean we can consistently estimate the finite constant of the number of latent factors that span the cross section of returns.} and, second, we assume these true latent factors can be decomposed into the latent-factor risk premia $\g \in \R^k$ (i.e., unknown parameters of the long-run average excess return) and latent-factor innovations $v_{t+1} \in \R^k$ (i.e., mean zero risk factor random variable), that is, $f_{t+1} \coloneqq \g + v_{t+1}.$ Then, we specify the observable factor $g_{t+1}$ as potentially linearly correlated with the latent factors through

equation*[equation* omitted — 51 chars of source]

where $\h \in \R^{k}$ is an unknown parameter mapping the relation between the latent-factor innovations and the observable factors, and $\e^g_{t+1} \in \R$ is measurement error in $g_{t+1}.$ This specification allows the observable factors to be either simply some component(s) of the true latent factors (e.g., setting $\e^g_{t+1}$ to zero with $\h$ set to $(1,0,0,\dots,0$) or, more generally, some unknown linear function $\h$ of the true latent factors and thus still carry a risk premium. To recover the tradable portfolio representing the nontradable observable risk factor, we map $g_{t+1}$ through $\h$ onto the column space of the true latent factors.

Precisely, the risk premium of an observable factor---our target parameter in this extension---is defined to be the expected excess return of a portfolio with loading (i.e., beta) of 1 with respect to this factor in $g_{t+1}$ and zero loadings on all other factors; in this model, that parameter is $\g_g \coloneqq \h^\top \g.$

\paragraph{Goal} Under the novel asymptotic assumption for this setting of $p,N,T \rightarrow \infty$, we aim to develop estimation procedures for the latent loadings $\G_\b$ and factors $f_{t+1}, \ \forall t;$ in addition, we aim to estimate and conduct inference on $\g_g$ under the novel use of a dynamic latent-factor model and the aforementioned novel high-dimensional asset characteristics.

Estimation

\paragraph{Motivating DSLFM Estimation} The goal is to jointly estimate the loading matrix $\G_\b$ and latent factors $f_{t+1},$ which are not separately identifiable without further restrictions, to be discussed bai2013principal. However, to begin, given the model takes the form

equation*[equation* omitted — 72 chars of source]

where $\e_{i,t+1} = (\e^\b_{i,t})^\top f_{t+1} + \e^r_{i,t+1}$ is the composite idiosyncratic error, we observe that our setting is high-dimensional panel data where we project $\{ r_{i,t+1} \}_{i=1,t=1}^{i=N,t=T}$ onto the column space of $\{ z_{i,t} \}_{i=1}^{N}$ for each time period to estimate each $p$ dimensional time-varying vector $\{ \widehat{\G_\b f_{t+1} } \}_{t=1}^{T}$ where we have to address $p \sim \max(N,T)$ or $p >> \max(N,T).$

Thus, if we utilize the objective function to minimize over the $p-$dimensional choice vector $\G_\b f_{t+1}$ the mean-squared error of $\sum_{i} (r_{i,t+1} - z_{i,t}^\top \G_\b f_{t+1})^2$, we will not only have a noisily estimated design matrix when $p \sim \max(N,T),$ (or, at worst, a nonsingular design matrix when $p>\max(N,T)$) but also a non-convex objective given the interaction between minimization arguments $f_{t+1}$ and $\G_\b$. This rules out implementing low-dimensional (in $p$) methods.

One potential solution would be to introduce sparsity in $z_{i,t},$ given that empirical estimates show, ex post, few covariates contribute the vast majority of the signal, as stated earlier.\footnote{The online implementation of IPCA kelly2019characteristics does indeed offer an $\ell_1$ regularization to their MSE objective, which is not discussed in the econometric theory of kelly2020instrumented.} This would amount to adding a regularization parameter to the aforementioned objective to combat the curse of dimensionality from $z_{i,t}$ \footnote{One of several ways to interpret the curse of dimensionality is that as the number of covariates increases linearly, the volume of the parameter space to estimate grows nonlinearly; hence the density of the data falls.} However, although potentially helpful for minimizing MSE by decreasing the variance of the estimator, this regularization introduces a bias in estimation, which will lead to invalid asymptotic inference for asset pricing tests, defeating a goal of this work.

We therefore adapt for our purpose the Double Selection Lasso estimator introduced by belloni2014inference. The key insight from their work was to introduce an orthogonality wherein, assuming $\G_\b$ is row sparse, the regularization bias from the LASSO first-stage estimation does not pass through to the target parameter of interest when conducting inference.\footnote{The ideas developed in the Double Selection Lasso paper for inference in partially linear models with high-dimensional controls were the basis for the generalization of this idea in the DML procedures as developed in chernozhukov2018double. It would likely be closer to the empirical reality to maintain a nonparametric loading fan2016projected. As aforementioned, there is thus a natural extension of the work herein to use DML wherein the target variable $c_{t,j}$ is linear while the controls are nonparametrically estimated via a machine learning method, which would require the development of a Neyman Orthogonal score for this panel data setting, perhaps in a similar fashion to semenova2021debiased. However, the econometric theory is unknown for inference in the nonparametric setting. Moreover, given the high-dimension asset characteristics, using a conditional independence assumption to obtain a causal parameter may be the most fruitful path toward a causal factor model, a major area of future work (e.g., lopez2022causal). We explore the out-of-sample predictive ability of a non-parametric model in the last section of this manuscript.} First, we estimate for each time period $t$ and each characteristic $j$ the scalar $\G_{\b,j}^\top f_{t+1}$ using DSL; then stacking these estimates into a $T \times p$ matrix, we use PCA to obtain separate estimates for latent loadings $\wh{\G}_\b$ and factors $\{ \wh{f}_{t+1} \}_{t=1}^T$; and, finally, soft-threshold $\wh{\G}_\b$ to set the majority of the rows to zero given the assumption of sparsity.

We rewrite the DSLFM model and introduce a first-stage:

equation[equation omitted — 322 chars of source]

where $c_{t+1,j}$ refers to the $j \in \{1,\dots,p \}$ component of $c_{t+1} := \G_{\b} f_{t+1}$ while $-j$ refers to the remaining $p-1$ elements of the vector; $\d_{t,j} \in \R^{(p-1)}$ is an unknown, possibly time-varying, parameter; and, $\e^z_{i,t,j}$ is an unknown scalar random error. $c_{t+1,j}$ is an asset return when its $j-$th characteristic is set to 1 and all other characteristics are set to zero, less its idiosyncratic return $\e_{i,t+1}.$

There are several ways to interpret and justify the first-stage equation as discussed in belloni2014inference. Intuitively, the procedure does not rely on perfect model selection for valid inference as instead we not only recover controls $z_{i,t,-j}$ in the second-stage equation for their pricing ability in the cross section of returns but also recover controls with high correlation to our target variable $z_{i,t,j}.$ From a theoretical perspective, the first-stage equation accounts for potential omitted variable bias if one estimated only the second-stage equation. That is, the set of potentially relevant asset covariates is enormous (chen2020deep and bianchi2022dynamics), and thus a researcher may be motivated to select their preferred subset to ameliorate the curse of dimensionality, which could introduce model selection mistakes. Moreover, it is known LASSO can have poor finite sample model selection performance chernozhukov2015valid. Thus, selecting covariates with only the second-stage equation could fail to include relevant controls.

\paragraph{DSLFM Estimation Procedure} Our estimation procedure for $\{ f_{t+1} \}_{t=1}^T$ and $\G_\b$ has three steps: Double Selection Lasso, PCA, and soft-threshold step.

enumerate• DSL: To estimate $\hat{c}_{t+1,j},$ run $T \times p$ Double Selection Lasso cross-sectional regressions. \footnote{To set the penalty parameter in the LASSO implementations, one can follow the analytic method developed for heteroskedastic, non-Gaussian settings detailed in Appendix A, Algorithm 1 of belloni2014inference. For a more modern approach, one can use the bootstrap-after-cross-validation method of chetverikov2021analytic. In practice, we use cross validation.} \begin{itemize} • Run LASSO of $\{r_{i,t+1} \}_{i=1}^N$ on $\{ z_{i,t}\}_{i=1}^N$ for $\hat{c}_{t+1,j}$ and $\hat{c}_{t+1,-j}.$ \begin{itemize} • Let $\hat{I}_1$ denote the nonzero elements of $\hat{c}_{t+1,-j}.$ \end{itemize} • Run LASSO $\{ z_{i,t,j} \}_{i=1}^N$ on $\{ z_{i,t,-j} \}_{i=1}^N$ for $\hat{\d}_{t,j}.$ \begin{itemize} • Let $\hat{I}_2$ denote the nonzero elements of $\hat{\d}_{t,j}.$ \end{itemize} • Define the set $\hat{I} = \hat{I}_1 \cup \hat{I}_2 \cup \hat{I}_3$ where $\hat{I}_3$ is the set of controls in $z_{i,t,-j}$ not included in the first two LASSOs that the econometrician thinks are important for ensuring robustness, termed the amelioration set. • Run OLS of $\{r_{i,t+1} \}_{i=1}^N$ on $\{z_{i,t,j}, \tilde{z}_{i,t,-j}\}_{i=1}^N$ where $\tilde{z}_{i,t,-j}$ includes only elements of $z_{i,t,-j}$ in $\hat{I}$. That is, \end{itemize} \begin{equation*} (\hat{c}_{t+1,j}, \hat{c}_{t+1,-j}) \coloneqq \arg \min_{c_{j}, c_{-j}} \{ \mathbb{E}_N [(r_{i,t+1} - z_{i,t,j}c_{t+1,j} - z_{i,t,-j}^\top c_{t+1,-j})] \ : \ c_{t+1,-j,l} = 0, \forall l \notin \hat{I} \}. \end{equation*} • PCA: To estimate $\G_\b$ and $f_{t+1},$ run PCA on $\hat{C}=\wh{F} \wh{\G}_\b^\top$---a $T \times p$ matrix---to decompose it into $p \times k$ and $T \times k$ matrices $\hat{\G}_\b$ and $\wh{F}.$ • Soft-threshold: Given the assumed exact row sparsity of $\G_\b,$ we set to zero all rows of $\hat{\G}_\b$ whose row $\ell_1$ norm is below a cross-validated hyperparameter $\lambda$. That is, \begin{equation} \check{\G}_{\b,j} = \left( \norm{ \widehat{\G}_{\b,j} }_1 - \lambda \right)_+ sign (\norm{ \widehat{\G}_{\b,j} }_1), \ \ j \in \{1,\dots,p\}. \end{equation}

This does require running $T \times p$ versions of the cross-sectional Double Selection Lasso regressions, which can be in the thousands in empirical settings; however, these regressions are all computationally light and can be run in parallel. Moreover, this allows for unbalanced panels as each cross-section can have a different number of assets. \footnote{In empirical practice, we find the entire estimation procedure is on the order of ten minutes.} Additionally, these cross-sectional regressions, followed by estimations with the entire panel, mirror the effort of the most commonly used estimation procedure in the factor model setting, namely two-pass Fama-MacBeth regressions.

The high dimensionality of the PCA procedure, given we have a $p \times T$ matrix, is adapted from the existing literature using $N \times T$ matrices bai2003inferential. The estimated factor matrix $\widehat{F}$ is the product of $\sqrt{T}$ and the eigenvectors corresponding to the k largest eigenvalues of the $T \times T$ matrix $(Tp)^{-1} \widehat{C} \widehat{C}^\top.$ The estimated factors are normalized such that $\widehat{F}^\top \widehat{F} = I_{k \times k},$ a standard approach. The estimated loadings are $\widehat{\G}_\b = T^{-1} \widehat{C}^\top \widehat{F}.$ We thus see the main challenge in deriving the large-sample theory will be handling the estimation error in using $\widehat{C}$ instead of the unobserved $C.$

The final soft-thresholding step (ref) in our estimation procedure exploits the sparsity in $\G_\b$ to not only reduce the dimensionality of the characteristic space $s$ ($ < < p$) but also remove noise from the characteristics that have low signal-to-noise ratios. belloni2018high discuss the general theoretical properties of the soft-threshold estimator with theory-based hyperparameter selection, and its close relation, the better known LASSO and Dantzig selector estimators. We further discuss constraints and selection of the hyperparameter in Appendix (ref).

\paragraph{Estimating the Risk Premium of an Observable Factor} Under the richer setting that includes the observable factor $g_{t+1},$ our model has an additional specification and moment conditions.

equation[equation omitted — 342 chars of source]

Our goal is to estimate and conduct inference on $\g_g \coloneqq \h^\top \g.$ Given the latent factors are unobserved, we cannot jointly estimate $v_{t+1}$ and $\h$ without further restrictions. We would have to invoke one of the three classic identification approaches of bai2013principal; however, by using the key rotation invariance result of giglio2021asset, we can estimate the latent factors up to an invertible rotation matrix $H \in \R^{k \times k}$ and still maintain identification of our target parameter $\g_g.$ That is, both of the underlying parameters will be identified up to this rotation matrix: $\h^\top = \h_0^\top H^{-1}$ and $\g = H \g_0.$ Thus, the target parameter is rotationally invariant to $H:$ $\g_g = \h_0^\top H^{-1} H \g_0 = \h^\top \g.$

For our estimation procedure, we replace the first PCA step of giglio2021asset with our above procedure---augmented to use return innovations---to estimate the latent loadings $\check{\G}_\b$ and factor innovations $\hat{v}_{t+1}$ for all $t.$ We then proceed with the latter two steps of the authors' procedure to obtain our target estimator.

enumerate• To estimate latent-factor risk premia $\hat{\g},$ run cross-sectional OLS of average returns $\bar{r} \in \R^N$ on averaged estimated latent-factor loadings $\bar{\hat{\b}} = \bar{Z}^\top \hat{\G}_\b \in \R^{N \times k}.$ • To estimate latent to observable factor mapping $\hat{\h},$ run a time series OLS regression of $\{ g_{t+1} \}_{t=1}^T$ on factor innovations $\hat{V} \in \R^{T \times k}.$

We can thus form our estimator of the risk premium for the observable factors $g_{t+1}$ by combining these estimators into $\hat{\g}_g = \hat{\h}^\top \hat{\g}.$

This procedure extends the estimation in giglio2021test to dynamic loadings and high-dimensional asset characteristics, while inheriting the rotation invariance and the specification consistent with two-pass estimators in this literature. Again, simply applying IPCA instead of PCA in the first step of giglio2021asset would not be feasible with $p > \max{N,T}$ or would yield biased inference if an $\ell_1$ penalty were simply added to the IPCA objective. The cross-sectional OLS of average returns on the estimated latent-factor loadings is the standard second step in two-pass Fama-MacBeth regressions, which could be replaced with generalized least squares or weighted least squares to explore asymptotic efficiency gains. The final time series regression is critical to translate the uninterpretable risk premia of latent factors to those of factors proposed by economic theory. Moreover, this procedure handles omitted variable bias which we now briefly discuss.

To illustrate, assume we have a scalar observable factor $g_{t+1},$ which is the first component of a two-dimensional latent-factor innovation vector: $v_{t+1} = (g_{t+1}, v_{2,t+1})^\top$ (i.e. $\h = (1,0)$). The vector-version of our model is thus

equation*[equation* omitted — 109 chars of source]

Using the standard Fama-MacBeth two-pass regressions fama1973risk will produce bias in estimating $\g_g$ if $v_{2,t+1}$ is omitted. The first step of a time series regression of asset excess returns on $g_{t+1}$ will give a biased estimate of $\hat{\b}_1$ as long as $v_{2,t+1}$ is correlated with both $g_{t+1}$ and $r_{t+1},$ per the standard OVB term: the covariance between between the outcome and the excluded regressor times the covariance between the included and excluded regressor, up to scale. Moreover, in the second step of a cross-sectional regression of average returns on estimated loadings, a second omitted variable bias is introduced if the loading of the omitted factor $\hat{\b}_2$ is correlated with both $\hat{\b}_1$ and $\bar{r}_{t+1}.$

Estimating the latent factors via the DSLFM procedure resolves this issue of omitting a potentially relevant factor given one can utilize a consistent estimator of the true number of latent factors, which we assume spans the true factor space.\footnote{The DSLFM could be further extended to estimate the zero-beta rate (i.e., alpha) using a very similar approach to that discussed in Online appendix I.2 of giglio2021asset.}

Asymptotic Theory

In this section, we present the asymptotic results for consistent estimation of the latent factors and loadings and the large sample distribution of the nontradable observable factor risk premium estimator under the assumed setting discussed in Section (ref) and using estimation procedures discussed in Section (ref) for models (ref) and (ref). We first provide the regularity conditions sufficient for the validity of the estimation and inference results. For clarity of exposition, we focus on motivating the assumptions and interpreting the results, while theoretical details and mathematical proofs are provided in Appendix (ref).

Throughout, let $\norm{ A } = [ tr(A^\top A)]^{1/2}$ denote the Frobenius norm of matrix A or $\norm{ x } = \left( \sum_i x_i^2\right)^{1/2}$ for the $\ell_2$ norm of a vector $x$. Let $\norm{ x }_0 $ and $\norm{ x }_1$ be the usual $\ell_0$ and $\ell_1$ norms, respectively. All limits are simultaneous where we will restrict the rates among $p,T,N$, to allow $p\rightarrow\infty,$ as discussed below.

Regularity Conditions

\paragraph{Consistent Estimators for the Latent-Factor Model} The following assumptions enable the consistent estimation of the factors $\{ f_{t+1} \}_{t=1}^T$ and the loadings $\G_\b$. Let $f_{t+1}^0$ and $\G_\b^0$ be the true factors and loadings such that $f_{t+1} = H f_{t+1}^0$ and $\G_\b = \G_\b^0 H^{-1}$ where $H$ is an unobserved $k \times k$ invertible rotation matrix.

In regard to identification, our results do not require the identification of the true factors $f_{t+1}^0$ and loadings $\G_\b^0$ but rather simply factors (loadings) that span the true factors (loadings) up to the rotation matrix $H.$ bai2013principal show identification results for PCA under three different sets of assumptions to pin down the $k \times k$ elements in $H,$ which requires pinning down the covariance matrices of the factor loadings and factors to be diagonal matrices or identity matrices to provide $k(k-1)/2 + k(k+1)/2 = k^2$ restrictions. The researcher can choose which asymptotic covariance matrix to restrict. As we will discuss, we will additionally not need these identification restrictions for the observable factor risk premia given the aforementioned rotation invariance result of the target parameter.

assumption[Consistency of DSL] \begin{enumerate} • Bounded Characteristic Portfolios: For a finite absolute constant $M$ and $\forall t,j,$ $\left|c_{t+1,j}\right| = \left|\G_{\b,j}^\top f_{t+1} \right| < M.$ • Sparse Loading: Loading matrix $\G_\b$ admits an exactly sparse form. That is, for $\exists s \in \mathbb{N}_+, i.e. p > s \geq 1$, $\G_\b$ has at most $s$ nonzero rows: $\sum_{j=1}^p \mathbb{1} \Bigl\{ \norm{\G_{\b,j} }_1 > 0 \Bigr\} \leq s.$ \end{enumerate}

These are two critical assumptions for DSL consistency with the additional standard and technical DSL assumptions in Appendix (ref). Assumption (ref)(i) converts the bounded target parameter, in the traditional DSL context, to the DSLFM context where we require realizations of $c_{t+1,j}$ to be finite-sample bounded by a constant that does not depend on $p,T,N.$ This imposes a bound on the return of characteristic portfolios, that is, the return of a portfolio with characteristic $j$ set to 1 and all other characteristics set to 0. We could instead assume returns are bounded random variable to impose Assumption (ref)(i).

Assumption (ref)(ii) is the key LASSO assumption that the parameter on the control regressors admits an exactly sparse form, which follows from our assumption such that $\forall t,j, \ \norm{c_{t+1,-j} }_0 = \norm{ \G_{\b,-j} f_{t+1} }_0 \leq s$. This sparsity of the loading matrix is supported empirically in asset pricing given the relevance of only a small number of asset characteristics, which we corroborate in our empirical setting. We have thus adapted the classic LASSO sparsity assumption to the empirical reality of cross-sectional asset pricing using high-dimensional asset characteristics. Exact sparsity could be relaxed to approximate sparsity with a similar but alternative high-dimensional econometrics toolkit.

We next turn to assumptions for consistently estimating the latent factors and loadings. The focus in our work is controlling the estimation error between the infeasible eigendecomposition of $(Tp)^{-1} CC^\top$ and the feasible eigendecomposition of $(Tp)^{-1} \wh{C}\wh{C}^\top,$ given we do not observe $C=F \G_\b^\top$ but instead estimate each element via DSL and then eigendecompose using standard PCA estimators as discussed in Section (ref).

assumption[Consistency of Latent-Factor Model] \begin{enumerate} • Factors: $\mathbb{E} \norm{ f_{t+1}^{0} }^4 \leq M < \infty$ and $T^{-1} \sum_t f_{t+1}^0 f_{t+1}^{0\top} \rightarrow_p \Sigma_f$ for some $k \times k$ positive definite matrix $\Sigma_f.$ • Factor Loadings: $\forall j, \ \norm{ \Gamma_{\beta,j} } \leq M < \infty$ and $\norm{ \Gamma_\beta^\top \Gamma_\beta / p - \Sigma_\Gamma } \rightarrow 0$ for some $k \times k$ positive definite matrix $\Sigma_\Gamma.$ • Nonzero and distinct eigenvalues: from the infeasible eigendecomposition, the $k$ largest eigenvalues $\l_i$ for $i \in \{1, \dots, k\}$ are bounded away from zero. Moreover, the $k$ largest infeasible eigenvalues are distinct, that is, $$\min_{i:i\neq \kappa} |\lambda_\kappa - \lambda_i | > 0.$$ \end{enumerate}

Assumptions (ref)(i)-(ii) are standard for factor models where the literature is styled after Assumptions A, B, and C of bai2003inferential. Assumption (ref)(i) does not impose i.i.d. factors, as in the classical factor analysis literature, but instead imposes the factors are stationary, strong mixing, and satisfy moment conditions. Assumption (ref)(ii) ensures each latent factor contributes to the second moment of $c_{t+1};$ that is, it imposes all factors are pervasive and excludes weak factors. See giglio2021test for adjustments for weak factors. The PCA estimation herein does not require the Assumption C of bai2003inferential given our target matrix $C = F \G_\b^\top$ is without an error term; we instead are controlling cross-sectional and temporal dependence using the moment conditions of DSL given in model (ref) and more technical assumptions in Appendix (ref).

Assumption (ref)(iii) assumes the $k-$largest eigenvalues from the infeasible and feasible eigendecompositions remain nonzero asymptotically. In finite sample, these are real and nonzero eigenvalues given we are taking the eigendecomposition of a rank $k$ symmetric matrix. It is reasonable to assume we have distinct eigenvalues given, for this not to hold, there would have to be two or more dimensions in the $k-$largest of the $T \times T$ matrix $CC^\top$ that have precisely the same variability.

\paragraph{Estimating Number of Factors} Given the focus of this work is on the consistency of the main estimators and the asymptotic distribution of the risk premium estimator, we assume $k=k^0$ is known.\footnote{The asymptotic distribution of the risk premium estimator is unaffected when the number of factors is estimated because

equation*[equation* omitted — 431 chars of source]

}

assumption[Consistent Estimator for Number of True Factors] For $\bar{k} > k^0,$ let $\wh{k} \coloneqq \arg \min_{0\leq k \leq \bar{k}} IC(k)$ where \begin{equation*} \begin{aligned} IC(k) &\coloneqq \log (V(k)) + k \left( \frac{p + T}{pT}\right) \log \left( \frac{pT}{p+T} \right) \\ V(k) &\coloneqq \min_{\G_\b,F} \left(pT\right)^{-1} \sum_{j,T} \left(c_{j,t+t} - \G_{\b,j}^\top f_{t+1} \right)^2. \end{aligned} \end{equation*} Assume $\wh{k} \rightarrow_p k^0$ without further restriction on the growth rates among $p,T,N$ and $k=k^0$ is known.

Assumptions (ref) and (ref) can be shown to be sufficient for consistently estimating, with the above Information Criterion, the number of true factors $k^0$ using the results of Appendix (ref) as in bai2002determining and bai2003inferential. Although providing this assumption to show the estimator to be studied in simulation, we are instead choosing to impose Assumption (ref) in the asymptotic to focus on the main results of this work. Note that commonly used model selection criteria (e.g., AIC or BIC) will not yield consistent estimators, hence the specification above using the contribution of bai2002determining.

\paragraph{Inference on Nontradable Observable Factor Risk Premia} The final assumptions are needed to derive the limiting distribution of the risk premium estimator.

assumption[Inference] There exists a generic absolute constant $M<\infty$ such that for all $p,T,N:$ \begin{enumerate} • Bounded idiosyncratic errors: $\E [ ( \sum_t \e_{i,t+1} )^2 ] \leq TM.$ • Bounded scaled factor innovations: $\E [ ( \sum_t z_{i,t}^\top \G_\b^0 v_{t+1}^0 )^2 ] \leq sTM.$ • Bounded measurement errors: $\E [ ( \epsilon^g_{t+1} )^2 ] \leq M.$ • Convergence of characteristics: $\frac{1}{NT} \sum_i \sum_{t'} \E[z_{i,t,j}] z_{i,t',j'} \rightarrow_p \mathcal{Z}_{t,j,j'}$ uniformly over $t,j,j'$ for $j,j' \in \{1,2,\dots,p\}$ and a nonstochastic finite constant $\mathcal{Z}_{t,j,j'} \in \R.$ • CLT: As $T\rightarrow \infty,$ the following joint central limit theorem holds: $$ \frac{\sqrt{T}}{T} \sum_t \left( \begin{matrix} v_{t+1}^0 \epsilon_{t+1}^g \\ \Pi_t v_{t+1}^0 \end{matrix} \right) \xrightarrow{d} \mathcal{N} (0, \Phi)$$ where random matrix $\Pi_t \in \R^{k \times k}$ and nonstochastic matrix $\Phi \in \R^{2k \times 2k}$ are defined in Appendix (ref). \end{enumerate}

Assumption (ref)(i) bounds the second contemporaneous and cross-moments of the idiosyncratic errors, aligning with the time and cross-section dependence assumptions of bai2003inferential Assumption C. The assumption would hold if we assumed $\e_{i,t+1}$ are uncorrelated across $t$, which is a simplified yet plausible assumption given the low signal-to-noise environment of asset pricing. We have thus relaxed the temporal dependence to the specified rate $T$.

Assumption (ref)(ii) bounds the squared time series average of the factor innovations scaled by the factor loadings. In the static factor model context of giglio2021asset, this holds in large sample by a simple LLN argument given the static loadings are not a function of $t$ and the factor innovations are mean zero random variables. Thus, we are ensuring the $\G_\b^0$ selected columns of $Z_t$ keep the scaled $v_{t+1}^0$ sufficiently small.

Assumption (ref)(iii) bounds the second moment of the observable factor measurement errors for use in proving $\norm{ \e^g } = O_p (\sqrt{T}).$ It is not a stringent assumption because we are bounding a zero mean scalar random variable. This is nearly an identical assumption and usage to giglio2021asset Assumption A8.

Assumption (ref)(iv) provides a convergence result such that the squared first moment for two different characteristics averaged over time and across assets is a nonstochastic finite constant. This is a weaker assumption on the distribution of characteristics than the DSL moment conditions discussed in Appendix (ref).

Assumption (ref)(v) is the assumed central limit theorem for the $2k$ (low) dimensional mean zero random variable given the models' (ref) and (ref) moment assumptions, which is satisfied by various mixing processes. The second moments of the later $2k$ random variables are bound already in Assumptions (ref)(i)-(ii). We nevertheless directly assume the needed CLT. This extends for our inference result the assumed CLT at the same rate in Assumption F4 of bai2003inferential and the assumed CLT at the same rate in Assumption A11 of giglio2021asset. Note that although we have the same two mean zero random vectors, our factor innovations are scaled by $\Pi_t$ instead of a constant $1$ given the dynamic factor loadings of our model.

Theory Results

This section presents the three main theoretical results.

\paragraph{Consistent Estimators for the Latent-Factor Model} We present the first two results showing the consistency of the latent-factor model estimators.

proposition[Consistency of Latent Factors] Under the model (ref), Assumptions (ref), (ref), (ref), and DSL Assumptions in Appendix (ref).2 where $T,N,p \rightarrow \infty,$ then for all $t$ the latent-factor estimator described above has the property that $$ \wh{f}_{t+1} - H^\top f^0_{t+1} = O_p \left( \sqrt{\frac{s \log(Tp)}{N}} \right).$$

The proof is in Appendix (ref). This result establishes the convergence rate of the latent factor estimator in a dynamic latent-factor model with high-dimensional characteristics. If the factor loadings were static and known, $\beta_i^0$ for all $i,$ then, $f_{t+1}^0$ would be estimated via a cross-sectional least squares with a convergence rate of $\sqrt{N}.$ bai2003inferential establishes in Theorem 1(ii), under $N,T\rightarrow \infty$ for a static latent-factor model, the foundational result of a convergence rate of $\min \left( \sqrt{N}, T \right)$ for the consistency of the latent-factor estimator for the rotated true factors $H^\top f_{t+1}^0.$ Incorporating dynamic loadings parameterized by high dimensional characteristics comes at the cost of slowing the rate to $\sqrt{pN / s \log(Tp)},$ which is nevertheless still reasonable for typical values of $p,T,N.$ Additional standard DSL rates, which are less restrictive, are in Appendix (ref).

Our rate is primarily driven by the $\sqrt{N / \log(Tp) }$ rate uniform consistency over $t$ and $j$ of the DSL estimation error $| \wh{c}_{t+1,j} - c_{t+1},j |$ as shown in Lemma A1 in Appendix (ref). Given the model for $C = F^0 \G_\b^{0 \top}$ contains no error, the eigendecomposition of the unobserved $C$ is exact for $F^0 H$ as shown in Lemma A6; and, thus, the estimation error from using $\wh{C}$ instead of $C$ drives this first main result. The assumed sparsity in $\G_\b^0$ does improve the rate with the $p / s$ ratio.

It is worth reiterating that under our setting $F^0$ and $\G^0_\b$ are not separately identifiable, hence the $k \times k$ invertible matrix transformation $H$ appears in each asymptotic result. Similarly, $\wh{F} \wh{\G}_\b^\top$ is an estimator of the identifiable, rotation invariant common component $C,$ which is corroborated by simulation results. Moreover, in many cases knowing $F^0 H$ is equivalent to knowing $F^0;$ for example, the regressor $F^0$ will give the same predicted values as using $F^0 H$ as a regressor given they have the same column space.

proposition[Consistency of Latent-Factor Loadings] Under the model (ref), Assumptions (ref), (ref), (ref), and DSL Assumptions in Appendix (ref).2 where $T,N,p \rightarrow \infty,$ then the latent loading estimator described above has the property that $$\left( \check{\G}_{\b} - \G_{\b}^0 H^{-1} \right) = O_p \left( \sqrt{\frac{s \log(Tp)}{N}} \right).$$

The proof is in Appendix (ref). This result establishes the convergence rate of the latent loading estimator in a dynamic latent-factor model with high-dimensional characteristics. When the factors, $f_{t+1}^0,$ for all $t,$ are observable, static loadings $\beta_i^0$ can be estimated by a time series regression with a convergence rate of $\sqrt{T}.$ bai2003inferential establishes in Theorem 2(ii), under $N,T\rightarrow \infty$ for a static latent-factor model, the foundational result of a convergence rate of $\min \left( N, \sqrt{T} \right)$ for the consistency of the latent-factor loading estimator $\wh{\b}_i$ for the rotated true factor loadings $H^{-1} \b_i^0.$ Incorporating dynamic loadings parameterized by high dimensional characteristics comes at the cost of slowing the rate to $\sqrt{N / s \log(Tp)},$ which is nevertheless still reasonable for typical values of $p,T,N.$

The rate follows similar reasoning to that of the latent-factor estimator in Proposition 1. However, here we are presenting the rate of the final soft-threshold estimator---derived using recent results in high dimensional econometrics belloni2018high---wherein we use the uniform estimation error between the eigendecomposition of $\wh{C}$ for $\wh{\b}_{\b,j}$ and the infeasible loading $\wt{\b}_{\b,j}$ from decomposing the unobserved $C,$ which eliminates $p$ in our rate from Proposition 1. The $\sqrt{N / log (Tp) }$ is similarly driven by the uniform consistency over $t$ and $j$ of the DSL estimation error $| \wh{c}_{t+1,j} - c_{t+1},j |,$ which is the key result used to establish these consistency propositions along with typical high dimensional random matrix theory (e.g., Davis Kahan Theorem, Weyl Inequality, and recent tools in high dimensional econometric theory found in belloni2018high).

\paragraph{Inference on Nontradable Observable Factor Risk Premium} Finally, we present the asymptotic normality of the nontradable observable factor risk premium estimator.

theorem[Normality of Observable Factor Risk Premium] Under the models (ref) and (ref); Assumptions (ref), (ref), (ref), (ref); DSL Assumptions in Appendix (ref).2; and if $T s^2 \log (Tp) / N \rightarrow 0,$ then as $T,N,p \rightarrow \infty$ the estimator $\hat{\g}_{g}$ obeys $$\sqrt{T} \frac{(\hat{\g}_{g} - \g_{g})}{\s_{g}} \xrightarrow{d} \Nc (0,1),$$ where $\s_{g}$ is defined in Appendix (ref).

The proof is in Appendix (ref). This result establishes $\sqrt{T}$ asymptotic normality of the nontradable observable factor risk premium estimator from giglio2021asset extended to the setting of dynamic factor loadings with high dimensionality characteristics. At a high level, our proof follows a similar approach yielding, as seen in the proof of Theorem 1, the same two asymptotically nonnegligible terms as in giglio2021asset. The first term arises from the time-series regression of the observed factor on the latent loadings where again the latent loading estimation error is higher order. The latter term in the $2k$ random vector in Assumption (ref)(v) arises from the cross-sectional regression of averaged asset excess returns on averaged factor loadings, where the factor loading estimation error and idiosyncratic error term are higher order. Although giglio2021asset have this same second term, ours is more complicated given the dynamic loadings, which necessitates the convergence Assumption (ref)(iv). A direct application of the delta method on the sum of these two terms yields the result in Theorem 1.

The crucial rate assumption is $T s^2 \log (Tp) / N \rightarrow 0,$ which controls the estimation error for the unobserved averaged latent-factor loadings $T^{-1} \sum_t \b_{i,t}.$ This is similar to bai2003inferential and giglio2021asset, which require $T/N \rightarrow 0$ to use the estimated factors or loadings as generated regressors. However, we have slowed the rate again due to the high-dimensionality in $p.$ This is our slowest required rate.

Given the rotation invariance of the target parameter $\g_g^0$, the unobserved rotation matrix $H$ does not appear in the asymptotic distribution, in contrast to the consistency results. In finite sample, we corroborate this result in the to-be-discussed simulations. In a similar vein, it warrants noting the asymptotic efficiency loss due to not observing the factors and loadings could be large when $N$ is relatively small. Also, we have assumed we directly observe the number of tree factors $k^0,$ which would require estimation in practice and thus likely contribute estimation error to affect finite sample performance.

In simulation, we use the plug-in estimator for $\s_g$, which has satisfactory finite-sample coverage properties. However, one can establish a consistent variance estimator using a newey1987hypothesis style plug-in estimator of the asymptotic variance $\s_g$ with lag corrections to account for temporal dependence as in giglio2021asset Section IV Part E.

Asset Pricing Tests

In this section we develop three tests central to our empirical analysis. The first uses the asymptotic normality of the observable factor risk premium for a statistical test of nonzero risk compensation. The second statistic informs the incremental significance of any specific asset characteristic. The third and final discusses how we empirically measure whether the DSLFM contributes predictive signal above and beyond a random walk.

\paragraph{Testing Nontradable Observable Factor Risk Premium} An empirical application of the DSLFM model will address whether a nontradable observable factor, namely, inflation, carries a nonzero risk premium in the crypto asset class. The target parameter $\g_g$ captures the risk premium of the (inflation) factor-mimicking portfolio within the crypto asset class as recovered by the estimated dynamic latent-factor model. We are interested not only in the sign of the parameter, but also, in practical settings, in whether a confidence interval suggests a risk premium of economic significance.

We test the hypothesis $H_0: \g_g = 0 \ \ \ \text{vs.} \ \ \ H_1: \g_g \neq 0$ using the risk premium estimation procedure described in Section (ref) with a plug-in variance estimator $\wh{\s}_g$ for $\s_g.$ Given the asymptotic normality of Theorem 1, we form a confidence interval $$\g_g \in \left[\wh{\g}_g - c(1-\alpha/2) \wh{\s}_g, \wh{\g}_g + c(1-\alpha/2) \wh{\s}_g \right]$$ where the critical value $c(1-\alpha/2)$ is the $1-\alpha/2$ quantile of a $N(0,1)$ distribution for the researcher-specified level of the test $\alpha.$ We find in the coming simulation acceptable finite sample coverage for this confidence interval.

\paragraph{Testing Characteristic Significance} The large-sample distribution of the latent loading is unknown given the DSL regularization. Even inference in simple cross-sectional LASSO is complicated lee2016exact. Instead, we develop a simple bootstrap procedure to infer whether a specific characteristic significantly contributes to loading $\G_{\b}$. We leave for subsequent research developing the supporting theory of this bootstrap procedure or develop the asymptotic distribution of a consistent latent loading estimator in this setting. We test the hypotheses $\forall j$

equation*[equation* omitted — 198 chars of source]

That is, we ask whether characteristic $j$ contributes to the factor loading through $k \times 1$ mapping vector $\G_{\b,j}.$ This allows the researcher, using a large number of characteristics, to systematically ask what characteristics contribute to the latent-factor model, instead of an ad hoc selection. We thus set the entire $k \times 1$ vector to zero so the characteristic contributes to predicting the variation in returns through none of the $k$ factors.

Our procedure is to test the alternative hypothesis model, with the unconstrained characteristic $j,$ and then form the test statistic $$W_{\G,j} = \G_{\b,j}^\top \G_{\b,j}.$$ Using bootstrapped standard errors, we assess whether this test $W_{\G,j}$ statistic is statistically distinguishable from zero.

\paragraph{Testing Out-of-Sample Performance} To study the out-of-sample pricing ability of the DSLFM, we use the “predictive $R^2$” defined as

$$\textrm{Predictive} \ R^2 = \frac{\sum_{i,t} \left(r_{i,t+1} - z_{i,t}^\top \Check{\G}_\b \wh{\lambda}_t \right)^2}{\sum_{i,t} r_{i,t+1}^2}$$ where $\wh{\lambda}$ is the moving average of the estimated factors in previous time periods over a cross-validated window size. This measure captures whether the model forecasts realized returns better than a random walk; or, said differently, it represents the fraction of realized return variation explained by the model's description of expected returns through exposure to systematic risk. This specification allows the model's estimated conditional expected returns to be driven not just by the dynamic factor loadings, estimated using high-dimensional asset characteristics, but also by time-varying risk prices $\lambda_t.$

Simulations

This section presents a brief study of the finite-sample performance of the dynamic latent-factor model estimators and coverage properties of the inference procedure using Monte Carlo simulations. To summarize, we find the estimation errors for factors and loadings are comparable to IPCA and the Three Pass estimator of giglio2021asset in low-dimensional settings while superior in high-dimensional settings. This holds even in rather small samples with low signal to noise ratios, reflecting the empirical reality of cross-sectional asset pricing. Moreover, we find estimation errors and coverage properties for the observable-factor risk premium to be comparable to giglio2021asset in low-dimensional settings while superior in high-dimensional settings. We now present the design, followed by the results.

Simulation Design

First, we describe the data-generating process for given $N,T,k$ where we follow the finite sample simulation study of IPCA kelly2020instrumented. That is, the DGP is favorable to IPCA. We calibrated the simulated data to parameter estimates from IPCA fit to our weekly panel of crypto asset data using all sixty three asset characteristics.

Latent factors $f_{t+1}$ are simulated from a $VAR(1)$ model employing normal innovations that was fit to the estimated IPCA factors. Asset characteristics are simulated from a $p$ variable panel $VAR(1)$ model with normal innovations, which was fit to the demeaned empirical weekly panel of randomly selected, without replacement, $p$ asset characteristics. For each asset, we set the means of the characteristics to a bootstrap sample from the empirical distribution of time series asset characteristic means. The idiosyncratic error $\e_{i,t+1}$ is simulated from an i.i.d. normal distribution whose variance is calibrated such that the population $R^2$ of the model is approximately 20%, matching the empirically estimated value from fitting IPCA. The measurement error $\e^g_{t+1}$ is simulated in a simple fashion but the $R^2=1-\E [e^g]/\E[g]$ is calibrated to approximately 40%. $\eta = (1,0,\dots,0)$ and the loadings $\G_\b$ are set to the empirically estimated values where $p-s$ rows are set to zero where $s=p/10.$ Finally, observable factors and returns are generated according to models (ref) and (ref).

The simulation studies results across $S=200$ Monte Carlo draws. Hyperparameters are fixed at $N=500, \ T=100,$ and $k=3.$ To compare the performance of estimators under low-dimensional and high-dimensional characteristics, results are generated for $p=10$ and $p=50.$ We report results for a variety of estimators, including latent loadings $\G_\b$, latent factors $F,$ average factor loadings $\bar{\b},$ latent matrix $C=F \G_\b^\top,$ and observable factor risk premium $\g_g.$

The benchmark estimation and inference procedures are IPCA and the three-pass estimator of giglio2021asset, given DSLFM's basis on these foundational models.\footnote{Many thanks to Matthias Buechner and Leland Bybee for the IPCA implementation \url{https://github.com/bkelly-lab/ipca}.} We focus on two comparisons: first, the estimation error of theoretically consistent latent loading $\G_\b$ and latent factor $\{ f_{t+1} \}_{t=1}^T$ IPCA and DSLFM estimators; and, second, coverage properties of the observable factor risk premium estimator. We do study estimation errors for additional estimands as relevant (e.g., the three-pass estimator does not estimate latent loadings $\G_\b$ nor latent matrix $C$, while IPCA does not have an observable factor risk premium estimation nor inference procedure).

Simulation Results

Table (ref) reports results. In the low-dimensional setting of $N=500,T=100,p=10,s=1,$ we find DSLFM to obtain smaller estimation errors for $\G_\b$ as compared to IPCA; however, IPCA has an order of magnitude lower estimation errors for $F.$ DSLFM's outperformance for the latent loading is driven by taking on higher bias yet substantially lower variance of the estimator; this is obtained from soft-thresholding many of the rows. In comparing other auxiliary estimands, DSLFM obtains lower estimation error for the time-series averaged factor loadings $\bar{\b}$ although higher error for the latent matrix $C = F \G_\b^\top.$ We attribute these results to IPCA fitting data simulated to match the fits of an empirically estimated IPCA model, yet we employ an exact row-sparsity structure in the true latent loadings $\G_\b.$

DSLFM slightly under-covers the 90% and 95% confidence intervals in the low-dimensional setting. The three-pass estimator obtains similar estimation error for the target parameter of the observable factor risk premium, but has finite-sample intervals that slightly over-cover.

Moving to the high-dimensional setting of $p=50,$ we find DSLFM to again obtain smaller estimation errors for latent loadings, yet now the estimation errors for the latent factors are of the same order as IPCA. In both cases, DSLFM takes on bias from its regularization methods, although DSLFM's latent factor estimator's variance is still higher than IPCA. DSLFM is now an order of magnitude improvement for the average factor loadings and a factor of two improvement for the latent matrix $C.$ Finally, coverage proprieties of the risk premium estimand for both the three-pass estimator and DSLFM estimators are degraded under high-dimensionality.

We hope to add even higher dimensional results with hyperparameters closer to the empirical values. Do note, given the DSLFM's large sample theory, it is constrained by $N > T,$ which although is the case for the panel of crypto asset returns, this is not the case for all cross-sectional asset pricing settings. Moreover, the DSLFM's performance was boosted by the assumed exact sparsity in the latent loadings; we hope to add results for approximate sparsity, which is likely closer to the empirical reality. Nevertheless, the DSLFM performs well as compared to state-of-the-art benchmark methods, especially under the setting for which it was developed: high-dimensional asset characteristics.

Empirical Applications

In this section, we first establish the pricing ability of a variety of benchmark factor models and compare these results to those of the DSLFM. With the DSLFM, we then utilize the asset pricing tests developed in Section (ref) to elucidate the drivers of returns and conduct inference for crypto's inflation risk premium.

See XYZ for the description of the data, which is the same panel we use in this paper. Tables (ref) and (ref) report descriptive statistics for the panel's dependent variable, asset excess returns over the subsequent week, and the set of sixty three asset characteristics.

\paragraph{Multivariate Observable Factor Models} We now turn to the out of sample performance of low dimensional factor models in estimating risk premia of crypto asset excess returns, beginning with multivariate observable factor models. Instead of selecting a small number of observable factors based on individual univariate performance, we instead form, using the sixty three characteristics, all combinations of one-, two-, and three-factor models to select the best model of each size based on its performance in combined period of the second half of 2021 and first half of 2022.

In detail, to select multivariate factor models, we perform the following procedure. First, we form strictly time varying risk factors using each of the sixty three asset characteristics as the top minus bottom value-weighted quintile portfolio excess returns. We thus do not normalize characteristics. To form predicted asset returns, we next estimate each asset's static factor loading as the contemporaneous in sample time series regression of excess returns on factor(s); estimate risk factor conditional means as the time series average of the factor in sample; and, predict the asset's return for the next week using the dot product. We use this procedure to fit models in 2018 through the first half of 2021 and, with an expanding window, predict week by week for the combined validation period of the second half of 2021 and the first half of 2022. For all one-factor models, the sixty three choose two two-factor models, and the sixty three choose three three-factor models, we then select the best model of each size based on the predictive $R^2$ for the fifty two weeks in the validation period. Throughout, we have to reform the panel each month for the relevant included assets. To compare to the literature, we also formed the liu2022common Fama-French style three-factor model using their CMKT, CSMB, and CMOM risk factors.\footnote{Code to replicate can be found at \url{https://github.com/adambaybutt/crypto_asset_pricing/blob/main/code/10a-low_dim_fm_multi.ipynb}.}

Table (ref) presents results for the test period, i.e., the second half of 2022. The best multivariate observable factor models were size; illiquidity and size; and, size, one month momentum, and three month volatility. Interestingly, the model selection process incorporated size in all three models. The predictive $R^2$ for all three models and the benchmark model were all negative, performing worse in MSE pricing ability than a random walk. However, although statistically insignificant, all three multivariate observable factor models had economically significant weekly time series average excess return spreads of about 1% with associated Sharpe ratios of 1.31-1.72, which all beat the benchmark model with a return spread of 0.5% and a Sharpe ratio of 0.65. All four associated alphas show these long-short strategy returns are meaningfully uncorrelated. The illiquidity and size two-factor model achieved the highest Sharpe of 1.72, which is suggestive evidence of a low number of factors being optimal. Although the strategies replicated out of sample, we should note these results are again before transaction costs and are for a very small test period.

\paragraph{Static Latent-Factor Model} We next study the out of sample performance of a static latent-factor model in estimating risk premia of crypto asset excess returns. We are not only interested in how learning the factors from the data changes the out of sample pricing ability, but also in developing benchmarks for the high-dimensional dynamic latent-factor model. We use the classic approach of PCA bai2003inferential to estimate one- through five-latent factor models.

In detail, for each month in the test period (i.e., the second half of 2022), we form factor(s) using the matrix of contemporaneous excess asset returns for the relevant included assets; for each asset, we run a time series predictive regression of its weekly excess returns on the factor(s) to obtain its factor loading(s); and, we use each asset's factor loadings and the PCA-estimated factors to generate predicted returns for the out of sample month.\footnote{Code to replicate can be found at \url{https://github.com/adambaybutt/crypto_asset_pricing/blob/main/code/10b-low_dim_fm_pca.ipynb}.}

Table (ref) reports results for the test period, i.e., the second half of 2022. The predictive $R^2$ for all five latent-factor models were all negative, performing worse in MSE pricing ability than a random walk. Although all five models had positive return spreads, only three of the five models had Sharpe ratios---0.78, 1.24, and 1.34---in the range of the observable factor models; however, there was not a clear pattern across the number of latent factors. This is suggestive evidence of latent factors not offering a clear benefit over the multivariate observable factor models. Their out-of-sample performance yielding uniformly positive Sharpe, generated without using asset characteristics, maintains a benchmark for the richer models.

\paragraph{Dynamic Latent-Factor Model with Low-Dimensional Characteristics} We close our study of the out of sample performance of low dimensional factor models in estimating risk premia of crypto asset excess returns. We investigate the performance of IPCA, a dynamic latent-factor model where the number of asset characteristics must be smaller than the number of assets and time periods.\footnote{We are grateful to Matthias Buechner and Leland Bybee for the IPCA implementation at \url{https://github.com/bkelly-lab/ipca}.} Given the small number of assets in the panel (i.e., there are less than two dozen for the majority of the weeks), we have to, outside of IPCA, select features from the sixty three asset characteristics. We chose to just use the characteristics listed in Table (ref).\footnote{We realized after doing this that this biases IPCA favorably as the univariate results used the IPCA test period, which is one of several reasons that we have not even formed 2023 data to repeat our out-of-sample exercises in fresh and larger data.} We again reform the panel each month and normalize period-by-period features to linearly spaced on $[0,1]$. Finally, we do not specify a constant to allow for mispricing effects but rather explain variation in expected returns using exposure to common latent risk factors. \footnote{Code to replicate our results can be found at \url{https://github.com/adambaybutt/crypto_asset_pricing/blob/main/code/10c-low_dim_fm_ipca.ipynb}.}

Table (ref) reports results for the test period, i.e., the second half of 2022. Predictive $R^2$ are positive except for a five factor specification with the maximum predictive $R^2$ of 0.18% for the three-factor model. All five models have economically significant weekly excess return spreads, from 1.5% to 3.1%, and associated annualized Sharpe ratios, from 2.07 to 4.07. The one- through four-factor models have statistically significant time-series average weekly excess return spreads for the zero-investment long-short strategies of, respectively, 2.8%, 2.9%, 3.1%, and 2.4%. The alphas remain statistically and economically significant with little return lost to the market. Remarkably, there are only two quintile portfolios out of twenty five that break monotonicity. Sharpe ratios nearly monotonically decline with the number of factors; the three factor model edges out the two factor model by a difference of 0.13. Although again these results do not account for transaction costs and the test period is short, the dynamic latent-factor model estimated with IPCA dominates the static factor models.

\paragraph{Dynamic Latent-Factor Model with High-Dimensional Characteristics} We study three questions using the DSLFM. We first compare the out of sample predictability in the same test period to that of the previous factor models, in addition to understanding the characteristics driving returns and to estimating the inflation risk premium in the crypto asset class. The setting is the same weekly panel, reformed each month with the relevant assets, with asset characteristics normalized to 0 to 1.\footnote{Code to replicate can be found at \url{https://github.com/adambaybutt/crypto_asset_pricing/blob/main/code/12-dslfm.ipynb}.}

The DSLFM estimation procedure is outlined in Section (ref) and the test procedures are defined in Section (ref). There are several hyperparameters for the statistician to chose, for which we are empirically motivated to use cross validation. Specifically, we generate predicted returns week by week in the same validation period of the second half of 2021 through the first half of 2022 by using an expanding window training data set using all previous weeks from the start of the panel. For models with one to five latent-factors, we cross validate the relevant hyperparameters, including the soft thresholding hyperparameter, the lasso penalty parameter, and the number of trailing weeks to average fitted latent factors over to form predicted factors. We present results for models with the best predictive $R^2$ in the validation period.

Table (ref) presents out of sample test period results for the DSLFM. Only one out of the five models had a positive predictive $R^2.$ Eight out of ten models, when forming equal-weighted and value-weighted portfolios within quintile, had economically significant time series average returns for the long-short strategies, which were maintained when studying the associated alphas. The equal-weighted portfolios had in all but one case, the five latent-factor specification, superior Sharpe ratios, driven by both improved return spreads and lower volatility.

Given the small panel, the poor pricing ability of the DSLFM with more factors is perhaps not surprising given it could be over parameterized and noisily estimated. The factor loading matrix grows by $p$ with each additional specified latent factor. Nevertheless, the large majority of the long-short strategies obtained economically significant Sharpe ratios and associated alphas, representing an improvement over the observable factor models and PCA. However, IPCA outperformed, which is perhaps driven by a meaningful signal to noise ratio improvement through the feature selection done before fitting IPCA. We will explore this in 2023 data, which will yield much more data given the wide cross-section relative to preceding years.

Table (ref) presents bootstrapped results on asset characteristic importance to understand the drivers of returns. Exchange inflows and outflows were the two statistically significant characteristics with point estimates on their importance more than an order of magnitude larger than the next characteristic. Again, we observe the importance of onchain data. This empirically supports approximate sparsity as a reasonable assumption, given the fast decay with a long tail on the importance of these characteristics. Our theory, although it would accommodate approximate, assumes exact sparsity for simplicity. Interestingly, none of the statistically significant univariate factor strategies were significant in the DSLFM characteristic importance. However, all six were at least in the top half, and, in practice, the importance of studying exchange flows is well known. In future work, we will compare these results to the importance measures available in IPCA.

Finally, to demonstrate the extensibility of the DSLFM, we conduct inference on an observable factor risk premium, namely testing for a nonzero premium for ten year expected inflation risk within the crypto asset class. This has been a long-standing research question to understand the relationship between crypto's returns and inflation. Early proponents of Bitcoin and other cryptocurrencies framed these as an outside option or hedge against traditional fiat currencies. To study this question, we use our extended model with one factor and the associated estimation procedure---as described in Section (ref) in how we extend giglio2021asset---to recover the 10-year expected inflation mimicking portfolio and measure its risk premium.

The inflation risk premium was estimated to be a statistically significant 1.4 bps with a standard error of 0.0097 bps. This translates to a 7.3% annual excess return, suggestive of positive compensation for investors holding an inflation-hedged crypto portfolio, ceteris paribus. The result corroborates similar findings using more simple methods, detailed in Section (ref), with a dynamic latent factor model with superior pricing ability. One could attribute this to several aspects of the DSLFM, for example, it allows for regime changes with time-varying loadings, it incorporates rich structure with the full asset characteristics, among other reasons. There are limitations however, including the slow asymptotic rate with this inference procedure, as discussed in Section (ref), which is exacerbated by the small cross-section in our setting. Thus, although significant, we should interpret this result as suggestive and seek replication.

References