EconBase
← Back to paper

Overparametrized models with posterior drift

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

52,042 characters · 14 sections · 73 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Overparametrized models with posterior drift

abstractThis paper investigates the impact of posterior drift on out-of-sample forecasting accuracy in overparametrized machine learning models. We document the loss in performance when the loadings of the data generating process change between the training and testing samples. This matters crucially in settings in which regime changes are likely to occur, for instance, in financial markets. Applied to equity premium forecasting, our results underline the sensitivity of a market timing strategy to sub-periods and to the bandwidth parameters that control the complexity of the model. For the average investor, we find that focusing on holding periods of 15 years can generate very heterogeneous returns, especially for small bandwidths. Large bandwidths yield much more consistent outcomes, but are far less appealing from a risk-adjusted return standpoint. All in all, our findings tend to recommend cautiousness when resorting to large linear models for stock market predictions.

Introduction

Recently, the literature in machine learning (ML) has investigated the notion of “double descent”, whereby linear overparametrized models (i.e., with more parameters than observations) can have surprising out-of-sample benefits.\footnote{This is sometimes also referred to as “benign overfitting” (bartlett2020benign), or “grokking” (power2022grokking, varma2023explaining), though these notions do not necessarily perfectly overlap.} As is customary in most machine learning contributions, including on double descent, a key assumption to derive analytical results is the invariance in distributions between the training and testing phases: the data generating process (DGP) is assumed to remain the same once the model has been calibrated. Unfortunately, this may not be the case in practice, especially in financial markets which can be subject to sudden regime changes. If we write $P_{y,X}=P_{y | X}P_X$ for the joint law of $X$ and $y$, we see that there are potentially two drivers for the change of distribution in the DGP: the conditional one, $P_{y | X}$, and the unconditional one, $P_X$. When the law of $X$ changes, we refer to covariate shift, whereas when the link between $X$ and $y$ evolves, we will talk of posterior drift (sometimes also referred to as concept drift, see, e.g. gama2014survey and lu2018learning). Interestingly, the notion of posterior drift has also witnessed a surge in interest lately, in particular from the ML community (see, e.g., maity2024linear, hu2025transfer and wang2025conformal).

The goal of the present paper is to investigate how posterior drift can be detrimental to out-of-sample forecasting accuracy in the context of overparametrized models.\footnote{The topic of covariate shift is already partly covered in Section 5 of hastie2022surprises (via what they call the misspecified case), so that the room for novel contributions on the matter is limited.} Our practical motivation originates in the seminal paper by kelly2023virtue which argues that the performance of aggregate market timing with linear models increases with the number of parameters. Since complexity is the ratio between the number of parameters and the sample size, the authors conclude that complexity is virtuous for equity premium prediction. However, several contributions have since then challenged this point of view, all from different angles. For instance, berk2023comment underlines the lack of financial grounding (equilibrium consistency), whereas cartea2025limited argue that the effect of noise is underestimated in kelly2023virtue. Indeed, as the number of factors (i.e., predictors) increases, the amount of noise may also increase, thereby impairing the models' accuracy. nagel2025seemingly demonstrates that the random Fourier feature (RFF) trick used to artificially increase the number of predictors in fact collapses to a low-complexity kernel ridge regression. Resorting to RFFs is also problematic because they often require in-sample rescaling that violates the shift-invariance property required in theoretical results (fallahgoul2025high). Finally, buncic2025simplified points to another important shortcoming: adding a constant in the set of predictors reverses the pattern and performance then decreases with complexity. As a consequence, it turns out that Sharpe ratios obtained with low dimensional predictors are much higher than those reported in kelly2023virtue for large levels of complexity.

In focusing on the perils of regime changes, we shed light on another explanation of the ambiguous role of overparametrization in the forecasting efficiency of linear models. Indeed, we contend that the efficacy of complex models in equity premium prediction can be severely jeopardized by changing economic environments. In practice, the relationships that are inferred during the estimation phase may change due to unpredictable shocks, and, in this case, the out-of-sample precision of predictions can be substantially attenuated. Another representation of such phenomena can be made through the lens of the signal-to-noise ratio. In the presence of posterior drift, the information from predictors wanes, and, when signals are too weak, recent results indicate that ridgeless estimators perform worse than models that ignore the data completely (see Theorem 2 in shen2024can and Corollary 1 in fallahgoul2025high for instance).

Our contributions are twofold. First, in Section (ref), we extend the misspecified isotropic results of hastie2022surprises to the case of non-i.i.d. data and derive the expected return of a market-timing strategy when posterior drift is taken into account. Second, in Section (ref), we corroborate our theoretical findings with Monte-Carlo simulations. Third, in Section (ref), we empirically confirm the intuitions. We first provide in Section (ref) new evidence of time-variation in links between predictors and aggregate returns (betas). We then replicate in Section (ref) the study of kelly2023virtue with some subtleties. We test their procedure over several sub-periods of 15 years (the horizon for a representative investor) and across a range of bandwidth parameters, which we define below.

Our results indicate strong discrepancies in both dimensions (periods and bandwidths). Plainly speaking, this means that even with the bandwidth chosen in kelly2023virtue, the performance can be either well above (e.g., +7% monthly in 2005-2019) or well below (+0.5% in 1975-1989) than the one that we report for the full sample (+4%). It is possible to reduce this uncertainty by increasing the bandwidth value, but in this case, the average return decreases invariably towards zero, making complexity much less appealing.

In sum, while it is likely that sophistication can bring value in forecasting models, our findings suggest that analysts who rely on overparametrized models should pay particular attention to the stability and sensitivity of performance in their backtests. \\

Notations. $n,p \in \mathbb{N}_{>0}$ and $q \in \mathbb{N}_{\geq 0}$ are large dimensional parameters, possibly tending to infinity. $n$ is the number of observations and $p+q$ the number of independent variables in linear models. We use $C$ and $D$ (resp. $\tau$) for arbitrary large (resp. small) positive constants. Let $\langle \cdot, \cdot \rangle$ stand for the scalar product, i.e. for any vectors $u$, $v \in \mathbb{R}^p$, $\langle u, v \rangle = u'v$. Denote by $\| \cdot \| = \| \cdot \|_2$ the Euclidean norm of vectors. For any matrix $A \in \mathbb{R}^{p \times p}$ and any vector $v \in \mathbb{R}^p$, we denote by $\| v \|_A = \sqrt{v'Av}$, a reweighted version of the Euclidean norm. Similarly, for any matrix $A \in \mathbb{R}^{p \times p}$ and any vector $v,u \in \mathbb{R}^p$, define $\langle u, v \rangle_A = u'Av$. We use $\| \cdot \|_{\text{op}}$ to denote the Euclidean $2$-norm of a matrix, namely, for any matrix $A \in \mathbb{R}^{n \times p}$, $\| A \|_{\text{op}}$ is the largest singular value of $A$.

Theoretical grounding

Setup: misspecification and posterior drift

This paper pertains to large models which, as in most of the recent literature, are linear, see, e.g., bartlett2020benign, and kelly2023virtue. hastie2022surprises do cover some non-linear cases but the corresponding require very lengthy proofs which is why the present paper focuses on extending linear combinations of features. Henceforth, our modelling framework follows that of hastie2022surprises closely. We assume a data generating process (DGP) of the form

align[align omitted — 185 chars of source]

where the $n$ random draws are independent. $P_{x,w}$ is a distribution on $\mathbb{R}^{p+q}$ with $\mathbb{E}[x_i]=0$ and $\mathbb{E}[w_i]=0$ and

align*[align* omitted — 179 chars of source]

$P_e$ is a distribution on $\mathbb{R}$ with zero mean and variance $\sigma^2$. In our empirical application, as in kelly2023virtue, $y$ will be the equity premium and $x$ a set of macroeconomic predictors that are expected to have some predictive power over the premium.

The agent has access to the $x_i$ but not to the $w_i$, and hence only sees an incomplete picture, this is why her model is labeled as misspecified. This corresponds to the assumptions in Section 5 of hastie2022surprises and, to some extent, to Section IV of kelly2023virtue. We write $(y_i, x_i)\in \mathbb{R} \times \mathbb{R}^p $ for one observation in the training data, and $(y,X)$ for the corresponding aggregated vector (in $\mathbb{R}^n$) and matrix (in $\mathbb{R}^{n\times p}$).

As in hastie2022surprises, we define out-of-sample prediction risk of the well specified model, i.e. the model in which all the covariates are observed, as

equation[equation omitted — 155 chars of source]

where $x_0$ is a random draw of $P_x$ that is independent from the training data. In the equation, the bias is $B_X(\hat{\beta}, \beta) =\|\mathbb{E}[\hat{\beta}|X]-\beta \|_\Sigma^2$, and the variance is $V_X(\hat{\beta}, \beta)=\text{Tr}[\mathbb{C}\text{ov}(\hat{\beta}|X)\Sigma]$.

The main departure from existing results is that we now discriminate between the $(\beta',\theta')'$ vector that is used for the two data generating processes in Equation (ref): one vector for the training data ($(\beta_{\text{is}}',\theta_{\text{is}}')'$) and another one for the testing data ($(\beta_{\text{oos}}',\theta_{\text{oos}}')'$). In doing so, we include posterior drift in the model: we postulate a change in the link between $y$ and $(x',w')'$. Empirical support for this assumption will be provided in Section (ref) below.

The estimator $\hat{\beta}$ is evaluated with the in-sample (is) or out-of-sample (oos) generated data:

equation[equation omitted — 181 chars of source]

where $\alpha > 0$ and $u \in \{\text{is}, \text{oos} \}$. Crucially, the above estimator relies on shrinkage with intensity $z_n = \alpha / n$ towards the identity matrix. In the remainder of the paper, we omit the dependence of the regularization parameter on the sample size $n$ and set $z := z_n$. If $n<p$ and $z \to 0$, the pseudo-inverse is taken for the inverse matrix. We have also set $\hat{\Sigma}=X'X/n$. Below, we explicit the terms for the quadratic error (ref) that is made when using the estimator $\hat{\beta}_{\text{is}}$ (from the data generated with $\beta_{\text{is}}$) when compared to the changed (or realized) $\beta_{\text{oos}}$. We are interested in this subsection in results for ridgeless estimators (when $z \to 0^+$).

propositionUnder model (ref) and with $z \to 0$, the prediction risk of the misspecified model is \begin{align*} R_X^m(\hat{\beta}_{is},\beta_{oos}, \theta_{oos}) & = \mathbb{E}[(x_0'(\hat{\beta}_{is}-\beta_{oos})-w_0'\theta_{oos})^2|X] \\ & = R_X(\hat{\beta}_{is},\beta_{oos}) + M(\theta_{oos}), \end{align*} where $M(\theta_{oos}) =\mathbb{E}[(w_0'\theta_{oos}-\mathbb{E}[w_0'\theta_{oos}|x_0])^2]$ is the \textit{misspecification bias}, i.e., the signal embedded in the unobserved features (following the notation of hastie2022surprises). Moreover, in the isotropic case $\Sigma = I$, \begin{align} \underset{n,p \rightarrow \infty}{\lim}R_X^m(\hat{\beta}_{\text{is}},\beta_{\text{oos}},\theta_{oos}) &= \left\{ \begin{array}{c l} (\sigma^2 + \| \theta_{is} \|_2^2) \frac{c}{1-c} + \|{\beta}_{\text{oos}}-{\beta}_{\text{is}}\|_2^2 + \| \theta_{oos} \|_2^2 & \quad p/n \rightarrow c < 1 \\ (\sigma^2 + \| \theta_{is} \|_2^2)(c-1)^{-1} +c^{-1}\|\beta_{\text{is}}-\beta_{\text{oos}} \|^2_2 &\quad p/n \rightarrow c > 1 \\+ (1-c^{-1}) \|\beta_{\text{oos}} \|_2^2 + \| \theta_{oos} \|_2^2 . \end{array} \right. \end{align}

The proof of the proposition is postponed to Appendix (ref). The limiting expressions generalize Theorem 4 in hastie2022surprises by adding the posterior drift component. Plainly, both expressions for the risk are increasing in the misspecification risk $\| \theta_{oos} \|_2^2$ due to the unobserved features and in the drift risk $\|\beta_{\text{is}}-\beta_{\text{oos}} \|_2^2$ that comes from the change in loadings.

Portfolio performance under posterior drift

Proposition (ref) quantifies the impact of the presence of posterior drift on the risk of out-of-sample prediction of linear (ridge) regressions in the over-parameterized setting. We now focus on the impact on the financial performance of a market timing strategy. As in kelly2023virtue, we analyze the returns of the strategy which consists of adjusting the position in an asset according to its estimated expected returns. Formally, the strategy's return can be written as

equation[equation omitted — 79 chars of source]

where $\hat{\pi}_t(z) = \sum_{k=1}^p \hat{\beta}_{\text{is}}^{(k)} (z)x_t^{(k)}$ is the adjusted position in the asset and $r_{t+1}$ is its realized monthly return (the superscript $^{(s)}$ in the notation stands for strategy). The estimated loadings in this case can be shrunk as in the general case in Equation (ref).

For brevity, we directly consider the misspecified framework described in (ref) and without loss of generality, we assume $\sigma^2 = 1$. We provide similar results within the well-specified framework in Appendix (ref) and (ref).

The following assumption establishes the linear model from which the covariates are considered to be drawn.

assumpThe covariates have the form $x_i = \Sigma^{1/2} z_i$ with $z_i = (z_{i1}, z_{i2}, \dots, z_{ip})$ having independent and identically distributed entries. Moreover, for all $j \in \{1,\dots,p\}$, $\mathbb{E}[z_{ij}] = 0$, $\mathbb{E}[z_{ij}^2] = 1$ and $\mathbb{E}[|z_{ij}|^{4+a}] \leq C < \infty$ for $a > 0$.

The out-of-sample expected return of the timing strategy under posterior drift can be decomposed as follows.

propositionLet $z > 0$, $\xi = (e_{1}, e_{2}, \dots , e_{n})'$ with $(e_i)_i$ as in (ref) and Assumption ((ref)) hold. Furthermore, denote by $r^{(s,w)}_{t+1}(z)$ the return of the strategy if the model is well-specified, i.e. there are no unobserved features, $w_i$, in the DGP (ref). Then, \begin{align} \mathbb{E}\left[ r^{(s)}_{t+1}(z) \big| X \right] = \mathbb{E}\left[ r^{(s,w)}_{t+1}(z) \big| X \right] + \mathcal{M}(z) \end{align} where \begin{align*} \mathbb{E}\left[ r^{(s,w)}_{t+1}(z) \big| X \right] &= \beta_{oos}' \Sigma_x (zI + \hat{\Sigma}_x)^{-1} \big( \hat{\Sigma}_x \beta_{is} + \frac{1}{n} X' \xi \big) \\ \mathcal{M}(z) & = \big( \beta_{oos}' \Sigma_x + \theta_{oos}' \Sigma_{xw}' \big) (zI + \hat{\Sigma}_x)^{-1} \mathbb{E} \Big[\hat{\Sigma}_{xw}\Big| X \Big] \theta_{is} \\ & \quad + \theta_{oos}' \Sigma_{xw}' (zI + \hat{\Sigma}_x)^{-1} \big( \hat{\Sigma}_x \beta_{is} + \frac{1}{n} X' \xi \big) . \end{align*}

The proof is postponed to Section (ref) in the Appendix. The proposition shows that the out-of-sample expected return of the timing strategy under posterior drift crucially depends on the interactions between the observed and the unobserved covariates. These complex relations are encapsulated in the quantity $\mathcal{M}(z)$. Therefore, in order to provide more qualitative qualitative insights and intuition, we study different models of interest that allow to make this term more explicit.

In what follows, we study the simpler case of i.i.d. features with unit variance, which entails $\Sigma=I$.

We consider the following dimensional ratios.

assumpFor some $\tau > 0$, $\tau \leq c \phi \leq \tau^{-1}$, with $c = (p+q)/n$, $\phi = p/(p+q) \in (0,1]$.

This assumption is standard in statistical theory and characterizes the proportional regime in which our study takes place. The constant $c$ stands for the complexity of the true models while $\phi$ is the misspecification ratio. The product $c\phi$ is the complexity of the misspecified model.

The following proposition gives the limiting out-of-sample expected return of the timing strategy under posterior drift and i.i.d. features, $\mathcal{E}^{(d)}(z;c\phi)$. Moreover, we introduce an important benchmark, namely the limiting zero-drift expected return from Proposition 3 in kelly2023virtue:

equation[equation omitted — 245 chars of source]
propositionLet $z>0$ and assume $(x_i,w_i)'$ has independent and identically distributed entries with unit variance and finite moments of order $4+a$, for some $a > 0$. Then, \begin{equation} \mathbb{E}\left[ r^{(s)}_{t+1}(z) \big| X \right] \xrightarrow[\substack{\mathstrut n,p \rightarrow \infty \\ p/n \rightarrow c\phi }]{\mathds{P}} \quad f(z;c\phi) \langle \beta_{is}, \beta_{oos} \rangle = \mathcal{E}^{(d)}(z;c\phi), \end{equation} where $f(z;c\phi)$ and $f(c\phi):=\lim_{z \to 0} f(z;c\phi)$ are defined in Equations (ref) and (ref) in the Appendix.\footnote{Propositions (ref) and (ref) in the Appendix provide non-asymptotic bounds (under stronger conditions) for the approximation of the expected return. However, these bounds are additive in the errors and, therefore, not relative to the scale of the return. Hence, given the typical order of magnitude of financial returns, it would make sense to establish multiplicative bounds similar to those developed by cheng2022dimension, misiakiewicz2024non or defilippis2024dimension for the approximation of the test error of different types of ridge regression. We leave this avenue open for further research.} In particular, if $\| \beta_{is} \|^2 = \| \beta_{oos} \|^2$, then \begin{equation*} \mathbb{E}\left[ r^{(s)}_{t+1}(z) \big| X \right] \xrightarrow[\substack{\mathstrut n,p \rightarrow \infty \\ p/n \rightarrow c\phi }]{\mathds{P}} \quad \mathcal{E}^{(d)}(z;c\phi) = \mathcal{E}(z;c\phi) - \frac{1}{2} f(z;c\phi) \| \beta_{is} - \beta_{oos} \|^2, \end{equation*} Furthermore, $z \mapsto f(z;c\phi) \langle \beta_{is}, \beta_{oos} \rangle$ is monotone decreasing (resp. increasing) in $z$ when \\ $\langle \beta_{is}, \beta_{oos} \rangle \geq 0$ (resp. $\langle \beta_{is}, \beta_{oos} \rangle \leq 0$).

This proposition, the proof of which is located in Section (ref), calls for several remarks. First, the above expression generalizes the results of kelly2023virtue. Indeed, the above formulation allows to recover Figure 5 in their theoretical section when setting $c=10$ and $\beta = \beta_{is} = \beta_{oos}$ with $\beta_i = \beta_j$ for any $i,j \in \{1, \dots, p \}$ and such that $\| \beta \|^2 = 0.2$ when $q=0$. The pattern exhibited in kelly2023virtue is a special case when the scalar product $\langle \beta_{is}, \beta_{oos} \rangle $ is fixed to a particular value.

The shape in the r.h.s. of (ref) is a multiple of the scalar product $\langle \beta_{is}, \beta_{oos} \rangle $. Simply put, the expected performance is proportional to the alignment between the in-sample loadings and the out-of-sample ones. Given the fact that the mean of the estimated loadings is very close to zero (upon the Random Fourier transform), the scalar product can also be interpreted as the sample correlation between $\beta_{is}$ and $\beta_{oos}$. In particular, if the two vectors are orthogonal (with null correlation), then the average return is simply zero. This makes sense, as it is in this case impossible to extract any meaningful information from the training set.

The proportional form in (ref) is in fact very general and can also be obtained with a DGP that includes latent features, as in hastie2020surprises. The corresponding results and proofs are postponed to Appendix (ref).

remarkThe limiting expression in ((ref)) can be adapted to any specific form of $\beta_{oos}$. For instance, let $\beta_{oos}$ be a linear transformation of $\beta_{is}$, i.e. $\beta_{oos} = W \beta_{is}$ with $W \in \mathbb{R}^{p \times p}$ deterministic. If we denote by $(\xi_k)_{1 \leq k \leq p}$ the eigenvalues of $W$ and $(v_i)_{1 \leq k \leq p}$ their associated eigenvectors, then the (asymptotic) expected return of the strategy converges to \begin{align*} \mathcal{E}^{(d)}(z;c\phi) = h(z;c\phi,\beta_{is},W) \|\beta_{is}\|^2, \end{align*} where $h(z;c\phi,\beta_{is},W) := f(z;c\phi) \sum_{k=1}^p \xi_k \text{cosim}(\beta_{is}, v_k)^2$ and $ \text{cosim}(\beta_{is}, v_k) := \langle \beta_{is}, v_k \rangle / \|\beta_{is}\|$ is the cosine similarity between $\beta_{is}$ and $v_k$, which lies between $-1$ and $1$. In the Appendix, we also show that $f \in [0,1]$, thus it is straightforward that $\mathcal{E}^{(d)}(z;c\phi) \geq 0$ if $W$ is positive semidefinite. Otherwise, if $W$ is not positive semidefinite, the picture is more nuanced and the expected return can be negative.

More importantly, when the covariates are isotropic, the misspecification term $\mathcal{M}(z)$ in (ref) vanishes, and the out-of-sample expected return behaves as if the model was well-specified with a true complexity equal to $c\phi = p/n$. This result leads us to draw the following remark which establishes the conditions under which the presence of posterior drift deteriorates the expected returns of the market-timing strategy.

remarkLet $z>0$ and assume that $(x_i,w_i)'$ has independent and identically distributed entries with unit variance and finite moments of order $4+a$, for some $a > 0$. Then, \begin{equation*} \mathcal{E}^{(d)}(z;c\phi) < \mathcal{E}(z;c\phi) \end{equation*} if and only if \begin{equation} \langle (\beta_{oos} - \beta_{is}), \beta_{is} \rangle < 0, \end{equation} or alternatively, if and only if, \begin{equation} \| \beta_{oos} \|^2 - \| \beta_{is} \|^2 < \| \beta_{is} - \beta_{oos} \|^2. \end{equation}

Consequently, the loss of financial performance depends on the scalar product between the vector of in-sample regression coefficients, $\beta_{is}$, and a residual vector, $\beta_{oos} - \beta_{is}$, arising from the change in DGP. In particular, when the amount of signal contained in the first $p$ features is preserved or is less important out-of-sample than in-sample, i.e. $\| \beta_{oos} \|^2 \leq \| \beta_{is} \|^2$, Remark (ref) implies that the presence of posterior drift not only increases the out-of-sample prediction risk (as documented in our first proposition), but, more interestingly from a financial standpoint, it also systematically deteriorates the returns of market timing strategies based on these predictions.

Simulations

The proofs of Propositions (ref) and (ref) are somewhat cumbersome. This section is dedicated to the verification of these theoretical results via Monte-Carlo simulations. To this end, we generate data as per the DGP of Equation (ref). Let $p+q=300$. We draw $n=100$ samples of covariates from a $300$-dimensional standard multivariate normal distribution, i.e. $(x_i,w_i) \sim N(0,\Sigma)$ with $\Sigma = I_{300}$. The error term is a white Gaussian noise, viz. $e_i \sim N(0,1)$.

The in-sample and out-of-sample true parameters vectors for the observed data, $\beta_{is}$ and $\beta_{oos}$ respectively, are equally distributed on the eigenvectors of $\Sigma$, namely, on the canonical basis vectors.\footnote{In Appendix (ref) , we perform the same analysis considering true parameters vectors concentrated on some directions of $\Sigma$.} $\beta_{is}$ has unit signal, that is $\| \beta_{is} \|^2 = 1$, while we consider a sequence of $(\beta_{oos,k})_k$ with decreasing signal. In particular, $\beta_{is} = p^{-1/2}(1,1,\dots,1,1)'$ and $\beta_{oos,k} = k p^{-1/2}(1,1,\dots,1,1)'$ with $k \in \{ 0.2, 0.4, 0.6, 0.8, 1\}$, which yields $\| \beta_{oos,k} \|^2 = k^2 \leq 1$ and $\langle \beta_{is}, \beta_{oos} \rangle = k$.

We analogously generate the true parameters vectors for the unobserved data, $\theta_{is}$ and $\theta_{oos}$, setting only a different number of coefficients, viz. $q$ instead of $p$. Thus, the signal of the model in-sample is always equal to $2$ while it is set to $2k^2$ out-of-sample, for $k \in \{ 0.2, 0.4, 0.6, 0.8, 1\}$. Finally, we simulate the performance of two market-timing strategies based on ridge estimators with shrinkage intensity $z \in \{0.01,0.1\}$.

figure[figure omitted — 658 chars of source]

As expected from Proposition (ref), the strategy's expected return is linear in the scalar product of $\beta_{is}$ and $\beta_{oos}$. This is verified in Figure (ref). This pattern arises whether the model is misspecified and underparameterized (Fig. (ref), left), misspecified and overparameterized (Fig. (ref), center) or well-specified (Fig. (ref), right). Note that the rightmost point on each plot corresponds to the expected return in the no-drift scenario.

We defer the comparison of simulated and theoretical expected returns of the market-timing strategy under strong regularization ($z \in \{10,100\}$) to Appendix (ref). And, once again, these experiments confirm the claims stated in Proposition (ref).

Empirical evidence

Data

We rely on the same data as kelly2023virtue, which comes from goyal2023comprehensive and is available on \href{https://sites.google.com/view/agoyal145\string#Amit Goyal's website}{Amit Goyal's website}. We use the CRSP value-weighted returns minus the T-bill rate as dependent variable. The predictors consist of the 14 macro-economic indicators used in kelly2023virtue, to which we add lagged returns.

In Section (ref), we will compute betas from predictive regressions:

equation[equation omitted — 92 chars of source]

where the $x_t^{(k)}$ are all $p$ predictors. We estimate $\hat{\beta}$ every month on rolling windows of 15 years (180 points), which is our reference horizon. This generates time-series of loadings $\hat{\beta}_t$ that characterize the predictive link between the predictors and the returns from the CRSP portfolio $r_{t+1}$. These series are plotted in Figure (ref) below and therein we have included lagged returns up to order 6.

figure[figure omitted — 686 chars of source]

Time-varying betas

There is considerable evidence supporting the time-varying nature of the coefficients in asset pricing models. Such effects have been documented on the performance of predictive regressions (dangl2012predictive, coulombe2023time, farmer2023pockets), in factor models (ghysels1998stable, bollerslev2023optimal) and for risk premia estimation in particular (gagliardini2016time, bakalli2023penalized).

Empirically, this stylized fact can be revealed by change point detection. Such techniques seek to uncover sharp changes in the distribution of ordered observations (time series). They can take two forms, parametric and non-parametric. Parametric change point detection constrains the model to learn a particular functional form, i.e., observations between the change points are assumed to be drawn from a specific family of distribution (see Frick2014MultCPD, roy2016cpdMarkov, Enikeeva2019hdCpd). This assumption is relaxed in the case of non-parametric change point detection and we refer to aminikhanghahi2017survey for a complete survey on change point detection applied to time series. Recent contributions evaluate the dissimilarity of distributions using kernel distances (garreau2018consistent, arlot2019kernel) or rank (lung2015homogeneity).

We choose to rely on the multivariate nonparametric multiple change point detection method recently proposed by londschien2023RfDetection and that builds upon the use of classifiers. Notably, this will allow us to study betas in a joint manner: the entire time series structure of the coefficients is used to detect change points. Specifically, our analysis of time-variation is carried out with random forests via the changeforest algorithm\footnote{The repository is available at \href{https://github.com/mlondschien/changeforest.git\string\#https://github.com/mlondschien/changeforest.git}{https://github.com/mlondschien/changeforest.git}} whose technicalities we briefly outline below.

The idea behind the algorithm is to determine recursive binary splits at the points maximizing the nonparametric classifier log-likelihood ratio. This is shown in Figure (ref).\footnote{We stick to the default hyperparameters of the algorithm with two exceptions. First, we set the maximum tree depth to 2 since, by default, $p^{1/2} = 4$ features are selected for each tree. Second, we arbitrarily choose a minimal relative segment length of 0.125 so the results appear sparser. By setting a smaller minimum relative length, the algorithm detects an abundance of change points.} Therein, we have fed the time-series of the betas (the $\hat{\beta}_t^{(p)}$ from Figure (ref)) to changeforest.\footnote{We exclude the variables "lty", "dp", "dy" and "tms" from the regressions as they introduced too much multicollinearity, yielding some estimates to be overly positive while others compensating with large negative values.} The top panel of the graph shows that the most relevant split occurs in December 1982. It is the one that maximizes the gain of the classifier when deciding between two regimes (two classes). The second horizontal panel show the following two splits, in November 1963 and January 2002. Lastly, the bottom panel provides three new splits. Indeed, we see that the first and last periods are not split in two because the corresponding tests were not conclusive. This procedure tests if the split significantly improves the gain of the classifier. When the gain is too marginal, the algorithm stops, as for the pruning of simple decision trees. The final result clearly identifies several break points in the time-series of loadings, which justifies our assumption that changes occur between the training (in-sample) phase and the testing (out-of-sample) stage.

figure[figure omitted — 458 chars of source]

Protocol

We now closely follow the empirical protocol of kelly2023virtue and we keep the notations therein. The idea is to model aggregate returns as a linear function of a large number of predictors which are created from a smaller set of macro-economic indicators, namely those analyzed in welch2008comprehensive. To expand the predictor space, kelly2023virtue resort to Random Fourier Features (RFFs). More precisely, if $G_t$ is the $15\times 1 $ vector of the original predictors, we will use

equation[equation omitted — 169 chars of source]

as our independent variables in the predictive regression $r_{t+1}=S_{i,t}'\beta_i+e_{t+1}$. Importantly, $i$ is the index of a random draw of the stochastic weights $w_i$, which are i.i.d. Because the predictors will be random, it is important to cancel the noise by averaging across many independent draws of $w_i$ - usually several hundreds. Moreover, we are interested in the overparametrized regime, hence all our results below pertain to the case $n=12$ and $p=600$, so that the complexity level is equal to $c=50$. In kelly2023virtue, the benefits from complexity vanish theoretically for $c>10$ and empirically for $c>50$. This is why we henceforth work with a high and fixed level of complexity $c=50$.

In Equation (ref), an important parameter is $\gamma$, the bandwidth of the kernel, which determines the level of nonlinearity of the model induced by the RFFs. In kelly2023virtue, the authors show that their approach is insensitive to this choice, at least when $\gamma$ lies between 0.5 and 2. Our findings below paint a more nuanced picture.

Each month, given a random draw $w_i$ and the features from Equation (ref), we estimate the $\hat{\beta}$ as

equation[equation omitted — 140 chars of source]

where $RF$ stands for "random features" and $z>0$ is a regularization parameter which ensures that the inverse matrix is well defined when the number of predictor exceeds the sample size $n$. The weight assigned to the market portfolio is then simply this estimator defined above and the corresponding realized return is evaluated after one month. The procedure is then repeated until the end of the sample and monthly returns are averaged over alternative periods of time.

For comparison purposes, we also report the average returns of the strategy when the raw features, $G_t$, are used to compute the the position, instead of the RFFs, $S_{i,t}$. The portfolio weight of this strategy is hence given by

equation[equation omitted — 153 chars of source]

where the superscript $L$ stands for "linear". In Figure (ref), the corresponding returns are shown with horizontal dashed lines (they do not depend on $\gamma$).

Results: sensitivity to bandwidths and periods

In Figure (ref), we plot the average return obtained when using $\hat{\beta}^{RF}_i(z)$ as portfolio weight - averaged across 500 draws of $w_i$. We average returns over periods of 15 years for two reasons. First, because 15 years is already a long horizon for an investor and second because this allows to span several types of economic conditions: both periods of growth and slowdown. The values are reported for two levels of regularization: $z=0.01$ (left) and $z=100$ (right).

figure[figure omitted — 864 chars of source]

For a small regularization ($z=0.01$) and $\gamma = 2$, as in the original paper, the return for the period 2005-2019 is 7.24% but it shrinks to 0.52% for 1975-1989 - a reduction by a factor 14. An analyst in 1974 would probably choose $\gamma=1$ (from the red curve with $z = 0.01$), and hope for a return of 5%. But the performance in the following 15 years would prove very disappointing (less than 1%: again, a significant contraction).

For the sake of completeness, we also depict the sensitivity of the annualized Sharpe ratio in Figure (ref). Over the full sample, we find a maximum value of 0.3, slightly lower than the roughly 0.4 reported by kelly2023virtue in their Figure 8 (Panel A). These values are somewhat underwhelming if we put them in perspective with the 0.44 Sharpe ratio of the excess returns on the market index documented in pav2021sharpe over the 1926-2020 period - from the \href{https://mba.tuck.dartmouth.edu/pages/faculty/ken.french/data_library.html\string#Ken French data library}{Ken French data library}. For the raw market portfolio, when the risk-free rate is not subtracted, the Sharpe ratio increases to 0.62. This is even before taking trading costs into account: holding the market portfolio entails minimum trading costs, which is not the case of the timing strategy presented in kelly2023virtue.

figure[figure omitted — 766 chars of source]

In fact, a closer look reveals the unreasonably high levels of volatility that are implied by these Sharpe ratios. A monthly return of 0.03 (reported both in Figure (ref) and in kelly2023virtue for a bandwidth of two) translates arithmetically to 0.36 on an annual basis, thus a Sharpe ratio of 0.3 suggests a yearly volatility above 100%, which is prohibitively high, even for risk-prone investors (e.g., hedge funds).

Lastly, Figure (ref) overwhelmingly indicates that the best choice of bandwidth is indeed close to two. Except for one period (posterior to World War II), all curves reach their maximum when the bandwidth is in the vicinity of two. This is a stable and consistent pattern - which may explain why it is presented as the baseline case in kelly2023virtue.

Time-varying bandwidths

In Figure (ref), we note that the optimal choice of $\gamma$ is time-varying. It lies between $\gamma^*=0$ in the first period (left graph) or in the second period (right graph) and $\gamma^*=1.5$ in 1990-2004 or 2005-2019 (left panel) and in 1990-2004 (right panel). Choosing larger values for $\gamma$ reduces uncertainty, but at the cost of lower returns.

We now want to account for the fact that investors can adjust $\gamma$ based on past performance. In a new exercise, we pick, for a given period of 15 years, the bandwidth that yielded the best results in the previous period (we call this the “feasible” strategy), but also the bandwidth that proved the best ex-post, i.e., in hindsight. The corresponding returns are gathered in Table (ref).

As expected, when $\gamma$ is chosen in hindsight (rightmost columns), the strategy based on RFFs always outperforms the feasible one. Note that this scheme is unfeasible in practice. However, there are periods during which the RFF-based return is dominated by the simpler linear strategy. Therefore, this further underlines some limitations of the more sophisticated approach. With enough regularization, we see that the simpler model reaches an unconditional (full sample) return that is higher than the high-complexity strategy. This is in line with the conclusions of shen2024can.

table[table omitted — 2,381 chars of source]

Counterfactual returns

Our most important theoretical result is Proposition (ref), which quantifies the loss in performance based on the misalignment between the in-sample loadings and the out-of-sample loadings. In this last section, we thus propose to assess the strategy's return in the hypothetical scenario where the data would not suffer from posterior drift. This will allow us to quantify the loss from Proposition (ref). To this end, we generate counterfactual returns using the time series of estimated loadings $\hat{\beta}_t$ obtained by regressing the RFFs onto the real returns $r_{t+1}$. Formally, the counterfactual returns are sampled as follows:

equation[equation omitted — 101 chars of source]

where the $x_t^{(p)}$ are all $p$ predictors (i.e., the RFFs, not the original signals). We underline that this assumes no misspecification as well, i.e., $\theta=0$ in Equation (ref). We opt for simplicity and adding misspecification would artificially burden the exercise with an additional source of errors. Moreover, this route is already explored in kelly2023virtue, whereas we focus here on posterior drift.

Once they are generated, the hypothetical returns have low average returns as well as low volatility over the whole period (90 years). This was expected because the ridgeless algorithm selects the solution with the smallest norm. Therefore, we scale these returns in order to match the first two empirical moments of the historical realized returns, i.e.,

equation[equation omitted — 135 chars of source]

where $\Bar{r}_c$ (resp. $\Bar{r}$) is the sample mean of the counterfactual (resp. realized) returns and $\sigma_c$ (resp. $\sigma$) is their sample standard deviation.

The average returns are plotted in Figure (ref) and confirm the theoretical results of Section (ref). The performance of the strategy is substantially stronger without posterior drift (dashed lines). This difference shown for $z=0.01$ is more pronounced when the bandwidth is small. As the latter increases, all returns slowly shrink to zero. The case of strong regularization ($z=100$) corroborate these patterns and is postponed to Figure (ref) in the Appendix. Overall, it is clear that when the loadings do not change, the model that is learned allows to generate substantial profits. But when the DGP changes (via the posterior drift), then returns are curtailed.

figure[figure omitted — 758 chars of source]

Conclusion

This article seeks to document the sensitivity of the accuracy of overparametrized linear models in equity premium prediction. We find that the results can be promising, but also unstable. They depend on market conditions, even on long horizons, and on a critical parameter that is not straightforward to tune. We attribute these limitations to changes in links between premia and macro variables. These changes strongly attenuate the signal contained in predictors. In sum, large models can bring value, but not unconditionally, and one should resort to them while keeping these caveats in mind.

Replication package

The python code to replicate Figures (ref) and (ref) is available \href{https://colab.research.google.com/drive/1zAQQVQHRab9n8-VoI-8qmS3TTTUwqv1J?usp=sharing\string#here}{here}.

The python code to replicate Figures (ref), (ref), (ref) and (ref) is available \href{https://colab.research.google.com/drive/13gGPOZfe6zMN7DIWrZWXDytvEVXzuiU-?usp=sharing\string#here}{here}.

The python code to replicate Figures (ref), (ref), (ref) and (ref) is available \href{https://colab.research.google.com/drive/1n2x5EoAQMOiT7hZfxwpz4cYIOTBdMTby?usp=sharing\string#here}{here}.

The python code to replicate Figure (ref) is available \href{https://colab.research.google.com/drive/161ML8ElvhTcv9auAXnSICvFEsqu2Tplr?usp=sharing\string#here}{here}.