Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
15,452 characters · 12 sections · 12 citation commands
A Nontrivial Upper Bound on the Out-of-Sample R2 in Return Forecasting
In this study, we aim to explore a nontrivial upper bound on the out-of-sample $R^2$ ($R^2_{\text{OOS}}$) in return forecasting. Prior studies have shown that predictive models typically perform worse than the naive baseline in terms of various error metrics Meese1983, Kilian2003, Campbell2008, Moosa2013, FotiosPetropoulos2022, Ellwanger2023, raising the question of whether one can continually improve out-of-sample performance by using more advanced predictive models. Given that the complexity of predictive models contributes little to $R^2_{\text{OOS}}$ values Welch2008, FotiosPetropoulos2022, Farmer2023, we argue that a nontrivial upper bound other than $R^2_{\text{OOS}} = 1$ exists.
Since the performance of the unconditional MSE-optimal forecast is intractable, we define a coin-flip oracle model as a proxy for the theoretically best predictive model. In particular, the oracle forecast uses the true conditional expected absolute return at each step, and its predicted sign is generated by a Bernoulli process with a constant probability of sign correctness. Under the same directional accuracy, it theoretically outperforms practical models in terms of MSE. Consequently, the $R^2_{\text{OOS}}$ of this oracle forecast, whose analytical expression is a quadratic function of directional accuracy, provides a tractable upper bound for real-world predictive models.
By juxtaposing the performance of various predictive models across multiple forecasting scenarios, we observe that the $R^2_{\text{OOS}}$ values of practical models are fundamentally bounded by this quadratic function. The findings of this study also offer a novel perspective on the dependency between conditional mean predictability and sign predictability.
Let $r_t = s_t |r_t|$ denote the log return of a financial asset at time $t$, where $s_t \in \{-1, 1\}$ denotes the sign of $r_t$, with zero returns assigned a positive sign. The forecast of a practical model, denoted as $\hat{r}_t^{\text{practical}}$, is given by
where $\hat{s}_t \in \{-1, 1\}$ denotes the sign of $\hat{r}_t^{\text{practical}}$, and $\hat{m}_t$ denotes the predicted magnitude. Let $\mathbb{I}_t^{\text{practical}}$ denote the indicator of sign correctness for $\hat{s}_t$. Accordingly, the conditional probability of sign correctness, $p_t$, satisfies $\mathbb{P}(\mathbb{I}_t^{\text{practical}} = 1 \mid \Omega_{t-1}) = p_t$, where $\Omega_{t-1}$ denotes the information set available at time $t-1$.
We then define an oracle forecast $\hat{r}_t^{\text{oracle}}$, whose sign forecast $\hat{s}_t^{\text{oracle}}$ is generated by a Bernoulli process with a constant probability $p$ ($p \ge 0.5$) such that $\mathbb{P}(\mathbb{I}_t^{\text{oracle}} = 1 \mid \Omega_{t-1}) = p = \mathbb{E}[p_t]$, where $\mathbb{I}_t^{\text{oracle}}$ is the indicator of sign correctness for $\hat{s}_t^{\text{oracle}}$. The magnitude of $\hat{r}_t^{\text{oracle}}$ under MSE loss is $(2p - 1)\psi_t$, where $\psi_t$ denotes the conditional expected absolute return $\mathbb{E}[|r_t| \mid \Omega_{t-1}]$. Therefore, $\hat{r}_t^{\text{oracle}}$ has the following form:
Given that $p = \mathbb{E}[p_t]$, we can compare the MSEs of the two types of forecasts under the same directional accuracy. Since $s_t$ and $|r_t|$ are considered conditionally independent given $\Omega_{t-1}$ Anatolyev2010, the MSE of $\hat{r}_t^{\text{practical}}$, denoted as $\text{MSE}^{\text{practical}}$, is given by
Moreover, the MSE of $\hat{r}_t^{\text{oracle}}$, denoted as $\text{MSE}^{\text{oracle}}$, is given by
Based on (ref), we can derive the difference between the two MSEs as follows:
Since high volatility inflates expected return magnitudes Merton1980, French1987 while reducing sign predictability Christoffersen2006, $p_t$ and $\psi_t \hat{m}_t$ move in opposite directions in response to volatility. Therefore, we have $\operatorname{Cov}(p_t, \psi_t \hat{m}_t) \le 0$, which ensures that $\text{MSE}^{\text{practical}} - \text{MSE}^{\text{oracle}} \ge 0$. Thus, given a directional accuracy $p$, the oracle model theoretically outperforms practical models in terms of MSE.
According to Welch2008 and Gu2020, as the out-of-sample size approaches infinity (with the zero-return prediction serving as the baseline), the $R^2_{\text{OOS}}$ of $\hat{r}_t^{\text{oracle}}$ can be expressed as follows:
Since the oracle forecast error is unconditionally orthogonal to the forecast itself, the expected squared realized return can be decomposed into the expected squared forecast and the MSE of $\hat{r}_t^{\text{oracle}}$:
Substituting (ref) back into (ref) simplifies $\operatorname{plim} R^2_{\text{OOS}}$ to:
We further specify the squared realized return as $r_t^2 = \sigma_t^2 \varepsilon_t$, where $\sigma_t$ is the $\Omega_{t-1}$-measurable conditional volatility and $\varepsilon_t$ is a positive multiplicative error term assumed to be i.i.d. Granger1995, Engle2006. Accordingly, we have $\psi_t = \sigma_t \mathbb{E}[\varepsilon_t^{1/2}]$ and $\psi_t^2 = \sigma_t^2 \big( \mathbb{E}[\varepsilon_t^{1/2}] \big)^2$. Based on (ref), $\mathbb{E}[(\hat{r}_t^{\text{oracle}})^2]$ can be expressed as follows:
Using the law of total expectation, we can express $\mathbb{E}[r_t^2]$ as follows:
By substituting (ref) back into (ref), we can analytically express the $R^2_{\text{OOS}}$ of the oracle forecast as a quadratic function of directional accuracy $p$:
where $\kappa = \big( \mathbb{E}[\varepsilon_t^{1/2}] \big)^2 \big( \mathbb{E}[\varepsilon_t] \big)^{-1}$.
In the empirical analysis, the $R^2_{\text{OOS}}$ of the oracle forecast can be estimated as follows:
where DA represents the realized out-of-sample directional accuracy, and $\hat{\kappa}$ is the sample estimate of $\kappa$ computed over the out-of-sample period of $T$ steps:
In (ref), $\hat{\varepsilon}_t$ is the estimated multiplicative error, given by $\hat{\varepsilon}_t = r_t^2 \hat{\sigma}_t^{-2}$, where $\hat{\sigma}_t$ represents the out-of-sample conditional volatility, which can be estimated using a conditional volatility model such as GARCH(1,1) Bollerslev2023.
With the quadratic function provided by (ref) as the upper bound, we now turn to actual financial data to examine whether the performance of practical models is bounded by this theoretical limit. Since sign dynamics are most prevalent at intermediate frequencies Christoffersen2006, we retrieve 14 financial time series from Yahoo Finance, each containing weekly closing prices. The details of each time series are provided in (ref). Moreover, each dataset is split into in-sample and out-of-sample sets at varying ratios ranging from 80:20 to 60:40. For each data splitting ratio, one $\hat{\kappa}$ is computed according to (ref). The values of $\hat{\kappa}$ for each forecasting scenario are provided in (ref). We then employ ten conventional predictive models to generate out-of-sample log return forecasts. The model details are provided in (ref). The performance of each model is evaluated using $R^2_{\text{OOS}}$ and DA. In addition, the top 2% of the absolute log returns in each out-of-sample set are excluded to mitigate the impact of sample noise on performance evaluation Gu2020.
We represent each model's performance as a coordinate pair $\big((2\text{DA} - 1)^2, R^2_{\text{OOS}} / \hat{\kappa}\big)$ and plot all pairs collectively in a single two-dimensional space. This allows the juxtaposition of model performances to cover a wider range of directional accuracies, as shown in (ref). The reference line represents the nontrivial upper bound given by the oracle model.
Several observations can be made based on (ref). First, the data points are fundamentally bounded by the reference line, indicating that model performance is constrained by the quadratic function. Performance falling below the reference line can be attributed to model misspecification or sample variation. Second, many data points have negative y-axis values alongside positive x-axis values, indicating that negative $R^2_{\text{OOS}}$ values are accompanied by modest directional accuracies, which is consistent with the metric disconnect phenomenon reported in empirical studies Leitch1991, Pesaran1995. Third, the models can outperform the zero-return baseline when directional accuracy is high. The higher the directional accuracy, the greater the potential $R^2_{\text{OOS}}$ improvement over the naive baseline. Fourth, the results evaluated on the trimmed data show that as sample variation is reduced, spurious deviations above the theoretical bound are largely removed.
While $R^2_{\text{OOS}}$ measures the goodness-of-fit of return forecasts, this study shows that it is fundamentally constrained by the nature of the data as well as directional accuracy. Given the quadratic link between $R^2_{\text{OOS}}$ and directional accuracy, minimizing magnitude-based error metrics and maximizing directional accuracy emerge as aligned optimization objectives. Sign predictability does not depend on conditional mean predictability; however, the reverse relationship holds.
The data and code are available at
\url{https://github.com/Zhang-Cheng-76200/R2DA}.
The author received no specific funding for this research.
The author reports that there are no competing interests to declare.
During the preparation of this work, the author used Gemini 3 to improve the readability of the manuscript. After using this tool, the author reviewed and edited the content as needed and takes full responsibility for the content of the published article.