EconBase
← Back to paper

A Quadratic Link between Out-of-Sample $R^2$ and Directional Accuracy

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

15,452 characters · 12 sections · 12 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Nontrivial Upper Bound on the Out-of-Sample R2 in Return Forecasting

frontmatter\ead{[email removed], [email removed]} \cortext[cor1]{Corresponding author} \begin{abstract} This study establishes a nontrivial upper bound on the out-of-sample $R^2$ ($R^2_{\text{OOS}}$) in return forecasting. In particular, we define a coin-flip oracle model that, under the same directional accuracy, theoretically outperforms practical models in terms of MSE. The $R^2_{\text{OOS}}$ of the oracle model, whose analytical expression is a quadratic function of directional accuracy, can therefore serve as a tractable upper bound on the actual $R^2_{\text{OOS}}$. Empirical analyses across multiple forecasting scenarios reveal that the $R^2_{\text{OOS}}$ values of common predictive models are fundamentally bounded by this quadratic function. \end{abstract} \begin{keyword} Nontrivial Upper Bound \sep Out-of-Sample $R^2$ \sep Return Forecasting \sep Directional Accuracy \sep Metric Disconnect \JEL C52 \sep C53 \sep G17 \end{keyword}

Introduction

In this study, we aim to explore a nontrivial upper bound on the out-of-sample $R^2$ ($R^2_{\text{OOS}}$) in return forecasting. Prior studies have shown that predictive models typically perform worse than the naive baseline in terms of various error metrics Meese1983, Kilian2003, Campbell2008, Moosa2013, FotiosPetropoulos2022, Ellwanger2023, raising the question of whether one can continually improve out-of-sample performance by using more advanced predictive models. Given that the complexity of predictive models contributes little to $R^2_{\text{OOS}}$ values Welch2008, FotiosPetropoulos2022, Farmer2023, we argue that a nontrivial upper bound other than $R^2_{\text{OOS}} = 1$ exists.

Since the performance of the unconditional MSE-optimal forecast is intractable, we define a coin-flip oracle model as a proxy for the theoretically best predictive model. In particular, the oracle forecast uses the true conditional expected absolute return at each step, and its predicted sign is generated by a Bernoulli process with a constant probability of sign correctness. Under the same directional accuracy, it theoretically outperforms practical models in terms of MSE. Consequently, the $R^2_{\text{OOS}}$ of this oracle forecast, whose analytical expression is a quadratic function of directional accuracy, provides a tractable upper bound for real-world predictive models.

By juxtaposing the performance of various predictive models across multiple forecasting scenarios, we observe that the $R^2_{\text{OOS}}$ values of practical models are fundamentally bounded by this quadratic function. The findings of this study also offer a novel perspective on the dependency between conditional mean predictability and sign predictability.

Derivation of the Upper Bound

The Coin-Flip Oracle Model

Let $r_t = s_t |r_t|$ denote the log return of a financial asset at time $t$, where $s_t \in \{-1, 1\}$ denotes the sign of $r_t$, with zero returns assigned a positive sign. The forecast of a practical model, denoted as $\hat{r}_t^{\text{practical}}$, is given by

equation[equation omitted — 94 chars of source]

where $\hat{s}_t \in \{-1, 1\}$ denotes the sign of $\hat{r}_t^{\text{practical}}$, and $\hat{m}_t$ denotes the predicted magnitude. Let $\mathbb{I}_t^{\text{practical}}$ denote the indicator of sign correctness for $\hat{s}_t$. Accordingly, the conditional probability of sign correctness, $p_t$, satisfies $\mathbb{P}(\mathbb{I}_t^{\text{practical}} = 1 \mid \Omega_{t-1}) = p_t$, where $\Omega_{t-1}$ denotes the information set available at time $t-1$.

We then define an oracle forecast $\hat{r}_t^{\text{oracle}}$, whose sign forecast $\hat{s}_t^{\text{oracle}}$ is generated by a Bernoulli process with a constant probability $p$ ($p \ge 0.5$) such that $\mathbb{P}(\mathbb{I}_t^{\text{oracle}} = 1 \mid \Omega_{t-1}) = p = \mathbb{E}[p_t]$, where $\mathbb{I}_t^{\text{oracle}}$ is the indicator of sign correctness for $\hat{s}_t^{\text{oracle}}$. The magnitude of $\hat{r}_t^{\text{oracle}}$ under MSE loss is $(2p - 1)\psi_t$, where $\psi_t$ denotes the conditional expected absolute return $\mathbb{E}[|r_t| \mid \Omega_{t-1}]$. Therefore, $\hat{r}_t^{\text{oracle}}$ has the following form:

equation[equation omitted — 108 chars of source]

Given that $p = \mathbb{E}[p_t]$, we can compare the MSEs of the two types of forecasts under the same directional accuracy. Since $s_t$ and $|r_t|$ are considered conditionally independent given $\Omega_{t-1}$ Anatolyev2010, the MSE of $\hat{r}_t^{\text{practical}}$, denoted as $\text{MSE}^{\text{practical}}$, is given by

equation[equation omitted — 548 chars of source]

Moreover, the MSE of $\hat{r}_t^{\text{oracle}}$, denoted as $\text{MSE}^{\text{oracle}}$, is given by

equation[equation omitted — 386 chars of source]

Based on (ref), we can derive the difference between the two MSEs as follows:

equation[equation omitted — 207 chars of source]

Since high volatility inflates expected return magnitudes Merton1980, French1987 while reducing sign predictability Christoffersen2006, $p_t$ and $\psi_t \hat{m}_t$ move in opposite directions in response to volatility. Therefore, we have $\operatorname{Cov}(p_t, \psi_t \hat{m}_t) \le 0$, which ensures that $\text{MSE}^{\text{practical}} - \text{MSE}^{\text{oracle}} \ge 0$. Thus, given a directional accuracy $p$, the oracle model theoretically outperforms practical models in terms of MSE.

The Out-of-Sample \texorpdfstring{$R^2$}{R2} of the Oracle Model

According to Welch2008 and Gu2020, as the out-of-sample size approaches infinity (with the zero-return prediction serving as the baseline), the $R^2_{\text{OOS}}$ of $\hat{r}_t^{\text{oracle}}$ can be expressed as follows:

equation[equation omitted — 233 chars of source]

Since the oracle forecast error is unconditionally orthogonal to the forecast itself, the expected squared realized return can be decomposed into the expected squared forecast and the MSE of $\hat{r}_t^{\text{oracle}}$:

equation[equation omitted — 162 chars of source]

Substituting (ref) back into (ref) simplifies $\operatorname{plim} R^2_{\text{OOS}}$ to:

equation[equation omitted — 143 chars of source]

We further specify the squared realized return as $r_t^2 = \sigma_t^2 \varepsilon_t$, where $\sigma_t$ is the $\Omega_{t-1}$-measurable conditional volatility and $\varepsilon_t$ is a positive multiplicative error term assumed to be i.i.d. Granger1995, Engle2006. Accordingly, we have $\psi_t = \sigma_t \mathbb{E}[\varepsilon_t^{1/2}]$ and $\psi_t^2 = \sigma_t^2 \big( \mathbb{E}[\varepsilon_t^{1/2}] \big)^2$. Based on (ref), $\mathbb{E}[(\hat{r}_t^{\text{oracle}})^2]$ can be expressed as follows:

equation[equation omitted — 232 chars of source]

Using the law of total expectation, we can express $\mathbb{E}[r_t^2]$ as follows:

equation[equation omitted — 223 chars of source]

By substituting (ref) back into (ref), we can analytically express the $R^2_{\text{OOS}}$ of the oracle forecast as a quadratic function of directional accuracy $p$:

equation[equation omitted — 89 chars of source]

where $\kappa = \big( \mathbb{E}[\varepsilon_t^{1/2}] \big)^2 \big( \mathbb{E}[\varepsilon_t] \big)^{-1}$.

In the empirical analysis, the $R^2_{\text{OOS}}$ of the oracle forecast can be estimated as follows:

equation[equation omitted — 88 chars of source]

where DA represents the realized out-of-sample directional accuracy, and $\hat{\kappa}$ is the sample estimate of $\kappa$ computed over the out-of-sample period of $T$ steps:

equation[equation omitted — 183 chars of source]

In (ref), $\hat{\varepsilon}_t$ is the estimated multiplicative error, given by $\hat{\varepsilon}_t = r_t^2 \hat{\sigma}_t^{-2}$, where $\hat{\sigma}_t$ represents the out-of-sample conditional volatility, which can be estimated using a conditional volatility model such as GARCH(1,1) Bollerslev2023.

Empirical Analysis

Data

With the quadratic function provided by (ref) as the upper bound, we now turn to actual financial data to examine whether the performance of practical models is bounded by this theoretical limit. Since sign dynamics are most prevalent at intermediate frequencies Christoffersen2006, we retrieve 14 financial time series from Yahoo Finance, each containing weekly closing prices. The details of each time series are provided in (ref). Moreover, each dataset is split into in-sample and out-of-sample sets at varying ratios ranging from 80:20 to 60:40. For each data splitting ratio, one $\hat{\kappa}$ is computed according to (ref). The values of $\hat{\kappa}$ for each forecasting scenario are provided in (ref). We then employ ten conventional predictive models to generate out-of-sample log return forecasts. The model details are provided in (ref). The performance of each model is evaluated using $R^2_{\text{OOS}}$ and DA. In addition, the top 2% of the absolute log returns in each out-of-sample set are excluded to mitigate the impact of sample noise on performance evaluation Gu2020.

Results

We represent each model's performance as a coordinate pair $\big((2\text{DA} - 1)^2, R^2_{\text{OOS}} / \hat{\kappa}\big)$ and plot all pairs collectively in a single two-dimensional space. This allows the juxtaposition of model performances to cover a wider range of directional accuracies, as shown in (ref). The reference line represents the nontrivial upper bound given by the oracle model.

figure[figure omitted — 359 chars of source]

Several observations can be made based on (ref). First, the data points are fundamentally bounded by the reference line, indicating that model performance is constrained by the quadratic function. Performance falling below the reference line can be attributed to model misspecification or sample variation. Second, many data points have negative y-axis values alongside positive x-axis values, indicating that negative $R^2_{\text{OOS}}$ values are accompanied by modest directional accuracies, which is consistent with the metric disconnect phenomenon reported in empirical studies Leitch1991, Pesaran1995. Third, the models can outperform the zero-return baseline when directional accuracy is high. The higher the directional accuracy, the greater the potential $R^2_{\text{OOS}}$ improvement over the naive baseline. Fourth, the results evaluated on the trimmed data show that as sample variation is reduced, spurious deviations above the theoretical bound are largely removed.

Conclusion

While $R^2_{\text{OOS}}$ measures the goodness-of-fit of return forecasts, this study shows that it is fundamentally constrained by the nature of the data as well as directional accuracy. Given the quadratic link between $R^2_{\text{OOS}}$ and directional accuracy, minimizing magnitude-based error metrics and maximizing directional accuracy emerge as aligned optimization objectives. Sign predictability does not depend on conditional mean predictability; however, the reverse relationship holds.

Data Availability

The data and code are available at

\url{https://github.com/Zhang-Cheng-76200/R2DA}.

Acknowledgments

The author received no specific funding for this research.

Declaration of interest statement

The author reports that there are no competing interests to declare.

Declaration of generative AI and AI-assisted technologies in the manuscript preparation process

During the preparation of this work, the author used Gemini 3 to improve the readability of the manuscript. After using this tool, the author reviewed and edited the content as needed and takes full responsibility for the content of the published article.