EconBase
← Back to paper

Probabilistic Forecasting for Day-ahead Electricity Prices, Battery Trading Strategies and the Economic Evaluation of Predictive Accuracy

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

134,500 characters · 26 sections · 100 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Probabilistic Forecasting for Day-ahead Electricity Prices, Battery Trading Strategies and the Economic Evaluation of Predictive Accuracy

abstractElectricity price forecasting supports decision-making in energy markets and asset operation. Probabilistic forecasts are increasingly adopted to explicitly quantify uncertainty, typically issued as quantile predictions or ensembles of the full predictive distribution. However, how improvements in statistical forecast quality translate into economic value remains unclear. Battery storage arbitrage in day-ahead markets is a popular application-based benchmark for this purpose. We analyze quantile-based trading strategies (QBTS) and identify two critical flaws: they do not incentivize honest probabilistic forecasting and they ignore the intertemporal dependence structure of electricity prices. We therefore frame battery optimization as a stochastic program based on fully probabilistic forecasts and examine decision quality measurement for risk-neutral and risk-averse settings under different uncertainty models. Our discussion touches both sides of the coin: How reliable is the economic evaluation of forecasting models though (simplified) application studies -- and how do improvements in statistical forecast quality for stochastic programs relate to the decision-quality and economic performance? We provide theoretical justification and empirical evidence from a case study on the German electricity market. Our results highlight the pitfalls of ranking forecasting models through battery trading strategies. We conclude with implications for evaluation practice and directions for future research in application-based forecast assessment.

Keywords: Probabilistic Forecasting; Scoring Rules; Battery Optimization; Stochastic Programming; Decision Quality; Forecast Evaluation

Introduction

Forecasting electricity prices is crucial for informed decision-making in the daily operations of energy companies. Over the last years, both industry and research have focused on probabilistic forecasting, which “has a lot to offer, in particular, improved assessment of future uncertainty, ability to plan different strategies for the range of possible outcomes, increased effectiveness of submitted bids” nowotarski2018recent. A pressing question in practice is whether improved forecasting performance, in terms of scoring rules or loss functions, improves decision-making, measured in monetary terms granger2000economic,yardley2021beyond. To address this need, various application studies have been proposed in the energy forecasting literature, nitka2023combining, vogler2021event. A common showcase for the economic benefits of better forecasts is the simplified operation of grid-scale batteries energy storage systems (BESS), which can buy electricity (and charge) when prices are low and sell electricity (and discharge) when prices are high. Better anticipation of prices should translate into higher profits. This storage arbitrage example has been adopted for both point and probabilistic forecasting studies sang2022electricity, chȩc2025extrapolating, serafin2025loss, maciejowska2025statistical, uniejewski2025smoothing, o2025conformal, o2025optimising.

It is widely agreed on that forecast evaluation and model selection should be conducted using proper scoring rules. Proper scoring rules incentivize honest forecasting, that is, the optimal score is attained for predicting the true distribution. Scoring rules are called strictly proper if the optimum is unique. Strictly proper scoring rules are thus the gold standard for forecast evaluation and model selection gneiting2007strictly. However, the strict propriety alone does not guarantee that scoring rules penalize forecast errors symmetrically buchweitz2025asymmetric, provide meaningful discrimination alexander2024evaluating,marcotte2023regions or that they are aligned with the decision-making process at hand. To bridge this gap, the properties of scoring rules and their modification through weighting, thresholding and censoring is an active area of research allen2023weighted, de2025localizing, shahroudi2025aligning.

Consequently, some questions arise: (How) can we infer the quality of forecasts based on the quality of the decisions? How can we assess the quality of decisions under competing models of uncertainty? Can battery trading strategies provide meaningful discrimination between forecasts of different quality? Are battery trading strategies aligned with the concept of proper scoring rules? How are statistical scores related to the economic performance? These questions will guide us throughout this paper.

The majority of works concerning probabilistic forecasting of electricity prices focuses on the marginal distribution of hourly prices. Based on this, quantile-based trading strategies (QBTS), originally introduced by uniejewski2025smoothing, employ quantile-based limit orders on the day-ahead market. Briefly, a trader places limit buy and limit sell orders based on $\alpha$ and ($1-\alpha$) -- quantiles of the predicted distribution, thereby controlling a trade-off between frequency of trading and anticipated profit of each trade. The approach has been used in various subsequent studies (see Table (ref) for an overview) and extended by o2025optimising,o2025conformal,o2024conformal. While instinctively intuitive, we provide theoretical and empirical evidence that the approach can be gamed by providing systematically biased forecasts and is not aligned with the concept of proper scoring rules. Moreover, we show that the approach does not anticipate diversification benefits by neglecting the dependence structure of electricity prices.

Keeping the battery example, a natural further starting point is to ask whether battery trading strategies based on a fully multivariate probabilistic forecasts, optimized for common risk-measures such as expected profits or the conditional value-at-risk (CVAR) can be used to evaluate the quality of forecasts. We show that even in this case, the resulting profits are not necessarily a strictly proper scoring rule and have potentially low discriminatory power. For point forecasts, maciejowska2025statistical empirically evaluate the relationship between the statistical quality of point forecasts and the economic performance of battery trading strategies. However, formal analysis and the evaluation of the probabilistic and risk-averse case is notably absent in the literature. We contribute to this gap by providing a theoretical and empirical analysis of the relationship between forecast quality and a notion of decision quality, measured by the predictive accuracy of the objective value of our stochastic program under competing forecasts. We relate decision quality, statistical forecast and economic forecast evaluation.

Our contribution can be summarized in the following points:

itemize• We analyze the theoretical properties of popular QBTS and show that they do not promote honest forecasting, as they can be gamed by providing overdispersed forecasts. We provide formal justification and a simulation study for this claim in Section (ref). • We discuss more generally the evaluation of battery trading strategies framed as stochastic programs based on fully multivariate probabilistic forecasts (Section (ref)). In particular: \begin{itemize} • We show that forecast evaluation using battery trading strategies is no strictly proper scoring rule and has potentially low discriminatory power. Low discriminatory power is an issue, as it reduces robustness of results across changing application settings (Section (ref)). • We develop options for the evaluation of decision quality under different probabilistic forecasts in a risk-averse battery optimization setting (Section (ref)). \end{itemize} • As a by-product, we extend the use of association measures for forecast evaluation introduced by maciejowska2025statistical to the probabilistic case and present a proper scoring rule based on Kendall's $\tau$ in Section (ref). • We provide an extensive empirical evaluation in Sections along a wide range of statistical and economic metrics, as well as different battery configurations in Section (ref).

We believe that our findings are of interest to multiple research communities, including the forecasting community, the energy economics community and the stochastic optimization/programming community.

itemize• The primary audience of this paper is the (electricity price) forecasting community. However, we believe that understanding the downstream impact of forecast on realistic and complex decision-making processes is also crucial for other forecasting applications such as supply chain and inventory management. • Our results are also of interest to the stochastic optimization/programming community. Much research has been done on scenario reduction methods to reduce computational complexity, but less work has been done on the influence of the underlying scenarios on the solution quality dupavcova2003scenario, lohndorf2016empirical, ziel2021energy, especially with respect to different underlying models of the uncertainty. • A better understanding of the relationship between scoring rules and economic performance can also inform the development of new loss functions for the integration of forecasting and optimization, as strong theory exist for learning using scoring rules dawid2016minimum,gneiting2007strictly, while custom loss functions are often developed ad-hoc sang2022electricity, sbaraglia2024optimizing,serafin2025loss.

On the other hand, we acknowledge some limitations a priori: this work aims at deepening the understanding of the relationship between forecast quality and decision quality in the context of battery trading strategies. We do not focus on the absolute profitability of BESS investments in energy markets spodniak2021profitability, staffell2016maximising or on sophisticated, multi-market optimization lohndorf2023value, kraft2023stochastic, ghadimi2024stochastic.

The remainder of this paper is structured as follows: Section (ref) formalizes our claim and analyzes the objective value of quantile-based optimization from a theoretical vantage point. Section (ref) presents and analyzes the proposed multivariate dynamic programming approach. Section (ref) gives a case study. The last section concludes and identifies avenues for further research.

Quantile-based Battery Trading Strategy

Quantile-based trading strategies (QBTS) are a popular approach to evaluate the economic value of probabilistic forecasts in the context of battery trading. The main idea is to place limit orders based on quantiles of the predicted distribution of electricity prices. By adjusting the quantile level $\alpha$, the trader can control the trade-off between the frequency of trading and the anticipated profit of each trade. Two main variants exist in the literature, the quantile-based trading strategies based on the works of uniejewski2025smoothing and the TS-1/2/3 strategies introduced by o2025optimising,o2024conformal,o2025conformal.

table[table omitted — 2,830 chars of source]

Optimization Algorithm

We briefly describe QBTS and their variations in the literature. Assume the trader has access to a probabilistic forecast of the day-ahead electricity prices, which is issued as quantile predictions for each delivery hour. The trader aims to maximize the expected returns from trading a battery storage asset on the day-ahead market. The battery has a certain capacity $\kappa$ and efficiency $\eta$. The strategy consists of two steps:

enumerate• Based on the median forecast, the trader identifies the hours with the lowest and highest predicted prices, denoted as $b$ and $s$. We ensure that $b \neq s$ to avoid buying and selling in the same hour. • For a given level $\alpha$, the trader places a limit buy order at hour $b$ with a price based on the $(1-\alpha)$-quantile of the predicted distribution for hour $b$, and a limit sell order at hour $s$ with a price based on the $\alpha$-quantile of the predicted distribution for hour $s$.

In this case, we assume that the execution of the buy and sell bid is coupled, i.e. either both are executed or neither is executed. Thereby, the battery is always in a valid state and no additional orders need to be placed in other markets to balance the battery.\footnote{This assumption is closely related to the so-called loop bid, a product available at the EPEX power exchange that guarantees joint or no execution of coupled buy and sell bids. Contrary to individual directional limit prices however, loop bids are conditional on the total cashflow of the linked bids epexspot2025trading. Framing as a loop bid would further move away from the original QBTS.} Alternatively, the trader might be faced with partial execution, that is, only the buy or only the sell trade is executed. There are multiple ways to handle this, the most prominent ones are described in the following and an overview is given in Table (ref).

itemize• The trader needs to reserve capacity in the battery to always remain in physically valid states or place additional orders in the balancing market to ensure that the battery is on half-full at the end of the day. This approach has been used by uniejewski2025smoothing. • The trader can place additional orders in the balancing market to balance the battery at the end of the day. This approach has been used by maciejowska2024probabilistic.

Both options introduce additional risk and make the (predicted) distribution of battery returns intractable. Therefore, we focus on the simplified approach in the following theoretical analysis and empirical evaluation. Figure (ref) summarizes the QBTS approach. Algorithmic descriptions of all battery trading strategies are given in the supplementary material.

figure[figure omitted — 502 chars of source]

The TS-1 strategy as proposed by o2025optimising is based on a similar algorithm. However, instead of placing limit orders based on quantiles, quantiles are used to choose the buying and selling hours and unlimited bids are placed. The objective function of the TS-1 strategy is given by

equation[equation omitted — 124 chars of source]

The main difference to QBTS is that the TS-1 strategy does not place limit orders but unlimited orders at the selected hours and hence bids are always accepted. The TS-2 strategy extends the TS-1 strategy by placing additional orders in the balancing market to balance the battery at the end of the day, while the TS-3 strategy extends the TS-1 strategy by considering ramp rate constraints. We focus on the TS-1 strategy in the following, as it is the most similar to QBTS and allows for a more direct comparison.

Theoretical Analysis

We begin the theoretical discussion by recapitulating the definition of a strictly proper scoring.

definition[Strictly Proper Scoring Rule, gneiting2007strictly] For a predicted distribution $\mathcal{F}$ and the realized value $x$ taken from the true distribution $\mathcal{D}$, the scoring rule $S(\mathcal{F}, x) \rightarrow \mathbb{R}^+$ denotes the forecaster's reward. As a convention, we have lower scores corresponding to better forecasts. For a proper scoring rule we require: \begin{equation} \mathbb{E}[S(\mathcal{D}, x)] \leq \mathbb{E}[S(\mathcal{F}, x)] \end{equation} that is, the score is only minimized when forecasting the true distribution. A scoring rule is called strictly proper if the minimum is unique for the true distribution, i.e. $\mathbb{E}[S(\mathcal{D}, x)] = \mathbb{E}[S(\mathcal{F}, x)]$ iff $\mathcal{F} = \mathcal{D}$.

We denote the unknown, true distribution of day-ahead electricity prices $\boldsymbol{P}_d = (P_{d,0}, ..., P_{d, 23}) \sim \mathcal{D}_d$ on day $d$ as $\mathcal{D}_d$. Denote the realized prices as $\ensuremath{\bm{\mathrm{p}}}_d = (p_{d,0}, ..., p_{d, 23})$. We describe the 24-dimensional vector of decisions as $\ensuremath{\bm{\mathrm{a}}}(b, s) = (a_0, ..., a_h, ..., a_H)$, where $b$ and $s$ denote the hour(s) in which electricity is bought and sold. We will omit the $d$ in the following as the optimization is repeated every day. Denote the efficiency as $\eta$ and the storage capacity as $\kappa$ and hence we have

equation[equation omitted — 171 chars of source]

For the first part we constrain ourselves to only one buy and one selling bid. We implicitly assume that the charge/discharge capacity $\xi = \kappa$. Without loss of generality we assume that $b < s$, hence we buy (and charge) before we sell (and discharge). The revenues $R_d$ for an unconstrained bid $\ensuremath{\bm{\mathrm{a}}}(b, s)$ are given by

equation[equation omitted — 163 chars of source]

and for a bid with limit prices $Q_b^{1-\alpha}$ and $Q_s^{\alpha}$, we have:

equation[equation omitted — 229 chars of source]

This allows us to state our main result regarding quantile-based trading strategies.

prop[QBTS profits as scoring rule] For given risk-control parameter $\alpha$, the expected profits of the QBTS are given by \begin{small}\begin{align} \mathbb{E}[R(b, s, \alpha)] &= \underbrace{\mathbb{P}(P_{b} \leq {Q}_b^{1-\alpha}; \; P_{s} \geq Q_s^{\alpha})}_Acceptance probability. \underbrace{\left(-\frac{1}{\eta} \kappa \mathbb{E}[P_b | P_{b} \leq {Q}_b^{1-\alpha}; P_s > Q_s^{\alpha}] + \eta\kappa\mathbb{E}[P_s | P_{b} < Q_b^{1-\alpha}; P_{s} \geq Q_s^{\alpha}] \right)}_Expected profits if accepted (EP). \\ &+ \left(1 - \mathbb{P}(P_{b} \leq {Q}_b^{1-\alpha}; \; P_{s} \geq Q_s^{\alpha})\right) 0 \nonumber \end{align}\end{small} This expression is not uniquely maximized for the true forecast, but is potentially maximized by an overdispersed respectively underdispersed forecast. Providing an overdispersed forecast increases the acceptance probability (AP), but decreases the expected profits (EP) if accepted. Vice versa, providing an underdispersed forecast decreases the AP, but increases the expected profits if accepted. The relative strength of the effects depends on the underlying price distribution and the risk-control parameter $\alpha$.

We show that an overdispersed forecast can lead to higher expected profits than the perfect forecast. Assume the price distribution is symmetric and we have an over-dispersed forecasts $\widetilde{\boldsymbol{P}} = b \cdot \boldsymbol{P}$, where $\boldsymbol{P} \sim \mathcal{D}$ is the true price distribution and $b > 1$ is a dispersion parameter. We have $\widetilde{Q}_h^q = b \cdot Q_h^q$, where $Q_h^q$ is the quantile function of the true distribution. For simplificity, we assume that delivery periods are independent. The mean and median forecast are unbiased, $\widetilde{Q}_h^{0.5} = Q_h^{0.5}$. We therefore have:

align[align omitted — 528 chars of source]

We are interested in the difference between the expected profits of the over-dispersed forecast and the perfect forecast, that is, $\mathbb{E}[\widetilde{R}(b, s, \alpha)] - \mathbb{E}[R(b, s, \alpha)]$. We look at the probability of bid acceptance and the expected value of the returns separately. We assume independence of the delivery periods, hence we have $\text{AP} = \mathbb{P}(P_{b} \leq Q_b^{1-\alpha}; \; P_{s} \geq Q_s^{\alpha}) = (1 - \alpha)^2$ for the perfect forecast and $\mathbb{P}(P_{b} \leq \widetilde{Q}_b^{1-\alpha}) = 1 - \widetilde{\alpha} \geq 1 - \alpha$ and hence:

align[align omitted — 253 chars of source]

The intuition can be seen in the following Figure (ref), which shows the true acceptance region, which corresponds to the acceptance region of a perfect forecast, and the acceptance region of an overdispersed forecast. We can also see that it is sufficient if one side (buy or sell) is provided with an overdispersed forecast.

wrapfigure[wrapfigure omitted — 1,781 chars of source]

The difference in the expected profits (EP) of the returns is given by:

small\begin{align*} \Delta EP &= \left( -\frac{1}{\eta} \mathbb{E}[P_b | P_{b} \leq \widetilde{Q}_b^{1-\alpha}; P_{s} \geq \widetilde{Q}_s^{\alpha}] + \eta\mathbb{E}[P_s | P_b \leq \widetilde{Q}_b^{1-\alpha}; P_{s} \geq \widetilde{Q}_s^{\alpha}] \right) \\ &- \left( -\frac{1}{\eta} \mathbb{E}[P_b | P_{b} \leq {Q}_b^{1-\alpha}; P_{s} \geq Q_s^{\alpha}] + \eta\mathbb{E}[P_s | P_b \leq {Q}_b^{1-\alpha}; P_{s} \geq Q_s^{\alpha}] \right) \\ &= -\frac{1}{\eta} \left( \mathbb{E}[P_b | P_{b} \leq \widetilde{Q}_b^{1-\alpha}; P_{s} \geq \widetilde{Q}_s^{\alpha}] - \mathbb{E}[P_b | P_{b} \leq Q_b^{1-\alpha}; P_{s} \geq Q_s^{\alpha}] \right) \\ &+ \eta\left( \mathbb{E}[P_s | P_{b} \leq \widetilde{Q}_b^{1-\alpha}; P_{s} \geq \widetilde{Q}_s^{\alpha}] - \mathbb{E}[P_s | P_{b} \leq Q_b^{1-\alpha}; P_{s} \geq Q_s^{\alpha}] \right) \leq 0 \end{align*}

since we have

small\begin{align*} &\mathbb{E}[P_b | P_{b} \leq \widetilde{Q}_b^{1-\alpha}; P_{s} \geq \widetilde{Q}_s^{\alpha}] > \mathbb{E}[P_b | P_{b} \leq Q_b^{1-\alpha}; P_{s} \geq Q_s^{\alpha}] \quad and \\ &\mathbb{E}[P_s | P_{s} \geq \widetilde{Q}_s^{\alpha}; P_{b} \leq \widetilde{Q}_b^{1-\alpha}] < \mathbb{E}[P_s | P_{s} \geq Q_s^{\alpha}; P_{b} \leq Q_b^{1-\alpha}] \end{align*}

as the conditioning sets of both expectations are larger for the over-dispersed forecast. Intuitively, we accept more trades and therefore more trades that are less profitable on average, i.e. higher buy prices and lower sell prices. Concluding, the difference in the expected value of the returns is negative. The overall effect therefore depends on the question whether the increase in the AP outweights the decrease in EP

align[align omitted — 321 chars of source]

For low values of $\alpha$, the AP is high and hence the first dominates. For high values of $\alpha$, the AP is low and hence the second dominates. Additionally, for larger spreads between the buy and sell price, the EP is higher and hence the first dominates. For intermediate values of $\alpha$, the effect depends on the underlying price distribution is hard to gauge. We show the implications of this result in a small simulation study: We simulate the true prices as Gaussian random variables with means $\mu_b$ and $\mu_s$. We over- and underdisperse the forecasts by multiplying the true $\sigma$ with a dispersion parameter $b$. We then compute the expected profits for different values of $b$ and $\alpha$. The results are shown in Figure (ref). First, we see that overdispersed forecasts lead mechanically to higher acceptance ratios (top row). However, the effect on the expected profits is not monotone. For high spreads ($\mu_b = 50, \mu_s = 100$) and low volatility ($\sigma = 10$), overdispersion leads to higher profits. For smaller spreads, and higher volatilty, profits decrease for overdispersion and low values of $\alpha$, since we start to accept more unprofitable trades.

figure[figure omitted — 701 chars of source]
remark[Bid acceptance of QBTS and the Dependence Structure of Electricity Prices] Assume that the true and the predictive distribution of the day-ahead electricity prices are continuous and equal (perfect forecast). The probability of acceptance of the bid is crucially dependent on the dependence structure of the electricity prices. Denote $\mathcal{D}(p_0, ..., p_{23})$ as the joint CDF of the electricity prices and $\mathcal{D}_{h}(p_h)$ as the marginal distribution. We have: \begin{align} \mathbb{P}(P_{b} \leq {Q}_b^{1-\alpha}; \; P_{s} \geq Q_s^{\alpha}) & = \mathbb{P}\left(P_{b} \leq {Q}_b^{1-\alpha}\right) - \mathbb{P}\left(P_{b} \leq {Q}_b^{1-\alpha}; \; P_{s} \leq Q_s^{\alpha}\right) \\ & = (1 - \alpha) - \mathcal{D} \left(\mathcal{D}_{b}^{-1}(1-\alpha), \mathcal{D}_{s}^{-1}(\alpha), ... \right) = (1 - \alpha) - \mathcal{C}_{b, s}(1 - \alpha, \alpha) \end{align} where $\mathcal{C}(\cdot, \cdot)$ is the Copula of the price distribution. Employing a copula-based argument here allows the argument to be margin-independent. It should be noted that $\mathcal{C}_{b, s}(1 - \alpha, \alpha)$ is not necessarily easy to compute for arbitrary dependence structures. For the Gaussian Copula $\mathcal{C}_{b, s}(1 - \alpha, \alpha)$ depends only on the pairwise Copula of $b$ and $s$. However, for other dependence structures, we can have an implicit dependence on all variables in between, since: $ \mathcal{C}_{b, s}(1 - \alpha, \alpha) = \int_{[0, 1]^{H-2}} \frac{\partial \mathcal{C}(u)}{\partial u_0 \dots \partial u_H} \text{d}u_{h \neq b, h \neq s}. $

In a second simulation study, we show the implications of this result. We simulate true prices as multivariate Gaussian random variables with means $\mu_b = 90$ and $\mu_s = 100$ and standard deviation $\sigma = 20$ (same as Panel 3 in Figure (ref)). We vary the correlation between $b$ and $s$, taking $\rho = \{0, 0.4, 0.8\}$ in the according covariance matrices $(\Sigma_0, \Sigma_1, \ensuremath{\bm{\mathrm{\Sigma}}}_2)$. Again, we show the AP and the expected profits for different values of $\alpha$ in Figure (ref). The AP decreases for higher correlation. Profits decrease for higher correlation and higher values of $\alpha$. For overdispersed forecasts, the effect is not as strong. This raises two important points: First, in practice, electricity prices are positively correlated across delivery hours, which further weakens the validity of QTBS for forecast evaluation. Second, the positive dependence should potentially decrease the risk of battery trading strategies, however, the QTBS approach does not capture this effect and hence can lead to suboptimal bids.

figure[figure omitted — 343 chars of source]

Although formally more involved to analyze due to the multi-market setting, this results highlights an important consideration for the TS-1/2/3 strategies o2024conformal,o2025conformal,o2025optimising, which directly optimizes for the difference of the quantile forecasts: Treating the hourly quantile forecasts as independent neglects the dependence structure of electricity prices and hence can lead to suboptimal bids.

Battery Trading using Multivariate Probabilistic Forecasts

Let us formalize the notation. We denote with $\ensuremath{\bm{\mathrm{a}}} = (a_0, .., a_H)$ a (fixed) vector of bids.\footnote{ Strictly speaking $\ensuremath{\bm{\mathrm{a}}}$ is a random variable, since the trading decision depend on the forecasted distribution $\mathcal{F}$, which in turn depends on learned coefficients and parameters. Thus, $\ensuremath{\bm{\mathrm{a}}}$ might be biased. We generally treat $\ensuremath{\bm{\mathrm{a}}}$ as deterministic and return to the issue briefly in the following Section (ref). } We have $\ensuremath{\ensuremath{\bm{\mathrm{a}}}^\top}\boldsymbol{P} \sim \mathcal{R}$ as the distribution of the revenues for bid vector $\ensuremath{\bm{\mathrm{a}}}$. Given the forecast $\boldsymbol{F} \sim \mathcal{F}$, usually issued as ensemble or scenario forecast $\widehat{\ensuremath{\bm{\mathrm{F}}}}$, we have $\ensuremath{\ensuremath{\bm{\mathrm{a}}}^\top}\boldsymbol{F} \sim \mathcal{P}^{\ensuremath{\bm{\mathrm{a}}}}$ as the forecasted distribution of the revenues for bid vector $\ensuremath{\bm{\mathrm{a}}}$. We can then write the battery optimization problem as $\max_{\ensuremath{\bm{\mathrm{a}}}} \rho(\mathcal{P}^{\ensuremath{\bm{\mathrm{a}}}})$, where $\rho(\cdot)$ is the risk measure we want to optimize. The realized revenues are given by $r_d = \ensuremath{\ensuremath{\bm{\mathrm{a}}}^\top}\ensuremath{\bm{\mathrm{p}}}_d$.

An Intuitive Optimization Algorithm

We approach the issue from the objective function. Commonly, we want to optimize some kind of risk measure $\rho(\cdot)$ of the returns $$ \max_{b,s} \quad \rho\left(R(\ensuremath{\bm{\mathrm{a}}}(b, s), \mathcal{F})\right) $$ this can be expected profits, or the (conditional) value-at-risk, a mixture of both or some specific functions. We assume that the forecaster equips the trader with a multivariate predictive distribution $\mathcal{F}$, on which decisions are based. $\mathcal{F}$ is commonly represented through a (large) number of samples $\widehat{\ensuremath{\bm{\mathrm{F}}}}$. Intuitively, we can use the following dynamic programming approach:

enumerate[noitemsep] • For each scenario $m = 1, ..., M$ calculate returns for each buy/sell pair $\widehat{\ensuremath{\bm{\mathrm{P}}}}_{bsm} = -\ensuremath{\ensuremath{\bm{\mathrm{a}}}(b ,s)^\top}\widehat{\ensuremath{\bm{\mathrm{F}}}}_m$ • Calculate the discretized version of $\widehat{\ensuremath{\bm{\mathrm{V}}}}_{b,s} = \rho(\widehat{\ensuremath{\bm{\mathrm{P}}}}_{bsm})$ for all $b, s$ pairs, where $b < s$. Taking no action corresponds to $\widehat{\ensuremath{\bm{\mathrm{V}}}}_{0,0} = \rho(0)$. • Choose $b, s = \arg \max (\widehat{\ensuremath{\bm{\mathrm{V}}}}_{b,s})$, define the optimal decision $\ensuremath{\bm{\mathrm{a}}}^* = \ensuremath{\bm{\mathrm{a}}}(b^*, s^*)$ and place unrestricted bids.

It is straightforward to see that, in this setting, we can directly optimize any risk measure. However, the approach is computationally demanding and does not scale well in the number of bids placed and the relative size of the bids. In the following Section (ref), we discuss a more scalable approach based on mixed-integer linear programming. Nevertheless, this mental model provides a clear and intuitive understanding of how to employ multivariate probabilistic forecasts for battery optimization, can serve as a straightforward benchmark and highlights the importance of the dependence structure. Figure (ref) visualizes the approach.

remark[Connection to Markowitz Portfolio Optimization] Note that the battery optimization can be seen as a long-short portfolio management problem with specific constraints on the weights, that is, the sum of the weights needs to be zero and the absolute weight of each asset is constrained by the battery capacity and depends on the order of assets (hours). In the classical Markowitz mean-variance optimization markowitz1952portfolio,markowitz1952utility, the (random) returns are commonly denoted as $\boldsymbol{r} \sim \mathcal{N}(\ensuremath{\bm{\mathrm{\mu}}}, \ensuremath{\bm{\mathrm{\Sigma}}})$ and the portfolio weights as $\ensuremath{\bm{\mathrm{w}}}$, which corresponds to the (random) prices $\boldsymbol{P}$ and the bid vector $\ensuremath{\bm{\mathrm{a}}}$ in our setting kan2007optimal.

This connection to portfolio optimization allows us to discuss two interesting points.

itemize• A somewhat counterintuitive property of QBTS is described in Result (ref): The expected return of QBTS decreases with increasing correlation between the day-ahead prices. On the contrary, for long-short portfolio optimization, positive correlation between the positions is desirable, since it reduces the variance of the portfolio returns and hence increases risk-adjusted returns. QBTS is not able to exploit this property. • It would be interesting to further analyze the distribution of the optimal bid vector $\ensuremath{\bm{\mathrm{a}}}^*$, which itself is a random variable. This is analogous to the distribution of the optimal portfolio weights in the context of portfolio optimization, which has been analyzed in the context of mean-variance optimization kan2007optimal. However, the distribution of the optimal bid vector $\ensuremath{\bm{\mathrm{a}}}^*$ is more complex, since the optimization problem is non-convex and the solution space is not continuous. Hence, we leave this analysis for future research.
figure[figure omitted — 2,142 chars of source]

Formulation as Mixed-Integer Linear Problem

While the dynamic programming approach is intuitive, it does not scale well. Hence, we propose an alternative formulation as a mixed-integer linear programming problem. A drawback of this approach is that only the expected value and the conditional value at risk can be optimized directly. Mean-variance and value-at-risk objective functions cannot be reformulated as MILP problems rockafellar2000optimization, uryasev2001conditional. We will discuss this difference in the simulation study in Section (ref).

We again assume that we have $M$ scenarios $\widehat{\ensuremath{\bm{\mathrm{F}}}}_{h,m}$ for each hour $h$ and scenario $m$. We further relax some assumptions from the section before and allow for multiple buy and sell bids. Denote the number of buy bids as $N_b$ and the number of sell bids as $N_s$. We treat the bid volume as decision variable $a_h^b > 0$ and $a_h^s >0$. We have

align[align omitted — 1,101 chars of source]

where $\kappa$ is the storage capacity and $\xi$ is the charge/discharge capacity (also called energy or rated capacity). In light of the recent market change to quarter-hourly trading in the single day-ahead coupling (SDAC), the use of such dynamic programming algorithms becomes computationally demanding and as such, the use of mixed-integer linear programming (MILP) solvers is advisable. Our implementation uses pyomo and the HIGHS solver hart2011pyomo,huangfu2018parallelizing.

Battery Optimization as Scoring Rules

In the following, we take the above formulation and discuss whether the multivariate probabilistic forecast can be scored through the battery optimization. Using a simple example, we show that it is possible to construct different forecast distributions that lead to the same objective value $\rho(\cdot)$ for the battery optimization. This raises the issue of discriminatory power of battery optimization for the evaluation of multivariate probabilistic forecasts.

The reason for this issue is the number of degrees of freedom within the battery optimization problem. In the classic forecast evaluation setting, we are interested in ranking a number $\mathcal{M}$ different models $h_m(\ensuremath{\bm{\mathrm{X}}}) \rightarrow \mathcal{F}_m$ based on their forecasted distribution $\mathcal{F}_m$ and the realized values $y$ using scores $s_m = S(\mathcal{F}_m, \ensuremath{\bm{\mathrm{p}}})$. In the battery optimization case, we rank different models based on the outcome of an optimization problem that takes the forecasted distribution as input. Therefore, we have

equation[equation omitted — 241 chars of source]

where $\ensuremath{\ensuremath{\bm{\mathrm{a}}}^\top} \boldsymbol{F}_m \sim \mathcal{P}^{\ensuremath{\bm{\mathrm{a}}}}_{m}$ is the forecasted distribution of the revenues for model $m$ and $R(\ensuremath{\bm{\mathrm{a}}}_m^*, \ensuremath{\bm{\mathrm{p}}})$ is the realized revenues for the optimal bid $\ensuremath{\bm{\mathrm{a}}}_m^*$ and the realized prices $\ensuremath{\bm{\mathrm{p}}}$. We can then rank the models based on $s_m$ and $r_m$, respectively, which raises the question of whether the ranking of the models based on $s_m$ and $r_m$ is equivalent.

itemize• From a decision-maker's perspective, it is desirable that both rankings coincide, as this ensures that the best model according to the forecast evaluation is also the best model for the decision-making problem at hand. However, this equivalence is not guaranteed, as we will show in the following. • Often, the ultimate goal is the application, so the ranking based on $r_m$ seems more relevant. However, ranking based on $s_m$ is more general, is based on well-established theory and allows for a more comprehensive evaluation of the forecast quality across different applications. It is reasonable to assume that a forecast that scores well in terms of $s_m$ will also perform well in terms of $r_m$ across a range of decision-making problems.

In the following, we will show through a simple example that different forecast distributions can lead to the same ranking of bids in the battery optimization problem, hence placing a fundamental limit on forecast evaluation through battery optimization.

prop[Battery Trading Strategies are not strictly proper scoring rules] Battery trading strategies compress the distributional information. Therefore, different forecasts can indistinguishably lead to the same optimal bids, given the risk measure $\rho(\cdot)$ employed in the battery optimization. Assume $\mathcal{F}_1$ is a perfect forecast, then there can exist different forecast distributions $\mathcal{F}_2$, where $\mathcal{F}_1 \neq \mathcal{F}_2$, such that\begin{equation} \rho(\mathcal{P}^{\ensuremath{\bm{\mathrm{a}}}}_{1}) = \rho(\mathcal{P}^{\ensuremath{\bm{\mathrm{a}}}}_{2}) \end{equation} and the ranking of all pairs $(b,s)$ does not change. Hence, the battery optimization using multivariate probabilistic forecasts cannot be a strictly proper scoring rule $\square$

From a practical perspective, this implies that economic backtest performance can be decision-relevant, but it is generally unsuitable as a stand-alone strictly proper scoring rule for model selection over full predictive distributions. We illustrate the issue of information compression through the battery optimization problem in the following examples.

example[Misspecified Mean Structure] A simple example can be constructed for optimizing the expected revenues $\rho(\cdot) = \mathbb{E}[\cdot]$ with relatively few assumptions. We have a perfect forecast $\mathcal{F}_1$ with $\ensuremath{\bm{\mathrm{\mu}}}_1$ and a second forecast $\mathcal{F}_2$ with $\widehat{\ensuremath{\bm{\mathrm{\mu}}}}$. For a 1-hour battery, we have the optimal bids $(b, s)$ where \begin{equation} b = \arg \min_h \ensuremath{\bm{\mathrm{\mu}}}_1 \quad and \quad s = \arg \max_h \ensuremath{\bm{\mathrm{\mu}}}_1 \end{equation} that is, we want to buy in the cheapest hour and sell in the most expensive hour. Conversely, we can have an arbitrary forecast $\mathcal{F}_2$ with \begin{equation} \ensuremath{\bm{\mathrm{\mu}}}_{2,s} = \ensuremath{\bm{\mathrm{\mu}}}_{1,s} \quad and \quad \ensuremath{\bm{\mathrm{\mu}}}_{2,b} = \ensuremath{\bm{\mathrm{\mu}}}_{1,b} \end{equation} and arbitrary values $\ensuremath{\bm{\mathrm{\mu}}}_{1,b} < \ensuremath{\bm{\mathrm{\mu}}}_{2,h} < \ensuremath{\bm{\mathrm{\mu}}}_{1,s} \; \forall h \in \{1, \ldots, 24\} \setminus \{b,s\}$. This forecast will lead to the same optimal bids and hence the same expected revenues, despite being a different forecast. The formulation can be extended to ranks, i.e. if the ranking of all hours is the same, the expected revenues will be the same. This allows us to construct forecasts that are arbitrarily far away from the true distribution, but still score perfectly and highlights the issue of information compression in the battery optimization example. This is further illustrated in Figure (ref), which shows different forecasts that lead to the same expected revenues.
figure[figure omitted — 772 chars of source]
example[Misspecified Covariance Structure] Assume the true and forecasted distribution of prices follow a multivariate normal distribution $\ensuremath{\bm{\mathrm{p}}} \sim \mathcal{N}(\ensuremath{\bm{\mathrm{\mu}}}, \ensuremath{\bm{\mathrm{\Sigma}}})$ and $\mathcal{N}(\widehat{\ensuremath{\bm{\mathrm{\mu}}}}, \widehat{\ensuremath{\bm{\mathrm{\Sigma}}}})$ with potentially misspecified marginal variances and correlation structure. For a bid pair $(b, s)$, the profit distribution is given as \begin{eqnarray} \mathcal{P} = \mathcal{N}(\mu_R, \sigma^2_R) & \mu_R = \eta^2 \kappa(\mu_s - \mu_b) & \sigma^2_R = (1/\eta)^2 \kappa^2\sigma^2_b + \eta^2 \kappa^2 \sigma^2_s + 2 \kappa^2\rho_{b,s}\sigma_b \sigma_s\\ \widehat{\mathcal{P}}= \mathcal{N}(\widehat{\mu}_R, \widehat{\sigma}^2_R) & \widehat{\mu}_R = \eta^2 \kappa(\widehat{\mu}_s - \widehat{\mu}_b) & \widehat{\sigma}_R^2 = (1/\eta)^2 \kappa^2\widehat{\sigma}^2_b + \eta^2 \kappa^2 \widehat{\sigma}^2_s + 2 \kappa^2\widehat{\rho}_{b,s}\widehat{\sigma}_b \widehat{\sigma}_s \end{eqnarray} Trivially, for the perfect forecast, the forecasted and true revenue distribution are equal. For calibrated mean forecasts, we have $\widehat{\mu}_R = \mu_R$. Focusing on the variance, we have: \begin{align} \sigma_R^2 - \widehat{\sigma}_R^2 &= (1/\eta)^2 \kappa^2\sigma^2_b + \eta^2 \kappa^2 \sigma^2_s - 2 \kappa^2\rho_{b,s}\sigma_b \sigma_s - \left((1/\eta)^2 \kappa^2\widehat{\sigma}^2_b - \eta^2 \kappa^2 \widehat{\sigma}^2_s - 2 \kappa^2\widehat{\rho}_{b,s}\widehat{\sigma}_b \widehat{\sigma}_s\right) \\ \nonumber &= (1/\eta)^2 \kappa^2(\sigma^2_b - \widehat{\sigma}^2_b) + \eta^2 \kappa^2 (\sigma^2_s - \widehat{\sigma}^2_s) - 2 \kappa^2 (\rho_{b,s}\sigma_b \sigma_s - \widehat{\rho}_{b,s}\widehat{\sigma}_b \widehat{\sigma}_s) \end{align} setting the right-hand side of Equation (ref) to zero gives us a relationship between the forecasted variances and correlation such that the variance of the profit distribution is identical between true and forecasted distribution: \begin{equation} \widehat{\sigma}_b^2 = \frac{ (1/\eta)^2\sigma^2_b + 2 \sigma^2_s \widehat{\sigma}_s^2 \rho_{b,s} - \eta^2 (\sigma^2_s - \widehat{\sigma}^2_s) }{ (1/\eta)^2 - 2\rho_{b,s}\sigma^2_b } \end{equation} which allows us to construct different forecast distributions with identical risk measure outcomes for the battery optimization, assuming that the differences between the true and predicted variances are sufficiently small such that the ranking of all pairs $(b,s)$ does not change.

These two examples already hint at potentially low discriminatory power of the battery optimization for forecast evaluation. It is important to be aware of this issue when evaluating forecasts through battery optimization. For optimizing the expected revenues, the issue has been discussed empirically by maciejowska2025statistical. The main risk in the application is that, for slight changes in the application, a forecast that has previously performed well cannot be expected to perform well anymore. For example, a forecast that has performed well for a 1-hour battery might not perform well for a 2-hour or 4-hour battery, since the optimal bids can change and hence, the relevant part of the forecast distribution can change.

remark[Relationship to Weighted Scoring Rules] The optimization can be seen as a censoring of the forecast distribution. allen2023weighted and de2025localizing discuss ways to create (strictly proper) scoring rules focused on specific parts of the distribution through weighted scoring rules. For example, de2025localizing define the region of interest as $A_w = \{y \in \mathcal{Y} : w(y) > 0\}$, where $w(\cdot)$ is a weighting function. However, the weighting function introduced by the battery optimization does not depend on the outcome variable alone, but also on the forecast distribution itself, which is not covered by the existing literature on weighted scoring rules.

Based on these results, two extreme points of view can be taken:

itemize• One only cares about the decision-making problem at hand and hence, the ranking based on the application outcome is the only relevant one. In this case, one can even argue that the forecast ranking based on statistical measures is obsolete and the predictive modeling step can be fully integrated into the decision-making problem, e.g. through end-to-end learning approaches donti2017task, elmachtoub2022smart. This argument is strongest if the application is stable. If the application is subject to change, e.g. due to changes in the asset portfolio, a general-purpose forecast is more desirable and flexible. • One focuses on delivering the best general-purpose forecast in terms of a general scoring rule $S(\cdot, \cdot)$. There is a strong case that for strictly proper scoring rules, a forecast that performs well in terms of $s_m$ will also perform well in terms of $r_m$ across a range of decision-making problems and will be robust to changes in the application.

There is a wide gray area in between these two extreme points of view. A plausible option is using weighted combinations of $s_m$ and $r_m$ to rank models. If rankings coincide, this gives us more confidence in the model choice and employ additional model diagnostics. Somewhat turning things around, it can be argued that this reliance on few parts of the distribution is advantageous. If we need only a certain part of the distribution to be (approximately) correct to create the optimal bids, we can utilize this for reducing the dimensionality of the problem. This can be hard to achieve a priori, as the optimal bids $\ensuremath{\bm{\mathrm{a}}}_m^*$ are not known. In hindsight, we can analyze which parts of the forecast distribution were relevant for the decision-making problem and focus on improving these parts in future forecasts.

Evaluation of Decision Quality

In the following, we discuss the evaluation of decision quality given competing forecasts. It is desirable that the evaluation of the decision quality is consistent with the objective function employed in the optimization.

itemize• This ensures that the ranking is consistent with the decision-making problem at hand. However, this requires the risk measure to be elicitable itself. For example, the CVAR is not elicitable on its own, but only jointly with the value-at-risk fissler2016higher,ziegel2020robust,fissler2015expected. • Under different competing forecasts $\mathcal{F}_m$, the optimal bids $\ensuremath{\bm{\mathrm{a}}}_m^* = \arg \max_{\ensuremath{\bm{\mathrm{a}}}} \rho(\mathcal{F}_m)$ might differ. This adds another layer of complexity to the evaluation of the forecasts based on battery optimization. This is in contrast to, e.g. the backtesting CVAR forecasts of GARCH models, where the implicit assumption is that the forecaster is long/short a fixed position in the asset.

Therefore, the predicted objective value of the battery optimization for the forecast model $m$ as $\omega^*_m = \rho(\mathcal{P}^{\ensuremath{\bm{\mathrm{a}}}_m^*}_m)$ and the realized returns of $r_m = \ensuremath{(\ensuremath{\bm{\mathrm{a}}}_m^*)^\top} \ensuremath{\bm{\mathrm{p}}}$ cannot be evaluated across different models. For a strictly proper scoring rule $S(\cdot, \cdot)$ of $\rho(\cdot)$ we have

align[align omitted — 210 chars of source]

which shows that the forecast optimal bids of forecast $\mathcal{F}_m$, which we would like to evaluate, enters the scoring rule on both sides of the equation. As possible workaround, for a number of models $m \in \mathcal{M}$, we can evaluate the scoring rule for all optimal bids $\ensuremath{\bm{\mathrm{a}}}_m^*$ and all forecasts $\mathcal{F}_m$ and analyze the resulting scores. If the forecast $\mathcal{F}_m$ has the lowest score of all models for the risk measure for bids $\ensuremath{\bm{\mathrm{a}}}_m^*$, then this should a good indication that actually this decision is good. If many/all models $i = 1, ... ,M; i \ne{} m$ have better scores on the objective value $\omega_m$ than $\mathcal{F}_m$, this is an indication that the decision $\ensuremath{\bm{\mathrm{a}}}_m^*$ is not globally optimal. This also gives rise to a formal test by comparing the scores of the models $i$ with the scores of model $m$ for the optimal bids $\ensuremath{\bm{\mathrm{a}}}_m^*$, which can be done through a Diebold-Mariano test diebold2002comparing, diebold2015comparing. We test for the hypothesis

align[align omitted — 232 chars of source]

If we can reject the null hypothesis, this implies that model $i$ was able to yield significantly better predictions of the objective value for this decision and hence an indication that the decision $\ensuremath{\bm{\mathrm{a}}}_m^*$ based on forecast $\mathcal{F}_m$ should not be trusted. Importantly, this only addresses one side of the optimization problem, since we are only evaluating how well the forecast $\mathcal{F}_m$ performs for the optimal bids $\ensuremath{\bm{\mathrm{a}}}_m^*$ relative to the other forecasts for the hours in which $m$ places bids, which is only a subset of hours. The predictive accuracy in the other hours is not evaluated, even though the predictive performance in these hours is relevant for the decision-making problem.

Forecasting and Simulation Study

In the following section we run a small forecasting experiment and analyze the impact of different forecast errors on the results of the optimization. We employ seven different forecasting models with different modeling strategies and distributional assumptions. These models range from the simple climatology and naive models, the lasso-estimated autoregressive model lago2021forecasting and distributional regression approaches muniain2020probabilistic, hirsch2024online,hirsch2025online. It is important to note that we have preferred models that provide either an option for bootstrapping or a parametric distributional assumption in order to draw samples for the stochastic program. For this reason, we have not included quantile regression or conformal prediction approaches uniejewski2025smoothing,o2025conformal. As the main aim of the case study is to relate statistical measures and economic evaluation on a diverse pool of models, we acknowledge that we have not undertaken extensive feature engineering or hyperparameter optimization, but chose a configuration that has been shown to work well in previous work.

Data and Market

The day-ahead market is the primary exchange for electricity in Germany. The market is organized as a uniform price auction, where the price for each hour of the next day is determined by the intersection of the aggregated supply and demand curves. The market is cleared at 12:00 CET for the delivery of electricity on the next day. We use data from ENTSO-E and \url{investing.com} provided by lipiecki2024postprocessing. The data covers the period from 2015 to 2023 and includes the day-ahead prices, the residual load, the marginal costs of gas, coal and oil fired production and the associated emission allowance costs. We use 2015-01-15 to 2018-12-26 for model training and 2018-12-27 to 2023-12-31 for testing, in line with other works using the same data set hirsch2024online,hirsch2025online,marcjasz2023distributional.

Forecasting Model

We employ a distributional regression model or generalized additive model for location, scale and shape rigby2005generalized based on an extended expert model using Student's $t$-distribution for the marginal distribution of the electricity prices $$P_{d,h} \sim t(\mu_{d,h}, \sigma_{d,h}, \nu_{d,h})$$ Distributional models are popular for electricity price forecasting serinaldi2011distributional, abramova2020forecasting, muniain2020probabilistic, klein2023deep,marcjasz2023distributional. We extend the well-known LASSO-estimated autoregressive model with exogenous variables lago2021forecasting model by non-linear effects representing the merit-order effects of residual load with the marginal costs (MC) of gas, coal and oil fired production and the associated emission allowance costs.

align[align omitted — 1,314 chars of source]

where $\eta^\ensuremath{\text{Gas}} = 0.2, \eta^\ensuremath{\text{Coal}} = 0.35$ and $\eta^\ensuremath{\text{Oil}} = 0.3$ are emission factors and $\eta$ is a conversion factor between the fuel's unit and MWh thermal energy ghelasi2025data, ghelasi2025day and $f(\cdot)$ denotes a b-Spline basis of degree 2 and 4 knots. Weekday effects are denoted with $\mathcal{W} = \{\text{Mon}, \text{Tue}, \text{Thu}, \text{Fri}, \text{Sat}, \text{Sun}, \text{Hol}\}$; holidays are encoded separately ziel2018modeling. For the scale parameter $\sigma_{d,h}$ of the price distribution, we employ a similar structure, but with linear equations for residual load and the fundamental fuel prices. The degrees of freedom for the $t$-distribution are estimated as intercept. The equations can be found in the Appendix (ref). We name the model distributional LASSO-estimated non-linear autoregressive model (DLENAR). The model is estimated in a expanding window fashion using the online learning algorithm developed in hirsch2024online to save computational time. We estimate separate models for each hour of the day. We model the dependence structure using a Gaussian Copula by the inference-for-margins approach. We convert the in-sample observations to the $\mathcal{U}(0, 1)$-space by the probability integral transformation (PIT) and subsequently to the $\mathcal{N}(0, 1)$ space by the inverse quantile transformation $\ensuremath{\bm{\mathrm{g}}} = \Phi^{-1}(\ensuremath{\bm{\mathrm{u}}})$. On the Gaussian space, we fit the correlation matrix.\footnote{Note that we additionally employ a ranking step to enforce strict uniformity on the residuals. This step does not alter the ordering of $\ensuremath{\bm{\mathrm{u}}}$, but the conversion to ranks makes observations equally spaced.} We fit the three different models for the dependence structure:

itemize• Independent (IND): We assume independence $\ensuremath{\bm{\mathrm{\Omega}}} = \ensuremath{\bm{\mathrm{I}}}$. • Empirical (DEP): We use the empirical correlation matrix $\ensuremath{\bm{\mathrm{\Omega}}} = \widehat{\operatorname{corr}}(\ensuremath{\bm{\mathrm{g}}})$. • Weekday (DWD): We estimate separate correlation matrices for each weekday $\ensuremath{\bm{\mathrm{\Omega}}}_{\text{WD}} = \widehat{\operatorname{corr}}(\ensuremath{\bm{\mathrm{g}}} | \text{WD})$.

We draw scenarios by sampling from the uniform distribution and applying a rank-reordering to ensure that the marginal distributions are exactly equal across the different dependence structures. Additionally, we employ a number of benchmark models from the literature to obtain a diverse set of models.

itemize• Climatology: We take a random sample from the in-sample observations as forecast for the next day. This is a common benchmark model for probabilistic forecasts gneiting2007strictly. • Naive Bootstrap (Naive-BS): The take the previous week, same weekday as the price forecast and sample from the in-sample residuals (of the same weekday). This is common benchmark model for seasonal time series such as electricity prices ziel2018modeling. • The widely used LASSO-estimated autoregressive model lago2021forecasting model with Gaussian errors $\mathcal{N}(0, \Sigma)$ estimated on the residuals of the mean model. The model is denoted as LEAR-N(0, $\Sigma$). • The LEAR model with bootstrap residuals: we sample from the in-sample residuals of the mean model. We somewhat improve the bootstrap approach by clustering the days and predicting cluster membership by a $k$-nearest neighbor classifier cunningham2021k

For all bootstrap approaches, we always sample full error trajectories to ensure that we (implicitly) capture the dependence structure of the errors. Model equations and further details can be found in the Appendix (ref).

Battery Optimization Experiments

We run a number of battery optimization experiments using different risk measures and correlation structures. Our experiments can largely be grouped in three categories. Generally, our experiments assume a storage of 10 MWh, a round-trip efficiency of $\eta^2 = 0.95^2$. The duration of the storage, that is, the ratio between storage capacity and maximum charge/discharge power is set to 1 hour unless otherwise stated. For larger battery capacities, we would need to take market impact into account narajewski2022optimal, which is beyond the scope of this paper.

table[table omitted — 716 chars of source]

\paragraph{Multivariate Probabilistic Scoring Rules and Battery Revenues} In this experiment, we evaluate the relationship between multivariate probabilistic scoring rules and battery revenues. We discuss the results in light of Result (ref) and analyze the relationship between the scoring rules, decision quality and battery revenues. We run the battery optimizations maximizing the expected profits and the CVAR at levels $\alpha \in \{0.90, 0.75, 0.50\}$ using the MILP approach for a 1, 2 and 4 hour battery. The results of this experiment are discussed in Section (ref).

\paragraph{Multivariate Forecasts and Optimization Algorithms} In Sections (ref) and (ref) we have described an intuitive dynamic programming approach and a more scalable mixed-integer linear programming (MILP) approach to optimize battery trading strategies using multivariate probabilistic forecasts. Here, we briefly evaluate the impact of placing only one buy and sell pair compared to the MILP approach allowing for multiple bids. To this end, we compare results for a 1-hour battery, maximizing the expected profits and the CVAR at levels $\alpha \in \{0.90, 0.75, 0.50\}$ using both optimization approaches. The results of the experiment are discussed in Section (ref).

\paragraph{Quantile-based Trading Strategies} We evaluate the quantile-based trading strategy (QBTS) as described in Section (ref) using different forecast models. We evaluate the total profits, profits per MWh traded and the acceptance probability for different nominal prediction interval widths $\alpha$. Here, we focus providing an intuitive understanding of the Proposition (ref) and Remark (ref) through numerical examples and contrast the results with the stochastic programming approach. The results of this experiment are discussed in Section (ref).

Economic Performance Measures

We evaluate the economic performance of the different forecast models based on the daily returns from the battery optimization experiments. The daily return for day $d$ and model $m$ is defined as $r_{d,m} = \sum_{h=0}^H a_{d,h,m} P_{d,h}$ where $a_{d,h,m}$ are the optimal actions derived from the battery optimization using forecast model $m$ for day $d$. For daily returns $r_{d, m}$ of $m$ and day $d$ we calculate the total profits and the Sharpe ratio sharpe1966mutual as

equation[equation omitted — 178 chars of source]

For the CVAR optimization, we also evaluate the ratio of Value-at-Risk (VaR) exceedances. The VaR at level $\alpha$ is defined as the ($1-\alpha$) -- quantile of the return distribution, that is, $\text{VaR}_\alpha = \inf \{x \in \mathbb{R} : P(r_{d,m} \le x) \ge \alpha\}$. The VaR exceedance ratio is then defined as

equation[equation omitted — 122 chars of source]

This can be both seen as a measure of the economic performance, and as a measure of the calibration of the CVAR forecasts, as for a well-calibrated CVAR forecast, the VaR exceedance ratio should be close to $1-\alpha$.

Forecast Evaluation

We evaluate the forecast models using standard probabilistic scoring rules and economic metrics derived from the battery optimization experiments. We follow the general principle of sharpness subject to calibration gneiting2007probabilistic. The following sections provide an overview of the measures of calibration and scoring rules. We denote the true value with $\ensuremath{\bm{\mathrm{y}}}$ and the forecasted distribution with $\mathcal{F}$, represented by an ensemble of $M$ scenarios $\widehat{\ensuremath{\bm{\mathrm{F}}}} \sim \mathcal{F}$ for $m = 1, \ldots, M$. We denote the mean and median forecasts as $\widehat{\ensuremath{\bm{\mathrm{\mu}}}}$ and $\bar{\ensuremath{\bm{\mathrm{\mu}}}}$. The data set of prices has $D$ days and $H$ hours per day and accordingly we aggregate by averaging. We test the statistical significance of differences in the scores by a Diebold-Mariano test diebold2002comparing, diebold2015comparing, nowotarski2018recent.

Scoring Rules for Continuous (Multivariate) Distributions

A calibrated probabilistic forecast correctly represents the probabilities of the observed data. For univariate and quantile forecast, marginal calibration can be easily checked by comparing expected and observed frequencies.

equation[equation omitted — 136 chars of source]

where $\hat{Q}_\alpha$ is the predicted quantile at level $\alpha$. For a well-calibrated forecast, the marginal calibration should be close to zero for all $\alpha$. Visually, this can be checked by plotting the marginal calibration curve, which plots the observed frequency of the event $\ensuremath{\bm{\mathrm{y}}}_t < \hat{Q}_t^{\alpha}$ against the predicted quantile level $\alpha$. For a well-calibrated forecast, the marginal calibration curve should be close to the diagonal.

The following scoring rules are used throughout the empirical evaluation. The MSE, MAE, CRPS and ES are strictly proper scoring rules for the mean, median, univariate and multivariate probabilistic forecasts. The DSS is only (strictly) proper for the mean-covariance structure of the predictive distribution. The VSS is not a strictly proper scoring rule. We keep the discussion of these scoring rules brief, as they are commonly used in PEPF. We have:

equation[equation omitted — 330 chars of source]
equation[equation omitted — 276 chars of source]
equation[equation omitted — 276 chars of source]
equation[equation omitted — 238 chars of source]
equation[equation omitted — 287 chars of source]

where $\widehat{\ensuremath{\bm{\mathrm{\mu}}}}$ and $\widehat{\ensuremath{\bm{\mathrm{\Sigma}}}}$ are the sample mean and covariance matrix of the ensemble forecasts $\widehat{\ensuremath{\bm{\mathrm{F}}}}$. The CRPS and the ES are estimated using the energy estimator gneiting2007strictly,zamo2018estimation. For further information on the scoring rules we refer to the original papers scheuerer2015variogram,dawid1999coherent,gneiting2007strictly and recent works on discriminatory power of scoring rules pinson2013discrimination, marcotte2023regions, alexander2024evaluating and potential asymmetric effects buchweitz2025asymmetric.

A particular issue is the evaluation of CVAR forecasts, as the CVAR is not elicitable on its own, but only jointly with the value-at-risk fissler2016higher, fissler2015expected. For the evaluation of CVAR forecasts, we therefore use the strictly proper joint scoring rule for the (VaR, CVAR) proposed by fissler2015expected. For a forecast of the VaR and CVAR at level $\alpha$, denoted as $(v, e)$, the scoring rule is defined as

equation[equation omitted — 248 chars of source]

where $G_1(v) = v$, $G_2(e) = \exp(e)/(1 + \exp(e))$ and $\mathcal{G}'_2(e) = G_2(e)$ (see fissler2015expected for details). The first term corresponds to the pinball score for the VaR forecast, while the second and third term correspond to the scoring rule for the CVAR forecast.

For the case of point forecasting and battery trading strategies, the use of correlation measures has been proposed as scoring rules by maciejowska2025statistical. For probabilistic forecasts, we propose the using a association-based scoring rule based on Kendall's $\tau$ kendall1938new, defined as

equation[equation omitted — 366 chars of source]

where $\tau(\ensuremath{\bm{\mathrm{v}}}, \ensuremath{\bm{\mathrm{w}}})$ is the Kendall's $\tau$ correlation between two vectors $\ensuremath{\bm{\mathrm{v}}}$ and $\ensuremath{\bm{\mathrm{w}}}$. Note that we can calculate the score either on the price ensembles or on the rank ensembles, since Kendall's $\tau$ is invariant under strictly increasing transformations. Kendall's $\tau$ also relates to the dependence structure of the forecast's copula and thus provides a natural way to evaluate the dependence structure gijbels2011conditional,ziel2019multivariate.

propThe scoring rule proposed in Equation (ref) is proper. The construction follows the kernel score framework gneiting2007strictly, which requires that the kernel is a positive definite kernel. jiao2015kendall show that Kendall's $\tau$ is a positive definite kernel $\square$

maciejowska2025statistical propose two BESS oriented scores for point forecasts: The mean profit deviation (MPD) and the mean hour deviation (MHD), defined as

equation[equation omitted — 234 chars of source]

where $h_\text{min}$ and $h_\text{max}$ are the hours with the lowest and highest price. The MPD can be seen as a simplification of the returns of a 1-hour battery, relaxing the assumption of charging before discharging. The MHD is closely related to rank-based scoring rules, which we will treat in the following section.

Scoring Rules for Rankings

We denote the ranks of the observed prices $\ensuremath{\bm{\mathrm{p}}}_d$ as $\ensuremath{\bm{\mathrm{\pi}}}_d$ and the rank forecast derived from the predicted ensemble forecast $\widehat{\ensuremath{\bm{\mathrm{F}}}}_d$ as $\widehat{\ensuremath{\bm{\mathrm{\Pi}}}}_d$. We therefore have a predictive distribution over the ranks of the observed prices. Note that we have as many ranks as hours, i.e. $K = H + 1$. In the marginal case, rank based scoring rules can take two views:

itemize• (Item view) For given hour $h$, did we correctly predict the rank $\pi_{h}$ of the observed price $\ensuremath{\bm{\mathrm{y}}}_{h}$? • (Rank view) For given rank $k$, did we correctly predict which hour $h$ corresponds to this rank?

In the following, we discuss scoring rules for the item and rank view and propose a new scoring rule for the joint distribution of the ranks based on the kernel scoring rule framework. The Brier score for the rank and item view is defined as

align[align omitted — 501 chars of source]

Note that is formulation corresponds to the multi-class Brier score, which is bounded in [0, 2]. For the marginal distribution of the ranks (item view), the ranked probability score (RPS) is a strictly proper scoring rule epstein1969scoring. The RPS is defined as

equation[equation omitted — 247 chars of source]

Again, we refer the reader to the original papers for further details on the scoring rules gneiting2007strictly,epstein1969scoring and recents works on the properties of rank scoring rules constantinou2012solving,du2021beyond, wheatcroft2021evaluating.

It is common to look at the performance of ranking measures for the top-$k$ and bottom-$k$ ranks. In the context of battery trading strategies, seems natural to define the relevant top-$k$ and bottom-$k$ as the product of the duration of the battery and the number of (ceiled) cycles per day. For example, for a 2-hour battery with a with one cycle per day, we would be interested in the top-2 and bottom-2 ranks. For a 4-hour battery, we would be interested in the top-4 and bottom-4 ranks -- the same as we would be for a 2-hour battery with two cycles per day. It is important to note that in this case, we have to consider the relevant structure of the battery optimization problem, that is, we can only discharge after charging and are thus interested in the top-$k$ ranks for the hours following the bottom-$k$ ranks. We therefore evaluate the top-$k$ and bottom-$k$ ordered by ranks, that is, we look at the top-$k$ which are preceded by at least $k$ hours with the bottom-$k$ ranks, which we denote as the BESS-$k$ scores.

Results

This section presents the results. We discuss the results of the statistical forecast evaluation and subsequently the results of the economic evaluation. To make matters easier for the reader, we consistently use the color scheme for the for tables/heatmaps and for the models.

Statistical Forecast Evaluation

This section briefly discusses the results of the forecast evaluation in terms of the scoring rules. Table (ref) shows the scores for the different forecast models and scoring rules. Table (ref) shows the top-$k$ Brier scores for the different forecast models and different definitions of top-$k$. Figure (ref) shows the Brier scores by rank and hour. Marginal calibration is shown in the left panel of Figure (ref), while the right panel of Figure (ref) shows the marginal calibration curve.

table[table omitted — 5,376 chars of source]
figure[figure omitted — 308 chars of source]

Generally, the results follow an expected pattern, with the distributional regression model with the fitted dependence structure (DLENAR-DEP) performing best across all scoring rules, followed by the distributional regression model with the fitted dependence structure (DLENAR-DWD) and the distributional regression model with independence (DLENAR-IND). The LEAR-based benchmark models follow, with the LEAR-N(0, $\Sigma$) performing slightly better than the LEAR-BS. The climatology and Naive-BS perform worst. For the DLENAR-IND and the DLENAR-DWD and DLENAR-DEP, the difference between the ES is not as pronounced as for the other scoring rules, a potential indication that the ES is less sensitive to the dependence structure than other multivariate scoring rules alexander2024evaluating, as the marginal model for these three is the same.

table[table omitted — 6,870 chars of source]

For the rank-based scores, the RPS and KS scores show a similar pattern as the ES, VS and CRPS. The Brier score shows a somewhat different pattern, as it yields worse scores for the Naive-BS than for the climatology model. This difference can be explained by the fact that the Naive-BS takes last weeks value as the forecast, while the climatology model yields an average over the training set. The predictions of the Naive-BS are hence somewhat more arbitrary, since the weather conditions between this week and last week can change drastically, while the climatology model captures at least the seasonal average. Figure (ref) shows the Brier scores by rank and hour. The scores decrease quite steeply for the lowest ranks and then level off. For the hour/item view, we see that performance is rather stable over the hours, while hour 19 seems to be the easiest to predict, presumably since it is naturally the most expensive hour. For the top-$k$ scores in Table (ref), we see that the BESS-$k$ scores are worse than the top-$k$ and bottom-$k$ scores, which is in line with the intuition that the middle ranks are harder to predict than the lowest and highest ranks. The differences between the different forecast models are not that pronounced, but we can see that the DLENAR-DEP-WD and DLENAR-DEP models perform better than the DLENAR-IND model, which is in line with the results for the other scoring rules.

figure[figure omitted — 261 chars of source]

Economic Forecast Evaluation

In the following, we present an economic evaluation based on multivariate probabilistic forecasts. Given the fact that risk-neutral optimization approaches do not profit from probabilistic forecasting, we focus our discussion on the risk-averse case. As in the discussion of simplified BESS optimization before, we report the total profits and the Sharpe ratio for the combination of the battery configurations and forecast models. These results are shown in Figures (ref) and (ref). We also report the Value-at-Risk (VaR) exceedance rates for the different models, which are shown in Table (ref). Figures (ref) and (ref) give the cross-evaluation of the optimal bids' objective values for the different forecast models for the 1-hour, 1 cycle case and the 4-hour, 2 cycle case. The results for the other cases are similar and can be found in the supplementary material. The results for the economic performance measures are as follows.

itemize• For the optimization with respect to the expected profits, the DLENAR models yield the highest profits. This is in line with all the scoring rules, which rank the DLENAR models highest. • For the risk-averse optimization, the LEAR-BS yields the highest profits for the CVAR optimization at the 90% and 75% level. For the CVAR at the 50% level, the DLENAR-DEP model yields the highest profits. The results are relatively similar for the 1 and 2 hour duration. • The climatology forecast yields surprisingly high Sharpe ratios for the risk-averse optimization. This is likely due to the fact that the optimal bids derived from the climatology forecast are very conservative, which leads to low profits but also low volatility of the profits, which results in a high Sharpe ratio.
figure[figure omitted — 293 chars of source]
figure[figure omitted — 291 chars of source]
table[table omitted — 7,976 chars of source]

For the risk-averse optimization, we discuss the results in more detail in the following.

itemize• The VaR exceedance rates for the CVAR optimization are closest to the nominal level for the DLENAR model. Interestingly, we see that for the rather low risk aversion of $\alpha =0.5$ and small battery configuration ($d \leq 2, c \leq 2$), the LEAR-based optimization yields the best calibrated VaR forecasts. For higher risk aversion, and larger battery configurations, modelling the dependence structure becomes more important and hence the DLENAR-DEP and DLENAR-DWD yield the best calibrated VaR forecasts. • The LEAR-BS model, which yields the highest profits for the CVAR optimization at the 90% and 75% level (see Figure (ref)), also has exceedance rates of 11% to 23% above the nominal level. The LEAR-N(0, $\Sigma$) model, which yields the second highest profits for the CVAR optimization at the 90% and 75% level, has exceedance rates of 5% to 15% above the nominal level. This stark difference is striking, as the two models are rated similar in terms of the CRPS and the ES, and the LEAR-BS model even has better scores in terms of the VS and DSS (see Table (ref)). • The climatology model has VaR exceedance rates below the nominal level and, for $\alpha = 0.9$, also very close to the nominal level, which gives an indication that the derived bids are very conservative. This is in line with its high Sharpe ratios and relatively high profits considering the simplicity of the model.

Additionally, we evaluate the decision quality by cross-scoring of the optimal bids derived from the different models. Figures (ref) and (ref) show the results for the 1-hour, 1 cycle battery configuration and the 4-hour, 2 cycle battery configuration. We choose these configurations, as the first one represents a common reference case and the latter can be seen as a reference for hydro-pumped storage lohndorf2023value and as a shorter-duration battery with additional commitments to balancing markets kraft2023stochastic. The upper panel show the pinball loss on the VaR forecast, the lower joint the joint score on the (VaR, CVAR) forecast. The columns denote the model whose optimal bids are evaluated and the rows denote the model whose forecasts are used for the evaluation. The diagonal hence shows the scores for the optimal bids derived from model $m$ evaluated on the forecasts of model $m$. The off-diagonal elements show the scores for the optimal bids derived from model $m$ evaluated on the forecasts of model $m' \neq m$. Figures for the remaining configurations are shown in the supplementary material and show similar results. The main conclusions from these results are as follows.

itemize• The results for the pinball scores for the VaR forecasts and the joint scores for the (VaR, CVAR) forecasts are similar, which is an indication that the results are not driven by a specific choice of the score function, but rather reflect a more general pattern. • For decreasing risk aversion, the differences between the different models become smaller, which is in line with the fact that for risk-neutral optimization, the optimal bids are not derived from the probabilistic forecasts, but rather from the point forecasts. Hence, the scores of LEAR-BS and LEAR-N(0, $\Sigma$) and the scores within the DLENAR models converge. We also see less statistically significant differences. • The DLENAR models have the lowest scores for the optimal bids derived from all models. Within the group of DLENAR models, for the 1 and 2 hour battery, the models are somewhat over-cross, the DLENAR-IND has the best forecasts for the bids derived from the DLENAR-DEP and vice versa. For the 4 hour battery, we see the DLENAR models have the best decision quality for the optimal bids derived from themselves, which is in line with the fact that for higher risk aversion and larger battery configurations, modelling the dependence structure becomes more important. • The Naive-BS model has the worst scores for the optimal bids derived from itself, but also for all other models' optimal bids. The climatology model has surprisingly good scores not only for its own optimal bids, but also for the optimal bids derived from the LEAR-based models, especially for the very risk-averse optimization. This is in contrast to the statistical scoring rules, where the climatology model performs worst.
figure[figure omitted — 501 chars of source]
figure[figure omitted — 495 chars of source]

Concluding, our economic evaluation of different forecasts in a risk-averse setting yields some interesting conclusions. While delivering the best profits and high Sharpe ratios, the LEAR-based models fail to deliver well-calibrated VaR forecasts and show shortcomings in the evaluation of the objective value of the optimal bids, when compared to simple benchmarks such as the climatology model. This is a stark reminder that total achieved profits are no suitable measure to compare model performance in a risk-averse setting. The stark differences in the VaR exceedance rates and the cross-scoring are not reflected in the total-profit ranking alone.

Economic Evaluation based on simplified BESS

Simplified case studies are popular for the monetary evaluation of forecasts through battery trading strategies. The predominant battery configuration is a 1-hour battery with one cycle per day, allowing one trade per day maciejowska2025statistical, chȩc2025extrapolating,serafin2025data. This last constraint is crucial, as it allows us to use a simple dynamic programming approach to optimize the battery trading strategy. Arguably, for risk-neutral optimization, this constraint does not lead to suboptimal strategies, as we just select the hours with the highest spreads. However, for risk-averse optimization, this constraint can lead to suboptimal strategies, as we might want to place multiple bids to diversify the risk.

Figure (ref) shows the total profits, Sharpe ratio and the number of no-bid days for the DP approach and the MILP approach for a 1-hour battery with one cycle per day. The results are shown for the expected profit optimization and the CVAR optimization at levels $\alpha \in \{0.90, 0.75, 0.50\}$. In the following, we discuss some key insights from these results. We denote with DP-1 the stylized optimization, with MILP-1 the MILP optimization with the same constraint and with MILP-24 the MILP optimization allowing for up to 24 bids per day.

itemize• For the expected profit optimization, we see that the profits are exactly the same for the simplified BESS and the MILP approach. This is in line with the theory, as for a 1-hour battery, we just select the highest spread hours. • For the risk-averse optimization, we see that the profits are higher for the MILP-24 approach. The effect is most prominent for the CVAR ($\alpha=0.9$) optimization and the DLENAR-IND and the climatology model. The simplified approach does not allow for diversification gains on the marginal, but only on the dependence structure. The MILP-24 approach also rewards diversification along the marginal. Intuitively, the simplified approach only allows us to trade the largest spread, or not at all, while the MILP-24 approach allows us to trade the other spreads as well, which can yield diversification gains even if the dependence structure is not captured well. This effect is also visible in the number of no-bid days, where the simplified approach has more no-bid days than the MILP-24 approach.

In this case, the simplified approach overstates the effect of modelling the dependence structure for the DLENAR models, and, at the same time, understates the performance of naive strategies such as the climatology model, which can still yield good diversification gains on the marginal. This is an indication that the simplified approach might lead to spurious conclusions about the performance of the forecasts, especially for risk-averse optimization.

figure[figure omitted — 564 chars of source]

Economic Evaluation based on QBTS

In this section, we briefly discuss the results of the QBTS for the different forecast models and potentially spurious conclusions. The results are shown in Figures (ref) and (ref), giving the total and per MWh profits and the acceptance probabilities (AP) for the different forecast models. We derive the AP based on the ensemble forecasts, taking the dependence structure into account: $ \widehat{\text{AP}} = \frac{1}{M} \sum_{m=1}^M \mathbb{I}\{ \hat{Q}^{1-\alpha}_s \leq \widehat{\ensuremath{\bm{\mathrm{F}}}}_{s,m} \cap \hat{Q}^{\alpha}_b \geq \widehat{\ensuremath{\bm{\mathrm{F}}}}_{b,m} \}. $ We denote the nominal prediction interval coverage as $\lambda = 1 - 2 \alpha$, as we have $Q_b^{1-\alpha}$ and $Q_s^{\alpha}$ as upper and lower bounds for the QBTS. The DLENAR models yield the highest total profits for the QBTS, followed by the LEAR-based models. The LEAR-BS and the Naive-BS are close, while the climatology model yields the lowest profits across all values of $\alpha$. In terms of the per-MWh profits, the LEAR-based models yield higher profits for small nominal prediction intervals, while the DLENAR models yield higher profits for larger nominal prediction intervals. The last panel gives the difference between expected profits and realized profits. Here, most models yield higher than expected profits. This effect is, however, defined by three effects: the level accuracy, the over- and underdispersion of the forecasts and the specification of the dependence structure. The difference is the largest for low nominal coverage and decreases with higher coverage. Interestingly, for the climatology model, the difference increases again for $\lambda > 0.8$, potentially due to the insufficient point accuracy of the model.

In Figure (ref) we compare the realized acceptance probabilities with the expected AP under independence, and the expected AP under the dependence structure of the forecasts. The realized AP are below the expected AP under independence, which is in line with the fact that the QBTS does not take the dependence structure into account and hence overestimates the AP. This is in line with the results from Section (ref), especially the Remark (ref) and Figure (ref). Especially for the Climatology model, we can also see the clear correlation between the AP and the profits. Around $\lambda = 0.65$, the empirical AP increases and the profits increase as well, but due to overdispersion, also more loss-making trades are accepted and the profits per MWh decrease. The difference between the expected AP under the forecast $\widehat{\ensuremath{\bm{\mathrm{F}}}}$ and the realized AP in the second panel of Figure (ref) is defined by three sources of error. For the DLENAR-IND model, we see that the difference is close to the theoretical AP in the left panel, since the model assumes independence. The DLENAR-DWD and DLENAR-DEP have the lowest difference between expected and empirical AP. The largest differences appear for the Climatology, the LEAR-BS and the Naive-BS, while the LEAR-N(0, $\Sigma$) is somewhat in the middle. Generally, we see that for higher nominal prediction interval coverage, the difference increases.

These results emphasize that the QBTS can lead to hard-to-interpret and potentially spurious conclusions about the performance of probabilistic forecasts. Disentangling the different effects of level accuracy, over- and underdispersion and the dependence structure is not straightforward, but crucial for a proper interpretation of the results.

figure[figure omitted — 382 chars of source]
figure[figure omitted — 286 chars of source]

Conclusion

Probabilistic forecasts are often proposed to enhance decision making, especially under risk. The economic evaluation of (probabilistic) forecasts is increasingly popular in electricity price forecasting and battery storage arbitrage has emerged as natural showcase application for the economic gains of better forecasts. However, the economic evaluation of forecasts, especially of probabilistic forecasts in a risk-averse setting, is not straightforward. This paper addresses two issues:

itemize• We analyze the properties of the popular quantile-based trading strategy uniejewski2025smoothing,o2025optimising and show that it is not a proper scoring rule for the evaluation of forecasts, as systematic distortions of the forecasts can lead to higher profits. We further show that QBTS cannot anticipate diversification benefits through the correlation structure of the forecasts, which can lead to suboptimal decisions and thus to spurious conclusions about the performance of forecasts. • Based on this insight, we discuss the evaluation of fully multivariate, probabilistic forecasts through battery trading strategies, framed as stochastic programming problems. We show that battery trading strategies can have low discriminatory power for the evaluation of forecasts. We discuss measures to analyze the decision quality in a risk-averse setting and use these to analyze the performance of different multivariate forecasts in a large-scale simulation study across various battery configurations.

Our results shed light on the connection between statistical scoring rules for probabilistic forecasts and measures of decision quality for stochastic, risk-averse optimization problems. Similar to previous work on point forecasting nitka2023combining,serafin2025data, we see that measures of decision quality and scoring rules disagree also in the probabilistic setting. While we acknowledge that the results of this paper are based on a specific case study, we believe that the insights are more general and can be applied to other settings as well. Our research touches upon a number of potential avenues for future research.

itemize• Strictly proper scoring rules give a principled way for the estimation of statistical models and better scores should translate into better decision quality. One avenue of further research is the development of new, objective-aligned evaluation measures, such as proposed by maciejowska2025statistical for point forecasts. On the other side, the properties of established scoring rules such as the CRPS, DSS, and ES with respect to locality, symmetry and discrimination are continuously studied buchweitz2025asymmetric,marcotte2023regions,alexander2024evaluating and their connection to decision quality is not yet fully understood. • The evaluation of the decision is based on trades and their profits. However, the decision not to trade (in a certain hour) can also be a valuable decision, but is not as prominently captured in the evaluation. This hides part of the forecaster's skill, as one needs a good forecast to decide not to trade. The development of diagnostic tools to capture the value-add of improved forecasts in this area is an interesting avenue for future research. • Our findings have further implications for the integration of decision and forecasting, e.g. in the setting of end-to-end learning donti2017task,elmachtoub2022smart, wen2025value, where the low discriminatory power of the decision-based evaluation can hint at potentially demanding learning surfaces. On the other hand, the fact that only a small part of the full price distribution is relevant for the decision can be used to develop more efficient learning algorithms if the relevant part of the distribution can be identified a priori. The connection between the rank- and association-based scoring rules and the relationship between battery optimization and portfolio management suggest that the application of learning-to-rank algorithms could be an interesting avenue for future research as well song2017stock, zhang2022constructing.

We conclude by summarizing the key findings of this paper in a recommendation: the economic evaluation of probabilistic forecasts should be treated as an additional layer, not only as forecast evaluation, but as decision quality evaluation. A robust procedure therefore combines (a) forecast evaluation using (strictly) proper scoring rules, and (b) objective-aligned decision diagnostics, and (c) the evaluation of economic performance measures. Agreement across all measures gives strong evidence for model quality, while disagreement can give insights into the strengths and weaknesses of the different models.

Acknowledgments

Simon Hirsch is employed as industrial Ph.D. student with Statkraft Trading GmbH and gratefully acknowledges the funding and support received. Florian Ziel acknowledges funding in the course of TRR 391 Spatio-temporal Statistics for the Transition of Energy and Transport (520388526) by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation). The authors are grateful for interesting discussions with Daniel Gruhlke and David Wozabal.

Declaration of generative AI and AI-assisted technologies in the manuscript preparation process

During the preparation of this work the authors used Github Copilot (multiple LLM models) in order to improve the quality of the code, figures, tables and text. After using this tool/service, the authors reviewed and edited the content as needed and take full responsibility for the content of the published article.

CRediT Authorship Contribution Statement

Simon Hirsch: Conceptualization; Data curation; Formal analysis; Investigation; Methodology; Software; Validation; Visualization; Roles/Writing - original draft.

Florian Ziel: Conceptualization; Funding acquisition; Methodology; Project administration; Resources; Supervision; Validation; Writing - review & editing.

{0pt}