Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
134,500 characters · 26 sections · 100 citation commands
Probabilistic Forecasting for Day-ahead Electricity Prices, Battery Trading Strategies and the Economic Evaluation of Predictive Accuracy
Keywords: Probabilistic Forecasting; Scoring Rules; Battery Optimization; Stochastic Programming; Decision Quality; Forecast Evaluation
Forecasting electricity prices is crucial for informed decision-making in the daily operations of energy companies. Over the last years, both industry and research have focused on probabilistic forecasting, which “has a lot to offer, in particular, improved assessment of future uncertainty, ability to plan different strategies for the range of possible outcomes, increased effectiveness of submitted bids” nowotarski2018recent. A pressing question in practice is whether improved forecasting performance, in terms of scoring rules or loss functions, improves decision-making, measured in monetary terms granger2000economic,yardley2021beyond. To address this need, various application studies have been proposed in the energy forecasting literature, nitka2023combining, vogler2021event. A common showcase for the economic benefits of better forecasts is the simplified operation of grid-scale batteries energy storage systems (BESS), which can buy electricity (and charge) when prices are low and sell electricity (and discharge) when prices are high. Better anticipation of prices should translate into higher profits. This storage arbitrage example has been adopted for both point and probabilistic forecasting studies sang2022electricity, chȩc2025extrapolating, serafin2025loss, maciejowska2025statistical, uniejewski2025smoothing, o2025conformal, o2025optimising.
It is widely agreed on that forecast evaluation and model selection should be conducted using proper scoring rules. Proper scoring rules incentivize honest forecasting, that is, the optimal score is attained for predicting the true distribution. Scoring rules are called strictly proper if the optimum is unique. Strictly proper scoring rules are thus the gold standard for forecast evaluation and model selection gneiting2007strictly. However, the strict propriety alone does not guarantee that scoring rules penalize forecast errors symmetrically buchweitz2025asymmetric, provide meaningful discrimination alexander2024evaluating,marcotte2023regions or that they are aligned with the decision-making process at hand. To bridge this gap, the properties of scoring rules and their modification through weighting, thresholding and censoring is an active area of research allen2023weighted, de2025localizing, shahroudi2025aligning.
Consequently, some questions arise: (How) can we infer the quality of forecasts based on the quality of the decisions? How can we assess the quality of decisions under competing models of uncertainty? Can battery trading strategies provide meaningful discrimination between forecasts of different quality? Are battery trading strategies aligned with the concept of proper scoring rules? How are statistical scores related to the economic performance? These questions will guide us throughout this paper.
The majority of works concerning probabilistic forecasting of electricity prices focuses on the marginal distribution of hourly prices. Based on this, quantile-based trading strategies (QBTS), originally introduced by uniejewski2025smoothing, employ quantile-based limit orders on the day-ahead market. Briefly, a trader places limit buy and limit sell orders based on $\alpha$ and ($1-\alpha$) -- quantiles of the predicted distribution, thereby controlling a trade-off between frequency of trading and anticipated profit of each trade. The approach has been used in various subsequent studies (see Table (ref) for an overview) and extended by o2025optimising,o2025conformal,o2024conformal. While instinctively intuitive, we provide theoretical and empirical evidence that the approach can be gamed by providing systematically biased forecasts and is not aligned with the concept of proper scoring rules. Moreover, we show that the approach does not anticipate diversification benefits by neglecting the dependence structure of electricity prices.
Keeping the battery example, a natural further starting point is to ask whether battery trading strategies based on a fully multivariate probabilistic forecasts, optimized for common risk-measures such as expected profits or the conditional value-at-risk (CVAR) can be used to evaluate the quality of forecasts. We show that even in this case, the resulting profits are not necessarily a strictly proper scoring rule and have potentially low discriminatory power. For point forecasts, maciejowska2025statistical empirically evaluate the relationship between the statistical quality of point forecasts and the economic performance of battery trading strategies. However, formal analysis and the evaluation of the probabilistic and risk-averse case is notably absent in the literature. We contribute to this gap by providing a theoretical and empirical analysis of the relationship between forecast quality and a notion of decision quality, measured by the predictive accuracy of the objective value of our stochastic program under competing forecasts. We relate decision quality, statistical forecast and economic forecast evaluation.
Our contribution can be summarized in the following points:
We believe that our findings are of interest to multiple research communities, including the forecasting community, the energy economics community and the stochastic optimization/programming community.
On the other hand, we acknowledge some limitations a priori: this work aims at deepening the understanding of the relationship between forecast quality and decision quality in the context of battery trading strategies. We do not focus on the absolute profitability of BESS investments in energy markets spodniak2021profitability, staffell2016maximising or on sophisticated, multi-market optimization lohndorf2023value, kraft2023stochastic, ghadimi2024stochastic.
The remainder of this paper is structured as follows: Section (ref) formalizes our claim and analyzes the objective value of quantile-based optimization from a theoretical vantage point. Section (ref) presents and analyzes the proposed multivariate dynamic programming approach. Section (ref) gives a case study. The last section concludes and identifies avenues for further research.
Quantile-based trading strategies (QBTS) are a popular approach to evaluate the economic value of probabilistic forecasts in the context of battery trading. The main idea is to place limit orders based on quantiles of the predicted distribution of electricity prices. By adjusting the quantile level $\alpha$, the trader can control the trade-off between the frequency of trading and the anticipated profit of each trade. Two main variants exist in the literature, the quantile-based trading strategies based on the works of uniejewski2025smoothing and the TS-1/2/3 strategies introduced by o2025optimising,o2024conformal,o2025conformal.
We briefly describe QBTS and their variations in the literature. Assume the trader has access to a probabilistic forecast of the day-ahead electricity prices, which is issued as quantile predictions for each delivery hour. The trader aims to maximize the expected returns from trading a battery storage asset on the day-ahead market. The battery has a certain capacity $\kappa$ and efficiency $\eta$. The strategy consists of two steps:
In this case, we assume that the execution of the buy and sell bid is coupled, i.e. either both are executed or neither is executed. Thereby, the battery is always in a valid state and no additional orders need to be placed in other markets to balance the battery.\footnote{This assumption is closely related to the so-called loop bid, a product available at the EPEX power exchange that guarantees joint or no execution of coupled buy and sell bids. Contrary to individual directional limit prices however, loop bids are conditional on the total cashflow of the linked bids epexspot2025trading. Framing as a loop bid would further move away from the original QBTS.} Alternatively, the trader might be faced with partial execution, that is, only the buy or only the sell trade is executed. There are multiple ways to handle this, the most prominent ones are described in the following and an overview is given in Table (ref).
Both options introduce additional risk and make the (predicted) distribution of battery returns intractable. Therefore, we focus on the simplified approach in the following theoretical analysis and empirical evaluation. Figure (ref) summarizes the QBTS approach. Algorithmic descriptions of all battery trading strategies are given in the supplementary material.
The TS-1 strategy as proposed by o2025optimising is based on a similar algorithm. However, instead of placing limit orders based on quantiles, quantiles are used to choose the buying and selling hours and unlimited bids are placed. The objective function of the TS-1 strategy is given by
The main difference to QBTS is that the TS-1 strategy does not place limit orders but unlimited orders at the selected hours and hence bids are always accepted. The TS-2 strategy extends the TS-1 strategy by placing additional orders in the balancing market to balance the battery at the end of the day, while the TS-3 strategy extends the TS-1 strategy by considering ramp rate constraints. We focus on the TS-1 strategy in the following, as it is the most similar to QBTS and allows for a more direct comparison.
We begin the theoretical discussion by recapitulating the definition of a strictly proper scoring.
We denote the unknown, true distribution of day-ahead electricity prices $\boldsymbol{P}_d = (P_{d,0}, ..., P_{d, 23}) \sim \mathcal{D}_d$ on day $d$ as $\mathcal{D}_d$. Denote the realized prices as $\ensuremath{\bm{\mathrm{p}}}_d = (p_{d,0}, ..., p_{d, 23})$. We describe the 24-dimensional vector of decisions as $\ensuremath{\bm{\mathrm{a}}}(b, s) = (a_0, ..., a_h, ..., a_H)$, where $b$ and $s$ denote the hour(s) in which electricity is bought and sold. We will omit the $d$ in the following as the optimization is repeated every day. Denote the efficiency as $\eta$ and the storage capacity as $\kappa$ and hence we have
For the first part we constrain ourselves to only one buy and one selling bid. We implicitly assume that the charge/discharge capacity $\xi = \kappa$. Without loss of generality we assume that $b < s$, hence we buy (and charge) before we sell (and discharge). The revenues $R_d$ for an unconstrained bid $\ensuremath{\bm{\mathrm{a}}}(b, s)$ are given by
and for a bid with limit prices $Q_b^{1-\alpha}$ and $Q_s^{\alpha}$, we have:
This allows us to state our main result regarding quantile-based trading strategies.
We show that an overdispersed forecast can lead to higher expected profits than the perfect forecast. Assume the price distribution is symmetric and we have an over-dispersed forecasts $\widetilde{\boldsymbol{P}} = b \cdot \boldsymbol{P}$, where $\boldsymbol{P} \sim \mathcal{D}$ is the true price distribution and $b > 1$ is a dispersion parameter. We have $\widetilde{Q}_h^q = b \cdot Q_h^q$, where $Q_h^q$ is the quantile function of the true distribution. For simplificity, we assume that delivery periods are independent. The mean and median forecast are unbiased, $\widetilde{Q}_h^{0.5} = Q_h^{0.5}$. We therefore have:
We are interested in the difference between the expected profits of the over-dispersed forecast and the perfect forecast, that is, $\mathbb{E}[\widetilde{R}(b, s, \alpha)] - \mathbb{E}[R(b, s, \alpha)]$. We look at the probability of bid acceptance and the expected value of the returns separately. We assume independence of the delivery periods, hence we have $\text{AP} = \mathbb{P}(P_{b} \leq Q_b^{1-\alpha}; \; P_{s} \geq Q_s^{\alpha}) = (1 - \alpha)^2$ for the perfect forecast and $\mathbb{P}(P_{b} \leq \widetilde{Q}_b^{1-\alpha}) = 1 - \widetilde{\alpha} \geq 1 - \alpha$ and hence:
The intuition can be seen in the following Figure (ref), which shows the true acceptance region, which corresponds to the acceptance region of a perfect forecast, and the acceptance region of an overdispersed forecast. We can also see that it is sufficient if one side (buy or sell) is provided with an overdispersed forecast.
The difference in the expected profits (EP) of the returns is given by:
since we have
as the conditioning sets of both expectations are larger for the over-dispersed forecast. Intuitively, we accept more trades and therefore more trades that are less profitable on average, i.e. higher buy prices and lower sell prices. Concluding, the difference in the expected value of the returns is negative. The overall effect therefore depends on the question whether the increase in the AP outweights the decrease in EP
For low values of $\alpha$, the AP is high and hence the first dominates. For high values of $\alpha$, the AP is low and hence the second dominates. Additionally, for larger spreads between the buy and sell price, the EP is higher and hence the first dominates. For intermediate values of $\alpha$, the effect depends on the underlying price distribution is hard to gauge. We show the implications of this result in a small simulation study: We simulate the true prices as Gaussian random variables with means $\mu_b$ and $\mu_s$. We over- and underdisperse the forecasts by multiplying the true $\sigma$ with a dispersion parameter $b$. We then compute the expected profits for different values of $b$ and $\alpha$. The results are shown in Figure (ref). First, we see that overdispersed forecasts lead mechanically to higher acceptance ratios (top row). However, the effect on the expected profits is not monotone. For high spreads ($\mu_b = 50, \mu_s = 100$) and low volatility ($\sigma = 10$), overdispersion leads to higher profits. For smaller spreads, and higher volatilty, profits decrease for overdispersion and low values of $\alpha$, since we start to accept more unprofitable trades.
In a second simulation study, we show the implications of this result. We simulate true prices as multivariate Gaussian random variables with means $\mu_b = 90$ and $\mu_s = 100$ and standard deviation $\sigma = 20$ (same as Panel 3 in Figure (ref)). We vary the correlation between $b$ and $s$, taking $\rho = \{0, 0.4, 0.8\}$ in the according covariance matrices $(\Sigma_0, \Sigma_1, \ensuremath{\bm{\mathrm{\Sigma}}}_2)$. Again, we show the AP and the expected profits for different values of $\alpha$ in Figure (ref). The AP decreases for higher correlation. Profits decrease for higher correlation and higher values of $\alpha$. For overdispersed forecasts, the effect is not as strong. This raises two important points: First, in practice, electricity prices are positively correlated across delivery hours, which further weakens the validity of QTBS for forecast evaluation. Second, the positive dependence should potentially decrease the risk of battery trading strategies, however, the QTBS approach does not capture this effect and hence can lead to suboptimal bids.
Although formally more involved to analyze due to the multi-market setting, this results highlights an important consideration for the TS-1/2/3 strategies o2024conformal,o2025conformal,o2025optimising, which directly optimizes for the difference of the quantile forecasts: Treating the hourly quantile forecasts as independent neglects the dependence structure of electricity prices and hence can lead to suboptimal bids.
Let us formalize the notation. We denote with $\ensuremath{\bm{\mathrm{a}}} = (a_0, .., a_H)$ a (fixed) vector of bids.\footnote{ Strictly speaking $\ensuremath{\bm{\mathrm{a}}}$ is a random variable, since the trading decision depend on the forecasted distribution $\mathcal{F}$, which in turn depends on learned coefficients and parameters. Thus, $\ensuremath{\bm{\mathrm{a}}}$ might be biased. We generally treat $\ensuremath{\bm{\mathrm{a}}}$ as deterministic and return to the issue briefly in the following Section (ref). } We have $\ensuremath{\ensuremath{\bm{\mathrm{a}}}^\top}\boldsymbol{P} \sim \mathcal{R}$ as the distribution of the revenues for bid vector $\ensuremath{\bm{\mathrm{a}}}$. Given the forecast $\boldsymbol{F} \sim \mathcal{F}$, usually issued as ensemble or scenario forecast $\widehat{\ensuremath{\bm{\mathrm{F}}}}$, we have $\ensuremath{\ensuremath{\bm{\mathrm{a}}}^\top}\boldsymbol{F} \sim \mathcal{P}^{\ensuremath{\bm{\mathrm{a}}}}$ as the forecasted distribution of the revenues for bid vector $\ensuremath{\bm{\mathrm{a}}}$. We can then write the battery optimization problem as $\max_{\ensuremath{\bm{\mathrm{a}}}} \rho(\mathcal{P}^{\ensuremath{\bm{\mathrm{a}}}})$, where $\rho(\cdot)$ is the risk measure we want to optimize. The realized revenues are given by $r_d = \ensuremath{\ensuremath{\bm{\mathrm{a}}}^\top}\ensuremath{\bm{\mathrm{p}}}_d$.
We approach the issue from the objective function. Commonly, we want to optimize some kind of risk measure $\rho(\cdot)$ of the returns $$ \max_{b,s} \quad \rho\left(R(\ensuremath{\bm{\mathrm{a}}}(b, s), \mathcal{F})\right) $$ this can be expected profits, or the (conditional) value-at-risk, a mixture of both or some specific functions. We assume that the forecaster equips the trader with a multivariate predictive distribution $\mathcal{F}$, on which decisions are based. $\mathcal{F}$ is commonly represented through a (large) number of samples $\widehat{\ensuremath{\bm{\mathrm{F}}}}$. Intuitively, we can use the following dynamic programming approach:
It is straightforward to see that, in this setting, we can directly optimize any risk measure. However, the approach is computationally demanding and does not scale well in the number of bids placed and the relative size of the bids. In the following Section (ref), we discuss a more scalable approach based on mixed-integer linear programming. Nevertheless, this mental model provides a clear and intuitive understanding of how to employ multivariate probabilistic forecasts for battery optimization, can serve as a straightforward benchmark and highlights the importance of the dependence structure. Figure (ref) visualizes the approach.
This connection to portfolio optimization allows us to discuss two interesting points.
While the dynamic programming approach is intuitive, it does not scale well. Hence, we propose an alternative formulation as a mixed-integer linear programming problem. A drawback of this approach is that only the expected value and the conditional value at risk can be optimized directly. Mean-variance and value-at-risk objective functions cannot be reformulated as MILP problems rockafellar2000optimization, uryasev2001conditional. We will discuss this difference in the simulation study in Section (ref).
We again assume that we have $M$ scenarios $\widehat{\ensuremath{\bm{\mathrm{F}}}}_{h,m}$ for each hour $h$ and scenario $m$. We further relax some assumptions from the section before and allow for multiple buy and sell bids. Denote the number of buy bids as $N_b$ and the number of sell bids as $N_s$. We treat the bid volume as decision variable $a_h^b > 0$ and $a_h^s >0$. We have
where $\kappa$ is the storage capacity and $\xi$ is the charge/discharge capacity (also called energy or rated capacity). In light of the recent market change to quarter-hourly trading in the single day-ahead coupling (SDAC), the use of such dynamic programming algorithms becomes computationally demanding and as such, the use of mixed-integer linear programming (MILP) solvers is advisable. Our implementation uses pyomo and the HIGHS solver hart2011pyomo,huangfu2018parallelizing.
In the following, we take the above formulation and discuss whether the multivariate probabilistic forecast can be scored through the battery optimization. Using a simple example, we show that it is possible to construct different forecast distributions that lead to the same objective value $\rho(\cdot)$ for the battery optimization. This raises the issue of discriminatory power of battery optimization for the evaluation of multivariate probabilistic forecasts.
The reason for this issue is the number of degrees of freedom within the battery optimization problem. In the classic forecast evaluation setting, we are interested in ranking a number $\mathcal{M}$ different models $h_m(\ensuremath{\bm{\mathrm{X}}}) \rightarrow \mathcal{F}_m$ based on their forecasted distribution $\mathcal{F}_m$ and the realized values $y$ using scores $s_m = S(\mathcal{F}_m, \ensuremath{\bm{\mathrm{p}}})$. In the battery optimization case, we rank different models based on the outcome of an optimization problem that takes the forecasted distribution as input. Therefore, we have
where $\ensuremath{\ensuremath{\bm{\mathrm{a}}}^\top} \boldsymbol{F}_m \sim \mathcal{P}^{\ensuremath{\bm{\mathrm{a}}}}_{m}$ is the forecasted distribution of the revenues for model $m$ and $R(\ensuremath{\bm{\mathrm{a}}}_m^*, \ensuremath{\bm{\mathrm{p}}})$ is the realized revenues for the optimal bid $\ensuremath{\bm{\mathrm{a}}}_m^*$ and the realized prices $\ensuremath{\bm{\mathrm{p}}}$. We can then rank the models based on $s_m$ and $r_m$, respectively, which raises the question of whether the ranking of the models based on $s_m$ and $r_m$ is equivalent.
In the following, we will show through a simple example that different forecast distributions can lead to the same ranking of bids in the battery optimization problem, hence placing a fundamental limit on forecast evaluation through battery optimization.
From a practical perspective, this implies that economic backtest performance can be decision-relevant, but it is generally unsuitable as a stand-alone strictly proper scoring rule for model selection over full predictive distributions. We illustrate the issue of information compression through the battery optimization problem in the following examples.
These two examples already hint at potentially low discriminatory power of the battery optimization for forecast evaluation. It is important to be aware of this issue when evaluating forecasts through battery optimization. For optimizing the expected revenues, the issue has been discussed empirically by maciejowska2025statistical. The main risk in the application is that, for slight changes in the application, a forecast that has previously performed well cannot be expected to perform well anymore. For example, a forecast that has performed well for a 1-hour battery might not perform well for a 2-hour or 4-hour battery, since the optimal bids can change and hence, the relevant part of the forecast distribution can change.
Based on these results, two extreme points of view can be taken:
There is a wide gray area in between these two extreme points of view. A plausible option is using weighted combinations of $s_m$ and $r_m$ to rank models. If rankings coincide, this gives us more confidence in the model choice and employ additional model diagnostics. Somewhat turning things around, it can be argued that this reliance on few parts of the distribution is advantageous. If we need only a certain part of the distribution to be (approximately) correct to create the optimal bids, we can utilize this for reducing the dimensionality of the problem. This can be hard to achieve a priori, as the optimal bids $\ensuremath{\bm{\mathrm{a}}}_m^*$ are not known. In hindsight, we can analyze which parts of the forecast distribution were relevant for the decision-making problem and focus on improving these parts in future forecasts.
In the following, we discuss the evaluation of decision quality given competing forecasts. It is desirable that the evaluation of the decision quality is consistent with the objective function employed in the optimization.
Therefore, the predicted objective value of the battery optimization for the forecast model $m$ as $\omega^*_m = \rho(\mathcal{P}^{\ensuremath{\bm{\mathrm{a}}}_m^*}_m)$ and the realized returns of $r_m = \ensuremath{(\ensuremath{\bm{\mathrm{a}}}_m^*)^\top} \ensuremath{\bm{\mathrm{p}}}$ cannot be evaluated across different models. For a strictly proper scoring rule $S(\cdot, \cdot)$ of $\rho(\cdot)$ we have
which shows that the forecast optimal bids of forecast $\mathcal{F}_m$, which we would like to evaluate, enters the scoring rule on both sides of the equation. As possible workaround, for a number of models $m \in \mathcal{M}$, we can evaluate the scoring rule for all optimal bids $\ensuremath{\bm{\mathrm{a}}}_m^*$ and all forecasts $\mathcal{F}_m$ and analyze the resulting scores. If the forecast $\mathcal{F}_m$ has the lowest score of all models for the risk measure for bids $\ensuremath{\bm{\mathrm{a}}}_m^*$, then this should a good indication that actually this decision is good. If many/all models $i = 1, ... ,M; i \ne{} m$ have better scores on the objective value $\omega_m$ than $\mathcal{F}_m$, this is an indication that the decision $\ensuremath{\bm{\mathrm{a}}}_m^*$ is not globally optimal. This also gives rise to a formal test by comparing the scores of the models $i$ with the scores of model $m$ for the optimal bids $\ensuremath{\bm{\mathrm{a}}}_m^*$, which can be done through a Diebold-Mariano test diebold2002comparing, diebold2015comparing. We test for the hypothesis
If we can reject the null hypothesis, this implies that model $i$ was able to yield significantly better predictions of the objective value for this decision and hence an indication that the decision $\ensuremath{\bm{\mathrm{a}}}_m^*$ based on forecast $\mathcal{F}_m$ should not be trusted. Importantly, this only addresses one side of the optimization problem, since we are only evaluating how well the forecast $\mathcal{F}_m$ performs for the optimal bids $\ensuremath{\bm{\mathrm{a}}}_m^*$ relative to the other forecasts for the hours in which $m$ places bids, which is only a subset of hours. The predictive accuracy in the other hours is not evaluated, even though the predictive performance in these hours is relevant for the decision-making problem.
In the following section we run a small forecasting experiment and analyze the impact of different forecast errors on the results of the optimization. We employ seven different forecasting models with different modeling strategies and distributional assumptions. These models range from the simple climatology and naive models, the lasso-estimated autoregressive model lago2021forecasting and distributional regression approaches muniain2020probabilistic, hirsch2024online,hirsch2025online. It is important to note that we have preferred models that provide either an option for bootstrapping or a parametric distributional assumption in order to draw samples for the stochastic program. For this reason, we have not included quantile regression or conformal prediction approaches uniejewski2025smoothing,o2025conformal. As the main aim of the case study is to relate statistical measures and economic evaluation on a diverse pool of models, we acknowledge that we have not undertaken extensive feature engineering or hyperparameter optimization, but chose a configuration that has been shown to work well in previous work.
The day-ahead market is the primary exchange for electricity in Germany. The market is organized as a uniform price auction, where the price for each hour of the next day is determined by the intersection of the aggregated supply and demand curves. The market is cleared at 12:00 CET for the delivery of electricity on the next day. We use data from ENTSO-E and \url{investing.com} provided by lipiecki2024postprocessing. The data covers the period from 2015 to 2023 and includes the day-ahead prices, the residual load, the marginal costs of gas, coal and oil fired production and the associated emission allowance costs. We use 2015-01-15 to 2018-12-26 for model training and 2018-12-27 to 2023-12-31 for testing, in line with other works using the same data set hirsch2024online,hirsch2025online,marcjasz2023distributional.
We employ a distributional regression model or generalized additive model for location, scale and shape rigby2005generalized based on an extended expert model using Student's $t$-distribution for the marginal distribution of the electricity prices $$P_{d,h} \sim t(\mu_{d,h}, \sigma_{d,h}, \nu_{d,h})$$ Distributional models are popular for electricity price forecasting serinaldi2011distributional, abramova2020forecasting, muniain2020probabilistic, klein2023deep,marcjasz2023distributional. We extend the well-known LASSO-estimated autoregressive model with exogenous variables lago2021forecasting model by non-linear effects representing the merit-order effects of residual load with the marginal costs (MC) of gas, coal and oil fired production and the associated emission allowance costs.
where $\eta^\ensuremath{\text{Gas}} = 0.2, \eta^\ensuremath{\text{Coal}} = 0.35$ and $\eta^\ensuremath{\text{Oil}} = 0.3$ are emission factors and $\eta$ is a conversion factor between the fuel's unit and MWh thermal energy ghelasi2025data, ghelasi2025day and $f(\cdot)$ denotes a b-Spline basis of degree 2 and 4 knots. Weekday effects are denoted with $\mathcal{W} = \{\text{Mon}, \text{Tue}, \text{Thu}, \text{Fri}, \text{Sat}, \text{Sun}, \text{Hol}\}$; holidays are encoded separately ziel2018modeling. For the scale parameter $\sigma_{d,h}$ of the price distribution, we employ a similar structure, but with linear equations for residual load and the fundamental fuel prices. The degrees of freedom for the $t$-distribution are estimated as intercept. The equations can be found in the Appendix (ref). We name the model distributional LASSO-estimated non-linear autoregressive model (DLENAR). The model is estimated in a expanding window fashion using the online learning algorithm developed in hirsch2024online to save computational time. We estimate separate models for each hour of the day. We model the dependence structure using a Gaussian Copula by the inference-for-margins approach. We convert the in-sample observations to the $\mathcal{U}(0, 1)$-space by the probability integral transformation (PIT) and subsequently to the $\mathcal{N}(0, 1)$ space by the inverse quantile transformation $\ensuremath{\bm{\mathrm{g}}} = \Phi^{-1}(\ensuremath{\bm{\mathrm{u}}})$. On the Gaussian space, we fit the correlation matrix.\footnote{Note that we additionally employ a ranking step to enforce strict uniformity on the residuals. This step does not alter the ordering of $\ensuremath{\bm{\mathrm{u}}}$, but the conversion to ranks makes observations equally spaced.} We fit the three different models for the dependence structure:
We draw scenarios by sampling from the uniform distribution and applying a rank-reordering to ensure that the marginal distributions are exactly equal across the different dependence structures. Additionally, we employ a number of benchmark models from the literature to obtain a diverse set of models.
For all bootstrap approaches, we always sample full error trajectories to ensure that we (implicitly) capture the dependence structure of the errors. Model equations and further details can be found in the Appendix (ref).
We run a number of battery optimization experiments using different risk measures and correlation structures. Our experiments can largely be grouped in three categories. Generally, our experiments assume a storage of 10 MWh, a round-trip efficiency of $\eta^2 = 0.95^2$. The duration of the storage, that is, the ratio between storage capacity and maximum charge/discharge power is set to 1 hour unless otherwise stated. For larger battery capacities, we would need to take market impact into account narajewski2022optimal, which is beyond the scope of this paper.
\paragraph{Multivariate Probabilistic Scoring Rules and Battery Revenues} In this experiment, we evaluate the relationship between multivariate probabilistic scoring rules and battery revenues. We discuss the results in light of Result (ref) and analyze the relationship between the scoring rules, decision quality and battery revenues. We run the battery optimizations maximizing the expected profits and the CVAR at levels $\alpha \in \{0.90, 0.75, 0.50\}$ using the MILP approach for a 1, 2 and 4 hour battery. The results of this experiment are discussed in Section (ref).
\paragraph{Multivariate Forecasts and Optimization Algorithms} In Sections (ref) and (ref) we have described an intuitive dynamic programming approach and a more scalable mixed-integer linear programming (MILP) approach to optimize battery trading strategies using multivariate probabilistic forecasts. Here, we briefly evaluate the impact of placing only one buy and sell pair compared to the MILP approach allowing for multiple bids. To this end, we compare results for a 1-hour battery, maximizing the expected profits and the CVAR at levels $\alpha \in \{0.90, 0.75, 0.50\}$ using both optimization approaches. The results of the experiment are discussed in Section (ref).
\paragraph{Quantile-based Trading Strategies} We evaluate the quantile-based trading strategy (QBTS) as described in Section (ref) using different forecast models. We evaluate the total profits, profits per MWh traded and the acceptance probability for different nominal prediction interval widths $\alpha$. Here, we focus providing an intuitive understanding of the Proposition (ref) and Remark (ref) through numerical examples and contrast the results with the stochastic programming approach. The results of this experiment are discussed in Section (ref).
We evaluate the economic performance of the different forecast models based on the daily returns from the battery optimization experiments. The daily return for day $d$ and model $m$ is defined as $r_{d,m} = \sum_{h=0}^H a_{d,h,m} P_{d,h}$ where $a_{d,h,m}$ are the optimal actions derived from the battery optimization using forecast model $m$ for day $d$. For daily returns $r_{d, m}$ of $m$ and day $d$ we calculate the total profits and the Sharpe ratio sharpe1966mutual as
For the CVAR optimization, we also evaluate the ratio of Value-at-Risk (VaR) exceedances. The VaR at level $\alpha$ is defined as the ($1-\alpha$) -- quantile of the return distribution, that is, $\text{VaR}_\alpha = \inf \{x \in \mathbb{R} : P(r_{d,m} \le x) \ge \alpha\}$. The VaR exceedance ratio is then defined as
This can be both seen as a measure of the economic performance, and as a measure of the calibration of the CVAR forecasts, as for a well-calibrated CVAR forecast, the VaR exceedance ratio should be close to $1-\alpha$.
We evaluate the forecast models using standard probabilistic scoring rules and economic metrics derived from the battery optimization experiments. We follow the general principle of sharpness subject to calibration gneiting2007probabilistic. The following sections provide an overview of the measures of calibration and scoring rules. We denote the true value with $\ensuremath{\bm{\mathrm{y}}}$ and the forecasted distribution with $\mathcal{F}$, represented by an ensemble of $M$ scenarios $\widehat{\ensuremath{\bm{\mathrm{F}}}} \sim \mathcal{F}$ for $m = 1, \ldots, M$. We denote the mean and median forecasts as $\widehat{\ensuremath{\bm{\mathrm{\mu}}}}$ and $\bar{\ensuremath{\bm{\mathrm{\mu}}}}$. The data set of prices has $D$ days and $H$ hours per day and accordingly we aggregate by averaging. We test the statistical significance of differences in the scores by a Diebold-Mariano test diebold2002comparing, diebold2015comparing, nowotarski2018recent.
A calibrated probabilistic forecast correctly represents the probabilities of the observed data. For univariate and quantile forecast, marginal calibration can be easily checked by comparing expected and observed frequencies.
where $\hat{Q}_\alpha$ is the predicted quantile at level $\alpha$. For a well-calibrated forecast, the marginal calibration should be close to zero for all $\alpha$. Visually, this can be checked by plotting the marginal calibration curve, which plots the observed frequency of the event $\ensuremath{\bm{\mathrm{y}}}_t < \hat{Q}_t^{\alpha}$ against the predicted quantile level $\alpha$. For a well-calibrated forecast, the marginal calibration curve should be close to the diagonal.
The following scoring rules are used throughout the empirical evaluation. The MSE, MAE, CRPS and ES are strictly proper scoring rules for the mean, median, univariate and multivariate probabilistic forecasts. The DSS is only (strictly) proper for the mean-covariance structure of the predictive distribution. The VSS is not a strictly proper scoring rule. We keep the discussion of these scoring rules brief, as they are commonly used in PEPF. We have:
where $\widehat{\ensuremath{\bm{\mathrm{\mu}}}}$ and $\widehat{\ensuremath{\bm{\mathrm{\Sigma}}}}$ are the sample mean and covariance matrix of the ensemble forecasts $\widehat{\ensuremath{\bm{\mathrm{F}}}}$. The CRPS and the ES are estimated using the energy estimator gneiting2007strictly,zamo2018estimation. For further information on the scoring rules we refer to the original papers scheuerer2015variogram,dawid1999coherent,gneiting2007strictly and recent works on discriminatory power of scoring rules pinson2013discrimination, marcotte2023regions, alexander2024evaluating and potential asymmetric effects buchweitz2025asymmetric.
A particular issue is the evaluation of CVAR forecasts, as the CVAR is not elicitable on its own, but only jointly with the value-at-risk fissler2016higher, fissler2015expected. For the evaluation of CVAR forecasts, we therefore use the strictly proper joint scoring rule for the (VaR, CVAR) proposed by fissler2015expected. For a forecast of the VaR and CVAR at level $\alpha$, denoted as $(v, e)$, the scoring rule is defined as
where $G_1(v) = v$, $G_2(e) = \exp(e)/(1 + \exp(e))$ and $\mathcal{G}'_2(e) = G_2(e)$ (see fissler2015expected for details). The first term corresponds to the pinball score for the VaR forecast, while the second and third term correspond to the scoring rule for the CVAR forecast.
For the case of point forecasting and battery trading strategies, the use of correlation measures has been proposed as scoring rules by maciejowska2025statistical. For probabilistic forecasts, we propose the using a association-based scoring rule based on Kendall's $\tau$ kendall1938new, defined as
where $\tau(\ensuremath{\bm{\mathrm{v}}}, \ensuremath{\bm{\mathrm{w}}})$ is the Kendall's $\tau$ correlation between two vectors $\ensuremath{\bm{\mathrm{v}}}$ and $\ensuremath{\bm{\mathrm{w}}}$. Note that we can calculate the score either on the price ensembles or on the rank ensembles, since Kendall's $\tau$ is invariant under strictly increasing transformations. Kendall's $\tau$ also relates to the dependence structure of the forecast's copula and thus provides a natural way to evaluate the dependence structure gijbels2011conditional,ziel2019multivariate.
maciejowska2025statistical propose two BESS oriented scores for point forecasts: The mean profit deviation (MPD) and the mean hour deviation (MHD), defined as
where $h_\text{min}$ and $h_\text{max}$ are the hours with the lowest and highest price. The MPD can be seen as a simplification of the returns of a 1-hour battery, relaxing the assumption of charging before discharging. The MHD is closely related to rank-based scoring rules, which we will treat in the following section.
We denote the ranks of the observed prices $\ensuremath{\bm{\mathrm{p}}}_d$ as $\ensuremath{\bm{\mathrm{\pi}}}_d$ and the rank forecast derived from the predicted ensemble forecast $\widehat{\ensuremath{\bm{\mathrm{F}}}}_d$ as $\widehat{\ensuremath{\bm{\mathrm{\Pi}}}}_d$. We therefore have a predictive distribution over the ranks of the observed prices. Note that we have as many ranks as hours, i.e. $K = H + 1$. In the marginal case, rank based scoring rules can take two views:
In the following, we discuss scoring rules for the item and rank view and propose a new scoring rule for the joint distribution of the ranks based on the kernel scoring rule framework. The Brier score for the rank and item view is defined as
Note that is formulation corresponds to the multi-class Brier score, which is bounded in [0, 2]. For the marginal distribution of the ranks (item view), the ranked probability score (RPS) is a strictly proper scoring rule epstein1969scoring. The RPS is defined as
Again, we refer the reader to the original papers for further details on the scoring rules gneiting2007strictly,epstein1969scoring and recents works on the properties of rank scoring rules constantinou2012solving,du2021beyond, wheatcroft2021evaluating.
It is common to look at the performance of ranking measures for the top-$k$ and bottom-$k$ ranks. In the context of battery trading strategies, seems natural to define the relevant top-$k$ and bottom-$k$ as the product of the duration of the battery and the number of (ceiled) cycles per day. For example, for a 2-hour battery with a with one cycle per day, we would be interested in the top-2 and bottom-2 ranks. For a 4-hour battery, we would be interested in the top-4 and bottom-4 ranks -- the same as we would be for a 2-hour battery with two cycles per day. It is important to note that in this case, we have to consider the relevant structure of the battery optimization problem, that is, we can only discharge after charging and are thus interested in the top-$k$ ranks for the hours following the bottom-$k$ ranks. We therefore evaluate the top-$k$ and bottom-$k$ ordered by ranks, that is, we look at the top-$k$ which are preceded by at least $k$ hours with the bottom-$k$ ranks, which we denote as the BESS-$k$ scores.
This section presents the results. We discuss the results of the statistical forecast evaluation and subsequently the results of the economic evaluation. To make matters easier for the reader, we consistently use the color scheme for the for tables/heatmaps and for the models.
This section briefly discusses the results of the forecast evaluation in terms of the scoring rules. Table (ref) shows the scores for the different forecast models and scoring rules. Table (ref) shows the top-$k$ Brier scores for the different forecast models and different definitions of top-$k$. Figure (ref) shows the Brier scores by rank and hour. Marginal calibration is shown in the left panel of Figure (ref), while the right panel of Figure (ref) shows the marginal calibration curve.
Generally, the results follow an expected pattern, with the distributional regression model with the fitted dependence structure (DLENAR-DEP) performing best across all scoring rules, followed by the distributional regression model with the fitted dependence structure (DLENAR-DWD) and the distributional regression model with independence (DLENAR-IND). The LEAR-based benchmark models follow, with the LEAR-N(0, $\Sigma$) performing slightly better than the LEAR-BS. The climatology and Naive-BS perform worst. For the DLENAR-IND and the DLENAR-DWD and DLENAR-DEP, the difference between the ES is not as pronounced as for the other scoring rules, a potential indication that the ES is less sensitive to the dependence structure than other multivariate scoring rules alexander2024evaluating, as the marginal model for these three is the same.
For the rank-based scores, the RPS and KS scores show a similar pattern as the ES, VS and CRPS. The Brier score shows a somewhat different pattern, as it yields worse scores for the Naive-BS than for the climatology model. This difference can be explained by the fact that the Naive-BS takes last weeks value as the forecast, while the climatology model yields an average over the training set. The predictions of the Naive-BS are hence somewhat more arbitrary, since the weather conditions between this week and last week can change drastically, while the climatology model captures at least the seasonal average. Figure (ref) shows the Brier scores by rank and hour. The scores decrease quite steeply for the lowest ranks and then level off. For the hour/item view, we see that performance is rather stable over the hours, while hour 19 seems to be the easiest to predict, presumably since it is naturally the most expensive hour. For the top-$k$ scores in Table (ref), we see that the BESS-$k$ scores are worse than the top-$k$ and bottom-$k$ scores, which is in line with the intuition that the middle ranks are harder to predict than the lowest and highest ranks. The differences between the different forecast models are not that pronounced, but we can see that the DLENAR-DEP-WD and DLENAR-DEP models perform better than the DLENAR-IND model, which is in line with the results for the other scoring rules.
In the following, we present an economic evaluation based on multivariate probabilistic forecasts. Given the fact that risk-neutral optimization approaches do not profit from probabilistic forecasting, we focus our discussion on the risk-averse case. As in the discussion of simplified BESS optimization before, we report the total profits and the Sharpe ratio for the combination of the battery configurations and forecast models. These results are shown in Figures (ref) and (ref). We also report the Value-at-Risk (VaR) exceedance rates for the different models, which are shown in Table (ref). Figures (ref) and (ref) give the cross-evaluation of the optimal bids' objective values for the different forecast models for the 1-hour, 1 cycle case and the 4-hour, 2 cycle case. The results for the other cases are similar and can be found in the supplementary material. The results for the economic performance measures are as follows.
For the risk-averse optimization, we discuss the results in more detail in the following.
Additionally, we evaluate the decision quality by cross-scoring of the optimal bids derived from the different models. Figures (ref) and (ref) show the results for the 1-hour, 1 cycle battery configuration and the 4-hour, 2 cycle battery configuration. We choose these configurations, as the first one represents a common reference case and the latter can be seen as a reference for hydro-pumped storage lohndorf2023value and as a shorter-duration battery with additional commitments to balancing markets kraft2023stochastic. The upper panel show the pinball loss on the VaR forecast, the lower joint the joint score on the (VaR, CVAR) forecast. The columns denote the model whose optimal bids are evaluated and the rows denote the model whose forecasts are used for the evaluation. The diagonal hence shows the scores for the optimal bids derived from model $m$ evaluated on the forecasts of model $m$. The off-diagonal elements show the scores for the optimal bids derived from model $m$ evaluated on the forecasts of model $m' \neq m$. Figures for the remaining configurations are shown in the supplementary material and show similar results. The main conclusions from these results are as follows.
Concluding, our economic evaluation of different forecasts in a risk-averse setting yields some interesting conclusions. While delivering the best profits and high Sharpe ratios, the LEAR-based models fail to deliver well-calibrated VaR forecasts and show shortcomings in the evaluation of the objective value of the optimal bids, when compared to simple benchmarks such as the climatology model. This is a stark reminder that total achieved profits are no suitable measure to compare model performance in a risk-averse setting. The stark differences in the VaR exceedance rates and the cross-scoring are not reflected in the total-profit ranking alone.
Simplified case studies are popular for the monetary evaluation of forecasts through battery trading strategies. The predominant battery configuration is a 1-hour battery with one cycle per day, allowing one trade per day maciejowska2025statistical, chȩc2025extrapolating,serafin2025data. This last constraint is crucial, as it allows us to use a simple dynamic programming approach to optimize the battery trading strategy. Arguably, for risk-neutral optimization, this constraint does not lead to suboptimal strategies, as we just select the hours with the highest spreads. However, for risk-averse optimization, this constraint can lead to suboptimal strategies, as we might want to place multiple bids to diversify the risk.
Figure (ref) shows the total profits, Sharpe ratio and the number of no-bid days for the DP approach and the MILP approach for a 1-hour battery with one cycle per day. The results are shown for the expected profit optimization and the CVAR optimization at levels $\alpha \in \{0.90, 0.75, 0.50\}$. In the following, we discuss some key insights from these results. We denote with DP-1 the stylized optimization, with MILP-1 the MILP optimization with the same constraint and with MILP-24 the MILP optimization allowing for up to 24 bids per day.
In this case, the simplified approach overstates the effect of modelling the dependence structure for the DLENAR models, and, at the same time, understates the performance of naive strategies such as the climatology model, which can still yield good diversification gains on the marginal. This is an indication that the simplified approach might lead to spurious conclusions about the performance of the forecasts, especially for risk-averse optimization.
In this section, we briefly discuss the results of the QBTS for the different forecast models and potentially spurious conclusions. The results are shown in Figures (ref) and (ref), giving the total and per MWh profits and the acceptance probabilities (AP) for the different forecast models. We derive the AP based on the ensemble forecasts, taking the dependence structure into account: $ \widehat{\text{AP}} = \frac{1}{M} \sum_{m=1}^M \mathbb{I}\{ \hat{Q}^{1-\alpha}_s \leq \widehat{\ensuremath{\bm{\mathrm{F}}}}_{s,m} \cap \hat{Q}^{\alpha}_b \geq \widehat{\ensuremath{\bm{\mathrm{F}}}}_{b,m} \}. $ We denote the nominal prediction interval coverage as $\lambda = 1 - 2 \alpha$, as we have $Q_b^{1-\alpha}$ and $Q_s^{\alpha}$ as upper and lower bounds for the QBTS. The DLENAR models yield the highest total profits for the QBTS, followed by the LEAR-based models. The LEAR-BS and the Naive-BS are close, while the climatology model yields the lowest profits across all values of $\alpha$. In terms of the per-MWh profits, the LEAR-based models yield higher profits for small nominal prediction intervals, while the DLENAR models yield higher profits for larger nominal prediction intervals. The last panel gives the difference between expected profits and realized profits. Here, most models yield higher than expected profits. This effect is, however, defined by three effects: the level accuracy, the over- and underdispersion of the forecasts and the specification of the dependence structure. The difference is the largest for low nominal coverage and decreases with higher coverage. Interestingly, for the climatology model, the difference increases again for $\lambda > 0.8$, potentially due to the insufficient point accuracy of the model.
In Figure (ref) we compare the realized acceptance probabilities with the expected AP under independence, and the expected AP under the dependence structure of the forecasts. The realized AP are below the expected AP under independence, which is in line with the fact that the QBTS does not take the dependence structure into account and hence overestimates the AP. This is in line with the results from Section (ref), especially the Remark (ref) and Figure (ref). Especially for the Climatology model, we can also see the clear correlation between the AP and the profits. Around $\lambda = 0.65$, the empirical AP increases and the profits increase as well, but due to overdispersion, also more loss-making trades are accepted and the profits per MWh decrease. The difference between the expected AP under the forecast $\widehat{\ensuremath{\bm{\mathrm{F}}}}$ and the realized AP in the second panel of Figure (ref) is defined by three sources of error. For the DLENAR-IND model, we see that the difference is close to the theoretical AP in the left panel, since the model assumes independence. The DLENAR-DWD and DLENAR-DEP have the lowest difference between expected and empirical AP. The largest differences appear for the Climatology, the LEAR-BS and the Naive-BS, while the LEAR-N(0, $\Sigma$) is somewhat in the middle. Generally, we see that for higher nominal prediction interval coverage, the difference increases.
These results emphasize that the QBTS can lead to hard-to-interpret and potentially spurious conclusions about the performance of probabilistic forecasts. Disentangling the different effects of level accuracy, over- and underdispersion and the dependence structure is not straightforward, but crucial for a proper interpretation of the results.
Probabilistic forecasts are often proposed to enhance decision making, especially under risk. The economic evaluation of (probabilistic) forecasts is increasingly popular in electricity price forecasting and battery storage arbitrage has emerged as natural showcase application for the economic gains of better forecasts. However, the economic evaluation of forecasts, especially of probabilistic forecasts in a risk-averse setting, is not straightforward. This paper addresses two issues:
Our results shed light on the connection between statistical scoring rules for probabilistic forecasts and measures of decision quality for stochastic, risk-averse optimization problems. Similar to previous work on point forecasting nitka2023combining,serafin2025data, we see that measures of decision quality and scoring rules disagree also in the probabilistic setting. While we acknowledge that the results of this paper are based on a specific case study, we believe that the insights are more general and can be applied to other settings as well. Our research touches upon a number of potential avenues for future research.
We conclude by summarizing the key findings of this paper in a recommendation: the economic evaluation of probabilistic forecasts should be treated as an additional layer, not only as forecast evaluation, but as decision quality evaluation. A robust procedure therefore combines (a) forecast evaluation using (strictly) proper scoring rules, and (b) objective-aligned decision diagnostics, and (c) the evaluation of economic performance measures. Agreement across all measures gives strong evidence for model quality, while disagreement can give insights into the strengths and weaknesses of the different models.
Simon Hirsch is employed as industrial Ph.D. student with Statkraft Trading GmbH and gratefully acknowledges the funding and support received. Florian Ziel acknowledges funding in the course of TRR 391 Spatio-temporal Statistics for the Transition of Energy and Transport (520388526) by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation). The authors are grateful for interesting discussions with Daniel Gruhlke and David Wozabal.
During the preparation of this work the authors used Github Copilot (multiple LLM models) in order to improve the quality of the code, figures, tables and text. After using this tool/service, the authors reviewed and edited the content as needed and take full responsibility for the content of the published article.
Simon Hirsch: Conceptualization; Data curation; Formal analysis; Investigation; Methodology; Software; Validation; Visualization; Roles/Writing - original draft.
Florian Ziel: Conceptualization; Funding acquisition; Methodology; Project administration; Resources; Supervision; Validation; Writing - review & editing.
{0pt}