Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
97,026 characters · 22 sections · 147 citation commands
From day-ahead to mid and long-term horizons with econometric electricity price forecasting models
\def\spacingset#1{ {#1}} \spacingset{1}
\if00 \fi
\if10 {
} \fi
{\it Keywords:} Electricity, power, price, forecasting, mid-term, long-term, renewables, load, energy crisis, econometric models, unit root, lasso.
\spacingset{1.45}
The global energy crisis that began in mid-2021 led to record-high european power prices by 2022 prompting regulators to reconsider energy policies and many investment projects to be reevaluated amid spiked future uncertainty (Figure (ref)) emiliozzi2023european, adolfsen2024gas, neri2023energy. Although prices have somewhat stabilized since then, they remain on average around twice the pre-crisis levels and exhibit much higher volatility segarra2024electricity.
This unprecedented situation raises critical questions: To what extent can existing power price forecasting models still be applied in the current environment? Can they be extended to forecast even further into the future, and if so, how far? These questions are vital also from a practical perspective as the crisis has heightened interest in mid to long-term forecasts among multiple market actors. Governments need such models for policy making reasons such as inflation containment and revision of subsidy schemes fabra2023reforming, kotek2023can, neri2023energy, adolfsen2024gas. Industrial customers, producers, investors and asset managers are also highly interested in longer term forecasts for power purchase agreement (PPA) contracting, hedging, and valuation purposes segarra2024electricity. Accurate forecasts are essential as they can reduce perceived investment risk and enhance the liquidity of futures markets.
Electricity price forecasting methods can be broadly classified into two categories: data-driven and fundamental models weron2014electricity, ziel2016electricity. Figure (ref) illustrates the relationships between these two types of models, highlighting their strengths and weaknesses.
Data-driven models, such as econometric, machine-learning, or artificial intelligence (AI) models, utilize historical data to forecast power prices and are primarily employed for short-term forecasting cuaresma2004forecasting, weron2014electricity, uniejewski2019understanding, ziel2018probabilistic, narajewski2020econometric. They have high short-term forecasting accuracy mainly due to autoregressive terms, as they can capture strategic and speculative behavior bello2016medium. However, they might have difficulty in incorporating market mechanisms and dealing with structural changes marcos2019electricity. In contrast, fundamental models replicate the market mechanisms that lead to price formation. They take into account the expected market conditions and power plant fleet within a geographic zone, deriving power prices from intersections of supply and demand curves or as shadow prices of the demand constraint kallabis2016plunge, beran2019modelling, pape2016fundamentals. These models are typically used for long-term forecasting, as they require detailed information on the power plant fleet, including technical parameters, fuel costs, availabilities and complex production schedules ringkjob2018review. They are mostly not open source and require significant computational resources and technical expertise to implement. They are not as accurate in the short-term as data-driven models as they tend to underestimate volatility beran2021multi. However, they are highly interpretable and causal conclusions can be drawn from them, which are essential properties for correct model assessment in the context of long-term forecasting.
This study explores the potential of extending data-driven models to long-term forecasting. Since interpretability and model transparency are crucial for such horizons we restricted our analysis to regularized linear regression models. In doing so we encountered several key obstacles, each addressed with a novel combination of approaches:
While there is an abundance of literature on short-term electricity price forecasting, research on mid to long-term forecasting remains relatively scarce, especially for an hourly resolution and for the energy crisis period between 2021-2023. ziel2018probabilistic compiled an exhaustive list of papers covering econometric short, mid and long-term power price forecasting published between 2011 and 2017. Among these, only four papers, including the cited one, addressed long-term forecasting, while thirteen papers focused on mid-term forecasting. The authors refer to forecasting horizons from one month to one year ahead as mid-term, while long-term are horizons beyond one year. We will be using this definition as well.
Since then, more papers in mid to long-term forecasting of power prices have been published, which are systemically summarized in Tables (ref) and (ref). The papers were found using the Scopus search engine with queries as in ziel2018probabilistic, which can be found in the Appendix (ref). We report five key aspects of each paper: the market, the forecasting horizon, the period covered, the regressors used, and the forecasting method. The most common forecasting methods are autoregressive models, neural networks, and support vector machines. The most common regressors are autoregressive terms, with many papers not including external regressors at all. The markets covered are mainly European and US markets, with the most commonly ones being the PJM and ISO-NE regions in the US, Germany, and the Iberian Electricity Market (MIBEL). The forecasting horizons range from one day to five years, with most papers focusing on horizons between one week and one month ahead. The periods covered range from 2002 to 2023, with most papers focusing on the years 2016-2021. Even recent papers tend to use data before the recent energy crisis.
Some results of the query included false positives which we removed. For example najafi2023application forecast day-ahead prices but refer to them as mid-term forecasts, and aruldoss2021week refer to them as weekly forecasts. Other papers do not focus on market power prices, such as yang2022electricity who analyze company-level prices, and others such as jkedrzejewski2021importance and marcjasz2019importance forecast day-ahead prices but include long-term components as regressors. The majority of papers focus on mid-term forecasting with forecasting horizons under a year, many under a month. Many studies model average monthly, weekly or yearly prices. None of the papers tackle forecasting power prices on an hourly basis around the 2022 energy crisis and the period thereafter.
To illustrate the obstacles encountered with econometric models when forecasting far into the future, consider the following established day-ahead power price model, which we will refer to as the expert model chai2023forecasting, bille2023forecasting, marcjasz2023distributional, maciejowska2020pca, fezzi2020size, serafin2019averaging, ziel2018day, uniejewski2017variance:
where $h$ is the forecasting horizon and $h=1$ corresponds to forecasting one day ahead.
Figure (ref) depicts the estimated coefficients of this model across various forecasting horizons, ranging from one day to 360 days ahead, before and after the energy crisis. The estimates exhibit high instability with increasing forecasting horizon $h$, particularly following the energy crisis. The effects of renewables infeed, load, and autoregressive terms initially diminish rapidly, in line with expectations. However, they sometimes resurface at higher horizons without a discernible pattern and with different signs. The coefficient estimates of the commodity futures prices are also sometimes negative, which is not in line with energy economics theory, since these represent fuel prices of power plants. All of these can be traced back to the obstacles listed in the beginning.
While some of these issues are addressed in a few papers (see Tables (ref) and (ref)), they are not considered as a whole. The provided solutions are often very specific to the pre-crisis period or the considered forecasting horizon, making it questionable if they remain robust when varying these two factors.
For incorporating fundamental information, gonzalez2011forecasting, marcos2019electricity, beran2021multi, and gabrielli2022data use a hybrid approach by including the price returned by a fundamental model as an explanatory variable into the econometric model. We incorporate fundamentally-derived information by setting coefficient bounds, thus stabilizing the model and reducing spurious effects.
Regarding the inclusion of short-term regressors, ziel2018probabilistic simulate wind and solar generation averages for long-term forecasts. gabrielli2022data use yearly wind and solar data. wagner2022short include renewables for short-term but not for long-term forecasts. Other studies that model horizons up to a month use renewables as lagged regressors. We provide a method to incorporate variables such as renewables by generating seasonal forecasts and projecting them into the future.
Studies such as nowotarski2016importance, marcjasz2019importance, and jkedrzejewski2021importance examine long-term seasonal components of power prices for day-ahead forecasts, but do not address unit root behavior drivers and their implications for mid to long-term forecasting. We show that unit root behavior is explained by commodity prices, though this diminishes with increasing horizons. We also show that after a certain threshold the best approach is to estimate the current unlagged relationship between power prices and their regressors, and projecting it into the future.
This paper is organized as follows. In the next section (ref), we provide a brief overview of the data used and present the basic model and study design. In section (ref), we tackle each of the challenges presented above and present modifications of the basic model to address them. In section (ref) we report on the accuracy of the new proposed models and compared them to some benchmarks. We offer a guideline of which input variables are important and what methodology performs best for which forecasting horizon. Lastly, in section (ref) we conclude and present directions for further research.
In this study we make use of the following data pertaining to Europe's largest energy market, Germany: day-ahead power prices, infeed from solar, onshore wind and offshore wind plants, and load data. This data is in hourly resolution and is freely available on the ENTSOE Transparency platform. We also use daily closing prices for futures of the commodities gas, coal, oil and european emission allowances (EUA), further referred to as $\textrm{CO}_2$, which were provided by the information platform Refinitiv Eikon. The data spans 2015 to 2024. We apply standard clock-change adjustment to the hourly data (interpolation in spring, averaging in fall).
Figure (ref) shows power prices, renewables infeed broken down by solar, onshore wind and offshore wind, and load (demand) data for two weeks of 2024. Typical power price characteristics can be observed: first, the weekly seasonality as a result of lower load during weekends, second the daily seasonality with lower prices during mid-day as a result of solar infeed, and third the strong impact of renewables infeed on power prices, also known as the merit order effect of renewables. This states that when renewables infeed, especially infeed from wind, is high, then power prices become lower often reaching negative values.
Figure (ref) illustrates mean and standard deviation statistics of power and commodity prices highlighting some of the same characteristics. It is evident that, until the energy crisis started in late 2021, power prices remained relatively stable around a mean of 40 EUR/MWh. With the onset of the energy crisis they surged to over 800 EUR/MWh by 2022 before subsiding to a mean slightly below 100 EUR/MWh since 2023. Nevertheless, the volatility remains at over 25 EUR/MWh which is notably higher than pre-crisis levels. A similar pattern is observed for the futures prices of $\textrm{CO}_2$, gas, coal, and oil. As these commodities are used as fuels in the generation of electricity, their prices are key drivers of power prices. Indeed, their long-term co-movement with the power prices is evident in the time series, especially for gas and coal (see Figure (ref)).
The recent energy crisis which started in late 2021, its timeline and effects are well documented di2022natural, medzhidova2022return, szpilko2022european, emiliozzi2023european. In summary, it was triggered by perceived scarcity of gas supply in Europe mainly as a result of Russia's reduction of gas supply amid rising geopolitical tensions which ultimatively led to the currently ongoing Russia-Ukraine conflict. Nevertheless, the energy crisis comes in the more complex context of the post-pandemic recovery and the unfolding of the European Green Deal.
Model structures for short-term power price forecasting are well-researched and regressors with high explanatory power have been identified in the literature chai2023forecasting, bille2023forecasting, marcjasz2023distributional, maciejowska2020pca, fezzi2020size, ziel2018day, uniejewski2017variance. These models are referred to as expert models as they are current state-of-the art models with solid theoretic background and empirical backup weron2019electricity. For day-ahead forecasts variables with high explanatory power include:
The expert model ((ref)) is a representative expert model with all of these regressors, except for the highest and lowest prices of the previous day. The problems with the expert model become apparent when looking at Figure (ref). The coefficients are as expected for the day-ahead forecast, but become unstable with increasing horizons and many apparently spurious effects can be observed. Hence the performance of the expert model is expected to decline drastically with increasing forecasting horizons.
Throughout this study we consider variations of the expert model ((ref)). For every considered model we conduct a rolling window forecasting study with a window size of $365 \times 3$. Here the window size refers to the number of rows of the regressor matrix. The window is rolled forward by one day after each forecast for an evaluation set of 6 years from April 2018 to April 2024. For each day 24 models are used, one for each hour of the day. Hence, at every rolling window step 24 models and sets of coefficients are estimated. We fit the models on normalized\footnotemark[4]\footnotetext[4]{By normalization we refer to substracting the sample mean and dividing by the sample standard deviation of a variable in order to get zero mean and unit variance.} data using elastic net as implemented in the glmnet R package glmnet1. The elastic net penalty parameter $\alpha$ is set to 0.5, which is a balanced mix of LASSO and ridge regression. The regularization parameter $\lambda$ is chosen using the Bayesian information criterion (BIC) from an exponential grid of $\lambda$. The models are evaluated using the root mean squared error (RMSE) and the mean absolute error (MAE) as performance measures.
In this section we will thoroughly discuss each of the challenges associated with forecasting power prices beyond day ahead (see Figure (ref)) and inspect ways of addressing these.
In the absence of restrictions which would force the expert model to adhere to energy economic principles, the estimated coefficients become highly unstable (see Figure (ref)). Fundamental information refers to the underlying economic and physical principles that drive power prices. In the context of power markets this includes how the day-ahead market mechanism works to form the power price. The simplest model of a power market is the supply stack model, also referred to as the merit order model. According to this model the marginal costs of power plants are ordered from lowest to highest generating the merit order curve and the power price is determined as the intersection of this curve with the inelastic load. In the following we will show that the expert model has an implicit merit-order model representation and hence that its coefficients are closely related to some technical parameters of the power plants. Thus coefficients can be constrained to be consistent with fundamental theory and as a result provide much needed stability and interpretability.
To illustrate how the expert model relates to the merit order, we consider a market where power is produced only by a single gas power plant. The simplified variable costs of this plant depend on the gas price, the $\textrm{CO}_2$ price, and on the characteristics of the power plant. They are calculated as:
where $\eta_{Gas}$ is the efficiency and $\varepsilon_{Gas}$ is the $\textrm{CO}_2$ intensity factor of the power plant. According to the merit order model, the power price is equal to the marginal costs of the price-setting technology. In this case, the only technology is gas, hence $\textrm{Price}_t = \textrm{VarCost}_t$ together with ((ref)) results in
Now consider a simplified econometric model for the power price with only two explanatory variables, namely gas and $\textrm{CO}_2$ prices:
Since this is equivalent to ((ref)), the corresponding coefficients $\beta_{Gas} = \frac{1}{\eta_{Gas}}$ and $\beta_{\textrm{CO}_{2}} = \frac{\varepsilon_{Gas}}{\eta_{Gas}} $ are equal to the inverse of the power plant efficiency, also known as the heat rate, and the ratio of $\textrm{CO}_2$ intensity to efficiency respectively.
Figure (ref) shows the merit order representation of the same illustrative example for an efficient gas power plant, an inefficient one, and for a coal power plant. Here the components of the power price are shown as colored areas. Note that for the more inefficient gas plant shown in the middle, the gas price component is higher than for the efficient plant, since more gas is needed to produce the same amount of electricity. For the coal plant on the right the $\textrm{CO}_2$ component is higher than for the gas plants, since coal produces more $\textrm{CO}_2$ per MWh as illustrated by the higher $\textrm{CO}_2$ intensity factor.
Now consider a market where all of the three technologies from Figure (ref) are simultaneously active. We consider the simplified power price model:
Similar to ((ref)), $\beta_{Gas}$ will correspond to the average heat rate of the gas power plants weighted by the corresponding capacities, and $\beta_{\textrm{CO}_{2}}$ will correspond to the weighted average $\textrm{CO}_2$-to-efficiency ratio of all plants. However, since these are averages, $\beta_{Gas}$ will never exceed the highest individual heat rate nor will $\beta_{\textrm{CO}_{2}}$ exceed the highest individual $\textrm{CO}_2$-to-efficiency ratio. Hence, using the notations from Figure (ref), we can derive upper bounds for these coefficients:
Doing such calculations for the expert model ((ref)) requires knowledge of the power plant characteristics, which were summarized by beran2019modelling and are also openly published by ENTOSE in the European Resource Adequacy Assessment (ERAA) report \footnotemark[5]\footnotetext[5]{Link: https://www.entsoe.eu/outlooks/eraa/2023/eraa-downloads/}. We summarize these characteristics in Table (ref). Note that the units of the coefficients must be chosen in such a way that the unit of the variable multiplied by the unit of the coefficient equals the unit of the power price (EUR/MWh). Now, upper bounds for the fuel coefficients can be derived similar to ((ref)). The resulting lower and upper bounds for the coefficients are shown in Table (ref) and the corresponding calculations can be found in the Appendix (ref). In addition to the fuel prices we also restrict RES to only have a negative effect on the power price, since the marginal costs of renewables are close to zero and hence they push the power price downwards according to the merit order model. Similarly, we restrict the load coefficient to be positive, since increasing demand should increase the equilibrium price according to standard economic theory. Autoregressive effects were also restricted to be positive. This is because we noticed that fuel and autoregressive coefficients can be replaced by one another. For example, when fuel coefficient is unconstrained, the autoregressive coefficient is small. When the fuel coefficient is constrained, the autoregressive coefficient becomes as relevant as the previously unconstrained fuel coefficient. This is expected since the time series strongly co-move in the long term. To mitigate this we impose the positivity constraint, which is plausible in this context since autoregressive terms should not be relevant too far out into the future once seasonalities are accounted for.
Figure (ref) shows the estimated coefficients of the constrained expert model (ref) using these bounds. Much more stable coefficients can be observed. This allows for some interpretation. First, the coefficients of short-term regressors such as renewables, load and autoregressive terms rapidly diminish to zero with increasing forecasting horizon. Nevertheless, these terms reappear at higher horizons without clear patterns, hinting towards remaining spurious effects. Second, the weekend and seasonal effects become relevant with increasing horizons. This stems from the fact that seasonal effects are mainly driven by load and renewables, for which there are no accurate forecasts. Third, gas and EUA seem to be most relevant for short forecasting horizons, while coal and oil become more relevant towards middle horizons. Towards the end EUA seems to be the most relevant fuel variable. However, it is unclear if these are causal effects or spurious correlations. We will tackle this aspect in the corresponding section below.
In the constrained expert model we used the front-month futures of the commodities $\textrm{CO}_2$, gas, coal, and oil. However, there are multiple futures available each day with maturities ranging from of one day up to several years. These are the prices that can be locked in for selling or buying the commodity at fixed time points in the future. They are highly correlated as shown in Figure (ref) and theoretically represent the market participants' expectations of future prices under a risk-neutral measure. Sometimes the longer maturity futures are traded at higher prices than the shorter maturities or vice-versa leading to upward (contagos) or downwards (backwardation) sloping futures curves. This information is expected to be relevant for power price forecasting, as market participants often hedge their fuel costs by entering into futures contracts thus forming a portfolio caldana2017electricity, steinert2019short. We define $D0$ to be the delivery period corresponding to the forecasting horizon $h$ such that the model
represents the power price $h$ days ahead explained by average gas future prices of the last 12 months with delivery in $h$ days from today. We can extend the model to include also the trailing $D-1$ and leading $D+1$ months to delivery. Intuitively we would expect the future with a maturity corresponding to the forecasting horizon to be most relevant. However, when we extend the constrained expert model by including the terms in ((ref)) with additional $D-1$ and $D+1$, we observe that the $D0$ future is not always the most relevant, but rather the most recent and liquid one. For example, it seems that when forecasting one month ahead, the current front-month future $D0$ is the most relevant. When forecasting 4 months ahead, it seems that the current trailing month future $D-1$, which is the 3-month future, is the most relevant, because it is closer to the present and hence more liquid and more informative. The lag structure used and the results are included in the Appendix in Tables (ref) and (ref) respectively. This observation aligns with the concept of an efficient market, where all available information is rapidly incorporated into the latest prices resembling a Markov chain dynamic. This represents some evidence in favor of always using front-month futures for every forecasting horizon. Furthermore, the expectations implied by the shape of the futures curve rarely happened. For example, at the start of the energy crisis in late 2021, the 12-month coal and gas futures were traded at a discount compared to the front-month futures, which implied that the market expected prices to decrease in the future. However, this was not the case, as prices increased dramatically over the next year. The effect was observed in early 2023. This gives some intuition as to why including multiple futures might not be as relevant as expected in a price forecasting model.
Short-term regressors refers to explanatory variables that can not be accurately forecasted beyond a few days. These are variables derived from meteorological data such as infeed from solar and wind, but also load since it is partly influenced by temperature and wind speed. Infeed from RES and load have very strong impacts on the day-ahead power price and are crucial for explaining effects such as negative prices, the merit order effect of renewables, and the weekend effect of load (see Figure (ref)). However, since it is only possible to accurately forecast these variables one day in advance, power price models beyond this meteorological forecasting horizon do not benefit from these variables. This is evident in Figure (ref)a, where the coefficients of renewables infeed and load are essentially only relevant for the day-ahead forecast. This is because the model only includes day-ahead forecasts of these variables.
It is however possible to forecast the daily, weekly and annual seasonalities as shown in Figure (ref). Here, short-term shocks can not be foreseen, but rather only broad observations such as that there is higher solar irradiation around noon and during summer, there is stronger wind during winter and spring, that load is higher on weekdays and in winter, and lower on weekends and in summer. Furthermore, the expected expansion of renewable capacity and development of load can be included in the forecast, which we did in the form of linear trends.
We can now include these forecasts in the expert model. However, simply regressing the power price on these forecasts would still greatly underestimate their contribution as some effects, such as negative price shocks due to sudden high RES infeed, would not be explainable. To address this, the RES and load coefficients can be estimated by regressing the power price on their actual values and then using the seasonal forecasts to generate power price forecasts. We refer to this approach as the current model and introduce it in the next section.
A widely accepted fact in finance is that stock prices exhibit unit root behavior, meaning that they are non-stationary and possess a stochastic trend. This assumption is evident in the modeling of stocks as Brownian motions, also referred to as Gaussian or Wiener processes, or functions thereof for structuring and valuation purposes, such as pricing derivatives black1973pricing, schwartz2000short, hull2009optionen. This unit-root behavior extends to commodities as well schwartz2000short. In fact, unit root tests for $\textrm{CO}_2$, gas, coal and oil futures fail to reject the null hypothesis of having a unit root (see Figure (ref) in the Appendix) berrisch2023modeling. As these commodities represent fuels for power generation, it is reasonable to assume that the power price inherits some unit root behavior from them, which is particularly visible in the long run (see Figure (ref)).
Some evidence for this is provided by considering the differenced expert model, where instead of modeling the levels of each component of the expert model we model their one-period differences, which is standard practice for eliminating a stochatic trend. The results are shown in Table (ref). There are two important observations. First, differentiation when fuels are not included sometimes improves the model, especially for the energy crisis period, indicating that power price do have a unit root behavior. Second, for the day-ahead case we observe that differentiation does not bring any improvement when fuels are included in the model, indicating that the unit root behavior is captured by the fuel prices. Also, the inclusion of fuels tends to visibly improve the forecast. This effect fades with increasing forecasting horizons since the power price drifts further away from its current value at $t$, which can not be explained anymore by the current values of the fuel prices.
Power and fuels prices could be cointegrated for short forecasting horizons, as they do seem to move together in the long run moutinho2022examining. However, the cointegration relationship might be complex and might change both over time due to complex effects such as fuel switches and supply-stack non-linearities. Furthermore, when the forecasting horizon is large, it is unlikely that a long-term equilibrium relationship even exists. Lastly, our 24-dimensional model structure, where we model each hour individually, and the usage of external regressors does not fit in the classical framework of cointegration and vector error correction models (VECMs), making standard statistical tests and theory unapplicable.
A well-known effect in time series analysis is that regressing a unit root process on another independent unit root process can lead to spurious regression results granger1974spurious, lutkepohl2005new. In our context, we have seen that both the response, which is the power price, and some explanatory variables, such as the fuel prices, exhibit unit root behavior. There is some evidence that load also exhibits long-term unit root behavior smyth2013fluctuations. However, these explanatory variables are not independent of the power price. For short forecasting horizons such as for day-ahead they are important variables with a lot of explanatory power, hence there is little risk of spurious effects. However, with increasing forecasting horizons, the model rapidly approaches the case of a spurious regression.
This effect is illustrated in Figure (ref) where 4 randomly generated brownian motions and white noises were included as regressors into the constrained expert model with positivity constraints.
It becomes apparent that for higher horizons the coefficients of the brownian motions are not always shrunk to zero, indicating that the model is struggling to filter out the stochastic trend from the regressors. This effect is especially prominent for the period after the energy crisis, where the coefficients of the brownian motions are at times exclusively chosen. This is because as the temporal gap between response and regressor widens the explanatory variables stop being informative. Essentially we experience a loss of explanatory power. For example, the RES infeed of today will have some correlation with the infeed of tomorrow but almost none with the infeed one year from now. The same applies to the fuels, but given that they have a unit root behavior they are prone to strong spurious effects, similar to the brownian motions above.
In the previous section we explained how the lag between the response and regressors leads to spurious effects. A natural solution would be to close this gap and estimate the same-day relationships between the response and regressors. This way the model would be able to capture the true relationship between the response and regressors, avoiding spurious effects. However, this approach alone would not be feasible for forecasting, since the regressors at time $t+h$ would be needed, which are not known at time $t$. Therefore, for the forecast we would need to substitute the regressors with their forecasts. The forecasts for RES infeed and load will be their seasonal averages (see Figure (ref)). The forecasts for the commodities $\textrm{CO}_2$, gas, coal and oil will be the futures with maturities corresponding to the forecasting horizon. By doing so the current relationship between response and regressors is captured and projected into the future where it is assumed to still hold. This effectively disentagles estimation from forecasting leading to a two-step approach which we refer to as the current model.
where $(01)$ is the front-month future, which is the most recent available price, and $(h)$ is the future with a maturity corresponding to $h$ as of $T$. Highlighted in red are the main changes from the expert model in ((ref)).
Figure (ref) shows the coefficients estimated from the current model. These are very stable accross the horizons, with oil completely missing and coal only being active for the pre-crisis period. RES and load become major components for all horizons, similar to the day-ahead case of the expert model (see Figure (ref)). Consequently, the weekday dummies become less relevant, as weekly seasonalities are captured by load. Such results are what would be fundamentally expected, as the price-setting technology remains stable for the same day regardless of when it was forecasted. In the first plot in Figure (ref)a price setting plants are coal and gas, while after the start of the energy crisis gas becomes the only price setting technology for this hour. Contrast this to previous models such as Figure (ref)a where the price setting technology switches from coal and gas to oil as the forecasting horizon increases, even though the same day is forecasted every time. For example, according this figure the price setting technology switches between $h=120$ to $h=150$ from less carbon intense oil plants to very high carbon intense oil plants, as the $\textrm{CO}_2$ coefficient is suddenly very high. This is most likely a spurious effect, as also suggested by the high coefficient of the brownian motion at this point. Hence the current model offers much more interpretable results.
In the previous section we introduced three different extensions of the original merit order model. To evaluate their performance and pinpoint their contribution to the forecasting accuracy we consider different combinations of these methods. We also include two benchmark models. First, we include the naive model, which implies a random walk, as it simply returns the last available price if the day to be forecasted is Tuesday through Thursday, or the last available price of Monday, Friday, Saturday or Sunday if the day to be forecasted is one of these. The differentiation by days is done to preserve the weekly seasonality. Second, we include and the weekday dummies model (WD) which represents the average weekday prices. As evaluation criteria we use the mean absolute error (MAE) and the root mean squared error (RMSE) which are defined as:
where $\widehat \varepsilon_{t,h}$ is the forecasting error at time $t$ for hour $h$. Table (ref) shows the RMSE results of the considered models. Table (ref) in the Appendix shows similar results for the MAE.
We observe that while the expert model performs best for the day-ahead setting, for which it was originally designed, other models outperform it for higher horizons. To disentangle the drivers of what makes some models better or worse for differen horizons we looked at different combinations, referred to in Table (ref) as the WD+ models. There are important conclusions that we can draw from the results in Table (ref). The corresponding Diebold-Mariano (DM) tests are shown in Figure (ref). We summarize these findings in Figure X as a decision guideline on what regressors and methods to use depending on the forecasting horizon. For ease of description we split the considered forecasting horizons into three groups: short horizons up to 30 days or 1 month ahead, mid horizons for 1 month up to 180 days or 6 months ahead, and long horizons for 6 months up to 1 year ahead.
First, for short horizons fuels and autoregressive effects are highly important. Models that do not include either perform consistently worse such as WD, WD+RL and WD+RL+C. On the other hand, models WD+F and WD+F+C which only use fuels also do not perform well, but much better in comparison to those that do not. It appears that a combination of RES, load and fuels are important for short horizons. Using fundamental constraints leads to significant improvements compared to the unconstrained model starting from 2 weeks onward.
Second, for mid horizons the current model seems to outperform the rest on average. This is an important inflection point as this is the period when the models that do not include fuels start to outperform models that do. This can be interpreted as the critical horizon after which the unit root behavior of the commodities becomes dominant making them exhibit spurious effects as they become uninformative for future power prices. The current model mitigates this effect by forcing current relationships between the commodities and the power price and projecting them into the future. In essence, at this horizon modelling the power price structure becomes more important than forecasting it. For this horizon all regressors seem to be relevant, but the modelling method is different.
Third, for long horizons models without fuels and autoregressive terms perform best on average. This is in contrast to current literature, where all models include autoregressive terms (see Tables (ref) and (ref)). In fact, all models that simply include fuels as lags perform worse than the WD model. This is a clear indication that the commodities are not informative. The current model performs slightly better than the WD model for two main reasons: the large RES and load coefficients act as model stabilizers (see Figure (ref)) and the fuels coefficients are regressed on actual values thus avoiding spurious estimations. The models WD+RL and WD+RL+C which include only weekday dummies and forecasts for renewables and load. It is interesting to see that simply adding an autoregressive effect as in \textbf{WD+ARL+C} makes the model much worse and has very similar effects to adding fuels (see \textbf{WD+F} and \textbf{WD+F+C}).
In addition to this, we can also inspect the components of the forecasts. Figure (ref) displays the contributions of the regressors to the forecasts for two selected weeks. The sum of the components at each time point equals the forecast. As expected, in the short term we are able to capture negative price shocks due to strong RES infeed while in the longer term we predict average seasonal patterns of the power price.
This study showcases the challenges associated with forecasting beyond short-term horizons in the power market using data-driven econometric models. First, it shows that fundamentally-derived bounds for the coefficients are needed for ensuring model stability and alignment with energy economic principles. It also illustrate how econometric models relate to fundamental models such as the merit order model and that some coefficients correspond to certain technical power plant characteristics. Second, it explains how regressors that are meteorologically limited in their forecasting accuracy to less than a few days ahead, such as renewables, load and temperature, can still be used in a long-term forecasting framework. This is done by generating seasonal renewables infeed and load forecasts which take into account buildout of renewable capacity and other macroeconomic trends. Third, it demonstrates that power prices exhibit a unit root behavior in the long run which is driven by the commodities prices $\textrm{CO}_2$ gas, coal and oil. It lays open the issue of spurious effects becoming highly relevant with increasing forecasting horizons especially since the onset of the energy crisis in 2021. Lastly, it also discusses the idea of portfolio effects where the power price is influenced by a mix of different futures of the same commodity.
To mitigate these issues we propose the current model together with fundamental constraints which is a highly interpretable model robust to spurious effects during estimation. This model can be used for forecasting power prices starting one month ahead. We also provide a guideline as to which explanatory variables are important and which methods perform the best for different forecasting horizons.
Further research could focus on extending the model to include probabilistic forecasts. It would require simulating future values of the regressors, particularly commodity prices, while accounting for their correlation or cointegration relationships. Finding a distribution that would accurately represent the power prices would be challening as it would most likely be multi-modal, with one mode around zero and a right skeweness that drastically changes over the years. Another aspect would be incorporating more complex relationships between the regressors and power prices, such as interactions between fuels. This could account for effects like fuel switches, where two or more fuels change order in the supply stack model. Furthermore, inspection of extending standard cointegration theory and vector error correction to include a 24-hour price model would be a nother area of extension. This could allow better capturing of the relationship between power price and fuel commodities, and potentially offer deeper insight into unit root behavior for longer forecasting horizons.
During the preparation of this work the authors used GitHub Copilot and OpenAI ChatGPT-3.5 in order to generate rephrasing suggestions to improve the language and readability of this paper. After using this tool/service, the authors reviewed and edited the content as needed and take full responsibility for the content of the published article.
Funding: This work was supported by the German Research Foundation (DFG, Deutsche Forschungsgemeinschaft) project number 505565850 to P.G. and F.Z.
The raw data required to reproduce the findings in this work include freely available resources such as day-ahead power prices, actual values and day-ahead forecasts for load, solar, onshore wind, and offshore wind infeed. This data can be downloaded from the ENTSOE Transparency Platform (https://transparency.entsoe.eu/). Data not freely available include daily closing prices for futures of gas, coal, oil, and European Emission Allowances (EUA), accessible via the information platform Refinitiv Eikon at a cost (https://eikon.refinitiv.com/).