EconBase
← Back to paper

The fine structure of electricity price volatility

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

87,310 characters · 19 sections · 27 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

The fine structure of electricity price volatility

abstractWe conduct the first rigorous study of electricity price volatility for the full panel of electricity prices across three European generation zones. By interpreting the observed day-ahead prices as local averages of a latent price process governed by a stochastic partial differential equation, we develop estimators of the weekly integrated variance. The inherently infinite dimensional setting introduce several complications that are not relevant in the conventional finite dimensional semimartingale setting, and we spend considerable effort in dealing with these. In particular, we must account for both mean-reversion in prices and semigroup-smoothing in the estimated variance. We provide a detailed decomposition and interpretation of the empirical estimates across three vastly different European generation zones, namely Germany, Norway, and Spain. Our findings indicate that each zone has very different drivers of volatility, and that the impact of generation variables differs considerably. We document that leverage effects appear to be present at first sight, but disappear once we condition on suitable state variables, thereby showing that electricity price volatility does not generally exhibit asymmetric responses to price shocks.

\noindentKeywords: Electricity markets, realized volatility, functional data analysis, quadratic variation, mean-reversion, leverage effect

Introduction

Volatility is a ubiquitous term in finance and typically refers to the second order structure of the conditional return distribution. It is generally regarded as being the most dominant time-varying feature of the distribution and plays a critical role in, e.g., portfolio management, derivatives pricing, and trading impact AndersenBollerslev1998,AndersenBollerslevDieboldEbens2001,AndersenBollerslevDieboldLabys2003,BNS2002_RV,ChristensenKinnebrockPodolskij2010,ChristensenOomenPodolskij2010,ChristensenPodolskij2012,ChristoffersenJacobsMimouni2010,FrenchSchwertStambaugh1987,RealizedRange2008. Depending on the horizon and data availability, different proxies for the second order structure can be considered, but most modern measures rely on quadratic variation theory.

In this work, the goal is to investigate the conditional second order structure in European day-ahead (spot) electricity markets in a systematic and rigorous way, down to the individual delivery period. While there is a large literature on stylized facts and forecasting of spot price levels, there seems to be less focus on the volatility of the panel of prices. The market for electricity is very different from classical equity markets and it therefore requires different interpretations and modeling tools BenthMonograph,Weron2014. The most notable feature of electricity spot prices is that they evolve as a high-dimensional multivariate time series $P_t\in\mathbb{R}^d$, where each component $P_t^{(i)}$ represents the price of consuming 1MWh of electricity during the $i$'th delivery period. Typically, it is the case that $d=24$ or $d=96$, which, respectively, corresponds to hourly and quarter-hourly delivery periods throughout the day. Unlike stock prices, electricity spot prices can also be negative, which means that log-transformations are not generally feasible, as is otherwise common in the literature on volatility. Other unique features of electricity spot prices that distinguish them from traditional asset prices are the presence of mean-reversion, short-lived spikes, and seasonality effects. Furthermore, the cross-sectional dependence structure of electricity prices exhibits a very particular diurnal pattern as shown in Kloster2026; in particular, electricity spot prices exhibit “cyclicality” in the sense that the earliest and the latest delivery period of the day are more highly correlated than, e.g., the earliest and the mid-day delivery periods. This is generally ascribed to the strong correlation of the underlying supply and demand at these periods, which is largely agnostic to when the price is determined.

These characteristics can be modeled by imposing suitable assumptions within a multivariate model, but it is illustrated in Kloster2026 that one obtains a flexible and parsimonious framework by letting $P_t$ evolve as a function-valued process, such that the daily cross section of prices corresponds to observations along a latent price curve. More precisely, it is shown that when $P_t$ takes values in the separable Hilbert space of square-integrable functions on the unit circle, denoted by $L^2(\mathbb{S}^1)$ where $\mathbb{S}^1$ is the unit circle, then cyclicality of prices is essentially a model-free implication. The functional time series perspective is not new in the context of electricity spot markets EPF_functional1,EPF_functional2,EPF_functional3, but the choice of $\mathbb{S}^1$ as an index for the cross sectional dimension introduced in Kloster2026 yields a particularly tractable framework in which we shall operate. The infinite dimensional perspective have several convenient implications that cannot be modelled by conventional methods, in addition to the diurnal periodic structure induced in the prices from the embedding into $L^2(\mathbb{S}^1)$. It allows for modeling arbitrary and changing partitions of the day into delivery periods, which is relevant as these are not fixed throughout time or generation zone in practice. It also allows for the interpretation of spot prices as functionals of the latent daily price curve, which is a perspective we adopt throughout. Furthermore, the interpretation of both time and cross section as a continuum is very tractable for the purpose of derivatives pricing, which is an area where infinite dimensional models have generally been used to great effect BenthParaschiv2018,AmbitFutures2014,CuchieroEnergy2024, and an area where volatility modeling is particularly important.

To illustrate our framework, suppose that $\lbrace [h_{i-1},h_{i})\rbrace_{i=1}^{d}$ is the partition of delivery periods where $h_{0}=0<h_{1}<\cdots <h_{d}=2\pi$. Each $h_i$ represents a unique time of the day via the canonical mapping onto the circle $h_{i}\mapsto(\cos(h_i),\sin(h_i))=(\cos (i/(2\pi d)),\sin(i/(2\pi d)))$. Let $X_t(h)$ be the instantaneous latent price of electricity on time $h$ of day $t$ and suppose that the function-valued process $X_t(\cdot)$ takes values in $L^2(\mathbb{S}^1)$. Since electricity can be consumed at any point during the delivery period, it is economically reasonable that the observed price, $P_t^{(i)}$, corresponds to the average of the true, but unobserved, price over that period. More precisely, we find it natural to impose that

equation[equation omitted — 110 chars of source]

In this view, each observed price $P_{t}^{(i)}$ is a bounded linear functional of $X_t$, which has several convenient statistical implications. In particular, we do not a priori require the latent process $X_t$ to take values in a “nice” subspace of $L^2(\mathbb{S}^1)$, which is often required to make the evaluation functional $\delta_x f=f(x)$ continuous. By formulating the evolution of $X_t(\cdot)$ as a stochastic partial differential equation (SPDE) in $L^2(\mathbb{S}^1)$, we may construct an estimate of the integrated variance -- which is an operator on $L^2(\mathbb{S}^1)$ -- via methods related to quadratic variation theory. Our methods and results are similar to those of BenthSchroersVeraart2022,BenthSchroersVeraart2024, who study realized covariation for a class of semilinear SPDE's in a high frequency setting. They introduce the so-called semigroup-adjusted realized covariation, which generalizes the multivariate realized covariation of BNS_2004_RCV. Their semigroup adjustment is a natural consequence of the functional setting, and must be accounted for to ensure that certain propagation effects disappear in the infill limit. A similar adjustment turns out to play a role in our long-span setting, and we discuss the relation to the high frequency case as well as classical realized covariation in detail.

To the best of our knowledge, the literature on volatility estimation in electricity markets is rather sparse, likely due to the complexity of the task. An approach for volatility estimation at the level of individual delivery periods is presented in Erdogu2016, where the price in each delivery period is treated as a separate and independent market. Based on our integrated variance estimator, we construct estimates of the weekly integrated variance in several European markets that, in contrast, explicitly and non-parametrically account for the strong dependence between individual prices and we study how the daily “term structures” of volatility behave via principal components analysis. We document that so-called propagation effects such as price mean-reversion make up more than $40\%$ of the total price variation and it is thus important to handle these effects in a robust manner. We show that our realized covariation estimator correlates highly and positively with the price level throughout several generation zones, but that the correlation with generation variables (such as wind or solar power generation) varies substantially with the generation zone. We also consider the leverage effect, which has occasionally been described as an inverse leverage effect in electricity markets, since high spot prices are associated to high volatility. We show that there is evidence of an unconditional weak form of inverse leverage, but that the asymmetric effect between increases and decreases in the price are largely explained by the price level itself and mean-reversion in volatility.

The rest of the paper is structured as follows. In Section (ref), we set up the model and construct the realized covariation estimator. Section (ref) considers the resulting volatility estimates from data on three generation zones and how they relate to several structural variables. Section (ref) contains more detailed analyses of some aspects of the realized covariation, Section (ref) studies the inverse leverage effect in detail, and Section (ref) concludes.

Model formulation

We shall model the daily cross section of latent prices $X_t$ as the mild solution to an SPDE of the form DaPratoZabczyk2014

equation[equation omitted — 83 chars of source]

where $W_t$ is a cylindrical Wiener process on $L^{2}(\mathbb{S}^1)$ generating a filtration $\mathcal{F}_t=\sigma(W_{t}:t\geq 0)$ and $\mathcal{A}$ is an unbounded operator on $\mathbb{S}^1$ generating a $C_0$-semigroup $\mathcal{S}$. The process $\mu_t$ represents a stochastic drift with state-space $L^2(\mathbb{S}^1)$ and we impose that this is $\mathcal{F}_t$-predictable and Bochner integrable. The process $\sigma_t$ represents the instantaneous volatility process and we impose that this is $\mathcal{F}_t$-predictable and takes values in the space of Hilbert-Schmidt operators on $L^2(\mathbb{S}^1)$. As we are interested in characterizing the second order structure of prices, we impose in all of the following that $\sup_{t\in [0,T]}\mathbb{E}\left[\lVert X_t \rVert^2_{L^2(\mathbb{S}^1)}\right]<\infty$, where $T>0$ is the maximum time horizon of interest. Ignoring the drift for simplicity, a sufficient condition is that

equation[equation omitted — 138 chars of source]

where $\lVert \cdot \rVert_{HS}$ denotes the Hilbert-Schmidt norm. The mild solution to the SPDE (ref) can be written in terms of the semigroup $\mathcal{S}$ as

equation[equation omitted — 155 chars of source]

If $\mathcal{S}$ admits an integral kernel $p(t-s,x,y)$ such that $(\mathcal{S}(t)f)(x)=\int p(t,x,y)f(y)dy$, then the price process $X_t$ has the martingale measure representation

equation[equation omitted — 222 chars of source]

where $W$ is a standard white noise on $\mathbb{R}_+\times \mathbb{S}^1$ of Walsh1986. Note that $W$ corresponds to a Gaussian Lévy basis on $\mathbb{R}_+\times \mathbb{S}^1$ with zero mean, unit variance and control measure equal to the canonical volume measure on $\mathbb{R}_+\times \mathbb{S}^1$. The martingale measure formulation (ref) therefore shows that the model is a special case of the time-stationary ambit field proposed in Kloster2026, with a particular choice of kernel function. In contrast to the ambit field framework, we have much less freedom to choose $p(t,x,y)$, but we note that $p$ can still exhibit temporal singularities, which have been argued for in electricity spot prices in Bennedsen2017.

The semigroup $\mathcal{S}$ is a model ingredient which should be specified. It may be known in certain cases, such as in the modeling of forward rates or electricity futures, where the shift semigroup arises naturally. In our case of electricity spot prices, it is not obvious what $\mathcal{S}$ should be and even a classical and tractable choice such as the heat semigroup is only known up to a scaling constant. In the following, we shall see that our measure of realized covariation is, in the best case, only identifiable up to smoothing by $\mathcal{S}$. This does not necessarily constitute a problem, but it changes the interpretation of the measured volatility proxy. In future applications, it will be interesting to study appropriate choices of $\mathcal{S}$ in more detail.

remarkThe model (ref) generalizes Ornstein-Uhlenbeck type models with stochastic volatility. Indeed, if the spatial dimension is collapsed to a single point, then we are in a standard one-dimensional setting and the only possible choice of $C_0$-semigroup is of the form $\mathcal{S}(t)x=e^{\lambda t}x$ with generator $\mathcal{A}x=\lambda x$ for some $\lambda\in\mathbb{R}$. Hence the model takes the familiar form \[ dX_t = \mu_t dt + \lambda X_tdt + \sigma_t dW_t, \] where $W_t$ is a standard Brownian motion.

Volatility estimation

Since electricity prices are determined and revealed at a daily frequency, this imposes a fixed upper bound on the sampling frequency. We can therefore not rely on high-frequency infill type results as developed in BenthSchroersVeraart2022,BenthSchroersVeraart2024, but are instead constrained to a long-span setting where we cannot expect the discretization error of the SPDE (ref) to vanish. In the following, whenever $A$ is a finite dimensional vector or matrix, we denote by $\lVert A \rVert$ the Euclidean norm on $\mathbb{R}^d$ or the Frobenius norm of $A$, and we denote by $\Sigma_t=\sigma_t\sigma_t^\ast$ the instantaneous variance operator.

Under the local average interpretation of prices in (ref), the individual daily observations are of the form $\widetilde{X}_t^{(i)}$, where \[ \widetilde{X}^{(i)}_t = \langle X_t,g_i\rangle_{L^2(\mathbb{S}^1)}, \quad g_{i}=\frac{1}{(h_{i}-h_{i-1})}\mathbf{1}_{[h_{i-1},h_i)}. \] Suppose that we have $0=t_{0}<t_{1}<\cdots t_{N}$ daily cross sectional observations and collect these in vectors $\widetilde{X}_{t_n}=(\widetilde{X}_{t_n}^{(1)},\ldots ,\widetilde{X}_{t_n}^{(d)})^\top$ and define \[ \Delta_{n}\widetilde{X} := \widetilde{X}_{t_{n}}-\widetilde{X}_{t_{n-1}}\in \mathbb{R}^d. \]

propositionDefine the observation operator $A:L^2(\mathbb{S}^1)\to \mathbb{R}^d$ by \[ (A x)_i := \langle x,g_i\rangle_{L^2(\mathbb{S}^1)}, \] such that the observed vector of local averages is $\widetilde{X}_t=AX_t$ with $X_t$ as in (ref). Then it holds that \begin{equation} \begin{aligned} \mathbb E\left[(\Delta_n\widetilde X)(\Delta_n\widetilde X)^\top\mid \mathcal{F}_{t_{n-1}}\right] &= \mathbb E\big[A(B_n+D_n)(B_n+D_n)^\ast A^\ast\mid \mathcal{F}_{t_{n-1}}\big] \\ &\quad \quad + A\mathbb E\left[\int_{t_{n-1}}^{t_{n}} \mathcal{S}(t_n-s)\Sigma_s \mathcal{S}(t_n-s)^\ast ds\mid\mathcal{F}_{t_{n-1}}\right]A^\ast, \end{aligned} \end{equation} where $B_n,D_n$ are, respectively, the propagation and the drift terms \[ B_n = (S(\delta)-I)X_{t_{n-1}}, \quad D_n = \int_{t_{n-1}}^{t_{n}}\mathcal{S}(t_{n} -s)\mu_sds. \] Equivalently, for $i,j\in\{1,\dots,d\}$, we have \begin{align*} \mathbb E\big[\Delta_n\widetilde X^{(i)}\Delta_n\widetilde X^{(j)}\big] &= \mathbb E\big[\langle B_n+D_n,g_i\rangle_{L^2(\mathbb{S}^1)}\,\langle B_n+D_n,g_j\rangle_{L^2(\mathbb{S}^1)}\big] \\ &\quad+ \mathbb E\left[\int_{t_{n-1}}^{t_n} \Big\langle \Sigma_s\,\mathcal{S}(t_n-s)^\ast g_j,\; \mathcal{S}(t_n-s)^\ast g_i\Big\rangle_{L^2(\mathbb{S}^1)}\,ds\right]. \end{align*}

Since observations are fixed at a daily frequency, we can set $\delta=(t_n-t_{n-1})$ and readily define the annualized realized covariation over some window of $w$ days as

equation[equation omitted — 176 chars of source]

When the period length $w$ is implicit, we shall simply refer to the estimator $\widehat{\Sigma}_j^w$ in (ref) as the RCV as shorthand. An application of Proposition (ref), the tower property, and a change of variables yields that

equation[equation omitted — 649 chars of source]

This has the intuitive interpretation that the annualized realized covariation over the next $w$ days is the average (conditionally) expected propagation and drift bias, plus the expected semigroup-weighted average variance over the period. As such, the realized covariation (ref) is a noisy proxy of the conditional average variance, in complete analogy to the finite dimensional setting, but with added bias coming from the semigroup. Supposing that all drift and propagation bias can be removed (i.e., $B_n=D_n=0$), we end up with an estimate of the form \[ \mathbb{E}\left[ \widehat{\Sigma}_j^w\mid \mathcal{F}_{t_{w(j-1)}} \right] = \frac{1}{\delta w}A \mathbb{E}\left[ \int_{0}^{\delta}\mathcal{S}(u)\left(\sum_{\ell=1}^{w}\Sigma_{w(t_j-1)+\ell-u}\right)\mathcal{S}(u)^\ast du \mid \mathcal{F}_{t_{w(j-1)}}\right]A^\ast. \] This estimate does not generally allow us to extract the raw expected average of $\Sigma_t$ due to the semigroup-weighting. We may, however, still think of the semigroup-weighted average variance as an “innovation term”, which is only related to the stochastic forcing $\sigma_tdW_t$ in (ref) and thus uniquely contains information on the shocks to prices. We note that in the high-frequency infill asymptotic setting of BenthSchroersVeraart2022,BenthSchroersVeraart2024, the so-called semigroup adjusted realized covariation is used to adjust the naive realized covariation estimator, such that the propagation bias $B_n$ disappears in the limit when the mesh between observations tends to zero.

In order to ensure convergence of the average sample realized covariation to a fixed limit, we require stationarity and ergodicity of the sampling sequence. For the series of local averages $(\Delta_n\widetilde{X})$, this is inherited from stationarity and ergodicity of the underlying price differences $(\Delta_n X)$ or the latent process $(X_t,\sigma_t)$ itself. Standard criteria for stationarity of semilinear SPDE's can be found in DaPratoGatarekZabczyk1992, which typically boil down to integrability conditions of the stochastic convolution appearing in (ref) and stability of the semigroup $\mathcal{S}$. Ergodicity typically requires stronger stability conditions on the semigroup $\mathcal{S}$, and some rather technical criteria on the SPDE dynamics are given in, e.g., HairerMattingly2011. A corresponding central limit theorem can be derived by imposing suitable mixing conditions on the SPDE, but we will not pursue this here.

assumptionThe sequence of differenced local averages $(\Delta_n \widetilde{X})_{n\in\mathbb{N}}$ is strictly stationary and ergodic.
propositionLet $\delta = t_{n}-t_{n-1}$ and let $w\in\mathbb{N}$ denote a window length and suppose that we can partition the observations $(\widetilde{X}_{t_n})_{n=1}^{N}$ into $N_w$ disjoint windows of length $w$. Suppose that the model (ref) satisfies Assumption (ref) and that $t_0=0$. Then \begin{equation} \frac{1}{N_w}\sum_{j=1}^{N_w}\widehat{\Sigma}_{j}^{w} \xrightarrow[N_w\to\infty]{a.s.} \mathbb{E}\left[ A(B_1+D_1)(B_1+D_1)^\ast A^\ast \right] + A\mathbb{E}\left[ \int_{0}^{\delta}\mathcal{S}(\delta -s)\Sigma_s\mathcal{S}(\delta -s)^\ast ds \right]A^\ast, \end{equation} where $\widehat{\Sigma}_j^w$ is defined by (ref) and $B_1,D_1$ in Proposition (ref).

The estimator (ref) and the long-span limit of Proposition (ref) are standard constructions, but obscured by extra terms, which are not present in the finite dimensional setting. In particular, the propagation terms $B_n$ and $B_1$ induced by the semigroup are normally not present. In Corollary (ref) we show that this contributes with more bias than we expect from the drift alone and in Proposition (ref) we show how this bias can be corrected when the semigroup is known. Proposition (ref) gives a feasible plug-in estimator of the one-step semigroup on the observation space, and hence a feasible correction of the propagation bias when the semigroup is unknown. In particular, we obtain a plugin estimator such that the adjusted RCV estimator $\widehat{\Sigma}_{j,\mathrm{pl}}^{w}$ estimates the conditional second moment of the semigroup-adjusted increment of $\widetilde{X}_{t_j}$, which in general does not eliminate the propagation bias entirely, since we must first estimate the semigroup.

propositionLet the setting be as in Proposition (ref) and suppose that the drift has bounded second moment, i.e., $\sup_{s\geq 0}\mathbb{E}\left[ \lVert \mu_s \rVert^2_{L^2(\mathbb{S}^1)}\right]<C_\mu<\infty$. Then \[ \frac{1}{N_w}\sum_{j=1}^{N_w}\widehat{\Sigma}_{j}^{w} - \frac{1}{\delta}\mathbb{E}\left[ A B_1 B_1^\ast A^\ast \right] \xrightarrow[N_w\to\infty]{a.s.} A\mathbb{E}\left[ \int_{0}^{\delta}\mathcal{S}(\delta-s)\Sigma_s\mathcal{S}(\delta-s)^\ast ds \right]A^\ast + O(\delta^2), \]
propositionFor notational simplicity, write $\widetilde X_n:=\widetilde X_{t_n}$. Suppose that the sequence $(\widetilde X_n)_{n\in\mathbb{N}}$ is strictly stationary, ergodic, and satisfies \[ \mathbb E[\widetilde X_0]=0, \qquad \mathbb E[\lVert\widetilde X_0\rVert^2]<\infty. \] Set $\Gamma:=\mathbb E[\widetilde X_0\widetilde X_0^\top]$, and suppose that $\Gamma$ is positive definite. Define the effective one-step semigroup on the observation space by \[ \mathcal S_\delta := \mathbb E[\widetilde X_1\widetilde X_0^\top]\Gamma^{-1}\in\mathbb R^{d\times d}, \] and its sample estimator by \begin{equation} \widehat{\mathcal{S}}_{\delta,N} :=\Big(\sum_{n=1}^N \widetilde{X}_n\widetilde{X}_{n-1}^\top\Big) \Big(\sum_{n=1}^N \widetilde{X}_{n-1}\widetilde{X}_{n-1}^\top\Big)^{-1}. \end{equation} Then it holds that $\widehat{\mathcal{S}}_{\delta,N}\xrightarrow[N\to\infty]{a.s.}\mathcal{S}_\delta$. Define the plug-in residuals by \[ \widehat{\varepsilon}_n^{\mathrm{pl}}:= \widetilde{X}_n-\widehat{\mathcal{S}}_{\delta,N}\widetilde{X}_{n-1} = \Delta_n\widetilde{X}-(\widehat{\mathcal{S}}_{\delta,N}-I_d)\widetilde{X}_{n-1}, \] and the residual-based feasible realized covariation by \begin{equation} \widehat{\Sigma}_{j,\mathrm{pl}}^{w} :=\frac{1}{\delta w}\sum_{\ell=1}^w\widehat{\varepsilon}_{(j-1)w+\ell}^{\mathrm{pl}} \big(\widehat{\varepsilon}_{(j-1)w+\ell}^{\mathrm{pl}}\big)^\top. \end{equation} Then each $\widehat\Sigma_{j,\mathrm{pl}}^{\,w}$ is positive semidefinite. Moreover, setting $N=wN_w$, it holds that \[ \frac1{N_w}\sum_{j=1}^{N_w}\widehat\Sigma_{j,\mathrm{pl}}^{w} \xrightarrow[N_w\to\infty]{a.s.} \frac{1}{\delta}\, \mathbb{E}\Big[(\widetilde X_1-\mathcal S_\delta \widetilde X_0) (\widetilde{X}_1-\mathcal{S}_\delta \widetilde{X}_0)^\top \Big], \] Equivalently, \begin{equation} \frac1{N_w}\sum_{j=1}^{N_w}\widehat{\Sigma}_{j,\mathrm{pl}}^{w} \xrightarrow[N_w\to\infty]{a.s.} \frac{1}{\delta}\mathbb{E}\left[(\Delta_1\widetilde X)(\Delta_1\widetilde X)^\top\right] - \frac{1}{\delta}\mathbb E\Big[ (\mathcal S_\delta-I_d)\widetilde X_0\widetilde X_0^\top(\mathcal S_\delta-I_d)^\top \Big]. \end{equation}
corollaryLet the setting be as in Proposition (ref) and assume in addition that there exists a matrix $S_{\delta,A}\in\mathbb R^{d\times d}$ such that \begin{equation} A\mathcal{S}_\delta=S_{\delta,A}A, \end{equation} and that either $\mu\equiv 0$ or the data have been de-drifted so that \[ \mathbb E[\widetilde{X}_n\mid \mathcal F_{t_{n-1}}]=S_{\delta,A}\widetilde{X}_{n-1}. \] Then $\mathcal{S}_\delta=S_{\delta,A}$, and therefore \[ \frac{1}{N_w}\sum_{j=1}^{N_w}\widehat\Sigma_{j,\mathrm{pl}}^{\,w} \xrightarrow[N_w\to\infty]{a.s.} A\mathbb{E}\Big[\int_0^\delta \mathcal{S}(\delta-s)\Sigma_s\mathcal S(\delta-s)^\ast\,ds\Big]A^\ast. \]
remarkWithout the closure condition (ref), the sample estimator (ref) still converges to a matrix representation on the observation space. Indeed, the population limit \[ \mathcal{S}_{\delta}:= \mathbb{E}[\widetilde{X}_{t_1}\widetilde{X}_0^\top]\Gamma^{-1}, \qquad \Gamma:=\mathbb{E}[\widetilde{X}_0\widetilde{X}_0^\top], \] is the unique solution to the least-squares problem \[ \mathcal{S}_{\delta}= \arg\min_{M\in\mathbb R^{d\times d}} \mathbb E\left[\|\widetilde{X}_{t_1}-M\widetilde{X}_0\|^2\right]. \] In other words, $\mathcal{S}_{\delta}\widetilde{X}_t$ is the best linear one-step predictor of the next vector of local averages, $\widetilde{X}_{t+\delta}$, from the current vector of local averages, $\widetilde X_t$. The sample estimator $\widehat{\mathcal{S}}_{\delta,N}$ is a feasible estimator of this population predictor.
corollaryAssume that the drift has bounded second moment and that the semigroup $\mathcal S$ is analytic with $X_{t_0}\in\mathcal D((-\mathcal A)^\alpha)$ for some $\alpha\in(0,1]$. Then \[ \frac{\delta}{N_w}\sum_{k=1}^{N_w}\widehat{\Sigma}^{w}_k\xrightarrow[N_w\to\infty]{a.s.} A\,\mathbb{E}\left[\int_{0}^{\delta} \mathcal S(\delta-s)\,\Sigma_s\,\mathcal S(\delta-s)^\ast\,ds\right]A^\ast \;+\;O(\delta^{2\alpha}). \]
remarkIf, in addition to the conditions of Corollary (ref), $\mathcal{S}(u)$ is close to the identity on $[0,\delta]$ in the relevant sense (e.g. $\lVert\mathcal{S}(u)-I\rVert_{\mathrm{op}}\lesssim u$ and $\Sigma_s$ is sufficiently regular), then the leading term in Corollary (ref) above can be approximated by $A\,\mathbb{E}[\int_{0}^{\delta}\Sigma_s\,ds]\,A^\ast$ up to an additional $O(\delta)$ error.
remarkThe condition $X_t\in\mathcal{D}((-\mathcal{A})^{\alpha})$ required in Corollary (ref) is a spatial regularity condition. Applying $(-\mathcal{A})^\alpha$ to the mild solution (ref) we see that $X_t\in\mathcal{D}((-\mathcal{A})^\alpha)$ if it holds that \begin{equation} \mathbb{E}\left[ \int_{0}^{T}\lVert (-\mathcal{A})^\alpha \mathcal{S}(t-s)\sigma_s \rVert^2_{HS} ds \right] < \infty. \end{equation}

Relation to classical realized covariation

In the one-dimensional setting of Remark (ref), the corresponding result to that of Proposition (ref) is that

equation[equation omitted — 309 chars of source]

where the right-hand-side is exactly a noisy proxy of the annualized integrated variance BNS_2004_RCV. In this case, the semigroup weighting is explicitly known, and a small-time expansion of the integrand on the right hand side of (ref) can, for small $\delta$, recover $\frac{1}{\delta}\int_{0}^{\delta}\mathbb{E}\left[\sigma_t^2\right]dt$ up to an additional $O(\delta)$ term. A severe complication of the genuinely infinite dimensional setting is that the semigroup is much less explicitly known. A simplification is possible if the mild solution (ref) is a strong Itô semimartingale in $L^2(\mathbb{S}^1)$, admitting the strong solution

equation[equation omitted — 132 chars of source]

which does not directly involve the semigroup $\mathcal{S}$. In this case, we have $X_t\in\mathcal{D}(\mathcal{A})$ and we obtain that the limit in Proposition (ref) is of the form \[ \frac{1}{N_w}\sum_{j=1}^{N_w}\widehat{\Sigma}_j^w - \frac{1}{\delta}A\mathbb{E}\left[ \left(\int_{0}^{\delta}\mathcal{A}X_sds\right)\left(\int_{0}^{\delta}\mathcal{A}X_sds\right)^\ast\right]A^\ast\xrightarrow[N_w\to\infty]{a.s.}\frac{1}{\delta}A\mathbb{E}\left[\int_{0}^{\delta}\sigma_t\sigma^\ast_tdt\right]A^\ast + O(\delta), \] in analogue to the classical case of realized covariation and in accordance with Corollary (ref). This situation is similar to that of BenthSchroersVeraart2022,BenthSchroersVeraart2024, who derive quite explicit conditions to characterize when their semigroup adjustment is needed and when the “naive” realized covariation is appropriate. If the Hilbert space in which $X_t$ takes values is finite dimensional, then $X_t$ is always a semimartingale and our estimator collapses to the usual realized covariation of the observed process $\widetilde{X}_t=AX_t$. To this end, note that without the semigroup-weighting, the innovation term in (ref) can be written entrywise as \[ \left( A\mathbb{E}\left[ \int_{0}^{\delta}\Sigma_sds\right]A^\ast\right)_{i,j} = \int_{0}^{\delta}\mathbb{E}\left[ \langle \Sigma_s g_j,g_i\rangle_{L^2(\mathbb{S}^1)} \right]ds. \] If the operator $\Sigma_s$ itself admits a kernel, $c$, then we have \[ \left( A\mathbb{E}\left[ \int_{0}^{\delta}\Sigma_sds\right]A^\ast\right)_{i,j} = \int_{0}^{\delta}\mathbb{E}\left[ \frac{1}{\lvert I_i\rvert \lvert I_j\rvert }\int_{I_i}\int_{I_j}c_s(h,u) du dh \right] ds, \] where $I_i,I_j$ represent the bin intervals. Hence we have, in the semimartingale case (ref), the concrete interpretation as the time integral of the average covariance kernel over spatial bins $i$ and $j$.

Realized variance of the average daily price

The object of study is often the average daily spot price, corresponding to modeling $\overline{P}_t=\frac{1}{2\pi}\int_{0}^{2\pi}X_t(h)dh$. This object arises as a natural univariate representation of “the price of electricity” and it forms the basis of the underlying in conventional futures contracts. If $X_t$ is a strong semimartingale solution to (ref), then the process $\overline{P}_t$ has quadratic variation (ignoring propagation and drift terms for simplicity)

equation[equation omitted — 235 chars of source]

where $\overline{g}(h)=\frac{1}{2\pi}\mathbf{1}_{[0,2\pi)}(h)$. In particular, it holds that the $w$-period realized variance $\widehat{RV}_j^w$ of $\overline{P}$ is given by (even with propagation and drift terms included)

equation[equation omitted — 196 chars of source]

where $\mathbf{1}=(1,\ldots ,1)^\top \in \mathbb{R}^d$. In practice, this means we may easily obtain the realized variance of the average daily price from our RCV estimator. This highlights one of the strengths of the continuous local average interpretation (ref), since the realized variance of the average price and the realized covariation of the panel of prices are connected without imposing any restrictions on the sampling scheme in time or space. For example, the relation (ref) is easily adapted to the case of spatial averaging over intervals of different lengths, by incorporating non-equal weights.

remarkNotice that if $X$ is only a mild solution to (ref) of the form (ref), then the semigroup smoothing disappears from the quadratic variation (ref) whenever it holds that $\mathcal{S}(t)^\ast \overline{g} = \overline{g}$ for all $t\geq 0$. This is often the case, such as for diffusion type semigroups on $\mathbb{S}^1$.

Realized covariation of electricity spot prices

In this Section, we apply the RCV estimator (ref) and the propagation adjusted version (ref) to construct proxies of the weekly integrated variance in three European electricity markets. We consider three different generation zones that are characterized by different underlying characteristics: Germany, Norway (NO2), and Spain. Germany is the largest market for electricity and related products in Europe, and owing to their central geographical placement in Europe, their electricity production is affected by many adjacent generation zones via, e.g., interconnectors. Norway is characterized by having a large share of hydropower generation, which is expected to lead to more stable prices, since they can regulate their production according to price signals. Spain is characterized by a large amount of solar energy, which may cause volatile price behaviour due to the nature of renewable generation. To save space throughout the main body of the paper, we shall mainly depict the results for Germany, with the corresponding results for Norway and Spain provided in Appendix (ref).

In the notation of Section (ref), we choose windows of $w=7$ days, such that the RCV is a noisy estimate of the annualized (i.e. $\delta=\tfrac{1}{365}$) weekly (semigroup weighted) integrated variance. The choice of one week is made to eliminate subtle weekly periodicities such as weekend effects. We consider data from October 1st 2018 to September 30th 2025. Prior to this period, the German zone was coupled with Austria and hence represents a structurally different market, and after this period most European zones converted from 24 daily delivery periods (hourly) to 96 daily delivery periods (15 minutes). Due to the local average interpretation (ref), it is in principle possible to still compute the hourly local averages from the quarterly prices, but we refrain from doing this due to potential market reactions to the transition. This leaves us with $N=2551$ daily observations per zone and a total of 61368 data points per generation zone. Aggregation into disjoint weekly bins leaves us with 365 weeks of data. To obtain more data and, by extension, more stable results in the following, we therefore consider a 7-day rolling RCV estimate according to (ref) or (ref), such that we have 2557 RCV estimates per zone. This induces a degree of temporal smoothing, but the interpretation as a noisy proxy for the 1-week integrated variance remains, and the result of Proposition (ref) is unchanged for the rolling sequence of price differences. Additionally, since each 7-day interval contains exactly one of each weekday, the issue of deterministic weekly periodicities is still avoided. On daylight-savings transition days, we impute the data as follows. If the affected day has 23 hours, we let the missing hour have a price given by the average of the adjacent hours. If the affected day has 25 hours, we have two “overlapping hours”; the prices in these two hours are merged into one price as their average.

Before considering the estimates obtained from data, we first deal with the terms $D_n,B_n$ of (ref). In Section (ref), we argue that the sequence of differences $(\Delta_n\widetilde{X})_{n\in\mathbb{N}}$ obtained from the data are consistent with the assumption of stationarity and do not exhibit any unit root type effects. Hence we can reasonably assume that $D_n\approx 0$ and that any non-stationarities are slowly moving trends. In Section (ref), we apply the estimator $\widehat{\Sigma}_{j,\mathrm{pl}}^w$ to eliminate propagation effects and show that this differs considerably from the raw estimator $\widehat{\Sigma}_{j}^w$. It follows that the propagation terms $B_n$ contribute with substantial variation in prices and should not be ignored. Furthermore, we consider the estimated matrix representation of the semigroup $\mathcal{S}(\delta)$, which reveals an intricate structure.

figure[figure omitted — 1,448 chars of source]

Handling non-stationarities

It is well-known that electricity price levels tend to exhibit non-stationary behaviour in the form of predictable trends and seasonal periodicities. Apart from potentially contradicting Assumption (ref), this also leads to larger bias and contamination in the estimates of weekly realized covariation. The procedure of removing non-stationarities in electricity prices is often associated to quite intricate filtering procedures, with an overview of the one-dimensional case given in, e.g., JanczuraJoanna2013Isas. We are, however, primarily interested in the sequence of differences, $(\Delta\widetilde{X})_{n=1}^{N}$ and so, if the non-stationarities in price levels can be interpreted as slow trends or unit roots, then the sequence of price differences should be suitably stationary.

We carry out several stationarity assessments for the 24-dimensional series of observed local averages $(P_n)_{n=1}^{N}$ in each country. As the primary test, we consider the multivariate KPSS type test for stationarity proposed in NyblomHarvey2000. We supply this with the standard univariate KPSS test for stationarity on the individual price series, as well as the augmented Dickey-Fuller (ADF) test for unit roots on the individual price series. The test statistics and their critical values are tabulated in Table (ref) in Appendix (ref). It is apparent from the KPSS tests that we reject stationarity in price levels across all countries at any reasonable level of significance, both jointly and in the individual price series. Interestingly, however, the ADF test rejects the presence of a unit root in most of the individual price series, which fits well with the notion that electricity spot prices are mean-reverting and hence shocks are not persistent. In Table (ref) in Appendix (ref), we state the corresponding test statistics for the differenced price series, $(\Delta\widetilde{X}_n)_{n=1}^{N}$, which unanimously favours the assumption of stationarity at any reasonable level of significance. This suggests that the non-stationarities in price levels are slowly evolving trends that can largely be differenced out. Interestingly, any periodic effects seem to also be eliminated by differencing, at least to the extent at which the applied tests can detect them. To this end, we note that the KPSS alternative can be quite weak to some forms of non-stationarity in the covariance structure. Overall, we treat the price differences $(\Delta\widetilde{X}_n)_{n=1}^{N}$ as being largely trend-stationary, but noting that there may still be some non-stationary effects which are hard to pick up. This would potentially invalidate the long run estimator of Proposition (ref), but does not affect the conditional estimates of the RCV estimator (ref). In particular, we note that the differencing means that the drift contribution is small, such that $D_n\approx 0$. Hence we expect that most of the bias arises from the propagation term, $B_n$, which can then be accounted for via Proposition (ref).

Filtering of the propagation term

In addition to non-stationary predictable effects, it is well-documented that electricity prices exhibit mean-reversion, in the sense that price shocks do not persist indefinitely and tend to gradually diminish over time. In HuismanHuurmanMahieu2007, it is found that prices in separate hours tend to mean-revert at different speeds. In the model (ref), this is captured by the term $\mathcal{A}X_tdt$ and leads to the additional propagation terms $\mathbb{E}[A B_nB_n^\ast A^\ast]$ in the estimator (ref). As we are mainly interested in the second order structure of the innovations/shocks to prices, we seek to filter out the deterministic propagation part by means of the plugin estimator $\widehat{\Sigma}^w_{j,\mathrm{pl}}$ of Proposition (ref). Given the results of Section (ref), we assume that the drift admits the decomposition \[ \mu_t = b_t+\widetilde{\mu}_t, \] where $b_t$ is a predictable process such that \[ \frac{1}{N}\sum_{n=1}^{N}\lVert m_n-m_{n-1}\rVert^2 \xrightarrow[N\to\infty]{}0, \quad m_n = A\int_{0}^{t_n}\mathcal{S}(t_n-s)b_sds, \] and the joint process $(\widetilde{\mu},\sigma)$ is strictly stationary. We can then estimate $\widehat{m}_n=\widehat{a}_n$, where

equation[equation omitted — 175 chars of source]

where $K$ is a kernel and $\tau$ the bandwidth. For our applications, we choose an Epanechnikov kernel with a bandwidth of 90 days, corresponding to $\tau = 0.246$; in order to eliminate weekly seasonality, we also include day-of-week dummies in the regression (ref). The process $(\widetilde{X}_{t_n}-\widehat{m}_n)_{n\in\mathbb{N}}$ is then strictly stationary and with mean zero, such that Proposition (ref) applies and the resulting estimator $\widehat{\Sigma}^w_{j,\mathrm{pl}}$ removes some of the predictable propagation effect.

The semigroup estimator $\widehat{\mathcal{S}}_\delta$ in (ref) is, according to Remark (ref), such that $\widehat{\mathcal{S}}_{\delta}\widetilde{X}_t$ is the best linear one-step predictor of the de-meaned vector of prices. The $(i,j)$'th entry of $\widehat{\mathcal S}_\delta$ can thus be interpreted as the contribution of today's hour $j$ in predicting tomorrow's hour $i$. Hence if the hourly price series $P_t^{(h)}$ are independent across $h$, the off-diagonal elements of $\widehat{\mathcal S}_\delta$ vanish, and the $i$'th diagonal entry equals exactly the AR(1) coefficient of the time series $P_t^{(i)}$ (this intuition breaks down whenever the off-diagonal elements are non-zero, but it is worth keeping in mind).

On Figure (ref), we depict the estimated $\widehat{\mathcal{S}}_\delta$ from the full data from German generation zone along with the corresponding eigenvalue spectrum. The estimated off-diagonal matrix is clearly non-zero and quite heterogeneous, supporting the findings of cross-sectionally varying mean-reversion rates in HuismanHuurmanMahieu2007. In particular, we note that column 24 is positive and of larger magnitude than the other hours, with a slight decay from the evening to morning hours. This is evidence of price cyclicality as also documented in Kloster2026, where the last price of today contributes with significant predictive power for the first few hours of tomorrow. The largest eigenvalue is $\lambda_1\approx0.8$, corresponding to common shocks across all hours to have a half-life of roughly $-\log(2)/\log(0.8)\approx 3$ days. The corresponding plots for Norway and Spain are depicted in Figures (ref) and (ref) in Appendix (ref) and we note that the overall structure is strikingly similar although the common shock half-life is ca. 6 days in Norway. In practice, we cannot rely on the semigroup matrix $\widehat{\mathcal{S}}_\delta$ estimated from the full dataset, as this creates look-ahead bias. Whenever it is relevant to maintain causality in the analyses throughout the paper, we re-estimate $\widehat{\mathcal{S}}_\delta$ based on the available backward-looking data every four weeks (28 days) starting with a 52 week burn-in.

figure[figure omitted — 671 chars of source]

Preliminary findings

In order to assess the magnitude of the propagation terms $B_j$ in the estimator $\widehat{\Sigma}^w_j$, we construct the propagation share, denoted $PS$, designed to measure how much variation in the naive RCV estimator $\widehat{\Sigma}_j^w$ can be explained by the deterministic propagation via $\mathcal{A}$. Supposing that the sequence $(\widetilde{X}_{t_n})_{n}$ has been de-meaned via (ref), it follows from Proposition (ref) that $\Delta_n\widetilde{X}$ decomposes as $\Delta_n\widetilde{X} = B_n+M_n$, where $B_n$ and $M_n$ are orthogonal. Hence \[ \sum_n \lVert \Delta_n\widetilde{X}\rVert^2 = \sum_{n}\lVert B_n\rVert^2 + \sum_{n}\lVert M_n\rVert^2, \] and we can define $PS$ as the following coefficient of determination \[ PS = \frac{\sum_{n}\lVert \widehat{B}_n\rVert^2}{\sum_n \lVert \Delta_n\widetilde{X}\rVert^2}, \] where $\widehat{B}_n=(\widehat{\mathcal{S}}_\delta - I)\widetilde{X}_{t_{n-1}}$. Across the three zones, the propagation share is $PS\approx 0.4$, indicating that approximately $40\%$ of the observed variation across all hours can be attributed to mean-reversion. Interestingly, when inspecting the propagation shares of individual hours $PS_h$, corresponding to the ratio of column-sums of squared components \[ PS_h = \frac{\sum_n (\widehat{B}_n^{(h)})^2}{\sum_n (\Delta_n\widetilde X^{(h)})^2}, \qquad h = 1, \ldots, d, \] we find that there is a clear decreasing proportion as $h$ increases. This means that, in all three zones under consideration, the late hours tend to exhibit a larger share of variation coming from raw noise than the early hours. In other words, the early hours tend to be relatively more predictable from the current price level itself. We depict the propagation share across hours on Figure (ref) for Germany, and on for Norway and Spain on Figures (ref) and (ref) in Appendix (ref). An interesting question arising from the observed monotone decay in $h\mapsto PS_h$ is whether this can be attributed to the temporal distance between the time at which prices are determined and the time of delivery. Indeed, when prices are determined, the price $P^{(d)}_t$ is valid many hours later than $P_t^{(1)}$ and there may thus be more uncertainty about the underlying supply and demand further in the future, leading to less predictable prices. We further investigate this effect in Section (ref)

On Figure (ref), we depict the unconditional estimates of the rolling weekly RCV in each market, constructed according to Proposition (ref). We include both the estimates obtained from $\widehat{\Sigma}_j^w$ and $\widehat{\Sigma}^w_{j,\mathrm{pl}}$. It is apparent that it is very important to account for the deterministic mean-reversion, as this removes a substantial amount of variation in prices. On Figure (ref), we depict the corresponding standardized “realized correlations”, computed as

equation[equation omitted — 98 chars of source]

where $\widehat{\Sigma}^w=\frac{1}{N_w}\sum_{j=1}^{N_w}\widehat{\Sigma}_j^w$ and $\widehat{Q} = \mathrm{diag}(\widehat{\Sigma}^w)^{1/2}$, and similarly for the propagation adjusted version. This realized correlation serves as a normalized measure of the estimated RCV, to better visually gauge patterns. Also here, it is apparent that the propagation adjustment leads to substantially different results. Most strikingly, we observe that prior to the adjustment there is a cluster of negative values, whereas all entries across all zones are positive after the adjustment. The correlations that appear negative at first sight are between early and late hours; in Section (ref), we show that this apparent negative correlation arises from the larger propagation share of early hours where most variation arise from the deterministic propagation which tends to drag early hour prices in the same direction, regardless of the shocks to other hours.

figure[figure omitted — 281 chars of source]

As an illustration of the conditional integrated variance, we depict the diagonal $\mathrm{diag}(\widehat{\Sigma}_{j,\mathrm{pl}}^w)$ over time, corresponding to the vector of integrated variances across all hours. Figure (ref) shows the results for the German zone, while Figures (ref) and (ref) in Appendix (ref) show the results of the Norwegian and Spanish zone. To maintain a reasonable scale, we plot the logarithm of the estimated conditional integrated variances, which in particular keeps the extreme spikes during 2022 at reasonable levels. It is clear that the overall volatility varies wildly over time with apparent clustering effects. Furthermore, we note that periods of high volatility seems to often coincide with periods of high price levels. We shall return to this issue in Section (ref).

figure[figure omitted — 1,363 chars of source]

Decomposing the realized covariation

The raw estimate of the realized covariation matrix $\widehat{\Sigma}^w_{\mathrm{pl}}=\frac{1}{N_w}\sum_{j=1}^{N_w}\widehat{\Sigma}_{j,\mathrm{pl}}^w\in\mathbb{R}^{d\times d}$ reveals an interesting pattern. In order to assess the structure in more detail, we carry out a principal component decomposition of the form \[ \widehat{\Sigma}_{\mathrm{pl}}^w=Q\Lambda Q^\top, \] where $Q^\top Q=Q Q^{\top}=I_d$, $\Lambda = \mathrm{diag}(\lambda_1,\ldots ,\lambda_d)$ with eigenvalues $\lambda_1\geq \cdots\geq \lambda_d\geq 0$, and $\widehat{\Sigma}_{\mathrm{pl}}^wq_k=\lambda_k q_k$. Equivalently, \[ \widehat{\Sigma}^w_{\mathrm{pl}} = \sum_{k=1}^{d}\lambda_kq_kq_k^\top, \] where $q_k\in\mathbb{R}^d$ are the principal directions and $\lambda_k$ the variance explained along that direction. The factor loadings are then defined as $\sqrt{\lambda_k}q_k$ for $k=1,\ldots ,d$, ranked by the proportion of variance that they explain. The factor scores are then the coordinates of each observation, $x_t$, in the factor directions such that the $k$'th factor score is $S_{k}(t)=q_k^\top x_t$. For all three generation zones, we find that between six and eight principal components are able to reproduce $95\%$ of the variance, indicating that a few common factors drive most of the volatility. We note that this amount of factors is relatively high compared to more traditional areas of finance, such as yield curve modeling, where three factors often suffice. This is not too surprising, since yield curves evolve much less erratically than electricity prices.

The factor loadings of the estimated RCV in the German generation zone are depicted on Figure (ref) for the first four factor loadings. These reveal an interesting structure which is remarkably similar across the three zones, where the corresponding plots for Norway and Spain are provided on Figures (ref) and (ref) in Appendix (ref). The first factor explains more than $60\%$ of the total variation in $\widehat{\Sigma}_{\mathrm{pl}}^w$ in all three zones, and we note that this can be interpreted as a “level factor”, that describes the overall shape of the weekly RCV. The interpretation of a level factor is slightly different from many conventional settings, since $q_1$ is not flat across hours. Indeed, $q_1$ shows a distinct hump shape that rises from the early hours, peaks around 18-19 o'clock, and slightly decreases towards the late hours. This indicates that the mid-day and late hours generally exhibit higher levels of volatility. All of the entries of $q_1$ do, however, share the same sign, meaning that the first factor determines co-movements of the volatility across all hours.

table[table omitted — 2,570 chars of source]

The remaining factors all have sign changes and therefore induce more sophisticated co-movement patterns. For example, we see that the third principal component in the German zone loads primarily onto the early hours, thus describing a substantial part of their variation. In contrast to the level factor, the third principal component induces negative correlation between the early and later hours. Interestingly, the second principal component seems to exhibit a shape close to the so-called duck curve, which refers to the conventional shape of electricity demand, less the demand supplied by renewable resources. These patterns are similar in the Norwegian zone, whereas in the Spanish zone the role of the second and third principal components are seemingly swapped. These slight discrepancies could likely arise from structural and market specific features such as price coupling and generation variables, and a more detailed analysis would be interesting to pursue in the future.

figure[figure omitted — 390 chars of source]

Factor loading stability

The principal components with loadings depicted on Figure (ref) are extracted based on the full sample. It is therefore not guaranteed that any interpretations based on these loadings are valid at all points in time. On Figure (ref), we plot the weekly estimated factor loadings over time from the German generation zone and corresponding to the first four principal components (aligned to have same sign throughout). The same figure is reproduced for Norway and Spain on Figures (ref) and (ref) in Appendix (ref). Although there is some -- seemingly seasonal -- variation in the factor loadings, their overall shapes are very consistent over time, especially for the first three factors. This suggests that the interpretations of the loadings are stable over time.

The level factor

Since the level factor explains an excess of 60% of the variation in the weekly RCV across all three zones under consideration, we consider this as the main driver of volatility. In order to assess what drives the level factor, and by extension volatility, we consider the factor score, $S_1(t)$, and regress $\log (S_1(t))$ on various variables relating to expected electricity demand and production, as well as the overall price level. The log-transform is chosen to suppress outliers, as the score has several large spikes; the log-transform is valid as the score is always positive throughout our datasets. The data is described in Appendix (ref). Each zone has day-ahead forecasted wind and solar generation (except NO2 which has no solar generation) and the actual realized production mix. We include rolling weekly averages of all these, as well as the PC1 factor scores of wind and solar forecasts in the regression, where we note that the first principal component of wind and solar forecasts explain 86% and 95% of their respective variance. We then control for price-scaling effects by including the rolling weekly average price $\overline{P}$ as well as $\overline{P}^2$. We also include the average forecast errors on wind and solar generation. We note that the regressions do not contain any information on what predicts volatility, but merely illustrates contemporaneous relationships between volatility and various generation and price variables. The results are depicted for the German generation zone in Table (ref), while similar tables are depicted for Norway and Spain in Appendix (ref). We do not add seasonal dummies, as these are largely absorbed by the price level controls. We also note that the coefficients are not comparable across zones, as the amount of electricity generated by the same type of fuel varies greatly.

figure[figure omitted — 907 chars of source]

The regression results reveal that the level of weekly RCV is very strongly tied to the price level itself in all zones. We refer to this as price/level-scaling and return to this issue when investigating leverage effects in Section (ref). Both $\overline{P}$ and $\overline{P}^2$ are highly significant in all zones, and combined explain 54%, 52%, and 33% of the variation in $\log(S_1(t))$ in Germany, Norway, and Spain, respectively. We note that in the full regression, almost all variables relating to renewable production are statistically insignificant, even the forecast errors. A notable exception is that of hydro generation, which is highly significant and adds a large amount of explainability in the Norwegian zone in particular. Overall, we achieve very high $R^2$ statistics in all zones, particularly Germany and Spain which both have $R^2>80\%$. By investigating the sign of the regression coefficients, we see that high forecasts on renewable energy sources are associated with higher price volatility, but that high realized renewable production contributes with different effects to the price volatility. For example, high realized wind production generally has a negative effect on the level of volatility. The remaining renewable sources such as hydro, biomass, and geothermal energy have varying effects and significance across the zones, which reflect their relative importance in the specific markets. Conventional fuel sources such as oil and gas contribute with opposite signs; oil generation is generally associated with lower volatility while gas generation is generally associated with higher volatility.

We conclude that, overall, it seems electricity price volatility is highly and positively correlated with the spot price level itself and that the realized production mix is important in explaining volatility dynamics, but the relative importance and contributions of the production mix components are highly individual to the market under consideration.

Investigating the propagation share

In this Section, we briefly turn to the propagation share across hours, which exhibits a decreasing pattern. This does not inherently mean that early hours are less volatile in absolute terms and we therefore depict the sum $(\widehat{\Sigma}^w_{\mathrm{pl}})^{(h,h)}+\frac{1}{N}\sum_{n}\widehat{B}_n^{(h)}$ on Figure (ref). This shows both the relative and absolute share of volatility versus propagation/mean-reversion and clearly indicates that the higher propagation share of early hours coincides with lower volatility levels in the German generation zone. Similar plots for Norway and Spain are provided on Figures (ref) and (ref) in Appendix (ref). As hinted earlier, a possible explanation for the decreasing propagation share and predictability of volatility over the hours of the day could be the temporal distance from the time of price settlement to the time of delivery, which inherently induces more uncertainty.

figure[figure omitted — 830 chars of source]

To assess whether this is the case, we run the following regression \[ PS_h = \alpha +\beta_0h + \beta_1e_h^{\mathrm{wind}}+\beta_2e_h^{\mathrm{solar}}, \] where $e_h^{\mathrm{wind}}$ is the mean squared error of the standardized (zero mean, unit variance) wind production forecast versus actual wind production in hour $h$ across the whole dataset, and similar for solar. It turns out that the wind forecast error is almost perfectly collinear with the hour, and as such this adds little value in the regression. The solar forecast errors are concentrated around the peak solar generation hours (mid-day) and as such adds value beyond just the hourly ordinals. Since the Norwegian zone has no solar generation, we only consider Germany and Spain with the regression results tabulated in Table (ref). The hour ordinal alone explains 96% of the total variation while the marginal contribution of solar production after controlling for hour is robust and identifies a separable role for forecast-error uncertainty. In both the German and Spanish zones, a one standard deviation increase in the per-hour solar MSE is associated to an decrease of approximately 4 percentage points in the propagation share. While the simple regression is not conclusive evidence, we conjecture that forecast uncertainty and time-to-delivery effects are important in explaining the increasing volatility and unpredictability of prices over the day.

table[table omitted — 1,417 chars of source]

The inverse leverage effect

The leverage effect typically refers to the negative correlation between an asset return and its volatility. In particular, this means that there is an asymmetric association between increases and decreases in volatility. It has been suggested in the literature that electricity prices in some markets seem to exhibit an inverse leverage effect, which is to say that price shocks are positively correlated with shocks to volatility InverseLeverage,JanczuraWeron2010. To the best of our knowledge, most findings of inverse leverage seem to be from estimation or calibration of a specific model and has not been studied via rigorous and non-parametric volatility estimates, and we therefore consider this effect in some detail. We distinguish between two types of leverage effect, which we refer to as weak and strong form, respectively:

itemize• Weak form: Changes in integrated variance responds asymmetrically to price changes. • Strong form: The asymmetry is not explained by price-level scaling or mean-reversion in volatility.

The strong form of leverage is essentially a conditional version of the weak form. Indeed, given the findings in Section (ref), we expect high absolute price levels to lead to large deviations around these levels, which we refer to as price-level scaling of the volatility. Similarly, when the price level is large, we expect that it is more likely to shrink fast towards the long-run mean. The strong form leverage imposes that we observe a correlation between volatility and price changes even after accounting for these effects.

Leverage effects in the level factor

As in Subsection (ref), we consider the score of the level factor as a proxy for the overall level of volatility. Denoting by $\overline{P}_t=\frac{1}{7d}\sum_{i=1}^d\sum_{j=1}^{7}P_{t+1-j}^{(i)}$ the rolling weekly average and by $\Delta\overline{P}_t=\overline{P}_t-\overline{P}_{t-7}$ the non-overlapping weekly difference, we then construct estimates of $\Delta\log(\widehat{S}_1(t))=\mathbb{E}\left[ \log(S_1(t))-\log(S_1(t-7))\mid \Delta \overline{P}_{t-7} \right]$ by partitioning the range of lagged price changes into 20 equal-frequency bins $B_b$ for $b=1,\ldots ,20$ and estimating the conditional mean within each bin, $\mu_b$ via the regression \[ \Delta\log (S_1(t)) = \sum_{b=1}^{20} \mu_b\mathbf{1}_{\lbrace \Delta\overline{P}_{(t-7)} \in \mathcal{B}_b\rbrace} + \varepsilon(t). \] Hence, $\Delta\log(\widehat{S}_1(t))$ estimates the expected change in the log-level of RCV one week ahead, conditional on the price change observed the previous week. Standard errors are computed using the Newey-West (HAC) estimator with 14 lags to account for serial dependence induced by the overlapping rolling window. This is similar to the “news impact curve” construction in, e.g., EngleNg1993. We construct four progressive specifications of news impact curves to isolate the leverage effect from confounding dynamics. In each case, both $\Delta \log(\widehat{S}_1(t))$ and $\Delta \overline{P}_{t-7}$ are residualized on a common subset of controls before estimation as follows (residualizing the price level ensures that the bins carry no information about the controls)

enumerate• No controls. • Price/level-scaling control via $\sinh^{-1}(\overline{P}_{t-7})$. • Mean-reversion control via $\log (S_1(t-7))$. • Joint price/level-scaling and mean-reversion control.

The results of this construction are depicted for the German market in Figure (ref), where we also include a regression line of the form

equation[equation omitted — 217 chars of source]

The figure reveals that, unconditionally, there seems to be some asymmetry, but that many of the estimated $\widehat{\mu}_b$ for $\overline{P}_{t-7}>0$ are not statistically different from 0. A standard Wald test rejects $\beta_+=\beta_-$ at the 5% level, suggesting some asymmetry. The price/level-scaling control dampens some of the effect and seemingly induces asymmetry that favours co-movements between price and volatility, while the mean-reversion control induces a very clear asymmetry that favours an inverse leverage effect. The joint control, however, essentially flattens the news impact curve to the point where there is no asymmetry, thus indicating that, upon controlling for the current state, most of the price/volatility co-movement disappears. The finding of no asymmetry is supported by the Wald test statistic, which is now unable to reject $\beta_+=\beta_-$. The corresponding news impact curve plots are depicted for Norway and Spain on Figures (ref) and (ref) in Appendix (ref), which reveal different patterns, but with the same ultimate conclusion that, after the joint control, any asymmetry is eliminated. These findings are very interesting, since it seems that asymmetric relationships between volatility and price changes in electricity markets can largely be explained by the price and volatility levels themselves at the aggregate level. Hence, when modeling price averages and their volatility, it seems that state-dependence in price and volatility is more relevant than correlation of innovations.

remarkOne may wonder if negative prices exhibit a different kind of leverage effect than positive prices. Although negative prices occur somewhat frequently at the hourly level in some generation zones, the average daily price is never negative in any of the zones throughout the dataset. Hence we can get no such information, since there are no negative values to study.
figure[figure omitted — 1,381 chars of source]

Functional leverage curves

So far, we have only viewed the leverage effect through the lens of the factor scores. While informative, this approach is unable to detect more subtle effects, such as leverage effects in individual hours, which may be washed away in the aggregation. In this Section, we therefore take a closer look at the finer structure of leverage on the level of individual hours. To do so, we let $IV_{j}^{(h)}=\widehat{\Sigma}_{\mathrm{pl}}^w(h,h)$ and run the regression \[ IV^{(h)}_j = \beta_0^{(h)} + \beta_+^{(h)} \max\lbrace\Delta\overline{P}_{j-w},0\rbrace + \beta_{-}^{(h)} \min \lbrace \Delta\overline{P}_{j-w},0\rbrace + \varepsilon_h(j), \] where $\overline{P}$ again denotes the weekly average price. From this we construct a functional leverage curve $h \mapsto \left(\beta_+^{(h)}, \beta_-^{(h)}\right)$ showing the diurnal pattern by which volatility across the day responds to price shocks. In the absence of leverage, we expect these curves to be overlaying, and if there is no difference across the hours of the day, we expect the curves to be flat. Keeping in line with the previous section, we conduct an unconditional weak form regression and a conditional strong form regression, where the hourly price level and mean-reversion effects have been partialled out. We depict the functional leverage curves on Figure (ref), where a similar pattern as observed in the aggregated case emerges. Only few hours exhibit weak form leverage, and any asymmetry disappears after controlling for price/level-scaling and volatility mean-reversion across all zones. Furthermore, most estimates of $\beta_+^{(h)}$ and $\beta_-^{(h)}$ are not statistically different from zero. This supports the findings of Section (ref) and suggests that leverage effects generally can be explained by other mechanical effects and that there is no meaningful difference across delivery hours.

figure[figure omitted — 1,627 chars of source]

Conclusion

This paper tackles the challenging problem of volatility estimation in the panel of electricity spot prices. By interpreting the panel as local averages of a latent function-valued price process, we construct an easy to implement estimator of the integrated conditional covariance operator that admits a simple decomposition into drift, propagation, and innovation terms. The estimator reveals a complex structure in the weekly realized covariation of electricity spot prices throughout several generation zones. Mean-reversion effects play a significant role in the weekly price variation and must be accounted for; in doing so, we find that mean-reversion is seemingly stronger in the early hours of the day. Peak-load and evening periods are associated with higher levels of overall volatility and we have documented that this correlates highly with the price level itself. Additionally, the realized production mix is important in explaining volatility dynamics, but the relative importance of renewables, fossil fuels, etc., varies greatly between generation zones. Consequently, each generation zone requires individual modeling tools and there is no “one size fits all” model across markets. We have zoomed in on the so-called inverse leverage effect, which is the notion that shocks to spot prices are positively correlated with shocks to volatility. We find evidence that such inverse leverage effects can be spurious and naturally arise from the strong state-dependence between volatility and price levels as well as mean-reversion in volatility itself.

Our findings open up several venues of future research into the volatility of electricity markets. In particular, it will be interesting to look deeper into the interplay of volatility and renewable generation, seasonal and diurnal patterns, and forecasting of the RCV estimator.

Disclosure statement

The authors declare no conflicts of interest.

Funding

T. K. Kloster gratefully acknowledges financial support from the Center of Research in Energy: Economics and Markets and The Danish Council of Independent Research under DFF grant 10.46540/5247-00005B. F. E. Benth gratefully acknowledges financial support from the SURE-AI Centre grant 357482, Research Council of Norway.

\printbibliography