Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Dynamically Consistent Analysis of Realized Covariations in Term Structure Models
scriptsize\begin{abstract}
In this article we show how to analyze the covariation of bond prices nonparametrically and robustly, staying consistent with a general no-arbitrage setting. This is, in particular, motivated by the problem of identifying the number of statistically relevant factors in the bond market under minimal conditions. We apply this method in an empirical study which suggests that a high number of factors is needed to describe the term structure evolution and that the term structure of volatility varies over time.
\end{abstract}
Introduction
We present a nonparametric method to measure covariations in a general arbitrage-free term structure setting in the spirit of bjork1999.
A motivation is to determine the number of statistically relevant random drivers needed to describe bond market dynamics under minimal assumptions and in a dynamically consistent manner. That is, by working coherently in the abstract setting of bjork1999, we circumenvent the well-known consistency problems of finding arbitrage-free finite-dimensional dynamics that reflect the empirical observations.
In the bond market the notion of a zero coupon bond is fundamental. A zero coupon bond guarantees its holder at time $t$ a fixed amount of money at some time $t+x$ in the future. The cost $P_t(x)$ of entering this contract at time $t$ is the price of this bond, which depends on the time to maturity $x$, implying a price curve $x\mapsto P_t(x)$ at each time $t\geq 0$ called the discount curve.
We assume to observe bond price or yield curve data, potentially derived by smoothing as in FPY2022 or liu2021, such that we can recover log bond prices
$$P_{i,j}^n:=\log P_{i\Delta_n}(j\Delta_n)\quad j=0,1,...,\lfloor M/\Delta_n\rfloor, \,\, i=0,1,...,\lfloor T\Delta_n\rfloor$$
for a resolution $\Delta_n=1/n$ and
with different maturities
$j\Delta_n $ and on different time points $i\Delta_n $.
Here $M>0$ is some maximal time to maturity (e.g. $M=10$ or $M=30$ when time is measured in years) and $T$ is the time until which the data are observed or considered.
Often, risk factor analyses are conducted on the basis of transformations of the discount curve, such as yield differences or excess returns,
which typically suggests that three factors explain a large amount of variation in bond market dynamics, c.f. Litterman1991.
Recently Crump2022, raised the concern that these low-dimensional factor structures are obtained irrespectively of the data generating process due to the high correlation of bond prices with close maturities.
To remedy this effect, dimension reduction could be based on difference returns, which are the returns of the trading strategy of buying an $x+\Delta$-maturity bond and shorting an $x$-maturity bond. Precisely, difference returns
are defined for $i=1,...,\ensuremath{\lfloor T/\Delta_n\rfloor}-1$ and $j=1,...,\lfloor M/\Delta_n\rfloor-1$
by
align[align omitted — 103 chars of source]
In this article we develop an asymptotic econometric theory for the realized covariations of these difference returns, that is, for $j_1,j_2=1,...,\lfloor M/\Delta_n\rfloor$ we analyse the covariations
$$\hat q_T^n(j_1,j_2):=
\sum_{i=1}^{\ensuremath{\lfloor T/\Delta_n\rfloor}} d_i(j_1)d_i(j_2).$$
Importantly, while $\ensuremath{\lfloor T/\Delta_n\rfloor}^{-1} q_T^n$ is the empirical covariance of the data $d_1^n,...d_{\ensuremath{\lfloor T/\Delta_n\rfloor}-1}^n$, assuming w.l.o.g. $\mathbb E[d_i^n]=0$, we do not consider it as an estimator of the population covariance of difference returns. Such an interpretation
is
not invariant with respect to the resolution $\Delta_n$ and requires the restrictive assumption that difference returns are i.i.d. or at least covariance stationary and ergodic. More importantly, it is not clear how dimension reduction can be conducted without entailing arbitrage opportunities.
General arbitrage-free term structure models in the sense of bjork1997 and FilipovicTappeTeichmann2010b require that forward rates
$ f_t(x)=-\partial_x\log(P_t(x))$ for $ x,t\geq 0,$
satisfy dynamics of the form
align[align omitted — 74 chars of source]
where the equation holds in an appropriate function space.
The latent process
$X$ is a possibly infinite-dimensional It{\^o} semimartingale
$$X_t:= \int_0^t \alpha_s ds+\int_0^t\sigma_s dW_s+J_t.
$$
where $\alpha$ is a curve-valued drift process, $\sigma$ is the (in general operator-valued) volatility, $W$ is an (in principle infinite-dimensional) Wiener process and $J$ is a jump process that we assume to model rare extreme events (c.f. Section (ref) for the details). Under a risk-neutral measure, the drift $\alpha$ can in addition be characterized as a deterministic function of the volatility $\sigma$ and characteristics of the jump process $J$ (c.f. bjork1997, FilipovicTappeTeichmann2010b). Parametrizations of forward cuves are then required to be viable in the dynamic setting (ref) to avoid the introduction of arbitrage opportunities to the model. This is known to be an intricate problem in term structure modelling (see e.g. bjork1999, bjork2001,
filipovic2003,
Filipovic2000, filipovic2000b
) and some frequently employed parametrizations of forward curves are incompatible with arbitrage-free dynamics or induce restrictive additional conditions (see e.g. filipovic1999).
The concerns on realized covariations of difference returns can be resolved when we consider infill asymptotics ($n\to \infty$). Precisely,
we show without imposing further assumptions and interpreting $\hat q_T^n$ as a piecewise constant kernel
that
equation[equation omitted — 238 chars of source]
where the limits hold in $L^2([0,M]^2)$. The right hand side describes the quadratic covariation of the latent driver $X$, which always exists (see e.g. Schroers2024) and can equivalently be described as the limit of covariance operators in $L^2(\mathbb R_+)$ by
equation[equation omitted — 209 chars of source]
where $\langle \cdot, \cdot\rangle$ is the inner product of $L^2(\mathbb R_+)$ and $\Delta_i^n X:= X_{i\Delta_n}-X_{(i-1)\Delta_n}$ denotes the $i$'th increment of $X$.
The limit (ref) implies that it is possible to infer on the number of random drivers in the bond market on the basis of discrete bond price data without further assumptions on moments, stationarity and ergodicity.
In fact, if the quadratic variation of $X$ is $d$-dimensional for $d\in \mathbb N$, then $X$ is $d$-dimensional and its state space (until time $T$) is spanned by the eigenvectors corresponding to the nonzero eigenvalues of $[X,X]_T$. The eigenvalues of $[X,X]_T$ indicate the amount of variation that the corresponding random factor explained of $X$ up to time $T$.
This is in contrast to the explained variation of factors derived from a covariance of $f$ or its increments, which is a priori not informative on the number of random drivers. In fact, it is possible that $f$ is infinite-dimensional, while $X$ is a one-dimensional process (c.f. Example 4.1 in Schroers2024).
The most striking advantage of the interpretation of the realized covariation of difference returns in (ref) is, however, that exchanging $X$ by an arbitrary finite-dimensional semimartingale (with the correct form of the drift) in the formulation of the dynamics in (ref) does not affect the capability of the model to be free of arbitrage, such that an investigation of the number of relevant factors can be conducted independently of further consistency conditions. This underlines that our abstract infinite-dimensional setting relaxes the analysis of the term structure when compared to models in which state spaces are assumed to be finite-dimensional a priori. An additional advantage of quadratic variations is that they are naturally interpreted as time-varying objects enabling their temporal analysis.
In practice, the distortion of the measurements due to ouliers can bias the analysis of the relevant factors.
For instance, in the context of sudden interest rate movements during an economic crisis it is possible that a single outlying difference return impacts the measurement of covariations and, thus, the measured dimensionality of the driver $X$ substantially.
For this reason, besides $[X,X]$,
the continuous part $[X^C,X^C]$ of the quadratic variation where $X^C_t= \int_0^t \alpha_s ds+\int_0^t \sigma_s dW_s$ is central for the task of identifying the statistically relevant number of random drivers, considering the jump part to model outlying events.
We describe how estimation of the continuous quadratic covariation is possible by a truncation technique, which sets outlying difference returns in the realized covariation on the left of (ref) to $0$.
We derive rates of convergence and a central limit theorem for these estimators and also show how the long-time limit of $[X^C,X^C]_T/T$ as $T\to \infty$ can be estimated, if it exists.
Our limit theory holds under weak assumptions, which mainly reflect those for finite-dimensional semimartingales, although, $f$ does not need to be a semimartingale.
We conduct an emprical study on the relevant drivers in the market via this covariation estimates based on real bond market data. The procedure is numerically equivalent to a principal component analysis based on the (truncated) empirical covariance of difference returns with a daily resolution and mean zero.
However, the classification of jumps takes into account the abstract setting (ref).
By investigation of a truncated version of the realized covariations $q_j^*(x,y):=\Delta_n^{-2}
\sum_{i=\lfloor ( j-1)/\Delta_n \rfloor}^{\lfloor j/\Delta_n \rfloor} d_i(\lfloor x/\Delta_n\rfloor)d_i(\lfloor y/\Delta_n\rfloor)$ for any year $j$ from 1990 to 2022, we find evidence for the dimension of the driver to vary from year to year but also to be consistently high (in each year more than $8$ drivers are needed to explain at least $99\% $ of the variation).
We further observe that quadratic variations vary in shape over time and not just their level.
We provide Monte-Carlo evidence for the validity of the limit theory in the context of sparse and noisy data.
Formal validity of our method is guaranteed by relating difference returns via cross-sectional and temproal discretization to the abstract setting (ref) and then apply the
results from the article Schroers2024.
The truncation procedure is inspired from the truncated realized variation estimators of Mancini2001, Mancini2004, Mancini2009 and Jacod2008 for finite-dimensional semimartingales.
Here, we also provide a data-driven variant of the truncation rule, to account for functional outliers in a similar way as the trimmed least squares method in ren2017.
Naturally, our asymptotic theory also allows for nonparametric estimation of characteristics of infinite-dimensional volatility models in continuous time employed for term structure modeling (c.f. BenthRudigerSuss2018, BenthSimonsen2018, BenthSgarra2021, Cox2020, Cox2021, Petersson2022 and cox2023).
The article is structured as follows: Section (ref) describes the general bond market setting that we consider for this article. Section
(ref) presents the estimation theory for the central application of term structure models. Identification for the quadratic variations of $X$, $X^C$ and $J$ on the basis of difference return variations can be found in Section (ref), while rates of convergence for estimating $[X^C,X^C]$ and a central limit theorem can be found in Section (ref). Section (ref) discusses long-time asynmptotics for estimation of a stationary mean of $[X^C,X^C]_T/T$. Section
(ref) contains practical considerations on smoothing of discrete bond price data and presents a data-driven truncation rule for robust estimation.
Section (ref) provides a simulation study.
Finally, we apply our theory to bond market data in Section (ref).
Technical proofs of our results along with further remarks on the simulation scheme and additional empirical results can be found in the appendix.
Technical preliminaries and notation
Let $I$ be an interval in $\mathbb R$.
We write $\langle h,g\rangle= \int_I h(x)g(x) dx$ for the $L^2$-scalar product of two elements $h,g\in L^2(I)$ as well as $\|h\|=\sqrt{\langle h,h\rangle}$ for the norm.
We write $L_{\text{HS}}(L^2(I))$ for the Hilbert space of Hilbert-Schmidt operators from $L^2(I)$ into itself and $\|\mathcal T_k\|_{\text{HS}}$ for the Hilbert-Schmidt norm of an $\mathcal T_k\in L_{\text{HS}}(L^2(I))$.
Recall, that
a Hilbert-Schmidt operator $\mathcal T_k:L^2(I)\to L^2(I)$ can be uniquely associated to a kernel $k\in L^2(I^2)$ such that $\|\mathcal T_k\|_{\text{HS}}=\|k\|_{L^2(I^2)}$ and
equation[equation omitted — 124 chars of source]
Importantly, the operator $h\otimes g:= \langle h,\cdot\rangle g$ is Hilbert-Schmidt for two elements $h,g \in L^2(I)$.
We shortly write $h^{\otimes 2}=h\otimes h$.
Finally, for $L^2(I)$-valued processes $X^n, n\in \mathbb N$, $X$, we write $X^n\stackrel{u.c.p.}{\longrightarrow}{X}\quad \text{ as }n\to\infty$ for the convergence uniformly on compacts in probability, i.e. it is $\mathbb P[\sup_{t\in [0,T]} \|X^n(t)-X(t)\|>\epsilon]\to 0$ for all $\epsilon, T>0$.
General Arbitrgae-Free Bond Market-Dynamics
Let $(\Omega,\mathcal F,(\mathcal F_t)_{t\geq 0},\mathbb P)$ be a filtered probability space with right-continuous filtration.
From here on, we assume that the forward rate process $(f_t)_{t\geq 0}$ is an $L^2(\mathbb R_+)$-valued stochastic process that is the mild solution to the stochastic partial differential equation (ref), defined on $(\Omega,\mathcal F,(\mathcal F_t)_{t\geq 0},\mathbb P)$.
That is,
align[align omitted — 164 chars of source]
where $ \mathcal S(t)f(x)=f(x+t)$ for $t\geq 0 $ and $f\in L^2(\mathbb R_+)$ defines the left-shift operator semigroup.
We relegate all further technical discussions on $X$, $\alpha$, $\sigma$, $W$ and $J$ and
various related technical assumptions that we need in for the validity of our limit theory to Sections (ref) in the appendix. We remark, that under arbitrage-free dynamics, that is, under an equivalent local martingale measure the drift is necessarily a deterministic function of $\sigma$ and $\gamma$ (c.f. bjork1997),
which was in the continuous case the original inside leading to the popular Heath-Jarrow-Morton framework of HJMoriginal for pricing bonds and interest rate sensitive contingent claims. However, this will not be of particular importance for our purposes, as the drift later vanishes asymptotically in our limit theory. We will discuss however some important examples subsequently. Before that, we make a remark on the choice $L^2(\mathbb R_+)$ as the state space of forward rate curves.
remark[On the forward curve space]
There are other choices for the state space of $f$ than $L^2(\mathbb R_+)$ such as the forward curve space of Filipovic2000.
We choose, however, to work in an $L^2$-setting because in that way we do not impose further regularity assumptions on the forward curves.
A supposed restriction of the state space $L^2(\mathbb R_+)$ is that the so-called long-rates $\lim_{x\to \infty}f_t(x)$ are equal to $0$. This is undesirable from a financial point of view and could easily be fixed in several ways. For instance, we could consider the Hilbert space
$H:=\mathbb R \oplus L^2(\mathbb R_+)=\{f:\mathbb R_+\to \mathbb R: f(x)=a+h(x), a\in\mathbb R, h\in L^2(\mathbb R_+) \}$, for which the first component models the long-rate. In this case, the forward curve space of Filipovic2000 would be contained as a subspace.
We then might just assume that the state spaces of $X^C$ and $J$ belong to $\{0\}\times L^2(\mathbb R_+)\equiv L^2(\mathbb R_+)$, which
is in line with the assumptions on the volatilities in FilipovicTappeTeichmann2010b. As $[X,X]$, $[X^C,X^C]$ and $[J,J]$ do not depend on the drift and the initial condition, the respective limit theory would be exactly the same. Another reason that justifies our choice is that in practice, our asymptotic analysis just
takes into account bond price data $P_t(x)$ with $(t,x)\in [0,T]\times [0,M]\subset \mathbb R_+^2$ for some maximal time to maturity $M<\infty$ and the behavior of the forward curves for $x\to \infty$ is of minor importance.
To relax the notation, we stick to the state space $L^2(\mathbb R_+)$ without loss of generality.
Let us now discuss some important simple examples.
example[Sum of $Q$-Wiener and compound Poisson process]
As a simple example assume $\alpha_t=\alpha \in L^2(\mathbb R_+)$ and $\sigma_t=\sigma \in L_{\text{HS}}(\mathbb R_+)$ to be constant. Then, $X^C_t$ is a Gaussian random variable in $L^2(\mathbb R_+)$ with mean $t\alpha$ and covariance $tQ:=t\sigma\sigma^*$, where $\sigma^*$ is the Hilbert space adjoint of $\sigma$. equivalently, $X_t^C$ has covariance kernel $q$ given by $\int_{\mathbb R_+} q(x,y)f(y) dy=(Q f)(x)$ for $x\geq 0$. Since we want use the jump process to model rare outliers, a reasonale model would be a compound Poisson process
$$J_t:=\sum_{i=1}^{N_t} \chi_i,$$
for an i.i.d. sequence $(\chi_i)_{i\in \mathbb N}$ of random variables in $L^2(\mathbb R_+)$ with law $F$ and finite second moment ($\mathbb E[\|\chi_i\|^2]<\infty$) and a Poisson process $N$ with intensity $\lambda>0$.
Since in the technical Section (ref) we require the jump process to be a martingale, this is not immediately a valid choice, but we can rewrite the dynamics accordingly (c.f. Example (ref) in the appendix).
The quadratic covariation of this semimartingale is then
$$[X,X]_t=[X^C,X^C]_t+[J,J]_t= t Q + \sum_{i=1}^{N_t} \chi^{\otimes 2}_i.$$
The term structure setting (ref) contains the vast majority of existing arbitrage-free term structure models considered in the literature. Among them is the class of affine term structure models, which are widely appreciated for their parsimony.
example[Affine term structure models]
In an affine term structure model, the state space for forward curves is spanned by a finite amount of factors, such that
\begin{align}
f_t(x)=g_0(x)+g_1(x) x_t^1+...+g_d(x)x_t^d,
\end{align}
where $g:\mathbb R_+\to\mathbb R$ are some particularly suitable functions and the process $x=(x^1,...,x^d)$ is an affine process (c.f. Duffie2003). To guarantee that a factor structure like (ref) can be in line with
general no arbitrage dynamics of the form (ref) one has to impose restrictions on both the functions $g_0,g_1,...,g_d$ as well as the multivariate semimartingale $x$ (c.f. e.g. Filipovic2000, Section 7.4 for a description in a continuous setting).
If $[x^c,x^c]$ and $[x^d,x^d]$ denote the continuous and discontinuous part of the multivariate quadratic variation of $x$, it is
$$[X^C,X^C]_tt= \sum_{i=1}^d[x_i^c,x_j^c]_t g_i\otimes g_j,\qquad [J,J]_t= \sum_{i=1}^d[x_i^d,x_j^d]_t g_i\otimes g_j.$$
If the respective bond market data are in line with an affine model such as (ref), our theory in Section (ref) identifies this structure asymptotically.
Let us outline two classical special cases when $d=1$ and when there are no jumps:
\begin{itemize}
• (Va{\v s}i{\v c}ek model) In the Va{\v s}i{\v c}ek model it is
$$dx^1_t= (b-ax_t^1)dt+\sigma_0 d\beta_t$$
for $b,a,\sigma_0> 0$ and a one-dimensional Brownian motion $\beta$. The functions $g_0,g_1$ are then given by $g_1(x)=e^{-ax}$ and $g_0(x)=b\int_0^x g_1(y)dy-(\sigma_0^2/2)(\int_0^x g_1(y)dy)^2$ (c.f. Section 7.4.2 in Filipovic2000).
In this case, the quadratic variation of the latent driving semimartingale $X$ is
$$ [X,X]_t=[f,f]_t= t \sigma_0^2(e^{-a \cdot})^{\otimes 2},\qquad t\geq 0.$$
With $Q=\sigma_0^2(e^{-a \cdot})^{\otimes 2}$, this is a special case of Example (ref) without jumps.
• (CIR model) In the CIR model we have short rate dynamics of the form $$dx_t^1=(b-ar_t)dt+\sigma_0 \sqrt {x^1_t}d\beta_t $$
for $b, a, \sigma_0 \geq 0$ and a standard univariate Brownian motion $\beta$. The function $g_1$
is given as the derivative of the function $x\mapsto G_1(x):=2 (e^{cx}-1)/((c-a)(e^{cx}-1)+2c)$ with $c=\sqrt{a^2-2\sigma_0}$ and the function $g_0$ is given as $g_0=aG_1$ (c.f. Section 7.4.1) in Filipovic2000).
In this case the quadratic variation of the latent driving semimartingale $X$ is
$$[X,X]_t= [f,f]_t=\sigma_0^2\left(\int_0^t x^1_t ds\right) g_1^{\otimes 2},\qquad t\geq 0.$$
\end{itemize}
Many simple ways to model term structures lead to nonaffine dynamics, such as
example[Volterra spot rate models]
For many term structure models the quadratic variation of $f$ is not necessarily well-defined, as it must not be a semimartingale. For instance, take the forward rate dynamics of the form
\begin{align}
f_t=f_0+\int_0^t \alpha_s(\cdot+t-s) ds +\sum_{i=1}^{d}\int_0^t k_i(\cdot+t-s) \sigma_s^i d\beta_s^i
\end{align}
for a multivariate standard Brownian motion $(\beta^1,...,\beta^d) $ for some $d\in \mathbb N$, a $d$-dimensional volatility $(\sigma^1,...,\sigma^d)$ and deterministic kernels $k^1,...,k^d\in L^2(\mathbb R_+)$ as well as a drift $\alpha$ which satisfies the HJM condition (c.f. Filipovic2000).
In this scenario, the underlying driving semimartingale has quadratic variation equalling
$$[X,X]_t= \sum_{i=1}^{d} \left(\int_0^t (\sigma_s^i)^2ds\right)k_i^{\otimes 2}.$$
Thus, if $\sigma_s^i= 1$ for all $s\geq 0$ and $i\in \mathbb N$, this is a special case of Example (ref) with $Q=\sum_{i=1}^{d}k_i^{\otimes 2}$ and without jumps, but it does not always correspond to an affine term structure.
In the energy market, for instance, fractional kernels such that $k_1(t)=\mathcal O(t^H)$ for $t\to 0$ and $k_i\equiv 0$ for $i\geq 2$ are used to model energy spot prices (c.f. BENNEDSEN2017 or BNBV2013). If we assume that $k_1(t)=t^H$, $\sigma^1\equiv 1$ and $\alpha_s\equiv 0$ for $t\in [0,1]$ one can prove that
$f_t$ is not a semimartingale in $L^2(\mathbb R_+)$ and the quadratic variation does not converge as we show in Appendix (ref).
Example (ref) (and also Example 3.16 in BSV2022) shows that the process $(f_t)_{t\geq 0}$ is in general not an $L^2(\mathbb R_+)$-valued semimartingale. However, all implied bond prices $(P_t(T-t))_{0\leq t\leq T}$ are semimartingales for all $T>0$, which is necessary to guarantee the absence of arbitrage in the bond market.
Moreover, while the quadratic covariations of $f$ must not be convergent, we show in the next section, in which we present our main results, that the realized covariation of difference returns measures quadratic covariations of the latent driver asymptotically and without further conditions.
Estimation of Quadratic Covariations
In this section we present our asymptotic theory for estimation of quadratic variations. We start with the identifiability of $[X,X]$, $[X^C,X^C]$ and $[J,J]$.
Identification of the quadratic covariation of the latent semimartingale
We rely on infill asymptotics
$\Delta_n\to 0$ as $n\to \infty$
and recall the definition of the realized covariation $(\hat q_t^n)_{t\geq 0}$ as a piecewise constant kernel. That is,
for $x\in [(j_1-1)\Delta_n,j_1\Delta_n] $ and $y\in [(j_2-1)\Delta_n,j_2\Delta_n]$, $j_1,j_2\in\mathbb N$ and $t\geq 0$ we define
equation[equation omitted — 158 chars of source]
For each $n\in \mathbb N$ and $t\geq 0$, we have that $\hat q_t^n\in L^2(\mathbb R_+^2)$, which follows from the Assumption that forward curves are elements in $L^2(\mathbb R_+)$ (c.f. Remark (ref) below).
We now state the general identifiability result for the quadratic covariation of $X$.
theoremIt is as $n\to \infty$ and w.r.t. the Hilbert-Schmidt norm and $\mathcal T_{\hat q^n}$ as in (ref)
\begin{equation}
\Delta_n^{-2}\mathcal T_{\hat q^n}\overset{u.c.p.}{\longrightarrow} [X,X].
\end{equation}
As in the case of finite-dimensional semimartingales, we do not have to impose any further conditions to identify the quadratic covariations of the driving semimartingale $X$ although the observable process $(f_t)_{t\geq 0}$ is not necessarily an $L^2(\mathbb R_+)$-valued semimartingale. This is due to the relation of difference returns to semigroup-adjusted forward rate returns, which where shown in Schroers2024, Benth2022 and BSV2022 to be well-suited for volatility estimation for processes of the form (ref). This relationship is made clear in the next remark.
remark[Difference returns are discretized semigroup-adjusted increments]
The reason for economically motivated difference returns to lead to such a general identifiability result is that difference returns
coincide with orthonormal projections onto semigroup-adjusted increments of forward rate curves. That is, we have
$$
d_i^n(j)=-\langle \tilde \Delta_i^n f,\ensuremath{\mathbb{I}}_{[(j-1)\Delta_n,j\Delta_n]}\rangle.$$
where $\tilde \Delta_i^n f$ denotes the semigroup-adjusted forward rate increment
$$\tilde \Delta_i^n f:= f_{i\Delta_n}-\mathcal S(\Delta_n) f_{(i-1)\Delta_n}\quad i=1,...,\ensuremath{\lfloor T/\Delta_n\rfloor}.$$
and $(\mathcal S(t))_{t\geq 0}$ is the left shift semigroup on $L^2(\mathbb R_+)$.
Define
\begin{equation}
\Pi_{n,M} h:= n \sum_{j=1}^{\lfloor M/\Delta_n\rfloor}\langle h, \ensuremath{\mathbb{I}}_{[(j-1)\Delta_n,j\Delta_n]}\rangle \ensuremath{\mathbb{I}}_{[(j-1)\Delta_n,j\Delta_n]}
\end{equation}
the projection onto $span(\ensuremath{\mathbb{I}}_{[(j-1)\Delta_n,j\Delta_n]}:j=1,...,\lfloor M/\Delta_n\rfloor)$ and observe that
\begin{align*}
\Delta_n^{-2} \mathcal T_{\hat q^{n,M}_t}=
\Pi_{n,M}(SARCV_t^n)\Pi_{n,M},
\end{align*}
where $SARCV_t^n=\sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} \tilde \Delta_i^n f^{\otimes 2}$ is the semigroup-adjusted realized covariation, which was shown to be a consistent estimator of
$[X,X]_t$ in Schroers2024 in the presence of jumps
and a consistent and asymptotically normal estimator of $[X^C,X^C]$ in Benth2022 and BSV2022 when $J\equiv 0$. This characterization of the realized covariation $\hat q^n$ also explains the appearance of the scalar $\Delta_n^{-2}$ in front of the covariation in (ref).
Next, we examine how to identify the continuous part of the quadratic covariation.
Identification of $[X^C,X^C]$ and $[J,J]$ via truncated covariation estimators
We now turn to the estimation of the continuous part of the quadratic covariation. We will derive jump robust estimators by a truncated form of $\hat q_t^{n}$ defined
by the piecewise constant kernel
$\hat q^{n,-}_t$ given for $x\in [(j_1-1)\Delta_n,j_1\Delta_n]$, $y\in [(j_2-1)\Delta_n, j_2\Delta_n]$, $j_1,j_2\in \mathbb N$ and $t\geq 0$ by
align[align omitted — 194 chars of source]
for $u_n=\alpha \Delta_n^{w}$, with $w\in (0,1/2)$ and $\alpha>0$
and a particular sequence of truncation functions $g_{n}$ that takes into account only the discrete data $d^i_n(j)$ for $j\in \mathbb N$.
Precisely, the corresponding sequence of truncation functions $g_n: l^2\to \mathbb R_+$ must satisfy for constants $c,C>0$ and for all $f,h\in l^2$ and all $n\in \mathbb N$
align[align omitted — 124 chars of source]
While the particular choice of the functions $g_n$ will not play a role for the asymptotic behavior of $\hat q^{n,-}_t$, it is important to modify it in practice. For the moment, one can take in mind the legitimate choice $g_n=\|\cdot\|_{l^2}$ for all $n\in \mathbb N$ for which we have with $\Pi_{n,\infty}$ defined as in (ref) for $M=\infty$ that
$\|d_i^n/\Delta_n\|_{l^2}= \|
\Pi_{n,\infty} \tilde \Delta_i^n f\|_{L^2(\mathbb R_+)}.$
We will discuss a data-driven specification of $g_n$ and the truncation level in Section (ref).
The next result states that $\hat q^{n,-}$ consistently estimates the quadratic covariation of $X^C$.
theoremUnder Assumption (ref)(2) and with the notation of (ref) it is as $n\to\infty$
$$\Delta_n^{-2}\mathcal T_{\hat q^{n,-}_{\cdot}} \overset{u.c.p.}{\longrightarrow}[X^C,X^C]. $$
Let us make a remark on the feasibility of the estimator.
remarkIn practice, we do not observe the $d_i^n(j)$ for all $j\in \mathbb N$ but rather up to a finite maturity $M$, that is ,for $j\in 1,...,\lfloor M/\Delta_n\rfloor-1$.
Consistency of $\hat q^{n}$ from Theorem (ref) implies the consistency of $\hat q^{n}\big|_{[0,M]^2}$ for each $M>0$, so there is no problem when we do not consider truncation.
However, $\hat q^{n,-}\big|_{[0,M]^2}$ is not a feasible estimator in this context, since it uses in the truncation function the whole infinitely long vector $(d^i_n(j))_{j\in\mathbb N}$.
A feasible estimator $\hat q^{n,M,-}$ is defined for $x\in [(j_1-1)\Delta_n,j_1\Delta_n]$, $y\in [(j_2-1)\Delta_n, j_2\Delta_n]$, $j_1,j_2\in \mathbb N$ and $t\geq 0$ by
\begin{align}
\Delta_n^{-2}\hat q^{n,M,-}_t (x,y):=
& \Delta_n^{-2}\sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}}d_i^n(j_1)d_i^n(j_2)\ensuremath{\mathbb{I}}_{ g_n\left((
\ensuremath{\mathbb{I}}_{[0,M]}(j\Delta_n)
d_i^n(j)/\Delta_n)_{j\geq 0}\right)\leq u_n}.
\end{align}
To relax the notation, we do not present the limit theorems for $\hat q^{n,M,-}$ in this Section.
However, all results that we state for $\mathcal T_{\hat q^{n,-}}$ with limit $[X^C,X^C]$, that is, Theorems (ref), (ref). (ref) hold for $\mathcal T_{\hat q^{n,-}}$ with limit $[\Pi_MX^C,\Pi_MX^C]$ and Theorem (ref) holds for $\mathcal T_{\hat q^{n,-}_T}/T$ with limit $\Pi_M \mathcal C\Pi_M$ where $\Pi_M h(x)= \ensuremath{\mathbb{I}}_{[0,M]}(x)h(x)$ for all $h\in L^2(\mathbb R_+)$. The formal proof for that can be found in the appendix.
Observe that we can also define an upward truncated estimator
$ \Delta_n^{-2}\hat q^{n,+}$ given for $t\geq 0$, $x\in [(j_1-1)\Delta_n,j_1\Delta_n]$, $y\in [(j_2-1)\Delta_n, j_2\Delta_n]$ and $j_1,j_2\in \mathbb N$ by
align*[align* omitted — 167 chars of source]
Obviously,
$\hat q^n=\hat q^{n,-}+\hat q^{n,+} $ and $\mathcal T_{\hat q^n}=\mathcal T_{\hat q^{n,-}}+\mathcal T_{\hat q^{n,+} }$.
Then, combining Theorem (ref) and Theorem (ref), we also obtain
corollaryIf Assumption (ref)(2) holds, we have as $n\to\infty$ that
$$ \Delta_n^{-2}\mathcal T_{\hat q^{n,+}_{\cdot}}\overset{u.c.p.}{\longrightarrow} [J,J].$$
This result shows that the quadratic covariation corresponding to the jump part is identifiable in the context of general bond market models.
However, the finer analysis of jumps is not part of this paper and relegated to future work. Instead,
we derive convergence rates for the estimation of the continuous part of the quadratic covariation in the next section.
Convergence rates and central limit theorem for estimation for $\hat q^{n,-}$
In order to derive rates of convergence and a central limit theorem for estimating the continuous part of the quadratic variation, we need to impose further regularity Assumptions, which depend on the smoothness of the kernel corresponding to the operators $[X^C,X^C]_t$. For the error bounds, this is Assumption (ref), which is discussed in detail in Section (ref). We discuss these Assumptions in the context of Example (ref) right below the subsequent abstract result.
theoremIf Assumptions (ref)(r) and (ref)($\gamma$) hold for some $r\in (0,2)$, $\gamma \in (0,1/2]$, it is for all $\rho<(2-r)w$ \begin{equation}
\sup_{t\in [0,T]}\left\|\Delta_n^{-2}\mathcal T_{\hat q^{n,-}_t}-[X^C,X^C]_t\right\|_{HS}=\mathcal O_p\left(\Delta_n^{\min(\rho,\gamma)}\right).
\end{equation}
In particular, if $r<2(1-\gamma)$, $w\in (\gamma/(2-r),1/2]$ it is
\begin{equation}
\sup_{t\in [0,T]}\left\|\Delta_n^{-2}\mathcal T_{\hat q^{n,-}_t}-[X^C,X^C]_t\right\|_{HS}=\mathcal O_p(\Delta_n^{\gamma}).
\end{equation}
The rate implied by Theorem (ref) is at most $\mathcal O_p(\Delta_n^{ 1/2})$, which is achieved if Assumptions (ref)(r) and Assumption (ref)($\gamma$) hold for $\gamma = 1/2$ and $r<1$.
Let us now discuss Theorem (ref) and Assumptions (ref)($\gamma$) and (ref)($r$) in the context of Example (ref).
example[Example (ref) ctn.]
Let us again assume that $\sigma$ is constant, write $Q:=\sigma\sigma^*$ and let $J$ be a compound Poisson process. Since $Q$ is Hilbert-Schmidt, it can be written as an integral operator $\mathcal T_{q}$ corresponding to a kernel $q\in L^2(\mathbb R_+^2)$. The regularity Assumption (ref)($\gamma$) for some $\gamma\in [0,1/2]$ is then guaranteed if for all $M>0$ it is
$$\sup_{r>0}\int_0^{M}\int_0^{M} \frac{(q(r+x,y)-q(x,y))^2}{r^{2\gamma}}dxdy<\infty.$$
This is the case, for instance, if $q$, as a function on $\mathbb R^2_+$ is locally $\gamma$-Hölder continuous.
The regularity Assumption (ref)(r) is foremost an Assumption on the jump activity. Indeed, in our case, in which the jumps correspond to a compound Poisson process with jump-distribution $F$, we always have that
$\int_{H\setminus \{0\}} (z\wedge 1)^r F(dz)\leq F(H\setminus \{0\})=1<\infty$ does hold for all $r> 0$, and hence, the Assumption holds for all $r\in (0,2)$. In particular, we can choose $w\in [\gamma ,\frac 12]$ and $r<2(1-\gamma)$ to derive the rate of convergence in (ref).
If we assume a slightly stronger Assumption than (ref)($1/2$), which can also be found in Section (ref), we can even obtain a stable CLT in the next result, where stable convergence in law is denoted by $\overset{st.}{\longrightarrow}$ \footnote{ Recall that a sequence of random variables $(X_n)_{n\in\mathbb N}$ defined on a probability space $(\Omega, \mathcal F,\mathbb P)$ and with values in a Hilbert space $H$ converges stably in law to a random variable $X$ defined on an extension $(\tilde{\Omega}, \tilde{\mathcal F},\tilde{\mathbb P})$ of $(\Omega, \mathcal F,\mathbb P)$ with values in $H$, if for all bounded continuous $f:H\to \mathbb R$ and all bounded random variables $Y$ on $(\Omega,\mathcal F)$ we have
$\mathbb E[Y f(X_n)]\to\tilde {\mathbb E}[Y f(X)]$ as $n\to\infty$,
where $\tilde {\mathbb E}$ denotes the expectation w.r.t. $\tilde{\mathbb P}$.}.
theoremLet Assumption (ref) hold.
Then Assumption (ref)($1/2$) holds. Moreover, let Assumptions (ref)(r) hold for $r<1$. Then for $w\in [1/(2-r),1/2]$ we have for every $t\geq 0$ that
$$\sqrt n \left(\Delta_n^{-2}\mathcal T_{\hat q^{n,-}_t}-[X^C,X^C]_t\right)\overset{st.}{\longrightarrow} \mathcal N(0,\mathfrak Q_t),$$
where $\mathcal N(0,\mathfrak Q_t)$ is for each $t\geq 0$ a Gaussian random variable in $L_{\text{HS}}(L^2(\mathbb R_+))$ defined on a very good filtered extension\footnote{See Section 2.4.1 in JacodProtter2012 for the definition of very good filtered extensions.} $(\tilde{\Omega},\tilde{\mathcal F},\tilde{\mathcal F}_t,\tilde {\mathbb P})$ of $(\Omega,\mathcal F,\mathcal F_t, \mathbb P)$ with mean $0$
and covariance process $\mathfrak Q_t:L_{\text{HS}}(L^2(\mathbb R_+))\to L_{\text{HS}}(L^2(\mathbb R_+))$ given as
$$\mathfrak Q_t K =\int_0^t \Sigma_s\left(K+K^*\right) \Sigma_s ds.$$
Here $\Sigma_t:=\sigma_t\sigma_t^*=\partial_t [X,X]_t$ is the squared volatility operator.
The partial derivative $\Sigma_t:=\partial_t [X,X]_t$ has to be interpreted as a Frechet-derivative and does always exists, due to the Assumption on $X$ being an It{\^o} semimartingale (c.f. Section (ref)).
Let us derive the form of $\mathfrak Q_t$ in the context of Example (ref):
example[Example (ref) ctn.]
in the setting of example (ref) the asymptotic covariance operator $\mathfrak Q_t$ has the form
$$\mathfrak Q_t=t (Q(\cdot+\cdot^*)Q).$$
Equivalently, $\mathfrak Q_t$ can be interpreted as a kernel operator on $L^2(\mathbb R_+^2)$ with kernel
$$\mathfrak q_t(x,z,w,y):= t\left(q(x,z)q(w,y)+q(x, w)q(z,y)\right).$$
In particular, $\mathfrak q_t(x,z,w,y)$ can be consistently estimated by the plug-in estimator
$\hat{\mathfrak q}_T^n(x,z,q,y):= \Delta_n^{-4}T^{-1}(\hat q_T^{n,-}(x,z)\hat q_T^{n,-}(w,y)+\hat q_T^{n,-}(x,w)\hat q_T^{n,-}(z,y))$.
Assumption (ref) holds, for instance, if $q$ is locally $\gamma$-Hölder continuous for some $\gamma >1/2$ except on finitely many discontinuity points. In particular, the CLT holds if $q$ is smooth.
So far we have discussed limit theorems for infill asymptotics leaving $T$ fixed. In the next section, we outline how it is possible under additional assumptions and letting $T\to\infty$ to make use of all available data to estimate the stationary instantaneous covariance for difference returns.
Long-time volatility estimation
The truncated estimation procedure described previously enables estimations of a time series of the integrated volatilities $\int_{i}^{i+1}\Sigma_s ds=[X^C,X^C]_i-[X^C,X^C]_{i-1} $ for $i\in \mathbb N$.
If the aim is to derive a time-invariant mean for the volatility, we have to impose further conditions, which are described in detail in Section (ref).
These Assumptions are much stricter than the ones we considered in the previous section for the infill asymptotics on finite intervals and in particular imply that the mean
$$\mathcal C:=\frac 1 T\mathbb E[[X,X]_T]$$
is independent of $T$.
However, they
allow us to derive a stationary mean of $\Sigma_t$ via large $T$ asymptotics.
theoremLet Assumptions (ref), (ref)(p,r) and (ref)($\gamma$) hold for some $r\in (0,2)$, $\gamma \in (0,1/2]$
and $p>\max(2/(1-2w),(1-wr)/(2w-rw))$.
Then we have as $n,T\to\infty$ that
$$\Delta_n^{-2}T^{-1}\mathcal T_{\hat q_T^{n,-}}\overset{p}{\longrightarrow}\mathcal C.
$$ If even $r<2(1-\gamma)$ and $w\in(\gamma/(1-2w),1/2)$ and additionally $p\geq 4$ we have with $a_T= \| [X^C,X^C]/T- \mathcal C\|_{\text{HS}}$ that
$$\left\| \Delta_n^{-2}T^{-1}\mathcal T_{\hat q_T^{n,-}}-\mathcal C\right\|_{L^2(\mathbb R_+^2)}=\mathcal O_p(\Delta_n^{\gamma}+a_T).
$$
Assumption (ref) does not impose very strong Assumptions on the dynamics of the volatility and is satisfied by most stochastic volatility models.
To verify this condition for particular models for the infinite-dimensional volatility process one might investigate the vast literature for ergodic properties of Hilbert space-valued processes and, in particular, SPDEs (c.f. DPZ2014 or PZ2007). For the existence of invariant measures for term structure models, we further mention vargiolu1999, Tehranchi2005, marinelli2010, rusinek2010,
and FFRS2020. Recently, friesen2022 examined the long-time behavior of infinite-dimensional affine volatility processes.
Here, we only review the validity of the Assumptions employed in Theorem (ref) in the context of our running Example (ref):
example[Example (ref) ctn.]
Once more, consdier the setting of Example (ref). Assumption (ref) requires stationarity and mean ergodicity on the continuous part of the quadratic variation, which is trivially fulfilled, since $\mathcal C=[X,X]_T/T=Q$ for all $T>0$. Assumption (ref)(p,r) is valid for all $p>0$ and $r>0$ since all coefficients of the semimartingale $X$ are deterministic and constant
Assumption (ref)($\gamma$)
holds for $\gamma\in (0,\frac 12]$ under analogous conditions in as Example (ref).
While $Q$ can be estimated without the long-time regime, it is simple to find situations when long time asymptotics provide additional information such as for the estimation of HEIDIH models from Petersson2022, which is described in Schroers2024. Another example, which is implemented in the simulation study in Section (ref) is that $\Sigma_s = x_s Q$ for a positive scalar mean reverting process $x$, which models the changing magnitude of volatility over time. Then the long time estimator can be used to determine the mean-reversion level of $x$.
Practical considerations
In this section, we discuss some practical complications on the implementations of the estimator. Namely, we present a data-driven truncation rule and comment on the use of nonparametrically smoothed yield or bond price curve data.
We start with a data driven choice of the truncation function $g_n$ and the tuning parameters $\alpha$ and $w$.
Truncation in practice
While the asymptotic theory of Section (ref) justifies the use of truncated estimators, the choice of the truncation level and functions remains a practical issue. Even in finite dimensions, this can be challenging and we refrain from finding optimal choices. However, we outline how the truncation rule can be reasonably implemented.
Truncation rules in finite dimensions often necessitate preliminary estimators for the average realized variance in the corresponding interval of interest (c.f. Mancini2012, page 418, for an overview of some truncation rules). One sorts
out a large amount of data first, to obtain a preliminary estimator of $[X^C,X^C]_T$. As this can be interpreted as the average covariation of the increments in the interval $[0,T]$, one then chooses truncation levels in terms of multiples of standard deviations as measured by the preliminary estimate.
In our infinite-dimensional setting, we mimic this procedure,
but it is harder to distinguish typical increments and outliers as we cannot argue componentwise.
While the choice $g_n\equiv \|\cdot\|$ leads to consistent estimators in terms of the limit theory developed in Section (ref), it is not necessarily a good choice in the context of finite data since the continuous martingale might vary considerably more in one direction
than another
.
We Therefore present a method that is based on a measure of functional outlyingness in the spirit of ren2017: Assuming that $\Sigma$ is independent of the driving Wiener process and does not vary too wildly on the interval $[0,T]$, and that no jumps exist, we have approximately that $\Sigma_t \approx \frac 1T\int_0^T \Sigma_s ds$ for $t\in [0,T]$ and $\tilde \Delta_i^n f/\sqrt{\Delta_n}|\Sigma\sim N(0, \frac 1T\int_0^T \Sigma_s ds)$.
If the largest $d$ eigenvalues $e_1,...,e_d$ of $\int_0^T \Sigma_s ds/T$ account for a large amount of the variation as measured by the summed eigenvalues of $\int_0^T \Sigma_s ds/T$ (e.g. 90 percent), we know that $P_d \tilde \Delta_t^n f$ with $P_d=\sum_{i=1}^d e_i^{\otimes 2}$ is a linearly optimal approximation of $\tilde \Delta_t^n f$ in the $L^2(\mathbb R_+)$-norm. We can also define
$P_d(\int_0^T \Sigma_s ds/T)^{-1}P_d =\sum_{i=1}^d(1/\lambda_i)e_i^{\otimes 2}$ and define $g^d(h):=\| P_d(\int_0^T \Sigma_s ds/T)^{-1}P_d h\|^2$.
This distance resembles the measure proposed in ren2017, however, it is not a valid truncation function, since (ref) cannot hold. Further, if a truncation at level $d$ is made, outliers impacting the higher eigenfactors might be overlooked.
We Therefore propose an adjusted method defining $g(x)=g^d(x)+ \frac{\|(I-P_d)x\|^2}{\sum_{i=d+1}^{\infty} \lambda_i} $. Then, for $p_n:l^2\to L^2(\mathbb R_+)$ given by
$p_n x:= \sum_{j=1}^{\infty} x_j \ensuremath{\mathbb{I}}_{[(j-1)\Delta_n,j\Delta_n]}$
we define the sequence of truncation functions $(g_n)_{n\in \mathbb N}$ via
equation[equation omitted — 243 chars of source]
where the index $d$ can be chosen freely as long as $\lambda_{d+1}>0$. E.g. we can choose $d$ such that the first $d$ eigenfactors for $\mathcal C$ explain $90$ % of the variation.
It is then with $\lambda_i^d:=\lambda_i$ for $i\leq d$ and $\lambda_i^d:=\sum_{j=d+1}^{\infty} \lambda_j$ for $i\geq d$ and with $\Delta_n$ small
align*[align* omitted — 256 chars of source]
In practice, we do not know the eigenvalues and eigenfunctions of $\int_0^T \Sigma_s ds/T$ and derive them from a preliminary estimate.
We suggest a simple truncation procedure in two steps:
itemize• First we have to specify a preliminary estimator which can be found as follows: For fixed $T$ choose a truncation level $u$, such that a large amount, say $0.25$, of the increments is sorted out by $\hat q_T^{n,-}$ with the choice $g_n(x)=\|\cdot\|_{l^2}$. That is, $25$ percent of the increments satisfy $g_n\left(
d_n^i/\Delta_n\right)
>u$. Then we define the preliminary estimate
\begin{align*}
\rho^*\Delta_n^{-2}T^{-1} \hat q^{n,-}_t \approx\frac 1T \int_0^Tq_t^cdt, \end{align*}
where $\rho^*>0$ properly rescales the preliminary estimator (one rescaling procedure is outlined in the appendix.
• Now set $g_n$ as in (ref), $d$ such that the first $d$ eigenvalues of the operator corresponding to the kernel $\rho^*\Delta_n^{-2}T^{-1} \hat q^{n,-}_t$ explain $90\%$ of the variation measured by the sum of eigenvalues of this operator and choose $u_n= l \sqrt{d+1}\Delta_n^{0.49}$ for an $l\in \mathbb N$. E.g. we might take $l=3,4$ or $5$. Observe that for $d$ large enough $g_n(d_n^i/\Delta_n)^2/\Delta_n\approx g^d(p_n(d_n^i/\Delta_n)^2/\Delta_n$ is under the above local normality assumptions approximately $\chi^2$ distributed with $d$ degrees of freedom. Hence, the probability that $g_n(d_n^i/\Delta_n)<u_n$ can be approximated by the cumulative distribution function of a $\chi^2$-distribution with $d$ degrees of freedom. For instance, if $d=4$ we have that $g_n(d_n^i/\Delta_n)<u_n$ with $l=4$ would hold for approximately 98.26% of the increments.
Then we can implement the estimators of Section (ref) with these choices for truncation function and level.
Arguably, there can be many other methods for deriving truncation rules, which however have to take into account the infinite dimensionality of the data and deal with the subtlety of functional outliers. The simulation study in Section (ref) shows the good performance of our method.
Presmoothing bond market data
It is rarely the case that term structure data are observed in the same resolution in time as in the maturity dimension. For bond market data, points on the discount curve $x\mapsto P_{i\Delta_n}(x)$ are observed irregularly with a lower resolution than daily along the maturity dimension. Additionally, information on the discount curve is sometimes latent as bonds are often coupon-bearing and assumed to be corrupted by market microstructure noise.
To account for these difficulties and in accordance with the classical “smoothing first, then estimation" procedure for functional data analysis advocated in Ramsay2005
we pursue the simple yet effective approach of presmoothing the data. We derive smoothed yield or discount curves, as described, for instance, in FPY2022, liu2021 or Linton2001, which allows us to derive approximate zero coupon bond prices for any desired maturity and for which the impact of market microstructure noise is mitigated.
While some theoretical guarantees in terms of asymptotic equivalence of discrete and noisy to perfect curve observation schemes could be derived for certain smoothing techniques and the task of estimating means and covariances of i.i.d. functional data (c.f. Zhang2007), in our case they would depend on the respective smoothing technique, the volatility and the semigroup as well as the magnitude of distorting market microstructure noise. A detailed theoretical analysis in that regard is beyond the scope of this article and instead, we showcase the robustness of our approach in the context of sparse, irregular and noisy bond price data within a simulation study in Section (ref).
Simulation study
In our simulation study we examine the performance of the truncated estimator $\mathcal T_{\hat q^{n,-}_{\cdot}}$ defined in (ref) as a measure of the continuous part $[X^C,X^C]$ of the quadratic covariation of the latent driver. As an important application of our theory is the identification of the number of statistically relevant drivers, we also examine how reliable the estimator can be used to determine the effective dimensions of $X^C$.
In this context, we also want to assess the robustness of our estimator concerning three important aspects: First, we need to confirm the robustness of the truncated estimator to jumps. Second, we study the effect of the common practice of presmoothing sparse, noisy, and irregular bond price data on the estimator's performance. Moreover, we examine how the routine of projecting these data onto a small finite set of linear factors (c.f. for instance Litterman1991 or the survey piazzesi2010), influences conclusions on the quadratic covariation.
For that,
we simulate log bond prices for some sampling size $m\leq 1000$, and $n=100$ time points, that is,
$$ P_{i,l}:=\log P_{i\Delta_n}(j_{i,l}\Delta_n)+\epsilon_{i,l}, \qquad i=1,...,100,\quad l=1,...,m,$$
where $(\epsilon_{i,j})_{i,j=1,...,100}$ are i.i.d centered Gaussian errors with variance $\sigma_{\epsilon }>0$ and the $j_{i,r}$ are drawn randomly from $\{0,...,1000\}$ without replacement for $r=1,...,m$.
We distinguish two cases: First, as a benchmark, we observe the data densely and without noise such that $\sigma_{\epsilon}=0$ and $m=1000$ and, second, we observe the data with noise $\sigma_{\epsilon}=0.01$ and sparsely with $m=100$.
In this case, the prices for all maturities $j\Delta_n, j=1,...,1000$ are recovered by quintic spline smoothing.
The roughness penalty for the smoothing splines is for each date chosen by a Bayesian information criterion and implemented via the ss-function from the npreg package in R.
Using quintic splines and a Bayesian information criterion induces smooth implied forward curves.
To analyze the impact of the customary procedure
of projecting the bond price data onto a low-dimensional linear subspace, we conduct our experiments in two scenarios. Scenario 1 in which we do not project the log bond prices and Scenario 2 in which
we project the log bond prices onto the first three eigenvalues of their covariance $c_{logbond}\equiv\frac 1{100}\sum_{i=2}^{100} (P_{i,\cdot}-P_{i-1,\cdot})(P_{i,\cdot}-P_{i-1,\cdot})'$ before we calculate $\hat q^{n,-}_t$. Indeed, as usual for bond market data, the first three eigenvectors of $c_{logbond}$ explain over 99% the variation in the log bond prices.
The log bond prices are derived from simulated instantaneous forward rates
$F_{i,j}:=\langle \ensuremath{\mathbb{I}}_{[(j-1)\Delta_n, j\Delta_n]}, f_{i\Delta_n}\rangle_{L^2(\mathbb R_+)}$
for $n=100$, $i=1,...,100$ and $j=1,...,1000$
from a forward rate process driven by a semimartingale $X$.
Precisely, we define
$X_t= \int_0^t \sqrt{\Sigma_s}dW_s+ J_t$ where
align*[align* omitted — 123 chars of source]
Here $x$ is a univariate mean-reverting square root process
$$dx(t)=1.5\left(0.058- x(t)\right)dt+ 0.05 \sqrt{x(t)}d\beta(t), \quad t\geq 0,\,\,x(0)=0.058$$
and $Q_{a}$ is a covariance operator on $L^2(\mathbb R_+)$ such that the corresponding covariance kernel $q_a$ (s.t. $Q_a=\mathcal T_{q_a}$) restricted to $[0,M]$ is a Gaussian covariance kernel $q_{a}(x,y) \propto \exp(-a(x-y)^2)$ for some $a>0$ and $\|q_{a}\big|_{[0,M]^2}\|_{L^2([0,M]^2)}=1$. The jumps are specified by two Poisson processes $N^1,N^2$ with intensities $\lambda_1, \lambda_2>0$ and jump distributions $\chi_i^2 \sim N(0,\rho_2 Q_{0.01})$ and $\chi_i^1 \sim N(0,\rho_1K)$ for $\rho_1,\rho_2 \geq 0$ and where $K$ is another covariance operator with kernel $k(x,y)\propto e^{-(x+y)}$ and $\|k\big|_{[0,M]^2}\|=1$.
We specify the corresponding parameters of this infinite-dimensional model as follows: We choose $a=50$ reflecting a high dimensional setting since the decay rate of the eigenvalues of $Q_a$ is slow (10 eigencomponents are needed to explain 99% of the variation of $X^C$).
The mean reversion level $0.058$ of the square root process corresponds to the Hilbert-Schmidt norm of the long-time estimator of volatility derived from bond market data as discussed in the next section.
Jumps corresponding to the first component are considered large and rare outliers reflected by a high $\rho_1=0.0116$ and low $\lambda_1=1$. The second component describes outliers which are more frequent and smaller in norm reflected by a lower $\rho_2=0.0029$ and higher $\lambda_2=4$ but correspond to changes of the shape of the forward curves. Both jump processes are chosen such that their Hilbert-Schmidt norm accounts for approximately 10 % of the quadratic variation.
We also consider cases in which no jumps are present (corresponding to the parameter choices $\lambda_1=\lambda_2=0$).
In each considered scenario, we compute the estimator $\Delta_n^{-2}\hat q_1^{n,10,-}$ (as defined in Remark (ref)) for $q_a\cdot \int_0^1 x(s)ds$.
In the cases in which jumps are present,
we consider the truncated estimator via the data-driven truncation rule of Section (ref) with different values of $l=3,4,5$ for the truncation level $u_n=l\sqrt{d+1}\Delta_n^{0.49}$. Here $d$ is chosen as the smallest value such that the first $d$ eigencomponents account for 90% of the variation as measured by the preliminary covariance estimator for which we truncate at the 0.75-quantile of the sequence of difference return
curves as measured in their $l^2$ norm. Models M1 and M2 for which no jumps are present serve as benchmarks for the truncated estimators and no truncation is conducted ($l=\infty$).
We assess the performance of the respective estimator in the context of two criteria of which each reflects an important application of our estimator. First, we measure the relative approximation error $rE(\Delta_n^{-2} \hat q_1^{n,10,-}, IV)$ where
for $x\in [(j_1-1)\Delta_n,j_1\Delta_n] $ and $y\in [(j_2-1)\Delta_n,j_2\Delta_n]$ the $\Delta_n$-resolution of the integrated volatility is $IV(x,y):=\int_0^1x(s)ds \cdot n^2\int_{(j_1-1)\Delta_n}^{j_1\Delta_n}\int_{(j_2-1)\Delta_n}^{j_2\Delta_n}q_{a}(z_1,z_2) dz_1,dz_2$ and
$$rE(k_1,k_2)
:
\frac{\| k_1- k_2 \|_{L^2([0,10]^2)}}{\| k_2 \|_{L^2([0,10]^2)}}\qquad k_1,k_2\in L^2([0,M]^2).$$
Second, we will investigate how reliably the estimator can be used to determine the number of factors needed to explain certain amounts of variation of the latent driving semimartingale. For that, we define
equation[equation omitted — 205 chars of source]
for a symmetric positive nuclear operator $ C$ and an orthonormal basis $e=(e_i)_{i\geq 0}\subset L^2([0,M])$. Let $\hat e=(\hat e_i)_{i\in \mathbb N}$ denote the eigenfunctions of $\mathcal T_{\hat q_1^{n,10,-}}$ ordered by the magnitude of the respective eigenvalues. We report the numbers $D_C^e$ for
$C=\mathcal T_{\hat q_1^{n,10,-}}$, $e=\hat e$ and $p=0.85$, $0.90$, $0.95$ and $0.99$, which are the numbers of factors needed to explain respectively $85\%$, $90\%$, $95\%$ and $99\%$ of the variation.
Table (ref) shows the results of the simulation study based on $500$ Monte-Carlo iterations.
For Scenario S1, reflecting our proposed fully infinite-dimensional estimation procedure, the log bond prices are not projected onto a finite-dimensional subspace a priori.
In this scenario, at least if jumps are truncated at a low level ($l=3$), the medians of relative errors are of a comparable magnitude when using either nonparametrically smoothed data (M2,M4) or perfect observations (M1,M3). While some jumps are overlooked by the truncation rule,
the medians of the relative errors in the cases with jumps (M3,M4) just moderately increased compared to the respective cases in which no jumps appeared (M1,M2), at least for a low truncation level.
The reported dimensions needed to explain the various levels of explained variation are estimated quite reliably (observe that the true thresholds
are respectively $5,6,7$ and $10$). In the noisy and irregular settings (M2, M4) the estimators tend to add a dimension compared to the perfectly observed settings in the median but have low interquartile ranges, which contain the correct dimension.
We conclude that the measurement of dimensions can be conducted accurately under realistic conditions.
Comparing scenarios S1 and S2 it becomes evident that the customary finite-dimensional projections of log bond prices (S2) affected the estimator's performance significantly.
All of the medians of relative errors are significantly higher compared to the case in which no projection was conducted, while for the practically important case in which data were smoothed from irregular sparse and noisy observations ($M4$) and jumps were truncated at level $l=3$, the error more than doubled.
For all considered thresholds of explained variation ($85\%$, $90\%$, $95\%$ and $99\%$) the reported dimension is constantly $3$, where we just reported the results for the $99\%$ threshold in the table. This is not surprising, since we started from a three-factor model for the log bond prices, but it demonstrates, that the common practice of projecting price or yield curves onto a few linear factors can disguise statistically important information, despite their high explanatory power for the variation of log bond prices, which is in line with Crump2022.
table[table omitted — 2,485 chars of source]
Empirical analysis of bond market data
In this section, we apply our theory to bond market data. In particular, we investigate the influence of jumps on the estimators and the dimensions of the integrated volatility, that is, the continuous part of the quadratic covariation, to determine how many random drivers are statistically relevant.
We consider nonparametrically smoothed yield curve data from FPY2022. For constructing smooth curves on each day, the authors of FPY2022 use a kernel ridge-regression approach based on the theory of reproducing kernel Hilbert spaces.
We measure time in years and the data are available for approximately $250$ trading days in each year, yielding $n\approx 250$, using a day count convention in trading days, with a daily resolution in the maturity direction, where we consider a maximal time of $M=10$ years to maturity. The data are given as yields, which we first transform to zero coupon bond prices and then derive the difference returns $d_i^n(j)$ for $i=1,...,\ensuremath{\lfloor T/\Delta_n\rfloor}$ and $j=1,...,\lfloor M/\Delta_n\rfloor$ by formula (ref).
We consider data from the first trading day of the year $1990$ ($i=1$) to the last trading day of the year $2022$ ($i=33$).
We then derive the estimators $\hat q_{i}^n\big|_{[0,10]},$ and $\hat q_{i}^{n,10,-}$ (as defined in Remark (ref)) for $i=1,...,33$ and derive the yearwise covariation kernels
$$\hat q_{i}^{*,-}:=\Delta_n^{-2}(\hat q_{i}^{n,10,-}-\hat q_{i-1}^{n,10,-})\quad \text{ and }\quad \hat q_{i}^*:=\Delta_n^{-2}(\hat q_{i}^n\big|_{[0,10]}-\hat q_{i-1}^n\big|_{[0,10]})$$ for $i=1,2,...,33$ and $q_{0}^n=0$.
The data-driven truncation rule described in Section (ref) is applied for a preliminary truncation at the $0.75$-quantile of the sequence of difference return
curves as measured in their $l^2$ norm and $d$ is chosen as the smallest value such that $d$ eigencomponents explain 90% of the variation of the preliminary estimator. Importantly, the truncation rule is conducted for different $l=3,4,5$ and for each year separately and only takes into account data within the respcetive year.
We also consider an estimator for a potential long-term volatility given by
$$\hat q_{long}^*:= \frac 1{33}\sum_{i=1}^{33} q_i^{*,-}.$$
Under the Assumption of Section (ref), this is an estimator for a stationary volatility kernel.
The results suggest that quadratic covariations in each year are rather complex in the sense that they exhibit a slow relative eigenvalue decay, unveil a varying shape and magnitude over time, and often differ quite substantially from measured quadratic covariations due to the existence of jumps.
Subsequently, we provide a thorough discussion of these observations.
A table containing all results of the analysis is contained in the appendix
Impact of jumps
table[table omitted — 1,332 chars of source]
On one hand, jumps that have a moderate impact on the magnitude of the overall quadratic covariation can visually distort the shape of the volatility. Figure (ref)
depicts plots of the graphs of the estimated truncated kernels $q_i^{*,-}$ (with the truncation level $l=3$) and nontruncated kernels $q^*_i$ for the years 2005, 2006 and 2007. In 2006, which is also a year in which the yield curve inverted before the financial crisis in 2007, two jumps
had a visible impact on the shape of the measured quadratic covariation kernels although they together accounted for less than 3 % of the magnitude of the quadratic covariation. This is due to a higher emphasis on the variation in difference returns with short maturities where one should note the different scalings in the plots.
Removing these two jumps leads to a more time-homogeneous shape of the integrated volatility surfaces in the sense of the relation of the variation in the shorter maturities to the variation in higher maturities.
On the other hand, jumps influence the magnitude of the quadratic variation. For instance,
nine increments in the year 2020 (Covid-19 outbreak) sorted out by the truncation rule for $l=3$ accounted for more than 50% of the overall quadratic variation in the data as measured by its norm. The statistics for jumps in all years can be found in the supplment to this article.
Interestingly, our measurements suggest that jumps tend to cluster.
figure[figure omitted — 942 chars of source]
Dimensionality of the continuous part of the quadratic variations
We now examine the statistically relevant number of random processes that are driving the continuous part of the forward curve dynamics by investigating the dimensionality of the continuous quadratic covariation of the latent driving semimartingale via the estimator $\mathcal T_{\hat q_i^{*,-}}$.
Table (ref) reports the number of eigenfunctions of $\mathcal T_{\hat q_i^{*,-}}$ that are needed to explain resp. $85\%$, $90\%$, $95\%$, and $99\%$ of the continuous covariation in the years 2005, 2006 and 2007 showing that to explain 99% of the variation at least 10 factors are needed in each year. The situation looks similar for all other years from 1990-2022, while the detailed results were relegated to the appendix.
We find that the complexity of the covariation seems to have decreased over the years, indicating a time-varying pattern of the volatility term structure that goes beyond its overall level.
It is noteworthy that in almost every year (30 out of 33), the number of linear factors needed to explain at least 99% of the truncated variation of the data is at least $10$.
These dimensions even increase if we employ the static factors $(e_j^{long})_{j\in \mathbb N}$ of eigenvectors of $\mathcal T_{\hat q_{long}^*}$ and do not update them in each year. In that case, in 27 out of 33 years at least 12 factors are needed to explain at least $99\%$ of the variation in each year.
Importance of higher-order factors for short term trading strategies
A natural question is if the higher-order factors
indicated by the analysis of real bond market data in Section (ref)
are of economic significance beyond capturing variation in difference returns. Therefore, we investigate whether the high dimensionality of the continuous quadratic varitations indicated by the estimators $\hat q_i^*$ and $q^*_{long}$ are important for other short term trading strategies than difference returns.
Precisely, Define the daily return $(d_L)_{i}^n(j)$ of the trading strategy of buying an $(j+L)\Delta_n$ bond and shorting an $j\Delta_n$ bond
$$(d_L)_{i}^n(j):= \sum_{l=0}^{L}d_i^n(j+l)=\tilde\Delta^n_{i\Delta_n} \log P((j+L)\Delta_n)-\tilde\Delta_{i\Delta_n}^n \log P(j\Delta_n)$$
where $\tilde\Delta_t^n \log P(x)=\log P_{t+\Delta_n}(x)-\log P_{t+\Delta_n}(x-\Delta_n)$. Evidently, we can derive them as linear functionals of either log bond price returns or difference returns.
We want to
determine the adequacy of approximation of these higher-order difference returns when they are derived either from approximated log price curves, which are projected onto its leading principal components or when they are derived from difference returns, which are projected onto the leading eigencomponents of the long-term volatility estimator.
For that, we calculate the
relative mean absolute error ($RMAE$) for a set $\mathcal V=\{i_1 \Delta_n^{val},...,i_{825}^{val}\Delta_n\}$ of dates (where in each year from 1990 to 2022 we randomly draw 25 dates making a total of 825 validation dates). That is, defining the piecewise constant kernels $(\tilde d_L)_{i}^n=\sum_{j=1}^{\lfloor M/\Delta_n\rfloor}(d_L)_i(j)\ensuremath{\mathbb{I}}_{[j\Delta_n,j\Delta_n)}$
we calculate
$$RMAE_L(f_1,...,f_d):=\frac 1{825}\sum_{l=1}^{825}\frac{\left\|(\tilde d_L)_{i_l}^n-\mathcal P_{f_1,...,f_d}(\tilde d_L)_{i_l}^n \right\|_{L^2(0,10)}}{\left\|(\tilde d_L)_{i_l}^n\right\|_{L^2(0,10)}}.
$$
where $\mathcal P_{f_1,...,f_d}:= \sum_{i=1}^d f_i^{\otimes 2}$.
The factors are derived in two different ways. In the first scenario (S1), the factors $f_1,...,f_d$ correspond to principal components of the empirical covariance of log-price differences $\tilde \Delta \log P_{i\Delta_n}$ for $i\notin \mathcal V$ and in the second scenario (S2) the factors $f_1,...,f_d$ correspond to the leading eigenfunctions of the estimated stationary volatility kernel $\hat q_{long}^*$
where as before the truncation of jumps is conductcted yearwise with truncation level $l=3$ according to the truncation procedure described in Section (ref).
We compare the results for lags of 7, 30, 90 and 180 days, since they approximately correspond to the returns of buying a bond and shorting another bond with time to maturity that is resp. a week, a month, a quarter or half a year higher. The $RMAE$s can be found in Table (ref).
table[table omitted — 2,010 chars of source]
It can be observed that a high number of factors is needed to approximate the lagged difference returns precisely and that approximations based on low factor structures as indicated by the covariance of difference returns imply high approximation errors. While it is not surprising that the approximation gets better if we use more factors, the high discrepancy of the approximation errors is noteworthy. The errors for a typically chosen three factor model based on log price differences (the factors correspond to level, slope and curvature), which explain more than 99,7% of the variation in log-price returns is for all lags higher than $0.4$, whereas for the approximation error for 14 factors, which we would need to explain 99% of the variation in difference returns as measured by $\hat q_{long}^*$ is never higher than $0.11$.
Interestingly, for all lags, choosing the factors equal to the leading eigenfunctions of the long term volatility $\hat q_{long}^*$ instead of the ones indicated by log-price differences can reduce the error for the higher-order approximations quite significantly and for $d=16$ and for lags not higher than 90 days by almost 50 %.
Higher-order factors of volatility can, thus, not easily be ignored and might carry important economic information.
Concluding remarks on the empirical study
We conclude that the reported dimensions are overall quite high compared to the few factors needed to explain a large amount of the variation in yield and discount curves. This suggests that low-dimensional factor models are not able to capture all statistically relevant codependencies of bond prices.
Still, exact magnitudes of explained variations of the higher order components have to be interpreted cautiously and conditional on the smoothing technique that was employed to derive yield or discount curves.
However, higher-order factors seem to be economically relevant for capturing variations in short term trading strategies as indicated by
the out-of-sample study of Section (ref).
Underestimation of the number of statistically relevant random drivers can have undesirable effects. For instance, Crump2022 showcase the potential economic impact on mean-variance optimal portfolio choices and hedging errors.
At the same time, not every model that is parsimonious in its parameters needs to
entail a low-dimensional factor structure such as the simple volatility model of Section (ref).
It seems desirable to derive parsimonious models that match the empirical observation of high or infinite-dimensional covariations and reflect the characteristics of their dynamic evolution.
Acknowledgements
I would like to thank Dominik Liebl, Fred Espen Benth, Alois Kneip and Andreas Petersson
for helpful comments and discussions. Funding by the
Argelander program of the University of Bonn is gratefully acknowledged.
thebibliography{10}
\bibitem{BNBV2013}
O. E. Barndorff-Nielsen, Fred Espen Benth, and Almut E. D. Veraart.
\newblock {Modelling energy spot prices by volatility modulated L{\'e}vy-driven
Volterra processes}.
\newblock {\em Bernoulli}, 19(3):803 -- 845, 2013.
\bibitem{BENNEDSEN2017}
M. Bennedsen.
\newblock A rough multi-factor model of electricity spot prices.
\newblock {\em Energy Econ.}, 63:301--313, 2017.
\bibitem{Petersson2022}
F. Benth, G. Lord, and A. Petersson.
\newblock The heat modulated infinite dimensional {H}eston model and its
numerical approximation.
\newblock {\em Available at ArXiv:2206.10166}, 2022.
\bibitem{BenthRudigerSuss2018}
Fred Espen Benth, Barbara R{\"u}diger, and Andre S{\"u}ss.
\newblock Ornstein--{U}hlenbeck processes in {H}ilbert space with
non-{G}aussian stochastic volatility.
\newblock {\em Stoch. Proc. Applic.}, 128(2):461--486, 2018.
\bibitem{BSV2022}
Fred Espen Benth, Dennis Schroers, and A. E. D. Veraart.
\newblock A feasible central limit theorem for realised covariation of spdes in
the context of functional data.
\newblock {\em Ann. Appl. Probab.}, 34(2):2208--2242, 2024.
\bibitem{Benth2022}
Fred Espen Benth, Dennis Schroers, and Almut E.D. Veraart.
\newblock A weak law of large numbers for realised covariation in a {H}ilbert
space setting.
\newblock {\em Stoch. Proc. Applic.}, 145:241--268, 2022.
\bibitem{BenthSgarra2021}
Fred Espen Benth and Carlo Sgarra.
\newblock A {B}arndorff-{N}ielsen and {S}hephard model with leverage in
{H}ilbert space for commodity forward markets.
\newblock {\em Available at SSRN 3835053}, 2021.
\bibitem{BenthSimonsen2018}
Fred Espen Benth and Iben Cathrine Simonsen.
\newblock The {H}eston stochastic volatility model in {H}ilbert space.
\newblock {\em Stoch. Analysis Applic.}, 36(4):733--750, 2018.
\bibitem{bjork1999}
T. Bj{\"o}rk and B. Christensen.
\newblock Interest rate dynamics and consistent forward rate curves.
\newblock {\em Math. Finance}, 9(4):323--348, 1999.
\bibitem{bjork1997}
T. Bj{\"o}rk, G. Di Masi, Y. Kabanov, and W. Runggaldier.
\newblock Towards a general theory of bond markets.
\newblock {\em Finance Stoch.}, 1:141--174, 1997.
\bibitem{bjork2001}
T. Bj{\"o}rk and L. Svensson.
\newblock On the existence of finite-dimensional realizations for nonlinear
forward rate models.
\newblock {\em Math. Finance}, 11(2):205--243, 2001.
\bibitem{cox2023}
S. Cox, C. Cuchiero, and A. Khedher.
\newblock Infinite-dimensional {W}ishart-processes.
\newblock {\em arXiv preprint arXiv:2304.03490}, 2023.
\bibitem{Cox2021}
S. Cox, S. Karbach, and A. Khedher.
\newblock An infinite-dimensional affine stochastic volatility model.
\newblock {\em Math. Finance}, 2021.
\bibitem{Cox2020}
S. Cox, S. Karbach, and A. Khedher.
\newblock Affine pure-jump processes on positive {H}ilbert-{S}chmidt operators.
\newblock {\em Stoch. Proc. Applic.}, 151:191--229, 2022.
\bibitem{Crump2022}
R. K. Crump and N. Gospodinov.
\newblock On the factor structure of bond returns.
\newblock {\em Econometrica}, 90(1):295--314, 2022.
\bibitem{DPZ2014}
Giuseppe Da Prato and Jerzy Zabczyk.
\newblock {\em Stochastic Equations in Infinite Dimensions}, volume 152 of {\em
Encyclopedia of Mathematics and its Applications}.
\newblock Cambridge University Press, Cambridge, second edition, 2014.
\bibitem{Duffie2003}
D. Duffie, D. Filipovi{\'c}, and W. Schachermayer.
\newblock {Affine processes and applications in finance}.
\newblock {\em Ann. Appl. Probab.}, 13(3):984 -- 1053, 2003.
\bibitem{FFRS2020}
B. F{\'a}rkas, M. Friesen, B. R{\"u}diger, and D. Schroers.
\newblock On a class of stochastic partial differential equations with multiple
invariant measures.
\newblock {\em NoDEA}, 28, 2020.
\bibitem{filipovic1999}
D. Filipovi{\'c}.
\newblock A note on the {N}elson--{S}iegel family.
\newblock {\em Math. finance}, 9(4):349--359, 1999.
\bibitem{filipovic2000b}
D. Filipovi{\'c}.
\newblock Exponential-polynomial families and the term structure of interest
rates.
\newblock {\em Bernoulli}, pages 1081--1107, 2000.
\bibitem{Filipovic2000}
D. Filipovi{\'c}.
\newblock {\em Consistency Problems for HJM Interest Rate Models}, volume 1760
of {\em Lecture Notes in Mathematics}.
\newblock Springer, Berlin, 2001.
\bibitem{FPY2022}
D. Filipovi{\'c}, M. Pelger, and Y. Ye.
\newblock Stripping the discount curve - a robust machine learning approach.
\newblock {\em Swiss Finance Institute Research Paper}, 2022.
\bibitem{FilipovicTappeTeichmann2010b}
D. Filipovi{\'c}, S. Tappe, and J. Teichmann.
\newblock Jump-diffusions in {H}ilbert spaces: existence, stability and
numerics.
\newblock {\em Stochastics}, 82(5):475--520, 2010.
\bibitem{filipovic2003}
D. Filipovi{\'c} and J. Teichmann.
\newblock Existence of invariant manifolds for stochastic equations in infinite
dimension.
\newblock {\em J. Funct. Anal.}, 197(2):398--432, 2003.
\bibitem{friesen2022}
M. Friesen and S. Karbach.
\newblock Stationary covariance regime for affine stochastic covariance models
in hilbert spaces.
\newblock {\em arXiv preprint arXiv:2203.14750}, 2022.
\bibitem{HJMoriginal}
D. Heath, R. Jarrow, and A. Morton.
\newblock Bond pricing and the term structure of interest rates: A new
methodology for contingent claims valuation.
\newblock {\em Econometrica}, 60(1):77--105, 1992.
\bibitem{Jacod2008}
J. Jacod.
\newblock Asymptotic properties of realized power variations and related
functionals of semimartingales.
\newblock {\em Stoch. Proc. Applic.}, 118(4):517--559, 2008.
\bibitem{JacodProtter2012}
J. Jacod and P. Protter.
\newblock {\em Discretization of Processes}, volume 67 of {\em Stochastic
Modelling and Applied Probability}.
\newblock Springer, Heidelberg, 2012.
\bibitem{Linton2001}
O. Linton, E. Mammen, J.P. Nielsen, and C. Tanggaard.
\newblock Yield curve estimation by kernel smoothing methods.
\newblock {\em J. Econ.}, 105(1):185--223, 2001.
\bibitem{Litterman1991}
R. Litterman and J. Scheinkman.
\newblock Common factors affecting bond returns.
\newblock {\em J. Fixed Income}, 1:62--74, 1991.
\bibitem{liu2021}
Y. Liu and C. Wu.
\newblock Reconstructing the yield curve.
\newblock {\em J. financ. econ.}, 142(3):1395--1425, 2021.
\bibitem{Mancini2001}
C. Mancini.
\newblock Disentangling the jumps of the diffusion in a geometric jumping
brownian motion.
\newblock {\em Giornale dell’Istituto Italiano degli Attuari}, 64:19--47,
2001.
\bibitem{Mancini2004}
C. Mancini.
\newblock Estimating the integrated volatility in stochastic volatility models
with l{\'e}vy type jumps.
\newblock {\em Tech. rep., Universit{\`a} di Firenze}, 2004.
\bibitem{Mancini2009}
C. Mancini.
\newblock Nonparametric threshold estimation for models with stochastic
diffusion coefficient and jumps.
\newblock {\em Scand. J. Statist.}, 36:270--296, 2009.
\bibitem{Mancini2012}
C. Mancini and F. Calvori.
\newblock {\em Jumps}, chapter Seventeen, pages 403--445.
\newblock John Wiley & Sons, Ltd, 2012.
\bibitem{mandrekar2015}
V. Mandrekar and B. R{\"u}diger.
\newblock {\em {Stochastic integration in Banach spaces}}, volume 73 of {\em
Probability Theory and Stochastic Modelling}.
\newblock Springer, 2015.
\bibitem{marinelli2010}
C. Marinelli.
\newblock Well-posedness and invariant measures for hjm models with
deterministic volatility and l{\'e}vy noise.
\newblock {\em Quant. Finance}, 10(1):39--47, 2010.
\bibitem{Panaretos2019}
V. Masarotto, V. M. Panaretos, and Y. Zemel.
\newblock Procrustes metrics on covariance operators and optimal transportation
of gaussian processes.
\newblock {\em Sankhya A}, 81(1):172--213, 2019.
\bibitem{PZ2007}
S. Peszat and J. Zabczyk.
\newblock {\em Stochastic Partial Differential Equations with {L}\'{e}vy
Noise}, volume 113 of {\em Encyclopedia of Mathematics and its Applications}.
\newblock Cambridge University Press, Cambridge, 2007.
\bibitem{piazzesi2010}
M. Piazzesi.
\newblock Affine term structure models.
\newblock In {\em Handbook of financial econometrics: Tools and Techniques},
pages 691--766. Elsevier, 2010.
\bibitem{Ramsay2005}
James Ramsay and B. W. Silverman.
\newblock {\em Functional Data Analysis}.
\newblock Springer Series in Statistics. Springer, second edition, 2005.
\bibitem{ren2017}
H. Ren, N. Chen, and C. Zou.
\newblock Projection-based outlier detection in functional data.
\newblock {\em Biometrika}, 104(2):411--423, 2017.
\bibitem{rusinek2010}
A. Rusinek.
\newblock Mean reversion for hjmm forward rate models.
\newblock {\em Adv. in Appl. Probab.}, 42(2):371--391, 2010.
\bibitem{Schroers2024}
D. Schroers.
\newblock Robust functional data analysis for stochastic evolution equations in
infinite dimensions.
\newblock {\em arXiv preprint}, 2024.
\bibitem{Tehranchi2005}
M. Tehranchi.
\newblock {{A Note on Invariant Measures for {HJM} Models}}.
\newblock {\em Finance Stoch.}, 9(3):389--398, 2005.
\bibitem{vargiolu1999}
T. Vargiolu.
\newblock Invariant measures for the {M}usiela equation with deterministic
diffusion term.
\newblock {\em Finance Stoch.}, 3(4):483--492, 1999.
\bibitem{Zhang2007}
J. Zhang and J. Chen.
\newblock {Statistical inferences for functional data}.
\newblock {\em Ann. Statist.}, 35(3):1052 -- 1079, 2007.
appendix\section{It{\^o} semimartingales in Hilbert spaces}
In this appendix, we provide an introduction and technical details for the class of It{\^o} semimartingales that we consider throughout the paper.
First, we specify the components of the driver $X$ which is an $L^2(\mathbb R_+)$-valued right-continuous process with left-limits (c{\`a}dl{\`a}g) that can be decomposed as
$$X_t:= X_t^C+J_t:=(A_t+M_t^C)+J_t\quad t\geq 0.
$$
Here, $A$ is a continuous process of finite variation, $M^C$ is a continuous martingale and $J$ is another martingale modeling the jumps of $X$. We assume that $X$ is an It{\^o} semimartingale for which the components have integral representations
\begin{equation}
A_t:=\int_0^t \alpha_s ds,\quad M_t^C:=\int_0^t\sigma_s dW_s, \quad J_t:=\int_0^t\int_{H\setminus \{0\}}\gamma_s(z) (N-\nu)(dz,ds).
\end{equation}
For the first part $(\alpha_t)_{t\geq 0}$ is an $H$-valued and and almost surely integrable (w.r.t. $\|\cdot\|_{L^2(\mathbb R_+)}$) process that is adapted to the filtration $(\mathcal F_t)_{t\geq 0}$.
The volatility process $(\sigma_t)_{t\geq 0}$ is predictable and takes values in the space of Hilbert-Schmidt operators $L_{\text{HS}}(U,L^2(\mathbb R_+))$ from a separable Hilbert space $U$ into $L^2(\mathbb R_+)$. Moreover, we have $\mathbb P[\int_0^T\|\sigma_s\|_{\text{HS}}^2ds<\infty]=1$. The space $U$ is left unspecified, as it is just formally the space on which the Wiener process $W$ is defined and does not affect the distribution of $X$. The cylindrical Wiener process $W$ is a weakly defined Gaussian process with independent stationary increments and covariance $I_U$, the identity on $U$. One might consult the standard textbooks DPZ2014, mandrekar2015 or PZ2007 for the integration theory w.r.t. $W$.
For the jump process $J$, we define a homogeneous Poisson random measure $N$ on $\mathbb R_+\times H\setminus \{0\}$ and its compensator measure $\nu$ which is of the form $\nu(dz,dt)=F(dz)\otimes dt$ for a $\sigma$-finite measure $F$ on $\mathcal B(H\setminus \{0\})$. The process $\gamma_s(z))_{s\geq 0, z\in H\setminus \{ 0\}}$ is the $l^2(\mathbb R_+)$-valued jump volatility process and is predictable and stochastically integrable w.r.t. the compensated Poisson random measure $\tilde N:=(N-\nu)$. For a detailed account on stochastic integration w.r.t. compensated Poisson random measures in Hilbert spaces, we refer to mandrekar2015 or PZ2007.
Let us now rewrite the quadratic covariation (ref) of $X$ in terms of the volatility $\sigma$ and the jumps of the process as
\begin{align}
[X,X]_t= [X^C,X^C]_t+[J,J]_t= \int_0^t \Sigma_s ds+ \sum_{s\leq t} (X_s-X_{s-})^{\otimes 2},
\end{align}
where $\Sigma_s = \sigma_s\sigma_s^*$ (where $\sigma_s^*$ is the Hilbert space adjoint) and
$X_{t-}:= \lim_{s\uparrow t}X_s$ is the left limit of $(X_t)_{t\geq 0}$ at $t$, which is well-defined, since $X_t$ has c{\`a}dl{\`a}g paths. This characterization follows as a special case of Theorem 3.1 in Schroers2024
Let us now reconsider Example (ref).
\begin{example}[Rewriting an $L^2(\mathbb R_+)$-valued Poisson random measure in compensated form]
In Example (ref) it was remarked that a compound Poisson process $J_t=\sum_{i=1}^{N_t} \chi_i$ is strictly speaking not a valid choice for the jump process, since it is not a martingale. Here we show that the semimartingale in the example can be easily rewritten to have the desired form:
For that, define the Poisson random measure
$N(B,[0,t]):= \#\{i\leq N_t: \chi_i\in B\} $ for $B \in \mathcal B(H\setminus \{0\}), t\geq 0.$
This has compensator measure $\nu= \lambda dt\otimes F(dz)$, so we can redefine $J$ in a formally correct manner by
$J_t=\sum_{i=1}^{N_t} \chi_i-\lambda t \mathbb E[\chi_1]=\int_0^t \int_{L^2(\mathbb R_+) \setminus \{0\}} z (N-\nu)(dz,ds)$ and set $A_t=(a+\lambda \mathbb E[X_1])t$.
\end{example}
\section{Technical Assumptions}
This section contains the technical Assumptions that are needed for the validity of Theorems (ref), (ref), (ref), (ref) and (ref).
\subsection{Assumption for derivation of idenifiability of $[X^C,X^C]$ and $[J,J]$ }
To derive asymptotic results for $\hat q^{n,-}_t$ in Theorem (ref), we introduce
\begin{assumption}[r]
$\alpha$ is locally bounded, $\sigma$ is c{\`a}dl{\`a}g and there is a localizing sequence of stopping times $(\tau_n)_{n\in \mathbb N}$ and for each $n\in \mathbb N$ a real valued function $\Gamma_n:H\setminus \{0\}\to \mathbb R$ such that $\|\gamma_t(z)(\omega)\|\wedge 1\leq \Gamma_n(z)$ whenever $t\leq \tau_n(\omega)$ and $\int_{L^2(\mathbb R_+)\setminus \{0\}}\Gamma_n(z)^rF(dz)<\infty$.
\end{assumption}
Assumption (ref) used in Theorem (ref) is a direct generalization of Assumption (H-r) in JacodProtter2012. It implies that for $r<2$, the jumps of the process are $r$-summable, that is, we have
$$\sum_{s\leq t}\|X_s-X_{s-}\|^{l}<\infty\quad \forall l>r.$$
\subsection{Assumption for derivation of convergence rates}
For the derivation of convergence rates in Theorem (ref), observe that, since $\Sigma_t=\sigma_t\sigma_t^*$ is for each $t\geq 0$ a Hilbert-Schmidt operator, we can find a process of kernels
\begin{equation}
q_t^C, such that \Sigma_t= \mathcal T_{q_t^C}\quad \forall t\geq 0.
\end{equation}
It is seems natural to impose H{\"o}lder-regularity assumptions on the volatility kernel $q_t^C$ for $t\geq 0$ to derive the error bounds.
For instance, one might consider a H{\"o}lder continuous volatility kernel, such that $q_t^C\in C^{\gamma}(\mathbb R_+^2)$ for
$$C^{\gamma}(\mathbb R_+^2):=\left\{q:\mathbb R_+^2\to \mathbb R:
\sup_{x,y,x',y'\leq M}\frac {|q(x,y)-q(x',y')|}{\|(x,y)-(x',y')\|_{\mathbb R^2}^{\gamma}}<\infty\quad \forall M\geq 0\right\}.$$
However, we can consider weaker regularity conditions, which do not necessarily assume the kernels to be continuous.
Namely, we require $q_t^C\in \mathfrak F_{\gamma}$ where
\begin{align*}
\mathfrak F_{\gamma}
:= & \left\{q \in L^2(\mathbb R_+^2): \|q\|^2_{\mathfrak F_{\gamma}(\mathbb R_+^2)}:=\sup_{r>0}\int_{\mathbb R_+^2} \frac{(q(r+x,y)-q(x,y))^2}{r^{2\gamma}}dxdy<\infty\right\}.
\end{align*}
The classes $\mathfrak F_{\gamma}$ might appear abstract but, in particular, it contains H{\"o}lder spaces, that is,
\begin{equation}
C^{\gamma}(\mathbb R_+^2)\subset \mathfrak F_{\gamma}.
\end{equation}
Vice versa, $ \mathfrak F_{\gamma}$ is not a subset of $ C^{\gamma}$ but it is strictly larger, allowing for discontinuities in volatility kernels:
Let $g(x,y):= \ensuremath{\mathbb{I}}_{[a,b]}(x)\ensuremath{\mathbb{I}}_{[a,b]}(y)$ for an interval $[a,b]\subset \mathbb R_+$. Then clearly, $g$ is not an element of $C^{\frac 12}(\mathbb R_+^2)$ as it is discontinuous. However, it is
$ \|g\|_{\mathfrak F_{1/2}}
= 2(b-a) <\infty.$
Hence $g \in \mathfrak F_{1/2}$, while $g\notin \mathfrak F_{\rho}$ for any $\rho>1/2$.
We now state our formal regularity assumption.
\begin{assumption}[$\gamma$]
Let $\gamma \in (0,1/2]$. We have $q^C_t\in \mathfrak F_{\gamma}$ $\mathbb P\otimes dt$-almost everywhere and
\begin{equation}
\mathbb P\left[\int_0^T \|q_s^C\|_{\mathfrak F_{\gamma}}ds<\infty\right]=1,\quad T>0.
\end{equation}
\end{assumption}
\begin{remark}
Regularity Assumption (ref) is sharp in Theorem (ref) in the sense that for every $\gamma'<\gamma$ we can always specify a squared volatility process $(\Sigma_t)_{t\geq 0}$ such that in probability $\Delta_n^{-\gamma}\sup_{t\in [0,T]}\left\|\mathcal T_{\hat q^{n,-}_t}-[X^C,X^C]\right\|_{\text{HS}}$
diverges but the process of kernels $(q_t^C)_{t\geq 0}$ fulfills Assumption (ref) for $\gamma'$ and (and not for $\gamma$)
(c.f. Example 3.6 in BSV2022).
\end{remark}
As a result of Theorem (ref) and (ref), we can derive rates of convergence also under H{\"o}lder regularity assumptions.
\begin{corollary}
If Assumption (ref)(r)
holds for some $r\in (0,2)$ and for all $t\geq 0$ it is $q_t\in C^{\gamma}(\mathbb R_+^2)$ $\mathbb P\otimes dt$-almost everywhere for some $\gamma\in (0,1/2]$, then (ref) holds for all $\rho<(2-r)w$ and (ref) holds if $r<2(1-\gamma)$ and $w\in [\gamma/(2-r),1/2]$.
\end{corollary}
For the central limit theorem, we further need
\begin{assumption}
It is almost surely
\begin{equation}\int_0^T\sup_{r\geq 0}\frac{\|(I-\mathcal S(r))\sigma_s\|_{op}^2}{r} ds <\infty,\quad T>0.
\end{equation}
\end{assumption}
\subsection{Assumptions for Long-time estimators}
We introduce
\begin{assumption}
The process $(\Sigma_t)_{t\geq 0}$ is mean stationary and mean ergodic, in the sense that $\mathbb E\left[\|\sigma_s\|^2_{L_{\text{HS}}(U,L^2(\mathbb R_+))}\right]<\infty$ and there is an operator $\mathcal C$ such that for all $t$ it is $\mathcal C=\mathbb E[ \Sigma_t]$ and as $T\to\infty$ we have in probability an w.r.t. the Hilbert-Schmidt norm that
\begin{equation}
\frac 1T \int_0^T\Sigma_s ds=\frac{[X^C,X^C]_T}T\to \mathcal C.
\end{equation}
\end{assumption}
Under Assumption (ref) we have that
$\mathbb E[ (M_t^C)^{\otimes 2}]=\mathbb E[ (\int_0^t \sigma_s dW_s)^{\otimes 2}] = t \mathcal C\quad \forall t\geq 0. $
Hence, $\mathcal C$ is the covariance of the driving continuous martingale $M^C$ (scaled by time).
Hence, as for regular functional principal component analyzes, we can find approximately a linearly optimal finite-dimensional approximation of the driving martingale, by projecting onto the eigencomponents of $\mathcal C$. Even more, $\mathcal C$ is the instantaneous covariance of the process $f$ in the sense that
$\mathcal C=\lim_{n\to\infty}\mathbb E[(f_{t+\Delta_n}-\mathcal S(\Delta_n)f_t)^{\otimes 2}]/\Delta_n.$
To estimate $\mathcal C$, we make use of a moment assumption for the coefficients.
\begin{assumption}[p,r]
For $p,r>0$ such that $\mathbb E\left[\|\gamma_s(z)\|^r\right]=\Gamma(z)$ independent of $s$ for all $s\geq 0$ and
there is a constant $A>0$ such that for all $s\geq 0$ it is
$$ \mathbb E\left[\|\alpha_s\|_{L^2(\mathbb R_+)}^p+\|\sigma_s \|_{\text{HS}}^p+\int_{L^2(\mathbb R_+)\setminus \{0\}}\|\gamma_s(z)\|^r\nu(dz)\right] \leq A.$$
\end{assumption}
Moreover, we also make an assumption on the regularity of the volatility.
\begin{assumption}[$\gamma$]
With the notation (ref) we have for
$\gamma\in (0,\frac 12]$ that
there is a constant $A>0$ such that for all $s\geq 0$ it is
$$ \mathbb E\left[\|q_s^C\|_{\mathfrak F_{\gamma}}\right] \leq A.$$
\end{assumption}
\section{Proofs of section (ref)}
\begin{proof}[Proof of the general nonsemimartingality of models in Example (ref)]
We need to prove that
\begin{align}
f_t=f_0 +\int_0^t k(\cdot+t-s) d\beta_s
\end{align}
is not a continuous semimartingale where $k\in L^2(\mathbb R_+)$, $\beta$ is a univariate standard Brownian motion.
Therefore, assume that $f_t$ defines a semimartingale in $L^2(\mathbb R_+)$ of the form
$f_t= A_t+M_t $
for an $H$-valued continuous martingale $M$ and a finite variation process $A$.
Observe that we also have that $f$ is a weak solution to the stochastic partial differential equation
$$\frac d{dx} f_t dt+ (e\otimes k) dW_t, \quad t\geq 0,$$
for a cylindrical Wiener process $W$ such that $\beta= \langle e, W\rangle$. Hence, for an orthonormal basis $(e_j)_{j\in \mathbb N}\subset D(d/dx)$ we find that
$$\langle f_t, e_j \rangle = \langle f_0,e_j\rangle +\int_0^t \langle f_s, \left(\frac d{dx}\right)^* e_j\rangle ds + \langle k,e_j\rangle \beta_t.$$
These are a one-dimensional semimartingales for which the first integral is of finite variation and the second part is of quadratic variation. As the decomposition of a continuous semimartingale into a continuous part with finite variation and a continuous martingale (which vanishes at $0$) with quadratic variation is unique up to $\mathbb P\otimes dt$ nullsets, we obtain that $\mathbb P\otimes dt$-almost everywhere
$$\langle A_t,e_j\rangle= \int_0^t \langle f_s ,\left(\frac d{dx}\right)^*e_j\rangle\qquad \langle M_t,e_j\rangle= \langle k, e_j\rangle \beta_t\quad \forall t\geq 0, j\in \mathbb N.$$
Therefore, we must have $M_t=\beta_t k$ and we must have $\sum_{i=1}^n \Delta_i^n f^{\otimes 2}=\sum_{i=1}^n \Delta_i^n M_t^{\otimes 2}\to k^{\otimes 2}$ in probability as $n\to \infty$.
Defining
$$S_t^n := \sqrt n\langle f_t, \ensuremath{\mathbb{I}}_{[0,\Delta_n]}\rangle,$$
we also obtain that in probability
$$\left|\sum_{i=1}^n (\Delta_i^n S^n)^2-n\langle k, \ensuremath{\mathbb{I}}_{[0,\Delta_n]}\rangle^2\right|\leq
\|\sum_{i=1}^n \Delta_i^n f^{\otimes 2}-k^{\otimes 2}\|_{L_{\text{HS}}(L^2(\mathbb R_+)}\to 0 $$
and since as $n\to \infty$ it is
$\sqrt n \langle k, \ensuremath{\mathbb{I}}_{[0,\Delta_n]}\rangle
=\Delta_n^{1/2+H}/(H+1)\to 0$
we also find that as $n\to\infty$ and in probability that
$\sum_{i=1}^n (\Delta_i^n S^n)^2\to 0$
must hold.
Moreover, we find, since the kernel $k$ is square integrable and $k(t)=t^H$ on $t\in [0,1]$ that by the Burkholder-Davis-Gundy inequality for $\epsilon>0$
\begin{align*}
& \mathbb E[(\Delta_i^n S^n)^{2+\epsilon}]^{\frac 2{2+\epsilon}} \\
\leq & 2\int_{(i-1)\Delta_n}^{i\Delta_n} \|k(i\Delta_n+\cdot-s)\|^2 ds\\
&\qquad+2n\int_0^{(i-1)\Delta_n}\langle k(i\Delta_n+\cdot-s) -k((i-1)\Delta_n+\cdot-s) ,\ensuremath{\mathbb{I}}_{[0,\Delta_n]}\rangle^2 ds\\
\leq &2 \|k\|^2\Delta_n
+2n\int_0^{(i-1)\Delta_n}\left(\int_0^{\Delta_n}(i\Delta_n+y-s)^H -((i-1)\Delta_n+y-s)^H dy\right)^2 ds\\
\leq &2 \|k\|^2\Delta_n
+2n\int_0^{(i-1)\Delta_n}\frac 4{(H+1)^2}\left((i\Delta_n-s)^{H+1}-((i-1)\Delta_n-s)^{H+1}\right)^2 ds.
\end{align*}
Now, using the mean value theorem and since $t^H$ is decreasing in $t$ we find
\begin{align*}
\mathbb E[(\Delta_i^n S^n)^{2+\epsilon}]^{\frac 2{2+\epsilon}}
\leq &
2\|k\|^2\Delta_n
+8 \Delta_n\int_0^{(i-1)\Delta_n} ((i-1)\Delta_n-s)^{2H} ds\\
\leq & 2\|k\|^2\Delta_n
+8\Delta_n\frac 1{2H+1}.
\end{align*}
This shows in particular, that by Jensen's inequality we have
\begin{align*}
\mathbb E\left[\left(\sum_{i=1}^n(\Delta_i^n S^n)^2\right)^{\frac {2+\epsilon}2}\right]\leq & n^{\frac {2+\epsilon}2-1}\sum_{i=1}^n\mathbb E[(\Delta_i^n S^n)^{2+\epsilon}]
\leq \left(\|k\|^2
+\frac 8{2H+1} \right)^{\frac {2+\epsilon}2},
\end{align*}
which shows that $\sum_{i=1}^n(\Delta_i^n S^n)^2$ is uniformly integrable. Thus, convergence in probability must imply convergence of the mean and we must have
$$ \mathbb E[\sum_{i=1}^n(\Delta_i^n S^n)^2]\to 0\quad \text{ as } n\to\infty.$$
However, we can show similarly to the calculations before that using the mean value theorem it is
\begin{align*}
\mathbb E\left[\sum_{i=1}^n(\Delta_i^n S^n)^2\right]
\geq & n\int_0^{(i-1)\Delta_n}\langle k(i\Delta_n+\cdot-s) -k((i-1)\Delta_n+\cdot-s),\ensuremath{\mathbb{I}}_{[0,\Delta_n]}\rangle^2 ds\\
\geq & 4 \Delta_n\int_0^{(i-1)\Delta_n}((1+i)\Delta_n-s)^{2H}) ds\\
= & \frac 4{2H+1} \Delta_n^{2H+2}((1+i)^{2H+1}-2^{2H+1}). \\
\end{align*}
Thus, writing $K= 4/(2H+1)H^2(H+1)^2$ we find
$$ \mathbb E[\sum_{i=1}^n(\Delta_i^n S^n)^2]\geq K \Delta_n^{2H+2} \sum_{i=1}^n(1+i)^{2H+1}-K \Delta_n^{2H+2} \sum_{i=1}^n 2^{2H+1}.$$
While the second term is $o(1)$, for the first term it is
$$K \Delta_n^{2H+2} \sum_{i=1}^n(1+i)^{2H+1}
\geq K \Delta_n^{2H+2} \int_0^{n+1} x^{2H+1} dx=\frac{K}{2H+2} \left(\frac {n+1}n\right)^{2H+2}\geq \frac{K}{2H+2}.$$
This cannot hold, since by the uniform integrability of the sequence $\sum_{i=1}^n(\Delta_i^n S^n)^2$ we have that the mean $ \mathbb E[\sum_{i=1}^n(\Delta_i^n S^n)^2]$ must converge to $0$.
\end{proof}
\section{Proofs of Section (ref)}
In this Section we prove the results of Section (ref). For that, we first prove an abstract limit theory for general evolution equations in Section (ref). We then derive the results of Section (ref) using this abstract result in Section (ref)
\subsection{An Abstract limit theorem}
The asymptotic theory elaborated in the article follows by an abstract result for abstract evolution equations in Hilbert spaces, which we present and prove in this section. Roughly speaking, we prove that the results in Schroers2024 are valid, also when we discretized the functional data also in the cross-section in a particular manner, that we will make precise next. For now let $f$ be a mild solution to a stochastic evolution equation of the for described in (ref).
We also introduce the notation
$$H:=L^2(\mathbb R_+).$$
We do this, because the subsequent Theorem (ref) holds under much more general conditions than for the term structure setting and with this notation it becomes simple to appreciate this generality. That is, Theorem (ref) holds for general separable Hilbert spaces $H$, semigroups $\mathcal S$ and general $H$-valued It{\^o} semimartingale as described in Schroers2024. To be consistent with the notation and since we do not want to restate all Assumptions for the abstract case, (they can be found in Schroers2024 we formally chose to state the theorem and its proofs for the term structure setting only.
For the cross-sectional discretization we introduce a sequence of projections $(\Pi_m)_{m\in \mathbb N}$ that coverges strongly to a projection operator $\Pi:H\to H$, which is not necessarily the identity. In the case of term structure models, $\Pi_m$ is defined as in (ref) for which $\Pi f (x)=f(x)\ensuremath{\mathbb{I}}_{[0,M]}(x)$.
We define the discretized truncated semigroup-adjusted realized covariation as
\begin{equation}
SARCV_t^n(u_n,-,m):= \sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} \Pi_m\tilde{\Delta}_i^n f^{\otimes 2}\ensuremath{\mathbb{I}}_{g_{n}(\Pi_m\tilde{\Delta}_i^n f)\leq u_n}
\end{equation}
for $m,n\in \mathbb N \cup\{\infty\}$ and a sequence $(u_n)_{n\in \mathbb N}\subset \mathbb R\cup\{\infty\}$ and a sequence of truncation functions $g_{n}: L^2(\mathbb R_+) \to\mathbb R_+$, such that there are constants $c,C>0$ such that for all $f\in H$ we have
\begin{align}
c\|f\|_{H}\leq g_n(f)\leq C\|f\|_{H}, \quad g_n(f+h)\leq g_n(h)+g_n(f)
\end{align}
Observe that if $\Pi=I$ is the identity on $H$, it is $SARCV(u_n,-,\infty)=SARCV(u_n,-)$ as in the previous section.
As a consequence of the possibile noncommutativity of the semigroup and the projections $\Pi_m$, the rates of convergence also depends on
\begin{equation}
b_m^T:= \int_0^T \|\Pi\Sigma\Pi-\Pi_m\Sigma_s\Pi_m\|_{HS}ds.
\end{equation}
Here we again use the notation $\Sigma_t=\sigma_t\sigma_t^*$ for $t\geq 0$.
That $b_m^T$ indeed converges to $0$ almost surely as $m\to\infty$ is a Corollary of Proposition 4 and Lemma 5 in Panaretos2019.
\begin{theorem}
\begin{itemize}
•
As $n,m\to \infty$ and w.r.t. the Hilbert-Schmidt norm it is
$$ SARCV_t^n(\infty,-,m)\overset{u.c.p.}{\longrightarrow}\Pi[X,X]_t\Pi=\int_0^t \Pi\Sigma_s\Pi ds+\sum_{s\leq t}\left(\Pi X_s-\Pi X_{s-}\right)^{\otimes 2}. $$
• Under Assumption (ref)(2) and w.r.t. the Hilbert-Schmidt norm and as $n,m\to\infty$ it is
$$SARCV_t^n(u_n,-,m)\overset{u.c.p.}{\longrightarrow}\Pi[X^C,X^C]_t\Pi=\int_0^t \Pi\Sigma_s \Pi ds.$$
•
Let Assumptions (ref)(r) hold for some $r\in (0,2)$ and Assumption (ref)($\gamma$) hold for some $\gamma \in (0,1/2]$.
Then it is for each $\rho<(2-r)w$, $T\geq 0$ as $n,m\to \infty$
$$\sup_{t\in [0,T]}\left\|SARCV_t^n(u_n,-,m)-\Pi[X^C,X^C]_t\Pi\right\|_{text{HS}}=\mathcal O_p(\Delta_n^{\min(\gamma,\rho)}+b_m^T)$$
In particular, if $r<2(1-\gamma)$ and $w\in [\gamma/(2-r),1/2]$ we have
$$\sup_{t\in [0,T]}\left\|SARCV_t^n(u_n,-,m)-\Pi[X^C,X^C]_t\Pi\right\|_{\text{HS}}=\mathcal O_p(\Delta_n^{\gamma}+b_m^T)$$
• Assume that
\begin{align}
\mathbb P\left[ \int_0^T\sup_{r\geq 0}\frac{\|(I-\mathcal S(r))\sigma_s\|_{op}^2}{r} ds<\infty\right]=1.
\end{align}
Then Assumption (ref)($1/2$) holds. Let, moreover, Assumption (ref)(r) hold for $r<1$, let $w\in [1/(2-r),1/2]$ and assume that $b_m^T=o_p(\Delta_n^{\frac 12})$.
Then we have w.r.t. the $\|\cdot \|_{L_{\text{HS}}(H)}$ norm and as $n,m\to\infty$ that
$$\sqrt n \left( SARCV_t^n(u_n,-,m)_t^n-\Pi[X^C,X^C]_t\Pi\right)\overset{st.}{\longrightarrow} \Pi\mathcal N(0,\mathfrak Q_t)\Pi,$$
where $\mathcal N(0,\mathfrak G_t)$ is for each $t\geq 0$ a Gaussian random variable in $L_{\text{HS}}(H)$ defined on a very good filtered extension $(\tilde{\Omega},\tilde{\mathcal F},\tilde{\mathcal F}_t,\tilde {\mathbb P})$ of $(\Omega,\mathcal F,\mathcal F_t, \mathbb P)$ with mean $0$
and covariance given for each $t\geq 0$ by a linear operator $\mathfrak Q_t:L_{\text{HS}}(H)\to L_{\text{HS}}(H)$ such that
$$\mathfrak Q_t =\int_0^t \Sigma_s (\cdot+\cdot^*) \Sigma_s ds.$$
•
Let Assumption (ref) hold and $\mathcal C=\mathbb E[ \Sigma_t]$ denote the global covariance of the continuous driving martingale.
Let furthermore Assumption (ref)(p,r) and (ref)($\gamma$) hold (for the abstract semigroup $\mathcal S$) for some $r\in (0,2)$, $\gamma \in (0,1/2]$
and $p>\max(2/(1-2w),(1-wr)/(2w-rw))$.
Then we have w.r.t. the Hilbert-Schmidt norm that as $n,m,T\to\infty$
$$\frac 1T SARCV_T^{n}(u_n,-,m)\overset{p}{\longrightarrow}\Pi\mathcal C\Pi.$$
If $r<2(1-\gamma)$ and $w\in(\gamma/(1-2w),1/2)$, $p\geq 4$ and observing that $\varphi_m= \text{tr}((I-\Pi_m)\mathbb E[\Sigma_1](I-\Pi_m))$ converges to $0$ as $m\to\infty$ (where $\text{tr}$ denotes the trace operation) we have with $a_T= \| [X^C,X^C]/T- \mathcal C\|_{\text{HS}}$ that
$$ \left\|\frac 1T SARCV_T^{n,m}(-)- \Pi \mathcal C\Pi\right\|_{\mathcal H}=\mathcal O_p(\Delta_n^{\gamma}+\varphi_m+a_T).$$
\end{itemize}
\end{theorem}
To prove this abstract result, we make use of the limit theory established in Schroers2024. However, Theorem (ref) is not a direct corollary of these results, since we have to take into account that jump-truncation rules now also depend on possible discrete approximations.
The key result to bridge this gap is
\begin{lemma}
Assume that Assumptions (ref)(p,r) holds for $r\in (0,2]$ and $p> (1-\rho/((2-r)w))^{-1}$ for some $\rho<(2-r)w$ when $r<2$ or $\rho = 0$ if $r=2$. Then we have
\begin{align}
\mathbb E\left[\left\|\sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} (\Pi_m \tilde \Delta_i^n f)^{\otimes 2} \ensuremath{\mathbb{I}}_{g_n(\Pi_m \tilde \Delta_i^n f)\leq u_n}-\sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} (\Pi_m \tilde \Delta_i^n f)^{\otimes 2} \ensuremath{\mathbb{I}}_{g_n( \tilde \Delta_i^n f)\leq u_n}\right\|\right]=Kt\Delta_n^{\rho}\phi_n
\end{align}
for a real sequence $(\phi_n)_{n\in \mathbb N}$ converging to $0$ and a constant $K>0$.
If Assumption (ref) holds, it is
\begin{align}
\left\|\sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} (\Pi_m \tilde \Delta_i^n f)^{\otimes 2} \ensuremath{\mathbb{I}}_{g_n(\Pi_m \tilde \Delta_i^n f)\leq u_n}-\sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} (\Pi_m \tilde \Delta_i^n f)^{\otimes 2} \ensuremath{\mathbb{I}}_{g_n( \tilde \Delta_i^n f)\leq u_n}\right\|=o_p(\Delta_n^{\rho})
\end{align}
\end{lemma}
Before we prove this Lemma, let us introduce some notation.
In the case that Assumption (ref)(r) is valid for $0<r\leq 1$ we write
\begin{align*}
f_t= \mathcal S(t)f_0+\int_0^t \mathcal S(t-s)\alpha_s' ds+\int_0^t \mathcal S(t-s)\sigma_s dW_s+\int_0^t \int_{H\setminus \{0\}} \mathcal S(t-s)\gamma_s(z) N(dz,ds),
\end{align*}
where
$$\alpha_s'= \alpha_s-\int_{H\setminus \{0\}} \gamma_s(z) F(dz)$$
and the integral w.r.t. the (not compensated) Poisson random measure $N$ is well defined (for the second term recall the definition of the integral e.g. from PZ2007).
We then define
\begin{align}
& f_t':=\mathcal S(t) f_0+\int_0^t \mathcal S(t-s)\alpha_s'+\int_0^t \mathcal S(t-s)\sigma_sdW_s,\\
&f_t”:=\int_0^t \int_{H\setminus \{ 0\}}\mathcal S(t-s)\gamma_s(z) N(dz,ds).\notag
\end{align}
If Assumption (ref)(r) holds for $r\in (1,2)$, we define
\begin{align}
& f_t':=\mathcal S(t) f_0+\int_0^t \mathcal S(t-s)\alpha_s+\int_0^t \mathcal S(t-s)\sigma_sdW_s,\\
&f_t”:=\int_0^t \int_{H\setminus \{ 0\}}\mathcal S(t-s)\gamma_s(z) (N-\nu)(dz,ds).\notag
\end{align}
\begin{proof}
We start with the case that Assumption (ref)(p,r) holds for $r\in (0,2]$ and $p>(1-\rho/((2-r)w))^{-1}$
Observe that
\begin{align}
& \left\|\sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} (\Pi_m \tilde \Delta_i^n f)^{\otimes 2} \ensuremath{\mathbb{I}}_{g_n(\Pi_m \tilde \Delta_i^n f)\leq u_n}-\sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} (\Pi_m \tilde \Delta_i^n f)^{\otimes 2} \ensuremath{\mathbb{I}}_{g_n( \tilde \Delta_i^n f)\leq u_n}\right\|\notag\\
\leq & \sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} \left\|\Pi_m \tilde \Delta_i^n f\right\|^2 \left(\ensuremath{\mathbb{I}}_{g_n(\Pi_m \tilde \Delta_i^n f)\leq u_n<g_n( \tilde \Delta_i^n f)}+\ensuremath{\mathbb{I}}_{g_n( \tilde \Delta_i^n f)\leq u_n<g_n(\Pi_m \tilde \Delta_i^n f)}\right)\notag\\
\leq & \sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} \left\|\Pi_m \tilde \Delta_i^n f\right\|^2 \left(\ensuremath{\mathbb{I}}_{c\|\Pi_m \tilde \Delta_i^n f\|\leq u_n<C\|\tilde \Delta_i^n f\|}+\ensuremath{\mathbb{I}}_{c\|\Pi_m \tilde \Delta_i^n f\|\leq u_n< C\|\Pi_m \tilde \Delta_i^n f\|}\right)\notag\\
\leq & 2 c^2u_n^2 \sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} \left(1\wedge \frac{\left\|\Pi_m \tilde \Delta_i^n f\right\|^2}{c^2u_n^2}\right) \ensuremath{\mathbb{I}}_{c\|\Pi_m\tilde \Delta_i^n f\|\leq u_n<C\|\tilde \Delta_i^n f\|}\notag\\
\leq & 2 \sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} \left\|\Pi_m \tilde \Delta_i^n f'\right\|^2 \ensuremath{\mathbb{I}}_{c\|\Pi_m\tilde \Delta_i^n f\|\leq u_n<C\|\tilde \Delta_i^n f\|, c\|\tilde \Delta_i^n f'\|\leq u_n}\\
&\qquad+2 \sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} \left\|\Pi_m \tilde \Delta_i^n f'\right\|^2 \ensuremath{\mathbb{I}}_{c\|\Pi_m\tilde \Delta_i^n f\|\leq u_n<C\|\tilde \Delta_i^n f\|, c\|\tilde \Delta_i^n f'\|> u_n}\\
&+ 2 c^2u_n^2 \sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} \left(1\wedge \frac{\left\|\Pi_m \tilde \Delta_i^n f”\right\|^2}{c^2u_n^2}\right) \ensuremath{\mathbb{I}}_{c\|\Pi_m\tilde \Delta_i^n f\|\leq u_n<C\|\tilde \Delta_i^n f\|, \|\tilde \Delta_i^n f'\|\leq u_n}\\
&\qquad+2 c^2u_n^2 \sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} \left(1\wedge \frac{\left\|\Pi_m \tilde \Delta_i^n f”\right\|^2}{c^2u_n^2}\right) \ensuremath{\mathbb{I}}_{c\|\Pi_m\tilde \Delta_i^n f\|\leq u_n<C\|\tilde \Delta_i^n f\|, \|\tilde \Delta_i^n f'\|> u_n}
\end{align}
We show for all summands (ref), (ref), (ref) and (ref) that they are are bounded by $Kt\Delta_n^{\rho}\phi_n$ for a real sequence $(\phi_n)_{n\in \mathbb N}$ converging to $0$ and a constant $K>0$.
We start with (ref). Since Assumption (ref) holds, we can use Lemma A.1 from Schroers2024. Since $\|\Pi_m\tilde\Delta_i^n f\| \leq u_n$ and $\|\Pi_m\tilde\Delta_i^n f'\|\leq \|\tilde\Delta_i^n f'\|\leq u_n$ implies that $\|\Pi_m\tilde\Delta_i^n f''\| \leq 2 u_n$ we find a constant $K>0$ such that
\begin{align*}
&\mathbb E\left[\sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} \left\|\Pi_m \tilde \Delta_i^n f'\right\|^2 \ensuremath{\mathbb{I}}_{c\|\Pi_m\tilde \Delta_i^n f\|\leq u_n<C\|\tilde \Delta_i^n f\|, c\|\tilde \Delta_i^n f'\|\leq u_n}\right]\\
\leq &\sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} \mathbb E\left[\left\|\Pi_m \tilde \Delta_i^n f'\right\|^p\right]^{\frac 2p} \mathbb E\left[ \left(1\wedge \frac{\|\tilde \Delta_i^n f”\|}{2u_n}\right)\right]^{1-\frac 2p}\\
\leq & Kt \Delta_n^{\rho}\phi_n
\end{align*}
For the second summand (ref), we apply Markov's inequality, choose $l=(2-2rw)/(2-4w)>1$ and again Lemma A.1 from Schroers2024 to obtain a constant $K>0$ such that
\begin{align*}
& \mathbb E\left[ \sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} \left\|\Pi_m \tilde \Delta_i^n f'\right\|^2 \ensuremath{\mathbb{I}}_{c\|\Pi_m\tilde \Delta_i^n f\|\leq u_n<C\|\tilde \Delta_i^n f\|, c\|\tilde \Delta_i^n f'\|> u_n}\right]\\
\leq & \sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} \mathbb E\left[\left\|\Pi_m \tilde \Delta_i^n f'\right\|^p\right]^{\frac 2p}\mathbb P\left[c\|\tilde \Delta_i^n f'\|> u_n\right]^{\frac{p-2}{p}}\\
\leq & Kt \Delta_n^{\rho}\phi_n
\end{align*}
Turning to the third summand, we again make use of Lemma A.1 from Schroers2024 to obtain a constanr $K>0$ and a real sequence $(\phi_n)_{n\in \mathbb N}$ convrging to $0$ such that
\begin{align*}
& \mathbb E\left[c^2u_n^2 \sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} \left(1\wedge \frac{\left\|\Pi_m \tilde \Delta_i^n f”\right\|^2}{c^2u_n^2}\right) \ensuremath{\mathbb{I}}_{c\|\Pi_m\tilde \Delta_i^n f\|\leq u_n<C\|\tilde \Delta_i^n f\|, \|\tilde \Delta_i^n f'\|\leq u_n}\right]\\
\leq &\mathbb E\left[c^2u_n^2 \sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} \left(1\wedge \frac{\left\| \tilde \Delta_i^n f”\right\|^2}{c^2u_n^2}\right)^2\right]\\
\leq & K c^2t\Delta_n^{\rho}\phi_n.
\end{align*}
For the fourth summand we find for $1<q=(2-r)w/\rho$ if $r<1$ and $1<q$ arbitrary if $r=2$ and use once more Lemma A.1 from Schroers2024 to obtain a constant $K>0$ and a real sequence $(\phi_n)_{n\in \mathbb N}$ converging to $0$ such that
\begin{align*}
& c^2u_n^2 \sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} \left(1\wedge \frac{\left\|\Pi_m \tilde \Delta_i^n f”\right\|^2}{c^2u_n^2}\right) \ensuremath{\mathbb{I}}_{c\|\Pi_m\tilde \Delta_i^n f\|\leq u_n<C\|\tilde \Delta_i^n f\|, \|\tilde \Delta_i^n f'\|> u_n}\\
\leq & c^2u_n^2 \sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} \mathbb E\left[\left(1\wedge \frac{\left\|\Pi_m \tilde \Delta_i^n f”\right\|}{cu_n}\right)^q\right]^{\frac 1q}\mathbb P[|\tilde \Delta_i^n f'\|> u_n]^{\frac {q-1}q}\\
\leq & K t c^2 \Delta_n^{\rho}\phi_n.
\end{align*}
Summing up, we proved (ref).
Let us now turn to the case that only Assumption (ref) holds.
Assumption (ref) implies that there is a localizing sequence of stopping times $(\rho_n)_{n\in\mathbb N}$ such that $\alpha_{t\wedge \rho_n}$ is bounded for each $n\in \mathbb N$. As $\sigma$ and $f$ are c{\`a}dl{\`a}g, the sequence of stopping times $\theta_n:=\inf\{s: \|f_s\|+\|\sigma_s\|_{\text{HS}}\geq n\}$ are localizing as well. If $(\tau_n)_{n\in\mathbb N}$ is the sequence of stopping times for the jump part as described in Assumption (ref), we can define
$\varphi_n:=\rho_n\wedge \theta_n\wedge \tau_n,\quad n\in\mathbb N.$
This defines a localizing sequence of stopping times, for which the coefficients $\alpha_s \ensuremath{\mathbb{I}}_{s\leq \varphi_n}$, $\sigma_s \ensuremath{\mathbb{I}}_{s\leq \varphi_n}$ and $\gamma_s(z) \ensuremath{\mathbb{I}}_{s\leq \varphi_n}$ satisfy Assumption (ref)(p,r) for $r\in (0,2)$ and all $p>0$.
Now define
\begin{align*}
Z_n(t):=& \Delta_n^{-\rho}\sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} \left\|\Pi_m \tilde \Delta_i^n f\right\|^2 (\ensuremath{\mathbb{I}}_{g_n(\Pi_m \tilde \Delta_i^n f)\leq u_n}- \ensuremath{\mathbb{I}}_{g_n( \tilde \Delta_i^n f)\leq u_n}). \\
\geq & \Delta_n^{-\rho}\left\|\sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} (\Pi_m \tilde \Delta_i^n f)^{\otimes 2} \ensuremath{\mathbb{I}}_{g_n(\Pi_m \tilde \Delta_i^n f)\leq u_n}-\sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} (\Pi_m \tilde \Delta_i^n f)^{\otimes 2} \ensuremath{\mathbb{I}}_{g_n( \tilde \Delta_i^n f)\leq u_n}\right\|.
\end{align*}
If $\varphi_n\geq t+1$, we have
\begin{align*}
Z_n(t\wedge \varphi_N)\leq & \Delta_n^{-\rho}\sum_{i=1}^{\lfloor (t+1)/\Delta_n\rfloor} \left\|\Pi_m \tilde \Delta_i^n f_{\cdot\wedge\varphi_N}\right\|^2 (\ensuremath{\mathbb{I}}_{g_n(\Pi_m \tilde \Delta_i^n f_{\cdot\wedge\varphi_N})\leq u_n}- \ensuremath{\mathbb{I}}_{g_n( \tilde \Delta_i^n f_{\cdot\wedge\varphi_N})\leq u_n}).
\end{align*}
We obtain as $n\to \infty$
\begin{align*}
\lim_{n\to \infty} \mathbb P\left[\sup_{t\in [0,T]}\mathcal Z_n^i(t)\geq \epsilon\right]
\leq &\lim_{n\to \infty} \mathbb P\left[\sup_{t\in [0,T]}\mathcal Z_n^i(t\wedge \varphi_N)
\geq \epsilon, T< \varphi_N\right]+ \lim_{n\to \infty}\mathbb P\left[T \geq\varphi_N \right]=0
\end{align*}
where the convergence in the last line is due to (ref) and Markov's inequality and since we know that $\Delta_n^{-\rho}\sum_{i=1}^{\lfloor (t+1)/\Delta_n\rfloor} \left\|\Pi_m \tilde \Delta_i^n f_{\cdot\wedge\varphi_N}\right\|^2 (\ensuremath{\mathbb{I}}_{g_n(\Pi_m \tilde \Delta_i^n f_{\cdot\wedge\varphi_N})\leq u_n}- \ensuremath{\mathbb{I}}_{g_n( \tilde \Delta_i^n f_{\cdot\wedge\varphi_N})\leq u_n})$ converges to $0$ uniformly on compacts by (ref). This yields (ref).
\end{proof}
Now we are able to prove Theorem (ref) as a Corollary of the results in Schroers2024 and Lemma (ref).
\begin{proof}[Proof of Theorem (ref)]
We start with assertion (i). For that, we observe that
\begin{align*}
\Pi_m SARCV_t^n \Pi_m-[\Pi X,\Pi X]_t= \left( \Pi_m SARCV_t^n\Pi_m-[\Pi_m X,\Pi_m X] \right)+\left([\Pi_m X,\Pi_m X]-[\Pi X,\Pi X]\right)
\end{align*}
For the first summand it is
\begin{align*}
\left\| \Pi_m SARCV_t^n\Pi_m-[\Pi_m X,\Pi_m X]_t \right\|= \left\| \Pi_m\left( SARCV_t^n-[ X,X]_t \right) \Pi_m\right\|\leq \left\| SARCV_t^n-[ X,X]_t \right\|,
\end{align*}
which converges to $0$ as $n\to \infty$ uniformly on compacts in probability by Theorem 3.1 in Schroers2024. For the second summand, we have
\begin{align*}
& \sup_{t\in [0,T]} \left\|[\Pi_m X,\Pi_m X]-[\Pi X,\Pi X]\right\|\\
\leq & \int_0^T \|\Pi_m \Sigma_s \Pi_m-\Sigma_s\| ds+ \sum_{s\leq T} \|\Pi_m (X_s-X_{s-})^{\otimes 2}\Pi_m-(X_s-X_{s-})^{\otimes 2}\|.
\end{align*}
By dominated convergence, if we can prove that for all $s\geq 0$ it is as $m\to\infty$ and in probability that
\begin{align}
\|\Pi_m \Sigma_s \Pi_m-\Sigma_s\|\to 0 and \|\Pi_m (X_s-X_{s-})^{\otimes 2}\Pi_m-(X_s-X_{s-})^{\otimes 2}\|\to 0,
\end{align}
the proof follows. But this holds true even as almost sure convergence, by Proposition 4 and Lemma 5 in Panaretos2019.
Before we prove the remaining assertions, let us observe the subsequent error decomposition
\begin{align}
&SARCV(u_n,,-,m)_t^n-[\Pi X^C,[\Pi X^C]_t \notag\\
\leq &SARCV(u_n,,-,m)_t^n-\sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} (\Pi_m\tilde \Delta_i^n f)^{\otimes 2}\ensuremath{\mathbb{I}}_{g_n(\tilde \Delta_i^n f)\leq u_n}\\
& \qquad +\Pi_m\left(\sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} (\tilde \Delta_i^n f)^{\otimes 2}\ensuremath{\mathbb{I}}_{g_n(\tilde \Delta_i^n f)\leq u_n}-[X^C,X^C]_t\right)\Pi_m\\
& \qquad+ [\Pi_m X^C,\Pi_m X^C]_t-[\Pi X^C,\Pi X^C]_t
\end{align}
We proceed with the proof of (ii). By Lemma (ref), (ref) converges to $0$. The second summand (ref) converges to $0$ by Theorem 3.2 in Schroers2024. The last summand (ref) is bounded by $b_m^T$, which converges to $0$ as $m\to \infty$.
Let us now turn to the proof of (iii), which works analogous to the proof of (ii), by employing the decompositiion of the approximation error into (ref), (ref) and (ref). Indeed, Lemma (ref) yields that (ref) is $o_p(\Delta_n^{\rho})$ with respect to the Hilbert-Schmidt norm, while Theorem 3.3 in Schroers2024 yields that the second summand is $\mathcal O_p(\Delta_n^{\min(\rho,\gamma)})$,with respect to the Hilbert-Schmidt norm, which shows (iii).
Now let us prove the central limit theorem (iv). Again employing the error decomposition into (ref), (ref) and (ref), we find that, Lemma (ref) yields that (ref) is $o_p(\Delta_n^{\rho})=o_p(\Delta_n^{1/2})$ and by Assumption, the same holds for (ref), since it is bounded by .$b_m^T$. Hence, we find that under the Assumptions imposed in (iv), it is
\begin{align*}
& \sqrt n\left(SARCV(u_n,,-,m)_t^n-[\Pi X^C,[\Pi X^C]_t\right)\\
= & \Pi_m \left(\sqrt n\left(\sum_{i=1}^{\ensuremath{\lfloor t/\Delta_n\rfloor}} (\tilde \Delta_i^n f)^{\otimes 2}\ensuremath{\mathbb{I}}_{g_n(\tilde \Delta_i^n f)\leq u_n}-[X^C,X^C]_t\right)\right)+o_p(1)
\end{align*}
Now (iv) follows directly from Theorem 3.5 in Schroers2024.
We conclude the proof by showing (v). For that we introduce the decomposition
\begin{align}
&\frac 1TSARCV(u_n,,-,m)_T^n-\Pi \mathcal C \Pi \notag\\
\leq & \frac 1T SARCV(u_n,,-,m)_T^n-\frac 1T\sum_{i=1}^{\ensuremath{\lfloor T/\Delta_n\rfloor}} (\Pi_m\tilde \Delta_i^n f)^{\otimes 2}\ensuremath{\mathbb{I}}_{g_n(\tilde \Delta_i^n f)\leq u_n}\\
& \qquad +\Pi_m\left(\frac 1T\sum_{i=1}^{\ensuremath{\lfloor T/\Delta_n\rfloor}} (\tilde \Delta_i^n f)^{\otimes 2}\ensuremath{\mathbb{I}}_{g_n(\tilde \Delta_i^n f)\leq u_n}-\mathcal C\right)\Pi_m\\
& \qquad+\Pi_m \mathcal C \Pi_m-\Pi \mathcal C \Pi.
\end{align}
By (ref), the first summand (ref) is $ o_p(\Delta_n^{\rho})$. The second summand (ref) converges to $0$ by Theorem 3.6 in Schroers2024 and the third summand (ref) converges to $0$ as $m\to\infty$. We obtain the rates of convergence also from Theorem 3.6 in Schroers2024 applied to (ref) and since $\rho$ can be chosen larger than $1/2$ if $r<1$ and the last summand equals $\text{tr}((\Pi-\Pi_m)\mathcal C(\Pi-\Pi_m))$.
\end{proof}
\subsection{Formal proofs of Section (ref)}
We will now show how (ref), Theorem (ref), (ref) and (ref) can be deduced from Theorem (ref).
Let us begin with the general identifiability results.
\begin{proof}[Proof of Theorem (ref)]
We have that the integral operator $\Delta_n^{-2}\mathcal T_{\hat q_t^{n}}$ corresponding to the piecewise constant kernel $\Delta_n^{-2}\hat q_t^{n}$, according to Remark (ref) is given by
$$\Delta_n^{-2}\mathcal T_{\hat q_t^{n}}= \Pi_{n,M}(SARCV_t^n)\Pi_{n,M}.$$
where $\Pi_{n,M}$ is defined as in (ref). Setting $\Pi_m= \Pi_{m,M}$ and $\Pi=I$ if $M=\infty$, or resp.$\Pi f(x)=\ensuremath{\mathbb{I}}_{[0,M]}(x)f(x)$ and $M<\infty$ for $f\in L^2(\mathbb R_+)$, the result follows immediately from Theorem (ref)(i)
\end{proof}
\begin{proof}[Proof of Theorem (ref)]
Again using the notation of Remark (ref), we can observe that the integral operator $\Delta_n^{-2}\mathcal T_{\hat q_t^{n,M,-}}$ corresponding to the kernel $\Delta_n^{-2}q_t^{n,M,-}$ defined in Remark (ref) is (with $\Pi_{n,M}$ as in Remark (ref)) given by
$$\Delta_n^{-2}\mathcal T_{\hat q_t^{n,M,-}}= (SARCV_t^n(u_n,-,n).$$
Thus, setting $\Pi_m= \Pi_{m,M}$ and $\Pi=I$ if $M=\infty$, or resp.$\Pi f(x)=\ensuremath{\mathbb{I}}_{[0,M]}(x)f(x)$ and $M<\infty$ for $f\in L^2(\mathbb R_+)$, the result follows immediately from (ref)(ii).
\end{proof}
Before proving Theorems (ref) and (ref), we observe that we can quantify the spatial discretization error now also in terms of the regularity of the semigroup.
\begin{theorem}
We have for all $\gamma >0$ that
\begin{align}
\left\|\Pi_{n,M}\Sigma_s\Pi_{n,M}- \Sigma_s \right\|_{HS} \leq 2 \|\sigma_s\|_{\text{op}} \left\|\Pi_{n,M}\sigma_s- \sigma_s\right\|_{\text{HS}}\notag
\leq & 2\|\sigma_s\|_{\text{op}}\Delta_m^{\gamma} \sup_{r\leq \Delta_n}\frac{\left\|\left(\mathcal S(r)-I\right)\sigma_s\right\|_{\text{HS}} }{r^{\gamma}}.
\end{align}
Hence, if Assumption (ref)($\gamma$) for $\gamma\in (0,1/2]$ is valid, we find (with $\Pi=I$ if $M=\infty$, or resp.$\Pi f(x)=\ensuremath{\mathbb{I}}_{[0,M]}(x)f(x)$ and $M<\infty$ for $f\in L^2(\mathbb R_+)$)
\begin{equation}
\left\|\Pi_{n,M} \int_0^t \Sigma_s ds \Pi_{n,M}-\int_0^t \Pi\Sigma_s \Pi ds\right\|_{\text{HS}} =\mathcal O_p(\Delta_n^{\gamma})
\end{equation}
If even Assumption (ref) holds, we find a constant $K$, which is independent of $T$ and $m$ such that
\begin{equation}
\mathbb E\left[\sup_{t\in [0,T]} \left\|\Pi_{n,M} \int_0^t \Sigma_s ds \Pi_{n,M}-\int_0^t \Pi\Sigma_s\Pi ds\right\|_{\text{HS}} \right]\leq KT\Delta_m^{\gamma}.
\end{equation}
\end{theorem}
\begin{proof} Let $q_s^{\sigma}\in L^2(\mathbb R_+^2)$ denote the integral kernel such that for all $f\in L^2(\mathbb R_+)$ and $x\geq 0$ it is
$$\sigma_s f(x)= \int_{\mathbb R_+} q_s^{\sigma}(x,y)f(y) dy.$$
Without loss of generality, choose $q_s^{\sigma}$ to be symmetric.
Then for $M=\infty$ and $M<\infty$ it is
\begin{align*}
\left\|(\Pi_{m,M}-\Pi) \sigma_s \right\|_{\text{HS}}^2
\leq &\int_{[0,M]} \sum_{j=1}^{\lfloor M/\Delta_m\rfloor}\int_{(j-1)\Delta_m}^{j\Delta_m} \Delta_m^{-1}\int_{(j-1)\Delta_m}^{j\Delta_m} \left(q_s^{\sigma}(x',y)-q_s^{\sigma}(x,y) \right)^2dx' dx dy\\
\leq &2\Delta_m^{-1}\int_{[0,M]} \sum_{j=1}^{\lfloor M/\Delta_m\rfloor}\int_{(j-1)\Delta_m}^{j\Delta_m} \int_{x}^{j\Delta_m} \left(q_s^{\sigma}(x',y)-q_s^{\sigma}(x,y) \right)^2dx' dx dy\\
= &2\Delta_m^{-1}\int_{0}^{\Delta_m}\int_{[0,M]} \sum_{j=1}^{\lfloor M/\Delta_m\rfloor}\int_{(j-1)\Delta_m}^{j\Delta_m} \left(q_s^{\sigma}(x'+x,y)-q_s^{\sigma}(x,y) \right)^2 dx dydx'. \end{align*}
Hence,
\begin{align*}
\left\|(\Pi_{m,M}-\Pi) \sigma_s \right\|_{\text{HS}}^2
\leq & 2 \sup_{x\leq \Delta_m} \left(\Delta_m^{-1}\int_{[0,M]}\sum_{j=1}^{\lfloor M/\Delta_m\rfloor}\int_{(j-1)\Delta_m}^{j\Delta_m} \left(((\mathcal S(x)-I)q_s(\cdot, y'))(x'))\right)^2 dx'dy'\right)\\
\leq & 2 \sup_{x\leq \Delta_m} \left(\int_{[0,M]}\int_{[0,M]} \left(\frac{((\mathcal S(x)-I)q_s(\cdot, y'))(x'))}x\right)^2 dx'dy'\right).
\end{align*}
This proves the claim.
\end{proof}
It is left to show Theorem (ref), Theorem (ref) and Theorem (ref). We start with the
\begin{proof}[Proof of Theorem (ref)]
Using again notation (ref) it is (with $\Pi_{n,M}$ as in Remark (ref))
$\Delta_n^{-2}\mathcal T_{\hat q_t^{n,M,-}}= SARCV_t^n(u_n,-,n)$ where $\hat q_t^{n,M,-}$ is defined in Remark (ref) and we obtain from Theorem (ref)(iii)
that
\begin{equation*}
\sup_{m\in \mathbb N\cup\{\infty\}}\sup_{t\in [0,T]}\left\|SARCV_t^n(u_n,-,n)-\int_0^t \Pi_{n,M}\Sigma_s \Pi_{n,M}ds\right\|_{\text{HS}}=\mathcal O_p\left(\Delta_n^{\min(\gamma,\rho)}\right).
\end{equation*}
Moreover, due to Theorem (ref) we obtain that with $\Pi=I$ if $M=\infty$, or resp.$\Pi f(x)=\ensuremath{\mathbb{I}}_{[0,M]}(x)f(x)$ and $M<\infty$ for $f\in L^2(\mathbb R_+)$ it is
\begin{equation*}
\left\|\Pi_{n,M}\int_0^t \Sigma_s ds \Pi_{n,M}-\Pi\int_0^t \Sigma_s ds\Pi\right\|_{\text{HS}} = \mathcal O_p(\Delta_m^{\gamma}),
\end{equation*}
which proves the claim.
\end{proof}
We continue with the
\begin{proof}[Proof of Theorem (ref)]
We first prove that (ref) implies Assumption (ref)($1/2$). It is by H{\"o}lder's inequality and the basic inequality $=\|AB^*\|_{\text{HS}}\leq \|A\|_{\text{HS}}\|B\|_{\text{op}}$
for a Hilbert-Schmidt operator $A$ and a bounded linear operator $B$,
\begin{align*}
\int_0^T \|q_s^C\|_{\mathfrak F_{\gamma}}ds
\leq &\left(\int_0^T \sup_{r>0}\frac{\|(I-\mathcal S(r))\sigma_s\|_{\text{op}}^2}{r}ds\right)^{\frac 12}\left(\int_0^T \|\sigma_s\|_{\text{HS}}^2ds\right)^{\frac 12}.
\end{align*}
Now (ref) implies that the factor on the left is finite almost surely, whereas the factor on the right is finite almost surely, due to the stochastic integrability of the volatility. This implies that Assumption (ref)($1/2$) is valid.
We now continue to derive Theorem (ref) from Theorem (ref)(iv).
We only have to show that $b_n^T=\int_0^T (\Pi-\Pi_{n,M})\Sigma_s ds=o_p(\Delta_n^{ 1/2})$.
For that, observe that
\begin{align*}
& \Delta_n^{-\frac 12}\left\|\Pi_{n,M} \int_0^t \Sigma_s ds \Pi_{n,M}-\int_0^t \Pi\Sigma_s \Pi ds\right\|_{\text{HS}} \\
\leq & \Delta_n^{-\frac 12}\int_0^t\left\|(\Pi_{n,M}-\Pi) \Sigma_s \right\|_{\text{HS}}ds+ \Delta_n^{-\frac 12}\int_0^t\left\| \Sigma_s (\Pi_{n,M}-\Pi)\right\|_{\text{HS}}ds.
\end{align*}
We will prove convergence of the first summand to $0$ as $n\to \infty$, while for the second summand, the proof is analogous.
We define the orthonormal basis $(e_j)_{j\in \mathbb N}\subset C_c(I)\subset L^2(I)$ where either $I=\mathbb R_+$ if $M=\infty$ and $I=[0,M]$ if $M<\infty$ and $C_c(I)$ is the set of compactly supported infinitely differentiable functions (which is dense in $L^2(I)$). Then, obviously for each $j$ we can find a constant $k_j$ such that $\sup_{|x-y|\leq \Delta_n}|e_j(x)-e_j(y)|\leq k_j \Delta_n$ and, thus, if $K_j\in \mathbb N$ such that $e_j(x)=0$ for all $x\geq K_j$
\begin{align*}
\|(\Pi_{n,M}-\Pi)e_j\|^2=&\int_0^{K_j}\left(\sum_{i=1}^{\infty} n\int_{(i-1)\Delta_n}^{i\Delta_n} e_j(y)dy\ensuremath{\mathbb{I}}_{[(i-1)\Delta_n,i\Delta_n]}(x)-e_j(x)\right)^2dx
\leq k_j^2 \Delta_n^2 K_j.
\end{align*}
Let $P_N$ denote the orthonormal projection onto $span(e_i\otimes e_j:i,j=1,...,N)$. We can decompose
\begin{align*}
& \Delta_n^{-\frac 12}\int_0^t\left\|(\Pi_{n,M}-\Pi) \Sigma_s \right\|_{\text{HS}}ds\\
\leq & \Delta_n^{-\frac 12}\int_0^t\left\| P_N(\Pi_{n,M}-\Pi)\Sigma_s \right\|_{\text{HS}}ds+\Delta_n^{-\frac 12}\int_0^t\left\| (I-P_N)(\Pi_{n,M}-\Pi)\Sigma_s \right\|_{\text{HS}}ds
\end{align*}
It is simple to see that $\|P_N A\|_{\text{HS}}\leq \sum_{i,j=1}^N |\langle Ae_i,e_j\rangle|$ and, hence,
For the first part we find
\begin{align*}
\Delta_n^{-\frac 12}\int_0^t\left\| P_N(\Pi_{n,M}-\Pi)\Sigma_s \right\|_{\text{HS}}ds
\leq & \Delta_n^{\frac 12}\int_0^t\| \Sigma_s \|_{\text{nuc}}ds\left( \sum_{j=1}^Nk_j \sqrt K_j\right).
\end{align*}
This converges to $0$ as $n\to\infty$ for all $N\in \mathbb N$.
For the second summand we observe that $I-P_N$ is the orthonormal projection onto $\overline{span(e_i\otimes e_j.i,j\geq N+1)}$ and hence can be written as $I-P_N=(I-p_N)(\cdot)(I-p_N)$ where $I-p_N=\sum_{i=N+1}^{\infty} e_i^{\otimes 2}$.
We find by H{\"o}lder's inequality that
\begin{align*}
& \Delta_n^{-\frac 12}\int_0^t\left\| (I-P_N)(\Pi_{n,M}-\Pi)\Sigma_s \right\|_{\text{HS}}ds\\
= & \Delta_n^{-\frac 12}\int_0^t\left\| (I-p_N)(\Pi_{n,M}-\Pi)\Sigma_s (I-p_N)\right\|_{\text{HS}}ds
\\
\leq & \Delta_n^{-\frac 12}\left(\int_0^t\left\| (I-p_N)(\Pi_{n,M}-\Pi)\sigma_s\right\|_{\text{op}}^2ds\right)^{\frac 12}\left(\int_0^t\|\sigma_s^*(I-p_N)\|_{\text{HS}}^2ds\right)^{\frac 12}
\end{align*}
The second factor converges to $0$ as $N\to\infty$ since
$$\| \sigma_s^* (I-p_N)\|_{\text{HS}}^2= \sum_{i=1}^{\infty}\| \sigma_s(I-p_N) e_j\|^2=\sum_{i=N+1}^{\infty}\| \sigma_s e_j\|^2$$
converges to $0$ as $N\to \infty$ and then the dominated convergence theorem applies. The first factor is bounded, since
\begin{align*}
\int_0^t\left\| (I-p_N)(\Pi_{n,M\Pi}-\Pi)\sigma_s\right\|_{\text{op}}^2ds \leq & \int_0^t\left\| (\Pi_{n,M}-\Pi)\sigma_s\right\|_{\text{op}}^2ds\\
\leq &2\Delta_n \int_0^t \sup_{r\leq \Delta_n}\frac{\left\|\left(\mathcal S(r)-I\right)\sigma_s\right\|_{L_{\text{HS}}(L^2(\mathbb R_+))} }{r}ds.
\end{align*}
This is finite by Assumption and summing up we obtain that as $N\to \infty$
$$ \sup_{n\in \mathbb N}\Delta_n^{-\frac 12}\int_0^t\left\| (I-P_N)(\Pi_{n,M}-\Pi)\Sigma_s \right\|_{L_{\text{HS}}(L^2(\mathbb R_+))}ds\to 0.$$
\end{proof}
Let us now conclude with the
\begin{proof}[Proof of Theorem (ref)]
We use that as before, the for integral operator $\Delta_n^{-2}\mathcal T_{\hat q_t^{n,M,-}}$ corresponding to the kernel $\Delta_n^{-2}q_t^{n,M,-}$ defined in Remark (ref) it is $\Delta_n^{-2}\mathcal T_{\hat q_t^{n,M,-}}= (SARCV_t^n(u_n,-,n)$(with $\Pi_{n,M}$ as in Remark (ref)). We obtain under the Assumption of Theorem (ref) that by Theorems (ref)(v) and Theorem (ref)
there is a constant $K>0$, which is independent of $T$ and $n$ such that
\begin{equation*}
\mathbb E\left[\sup_{m\in \mathbb N\cup\{\infty\}}\sup_{t\in [0,T]}\left\|SARCV_t^n(u_n,-,n)-\int_0^t \Pi_{n,M}\Sigma_s \Pi_{n,M} ds\right\|\right]\leq K T \Delta_n^{\gamma}.
\end{equation*}
and
\begin{equation*}
\mathbb E\left[\sup_{t\in [0,T]} \left\|\Pi_{n,M} \int_0^t \Sigma_s ds \Pi_{n,M}-\int_0^t \Pi\Sigma_s \Pi ds\right\|_{L_{\text{HS}}(L^2(0,M))} \right]\leq KT\Delta_n^{\gamma}.
\end{equation*}
Moreover, by Assumption we have that as $T\to\infty$
$$\frac 1T \int_0^T \Pi\Sigma_s \Pi ds\overset{p}{\longrightarrow}\Pi\mathcal C\Pi.$$
Hence, the claim follows since we can decompose
\begin{align*}
&\frac 1T SARCV_T^n(u_n,-,n)-\mathcal C\\
=&\frac 1T \left(SARCV_T^n(u_n,-,n)-\int_0^t \Pi_{n,M}\Sigma_s \Pi_{n,M} ds\right)+\frac 1T \int_0^T \Pi_{n,M}\Sigma_s \Pi_{n,M} -\Pi\Sigma_s \Pi ds\\ &\quad+ \frac 1T \int_0^t \Pi\Sigma_s \Pi ds -\Pi\mathcal C\Pi.
\end{align*}
\end{proof}
\section{Further Practical considerations}
We now make some considerations for the practical implementation of the estimator here. Precisely, we discuss the effects of smoothing the data a posteriori in the cross-sectional dimension and showcase a possible rescaling procedure for the truncation rule described in Section (ref).
\subsection{Ex-post smoothing}
For term structure models, we might have strong beliefs that forward curves are continuous or even differentiable. While such smoothness Assumptions are reflected by better rates of convergence, the estimator $\hat q^n$ is discontinuous and we might want to derive a smooth approximation instead. A possible way to achieve this is to smooth the estimators a posteriori. This can also serve the purpose of an ex-post regularization to obtain more pleasing visual results or can favor the computational tractability of the estimator (a difference return curve with a daily resolution and 10 years maximally considered maturity needs to store approximately 2500 data points).
Hence, we might want to reduce the number of data points in the maturity direction in the sense of functional data analysis. That is, let $P_m$ be an orthonormal projection onto a finite-dimensional subspace of $L^2(0,M)$ which is spanned by the orthonormal vectors $e_1,....,e_m$. For instance, we could consider a spline basis, Fourier bases or just a lower resolution than daily (e.g. monthly) and let $P_m$ be the projection onto $\{\ensuremath{\mathbb{I}}_{[(j-1)\Delta_m,j\Delta_m]}/\sqrt{\Delta_{m}}:j=1,...,\lfloor M/\Delta_m\rfloor\}$ for $m=n*l$ for some $l\in \mathbb N$. In general, if $P_m$ is a continuous linear projection, we have
$$\sup_{t\in [0,T]}\| \Delta_n^{-2} P_m \mathcal T_{\hat q_t^{n,-}} P_m-\int_0^t\Sigma_s ds\|\leq \sup_{t\in [0,T]}\|\Delta_n^{-2}\mathcal T_{\hat q_t^{n,-}}-\int_0^t\Sigma_s ds\|+\int_0^T\| P_m\Sigma_s P_m-\Sigma_s\| ds,$$
so the additional error is quantified by the second summand on the left.
As long as $\Pi_m \to I$ strongly, this converges to $0$ as $m\to \infty$ by Proposition 4 and Lemma 5 in Panaretos2019.
The exact rate of convergence depends on the particular projection as well as the regularity of the volatility operator. It can be quantified by imposing further regularity assumptions on $(\Sigma_t)_{t\geq 0}$.
An example is given next.
\begin{example}[Forward curves in reproducing kernel Hilbert spaces]
Assume that $\Sigma_s$ maps into a reproducing kernel Hilbert space $H_k= \mathcal T_k^{\frac 12} L^2(\mathbb R_+^2)\subset L^2(\mathbb R_+^2)$ where $k\in L^2(\mathbb R_+^2)$ is a kernel and $\mathcal T_k$ is the corresponding positive definite integral operator with kernel $k$.
The space $H_k$ can be equipped with the norm
$\|f\|_{H_k}=\|\mathcal T_k^{-\frac 12}f\|_{L^2(\mathbb R_+)}$. For instance, we might assume that $k(x,y)=\frac 1{a}(1+e^{-a\min(x,y)})$ for some $a>0$ corresponding to the forward curve space introduced by Filipovic2000, which is also the space in which the nonparametrically smoothed yield curve data from FPY2022 are taken that we use for our empirical analysis in Section (ref) . Such a kernel has a Mercer decomposition
$k(x,y)=\sum_{i=1}^{\infty}\lambda_i e_i(s)e_i(t)$ for an orthonormal basis $(e_i)_{i\in \mathbb N}$ of $L^2(\mathbb R_+)$ and corresponding positive eigenvalues $(\lambda_i)_{i\in \mathbb N}$. We might specify $P_m=\sum_{i=1}^m e_i^{\otimes 2}$ to be the orthonormal projection onto these basis functions.
If we even have that $\Sigma_s\in L_{\text{HS}}(L^2(\mathbb R_+),H_k)$, and $\int_0^T \|\Sigma_s \|_{L_{\text{HS}}(L^2(\mathbb R_+),H_k)} ds<\infty$ almost surely, we obtain
\begin{align*}
\int_0^T\| P_m\Sigma_s P_m-\Sigma_s\|_{L_{\text{HS}(L^2(\mathbb R_+))}} ds
\leq \lambda_{m+1}^{\frac 12}\int_0^T \|\Sigma_s\|_{L_{\text{HS}(L^2(\mathbb R_+),H_k)}}ds,
\end{align*}
which yields an additional $\mathcal O_p(\lambda_{m+1}^{\frac 12})$-error.
\end{example}
\subsection{Remarks on the scaling factor for preliminary estimators of the quadratic variation}
In Section (ref) we adjusted the truncated estimator $q_t^n(-)$ in the preliminary step by some $\rho^*>0$.
As we do not know $\Sigma$, this correct scaling can be conducted in several ways. One reasonable possibility is to choose $\rho^*$ in such a way that the scaled truncated estimator coincides with another robust variance estimate for the data projected onto a particular linear functional. In the simple framework without drift and jumps and where $\Sigma$ is constant and independent of the driving Wiener process and $\Delta_n$ small, we have that $ \Delta_n^{-2}\sum_{i=1}^{\lfloor M/\Delta_n\rfloor} \tilde \Delta_{i\Delta_n} d(j\Delta_n)\ensuremath{\mathbb{I}}_{[(j-1)\Delta_n,j\Delta_n]}\approx \tilde \Delta_i^n f\overset{approx.}{\sim} N(0,\frac 1T \int_0^T \Sigma_s ds)$.
Hence, we choose
$$\rho^*=\frac {\left(q_{.75}-q_{.25}\right)^2}{4\Phi^{-1}(0.75)^2\Delta_n\hat\lambda_1}$$
where $q_{.75}$, and resp. the $q_{.25}$, is the $0.75$-quantile and resp. the $0.25$-quantile, of the data $\sum_{i=1}^{\lfloor M/\Delta_n\rfloor} \tilde \Delta_{i\Delta_n} d(j\Delta_n)\langle \ensuremath{\mathbb{I}}_{[(j-1)\Delta_n,j\Delta_n]}, \hat e_1\rangle $, $i=1,...,\ensuremath{\lfloor T/\Delta_n\rfloor}$, $\hat \lambda_1$ and $\hat e_1$ are respectively the first eigenvalue and the first eigenvector of the preliminary estimator $\hat q_t^n$ and $\Phi^{-1}(0.75)$ is the $.75$ quantile of the standard normal distribution.
In this way,
the rescaled estimator $\rho^* \hat q_t^n (-)$ projected onto $\hat e_1^{\otimes 2}$ corresponds to the interquartile estimator of the variance of the factor loadings of the first eigenvector $\hat e_1$, that is, $\hat \lambda_1$
corresponds to the normalized interquartile range estimator
$$NIQR^2=\left(\frac {(q_{.75}-q_{.25})}{2\sqrt\Delta_n\Phi^{-1}(0.75)}\right)^2.$$
\section{Remarks on the simulation scheme}
We here describe how to sample local averages
$F_{i,j}:=\langle \ensuremath{\mathbb{I}}_{[(j-1)\Delta_n, j\Delta_n]}, f_{i\Delta_n}\rangle_{L^2([0,10]}$
for $n=100$, $i=1,...,100$ and $j=1,...,1000$
of the forward curve process
described in section (ref).
We use that for $i\geq 1$ it is
\begin{align*}
F_{i,\cdot}\equiv & \Pi_{10} f_{i\Delta_n}= \Pi_{10} \mathcal S(\Delta_n)f_{(i-1)\Delta_n}+\Pi_{10} \int_{(i-1)\Delta_n}^{i\Delta_n} \mathcal S(i\Delta_n-s) dX_s= F_{i-1,\cdot+1}+\Pi_{10} \tilde \Delta_i^n f
\end{align*}
where $\tilde \Delta_i^n f= f_{i\Delta_n}-\mathcal S(\Delta_n)f_{(i-1)\Delta_n}$ as before. Conditional on the Ornstein Uhlenbeck process $x$, the adjusted increments are independent and we can simulate $F_{i,\cdot}, i=1,..., n$ iteratively by simulating $x$ and the increments $\Pi_{10} \tilde \Delta_i^n f$. For the latter we have in distribution (conditional on $x$)
\begin{align*}
\Pi_{10} \tilde \Delta_i^n f
\overset{d}{=} & \sqrt{\int_{(i-1)\Delta_n}^{i\Delta_n} x^2(s)ds}N\left(0, \Pi_{10}Q_a\Pi_{10}\right)+ \Pi_{10}\int_{(i-1)\Delta_n}^{i\Delta_n} \mathcal S(i\Delta_n-s) dJ_s
\end{align*}
where we used that $\mathcal S(t)Q_a\mathcal S(t)^*=Q_a$ for all $t\geq 0$. Moreover, we can identify the covariance $\Pi_{10}Q\Pi_{10}$ with covariance matrix
\begin{equation}
\Pi_{10} Q_1 \Pi_{10}\equiv \left[\int_{(j-1-1)\Delta_n}^{j_1\Delta_n}\int_{(j-2-1)\Delta_n}^{j_2\Delta_n} e^{-a(x-y)^2} dxdy \right]_{j_1,j_2=1,...,1000}.
\end{equation}
To have a good approximation of the integrals $\int_{(i-1)\Delta_n}^{i\Delta_n} x^2(s)ds$ we simulate the square root-process $x$ on a resolution of $10000$, allowing us to make an approximation of the integrals of $x$ with a Riemann sum of length $100$.
For the jump part, we have $J=J_1+J_2$ where $J_1,J_2$ are two $L^2([0,10])$-valued compound Poisson processes, that is
$J_t^i= \sum_{l=1}^{N^i_t} \chi^i_l,$ for $ i=1,2,$ and $t \geq 0$
where $N^i$ are compound Poisson processes with intensities $\lambda_i$ and jumps $\chi_i\sim N(0,Q^{jump}_i)$, where
$Q^{jump}_1= Q_{0.01}$ and $Q^{jump}_2= K$ as described in section (ref). It is then, again since $\mathcal S(t)Q_1\mathcal S(t)^*=Q_1$
$$\Pi_{10}\int_{(i-1)\Delta_n}^{i\Delta_n} \mathcal S(i\Delta_n-s) dJ_s^1\overset{d}{=} \sum_{i=1}^{N_{\Delta_n}^1} \Pi_{10}\chi_i^1$$
and
$$\Pi_{10}\int_{(i-1)\Delta_n}^{i\Delta_n} \mathcal S(i\Delta_n-s) dJ_s^2\overset{d}{=} \sum_{i=1}^{N_{\Delta_n}^2} \Pi_{10}\chi_i^2(\cdot+ \Delta_n-\tau_i)=\sum_{i=1}^{N_{\Delta_n}^2} e^{-(\Delta-\tau_i)}\Pi_{10}\chi_i^2$$
where $\Pi_{10}\chi_i^1\sim N(0,\Pi_{10}Q_{0.01}\Pi_{10})$ and $\Pi_{10}\chi_i^2\sim N(0,\Pi_{10}K\Pi_{10})$ and where $\tau_i$ are the jump times at which $N_{\tau_i}-N_{\tau_i-}>0$. As $\Pi_{10}Q_{0.01}\Pi_{10}$ can be identified with a matrix analogously to (ref) and $\Pi_{10}K\Pi_{10}$, disregarding a normalization constant, with the matrix
\begin{align*}
\Pi_{10}K\Pi_{10}\equiv &\left[\int_{(j_1-1)\Delta_n}^{j_1\Delta_n}\int_{(j_2-1)\Delta_n}^{j_2\Delta_n} e^{-(x+y)} dxdy \right]_{j_1,j_2}
=\left[\left(\frac{1-e^{-10 \Delta_n}}{10}\right)^2 e^{-10(i+j-2)\Delta_n}\right]_{j_1,j_2}
\end{align*}
Hence, the adjusted increments $\Pi_{10} (f_{i\Delta_n}-\mathcal S(\Delta_n)f_{i\Delta_n})$ can be simulated exactly.
\section{Detailed results for the empirical analysis }
We here provide the detailed results for the empirical study of Section (ref) in Tables (ref) and (ref).
\begin{table}\caption{
Columns $2$ to $4$ report the number of jumps detected by the estimator and the ratio of the norms of the truncated estimator $\hat q_i^{*,-}$ to the quadratic variation estimator $\hat q_i^*$, which indicates how large the impact of jumps was on the quadratic variation in each year. The number is bold, if at least one jump was detected. The norms of the quadratic variation estimators are reported in column $5$.
}
\begin{tabular}{ c ccc c }
\toprule
Year &\multicolumn{3}{c}{Trunc. increments,$\frac{\|\hat q_i^{*,-}\|_{L^2}}{\|\hat q_i^*\|_{L^2}}$} & $\|\hat q_i^*\|_{L^2}$ \\
\cmidrule(lr{1em}){2-4}
&$l=3$& $l=4$& $l=5$ &\\
\midrule
1990 & {\bf 1, 0.88} & 0, 1.00 & 0, 1.00 & {\bf 0.00096}\\
1991 & 0, 1.00 & 0, 1.00 & 0, 1.00 & 0.00061\\
1992 & 0, 1.00 & 0, 1.00 & 0, 1.00 & 0.00073 \\
1993 & 0, 1.00 & 0, 1.00 & 0, 1.00 & 0.00067 \\
1994 & {\bf 3, 0.85} & {\bf 2, 0.96} & {\bf 2, 0.96} & {\bf 0.00123} \\
1995 & {\bf 1, 0.97} & 0, 1.00 & 0, 1.00 & {\bf 0.00074}\\
1996 & {\bf 2, 0.89} & 0, 1.00 & 0, 1.00&{\bf 0.00108}\\
1997 & {\bf 2, 0.98} & {\bf 2, 0.98} & 0, 1.00& {\bf 0.00066}\\
1998 & {\bf 15, 0.55} & {\bf 9, 0.65} & {\bf6, 0.69}& {\bf 0.00110}\\
1999 & 0, 1.00 & 0, 1.00 & 0, 1.00&0.00091\\
2000 & 0, 1.00 &0, 1.00 & 0, 1.00& 0.00069\\
2001 & {\bf 2, 0.94} & {\bf 2, 0.94} &{\bf 2, 0.94}&{\bf 0.00116}\\
2002 & {\bf 2, 0.98} & 0, 1.00 & 0, 1.00&{\bf 0.00118}\\
2003 & 0, 1.00 & 0, 1.00 & 0, 1.00& 0.00129\\
2004 & 0, 1.00 & 0, 1.00 & 0, 1.00&0.00087\\
2005 & 0, 1.00 & 0, 1.00 & 0, 1.00&0.00059\\
2006 & {\bf 2, 0.97} & {\bf 2, 0.97} & {\bf 2, 0.97}&{\bf 0.00038}\\
2007 & {\bf 4, 0.99} & 0, 1.00 & 0, 1.00&{\bf 0.00074}\\
2008 & {\bf 4, 0.91} & 0, 1.00 & 0, 1.00&{\bf 0.00230}\\
2009 & {\bf 1, 0.88} & 0, 1.00 & 0, 1.00&{\bf 0.00208}\\
2010 & 0, 1.00 & 0, 1.00 & 0, 1.00& 0.00133\\
2011 & 0, 1.00 & 0, 1.00 & 0, 1.00& 0.00157\\
2012 & 0, 1.00 & 0, 1.00 & 0, 1.00&0.00071\\
2013 & {\bf 2, 0.87} & 0, 1.00 & 0, 1.00&{\bf 0.00079}\\
2014 & 0, 1.00 & 0, 1.00 & 0, 1.00&0.00047\\
2015 & 0, 1.00 & 0, 1.00 & 0, 1.00&0.00085\\
2016 & 0, 1.00 & 0, 1.00 & 0, 1.00&0.00059\\
2017 & 0, 1.00 & 0, 1.00 & 0, 1.00&0.00037\\
2018 & 0, 1.00 & 0, 1.00 & 0, 1.00&0.00033\\
2019 & 0, 1.00 & 0, 1.00 & 0, 1.00&0.00050\\
2020 & {\bf 9, 0.49} & {\bf 3, 0.73} & 0, 1.00&{\bf 0.00109}\\
2021 & {\bf 2, 0.94} & 0, 1.00 & 0, 1.00&{\bf0.00056}\\
2022 & 0, 1.00 & 0, 1.00 & 0, 1.00& 0.001561\\
\bottomrule
\end{tabular}
\end{table}
\begin{table}\caption{Columns $2$ to $5$ report the numbers $D_{C}^{\hat e^{*,i}}(p)$ for $C=\mathcal T_{q_i^{*,-}},$
defined in (ref) of linear factors needed in each year to explain $p=85\%, 90\%, 95\%, 99\%$ of the variation of difference returns as measured by the truncated variation estimators $\hat q_i^{*,-}$ where the truncation rule was conducted with $l=3$ and $\hat e^{*,i}=(\hat e_1^{*,i},\hat e_2^{*,i},...)$ is the basis of eigenfunctions corresponding to the kernel $\hat q_i^*$. Columns $6$ to $9$ report $D_{C}^{\hat e^{long}}(p)$ for $C=\mathcal T_{q_i^{*,-}}$, which explain how many leading eigenvectors of the static estimator $\hat q_{long}^*$ are needed as approximating factors to explain the variation in all years separately.
}
\begin{tabular}{ c cccccccc }
\toprule
Year & \multicolumn{4}{c}{ $D_{\mathcal T_{\hat q_i^{*,-}}}^{\hat e^{*,i}}(p)$}& \multicolumn{4}{c}{$D_{\mathcal T_{\hat q_i^{*,-}}}^{\hat e^{long}}(p)$} \\
\cmidrule(lr{1em}){2-5} \cmidrule(lr{1em}){6-9}
&$0.85$&$0.90$&$0.95$& $0.99$ &$0.85$&$0.90$&$0.95$& $0.99$ \\
\midrule
1990 &4& 6 & 9 & 15&$5$&$7$&$10$& $15$ \\
1991 &5& 6 & 9 & 15&$5$&$7$&$10$& $15$ \\
1992 &5& 6 & 8 & 14&$6$&$7$&$9$& $15$ \\
1993 &4& 5 & 7 & 14&$4$&$5$&$9$& $14$ \\
1994 &3& 5 & 8 & 14&$3$&$5$&$9$& $15$ \\
1995 &4& 5 & 8 & 14&$4$&$6$&$9$& $15$ \\
1996 &4& 5 & 8 & 13&$4$&$6$&$8$& $15$ \\
1997 &3& 4 & 7 & 13&$3$&$4$&$9$& $14$\\
1998 &5& 6 & 8 & 14&$5$&$6$&$10$& $15$ \\
1999 &3& 5 & 8 & 14&$4$&$6$&$10$& $16$ \\
2000 &4& 5 & 8 & 13&$4$&$6$&$10$& $14$\\
2001 &4& 5 & 8 & 13&$5$&$6$&$9$& $13$ \\
2002 &3& 5 & 7 & 12&$4$&$6$&$8$& $14$\\
2003 &2& 4 & 6 & 10&$2$&$4$&$6$& $12$ \\
2004 &2& 3 & 6 & 11&$2$&$4$&$7$& $13$\\
2005 &2& 3 & 5 & 11&$2$&$3$&$8$& $12$ \\
2006 &2& 2 & 4 & 10&$2$&$2$&$7$& $11$ \\
2007 &2& 3 & 6 & 10&$3$&$6$&$10$& $13$ \\
2008 & 3& 4 & 7 & 12&$3$&$5$&$9$& $13$ \\
2009 &3& 4 & 6 & 11&$4$&$5$&$7$& $13$\\
2010 &2& 3 & 6 & 11&$3$&$4$&$7$& $13$ \\
2011 &2& 3 & 5 & 10&$2$&$4$&$6$& $12$\\
2012 &1& 2 & 3 & 8 &$2$&$3$&$4$& $10$\\
2013 &2& 2 & 3 & 8 &$2$&$3$&$5$& $10$\\
2014 &2& 2 & 4 & 10 &$2$&$3$&$6$& $11$\\
2015 & 2& 2 & 3 & 10 &$2$&$2$&$5$& $11$\\
2016 & 2& 2 & 4 & 11 &$2$&$2$&$5$& $12$\\
2017 &2& 2 & 5 & 11 &$2$&$3$&$6$& $12$\\
2018 & 2& 2 & 5 & 12 &$2$&$3$&$7$& $13$\\
2019 &2& 2 & 5 & 11 &$2$&$2$&$7$& $12$\\
2020 &2& 3 & 6 & 12 &$2$&$4$&$8$& $13$\\
2021 &2& 2 & 4 & 10 &$2$&$3$&$6$& $12$\\
2022 &2& 2 & 4 & 8&$2$&$2$&$5$& $11$ \\
\bottomrule
\end{tabular}
\end{table}