Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
82,862 characters · 12 sections · 92 citation commands
Optimal Estimation for General Gaussian Processes
Gaussian processes have been widely applied across a broad range of scientific and applied disciplines, including economics, finance, physics, hydrology, and telecommunications. One of their most extensively studied features is the long-memory property, which captures long-range dependence. The autoregressive fractionally integrated moving average (ARFIMA) model was independently introduced by Granger-1980 and Hosking-1981 to model this feature. In economics and finance, long memory has been examined in a wide array of time series, including real output Diebold-Rudebusch-1989, income Diebold-Rudebusch-1991, stock returns Lo-1991, Liu-Jing-2018, interest rates Shea-1991, exchange rates Diebold-Husted-Rush-1991, Cheung-1993, volatility Ding-Granger-Engle-1993, Andersen-Bollerslev-1997-b, Andersen-Bollerslev-Diebold-Labys-2003 and bubble detection Lui-Phillips-Yu-2024. Various mechanisms have been proposed to explain the emergence of long memory, including cross-sectional aggregation Granger-1980, Abadir-2002, regime switching Potter-1976, Diebold-Inoue-2001, marginalization Chevillon-2018, and network effects Schennach-2018.
More recently, a rapidly growing strand of literature has focused on continuous-time Gaussian processes, which can characterize the local behavior and reproduce the rough sample paths observed in volatility and trading volume Gatheral-Jaisson-Rosenbaum-2018, Fukasawa-Takabatake-Westphal-2022, Wang-Xiao-Yu-2023-JoE, Bolko-Christensen-Pakkanen-Veliyev-2023, Shi-Yu-Zhang-2024-b, Chong-Todorov-2025. Two prominent models in this class are fractional Brownian motion (fBm)Mandelbrot-1965, mandelbrot1968, Gatheral-Jaisson-Rosenbaum-2018 and the fractional Ornstein–Uhlenbeck (fOU) process Cheridito-Kawaguchi-Maejima-2003, Wang-Xiao-Yu-2023-JoE. When applied to data, fractional Gaussian noise (fGn, the first-order difference of fBm) and fOU—exhibit anti-persistence Gatheral-Jaisson-Rosenbaum-2018, Fukasawa-Takabatake-2019, Shi-Yu-Zhang-2024 and short memory Shi-Yu-Zhang-2024-b, Wang-Xiao-Yu-Zhang-2023. Several studies have begun to investigate the micro-level origins of this roughness. For example, eleuch2018 show that in highly endogenous markets, rough volatility may arise from a large number of split orders, while Jusselin-Rosenbaum-2020 demonstrate that rough volatility emerges naturally under the no-arbitrage condition.
The ML estimation method of non-centered stationary discrete- and continuous-time Gaussian models with long memory, short memory, or anti-persistence, referred to as general Gaussian processes, is the focus of this paper. Since these memory properties are relevant across various applications, our goal is to develop an estimation method that does not impose prior restrictions on the memory type of the process. The choice of ML estimation is motivated by the practical need for accurate estimation of all model parameters, particularly when computing impulse response functions or performing forecasts. In such cases, suboptimal estimators---such as the semi-parametric methods of geweke1983, Robinson-1995-LWE, phillips2004, shimotsu2010, the method of moments in Wang-Xiao-Yu-2023-JoE, and the composite likelihood approach in Bennedsen-2024---are not recommended. For example, corsi2009 criticized semi-parametric methods for producing significantly biased and inefficient estimates in forecasting applications with ARFIMA models. Moreover, although the ML and Whittle ML estimators are asymptotically equivalent, the ML estimation method generally demonstrates superior finite-sample performance Cheung-Diebold-1994.\footnote{Rao-Yang-2021 propose new frequency-domain quasi-likelihoods that improve the finite-sample behavior of the Whittle ML estimator for short-memory Gaussian processes.}
Considerable progress has been made in developing ML estimation methods, extending from specific parametric models to general Gaussian processes based on discrete-time observations.\footnote{A parallel literature studies parameter estimation for continuous-time fractional models under continuous-record observations; see Kleptsyna2002.} For example, Yajima-1985 established the consistency and asymptotic normality of the MLE for the ARFIMA$(0,d,0)$ model with $d \in (0, 0.5)$, representing the long-memory case. These results were subsequently extended to stationary Gaussian processes with long memory by Dahlhaus-1989, Dahlhaus-2006, and further generalized to general Gaussian processes by Lieberman-2012. However, these existing methods adopt a two-stage procedure in which $\mu$ is first estimated by the sample mean and then substituted into the likelihood, yielding a so-called plug-in MLE for the remaining parameters. Some implementations apply the MLE to demeaned data Tsai-Chan-2005, shi2022volatility, which is essentially equivalent to the plug-in MLE procedure. Regarding the optimality of this procedure, the sample mean is clearly not the efficient estimator for $\mu$. While Dahlhaus-1989, Dahlhaus-2006 argued for the efficiency of the plug-in MLE by showing that its asymptotic covariance matrix equals the inverse of the Fisher information matrix, they did not explicitly establish the existence of a Crame\'r--Rao lower bound. Cohen-Gamboa-Lacaux-Loubes-2013 derived the LAN property for centered stationary Gaussian processes with long memory, short memory, or anti-persistence, providing a minimax lower bound for general estimators; however, their results do not imply the asymptotic efficiency of plug-in MLE. Another concern is that the inefficiency in estimating $\mu$ may impair the finite-sample performance of the plug-in MLE. Cheung-Diebold-1994 showed that when the $\mu$ is unknown, the finite-sample performance of the MLE for the other parameters deteriorates, even though their asymptotic variances remain the same as in the known-mean case. To date, the problem of obtaining theoretically optimal estimators for all parameters for general Gaussian processes within the ML framework remains unresolved. Only one exception is Wang-Xiao-Yu-Zhang-2023, who established the consistency and asymptotic normality of the MLE for all parameters—including $\mu$—in the fOU process. Nonetheless, their framework is model-specific and not readily applicable to other fractional models. Moreover, the efficiency of the MLE was not addressed.
We introduce a novel exact ML method, a term we adopt to distinguish it from the plug-in ML method, for general Gaussian processes, where all the parameters are estimated jointly. We establish the consistency and asymptotic normality of the exact MLE. These results extend those in Wang-Xiao-Yu-Zhang-2023 from fOU to general Gaussian processes. Moreover, we establish the LAN property of the sequence of statistical experiments in the Le Cam sense. This result extends that in Cohen-Gamboa-Lacaux-Loubes-2013 from centered stationary Gaussian processes to “non-centered” ones within the long span asymptotics, which is different from several extensions under the high-frequency asymptotics recently made by Brouste-Fukasawa-2018,Fukasawa-Takabatake-2019,Szymanski-2024,Szymanski-Takabatake-2023-Additive,Chong-Mies-2025. The LAN property directly yields the efficiency of the exact MLE. The results rely solely on the asymptotic behavior of the spectral density near zero for a discrete record of observations, which allows for broad applicability. \footnote{To ensure our theoretical results are broadly applicable and consistent with discrete observations, we work with the spectral density corresponding to discrete-time data, regardless of whether the underlying process is continuous- or discrete-time. In TYZ (working paper), we demonstrate how to verify the necessary conditions for the discrete spectral density using the continuous-time spectral density for a wide range of continuous-time processes.}
To demonstrate the practical applicability of the proposed method, we conduct three Monte Carlo simulation studies in which our exact ML estimator is applied to three widely studied non-centered processes: the ARFIMA$(0,d,0)$ model, fractional Gaussian noise (fGn), and the fractional Ornstein–Uhlenbeck (fOU) process. Overall, our exact estimator for $\mu$ improves upon the performance of the sample mean. Regarding the performance of plug-in MLE, our simulation results show that the plug-in MLE performs nearly as well as the exact MLE, alleviating concerns that inefficient estimation of $\mu$ would compromise the efficiency of the remaining parameter estimates. In particular, for the ARFIMA$(0,d,0)$ model, the gain in efficiency for estimating $\mu$ aligns with the theoretical result of Adenstedt-1974. We also conduct a forecasting horse race for realized volatility using the fOU process with three alternative estimators: the exact MLE, the plug-in MLE, and the change-of-frequency (CoF) estimator by Wang-Xiao-Yu-2023-JoE. As expected, the exact MLE delivers the best forecasting performance, followed by the plug-in MLE, which performs satisfactorily, and then the CoF estimator.
To sum up, we contribute to the literature in the following aspects. First, we propose a novel exact ML estimation method for all parameters in a general stationary Gaussian process, establishing its consistency and asymptotic normality. Second, we prove the LAN property of the sequence of statistical experiments for general stationary Gaussian processes in the Le Cam sense, providing a theoretical foundation of optimal estimation. The LAN property we have established is also essential for building asymptotic optimality of statistical tests and selecting the order of models based on the likelihood function, see Remark (ref) for further references. Third, our method serves as a benchmark for evaluating the finite-sample performance of the existing plug-in MLE. Although the performance gap between the plug-in MLE and the MLE with known \( \mu \) can be substantial in finite samples Cheung-Diebold-1994, our analysis shows that this difference is not driven by inefficiencies in estimating \( \mu \).
The remainder of this paper is organized as follows. Section (ref) presents the exact ML method and develops asymptotic Properties. In Section (ref), we provide several examples to which our results can be applied. Section (ref) presents a Monte Carlo study to assess the performance of the estimation method. Section (ref) concludes the paper.
Let $\Theta_{\xi}$ be a convex domain of $\mathbb{R}^{p-1}$ with compact closure and set $\Theta:=\Theta_{\xi}\times(0,\infty)$. Write $\theta=(\xi,\sigma)^{\top}\in\Theta$ and $\vartheta=(\theta,\mu)^{\top}=(\xi,\sigma,\mu)^{\top}\in\Theta\times\mathbb{R}$. Denote by $\partial_{z}=\partial/\partial{z}$, $\partial_{\omega}=\partial/\partial{\omega}$ and $\partial_{j} := \partial/\partial{\theta_{j}}$ for $j\in\{1,\cdots,p+1\}$. For notational simplicity, $\partial_{0}$, $\partial_{z}^{0}$ and $\partial_{\omega}^{0}$ denote the identify operator. The derivative operators $\partial_{j_{1},\cdots,j_{k}}^{k}$ are recursively defined by $\partial_{j_{1},\cdots,j_{k}}^{k} := \partial_{j_{1}} \circ \partial_{j_{2},\cdots,j_{k}}^{k-1}$ for $j_{1},\cdots,j_{k}\in\{0,1,\cdots,p+1\}$ and $k\in\mathbb{N}$. Moreover, $\mathbf{1}_{n}$ denotes a $n$-dimensional vector whose all elements are equal to $1$ and, for an integrable function $f$ on $[-\pi,\pi]$, $\Sigma_{n}(f)$ denotes the symmetric Toeplitz matrix whose $(i,j)$th elements are equal to the $(i-j)$th Fourier coefficients of $f$.\\
Let us consider a stationary Gaussian time series $\{X^{\vartheta}_{j}\}_{j\in\mathbb{Z}}$ with mean $\mu$ and spectral density function $s^{X}(\omega,\theta)$. We may write $s_{\theta}^{X}(\omega):=s^{X}(\omega,\theta)$. Let us denote by $\vartheta_{0}=(\xi_{0},\sigma_{0},\mu_{0})^{\top}$ an interior point of $\Theta\times\mathbb{R}$, which may call a true value of the parameter $\vartheta$, and we assume that we observe a realization of $X_{1}^{\vartheta_{0}},\cdots,X_{n}^{\vartheta_{0}}$. Let $s_{\xi}^{X}(\omega)= s_{\theta}^{X}(\omega)/\sigma^2$. Then, for each $\vartheta=(\theta,\mu)^{\top}\in\Theta\times\mathbb{R}$ and $n\in\mathbb{N}$, we denote by $\mathbb{P}_{\vartheta}^{n}$ a distribution on the Borel space $(\mathbb{R}^{n},\mathcal{B}(\mathbb{R}^{n}))$ under which a random vector $\mathbf{X}_{n} =(X_{1},\cdots,X_{n})^{\top}$ follows a $n$-dimensional Gaussian vector with mean vector $\mu\mathbf{1}_{n}$ and variance-covariance matrix $\Sigma_{n}(s_{\theta}^{X})$. We also denote by $\gamma_{\theta}^{X}(\cdot)$ the autocovariance function of $X^{\vartheta}$. \\
Denote by $\ell_{n}(\vartheta)\equiv\ell_n(\xi,\sigma,\mu)$ the Gaussian log-likelihood function of the observations $\mathbf{X}_{n}$ under the distribution $\mathbb{P}_{\vartheta}^{n}$, which is given by
and then the maximum likelihood estimator (MLE)\footnote{We have derived alternative expression of the likelihood function using the conditional likelihood based on the Bayes formula, which improve the computational efficiency by avoiding direct computations of the inverse and determinant of large-scale covariance matrix. See (ref) for details.} is defined by
Notice that, from the definition of the MLE, the MLE satisfies the estimating equations
which conclude the equations
hold for any $(\xi,\sigma,\mu)^{\top}\in\Theta_{\xi}\times(0,\infty)\times\mathbb{R}$. Then the MLE $\widehat{\xi}_{n}^{\mathrm{MLE}}$ would be a maximizer of the function $\bar{\ell}_{n}(\xi):=\ell_{n}(\xi,\bar{\sigma}_{n}(\xi),\mu_{n}(\xi))$, where $\bar{\sigma}^{2}_{n}(\xi):=\sigma_{n}^{2}(\xi,\mu_{n}(\xi))$ and $\bar{\sigma}_{n}(\xi):=\sqrt{\bar{\sigma}^{2}_{n}(\xi)}$, over the parameter $\xi\in\Theta_{\xi}$ and the MLEs $\widehat{\mu}_{n}^{\mathrm{MLE}}$ and $\widehat{\sigma}_{n}^{\mathrm{MLE}}$ satisfy the equations
Therefore, we define our proposed estimator $\widehat{\vartheta}_{n} :=(\widehat{\xi}_{n},\widehat{\sigma}_{n},\widehat{\mu}_{n})^{\top}$ by
and we call $\widehat{\vartheta}_{n} =(\widehat{\xi}_{n},\widehat{\sigma}_{n},\widehat{\mu}_{n})^{\top}$ the {\it exact MLE} throughout this paper. In subsequent sections, we investigate the asymptotic properties of the exact MLE.
We firstly introduce several conditions on the spectral density function $s_{\theta}^{X}(\omega)$ summarized in the following assumption that is used to prove asymptotic properties of the exact MLE and the likelihood ratio process; see Sections (ref) and (ref) for details.
Assumption (ref) is the usual conditions on the “discrete-time” spectral density function for stationary Gaussian time series with long/short/anti-persistent memory used in the literature; see the assumptions in Fox-Taqqu-1986, Dahlhaus-1989, Dahlhaus-2006, Lieberman-2012, Cohen-Gamboa-Lacaux-Loubes-2013 and Fukasawa-Takabatake-2019 as references. $X_j^{\vartheta}$ is said to have long memory (or long-range dependence) if $0 < \alpha_{X}(\xi) < 1$, short memory if $\alpha_{X}(\xi) = 0$, and anti-persistence if $\alpha_{X}(\xi)< 0$. The range $\alpha_{X}(\xi) \leq -1$ corresponds to noninvertibility, and our results cover this case as well. Since these memory properties are relevant across various applications, we do not impose prior restrictions on the memory type of the process.
Before stating our main results, we introduce additional notation. We write $\widehat{\theta}_{n}:=(\widehat{\xi}_{n},\widehat{\sigma}_{n})^{\top}$ and then we may write $\widehat{\vartheta}_{n}=(\widehat{\theta}_{n},\widehat{\mu}_{n})^{\top}$. Define $p\times p$ dimensional matrix $\mathcal{F}_{p}(\theta)$ by
where
\color{black}
Then we assume the following condition on $\mathcal{F}_{p}(\xi)$.
Our first main result is a (weak) consistency and an asymptotic normality of the sequence of the MLEs $\{\widehat{\theta}_{n}\}_{n\in\mathbb{N}}$ defined in (ref), that is a generalization of, for example, Theorems 3.1 and 3.2 in Dahlhaus-1989 and Theorem 1 in Lieberman-2012 to the case of general Gaussian processes using the multi-step estimation procedure based on the exact MLE defined in (ref).
See Section (ref) for the proof of consistency and Section (ref) for the proof of asymptotic normality.
To prove the asymptotic normality of the MLEs for the joint estimation $\vartheta=(\theta,\mu)^{\top}$ as well as the MLE for $\mu$, we need to further assume the precise asymptotic behavior of the spectral density function $s_{\theta}^{X}(\omega)$ around the frequency $\omega=0$ given in the following assumption.
Based on $\mathcal{F}_{p}(\xi)$ defined in (ref), we further define
where $B(\alpha,\beta)$ is the beta-function. Note that, under Assumption (ref), the matrix $\mathcal{I}(\xi)$ is also invertible for each $\xi\in\Theta_{\xi}$. Moreover, we also introduce a normalized score function $\zeta_{n}(\vartheta)$ and an observed Fisher information matrix $\mathcal{I}_{n}(\vartheta)$, respectively defined by
Now we can state our second main result of the asymptotic normality of the MLE for the joint parameter $\vartheta=(\theta,\mu)^{\top}$, summarized in the following theorem with its proof given in Section (ref).
The LAN property is a fundamental concept in asymptotic statistics, originally introduced by Wald-1943 and further developed by LeCam-1960. It plays a central role not only in establishing the asymptotic optimality of estimators but also in facilitating statistical inference. In this section, we establish the LAN property of the sequence of statistical experiments ${(\mathbb{R}^{n}, \mathcal{B}(\mathbb{R}^{n}), {\mathbb{P}{\vartheta}^{n}}{\vartheta \in \Theta \times \mathbb{R}})}_{n \in \mathbb{N}}$ for general Gaussian processes. We then provide several results concerning the local asymptotic minimax efficiency of the exact MLE. Additional applications of the LAN property are also discussed.
The proof of Theorem (ref) is left to Section (ref). In the relevant literature, Theorem 2.4 in Cohen-Gamboa-Lacaux-Loubes-2013 provides the LAN property for centered stationary Gaussian time series with long/short/anti-persistent memory property under Assumptions (ref) and (ref). Theorem (ref) is a generalization of Theorem 2.4 in Cohen-Gamboa-Lacaux-Loubes-2013 to general Gaussian processes within the long span asymptotics, which is different from several extensions under the high-frequency asymptotics recently made by Brouste-Fukasawa-2018,Fukasawa-Takabatake-2019,Szymanski-2024,Szymanski-Takabatake-2023-Additive,Chong-Mies-2025.
The LAN property is applied to derive asymptotically efficient rates and variances for estimating $\vartheta=(\theta,\mu)^{\top}$. These results are derived from lower bounds of estimation provided by H\'ajek's convolution theorem (see Hajek-1972, Ibragimov-Hasminski-1981,LeCam-1972), recalled below.
Notice that we have already proved that the sequence of the exact MLEs $\{\widehat{\vartheta}_{n}\}_{n\in\mathbb{N}}$ defined in (ref) satisfies the coupling property (ref) in Theorem $\ref{Thm:MLE2}$ so that, using the result in Section 7.12.(b) of Hopfner-2014-Book in addition to Theorem (ref), we can conclude that the sequence of exact MLEs is asymptotically efficient in the local asymptotic minimax sense as well as in the Fisher sense, and then it attains the local asymptotic minimax bound of estimation given in Theorem (ref). We summarize the aforementioned result in the following corollary.
As an alternative estimator of $\theta$ in the literature, the plug-in MLE (PMLE) is defined by
using some compact set $\Theta_{\ast}\subset\Theta$ and an estimator $\widetilde{\mu}_{n}$. Under similar assumptions to Assumption (ref), Dahlhaus-1989, Dahlhaus-2006 and Lieberman-2012 show that the plug-in MLE is weakly consistent, asymptotically normal with the asymptotic variance $\mathcal{F}_{p}(\theta)^{-1}$ at the point $\vartheta_{0}$? (should be $\theta_0$) when the plug-in estimator $\widetilde{\mu}_{n}$ satisfies the assumption
The assumption (ref) corresponds to Assumption 5 of Dahlhaus-1989 for the long memory case $\alpha_{X}(\xi_{0})\in(0,1)$ and Assumption 5 of Lieberman-2012 for the long/short/anti-persistent memory case $\alpha_{X}(\xi_{0})\in(-\infty,1)$. The plug-in MLE shares the same convergence rate and asymptotic variance as our exact MLE, implying that it is also asymptotically efficient.
Notice that Theorem (ref) combining with Theorem (ref) yields the asymptotic minimax lower bound
for any $r>0$ and any sequences of estimators $\widetilde{\mu}_{n}$, which implies the convergence rate of the estimator satisfying the assumption (ref) is the same as that of the optimal estimators given in (ref), that is, the minimax optimal rate. Although Adenstedt-1974 proved that the best linear unbiased estimator (BLUE) of $\mu$ satisfies the assumption (ref) for all $\alpha_{X}(\xi_{0})\in(-\infty,1)$, the BLUE of $\mu$ is infeasible without knowing the true value of $\xi_{0}$. Although the assumption (ref) is satisfied by the widely used sample mean, its asymptotic variance does not achieve the minimax optimal bound. For example, Adenstedt-1974 quantified the asymptotic inefficiency of the sample mean for the ARFIMA$(0,d,0)$ model by computing the ratio of the asymptotic variance of an efficient estimator for $\mu$ to that of the sample mean:
This metric is always less than $1$, except when $d = 0$. This plug-in approach is equivalent to the implementation that applies the MLE to demeaned data. Lieberman-2012 proposed an alternative estimator of the form
where $s_\ast:=s_{\theta_\ast}$ with any $\theta_\ast=(\xi_\ast,\sigma_\ast)^\top\in\Theta$ satisfying $\alpha(\xi_\ast)=\inf_{\xi\in\Theta_\xi}\alpha(\xi)$ (by compactness of $\Theta_\xi$ there exists at least one such value), or even $s_\ast(\omega):=(1-\cos(\omega))^{\frac{\alpha_\ast}{2}}$ with $\alpha_\ast\leq\inf_{\xi\in\Theta_\xi}\alpha(\xi)$, and proved that the estimator in (ref) satisfies the assumption (ref) for all $\alpha_{X}(\xi_{0})\in(-\infty,1)$ using the results in Adenstedt-1974. This estimator also fails to attain the minimax optimal bound, as it inherently relies on a misspecified structure of the autocovariance function embedded in the Toeplitz matrix $\Sigma_n(s_\ast)$.
Our assumptions on the spectral density are very general and apply to many well-known processes, including but not limited to the ARFIMA$(p,d,q)$ process with $|d|<1/2$, the fractional Gaussian noise with an unknown mean and the fOU process.\footnote{For other continuous-time processes—such as the continuous-time autoregressive fractionally integrated moving average (CARFIMA) model—we refer the reader to Tetsuya et al. (WP) for a detailed analysis of their spectral densities.}
The {\it non-centered} ARFIMA$(p,d,q)$ process was introduce by Granger-1980 and Hosking-1981 independently. For notation simplicity, we start with the ARFIMA$(0,d,0)$ model. The non-centered ARFIMA$(0,d,0)$ model is specified as
where $L$ is the lag operator, $(1-L)^{-d}$ is the fractional difference operator with the memory parameter $d$ and $\epsilon_{t}\overset{\mathrm{iid}}{\sim} N(0,1)$. It reduces to a Gaussian white noise when $d=0$. When $d\in (-1/2,1/2)$, the ARFIMA$(0,d,0)$ process is stationary and invertible bloomfield1985. Let $u_{t}:=(1-L)^{-d}\epsilon _{t}$ be the fractionally integrated process and $\gamma _{u}(j):=\mathrm{Cov}[u_{t},u_{t-j}]$ be its $j$th order auto-covariance. According to Hosking-1981, the auto-covariance function of $ u=\{u_{t}\}_{t\in\mathbb{Z}}$ is expressed by
The long-run variance covariance $\sum_{j=-\infty }^{\infty }\gamma _{u}(j)$ $= \infty $ when $d\in \left( 0,1/2\right) $ and $\sum_{j=-\infty }^{\infty }\gamma_{u}(j)=0$ when $d\in \left( -1/2,0\right) $. Therefore, $ u_{t}$ has a long memory if $d\in \left( 0,1/2\right) $ and is anti-persistent if $d\in \left( -1/2,0\right) $. The spectral density of the model is given by \[s_{\theta}^X(\omega)=\frac{\sigma^2}{2\pi} \sim \frac{\sigma^2}{2\pi}|\omega|^{-2d}\ \ \text{ as }|\omega|\rightarrow 0.\] In this case, $\alpha_X(\xi)=2d, c_X(\xi)=(2\pi)^{-1}$.
Let $p,q\in\mathbb{N}\cup\{0\}$ and $\xi:=(\phi_1,\ldots,\phi_p,\psi_1,\ldots,\psi_q)\in\mathbb{R}^{p+q}$. The {\it non-centered} ARFIMA$(p,d,q)$ process is defined by
where $\phi_\xi(z):=1-\phi_1z-\cdots-\phi_pz^p$ and $\psi_\xi(z):=1+\psi_1z+\cdots+\psi_qz^q$. Assume that for each \( \xi \), the functions \( \phi_\xi(z) \) and \( \psi_\xi(z) \) have no common roots in \( \mathbb{C} \), and that all their roots lie outside the unit circle. This implies that \( \phi_\xi(z) \neq 0 \) and \( \psi_\xi(z) \neq 0 \) for \( |z| \leq 1 \). Then, for \( |d| < 1/2 \), the difference equation in (ref) admits a unique stationary solution \( X = \{X_t\}_{t\in\mathbb{Z}} \) of the form \[ X_t = \mu + \sigma \phi_\xi(L)^{-1} \psi_\xi(L) (1-L)^{-d} \epsilon_t, \quad t \in \mathbb{Z}. \]
The spectral density function of ARFIMA$(p,d,q)$ is given by \[ s_{\theta}^X(\omega) = \frac{\sigma^2}{2\pi} |1 - e^{-i\omega}|^{-2d} \frac{|\psi_\xi(e^{-i\omega})|^2}{|\phi_\xi(e^{-i\omega})|^2}. \] for $\omega \in (-\pi, \pi]$. It can be shown that \[ s_{\theta}^X(\omega) =\frac{\sigma^{2}}{2\pi}\left( 2-2\cos\left( \lambda\right) \right) ^{-d}\frac{\left( 1+\sum_{j=1} ^{q}\psi_{j}\cos\left( j\omega\right) \right) ^{2}+\left( \sum_{j=1} ^{q}\psi_j\sin\left( j\omega\right) \right) ^{2}}{\left( 1-\sum_{j=1}^{p} \phi_{j}\cos\left( j\omega\right) \right) ^{2}+\left( \sum_{j=1} ^{p}\phi_{j}\sin\left( j\omega\right) \right) ^{2}}. \] Since the assumption \( \phi_\xi(z) \neq 0 \) for \( |z| \leq 1 \) ensures that the spectral density function of the ARMA\( (p,q) \) process, \[ f_{\mathrm{ARMA}}(\omega) := \frac{\sigma^2}{2\pi} \frac{|\psi_\xi(e^{-i\omega})|^2}{|\phi_\xi(e^{-i\omega})|^2}, \] is bounded away from zero on \( [-\pi, \pi] \), the singularity of the spectral density of \( X \) in the vicinity of zero frequency is governed by the ARFIMA\( (0,d,0) \) factor \( |1 - e^{-i\omega}|^{-2d} \). As \( \omega \to 0 \), it exhibits the asymptotic behavior \[ s_{\theta}^X(\omega) \sim \sigma^2c_X(\theta)|\omega|^{-\alpha_X(\xi)}. \]
In this case, $\alpha_X(\xi) = 2d, c_X(\xi) = (2\pi)^{-1} \left| \frac{\psi_\xi(1)}{\phi_\xi(1)} \right|^2$ in Assumptions (ref) and (ref). Hence, our results are applicable to the non-centered Gaussian ARFIMA$(p, d, q)$ process. According to Theorem (ref), when $d < 0$, the convergence rate for the exact MLE of $\mu$ is $\frac{1}{2}(1 - \alpha_X(\xi)) > 1/2$, indicating super-consistency.
The fractional Brownian motion (fBm) with Hurst index $H\in(0,1)$, denoted by $B^H = \{B_t^H\}_{t \in \mathbb{R}}$, is a unique centered Gaussian process that is almost surely equal to zero at $t = 0$ and possesses both stationary increments and $H$-self-similarity properties. Specifically, these properties are expressed as
for any $s, t \in \mathbb{R}$ and $c > 0$, where $\overset{d}{=}$ denotes equality in distribution. mandelbrot1968 demonstrated that the fBm can be represented as a causal moving average process involving the past differential increments of a (two-sided) standard Brownian motion $B = \{B_t\}_{t \in \mathbb{R}}$. This representation is given by
where $\Gamma(x)$ denotes the gamma function, \footnote{ This is also referred to as the Type I fBm.} which implies that the fBm reduces to the standard Brownian motion when $H=0.5$.
The fBm is a Gaussian process with mean zero and covariance
The increment of fBm is fGn and denoted by $y_{t}$. Using discrete time notations, we have
We extend the fGn to the non-centered case by defining its expectation as $\mu$. The process is given by:
The auto-covariance function of $X$ is, $\forall k\geq 0$,
where the approximation is based on the Taylor expansion. When $H\in (0.5,1)$, the asymptotic behavior in (ref) implies that the sequence of the auto-covariances of fGn is not absolutely summable so that the fGn has the long memory property. When $H\in (0,0.5)$, we can also verify that $\forall k\neq 0$, $ \mathrm{Cov}[X_{t},X_{t+k}] <0$ and $\sum_{k=-\infty }^{\infty}\mathrm{Cov}[X_{t},X_{t+k}] =0$ so that the fGn has the anti-persistent memory property.
The spectral density of (non-centered) fGn is given by sinai1976:
where $C_{H}:=\left( 2\pi \right) ^{-1}\Gamma (2H+1)\sin (\pi H)$. It can be shown that \[s_{\theta}^X(\omega ) \sim \sigma^2C_H|\omega|^{1-2H},\,\, \text{when}\,\, \omega \rightarrow 0.\] In this case, the functions $\alpha_X(\xi)$ and $c_X(\xi)$ in Assumptions (ref) and (ref) are $\alpha_X(\xi)=2H-1$ and $c_X(\xi)=C_H$. Hence, our results are applicable to the non-centered fractional Gaussian noise. According to Theorem (ref), when $H<0.5$, the convergence rate for the exact MLE of $\mu$ is $\frac{1}{2}(1-\alpha_X(\xi))>1/2$, implying super-consistency. This result echoes that for ARFIMA.
The fractional Ornstein-Uhlenbeck (fOU) process is an extension of the classical Ornstein-Uhlenbeck (OU) process, where the driving noise is replaced by a fBm with Hurst index $H \in (0,1)$. This process is particularly useful for modeling systems that exhibit long-range dependence and locally self-similarity, which cannot be captured by the classical OU process. The stationary fOU process with a long-run mean $\mu$ has applications in various fields, including mathematical finance, physics, and time series modeling.
The fOU process $Y=\{Y_t\}_{t\in\mathbb{R}}$ is defined by a unique solution of the following linear SDE (Stochastic Differential Equation):
with initial condition $Y_0$, where $B^H = \{B_t^H\}_{t\in\mathbb{R}}$ is a fBm with Hurst index $H$. The explicit solution of this SDE is given by
where the above stochastic integral can be interpreted as the pathwise Riemann-Stieltjes integral or the Wiener integral associated with the fBm for any $H\in(0,1)$. The fOU process reduces to the classical OU process when $H=0.5$ and to the fBm when $\kappa =0$. When $\kappa >0$, the stationary solution of the SDE (ref), denoted by $\bar{Y}=\{\bar{Y}_t\}_{t\in\mathbb{R}}$, is given by
For $t\geq 0$, the unique solution of the SDE (ref) with the initial condition
is exactly equal to the stationary solution $\bar{Y}$ in (ref), and then the error between $\bar{Y}_t$ and $Y_t$ with arbitrary initial condition $Y_0$ is expressed by
which implies that the error between the solutions (ref) and (ref) converges to zero exponentially as $t\to\infty$ for arbitrary initial condition $Y_0$. In the rest of this section, we consider the case where a data generating process is the discretely and equidistantly observed time series from the stationary solution given in (ref).
Consider a stationary time series $X=\{X_j\}_{j\in\mathbb{Z}}$ of the form $X_j:=Y_{j\Delta}$ for $j\in\mathbb{Z}$ with the sampling frequency $\Delta$. Notice that the time series $X$ is stationary and its auto-covariance is available from garnier2018 when $\kappa >0$:
From Cheridito-Kawaguchi-Maejima-2003, the autocovariance function of the fOU process exhibits the same order of decay as fGn, decaying hyperbolically for $H \neq 1/2$. Cheridito-Kawaguchi-Maejima-2003 and Hult-2003-PhDThesis provide the spectral density function of the stationary solution $\bar{Y}$ given by
For discrete observations $X$, the spectral density is given by Hult-2003-PhDThesis
It can be shown that \[ s_{\theta}^X(\omega) \sim
\] Shi-Yu-Zhang-2024-b. Hence, for fOU, we have \[ \alpha_X(\xi) =
\] and \[ c_X(\xi) =
\] in Assumptions (ref) and (ref). Hence, our results are applicable to the fOU process. However, the function $\alpha_X(\xi)$ exhibits a sharp contrast compared to that of ARFIMA and fGn. According to Theorem (ref), the convergence rate for the exact MLE of $\mu$ in fOU is $\sqrt{n}$ when $H\leq 1/2$, as recently reported in Wang-Xiao-Yu-Zhang-2023.
We consider three data-generating processes (DGPs): ARFIMA$(0,d,0)$, fGn and fOU. For simplicity, the long-run mean $\mu$ is set to 0. The results for $\mu \pm 1$ are provided in Appendix (ref). The parameter $d$ in the ARFIMA$(0,d,0)$ model takes 9 values: $\{-0.4, -0.3, \dots, 0, \dots, 0.4\}$, while the parameter $H$ in the fGn and fOU takes 9 values: $H = 0.1, 0.2, \dots, 0.9$. The fOU has an additional parameter $\kappa=10$. The scale parameter $\sigma$ is set to $1$ in all cases. For each DGP, we compare our exact MLE with the two MLEs considered in Cheung-Diebold-1994. MLE1 refers to the case where $\mu$ is known, MLE2 refers to our exact MLE, and MLE3 refers to the plug-in MLE, where the sample mean is used as an estimator for $\mu$. The number of replications is set to 1000. The sample size is set to $250$ or $1000$. Reported are the bias, standard error (Std), and root mean squared error (RMSE) across all replications for each method.
Table (ref) reports the results for ARFIMA and Table (ref) reports the results for fGn. From Tables (ref)-(ref), we observe the following findings. First, in terms of convergence rates, the performance of MLE2 aligns well with our asymptotic theory. For $\mu$, the convergence rate is $n^{-(1 - \alpha_X(\xi))/2}$, which becomes slower as $d$ increases toward $1/2$ from $-1/2$ (or as $H$ increases toward $1$ from $0$). A similar pattern can be observed for the sample mean, as it shares the same convergence rate as given in ((ref)) and ((ref)). For the remaining parameters, the convergence rate remains at the root-$n$. Second, MLE2 always performs better than MLE3, except when $d = 0$ in ARFIMA or $H = 1/2$ in fGn, where the two methods perform nearly identically. Third, interestingly, this superior performance in estimating $\mu$ by MLE2 does not translate into better performance in estimating other parameters. Using the true value of $\mu$, MLE1 does not lead to a better performance in estimating other parameters. The three ML methods lead to a similar finite sample performance for parameters other than $\mu$.
To see how the relative inefficiency of the sample mean over MLE2 of $\mu$, the two dashed lines in Figure (ref) plot the ratio of the sample variance of MLE2 for $\mu$ to that of MLE3 as a function of $d$ for the two models when $n = 1000$. Clearly, the relative inefficiency goes up rapidly as $d$ nears $-0.5$ in ARFIMA and fGn. For comparison, also plotted by the solid line is the theoretical asymptotic inefficiency given in ((ref)). The red dashed line is closely aligned with the theory.
Tables (ref)-(ref) report the results for fOU, from which we observe the following findings. First, in terms of convergence rates, the performance of MLE2 aligns well with our asymptotic theory. For $\mu$, the convergence rate is root-$n$ when $H \leq 1/2$, and transitions to $n^{1 - H}$ when $H > 1/2$. The standard deviation of the estimator for $\mu$ decreases substantially as the sample size increases from $250$ to $1000$ when $H \leq 1/2$; however, as $H$ approaches $1$, the percentage reduction becomes markedly smaller. A similar pattern can be observed for the sample mean, as it shares the same convergence rate as given in ((ref)) and ((ref)). For the remaining parameters, the convergence rate remains at the root-$n$. Second, we see a clear dominance of MLE2 over MLE3 in terms of finite sample performance of estimates of $\mu$, when both $n$ is large and $H$ is near either zero or one. Third, this superior performance in estimating $\mu$ by MLE2 does not translate into a better performance in estimating other parameters. Using the true value of $\mu$, MLE1 does not lead to better performance in estimating the other 3 parameters. The three ML methods lead to a similar finite sample performance for parameters other than $\mu$.\footnote{In the online supplement (Section (ref)), we conduct a forecasting horse race for realized volatility using the fOU process with three alternative estimators: MLE2, MLE3, and the CoF estimator. As expected, MLE2 delivers the best forecasting performance, followed by MLE3 and then the CoF estimator.}
Gaussian processes have gained significant attention due to their broad applicability across various scientific and applied disciplines. To obtain the MLE, two common approaches are typically employed. The first approach maximizes the likelihood assuming \( \mu \) is known and set to $0$, which results in an unrealistic MLE. The second approach uses the sample mean as an estimator for \( \mu \), leading to a plug-in MLE. However, both methods fail to address the inefficiency of the estimator for \( \mu \), and concerns have been raised about the finite sample performance of the plug-in MLE. Adenstedt-1974 proposed an efficient but infeasible estimator for \( \mu \).
In this paper, we introduce a novel exact ML method for all the parameters for general Gaussian processes with long-memory, short-memory, or anti-persistence properties. We prove that the exact MLE exhibits the properties of consistency and asymptotic normality. We also establish the LAN property of the sequence of statistical experiments for general Gaussian processes in the sense of Le Cam, which directly yields efficiency. Our method offers a comprehensive understanding of MLE for fractional Gaussian models: first, we show that the estimators for all parameters are optimal, effectively complementing the infeasible estimator for $\mu$ proposed by Adenstedt-1974; second, we evaluate the difference between the plug-in MLE, exact MLE, and the MLE with known $\mu$. The plug-in MLE performs as good as the exact MLE for all parameters except for \( \mu \). The discrepancy between plug-in MLE and the MLE with known $\mu$ is not due to an inefficient estimator for \( \mu \).
The Whittle MLE is asymptotically equivalent to the exact MLE under certain regularity conditions. Although its finite-sample performance is generally inferior to that of the exact MLE, the performance gap narrows as the sample size increases. At the same time, the computational burden of the exact MLE increases substantially due to the need to invert the covariance matrix at each evaluation of the likelihood function—a step that the Whittle method avoids. The Whittle ML method remains an attractive alternative. However, existing theoretical results for the Whittle MLE primarily pertain to ARFIMA models and do not extend to continuous-time models. In future work, we aim to investigate the optimality of the Whittle MLE for general Gaussian processes.