EconBase
← Back to paper

Optimal Estimation for General Gaussian Processes

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

82,862 characters · 12 sections · 92 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Optimal Estimation for General Gaussian Processes

abstractThis paper proposes a novel exact maximum likelihood (ML) estimation method for general Gaussian processes, where all parameters are estimated jointly. The exact ML estimator (MLE) is consistent and asymptotically normally distributed. We prove the local asymptotic normality (LAN) property of the sequence of statistical experiments for general Gaussian processes in the sense of Le Cam, thereby enabling optimal estimation and facilitating statistical inference. The results rely solely on the asymptotic behavior of the spectral density near zero, allowing them to be widely applied. The established optimality not only addresses the gap left by Adenstedt-1974, who proposed an efficient but infeasible estimator for the long-run mean $\mu$, but also enables us to evaluate the finite-sample performance of the existing method— the commonly used plug-in MLE, in which the sample mean is substituted into the likelihood. Our simulation results show that the plug-in MLE performs nearly as well as the exact MLE, alleviating concerns that inefficient estimation of $\mu$ would compromise the efficiency of the remaining parameter estimates.

Introduction

Gaussian processes have been widely applied across a broad range of scientific and applied disciplines, including economics, finance, physics, hydrology, and telecommunications. One of their most extensively studied features is the long-memory property, which captures long-range dependence. The autoregressive fractionally integrated moving average (ARFIMA) model was independently introduced by Granger-1980 and Hosking-1981 to model this feature. In economics and finance, long memory has been examined in a wide array of time series, including real output Diebold-Rudebusch-1989, income Diebold-Rudebusch-1991, stock returns Lo-1991, Liu-Jing-2018, interest rates Shea-1991, exchange rates Diebold-Husted-Rush-1991, Cheung-1993, volatility Ding-Granger-Engle-1993, Andersen-Bollerslev-1997-b, Andersen-Bollerslev-Diebold-Labys-2003 and bubble detection Lui-Phillips-Yu-2024. Various mechanisms have been proposed to explain the emergence of long memory, including cross-sectional aggregation Granger-1980, Abadir-2002, regime switching Potter-1976, Diebold-Inoue-2001, marginalization Chevillon-2018, and network effects Schennach-2018.

More recently, a rapidly growing strand of literature has focused on continuous-time Gaussian processes, which can characterize the local behavior and reproduce the rough sample paths observed in volatility and trading volume Gatheral-Jaisson-Rosenbaum-2018, Fukasawa-Takabatake-Westphal-2022, Wang-Xiao-Yu-2023-JoE, Bolko-Christensen-Pakkanen-Veliyev-2023, Shi-Yu-Zhang-2024-b, Chong-Todorov-2025. Two prominent models in this class are fractional Brownian motion (fBm)Mandelbrot-1965, mandelbrot1968, Gatheral-Jaisson-Rosenbaum-2018 and the fractional Ornstein–Uhlenbeck (fOU) process Cheridito-Kawaguchi-Maejima-2003, Wang-Xiao-Yu-2023-JoE. When applied to data, fractional Gaussian noise (fGn, the first-order difference of fBm) and fOU—exhibit anti-persistence Gatheral-Jaisson-Rosenbaum-2018, Fukasawa-Takabatake-2019, Shi-Yu-Zhang-2024 and short memory Shi-Yu-Zhang-2024-b, Wang-Xiao-Yu-Zhang-2023. Several studies have begun to investigate the micro-level origins of this roughness. For example, eleuch2018 show that in highly endogenous markets, rough volatility may arise from a large number of split orders, while Jusselin-Rosenbaum-2020 demonstrate that rough volatility emerges naturally under the no-arbitrage condition.

The ML estimation method of non-centered stationary discrete- and continuous-time Gaussian models with long memory, short memory, or anti-persistence, referred to as general Gaussian processes, is the focus of this paper. Since these memory properties are relevant across various applications, our goal is to develop an estimation method that does not impose prior restrictions on the memory type of the process. The choice of ML estimation is motivated by the practical need for accurate estimation of all model parameters, particularly when computing impulse response functions or performing forecasts. In such cases, suboptimal estimators---such as the semi-parametric methods of geweke1983, Robinson-1995-LWE, phillips2004, shimotsu2010, the method of moments in Wang-Xiao-Yu-2023-JoE, and the composite likelihood approach in Bennedsen-2024---are not recommended. For example, corsi2009 criticized semi-parametric methods for producing significantly biased and inefficient estimates in forecasting applications with ARFIMA models. Moreover, although the ML and Whittle ML estimators are asymptotically equivalent, the ML estimation method generally demonstrates superior finite-sample performance Cheung-Diebold-1994.\footnote{Rao-Yang-2021 propose new frequency-domain quasi-likelihoods that improve the finite-sample behavior of the Whittle ML estimator for short-memory Gaussian processes.}

Considerable progress has been made in developing ML estimation methods, extending from specific parametric models to general Gaussian processes based on discrete-time observations.\footnote{A parallel literature studies parameter estimation for continuous-time fractional models under continuous-record observations; see Kleptsyna2002.} For example, Yajima-1985 established the consistency and asymptotic normality of the MLE for the ARFIMA$(0,d,0)$ model with $d \in (0, 0.5)$, representing the long-memory case. These results were subsequently extended to stationary Gaussian processes with long memory by Dahlhaus-1989, Dahlhaus-2006, and further generalized to general Gaussian processes by Lieberman-2012. However, these existing methods adopt a two-stage procedure in which $\mu$ is first estimated by the sample mean and then substituted into the likelihood, yielding a so-called plug-in MLE for the remaining parameters. Some implementations apply the MLE to demeaned data Tsai-Chan-2005, shi2022volatility, which is essentially equivalent to the plug-in MLE procedure. Regarding the optimality of this procedure, the sample mean is clearly not the efficient estimator for $\mu$. While Dahlhaus-1989, Dahlhaus-2006 argued for the efficiency of the plug-in MLE by showing that its asymptotic covariance matrix equals the inverse of the Fisher information matrix, they did not explicitly establish the existence of a Crame\'r--Rao lower bound. Cohen-Gamboa-Lacaux-Loubes-2013 derived the LAN property for centered stationary Gaussian processes with long memory, short memory, or anti-persistence, providing a minimax lower bound for general estimators; however, their results do not imply the asymptotic efficiency of plug-in MLE. Another concern is that the inefficiency in estimating $\mu$ may impair the finite-sample performance of the plug-in MLE. Cheung-Diebold-1994 showed that when the $\mu$ is unknown, the finite-sample performance of the MLE for the other parameters deteriorates, even though their asymptotic variances remain the same as in the known-mean case. To date, the problem of obtaining theoretically optimal estimators for all parameters for general Gaussian processes within the ML framework remains unresolved. Only one exception is Wang-Xiao-Yu-Zhang-2023, who established the consistency and asymptotic normality of the MLE for all parameters—including $\mu$—in the fOU process. Nonetheless, their framework is model-specific and not readily applicable to other fractional models. Moreover, the efficiency of the MLE was not addressed.

We introduce a novel exact ML method, a term we adopt to distinguish it from the plug-in ML method, for general Gaussian processes, where all the parameters are estimated jointly. We establish the consistency and asymptotic normality of the exact MLE. These results extend those in Wang-Xiao-Yu-Zhang-2023 from fOU to general Gaussian processes. Moreover, we establish the LAN property of the sequence of statistical experiments in the Le Cam sense. This result extends that in Cohen-Gamboa-Lacaux-Loubes-2013 from centered stationary Gaussian processes to “non-centered” ones within the long span asymptotics, which is different from several extensions under the high-frequency asymptotics recently made by Brouste-Fukasawa-2018,Fukasawa-Takabatake-2019,Szymanski-2024,Szymanski-Takabatake-2023-Additive,Chong-Mies-2025. The LAN property directly yields the efficiency of the exact MLE. The results rely solely on the asymptotic behavior of the spectral density near zero for a discrete record of observations, which allows for broad applicability. \footnote{To ensure our theoretical results are broadly applicable and consistent with discrete observations, we work with the spectral density corresponding to discrete-time data, regardless of whether the underlying process is continuous- or discrete-time. In TYZ (working paper), we demonstrate how to verify the necessary conditions for the discrete spectral density using the continuous-time spectral density for a wide range of continuous-time processes.}

To demonstrate the practical applicability of the proposed method, we conduct three Monte Carlo simulation studies in which our exact ML estimator is applied to three widely studied non-centered processes: the ARFIMA$(0,d,0)$ model, fractional Gaussian noise (fGn), and the fractional Ornstein–Uhlenbeck (fOU) process. Overall, our exact estimator for $\mu$ improves upon the performance of the sample mean. Regarding the performance of plug-in MLE, our simulation results show that the plug-in MLE performs nearly as well as the exact MLE, alleviating concerns that inefficient estimation of $\mu$ would compromise the efficiency of the remaining parameter estimates. In particular, for the ARFIMA$(0,d,0)$ model, the gain in efficiency for estimating $\mu$ aligns with the theoretical result of Adenstedt-1974. We also conduct a forecasting horse race for realized volatility using the fOU process with three alternative estimators: the exact MLE, the plug-in MLE, and the change-of-frequency (CoF) estimator by Wang-Xiao-Yu-2023-JoE. As expected, the exact MLE delivers the best forecasting performance, followed by the plug-in MLE, which performs satisfactorily, and then the CoF estimator.

To sum up, we contribute to the literature in the following aspects. First, we propose a novel exact ML estimation method for all parameters in a general stationary Gaussian process, establishing its consistency and asymptotic normality. Second, we prove the LAN property of the sequence of statistical experiments for general stationary Gaussian processes in the Le Cam sense, providing a theoretical foundation of optimal estimation. The LAN property we have established is also essential for building asymptotic optimality of statistical tests and selecting the order of models based on the likelihood function, see Remark (ref) for further references. Third, our method serves as a benchmark for evaluating the finite-sample performance of the existing plug-in MLE. Although the performance gap between the plug-in MLE and the MLE with known \( \mu \) can be substantial in finite samples Cheung-Diebold-1994, our analysis shows that this difference is not driven by inefficiencies in estimating \( \mu \).

The remainder of this paper is organized as follows. Section (ref) presents the exact ML method and develops asymptotic Properties. In Section (ref), we provide several examples to which our results can be applied. Section (ref) presents a Monte Carlo study to assess the performance of the estimation method. Section (ref) concludes the paper.

Exact MLE and Asymptotic Properties

Notation and Exact MLE

Let $\Theta_{\xi}$ be a convex domain of $\mathbb{R}^{p-1}$ with compact closure and set $\Theta:=\Theta_{\xi}\times(0,\infty)$. Write $\theta=(\xi,\sigma)^{\top}\in\Theta$ and $\vartheta=(\theta,\mu)^{\top}=(\xi,\sigma,\mu)^{\top}\in\Theta\times\mathbb{R}$. Denote by $\partial_{z}=\partial/\partial{z}$, $\partial_{\omega}=\partial/\partial{\omega}$ and $\partial_{j} := \partial/\partial{\theta_{j}}$ for $j\in\{1,\cdots,p+1\}$. For notational simplicity, $\partial_{0}$, $\partial_{z}^{0}$ and $\partial_{\omega}^{0}$ denote the identify operator. The derivative operators $\partial_{j_{1},\cdots,j_{k}}^{k}$ are recursively defined by $\partial_{j_{1},\cdots,j_{k}}^{k} := \partial_{j_{1}} \circ \partial_{j_{2},\cdots,j_{k}}^{k-1}$ for $j_{1},\cdots,j_{k}\in\{0,1,\cdots,p+1\}$ and $k\in\mathbb{N}$. Moreover, $\mathbf{1}_{n}$ denotes a $n$-dimensional vector whose all elements are equal to $1$ and, for an integrable function $f$ on $[-\pi,\pi]$, $\Sigma_{n}(f)$ denotes the symmetric Toeplitz matrix whose $(i,j)$th elements are equal to the $(i-j)$th Fourier coefficients of $f$.\\

Let us consider a stationary Gaussian time series $\{X^{\vartheta}_{j}\}_{j\in\mathbb{Z}}$ with mean $\mu$ and spectral density function $s^{X}(\omega,\theta)$. We may write $s_{\theta}^{X}(\omega):=s^{X}(\omega,\theta)$. Let us denote by $\vartheta_{0}=(\xi_{0},\sigma_{0},\mu_{0})^{\top}$ an interior point of $\Theta\times\mathbb{R}$, which may call a true value of the parameter $\vartheta$, and we assume that we observe a realization of $X_{1}^{\vartheta_{0}},\cdots,X_{n}^{\vartheta_{0}}$. Let $s_{\xi}^{X}(\omega)= s_{\theta}^{X}(\omega)/\sigma^2$. Then, for each $\vartheta=(\theta,\mu)^{\top}\in\Theta\times\mathbb{R}$ and $n\in\mathbb{N}$, we denote by $\mathbb{P}_{\vartheta}^{n}$ a distribution on the Borel space $(\mathbb{R}^{n},\mathcal{B}(\mathbb{R}^{n}))$ under which a random vector $\mathbf{X}_{n} =(X_{1},\cdots,X_{n})^{\top}$ follows a $n$-dimensional Gaussian vector with mean vector $\mu\mathbf{1}_{n}$ and variance-covariance matrix $\Sigma_{n}(s_{\theta}^{X})$. We also denote by $\gamma_{\theta}^{X}(\cdot)$ the autocovariance function of $X^{\vartheta}$. \\

Denote by $\ell_{n}(\vartheta)\equiv\ell_n(\xi,\sigma,\mu)$ the Gaussian log-likelihood function of the observations $\mathbf{X}_{n}$ under the distribution $\mathbb{P}_{\vartheta}^{n}$, which is given by

align[align omitted — 354 chars of source]

and then the maximum likelihood estimator (MLE)\footnote{We have derived alternative expression of the likelihood function using the conditional likelihood based on the Bayes formula, which improve the computational efficiency by avoiding direct computations of the inverse and determinant of large-scale covariance matrix. See (ref) for details.} is defined by

equation*[equation* omitted — 297 chars of source]

Notice that, from the definition of the MLE, the MLE satisfies the estimating equations

align*[align* omitted — 114 chars of source]

which conclude the equations

align*[align* omitted — 382 chars of source]

hold for any $(\xi,\sigma,\mu)^{\top}\in\Theta_{\xi}\times(0,\infty)\times\mathbb{R}$. Then the MLE $\widehat{\xi}_{n}^{\mathrm{MLE}}$ would be a maximizer of the function $\bar{\ell}_{n}(\xi):=\ell_{n}(\xi,\bar{\sigma}_{n}(\xi),\mu_{n}(\xi))$, where $\bar{\sigma}^{2}_{n}(\xi):=\sigma_{n}^{2}(\xi,\mu_{n}(\xi))$ and $\bar{\sigma}_{n}(\xi):=\sqrt{\bar{\sigma}^{2}_{n}(\xi)}$, over the parameter $\xi\in\Theta_{\xi}$ and the MLEs $\widehat{\mu}_{n}^{\mathrm{MLE}}$ and $\widehat{\sigma}_{n}^{\mathrm{MLE}}$ satisfy the equations

align*[align* omitted — 295 chars of source]

Therefore, we define our proposed estimator $\widehat{\vartheta}_{n} :=(\widehat{\xi}_{n},\widehat{\sigma}_{n},\widehat{\mu}_{n})^{\top}$ by

align[align omitted — 252 chars of source]

and we call $\widehat{\vartheta}_{n} =(\widehat{\xi}_{n},\widehat{\sigma}_{n},\widehat{\mu}_{n})^{\top}$ the {\it exact MLE} throughout this paper. In subsequent sections, we investigate the asymptotic properties of the exact MLE.

Consistency and Asymptotic Normality of Exact MLE

We firstly introduce several conditions on the spectral density function $s_{\theta}^{X}(\omega)$ summarized in the following assumption that is used to prove asymptotic properties of the exact MLE and the likelihood ratio process; see Sections (ref) and (ref) for details.

assumption\begin{enumerate}[label=$(\arabic*)$] • For each $\theta\in\Theta$, $s_{\theta}^{X}(\omega)$ is a non-negative integrable even function in $\omega$ on $[-\pi,\pi]$ with $2\pi$-periodicity. Moreover, it satisfies \begin{itemize} • for each $\omega\in[-\pi,\pi]\backslash\{0\}$, $s_{\theta}^{X}(\omega)$ is three times continuously differentiable in $\theta$ on the interior of $\Theta$, • for each $\theta\in\Theta$ and $j\in\{1,\cdots,p\}$, $s_{\theta}^{X}(\omega)$ and $\partial_{j}s_{\theta}^{X}(\omega)$ are continuously differentiable in $\omega$ on $[-\pi,\pi]\setminus\{0\}$. \end{itemize} • If $\theta_{1}$ and $\theta_{2}$ are distinct elements of $\Theta$, the set $\{\omega\in [-\pi,\pi]: s_{\theta_{1}}^{X}(\omega)\neq s_{\theta_{2}}^{X}(\omega)\}$ has a positive Lebesgue measure. • There exists a continuous function $\alpha_X:\Theta_\xi\to(-\infty,1)$ such that for any $\iota>0$ and some constants $c_{1,\iota},c_{2,\iota},c_{3,\iota}>0$, which only depends on $\iota$, the following conditions hold for every $(\theta,\omega)\in\Theta\times[-\pi, \pi]\backslash\{0\}$: \begin{enumerate}[label=$(\alph*)$] • $c_{1,\iota}|\omega|^{-\alpha_{X}(\xi)+\iota}\leq s_{\theta}^{X}(\omega)\leq c_{2,\iota}|\omega|^{-\alpha_{X}(\xi)-\iota}$. • For any $j_{1},j_{2},j_{3}\in\{0,1,\cdots,p\}$, \begin{equation*} \left| \partial_{j_{1},j_{2},j_{3}}^{3} s_{\theta}^{X}(\omega)\right|\leq c_{3, \iota}|\omega|^{-\alpha_{X}(\xi)-\iota} \ \ and\ \ \left| \partial_{\omega} \partial_{j_{1}} s_{\theta}^{X}(\omega)\right|\leq c_{3, \iota}|\omega|^{-\alpha_{X}(\xi)-1-\iota}. \end{equation*} \end{enumerate} \end{enumerate}

Assumption (ref) is the usual conditions on the “discrete-time” spectral density function for stationary Gaussian time series with long/short/anti-persistent memory used in the literature; see the assumptions in Fox-Taqqu-1986, Dahlhaus-1989, Dahlhaus-2006, Lieberman-2012, Cohen-Gamboa-Lacaux-Loubes-2013 and Fukasawa-Takabatake-2019 as references. $X_j^{\vartheta}$ is said to have long memory (or long-range dependence) if $0 < \alpha_{X}(\xi) < 1$, short memory if $\alpha_{X}(\xi) = 0$, and anti-persistence if $\alpha_{X}(\xi)< 0$. The range $\alpha_{X}(\xi) \leq -1$ corresponds to noninvertibility, and our results cover this case as well. Since these memory properties are relevant across various applications, we do not impose prior restrictions on the memory type of the process.

Before stating our main results, we introduce additional notation. We write $\widehat{\theta}_{n}:=(\widehat{\xi}_{n},\widehat{\sigma}_{n})^{\top}$ and then we may write $\widehat{\vartheta}_{n}=(\widehat{\theta}_{n},\widehat{\mu}_{n})^{\top}$. Define $p\times p$ dimensional matrix $\mathcal{F}_{p}(\theta)$ by

align[align omitted — 411 chars of source]

where

align*[align* omitted — 369 chars of source]

\color{black}

Then we assume the following condition on $\mathcal{F}_{p}(\xi)$.

assumptionThe matrix $\mathcal{F}_{p}(\xi)$ is invertible for each $\xi\in\Theta_{\xi}$.

Our first main result is a (weak) consistency and an asymptotic normality of the sequence of the MLEs $\{\widehat{\theta}_{n}\}_{n\in\mathbb{N}}$ defined in (ref), that is a generalization of, for example, Theorems 3.1 and 3.2 in Dahlhaus-1989 and Theorem 1 in Lieberman-2012 to the case of general Gaussian processes using the multi-step estimation procedure based on the exact MLE defined in (ref).

theoremUnder Assumptions $\ref{Assump:DSPD1}$ and $\ref{Assump:mat-Fp-invertible}$, the sequence of the exact MLEs $\{\widehat{\theta}_{n}\}_{n\in\mathbb{N}}$ is weakly consistent and asymptotically normal, that is, it holds that \begin{align*} \sqrt{n} (\widehat{\theta}_{n}-\theta) \rightarrow \mathcal{N}_{p}( \mathbf{0}_{p},\mathcal{F}_{p}(\theta)^{-1} ) \ \ as $n\to\infty$ \end{align*} in law under the distribution $\mathbb{P}_{\vartheta}^{n}$, where $\mathcal{F}_{p}(\theta)$ is the non-singular matrix defined in (ref).

See Section (ref) for the proof of consistency and Section (ref) for the proof of asymptotic normality.

To prove the asymptotic normality of the MLEs for the joint estimation $\vartheta=(\theta,\mu)^{\top}$ as well as the MLE for $\mu$, we need to further assume the precise asymptotic behavior of the spectral density function $s_{\theta}^{X}(\omega)$ around the frequency $\omega=0$ given in the following assumption.

assumptionIn addition to Assumption $\ref{Assump:DSPD1}$, we further assume that there exists a continuous function $c_{X}:\Theta_{\xi}\to(0,\infty)$ such that for each $\xi\in\Theta_{\xi}$, \begin{align*} s_{\theta}^{X}(\omega) \sim \sigma^{2}c_{X}(\xi) |\omega|^{-\alpha_{X}(\xi)}\ \ as $|\omega|\to 0$. \end{align*}

Based on $\mathcal{F}_{p}(\xi)$ defined in (ref), we further define

align[align omitted — 479 chars of source]

where $B(\alpha,\beta)$ is the beta-function. Note that, under Assumption (ref), the matrix $\mathcal{I}(\xi)$ is also invertible for each $\xi\in\Theta_{\xi}$. Moreover, we also introduce a normalized score function $\zeta_{n}(\vartheta)$ and an observed Fisher information matrix $\mathcal{I}_{n}(\vartheta)$, respectively defined by

align[align omitted — 268 chars of source]

Now we can state our second main result of the asymptotic normality of the MLE for the joint parameter $\vartheta=(\theta,\mu)^{\top}$, summarized in the following theorem with its proof given in Section (ref).

theoremUnder Assumptions $\ref{Assump:mat-Fp-invertible}$ and $\ref{Assump:DSPD2}$, the sequence of the MLEs $\{\widehat{\vartheta}_{n}\}_{n\in\mathbb{N}}$ satisfies the following asymptotic normality: \begin{align} \Phi_{n}(\vartheta)^{-1} (\widehat{\vartheta}_{n}-\vartheta) = \mathcal{I}_{n}(\theta)^{-1}\rm{1}\rm{l}_{\{\det[\mathcal{I}_{n}(\theta)]>0\}}\zeta_{n}(\vartheta) + o_{\mathbb{P}_{\vartheta}^{n}}(1) \rightarrow \mathcal{N}_{p+1}\left(\mathbf{0}_{p+1},\mathcal{I}(\theta)^{-1}\right) \end{align} in law under the distribution $\mathbb{P}_{\vartheta}^{n}$ as $n\to\infty$, where $\zeta_{n}(\vartheta)$ and $\mathcal{I}_{n}(\theta)$ are the normalized score function and the observed Fisher information matrix defined in (ref).
remarkWang-Xiao-Yu-Zhang-2023 established the consistency and asymptotic normality of the MLE for all parameters in the fOU process. Nonetheless, their proof is model-specific, and their results are encompassed by our more general framework.
remarkThe asymptotic normality properties in Theorems $\ref{Thm:MLE}$ and $\ref{Thm:MLE2}$ show that the sequences of plug-in MLEs for the estimation of $\theta$ with nuisance parameter $\mu$ and the exact MLEs for the estimation of $\vartheta=(\theta,\mu)^\top$ are respectively asymptotically efficient in the Fisher sense, in that their limiting covariance matrices equal the inverse of the Fisher information matrices, given by the limits of the sequences of the matrices $-n^{-1}\partial_{\theta}^{2}\ell_{n}((\theta,\mu_0)^\top)$ and $\mathcal{I}_n(\vartheta)$ defined in (ref). These Fisher efficiencies have also been discussed in Dahlhaus-1989,Lieberman-2012 for the plug-in MLE and in Wang-Xiao-Yu-Zhang-2023 for the exact MLE. However, these studies do not establish the minimax optimality proved later. One might wonder whether this minimax optimality could already be deduced by passing to the limit from the Cram\'er–Rao inequality formulated for possibly biased estimators. However, this is not the case. Such finite-sample inequalities control pointwise variances but do not rule out the existence of super-efficient points. A classical example is Hodges’s estimator in the i.i.d. Gaussian location model, which is $\sqrt{n}$-consistent and has the same asymptotic variance as the MLE except at a single point where it is super-efficient. This example illustrates that Cram\'er–Rao–type arguments alone are insufficient to rule out super-efficient points, indicating that the existence of general lower bounds for estimators cannot be derived solely from such inequalities. LeCam-1953 showed that the set of super-efficient points has Lebesgue measure zero, with subsequent extensions by Bahadur-1964 and Pfanzagl-1970. These results, however, rely on parametric i.i.d. models or other regularity assumptions, and to the best of our knowledge, no extensions of these results to statistical experiments induced by general Gaussian processes have been established. Therefore, one cannot rely on the Fisher efficiencies or Cram\'er–Rao–type inequalities alone to establish the minimax optimality in local neighborhoods of the true parameter. To overcome this limitation, one needs the LAN property together with the H\'ajek–Le Cam’s local asymptotic minimax theorem Hajek-1972,LeCam-1972, which ensures that no estimator can asymptotically achieve a smaller risk than the bound determined by the Fisher information in shrinking neighborhoods of the true parameter. This motivates the next subsection, where we establish the LAN property for statistical experiments induced by general Gaussian processes and then derive the minimax efficiency of the exact MLE as well as the plug-in MLE.

LAN property and Asymptotic Efficiency of the Exact MLE

The LAN property is a fundamental concept in asymptotic statistics, originally introduced by Wald-1943 and further developed by LeCam-1960. It plays a central role not only in establishing the asymptotic optimality of estimators but also in facilitating statistical inference. In this section, we establish the LAN property of the sequence of statistical experiments ${(\mathbb{R}^{n}, \mathcal{B}(\mathbb{R}^{n}), {\mathbb{P}{\vartheta}^{n}}{\vartheta \in \Theta \times \mathbb{R}})}_{n \in \mathbb{N}}$ for general Gaussian processes. We then provide several results concerning the local asymptotic minimax efficiency of the exact MLE. Additional applications of the LAN property are also discussed.

theoremConsider the sequence of rate matrices $\{\Phi_{n}(\vartheta)\}_{n\in\mathbb{N}}$ defined in (ref). Under Assumptions $\ref{Assump:mat-Fp-invertible}$ and $\ref{Assump:DSPD2}$, the family of distributions $\{\mathbb{P}^{n}_{\vartheta}\}_{\vartheta\in\Theta\times\mathbb{R}}$ satisfies the following LAN property at each interior point $\vartheta=(\theta,\mu)^{\top}$ of $\Theta\times\mathbb{R}$: \begin{align*} \left| \log\frac{\mathrm{d}\mathbb{P}^{n}_{\vartheta+\Phi_{n}(\vartheta)u}}{\mathrm{d}\mathbb{P}^{n}_{\vartheta}} -\left( u^{\top} \zeta_{n}(\vartheta)-\frac{1}{2}u^{\top}\mathcal{I}(\theta)u \right)\right|=o_{\mathbb{P}^{n}_{\theta}}(1)\ \ as $n\to\infty$, \end{align*} where the invertible matrix $\mathcal{I}(\theta)$ is defined in $\eqref{Def:mat-I}$ and the normalized score function $\zeta_{n}(\vartheta)=\Phi_{n}(\vartheta)^{\top}\ell_{n}(\vartheta)$ satisfies the convergence \begin{equation*} \zeta_{n}(\vartheta)\rightarrow\mathcal{N}(0,\mathcal{I}(\theta))\ \ as $n\to\infty$ \end{equation*} in law under the distribution $\mathbb{P}_{\vartheta}^{n}$.

The proof of Theorem (ref) is left to Section (ref). In the relevant literature, Theorem 2.4 in Cohen-Gamboa-Lacaux-Loubes-2013 provides the LAN property for centered stationary Gaussian time series with long/short/anti-persistent memory property under Assumptions (ref) and (ref). Theorem (ref) is a generalization of Theorem 2.4 in Cohen-Gamboa-Lacaux-Loubes-2013 to general Gaussian processes within the long span asymptotics, which is different from several extensions under the high-frequency asymptotics recently made by Brouste-Fukasawa-2018,Fukasawa-Takabatake-2019,Szymanski-2024,Szymanski-Takabatake-2023-Additive,Chong-Mies-2025.

The LAN property is applied to derive asymptotically efficient rates and variances for estimating $\vartheta=(\theta,\mu)^{\top}$. These results are derived from lower bounds of estimation provided by H\'ajek's convolution theorem (see Hajek-1972, Ibragimov-Hasminski-1981,LeCam-1972), recalled below.

theorem[Theorem II.12.1 in Ibragimov-Hasminski-1981] Suppose that a family of distributions $\{\mathbb{P}^{n}_{\theta}\}_{\theta\in\Theta}$ satisfies the LAN property at the interior point $\theta$ of $\Theta \subset \mathbb R^{d}$ for a sequence of $d\times d$-rate matrices $\{\Phi_{n}(\vartheta)\}_{n\in\mathbb{N}}$ and a $d\times d$-invertible matrix $\mathcal{I}(\theta)$. Then, for any sequence of estimators $\widehat{\theta}_n$ and any symmetric nonnegative quasi-convex function $L$ on $\mathbb{R}^{d}$ such that $e^{-\varepsilon \|z\|_{\mathbb{R}^{d}}^2} L(z) \to 0$ as $\|z\|_{\mathbb{R}^{d}} \to \infty$ for any $\varepsilon > 0$, we have \begin{align*} \varliminf_{c\to\infty}\varliminf_{n \to \infty} \sup_{\theta^{\prime}\in\Theta: \left\|\Phi_{n}(\vartheta)^{-1}(\theta^{\prime}-\theta)\right\|_{\mathbb{R}^{d}}\leq c} \mathbb{E}_{\theta^{\prime}}^{n}\left[ L\left(\Phi_{n}(\vartheta)^{-1}(\widehat{\theta}_{n}-\theta^{\prime})\right) \right] \geq (2\pi)^{-\frac{d}{2}}\int_{\mathbb{R}^d} L\left(\mathcal{I}(\theta)^{-1/2}z\right) \exp\left(-\frac{|z|^2}{2}\right)\,\mathrm{d}z. \end{align*}

Notice that we have already proved that the sequence of the exact MLEs $\{\widehat{\vartheta}_{n}\}_{n\in\mathbb{N}}$ defined in (ref) satisfies the coupling property (ref) in Theorem $\ref{Thm:MLE2}$ so that, using the result in Section 7.12.(b) of Hopfner-2014-Book in addition to Theorem (ref), we can conclude that the sequence of exact MLEs is asymptotically efficient in the local asymptotic minimax sense as well as in the Fisher sense, and then it attains the local asymptotic minimax bound of estimation given in Theorem (ref). We summarize the aforementioned result in the following corollary.

corollary[Asymptotic Minimax Optimality] Consider the sequence of rate matrices $\{\Phi_{n}(\vartheta)\}_{n\in\mathbb{N}}$ defined in (ref) and the matrix $\mathcal{I}(\theta)$ defined in (ref). Under Assumptions $\ref{Assump:mat-Fp-invertible}$ and $\ref{Assump:DSPD2}$, the sequence of the exact MLEs $\{\widehat{\vartheta}_{n}\}_{n\in\mathbb{N}}$ defined in (ref) attains the local asymptotic minimax bound given in Theorem $\ref{Thm:Hajek-LeCam}$ at each interior point $\vartheta=(\theta,\mu)^{\top}$ of $\Theta\times\mathbb{R}$. Namely, for any symmetric nonnegative quasi-convex function $L$ on $\mathbb{R}^{p+1}$ such that $e^{-\varepsilon \|z\|_{\mathbb{R}^{p+1}}^2} L(z) \to 0$ as $\|z\|_{\mathbb{R}^{p+1}} \to \infty$ for any $\varepsilon > 0$, we obtain \begin{align*} &\varliminf_{c\to\infty}\varliminf_{n \to \infty} \sup_{\vartheta\in\Theta\times\mathbb{R}: \left\|\Phi_{n}(\vartheta_{0})^{-1}(\vartheta-\vartheta_{0})\right\|_{\mathbb{R}^{p+1}}\leq c} \mathbb{E}_{\vartheta}^{n}\left[ L\left(\Phi_{n}(\vartheta_{0})^{-1}(\widehat{\vartheta}_{n}-\vartheta)\right) \right] \\ & = (2\pi)^{-\frac{1}{2}(p+1)} \int_{\mathbb{R}^{p+1}} L\left(\mathcal{I}(\xi_{0})^{-1/2}z\right) \exp\left(-\frac{|z|^2}{2}\right)\,\mathrm{d} z. \end{align*}
remarkRegarding the efficiency of the two-stage procedure, the sample mean is clearly not an efficient estimator for $\mu$. For the remaining parameters, Dahlhaus-1989, Dahlhaus-2006 showed that the asymptotic covariance matrix of the plug-in MLE equals the inverse of the Fisher information matrix; however, they did not explicitly establish the existence of a Cramér--Rao lower bound. Although the LAN property has been established by Cohen-Gamboa-Lacaux-Loubes-2013 for centered stationary Gaussian processes with long memory, short memory, or anti-persistence, their results do not imply the asymptotic efficiency of the plug-in MLE.
remarkUnder additional technical conditions, the LAN property established in this paper can be used to verify, for instance, the asymptotically uniformly most powerful unbiased (AUMPU) property of the likelihood ratio test Choi-Hall-Schick-1996, as well as asymptotic properties of model selection criteria—such as the (weak) consistency of model selection based on the (quasi) Bayesian information criterion (BIC) Eguchi-Masuda-2018.

Comparison between Exact MLE and Plug-in MLE

As an alternative estimator of $\theta$ in the literature, the plug-in MLE (PMLE) is defined by

align[align omitted — 171 chars of source]

using some compact set $\Theta_{\ast}\subset\Theta$ and an estimator $\widetilde{\mu}_{n}$. Under similar assumptions to Assumption (ref), Dahlhaus-1989, Dahlhaus-2006 and Lieberman-2012 show that the plug-in MLE is weakly consistent, asymptotically normal with the asymptotic variance $\mathcal{F}_{p}(\theta)^{-1}$ at the point $\vartheta_{0}$? (should be $\theta_0$) when the plug-in estimator $\widetilde{\mu}_{n}$ satisfies the assumption

align[align omitted — 170 chars of source]

The assumption (ref) corresponds to Assumption 5 of Dahlhaus-1989 for the long memory case $\alpha_{X}(\xi_{0})\in(0,1)$ and Assumption 5 of Lieberman-2012 for the long/short/anti-persistent memory case $\alpha_{X}(\xi_{0})\in(-\infty,1)$. The plug-in MLE shares the same convergence rate and asymptotic variance as our exact MLE, implying that it is also asymptotically efficient.

Notice that Theorem (ref) combining with Theorem (ref) yields the asymptotic minimax lower bound

align[align omitted — 347 chars of source]

for any $r>0$ and any sequences of estimators $\widetilde{\mu}_{n}$, which implies the convergence rate of the estimator satisfying the assumption (ref) is the same as that of the optimal estimators given in (ref), that is, the minimax optimal rate. Although Adenstedt-1974 proved that the best linear unbiased estimator (BLUE) of $\mu$ satisfies the assumption (ref) for all $\alpha_{X}(\xi_{0})\in(-\infty,1)$, the BLUE of $\mu$ is infeasible without knowing the true value of $\xi_{0}$. Although the assumption (ref) is satisfied by the widely used sample mean, its asymptotic variance does not achieve the minimax optimal bound. For example, Adenstedt-1974 quantified the asymptotic inefficiency of the sample mean for the ARFIMA$(0,d,0)$ model by computing the ratio of the asymptotic variance of an efficient estimator for $\mu$ to that of the sample mean:

align[align omitted — 114 chars of source]

This metric is always less than $1$, except when $d = 0$. This plug-in approach is equivalent to the implementation that applies the MLE to demeaned data. Lieberman-2012 proposed an alternative estimator of the form

align[align omitted — 189 chars of source]

where $s_\ast:=s_{\theta_\ast}$ with any $\theta_\ast=(\xi_\ast,\sigma_\ast)^\top\in\Theta$ satisfying $\alpha(\xi_\ast)=\inf_{\xi\in\Theta_\xi}\alpha(\xi)$ (by compactness of $\Theta_\xi$ there exists at least one such value), or even $s_\ast(\omega):=(1-\cos(\omega))^{\frac{\alpha_\ast}{2}}$ with $\alpha_\ast\leq\inf_{\xi\in\Theta_\xi}\alpha(\xi)$, and proved that the estimator in (ref) satisfies the assumption (ref) for all $\alpha_{X}(\xi_{0})\in(-\infty,1)$ using the results in Adenstedt-1974. This estimator also fails to attain the minimax optimal bound, as it inherently relies on a misspecified structure of the autocovariance function embedded in the Toeplitz matrix $\Sigma_n(s_\ast)$.

Examples

Our assumptions on the spectral density are very general and apply to many well-known processes, including but not limited to the ARFIMA$(p,d,q)$ process with $|d|<1/2$, the fractional Gaussian noise with an unknown mean and the fOU process.\footnote{For other continuous-time processes—such as the continuous-time autoregressive fractionally integrated moving average (CARFIMA) model—we refer the reader to Tetsuya et al. (WP) for a detailed analysis of their spectral densities.}

Non-Centered Gaussian ARFIMA$(p,d,q)$ Process

The {\it non-centered} ARFIMA$(p,d,q)$ process was introduce by Granger-1980 and Hosking-1981 independently. For notation simplicity, we start with the ARFIMA$(0,d,0)$ model. The non-centered ARFIMA$(0,d,0)$ model is specified as

equation[equation omitted — 117 chars of source]

where $L$ is the lag operator, $(1-L)^{-d}$ is the fractional difference operator with the memory parameter $d$ and $\epsilon_{t}\overset{\mathrm{iid}}{\sim} N(0,1)$. It reduces to a Gaussian white noise when $d=0$. When $d\in (-1/2,1/2)$, the ARFIMA$(0,d,0)$ process is stationary and invertible bloomfield1985. Let $u_{t}:=(1-L)^{-d}\epsilon _{t}$ be the fractionally integrated process and $\gamma _{u}(j):=\mathrm{Cov}[u_{t},u_{t-j}]$ be its $j$th order auto-covariance. According to Hosking-1981, the auto-covariance function of $ u=\{u_{t}\}_{t\in\mathbb{Z}}$ is expressed by

equation[equation omitted — 108 chars of source]

The long-run variance covariance $\sum_{j=-\infty }^{\infty }\gamma _{u}(j)$ $= \infty $ when $d\in \left( 0,1/2\right) $ and $\sum_{j=-\infty }^{\infty }\gamma_{u}(j)=0$ when $d\in \left( -1/2,0\right) $. Therefore, $ u_{t}$ has a long memory if $d\in \left( 0,1/2\right) $ and is anti-persistent if $d\in \left( -1/2,0\right) $. The spectral density of the model is given by \[s_{\theta}^X(\omega)=\frac{\sigma^2}{2\pi} \sim \frac{\sigma^2}{2\pi}|\omega|^{-2d}\ \ \text{ as }|\omega|\rightarrow 0.\] In this case, $\alpha_X(\xi)=2d, c_X(\xi)=(2\pi)^{-1}$.

Let $p,q\in\mathbb{N}\cup\{0\}$ and $\xi:=(\phi_1,\ldots,\phi_p,\psi_1,\ldots,\psi_q)\in\mathbb{R}^{p+q}$. The {\it non-centered} ARFIMA$(p,d,q)$ process is defined by

align[align omitted — 101 chars of source]

where $\phi_\xi(z):=1-\phi_1z-\cdots-\phi_pz^p$ and $\psi_\xi(z):=1+\psi_1z+\cdots+\psi_qz^q$. Assume that for each \( \xi \), the functions \( \phi_\xi(z) \) and \( \psi_\xi(z) \) have no common roots in \( \mathbb{C} \), and that all their roots lie outside the unit circle. This implies that \( \phi_\xi(z) \neq 0 \) and \( \psi_\xi(z) \neq 0 \) for \( |z| \leq 1 \). Then, for \( |d| < 1/2 \), the difference equation in (ref) admits a unique stationary solution \( X = \{X_t\}_{t\in\mathbb{Z}} \) of the form \[ X_t = \mu + \sigma \phi_\xi(L)^{-1} \psi_\xi(L) (1-L)^{-d} \epsilon_t, \quad t \in \mathbb{Z}. \]

The spectral density function of ARFIMA$(p,d,q)$ is given by \[ s_{\theta}^X(\omega) = \frac{\sigma^2}{2\pi} |1 - e^{-i\omega}|^{-2d} \frac{|\psi_\xi(e^{-i\omega})|^2}{|\phi_\xi(e^{-i\omega})|^2}. \] for $\omega \in (-\pi, \pi]$. It can be shown that \[ s_{\theta}^X(\omega) =\frac{\sigma^{2}}{2\pi}\left( 2-2\cos\left( \lambda\right) \right) ^{-d}\frac{\left( 1+\sum_{j=1} ^{q}\psi_{j}\cos\left( j\omega\right) \right) ^{2}+\left( \sum_{j=1} ^{q}\psi_j\sin\left( j\omega\right) \right) ^{2}}{\left( 1-\sum_{j=1}^{p} \phi_{j}\cos\left( j\omega\right) \right) ^{2}+\left( \sum_{j=1} ^{p}\phi_{j}\sin\left( j\omega\right) \right) ^{2}}. \] Since the assumption \( \phi_\xi(z) \neq 0 \) for \( |z| \leq 1 \) ensures that the spectral density function of the ARMA\( (p,q) \) process, \[ f_{\mathrm{ARMA}}(\omega) := \frac{\sigma^2}{2\pi} \frac{|\psi_\xi(e^{-i\omega})|^2}{|\phi_\xi(e^{-i\omega})|^2}, \] is bounded away from zero on \( [-\pi, \pi] \), the singularity of the spectral density of \( X \) in the vicinity of zero frequency is governed by the ARFIMA\( (0,d,0) \) factor \( |1 - e^{-i\omega}|^{-2d} \). As \( \omega \to 0 \), it exhibits the asymptotic behavior \[ s_{\theta}^X(\omega) \sim \sigma^2c_X(\theta)|\omega|^{-\alpha_X(\xi)}. \]

In this case, $\alpha_X(\xi) = 2d, c_X(\xi) = (2\pi)^{-1} \left| \frac{\psi_\xi(1)}{\phi_\xi(1)} \right|^2$ in Assumptions (ref) and (ref). Hence, our results are applicable to the non-centered Gaussian ARFIMA$(p, d, q)$ process. According to Theorem (ref), when $d < 0$, the convergence rate for the exact MLE of $\mu$ is $\frac{1}{2}(1 - \alpha_X(\xi)) > 1/2$, indicating super-consistency.

Non-Centered Fractional Gaussian Noise

The fractional Brownian motion (fBm) with Hurst index $H\in(0,1)$, denoted by $B^H = \{B_t^H\}_{t \in \mathbb{R}}$, is a unique centered Gaussian process that is almost surely equal to zero at $t = 0$ and possesses both stationary increments and $H$-self-similarity properties. Specifically, these properties are expressed as

align*[align* omitted — 115 chars of source]

for any $s, t \in \mathbb{R}$ and $c > 0$, where $\overset{d}{=}$ denotes equality in distribution. mandelbrot1968 demonstrated that the fBm can be represented as a causal moving average process involving the past differential increments of a (two-sided) standard Brownian motion $B = \{B_t\}_{t \in \mathbb{R}}$. This representation is given by

equation*[equation* omitted — 219 chars of source]

where $\Gamma(x)$ denotes the gamma function, \footnote{ This is also referred to as the Type I fBm.} which implies that the fBm reduces to the standard Brownian motion when $H=0.5$.

The fBm is a Gaussian process with mean zero and covariance

equation[equation omitted — 207 chars of source]

The increment of fBm is fGn and denoted by $y_{t}$. Using discrete time notations, we have

equation*[equation* omitted — 63 chars of source]

We extend the fGn to the non-centered case by defining its expectation as $\mu$. The process is given by:

equation*[equation* omitted — 79 chars of source]

The auto-covariance function of $X$ is, $\forall k\geq 0$,

equation[equation omitted — 195 chars of source]

where the approximation is based on the Taylor expansion. When $H\in (0.5,1)$, the asymptotic behavior in (ref) implies that the sequence of the auto-covariances of fGn is not absolutely summable so that the fGn has the long memory property. When $H\in (0,0.5)$, we can also verify that $\forall k\neq 0$, $ \mathrm{Cov}[X_{t},X_{t+k}] <0$ and $\sum_{k=-\infty }^{\infty}\mathrm{Cov}[X_{t},X_{t+k}] =0$ so that the fGn has the anti-persistent memory property.

The spectral density of (non-centered) fGn is given by sinai1976:

equation[equation omitted — 201 chars of source]

where $C_{H}:=\left( 2\pi \right) ^{-1}\Gamma (2H+1)\sin (\pi H)$. It can be shown that \[s_{\theta}^X(\omega ) \sim \sigma^2C_H|\omega|^{1-2H},\,\, \text{when}\,\, \omega \rightarrow 0.\] In this case, the functions $\alpha_X(\xi)$ and $c_X(\xi)$ in Assumptions (ref) and (ref) are $\alpha_X(\xi)=2H-1$ and $c_X(\xi)=C_H$. Hence, our results are applicable to the non-centered fractional Gaussian noise. According to Theorem (ref), when $H<0.5$, the convergence rate for the exact MLE of $\mu$ is $\frac{1}{2}(1-\alpha_X(\xi))>1/2$, implying super-consistency. This result echoes that for ARFIMA.

Fractional Ornstein-Uhlenbeck Process

The fractional Ornstein-Uhlenbeck (fOU) process is an extension of the classical Ornstein-Uhlenbeck (OU) process, where the driving noise is replaced by a fBm with Hurst index $H \in (0,1)$. This process is particularly useful for modeling systems that exhibit long-range dependence and locally self-similarity, which cannot be captured by the classical OU process. The stationary fOU process with a long-run mean $\mu$ has applications in various fields, including mathematical finance, physics, and time series modeling.

The fOU process $Y=\{Y_t\}_{t\in\mathbb{R}}$ is defined by a unique solution of the following linear SDE (Stochastic Differential Equation):

equation[equation omitted — 132 chars of source]

with initial condition $Y_0$, where $B^H = \{B_t^H\}_{t\in\mathbb{R}}$ is a fBm with Hurst index $H$. The explicit solution of this SDE is given by

equation[equation omitted — 159 chars of source]

where the above stochastic integral can be interpreted as the pathwise Riemann-Stieltjes integral or the Wiener integral associated with the fBm for any $H\in(0,1)$. The fOU process reduces to the classical OU process when $H=0.5$ and to the fBm when $\kappa =0$. When $\kappa >0$, the stationary solution of the SDE (ref), denoted by $\bar{Y}=\{\bar{Y}_t\}_{t\in\mathbb{R}}$, is given by

align[align omitted — 147 chars of source]

For $t\geq 0$, the unique solution of the SDE (ref) with the initial condition

align*[align* omitted — 79 chars of source]

is exactly equal to the stationary solution $\bar{Y}$ in (ref), and then the error between $\bar{Y}_t$ and $Y_t$ with arbitrary initial condition $Y_0$ is expressed by

align*[align* omitted — 75 chars of source]

which implies that the error between the solutions (ref) and (ref) converges to zero exponentially as $t\to\infty$ for arbitrary initial condition $Y_0$. In the rest of this section, we consider the case where a data generating process is the discretely and equidistantly observed time series from the stationary solution given in (ref).

Consider a stationary time series $X=\{X_j\}_{j\in\mathbb{Z}}$ of the form $X_j:=Y_{j\Delta}$ for $j\in\mathbb{Z}$ with the sampling frequency $\Delta$. Notice that the time series $X$ is stationary and its auto-covariance is available from garnier2018 when $\kappa >0$:

align[align omitted — 246 chars of source]

From Cheridito-Kawaguchi-Maejima-2003, the autocovariance function of the fOU process exhibits the same order of decay as fGn, decaying hyperbolically for $H \neq 1/2$. Cheridito-Kawaguchi-Maejima-2003 and Hult-2003-PhDThesis provide the spectral density function of the stationary solution $\bar{Y}$ given by

equation[equation omitted — 165 chars of source]

For discrete observations $X$, the spectral density is given by Hult-2003-PhDThesis

equation[equation omitted — 273 chars of source]

It can be shown that \[ s_{\theta}^X(\omega) \sim

cases\sigma^2C_{H} \Delta ^{2H} \sum_{k=-\infty }^{\infty }\frac{ \left\vert 2\pi k\right\vert ^{1-2H}}{(\kappa \Delta )^{2}+\left(2\pi k\right) ^{2}}, & when \omega \rightarrow 0 and 0 < H \leq \frac{1}{2}, \\ \sigma^2C_{H} \Delta ^{2H-2}\kappa^{-2} |\omega|^{1 - 2H}, & when \omega \rightarrow 0 and \frac{1}{2} < H < 1,

\] Shi-Yu-Zhang-2024-b. Hence, for fOU, we have \[ \alpha_X(\xi) =

cases0, & when H \leq \frac{1}{2},\\ 2H - 1, & when H > \frac{1}{2},

\] and \[ c_X(\xi) =

casesC_{H} \Delta ^{2H} \sum_{k=-\infty }^{\infty }\frac{ \left\vert 2\pi k\right\vert ^{1-2H}}{(\kappa \Delta )^{2}+\left(2\pi k\right) ^{2}}, & when H \leq \frac{1}{2},\\ C_{H} \Delta ^{2H-2}\kappa^{-2} & when H > \frac{1}{2},

\] in Assumptions (ref) and (ref). Hence, our results are applicable to the fOU process. However, the function $\alpha_X(\xi)$ exhibits a sharp contrast compared to that of ARFIMA and fGn. According to Theorem (ref), the convergence rate for the exact MLE of $\mu$ in fOU is $\sqrt{n}$ when $H\leq 1/2$, as recently reported in Wang-Xiao-Yu-Zhang-2023.

Mont Carlo Study

We consider three data-generating processes (DGPs): ARFIMA$(0,d,0)$, fGn and fOU. For simplicity, the long-run mean $\mu$ is set to 0. The results for $\mu \pm 1$ are provided in Appendix (ref). The parameter $d$ in the ARFIMA$(0,d,0)$ model takes 9 values: $\{-0.4, -0.3, \dots, 0, \dots, 0.4\}$, while the parameter $H$ in the fGn and fOU takes 9 values: $H = 0.1, 0.2, \dots, 0.9$. The fOU has an additional parameter $\kappa=10$. The scale parameter $\sigma$ is set to $1$ in all cases. For each DGP, we compare our exact MLE with the two MLEs considered in Cheung-Diebold-1994. MLE1 refers to the case where $\mu$ is known, MLE2 refers to our exact MLE, and MLE3 refers to the plug-in MLE, where the sample mean is used as an estimator for $\mu$. The number of replications is set to 1000. The sample size is set to $250$ or $1000$. Reported are the bias, standard error (Std), and root mean squared error (RMSE) across all replications for each method.

table[table omitted — 5,547 chars of source]
table[table omitted — 5,523 chars of source]

Table (ref) reports the results for ARFIMA and Table (ref) reports the results for fGn. From Tables (ref)-(ref), we observe the following findings. First, in terms of convergence rates, the performance of MLE2 aligns well with our asymptotic theory. For $\mu$, the convergence rate is $n^{-(1 - \alpha_X(\xi))/2}$, which becomes slower as $d$ increases toward $1/2$ from $-1/2$ (or as $H$ increases toward $1$ from $0$). A similar pattern can be observed for the sample mean, as it shares the same convergence rate as given in ((ref)) and ((ref)). For the remaining parameters, the convergence rate remains at the root-$n$. Second, MLE2 always performs better than MLE3, except when $d = 0$ in ARFIMA or $H = 1/2$ in fGn, where the two methods perform nearly identically. Third, interestingly, this superior performance in estimating $\mu$ by MLE2 does not translate into better performance in estimating other parameters. Using the true value of $\mu$, MLE1 does not lead to a better performance in estimating other parameters. The three ML methods lead to a similar finite sample performance for parameters other than $\mu$.

To see how the relative inefficiency of the sample mean over MLE2 of $\mu$, the two dashed lines in Figure (ref) plot the ratio of the sample variance of MLE2 for $\mu$ to that of MLE3 as a function of $d$ for the two models when $n = 1000$. Clearly, the relative inefficiency goes up rapidly as $d$ nears $-0.5$ in ARFIMA and fGn. For comparison, also plotted by the solid line is the theoretical asymptotic inefficiency given in ((ref)). The red dashed line is closely aligned with the theory.

figure[figure omitted — 266 chars of source]
table[table omitted — 4,736 chars of source]
table[table omitted — 4,714 chars of source]

Tables (ref)-(ref) report the results for fOU, from which we observe the following findings. First, in terms of convergence rates, the performance of MLE2 aligns well with our asymptotic theory. For $\mu$, the convergence rate is root-$n$ when $H \leq 1/2$, and transitions to $n^{1 - H}$ when $H > 1/2$. The standard deviation of the estimator for $\mu$ decreases substantially as the sample size increases from $250$ to $1000$ when $H \leq 1/2$; however, as $H$ approaches $1$, the percentage reduction becomes markedly smaller. A similar pattern can be observed for the sample mean, as it shares the same convergence rate as given in ((ref)) and ((ref)). For the remaining parameters, the convergence rate remains at the root-$n$. Second, we see a clear dominance of MLE2 over MLE3 in terms of finite sample performance of estimates of $\mu$, when both $n$ is large and $H$ is near either zero or one. Third, this superior performance in estimating $\mu$ by MLE2 does not translate into a better performance in estimating other parameters. Using the true value of $\mu$, MLE1 does not lead to better performance in estimating the other 3 parameters. The three ML methods lead to a similar finite sample performance for parameters other than $\mu$.\footnote{In the online supplement (Section (ref)), we conduct a forecasting horse race for realized volatility using the fOU process with three alternative estimators: MLE2, MLE3, and the CoF estimator. As expected, MLE2 delivers the best forecasting performance, followed by MLE3 and then the CoF estimator.}

Conclusion

Gaussian processes have gained significant attention due to their broad applicability across various scientific and applied disciplines. To obtain the MLE, two common approaches are typically employed. The first approach maximizes the likelihood assuming \( \mu \) is known and set to $0$, which results in an unrealistic MLE. The second approach uses the sample mean as an estimator for \( \mu \), leading to a plug-in MLE. However, both methods fail to address the inefficiency of the estimator for \( \mu \), and concerns have been raised about the finite sample performance of the plug-in MLE. Adenstedt-1974 proposed an efficient but infeasible estimator for \( \mu \).

In this paper, we introduce a novel exact ML method for all the parameters for general Gaussian processes with long-memory, short-memory, or anti-persistence properties. We prove that the exact MLE exhibits the properties of consistency and asymptotic normality. We also establish the LAN property of the sequence of statistical experiments for general Gaussian processes in the sense of Le Cam, which directly yields efficiency. Our method offers a comprehensive understanding of MLE for fractional Gaussian models: first, we show that the estimators for all parameters are optimal, effectively complementing the infeasible estimator for $\mu$ proposed by Adenstedt-1974; second, we evaluate the difference between the plug-in MLE, exact MLE, and the MLE with known $\mu$. The plug-in MLE performs as good as the exact MLE for all parameters except for \( \mu \). The discrepancy between plug-in MLE and the MLE with known $\mu$ is not due to an inefficient estimator for \( \mu \).

The Whittle MLE is asymptotically equivalent to the exact MLE under certain regularity conditions. Although its finite-sample performance is generally inferior to that of the exact MLE, the performance gap narrows as the sample size increases. At the same time, the computational burden of the exact MLE increases substantially due to the need to invert the covariance matrix at each evaluation of the likelihood function—a step that the Whittle method avoids. The Whittle ML method remains an attractive alternative. However, existing theoretical results for the Whittle MLE primarily pertain to ARFIMA models and do not extend to continuous-time models. In future work, we aim to investigate the optimality of the Whittle MLE for general Gaussian processes.