Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
22,792 characters · 4 sections · 8 citation commands
Estimating sample paths of Gauss-Markovprocesses from noisy data
Suppose we observe data $\mathcal{D}=\{(x_i,y_i)\}_{i=1}^n$ generated by the process
where $x_i\ge0$ is non-decreasing in $i$, where $f:[0,\infty)\to\mathbb{R}$ is unknown, and where the errors $\varepsilon_i$ are jointly normally distributed (hereafter “Gaussian”) with $\operatorname{E}[\varepsilon_i\mid x_1,x_2\ldots,x_n]=0$ independently of $\{f(x)\}_{x\ge0}$. We use $\mathcal{D}$ to construct pointwise estimates \[ \hat{f}_\mathcal{D}(x)\equiv\operatorname{E}[f(x)\mid\mathcal{D}] \] with mean squared error (MSE)
In this note, I derive expressions for $\operatorname{E}[f(x)\mid\mathcal{D}]$ and $\operatorname{Var}(f(x)\mid\mathcal{D})$ when $\{f(x)\}_{x\ge0}$ is a sample path of a Gauss-Markov process. Such processes have two defining properties:
Property (G) implies that $f(x)\mid\mathcal{D}$ is Gaussian, and so its distribution is fully determined by its mean $\operatorname{E}[f(x)\mid\mathcal{D}]$ and variance $\operatorname{Var}(f(x)\mid\mathcal{D})$. Theorem (ref) expresses these moments in terms of the means and (co)variances of the $f(x_i)\mid\mathcal{D}$. This allows me to construct the estimate $\hat{f}_\mathcal{D}(x)$ and its MSE at all points $x\ge0$. This estimate optimally extrapolates from, or interpolates between, the observations $(x_i,y_i)$ in $\mathcal{D}$. \footnote{ The estimate is “optimal” in that it minimizes the MSE (ref) for all $x\ge0$. }
I let these observations be noisy, with Gaussian errors $\varepsilon_i=y_i-f(x_i)$. This allows me to extend analyses that assume observations have no noise Bardhi-2024-ECTA,Callander-2011-AER,Carnehl-Schneider-2023-. \footnote{ Rasmussen-Williams-2006- derive expressions for $\operatorname{E}[f(x)\mid\mathcal{D}]$ and $\operatorname{Var}(f(x)\mid\mathcal{D})$ when $\{f(x)\}_{x\ge0}$ follows a Gaussian (but not necessarily Markov) process and the errors $\varepsilon_i$ are iid. I impose the Markov property (M) to obtain (relatively) closed-form expressions for $\operatorname{E}[f(x)\mid\mathcal{D}]$ and $\operatorname{Var}(f(x)\mid\mathcal{D})$. I also allow for arbitrary (co)variances in the $\varepsilon_i$. } I also let the sampled points $x_i$ be less or greater than the target point $x$. This contrasts with Davies-2024-, who studies sequential learning from noisy observations of a sample path. Such learning always involves extrapolation, whereas I allow for interpolation.
The data $\mathcal{D}=\{(x_i,y_i)\}_{i=1}^n$ contain noisy observations $y_i=f(x_i)+\varepsilon_i$ of the values $f(x_i)$. These observations equal the sum of two Gaussian random variables and so are Gaussian too. Moreover, by property (G), the vector $(f(x_1),f(x_2),\ldots,f(x_n),f(x))$ is multivariate Gaussian for all $x\ge0$. It follows that \[ (y_1,y_2,\ldots,y_n,f(x))=(f(x_1),f(x_2),\ldots,f(x_n),f(x))+(\varepsilon_1,\varepsilon_2,\ldots,\varepsilon_n,f(x)) \] is also multivariate Gaussian. Consequently, we can construct the conditional distribution of $f(x)$ given $y_1,y_2,\ldots,y_n$ using a well-known result about multivariate Gaussian variables:
See Bishop-2006- or DeGroot-2004- for proofs of this lemma, and Appendix (ref) for proofs of my other results.
Substituting $z_1=f(x)$ and $z_2=(y_1,y_2,\ldots,y_n)$ into (ref) provides expressions for the moments of $f(x)\mid\mathcal{D}=z_1\mid z_2$. I refine these expressions by imposing properties (G) and (M).
Property (G) comes from $\{f(x)\}_{x\ge0}$ being a sample path of a Gaussian process. \footnote{ Gaussian processes are stochastic processes satisfying by property (G). For more information on these processes and their applications, see Section 6.4 of Bishop-2006- or Chapter 2 of Rasmussen-Williams-2006-. } This process can be characterized by
For convenience, I define a variance function $V:[0,\infty)\to\mathbb{R}$ by $V(x)\equiv C(x,x)$ for all $x\ge0$. The values of $m(x)$, $V(x)$, and $C(x,x')$ are known for all $x,x'\ge0$, but the values of $f(x)$ are not.
Property (M) comes from $\{f(x)\}_{x\ge0}$ being a sample path of a Markov process. It allows me to focus on the conditional distribution of $f(x)$ given at most two values $f(x_i)$: those with $x_i$ closest to $x$. Lemma (ref) characterizes this conditional distribution.
\ifbodyproofs\fi
For example, suppose $x_k\le x\le x_{k+1}$ for some $k<n$. Lemma (ref) characterizes the distributions of $f(x)\mid f(x_{k+1})$ and $f(x)\mid f(x_k),f(x_{k+1})$ when the values of $f(x_k)$ and $f(x_{k+1})$ are known. However, variation in the errors $\varepsilon_i$ makes the values of $f(x_k)$ and $f(x_{k+1})$ unknown. So the conditional means of $f(x)\mid f(x_{k+1})$ and $f(x)\mid f(x_k),f(x_{k+1})$ given $\mathcal{D}$ are random. But their conditional variances given $\mathcal{D}$ are not random; by Lemma (ref), these variances depend on only the known values of the variance function $V:[0,\infty)\to\mathbb{R}$ and covariance function $C:[0,\infty)\to\mathbb{R}$.
Theorem (ref) refines Lemma (ref) by imposing properties (G) and (M). Specifically, it characterizes the mean and variance of $f(x)\mid\mathcal{D}$ when $\{f(x)\}_{x\ge0}$ is a sample path of an arbitrary Gauss-Markov process. These moments depend on the location of $x$ relative to the points $x_i$ at which $\mathcal{D}$ contains noisy observations $y_i$ of $f(x_i)$.
\ifbodyproofs\fi
If $x\le x_1$, then the estimate $\hat{f}_\mathcal{D}(x)\equiv\operatorname{E}[f(x)\mid\mathcal{D}]$ of $f(x)$ is linear in the estimate $\hat{f}_\mathcal{D}(x_1)$ of $f(x_1)$. Conversely, if $x\ge x_n$, then $\hat{f}_\mathcal{D}(x)$ is linear in the estimate $\hat{f}_\mathcal{D}(x_n)$ of $f(x_n)$. In both cases, we can construct $\hat{f}_\mathcal{D}(x)$ in two steps:
For example, suppose $\mathcal{D}=\{(x_1,y_1)\}$ contains one observation $y_1=f(x_1)+\varepsilon_1$ with error $\varepsilon_1\sim\mathcal{N}(0,\sigma_\varepsilon^2)$. Then \[ f(x_1)\mid\mathcal{D}_1\sim\mathcal{N}\mathopen{}\mathclose\bgroup\oldleft(m(x_1)+\frac{V(x_1)}{V(x_1)+\sigma_\varepsilon^2}(y_1-m(x_1)),\ \mathopen{}\mathclose\bgroup\oldleft(\frac{1}{V(x_1)}+\frac{1}{\sigma_\varepsilon^2}\aftergroup\egroup\oldright)^{-1}\aftergroup\egroup\oldright) \] by Lemma (ref), and so
is linear in the deviation $(y_1-m(x_1))$ of $y_1$ from its mean $m(x_1)$. Intuitively, this deviation provides information about the difference $(f(x_1)-m(x_1))$, which the factor $C(x_1,x)/(V(x)+\sigma_\varepsilon^2)$ translates into information about the difference $(f(x)-m(x))$. This factor is larger when the covariance $C(x_1,x)$ of $y_1$ and $f(x)$ is larger, and when the variance $V(x)+\sigma_\varepsilon^2$ of $y_1$ is smaller.
If $n>1$ and $x_1\le x\le x_n$, then we can construct $\hat{f}_\mathcal{D}(x)$ in three steps:
The interpolant $\hat{f}_\mathcal{D}(x)$ is a weighted sum of the mean $m(x)$, the deviation $(\hat{f}_\mathcal{D}(x_k)-m(x_k))$, and the deviation $(\hat{f}_\mathcal{D}(x_{k+1})-m(x_{k+1}))$. The weights on the two deviations depend on the (co)variances of $f(x)$, $f(x_k)$, and $f(x_{k+1})$, as well as the (co)variances of $\varepsilon_k$ and $\varepsilon_{k+1}$.
Theorem (ref) expresses the moments of $f(x)\mid\mathcal{D}$ in terms of the moments of the $f(x_i)\mid\mathcal{D}$. The latter moments can be computed analytically if $n$ is small (e.g., as in the case with $n=1$ above), or numerically if $n$ is large. The expressions in Theorem (ref) reveal how changing the moments of the $f(x_i)\mid\mathcal{D}$ changes the moments of $f(x)\mid\mathcal{D}$.
I now consider a specific Gauss-Markov process: a Brownian motion with known drift $\mu\in\mathbb{R}$ and scale $\sigma\ge0$, and unknown initial value $f(0)\sim\mathcal{N}(\mu_0,\sigma_0^2)$. \footnote{ So the sample path $\{f(x)\}_{x\ge0}$ solves the stochastic differential equation \[ \mathrm{d} f(x)=\mu\,\mathrm{d} x+\sigma\,\mathrm{d} W(x), \] where $\{W(x)\}_{x\ge0}$ is an unknown sample path of a (standard) Wiener process. This process has initial value $W(0)=0$ and iid Gaussian increments $\mathrm{d} W(x)\equiv W(x+\mathrm{d} x)-W(x)\sim\mathcal{N}(0,\mathrm{d} x)$. } This process has mean and covariance functions defined by \[ m(x)=\mu_0+\mu x \] and \[ C(x,x')=\sigma_0^2+\sigma^2\min\{x,x'\} \] for all $x,x'\ge0$. \footnote{ Thus $V(x)=\sigma_0^2+\sigma^2 x$ for all $x\ge0$. } Substituting these expressions into Theorem (ref) yields the following result.
\ifbodyproofs\fi
If $\{f(x)\}_{x\ge0}$ is a sample path of a Brownian motion, then the estimate $\hat{f}_\mathcal{D}(x)\equiv\operatorname{E}[f(x)\mid\mathcal{D}]$ of $f(x)$ is piecewise linear in $x\ge0$. For example, consider the canonical case in which the initial value $f(0)=\mu_0$ is known. As $x$ increases from zero, the estimate \[ \hat{f}_\mathcal{D}(x)=
\] interpolates linearly between $(0,\mu_0)$ and $(x_1,\hat{f}_\mathcal{D}(x_1))$, then between $(x_1,\hat{f}_\mathcal{D}(x_1))$ and $(x_2,\hat{f}_\mathcal{D}(x_2))$, then between $(x_2,\hat{f}_\mathcal{D}(x_2))$ and $(x_3,\hat{f}_\mathcal{D}(x_3))$, and so on until $(x_n,\hat{f}_\mathcal{D}(x_n))$. Beyond this point, the data $\mathcal{D}$ provide no information about $f(x)$. Consequently, the estimate $\hat{f}_\mathcal{D}(x)$ treats $\{f(x)\}_{x\ge x_n}$ as a Brownian motion with drift $\mu$, scale $\sigma$, and (possibly random) initial value $\hat{f}_\mathcal{D}(x_n)$.
I illustrate this behavior in Figure (ref).
It shows how $\hat{f}_\mathcal{D}(x)$ varies with $x\ge0$ when $(\mu_0,\mu,\sigma_0,\sigma)=(0,0,0,1)$ and the data contain $n=4$ observations with iid errors. Figure (ref) also shows the 90% confidence interval around $\hat{f}_\mathcal{D}(x)$, constructed analytically from the conditional variances defined in Corollary (ref). This interval expands as $x$ moves away from the sampled points $x_i$. \footnote{ If $f(0)$ is known (i.e., $\sigma_0^2=0$) and the observations $y_i$ have no noise (i.e., $\varepsilon_i=0$ for each $i$), then the MSE \[ \operatorname{E}\mathopen\mathclose\bgroup\oldleft[\mathopen\mathclose\bgroup\oldleft(f(x)-\hat{f}_\mathcal{D}(x)\aftergroup\egroup\oldright)^2\mid\mathcal{D}\aftergroup\egroup\oldright]=\sigma^2
\] attains its piecewise maxima at the midpoint of each piece. }
Suppose the data are not noisy (i.e., $\varepsilon_i=0$ for each $i$) and $x_k\le x\le x_{k+1}$ for some $k<n$. Then the distribution of $f(x)\mid\mathcal{D}$ coincides with the distribution obtained by assuming $\{f(x)\}_{x_k\le x\le x_{k+1}}$ is a sample path of a Brownian bridge with scale $\sigma$, known initial value $f(x_k)$, and known terminal value $f(x_{k+1})$. \footnote{ See, e.g., Section 5.6.B of Karatzas-Shreve-1988- for more information about Brownian bridges. } Corollary (ref)(ii) generalizes to Brownian bridges with unknown initial and terminal values. For example, suppose $x_1\le x\le x_2$ and that \[
\mid\mathcal{D}\sim\mathcal{N}\mathopen\mathclose\bgroup\oldleft(
,\,
\aftergroup\egroup\oldright) \] for some means $\mu_1,\mu_2\in\mathbb{R}$, variances $\sigma_1^2,\sigma_2^2\ge0$, and correlation $\rho\in[-1,1]$. Then $f(x)\mid\mathcal{D}$ has mean \[ \operatorname{E}[f(x)\mid\mathcal{D}]=\mu_1+\frac{x-x_1}{x_2-x_1}\mathopen{}\mathclose\bgroup\oldleft(\mu_2-\mu_1\aftergroup\egroup\oldright) \] and variance
Choosing $\sigma_1=0$ and $\sigma_2=0$ yields the mean and variance for the Brownian bridge on $[x_1,x_2]$ with $f(x_1)=\mu_1$ and $f(x_2)=\mu_2$.