EconBase
← Back to paper

Continuity of the Distribution Function of the argmax of a Gaussian Process

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

55,570 characters · 17 sections · 46 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Continuity of the Distribution Function of the $arg\,max$ of a Gaussian Process

\setcounter{page}{0}\thispagestyle{empty}

abstractCertain extremum estimators have asymptotic distributions that are non-Gaussian, yet characterizable as the distribution of the $\operatorname*{arg\,max}$ of a Gaussian process. This paper presents high-level sufficient conditions under which such asymptotic distributions admit a continuous distribution function. The plausibility of the sufficient conditions is demonstrated by verifying them in three examples, namely maximum score estimation, empirical risk minimization, and threshold regression estimation. In turn, the continuity result buttresses several recently proposed inference procedures whose validity seems to require a result of the kind established herein. A notable feature of the high-level assumptions is that one of them is designed to enable us to employ the Cameron-Martin theorem. In a leading special case, the assumption in question is demonstrably weak and appears to be close to minimal.

Keywords: Cameron-Martin Theorem, Cube Root Asymptotics, Gaussian Processes, Reproducing Kernel Hilbert Space.

Introduction

Certain extremum estimators have asymptotic distributions that are non-Gaussian, yet characterizable as the distribution of the $\operatorname*{arg\,max}$ of a Gaussian process. To fix ideas, letting $\bm{\theta}_0\in\mathbb{R}^d$ denote a parameter (vector) of interest, the estimators $\hat{\bm{\theta}}_n$ in question satisfy

equation[equation omitted — 194 chars of source]

where $n$ is the sample size, $r_n$ is a rate of convergence, $\rightsquigarrow$ denotes weak convergence (as $n\to\infty$), and $\mathcal{G}$ is a Gaussian process admitting a unique maximizer (over $\mathbb{R}^d$) whose distribution is non-Gaussian. The seminal work of Kim-Pollard_1990_AoS was concerned with (cube root asymptotic) cases where $r_n=\sqrt[3]{n}$ and the mean function of $\mathcal{G}$ is a quadratic form, but subsequent work \citep*[e.g.,][]{Hansen_2000_ECMA,Lai-Lee_2005_JASA,Lee-Liao-Seo-Shin_2021_AoS,Lee-Pun_2006_JASA,Lee-Yang_2020_AoS,Westling-Carone_2020_AoS,Yu-Fan_2021_JBES} has documented the relevance of allowing for the extra flexibility afforded by the more general formulation in (ref).

Letting $\mu$ and $\mathcal{C}$ denote the mean function and covariance kernel of $\mathcal{G}$ and defining

equation[equation omitted — 233 chars of source]

our goal in this paper is to give conditions on $\mu$ and $\mathcal{C}$ that imply continuity of $F_{\hat{\mathbf{s}}}$. Continuity of $F_{\hat{\mathbf{s}}}$ is useful when the goal is to use $\hat{\bm{\theta}}_n$ to construct confidence regions. For instance, vanderVaart_1998_book assumes continuity when establishing validity of bootstrap-based confidence intervals; see also \citet*[Section 1.2]{Politis-Romano-Wolf_1999_Book}. Moreover, and relatedly, it follows from Polya's theorem that if (ref) holds and if $F_{\hat{\mathbf{s}}}$ is continuous, then the the probability laws of $r_n(\hat{\bm{\theta}}_n-\bm{\theta}_0)$ converge to the law with distribution function $F_{\hat{\mathbf{s}}}$ not only in the bounded Lipschitz metric (or any other metric metrizing weak convergence), but also in the Kolmogorov metric; that is, we have a result of the form

equation[equation omitted — 230 chars of source]

When $\mu$ is a quadratic form and $\mathcal{C}$ is a bilinear form, the distribution of $\hat{\mathbf{s}}$ is Gaussian. More generally, under mild conditions on $\mu$ the distribution of $\hat{\mathbf{s}}$ is that of a transformation of a Gaussian vector when $\mathcal{C}$ is a bilinear form, implying in particular that the properties of $F_{\hat{\mathbf{s}}}$ can be deduced by means of a change of variables argument. Two other special cases where a complete characterization of $F_{\hat{\mathbf{s}}}$ is available are when $d=1$, $\mathcal{C}$ is the covariance kernel of a two-sided Brownian motion, and $\mu$ is proportional to either the absolute value function or the square function. In both cases, the distribution of $\hat{\mathbf{s}}$ is that of a scalar multiple of a random variable with a well-known continuous distribution. Somewhat more generally, \citet*[Lemma A.2]{Cattaneo-Jansson-Nagasawa_2024_AoS} gave conditions on $\mu$ under which $F_{\hat{\mathbf{s}}}$ is continuous when $d=1$ and $\mathcal{C}$ is the covariance kernel of a two-sided Brownian motion. On the other hand, little (if anything) appears to be known about the properties of $F_{\hat{\mathbf{s}}}$ when $d>1$ and $\mathcal{C}$ is not bilinear.

In this paper we close this gap by presenting sufficient conditions for continuity of $F_{\hat{\mathbf{s}}}$ that do not require $d=1$ and are applicable (only) when $\mathcal{C}$ is not bilinear. Proceeding under the assumption that $d=1$, the proof of Cattaneo-Jansson-Nagasawa_2024_AoS establishes continuity of $F_{\hat{\mathbf{s}}}$ by showing that the distribution of $\hat{\mathbf{s}}$ is atomless (under the additional assumptions of the lemma). The method of proof can be adapted to give conditions under which the distribution of $\hat{\mathbf{s}}$ is atomless also when $d>1$, but when $d>1$ a distribution can be atomless even if the associated distribution function is discontinuous. Establishing continuity of $F_{\hat{\mathbf{s}}}$ when $d>1$ therefore requires a fundamentally different method of proof than that employed by Cattaneo-Jansson-Nagasawa_2024_AoS. The differences in proof strategies are reflected also in the assumptions under which the proofs proceed. Notably, one of the conditions imposed in this paper explicitly involves both $\mu$ and $\mathcal{C}$ and requires that for every $N\in\mathbb{N}$, restriction of $\mathcal{G}$ to $[-N,N]^d$ has a mean function that belongs to the reproducing kernel Hilbert space (RKHS) of its covariance kernel. By the Cameron-Martin theorem, if a Gaussian process has a mean belonging to the RKHS of its covariance kernel, then its induced probability measure and the probability measure induced by its centered version are mutually absolutely continuous. The proof of our main result uses this fact and an assumed shift equivariance property of the covariance kernel to deduce continuity of $F_{\hat{\mathbf{s}}}$.

The usefulness of our main result is illustrated by applying it to three examples: maximum score estimation, empirical risk minimization, and threshold regression estimation. Each example involves an estimator satisfying (ref) with $d$ possibly greater than one and a covariance kernel that is not bilinear. Although distinct in several ways, the examples enjoy the common feature that continuity of $F_{\hat{\mathbf{s}}}$ can be shown by verifying the conditions of our main result. In particular, the condition that the mean function belongs to the RKHS of the covariance kernel can be verified by following a general strategy outlined in Lifshits_1995_Book.

In addition to facilitating the justification of certain large-sample inference procedures based on distributional approximations of the form (ref), our paper sheds new light on the canonical problem of characterizing the distributional properties of the $\operatorname*{arg\,max}$ of a Gaussian process. That problem is substantially different from the well-studied problem of understanding the distributional properties of the maximum itself, where the $d=1$ case is mostly settled Lifshits_1995_Book, the multidimensional case is fairly well understood Azais-Wschebor_2005_AOAP, and where, more generally, continuity of the distribution function of the maximum can be established with the help of anti-concentration results \citep*[e.g.,][]{Chernozhukov-Chetverikov-Kato_2015_PTRF}. However, as noted by Samorodnitsky-Shen_2013_AOP, “very little is known about the random location of the supremum” of a Gaussian process. In the multidimensional case, we are only aware of Azais-Chassan_2020_SPA, which shows that the distribution admits a density under the assumption that the sample paths are twice differentiable.

The remainder of the paper proceeds as follows. Section (ref) introduces our three examples. Our main result is presented in Section (ref), while Section (ref) outlines a general strategy for verifying the conditions of the main result and demonstrates how to apply it in the examples. Finally, Section (ref) compares our results with the known results alluded to in the third paragraph of this section.

Motivating Examples

The class of estimators satisfying (ref) is rich, containing examples in econometrics, statistics, and other data science disciplines. To further motivate our work, this section presents three representative examples.

Maximum Score

Suppose $\{(y_i,w_i,\mathbf{x}_i')'\}_{i=1}^n$ is a random sample from the distribution of a vector $(y,w,\mathbf{x}')'$ generated by the semiparametric binary response model

equation*[equation* omitted — 116 chars of source]

where $\mathbbm{1}\{\cdot\}$ is the indicator function, $w,u\in\mathbb{R}$ and $\mathbf{x}\in\mathbb{R}^d$ are random variables, and $\bm{\theta}_0\in\Theta\subseteq\mathbb{R}^d$ is the parameter of interest. Manski_1975_JoE introduced the maximum score estimator of $\bm{\theta}_0$, which is any maximizer $\hat\bm{\theta}_n$ of

equation*[equation* omitted — 92 chars of source]

with respect to $\bm{\theta}\in\Theta$. Using the methods of Kim-Pollard_1990_AoS, Abrevaya-Huang_2005_ECMA gave regularity conditions under which (ref) holds with $r_n=\sqrt[3]{n}$ and $\mathcal{G}$ being a Gaussian process whose mean function and covariance kernel take the form

equation*[equation* omitted — 215 chars of source]

and

equation*[equation* omitted — 204 chars of source]

respectively, where $f_{u|w,\mathbf{x}}$ and $f_{w|\mathbf{x}}$ denote conditional (Lebesgue) densities, and where $\mathcal{C}_{\mathtt{BM}}$ is the covariance kernel of a two-sided standard Brownian motion; that is,

equation*[equation* omitted — 128 chars of source]

with $\operatorname*{sgn}(\cdot)$ denoting the sign function.

When $d=1$, because $\mu$ is quadratic and $\mathcal{C}$ is a scalar multiple of $\mathcal{C}_{\mathtt{BM}}$, it follows from vanderVaart-Wellner_2023_book that the distribution of $\hat{\mathbf{s}}$ is that of a scalar multiple of a random variable with a well-known continuous distribution, namely the Chernoff_1964_AISM distribution. For $d>1$, on the other hand, it would appear to be an open question whether $F_{\hat{\mathbf{s}}}$ is continuous. We provide an affirmative answer to that question below, hereby buttressing a variety of inference procedures based on the maximum score estimator.

For specificity, consider the procedure of Cattaneo-Jansson-Nagasawa_2020_ECMA. That paper proposed a bootstrap-based estimator $\tilde{\bm{\theta}}_n^*$ and gave conditions under which this estimator satisfies

equation*[equation* omitted — 173 chars of source]

where $\rightsquigarrow_\mathbb{P}$ denotes weak convergence in probability. Because $F_{\hat{\mathbf{s}}}$ is continuous, the displayed result can be combined with (ref) to yield a bootstrap consistency result of the form

equation*[equation* omitted — 238 chars of source]

where $\mathbb{P}_n^*$ is the bootstrap probability measure. As a consequence, for any $\bm{\lambda}\in\mathbb{R}^d$, defining

equation*[equation* omitted — 208 chars of source]

vanderVaart_1998_book shows that the equal-tailed “percentile” interval

equation*[equation* omitted — 215 chars of source]

is a confidence interval (for $\bm{\lambda}'\bm{\theta}_0$) of asymptotic level $1-\alpha$:

equation*[equation* omitted — 151 chars of source]

With minor modifications, similar comments apply to the inference procedures proposed by Delgado-RodriguezPoo-Wolf_2001_EL, Hong-Li_2020_AoS, Jun-Pinkse-Wan_2015_JoE, Lee-Yang_2020_AoS, and Patra-Seijo-Sen_2018_JoE.

Collectively, the inference procedures mentioned in the previous paragraph therefore constitute asymptotically valid alternatives to inference procedures based on the smoothed maximum score estimator of Horowitz_1992_ECMA. (A finite-sample inference procedure for the semiparametric binary response model has recently been proposed by Rosen-Ura_2025_REStud.)

Empirical Risk Minimization

Mohammadi-vandeGeer_2005_JMLR considered the classification problem of estimating the minimizer $\bm{\theta}_0\in\Theta\subseteq\mathbb{R}^d$ of the classification error $\mathbb{P}[y\neq h_{\bm{\theta}}(x)]$ with respect to $\bm{\theta}\in\Theta$, where $y\in \{-1,1\}$ is a binary outcome, $x\in\mathcal{X}\subseteq\mathbb{R}$ is a scalar feature, and $\{h_{\bm{\theta}}:\bm{\theta}\in\Theta\}$ is a collection of classifiers. Given a random sample $\{(y_i,x_i)\}_{i=1}^n$ from the distribution of $(y,x)$, an empirical risk minimizer is a minimizer $\hat{\bm{\theta}}_n$ of

equation*[equation* omitted — 79 chars of source]

Setting $\mathcal{X}=[0,1]$ and specializing to the case where the classifiers are of the form

equation*[equation* omitted — 122 chars of source]

for $\bm{\theta}=(\theta_1,\dots,\theta_d)'\in\Theta=\{\bm{\theta}\in [0,1]^d: 0=\theta_0\leq \theta_1\leq \dots\leq \theta_d\leq \theta_{d+1}=1\}$, Mohammadi-vandeGeer_2005_JMLR gave conditions under which (ref) holds with $r_n=\sqrt[3]{n}$ and $\mathcal{G}$ being a Gaussian process whose mean function and covariance kernel take the form

equation[equation omitted — 187 chars of source]

and

equation[equation omitted — 241 chars of source]

respectively, where $\bm{\theta}_0=(\theta_{0,1},\dots,\theta_{0,d})',\mathbf{s}=(s_1,\dots,s_d)',\mathbf{t}=(t_1,\dots,t_d)'$, $f$ is a Lebesgue density of $x$, $p(x)= \text{d} \mathbb{P}[y=1|x]/\text{d}x$, and where the assumptions imposed on the model ensure that $(-1)^{\ell}p(\theta_{0,\ell})f(\theta_{0,\ell})<0$ for every $\ell=1,\dots,d$.

This example is similar to the maximum score example insofar as when $d=1$, the distribution of $\hat{\mathbf{s}}$ is that of a scalar multiple of a random variable with a Chernoff distribution. In fact, also when $d>1$, the elements of $\hat{\mathbf{s}}=(\hat{s}_1,\dots,\hat{s}_d)'$ are mutually independent, each having a distribution which is that of a scalar multiple of a random variable with a Chernoff distribution. Indeed, letting $\mathcal{G}_1,\dots,\mathcal{G}_d$ be mutually independent Gaussian processes with mean functions $\mu_1,\dots,\mu_d$ and covariance kernels $\mathcal{C}_1,\dots,\mathcal{C}_d$, respectively, $\mathcal{G}$ admits the representation $\mathcal{G}(\mathbf{s})=\sum_{\ell=1}^d \mathcal{G}_\ell(s_\ell)$, implying in particular that

equation*[equation* omitted — 148 chars of source]

In other words, this example has special structure that can be used to obtain a definitive characterization of the distribution of $\hat{\mathbf{s}}$ without utilizing new tools. It is nevertheless of interest to explore the ease with which the technology developed in this paper can be deployed to establish continuity of $F_{\hat{\mathbf{s}}}$ in this example. In particular, it is of interest to explore whether the (effective) “dimension reduction” permitted by this example can be leveraged when verifying the conditions of Theorem (ref) below.

Threshold Regression

Consider the threshold regression model

equation*[equation* omitted — 164 chars of source]

where $y\in\mathbb{R}$ is a dependent variable, $\mathbf{x}\in\mathbb{R}^k$ is a (possibly) vector-valued regressor, $q$ is a threshold variable, $\mathbf{w}\in\mathbb{R}^d$ is a (possibly) vector-valued factor governing the threshold cutoff, and where, borrowing ideas from the change-point literature Bai_1997_REStat, $\bm{\delta}_n$ is a “threshold effect” whose magnitude vanishes with $n$. The present model (as well as distinct generalizations thereof) has been studied by Lee-Liao-Seo-Shin_2021_AoS and Yu-Fan_2021_JBES, and differs from the model considered in Hansen_2000_ECMA by allowing the factor $\mathbf{w}$ to be non-constant.

Given a random sample $\{(y_i,\mathbf{x}_i',q_i,\mathbf{w}_i')'\}_{i=1}^n$ from the distribution of $(y,\mathbf{x}',q,\mathbf{w}')'$, a least squares estimator $(\hat{\bm{\beta}}_n',\hat{\bm{\delta}}_n',\hat{\bm{\theta}}_n')'$ of $(\bm{\beta}_0',\bm{\delta}_n',\bm{\theta}_0')'$ is a minimizer of

equation*[equation* omitted — 144 chars of source]

over $(\bm{\beta}',\bm{\delta}',\bm{\theta}')'\in\mathbb{R}^{2k+d}$. Assuming $\|\bm{\delta}_n\|\rightarrow 0$, $n\|\bm{\delta}_n\|^2\rightarrow\infty$, and $\bar{\bm{\delta}}_n=\bm{\delta}_n/\|\bm{\delta}_n\|\rightarrow \bar{\bm{\delta}}$ (for some $\bar{\bm{\delta}}\in \mathbb{S}^{d-1}=\{\bar{\bm{\delta}}\in\mathbb{R}^d:\|\bar{\bm{\delta}}\|=1\}$), Yu-Fan_2021_JBES gave conditions under which (ref) holds with $r_n=n\|\bm{\delta}_n\|^2$ and $\mathcal{G}=\mathcal{G}(\cdot;\bar{\bm{\delta}})$ being a Gaussian process whose mean function and covariance kernel take the form

equation*[equation* omitted — 307 chars of source]

and

equation*[equation* omitted — 384 chars of source]

respectively, where $\mathbb{E}_{\cdot|q,\mathbf{w}}$ denotes conditional expectation.

When $d=1$, the distribution of $\hat{\mathbf{s}}$ is that of a scalar multiple of a random variable with a (known) continuous distribution Hansen_2000_ECMA. For $d>1$, on the other hand, it would appear to be an open question whether $F_{\hat{\mathbf{s}}}$ is continuous. We provide an affirmative answer to that question below.

\paragraph*{Remark.} As a by-product, continuity of the limiting distribution function can be used to show that if $\|\bm{\delta}_n\|\rightarrow 0$ and if $n\|\bm{\delta}_n\|^2\rightarrow\infty$, then (whether or not $\bar{\bm{\delta}}_n$ is convergent in $\mathbb{S}^{d-1}$) we have

equation*[equation* omitted — 270 chars of source]

Main Result

As before, let $\mathcal{G}$ be a Gaussian process on $\mathbb{R}^d$ with mean function $\mu$ and covariance kernel $\mathcal{C}$. Also, for any $N\in\mathbb{N}$, let $\mathcal{G}_N$ be the restriction of $\mathcal{G}$ to $[-N, N]^d$, let $\mu_N$ and $\mathcal{C}_N$ be the mean function and covariance kernel of $\mathcal{G}_N$, and let $\mathscr{H}_N$ be the RKHS of $\mathcal{C}_N$ Gine-Nickl_2016_Book.

The following high-level assumption holds in each of the examples of Section (ref).

assumption\begin{enumerate}[label=(\roman*)] • With probability one, $\mathcal{G}$ has continuous sample paths and admits a maximizer over $\mathbb{R}^d$. • For any $\mathbf{s},\mathbf{t}\in\mathbb{R}^d$ and any $\mathbf{h}\in\mathbb{R}^{d}\backslash\{\mathbf{0}\}$, $\mathcal{C}(\mathbf{h},\mathbf{h})>0$ and \begin{equation*} \mathcal{C}(\mathbf{h}+\mathbf{s},\mathbf{h}+\mathbf{t})-\mathcal{C}(\mathbf{h}+\mathbf{s},\mathbf{h})-\mathcal{C}(\mathbf{h},\mathbf{h}+\mathbf{t})+\mathcal{C}(\mathbf{h},\mathbf{h})=\mathcal{C}(\mathbf{s},\mathbf{t}). \end{equation*} • For any $N\in\mathbb{N}$, $\mu_N\in \mathscr{H}_N$. \end{enumerate}

Part (ref) guarantees existence (with probability one) of a maximizer of $\mathcal{G}$ over $\mathbb{R}^d$, and over any compact set $S\subset\mathbb{R}^d$. By Kim-Pollard_1990_AoS, these maximizers are unique provided $\mathbb{V}[\mathcal{G}(\mathbf{s})-\mathcal{G}(\mathbf{t})]\neq0$ for every $\mathbf{s}\neq\mathbf{t}$. Part (ref) gives sufficient conditions for this non-degeneracy condition to hold and furthermore implies that the centered process $\mathcal{G}^\mu=\mathcal{G}-\mu$ is shift equivariant in the sense that the law of the process $\mathcal{G}^\mu(\mathbf{h}+\cdot)-\mathcal{G}^\mu(\mathbf{h})$ is the same for every $\mathbf{h}\in\mathbb{R}^d$. An immediate implication of shift equivariance is that $\mathbb{V}[\mathcal{G}(0)]=0$. A slightly more subtle implication is recorded in the following lemma.

lemmaSuppose $\mathcal{G}$ has continuous sample paths and suppose Assumption (ref)(ref) holds. For any $\mathbf{h}\in\mathbb{R}^d$, any measurable set $T\subseteq\mathbb{R}^d$, and any compact set $S\subset\mathbb{R}^d$, \begin{equation*} \mathbb{P}\left[\operatorname*{arg\,max}_{\mathbf{s}\in S} \mathcal{G}^\mu(\mathbf{s}) \in T\right] = \mathbb{P}\left[\operatorname*{arg\,max}_{\mathbf{s}\in S+\mathbf{h}} \mathcal{G}^\mu(\mathbf{s}) \in T +\mathbf{h}\right]. \end{equation*}
myproof{Lemma (ref)} By change of variables and shifting by a constant, \begin{equation*} \operatorname*{arg\,max}_{\mathbf{s}\in S +\mathbf{h}} \mathcal{G}^\mu(\mathbf{s}) =\operatorname*{arg\,max}_{\mathbf{s}\in S} \mathcal{G}^\mu(\mathbf{h}+\mathbf{s}) +\mathbf{h} =\operatorname*{arg\,max}_{\mathbf{s}\in S} \{\mathcal{G}^\mu(\mathbf{h}+\mathbf{s})-\mathcal{G}^\mu(\mathbf{h})\}+\mathbf{h}. \end{equation*} The desired conclusion follows from shift equivariance of $\mathcal{G}^\mu$.

By the Cameron-Martin theorem Gine-Nickl_2016_Book, part (ref) of Assumption (ref) ensures that the probability measures associated with $\mathcal{G}_N$ and $\mathcal{G}^\mu_N=\mathcal{G}_N-\mu_N$ are mutually absolutely continuous for any $N\in\mathbb{N}$. Along with Lemma (ref), that property plays a key role in our proof of the following result.

theoremUnder Assumption (ref), $F_{\hat{\mathbf{s}}}$ in (ref) is continuous.
myproof{Theorem (ref)} For any compact set $S\subset\mathbb{R}^d$, parts (ref) and (ref) of Assumption (ref) guarantee existence (with probability one) of a unique maximizer of $\mathcal{G}^\mu$ over $S$, uniqueness being a consequence of Kim-Pollard_1990_AoS. This observation will be used repeatedly without further mention. The (joint) distribution function of $\hat{\mathbf{s}}=(\hat{s}_1,\dots,\hat{s}_d)'$ is continuous if and only if each of its marginal distribution functions is continuous. Fixing $\ell\in\{1,\dots,d\}$ and $t\in\mathbb{R}$, the proof can therefore be completed by showing that \begin{equation} \mathbb{P}\left[\hat{s}_\ell=t\right]=0. \end{equation} Defining $N_t=\lceil|t|\rceil+1$, letting $\mathbf{e}_\ell$ denote the $\ell$th standard basis vector of $\mathbb{R}^d$, and noting that \begin{equation*} \left\{\hat{s}_\ell=t\right\} = \left\{\mathbf{e}_\ell'\operatorname*{arg\,max}_{\mathbf{s}\in\mathbb{R}^d}\mathcal{G}(\mathbf{s})=t\right\} \subseteq \bigcup_{N=N_t}^{\infty} \left\{\mathbf{e}_\ell'\operatorname*{arg\,max}_{\mathbf{s}\in[-N,N]^d}\mathcal{G}(\mathbf{s})=t\right\}, \end{equation*} a sufficient condition for (ref) to hold is that \begin{equation*} \mathbb{P}\left[\mathbf{e}_\ell'\operatorname*{arg\,max}_{\mathbf{s}\in[-N,N]^d}\mathcal{G}(\mathbf{s})=t\right]=0 \qquad for every N \geq N_t. \end{equation*} By the Cameron-Martin theorem, under Assumption (ref)(ref) the displayed condition is equivalent to \begin{equation} \mathbb{P}\left[\mathbf{e}_\ell'\operatorname*{arg\,max}_{\mathbf{s}\in[-N,N]^d}\mathcal{G}^\mu(\mathbf{s})=t\right]=0 \qquad for every N \geq N_t. \end{equation} Fixing $N\geq N_t$ and $J\geq 2$, let \begin{equation*} S_j=\left\{(s_1,\dots,s_d)\in[-N,N]^d:-N_t+\frac{1}{2}\frac{j-1}{J-1} \leq s_\ell \leq N_t+\frac{1}{2}\left(\frac{j-1}{J-1}-1\right)\right\} \end{equation*} for $j\in\{1,\dots,J\}$. Noting that \begin{align*} [-N,N]^d &\supseteq \bar{S} = \{(s_1,\dots,s_d)\in[-N,N]^d:-N_t \leq s_\ell \leq N_t\}=\cup_{j=1}^{J}S_j\\ &\supset\cap_{j=1}^{J}S_j=\{(s_1,\dots,s_d)\in[-N,N]^d:-N_t+1/2 \leq s_\ell \leq N_t-1/2\}=S, \end{align*} we have \begin{align*} \mathbb{P}\left[\mathbf{e}_\ell'\operatorname*{arg\,max}_{\mathbf{s}\in[-N,N]^d}\mathcal{G}^\mu(\mathbf{s})=t\right] &\leq \mathbb{P}\left[\mathbf{e}_\ell'\operatorname*{arg\,max}_{\mathbf{s}\in\bar{S}}\mathcal{G}^\mu(\mathbf{s})=t\right] \leq \mathbb{P}\left[\mathbf{e}_\ell'\operatorname*{arg\,max}_{\mathbf{s}\in S_1}\mathcal{G}^\mu(\mathbf{s})=t\right]\\ &= \frac{1}{J}\sum_{j=1}^{J}\mathbb{P}\left[\mathbf{e}_\ell'\operatorname*{arg\,max}_{\mathbf{s}\in S_j}\mathcal{G}^\mu(\mathbf{s})=t+\frac{1}{2}\frac{j-1}{J-1}\right]\\ &\leq \frac{1}{J}\sum_{j=1}^{J}\mathbb{P}\left[\mathbf{e}_\ell'\operatorname*{arg\,max}_{\mathbf{s}\in \underline{S}}\mathcal{G}^\mu(\mathbf{s})=t+\frac{1}{2}\frac{j-1}{J-1}\right]\\ &= \frac{1}{J}\mathbb{P}\left[\mathbf{e}_\ell'\operatorname*{arg\,max}_{\mathbf{s}\in \underline{S}}\mathcal{G}^\mu(\mathbf{s})\in\left\{t+\frac{1}{2}\frac{j-1}{J-1}:1\leq j \leq J\right\}\right]\leq \frac{1}{J}, \end{align*} where the first equality uses Lemma (ref) and the second equality uses uniqueness of the maximizer of $\mathcal{G}^\mu$ over $\underline{S}$. Since $J \geq 2$ was arbitrary, (ref) follows.

Verification of Assumption (ref)

Assumption (ref)(ref)

The continuity part of Assumption (ref)(ref) is mild and usually trivial to verify. Under continuity and assuming that $\mathbb{P}[\mathcal{G}(\mathbf{0})=0]=1$, a high-level sufficient condition for existence of a maximizer of $\mathcal{G}$ over $\mathbb{R}^d$ is that

equation[equation omitted — 146 chars of source]

In turn, proceeding as in the proof of Kim-Pollard_1990_AoS it can be shown that if the covariance kernel satisfies the (self-similarity) property that for some $H>0$,

equation[equation omitted — 221 chars of source]

then (ref) is implied by the following mild condition on the mean function:

equation[equation omitted — 189 chars of source]

The assumption $\mathbb{P}[\mathcal{G}(\mathbf{0})=0]=1$ holds if (and only if) $\mu(\mathbf{0})=0=\mathcal{C}(\mathbf{0},\mathbf{0})$ and is therefore satisfied in each of the examples of Section (ref). Likewise, the conditions (ref) and (ref) are both fairly primitive. Also, by inspection, (ref) can be seen to hold with $H=1/2$ in each of the examples of Section (ref). Moreover, setting $H=1/2$, (ref) can be seen to hold with $\epsilon=3/2$ in the maximum score and empirical risk minimization examples, and with $\epsilon=1/2$ in the threshold regression example.

To explain why it is no coincidence that the self-similarity property of $\mathcal{C}$ holds in the examples of Section (ref), it may be helpful to note that in each case $\mathcal{G}$ is the weak limit of a process of the form

equation*[equation* omitted — 157 chars of source]

where $\{\mathbf{z}_i\}_{i=1}^n$ is a random sample and $m_n$ is some function (possibly depending on $n$). For instance, in the threshold regression example we have

equation*[equation* omitted — 232 chars of source]

It therefore stands to reason that $\mathcal{C}$ can be characterized as follows:

equation*[equation* omitted — 253 chars of source]

To further conclude that (ref) holds with $H=1/2$, it suffices to assume that the preceding display admits the following strengthening: for any $\eta_n>0$ with $\eta_n=O(r_n^{-1})$,

equation[equation omitted — 323 chars of source]

A characterization of the form (ref) is valid in each of the examples of Section (ref).

Assumption (ref)(ref)

By inspection, the displayed part of Assumption (ref)(ref) holds in each of the examples of Section (ref). To explain why this is no coincidence, observe that upon defining $\bm{\theta}_n=\bm{\theta}_0+\eta_n\mathbf{h}$ we have

align*[align* omitted — 1,064 chars of source]

The displayed part of Assumption (ref)(ref) is therefore valid whenever the following “local uniform” version of (ref) is valid: for any $\eta_n>0$ with $\eta_n=O(r_n^{-1})$ and any $\bm{\theta}_n=\bm{\theta}_0+O(\eta_n)$,

equation[equation omitted — 327 chars of source]

In turn, a characterization of the form (ref) is valid in each of the examples of Section (ref).

Assumption (ref)(ref)

Evaluating Assumption (ref)(ref) is usually straightforward when $\mathscr{H}_N$ is known. More generally, a viable strategy for verifying Assumption (ref)(ref) can be based on Lifshits_1995_Book. This subsection first outlines that strategy, and then demonstrates its usefulness by employing it in each of the examples in Section (ref).

General Strategy

For $N\in\mathbb{N}$, suppose $\mathscr{E}_N=\{e_N(\cdot;\mathbf{s}):\mathbf{s}\in [-N,N]^d\}$ is a model of the covariance kernel $\mathcal{C}_N$ Lifshits_1995_Book; that is, suppose that for some measure space $(\Omega_N,\mathscr{B}_N,\nu_N)$, $\mathscr{E}_N$ is a collection of elements of $L_2(\Omega_N,\mathscr{B}_N,\nu_N)$ satisfying

equation[equation omitted — 249 chars of source]

Then, as discussed in Lifshits_1995_Book, the mean function $\mu_N$ belongs to $\mathscr{H}_N$ if (and only if) it admits a function $l_N \in L_2(\Omega_N,\mathscr{B}_N,\nu_N)$ satisfying

equation[equation omitted — 202 chars of source]

Lifshits_1995_Book demonstrates how this strategy can be used to characterize $\mathscr{H}_N$ for several examples of Gaussian processes. There is no general blueprint for defining $\mathscr{E}_N$ and $l_N$ satisfying (ref)-(ref). We modify the arguments used in Lifshits_1995_Book to cover our examples.

Example: Maximum Score (Continued)

For $N\in\mathbb{N}$, let $\mathscr{B}_N$ be the Borel $\sigma$-algebra on $\Omega_N=\mathbb{R}^{1+d}$ and let $\nu_N = \lambda\times \mathbb{P}_\mathbf{x}$, where $\lambda$ is the Lebesgue measure on $\mathbb{R}$ and $\mathbb{P}_\mathbf{x}$ is the probability measure induced by $\mathbf{x}$. A direct calculation shows that (ref)-(ref) hold with $\bm{\omega}=(\omega_1,\mathbf{x}')'$,

equation*[equation* omitted — 239 chars of source]

and

equation*[equation* omitted — 229 chars of source]

Example: Empirical Risk Minimization (Continued)

For $N\in\mathbb{N}$, let $\mathscr{B}_N$ be the Borel $\sigma$-algebra on $\Omega_N=\mathbb{R}^d$ and let $\nu_N$ be the Lebesgue measure on $\mathbb{R}^d$. Then (ref)-(ref) hold with $\bm{\omega}=(\omega_1,\dots,\omega_d)'$,

equation*[equation* omitted — 307 chars of source]

and

equation*[equation* omitted — 179 chars of source]

where $u_N$ is a Lebesgue density of the uniform distribution on $[-N,N]^d$.

The strategy described in Section (ref) and followed in the previous paragraph is general and is not designed to leverage the special structure highlighted in Section (ref). Because it stands to reason that the additive separability of $\mathcal{C}_N$ induces an analogous simplification of the associated $\mathscr{H}_N$, it seems natural to ask whether such a “tensorization” (is materialized and) can in turn be exploited when verifying Assumption (ref)(ref).

The special structure (ref) of $\mathcal{C}_N$ implies that $\mathscr{H}_N$ consists of those functions that are of the form $h_N(\mathbf{s})= \sum_{\ell=1}^d h_{\ell,N}(s_\ell)$, with $h_{\ell,N}$ belonging to the RKHS of $\mathcal{C}_{\ell,N}$ for each $\ell$. Now, proceeding as in Gine-Nickl_2016_Book it can be shown that the RKHSs of $\mathcal{C}_{1,N},\dots,\mathcal{C}_{d,N}$ all (coincide and) consist of those functions on $[-N,N]$ that are zero-preserving and absolutely continuous with a square integrable (weak) derivative. In particular, $\mu_N\in\mathscr{H}_N$ when $\mu$ is given by (ref).

Example: Threshold Regression (Continued)

For $N\in\mathbb{N}$, let $\mathscr{B}_N$ be the Borel $\sigma$-algebra on $\Omega_N=\mathbb{R}^{1+d}$ and let $\nu_N = \lambda\times \mathbb{P}_{\mathbf{w}}$, where $\lambda$ is the Lebesgue measure on $\mathbb{R}$ and $\mathbb{P}_{\mathbf{w}}$ is the probability measure induced by $\mathbf{w}$. Then (ref)-(ref) hold with $\bm{\omega}=(\omega_1,\mathbf{w}')'$,

equation*[equation* omitted — 332 chars of source]

and

align*[align* omitted — 382 chars of source]

Discussion of Assumption (ref)(ref)

Interpreted as a condition on the mean function $\mu$, the strength of Assumption (ref) (ref) is inversely related to the richness of the RKHSs $\mathscr{H}_N$ generated by the covariance kernel $\mathcal{C}$. It seems natural to ask, therefore, whether certain covariance kernels are so simple that Theorem (ref) is silent about the continuity properties of $F_{\hat{\mathbf{s}}}$ for many (possibly most) interesting mean functions. Revisiting a particularly simple covariance kernel, Section (ref) provides an affirmative answer to that question.

A related, but arguably more interesting, question is whether Assumption (ref) (ref) is likely to be “close” to minimal in cases where the covariance kernel generates RKHSs that are sufficiently rich to contain many (possibly most) interesting mean functions. Revisiting another well-known covariance kernel, Section (ref) provides evidence suggesting that the answer to that question will be affirmative in certain special cases.

Bilinear Covariance Kernel

Suppose $\mathcal{C}(\mathbf{s},\mathbf{t})=\mathbf{s}'\mathbf{\Sigma}\mathbf{t}$ for some (symmetric and) positive definite $\mathbf{\Sigma}$; that is, suppose $\mathcal{C}$ is a bilinear form. Then $\mathcal{G}^{\mu}(\mathbf{s})=\mathbf{s}'\mathbf{\dot{\mathcal{G}}^{\mu}}$, where $\mathbf{\dot{\mathcal{G}}^{\mu}}=(\mathcal{G}^{\mu}(\mathbf{e}_1),\dots,\mathcal{G}^{\mu}(\mathbf{e}_d))'\thicksim \mathcal{N}(\mathbf{0},\mathbf{\Sigma})$. If also $\mu$ is a quadratic form $\mu(\mathbf{s})=-\mathbf{s}'\mathbf{\Gamma}\mathbf{s}/2$ (for some symmetric and positive definite $\mathbf{\Gamma}$), then

equation*[equation* omitted — 344 chars of source]

Thus, this case covers asymptotically normal estimators by writing the normal random limit as the argmax of a Gaussian process. More generally, under regularity conditions including invertibility of the gradient $\mathbf{\dot{\mu}}$ of $\mu$, we have $\hat{\mathbf{s}}=\mathbf{\dot{\mu}}^{-1}(-\mathbf{\dot{\mathcal{G}}^{\mu}})$, implying in turn that the distributional properties of $\hat{\mathbf{s}}$ can be deduced with the help of standard tools.

In other words, if $\mathcal{C}$ is bilinear, then conditions for continuity of $F_{\hat{\mathbf{s}}}$ can be formulated without invoking the results of this paper. In fact, it turns out that Theorem (ref) is completely silent about the case where $\mathcal{C}$ is bilinear because in that case $\mu$ satisfies Assumption (ref)(ref) if and only if it is a linear form $\mu(\mathbf{s})=\mathbf{s}'\mathbf{\dot{\mathbf{\mu}}}$ (for some $\dot{\mathbf{\mu}}\in\mathbb{R}^d$) , in which case Assumption (ref)(ref) fails because $\mathcal{G}(\mathbf{s})=\mathbf{s}'(\mathbf{\dot{\mathbf{\mu}}}+\mathbf{\dot{\mathbf{\mathcal{G}}}^{\mu}})$ does not admit a maximizer over $\mathbb{R}^d$. To summarize, our results complement existing techniques, Assumption (ref)(ref) being very restrictive precisely when the covariance kernel of $\mathcal{G}$ is so simple that no new methods are needed in order to analyze the distribution of $\hat{\mathbf{s}}$.

Two-Sided Brownian Motion

Our motivating examples have the common feature that if $d=1$, then $\mathcal{C}$ is proportional to $\mathcal{C}_{\mathtt{BM}}$. More generally, the examples have the feature that for any $d$, $\mathcal{C}$ is a (linear) functional of $\mathcal{C}_{\mathtt{BM}}$, a feature which in turn would appear to be shared by most other examples of estimators satisfying (ref) with a covariance kernel that is not bilinear. It is therefore of interest to further investigate the continuity properties of $F_{\hat{\mathbf{s}}}$ in the special case where $\mathcal{G}^{\mu}$ is a two-sided Brownian motion.

Accordingly, suppose $d=1$ and suppose $\mathcal{C}=\sigma^2\mathcal{C}_{\mathtt{BM}}$ for some $\sigma^2>0$. Also in this case Assumption (ref)(ref) reduces to a primitive condition on $\mu$. Indeed, proceeding as in Gine-Nickl_2016_Book it can be shown that Assumption (ref)(ref) holds if and only if $\mu$ is zero-preserving and absolutely continuous with a locally square integrable (weak) derivative.

In the leading special case where

equation[equation omitted — 102 chars of source]

Theorem (ref) therefore implies that $F_{\hat{\mathbf{s}}}$ is continuous whenever $\gamma>1/2$. The same condition on $\gamma$ is necessary and sufficient in order to deduce continuity of $F_{\hat{\mathbf{s}}}$ by applying Cattaneo-Jansson-Nagasawa_2024_AoS, which replaces Assumption (ref)(ref) with the Brownian motion-specific assumption

equation[equation omitted — 158 chars of source]

Maintaining the assumption that $\mathcal{G}^{\mu}$ is a two-sided Brownian motion, but looking beyond mean functions of the form (ref), the assumption (ref) is slightly more general than Assumption (ref)(ref). To see this, notice on the one hand that if $\mu$ is absolutely continuous with (weak) derivative $\dot{\mu}$, then

equation*[equation* omitted — 195 chars of source]

by the Cauchy-Schwarz inequality, so (ref) holds if $\dot{\mu}$ is locally square integrable. On the other hand, the function

equation*[equation* omitted — 79 chars of source]

satisfies (ref), but is not absolutely continuous (on intervals containing zero).

In other words, in the special case of a two-sided Brownian motion, Assumption (ref)(ref) can be replaced by a slightly weaker assumption, namely (ref), which can accommodate departures from absolute continuity. What is less clear, but arguably more interesting, is whether Assumption (ref)(ref) imposes unduly restrictive constraints on $\gamma$ also when $\mu$ is of the form (ref) near zero. In the remainder of this section, we attempt to shed light on that question.

When $\mu$ is of the form (ref), the condition $\gamma>1/2$ serves dual purposes when verifying the fact that $F_{\hat{\mathbf{s}}}$ is continuous. On the one hand, because (ref) holds with $H=1/2$, the condition $\gamma>1/2$ ensures that the tail behavior of $\mu$ is such that (ref) holds. In addition, $\gamma>1/2$ is necessary to ensure that $\mu$ is sufficiently well behaved near zero that the weak derivative of $\mu$ is locally square integrable. To shed more light on Assumption (ref)(ref), we disentangle the dual implications of $\gamma>1/2$ in (ref) by considering the mean function

equation[equation omitted — 128 chars of source]

which automatically satisfies (ref), but satisfies Assumption (ref)(ref) only when $\gamma>1/2$. The following example shows that every $\gamma<1/2$ admits a $c=c(\gamma,\sigma^2)>0$ such that if $\mu$ is given by (ref), then $F_{\hat{\mathbf{s}}}$ is discontinuous, suggesting in turn that for this canonical covariance kernel at least, Assumption (ref)(ref) is close to minimal in Theorem (ref).

exampFix $\gamma\in(0,1/2)$ and note that with probability one the sample paths of $\mathcal{G}^{\mu}$ are $\gamma$-H\"older continuous on $[0,1]$ Kallenberg_2021_Book. As a consequence, there exists a constant $c=c(\gamma,\sigma^2)$ such that \begin{equation} \mathbb{P}\left[\sup_{s\in(0,1]} s^{-\gamma}\mathcal{G}^{\mu}(s)\geq c\right]<1/4. \end{equation} Fixing any such $c$, let $\mathcal{G}=\mathcal{G}^{\mu}+\mu$, where $\mu$ is defined as in (ref). By (ref), we have \begin{equation*} \mathbb{P}[\hat{\mathbf{s}}\in(0,1]] \leq \mathbb{P}\left[\sup_{s\in(0,1]} \mathcal{G}(s)\geq 0\right] = \mathbb{P}\left[\sup_{s\in(0,1]} s^{-\gamma}\mathcal{G}^{\mu}(s)\geq c\right]< 1/4, \end{equation*} where the first inequality uses $\mathcal{G}(0)=0$. Similarly, \begin{align*} \mathbb{P}[\hat{\mathbf{s}}\in[1,\infty)] &\leq \mathbb{P}\left[\sup_{s\in[1,\infty)} \mathcal{G}(s)\geq 0\right] = \mathbb{P}\left[\sup_{s\in[1,\infty)} s^{-1}\mathcal{G}^{\mu}(s) \geq c\right] \\ &\leq \mathbb{P}\left[\sup_{s\in[1,\infty)} s^{\gamma-1}\mathcal{G}^{\mu}(s)\geq c\right] = \mathbb{P}\left[\sup_{s\in(0,1]} s^{1-\gamma}\mathcal{G}^{\mu}(1/s)\geq c\right] < 1/4, \end{align*} where the first inequality uses $\mathcal{G}(0)=0$, the second inequality uses $\gamma\geq0$, and the third inequality uses (ref) and the time inversion property of Brownian motion Kallenberg_2021_Book. Therefore, $\mathbb{P}[\hat{\mathbf{s}}>0] \leq \mathbb{P}[\hat{\mathbf{s}}\in(0,1]] + \mathbb{P}[\hat{\mathbf{s}}\in[1,\infty)] < 1/2$. Likewise, $\mathbb{P}\left[\hat{\mathbf{s}}<0\right] < 1/2$, so $\mathbb{P}[\hat{\mathbf{s}}=0] = 1 - \mathbb{P}[\hat{\mathbf{s}}<0] - \mathbb{P}[\hat{\mathbf{s}}>0] > 0$, implying in particular that $F_{\hat{\mathbf{s}}}$ is discontinuous at zero.