Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Bridging Root-$n$ and Non-standard Asymptotics: Adaptive Inference in M-Estimation
abstractThis manuscript studies a general approach to construct confidence sets for the solution of population-level optimization, commonly referred to as M-estimation. Statistical inference for M-estimation poses significant challenges due to the non-standard limiting behaviors of the corresponding estimator, which arise in settings with increasing dimension of parameters, non-smooth objectives, or constraints. We propose a simple and unified method that guarantees validity in both regular and irregular cases. Moreover, we provide a comprehensive width analysis of the proposed confidence set, showing that the convergence rate of the diameter is adaptive to the unknown degree of instance-specific regularity. We apply the proposed method to several high-dimensional and irregular statistical problems.
Introduction
The present study examines the inference for the parameter defined as the solution to an optimization problem, commonly referred to as M-estimation, which arises in broad statistical applications. Let $\mathcal{P}$ be a set of probability measures on a measurable space $(\Omega, \mathcal{S})$ with a $\sigma$-algebra $\mathcal{S}$. Let $Z_1, \ldots, Z_n \in \mathcal{Z}$ be a sequence of identically distributed random variables, following an unknown data-generating distribution $P \in\mathcal{P}$. We emphasize that independence is not assumed unless explicitly stated otherwise. Given a metric space $(\Theta, \|\cdot\|)$ and a “criterion” function $\operatorname{\mathbb{M}}: \Theta \times \mathcal{P} \mapsto \mathbb{R}$, the goal of an M-estimation problem is to identify an element $\theta(P) \in \Theta$, which minimizes (or maximizes) the mapping $\theta \mapsto \operatorname{\mathbb{M}}(\theta, P)$. Equivalently, the aim is to estimate
equation[equation omitted — 143 chars of source]
The uniqueness of the solution has not yet been assumed, and $\theta(P)$ in (ref) denotes the set of minimizers. The primary objective of this manuscript is the construction of an honest confidence set Li1989, Ptscher2002Lower for the $P$-dependent minimizer such that
equation[equation omitted — 204 chars of source]
where $\mathbb{P}_P(\cdot)$ denotes the probability of an event under distribution $P$. A finite-sample guarantee such as (ref) is impossible without very strong assumptions on the class of distributions $\mathcal{P}$ bahadur1956nonexistence. One may instead follow the convention in large sample theory and aim for the alternative asymptotic validity guarantee:
equation[equation omitted — 203 chars of source]
In what follows, to reduce notational burden, we use the following convention:
\[
\mathbb{P}_P(\theta(P)\in\widehat{\mathrm{CI}}_{n,\alpha}) ~:=~ \inf_{\theta^*\in\theta(P)}\,\mathbb{P}_P(\theta^*\in\widehat{\mathrm{CI}}_{n,\alpha}).
\]
The method studied in this article satisfies the asymptotic validity guarantee (ref) and additionally retains validity regardless of the dimension/complexity of the functional $\theta(P)$, including the cases where the dimension is comparable to the sample size. This property is termed dimension-agnostic by kim2020dimension.
Commonly used confidence set procedures include (1) the Wald methods based on the limiting distribution of a studied estimator, and (2) the resampling approaches. Both methods require an estimator $\widehat{\theta}_n$ such that for a suitable rate of convergence $r_n$, diverging to $\infty$, $r_n(\widehat{\theta}_n - \theta(P))$ converges in distribution. The Wald methods assume a parametric structure on the limiting distribution and use the quantiles of the estimated parametric limiting distribution to construct the confidence sets. The second (resampling) approach non-parametrically estimates the limiting distribution by resampling the available data. From the extensive study of both approaches in the literature, we know of numerous settings in which the corresponding confidence sets might not satisfy the guarantee of honest inference (ref). In particular, if the weak convergence of the normalized estimator is not “continuous” in $P$, the honest validity guarantee may fail for both Wald and resampling techniques Andrews2000, andrews2010asymptotic, cattaneo2020bootstrap, cattaneo2023bootstrap. An example of such “continuity” condition is the regularity of an estimator van2000asymptotic. We refer to the settings where such “continuity” condition fails (or equivalently, the estimator is ill-behaved) as irregular problems. It may be helpful to clarify that we consider cases where the functional $\theta(P)$ or estimator $\widehat{\theta}_n$ considered is ill-behaved to be irregular.
Inference for M-estimation problems that induce non-standard or irregular asymptotics is a particularly active area of research in econometrics geyer1994asymptotics, Ketz2018,Horowitz2019, Hsieh2022, Li2024. This list is far from exhaustive. To the best of our knowledge, existing methods typically are tailored to specific classes of regularities. For instance, Li2024 proposed a general inferential framework for stochastic optimization, allowing for potentially nonsmooth and non-convex objective functions; however, their approach requires knowledge of the estimator's convergence rate, which depends on unknown regularity, and the quantile of a complex random object must be estimated. vogel2008universal proposes a method closely related to the one in this manuscript; however, their corresponding confidence set often requires stronger regularity assumptions for validity, particularly, with respect to the complexity of $\Theta$. dey2024anytime and park2023robust consider inferential procedures based on the sample-splitting; however, they both require stronger assumptions than those in this manuscript for validity.
This manuscript proposes a simpler approach to inference based on the defining property of the functional and sample-splitting. Sample-splitting has seen renewed interest in recent times for complicated inference problems in several works Robins2006, chakravarti2019gaussian, wasserman2020universal, park2023robust, kim2020dimension, dey2024anytime. By the defining property, we mean (ref), which yields the functional as a minimizer of some population quantity. The method studied at an intuitive level can be derived in two steps as follows: (1) Because $\theta(P)$ minimizes $\theta\mapsto \mathbb{M}(\theta, P)$, we know that $\theta(P)$ is contained in $\{\theta:\, \mathbb{M}(\theta, P) \le \mathbb{M}(\widehat{\theta}, P)\}$ for any $\widehat{\theta}$ in the parameter space; (2) if $\theta\mapsto\widehat{\mathbb{M}}_n(\theta)$ is an estimator of $\mathbb{M}(\theta, P)$, then it is reasonable to expect that $\theta(P)$ would also be contained in $\{\theta:\, \widehat{\mathbb{M}}_n(\theta) \le \widehat{\mathbb{M}}_n(\widehat{\theta}_n) + \gamma_n\}$ for some “appropriate” $\gamma_n$. See Section (ref) for a precise description of the method.
This idea is not new and can be found, for example, in Robins2006 and vogel2008universal; see Section (ref) for some historical developments of this idea. Our contribution to this idea is two-fold. First, we provide a systematic way to obtain a dimension/complexity-agnostic validity guarantee; this is made possible by constructing $\widehat{\mathbb{M}}_n(\cdot)$ and $\widehat{\theta}_n$ on two independent datasets. Second, we analyze the width/diameter of the proposed confidence set under mild conditions and use the general result to prove the rate adaptivity in some examples. The conditions we adopt for width analysis are comparable to those employed in studying the convergence rate of the M-estimator. We believe these results represent a significant advancement in the field of honest and adaptive inference.
The proposed approach offers great flexibility and robustness compared to traditional methods of inference. First, the validity of the method is agnostic to the choice of the initial estimator $\widehat\theta_n$, and more importantly, it does not require any guarantees on its convergence rate or the existence of a limiting distribution. Second, the proposed method does not require knowledge of the convergence rate of the estimator. This is a significant improvement over traditional methods because in irregular problems the rate of convergence is impossible to estimate uniformly consistently. Third, it can accommodate constrained parameter spaces or (data-independent) regularization penalties with no additional modifications; this is because the proposed method is based on thresholding the (real-valued) objective function. Note that the limiting distributions of constrained M-estimators can be significantly complicated wang1996asymptotics.
This flexibility and generality come with a price in two aspects. Although we provide conditions under which the proposed confidence set shrinks to a singleton at the optimal rate adaptively, the proposed confidence set can be larger than the traditional ones due to sample splitting. Moreover, the shape of the confidence set is controlled by the shape of the objective function. Unlike traditional methods that control the shape of the confidence set by considering an appropriate statistic, the proposed confidence set can be non-convex or even disconnected depending on the (estimated) objective function. It might be worth pointing out that, in regular cases, the proposed confidence set will approximately be an ellipsoid in similarity to the likelihood ratio confidence set. Furthermore, we note that the universal inference procedure of wasserman2020universal also shares the same drawbacks.
In summary, the confidence set proposed in this manuscript remains valid even in high-dimensional or irregular problems where the standard approach exhibits non-standard asymptotics. The following statistical problems are a few examples in which inference remains difficult to date and the proposed confidence set provides a simple solution.
enumerate• High-dimensional Linear Regression: Inference for ordinary least squares (OLS) remains challenging when the dimension $d$ increases with the sample size $n$. In particular, mammen1993bootstrap and cattaneo2019two establish that the standard OLS estimator has a bias of order $d/n^{1/2}$, resulting in a shortage of valid inferential methods when $d \gg n^{1/2}$. Recently, cattaneo2019two and chang2023inference have proposed methods based on explicit bias correction, regaining validity in some regimes $d \gg n^{1/2}$; however, these methods still impose some constraints on the growth condition of $d$. Sections (ref) and (ref) show that the proposed confidence set is asymptotically valid for any growth rate of the dimension with a width tending to zero as long as $d/n\to 0$.
• Cube-root Estimators: Certain families of M-estimation share a common structure known as cube-root asymptotics kim1990cube. Notable examples include Manski’s maximum score estimator manski1975maximum, manski1985semiparametric,horowitz1992smoothed,delgado2001subsampling, the Grenander estimator grenander1956theory, sen2010inconsistency, westling2020unified, cattaneo2023bootstrap, and classification in machine learning mohammadi2005asymptotics. In these problems, the M-estimator converges at $n^{-1/3}$ rate whose limit process involves unknown infinite-dimensional objects, making inference difficult. In particular, classical empirical bootstrap is known to be inconsistent for these problems sen2010inconsistency, patra2018consistent; cattaneo2020bootstrap, cattaneo2023bootstrap recently proposed a modified resampling procedure for cube root problems. Section (ref) provides new inferential results related to a prototypical example in this class.
• Non-smooth Objective: Many common criterion functions in (ref) can be written as $\operatorname{\mathbb{M}}(\theta, P) \equiv \operatorname{\mathbb{E}}_P[m_{\theta}(Z)]$ where $m_{\theta}(Z)$ is often referred to as a “loss" function. When $\theta \mapsto m_{\theta}$ is non-smooth, the limiting distribution of an estimator can be non-standard unless additional regularity conditions hold smirnov1952limit, knight1998limiting. One well-known example is quantile estimation whose limiting distribution depends on the unknown smoothness of the cumulative distribution function (CDF) associated with $P$. While distribution-free finite sample valid confidence intervals exist for the quantiles, we study the behavior of the proposed confidence set in this problem in Section (ref).
• Constrained Optimization: The parameter space $\Theta$ in (ref) can incorporate structural constraints, such as sparsity, monotonicity, convexity, and boundedness wang1996asymptotics, candes2007dantzig, li2015geometric, royset2020variational. Confidence sets under such constraints have been explored in the literature geyer1994asymptotics, particularly within the operations research literature vogel2008confidence, vogel2008universal, vogel2017confidence, vogel2019universal. The proposed confidence set remains valid under such structural constraints.
While we prove the confidence set shrinks at an adaptive rate, it may not be the best set in terms of constants. Given the level of generality achieved, we do not know if one procedure can be adaptive and also be sharp in constants. We leave this aspect for the future research.
The remainder of this manuscript is organized as follows. Section (ref) formally defines the proposed procedure within a general optimization framework. Section (ref) establishes the foundational theorems on the validity and width of the proposed confidence set. Section (ref) provides an analysis of the confidence set proposed in statistical applications whose inference has been considered challenging. Section (ref) provides numerical results. We end the manuscript with a few concluding remarks in Section (ref).
\paragraph{Notation.} We adopt the following convention. For two real numbers $a$ and $b$, we set $a \vee b = \max\{a, b\}$ and $a \wedge b = \min\{a, b\}$. For $x \in \mathbb{R}^d$, we write $\|x\|_2 = \sqrt{x^\top x}$. In particular, we define the unit sphere with respect to $\|\cdot\|_2$ such that $\mathbb{S}^{d-1} = \{u \in \mathbb{R}^d \, : \, \|u\|_2 =1\}$. Given a square matrix $A \in \mathbb{R}^{d\times d}$, its trace, the smallest and the largest eigenvalues are denoted by $\mathrm{tr}(A)$, $\lambda_{\min}(A)$ and $\lambda_{\max}(A)$ respectively. For a real-valued function $f : \Omega \mapsto \mathbb{R}$, its supremum norm is denoted by $\|f\|_\infty := \sup_{x \in\Omega}|f(x)|$. A standard indicator function is denoted by $\mathbf{1}\{\cdot\}$, i.e., $\mathbf{1}\{x\in A\} = 1$ if $x\in A$ and $0$ if $x\notin A$. For any deterministic sequences $\{x_n\}_{n \ge 1}$ and $\{r_n\}_{n \ge 1}$, we denote $x_n = O(r_n)$ if there exists a universal constant $C>0$ such that $|x_n| \le C|r_n|$ for all $n$ larger than some $N$. Similarly, we denote $x_n = O_P(r_n)$ if, for any $\varepsilon>0$, there exists a constant $C_\varepsilon>0$ such that $\mathbb{P}(|x_n| \le C_\varepsilon|r_n|) \le \varepsilon$ for all $n$ larger than some $N_\varepsilon$.
We denote $x_n = o(r_n)$ if $x_n/r_n \to 0$ and $x_n = o_p(r_n)$ if $x_n/r_n \overset{p}{\to} 0$ where $\overset{p}{\to}$ denotes convergence in probability.
Construction of the Confidence Set
Given an identically, but not necessarily independently, distributed observation $Z_1, \ldots, Z_N$, we construct two sets of observations $D_1:= \{Z_i: i \in \mathcal{I}_1\}$ and $D_2:= \{Z_i : i \in \mathcal{I}_2\}$, where $\mathcal{I}_1$ and $\mathcal{I}_2$ are the disjoint partitions of $\{1, \ldots, N\}$. First, we construct any estimator of $\theta(P)$ using $D_1$, defining $\widehat \theta_1 := \widehat \theta_1(D_1) \in \Theta$. We use the first data only to obtain an initial estimator and remain agnostic to the choice of $\widehat \theta_1$. Given this estimator $\widehat \theta_1$, we construct a confidence set using the second data $D_2$. Throughout the manuscript, we denote by $n$ the cardinality of $D_2$ as we primarily use $D_2$ for the inferential task.
As stated in Section (ref), the proposed confidence set is based on the defining property of $\theta(P)$. An ideal (unactionable) confidence set is
\[
\widetilde{\mathrm{CI}} := \left\{\theta\in\Theta:\, \mathbb{M}(\theta, P) - \mathbb{M}(\widehat{\theta}_1, P) \le 0\right\}.
\]
From this, it might be tempting to consider the set
equation[equation omitted — 212 chars of source]
where $\theta\mapsto\widehat{\mathbb{M}}_n(\theta)$ is an estimator of $\mathbb{M}(\theta, P)$ based on $D_2$. If $\widehat{\theta}_1$ is consistent for $\theta(P)$, then the confidence set in (ref) may not have valid coverage, and if $\widehat{\theta}_1$ is not consistent for $\theta(P)$, then this confidence set may not shrink to a singleton as the sample size increases. Nevertheless, we can prove the following result on the coverage validity of the confidence set in (ref), and this can be useful to find a “small” set that contains $\theta(P)$. To state the result, for any $\theta, \theta'\in\Theta$, consider
equation[equation omitted — 342 chars of source]
The quantity $\mathbb{C}_P(\theta)$ represents the curvature of the M-estimation problem. It quantifies the hardness of “estimating” $\theta(P)$. Note that $\mathbb{V}_P(\theta, \theta')$ and $\mathbb{C}_P(\theta)$ are defined for non-stochastic $\theta, \theta'\in\Theta$ and if evaluated at a random point $\widehat{\theta}_1\in\Theta$, these should be considered as random variables.
theoremFor any initial estimator $\widehat{\theta}_1$ computed on $D_1$, and any estimator $\widehat{\mathbb{M}}_n(\cdot)$ of $\mathbb{M}(\cdot, P)$ computed on $D_2$, we have
\[
\mathbb{P}_P\left(\theta(P) \notin \widehat{\mathrm{CI}}_{n}^{\dagger}\right) ~\le~ \mathbb{E}_P\left[\frac{\mathbb{V}_P(\theta(P), \widehat{\theta}_1)}{\mathbb{V}_P(\theta(P), \widehat{\theta}_1) + \mathbb{C}_P^2(\widehat{\theta}_1)}\right] \le 1.
\]
In particular, if $\mathbb{V}_P(\theta(P), \widehat{\theta}_1)/\mathbb{C}_P^2(\widehat{\theta}_1) = o_p(1)$ uniformly over all $P\in\mathcal{P}$, then $\widehat{\mathrm{CI}}_n^{\dagger}$ is an asymptotically uniformly valid confidence interval of confidence 1.
The proof of Theorem (ref) is provided in Section (ref) and is based on Cantelli's inequality pinelis2010between. Observe that uniqueness of $\theta(P)$ is not required in Theorem (ref). Moreover, the miscoverage bound depends naturally on two aspects of the M-estimation problem: (1) the estimation error of $\widehat{\mathbb{M}}_n(\cdot)$ captured by the mean squared error $\mathbb{V}_P(\cdot, \cdot)$; and (2) the curvature of the problem $\mathbb{C}_P(\cdot)$. Note that, by the definition of $\theta(P)$, $\mathbb{C}_P(\theta) \ge 0$ for all $\theta\in\Theta$. The miscoverage bound is non-decreasing in the estimation error and non-increasing in the curvature of the problem. Because the definition of the confidence set does not depend on any target coverage, the confidence set provides an agnostic bound. To understand the behavior of the miscoverage bound, consider the case when $\mathbb{M}(\theta, P) = \mathbb{E}_P[m_{\theta}(Z)]$ for a loss function $m_{\theta}(Z)$ and $\widehat{\mathbb{M}}_n(\theta) = \sum_{i=1}^n m_{\theta}(Z_i)/n$, $\mathbb{V}_P(\theta, \theta') = \mbox{Var}(m_{\theta}(Z) - m_{\theta'}(Z))/n$. If $n\mathbb{C}_P^2(\widehat{\theta}_1)$ diverges to infinity in probability, while $\mbox{Var}(m_{\theta}(Z) - m_{\widehat{\theta}_1}(Z)|\widehat{\theta}_1)$ is bounded away from zero as $n\to\infty$, then the confidence set in (ref) has an asymptotic confidence of 1. For the special case of loss functions satisfying the so-called Bernstein condition bartlett2006empirical with parameters $(\beta, B)$ with $\beta\in(0, 1]$, i.e., $\mathbb{E}_P[(m_{\theta}(Z) - m_{\theta(P)}(Z))^2] \le B(\mathbb{C}_P(\theta))^{\beta}$, this divergence condition is satisfied if $(n/B)^{1/(2-\beta)}\mathbb{C}_P(\widehat{\theta}_1)$ diverges to infinity in probability as $n\to\infty.$ Note that this divergence condition can be trivially satisfied by taking an inconsistent estimator $\widehat{\theta}_1$ of $\theta(P)$, and the resulting confidence set would be a (non-shrinking) bounded set, in most cases.
For finer control on the coverage guarantee, we modify the confidence set in (ref) as follows.
Define random mappings $\theta \mapsto L_{n,\alpha_1}(\theta, \widehat \theta_1; D_2)$ and $\theta \mapsto U_{n,\alpha_2}(\theta, \widehat \theta_1; D_2)$ based on the dataset $D_2$, satisfying the following properties:
align[align omitted — 204 chars of source]
and
align[align omitted — 202 chars of source]
for any $n\ge 1$ and $\alpha_1, \alpha_2 \in (0,1)$. The final confidence set for $\theta(P)$ is defined as:
equation[equation omitted — 231 chars of source]
where $\alpha := (\alpha_1, \alpha_2)$. In particular, if $\alpha_2 = 0$, then $U_{n,\alpha_2}(\theta(P), \widehat{\theta}_1; D_2)$ can be taken to be zero because the definition of $\theta(P)$ implies $-\mathbb{C}_P(\widehat{\theta}_1) \le 0$ almost surely, and this yields
equation[equation omitted — 171 chars of source]
The following theorem establishes the validity guarantee for the proposed confidence set.
theoremFix $n\ge1$, and assume (ref) and (ref) to hold. Then,
\begin{align*}
\inf_{P\in\mathcal{P}}\, \mathbb{P}_P(\theta(P) \in \widehat{\mathrm{CI}}_{n,\alpha} \mid D_1) \ge 1-\alpha_1-\alpha_2.
\end{align*}
The proof of Theorem (ref) can be found in Section (ref). In situations where constructing a confidence set for $\theta(P)$ seems challenging, it may initially appear puzzling that the proposed approach involves building a (lower) confidence bound for $-\mathbb{C}_P(\widehat{\theta}_1) = \mathbb{M}(\theta(P), P) - \mathbb{M}(\widehat{\theta}_1, P)$. However, recognizing that the confidence set (ref) relies on the random mapping $\theta \mapsto L_{n,\alpha_1}(\theta, \widehat{\theta}_1; D_2)$, it is helpful to reinterpret condition (ref) as
equation[equation omitted — 285 chars of source]
In other words, we need to construct a lower confidence bound for $\mathbb{M}(\theta, P) - \mathbb{M}(\theta', P)$ for every (non-stochastic) pair $\theta, \theta'\in\Theta$. For instance, if $\mathbb{M}(\theta, P) = \mathbb{E}_P[m_{\theta}(Z)]$, then $\mathbb{M}(\theta, P) - \mathbb{M}(\theta', P) = \mathbb{E}_P[m_{\theta}(Z) - m_{\theta'}(Z)]$ for which lower confidence bound can be constructed by the central limit theorem.
Because the mapping $\theta \mapsto \operatorname{\mathbb{M}}(\theta, P) - \operatorname{\mathbb{M}}(\widehat \theta_1, P)$ is always real-valued regardless of the complexity of $\Theta$, it is conceivable that the lower confidence limit in (ref) can be constructed without any reference to the complexity of $\Theta$. This is in stark contrast to the “classical” inferential approach based on the weak convergence of $r_n(\widehat \theta_1-\theta(P))$ to some limit process, which often depends heavily on the complexity of $\Theta$.
Comparing the definitions of $\widehat{\mathrm{CI}}_{n}^{\dagger}$ in (ref) and $\widehat{\mathrm{CI}}_{n,\alpha}$ in (ref), one can interpret the naive confidence set in (ref) as using the estimator $\widehat{\mathbb{M}}_n(\theta) - \widehat{\mathbb{M}}_n(\widehat{\theta}_1)$ as $L_{n,\alpha}(\theta, \widehat{\theta}_1; D_2)$. This, in general, does not satisfy the guarantee (ref). In the following section, we consider two methods for constructing lower confidence bounds satisfying (ref): one based on concentration inequalities under the assumption of bounded loss function and the other based on central limit theorem. Both lower bounds are of the form
\[
L_{n,\alpha}(\theta, \widehat{\theta}_1; D_2) = \widehat{\mathbb{M}}_n(\theta) - \widehat{\mathbb{M}}_n(\widehat{\theta}_1) - t(\alpha, \theta, \theta'),
\]
for some non-negative function $t(\cdot, \cdot, \cdot)$. With such non-negativity, it is clear that $\widehat{\mathrm{CI}}_{n,\alpha} \supseteq \widehat{\mathrm{CI}}_{n}^{\dagger}$.
remarkThe upper and lower confidence bounds play asymmetric roles. While the lower confidence bound is crucial for the validity, the upper confidence bound serves to improve the statistical power. As noted above, the upper confidence bound (ref), in fact, holds trivially by setting $U_{n,\alpha_2}(\theta(P), \widehat \theta_1; D_2) = 0$ and $\alpha_2=0$. We emphasize that (ref) and (ref) are only required to hold at $\theta(P)$, and it is generally true that $L_{n,\alpha_1}(\theta, \widehat \theta_1; D_2) > U_{n,\alpha_2}(\theta, \widehat \theta_1; D_2)$ for some $\theta \in \Theta$. As a result, we have $ \{\theta \in \Theta : L_{n,\alpha_1}(\theta, \widehat \theta_1; D_2) \le U_{n,\alpha_2}(\theta, \widehat \theta_1; D_2)\} \subsetneq \Theta$, and thus
\begin{align}
\widehat{\mathrm{CI}}_{n,\alpha} &:= \{L_{n,\alpha_1}(\theta, \widehat \theta_1; D_2) \le 0\} \cap \{L_{n,\alpha_1}(\theta, \widehat \theta_1; D_2) \le U_{n,\alpha_2}(\theta, \widehat \theta_1; D_2)\} \nonumber\\
&\subsetneq \{L_{n,\alpha_1}(\theta, \widehat \theta_1; D_2) \le 0\},\nonumber
\end{align}
leading to a non-trivial shrinkage of the confidence set when $U_{n,\alpha_2}(\theta(P), \widehat \theta_1; D_2) < 0$.
remarkThe validity of the confidence set is established even when the optimizer $\theta(P)$ in (ref) is not uniquely identified. In such cases, the resulting confidence set will contain any point that minimizes (ref), and consequently will not converge to a singleton set. In order to establish the convergence rate of the diameter of the confidence set, we assume the uniqueness of the optimizer. Formal statements are provided as Theorems (ref) and (ref) below.
Construction of Lower Confidence Bounds
Section (ref) establishes that the inference on $\theta(P)$, as defined in (ref), can be reduced to the construction of the upper and lower confidence bounds, satisfying (ref) and (ref). Furthermore, Remark (ref) suggests that a valid confidence set can be obtained without estimating the upper confidence bound as the upper bound only plays a role in improving the statistical power. Therefore, the primary focus is on constructing the lower confidence bound, which is essential for ensuring validity. This section presents two general approaches for developing the lower confidence bounds.
Throughout, we introduce an additional structure to the M-estimation shared by many problems such that $\operatorname{\mathbb{M}}(\theta, P)$ corresponds to the expectation of some “loss" function. Formally, we define a measurable function $m_\theta : \mathcal{Z} \mapsto \mathbb{R}$ indexed by $\theta \in \Theta$. We consider the minimization of $m_\theta$ under the expectation with respect to $P$ such that
align*[align* omitted — 185 chars of source]
For example, taking $m_{\theta}(Y,X) := (Y-\theta^\top X)^2$ with $\Theta \equiv \mathbb{R}^d$ corresponds to linear regression while taking $m_{\theta}(Z) := -\log P(Z; \theta)$ for a parametrized family of likelihood functions $P(Z; \theta)$ yields maximum likelihood estimation. This definition arises in many popular situations; however, it does exclude certain classes of M-estimation problems. For example, there are problems when $\operatorname{\mathbb{M}}(\theta, P)$ is defined as U-statistics or higher-order U-statistics bose2018u, diciccio2022clt, U-quantile choudhury1988generalized, and other problems where $\operatorname{\mathbb{M}}(\theta, P)$ involves nuisance parameters, such as Cox proportional hazard models cox1972regression. Although the general results in Section (ref) still apply to these problems, detailed investigations of these applications are left for future research.
We introduce the definitions and notation that we frequently refer to. For any $n \in \mathbb{N}$ and the identically distributed observation $Z_1, \ldots, Z_n \in D_2$, the empirical measure is defined as $\mathbb{P}_n := n^{-1}\sum_{i=1}^n \delta_{Z_i}$ where $\delta_z$ is the Dirac measure at $z$. For any measure $P$ and $P$-integrable function, we set $Pf = \int f dP$. In particular, $\mathbb{P}_n f$ means $n^{-1} \sum_{i=1}^n f(Z_i)$. The empirical process is defined as the centered and normalized process, which is denoted by $ \mathbb{G}_n f := n^{1/2}(\mathbb{P}_n-P)f$. For notational convenience, we also denote by $\operatorname{\mathbb{E}}_P[\cdot]$ the expectation under $P$. Given an arbitrary initial estimator $\widehat\theta_1$, it follows that
align*[align* omitted — 320 chars of source]
Construction by concentration inequalities
As discussed in Section (ref), we can derive the lower confidence bound by concentration inequalities. In particular, replacing $\operatorname{\mathbb{M}}(\theta, P)$ with $P m_{\theta}$ and $\widehat{\mathbb{M}}_n(\theta)$ with $\mathbb{P}_n m_{\theta}$, it suffices to establish the following inequality:
align*[align* omitted — 155 chars of source]
There is a wide range of concentration inequalities for this purpose. We refer to Boucheron2013Concentration for a glossary of classical results and hao2019bootstrapping, ramdas2023randomized, waudby2024estimating, bates2021distribution for more recent developments. As an illustration, we may consider the one-sided empirical Bernstein inequality maurer2009empirical, established below:
example[Empirical Bernstein inequality]
Suppose $Z_1, \ldots, Z_n$ are independent and identically distributed (IID) according to $P$, and
\begin{align}
\|m_{\theta_1}-m_{\theta_2}\|_\infty \le B_0 \quad for all \quad \theta_1, \theta_2 \in \Theta.
\end{align}
Denoting the sample variance of $m_{\theta_1}-m_{\theta_2}$ by
\begin{align}
\widehat \sigma_{\theta_1, \theta_2}^2 := \frac{1}{n-1}\sum_{i=1}^n \left\{(m_{\theta_1}-m_{\theta_2})(Z_i) - \mathbb{P}_n(m_{\theta_1}-m_{\theta_2})\right\}^2,
\end{align}
the one-sided empirical Bernstein inequality implies that
\begin{align}
\mathbb{P}_P\left((\mathbb{P}_n-P)(m_{\theta(P)}-m_{\widehat \theta_1}) > \sqrt{\frac{2\widehat \sigma^2_{\theta, \widehat\theta_1} \log(2/\alpha_1)}{n}} + \frac{7B_0\log(2/\alpha_1)}{3(n-1)}\mid D_1\right) \le \alpha_1.
\end{align}
Hence, the concentration inequality-based confidence set becomes
\begin{align}
&\widehat{\mathrm{CI}}^{\mathrm{E.B.}}_{n,\alpha} \nonumber\\
&\quad := \left\{\theta \in \Theta \, : \, \mathbb{P}_n(m_{\theta}-m_{\widehat \theta_1})\le \sqrt{\frac{2\widehat \sigma^2_{\theta, \widehat\theta_1} \log(2/\alpha_1)}{n}} + \frac{7B_0\log(2/\alpha_1)}{3(n-1)}+\big(0 \wedge U_{n,\alpha_2}(\theta, \widehat \theta_1; D_2)\big)\right\}.
\end{align}
The validity of the confidence set can be established without any additional assumptions, as shown in the following theorem.
theoremAssume that $Z_1, \ldots, Z_n$ are IID according to $P \in \mathcal{P}$. Suppose (ref) and (ref) hold.
Then for any fixed $n\ge 1$, and $\alpha_1, \alpha_2\in(0, 1)$,
\begin{align*}
\inf_{P\in\mathcal{P}}\, \mathbb{P}_P(\theta(P) \in \widehat{\mathrm{CI}}^{\mathrm{E.B.}}_{n,\alpha}) \ge 1-\alpha_1-\alpha_2.
\end{align*}
This theorem is a simple corollary of Theorem (ref). The conditional probability $\mathbb{P}_P(\cdot|D_1)$ is replaced by the marginal probability $\mathbb{P}_P(\cdot)$ since the lower bound no longer depends on $D_1$.
Next, we establish the rate at which the confidence set shrinks. As discussed in Remark (ref), the convergence rate of the set requires the uniqueness of the optimizer $\theta(P)$. Below, we assume that $\theta(P)$ is a unique point in $\Theta$ such that the following assumptions hold:
enumerate[label=(A\arabic*),leftmargin=2cm]
• There exist constants $c_0, \beta \ge 0$ such that
\begin{align*}
\mathbb{C}_P(\theta) = P (m_\theta - m_{\theta(P)}) \ge c_0\|\theta-\theta(P)\|^{1+\beta}\quad for all \quad \theta \in \Theta.
\end{align*}
• There exists a function $\phi_n : \mathbb{R}_+ \mapsto \mathbb{R}$ such that
\begin{align}
\left\{\operatorname{\mathbb{E}}_P \left[\sup_{\|\theta-\theta(P)\| < \delta}|\mathbb{G}_n (m_\theta - m_{\theta(P)})| \right] \vee \sup_{\|\theta-\theta(P)\| \le \delta }\left(P(m_{\theta}-m_{\theta(P)})^2 \right)^{1/2}\right\}\le \phi_n(\delta)
\end{align}
for every $n\ge 1$. Furthermore, $\phi_n(x)/x^q$ is non-increasing for some $q < 1+\beta$.
• For every $n \ge 1$, there exists a random variable $s_n$ such that the initial estimator satisfies
\begin{align}
n^{-1/2}\sqrt{P (m_{\widehat\theta_1}-m_{\theta(P)})^2} + \mathbb{C}_P(\widehat{\theta}_1) \le s_n.
\end{align}
The parameter $\beta$ in (ref) links the optimization problem at hand to the curvature within the parameter space. This condition also implies that $\theta(P)$ is a strong global minimizer of $\theta\mapsto\mathbb{M}(\theta, P)$ drusvyatskiy2013tilt. The modulus $\phi_n$ in (ref) quantifies the complexity of the parameter space $\Theta$ and serves as an upper bound on the variance of the loss function near $\theta(P)$. In particular, the first part of inequality (ref) is known in the literature as the maximal inequality with its upper bound $\phi_n(\cdot)$ reflecting the complexity of $\Theta$. Finally, (ref) pertains to the rate of convergence of the initial estimator. The convergence rate of the proposed confidence set depends on all these terms; such geometric structures are commonly used to establish the convergence rates of M-estimators, as discussed in Theorem 3.2.5 of van1996weak. These conditions are also used in kim1990cube in the estimation context to explain the differences between M-estimators in the regular and irregular cases.
remarkThe “curvature assumption" (ref) is stated for all $\theta \in \Theta$. However, in most cases, this holds only locally in some neighborhood of the optimum $\theta(P)$---See Section (ref) for a concrete example. Relaxing (ref) to its local analog is not trivial. For example, when (ref) holds locally and $\Theta$ is unbounded, the conditions above may not be sufficient to claim that the proposed confidence set is bounded. To this end, we envision the use of the naive confidence set $\widehat{\mathrm{CI}}_n^{\dagger}$ in (ref) with a potentially inconsistent $\widehat{\theta}_1$ to first obtain a bounded confidence set and then consider the intersection of $\widehat{\mathrm{CI}}_{n,\alpha}$ with $\widehat{\mathrm{CI}}_n^{\dagger}$ to obtain a provably bounded confidence set. This intersection allows us to weaken assumptions (ref) and (ref) by restricting to a bounded subset of $\Theta$, even if $\Theta$ is unbounded to start with.
We now establish a bound on the diameter of the confidence set. For any set $A$ equipped with a metric $\|\cdot\|$, the diameter of $A$ is denoted by $\mathrm{Diam}_{\|\cdot\|}(A) := \sup\{\|a-b\|\, :\, a, b\in A\}$. The following theorem provides the high probability bound on the diameter of the proposed confidence set constructed using a concentration inequality-based lower confidence bound.
theoremAssume that $Z_1, \ldots, Z_n$ are IID according to $P \in \mathcal{P}$. Assume $\theta(P)$ is a unique solution of (ref) that satisfies (ref)--(ref). Given $c_0$ and $\beta$ in (ref) and the modulus $\phi_n$ in (ref), we define $r_n$ as any value that satisfies
\begin{align}
r_n^{-2} \phi_n(c_0^{-1/(1+\beta)}r_n^{2/(1+\beta)}) \le n^{1/2}.
\end{align}
Then, for any $n\ge 1$ and $\varepsilon > 0$, there exists a constant $\mathfrak{C}$, depending on $\alpha, \beta, B_0, \varepsilon$ such that
\begin{align*}
\mathbb{P}_P\left(\mathrm{Diam}_{\|\cdot\|}\big(\widehat{\mathrm{CI}}^{\mathrm{E.B.}}_{n,\alpha}\big)\le \mathfrak{C}c_0^{-1/(1+\beta)}(r_n^{2/(1+\beta)} + (B_0/n)^{1/(1+\beta)} + s_n^{1/(1+\beta)})\right) \ge 1-\varepsilon.
\end{align*}
The proof is provided in Section (ref) of the supplement. In summary, Theorem (ref) states that the diameter of the set $\widehat{\mathrm{CI}}^{\mathrm{E.B.}}_{n,\alpha}$ depends on four key factors: (1) the curvature of the optimization problem, (2) the complexity of $\Theta$, (3) the variance of the loss function near $\theta(P)$, and (4) the quality of $\widehat \theta_1$.
Construction by the central limit theorem
This section discusses an alternative approach to construct the lower confidence bound, given by (ref), based on the central limit theorem (CLT). Let $\widehat \sigma^{2}_{\theta(P), \widehat\theta_1}$ be the sample variance of $(m_{\theta(P)}-m_{\widehat \theta_1})(Z_i)$, given by (ref), and $\Phi(t)$ be the cumulative distribution function of the standard normal random variable. We define the Kolmogorov-Smirnov distance as
align[align omitted — 241 chars of source]
We will shortly discuss the upper bound for $\Delta_{n,P}$ based on the Berry–Esseen bound for studentized statistics bentkus1996berry, Bentkus1996. The lower confidence bound is defined as
align*[align* omitted — 189 chars of source]
where $z_{\alpha}$ is the $1-\alpha$-th quantile of a standard normal distribution. The corresponding confidence set is given by
equation[equation omitted — 317 chars of source]
The validity of this set follows immediately in view of (ref) and Theorem (ref).
theoremAssume that $Z_1, \ldots, Z_n$ are IID according to $P \in \mathcal{P}$. Suppose (ref) holds. Then for any fixed $n\ge 1$,
\begin{align*}
\inf_{P\in\mathcal{P}}\, \mathbb{P}_P(\theta(P) \in \widehat{\mathrm{CI}}^{\mathrm{CLT}}_{n,\alpha}|D_1) \ge 1-\alpha_1-\alpha_2- \sup_{P\in\mathcal{P}}\,\Delta_{n,P}.
\end{align*}
The proof is provided in Section (ref) of the supplement. Comparing two expressions (ref) and (ref), we can observe that the CLT-based method always yields a smaller confidence set, since the leading constant for the $n^{-1/2}$ term in (ref) is $\sqrt{2\log(2/\alpha)} > z_{\alpha}$ for any $\alpha \in (0,1)$. This shows a strong practical advantage of the CLT-based confidence set. However, the CLT-based method compromises the validity for a finite sample size because $\sup_{P\in\mathcal{P}}\Delta_{n,P}$ might at best be close to zero as $n\to\infty$. In the following, we discuss sufficient conditions, based on katz1963note, bentkus1996berry, and Bentkus1996, under which $\sup_{P\in\mathcal{P}}\Delta_{n,P}$ converges to zero.
lemmaAssume $Z_1, \ldots, Z_n$ is IID, generated from $P$. Define a centered random variable
\begin{align}
W_i := m_{\theta(P)}(Z_i)-m_{\widehat \theta_1}(Z_i) - P(m_{\theta(P)} - m_{\widehat \theta_1}) \quad for\quad 1 \le i \le n
\end{align}
and let $B^2 := \operatorname{\mathbb{E}}_P[W_1^2|D_1]$. Then
\begin{align}
\Delta_{n,P} &\le \min\left\{1, C \operatorname{\mathbb{E}}_P\left[\frac{W_1^2}{B^2} \min\left\{1, \frac{|W_1|}{n^{1/2}B}\right\}\bigg| D_1 \right]\right\}
\end{align}
for some universal constant $C > 0$.
This result is a direct consequence of Corollary 1.1 of bentkus1996berry. To ensure the validity of the proposal, we require that the right-hand term of (ref) converge to zero.
theoremAssume that $Z_1, \ldots, Z_n$ are IID according to $P \in \mathcal{P}$. Suppose (ref) holds, and
\begin{align}
\sup_{P\in\mathcal{P}}\,\operatorname{\mathbb{E}}_P\left[\frac{W_1^2}{B^2} \min\left\{1, \frac{|W_1|}{n^{1/2}B}\right\}\, \bigg| D_1\right] = o_p(1)\quad as $n \to \infty$.
\end{align}
Then
\begin{align*}
\inf_{P \in \mathcal{P}}\, \mathbb{P}_P(\theta(P) \in \widehat{\mathrm{CI}}^{\mathrm{CLT}}_{n,\alpha}) \ge 1-\alpha_1-\alpha_2 - o(1)\quadas\quad n\to\infty.
\end{align*}
This theorem is an immediate consequence of combining Theorem (ref) and Lemma (ref). In particular, the conditional probability $\mathbb{P}_P(\cdot | D_1)$ can be replaced by the marginal probability $\mathbb{P}_P(\cdot)$ because $\Delta_{n,P}$ can be bounded by the minimum of one and the left-hand side of (ref). Section (ref) provides a further discussion on condition (ref).
remarkThe requirement that $\sup_{P\in\mathcal{P}}\Delta_{n,P}$ converges to zero can be established under conditions considerably weaker than those imposed by (ref). This is a well-known benefit of studentization. In particular, we can relax the assumption that the variance of $W$ is finite, provided that $W$ belongs to the domain of attraction of the normal law. For equivalent conditions without finite variance, see Corollary 1.5 of Bentkus1996.
We now establish bounds on the diameter of the CLT-based confidence set. First, we introduce an additional assumption to allow for unbounded loss functions.
enumerate[label=(A\arabic*),leftmargin=2cm]\setcounter{enumi}{3}
• There exists a function $\omega_n : \mathbb{R}_+ \mapsto \mathbb{R}$ such that
\begin{align}
\operatorname{\mathbb{E}}_P \left[\sup_{\|\theta-\theta(P)\| < \delta}|\mathbb{G}_n (m_\theta - m_{\theta(P)})^2| \right]\le \omega^2_n(\delta)
\end{align}
for every $n\ge 1$. Furthermore, $\omega_n(x)/x^q$ is non-increasing for some $q < 1+\beta$.
Assumption (ref) resembles (ref), but instead, quantifies the growth rate of the expectation of the localized squared empirical processes. As $m_{\theta}$ is unbounded, the growth rates of the moduli $\phi_n$ and $\omega_n$ may depend on the concentration properties of $(m_\theta - m_{\theta(P)})(Z_i)$, characterized through, for instance, sub-Weibull tails or the number of the available moments. Such arguments often require specific case-by-case analysis. Following Section 3.2 of van1996weak, we present the general result in terms of the growth rate of $\omega_n$, allowing for the theorem to apply without specific conditions on $m_\theta - m_{\theta(P)}$. The theorem is provided as follows:
theoremAssume that $Z_1, \ldots, Z_n$ are IID according to $P\in \mathcal{P}$. Assume $\theta(P)$ is a unique solution of (ref) that satisfies (ref)--(ref). Given $c_0$ and $\beta$ in (ref), the moduli $\phi_n$ and $\omega_n$ in (ref) and (ref), we define $r_n$ and $u_n$ as any values that satisfy
\begin{align}
r_n^{-2} \phi_n(c_0^{-1/(1+\beta)}r_n^{2/(1+\beta)}) \le n^{1/2}\quad and \quad u_n^{-2} \omega_n(c_0^{-1/(1+\beta)}u_n^{2/(1+\beta)}) \le n^{3/4}.
\end{align}
Then, for any $n\ge 1$ and $\varepsilon > 0$, there exists a constant $\mathfrak{C}$, depending on $\alpha, \beta, \varepsilon$ such that
\begin{align*}
\mathbb{P}_P\left(\mathrm{Diam}_{\|\cdot\|}\big(\widehat{\mathrm{CI}}^{\mathrm{CLT}}_{n,\alpha}\big)\le \mathfrak{C}c_0^{-1/(1+\beta)}(r_n^{2/(1+\beta)} + u_n^{2/(1+\beta)} + s_n^{1/(1+\beta)})\right) \ge 1-\varepsilon.
\end{align*}
The proof is provided in Section (ref) of the supplement.
remarkThe CLT-based confidence set remains useful even when $m_{\theta}$ is uniformly bounded. In such cases, Theorem (ref) yields the same rate as Theorem (ref) since we can simply take $\omega_n = \phi_n$ using the contraction inequality (Theorem 4.12 of ledoux2013probability). Although the rates are identical, the validity must be verified since there exist uniformly bounded random variables for which the CLT does not hold. A classical example is when $W_1, \ldots, W_n$ are IID Bernoulli random variables with success probability $p$ such that $np\to \lambda$. In this case, the condition (ref) fails.
remarkWhile it is common to impose light-tail assumptions, such as the sub-Gaussian tail for $m_{\theta}-m_{\theta(P)}$, to obtain $\omega_n$ for unbounded processes, Lemma (ref) in Section (ref) of the supplement offers general construction of $\omega_n$ under much weaker assumptions. This is achieved through a truncation argument, which is a proof device frequently employed in the literature. See similar results, for instance, Proposition 3.1 of gine2000exponential and Proposition B.1 of kuchibhotla2022least.
On the validity of the CLT-based method
Section (ref) outlines the construction of the confidence set based on the CLT whose (asymptotic) validity is guaranteed so long as the Gaussian approximation tends to zero as $n \to \infty$.
Verification of (ref) can be difficult when $\widehat{\theta}_1$ is consistent for $\theta(P)$ because the variance $B^2$ in Lemma (ref) converges to zero under consistency of $\widehat{\theta}_1$.
This section provides sufficient conditions for the validity to hold. The results of this section are frequently used in the statistical applications presented in
Section (ref).
For any real-valued random variable $X$, we write $\|X\|_p = (\mathbb{E}[|X|^p])^{1/p}$ for any $p > 0$.
definition[Uniform Lindeberg condition]
A distribution $Q$ supported on (a subset of) $\mathbb{R}^d, d\ge1$ is said to satisfy the {\em uniform Lindeberg condition} (ULC) if a random variable $H\sim Q$ satisfies
\begin{align}
\lim_{\kappa \to \infty}\, \sup_{t \in \mathbb{R}^d}\, \operatorname{\mathbb{E}}\left[\frac{\langle t, H \rangle^2}{\|\langle t, H \rangle\|_{2}^2} \min\left\{1, \frac{|\langle t, H \rangle|}{\kappa \|\langle t, H \rangle\|_{2}}\right\}\right] = 0.
\end{align}
In particular, a class of distributions $\mathcal{Q}$ is said to satisfy the ULC if
\begin{align}
\lim_{\kappa \to \infty}\,\sup_{Q\in\mathcal{Q}}\, \sup_{t \in \mathbb{R}^d}\, \operatorname{\mathbb{E}}_Q\left[\frac{\langle t, H \rangle^2}{\|\langle t, H \rangle\|_{2}^2} \min\left\{1, \frac{|\langle t,H\rangle|}{\kappa \|\langle t, H \rangle\|_{2}}\right\}\right] = 0.
\end{align}
Importantly, the uniform Lindeberg condition does not require strong moment assumptions on $Q$ such as the finite third moments.
The proposition below provides the first-order approximation result, which becomes useful for many statistical applications.
propositionSuppose that there exists a $\delta_0 > 0$ and a $P$-dependent\footnote{We mean that the distribution of $H_i$ depends on $P$.} mean-zero random variable/vector $H_i$ such that
\begin{align}
\frac{\operatorname{\mathbb{E}}_P[W_1- \langle \theta - \theta(P), H_1 \rangle ]^2}{\operatorname{\mathbb{E}}_P[\langle \theta - \theta(P), H_1 \rangle^2]} \le \varphi(\|\theta - \theta(P)\|) \quad for all \quad \|\theta - \theta(P)\| < \delta_0
\end{align}
where $\varphi:\mathbb{R}_+\to\mathbb{R}_+$ is continuous and $\varphi(0) = 0$. Then, for any $P\in\mathcal{P}$,
\begin{align*}
\operatorname{\mathbb{E}}_P\left[\frac{W_1^2}{B^2} \min\left\{1, \frac{|W_1|}{n^{1/2}B}\right\}\right] &\lesssim \inf_{\delta < \delta_0}\left\{\varphi(\delta) + \varphi^{1/2}(\delta) + \mathbb{P}_P(\|\widehat\theta_1 - \theta(P)\| > \delta)\right\}\\
&\quad + \sup_{t \in \mathbb{R}^d}\,\operatorname{\mathbb{E}}_P\left[\frac{\langle t, H_1 \rangle^2}{\|\langle t, H_1 \rangle\|_{2}^2} \min\left\{1, \frac{|\langle t, H_1 \rangle|}{n^{1/2}\|\langle t, H_1 \rangle\|_{2}}\right\}\right]
\end{align*}
where $B^2 = \operatorname{\mathbb{E}}_P[W^2_1|D_1]$. In particular, if $\widehat{\theta}_1$ is uniformly (in $P\in\mathcal{P}$) consistent for $\theta(P)$ and $H$ satisfies ULC uniformly over $P\in\mathcal{P}$, then the left hand side converges to zero as $n\to\infty$.
The proof is provided in Section (ref) of the supplement. Intuitively, the random variable $H_1$ serves as the gradient of $W_1$ with respect to $\theta$ (evaluated at $\theta(P)$) so that $W_1 \approx \langle \theta - \theta(P), H_1 \rangle$. The approximation defined in (ref) is a quantitative version of the quadratic mean differentiability van1996weak, which is weaker than the pointwise differentiability. The role of the uniform Lindeberg condition is clearer in this context: when the first-order approximation in the sense of (ref) is available, it suffices to verify the uniform Lindeberg condition for the random sequence $\{H_i\}_{i=1}^n$ instead of directly inspecting $\{W_i\}_{i=1}^n$. The upper bound of this proposition no longer involves conditioning on $D_1$, but instead, requires the consistency of the initial estimator $\widehat \theta_1$. Although this aspect may be stringent compared to the confidence set based on concentration inequalities, which remained valid even when $\widehat\theta_1$ is not consistent, such consistency is not a necessary condition for the validity of the CLT-based confidence sets---see Remark (ref) in this regard. The following proposition provides sufficient conditions under which uniform Lindeberg condition is satisfied.
definition[$L_{2+\delta}$-$L_2$ norm equivalence]
A mean-zero random variable/vector $H$ satisfies the uniform $L_{2+\delta}$-$L_2$ norm equivalence with a constant $L \ge 1$ if
\begin{align}
\sup_{P \in \mathcal{P}}\,\sup_{t \in \mathbb{R}^d}\, \frac{\|\langle t, H \rangle\|_{2+\delta}}{\|\langle t, H \rangle\|_{2}} \le L \quad for some \quad \delta \in (0,1].
\end{align}
definition[Uniform integrability under standardization]
A mean-zero random variable/vector $H$ is uniformly integrable under standardization if it satisfies
\begin{align}
\lim_{\kappa \to \infty}\, \sup_{P \in \mathcal{P}}\,\sup_{t \in \mathbb{R}^d}\,\operatorname{\mathbb{E}}_P\left[\frac{\langle t, H\rangle^2}{\|\langle t, H \rangle\|_{2}^2}\mathbf{1}\left\{\frac{\langle t, H\rangle^2}{\|\langle t, H \rangle\|_{2}^2} > \kappa \right\}\right] \to 0.
\end{align}
propositionThe following statement holds:
\begin{align*}
(ref) \Longrightarrow (ref) \Longrightarrow \mathrm{Uniform \, Lindeberg \, condition}.
\end{align*}
Proposition (ref) is standard in the CLT literature, and for the reader's convenience, we provide proof in Section (ref) of the supplement.
remark[On $L_{2+\delta}$-$L_2$ norm equivalence assumption]The $L_{2+\delta}$-$L_2$ norm equivalence, with a particular emphasis on $\delta=2$, is a widely employed structure in the literature of high-dimensional covariance matrix estimation minsker2018sub, mendelson2020robust and high-dimensional least squares oliveira2016lower, catoni2016pac, mourtada2022distribution.
This assumption is considerably less restrictive than imposing the
sub-Gaussianity of $X$ since any such $X$ satisfies the $L_{2+\delta}$-$L_2$ norm equivalence with $\delta\ge2$. Remarks 2.19, 2.20 and Figure S.7 of patil2022mitigating provide useful discussion and visual comparison of different norm equivalence assumptions.
remark[On consistency of the initial estimator]
Proposition (ref) establishes the validity of the confidence set in terms of the assumption on $H$, which is often easily verified in many statistical problems. This is achieved by requiring the consistency of the initial estimator. However, the existence of such an estimator may be harder to establish depending on the specific problem at hand royset2020variational. Importantly, we emphasize that the consistency of the estimator, even for some applications provided below, is not necessary. This requirement can be relaxed under alternative assumptions on the data-generating distribution that may be slightly more difficult to verify.
Statistical Applications
In this section, we present statistical applications of the proposed method and provide sufficient conditions for validity, along with the convergence rates of the corresponding confidence sets. For simplicity, we focus on the upper confidence bound $U_{n,\alpha_2}(\theta(P), \widehat{\theta}_1; D_2) = 0$ with $\alpha_2 = 0$. As discussed in Remark (ref), this choice does not affect the validity of the confidence sets. The width analysis provided here demonstrates that the simple choice of $U_{n,\alpha_2}(\theta(P), \widehat{\theta}_1; D_2) = 0$ is sufficient for rate-optimality, while alternative choices for $U_{n,\alpha_2}(\theta(P), \widehat{\theta}_1; D_2) < 0$ will only improve upon the constant factor. For each application, we specify the choice of the loss function $m_\theta$. Once such $m_\theta$ is defined, the confidence set derived from the following expression will be referred to as the
empirical Bernstein-based confidence set:
align[align omitted — 301 chars of source]
where $B_0$ and $\widehat \sigma^2_{\theta, \widehat\theta_1}$ are defined in Example (ref). We similarly define the following set as the CLT-based confidence set:
equation[equation omitted — 245 chars of source]
Throughout this section, the $N$ observations are split into two sets $D_1$ and $D_2$ such that there exist $c$ and $C$ with $0 < c < |D_1|/|D_2| \le C < \infty$. All relevant proofs are provided in Section (ref) of the supplement.
High-dimensional mean estimation
Consider an IID observation $X_1, \ldots, X_N \in \mathbb{R}^d$ generated from $P \in \mathcal{P}$. The dimension $d$ is allowed to grow with $N$. In this problem, the inference of interest is the expectation of $X$ under $P$, which can be also written as
align*[align* omitted — 128 chars of source]
The covariance matrix of $X$ is denoted by $\Sigma := \operatorname{\mathbb{E}}_P(X-\operatorname{\mathbb{E}}_P[X])(X-\operatorname{\mathbb{E}}_P[X])^\top$. Although this problem may seem trivial, mean estimation and inference in growing dimensions under weak distributional assumptions remains an active area of research lugosi2019mean. We provide results under assumptions, requiring only the existence of the covariance matrix and the uniform Lindeberg condition (ref). Below, we denote by $\widehat{\mathrm{CI}}_{n,\alpha}^{\mathrm{mean}}$ the CLT-based confidence set (ref) with $m_{\theta}(x) = \|x-\theta\|_2^2$. We now provide the validity statement for this confidence set.
theoremFor any $n\ge 1$, it holds
\begin{align*}
\inf_{P\in\mathcal{P}}\, \mathbb{P}_P(\theta(P) \in \widehat{\mathrm{CI}}_{n,\alpha}^{\mathrm{mean}}) \ge 1 - \alpha -\sup_{P\in\mathcal{P}}\Delta_{n,P}
\end{align*}
where it holds for some universal constant $C_0$,
\begin{align*}
\sup_{P\in\mathcal{P}}\Delta_{n,P} \le \min \left\{1, C_0\sup_{P\in\mathcal{P}}\,\sup_{t\in\mathbb{S}^{d-1}}\,\operatorname{\mathbb{E}}_P\left[\frac{\langle t, X-\theta(P)\rangle^2}{\|\langle t, X-\theta(P)\rangle\|_{2}^2} \min \left\{1, \frac{|\langle t, X-\theta(P)\rangle|}{n^{1/2}\|\langle t, X-\theta(P)\rangle\|_{2}}\right\}\right]\right\}.
\end{align*}
Furthermore, assume that $X-\theta(P)$ satisfies the uniform Lindeberg condition (ref), then
\begin{align*}
\liminf_{n\to \infty}\, \inf_{P\in\mathcal{P}}\, \mathbb{P}_P(\theta(P) \in \widehat{\mathrm{CI}}_{n,\alpha}^{\mathrm{mean}}) \ge 1 - \alpha.
\end{align*}
Next, we demonstrate the convergence rate of the confidence set in $L_2$-norm.
theoremLet $s_n$ be the random variable defined as (ref). For any $\varepsilon > 0$ and $\mathrm{tr}(\Sigma) \le n$, it follows
\begin{align*}
\mathbb{P}_P\left(\mathrm{Diam}_{\|\cdot\|_2}\big(\widehat{\mathrm{CI}}_{n,\alpha}^{\mathrm{mean}}\big)\le \mathfrak{C}\left\{\left(\frac{\mathrm{tr}(\Sigma)}{n}\right)^{1/2} + s_n^{1/2}\right\}\right) \ge 1-\varepsilon
\end{align*}
where $\mathfrak{C}$ only depends on $\varepsilon$ and $\alpha$. Furthermore, assume that the initial estimator satisfies
\begin{align}
\mathbb{P}_P\left(\|\widehat\theta_1 - \theta(P)\|_2^2 \le \frac{C_\varepsilon \mathrm{tr}(\Sigma)}{n} \right) \ge 1-\varepsilon
\end{align}
for any $\varepsilon > 0$ with a constant $C_\varepsilon$ depending on $\varepsilon$. It then implies
\begin{align*}
\mathbb{P}_P\left(\mathrm{Diam}_{\|\cdot\|_2}\big(\widehat{\mathrm{CI}}_{n,\alpha}^{\mathrm{mean}}\big)\le \mathfrak{C}\left(\frac{\mathrm{tr}(\Sigma)}{n}\right)^{1/2}\right) \ge 1-\varepsilon
\end{align*}
for all $n \ge N_\varepsilon$ with $N_\varepsilon$ depending on $\varepsilon$ and $\mathfrak{C}$ depending on $\varepsilon$ and $\alpha$.
The requirement (ref) is satisfied when $\widehat{\theta}_1$ is the sample mean. We emphasize that Theorem (ref) imposes no restrictions on the dimension $d$---hence it is typically called dimension agnostic. The dependency on $\sqrt{\mathrm{tr}(\Sigma)/n}$ is, in fact, not improvable since this is the exact risk of the mean estimation under a multivariate Gaussian distribution; see the formal minimax argument in Section 5 of lee2022optimal.
The results in this subsection can be extended to inference for the Fr\'{e}chet mean where $\Theta$ is a general metric space. In the literature, assumption (ref) is referred to as a growth condition, or variance inequality. The quadruple condition studied in Schotz2019convergence can be used to establish (ref). We anticipate that the proposed confidence set for this problem will yield a width that shrinks at a rate determined by the growth of the entropy bounds on the space $\Theta$. A detailed exploration of this extension is deferred to future work.
Misspecified linear regression
Consider an IID observation $(Y_1, X_1^\top)^\top, \ldots, (Y_N, X_N^\top)^\top \in \mathbb{R}\times\mathbb{R}^d$ generated from the following model:
align*[align* omitted — 162 chars of source]
We assume that the gram matrix $\Gamma_P := \operatorname{\mathbb{E}}_P[XX^\top]$ is invertible such that $\theta_P \in \mathbb{R}^d$ exists even when the regression function $\operatorname{\mathbb{E}}_P[Y_i| X_i]$ is not linear. In this problem, the inference of interest is $\theta_P$, which can be also written as
equation*[equation* omitted — 135 chars of source]
without making the linearity assumption for $\operatorname{\mathbb{E}}_P[Y_i| X_i]$. Below, we denote by $\widehat{\mathrm{CI}}_{n,\alpha}^{\mathrm{LR}}$ the CLT-based confidence set (ref) with $m_{\theta}(y,x) = (y-\theta^\top x)^2$. We introduce the following assumptions for the validity statement:
enumerate[label=(B\arabic*),leftmargin=2cm]
• There exist constants $q_x \ge 4$ and $L \ge 1$ such that
\begin{align*}
\|t^\top \Gamma_P^{-1/2}X_i\|_{q_x}\le L \quad for all \quad t \in \mathbb{S}^{d-1}.
\end{align*}
• There exist positive constants $\underbar{$\sigma$}, \overline{\sigma}$ such that $\underbar{$\sigma$} < \sigma_i < \overline{\sigma}$ for all $1 \le i \le n$.
(ref) requires the $L_{q_x}$-$L_2$ norm equivalence on $X$---see Remark (ref). (ref) ensures that the error variables $\xi$ do not become degenerate or possess infinite variance.
The following result is obtained:
theoremAssume (ref) and (ref). Then for any $n\ge1$ and $\varepsilon > 0$, it holds
\begin{align*}
\inf_{P\in\mathcal{P}}\, \mathbb{P}_P(\theta_P \in \widehat{\mathrm{CI}}_{n,\alpha}^{\mathrm{LR}}) \ge 1 - \alpha - \sup_{P\in\mathcal{P}}\Delta_{n,P}
\end{align*}
where it holds for some universal constant $C_0$,
\begin{align*}
\sup_{P\in\mathcal{P}}\Delta_{n,P}&\le \min \left\{1, C_0\sup_{P\in\mathcal{P}}\bigg[\inf_{\varepsilon > 0}\left\{\{1 \wedge \underbar{$\sigma$}^{-1}L^2 \lambda^{1/2}_{\max}(\Gamma_P)\varepsilon\} + \mathbb{P}_P(\|\widehat\theta_1 - \theta_P\| > \varepsilon) \right\}\right.\\
&\left.\qquad + \sup_{\|t\|=1}\,\operatorname{\mathbb{E}}_P\left[\frac{\langle t, X\xi \rangle^2}{\|\langle t, X\xi \rangle\|^2_{2}} \min\left\{1, \frac{|\langle t, X\xi \rangle|}{n^{1/2}\|\langle t, X\xi \rangle\|_{ 2}}\right\}\right]\bigg]\right\}
\end{align*}
and $\lambda_{\max}(\Gamma_P)$ denotes the largest eigenvalue of the gram matrix, which depends on $P$. Furthermore, assume the existence of a positive constant $\overline{\lambda} < \infty$ such that $\sup_{P\in\mathcal{P}}\lambda_{\max}(\Gamma_P) \le \overline{\lambda}$, the uniform Lindeberg condition (ref) on the mean-zero random variable $\xi_i X_i$, and the uniform consistency of the initial estimator such that
\begin{align*}
\lim_{n\to \infty}\, \sup_{P\in\mathcal{P}}\,\mathbb{P}_P(\|\widehat\theta_1 - \theta_P\| > \varepsilon) = 0.
\end{align*}
Then, we obtain
\begin{align*}
\liminf_{n\to \infty}\,\inf_{P\in\mathcal{P}}\, \mathbb{P}_P(\theta_P \in \widehat{\mathrm{CI}}_{n,\alpha}^{\mathrm{LR}}) \ge 1 - \alpha.
\end{align*}
Next, we provide the width of the confidence set in terms of the matrix-norm $\|\cdot\|_{\Gamma_P}$:
theoremLet $s_n$ be the random variable defined as (ref). Assume (ref) and (ref). For any $n \ge \left(L^2d^{2/q_x}\log(d)\right)^{q_x/(q_x-2)}$, it follows
\begin{align*}
\mathbb{P}_P\left(\mathrm{Diam}_{\|\cdot\|_{\Gamma_P}}\big(\widehat{\mathrm{CI}}_{n,\alpha}^{\mathrm{LR}}\big)\le \mathfrak{C}\left\{\left(\frac{\overline{\sigma}^2 d}{n}\right)^{1/2} + s_n^{1/2}\right\}\right) \ge 1-\varepsilon
\end{align*}
where $\mathfrak{C}$ depends on $\varepsilon$ and $\alpha$. Furthermore, assume that the initial estimator satisfies
\begin{align}
\mathbb{P}_P\left(\|\widehat\theta_1 - \theta(P)\|_{\Gamma_P}^2 \le \frac{C_\varepsilon\overline{\sigma}^2d}{n} \right) \ge 1-\varepsilon
\end{align}
for any $\varepsilon > 0$ with a constant $C_\varepsilon$ depending on $\varepsilon$. It then implies
\begin{align*}
\mathbb{P}_P\left(\mathrm{Diam}_{\|\cdot\|_{\Gamma_P}}\big(\widehat{\mathrm{CI}}_{n,\alpha}^{\mathrm{LR}}\big)\le \mathfrak{C}\left(\frac{L^2\overline{\sigma}^2d}{n}\right)^{1/2}\right) \ge 1-\varepsilon
\end{align*}
where $\mathfrak{C}$ depends on $\varepsilon$ and $\alpha$.
The requirement (ref) is satisfied when $\widehat{\theta}_1$ is the ordinary least squares (OLS). The dependency on $\overline{\sigma}^2 d/n$ in the rate is minimax optimal as shown by Theorem 1 of mourtada2022exact. Recently, chang2024confidence provided confidence sets for high-dimensional OLS under similar conditions as above. Their proposal is based on the formulation of the OLS as the root of an estimating equation, or $Z$-estimation. Theorem 10 of chang2024confidence establishes that their proposed confidence set converges at the optimal rate of $\sqrt{d/n}$; however, they require $d = o(n^{1/2})$ under the equivalent condition as (ref). chang2023inference provides a method based on one-step bias-correction, which requires $d = o(n^{2/3})$. On the other hand, Theorem (ref) holds as long as $C d(\log d)^{2} \le n$ (when $q_x = 4$) for some constant $C$ depending on $P$ but not on $n$ or $d$.
remarkOne can observe that
\begin{align*}
\mathbb{C}_P(\theta) = \|\theta - \theta_P\|^2_{\Gamma_P} \ge \lambda_{\min}(\Gamma_P)\|\theta - \theta_P\|^2_2
\end{align*}
and (ref) holds with $c_0 = \lambda_{\min}(\Gamma_P)$ and $\beta=1$. Theorem (ref) thus can be stated in terms of the $L_2$-norm with an additional assumption that $\lambda_{\min}(\Gamma_P) > \underbar{$\lambda$}$ for some constant $\underbar{$\lambda$} > 0$.
remarkThe proposed framework can be easily extended to penalized least squares, where the minimizer is defined as\begin{equation*}
\theta_P := \operatorname*{arg\,min}_{\theta \in \mathbb{R}^d}\, \operatorname{\mathbb{E}}_P[(Y-\theta^\top X)^2] + \lambda(\theta),
\end{equation*}
and $\lambda : \Theta \mapsto \mathbb{R}_+\cup\{+\infty\}$ is a penalization term or a constraint that may depend on $n$ (but not on the data). In this case, we can construct the confidence set as
\begin{align*}
\left\{\theta \in \Theta \, : \, \mathbb{P}_n(m_{\theta}-m_{\widehat \theta_1}) + \lambda(\theta) - \lambda(\widehat \theta_1)\le n^{-1/2}z_{\alpha} \widehat\sigma_{\theta, \widehat\theta_1}\right\}
\end{align*}
with $m_{\theta}(y,x) = (y-\theta^\top x)^2$ and $\widehat\sigma_{\theta, \widehat\theta_1}$ is defined as in (ref). The validity of this set holds under the same assumptions as Theorem (ref). However, the corresponding width analysis requires considerably more effort, as it involves the limiting behavior of the sequence $\{\lambda(\theta_P) - \lambda(\widehat \theta_1)\}$ as $\|\widehat \theta_1 - \theta_P\| = o_P(1)$.
Manski's discrete choice model
Consider an IID observation $(Y_1, X_1^\top)^\top, \ldots (Y_N, X_N^\top)^\top \in \{-1,1\} \times \mathbb{R}^d$ generated from the following binary response model:
align[align omitted — 140 chars of source]
Here, the error variable $\varepsilon_i$ has zero conditional median given covariates, i.e., $\mathrm{med}(\varepsilon_i | X_i)=0$, but otherwise allowed to depend on $X_i$. In this problem, the inference of interest is $\theta_P$, which can be also written as
align*[align* omitted — 149 chars of source]
A popular estimator for $\theta_P$ is defined as the following maximum score estimator, proposed by manski1975maximum, which solves
align[align omitted — 189 chars of source]
The asymptotic behavior of $\widehat \theta_n$ exhibits a non-standard limit, studied by manski1985semiparametric, kim1990cube. The inference for this problem is known to be challenging, and cattaneo2020bootstrap proposes a bootstrap-based approach. By the fact that $Y_i\, \mathrm{sgn}(\theta^\top _i X_i)$ is uniformly bounded by one, we can use the confidence set based on the concentration inequality with $B_0 = 2$. Below, we denote by $\widehat{\mathrm{CI}}^{\mathrm{Manski}}_{n,\alpha}$ the empirical-Bernstein-based confidence set (ref) with $m_{\theta}(x,y) = -y\cdot\mathrm{sgn}(\theta^\top x)$. We now provide the validity statement for this confidence set.
theoremFor any fixed $n\ge 1$,
\begin{align*}
\inf_{P\in\mathcal{P}}\, \mathbb{P}_P(\theta_P \in \widehat{\mathrm{CI}}^{\mathrm{Manski}}_{n,\alpha}) \ge 1-\alpha.
\end{align*}
proof[{Proof of Theorem (ref)}]
The validity follows immediately in view of Theorem (ref).
Next, we demonstrate the width of the confidence set in $L_2$-norm. Below, we introduce additional assumptions on the joint distribution on $X$ and $\varepsilon$. Both assumptions are standard in the literature.
enumerate[label=(B\arabic*),leftmargin=2cm]
\setcounter{enumi}{2}
• Let $P \in \mathcal{P}$ be the joint distribution on $X, \varepsilon$ and $\eta_P(x) := \mathbb{P}_P(Y=1 \mid X=x)$. There exist constants $C_0$ and $0 < t^* < 1/2$ with $C_0t^* > 1$, and $\beta < \infty$, such that
\begin{align*}
\mathbb{P}_P\left(\left|\eta_P(X)- \frac{1}{2}\right| < t\right) \le C_0t^{1/\beta} \quad for all \quad 0 < t < t^*.
\end{align*}
• There exists a constant $c_1$, not depending on $n$ or $d$, such that
\begin{align*}
c_1\|\theta-\theta_P\|_2 \le \mathbb{P}_X\left(\mathrm{sgn}(\theta^\top X) \neq \mathrm{sgn}(\theta_P^\top X)\right)
\end{align*}
for all $\theta \in \mathbb{S}^{d-1}$ where $\mathbb{P}_X$ denotes the probability measure under the marginal distribution of $X$.
Both assumptions are identical to those considered in mukherjee2019nonstandard, mukherjee2021optimal. (ref) is often called the low noise (the margin) assumption in the classification literature mammen1999smooth, tsybakov2004optimal. This condition quantifies the deviation of the conditional class probability from $1/2$ near the decision boundary of the Bayes' classifier, i.e., $\mathbb{P}_P(Y=1|X)\ge 1/2$. As $\beta \to 0$, the decision boundary becomes bounded away from $1/2$, representing the most favorable situation for the classification. (ref) was introduced to relate the distribution of covariates $X$ to the underlying geometry in the parameter space $\mathbb{S}^{d-1}$. See mukherjee2021optimal for further discussion.
We provide the width of the confidence set in terms of the $L_2$-norm:
theoremLet $s_n$ be the random variable defines as (ref). Assume (ref) and (ref). Then for any $d \le n$, it follows
\begin{align*}
\mathbb{P}_P\left(\mathrm{Diam}_{\|\cdot\|_2}\big(\widehat{\mathrm{CI}}^{\mathrm{Manski}}_{n,\alpha}\big)\le \mathfrak{C}\left\{\left(\frac{d\log(n/d)}{n}\right)^{1/(1+2\beta)} + s_n^{1/(1+\beta)}\right\}\right) \ge 1-\varepsilon
\end{align*}
where $\mathfrak{C}$ depends on $\alpha, \varepsilon, c_1,\beta,t^*$ and $C_0$. Furthermore, assume that the initial estimator satisfies
\begin{align}
\mathbb{P}_P\left(\|\widehat\theta_1-\theta(P)\|_2 \le C_\varepsilon\left(\frac{d\log(n/d)}{n}\right)^{1/(1+2\beta)}\right) \ge 1-\varepsilon
\end{align}
for any $\varepsilon > 0$ with a constant $C_\varepsilon$ depending on $\varepsilon$. It then implies
\begin{align*}
\mathbb{P}_P\left(\mathrm{Diam}_{\|\cdot\|_2}\big(\widehat{\mathrm{CI}}^{\mathrm{Manski}}_{n,\alpha}\big)\le \mathfrak{C}\left(\frac{d\log(n/d)}{n}\right)^{1/(1+2\beta)}\right) \ge 1-\varepsilon
\end{align*}
for all $n \ge N_\varepsilon$ with $N_\varepsilon$ depending on $\varepsilon$ and $\mathfrak{C}$ depending on $\varepsilon$ and $\alpha$.
Theorem 3.2 of mukherjee2019nonstandard shows that the rate requirement (ref) is satisfied by the standard maximum score estimator defined as (ref). The rate of convergence matches that of the minimax lower bound up to the logarithmic factor (See Theorem 3.4 of mukherjee2019nonstandard).
Quantile estimation without positive densities
Consider an IID observation $X_1, \ldots, X_N \in \mathbb{R}$ generated from $P \in \mathcal{P}$ where the inference of interest is the $\gamma$-quantile defined as
align*[align* omitted — 112 chars of source]
and $F_P(t) := \mathbb{P}_P(X\le t)$. It is well-known that $\theta(P)$ minimizes the following “quantile" loss:
align*[align* omitted — 154 chars of source]
where $(t)_+ = \max(t, 0)$. The sample quantile centered at $\theta(P)$ converges to a Gaussian distribution when scaled by $n^{1/2}$ if the distribution of $X$ has a strictly positive density at $\theta(P)$. If the density at $\theta(P)$ is zero or non-existent, however, the sample quantile converges at a rate depending on the H\"{o}lder smoothness of the $F_P(t)$ in the neighborhood of $\theta(P)$. In this case, the limiting distribution is no longer Gaussian and also depends on the H\"{o}lder smoothness of the $F_P(t)$ in the neighborhood of $\theta(P)$ smirnov1952limit. Although finite-sample valid, distribution-free confidence intervals for quantiles already exist scheffe1945non, we present this result to illustrate the behavior of the proposed method in irregular settings.
Below, we construct the CLT-based confidence set (ref) with $m_{\theta}(x) = \gamma(x-\theta)_+ + (1-\gamma)(\theta-x)_+$, and denote the corresponding set by $\widehat{\mathrm{CI}}_{n,\alpha}^{\gamma}$. We quantity the smoothness of the $F_P(t)$ near $\theta(P)$ as follows:
enumerate[label=(B\arabic*),leftmargin=2cm]
\setcounter{enumi}{4}
• There exist $\delta >0$, $M_0,M_1 \in (0,\infty)$ and $M_0 >M_1$ such that
\begin{align*}
|F(\theta) - F(\theta(P)) - M_0|\theta - \theta(P)|^{\beta} \mathrm{sgn}(\theta-\theta(P))| \le M_1|\theta-\theta(P)|^\beta
\end{align*}
for all $\theta$ such that $|\theta-\theta(P)| \le \delta$.
The H\"{o}lder smoothness as described in (ref) should be compared to knight1998limiting. When $\beta=1$, this assumption becomes equivalent to requiring that the density at the true $\gamma$-quantile is bounded away from zero. We now proceed to the validity statement.
theoremLet $\{q_n\}$ be a deterministic and non-decreasing sequence such that $q_n \ge 1$ for all $n$, and
\begin{align*}
\lim_{n\to \infty}\,\sup_{P \in \mathcal{P}}\mathbb{P}_P(q_n|\widehat \theta_1 - \theta(P)| > \delta) = 0
\end{align*}
where $\delta$ corresponds the one defined in (ref). Then assuming $\{n\gamma(1-\gamma)\}^{-1} = o(1)$ and $\{q_n^\beta\gamma(1-\gamma)\}^{-1}= O(1)$, it follows that
\begin{align*}
\liminf_{n\to \infty}\,\inf_{P\in\mathcal{P}}\, \mathbb{P}_P(\theta(P) \in \widehat{\mathrm{CI}}_{n,\alpha}^{\gamma}) \ge 1 - \alpha.
\end{align*}
The result of Theorem (ref) allows for $\gamma \equiv \gamma_n$ to depend on the sample size $n$. When there exist constants $c$ and $C$ such that $0 < c \le \gamma_n \le C < 1$ for all $n$, the validity holds under the consistency of the initial estimator. When $\gamma_n \to 0$ or $\gamma_n \to 1$, then there is a restriction on how quickly $\gamma_n$ can tend to these extreme values depending on $\beta$ and the convergence rate of $\widehat{\theta}_1$. This requirement may be relaxed under the alternative, but possibly stronger, assumptions on $P$---see Remark (ref).
Finally, the width of the confidence set is obtained as follows:
theoremLet $s_n$ be the random variable defined as (ref). Then for all $n \ge C_1 \delta^{-2\beta}$ with $\delta$ corresponding to the one in (ref), it holds that
\begin{align*}
\mathbb{P}_P\left(\mathrm{Diam}_{|\cdot|}\big(\widehat{\mathrm{CI}}_{n,\alpha}^{\gamma}\big)\le \mathfrak{C}\tau_n\right) \ge 1-\varepsilon
\end{align*}
where
\begin{align*}
\tau_n &:= \left(\frac{1}{n^{1/2\beta}} + s_n^{1/(1+\beta)}\right)\mathbb{P}_P(|\widehat\theta_1-\theta_0|<C_2\delta^{1+\beta}) + \delta^{-\beta}\operatorname{\mathbb{E}}_P[|\widehat\theta_1 - \theta_0|\mathbf{1}\{|\widehat\theta_1-\theta_0|\ge C_2\delta^{1+\beta}\}]
\end{align*}
and $\mathfrak{C}$, $C_1$, and $C_2$ depend on $\varepsilon$, $M_0$, $M_1$, $\beta$ and $\alpha$. Furthermore, assume that the initial estimator satisfies
\begin{align}
\mathbb{P}_P\left(|\widehat\theta_1 - \theta(P)|^2 \le C_\varepsilon n^{-1/\beta}\right) \ge 1-\varepsilon
\end{align}
for any $\varepsilon > 0$ with a constant $C_\varepsilon$ depending on $\varepsilon$. It then implies for all $n \ge N_{\varepsilon,\delta}$ with $N_{\varepsilon,\delta}$ depending on $\varepsilon$, $M_0$, $M_1$, $\beta$, $\alpha$ and $\delta$,
\begin{align*}
\mathbb{P}_P\left(\mathrm{Diam}_{|\cdot|}\big(\widehat{\mathrm{CI}}_{n,\alpha}^{\gamma}\big)\le \mathfrak{C}n^{-1/(2\beta)}\right) \ge 1-\varepsilon
\end{align*}
where $\mathfrak{C}$ only depends on $\varepsilon$, $\alpha$, $\beta$ $M_0$ and $M_1$.
The requirement in (ref) is satisfied when $\widehat{\theta}_1$ is the sample quantile. The convergence rate of the confidence set indicated by Theorem (ref) matches that of the sample quantile under (ref) as given by Example 1 of knight1998limiting. In fact, knight1998limiting provides the limiting distribution of the sample quantile under (ref); however, the traditional approach requires prior knowledge of $\beta$ to perform asymptotic inference. The proposed confidence set of this manuscript, in contrast, converges at the rate $n^{-1/(2\beta)}$ automatically without the knowledge of $\beta$. In the special case of $\beta=1$, i.e., the density is bounded away from zero at the $\gamma$-quantile, and the confidence shrinks at the parametric rate of $n^{-1/2}$.
remark[The proof of Theorem (ref)]
The width analyses in this section are mostly performed as the direct applications of Theorem (ref) or Theorem (ref). The proof of Theorem (ref), however, differs significantly since the curvature assumption (ref) only holds locally for $\theta$ such that $\|\theta-\theta(P)\| < \delta$. As a result, Theorem (ref) cannot be applied for the parameter outside of this neighborhood. While the proof of Theorem (ref) crucially relies on the convexity and Lipschitzness of the quantile loss, a general result in the spirit of Theorem (ref) may be useful under the local analog to (ref). Towards this task, one may need to extend the ratio-type empirical process to the unbounded function spaces gine2006concentration. We are currently investigating this direction.
Discrete argmin inference
This application is motivated by the recent manuscript of zhang2024winners, which studies a prototypical problem in general model selection. The result in this subsection can be extended to constructing a confidence set that contains the best “predictor" minimizing the population risk among $\{f_1,\ldots, f_d\}$ where $f_i$ corresponds to machine learning methods developed from the same data, and consequently $f_i$ and $f_j$ for $i\neq j$ can be highly correlated.
Consider an IID observation $X_1, \ldots, X_N \in \mathbb{R}^d$ generated from $P \in \mathcal{P}$ where $d$ may grow with $N$. In this problem, the inference of interest is the index of the coordinate corresponding to the minimum marginal mean. In other words, the population parameter is defined as
align*[align* omitted — 123 chars of source]
Following zhang2024winners, we do not assume uniqueness of $\theta(P)$, allowing the proposed confidence set to accommodate ties. Below, we denote by $\widehat{\mathrm{CI}}_{n,\alpha}^{\mathrm{argmin}}$ the CLT-based confidence set (ref) with $m_{\theta}(x) = e_\theta^\top x$ for $\theta \in \{1, \ldots, d\}$. We introduce the following assumption to the class of distributions $\mathcal{P}$:
enumerate[label=(B\arabic*),leftmargin=2cm]
\setcounter{enumi}{5}
• For any $i \neq j $, it holds that
\begin{align*}
\sup_{P\in\mathcal{P}}\, \operatorname{\mathbb{E}}_P\left[\frac{W_{ij}^2}{\operatorname{\mathbb{E}}_P [W_{ij}^2]} \min\left\{1, \frac{|W_{ij}|}{n^{1/2}(\operatorname{\mathbb{E}}_P [W_{ij}^2])^{1/2}}\right\} \right] = o(1)
\end{align*}
as $n \to\infty$ where $W_{ij} = e_i^\top (X-\operatorname{\mathbb{E}}_P[X])- e_j^\top (X-\operatorname{\mathbb{E}}_P[X])$.
zhang2024winners establish the validity of their method under the assumption that the smallest eigenvalue of the covariance matrix of $X$ is bounded away from zero (See Theorem 3.1 of zhang2024winners). However, as pointed out in their work, this assumption may be violated in practice when the components of $X$ are highly correlated, such as in model selection for LASSO (see Section 6.2 of zhang2024winners). In contrast, (ref) avoids imposing a condition on the smallest eigenvalue, thereby allowing for scenarios where the components of $X$ are strongly correlated. Additionally, zhang2024winners develop their theoretical results under the assumption that $e_j^\top X_i$ is almost surely uniformly bounded for all $1\le j \le d$ and $1 \le i \le n$ for some constant. Once such a constant becomes available, we can construct a confidence set based on concentration inequalities where the finite-sample validity can be established without assuming (ref).
We now establish the validity of this confidence set.
theoremAssume (ref). Then, we obtain
\begin{align*}
\liminf_{n\to \infty}\,\inf_{P\in\mathcal{P}}\, \mathbb{P}_P(\theta(P) \in \widehat{\mathrm{CI}}_{n,\alpha}^{\mathrm{argmin}}) \ge 1 - \alpha.
\end{align*}
The corresponding width analysis requires a new framework, as the parameter space $\Theta$ is discrete and not continuous. One potential approach to establish the convergence rate is through the analysis of the non-central t-statistic Bentkus2007Limiting. Interestingly, the limiting law of the non-central
t-statistic can be highly non-standard, depending on the underlying moment assumptions and the rate at which the population mean decays (See Theorems 2.4 and 2.5 of Bentkus2007Limiting). Using these results, one may be able to establish a result akin to Theorem 4.1 of zhang2024winners, where the confidence set's rejection probability is analyzed under different rates at which the curvature $\mathbb{C}_P(\widehat \theta_1)$ converges to zero.
Numerical Illustration
This section provides an empirical illustration of the proposed method for inference in high-dimensional settings. The aim of this section is to demonstrate the robust coverage of the proposed methods in high dimensions, where Wald intervals may be anti-conservative. Additionally, we compare the volume of the proposed confidence set with that of the Wald intervals, highlighting that the improved robustness comes at a minor constant-factor cost.
The scenario under consideration is identical to the high-dimensional mean estimation setup from subsection (ref). For $N=500$ and $2 \le d \le 100$, independent observations $X_i \in \mathbb{R}^d$, $1 \le i \le N$ are generated from a multivariate Normal distribution: $X_i \sim N(0_d, I_d)$. We compare three methods. The first method, Asymptotic, corresponds to the Wald interval based on the asymptotic distribution. Formally, the confidence set is defined as
align*[align* omitted — 223 chars of source]
where $\overline X_N$ is the sample mean and $\widehat\Sigma_N$ is the sample covariance matrix
computed from the $N$ observations. The cutoff value $\chi^2_{d,\alpha}$ represents the ($1-\alpha$)th-quantile of the chi-square distribution with
$d$ degrees of freedom. The second method, Sample split, corresponds to the proposed approach theoretically analyzed in subsection (ref). To recall, we split the observations into two parts \[D_1 := \{X_i \, : \, 1\le i \le \lfloor N/2\rfloor \} \text{ and } D_2 := \{X_i \, : \, \lfloor N/2\rfloor + 1 \le i \le N\}.\] Let $\widehat \theta_1$ denote the sample mean computed from $D_1$, and denote $n = |D_2|$. The confidence set, Sample split, is then given by
align*[align* omitted — 267 chars of source]
where $\widehat \sigma_{\theta, \widehat \theta_1}$ is the sample standard deviation of $\|X_i -\theta\|^2 - \|X_i -\widehat \theta_1\|^2$ computed using $D_2$. The final method slightly modifies the sample-splitting approach by incorporating the upper bound from (ref). In particular, for mean estimation, it holds that $-\mathbb{C}_P(\widehat{\theta}_1) = -\|\widehat{\theta}_1 - \theta(P)\|^2$ and thus
align*[align* omitted — 134 chars of source]
The confidence set, Sample split + Upper bound, is obtained by the expression (ref):
align*[align* omitted — 304 chars of source]
Figure (ref) displays the empirical coverage of the $95\%$ confidence sets, calculated from $1000$ replications. We observe that the coverage accuracy of the Wald intervals deteriorates as the dimension increases. By $d \approx 100$, the average coverage of the $95\%$ confidence set drops to approximately $50\%$. In contrast, the proposed sample-splitting methods remain their validity across all dimensions we investigated. The simulation results further highlight the empirical advantages of estimating the upper bounds of the curvature $-\mathbb{C}_P(\widehat\theta_1)$. While the confidence set $\widehat{\mathrm{CI}}^{\mathrm{SS}}_{N,\alpha}$, analyzed in subsection (ref) to show the optimal shrinkage rate, its empirical coverage suggests that the method is conservative in practice. Meanwhile, incorporating the upper bound significantly improves coverage accuracy, bringing it close to the nominal $95\%$ level across all dimensions.
figure[figure omitted — 926 chars of source]
Next, we analyze the size of the proposed confidence set in comparison to the Wald intervals. As illustrated in Figure (ref) in the supplement, the proposed confidence set is non-convex and the exact computation of the volume is challenging. To address this, we adopt the following approximation. First, we can establish that there exists a random variable $r \in \mathbb{R}_+$ such that
align*[align* omitted — 130 chars of source]
Meanwhile, Lemma (ref) in Section (ref) of the supplement implies that the volume of the Wald interval is given by
align*[align* omitted — 227 chars of source]
Here, $\Gamma$ denotes the Euler's gamma functions and $\lambda_{k}(\widehat \Sigma_N)$ denotes the $k$th smallest eigenvalue of the sample covariance matrix $\widehat \Sigma_N$ computed from all $N$ observations. Then, it is straightforward to see that
align*[align* omitted — 389 chars of source]
The final expression represents the ratio between the radius $r$ and the geometric mean of the ellipsoid's radii $(r_1, \ldots, r_d)$. A detailed derivation is provided in Section (ref) of the supplement. Figure (ref) presents the average ratio of the radii computed from $1000$ replications across different dimensions. The results roughly state that the proposed confidence set enlarges the Wald intervals in each of the $d$-dimensional axes by at most a factor of $2$. We emphasize that this is an upper bound of the ratio and two confidence sets can be much closer in size in practice. See Figure (ref) in the supplement for visual illustration. This enlargement is expected since the proposed methods employ sample-splitting, which reduces sample efficiency but enhances the validity of inference in high-dimensional settings.
figure[figure omitted — 802 chars of source]
Concluding Remarks
This manuscript introduces a general approach to constructing confidence sets for M-estimation, addressing longstanding challenges in statistical inference due to inherent irregularities in these tasks. The proposed method employs sample splitting, facilitating validity across regular and irregular settings. In particular, the method offers a dimension-agnostic solution that is particularly valuable for high-dimensional problems where inferential tools are currently limited or unavailable. The general framework is illustrated through two approaches: one based on concentration inequalities and the other on the central limit theorem (CLT). The manuscript provides foundational theorems for each case that guarantee validity and characterize their convergence rates.
The theoretical properties of the proposed methods are demonstrated through statistical applications where inference is challenging. The first application considers mean estimation and misspecified linear regression in growing dimensions. The proposed methods remain valid even when the dimension $d$ exceeds $n^{1/2}$, a regime where many existing inferential tools fail. The convergence rates of the proposed method are also dimension-agnostic and match those of the known minimax lower bounds---these properties have been previously established in estimation but have been less known for inference. The second application extends the proposed methods to irregular settings, such as cube-root asymptotic, whose convergence rates are shown to adapt to the underlying geometry of the optimization problem at hand.
The authors are currently exploring several directions for extending the proposed method. One key area is the constrained optimization problems, where similar irregularities emerge when the solution lies on the boundary of the constrained set. This issue includes problems involving shape constraints and sparsity. In particular, there is a lack of inferential tools for LASSO and Dantzig selectors candes2007dantzig despite their widespread use, making the investigation in this area of significant interest.
Several open problems remain where the proposed methods may be extended. First, inference for generalized linear models in growing dimensions, including logistic regression and Poisson regression, remains underdeveloped. Second, the extension to the Cox proportional hazard model may be interesting. This problem involves nuisance parameters, requiring the optimization in the form of $\operatorname{\mathbb{M}}(\theta, \eta, P)$ for some unknown $\eta \in \mathcal{H}$, where $\mathcal{H}$ can be a nonparametric class of functions. Achieving efficiency in estimating $\operatorname{\mathbb{M}}(\theta, \eta, P)$ will likely require the first-order bias correction from the estimation error of $\eta$, which is an interesting area to investigate. Third, the manuscript did not focus on optimization problems involving U-statistics or U-quantiles. Given that the CLT for U-statistics is well-established, we anticipate the general framework to be applicable to these problems as well. Finally, this manuscript considered scenarios where the sample size $n$ is fixed and not data-dependent. We envision extending this framework to data-dependent stopping rules, or anytime-valid inference, can be achieved when the loss function is equipped with certain concentration properties, such as sub-Gaussian tails schreuder2020nonasymptotic or under the asymptotic confidence sequence framework formalized by waudby2024time. Pursuing these extensions will require considerable additional effort and represent substantial methodological advances.
Acknowledgements
The first author gratefully acknowledges Woonyoung Chang for the series of helpful discussions. We also thank Christof Sch\"{o}tz for providing us comments on the proof of Theorem (ref) in the initial manuscript and informing us of the application to Fr\'{e}chet means.
\setcounter{section}{0}
\setcounter{equation}{0}
\setcounter{figure}{0}
center[center omitted — 133 chars of source]
abstractThis supplement contains the proofs of all the main results in the paper and some supporting lemmas.
Review of History
This section summarizes the historical developments of the key idea behind the content of this manuscript.
itemize• Inverting the asymptotic risk of (irregular) estimators to construct confidence sets has a long history. The idea dates back at least to Stein1981, who foreshadowed the possibility of the inference for a shrinkage estimator of the multivariate Gaussian mean in large dimensions.
• Later, Li1989 applied a similar risk inversion framework to nonparametric regression with Gaussian errors and introduced the concept of “honest” confidence sets, which subsequently sparked developments in adaptive nonparametric inference. beran1996 and Beran1998 extended these ideas, explicitly crediting Stein1981, and proposed the inversion of the asymptotic normality based on the central limit theorem (CLT), which they call modulation of estimators. Here, sample-splitting was not considered.
• Parallel developments in stochastic programming analyzed the risk of constrained optimization problems. Early references include Shapiro1989 and geyer1994asymptotics, both of whom leveraged the CLT to establish asymptotic distributions. Confidence set construction in this setting was explicitely mentioned by geyer1996asymptotics. The asymptotic behavior of the risk under general loss functions and constraints was studied in great generality by pflug1991asymptotic, pflug1995asymptotic, pflug2003stochastic. Here, as well, sample splitting was not considered.
• Robins2006 were among the first to combine sample-splitting with with risk inversion based on the CLT. Their Theorem 3.4 established that the validity of the CLT depended only on the sample size tending to infinity and a Feller condition on the univariate risk space, making their result effectively “dimension-agnostic".
During this period, statistical literature primarily focused on squared error loss in nonparametric regression, with exceptions such as Hoffmann2011, Carpentier2013. Meanwhile, operations research literature examined inference for constrained optimization solutions with general loss functions. Inspired by the series of works by pflug1991asymptotic, pflug1995asymptotic, pflug2003stochastic, vogel2008universal investigated risk inversion for confidence sets, but without incorporating sample-splitting. Related works include vogel2008confidence, vogel2017confidence and guigues2017non.
• kim2020dimension later introduced the term “dimension-agnostic" to describe properties similar to those established by Robins2006, though they did not cite the earlier work. For instance, Theorem 4.2 of kim2020dimension should be compared to Theorem 3.4 of Robins2006. They applied sample-splitting and CLT-based inversion to high-dimensional hypothesis testing problems, such as goodness-of-fit testing using Gaussian maximum mean discrepancy (MMD). chakravarti2019gaussian uses the similar methodology based on sample-splitting and CLT for testing the relative fit of Gaussian mixtures. park2023robust also employ sample-splitting and CLT for the inference on population maximum likelihood estimation under model misspecification, among other techniques. Neither chakravarti2019gaussian, kim2020dimension nor park2023robust developed a general theory for M-estimation such as width/diameter of the resulting confidence sets.
• Although the explicit application of sample-splitting and CLT inversion to general M-estimation has not been previously explored, the methodological approach is a natural consequence of prior work, including Beran1998, Robins2006, and vogel2008confidence. Consequently, we do not claim innovation in methodological front, as such confidence sets would likely have emerged given the historical trajectory of the field. Instead, our contribution lies in analyzing the properties of these confidence sets, including their validity and width.
• M-estimators are known to exhibit locally adaptive rates of convergence, depending on problem-specific geometric factors such as curvatures kim1990cube, van1996weak. This notion of adaptivity has not been investigated within the “adaptive" inference literature on nonparametric submodels, such as Robins2006 and Patschkowski2019. The concept of locally adaptive confidence sets is particularly relevant to the M-estimation framework, holding both methodological and theoretical significance. While some adaptive confidence sets have been studied in shape-restricted regression Yang2019, dumbgen2003optimal, Bellec2021, these approaches typically assume strong distributional conditions such as (sub-)Gaussian errors. One of the key contributions of this work is the establishment of adaptive confidence sets for general M-estimation under weaker distributional assumptions.
Research on adaptive confidence sets for nonparametric models was particularly active from the 1990s to the 2010s. These studies generally relied on concentration inequalities to establish validity, requiring precise error quantification for adaptive nonparametric estimators. Additional historical developments can be found in Chapter 8.4 of gine2021mathematical.
Proof of Theorem (ref)
We denote by $\operatorname{\mathbb{M}}(\theta) \equiv \operatorname{\mathbb{M}}(\theta, P)$ for any $\theta \in \Theta$. Note that
align*[align* omitted — 840 chars of source]
where the last inequality follows from Cantelli's inequality pinelis2010between.
Proof of Theorem (ref)
Let $P$ be an arbitrary distribution in $\mathcal{P}$.
Conditioning on $D_1$, such that $\widehat \theta_1$ is considered deterministic, we define following events:
align[align omitted — 390 chars of source]
We further observe that
equation[equation omitted — 528 chars of source]
The last event has probability zero under $\mathbb{P}_{P}( \cdot | D_1)$ since $\theta(P)$ is the minimizer of $\theta \mapsto \mathbb{M}(\theta, P)$ and thus $\mathbb{M}(\theta(P), P) \le \mathbb{M}(\theta, P)$ holds $P$-almost surely for any $\theta \in \Theta$. When $U_{n,\alpha_2}(\theta(P), \widehat \theta_1; D_2) < 0$, the probability under $\mathbb{P}_{P}(\cdot | D_1)$ remains zero after taking the intersection with $\mathcal{B}_n$. Hence, we have
align*[align* omitted — 390 chars of source]
Finally, we conclude the claim by the fact that $\mathbb{P}_{P}(\mathcal{A}_n^c\mid D_1) \le \alpha_1$ by (ref) and $\mathbb{P}_{P}(\mathcal{B}_n^c\mid D_1) \le \alpha_2$ by (ref) uniformly for all $P \in \mathcal{P}$.
Proof of Theorem (ref)
Without loss of generality, we set $U_{n,\alpha_2}(\theta, \widehat \theta_1; D_2) = 0$, which only enlarges the confidence set and hence the following result remains to hold. Any element $\theta \in \Theta$ in the confidence set, defined as (ref), satisfies the following:
align*[align* omitted — 500 chars of source]
Now the original confidence set is contained $P$-almost surely in the following supersets:
align*[align* omitted — 935 chars of source]
where
\[\Gamma_n(\widehat\theta_1, \theta(P)) := |\mathbb{P}_n(m_{\widehat{\theta}_1}-m_{\theta(P)}) |+\sqrt{\frac{2\widehat \sigma^2_{\widehat\theta_1, \theta(P)} \log(2/\alpha)}{n}}\]
and the last step uses (ref).
Furthermore, we have
align*[align* omitted — 101 chars of source]
and hence
align*[align* omitted — 179 chars of source]
Using this result, the confidence set can be further contained in
align*[align* omitted — 827 chars of source]
We analyze the superset $\overline{\mathrm{CI}}_{n, \alpha}$. Given $c_0$ and $\beta$ in (ref) and the function $\phi_n$ in (ref), we define $r_n$ as any value that satisfies
align*[align* omitted — 83 chars of source]
We define
align*[align* omitted — 211 chars of source]
where
align*[align* omitted — 134 chars of source]
The constant $M > 0$ will be specified later. We denote by $B_{\|\cdot\|}(\theta; r)$ a $\|\cdot\|$-ball centered at $\theta$ with radius $r \ge 0$.
We consider the partition of the parameter space $\Theta$ into the intersection with the ball
align*[align* omitted — 145 chars of source]
where
align*[align* omitted — 126 chars of source]
Thus far, we have shown that $\widehat{\mathrm{CI}}^{\mathrm{E.B.}}_{n,\alpha} \subseteq \overline{\mathrm{CI}}_{n,\alpha}$. We now consider the following elementary result:
align*[align* omitted — 277 chars of source]
Therefore to claim that, with high probability, the set $\widehat{\mathrm{CI}}^{\mathrm{E.B.}}_{n,\alpha}$ is contained in the ball $B_{\|\cdot\|}(\theta(P); \mathrm{R}_M)$ (and hence the diameter is less than $2\mathrm{R}_M$), we need to establish that $\overline{\mathrm{CI}}_{n,\alpha}$ intersects with $B^c_{\|\cdot\|}(\theta(P); \mathrm{R}_M)$ with an arbitrarily small probability.
Below, we adopt the notation $\mathbb{P}^{1}_P \equiv \mathbb{P}_P(\cdot \mid D_1)$, meaning that probability should be regarded as conditioning on $D_1$ under the distribution $P$. We now establish the existence of $M$ large enough such that for
$\|\theta-\theta(P)\| > \mathrm{R}_M$,
align*[align* omitted — 351 chars of source]
It follows that
align*[align* omitted — 1,042 chars of source]
The last term is trivially controlled since
align[align omitted — 454 chars of source]
by Markov inequality and the fact that $\Gamma_n'(\widehat\theta_1, \theta(P))$ is deterministic under $\mathbb{P}^{1}_P(\cdot)$.
The last display becomes less than $\varepsilon/8$ for $M$ large enough.
Moving onto the term $\mathbf{I}$, we first observe that
align*[align* omitted — 422 chars of source]
We define the “shell":
\[S_j = \{\theta \in \Theta : 2^{j}c_0^{-1/(1+\beta)}r_n^{2/(1+\beta)} \le \|\theta-\theta(P)\| < 2^{j+1}c_0^{-1/(1+\beta)}r_n^{2/(1+\beta)}\}\]
for each $j \in \{0\} \cup \mathbb{N}$. It then follows that
align[align omitted — 1,308 chars of source]
By assumption (ref) that $q < 1+\beta$, the last display becomes less than $\varepsilon/8$ for $M$ large enough, depending on $\beta$ and $q$.
The term $\mathbf{II}$ can similarly be controlled as:
align*[align* omitted — 929 chars of source]
We denote by $\bar{r}_n := (r_n \vee n^{-1/2})$ where $r_n$ is the solution of the equation given by (ref). We now define the “shell":
align*[align* omitted — 260 chars of source]
for each $j \in \{0\} \cup \mathbb{N}$. It follows that
align*[align* omitted — 1,201 chars of source]
Now, we observe that
align*[align* omitted — 242 chars of source]
and the mapping $t\mapsto t^2$ is $2$-Lipschitz on $t \in [-1,1]$. Hence, by the contraction inequality, for instance, Theorem 4.12 of ledoux2013probability or Corollary 3.2.2 of gine2021mathematical, we obtain
align*[align* omitted — 260 chars of source]
To complete the argument, we have
align[align omitted — 398 chars of source]
We now have that
align*[align* omitted — 424 chars of source]
When $r_n \le n^{-1/2}$, we have $\bar{r}_n = n^{-1/2}$ and thus
align*[align* omitted — 295 chars of source]
where the last inequality follows since
$\phi_n(t^{2/(1+\beta)})/t^2$ is strictly decreasing and thus
align*[align* omitted — 138 chars of source]
Hence $ n\phi_n(c_0^{-1/(1+\beta)}n^{-1/(1+\beta)}) \le n^{1/2}$.
On the other hand, when $r_n \ge n^{-1/2}$, we have $\bar{r}_n = r_n$. It further follows that $\bar{r}_n \ge n^{-1/2} \Leftrightarrow n^{1/2}\bar{r}_n \ge 1$, and we can bound $n^{-q/(2+2\beta)} r_n^{-q/(1+\beta)} = (n^{1/2}r_n)^{-q/(1+\beta)}$ by one. The rest of the argument is identical to the earlier derivation and the last summation in (ref) can be bounded as
align*[align* omitted — 281 chars of source]
which can be made smaller than $\varepsilon/8$ for $M$ large enough, depending on $\alpha$, $\beta$ and $B_0$.
Finally, the third term can be bounded as
align*[align* omitted — 491 chars of source]
By repeating the analogous peeling argument over the shell
\[S_j = \{\theta \in \Theta : 2^{j}c_0^{-1/(1+\beta)}r_n^{2/(1+\beta)} \le \|\theta-\theta(P)\| < 2^{j+1}c_0^{-1/(1+\beta)}r_n^{2/(1+\beta)}\}\]
for each $j \in \{0\} \cup \mathbb{N}$, it follows that
align[align omitted — 566 chars of source]
Since $q < 1+\beta$, the last expression can be smaller than $\varepsilon/8$ for large $M$ enough, depending on $\alpha$ and $\beta$.
Finally, let $M$ be a large constant only depending on $\varepsilon$, $\alpha$, $\beta$ and $B_0$ such that (ref),(ref),(ref) and (ref) become smaller than $\varepsilon/8$ respectively. Then with probability at least $1-\varepsilon/2$, it holds that
align*[align* omitted — 526 chars of source]
where the constant $\mathfrak{C}$ depends on $\varepsilon$, $\alpha$, $\beta$ and $B_0$. Since all terms in the last display are non-negative, we have
align*[align* omitted — 355 chars of source]
Furthermore, it follows
align*[align* omitted — 181 chars of source]
since the middle term is always smaller than one of the other two. Putting all results together, we conclude the claim such that \[\widehat{\mathrm{CI}}^{\mathrm{E.B.}}_{n,\alpha} \subseteq B_{\|\cdot\|}(\theta(P); \mathrm{R}_n)\]
with probability greater than $1-\varepsilon/2$
where
align*[align* omitted — 169 chars of source]
This result is stated under the conditional probability on $D_1$. The only term that remains random is $\Gamma^{1/(1+\beta)}_n(\widehat\theta, \theta(P))$. We establish that there exists $S_\varepsilon$ such that
align*[align* omitted — 114 chars of source]
We observe that
align*[align* omitted — 410 chars of source]
The first term can be bounded by the Markov's inequality as
align*[align* omitted — 604 chars of source]
Similarly, we have
align*[align* omitted — 207 chars of source]
Hence by defining $S_\varepsilon = C_\varepsilon s_n$ where $s_n$ is defined as (ref) and $C_\varepsilon$ is a constant, only depending on the $\varepsilon$, it follows that
align*[align* omitted — 275 chars of source]
Choosing $C_\varepsilon$ large enough, we can obtain
align*[align* omitted — 202 chars of source]
Hence, we conclude the claim such that \[\widehat{\mathrm{CI}}^{\mathrm{E.B.}}_{n,\alpha} \subseteq B_{\|\cdot\|}(\theta(P); \mathrm{R}_n)\]
with probability greater than $1-\varepsilon$
where
align*[align* omitted — 136 chars of source]
Proof of Theorem (ref)
We verify the condition (ref), which follows as
align[align omitted — 1,165 chars of source]
The result follows from Theorem (ref).
Proof of Theorem (ref)
The general structure of the proof is identical to the proof of Theorem (ref). By an analogous argument, the CLT-based confidence is contained almost surely by the superset as follows:
align*[align* omitted — 471 chars of source]
where
\[\Gamma_n(\widehat\theta_1, \theta(P)) := |\mathbb{P}_n(m_{\widehat{\theta}_1}-m_{\theta(P)}) |+\sqrt{\frac{z_{\alpha}\widehat \sigma^2_{\widehat\theta_1, \theta(P)} }{n}}.\]
The main difference from the proof of Theorem (ref) is that we can no longer use the contraction inequality to control the squared process. We now define the “shell":
\[S'''_j = \{\theta \in \Theta : 2^{j}c_0^{-1/(1+\beta)}u_n^{2/(1+\beta)} \le \|\theta-\theta(P)\| < 2^{j+1}c_0^{-1/(1+\beta)}u_n^{2/(1+\beta)}\}\]
for each $j \in \{0\} \cup \mathbb{N}$. Then, it follows that
align*[align* omitted — 1,329 chars of source]
By the assumption $q < 1-\beta$, the last series is summable and thus can be made smaller than $\varepsilon$ for $M$ large enough. Then with probability at least $1-\varepsilon/2$, it holds that
align*[align* omitted — 234 chars of source]
where the constant $\mathfrak{C}$ depends on $\varepsilon$, $\alpha$ and $\beta$. We note that the dependence of $\mathfrak{C}$
on $B_0$ is dropped as this was introduced in the proof of Theorem (ref) as a consequence of the contraction inequality. The rest of the proof is identical to that of Theorem (ref).
Proof of Proposition (ref)
Fix $P \in \mathcal{P}$.
Beginning with the statement of Theorem (ref), it is immediate that
align*[align* omitted — 325 chars of source]
where $C$ is a universal constant and $B^2 = \operatorname{\mathbb{E}}_P^1[W_1^2]$ and the second inequality follows from Proposition (ref). We now provide the upper bounds on two terms in the parenthesis.
First, by adding and subtracting terms, we obtain
align*[align* omitted — 605 chars of source]
We invoke technical Lemma (ref) provided below to control the approximation term. The equation (ref) implies that there exists $\varphi$ such that
align*[align* omitted — 195 chars of source]
holds for all $\widehat{\theta}_1$ such that $\|\widehat{\theta}_1-\theta(P)\| < \delta_0$. Hence, by the second statement of Lemma (ref),
align*[align* omitted — 342 chars of source]
assuming $\|\widehat{\theta}_1-\theta(P)\| < \delta$ for any $0 < \delta < \delta_0$. Similarly, by the third statement of Lemma (ref),
align*[align* omitted — 395 chars of source]
We thus obtain for any $0 < \delta < \delta_0$,
align*[align* omitted — 1,077 chars of source]
where we use the trivial bound of $\Delta_{n,P} \le 1$ in the case when $\|\widehat{\theta}_1-\theta(P)\| \ge \delta$. By taking the supremum over $\|t\|=1$, we conclude
align*[align* omitted — 470 chars of source]
for any $0 \le \delta < \delta_0$.
Proof of Proposition (ref)
We fix $P \in \mathcal{P}$ and $t \in \mathbb{R}^d$. We denote by $Z_t:=\langle t, Z\rangle$, which is a real-valued mean-zero random variable and $B_t = (\operatorname{\mathbb{E}}_P[Z_t^2])^{1/2}$. First, we prove $\eqref{eq:unif-integrable-def}\Longrightarrow$ uniform Lindeberg condition. We observe that
align[align omitted — 1,435 chars of source]
for any $\varepsilon > 0$. Hence we have
align*[align* omitted — 363 chars of source]
Assuming uniform integrability under standardization, the last display tends to zero as $\kappa \to \infty$ for fixed $\varepsilon > 0$, and then we take $\varepsilon \to 0$.
Next, we prove $\eqref{eq:norm-equivalence-def} \Longrightarrow \eqref{eq:unif-integrable-def}$. This follows since
align*[align* omitted — 576 chars of source]
as $\kappa \to \infty$. Hence the uniform integrability is implied. These two results establish
align*[align* omitted — 164 chars of source]
In fact, (ref) implies a stronger result than the uniform Lindeberg condition. To see this, we begin from (ref) above and obtain
align*[align* omitted — 400 chars of source]
Hence we have
align*[align* omitted — 210 chars of source]
as $\kappa \to \infty$ under the $L_{2+\delta}$-$L_2$ norm equivalence. Hence not only does the last expression tend to zero, but the rate of convergence is given by $O(\kappa^{-\delta})$.
Statistical Applications
This section contains all proofs associated with the statistical applications. We first provide notations to which we frequently refer. For any set $\Theta$ equipped with a metric $\|\cdot\|$, and any $\varepsilon > 0$, an $\varepsilon$-covering number $\mathcal{N}(\varepsilon,\Theta, \|\cdot\|)$ of $\Theta$ relative to the metric $\|\cdot\|$ is defined as the minimal number of $\|\cdot\|$-balls of radius less than or equal to $\varepsilon$ required for covering $\Theta$. On the other hands, the $\varepsilon$-bracketing number $\mathcal{N}_{[\,]}(\varepsilon,\Theta, \|\cdot\|)$ is the minimal number of “brackets" such that $[L, U] := \{f : L \le f \le U\}$ of size $\|U-L\|\le \varepsilon$ required for covering $\Theta$. In particular, we consider when $\Theta$ contains measurable functions of observations $Z_1\, \ldots, Z_n \in \mathcal{Z}$ and let $Q$ be any discrete probability measure on $Z_1\, \ldots, Z_n$. We define an envelop function $F$ of the class $\Theta$ as $F := z\mapsto \sup_{f\in\Theta}|f(z)|$. The uniform entropy numbers of $\Theta$ relative to $L_r$ is defined as
align*[align* omitted — 141 chars of source]
where $\|f\|_{Q,r} := \left(\sum_{i=1}^n f^r(z_i)Q(z_i)\right)^{1/r}$. Similarly, the bracketing entropy integral is defined as
align*[align* omitted — 146 chars of source]
crucially without taking the supremum over $Q$ and $\varepsilon$ is not normalized by the norm of the envelop function. We use the following result from van2011local:
theorem[Theorem 2.1 of van2011local]
Let $\mathcal{F}$ be a collection of $P$-square integrable functions equipped with an envelop function $F \le 1$. If $\mathbb{E}_P f^2 \le t^2 \mathbb{E}_P F^2$, for every $f$ and some $t \in (0,1)$, then
\begin{align*}
\operatorname{\mathbb{E}}_P\, \left[\sup_{f \in \mathcal{F}}\, |\mathbb{G}_n f|\right] \lesssim J(t, \mathcal{F}, L_2)\left(1+ \frac{J(t, \mathcal{F}, L_2)}{t^2 \sqrt{n}\|F\|_{P,2}}\right) \|F\|_{P, 2}
\end{align*}
where the expectation should be regarded as an outer expectation (Chapter 1.2 of van1996weak) when the content inside is not measurable.
An analogous result under the bracketing entropy integral is also available:
theorem[Theorem 2.14.17' of van1996weak]
Let $\mathcal{F}$ be a collection of $P$-square integrable functions equipped with an envelop function $F \le M$. If $\mathbb{E}_P f^2 \le t^2 \mathbb{E}_P F^2$, for every $f$ and some $t \in (0,1)$, then
\begin{align*}
\operatorname{\mathbb{E}}_P\, \left[\sup_{f \in \mathcal{F}}\, |\mathbb{G}_n f|\right] \lesssim J_{[\, ]}(t, \mathcal{F}, L_2(P))\left(1+ \frac{J_{[\, ]}(t, \mathcal{F}, L_2(P))}{t^2 \sqrt{n}}M\right)
\end{align*}
where the expectation should be regarded as an outer expectation (Chapter 1.2 of van1996weak) when the content inside is not measurable.
High-dimensional mean estimation
proof[{Proof of Theorem (ref)}]
Throughout, we treat $\|\cdot\|\equiv\|\cdot\|_2$. We provide the sufficient condition under which $\Delta_{n,P}$ tends to zero as $n \to \infty$. Observe that
\begin{align*}
W_i &= m_{\widehat\theta_1}-m_{\theta(P)} - P(m_{\widehat\theta_1}-m_{\theta(P)})\\
&= 2(X-\theta(P))^\top(\widehat\theta_1-\theta(P))\\
&= 2t^\top (X-\theta(P))\|\widehat\theta_1-\theta(P)\|\quad for \quad t \in \mathbb{S}^{d-1}
\end{align*}
and $B^2 = \operatorname{\mathbb{E}}_P[W_i^2] = 4\|\widehat\theta_1-\theta(P)\|^2 P(t^\top (X-\theta(P)))^2$. This is an example where the linearization in the form of (ref) holds exactly. In the context of Proposition (ref), we can take $\varphi \equiv 0$ and $\delta \to \infty$. Alternatively, by directly inspecting the upper bound in Lemma (ref), we obtain
\begin{align*}
\Delta_{n,P} &\le C_0 \operatorname{\mathbb{E}}_P\left[\frac{W_1^2}{B^2} \left\{1 \wedge n^{-1/2}\frac{|W_1|}{B}\right\}\bigg|D_1\right] \\
&=C_0 \operatorname{\mathbb{E}}_P\left[\frac{\langle t, X-\theta(P)\rangle^2}{B_t^2} \left\{1 \wedge n^{-1/2}\frac{|\langle t, X-\theta(P)\rangle|}{B_t}\right\}\bigg|D_1\right] \\
&\le C_0 \sup_{t\in\mathbb{S}^{d-1}}\,\operatorname{\mathbb{E}}_P\left[\frac{\langle t, X-\theta(P)\rangle^2}{B_t^2} \left\{1 \wedge n^{-1/2}\frac{|\langle t, X-\theta(P)\rangle|}{B_t}\right\}\right]
\end{align*}
where $B_t = \operatorname{\mathbb{E}}_P\langle t, X-\theta(P)\rangle^2$
for some universal constant $C_0 > 0$. Hence under the uniform Lindeberg condition on $X-\theta(P)$, we obtain $\sup_{P\in\mathcal{P}}\, \Delta_{n,P} \to 0$ as $n \to \infty$. This result does not require the consistency of the initial estimator. We conclude the claim in view of Theorem (ref).
proof[{Proof of Theorem (ref)}]
Throughout, we treat $\|\cdot\|\equiv\|\cdot\|_2$. The proof proceeds by verifying conditions required for Theorem (ref). First, we check (ref). We observe for any $\theta \in \Theta$,
\begin{align*}
m_{\theta}-m_{\theta(P)} := \|X-\theta\|^2 - \|X-\theta(P)\|^2 = 2(X-\theta)^\top(\theta-\theta(P)) + \|\theta-\theta(P)\|^2
\end{align*}
and $P(m_{\theta}-m_{\theta(P)}) = \|\theta-\theta(P)\|^2$. Thus (ref) holds (with an equality) with $\beta=1$ and $c_0=1$. Next, we check (ref). For any $\theta$ such that $\|\theta-\theta(P)\| \le \delta$, we have
\begin{align*}
\sup_{\|\theta-\theta(P)\| \le \delta}\, |m_{\theta}-m_{\theta(P)} - P(m_{\theta}-m_{\theta(P)})| &\le \sup_{\|\theta-\theta(P)\| \le \delta}\, 2|\big(X-\theta(P)\big)^\top (\theta(P)-\theta)| \\
&= 2\delta \sup_{u \in \mathbb{S}^{d-1}}\,|u^\top (X-\theta(P))| \\
&= 2\delta \|X-\theta(P)\|.
\end{align*}
Hence, we obtain
\begin{align*}
\operatorname{\mathbb{E}}_P \left[\sup_{\|\theta-\theta(P)\| < \delta}|\mathbb{G}_n (m_\theta - m_{\theta(P)})| \right]&\le n^{-1/2}\operatorname{\mathbb{E}}_P \left[\left|\sum_{i=1}^n 2\delta \|X_i-\theta(P)\|\right| \right]\\
&\le2\delta\left(\operatorname{\mathbb{E}}_P\|X-\theta(P)\|^2\right)^{1/2}.
\end{align*}
Similarly, we have
\begin{align*}
P(m_{\theta}-m_{\theta(P)})^2 &\le 4\operatorname{\mathbb{E}}_P|(X-\theta)^\top(\theta-\theta(P))|^2 + 2\|\theta-\theta(P)\|^4\\
&\le 4\|\theta-\theta(P)\|^2\operatorname{\mathbb{E}}_P\|X-\theta(P)\|^2 + 2\|\theta-\theta(P)\|^4.
\end{align*}
As a result, we can choose $\phi_n$ in (ref) as
\begin{align*}
&\operatorname{\mathbb{E}}_P \left[\sup_{\|\theta-\theta(P)\| < \delta}|\mathbb{G}_n (m_\theta - m_{\theta(P)})| \right] \vee \sqrt{P(m_{\theta}-m_{\theta(P)})^2}\\
&\qquad \le 2\delta\left(\operatorname{\mathbb{E}}_P\|X-\theta(P)\|^2\right)^{1/2} + \sqrt{2}\delta^2 = \phi_n(\delta).
\end{align*}
For $\omega_n$ in (ref), we invoke Lemma (ref) in Section (ref).
From the earlier derivation, it is immediate that the local envelope can be defined as
\begin{align*}
M_\delta(x)= \sup_{\|\theta-\theta(P)\| \le \delta}\, |(m_{\theta}-m_{\theta(P)})(x)|\le 2\delta \|x-\theta(P)\| + \delta^2.
\end{align*}
Lemma (ref) implies
\begin{align*}
&\operatorname{\mathbb{E}}_P\left[\sup_{\|\theta-\theta(P)\| < \delta}\, |\mathbb{G}_n(m_{\theta}-m_{\theta(P)} )^2|\right] \\
&\qquad \le C\left\{ n^{1/2} \left(\mathbb{E}_P M_\delta^2 \right)+ \left(\mathbb{E}_PM_\delta^2\right)^{1/2}\operatorname{\mathbb{E}}_P \left[\sup_{\|\theta-\theta(P)\| < \delta}\, \left|\sum_{i=1}^n \varepsilon_i (m_{\theta}-m_{\theta(P)})\right|\right]\right\}\\
&\qquad \le C\left\{ n^{1/2} \left(\mathbb{E}_P M_\delta^2 \right)+ \left(n\mathbb{E}_PM_\delta^2\right)^{1/2}\delta\left(\operatorname{\mathbb{E}}_P\|X-\theta(P)\|^2\right)^{1/2}\right\}\\
&\qquad \le Cn^{1/2} \left(\delta^2 \operatorname{\mathbb{E}}_P\|X-\theta(P)\|^2 + \delta^4\right).
\end{align*}
As a result, we can choose $\omega_n$ in (ref) as
\begin{align*}
\omega_n(\delta) = \sqrt{Cn^{1/2} \left(\delta^2 \operatorname{\mathbb{E}}_P\|X-\theta(P)\|^2 + \delta^4\right)}.
\end{align*}
It now remains the solve the inequalities in Theorem (ref). First,
\begin{align*}
2r_n^{-1}\left(\operatorname{\mathbb{E}}_P\|X-\theta(P)\|^2\right)^{1/2} \le n^{1/2} \implies r_n \ge 2\left(\frac{\operatorname{\mathbb{E}}_P\|X-\theta(P)\|^2}{n}\right)^{1/2} = 2\left(\frac{\mathrm{tr}(\Sigma)}{n}\right)^{1/2}.
\end{align*}
When $r_n \le 1$, this term dominates (up to a constant) and thus the inequality in (ref) is satisfied as long as $\mathrm{tr}(\Sigma) \le n$. Similarly, we arrive at
\begin{align*}
&u_n^{-2}\sqrt{Cn^{1/2} \left(u_n^2 \operatorname{\mathbb{E}}_P\|X-\theta(P)\|^2 + u_n^4\right)} \le n^{3/4} \\
&\qquad\implies u_n \ge \left(\frac{C\operatorname{\mathbb{E}}_P\|X-\theta(P)\|^2}{n}\right)^{1/2} = \left(\frac{C\mathrm{tr}(\Sigma)}{n}\right)^{1/2}
\end{align*}
assuming $\mathrm{tr}(\Sigma) \le n$. Thus by Theorem (ref),
\begin{align*}
&r_n^{2/(1+\beta)} + u_n^{2/(1+\beta)} +s_n^{1/(1+\beta)} = C\left\{\left(\frac{\mathrm{tr}(\Sigma) }{n}\right)^{1/2} +s_n^{1/2}\right\}.
\end{align*}
This concludes the first result.
Next, we assess the requirement for $s_n$ as defined in (ref). First, for fixed $\widehat \theta_1$, we have
\begin{align*}
\left(P(m_{\widehat\theta_1}-m_{\theta(P)})^2\right)^{1/2} &\le 2\left(\mathrm{tr}(\Sigma)\|\widehat\theta_1-\theta(P)\|^2\right)^{1/2} + \left(2\|\widehat\theta_1-\theta(P)\|^4\right)^{1/2}.
\end{align*}
Assuming that $\mathbb{P}_P(\|\widehat\theta_1-\theta(P)\|^2 \ge C_\varepsilon \mathrm{tr}(\Sigma)/n) \le \varepsilon$, we choose $C_\varepsilon$ such that an event $\Omega := \{\|\widehat\theta_1-\theta(P)\|^2 < C_\varepsilon \mathrm{tr}(\Sigma)/n\}$ holds with probability greater than $1-\varepsilon/2$. Then, conditioning on $\Omega$, we have
\begin{align*}
\operatorname{\mathbb{E}}_P\|X-\widehat\theta_1\|^2 - \operatorname{\mathbb{E}}_P\|X-\theta(P)\|^2 = \|\widehat\theta_1-\theta(P)\|^2 \le \frac{C_1\mathrm{tr}(\Sigma)}{n},
\end{align*}
and this implies
\begin{align*}
s_n^{1/2} \le 2\left(\frac{\mathrm{tr}(\Sigma)\|\widehat\theta_1-\theta(P)\|^2}{n}\right)^{1/4} + \left(\frac{2\|\widehat\theta_1-\theta(P)\|^4}{n}\right)^{1/4} + \left(\frac{C_1\mathrm{tr}(\Sigma)}{n}\right)^{1/2}
\end{align*}
with probability greater than $1-\varepsilon/2$. Thus as long as the initial estimation satisfies
\begin{align*}
\|\widehat\theta_1-\theta(P)\|^2 \lesssim \frac{\mathrm{tr}(\Sigma)}{n} \quad and \quad \|\widehat\theta_1-\theta(P)\|^4 \lesssim \frac{\mathrm{tr}^2(\Sigma)}{n},
\end{align*}
in high probability, we can claim $ s_n^{1/2} =O_P( \sqrt{\mathrm{tr}(\Sigma)/n})$. In particular, the requirement of $\|\widehat\theta_1-\theta(P)\|^2 \lesssim \mathrm{tr}(\Sigma)/n$ in high probability suffices. To see this, we obtain the following from the first result of this theorem:
\begin{align*}
\mathbb{P}_P\left(\mathrm{Diam}_{\|\cdot\|}\big(\widehat{\mathrm{CI}}_{n,\alpha}^{\mathrm{mean}}\big)\le \mathfrak{C}\left\{\left(\frac{\mathrm{tr}(\Sigma)}{n}\right)^{1/2} + s_n^{1/2}\right\}\right) \ge 1-\varepsilon/2.
\end{align*}
Then, it follows that
\begin{align*}
&\mathbb{P}_P\left(\mathrm{Diam}_{\|\cdot\|}\big(\widehat{\mathrm{CI}}_{n,\alpha}^{\mathrm{mean}}\big)\le \mathfrak{C}\left\{\left(\frac{\mathrm{tr}(\Sigma)}{n}\right)^{1/2} + s_n^{1/2}\right\}\right) \\
&\qquad \le \mathbb{P}_P\left(\mathrm{Diam}_{\|\cdot\|}\big(\widehat{\mathrm{CI}}_{n,\alpha}^{\mathrm{mean}}\big)\le \mathfrak{C}\left\{\left(\frac{\mathrm{tr}(\Sigma)}{n}\right)^{1/2} + s_n^{1/2}\right\}\bigg| \Omega\right)\mathbb{P}_P(\Omega) + \mathbb{P}_P(\Omega^c) \\
&\qquad \le \mathbb{P}_P\left(\mathrm{Diam}_{\|\cdot\|}\big(\widehat{\mathrm{CI}}_{n,\alpha}^{\mathrm{mean}}\big)\le \mathfrak{C}\left\{\left(\frac{\mathrm{tr}(\Sigma)}{n}\right)^{1/2} + s_n^{1/2}\right\}\bigg| \Omega\right) + \frac{\varepsilon}{2}.
\end{align*}
On the event $\Omega$ for $n \ge C_\varepsilon$, it is implies that
\begin{align*}
\|\widehat\theta_1-\theta(P)\|^4 \le \mathrm{tr}(\Sigma)\|\widehat\theta_1-\theta(P)\|^2 \le \frac{C_\varepsilon \mathrm{tr}^2(\Sigma)}{n}
\end{align*}
hence the required probabilistic upper bound on $\|\widehat\theta_1-\theta(P)\|^4$ is implied. We thus conclude
\begin{align*}
\mathbb{P}_P\left(\mathrm{Diam}_{\|\cdot\|}\big(\widehat{\mathrm{CI}}_{n,\alpha}^{\mathrm{mean}}\big)\le \mathfrak{C}\left(\frac{\mathrm{tr}(\Sigma)}{n}\right)^{1/2}\right) \ge 1-\varepsilon
\end{align*}
when the requirements for $\widehat \theta_1$ are satisfied.
Misspecified linear regression
proof[{Proof of Theorem (ref)}]
Throughout, we treat $\|\cdot\|\equiv\|\cdot\|_2$. First, observe that
\begin{align*}
W_i&=m_{\theta_P} - m_{\widehat \theta_1} - P(m_{\theta_P} - m_{\widehat \theta_1})\\
& = (Y_i-\theta_P^\top X_i)^2-(Y_i-\widehat \theta_1^\top X_i)^2-P(m_{\theta_P}-m_{\widehat\theta_1})\\
& = -2(Y_i-\theta_P^\top X)(\theta_P - \widehat\theta_1)^\top X-\{(\theta_P-\widehat \theta_1)^\top X_i\}^2+\operatorname{\mathbb{E}}_P \{X_i^\top(\theta_P - \widehat \theta_1)\}^2.
\end{align*}
Taking $H := -2X(Y-\theta_P^\top X)$, the equation (ref) corresponds to
\begin{align*}
&\frac{\operatorname{\mathbb{E}}_P[m_{\widehat\theta_1} - m_{\theta_P} -P(m_{\widehat\theta_1}-m_{\theta_P})- \langle \widehat\theta_1 - \theta_P, H \rangle ]^2}{\operatorname{\mathbb{E}}_P \langle \widehat\theta_1-\theta_P, H \rangle^2 } \\
&\qquad \le \frac{\operatorname{\mathbb{E}}_P\big[\{(\theta_P- \widehat\theta_1)^\top X\}^2+\operatorname{\mathbb{E}}_P [\{X^\top(\theta_P - \widehat\theta_1)\}^2] \big]^2}{4\operatorname{\mathbb{E}}_P[\operatorname{\mathbb{E}}_P[\xi_i^2|X_i](X^\top(\theta_P - \widehat\theta_1))^2]} \\
&\qquad \le \frac{\operatorname{\mathbb{E}}_P[\{(\theta_P- \widehat\theta_1)^\top X\}^4]}{\underbar{$\sigma$}^2\operatorname{\mathbb{E}}_P[\{(\theta_P- \widehat\theta_1)^\top X\}^2]}\le \underbar{$\sigma$}^{-2}L^4 \|\theta_P- \widehat\theta_1\|^2\lambda_{\max}(\Gamma_P)
\end{align*}
where the last inequality uses (ref). With the choice
\begin{align*}
\varphi(\|\theta_P-\widehat\theta_1\|) = \underbar{$\sigma$}^{-2}L^4 \lambda_{\max}(\Gamma_P)\|\widehat\theta_1-\theta_P\|^2,
\end{align*}
Proposition (ref) holds with $H :=-2X\xi$, which states that
\begin{align*}
\sup_{P\in\mathcal{P}}\, \Delta_{n,P} &\lesssim \underbar{$\sigma$}^{-2}L^4 \lambda_{\max}(\Gamma_P)\varepsilon^2 +\underbar{$\sigma$}^{-1}L^2 \lambda^{1/2}_{\max}(\Gamma_P)\varepsilon + \sup_{P\in\mathcal{P}}\,\mathbb{P}_P(\|\widehat\theta_1 - \theta_P\| > \varepsilon) \\
&\qquad + \sup_{P\in\mathcal{P}}\, \sup_{\|t\|=1}\,\operatorname{\mathbb{E}}_P\left[\frac{\langle t, 2X\xi \rangle^2}{\operatorname{\mathbb{E}}_P\langle t, 2X\xi \rangle^2} \left\{1 \wedge n^{-1/2}\frac{|\langle t, 2X\xi \rangle|}{\operatorname{\mathbb{E}}_P\langle t, 2X\xi \rangle}\right\}\right].
\end{align*}
Thus under the uniform Lindeberg condition on $X\xi$ and the assumptions stated in the theorem, we conclude $ \sup_{P\in\mathcal{P}}\, \Delta_{n,P} \to 0$ by taking $n \to \infty$ and then $\varepsilon \to 0$. We conclude the claim in view of Theorem (ref).
proof[{Proof of Theorem (ref)}]
We verify conditions required for Theorem (ref). First, we check (ref). We observe that
\begin{align*}
m_{\theta}-m_{\theta_P} = (Y-\theta^\top X)^2 - (Y-\theta_P^\top X)^2 &= 2(Y-\theta_P^\top X)(\theta_P^\top X-\theta^\top X) + (\theta_P^\top X-\theta^\top X)^2\\
&= 2\xi X^{\top}(\theta_P - \theta) + ((\theta_P-\theta)^{\top}X)^2.
\end{align*}
Taking the expectation under $P$ and the fact that $\operatorname{\mathbb{E}}_P(Y-\theta_P^\top X)X = \operatorname{\mathbb{E}}_P[\xi X] = 0$, we obtain
\begin{align*}
P(m_{\theta}-m_{\theta_P}) = \operatorname{\mathbb{E}}_P (\theta_P^\top X-\theta^\top X)^2 = \|\theta_P-\theta\|^2_{\Gamma_P}
\end{align*}
which implies that (ref) is satisfied with $\beta=1$ and $c_0 = 1$.
Next, we check (ref). First, we note that
\begin{align*}
&\sup_{\|\theta-\theta_P\|_{\Gamma_P}\le \delta} |(Y-\theta^\top X)^2 - (Y-\theta_P^\top X)^2| \\
&\qquad \le \sup_{\|\theta-\theta_P\|_{\Gamma_P}\le \delta}\left\{2|\xi(\theta_P^\top X-\theta^\top X) | + |(\theta_P^\top X-\theta^\top X)^2|\right\}\\
&\qquad \le \sup_{\|\theta-\theta_P\|_{\Gamma_P}\le \delta}\left\{2|(\theta_P-\theta)^\top \Gamma_P^{1/2}\Gamma_P^{-1/2}\xi X |\right\} + \sup_{\|\theta-\theta_P\|_{\Gamma_P}\le \delta}\left\{|(\theta_P-\theta)^\top \Gamma_P^{1/2}\Gamma_P^{-1/2}X|^2\right\}\\
&\qquad = 2\delta\|\Gamma_P^{-1/2}\xi X\|+ \delta^2 \|\Gamma_P^{-1/2}XX^\top\Gamma_P^{-1/2}\|_{\mathrm{op}}.
\end{align*}
Here, $\|\cdot\|_{\mathrm{op}}$ denotes the operator norm. Note that the last expression can be taken as our local envelope function. It then follows that
\begin{align*}
&\operatorname{\mathbb{E}}_P \left[\sup_{\|\theta-\theta_P\|_{\Gamma_P}< \delta}|\mathbb{G}_n (m_\theta - m_{\theta_P})| \right]
\\
&\qquad\le 2\operatorname{\mathbb{E}}_P\left[\sup_{\|\theta-\theta_P\|_{\Gamma_P} < \delta}|\mathbb{G}_n[\xi X]^{\top}\Gamma_P^{-1/2}\Gamma_P^{1/2}(\theta - \theta_P)|\right]\\
&\qquad\qquad+ \operatorname{\mathbb{E}}_P\left[\sup_{\|\theta - \theta_P\|_{\Gamma_P} < \delta}|(\theta - \theta_P)^{\top}\Gamma_P^{1/2}\Gamma_P^{-1/2}\mathbb{G}_n[XX^{\top}]\Gamma_P^{1/2}\Gamma_P^{-1/2}(\theta - \theta_P)|\right]\\
&\qquad= 2\delta\operatorname{\mathbb{E}}_P[\|\mathbb{G}_n[\Gamma_P^{-1/2}\xi X]\|] + \delta^2\operatorname{\mathbb{E}}_P[\|\Gamma_P^{-1/2}\mathbb{G}_n[XX^{\top}]\Gamma_P^{-1/2}\|_{\mathrm{op}}].
\end{align*}
Now observe that
\begin{align*}
\mathbb{E}_P[\|\mathbb{G}_n[\Gamma_P^{-1/2}\xi X]\|] &\le (\mathbb{E}_P[\|\mathbb{G}_n[\xi \Gamma_P^{-1/2}X]\|^2])^{1/2} \\
&\le (\mathbb{E}_P[|\xi|^2\|\Gamma_P^{-1/2}X\|^2])^{1/2} = \overline{\sigma}(tr(\Gamma_P^{-1/2}\Gamma_P\Gamma_P^{-1/2}))^{1/2}=\overline{\sigma}d^{1/2}.
\end{align*}
From Theorem I of tropp2016expected, we get
\begin{equation}
\mathbb{E}_P[\|\mathbb{G}_n[\Gamma_P^{-1/2}XX^{\top}\Gamma_P^{-1/2}]\|_{\mathrm{op}}] \le C\sqrt{\log(d)v(X)} + C\frac{\log(d)}{n^{1/2}}\left(\mathbb{E}_P[\max_{1\le i\le n}\|\Gamma_P^{-1/2}X_i\|^4]\right)^{1/2}
\end{equation}
for some universal constant $C$. Here
\[
v(X) = \left\|\frac{1}{n}\sum_{i=1}^n \mathbb{E}_P\left[(\Gamma_P^{-1/2}X_iX_i^{\top}\Gamma_P^{-1/2} - I_d)(\Gamma_P^{-1/2}X_iX_i^{\top}\Gamma_P^{-1/2}-I_d)\right]\right\|_{\mathrm{op}}.
\]
We bound two terms of the upper bound in (ref).
For the second term, we observe
\begin{align*}\left(\mathbb{E}_P[\max_{1\le i\le n}\|\Gamma_P^{-1/2}X_i\|^4]\right)^{1/2} &\le \left(\mathbb{E}_P\left[\max_{1\le i\le n}\|\Gamma_P^{-1/2}X_i\|^{q_x}\right]\right)^{2/q_x} \le n^{2/q_x}\left(\mathbb{E}_P[\|\Gamma_P^{-1/2}X_i\|^{q_x}]\right)^{2/q_x}\\
&= n^{2/q_x}\left(\mathbb{E}_P[\sum_{j=1}^d (e_j^\top \Gamma_P^{-1/2}X_i)^{q_x}]\right)^{2/q_x} \le n^{2/q_x}d^{2/q_x}L^2.
\end{align*}
where the last step follows by (ref). Next, we observe
\begin{align*}
&\mathbb{E}_P\left[(\Gamma_P^{-1/2}X_iX_i^{\top}\Gamma_P^{-1/2} - I_d)(\Gamma_P^{-1/2}X_iX_i^{\top}\Gamma_P^{-1/2}-I_d)\right] \\
&\qquad = \mathbb{E}_P\left[\Gamma_P^{-1/2}X_iX_i^{\top}\Gamma_P^{-1/2}\Gamma_P^{-1/2}X_iX_i^{\top}\Gamma_P^{-1/2} - 2\Gamma_P^{-1/2}X_iX_i^{\top}\Gamma_P^{-1/2} + I_d\right]\\
&\qquad = \mathbb{E}_P\left[\|\Gamma_P^{-1/2}X_i\|^2\Gamma_P^{-1/2}X_iX_i^{\top} \Gamma_P^{-1/2}- I_d\right].
\end{align*}
The operator norm can be controlled as
\begin{align*}
&\left\|\frac{1}{n}\sum_{i=1}^n \mathbb{E}\left[(\Gamma_P^{-1/2}X_iX_i^{\top}\Gamma_P^{-1/2} - I_d)(\Gamma_P^{-1/2}X_iX_i^{\top}\Gamma_P^{-1/2}-I_d)\right]\right\|_{\mathrm{op}}\\
&\qquad \le \frac{1}{n}\sum_{i=1}^n \left\|\mathbb{E}\left[(\Gamma_P^{-1/2}X_iX_i^{\top}\Gamma_P^{-1/2} - I_d)(\Gamma_P^{-1/2}X_iX_i^{\top}\Gamma_P^{-1/2}-I_d)\right]\right\|_{\mathrm{op}}\\
&\qquad= \frac{1}{n}\sum_{i=1}^n \sup_{\|u\|=1}| \mathbb{E}\left[\|\Gamma_P^{-1/2}X_i\|^2|u^\top \Gamma_P^{-1/2}X_i|^2\right] - u^\top I_du|\\
&\qquad\le \frac{1}{n}\sum_{i=1}^n \sup_{\|u\|=1}\sqrt{ \mathbb{E}\|\Gamma_P^{-1/2}X\|^4\mathbb{E}|u^\top \Gamma_P^{-1/2}X_i|^4}\le dL^4,
\end{align*}
where the last step follows by (ref) since $q_x \ge 4$. We thus obtain
\begin{align*}
\mathbb{E}_P[\|\mathbb{G}_n[\Gamma_P^{-1/2}XX^{\top}\Gamma_P^{-1/2}]\|_{\mathrm{op}}]&\le C\sqrt{d \log d}L^2 + C\frac{d^{2/q_x} \log d}{n^{1/2 - 2/q_x}}L^2.
\end{align*}
From the earlier derivation, it also follows that
\begin{align*}
P(m_{\theta}-m_{\theta_P})^2 &\le 4\operatorname{\mathbb{E}}_P|\xi(\theta_P^\top X-\theta^\top X) |^2 + 2 \operatorname{\mathbb{E}}_P|(\theta_P-\theta)^\top X|^4 \\
&\le 4\|\theta_P-\theta\|_{\Gamma_P}^2\sup_{u\in\mathbb{S}^{d-1}}\,\operatorname{\mathbb{E}}_P |u^\top \xi \Gamma_P^{-1/2}X|^2 + 2\|\theta_P-\theta\|_{\Gamma_P}^4\sup_{u\in\mathbb{S}^{d-1}}\, \operatorname{\mathbb{E}}_P (u^\top \Gamma_P^{-1/2}X)^4 \\
&\le 4\|\theta_P-\theta\|_{\Gamma_P}^2\operatorname{\mathbb{E}}_P \|\xi\Gamma_P^{-1/2} X\|^2 + 2\|\theta_P-\theta\|_{\Gamma_P}^4 L^4.
\end{align*}
Hence, we can take
\begin{align*}
\phi(\delta) = C\overline{\sigma}d^{1/2} \delta+ CL^2\delta^2\sqrt{d\log d}\left[1 + \frac{d^{2/q_x-1/2}\sqrt{ \log d}}{n^{1/2 - 2/q_x}}\right]
\end{align*}
in (ref). The first term dominates when
\begin{align*}
&\overline{\sigma}d^{1/2} \delta \ge \sqrt{d\log d}L^2\delta^2\left[1 + \sqrt{\frac{\log d}{(nd)^{1 - 4/q_x}}}\right]\&\qquad\Rightarrow \delta \le \left(\frac{\overline{\sigma}d^{1/2}}{L^2}\right)\left(\frac{1}{\sqrt{d\log d}(1 + \sqrt{\log d/(nd)^{1 - 4/q_x}})}\right)\&\qquad\Rightarrow \delta \le \frac{\overline{\sigma}}{L^2}\left(\frac{(nd)^{1/2 - 2/q_x}}{\log d}\right) = \mathfrak{C}_P\left(\frac{(nd)^{1/2 - 2/q_x}}{\log d}\right)
\end{align*}
where $\mathfrak{C}_P = \overline{\sigma}/L^2$. We verify this condition at the end for the choice of $\delta$. Provided that $\delta$ satisfies this inequality, we can take $\phi(\delta) = C\overline{\sigma}d^{1/2} \delta$.
Finally, we verify (ref) by invoking Lemma (ref). The local envelope is defined as
\begin{align*}
M_\delta(x)= \sup_{\|\theta-\theta(P)\|_{\Gamma_P} \le \delta}\, |(m_{\theta}-m_{\theta(P)})(x)|\le 2\delta\|\Gamma_P^{-1/2}\xi X\|+ \delta^2 \|\Gamma_P^{-1/2}XX^\top \Gamma_P^{-1/2}\|_{\mathrm{op}},
\end{align*}
and
\begin{align*}
\operatorname{\mathbb{E}}_PM_\delta^2 \le 4\delta^2\operatorname{\mathbb{E}}_P\|\Gamma_P^{-1/2}\xi X\|^2+ 2\delta^4 \operatorname{\mathbb{E}}_P\|\Gamma_P^{-1/2}XX^\top\Gamma_P^{-1/2}\|^2_{\mathrm{op}}.
\end{align*}
From assumption (ref) and the earlier derivation,
\[
\operatorname{\mathbb{E}}_P\|\Gamma_P^{-1/2}XX^\top\Gamma_P^{-1/2}\|^2_{\mathrm{op}}\le C\left(d \log dL^4 + d^{4/q_x} (\log d)^2L^4\right)
\]
By Lemma (ref),
\begin{align*}
\omega^2_n(\delta)& \le Cn^{1/2} \left\{ \delta^2\overline{\sigma}^2d+\delta^4(\log d)^2dL^4+\delta\overline{\sigma}d^{1/2}\phi(\delta) +\delta^2(\log d)d^{1/2}L^2\phi(\delta)\right\}\\
& \le Cn^{1/2} \left\{ \delta^2\overline{\sigma}^2d+\delta\overline{\sigma}d^{1/2}\phi(\delta) \right\}\\
& \le 2Cn^{1/2} \delta^2\overline{\sigma}^2d
\end{align*}
where the second and third lines follow provided that $\mathfrak{C}_P (nd)^{1/2-2/q_x}\log^{-1}(d)\ge\delta$.
It thus remains to solve equations (ref) to derive the convergence rate. First, recalling that $\beta=1$ and $c_0 = 1$, we have
\begin{align*}
&C\overline{\sigma}d^{1/2} r_n^{-1} \le n^{1/2}
\Longleftrightarrow r_n \ge C\overline{\sigma}\left(d/n\right)^{1/2}
\end{align*}
and
\begin{align*}
&C\overline{\sigma}d^{1/2} u_n^{-1} \le n^{1/2}
\Longleftrightarrow u_n \ge C\overline{\sigma}\left(d/n\right)^{1/2}
\end{align*}
Finally, we verify $r_n, u_n \le \mathfrak{C}_P (nd)^{1/2-2/q_x}\log^{-1}(d)$. This is satisfied when
\begin{align*}
&\overline{\sigma}\left(d/n\right)^{1/2} \le \mathfrak{C}_P (nd)^{1/2-2/q_x}\log^{-1}(d)\Rightarrow \left(L^2d^{2/q_x}\log(d)\right)^{q_x/(q_x-2)}\le n.
\end{align*}
Thus the inequality in (ref) is satisfied when $n \ge \left(L^2d^{2/q_x}\log(d)\right)^{q_x/(q_x-2)}$. We note from the fact that $q_x \ge 4$, we can derive the condition under the least-favorable case, which is $n \ge L^2d (\log d)^2$.
Finally by Theorem (ref),
\begin{align*}
&c_0^{-1/(1+\beta)}\left(r_n^{2/(1+\beta)} + u_n^{2/(1+\beta)}+s_n^{1/(1+\beta)}\right)= C\left\{\left(\frac{\overline{\sigma}^2 d}{n}\right)^{1/2}+s_n^{1/2}\right\}
\end{align*}
for some universal constant $C$. This concludes the first result.
We consider the requirement for $s_n$ defined as (ref). First, for fixed $\widehat\theta_1$,
\begin{align*}
P(m_{\widehat\theta_1}-m_{\theta_P})^2 &\lesssim \|\widehat\theta_1-\theta_P\|^2_{\Gamma_P}\operatorname{\mathbb{E}}_P \|\xi \Gamma_P^{-1/2} X\|^2 +L^4 \|\widehat\theta_1-\theta_P\|^4_{\Gamma_P}.
\end{align*}
We also have
\begin{align*}
\operatorname{\mathbb{E}}_P(Y-\widehat\theta_1^\top X)^2 - \operatorname{\mathbb{E}}_P(Y-\theta_P^\top X)^2 &= \|\widehat\theta_1-\theta(P)\|^2_{\Gamma_P}.
\end{align*}
The remaining argument is analogous to the proof of Theorem (ref).
\begin{align*}
&n^{-1/2}\sqrt{P (m_{\widehat\theta_1}-m_{\theta(P)})^2} + \mathbb{C}_P(\widehat{\theta}_1) \\
&\qquad \le n^{-1/2} \|\widehat\theta_1-\theta_P\|_{\Gamma_P}\big(\operatorname{\mathbb{E}}_P \|\xi \Gamma_P^{-1/2}X\|^2\big)^{1/2} + (n^{-1/2}L^2+1)\|\widehat\theta_1-\theta_P\|_{\Gamma_P}^2 \\
&\qquad \le n^{-1/2} \overline{\sigma}d^{1/2}\|\widehat\theta_1-\theta_P\|_{\Gamma_P} + 2L^2\|\widehat\theta_1-\theta_P\|_{\Gamma_P}^2 := s_n.
\end{align*}
Conditioning on the event where $\|\widehat\theta_1 - \theta_P\|_{\Gamma_P}^2 \le C_\varepsilon\overline{\sigma}^2d/n$, we obtain
\begin{align*}
s_n &=\frac{ \overline{\sigma}d^{1/2}\|\widehat\theta_1-\theta_P\|_{\Gamma_P} }{n^{1/2}} + 2L^2\|\widehat\theta_1-\theta_P\|_{\Gamma_P}^2\\
&\le\frac{ C_\varepsilon^{1/2} \overline{\sigma}^2d }{n} + 2L^2\frac{ C_\varepsilon \overline{\sigma}^2d }{n}
\end{align*}
Then for $C_\varepsilon$ large enough, we conclude
\begin{align*}
\mathbb{P}_P\left(\mathrm{Diam}_{\|\cdot\|}\big(\widehat{\mathrm{CI}}_{n,\alpha}^{\mathrm{LR}}\big)\le \mathfrak{C}\left(\frac{L^2\overline{\sigma}^2d}{n}\right)^{1/2}\right) \ge 1-\varepsilon
\end{align*}
when the requirements for $\widehat \theta_1$ are satisfied.
Manski's discrete choice model
In this example, we take
\[m_{\theta} := (y,x) \mapsto -\frac{1}{2}\, y \cdot \mathrm{sgn}(\theta^\top x).\]
whose envelop function for $m_{\theta}-m_{\theta_P}$ is trivially given by $M=1$. The leading constant $1/2$ is introduced without loss of generality.
proof[{Proof of Theorem (ref)}]
For any $\theta_1, \theta_2 \in \mathbb{S}^{d-1}$, we define our pseudo-metric as
\begin{align*}
d_\Delta(\theta_1, \theta_2) := \mathbb{P}_P(\mathrm{sgn}(\theta_1^\top X) \neq \mathrm{sgn}(\theta_2^\top X)).
\end{align*}
Now we consider the following collection of “localized” functions:
\begin{align*}
\mathcal{M}_{\delta}^\Delta := \left\{ \frac{1}{2}\, y \left(\mathrm{sgn}(\theta_P^\top x)-\mathrm{sgn}(\theta^\top x)\right) for all \theta i.e., d_\Delta(\theta, \theta_P) \le \delta and \theta \in \mathbb{S}^{d-1}\right\}.
\end{align*}
Below, we provide the confidence width in terms of $d_\Delta(\theta_1, \theta_2)$ using Theorem (ref). First, we check (ref). Following Proposition 1 of tsybakov2004optimal (and similarly for Proposition 2.4 of mukherjee2021optimal), we define the set $\mathcal{A}(\theta) := \{x : \mathrm{sgn}(\theta_P^\top X) \neq \mathrm{sgn}(\theta^\top X)\}$ for each $\theta$. It then follows that
\begin{align*}
P(m_{\theta}-m_{\theta_P}) &=\frac{1}{2}\operatorname{\mathbb{E}}_P\left[ Y \left(\mathrm{sgn}(\theta_P^\top X)-\mathrm{sgn}(\theta^\top X)\right)\right] \\
&=\int_{\mathcal{A}(\theta)}\, |\operatorname{\mathbb{E}}_P[Y |X=x] |P_X(x)\, dx\\
&\ge 2 \int_{\mathcal{A}(\theta)}\, |\eta_P(x)-1/2|P_X(x)\, dx \\
&\ge 2\sup_{0 \le t \le t^*}t\mathbb{P}_P\left(|\eta_P(X)-1/2| \ge t \cap X \in \mathcal{A}(\theta)\right)\\
&\ge 2\sup_{0 \le t \le t^*}t\left(d_\Delta(\theta, \theta_P)-\mathbb{P}_P\left(|\eta_P(X)-0.5| \le t \right)\right)\\
&\ge 2\sup_{0 \le t \le t^*}t\left(d_\Delta(\theta, \theta_P)- C_0t^{1/\beta} \right).
\end{align*}
The last inequality uses (ref). The optimal choice of $t$ is given by
\begin{align*}
t = \begin{cases}
(1+1/\beta)^{-\beta}C_0^{-\beta}d^\beta_\Delta(\theta, \theta_P) & when\quad d_\Delta(\theta, \theta_P)\le (1+1/\beta)C_0(t^*)^{1/\beta}\\
t^* & otherwise.
\end{cases}
\end{align*}
Putting together, it follows for all $\theta \in \mathbb{S}^{d-1}$,
\begin{align*}
P(m_{\theta}-m_{\theta_P}) &\ge \left(\frac{2}{1+\beta}\right)\frac{d^{1+\beta}_\Delta(\theta, \theta_P)}{ (1+1/\beta)^{\beta}C_0^\beta}\mathbf{1}\left\{d_\Delta(\theta, \theta_P)\le (1+1/\beta)C_0(t^*)^{1/\beta}\right\}\\
&\qquad + \left(\frac{2}{1+\beta}\right)t^*d_\Delta(\theta, \theta_P)\mathbf{1}\left\{d_\Delta(\theta, \theta_P)>(1+1/\beta)C_0(t^*)^{1/\beta}\right\}.
\end{align*}
Furthermore, assuming $C_0(t^*)^{1/\beta} > 1$, we conclude
\begin{align*}
P(m_{\theta}-m_{\theta_P})\ge \mathfrak{C} d_\Delta(\theta, \theta_P)^{1+\beta}
\end{align*}
for all $\theta\in \mathbb{S}^{d-1}$ and $\mathfrak{C}$ depending on $\beta$ and $C_0$. Next, we derive $\phi_n$ in (ref). We observe that
\begin{align*}
&\operatorname{\mathbb{E}}_P \left[\sup_{d_\Delta(\theta, \theta_P) < \delta}|\mathbb{G}_n (m_\theta - m_{\theta_P})| \right] \\
&\qquad = \operatorname{\mathbb{E}}_P\left[\sup_{d_\Delta(\theta, \theta_P) < \delta}\left|\mathbb{G}_n \frac{ y \left(\mathrm{sgn}(\theta_P^\top x)-\mathrm{sgn}(\theta^\top x)\right)}{2}\right| \right]= \operatorname{\mathbb{E}}_P\left[\sup_{m \in \mathcal{M}^\Delta_\delta}\left|\mathbb{G}_n m\right| \right],
\end{align*}
which we employ Theorem (ref) to control. First, we relate the covering number of $\mathcal{M}^\Delta_\delta$ to the VC dimension of the subgraphs of the functions in $\mathcal{M}^\Delta_\delta$. We observe that the subgraph of a function $(x,y) \mapsto y\cdot \{\mathrm{sgn}(\theta_P^\top x)-\mathrm{sgn}(\theta^\top x)\}$ for $y \in \{-1,1\}$ is contained in
\begin{align*}
&\left\{(y,x,t) \, :\, y\{\mathrm{sgn}(\theta_P^\top x)-\mathrm{sgn}(\theta^\top x)\} \ge t\right\} \\
&\qquad = \{(1, x,t) \, :\,\mathrm{sgn}(\theta_P^\top x)-\mathrm{sgn}(\theta^\top x) \ge t\} \cup \{(-1,x,t) \, :\,\mathrm{sgn}(\theta^\top x)-\mathrm{sgn}(\theta_P^\top x)\ge t\} \\
&\qquad= \{(1,x,t) \, :\,-\mathrm{sgn}(\theta^\top x) \ge t-\mathrm{sgn}(\theta_P^\top x)\} \cup \{(-1,x,t) \, :\,\mathrm{sgn}(\theta^\top x)\ge t+\mathrm{sgn}(\theta_P^\top x)\}\\
&\qquad \subseteq \{(1,x,t) \, :\,-\mathrm{sgn}(\theta^\top x) \ge t\} \cup \{(-1,x,t) \, :\,\mathrm{sgn}(\theta^\top x)\ge t\}\\
&\qquad \subseteq \{(1, x,t) \, :\,\theta^\top x \ge t\} \cup \{(-1, x,t) \, :\,\theta^\top x \ge t\}.
\end{align*}
Hence, the set of points that the subgraph of the function space $\mathcal{M}_\delta^\Delta$ can shatter is contained in the set of points that a half-space in $\mathbb{R}^d$ can shatter. Furthermore, the covering number of the VC functions (i.e., whose subgraphs form VC-class of sets) is given by
\begin{align*}
N(\varepsilon \|M\|_{L_2(Q)}, \mathcal{M}_\delta^\Delta, L_2(Q)) \le C d (16e)^d \left(\frac{1}{\varepsilon}\right)^{d}
\end{align*}
by Theorem 2.6.7 of van1996weak for any probabiliry measure $Q$ and $\varepsilon \in (0,1)$ and $C$ is a universal constant. We thus obtain
\begin{align*}
t\mapsto J(t, \mathcal{M}^\Delta_{\delta}, L_2) &= \sup_{Q} \int_{0}^t \sqrt{1+ \log\left(C d (16e)^d \left(\frac{1}{\varepsilon}\right)^{2d}\right)}\, d\varepsilon\\
&\le \sup_{Q} \int_{0}^t \sqrt{\mathfrak{C} d + 2d\log\left(\frac{1}{\varepsilon}\right)}\, d\varepsilon \le \mathfrak{C}t \sqrt{d\log(1/t)}
\end{align*}
where $\mathfrak{C}$ is a universal constants that may change line by line. Furthermore, we note that
\begin{align*}
\operatorname{\mathbb{E}}_P\left(\frac{y\left(\mathrm{sgn}(\theta_P^\top x)-\mathrm{sgn}(\theta^\top x)\right)}{2}\right)^2 = \mathbb{P}_P\left(\mathrm{sgn}(\theta_P^\top x) \neq \mathrm{sgn}(\theta^\top x)\right) \le \delta.
\end{align*}
Thus the condition of Theorem (ref) holds with $t^2 = \delta$ and $F=1$. By Theorem (ref), we obtain
\begin{align}
\operatorname{\mathbb{E}}_P\, \left[\sup_{m \in \mathcal{M}_\delta^\Delta}\, |\mathbb{G}_n m|\right]\lesssim \sqrt{\delta d\log(1/\delta)}\left(1+ \frac{ \sqrt{\delta d\log(1/\delta)}}{\delta \sqrt{n}}\right) = \sqrt{\delta d\log(1/\delta)}+ \frac{d\log(1/\delta)}{\sqrt{n}}.\nonumber
\end{align}
We can thus take $\phi_n$ in (ref) as
\begin{align*}
\delta \mapsto \phi_n(\delta) = \sqrt{\delta d\log(1/\delta)}+ \frac{d\log(1/\delta)}{\sqrt{n}} + \sqrt{\delta} \lesssim \sqrt{\delta d\log(1/\delta)}+ \frac{d\log(1/\delta)}{\sqrt{n}}.
\end{align*}
Solving the inequality in Theorem (ref),
\begin{align*}
&r_n^{-2} \left(r_n^{1/(1+\beta)} \sqrt{d\log(1/r_n)}+ \frac{d\log(1/r_n)}{\sqrt{n}}\right) \le n^{1/2} \\
&\qquad \implies r_n \ge \left(\frac{d\log(n/d)}{n}\right)^{(1+\beta)/(2+4\beta)} \vee \left(\frac{d\log(n/d)}{n}\right)^{1/2}\\
&\qquad \implies r_n^{2/(1+\beta)} \ge \left(\frac{d\log(n/d)}{n}\right)^{1/(1+2\beta)} \vee \left(\frac{d\log(n/d)}{n}\right)^{1/(1+\beta)}.
\end{align*}
Now by Theorem (ref), we conclude
\begin{align}
\mathrm{Diam}_{d_\Delta(\cdot)}\big(\widehat{\mathrm{CI}}^{\mathrm{Manski}}_{n,\alpha}\big)\le \mathfrak{C}\left\{\left(\frac{d\log(n/d)}{n}\right)^{1/(1+2\beta)}+ s_n^{1/(1+\beta)}\right\}
\end{align}
with probability greater than $1-\varepsilon$.
Finally, we relate this result to $\|\theta-\theta_P\|_2$ using (ref). The result established thus far and (ref) together imply that for any $\theta \in \widehat{\mathrm{CI}}^{\mathrm{Manski}}_{n,\alpha}$,
\begin{align*}
\|\theta - \theta(P)\|_2 \le c_0^{-1}d_{\Delta}(\theta_1, \theta(P)) \le \mathfrak{C}\left\{\left(\frac{d\log(n/d)}{n}\right)^{1/(1+2\beta)}+ s_n^{1/(1+\beta)}\right\},
\end{align*}
with high probability. This proves the result.
Quantile estimation without positive densities
For given $\gamma \in (0,1)$, we consider the quantile loss defined as
align*[align* omitted — 79 chars of source]
proof[{Proof of Theorem (ref)}]
Consider the case when $\theta > \theta_P$, then
\begin{align*}
m_{\theta} - m_{\theta_P}
& = \gamma(X-\theta)_+ + (1-\gamma)(\theta-X)_+ - \gamma(X-\theta_P)_+ - (1-\gamma)(\theta_P-X)_+\\
& = (1-\gamma)(\theta-\theta_P) \mathbf{1}\{X \le \theta_P\}
-\gamma(\theta-\theta_P)\mathbf{1}\{\theta < X\} \\
&\qquad\qquad+ \left\{(1-\gamma)(\theta-X)-\gamma(X-\theta_P)\right\} \mathbf{1}\{\theta_P < X \le \theta\}\\
&=(\theta-\theta_P) \mathbf{1}\{X \le \theta_P\}
-\gamma(\theta-\theta_P)(1-\mathbf{1}\{X > \theta_P\})+\gamma(\theta-\theta_P)\mathbf{1}\{\theta_P < X \le \theta\} \\
&\qquad+ \left\{\theta_P-X +(1-\gamma) (\theta - \theta_P)\right\} \mathbf{1}\{\theta_P < X \le \theta\}\\
& = (\theta-\theta_P) \mathbf{1}\{X \le \theta_P\}
-\gamma(\theta-\theta_P)+\gamma(\theta-\theta_P)\mathbf{1}\{\theta_P < X \le \theta\} \\
&\qquad+ \left\{\theta_P-X +(1-\gamma) (\theta - \theta_P)\right\} \mathbf{1}\{\theta_P < X \le \theta\}\\
& = (\theta-\theta_P) \mathbf{1}\{X \le \theta_P\}
-\gamma(\theta-\theta_P)+ (\theta-X) \mathbf{1}\{\theta_P < X \le \theta\}.
\end{align*}
Analogously, we have
\[m_{\theta} - m_{\theta_P} = \gamma(\theta_P-\theta) - (\theta_P-\theta) \mathbf{1}\{X \le \theta_P\} + (X-\theta) \mathbf{1}\{\theta < X \le \theta_P\}\]
when $\theta_P > \theta$. Taking expectations, we obtain
\begin{align*}
P(m_{\theta} - m_{\theta_P}) = \operatorname{\mathbb{E}}_P[ (\theta-X) \mathbf{1}\{\theta_P < X \le \theta\}] + \operatorname{\mathbb{E}}_P[ (X-\theta) \mathbf{1}\{\theta < X \le \theta_P\}].
\end{align*}
We now define the centered random variable,
\begin{align*}
W_i &= m_{\theta}(X_i)-m_{\theta_P}(X_i) - P( m_{\theta}-m_{\theta_P})\\
&= (\theta-\theta_P)\left(\mathbf{1}\{X \le \theta_P\} - \gamma\right) + (\theta-X) \mathbf{1}\{\theta_P < X \le \theta\} + (X-\theta) \mathbf{1}\{\theta < X \le \theta_P\}\\
&\qquad- \mathbb{E}_P[(\theta-X) \mathbf{1}\{\theta_P < X \le \theta\}]- \mathbb{E}_P[(X-\theta) \mathbf{1}\{\theta < X \le \theta_P\}].
\end{align*}
We take $H := \mathbf{1}\{X \le \theta_P\} - \gamma$ then check (ref) in Proposition (ref). For $\theta > \theta_P$, this follows
\begin{align*}
&\frac{\mathbb{E}_P|W - (\theta-\theta_P)(\mathbf{1}\{X \le \theta_P\} - \gamma)|^2}{|\theta-\theta_P|^2\operatorname{\mathbb{E}}_P(\mathbf{1}\{X \le \theta_P\} - \gamma)^2} \\
&\qquad=\frac{\mathbb{E}_P|(\theta-X) \mathbf{1}\{\theta_P < X \le \theta\} - \mathbb{E}_P[(\theta-X) \mathbf{1}\{\theta_P < X \le \theta\}|^2}{|\theta-\theta_P|^2\operatorname{\mathbb{E}}_P(\mathbf{1}\{X \le \theta_P\} - \gamma)^2} \\
&\qquad \le \frac{2(\theta-\theta_P)^2\mathbb{P}_P(\theta_P < X \le \theta) }{|\theta-\theta_P|^2\gamma(1-\gamma)}\\
&\qquad\le \frac{2|\mathbb{P}_P(\theta_P < X \le \theta)-M_0|\theta - \theta_P|^{\beta} \mathrm{sgn}(\theta-\theta_P) +M_0|\theta - \theta_P|^{\beta} \mathrm{sgn}(\theta-\theta_P)|}{\gamma(1-\gamma)}\\
&\qquad\le \frac{2M_1|\theta-\theta_P|^\beta +2M_0|\theta - \theta_P|^{\beta}}{\gamma(1-\gamma)}\le \frac{4M_0|\theta-\theta_P|^\beta }{\gamma(1-\gamma)}
\end{align*}
since $M_0 > M_1$. We then repeat the identical argument with $\theta < \theta_P$.
We use the last display as $\varphi(|\theta-\theta_P|)$ in (ref).
Now we introduce the event $\mathcal{E} := \{|\widehat \theta_1 - \theta(P)| < \delta\}$. Then we have
\begin{align*}
\mathbb{P}_P(\theta(P) \in \widehat{\mathrm{CI}}^\gamma_{n,\alpha}) &\ge\mathbb{P}_P(\theta(P) \in \widehat{\mathrm{CI}}^\gamma_{n,\alpha} \cap \mathcal{E})-\mathbb{P}_P(\theta(P) \in \widehat{\mathrm{CI}}^\gamma_{n,\alpha} \cap \mathcal{E}^c)\\
&\ge\mathbb{P}_P(\theta(P) \in \widehat{\mathrm{CI}}^\gamma_{n,\alpha} \cap \mathcal{E})-\mathbb{P}_P(\mathcal{E}^c).
\end{align*}
By invoking Proposition (ref) on the event $\mathcal{E} := \{|\widehat \theta_1 - \theta(P)| < \delta\}$, we have for any $\varepsilon < \delta$
\begin{align*}
\sup_{P\in\mathcal{P}}\, \Delta_{n,P} &\lesssim \frac{4M_0\varepsilon^{\beta}}{\gamma(1-\gamma)} +\left(\frac{4M_0\varepsilon^{\beta}}{\gamma(1-\gamma)}\right)^{1/2} + \sup_{P\in\mathcal{P}}\,\mathbb{P}_P(\|\widehat\theta_1 - \theta_P\| > \varepsilon) \\
&\qquad + \sup_{P\in\mathcal{P}}\,\operatorname{\mathbb{E}}_P\left[\frac{|\mathbf{1}\{X \le \theta_P\} - \gamma|^2}{\gamma(1-\gamma)} \left\{1 \wedge \sqrt{\frac{|\mathbf{1}\{X \le \theta_P\} - \gamma|^2}{n\gamma(1-\gamma)}}\right\}\right]
\end{align*}
since $B^2 = \operatorname{\mathbb{E}}_P(\mathbf{1}\{X \le \theta_P\} - \gamma)^2 = \gamma(1-\gamma)$. By the trivial upper bound $\Delta_{n,P} \le 1$, it suffices to consider $(4M_0\varepsilon^{\beta})/(\gamma(1-\gamma))\le 1$ and by the assumption that $\delta > \varepsilon$, we have
\begin{align*}
\sup_{P\in\mathcal{P}}\,\mathbb{P}_P(\|\widehat\theta_1 - \theta_P\| > \delta) \le \sup_{P\in\mathcal{P}}\,\mathbb{P}_P(\|\widehat\theta_1 - \theta_P\| > \varepsilon).
\end{align*}
Thus we obtain
\begin{align*}
\inf_{P\in\mathcal{P}}\,\mathbb{P}_P(\theta(P) \in \widehat{\mathrm{CI}}^\gamma_{n,\alpha}) \ge 1-\alpha -\sup_{P\in\mathcal{P}}\,\Delta_{n,P}
\end{align*}
where
\begin{align*}
\sup_{P\in\mathcal{P}}\Delta_{n,P} &\le C \left( 1\wedge \left(\frac{\varepsilon^{\beta}}{\gamma(1-\gamma)}\right)^{1/2} + \sup_{P\in\mathcal{P}}\,\mathbb{P}_P(\|\widehat\theta_1 - \theta_P\| > \varepsilon)\right. \\
&\left.\qquad + \sup_{P\in\mathcal{P}}\,\operatorname{\mathbb{E}}_P\left[\frac{|\mathbf{1}\{X \le \theta_P\} - \gamma|^2}{\gamma(1-\gamma)} \left\{1 \wedge \sqrt{\frac{|\mathbf{1}\{X \le \theta_P\} - \gamma|^2}{n\gamma(1-\gamma)}}\right\}\right]\right).
\end{align*}
Let $\{q_n\}$ be a positive, non-decreasing sequence such that $q_n \ge 1$ for all $n$ and
\begin{align*}
\lim_{n\to \infty}\sup_{P\in\mathcal{P}}\,\mathbb{P}_P( q_n\|\widehat\theta_1 - \theta_P\| > \varepsilon) = 0.
\end{align*}
Then, for any fixed $0 < \varepsilon < \delta \le 1$, we have
\begin{align*}
\Delta_{n,P} &\le C\left( 1\wedge \left(\frac{\varepsilon^{\beta}}{q_n^{\beta}\gamma(1-\gamma)}\right)^{1/2} + \sup_{P\in\mathcal{P}}\,\mathbb{P}_P(q_n\|\widehat\theta_1 - \theta_P\| > \varepsilon) \right.\\
&\qquad \left.+ \sup_{P\in\mathcal{P}}\,\operatorname{\mathbb{E}}_P\left[\frac{|\mathbf{1}\{X \le \theta_P\} - \gamma|^2}{\gamma(1-\gamma)} \left\{1 \wedge \sqrt{\frac{|\mathbf{1}\{X \le \theta_P\} - \gamma|^2}{n\gamma(1-\gamma)}}\right\}\right]\right).
\end{align*}
Under the stated assumptions on $\gamma$ and $q_n$, this term tends to zero by taking $n \to \infty$ and then taking $\varepsilon \to 0$.
proof[{Proof of Theorem (ref)}]
First, we condition on $D_1$ and we take the expectation over $D_1$ at the end. The proof is split into two cases: (1) $|\theta-\theta_P| \le \delta$ and (2) $|\theta-\theta_P| > \delta$ where $\delta > 0$ corresponds to the value defined in (ref). Without loss of generality, we assume $\delta < 1$. When (ref) holds with $\delta > 1$, we set $\delta = 1$.
First we consider the case (1). We have shown that
\begin{align*}
P(m_{\theta} - m_{\theta_P}) = \operatorname{\mathbb{E}}_P[ (\theta-X) \mathbf{1}\{\theta_P < X \le \theta\}] + \operatorname{\mathbb{E}}_P[ (X-\theta) \mathbf{1}\{\theta < X \le \theta_P\}].
\end{align*}
We observe that $(\theta-X) \mathbf{1}\{\theta_P < X \le \theta\}$ is a non-negative random variable, taking values from $0$ to $\theta-\theta_P$ (when $\theta > \theta_P$). Let $R \in (0,1)$ be an arbitrary constant. Then
\begin{align*}
(\theta-X) \mathbf{1}\{\theta_P < X \le \theta\} \ge R(\theta-\theta_P)\mathbf{1}\{\theta_P < X \le \theta_P + (1-R)(\theta-\theta_P)\}
\end{align*}
and thus
\begin{align*}
\operatorname{\mathbb{E}}_P[ (\theta-X) \mathbf{1}\{\theta_P < X \le \theta\}]\ge R(\theta-\theta_P)\{F(\theta_P + (1-R)(\theta-\theta_P)) - F(\theta_P)\}.
\end{align*}
When $\theta-\theta_P \le \delta$,
by (ref), we can further lower bound the expectation as
\begin{align*}
P(m_{\theta} - m_{\theta_P}) &\ge RM_0 (1-R)^{\beta}|\theta-\theta_P|^{1+\beta} - M_1R(1-R)^\beta|\theta-\theta_P|^{1+\beta}\\
&= (M_0-M_1)R (1-R)^{\beta}|\theta-\theta_P|^{1+\beta}.
\end{align*}
By the assumption that $M_0 > M_1$, we can conclude that
\begin{align*}
P(m_{\theta} - m_{\theta_P}) \ge \mathfrak{C} (\theta-\theta_P)^{1+\beta}
\end{align*}
where $\mathfrak{C}$ depends on $M_0$ and $M_1$.
Repeating the analogous argument for the case with $\theta_P > \theta$, we conclude that
\begin{align*}
P(m_{\theta} - m_{\theta_P}) \ge \mathfrak{C} |\theta-\theta_P|^{1+\beta}.
\end{align*}
The rest of the proof is similar to that of Theorem (ref) with small modification that does not alter the main result. First, observe that
\begin{align*}
(m_{\theta} - m_{\theta_P})^2 &= \big\{\gamma(\theta_P-\theta) - (\theta_P-\theta) \mathbf{1}\{X \le \theta_P\} + (X-\theta) \mathbf{1}\{\theta < X \le \theta_P\}\big\}^2 \\
&\le 2(\gamma^2 +2)|\theta_P-\theta|^2,
\end{align*}
and thus we obtain
\begin{align*}
\operatorname{\mathbb{E}}_P \left[\sup_{\|\theta-\theta_P\| < \delta}|\mathbb{G}_n (m_\theta - m_{\theta_P})| \right] \vee \sup_{\|\theta-\theta_P\| < \delta}\, \sqrt{P(m_{\theta} - m_{\theta_P})^2}\le\sqrt{2}(2+\gamma)\delta =: \phi(\delta).
\end{align*}
By Lemma (ref), we also obtain
\begin{align*}
\omega^2(\delta) = C\left\{ n^{1/2} \left(\mathbb{E}_P M_\delta^2 \right)+ \left(\mathbb{E}_PM_\delta^2\right)^{1/2}\operatorname{\mathbb{E}}_P \left[\sup_{m \in \mathcal{M}_{\delta}}\, \left|\sum_{i=1}^n \varepsilon_i m(X_i)\right|\right]\right\} = C\sqrt{2}(2+\gamma)^2n^{1/2}\delta^2.
\end{align*}
It remains to solve equations in Theorem (ref) to derive the rate of convergence. Dropping constants,
\begin{align*}
(2+\gamma)r_n^{2/(1+\beta)}r_n^{-2} \le n^{1/2} \Longrightarrow (2+\gamma)^{(1+\beta)/(2\beta)}n^{-(1+\beta)/(4\beta)} \le r_n
\end{align*}
and
\begin{align*}
C^{1/2}(2+\gamma)n^{1/4}u_n^{2/(1+\beta)}u_n^{-2} \le n^{3/4} \Longrightarrow C^{(1+\beta)/(4\beta)}(2+\gamma)^{(1+\beta)/(2\beta)}n^{-(1+\beta)/(4\beta)} \le u_n.
\end{align*}
While $\gamma$ may depend on $n$, we observe that
\begin{align*}
2^{(1+\beta)/(2\beta)} \le (2+\gamma)^{(1+\beta)/(2\beta)} \le 3^{(1+\beta)/(2\beta)},
\end{align*}
and hence the leading constant does not depend on $n$. Finally by Theorem (ref), we conclude the width of the confidence set is bounded with probability $1-\varepsilon$ by
\begin{align*}
&\mathfrak{C}(r_n^{2/(1+\beta)} + u_n^{2/(1+\beta)}+s_n^{1/(1+\beta)})\le \mathfrak{C}\left(\frac{1}{n^{1/2\beta}} + s_n^{1/(1+\beta)}\right)
\end{align*}
with $\mathfrak{C}$ depending on $\alpha, \beta, M_0, M_1$ and $\varepsilon$.
We now consider the case (2). For any $\theta = \theta_P + \ell u$ such that $u \in \{-1, 1\}$ and $\ell > \delta$, we define $\bar\theta = \theta_P + \ell u$. Since $\bar\theta = (1-\delta/\ell)\theta_P + \delta/\ell \theta$, it follows by the convexity of $\theta \mapsto Pm_{\theta}$,
\begin{align*}
(1-\delta/\ell) Pm_{\theta_P} - \delta/\ell Pm_{\theta} \ge Pm_{\bar \theta} \Longleftrightarrow Pm_{\theta}-Pm_{\theta_P} \ge (\ell/\delta)(Pm_{\bar \theta} -Pm_{\theta_P}) \ge\mathfrak{C} |\theta-\theta_P|\delta^{\beta}.
\end{align*}
Next, recall for any $\theta \in \Theta$,
\begin{align*}
|m_{\theta}-m_{\theta_P}|\le \sqrt{2}(2+\gamma)|\theta-\theta_P|
\end{align*}
and this implies
\begin{align*}
\widehat \sigma_{\theta, \theta_P}^2 \le \mathbb{P}_n|m_{\theta}-m_{\theta_P}|^2 \le 2(2+\gamma)^2|\theta-\theta_P|^2 \quadand\quad \widehat \sigma_{\widehat\theta_1, \theta_P}^2 \le 2(2+\gamma)^2|\widehat\theta_1-\theta_P|^2.
\end{align*}
Any $\theta \in \widehat{\mathrm{CI}}_{n,\alpha}^{\gamma}$ satisfies
\begin{align*}
&\mathbb{P}_n(m_{\theta} - m_{\widehat{\theta}_1}) - q_{1-\alpha}n^{-1/2}\widehat\sigma_{\theta, \widehat{\theta}_1} \le 0 \\
&\qquad \Longleftrightarrow P(m_{\theta} - m_{\theta_P}) + (\mathbb{P}_n-P)(m_{\theta} - m_{\theta_P}) \le \mathbb{P}_n(m_{\widehat{\theta}_1} - m_{\theta_P}) + q_{1-\alpha}n^{-1/2}\widehat\sigma_{\widehat{\theta}_1,\theta}\\
&\qquad \Longrightarrow P(m_{\theta} - m_{\theta_P}) + (\mathbb{P}_n-P)(m_{\theta} - m_{\theta_P}) \\
&\qquad\qquad \le \mathbb{P}_n(m_{\widehat{\theta}_1} - m_{\theta_P}) + q_{1-\alpha}n^{-1/2}\sqrt{2}(2+\gamma)\left(|\theta-\theta_P|+ |\widehat{\theta}_1-\theta_P|\right)\\
&\qquad \Longrightarrow P(m_{\theta} - m_{\theta_P}) + (\mathbb{P}_n-P)(m_{\theta} - m_{\theta_P}) \\
&\qquad\qquad \le \mathbb{P}_n(m_{\widehat{\theta}_1} - m_{\theta_P}) + q_{1-\alpha}n^{-1/2}\sqrt{2}(2+\gamma)\left(|\theta-\theta_P|+ |\widehat{\theta}_1-\theta_P|\right).
\end{align*}
When $|\theta-\theta_P| \ge \delta$, we have $P(m_{\theta} - m_{\theta_P}) > \mathfrak{C}\delta^{\beta+1}$. The lower bound of the above inequality can be written out as
\begin{align}
P(m_{\theta} - m_{\theta_P}) + (\mathbb{P}_n-P)(m_{\theta} - m_{\theta_P}) = P(m_{\theta} - m_{\theta_P})\left(1 + \frac{(\mathbb{P}_n-P)(m_{\theta} - m_{\theta_P})}{P(m_{\theta} - m_{\theta_P})}\right).
\end{align}
We claim that the fraction inside of the parenthesis is stochastically bounded when $|\theta-\theta_P| > \delta$. First, we observe that
\begin{align*}
\mathbb{E}\left[\sup_{\theta \in \Theta; |\theta-\theta_P| > \delta}\left|\frac{(\mathbb{P}_n-P)(m_{\theta} - m_{\theta_P})}{P(m_{\theta} - m_{\theta_P})}\right|\right]&\le \mathbb{E}\left[\sup_{\theta \in \Theta; |\theta-\theta_P| > \delta}\left|\frac{(\mathbb{P}_n-P)(m_{\theta} - m_{\theta_P})}{\mathfrak{C}|\theta-\theta_P|\delta^{\beta}}\right|\right] \\
&\le \sqrt{\mathbb{E}\left[\sup_{\theta \in \Theta; |\theta-\theta_P| > \delta}\left|\frac{(\mathbb{P}_n-P)(m_{\theta} - m_{\theta_P})}{\mathfrak{C}|\theta-\theta_P|\delta^{\beta}}\right|^2\right]}\\
&\le n^{-1/2}\mathfrak{C}^{-1}\delta^{-\beta}\left(\operatorname{\mathbb{E}} \left[\sup_{\theta \in \Theta; |\theta-\theta_P| > \delta}\frac{|m_{\theta} - m_{\theta_P}|^2}{|\theta-\theta_P|^2}\right]\right)^{1/2}\\
&\le n^{-1/2}\mathfrak{C}^{-1}\delta^{-\beta}(2+\gamma).
\end{align*}
Hence by Markov's inequality, we have
\begin{align}
\sup_{\theta \in \Theta; |\theta-\theta_P| > \delta}\left|\frac{(\mathbb{P}_n-P)(m_{\theta} - m_{\theta_P})}{P(m_{\theta} - m_{\theta_P})}\right| \le C_{\varepsilon, \beta, M_0, M_1} n^{-1/2}\delta^{-\beta}
\end{align}
with probability greater than $1-\varepsilon$ for $C_{\varepsilon, \beta, M_0, M_1}$ large enough. Coming back to the inequality, we have
\begin{align*}
&P(m_{\theta} - m_{\theta_P}) + (\mathbb{P}_n-P)(m_{\theta} - m_{\theta_P}) \&\qquad= P(m_{\theta} - m_{\theta_P})\left(1 + \frac{(\mathbb{P}_n-P)(m_{\theta} - m_{\theta_P})}{P(m_{\theta} - m_{\theta_P})}\right)\\
&\qquad\ge P(m_{\theta} - m_{\theta_P})\left(1 - \sup_{|\theta-\theta_P| > \delta}\left|\frac{(\mathbb{P}_n-P)(m_{\theta} - m_{\theta_P})}{P(m_{\theta} - m_{\theta_P})}\right|\right)_+
\\
&\qquad\ge P(m_{\theta} - m_{\theta_P})\left(1 - C_{\varepsilon, \beta, M_0, M_1} n^{-1/2}\delta^{-\beta}\right)_+.
\end{align*}
On the event the inequality holds, and $n \ge 4C^2_{\varepsilon, \beta, M_0, M_1}\delta^{-2\beta}$, we have
\[P(m_{\theta} - m_{\theta_P}) + (\mathbb{P}_n-P)(m_{\theta} - m_{\theta_P}) \ge 2^{-1}P(m_{\theta} - m_{\theta_P}) \ge2^{-1}\mathfrak{C}\delta^\beta|\theta-\theta_P|. \]
Plugging this result into the expression for the confidence set, we arrive at
\begin{align*}
&\mathbb{P}_n(m_{\theta} - m_{\widehat{\theta}_1}) - q_{1-\alpha}n^{-1/2}\widehat\sigma_{\theta, \widehat{\theta}} \le 0 \\
&\qquad \Longrightarrow 2^{-1}\mathfrak{C}\delta^\beta|\theta-\theta_P| \le \mathbb{P}_n(m_{\widehat{\theta}} - m_{\theta_P}) + q_{1-\alpha}n^{-1/2}\sqrt{2}(2+\gamma)\left(|\theta-\theta_P|+ |\widehat{\theta}-\theta_P|\right)\\
&\qquad \Longrightarrow \left(2^{-1}\mathfrak{C}\delta^\beta-q_{1-\alpha}n^{-1/2}\sqrt{2}(2+\gamma)\right)_+|\theta-\theta_P| \\
&\qquad\qquad\le \mathbb{P}_n(m_{\widehat{\theta}} - m_{\theta_P}) + q_{1-\alpha}n^{-1/2}\sqrt{2}(2+\gamma)|\widehat{\theta}-\theta_P|.
\end{align*}
Furthermore when $n \ge 16\mathfrak{C}^{-2}\delta^{-2\beta}q^2_{1-\alpha}\sqrt{2}(2+\gamma)^2$, the lower bound becomes strictly positive. Hence, it implies
\begin{align}
\left(2^{-1}\mathfrak{C}\delta^\beta-q_{1-\alpha}n^{-1/2}\sqrt{2}(2+\gamma)\right)_+|\theta-\theta_P| \ge 2^{-2}\mathfrak{C}\delta^\beta |\theta-\theta_P|.
\end{align}
Putting together, we conclude that for $n \ge (32\mathfrak{C}^{-2}\delta^{-2\beta}q^2_{1-\alpha}(2+\gamma)^2 \vee 4C^2_{\varepsilon, \beta, M_0, M_1}\delta^{-2\beta})$, we have
\begin{align*}
|\theta-\theta_P| &\le 4\mathfrak{C}^{-1}\delta^{-\beta}\left(\mathbb{P}_n(m_{\widehat{\theta}} - m_{\theta_P}) + q_{1-\alpha}n^{-1/2}\sqrt{2}(2+\gamma)|\widehat{\theta}-\theta_P|\right)\\
&\le 8\mathfrak{C}^{-1}\delta^{-\beta} q_{1-\alpha}\sqrt{2}(2+\gamma)|\widehat{\theta}-\theta_P|
\end{align*}
with probability greater than $1-\varepsilon$.
Summarizing the results we obtained so far, we have shown that there exist constants $C_1, C_2, C_3$ all different but only depend on $\varepsilon, M_0, M_1, \beta$ such that for all $n \ge C_1 \delta^{-2\beta}$, we have
\begin{align*}
\mathbb{P}_P\left(\mathrm{Diam}_{\|\cdot\|}\big(\widehat{\mathrm{CI}}_{n,\alpha}^{\gamma}\big)\le C_3\tau_n \right) \ge 1-\varepsilon
\end{align*}
where
\begin{align*}
\tau_n = \left(\frac{1}{n^{1/2\beta}} + s_n^{1/(1+\beta)}\right)\mathbf{1}\{|\widehat{\theta}_1-\theta_P| \le C_2 \delta^{1+\beta}\} + \delta^{-\beta}|\widehat{\theta}_1-\theta_P|\mathbf{1}\{|\widehat{\theta}_1-\theta_P| > C_2 \delta^{1+\beta}\}.
\end{align*}
We conclude the first claim.
Next, we verify the requirements for $\widehat\theta_1$. Recalling that $|m_{\theta}-m_{\theta_P}|\lesssim |\theta-\theta_P|$ for any $\theta$. On the other hand, we have $P(m_{\theta} - m_{\theta_P}) \le |\widehat \theta_1 - \theta_P|\mathbb{P}_P(\widehat\theta < X \le \theta_P)$ when $|\widehat \theta_1 - \theta_P | < \delta$. Putting together implies
\begin{align*}
|m_{\theta}-m_{\theta_P}| + n^{-1/2}\sqrt{P(m_{\widehat\theta_1}-m_{\theta_P})^2} \lesssim |\widehat\theta_1-\theta_P|^{1+\beta} + \left(\frac{|\widehat\theta_1-\theta_P|^2}{n}\right)^{1/2} = s_n.
\end{align*}
Introducing an event $\mathcal{E}=:\{(\widehat\theta_1-\theta_P)^2 \lesssim n^{-1/\beta}\}$ such that $\mathbb{P}(\mathcal{E}) \ge 1-\varepsilon/2$. On the event $\Omega$, we take $n$ large enough so $|\widehat \theta_1 - \theta_P | < \delta$. We then conclude that $s_n^{1/(1+\beta)} \lesssim n^{-1/2\beta}$ with probability greater than $1-\varepsilon/2$.
Finally by the first result of this theorem, it holds that
\begin{align*}
\mathbb{P}_P\left(\mathrm{Diam}_{\|\cdot\|}\big(\widehat{\mathrm{CI}}_{n,\alpha}^{\gamma}\big)\le C_3\tau_n\right) \ge 1-\varepsilon/2
\end{align*}
Furthermore, observe that
\begin{align*}
\mathbb{P}_P\left(\mathrm{Diam}_{\|\cdot\|}\big(\widehat{\mathrm{CI}}_{n,\alpha}^{\gamma}\big)\le C_3\tau_n \right) &\le \mathbb{P}_P\left(\mathrm{Diam}_{\|\cdot\|}\big(\widehat{\mathrm{CI}}_{n,\alpha}^{\gamma}\big)\le C_3\tau_n \bigg | \mathcal{E}\right)\mathbb{P}(\mathcal{E}) + \varepsilon/2 \\
&=\mathbb{P}_P\left(\mathrm{Diam}_{\|\cdot\|}\big(\widehat{\mathrm{CI}}_{n,\alpha}^{\gamma}\big)\le \frac{\mathfrak{C}}{n^{1/2\beta}} \bigg | \mathcal{E}\right)\mathbb{P}(\mathcal{E}) + \varepsilon/2
\end{align*}
where the second equality holds for all $n \ge (C_2\delta)^{-\beta}$ on $\mathcal{E}$. This concludes the claim.
\subsection{Discrete argmin inference}
\begin{proof}[{Proof of Theorem (ref)}]
Let $\Omega$ be an event where $\widehat{\theta}_1 \in \theta(P)$, that is, the initial estimator correctly identifies one of the elements in the argmin set. By the definition of the confidence set $\widehat{\theta}_1 \in \widehat{\mathrm{CI}}_{n,\alpha}^{\mathrm{argmin}}$ almost surely. It follows that
\begin{align*}
\mathbb{P}_P(\theta(P) \notin \widehat{\mathrm{CI}}_{n,\alpha}^{\mathrm{argmin}} \mid D_1) &= \mathbb{P}_P(\theta(P) \notin \widehat{\mathrm{CI}}_{n,\alpha}^{\mathrm{argmin}} \cap \Omega \mid D_1) + \mathbb{P}_P(\theta(P) \notin \widehat{\mathrm{CI}}_{n,\alpha}^{\mathrm{argmin}} \cap \Omega^c \mid D_1)\\
&= \mathbb{P}_P(\theta(P) \notin \widehat{\mathrm{CI}}_{n,\alpha}^{\mathrm{argmin}} \cap \Omega^c \mid D_1).
\end{align*}
Hence, it suffices to analyze the case when $\widehat{\theta}_1 \neq \theta(P)$. We provide the sufficient condition under which $\Delta_{n,P}$ tends to zero as $n \to \infty$. Observe that
\begin{align*}
W = m_{\widehat{\theta}_1} - m_{\theta(P)} - P(m_{\widehat{\theta}_1} - m_{\theta(P)})= e_{\widehat{\theta}_1}^\top X - e_{\theta(P)}^\top X - (e_{\widehat{\theta}_1}^\top \operatorname{\mathbb{E}}_P[X] - e_{\theta(P)}^\top \operatorname{\mathbb{E}}_P[X])
\end{align*}
By directly inspecting the upper bound in Lemma (ref), we obtain
\begin{align*}
\Delta_{n,P} &\le C_0 \operatorname{\mathbb{E}}_P\left[\frac{W^2}{B^2} \left\{1 \wedge n^{-1/2}\frac{|W|}{B}\right\}\bigg|D_1\right] \\
&\le C_0 \sup_{i\neq j}\,\operatorname{\mathbb{E}}_P\left[\frac{W_{ij}^2}{\operatorname{\mathbb{E}}_P [W_{ij}^2]} \min\left\{1, \frac{|W_{ij}|}{n^{1/2}(\operatorname{\mathbb{E}}_P [W_{ij}^2])^{1/2}}\right\} \right]
\end{align*}
where $W_{i, j} = e_{i}^\top (X - \operatorname{\mathbb{E}}_P[X]) - e_{j}^\top (X - \operatorname{\mathbb{E}}_P[X])$
for some universal constant $C_0 > 0$. Hence assuming (ref), we conclude the claim in view of Theorem (ref). This result does not require the consistency of the initial estimator.
\end{proof}
Technical Lemma
We denote the localized collection of functions by
align*[align* omitted — 113 chars of source]
and let the envelope be denoted by $M_\delta := z \mapsto \sup_{m \in \mathcal{M}_{\delta}}|m(z)|$. The following result provides the upper bound on the expectation of the squared empirical process without requiring the boundedness assumption on the $\mathcal{M}_\delta$.
lemmaGiven $n$ IID random variables, $Z_1, \ldots, Z_n$, the following holds:
\begin{align*}
\operatorname{\mathbb{E}}_P \left[\sup_{m\in\mathcal{M}_{\delta}}\, |\mathbb{G}_n m^2 |\right] \le C\left\{ n^{1/2} \left(\mathbb{E}_P M_\delta^2 \right)+ \left(\mathbb{E}_PM_\delta^2\right)^{1/2}\operatorname{\mathbb{E}}_P \left[\sup_{m \in \mathcal{M}_{\delta}}\, \left|\sum_{i=1}^n \varepsilon_i m(Z_i)\right|\right]\right\}
\end{align*}
where $\{\varepsilon_i\}_{i=1}^n$ is a sequence of independent Rademacher random variables and $C$ is a universal constant.
proof[{Proof of Lemma (ref)}]
Define a truncation parameter $B := 8 n^{1/2} (\operatorname{\mathbb{E}}_P M_\delta^2)^{1/2}$ and let $\{\varepsilon_i\}_{i=1}^n$ be a sequence of independent Rademacher random variables. Then, we have
\begin{align*}
\operatorname{\mathbb{E}}_P \left[\sup_{m \in \mathcal{M}_\delta}\, |\mathbb{G}_n m^2| \right]& = n^{-1/2}\operatorname{\mathbb{E}}_P \left[\sup_{m\in\mathcal{M}_{\delta}}\, \left|\sum_{i=1}^n m^2(Z_i) - \mathbb{E}_P m^2\right| \right]\\
& \le 2n^{-1/2}\operatorname{\mathbb{E}}_{P\times \varepsilon} \left[\sup_{m \in \mathcal{M}_\delta}\, \left|\sum_{i=1}^n \varepsilon_i m^2(Z_i)\right|\right]\\
& \le 2n^{-1/2}\operatorname{\mathbb{E}}_{P\times \varepsilon} \left[\sup_{m \in \mathcal{M}_\delta}\, \left|\sum_{i=1}^n \varepsilon_i m^2(Z_i)\mathbf{1}\{M_\delta
> B\}\right|\right] \\
&\qquad+ 2n^{-1/2}\operatorname{\mathbb{E}}_{P\times \varepsilon} \left[\sup_{m \in \mathcal{M}_\delta}\, \left|\sum_{i=1}^n \varepsilon_i m^2(Z_i)\mathbf{1}\{M_\delta
\le B\}\right|\right].
\end{align*}
where the second inequality follows from symmetrization (see for instance, Lemma 2.3.1 of van1996weak). We now handle two terms separately. For the unbounded part, we have
\begin{align*}
&\operatorname{\mathbb{E}}_P \left[\sup_{m \in \mathcal{M}_\delta}\, \left|\sum_{i=1}^n \varepsilon_i m^2(Z_i)\mathbf{1}\{M_\delta
> B\}\right|\right] \le \operatorname{\mathbb{E}}_P \left[\left|\sum_{i=1}^n M_\delta^2(Z_i) \mathbf{1}\{M_\delta
> B\}\right|\right].
\end{align*}
We apply the Hoffmann-J$\o$rgensen inequality (See Proposition 6.8 of ledoux2013probability with $p=1$), which states
\begin{align*}
\operatorname{\mathbb{E}}_P \left[\left|\sum_{i=1}^n M_\delta^2(Z_i)\mathbf{1}\{M_\delta
> B\}\right|\right] \le 8 \left(\mathbb{E}_P\left[\max_{1\le i \le n} M_\delta^2(Z_i)\right] + t_0^2\right)
\end{align*}
for any $t_0$ such that
\begin{align}
\mathbb{P}_P\left(\sum_{i=1}^n M_\delta^2(Z_i) \mathbf{1}\{M_\delta
> B\} > t_0\right) \le 1/8.
\end{align}
At our truncation level $B$, the result follows with $t_0=0$. We can indeed verify (ref) by observing
\begin{align*}
\mathbb{P}_P\left(\sum_{i=1}^n M_\delta^2 \mathbf{1}\{M_\delta
> B\} > 0\right) &\le \mathbb{P}_P\left(\max_{1\le i\le n}M_\delta(Z_i)
> B\right) \\
&\le \frac{\mathbb{E}_P[\max_{1\le i\le n}M_\delta(Z_i) ]}{B}\\
&\le \frac{\left(\mathbb{E}_P[\sum_{i=1}^n M^2_\delta(Z_i) ]\right)^{1/2}}{B} \le 1/8.
\end{align*}
Hence by the Hoffmann-J$\o$rgensen inequality, we conclude
\begin{align*}
2n^{-1/2}\operatorname{\mathbb{E}}_P \left[\sup_{m \in \mathcal{M}_\delta}\, \left|\sum_{i=1}^n \varepsilon_i m^2(Z_i)\mathbf{1}\{M_\delta
> B\}\right|\right] &\le 16n^{-1/2} \mathbb{E}_P\left[\max_{1\le i \le n} |M_\delta(Z_i)|^2\right] \\
&\le 16n^{1/2} \left(\mathbb{E}_P M_\delta^2\right).
\end{align*}
For the second term, we observe that the entire process is uniformly bounded by $B$. We can thus apply the contraction inequality, such as, Theorem 4.12 of ledoux2013probability or Corollary 3.2.2 of gine2021mathematical. This in tern implies that
\begin{align*}
2n^{-1/2}\operatorname{\mathbb{E}}_P \left[\sup_{m \in \mathcal{M}_\delta}\,\left|\sum_{i=1}^n \varepsilon_i m^2(Z_i)\mathbf{1}\{M_\delta
\le B\}\right|\right] &\le 4B n^{-1/2}\operatorname{\mathbb{E}}_P \left[\sup_{m \in \mathcal{M}_\delta}\,\left|\sum_{i=1}^n \varepsilon_i m(Z_i)\right|\right]\\
&= 4\left(\mathbb{E}_PM_\delta^2\right)^{1/2}\operatorname{\mathbb{E}}_P \left[\sup_{m \in \mathcal{M}_\delta}\,\left|\sum_{i=1}^n \varepsilon_i m(Z_i)\right|\right],
\end{align*}
thus we conclude the claim.
lemmaSuppose that $X$ and $Y$ are the random variables defined on a common measurable space, satisfying $\mathbb{E}|X-Y|^2 \le C \mathbb{E}[X^2]$ with some constant $C>0$.
Then we have the following three upper bounds:
\begin{enumerate}
• \[\mathbb{E}\left[\left|\frac{X^2}{\mathbb{E}[X^2]}-\frac{Y^2}{\mathbb{E}[Y^2]}\right|\right] \leq 4 (C+ C^{1 / 2})\]
• For any $\kappa > 0$,
\[\mathbb{E}{\left[\frac{X^2}{\mathbb{E}[X^2]} \mathbf{1}\{X^2>\kappa\mathbb{E}[X^2]\}\right]-\mathbb{E}\left[\frac{Y^2}{\mathbb{E}[Y^2]} \mathbf{1}\{Y^2>\kappa\mathbb{E}[Y^2]\}\right] } \leq 16(C+ C^{1 / 2})\]
• For any $\kappa > 0$,
\[\mathbb{E} {\left[\left(\frac{X^2}{\mathbb{E}[X^2]}\right)^{3/2} \mathbf{1}\{X^2\le \kappa\mathbb{E}[X^2]\}\right]-\mathbb{E}\left[\left(\frac{Y^2}{\mathbb{E}[Y^2]}\right)^{3/2} \mathbf{1}\{Y^2\le \kappa\mathbb{E}[Y^2]\}\right] } \leq 16 \kappa^{1/2} (C+ C^{1 / 2}).\]
\end{enumerate}
proof[{Proof of Lemma (ref)}]
Suppose $X$ and $Y$ are random variables such that $\mathbb{E}|X-Y|^2 \le C \mathbb{E}[X^2]$.
To obtain the first result, we observe that
\begin{align*}
\mathbb{E}\left[\left|\frac{X^2}{\mathbb{E}[X^2]}-\frac{Y^2}{\mathbb{E}[Y^2]}\right|\right]
&=\mathbb{E}\left[\left|\frac{X^2 \mathbb{E}[Y^2]-\mathbb{E}[X^2] Y^2}{\mathbb{E}[X^2] \mathbb{E}[Y^2]}\right|\right] \\
&\le \mathbb{E}\left[\left|\frac{(X^2-Y^2) \mathbb{E}[Y^2]}{\mathbb{E}[X^2] \mathbb{E}[Y^2]}\right|\right]+\mathbb{E}\left[\left|\frac{Y^2(\mathbb{E}[Y^2]-\mathbb{E}[X^2])}{\mathbb{E}[X^2] \mathbb{E}[Y^2]}\right|\right] \\
& =\frac{\mathbb{E}[|X^2-Y^2|]}{\mathbb{E}[X^2]}+\frac{|\mathbb{E}[Y^2-X^2]|}{\mathbb{E}[X^2]} \\
&\leq \frac{2\mathbb{E}[|Y^2-X^2|]}{\mathbb{E}[X^2]}
\end{align*}
by Jensen's inequality. Since $Y^2=(Y-X)^2+2 X(Y-X)+X^2$, we obtain
\begin{align*}
\mathbb{E}[| Y^2-X^2|] &= \mathbb{E}[|Y-X|^2] +2\mathbb{E}[ X(Y-X)]\\
&\leq \mathbb{E}[|Y-X|^2]+2\left(\mathbb{E}[X^2]\right)^{1/2}\left(\mathbb{E}[|Y-X|^2]\right)^{1 / 2}\\
&\le C \mathbb{E}[X^2]+2C^{1/2}\mathbb{E}[X^2]
\end{align*}
by Cauchy-Schwarz inequality. Hence we conclude
\begin{align*}
\mathbb{E}\left[\left|\frac{X^2}{\mathbb{E}[X^2]}-\frac{Y^2}{\mathbb{E}[Y^2]}\right|\right]
\leq \frac{2\mathbb{E}[|Y^2-X^2|]}{\mathbb{E}[X^2]} \le 2C +4C^{1/2}.
\end{align*}
Moving onto the second result, we first denote by $\widetilde{X}=X^2 / \mathbb{E}[X^2]$ and $\widetilde{Y}=Y^2 / \mathbb{E}[Y^2]$. It then follows that
\begin{align*}
& \mathbb{E}[\widetilde{X} \mathbf{1}\{\widetilde{X}>\kappa\}]-\mathbb{E}[\widetilde{Y}\mathbf{1}\{\widetilde{Y}>\kappa\}] \\
& \qquad \leq \mathbb{E}[\widetilde{X} \mathbf{1}\{\widetilde{X}>\kappa / 2\}]-\mathbb{E}[\widetilde{Y}\mathbf{1}\{\widetilde{Y}>\kappa\}] \\
& \qquad\leq \mathbb{E}[\widetilde{X} \mathbf{1}\{\widetilde{X}>\kappa / 2, \widetilde{Y}>\kappa\}]+\mathbb{E}[\widetilde{X} \mathbf{1}\{\widetilde{X}>\kappa / 2, \widetilde{Y}\leq \kappa\}] \\
&\qquad\qquad -\mathbb{E}[\widetilde{Y}\mathbf{1}\{\widetilde{Y}>\kappa,\widetilde{X}>\kappa / 2\}]-\mathbb{E}[\widetilde{Y}\mathbf{1}\{\widetilde{Y}>\kappa,\widetilde{X} \leq \kappa / 2\}] \\
& \qquad\leq \mathbb{E}[| \widetilde{X}-\widetilde{Y}|]+\mathbb{E}[\widetilde{X} \mathbf{1}\{\widetilde{X}>\kappa / 2, \widetilde{Y}\leq \kappa\}] \\
&\qquad \leq \mathbb{E}[|\widetilde{X}-\widetilde{Y}|]+\mathbb{E}[(\widetilde{X}-\widetilde{Y}) \mathbf{1}\{\widetilde{X}>\kappa / 2,\widetilde{Y}\leq \kappa\}]+\mathbb{E}[\widetilde{Y} \mathbf{1}\{\widetilde{X}>\kappa / 2,\widetilde{Y}\leq \kappa\}] \\
&\qquad \leq 2 \mathbb{E}[|\widetilde{X}-\widetilde{Y}|]+\kappa \mathbb{E}[\mathbf{1}\{|\widetilde{X}-\widetilde{Y}|>\kappa / 2\}] \\
& \qquad\leq 4 \mathbb{E}[|\widetilde{X}-\widetilde{Y}|].
\end{align*}
Using the first result, the second claim is obtained.
Finally, it follows that
\begin{align*}
& \mathbb{E}[\widetilde{X}^{3 / 2} \mathbf{1}\{\widetilde{X} \leq \kappa\}]-\mathbb{E}[\widetilde{Y}^{3 / 2} \mathbf{1}\{\widetilde{Y}\leq \kappa\}] \\
& \qquad\leq \mathbb{E}[\widetilde{X}^{3 / 2} \mathbf{1}\{\widetilde{X} \leq \kappa\}]-\mathbb{E}[\widetilde{Y}^{3 / 2} \mathbf{1}\{\widetilde{Y}\leq \kappa / 2\}] \\
&\qquad\leq \mathbb{E}[(\widetilde{X}^{3 / 2}-\widetilde{Y}^{3 / 2}) \mathbf{1}\{\widetilde{X} \leq \kappa, \widetilde{Y}\leq \kappa / 2\}] \\
&\qquad\qquad +\mathbb{E}[\widetilde{X}^{3 / 2} \mathbf{1}\{\widetilde{X} \leq \kappa,\widetilde{Y}>\kappa / 2\}]-\mathbb{E}[\widetilde{Y}^{3 / 2} \mathbf{1}\{\widetilde{X}>\kappa,\widetilde{Y}\leq \kappa / 2\}] \\
&\qquad\leq \frac{3}{2} \kappa^{1 / 2} \mathbb{E}[|\widetilde{Y}-\widetilde{X}|]+\kappa^{3 / 2} \mathbb{P}(|\widetilde{X}-\widetilde{Y}|>\kappa / 2) \\
& \qquad\leq 4 \kappa^{1 / 2} \mathbb{E}[|\widetilde{Y}-\widetilde{X}|].
\end{align*}
Using the first result, the third claim is also obtained.
Additional Details on Numerical Results
This section presents supplementary results complementing those in Section (ref). We begin by stating and proving a lemma related to the confidence set $\widehat{\mathrm{CI}}^{\mathrm{SS+U}}_{N,\alpha}$. To recall, the set $\widehat{\mathrm{CI}}^{\mathrm{SS+U}}_{N,\alpha}$ is given by
align*[align* omitted — 283 chars of source]
lemmaLet $\overline{X}_n$ be the sample mean of $X_i$ and let $\widehat \theta_1$ be an estimator computed from an independent set. It then follows that
\begin{align*}
\widehat{\mathrm{CI}}^{\mathrm{SS+U}}_{N,\alpha} \subseteq B_d(\widehat \theta_1; h)\quad almost surely
\end{align*}
with $h$ given by
\begin{align*}
h = n^{-1/2} z_\alpha \lambda_{\max}^{1/2}(\widehat\Sigma) + \lambda_{\max}^{1/2}((\widehat \theta_1-\overline{X}_n)(\widehat \theta_1-\overline{X}_n)^\top)
\end{align*}
where $\widehat\Sigma$ is the sample covariance matrix of $X_i$.
proof[{Proof of Lemma (ref)}]
We rearrange to simplify the expression for $\widehat{\mathrm{CI}}^{\mathrm{SS+U}}_{N,\alpha}$. First, we observe that
\begin{align*}
\|X_i - \theta\|^2 -\|X_i -\widehat\theta_1\|^2 &= \|X_i-\overline{X}_n+\overline{X}_n - \theta\|^2 -\|X_i-\overline{X}_n+\overline{X}_n - \widehat\theta_1\|^2 \\
&= \|X_i-\overline{X}_n\|^2 + 2(X_i-\overline{X}_n)^\top (\overline{X}_n - \theta) + \|\overline{X}_n - \theta\|^2\\
&\qquad -\left(\|X_i-\overline{X}_n\|^2 + 2(X_i-\overline{X}_n)^\top (\overline{X}_n - \widehat\theta_1) + \|\overline{X}_n - \widehat\theta_1\|^2\right)\\
&= 2(X_i-\overline{X}_n)^\top (\widehat\theta_1- \theta) + \|\overline{X}_n - \theta\|^2-\|\overline{X}_n - \widehat\theta_1\|^2.
\end{align*}
Next, we have
\begin{align*}
n^{-1}\sum_{i=1}^n (\|X_i - \theta\|^2 -\|X_i -\widehat\theta_1\|^2) &= \|\overline{X}_n - \theta\|^2-\|\overline{X}_n - \widehat\theta_1\|^2 \\
&= \|\overline{X}_n -\widehat\theta_1+\widehat\theta_1- \theta\|^2-\|\overline{X}_n - \widehat\theta_1\|^2 \\
&= 2(\overline{X}_n -\widehat\theta_1)^\top (\widehat\theta_1- \theta) + \|\widehat\theta_1- \theta\|^2.
\end{align*}
Finally,
\begin{align*}
\widehat \sigma_{\theta, \widehat \theta_1} = \frac{1}{n-1}\sum_{i=1}^n 4(\widehat\theta_1- \theta)^\top (X_i-\overline{X}_n)(X_i-\overline{X}_n)^\top (\widehat\theta_1- \theta) = 4 \|\widehat\theta_1- \theta\|^2_{\widehat \Sigma}.
\end{align*}
In summary, we have obtained that
\begin{align}
\widehat{\mathrm{CI}}^{\mathrm{SS+U}}_{N,\alpha} &= \left\{\theta \in \mathbb{R}^d \, : n^{-1}\sum_{i=1}^n \|X_i -\theta\|^2 - \|X_i -\widehat \theta_1\|^2 \le n^{1/2} z_\alpha \widehat \sigma_{\theta, \widehat \theta_1}-\|\widehat{\theta}_1 - \theta\|^2 \right\}\nonumber\\
&= \left\{\theta \in \mathbb{R}^d \, : 2(\overline{X}_n -\widehat\theta_1)^\top (\widehat\theta_1- \theta) + \|\widehat\theta_1- \theta\|^2 \le 2n^{1/2} z_\alpha \|\widehat\theta_1- \theta\|_{\widehat \Sigma}-\|\widehat{\theta}_1 - \theta\|^2 \right\}\nonumber\\
&= \left\{\theta \in \mathbb{R}^d \, : (\overline{X}_n -\widehat\theta_1)^\top (\widehat\theta_1- \theta) + \|\widehat{\theta}_1 - \theta\|^2 \le n^{1/2} z_\alpha \|\widehat\theta_1- \theta\|_{\widehat \Sigma} \right\}\nonumber\\
&= \left\{\theta \in \mathbb{R}^d \, : (\overline{X}_n -\theta)^\top (\widehat\theta_1- \theta) \le n^{1/2} z_\alpha \|\widehat\theta_1- \theta\|_{\widehat \Sigma} \right\}.
\end{align}
Let $\theta = \widehat \theta_1 + hu$ where $u \in \mathbb{S}^{d-1}$. We derive the upper bound on $|h|$ such that (ref) holds uniformly in all $u \in \mathbb{S}^{d-1}$. By plugging in, we obtain
\begin{align*}
&(\theta-\overline{X}_n)^\top (\theta-\widehat\theta_1) \le n^{-1/2} z_\alpha \|\widehat\theta_1- \theta\|_{\widehat \Sigma} \\
&\qquad \Leftrightarrow (\widehat \theta_1 + hu-\overline{X}_n)^\top hu\le n^{-1/2} z_\alpha |h| (u^\top \widehat \Sigma u)^{1/2}\\
&\qquad \Leftrightarrow h^2 + h(\widehat \theta_1-\overline{X}_n)^\top u\le n^{-1/2} z_\alpha |h| (u^\top \widehat \Sigma u)^{1/2}\\
&\qquad \Rightarrow h^2 \le n^{-1/2} z_\alpha |h| \sup_{u \in \mathbb{S}^{d-1}}\,(u^\top \widehat \Sigma u)^{1/2} + |h|\sup_{u \in \mathbb{S}^{d-1}}\,(u^\top (\widehat \theta_1-\overline{X}_n)(\widehat \theta_1-\overline{X}_n)^\top u)^{1/2}\\
&\qquad \Leftrightarrow |h| \le n^{-1/2} z_\alpha \lambda_{\max}^{1/2}(\widehat\Sigma) + \lambda_{\max}^{1/2}((\widehat \theta_1-\overline{X}_n)(\widehat \theta_1-\overline{X}_n)^\top).
\end{align*}
This conclude the claim.
We provide a short result to derive the semi-axes of an ellipsoid given by the Wald interval based on the asymptotic distribution. To recall, the confidence set is defined as
align*[align* omitted — 223 chars of source]
lemmaThe volume of the Wald interval $\widehat{\mathrm{CI}}^{\mathrm{asympt}}_{N,\alpha}$ is given by
\begin{align*}
\mathrm{Vol}(\widehat{\mathrm{CI}}^{\mathrm{asympt}}_{N,\alpha}) = \frac{\pi^{d/2}}{\Gamma(d/2+1)} \prod_{i=1}^n r_i \quad where \quad r_k = \sqrt{N^{-1}\chi^2_{d,\alpha}\lambda_{k}(\widehat \Sigma_N)}
\end{align*}
and $\lambda_k(A)$ is the $k$th smallest eigenvalue of the matrix $A$.
proof[{Proof of Lemma (ref)}]
Let $\theta = \overline X_N + u$. Then $\theta \in \widehat{\mathrm{CI}}^{\mathrm{asympt}}_{N,\alpha}$ if and only if
\begin{align*}
u^\top \widehat\Sigma_N^{-1} u \le N^{-1} \chi^2_{d,\alpha} \Leftrightarrow u^\top v \Lambda v^\top u \le 1
\end{align*}
where $v$ is the eigenvectors of $\widehat\Sigma_N^{-1}$ and
\begin{align*}
\Lambda = \begin{bmatrix}N (\chi^2_{d,\alpha})^{-1}\lambda_1(\widehat\Sigma_N^{-1}) & &0 \\ & \ddots & \\0 & & N (\chi^2_{d,\alpha})^{-1}\lambda_d(\widehat\Sigma_N^{-1})\end{bmatrix}.
\end{align*}
Since the volume of the ellipsoid does not change under rotational transformation, this concludes the result.
Finally, we provide a qualitative illustration of different confidence sets by visualizing the confidence regions in a bivariate setting. We generate $X_1, \ldots, X_{100}$ from a bivariate normal distribution $N(0, \Sigma)$ where
align*[align* omitted — 92 chars of source]
Figure (ref) displays the confidence regions three methods defined in Section (ref) for confidence levels of $95\%$, $85\%$, and $75\%$. This visualization highlights key qualitative differences between the methods. As shown, the proposed confidence sets (middle and right panels) are non-convex, whereas the Wald interval (left panel) forms an ellipse. Additionally, the proposed methods yield slighly enlarged confidence regions. Figure (ref) estimates that the confidence set in the right panel does not enlarge the Wald interval in the left panel by more than a factor of 2. We observe that this bound is indeed a moderate estimate, as the actual confidence sets may be much closer in size. A close inspection of the confidence set presented in the right panel, given by expression (ref), reveals that it coincides with the adaptive confidence sets by Robins2006 and dimension-agnostic confidence sets based on cross U-statistics proposed by kim2020dimension in their Appendix D (See the left panel of their Figure 3), though they did not provide a rigorous analysis. For mean estimation, these methods become identical.
figure[figure omitted — 570 chars of source]