EconBase
← Back to paper

Bootstrap-Assisted Inference for Generalized Grenander-type Estimators

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

175,670 characters · 28 sections · 47 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Supplement to “Bootstrap-Assisted Inference for Generalized Grenander-type Estimators”

frontmatter\runtitle{Bootstrap-Assisted Inference for Monotone Estimators} \begin{aug} , \and \address[A]{Department of Operations Research and Financial Engineering, Princeton University\printead[presep={,\ }]{e1}} \address[B]{Department of Economics, University of California at Berkeley\printead[presep={,\ }]{e2}} \address[C]{Department of Economics, University of Warwick\printead[presep={,\ }]{e3}} \end{aug} \begin{keyword}[class=MSC] \kwd[Primary ]{62G09} \kwd{62G20} \kwd[; secondary ]{62G07} \kwd{62G08} \end{keyword} \begin{keyword} \kwd{Monotone estimation} \kwd{bootstrapping} \kwd{robust inference} \end{keyword}

Proofs

Proof of Lemma (ref)

Lemma 4.2 of vanderVaart-vanderLaan_2006_IJB implies that, on $(l,u)$,

equation*[equation* omitted — 124 chars of source]

Since $\Phi(\mathsf{x}) \in (l,u)$, we therefore have

equation*[equation* omitted — 217 chars of source]

so Lemma 4.1 of vanderVaart-vanderLaan_2006_IJB implies that

equation*[equation* omitted — 90 chars of source]

where

equation*[equation* omitted — 132 chars of source]

Suppose $y^{\star} \in \Phi(I)$. Then $y^{\star} = \Phi(x^{\star})$, where

equation*[equation* omitted — 163 chars of source]

In particular,

equation*[equation* omitted — 164 chars of source]

where, in fact,

equation*[equation* omitted — 142 chars of source]

because if

equation*[equation* omitted — 145 chars of source]

then, contradicting the definition of $y^{\star}$, we have

equation*[equation* omitted — 175 chars of source]

The proof can therefore be completed by showing that $y^{\star}\in \Phi(I)$.

If $\Phi(I) \supseteq [l,u]$, then there is nothing to show, so suppose $\Phi(I) \not\supseteq [l,u]$. If $y \in \Phi(I)^c \cap [l,u]$, then, since $\Phi(I) \cap [l,u]$ is closed and $l,u \in \Phi(I)$, we have $[y-\eta,y+\eta] \cap \Phi(I) = \emptyset$ for some $\eta > 0$ with $[y-\eta,y+\eta] \subset [l,u]$. Therefore, the function $\mathsf{LSC} (\Gamma\circ\Phi^{-})$ is constant on the interval $[y-\eta/2,y+\eta/2]$, implying in particular that

equation*[equation* omitted — 149 chars of source]

Proof of Lemma (ref)

$\mathbbm{G}(v)$ is a Gaussian process with $\mathbbm{E}[\mathbbm{G}(v)]=\mu(v)$ for all $v\in \mathbb{R}$ and $\mathbbm{C}\mathrm{ov}[\mathbbm{G}(v),\mathbbm{G}(v')] = C \min\{|v|,|v'|\}\mathbbm{1}(\operatorname*{sign}(v)=\operatorname*{sign}(v')) =: \mathcal{K}(v,v')$ for some positive constant $C$ and for all $v,v' \in \mathbb{R}$. The form of the covariance kernel $\mathcal{K}$ implies that for every $\tau> 0$ and every $v,v' \in \mathbb{R}$,

equation*[equation* omitted — 73 chars of source]

and \[\mathcal{K}(v + v',v + v') - 2 \mathcal{K}(v + v',v') + \mathcal{K}(v',v') = \mathcal{K}(v,v).\] In addition, $\mathcal{K}(1,1) > 0$ and $\lim_{v \downarrow 0}\mathcal{K}(1,v) / \sqrt{v} = 0$.\footnote{As a side note, if the covariance kernel satisfies the two displays, then the Gaussian process is two-sided Brownian motion, up to mean shift and scaling by a positive constant.}

We begin by adapting the arguments of Kim-Pollard_1990_AoS to show that a maximizer of $\mathbbm{G}(v)$ over $v \in \mathbb{R}$ exists and is unique with probability one. Let $\mathbbm{G}^{\mu}(v) = \mathbbm{G}(v) - \mu(v)$ be the centered process and suppose that, for the same $c>1/2$ as in the hypothesis,

equation[equation omitted — 176 chars of source]

Then, with probability one, $\mathbbm{G}(v) \to -\infty$ as $|v| \to \infty$, implying in turn that a maximizer of $\mathbbm{G}(v)$ exists (because sample paths are continuous). Also, since

equation*[equation* omitted — 156 chars of source]

for $v \neq v'$, Lemma 2.6 of Kim-Pollard_1990_AoS implies that this maximizer is unique with probability one. In turn, (ref) follows from the Borel-Cantelli lemma because

align*[align* omitted — 547 chars of source]

where the equality uses the rescaling property $\mathcal{K}(v\tau,v'\tau) = \tau\mathcal{K}(v,v')$ and where the last inequality uses Jain-Marcus_1978_BookCh.

To show continuity of the function $x \mapsto \mathbbm{P}[\operatorname*{argmax}_{v \in \mathbb{R}}\{\mathbbm{G}(v)\} \leq x]$, it suffices to show that $\mathbbm{P}[\operatorname*{argmax}_{v \in \mathbb{R}}\{\mathbbm{G}(v)\} = x] = 0$ for every $x \in \mathbb{R}$. Fix $x \in \mathbb{R}$ and define $\widetilde{\mathbbm{G}}_x(0) = 0$ and

equation*[equation* omitted — 136 chars of source]

Then $\max_{v \in \mathbb{R}}\widetilde{\mathbbm{G}}_x(v) \geq 0$ and, for any set $\mathcal{V} \subset \mathbb{R}$,

equation[equation omitted — 340 chars of source]

In the sequel, we show that the majorant in (ref) can be made arbitrarily small by choice of $\mathcal{V}$. In particular, for $\varepsilon \in (0,1)$, we construct $v_{\varepsilon} > 0$ such that

equation[equation omitted — 177 chars of source]

and

equation[equation omitted — 248 chars of source]

For any $N \in \mathbb{N}$, defining $\mathcal{V}_{\varepsilon,N} = \{v_{\varepsilon},\dots,v_{\varepsilon}^N\}$, we therefore have

equation*[equation* omitted — 338 chars of source]

where the second inequality uses the fact that convergence of means and covariances of normal random vectors implies convergence in distribution. Since $N$ is arbitrary, the left-hand side in the preceding display is zero. Letting $\varepsilon \in (0,1)$ be given, the proof can therefore be completed by exhibiting $v_{\varepsilon}$ satisfying (ref) and (ref).

Because $\mathcal{K}(\tau,\tau) = \tau \mathcal{K}(1,1)$ and $\lim_{\tau \downarrow 0}[\mu(x + \tau) - \mu(x)]/\sqrt{\tau} = 0$, there exists $\bar{\tau}_\varepsilon^{'} \in (0,1)$ such that

equation*[equation* omitted — 235 chars of source]

for every $\tau \in (0,\bar{\tau}_\varepsilon^{'}]$. Also, because $\mathcal{K}(\tau_i,\tau_j) = \tau_i\mathcal{K}(1,\tau_j/\tau_i)$ and $\lim_{\tau \downarrow 0}\mathcal{K}(1,\tau )/\sqrt{\tau} = 0$, there exists $\bar{\tau}_\varepsilon^{''} \in (0,1)$ such that

equation*[equation* omitted — 330 chars of source]

for all $\tau_i,\tau_j > 0$ with $\tau_j/\tau_i \leq \bar{\tau}_\varepsilon^{''}$. If $v_{\varepsilon} \in (0, \bar{\tau}_\varepsilon^{'} \land \bar{\tau}_\varepsilon^{''})$, then (ref) and (ref) are satisfied.\qed

Technical Lemmas

In preparation for the proof of Theorem (ref), this section presents six technical lemmas. The first lemma is a switching lemma, which will be used when characterizing the limiting distributions obtained in Theorem (ref).

lemmaLet $\Gamma: \mathbb{R} \to \mathbb{R}$ be a lower semi-continuous function that is bounded from below and satisfies $\lim_{|v| \to \infty}\Gamma(v) / |v| = \infty$. Then, for any $\mathsf{x},t \in \mathbb{R}$, \begin{equation*} \partial_{-} \mathsf{GCM}_{\mathbb{R}}(\Gamma)(\mathsf{x}) > t \qquad \iff \qquad \max \operatorname*{argmax}_{v \in \mathbb{R}} \left\{ vt - \Gamma(v) \right\} < \mathsf{x}. \end{equation*}
proofBecause $\lim_{|v| \to \infty}\Gamma(v) / |v| = \infty$, there exists a $K > |\mathsf{x}|$ such that if $|v| \geq K$, then $vt - \Gamma(v) < -\Gamma(0)$, implying in particular that \begin{equation*} \operatorname*{argmax}_{v \in \mathbb{R}} \left\{ vt - \Gamma(v) \right\} = \operatorname*{argmax}_{v \in [-c,c]} \left\{vt - \Gamma(v) \right\} \end{equation*} for every $c \geq K$. Also, by Lemma A.1.\ of Sen-Banerjee-Woodroofe_2010_AoS there exists a $c > K$ such that $\mathsf{GCM}_{\mathbb{R}}(\Gamma) = \mathsf{GCM}_{[-c,c]}(\Gamma)$ on $[-K,K]$, implying in particular that \begin{equation*} \partial_{-}\mathsf{GCM}_{\mathbb{R}}(\Gamma)(\mathsf{x}) = \partial_{-}\mathsf{GCM}_{[-c,c]}(\Gamma)(\mathsf{x}). \end{equation*} For any such $c$, the conclusion of the lemma is equivalent to the statement \begin{equation*} \partial_{-}\mathsf{GCM}_{[-c,c]}(\Gamma)(\mathsf{x}) > t \qquad \iff \qquad \max \operatorname*{argmax}_{v \in [-c,c]}\left\{vt - \Gamma(v)\right\} < \mathsf{x}, \end{equation*} whose validity follows from Lemma 4.1 of vanderVaart-vanderLaan_2006_IJB.

The proof of Theorem (ref) furthermore employs various approximations to functionals of the form $\mathsf{LSC}_\Phi(f)$. The approximations in question are obtained using Lemmas (ref), (ref), (ref), and (ref). In all cases, the approximations are based on the representation

equation[equation omitted — 165 chars of source]

where $\mathcal{X}_\Phi^\epsilon(x) = \left(\Phi^-(\Phi(x) - \epsilon),\Phi^-(\Phi(x))\right] \cup \left(\Phi^-(\Phi(x)+),\Phi^-(\Phi(x) + \epsilon) \right)$.

The following lemma uses ((ref)) and the special structure of $\Gamma_0$ to obtain a simple “global” bound on the error of the approximation $\mathsf{LSC}_{\Phi}(\Gamma_0) \approx \Gamma_0$.

lemmaSuppose Assumption (ref) holds and suppose $\Phi$ is non-decreasing and right-continuous on $I$. Then, for every $x \in I$, \begin{equation*} |\mathsf{LSC}_{\Phi} (\Gamma_0)(x) - \Gamma_0(x)| \leq 2 \left( \sup_{x' \in I} |\theta_0 (x')| \right) \sup_{x' \in I} |\Phi(x') - \Phi_0 (x')|. \end{equation*}
proofBy ((ref)) and continuity of $\Gamma_0$, \begin{equation*} \mathsf{LSC}_{\Phi} (\Gamma_0)(x) = \min \left[ \Gamma_0(\Phi^-(\Phi(x))),\Gamma_0(\Phi^-(\Phi(x)+)) \right], \end{equation*} while, by Assumption (ref), \begin{equation*} |\Gamma_0(x')-\Gamma_0(x)| \leq \left( \sup_{x” \in I} |\theta_0 (x”)| \right) |\Phi_0(x') - \Phi_0(x)|. \end{equation*} Now, using $\Phi\circ\Phi^-\circ\Phi = \Phi$, \begin{align*} |\Phi_0(\Phi^-(\Phi(x))) - \Phi_0(x)| &= |\Phi_0(\Phi^-(\Phi(x))) - \Phi(\Phi^-(\Phi(x))) + \Phi(x) - \Phi_0(x)| \\ &\leq 2 \sup_{x' \in I}|\Phi(x') - \Phi_0 (x')|. \end{align*} Also, \begin{equation*} |\Phi_0(\Phi^-(\Phi(x)+)) - \Phi_0(x)| \leq 2 \sup_{x' \in I} |\Phi(x') - \Phi_0 (x')| \end{equation*} because, for every $\eta > 0$, \begin{align*} 0 &\leq \Phi_0(\Phi^-(\Phi(x)+)) - \Phi_0(x) \leq \Phi_0(\Phi^-(\Phi(x) + \eta)) - \Phi_0(x) \\ &\leq \Phi_0\left( \Phi^- \left( \Phi_0(x) + \sup_{x' \in I} |\Phi(x') - \Phi_0 (x')| + \eta \right) \right) - \Phi_0(x) \\ &\leq \Phi_0 \left( \Phi_0^- \left(\Phi_0(x)+2\sup_{x' \in I} |\Phi(x') - \Phi_0 (x')| + \eta \right) \right) - \Phi_0(x) \\ &\leq 2 \sup_{x' \in I} |\Phi(x') - \Phi_0 (x')| + \eta, \end{align*} where the last inequality uses continuity of $\Phi_0$.

A simple “global” bound on the error of the approximation $\mathsf{LSC}_{\Phi}(f) \approx f$ is available also in the important special case where $f$ is proportional to $\Phi$.

lemmaSuppose Assumption (ref) holds and suppose $\Phi$ is non-decreasing and right-continuous on $I$. Then, for every $x \in I$ and every $\theta \in \mathbb{R}$, \begin{equation*} |\mathsf{LSC}_{\Phi}(\theta \Phi)(x) - \theta \Phi(x)| \leq |\theta| \sup_{x' \in I} |\Phi(x') - \Phi(x'-)|. \end{equation*}
proofFirst, if $\theta > 0$, then the result follows from the fact that, by ((ref)), \begin{equation*} \mathsf{LSC}_{\Phi}(\theta \Phi)(x) = \theta \Phi(\Phi^-(\Phi(x))-) \leq \theta \Phi(\Phi^-(\Phi(x))) = \theta \Phi(x), \end{equation*} where \begin{equation*} \Phi(x) = \Phi(\Phi^-(\Phi(x))) \leq \Phi(\Phi^-(\Phi(x))-) + \sup_{x' \in I} |\Phi(x') - \Phi(x'-)|. \end{equation*} Next, if $\theta < 0$, then the result follows from the fact that, by ((ref)), \begin{equation*} \mathsf{LSC}_{\Phi}(\theta \Phi)(x) = \theta \Phi(\Phi^-(\Phi(x)+)) \leq \theta \Phi(\Phi^-(\Phi(x))) = \theta \Phi(x), \end{equation*} where \begin{align*} \Phi(\Phi^-(\Phi(x)+)) &\leq \liminf_{\eta \downarrow 0} \Phi(\Phi^-(\Phi(x) + \eta)-) + \sup_{x' \in I}|\Phi(x') - \Phi(x'-)| \\ &\leq \Phi(x)+ \sup_{x' \in I}|\Phi(x') - \Phi(x'-)|. \end{align*}

Next, we give a “local” approximation to $\Phi^- \circ \Phi$. That approximation will later be used in combination with ((ref)) to obtain “local” approximations to $\mathsf{LSC}_\Phi(f)$, but the approximation is also useful in its own right and we therefore state it as a separate lemma.

lemmaSuppose Assumption (ref) holds and suppose $\Phi$ is non-decreasing and right-continuous on $I$. Also, suppose \begin{equation*} \frac{2 \sup_{x' \in I_{\mathsf{x}}^{\delta + \epsilon}} |\Phi(x') - \Phi_0(x')|}{\inf_{x' \in I_{\mathsf{x}}^{\delta + \epsilon}}\partial\Phi_0(x')} < \epsilon \end{equation*} for some $\delta,\epsilon > 0$ with $I_{\mathsf{x}}^{\delta + \epsilon} \subseteq I$. Then, for every $x \in I_{\mathsf{x}}^{\delta}$, \begin{equation*} |\Phi^-(\Phi(x)) - x| < \epsilon \qquad and \qquad |\Phi^-(\Phi(x)+) - x| < \epsilon. \end{equation*}
proofFirst, suppose $|\Phi^-(\Phi(x))-x| \geq \epsilon$. Then $\Phi^-(\Phi(x)) \leq x - \epsilon$, implying in particular that $\Phi(x-\epsilon) = \Phi(x)$ and therefore also \begin{equation*} \left[\Phi(x - \epsilon) - \Phi_0(x - \epsilon)\right] - \left[\Phi(x) - \Phi_0(x)\right] = \Phi_0(x) - \Phi_0(x - \epsilon). \end{equation*} Now, if $x \in I_{\mathsf{x}}^{\delta}$, then $[x-\epsilon,x] \subseteq I_{\mathsf{x}}^{\delta + \epsilon}$, so \begin{equation*} \Phi_0(x) - \Phi_0(x - \epsilon) \geq \epsilon \inf_{x' \in I_{\mathsf{x}}^{\delta + \epsilon}} \partial \Phi_0(x') > 2 \sup_{x' \in I_{\mathsf{x}}^{\delta + \epsilon}}|\Phi(x') - \Phi_0(x')|, \end{equation*} whereas \begin{equation*} \left| \left[ \Phi(x - \epsilon) - \Phi_0(x - \epsilon) \right] - \left[ \Phi(x) - \Phi_0(x) \right] \right| \leq 2 \sup_{x' \in I_{\mathsf{x}}^{\delta + \epsilon}} |\Phi(x') - \Phi_0(x')|. \end{equation*} In other words, $x \not \in I_{\mathsf{x}}^{\delta}$. Next, suppose $|\Phi^-(\Phi(x)+) - x| \geq \epsilon$. Then, for every $\eta,\eta' > 0$, $\Phi^-(\Phi(x) + \eta) \geq x + \epsilon$ and therefore \begin{equation*} \Phi(x + \epsilon - \eta')-\sup_{x' \in I_{\mathsf{x}}^{\delta + \epsilon}}|\Phi(x') - \Phi_0(x')| - \eta < \Phi(x)-\sup_{x' \in I_{\mathsf{x}}^{\delta + \epsilon}}|\Phi(x') - \Phi_0(x')|, \end{equation*} where, for $x \in I_{\mathsf{x}}^{\delta}$, \begin{equation*} \Phi(x) - \sup_{x' \in I_{\mathsf{x}}^{\delta + \epsilon}}|\Phi(x') - \Phi_0(x')| \leq \Phi_0(x), \end{equation*} whereas \begin{align*} &\liminf_{\eta,\eta' \downarrow 0} \left[ \Phi(x + \epsilon - \eta') - \sup_{x' \in I_{\mathsf{x}}^{\delta + \epsilon}} |\Phi(x') - \Phi_0(x')| - \eta \right] \\ &\geq \liminf_{\eta' \downarrow 0} \left[ \Phi_0(x + \epsilon - \eta') - 2 \sup_{x' \in I_{\mathsf{x}}^{\delta + \epsilon}} |\Phi(x') - \Phi_0(x')| \right] \\ &\geq \Phi_0(x) + \left( \inf_{x' \in I_{\mathsf{x}}^{\delta + \epsilon}} \partial \Phi_0(x') \right) \liminf_{\eta' \downarrow 0} \left[ \epsilon - \frac{2 \sup_{x' \in I_{\mathsf{x}}^{\delta + \epsilon}} |\Phi(x') - \Phi_0(x')|} {\inf_{x' \in I_{\mathsf{x}}^{\delta + \epsilon}} \partial \Phi_0 x')} - \eta' \right] > \Phi_0(x). \end{align*} In other words, $x \not \in I_{\mathsf{x}}^{\delta}$.

Next, we obtain two “local” approximations to $\mathsf{LSC}_\Phi(f)$. The first of these is a generic approximation obtained by simply combining ((ref)) and Lemma (ref), but for later reference we state the result as a separate lemma.

lemmaSuppose the assumptions of Lemma (ref) hold. Then, for every $x \in I_{\mathsf{x}}^{\delta}$ and every $f : I \to \mathbb{R}$, \begin{equation*} \inf_{|x' - x| \leq \epsilon}f(x') \leq \mathsf{LSC}_{\Phi}(f)(x) \leq \sup_{|x' - x| \leq \epsilon}f(x') \end{equation*} and \begin{equation*} |\mathsf{LSC}_{\Phi}(f)(x) - f(x)| \leq \sup_{|x' - x| \leq \epsilon} |f(x') - f(x)|. \end{equation*}

The final lemma is concerned with the special case where $f$ is proportional to $\Phi$. In that case, the following “local” analog of Lemma (ref) shows that the bound(s) obtained in Lemma (ref) can be improved under mild conditions on $\Phi$.

lemmaSuppose the assumptions of Lemma (ref) hold. Then, for every $x \in I_{\mathsf{x}}^{\delta}$ and every $\theta \in \mathbb{R}$, \begin{equation*} |\mathsf{LSC}_{\Phi} (\theta \Phi) (x) - \theta \Phi(x)| \leq |\theta| \sup_{x' \in I_{\mathsf{x}}^{\delta + \epsilon}}|\Phi(x') - \Phi(x'-)|. \end{equation*}
proofFirst, if $\theta > 0$, then the result follows from the fact that, by ((ref)), \begin{equation*} \mathsf{LSC}_{\Phi} (\theta \Phi) (x) = \theta \Phi(\Phi^-(\Phi(x))-) \leq \theta \Phi(\Phi^-(\Phi(x))) = \theta \Phi(x), \end{equation*} where, if $x \in I_{\mathsf{x}}^{\delta}$, then $\Phi^-(\Phi(x)) \in I_{\mathsf{x}}^{\delta + \epsilon}$ by Lemma (ref), and therefore \begin{equation*} \Phi(x) = \Phi(\Phi^-(\Phi(x))) \leq \Phi(\Phi^-(\Phi(x))-) + \sup_{x' \in I_{\mathsf{x}}^{\delta + \epsilon}}|\Phi(x') - \Phi(x'-)|. \end{equation*} Next, if $\theta < 0$, then the result follows from the fact that, by ((ref)), \begin{equation*} \mathsf{LSC}_{\Phi} (\theta \Phi)(x) = \theta \Phi(\Phi^-(\Phi(x)+)) \leq \theta \Phi(\Phi^-(\Phi(x))) = \theta \Phi(x), \end{equation*} where, if $x \in I_{\mathsf{x}}^{\delta}$, then $\Phi^-(\Phi(x)+) \in I_{\mathsf{x}}^{\delta + \epsilon}$ by Lemma (ref), and therefore \begin{align*} \Phi(\Phi^-(\Phi(x)+)) &\leq \liminf_{\eta \downarrow 0} \Phi(\Phi^-(\Phi(x) + \eta)-) + \sup_{x' \in I_{\mathsf{x}}^{\delta + \epsilon}}|\Phi(x') - \Phi(x'-)| \\ &\leq \Phi(x)+ \sup_{x' \in I_{\mathsf{x}}^{\delta + \epsilon}}|\Phi(x') - \Phi(x'-)|. \end{align*}

Proof of Theorem (ref)

\paragraph*{Proof of (ref)} Let $t \in \mathbb{R}$ be given. By Lemma (ref) and change of variables,

align*[align* omitted — 618 chars of source]

where

equation*[equation* omitted — 153 chars of source]
equation*[equation* omitted — 368 chars of source]

and

equation*[equation* omitted — 156 chars of source]

with

align*[align* omitted — 812 chars of source]

By (ref) and Lemma (ref), $\widehat{Z}_{\mathsf{x},n}^{\mathfrak{q}} = o_{\mathbbm{P}}(1)$. Suppose also that

equation[equation omitted — 280 chars of source]

where $\mathcal{H}_{\mathsf{x}}^{\mathfrak{q}}(v;t) = \mathcal{G}_{\mathsf{x}}(v) + \mathcal{M}_{\mathsf{x}}^{\mathfrak{q}}(v) - t \partial \Phi_0(\mathsf{x}) v$. Then

align*[align* omitted — 626 chars of source]

where the second line uses Lemma (ref) and where the last equality uses Lemma (ref). The proof of (ref) can therefore be completed by showing (ref).

We shall do so by means of the argmax continuous mapping theorem of Cox_2022. To be specific, using that theorem it can be shown that (ref) holds if

equation[equation omitted — 214 chars of source]

and if

equation[equation omitted — 167 chars of source]

We begin by showing (ref). First, by (ref), $\widehat{G}_{\mathsf{x},n}^{\mathfrak{q}} \leadsto \mathcal{G}_{\mathsf{x}}$. Also, by (ref) and (ref), as $u \to 0$,

align*[align* omitted — 443 chars of source]

where the first equality uses L'H\^{o}pital's rule and

equation*[equation* omitted — 273 chars of source]

As a consequence, $M_{\mathsf{x},n}^{\mathfrak{q}} \leadsto \mathcal{M}_{\mathsf{x}}^{\mathfrak{q}}$. Moreover, $\widehat{L}_{\mathsf{x},n}^{\mathfrak{q}} - L_{\mathsf{x},n}^{\mathfrak{q}} \leadsto 0$ by (ref) and $L_{\mathsf{x},n}^{\mathfrak{q}} \leadsto \mathcal{L}_{\mathsf{x}}$ by (ref), where $L_{\mathsf{x},n}^{\mathfrak{q}}(v) = a_n \left[\Phi_0(\mathsf{x} + v a_n^{-1}) - \Phi_0(\mathsf{x})\right]$. In particular, $\widehat{L}_{\mathsf{x},n}^{\mathfrak{q}} \leadsto \mathcal{L}_{\mathsf{x}}$ and therefore

equation*[equation* omitted — 240 chars of source]

Because $\widehat{G}_{\mathsf{x},n}^{\mathfrak{q}}+M_{\mathsf{x},n}^{\mathfrak{q}}$ is asymptotically equicontinuous,

equation*[equation* omitted — 267 chars of source]

by (ref) and Lemma (ref). Also, by (ref) and Lemma (ref),

equation*[equation* omitted — 251 chars of source]

The result (ref) follows from the three preceding displays and the fact that, on $\widehat{V}_{\mathsf{x},n}^{\mathfrak{q}}$,

align*[align* omitted — 873 chars of source]

Next, to show (ref), we first define $\theta_n(\mathsf{x};t) = \theta_0(\mathsf{x}) + t r_n^{-1}$ and note that

equation*[equation* omitted — 271 chars of source]

Now, if $|\widehat{v}_n(t)| > a_n \delta > 0,$ then

equation*[equation* omitted — 325 chars of source]

where $|\theta_n(\mathsf{x};t) - \theta_0(\mathsf{x})| = O(r_n^{-1}) = o(1)$, and, by (ref),

equation*[equation* omitted — 101 chars of source]

Also, using (ref), (ref), and Lemma (ref),

align*[align* omitted — 582 chars of source]

As a consequence, $\widehat{v}_n(t) = o_{\mathbbm{P}}(a_n)$: For any $\delta > 0$,

align*[align* omitted — 312 chars of source]

where the equality uses the fact, noted by Westling-Carone_2020_AoS, that the function $v \mapsto \theta_0(\mathsf{x}) \Phi_0(v) - \Gamma_0(v)$ is unimodal and maximized at $v = \mathsf{x}$.

Next, defining $\widehat{V}_{\mathsf{x},n}^{\mathfrak{q}}(j) = \{v\in \widehat{V}_{\mathsf{x},n}^{\mathfrak{q}} : 2^j < |v| \leq 2^{j+1}\}$ and using $\widehat{v}_n(t) = o_{\mathbbm{P}}(a_n)$, we have, for any $K$, any positive $\delta'$, and any sequence of events $\{\mathcal{A}_n'\}$ with $\lim_{n \to \infty}\mathbbm{P}[\mathcal{A}_n'] = 1$,

align*[align* omitted — 525 chars of source]

The proof of (ref) can therefore be completed by showing that the majorant side in the display can be made arbitrarily small by choice of $K$, $\delta'$, and $\{\mathcal{A}_n'\}$.

To do so, we begin by analyzing each term in the basic bound

align*[align* omitted — 664 chars of source]

Because $\widehat{H}_{\mathsf{x},n}^{\mathfrak{q}}(0;t) \leadsto \mathcal{H}_{\mathsf{x}}^{\mathfrak{q}}(0;t) = 0$ and because, by (ref) and Lemma (ref), there is a positive $\delta'$ such that

equation*[equation* omitted — 298 chars of source]

we may assume that, on $\{\mathcal{A}_n'\}$ and for some $C_0$,

equation*[equation* omitted — 337 chars of source]

Also, because, by (ref) and (ref), there is a positive $\delta'$ such that

equation*[equation* omitted — 324 chars of source]

we may assume that, on $\{\mathcal{A}_n'\}$ and for some $C_L$,

equation*[equation* omitted — 132 chars of source]

Next, by (ref) and Lemma (ref), with probability approaching one,

equation*[equation* omitted — 239 chars of source]

while, by (ref) and (ref), there is a positive $\delta'$ such that

equation*[equation* omitted — 151 chars of source]

We may therefore assume that, on $\{\mathcal{A}_n'\}$ and for some positive $C_M$,

equation*[equation* omitted — 199 chars of source]

Finally, by (ref) and Lemma (ref), with probability approaching one,

equation*[equation* omitted — 292 chars of source]

and we may therefore assume that, on $\{\mathcal{A}_n'\}$,

equation*[equation* omitted — 365 chars of source]

where $V_\eta(j) = \{v \in \mathbb{R}: \eta^{-1}2^j \leq |v| \leq \eta 2^{j+1}\}$.

As a consequence, by the Markov inequality,

align*[align* omitted — 740 chars of source]

where, by (ref), we may assume that, for some $C_G$,

equation*[equation* omitted — 348 chars of source]

and where, for all sufficiently large $j$,

equation*[equation* omitted — 138 chars of source]

In other words, for large $K$,

equation*[equation* omitted — 359 chars of source]

which can be made arbitrarily small by choice of $K$.

\paragraph*{Proof of (ref)} We proceed as in the proof of (ref). Let $t \in \mathbb{R}$ be given. By Lemma (ref) and change of variables,

align*[align* omitted — 671 chars of source]

where

equation*[equation* omitted — 158 chars of source]
equation*[equation* omitted — 411 chars of source]

and

equation*[equation* omitted — 162 chars of source]

with

align*[align* omitted — 758 chars of source]

By (ref) and Lemma (ref), $\widehat{Z}_{\mathsf{x},n}^{\mathfrak{q},*} = o_{\mathbbm{P}}(1)$. Suppose also that

equation[equation omitted — 308 chars of source]

Then, as in the proof of (ref),

equation*[equation* omitted — 323 chars of source]

The proof of (ref) can therefore be completed by showing (ref).

We shall do so by showing that

equation[equation omitted — 224 chars of source]

and

equation[equation omitted — 187 chars of source]

We begin by showing (ref). First, by (ref), $\widehat{G}_{\mathsf{x},n}^{\mathfrak{q},*} \leadsto_{\mathbbm{P}} \mathcal{G}_{\mathsf{x}}$. Also, by Assumption (ref), $\widetilde{M}_{\mathsf{x},n}^{\mathfrak{q},*} \leadsto_{\mathbbm{P}} \mathcal{M}_{\mathsf{x}}^{\mathfrak{q}}$. Moreover, $\widehat{L}_{\mathsf{x},n}^{\mathfrak{q},*} - \widehat{L}_{\mathsf{x},n}^{\mathfrak{q}} \leadsto_{\mathbbm{P}} 0$ by (ref), where, as shown in the proof of ((ref)), $\widehat{L}_{\mathsf{x},n}^{\mathfrak{q}} \leadsto \mathcal{L}_{\mathsf{x}}$. In particular, $\widehat{L}_{\mathsf{x},n}^{\mathfrak{q},*} \leadsto_{\mathbbm{P}} \mathcal{L}_{\mathsf{x}}$ and therefore

equation*[equation* omitted — 293 chars of source]

Because $\widehat{G}_{\mathsf{x},n}^{\mathfrak{q},*} + \widetilde{M}_{\mathsf{x},n}^{\mathfrak{q}}$ is asymptotically equicontinuous,

equation*[equation* omitted — 316 chars of source]

by (ref) and Lemma (ref). Also, by (ref) and Lemma (ref),

equation*[equation* omitted — 296 chars of source]

The result (ref) follows from the three preceding displays and the fact that, on $\widehat{V}_{\mathsf{x},n}^{\mathfrak{q},*}$,

align*[align* omitted — 1,013 chars of source]

Next, to show (ref), we first define $\widehat{\theta}_n(\mathsf{x};t) = \widehat{\theta}_n(\mathsf{x}) + t r_n^{-1}$ and note that

equation*[equation* omitted — 300 chars of source]

Now, if $|\widehat{v}_n^*(t)| > a_n \delta > 0,$ then

equation*[equation* omitted — 368 chars of source]

where $|\widehat{\theta}_n(\mathsf{x};t) - \theta_0(\mathsf{x})| = O_{\mathbbm{P}}(r_n^{-1}) = o_{\mathbbm{P}}(1)$, and, by (ref),

equation*[equation* omitted — 106 chars of source]

Therefore, defining $\widehat{\mathsf{x}}_n^* = \widehat{\Phi}_n^{*-}(\widehat{\Phi}_n^{*}(\mathsf{x})) = \mathsf{x} + o_{\mathbbm{P}}(1)$ and using ((ref)),

align*[align* omitted — 669 chars of source]

where the last equality uses (ref), $\widehat{Z}_{\mathsf{x},n}^{\mathfrak{q},*} = o_{\mathbbm{P}}(1)$, and Assumption (ref). Also, using (ref), (ref), and Assumption (ref), we have, uniformly in $x \not \in I_\mathsf{x}^\delta$ and for some $c > 0$,

align*[align* omitted — 340 chars of source]

and therefore, by Lemma (ref),

equation*[equation* omitted — 269 chars of source]

As a consequence, $\widehat{v}_n^*(t) = o_{\mathbbm{P}}(a_n)$: For any $\delta > 0$,

equation*[equation* omitted — 174 chars of source]

Next, defining $\widehat{V}_{\mathsf{x},n}^{\mathfrak{q},*}(j) = \{v\in \widehat{V}_{\mathsf{x},n}^{\mathfrak{q},*} : 2^j < |v| \leq 2^{j+1}\}$ and using $\widehat{v}_n^*(t) = o_{\mathbbm{P}}(a_n)$, we have, for any $K$, any positive $\delta'$, and any sequence of events $\{\mathcal{A}_n'\}$ with $\lim_{n \to \infty}\mathbbm{P}[\mathcal{A}_n'] = 1$,

align*[align* omitted — 538 chars of source]

The proof of (ref) can therefore be completed by showing that the majorant side in the display can be made arbitrarily small by choice of $K$, $\delta'$, and $\{\mathcal{A}_n'\}$.

To do so, we begin by analyzing each term in the basic bound

align*[align* omitted — 728 chars of source]

Because $\widehat{H}_{\mathsf{x},n}^{\mathfrak{q},*}(0;t) \leadsto_{\mathbbm{P}} \mathcal{H}_{\mathsf{x}}^{\mathfrak{q}}(0;t) = 0$ and because, by (ref) and Lemma (ref), there is a positive $\delta'$ such that

equation*[equation* omitted — 319 chars of source]

we may assume that, on $\{\mathcal{A}_n'\}$ and for some $C_0$,

equation*[equation* omitted — 365 chars of source]

Also, because, by (ref) and (ref), there is a positive $\delta'$ such that

equation*[equation* omitted — 326 chars of source]

we may assume that, on $\{\mathcal{A}_n'\}$ and for some $C_L$,

equation*[equation* omitted — 137 chars of source]

Next, by (ref) and Lemma (ref), with probability approaching one,

equation*[equation* omitted — 264 chars of source]

while, by Assumption (ref), there is a positive $c$ such that, with probability approaching one,

equation*[equation* omitted — 118 chars of source]

We may therefore assume that, on $\{\mathcal{A}_n'\}$ and for some positive $C_M$,

equation*[equation* omitted — 213 chars of source]

Finally, by (ref) and Lemma (ref), with probability approaching one,

equation*[equation* omitted — 298 chars of source]

and we may therefore assume that, on $\{\mathcal{A}_n'\}$,

equation*[equation* omitted — 373 chars of source]

As a consequence, by the Markov inequality,

align*[align* omitted — 751 chars of source]

where, by (ref), we may assume that, for some $C_G$,

equation*[equation* omitted — 352 chars of source]

and where, for all sufficiently large $j$,

equation*[equation* omitted — 138 chars of source]

In other words, for large $K$,

equation*[equation* omitted — 365 chars of source]

which can be made arbitrarily small by choice of $K$.

\paragraph*{Proof of (ref)} The bootstrap consistency result (ref) follows from (ref), (ref), Polya's theorem, and the fact that, by Lemma (ref), the limiting distribution in (ref) and (ref) has a continuous cdf. \qed

Proof of Lemma (ref)

For the monomial approximation estimator, we have

align*[align* omitted — 1,049 chars of source]

where the second equality uses $\epsilon_n \to 0$ and the last equality uses $n \epsilon_n^{1 + 2 \mathfrak{q}} \to \infty$.

Similarly, for the forward difference estimator, we have

align*[align* omitted — 1,332 chars of source]

Proof of Lemma (ref)

Proceeding as in the proof of Lemma (ref), we have

align*[align* omitted — 1,253 chars of source]

where the second equality uses $\epsilon_n \to 0$ and the defining property of $\{\lambda_j^{\mathtt{BR}}(k) : k = 1,\dots,\underline{\mathfrak{s}}\}$.

The second part of the lemma follows from the fact that if

equation*[equation* omitted — 213 chars of source]

then

equation*[equation* omitted — 166 chars of source]

\qed

Higher-order expansion of the bias-reduced estimator

In addition to the assumptions of Lemma (ref), suppose that $\widehat{R}_{\mathsf{x},n}(1;\eta_n) = O_{\mathbbm{P}}(a_n^{-1/2})$ for $a_n^{-1}\eta_n^{-1}=O(1)$ and that, for some $\delta > 0$, $\theta_0$ is $(\underline{\mathfrak{s}} + 1)$-times continuously differentiable and $\Phi_0$ is $(\underline{\mathfrak{s}} + 2)$-times continuously differentiable on $I_{\mathsf{x}}^{\delta}$. Then, the first term in the stochastic expansion of $\widetilde{\mathcal{D}}_{j,n}^{\mathtt{BR}}(\mathsf{x})$ satisfies

align*[align* omitted — 693 chars of source]

Also, the approximate variance of

equation*[equation* omitted — 166 chars of source]

is

equation*[equation* omitted — 307 chars of source]

Finally, the third term in the stochastic expansion of $\widetilde{\mathcal{D}}_{j,n}^{\mathtt{BR}}(\mathsf{x})$ is asymptotically negligible under the condition that $\widehat{R}_{\mathsf{x},n}(1;\eta_n) = O_{\mathbbm{P}}(a_n^{-1/2})$ for $a_n^{-1}\eta_n^{-1}=O(1)$, while the fourth term exhibits only a higher-order dependence on $\epsilon_n$ (relative to the dependence exhibited by the first two terms).

Proof of Lemma (ref)

We verify that Assumptions (ref) and (ref) imply Assumptions (ref)-(ref). Define \[\bar{\Gamma}_n^*(x) = \frac{1}{n} \sum_{i = 1}^nW_{i,n} \gamma_0(x;\mathbf{Z}_i) \qquad \text{and} \qquad \bar{\Phi}_n^*(x) = \frac{1}{n} \sum_{i = 1}^nW_{i,n} \phi_0(x;\mathbf{Z}_i).\]

Verifying Assumption (ref)

\paragraph*{Non-bootstrap weak convergence} We first prove $\widehat{G}_{\mathsf{x},n}^{\mathfrak{q}}\rightsquigarrow \mathcal{G}_{\mathsf{x}}$. By Assumption (ref)-(ref),

equation*[equation* omitted — 227 chars of source]
equation*[equation* omitted — 218 chars of source]

for each $K>0$, and thus, defining $\bar{\psi}_{\mathsf{x},n}(v;\mathbf{Z}_i) = \sqrt{a_n}\psi_{\mathsf{x}}(va_n^{-1};\mathbf{Z}_i)$,

equation*[equation* omitted — 266 chars of source]

Next, we prove weak convergence of the empirical process indexed by $\{\bar{\psi}_{\mathsf{x},n}(v;\cdot):|v|\leq K\}$ by verifying finite-dimensional weak convergence and stochastic equicontinuity.

Letting $\eta_n = Ka_n^{-1}$,

equation*[equation* omitted — 253 chars of source]

Also, convergence of the covariance kernel is imposed in Assumption (ref). Thus, the Lyapunov central limit theorem implies finite-dimensional weak convergence.

For stochastic equicontinuity, proceeding as in Kim-Pollard_1990_AoS and using $a_n\mathbbm{E}[\bar{D}_{\gamma}^{\eta_n} (\mathbf{Z})^2 + \bar{D}_{\phi}^{\eta_n} (\mathbf{Z})^2]=O(1)$, it suffices to show that

equation*[equation* omitted — 213 chars of source]

for any $\epsilon_n=o(1)$. For any $M>0$, using $|\bar{\psi}_{\mathsf{x},n}(v;\mathbf{z})|\leq \sqrt{a_n}(1+\theta_0(\mathsf{x}))(\bar{D}_{\gamma}^{\eta_n} (\mathbf{z}) + \bar{D}_{\phi}^{\eta_n} (\mathbf{z}) )$

align*[align* omitted — 986 chars of source]

where $\bar{M}=2M(1+|\theta_0(\mathsf{x})|)$, the expectation of the first term after the inequality can be made arbitrarily small by making $M$ large using $a_n \mathbbm{E}[ \bar{D}_{\gamma}^{\eta_n} (\mathbf{Z})^4 + \bar{D}_{\phi}^{\eta_n} (\mathbf{Z})^4] = O(1)$, the second term is $o(1)$ by Assumption (ref), and the third term is $O_{\mathbbm{P}}( \sqrt{a_n/n} )$ by Theorem 4.2 of Pollard_1989_SS.

\paragraph*{Bootstrap weak convergence} We next prove $\widehat{G}_{\mathsf{x},n}^{\mathfrak{q}}\rightsquigarrow_{\mathbbm{P}}\mathcal{G}_{\mathsf{x}}$. As shown below, we have

align[align omitted — 513 chars of source]

for each $K>0$. Therefore, using $\widehat{\theta}_n(\mathsf{x})\to_{\mathbbm{P}}\theta_0(\mathsf{x})$ and the fact that, uniformly over $| v| \leq K$,

equation*[equation* omitted — 266 chars of source]

we obtain

equation*[equation* omitted — 205 chars of source]

where $\widehat{\psi}_{\mathsf{x},n}(v;\mathbf{Z}) = \bar{\psi}_{\mathsf{x},n}(v;\mathbf{Z}) - n^{-1}\sum_{j=1}^n \bar{\psi}_{\mathsf{x},n}(v;\mathbf{Z}_j)$.

To prove finite-dimensional weak convergence, we apply Lemma 3.6.15 of vanderVaart-Wellner_1996_Book. Assumption (ref) implies that $n^{-1}\sum_{i=1}^n(W_{i,n}-1)^2\to_{\mathbbm{P}}1$ and $n^{-1}\max_{1\leq i \leq n}W_{i,n}^2=o_{\mathbbm{P}}(1)$. Since

align*[align* omitted — 394 chars of source]

and $\sup_{|v|\leq\eta } |\psi_{\mathsf{x}}(v;\mathbf{Z})|\leq \bar{D}_{\gamma}^{\eta}(\mathbf{Z}) + |\theta_0(\mathsf{x})|\bar{D}_{\phi}^{\eta}(\mathbf{Z})$, for any $v,u\in\mathbb{R}$,

equation*[equation* omitted — 251 chars of source]

Also, $n^{-1}\sum_{i=1}^n \widehat{\psi}_{\mathsf{x},n}^4(v;\mathbf{Z}_i)=O_{\mathbbm{P}}(1)$, verifying the hypothesis of the lemma.

For stochastic equicontinuity, let $\epsilon_n=o(1)$ and $\eta_n=K a_n^{-1}$. Lemma 3.6.7 of vanderVaart-Wellner_1996_Book implies that for any $n_0 \in \{1,\dots, n\}$, there is a fixed constant $C>0$ such that

align*[align* omitted — 846 chars of source]

where $(R_1,\dots,R_n)$ is uniformly distributed on the set of all permutations of $\{1,\dots,n\}$, independent of $\{\mathbf{Z}_i\}_{i=1}^n$. Choose $n_0$ such that $n^{1/2-1/\mathfrak{r}}/n_0\to \infty$ and $n_0/a_n\to \infty$ (which is possible by $\mathfrak{r}>(4\mathfrak{q}+2)/(2\mathfrak{q}-1)$), and the first term after the inequality in the above display is $o_{\mathbbm{P}}(1)$. For the second term, following the argument of vanderVaart-Wellner_1996_Book, it suffices to bound

equation*[equation* omitted — 295 chars of source]

where $\{\mathbf{Z}_i^*\}_{i=1}^k$ denotes a random sample from the empirical cdf and $\mathbbm{E}_n^*$ is the expectation under this empirical bootstrap law. Following the argument of Kim-Pollard_1990_AoS, it suffices to show that

equation*[equation* omitted — 296 chars of source]

For $k\in \{n_0,\dots,n\}$ and $M>0$,

align*[align* omitted — 699 chars of source]

where the second term after the inequality is shown to be $o_{\mathbbm{P}}(1)$ in the non-bootstrap case above. For the first term after the inequality,

align*[align* omitted — 1,103 chars of source]

where $\bar{M}=2M(1+|\theta_0(\mathsf{x})|)$. The first term after the inequality does not depend on $k$ and its expectation can be made arbitrarily small by taking $M$ sufficiently large. The second term is independent of $k$ and we can handle this term by adding and subtracting the expectation inside the summation. For the third term, applying Theorem 4.2 of Pollard_1989_SS again, it is bounded by a constant multiple of

equation*[equation* omitted — 221 chars of source]

which is $o_{\mathbbm{P}}(1)$ by the choice of $n_0$.

\paragraph*{Verifying (ref)} We focus on the first display. By adding and subtracting the bootstrap means,

align*[align* omitted — 395 chars of source]

where

align*[align* omitted — 384 chars of source]

By Assumption (ref),

equation*[equation* omitted — 214 chars of source]

Identical to above, Lemma 3.6.7 and the argument in Theorem 3.6.13 of vanderVaart-Wellner_1996_Book imply that for some fixed $C>0$,

align*[align* omitted — 580 chars of source]

By Assumption (ref),

align*[align* omitted — 572 chars of source]

Also, Corollary 4.3 of Pollard_1989_SS implies that for some fixed $C>0$,

align*[align* omitted — 390 chars of source]

which is $o_{\mathbbm{P}}(a_n^{-1})$ by Assumption (ref).

Verifying Assumption (ref)

Defining

align*[align* omitted — 403 chars of source]

we have

align*[align* omitted — 484 chars of source]

Assumptions (ref) and (ref) imply that for $V \in [1,a_n \delta]$,

equation*[equation* omitted — 239 chars of source]

and

equation*[equation* omitted — 232 chars of source]

where $\beta = \beta_\gamma \lor \beta_{\phi} , A_n = A_{\gamma,n} \lor A_{\phi,n} = o_{\mathbbm{P}}(1)$, and $B_n = B_{\gamma,n} \lor B_{\phi,n} = o_{\mathbbm{P}}(a_n^{\beta})$. As a consequence, there exists $\eta_n' = o(1)$ such that

equation*[equation* omitted — 252 chars of source]

we take the event in the display to be $\mathcal{A}_n$. Also,

equation*[equation* omitted — 168 chars of source]

and, using Corollary 4.3 of Pollard_1989_SS,

equation*[equation* omitted — 162 chars of source]

Therefore,

equation*[equation* omitted — 184 chars of source]

implying in particular that

equation*[equation* omitted — 185 chars of source]

For the bootstrap counterpart, defining

align*[align* omitted — 437 chars of source]

we have

align[align omitted — 1,259 chars of source]

where bounds for the third and fourth terms were derived in the non-bootstrap case and where the last term is $o_{\mathbbm{P}}(1)$ uniformly over $| v| \leq a_n \delta$ as $\sqrt{a_n}[\widehat{\theta}_n(\mathsf{x}) - \theta_0(\mathsf{x})]=o_{\mathbbm{P}}(1)$ and $\sqrt{n} \sup_{| v|\leq \delta}|\bar{\Phi}_n^*(\mathsf{x} + v) - \bar{\Phi}_n(\mathsf{x} +v ) |=O_{\mathbbm{P}}(1)$. For the first term after the equality in (ref),

align*[align* omitted — 453 chars of source]

where, by Assumption (ref),

equation*[equation* omitted — 225 chars of source]

uniformly over $V\in [1,a_n\delta]$ and where, applying Lemma 3.6.7 of vanderVaart-Wellner_1996_Book, we have, for any $n_0\in \{1,\dots, n\}$,

align*[align* omitted — 590 chars of source]

where $C>0$ is a fixed constant. Letting $n_0$ be a diverging sequence (dependent on $n$) such that $n_0 n^{\mathfrak{r}}\sqrt{a_n/n} =o(1)$, the first term on the majorant side is bounded by

equation*[equation* omitted — 260 chars of source]

and, by Corollary 4.3 of Pollard_1989_SS, the second term is bounded (up to a constant) by

equation*[equation* omitted — 181 chars of source]

which is $o_{\mathbbm{P}}(1)$ by Assumption (ref). As a consequence, there exists a sequence of random variables $A_n'=o_{\mathbbm{P}}(1)$ such that for $V\in [1,a_n\delta]$,

equation*[equation* omitted — 244 chars of source]

By identical arguments, an analogous bound holds for the second term after the inequality in (ref). Therefore, there exists $\eta_n'=o(1)$ and events $\mathcal{A}_n$ such that $\lim_{n\to\infty}\mathbbm{P}[\mathcal{A}_n]=1$ and

equation*[equation* omitted — 226 chars of source]

for any $V\in [1,a_n\delta]$. Finally, proceeding as in the proof of stochastic equicontinuity, we obtain the bound $\mathbbm{E}[\sup_{| v|\in [V,2V] }|\bar{G}_{\mathsf{x},n}^{\mathfrak{q},*}(v)|]\leq C\sqrt{V}$.

Verifying Assumptions (ref)-(ref)

We have

equation*[equation* omitted — 219 chars of source]

where the first term after the inequality is assumed to be $o_{\mathbbm{P}}(1)$ and the second term is $o_{\mathbbm{P}}(1)$ by standard arguments. Similarly, $\sup_{x\in I}|\widehat{\Phi}_n(x)-\Phi_0(x)|=o_{\mathbbm{P}}(1)$. Also,

align*[align* omitted — 421 chars of source]

where the last two terms are $o_{\mathbbm{P}}(a_n^{-1})$ and where Assumption (ref) implies

equation*[equation* omitted — 263 chars of source]

the last equality using $B_{\phi,n}=o_{\mathbbm{P}}(a_n^{\beta_{\phi}})$ and $\beta_{\phi}\leq \mathfrak{q}$.

Now we look at the bootstrap objects. For $\widehat{\Gamma}_n^*$,

equation*[equation* omitted — 320 chars of source]

where the last term is assumed to be $o_{\mathbbm{P}}(1)$ and where

align*[align* omitted — 424 chars of source]

For $\bar{\Gamma}_n^* -\bar{\Gamma}_n$, using Lemma 3.6.7 of vanderVaart-Wellner_1996_Book and the same argument as for verifying Assumption (ref), it suffices to show that

equation*[equation* omitted — 218 chars of source]

as can be done using Corollary 4.3 of Pollard_1989_SS.

For $\widehat{\Phi}_n^*$, $\sup_{x\in I}|\widehat{\Phi}_n^*(x)-\widehat{\Phi}_n(x)|=o_{\mathbbm{P}}(1)$ follows from the same argument as for $\widehat{\Gamma}_n^*$. For $a_n \sup_{x\in I_{\mathsf{x}}^{\delta}}|\widehat{\Phi}_n^*(x) -\widehat{\Phi}_n(x)|=o_{\mathbbm{P}}(1)$,

align*[align* omitted — 481 chars of source]

where the last term is $o_{\mathbbm{P}}(a_n^{-1})$ as shown above and the second and third terms after the inequality are $O_{\mathbbm{P}}(n^{-1/2})$ by standard arguments. For the remaining term,

align*[align* omitted — 462 chars of source]

where $\Breve{\phi}_n(x;\mathbf{Z}) = \widehat{\phi}_n(x;\mathbf{Z})-\phi_0(x;\mathbf{Z}) - [\check{\Phi}_n(x)-\bar{\Phi}_n(x)]$. The last two terms are $o_{\mathbbm{P}}(a_n^{-1})$ by Assumption (ref). Using Lemma 3.6.7 of vanderVaart-Wellner_1996_Book and the argument similar to above, the remaining term is $o_{\mathbbm{P}}(a_n^{-1})$.

Proof of Theorem (ref)

The proof is by contradiction and follows Kosorok_2008_BookCh. We omit some details in cases where the arguments are almost identical to those for Theorem (ref) and Lemma (ref).

Suppose that the bootstrap approximation is consistent; that is, suppose

equation*[equation* omitted — 280 chars of source]

Then, by Theorem 2.2 of Kosorok_2008_BookCh, we have

equation[equation omitted — 160 chars of source]

where $=_d$ denotes the distributional equality, $Y_1$ and $Y_2$ are independent copies of $Y$, and where the convergence in distribution is unconditional.

Using the switching lemma, $\mathbbm{P} \left[r_n \left( \widehat{\theta}_n^*(\mathsf{x}) - \theta_0(\mathsf{x}) \right) > t \right]$ equals

equation*[equation* omitted — 340 chars of source]

By the arguments used in the proof of Theorem (ref), to characterize the limiting distribution of $r_n(\widehat{\theta}_n^*(\mathsf{x}) - \theta_0(\mathsf{x}))$, it suffices to look at

equation*[equation* omitted — 161 chars of source]

where

align*[align* omitted — 448 chars of source]

and

equation*[equation* omitted — 253 chars of source]

It can be shown that $\check{G}_{\mathsf{x},n}^{\mathfrak{q},*} \leadsto_{\mathbbm{P}} \mathcal{G}_{\mathsf{x}}, \widehat{L}_{\mathsf{x},n}^{\mathfrak{q},*} \leadsto \mathcal{L}_{\mathsf{x}}$, and that $\check{M}_{\mathsf{x},n}^{\mathfrak{q}} \leadsto \mathcal{G}_{\mathsf{x}} + \mathcal{M}_{\mathsf{x}}^{\mathfrak{q}}$. Thus,

equation*[equation* omitted — 360 chars of source]

where $\mathcal{G}_{\mathsf{x},1}$ and $\mathcal{G}_{\mathsf{x},2}$ are independent copies of $\mathcal{G}_{\mathsf{x}}$. Noting that $\mathcal{G}_{\mathsf{x}}(a v)=_d \sqrt{|a|} \mathcal{G}_{\mathsf{x}}(v)$ and using the change of variable $v=u2^{\frac{1}{2\mathfrak{q}+1}}$, the limit distribution equals

align*[align* omitted — 517 chars of source]

As a consequence,

equation*[equation* omitted — 290 chars of source]

contradicting (ref) because $2^{\frac{\mathfrak{q}}{2\mathfrak{q}+1}} \neq \sqrt{2}$.

In other words, the bootstrap estimator $\widehat{\theta}_n^*(\mathsf{x})$ fails to approximate the limit distribution. $\qedsymbol$

Remarks on verifying conditions in applications

Below we verify the hypothesis of Theorem (ref) for various examples. For this purpose, it suffices to verify Assumptions (ref), (ref)-(ref), and (ref) since Assumption (ref) implies (ref)-(ref) by Lemma (ref).

When $\gamma_0$ is known, it is natural to take $\widehat{\Gamma}_n=\check{\Gamma}_n=\bar{\Gamma}_n$, in which case (ref) reduces to the requirement that, for some $\rho_{\gamma}\in (0,2)$,

equation[equation omitted — 380 chars of source]

An identical remark applies to $\phi_0$ and (ref).

Proof of Corollary (ref)

Assumption (ref) and (ref)-(ref) follow from the hypothesis. \newline (ref): In this example, $\gamma_0(x;\mathbf{Z})=\mathbbm{1}(X\leq x)$ is known, so it suffices to verify (ref). The uniform covering number of $\{\mathbbm{1}(\cdot\leq x) :x\in\mathbb{R}\}$ grows linearly, and an envelope function can be taken to be $1$. For an envelope function of $\{\mathbbm{1}(\cdot\leq x)-\mathbbm{1}(\cdot\leq \mathsf{x}) : |x-\mathsf{x}|\leq \eta \}$, we can take $\mathbbm{1}(-\eta +\mathsf{x}\leq \cdot\leq \mathsf{x}+ \eta )$ and the moment bound is satisfied as $\mathbbm{E}[\mathbbm{1}(-\eta +\mathsf{x}\leq X\leq \mathsf{x}+ \eta )]\leq C \eta$. \newline (ref) trivially holds as $\widehat{\Phi}_n(x)=\widehat{\Phi}_n^*(x)=x$. \newline (ref): Here $\psi_{\mathsf{x}}(v;\mathbf{Z})=\mathbbm{1}(X\leq \mathsf{x}+v)-\mathbbm{1}(X\leq\mathsf{x}) -\Phi_0(\mathsf{x})v$. Then,

equation*[equation* omitted — 240 chars of source]

Also, $\psi_{\mathsf{x}_n}(s\eta_n;\mathbf{Z})=\mathbbm{1}(\mathsf{x}_n \land (\mathsf{x}_n+s\eta_n) < X\leq \mathsf{x}_n \lor (\mathsf{x}_n+s\eta_n) ) - f_0(\mathsf{x}) s\eta_n$ and

align*[align* omitted — 625 chars of source]

Then, for any $s,t\in\mathbb{R}$ and $\mathsf{x}_n\to\mathsf{x}$, using continuity of $f_0$ at $\mathsf{x}$,

align*[align* omitted — 553 chars of source]

(ref) holds since $\widehat{u}_n=\widehat{u}_n^*$ converges in probability to $u_0$, the supremum of the support of $X$ by i.i.d.\ assumption. \newline (ref) and (ref) hold trivially since $\widehat{\Phi}_n$ and $\widehat{\Phi}_n^*$ are the identity map. \newline Assumption (ref) follows from (ref) and empirical process theory arguments.

Proof of Corollary (ref)

Assumption (ref) and (ref)-(ref) follow from the hypothesis. \newline (ref): We have $\widehat{\Gamma}_n=1-\widehat{S}_n$ with $\widehat{S}_n$ the Kaplan-Meier estimator. By Theorem 1 of Lo-Singh_1986_PTRF,

equation*[equation* omitted — 178 chars of source]

Since $\sqrt{n a_n} \leq n^{2/3}$ for $\mathfrak{q}\geq 1$, $\sup_{x \in I}|\widehat{\Gamma}_n(x)-\Gamma_0(x)|=o_{\mathbbm{P}}(1)$ and $\sqrt{na}\sup_{|v|\leq \delta}|\widehat{\Gamma}_n(\mathsf{x}+v)-\widehat{\Gamma}_n(\mathsf{x}) - \bar{\Gamma}_n(\mathsf{x}+v)+\bar{\Gamma}_n(\mathsf{x})|=o_{\mathbbm{P}}(1)$ hold.

We have

align*[align* omitted — 826 chars of source]

By $S_0(u_0)G_0(u_0)>0$, we have $\sqrt{n}\sup_{x\in I}|\widehat{S}_n(x)-S_0(x)|=O_{\mathbbm{P}}(1)$, $\sqrt{n}\sup_{x\in I}|\widehat{G}_n(x)-G_0(x)|=O_{\mathbbm{P}}(1)$, and $\sqrt{n}\sup_{x\in I}|\widehat{\Lambda}_n(x)-\Lambda_0(x)|=O_{\mathbbm{P}}(1)$, which in turn implies

equation*[equation* omitted — 305 chars of source]

Let $\delta = \min\{\mathsf{x}, (u_0-\mathsf{x})\}/4$, $R_{1n}(v)=| \widehat{F}_n(\mathsf{x}+v)-\widehat{F}_n(\mathsf{x}) - F_0(\mathsf{x}+v)+F_0(\mathsf{x})|$, $R_{2n}=| \widehat{F}_n(\mathsf{x})-F_0(\mathsf{x})|$, $R_{3n}= \sup_{x\in I}| [\widehat{S}_n(x)\widehat{G}_n(x)]^{-1} - [S_0(x)G_0(x)]^{-1} |$, and $R_{4n}(x_1,x_2) = |\int_{x_1}^{x_2} \frac{\widehat{\Lambda}_n(du)}{\widehat{S}_n(u)\widehat{G}_n(u)} -\int_{x_1}^{x_2} \frac{\Lambda_0(du)}{S_0(u)G_0(u)}| $. For $|v|\leq \delta$,

align*[align* omitted — 840 chars of source]

As noted above, $\sup_{|v|\leq \delta}R_{1n}(v)=o_{\mathbbm{P}}( (na_n)^{-1/2})$, $R_{2n}=O_{\mathbbm{P}}(n^{-1/2})$, and $R_{3n}=O_{\mathbbm{P}}(n^{-1/2})$ by $S_0(u_0)G_0(u_0)>0$. Also, uniformly over $V\in (0, 2\delta]$, $\sup_{| v|\leq V}| \frac{1}{n}\sum_{i=1}^n (\mathbbm{1}(X_i\leq \mathsf{x} +v) - \mathbbm{1}(X_i\leq \mathsf{x}) )| \leq C V + O_{\mathbbm{P}}(n^{-1/2})$. If

equation[equation omitted — 196 chars of source]

and

equation[equation omitted — 252 chars of source]

uniformly over $|v|\leq 2\delta$, then there exist random variables $A_n=o_{\mathbbm{P}}(1)$ and $B_n=O_{\mathbbm{P}}(\sqrt{a_n})$ independent of $v$ such that for $V\in (0, 2\delta]$,

equation*[equation* omitted — 197 chars of source]

that is,\ $\beta_{\gamma}=1$ in the notation of (ref). To show (ref) and (ref), for $x_1,x_2\in I$,

equation*[equation* omitted — 216 chars of source]

Let $J_0(u)= 1/[S_0(u)G_0(u)]$ and integration by parts implies

align*[align* omitted — 592 chars of source]

The first term after the second equality is bounded by $O_{\mathbbm{P}}(n^{-1/2}) \vert x_2-x_1\vert$. For the second term, Theorem 1 of Burketal implies that on a suitable probability space there exists a sequence of standard Brownian motion $W_n$ such that $\sqrt{a_n}\sup_{\vert x_1-x_2\vert\leq v} \vert \sqrt{n}[\widehat{\Lambda}_n(x_1)-\Lambda_0(x_2) -\widehat{\Lambda}_n(x_1)+\Lambda_0(x_2)] - W_n(d(x_1))+W_n(d(x_2))\vert = o_{\mathbbm{P}}(1)$, where $d(x)=\int_0^x \frac{F_0(du)}{S_0(u)^2G_0(u)}$. By Theorem 3.2 of Pollard_1989_SS, there is some fixed constant $C>0$ such that $\mathbbm{E}[\sup_{\vert x_1-x_2\vert\leq v}\vert W_n(d(x_1))-W_n(d(x_2))\vert]\leq C v$. Finally, $\vert \int_{x_1}^{x_2} [\widehat{\Lambda}_n(u)-\Lambda_0(u)]J_0(du)\vert\leq \sup_{x'\in [x_1,x_2]}\vert\widehat{\Lambda}_n(x')-\Lambda_0(x')\vert [J_0(x_2)-J_0(x_1)]\leq O_{\mathbbm{P}}(n^{-1/2}) \vert x_2-x_1\vert$. Thus, (ref) and (ref) hold.

For the function class $\mathfrak{F}_{\gamma}$, we can take $\bar{F}_{\gamma}(\mathbf{Z}) = 1 + [S_0(u_0)G_0(u_0)]^{-1}[1 + \Lambda_0(u_0)]$ as a constant envelope. For the function class $\{S_0(x):x\in I\}$, given $m\in\mathbbm{N}$, there exists $\{x_1,\dots, x_{m+1}\}\subset I$ such that $\sup_{x\in I}\min_{l=1,\dots,m+1}|S_0(x_l)-S_0(x)|\leq 1/m$, which implies the uniform covering number is bounded by a linear function. The covering numbers of $\{\mathbbm{1}(\cdot\leq s):s\in I\}$ and $\{\int_0^{\cdot\land s}[S_0(u)G_0(u)]^{-1}\Lambda_0(du):s\in I\}$ are also bounded by a linear function. By Lemma 5.1 of vanderVaart-vanderLaan_2006_IJB, there exists $\rho\in (0,2)$ such that $\limsup_{\eta\downarrow 0}\log N_U(\eta,\mathfrak{F}_{\gamma})\eta^{\rho} < \infty$ holds.

Now consider the uniform covering number of $\hat{\mathfrak{F}}_{\gamma}$. Given a realization of $(\widehat{S}_n,\widehat{G}_n)$, the mapping $x\mapsto\int_0^{x\land s}[\widehat{S}_n(u)\widehat{G}_n(u)]^{-1}\widehat{\Lambda}_n(du)$ is a composition of $x\mapsto x\land s$ and $x\mapsto \int_0^x [\widehat{S}_n(u)\widehat{G}_n(u)]^{-1}\widehat{\Lambda}_n(du)$. The latter mapping is monotone, and the first mapping is a VC-subgraph class, and Lemma 2.6.18 of vanderVaart-Wellner_1996_Book implies $\{\int_0^{\cdot\land s}[\widehat{S}_n(u)\widehat{G}_n(u)]^{-1}\widehat{\Lambda}_n(du):s\in I\}$ is a VC-subgraph class. Note that since $S_0,G_0$ are bounded away from zero, $\widehat{S}_n,\widehat{G}_n$ are bounded away from zero with probability approaching one. Thus, for some $\rho\in (0,2)$, $\limsup_{\eta\downarrow 0}\log N_U(\eta,\hat{\mathfrak{F}}_{\gamma})\eta^{\rho} = O_{\mathbbm{P}}(1)$ holds.

For $s\leq t\in I$,

equation*[equation* omitted — 239 chars of source]

and we can take $D_{\gamma}^{\eta}(\mathbf{Z})$ to be a constant multiple of $\sup_{|s|\leq \eta}|F_0(\mathsf{x} +s)-F_0(\mathsf{x})| + \Delta \mathbbm{1}(|\check{X}-\mathsf{x}|\leq \eta) + \int_{ x-\eta}^{x+\eta}\Lambda_0(du)/S_0(u)G_0(u)$. For $\eta >0$ small enough, there is some fixed $C>0$ with

equation*[equation* omitted — 133 chars of source]

(ref) trivially holds as $\widehat{\Phi}_n(x)=\widehat{\Phi}_n^*(x)=x$. \newline (ref): We have

equation*[equation* omitted — 203 chars of source]

where $O(|v|)$ is uniformly over small enough $|v|$. Since

equation*[equation* omitted — 206 chars of source]

the first display in (ref) is satisfied. For the covariance kernel,

align*[align* omitted — 670 chars of source]

where the last equality uses continuity of $(S_0,G_0,f_0)$ at $\mathsf{x}$ i.e.,\ $\int_{\mathsf{x}_n}^{\mathsf{x}_n +\eta_n} [\frac{f_0(u)}{S_0(u)^2G_0(u)}-\frac{f_0(\mathsf{x})}{S_0(\mathsf{x})^2G_0(\mathsf{x})}]du =o(1)\eta_n$.

(ref), (ref), and (ref) hold since in this example, $\widehat{u}_n=\widehat{u}_n^*=u_0$ and $\widehat{\Phi}_n,\widehat{\Phi}_n^*$ are the identity map.

\paragraph*{Assumption (ref)} As noted when verifying (ref), $\widehat{G}_n(1;\eta_n)=o_{\mathbbm{P}}(1)$ for any $\eta_n=o(1)$ with $a_n^{-1}\eta_n^{-1}=O(1)$. $\widehat{\Phi}_n=\Phi_0$ is the identity map and the desired result holds.

Proof of Corollary (ref)

Assumption (ref) and (ref)-(ref) follow from the hypothesis. \newline (ref): In this example, $\gamma_0(x;\mathbf{Z})=Y\mathbbm{1}(X\leq x)$ is known, so it suffices to verify (ref). The uniform covering number bound is straightforward as $\{\mathbbm{1}(\cdot\leq x):x\in\mathbb{R}\}$ is a VC-subgraph class. An envelope function is $|Y|$, whose second moment is finite. For $x\in I_{\mathsf{x}}^{\eta}$, $|\gamma_0(x;\mathbf{Z})-\gamma_0(\mathsf{x};\mathbf{Z})|\leq |Y| \mathbbm{1}(\mathsf{x}-\eta \leq X\leq \mathsf{x} +\eta)$, which we can take as $\bar{D}_{\gamma}^{\eta}(\mathbf{Z})$. Then, for $j=2,4$,

equation*[equation* omitted — 201 chars of source]

and the desired bound holds. \newline (ref): $\phi_0(x;\mathbf{Z})=\mathbbm{1}(X\leq x)$ is known, so it suffices to verify the analogue of (ref). The argument is the same as for checking (ref) in monotone density estimation with no censoring.

(ref): We have

equation*[equation* omitted — 228 chars of source]

Then,

equation*[equation* omitted — 224 chars of source]

and the first display holds. For the covariance kernel, note $|(\mu_0(X)-\mu_0(\mathsf{x}_n))(\mathbbm{1}(X\leq \mathsf{x}_n+v)-\mathbbm{1}(X\leq \mathsf{x}_n)|\leq |v| \sup_{|x-\mathsf{x}|\leq 2\eta} |\partial\mu_0(x)|$ for $|x_n-\mathsf{x}|\lor |v|\leq \eta$ for $\eta>0$ small enough. Then,

align*[align* omitted — 537 chars of source]

and

align*[align* omitted — 369 chars of source]

as desired.

(ref) trivially holds since $\widehat{u}_n=\widehat{u}_n^*=1$ in this example.

(ref): $\widehat{\Phi}_n,\widehat{\Phi}_n^*$ are (empirical) cdfs, so they are non-negative, non-decreasing, and right-continuous. $\{0,\widehat{u}_n\}\subset \widehat{\Phi}_n(I)$ and $\{0,\widehat{u}_n^*\}\subset \widehat{\Phi}_n^*(I)$ hold as $\widehat{u}_n=\widehat{u}_n^*=1$, $\widehat{\Phi}_n(\min_iX_i-)=0=\widehat{\Phi}_n^*(\min_iX_i-)$, and $\widehat{\Phi}_n(\max_iX_i)=1=\widehat{\Phi}_n^*(\max_iX_i)$. The sets $\widehat{\Phi}_n(I),\widehat{\Phi}_n^*(I)$ are finite and thus closed.

(ref): With probability one, all $X_i$'s are distinct. If $x$ is one of $X_i$'s, $\widehat{\Phi}_n^*( x ) - \widehat{\Phi}_n^*( x- ) = n^{-1}W_{j,n}$ for some $j$, and $\widehat{\Phi}_n^*( x ) - \widehat{\Phi}_n^*( x- ) =0$ otherwise.

Given $\mathbbm{E}|W_{1,n}|^{\mathfrak{r}}<\infty$, $n^{-1}\max_{1\leq i\leq n}|W_{i,n}|=o_{\mathbbm{P}}(n^{-5/6})$, which implies $$\sqrt{na_n}\sup_{x\in I}|\widehat{\Phi}_n^*( x ) - \widehat{\Phi}_n^*( x- ) |=o_{\mathbbm{P}}(1).$$ The argument for $\widehat{\Phi}_n$ is similar. \newline Assumption (ref) follows from (ref)-(ref) and empirical process theory arguments.

Proof of Corollary (ref)

Assumption (ref) and (ref)-(ref) follow from the hypothesis.

(ref): In this example, $\check{\Gamma}_n=\widehat{\Gamma}_n$.

align*[align* omitted — 453 chars of source]

The last sum is bounded by $\sup_{x\in I}|\frac{1}{n}\sum_{j=1}^n\mu_0(x,\mathbf{A}_j)-\theta_0(x)|$, and this object is $O_{\mathbbm{P}}(n^{-1/2})$: to see this claim, first note that Assumption MRC (iv) and Theorem 2.7.11 of vanderVaart-Wellner_1996_Book imply $\limsup_{\epsilon\downarrow0}\log N_U(\epsilon,\{\mu(x,\cdot):x\in I\}) \epsilon^V <\infty$ for some $V\in (0,2)$ and Theorem 4.2 of Pollard_1989_SS implies $\sup_{x\in I}|\frac{1}{n}\sum_{j=1}^n\mu_0(x,\mathbf{A}_j)-\theta_0(x)|=O_{\mathbbm{P}}(n^{-1/2})$. Together with Assumption MRC (iii), $a_n \frac{1}{n}\sum_{i=1}^n \sup_{x\in I}|\widehat{\gamma}_n(x;\mathbf{Z})-\gamma_0(x;\mathbf{Z})|^2=o_{\mathbbm{P}}(1)$ holds.

By Assumption MRC (iii), uniformly over $V\in (0,2\delta]$

equation*[equation* omitted — 243 chars of source]

and the desired inequality holds.

The uniform covering numbers of $\mathfrak{F}_{\gamma},\hat{\mathfrak{F}}_{\gamma,n}$ are the same order as for $\{\mathbbm{1}(\cdot\leq x):x\in I\}$. For $x \in I_{\mathsf{x}}^{\eta}$, $|\gamma_0(x;\mathbf{Z})-\gamma_0(\mathsf{x};\mathbf{Z})|\leq \mathbbm{1}(\mathsf{x}-\eta\leq X\leq \mathsf{x}+\eta)(|\varepsilon| c^{-1}+\theta_0(\mathsf{x}+\eta) )$. Then, $\limsup_{\eta\downarrow0}\mathbbm{E}[\bar{D}_{\gamma}^{\eta}(\mathbf{Z})^j]\eta^{-1} <\infty$ holds for $j=2,4$.

(ref): $\phi_0(x;\mathbf{Z})=\mathbbm{1}(X\leq x)$ is known and the same as in the classical case, so the same argument applies.

(ref): We have

equation*[equation* omitted — 204 chars of source]

Then, for $v,v'\in [-\eta,\eta]$ with sufficiently small $\eta>0$,

align*[align* omitted — 264 chars of source]

and $\sup_{v\neq v'\in [-\eta_n,\eta_n]}\mathbbm{E}[|\psi_{\mathsf{x}}(v;\mathbf{Z}) - \psi_{\mathsf{x}}(v';\mathbf{Z})|]/|v-v'|=O(1)$ holds.

For $s\eta_n$ small enough, $\psi_{\mathsf{x}}(s\eta_n;\mathbf{Z})= (\mathbbm{1}(X\leq \mathsf{x}+ s\eta_n)-\mathbbm{1}(X\leq \mathsf{x}))\varepsilon g_0(X,\mathbf{A})^{-1} + O(\eta_n)$ and

align*[align* omitted — 347 chars of source]

and

align*[align* omitted — 863 chars of source]

Since $\frac{f_{X|\mathbf{A}}(x|\mathbf{A})}{g_0(x,\mathbf{A})^2} = \frac{f_0(x)}{g_0(x,\mathbf{A})}$, we have

equation*[equation* omitted — 305 chars of source]

as desired.

(ref) (ref) (ref): Verifying these conditions is the same as in the classical monotone regression case.

Primitive sufficient conditions for Assumption MRC (iii)

Here we provide primitive sufficient conditions for Assumption MRC (iii) by focusing on specific estimators $\widehat{\mu}_n$ and $\widehat{g}_n$. As discussed by Westling-Gilbert-Carone_2020_JRRSB, cross-fitting avoids restrictions on uniform entropy, allowing for a large class of flexible preliminary estimators. Here we use sample splitting to simplify exposition, but the proposed procedure can be straightforwardly modified for cross-fitting.

Suppose there is a separate random sample $\tilde{\mathbf{Z}}_1,\dots,\tilde{\mathbf{Z}}_n$ drawn from the distribution of $\mathbf{Z}$, which is independent of $\mathbf{Z}_1,\dots,\mathbf{Z}_n$. Preliminary estimators $\widehat{\mu}_n$ and $\widehat{g}_n$ are constructed from $\tilde{\mathbf{Z}}_1,\dots,\tilde{\mathbf{Z}}_n$. For concreteness, we consider a partitioning-based least squares estimator $\widehat{\mu}_n$ Cattaneo-Farrell-Feng_2020_AoS and local polynomial kernel-based estimators $\widehat{f}_{X|\mathbf{A},n}(x|\mathbf{a})$ and $\widehat{f}_n(x)$ of $f_{X|\mathbf{A}}(x|\mathbf{a})$ and $f_0(x)$ \citep*{Cattaneo-Jansson-Ma_2020_JASA,Cattaneo-Chandak-Jansson-Ma_2024}, from which we construct $\widehat{g}_n(x,\mathbf{a})=\widehat{f}_{X|\mathbf{A},n}(x|\mathbf{a})/\widehat{f}_n(x)$.

Let $d=\mathrm{dim}(\mathbf{A})$. For simplicity, suppose the support of $(X,\mathbf{A}')'$ equals $[0,1]^{1+d}$. Let $\mathbf{p}(x,\mathbf{a})$ be a $k_n$-dimensional vector of bounded basis functions of order $m$ on $\mathcal{S}$ which are locally supported e.g.,\ splines Cattaneo-Farrell-Feng_2020_AoS. We consider the estimator

equation*[equation* omitted — 267 chars of source]

For the estimator of $f_{X|\mathbf{A}}(x\vert\mathbf{a})$, letting $\widehat{F}_{X|\mathbf{A},n}(\cdot\vert \mathbf{a})$ be an estimator of $\mathbbm{P}[X\leq \cdot \vert \mathbf{A}=\mathbf{a}]$ specified below, $\widehat{f}_{X|\mathbf{A},n}(x|\mathbf{a})$ is obtained by local polynomial regression:

equation*[equation* omitted — 282 chars of source]

where $\mathfrak{p}_1\geq 1$ is the order of the polynomial basis $\mathbf{q}_1(x)=(1,x/1!, x^2/2!,\dots,x^{\mathfrak{p}_1}/\mathfrak{p}_1!)'$, $\mathbf{e}_l$ is the conformable unit vector whose $l$th element is unity, and $K_h(x)=K(x/h)/h$ for some kernel function $K$ and some positive bandwidth $h$. The estimator $\widehat{F}_{X|\mathbf{A},n}(x|\mathbf{a})$ is constructed via local polynomial regression of order $\mathfrak{p}_2=\mathfrak{p}_1-1$:

equation*[equation* omitted — 298 chars of source]

where, using standard multi-index notation, $\mathbf{q}_2(\mathbf{a})$ denotes the $k_{\mathfrak{p}_2}$-dimensional vector collecting the polynomials $\mathbf{a}^{\mathbf{m}}/\mathbf{m}!$ for $0\leq \vert \mathbf{m}\vert\leq \mathfrak{p}_2$ with $\mathbf{a}^{\mathbf{m}} = a_1^{m_1}a_2^{m_2}\dots a_d^{m_d}$, $\vert\mathbf{m}\vert=\sum_{j=1}^dm_j$, and $k_{\mathfrak{p}_2}=\frac{(d+\mathfrak{p}_2)!}{d!\mathfrak{p}_2!}+1$, and $L_h(\mathbf{a})=L(\mathbf{a}/h)/h^d$ for $L(\mathbf{a})=\prod_{j=1}^dK(a_j)$ i.e.,\ product kernel. The estimator $\widehat{f}_n(x)$ is constructed in a similar manner. First, the empirical cdf $\widehat{F}_n$ of $\{\tilde{X}_i\}$ is constructed and then $\widehat{f}_n(x)$ is formed via local polynomial regression:

equation*[equation* omitted — 227 chars of source]

where $b>0$ is some bandwidth.

Now we state sufficient conditions for Assumption MRC (ii) based on the partitioning-based series estimator $\widehat{\mu}_n$ and the kernel-based estimator $\widehat{g}_n$.

description• • $\tilde{\mathbf{Z}}_1,\dots,\tilde{\mathbf{Z}}_n$ are independent of $\mathbf{Z}_1,\dots,\mathbf{Z}_n$. • The support of $(X,\mathbf{A}')'$ is $[0,1]^{1+d}$, and the distribution of $(X,\mathbf{A}')'$ is absolutely continuous. The Lebesgue density of $(X,\mathbf{A}')'$ and the conditional variance of $Y$ given $(X,\mathbf{A}')'$ are bounded away from zero and continuous on $[0,1]^{1+d}$. $\mu_0$ is $(m+1)$-times continuously differentiable on $[0,1]^{1+d}$. • The vector of basis functions $p$ satisfies Assumptions 2, 3, and 4 of Cattaneo-Farrell-Feng_2020_AoS. • $f_{X|\mathbf{A}}(x|\mathbf{a})$ and $f_0(x)$ are $\mathfrak{p}_1$-times continuously differentiable in $x$, and $f_{X|\mathbf{A}}(x|\mathbf{a})$ is $\mathfrak{p}_1$-times continuously differentiable in $\mathbf{a}$. • $K$ is a symmetric, Lipschitz continuous probability density function supported on $[-1,1]$.

As verified by Cattaneo-Farrell-Feng_2020_AoS, (iii) holds for widely used local basis functions such as splines and wavelets.

lemmaSuppose $\tilde{\mathbf{Z}}_1,\dots,\tilde{\mathbf{Z}}_n$ is a random sample drawn from the distribution of $\mathbf{Z}$ and Primitive Conditions MRC hold. In addition, with $\tau_n= n^{-1}\log n$, $k_n=O(\tau_n^{-\frac{d+1}{2m+d+1}})$, $h=O(\tau_n^{\frac{d+1}{2\mathfrak{p}_1+d+1}})$, $b=O(\tau_n^{\frac{1}{2\mathfrak{p}_1+1}})$, $\frac{m}{2m+d+1}+\frac{\mathfrak{p}_1}{2\mathfrak{p}_1+d+1} \geq \frac{1}{2}$, and $\min\{\frac{2m}{2m+d+1},\frac{2\mathfrak{p}_1}{2\mathfrak{p}_1+d+1} \} > \frac{1}{2\mathfrak{q}+1}$. Then, $\widehat{\mu}_n$ and $\widehat{g}_n$ described above and $\widehat{\Gamma}_n$ based on the $\widehat{\mu}_n,\widehat{g}_n$ satisfy Assumption MRC (iii) with $\delta = \min\{\mathsf{x},1-\mathsf{x}\}/4$. In particular, \begin{equation} \sqrt{n} \sup_{|v|\leq V}\big\vert \widehat{\Gamma}_n(\mathsf{x}+v) - \widehat{\Gamma}_n(\mathsf{x}) - \bar{\Gamma}_n(\mathsf{x}+v) +\bar{\Gamma}_n(\mathsf{x})\big\vert \leq V O_{\mathbbm{P}}(1) + o_{\mathbbm{P}}(a_n^{-1}) \end{equation} uniformly over $V\in (0, 2\delta]$.
proofBy Theorem 4.3 of Cattaneo-Farrell-Feng_2020_AoS, \begin{equation*} \sup_{(x,\mathbf{a})\in [0,1]^{1+d}} \vert \widehat{\mu}_n(x,\mathbf{a})-\mu_0(x,\mathbf{a})\vert = O_{\mathbbm{P}}\bigg( \sqrt{\frac{k_n\log n}{n}} + k_n^{-\frac{m}{d+1}}\bigg) \end{equation*} and by Theorem 1 of \citet*{Cattaneo-Chandak-Jansson-Ma_2024}, \begin{equation*} \sup_{(x,\mathbf{a})\in [0,1]^{1+d}}\vert\widehat{f}_{X|\mathbf{A},n}(x|\mathbf{a})- f_{X|\mathbf{A}}(x|\mathbf{a})\vert = O_{\mathbbm{P}}\bigg( \sqrt{\frac{\log n}{n h^{1+d}}} + h^{\mathfrak{p}_1}\bigg). \end{equation*} Also, one can show \begin{equation*} \sup_{x\in [0,1]}\vert\widehat{f}_n(x) -f_0(x)\vert = O_{\mathbbm{P}}\bigg( \sqrt{\frac{\log n}{n b}} + b^{\mathfrak{p}_1}\bigg). \end{equation*} Then, with the specified rate of $k_n,h,b$, \begin{equation*} \sup_{(x,\mathbf{a})\in [0,1]^{1+d}} \vert \widehat{\mu}_n(x,\mathbf{a})-\mu_0(x,\mathbf{a})\vert = O_{\mathbbm{P}}\Big( \tau_n^{\frac{m}{2m+d+1}} \Big), \end{equation*} \begin{equation*} \sup_{(x,\mathbf{a})\in [0,1]^{1+d}} \vert \widehat{g}_n(x,\mathbf{a})-g_0(x,\mathbf{a})\vert = O_{\mathbbm{P}}\Big( \tau_n^{\frac{\mathfrak{p}_1}{2\mathfrak{p}_1+d+1}} \Big). \end{equation*} Since $\min\{\frac{2m}{2m+d+1},\frac{2\mathfrak{p}_1}{2\mathfrak{p}_1+d+1} \} > \frac{1}{2\mathfrak{q}+1}$, it follows $a_n\frac{1}{n}\sum_{i=1}^n|\widehat{\mu}_n(X_i,\mathbf{A}_i)-\mu_0(X_i,\mathbf{A}_i)|^2=o_{\mathbbm{P}}(1)$, $a_n\frac{1}{n^2}\sum_{i=1}^n\sum_{j=1}^n|\widehat{\mu}_n(X_i,\mathbf{A}_j)-\mu_0(X_i,\mathbf{A}_j)|^2=o_{\mathbbm{P}}(1)$, and $a_n\frac{1}{n}\sum_{i=1}^n\varepsilon_i^2|\widehat{g}_n(X_i,\mathbf{A}_i)-g_0(X_i,\mathbf{A}_i)|^2=o_{\mathbbm{P}}(1)$. Also, by $\frac{m}{2m+d+1}+\frac{\mathfrak{p}_1}{2\mathfrak{p}_1+d+1} \geq \frac{1}{2}$, \begin{align} \sup_{(x,\mathbf{a})\in [0,1]^{1+d}} \vert \widehat{\mu}_n(x,\mathbf{a})-\mu_0(x,\mathbf{a})\vert \sup_{(x,\mathbf{a})\in [0,1]^{1+d}} \vert \widehat{g}_n(x,\mathbf{a})-g_0(x,\mathbf{a})\vert = O_{\mathbbm{P}}\Big( n^{-1/2}\Big). \end{align} Decompose $\widehat{\gamma}_n$ into $\widehat{\gamma}_{1,n}$ and $\widehat{\gamma}_{2,n}$ where \begin{equation*} \widehat{\gamma}_{1,n}(x;\mathbf{Z}) = \mathbbm{1}(X\leq x) \frac{Y-\widehat{\mu}_n(X,\mathbf{A})}{\widehat{g}_n(X,\mathbf{A})},\quad \widehat{\gamma}_{2,n}(x;\mathbf{Z}) = \mathbbm{1}(X\leq x) \frac{1}{n}\sum_{j=1}^n\widehat{\mu}_n(X,\mathbf{A}_j) \end{equation*} and let $\widehat{\Gamma}_{k,n}(x)=\frac{1}{n}\sum_{i=1}^n \widehat{\gamma}_{1,n}(x;\mathbf{Z}_i)$ for $k=1,2$. Define $\gamma_{k,0},\bar{\Gamma}_{k,n}$ $k=1,2$ in the same manner. Letting $\tilde{\mathfrak{Z}}_n$ be the $\sigma$-field generated by $\tilde{\mathbf{Z}}_1,\dots,\tilde{\mathbf{Z}}_n$, \begin{align*} &\mathbbm{V}\Big[ \widehat{\Gamma}_{1,n}(\mathsf{x}+v) - \widehat{\Gamma}_{1,n}(\mathsf{x}) - \bar{\Gamma}_{1,n}(\mathsf{x}+v) + \bar{\Gamma}_{1,n}(\mathsf{x}) \Big\vert \tilde{\mathfrak{Z}}_n\Big] \\ &\leq n^{-1} \mathbbm{E}\left[\left. \big(\mathbbm{1}(X\leq\mathsf{x}+v) -\mathbbm{1}(X\leq \mathsf{x}) \big)^2 \bigg(\frac{Y-\widehat{\mu}_n(X,\mathbf{A})}{\widehat{g}_n(X,\mathbf{A})} - \frac{Y-\mu_0(X,\mathbf{A})}{g_0(X,\mathbf{A})}\bigg)^2 \right\vert\tilde{\mathfrak{Z}}_n\right] \\ &\leq n^{-1} C \mathbbm{E}\big[\big\vert\mathbbm{1}(X\leq\mathsf{x}+v) -\mathbbm{1}(X\leq \mathsf{x}) \big\vert \varepsilon^2\big] \sup_{\vert x-\mathsf{x}\vert\leq \vert v\vert}\sup_{\mathbf{a}\in [0,1]^d} \vert \widehat{g}_n(x,\mathbf{a})-g_0(x,\mathbf{a})\vert^2 \\ &\quad + n^{-1} C \mathbbm{E}\big[\big\vert\mathbbm{1}(X\leq\mathsf{x}+v) -\mathbbm{1}(X\leq \mathsf{x}) \big\vert \big]\sup_{\vert x-\mathsf{x}\vert\leq \vert v\vert}\sup_{\mathbf{a}\in [0,1]^d} \vert \widehat{\mu}_n(x,\mathbf{a})-\mu_0(x,\mathbf{a})\vert^2\\ &= o_{\mathbbm{P}}\big( (n a_n)^{-1} \big) \end{align*} where the inequalities hold with probability one. Note $\widehat{\Gamma}_{2,n}(x) = \frac{1}{n^2}\sum_{1\leq i\neq j\leq n} \mathbbm{1}(X_i\leq x)\widehat{\mu}_n(X_i,\mathbf{A}_j) + O_{\mathbbm{P}}(n^{-1})$, and \begin{align*} &\mathbbm{V}\left[ \left.\frac{1}{n(n-1)}\sum_{1\leq i\neq j\leq n} \big(\mathbbm{1}(X_i\leq \mathsf{x}+v)-\mathbbm{1}(X_i\leq \mathsf{x})\big)\big(\widehat{\mu}_n(X_i,\mathbf{A}_j) -\mu_0(X_i,\mathbf{A}_j)\big) \right\vert \tilde{\mathfrak{Z}}_n\right] \\ &\leq n^{-1}C \vert v\vert \sup_{\vert x-\mathsf{x}\vert\leq \vert v\vert}\sup_{\mathbf{a}\in [0,1]^d} \vert \widehat{\mu}_n(x,\mathbf{a})-\mu_0(x,\mathbf{a})\vert^2 \\ &\quad + C n^{-2}\vert v\vert \sup_{\vert x-\mathsf{x}\vert\leq \vert v\vert}\sup_{\mathbf{a}\in [0,1]^d} \vert \widehat{\mu}_n(x,\mathbf{a})-\mu_0(x,\mathbf{a})\vert^2 = o_{\mathbbm{P}}\big((n a_n)^{-1}\big) \end{align*} where we use Hoeffding decomposition and the inequality holds with probability one. To complete the proof, it suffices to show that there is a sequence of random variables $A_n'=O_{\mathbbm{P}}(1)$ such that for $V \in (0,a_n\delta]$, \begin{equation*} \sqrt{n } \sup_{\vert v\vert \leq V}\left\vert \mathbbm{E}\big[\widehat{\Gamma}_{n}(\mathsf{x}+v) - \widehat{\Gamma}_{n}(\mathsf{x}) - \bar{\Gamma}_{n}(\mathsf{x}+v) + \bar{\Gamma}_{n}(\mathsf{x})\big\vert \tilde{\mathfrak{Z}}_n\big] \right\vert \leq A_n' V. \end{equation*} With $f_{\mathbf{A}}(\mathbf{a})$ denoting the Lebesgue density of $\mathbf{A}$, \begin{align*} &\mathbbm{E}\big[\widehat{\Gamma}_{n}(\mathsf{x}+v) - \widehat{\Gamma}_{n}(\mathsf{x}) \big\vert \tilde{\mathfrak{Z}}_n\big] \\ &= \int \int \big(\mathbbm{1}(u\leq \mathsf{x}+v)-\mathbbm{1}(u\leq \mathsf{x})\big) \frac{g_0(u,\mathbf{a})}{\widehat{g}_n(u,\mathbf{a})}\big[\mu_0(u,\mathbf{a})-\widehat{\mu}_n(u,\mathbf{a})\big]f_0(u)du f_{\mathbf{A}}(\mathbf{a})d\mathbf{a} \\ &\qquad+ \int \big(\mathbbm{1}(u\leq \mathsf{x}+v)-\mathbbm{1}(u\leq \mathsf{x})\big)\int\widehat{\mu}_n(u,\mathbf{a})f_{\mathbf{A}}(\mathbf{a})d\mathbf{a} du\\ & \mathbbm{E}\big[\bar{\Gamma}_{n}(\mathsf{x}+v) - \bar{\Gamma}_{n}(\mathsf{x})\big\vert \tilde{\mathfrak{Z}}_n\big]\\ &=\int\big(\mathbbm{1}(u\leq \mathsf{x}+v)-\mathbbm{1}(u\leq \mathsf{x})\big)\int \mu_0(u,\mathbf{a})f_{\mathbf{A}}(\mathbf{a})d\mathbf{a}f_0(u)du. \end{align*} Then, for $v\geq 0$, \begin{align*} &\mathbbm{E}\big[\widehat{\Gamma}_{n}(\mathsf{x}+v) - \widehat{\Gamma}_{n}(\mathsf{x}) - \bar{\Gamma}_{n}(\mathsf{x}+v) + \bar{\Gamma}_{n}(\mathsf{x})\big\vert \tilde{\mathfrak{Z}}_n\big]\\ &= \int_{\mathsf{x}\leq u\leq \mathsf{x} +v} \big[\widehat{\mu}_n(u,\mathbf{a})-\mu_0(u,\mathbf{a})\big] \Big[1-\frac{g_0(u,\mathbf{a})}{\widehat{g}_n(u,\mathbf{a})}\Big]f_0(u) f_{\mathbf{A}}(\mathbf{a})d(u,\mathbf{a}). \end{align*} A similar expression holds for $v<0$. Then, for some fixed $C>0$, \begin{align*} &\left\vert \mathbbm{E}\big[\widehat{\Gamma}_{n}(\mathsf{x}+v) - \widehat{\Gamma}_{n}(\mathsf{x}) - \bar{\Gamma}_{n}(\mathsf{x}+v) + \bar{\Gamma}_{n}(\mathsf{x})\big\vert \tilde{\mathfrak{Z}}_n\big]\right\vert \\ &\leq C \vert v\vert \sup_{\vert x-\mathsf{x}\vert\leq \vert v\vert}\sup_{\mathbf{a}\in [0,1]^d} \vert \widehat{\mu}_n(x,\mathbf{a})-\mu_0(x,\mathbf{a})\vert \sup_{\vert x-\mathsf{x}\vert\leq \vert v\vert}\sup_{\mathbf{a}\in [0,1]^d} \vert \widehat{g}_n(x,\mathbf{a})-g_0(x,\mathbf{a})\vert \end{align*} and the desired result follows from (ref).

\paragraph*{Remark} Using (ref), one can verify the first part of Assumption (ref) i.e.,\ $\widehat{G}_{\mathsf{x},n}(1;\eta_n)=O_{\mathbbm{P}}(1)$. The second part is easy to verify using standard empirical process theory arguments.

Additional example: monotone density function with conditionally independent right-censoring

We consider the problem of estimating the density of a non-negative, continuously distributed random variable with censoring. We use the same notation as in Section (ref) of the main paper. Relative to Section (ref), we consider the additional complication of censoring being informative about the “survival time” $X$ i.e.,\ $X\not\protect\mathpalette{\protect\independenT}{\perp} C$. With covariates $\mathbf{A}$, we consider the setting of censoring at random: $X \protect\mathpalette{\protect\independenT}{\perp} C|\mathbf{A}$. See vanderlaan-Robins_2003_Book,Zeng_2004_AoS and references therein for existing analysis of this problem. We have

equation*[equation* omitted — 289 chars of source]

where $F_0(x|A)=1-S_0(x|A)$, $S_0(x|\mathbf{A})=\mathbbm{P}[X>x|\mathbf{A}]$, $G_0(c|\mathbf{A})=\mathbbm{P}[C>c|\mathbf{A}]$, and $\Lambda_0(x|\mathbf{A}) = \int_0^{x} \frac{f_0(u|\mathbf{A})}{S_0(u|\mathbf{A})}du$ with $f_0$ being the Lebesgue density of $X$. Denote by $\widehat{S}_n(\cdot|\cdot)$, $\widehat{G}_n(\cdot|\cdot)$, $\widehat{\Lambda}_n(\cdot|\cdot)$ preliminary estimates of $S_0,G_0,\Lambda_0$, respectively.

\paragraph*{Assumption SA.\arabic{section}} Let $\mathfrak{S}_n$, $\mathfrak{G}_n,\mathfrak{L}_n$ be sequences of function classes that contain $S_0(\cdot|\cdot)$, $G_0(\cdot|\cdot)$, $\Lambda_0(\cdot|\cdot)$, respectively.

enumerate[label=\normalfont(\roman*),noitemsep,itemindent=*] • $\mathsf{x}$ is in the interior of $I=[0,u_0]$, $X\protect\mathpalette{\protect\independenT}{\perp} C |A$, and $\theta_0=f_0$ satisfies Assumption (ref). • There exist $c,c_1,c_2>0,\rho_{\gamma}\in (0,2)$such that for $n\geq 1$, for any $S\in\mathfrak{S}_n$, $G\in \mathfrak{G}_n$, and $\Lambda \in \mathfrak{L}_n$, the following hold: $\log N_U(\varepsilon,\{S(x|\cdot):x\in I\}) \leq c \varepsilon^{-\rho_{\gamma}}$ for $\varepsilon\in (0,1)$, where $N_U$ is as defined in Section 4.2 of the main paper, and $c_1\leq S(x|\mathbf{A})\leq c_2$, $c_1\leq G(x|\mathbf{A})\leq c_2$ for $x\in I$, and $\Lambda(u_0|\mathbf{A})\leq c_2$ with probability one. • There exist $\delta>0, \beta_{\gamma} \in [1/2,2)$ such that for $V \in (0,2\delta]$, $\sqrt{n a_n}\sup_{|v|\leq V}|\widehat{\Gamma}_n(\mathsf{x}+v)-\widehat{\Gamma}_n(\mathsf{x})-\bar{\Gamma}_n(\mathsf{x}+v)+\bar{\Gamma}_n(\mathsf{x})|\leq o_{\mathbbm{P}}(1) +V^{\beta_{\gamma}}o_{\mathbbm{P}}(a_n^{\beta_{\gamma}})$ where $o_{\mathbbm{P}}$ terms do not depend on $V$. • With probability approaching one, $\widehat{S}_n\in\mathfrak{S}_n$, $\widehat{G}_{n}\in\mathfrak{G}_n$, $\widehat{\Lambda}_n\in\mathfrak{L}_n$. For $(\widehat{h}_n,h_0) \in \{(\widehat{S}_n,S_0),(\widehat{G}_n,G_0),(\widehat{\Lambda}_n,\Lambda_0) \}$, \begin{equation*} a_n\frac{1}{n}\sum_{i=1}^n\sup_{x\in I}|\widehat{h}_n(x|\mathbf{A}_i)-h_0(x|\mathbf{A}_i)|^2=o_{\mathbbm{P}}(1). \end{equation*} • The conditional distribution of $X$ given $\mathbf{A}$ has bounded Lebesgue density $f_{X|\mathbf{A}}$, $\mathbbm{E}[\frac{f_{X|\mathbf{A}}(\mathsf{x}|\mathbf{A})}{G_0(\mathsf{x}|\mathbf{A})}]>0$, and there are real-valued functions $B,\omega$ such that $\mathbbm{E}[B(\mathbf{A})]<\infty$, $\lim_{\eta\downarrow0}\omega(\eta)=0$, and for $|x-\mathsf{x}|$ sufficiently small, $|\frac{f_{X|\mathbf{A}}(x|\mathbf{A})}{S_0(x|\mathbf{A})G_0(x|\mathbf{A})}-\frac{f_{X|\mathbf{A}}(\mathsf{x}|\mathbf{A})}{S_0(\mathsf{x}|\mathbf{A})G_0(\mathsf{x}|\mathbf{A})}|\leq \omega(|x-\mathsf{x}|)B(\mathbf{A})$.

The condition (ref) is high-level, and there are a few different approaches to verify them. See Westling-Carone_2020_AoS for details.

corollaryUnder Assumption (ref), Assumptions (ref) and (ref) hold with \begin{equation*} \widehat{\Gamma}_n(x) = \frac{1}{n}\sum_{i=1}^n \widehat{\gamma}_n(x;\mathbf{Z}_i),\quad \widehat{\Gamma}_n^{*}(x) = \frac{1}{n}\sum_{i=1}^nW_{i,n} \widehat{\gamma}_n(x;\mathbf{Z}_i) \end{equation*} \[ \widehat{\gamma}_n(x;\mathbf{Z}) = \widehat{F}_n(x|\mathbf{A}) + \widehat{S}_n(x|\mathbf{A}) \bigg[\frac{\Delta\mathbbm{1}(\check{X}\leq x)}{\widehat{S}_n(\check{X}|\mathbf{A})\widehat{G}_n(\check{X}|\mathbf{A})} - \int_0^{\check{X}\land x} \frac{\widehat{\Lambda}_n(du|\mathbf{A})}{\widehat{S}_n(u|\mathbf{A})\widehat{G}_n(u|\mathbf{A})}\bigg], \] where $\widehat{F}_n=1-\widehat{S}_n$, \begin{equation*} \widehat{\Phi}_n(x)=\widehat{\Phi}_n^*(x)=x,\quad \widehat{u}_n=\widehat{u}_n^*=u_0, \end{equation*} \begin{equation*} \mathcal{C}_{\mathsf{x}}(s,t) = \mathbbm{E}\Big[\frac{f_{X|\mathbf{A}}(\mathsf{x}|\mathbf{A})}{G_0(\mathsf{x}|\mathbf{A})}\Big] (|s| \land |t|) \mathbbm{1}(\operatorname*{sign}(s)=\operatorname*{sign}(t)), \quad \mathcal{D}_{\mathfrak{q}}(\mathsf{x}) = \frac{\partial^\mathfrak{q} f_0(\mathsf{x})}{(\mathfrak{q}+1)!}. \end{equation*}

Proof of Corollary (ref)

In this example, $\widehat{\Phi}_n(x)=\widehat{\Phi}_n^*(x)=x=\Phi_0(x)$. Assumptions (ref) and (ref)-(ref) follow from the hypothesis. \newline (ref): In this example, $\check{\Gamma}_n=\widehat{\Gamma}_n$. For $x\in I$,

align*[align* omitted — 952 chars of source]

and

align*[align* omitted — 940 chars of source]

using integration by parts, where $J_0(u|\mathbf{a}) = [S_0(u|\mathbf{a})G_0(u|\mathbf{a})]^{-1}$. Thus, there is a fixed $C>0$ such that

align*[align* omitted — 308 chars of source]

From the hypothesis,

equation*[equation* omitted — 298 chars of source]

follow.

For uniform covering numbers, it suffices to show that each of $\{S(x|\cdot):x\in I\}$, $\{\mathbbm{1}(\cdot\leq x):x\in I\}$, and $\{\int_0^{\cdot\land x} \frac{\Lambda(du|\cdot)}{S(u|\cdot)G(u|\cdot)} : x\in I\}$ has an appropriate bound on the uniform covering number by Lemma 5.1 of vanderVaart-vanderLaan_2006_IJB (see examples after the lemma). For $\{\int_0^{\cdot\land x} \frac{\Lambda(du|\cdot)}{S(u|\cdot)G(u|\cdot)} : x\in I\}$ with $(S,G,\Lambda)\in \mathfrak{S}_n\times \mathfrak{G}_n\times\mathfrak{L}_n$, the mapping $x\mapsto \int_0^x \frac{\Lambda(du|\cdot)}{S(u|\cdot)G(u|\cdot)}$ is monotone (by the non-decreasing property of $\Lambda$ and $S,G\geq c_1 >0 $) and Lemma 2.6.18 of vanderVaart-Wellner_1996_Book implies the desired result.

There is a fixed $C>0$ such that for $x\in I_{\mathsf{x}}^{\eta}$,

align*[align* omitted — 221 chars of source]

and using $1-S_0(x|\cdot) =\int_0^x f_{X|\mathbf{A}}(u|\cdot) du$ with $f_{X|\mathbf{A}}$ being bounded, we can take

equation*[equation* omitted — 141 chars of source]

which satisfies the desired bound condition. \newline (ref) trivially holds since $\widehat{\Phi}_n(x)=\widehat{\Phi}_n^*(x)=x$.

(ref): We have

equation*[equation* omitted — 236 chars of source]

and the first display follows as in the independent censoring case. For the covariance kernel,

align*[align* omitted — 546 chars of source]

and $\eta_n^{-1}\mathbbm{E}[\psi_{\mathsf{x}_n}(s\eta_n;\mathbf{Z})\psi_{\mathsf{x}_n}(t\eta_n;\mathbf{Z})]$ converges to $\mathbbm{E}[\frac{f_{X|\mathbf{A}}(\mathsf{x}|\mathbf{A})}{G_0(\mathsf{x}|\mathbf{A})}] (|s|\land|t|) \mathbbm{1}(\operatorname*{sign}(s)=\operatorname*{sign}(t))$.

(ref), (ref), and (ref) hold since in this example, $\widehat{u}_n=\widehat{u}_n^*=u_0$ and $\widehat{\Phi}_n,\widehat{\Phi}_n^*$ are the identity map.

Additional example: monotone hazard function

Let $X$ be a non-negative random variable, $f_0$ be its Lebesgue density, and $S_0(x)=\mathbbm{P}[X>x]$ be its survival function. We consider the parameter of estimating the hazard function of $X$, $\theta_0(\mathsf{x})=f_0(\mathsf{x})/S_0(\mathsf{x})$, with possible right-censoring as in the monotone density function example. Observations $\mathbf{Z}_1,\dots,\mathbf{Z}_n$ come from a random sample of $\mathbf{Z}=(\check{X},\Delta)'$ where $\check{X}=\min\{X,C\}$ and $\Delta=\mathbbm{1}(X\leq C)$, $C$ being a random censoring time. As pointed out by Westling-Carone_2020_AoS, with strictly increasing $\Phi_0$ with $\Phi_0(0)=0$, the function $\Gamma_0$ takes the form $\Gamma_0(x) = \int_0^x \frac{f_0(u)}{S_0(u)}\Phi_0(du)$, and by taking $\Phi_0(x)=\int_0^xS_0(u)du$, $\Gamma_0(x) =F_0(x)=\mathbbm{P}[X\leq x]$. Since $\Gamma_0$ is identical to the monotone density case with the choice $\Phi_0=\int_0^xS_0(u)du$, we can leverage the analysis for the monotone density. The interval $I$ equals $[0,u_0^{\mathtt{MD}}]$ where $u_0^{\mathtt{MD}}$ is $u_0$ in the monotone density example. The $u_0$ for the monotone hazard function estimation is $u_0=\Phi_0(u_0^{\mathtt{MD}})$.

Consider the case of completely random censoring i.e.,\ $X\protect\mathpalette{\protect\independenT}{\perp} C$. As in the setup for Corollary (ref), let $\widehat{S}_n(x)$ be the Kaplan-Meier estimator for $S_0(x)=1-F_0(x)=\mathbbm{P}[X> x]$, $\widehat{F}_n=1-\widehat{S}_n$, and $\widehat{G}_n$ be the Kaplan-Meier estimator for $G_0(x)=\mathbbm{P}[C>x]$. Also,

equation*[equation* omitted — 208 chars of source]

and $\phi_0(x;\mathbf{Z}) =x-\int_0^x \gamma_0(u;\mathbf{Z})du$.

corollarySuppose that the hypothesis of Corollary (ref) and Assumption BW hold. Then, Assumptions (ref) and (ref) hold with \begin{equation*} \widehat{\Gamma}_n(x)=1-\widehat{S}_n(x),\quad \widehat{\Gamma}_n^*(x)=\frac{1}{n}\sum_{i=1}^nW_{i,n}\widehat{\gamma}_n(x;\mathbf{Z}_i), \end{equation*} \begin{equation*} \widehat{\gamma}_n(x;\mathbf{Z}) = \widehat{F}_n(x) + \widehat{S}_n(x)\left[\frac{\Delta\mathbbm{1}(\check{X}\leq x)}{\widehat{S}_n(\check{X})\widehat{G}_n(\check{X})} - \int_0^{\check{X}\land x} \frac{\widehat{\Lambda}_n(du)}{\widehat{S}_n(u)\widehat{G}_n(u)}\right], \end{equation*} \begin{equation*} \widehat{\Phi}_n(x)= \int_0^x \widehat{F}_n(u) du,\quad \widehat{\Phi}_n^*(x)=\int_0^x [1-\widehat{\Gamma}_n^*(u)]du = \frac{1}{n}\sum_{i=1}^n W_{i,n} \widehat{\phi}_n(x;\mathbf{Z}_i), \end{equation*} \begin{equation*} \widehat{\phi}_n(x;\mathbf{Z}) = x-\int_0^x \widehat{\gamma}_n(u;\mathbf{Z})du,\quad \widehat{u}_n=\widehat{u}_n^*=\widehat{\Phi}_n(u_0^{\mathtt{MD}}), \end{equation*} \begin{equation*} \mathcal{C}_{\mathsf{x}}(s,t) = \frac{f_0(\mathsf{x})}{G_0(\mathsf{x})} (|s| \land |t|) \mathbbm{1}(\operatorname*{sign}(s)=\operatorname*{sign}(t)),\quad \mathcal{D}_{\mathfrak{q}}(\mathsf{x}) = \frac{S_0(\mathsf{x})\partial^\mathfrak{q} f_0(\mathsf{x})}{(\mathfrak{q}+1)!}. \end{equation*}

Proof of Corollary (ref)

We use the same $\widehat{\gamma}_n$ function and assumptions as in the monotone density setting. Also, the covariance kernels are the same as in the monotone density case. We focus on (ref)-(ref) and (ref)-(ref).

(ref): Since

equation*[equation* omitted — 145 chars of source]
equation*[equation* omitted — 278 chars of source]

follow from $a_n\frac{1}{n}\sum_{i=1}^n\sup_{x\in I}\vert\widehat{\gamma}_n(x;\mathbf{Z}_i) - \gamma_0(x;\mathbf{Z}_i)\vert^2$, which was verified in Section (ref). To check $\sup_{x\in I}|\widehat{\Phi}_n(x)-\Phi_0(x)|=o_{\mathbbm{P}}(1)$, $\sup_{x\in I}|\frac{1}{n}\sum_{i=1}^n\phi_0(x;\mathbf{Z}_i)-\Phi_0(x)|=o_{\mathbbm{P}}(1)$ follows from Glivenko-Cantelli, and

equation*[equation* omitted — 260 chars of source]

where the last equality follows from $\frac{1}{n}\sum_{i=1}^n\sup_{x\in I}|\widehat{\gamma}_n(x;\mathbf{Z}_i)-\gamma_0(x;\mathbf{Z}_i)|^2=o_{\mathbbm{P}}(1)$. Now $\sup_{x\in I}|\widehat{\Phi}_n(x)-\Phi_0(x)|=o_{\mathbbm{P}}(1)$ follows by the triangle inequality.

For $|v|\leq V$

align*[align* omitted — 544 chars of source]

Using the argument in Section (ref), we can bound $\sup_{\vert v\vert\leq \vert V\vert } \vert \widehat{\Gamma}_n(\mathsf{x}+v)-\widehat{\Gamma}_n(\mathsf{x}) - \bar{\Gamma}_n(\mathsf{x}+v) + \bar{\Gamma}_n(\mathsf{x})\vert$. Then, $\sqrt{n} \vert \widehat{\Gamma}_n(\mathsf{x})-\bar{\Gamma}_n(\mathsf{x})\vert=O_{\mathbbm{P}}(1)$ implies

equation*[equation* omitted — 252 chars of source]

uniformly over $V\in (0,2\delta]$. Theorem 1 of Lo-Singh_1986_PTRF implies $\sqrt{n a_n}\sup_{x\in I}\vert \check{\Phi}_n(x ) - \check{\Phi}_n(\mathsf{x} ) - \bar{\Phi}_n(x)+\bar{\Phi}_n(\mathsf{x} ) \vert =o_{\mathbbm{P}}(1)$.

The conditions on the uniform covering number hold because $\gamma_0$ and $\widehat{\gamma}_n$ are bounded (for $\widehat{\gamma}_n$, with probability approaching one) and thus $|\phi_0(x_1;\mathbf{Z})-\phi_0(x_2;\mathbf{Z})|\leq C |x_1-x_2|$ and $|\widehat{\phi}_n(x_1;\mathbf{Z})-\widehat{\phi}_n(x_2;\mathbf{Z})|\leq C |x_1-x_2|$ with probability approaching one. By this Lipschitz property, the condition on $\bar{D}_{\phi}^{\eta}(\mathbf{Z})$ also holds.

(ref): Let $\psi_{\mathsf{x}}^{\mathtt{MD}}(v;\mathbf{Z})=\gamma_0(\mathsf{x} +v;\mathbf{Z})-\gamma_0(\mathsf{x};\mathbf{Z}) - \theta_0(\mathsf{x})v$ be the $\psi_{\mathsf{x}}$ function for the monotone density. Then, for $x$ sufficiently close to $\mathsf{x}$ and $|v|$ small enough,

align*[align* omitted — 298 chars of source]

Then, the same argument as in the monotone density case implies the desired result. \newline (ref) follows from consistency of $\widehat{\Phi}_n$ and $\widehat{\Phi}_n^*$.

(ref): $\widehat{\Phi}_n(x)=\int_0^x \widehat{F}_n(u)du,\widehat{\Phi}_n^*(x)=\int_0^x 1-\widehat{\Gamma}_n^*(u)du$ are non-negative since $\widehat{F}_n\geq 0$ and $1-\widehat{\Gamma}_n^*\geq 0$ with probability approaching one. This property also implies the non-decreasing property as $\widehat{\Phi}_n,\widehat{\Phi}_n^*$ are integrals. The continuity property also follows from the integral representation. By definition, $\widehat{\Phi}_n(0)=0=\widehat{\Phi}_n^*(0)$ and $\widehat{\Phi}_n(u_0^{\mathtt{MD}})=\widehat{u}_n=\widehat{u}_n^*=\widehat{\Phi}_n^*(u_0^{\mathtt{MD}})$ with $I=[0,u_0^{\mathtt{MD}}]$. The closedness of the range follows from continuity and $I$ being a compact interval. \newline (ref) follows from continuity of $\widehat{\Phi}_n$ and $\widehat{\Phi}_n^*$.

Additional example: distribution function estimation with current status data

We consider the problem of estimating the cdf of $X$ at $\mathsf{x}$, $\theta_0(\mathsf{x})=F_0(\mathsf{x})$. Observations $\mathbf{Z}_1,\dots,\mathbf{Z}_n$ come from a random sample of $\mathbf{Z}=(\Delta,C,\mathbf{A}')'$ where $\Delta=\mathbbm{1}(X\leq C)$, $C$ is a random censoring time, and $\mathbf{A}$ is a vector of covariates. In this example, we do not observe $\check{X}=X\land C$. Instead, we observe the censoring time and whether the observation was censored. This setup is often referred to as current status data. Let $H_0(x)=\mathbbm{P}[C\leq x]$ be the cdf of $C$. We can use $\Gamma_0(x) = \int_0^x F_0(u) H_0(du)$ and $\Phi_0(x)=H_0(x)$. The interval $I$ is the support of $X$ and $u_0=1$. We also assume $H_0$ admits a Lebesgue density $h_0$. The structure of the estimation problem turns out to be identical to the one for the monotone regression example, and we can leverage the common structure.

Independent right-censoring

First we consider the case of completely at random censoring $X\protect\mathpalette{\protect\independenT}{\perp} C$. See Groeneboom-Wellner_1992_Book for existing analysis. In this exaple, we do not use covariates $\mathbf{A}$. We set $\gamma_0(x;\mathbf{Z})=\Delta \mathbbm{1}(C\leq x)$ and $\phi_0(x;\mathbf{Z})=\mathbbm{1}(C\leq x)$. Note that if the notation is mapped by $(\Delta,C)\leftrightarrow (Y,X)$, then these functions are identical to those of the classical monotone regression problem (Corollary (ref)). Thus, the following result is identical to Corollary (ref), up to notation and some changes due to boundedness of $\Delta$.

corollaryLet $\varepsilon=\Delta-\mathbbm{E}[\Delta|C]$ and $\mathsf{x}$ be an interior point of $I$. Suppose that Assumption BW holds, $\theta_0=F_0$ satisfies Assumption (ref), the cdf $\Phi_0=H_0$ satisfies Assumption (ref), and $\sigma_0^2(x) = \mathbbm{E}[\varepsilon^2|C=x]$ is continuous and positive at $\mathsf{x}$. Then Assumptions (ref) and (ref) hold with \begin{equation*} \widehat{\Gamma}_n(x) = \frac{1}{n}\sum_{i=1}^n \Delta \mathbbm{1}(C\leq x),\quad \widehat{\Gamma}_n^*(x) = \frac{1}{n}\sum_{i=1}^n W_{i,n}\Delta \mathbbm{1}(C\leq x), \end{equation*} \begin{equation*} \widehat{\Phi}_n(x) = \frac{1}{n}\sum_{i=1}^n \mathbbm{1}(C\leq x),\quad \widehat{\Phi}_n^*(x) = \frac{1}{n}\sum_{i=1}^n W_{i,n}\mathbbm{1}(C\leq x),\quad \widehat{u}_n=\widehat{u}_n^*=1, \end{equation*} \begin{equation*} \mathcal{C}_{\mathsf{x}}(s,t) = h_0(\mathsf{x}) \sigma_0^2(\mathsf{x}) (|s| \land |t|)\mathbbm{1}(\operatorname*{sign}(s)=\operatorname*{sign}(t)),\qquad \mathcal{D}_{\mathfrak{q}}(\mathsf{x}) = \frac{h_0(\mathsf{x})\partial^{\mathfrak{q}}F_0(\mathsf{x})}{(\mathfrak{q}+1)!}. \end{equation*}

Conditionally independent right-censoring

We consider the case where right-censoring is conditionally independent i.e.,\ $X\protect\mathpalette{\protect\independenT}{\perp} C|\mathbf{A}$. vanderVaart-vanderLaan_2006_IJB analyzed this example as well as settings with time-varying covariates. We are focusing on time-invariant covariates. Define $F_0(C,\mathbf{A})=\mathbbm{E}[\Delta|C,\mathbf{A}]$ and $g_0(C,\mathbf{A})=\frac{h_{C|\mathbf{A}}(C|\mathbf{A})}{h_0(C)}$ where $h_{C|\mathbf{A}}$ is a conditional Lebesgue density of $C$ given $\mathbf{A}$ and $h_0$ is a Lebesgue density of $C$. Let $\widehat{F}_n(c,\mathbf{a})$ and $\widehat{g}_n(c,\mathbf{a})$ be preliminary estimators for $F_0(c,\mathbf{a})$ and $g_0(c,\mathbf{a})$, respectively.

Identical to the censoring completely at random case, with appropriate changes in the notation (i.e.,\ $(\Delta, C)\leftrightarrow (Y,X)$), the setup is equivalent to that of the monotone regression with covariates.

\paragraph*{Assumption SA.\arabic{section}.\arabic{subsection}} Let $\varepsilon=\Delta-\mathbbm{E}[\Delta|C,\mathbf{A}]$, $\sigma_0^2(C,\mathbf{A})=\mathbbm{E}[\varepsilon^2|C,\mathbf{A}]$, and $\delta >0$ be some fixed number.

enumerate[label=\normalfont(\roman*),noitemsep,itemindent=*] • $\mathbbm{E}[\frac{\sigma_0^2(\mathsf{x},\mathbf{A})}{g_0(\mathsf{x},\mathbf{A})}]>0$. • $h_{C|A}$ is bounded and $g_0$ is bounded away from zero. • There exist random variables $A_n=o_{\mathbbm{P}}(1)$ and $B_n=O_{\mathbbm{P}}(a_n^{1/2})$ such that $$\sqrt{na_n}\sup_{|v|\leq V}|\widehat{\Gamma}_n(\mathsf{x}+v)-\widehat{\Gamma}_n(\mathsf{x})-\Gamma_0(\mathsf{x}+v)+\Gamma_0(\mathsf{x})|\leq A_n +B_nV,\qquad V\in (0,2\delta].$$ In addition, $$\frac{a_n}{n}\sum_{i=1}^n|\widehat{F}_n(C_i,\mathbf{A}_i)-F_0(C_i,\mathbf{A}_i)|^2=o_{\mathbbm{P}}(1),$$ $$\frac{a_n}{n^2}\sum_{i=1}^n\sum_{j=1}^n|\widehat{F}_n(C_i,\mathbf{A}_j)-F_0(C_i,\mathbf{A}_j)|^2=o_{\mathbbm{P}}(1),$$ and $$\frac{a_n}{n}\sum_{i=1}^n|\widehat{g}_n(C_i,\mathbf{A}_i)-g_0(C_i,\mathbf{A}_i)|^2=o_{\mathbbm{P}}(1).$$$\mathbbm{E}[\bar{F}(\mathbf{A})^2]<\infty$, where $$\bar{F}(\mathbf{A})=\sup_{|c-c'|\leq\delta}\frac{|F_0(c,\mathbf{A})-F_0(c',\mathbf{A})|}{|c-c'|}.$$$\mathbbm{E}[\bar{\sigma}(\mathbf{A})^2]<\infty$, where, for some function $\omega$ with $\lim_{\eta\downarrow0}\omega(\eta)=0$, $$\left\vert\frac{\sigma_0^2(x,\mathbf{A})h_{0}(x)}{g_0(x,\mathbf{A})} -\frac{\sigma_0^2(\mathsf{x},\mathbf{A})h_{0}(x)}{g_0(\mathsf{x},\mathbf{A})}\right\vert \leq\omega(|x-\mathsf{x}|) \bar{\sigma}(\mathbf{A}),\qquad x\in I_{\mathsf{x}}^{\delta}.$$

Note that $|\varepsilon|\leq 1$.

corollarySuppose that $\mathsf{x}$ is in the interior of $I$, $\theta_0$ satisfies (ref), $\Phi_0$ satisfies (ref), and Assumption BW holds. If Assumption Assumption SA.\arabic{section}.\arabic{subsection} \ holds, Assumptions (ref) and (ref) hold with \begin{equation*} \widehat{\Gamma}_n(x)= \frac{1}{n}\sum_{i=1}^n \widehat{\gamma}_n(x;\mathbf{Z}_i),\quad \widehat{\Gamma}_n^*(x)= \frac{1}{n}\sum_{i=1}^nW_{i,n} \widehat{\gamma}_n(x;\mathbf{Z}_i), \end{equation*} \begin{equation*} \widehat{\gamma}_n(x;\mathbf{Z})=\mathbbm{1}(C\leq x)\bigg[\frac{\Delta-\widehat{F}_n(C,\mathbf{A})}{\widehat{g}_n(C,\mathbf{A})} + \frac{1}{n}\sum_{j=1}^n\widehat{F}_n(C,\mathbf{A}_j)\bigg], \end{equation*} \begin{equation*} \widehat{\Phi}_n(x) = \frac{1}{n}\sum_{i=1}^n \mathbbm{1}(C_i\leq x),\quad \widehat{\Phi}_n^*(x) = \frac{1}{n}\sum_{i=1}^nW_{i,n} \mathbbm{1}(C_i\leq x),\quad \widehat{u}_n=\widehat{u}_n^*=1, \end{equation*} \begin{equation*} \mathcal{C}_{\mathsf{x}}(s,t) = h_0(\mathsf{x}) \mathbbm{E}\bigg[\frac{\sigma_0^2(\mathsf{x},\mathbf{A})}{g_0(\mathsf{x},\mathbf{A})}\bigg] (|s| \land |t|) \mathbbm{1}(\operatorname*{sign}(s)=\operatorname*{sign}(t)),\qquad \mathcal{D}_{\mathfrak{q}}(\mathsf{x}) = \frac{h_0(\mathsf{x})\partial^{\mathfrak{q}}F_0(\mathsf{x})}{(\mathfrak{q}+1)!}. \end{equation*}

Proof of Corollaries (ref) and (ref)

As noted above, by mapping the notation $(\Delta, C)\leftrightarrow (Y,X)$, the arguments in Sections (ref) and (ref) directly apply to the current status estimators.

Rule-of-thumb step size selection

Here we develop a rule-of-thumb procedure to choose a step size for the bias-reduced numerical derivative estimator in the context of isotonic regression without covariates. Specifically, we consider the numerical derivative estimator

equation*[equation* omitted — 245 chars of source]

with $\underline{\mathfrak{s}}=3$, $c_1=1,c_2=-1,c_3=2,c_4=-2$. Then,

equation*[equation* omitted — 161 chars of source]
equation*[equation* omitted — 161 chars of source]

We use the (asymptotic) MSE-optimal step size discussed in the main paper. See also (ref). Yet, with the choice of $c_k$'s, part of the bias constant $\sum_{k=1}^{\underline{\mathfrak{s}}+1}\lambda_j^{\mathtt{BR}}(k)c_k^{\underline{s}+2}$ equals zero, and we need to turn to the next leading term of the bias, which is

equation*[equation* omitted — 248 chars of source]

Then, letting $\mathsf{B}_j^{\mathtt{BR}}(\mathsf{x}) = \frac{\partial^{6}\Upsilon_0(\mathsf{x})}{6!} \sum_{k=1}^{4}\lambda_j^{\mathtt{BR}}(k)c_k^{6}$, the MSE-optimal step size is

equation*[equation* omitted — 182 chars of source]

The bias and variance constants depend on unknown features of the data generating process. Specifically, $\mathsf{B}_j^{\mathtt{BR}}(\mathsf{x})$ depends on the regression function $\theta_0$, the Lebesgue density of $X$, and their derivatives at $X=\mathsf{x}$ while $\mathsf{V}_j^{\mathtt{BR}}(\mathsf{x})$ is determined by the density of $X$ and the conditional variance of the regression error $\varepsilon=Y-\theta_0(X)$ at $X=\mathsf{x}$. To implement the construction of the step size, we posit a simple parametric model:

equation*[equation* omitted — 133 chars of source]

where $\{\gamma_0,\gamma_1,\gamma_2,\gamma_4,\gamma_4,\gamma_5,\mu,\sigma\}$ are parameters to be estimated. Once we estimate the parameters of this reference model, we can construct a rule-of-thumb step size $\epsilon_{j,n}^{\mathtt{ROT}}$ by replacing $\mathsf{B}_j^{\mathtt{BR}}(\mathsf{x})$ and $\mathsf{V}_j^{\mathtt{BR}}(\mathsf{x})$ with their estimates. Note that although the bias and variance constant estimators may not be consistent for the true $\mathsf{B}_j^{\mathtt{BR}}(\mathsf{x})$ and $\mathsf{V}_j^{\mathtt{BR}}(\mathsf{x})$, the rate of $\epsilon_{j,n}^{\mathtt{ROT}}$ is MSE-optimal, and the numerical derivative estimator converges to $\mathcal{D}_j(\mathsf{x})$ sufficiently fast to satisfy Equation (ref) in the main paper.

\makeatletter\@input{xx.tex}\makeatother