Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
175,670 characters · 28 sections · 47 citation commands
Supplement to “Bootstrap-Assisted Inference for Generalized Grenander-type Estimators”
Lemma 4.2 of vanderVaart-vanderLaan_2006_IJB implies that, on $(l,u)$,
Since $\Phi(\mathsf{x}) \in (l,u)$, we therefore have
so Lemma 4.1 of vanderVaart-vanderLaan_2006_IJB implies that
where
Suppose $y^{\star} \in \Phi(I)$. Then $y^{\star} = \Phi(x^{\star})$, where
In particular,
where, in fact,
because if
then, contradicting the definition of $y^{\star}$, we have
The proof can therefore be completed by showing that $y^{\star}\in \Phi(I)$.
If $\Phi(I) \supseteq [l,u]$, then there is nothing to show, so suppose $\Phi(I) \not\supseteq [l,u]$. If $y \in \Phi(I)^c \cap [l,u]$, then, since $\Phi(I) \cap [l,u]$ is closed and $l,u \in \Phi(I)$, we have $[y-\eta,y+\eta] \cap \Phi(I) = \emptyset$ for some $\eta > 0$ with $[y-\eta,y+\eta] \subset [l,u]$. Therefore, the function $\mathsf{LSC} (\Gamma\circ\Phi^{-})$ is constant on the interval $[y-\eta/2,y+\eta/2]$, implying in particular that
$\mathbbm{G}(v)$ is a Gaussian process with $\mathbbm{E}[\mathbbm{G}(v)]=\mu(v)$ for all $v\in \mathbb{R}$ and $\mathbbm{C}\mathrm{ov}[\mathbbm{G}(v),\mathbbm{G}(v')] = C \min\{|v|,|v'|\}\mathbbm{1}(\operatorname*{sign}(v)=\operatorname*{sign}(v')) =: \mathcal{K}(v,v')$ for some positive constant $C$ and for all $v,v' \in \mathbb{R}$. The form of the covariance kernel $\mathcal{K}$ implies that for every $\tau> 0$ and every $v,v' \in \mathbb{R}$,
and \[\mathcal{K}(v + v',v + v') - 2 \mathcal{K}(v + v',v') + \mathcal{K}(v',v') = \mathcal{K}(v,v).\] In addition, $\mathcal{K}(1,1) > 0$ and $\lim_{v \downarrow 0}\mathcal{K}(1,v) / \sqrt{v} = 0$.\footnote{As a side note, if the covariance kernel satisfies the two displays, then the Gaussian process is two-sided Brownian motion, up to mean shift and scaling by a positive constant.}
We begin by adapting the arguments of Kim-Pollard_1990_AoS to show that a maximizer of $\mathbbm{G}(v)$ over $v \in \mathbb{R}$ exists and is unique with probability one. Let $\mathbbm{G}^{\mu}(v) = \mathbbm{G}(v) - \mu(v)$ be the centered process and suppose that, for the same $c>1/2$ as in the hypothesis,
Then, with probability one, $\mathbbm{G}(v) \to -\infty$ as $|v| \to \infty$, implying in turn that a maximizer of $\mathbbm{G}(v)$ exists (because sample paths are continuous). Also, since
for $v \neq v'$, Lemma 2.6 of Kim-Pollard_1990_AoS implies that this maximizer is unique with probability one. In turn, (ref) follows from the Borel-Cantelli lemma because
where the equality uses the rescaling property $\mathcal{K}(v\tau,v'\tau) = \tau\mathcal{K}(v,v')$ and where the last inequality uses Jain-Marcus_1978_BookCh.
To show continuity of the function $x \mapsto \mathbbm{P}[\operatorname*{argmax}_{v \in \mathbb{R}}\{\mathbbm{G}(v)\} \leq x]$, it suffices to show that $\mathbbm{P}[\operatorname*{argmax}_{v \in \mathbb{R}}\{\mathbbm{G}(v)\} = x] = 0$ for every $x \in \mathbb{R}$. Fix $x \in \mathbb{R}$ and define $\widetilde{\mathbbm{G}}_x(0) = 0$ and
Then $\max_{v \in \mathbb{R}}\widetilde{\mathbbm{G}}_x(v) \geq 0$ and, for any set $\mathcal{V} \subset \mathbb{R}$,
In the sequel, we show that the majorant in (ref) can be made arbitrarily small by choice of $\mathcal{V}$. In particular, for $\varepsilon \in (0,1)$, we construct $v_{\varepsilon} > 0$ such that
and
For any $N \in \mathbb{N}$, defining $\mathcal{V}_{\varepsilon,N} = \{v_{\varepsilon},\dots,v_{\varepsilon}^N\}$, we therefore have
where the second inequality uses the fact that convergence of means and covariances of normal random vectors implies convergence in distribution. Since $N$ is arbitrary, the left-hand side in the preceding display is zero. Letting $\varepsilon \in (0,1)$ be given, the proof can therefore be completed by exhibiting $v_{\varepsilon}$ satisfying (ref) and (ref).
Because $\mathcal{K}(\tau,\tau) = \tau \mathcal{K}(1,1)$ and $\lim_{\tau \downarrow 0}[\mu(x + \tau) - \mu(x)]/\sqrt{\tau} = 0$, there exists $\bar{\tau}_\varepsilon^{'} \in (0,1)$ such that
for every $\tau \in (0,\bar{\tau}_\varepsilon^{'}]$. Also, because $\mathcal{K}(\tau_i,\tau_j) = \tau_i\mathcal{K}(1,\tau_j/\tau_i)$ and $\lim_{\tau \downarrow 0}\mathcal{K}(1,\tau )/\sqrt{\tau} = 0$, there exists $\bar{\tau}_\varepsilon^{''} \in (0,1)$ such that
for all $\tau_i,\tau_j > 0$ with $\tau_j/\tau_i \leq \bar{\tau}_\varepsilon^{''}$. If $v_{\varepsilon} \in (0, \bar{\tau}_\varepsilon^{'} \land \bar{\tau}_\varepsilon^{''})$, then (ref) and (ref) are satisfied.\qed
In preparation for the proof of Theorem (ref), this section presents six technical lemmas. The first lemma is a switching lemma, which will be used when characterizing the limiting distributions obtained in Theorem (ref).
The proof of Theorem (ref) furthermore employs various approximations to functionals of the form $\mathsf{LSC}_\Phi(f)$. The approximations in question are obtained using Lemmas (ref), (ref), (ref), and (ref). In all cases, the approximations are based on the representation
where $\mathcal{X}_\Phi^\epsilon(x) = \left(\Phi^-(\Phi(x) - \epsilon),\Phi^-(\Phi(x))\right] \cup \left(\Phi^-(\Phi(x)+),\Phi^-(\Phi(x) + \epsilon) \right)$.
The following lemma uses ((ref)) and the special structure of $\Gamma_0$ to obtain a simple “global” bound on the error of the approximation $\mathsf{LSC}_{\Phi}(\Gamma_0) \approx \Gamma_0$.
A simple “global” bound on the error of the approximation $\mathsf{LSC}_{\Phi}(f) \approx f$ is available also in the important special case where $f$ is proportional to $\Phi$.
Next, we give a “local” approximation to $\Phi^- \circ \Phi$. That approximation will later be used in combination with ((ref)) to obtain “local” approximations to $\mathsf{LSC}_\Phi(f)$, but the approximation is also useful in its own right and we therefore state it as a separate lemma.
Next, we obtain two “local” approximations to $\mathsf{LSC}_\Phi(f)$. The first of these is a generic approximation obtained by simply combining ((ref)) and Lemma (ref), but for later reference we state the result as a separate lemma.
The final lemma is concerned with the special case where $f$ is proportional to $\Phi$. In that case, the following “local” analog of Lemma (ref) shows that the bound(s) obtained in Lemma (ref) can be improved under mild conditions on $\Phi$.
\paragraph*{Proof of (ref)} Let $t \in \mathbb{R}$ be given. By Lemma (ref) and change of variables,
where
and
with
By (ref) and Lemma (ref), $\widehat{Z}_{\mathsf{x},n}^{\mathfrak{q}} = o_{\mathbbm{P}}(1)$. Suppose also that
where $\mathcal{H}_{\mathsf{x}}^{\mathfrak{q}}(v;t) = \mathcal{G}_{\mathsf{x}}(v) + \mathcal{M}_{\mathsf{x}}^{\mathfrak{q}}(v) - t \partial \Phi_0(\mathsf{x}) v$. Then
where the second line uses Lemma (ref) and where the last equality uses Lemma (ref). The proof of (ref) can therefore be completed by showing (ref).
We shall do so by means of the argmax continuous mapping theorem of Cox_2022. To be specific, using that theorem it can be shown that (ref) holds if
and if
We begin by showing (ref). First, by (ref), $\widehat{G}_{\mathsf{x},n}^{\mathfrak{q}} \leadsto \mathcal{G}_{\mathsf{x}}$. Also, by (ref) and (ref), as $u \to 0$,
where the first equality uses L'H\^{o}pital's rule and
As a consequence, $M_{\mathsf{x},n}^{\mathfrak{q}} \leadsto \mathcal{M}_{\mathsf{x}}^{\mathfrak{q}}$. Moreover, $\widehat{L}_{\mathsf{x},n}^{\mathfrak{q}} - L_{\mathsf{x},n}^{\mathfrak{q}} \leadsto 0$ by (ref) and $L_{\mathsf{x},n}^{\mathfrak{q}} \leadsto \mathcal{L}_{\mathsf{x}}$ by (ref), where $L_{\mathsf{x},n}^{\mathfrak{q}}(v) = a_n \left[\Phi_0(\mathsf{x} + v a_n^{-1}) - \Phi_0(\mathsf{x})\right]$. In particular, $\widehat{L}_{\mathsf{x},n}^{\mathfrak{q}} \leadsto \mathcal{L}_{\mathsf{x}}$ and therefore
Because $\widehat{G}_{\mathsf{x},n}^{\mathfrak{q}}+M_{\mathsf{x},n}^{\mathfrak{q}}$ is asymptotically equicontinuous,
by (ref) and Lemma (ref). Also, by (ref) and Lemma (ref),
The result (ref) follows from the three preceding displays and the fact that, on $\widehat{V}_{\mathsf{x},n}^{\mathfrak{q}}$,
Next, to show (ref), we first define $\theta_n(\mathsf{x};t) = \theta_0(\mathsf{x}) + t r_n^{-1}$ and note that
Now, if $|\widehat{v}_n(t)| > a_n \delta > 0,$ then
where $|\theta_n(\mathsf{x};t) - \theta_0(\mathsf{x})| = O(r_n^{-1}) = o(1)$, and, by (ref),
Also, using (ref), (ref), and Lemma (ref),
As a consequence, $\widehat{v}_n(t) = o_{\mathbbm{P}}(a_n)$: For any $\delta > 0$,
where the equality uses the fact, noted by Westling-Carone_2020_AoS, that the function $v \mapsto \theta_0(\mathsf{x}) \Phi_0(v) - \Gamma_0(v)$ is unimodal and maximized at $v = \mathsf{x}$.
Next, defining $\widehat{V}_{\mathsf{x},n}^{\mathfrak{q}}(j) = \{v\in \widehat{V}_{\mathsf{x},n}^{\mathfrak{q}} : 2^j < |v| \leq 2^{j+1}\}$ and using $\widehat{v}_n(t) = o_{\mathbbm{P}}(a_n)$, we have, for any $K$, any positive $\delta'$, and any sequence of events $\{\mathcal{A}_n'\}$ with $\lim_{n \to \infty}\mathbbm{P}[\mathcal{A}_n'] = 1$,
The proof of (ref) can therefore be completed by showing that the majorant side in the display can be made arbitrarily small by choice of $K$, $\delta'$, and $\{\mathcal{A}_n'\}$.
To do so, we begin by analyzing each term in the basic bound
Because $\widehat{H}_{\mathsf{x},n}^{\mathfrak{q}}(0;t) \leadsto \mathcal{H}_{\mathsf{x}}^{\mathfrak{q}}(0;t) = 0$ and because, by (ref) and Lemma (ref), there is a positive $\delta'$ such that
we may assume that, on $\{\mathcal{A}_n'\}$ and for some $C_0$,
Also, because, by (ref) and (ref), there is a positive $\delta'$ such that
we may assume that, on $\{\mathcal{A}_n'\}$ and for some $C_L$,
Next, by (ref) and Lemma (ref), with probability approaching one,
while, by (ref) and (ref), there is a positive $\delta'$ such that
We may therefore assume that, on $\{\mathcal{A}_n'\}$ and for some positive $C_M$,
Finally, by (ref) and Lemma (ref), with probability approaching one,
and we may therefore assume that, on $\{\mathcal{A}_n'\}$,
where $V_\eta(j) = \{v \in \mathbb{R}: \eta^{-1}2^j \leq |v| \leq \eta 2^{j+1}\}$.
As a consequence, by the Markov inequality,
where, by (ref), we may assume that, for some $C_G$,
and where, for all sufficiently large $j$,
In other words, for large $K$,
which can be made arbitrarily small by choice of $K$.
\paragraph*{Proof of (ref)} We proceed as in the proof of (ref). Let $t \in \mathbb{R}$ be given. By Lemma (ref) and change of variables,
where
and
with
By (ref) and Lemma (ref), $\widehat{Z}_{\mathsf{x},n}^{\mathfrak{q},*} = o_{\mathbbm{P}}(1)$. Suppose also that
Then, as in the proof of (ref),
The proof of (ref) can therefore be completed by showing (ref).
We shall do so by showing that
and
We begin by showing (ref). First, by (ref), $\widehat{G}_{\mathsf{x},n}^{\mathfrak{q},*} \leadsto_{\mathbbm{P}} \mathcal{G}_{\mathsf{x}}$. Also, by Assumption (ref), $\widetilde{M}_{\mathsf{x},n}^{\mathfrak{q},*} \leadsto_{\mathbbm{P}} \mathcal{M}_{\mathsf{x}}^{\mathfrak{q}}$. Moreover, $\widehat{L}_{\mathsf{x},n}^{\mathfrak{q},*} - \widehat{L}_{\mathsf{x},n}^{\mathfrak{q}} \leadsto_{\mathbbm{P}} 0$ by (ref), where, as shown in the proof of ((ref)), $\widehat{L}_{\mathsf{x},n}^{\mathfrak{q}} \leadsto \mathcal{L}_{\mathsf{x}}$. In particular, $\widehat{L}_{\mathsf{x},n}^{\mathfrak{q},*} \leadsto_{\mathbbm{P}} \mathcal{L}_{\mathsf{x}}$ and therefore
Because $\widehat{G}_{\mathsf{x},n}^{\mathfrak{q},*} + \widetilde{M}_{\mathsf{x},n}^{\mathfrak{q}}$ is asymptotically equicontinuous,
by (ref) and Lemma (ref). Also, by (ref) and Lemma (ref),
The result (ref) follows from the three preceding displays and the fact that, on $\widehat{V}_{\mathsf{x},n}^{\mathfrak{q},*}$,
Next, to show (ref), we first define $\widehat{\theta}_n(\mathsf{x};t) = \widehat{\theta}_n(\mathsf{x}) + t r_n^{-1}$ and note that
Now, if $|\widehat{v}_n^*(t)| > a_n \delta > 0,$ then
where $|\widehat{\theta}_n(\mathsf{x};t) - \theta_0(\mathsf{x})| = O_{\mathbbm{P}}(r_n^{-1}) = o_{\mathbbm{P}}(1)$, and, by (ref),
Therefore, defining $\widehat{\mathsf{x}}_n^* = \widehat{\Phi}_n^{*-}(\widehat{\Phi}_n^{*}(\mathsf{x})) = \mathsf{x} + o_{\mathbbm{P}}(1)$ and using ((ref)),
where the last equality uses (ref), $\widehat{Z}_{\mathsf{x},n}^{\mathfrak{q},*} = o_{\mathbbm{P}}(1)$, and Assumption (ref). Also, using (ref), (ref), and Assumption (ref), we have, uniformly in $x \not \in I_\mathsf{x}^\delta$ and for some $c > 0$,
and therefore, by Lemma (ref),
As a consequence, $\widehat{v}_n^*(t) = o_{\mathbbm{P}}(a_n)$: For any $\delta > 0$,
Next, defining $\widehat{V}_{\mathsf{x},n}^{\mathfrak{q},*}(j) = \{v\in \widehat{V}_{\mathsf{x},n}^{\mathfrak{q},*} : 2^j < |v| \leq 2^{j+1}\}$ and using $\widehat{v}_n^*(t) = o_{\mathbbm{P}}(a_n)$, we have, for any $K$, any positive $\delta'$, and any sequence of events $\{\mathcal{A}_n'\}$ with $\lim_{n \to \infty}\mathbbm{P}[\mathcal{A}_n'] = 1$,
The proof of (ref) can therefore be completed by showing that the majorant side in the display can be made arbitrarily small by choice of $K$, $\delta'$, and $\{\mathcal{A}_n'\}$.
To do so, we begin by analyzing each term in the basic bound
Because $\widehat{H}_{\mathsf{x},n}^{\mathfrak{q},*}(0;t) \leadsto_{\mathbbm{P}} \mathcal{H}_{\mathsf{x}}^{\mathfrak{q}}(0;t) = 0$ and because, by (ref) and Lemma (ref), there is a positive $\delta'$ such that
we may assume that, on $\{\mathcal{A}_n'\}$ and for some $C_0$,
Also, because, by (ref) and (ref), there is a positive $\delta'$ such that
we may assume that, on $\{\mathcal{A}_n'\}$ and for some $C_L$,
Next, by (ref) and Lemma (ref), with probability approaching one,
while, by Assumption (ref), there is a positive $c$ such that, with probability approaching one,
We may therefore assume that, on $\{\mathcal{A}_n'\}$ and for some positive $C_M$,
Finally, by (ref) and Lemma (ref), with probability approaching one,
and we may therefore assume that, on $\{\mathcal{A}_n'\}$,
As a consequence, by the Markov inequality,
where, by (ref), we may assume that, for some $C_G$,
and where, for all sufficiently large $j$,
In other words, for large $K$,
which can be made arbitrarily small by choice of $K$.
\paragraph*{Proof of (ref)} The bootstrap consistency result (ref) follows from (ref), (ref), Polya's theorem, and the fact that, by Lemma (ref), the limiting distribution in (ref) and (ref) has a continuous cdf. \qed
For the monomial approximation estimator, we have
where the second equality uses $\epsilon_n \to 0$ and the last equality uses $n \epsilon_n^{1 + 2 \mathfrak{q}} \to \infty$.
Similarly, for the forward difference estimator, we have
Proceeding as in the proof of Lemma (ref), we have
where the second equality uses $\epsilon_n \to 0$ and the defining property of $\{\lambda_j^{\mathtt{BR}}(k) : k = 1,\dots,\underline{\mathfrak{s}}\}$.
The second part of the lemma follows from the fact that if
then
\qed
In addition to the assumptions of Lemma (ref), suppose that $\widehat{R}_{\mathsf{x},n}(1;\eta_n) = O_{\mathbbm{P}}(a_n^{-1/2})$ for $a_n^{-1}\eta_n^{-1}=O(1)$ and that, for some $\delta > 0$, $\theta_0$ is $(\underline{\mathfrak{s}} + 1)$-times continuously differentiable and $\Phi_0$ is $(\underline{\mathfrak{s}} + 2)$-times continuously differentiable on $I_{\mathsf{x}}^{\delta}$. Then, the first term in the stochastic expansion of $\widetilde{\mathcal{D}}_{j,n}^{\mathtt{BR}}(\mathsf{x})$ satisfies
Also, the approximate variance of
is
Finally, the third term in the stochastic expansion of $\widetilde{\mathcal{D}}_{j,n}^{\mathtt{BR}}(\mathsf{x})$ is asymptotically negligible under the condition that $\widehat{R}_{\mathsf{x},n}(1;\eta_n) = O_{\mathbbm{P}}(a_n^{-1/2})$ for $a_n^{-1}\eta_n^{-1}=O(1)$, while the fourth term exhibits only a higher-order dependence on $\epsilon_n$ (relative to the dependence exhibited by the first two terms).
We verify that Assumptions (ref) and (ref) imply Assumptions (ref)-(ref). Define \[\bar{\Gamma}_n^*(x) = \frac{1}{n} \sum_{i = 1}^nW_{i,n} \gamma_0(x;\mathbf{Z}_i) \qquad \text{and} \qquad \bar{\Phi}_n^*(x) = \frac{1}{n} \sum_{i = 1}^nW_{i,n} \phi_0(x;\mathbf{Z}_i).\]
\paragraph*{Non-bootstrap weak convergence} We first prove $\widehat{G}_{\mathsf{x},n}^{\mathfrak{q}}\rightsquigarrow \mathcal{G}_{\mathsf{x}}$. By Assumption (ref)-(ref),
for each $K>0$, and thus, defining $\bar{\psi}_{\mathsf{x},n}(v;\mathbf{Z}_i) = \sqrt{a_n}\psi_{\mathsf{x}}(va_n^{-1};\mathbf{Z}_i)$,
Next, we prove weak convergence of the empirical process indexed by $\{\bar{\psi}_{\mathsf{x},n}(v;\cdot):|v|\leq K\}$ by verifying finite-dimensional weak convergence and stochastic equicontinuity.
Letting $\eta_n = Ka_n^{-1}$,
Also, convergence of the covariance kernel is imposed in Assumption (ref). Thus, the Lyapunov central limit theorem implies finite-dimensional weak convergence.
For stochastic equicontinuity, proceeding as in Kim-Pollard_1990_AoS and using $a_n\mathbbm{E}[\bar{D}_{\gamma}^{\eta_n} (\mathbf{Z})^2 + \bar{D}_{\phi}^{\eta_n} (\mathbf{Z})^2]=O(1)$, it suffices to show that
for any $\epsilon_n=o(1)$. For any $M>0$, using $|\bar{\psi}_{\mathsf{x},n}(v;\mathbf{z})|\leq \sqrt{a_n}(1+\theta_0(\mathsf{x}))(\bar{D}_{\gamma}^{\eta_n} (\mathbf{z}) + \bar{D}_{\phi}^{\eta_n} (\mathbf{z}) )$
where $\bar{M}=2M(1+|\theta_0(\mathsf{x})|)$, the expectation of the first term after the inequality can be made arbitrarily small by making $M$ large using $a_n \mathbbm{E}[ \bar{D}_{\gamma}^{\eta_n} (\mathbf{Z})^4 + \bar{D}_{\phi}^{\eta_n} (\mathbf{Z})^4] = O(1)$, the second term is $o(1)$ by Assumption (ref), and the third term is $O_{\mathbbm{P}}( \sqrt{a_n/n} )$ by Theorem 4.2 of Pollard_1989_SS.
\paragraph*{Bootstrap weak convergence} We next prove $\widehat{G}_{\mathsf{x},n}^{\mathfrak{q}}\rightsquigarrow_{\mathbbm{P}}\mathcal{G}_{\mathsf{x}}$. As shown below, we have
for each $K>0$. Therefore, using $\widehat{\theta}_n(\mathsf{x})\to_{\mathbbm{P}}\theta_0(\mathsf{x})$ and the fact that, uniformly over $| v| \leq K$,
we obtain
where $\widehat{\psi}_{\mathsf{x},n}(v;\mathbf{Z}) = \bar{\psi}_{\mathsf{x},n}(v;\mathbf{Z}) - n^{-1}\sum_{j=1}^n \bar{\psi}_{\mathsf{x},n}(v;\mathbf{Z}_j)$.
To prove finite-dimensional weak convergence, we apply Lemma 3.6.15 of vanderVaart-Wellner_1996_Book. Assumption (ref) implies that $n^{-1}\sum_{i=1}^n(W_{i,n}-1)^2\to_{\mathbbm{P}}1$ and $n^{-1}\max_{1\leq i \leq n}W_{i,n}^2=o_{\mathbbm{P}}(1)$. Since
and $\sup_{|v|\leq\eta } |\psi_{\mathsf{x}}(v;\mathbf{Z})|\leq \bar{D}_{\gamma}^{\eta}(\mathbf{Z}) + |\theta_0(\mathsf{x})|\bar{D}_{\phi}^{\eta}(\mathbf{Z})$, for any $v,u\in\mathbb{R}$,
Also, $n^{-1}\sum_{i=1}^n \widehat{\psi}_{\mathsf{x},n}^4(v;\mathbf{Z}_i)=O_{\mathbbm{P}}(1)$, verifying the hypothesis of the lemma.
For stochastic equicontinuity, let $\epsilon_n=o(1)$ and $\eta_n=K a_n^{-1}$. Lemma 3.6.7 of vanderVaart-Wellner_1996_Book implies that for any $n_0 \in \{1,\dots, n\}$, there is a fixed constant $C>0$ such that
where $(R_1,\dots,R_n)$ is uniformly distributed on the set of all permutations of $\{1,\dots,n\}$, independent of $\{\mathbf{Z}_i\}_{i=1}^n$. Choose $n_0$ such that $n^{1/2-1/\mathfrak{r}}/n_0\to \infty$ and $n_0/a_n\to \infty$ (which is possible by $\mathfrak{r}>(4\mathfrak{q}+2)/(2\mathfrak{q}-1)$), and the first term after the inequality in the above display is $o_{\mathbbm{P}}(1)$. For the second term, following the argument of vanderVaart-Wellner_1996_Book, it suffices to bound
where $\{\mathbf{Z}_i^*\}_{i=1}^k$ denotes a random sample from the empirical cdf and $\mathbbm{E}_n^*$ is the expectation under this empirical bootstrap law. Following the argument of Kim-Pollard_1990_AoS, it suffices to show that
For $k\in \{n_0,\dots,n\}$ and $M>0$,
where the second term after the inequality is shown to be $o_{\mathbbm{P}}(1)$ in the non-bootstrap case above. For the first term after the inequality,
where $\bar{M}=2M(1+|\theta_0(\mathsf{x})|)$. The first term after the inequality does not depend on $k$ and its expectation can be made arbitrarily small by taking $M$ sufficiently large. The second term is independent of $k$ and we can handle this term by adding and subtracting the expectation inside the summation. For the third term, applying Theorem 4.2 of Pollard_1989_SS again, it is bounded by a constant multiple of
which is $o_{\mathbbm{P}}(1)$ by the choice of $n_0$.
\paragraph*{Verifying (ref)} We focus on the first display. By adding and subtracting the bootstrap means,
where
By Assumption (ref),
Identical to above, Lemma 3.6.7 and the argument in Theorem 3.6.13 of vanderVaart-Wellner_1996_Book imply that for some fixed $C>0$,
By Assumption (ref),
Also, Corollary 4.3 of Pollard_1989_SS implies that for some fixed $C>0$,
which is $o_{\mathbbm{P}}(a_n^{-1})$ by Assumption (ref).
Defining
we have
Assumptions (ref) and (ref) imply that for $V \in [1,a_n \delta]$,
and
where $\beta = \beta_\gamma \lor \beta_{\phi} , A_n = A_{\gamma,n} \lor A_{\phi,n} = o_{\mathbbm{P}}(1)$, and $B_n = B_{\gamma,n} \lor B_{\phi,n} = o_{\mathbbm{P}}(a_n^{\beta})$. As a consequence, there exists $\eta_n' = o(1)$ such that
we take the event in the display to be $\mathcal{A}_n$. Also,
and, using Corollary 4.3 of Pollard_1989_SS,
Therefore,
implying in particular that
For the bootstrap counterpart, defining
we have
where bounds for the third and fourth terms were derived in the non-bootstrap case and where the last term is $o_{\mathbbm{P}}(1)$ uniformly over $| v| \leq a_n \delta$ as $\sqrt{a_n}[\widehat{\theta}_n(\mathsf{x}) - \theta_0(\mathsf{x})]=o_{\mathbbm{P}}(1)$ and $\sqrt{n} \sup_{| v|\leq \delta}|\bar{\Phi}_n^*(\mathsf{x} + v) - \bar{\Phi}_n(\mathsf{x} +v ) |=O_{\mathbbm{P}}(1)$. For the first term after the equality in (ref),
where, by Assumption (ref),
uniformly over $V\in [1,a_n\delta]$ and where, applying Lemma 3.6.7 of vanderVaart-Wellner_1996_Book, we have, for any $n_0\in \{1,\dots, n\}$,
where $C>0$ is a fixed constant. Letting $n_0$ be a diverging sequence (dependent on $n$) such that $n_0 n^{\mathfrak{r}}\sqrt{a_n/n} =o(1)$, the first term on the majorant side is bounded by
and, by Corollary 4.3 of Pollard_1989_SS, the second term is bounded (up to a constant) by
which is $o_{\mathbbm{P}}(1)$ by Assumption (ref). As a consequence, there exists a sequence of random variables $A_n'=o_{\mathbbm{P}}(1)$ such that for $V\in [1,a_n\delta]$,
By identical arguments, an analogous bound holds for the second term after the inequality in (ref). Therefore, there exists $\eta_n'=o(1)$ and events $\mathcal{A}_n$ such that $\lim_{n\to\infty}\mathbbm{P}[\mathcal{A}_n]=1$ and
for any $V\in [1,a_n\delta]$. Finally, proceeding as in the proof of stochastic equicontinuity, we obtain the bound $\mathbbm{E}[\sup_{| v|\in [V,2V] }|\bar{G}_{\mathsf{x},n}^{\mathfrak{q},*}(v)|]\leq C\sqrt{V}$.
We have
where the first term after the inequality is assumed to be $o_{\mathbbm{P}}(1)$ and the second term is $o_{\mathbbm{P}}(1)$ by standard arguments. Similarly, $\sup_{x\in I}|\widehat{\Phi}_n(x)-\Phi_0(x)|=o_{\mathbbm{P}}(1)$. Also,
where the last two terms are $o_{\mathbbm{P}}(a_n^{-1})$ and where Assumption (ref) implies
the last equality using $B_{\phi,n}=o_{\mathbbm{P}}(a_n^{\beta_{\phi}})$ and $\beta_{\phi}\leq \mathfrak{q}$.
Now we look at the bootstrap objects. For $\widehat{\Gamma}_n^*$,
where the last term is assumed to be $o_{\mathbbm{P}}(1)$ and where
For $\bar{\Gamma}_n^* -\bar{\Gamma}_n$, using Lemma 3.6.7 of vanderVaart-Wellner_1996_Book and the same argument as for verifying Assumption (ref), it suffices to show that
as can be done using Corollary 4.3 of Pollard_1989_SS.
For $\widehat{\Phi}_n^*$, $\sup_{x\in I}|\widehat{\Phi}_n^*(x)-\widehat{\Phi}_n(x)|=o_{\mathbbm{P}}(1)$ follows from the same argument as for $\widehat{\Gamma}_n^*$. For $a_n \sup_{x\in I_{\mathsf{x}}^{\delta}}|\widehat{\Phi}_n^*(x) -\widehat{\Phi}_n(x)|=o_{\mathbbm{P}}(1)$,
where the last term is $o_{\mathbbm{P}}(a_n^{-1})$ as shown above and the second and third terms after the inequality are $O_{\mathbbm{P}}(n^{-1/2})$ by standard arguments. For the remaining term,
where $\Breve{\phi}_n(x;\mathbf{Z}) = \widehat{\phi}_n(x;\mathbf{Z})-\phi_0(x;\mathbf{Z}) - [\check{\Phi}_n(x)-\bar{\Phi}_n(x)]$. The last two terms are $o_{\mathbbm{P}}(a_n^{-1})$ by Assumption (ref). Using Lemma 3.6.7 of vanderVaart-Wellner_1996_Book and the argument similar to above, the remaining term is $o_{\mathbbm{P}}(a_n^{-1})$.
The proof is by contradiction and follows Kosorok_2008_BookCh. We omit some details in cases where the arguments are almost identical to those for Theorem (ref) and Lemma (ref).
Suppose that the bootstrap approximation is consistent; that is, suppose
Then, by Theorem 2.2 of Kosorok_2008_BookCh, we have
where $=_d$ denotes the distributional equality, $Y_1$ and $Y_2$ are independent copies of $Y$, and where the convergence in distribution is unconditional.
Using the switching lemma, $\mathbbm{P} \left[r_n \left( \widehat{\theta}_n^*(\mathsf{x}) - \theta_0(\mathsf{x}) \right) > t \right]$ equals
By the arguments used in the proof of Theorem (ref), to characterize the limiting distribution of $r_n(\widehat{\theta}_n^*(\mathsf{x}) - \theta_0(\mathsf{x}))$, it suffices to look at
where
and
It can be shown that $\check{G}_{\mathsf{x},n}^{\mathfrak{q},*} \leadsto_{\mathbbm{P}} \mathcal{G}_{\mathsf{x}}, \widehat{L}_{\mathsf{x},n}^{\mathfrak{q},*} \leadsto \mathcal{L}_{\mathsf{x}}$, and that $\check{M}_{\mathsf{x},n}^{\mathfrak{q}} \leadsto \mathcal{G}_{\mathsf{x}} + \mathcal{M}_{\mathsf{x}}^{\mathfrak{q}}$. Thus,
where $\mathcal{G}_{\mathsf{x},1}$ and $\mathcal{G}_{\mathsf{x},2}$ are independent copies of $\mathcal{G}_{\mathsf{x}}$. Noting that $\mathcal{G}_{\mathsf{x}}(a v)=_d \sqrt{|a|} \mathcal{G}_{\mathsf{x}}(v)$ and using the change of variable $v=u2^{\frac{1}{2\mathfrak{q}+1}}$, the limit distribution equals
As a consequence,
contradicting (ref) because $2^{\frac{\mathfrak{q}}{2\mathfrak{q}+1}} \neq \sqrt{2}$.
In other words, the bootstrap estimator $\widehat{\theta}_n^*(\mathsf{x})$ fails to approximate the limit distribution. $\qedsymbol$
Below we verify the hypothesis of Theorem (ref) for various examples. For this purpose, it suffices to verify Assumptions (ref), (ref)-(ref), and (ref) since Assumption (ref) implies (ref)-(ref) by Lemma (ref).
When $\gamma_0$ is known, it is natural to take $\widehat{\Gamma}_n=\check{\Gamma}_n=\bar{\Gamma}_n$, in which case (ref) reduces to the requirement that, for some $\rho_{\gamma}\in (0,2)$,
An identical remark applies to $\phi_0$ and (ref).
Assumption (ref) and (ref)-(ref) follow from the hypothesis. \newline (ref): In this example, $\gamma_0(x;\mathbf{Z})=\mathbbm{1}(X\leq x)$ is known, so it suffices to verify (ref). The uniform covering number of $\{\mathbbm{1}(\cdot\leq x) :x\in\mathbb{R}\}$ grows linearly, and an envelope function can be taken to be $1$. For an envelope function of $\{\mathbbm{1}(\cdot\leq x)-\mathbbm{1}(\cdot\leq \mathsf{x}) : |x-\mathsf{x}|\leq \eta \}$, we can take $\mathbbm{1}(-\eta +\mathsf{x}\leq \cdot\leq \mathsf{x}+ \eta )$ and the moment bound is satisfied as $\mathbbm{E}[\mathbbm{1}(-\eta +\mathsf{x}\leq X\leq \mathsf{x}+ \eta )]\leq C \eta$. \newline (ref) trivially holds as $\widehat{\Phi}_n(x)=\widehat{\Phi}_n^*(x)=x$. \newline (ref): Here $\psi_{\mathsf{x}}(v;\mathbf{Z})=\mathbbm{1}(X\leq \mathsf{x}+v)-\mathbbm{1}(X\leq\mathsf{x}) -\Phi_0(\mathsf{x})v$. Then,
Also, $\psi_{\mathsf{x}_n}(s\eta_n;\mathbf{Z})=\mathbbm{1}(\mathsf{x}_n \land (\mathsf{x}_n+s\eta_n) < X\leq \mathsf{x}_n \lor (\mathsf{x}_n+s\eta_n) ) - f_0(\mathsf{x}) s\eta_n$ and
Then, for any $s,t\in\mathbb{R}$ and $\mathsf{x}_n\to\mathsf{x}$, using continuity of $f_0$ at $\mathsf{x}$,
(ref) holds since $\widehat{u}_n=\widehat{u}_n^*$ converges in probability to $u_0$, the supremum of the support of $X$ by i.i.d.\ assumption. \newline (ref) and (ref) hold trivially since $\widehat{\Phi}_n$ and $\widehat{\Phi}_n^*$ are the identity map. \newline Assumption (ref) follows from (ref) and empirical process theory arguments.
Assumption (ref) and (ref)-(ref) follow from the hypothesis. \newline (ref): We have $\widehat{\Gamma}_n=1-\widehat{S}_n$ with $\widehat{S}_n$ the Kaplan-Meier estimator. By Theorem 1 of Lo-Singh_1986_PTRF,
Since $\sqrt{n a_n} \leq n^{2/3}$ for $\mathfrak{q}\geq 1$, $\sup_{x \in I}|\widehat{\Gamma}_n(x)-\Gamma_0(x)|=o_{\mathbbm{P}}(1)$ and $\sqrt{na}\sup_{|v|\leq \delta}|\widehat{\Gamma}_n(\mathsf{x}+v)-\widehat{\Gamma}_n(\mathsf{x}) - \bar{\Gamma}_n(\mathsf{x}+v)+\bar{\Gamma}_n(\mathsf{x})|=o_{\mathbbm{P}}(1)$ hold.
We have
By $S_0(u_0)G_0(u_0)>0$, we have $\sqrt{n}\sup_{x\in I}|\widehat{S}_n(x)-S_0(x)|=O_{\mathbbm{P}}(1)$, $\sqrt{n}\sup_{x\in I}|\widehat{G}_n(x)-G_0(x)|=O_{\mathbbm{P}}(1)$, and $\sqrt{n}\sup_{x\in I}|\widehat{\Lambda}_n(x)-\Lambda_0(x)|=O_{\mathbbm{P}}(1)$, which in turn implies
Let $\delta = \min\{\mathsf{x}, (u_0-\mathsf{x})\}/4$, $R_{1n}(v)=| \widehat{F}_n(\mathsf{x}+v)-\widehat{F}_n(\mathsf{x}) - F_0(\mathsf{x}+v)+F_0(\mathsf{x})|$, $R_{2n}=| \widehat{F}_n(\mathsf{x})-F_0(\mathsf{x})|$, $R_{3n}= \sup_{x\in I}| [\widehat{S}_n(x)\widehat{G}_n(x)]^{-1} - [S_0(x)G_0(x)]^{-1} |$, and $R_{4n}(x_1,x_2) = |\int_{x_1}^{x_2} \frac{\widehat{\Lambda}_n(du)}{\widehat{S}_n(u)\widehat{G}_n(u)} -\int_{x_1}^{x_2} \frac{\Lambda_0(du)}{S_0(u)G_0(u)}| $. For $|v|\leq \delta$,
As noted above, $\sup_{|v|\leq \delta}R_{1n}(v)=o_{\mathbbm{P}}( (na_n)^{-1/2})$, $R_{2n}=O_{\mathbbm{P}}(n^{-1/2})$, and $R_{3n}=O_{\mathbbm{P}}(n^{-1/2})$ by $S_0(u_0)G_0(u_0)>0$. Also, uniformly over $V\in (0, 2\delta]$, $\sup_{| v|\leq V}| \frac{1}{n}\sum_{i=1}^n (\mathbbm{1}(X_i\leq \mathsf{x} +v) - \mathbbm{1}(X_i\leq \mathsf{x}) )| \leq C V + O_{\mathbbm{P}}(n^{-1/2})$. If
and
uniformly over $|v|\leq 2\delta$, then there exist random variables $A_n=o_{\mathbbm{P}}(1)$ and $B_n=O_{\mathbbm{P}}(\sqrt{a_n})$ independent of $v$ such that for $V\in (0, 2\delta]$,
that is,\ $\beta_{\gamma}=1$ in the notation of (ref). To show (ref) and (ref), for $x_1,x_2\in I$,
Let $J_0(u)= 1/[S_0(u)G_0(u)]$ and integration by parts implies
The first term after the second equality is bounded by $O_{\mathbbm{P}}(n^{-1/2}) \vert x_2-x_1\vert$. For the second term, Theorem 1 of Burketal implies that on a suitable probability space there exists a sequence of standard Brownian motion $W_n$ such that $\sqrt{a_n}\sup_{\vert x_1-x_2\vert\leq v} \vert \sqrt{n}[\widehat{\Lambda}_n(x_1)-\Lambda_0(x_2) -\widehat{\Lambda}_n(x_1)+\Lambda_0(x_2)] - W_n(d(x_1))+W_n(d(x_2))\vert = o_{\mathbbm{P}}(1)$, where $d(x)=\int_0^x \frac{F_0(du)}{S_0(u)^2G_0(u)}$. By Theorem 3.2 of Pollard_1989_SS, there is some fixed constant $C>0$ such that $\mathbbm{E}[\sup_{\vert x_1-x_2\vert\leq v}\vert W_n(d(x_1))-W_n(d(x_2))\vert]\leq C v$. Finally, $\vert \int_{x_1}^{x_2} [\widehat{\Lambda}_n(u)-\Lambda_0(u)]J_0(du)\vert\leq \sup_{x'\in [x_1,x_2]}\vert\widehat{\Lambda}_n(x')-\Lambda_0(x')\vert [J_0(x_2)-J_0(x_1)]\leq O_{\mathbbm{P}}(n^{-1/2}) \vert x_2-x_1\vert$. Thus, (ref) and (ref) hold.
For the function class $\mathfrak{F}_{\gamma}$, we can take $\bar{F}_{\gamma}(\mathbf{Z}) = 1 + [S_0(u_0)G_0(u_0)]^{-1}[1 + \Lambda_0(u_0)]$ as a constant envelope. For the function class $\{S_0(x):x\in I\}$, given $m\in\mathbbm{N}$, there exists $\{x_1,\dots, x_{m+1}\}\subset I$ such that $\sup_{x\in I}\min_{l=1,\dots,m+1}|S_0(x_l)-S_0(x)|\leq 1/m$, which implies the uniform covering number is bounded by a linear function. The covering numbers of $\{\mathbbm{1}(\cdot\leq s):s\in I\}$ and $\{\int_0^{\cdot\land s}[S_0(u)G_0(u)]^{-1}\Lambda_0(du):s\in I\}$ are also bounded by a linear function. By Lemma 5.1 of vanderVaart-vanderLaan_2006_IJB, there exists $\rho\in (0,2)$ such that $\limsup_{\eta\downarrow 0}\log N_U(\eta,\mathfrak{F}_{\gamma})\eta^{\rho} < \infty$ holds.
Now consider the uniform covering number of $\hat{\mathfrak{F}}_{\gamma}$. Given a realization of $(\widehat{S}_n,\widehat{G}_n)$, the mapping $x\mapsto\int_0^{x\land s}[\widehat{S}_n(u)\widehat{G}_n(u)]^{-1}\widehat{\Lambda}_n(du)$ is a composition of $x\mapsto x\land s$ and $x\mapsto \int_0^x [\widehat{S}_n(u)\widehat{G}_n(u)]^{-1}\widehat{\Lambda}_n(du)$. The latter mapping is monotone, and the first mapping is a VC-subgraph class, and Lemma 2.6.18 of vanderVaart-Wellner_1996_Book implies $\{\int_0^{\cdot\land s}[\widehat{S}_n(u)\widehat{G}_n(u)]^{-1}\widehat{\Lambda}_n(du):s\in I\}$ is a VC-subgraph class. Note that since $S_0,G_0$ are bounded away from zero, $\widehat{S}_n,\widehat{G}_n$ are bounded away from zero with probability approaching one. Thus, for some $\rho\in (0,2)$, $\limsup_{\eta\downarrow 0}\log N_U(\eta,\hat{\mathfrak{F}}_{\gamma})\eta^{\rho} = O_{\mathbbm{P}}(1)$ holds.
For $s\leq t\in I$,
and we can take $D_{\gamma}^{\eta}(\mathbf{Z})$ to be a constant multiple of $\sup_{|s|\leq \eta}|F_0(\mathsf{x} +s)-F_0(\mathsf{x})| + \Delta \mathbbm{1}(|\check{X}-\mathsf{x}|\leq \eta) + \int_{ x-\eta}^{x+\eta}\Lambda_0(du)/S_0(u)G_0(u)$. For $\eta >0$ small enough, there is some fixed $C>0$ with
(ref) trivially holds as $\widehat{\Phi}_n(x)=\widehat{\Phi}_n^*(x)=x$. \newline (ref): We have
where $O(|v|)$ is uniformly over small enough $|v|$. Since
the first display in (ref) is satisfied. For the covariance kernel,
where the last equality uses continuity of $(S_0,G_0,f_0)$ at $\mathsf{x}$ i.e.,\ $\int_{\mathsf{x}_n}^{\mathsf{x}_n +\eta_n} [\frac{f_0(u)}{S_0(u)^2G_0(u)}-\frac{f_0(\mathsf{x})}{S_0(\mathsf{x})^2G_0(\mathsf{x})}]du =o(1)\eta_n$.
(ref), (ref), and (ref) hold since in this example, $\widehat{u}_n=\widehat{u}_n^*=u_0$ and $\widehat{\Phi}_n,\widehat{\Phi}_n^*$ are the identity map.
\paragraph*{Assumption (ref)} As noted when verifying (ref), $\widehat{G}_n(1;\eta_n)=o_{\mathbbm{P}}(1)$ for any $\eta_n=o(1)$ with $a_n^{-1}\eta_n^{-1}=O(1)$. $\widehat{\Phi}_n=\Phi_0$ is the identity map and the desired result holds.
Assumption (ref) and (ref)-(ref) follow from the hypothesis. \newline (ref): In this example, $\gamma_0(x;\mathbf{Z})=Y\mathbbm{1}(X\leq x)$ is known, so it suffices to verify (ref). The uniform covering number bound is straightforward as $\{\mathbbm{1}(\cdot\leq x):x\in\mathbb{R}\}$ is a VC-subgraph class. An envelope function is $|Y|$, whose second moment is finite. For $x\in I_{\mathsf{x}}^{\eta}$, $|\gamma_0(x;\mathbf{Z})-\gamma_0(\mathsf{x};\mathbf{Z})|\leq |Y| \mathbbm{1}(\mathsf{x}-\eta \leq X\leq \mathsf{x} +\eta)$, which we can take as $\bar{D}_{\gamma}^{\eta}(\mathbf{Z})$. Then, for $j=2,4$,
and the desired bound holds. \newline (ref): $\phi_0(x;\mathbf{Z})=\mathbbm{1}(X\leq x)$ is known, so it suffices to verify the analogue of (ref). The argument is the same as for checking (ref) in monotone density estimation with no censoring.
(ref): We have
Then,
and the first display holds. For the covariance kernel, note $|(\mu_0(X)-\mu_0(\mathsf{x}_n))(\mathbbm{1}(X\leq \mathsf{x}_n+v)-\mathbbm{1}(X\leq \mathsf{x}_n)|\leq |v| \sup_{|x-\mathsf{x}|\leq 2\eta} |\partial\mu_0(x)|$ for $|x_n-\mathsf{x}|\lor |v|\leq \eta$ for $\eta>0$ small enough. Then,
and
as desired.
(ref) trivially holds since $\widehat{u}_n=\widehat{u}_n^*=1$ in this example.
(ref): $\widehat{\Phi}_n,\widehat{\Phi}_n^*$ are (empirical) cdfs, so they are non-negative, non-decreasing, and right-continuous. $\{0,\widehat{u}_n\}\subset \widehat{\Phi}_n(I)$ and $\{0,\widehat{u}_n^*\}\subset \widehat{\Phi}_n^*(I)$ hold as $\widehat{u}_n=\widehat{u}_n^*=1$, $\widehat{\Phi}_n(\min_iX_i-)=0=\widehat{\Phi}_n^*(\min_iX_i-)$, and $\widehat{\Phi}_n(\max_iX_i)=1=\widehat{\Phi}_n^*(\max_iX_i)$. The sets $\widehat{\Phi}_n(I),\widehat{\Phi}_n^*(I)$ are finite and thus closed.
(ref): With probability one, all $X_i$'s are distinct. If $x$ is one of $X_i$'s, $\widehat{\Phi}_n^*( x ) - \widehat{\Phi}_n^*( x- ) = n^{-1}W_{j,n}$ for some $j$, and $\widehat{\Phi}_n^*( x ) - \widehat{\Phi}_n^*( x- ) =0$ otherwise.
Given $\mathbbm{E}|W_{1,n}|^{\mathfrak{r}}<\infty$, $n^{-1}\max_{1\leq i\leq n}|W_{i,n}|=o_{\mathbbm{P}}(n^{-5/6})$, which implies $$\sqrt{na_n}\sup_{x\in I}|\widehat{\Phi}_n^*( x ) - \widehat{\Phi}_n^*( x- ) |=o_{\mathbbm{P}}(1).$$ The argument for $\widehat{\Phi}_n$ is similar. \newline Assumption (ref) follows from (ref)-(ref) and empirical process theory arguments.
Assumption (ref) and (ref)-(ref) follow from the hypothesis.
(ref): In this example, $\check{\Gamma}_n=\widehat{\Gamma}_n$.
The last sum is bounded by $\sup_{x\in I}|\frac{1}{n}\sum_{j=1}^n\mu_0(x,\mathbf{A}_j)-\theta_0(x)|$, and this object is $O_{\mathbbm{P}}(n^{-1/2})$: to see this claim, first note that Assumption MRC (iv) and Theorem 2.7.11 of vanderVaart-Wellner_1996_Book imply $\limsup_{\epsilon\downarrow0}\log N_U(\epsilon,\{\mu(x,\cdot):x\in I\}) \epsilon^V <\infty$ for some $V\in (0,2)$ and Theorem 4.2 of Pollard_1989_SS implies $\sup_{x\in I}|\frac{1}{n}\sum_{j=1}^n\mu_0(x,\mathbf{A}_j)-\theta_0(x)|=O_{\mathbbm{P}}(n^{-1/2})$. Together with Assumption MRC (iii), $a_n \frac{1}{n}\sum_{i=1}^n \sup_{x\in I}|\widehat{\gamma}_n(x;\mathbf{Z})-\gamma_0(x;\mathbf{Z})|^2=o_{\mathbbm{P}}(1)$ holds.
By Assumption MRC (iii), uniformly over $V\in (0,2\delta]$
and the desired inequality holds.
The uniform covering numbers of $\mathfrak{F}_{\gamma},\hat{\mathfrak{F}}_{\gamma,n}$ are the same order as for $\{\mathbbm{1}(\cdot\leq x):x\in I\}$. For $x \in I_{\mathsf{x}}^{\eta}$, $|\gamma_0(x;\mathbf{Z})-\gamma_0(\mathsf{x};\mathbf{Z})|\leq \mathbbm{1}(\mathsf{x}-\eta\leq X\leq \mathsf{x}+\eta)(|\varepsilon| c^{-1}+\theta_0(\mathsf{x}+\eta) )$. Then, $\limsup_{\eta\downarrow0}\mathbbm{E}[\bar{D}_{\gamma}^{\eta}(\mathbf{Z})^j]\eta^{-1} <\infty$ holds for $j=2,4$.
(ref): $\phi_0(x;\mathbf{Z})=\mathbbm{1}(X\leq x)$ is known and the same as in the classical case, so the same argument applies.
(ref): We have
Then, for $v,v'\in [-\eta,\eta]$ with sufficiently small $\eta>0$,
and $\sup_{v\neq v'\in [-\eta_n,\eta_n]}\mathbbm{E}[|\psi_{\mathsf{x}}(v;\mathbf{Z}) - \psi_{\mathsf{x}}(v';\mathbf{Z})|]/|v-v'|=O(1)$ holds.
For $s\eta_n$ small enough, $\psi_{\mathsf{x}}(s\eta_n;\mathbf{Z})= (\mathbbm{1}(X\leq \mathsf{x}+ s\eta_n)-\mathbbm{1}(X\leq \mathsf{x}))\varepsilon g_0(X,\mathbf{A})^{-1} + O(\eta_n)$ and
and
Since $\frac{f_{X|\mathbf{A}}(x|\mathbf{A})}{g_0(x,\mathbf{A})^2} = \frac{f_0(x)}{g_0(x,\mathbf{A})}$, we have
as desired.
(ref) (ref) (ref): Verifying these conditions is the same as in the classical monotone regression case.
Here we provide primitive sufficient conditions for Assumption MRC (iii) by focusing on specific estimators $\widehat{\mu}_n$ and $\widehat{g}_n$. As discussed by Westling-Gilbert-Carone_2020_JRRSB, cross-fitting avoids restrictions on uniform entropy, allowing for a large class of flexible preliminary estimators. Here we use sample splitting to simplify exposition, but the proposed procedure can be straightforwardly modified for cross-fitting.
Suppose there is a separate random sample $\tilde{\mathbf{Z}}_1,\dots,\tilde{\mathbf{Z}}_n$ drawn from the distribution of $\mathbf{Z}$, which is independent of $\mathbf{Z}_1,\dots,\mathbf{Z}_n$. Preliminary estimators $\widehat{\mu}_n$ and $\widehat{g}_n$ are constructed from $\tilde{\mathbf{Z}}_1,\dots,\tilde{\mathbf{Z}}_n$. For concreteness, we consider a partitioning-based least squares estimator $\widehat{\mu}_n$ Cattaneo-Farrell-Feng_2020_AoS and local polynomial kernel-based estimators $\widehat{f}_{X|\mathbf{A},n}(x|\mathbf{a})$ and $\widehat{f}_n(x)$ of $f_{X|\mathbf{A}}(x|\mathbf{a})$ and $f_0(x)$ \citep*{Cattaneo-Jansson-Ma_2020_JASA,Cattaneo-Chandak-Jansson-Ma_2024}, from which we construct $\widehat{g}_n(x,\mathbf{a})=\widehat{f}_{X|\mathbf{A},n}(x|\mathbf{a})/\widehat{f}_n(x)$.
Let $d=\mathrm{dim}(\mathbf{A})$. For simplicity, suppose the support of $(X,\mathbf{A}')'$ equals $[0,1]^{1+d}$. Let $\mathbf{p}(x,\mathbf{a})$ be a $k_n$-dimensional vector of bounded basis functions of order $m$ on $\mathcal{S}$ which are locally supported e.g.,\ splines Cattaneo-Farrell-Feng_2020_AoS. We consider the estimator
For the estimator of $f_{X|\mathbf{A}}(x\vert\mathbf{a})$, letting $\widehat{F}_{X|\mathbf{A},n}(\cdot\vert \mathbf{a})$ be an estimator of $\mathbbm{P}[X\leq \cdot \vert \mathbf{A}=\mathbf{a}]$ specified below, $\widehat{f}_{X|\mathbf{A},n}(x|\mathbf{a})$ is obtained by local polynomial regression:
where $\mathfrak{p}_1\geq 1$ is the order of the polynomial basis $\mathbf{q}_1(x)=(1,x/1!, x^2/2!,\dots,x^{\mathfrak{p}_1}/\mathfrak{p}_1!)'$, $\mathbf{e}_l$ is the conformable unit vector whose $l$th element is unity, and $K_h(x)=K(x/h)/h$ for some kernel function $K$ and some positive bandwidth $h$. The estimator $\widehat{F}_{X|\mathbf{A},n}(x|\mathbf{a})$ is constructed via local polynomial regression of order $\mathfrak{p}_2=\mathfrak{p}_1-1$:
where, using standard multi-index notation, $\mathbf{q}_2(\mathbf{a})$ denotes the $k_{\mathfrak{p}_2}$-dimensional vector collecting the polynomials $\mathbf{a}^{\mathbf{m}}/\mathbf{m}!$ for $0\leq \vert \mathbf{m}\vert\leq \mathfrak{p}_2$ with $\mathbf{a}^{\mathbf{m}} = a_1^{m_1}a_2^{m_2}\dots a_d^{m_d}$, $\vert\mathbf{m}\vert=\sum_{j=1}^dm_j$, and $k_{\mathfrak{p}_2}=\frac{(d+\mathfrak{p}_2)!}{d!\mathfrak{p}_2!}+1$, and $L_h(\mathbf{a})=L(\mathbf{a}/h)/h^d$ for $L(\mathbf{a})=\prod_{j=1}^dK(a_j)$ i.e.,\ product kernel. The estimator $\widehat{f}_n(x)$ is constructed in a similar manner. First, the empirical cdf $\widehat{F}_n$ of $\{\tilde{X}_i\}$ is constructed and then $\widehat{f}_n(x)$ is formed via local polynomial regression:
where $b>0$ is some bandwidth.
Now we state sufficient conditions for Assumption MRC (ii) based on the partitioning-based series estimator $\widehat{\mu}_n$ and the kernel-based estimator $\widehat{g}_n$.
As verified by Cattaneo-Farrell-Feng_2020_AoS, (iii) holds for widely used local basis functions such as splines and wavelets.
\paragraph*{Remark} Using (ref), one can verify the first part of Assumption (ref) i.e.,\ $\widehat{G}_{\mathsf{x},n}(1;\eta_n)=O_{\mathbbm{P}}(1)$. The second part is easy to verify using standard empirical process theory arguments.
We consider the problem of estimating the density of a non-negative, continuously distributed random variable with censoring. We use the same notation as in Section (ref) of the main paper. Relative to Section (ref), we consider the additional complication of censoring being informative about the “survival time” $X$ i.e.,\ $X\not\protect\mathpalette{\protect\independenT}{\perp} C$. With covariates $\mathbf{A}$, we consider the setting of censoring at random: $X \protect\mathpalette{\protect\independenT}{\perp} C|\mathbf{A}$. See vanderlaan-Robins_2003_Book,Zeng_2004_AoS and references therein for existing analysis of this problem. We have
where $F_0(x|A)=1-S_0(x|A)$, $S_0(x|\mathbf{A})=\mathbbm{P}[X>x|\mathbf{A}]$, $G_0(c|\mathbf{A})=\mathbbm{P}[C>c|\mathbf{A}]$, and $\Lambda_0(x|\mathbf{A}) = \int_0^{x} \frac{f_0(u|\mathbf{A})}{S_0(u|\mathbf{A})}du$ with $f_0$ being the Lebesgue density of $X$. Denote by $\widehat{S}_n(\cdot|\cdot)$, $\widehat{G}_n(\cdot|\cdot)$, $\widehat{\Lambda}_n(\cdot|\cdot)$ preliminary estimates of $S_0,G_0,\Lambda_0$, respectively.
\paragraph*{Assumption SA.\arabic{section}} Let $\mathfrak{S}_n$, $\mathfrak{G}_n,\mathfrak{L}_n$ be sequences of function classes that contain $S_0(\cdot|\cdot)$, $G_0(\cdot|\cdot)$, $\Lambda_0(\cdot|\cdot)$, respectively.
The condition (ref) is high-level, and there are a few different approaches to verify them. See Westling-Carone_2020_AoS for details.
In this example, $\widehat{\Phi}_n(x)=\widehat{\Phi}_n^*(x)=x=\Phi_0(x)$. Assumptions (ref) and (ref)-(ref) follow from the hypothesis. \newline (ref): In this example, $\check{\Gamma}_n=\widehat{\Gamma}_n$. For $x\in I$,
and
using integration by parts, where $J_0(u|\mathbf{a}) = [S_0(u|\mathbf{a})G_0(u|\mathbf{a})]^{-1}$. Thus, there is a fixed $C>0$ such that
From the hypothesis,
follow.
For uniform covering numbers, it suffices to show that each of $\{S(x|\cdot):x\in I\}$, $\{\mathbbm{1}(\cdot\leq x):x\in I\}$, and $\{\int_0^{\cdot\land x} \frac{\Lambda(du|\cdot)}{S(u|\cdot)G(u|\cdot)} : x\in I\}$ has an appropriate bound on the uniform covering number by Lemma 5.1 of vanderVaart-vanderLaan_2006_IJB (see examples after the lemma). For $\{\int_0^{\cdot\land x} \frac{\Lambda(du|\cdot)}{S(u|\cdot)G(u|\cdot)} : x\in I\}$ with $(S,G,\Lambda)\in \mathfrak{S}_n\times \mathfrak{G}_n\times\mathfrak{L}_n$, the mapping $x\mapsto \int_0^x \frac{\Lambda(du|\cdot)}{S(u|\cdot)G(u|\cdot)}$ is monotone (by the non-decreasing property of $\Lambda$ and $S,G\geq c_1 >0 $) and Lemma 2.6.18 of vanderVaart-Wellner_1996_Book implies the desired result.
There is a fixed $C>0$ such that for $x\in I_{\mathsf{x}}^{\eta}$,
and using $1-S_0(x|\cdot) =\int_0^x f_{X|\mathbf{A}}(u|\cdot) du$ with $f_{X|\mathbf{A}}$ being bounded, we can take
which satisfies the desired bound condition. \newline (ref) trivially holds since $\widehat{\Phi}_n(x)=\widehat{\Phi}_n^*(x)=x$.
(ref): We have
and the first display follows as in the independent censoring case. For the covariance kernel,
and $\eta_n^{-1}\mathbbm{E}[\psi_{\mathsf{x}_n}(s\eta_n;\mathbf{Z})\psi_{\mathsf{x}_n}(t\eta_n;\mathbf{Z})]$ converges to $\mathbbm{E}[\frac{f_{X|\mathbf{A}}(\mathsf{x}|\mathbf{A})}{G_0(\mathsf{x}|\mathbf{A})}] (|s|\land|t|) \mathbbm{1}(\operatorname*{sign}(s)=\operatorname*{sign}(t))$.
(ref), (ref), and (ref) hold since in this example, $\widehat{u}_n=\widehat{u}_n^*=u_0$ and $\widehat{\Phi}_n,\widehat{\Phi}_n^*$ are the identity map.
Let $X$ be a non-negative random variable, $f_0$ be its Lebesgue density, and $S_0(x)=\mathbbm{P}[X>x]$ be its survival function. We consider the parameter of estimating the hazard function of $X$, $\theta_0(\mathsf{x})=f_0(\mathsf{x})/S_0(\mathsf{x})$, with possible right-censoring as in the monotone density function example. Observations $\mathbf{Z}_1,\dots,\mathbf{Z}_n$ come from a random sample of $\mathbf{Z}=(\check{X},\Delta)'$ where $\check{X}=\min\{X,C\}$ and $\Delta=\mathbbm{1}(X\leq C)$, $C$ being a random censoring time. As pointed out by Westling-Carone_2020_AoS, with strictly increasing $\Phi_0$ with $\Phi_0(0)=0$, the function $\Gamma_0$ takes the form $\Gamma_0(x) = \int_0^x \frac{f_0(u)}{S_0(u)}\Phi_0(du)$, and by taking $\Phi_0(x)=\int_0^xS_0(u)du$, $\Gamma_0(x) =F_0(x)=\mathbbm{P}[X\leq x]$. Since $\Gamma_0$ is identical to the monotone density case with the choice $\Phi_0=\int_0^xS_0(u)du$, we can leverage the analysis for the monotone density. The interval $I$ equals $[0,u_0^{\mathtt{MD}}]$ where $u_0^{\mathtt{MD}}$ is $u_0$ in the monotone density example. The $u_0$ for the monotone hazard function estimation is $u_0=\Phi_0(u_0^{\mathtt{MD}})$.
Consider the case of completely random censoring i.e.,\ $X\protect\mathpalette{\protect\independenT}{\perp} C$. As in the setup for Corollary (ref), let $\widehat{S}_n(x)$ be the Kaplan-Meier estimator for $S_0(x)=1-F_0(x)=\mathbbm{P}[X> x]$, $\widehat{F}_n=1-\widehat{S}_n$, and $\widehat{G}_n$ be the Kaplan-Meier estimator for $G_0(x)=\mathbbm{P}[C>x]$. Also,
and $\phi_0(x;\mathbf{Z}) =x-\int_0^x \gamma_0(u;\mathbf{Z})du$.
We use the same $\widehat{\gamma}_n$ function and assumptions as in the monotone density setting. Also, the covariance kernels are the same as in the monotone density case. We focus on (ref)-(ref) and (ref)-(ref).
(ref): Since
follow from $a_n\frac{1}{n}\sum_{i=1}^n\sup_{x\in I}\vert\widehat{\gamma}_n(x;\mathbf{Z}_i) - \gamma_0(x;\mathbf{Z}_i)\vert^2$, which was verified in Section (ref). To check $\sup_{x\in I}|\widehat{\Phi}_n(x)-\Phi_0(x)|=o_{\mathbbm{P}}(1)$, $\sup_{x\in I}|\frac{1}{n}\sum_{i=1}^n\phi_0(x;\mathbf{Z}_i)-\Phi_0(x)|=o_{\mathbbm{P}}(1)$ follows from Glivenko-Cantelli, and
where the last equality follows from $\frac{1}{n}\sum_{i=1}^n\sup_{x\in I}|\widehat{\gamma}_n(x;\mathbf{Z}_i)-\gamma_0(x;\mathbf{Z}_i)|^2=o_{\mathbbm{P}}(1)$. Now $\sup_{x\in I}|\widehat{\Phi}_n(x)-\Phi_0(x)|=o_{\mathbbm{P}}(1)$ follows by the triangle inequality.
For $|v|\leq V$
Using the argument in Section (ref), we can bound $\sup_{\vert v\vert\leq \vert V\vert } \vert \widehat{\Gamma}_n(\mathsf{x}+v)-\widehat{\Gamma}_n(\mathsf{x}) - \bar{\Gamma}_n(\mathsf{x}+v) + \bar{\Gamma}_n(\mathsf{x})\vert$. Then, $\sqrt{n} \vert \widehat{\Gamma}_n(\mathsf{x})-\bar{\Gamma}_n(\mathsf{x})\vert=O_{\mathbbm{P}}(1)$ implies
uniformly over $V\in (0,2\delta]$. Theorem 1 of Lo-Singh_1986_PTRF implies $\sqrt{n a_n}\sup_{x\in I}\vert \check{\Phi}_n(x ) - \check{\Phi}_n(\mathsf{x} ) - \bar{\Phi}_n(x)+\bar{\Phi}_n(\mathsf{x} ) \vert =o_{\mathbbm{P}}(1)$.
The conditions on the uniform covering number hold because $\gamma_0$ and $\widehat{\gamma}_n$ are bounded (for $\widehat{\gamma}_n$, with probability approaching one) and thus $|\phi_0(x_1;\mathbf{Z})-\phi_0(x_2;\mathbf{Z})|\leq C |x_1-x_2|$ and $|\widehat{\phi}_n(x_1;\mathbf{Z})-\widehat{\phi}_n(x_2;\mathbf{Z})|\leq C |x_1-x_2|$ with probability approaching one. By this Lipschitz property, the condition on $\bar{D}_{\phi}^{\eta}(\mathbf{Z})$ also holds.
(ref): Let $\psi_{\mathsf{x}}^{\mathtt{MD}}(v;\mathbf{Z})=\gamma_0(\mathsf{x} +v;\mathbf{Z})-\gamma_0(\mathsf{x};\mathbf{Z}) - \theta_0(\mathsf{x})v$ be the $\psi_{\mathsf{x}}$ function for the monotone density. Then, for $x$ sufficiently close to $\mathsf{x}$ and $|v|$ small enough,
Then, the same argument as in the monotone density case implies the desired result. \newline (ref) follows from consistency of $\widehat{\Phi}_n$ and $\widehat{\Phi}_n^*$.
(ref): $\widehat{\Phi}_n(x)=\int_0^x \widehat{F}_n(u)du,\widehat{\Phi}_n^*(x)=\int_0^x 1-\widehat{\Gamma}_n^*(u)du$ are non-negative since $\widehat{F}_n\geq 0$ and $1-\widehat{\Gamma}_n^*\geq 0$ with probability approaching one. This property also implies the non-decreasing property as $\widehat{\Phi}_n,\widehat{\Phi}_n^*$ are integrals. The continuity property also follows from the integral representation. By definition, $\widehat{\Phi}_n(0)=0=\widehat{\Phi}_n^*(0)$ and $\widehat{\Phi}_n(u_0^{\mathtt{MD}})=\widehat{u}_n=\widehat{u}_n^*=\widehat{\Phi}_n^*(u_0^{\mathtt{MD}})$ with $I=[0,u_0^{\mathtt{MD}}]$. The closedness of the range follows from continuity and $I$ being a compact interval. \newline (ref) follows from continuity of $\widehat{\Phi}_n$ and $\widehat{\Phi}_n^*$.
We consider the problem of estimating the cdf of $X$ at $\mathsf{x}$, $\theta_0(\mathsf{x})=F_0(\mathsf{x})$. Observations $\mathbf{Z}_1,\dots,\mathbf{Z}_n$ come from a random sample of $\mathbf{Z}=(\Delta,C,\mathbf{A}')'$ where $\Delta=\mathbbm{1}(X\leq C)$, $C$ is a random censoring time, and $\mathbf{A}$ is a vector of covariates. In this example, we do not observe $\check{X}=X\land C$. Instead, we observe the censoring time and whether the observation was censored. This setup is often referred to as current status data. Let $H_0(x)=\mathbbm{P}[C\leq x]$ be the cdf of $C$. We can use $\Gamma_0(x) = \int_0^x F_0(u) H_0(du)$ and $\Phi_0(x)=H_0(x)$. The interval $I$ is the support of $X$ and $u_0=1$. We also assume $H_0$ admits a Lebesgue density $h_0$. The structure of the estimation problem turns out to be identical to the one for the monotone regression example, and we can leverage the common structure.
First we consider the case of completely at random censoring $X\protect\mathpalette{\protect\independenT}{\perp} C$. See Groeneboom-Wellner_1992_Book for existing analysis. In this exaple, we do not use covariates $\mathbf{A}$. We set $\gamma_0(x;\mathbf{Z})=\Delta \mathbbm{1}(C\leq x)$ and $\phi_0(x;\mathbf{Z})=\mathbbm{1}(C\leq x)$. Note that if the notation is mapped by $(\Delta,C)\leftrightarrow (Y,X)$, then these functions are identical to those of the classical monotone regression problem (Corollary (ref)). Thus, the following result is identical to Corollary (ref), up to notation and some changes due to boundedness of $\Delta$.
We consider the case where right-censoring is conditionally independent i.e.,\ $X\protect\mathpalette{\protect\independenT}{\perp} C|\mathbf{A}$. vanderVaart-vanderLaan_2006_IJB analyzed this example as well as settings with time-varying covariates. We are focusing on time-invariant covariates. Define $F_0(C,\mathbf{A})=\mathbbm{E}[\Delta|C,\mathbf{A}]$ and $g_0(C,\mathbf{A})=\frac{h_{C|\mathbf{A}}(C|\mathbf{A})}{h_0(C)}$ where $h_{C|\mathbf{A}}$ is a conditional Lebesgue density of $C$ given $\mathbf{A}$ and $h_0$ is a Lebesgue density of $C$. Let $\widehat{F}_n(c,\mathbf{a})$ and $\widehat{g}_n(c,\mathbf{a})$ be preliminary estimators for $F_0(c,\mathbf{a})$ and $g_0(c,\mathbf{a})$, respectively.
Identical to the censoring completely at random case, with appropriate changes in the notation (i.e.,\ $(\Delta, C)\leftrightarrow (Y,X)$), the setup is equivalent to that of the monotone regression with covariates.
\paragraph*{Assumption SA.\arabic{section}.\arabic{subsection}} Let $\varepsilon=\Delta-\mathbbm{E}[\Delta|C,\mathbf{A}]$, $\sigma_0^2(C,\mathbf{A})=\mathbbm{E}[\varepsilon^2|C,\mathbf{A}]$, and $\delta >0$ be some fixed number.
Note that $|\varepsilon|\leq 1$.
As noted above, by mapping the notation $(\Delta, C)\leftrightarrow (Y,X)$, the arguments in Sections (ref) and (ref) directly apply to the current status estimators.
Here we develop a rule-of-thumb procedure to choose a step size for the bias-reduced numerical derivative estimator in the context of isotonic regression without covariates. Specifically, we consider the numerical derivative estimator
with $\underline{\mathfrak{s}}=3$, $c_1=1,c_2=-1,c_3=2,c_4=-2$. Then,
We use the (asymptotic) MSE-optimal step size discussed in the main paper. See also (ref). Yet, with the choice of $c_k$'s, part of the bias constant $\sum_{k=1}^{\underline{\mathfrak{s}}+1}\lambda_j^{\mathtt{BR}}(k)c_k^{\underline{s}+2}$ equals zero, and we need to turn to the next leading term of the bias, which is
Then, letting $\mathsf{B}_j^{\mathtt{BR}}(\mathsf{x}) = \frac{\partial^{6}\Upsilon_0(\mathsf{x})}{6!} \sum_{k=1}^{4}\lambda_j^{\mathtt{BR}}(k)c_k^{6}$, the MSE-optimal step size is
The bias and variance constants depend on unknown features of the data generating process. Specifically, $\mathsf{B}_j^{\mathtt{BR}}(\mathsf{x})$ depends on the regression function $\theta_0$, the Lebesgue density of $X$, and their derivatives at $X=\mathsf{x}$ while $\mathsf{V}_j^{\mathtt{BR}}(\mathsf{x})$ is determined by the density of $X$ and the conditional variance of the regression error $\varepsilon=Y-\theta_0(X)$ at $X=\mathsf{x}$. To implement the construction of the step size, we posit a simple parametric model:
where $\{\gamma_0,\gamma_1,\gamma_2,\gamma_4,\gamma_4,\gamma_5,\mu,\sigma\}$ are parameters to be estimated. Once we estimate the parameters of this reference model, we can construct a rule-of-thumb step size $\epsilon_{j,n}^{\mathtt{ROT}}$ by replacing $\mathsf{B}_j^{\mathtt{BR}}(\mathsf{x})$ and $\mathsf{V}_j^{\mathtt{BR}}(\mathsf{x})$ with their estimates. Note that although the bias and variance constant estimators may not be consistent for the true $\mathsf{B}_j^{\mathtt{BR}}(\mathsf{x})$ and $\mathsf{V}_j^{\mathtt{BR}}(\mathsf{x})$, the rate of $\epsilon_{j,n}^{\mathtt{ROT}}$ is MSE-optimal, and the numerical derivative estimator converges to $\mathcal{D}_j(\mathsf{x})$ sufficiently fast to satisfy Equation (ref) in the main paper.
\makeatletter\@input{xx.tex}\makeatother