Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Asymptotic Theory of $L$Statistics and Integrable Empirical Processes
frontmatter\runtitle{$L$\hypStatistics and Integrable Processes}
\thankstext{T1}{The previous version was circulated with the title “Switching to the New Norm: From Heuristics to Formal Tests using Integrable Empirical Processes.”}
\begin{aug}
\runauthor{T. Kaji}
\address{The University of Chicago\\Booth School of Business\\
5807 South Woodlawn Avenue\\Chicago, IL 60637\\
\printead{e1}\\
\phantom{E-mail:\ }}
\end{aug}
\begin{abstract}
This paper develops asymptotic theory of integrals of empirical quantile functions with respect to random weight functions, which is an extension of classical $L$\hypstatistics.
They appear when sample trimming or Winsorization is applied to asymptotically linear estimators.
The key idea is to consider empirical processes in the spaces appropriate for integration.
First, we characterize weak convergence of empirical distribution functions and random weight functions in the space of {\em bounded integrable} functions.
Second, we establish the delta method for empirical quantile functions as {\em integrable} functions.
Third, we derive the delta method for $L$\hypstatistics.
Finally, we prove weak convergence of their bootstrap processes, showing validity of nonparametric bootstrap.
\end{abstract}
\begin{keyword}[class=MSC]
\kwd{62E20}
\kwd{62G30}
\end{keyword}
\begin{keyword}
\kwd{$L$\hypstatistics}
\kwd{empirical processes}
\kwd{quantile processes}
\kwd{functional delta methods}
\kwd{nonparametric bootstrap}
\end{keyword}
Introduction
We derive the asymptotic distribution of the statistics of the form
\[
\int_0^1m(\mathbb{Q}_n)d\mathbb{K}_n,
\]
where $m:\mathbb{R}\to\mathbb{R}$ is a known continuously differentiable function, $\mathbb{Q}_n:(0,1)\to\mathbb{R}$ an empirical quantile function of a random variable $X_i$, and $\mathbb{K}_n:(0,1)\to\mathbb{R}$ a random Lipschitz function that depends on $\{X_i\}$.
This is a generalization of the classical $L$\hypstatistics ms1992,s1997,s2017,sw1986 to allow for integration with respect to random processes $\mathbb{K}_n$.
\footnote{s1997 allows integration on a random interval but not with respect to a random process.}
This type of statistics appears, for example, when sample trimming or Winsorization is applied to asymptotically linear estimators.
Let us collectively call sample trimming and Winsorization {\em sample adjustments}.
If sample adjustments are made conditional on the values of $X_i$, $\mathbb{K}_n$ is a nonrandom function and it falls within the framework of classical $L$\hypstatistics.
If sample adjustments are made on variables other than $X_i$, $\mathbb{K}_n$ becomes random and it affects the asymptotic distribution of the $L$\hypstatistics.
In economics, this occurs as the parameters of interest (what $L$\hypstatistics estimate) often differ from the variables whose outliers we are concerned.
In such cases, dependence of $\mathbb{K}_n$ can be difficult to handle directly.
This paper gives both high\hyplevel and low\hyplevel conditions for weak convergence of the $L$\hypstatistics, derives the asymptotic distribution formula, and verifies validity of nonparametric bootstrap.
The innovation of this paper lies in considering empirical processes in the space of integrable functions.
The literature on empirical processes has largely focused on uniform convergence irrespective of the intended statistical application.
As $L$\hypstatistics are integrals of empirical processes, we (partly) renounce uniform convergence and instead require integrability, which buys us substantial benefits in dealing with $L$\hypstatistics.
\footnote{In applying the empirical process theory to $L$\hypstatistics, Van der Vaart v1998 states that “[this approach] is preferable in that it applies to more general statistics, but it{\ldots}does not cover the simplest $L$\hypstatistic: the sample mean.” Our empirical process theory overcomes this problem.}
Our theoretical development is summarized as follows.
By integration by parts, we expect
multline*[multline* omitted — 292 chars of source]
First, we consider $\sqrt{n}(\mathbb{F}_n-F)$ and $\sqrt{n}(\mathbb{K}_n-K)$ as elements in the space of bounded integrable functions with respect to appropriate measures and derive conditions for weak convergence therein ((ref)).
Second, we establish the functional delta method for the “inverse map,” $F\mapsto m(F^{-1})=m(Q)$, from the space of bounded integrable functions to the space of integrable functions, which shows weak convergence of $\sqrt{n}[m(\mathbb{Q}_n)-m(Q)]$ as an integrable process ((ref)).
\footnote{This paper is presumably the first to show weak convergence of (possibly unbounded) empirical quantile processes in $L_1$ on the untruncated domain $(0,1)$.}
Third, we develop the functional delta method for the map, $(Q,K)\mapsto\int m(Q)dK$, from the spaces of integrable and bounded integrable functions to a Euclidean space, establishing weak convergence of $L$\hypstatistics ((ref)).
Finally, we develop conditions for nonparametric bootstrap for the processes and $L$\hypstatistics ((ref)).
The theory of this paper was originally motivated by the following problem of formalizing outlier robustness analyses in economics.
exa[Outlier Robustness Analysis]
Applied researchers often want to examine whether a small portion of outliers affect the regression outcomes ajkkm2016,ajr2001,anrr2016,abb2010,frw2007.
The common heuristic practice in economics is to compare two estimators $\hat{\beta}_1$ and $\hat{\beta}_2$, where $\hat{\beta}_1$ is estimated with the full sample and $\hat{\beta}_2$ with the sample that excludes outliers, against the standard error of $\hat{\beta}_1$.
However, since $\hat{\beta}_1$ and $\hat{\beta}_2$ share largely overlapping samples, their difference tends to be small simply because of their strong positive correlation.
To account for this, it is more appropriate to compare the difference $\hat{\beta}_1-\hat{\beta}_2$ to its own variance, as opposed to the marginal variance of $\hat{\beta}_1$.
This calls for the joint distribution of $\hat{\beta}_1$ and $\hat{\beta}_2$.
Consider linear regression $y_i=x_i\beta+\varepsilon_i$ with $\mathbb{E}[x_i\varepsilon_i]=0$.
The ordinary least squares (OLS) estimator of $\beta$ is $\hat{\beta}_1=\bigl(\frac{1}{n}\sum_{i=1}^n x_i^2\bigr)^{-1}\frac{1}{n}\sum_{i=1}^n x_iy_i$, so its asymptotic distribution depends on that of the average of $x_iy_i$.
However, $x_iy_i$ is usually not the quantity whose outliers are of natural concern, but rather, $x_i$ ajr2001, $y_i$ ajkkm2016,bdh2014, or $\hat{\varepsilon}_i$ anrr2016 is.
Then, conditional on the value of $x_iy_i$, the probability that the observation is deemed as an outlier is probabilistic.
Suppose we remove the 2% tail observations of $x_i$ and $y_i$.
Let $w_i$ be $1$ if $x_{(\lceil0.02n\rceil)}\leq x_i\leq x_{(\lceil0.98n\rceil)}$ and $y_{(\lceil0.02n\rceil)}\leq y_i\leq y_{(\lceil0.98n\rceil)}$, and $0$ otherwise.
\footnote{Winsorization can also be accommodated by appropriately defining $w_i$.}
The outlier\hypremoved estimator $\hat{\beta}_2$ is $\bigl(\frac{1}{n}\sum_{i=1}^n x_i^2w_i\bigr)^{-1}\frac{1}{n}\sum_{i=1}^n x_iy_iw_i$.
Through the quantile transform, we can write
\[
\hat{\beta}_1=\biggl(\frac{1}{n}\sum_{i=1}^nx_i^2\biggr)^{-1}\int_0^1\mathbb{Q}_n(u)du, \quad
\hat{\beta}_2=\biggl(\frac{1}{n}\sum_{i=1}^nx_i^2w_i\biggr)^{-1}\int_0^1\mathbb{Q}_n(u)d\mathbb{K}_n(u),
\]
where $\mathbb{Q}_n$ is the empirical quantile function of $x_iy_i$ and $\mathbb{K}_n$ a random weight function whose derivative is $w_i$ for $u\in(\mathbb{F}_n(x_iy_i)-1/n,\mathbb{F}_n(x_iy_i)]$ for the empirical distribution function $\mathbb{F}_n$ of $x_iy_i$.
Then $\mathbb{K}_n$ is random for each fixed value of $x_iy_i$, which affects the asymptotic distribution of the integrals.
In (ref), we revisit the outlier robustness analysis in anrr2016.
The rest of the paper is organized as follows.
(ref) defines the setup.
(ref) develops the theory of weak convergence of bounded integrable processes.
(ref) establishes Hadamard differentiability of the inverse map.
(ref) shows Hadamard differentiability of the $L$\hypstatistics.
(ref) verifies validity of nonparametric bootstrap.
(ref) contains proofs.
(ref) contains supporting lemmas and an empirical application.
The Setup
Let $X_i$ be i.i.d.\ scalar random variables and $w_i$ be possibly random weights whose distribution is bounded but can depend on all of $\{X_i\}$. Consider a statistic of the form
\(
\hat{\beta}\vcentcolon=\frac{1}{n}\sum_{i=1}^n m(X_i)w_i=\frac{1}{n}\sum_{i=1}^n m(X_{(i)})w_{(i)},
\)
where $m$ is a continuously differentiable function, $X_{(i)}$ is an order statistic such that $X_{(1)}\leq X_{(2)}\leq\cdots\leq X_{(n)}$, and $w_{(i)}$ is ordered according to the order of $X_i$.
Let $\mathbb{Q}_n(u)\vcentcolon=X_{(i)}$ and $d\mathbb{K}_n(u)\vcentcolon=w_{(i)}$, $u\in(\frac{i-1}{n},\frac{i}{n}]$, be the empirical quantile function of $X_i$ and the random weight function.
With these,
\[
\hat{\beta}=\int_0^1 m(\mathbb{Q}_n(u))d\mathbb{K}_n(u).
\]
Denote by $\mathbb{F}_n(x)\vcentcolon=\frac{1}{n}\sum_{i=1}^n\mathbbm{1}\{X_i\leq x\}$ the empirical distribution function of $X_i$ and define the inverse of a nondecreasing function $f:\mathbb{R}\to\mathbb{R}$ by $f^{-1}(y)\vcentcolon=\inf\{x\in\mathbb{R}:f(x)\geq y\}$.
Then, $\mathbb{Q}_n$ equals $\mathbb{F}_n^{-1}$.
The aim of this paper is to derive the joint distribution of finitely many such quantities $(\hat{\beta}_1,\dots,\hat{\beta}_d)$ for possibly different $m$, $\{X_i\}$, and $\{w_i\}$.
For this, we proceed in four steps:
enumerate[i.]
• Give conditions for convergence of $\sqrt{n}(\mathbb{F}_n-F)$ and $\sqrt{n}(\mathbb{K}_n-K)$ to Gaussian processes as bounded integrable processes.
• Show convergence of $\sqrt{n}(\mathbb{Q}_n-Q)$ to a Gaussian process as an integrable process via a functional delta method from $\mathbb{F}_n$ to $\mathbb{Q}_n$.
• Show convergence of $L$\hypstatistics via a functional delta method from $(\mathbb{Q}_n,\mathbb{K}_n)$ to $\int m(\mathbb{Q}_n)d\mathbb{K}_n$.
• Show bootstrap convergence for $\sqrt{n}(\mathbb{F}_n-F)$ and $\sqrt{n}(\mathbb{K}_n-K)$.
Convergence of Bounded Integrable Processes
Define the space of bounded integrable functions as follows.
defnLet $(T, \mathcal{T}, \mu)$ be a measure space where $T$ is an arbitrary set, $\mathcal{T}$ a $\sigma$\hypfield on $T$, and $\mu$ a $\sigma$\hypfinite signed measure on $\mathcal{T}$.
Let $\mathbb{L}_\mu$ be the space of bounded and $\mu$\hypintegrable functions $z : T \to \mathbb{R}$ with the norm
\[
\|z\|_{\mathbb{L}_\mu} \vcentcolon= \|z\|_T \vee \|z\|_\mu \vcentcolon= \biggl( \sup_{t\in T} |z(t)| \biggr) \vee \biggl( \int_T |z| |d\mu| \biggr),
\]
where $|d\mu|$ represents integration with respect to the total variation measure.
For sums of i.i.d.\ random variables such as $\sqrt{n}(\mathbb{F}_n-F)$, it is straightforward to prove weak convergence in $\mathbb{L}_\mu$ by the combination of classical central limit theorems (CLTs) vw1996.
propLet $(\mathbb{R},\mathfrak{B}(\mathbb{R}),\mu)$ be a $\sigma$-finite Borel measure on $\mathbb{R}$.
For a probability distribution $F$ on $\mathbb{R}$ such that $\int_{\mathbb{R}}\sqrt{F(1-F)}|d\mu|<\infty$, the empirical process $\sqrt{n}(\mathbb{F}_n-F)$ converges weakly in $\mathbb{L}_\mu$ to a Gaussian process with mean zero and covariance function $\operatorname{Cov}(x,y)=F(x\wedge y)-F(x)F(y)$.
remFor an increasing function $m$, $\int_{-\infty}^\infty\sqrt{F(1-F)}dm<\infty$ is equivalent to $\|m(X)\|_{2,1}\vcentcolon=\int_0^\infty\sqrt{\Pr(|m(X)|>t)}dt<\infty$ bgm1999.
Moreover, if $m(X)$ has a $(2+c)$th moment for some $c>0$, we have $\|m(X)\|_{2,1}<\infty$.
For processes not given as sums of i.i.d.\ variables such as $\sqrt{n}(\mathbb{K}_n-K)$, we need direct conditions for weak convergence.
As in classical literature, we characterize weak convergence in $\mathbb{L}_\mu$ by asymptotic tightness plus marginal convergence.
Following vw1996, we consider a {\em net} $X_\alpha$ indexed by an arbitrary directed set, rather than a sequence $X_n$ indexed by natural numbers.
We also allow the sample space to be different for each element in a net, $X_\alpha:\Omega_\alpha\to\mathbb{L}_\mu$.
Finally, we allow each element in the net to be not necessarily measurable.
When we write $X(t)$ for a map $X : \Omega \to \mathbb{L}_\mu$, $t$ is understood to be an element of $T$ and we regard $X(t)$ as a map from $\Omega$ to $\mathbb{R}$ indexed by $T$; when we explicitly use $\omega \in \Omega$ in the discussion, we write $X(t,\omega)$.
thmLet $X_\alpha : \Omega_\alpha \to \mathbb{L}_\mu$ be arbitrary. Then, $X_\alpha$ converges weakly to a tight limit if and only if $X_\alpha$ is asymptotically tight and marginals $(X_\alpha(t_1), \dots, X_\alpha(t_k))$ converge weakly for every finite subset $t_1, \dots, t_k$ of $T$.
If $X_\alpha$ is asymptotically tight and its marginals converge weakly to the marginals $(X(t_1), \dots, X(t_k))$ of a stochastic process $X$, then there is a version of $X$ with sample paths in $\mathbb{L}_\mu$ and $X_\alpha \leadsto X$.
Weak convergence of marginals can be established by classical results such as CLTs in Euclidean spaces.
The question is asymptotic tightness. We characterize this with uniform equicontinuity and equiintegrability.
defnFor a $\mu$\hypmeasurable semimetric $\rho$ on $T$,
\footnote{We call a semimetric {\em $\mu$\hypmeasurable} if every open set induced is measurable with respect to $\mu$.}
the net $X_\alpha : \Omega_\alpha \to \mathbb{L}_\mu$ is {\em asymptotically uniformly $\rho$\hypequicontinuous and $(\rho,\mu)$\hypequiintegrable in probability} if for every $\varepsilon,\eta>0$ there exists $\delta>0$ such that
\begin{multline*}
\limsup_\alpha P^\ast \biggl( \sup_{t\in T} \biggl[ \biggl( \sup_{\rho(s,t)<\delta} |X_\alpha(s)-X_\alpha(t)| \biggr) \\
\vee \biggl( \int_{0<\rho(s,t)<\delta} |X_\alpha(s)| |d\mu(s)| \biggr) \biggr] > \varepsilon \biggr) < \eta.
\end{multline*}
The following result characterizes asymptotic tightness in $\mathbb{L}_\mu$.
thmThe following are equivalent.
\begin{enumerate}[i.]
• A net $X_\alpha : \Omega_\alpha \to \mathbb{L}_\mu$ is asymptotically tight.
• $X_\alpha(t)$ is asymptotically tight in $\mathbb{R}$ for every $t \in T$, $\|X_\alpha\|_\mu$ is asymptotically tight in $\mathbb{R}$, and for every $\varepsilon,\eta>0$ there exists a finite $\mu$\hypmeasurable partition $T=\bigcup_{i=1}^k T_i$ such that
\begin{multline}
\limsup_\alpha P^\ast \Biggl( \biggl[ \sup_{1\leq i\leq k} \sup_{s,t\in T_i} |X_\alpha(s)-X_\alpha(t)| \biggr] \\
\vee \sum_{i=1}^k \inf_{x\in\mathbb{R}} \int_{T_i} |X_\alpha-x| |d\mu| > \varepsilon \Biggr) < \eta.
\end{multline}
• $X_\alpha(t)$ is asymptotically tight in $\mathbb{R}$ for every $t \in T$ and there exists a $\mu$\hypmeasurable semimetric $\rho$ on $T$ such that $(T,\rho)$ is totally bounded and $X_\alpha$ is asymptotically uniformly $\rho$\hypequicontinuous and $(\rho,\mu)$\hypequiintegrable in probability.
\end{enumerate}
remThe condition “$0 < \rho(s,t)$” allows for the point masses in $\mu$ and plateaus in $X_\alpha$.
In ((ref)), this corresponds to “$-x$.”
Now we turn to conditions for $\sqrt{n}(\mathbb{K}_n-K)$.
The following is a special case of $\mathbb{L}_\mu$ suitable for $\mathbb{K}_n$.
defnLet $Q:(0,1)\to\mathbb{R}$ be an integrable increasing function and let $\mathbb{L}_Q$ be the space of functions $\kappa:(0,1)\to\mathbb{R}$ with the norm
\[
\|\kappa\|_{\mathbb{L}_Q}\vcentcolon=\|\kappa\|_{Q,\infty}\vee\|\kappa\|_Q\vcentcolon=\biggl(\sup_{u\in(0,1)}|(|Q|\vee1)(u)\kappa(u)|\biggr)\vee\biggl(\int_0^1|\kappa|dQ\biggr).
\]
Let $\mathbb{L}_{Q,M}\subset\mathbb{L}_Q$ be the subset of Lipschitz functions with Lipschitz constants bounded by $M$.
The following lemma gives a low\hyplevel condition for $\sqrt{n}(\mathbb{K}_n-K)$ to converge in $\mathbb{L}_Q$.
Roughly, if $|Q|^{2+c}$ is integrable, then $\frac{X_\alpha}{u^r(1-u)^r}\leadsto\frac{X}{u^r(1-u)^r}$ in the uniform norm for some $r>\frac{1}{2+c}$ implies $X_\alpha\leadsto X$ in $\mathbb{L}_Q$.
lemLet $Q : (0,1) \to \mathbb{R}$ be an increasing function in $L_{2+c}$ for some $c>0$.
If for a net of processes $X_\alpha : \Omega_\alpha \to \mathbb{L}_Q$ there exists $r>\frac{1}{2+c}$ such that for every $\eta>0$ there exists $M$ satisfying
\[
\limsup_\alpha P^\ast \biggl( \biggl\| \frac{X_\alpha}{u^r(1-u)^r} \biggr\|_\infty > M \biggr) < \eta,
\]
then there exists a semimetric $\rho$ on $(0,1)$ such that $(0,1)$ is totally bounded, $X_\alpha$ is asymptotically uniformly $\rho$\hypequicontinuous in probability, and $X_\alpha$ is asymptotically $(\rho,Q)$\hypequiintegrable in probability.
This implies that sample adjustments based on fixed quantiles satisfy the condition.
For example, let $X_1,\dots,X_n$ be i.i.d.\ continuous random variables and $X_{1,n},\dots,X_{m,n}$ be their subset selected by some (possibly random) criterion.
Then, if the empirical process of the subset converges weakly uniformly to a smooth distribution, then $\sqrt{n}(\mathbb{K}_n-K)$ converges weakly in $\mathbb{L}_Q$.
propLet $U_1, \dots, U_n$ be independent uniformly distributed random variables on $(0,1)$ and $w_{1,n}, \dots, w_{n,n}$ random variables bounded by $M$ whose distribution can depend on $U_1, \dots, U_n$ and $n$.
Define
\(
\mathbb{F}_n(u) \vcentcolon= \frac{1}{n} \sum_{i=1}^n \mathbbm{1}\{U_i \leq u\}$ and $
\mathbb{G}_n(u) \vcentcolon= \frac{1}{n} \sum_{i=1}^n w_{i,n} \mathbbm{1}\{U_i \leq u\}.
\)
Let $I(u)\vcentcolon=u$ and assume that $K(u)\vcentcolon=\lim_{n\to\infty} \mathbb{E}[\mathbb{G}_n(u)]$ exists and is Lipschitz differentiable.
If $\sqrt{n}(\mathbb{F}_n-I)$ and $\sqrt{n}(\mathbb{G}_n-K)$ converges weakly jointly in $L_\infty$, then for
\[
\mathbb{K}_n(u) \vcentcolon= \frac{1}{n} \sum_{i=1}^n w_{i,n} \mathbbm{1} \bigl\{ 0 \vee \bigl( nu - n\mathbb{F}_n(U_i) + 1 \bigr) \wedge 1 \bigr\},
\]
we have $\sqrt{n}(\mathbb{K}_n-K)$ converge weakly in $\mathbb{L}_Q$ for every increasing function $Q\in L_{2+c}$ for every $c>0$.
Convergence of Quantile Processes as Integrable Processes
For a smooth function $m$ for which $m(X)$ has sufficient moments, we establish weak convergence of $\sqrt{n}(m(\mathbb{Q}_n)-m(Q))$ to a Gaussian process.
If $m$ is identity, the (unweighted) empirical quantile process converges weakly in $L_1$ on the entire domain $(0,1)$, without truncating the tails, even if $Q$ is an unbounded function.
Interestingly, this point has been overlooked in the literature, which mostly concerned uniform convergence of either bounded or weighted quantile processes k2002,ch1993,chs1993,cchm1986,sw1986.
In particular, we show differentiability of the inverse map as a functional from $\mathbb{L}_\mu$ to $L_1$.
Note that $\mathbb{E}[m(X)]=\int mdF=-\int Fdm$ in terms of $F$ and $\mathbb{E}[m(X)]=\int m(Q)du$ in terms of $Q$.
Therefore, the appropriate space for $F$ is the following special case of $\mathbb{L}_\mu$ while the space for $Q$ is a standard $L_1$.
defnLet $m:\mathbb{R}\to\mathbb{R}$ be a nondecreasing continuously differentiable function.
Let $\mathbb{L}_m$ be the space of Borel-measurable functions $z:\mathbb{R}\to\mathbb{R}$ with limits $z(\pm\infty) \vcentcolon= \lim_{x\to\pm\infty} z(x)$ and the norm
\[
\|z\|_{\mathbb{L}_m} \vcentcolon= \|z\|_\infty \vee \|z\|_m \vcentcolon= \biggl( \sup_{x\in\mathbb{R}} |z(x)| \biggr) \vee \biggl( \int_{-\infty}^\infty |\tilde{z}| dm \biggr)
\]
where $\tilde{z}(x)\vcentcolon=z(x)-z(-\infty)\mathbbm{1}\{x<0\}-z(+\infty)\mathbbm{1}\{x\geq0\}$.
Denote by $\mathbb{L}_{m,\phi}$ the subset of $\mathbb{L}_m$ of monotone cadlag functions with $z(-\infty)=0$ and $z(+\infty)=1$.
defnLet $\mathbb{B}$ be the space of ladcag functions $z:(0,1)\to\mathbb{R}$ with the norm
\(
\|z\|_{\mathbb{B}} \vcentcolon= \int_0^1 |z(u)| du.
\)
thm[Inverse map]
Let $m : \mathbb{R} \to \mathbb{R}$ be a continuously differentiable function and $F \in \mathbb{L}_{m,\phi}$ a distribution function on (an interval of) $\mathbb{R}$ that has at most finitely many jumps and is otherwise continuously differentiable with strictly positive density $f$.
Then, the map $\phi \circ \psi : \mathbb{L}_{m,\phi} \to \mathbb{B}$, $\phi \circ \psi(F) \vcentcolon= m(Q)$, is Hadamard differentiable at $F$ tangentially to the set $\mathbb{L}_{m,0}$ of all continuous functions in $\mathbb{L}_m$. The derivative is given by
\(
(\phi\circ\psi)_F'(z)\vcentcolon=-(m' z/f)\circ Q.
\)
The main conclusion of this section is summarized as follows.
propLet $m : \mathbb{R} \to \mathbb{R}$ be a continuously differentiable function.
For a distribution function $F$ on (an interval of) $\mathbb{R}$ that has at most finitely many jumps and is otherwise continuously differentiable with strictly positive density $f$ such that $\int_{\mathbb{R}}\sqrt{F(1-F)}|dm|<\infty$, the process $\sqrt{n}(m(\mathbb{Q}_n)-m(Q))$ converges weakly in $\mathbb{B}$ to a Gaussian process with mean zero and covariance $\operatorname{Cov}(s,t) = m'(Q(s))Q'(s)m'(Q(t))Q'(t) (s\wedge t-st)$.
Convergence of $L$\hypstatistics
We seek conditions under which the integral of a stochastic process with respect to another stochastic process converges weakly.
This is an extension of Wilcoxon statistics vw1996 that allows unbounded integrands.
thm[Wilcoxon statistic]
For each fixed $M$, the maps $\lambda:\mathbb{B}\times\mathbb{L}_{Q,M}\to\mathbb{R}$ and $\tilde{\lambda}:\mathbb{B}\times\mathbb{L}_{Q,M}\to L_\infty(0,1)^2$,
\(
\lambda(Q,K)\vcentcolon=\int_0^1QdK $ and $
\tilde{\lambda}(Q,K)(s,t)\vcentcolon=\int_s^tQdK,
\)
are Hadamard differentiable at every $(Q,K)\in\mathbb{B}\times\mathbb{L}_{Q,M}$ uniformly over $\mathbb{L}_{Q,M}$.
The derivative maps are
\(
\lambda_{Q,K}'(z,\kappa)\vcentcolon=\int_0^1Qd\kappa+\int_0^1zdK $ and $
\tilde{\lambda}_{Q,K}'(z,\kappa)(s,t)\vcentcolon=\int_s^tQd\kappa+\int_s^tzdK,
\)
where $\int Qd\kappa$ is defined via integration by parts if $\kappa$ is of unbounded variation.
Now we are ready to give the main conclusion of this paper.
prop[$L$\hypstatistic]
Let $m_1, m_2 : \mathbb{R} \to \mathbb{R}$ be continuously differentiable functions and $F : \mathbb{R}^2 \to [0,1]$ be a distribution function on (a rectangular of) $\mathbb{R}^2$ with marginal distributions $(F_1, F_2)$ that have at most finitely many jumps and are otherwise continuously differentiable with strictly positive marginal densities $(f_1, f_2)$ such that $m_1(X_1)$ and $m_2(X_2)$, $(X_1,X_2) \sim F$, have $(2+c)$th moments for some $c>0$.
Along with i.i.d.\ random variables $X_{1,1}, \dots, X_{n,1}$ and $X_{1,2}, \dots, X_{n,2}$, let $w_{1,n,1}, \dots, w_{n,n,1}$ and $w_{1,n,2}, \dots, w_{n,n,2}$ be random variables bounded by $M$ whose distribution can depend on $n$, $X_{1,1}, \dots, X_{n,1}$, and $X_{1,2}, \dots, X_{n,2}$ such that the empirical distributions of $X_{i,1}$, $X_{i,2}$, $w_{i,n,1} X_{i,1}$, and $w_{i,n,2} X_{i,2}$ converge uniformly jointly to continuously differentiable functions.
Then,
\begin{multline*}
\sqrt{n} \begin{pmatrix} \mathbb{E}_n[m_1(X_{i,1}) w_{i,n,1}] - \mathbb{E}[m_1(X_{i,1}) w_{i,n,1}] \\ \mathbb{E}_n[m_2(X_{i,2}) w_{i,n,2}] - \mathbb{E}[m_2(X_{i,2}) w_{i,n,2}] \end{pmatrix} \\
= \sqrt{n} \begin{pmatrix} \int_0^1 m_1(\mathbb{Q}_{n,1}) d\mathbb{K}_{n,1} - \int_0^1 m_1(Q_1) dK_1 \\ \int_0^1 m_2(\mathbb{Q}_{n,2}) d\mathbb{K}_{n,2} - \int_0^1 m_2(Q_2) dK_2 \end{pmatrix}
\end{multline*}
where
\(
\mathbb{K}_{n,j}(u) \vcentcolon= \frac{1}{n} \sum_{i=1}^n w_{i,n,j} \mathbbm{1} \bigl\{ 0 \vee \bigl( nu - n\mathbb{F}_{n,j}(X_i) + 1 \bigr) \wedge 1 \bigr\}
\)
and
\(
K_j(u) \vcentcolon= \lim_{n\to\infty} \mathbb{E}[w_{i,n,j} \mid F_j(X_{i,j}) \leq u],
\)
converge weakly in $\mathbb{R}^2$ to a normal vector $(\xi_1,\xi_2)$ with mean zero and (co)variance
\begin{multline*}
\operatorname{Cov}(\xi_j, \xi_k) = \int_0^1 \int_0^1 m_j'(Q_j(s))Q_j'(s)m_k'(Q_k(t))Q_k'(t) \times \\
\Bigl( [F_{jk}^Q(s,t)-st] + [K_{jk}(s,t) F_{jk}^Q(s,t)-stK_j(s)K_k(t)] \\
- K_j(s) [F_{jk}^Q(s,t)-st] - K_k(t) [F_{jk}^Q(s,t)-st] \Bigr) ds dt,
\end{multline*}
where $F_{jk}^Q(s,t)\vcentcolon=\Pr(X_{i,j}\leq Q_j(s), X_{i,k}\leq Q_k(t))$ and $K_{jk}(s,t)\vcentcolon=\lim_{n\to\infty}$ $\mathbb{E}[w_{i,n,j}w_{i,n,k} \mid X_{i,j} \leq Q_j(s), X_{i,k} \leq Q_k(t)]$.
If $F$ has no jumps, this equals
\begin{multline*}
\operatorname{Cov}(\xi_j, \xi_k) = \int_{-\infty}^\infty \int_{-\infty}^\infty \Bigl( [1-K_j^F(x)-K_k^F(y)][F_{jk}(x,y)-F_j(x)F_k(y)] \\
+ [K_{jk}^F(x,y) F_{jk}(x,y)-K_j^F(x)K_k^F(y)F_j(x)F_k(y)] \Bigr) dm_j(x) dm_k(y),
\end{multline*}
where $F_{jk}(x,y)\vcentcolon=\Pr(X_{i,j}\leq x, X_{i,k}\leq y)$ and $K_{jk}^F(x,y)\vcentcolon=\lim_{n\to\infty}\mathbb{E}[w_{i,n,j}$ $w_{i,n,k} \mid X_{i,j} \leq x, X_{i,k} \leq y]$.
If $m_j$ and $m_k$ are known, this can be consistently estimated by its sample analogue
\begin{gather*}
\widehat{\operatorname{Cov}(\xi_j, \xi_k)} = \int_{-\infty}^\infty \int_{-\infty}^\infty \Bigl( [1-\mathbb{K}_{n,j}^F(x)-\mathbb{K}_{n,k}^F(y)][\mathbb{F}_{n,jk}(x,y)-\mathbb{F}_{n,j}(x)\mathbb{F}_{n,k}(y)] \\
+ [\mathbb{K}_{n,jk}^F(x,y) \mathbb{F}_{n,jk}(x,y)-\mathbb{K}_{n,j}^F(x)\mathbb{K}_{n,k}^F(y)\mathbb{F}_{n,j}(x)\mathbb{F}_{n,k}(y)] \Bigr) dm_j(x) dm_k(y),
\end{gather*}
where $\mathbb{F}_{n,jk}(x,y)\vcentcolon=\mathbb{E}_n[\mathbbm{1}\{X_{i,j}\leq x, X_{i,k}\leq y\}]$ and $\mathbb{K}_{n,jk}^F(x,y)\vcentcolon=\mathbb{E}_n[w_{i,n,j}$ $w_{i,n,k} \mid X_{i,j} \leq x, X_{i,k} \leq y]$.
Convergence of Bootstrap Processes
We establish validity of nonparametric bootstrap, viz., conditional weak convergence of the bootstrap processes. The {\em bootstrap process} for $\mathbb{F}_n$ is given by
align*[align* omitted — 260 chars of source]
where $M_{ni}$ is the number of times $X_i$ is drawn in the bootstrap sample.
We show that $\hat{\mathbb{Z}}_n$ converges weakly to the same limit as $\mathbb{Z}_n\vcentcolon=\sqrt{n}(\mathbb{F}_n-F)$ conditional on $\{X_i\}$.
As in vw1996, we proceed as follows: since $M_{ni}$ sums up to $n$, it is slightly dependent on each other; we replace $M_{ni}$ with {\em independent} Poisson random variables $\xi_i$ by showing equivalence of weak convergence of $\hat{\mathbb{Z}}_n$ and of the {\em multiplier process} $\mathbb{Z}_n'\vcentcolon=n^{-1/2}\sum\xi_i(\mathbbm{1}\{X_i\leq x\}-F)$ ((ref)); then, we prove unconditional convergence of $\mathbb{Z}_n'$ (randomness comes from both $X_i$ and $\xi_i$) by symmetrization ((ref)); finally, we show convergence of $\mathbb{Z}_n'$ conditional on $\mathbb{Z}_n$ (randomness only comes from $\xi_i$) by discretizing $\mathbb{Z}_n'$ ((ref)).
We observe that many proofs in vw1996 carry over to $\mathbb{L}_\mu$, so we will not reproduce the entire argument but prove steps that require modification.
In addition, we establish conditional weak convergence of the bootstrap process for $\mathbb{K}_n$. We restrict attention to sample adjustments by quantiles and write its bootstrap process in terms of empirical processes ((ref)).
The following shows conditional convergence of $\mathbb{Z}_n'$ as in vw1996.
Other lemmas are given in (ref).
lemLet $\xi_1,\dots,\xi_n$ be i.i.d.\ random variables with mean $0$, variance $1$, and $\|\xi\|_{2,1}<\infty$, independent of $X_1,\dots,X_n$.
For a probability distribution $F$ on $\mathbb{R}$ such that $\int_{\mathbb{R}}\sqrt{F(1-F)}|d\mu|<\infty$, the process $\mathbb{Z}_n'(x)=n^{-1/2}\sum_{i=1}^n\xi_i[\mathbbm{1}\{X_i\leq x\}-F(x)]$ satisfies
\(
\sup_{h\in\text{\rm BL}_1(\mathbb{L}_\mu)}\bigl|\mathbb{E}_\xi h(\mathbb{Z}_n')-\mathbb{E}h(\mathbb{Z})\bigr|\operatorname*{\mathchoice{
\,\longrightarrow\,}{
\rightarrow}{
\rightarrow}{
\rightarrow}
}0
\)
in outer probability, and the sequence $\mathbb{Z}_n'$ is asymptotically measurable.
These results show that nonparametric bootstrap works for $\sqrt{n}(\mathbb{F}_n-F)$ and $\sqrt{n}(\mathbb{Q}_n-Q)$.
We also show validity for $\sqrt{n}(\mathbb{K}_n-K)$ by representing $\mathbb{K}_n$ as a function of “$\mathbb{F}_n$" and “$\mathbb{G}_n$” in (ref).
lemLet $U_1,\dots,U_n$ be independent uniformly distributed random variables on $(0,1)$ and $\xi_1,\dots,\xi_n$ be i.i.d.\ random variables with mean $0$, variance $1$, and $\|\xi\|_{2,1}<\infty$, independent of $U_1,\dots,U_n$.
Define the bootstrap empirical process of $U$ by
\(
\mathbb{F}_n'(u)\vcentcolon=\frac{1}{n}\sum_{i=1}^n\xi_i\mathbbm{1}\{U_i\leq u\},
\)
and let $w_{i,n}'$ be the indicator of whether $U_i$ is above the $\alpha$\hypquantile of the bootstrap sample, that is, $w_{i,n}'\vcentcolon=\mathbbm{1}\{U_i>\mathbb{F}_n'^{-1}(\alpha)\}$.
Define
\(
\mathbb{G}_n'(u)\vcentcolon=\frac{1}{n}\sum_{i=1}^n\xi_iw_{i,n}'\mathbbm{1}\{U_i\leq u\}.
\)
Then, for $F(u)=0\vee u\wedge1$ and $G(u)=0\vee(u-\alpha)\wedge(1-\alpha)$,
\begin{gather*}
\sup_{h\inBL_1(L_\infty)}\bigl|\mathbb{E}_\xi h\bigl(\sqrt{n}(\mathbb{F}_n'-F)\bigr)-\mathbb{E}h\bigl(\sqrt{n}(\mathbb{F}_n-F)\bigr)\bigr|\operatorname*{\mathchoice{
\,\longrightarrow\,}{
\rightarrow}{
\rightarrow}{
\rightarrow}
}0, \\
\sup_{h\inBL_1(L_\infty)}\bigl|\mathbb{E}_\xi h\bigl(\sqrt{n}(\mathbb{G}_n'-G)\bigr)-\mathbb{E}h\bigl(\sqrt{n}(\mathbb{G}_n-G)\bigr)\bigr|\operatorname*{\mathchoice{
\,\longrightarrow\,}{
\rightarrow}{
\rightarrow}{
\rightarrow}
}0
\end{gather*}
in outer probability, and $\sqrt{n}(\mathbb{F}_n'-F)$ and $\sqrt{n}(\mathbb{G}_n'-G)$ are asymptotically measurable.
Altogether, nonparametric bootstrap works for $L$\hypstatistics when sample adjustment is based on empirical quantiles.
prop[Validity of nonparametric bootstrap]
In addition to assumptions in (ref)}, assume that $w_{i,n,j}$ represents sample adjustments based on a finite number of fixed quantiles.
\footnote{The assumption on convergence must be extended to jointly over all processes.}
Then, the joint distribution of $(\hat{\beta}_1, \dots, \hat{\beta}_d)$ can be consistently estimated by nonparametric bootstrap.
Proofs
Convergence of Bounded Integrable Processes
proof[Proof of (ref)]
Marginal convergence is trivial.
By vw1996, $\sqrt{n}(\mathbb{F}_n-F)$ converges weakly in $L_\infty$.
In light of vw1996 and (ref), it suffices to show that for $Z(x)\vcentcolon=\mathbbm{1}\{X\leq x\}-F(x)$, (i)
\(
\Pr(\|Z\|_\mu > t) = o(t^{-2})
\)
and (ii)
\(
\int_{\mathbb{R}} \sqrt{\mathbb{E}[Z^2]} |d\mu| < \infty.
\)
(ii) follows since $\int_{\mathbb{R}}\sqrt{\mathbb{E}[Z^2]}|d\mu|=\int_{\mathbb{R}}\sqrt{F(1-F)}|d\mu|$.
Let $m(x)\vcentcolon=\int_{(-\infty,x]}|d\mu|$ and $\tilde{F}(x)\vcentcolon=F(x)-\mathbbm{1}\{0\leq x\}$.
Note that (ii) implies that $m(X)$ has variance.
Writing $Z(x)=\mathbbm{1}\{X \leq x\}-\mathbbm{1}\{0\leq x\}-\tilde{F}(x)$, we find
\(
\|Z\|_\mu \leq |m(X)-m(0)| + \int_{\mathbb{R}} |\tilde{F}| |d\mu|.
\)
The second term is a finite constant if (ii) holds.
Thus, (i) holds if $\tilde{F} \circ m^{-1}(t) = o(t^{-2})$, which is the case if $m(X)$ has variance.
Thus, $\sqrt{n}(\mathbb{F}_n-F)$ converges weakly in $L_1(\mu)$.
lemIf $X_\alpha : \Omega_\alpha \to \mathbb{L}_\mu$ is asymptotically tight, it is asymptotically measurable if and only if $X_\alpha(t)$ is asymptotically measurable for every $t\in T$.
lemIf $X$ and $Y$ are tight Borel measurable maps into $\mathbb{L}_\mu$,
then $X$ and $Y$ are equal in law if and only if every marginal of $X$ and $Y$ is equal in law.
proof[Proofs of (ref)]
These claims are not corollaries of vw1996 since $C_b(\mathbb{L}_\mu)$ is bigger than $C_b(\mathbb{L}_T)$ and $C_b(\mathbb{L}_1)$, but they follow by the same logic.
proof[Proof of (ref)]
Necessity is immediate. We prove sufficiency.
If $X_\alpha$ is asymptotically tight and its marginals converge weakly, then $X_\alpha$ is asymptotically measurable by (ref).
By Prohorov's theorem vw1996, $X_\alpha$ is relatively compact.
Take any subnet in $X_\alpha$ that is convergent.
Its limit point is unique by (ref) and the assumption that every marginal converges weakly.
Thus, $X_\alpha$ converges weakly.
The last statement is another consequence of Prohorov's theorem.
proof[Proof of (ref)]
We proceed ((ref)) $\Rightarrow$ ((ref)) $\Rightarrow$ ((ref)) $\Rightarrow$ ((ref)).
((ref)) $\Rightarrow$ ((ref)). Fix $\varepsilon,\eta>0$.
Pick one $t_i$ from each $T_i$.
Then, $\|X_\alpha\|_T \leq \max_i |X_\alpha(t_i)|+\varepsilon$ with inner probability at least $1-\eta$.
Since the maximum of finitely many tight nets of real variables is tight and $\|X_\alpha\|_\mu$ is assumed to be tight, it follows that the net $\|X_\alpha\|_{\mathbb{L}_\mu}$ is asymptotically tight in $\mathbb{R}$.
Fix $\zeta>0$ and take $\varepsilon_m \searrow 0$. Let $M$ satisfy $\limsup P^\ast(\|X_\alpha\|_{\mathbb{L}_\mu}>M)<\zeta$. Taking $(\varepsilon,\eta)$ in ((ref)) as $(\varepsilon_m,2^{-m}\zeta)$, we obtain for each $m$ a measurable partition $T=\bigcup_{i=1}^k T_i$ (suppressing dependence on $m$). For each $T_i$, enumerate all of the finitely many values $0=a_{i,0}\leq a_{i,1}\leq\cdots\leq a_{i,p}\leq M$ such that
\[
\int_{T_i} (a_{i,j}-a_{i,j})|d\mu| \leq \frac{\varepsilon_m}{k} \quad \text{for} \quad j=1,\dots,p \quad \text{and} \quad
\int_{T_i} a_{i,p}|d\mu| \leq M.
\]
Since $\mu$ is not necessarily finite on the whole $T$, on some partition $T_i$ the only choice of $a_{i,j}$ may be $0$.
Let $z_1,\dots,z_q$ be the finite exhaustion of all functions in $\mathbb{L}_\mu$ that are constant on each $T_i$ and take values on
\[
0, \pm\varepsilon_m, \dots, \pm\lfloor M/\varepsilon_m \rfloor\varepsilon_m, \quad
\pm a_{1,1}, \dots, \pm a_{1,p}, \quad
\dots, \quad
\pm a_{k,1}, \dots, \pm a_{k,p}.
\]
Let $K_m$ be the union of $q$ closed balls of radius $2\varepsilon_m$ around each $z_i$.
Then, since $\inf_j \int_{T_i} |X_\alpha-a_{i,j}||d\mu|\leq\frac{\varepsilon_m}{k}+\inf_x \int_{T_i} |X_\alpha-x||d\mu|$, the three conditions
\(
\|X_\alpha\|_T \leq M,
\)
\(
\sup_i \sup_{s,t\in T_i} |X_\alpha(s)-X_\alpha(t)| \leq \varepsilon_m,
\)
and
\(
\sum_i \inf_x \int_{T_i} |X_\alpha-x| |d\mu| \leq \varepsilon_m
\)
imply that $X_\alpha\in K_m$. This holds for each $m$.
Let $K=\bigcap_{m=1}^\infty K_m$, which is closed, totally bounded, and therefore compact.
Moreover, we argue that for every $\delta>0$ there exists $m$ with $K^\delta \supset \bigcap_{j=1}^m K_j$. Suppose not.
Then there is a sequence $z_m$ not in $K^\delta$, but with $z_m \in \bigcap_{j=1}^m K_j$ for every $m$.
This has a subsequence contained in only one of the closed balls constituting $K_1$, and a further subsequence contained in only one of the balls constituting $K_2$, and so on.
The diagonal sequence of such subsequences would eventually be contained in a ball of radius $2\varepsilon_m$ for every $m$.
Therefore, it is Cauchy and its limit should be in $K$, which is a contradiction to the supposition $d(z_m,K)\geq\delta$ for every $m$.
Thus, if $X_\alpha$ is not in $K^\delta$, it is not in $\bigcap_{j=1}^m K_j$ for some $m$. Therefore,
\begin{multline*}
P^\ast(X_\alpha \notin K^\delta) \leq P^\ast \Biggl( X_\alpha \notin \bigcap_{j=1}^m K_j \Biggr)
\leq P^\ast(\|X_\alpha\|_{\mathbb{L}_\mu} > M) \\
+ \sum_{j=1}^m P^\ast \biggl( \biggl[ \sup_i \sup_{s,t \in T_i} |X_\alpha(s) - X_\alpha(t)| \biggr] \vee \sum_i \inf_x \int_{T_i} |X_\alpha - x| |d\mu| > \varepsilon_j \biggr) \\
\leq \zeta + \sum_{j=1}^m \zeta 2^{-j} < 2 \zeta.
\end{multline*}
Hence, we obtain $\limsup_\alpha P^\ast(X_\alpha \notin K^\delta) < 2 \zeta$, as asserted.
((ref)) $\Rightarrow$ ((ref)).
If $X_\alpha$ is asymptotically tight, then so is each coordinate projection.
Therefore, $X_\alpha(t)$ is asymptotically tight in $\mathbb{R}$ for every $t\in T$.
Let $K_1 \subset K_2 \subset \cdots$ be a sequence of compact sets such that $\liminf P_\ast(X_\alpha \in K_m^\varepsilon) \geq 1-1/m$ for every $\varepsilon>0$. Define a semimetric $d$ on $T$ induced by $z$ by
\(
d(s,t;z) \vcentcolon= |z(s) - z(t)|
\vee \int_T |z| \mathbbm{1}\{z(s) \wedge z(t) \leq z \leq z(s) \vee z(t)\} \mathbbm{1}\{z(s) \neq z(t)\} |d\mu|.
\)
Observe that $d(s,s;z)=0$ and that $d$ is measurable with respect to $\mu$.
\footnote{$T$ is not necessarily complete with respect to $d$.}
Now for every $m$, define a semimetric $\rho_m$ on $T$ by
\(
\rho_m(s,t) \vcentcolon= \sup_{z\in K_m} d(s,t;z).
\)
We argue that $(T,\rho_m)$ is totally bounded. For $\eta>0$, cover $K_m$ by finitely many balls of radius $\eta$ centered at $z_1, \dots, z_k$.
Consider the partition of $\mathbb{R}^{2k}$ into cubes of edge length $\eta$.
For each cube, if there exists $t\in T$ such that the following $2k$\hyptuple is in the cube,
\begin{multline*}
r(t) \vcentcolon= \biggl( z_1(t), \ \int_T z_1 \mathbbm{1} \{0 \wedge z_1(t) \leq z_1 \leq 0 \vee z_1(t)\} |d\mu|, \quad \dots, \\
z_k(t), \ \int_T z_k \mathbbm{1}\{0 \wedge z_k(t) \leq z_k \leq 0 \vee z_k(t)\} |d\mu| \biggr),
\end{multline*}
then pick one such $t$.
Since $\|z_j\|_{\mathbb{L}_\mu}$ is finite for every $j$ (i.e., the diameter of $T$ measured by each $d(\cdot,\cdot;z_j)$ is finite), this gives finitely many points $t_1,\dots,t_p$.
Notice that the balls $\{t : \rho_m(t,t_i)<3\eta\}$ cover $T$, that is, $t$ is in the ball around $t_i$ for which $r(t)$ and $r(t_i)$ are in the same cube; this follows because $\rho_m(t,t_i)$ can be bounded by
\(
2 \sup_{z \in K_m} \inf_j \|z-z_j\|_{\mathbb{L}_\mu} +{} \sup_j d(t,t_i;z_j) < 3 \eta.
\)
The first term is the error of approximating $z(t)$ and $z(t_i)$ by $z_j(t)$ and $z_j(t_i)$; the second is the distance of $t$ and $t_i$ measured by $d(\cdot, \cdot; z_j)$.
Define the semimetric $\rho$ by
\(
\rho(s,t) \vcentcolon= \sum_{m=1}^\infty 2^{-m} \bigl( \rho_m(s,t) \wedge 1 \bigr).
\)
We show that $(T,\rho)$ is still totally bounded. For $\eta>0$ take $m$ such that $2^{-m}<\eta$.
Since $T$ is totally bounded in $\rho_m$, we may cover $T$ with finitely many $\rho_m$\hypballs of radius $\eta$.
Denote by $t_1,\dots,t_p$ the centers of such a cover.
Since $K_m$ is nested, we have $\rho_1 \leq \rho_2 \leq \cdots$.
Since we also have $\rho_m(t,t_i)<\eta$, for every $t$ there exists $t_i$ such that $\rho(t,t_i)\leq\sum_{k=1}^m 2^{-k}\rho_k(t,t_i)+2^{-m}<2\eta$.
Therefore, $(T,\rho)$ is totally bounded.
By definition we have $d(s,t;z)\leq\rho_m(s,t)$ for every $z\in K_m$ and $\rho_m(s,t)\wedge1\leq2^m\rho(s,t)$.
And if $\|z_0-z\|_{\mathbb{L}_\mu}<\varepsilon$ for $z\in K_m$, then $d(s,t;z_0)<2\varepsilon+d(s,t;z)$ for every pair $(s,t)$.
Hence, we conclude
\(
K_m^\varepsilon \subset \bigl\{ z : \sup_{\rho(s,t)<2^{-m}\varepsilon}$ $d(s,t;z) \leq 3\varepsilon \bigr\}.
\)
Therefore, for $\delta<2^{-m}\varepsilon$,
\begin{multline*}
\liminf_\alpha P_\ast \biggl( \sup_{\rho(s,t)<\delta} d(s,t;X_\alpha) \leq 3 \varepsilon \biggr) \\
\begin{aligned}
&\geq \liminf_\alpha P_\ast \biggl( \sup_{t \in T} \biggl[ \sup_{\rho(s,t)<\delta} |X_\alpha(s) - X_\alpha(t)| \vee \int_{0<\rho(s,t)<\delta} |X_\alpha(s)| |d\mu| \biggr] \leq 3 \varepsilon \biggr) \\
&\geq 1 - \frac{1}{m}.
\end{aligned}
\end{multline*}
((ref)) $\Rightarrow$ ((ref)).
For $\varepsilon, \eta > 0$, take $\delta > 0$ as given.
Since $T$ is totally bounded, it can be covered with finitely many balls of radius $\delta$; let $t_1, \dots, t_K$ be their centers.
Disjointify the balls to obtain $\{T_i^\varepsilon\}$.
If $\int_{\{t_i\}} |X_\alpha| |d\mu| > 0$, then separate the partition $T_i^\varepsilon$ into $\{t_i\}$ and $T_i^\varepsilon \setminus \{t_i\}$.
There are three types of components in the partition: (a) singleton components (mass points) of $\mu$, (b) components with $|\mu|(T_i^\varepsilon) = \infty$, and (c) components with $|\mu|(T_i^\varepsilon) < \infty$.
The size of (a) is controlled by construction, so we control (b) and (c).
Clearly,
\begin{equation}
\sup_{s,t \in T_i^{\varepsilon}} |X_\alpha(s)-X_\alpha(t)| \leq 2 \sup_{\rho(s,t_i)<\delta} |X_\alpha(s)-X_\alpha(t_i)| \leq 2 \varepsilon.
\end{equation}
Denote by $i_\infty$ the index for which $|\mu|(T_{i_\infty}^\varepsilon) = \infty$.
Now we argue that $\sum_{i_\infty} \int_{T_{i_\infty}^\varepsilon} |X_\alpha| |d\mu|$ can be arbitrarily small (with inner probability at least $1-\eta$) for sufficiently small $\varepsilon$.
By construction, $\sup_{s \in T_{i_\infty}^\varepsilon} |X_\alpha(s)| \leq 2\varepsilon$.
\footnote{This follows because $\inf_{T_{i_\infty}^\varepsilon} |X_\alpha| = 0$ given that $\int_{T_{i_\infty}^\varepsilon} |X_\alpha| |d\mu| < \infty$.}
Thus, $\sum_{i_\infty} \int_{T_{i_\infty}^\varepsilon} |X_\alpha| |d\mu| \leq \int_T |X_\alpha| \mathbbm{1}\{|X_\alpha| \leq 2\varepsilon\} |d\mu|$.
Since $T$ is totally bounded, $\int_T |X_\alpha| |d\mu|$ is bounded by $K\varepsilon$ with inner probability at least $1-\eta$ (proving asymptotic tightness of $\|X_\alpha\|_\mu$), and hence the previous integral must be arbitrarily small for small $\varepsilon$.
Now we turn to (c).
Let $\varepsilon'$ be such that
\begin{equation}
\limsup_\alpha P^\ast \biggl( \int_T |X_\alpha| \mathbbm{1}\{|X_\alpha| \leq 3 \varepsilon'\} |d\mu| > \varepsilon \biggr) < 1 - \eta.
\end{equation}
For each $T_i^\varepsilon$ with $|\mu|(T_i^\varepsilon) < \infty$, construct a further partition of it $\{T_j^{\varepsilon'}\}$ with this $\varepsilon'$.
Note that $\{T_{i_\infty}^\varepsilon\} \cup \{T_j^{\varepsilon'}\}$ defines another finite partition of $T$.
If there exists $s \in T_j^{\varepsilon'}$ such that $|X_\alpha(s)| \leq \varepsilon'$, then by construction $\sup_{t \in T_j^{\varepsilon'}} |X_\alpha(t)| \leq 3 \varepsilon'$.
The contrapositive of this is also true.
Thus, observing
\(
\sum_j \inf_x \int_{T_j^{\varepsilon'}} |X_\alpha - x| |d\mu|
\leq \sum_j \inf_x \int_{T_j^{\varepsilon'}} |X_\alpha - x| \mathbbm{1}\{|X_\alpha| > \varepsilon'\} |d\mu| + \int_T |X_\alpha| \mathbbm{1}\{|X_\alpha| \leq 3 \varepsilon'\} |d\mu|,
\)
we may assume $\inf_{T'} |X_\alpha(s)| \geq \varepsilon' > 0$ at the cost of one more $\varepsilon$.
Then, we also have $\int_{T'} |d\mu| \leq K\varepsilon/\varepsilon'$ since $\varepsilon' \int_{T'} |d\mu| \leq \int_T |X_\alpha| |d\mu|$.
For the partition $T_j^{\varepsilon'}$ of $T'$, further construct a nested finite partition $T_k^{\varepsilon'/K}$. Now
\begin{multline*}
\sum_k \inf_x \int_{T_k^{\varepsilon'/K}} |X_\alpha - x| |d\mu| \\
\leq \sum_k \sup_{s, t \in T_k^{\varepsilon'/K}} |X_\alpha(s) - X_\alpha(t)| \int_{T_k^{\varepsilon'/K}} |d\mu|
\leq \frac{\varepsilon'}{K} \int_{T'} |d\mu| \leq \varepsilon
\end{multline*}
with inner probability at least $1-\eta$.
This, ((ref)), and ((ref)) yield the result.
proof[Proof of (ref)]
We first work on the case $r<1$.
Define
\(
\rho(s,t) \vcentcolon= \int_{(s,t)} u^r(1-u)^r dQ.
\)
We show that $(0,1)$ is totally bounded with respect to $\rho$.
Observe that (ref) ((ref)) and $r>\frac{1}{2+c}$ imply $u^r(1-u)^r Q(u) \to 0$ as $u\to\{0,1\}$.
Therefore, integrating by parts,
\[
\rho(0,1) \leq \int_{(0,1)} u^r \wedge(1-u)^r dQ \leq |Q|\Bigl(\frac{1}{2}\Bigr) + \int_0^{\frac{1}{2}} u^{r-1} |Q| du + \int_{\frac{1}{2}}^1 (1-u)^{r-1} |Q| du.
\]
Since $Q \in L_{2+c}$ and $u^{r-1}\wedge(1-u)^{r-1} \in L_q$ for every $q<1/(1-r)$, in particular for $q=(2+c)/(1+c)$, this integral is finite by H\"older's inequality.
This means the diameter of $(0,1)$ is finite, so $(0,1)$ is totally bounded.
Note that $|Q|$ is eventually smaller than $u^{-r}(1-u)^{-r}$ near $0$ and $1$, so that for every $\eta>0$ there exists $M$ such that
\[
\limsup_\alpha P^\ast(\|(|Q|\vee1)X_\alpha\|_\infty>M)\leq\limsup_\alpha P^\ast\biggl(\biggl\|\frac{X_\alpha}{u^r(1-u)^r}\biggr\|_\infty>M\biggr)<\eta.
\]
This shows uniform equicontinuity.
Next, for every $0<s\leq t<1$,
\[
\int_{(s,t)}|X_\alpha|dQ\leq\sup_{u\in(0,1)}\frac{|X_\alpha(u)|}{u^r(1-u)^r}\int_{(s,t)}v^r(1-v)^r dQ
\leq\biggl\|\frac{X_\alpha}{u^r(1-u)^r}\biggr\|_\infty\rho(s,t).
\]
Therefore,
\[
P^\ast \biggl( \sup_{t\in(0,1)} \int_{0<\rho(s,t)<\delta} |X_\alpha| dQ(s) > \varepsilon \biggr)
\leq P^\ast \biggl( \biggl\| \frac{X_\alpha}{u^r(1-u)^r} \biggr\|_\infty > \frac{\varepsilon}{\delta} \biggr).
\]
By assumption, this can be however small by the choice of $\delta$.
Conclude that $X_\alpha$ is asymptotically $(\rho,Q)$\hypequiintegrable in probability.
Finally, for $r \geq 1$, replace every $r$ by $1/2$. Then the result follows since $\bigl\| \frac{X_\alpha}{u^r(1-r)^r} \bigr\|_\infty \geq \bigl\| \frac{X_\alpha}{u^{1/2}(1-u)^{1/2}} \bigr\|_\infty$.
proof[Proof of (ref)]
Assume without loss of generality $M=1$.
Define $U_{(0)} \vcentcolon= 0$. Let $\tilde{\mathbb{F}}_n$ and $\tilde{\mathbb{G}}_n$ be the continuous linear interpolations of $\mathbb{F}_n$ and $\mathbb{G}_n$, that is, for $U_{(i-1)} \leq u < U_{(i)}$,
\(
\tilde{\mathbb{F}}_n(u) \vcentcolon= \frac{i-1}{n} + \frac{u-U_{(i-1)}}{n(U_{(i)}-U_{(i-1)})}
\)
and
\(
\tilde{\mathbb{G}}_n(u) \vcentcolon= \frac{1}{n} \sum_{i=1}^n w_{i,n} \mathbbm{1}\{U_i\leq u\} + \frac{w_{i,n}(u-U_{(i-1)})}{n(U_{(i)}-U_{(i-1)})},
\)
and for $u\geq U_{(n)}$, $\tilde{\mathbb{F}}_n(u) \vcentcolon= 1$ and $\tilde{\mathbb{G}}_n(u) \vcentcolon= \frac{1}{n} \sum w_{i,n}$.
Observe that $\mathbb{K}_n(u) = \tilde{\mathbb{G}}_n(\tilde{\mathbb{F}}_n^{-1}(u))$.
By (ref) it suffices to show that $\sqrt{n}(\tilde{\mathbb{F}}_n-I)$ and $\sqrt{n}(\tilde{\mathbb{G}}_n-K)$ converge weakly jointly in $\mathbb{L}_Q$.
Note that
\(
\|\mathbb{F}_n-I\|_\infty - \frac{1}{n} \leq \| \tilde{\mathbb{F}}_n-I \|_\infty \leq \|\mathbb{F}_n-I\|_\infty + \frac{1}{n}
\)
and
\(
\|\mathbb{F}_n-I\|_Q - C \leq \|\tilde{\mathbb{F}}_n-I\|_Q \leq \|\mathbb{F}_n-I\|_Q + C
\)
for $C\vcentcolon=\int (\tilde{I}-\lfloor n\tilde{I}\rfloor/n) dQ=O(1/n)$.
Thus, $\sqrt{n}(\tilde{\mathbb{F}}_n-I)$ converges weakly in $\mathbb{L}_Q$ if and only if $\sqrt{n}(\mathbb{F}_n-I)$ does, and they share the same limit.
The same is true for $\tilde{\mathbb{G}}_n$ and $\mathbb{G}_n$.
The classical results imply that $\sqrt{n}(\mathbb{F}_n-I)$ converges weakly in $L_\infty$ to a Brownian bridge and
\(
\bigl\| \frac{\sqrt{n}(\mathbb{F}_n-I)}{u^r(1-u)^r} \bigr\|_\infty = O_P(1)
\)
for every $r<1/2$ ch1993.
By (ref), it follows that $\sqrt{n}(\mathbb{F}_n-I)$ converges weakly in $\mathbb{L}_Q$.
By assumption $\sqrt{n}(\mathbb{G}_n-K)$ converges weakly in $L_\infty$ jointly with $\sqrt{n}(\mathbb{F}_n-I)$, and since $M|\mathbb{F}_n-I|\geq|\mathbb{G}_n-K|$, conclude that $\sqrt{n}(\mathbb{G}_n-K)$ converges weakly in $\mathbb{L}_Q$ jointly with $\sqrt{n}(\mathbb{F}_n-I)$.
lem[Inverse composition map]
Let $\mathbb{L}_Q$ contain the identity map $I(u)\vcentcolon=u$.
Let $\mathcal{D}$ be the subset of $\mathbb{L}_Q \times \mathbb{L}_Q$ such that every $(A,B) \in \mathcal{D}$ satisfies $A(u_1)-A(u_2)\geq B(u_1)-B(u_2)\geq0$ for every $u_1\geq u_2$, the range of $A$ contains $(0,1)$, and $B$ is differentiable and Lipschitz.
Let $\mathbb{L}_{Q,\textrm{UC}}$ be the subset of $\mathbb{L}_Q$ of uniformly continuous functions.
Then, the map $\chi : \mathcal{D} \to \mathbb{L}_Q$, $\chi(A,B) \vcentcolon= B\circ A^{-1}$, is Hadamard differentiable at $(A,B) \in \mathcal{D}$ for $A=I$ tangentially to $\mathbb{L}_Q \times \mathbb{L}_{Q,\textrm{UC}}$. The derivative is given by
\(
\chi_{I,B}'(a,b)(u) = b(u) + B'(u) a(u)$ for $u \in (0,1).
\)
proofFor $(A,B)\in\mathcal{D}$ and $u_1 \geq u_2$, denote $v_1\vcentcolon=A(u_1)$ and $v_2\vcentcolon=A(u_2)$.
By assumption we have $v_1-v_2\geq B(A^{-1}(v_1))-B(A^{-1}(v_2))\geq0$ for every $v_1\geq v_2$.
Therefore, $B\circ A^{-1}$ is monotone and bounded by the identity map up to a constant.
This implies
\(
\int_{(0,1)} \bigl| \widetilde{B \circ A^{-1}} \bigr| dQ \leq \int_{(0,1)} |\tilde{I}| dQ < \infty
\)
and $\|Q(B\circ A^{-1})\|_\infty<\infty$; it follows that $B \circ A^{-1}$ is in $\mathbb{L}_{Q,1}$.
Let $a_t \to a$ and $b_t \to b$ in $\mathbb{L}_Q$ and $(A_t,B_t)\vcentcolon=(I+ta_t,B+tb_t) \in \mathcal{D}$.
We want to show
\(
\bigl\| \frac{B_t\circ A_t^{-1}-B\circ I^{-1}}{t} - b - B' a \bigr\|_{\mathbb{L}_Q} \operatorname*{\mathchoice{
\,\longrightarrow\,}{
\rightarrow}{
\rightarrow}{
\rightarrow}
} 0
\)
as $t\to0$.
That $\|\cdot\|_{Q,\infty} \to 0$ follows by applying vw1996 to $(A^{-1},QB)$ as elements in $L_\infty$.
Thus, it remains to show $\|\cdot\|_Q \to 0$.
In the assumed inequality, substitute $(u_1,u_2)$ by $(u,A_t^{-1}(u))$ to find that
\(
|A_t(u)-u| \geq |B_t(A_t^{-1}(u))-B_t(u)| \geq 0.
\)
Therefore, the following inequality holds pointwise:
\(
|B_t\circ A_t^{-1}-B| \leq |B_t\circ A_t^{-1}-B_t|+|B_t-B| \leq |A_t-I|+|B_t-B|=|ta_t|+|tb_t|.
\)
For $\varepsilon>0$, write $\|\cdot\|_Q$ as
\(
\bigl( \int_0^{\varepsilon} + \int_\varepsilon^{1-\varepsilon} + \int_{1-\varepsilon}^1 \bigr) \bigl| \frac{B_t\circ A_t^{-1}-B}{t} - b - B' a \bigr| dQ.
\)
For any fixed $\varepsilon>0$ the middle term vanishes as $t\to0$ since $\|\cdot\|_{Q,\infty}\to0$.
It remains to show that the first term can be arbitrarily small since then by symmetry the third term is also ignorable.
Using the above inequality, write
\(
\int_0^{\varepsilon} \bigl| \frac{B_t\circ A_t^{-1}-B}{t} - b - B' a \bigr| dQ \leq \int_0^{\varepsilon} (|a_t|+|b_t|+|b|+|B'a|) dQ.
\)
Since $\|a_t-a\|_Q\to0$ and $\|b_t-b\|_Q\to0$, this integral can be arbitrarily small by the choice of $\varepsilon$, as desired.
Convergence of Quantile Processes as Integrable Processes
If $m$ is an identity, we denote $\mathbb{L}_m$ and $\mathbb{L}_{m,\phi}$ by $\mathbb{L}$ and $\mathbb{L}_\phi$, and $\|\cdot\|_m$ by $\|\cdot\|_1$.
We first establish differentiability of the inverse map for distribution functions with finite first moments.
lem[Inverse map]
Let $F \in \mathbb{L}_\phi$ be a distribution function on (an interval of) $\mathbb{R}$ that has at most finitely many jumps and is otherwise continuously differentiable with strictly positive density $f$.
Then, the inverse map $\phi : \mathbb{L}_\phi \to \mathbb{B}$, $\phi(F)\vcentcolon=Q=F^{-1}$, is Hadamard differentiable at $F$ tangentially to the set $\mathbb{L}_0$ of all continuous functions in $\mathbb{L}$. The derivative is
\(
\phi'_F(z) = -(z\circ Q) Q'.
\)
proof[Proof of (ref)]
Take $z_t \to z$ in $\mathbb{L}$ and $F_t \vcentcolon= F+tz_t \in \mathbb{L}_\phi$.
We want to show
\(
\bigl\| \frac{\phi(F_t)-\phi(F)}{t}-\phi'_F(z) \bigr\|_{\mathbb{B}}=\int_0^1 \bigl| \frac{\phi(F_t)-\phi(F)}{t}-\phi'_F(z) \bigr| du \operatorname*{\mathchoice{
\,\longrightarrow\,}{
\rightarrow}{
\rightarrow}{
\rightarrow}
} 0
\)
as $t \to 0$.
Let $j\in\mathbb{R}$ be a point of jump of $F$.
For small $\varepsilon>0$, split the integral as
\[
\biggl( \int_0^{F(j-\varepsilon)} + \int_{F(j-\varepsilon)}^{F(j+\varepsilon)} + \int_{F(j+\varepsilon)}^1 \biggr) \biggl| \frac{\phi(F_t)-\phi(F)}{t}-\phi'_F(z) \biggr| du.
\]
Observe that the second integral is bounded by
\[
2\varepsilon\biggl\|\frac{F_t-F}{t}\biggr\|_\infty + \biggl( \int_{F(j-\varepsilon)}^{F(j-)} + \int_{F(j)}^{F(j+\varepsilon)} \biggr) |\phi_F'(z)|du.
\]
The first term of this equals $2\varepsilon\|z_t\|_\infty$ and can be arbitrarily small by the choice of $\varepsilon$.
If $\varepsilon$ is small enough that there is no other jump in $[j-\varepsilon,j+\varepsilon]$, by Fubini's theorem,
\(
\int_{F(j-\varepsilon)}^{F(j-)}|\phi_F'(z)|du = \int_{j-\varepsilon}^{j} \bigl|\frac{z}{f}\bigr|dF \leq \varepsilon\|z\|_\infty,
\)
which can be, again, arbitrarily small.
Similarly for the last integral.
Therefore, we ignore finitely many jumps of $F$ so that $f>0$ everywhere.
For $\varepsilon>0$ there exists $M$ such that $F(-M) < \varepsilon$ and $1-F(M) < \varepsilon$.
Write
\begin{multline*}
\biggl\| \frac{\phi(F_t)-\phi(F)}{t}-\phi'_F(z) \biggr\|_{\mathbb{B}}
\leq \int_{F(-M)+\varepsilon}^{F(M)-\varepsilon} \biggl| \frac{\phi(F_t)-\phi(F)}{t}-\phi'_F(z) \biggr| du \\
+ \biggl( \int_0^{2\varepsilon} + \int_{1-2\varepsilon}^1 \biggr) \biggl| \frac{\phi(F_t)-\phi(F)}{t}-\phi'_F(z) \biggr| du.
\end{multline*}
By vw1996, the integrand vanishes uniformly on $[F(-M)+\varepsilon,F(M)-\varepsilon]$.
As the first integral is bounded by
\(
\sup_{u\in[F(-M)+\varepsilon,F(M)-\varepsilon]}$ $\bigl|\frac{\phi(F_t)-\phi(F)}{t}-\phi_F'(z)\bigr|,
\)
it vanishes as $t\to0$.
Now turn to the second integral. The triangle inequality bounds it by
\(
\int_0^{2\varepsilon} \bigl| \frac{\phi(F_t)-\phi(F)}{t} \bigr| du + \int_0^{2\varepsilon} |\phi'_F(z)| du.
\)
Since $F$ and $F_t$ are nondecreasing, by Fubini's theorem,
\begin{multline*}
\int_0^{2\varepsilon} \biggl| \frac{\phi(F_t)-\phi(F)}{t} \biggr| du = \frac{1}{|t|} \int_0^{2\varepsilon} \bigl| F_t^{-1}-F^{-1} \bigr| du \\
\leq \frac{1}{|t|} \int_{-\infty}^{F^{-1}(2\varepsilon+\|tz_t\|_\infty)} |tz_t| dx \leq \|z_t-z\|_1 + \int_{-\infty}^{F^{-1}(2\varepsilon+t\|z_t\|_\infty)} |z| dx.
\end{multline*}
The first term goes to $0$ and the second term can be arbitrarily small by the choice of $\varepsilon$.
Finally, by the change of variables,
\(
\int_0^{2\varepsilon} |\phi'_F(z)| du = \int_{-\infty}^{F^{-1}(2\varepsilon)} \bigl| \frac{z}{f} \bigr| dF = \int_{-\infty}^{F^{-1}(2\varepsilon)} |z| dx,
\)
which can be arbitrarily small.
Likewise, the integral from $1-2\varepsilon$ to $1$ converges to $0$.
This completes the proof.
Now we allow transformations of locally bounded variation.
A function of locally bounded variation admits decomposition into the difference of two monotone functions.
Then, we exploit the relationship $m(F^{-1}) = (F \circ m^{-1})^{-1}$ for a monotone $m$ and use the chain rule.
proof[Proof of (ref)]
Since $m$ is of locally bounded variation, write $m(x) = m_1(x)-m_2(x)$ where $m_1$ and $m_2$ are increasing.
Moreover, $m_1$ and $m_2$ can be chosen to be continuously differentiable and strictly increasing, and for their corresponding Lebesgue\hypStieltjes measures $\mu_1$ and $\mu_2$, $F$ belongs to both $\mathbb{L}_{\mu_1,\phi}$ and $\mathbb{L}_{\mu_2,\phi}$.
Since the derivative formula is linear in $m'$, it suffices to show the claim for $m_1$ and $m_2$ separately.
Now observe that $z$ is in $\mathbb{L}_\mu$ (or $\mathbb{L}_{\mu,0}$) if and only if $z \circ m^{-1}$ is in $\mathbb{L}$ (or $\mathbb{L}_0$).
The assertion then follows by vw1996 applied to (ref).
lemLet $m : \mathbb{R} \to \mathbb{R}$ be a strictly increasing continuous function and $\mu$ be the associated Lebesgue\hypStieltjes measure.
Then, the map $\psi : \mathbb{L}_\mu \to \mathbb{L}$, $\psi(F) \vcentcolon= F \circ m^{-1}$, is
uniformly Fr\'echet differentiable with rate function $q \equiv 0$.
\footnote{A map $\psi : \mathbb{L} \to \mathbb{B}$ is {\em uniformly Fr\'echet differentiable with rate function $q$} if there exists a continuous linear map $\psi_F' : \mathbb{L} \to \mathbb{B}$ such that $\|\psi(F+z)-\psi(F)-\psi_F'(z)\|_{\mathbb{B}}=O(q(\|z\|_{\mathbb{L}}))$ uniformly over $F\in\mathbb{L}$ as $z \to 0$ and $q$ is monotone with $q(t)=o(t)$.}
The derivative is given by $\psi_F'(z) \vcentcolon= z \circ m^{-1}$.
proof[Proof of (ref)]
Observe that $\psi(F+z) - \psi(F) = (F+z)(m^{-1}) - F(m^{-1}) = z(m^{-1})$.
Therefore, $\psi(F+z)-\psi(F)-\phi_F'(z) = 0$.
proof[Proof of (ref)]
This follows from (ref).
Convergence of $L$\hypstatistics
proof[Proof of (ref)]
The derivative map is linear by construction; it is also continuous since
\(
|\lambda_{Q,K}'(z_1,\kappa_1)-\lambda_{Q,K}'(z_2,\kappa_2)|=\bigl|\int_0^1Qd(\kappa_1-\kappa_2)+\int_0^1(z_1-z_2)dK\bigr|
\leq\|\kappa_1-\kappa_2\|_{Q,\infty}+\|\kappa_1-\kappa_2\|_{Q}+M\|z_1-z_2\|_{\mathbb{B}},
\)
which vanishes as $\|z_1-z_2\|_{\mathbb{B}}\to0$ and $\|\kappa_1-\kappa_2\|_{\mathbb{L}_Q}\to0$.
Let $z_t\to z$ and $\kappa_t\to\kappa$ such that $Q_t\vcentcolon=Q+tz_t$ is in $\mathbb{B}$ and $K_t\vcentcolon=K+t\kappa_t$ is in $\mathbb{L}_{Q,M}$. Observe
\[
\frac{\lambda(Q_t,K_t)-\lambda(Q,K)}{t} - \lambda'_{Q,K}(z_t,\kappa_t) = \int (z_t-z) d(K_t-K) + \int z d(K_t-K).
\]
The first term vanishes since
\(
\bigl| \int (z_t-z) d(K_t-K) \bigr| \leq 2M \int |z_t-z| du = 2M \|z_t-z\|_{\mathbb{B}}.
\)
As $z$ is integrable, for every $\varepsilon>0$ there exists a small number $\delta>0$ such that
\(
\bigl(\int_0^\delta+\int_{1-\delta}^1\bigr) |z| du + \int_\delta^{1-\delta} (|z| - (|z| \wedge \delta^{-1})) du \leq \varepsilon.
\)
This gives
\begin{align*}
\biggl| \int z d(K_t-K) \biggr| &\leq \biggl| \int_\delta^{1-\delta} (-\delta^{-1} \vee z \wedge \delta^{-1}) d(K_t-K) \biggr| \\
&\hphantom{=} + \biggl| \int z d(K_t-K) - \int_\delta^{1-\delta} (-\delta^{-1} \vee z \wedge \delta^{-1}) d(K_t-K) \biggr| \\
&\leq \biggl| \int_\delta^{1-\delta} (-\delta^{-1} \vee z \wedge \delta^{-1}) d(K_t-K) \biggr| + 2M\varepsilon.
\end{align*}
Let $\tilde{z} \vcentcolon= -\delta^{-1} \vee z \wedge \delta^{-1}$.
Since $\tilde{z}$ is ladcag on $[\delta,1-\delta]$, there exists a partition $\delta=t_0<t_1<\cdots<t_m=1-\delta$ such that $\tilde{z}$ varies less than $\varepsilon$ on each interval $(t_{i-1},t_i]$.
Let $\bar{z}$ be the piecewise constant function that equals $\tilde{z}(t_i)$ on each interval $(t_{i-1},t_i]$. Then
\begin{multline*}
\biggl| \int_\delta^{1-\delta} \tilde{z} d(K_t-K) \biggr| \leq 2M \sup_{u\in[\delta,1-\delta]}|\tilde{z}-\bar{z}| + |\tilde{z}(\delta)| |(K_t-K)(\{\delta\})| \\
+ \sum_{i=1}^m |\tilde{z}(t_i)| |(K_t-K)((t_{i-1},t_i])|.
\end{multline*}
The first term is arbitrarily small by the choice of $\varepsilon$, and the second and third terms are collectively bounded by $(2m+1)\delta^{-1}\|K_t-K\|_\infty=(2m+1)\delta^{-1}t\|\kappa_t\|_\infty$, which converges to $0$ regardless of the choice of $K$.
The proof for $\tilde{\lambda}$ is basically the same.
proof[Proof of (ref)]
Weak convergence follows from (ref).
The derivative formulas give us
\begin{align*}
&\operatorname{Cov}(\xi_j, \xi_k) = \int_0^1 \int_0^1 (m_j'\circ Q_j)Q_j'(s)(m_k'\circ Q_k)Q_k'(t) [F_{ik}^Q(s,t)-st] ds dt \\
&+ \int_0^1 \int_0^1 (m_j'\circ Q_j)Q_j'(s)(m_k'\circ Q_k)Q_k'(t) [K_{jk}F_{jk}^Q(s,t)-stK_j(s)K_k(t)] ds dt \\
&- \int_0^1 \int_0^1 (m_j'\circ Q_j)Q_j'(s)(m_k'\circ Q_k)Q_k'(t) K_j(s) [F_{jk}^Q(s,t)-st] ds dt \\
&- \int_0^1 \int_0^1 )m_j'\circ Q_j)Q_j'(s)(m_k'\circ Q_k)Q_k'(t) K_k(t) [F_{jk}^Q(s,t)-st] ds dt.
\end{align*}
Consistency of the sample analogue estimator follows from uniform convergence of $\mathbb{K}_{n,j}^F$ and $\mathbb{K}_{n,k}^F$ and (ref).
lemLet $m : \mathbb{R} \to \mathbb{R}$ be a ladcag increasing function. For a probability measure $F$ on $\mathbb{R}$ such that $\mathbb{E}[m(X)]<\infty$, $X\sim F$, we have
\begin{gather*}
\bigl\| m(t) \mathbb{F}_n(t) - m(t) F(t) \bigr\|_\infty \operatorname*{\mathchoice{
\,\longrightarrow\,}{
\rightarrow}{
\rightarrow}{
\rightarrow}
}^{\operatorname{as\ast}} 0, \quad
\biggl\| \int_{[s,t]} |m| d\mathbb{F}_n - \int_{[s,t]} |m| dF \biggr\|_\infty \operatorname*{\mathchoice{
\,\longrightarrow\,}{
\rightarrow}{
\rightarrow}{
\rightarrow}
}^{\operatorname{as\ast}} 0, \\
\biggl\| \int_{[s,t]} |\tilde{\mathbb{F}}_n| dm - \int_{[s,t]} |\tilde{F}| dm \biggr\|_\infty \operatorname*{\mathchoice{
\,\longrightarrow\,}{
\rightarrow}{
\rightarrow}{
\rightarrow}
}^{\operatorname{as}} 0, \qquad
\int_{\mathbb{R}} |\mathbb{F}_n - F| dm \operatorname*{\mathchoice{
\,\longrightarrow\,}{
\rightarrow}{
\rightarrow}{
\rightarrow}
}^{\operatorname{as}} 0,
\end{gather*}
where the suprema are each taken over $t \in \mathbb{R}$, $(s,t) \in \overline{\mathbb{R}}^2$, and $(s,t) \in \overline{\mathbb{R}}^2$.
proof[Proof of (ref)]
We assume $m(0)=0$ without loss of generality.
In view of vw1996, the first two claims follow if
\begin{gather*}
\mathcal{F} = \bigl\{ f_t : \mathbb{R} \to \mathbb{R} : t \in \overline{\mathbb{R}}, \, f_t(x) = m(t) \mathbbm{1}\{x \leq t\} \bigr\}, \\
\mathcal{G} = \bigl\{ g_{s,t} : \mathbb{R} \to \mathbb{R} : s, t \in \overline{\mathbb{R}}, \, g_{s,t}(x) = |m(x)| \mathbbm{1}\{s \leq x \leq t\} \bigr\}
\end{gather*}
have finite bracketing numbers with respect to $L_1(P)$.
For $\mathcal{F}$ take $-\infty = t_0 < t_1 < \cdots < t_m = \infty$ such that $|\int (f_{t_{i+1}} - f_{t_i}) dF| < \varepsilon$ for each $i$ and consider the brackets $\{f_{t_i}\}$.
\footnote{If $F$ has a probability mass at $t$, then for small $\varepsilon$ take, instead of $f_t$, $\tilde{f}_{t,c}(x) = m(t) [c \mathbbm{1}\{x \leq t\} + (1-c) \mathbbm{1}\{x < t\}]$ for appropriately chosen $c$.}
This partition is finite by (ref) and $\mathbb{E}[m(X)]<\infty$.
For $\mathcal{G}$ take $-\infty = t_0 < t_1 < \cdots < t_m = \infty$ such that $\bigl| \int_{(-\infty,t_{i+1}]} |m| dF - \int_{(-\infty,t_i]} |m| dF \bigr| < \varepsilon$ for each $i$ and consider the brackets $\{g_{s,t}\}$ for every pair $s, t \in \{t_0, \dots, t_m\}$.
\footnote{Again, if $F$ has a mass, similar adjustments are needed.}
This partition is finite by $\mathbb{E}[m(X)]<\infty$.
For the third claim, observe that
\(
\int_{[s,t]} |m| d\mathbb{F}_n = \int_{[s,t]} |m| d\tilde{\mathbb{F}}_n
= \bigl[ |m| \tilde{\mathbb{F}}_n \bigr]_s^t + \int_{[s,t]} |\tilde{\mathbb{F}}_n| d\mu.
\)
Then the claim follows by the first two claims and the triangle inequality,
\(
\bigl\| \int_{[s,t]} |\tilde{\mathbb{F}}_n| d\mu - \int_{[s,t]} |\tilde{F}| d\mu \bigr\|_\infty
\leq 2 \bigl\| m(t) \mathbb{F}_n(t) - m(t) F(t) \bigr\|_\infty
+ \bigl\| \int_{[s,t]} |m| d\mathbb{F}_n - \int_{[s,t]} |m| dF \bigr\|_\infty.
\)
\footnote{Measurability of the sup on the LHS follows by the continuity of Lebesgue integrals.}
For the last claim, observe that (ref) and the preceding claim imply that for $\varepsilon > 0$ there exists $M < \infty$ such that
\(
\bigl( \int_{(-\infty,-M]} + \int_{[M,\infty)} \bigr) |\tilde{\mathbb{F}}_n| d\mu + \bigl( \int_{(-\infty,-M]} + \int_{[M,\infty)} \bigr) |\tilde{F}| d\mu < \varepsilon
\)
with probability tending to $1$. By the triangle inequality,
\(
\int_{\mathbb{R}} \bigl| \tilde{\mathbb{F}}_n - \tilde{F} \bigr| d\mu \leq \int_{(-M,M)} \bigl| \tilde{\mathbb{F}}_n - \tilde{F} \bigr| d\mu + \varepsilon
\leq \|\mathbb{F}_n - F\|_\infty \mu((-M,M)) + \varepsilon.
\)
Then the assertion follows by the Glivenko\hypCantelli theorem.
Validity of Nonparametric Bootstrap
We start with the key lemma in Poissonization, the counterpart of vw1996.
lemFor each $n$, let $(W_{n1},\dots,W_{nn})$ be an exchangeable nonnegative random vector independent of $X_1,X_2,\dots$ such that $\sum_{i=1}^nW_{ni}=1$ and $\max_{1\leq i\leq n}|W_{ni}|$ converges to zero in probability.
Let $F$ be a probability distribution on $\mathbb{R}$ such that $\int_{\mathbb{R}}\sqrt{F(1-F)}|d\mu|<\infty$.
Then, for every $\varepsilon>0$, as $n\to\infty$,
\(
{\Pr}_W\bigl(\bigl\|\sum_{i=1}^nW_{ni}\bigl(\mathbbm{1}\{X_i\leq x\}-F(x)\bigr)\bigr\|_\mu^\ast>\varepsilon\bigr)\operatorname*{\mathchoice{
\,\longrightarrow\,}{
\rightarrow}{
\rightarrow}{
\rightarrow}
}^{\operatorname{as\ast}}0.
\)
proof[Proof of (ref)]
Assume without loss of generality that $\mu$ is a positive measure and let $m(x)\vcentcolon=\mu([0,x))$ for $x\geq0$ and $\mu([x,0))$ for $x<0$.
Since vw1996 goes through with $\|\cdot\|_{\mathbb{L}_\mu}$, the proof of this lemma is almost identical to vw1996.
Essentially, the only part that requires modification is boundedness of $n^{-1}\sum_{i=1}^n\|\mathbbm{1}\{X_i\leq x\}-F(x)\|_\mu^r$ ($r<1$).
Note that
\(
|\mathbbm{1}\{X_i\leq x\}-F(x)|\leq|\mathbbm{1}\{X_i\leq x\}-\mathbbm{1}\{0\leq x\}|+|\tilde{F}(x)|.
\)
Therefore,
\(
\|\mathbbm{1}\{X_i\leq x\}-F(x)\|_\mu
\leq m(X_i)+\|\tilde{F}\|_\mu.
\)
Find that
\(
\frac{1}{n}\sum_{i=1}^n\|\mathbbm{1}\{X_i\leq x\}-F(x)\|_\mu^r\leq\frac{1}{n}\sum_{i=1}^n m(X_i)^r+\|\tilde{F}\|_\mu^r,
\)
which converges almost surely to $\mathbb{E}[m(X_i)^r]+\|\tilde{F}\|_\mu^r<\infty$.
Given this, we infer as in vw1996 that conditional weak convergence of $\hat{\mathbb{Z}}_n$ follows from conditional weak convergence of $\mathbb{Z}_n'$.
For the latter, we first need to show unconditional convergence of $\mathbb{Z}_n'$ in our norm.
The following is a modification of vw1996.
lemLet $\xi_1,\dots,\xi_n$ be i.i.d.\ random variables with mean zero, variance $1$, and $\|\xi\|_{2,1}<\infty$, independent of $X_1,\dots,X_n$.
For a probability distribution $F$ on $\mathbb{R}$ such that $m(X)$ has a $(2+c)$th moment for $X\sim F$ and some $c>0$, the process $\mathbb{Z}_n'(x)\vcentcolon=n^{-1/2}\sum_{i=1}^n\xi_i[\mathbbm{1}\{X_i\leq x\}-F(x)]$ converges weakly to a tight limit process in $\mathbb{L}_\mu$ if and only if $\mathbb{Z}_n\vcentcolon=n^{-1/2}\sum_{i=1}^n[\mathbbm{1}\{X_i\leq x\}-F(x)]$ does.
In that case, they share the same limit processes.
proof[Proof of (ref)]
Marginal convergence and asymptotic equicontinuity of $\mathbb{Z}_n'$ follow from (ref) and vw1996.
It remains to show the equivalence of asymptotic equiintegrability of $\mathbb{Z}_n'$ and $\mathbb{Z}_n$.
Note that the proofs of vw1996 do not depend on the specificity of the norm $\|\cdot\|_{\mathcal{F}}$, but they continue to hold with $\|\cdot\|_{\mathbb{L}_\mu}$.
Given this, vw1996 also holds with $\|\cdot\|_{\mathbb{L}_\mu}$ (and $\|\cdot\|_{\mathbb{L}_{\mu,\delta_n}}$).
Finally, rewriting the proof of vw1996 in terms of $\|\cdot\|_{\mathbb{L}_\mu}$ yields the proof of this lemma.
proof[Proof of (ref)]
By (ref), $\mathbb{Z}_n'$ is asymptotically measurable.
Define a semimetric on $\mathbb{R}$ by
\(
\rho(s,t)\vcentcolon=|F(s)-F(t)|\vee\int_s^t\sqrt{F(1-F)}|d\mu|.
\)
For $\delta>0$, $t_1<\cdots<t_p$ be such that $\rho(-\infty,t_1)\leq\delta$, $\rho(t_j,t_{j+1})\leq\delta$, and $\rho(t_p,\infty)\leq\delta$.
Define $\mathbb{Z}_\delta$ by
\[
\mathbb{Z}_\delta(x)\vcentcolon=\begin{cases} 0 & x<t_1\text{ or }x\geq t_p, \\ \mathbb{Z}(t_i) & t_i\leq x\leq t_{i+1}, \, i=1,\dots,p-1. \end{cases}
\]
Define $\mathbb{Z}_{n,\delta}'$ analogously.
By the continuity and integrability of the limit process $\mathbb{Z}$, we have $\mathbb{Z}_\delta\to\mathbb{Z}$ in $\mathbb{L}_\mu$ almost surely as $\delta\to0$.
Therefore,
\(
\sup_{h\in\text{BL}_1(\mathbb{L}_\mu)}\bigl|\mathbb{E}h(\mathbb{Z}_\delta)-\mathbb{E}h(\mathbb{Z})\bigr|\operatorname*{\mathchoice{
\,\longrightarrow\,}{
\rightarrow}{
\rightarrow}{
\rightarrow}
}0$ as $\delta\to0.
\)
Second, by vw1996,
\(
\sup_{h\in\text{BL}_1(\mathbb{L}_\mu)}\bigl|\mathbb{E}_\xi h(\mathbb{Z}_{n,\delta}')-\mathbb{E}h(\mathbb{Z}_\delta)\bigr|\operatorname*{\mathchoice{
\,\longrightarrow\,}{
\rightarrow}{
\rightarrow}{
\rightarrow}
}0$ as $n\to\infty
\)
for almost every sequence $X_1,X_2,\dots$ and fixed $\delta>0$.
Since $\mathbb{Z}_\delta$ and $\mathbb{Z}_{n,\delta}'$ take only on a finite number of values and their tail values are zero, one can replace the supremum over $\text{BL}_1(\mathbb{L}_\mu)$ with a supremum over $\text{BL}_1(\mathbb{R}^p)$.
Observe that $\text{BL}_1(\mathbb{R}^p)$ is separable with respect to the topology of uniform convergence on compact sets; this supremum is effectively over a countable set, hence measurable.
Third,
\(
\sup_{h\in\text{BL}_1(\mathbb{L}_\mu)}\bigl|\mathbb{E}_\xi h(\mathbb{Z}_{n,\delta}')-\mathbb{E}_\xi h(\mathbb{Z}_n')\bigr| \leq \sup_{h\in\text{BL}_1(\mathbb{L}_\mu)}\mathbb{E}_\xi\bigl|h(\mathbb{Z}_{n,\delta}')-h(\mathbb{Z}_n')\bigr|
\leq \mathbb{E}_\xi\|\mathbb{Z}_{n,\delta}'-\mathbb{Z}_n'\|_{\mathbb{L}_\mu}^\ast \leq \mathbb{E}_\xi\|\mathbb{Z}_n'\|_{\mathbb{L}_{\mu,\delta}^\ast}.
\)
This implies that its outer expectation is bounded by $\mathbb{E}^\ast\|\mathbb{Z}_n'\|_{\mathbb{L}_{\mu,\delta}}$, which vanishes as $n\to\infty$ by the modified vw1996 as discussed in (ref).
proof[Proof of (ref)]
Noting $\mathbb{G}_n'(u)=0\vee[\mathbb{F}_n'(u)-\mathbb{F}_n'\circ\mathbb{F}_n'^{-1}(\alpha)]$, weak convergence of $\sqrt{n}(\mathbb{F}_n'-F)$ and $\sqrt{n}(\mathbb{G}_n'-G)$ follows from (ref) (or vw1996) and (ref).
proof[Proof of (ref)]
With the remark below (ref), the proposition follows from (ref) and vw1996.