The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
66,201 characters
Closed-form estimation and uniform inference in additively separable triangular models with a nonseparable first stage
\title[Nonseparable first stage]{Closed-form estimation and uniform inference in additively separable
triangular models with a nonseparable first stage}
\author[K.~Sunada]{Keita Sunada}
\address{Aarhus Center for Econometrics, Aarhus University, Universitetsbyen
51, 8000 Aarhus C, Denmark.}
\email{[email removed]}
\thanks{I am grateful to Nese Yildiz and Bin Chen for their helpful comments.
This research was supported in part by the Aarhus Center for Econometrics
(ACE) funded by the Danish National Research Foundation grant number
DNRF186.}
\date{August 23, 2026}
\begin{abstract}
This paper studies the nonparametric identification and estimation
of additively separable triangular models with continuous endogenous
and instrumental variables, allowing for a nonseparable first-stage
equation. Under the independence of instrumental variables and unobservables,
we show that the outcome function possesses a closed-form expression
as a functional of conditional cumulative distribution functions.
The resulting plug-in estimators require no regularization and converge
at the rate $n^{-m/(2m+1)}$, where $m$ is the order of smoothness
the model imposes. Also, no estimator of the outcome function converges
faster. We use the empirical bootstrap to construct a uniform confidence
band that covers the outcome function at every point of a compact
set simultaneously.
\end{abstract}
\maketitle
\section{Introduction}
This paper studies the identification and estimation of a system of
structural equations that takes the following form:
\begin{align}
Y & =g\left(X\right)+\varepsilon,\label{eq:model}\\
X & =h\left(Z,\eta\right),\nonumber
\end{align}
where $X$ is a continuous scalar explanatory variable (treatment),
$\varepsilon$ and $\eta$ are scalar unobservables, $Z$ is a continuous
scalar instrumental variable, and $g$ and $h$ are unknown functions.
We assume that instrument $Z$ and unobservables $\left(\varepsilon,\eta\right)$
are independent.
A key feature of (\ref{eq:model}) is that the first-stage equation
$X=h(Z,\eta)$ allows for nonseparable interactions between the instrument
and the unobservable. This is empirically relevant because in many
applications, the validity of an instrument is derived from economic
theory, and the resulting first-stage relationship is inherently nonseparable.
For instance, in demand estimation, researchers commonly use demand
shifters or product characteristics as instruments for price. These
instruments affect price through the markup implied by the firm's
first-order condition, and without strong functional-form assumptions,
the instrument and the unobservable can interact in a nonseparable
way. More generally, when instruments arise from equilibrium conditions
of structural models, separability of the first stage is the exception
rather than the rule.
Our contributions are threefold. First, we show that $g$ has a closed-form
expression: it is identified as a functional of the conditional distribution
functions $F_{Y|X,Z}$ and $F_{X|Z}$. Second, replacing those two
functions by kernel estimators gives a plug-in estimator of $g$ that
needs no regularization and converges at the rate $n^{-m/(2m+1)}$
at a point, and is asymptotically normal at a slightly smaller bandwidth.
We also show that no estimator of $g$ converges faster, so the nonseparable
first stage does not lower the rate: it is the rate that would be
available if the endogeneity were absent. Third, we use the empirical
bootstrap to construct a uniform confidence band that covers $g$
at every point of a compact subset of the interior of the support
of $X$ simultaneously, the subset omitting the normalization point.
In the literature on identification and estimation of nonparametric
structural models, a closely related work is presented by \citet{NPV99-ECTA}.
Their estimation method relies on the additive separability of first-stage
equation $X=h\left(Z\right)+\eta$ under the mean independence condition
$\mathbb{E}\left[\varepsilon|\eta,Z\right]=\mathbb{E}\left[\varepsilon|\eta\right]$.
Thus their results are not applicable to (\ref{eq:model}). Their
estimator attains the slower of the two rates that \citet{Sto82-AoS}
shows to be optimal for the second stage and for the reduced form
separately. In model (\ref{eq:model}) both stages are univariate
and the same $m$ derivatives are assumed of each, so those two rates
coincide: theirs is $n^{-2m/(2m+1)}$ in mean square, whose square
root is the rate our estimator attains at a point. Separability of
the first stage therefore does not improve the rate at which $g$
is estimated.\footnote{Assumption \ref{assu:rate} allows a nonempty bandwidth window only
when the model has $m\ge4$ derivatives and the kernel is of order
$m$. The comparison therefore holds at $m\ge4$. Below that no bandwidth
sequence meets the assumption.} Fully nonseparable systems have also been investigated \citep{Che03-ECTA,IN09-ECTA,DF15-ECTA,Tor15-ECTA,Ish21-ET}.
Notably, \citet{IN09-ECTA} focus on quantile, average, and policy
effects rather than the direct estimation of $g$. \citet{DF15-ECTA}
and \citet{Tor15-ECTA} provide sufficient conditions for point-identification
of the nonseparable function $g(x,w,e)$ using binary instruments.
However, they do not offer clear guidance for estimation. Additionally,
\citet[Theorem S1]{Tor15-ECTA} establishes point-identification of
the function $g$ using continuous $Z$ in a non-constructive manner.
Since the model (\ref{eq:model}) emerges as a special case of the
model in \citet[Theorem S1]{Tor15-ECTA} under a level-normalization,
our first contribution lies not in establishing the identification
of $g$ but in providing the closed-form expression for $g$. \citet{torgovitsky2017minimum}
provides an estimation method for nonseparable triangular systems,
but he restricts his attention to the case that the function $g(x,w,e)$
is unknown up to a finite dimensional parameter.
Our identification result is close in spirit to those obtained by
\citet{Che03-ECTA} and \citet{CKK15-JoE}. \citet{Che03-ECTA} provides
the identification of the derivative of $g$ under a nonseparable
triangular system, which shares similarities with the form presented
in Proposition \ref{prop:identification} below.
Our estimator is more directly indebted to \citet{CKK15-JoE}. They
study nonparametric transformation models, a different model identified
under different conditions, but the estimator they build from their
identification result is one we can reuse. The difference is that
their identification result gives $g$ itself while ours gives $g^{\prime}$,
so that $g(x)$ is recovered by integrating from 0 to $x$. The endpoint
of that integral is not averaged away, and Proposition \ref{prop:lower_bound}
shows that no estimator of $g(x)$ reaches the parametric rate. Their
estimator is asymptotically linear, and ours cannot be, since asymptotic
linearity would give the parametric rate. The two asymptotic arguments
therefore share little beyond the averaging step. What is new here
is the identification result that supplies the expression to average,
and the bootstrap and the band of Section \ref{sec:Estimation-and-bootstrap}.
In the absence of assistance from the first-stage equations, one can
estimate $g$ by solving an integral equation $\mathbb{E}\left[Y|Z\right]=\mathbb{E}\left[g(X)|Z\right]$
with respect to $g$ under the mean independence condition $\mathbb{E}\left[\varepsilon|Z\right]=0$.
Given a well-known completeness condition, there exists a unique solution
to this problem. However, this approach necessitates a regularization
scheme to handle the ill-posed inverse problem, which leads to a slower
convergence rate of estimators in general \citep{NP03-ECTA,BCK07-ECTA,CFR07-Handbook,DFFR11-ECTA,Hor14-AnnRevEcon}.
Uniform confidence bands are available in that literature: \citet{HL12-JoE}
construct them from bootstrap critical values, \citet{CC18-QE} obtain
them for nonlinear functionals of a sieve estimator, and \citet{CCK25-ReStud}
choose the sieve dimension adaptively; see also \citet{Babii20-ET}.
All of them inherit the ill-posedness of the inverse problem, so the
bands rest on estimators whose rate it degrades; the band of Section
\ref{sec:Estimation-and-bootstrap} does not.
The remainder of the paper is organized as follows. Section \ref{sec:Identification}
provides the identification result. Section \ref{sec:Estimation-and-bootstrap}
proposes nonparametric estimators based on this identification result,
and studies their asymptotic behaviors. Section \ref{sec:Monte-carlo}
reports a Monte Carlo study. All proofs are deferred to the Appendix.
\section{Identification }\label{sec:Identification}
Let $\mathcal{A}$ be the support of random variable $A$, and $\mathcal{A}_{b}$
the conditional support of $A$ given $B=b$. Let $\mathcal{A}^{\circ}$
be the interior of $\mathcal{A}$. We assume that $\mathcal{Y},\mathcal{X},\mathcal{Z}\subset\mathbb{R}$.
We consider the model in (\ref{eq:model}) and impose the following
assumptions:
\begin{assumption}
\label{assu:rv}(i) The random variables $Y|X=x,Z=z$, $X|Z=z$, $Z$
are absolutely continuously distributed for each $x$, $z$. (ii)
$\mathcal{X}$ is convex.
\end{assumption}
This assumption is imposed by \citet[Assumption C]{Tor15-ECTA}. The
first condition requires $Y$, $X$ and $Z$ to be continuously distributed.
It covers $\eta$ as well, though it does not name it: $h(z,\cdot)$
is one-to-one under Assumption \ref{assu:IV}(ii), so an atom of $\eta$
at $a$ would put one of $X$ given $Z=z$ at $h(z,a)$. The proof
of Proposition \ref{prop:identification} uses $V(\cdot|z)$, which
is continuous for that reason.
\begin{assumption}
\label{assu: func_g} (i) $g$ is continuously differentiable on $\mathcal{X}$,
the derivative being one-sided at a boundary point. (ii) $0\in\mathcal{X}^{\circ}$
such that $g(0)=0$.
\end{assumption}
The first condition is imposed by \citet[Assumption G]{Tor15-ECTA}.
The second condition is a level normalization, which requires that
the constant be zero when the model is linear. This type of normalization
is also necessary for identification in nonseparable models.
\begin{assumption}
\label{assu:IV} (i) The instrument is independent of the unobservables:
$\left(\varepsilon,\eta\right)\mathop{\perp \! \! \! \perp} Z$. (ii) The function $h(z,\cdot)$
is strictly increasing for each $z$.
\end{assumption}
The first condition requires a full independence between instruments
and error terms, which is imposed by \citet{IN09-ECTA}, \citet{Tor15-ECTA},
and \citet{DF15-ECTA}. \citet{NPV99-ECTA} impose a weaker mean independence
condition $\mathbb{E}\left[\varepsilon|Z,\eta\right]=\mathbb{E}\left[\varepsilon|\eta\right]$.
Our identification result exploits the full independence, and cannot
be weakened to this mean independence. The monotonicity imposed in
the second condition is standard in the literature.
\begin{assumption}
\label{assu: local} Let $V(x|z)=F_{X|Z}(x|z)$. For every $x$ in
a dense subset $\mathcal{X}_{d}$ of $\mathcal{X}$ and $y\in\mathcal{Y}^{\circ}_{x}$,
there exists $z\in\mathcal{Z}^{\circ}$ such that (i) $y\in\mathcal{Y}^{\circ}_{x,z}$
and $x\in\mathcal{X}^{\circ}_{z^{\prime}}$ for every $z^{\prime}$
in a neighborhood of $z$, (ii) $V(x|\cdot)$ is continuously differentiable
on a neighborhood of $z$ with $\nabla_{z}V(x|z)\neq0$, and $\nabla_{x}V(x|z)$
exists, (iii) $\nabla_{y}F_{Y|X,Z}(\cdot|x,\cdot)$ and $\nabla_{z}F_{Y|X,Z}(\cdot|x,\cdot)$
exist and are continuous on a neighborhood of $(y,z)$, and $\nabla_{x}F_{Y|X,Z}(y|x,z)$
exists; write $f_{Y|X,Z}(y|x,z):=\nabla_{y}F_{Y|X,Z}(y|x,z)$. (iv)
$f_{Y|X,Z}(y|x,z)>0$.
\end{assumption}
The first and second conditions strengthen those imposed by \citet[Theorem S1]{Tor15-ECTA}.
The third and fourth are additionally needed. Conditional distributions
are determined only up to null sets, while the four conditions above
and the conclusion of Proposition \ref{prop:identification} are statements
at a point. Throughout, $F_{Y|X,Z}$ and $V$ denote the versions
continuous in their conditioning arguments, which is what parts (ii)
and (iii) are conditions on. Assumption \ref{assu: density} asks
the same on the region where estimation takes place. The conditions
are placed on a neighborhood of $z$ rather than at $z$ alone because
the proof inverts $z\mapsto V(x|z)$ near $z$ and then differentiates
a composition in which $x$ enters through both arguments; the two
partial derivatives at the point would not suffice for that.
\begin{assumption}
\label{assu:support} The supports satisfy $\mathcal{Y}_{x,z}=\mathcal{Y}$
and $\mathcal{X}_{z}=\mathcal{X}$ for each $x$ and $z$.
\end{assumption}
Proposition \ref{prop:identification} itself does not need this assumption.
Under it the support conditions of Assumption \ref{assu: local} (i)
hold at every $(x,y,z)$ with $x\in\mathcal{X}^{\circ}$ and $y\in\mathcal{Y}^{\circ}$,
so the identified expression holds at every $(y,z)$ at which the
derivatives in parts (ii) and (iii) of that assumption exist and $\nabla_{z}V(x|z)$
and $f_{Y|X,Z}(y|x,z)$ are nonzero. Assumption \ref{assu: density}
below imposes those conditions on the region over which the estimators
of Section \ref{sec:Estimation-and-bootstrap} average. Averaging
over $\left(y,z\right)$ reduces the asymptotic variance.
Our first result is stated as follows:
\begin{prop}
\label{prop:identification}Suppose Assumptions \ref{assu:rv}--\ref{assu:IV}
hold. Let $(x,y,z)\in\mathcal{X}\times\mathcal{Y}\times\mathcal{Z}$
satisfy conditions (i) to (iii) of Assumption \ref{assu: local}.
Then
\begin{equation}
\nabla_{x}g(x)f_{Y|X,Z}(y|x,z)=\frac{\nabla_{x}V(x|z)}{\nabla_{z}V(x|z)}\nabla_{z}F_{Y|X,Z}(y|x,z)-\nabla_{x}F_{Y|X,Z}(y|x,z).\label{eq:identification_prod}
\end{equation}
If condition (iv) also holds, the derivative of the outcome function
$g$ is identified at $x$ as
\begin{equation}
\nabla_{x}g(x)=\frac{\frac{\nabla_{x}V(x|z)}{\nabla_{z}V(x|z)}\nabla_{z}F_{Y|X,Z}(y|x,z)-\nabla_{x}F_{Y|X,Z}(y|x,z)}{f_{Y|X,Z}(y|x,z)}.\label{eq:identification}
\end{equation}
\end{prop}
\begin{proof}
See Appendix \ref{sec:lemma1}.
\end{proof}
\begin{rem}
If Assumption \ref{assu: local} holds, such a $(y,z)$ exists for
every $x\in\mathcal{X}_{d}$. Since $\nabla_{x}g$ is continuous and
$\mathcal{X}_{d}$ is dense in $\mathcal{X}$, the derivative is determined
on all of $\mathcal{X}$ by its values on $\mathcal{X}_{d}$. Determination
is not the same as an expression: (\ref{eq:identification}) holds
only where conditions (i) to (iv) hold, and condition (i) puts those
$x$ in $\mathcal{X}^{\circ}$, so it is never available at the edge
of the support. $\blacksquare$
\end{rem}
\begin{rem}
Under Assumptions \ref{assu:rv}--\ref{assu:IV} and parts (i) and
(ii) of Assumption \ref{assu: local}, the additively separable model
considered in this setting is a special case of the nonseparable models
studied by \citet{Tor15-ECTA}. Since \citet[Theorem S1]{Tor15-ECTA}
provides the identification of $g(x,e)$, the identification of $\nabla_{x}g(x)$
is also known where his conditions hold. However, his identification
of $g(x,e)$ is not constructive. Thus the contribution of Proposition
\ref{prop:identification} lies in providing the closed-form expression.
$\blacksquare$
\end{rem}
\begin{rem}
Model (\ref{eq:model}) has no covariates. Discrete covariates $W$
need no change: everything below applies within each cell $\{W=w\}$,
so $g(\cdot,w)$ is identified and estimated cell by cell, with the
rate governed by the cell sizes. Continuous covariates are a different
matter, since they enter the kernel estimators and raise their dimension;
the rate falls with each one added. $\blacksquare$
\end{rem}
As a corollary of Proposition \ref{prop:identification}, we have
$g(x)=S(y,x,z)$ where
\begin{equation}
S(y,x,z)=\int^{x}_{0}\frac{\frac{\nabla_{x}V(u|z)}{\nabla_{z}V(u|z)}\nabla_{z}F_{Y|X,Z}(y|u,z)-\nabla_{x}F_{Y|X,Z}(y|u,z)}{f_{Y|X,Z}(y|u,z)}\mathrm{d}u\label{eq:g_functional}
\end{equation}
at every $(y,z)$ such that conditions (i) to (iv) of Assumption \ref{assu: local}
hold at $(u,y,z)$ for every $u$ between $0$ and $x$. Assumption
\ref{assu:support} makes the supports invariant, and Assumption \ref{assu: density}
of Section \ref{sec:Estimation-and-bootstrap} places the remaining
conditions on the region over which the estimators average. The step
from $\nabla_{u}g$ to $g$ uses Assumption \ref{assu:rv}(ii), so
that the segment from $0$ to $x$ lies in $\mathcal{X}$, and Assumption
\ref{assu: func_g}, which makes the integrand continuous and fixes
the level. Note that the closed form holds at any triple satisfying
conditions (i) to (iv), whether or not $u$ lies in $\mathcal{X}_{d}$;
the dense subset enters Assumption \ref{assu: local} only through
the existence of a suitable $z$. We shall propose nonparametric estimators
based on this expression.
\section{Estimation and bootstrap inference }\label{sec:Estimation-and-bootstrap}
\subsection{Estimation}
Suppose we have a random sample $\left\{ Y_{i},X_{i},Z_{i}\right\} ^{n}_{i=1}$
drawn from the model. Let $\widehat{S}(y,x,z)$ be a preliminary nonparametric
estimator of $S(y,x,z)$. We follow \citet{CKK15-JoE} and consider
two estimators. The first one is a least-squares type estimator:
\[
\widehat{g}_{\mathrm{LS}}(x):=\int\int\mu(y,z)\widehat{S}(y,x,z)\mathrm{d}y\mathrm{d}z,
\]
where $\mu(y,z)$ is a weighting function. The second one is a smoothed
version of the least absolute deviation (LAD) type estimator by \citet{horowitz1998bootstrap}:
\begin{align*}
\hat{g}_{\mathrm{LAD}}(x) & =\arg\min_{q}Q_{b}(q|\hat{S}(\cdot,x,\cdot)),\\
Q_{b}(q|\hat{S}(\cdot,x,\cdot)) & =\int\int\mu(y,z)\left\{ \hat{S}(y,x,z)-q\right\} \left\{ 2\Lambda_{b}\left[\hat{S}(y,x,z)-q\right]-1\right\} \mathrm{d}y\mathrm{d}z,
\end{align*}
with $\Lambda_{b}:=\Lambda(\cdot/b)$ for a known distribution function
$\Lambda$ with median zero and some bandwidth $b>0$. $\hat{g}_{\mathrm{LAD}}$
is expected to be more robust to outliers in the sample than $\widehat{g}_{\mathrm{LS}}$.
\begin{rem}
\label{rmk:mu} Under Assumptions \ref{assu:support} and \ref{assu: density},
conditions (i) to (iv) of Assumption \ref{assu: local} hold at $(u,y,z)$
for every $y\in\mathcal{Y}_{\mu}$, $z\in\mathcal{Z}_{\mu}$ and $u\in\widetilde{\mathcal{X}}_{0}$,
so $S(y,x,z)=g(x)$ at every $(y,z)\in\mathcal{Y}_{\mu}\times\mathcal{Z}_{\mu}$
and every $x\in\mathcal{X}_{0}$. With $\int\int\mu=1$ from Assumption
\ref{assu: density}(ii) this gives $\int\int\mu(y,z)S(y,x,z)\mathrm{d}y\mathrm{d}z=g(x)$.
The estimand does not depend on $\mu$, which enters only through
the variance, much as the weight matrix does in generalized method
of moments. The same holds for $\hat{g}_{\mathrm{LAD}}$: the population
objective is $\varrho(g(x)-q)$ for $\varrho(t)=t\{2\Lambda_{b}(t)-1\}$,
which Assumption \ref{assu: kernel}(ii) makes uniquely minimized
at $q=g(x)$, whatever the weight. $\blacksquare$
\end{rem}
To estimate $S(y,x,z)$ nonparametrically, we use a kernel-based estimator.
First, observe that
\[
F_{Y\mid X,Z}(y|x,z)=\frac{\int^{y}_{-\infty}f_{Y,X,Z}(u,x,z)du}{f_{X,Z}(x,z)}=:\frac{\Phi(y,x,z)}{f(x,z)}.
\]
Then a natural estimator of the conditional distribution $F_{Y\mid X,Z}(y|x,z)$
is $\widehat{F}(y|x,z)=\widehat{\Phi}(y,x,z)/\widehat{f}(x,z)$ where
\begin{align*}
\widehat{\Phi}(y,x,z) & =\frac{1}{n}\sum^{n}_{i=1}\frac{1}{h_{x}h_{z}}\boldsymbol{1}\left\{ Y_{i}\le y\right\} K\left(\frac{X_{i}-x}{h_{x}}\right)K\left(\frac{Z_{i}-z}{h_{z}}\right),\\
\widehat{f}(x,z) & =\frac{1}{n}\sum^{n}_{i=1}\frac{1}{h_{x}h_{z}}K\left(\frac{X_{i}-x}{h_{x}}\right)K\left(\frac{Z_{i}-z}{h_{z}}\right).
\end{align*}
As an estimator for the density $\frac{\partial}{\partial y}F_{Y|X,Z}=f_{Y|X,Z}$,
we consider $\widehat{\Phi}_{y}(y,x,z)/\widehat{f}(x,z)$ where
\[
\widehat{\Phi}_{y}(y,x,z)=\frac{1}{n}\sum^{n}_{i=1}\frac{1}{h_{y}h_{x}h_{z}}K\left(\frac{Y_{i}-y}{h_{y}}\right)K\left(\frac{X_{i}-x}{h_{x}}\right)K\left(\frac{Z_{i}-z}{h_{z}}\right).
\]
A kernel-based estimator of the conditional distribution $V(x|z)$
is obtained by $\widehat{V}(x|z)=\widehat{\Psi}(x,z)/\widehat{p}(z)$
where
\[
\widehat{\Psi}(x,z)=\frac{1}{n}\sum^{n}_{i=1}\frac{1}{h_{z}}\boldsymbol{1}\left\{ X_{i}\le x\right\} K\left(\frac{Z_{i}-z}{h_{z}}\right),\qquad\widehat{p}(z)=\frac{1}{n}\sum^{n}_{i=1}\frac{1}{h_{z}}K\left(\frac{Z_{i}-z}{h_{z}}\right).
\]
Finally a kernel based estimator of $S(y,x,z)$ is obtained by
\begin{align*}
\widehat{S}(y,x,z) & =\int^{x}_{0}\frac{\frac{\nabla_{x}\widehat{V}(u|z)}{\nabla_{z}\widehat{V}(u|z)}\nabla_{z}\widehat{F}(y|u,z)-\nabla_{x}\widehat{F}(y|u,z)}{\nabla_{y}\widehat{F}(y|u,z)}\mathrm{d}u.
\end{align*}
We employ the following assumptions on the model, the kernel function,
and the weight function.
\begin{assumption}
\label{assu: kernel} (i) The univariate kernel $K$ is differentiable,
and there exist constants $C>0$ and $\nu>m+1$ such that $\left|K^{(i)}(z)\right|\le C\left|z\right|^{-\nu}$,
$\left|K^{(i)}(z)-K^{(i)}(z^{\prime})\right|\le C\left|z-z^{\prime}\right|$,
and $K^{(i)}$ is of bounded variation, for $i=0,1$, where $K^{(i)}(z)$
denotes $i$-th derivative of $K$. Furthermore, there exists $m\ge2$
such that $\int_{\mathbb{R}}K(z)\mathrm{d}z=1$, $\int_{\mathbb{R}}z^{j}K(z)\mathrm{d}z=0$
for $1\le j\le m-1$, and $\int_{\mathbb{R}}\left|z\right|^{m}\left|K(z)\right|\mathrm{d}z<\infty$.
(ii) The distribution function $\Lambda$ is strictly increasing and
three times continuously differentiable with bounded derivatives,
its density satisfies $\Lambda^{\prime}(0)>0$, and the bandwidth
$b>0$ for $\Lambda_{b}=\Lambda(\cdot/b)$ is held fixed.
\end{assumption}
The differentiability condition is required to ensure the existence
of derivatives of $\widehat{F}(y|x,z)$ and $\widehat{V}(x|z)$. Hence
we rule out uniform and Epanechnikov kernels. Part (ii) concerns $\hat{g}_{\mathrm{LAD}}$
alone. The condition $\Lambda^{\prime}(0)>0$ makes the smoothed objective
behave like an absolute-deviation criterion at its minimum, and it
is the only feature of $\Lambda$ that survives into the first-order
asymptotics: as the end of Appendix \ref{sec:thm1} shows, $\Lambda^{\prime}(0)/b$
enters the score and the Hessian in the same way and cancels between
them, so $b$ does not appear in the first-order asymptotics and need
not shrink with $n$. \citet{CKK15-JoE} hold $b$ fixed for the same
reason. Strict monotonicity is used elsewhere: with median zero it
makes the population objective in Appendix \ref{sec:thm1} uniquely
minimized at $g(x)$. The standard normal distribution used in Section
\ref{sec:Monte-carlo} satisfies both conditions.
\begin{assumption}
\label{assu: density} Let $\mathcal{X}_{0}\subset\mathcal{X}^{\circ}$
be a compact set with $0\notin\mathcal{X}_{0}$, let $\widetilde{\mathcal{X}}_{0}$
be the smallest interval containing $\mathcal{X}_{0}\cup\{0\}$, and
write $\mathcal{Y}_{\mu}\times\mathcal{Z}_{\mu}$ for the support
of $\mu$. (i) The joint density $f_{Y,X,Z}(y,x,z)$ is bounded and
$m$-times differentiable with bounded derivatives. There are open
sets $\mathcal{U}\supset\widetilde{\mathcal{X}}_{0}$ and $\mathcal{W}\supset\mathcal{Z}_{\mu}$
whose closures are compact subsets of $\mathcal{X}^{\circ}$ and $\mathcal{Z}^{\circ}$
such that $\int\sup_{(x,z)\in\mathcal{U}\times\mathcal{W}}\left|\partial^{j_{1}}_{x}\partial^{j_{2}}_{z}f_{Y,X,Z}(y,x,z)\right|\mathrm{d}y<\infty$
for $j_{1}+j_{2}\le m$. (ii) The weight function $\mu(y,z)$ is $m$-times
continuously differentiable for both arguments with compact support
$\mathcal{Y}_{\mu}\times\mathcal{Z}_{\mu}$ contained in $\mathcal{Y}^{\circ}\times\mathcal{Z}^{\circ}$
and with nonempty interior, and $\int\int\mu(y,z)\mathrm{d}y\mathrm{d}z=1$.
(iii) On a neighborhood of $\mathcal{Y}_{\mu}\times\widetilde{\mathcal{X}}_{0}\times\mathcal{Z}_{\mu}$
the derivatives $\nabla_{x}V$, $\nabla_{z}V$, $\nabla_{x}F_{Y|X,Z}$,
$\nabla_{y}F_{Y|X,Z}$ and $\nabla_{z}F_{Y|X,Z}$ exist and are continuous,
and the denominators are bounded away from zero: $\inf f(x,z)>0$,
$\inf f_{Y|X,Z}(y|x,z)>0$, $\inf\left|\nabla_{z}V(x|z)\right|>0$
and $\inf p(z)>0$, the infima taken over $(y,x,z)\in\mathcal{Y}_{\mu}\times\widetilde{\mathcal{X}}_{0}\times\mathcal{Z}_{\mu}$.
\end{assumption}
The integrability in (i) allows differentiating under the integral
sign, which the proofs of Lemma \ref{lem:boundary} and Lemma \ref{lem:plugin}
both do. Part (iii) makes conditions (ii) to (iv) of Assumption \ref{assu: local}
hold at every point of the region, which the discussion of Assumption
\ref{assu:support} defers to it. The tail exponent is taken above
$m+1$. It makes $\int K^{2}$, $\int\left|v\right|K(v)^{2}\mathrm{d}v$,
$\int\left|K\right|^{3}$ and $\int\left|K\right|^{4}$ finite, the
first two appearing in the asymptotic variance and in its distance
from the variance at a finite $n$, and it bounds the kernel away
from its centre wherever Appendix \ref{sec:thm1} needs that. The
kernel is not assumed to have compact support, so a remainder is left
from outside the region on which the density is smooth, and $\nu\ge m+1$
keeps it within the $O(h^{m}_{x})$ of the smoothing bias. Note that
the number of derivatives $m$ must align with the order of the kernel
function. Part (iii) collects the denominators. The estimators below
are ratios: $\widehat{F}=\widehat{\Phi}/\widehat{f}$ and $\widehat{V}=\widehat{\Psi}/\widehat{p}$
are ratios of kernel averages, and the integrand of (\ref{eq:g_functional})
divides in turn by $f_{Y|X,Z}$ and by $\nabla_{z}V$. A lower bound
on each turns uniform convergence of the kernel estimators into uniform
convergence of the ratios, on $\mathcal{Y}_{\mu}\times\widetilde{\mathcal{X}}_{0}\times\mathcal{Z}_{\mu}$.
The domain is $\widetilde{\mathcal{X}}_{0}$ rather than $\mathcal{X}_{0}$
because (\ref{eq:g_functional}) integrates from the normalization
point to $x$: estimating $g$ at a point of $\mathcal{X}_{0}$ uses
the integrand at every $u$ between $0$ and $x$, and those $u$
need not lie in $\mathcal{X}_{0}$. Assumption \ref{assu:rv} (ii)
makes $\mathcal{X}$ convex and Assumption \ref{assu: func_g} (ii)
places the normalization point in its interior, so $\widetilde{\mathcal{X}}_{0}$
is again a compact subset of $\mathcal{X}^{\circ}$. The restriction
to a compact $\mathcal{X}_{0}$ strictly inside $\mathcal{X}$ is
not a technicality: at the edge of the support of $X$ the bounds
on $\left|\nabla_{z}V\right|$ and on $f$ both fail.
\begin{assumption}
\label{assu:rate}$h_{v}=c_{v}n^{-a_{v}}\ (a_{v}>0)\text{ for }v\in\{y,x,z\}$,
$\sqrt{nh_{x}}h^{m}_{y}\to0$, $\sqrt{nh_{x}}h^{m}_{z}\to0$, $\log n/(\sqrt{n}h_{y}h^{1/2}_{x}h_{z})\to0$,
$\log n/(\sqrt{n}h^{5/2}_{x}h_{z})\to0$, and $\log n/(\sqrt{n}h^{1/2}_{x}h^{3}_{z})\to0$.
\end{assumption}
Write $\bar{h}:=\max\{h_{y},h_{x},h_{z}\}$ for the largest of the
three. The first restriction fixes the bandwidths at powers of $n$.
Section \ref{sec:Monte-carlo} uses that, and it makes $\log(1/h_{x})$
of order $\log n$, which the entropy bounds of Appendix \ref{sec:proof_band}
need. The next two make the smoothing bias of $(\hat{\Phi},\hat{\Psi},\hat{f},\hat{p})$
and of their derivatives $o((nh_{x})^{-1/2})$. Nothing is imposed
on $h^{m}_{x}$: the bias in the direction of $x$ is the estimator's
own rather than a nuisance error, and Theorems \ref{thm:pointwise}
and \ref{thm:band} control it by undersmoothing. The last three conditions
make the squares and products of the estimation errors $\widehat{\Phi}-\Phi,\widehat{\Psi}-\Psi,\widehat{f}-f,\widehat{p}-p$
and of the errors in their derivatives $o((nh_{x})^{-1/2})$. Squares
and products are the form in which these errors enter: $\widehat{S}$
is built from the ratios $\widehat{\Phi}/\widehat{f}$ and $\widehat{\Psi}/\widehat{p}$,
and expanding a ratio leaves a remainder quadratic in them. No condition
is placed on the errors themselves, which are of order $\sqrt{\log n/(nh_{y}h_{x}h_{z})}$
or larger and so are not $o_{p}((nh_{x})^{-1/2})$. They also enter
linearly, and the linear terms produce the limit distribution of Theorem
\ref{thm:pointwise}. Write $a_{n}\asymp b_{n}$ when $a_{n}/b_{n}$
is bounded above and away from zero. With $h_{y}\asymp h_{x}\asymp h_{z}\asymp n^{-a}$
these conditions, together with the undersmoothing that Theorems \ref{thm:pointwise}
and \ref{thm:band} add, leave $a\in(1/(2m+1),1/7)$, which is nonempty
once $m\ge4$.\footnote{With $h_{y}\asymp h_{x}\asymp h_{z}\asymp n^{-a}$ the last three
conditions are $n^{-1/2+5a/2}$, $n^{-1/2+7a/2}$ and $n^{-1/2+7a/2}$
up to the logarithm, so the last two bind and give $a<1/7$; $\sqrt{nh_{x}}h^{m}_{x}\rightarrow0$
gives $a>1/(2m+1)$.} \footnote{The three exponents need not agree, and the conditions bind on them
differently: $h_{x}$ enters the second with a high power because
two of the smoothed quantities, $\nabla_{x}\widehat{\Phi}$ and $\nabla_{x}\widehat{f}$,
are derivatives in $x$, while $h_{y}$ and $h_{z}$ are constrained
mainly through the bias. With $a_{x}=0.15$, for instance, the conditions
hold for $a_{y},a_{z}\in(0.106,0.125)$ at $m=4$. The lower endpoint
is the exponent at which the estimator attains the optimal rate, and
it is excluded: inference requires undersmoothing, and with it a rate
arbitrarily close to the optimum rather than equal to it. Section
\ref{sec:Monte-carlo} uses a fourth-order kernel.}
The following gives the limit distribution at a point for the nonparametric
estimators $\hat{g}_{\mathrm{LS}}$ and $\hat{g}_{\mathrm{LAD}}$.
\begin{thm}
\label{thm:pointwise} Suppose that Assumptions \ref{assu:rv}--\ref{assu:support}
and \ref{assu: kernel}--\ref{assu:rate} hold, that $\sigma^{2}(0)>0$,
and that the bandwidth undersmooths in the sense that $\sqrt{nh_{x}}h^{m}_{x}\rightarrow0$.
Then for each $x\in\mathcal{X}_{0}$
\[
\sqrt{nh_{x}}\left(\widehat{g}_{\mathrm{LS}}(x)-g(x)\right)\rightsquigarrow N\left(0,\sigma^{2}(x)+\sigma^{2}(0)\right),
\]
where $\sigma^{2}(a)$ is defined below. The same limit holds for
$\widehat{g}_{\mathrm{LAD}}(x)$.
\end{thm}
\begin{proof}
See Appendix \ref{sec:thm1}.
\end{proof}
Lemma \ref{lem:boundary} decomposes the deviation as
\[
\widehat{g}_{\mathrm{LS}}(x)-g(x)=T_{n}(x)-T_{n}(0)+o_{p}((nh_{x})^{-1/2}),
\]
where
\begin{align*}
T_{n}(a) & =\frac{1}{nh_{x}}\sum^{n}_{i=1}K\left(\frac{X_{i}-a}{h_{x}}\right)c_{a}(W_{i}),\\
c_{a}(W_{i}) & =\int\frac{\mu(y,Z_{i})\left[F_{Y|X,Z}(y|a,Z_{i})-\boldsymbol{1}\left\{ Y_{i}\le y\right\} \right]}{f_{Y|X,Z}(y|a,Z_{i})f(a,Z_{i})}\mathrm{d}y,\\
\sigma^{2}(a) & =\left[\int K(v)^{2}\mathrm{d}v\right]f_{X}(a)\mathbb{E}\left[c_{a}(W_{i})^{2}|X_{i}=a\right].
\end{align*}
Each $T_{n}(a)$ is a kernel average in the single argument $x$,
centered at $a$, and $\sigma^{2}(a)$ is the limit of $nh_{x}$ times
its variance. The term at $a=0$ is present because the estimator
integrates from the normalization point. Assumption \ref{assu: func_g}(ii)
sets $g(0)=0$, $\widehat{S}$ integrates from $0$ to $x$, and the
error made at the lower endpoint is the same at every $x$. For $x$
bounded away from the origin the two averages are asymptotically uncorrelated,
so their variances add. Section \ref{sec:Monte-carlo} measures the
component that the evaluation points share.
We need to exclude the normalization point from the evaluation set,
though not because anything degenerates there. Assumption \ref{assu: func_g}
(ii) sets $g(0)=0$, so $\hat{g}(0)=0$ by construction and the deviation
at that point is identically zero. Away from it, the two terms Lemma
\ref{lem:boundary} leaves are kernel averages centered at $x$ and
at $0$, and their covariance is $O\big((h_{x}/\left|x\right|)^{\nu}\big)$
after rescaling, which makes their variances add to $\sigma^{2}(x)+\sigma^{2}(0)$.
For that to hold uniformly, $\mathcal{X}_{0}$ has to be bounded away
from the origin. Theorem \ref{thm:pointwise} and the band of Theorem
\ref{thm:band} are therefore available on any compact set omitting
a neighborhood of the normalization point, but not on an interval
straddling it; in Section \ref{sec:Monte-carlo}, $\mathcal{X}_{0}$
is a union of two intervals bounded away from the origin.
The rate is $(nh_{x})^{-1/2}$ rather than $n^{-1/2}$ because $\widehat{S}$
differentiates in the direction it integrates over. A kernel in $x$
that stays inside $\int^{x}_{0}\mathrm{d}u$ is averaged by it and
contributes at $n^{-1/2}$; but $\nabla_{x}\widehat{V}$ and $\nabla_{x}\widehat{F}$
are derivatives in $u$, and integrating them by parts returns the
kernel at $u=0$ and $u=x$, where nothing averages it. $T_{n}(x)$
and $T_{n}(0)$ are those two terms, and a kernel at a point converges
at $(nh_{x})^{-1/2}$. Only $h_{x}$ appears in the limit, although
$\widehat{F}$ smooths in three arguments and $\widehat{V}$ in two.
The estimator integrates $\widehat{S}$ against $\mu$ in $y$ and
$z$, so those two directions are averaged rather than evaluated at
a point, and $h_{y}$ and $h_{z}$ enter through the bias and the
remainder alone. The leading term is therefore a kernel average in
$x$ alone, and the rate is the one-dimensional rate.
The rate $n^{-m/(2m+1)}$ and the limit distribution of Theorem \ref{thm:pointwise}
are attained at different bandwidths. At $h_{x}\asymp n^{-1/(2m+1)}$
the two terms in the bound $|\widehat{g}_{\mathrm{LS}}(x)-g(x)|=O_{p}(h^{m}_{x}+(nh_{x})^{-1/2})$
of Lemma \ref{lem:boundary} are of the same order, and the bound
is $O_{p}(n^{-m/(2m+1)})$. At that bandwidth the estimator attains
the rate Proposition \ref{prop:lower_bound} shows below cannot be
improved. Note that Theorem \ref{thm:pointwise} holds at smaller
bandwidths, those satisfying $\sqrt{nh_{x}}h^{m}_{x}\rightarrow0$.
That condition removes the bias from the limit distribution, and it
excludes $h_{x}\asymp n^{-1/(2m+1)}$. Appendix \ref{sec:thm1} is
written under the conditions of Theorem \ref{thm:pointwise}, so Lemma
\ref{lem:boundary} carries that condition too. The bound does not
use it. Parts (ii) to (iv) retain the bias terms rather than discarding
them, and the undersmoothing condition enters their proofs only in
remarks that a retained term is negligible, so they hold under Assumptions
\ref{assu:rv}--\ref{assu:support} and \ref{assu: kernel}--\ref{assu:rate}
alone. Thus, the bandwidths it allows give a rate slower than $n^{-m/(2m+1)}$.
Taking $a_{x}$ arbitrarily close to $1/(2m+1)$ makes the difference
as small as desired.
The next result bounds what any estimator of $g(x)$ can achieve,
using a result of \citet{Sto80-AoS}.
\begin{prop}
\label{prop:lower_bound} Let $\mathcal{P}$ be the set of distributions
of $\left(Y,X,Z\right)$ that (\ref{eq:model}) generates under Assumptions
\ref{assu:rv}--\ref{assu:support}, with the joint density bounded
and $m$ times differentiable with bounded derivatives, that bound
being held fixed across the set. Fix $x\in\mathcal{X}^{\circ}$ with
$x\neq0$. Then there is a $c>0$ such that
\[
\liminf_{n\rightarrow\infty}\inf_{\tilde{g}_{n}}\sup_{P\in\mathcal{P}}P^{n}\left(\left|\tilde{g}_{n}(x)-g_{P}(x)\right|\ge cn^{-m/(2m+1)}\right)>0,
\]
where the infimum runs over all estimators of $g(x)$ based on a sample
of size $n$.
\end{prop}
\begin{proof}
See Appendix \ref{sec:lower}.
\end{proof}
Imposing that the endogeneity is absent makes the problem easier,
so the best rate available under that restriction limits what any
estimator can achieve without it. \citet[p.~576]{NPV99-ECTA} argue
in the same way for their own estimator, comparing instead with the
problem in which the control function is known rather than estimated.
The proof restricts to the subset $\mathcal{P}_{0}$ of distributions
on which $\varepsilon$ is independent of $\left(\eta,Z\right)$.
The step that is not immediate is that $\mathcal{P}_{0}$ is still
inside $\mathcal{P}$. The restriction makes $F_{Y|X,Z}(y|x,z)$ free
of $z$, so $\nabla_{z}F_{Y|X,Z}(y|x,z)$ vanishes. The proof checks
Assumption \ref{assu: local} to confirm that a model with that property
still satisfies the assumptions that define $\mathcal{P}$. The proof
fixes the first stage at one admissible choice, but uses no property
of that choice beyond what Assumptions \ref{assu:rv}--\ref{assu:support}
require. Letting $g$ vary leaves the first stage untouched, and with
it $V(x|z)$, so every member of $\mathcal{P}_{0}$ has the first
stage that was fixed. Write $\mathcal{P}(\lambda)$ for the distributions
in $\mathcal{P}$ whose first stage satisfies $\inf\left|\nabla_{z}V(x|z)\right|\ge\lambda$.
Fixing a first stage that meets the bound puts $\mathcal{P}_{0}$
inside $\mathcal{P}(\lambda)$, and the argument above gives the same
lower bound there. The rate is therefore $n^{-m/(2m+1)}$ over every
nonempty $\mathcal{P}(\lambda)$, so a stronger instrument does not
raise the rate at which $g(x)$ can be estimated. Nothing in the construction
is particular to this way of indexing the first stage, and the same
holds for any restriction imposed on the first stage alone.
\subsection{Bootstrap inference}
An interval that covers $g(x)$ at each $x$ separately need not cover
at every $x$ at once. Only simultaneous coverage makes a band informative
about the shape of $g$: whether it is monotone, whether it is linear,
whether it departs from a candidate specification anywhere on $\mathcal{X}_{0}$.
We study a sup-$t$ band, with the critical value taken from the empirical
bootstrap.
Let $\left\{ W^{*}_{i}\right\} ^{n}_{i=1}$ be drawn i.i.d.\ with
replacement from the original sample $\left\{ W_{i}\right\} ^{n}_{i=1}$,
and let $\hat{F}^{*}$ and $\hat{V}^{*}$ be the kernel estimators
constructed from $\left\{ W^{*}_{i}\right\} ^{n}_{i=1}$ with the
same bandwidths. Let $\widehat{S}^{*}$ be obtained by substituting
$\hat{F}^{*}$ and $\hat{V}^{*}$ for $\hat{F}$ and $\hat{V}$ in
the definition of $\widehat{S}$ above, and let $\hat{g}^{*}(x)$
be $\hat{g}_{\mathrm{LS}}(x)$ or $\hat{g}_{\mathrm{LAD}}(x)$ computed
from $\widehat{S}^{*}$ in place of $\widehat{S}$.
Write $\bar{\sigma}(x):=\left\{ \sigma^{2}(x)+\sigma^{2}(0)\right\} ^{1/2}$
for the standard deviation in Theorem \ref{thm:pointwise}. Let $\hat{s}_{n}(x)>0$
estimate $\bar{\sigma}(x)$ from the sample and $\hat{s}^{*}_{n}(x)>0$
estimate it from the bootstrap sample. They need not differ: Section
\ref{sec:Monte-carlo} uses the same estimator for both. For $\alpha\in(0,1)$
define the conditional quantile of the studentized supremum,
\begin{equation}
c_{n}(1-\alpha):=\inf\left\{ c\ge0:P^{*}\left(\sup_{x\in\mathcal{X}_{0}}\frac{\sqrt{nh_{x}}\left|\hat{g}^{*}(x)-\hat{g}(x)\right|}{\hat{s}^{*}_{n}(x)}\le c\right)\ge1-\alpha\right\} ,\label{eq:supt_crit}
\end{equation}
and the band
\begin{equation}
\mathcal{C}_{n}(x):=\left[\hat{g}(x)-c_{n}(1-\alpha)\frac{\hat{s}_{n}(x)}{\sqrt{nh_{x}}},\;\hat{g}(x)+c_{n}(1-\alpha)\frac{\hat{s}_{n}(x)}{\sqrt{nh_{x}}}\right].\label{eq:supt_band}
\end{equation}
Since $\hat{s}_{n}$ is positive, $g(x)\in\mathcal{C}_{n}(x)$ for
every $x\in\mathcal{X}_{0}$ exactly when
\[
\sup_{x\in\mathcal{X}_{0}}\frac{\sqrt{nh_{x}}\left|\hat{g}(x)-g(x)\right|}{\hat{s}_{n}(x)}\le c_{n}(1-\alpha).
\]
The studentizer that sets the half-width is therefore the one that
appears in the coverage event, while the one used in (\ref{eq:supt_crit})
enters only through the critical value. The two can be chosen separately,
and Assumption \ref{assu: band} asks the same thing of each. Studentizing
at all lets the width track the local precision of $\hat{g}$, which
in this model varies by a factor of six across $\mathcal{X}_{0}$,
since $\hat{g}(x)$ is $x$ times the average of $\widehat{\nabla_{u}g}$
over $[0,x]$.
Two choices of the pair are natural. The first takes both from the
bootstrap,
\begin{equation}
\hat{s}_{n}(x)=\hat{s}^{*}_{n}(x)=\hat{\sigma}(x),\qquad\hat{\sigma}^{2}(x):=nh_{x}\mathbb{E}^{*}\left[\left(\hat{g}^{*}(x)-\mathbb{E}^{*}\hat{g}^{*}(x)\right)^{2}\right],\label{eq:boot_sd}
\end{equation}
where $\mathbb{E}^{*}$ is expectation with respect to the resampling
distribution conditional on the data. Section \ref{sec:Monte-carlo}
implements this. Since $\hat{\sigma}$ is a function of the sample
alone it is the same in every resample, and multiplying it by a constant
leaves the band unchanged, the critical value absorbing the factor;
what the band responds to is the shape of $\hat{\sigma}$ across $\mathcal{X}_{0}$
and not its level.
The second studentizes each resample separately: form
\begin{equation}
\hat{\sigma}^{2}_{\mathrm{pl}}(a):=\frac{1}{n}\sum^{n}_{i=1}\frac{1}{h_{x}}K\left(\frac{X_{i}-a}{h_{x}}\right)^{2}\hat{c}_{a}(W_{i})^{2},\label{eq:plugin}
\end{equation}
with $\hat{c}_{a}$ formed from $\widehat{F}$, $\nabla_{y}\widehat{F}$
and $\hat{f}$. This estimates $\sigma^{2}(a)$: it reuses $h_{x}$
and replaces $K$ by $K^{2}$. Assumption \ref{assu: band} constrains
the studentizer against $\bar{\sigma}(x)$, which combines $\sigma^{2}(x)$
and $\sigma^{2}(0)$, so the plug-in is taken at both points and summed,
\[
\hat{s}_{n}(x)^{2}:=\hat{\sigma}^{2}_{\mathrm{pl}}(x)+\hat{\sigma}^{2}_{\mathrm{pl}}(0).
\]
Evaluating at $0$ is allowed although $0$ is excluded from $\mathcal{X}_{0}$:
it lies in $\widetilde{\mathcal{X}}_{0}$, where Assumption \ref{assu: density}(iii)
keeps the estimated nuisance functions away from zero.
Taking $\hat{s}_{n}$ from the sample and $\hat{s}^{*}_{n}$ from
each resample makes the band a percentile-$t$ band. The denominator
of the bootstrap statistic then comes from the same resample as its
numerator, so a resample that is dispersed inflates both. The plug-in
costs one further pass over each resample. Studentizing each resample
by a bootstrap standard deviation instead would require a bootstrap
inside each resample.
Drawing the critical value from a bootstrap rather than from an extreme-value
limit is standard in nonparametric problems \citep[see, e.g.,][]{CCK14-AoS}.
In NPIV models, \citet{HL12-JoE} and \citet{CC18-QE} provide valid
uniform confidence bands for ill-posed inverse problems. The current
setup avoids ill-posedness by using a (nonseparable) control function.
\citet{HL12-JoE} employ the empirical bootstrap as we do and studentize
with an explicit estimator of the influence-function variance; (\ref{eq:plugin})
is the counterpart here. Their band rests on finitely many points:
they form joint intervals there and interpolate between them, which
requires $g$ to be piecewise monotone or Lipschitz and leaves a term
that vanishes only as the grid gets finer. Theorem \ref{thm:band}
is stated on $\mathcal{X}_{0}$ directly, so no interpolation step
is needed. In practice $\hat{s}_{n}$ and $c_{n}$ are computed from
$B$ bootstrap draws on a finite grid; see Section \ref{sec:Monte-carlo}.
\begin{assumption}
\label{assu: band} (i) $\sup_{x\in\mathcal{X}_{0}}\left|\hat{s}_{n}(x)/\bar{\sigma}(x)-1\right|=o_{p}(1/\log n)$;
and (ii) $\sup_{x\in\mathcal{X}_{0}}\left|\hat{s}^{*}_{n}(x)/\bar{\sigma}(x)-1\right|=o_{p}(1/\log n)$
conditionally on the data, in probability.
\end{assumption}
Assumption \ref{assu: band} needs $\inf_{x\in\mathcal{X}_{0}}\bar{\sigma}(x)$,
which is ensured because Theorem \ref{thm:pointwise} already requires
$\sigma^{2}(0)>0$, and since $\sigma^{2}(x)\ge0$ that gives $\inf_{x\in\mathcal{X}_{0}}\bar{\sigma}(x)\ge\sigma(0)>0$.
Parts (i) and (ii) are the price of studentizing, and they ask for
a rate rather than for consistency alone. The reason is that the class
over which the supremum is taken is not Donsker: its envelope is of
order $h^{-1/2}_{x}$ and dominates its members (Lemma \ref{lem:vc}),
so the supremum grows like $\sqrt{\log n}$ rather than being bounded
in probability. Write $\varepsilon$ for the supremum in parts (i)
and (ii), the largest relative error the studentizer makes on $\mathcal{X}_{0}$.
Dividing by the studentizer rather than by $\bar{\sigma}$ shifts
the statistic by $\varepsilon\sqrt{\log n}$. The anti-concentration
inequality of \citet{CCK14-AoS} turns that shift into a change in
probability, again with a factor of order $\sqrt{\log n}$. The two
factors compound, so $\varepsilon$ contributes a term of order $\varepsilon\log n$
to the coverage error. Parts (i) and (ii) require $\varepsilon=o_{p}(1/\log n)$,
which makes that term vanish. With $\varepsilon=o_{p}(1)$ alone,
it need not.
The plug-in (\ref{eq:plugin}) and the bootstrap standard deviation
(\ref{eq:boot_sd}) differ in what parts (i) and (ii) then require.
For the plug-in they can be verified, which is Lemma \ref{lem:plugin}.
For $\hat{\sigma}$ taken from the bootstrap, (i) asks for the convergence
of a conditional second moment, and a distributional approximation
does not deliver one on its own. It would follow from uniform square-integrability
of $\sqrt{nh_{x}}(\hat{g}^{*}-\hat{g})$ with a rate attached, which
is not established here. Assumption \ref{assu: band} is imposed on
the pair itself rather than on either construction, so the theorem
below holds for whichever of the two is used.
\begin{thm}
\label{thm:band} Suppose the conditions of Theorem \ref{thm:pointwise}
and Assumption \ref{assu: band} hold. Then the band (\ref{eq:supt_band})
covers $g$ at every point of $\mathcal{X}_{0}$ simultaneously with
asymptotic probability $1-\alpha$,
\[
P\left(g(x)\in\mathcal{C}_{n}(x)\;\text{for all }x\in\mathcal{X}_{0}\right)\rightarrow1-\alpha.
\]
The difference between the two sides is at most a constant multiple
of
\[
(h_{x}\log n)^{1/2}+\sqrt{nh_{x}}R_{n}(\log n)^{1/2}+\sqrt{nh_{x}}\bar{h}^{m}(\log n)^{1/2}+(nh_{x})^{-1/8+\upsilon}+\left(\Delta_{n}+\Delta^{*}_{n}\right)\log n,
\]
for every $\upsilon\in(0,1/8)$, the constant depending on $\upsilon$,
where $\Delta_{n}$ and $\Delta^{*}_{n}$ are deterministic sequences
dominating the two suprema in Assumption \ref{assu: band} (i) and
(ii) with probability tending to one. The sequence $R_{n}$ is the
square of the largest uniform rate the smoothed quantities attain,
given at (\ref{eq:Rn}).
\end{thm}
\begin{proof}
See Appendix \ref{sec:proof_band}.
\end{proof}
The rescaled process $\sqrt{nh_{x}}(\hat{g}(x)-g(x))$, indexed by
$x\in\mathcal{X}_{0}$, is not tight, because the class the supremum
runs over changes with $n$. The proof therefore cannot draw the critical
value from a tight limit process, or from an extreme-value law for
its supremum, as a band on a continuum usually does. It relies instead
on two results, which do different work. \citet{CCK16-SPA} supplies
couplings: on the sample side and on the bootstrap side, each supremum
is brought close to the supremum of one Gaussian process, indexed
by the same class and at the same $n$, and the two are compared through
that common intermediary. A coupling bounds the difference between
two suprema, and does not by itself bound the difference between their
distribution functions. The anti-concentration inequality of \citet{CCK14-AoS}
supplies that step, because it limits the mass the supremum of a Gaussian
process can place in a short interval, and the same inequality sets
the rate asked of the studentizers in Assumption \ref{assu: band}.
The coverage error is bounded at each $n$ rather than in the limit,
which is why the second display carries a rate.
\begin{rem}
A band is honest over a class of distributions if its coverage error
tends to zero uniformly over the class \citep[p.~1788]{CCK14-AoS}.
Theorem \ref{thm:band} bounds the coverage error at a single distribution,
and the constant in that bound depends on it, so it does not establish
honesty. The obstacle is not the construction of the band. Their Corollary
3.1 treats a band built by undersmoothing rather than by an explicit
bias correction, and when the smoothing parameter is deterministic
its coverage is asymptotically exact uniformly over the class. The
band here is of that construction, with deterministic bandwidths by
Assumption \ref{assu:rate}. What is missing is the class: the conditions
the theorem uses are imposed at one distribution, and honesty asks
them to hold uniformly over a set of them. Three of them would have
to change. First, Assumptions \ref{assu: kernel} and \ref{assu: density}
bound densities and derivatives without naming the bounds, so there
is no class to quantify over. Proposition \ref{prop:lower_bound}
defines one, $\mathcal{P}$, by holding those bounds fixed across
the set of distributions. Second, the condition $\sigma^{2}(0)>0$
is imposed at a single distribution, and would become an infimum over
the class. Third, parts (i) and (ii) of Assumption \ref{assu: band}
would have to hold uniformly over the class and at a polynomial rate,
which is what their Condition H4 asks of the scale estimator, in place
of the $o_{p}(1/\log n)$ in probability asked here. $\blacksquare$
\end{rem}
\section{Monte Carlo }\label{sec:Monte-carlo}
We study the finite-sample behavior of the estimators and of the band
in the design
\begin{align*}
Y_{i} & =\sin(X_{i})+\varepsilon_{i},\\
X_{i} & =6L(-Z_{i}+\eta_{i})-3,
\end{align*}
where $L(a)=(1+\exp(-a))^{-1}$. The support of $X_{i}$ is $(-3,3)$
and $g(0)=0$, as Assumption \ref{assu: func_g} requires. The marginal
distributions of $\varepsilon_{i}$ and $\eta_{i}$ are both $N(0,1)$
and their joint distribution is a Frank copula
\[
F_{\varepsilon,\eta}(e,n)=-\frac{1}{\lambda}\log\left(1+\frac{(\exp(-\lambda\Phi(e))-1)(\exp(-\lambda\Phi(n))-1)}{\exp(-\lambda)-1}\right),
\]
with $\lambda=2$, following \citet{torgovitsky2017minimum}. The
instrument $Z_{i}$ is drawn independently from $N(0,1)$, so that
$(\varepsilon,\eta)$ is independent of $Z$ as Assumption \ref{assu:IV}
requires. Each experiment uses 100 Monte Carlo replications.
The first stage is nonseparable in $(Z,\eta)$, which is the case
this paper is about. It is also a case in which the control function
used by \citet{NPV99-ECTA} is not valid, so their estimator gives
a natural comparison.
The support of the weight function is chosen in a data-driven way
to avoid the denominator problem:
\[
\mathcal{Y}_{\mu}\times\mathcal{Z}_{\mu}:=\left\{ (y,z):\begin{array}{c}
\nabla_{y}\widehat{F}(y|\overline{X},z)>\tau\\
\nabla_{y}\widehat{F}(y|\overline{X},z)\nabla_{z}\widehat{V}(\overline{X}|z)>\tau
\end{array}\text{ and }\begin{array}{c}
\hat{q}_{Y}(2.5)\le y\le\hat{q}_{Y}(97.5)\\
\hat{q}_{Z}(2.5)\le z\le\hat{q}_{Z}(97.5)
\end{array}\right\} ,
\]
where $\overline{X}$ is the sample mean of $X$ and $\hat{q}_{R}(\cdot)$
is the sample quantile function for $R\in\left\{ Y,Z\right\} $. Both
bounds are on denominators of the integrand in (\ref{eq:g_functional}),
which divides by $f_{Y|X,Z}$ and by $\nabla_{z}Vf_{Y|X,Z}$. We take
$\tau=10^{-4}$ in every experiment. The estimators are evaluated
as follows: (i) draw $\{(y_{j},z_{j})\}^{M}_{j=1}$ from $\mathcal{Y}_{\mu}\times\mathcal{Z}_{\mu}$
uniformly at random, (ii) for each $j$, compute
\[
\widehat{S}^{(j)}(y_{j},x,z_{j})=\int^{x}_{0}\frac{\frac{\nabla_{x}\widehat{V}(u|z_{j})}{\nabla_{z}\widehat{V}(u|z_{j})}\nabla_{z}\widehat{F}(y_{j}|u,z_{j})-\nabla_{x}\widehat{F}(y_{j}|u,z_{j})}{\nabla_{y}\widehat{F}(y_{j}|u,z_{j})}\mathrm{d}u,
\]
where the integral over $u$ is evaluated on a grid of $M_{u}$ midpoints,
and (iii) compute $\hat{g}_{\mathrm{LS}}(x)=M^{-1}\sum^{M}_{j=1}\widehat{S}^{(j)}(y_{j},x,z_{j})$
or $\hat{g}_{\mathrm{LAD}}(x)=\arg\min_{q}Q_{b}(q|\{\hat{S}^{(j)}\}^{M}_{j=1})$
with
\[
Q_{b}(q|\{\hat{S}^{(j)}\}^{M}_{j=1})=M^{-1}\sum^{M}_{j=1}\left\{ \hat{S}^{(j)}(y_{j},x,z_{j})-q\right\} \left\{ 2\Lambda_{b}\left[\hat{S}^{(j)}(y_{j},x,z_{j})-q\right]-1\right\} ,
\]
taking $b=0.01$ and $\Lambda$ the standard normal distribution function.
We use $M=300$ and $M_{u}=100$ in the point-estimation experiments
and $M=200$, $M_{u}=50$ in the bootstrap experiment, which costs
$B=500$ times as much per replication. The support of $\mu$ and
the draws $(y_{j},z_{j})$ are selected once from the original sample
and held fixed across bootstrap replications, so that $\hat{g}^{*}-\hat{g}$
does not carry integration noise that $\hat{g}-g$ does not.\footnote{Two features of the implementation lie outside the theory as stated;
a third, the bandwidth sequence, is taken up below. The first is the
data-driven support. Assumption \ref{assu: density} asks that $\mu$
be fixed and $m$-times continuously differentiable, which an indicator
of estimated derivatives crossing a threshold is not, so the theorems
do not cover the rule used here. Appendix \ref{sec:thm1} identifies
what the smoothness provides: it makes the boundary terms of the four
integrations by parts in $z$ vanish. With a weight that is not continuous
those terms survive at order $(nh_{z})^{-1/2}$, the order of the
leading term itself, so this departure is not one that merely moves
a constant. Remark \ref{rmk:mu} gives that the estimand is the same
for every $\mu$, so a data-driven choice cannot move the target;
it does not follow that plugging in a random $\hat{\mu}$ is first-order
negligible. The second feature is the finite $M$. There is no approximation
bias: $S(y,x,z)=g(x)$ at every $(y,z)$ in the support, so $M^{-1}\sum^{M}_{j=1}S(y_{j},x,z_{j})$
is exactly $g(x)$ whatever the draws. The first-order variance is
another matter: the average over $M$ draws replaces the integral
against $\mu$ by one against their empirical measure, changing the
function $c_{a}$ and with it the variance of Theorem \ref{thm:pointwise}.
The simulations condition on the draws.}
The estimators are evaluated at 21 equally spaced points on $[-2.5,2.5]$.
We do not evaluate at $\pm3$. At the edges of the support of $X$
one has $V(x|z)\to1$ for every $z$, so $\nabla_{z}V\to0$ and the
ratio that identifies $\nabla_{x}g$ has a vanishing denominator;
Assumption \ref{assu: density} and Theorem \ref{thm:pointwise} are
stated on a compact $\mathcal{X}_{0}$ strictly inside $\mathcal{X}$
for this reason. The point $x=0$ is the normalization point, where
both estimators are zero by construction; the band excludes it, for
the reason given after Theorem \ref{thm:pointwise}, so the band is
reported on the remaining 20 points.
We set $h_{v}(n)=c\,\hat{\sigma}_{v}\,(n/1000)^{-0.15}$ for $v\in\{y,x,z\}$,
where $\hat{\sigma}^{2}_{v}$ is the sample variance. The exponent
0.15 that the simulations use for all three bandwidths departs deliberately
from what Assumption \ref{assu:rate} allows, and it is worth saying
at once in what respect. With a kernel of order $m$ the common exponent
must lie in $(1/(2m+1),1/7)$, which is $(0.111,0.143)$ at $m=4$,
and that window belongs to the case of three equal exponents. The
value 0.15 meets every condition in which $h_{x}$ appears, including
the undersmoothing one. It fails the conditions on the first-stage
bandwidths, which at $a_{x}=0.15$ ask for $a_{y},a_{z}\in(0.106,0.125)$
rather than 0.15. Using 0.15 for all three is therefore a further
respect in which the implementation departs from the theory. The first-stage
bandwidths it gives differ by a few per cent from those an exponent
inside the window would give, and at $n=1000$ not at all, since the
factor $(n/1000)^{-0.15}$ equals one there. Assumption \ref{assu:rate}
restricts the exponents and leaves the constants free, so $c$ can
be chosen separately for point estimation and for the band. We choose
it by comparing values of $c$ on a grid at $n=1000$. Table \ref{tab:bw}
reports that grid for the LAD estimator. Two criteria pick different
constants. Root mean squared error is smallest at $c=0.60$. The ratio
$\max_{x}|\text{bias}|/\text{sd}$, which governs how far a confidence
interval centered at $\hat{g}$ is displaced relative to its own width,
is smallest at $c=0.40$. We use $c=0.60$ for point estimation and
$c=0.40$ for the band.
\begin{table}[t]
\centering
\begin{tabular}{lcccccccc}
\hline
$c$ & $0.20$ & $0.30$ & $0.40$ & $0.50$ & $0.60$ & $0.75$ & $0.90$ & $1.10$\\
\hline
RMSE & $0.336$ & $0.226$ & $0.175$ & $0.148$ & $\mathbf{0.135}$ & $0.146$ & $0.190$ & $0.248$\\
$\max_x|\text{bias}|/\text{sd}$ & $0.869$ & $0.544$ & $\mathbf{0.461}$ & $0.578$ & $0.869$ & $1.484$ & $3.143$ & $6.025$\\
\hline
\end{tabular}
\caption{Bandwidth selection for the LAD estimator, $n=1000$, 21 evaluation points, 100 replications. Smallest value in each row in bold.}
\label{tab:bw}
\end{table}
The bias is not monotone in $c$. Below $c=0.40$ the trimming in
the display above becomes active: the surviving $(y,z)$ set falls
from about 9,200 of the 10,201 grid points at $c=0.40$ to 6,400 at
$c=0.20$, and the ratio rises again. The choice of the bandwidth
shrinks $\mathcal{Y}_{\mu}\times\mathcal{Z}_{\mu}$ since a smaller
one makes the estimated denominators smaller and noisier.
The kernel is the fourth-order Gaussian $K(u)=\tfrac{1}{2}(3-u^{2})\varphi(u)$,
where $\varphi$ is the standard normal density. A fourth-order kernel
can take negative values, so the estimated densities can go non-positive
and the ratio defining $\nabla_{x}g$ can blow up. In this simulation,
the estimated densities stayed positive in all 100 replications at
every bandwidth in Table \ref{tab:bw}.
\subsection{Point estimation}
Table \ref{tab:point} reports bias, standard deviation and root mean
squared error, averaged over the 21 evaluation points, for the two
estimators of this paper and for the series estimator of \citet{NPV99-ECTA}
at series lengths $p=2,\dots,6$. The series estimator uses a power
series basis with the same $p$ at both stages: $\mathbb{E}[X|Z]$
is fitted by a polynomial of degree $p$ in $Z$, and $Y$ is then
fitted by a polynomial of total degree $p$ in $(X,\hat{\eta})$,
with the estimate of $g$ taken from the terms in $X$ alone. Both
of our estimators use $c=0.60$. For reference, the range of $g$
over the grid is $1.995$.
\begin{table}[t]
\centering
\begin{tabular}{lcccccc}
\hline
& \multicolumn{3}{c}{$n=500$} & \multicolumn{3}{c}{$n=1000$}\\
& $|\text{bias}|$ & sd & RMSE & $|\text{bias}|$ & sd & RMSE\\
\hline
LAD & $0.047$ & $0.152$ & $0.179$ & $0.048$ & $0.104$ & $\mathbf{0.133}$\\
LS & $0.075$ & $0.241$ & $0.294$ & $0.070$ & $0.185$ & $0.237$\\
\hline
NPV, $p=2$ & $0.309$ & $0.086$ & $0.352$ & $0.310$ & $0.068$ & $0.348$\\
NPV, $p=3$ & $0.056$ & $0.126$ & $\mathbf{0.162}$ & $0.065$ & $0.093$ & $\mathbf{0.133}$\\
NPV, $p=4$ & $0.052$ & $0.171$ & $0.214$ & $0.065$ & $0.123$ & $0.162$\\
NPV, $p=5$ & $0.073$ & $0.207$ & $0.265$ & $0.065$ & $0.143$ & $0.180$\\
NPV, $p=6$ & $0.077$ & $0.242$ & $0.315$ & $0.067$ & $0.169$ & $0.213$\\
\hline
\end{tabular}
\caption{Bias, standard deviation and RMSE, averaged over 21 points on $[-2.5,2.5]$, 100 replications. Smallest RMSE in each column in bold.}
\label{tab:point}
\end{table}
Three things are visible. First, LAD dominates LS by a wide margin
at the same bandwidth, by a factor of 1.6 in RMSE at $n=500$ and
1.8 at $n=1000$. The gap is variance, not bias. It narrows if LS
is given its own bandwidth --- at $n=1000$ the RMSE-minimizing value
for LS is $c=0.75$, giving $0.245$ --- but LAD remains ahead. We
report LAD as the main estimator from here on.
Second, the series estimator's RMSE varies by a factor of 2.6 across
$p=2,\dots,6$ at $n=1000$. However, one can use a cross-validation
to choose $p$, and it picks $p=3$ in 89 of 100 samples at $n=500$
and 90 at $n=1000$. No equally natural criterion is available for
the bandwidth. Held-out prediction of $Y$ is the closest analogue,
but it targets $\mathbb{E}[Y\mid X]=g(X)+\mathbb{E}[\varepsilon\mid X]$,
which differs from $g$ precisely because $X$ is endogenous. The
difficulty is not specific to this estimator: \citet[Section 2.2]{CCK25-ReStud}
argue that under endogeneity the cross-validation criterion need not
estimate the mean squared error even asymptotically, so it is not
a meaningful criterion for choosing the sieve dimension, and they
replace it by a data-driven rule. Nothing of the kind is available
for a bandwidth here, and its selection remains an open problem.
Third, we do not claim an advantage in root mean squared error. LAD
beats the series estimator at every $p$ except $p=3$, and loses
to $p=3$ by 11 per cent at $n=500$. At $n=1000$ the two are level:
the difference is $0.0008$.
It should be noted that a misspecified estimator often achieves lower
mean squared error than a consistent one at a given sample size \citep[see, e.g.,][]{AKS25-ECTA}.
Lower error at one sample size is therefore not a reason to use the
series estimator here. What separates the two is whether the error
vanishes as the sample grows. Table \ref{tab:cons} in Appendix \ref{sec:Consistency}
follows both estimators from $n=250$ to $n=2000$. The bias of our
estimators falls with the sample: $n^{-0.19}$ for LS and $n^{-0.28}$
for LAD. The bias of the series estimator does not fall, at $n^{-0.05}$
with $p$ chosen by cross-validation. Its standard deviation falls
fastest of all, at $n^{-0.56}$, and by $n=2000$ its bias is larger
than its standard deviation.
\subsection{The band}
We now turn to the band (\ref{eq:supt_band}), computed for the LAD
estimator at $n=1000$ with $c=0.40$ and $B=500$ bootstrap draws.
The critical value $c_{n}(1-\alpha)$ and the scale $\hat{\sigma}$
are computed on the 20 evaluation points other than $x=0$; $\hat{\sigma}$
is the bootstrap standard deviation at each point, not smoothed across
points. This is the first of the two studentizers of Section \ref{sec:Estimation-and-bootstrap},
that is, $\hat{s}_{n}=\hat{s}^{*}_{n}=\hat{\sigma}$.
\begin{figure}[t]
\begin{centering}
\includegraphics[width=0.8\textwidth]{band_sin_LAD}
\par\end{centering}
\caption{Uniform band and pointwise intervals at $\alpha=0.10$, LAD, $n=1000$,
$c=0.40$, $B=500$. Averages over 100 replications, with the bands
from individual replications in grey.}\label{fig:band}
\end{figure}
Figure \ref{fig:band} shows the band averaged over replications,
with the pointwise intervals and the individual replications behind
it. The width varies across $\mathcal{X}_{0}$, because $\hat{g}(x)$
carries the factor $x$: the band is about six times wider at the
outermost evaluation points $\pm2.5$ than at the innermost $\pm0.25$,
and narrows toward the normalization point, which is not itself one
of the evaluation points. That factor multiplies the bias as well,
but the smoothing bias is much smaller, so the mean of $\hat{g}$
still tracks $g$ where the band is widest. The individual replications
drawn behind the average share a common vertical displacement, which
$T_{n}(0)$ of Lemma \ref{lem:boundary} leaves: the deviation at
every $x$ carries the same $-T_{n}(0)$. The shared term makes the
deviations covary: for any two evaluation points, however far apart,
the covariance of $\hat{g}-g$ contains the variance of $T_{n}(0)$.
Without that term the covariance would fall to zero once the kernels
at the two points stop overlapping, at a separation of about two bandwidths.
Averaged over the 97 pairs of points more than 1.5 apart, it is 0.005,
against a variance of 0.029 at each point. About a fifth of the variance
at a point is therefore shared with every other point. The covariance
is positive in 95 of those 97 pairs, so the value is not noise around
zero. The comparison is unaffected by the smoothing bias, which is
common across replications and so drops out of a covariance.
\subsection{Coverage}
Table \ref{tab:cov} reports coverage over 100 replications. With
100 replications a coverage rate has a Monte Carlo standard error
of about $0.03$ at $0.90$. The row marked simultaneous asks how
often the pointwise intervals cover $g$ at all 20 points at once,
which is the quantity the band is designed to control. The two pointwise
rows differ only in centering. Write $q^{*}_{\tau}(x)$ for the $\tau$
quantile of the bootstrap replicates $\hat{g}^{*}(x)$. The row marked
Efron is $[q^{*}_{\alpha/2}(x),q^{*}_{1-\alpha/2}(x)]$. The row marked
percentile is its reflection about $\hat{g}(x)$, $[2\hat{g}(x)-q^{*}_{1-\alpha/2}(x),2\hat{g}(x)-q^{*}_{\alpha/2}(x)]$.
The two have the same width at every point, which is why the width
columns agree.
\begin{table}[t]
\centering
\begin{tabular}{lcccc}
\hline
& \multicolumn{2}{c}{$\alpha=0.10$} & \multicolumn{2}{c}{$\alpha=0.05$}\\
& coverage & width & coverage & width\\
\hline
uniform band & $0.950$ & $0.927$ & $0.980$ & $1.018$\\
\hline
pointwise, percentile & $0.808$ & $0.529$ & $0.888$ & $0.632$\\
pointwise, Efron & $0.915$ & $0.529$ & $0.964$ & $0.632$\\
pointwise, Efron, simultaneous & $0.550$ & --- & $0.790$ & ---\\
\hline
\end{tabular}
\caption{Coverage and average width, LAD, $n=1000$, $c=0.40$, $B=500$, 100 replications. Pointwise entries are averaged over the 20 points; the simultaneous row is the fraction of replications in which the pointwise intervals cover at all 20 points.}
\label{tab:cov}
\end{table}
There is no evidence of undercoverage: the band covers in $95$ replications
at $\alpha=0.1$ and 98 replications at $\alpha=0.05$. The pointwise
intervals cover at every point at once in 55 replications. The band
achieves this by being wider, by a factor of 1.75, the ratio of the
two critical values. That width is still useful because the band separates
the peak of the sine from its trough with room to spare. The band
also over-covers, at 0.950 against a nominal 0.90 and 0.980 against
0.95, and Theorem \ref{thm:band} does not predict that. Its error
bound is explicit but not numerically small here: at $n=1000$ with
$c=0.40$, so that $h_{x}\approx0.63$ and $\log n\approx6.9$, the
leading term $(h_{x}\log n)^{1/2}$ is about $2.1$ and the coupling
term about $2.4$, before any constant. As a numerical statement about
the coverage error the bound therefore says nothing at this sample
size; what the theorem gives is that the error vanishes, not that
it is small here.
The bootstrap supremum is larger than the sampling supremum it approximates.
Over 40 replications with 60 resamples, at the same $n$ and bandwidth,
the supremum of $\sqrt{nh_{x}}|\hat{g}^{*}-\hat{g}|$ over the resamples
exceeds the supremum of $\sqrt{nh_{x}}|\hat{g}-g|$ over the replications
by 12 per cent at the $0.95$ quantile (the common $\hat{\sigma}$
divides both suprema and so does not enter the comparison). Thus the
critical value made from the bootstrap distribution becomes larger,
and the band also becomes wider. Recomputing the plug-in standard
error inside each resample, which is the second studentizer of Section
\ref{sec:Estimation-and-bootstrap}, widens the gap to 23 per cent
instead of closing it.