Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Honest Confidence Sets in Nonparametric IV Regression and Other Ill-Posed Models
abstractThis paper develops inferential methods for a very general class of ill-posed models in econometrics encompassing the nonparametric instrumental variable regression, various functional regressions, and the density deconvolution. We focus on uniform confidence sets for the parameter of interest estimated with Tikhonov regularization, as in darolles2011nonparametric. Since it is impossible to have inferential methods based on the central limit theorem, we develop two alternative approaches relying on the concentration inequality and bootstrap approximations. We show that expected diameters and coverage properties of resulting sets have uniform validity over a large class of models, i.e., constructed confidence sets are honest. Monte Carlo experiments illustrate that introduced confidence sets have reasonable width and coverage properties. Using U.S. data, we provide uniform confidence sets for Engel curves for various commodities.
keywordsnonparametric IV regression, functional linear regression, honest uniform confidence sets, non-asymptotic inference, Tikhonov regularization.
jel\; C14, C36
Introduction
This paper develops honest and uniform confidence sets for the structural function $\varphi$ in a generic class of ill-posed models treated with the Tikhonov regularization. The leading example is a nonparametric instrumental variable regression (NPIV) studied in florens2003inverse, newey2003instrumental, darolles2011nonparametric, hall2005nonparametric, and blundell2007semi. The NPIV model is
equation[equation omitted — 87 chars of source]
where $Z\in\ensuremath{\mathbf{R}}^p$ is a vector of explanatory variables, $W\in\ensuremath{\mathbf{R}}^q$ is a vector of instruments, and $\varphi$ is the function of interest. The NPIV model is ill-posed in the sense that the map from the distribution of the data to the function $\varphi$ is not continuous. As a result, we need to introduce some amount of regularization that would smooth out discontinuities and yield a consistent estimator.
In empirical studies, the function $\varphi$ represents a structural economic relation, such as an Engel curve, cost function, or demand curve. It is not enough to estimate this function to infer possible economic effects. We can only infer the range of possible economic effects from confidence sets. This paper is the first to provide inferential methods for Tikhonov-regularized estimators. I focus on uniform inference, which amounts to constructing a set containing the entire function $\varphi$ with a high probability. Uniform inference allows one to assess global features of the estimated function and to quantify the range of possible economic effects compatible with the data. Global features may include the evidence for non-linearities, the amount of endogeneity bias compared to the local polynomial estimator, monotonicity, concavity/convexity, or other shape properties. In contrast, pointwise confidence intervals only contain the value $\varphi(z_0)$ at some particular point $z_0$ with high probability and do not provide a valid inference for the entire function $\varphi$. Another feature of confidence sets constructed in this paper is honesty in the sense of li1989honest. Honesty means that coverage properties have uniform validity over a large class of specified models.\footnote{See Eq. (ref) in Section (ref) for a formal definition.} Honesty is desirable since the underlying model is never known and coverage properties of dishonest sets may vary from one model to another.
Building uniform confidence sets for a function requires approximating the distribution of the supremum of a certain stochastic process. We show that for a broad class of ill-posed models treated with Tikhonov regularization, the stochastic process driving the distribution of the estimator does not converge weakly in the space of functions. As a result, it is not possible to build the uniform confidence set by relying on the uniform central limit theorem which calls for alternative approaches to inference.
To construct confidence sets with good coverage properties, we rely on two different approaches. The first inferential method developed in this paper relies on estimates of tail probabilities with a suitable concentration inequality. A simple illustration of this approach is to build a uniform confidence band for the empirical distribution function using the Dvoretzky-Kiefer-Wolfowitz inequality. Given the empirical distribution function $F_n(x)=\frac{1}{n}\sum_{i=1}^n\ensuremath{\mathds{1}}_{\{X_i\leq x\}}$ based on the i.i.d. sample $(X_i)_{i=1}^n$, the Dvoretzky-Kiefer-Wolfowitz tells us that for any finite sample size $n$ the probability of $F_n$ deviating from the true distribution function $F$ in the supremum norm declines at the exponential rate
equation*[equation* omitted — 92 chars of source]
Setting $x=\sqrt{\frac{\log (2/\gamma)}{2n}}$ and $\gamma\in(0,1)$, the inequality becomes
equation*[equation* omitted — 105 chars of source]
This allows us to build a uniform confidence band for $F(x)$ with a guaranteed coverage probability $1-\gamma$ taking $F_n(x)\pm\sqrt{\frac{\log (2/\gamma)}{2n}}$. In this simple example, there is no coverage error (the coverage is at least $1-\gamma$ for any finite sample size $n$), and the diameter of the set shrinks at the rate $1/\sqrt{n}$. This inferential approach is different from the alternative asymptotic approach based on the Donsker central limit theorem, which also leads to confidence sets with diameters shrinking at the rate $1/\sqrt{n}$, but having the coverage level $1-\gamma-o(1)$. The coverage error disappears only as sample size goes to infinity. The empirical distribution function is an unbiased estimator, while the majority of estimators of functions are biased. As a result, confidence sets for such estimators, including the ill-posed models considered in this paper, usually have coverage errors.
In this paper, we rely on a concentration inequality valid for more complex statistics than $\|F_n - F\|_\infty$, see boucheron2013concentration. In particular, we exploit a data-driven concentration inequality for the supremum of the variance of the estimator to approximate quantiles of the unknown distribution. This approach does not rely on the existence of the supremum of the Gaussian process approximating the supremum of the variance of the estimator. As a result, it is valid for a broad class of data-generating processes and is especially useful in settings where all other approaches to inference fail.
For confidence sets based on the concentration inequality, we characterize non-asymptotic rates of coverage errors and expected diameters explicitly. It is especially important to know both rates, since consistency alone may not be very informative for inference. For instance, the coverage error may decrease at a rate slower than the shrinking of the confidence set, requiring larger sample sizes to achieve good coverage. We show that the bias of the estimator, not the noise coming from the estimation of the operator, drives the coverage errors of our confidence sets. To the best of our knowledge, convergence rates for coverage errors have not been derived for the NPIV model or other ill-posed models considered in this paper.
As an alternative, we also study a more traditional bootstrap inference that relies on a non-trivial application of Gaussian and bootstrap coupling inequalities developed in a seminal series of papers, chernozhukov2014gaussian, chernozhukov2016empirical. To the best of our knowledge, our application of these coupling results to the nonparametric IV model and Tikhonov estimators is new.
Although the Tikhonov-regularized NPIV is the leading example, the inferential methods developed in this paper are valid for other ill-posed models, including functional regression models and the density deconvolution model. To the best of our knowledge, despite the extensive existing literature on $L_2$ results for the functional regression models, no uniform convergence rates or uniform inferential methods are currently available. Moreover, Tikhonov regularization plays a prominent role in the statistical learning theory. The inferential results obtained in this paper can also be applied in that setting.
\paragraph*{Contribution and related literature.} This paper is the first to develop honest uniform inferential methods in a general and unifying framework, encompassing different ill-posed models treated with Tikhonov regularization. This paper is also the first to provide uniform inference for several ill-posed inverse problems based on the concentration inequality and to provide the non-asymptotic analysis, including convergence rates of coverage errors. Uniform confidence sets are available only for sieve-type estimators of the NPIV model and the density deconvolution model. The literature is also silent about coverage errors of confidence sets. Moreover, it is not known what sorts of restrictions are needed to construct honest confidence sets for ill-posed inverse problems. In this paper, we show that the class of models should involve both the class of functions and the class of operators, which is different from classical direct nonparametric problems where we need only to restrict the function. horowitz2012uniform develop uniform confidence bands for sieve NPIV estimator by first constructing pointwise confidence intervals at a finite grid of points and then letting the number of grid points grow at a certain speed to achieve uniform coverage. chen2018optimal develop inferential methods for the sieve NPIV estimator without relying on discrete approximations. They focus on uniform inference for a collection of linear and nonlinear functionals using Yurinskii's coupling and obtain uniform confidence bands in the particular case of point evaluation functionals, see also belloni2015some for the conditional mean function and tao2014inference for general conditional moment restriction models. lounici2011global obtain uniform confidence sets for the wavelet deconvolution estimator by relying on Bousquet's version of Talagrand's concentration inequality.
kato2016uniform consider a more general case, where the density of the noise is estimated from an auxiliary sample. This paper studies the estimator based on Fourier inversion and builds on the coupling inequalities developed in chernozhukov2014gaussian and chernozhukov2016empirical.
We show that the confidence sets developed in this paper may enjoy polynomial convergence rates of coverage errors in mildly ill-posed, and some severely ill-posed, cases. An appealing feature of Tikhonov's and more general spectral regularizations is that the estimator is always well-defined in finite samples and converges to the best approximation to the structural parameter under various identification failures, see florens2011identification and babiiflorens2016b.
The theoretical validity of a uniform confidence set relies on uniform convergence rates of the corresponding estimator. gagliardini2012tikhonov obtain uniform convergence rates for the Tikhonov-regularized minimum distance estimator relying on the Sobolev embedding. chen2018optimal show that the sieve nonparametric IV estimator can attain optimal asymptotic uniform convergence rates under a somewhat different set of assumptions. babiiflorens2016b derive non-asymptotic $L_2$ and uniform risk bounds for more general spectrally regularized estimators not tied to a particular estimator of the conditional mean function and show that it is possible to attain polynomial convergence rates in some severely ill-posed cases.\footnote{The minimax-optimal uniform convergence rates in the severely ill-posed case are logarithmically slow, see chen2018optimal. However, in the more restricted smoothness class, when the estimated function can be described by the finite number of generalized Fourier coefficients, or more generally if its smoothness matches the ill-posedness of the operator, one can achieve faster polynomial uniform convergence rates.}
Some pointwise inferential results are available for spectral cut-off estimators when the ill-posed operator is known and is not estimated from the data, see carrasco2011spectral, gagliardini2012nonparametric, gautier2013nonparametric, and florenshorowitzkeielgom2016. At the same time, chen2015sieve provide pointwise inference and bootstrap confidence bands for conditional moment restriction models treated with a sieve approach and nesting the NPIV model as a special case. cardot2007clt provide results on the asymptotic normality of linear functionals, which solves the prediction problem for functional regression models. Lastly, carrasco2013asymptotic study asymptotic normality of inner products for Tikhonov estimators.
The structure of the paper is as follows. We introduce the notation in the remaining part of this section. Section (ref) describes the problem of constructing honest uniform confidence sets and introduces two inferential approaches in the special case of the nonparametric IV model. Under a general set of assumptions that are verified later on in each particular application, we establish convergence rates for coverage errors and diameters of constructed sets, uniform over a general set of models. Section (ref) extends these results to different functional regression models, and to the density deconvolution. In Section (ref) we show how to implement confidence sets in practice for the NPIV estimator and explore their finite-sample properties with Monte Carlo experiments. Section (ref) considers the empirical application to Engel curves, and Section (ref) concludes.
\paragraph{Notation.}
Let $L_2[0,1]^p$ denote the space of functions on some compact set $[0,1]^p\subset \ensuremath{\mathbf{R}}^p$, square integrable with respect to the Lebesgue measure $\lambda$. For $\varphi\in L_2[0,1]^p$, let $\|.\|$ denote the usual $L_2$ norm derived from the inner product $\langle .,.\rangle$. Let $C[a,b]^p$ denote the space of continuous functions endowed with supremum norm $\|.\|_\infty$. For some positive real number $\beta$, let $C^\beta_M[a,b]^p$ denote the class of $\beta$-H\"{o}lder functions on $(a,b)^p$ with $0<M<\infty$
equation*[equation* omitted — 283 chars of source]
where $k=(k_1,\dots,k_p)\in\ensuremath{\mathbf{N}}^p$ is a multi-index, $|k|=\sum_{j=1}^pk_j$, $\varphi^{(k)}(z)=\frac{\partial^{|k|}\varphi(z)}{\partial z_1^{k_1}\dots\partial z_p^{k_p}}$, and $\lfloor \beta\rfloor$ is the largest integer strictly smaller than $\beta$. For $A\subset\ensuremath{\mathbf{R}}$, let $BV(A)$ be a set of functions of bounded variation on $A$. Let $\ensuremath{\mathcal{L}}_{2}$, $\ensuremath{\mathcal{L}}_{2,\infty}$ and $\ensuremath{\mathcal{L}}_{\infty}$ be spaces of bounded linear operators from $L_2$ to $L_2$, from $L_2$ to $C$, and from $C$ to $C$. Sets on which functions are defined should be clear from the context. Spaces $\ensuremath{\mathcal{L}}_2,\ensuremath{\mathcal{L}}_{2,\infty}$, and $\ensuremath{\mathcal{L}}_{\infty}$ are endowed with standard operator norms, denoted by $\|K\|=\sup_{\|\varphi\|\leq 1}\|K\varphi\|$, $\|K\|_{2,\infty} = \sup_{\|\varphi\|\leq1}\|K\varphi\|_\infty$, and $\|K\|_\infty=\sup_{\|\varphi\|_\infty\leq1}\|K\varphi\|_\infty$. For $K\in\ensuremath{\mathcal{L}}_{2}$, let $K^*$ denote its Hilbert adjoint operator. Let $\ensuremath{\mathcal{R}}(T)$ and $\ensuremath{\mathcal{D}}(T)$ be the range and the domain of the operator $T$. For two real numbers $a$ and $b$, I denote $a\wedge b = \min\{a,b\}$ and $a\vee b = \max\{a,b\}$.
Honest uniform confidence sets
We focus on uniform confidence sets, honest to some class of models $\ensuremath{\mathcal{P}}$. The class of $\ensuremath{\mathcal{P}}$ consists of probability distributions of the data satisfying certain restrictions. For a given level $\gamma\in(0,1)$, the honest $1-\gamma$ uniform confidence set, denoted $C_{n,1-\gamma}=\left\{C_{n,1-\gamma}(z)=\left[C_l(z), C_u(z)\right],\;z\in[0,1]^p\right\}$, satisfies the following coverage
equation[equation omitted — 165 chars of source]
for some sequence $\delta_n\to 0$. Honesty is desirable for theoretical and practical reasons. On the theoretical side, it is well-known that focusing on a fixed model leads to inconsistent concepts of optimality for non-parametric models; see, e.g., tsybakov2009introduction, p.16-19. Honesty also ensures that for a sufficiently large sample size $n$, independent from the distribution of the data, the coverage level of the set will be close to $1-\gamma$, no matter what model in $\ensuremath{\mathcal{P}}$ nature gives us. Honesty is desirable in practice since the underlying probability distribution is unknown. In contrast, a dishonest set requires a weaker condition
equation*[equation* omitted — 155 chars of source]
and the sample size $n$ needed to achieve the coverage close to $1-\gamma$ will depend on the distribution of the data.
The second requirement for the confidence set is that its expected diameter under the supremum norm, denoted $|C_{n,1-\gamma}|_\infty=\|C_u-C_l\|_\infty$, shrinks at some rate $\rho_n\to0$
equation*[equation* omitted — 123 chars of source]
We will also consider a variation of this requirement when the diameter is bounded in probability only.
Nonparametric estimation usually involves a bias-variance trade-off for the risk of the estimator. This trade-off translates into a trade-off between the rate $\delta_n$ at which coverage errors tend to zero and the rate $\rho_n$ at which the expected diameter tends to zero. Roughly speaking, we show that the coverage error is driven by the bias of the estimator, while the variance determines its diameter. It is vital to know both rates and requiring only the limiting coverage $1-\gamma$ in the Eq. (ref) can be misleading, as the coverage error of the set may tend to zero arbitrarily slow. In the following sections, we show that the confidence sets based on the concentration inequality allow us to characterize coverage errors explicitly.
It is also worth mentioning that among confidence sets based on different $L_p,p\in[1,\infty]$ norms, only uniform confidence sets have appealing visualization and are easy to implement numerically. For example, if $\varphi$ is the function on the real line, the confidence set becomes a band on the plane, which contains the whole graph of $\varphi$ with high probability.
Nonparametric IV
The model
The nonparametric IV model
equation*[equation* omitted — 74 chars of source]
leads to the functional equation
equation[equation omitted — 102 chars of source]
which is an example of an ill-posed inverse problem. Even if we knew the distribution of the data $(Y,Z,W)$, obtaining $\varphi$ requires inverting the conditional expectation operator, which is typically not continuous. florens2003inverse and darolles2011nonparametric introduce Tikhonov regularization, see tikhonov1963solution, to smooth out discontinuities of inversion.
Multiplying Eq. (ref) by the marginal density function $f_W$
equation*[equation* omitted — 159 chars of source]
Following darolles2011nonparametric we focus on the kernel estimator, but other choices are possible. For simplicity, we use the product kernel and equal bandwidth parameters $h_n\to 0$ as $n\to\infty$ for all coordinates
equation*[equation* omitted — 343 chars of source]
The Tikhonov-regularized estimator solves the penalized least-squares problem
equation*[equation* omitted — 130 chars of source]
which admits a closed-form solution
equation*[equation* omitted — 81 chars of source]
where $\hat T^*$ is the adjoint operator of $\hat T$. The adjoint operator is a solution to $\langle \hat T\varphi,\psi \rangle = \langle \varphi,\hat T^*\psi\rangle$ and is computed by Fubini's theorem
equation*[equation* omitted — 102 chars of source]
Ill-posedness and regularization bias
We first introduce the class of functions and operators for which we wish to obtain the honest coverage.
assumptionThe structural function $\varphi$ and the 1-1 operator $T:C[0,1]^p\to C[0,1]^q$ belong to the following class
\begin{equation*}
\ensuremath{\mathcal{F}} = \ensuremath{\mathcal{F}}_{\beta,t,M,C} = \left\{(\varphi,T)\in C^t_M[0,1]^p\times \ensuremath{\mathcal{L}}_{2,\infty}:\; \varphi = (T^*T)^\beta T^*\psi,\;\kappa(\varphi,\psi,T,T^*)\leq C\right\},
\end{equation*}
where $\kappa(\varphi,\psi,T,T^*) = \|\varphi\|_\infty\vee \|T^*\|_{2,\infty}\vee\|\psi\|\vee \|T\|^{-1}$ and $0<\beta,t,M,C<\infty$.
Source conditions are at the heart of spectral regularization theory, and describe the regularity of the problem by restricting how ill-posed the operator $T$ is, compared to the smoothness of the parameter of interest $\varphi$ (see also chen2011rate for its comparability with other assumptions used in the literature). The present source condition is different from the one used to characterize $L_2$ rates, where we would require $\varphi = (T^*T)^\beta\psi$ only, see carrasco2007linear.
We show in the next section that under this source condition, our confidence sets enjoy polynomial convergence rates of coverage errors. In the severely ill-posed case, when singular values of the operator $T$ decay to zero exponentially fast, the source condition in Assumption (ref) requires that the function $\varphi$ is well-approximated by a small number of generalized Fourier coefficients with respect to the SVD basis of $T$. It is a remarkable fact that we can still achieve polynomial convergence rates in this case, since many functions encountered in practice have rapidly declining generalized Fourier coefficients.
Assumption (ref) allows us to control the regularization bias. Note that in the Proposition that follows, the bound is uniform over the source set, and this will be needed to establish honesty of our confidence sets.
propositionSuppose that Assumption (ref) is satisfied, then
\begin{equation*}
\sup_{(\varphi,T)\in\ensuremath{\mathcal{F}}}\left\|(\alpha_n I + T^*T)^{-1}T^*r - \varphi\right\|_\infty \leq R\alpha_n^{\beta\wedge 1},
\end{equation*}
where $R = C^2\left[\beta^\beta(1-\beta)^{1-\beta}\ensuremath{\mathds{1}}_{0<\beta<1} + C^{2(\beta-1)}\ensuremath{\mathds{1}}_{\beta\geq1}\right]$.
In terms of the bias, the regularization takes advantage of the smoothness up to $\beta=1$. This effect is somewhat reminiscent of the saturation of convergence of the kernel density estimator and can be avoided using the iterated or the extrapolated Tikhonov regularization similarly to using higher or infinite order kernels for the kernel density estimator, see darolles2011nonparametric and carrasco2007linear for more discussion.
Honest confidence sets
Put
equation*[equation* omitted — 149 chars of source]
where $\hat X_{ni}(w) = \frac{1}{\sqrt{h_n^q}}\hat U_iK_w\left(h_n^{-1}(W_i - w)\right)$, $\hat U_i = Y_i - \hat\varphi(Z_i)$, and $\eta_i$ are i.i.d. Rademacher random variables. Let $c_{1-\gamma}^*$ be $1-\gamma$ quantile of
equation*[equation* omitted — 129 chars of source]
conditionally on the data $\ensuremath{\mathcal{D}} = (Y_i,Z_i,W_i)_{i=1}^n$ and $\varepsilon_i$ are i.i.d. N(0,1).
The class of models $\ensuremath{\mathcal{P}}$ is restricted by the following assumptions.
assumption(i) $(Y_i,Z_i,W_i)_{i=1}^n$ is an i.i.d. sample of $(Y,Z,W)$; (ii) $f_{ZW}\in C^s([0,1]^{p+q}),s>0$ and there exists $\underline f,\bar f$ such that $0<\underline f\leq f_{ZW}\leq \bar f<\infty$; (iii) $K_z$ and $K_w$ are products of the kernel function $k:\ensuremath{\mathbf{R}}\to\ensuremath{\mathbf{R}}$ with $k\in C(\ensuremath{\mathbf{R}})\cap BV(\ensuremath{\mathbf{R}})$ having order $\lfloor s\rfloor\vee\lfloor t\rfloor$ and such that $k\in L_r(\ensuremath{\mathbf{R}}),r=1,2,3,4$ and $\int|u|^{s\vee t}|k(u)|\ensuremath{\mathrm{d}} u<\infty$; (iv) $\ensuremath{\mathds{E}}|Y|^4<\infty$, $\ensuremath{\mathds{E}}[|U|^2|W]\geq \underline \sigma_2>0$ a.s., and $|U|\leq F$; (v) the integral operator $T:C[0,1]^p\to C[0,1]^q$ is $1$-$1$; (vi) $h_n$ and $\alpha_n$ tend to $0$ polynomially fast and (a) $h_n^tn\log^2 n\to 0$, $\alpha_n^{2(\beta\wedge 1 + 1)}h_n^qn\log^2 n \to 0$, and $\alpha_nh_n^{2s+q}n\log^2 n\to 0$; (b) $\alpha_n\log^3 n/h_n^p\to0$, $h_n^{s-q/2}\log n/\alpha_n \to 0$, and $\alpha_n^{\beta\wedge 1}\log n/h_n^{q/2}\to 0$; (c) $\alpha_n^2nh_n^{2q}/\log n\to\infty$.
Most of these assumptions impose mild smoothness, boundedness, or moment conditions which are plausible in empirical applications. Conditions imposed on the kernel functions are extremely mild and are verified by all kernel functions commonly used in practice. Similar assumptions are usually made in other settings for uniform nonparametric estimation and inference. Assumption (ref) imposes mild regularity conditions on the data generating process. (i) is the standard sampling assumption and may be relaxed to time series data under some weak dependence conditions, e.g., see babii2016b; (ii) may rule out common elements between $Z$ and $W$, see discussion in Section (ref) how to deal with common elements; (iii) allows for boundary-corrected kernels, see darolles2011nonparametric\footnote{Otherwise, to avoid problems at end-points, we will make a restriction to the interior of $[0,1]^{p+q}$, and all results should be read as uniform over the interior of this set.} for further discussion. Assumption (v) is a completeness condition which may be questionable in empirical applications. Note, however, that in many cases the completeness condition can be relaxed, and it is possible to have valid inferences for various deviations from the completeness condition, see babiiflorens2016b for more details.
Assumption (ref) (vi) involves several conditions on tuning parameters and deserves special attention. It imposes implicitly that different components of the model have sufficient regularity with respect to the dimensions $p$ and $q$. To illustrate that these conditions are not contradictory, suppose that $h_n\sim n^{-c_1}$ and $\alpha_n\sim n^{-c_2}$ for some $c_1,c_2>0$ and that $\beta\geq 1$. Then Assumption (ref) reduces to the following 6 conditions: (1) $c_2>(p\vee q/2)c_1$, (2) $4c_2+qc_1>1$, (3) $2c_2+2qc_1<1$, (4) $c_1>1/t$, (5) $c_1(s-q/2)>c_2$, and (6) $c_2+c_1(2s+q)>1$. Conditions (1)-(3) define the set of feasible value for $c_1$ and $c_2$ depending on the dimension parameters $p$ and $q$. Then (4) and (5)-(6) simply require that smoothness parameters $t$ and respectively $s$ are sufficiently large. For instance, when $p=q=1$, we can take, $h_n\sim n^{-1/5}$ and $\alpha_n\sim n^{-1/4}$, and need $t>5$ and $s>7/4$.
The bootstrap and the concentration inequality-based confidence sets are described as
equation[equation omitted — 265 chars of source]
with
equation*[equation* omitted — 304 chars of source]
where the operator norm can be computed using Lemma (ref) as
equation*[equation* omitted — 121 chars of source]
The next result provides theoretical justification for our honest uniform confidence sets.
theoremSuppose that $\ensuremath{\mathcal{P}}$ consists of distributions satisfying Assumptions (ref) and (ref). Then
\begin{equation*}
\inf_{P\in\ensuremath{\mathcal{P}}}\Pr\left(\varphi\in C_{n,1-\gamma}\right) \geq 1 - \gamma - O(\delta_n)
\end{equation*}
with
\begin{equation*}
\begin{aligned}
\delta_n & = \left(\frac{h_n^\frac{t-q}{2}}{\alpha_n} + \frac{1}{\alpha_n^{1/2}}\left(\sqrt{\frac{\log h_n^{-1}}{nh_n^{p+q}}} + h_n^s\right) + \alpha_n^{\beta\wedge 1}\right)h_n^{q/2}\alpha_nn^{1/2}\log n + \frac{\log n}{\alpha_n}\left(\frac{1}{\sqrt{nh_n^q}} + h_n^s\right). \\
\end{aligned}
\end{equation*}
Moreover, if Assumption (ref) (v)-(vi) is satisfied
\begin{equation*}
\liminf_{n\to\infty}\inf_{P\in\ensuremath{\mathcal{P}}}\mathrm{Pr}\left(\varphi\in C_{n,1-\gamma}^*\right) \geq 1 - \gamma.
\end{equation*}
The proof of this result can be found in the Appendix and is a consequence of more general Theorems (ref) and (ref). In particular, Theorem (ref) derives coverage errors for concentration inequality-based confidence sets, while Theorem (ref) establishes the asymptotic validity of bootstrap confidence sets for a general class of ill-posed inverse models. Theorem (ref) is based on the Gaussian and bootstrap coupling inequalities as well as the anti-concentration inequality of chernozhukov2014gaussian and chernozhukov2016empirical. Note that concentration-based confidence sets allow us to characterize explicitly the coverage error, which has not been known for confidence sets in the nonparametric IV model. Inspection of the proof of Theorem (ref) reveals that bootstrap confidence sets have additional coverage errors due to the Gaussian and bootstrap approximations.
To illustrate that conditions needed for $\delta_n\to 0$ are not contradictory, suppose that $h_n\sim n^{-c_1}$ and $\alpha_n\sim n^{-c_2}$ for some $c_1,c_2>0$ and that $\beta\geq 1$. Then we need (1) $c_2>pc_1$, (2) $4c_2+qc_1>1$, (3) $2c_2+qc_1<1$, (4) $c_1>1/t$, (5) $c_1s>c_2$, and (6) $c_2+c_1(2s+q)>1$. Note that the set of feasible values of $c_1,c_2$ is larger than that allowed by Assumption (ref) (vi).
The following theorem describes the expected size of our confidence sets.
theoremSuppose that assumptions of Theorem (ref) are satisfied. Then
\begin{equation*}
\sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\left|C_{n,1-\gamma}\right|_\infty = O\left(\frac{1}{\alpha_n\sqrt{nh_n^q}}\right)\quad and\quad \left|C_{n,1-\gamma}^*\right|_\infty = O_p\left(\frac{1}{\alpha_nh_n^qn^{1/2}}\right),
\end{equation*}
uniformly over $P\in\ensuremath{\mathcal{P}}$.
Note that we need $\alpha_n^2nh_n^{q}\to \infty$ to ensure that the concentration-based confidence set shrinks and a slightly stronger condition $\alpha_n^2nh_n^{2q}\to \infty$ to ensure that the bootstrap confidence set shrinks, which is also needed to ensure that the coverage error tends to zero.
Functional regressions and density deconvolution
Functional regressions
In the functional linear regression, the real dependent variable $Y$ is explained by the continuous-time stochastic process $Z(t),t\in[0,1]^p$. The distinctive feature of this class of models is that it allows handling high-dimensional data without relying on the sparsity assumption, which may be restrictive in this setting. There is vast literature on functional regression models in statistics, see cardot2003testing, hall2007methodology, and references therein. There is also some growing literature in econometrics, see florens2015instrumental, benatia2017functional, and babii2016b. However, all these papers focus on the estimation in the $L_2$ norm and to the best of my knowledge, there are no currently available uniform inferential methods for functional regression models.
The functional IV regression model is described as
equation*[equation* omitted — 149 chars of source]
The slope parameter $\varphi$ measures the strength of the impact of the process $Z$ at different points $t\in[0,1]^p$. If $W=Z$, we obtain the classical functional linear regression model without endogeneity, see, for instance, hall2007methodology, while in the IV case of florens2015instrumental, $W$ is some functional instrumental variable, uncorrelated with the error term.
The moment restriction leads to the ill-posed equation
equation*[equation* omitted — 175 chars of source]
The slope parameter $\varphi$ is identified when the covariance operator is 1-1, which generalizes the non-singularity condition for the covariance matrix in the finite-dimensional linear regression model.
A variation of this model is studied in babii2016b, where the identification is achieved with real-valued instrumental variables through the conditional moment restriction $\ensuremath{\mathds{E}}[U|W]=0$. The identifying restriction is the linear completeness condition. Conditional mean-independence leads to the following ill-posed equation
equation[equation omitted — 149 chars of source]
which for appropriate families of functions $\Psi(s,.)$ can equivalently be written as
equation*[equation* omitted — 184 chars of source]
In what follows we denote by $W(s)$ either the regressor $Z(s)$, some functional IV $W(s)$, or the function $\Psi(s,W)$ of some real IV $W$. This notation allows encompassing all variations of the functional regression discussed above.
Unlike the NPIV, which requires a non-parametric estimation of the joint density function, all components of functional regression models are estimated at the parametric rate (in the $L_2$ norm) using sample analogs to population moments
equation*[equation* omitted — 120 chars of source]
Operator $T$ can be estimated as
equation*[equation* omitted — 99 chars of source]
The adjoint operator $\hat T^*$ can be obtained using Fubini's theorem
equation*[equation* omitted — 99 chars of source]
Put
equation*[equation* omitted — 137 chars of source]
where $\hat X_{ni}(s) = \hat U_iW_i(s)$, $\hat U_i = Y_i - \langle Z_i,\varphi\rangle$, and $\eta_i$ are i.i.d. Rademacher random variables.
Let $c_{1-\gamma}^*$ be $1-\gamma$ quantile of
equation*[equation* omitted — 129 chars of source]
and $\varepsilon_i\sim_{i.i.d.}N(0,1)$.
The following assumption restricts the class of models $\ensuremath{\mathcal{P}}$.
assumption(i) $(Y_i,Z_i,W_i)_{i=1}^n$ is an i.i.d. sample of $(Y,Z,W)$; (ii) $Z$ and $W$ have continuous trajectories, $\|Z\|_\infty\leq C$ and $\|UW\|_\infty\leq F$; (iii) $W\in C^s_M[0,1]^q$ and $ZW\in C_M^s[0,1]^{p+q}$, for some $s>(p+q)/2$; (iv) $\ensuremath{\mathds{E}}|UW(s)|\geq \underline\sigma>0$ for all $s\in[0,1]^q$; (v) the integral operator $T:C[0,1]^p\to C[0,1]^q$ is 1-1; (vi) $\alpha_n\to 0$ polynomially fast and $\alpha_nn/\log^2 n\to \infty$ and $\alpha_n^{2(\beta\wedge 1) + 2}n\log^2n\to0$.
Note that the functional linear regressions admit a simpler characterization and involve even less stringent assumptions than the nonparametric IV model.
The bootstrap and the concentration inequality-based confidence sets are
equation*[equation* omitted — 111 chars of source]
and
equation*[equation* omitted — 117 chars of source]
with
equation*[equation* omitted — 287 chars of source]
where the operator norm can be computed using Lemma (ref) as
equation*[equation* omitted — 116 chars of source]
The next result provides theoretical justification for confidence sets based on the concentration inequality and bootstrap confidence sets.
theoremSuppose that Assumptions (ref) and (ref) are satisfied. Then for any $\gamma\in(0,1)$
\begin{equation*}
\inf_{P\in\ensuremath{\mathcal{P}}}\mathrm{Pr}\left(\varphi\in C_{n,1-\gamma}\right) \geq 1 - \gamma - O\left(\left(\alpha_n^{\beta\wedge 1+1}n^{1/2} + \frac{1}{\sqrt{\alpha_nn}}\right)\log n\right).
\end{equation*}
Moreover, if Assumption (ref) (vi) is satisfied
\begin{equation*}
\liminf_{n\to\infty}\inf_{P\in\ensuremath{\mathcal{P}}}\mathrm{Pr}\left(\varphi\in C_{n,1-\gamma}^*\right) \geq 1 - \gamma.
\end{equation*}
Note that for concentration inequality-based confidence sets the coverage error is effectively driven by the bias induced by the regularization and kernel smoothing. For bootstrap confidence sets we have an additional coverage error due to the Gaussian and bootstrap approximations and need additional condition $nh_n^2/\log^7n\to \infty$. The following result describes the expected size of various confidence sets.
theoremSuppose that assumptions of Theorem (ref) are satisfied, then
\begin{equation*}
\sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\left|C_{n,1-\gamma}\right|_\infty = O\left(\frac{1}{\alpha_nn^{1/2}}\right)\qquad and\qquad \left|C_{n,1-\gamma}^*\right|_\infty = O_p\left(\frac{1}{\alpha_nn^{1/2}}\right),
\end{equation*}
uniformly over $P\in\ensuremath{\mathcal{P}}$.
To ensure that confidence sets shrink we need $\alpha_n^2n \to \infty$. In both cases, sets shrink at the same speed.
Density deconvolution
Often economic data are not measured precisely. Density deconvolution allows estimating the density of unobserved data from the data measured with errors. Density deconvolution is encountered in a variety of econometric applications, e.g. to the earning dynamics in bonhomme2010generalized, to panel data in evdokimov2010identification, or to the instrumental regression in adusumilli2018nonparametric. In the simplest example of this model, see carrasco2011spectral, we have some noisy scalar observations $Y$, of the latent variable $Z$, contaminated by measurement errors $U$
equation*[equation* omitted — 59 chars of source]
Distributions of both $Z$ and $U$ are assumed to be absolutely continuous with respect to the Lebesgue measure with densities $\varphi$ and $f$. The density of measurement errors $f$ is assumed to be known. The goal is to recover the density of the latent variable $Z$ from observing contaminated i.i.d. sample $(Y_i)_{i=1}^n$.
Independence and additivity of the noise imply that the density function $r$ of $Y$ satisfies the following convolution equation
equation[equation omitted — 111 chars of source]
where the operator $T:L_2\to L_2$. For simplicity of presentation and tractability of our results, let us assume that all densities are continuous, bounded and compactly supported, with support contained inside the interval $[0,1] \subset \ensuremath{\mathbf{R}}$. This assumption is not very restrictive, since most of the economic variables, e.g., reported earning or costs are bounded.
The adjoint operator to $T$ can be computed using Fubini's theorem
equation*[equation* omitted — 78 chars of source]
We estimate the density of $Y$ from the sample $(Y_i)_{i=1}^n$ using the kernel density estimator
equation*[equation* omitted — 90 chars of source]
Put
equation*[equation* omitted — 147 chars of source]
where $\hat X_{ni}(y) = \frac{1}{\sqrt{h_n}}K\left(\frac{Y_i - y}{h_n}\right) - \frac{1}{n\sqrt{h_n}}\sum_{i=1}^nK\left(\frac{Y_i - y}{h_n}\right)$ and $\eta_i$ are i.i.d. Rademacher random variables.
The class of models $\ensuremath{\mathcal{P}}$ is restricted by the following assumptions.
assumption(i) $(Y_i)_{i=1}^n$ is an i.i.d. sample of $Y$; (ii) $f,\varphi, r$ are continuous, compactly supported on some subsets of $[0,1]$, and the density $r$ is uniformly bounded away from zero and infinity by some finite constant; (iii) $r\in C_M^s[0,1]$, for some $s>0$; (iv) $K$ is a symmetric continuous square-integrable kernel function of bounded variation of order $\lfloor s\rfloor$ satisfying $\int|u|^s|K(u)|\ensuremath{\mathrm{d}} u<\infty$; (v) the integral operator $T:C[0,1]^p\to C[0,1]^q$ is 1-1; (vi) $h_n\to 0$ and $\log h_n^{-1}=O(\log n)$, $nh_n^2/\log^7n\to \infty$, $nh_n^{2s+1}\log^2 n\to 0$, and $\alpha_n^{2(\beta\wedge 1) + 2}nh_n\log^2n\to0$.
Note that for this model, operators $T$ and $T^*$ are known, which simplifies the analysis. Let $c_{1-\gamma}^*$ be $1-\gamma$ quantile of
equation*[equation* omitted — 129 chars of source]
and $\varepsilon_i\sim_{i.i.d.}N(0,1)$.
The bootstrap and the concentration inequality-based confidence sets are described as
equation*[equation* omitted — 111 chars of source]
and
equation*[equation* omitted — 117 chars of source]
with
equation*[equation* omitted — 289 chars of source]
where the operator norm can be computed using Lemma (ref) as
equation*[equation* omitted — 106 chars of source]
The next result provides theoretical justification for confidence sets based on the concentration inequality and bootstrap confidence sets.
theoremSuppose that Assumptions (ref) and (ref) are satisfied. Then for any $\gamma\in(0,1)$
\begin{equation*}
\inf_{P\in\ensuremath{\mathcal{P}}}\mathrm{Pr}\left(\varphi\in C_{n,1-\gamma}\right) \geq 1 - \gamma - O\left(\left(h_n^{s} + \alpha_n^{\beta\wedge 1+1}\right)\sqrt{nh_n}\log n + \frac{\log n}{n^{1/2}}\right).
\end{equation*}
Moreover, if Assumption (ref) (vi)-(vii) is satisfied
\begin{equation*}
\liminf_{n\to\infty}\inf_{P\in\ensuremath{\mathcal{P}}}\mathrm{Pr}\left(\varphi\in C_{n,1-\gamma}^*\right) \geq 1 - \gamma.
\end{equation*}
Note that for concentration inequality-based confidence sets the coverage error is effectively driven by the bias induced by the regularization and kernel smoothing. For bootstrap confidence sets we have an additional coverage error due to the Gaussian and bootstrap approximations and need additional condition $nh_n^2/\log^7n\to \infty$. The following result describes the expected size of various confidence sets.
theoremSuppose that assumptions of Theorem (ref) are satisfied, then
\begin{equation*}
\sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\left|C_{n,1-\gamma}\right|_\infty = O\left(\frac{1}{\alpha_n\sqrt{nh_n}}\right)\qquad and\qquad \left|C_{n,1-\gamma}^*\right|_\infty = O_p\left(\frac{1}{\alpha_nh_nn^{1/2}}\right),
\end{equation*}
uniformly over $P\in\ensuremath{\mathcal{P}}$.
To ensure that a concentration inequality-based confidence set shrinks, we need $\alpha_n^2h_nn \to \infty$ while for the bootstrap confidence set we need the slightly stronger requirement $\alpha_n^2h_n^2n\to \infty$.
Implementation and Monte Carlo experiments
This section reports results of Monte Carlo experiments for our confidence sets. We focus on the NPIV estimator. Samples of size $n\in\{1000,5000\}$ are generated as follows
equation*[equation* omitted — 405 chars of source]
where $\sigma_z = \sigma_w = 0.3$, $\sigma_u = \sqrt{0.03}, \sigma_{zu} = 0.04$, and $\rho=0.3$. To be consistent with the theory, we keep only observations inside a sufficiently large compact set. The strength of the instrument as measured by the correlation coefficient $\rho$ is calibrated to our empirical application.
Since the joint density of $(Z,W)$ is bivariate normal, this choice of functional forms resembles up to a constant the convolution of Gaussian densities. Therefore, with appropriate modifications, we can easily adapt the simulation design not only to functional regression but also to the deconvolution model. Since the underlying integral equation would be the same in all cases, to economize on space, we do not report MC experiments for other models. We also remark that the normal density leads to rapidly declining singular values. This design corresponds to the most difficult severely ill-posed setting. We consider $5000$ replications of each experiment.
We consider confidence sets described in Eq. (ref), where the operator norm $\|\hat T^*\|_{2,\infty}$ is computed using the mixed norm of the kernel function of $\hat T^*$ by Lemma (ref). The estimate of the density function $f_{ZW}$ is obtained using kernel smoothing. For simplicity of implementation, we do not optimize the performance with higher-order boundary kernels and take the product of second-order Epanechnikov kernels.
The estimator is discretized using a simple Riemann sum on the grid of $100$ equidistant points. This suffices to ensure that the numerical errors are negligible compared to the statistical noise in our setting. A higher number of grid points or cubature can give better approximation if needed. The discretized estimator has closed-form expression, and its computation boils down to solving a system of linear equations
equation*[equation* omitted — 139 chars of source]
where $\mathbf{I}$ is the $T\times T$ identity matrix, $\mathbf{r} = \left(\frac{1}{n}\sum_{i=1}^nY_ih_n^{-1}K\left(h_n^{-1}(W_i - w)\right)\right)_{1\leq j\leq T}$, $\mathbf{K}=\boldsymbol{f}\Delta$ with $\boldsymbol{f} = (\hat f_{ZW}(z_k,w_j))_{1\leq j,k\leq T}$, and $\Delta$ is the grid step. Therefore, the estimator is straightforward to implement and fast to compute.
\paragraph{Data-driven choice of tuning parameters.} As discussed in Section (ref), for confidence sets, the bias-variance trade-off for the risk of the estimator reduces to the trade-off between coverage errors and the diameter of confidence sets. For larger values of tuning parameters, the band becomes too narrow, and the bias starts to dominate, reducing uniform coverage. On the other side, even though smaller values of tuning parameters reduce bias and the estimator becomes closer to its population value, they also increase the size of the variance and lead to wider bands. The optimal choice of tuning parameters balances the two.
Existing data-driven rules for Tikhonov regularization aim to obtain an accurate estimator, see feve2010practice or centorrino2014data, and do not take into account the trade-off between the coverage and the diameter. Note that the risk-optimal estimation conflicts with optimal inference. In particular, our theory requires undersmoothing in order to obtain valid inferences. We suggest the following procedure inspired by bissantz2007non:
enumerate• Use some over-smoothing data-driven rule to select pilot tuning parameters. We consider the cross-validation studied in centorrino2014data which consists of solving the leave-one-out empirical counterpart to $\|\hat r - \hat T\hat\varphi_{\alpha,h}\|$
\begin{equation*}
(\hat\alpha_J,\hat h_J) = \operatorname*{arg\,min}_{\alpha,h>0}\frac{1}{n}\sum_{i=1}^n\left\{\frac{1}{n-1}\sum_{j\ne i}(Y_j - [K_h\ast\hat\varphi_{\alpha,h}](Z_j))\frac{1}{h}K\left(h^{-1}(W_j - W_i)\right) \right\}^2,
\end{equation*}
where $[K_h\ast g](x) = \frac{1}{h}\int g(z)K(h^{-1}(x-z))\ensuremath{\mathrm{d}} x$ denotes the convolution operator. To ensure that tuning parameters are over-smoothing, we scale pilot tuning parameters $(\kappa_a\hat\alpha_J,\kappa_h\hat h_J)$ with some $\kappa_a,\kappa_h\geq 1$.
• Compute $\hat\varphi_{h_j,\alpha_j}$ on a grid of points $(j\kappa_a\alpha_J/J,j\kappa_hh_J/J)_{j=1}^J$. Choose the largest $(\alpha_j,h_j)$ such that
\begin{equation*}
\|\hat\varphi_{\alpha_j,h_j} - \hat\varphi_{\alpha_{j-1},h_{j-1}}\|_\infty > \tau\|\hat\varphi_{\alpha_J,h_J} - \hat\varphi_{\alpha_{J-1},h_{J-1}}\|_\infty
\end{equation*}
Like bissantz2007non we find that $J=20$ and $\tau = 2$ work well in practice. We also find that $(\kappa_a,\kappa_h) = (3.5,1.5)$ is suitable for concentration-based confidence sets and $(\kappa_a,\kappa_h) = (1,1.5)$ works well for bootstrap confidence sets. Note that bissantz2007non have a single tuning parameter, and they find that the factor $1.5$ is suitable, which is what we find for the bandwidth parameter.
\paragraph{Simulation results.} In Figure (ref) we plot estimates with a $95\%$ confidence band based on bootstrap approximation, averaged over $5000$ Monte Carlo experiments. Figure (ref) represents the same plot for confidence bands based on concentration inequality. We also report MC coverage probabilities $\hat\gamma$. In our Monte Carlo experiments, bootstrap approximation leads to slightly better performance. Resulting confidence bands are narrower, and the estimator is centered more closely to the population value of the parameter of interest. Nonetheless, confidence sets based on the concentration inequality are also informative about global shape properties of the estimated function. For example, Figure (ref) (b) gives an excellent confidence band with a reasonably small amount of bias.
figure[figure omitted — 452 chars of source]
figure[figure omitted — 461 chars of source]
\paragraph{Exogenous covariates.}
In the presence of additional exogenous covariates, one can consider several avenues. As an example, suppose that one is interested in estimating the demand function, so $Z=(Z_1,Z_2)$ with $Z_1$ denoting the endogenous price and $Z_2$ denoting exogenous covariates, such as income and other control variables. One way to accommodate the exogenous covariates is to estimate the semiparametric partially linear specification allowing for non-linearities in prices, see yatchew2001household and florens2015instrumental.
Alternatively, one may want to keep the fully non-parametric specification. In this case, the model $Y=\varphi(Z_1,Z_2)+U,\ensuremath{\mathds{E}}[U|Z_2,W]=0$ leads to the following family of ill-posed integral equations
equation*[equation* omitted — 147 chars of source]
For any fixed value of $z_2$, we are back to our setting, and our uniform inferential methods would allow us to construct uniform confidence sets for $z_1\mapsto \varphi(z_1,z_2)$ for any fixed value of $z_2$. Uniform inference for the demand function for different income levels might be sufficient to reveal the amount of non-linearities and/or endogeneity. Constructing the joint uniform confidence set for the bivariate function $\varphi$ and the theoretical analysis of such sets is beyond the scope of the present paper.
Engel curves in the United States
In this section we estimate Engel curves using the NPIV approach and construct uniform confidence sets using the bootstrap. Engel curves describe how the demand for a commodity changes while the household's budget increases. Estimation of Engel curves is fundamental for the analysis of consumer behavior and has implications in different fields of empirical research. Interesting applications include the measurement of welfare losses associated with tax distortions in banks1997quadratic, estimation of growth and inflation in nakamura2014chinese, or estimation of income inequality across countries in almaas2012international.
Previously blundell2007semi estimated a shape-invariant system of Engel curves for food, alcohol, and fuel on UK data with a sieve approach. For simplicity, we focus on the non-parametric specification of the Engel curve.
Our dataset is drawn from the 2015 US Consumer Expenditure Survey data, and we estimate Engel curves for a set of goods, including food, tobacco, alcohol, gas and oil, and health. In our subsample, we have married couples with positive income during the past 12 months, including households with and without children. The dependent variable is the share of expenditures on the particular commodity. The log of total expenditures on non-durable goods is used as an independent variable. As in blundell2007semi we instrument the log of total expenditures with income before tax. However, in contrast to blundell2007semi, we also estimate Engel curves for tobacco and health care.
We plot NPIV estimates with a $95\%$ uniform confidence bands in Fig. (ref). We also plot the local linear estimator, which does not correct for the endogeneity bias. Estimates in Figure (ref) illustrate that there is significant endogeneity bias in the Engel curves. Indeed, the local polynomial estimator is often outside the confidence band. For instance, the local polynomial estimator largely overestimates the shape of the Engel curve for higher budgets. In most of the cases the estimated Engel curves exhibit more curvature than suggested by the local polynomial estimator. While it may look like Engel curves slightly increase at the beginning, this observation is not statistically significant and most likely can be attributed to boundary effects.
figure[figure omitted — 1,052 chars of source]
Conclusions
This paper studies uniform inferential methods in ill-posed models treated with Tikhonov regularization. Building uniform confidence sets in this setting is a difficult problem and requires approximating the distribution of the supremum of a complex empirical process. We show that it is not possible to establish functional convergence of this process for a very general class of ill-posed models known in econometrics and statistics. Nonetheless, we demonstrate that it is possible to obtain uniform inference relying on alternative methods.
We develop two approaches to uniform inference that lead to honest confidence sets. Honest confidence sets are of practical interest for several reasons. They ensure the existence of a certain sample size after which the coverage level will be not significantly smaller than the nominal coverage level, regardless of how complex the estimated function is within a given smoothness class. Moreover, it is widely recognized in non-parametric statistics, results for a fixed model can lead to inconsistent notions of optimality. Honest confidence sets, on the other hand, have the uniform validity.
The first approach developed in this paper relies on data-driven concentration inequality and allows for non-asymptotic characterization of coverage errors and expected diameters. These rates are new and have not been previously discussed in the literature. We also develop the asymptotic bootstrap approach to inference relying on the coupling inequalities of chernozhukov2016empirical and show that the resulting confidence sets have desirable theoretical properties.
Both methods demonstrate good performance in Monte Carlo experiments, and it seems that bootstrap confidence sets are slightly more narrow for a comparable coverage level. On the other hand, concentration-based confidence sets are computationally less demanding and still reveal lots of information about the global shape properties of functional parameters. Lastly, unlike bootstrap and/or Gaussian approximation-based inference, concentration inequalities do not impose any restrictions on indexing classes of empirical processes. Therefore, this approach can be easily adapted to other complex nonparametric estimators.
\setcounter{page}{1}
\setcounter{section}{0}
\setcounter{equation}{0}
\setcounter{table}{0}
\setcounter{figure}{0}
center[center omitted — 40 chars of source]
Concentration of the supremum of empirical processes
Let $(X_i)_{i\in\ensuremath{\mathbf{N}}}$ be a sequence of i.i.d. random variables taking values in some measure space $(S,\ensuremath{\mathcal{S}})$, and let $\ensuremath{\mathcal{H}}$ be a countable class of real functions defined on $S$. Consider the empirical process
equation*[equation* omitted — 132 chars of source]
and denote the symmetrized process by
equation*[equation* omitted — 70 chars of source]
where $(\eta_i)_{i\in\ensuremath{\mathbf{N}}}$ is a sequence of i.i.d. Rademacher random variables, independent from $(X_i)_{i\in\ensuremath{\mathbf{N}}}$. We use $\|.\|_\ensuremath{\mathcal{H}}$ to denote the supremum norm over the class of functions $\ensuremath{\mathcal{H}}$.
propositionThe following two inequalities hold
\begin{equation}
2^{-1}\ensuremath{\mathds{E}} \left\|\nu_n^\eta\right\|_\ensuremath{\mathcal{H}} - 2^{-1}n^{-1/2}\|Ph\|_\ensuremath{\mathcal{H}} \leq \ensuremath{\mathds{E}}\|\nu_n\|_\ensuremath{\mathcal{H}} \leq 2\ensuremath{\mathds{E}} \left\|\nu_n^\eta\right\|_\ensuremath{\mathcal{H}}.
\end{equation}
The second inequality in Eq. (ref), is a symmetrization inequality, while the first is the desymmetrization inequality. These inequalities allow us to establish uniform limit theorems and play an essential role in the empirical process theory, see van2000weak. We refer to koltchinskii2006local, p.7, and references therein for a proof of this modification of the symmetrization inequality.
Our construction of confidence sets relies on the concentration inequality for functions of bounded difference, also known as McDiarmid's inequality, see the excellent treatment of concentration inequalities in boucheron2013concentration. Such concentration inequalities tell us that the probability that the supremum of the empirical process $\|\nu_n\|_\ensuremath{\mathcal{H}}$ deviates from its expected value $\ensuremath{\mathds{E}}\|\nu_n\|_\ensuremath{\mathcal{H}}$ declines exponentially fast. While this expected value is unknown in practice, it can be estimated using the multiplier bootstrap with Rademacher multipliers. The symmetrization inequality allows us then to compare the expected value of this estimate $\|\nu_n^\eta\|_\ensuremath{\mathcal{H}}$ to the original expected value, which leads to the data-driven concentration inequality. This insightful idea comes from the statistical learning literature (see, e.g., koltchinskii2001rademacher).
The next proposition states the precise version of the data-driven concentration inequality used in the present paper.
propositionLet $(X_i)_{i=1}^n$, $(\eta_i)_{i=1}^n$, and $\ensuremath{\mathcal{H}}$ be defined as before. Suppose also that the absolute value of all $h\in\ensuremath{\mathcal{H}}$ is uniformly bounded by some $H$. Then for all $x>0$ and $n\in\ensuremath{\mathbf{N}}$
\begin{equation*}
\Pr\left(\|\nu_n\|_\ensuremath{\mathcal{H}} > 2\left\|\nu_n^\eta\right\|_\ensuremath{\mathcal{H}} + 3H\sqrt{\frac{2x}{n}}\right) \leq e^{-x}.
\end{equation*}
See koltchinskii2011oracle, Theorem 4.6, for the proof of this result.
General ill-posed inverse problems
Impossibility of weak convergence
We first discuss the impossibility of using the uniform central limit theorem to obtain the distribution of the estimator.
To fix notation through the rest of the paper, let $P$ be a probability measure to any of the ill-posed models introduced in the previous section. The model is described by the function equation $r=T\varphi$, where $\varphi:[0,1]^p\to\ensuremath{\mathbf{R}}$ is an infinite dimensional parameter of interest. We aim to construct a random set $C_{n,1-\gamma}$ containing the function $\varphi$ with probability of at least $1-\gamma$ for $\gamma\in(0,1)$ and such that its diameter shrinks as the sample size increases at a specific rate. We focus on confidence sets for the Tikhonov-regularized estimator defined as follows
equation*[equation* omitted — 83 chars of source]
where $\hat T,\hat T^*$, and $\hat r$ are appropriate estimators\footnote{If some operators are known, which is the case in the density deconvolution model, we replace estimators by known quantities.} and $\alpha_n$ is some positive sequence converging to zero as $n\to\infty$. This estimator belongs to the general family of spectral regularization schemes, see carrasco2007linear.
For models considered in this paper, the dominating stochastic component of the Tikhonov-regularized estimator is driven by a sequence of i.i.d. centered random functions
equation[equation omitted — 102 chars of source]
Assume that $T$ is an integral operator with continuous kernel function, so that $T,T^*$, and $T^*T$ map to the space of continuous functions. Thus, we can also think of the operator $T$ as acting between $(C,\|.\|_\infty)$ spaces. By Lemma (ref) in the Appendix C, the operator $(\alpha_n I + T^*T)$ is invertible between spaces of continuous functions. Thus, trajectories of the process in Eq. (ref) belong to the space $(C,\|.\|_\infty)$.
To build uniform confidence sets, we need to approximate the distribution of the supremum of this process. The most straightforward route to achieve this would be to establish its weak convergence to some Gaussian process and then to rely on quantiles of Gaussian suprema to build uniform confidence sets. Unfortunately, the process in Eq. (ref) can not converge weakly as a random element in $(C,\|.\|_\infty)$ space. Indeed, we introduced regularization to smooth out discontinuities of the operator inversion. However, in the limit, as the regularization parameter tends to zero, we get back a discontinuous/unbounded operator. This discontinuity in the limit is a precise reason why the functional convergence fails.
More precisely, we can see that depending on the direction $\delta\in L_2$, the speed of convergence of inner products $\langle\hat\varphi - \varphi,\delta\rangle$ differs, see carrasco2007linear, which makes it impossible to converge weakly in $L_2$. In Proposition (ref) we show formally that the weak convergence in $(C,\|.\|_\infty)$ is impossible generalizing the result of cardot2007clt who focus on the weak convergence in $L_2$ in the special case of the functional linear regression without instrumental variables. Note that our result holds for the arbitrary ill-posed inverse model, including the nonparametric IV, functional IV regressions, and the density deconvolution. Note also that our result holds regardless of what type of the estimator is used in practice, be it kernel or series estimators in the nonparametric IV model. In particular, it is worth comparing it to the impossibility of weak convergence of the kernel density estimator, where inner products do converge at the same speed, but the only compatible limiting process is a white noise process, as in ruymgaart1998note.
propositionSuppose that the inverse of the operator $T^*T$ in Eq. (ref) is unbounded. Then there does not exist a normalizing sequence $r_n$ such that $r_n\nu_{n}$ would converge weakly in $(C,\|.\|_\infty)$ to a non-degenerate random process.
proof[Proof of Proposition (ref)]
Recall that for some normalizing sequence $r_n$, the weak convergence of $r_n\nu_{n}$ in $L_2$ to some random element requires $\langle r_n\nu_{n},\delta\rangle$ to converge weakly in $\ensuremath{\mathbf{R}}$ for all $\delta\in L_2$, van2000weak, Theorem 1.8.4. Since $(T^*T)^{-1}$ is unbounded with $\ensuremath{\mathcal{D}}[(T^*T)^{-1}]\subset L_2[0,1]^p$, for $\delta\in \ensuremath{\mathcal{D}}[(T^*T)^{-1}]$, we need to set $r_n=n^{1/2}$, since
\begin{equation*}
\left\langle r_n\nu_{n},\delta\right\rangle = \frac{1}{\sqrt{n}}\sum_{i=1}^n\left\langle T^*X_{ni},(\alpha_n I + T^*T)^{-1}\delta\right\rangle\xrightarrow{d} N\left(0,\ensuremath{\mathds{E}}\left\langle T^*X_1,(T^*T)^{-1}\delta\right\rangle^2\right).
\end{equation*}
On the other hand, if $\delta\not\in \ensuremath{\mathcal{D}}[(T^*T)^{-1}]$, $\|(\alpha_nI + T^*T)^{-1}\delta\|^2\to\infty$, making $\ensuremath{\mathds{E}}\langle T^*X_1,(\alpha_nI + T^*T)^{-1}\delta\rangle^2\to\infty$, and so $\langle n^{1/2}\nu_{n},\delta\rangle$ can not converge in distribution. This shows that it is not possible to converge weakly in $L_2$. Since bounded and continuous functionals on $L_2$ are bounded and continuous on $(C,\|.\|_\infty)$ for finite measure spaces, it follows from the definition of weak convergence that $r_n\nu_{n}$ does not converge weakly in $(C,\|.\|_\infty)$ for any choice of the normalizing sequence $r_n$.
Proposition (ref) tells us that it is not possible to rely on the asymptotic approximation with the conventional central limit theorem. As a result, we need to rely on alternative approaches to inference.
Bootstrap confidence sets
Consider the empirical process
equation*[equation* omitted — 119 chars of source]
indexed by some pointwise measurable pre-Gaussian class of functions $\ensuremath{\mathcal{G}}_n\subset L_2(P)$, where $L_2(P)$ is the set of square-integrable functions with respect to $P$ and $P_n$ is the empirical measure. We will always assume that the data are i.i.d. In our setting the empirical process $V_{n}$ can also be considered as a stochastic process
equation*[equation* omitted — 70 chars of source]
indexed by $w\in[0,1]^q$. Let
equation*[equation* omitted — 143 chars of source]
be the supremum of the empirical process and let
equation*[equation* omitted — 203 chars of source]
be the supremum of the tight centered Gaussian process with the same covariance structure as $V_{n}$, i.e., for all $g_1,g_2\in\ensuremath{\mathcal{G}}_n$
equation*[equation* omitted — 218 chars of source]
We shall note that we will work only with processes having continuous trajectories so that all suprema can be restricted to countable sets of rational numbers, and whence are measurable.
Note that by the linearity
equation*[equation* omitted — 156 chars of source]
so we can always drop the absolute value from the supremum extending the supremum to the enlarged class of functions. A similar observation is valid for the supremum of the Gaussian process dudley2014uniform, Theorem 2.1. We will typically work with suprema of absolute value and will use the above facts repeatedly.
Let $G_{n}$ be the envelope of the class $\ensuremath{\mathcal{G}}_n$, which we assume to be square-integrable and let $N(\ensuremath{\mathcal{G}}_n,\|.\|_{Q,2},\epsilon)$ be the $\epsilon$-covering of $(\ensuremath{\mathcal{G}}_n,\|.\|_{Q,2})$, with $\|g\|_{Q,q} = (Q|g|^q)^{1/q}$ for some probability measure $Q$. We assume that $\ensuremath{\mathcal{G}}_n$ is a VC-type class of functions.
assumptionSuppose that the class $\ensuremath{\mathcal{G}}_n$ is such that
\begin{equation*}
\sup_{Q}N(\ensuremath{\mathcal{G}}_n,\|.\|_{Q,2},\epsilon\|G_{n}\|_{Q,2})\leq \left(\frac{A}{\epsilon}\right)^v,\qquad 0<\epsilon\leq 1
\end{equation*}
for some universal constants $A\geq e$ and $v\geq 1$, where the supremum is taken over all probability measures $Q$ with $0<\|G_{n}\|_{Q,2} <\infty$.
For an i.i.d. sequence of Gaussian random variables $(\varepsilon_i)_{i=1}^n$, consider the supremum of the bootstrapped process
equation*[equation* omitted — 132 chars of source]
The process $X_{ni}$ is typically unobserved. Let $\hat X_{ni}$ be its empirical counterpart estimated from the data and let $c_{1-\gamma}^*$ be $1-\gamma$ conditional on the data $\ensuremath{\mathcal{D}}_n$ quantile of the supremum of the bootstrapped process
equation*[equation* omitted — 131 chars of source]
where $(\varepsilon_i)_{i=1}^n$ are i.i.d. $N(0,1)$ random variables independent from the sample.
assumption(i) there exist constants $0<\underline\sigma<\bar{\sigma}$ such that $\underline{\sigma}^2\leq \ensuremath{\mathrm{Var}}(X_{ni}(w)) \leq \bar{\sigma}^2$ for all $w\in[0,1]^q$; (ii) $\sup_{g\in\ensuremath{\mathcal{G}}_n}\|g\|_{P,k}^r\leq \sigma^2b_n^{r-2}<\infty,r=2,3,4$ and $\|G_{n}\|_{P,4}\leq b_n < \infty$ for some $b_n=O(n^c)$.
Assumptions (ref) and (ref) consist of several technical conditions needed for the Gaussian and bootstrap approximation of chernozhukov2014gaussian.
assumptionLet $\ensuremath{\mathcal{P}}$ be the class of models consisting of all probability distributions corresponding to the ill-posed model satisfying Assumptions (ref), (ref), (ref). Suppose additionally that $\ensuremath{\mathcal{P}}$ is such that
\begin{equation*}
\begin{aligned}
(C1) &\quad \sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\left\|\hat T^* - T^*\right\|_{2,\infty}^2 = O(\delta_{1n}^2), \\
(C2) &\quad \sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\left\|\hat V_n^\varepsilon - V^\varepsilon_n\right\|_\infty = O(\delta_{2n}), \\
(C3) &\quad \hat r - \hat T\varphi = \frac{u_n}{n}\sum_{i=1}^nX_{ni} + R_{1n},
\end{aligned}
\end{equation*}
where $R_{1n}$ is such that $\sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\|R_{1n}\|^2 = O(\delta_{3n}^2)$ for some sequences $\delta_{1n},\delta_{2n},\delta_{3n}\to 0$.
Assumption (C3) is asymptotic linearization of residuals in the model. In some cases this expansion is exact and $R_{1n}$ is trivially zero. (C1) controls the estimation error in the operator, while (C2) controls the error in the bootstrap process.
We focus on bootstrap confidence sets described as
equation[equation omitted — 149 chars of source]
with $q_n^* = \frac{c_{1-\gamma}^*\|\hat T^*\|_{2,\infty} + \log^{-1}n}{\alpha_nn^{1/2}}u_n$.
In the following theorem, we show that our bootstrap confidence sets are honest to $\ensuremath{\mathcal{P}}$ and characterize coverage errors explicitly. This result will be specialized to each particular model in the following sections.
theoremSuppose that Assumptions (ref), (ref), (ref), and (ref) are satisfied and $b_n^4\log^7n/n\to 0$. Then
\begin{equation*}
\inf_{P\in\ensuremath{\mathcal{P}}}\Pr(\varphi\in C_{n,1-\gamma}^*) \geq 1 - \gamma - O_p\left(\delta_n\right) - o_p(1)
\end{equation*}
with
\begin{equation*}
\delta_n = \left(\alpha_n^{-1/2}\delta_{1n} + \alpha_n^{-1}\delta_{3n} + \alpha_n^{\beta\wedge 1}\right)u_n^{-1}\alpha_nn^{1/2}\log n + \delta_{2n}\log n.
\end{equation*}
proofUsing Assumption (ref) (C3), decompose
\begin{equation*}
\hat\varphi - \varphi = \hat\nu_n + (\alpha_nI + \hat T^*\hat T)^{-1}\hat T^*R_{1n} + R_{2n} + R_{3n}
\end{equation*}
with
\begin{equation*}
\begin{aligned}
\hat\nu_n & = (\alpha_n I + \hat T^*\hat T)^{-1}\hat T^*\frac{u_n}{n}\sum_{i=1}^nX_{ni} \\
R_{2n} & = \left[(\alpha_n I + \hat T^*\hat T)^{-1}\hat T^*\hat T - (\alpha_n I + T^*T)^{-1}T^*T\right]\varphi. \\
R_{3n} & = (\alpha_n I + T^*T)^{-1}T^*T\varphi - \varphi.
\end{aligned}
\end{equation*}
By Lemma (ref) there exists a constant $C<\infty$ independent from $P\in\ensuremath{\mathcal{P}}$ such that
\begin{equation*}
\|R_{2n}\|_\infty \leq \frac{C}{\alpha_n^{1/2}}\left(\|\hat T^* - T^*\|_{2,\infty} + \|\hat T^* - T^*\|_{2,\infty}^2\right),
\end{equation*}
whence under Assumption (ref) (C1)
\begin{equation*}
\sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\|R_{2n}\|_\infty = O\left(\frac{\delta_{1n} + \delta_{1n}^2}{\alpha_n^{1/2}}\right).
\end{equation*}
Next, under Assumption (ref) by Proposition (ref) there exists a constant $R$ independent from $P\in\ensuremath{\mathcal{P}}$ such that
\begin{equation*}
\|R_{3n}\|_\infty \leq R\alpha_n^{\beta\wedge 1}.
\end{equation*}
Lastly,
\begin{equation*}
\begin{aligned}
\left\|(\alpha_n I + \hat T^*\hat T)^{-1}\hat T^*R_{1n}\right\|_\infty & = \left\|\hat T^*(\alpha_n I + \hat T\hat T^*)^{-1}R_{1n}\right\|_\infty \\
& \leq \frac{\|\hat T^*\|_{2,\infty}}{\alpha_n}\|R_{1n}\|.
\end{aligned}
\end{equation*}
Then under Assumption (ref) by Markov's inequality
\begin{equation*}
\begin{aligned}
& \Pr\left(\varphi\in C_{n,1-\gamma}^*\right) \\
& = \Pr\left(\|\hat \varphi - \varphi\|_\infty\leq q_n^*\right) \\
& \geq \Pr\left(\|\hat\nu_n\|_\infty + \frac{\|\hat T^*\|_{2,\infty}}{\alpha_n}\|R_{1n}\| + \|R_{2n}\|_\infty + \|R_{3n}\|_\infty\leq q_n^* \right) \\
& \geq \Pr\left(\left\|V_{n}\right\|_\infty \leq c_{1-\gamma}^* + (2\|\hat T^*\|_{2,\infty}\log n)^{-1}\right) \\
& \qquad - \Pr\left(\frac{\|\hat T^*\|_{2,\infty}}{\alpha_n}\|R_{1n}\| + \|R_{2n}\|_\infty + \|R_{3n}\|_\infty > \frac{u_n}{2\alpha_nn^{1/2}\log n}\right) \\
& \geq \Pr\left(\left\|V_{n}\right\|_\infty \leq c_{1-\gamma}^* + (2\|\hat T^*\|_{2,\infty}\log n)^{-1}\right)\\
& \qquad - \sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\left[\frac{\|\hat T^*\|_{2,\infty}}{\alpha_n}\|R_{1n}\| + \|R_{2n}\|_\infty + \|R_{3n}\|_\infty\right]u_n^{-1}\alpha_nn^{1/2}\log n \\
& \geq \Pr\left(\left\|V_{n}\right\|_\infty \leq c_{1-\gamma}^* + (2\|\hat T^*\|_{2,\infty}\log n)^{-1}\right) \\
& \qquad - O\left(\left(\alpha_n^{-1}\delta_{3n} + \alpha_n^{-1/2}\delta_{1n} + \alpha_n^{\beta\wedge 1}\right)u_n^{-1}\alpha_nn^{1/2}\log n\right). \\
\end{aligned}
\end{equation*}
Under Assumptions (ref) and (ref) by the coupling inequality in chernozhukov2016empirical, Theorem 2.1, for any $P\in\ensuremath{\mathcal{P}}$ one can approximate the supremum of the empirical process $V_{n}$ by the supremum of the Gaussian process $\ensuremath{\mathbb{G}}_n$ with the same covariance structure
\begin{equation*}
\begin{aligned}
& \Pr\left(\|V_{n}\|_\infty\leq c^*_{1-\gamma} + (2\|\hat T^*\|_{2,\infty}\log n)^{-1}\right) \\
& \geq \Pr\left(\left\|\ensuremath{\mathbb{G}}_n\right\|_{\ensuremath{\mathcal{G}}_n} \leq c_{1-\gamma}^* + (2\|\hat T^*\|_{2,\infty}\log n)^{-1} - C_1\left(\frac{b_n\log^{5/4} n}{n^{1/4}} + \frac{b_n^{1/3}\log n}{n^{1/6}}\right)\right) \\
& \qquad - \Pr\left(\left|\|V_n\|_\infty - \|\ensuremath{\mathbb{G}}_n\|_{\ensuremath{\mathcal{G}}_n}\right| > C_1\left(\frac{b_n\log^{5/4} n}{n^{1/4}} + \frac{b_n^{1/3}\log n}{n^{1/6}}\right) \right) \\
& \geq \Pr\left(\left\|\ensuremath{\mathbb{G}}_n\right\|_{\ensuremath{\mathcal{G}}_n} \leq c_{1-\gamma}^* + (2\|\hat T^*\|_{2,\infty}\log n)^{-1} - C_1\left(\frac{b_n\log^{5/4} n}{n^{1/4}} + \frac{b_n^{1/3}\log n}{n^{1/6}}\right)\right) - o(1),
\end{aligned}
\end{equation*}
where $C_1$ is some absolute constants and the $o(1)$ term does not depend on $P\in\ensuremath{\mathcal{P}}$. Next, for some sequence of polynomial order $\delta_{4n}$ (to be specified below), put
\begin{equation*}
\begin{aligned}
\delta_n^I & = \delta_{4n} + \frac{b_n\log^{5/4} n}{n^{1/4}} + \frac{b_n^{1/3}\log n}{n^{1/6}} \\
\delta_n^{II} & = \left(\alpha_n^{-1}\delta_{3n} + \alpha_n^{-1/2}\delta_{1n} + \alpha_n^{\beta\wedge 1}\right)u_n^{-1}\alpha_nn^{1/2}\log n.
\end{aligned}
\end{equation*}
Then
\begin{equation*}
\begin{aligned}
\Pr\left(\varphi\in C^*_{n,1-\gamma}\right)
& \geq \Pr\left(\|\ensuremath{\mathbb{G}}_n\|_{\ensuremath{\mathcal{G}}_n} \leq c^*_{1-\gamma} + (2\|\hat T^*\|_{2,\infty}\log n)^{-1} + \delta_{4n}\right) \\
& - \sup_{x\in\ensuremath{\mathbf{R}}}\Pr\left(|\|\ensuremath{\mathbb{G}}_n\|_{\ensuremath{\mathcal{G}}_n} - x|\leq \delta_n^{I}\right) - O\left(\delta^{II}_n\right) - o(1).
\end{aligned}
\end{equation*}
Under Assumption (ref) (i) by the anti-concentration inequality in chernozhukov2014gaussian, Lemma A.1, for every $\epsilon>0$
\begin{equation*}
\sup_{x\in\ensuremath{\mathbf{R}}}\Pr\left(\left|\|\ensuremath{\mathbb{G}}_n\|_{\ensuremath{\mathcal{G}}_n} - x\right|\leq \epsilon\right)\leq C_2\epsilon\left(\ensuremath{\mathds{E}}\|\ensuremath{\mathbb{G}}_n\|_{\ensuremath{\mathcal{G}}_n} + \sqrt{1\vee \log(\sigma/\epsilon)}\right),
\end{equation*}
where the constant $C_2$ depends only on $\underline{\sigma}$ and $\bar\sigma$.
Moreover, under Assumption (ref) by the Dudley-Sudakov's entropy bound\footnote{Note that $N(\ensuremath{\mathcal{G}}_n,\|.\|_{P,2},\epsilon)=1$ for all $\epsilon\geq \mathrm{diam}(\ensuremath{\mathcal{G}}_n)/2$ and $\mathrm{diam}(\ensuremath{\mathcal{G}}_n)\leq 2\sup_{g\in\ensuremath{\mathcal{G}}_{n}}\|g\|_{P,2}\leq 2\bar\sigma$.}, see e.g. dudley2016vn
\begin{equation*}
\begin{aligned}
\ensuremath{\mathds{E}}\|\ensuremath{\mathbb{G}}_n\|_{\ensuremath{\mathcal{G}}_n} & \leq 24\int_0^{\bar{\sigma}}\sqrt{\log N(\ensuremath{\mathcal{G}}_n,\|.\|_{P,2},\epsilon)}\ensuremath{\mathrm{d}} \epsilon =O\left(\log^{1/2} n\right),
\end{aligned}
\end{equation*}
where the last equality follows under Assumptions (ref) and (ref) (ii).
Next under Assumptions (ref) and (ref) by the coupling inequality in chernozhukov2016empirical, Theorem 2.2, we approximate the supremum of the Gaussian process by the supremum of the multiplier bootstrap process. To that end for some constant $C_3$ independent from $P\in\ensuremath{\mathcal{P}}$
\begin{equation*}
\begin{aligned}
&\Pr\left(\|\ensuremath{\mathbb{G}}_n\|_\infty \leq c^*_{1-\gamma} + (2\|\hat T^*\|_{2,\infty}\log n)^{-1} + C_3\left(\frac{b_n\log^{5/4} n}{n^{1/4}} + \frac{b_n^{1/2}\log n}{n^{1/4}} \right)\right) \\
& \geq \Pr\left(\|V_{n}^\varepsilon\|_\infty \leq c^*_{1-\gamma} + (2\|\hat T^*\|_{2,\infty}\log n)^{-1}| \ensuremath{\mathcal{D}}_n\right) - o_p(1) \\
& \geq \Pr\left(\|\hat V_{n}^\varepsilon\|_\infty \leq c^*_{1-\gamma}|\ensuremath{\mathcal{D}}_n\right) - 2\|\hat T^*\|_{2,\infty}\ensuremath{\mathds{E}}\left[\left\|\hat V_n^\varepsilon - V_n^\varepsilon\right\|_\infty|\ensuremath{\mathcal{D}}_n\right]\log n - o_p(1) \\
\end{aligned}
\end{equation*}
where we use Markov's inequality, $\ensuremath{\mathcal{D}}_n$ denotes the data, and the $o_p(1)$ term does not depend on $P\in\ensuremath{\mathcal{P}}$. Setting $\delta_{4n} = C_3\left(\frac{b_n\log^{5/4} n}{n^{1/4}} + \frac{b_n^{1/2}\log n}{n^{1/4}}\right)$ under Assumption (ref) (C2) we obtain
\begin{equation*}
\Pr\left(\varphi\in C^*_{n,1-\gamma}\right) \geq 1-\gamma - O_p\left(\delta_n^I\log^{1/2} n + \delta_n^{II} + \delta_{2n}\log n\right) - o_p(1),
\end{equation*}
whence the result follows under the assumption $b_n^4\log^7n/n\to 0$.
To characterize expected diameters of bootstrap confidence sets we need additional assumptions.
assumptionSuppose that (i) conditionally on the data $\|\hat V_n^\varepsilon\|_\infty$ is a supremum over the class of functions satisfying Assumption (ref);
\begin{equation*}
(ii)\;\sup_{P\in\ensuremath{\mathcal{P}}}\max_{1\leq i\leq n}\ensuremath{\mathds{E}}_P\|\hat X_{ni} - X_{ni}\|^2_\infty = O(1);\qquad (iii)\; \sup_{P\in\ensuremath{\mathcal{P}}}\max_{1\leq i\leq n}\ensuremath{\mathds{E}}_P\|X_{ni}\|^2_\infty = O(u_n^2),
\end{equation*}
and $u_n\to \infty$.
theoremSuppose that Assumption (ref) and assumptions of Theorem (ref) are satisfied. If $\delta_{1n}=O(1)$, then
\begin{equation*}
\left|C^*_{n,1-\gamma}\right| = O_p\left(\frac{u_n^2}{\alpha_nn^{1/2}}\right)
\end{equation*}
uniformly over $P\in\ensuremath{\mathcal{P}}$.
proof[Proof of Theorem (ref)]
First note that
\begin{equation*}
\begin{aligned}
\left|C_{n,1-\gamma}^*\right|_\infty & = 2\left|q_n^*\right| \\
& = \frac{2u_n}{\alpha_nn^{1/2}}\left\{\left[c^*_{1-\gamma}\|\hat T^*\|_{2,\infty}\right] + \log^{-1}n \right\}.
\end{aligned}
\end{equation*}
Recall that $c^*_{1-\gamma}$ is $1-\gamma$ conditional on the data $\ensuremath{\mathcal{D}}_n$ quantile of the supremum of the Gaussian process $\|\hat V_n^\varepsilon\|_\infty$. By the Gaussian concentration inequality, e.g., see boucheron2013concentration, Theorem 5.8, conditionally on the data
\begin{equation*}
c_{1-\gamma}^* \leq \ensuremath{\mathds{E}}_P\left[\|\hat V_n^\varepsilon\|_\infty | \ensuremath{\mathcal{D}}_n\right] + \sqrt{\frac{2}{n}\sum_{i=1}^n\|\hat X_{ni}\|_\infty^2\log(1/\gamma)}.
\end{equation*}
By Dudley-Sudakov's inequality, e.g., see dudley2016vn, and the change of variables under Assumption (ref)
\begin{equation*}
\ensuremath{\mathds{E}}_P\left[\|\hat V_n^\varepsilon\|_\infty | \ensuremath{\mathcal{D}}_n\right] \leq 24\sqrt{\frac{1}{n}\sum_{i=1}^n\|\hat X_{ni}\|_\infty^2}\int_0^1\sqrt{v\log\frac{A}{\varepsilon}}\ensuremath{\mathrm{d}}\varepsilon.
\end{equation*}
The result follows under Assumption (ref) (C1).
Concentration-based confidence sets
The confidence sets based on the concentration inequality are described as
equation*[equation* omitted — 134 chars of source]
where $\hat q_n = 2\left\|\hat \nu_{n}^\eta\right\|_\infty + \frac{3\|\hat T^*\|_{2,\infty}G\sqrt{2\log(1/\gamma)} + \log^{-1} n}{\alpha_nn^{1/2}}u_n$ and
equation*[equation* omitted — 166 chars of source]
denotes the estimate of the supremum of the symmetrized process. The operator norm $\|\hat T^*\|_{2,\infty}$ can be computed by Lemma (ref), while $G$ is selected so that $\|X_{ni}\|\leq G$.
For concentration inequality we need a somewhat different version of Assumption (ref).
assumptionLet $\ensuremath{\mathcal{P}}$ be the class of models consisting of all probability distributions corresponding to the ill-posed model satisfying Assumption (ref). Suppose additionally that $\ensuremath{\mathcal{P}}$ is such that
\begin{equation*}
\begin{aligned}
(C1) &\quad \sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\left\|\hat T^* - T^*\right\|_{2,\infty}^2 = O(\delta_{1n}^2), \\
(C2) &\quad \sup_{P\in\ensuremath{\mathcal{P}}}\max_{1\leq i\leq n}\ensuremath{\mathds{E}}_P\left\|\hat X_{ni} - X_{ni}\right\|^2 = O(\delta_{2n}^2), \\
(C3) &\quad \hat r - \hat T\varphi = \frac{u_n}{n}\sum_{i=1}^nX_{ni} + R_{1n},
\end{aligned}
\end{equation*}
where $R_{1n}$ is a remainder term such that $\sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\|R_{1n}\|^2 = O(\delta_{3n}^2)$ and $\delta_{1n},\delta_{2n},\delta_{3n}\to 0$ are some sequences.
theoremSuppose that Assumption (ref) is satisfied and $\|X_{ni}\|\leq G$ for some universal constant $G$ for all $P\in\ensuremath{\mathcal{P}}$. Then
\begin{equation*}
\inf_{P\in\ensuremath{\mathcal{P}}}\Pr\left(\varphi\in C_{n,1-\gamma}\right) \geq 1 - \gamma - O(\delta_n)
\end{equation*}
with $\delta_n = \left(\frac{\delta_{3n}}{\alpha_n} + \frac{\delta_{1n}}{\alpha_n^{1/2}} + \alpha_n^{\beta\wedge 1}\right)u_n^{-1}\alpha_nn^{1/2}\log n + \delta_{2n}\log n$.
proof[Proof of Theorem (ref)]
Under Assumption (ref) (C3), decompose
\begin{equation*}
\hat\varphi - \varphi = \hat\nu_n + (\alpha_nI + \hat T^*\hat T)^{-1}\hat T^*R_{1n} + R_{2n} + R_{3n}
\end{equation*}
with
\begin{equation*}
\begin{aligned}
\hat\nu_n & = (\alpha_n I + \hat T^*\hat T)^{-1}\hat T^*\frac{u_n}{n}\sum_{i=1}^nX_{ni} \\
R_{2n} & = \left[(\alpha_n I + \hat T^*\hat T)^{-1}\hat T^*\hat T - (\alpha_n I + T^*T)^{-1}T^*T\right]\varphi. \\
R_{3n} & = (\alpha_n I + T^*T)^{-1}T^*T\varphi - \varphi.
\end{aligned}
\end{equation*}
Recall that $\hat q_n = 2\|\nu_n^\eta\|_\infty + \frac{3\|\hat T^*\|_{2,\infty}G\sqrt{2\log(1/\gamma)} + \log^{-1} n}{\alpha_n n^{1/2}}u_n$ and we also denote
\begin{equation*}
\begin{aligned}
\nu_n & = (\alpha_n I + T^*T)^{-1}T^*\frac{u_n}{n}\sum_{i=1}^nX_{ni}, \\
\nu_n^\eta & = (\alpha_n I + T^*T)^{-1}T^*\frac{u_n}{n}\sum_{i=1}^n\eta_iX_{ni},
\end{aligned}
\end{equation*}
where $\varepsilon_i$ are i.i.d. Rademacher random variables. Then by the triangle and Markov's inequalities
\begin{equation}
\begin{aligned}
& \Pr\left(\varphi\in C_{n,1-\gamma}\right) \\
& = \Pr\left(\|\hat \varphi - \varphi\|_\infty\leq \hat q_{n}\right) \\
& \geq \Pr\left(\|\hat \nu_{n}\|_\infty + \frac{\|\hat T^*\|_{2,\infty}}{\alpha_n}\|R_{1n}\| + \|R_{2n}\|_\infty + \|R_{3n}\|_\infty \leq \hat q_n\right) \\
& \geq \Pr\left(\|\nu_{n}\|_\infty + \left\|\hat \nu_{n} - \nu_{n}\right\|_\infty + \frac{\|\hat T^*\|_{2,\infty}}{\alpha_n}\|R_{1n}\| + \|R_{2n}\|_\infty + \|R_{3n}\|_\infty\leq \hat q_n\right) \\
& \geq \Pr\left(\|\nu_{n}\|_\infty \leq 2\|\nu_n^\eta\|_\infty + \frac{3\|T^*\|_{2,\infty}G\sqrt{2\log(1/\gamma)}}{\alpha_nn^{1/2}}u_n\right) \\
& \qquad - \Pr\left(\left\|\hat \nu_{n} - \nu_{n}\right\|_\infty + \frac{\|\hat T^*\|_{2,\infty}}{\alpha_n}\|R_{1n}\| + \|R_{2n}\|_\infty + \|R_{3n}\|_\infty > \frac{u_n}{2\alpha_nn^{1/2}\log n}\right) \\
& \qquad - \Pr\left(\left\|\hat \nu_{n}^\eta - \nu_{n}^\eta\right\|_\infty + \frac{3G\sqrt{2\log(1/\gamma)}}{\alpha_nn^{1/2}}u_n\|\hat T^* - T^*\|_{2,\infty} > \frac{u_n}{2\alpha_nn^{1/2}\log n}\right) \\
& \geq \Pr\left(\|\nu_{n}\|_\infty \leq 2\|\nu_n^\eta\|_\infty + \frac{3\|T^*\|_{2,\infty}G\sqrt{2\log(1/\gamma)}}{\alpha_nn^{1/2}}u_n\right) \\
& \qquad - 2\sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\left[\left\|\hat \nu_{n} - \nu_{n}\right\|_\infty + \frac{\|\hat T^*\|_{2,\infty}}{\alpha_n}\|R_{1n}\| + \|R_{2n}\|_\infty + \|R_{3n}\|_\infty\right]u_n^{-1}\alpha_nn^{1/2}\log n \\
& \qquad - 2\sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\left[\left\|\hat \nu_{n}^\eta - \nu_{n}^\eta\right\|_\infty u_n^{-1}\alpha_nn^{1/2}\log n + 3G\sqrt{2\log(1/\gamma)}\|\hat T^* - T^*\|_{2,\infty}\log n\right]. \\
\end{aligned}
\end{equation}
Under Assumption (ref) (C1)-(C2) by Lemma (ref)
\begin{equation*}
\begin{aligned}
\sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\left\|\hat \nu_{n} - \nu_{n}\right\|_\infty & = O\left(\frac{u_n\delta_{1n}}{\alpha_n^{3/2}n^{1/2}}\right), \\
\sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\left\|\hat \nu_{n}^\eta - \nu_{n}^\eta\right\|_\infty & = O\left(\frac{u_n\delta_{1n}}{\alpha_n^{3/2}n^{1/2}} + \frac{u_n\delta_{2n}}{\alpha_nn^{1/2}}\right).
\end{aligned}
\end{equation*}
Therefore, another application of Assumption (ref) and the order of $R_{2n}$ and $R_{3n}$ from the proof of Theorem (ref) imply
\begin{equation*}
\begin{aligned}
\Pr\left(\varphi\in C_{n,1-\gamma}\right) & \geq \Pr\left(\|\nu_{n}\|_\infty \leq 2\|\nu_n^\eta\|_\infty + \frac{3\|T^*\|_{2,\infty}G\sqrt{2\log(1/\gamma)}}{\alpha_nn^{1/2}}u_n\right) \\
& \qquad - O\left(\left(\frac{\delta_{3n}}{\alpha_n} + \frac{\delta_{1n}}{\alpha_n^{1/2}} + \alpha_n^{\beta\wedge 1}\right)u_n^{-1}\alpha_nn^{1/2}\log n + \delta_{2n}\log n \right).
\end{aligned}
\end{equation*}
Next, $\|\nu_{n}\|_\infty$ is the supremum of empirical processes indexed by the class of functions
\begin{equation*}
\ensuremath{\mathcal{H}}_n = \left\{h:x\in C[0,1]^q\to u_n\left[(\alpha_nI + T^*T)^{-1}T^*x\right](z),\;z\in [0,1]^p\right\}.
\end{equation*}
For any $h\in\ensuremath{\mathcal{H}}_n$
\begin{equation*}
\|h(X_{ni})\|_\infty \leq u_n\|T^*\|_{2,\infty}\|(\alpha_nI + TT^*)^{-1}\|\|X_{ni}\| \leq \frac{\|T^*\|_{2,\infty}}{\alpha_n}Gu_n \triangleq H_n.
\end{equation*}
So $H_n$ is the envelope for the class $\ensuremath{\mathcal{H}}_n$. Then by Proposition (ref)
\begin{equation*}
\begin{aligned}
\Pr(\|\nu_n\|_\infty\leq \hat q_n) & = \Pr\left(\|\nu_n\|_\infty\leq 2\|\nu_n^\eta\|_\infty + 3H_n\sqrt{\frac{2\log(1/\gamma)}{n}} \right) \\
& \geq 1 - \gamma,
\end{aligned}
\end{equation*}
whence the result.
The next result describes the convergence rate of the expected diameter of the confidence set based on the concentration inequality.
theoremSuppose that assumptions of Theorem (ref) are satisfied. If $\delta_{1n}/\alpha_n^{1/2}=O(1)$, then
\begin{equation*}
\sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\left|C_{n,1-\gamma}\right|_\infty = O\left(\frac{u_n}{\alpha_n n^{1/2}}\right).
\end{equation*}
proof[Proof of Theorem (ref)]
Since $\ensuremath{\mathds{E}}_P\left|C_{n,1-\gamma}\right|_\infty = 2\ensuremath{\mathds{E}}_P \left|\hat q_n\right|$, under assumptions of Theorem (ref), there exists some universal constant $C$ such that
\begin{equation*}
\ensuremath{\mathds{E}}_P\left|C_{n,1-\gamma}\right|_\infty \leq C\left\{\ensuremath{\mathds{E}}_P\|\nu_n^\eta\|_\infty + \ensuremath{\mathds{E}}_P\|\hat \nu_n^\eta - \nu_n^\eta\|_\infty + \frac{u_n}{\alpha_nn^{1/2}}\ensuremath{\mathds{E}}_P\|\hat T^*\|_{2,\infty} \right\}.
\end{equation*}
By Jensen's inequality, since $\|X_{ni}\|\leq C$
\begin{equation*}
\begin{aligned}
\ensuremath{\mathds{E}}_P\|\nu_n^\eta\|_\infty & \leq \frac{\|T^*\|_{2,\infty}}{\alpha_n}\ensuremath{\mathds{E}}_P\left\|\frac{u_n}{n}\sum_{i=1}^n\eta_iX_{ni}\right\| \\
& \leq \frac{u_n\|T^*\|_{2,\infty}}{\alpha_nn^{1/2}}\sqrt{\ensuremath{\mathds{E}}_P\|X_{ni}\|^2} \\
& \leq \frac{u_n\|T^*\|_{2,\infty}}{\alpha_nn^{1/2}}G.
\end{aligned}
\end{equation*}
Therefore by Lemma (ref) under Assumptions (ref) and (ref) (C1)-(C2)
\begin{equation*}
\sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\left|C_{n,1-\gamma}\right|_\infty = O\left(\frac{u_n}{\alpha_n n^{1/2}} + \frac{u_n\delta_{1n}}{\alpha_n^{3/2}n^{1/2}} + \frac{u_n\delta_{2n}}{\alpha_nn^{1/2}}\right).
\end{equation*}
NPIV, functional regressions, and density deconvolution
NPIV model
propositionSuppose that Assumption (ref) (i)-(iii) is satisfied and the sequence of bandwidth parameters is such that $1/(nh_n^{p+q})=O(1)$, then for all $1<r<\infty$
\begin{equation*}
\left(\ensuremath{\mathds{E}}\left\|\hat T^* - T^*\right\|_{2,\infty}^r\right)^{1/r} \leq C\left\{\sqrt{\frac{\log h_n^{-1}}{nh_n^{p+q}}} + h_n^{s}\right\},
\end{equation*}
where the constant $C$ depends only on $s,M,r,\bar f,\|k\|_\infty$.
proofBy Lemma (ref),
\begin{equation*}
\begin{aligned}
\ensuremath{\mathds{E}}\left\|\hat T^* - T^*\right\|_{2,\infty}^r & = \ensuremath{\mathds{E}}\left[\sup_{z\in(0,1)^p}\left(\int_{[0,1]^q}\left|\hat f_{ZW}(z,w) - f_{ZW}(z,w)\right|^2\ensuremath{\mathrm{d}} w\right)^{1/2}\right]^r \\
& \leq \ensuremath{\mathds{E}}\left\|\hat f_{ZW} - f_{ZW}\right\|_\infty^{r}. \\
\end{aligned}
\end{equation*}
By Minkowski inequality,
\begin{equation*}
\left(\ensuremath{\mathds{E}}\left\|\hat f_{ZW} - f_{ZW}\right\|_\infty^{r}\right)^{1/r} \leq \left(\ensuremath{\mathds{E}}\left\|\hat f_{ZW} - \ensuremath{\mathds{E}}\hat f_{ZW}\right\|^r_\infty\right)^{1/r} + \left(\left\|\ensuremath{\mathds{E}}\hat f_{ZW} - f_{ZW}\right\|_\infty^r\right)^{1/r}.
\end{equation*}
The order of the bias follows by standard computations under Assumption (ref) (i) and (iii), see tsybakov2009introduction
\begin{equation*}
\left\|\ensuremath{\mathds{E}}\hat f_{ZW} - f_{ZW}\right\|_\infty = O\left(h_n^{s}\right),
\end{equation*}
where the big-O term is independent from the distribution of $(Z,W)$. For the variance term, we apply the moment inequality in gine2015mathematical, see Theorem 5.1.5 and Theorem 5.1.15, which gives
\begin{equation*}
\begin{aligned}
\left(\ensuremath{\mathds{E}}\left\|\hat f_{ZW} - \ensuremath{\mathds{E}}\hat f_{ZW}\right\|^r_\infty\right)^{1/r} = O\left(\sqrt{\frac{\log h_n^{-1}}{nh_n^{p+q}}} + \sqrt{\frac{1}{nh_n^{p+q}}} + \frac{1}{nh_n^{p+q}}\right).
\end{aligned}
\end{equation*}
Combining all estimates we obtain the result.
proof[Proof of Theorem (ref)]
We shall verify conditions of Theorems (ref) and (ref). We show first that in the nonparametric IV model Assumption (ref) (C3) is verified with $u_n=h_n^{-q/2}$ and $X_{ni}(w) = U_ih_n^{-q/2}K_w\left(h_n^{-1}(W_i-w)\right)$, where $[K_z\ast\varphi](Z) = h_n^{-p}\int\varphi(z)K_z\left(h_n^{-1}(Z-z)\right)\ensuremath{\mathrm{d}} z$. Indeed, it is easy to see that
\begin{equation*}
\begin{aligned}
(\hat r - \hat T\varphi)(w) & = \frac{h_n^{-q/2}}{n}\sum_{i=1}^nX_{ni}(w) + R_{1n}(w) \\
R_{1n}(w) & = \frac{1}{nh_n^q}\sum_{i=1}^n\left\{\varphi(Z_i) - [K_z\ast\varphi](Z_i)\right\}K_w\left(h_n^{-1}(W_i - w)\right).
\end{aligned}
\end{equation*}
Under Assumption (ref), $\varphi\in C_M^t[0,1]^p$, so that using standard bias computations, e.g. see tsybakov2009introduction, under Assumption (ref) (iii) there exists some constant $C<\infty$ such that
\begin{equation*}
\begin{aligned}
\left\|\varphi - [K_z\ast \varphi]\right\|_\infty \leq Ch_n^t,
\end{aligned}
\end{equation*}
Then under Assumption (ref) (i) and (ii)
\begin{equation*}
\begin{aligned}
\sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\|R_{1n}\|^2 & \leq
\sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\left\|\frac{1}{n}\sum_{i=1}^n\left|h_n^{-q}K_w\left(h_n^{-1}(W_i - .)\right)\right|\right\|^2 Ch_n^t \\
& \leq \sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\left\|h_n^{-q}K_w\left(h_n^{-1}(W_i - .)\right)\right\|^2Ch_n^t \\
& = O(h_n^{t-q}).
\end{aligned}
\end{equation*}
Therefore, Assumption (ref) (C3) is verified with $\delta_{3n} = h_n^{(t-q)/2}$.
Consider now the class of functions
\begin{equation*}
\ensuremath{\mathcal{G}}_{n} = \left\{(u,w)\mapsto u\frac{1}{h_n^{q/2}}K_w\left(h_n^{-1}(w-v)\right):\; v\in[0,1]^q\right\}.
\end{equation*}
Under Assumption (ref) (iii) and (iv) this class admits a square-integrable envelope $G_{n}(u,w) = h_n^{-q/2}F\|K\|_\infty$ with
\begin{equation*}
\|G_{n}\|_{P,2} = h_n^{-q/2}\|K\|_\infty F
\end{equation*}
for all $P\in\ensuremath{\mathcal{P}}$. Recall that $K_w$ is a product kernel of functions of bounded variation under Assumption (ref) (iii) and that the class of translations of functions of bounded variation is of the VC-type, so Assumption (ref) holds.
Next, for any $g_n\in\ensuremath{\mathcal{G}}_n$
\begin{equation*}
\ensuremath{\mathrm{Var}}\left(g_n(v)\right) = \ensuremath{\mathds{E}}\left|Uh_n^{-q/2}K_w\left(h_n^{-1}(W-v)\right)\right|^2,
\end{equation*}
so that by the law of iterated expectation and change of variables under Assumption (ref) (ii) and (iv)
\begin{equation*}
\sigma_2\|k\|^{2q}f \leq \ensuremath{\mathrm{Var}}\left(g_n(v)\right) \leq F\|k\|^{2q}\bar f.
\end{equation*}
This verifies Assumption (ref) (i). Next, under Assumption (ref) (iv)
\begin{equation*}
\begin{aligned}
\|G\|_{P,4} & = h_n^{-q/2}\|k\|_\infty^q\|U\|_{P,4} \\
& \leq Fh_n^{-q/2}\|k\|_\infty^q \\
P|g_n|^3 & = h_n^{-3q/2}\ensuremath{\mathds{E}}\left|UK_w\left(h_n^{-1}(W-v)\right)\right|^3 \\
& \leq Fh_n^{-q/2}\|k\|^{3q}_3 \\
P|g_n|^4 & = h_n^{-4q/2}\ensuremath{\mathds{E}}\left|UK_w\left(h_n^{-1}(W-v)\right)\right|^4 \\
& \leq Fh_n^{-q}\|k\|^{4q}_4.
\end{aligned}
\end{equation*}
Therefore Assumption (ref) (ii) is satisfied with $b_n \leq Ch_n^{-q/2}$ and some universal constant $C<\infty$. Note that $b_n^4\log^7 n/n\to 0$ under Assumption (ref) (vi) (c).
It remains to verify Assumption (ref) (C1) and (C2). The former follows from the Proposition (ref) under Assumption (ref) (i)-(iii) with $\delta_{1n} = \sqrt{\frac{\log h_n^{-1}}{nh_n^{p+q}}} + h_n^s$. Lastly, conditionally on the data $\ensuremath{\mathcal{D}}_n$
\begin{equation*}
\begin{aligned}
\ensuremath{\mathds{E}}_\varepsilon\left\|\hat V_n^\varepsilon - V_n^\varepsilon\right\|_\infty & = \ensuremath{\mathds{E}}_\varepsilon\left\|\frac{1}{\sqrt{n}}\sum_{i=1}^n\varepsilon_i(\hat{\varphi}(Z_i) - \varphi(Z_i))h_n^{-q/2}K_w\left(h_n^{-1}(W_i-.)\right)\right\|_\infty \\
\end{aligned}
\end{equation*}
is the expected value of the supremum of the Gaussian process indexed by the VC-type class. By Dudley-Sudakov's inequality, see dudley2016vn, and change of variables
\begin{equation*}
\begin{aligned}
\ensuremath{\mathds{E}}_\varepsilon\left\|\hat V_n^\varepsilon - V_n^\varepsilon\right\|_\infty & \leq 24\int_0^{\|H_n\|_{P_n,2}}\sqrt{\log N(\ensuremath{\mathcal{H}}_n,\|.\|_{P_n,2},\epsilon)}\ensuremath{\mathrm{d}}\epsilon \\
& \leq 24\|H_n\|_{P_n,2}\int_0^1\sqrt{v\log\left(\frac{A}{\epsilon}\right)}\ensuremath{\mathrm{d}}\epsilon,
\end{aligned}
\end{equation*}
where
\begin{equation*}
\begin{aligned}
\|H_n\|_{P_n,2}^2 & = \frac{1}{n}\sum_{i=1}^n(\hat\varphi(Z_i) - \varphi(Z_i))^2h_n^{-q}\|k\|_\infty^{2q}.
\end{aligned}
\end{equation*}
Therefore
\begin{equation*}
\sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\left\|\hat V_n^\varepsilon - V_n^\varepsilon\right\|_\infty = \sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\left\|\hat\varphi - \varphi\right\|_\infty O\left(h_n^{-q/2}\right).
\end{equation*}
Under Assumptions (ref) and (ref) (i)-(iv), it follows from babiiflorens2016b, Theorem 4.2, that
\begin{equation*}
\sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\left\|\hat\varphi - \varphi\right\|_\infty = O\left(\frac{1}{\alpha_n}\left(\frac{1}{\sqrt{nh_n^q}} + h_n^{s}\right) + \frac{1}{\alpha_n^{1/2}}\sqrt{\frac{\log h_n^{-1}}{nh_n^{p+q}}} + \alpha_n^{\beta\wedge 1}\right).
\end{equation*}
Therefore, Assumption (ref) (C2) is verified with
\begin{equation*}
\delta_{2n} = O\left(\frac{1}{\alpha_n}\left(\frac{1}{n^{1/2}h_n^q} + h_n^{s-q/2}\right) + \frac{1}{\alpha_n^{1/2}}\frac{\log^{1/2} h_n^{-1}}{n^{1/2}h_n^{p/2+q}} + \alpha_n^{\beta\wedge 1}h_n^{-q/2}\right).
\end{equation*}
Combining all estimates, it follows from Theorem (ref) that if $\frac{\log^7n}{nh_n^{2q}}\to 0$
\begin{equation*}
\inf_{P\in\ensuremath{\mathcal{P}}}\Pr(\varphi\in C_{n,1-\gamma}^*) \geq 1 - \gamma - O_p\left(\delta_n\right) - o_p(1)
\end{equation*}
with
\begin{equation*}
\delta_n = \left(\alpha_n^{-1/2}\delta_{1n} + \alpha_n^{-1}\delta_{3n} + \alpha_n^{\beta\wedge 1}\right)u_n^{-1}\alpha_nn^{1/2}\log n + \delta_{2n}\log n.
\end{equation*}
The result then follows under Assumption (ref) (vi) after taking the probability limit.
For confidence sets based on the concentration inequality, we note that Assumptions (ref) (C1) and (C3) hold with the same $\delta_{1n}$ and $\delta_{3n}$ sequences. Note also that $\|X_{ni}\|\leq F\|k\|^q\triangleq G$. It only remains to verify Assumption (ref) (C2). To that end note that
\begin{equation*}
\begin{aligned}
\ensuremath{\mathds{E}}_P\left\|\hat X_{ni} - X_{ni}\right\|^2 & = \ensuremath{\mathds{E}}_P\left\|\frac{1}{h_n^{q/2}}(\hat\varphi(Z_i) - \varphi(Z_i))K_w(h_n^{-1}(W_i - .))\right\|^2 \\
& \leq \ensuremath{\mathds{E}}_P\|\hat\varphi - \varphi\|^2_\infty\|K_w\|^2.
\end{aligned}
\end{equation*}
Inspection of the proof of the uniform risk bound in babiiflorens2016b reveals that for all $r\geq 2$ we also have
\begin{equation}
\left(\sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\|\hat\varphi - \varphi\|_\infty^r\right)^{1/r} = O\left(\frac{1}{\alpha_n}\left(\frac{1}{\sqrt{nh_n^q}} + h_n^{s}\right) + \frac{1}{\alpha_n^{1/2}}\sqrt{\frac{\log h_n^{-1}}{nh_n^{p+q}}} + \alpha_n^{\beta\wedge 1}\right),
\end{equation}
provided that
\begin{equation*}
\sup_{P\in\ensuremath{\mathcal{P}}}\left(\ensuremath{\mathds{E}}_P\left\|\hat T^* - T^*\right\|_{2,\infty}^{2r}\right)^{\frac{1}{2r}} = O\left(\sqrt{\frac{\log h_n^{-1}}{nh_n^{p+q}}}\right)
\end{equation*}
and
\begin{equation}
\sup_{P\in\ensuremath{\mathcal{P}}}\left(\|\hat r - \ensuremath{\mathds{E}}\hat r\|^{2r}\right)^{\frac{1}{2r}} = O\left(\frac{1}{\sqrt{nh_n^q}} + h_n^s\right).
\end{equation}
The former statement is established in Proposition (ref), while for the latter we use the following inequality valid for any i.i.d. sequence $(\xi_{ni})_{i=1}^n$
\begin{equation*}
\left(\ensuremath{\mathds{E}}\left\|\frac{1}{n}\sum_{i=1}^n\xi_{ni}\right\|^{r}\right)^\frac{1}{r} \leq C\left\{n^{-1/2}\left(\ensuremath{\mathds{E}}\|\xi_{ni}\|^2\right)^{1/2} + n^{1/r-1}\left(\ensuremath{\mathds{E}}\|\xi_{ni}\|^r \right)^{1/r} \right\}.
\end{equation*}
In particular, setting $\xi_{ni} = Y_ih_n^{-q}K_w(h_n^{-1}(W_i - w)) - \ensuremath{\mathds{E}}\left[Y_ih_n^{-q}K_w(h_n^{-1}(W_i - w))\right]$, under Assumption (ref) (iv) we obtain Eq. (ref) with $r=2$. Therefore, Assumption (ref) (C2) is verified with
\begin{equation*}
\delta_{2n} = \left(\frac{1}{\alpha_n}\left(\frac{1}{\sqrt{nh_n^q}} + h_n^s\right) + \frac{1}{\alpha_n^{1/2}}\sqrt{\frac{\log h_n^{-1}}{nh_n^{p+q}}} + \alpha_n^{\beta \wedge 1}\right)
\end{equation*}
and by Theorem (ref)
\begin{equation*}
\inf_{P\in\ensuremath{\mathcal{P}}}\Pr\left(\varphi\in C_{n,1-\gamma}\right) \geq 1 - \gamma - O(\delta_n)
\end{equation*}
with
\begin{equation*}
\delta_n = \delta_{2n}\log n + \left(\frac{\delta_{1n}}{\alpha_n^{1/2}} + \frac{\delta_{3n}}{\alpha_n} + \alpha_n^{\beta\wedge 1}\right)u_n^{-1}\alpha_nn^{1/2}\log n.
\end{equation*}
proof[Proof of Theorem (ref)]
For concentration-based confidence sets we apply Theorem (ref). In the proof of Theorem (ref) we show that $\delta_{1n} = \sqrt{\frac{\log n}{nh_n^{p+q}}} + h_n^{s}$. Under Assumption (ref) (vi) $\delta_{1n}/\alpha_n^{1/2}\to 0$, whence since $u_n = h_n^{-q/2}$, by Theorem (ref)
\begin{equation*}
\sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\left|C_{n,1-\gamma}\right|_\infty = O\left(\frac{1}{\alpha_n\sqrt{nh_n^q}}\right).
\end{equation*}
Next we verify Assumption (ref). Note that (i) holds, since the class is described by the convolution kernel of bounded variation. For (ii) we note that
\begin{equation*}
\ensuremath{\mathds{E}}_P\left\|\hat X_{ni} - X_{ni}\right\|^2_\infty \leq \ensuremath{\mathds{E}}_P\|\hat\varphi - \varphi\|^2_\infty\|K_w\|^2_\infty h_n^{-q},
\end{equation*}
which tends to zero due to Eq. (ref) and Assumption (ref) (vi). Lastly, (iii) follows since $\ensuremath{\mathds{E}}_P\|X_{ni}\|^2_\infty \leq F\|K_w\|_\infty^2h_n^{-q}$. Therefore, we obtain the second statement by Theorem (ref).
Functional regressions
proof[Proof of Theorem (ref)]
We verify assumptions of Theorems (ref) and (ref). Under Assumption (ref) (ii), by Hoffman-J\o rgensen's inequality, e.g., see gine2015mathematical, Theorem 3.1.15,
\begin{equation*}
\begin{aligned}
\left(\ensuremath{\mathds{E}}_P\left\|\hat T^* - T^*\right\|_{2,\infty}^2\right)^{1/2} & \leq \left(\ensuremath{\mathds{E}}_P\left\|\frac{1}{n}\sum_{i=1}^nZ_iW_i - \ensuremath{\mathds{E}} [Z_iW_i]\right\|_{\infty}^2\right)^{1/2} \\
& = 12\sqrt{3}\left(16\ensuremath{\mathds{E}}_P\left\|\frac{1}{n}\sum_{i=1}^nZ_iW_i - \ensuremath{\mathds{E}} [Z_iW_i]\right\|_{\infty} + \frac{2FC}{n}\right).\\
\end{aligned}
\end{equation*}
Next, under Assumption (ref) (iii), by the bracketing moment inequality, e.g., see gine2015mathematical, Propositions 3.5.15 and 4.3.36,
\begin{equation*}
\sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\left\|\frac{1}{n}\sum_{i=1}^nZ_iW_i - \ensuremath{\mathds{E}}[Z_iW_i]\right\|_\infty = O\left(\frac{1}{\sqrt{n}}\right).
\end{equation*}
Therefore, Assumptions (ref) (C1) and (ref) (C1) are satisfied with $\delta_{1n}=1/\sqrt{n}$.
Since
\begin{equation*}
\hat r - \hat T\varphi = \frac{1}{n}\sum_{i=1}^nU_iW_i,
\end{equation*}
Assumptions (ref) (C3) and (ref) (C3) are trivially satisfied with $X_{ni} = U_iW_i$, $u_n = 1$, and $\delta_{3n}=0$.
Note also that by the Cauchy-Schwartz inequality
\begin{equation*}
\begin{aligned}
\ensuremath{\mathds{E}}_P\left\|\hat X_{ni} - X_{ni}\right\|^2 & = \ensuremath{\mathds{E}}_P\left\|\frac{1}{n}\sum_{i=1}^n\langle\hat\varphi - \varphi,Z_i\rangle W_i\right\|^2 \\
& \leq \ensuremath{\mathds{E}}_P\|\hat\varphi - \varphi\|^2.
\end{aligned}
\end{equation*}
Under Assumptions (ref) (i)-(ii), it follows from babiiflorens2016b that
\begin{equation}
\sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\|\hat\varphi - \varphi\|^2 = O\left(\frac{1}{\alpha_nn} + \alpha_n^{(2\beta+1)\wedge 2}\right),
\end{equation}
whence Assumption (ref) (C2) is verified with $\delta_{2n} = \frac{1}{\sqrt{\alpha_nn}} + \alpha_n^{(\beta+1/2)\wedge 1}$. Lastly, $\|X_{ni}\|\leq 1 \triangleq G$. Then by Theorem (ref) for confidence sets based on the concentration inequality, we obtain
\begin{equation*}
\inf_{P\in\ensuremath{\mathcal{P}}}\mathrm{Pr}\left(\varphi\in C_{n,1-\gamma}\right) \geq 1 - \gamma - O\left(\left(\alpha_n^{\beta\wedge 1+1}n^{1/2} + \frac{1}{\sqrt{\alpha_nn}} + \alpha_n^{\beta/2\wedge 1}\right)\log n\right).
\end{equation*}
For bootstrap confidence sets we need additionally to verify Assumptions (ref), (ref), and (ref) (C2). The relevant class of functions is $\ensuremath{\mathcal{G}} = \left\{g_s(u,w) = uw(s):\; s\in[0,1]^p \right\}$.
This class is of VC-type under Assumption (ref) (iii), e.g., see gine2015mathematical, Proposition 3.6.12. Therefore Assumption (ref) is satisfied. Next under Assumption (ref) (iv)
\begin{equation*}
\ensuremath{\mathrm{Var}}(g_s) = \ensuremath{\mathds{E}}|UW(s)|^2
\end{equation*}
is uniformly bounded away from zero and infinity, which verifies Assumption (ref) (i). Next, under Assumption (ref) (ii)
\begin{equation*}
\begin{aligned}
\|G\|_{P,4} & = F \\
P|g_s|^3 & = \ensuremath{\mathds{E}}\left|UW(s)\right|^3 \leq F^3 \\
P|g_s|^4 & = \ensuremath{\mathds{E}}\left|UW(s)\right|^4 \leq F^4.
\end{aligned}
\end{equation*}
Therefore Assumption (ref) (ii) is satisfied with universal constants $b,\sigma$. Therefore, $b^4\log^7n/n\to 0$.
Lastly, note that for some universal constant $C$ depending only on $\|r\|_\infty$ and $K$
\begin{equation*}
\begin{aligned}
\ensuremath{\mathds{E}}_P\left\|\hat V_n^\varepsilon - V_n^\varepsilon \right\|_\infty & = \ensuremath{\mathds{E}}_P\ensuremath{\mathds{E}}_\varepsilon\left\|\frac{1}{\sqrt{n}}\sum_{i=1}^n\varepsilon_i\langle \hat\varphi - \varphi,Z_i\rangle W_i \right\|_\infty \\
& \leq CF\ensuremath{\mathds{E}}_P\|\hat\varphi - \varphi\|,
\end{aligned}
\end{equation*}
where the inequality follows by Dudley-Sudakov's bound. Therefore, Assumption (ref) (C2) is verified with the same $\delta_{2n} = \frac{1}{\sqrt{\alpha_nn}} + \alpha_n^{\beta/2\wedge 1}$. Then, by Theorem (ref)
\begin{equation*}
\inf_{P\in\ensuremath{\mathcal{P}}}\Pr(\varphi\in C_{n,1-\gamma}^*) \geq 1 - \gamma - O_p\left(\left(\alpha_n^{\beta\wedge 1+1}n^{1/2} + \frac{1}{\sqrt{\alpha_nn}} + \alpha_n^{\beta/2\wedge 1}\right)\log n\right) - o_p(1)
\end{equation*}
and the result follows by taking the probability limit.
proof[Proof of Theorem (ref)]
For concentration-based confidence sets we apply Theorem (ref). In the proof of Theorem (ref) we show that $\delta_{1n}=n^{-1/2}$. Under Assumption (ref) (vi), $\delta_{1n}/\alpha_n^{1/2}\to 0$ and $\delta_{2n}\to 0$, whence since $u_n = 1$, by Theorem (ref)
\begin{equation*}
\sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\left|C_{n,1-\gamma}\right|_\infty = O\left(\frac{1}{\alpha_nn^{1/2}}\right).
\end{equation*}
Next we verify Assumption (ref). (i) holds, under Assumption (ref) (iii). For (ii), under Assumption (ref) (ii)
\begin{equation*}
\ensuremath{\mathds{E}}_P\left\|\hat X_{ni} - X_{ni}\right\|^2_\infty \leq CF\ensuremath{\mathds{E}}_P\|\hat\varphi - \varphi\|^2,
\end{equation*}
which tends to zero due to Eq. (ref) and Assumption (ref) (vi). Lastly, (iii) follows trivially under Assumption (ref) (ii). Therefore, we obtain the second statement by Theorem (ref).
Density deconvolution
proof[Proof of Theorem (ref)]
We verify assumptions of Theorems (ref) and (ref). Note that Assumptions (ref) (C1) and (ref) (C1) are trivially satisfied with $\delta_{1n}=0$. Next,
\begin{equation*}
\hat r - T\varphi = \frac{1}{nh_n}\sum_{i=1}^nK\left(\frac{Y_i - y}{h_n}\right) - \frac{1}{h_n}\ensuremath{\mathds{E}}\left[K\left(\frac{Y_i - y}{h_n}\right)\right] + R_{1n}
\end{equation*}
with
\begin{equation*}
\begin{aligned}
\left\|R_{1n}\right\|^2 & = \int_0^1\left|\frac{1}{h_n}\ensuremath{\mathds{E}}\left[K\left(\frac{Y_i - y}{h_n}\right)\right] - r(y)\right|^2\ensuremath{\mathrm{d}} y.
\end{aligned}
\end{equation*}
Under Assumption (ref) (i)-(iv) by tsybakov2009introduction, Proposition 1.5, we have
\begin{equation*}
\sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\|R_{1n}\|^2 = O(h_n^{2s}).
\end{equation*}
Therefore, Assumptions (ref) (C3) and (ref) (C3) are verified with $X_{ni} = \frac{1}{h_n^{1/2}}K\left(\frac{Y_i - y}{h_n}\right) - \frac{1}{h_n^{1/2}}\ensuremath{\mathds{E}}\left[K\left(\frac{Y_i - y}{h_n}\right)\right]$, $u_n = h_n^{-1/2}$, and $\delta_{3n}=h_n^s$.
Note also that
\begin{equation*}
\begin{aligned}
\ensuremath{\mathds{E}}_P\left\|\hat X_{ni} - X_{ni}\right\|^2 & = \ensuremath{\mathds{E}}_P\left\|\frac{1}{nh_n^{1/2}}\sum_{i=1}^nK\left(\frac{Y_i - .}{h_n}\right) - \frac{1}{h_n^{1/2}}\ensuremath{\mathds{E}}\left[K\left(\frac{Y_1 - .}{h_n}\right)\right] \right\|^2 \\
& \leq \frac{\|K\|^2}{n}.
\end{aligned}
\end{equation*}
Therefore, Assumption (ref) (C2) is verified with $\delta_{2n} = n^{-1/2}$. Lastly, putting $[K_h\ast r](y) = h_n^{-1}\int K(h_n^{-1}(y-u))r(u)\ensuremath{\mathrm{d}} u$
\begin{equation*}
\begin{aligned}
\|X_{ni}\|^2 & = \left\|h_n^{-1/2}K\left(\frac{Y_i-.}{h_n}\right) - h_n^{1/2}[K_h\ast r]\right\|^2 \\
& \leq \frac{2}{h_n}\int K^2\left(\frac{Y_i - y}{h_n}\right)\ensuremath{\mathrm{d}} y + 2h_n\|K_h\ast r\|^2 \\
& \leq 4\|K\|^2 \triangleq G^2,
\end{aligned}
\end{equation*}
where the last inequality follows by Young's inequality for convolutions. Then by Theorem (ref) for confidence sets based on the concentration inequality, we obtain
\begin{equation*}
\inf_{P\in\ensuremath{\mathcal{P}}}\mathrm{Pr}\left(\varphi\in C_{n,1-\gamma}\right) \geq 1 - \gamma - O\left(\left(h_n^{s} + \alpha_n^{\beta\wedge 1+1}\right)\sqrt{nh_n}\log n + \frac{\log n}{n^{1/2}}\right).
\end{equation*}
For bootstrap confidence sets we also need to verify Assumptions (ref), (ref), and (ref) (C2). The relevant class of functions is
\begin{equation*}
\ensuremath{\mathcal{G}}_n = \left\{\tilde y \mapsto \frac{1}{h_n^{1/2}}K\left(\frac{y - \tilde y}{h_n}\right):\; y\in[0,1] \right\}.
\end{equation*}
This class admits the envelope $G(y) = h_n^{-1/2}\|K\|_\infty$ and is of VC-type under Assumption (ref) (vi), e.g., see gine2015mathematical, Proposition 3.6.12. Therefore Assumption (ref) is satisfied. For $g_n\in\ensuremath{\mathcal{G}}_n$
\begin{equation*}
\ensuremath{\mathrm{Var}}(g_n) = \ensuremath{\mathds{E}}\left|\frac{1}{h_n^{1/2}}K\left(\frac{Y_1 - y}{h_n}\right)\right|^2 - \left(\ensuremath{\mathds{E}}\left[\frac{1}{h_n^{1/2}}K\left(\frac{Y_1 - y}{h_n}\right)\right]\right)^2
\end{equation*}
so that for sufficiently large $n$
\begin{equation*}
0<\inf_{y\in[0,1]} r(y)\|K\|^2 - h_n\|K\|_1^2\|r\|_\infty^2 \leq \ensuremath{\mathrm{Var}}\left(g_n\right) \leq \|K\|^2\|r\|_\infty.
\end{equation*}
This verifies Assumption (ref) (i). Next, under Assumption (ref) (iv)
\begin{equation*}
\begin{aligned}
\|G\|_{P,4} & = h_n^{-1/2}\|K\|_\infty \\
P|g_n|^3 & = h_n^{-3/2}\ensuremath{\mathds{E}}\left|K\left(h_n^{-1}(Y-y)\right)\right|^3 \\
& \leq h_n^{-1/2}\|K\|^{3}_3\|r\|_\infty \\
P|g_n|^4 & = h_n^{-4/2}\ensuremath{\mathds{E}}\left|K\left(h_n^{-1}(Y-y)\right)\right|^4 \\
& \leq h_n^{-1}\|K\|^{4}_4\|r\|_\infty.
\end{aligned}
\end{equation*}
Therefore, Assumption (ref) (ii) is satisfied with $b_n = O(h_n^{-1/2})$ and some universal constant $\sigma$. Note also that $b_n^4\log^7n/n\to 0$ under Assumption (ref)
Lastly, note that for some universal constant $C$ depending only on $\|r\|_\infty$ and $K$
\begin{equation*}
\begin{aligned}
\ensuremath{\mathds{E}}_P\left\|\hat V_n^\varepsilon - V_n^\varepsilon \right\|_\infty & = \ensuremath{\mathds{E}}\left|\frac{1}{\sqrt{n}}\sum_{i=1}^n\varepsilon_i\right|\ensuremath{\mathds{E}}_P\left\|\frac{1}{nh_n^{1/2}}\sum_{i=1}^nK\left(\frac{Y_i - .}{h_n}\right) - \frac{1}{h_n^{1/2}}\ensuremath{\mathds{E}}\left[K\left(\frac{Y_1 - .}{h_n}\right)\right] \right\|_\infty \\
& \leq C\sqrt{\frac{\log h_n^{-1}}{nh_n}},
\end{aligned}
\end{equation*}
where the last line follows by gine2002rates, Theorem 2.1, under Assumption (ref) (i), (ii), (iv), and (vi). This verifies Assumption (ref) (C2) with $\delta_{2n} = \sqrt{\frac{\log h_n^{-1}}{nh_n}}$. Therefore, by Theorem (ref)
\begin{equation*}
\inf_{P\in\ensuremath{\mathcal{P}}}\Pr(\varphi\in C_{n,1-\gamma}^*) \geq 1 - \gamma - O_p\left(\left(h_n^s + \alpha_n^{\beta\wedge 1+1}\right)\sqrt{nh_n}\log n + \sqrt{\frac{\log h_n^{-1}}{nh_n}}\log n\right) - o_p(1)
\end{equation*}
and the result follows by taking the probability limit.
proof[Proof of Theorem (ref)]
For concentration-based confidence sets we apply Theorem (ref). In the proof of Theorem (ref) we show that $\delta_{1n}=0$. Therefore, by Theorem (ref) with $u_n=h_n^{-1/2}$
\begin{equation*}
\sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\left|C_{n,1-\gamma}\right|_\infty = O\left(\frac{1}{\alpha_n\sqrt{nh_n^q}}\right)
\end{equation*}
Next we verify Assumption (ref). Note that (i) holds, since the class is described by the translation of the kernel function of bounded variation. For (ii) we note that by Hoffman-J\o rgensen's inequality
\begin{equation*}
\begin{aligned}
& \ensuremath{\mathds{E}}_P\left\|\hat X_{ni} - X_{ni}\right\|^2_\infty \\
& \leq \ensuremath{\mathds{E}}_P\left\|\frac{1}{nh_n^{1/2}}\sum_{i=1}^nK\left(\frac{Y_i-y}{h_n}\right) - \frac{1}{h_n^{1/2}}\ensuremath{\mathds{E}}\left[K\left(\frac{Y_i-y}{h_n}\right)\right]\right\|^2_\infty \\
& \leq 432\left(\ensuremath{\mathds{E}}_P\left\|\frac{1}{nh_n^{1/2}}\sum_{i=1}^nK\left(\frac{Y_i-y}{h_n}\right) - \frac{1}{h_n^{1/2}}\ensuremath{\mathds{E}}\left[K\left(\frac{Y_i-y}{h_n}\right)\right]\right\|_\infty + \frac{2}{nh_n^{1/2}}\right)^2.
\end{aligned}
\end{equation*}
Therefore,
\begin{equation*}
\sup_{P\in\ensuremath{\mathcal{P}}}\ensuremath{\mathds{E}}_P\left\|\hat X_{ni} - X_{ni}\right\|^2_\infty = O\left(\frac{\log h_n^{-1}}{n} + h_n^{2s} + \frac{1}{n^2h_n}\right)
\end{equation*}
which tends to zero under Assumption (ref) (vi). Lastly, (iii) follows since $\ensuremath{\mathds{E}}_P\|X_{ni}\|^2_\infty \leq 2\|K\|_\infty^2h_n^{-1}$. Therefore, we obtain the second statement by Theorem (ref).
Technical lemmas
We first characterize the bias of the regularized estimator uniformly over our class of models.
proof[Proof of Proposition (ref)]
By Lemma (ref), $(\alpha_n I + T^*T)$ is an invertible operator between $(C,\|.\|_\infty)$ spaces. Using $f(T^*T)T^*=T^*f(TT^*)$ with $f(x)=(\alpha_n+x)^{-1}$, and factorizing the operator norm $\|T^*g(TT^*)\phi\|_\infty \leq \|T^*\|_{2,\infty}\|g(TT^*)\|\|\phi\|$ with $g(x)=\alpha_n(\alpha_n + x)^{-1}x^\beta$, under Assumption (ref) for any $(\varphi,T)\in\ensuremath{\mathcal{F}}$
\begin{equation}
\begin{aligned}
\left\|(\alpha_n I + T^*T)^{-1}T^*r - \varphi\right\|_\infty & = \left\|\left[(\alpha_n I + T^*T)^{-1}T^*T - I\right]\varphi\right\|_\infty \\
& = \left\|\alpha_n (\alpha_n I + T^*T)^{-1}\varphi\right\|_\infty \\
& = \left\|\alpha_n(\alpha_n I + T^*T)^{-1}(T^*T)^\beta T^*\psi\right\|_\infty \\
& = \left\|T^*\alpha_n(\alpha_n I + TT^*)^{-1}(TT^*)^\beta\psi\right\|_\infty \\
& \leq \|T^*\|_{2,\infty}\left\|\alpha_n(\alpha_n I + TT^*)^{-1}(TT^*)^\beta\right\|\|\psi\| \\
\end{aligned}
\end{equation}
Note that for $b\in(0,1)$, the function $\lambda\mapsto \frac{\lambda^b}{\alpha_n + \lambda}$ is strictly concave on $(0,\infty)$ admitting its maximum at $\lambda = \frac{b}{1-b}\alpha_n$. On the other hand, for $b\in[1,\infty)$, this function is strictly increasing on $[0,\|T\|^2]$, reaching its maximum at the end of this interval. Therefore, by isometry of functional calculus
\begin{equation*}
\left\|\alpha_n(\alpha_n I + TT^*)^{-1}(TT^*)^\beta\right\| = \alpha_n\sup_{\lambda\in[0,\|T\|^2]}\left|\frac{\lambda^\beta}{\alpha_n + \lambda}\right| \leq \tilde R\alpha_n^{\beta\wedge 1}
\end{equation*}
with $\tilde R = \beta^\beta(1-\beta)^{1-\beta}\ensuremath{\mathds{1}}_{0<\beta<1} + C^{2(\beta-1)}\ensuremath{\mathds{1}}_{\beta\geq1}$. Therefore,
\begin{equation*}
\sup_{(\varphi,T)\in\ensuremath{\mathcal{F}}}\left\|(\alpha_n I + T^*T)^{-1}T^*r - \varphi\right\|_\infty \leq R\alpha_n^{\beta\wedge 1}
\end{equation*}
with $R = C^2\tilde R$.
We also need the following inequality known in the theory of numerical ill-posed inverse problems.
lemmaSuppose that $T:L_2[a,b]^p\to L_2[a,b]^q$ is an integral operator with a continuous kernel function. Then for any $\alpha>0$ $(\alpha I + T^*T)$ is invertible as an operator from $\ensuremath{\mathcal{R}}(T^*T)\subset (C,\|.\|_\infty)$ to $(C,\|.\|_\infty)$ space and
\begin{equation*}
\left\|(\alpha I + T^*T)^{-1}\right\|_{\infty} \leq \frac{\|T^*\|_{2,\infty}/2 + \alpha^{1/2}}{\alpha^{3/2}}.
\end{equation*}
proofContinuity of the kernel function ensures that $T^*T$ maps to $C[a,b]^p$. Since for any $\alpha>0$, $\alpha I + T^*T$ is invertible as an operator on the $L_2$ space, see nair2009linear, Lemma 4.1, and $(C,\|.\|_\infty)\subset L_2$, it is also invertible as an operator between $(C,\|.\|_\infty)$.
Note that for any $\varphi\in C[a,b]^p$
\begin{equation*}
\left[(\alpha_nI + T^*T)^{-1}T^*T - I\right]\varphi = -\alpha_n(\alpha_n I + T^*T)^{-1}\varphi.
\end{equation*}
Using this identity
\begin{equation}
\left\|(\alpha_n I + T^*T)^{-1}\right\|_{\infty} = \sup_{\|\varphi\|_\infty=1}\left\|(\alpha_n I + T^*T)^{-1}\varphi\right\|_\infty \leq \frac{\|(\alpha_n I + T^*T)^{-1}T^*T\|_\infty + 1}{\alpha_n}.
\end{equation}
Factoring the norm as $\|T^*\psi\|_\infty\leq \|T^*\|_{2,\infty}\|\psi\|$, and using the isometry of functional calculus
\begin{equation*}
\begin{aligned}
\left\|(\alpha_n I + T^*T)^{-1}T^*T\right\|_\infty & = \left\|T^*(\alpha_n I + TT^*)^{-1}T\right\|_\infty \\
& \leq \|T^*\|_{2,\infty}\sup_{\lambda\in[0,\|T\|^2]}\left|\frac{\lambda^{1/2}}{\alpha_n + \lambda}\right| \\
& = \frac{\|T^*\|_{2,\infty}}{2\alpha_n^{1/2}}.
\end{aligned}
\end{equation*}
Combining this with Eq. (ref) gives the result.
For the next result, put
equation*[equation* omitted — 126 chars of source]
lemmaSuppose that $T$ and $\hat T$ are integral operators in $\ensuremath{\mathcal{L}}_2$ with continuous kernel functions mapping from $L_2[a,b]^p$ to a subset of $L_2[a,b]^q$. Then under Assumption (ref), there exists a constant $C<\infty$ that does not depend on $(\varphi,T)$ such that
\begin{equation*}
\|\xi_n\|_\infty \leq \frac{C}{\alpha_n^{1/2}}\left(\|\hat T^* - T^*\|_{2,\infty} + \|\hat T^* - T^*\|_{2,\infty}^2\right).
\end{equation*}
proofDecompose
\begin{equation*}
\begin{aligned}
\xi_n & = \left[(\alpha_n I + \hat T^*\hat T)^{-1}\hat T^*\hat T - (\alpha_n I + T^*T)^{-1}T^*T\right]\varphi \\
& = \left[\alpha_n(\alpha_n I + T^*T)^{-1} - \alpha_n(\alpha_n I + \hat T^*\hat T)^{-1}\right]\varphi \\
& = (\alpha_nI + \hat T^*\hat T)^{-1}\left[\hat T^*\hat T - T^*T\right]\alpha_n(\alpha_nI + T^*T)^{-1}\varphi \\
& \leq (\alpha_n I + \hat T^*\hat T)^{-1}\hat T^*(\hat T - T)\alpha_n(\alpha_n I + T^*T)^{-1}\varphi \\
& + (\alpha_n I + \hat T^*\hat T)^{-1}(\hat T^* - T^*)\alpha_nT(\alpha_n I + T^*T)^{-1}\varphi \\
& \equiv I_n + II_n.
\end{aligned}
\end{equation*}
Both terms involve the error from the estimation of the operators and regularization bias. For the first term, we factor the operator norm $\|\hat T^*\psi\|_\infty \leq \|\hat T^*\|_{2,\infty}\|\psi\|$
\begin{equation*}
\begin{aligned}
\|I_n\|_\infty & \leq \|\hat T^*\|_{2,\infty}\left\|(\alpha_n I + \hat T\hat T^*)^{-1}\right\|\|\hat T - T\|\left\|\alpha_n(\alpha_n I + T^*T)^{-1}\varphi\right\|
\end{aligned}
\end{equation*}
By isometry of functional calculus
\begin{equation*}
\left\|(\alpha_n I + \hat T^*\hat T)^{-1}\right\| = \sup_{\lambda\in[0,\|\hat T\|^2]}\left|\frac{1}{\alpha_n + \lambda}\right| \leq \frac{1}{\alpha_n},\qquad \mathrm{a.s.}
\end{equation*}
Under Assumption (ref), the last term is
\begin{equation*}
\left\|\alpha_n(\alpha_n I + T^*T)^{-1}(T^*T)^\beta T^*\psi\right\| \leq \alpha_n\sup_{\lambda\in[0,\|T\|^2]}\left|\frac{\lambda^{\beta+1/2}}{\alpha_n + \lambda}\right|C = O\left(\alpha_n^{(\beta+1/2)\wedge 1}\right).
\end{equation*}
Combining these findings with the following bound
\begin{equation*}
\|\hat T - T\|=\|\hat T^* - T^*\|\leq (b-a)^{p/2}\|\hat T^* - T^*\|_{2,\infty},
\end{equation*}
we have for all $(\varphi,T)\in\ensuremath{\mathcal{F}}$
\begin{equation*}
\|I_n\|_\infty = O\left(\frac{\alpha_n^{\beta\wedge 1/2}}{\alpha_n^{1/2}}\|\hat T^*-T^*\|_{2,\infty}\|\hat T^*\|_{2,\infty}\right).
\end{equation*}
For the second term, we factor the operator norm in the following way:
\begin{equation*}
\|II_n\|_\infty \leq \left\|(\alpha_n I + \hat T^*\hat T)^{-1}\right\|_\infty\left\|\hat T^* - T^*\right\|_{2,\infty}\left\|\alpha_nT(\alpha_n I + T^*T)^{-1}\varphi\right\|
\end{equation*}
By Lemma (ref)
\begin{equation*}
\left\|(\alpha_n I + \hat T^*\hat T)^{-1}\right\|_\infty \leq \frac{\|\hat T^*\|_{2,\infty}/2 + \alpha_n^{1/2}}{\alpha_n^{3/2}}
\end{equation*}
Under Assumption (ref),
\begin{equation*}
\left\|\alpha_nT(\alpha_n I + T^*T)^{-1}(T^*T)^\beta T^*\psi\right\| \leq\alpha_n\sup_{\lambda\in[0,\|T\|^2]}\left|\frac{\lambda^{\beta+1}}{\alpha_n + \lambda}\right| C = O(\alpha_n).
\end{equation*}
Combining all above findings, we have uniformly in $(\varphi, T)$
\begin{equation*}
\|II_n\|_\infty = O\left(\|\hat T^* - T^*\|_{2,\infty}\left(\frac{\|\hat T^*\|_{2,\infty}}{\alpha_n^{1/2}} + 1\right)\right),
\end{equation*}
and the conclusion follows from collecting all estimates and by triangle inequality.
For the following lemma denote
equation*[equation* omitted — 413 chars of source]
lemmaFor any $T,\hat T\in\ensuremath{\mathcal{L}}_2$ we have
\begin{equation*}
\begin{aligned}
\left\|\hat\nu_{n} - \nu_{n}\right\|_\infty & \leq \left(\frac{1}{\alpha_n} + \frac{1}{2\alpha_n^{3/2}}\|T^*\|_{2,\infty} \right)\|\hat T^* - T^*\|_{2,\infty}\left\|\frac{u_n}{n}\sum_{i=1}^nX_{ni}\right\|, \\
\left\|\hat\nu_{n}^\eta - \nu_{n}^\eta\right\|_\infty & \leq \left(\frac{1}{\alpha_n} + \frac{1}{2\alpha_n^{3/2}}\|T^*\|_{2,\infty} \right)\|\hat T^* - T^*\|_{2,\infty}\left\|\frac{u_n}{n}\sum_{i=1}^n\eta_i X_{ni}\right\| \\
& \qquad + \frac{\|\hat T^*\|_{2,\infty}}{\alpha_n}\left\|\frac{u_n}{n}\sum_{i=1}^n\eta_i(\hat X_{ni} - X_{ni})\right\|. \\
\end{aligned}
\end{equation*}
proofDecompose
\begin{equation*}
\begin{aligned}
& \hat\nu_n - \nu_n = \\
& = \left[\hat T^*(\alpha_n I + \hat T\hat T^*)^{-1} - T^*(\alpha_n I + TT^*)^{-1}\right]\frac{u_n}{n}\sum_{i=1}^nX_{ni} \\
& = \left[(\hat T^* - T^*)(\alpha_n I + \hat T\hat T^*)^{-1} - T^*(\alpha_n I + \hat T\hat T^*)^{-1}(\hat T\hat T^* - TT^*)(\alpha_n I + TT^*)^{-1}\right]\frac{u_n}{n}\sum_{i=1}^nX_{ni} \\
& = \left[(\hat T^* - T^*)(\alpha_n I + \hat T\hat T^*)^{-1} - T^*(\alpha_n I + \hat T\hat T^*)^{-1}\hat T(\hat T^* - T^*)(\alpha_n I + TT^*)^{-1}\right]\frac{u_n}{n}\sum_{i=1}^nX_{ni} \\
& \quad - T^*(\alpha_n I + \hat T\hat T^*)^{-1}(\hat T - T)T^*(\alpha_n I + TT^*)^{-1}\frac{u_n}{n}\sum_{i=1}^nX_{ni}.
\end{aligned}
\end{equation*}
Note that
\begin{equation*}
\|\hat T - T\| = \|\hat T^* - T^*\| \leq \|\hat T^* - T^*\|_{2,\infty}
\end{equation*}
and
\begin{equation*}
\|(\alpha_n I + \hat T\hat T^*)^{-1}\hat T\|\leq \frac{1}{2\alpha_n^{1/2}},\qquad \|T^*(\alpha_n I + TT^*)^{-1}\| \leq \frac{1}{2\alpha_n^{1/2}}.
\end{equation*}
Then
\begin{equation*}
\begin{aligned}
\|\hat\nu_n - \nu_n\|_\infty & \leq \left(\frac{1}{\alpha_n} + \frac{1}{2\alpha_n^{3/2}}\|T^*\|_{2,\infty} \right)\|\hat T^* - T^*\|_{2,\infty}\left\|\frac{u_n}{n}\sum_{i=1}^nX_{ni}\right\|. \\
\end{aligned}
\end{equation*}
For the second statement note that
\begin{equation*}
\begin{aligned}
\hat\nu_n^\eta - \nu_n^\eta & = \left[\hat T^*(\alpha_n I + \hat T\hat T^*)^{-1} - T^*(\alpha_n I + TT^*)^{-1}\right]\frac{u_n}{n}\sum_{i=1}^n\eta_iX_{ni} \\
& \qquad + (\alpha_nI + \hat T^*\hat T)^{-1}\hat T^*\frac{u_n}{n}\sum_{i=1}^n\eta_i(\hat X_{ni} - X_{ni}) \\
& \triangleq I_n + II_n.
\end{aligned}
\end{equation*}
Note that the bound on $I_n$ follows trivially from above computations if we replace $X_{ni}$ by $\eta_iX_{ni}$, while
\begin{equation*}
\|II_n\|_\infty \leq \frac{\|\hat T^*\|_{2,\infty}}{\alpha_n}\left\|\frac{u_n}{n}\sum_{i=1}^n\eta_i(\hat X_{ni} - X_{ni})\right\|.
\end{equation*}
lemmaSuppose that $T^*\in\ensuremath{\mathcal{L}}_{2,\infty}$ is an integral operator with continuous kernel function $k$, then
\begin{equation*}
\|T^*\|_{2,\infty} = \sup_{z\in[a,b]^p}\left(\int_{[a,b]^q}|k(z,w)|^2\ensuremath{\mathrm{d}} w\right)^{1/2}\equiv \|k\|_{2,\infty},
\end{equation*}
where $\|k\|_{2,\infty}$ is a mixed norm on the iterated space $L_\infty[a,b]^p(L_2[a,b]^q)$.
proofBy Cauchy-Schwartz inequality, we have $\|T^*\|_{2,\infty}\leq \|k\|_{2,\infty}$. On the other side, continuity of $k$ implies that $\exists z_0\in[a,b]^p$ such that $\|k\|_{2,\infty} = \left(\int_{[a,b]^q}|k(z_0,w)|^2\ensuremath{\mathrm{d}} w\right)^{1/2}$. Take $\psi(w) = \frac{k(z_0,w)}{\|k(z_0,.)\|}$. Then $\|\psi\|=1$, and
\begin{equation*}
\|T^*\|_{2,\infty} \geq \|T^*\psi\|_{\infty} \geq |(T^*\psi)(z_0)| = \|k(z_0,.)\| = \|k\|_{2,\infty}.
\end{equation*}