EconBase
← Back to paper

Nonlinear Binscatter Methods

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

214,314 characters · 37 sections · 45 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Nonlinear Binscatter Methods Supplemental Appendix

abstractThis supplement collects all technical proofs for more general theoretical results than those reported in the main paper. Several of our new theoretical results for nonlinear partitioning-based series estimation may be of independent interest. More details on methodological aspects of nonlinear binscatter are also provided. Companion general-purpose software and replication files are available at \url{https://nppackages.github.io/binsreg/}.

\thispagestyle{empty}

\pagestyle{plain}

\doublespacing \setcounter{page}{1}

{\onehalfspacing \setcounter{tocdepth}{2} \addtocontents{toc}{\thispagestyle{plain}} }

Introduction

This supplemental appendix is a comprehensive collection of all our new theoretical results for nonlinear binscatter estimators with semi-linear covariate-adjustment and random partitioning. Many of our results contribute to the broader literature on nonparametric estimation and inference, particularly when using series estimators, and are thus of independent interest outside of binned scatter plots. To help place our results in the literature, we include a remark labelled “Improvements over literature” at the end of each technical subsection that discusses in detail the technical improvements presented in that subsection and gives related references.

Here we give a brief summary of this appendix, including pointers to some of the major new results. We proceed as follows. The next subsection lists notation used throughout; further notation is defined throughout Section (ref) and at the outset of Section (ref). Then, Section (ref) describes the setup for nonlinear binscatter methods, including the statistical model, parameters of interest, and assumptions, as well as the (random) partitioning and estimation details. Specifically, Assumption (ref) imposes some basic conditions on the data generating process. Assumption (ref) imposes some technical conditions that characterizes and restricts the statistical model of interest. The loss function specified there is general enough to cover many practically important examples such as mean regression, quantile regression, logit/probit estimation, and Huber regression. Assumption (ref) imposes some mild high-level conditions on the estimation procedure. Assumption (ref) summarizes the key conditions on the partitioning scheme used in our theory. We allow for a large class of random partitions. Importantly, the “convergence” of the random partition (Assumption (ref)(ii)) is not necessary for most of our main theoretical results, thereby allowing for flexible data-driven partitioning methods, including certain recursive adaptive partitioning methods: see Devroye-etal2013_book, zhang2010recursive, and references therein.

Section (ref) presents some preliminary technical lemmas for analyzing nonlinear binscatter (and thus also partitioning-based estimators more broadly). New results include precise non-asymptotic concentration results related to Gram matrices (Lemma (ref)), asymptotic variances (Lemmas (ref) and (ref)), approximation errors (Lemma (ref)), and uniform convergence (Lemma (ref)). Sharp control of these objects is a crucial ingredient for obtaining results under weak conditions, as below.

Section (ref) presents a tight uniform (in $x$) Bahadur representation for nonlinear binscatter (Theorem (ref)). This is our first main result. We allow for random partitions and much weaker rate restrictions (on the tuning parameter $J$) than previously imposed in the literature, in addition to additional controls. The data-dependent partitioning means our series estimator uses random basis functions, and this is entirely new. In terms of tuning parameter rate restrictions, previous results required $J^4/n\to 0$ (up to $\log(n)$ terms) or something stricter, while our restriction is that $J^{\frac{2\nu}{\nu-1}}/n\to 0$ (up to $\log(n)$ terms), with $\nu>2$ denoting the number of finite moments of the “score”, and thus may be substantially weaker. Note that our class of models is often broader than prior work also. Importantly, our results can now allow for piecewise constant binscatters, i.e., with degree $p=0$, which is excluded by prior results in the literature (i.e., for previous technical results there was no sequence $J \to\infty$ such that the bias and variance are simultaneously controlled). In addition, employing our novel uniform Bahadur representation, we can establish the uniform convergence rates of nonlinear binscatter (Corollary (ref)) and variance estimators (Theorem (ref)) under similarly weak restrictions.

Section (ref) studies the pointwise distributional approximation for nonlinear binscatter estimators. These results are omitted from the main paper to save space, but are standard properties of interest in the nonparametrics literature and thus are included for completeness. The main result is Theorem (ref), which establishes pointwise asymptotic Normality for our point estimators, again allowing for random (and possibly “non-convergent”) partitions, and under mild rate restrictions similar to those for the (uniform) Bahadur representation.

Section (ref) presents a new Nagar-type approximate IMSE expansion for nonlinear binscatter estimators with semi-linear covariate-adjustment and random partitions (Theorem (ref)), which has no antecedent in the literature. Our results can be used to design data-driven procedures for selecting IMSE-optimal choices of tuning parameters for nonlinear binscatter. Again, these results are novel in their breadth, the weakness of the assumptions, and the conditions on the partitioning. Here we do require an extra assumption on the partitioning in order to characterize the leading terms in the expansion: intuitively, the random partitioning must “settle” to a population partition so that the leading constants of the expansion can be expressed. For example, sample quantiles converge to population quantiles, so this assumption is satisfied.

Uniform inference is dealt with in the next several sections of this appendix. First, Section (ref) establishes a uniform (in $x$) distributional approximation for nonlinear binscatter estimators. The two main results, which are combined into one in the main text (Theorem 2), are the (conditional) strong approximation in Theorem (ref) and the feasible implementation thereof in Theorem (ref). Again, We allow for a large class of random partitions, a broad class of (possibly) nonlinear and nonsmooth models, and additional controls. Here the partitions do not need to be “convergent” in any sense. These results are obtained under weak assumptions, including in particular mild rate restrictions, as in the case of the uniform Bahadur representation, all of which improves on the literature in various directions as explained in the text below. Finally, Theorem (ref) shows a distributional approximation for the suprema of the $t$-statistic processes in the case of the convergent partition (as in the previous paragraph).

Sections (ref)--(ref) employ the strong approximation results to study uniform inference for various parameters of interest in the specific context of nonlinear binscatter. These results rely on, and inherit the novelty of, Theorems (ref) and (ref). New results include valid uniform confidence bands (Theorem (ref)), consistent hypothesis tests about parametric specification (Theorem (ref)) and tests for shape restrictions (Theorem (ref)). All these results explicitly account for the possibly random partitioning scheme and semi-linear covariate-adjustment with random evaluation points.

Section (ref) discusses implementation details for nonlinear binscatter, including standard error computation, feasible data-driven number of bins selector, and choices of polynomial orders given a fixed number of bins. For a more explicit treatment of the package binsreg per se, see Cattaneo-Crump-Farrell-Feng_2024_Stata and \url{https://nppackages.github.io/binsreg/}.

Finally, Section (ref) contains the proofs for all the technical results in Section (ref).

Notation

See vandevarrt-Wellner_1996_book, Bhatia_2013_book, Gine-Nickl_2016_book, and references therein, for background definitions.

Matrices and Norms. For (column) vectors, $\|\cdot\|$ denotes the Euclidean norm, $\|\cdot\|_1$ denotes the $L_1$ norm, $\|\cdot\|_\infty$ denotes the sup-norm, and $\|\cdot\|_0$ denotes the number of nonzeros. For matrices, $\|\cdot\|$ is the operator matrix norm induced by the $L_2$ norm, and $\|\cdot\|_\infty$ is the matrix norm induced by the supremum norm, i.e., the maximum absolute row sum of a matrix. For a square matrix $\mathbf{A}$, $\lambda_{\max}(\mathbf{A})$ and $\lambda_{\min}(\mathbf{A})$ are the maximum and minimum eigenvalues of $\mathbf{A}$, respectively. $[\mathbf{A}]_{ij}$ denotes the $(i,j)$th entry of a generic matrix $\mathbf{A}$. We will use $\mathcal{S}^{L}$ to denote the unit circle in $\mathbb{R}^L$, i.e., $\|\mathbf{a}\|=1$ for any $\mathbf{a}\in \mathcal{S}^L$. For a real-valued function $g(\cdot)$ defined on a measure space $\mathcal{Z}$, let $\|g\|_{\mathbb{Q}, 2}:=(\int_\mathcal{Z} |g|^2d\mathbb{Q})^{1/2}$ be its $L_2$-norm with respect to the measure $\mathbb{Q}$. In addition, let $\|g\|_\infty=\sup_{z\in\mathcal{\mathcal{Z}}}|g(z)|$ be $L_\infty$-norm of $g(\cdot)$, and if $g$ is a univariate function, let $g^{(v)}(z)=d^vg(z)/dz^v$ be the $v$th derivative for $v\geq 0$.

Asymptotics. For sequences of numbers or random variables, we use $l_n\lesssim m_n$ to denote that $\limsup_n|l_n/m_n|$ is finite, $l_n\lesssim_\mathbb{P} m_n$ or $l_n=O_\mathbb{P}(m_n)$ to denote $\limsup_{\varepsilon\to\infty}\limsup_n\mathbb{P}[|l_n/m_n|\geq\varepsilon]=0$, $l_n=o(m_n)$ implies $l_n/m_n\to 0$, and $l_n=o_\mathbb{P}(m_n)$ implies that $l_n/m_n\to_\mathbb{P} 0$, where $\to_\mathbb{P}$ denotes convergence in probability. Accordingly, we write $l_n\gtrsim m_n$ if $m_n\lesssim l_n$, and $l_n\gtrsim_\mathbb{P} m_n$ if $m_n\lesssim_\mathbb{P} l_n$. $l_n\asymp m_n$ implies that $l_n\lesssim m_n$ and $m_n\lesssim l_n$.

Empirical Process. We employ standard empirical process notation: $\mathbb{E}_n[g(\mathbf{v}_i)]=\frac{1}{n}\sum_{i=1}^ng(\mathbf{v}_i)$, and $\mathbb{G}_n[g(\mathbf{v}_i)]=\frac{1}{\sqrt{n}}\sum_{i=1}^n(g(\mathbf{v}_i)-\mathbb{E}[g(\mathbf{v}_i)])$ for a sequence of random variables $\{\mathbf{v}_i\}_{i=1}^n$. In addition, we employ the notion of covering number extensively in the proofs. Specifically, given a measurable space $(A, \mathcal{A})$ and a suitably measurable class of functions $\mathcal{G}$ mapping $A$ to $\mathbb{R}$ equipped with a measurable envelop function $\bar{G}(z)\geq \sup_{g\in\mathcal{G}}|g(z)|$, the covering number of $N(\mathcal{G}, L_2(\mathbb{Q}), \varepsilon)$ is the minimal number of $L_2(\mathbb{Q})$-balls of radius $\varepsilon$ needed to cover $\mathcal{G}$ for a measure $\mathbb{Q}$. The covering number of $\mathcal{G}$ relative to the envelope is denoted as $N(\mathcal{G}, L_2(\mathbb{Q}), \varepsilon\|\bar{G}\|_{\mathbb{Q},2})$.

Other. $\lceil z \rceil$ outputs the smallest integer no less than $z$ and $a\wedge b=\min\{a,b\}$. “w.p.a. $1$” means “with probability approaching one”.

Setup

Suppose that $(y_i,x_i, \mathbf{w}_i')$, $1\leq i\leq n$, is a random sample where $y_i\in\mathcal{Y}$ is a scalar response variable, $x_i\in\mathcal{X}$ is a scalar covariate, and $\mathbf{w}_i\in\mathcal{W}$ is a vector of additional control variables of dimension $d$. Let $\mathbf{D}= [(y_i, x_i, \mathbf{w}_i')' : i=1,2,\dots,n]$.

For a loss function $\rho(\cdot; \cdot)$ and a strictly monotonic transformation function $\eta(\cdot)$, define

equation[equation omitted — 248 chars of source]

where $\mathcal{M}$ is a space of functions satisfying certain smoothness conditions to be specified later.

This setup is general. For example, consider $\boldsymbol{\gamma}_0=\mathbf{0}$. If $\rho(\cdot;\cdot)$ is a squared loss and $\eta(\cdot)$ is the identity function, $\mu_0(x)$ is the conditional expectation of $y_i$ given $x_i=x$. Let $\mathds{1}(\cdot)$ denote the indicator function. If $\rho(y;\eta)=(q-\mathds{1}(y<\eta))(y-\eta)$ for some $0<q<1$ and $\eta(\cdot)$ is an identity function, then $\mu_0(x)$ is the $q$th conditional quantile of $y_i$ given $x_i=x$. Introducing a transformation function $\eta(\cdot)$ is useful. For instance, it may accommodate logistic regression for binary responses. When $\boldsymbol{\gamma}_0\neq \mathbf{0}$, the parametric and the nonparametric components are additively separable, and thus (ref) becomes a generalized partially linear model.

Binscatter estimators are typically constructed based on a (possibly random) partition of the support of the covariate $x_i$. Specifically, the relevant support of $x_i$ is partitioned into $J$ disjoint intervals, leading to the partitioning scheme $\widehat{\Delta} = \{\widehat{\mathcal{B}}_1, \widehat{\mathcal{B}}_2, \dots, \widehat{\mathcal{B}}_J\}$, where \[ \widehat{\mathcal{B}}_j =

cases[\widehat{\tau}_{j-1}, \widehat{\tau}_j)& \qquad if j=1, \cdots, J-1\\ [\widehat{\tau}_{J-1}, \widehat{\tau}_J]& \qquad if j=J

, \] One popular choice in binscatter applications is the quantile-based partition: $\widehat\tau_j=\widehat{F}_X^{-1}(j/J)$ with $\widehat{F}_X(u)=n^{-1}\sum_{i=1}^n\mathds{1}(x_i\leq u)$ the empirical cumulative distribution function and $\widehat{F}_X^{-1}$ its generalized inverse. Our theory is general enough to cover other partitioning schemes satisfying certain regularity conditions specified below. An innovation herein is accounting for the additional randomness from the partition $\widehat\Delta$. The number of bins $J$ plays the role of the tuning parameter for the binscatter method, and is assumed to diverge: $J\to\infty$ as $n\to\infty$ throughout the supplement, unless explicitly stated otherwise.

The piecewise polynomial basis of degree $p$, for some choice of $p=0,1,2,\dots$, is defined as \[\Big[

array[array omitted — 148 chars of source]

\Big]' \otimes \Big[

array[array omitted — 36 chars of source]

\Big]', \] where $\mathds{1}_\mathcal{A}(x)=\mathds{1}(x\in \mathcal{A})$ and $\otimes$ is the Kronecker product operator. For convenience of later analysis, we use $\widehat{\mathbf{b}}_{p,0}(x)$ to denote a standardized rotated basis, the $j$th element of which is given by \[ \sqrt{J}\times\mathds{1}_{\widehat{\mathcal{B}}_{\bar{j}}}(x)\times\Big(\frac{x-\widehat{\tau}_{\bar{j}-1}}{\hat{h}_{\bar{j}}} \Big)^{j-1-(\bar{j}-1)(p+1)}, \quad j=1, \cdots, (p+1)J, \] where $\bar{j}=\lceil j/(p+1)\rceil$, $\lceil\cdot\rceil$ is the ceiling operator, and $\hat{h}_{\bar{j}}=\widehat\tau_{\bar{j}}-\widehat\tau_{\bar{j}-1}$. Thus, each local polynomial is centered at the start of each bin and scaled by the length of the bin. $\sqrt{J}$ is an additional scaling factor which helps simplify some expressions of our results. The standardized rotated basis $\widehat{\mathbf{b}}_{p,0}(x)$ is equivalent to the original piecewise polynomial basis in the sense that they represent the same (linear) function space.

To impose the restriction that the estimated function is $(s-1)$-times continuously differentiable for $1\leq s \leq p$, we introduce the following basis \[ \widehat{\mathbf{b}}_{p,s}(x)=\Big(\widehat{b}_{p,s,1}(x),\ldots, \widehat{b}_{p,s,K_{p,s}}(x)\Big)' =\widehat{\mathbf{T}}_s\widehat{\mathbf{b}}_{p,0}(x), \qquad K_{p,s}=(p+1)J-s(J-1), \] where $\widehat{\mathbf{T}}_s:=\widehat{\mathbf{T}}_s(\widehat{\Delta})$ is a $K_{p,s}\times(p+1)J$ matrix depending on $\widehat{\Delta}$, which transforms a piecewise polynomial basis into a smoothed binscatter basis. Some useful properties of $\widehat\mathbf{T}_s$ are given in Lemma (ref) in Section (ref), and the explicit representation of $\widehat\mathbf{T}_s$ is available in the proof of Lemma SA-3.2 in \citet*{Cattaneo-Crump-Farrell-Feng_2024_AER}. When $s=0$, we let $\widehat{\mathbf{T}}_0=\mathbf{I}_{(p+1)J}$, the identity matrix of dimension $(p+1)J$. When $s=p$, $\widehat{\mathbf{b}}_{p,s}(x)$ is the well-known $B$-spline basis of order $p+1$ with simple knots, which is $(p-1)$-times continuously differentiable. When $0<s<p$, they can be defined similarly as $B$-splines with knots of certain multiplicities. See Definition 4.1 in Section 4 of Schumaker_2007_book for more details about spline functions. We require $s\leq p$, since if $s=p+1$, $\widehat{\mathbf{b}}_{p,s}(x)$ reduces to a global polynomial basis of degree $p$.

Given a choice of basis, we consider the following generalized binscatter estimator:

equation[equation omitted — 459 chars of source]

where $\widehat{\mathbf{b}}_{p,s}^{(v)}(x)=\frac{d^v}{dx^v}\widehat{\mathbf{b}}_{p,s}(x)$ for some $v\in\mathbb{Z}_+$ such that $v\leq p$. This estimator can be written as:

equation[equation omitted — 462 chars of source]

The representation (ref) allows us to be more general and agnostic about the estimation of $\boldsymbol{\gamma}_0$, and also simplifies some of the proofs. More specifically, our theory requires only a sufficiently fast convergence rate of $\widehat{\boldsymbol{\gamma}}$ (see Assumption (ref) below), which in nonlinear estimation models cases can be justified in different ways, e.g., joint estimation, backfitting, profiling, and split-sampling, among other possibilities. Our software implementation \citep*{Cattaneo-Crump-Farrell-Feng_2024_Stata} relies on joint estimation, as done by the default base estimation packages in Python, R, and Stata.

In this supplement, we focus on estimation and inference of the following three parameters:

enumerate[label=(\roman*)] • the nonparametric component $\mu_0^{(v)}(x)$ for any $v\geq 0$, • the level function $\vartheta_0(x,\mathsf{w})=\eta(\mu_0(x)+\mathsf{w}'\boldsymbol{\gamma}_0)$, and • the marginal effect $\zeta_0(x,\mathsf{w})=\frac{\partial}{\partial x}\eta(\mu_0(x)+\mathsf{w}'\boldsymbol{\gamma}_0)$,

where $\mathsf{w}$ is a user-chosen evaluation point of the control variables, and thus these parameters are viewed as functions of $x$ only in our theory. Nevertheless, all our results are readily applied to other linear or nonlinear transformations of $\mu_0(x)$, such as the higher-order derivatives $\frac{\partial^v}{\partial x^v}\eta(\mu_0(x)+\mathsf{w}'\boldsymbol{\gamma}_0)$. Given the binscatter estimates $\widehat{\mu}_{p,s}(x)$ and $\widehat\boldsymbol{\gamma}$ in (ref), the estimators of the three parameters defined above are given by

equation*[equation* omitted — 382 chars of source]

respectively, for some consistent estimate $\widehat\mathsf{w}$ (non-random or generated based on $\{\mathbf{w}_i\}_{i=1}^n$) of the evaluation point $\mathsf{w}$. As a reminder, we need to require $p\geq v$ to get $\widehat\mu_{p,s}^{(v)}(x)$, $p\geq 0$ to get $\widehat\vartheta_{p,s}(x,\widehat\mathsf{w})$, and $p\geq 1$ to get $\widehat\zeta_{p,s}(x,\widehat\mathsf{w})$.

Recall that in the main text we always set $s=p$ and omit the dependence of estimators on $s$. Thus, $\widehat{\mu}_p^{(v)}(x)=\widehat{\mu}_{p,p}^{(v)}(x)$, $\widehat{\vartheta}_p(x,\widehat\mathsf{w})= \widehat{\vartheta}_{p,p}(x,\widehat\mathsf{w})$, and $\widehat{\zeta}_p(x,\widehat\mathsf{w})= \widehat{\zeta}_{p,p}(x,\widehat\mathsf{w})$. In this supplement, however, all our results hold for a general choice of the degree and the smoothness of the basis. For ease of notation, the subscripts $p$ and $s$ of the above estimators are dropped hereafter: \[ \widehat{\mu}^{(v)}(x):=\widehat{\mu}_{p,s}^{(v)}(x), \quad \widehat{\vartheta}(x,\widehat\mathsf{w}):=\widehat{\vartheta}_{p,s}(x,\widehat\mathsf{w}),\quad \text{and}\quad \widehat{\zeta}(x,\widehat\mathsf{w}):=\widehat{\zeta}_{p,s}(x,\widehat\mathsf{w}). \]

remark[Smoothness and Bias Correction] This supplemental appendix presents all results under general choices of the number of bins $J$, the degree of the basis $p$, and the smoothness of the basis $s$. By contrast, for simplicity, the main paper employs the basis with the maximum smoothness, i.e. choosing $s=p$, and considers the special case in which $J$ is taken to be the IMSE-optimal choice for a fixed $p$ (see Theorem (ref)), and inference is conducted based on the binscatter basis of degree $(p+1)$. Such a choice of $J$ guarantees that the smoothing bias of the binscatter estimator is negligible in inference under mild conditions and thus can be viewed as a bias correction strategy.

We first assume the following basic conditions on the data generating process.

assumDGP[Data Generating Process] \ \begin{enumerate}[label=(\roman*)] • $\{(y_{i}, x_{i}, \mathbf{w}_i'): 1\leq i\leq n\}$ are i.i.d. random vectors satisfying (ref) and supported on $\mathcal{Y} \times \mathcal{X} \times \mathcal{W}$, where $\mathcal{X}$ is a compact interval and $\mathcal{W}$ is a compact set. • $F_X(x):=\mathbb{P}[x_i\leq x]$ has a Lipschitz continuous (Lebesgue) density $f_X(x)$ bounded away from zero on $\mathcal{X}$. • $F_{Y|XW}(y|x_i,\mathbf{w}_i):=\mathbb{P}[y_i\leq y|x_i,\mathbf{w}_i]$ has a (conditional) density $f_{Y|XW}(y|x_i,\mathbf{w}_i)$ supported on $\mathcal{Y}_{x\mathbf{w}}$ with respect to some sigma-finite measure, and $\sup_{x\in\mathcal{X},\mathbf{w}\in\mathcal{W}}\sup_{y\in\mathcal{Y}_{x\mathbf{w}}} f_{Y|XW}(y|x,\mathbf{w}) \lesssim 1$. \end{enumerate}

Next, we impose several technical conditions related to the statistical model of interest.

assumGL[Statistical Model] \ \begin{enumerate}[label=(\roman*)] • $\rho(y;\eta)$ is absolutely continuous with respect to $\eta\in\mathbb{R}$ and admits a derivative $\psi(y,\eta) := \psi^\dagger(y-\eta)\psi^\ddagger(\eta)$ almost everywhere. $\psi^\ddagger(\cdot)$ is continuously differentiable and strictly positive or negative. $\psi^\dagger(\cdot)$ is Lipschitz continuous if $F_{Y|XW}(y|x_i,\mathbf{w}_i)$ does not have a Lebesgue density, or piecewise Lipschitz with finitely many discontinuity points otherwise. • $\rho(y;\eta(\theta))$ is convex with respect to $\theta$. $\eta(\cdot)$ is strictly monotonic and three-times continuously differentiable. • $\mathbb{E}[\psi(y_i,\eta(\mu_0(x_i)+\mathbf{w}_i'\boldsymbol{\gamma}_0))|x_i,\mathbf{w}_i]=0$. For $\sigma^2(x,\mathbf{w}):=\mathbb{E}[\psi(y_i,\eta(\mu_0(x_i)+\mathbf{w}_i'\boldsymbol{\gamma}_0))^2|x_i=x,\mathbf{w}_i=\mathbf{w}]$, $\inf_{x\in\mathcal{X},\mathbf{w}\in\mathcal{W}}\sigma^2(x,\mathbf{w})\gtrsim 1$. $\mathbb{E}[\eta^{(1)}(\mu_0(x_i)+\mathbf{w}_i'\boldsymbol{\gamma}_0)^2\sigma^2(x_i,\mathbf{w}_i)|x_i=x]$ is Lipschitz continuous on $\mathcal{X}$, and $\sup_{x\in\mathcal{X},\mathbf{w}\in\mathcal{W}}\mathbb{E}[|\psi(y_i,\eta(\mu_0(x_i)+\mathbf{w}_i'\boldsymbol{\gamma}_0))|^\nu|x_i=x, \mathbf{w}_i=\mathbf{w}] \lesssim 1$ for some $\nu>2$. $\mathbb{E}[\psi(y_i,\eta)|x_i=x, \mathbf{w}_i=\mathbf{w}]$ is twice continuously differentiable with respect to $\eta$. • $\inf_{x\in\mathcal{X},\mathbf{w}\in\mathcal{W}}\varkappa(x,\mathbf{w})\gtrsim 1$ and $\mathbb{E}[\varkappa(x_i,\mathbf{w}_i)|x_i=x]$ is Lipschitz continuous on $\mathcal{X}$ where $\varkappa(x,\mathbf{w}):=\Psi_1(x,\mathbf{w};\eta(\mu_0(x)+\mathbf{w}'\boldsymbol{\gamma}_0))(\eta^{(1)}(\mu_0(x)+\mathbf{w}'\boldsymbol{\gamma}_0))^2$, $\Psi_1(x,\mathbf{w};\eta):=\frac{\partial}{\partial \eta}\Psi(x,\mathbf{w};\eta)$, and $\Psi(x,\mathbf{w};\eta):=\mathbb{E}[\psi(y_i,\eta)|x_i=x, \mathbf{w}_i=\mathbf{w}]$. • $\mu_0(\cdot)$ is $\varsigma$-times continuously differentiable for some $\varsigma\geq p+1$. \end{enumerate}

Our next assumption imposes mild high-level conditions on the estimator $\widehat{\boldsymbol{\gamma}}$ of the coefficient vector $\boldsymbol{\gamma}_0$, the estimator $\widehat\mathsf{w}$ of the evaluation point $\mathsf{w}$ for control variables, and the estimator of the function $\Psi_1$ defined previously in Assumption (ref)(iv).

assumHLE[High-Level Estimation Conditions] \ \begin{enumerate}[label=(\roman*)] • $\|\widehat{\boldsymbol{\gamma}}-\boldsymbol{\gamma}_0\|\lesssim_\mathbb{P} \mathfrak{r}_\gamma$ for $\mathfrak{r}_\gamma=o(\sqrt{J/n}+J^{-p-1})$, and $\|\widehat\mathsf{w}-\mathsf{w}\|=o_\mathbb{P}(1)$. • For some estimator $\widehat{\Psi}_1$ of $\Psi_1$, $\|\mathbb{E}_n[\widehat{\mathbf{b}}_{p,s}(x_i)\widehat{\mathbf{b}}_{p,s}(x_i)' (\widehat{\varkappa}(x_i, \mathbf{w}_i)-\varkappa(x_i,\mathbf{w}_i))\|\lesssim_\mathbb{P} J^{-p-1}+\big(\frac{J\log n}{n^{1-2/\nu}}\big)^{1/2}$ where $\widehat{\varkappa}(x_i,\mathbf{w}_i)=\widehat{\Psi}_1(x_i, \mathbf{w}_i;\eta(\widehat{\mu}(x_i)+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}}))(\eta^{(1)}(\widehat{\mu}(x_i)+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}}))^2$. \end{enumerate}

Note that $\Upsilon(x,\mathbf{w})=\Psi_1(x, \mathbf{w};\eta(\mu_0(x)+\mathbf{w}'\boldsymbol{\gamma}_0))$ in the main paper to streamline the presentation. Part (i) is a mild condition on the convergence of $\widehat{\boldsymbol{\gamma}}$ and $\widehat\mathsf{w}$. Part (ii) is a high-level condition that ensures we have a valid feasible estimator of the Gram matrix ($\bar\mathbf{Q}$ or $\mathbf{Q}_0$ defined at the outset of Section (ref) below). Note that the convergence rate of $\eta^{(1)}(\widehat{\mu}(x_i)+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}})$ can be deduced from Corollary (ref) below. Thus, part (ii) can be largely viewed as a restriction on $\widehat{\Psi}_1$ only. Note that $\widehat{\Psi}_1$ does not have to be consistent for $\Psi_1$ in any sense; it suffices that the estimator $\mathbb{E}_n[\widehat{\mathbf{b}}_{p,s}(x_i)\widehat{\mathbf{b}}_{p,s}(x_i)' \widehat{\varkappa}(x_i, \mathbf{w}_i)]$ based on $\widehat{\Psi}_1$ as a whole is consistent. See Section (ref) for several examples of the estimator $\widehat{\Psi}_1$.

Partitions

We need some regularity conditions on the partitioning scheme, which can be verified in a case-by-case basis. We first define a family of “quasi-uniform” partitions for some absolute constant $C>0$:

equation[equation omitted — 161 chars of source]

where $h_j(\Delta)$ denotes the length of the $j$th bin in the partition $\Delta$. Roughly speaking, (ref) says that the bins in any $\Delta\in\Pi_C$ do not differ too much in length. Also, let $\mathbf{X} = [x_1, \ldots, x_n]'$, $\mathbf{W}=[\mathbf{w}_1,\cdots, \mathbf{w}_n]'$ and $\mathbf{Y}=[y_1, \cdots, y_n]'$.

assumRP[Random Partition] \leavevmode \begin{enumerate}[label=(\roman*)] • $\widehat\Delta\perp \!\!\! \perp \mathbf{Y} |(\mathbf{X}, \mathbf{W})$ and $\widehat\Delta\in\Pi_C$ w.p.a. $1$ for some absolute constant $C>0$. • There exists a non-random partition $\Delta_0=\{\mathcal{B}_1, \cdots, \mathcal{B}_J\}$ with $\mathcal{B}_j=[\tau_{j-1}, \tau_j)$ for $j\leq J-1$ and $\mathcal{B}_J=[\tau_{J-1}, \tau_J]$ such that $\frac{\max_{1\leq j\leq J}h_j}{\min_{1\leq j\leq J}h_j}\leq c_{\tt QU}$ for some absolute constant $c_{\tt QU}>0$, and $\max_{1\leq j\leq J}|\hat{h}_j-h_j|\lesssim_\mathbb{P} J^{-1}\mathfrak{r}_{\tt RP}$ for $\mathfrak{r}_{\tt RP}=o(1)$. \end{enumerate}

Part (i) is the key condition for our main results and will be imposed throughout. First, it requires that the possibly random partition $\widehat\Delta$ be independent of the outcome $\mathbf{Y}$ given the covariates $(\mathbf{X},\mathbf{W})$. This conditional independence assumption is trivially satisfied if $\widehat\Delta$ is deterministic (e.g., equally-spaced partition) or depends on $\mathbf{X}$ and $\mathbf{W}$ only (e.g., quantile-spaced partition based on $\mathbf{X}$). It also holds if a sample splitting scheme is used: a subsample (including the information about the outcome) is used for constructing the partition, and the other is employed to construct the binscatter estimator (so that $\widehat\Delta$ is independent of the data $(\mathbf{X}, \mathbf{W}, \mathbf{Y})$). Second, $\widehat\Delta$ is required to be “quasi-uniform” with large probability. It is trivially true for equally-spaced partitions and can be verified for quantile-spaced partitions under the mild conditions on the covariates density imposed before (see Lemma (ref)). However, this condition may be too restrictive for other modern machine-learning-based partitioning methods, in which case some additional regularization may be necessary to recover the quasi-uniformity property.

Part (ii) requires that the random partition $\widehat\Delta$ “stabilizes” to a fixed one in large samples. This is true if the partition is non-deterministic or generated by sample quantiles (since sample quantiles converge to population quantiles), but more generally, it is not always possible. Fortunately, this “convergence” requirement is not necessary for most of our main results (except Theorem (ref) and Theorem (ref)). Thus, we will always make it clear if part (ii) of Assumption (ref) is imposed.

Given the random partition $\widehat{\Delta}$, we use the notation $\mathbb{E}_{\widehat{\Delta}}[\cdot]$ to denote the expectation operator with the partition $\widehat{\Delta}$ viewed as fixed. To further simplify notation, let $\hat{h}_j=\hat{\tau}_j-\hat{\tau}_{j-1}$ be the width of the $j$th bin $\widehat{\mathcal{B}}_j$, and when the “limiting” partition $\Delta_0=\{\mathcal{B}_1, \cdots, \mathcal{B}_J\}$ is defined (Assumption (ref)(ii) holds), let $h_j$ be the width of $\mathcal{B}_j$. Analogously to $\widehat{\mathbf{b}}_{p,s}(x)$, $\mathbf{b}_{p,s}(x)$ denotes the binscatter basis of degree $p$ that is $(s-1)$-times continuously differentiable and is constructed based on the nonrandom partition $\Delta_0$. We sometimes write $\mathbf{b}_{p,s}(x;\Delta)=(b_{p,s,1}(x;\Delta), \ldots, b_{p,s,K_{p,s}}(x;\Delta))'$ to emphasize a binscatter basis is constructed based on a particular partition $\Delta$. Therefore, $\widehat{\mathbf{b}}_{p,s}(x)=\mathbf{b}_{p,s}(x;\widehat{\Delta})$ and $\mathbf{b}_{p,s}(x)=\mathbf{b}_{p,s}(x;\Delta_0)$. Accordingly, we use $\mathbf{T}_s$ to denote the transformation matrix based on the non-random partition $\Delta_0$ (which transforms $\mathbf{b}_{p,0}(x)$ to $\mathbf{b}_{p,s}(x)$).

Main Results

We introduce the following quantities that will be extensively used throughout the supplement:

align*[align* omitted — 4,518 chars of source]

Recall that in the main text we always set $s=p$ and omit the dependence on $s$ whenever there is no confusion. Thus,

align*[align* omitted — 1,044 chars of source]

In this supplement, however, all our results hold for a general choice of the degree and the smoothness of the basis. For ease of notation, the subscripts $p$ and $s$ of the above quantities are dropped hereafter:

align*[align* omitted — 1,014 chars of source]

In addition, given the family $\Pi_C$ of the quasi-uniform partitions defined in (ref), for any $\Delta\in \Pi$, we let $\boldsymbol{\beta}_0(\Delta)\in\mathbb{R}^{K_{p,s}}$ be any vector such that for every $v\leq p$,

equation*[equation* omitted — 146 chars of source]

Let $r_{0,v}(x;\Delta)=\mu_0^{(v)}(x)-\mathbf{b}_{p,s}^{(v)}(x;\Delta)'\boldsymbol{\beta}_0(\Delta)$ denote the corresponding approximation error. Accordingly, given the random partition $\widehat\Delta$, we let $\widehat\boldsymbol{\beta}_0:=\boldsymbol{\beta}_0(\widehat\Delta)$, and $\widehat{r}_{0,v}(x)=\mu_0^{(v)}(x)-\widehat\mathbf{b}_{p,s}^{(v)}(x)'\widehat\boldsymbol{\beta}_0$ denote the corresponding approximation error. The existence of such vectors is guaranteed by Assumptions (ref) and (ref)(v), and is verified in Lemma (ref) in Section (ref).

Preliminary Lemmas

lem[Gram] Suppose that Assumptions (ref), (ref), (ref) and (ref)(i) hold. If $\frac{J\log J}{n}=o(1)$, then \[ 1\lesssim \lambda_{\min}(\bar{\mathbf{Q}})\leq\lambda_{\max}(\bar{\mathbf{Q}})\lesssim 1,\quad [\bar{\mathbf{Q}}^{-1}]_{ij}\lesssim \varrho^{|i-j|} \quad \text{w.p.a. } 1,\quad \text{and}\quad \|\bar{\mathbf{Q}}^{-1}\|_\infty\lesssim_\mathbb{P} 1, \] where $\varrho\in(0,1)$ is some absolute constant. If, in addition, Assumption (ref)(ii) holds. Then, \begin{eqnarray*} &1\lesssim \lambda_{\min}(\mathbf{Q}_0)\leq \lambda_{\max}(\mathbf{Q}_0)\lesssim 1,\\ &\|\bar{\mathbf{Q}}-\mathbf{Q}_0\| \lesssim_\mathbb{P} \Big(\frac{J\log J}{n}\Big)^{1/2}+\mathfrak{r}_{\tt RP}, \quadand\quad \|\bar{\mathbf{Q}}^{-1}-\mathbf{Q}_0^{-1}\|_\infty\lesssim_\mathbb{P} \Big(\frac{J\log J}{n}\Big)^{1/2}+\mathfrak{r}_{\tt RP}. \end{eqnarray*}

The next lemma shows that the limiting variance is bounded from above and below.

lem[Asymptotic Variance] Suppose that Assumptions (ref), (ref), (ref) and (ref)(i) hold. If $\frac{J\log J}{n}=o(1)$, then w.p.a. $1$, \begin{eqnarray*} &J^{1+2v}\lesssim\inf_{x\in\mathcal{X}}\bar{\Omega}_{\mu^{(v)}}(x) \leq\sup_{x\in\mathcal{X}}\bar{\Omega}_{\mu^{(v)}}(x) \lesssim J^{1+2v},\\ &J\lesssim \inf_{x\in\mathcal{X}}\bar{\Omega}_{\vartheta}(x) \leq\sup_{x\in\mathcal{X}}\bar{\Omega}_{\vartheta}(x) \lesssim J,\\ &J^3\lesssim \inf_{x\in\mathcal{X}}\bar{\Omega}_{\zeta}(x) \leq\sup_{x\in\mathcal{X}}\bar{\Omega}_{\zeta}(x) \lesssim J^3. \end{eqnarray*} If, in addition, Assumption (ref)(ii) holds, then w.p.a. $1$, \begin{eqnarray*} &J^{1+2v}\lesssim\inf_{x\in\mathcal{X}}\Omega_{\mu^{(v)}}(x) \leq\sup_{x\in\mathcal{X}}\Omega_{\mu^{(v)}}(x) \lesssim J^{1+2v},\\ &J\lesssim\inf_{x\in\mathcal{X}}\Omega_{\vartheta}(x) \leq\sup_{x\in\mathcal{X}}\Omega_{\vartheta}(x) \lesssim J,\\ &J^3\lesssim\inf_{x\in\mathcal{X}}\Omega_{\zeta}(x) \leq\sup_{x\in\mathcal{X}}\Omega_{\zeta}(x) \lesssim J^3. \end{eqnarray*}

The next lemma gives a bound on the variance component of the nonlinear binscatter estimator.

lem[Uniform Convergence: Variance] Suppose that Assumptions (ref), (ref), (ref) and (ref)(i) hold. If $\frac{J^{\frac{\nu}{\nu-2}}\log J}{n}=o(1)$, then \[ \sup_{x\in\mathcal{X}}\Big|\widehat{\mathbf{b}}_{p,s}^{(v)}(x)'\bar{\mathbf{Q}}^{-1}\mathbb{E}_n[\widehat{\mathbf{b}}_{p,s}(x_i)\eta_{i,1}\psi(y_i,\eta_i)]\Big|\lesssim_\mathbb{P} J^v\Big(\frac{J\log J}{n}\Big)^{1/2}. \]
lem[Projection of Approximation Error] Suppose that Assumptions (ref), (ref), (ref) and (ref)(i) hold. If $\frac{J^{\frac{\nu}{\nu-2}}\log J}{n}=o(1)$, then \[ \begin{split} &\sup_{x\in\mathcal{X}} \Big|\widehat{\mathbf{b}}_{p,s}^{(v)}(x)'\bar{\mathbf{Q}}^{-1} \mathbb{E}_n\Big[\widehat{\mathbf{b}}_{p,s}(x_{i})\Big(\eta_{i,1}\psi(y_i,\eta_i)- \eta^{(1)}(\widehat{\mathbf{b}}_{p,s}(x_i)'\widehat{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0) \psi(y_i,\eta(\widehat{\mathbf{b}}_{p,s}(x_i)'\widehat{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0))\Big)\Big]\Big|\\ \lesssim_\mathbb{P}&\; J^{-p-1+v}+J^{\frac{2v-p-1}{2}}\Big(\frac{J\log J}{n}\Big)^{1/2}+\frac{J^{1+v}\log J}{n}. \end{split} \]
lem[Uniform Consistency] Suppose that Assumptions (ref), (ref), (ref) and (ref)(i) hold. If $\frac{J^{\frac{2\nu}{\nu-1}}(\log J)^{\frac{\nu}{\nu-1}}}{n}=o(1)$, then \[ \|\widehat{\boldsymbol{\beta}}-\widehat{\boldsymbol{\beta}}_0\|_\infty=o_\mathbb{P}(J^{-1/2})\quad \text{and }\; \sup_{x\in\mathcal{X}}\,|\widehat{\mu}(x)-\mu_0(x)|=o_\mathbb{P}(1). \]
remark[Side rate conditions] When $\nu\to\infty$, the rate restriction $\frac{J^{\frac{2\nu}{\nu-1}}(\log J)^{\frac{\nu}{\nu-1}}}{n}=o(1)$ tends to be $\frac{J^2\log J}{n}=o(1)$. We conjecture this rate restriction is stronger than needed. In fact, for piecewise polynomials (i.e., $s=0$), we can show that $\frac{J^{\frac{\nu}{\nu-1}}(\log J)^{\frac{\nu}{\nu-1}}}{n}=o(1)$ suffices to establish the uniform consistency of $\widehat{\boldsymbol{\beta}}$, and this restriction is redundant in our main theorems in view of the condition $\frac{J^{\frac{\nu}{\nu-2}}(\log n)^{\frac{\nu}{\nu-2}}}{n}=o(1)$ imposed below. In other words, in this special case ($s=0$), the condition $\frac{J^{\frac{2\nu}{\nu-1}}(\log J)^{\frac{\nu}{\nu-1}}}{n}=o(1)$ in all theorems below can be dropped.

Our result holds without imposing any smoothness restrictions on the estimation space. Specifically, the estimation procedure (ref) searches for solutions in $\mathbb{R}^{K_{p,s}}$, leading to an estimation space $\{\widehat{\mathbf{b}}_{p,s}(x)'\boldsymbol{\beta}: \boldsymbol{\beta}\in\mathbb{R}^{K_{p,s}}\}$. In contrast, many studies of series (or sieve) methods restrict the functions in the estimation space to satisfy certain smoothness conditions, e.g., Lipschitz continuity, to derive the uniform consistency. See, for example, Chernozhukov-Imbens-Newey_2007_JoE and references therein.

remark[Improvements over literature] Most of the results in this subsection are new to the literature, even in the case of non-random partitioning and without covariate-adjustments, because they take advantage of the specific binscatter structure (i.e., locally bounded series basis). The closest antecedent in the literature is \citet*{Belloni-Chernozhukov-Chetverikov-FernandezVal_2019_JoE}, while it focuses on series-based quantile regression only. Furthermore, relative to prior work, our results allow for random partitioning schemes, formally taking into account both the potential randomness of the partition and the semi-linear regression estimation structure. Importantly, we highlight the key conditions on the possibly random partition (Assumptions (ref)(i) and (ref)(ii)) used to derive various properties of the Gram matrix, asymptotic variance and other quantities.

Bahadur Representation

thm[Bahadur Representation] Suppose that Assumptions (ref), (ref), (ref) and (ref)(i) hold and $\frac{J^{\frac{\nu}{\nu-2}}\log n}{n}+\frac{J(\log n)^{7/3}}{n}+\frac{J^{\frac{2\nu}{\nu-1}}(\log n)^{\frac{\nu}{\nu-1}}}{n}+\frac{\log n}{J} =o(1)$. Then, \begin{enumerate}[label=(\roman*)] • $\widehat\mu^{(v)}(x)$ satisfies that \[ \begin{split} &\sup_{x\in\mathcal{X}}\Big|\widehat{\mu}^{(v)}(x)-\mu_0^{(v)}(x)+\widehat{\mathbf{b}}_{p,s}^{(v)}(x)'\bar{\mathbf{Q}}^{-1}\mathbb{E}_n[\widehat{\mathbf{b}}_{p,s}(x_i)\eta_{i,1}\psi(y_i,\eta_i)]\Big|\\ \lesssim_\mathbb{P}&\; J^v\Big\{\Big(\frac{J\log n}{n}\Big)^{3/4}\log n+ J^{-\frac{p+1}{2}}\Big(\frac{J\log^2 n}{n}\Big)^{1/2}+J^{-p-1}+\mathfrak{r}_\gamma\Big\}. \end{split} \]$\widehat\vartheta(x,\widehat\mathsf{w})$ satisfies that \begin{align*} &\sup_{x\in\mathcal{X}}\Big|\widehat\vartheta(x,\widehat\mathsf{w})-\vartheta_0(x,\mathsf{w})+ \eta^{(1)}(\mu_0(x)+\mathsf{w}'\boldsymbol{\gamma}_0)\widehat{\mathbf{b}}_{p,s}(x)'\bar{\mathbf{Q}}^{-1}\mathbb{E}_n[\widehat\mathbf{b}_{p,s}(x_i)\eta_{i,1}\psi(y_i,\eta_i)]\Big|\\ \lesssim_\mathbb{P}&\; \Big(\frac{J\log n}{n}\Big)^{3/4}\log n+ J^{-\frac{p+1}{2}}\Big(\frac{J\log^2 n}{n}\Big)^{1/2}+J^{-p-1}+\mathfrak{r}_\gamma+\|\widehat\mathsf{w}-\mathsf{w}\|. \end{align*} • $\widehat\zeta(x,\widehat\mathsf{w})$ satisfies that \begin{align*} &\sup_{x\in\mathcal{X}}\Big|\widehat\zeta(x,\widehat\mathsf{w})-\zeta_0(x,\mathsf{w})+\eta^{(1)}(\mu_0(x)+\mathsf{w}'\boldsymbol{\gamma}_0)\widehat{\mathbf{b}}_{p,s}^{(1)}(x)'\bar{\mathbf{Q}}^{-1}\mathbb{E}_n[\widehat\mathbf{b}_{p,s}(x_i)\eta_{i,1}\psi(y_i,\eta_i)] \Big|\\[.5em] \lesssim_\mathbb{P}&\; \Big(\frac{J\log n}{n}\Big)^{1/2}+J\Big\{\Big(\frac{J\log n}{n}\Big)^{3/4}\log n+ J^{-\frac{p+1}{2}}\Big(\frac{J\log^2 n}{n}\Big)^{1/2}+J^{-p-1}+\mathfrak{r}_\gamma\Big\}\\ &+\|\widehat\mathsf{w}-\mathsf{w}\|\Big(1+J\Big(\frac{J\log n}{n}\Big)^{1/2}\Big). \end{align*} \end{enumerate}

The following corollary is an immediate result of Lemma (ref) and Theorem (ref), and hence its proof is omitted.

coro[Uniform Convergence] Suppose that the conditions of Theorem (ref) hold and $\frac{J(\log n)^5}{n}\lesssim 1$. Then \[ \sup_{x\in\mathcal{X}}|\widehat{\mu}^{(v)}(x)-\mu_0^{(v)}(x)| \lesssim_\mathbb{P} J^{v}\Big(\Big(\frac{J\log n}{n}\Big)^{1/2}+J^{-p-1}\Big). \] If, in addition, $\|\widehat\mathsf{w}-\mathsf{w}\|\lesssim_\mathbb{P} \Big(\frac{J\log n}{n}\Big)^{1/2}+J^{-p-1}$, then \begin{align*} &\sup_{x\in\mathcal{X}}|\widehat{\vartheta}(x,\widehat\mathsf{w})-\vartheta_0(x,\mathsf{w})| \lesssim_\mathbb{P} \Big(\frac{J\log n}{n}\Big)^{1/2}+J^{-p-1}\quad and\\ & \sup_{x\in\mathcal{X}}|\widehat{\zeta}(x,\widehat\mathsf{w})-\zeta_0(x,\mathsf{w})| \lesssim_\mathbb{P} J\Big(\Big(\frac{J\log n}{n}\Big)^{1/2}+J^{-p-1}\Big). \end{align*}

The next theorem shows that the proposed variance estimator is consistent.

thm[Variance Estimate] Suppose that Assumptions (ref), (ref), (ref) and (ref)(i) hold. If $\frac{J^{\frac{\nu}{\nu-2}}(\log n)^{\frac{\nu}{\nu-2}}}{n}+\frac{J^{\frac{2\nu}{\nu-1}}(\log n)^{\frac{\nu}{\nu-1}}}{n}+\frac{J(\log n)^5}{n}+\frac{\log n}{J}=o(1)$ and $\|\widehat\mathsf{w}-\mathsf{w}\|\lesssim_\mathbb{P} \Big(\frac{J\log n}{n}\Big)^{1/2}+J^{-p-1}$, then \begin{align*} &\Big\|\widehat{\boldsymbol{\Sigma}}-\bar\boldsymbol{\Sigma}\Big\| \lesssim_\mathbb{P} J^{-p-1}+\Big(\frac{J\log n}{n^{1-\frac{2}{\nu}}}\Big)^{1/2},\\ &\sup_{x\in\mathcal{X}}\Big|\widehat{\Omega}_{\mu^{(v)}}(x)-\bar\Omega_{\mu^{(v)}}(x)\Big|\lesssim_\mathbb{P} J^{1+2v}\Big(J^{-p-1}+\Big(\frac{J\log n}{n^{1-\frac{2}{\nu}}}\Big)^{1/2}\Big),\\ &\sup_{x\in\mathcal{X}}\Big|\widehat{\Omega}_{\vartheta}(x)-\bar\Omega_{\vartheta}(x)\Big|\lesssim_\mathbb{P} J\Big(J^{-p-1}+\Big(\frac{J\log n}{n^{1-\frac{2}{\nu}}}\Big)^{1/2}\Big),\quadand\\ &\sup_{x\in\mathcal{X}}\Big|\widehat{\Omega}_\zeta(x)-\bar\Omega_\zeta(x)\Big|\lesssim_\mathbb{P} J^3\Big(J^{-p-1}+\Big(\frac{J\log n}{n^{1-\frac{2}{\nu}}}\Big)^{1/2}\Big). \end{align*} If, in addition, Assumption (ref)(ii) holds, then \begin{align*} &\Big\|\widehat{\boldsymbol{\Sigma}}-\boldsymbol{\Sigma}_0\Big\| \lesssim_\mathbb{P} J^{-p-1}+\Big(\frac{J\log n}{n^{1-\frac{2}{\nu}}}\Big)^{1/2}+\mathfrak{r}_{\tt RP},\\ &\sup_{x\in\mathcal{X}}\Big|\widehat{\Omega}_{\mu^{(v)}}(x)-\Omega_{\mu^{(v)}}(x)\Big|\lesssim_\mathbb{P} J^{1+2v}\Big(J^{-p-1}+\Big(\frac{J\log n}{n^{1-\frac{2}{\nu}}}\Big)^{1/2}+\mathfrak{r}_{\tt RP}\Big),\\ &\sup_{x\in\mathcal{X}}\Big|\widehat{\Omega}_{\vartheta}(x)-\Omega_{\vartheta}(x)\Big|\lesssim_\mathbb{P} J\Big(J^{-p-1}+\Big(\frac{J\log n}{n^{1-\frac{2}{\nu}}}\Big)^{1/2}+\mathfrak{r}_{\tt RP}\Big),\quad and\\ &\sup_{x\in\mathcal{X}}\Big|\widehat{\Omega}_\zeta(x)-\Omega_\zeta(x)\Big|\lesssim_\mathbb{P} J^3\Big(J^{-p-1}+\Big(\frac{J\log n}{n^{1-\frac{2}{\nu}}}\Big)^{1/2}+\mathfrak{r}_{\tt RP}\Big). \end{align*}
remark[Improvements over literature] Theorem (ref) and Corollary (ref) construct the Bahadur representation and uniform convergence of nonlinear binscatter-based M-estimators, which improve upon prior results in the literature in at least two aspects. First, our results allow for random partitioning schemes, and the key condition imposed on the partition is Assumption (ref)(i), i.e., the conditional independence between the partition and the outcome and the quasi-uniformity of the partition. The “convergence” of the random partition (Assumption (ref)(ii)) is not required, which implies that our results can accommodate more complex partitioning schemes other than evenly-spaced or empirical-quantile-spaced partitions. Second, our results are established under weaker rate restrictions. Specifically, we require $J^{\frac{8}{3}}/n=o(1)$ up to $\log n$ terms when $\nu\geq 4$, thus accommodating IMSE-optimal piecewise constant binscatter estimators. In fact, for piecewise polynomials ($s=0$), we can show that the Bahadur representation still holds under $J/n=o(1)$ up to $\log n$ terms when a subexponential moment restriction holds for the “score” $\psi(y_i, \eta_i)$, which is analogous to the result for kernel-based estimators in the literature Kong-Linton-Xia_2010_ET. For series estimators, similar results were established for particular choices of loss functions under more stringent conditions in the literature. For example, \citet*{Belloni-Chernozhukov-Chetverikov-FernandezVal_2019_JoE} considers series-based quantile regression, and Theorem 2 and Corollary 2 therein can be used to establish a Bahadur representation and uniform convergence of the resulting estimators under $J^4/n^{1-\varepsilon}=o(1)$ for some $\varepsilon>0$. The results in Belloni-Chernozhukov-Chetverikov-FernandezVal_2019_JoE are slightly stronger than our Theorem (ref) in the sense that the expansion holds uniformly over both the evaluation point $x\in\mathcal{X}$ and the desired quantiles $u\in\mathcal{U}$ for a compact set of quantile indices $\mathcal{U}\subset(0,1)$. Our results regarding Bahadur representation can be extended to achieve the same level of uniformity. In general, the parameter of interest (ref) and the estimator (ref) are defined for each particular choice of the loss function within a function class $\mathcal{F}$. For the class of check functions used in quantile regression or other function classes with low complexity, it can be shown that the Bahadur representation still holds uniformly over the evaluation point $x\in\mathcal{X}$ and the loss function $\rho\in\mathcal{F}$ under rate restrictions similar to those in Theorem (ref), thereby providing an improvement over the literature.

Pointwise Inference

Starting from this section, we consider statistical inference on $\mu_0^{(v)}(x)$, $\vartheta_0(x,\mathsf{w})$ and $\zeta_0(x,\mathsf{w})$ based on the following Studentized $t$-statistics:

align*[align* omitted — 417 chars of source]

The next theorem shows the pointwise asymptotic normality of the binscatter estimators.

thm[Pointwise Asymptotic Distribution] Suppose that Assumptions (ref), (ref), (ref) and (ref)(i) hold, $\sup_{x\in\mathcal{X}}\mathbb{E}[|\psi(y_i,\eta_i)|^\nu|x_i=x]\lesssim 1$ for some $\nu\geq 3$, and $\frac{J^{\frac{\nu}{\nu-2}}(\log n)^{\frac{\nu}{\nu-2}}}{n} +\frac{J^{\frac{2\nu}{\nu-1}}(\log n)^{\frac{\nu}{\nu-1}}}{n}+nJ^{-2p-3} =o(1)$. Then the following conclusions hold: \begin{enumerate}[label=(\roman*)] • For $\widehat\mu^{(v)}(x)$, \[ \sup_{u\in\mathbb{R}}\Big|\mathbb{P}(T_{\mu^{(v)},p}(x)\leq u)-\Phi(u)\Big|=o(1),\quad \text{for each } x\in\mathcal{X}. \] • For $\widehat\vartheta(x,\widehat\mathsf{w})$, if, in addition, $\|\widehat\mathsf{w}-\mathsf{w}\|=o_\mathbb{P}(\sqrt{J/n})$, then \[ \sup_{u\in\mathbb{R}}\Big|\mathbb{P}(T_{\vartheta,p}(x)\leq u)-\Phi(u)\Big|=o(1)\quad \text{for each } x\in\mathcal{X}. \] • For $\widehat\zeta(x,\widehat\mathsf{w})$, if, in addition, $\|\widehat\mathsf{w}-\mathsf{w}\|=o_\mathbb{P}(\sqrt{J^3/n}+(\log n)^{-1/2})$, then \[ \sup_{u\in\mathbb{R}}\Big|\mathbb{P}(T_{\zeta,p}(x)\leq u)-\Phi(u)\Big|=o(1)\quad \text{for each } x\in\mathcal{X}. \] \end{enumerate}
remark[Improvements over literature] The result in this subsection is new to the literature, even in the case of non-random partitioning and without covariate adjustments, because it takes advantage of the specific binscatter structure (i.e., locally bounded series basis). The closest antecedent in the literature is Belloni-Chernozhukov-Chetverikov-FernandezVal_2019_JoE, which focuses on series-based quantile regression only. Furthermore, relative to prior work, our results allow for more general partitioning schemes, formally take into account the potential randomness of the partition, and account for the semi-linear regression estimation structure. The key condition imposed on the partition for pointwise inference is Assumption (ref)(i), and the “convergence” of the random partition is not required.

Integrated Mean Squared Error

In this section we give a Nagar-type approximate IMSE expansion for each of the three estimators $\widehat\mu^{(v)}(x)$, $\widehat\vartheta(x,\widehat\mathsf{w})$ and $\widehat\zeta(x,\widehat\mathsf{w})$, with explicit characterization of the leading constants. Define

align*[align* omitted — 161 chars of source]

where $\mathscr{E}_m(\cdot)$ is the $m$th Bernoulli polynomial for each $m\in\mathbb{Z}_+$, $\tau_x^{\mathtt{L}}$ is the start of the interval in the non-random partition $\Delta_0$ containing $x$ and $h_x$ denotes its length.

thm[IMSE] Suppose that Assumptions (ref), (ref), (ref) and (ref) (including (ref)(ii)) hold. Let $\omega(x)$ be a continuous weighting function over $\mathcal{X}$ bounded away from zero. Also, assume that $\frac{J^{\frac{\nu}{\nu-2}}\log n}{n}+\frac{J^{\frac{2\nu}{\nu-1}}(\log n)^{\frac{\nu}{\nu-1}}}{n}+\frac{J(\log n)^7}{n} +\frac{(\log n)^2}{J}=o(1)$. \begin{enumerate}[label=(\roman*)] • For $\widehat{\mu}^{(v)}(x)$, \[ \int_{\mathcal{X}}\Big(\widehat{\mu}^{(v)}(x)-\mu_0^{(v)}(x)\Big)^2\omega(x)dx =\mathtt{AISE}_{\mu^{(v)}}+o_\mathbb{P}\Big(\frac{J^{1+2v}}{n}+J^{-2(p+1-v)}\Big) \] where \[\begin{split} &\mathbb{E}[\mathtt{AISE}_{\mu^{(v)}}|\mathbf{X}, \mathbf{W},\widehat\Delta]=\frac{J^{1+2v}}{n}\mathscr{V}_n(p,s,v)+J^{-2(p+1-v)}\mathscr{B}_n(p,s,v) +o_\mathbb{P}\Big(\frac{J^{1+2v}}{n}+J^{-2(p+1-v)}\Big),\\ &\mathscr{V}_n(p,s,v):=J^{-(1+2v)}\operatorname*{trace}\Big( \mathbf{Q}_0^{-1}\boldsymbol{\Sigma}_0\mathbf{Q}_0^{-1} \int_{\mathcal{X}}\mathbf{b}_{p,s}^{(v)}(x)\mathbf{b}_{p,s}^{(v)}(x)'\omega(x)dx\Big) \asymp 1, \\ &\mathscr{B}_n(p,s,v):=J^{2p+2-2v}\int_{\mathcal{X}} \Big(r^\star_{0,v}(x)-\mathbf{b}_{p,s}^{(v)}(x)'\mathbf{Q}_0^{-1}\mathbb{E}[\mathbf{b}_{p,s}(x_i)\varkappa(x_i,\mathbf{w}_i)r^\star_{0,0}(x_i)]\Big)^2\omega(x)dx\lesssim 1. \end{split} \] • For $\widehat\vartheta(x,\widehat\mathsf{w})$, if $\|\widehat\mathsf{w}-\mathsf{w}\|=o_\mathbb{P}(\sqrt{J/n}+J^{-p-1})$, then \[ \int_{\mathcal{X}}\Big(\widehat{\vartheta}(x,\widehat\mathsf{w})-\vartheta_0(x,\mathsf{w})\Big)^2\omega(x)dx =\mathtt{AISE}_{\vartheta}+o_\mathbb{P}\Big(\frac{J}{n}+J^{-2(p+1)}\Big) \] where \[\begin{split} &\mathbb{E}[\mathtt{AISE}_{\vartheta}|\mathbf{X}, \mathbf{W},\widehat\Delta]=\frac{J}{n}\mathscr{V}_n(p,s)+J^{-2(p+1)}\mathscr{B}_n(p,s) +o_\mathbb{P}\Big(\frac{J}{n}+J^{-2(p+1)}\Big),\\ &\mathscr{V}_n(p,s):=J^{-1}\operatorname*{trace}\Big( \mathbf{Q}_0^{-1}\boldsymbol{\Sigma}_0\mathbf{Q}_0^{-1} \int_{\mathcal{X}}\eta_{0,1}(x,\mathsf{w})^2 \mathbf{b}_{p,s}(x)\mathbf{b}_{p,s}(x)'\omega(x)dx\Big) \asymp 1, \\ &\mathscr{B}_n(p,s):=J^{2p+2}\int_{\mathcal{X}} \Big[\eta_{0,1}(x,\mathsf{w}) \Big(r^\star_{0,0}(x)-\mathbf{b}_{p,s}(x)'\mathbf{Q}_0^{-1}\mathbb{E}[\mathbf{b}_{p,s}(x_i)\varkappa(x_i,\mathbf{w}_i)r^\star_{0,0}(x_i)]\Big)\Big]^2\omega(x)dx\lesssim 1. \end{split} \] • For $\widehat\zeta(x,\widehat\mathsf{w})$, if $\|\widehat\mathsf{w}-\mathsf{w}\|=o_\mathbb{P}(\sqrt{J^3/n}+J^{-p}+(\log n)^{-1/2})$, then \[ \int_{\mathcal{X}}\Big(\widehat{\zeta}(x,\widehat\mathsf{w})-\zeta_0(x,\mathsf{w})\Big)^2\omega(x)dx =\mathtt{AISE}_\zeta+o_\mathbb{P}\Big(\frac{J^{3}}{n}+J^{-2p}\Big) \] where \[\begin{split} &\mathbb{E}[\mathtt{AISE}_\zeta|\mathbf{X}, \mathbf{W},\widehat\Delta]=\frac{J^{3}}{n}\mathscr{V}_n(p,s) +J^{-2p}\mathscr{B}_n(p,s) +o_\mathbb{P}\Big(\frac{J^{3}}{n}+J^{-2p}\Big),\\ &\mathscr{V}_n(p,s):=J^{-3}\operatorname*{trace}\Big( \mathbf{Q}_0^{-1}\boldsymbol{\Sigma}_0\mathbf{Q}_0^{-1} \int_{\mathcal{X}}\eta_{0,1}(x,\mathsf{w})^2\mathbf{b}_{p,s}^{(1)}(x)\mathbf{b}_{p,s}^{(1)}(x)'\omega(x)dx\Big) \asymp 1, \\ &\mathscr{B}_n(p,s):=J^{2p}\int_{\mathcal{X}} \Big[\eta_{0,1}(x,\mathsf{w}) \Big(r^\star_{0,1}(x)-\mathbf{b}_{p,s}^{(1)}(x)'\mathbf{Q}_0^{-1}\mathbb{E}[\mathbf{b}_{p,s}(x_i)\varkappa(x_i,\mathbf{w}_i)r^\star_{0,0}(x_i)]\Big)\Big]^2\omega(x)dx\lesssim 1. \end{split} \] \end{enumerate}

In general, $\mathscr{B}_n(p,s,v)\gtrsim 1$ (see Remark SA-3.7 in Cattaneo-Crump-Farrell-Feng_2024_AER), and thus the above theorem implies that the (approximate) IMSE-optimal number of bins satisfies that $J_{\mathtt{AIMSE}}\asymp n^{\frac{1}{2p+3}}$. Relying on the IMSE expansion in Theorem (ref), one may design a data-driven procedure to select the IMSE-optimal number of bins for nonlinear binscatter-based M-estimators.

remark[Improvements over literature] The results in this subsection are new to the literature, even in the case of non-random partitioning and without covariate-adjustments, for both nonlinear series estimators and binscatter (piecewise polynomials and splines) nonlinear series estimators in particular. Furthermore, our results allow for random partitioning schemes, formally take into account the potential randomness of the partition, and account for the semi-linear regression estimation structure. We highlight the key conditions imposed on the partition (Assumption (ref)) for the approximate IMSE expansion. The “convergence” of the random partition (Assumption (ref)(ii)) is needed to derive the non-random variance and bias constants $\mathscr{V}_n(p,s)$ and $\mathscr{B}_n(p,s)$.

Uniform Inference

Recall that $(a_n: n\geq 1)$ is a sequence of non-vanishing constants. We will first show that the (feasible) Studentized $t$-statistic processes $T_{\mu^{(v)},p}(\cdot)$, $T_{\vartheta,p}(\cdot)$ and $T_{\zeta,p}(\cdot)$ can be approximated by Gaussian processes in a proper sense at certain rate.

thm[Strong Approximation] Suppose that Assumptions (ref), (ref), (ref) and (ref)(i) hold, \[ \frac{J(\log n)^2}{n^{1-\frac{2}{\nu}}} +\Big(\frac{J(\log n)^7}{n}\Big)^{1/2} +nJ^{-2p-3} +\frac{(\log n)^2}{J^{p+1}}+nJ^{-1}\mathfrak{r}_\gamma^2 =o(a_n^{-2}) \quad \text{and}\quad\frac{J^{\frac{2\nu}{\nu-1}}(\log n)^{\frac{\nu}{\nu-1}}}{n}=o(1). \] Then the following conclusions hold: \begin{enumerate}[label=(\roman*)] • On a properly enriched probability space, there exists some $K_{p,s}$-dimensional standard normal random vector $\mathbf{N}_{K_{p,s}}$ such that for any $\xi>0$, \[ \mathbb{P}\Big(\sup_{x\in\mathcal{X}}|T_{\mu^{(v)},p}(x)- \bar{Z}_{\mu^{(v)},p}(x)|>\xi a_n^{-1}\Big)=o(1), \quad \bar{Z}_{\mu^{(v)},p}(x)=\frac{\widehat{\mathbf{b}}_{p,0}^{(v)}(x)'\widehat{\mathbf{T}}_s'\bar\mathbf{Q}^{-1}\bar\boldsymbol{\Sigma}^{1/2}}{\sqrt{\bar\Omega_{\mu^{(v)}}(x)}}\mathbf{N}_{K_{p,s}}. \] If Assumption (ref)(ii) also holds with $\mathfrak{r}_{\tt RP}=o(a_n^{-1}(\log n)^{-1/2})$, then \[ \mathbb{P}\Big(\sup_{x\in\mathcal{X}}|T_{\mu^{(v)},p}(x)-Z_{\mu^{(v)},p}(x)|>\xi a_n^{-1}\Big)=o(1), \quad Z_{\mu^{(v)},p}(x)=\frac{\widehat{\mathbf{b}}_{p,0}^{(v)}(x)'\mathbf{T}_s'\mathbf{Q}_0^{-1}\boldsymbol{\Sigma}_0^{1/2}}{\sqrt{\Omega_{\mu^{(v)}}(x)}}\mathbf{N}_{K_{p,s}}. \] • If $\|\widehat\mathsf{w}-\mathsf{w}\|=o_\mathbb{P}(a_n^{-1}\sqrt{J/n})$, then on a properly enriched probability space there exists some $K_{p,s}$-dimensional standard normal random vector $\mathbf{N}_{K_{p,s}}$ such that for any $\xi>0$, \[ \mathbb{P}\bigg(\sup_{x\in\mathcal{X}}|T_{\vartheta,p}(x)- \bar{Z}_{\vartheta,p}(x)|>\xi a_n^{-1}\bigg)=o(1),\quad \bar{Z}_{\vartheta,p}(x)=\frac{\widehat{\mathbf{b}}_{p,0}(x)'\widehat\mathbf{T}_s'\eta_{0,1}(x,\mathsf{w}) \bar\mathbf{Q}^{-1}}{\sqrt{\bar\Omega_\vartheta(x)}}\bar\boldsymbol{\Sigma}^{1/2}\mathbf{N}_{K_{p,s}}. \] If Assumption (ref)(ii) also holds with $\mathfrak{r}_{\tt RP}=o(a_n^{-1}(\log n)^{-1/2})$, then \[ \mathbb{P}\bigg(\sup_{x\in\mathcal{X}}|T_{\vartheta,p}(x)- Z_{\vartheta,p}(x)|>\xi a_n^{-1}\bigg)=o(1),\quad Z_{\vartheta,p}(x)=\frac{\widehat{\mathbf{b}}_{p,0}(x)'\mathbf{T}_s'\eta_{0,1}(x,\mathsf{w})\mathbf{Q}_0^{-1}}{\sqrt{\Omega_\vartheta(x)}}\boldsymbol{\Sigma}_0^{1/2}\mathbf{N}_{K_{p,s}}. \] • If $\|\widehat\mathsf{w}-\mathsf{w}\|= o_\mathbb{P}(a_n^{-1}(\sqrt{J^3/n}+(\log n)^{-1/2}))$, then on a properly enriched probability space there exists some $K_{p,s}$-dimensional standard normal random vector $\mathbf{N}_{K_{p,s}}$ such that for any $\xi>0$, \[ \mathbb{P}\bigg(\sup_{x\in\mathcal{X}}|T_{\zeta,p}(x)- \bar{Z}_{\zeta,p}(x)|>\xi a_n^{-1}\bigg)=o(1),\quad \bar{Z}_{\zeta,p}(x)=\frac{\widehat{\mathbf{b}}_{p,0}^{(1)}(x)'\widehat\mathbf{T}_s'\eta_{0,1}(x,\mathsf{w})\bar\mathbf{Q}^{-1}}{\sqrt{\bar\Omega_{\zeta}(x)}}\bar\boldsymbol{\Sigma}^{1/2}\mathbf{N}_{K_{p,s}}. \] If Assumption (ref)(ii) also holds with $\mathfrak{r}_{\tt RP}=o(a_n^{-1}(\log n)^{-1/2})$, then \[ \mathbb{P}\bigg(\sup_{x\in\mathcal{X}}|T_{\zeta,p}(x)- Z_{\zeta,p}(x)|>\xi a_n^{-1}\bigg)=o(1),\quad Z_{\zeta,p}(x)=\frac{\widehat{\mathbf{b}}_{p,0}^{(1)}(x)'\mathbf{T}_s'\eta_{0,1}(x,\mathsf{w})\mathbf{Q}_0^{-1}}{\sqrt{\Omega_{\zeta}(x)}}\boldsymbol{\Sigma}_0^{1/2}\mathbf{N}_{K_{p,s}}. \] \end{enumerate}

The approximating processes $\bar{Z}_{\mu^{(v)},p}(\cdot)$, $\bar{Z}_{\vartheta,p}(\cdot)$ and $\bar{Z}_{\zeta,p}(\cdot)$ are Gaussian processes conditional on $\mathbf{X}$, $\mathbf{W}$ and $\widehat{\Delta}$, and $Z_{\mu^{(v)},p}(\cdot)$, $Z_{\vartheta,p}(\cdot)$ and $Z_{\zeta,p}(\cdot)$ are Gaussian processes conditional on $\widehat{\Delta}$ by construction. In practice, one can replace all unknowns in $\bar{Z}_{\mu^{(v)},p}(\cdot)$, $\bar{Z}_{\vartheta, p}(\cdot)$ and $\bar{Z}_{\zeta, p}(\cdot)$ (or $Z_{\mu^{(v)},p}(\cdot)$, $Z_{\vartheta, p}(\cdot)$ and $Z_{\zeta, p}(\cdot)$) by their sample analogues, and then construct the following feasible (conditional) Gaussian processes:

align*[align* omitted — 1,331 chars of source]

where $\mathbf{N}_{K_{p,s}}^\star$ denotes a $K_{p,s}$-dimensional standard normal vector independent of the data $\mathbf{D}$ and the partition $\widehat{\Delta}$.

For ease of presentation, we will always require a fast convergence rate of $\widehat\mathsf{w}$ hereafter: $\|\widehat\mathsf{w}-\mathsf{w}\|=o_\mathbb{P}(a_n^{-1}\sqrt{J/n})$. Nevertheless, note that as shown in Theorem (ref), such a rate restriction on $\widehat\mathsf{w}$ can be different for inference of $\vartheta_0(x,\mathsf{w})$ and $\zeta_0(x,\mathsf{w})$ and are unnecessary for inference of $\mu_0^{(v)}(x)$.

thm[Plug-in Approximation] Suppose that Assumptions (ref), (ref), (ref) and (ref)(i) hold, \begin{eqnarray*} &\frac{J(\log n)^2}{n^{1-\frac{2}{\nu}}} +\Big(\frac{J(\log n)^7}{n}\Big)^{1/2} +nJ^{-2p-3} +\frac{(\log n)^2}{J^{p+1}}+nJ^{-1}\mathfrak{r}_\gamma^2 =o(a_n^{-2}),\\ &\frac{J^{\frac{2\nu}{\nu-1}}(\log n)^{\frac{\nu}{\nu-1}}}{n}=o(1), \quadand\quad \|\widehat\mathsf{w}-\mathsf{w}\|=o_\mathbb{P}(a_n^{-1}\sqrt{J/n}). \end{eqnarray*} Then on a properly enriched probability space, there exists a $K_{p,s}$-dimensional standard normal random vector $\mathbf{N}_{K_{p,s}}^\star$ independent of $\mathbf{D}$ and $\widehat{\Delta}$ such that for any $\xi>0$, \begin{enumerate}[label=(\roman*)] • $ \mathbb{P}\Big(\sup_{x\in\mathcal{X}}|\widehat{Z}_{\mu^{(v)},p}(x)-\bar{Z}_{\mu^{(v)},p}(x)|>\xi a_n^{-1}\Big|\mathbf{D},\widehat{\Delta}\Big)=o_\mathbb{P}(1), $$ \mathbb{P}\Big(\sup_{x\in\mathcal{X}}|\widehat{Z}_{\vartheta,p}(x)-\bar{Z}_{\vartheta,p}(x)|>\xi a_n^{-1}\Big|\mathbf{D},\widehat{\Delta}\Big)=o_\mathbb{P}(1), $$ \mathbb{P}\Big(\sup_{x\in\mathcal{X}}|\widehat{Z}_{\zeta,p}(x)-\bar{Z}_{\zeta,p}(x)|>\xi a_n^{-1}\Big|\mathbf{D},\widehat{\Delta}\Big)=o_\mathbb{P}(1). $ \end{enumerate} If Assumption (ref)(ii) also holds with $\mathfrak{r}_{\tt RP}=o(a_n^{-1}(\log n)^{-1/2})$, then \begin{enumerate}[label=(\roman*)]\setcounter{enumi}{3} • $ \mathbb{P}\Big(\sup_{x\in\mathcal{X}}|\widehat{Z}_{\mu^{(v)},p}(x)-Z_{\mu^{(v)},p}(x)|>\xi a_n^{-1}\Big|\mathbf{D},\widehat{\Delta}\Big)=o_\mathbb{P}(1), $$ \mathbb{P}\Big(\sup_{x\in\mathcal{X}}|\widehat{Z}_{\vartheta,p}(x)-Z_{\vartheta,p}(x)|>\xi a_n^{-1}\Big|\mathbf{D},\widehat{\Delta}\Big)=o_\mathbb{P}(1), $$ \mathbb{P}\Big(\sup_{x\in\mathcal{X}}|\widehat{Z}_{\zeta,p}(x)-Z_{\zeta,p}(x)|>\xi a_n^{-1}\Big|\mathbf{D},\widehat{\Delta}\Big)=o_\mathbb{P}(1). $ \end{enumerate}
remark[Improvements over literature] Theorems (ref) and (ref) provide empirical researchers with powerful tools for uniform inference based on binscatter methods. Importantly, we allow for random partitioning schemes, formally take into account the potential randomness of the partition, and construct a novel strong approximation of nonlinear binscatter-based M-estimators under mild rate restrictions. For $a_n=\sqrt{\log n}$ and $\nu\geq 4$, we require $J^{\frac{8}{3}}/n=o(1)$, up to $\log n$ terms. In the literature, similar results were only available in some special cases under stringent rate restrictions. For instance, Belloni-Chernozhukov-Chetverikov-FernandezVal_2019_JoE considers strong approximations of more general series-based quantile regression estimators. For the binscatter basis considered in this paper, their Theorem 11 can be applied to construct strong approximation of the $t$-statistic process based on pivotal coupling that achieves the approximation rate $a_n=n^{-\varepsilon'}$ under $J^4/n^{1-\varepsilon}=o(1)$ for some constants $\varepsilon, \varepsilon'>0$, whereas their Theorem 12 can be used to construct strong approximation based on Gaussian processes under $J^{5}/n^{1-\varepsilon}=o(1)$. It should be noted that their notion of strong approximation is stronger than ours in the sense that it holds uniformly over both the evaluation point $x\in\mathcal{X}$ and the desired quantile $u\in\mathcal{U}$ for a compact set of quantile indices $\mathcal{U}\subset(0,1)$. On the other hand, our methods allow for other loss functions (e.g., Huber regression), a large class of random partitions, and semi-linear covariate adjustment, leading to new results that were previously unavailable in the literature.

Theorems (ref) and (ref) offer a way to approximate the distribution of the whole $t$-statistic process based on $\widehat\mu^{(v)}(\cdot)$, $\widehat\vartheta(\cdot, \widehat\mathsf{w})$ or $\widehat{\zeta}(\cdot,\widehat\mathsf{w})$. A direct application of these results is the distributional approximations to the suprema of these $t$-statistic processes.

thm[Supremum Approximation] Suppose that Assumptions (ref), (ref), (ref) and (ref) (including (ref)(ii)) hold, \begin{eqnarray*} &\frac{J(\log n)^2}{n^{1-\frac{2}{\nu}}} +nJ^{-2p-3} +nJ^{-1}\mathfrak{r}_\gamma^2 =o((\log J)^{-1}),\\ &\frac{J^{\frac{2\nu}{\nu-1}}(\log n)^{\frac{\nu}{\nu-1}}}{n}=o(1),\quad \|\widehat\mathsf{w}-\mathsf{w}\|=o_\mathbb{P}\Big(\sqrt{\frac{J}{n\log J}}\Big), \quadand \quad \mathfrak{r}_{\tt RP}=o\Big(\frac{1}{\sqrt{\log n\,\log J}}\Big). \end{eqnarray*} Then, \begin{align*} &\sup_{u\in\mathbb{R}}\Big|\mathbb{P}\Big(\sup_{x\in\mathcal{X}} |T_{\mu^{(v)},p}(x)|\leq u\Big)- \mathbb{P}\Big(\sup_{x\in\mathcal{X}}|\widehat{Z}_{\mu^{(v)},p}(x)|\leq u\Big|\mathbf{D},\widehat{\Delta}\Big)\Big|=o_\mathbb{P}(1),\\ &\sup_{u\in\mathbb{R}}\Big|\mathbb{P}\Big(\sup_{x\in\mathcal{X}} |T_{\vartheta,p}(x)|\leq u\Big)- \mathbb{P}\Big(\sup_{x\in\mathcal{X}}|\widehat{Z}_{\vartheta,p}(x)|\leq u\Big|\mathbf{D},\widehat{\Delta}\Big)\Big|=o_\mathbb{P}(1),\quad and\\ &\sup_{u\in\mathbb{R}}\Big|\mathbb{P}\Big(\sup_{x\in\mathcal{X}} |T_{\zeta,p}(x)|\leq u\Big)- \mathbb{P}\Big(\sup_{x\in\mathcal{X}}|\widehat{Z}_{\zeta,p}(x)|\leq u\Big|\mathbf{D},\widehat{\Delta}\Big)\Big|=o_\mathbb{P}(1). \end{align*}

Confidence Bands

Let

align*[align* omitted — 467 chars of source]

be confidence bands for $\mu_0^{(v)}(\cdot)$, $\vartheta_0(\cdot,\mathsf{w})$ and $\zeta_0(\cdot,\mathsf{w})$ respectively, where $\mathfrak{c}_{\mu^{(v)}}$, $\mathfrak{c}_\vartheta$ and $\mathfrak{c}_\zeta$ are corresponding critical values to be specified. Recall that $\mathsf{w}$ here is taken as a fixed evaluation point for the control variables, and these bands are constructed based on a certain choice of $J$ and the $p$th-order binscatter basis. Using the previous results, we have the following theorem.

thmSuppose that Assumptions (ref), (ref), (ref) and (ref)(i) hold, \begin{eqnarray*} &\frac{J(\log n)^2}{n^{1-\frac{2}{\nu}}} +nJ^{-2p-3} +nJ^{-1}\mathfrak{r}_\gamma^2 =o((\log J)^{-1}),\\ &\frac{J^{\frac{2\nu}{\nu-1}}(\log n)^{\frac{\nu}{\nu-1}}}{n}=o(1),\quadand\quad\|\widehat\mathsf{w}-\mathsf{w}\|=o_\mathbb{P}\Big(\sqrt{\frac{J}{n\log J}}\Big). \end{eqnarray*} \begin{enumerate}[label=(\roman*)] • If $\mathfrak{c}_{\mu^{(v)}}=\inf\Big\{c\in\mathbb{R}_+:\mathbb{P}[\sup_{x\in\mathcal{X}}|\widehat{Z}_{\mu^{(v)},p}(x)|\leq c \;|\mathbf{D},\widehat{\Delta}]\geq 1-\alpha\Big\}$, then \[ \mathbb{P}\Big[\mu_0^{(v)}(x)\in\widehat{I}_{\mu^{(v)},p}(x),\text{ for all }x\in\mathcal{X}\Big]=1-\alpha+o(1). \] • If $\mathfrak{c}_\vartheta=\inf\Big\{c\in\mathbb{R}_+:\mathbb{P}[\sup_{x\in\mathcal{X}}|\widehat{Z}_{\vartheta,p}(x)|\leq c \;|\mathbf{D},\widehat{\Delta}]\geq 1-\alpha\Big\}$, then \[ \mathbb{P}\Big[\vartheta_0(x,\mathsf{w})\in\widehat{I}_{\vartheta,p}(x,\mathsf{w}),\text{ for all }x\in\mathcal{X}\Big]=1-\alpha+o(1). \] • If $\mathfrak{c}_\zeta=\inf\Big\{c\in\mathbb{R}_+:\mathbb{P}[\sup_{x\in\mathcal{X}}|\widehat{Z}_{\zeta,p}(x)|\leq c \;|\mathbf{D},\widehat{\Delta}]\geq 1-\alpha\Big\}$, then \[ \mathbb{P}\Big[\zeta_0(x,\mathsf{w})\in\widehat{I}_{\zeta,p}(x,\mathsf{w}),\text{ for all }x\in\mathcal{X}\Big]=1-\alpha+o(1). \] \end{enumerate}
remarkThe above results construct valid uniform confidence bands for nonlinear binscatter-based M-estimators under mild rate restrictions. Specifically, when $\nu\geq 4$, we require $J^{\frac{8}{3}}/n=o(1)$, up to $\log n$ terms. In contrast, Belloni-Chernozhukov-Chetverikov-FernandezVal_2019_JoE considers more general series-based quantile regression estimators, and Theorem 15 therein can be used to construct confidence bands for binscatter estimators via various resampling methods under $J^4/n^{1-\varepsilon}=o(1)$ for some $\varepsilon>0$. Furthermore, our results allow for random partitioning schemes, formally taking its randomness and generic structure. The key condition imposed on the partition for the validity of confidence bands is Assumption (ref)(i), but the “convergence” of the random partition (Assumption (ref)(ii)) is not necessary.

Parametric Specification Tests

As another application, we can test parametric specifications of $\mu_0^{(v)}(x)$, $\vartheta_0(x,\mathsf{w})$ and $\zeta_0(x,\mathsf{w})$. We introduce the following tests:

alignat*{2} &\dot{\mathsf{H}}_0^{\mu^{(v)}}:\quad &&\sup_{x\in\mathcal{X}} \Big|\mu_0^{(v)}(x) - m^{(v)}(x;\boldsymbol{\theta})\Big|=0, \quad for some \boldsymbol{\theta}, \qquad vs.\\ &\dot{\mathsf{H}}_A^{\mu^{(v)}}: &&\sup_{x\in\mathcal{X}} \Big|\mu_0^{(v)}(x) - m^{(v)}(x;\boldsymbol{\theta})\Big|>0, \quad for all \boldsymbol{\theta}.

where $m(x;\boldsymbol{\theta})$ is some known function depending on some finite dimensional parameter $\boldsymbol{\theta}$. This testing problem can be viewed as a two-sided test where the equality between two functions holds uniformly over $x\in\mathcal{X}$. In this case, we introduce $\widetilde{\boldsymbol{\theta}}$ and $\widetilde{\boldsymbol{\gamma}}$ as consistent estimators of $\boldsymbol{\theta}$ and $\boldsymbol{\gamma}_0$ under $\dot{\mathsf{H}}_0^{\mu^{(v)}}$. Then we rely on the following test statistic: \[ \dot{T}_{\mu^{(v)},p}(x) :=\frac{\widehat{\mu}^{(v)}(x)-m^{(v)}(x;\widetilde{\boldsymbol{\theta}})} {\sqrt{\widehat{\Omega}_{\mu^{(v)}}(x)/n}}. \] The null hypothesis is rejected if $\sup_{x\in\mathcal{X}}|\dot{T}_{\mu^{(v)},p}(x)|>\mathfrak{c}_{\mu^{(v)}}$ for some critical value $\mathfrak{c}_{\mu^{(v)}}$.

Similarly, to test the specification of $\vartheta_0(x,\mathsf{w})$, we introduce

alignat*{2} &\dot{\mathsf{H}}_0^\vartheta:\quad &&\sup_{x\in\mathcal{X}} \Big|\vartheta_0(x,\mathsf{w}) - M(x,\mathsf{w};\boldsymbol{\theta},\boldsymbol{\gamma}_0)\Big|=0, \quad for some \boldsymbol{\theta}, \qquad vs.\\ &\dot{\mathsf{H}}_A^\vartheta: &&\sup_{x\in\mathcal{X}} \Big|\vartheta_0(x,\mathsf{w}) - M(x,\mathsf{w};\boldsymbol{\theta},\boldsymbol{\gamma}_0)\Big|>0, \quad for all \boldsymbol{\theta}.

where $M(x,\mathsf{w};\boldsymbol{\theta},\boldsymbol{\gamma}_0)=\eta(m(x;\boldsymbol{\theta})+\mathsf{w}'\boldsymbol{\gamma}_0)$. We rely on the following test statistic: \[ \dot{T}_{\vartheta,p}(x) :=\frac{\widehat{\vartheta}(x,\widehat\mathsf{w})-M(x,\widehat\mathsf{w};\widetilde{\boldsymbol{\theta}},\widetilde\boldsymbol{\gamma})} {\sqrt{\widehat{\Omega}_\vartheta(x)/n}}. \] The null hypothesis is rejected if $\sup_{x\in\mathcal{X}}|\dot{T}_{\vartheta,p}(x)|>\mathfrak{c}_\vartheta$ for some critical value $\mathfrak{c}_\vartheta$.

To test the specification of $\zeta_0(x,\mathsf{w})$, we introduce

alignat*{2} &\dot{\mathsf{H}}_0^\zeta:\quad &&\sup_{x\in\mathcal{X}} \Big|\zeta_0(x,\mathsf{w}) - M^{(1)}(x,\mathsf{w};\boldsymbol{\theta},\boldsymbol{\gamma}_0)\Big|=0, \quad for some \boldsymbol{\theta}, \qquad vs.\\ &\dot{\mathsf{H}}_A^\zeta: &&\sup_{x\in\mathcal{X}} \Big|\zeta_0(x,\mathsf{w}) - M^{(1)}(x,\mathsf{w};\boldsymbol{\theta},\boldsymbol{\gamma}_0)\Big|>0, \quad for all \boldsymbol{\theta}.

where $M^{(1)}(x,\mathsf{w};\boldsymbol{\theta},\boldsymbol{\gamma}_0):=\eta^{(1)}(m(x;\boldsymbol{\theta})+\mathsf{w}'\boldsymbol{\gamma}_0)m^{(1)}(x;\boldsymbol{\theta})$. We rely on the following test statistic: \[ \dot{T}_{\zeta,p}(x) :=\frac{\widehat{\zeta}(x,\widehat\mathsf{w})-M^{(1)}(x,\widehat\mathsf{w};\widetilde{\boldsymbol{\theta}},\widetilde\boldsymbol{\gamma})} {\sqrt{\widehat{\Omega}_\zeta(x)/n}}. \] The null hypothesis is rejected if $\sup_{x\in\mathcal{X}}|\dot{T}_{\zeta,p}(x)|>\mathfrak{c}_\zeta$ for some critical value $\mathfrak{c}_\zeta$.

thm[Specification Tests] Suppose that the conditions in Theorem (ref) hold. \begin{enumerate}[label=(\roman*)] • Let $\mathfrak{c}_{\mu^{(v)}}=\inf\{c\in\mathbb{R}_+: \mathbb{P}[\sup_{x\in\mathcal{X}}|\widehat{Z}_{\mu^{(v)},p}(x)|\leq c |\mathbf{D},\widehat{\Delta}]\geq 1-\alpha \}$. Under $\dot{\mathsf{H}}_0^{\mu^{(v)}}$, if $\sup_{x\in\mathcal{X}} |\mu^{(v)}(x)-m^{(v)} (x;\widetilde{\boldsymbol{\theta}})|=o_\mathbb{P}\Big(\sqrt{\frac{J^{1+2v}}{n\log J}}\Big)$, then $$ \lim_{n\to\infty}\mathbb{P}\Big[\sup_{x\in\mathcal{X}}|\dot{T}_{\mu^{(v)},p}(x)|> \mathfrak{c}_{\mu^{(v)}}\Big]=\alpha. $$ Under $\dot{\mathsf{H}}_{\text{A}}^{\mu^{(v)}}$, if there exist some fixed $\bar{\boldsymbol{\theta}}$ such that $\sup_{x\in\mathcal{X}} |m^{(v)}(x;\widetilde{\boldsymbol{\theta}})-m^{(v)}(x;\bar{\boldsymbol{\theta}})|=o_\mathbb{P}(1)$, and $J^v\Big(\frac{J\log J}{n}\Big)^{1/2}=o(1)$, then \[ \lim_{n\to\infty}\mathbb{P}\Big[\sup_{x\in\mathcal{X}} |\dot{T}_{\mu^{(v)},p}(x)|>\mathfrak{c}_{\mu^{(v)}}\Big]=1. \] • Let $\mathfrak{c}_\vartheta=\inf\{c\in\mathbb{R}_+: \mathbb{P}[\sup_{x\in\mathcal{X}}|\widehat{Z}_{\vartheta,p}(x)|\leq c |\mathbf{D},\widehat{\Delta}]\geq 1-\alpha \}$. Under $\dot{\mathsf{H}}_0^\vartheta$, if $\sup_{x\in\mathcal{X}} |\vartheta_0(x,\mathsf{w})- M(x,\widehat\mathsf{w};\widetilde{\boldsymbol{\theta}},\widetilde{\boldsymbol{\gamma}})|=o_\mathbb{P}\Big(\sqrt{\frac{J^{1+2v}}{n\log J}}\Big)$, then $$ \lim_{n\to\infty}\mathbb{P}\Big[\sup_{x\in\mathcal{X}}|\dot{T}_{\vartheta,p}(x)|> \mathfrak{c}\Big]=\alpha. $$ Under $\dot{\mathsf{H}}_{\text{A}}^\vartheta$, if there exist some fixed $\bar{\boldsymbol{\theta}}$ and $\bar{\boldsymbol{\gamma}}$ such that $\sup_{x\in\mathcal{X}} |M(x,\widehat\mathsf{w};\widetilde{\boldsymbol{\theta}},\widetilde{\boldsymbol{\gamma}})-M(x,\mathsf{w};\bar{\boldsymbol{\theta}}, \bar{\boldsymbol{\gamma}})|=o_\mathbb{P}(1)$, and $J^v\Big(\frac{J\log J}{n}\Big)^{1/2}=o(1)$, then \[ \lim_{n\to\infty}\mathbb{P}\Big[\sup_{x\in\mathcal{X}} |\dot{T}_{\vartheta,p}(x)|>\mathfrak{c}\Big]=1. \] • Let $\mathfrak{c}_\zeta=\inf\{c\in\mathbb{R}_+: \mathbb{P}[\sup_{x\in\mathcal{X}}|\widehat{Z}_{\zeta,p}(x)|\leq c |\mathbf{D},\widehat{\Delta}]\geq 1-\alpha \}$. Under $\dot{\mathsf{H}}_0^\zeta$, if $\sup_{x\in\mathcal{X}} |\zeta_0(x,\mathsf{w})-M^{(1)} (x,\widehat\mathsf{w};\widetilde{\boldsymbol{\theta}},\widetilde{\boldsymbol{\gamma}})|=o_\mathbb{P}\Big(\sqrt{\frac{J^{1+2v}}{n\log J}}\Big)$, then $$ \lim_{n\to\infty}\mathbb{P}\Big[\sup_{x\in\mathcal{X}}|\dot{T}_{\zeta,p}(x)|> \mathfrak{c}\Big]=\alpha. $$ Under $\dot{\mathsf{H}}_{\text{A}}^\zeta$, if there exist some fixed $\bar{\boldsymbol{\theta}}$ and $\bar{\boldsymbol{\gamma}}$ such that $\sup_{x\in\mathcal{X}} |M^{(1)}(x,\widehat\mathsf{w};\widetilde{\boldsymbol{\theta}},\widetilde{\boldsymbol{\gamma}})-M^{(1)}(x,\mathsf{w};\bar{\boldsymbol{\theta}}, \bar{\boldsymbol{\gamma}})|=o_\mathbb{P}(1)$, and $J^v\Big(\frac{J\log J}{n}\Big)^{1/2}=o(1)$, then \[ \lim_{n\to\infty}\mathbb{P}\Big[\sup_{x\in\mathcal{X}} |\dot{T}_{\zeta,p}(x)|>\mathfrak{c}\Big]=1. \] \end{enumerate}

Shape Restriction Tests

The third application of our results is to test certain shape restrictions on $\mu_0^{(v)}(x)$, $\vartheta_0(x,\mathsf{w})$ and $\zeta_0(x,\mathsf{w})$. To be specific, consider the following problem: \[

split[split omitted — 442 chars of source]

\]

This testing problem can be viewed as a one-sided test where the inequality holds uniformly over $x\in\mathcal{X}$. Importantly, it should be noted that under both $\ddot{\mathsf{H}}_0^{\mu^{(v)}}$ and $\ddot{\mathsf{H}}_\text{A}^{\mu^{(v)}}$, we fix $\bar{\boldsymbol{\theta}}$ and $\bar{\boldsymbol{\gamma}}$ to be the same values in the parameter space. In such a case, we introduce $\widetilde{\boldsymbol{\theta}}$ and $\widetilde{\boldsymbol{\gamma}}$ as consistent estimators of $\bar{\boldsymbol{\theta}}$ and $\bar{\boldsymbol{\gamma}}$ under both $\ddot{\mathsf{H}}_0^{\mu^{(v)}}$ and $\ddot{\mathsf{H}}_\text{A}^{\mu^{(v)}}$. Then we will rely on the following test statistic: \[ \ddot{T}_{\mu^{(v)},p}(x):= \frac{\widehat{\mu}^{(v)}(x)-m^{(v)}(x;\widetilde{\boldsymbol{\theta}})} {\sqrt{\widehat{\Omega}_{\mu^{(v)}}(x)/n}}. \] The null hypothesis is rejected if $\sup_{x\in\mathcal{X}} \ddot{T}_{\mu^{(v)},p}(x)>\mathfrak{c}_{\mu^{(v)}}$ for some critical value $\mathfrak{c}_{\mu^{(v)}}$.

Similarly, define the test for the shape of $\vartheta_0(x,\mathsf{w})$: \[

split[split omitted — 526 chars of source]

\] We will rely on the following test statistic: \[ \ddot{T}_{\vartheta,p}(x):= \frac{\widehat{\vartheta}(x,\widehat\mathsf{w})-M(x,\widehat\mathsf{w};\widetilde{\boldsymbol{\theta}},\widetilde{\boldsymbol{\gamma}})} {\sqrt{\widehat{\Omega}_\vartheta(x)/n}}. \] The null hypothesis is rejected if $\sup_{x\in\mathcal{X}} \ddot{T}_{\vartheta,p}(x)>\mathfrak{c}_\vartheta$ for some critical value $\mathfrak{c}_\vartheta$.

Also, define the test for the shape of $\zeta_0(x,\mathsf{w})$: \[

split[split omitted — 522 chars of source]

\] We will rely on the following test statistic: \[ \ddot{T}_{\zeta,p}(x):= \frac{\widehat{\zeta}(x,\widehat\mathsf{w})-M^{(1)}(x,\widehat\mathsf{w};\widetilde{\boldsymbol{\theta}},\widetilde{\boldsymbol{\gamma}})} {\sqrt{\widehat{\Omega}_\zeta(x)/n}}. \] The null hypothesis is rejected if $\sup_{x\in\mathcal{X}} \ddot{T}_{\zeta,p}(x)>\mathfrak{c}_\zeta$ for some critical value $\mathfrak{c}_\zeta$.

The following theorem characterizes the size and power of such tests.

thm[Shape Restriction Tests] Suppose that the conditions in Theorem (ref) hold. \begin{enumerate}[label=(\roman*)] • Assume $\sup_{x\in\mathcal{X}} |m(x;\widetilde{\boldsymbol{\theta}})- m(x;\bar{\boldsymbol{\theta}})|= o_\mathbb{P}\Big(\sqrt{\frac{J^{1+2v}}{n\log J}}\Big)$. Let $\mathfrak{c}_{\mu^{(v)}}=\inf\{c\in\mathbb{R}_+: \mathbb{P}[\sup_{x\in\mathcal{X}}\widehat{Z}_{\mu^{(v)},p}(x)\leq c |\mathbf{D},\widehat{\Delta}]\geq 1-\alpha \}$. Under $\ddot{\mathsf{H}}_0^{\mu^{(v)}}$, \[ \lim_{n\to\infty}\mathbb{P}\Big[\sup_{x\in\mathcal{X}} \ddot{T}_{\mu^{(v)},p}(x)>\mathfrak{c}_{\mu^{(v)}}\Big]\leq \alpha. \] Under $\ddot{\mathsf{H}}_\text{A}^{\mu^{(v)}}$, if $J^v\Big(\frac{J\log J}{n}\Big)^{1/2}=o(1)$, \[ \lim_{n\to\infty}\mathbb{P}\Big[\sup_{x\in\mathcal{X}} \ddot{T}_{\mu^{(v)},p}(x)>\mathfrak{c}_{\mu^{(v)}}\Big]=1. \] • Assume $\sup_{x\in\mathcal{X}} |M(x,\widehat\mathsf{w};\widetilde{\boldsymbol{\theta}},\widetilde{\boldsymbol{\gamma}})- M(x,\mathsf{w};\bar{\boldsymbol{\theta}},\bar{\boldsymbol{\gamma}})|= o_\mathbb{P}\Big(\sqrt{\frac{J^{1+2v}}{n\log J}}\Big)$. Let $\mathfrak{c}_\vartheta=\inf\{c\in\mathbb{R}_+: \mathbb{P}[\sup_{x\in\mathcal{X}}\widehat{Z}_{\vartheta,p}(x)\leq c |\mathbf{D},\widehat{\Delta}]\geq 1-\alpha \}$. Under $\ddot{\mathsf{H}}_0^\vartheta$, \[ \lim_{n\to\infty}\mathbb{P}\Big[\sup_{x\in\mathcal{X}} \ddot{T}_{\vartheta,p}(x)>\mathfrak{c}_\vartheta\Big]\leq \alpha. \] Under $\ddot{\mathsf{H}}_\text{A}^\vartheta$, if $J^v\Big(\frac{J\log J}{n}\Big)^{1/2}=o(1)$, \[ \lim_{n\to\infty}\mathbb{P}\Big[\sup_{x\in\mathcal{X}} \ddot{T}_{\vartheta,p}(x)>\mathfrak{c}_\vartheta\Big]=1. \] • Assume $\sup_{x\in\mathcal{X}} |M^{(1)}(x,\widehat\mathsf{w};\widetilde{\boldsymbol{\theta}},\widetilde{\boldsymbol{\gamma}})- M^{(1)}(x,\mathsf{w};\bar{\boldsymbol{\theta}},\bar{\boldsymbol{\gamma}})|= o_\mathbb{P}\Big(\sqrt{\frac{J^{1+2v}}{n\log J}}\Big)$. Let $\mathfrak{c}_\zeta=\inf\{c\in\mathbb{R}_+: \mathbb{P}[\sup_{x\in\mathcal{X}}\widehat{Z}_{\zeta,p}(x)\leq c |\mathbf{D},\widehat{\Delta}]\geq 1-\alpha \}$. Under $\ddot{\mathsf{H}}_0^\zeta$, \[ \lim_{n\to\infty}\mathbb{P}\Big[\sup_{x\in\mathcal{X}} \ddot{T}_{\zeta,p}(x)>\mathfrak{c}_\zeta\Big]\leq \alpha. \] Under $\ddot{\mathsf{H}}_\text{A}^\zeta$, if $J^v\Big(\frac{J\log J}{n}\Big)^{1/2}=o(1)$, \[ \lim_{n\to\infty}\mathbb{P}\Big[\sup_{x\in\mathcal{X}} \ddot{T}_{\zeta,p}(x)>\mathfrak{c}_\zeta\Big]=1. \] \end{enumerate}
remark[Improvements over literature] The results in Sections (ref)--(ref) are new to the literature, even in the case of non-random partitioning and without covariate-adjustments, because they take advantage of the specific binscatter structure (i.e., locally bounded series basis). Furthermore, relative to prior work, our results allow for a large class of random partitioning schemes, formally take into account the potential randomness of the partition, account for the generalized semi-linear structure, and consider an array of possibly nonlinear estimation and inference problems. In particular, the approach taken in Theorems (ref) and (ref) to establish strong approximation and related distributional approximations for nonlinear binscatter statistics may be of independent interest. The key condition imposed on the partition for uniform inference (confidence bands and hypothesis testing) is Assumption (ref)(i), while “convergence” of the random partition (Assumption (ref)(ii)) is not required.

Implementation Details

Standard Error Computation

In Section (ref), we have given the variance formulas $\widehat\Omega_{\mu^{(v)}}(x)$, $\widehat\Omega_\vartheta(x)$ and $\widehat\Omega_\zeta(x)$ that can be used to obtain the standard errors of $\widehat{\mu}^{(v)}(x)$, $\widehat{\vartheta}(x,\widehat\mathsf{w})$ and $\widehat{\zeta}(x,\widehat\mathsf{w})$. Recall that the formula for the estimator $\widehat{\boldsymbol{\Sigma}}$ of $\boldsymbol{\Sigma}_0$ is $$ \widehat{\boldsymbol{\Sigma}}= \mathbb{E}_n\Big[\widehat{\mathbf{b}}_{p,s}(x_i)\widehat{\mathbf{b}}_{p,s}(x_i)' \psi(y_i, \widehat{\eta}_i)^2\eta^{(1)}(\widehat{\mu}(x_i)+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}})^2\Big]. $$ It only relies on known or estimable quantities such as the derivative of the loss function $\psi(\cdot)$, the derivative of the inverse link function $\eta^{(1)}(\cdot)$, the residual $\widehat{\epsilon}_i$ and the binscatter estimates $\widehat{\mu}(\cdot)$ and $\widehat\boldsymbol{\gamma}$. Thus, $\widehat{\boldsymbol{\Sigma}}$ and other types of heteroskedasticity-robust “meat” matrix estimators can be easily constructed using the data. Then, it remains to obtain an estimator $\widehat{\mathbf{Q}}$ of $\bar\mathbf{Q}$ (or $\mathbf{Q}_0$), which in general relies on an estimator $\widehat{\Psi}_1(\cdot)$ of $\Psi_1(\cdot)$ and can be constructed in a case-by-case basis. In the following we discuss several examples.

Example 1 (Least Squares Regression). For least squares regression, the loss function $\rho(y; \eta)=\frac{1}{2}(y-\eta)^2$ and the (inverse) link function $\eta(\theta)=\theta$. Therefore, $\psi(y_i,\eta_i)=-\epsilon_i$ and $\eta_{i,1}=1$. Thus, the formula for $\widehat{\mathbf{Q}}$ given in Section (ref) reduces to $\mathbb{E}_n[\widehat\mathbf{b}_{p,s}(x_i)\widehat\mathbf{b}_{p,s}(x_i)']$, which is immediately feasible in practice.

Example 2 (Logistic Regression). For logistic regression, the loss function is given by the corresponding likelihood function, i.e., $-\rho(y;\eta)=y\log \eta+(1-y)\log (1-\eta)$, and the inverse link is given by the logistic function $\eta(\theta)=\frac{e^\theta}{1+e^\theta}$. Accordingly, an estimator of $\mathbf{Q}_0$ is given by \[ \widehat{\mathbf{Q}}=\mathbb{E}_n\Big[\widehat{\mathbf{b}}_{p,s}(x_i)\widehat{\mathbf{b}}_{p,s}(x_i)'\widehat{\eta}_i(1-\widehat{\eta}_i)\Big], \quad \widehat{\eta}_i=\eta(\widehat{\mu}(x_i)+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}}). \]

Example 4 (Quantile Regression). For quantile regression, $\rho(y;\eta)=(q-\mathds{1}(y<\eta))(y-\eta)$ for some $q\in(0,1)$ and $\eta(\theta)=\theta$. Accordingly, $\psi(y_i,\eta_i)=\mathds{1}(\epsilon_i<0)-q$, and one needs to estimate \[ \mathbf{Q}_0=\mathbb{E}\Big[\mathbf{b}_{p,s}(x_i)\mathbf{b}_{p,s}(x_i)'f_{Y|XW}(\mu_0(x_i)+\mathbf{w}_i'\boldsymbol{\gamma}_0|x_i,\mathbf{w}_i)\Big]. \] The key is to estimate the conditional density $f_{Y|XW}(\cdot|x_i,\mathbf{w}_i)$ evaluated at the conditional quantile of interest $(\mu_0(x_i)+\mathbf{w}_i'\boldsymbol{\gamma}_0)$, whose reciprocal is termed “sparsity function” in the literature. Many different methods have been proposed. For example, the sparsity function is simply the derivative of the conditional quantile function with respect to the quantile, which can be estimated by using the difference quotient of the estimated conditional quantile function. Alternatively, $\mathbf{Q}_0$ can be viewed as a matrix-weighted density function, and one can construct a corresponding estimator based on kernel density estimation ideas. In addition, one can use bootstrapping methods to estimate the variance, avoiding the technical difficulty of estimating the sparsity function. See Section 3.4 and Section 3.9 of Koenker_2005_book for more discussion of variance estimation for quantile regression.

Number of Bins Selector

We discuss the implementation details for data-driven selection of the number of bins, based on the approximate integrated mean squared error expansion in Theorem (ref).

We offer two procedures for estimating the bias and variance constants, and once these estimates ($\widehat{\mathscr{B}}_n(p,s,v)$ and $\widehat{\mathscr{V}}_n(p,s,v)$) are available, the estimated optimal $J$ is \[ \widehat{J}_{\mathtt{IMSE}}=\bigg\lceil \bigg(\frac{2(p-v+1)\widehat{\mathscr{B}}_n(p,s,v)} {(1+2v)\widehat{\mathscr{V}}_n(p,s,v)}\bigg)^{\frac{1}{2p+3}} n^{\frac{1}{2p+3}} \bigg\rceil. \] We always let $\omega(x)=f_X(x)$ as weighting function for concreteness.

Rule-of-thumb Selector

A rule-of-thumb choice of $J$ can be obtained based on Corollary SA-3.2 in Cattaneo-Crump-Farrell-Feng_2024_AER, which gives an explicit characterization of the variance and bias constants for least squares binscatter using piecewise polynomials ($s=0$).

Specifically, the variance constant $\mathscr{V}(p,0,v)$ is estimated by \[ \widehat{\mathscr{V}}(p,0,v)=\operatorname*{trace}\Big\{\Big(\int_{0}^{1}\bm{\varphi}(z)\bm{\varphi}(z)'dz\Big)^{-1}\int_0^1\bm{\varphi}^{(v)}(z)\bm{\varphi}^{(v)}(z)'dz\Big\}\times \frac{1}{n}\sum_{i=1}^{n}\widehat{\sigma}^2(x_i,\mathbf{w}_i)\widehat{f}_X(x_i)^{2v} \] where $\bm{\varphi}(z)=(1, z, \ldots, z^p)'$, $\widehat{\sigma}^2(x_i,\mathbf{w}_i)$ is some estimate of the conditional variance $\mathbb{V}[y_i|x_i,\mathbf{w}_i]$ and $\widehat{f}_X(x_i)$ is some estimate of the density $f_X(x_i)$. On the other hand, the bias constant $\mathscr{B}(p,0,v)$ is estimated by \[ \widehat{\mathscr{B}}(p,0,v)=\frac{\int_{0}^{1}[\mathscr{B}_{p+1-v}(z)]^2 dz}{((p+1-v)!)^2}\times \frac{1}{n}\sum_{i=1}^{n}\frac{[\widehat{\mu}^{(p+1)}(x_i)]^2}{\widehat{f}_X(x_i)^{2p+2-2v}}. \] where $\mathscr{B}_p(z)=(-1)^p\sum_{k=0}^p\binom{p}{k}\binom{p+k}{k}(-z)^k/\binom{2p}{p}$ for each $p\in\mathbb{Z}_+$ and $\widehat\mu^{(p+1)}(x_i)$ is some preliminary estimate of $\mu_0^{(p+1)}(x_i)$. The details about getting the estimates $\widehat\sigma^2(x_i,\mathbf{w}_i)$, $\widehat{f}_X(x_i)$ and $\widehat\mu^{(p+1)}(x_i)$ can be found in Section SA-4.1 in Cattaneo-Crump-Farrell-Feng_2024_AER.

This procedure still yields a choice of $J$ with the correct rate, though the constant approximations are inconsistent for general loss.

Direct-plug-in Selector

The direct-plug-in selector is implemented based on nonlinear binscatter estimators, which applies to any user-specified $p$, $s$ and $v$. It requires a preliminary choice of $J$, for which the rule-of-thumb selector previously described can be used.

More generally, suppose that a preliminary choice $J_{\mathtt{pre}}$ is given, and then a binscatter basis $\widehat{\mathbf{b}}_{p,s}(x)$ (of order $p$) can be constructed immediately on the preliminary partition. Implementing a nonlinear binscatter estimation using this basis and partitioning, we can obtain the variance constant estimate using the variance matrix estimators discussed in Section (ref).

Regarding the bias constant, the key unknown in the expression of the leading approximation error $r_{0,v}^\star(x)$ in Theorem (ref) is $\mu_0^{(p+1)}(x)$, which can be estimated by implementing a nonlinear binscatter estimation of order $p+1$ (with the preliminary partition unchanged). Also, an estimate of $f_X(x_i)^{-1}$ in $r_{0,v}^\star(x)$ is $J\hat{h}_{x_i}$ where $\hat{h}_{x_i}$ denotes the length of the interval in $\widehat\Delta$ containing $x_i$. All other quantities in the expression of $\mathscr{B}(p,s,v)$ can be replaced by their sample analogues. Then, a bias constant estimate is available.

By this construction, the direct-plug-in selector employs the correct rate and consistent constant approximations for any nonlinear binscatter with any choice of $p$, $s$ and $v$.

Fixed $J$ and choice of polynomial order

Our main theory treats $J$ as diverging with the sample size. This reflects the standard approach wherein a researcher selects $p$ and $s$ in advance (often $s=p=0$ or $s=p=3$) and then selects $J$ given the data. The partition must get finer to remove the nonparametric smoothing bias in estimating the function $\mu_0(x)$ (and along with it, $\vartheta_0(x,\mathsf{w})$ or $\zeta_0(x,\mathsf{w})$). Correct recovery (either for estimation or visualization) of these functions is the primary use of binscatter. However, researchers sometimes prefer to pre-specify a fixed $J = {\tt J}$, and we also discuss implementation and interpretation of binscatter in this case.

Instead of modeling $J$ as diverging and searching for the optimal choice, a researcher may desire a fixed (often small and round) number of $J$, which we denote by ${\tt J}$. This is done either to make the estimate more visually appealing or because the results can be directly interpreted. In this case, instead of recovering the functions $\mu_0^{(v)}(x)$, $\vartheta_0(x,\mathsf{w})$, and $\zeta_0(x,\mathsf{w})$, the binscatter is interpreted as estimating their coarsened versions: the distribution of $y_i$ conditional on $x_i$ lying in a (fixed) bin, rather than at a single point. For some ${\tt J}$, this remains interpretable and all our inference results apply to this case. For example, in our application we can take ${\tt J} = 10$ and study the distribution of uninsured rate for each decile of income. The confidence bands then become pointwise confidence intervals that are simultaneously valid. For example, this could be used to examine inequality in health care access by asking if median uninsured rates are statistically different between the top and bottom decile.

A fixed ${\tt J}$ is also interpretable, and applicable, if $x_i$ is discrete. Then each mass point is given its own bin and the results apply to the conditional distribution of $y_i$ given $x_i = x$. Again, our theoretical results apply directly to this case and one obtains simultaneous inference over the set of points. Cattaneo-Crump-Farrell-Feng_2024_AER give further discussion and examples.

As a practical compromise between the visual appeal and interpretation of a small, fixed ${\tt J}$ and the demand for consistent estimation, we propose a novel, albeit ad-hoc, adjustment to the estimator aimed at addressing the smoothing bias left by fixing ${\tt J}$ by adjusting the choice of polynomial order $p$. The standard approach fixes $p$ in advance and selects $J$ based on the data, but we can invert this and search for the value of $p$ for which the fixed ${\tt J}$ would be IMSE-optimal. That is, we solve for

align[align omitted — 161 chars of source]

where in principle the set $\mathcal{P}$ is all nonnegative integers, but in practice $\mathcal{P} = \{p_\mathtt{min},p_\mathtt{min}+1,\dots,p_\mathtt{max}-1,p_\mathtt{max}\}$, for some integers $0 \leq p_\mathtt{min} \leq p_\mathtt{max}$. The (in)flexibility of fixed $J = {\tt J}$ is offset by changing the polynomial approximation. This may have some practical appeal, but our theoretical results in the next section continue in the standard case of fixed $p$ and diverging $J$.

To implement the data-driven choice $p_{\mathtt{IMSE}}({\tt J},v)$, users needs to specify the desired (often small) number of bins $\mathtt{J}$, the derivative order $v$ of interest, and a (finite) set $\mathcal{P}$ of acceptable polynomial orders. The size of $\mathcal{P}$ is usually small since in practice $p=3$ or $4$ often suffices to yield a small IMSE-optimal number of bins. Then, for each value of $p$ in $\mathcal{P}$, we can implement the rule-of-thumb or direct plug-in procedure as described in Section (ref) to obtain $J_{\mathtt{IMSE}}(p,v)$. The “optimal” choice $p_{\mathtt{IMSE}}({\tt J},v)$ is the value of $p$ with the resulting $J_{\mathtt{IMSE}}(p,v)$ closest to $\mathtt{J}$.

Proofs

We begin with a subsection collecting some technical lemmas used in the proofs of our main results. We then collect all the proof of the results presented in this supplemental appendix, which are in several cases more general than those discussed in the main text. Some of our technical results may be of more broad independent interest in the nonlinear series estimation literature.

Technical Lemmas

We first give several simple facts about $\widehat\Delta$ in the following lemma, which are immediate from Assumption (ref)(ii).

lem[Quasi-Uniformity] Suppose that Assumption (ref)(ii) holds. Then, (i) $J^{-1}\lesssim\min_{1\leq j\leq J}h_j\leq \max_{1\leq j\leq J}h_j\lesssim J^{-1}$, (ii) $\max_{1\leq j\leq J}|\hat{\tau}_j-\tau_j|\lesssim_\mathbb{P} \mathfrak{r}_{\tt RP}$, and (iii) $\widehat\Delta\in\Pi_{3c_{\tt QU}}$ w.p.a. $1$.
proofBy Assumption (ref)(ii), $\mathrm{len}(\mathcal{X})=\sum_{j=1}^Jh_j\geq J\min_{1\leq j\leq J}h_j\geq c_{\tt QU}^{-1}J\max_{1\leq j\leq J} h_j$ where $\mathrm{len}(\mathcal{X})$ denotes the length of $\mathcal{X}$ (which is a fixed number). On the other hand, $\mathrm{len}(\mathcal{X})\leq J\max_{1\leq j\leq J}h_j\\ \leq c_{\tt QU}J\min_{1\leq j\leq J} h_j$. Therefore, $c_{\tt QU}^{-1}J^{-1}\mathrm{len}(\mathcal{X})\leq \min_{1\leq j\leq J}h_j \leq \max_{1\leq j\leq J}h_j\leq c_{\tt QU}J^{-1}\mathrm{len}(\mathcal{X})$. Next, by Assumption (ref)(ii), $\max_{1\leq j\leq J}|\hat{\tau}_j-\tau_j|=\max_{1\leq j\leq J}|\sum_{l=1}^j(\hat{h}_l-h_l)|\leq J\max_{1\leq l\leq J}|\hat{h}_l-h_l|\lesssim \mathfrak{r}_{\tt RP}$. In addition, $\max_{1\leq j\leq J}|\hat{h}_j-h_j|\leq \frac{1}{2}c_{\tt QU}^{-1}J^{-1}\mathrm{len}(\mathcal{X}) \leq \frac{1}{2}\min_{1\leq j\leq J}h_j$ w.p.a. $1$, and thus \[ \frac{\max_{1\leq j\leq J}\hat{h}_j}{\min_{1\leq j\leq J}\hat{h}_j}= \frac{\max_{1\leq j\leq J}h_j+\max_{1\leq j\leq J}|\hat{h}_j-h_j|}{\min_{1\leq j\leq J}h_j-\max_{1\leq j\leq J}|\hat{h}_j-h_j|}\leq 3c_{\tt QU}, \quad \text{w.p.a.} 1. \] Then, the proof is complete.

The next lemma then verifies Assumption (ref)(ii) for the special case of quantile-spaced partitions. The proof is available in the supplemental appendix of Cattaneo-Crump-Farrell-Feng_2024_AER (see Section SA-3.1 therein) and thus omitted here.

lem[Quasi-Uniformity of Quantile-Spaced Partitions] Suppose that Assumption (ref)(i) and (ref)(ii) holds and $\widehat\Delta$ is generated by sample quantiles, i.e., $\hat{\tau}_j=\hat{F}_X^{-1}(j/J)$. If $\frac{J\log J}{n}=o(1)$ and $\frac{\log n}{J}=o(1)$, then Assumption (ref)(ii) holds with $\tau_j=F_X^{-1}(j/J)$ and $\mathfrak{r}_{\tt RP}=\Big(\frac{J\log J}{n}\Big)^{1/2}$.

The next three lemmas (ref)--(ref) concern the properties of binscatter basis functions. Their proofs are the same as those for quantile-based partitions that are available in the supplemental appendix of Cattaneo-Crump-Farrell-Feng_2024_AER (see Section SA-3.1 therein) and are omitted here to conserve space.

lem[Transformation Matrix] Suppose that Assumption (ref)(i) holds. Then $\widehat{\mathbf{b}}_{p,s}(x)=\widehat{\mathbf{T}}_s\widehat{\mathbf{b}}_{p,0}(x)$ with $\|\widehat{\mathbf{T}}_s\|_\infty\lesssim_\mathbb{P} 1$ and $\|\widehat{\mathbf{T}}_s\|\lesssim_\mathbb{P} 1$. If, in addition, Assumption (ref)(ii) holds, then $\|\widehat{\mathbf{T}}_s-\mathbf{T}_s\|_\infty\lesssim_\mathbb{P}\mathfrak{r}_{\tt RP}$ and $\|\widehat{\mathbf{T}}_s-\mathbf{T}_s\|\lesssim_\mathbb{P}\mathfrak{r}_{\tt RP}$.
lem[Local Basis] Suppose that Assumption (ref)(i) holds. Then $\sup_{x\in\mathcal{X}}\|\widehat{\mathbf{b}}_{p,s}^{(v)}(x)\|_0\leq (p+1)^2$ and $\sup_{x\in\mathcal{X}}\|\widehat{\mathbf{b}}_{p,s}^{(v)}(x)\|\lesssim_\mathbb{P} J^{\frac{1}{2}+v}$.

The following lemma provides a particular way to define $\boldsymbol{\beta}_0(\Delta)$ and $\widehat\boldsymbol{\beta}_0$ so that the required approximation rate is achieved. We define $$ \boldsymbol{\beta}_0^{\mathtt{LS}}(\Delta):=\operatorname*{arg\,min}_{\boldsymbol{\beta}\in\mathbb{R}^{K_{p,s}}}\; \mathbb{E}[(\mu_0(x_i)-\mathbf{b}_{p,s}(x_i;\Delta)'\boldsymbol{\beta})^2], \quad \widehat\boldsymbol{\beta}_0^{\mathtt{LS}}=\boldsymbol{\beta}_0^{\mathtt{LS}}(\widehat\Delta). $$

lem[Approximation Error] Suppose that Assumptions (ref)(i)(ii), (ref)(v) and (ref)(i) hold. Then \begin{align*} \sup_{\Delta\in\Pi}\sup_{x\in\mathcal{X}}|\mathbf{b}_{p,s}^{(v)}(x;\Delta)'\boldsymbol{\beta}_0^{\mathtt{LS}}(\Delta)-\mu_0^{(v)}(x)| \lesssim J^{-p-1+v}\quad and\quad \sup_{x\in\mathcal{X}}|\widehat{\mathbf{b}}_{p,s}^{(v)}(x)'\widehat{\boldsymbol{\beta}}_0^{\mathtt{LS}}-\mu_0^{(v)}(x)|\lesssim_\mathbb{P} J^{-p-1+v}. \end{align*}

Next, the following maximal inequality is useful in our analysis. Its proof is available in Cattaneo-Feng-Underwood_2024_jasa and thus omitted here.

lem[Maximal Inequality] Let $Z_1, \cdots, Z_n$ be independent but not necessarily identically distributed random variables taking values in a measurable space $(\mathcal{S}; \mathscr{S})$. Denote the joint distribution of $Z_1, \cdots, Z_n$ by $\mathbb{P}$ and the marginal distribution of $Z_i$ by $\mathbb{P}_i$, and let $\bar{\mathbb{P}}=\frac{1}{n}\sum_{i=1}^n\mathbb{P}_i$. Let $\mathcal{F}$ be a class of Borel measurable functions from $\mathcal{S}$ to $\mathbb{R}$ which is pointwise measurable. Let $\bar{F}$ be a measurable envelope function for $\mathcal{F}$. Suppose that $\|\bar{F}\|_{L_2(\bar{\mathbb{P}})}<\infty$. Let $\bar{\sigma}>0$ satisfy $\sup_{f\in\mathcal{F}}\|f\|_{L_2(\bar{\mathbb{P}})}\leq\bar{\sigma}\leq \|\bar{F}\|_{L_2(\bar{\mathbb{P}})}$ and define $\bar{\bar{F}}=\max_{1\leq i\leq n}\bar{F}(Z_i)$. Then, with $\delta=\bar{\sigma}/\|\bar{F}\|_{L_2(\bar{\mathbb{P}})}$, \[ \mathbb{E}\Big[\sup_{f\in\mathcal{F}}\Big|\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\Big(f(Z_i)-\mathbb{E}[f(Z_i)]\Big)\Big|\Big]\lesssim\|\bar{F}\|_{L_2(\bar{\mathbb{P}})}J(\delta, \mathcal{F}, \bar{F})+\frac{\|\bar{\bar{F}}\|_{L_2(\mathbb{P})}J(\delta, \mathcal{F}, \bar{F})^2}{\delta^2\sqrt{n}}, \] where \[ J(\delta, \mathcal{F}, \bar{F})=\int_0^{\delta}\sqrt{1+\sup_{\mathbb{Q}}\log N(\mathcal{F}, L_2(\mathbb{Q}), \varepsilon\|\bar{F}\|_{L_2(\mathbb{Q})})}d\varepsilon. \]

Proof of Lemma (ref)

proofWe write $\Psi_{i,1}:=\Psi_1(x_i,\mathbf{w}_i;\eta_i)$. (i) We first prove a convergence result for $\bar{\mathbf{Q}}$. In view of Lemma (ref), it suffices to show the convergence for $s=0$. Let $\mathcal{A}_n$ denote the event on which $\widehat{\Delta}\in\Pi$. By Assumption (ref)(i), $\mathbb{P}(\mathcal{A}_n^c)=o(1)$. On $\mathcal{A}_n$, \[ \begin{split} &\;\Big\|\mathbb{E}_n[\widehat{\mathbf{b}}_{p,0}(x_i)\widehat{\mathbf{b}}_{p,0}(x_i)'\Psi_{i,1}\eta_{i,1}^2]- \mathbb{E}_{\widehat{\Delta}}[\widehat{\mathbf{b}}_{p,0}(x_i)\widehat{\mathbf{b}}_{p,0}(x_i)'\Psi_{i,1}\eta_{i,1}^2]\Big\|\\ \leq & \sup_{\Delta\in\Pi}\|\mathbb{E}_n[\mathbf{b}_{p,0}(x_i;\Delta)\mathbf{b}_{p,0}(x_i;\Delta)'\Psi_{i,1}\eta_i^2]- \mathbb{E}[\mathbf{b}_{p,0}(x_i;\Delta)\mathbf{b}_{p,0}(x_i;\Delta)'\Psi_{i,1}\eta_i^2]\|_\infty. \end{split} \] Let $a_{kl}$ be a generic $(k,l)$th entry of the matrix inside the norm, i.e., \[ |a_{kl}|=\Big|\mathbb{E}_n[b_{p,0,k}(x_{i}; \Delta)b_{p,0,l}(x_{i};\Delta)'\Psi_{i,1}\eta_{i,1}^2]- \mathbb{E}[b_{p,0,k}(x_{i};\Delta)b_{p,0,l}(x_{i};\Delta)'\Psi_{i,1}\eta_{i,1}^2]\Big|. \] Clearly, if $b_{p,0,k}(\cdot \,;\Delta)$ and $b_{p,0,l}(\cdot \,;\Delta)$ are basis functions with different supports, $a_{kl}$ is zero. Now define the following function class \[\mathcal{G}=\Big\{(x_1, \mathbf{w}_1)\mapsto b_{p,0,k}(x_1;\Delta)b_{p,0,l}(x_1;\Delta)\Psi_i\eta_{i,1}^2: 1\leq k, l\leq J(p+1), \Delta\in\Pi\Big\}. \] We have $ \sup_{g\in\mathcal{G}}|g|_\infty\lesssim J$ and $ \sup_{g\in\mathcal{G}}\mathbb{V}[g]\leq \sup_{g\in\mathcal{G}}\mathbb{E}[g^2]\lesssim J, $ by Assumption (ref). Also, by Proposition 3.6.12 of Gine-Nickl_2016_book, the collection $\mathcal{G}$ is of VC type with a bounded index. Then, by Lemma (ref), $$\sup_{g\in\mathcal{G}}\Big|\frac{1}{n}\sum_{i=1}^{n}g(x_i)-\mathbb{E}[g(x_i)]\Big| \lesssim_\mathbb{P} \sqrt{J\log J/n},$$ which implies $\|\mathbb{E}_n[\widehat{\mathbf{b}}_{p,0}(x_i)\widehat{\mathbf{b}}_{p,0}(x_i)'\Psi_{i,1}\eta_{i,1}^2]- \mathbb{E}_{\widehat{\Delta}}[\widehat{\mathbf{b}}_{p,0}(x_i)\widehat{\mathbf{b}}_{p,0}(x_i)'\Psi_{i,1}\eta_{i,1}^2]\| \lesssim_\mathbb{P} \sqrt{J\log J/n}$. Then, the lower bound on the minimum eigenvalue of $\bar\mathbf{Q}$ follows by Theorem 4.42 of Schumaker_2007_book and Assumption (ref)(i). The upper bound immediately follows by Assumption (ref)(i) and Lemmas (ref) and (ref). Given the above fact, it follows that $\|\bar{\mathbf{Q}}^{-1}\|\lesssim_\mathbb{P} 1$. Notice that $\bar{\mathbf{Q}}$ is a banded matrix with a finite band width. Then, the bounds on the elements of $\bar{\mathbf{Q}}^{-1}$ and $\|\bar{\mathbf{Q}}^{-1}\|_\infty$ hold by Theorem 2.2 of Demko_1977_SIAM. (ii) By Assumption (ref) and (ref), $\Psi_{i,1}\eta_{i,1}^2$ is bounded and bounded away from zero uniformly over $1\leq i\leq n$. Then, $\mathbb{E}[\mathbf{b}_{p,s}(x_i)\mathbf{b}_{p,s}(x_i)']\lesssim \mathbf{Q}_0\lesssim\mathbb{E}[\mathbf{b}_{p,s}(x_i)\mathbf{b}_{p,s}(x_i)']$. The desired bounds on the minimum and maximum eigenvalues of $\mathbf{Q}_0$ follow from Lemma SA-3.5 of Cattaneo-Crump-Farrell-Feng_2024_AER. Next, we show the convergence of $\bar\mathbf{Q}$ to $\mathbf{Q}_0$. Let $\alpha_{kl}$ be a generic $(k,l)$th entry of $$ \mathbb{E}_{\widehat{\Delta}}[\widehat{\mathbf{b}}_{p,0}(x_{i})\widehat{\mathbf{b}}_{p,0}(x_{i})'\Psi_{i,1}\eta_{i,1}^2]/J - \mathbb{E}[\mathbf{b}_{p,0}(x_{i})\mathbf{b}_{p,0}(x_{i})'\Psi_{i,1}\eta_{i,1}^2]/J. $$ By definition, it is either equal to zero or \begin{align*} \alpha_{kl} =&\int_{\widehat{\mathcal{B}}_j}\Big(\frac{x-\hat{\tau}_j}{\hat{h}_j}\Big)^\ell \varphi(x_i) f_X(x)dx -\int_{\mathcal{B}_j}\Big(\frac{x-\tau_j}{h_j}\Big)^\ell \varphi(x_i) f_X(x)dx \\ =&\hat{h}_j\int_0^1z^\ell \varphi(z\hat{h}_j+\hat{\tau}_j) f_X(z\hat{h}_j+\hat{\tau}_j)dz -h_j\int_0^1 z^\ell \varphi(zh_j+\tau_j) f_X(zh_j+\tau_j)dz \nonumber\\ =&(\hat{h}_j-h_j)\int_0^1z^\ell \varphi(z\hat{h}_j+\hat{\tau}_j) f_X(z\hat{h}_j+\hat{\tau}_j)dz\\ &+h_j\int_0^1z^\ell\Big(\varphi(z\hat{h}_j+\hat{\tau}_j)f_X(z\hat{h}_j+\hat{\tau}_j)- \varphi(zh_j+\tau_j) f_X(zh_j+\tau_j)\Big)dz \end{align*} for some $1\leq j\leq J$ and $0\leq\ell\leq 2p$ and $\varphi(x_i)=\mathbb{E}[\varkappa(x_i,\mathbf{w}_i)|x_i]$. By Assumptions (ref) and (ref) and the argument in the proof of Lemma SA-3.5 of Cattaneo-Crump-Farrell-Feng_2024_AER, $$\|\mathbb{E}_{\widehat{\Delta}}[\widehat{\mathbf{b}}_{p,0}(x_i)\widehat{\mathbf{b}}_{p,0}(x_i)'\Psi_{i,1}\eta_{i,1}^2]-\mathbf{Q}_0\|\lesssim_\mathbb{P} \mathfrak{r}_{\tt RP}.$$ Since $\bar{\mathbf{Q}}$ and $\mathbf{Q}_0$ are banded matrices with finite band widths. Then, the bound $\|\bar{\mathbf{Q}}^{-1}-\mathbf{Q}_0^{-1}\|_\infty$ hold by Theorem 2.2 of Demko_1977_SIAM. This completes the proof.

Proof of Lemma (ref)

proofSince $\mathbb{E}[\psi(y_i, \eta_i)^2|x_i=x, \mathbf{w}_i=\mathbf{w}]$ and $(\eta^{(1)}(\mu_0(x)+\mathbf{w}'\boldsymbol{\gamma}_0))^2$ is bounded and bounded away from zero uniformly over $x\in\mathcal{X}$ and $\mathbf{w}\in\mathcal{W}$, $\mathbb{E}_n[\widehat{\mathbf{b}}_{p,s}(x_i)\widehat{\mathbf{b}}_{p,s}(x_i)']\lesssim \bar{\boldsymbol{\Sigma}}\lesssim\mathbb{E}_n[\widehat{\mathbf{b}}_{p,s}(x_i)\widehat{\mathbf{b}}_{p,s}(x_i)']$. By the same argument in the proof of Lemma (ref) (we can simply drop the additional term $\Psi_{i,1}\eta_{i,1}^2$ in $\bar\mathbf{Q}$), the eigenvalues of $\mathbb{E}_n[\widehat\mathbf{b}_{p,s}(x_i)\widehat\mathbf{b}_{p,s}(x_i)']$ and thus $\bar\boldsymbol{\Sigma}$ are bounded and bounded away from zero. Then, the desired results follow from Lemma (ref) and the fact that $\inf_{x\in\mathcal{X}}\|\widehat\mathbf{b}_{p,s}^{(v)}(x)\|\gtrsim J^{1/2+v}$ w.p.a. $1$ (it was shown in the proof of Lemma SA-3.6 of Cattaneo-Crump-Farrell-Feng_2024_AER).

Proof of Lemma (ref)

proofBy Lemmas (ref), (ref) and (ref), $\sup_{x\in\mathcal{X}}\|\widehat{\mathbf{b}}_{p,s}^{(v)}(x)\|_1\lesssim_\mathbb{P} J^{1/2+v}$, $\|\bar{\mathbf{Q}}^{-1}\|_\infty\lesssim_\mathbb{P} 1$ and $\|\widehat{\mathbf{T}}_s\|_\infty\lesssim_\mathbb{P} 1$. Recall that by Assumption (ref), $\psi(y_i,\eta_i)=\psi^\dagger(y_i-\eta_i)\psi^\ddagger(\eta_i)=\psi^\dagger(\epsilon_i)\psi^\ddagger(\eta_i)$. Define the following function class \[ \mathcal{G}=\Big\{(x_1, \mathbf{w}_1, \epsilon_1)\mapsto b_{p,0,l}(x_1;\Delta)\eta^{(1)} (\mu_0(x_1)+\mathbf{w}_1'\boldsymbol{\gamma}_0)\psi^\dagger(\epsilon_1)\psi^\ddagger(\eta_1):1\leq l\leq J(p+1), \Delta\in\Pi \Big\}. \] Then, $\sup_{g\in\mathcal{G}}|g|\lesssim \sqrt{J}|\psi^\dagger(\epsilon_1)|$, and hence take an envelop $\bar{G}=C\sqrt{J}|\psi^\dagger(\epsilon_1)|$ for some $C$ large enough. Moreover, $\sup_{g\in\mathcal{G}}\mathbb{V}[g]\lesssim 1$ and $\mathcal{G}$ is of VC type with a bounded index. By Proposition 6.1 of Belloni-Chernozhukov-Chetverikov-Kato_2015_JoE, \[ \sup_{g\in\mathcal{G}}\Big|\frac{1}{n}\sum_{i=1}^{n}g(x_i, \epsilon_i)\Big|\lesssim_\mathbb{P} \sqrt{\frac{\log J}{n}}+\frac{J^{\frac{\nu}{2(\nu-2)}}\log J}{n}\lesssim\sqrt{\frac{\log J}{n}}, \] and the desired result follows.

Proof of Lemma (ref)

proofLet $\hat{z}_i=\widehat{\mathbf{b}}_{p,s}(x_i)'\widehat{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0$ and $\mathfrak{r}(x_i,\mathbf{w}_i, y_i):=\mathfrak{r}(x_i, \mathbf{w}_i, y_i;\widehat{\Delta}):= \eta_{i,1}\psi(y_i,\eta_i)- \eta^{(1)}(\hat{z}_i) \psi(y_i, \eta(\hat{z}_i))\\ =A_1(x_i, \mathbf{w}_i, y_i)+A_2(x_i, \mathbf{w}_i, y_i)$ where \begin{align*} &A_1(x_i, \mathbf{w}_i, y_i):=A_1(x_i,\mathbf{w}_i, y_i;\widehat{\Delta}):= [\eta_{i,1}\psi^\ddagger(\eta_i)-\eta^{(1)}(\hat{z}_i)\psi^\ddagger(\eta(\hat{z}_i))] \psi^\dagger(y_i,\eta_i) \;\;and\\ &A_2(x_i, \mathbf{w}_i, y_i):=A_2(x_i,\mathbf{w}_i, y_i;\widehat{\Delta}):= \eta^{(1)}(\hat{z}_i)\psi^\ddagger(\eta(\hat{z}_i))[\psi^\dagger(y_i, \eta_i)-\psi^\dagger(y_i, \eta(\hat{z}_i))]. \end{align*} First, by Assumption (ref) and Lemma (ref), $\sup_{x\in\mathcal{X}, \mathbf{w}\in\mathcal{W}}|\eta_{i,1}\psi^\ddagger(\eta_i)-\eta^{(1)}(\hat{z}_i)\psi^\ddagger(\eta(\hat{z}_i))|\lesssim J^{-p-1}$ w.p.a. 1. Also, for every $1\leq l\leq K_{p,s}$ and $\Delta\in\Pi$, \[\begin{split} &b_{p,s,l}(x;\Delta)\Big(\eta_{i,1}\psi^\ddagger(\eta_i)- \eta^{(1)}(\mathbf{b}_{p,s}(x;\Delta)'\boldsymbol{\beta}_0(\Delta)+\mathbf{w}'\boldsymbol{\gamma}_0)\psi^\ddagger(\mathbf{b}_{p,s}(x;\Delta)'\boldsymbol{\beta}_0(\Delta)+\mathbf{w}'\boldsymbol{\gamma}_0)\Big)\\ =&\;b_{p,s,l}(x;\Delta)\eta_{i,1}\psi^\ddagger(\eta_i)-\\ &\;b_{p,s,l}(x;\Delta)\eta^{(1)}\bigg(\sum_{k=\underline{k}_l}^{\underline{k}_l+p}b_{p,s,k}(x;\Delta)\beta_{0,k}(\Delta)+\mathbf{w}'\boldsymbol{\gamma}_0\bigg)\psi^\ddagger\bigg(\sum_{k=\underline{k}_l}^{\underline{k}_l+p}b_{p,s,k}(x;\Delta)\beta_{0,k}(\Delta)+\mathbf{w}'\boldsymbol{\gamma}_0\bigg) \end{split}\] for some integer $\underline{k}_l\in[1, K_{p,s}]$ where $\beta_{0,k}(\Delta)$ denotes the $k$th element in $\boldsymbol{\beta}_0(\Delta)$. Then, the function class $\mathcal{G}=\{(x,\mathbf{w}, y)\mapsto b_{p,s,l}(x;\Delta)A_1(x,\mathbf{w}, y;\Delta): 1\leq l\leq K_{p,s}, \Delta\in\Pi\}$ is of VC type with a bounded index. By the same argument given in the proof of Lemma (ref), \[ \|\mathbb{E}_n[\widehat{\mathbf{b}}_{p,s}(x_i)A_1(x_i,\mathbf{w}_i, y_i)]\|_\infty\lesssim_\mathbb{P} J^{-p-1} \Big(\frac{\log J}{n}\Big)^{1/2}. \] Next, let $\mathscr{F}_{XW\Delta}$ be the $\sigma$-field generated by $\{(x_i, \mathbf{w}_i)\}_{i=1}^n$ and $\widehat\Delta$. Note that \[ \begin{split} \mathbb{E}_n[\widehat{\mathbf{b}}_{p,s}(x_i)A_2(x_i, \mathbf{w}_i, y_i)] =\;&\mathbb{E}_n[\mathbb{E}[\widehat{\mathbf{b}}_{p,s}(x_i)A_2(x_i, \mathbf{w}_i , y_i)|\mathscr{F}_{XW\Delta}]]+\\ &\mathbb{E}_n\Big[\widehat{\mathbf{b}}_{p,s}(x_i)A_2(x_i, \mathbf{w}_i, y_i)-\mathbb{E}[\widehat{\mathbf{b}}_{p,s}(x_i)A_2(x_i, \mathbf{w}_i, y_i)|\mathscr{F}_{XW\Delta}]\Big]. \end{split} \] By Assumption (ref)(iii) and Lemma (ref), \begin{align*} &\max_{1\leq i\leq n} |\mathbb{E}[A_2(x_i, \mathbf{w}_i, y_i)|\mathscr{F}_{XW\Delta}]|\\ =&\max_{1\leq i\leq n}|\eta^{(1)}(\widehat{\mathbf{b}}_{p,s}(x_i)'\widehat{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0) \Psi(x_i, \mathbf{w}_i;\eta(\hat{z}_i))| \lesssim_\mathbb{P} J^{-p-1}. \end{align*} Then, $\|\mathbb{E}_n[\mathbb{E}[\widehat{\mathbf{b}}_{p,s}(x_i)A_2(x_i, \mathbf{w}_i, y_i)|\mathscr{F}_{XW\Delta}]]\|_\infty \lesssim_\mathbb{P} J^{-p-1-1/2}$ by the same argument in the proof of Lemma (ref). On the other hand, define the following function class \[ \mathcal{G}:=\Big\{(x,\mathbf{w}, y)\mapsto b_{p,s,l}(x;\Delta)A_2(x, \mathbf{w}, y;\Delta):1\leq l\leq K_{p,s}, \Delta\in\Pi \Big\}. \] By Assumption (ref), $\sup_{g\in\mathcal{G}}\|g\|_\infty\lesssim J^{1/2}$, and $\sup_{g\in\mathcal{G}}\mathbb{V}[g(x_i,\mathbf{w}_i,y_i)]\lesssim J^{-p-1}$. By a similar argument given before, this function class is of VC type with a bounded index. Then, as in the proof of Lemma (ref), by Proposition 6.1 of Belloni-Chernozhukov-Chetverikov-FernandezVal_2019_JoE, \[ \sup_{g\in\mathcal{G}}\Big|\frac{1}{n}\sum_{i=1}^{n}(g(x_i, \mathbf{w}_i, y_i)-\mathbb{E}[g(x_i,\mathbf{w}_i,y_i)])\Big|\lesssim_\mathbb{P} J^{-\frac{p+1}{2}}\sqrt{\frac{\log J}{n}}+\frac{J^{1/2}\log J}{n}. \] Collecting these results, we conclude that \[ \widehat{\mathbf{b}}_{p,s}^{(v)}(x)'\bar{\mathbf{Q}}^{-1}\mathbb{E}[\widehat{\mathbf{b}}_{p,s}(x_i)\mathfrak{r}(x_i, \mathbf{w}_i, y_i)] \lesssim_\mathbb{P} J^{-p-1+v}+J^{\frac{2v-p-1}{2}}\Big(\frac{J\log J}{n}\Big)^{1/2}+\frac{J^{1+v}\log J}{n}. \] The proof is complete.

Proof of Lemma (ref)

proofBy convexity of $\rho(y;\eta(\cdot))$, we only need to consider $\boldsymbol{\beta}=\widehat{\boldsymbol{\beta}}_0+\varepsilon\bm{\alpha}/\sqrt{J}$ for any sufficiently small fixed $\varepsilon>0$ and $\bm{\alpha}\in\mathbb{R}^{K_{p,s}}$ such that $\|\bm{\alpha}\|=1$. For notational simplicity, let $\widehat{\mathbf{b}}_i:=\widehat{\mathbf{b}}_{p,s}(x_i)$. For this choice of $\boldsymbol{\beta}$ and $\boldsymbol{\gamma}\in\mathbb{R}^d$, \[ \begin{split} \delta_i(\boldsymbol{\beta}, \boldsymbol{\gamma})&=\rho(y_i;\eta(\widehat{\mathbf{b}}_i'\boldsymbol{\beta}+\mathbf{w}_i'\boldsymbol{\gamma}))- \rho(y_i;\eta(\widehat{\mathbf{b}}_i'\widehat{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}))\\ &=\int_0^{\varepsilon\widehat{\mathbf{b}}_i'\bm{\alpha}/\sqrt{J}}\psi\Big(y_i,\eta(\widehat{\mathbf{b}}_i'\widehat{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}+t)\Big)\eta^{(1)}(\widehat{\mathbf{b}}_i'\widehat{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}+t)dt. \end{split}\] Let $\mathscr{F}_{XW\Delta}$ be the $\sigma$-field generated by $\{(x_i,\mathbf{w}_i)\}_{i=1}^n$ and $\widehat\Delta$. We have \[ \mathbb{E}_n[\delta_i(\boldsymbol{\beta}, \widehat{\boldsymbol{\gamma}})] =\frac{1}{\sqrt{n}}\mathbb{G}_n[\delta_i(\boldsymbol{\beta}, \widehat{\boldsymbol{\gamma}})]+ \mathbb{E}_n\Big[\mathbb{E}[\delta_i(\boldsymbol{\beta}, \widehat{\boldsymbol{\gamma}})|\mathscr{F}_{XW\Delta}]\Big], \] where $\mathbb{G}_n[\cdot]$ denotes $\sqrt{n}(\mathbb{E}_n[\cdot]-\mathbb{E}[\cdot|\mathscr{F}_{XW\Delta}])$, and $\mathbb{E}[\delta_i(\boldsymbol{\beta}, \widehat{\boldsymbol{\gamma}})|\mathscr{F}_{XW\Delta}]:=\mathbb{E}[\delta_i(\boldsymbol{\beta}, \boldsymbol{\gamma})|\mathscr{F}_{XW\Delta}]|_{\boldsymbol{\gamma}=\widehat{\boldsymbol{\gamma}}}$, i.e., the conditional expectation with $\widehat{\boldsymbol{\gamma}}$ viewed as fixed. By Assumption (ref), \begin{align*} &\mathbb{E}[\delta_i(\boldsymbol{\beta}, \widehat{\boldsymbol{\gamma}})|\mathscr{F}_{XW\Delta}]= \int_0^{\varepsilon\widehat{\mathbf{b}}_i'\bm{\alpha}/\sqrt{J}}\Psi\Big(x_i,\mathbf{w}_i;\eta(\widehat{\mathbf{b}}_i'\widehat{\boldsymbol{\beta}}_0+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}}+t)\Big)\eta^{(1)}(\widehat{\mathbf{b}}_i'\widehat{\boldsymbol{\beta}}_0+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}}+t)dt\\ =\;&\int_0^{\varepsilon\widehat{\mathbf{b}}_i'\bm{\alpha}/\sqrt{J}} \Psi_1(x_i,\mathbf{w}_i;\xi_{i,t})(\eta(\widehat{\mathbf{b}}_i'\widehat{\boldsymbol{\beta}}_0+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}}+t)-\eta_{i})\eta^{(1)}(\widehat{\mathbf{b}}_i'\widehat{\boldsymbol{\beta}}_0+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}}+t)dt, \end{align*} where $\xi_{i,t}$ is between $\eta(\widehat{\mathbf{b}}_i'\widehat{\boldsymbol{\beta}}_0+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}}+t)$ and $\eta(\mu_0(x_i)+\mathbf{w}_i'\boldsymbol{\gamma}_0)$ and we use the fact that $\Psi(x,\mathbf{w}_i;\eta_i)=0$. By Lemma (ref), the fact that $\eta(\cdot)$ is strictly monotonic and $\widehat{\boldsymbol{\gamma}}-\boldsymbol{\gamma}_0=o_\mathbb{P}(\sqrt{J/n}+J^{-p-1})$ and the rate condition imposed, we have $\mathbb{E}_n[\mathbb{E}[\delta_i(\boldsymbol{\beta}, \widehat{\boldsymbol{\gamma}})|\mathscr{F}_{XW\Delta}]] \gtrsim_\mathbb{P}\varepsilon^2\bm{\alpha}'\mathbb{E}_n[\widehat{\mathbf{b}}_i\widehat{\mathbf{b}}_i']\bm{\alpha}/J \gtrsim_\mathbb{P} J^{-1}\varepsilon^2$. On the other hand, let $\mathcal{H}:=\{\boldsymbol{\gamma}:\|\boldsymbol{\gamma}-\boldsymbol{\gamma}_0\|\leq C\mathfrak{r}_\gamma\}$ and define the following function class \[ \mathcal{G}:=\Big\{(x_i,\mathbf{w}_i,y_i)\mapsto \delta_i(\boldsymbol{\beta}, \boldsymbol{\gamma}): \bm{\alpha}\in\mathcal{S}^{K_{p,s}}, \boldsymbol{\gamma}\in\mathcal{H}\Big\}. \] Note that \begin{align*} \delta_i(\boldsymbol{\beta}, \boldsymbol{\gamma})= &\int_0^{\varepsilon\widehat{\mathbf{b}}_i'\bm{\alpha}/\sqrt{J}}\Big(\psi(y_i,\eta(\widehat{\mathbf{b}}_i'\widehat{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}+t))-\psi(y_i,\eta_{i})\Big)\eta^{(1)}(\widehat{\mathbf{b}}_i'\widehat{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}+t)dt\;+\\ &\int_0^{\varepsilon\widehat{\mathbf{b}}_i'\bm{\alpha}/\sqrt{J}}\psi(y_i,\eta_i)\eta^{(1)}(\widehat{\mathbf{b}}_i'\widehat{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}+t)dt. \end{align*} By Assumption (ref), we have $\sup_{g\in\mathcal{G}}|g|\lesssim \varepsilon(1+|\psi(y_i,\eta_i)|)$, $\|\max_{1\leq i\leq n}|\psi(y_i,\eta_i)|\|_{L_2(\mathbb{P})}\lesssim n^{1/\nu}$, $\sup_{g\in\mathcal{G}}\mathbb{E}_n[\mathbb{E}[g^2|\mathscr{F}_{XW\Delta}]] \lesssim_\mathbb{P} J^{-1}\varepsilon^2$, and the VC-index of $\mathcal{G}$ is bounded by $C'K_{p,s}$ for an absolute constant $C'>0$. Therefore, by Lemma (ref) and the rate restriction, \[ \sup_{g\in\mathcal{G}} \Big|\frac{1}{\sqrt{n}}\mathbb{G}_n[\delta_i(\boldsymbol{\beta}, \boldsymbol{\gamma})]\Big| \lesssim_\mathbb{P} J^{-1}\Big(\frac{J^2\log J}{n}\Big)^{1/2 }\varepsilon+J^{-1}\frac{J^2\log J}{n^{1-\frac{1}{\nu}}}\varepsilon=o(\varepsilon/J). \] Thus, for any fixed (sufficiently small) $\varepsilon>0$, $\mathbb{E}_n[\delta_i(\boldsymbol{\beta}, \widehat{\boldsymbol{\gamma}})]>0$ when $n$ is sufficiently large. Thus, $\|\widehat{\boldsymbol{\beta}}-\widehat{\boldsymbol{\beta}}_0\|=o_\mathbb{P}(J^{-1/2})$, implying $\|\widehat{\boldsymbol{\beta}}-\widehat{\boldsymbol{\beta}}_0\|_\infty=o_\mathbb{P}(J^{-1/2})$ immediately.

Proof of Theorem (ref)

proofThe proof is long. We divide it into several steps. Step 0: We first prepare some notation and useful facts. To simplify the presentation, in this proof we drop the scaling factor $\sqrt{J}$ in the basis by defining $$\breve{\mathbf{b}}_i:=\widehat{\mathbf{b}}_{p,s}(x_i)/\sqrt{J}=(\widehat{b}_{p,s,1}(x_i), \cdots, \widehat{b}_{p,s,K_{p,s}}(x_i))'/\sqrt{J}\quad \text{and}\quad \breve{\boldsymbol{\beta}}_0=\sqrt{J}\widehat{\boldsymbol{\beta}}_0. $$ Throughout the proof, $C, c, C_1, c_1, C_2, c_2, \cdots$ denote (strictly positive) absolute constants, $\mathscr{F}_{XW\Delta}$ denotes the $\sigma$-field generated by $\{(x_i,\mathbf{w}_i)\}_{i=1}^n$ and $\widehat\Delta$, and $\operatorname*{supp}(g(\cdot))$ denotes the support of a generic function $g(\cdot)$. Moreover, define \[ \begin{split} &\mathcal{V}=\{(v_1, \cdots, v_{K_{p,s}})': \exists k\in\{1,\cdots, K_{p,s}\}, |v_{\ell}|\leq\varrho^{|k-\ell|}\varepsilon_n \text{ for }|\ell-k|\leq M_n \text{ and } v_\ell=0 \text{ otherwise} \},\\ &\mathcal{H}_l=\{\mathbf{v}\in\mathbb{R}^{K_{p,s}}:\|\mathbf{v}\|_\infty\leq r_{l,n}\}\;\; \text{for}\; l=1,2,\quad \text{and}\quad \mathcal{H}_3=\{\mathbf{v}\in\mathbb{R}^d: \|\mathbf{v}\|\leq r_{3,n}\}, \end{split} \] where $\varrho\in(0,1)$ is the constant given in Lemma (ref), $r_{1,n}=C_1[(J\log n/n)^{1/2}+J^{-p-1}]$, $r_{2,n}=\mathfrak{z}\mathfrak{r}_{2,n}$ for $\mathfrak{z}>0$, $\varepsilon_n=\mathfrak{z}'\mathfrak{r}_{2,n}$ for $\mathfrak{z}'>0$, $\mathfrak{r}_{2,n}=[(\frac{J\log n}{n})^{3/4}\log n+J^{-\frac{p+1}{2}}\sqrt{\frac{J}{n}}\log n+ J^{-2p-2}+\mathfrak{r}_\gamma]$, $r_{3,n}=C\mathfrak{r}_\gamma$, and $M_n=c_1\log n$. In the last step of the proof, we will consider $\mathfrak{z}=2^\ell$, $\ell=L, L+1, \cdots, \bar{L}$ where $\bar{L}$ is the smallest number such that $2^{\bar{L}}r_{2n}\geq c$ for some sufficiently small constant $c>0$, and $\varepsilon_n$ is a quantity that we can choose. By Assumption (ref), $\widehat{\boldsymbol{\gamma}}-\boldsymbol{\gamma}_0\in\mathcal{H}_3$ with probability approaching one for $C$ large enough, and by Lemma (ref), $\sqrt{J}\widehat{\boldsymbol{\beta}}-\breve{\boldsymbol{\beta}}_0\leq c$ with probability approaching one. For any $\boldsymbol{\beta}_1\in\mathcal{H}_1, \boldsymbol{\beta}_2\in\mathcal{H}_2$, $\boldsymbol{\upsilon}\in\mathcal{V}$ and $\boldsymbol{\gamma}:=\boldsymbol{\gamma}_0+\boldsymbol{\gamma}_1$ with $\boldsymbol{\gamma}_1\in\mathcal{H}_3$, define \[ \begin{split} \delta_i(\boldsymbol{\beta}_1, \boldsymbol{\beta}_2, \boldsymbol{\upsilon},\boldsymbol{\gamma}) =\;&\rho\Big(y_i; \eta(\breve{\mathbf{b}}_i'(\breve{\boldsymbol{\beta}}_0+\boldsymbol{\beta}_1+\boldsymbol{\beta}_2)+\mathbf{w}_i'\boldsymbol{\gamma})\Big)- \rho\Big(y_i;\eta(\breve{\mathbf{b}}_i'(\breve{\boldsymbol{\beta}}_0+\boldsymbol{\beta}_1+\boldsymbol{\beta}_2-\boldsymbol{\upsilon})+\mathbf{w}_i'\boldsymbol{\gamma})\Big)\\ &-\Big[\eta(\breve{\mathbf{b}}_i'(\breve{\boldsymbol{\beta}}_0+\boldsymbol{\beta}_1+\boldsymbol{\beta}_2)+\mathbf{w}_i'\boldsymbol{\gamma})- \eta(\breve{\mathbf{b}}_i'(\breve{\boldsymbol{\beta}}_0+\boldsymbol{\beta}_1+\boldsymbol{\beta}_2-\boldsymbol{\upsilon})+\mathbf{w}_i'\boldsymbol{\gamma})\Big]\\ &\hspace{21em}\times\psi(y_i,\eta(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0))\\ =\;&\int_{-\breve{\mathbf{b}}_i'\boldsymbol{\upsilon}}^{0} \Big[\psi\Big(y_i, \eta(\breve{\mathbf{b}}_i'(\breve{\boldsymbol{\beta}}_0+\boldsymbol{\beta}_1+\boldsymbol{\beta}_2)+\mathbf{w}_i'\boldsymbol{\gamma}+t)\Big) -\psi\Big(y_i,\eta(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0)\Big)\Big]\\ &\hspace{15em}\times\eta^{(1)}\Big(\breve{\mathbf{b}}_i'(\breve{\boldsymbol{\beta}}_0+\boldsymbol{\beta}_1+\boldsymbol{\beta}_2)+\mathbf{w}_i'\boldsymbol{\gamma}+t\Big)dt. \end{split} \] Note that $\delta_i(\boldsymbol{\beta}_1, \boldsymbol{\beta}_2, \boldsymbol{\upsilon}, \boldsymbol{\gamma})\neq 0$ only if $\breve{\mathbf{b}}_i'\boldsymbol{\upsilon}\neq 0$. For each $\boldsymbol{\upsilon}\in\mathcal{V}$, let $\mathcal{J}_{\boldsymbol{\upsilon}}=\{j: \upsilon_j\neq 0\}$. By construction, the cardinality of $\mathcal{J}_{\boldsymbol{\upsilon}}$ is bounded by $2M_n+1$. We have $\delta_i(\boldsymbol{\beta}_1, \boldsymbol{\beta}_2, \boldsymbol{\upsilon}, \boldsymbol{\gamma})\neq 0$ only if $\breve{b}_{j}(x_i)\neq 0$ for some $j\in\mathcal{J}_\upsilon$, which happens only when $x_i\in\operatorname*{supp}(\breve{b}_j(\cdot))$ for some $j\in\mathcal{J}_{\boldsymbol{\upsilon}}$. Let $\mathcal{I}_{\boldsymbol{\upsilon}}=\cup_{j\in\mathcal{J}_{\boldsymbol{\upsilon}}}\operatorname*{supp}(\breve{b}_j(\cdot))$. Since the basis functions are locally supported, $\mathcal{I}_{\boldsymbol{\upsilon}}$ includes at most $c_2M_n$ (connected) intervals for all $\boldsymbol{\upsilon}\in\mathcal{V}$. Moreover, at most $c_3M_n$ basis functions in $\breve{\mathbf{b}}(\cdot)$ have supports overlapping with $\mathcal{I}_{\boldsymbol{\upsilon}}$. Denote the set of indices for such basis functions by $\bar{\mathcal{J}}_{\boldsymbol{\upsilon}}$. Let $\breve{\beta}_{0,j}$, $\beta_{1,j}$ and $\beta_{2,j}$ be the $j$th entries of $\breve{\boldsymbol{\beta}}_0$, $\boldsymbol{\beta}_1$, and $\boldsymbol{\beta}_2$ respectively, and $\upsilon_j$ be the $j$th entry of $\boldsymbol{\upsilon}$. Based on the above observations, we have $\delta_i(\boldsymbol{\beta}_1, \boldsymbol{\beta}_2, \boldsymbol{\upsilon}, \boldsymbol{\gamma})\equiv\delta_i(\boldsymbol{\beta}_{1,\bar{\mathcal{J}}_{\boldsymbol{\upsilon}}}, \boldsymbol{\beta}_{2,\bar{\mathcal{J}}_{\boldsymbol{\upsilon}}}, \boldsymbol{\upsilon}, \boldsymbol{\gamma})$ where \[ \begin{split} \delta_i(\boldsymbol{\beta}_{1,\bar{\mathcal{J}}_{\boldsymbol{\upsilon}}}, \boldsymbol{\beta}_{2,\bar{\mathcal{J}}_{\boldsymbol{\upsilon}}}, \boldsymbol{\upsilon}, \boldsymbol{\gamma})&:= \int_{-\underset{j\in\mathcal{J}_{\boldsymbol{\upsilon}}}{\sum}\breve{b}_{i,j}\upsilon_j}^{0}\Big[ \psi\Big(y_i,\eta\Big(\sum_{l\in\bar{\mathcal{J}}_{\boldsymbol{\upsilon}}} \breve{b}_{i,l}(\breve{\beta}_{0,l}+\beta_{1,l}+\beta_{2,l})+\mathbf{w}_i'\boldsymbol{\gamma}+t\Big)\Big) \\ -\psi\Big(y_i,\,&\eta\Big(\sum_{l\in\bar{\mathcal{J}}_{\boldsymbol{\upsilon}}} \breve{b}_{i,l}\breve{\beta}_{0,l}+\mathbf{w}_i'\boldsymbol{\gamma}_0\Big)\Big)\Big] \times\eta^{(1)}\Big(\sum_{l\in\bar{\mathcal{J}}_{\boldsymbol{\upsilon}}}\breve{b}_{i,l}(\breve{\beta}_{0,l}+\beta_{1,l}+\beta_{2,l})+\mathbf{w}_i'\boldsymbol{\gamma}+t\Big)dt\mathds{1}_{i,\boldsymbol{\upsilon}}, \end{split} \] $\mathds{1}_{i,\boldsymbol{\upsilon}}=\mathds{1}(x_i\in\mathcal{I}_{\boldsymbol{\upsilon}})$, and $\boldsymbol{\beta}_{1,\bar{\mathcal{J}}_{\boldsymbol{\upsilon}}}$ and $\boldsymbol{\beta}_{2,\bar{\mathcal{J}}_{\boldsymbol{\upsilon}}}$ respectively denote the subvectors of $\boldsymbol{\beta}_1$ and $\boldsymbol{\beta}_2$ whose indices belong to $\bar{\mathcal{J}}_{\boldsymbol{\upsilon}}$. Accordingly, define the following function class \begin{align*} \mathcal{G}=\Big\{(x_i, \mathbf{w}_i, y_i)\mapsto \delta_i(\widetilde{\boldsymbol{\beta}}_1, \widetilde{\boldsymbol{\beta}}_2, \boldsymbol{\upsilon}, \boldsymbol{\gamma}): \boldsymbol{\upsilon}\in\mathcal{V}, \, &\widetilde{\boldsymbol{\beta}}_1\in\mathbb{R}^{c_3M_n}, \widetilde{\boldsymbol{\beta}}_2\in\mathbb{R}^{c_3M_n},\\ &\|\widetilde{\boldsymbol{\beta}}_1\|_\infty\leq r_{1,n}, \|\widetilde{\boldsymbol{\beta}}_2\|_\infty\leq r_{2,n}, \boldsymbol{\gamma}-\boldsymbol{\gamma}_0\in\mathcal{H}_3 \Big\}. \end{align*} Step 1: We bound $\sup_{g\in\mathcal{G}}|\mathbb{E}_n[g(x_i,\mathbf{w}_i,y_i)]-\mathbb{E}[g(x_i,\mathbf{w}_i,y_i)|\mathscr{F}_{XW\Delta}]|$ in this step. Let $a_i(t):=\eta(\sum_{l\in\bar{\mathcal{J}}_{\boldsymbol{\upsilon}}}\breve{b}_{i,l}'\breve{\beta}_{0,l}+\mathbf{w}_i'\boldsymbol{\gamma}_0+t)$. Define \[\begin{split} & \underline{a}_i=\min\Big\{a_i(0), a_i\Big(\sum_{l\in\bar{\mathcal{J}}_{\boldsymbol{\upsilon}}}\breve{b}_{i,l}(\beta_{1,l}+\beta_{2,l})+\mathbf{w}_i'\boldsymbol{\gamma}_1\Big), a_i\Big(\sum_{l\in\bar{\mathcal{J}}_{\boldsymbol{\upsilon}}}\breve{b}_{i,l}(\beta_{1,l}+\beta_{2,l})+\mathbf{w}_i'\boldsymbol{\gamma}_1+\sum_{j\in\mathcal{J}_{\boldsymbol{\upsilon}}}\breve{b}_{i,j}\upsilon_j\Big)\Big\} \text{ and} \\ &\bar{a}_i=\max\Big\{a_i(0), a_i\Big(\sum_{l\in\bar{\mathcal{J}}_{\boldsymbol{\upsilon}}}\breve{b}_{i,l}(\beta_{1,l}+\beta_{2,l})+\mathbf{w}_i'\boldsymbol{\gamma}_1\Big), a_i\Big(\sum_{l\in\bar{\mathcal{J}}_{\boldsymbol{\upsilon}}}\breve{b}_{i,l}(\beta_{1,l}+\beta_{2,l})+\mathbf{w}_i'\boldsymbol{\gamma}_1+\sum_{j\in\mathcal{J}_{\boldsymbol{\upsilon}}}\breve{b}_{i,j}\upsilon_j\Big)\Big\}. \end{split} \] Consider the following two cases. First, suppose that $(y_i-\bar{a}_i, y_i-\underline{a}_i)$ does not contain any discontinuity points. By Assumption (ref), for all $t$ in the interval of integration $[-\sum_{j\in\mathcal{J}_{\boldsymbol{\upsilon}}}\breve{b}_{i,j}\upsilon_j, 0]$ (or $[0, -\sum_{j\in\mathcal{J}_{\boldsymbol{\upsilon}}}\breve{b}_{i,j}\upsilon_j])$, $$\Big|\psi\Big(y_i, a_i\Big(\sum_{l\in\bar{\mathcal{J}}_{\boldsymbol{\upsilon}}}\breve{b}_{i,l}(\beta_{1,l}+\beta_{2,l})+\mathbf{w}_i'\boldsymbol{\gamma}+t\Big)\Big)-\psi(y_i, a_i(0))\Big|\lesssim r_{1, n}+r_{2,n}+\varepsilon_n+r_{3,n}.$$ Second, if $(y_i-\bar{a}_i, y_i-\underline{a}_i)$ contains at least one discontinuity point, say $\jmath$. For any $t$ in the interval of integration, by Assumption (ref), $$\Big|\psi\Big(y_i, a_i\Big(\sum_{l\in\bar{\mathcal{J}}_{\boldsymbol{\upsilon}}}\breve{b}_{i,l}(\beta_{1,l}+\beta_{2,l})+\mathbf{w}_i'\boldsymbol{\gamma}+t\Big)\Big)-\psi(y_i, a_i(0))\Big| \lesssim 1+r_{3,n}+(1+|\psi(y_i,\eta_i)|)(r_{1,n}+r_{2,n}+\varepsilon_n+r_{3,n})$$ for any $(x_i,\mathbf{w}_i,y_i)$, and in this case $y_i\in(\jmath+\underline{a}_i, \jmath+\bar{a}_i)$. By Assumption (ref), $$|\bar{a}_i-\underline{a}_i|\lesssim (r_{1,n}+r_{2,n}+r_{3,n}+\varepsilon_n)(|\eta_{i,1}|+r_{1,n}+r_{2,n}+r_{3,n}+\varepsilon_n).$$ By construction, for each $\boldsymbol{\upsilon}\in\mathcal{V}$, there exists some $k_{\boldsymbol{\upsilon}}$ such that $|\upsilon_\ell|\leq \varrho^{|\ell-k_{\boldsymbol{\upsilon}}|}\varepsilon_n$ for $|\ell-k_{\boldsymbol{\upsilon}}|\leq M_n$. Therefore, we can further write $\mathds{1}_{i,\boldsymbol{\upsilon}}=\sum_{j:\widehat{\mathcal{B}}_j\subset\mathcal{I}_{\boldsymbol{\upsilon}}}\mathds{1}_{i,\boldsymbol{\upsilon},j}$ where each $\mathds{1}_{i,\boldsymbol{\upsilon},j}$ is an indicator of the subinterval involved in $\mathcal{I}_{\boldsymbol{\upsilon}}$, and the above facts imply that for any $x_i\in\widehat{\mathcal{B}}_{l}$ for some $\widehat{\mathcal{B}}_{l}\subset\mathcal{I}_{\boldsymbol{\upsilon}}$, \[ \mathbb{V}[\delta_i(\boldsymbol{\beta}_1, \boldsymbol{\beta}_2, \boldsymbol{\upsilon},\boldsymbol{\gamma})|\mathscr{F}_{XW\Delta}]\lesssim\; \varrho^{2|(p-s+1)l-k_{\boldsymbol{\upsilon}}|}\varepsilon_n^2 (r_{1,n}+r_{2,n}+\varepsilon_n+r_{3,n})(|\eta_{i,1}|+r_{1,n}+r_{2,n}+\varepsilon_n+r_{3,n}). \] In addition, since $\delta_i(\boldsymbol{\beta}_1, \boldsymbol{\beta}_2, \boldsymbol{\upsilon}, \boldsymbol{\gamma})\neq 0$ only if $x_i\in\mathcal{I}_{\boldsymbol{\upsilon}}$, for all $g\in\mathcal{G}$ (each corresponds to a particular $\boldsymbol{\upsilon}$), \[ \mathbb{E}_n[\mathbb{V}[g(x_i,\mathbf{w}_i,y_i)|\mathscr{F}_{XW\Delta}]]\lesssim \varepsilon_n^2(r_{1,n}+r_{2,n}+\varepsilon_n+r_{3,n})\sum_{l:\widehat{\mathcal{B}}_l\subset\mathcal{I}_{\boldsymbol{\upsilon}}}\mathbb{E}_n[\mathds{1}_{i,\boldsymbol{\upsilon},l}]\varrho^{2|(p-s+1)l-k_{\boldsymbol{\upsilon}}|}. \] This inequality holds for any event in $\mathscr{F}_{XW\Delta}$. Define an event $\mathcal{A}_1$ on which $\sup_{1\leq j\leq J}\mathbb{E}_n[\mathds{1}_{i,j}]\leq C_2J^{-1}$ for some large enough $C_2>0$ where $\mathds{1}_{i,j}=\mathds{1}(x_i\in\widehat{\mathcal{B}}_j)$. By the argument in Lemma (ref), $\mathbb{P}(\mathcal{A}_1^c)\to 0$. On $\mathcal{A}_1$, $$ \bar{\sigma}^2:=\sup_{g\in\mathcal{G}}\;\mathbb{E}_n[\mathbb{V}[g(x_i,\mathbf{w}_i,y_i)|\mathscr{F}_{XW\Delta}]]\lesssim \varepsilon_n^2J^{-1}(r_{1,n}+r_{2,n}+\varepsilon_n+r_{3,n}).$$ On the other hand, $$\bar{G}:=\sup_{g\in\mathcal{G}}\;|g(x_i,\mathbf{w}_i, y_i)|\lesssim \varepsilon_n(1+r_{3,n}+|\psi(y_i,\eta_i)|(r_{1,n}+r_{2,n}+\varepsilon_n+r_{3,n}))(|\eta_{i,1}|+r_{1,n}+r_{2,n}+\varepsilon_n+r_{3,n}).$$ Also, for any $g,\tilde{g}\in\mathcal{G}$, denote the corresponding parameters defining $g$ and $\tilde{g}$ by $(\boldsymbol{\beta}_1, \boldsymbol{\beta}_2, \boldsymbol{\upsilon}, \boldsymbol{\gamma})$ and $(\tilde{\boldsymbol{\beta}}_1, \tilde{\boldsymbol{\beta}}_2, \tilde{\boldsymbol{\upsilon}}, \tilde{\boldsymbol{\gamma}})$. We have \begin{align*} \tilde{g}(x_i,\mathbf{w}_i,y_i)-g(x_i,\mathbf{w}_i,y_i)=& \int_{0}^{\varLambda_1}\Big[\psi(y_i,\eta(\breve{\mathbf{b}}_i'(\breve{\boldsymbol{\beta}}_0+\boldsymbol{\beta}_1+\boldsymbol{\beta}_2)+\mathbf{w}_i'\boldsymbol{\gamma}+t))\\ &-\psi(y_i,\eta(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0))\Big] \times\eta^{(1)}(\breve{\mathbf{b}}_i'(\breve{\boldsymbol{\beta}}_0+\boldsymbol{\beta}_1+\boldsymbol{\beta}_2)+\mathbf{w}_i'\boldsymbol{\gamma}+t)dt\\ &-\int_{0}^{\varLambda_2}\Big[\psi(y_i,\eta(\breve{\mathbf{b}}_i'(\breve{\boldsymbol{\beta}}_0+\boldsymbol{\beta}_1+\boldsymbol{\beta}_2-\boldsymbol{\upsilon})+\mathbf{w}_i'\boldsymbol{\gamma}+t))\\ &-\psi(y_i,\eta(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0))\Big] \times\eta^{(1)}(\breve{\mathbf{b}}_i'(\breve{\boldsymbol{\beta}}_0+\boldsymbol{\beta}_1+\boldsymbol{\beta}_2-\boldsymbol{\upsilon})+\mathbf{w}_i'\boldsymbol{\gamma}+t)dt\\ \lesssim&\; (1+\Lambda_1+\Lambda_2)(|\eta_{i,1}|+r_{1,n}+r_{2,n}+\varLambda_1+\varLambda_2+r_{3,n})\\ &\times(\|(\tilde{\boldsymbol{\beta}}_1-\boldsymbol{\beta}_1\|_\infty+\|\tilde{\boldsymbol{\beta}}_2-\boldsymbol{\beta}_2)\|_\infty+\|\tilde{\boldsymbol{\upsilon}}-\boldsymbol{\upsilon}\|_\infty+\|\tilde{\boldsymbol{\gamma}}-\boldsymbol{\gamma}\|), \end{align*} where $\varLambda_1=\breve{\mathbf{b}}_i'(\tilde{\boldsymbol{\beta}}_1+\tilde{\boldsymbol{\beta}}_2-\boldsymbol{\beta}_1-\boldsymbol{\beta}_2)+\mathbf{w}_i'(\tilde{\boldsymbol{\gamma}}-\boldsymbol{\gamma})$ and $\varLambda_2=\varLambda_1-\breve{\mathbf{b}}_i'(\tilde{\boldsymbol{\upsilon}}-\boldsymbol{\upsilon})$. Based on these observations, \[ \|\bar{G}\|_{\bar{\mathbb{P}},2}\int_0^{\frac{\bar\sigma}{\|\bar{G}\|_{\bar{\mathbb{P}},2}}}\sqrt{1+\sup_{\mathbb{Q}}\log N(\mathcal{G}, L_2(\mathbb{Q}), t\|\bar{G}\|_{\mathbb{Q},2})}dt\lesssim \bar{\sigma}\Big(\sqrt{\log J}+\sqrt{\log n\log \frac{1}{\bar{\sigma}}}\Big) \lesssim\bar{\sigma}\log n, \] where the supremum is taken over all finite discrete probability measures $\mathbb{Q}$. Then, by Lemma (ref), \[ \mathbb{E}\bigg[\sup_{g\in\mathcal{G}} \Big|\mathbb{G}_n[g(x_i, \mathbf{w}_i, y_i)]\Big|\bigg|\mathscr{F}_{XW\Delta}\bigg] \lesssim \bar{\sigma}\log n+\frac{\sqrt{\mathbb{E}[\bar{\bar{G}}^2]} \log^2 n}{\sqrt{n}}, \] where $\bar{\bar{G}}=\max_{1\leq i\leq n}\bar{G}(x_i,\mathbf{w}_i,y_i)$. Note that $(\mathbb{E}[\bar{\bar{G}}^2])^{1/2}\lesssim\varepsilon_n (1+n^{1/\nu}(r_{1,n}+r_{2,n}+r_{3,n}+\varepsilon_n))$. Therefore, on $\mathcal{A}_1$ (whose probability approaches one), \[ \begin{split} &\sup_{\boldsymbol{\beta}_1\in\mathcal{H}_1, \boldsymbol{\beta}_2\in\mathcal{H}_2, \boldsymbol{\upsilon}\in\mathcal{V}, \boldsymbol{\gamma}_1\in\mathcal{H}_3}\; \Big|\mathbb{E}_n\Big[\delta_i(\boldsymbol{\beta}_1, \boldsymbol{\beta}_2, \boldsymbol{\upsilon}, \boldsymbol{\gamma})\Big]- \mathbb{E}_n\Big[\mathbb{E}[\delta_i(\boldsymbol{\beta}_1, \boldsymbol{\beta}_2, \boldsymbol{\upsilon}, \boldsymbol{\gamma})|\mathscr{F}_{XW\Delta}]\Big]\Big|\\ \lesssim &\;\bigg(J^{-1}\varepsilon_n\sqrt{\mathfrak{L}_n} \sqrt{\frac{J}{n}}\log n+\frac{\varepsilon_n(1+n^{1/\nu}\mathfrak{L}_n)(\log n)^2}{n}\bigg) \end{split} \] for $\mathfrak{L}_n=r_{1,n}+r_{2,n}+r_{3,n}+\varepsilon_n$. Step 2: For $\widetilde{\mathbf{Q}}:= \mathbb{E}_n[\breve{\mathbf{b}}_i\breve{\mathbf{b}}_i' \Psi_1(x_i,\mathbf{w}_i;\eta(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0)) (\eta^{(1)}(\breve{\mathbf{b}}_i\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0))^2]$, by Assumption (ref) and the same argument in the proof of Lemma (ref), $\|\bar{\mathbf{Q}}-\widetilde{\mathbf{Q}}\|_\infty\vee\|\bar{\mathbf{Q}}-\widetilde{\mathbf{Q}}\| \lesssim J^{-p-1}J^{-1}$. Therefore, $$ \sup_{\boldsymbol{\beta}_1\in\mathcal{H}_1, \boldsymbol{\beta}_2\in\mathcal{H}_2, \boldsymbol{\upsilon}\in\mathcal{V}}|\boldsymbol{\upsilon}'(\widetilde{\mathbf{Q}}-\bar{\mathbf{Q}})(\boldsymbol{\beta}_1+\boldsymbol{\beta}_2)|\lesssim J^{-p-2}\varepsilon_n(r_{1,n}+r_{2,n}). $$ In addition, by Lemmas (ref) and (ref), $\|\bar{\boldsymbol{\beta}}\|_\infty\leq r_{1,n}$ with probability approaching one for $C_1$ large enough, where \[ \bar{\boldsymbol{\beta}}:=-\bar{\mathbf{Q}}^{-1}\mathbb{E}_n\Big[\breve{\mathbf{b}}_i\eta^{(1)}(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0)\psi\Big(y_i,\eta(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0)\Big)\Big]. \] Step 3: By Taylor expansion, we have \begin{align*} &\;\mathbb{E}_n\Big[\mathbb{E}[\delta_i(\boldsymbol{\beta}_1, \boldsymbol{\beta}_2, \boldsymbol{\upsilon}, \boldsymbol{\gamma})|\mathscr{F}_{XW\Delta}]\Big]\\ =&\;\mathbb{E}_n\bigg[\int_{-\breve{\mathbf{b}}_i'\boldsymbol{\upsilon}}^{0} \Big\{\Psi(x_i, \mathbf{w}_i; \eta(\breve{\mathbf{b}}_i'(\breve{\boldsymbol{\beta}}_0+\boldsymbol{\beta}_1+\boldsymbol{\beta}_2)+\mathbf{w}_i'\boldsymbol{\gamma}+t))\\ &-\Psi(x_i, \mathbf{w}_i;\eta(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0))\Big\} \times\eta^{(1)}\Big(\breve{\mathbf{b}}_i'(\breve{\boldsymbol{\beta}}_0+\boldsymbol{\beta}_1+\boldsymbol{\beta}_2)+\mathbf{w}_i'\boldsymbol{\gamma}+t\Big)dt\bigg]\\ =&\;\mathbb{E}_n\bigg[\int_{-\breve{\mathbf{b}}_i'\boldsymbol{\upsilon}}^{0}\Big\{ \Psi_1(x_i, \mathbf{w}_i;\eta(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0)) \Big(\eta^{(1)}(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0) (\breve{\mathbf{b}}_i'(\boldsymbol{\beta}_1+\boldsymbol{\beta}_2)+\mathbf{w}_i'\boldsymbol{\gamma}_1+t)\\ &+\frac{1}{2}\eta^{(2)}(\xi_{i,t})(\breve{\mathbf{b}}_i'(\boldsymbol{\beta}_1+\boldsymbol{\beta}_2)+\mathbf{w}_i'\boldsymbol{\gamma}_1+t)^2\Big)\\ &+\frac{1}{2}\Psi_2(x_i, \mathbf{w}_i; \tilde{\xi}_{i,t})\Big(\eta(\breve{\mathbf{b}}_i'(\breve{\boldsymbol{\beta}}_0+\boldsymbol{\beta}_1+\boldsymbol{\beta}_2)+\mathbf{w}_i'\boldsymbol{\gamma}+t)- \eta(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0)\Big)^2\Big\}\\ &\times \Big(\eta^{(1)}(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0)+ \eta^{(2)}(\check{\xi}_{i,t})(\breve{\mathbf{b}}_i'(\boldsymbol{\beta}_1+\boldsymbol{\beta}_2)+\mathbf{w}_i'\boldsymbol{\gamma}_1+t)\Big)dt\bigg]\\ =&\;\boldsymbol{\upsilon}'\widetilde{\mathbf{Q}}(\boldsymbol{\beta}_1+\boldsymbol{\beta}_2)+\boldsymbol{\upsilon}'\mathbb{E}_n[\mathbf{b}_i\widetilde{\varkappa}_i\mathbf{w}_i']\boldsymbol{\gamma}_1-\frac{1}{2}\boldsymbol{\upsilon}\widetilde{\mathbf{Q}}\boldsymbol{\upsilon}+\mathrm{I}+\mathrm{II}+\mathrm{III}, \end{align*} where $\xi_{i,t}$ and $\check{\xi}_{i,t}$ are between $\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0$ and $\breve{\mathbf{b}}_i'(\breve{\boldsymbol{\beta}}_0+\boldsymbol{\beta}_1+\boldsymbol{\beta}_2)+\mathbf{w}_i'\boldsymbol{\gamma}+t$, $\tilde{\xi}_{i,t}$ is between $\eta(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0)$ and $\eta(\breve{\mathbf{b}}_i'(\breve{\boldsymbol{\beta}}_0+\boldsymbol{\beta}_1+\boldsymbol{\beta}_2)+\mathbf{w}_i'\boldsymbol{\gamma}+t)$, $\Psi_2(x,\mathbf{w};\tau)=\frac{\partial^2}{\partial \tau^2}\Psi(x,\mathbf{w};\tau)$, $\widetilde{\varkappa}_i=\Psi_1(x_i,\mathbf{w}_i;\eta(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0)) (\eta^{(1)}(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0))^2$, $\boldsymbol{\upsilon}'\mathbb{E}_n[\mathbf{b}_i\widetilde{\varkappa}_i\mathbf{w}_i']\boldsymbol{\gamma}_1\lesssim\varepsilon_nr_{3,n}/J$, $-\frac{1}{2}\boldsymbol{\upsilon}\widetilde{\mathbf{Q}}\boldsymbol{\upsilon}\lesssim \varepsilon_n^2/J$, and $\mathrm{I}, \mathrm{II}$, and $\mathrm{III}$ are defined and bounded as follows: \begin{align*} \mathrm{I}&=\mathbb{E}_n\bigg[\int_{-\breve{\mathbf{b}}_i'\boldsymbol{\upsilon}}^{0} \Psi_1(x_i;\eta(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0)) \eta^{(1)}(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0)\\ &\times\eta^{(2)}(\check{\xi}_{i,t})(\breve{\mathbf{b}}_i'(\boldsymbol{\beta}_1+\boldsymbol{\beta}_2)+\mathbf{w}_i'\boldsymbol{\gamma}_1+t)^2dt\mathds{1}_{i,\boldsymbol{\upsilon}} \bigg] \lesssim\varepsilon_nJ^{-1}(r_{1,n}+r_{2,n}+\varepsilon_n+r_{3,n})^2,\\ \mathrm{II}&=\mathbb{E}_n\bigg[\int_{-\breve{\mathbf{b}}_i'\boldsymbol{\upsilon}}^{0} \Psi_1(x_i;\eta(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0))\times \frac{1}{2}\eta^{(2)}(\xi_{i,t})(\breve{\mathbf{b}}_i'(\boldsymbol{\beta}_1+\boldsymbol{\beta}_2)+\mathbf{w}_i'\boldsymbol{\gamma}_1+t)^2\\ &\times\eta^{(1)}\Big(\breve{\mathbf{b}}_i'(\breve{\boldsymbol{\beta}}_0+\boldsymbol{\beta}_1+\boldsymbol{\beta}_2)+\mathbf{w}_i'\boldsymbol{\gamma}+t\Big)dt\mathds{1}_{i,\boldsymbol{\upsilon}}\bigg] \lesssim\varepsilon_nJ^{-1}(r_{1,n}+r_{2,n}+\varepsilon_n+r_{3,n})^2,\\ \mathrm{III}&=\mathbb{E}_n\bigg[\int_{-\breve{\mathbf{b}}_i'\boldsymbol{\upsilon}}^{0} \frac{1}{2}\Psi_2(\tilde{\xi}_{i,t})\Big(\eta(\breve{\mathbf{b}}_i'(\breve{\boldsymbol{\beta}}_0+\boldsymbol{\beta}_1+\boldsymbol{\beta}_2)+\mathbf{w}_i'\boldsymbol{\gamma}+t)-\eta(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0)\Big)^2\\ &\times\eta^{(1)}\Big(\breve{\mathbf{b}}_i'(\breve{\boldsymbol{\beta}}_0+\boldsymbol{\beta}_1+\boldsymbol{\beta}_2)+\mathbf{w}_i'\boldsymbol{\gamma}+t\Big)dt\mathds{1}_{i,\boldsymbol{\upsilon}}\bigg] \lesssim\varepsilon_nJ^{-1}(r_{1,n}+r_{2,n}+\varepsilon_n+r_{3,n})^2. \end{align*} These bounds hold uniformly for $\boldsymbol{\upsilon}\in\mathcal{V}$, $\boldsymbol{\beta}_1\in\mathcal{H}_1$, $\boldsymbol{\beta}_2\in\mathcal{H}_2$ and $\boldsymbol{\gamma}_1\in\mathcal{H}_3$ (that is, uniformly over the function class $\mathcal{G}$), and on an event $\mathcal{A}_1\cap\mathcal{A}_2$ where $\mathcal{A}_2=\{\lambda_{\max}(\widetilde{\mathbf{Q}})\leq c_4J^{-1}\}$ for some large enough $c_4>0$. Note that $\mathbb{P}(\mathcal{A}_1\cap\mathcal{A}_2)\to 1$ by Lemma (ref). Step 4: By Assumption (ref) and Taylor's expansion, \begin{align*} \mathrm{IV}&=\mathbb{E}_n\bigg[\Big(\eta(\breve{\mathbf{b}}_i'(\breve{\boldsymbol{\beta}}_0+\boldsymbol{\beta}_1+\boldsymbol{\beta}_2)+\mathbf{w}_i'\boldsymbol{\gamma})- \eta(\breve{\mathbf{b}}_i'(\breve{\boldsymbol{\beta}}_0+\boldsymbol{\beta}_1+\boldsymbol{\beta}_2-\boldsymbol{\upsilon})+\mathbf{w}_i'\boldsymbol{\gamma})\Big) \psi(y_i,\eta(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0))\bigg]\\ &\qquad-\mathbb{E}_n\Big[\boldsymbol{\upsilon}'\breve{\mathbf{b}}_i\psi(y_i, \eta(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0)) \eta^{(1)}(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0)\Big]\\ &=\mathbb{E}_n\Big[\boldsymbol{\upsilon}'\breve{\mathbf{b}}_i\psi(y_i, \eta(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0))\Big( \eta^{(2)}(\xi_{i})(\breve{\mathbf{b}}_i'(\boldsymbol{\beta}_1+\boldsymbol{\beta}_2-\boldsymbol{\upsilon})+ \mathbf{w}_i'\boldsymbol{\gamma}_1)+ \frac{1}{2 }\eta^{(2)}(\tilde{\xi}_{i})\boldsymbol{\upsilon}'\breve{\mathbf{b}}_i\Big)\Big]\\ &\lesssim J^{-1}((J\log n/n)^{1/2}+J^{-p-1})(\varepsilon_n+r_{1,n}+r_{2,n}+r_{3,n})\varepsilon_n, \end{align*} where $\xi_i$ is between $\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0$ and $\breve{\mathbf{b}}_i'(\breve{\boldsymbol{\beta}}_0+\boldsymbol{\beta}_1+\boldsymbol{\beta}_2-\boldsymbol{\upsilon})+\mathbf{w}_i'\boldsymbol{\gamma}$ and $\tilde{\xi}_i$ is between $\breve{\mathbf{b}}_i'(\breve{\boldsymbol{\beta}}_0+\boldsymbol{\beta}_1+\boldsymbol{\beta}_2)+\mathbf{w}_i'\boldsymbol{\gamma}$ and $\breve{\mathbf{b}}_i'(\breve{\boldsymbol{\beta}}_0+\boldsymbol{\beta}_1+\boldsymbol{\beta}_2-\boldsymbol{\upsilon})+\mathbf{w}_i'\boldsymbol{\gamma}$. The last line holds on the event \[ \begin{split} \mathcal{A}_3=\bigg\{ &\sup\;\bigg(\Big\|\mathbb{E}_n\Big[\breve{\mathbf{b}}_i\breve{\mathbf{b}}_i'\psi(y_i,\eta(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0))\eta^{(2)}(\varpi_i)\Big]\Big\|_\infty+\\ &\qquad\;\;\Big\|\mathbb{E}_n\Big[\breve{\mathbf{b}}_i\psi(y_i,\eta(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0))\eta^{(2)}(\varpi_i)\mathbf{w}_i\Big]\Big\|_\infty\bigg) \lesssim J^{-1}\Big(\Big(\frac{J\log n}{n}\Big)^{1/2}+J^{-p-1}\Big) \bigg\}, \end{split} \] where the supremum is taken over $\boldsymbol{\beta}_1\in\mathcal{H}_1, \boldsymbol{\beta}_2\in\mathcal{H}_2, \boldsymbol{\upsilon}\in\mathcal{V}, \boldsymbol{\gamma}_1\in\mathcal{H}_3$ and $\varpi_i$ within the range of $\xi_i$ or $\tilde\xi_i$. Note that $\mathbb{E}[\psi(y_i, \eta_i)|\mathscr{F}_{XW\Delta}]=0$ and $\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0-\mu_0(x_i)\lesssim J^{-p-1}$. Then, we can use the argument in the proof of Lemmas (ref) and (ref) to obtain $\mathbb{P}(\mathcal{A}_3)\to 1$ by choosing $C_3>0$ sufficiently large. Step 5: Let $\bar\boldsymbol{\upsilon}=c_5\varepsilon_nJ^{-1}[\bar{\mathbf{Q}}^{-1}]_{k\cdot}$ for some $k$ such that $|\beta_{2,k}|=\|\boldsymbol{\beta}_2\|_\infty$ for some $c_5>0$ where $[\bar{\mathbf{Q}}^{-1}]_{k\cdot}$ denotes the $k$th row of $\bar{\mathbf{Q}}^{-1}$. Note that $\boldsymbol{\upsilon}'\bar{\mathbf{Q}}\boldsymbol{\beta}_2=\beta_{2,k}$. Take $\boldsymbol{\upsilon}=(\upsilon_1, \cdots, \upsilon_{K_{p,s}})$ where $\upsilon_j=\bar{\upsilon}_j$ for $|j-k|\leq M_n$ and zero otherwise. Clearly, $\boldsymbol{\upsilon}\in\mathcal{V}$ on an event $\mathcal{A}_4$ with $\mathbb{P}(\mathcal{A}_4)\to 1$. On $\mathcal{A}_2\cap\mathcal{A}_4$, \[ |(\boldsymbol{\upsilon}-\bar{\boldsymbol{\upsilon}})'\bar{\mathbf{Q}}\boldsymbol{\beta}_2|\lesssim \varepsilon_nJ^{-1}r_{2,n}n^{-c_6} \] for some large $c_6>0$ if we let $c_1$ be sufficiently large. \textbf{Step 6:} Finally, partition the whole parameter space into shells: $\mathcal{O}=\cup_{\ell=-\infty}^{\bar{L}}\mathcal{O}_{\ell}$ where $\mathcal{O}_{\ell}=\{\boldsymbol{\beta}\in\mathbb{R}^{K_{p,s}}: 2^{\ell-1} \mathfrak{r}_{2,n}\leq \|\boldsymbol{\beta}-\breve{\boldsymbol{\beta}}_0-\bar{\boldsymbol{\beta}}\|_\infty\leq 2^\ell \mathfrak{r}_{2,n}\}$ for the smallest $\bar{L}$ such that $2^{\bar{L}}r_{2,n}\geq c$, and $\bar{\mathbf{Q}}\bar{\boldsymbol{\beta}}=-\mathbb{E}_n[\breve{\mathbf{b}}_i\eta^{(1)}(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0)\psi(y_i,\eta(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0))]$. Define $\mathcal{A}=\cap_{j=1}^4\mathcal{A}_j$. Then, for some constant $L\leq \bar{L}$, we have by Lemma (ref) and the results given in the previous steps, \begin{align*} &\;\mathbb{P}(\|\breve{\boldsymbol{\beta}}-\breve{\boldsymbol{\beta}}_0-\bar{\boldsymbol{\beta}}\|_\infty\geq 2^Lr_{2,n}|\mathscr{F}_{XW\Delta})\\ \leq\;& \mathbb{P}\Big(\bigcup_{\ell=L}^{\bar{L}} \Big\{\inf_{\boldsymbol{\beta}\in\mathcal{O}_{\ell}}\sup_{\boldsymbol{\upsilon}\in\mathcal{V}} \;\mathbb{E}_n[\rho(y_i;\eta(\breve{\mathbf{b}}_i'\boldsymbol{\beta}+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}}))- \rho(y_i;\eta(\breve{\mathbf{b}}_i'(\boldsymbol{\beta}-\boldsymbol{\upsilon})+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}}))]<0\Big\}\Big|\mathscr{F}_{XW\Delta}\Big)+o_\mathbb{P}(1)\\ =\;& \mathbb{P}\Big(\bigcup_{\ell=L}^{\bar{L}}\Big\{\inf_{\boldsymbol{\beta}\in\mathcal{O}_{\ell}}\sup_{\boldsymbol{\upsilon}\in\mathcal{V}}\Big\{ \mathbb{E}\Big[\rho(y_i;\eta(\breve{\mathbf{b}}_i'\boldsymbol{\beta}+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}}))- \rho(y_i;\eta(\breve{\mathbf{b}}_i'(\boldsymbol{\beta}-\boldsymbol{\upsilon})+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}}))\\ &-[\eta(\breve{\mathbf{b}}_i'\boldsymbol{\beta}+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}})- \eta(\breve{\mathbf{b}}_i'(\boldsymbol{\beta}-\boldsymbol{\upsilon})+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}})]\psi(y_i,\eta(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}}))|\mathscr{F}_{XW\Delta}\Big]+\\ &\mathbb{E}_n\Big[(\eta(\breve{\mathbf{b}}_i'\boldsymbol{\beta}+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}})-\eta(\breve{\mathbf{b}}_i'(\boldsymbol{\beta}-\boldsymbol{\upsilon})+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}}))\psi(y_i,\eta(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}}))\Big]+\\ &\frac{1}{\sqrt{n}}\mathbb{G}_n\Big[\rho(y_i;\eta(\breve{\mathbf{b}}_i'\boldsymbol{\beta}+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}}))- \rho(y_i;\eta(\breve{\mathbf{b}}_i'(\boldsymbol{\beta}-\boldsymbol{\upsilon})+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}}))-\\ & [\eta(\breve{\mathbf{b}}_i'\boldsymbol{\beta}+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}})-\eta(\breve{\mathbf{b}}_i'(\boldsymbol{\beta}-\boldsymbol{\upsilon})+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}})] \psi(y_i,\eta(\breve{\mathbf{b}}_i'\breve{\boldsymbol{\beta}}_0+\mathbf{w}_i'\widehat{\boldsymbol{\gamma}}))\Big]\Big\}<0\Big\}\Big|\mathscr{F}_{XW\Delta}\Big)+o_\mathbb{P}(1)\\ \leq&\;\mathbb{P}\Big(\bigcup_{\ell=L}^{\bar{L}}\Big\{\sup_{\boldsymbol{\beta}_1\in\mathcal{H}_1}\sup_{\boldsymbol{\beta}_2\in\mathcal{H}_{2,\ell}}\sup_{\boldsymbol{\gamma}_1\in\mathcal{H}_3}\sup_{\boldsymbol{\upsilon}\in\mathcal{V}} \frac{1}{\sqrt{n}}\Big|(\mathds{1}(\mathcal{A}_1)+\mathds{1}(\mathcal{A}_1^c)) \mathbb{G}_n[\delta_i(\boldsymbol{\beta}_1, \boldsymbol{\beta}_2, \boldsymbol{\upsilon}, \boldsymbol{\gamma})]\Big|>\\ &C_4J^{-1}2^{\ell} r_{2,n}\varepsilon_n\Big\}\cap\mathcal{A}\Big|\mathscr{F}_{XW\Delta}\Big) +o_\mathbb{P}(1)\\ \leq &\sum_{\ell=L}^{\bar{L}}(C_6J^{-1}2^{\ell} \mathfrak{r} _{2,n}\varepsilon_n)^{-1} \mathds{1}(\mathcal{A}_1)\mathbb{E}\Big[\sup_{\boldsymbol{\beta}_1\in\mathcal{H}_1}\sup_{\boldsymbol{\beta}_2\in\mathcal{H}_{2,\ell}} \sup_{\boldsymbol{\gamma}_1\in\mathcal{H}_3}\sup_{\boldsymbol{\upsilon}\in\mathcal{V}}\frac{1}{\sqrt{n}}\mathbb{G}_n[\delta_i(\boldsymbol{\beta}_1, \boldsymbol{\beta}_2, \boldsymbol{\upsilon}, \boldsymbol{\gamma})]\Big|\mathscr{F}_{XW\Delta}\Big]+o_\mathbb{P}(1), \end{align*} where $\mathbb{G}_n[\cdot]$ is understood as $\sqrt{n}(\mathbb{E}_n[\cdot]-\mathbb{E}[\cdot|\mathscr{F}_{XW}])$ in the above, we let $\varepsilon_n=2^Lr_{2,n}$, and $\mathds{1}(\mathcal{A}_1)$ is an indicator of the event $\mathcal{A}_1$. Using the result in Step 1 and the rate condition, the first term in the last line can be made arbitrarily small by choosing $L$ large enough, when $n$ is sufficiently large. Then, the proof for part (i) is complete. \textbf{Step 7:} To show part (ii) and part (iii), by Taylor expansion and the result in part (i), \begin{align*} &\eta(\widehat{\mu}(x)+\widehat\mathsf{w}'\widehat{\boldsymbol{\gamma}})-\eta(\mu_0(x)+\mathsf{w}'\boldsymbol{\gamma}_0)\\ =\;&\eta^{(1)}(\mu_0(x)+\mathsf{w}'\boldsymbol{\gamma}_0) \Big(\widehat{\mathbf{b}}_{p,s}(x)'\widehat{\boldsymbol{\beta}}-\mu_0(x)\Big)\\ &+O_\mathbb{P}\Big(\|\widehat\mathsf{w}-\mathsf{w}\|+\|\widehat{\boldsymbol{\gamma}}-\boldsymbol{\gamma}_0\|+\frac{J\log n}{n}+J^{-2p-2}+\mathfrak{r}_{2,n}^2\Big)\\ =\;&-\eta^{(1)}(\mu_0(x)+\mathsf{w}'\boldsymbol{\gamma}_0)\widehat{\mathbf{b}}_{p,s}(x)'\bar{\mathbf{Q}}^{-1}\mathbb{E}_n[\widehat{\mathbf{b}}_{p,s}(x_i)\eta_{i,1}\psi(y_i,\eta_i)]\\ &+O_\mathbb{P}\Big(J^{-p-1}+ \Big(\frac{J\log n}{n}\Big)^{3/4}\log n+J^{-\frac{p+1}{2}}\Big(\frac{J\log^2 n}{n}\Big)^{1/2}+\mathfrak{r}_\gamma+ \|\widehat\mathsf{w}-\mathsf{w}\|\Big), \end{align*} and \begin{align*} &\eta^{(1)}(\widehat{\mu}(x)+\widehat\mathsf{w}'\widehat{\boldsymbol{\gamma}})\widehat{\mu}^{(1)}(x)- \eta^{(1)}(\mu_0(x)+\mathsf{w}'\boldsymbol{\gamma}_0)\mu_0^{(1)}(x)\\ =&\;\eta^{(1)}(\mu_0(x)+\mathsf{w}'\boldsymbol{\gamma}_0) \Big(\widehat{\mu}^{(1)}(x)-\mu_0^{(1)}(x)\Big)\\ &+O_\mathbb{P}\Big(\Big(\frac{J\log n}{n}\Big)^{1/2}+J^{-p-1}+\|\widehat\mathsf{w}-\mathsf{w}\|+\mathfrak{r}_{2,n}\Big) O_\mathbb{P}\Big(1+J\Big(\Big(\frac{J\log n}{n}\Big)^{1/2}+J^{-p-1}+\mathfrak{r}_{2,n}\Big)\Big)\\ =&\;-\eta^{(1)}(\mu_0(x)+\mathsf{w}'\boldsymbol{\gamma}_0)\widehat{\mathbf{b}}_{p,s}^{(1)}(x)'\bar{\mathbf{Q}}^{-1}\mathbb{E}_n[\widehat{\mathbf{b}}_{p,s}(x_i)\eta_{i,1}\psi(y_i,\eta_i)]+\\ &O_\mathbb{P}\Big(\Big(\frac{J\log n}{n}\Big)^{1/2}+J^{-p}+J\Big(\frac{J\log n}{n}\Big)^{3/4}\log n+J^{-\frac{p-1}{2}}\Big(\frac{J\log^2 n}{n}\Big)^{1/2}+J\mathfrak{r}_\gamma\\ &+ \|\widehat{\mathsf{w}}-\mathsf{w}\|\Big(1+\Big(\frac{J^3\log n}{n}\Big)^{1/2}\Big)\Big). \end{align*} In the above derivation the probability bound holds uniformly over $x\in\mathcal{X}$ as well. Then the proof is complete.

Proof of Theorem (ref)

proofSince $\widehat{\epsilon}_{i}:=\epsilon_{i}+\eta_i-\widehat{\eta}_i=:\epsilon_{i}+u_{i}$, we can write \begin{align*} &\mathbb{E}_n[\widehat{\mathbf{b}}_{p,s}(x_{i})\widehat{\mathbf{b}}_{p,s}(x_{i})'\widehat{\eta}_{i,1}^2\psi^\ddagger(\widehat{\eta}_i)^2\psi^\dagger(\widehat{\epsilon}_{i})^2] -\mathbb{E}[\mathbf{b}_{p,s}(x_{i})\mathbf{b}_{p,s}(x_{i})'\eta_{i,1}^2\sigma^2(x_i, \mathbf{w}_i)]\\ =\;&\mathbb{E}_n\Big[\widehat{\mathbf{b}}_{p,s}(x_{i})\widehat{\mathbf{b}}_{p,s}(x_{i})' \widehat{\eta}_{i,1}^2\psi^\ddagger(\widehat{\eta}_i)^2\Big(\psi^\dagger(\epsilon_i+u_i)^2-\psi^\dagger(\epsilon_i)^2\Big)\Big]\\ &+\mathbb{E}_n\Big[\widehat{\mathbf{b}}_{p,s}(x_{i}) \widehat{\mathbf{b}}_{p,s}(x_{i})'\Big(\widehat{\eta}_{i,1}^2\psi^\ddagger(\widehat{\eta}_i)^2-\eta_{i,1}^2\psi^\ddagger(\eta_i)^2\Big)\psi^\dagger(\epsilon_i)^2\Big]\\ &+\mathbb{E}_n[\widehat{\mathbf{b}}_{p,s}(x_{i})\widehat{\mathbf{b}}_{p,s}(x_{i})' \eta_{i,1}^2(\psi(y_i, \eta_i)^2-\sigma^2(x_i, \mathbf{w}_i))]\\ &+\Big(\mathbb{E}_n[\widehat{\mathbf{b}}_{p,s}(x_{i})\widehat{\mathbf{b}}_{p,s}(x_{i})'\eta_{i,1}^2\sigma^2(x_i, \mathbf{w}_i)]- \mathbb{E}[\mathbf{b}_{p,s}(x_{i})\mathbf{b}_{p,s}(x_{i})'\eta_{i,1}^2\sigma^2(x_i, \mathbf{w}_i)]\Big)\\ =:&\mathbf{V}_1+\mathbf{V}_2+\mathbf{V}_3+\mathbf{V}_4. \end{align*} We bound each term in the following. The first part of the theorem only concerns $\mathbf{V}_1+\mathbf{V}_2+\mathbf{V}_3$, and the second part needs a bound on $\mathbf{V}_4$ as well where the additional Assumption (ref)(ii) is used. Step 1: For $\mathbf{V}_1$, we further write $\mathbf{V}_1=\mathbf{V}_{11}+\mathbf{V}_{12}$ where \[ \begin{split} &\mathbf{V}_{11}:=\mathbb{E}_n\Big[\widehat{\mathbf{b}}_{p,s}(x_{i})\widehat{\mathbf{b}}_{p,s}(x_{i})' \eta_{i,1}^2\psi^\ddagger(\eta_i)^2\Big(\psi^\dagger(\epsilon_i+u_i)^2-\psi^\dagger(\epsilon_i)^2\Big)\Big],\\ &\mathbf{V}_{12}:=\mathbb{E}_n\Big[\widehat{\mathbf{b}}_{p,s}(x_{i})\widehat{\mathbf{b}}_{p,s}(x_{i})' \Big(\widehat{\eta}_{i,1}^2\psi^\ddagger(\widehat{\eta}_i)^2-\eta_{i,1}^2\psi^\ddagger(\eta_i)^2\Big)\Big(\psi^\dagger(\epsilon_i+u_i)^2-\psi^\dagger(\epsilon_i)^2\Big)\Big]. \end{split} \] Let $r_{1,n}=C_1(J\log n/n)^{1/2}+J^{-p-1}$ for a constant $C_1>0$. By Assumption (ref) and Corollary (ref), $\max_{1\leq i\leq n}|u_i|\leq r_{1,n}$ with arbitrarily large probability for $C_1$ sufficiently large. For $\mathbf{V}_{11}$, let $\mathcal{J}$ be the set of all discontinuity points of $\psi(\cdot)$. Define $\mathds{1}_{i,\mathcal{D}}:=\mathds{1}(\epsilon_i\in\mathcal{D})$ and $\mathds{1}_{i,\mathcal{D}^c}:=(1-\mathds{1}_{i,\mathcal{D}})$ where $\mathcal{D}:=\{a: |a-\jmath|\leq r_{1,n} \text{ for some } \jmath\in\mathcal{J} \}$. Define \begin{align*} &\mathbf{V}_{111}:=\mathbb{E}_n\Big[\widehat{\mathbf{b}}_{p,s}(x_{i})\widehat{\mathbf{b}}_{p,s}(x_{i})' \eta_{i,1}^2\psi^\ddagger(\eta_i)^2\Big(\psi^\dagger(\epsilon_i+u_i)^2-\psi^\dagger(\epsilon_i)^2\Big)\mathds{1}_{i,\mathcal{D}}\Big],\\ &\mathbf{V}_{112}:=\mathbb{E}_n\Big[\widehat{\mathbf{b}}_{p,s}(x_{i})\widehat{\mathbf{b}}_{p,s}(x_{i})' \eta_{i,1}^2\psi^\ddagger(\eta_i)^2\Big(\psi^\dagger(\epsilon_i+u_i)^2-\psi^\dagger(\epsilon_i)^2\Big)\mathds{1}_{i,\mathcal{D}^c}\Big]. \end{align*} By definition of $\mathcal{D}$ and Assumption (ref), \[ \|\mathbf{V}_{111}\| \lesssim\|\mathbb{E}_n[\widehat{\mathbf{b}}_{p,s}(x_{i})\widehat{\mathbf{b}}_{p,s}(x_{i})' \mathbb{E}[\mathds{1}_{i,\mathcal{D}}|\mathscr{F}_{XW\Delta}]]\|+ \|\mathbb{E}_n[\widehat{\mathbf{b}}_{p,s}(x_{i})\widehat{\mathbf{b}}_{p,s}(x_{i})' (\mathds{1}_{i,\mathcal{D}}-\mathbb{E}[\mathds{1}_{i,\mathcal{D}}|\mathscr{F}_{XW\Delta}])]\|. \] By Assumption (ref) and Lemma SA-3.5 of Cattaneo-Crump-Farrell-Feng_2024_AER, the first term on the right hand side is $O_\mathbb{P}(r_{1,n})$. For the second term, conditional on $\mathscr{F}_{XW\Delta}$, it is an independent sequence with mean zero. Thus, we can apply the argument given in Step 3 below and conclude that the second term is $O_\mathbb{P}(\sqrt{r_{1,n}J\log J/n}+J\log J/n)$. In this case, the indicator $\mathds{1}_{i,\mathcal{D}}$ is trivially bounded uniformly. On the other hand, by Assumption (ref), $$\|\mathbf{V}_{112}\|\lesssim r_{1,n} \|\mathbb{E}_n[\widehat{\mathbf{b}}_{p,s}(x_{i})\widehat{\mathbf{b}}_{p,s}(x_{i})' \eta_{i,1}^2\psi^\ddagger(\eta_i)^2|\psi^\dagger(\epsilon_i+u_i)+\psi^\dagger(\epsilon_i)|]\|. $$ Since $|c|\leq \frac{1}{2}(1+c^2)$ for any scalar $c$, we have \[ \mathbb{E}_n\Big[\widehat{\mathbf{b}}_{p,s}(x_{i})\widehat{\mathbf{b}}_{p,s}(x_{i})' \eta_{i,1}^2\psi^\ddagger(\eta_i)^2|\psi^\dagger(\epsilon_i)|\Big]\leq \frac{1}{2}\mathbb{E}_n\Big[\widehat{\mathbf{b}}_{p,s}(x_{i})\widehat{\mathbf{b}}_{p,s}(x_{i})' \eta_{i,1}^2\psi^\ddagger(\eta_i)^2(1+\psi^\dagger(\epsilon_i)^2)\Big]\lesssim_\mathbb{P} 1, \] by Lemma (ref) and the result in Step 3. In addition, we further write \begin{align*} &\mathbb{E}_n\Big[\widehat{\mathbf{b}}_{p,s}(x_{i})\widehat{\mathbf{b}}_{p,s}(x_{i})' \eta_{i,1}^2\psi^\ddagger(\eta_i)^2|\psi^\dagger(\epsilon_{i}+u_i)|\Big]\\ =&\,\mathbb{E}_n\Big[\widehat{\mathbf{b}}_{p,s}(x_{i})\widehat{\mathbf{b}}_{p,s}(x_{i})' \eta_{i,1}^2\psi^\ddagger(\eta_i)^2 |\psi^\dagger(\epsilon_i)+(\psi^\dagger(\epsilon_i+u_i)-\psi^\dagger(\epsilon_i))|\Big]. \end{align*} Repeat the previous argument to bound this term. We conclude that $\|\mathbf{V}_{11}\|\lesssim_\mathbb{P} r_{1,n}$. $\mathbf{V}_{12}$ can be treated using the previous argument combined with the argument given in Step 2 and the result in Step 3. It leads to $\|\mathbf{V}_{12}\|\lesssim_\mathbb{P} r_{1,n}$. Step 2: For $\mathbf{V}_2$, by Assumption (ref), Corollary (ref) and the argument given later in Step 3, we have $$\|\mathbf{V}_2\|\leq \max_{1\leq i\leq n}|\widehat{\eta}_{i,1}^2\psi^\ddagger(\widehat\eta_i)^2-\eta_{i,1}^2\psi^\ddagger(\eta_i)^2|\|\mathbb{E}_n[\widehat{\mathbf{b}}_{p,s}(x_i)\widehat{\mathbf{b}}_{p,s}(x_i)'\psi^\dagger(\epsilon_i)^2]\| \lesssim_\mathbb{P} (J\log n/n)^{1/2}+J^{-p-1}.$$ Step 3: For $\mathbf{V}_3$, in view of Lemmas (ref) and (ref), it suffices to show that \[\sup_{\Delta\in\Pi}\Big\|\mathbb{E}_n[\mathbf{b}_{p,0}(x_{i};\Delta) \mathbf{b}_{p,0}(x_{i};\Delta)'\eta_{i,1}^2(\psi(y_i, \eta_i)^2-\sigma^2(x_i, \mathbf{w}_i))]\Big\|\lesssim_\mathbb{P} \Big(\frac{J\log J}{n^{\frac{\nu-2}{\nu}}}\Big)^{1/2}. \] For notational simplicity, we write $\varphi_{i}=\psi(y_i, \eta_i)^2-\sigma^2(x_i, \mathbf{w}_i)$, $\varphi_{i}^-=\varphi_{i}\mathds{1}(|\varphi_{i}|\leq M)-\mathbb{E}[\varphi_{i}\mathds{1}(|\varphi_{i}|\leq M)|x_i,\mathbf{w}_i]$, $\varphi_{i}^+=\varphi_{i}\mathds{1}(|\varphi_{i}|> M)-\mathbb{E}[\varphi_{i}\mathds{1}(|\varphi_{i}|> M)|x_i,\mathbf{w}_i]$ for some $M>0$ to be specified later. Since $\mathbb{E}[\varphi_{i}|x_i, \mathbf{w}_i]=0$, $\varphi_{i}=\varphi_{i}^-+\varphi_{i}^+$. Then, define a function class \[\mathcal{G}=\Big\{(x_{1}, \mathbf{w}_1, \varphi_{1})\mapsto b_{p,0,l}(x_{1};\Delta)b_{p,0,k}(x_{1};\Delta)\eta_{i,1}^2\varphi_{1}:1\leq l\leq J(p+1), 1\leq k\leq J(p+1), \Delta\in\Pi\Big\}. \] For $g\in\mathcal{G}$, $\sum_{i=1}^{n}g(x_i, \mathbf{w}_i, \varphi_i) =\sum_{i=1}^{n}g(x_i, \mathbf{w}_i, \varphi_i^+)+\sum_{i=1}^{n}g(x_i, \mathbf{w}_i, \varphi_i^-)$. For the truncated piece, we have $\sup_{g\in\mathcal{G}}|g(x_i, \mathbf{w}_i, \varphi_i^-)|\lesssim JM$, and \begin{align*} \sup_{g\in\mathcal{G}}\mathbb{V}[g(x_{1}, \mathbf{w}_1, \varphi_{1}^-)] &\lesssim \sup_{x\in\mathcal{X}, \mathbf{w}\in\mathcal{W}}\mathbb{E}[(\varphi_{i}^-)^2|x_i=x, \mathbf{w}_i=\mathbf{w}]\sup_{\Delta\in\Pi}\sup_{1\leq l,k\leq J(p+1)}\mathbb{E}[b_{p,0,l}^2(x_i;\Delta)b_{p,0,k}^2(x_i;\Delta)\eta_{i,1}^4] \\ &\lesssim JM \sup_{x\in\mathcal{X}, \mathbf{w}\in\mathcal{W}} \mathbb{E}\Big[|\varphi_{1}|\Big|x_i=x\Big]\lesssim JM. \end{align*} The VC condition holds by the same argument given in the proof of Lemma (ref). Then, by Lemma (ref), \[ \mathbb{E}\Big[\sup_{g\in\mathcal{G}}\Big|\mathbb{E}_n[g(x_{i}, \mathbf{w}_i, \varphi_{i}^-)]\Big|\Big] \lesssim \sqrt{\frac{JM\log (JM)}{n}}+\frac{JM\log(JM)}{n}. \] Regarding the tail, we apply Theorem 2.14.1 of \citet*{vandevarrt-Wellner_1996_book} and obtain \begin{align*} \mathbb{E}\Big[\sup_{g\in\mathcal{G}}\Big|\mathbb{E}_n[g(x_{i}, \mathbf{w}_i, \varphi_{i}^+)]\Big|\Big] &\lesssim \frac{1}{\sqrt{n}}J \mathbb{E}\Big[\sqrt{\mathbb{E}_n[|\varphi_{i}^+|^2]}\Big]\\ &\leq \frac{1}{\sqrt{n}}J (\mathbb{E}[\max_{1\leq i\leq n}|\varphi_{i}^+|])^{1/2}(\mathbb{E}[\mathbb{E}_n[|\varphi_{i}^+|])^{1/2}\\ &\lesssim \frac{J}{\sqrt{n}}\cdot \frac{n^{\frac{1}{\nu}}}{M^{(\nu-2)/4}}, \end{align*} where the second line follows from Cauchy-Schwarz inequality and the third line uses the fact that \[ \mathbb{E}[\max_{1\leq i\leq n}|\varphi_{i}^+|]\lesssim \mathbb{E}[\max_{1\leq i\leq n}\psi(y_i, \eta_i)^2]\lesssim n^{2/\nu} \quad \text{and}\quad \mathbb{E}[\mathbb{E}_n[|\varphi_{i}^+|]]\leq \mathbb{E}[|\varphi_{1}|^+|]\lesssim \frac{\mathbb{E}[|\psi(y_1,\eta_1)|^{\nu}]}{M^{(\nu-2)/2}}. \] Then the desired result follows simply by setting $M=J^{\frac{2}{\nu-2}}$ and the sparsity of the basis. Step 4: For $\mathbf{V}_4$, since by Assumption (ref), $\sup_{x\in\mathcal{X}, \mathbf{w}\in\mathcal{W}}\mathbb{E}[\psi(y_i, \eta_i)^2|x_i=x]\lesssim 1$. Then, by the same argument given in the proof of Lemma (ref), \begin{align*} &\sup_{\Delta\in\Pi}\Big\|\frac{1}{\sqrt{n}}\mathbb{G}_n[\mathbf{b}_{p,s}(x_{i};\Delta)\mathbf{b}_{p,s}(x_{i};\Delta)' \eta_{i,1}^2\sigma^2(x_i, \mathbf{w}_i)]\Big\| \lesssim_\mathbb{P}\sqrt{J\log J/n} \quadand\\ &\Big\|\mathbb{E}_{\widehat{\Delta}}\Big[\widehat{\mathbf{b}}_{p,s}(x_{i})\widehat{\mathbf{b}}_{p,s}(x_{i})' \eta_{i,1}^2\psi(y_i, \eta_i)^2\Big]- \mathbb{E}\Big[\mathbf{b}_{p,s}(x_{i})\mathbf{b}_{p,s}(x_{i})'\eta_{i,1}^2\psi(y_i, \eta_i)^2\Big]\Big\|\lesssim_\mathbb{P} \sqrt{J\log J/n}+\mathfrak{r}_{\tt RP}. \end{align*} The proof for the first conclusion is complete. Step 5: The results about $\widehat\Omega_{\mu^{(v)}}(x)$, $\widehat\Omega_\vartheta(x)$ and $\widehat\Omega_\zeta(x)$ follow by Assumptions (ref) and (ref), Lemmas (ref) and (ref), and Corollary (ref). The proof is complete.

Proof of Theorem (ref)

proofWe first show that for each fixed $x\in\mathcal{X}$, \[ \bar{\Omega}_{\mu^{(v)}}(x)^{-1/2}\widehat{\mathbf{b}}_{p,s}^{(v)}(x)'\bar{\mathbf{Q}}^{-1}\mathbb{G}_n[\widehat{\mathbf{b}}_{p,s}(x_i)\eta_{i,1}\psi(y_i,\eta_i)]=:\mathbb{G}_n[a_{i}\psi(y_i, \eta_i)] \] is asymptotically normal. Conditional on $\mathscr{F}_{XW\Delta}$, the $\sigma$-field generated by $\{(x_i,\mathbf{w}_i)\}_{i=1}^n$ and $\widehat\Delta$, it is an independent mean-zero sequence over $i$ with variance equal to $1$. Then by Berry-Esseen inequality, \begin{align*} \sup_{u\in\mathbb{R}} \Big|\mathbb{P}(\mathbb{G}_n[a_i\psi(y_i,\eta_i)]\leq u|)-\Phi(u)\Big| \leq \min\bigg(1, \;\frac{\sum_{i=1}^{n}\mathbb{E}[|a_{i}\psi(y_i, \eta_i)|^3|\mathscr{F}_{XW\Delta}]}{n^{3/2}}\bigg). \end{align*} By Lemmas (ref), (ref) and (ref), \begin{align*} &\quad\;\frac{1}{n^{3/2}}\sum_{i=1}^{n}\mathbb{E}\Big[|a_{i}\psi(y_i, \eta_i)|^3\Big|\mathscr{F}_{XW\Delta}\Big]\\ &\lesssim \bar{\Omega}_{\mu^{(v)}}(x)^{-3/2}\frac{1}{n^{3/2}}\sum_{i=1}^{n}\mathbb{E}\Big[ |\widehat{\mathbf{b}}_{p,s}^{(v)}(x)'\bar{\mathbf{Q}}^{-1}\widehat{\mathbf{b}}_{p,s}(x_{i})\eta_{i,1}\psi(y_i, \eta_i)|^3\Big|\mathscr{F}_{XW\Delta}\Big]\\ &\lesssim \bar{\Omega}_{\mu^{(v)}}(x)^{-3/2}\frac{1}{n^{3/2}}\sum_{i=1}^n|\widehat{\mathbf{b}}_{p,s}^{(v)}(x)'\bar{\mathbf{Q}}^{-1}\widehat{\mathbf{b}}_{p,s}(x_{i})|^3\\ &\leq\bar{\Omega}_{\mu^{(v)}}(x)^{-3/2}\frac{\sup_{x\in\mathcal{X}}\sup_{z\in\mathcal{X}} |\widehat{\mathbf{b}}_{p,s}^{(v)}(x)'\bar{\mathbf{Q}}^{-1}\widehat{\mathbf{b}}_{p,s}(z)|}{n^{3/2}} \sum_{i=1}^{n} |\widehat{\mathbf{b}}_{p,s}^{(v)}(x)'\bar{\mathbf{Q}}^{-1}\widehat{\mathbf{b}}_{p,s}(x_{i})|^2\\ &\lesssim_\mathbb{P}\frac{1}{J^{3/2+3v}}\cdot\frac{J^{1+v}}{\sqrt{n}}\cdot J^{1+2v} \to 0 \end{align*} since $J/n=o(1)$. By Theorem (ref), the above weak convergence still holds if $\bar{\Omega}_{\mu^{(v)}}(x)$ is replaced by $\widehat{\Omega}_{\mu^{(v)}}(x)$. Then, the desired results follow by Theorem (ref).

Proof of Theorem (ref)

proofWe let $\widehat\boldsymbol{\beta}_0$ and $\widehat{r}_{0,v}$ be defined as in Lemma (ref). By Lemmas (ref) and (ref), Theorem (ref) and the results given in the proof of Lemma (ref), we have \[ \begin{split} \widehat{\mu}^{(v)}(x)-\mu_0^{(v)}(x) =&\,\widehat{\mathbf{b}}_{p,s}(x_i)'(\widehat{\boldsymbol{\beta}}-\widehat{\boldsymbol{\beta}}_0)-\widehat{r}_{0,v}(x)\\ =&-\widehat{\mathbf{b}}_{p,s}^{(v)}(x)'\mathbf{Q}_0^{-1}\mathbb{E}_n[\widehat{\mathbf{b}}_{p,s}(x_i)\eta_{i,1}\psi(y_i,\eta_i)] -\widehat{\mathbf{b}}_{p,s}^{(v)}(x)'\mathbf{Q}_0^{-1}\mathbb{E}_n[\widehat{\mathbf{b}}_{p,s}(x_i)\eta_{i,1} \Psi(x_i, \mathbf{w}_i;\check{\eta}_i)]\\ &\hspace{2em}-\widehat{r}_{0,v}(x)+ O_\mathbb{P}\Big(J^v\Big\{\Big(\frac{J\log n}{n}\Big)^{3/4}\sqrt{\log n}+ J^{-\frac{p+1}{2}}\Big(\frac{J\log^2 n}{n}\Big)^{1/2}+\mathfrak{r}_\gamma\Big\}\Big), \end{split} \] where $\check{\eta}_i=\eta(\widehat{\mathbf{b}}_{p,s}(x_i)'\widehat{\boldsymbol{\beta}}_0+\mathbf{w}_i'\boldsymbol{\gamma}_0)$. Recall that the $O_\mathbb{P}(\cdot)$ in the last line holds uniformly over $x\in\mathcal{X}$, and thus the integral of the squared remainder is $o_\mathbb{P}(J^{1+2v}/n+J^{-2(p+1-v)})$ by the rate condition imposed. Then, \begin{align*} \mathtt{AISE}_{\mu^{(v)}}=\int_{\mathcal{X}}\Big(&\widehat{\mathbf{b}}_{p,s}^{(v)}(x)'\mathbf{Q}_0^{-1}\mathbb{E}_n[\widehat{\mathbf{b}}_{p,s}(x_i)\eta_{i,1}\psi(y_i,\eta_i)]\\ &+\widehat{\mathbf{b}}_{p,s}^{(v)}(x)'\mathbf{Q}_0^{-1}\mathbb{E}_n[\widehat{\mathbf{b}}_{p,s}(x_i)\eta_{i,1} \Psi(x_i, \mathbf{w}_i;\check{\eta}_i)] +\widehat{r}_{0,v}(x)\Big)^2\omega(x)dx. \end{align*} Next, taking conditional expectation given $\mathbf{X}$, $\mathbf{W}$ and $\widehat\Delta$ and using the argument in the proof of Lemma (ref) again, we have \begin{align*} \mathbb{E}[\mathtt{AISE}_{\mu^{(v)}}|\mathbf{X}, \mathbf{W}, \widehat\Delta]=&\;\frac{1}{n}\operatorname*{trace}\Big(\mathbf{Q}_0^{-1}\boldsymbol{\Sigma}_0\mathbf{Q}_0^{-1}\int_{\mathcal{X}}\mathbf{b}_{p,s}^{(v)}(x)\mathbf{b}_{p,s}^{(v)}(x)'\omega(x)dx\Big)+o_\mathbb{P}(J^{2v+1}/n)\\ &+\int_{\mathcal{X}}\Big(\widehat{\mathbf{b}}_{p,s}^{(v)}(x)'\widehat{\boldsymbol{\beta}}_0-\mu_0^{(v)}(x)\Big)^2\omega(x)dx \\ &+\int_{\mathcal{X}}\Big(\widehat{\mathbf{b}}_{p,s}^{(v)}(x)'\mathbf{Q}_0^{-1}\mathbb{E}_n[\widehat{\mathbf{b}}_{p,s}(x_i)\eta_{i,1} \Psi(x_i, \mathbf{w}_i;\check{\eta}_i)]\Big)^2\omega(x)dx\\ &+2\int_{\mathcal{X}}\widehat{\mathbf{b}}_{p,s}^{(v)}(x)'\mathbf{Q}_0^{-1}\mathbb{E}_n[\widehat{\mathbf{b}}_{p,s}(x_i)\eta_{i,1} \Psi(x_i, \mathbf{w}_i;\check{\eta}_i)]\widehat{r}_{0,v}(x)\omega(x)dx. \end{align*} By Assumption (ref), $\Psi(x_i,\mathbf{w}_i;\check{\eta}_i)=-\Psi_1(x_i,\mathbf{w}_i;\eta_{i,0})\eta_{i,1}\widehat{r}_{0}(x_i)+O_\mathbb{P}(J^{-2p-2})$ where $O_\mathbb{P}(\cdot)$ holds uniformly over $i$. The terms in the last three lines correspond to the integrated squared bias. Also, using the same argument in the proof of Lemma (ref), $\mathbb{E}_n[\cdot]$ in the last two lines can be safely replaced by $\mathbb{E}_{\widehat\Delta}[\cdot]$, which only introduces some additional approximation error of order $o_\mathbb{P}(J^{-2p-2+2v})$. The proof of Theorem SA-3.4 in Cattaneo-Crump-Farrell-Feng_2024_AER shows that \begin{align*} \widehat{r}_{0,v}(x)=\;&\mu_0^{(v)}(x)-\widehat{\mathbf{b}}_{p,s}^{(v)}(x)'\widehat{\boldsymbol{\beta}}_0 \\ =&\frac{J^{-p-1+v}\mu_0^{(p+1)}(x)}{(p+1-v)!f_X(x)^{p+1-v}}\mathscr{E}_{p+1-v} \Big(\frac{x-\hat{\tau}_x^{\mathtt{L}}}{\hat{h}_x}\Big)\\ &-J^{-p-1}\widehat{\mathbf{b}}_{p,s}^{(v)}(x)'\mathbf{Q}_0^{-1}\mathbf{T}_s \mathbb{E}_{\widehat{\Delta}}\bigg[\widehat{\mathbf{b}}_{p,0}(x_i)\frac{\mu_0^{(p+1)}(x_i)}{(p+1)!f_X(x_i)^{p+1}}\mathscr{E}_{p+1} \Big(\frac{x_i-\hat{\tau}_{x_i}^{\mathtt{L}}}{\hat{h}_{x_i}}\Big)\bigg] +o_\mathbb{P}(J^{-p-1+v}), \end{align*} where $\hat\tau_{x}^{\mathtt{L}}$ is the start of the (random) interval in $\widehat\Delta$ containing $x$ and $\hat{h}_{x}$ denotes its length. Then, using the same argument as in the proof of Theorem SA-3.4 in Cattaneo-Crump-Farrell-Feng_2024_AER, we can approximate the integrated squared bias by the analogue based on the non-random partition $\Delta_0$, i.e., $\int_{\mathcal{X}}(r_{0,v}^\dagger(x)-\mathbf{b}_{p,s}^{(v)}(x)'\mathbf{Q}_0^{-1}\mathbb{E}[\mathbf{b}_{p,s}(x_i)\varkappa(x_i,\mathbf{w}_i)r_{0,0}^\dagger(x_i)])^2\omega(x)dx$ where \begin{align*} r_{0,v}^\dagger(x)=& \frac{J^{-p-1+v}\mu_0^{(p+1)}(x)}{(p+1-v)!f_X(x)^{p+1-v}}\mathscr{E}_{p+1-v} \Big(\frac{x-\tau_x^{\mathtt{L}}}{h_x}\Big)\\ &-J^{-p-1}\widehat{\mathbf{b}}_{p,s}^{(v)}(x)'\mathbf{Q}_0^{-1}\mathbf{T}_s \mathbb{E}\bigg[\mathbf{b}_{p,0}(x_i)\frac{\mu_0^{(p+1)}(x_i)}{(p+1)!f_X(x_i)^{p+1}}\mathscr{E}_{p+1} \Big(\frac{x_i-\tau_{x_i}^{\mathtt{L}}}{h_{x_i}}\Big)\bigg]. \end{align*} The expression of the bias term can be further simplified. For both $R_v(x)=r_{0,v}^\dagger(x)$ and $R_v(x)=r_{0,v}^\star(x)$, there exists some vector $\boldsymbol{\beta}$ such that $\sup_{x\in\mathcal{X}}|\mu_0(x)-\mathbf{b}_{p,s}(x_i)'\boldsymbol{\beta}-R_v(x)|=o(J^{-p-1+v})$ (see Lemma (ref) and Lemma SA-6.1 of Cattaneo-Farrell-Feng_2020_AoS). Define \[ r_{0,v}^{\tt P}(x)=\mu_0^{(v)}(x)-\mathbf{b}_{p,s}^{(v)}(x)'\mathbf{Q}_0^{-1}\mathbb{E}[\mathbf{b}_{p,s}(x_i)\varkappa(x_i,\mathbf{w}_i)\mu_0(x_i)]. \] Then, it follows that $r_{0,v}^{\tt P}(x)=R_v(x)-\mathbf{b}_{p,s}(x)'\mathbf{Q}_0^{-1}\mathbb{E}[\mathbf{b}_{p,s}(x_i)\varkappa(x_i,\mathbf{w}_i)R_0(x_i)]+o(J^{-p-1+v})$. Thus, \begin{align*} &\{r_{0,v}^\dagger(x)-\mathbf{b}_{p,s}^{(v)}(x)'\mathbf{Q}_0^{-1}\mathbb{E}[\mathbf{b}_{p,s}(x_i)\varkappa(x_i,\mathbf{w}_i)r_{0,0}^\dagger(x_i)]\}\\ -&\{[r_{0,v}^\star(x)-\mathbf{b}_{p,s}^{(v)}(x)'\mathbf{Q}_0^{-1}\mathbb{E}[\mathbf{b}_{p,s}(x_i)\varkappa(x_i,\mathbf{w}_i)r_{0,0}^\star(x_i)]]\}=o(J^{-p-1+v}). \end{align*} Therefore, the expression of $\mathscr{B}_n(p,s,v)$ given in the theorem holds. Finally, the desired results in part (ii) and part (iii) follow by Theorem (ref), the rate condition imposed and the same argument for part (i).

Proof of Theorem (ref)

proofThe proof is divided into several steps. Step 1: Note that \begin{align*} &\sup_{x\in\mathcal{X}}\bigg|\frac{\widehat{\mu}^{(v)}(x)-\mu_0^{(v)}(x)}{\sqrt{\widehat{\Omega}_{\mu^{(v)}}(x)/n}}-\frac{\widehat{\mu}^{(v)}(x)-\mu_0^{(v)}(x)}{\sqrt{\bar\Omega_{\mu^{(v)}}(x)/n}}\bigg|\\ \leq& \sup_{x\in\mathcal{X}}\bigg|\frac{\widehat{\mu}^{(v)}(x)-\mu_0^{(v)}(x)}{\sqrt{\bar\Omega_{\mu^{(v)}}(x)/n}}\bigg| \sup_{x\in\mathcal{X}}\bigg|\frac{\widehat{\Omega}_{\mu^{(v)}}(x)^{1/2}-\bar\Omega_{\mu^{(v)}}(x)^{1/2}}{\widehat{\Omega}_{\mu^{(v)}}(x)^{1/2}}\bigg|\\ \lesssim&_\mathbb{P} \Big(\sqrt{\log n}+\sqrt{n}J^{-p-1-1/2}\Big)\Big(J^{-p-1}+\sqrt{\frac{J\log n}{n^{1-\frac{2}{\nu}}}}\Big) \end{align*} where the last step uses Lemma (ref) and Corollary (ref). Then, in view of Lemmas (ref), (ref), Theorems (ref), (ref) and the rate restriction given in the lemma, we have \[ \sup_{x\in\mathcal{X}}\bigg|\frac{\widehat{\mu}^{(v)}(x)-\mu_0^{(v)}(x)} {\sqrt{\widehat{\Omega}_{\mu^{(v)}}(x)/n}} + \frac{\widehat{\mathbf{b}}_{p,s}^{(v)}(x)'\bar{\mathbf{Q}}^{-1}}{\sqrt{\bar\Omega_{\mu^{(v)}}(x)}}\mathbb{G}_n[\widehat{\mathbf{b}}_{p,s}(x_i)\eta_{i,1}\psi(y_i,\eta_i)]\bigg|=o_\mathbb{P}(a_n^{-1}). \] Step 2: Let us write $\mathscr{K}(x,x_i)=\Omega_{\mu^{(v)}}(x)^{-1/2}\widehat{\mathbf{b}}_{p,s}^{(v)}(x)'\bar{\mathbf{Q}}^{-1}\widehat{\mathbf{b}}_{p,s}(x_i)$ (the dependence of $\widehat{\mathbf{b}}_{p,s}^{(v)}(x)$, $\bar\mathbf{Q}$ and $\bar\Omega_{\mu^{(v)}}(x)$ on $\mathbf{X}$, $\mathbf{W}$ and $\widehat\Delta$ is omitted for simplicity), and $\tilde\sigma^2(x_i,\mathbf{w}_i)=\mathbb{E}[\psi^\dagger(\epsilon_i)^2|x_i,\mathbf{w}_i]$. Now we rearrange $\{x_i\}_{i=1}^n$ as a sequence of order statistics $\{x_{(i)}\}_{i=1}^n$, i.e., $x_{(1)}\leq\cdots\leq x_{(n)}$. Accordingly, $\{\epsilon_i\}_{i=1}^n$, $\{\mathbf{w}_i\}_{i=1}^n$ and $\{\tilde\sigma^2(x_i, \mathbf{w}_i)\}_{i=1}^n$ are ordered as concomitants $\{\epsilon_{[i]}\}_{i=1}^n$, $\{\mathbf{w}_{[i]}\}$ and $\{\tilde{\sigma}^2_{[i]}\}_{i=1}^n$ where $\tilde\sigma^2_{[i]}=\tilde\sigma^2(x_{(i)},\mathbf{w}_{[i]})$. Clearly, conditional on $\mathscr{F}_{XW\Delta}$ (the $\sigma$-field generated by $\{(x_i,\mathbf{w}_i)\}$ and $\widehat\Delta$), $\{\psi^\dagger(\epsilon_{[i]})\}_{i=1}^n$ is still an independent mean-zero sequence. Then by Assumptions (ref), (ref) and the result of Sakhanenko_1991_SAM, there exists a sequence of i.i.d. standard normal random variables $\{\zeta_{[i]}\}_{i=1}^n$ such that \begin{align*} \max_{1\leq \ell\leq n}|S_{\ell}|:= \max_{1\leq \ell \leq n}&\bigg|\sum_{i=1}^\ell\eta^{(1)}(\mu_0(x_{(i)})+\mathbf{w}_{[i]}'\boldsymbol{\gamma}_0)\psi^\ddagger(\eta(\mu_0(x_{(i)})+\mathbf{w}_{[i]}'\boldsymbol{\gamma}_0))\psi^\dagger(\epsilon_{[i]}) \\ &- \sum_{i=1}^\ell\eta^{(1)}(\mu_0(x_{(i)})+\mathbf{w}_{[i]}'\boldsymbol{\gamma}_0)\psi^\ddagger(\eta(\mu_0(x_{(i)})+\mathbf{w}_{[i]}'\boldsymbol{\gamma}_0)) \tilde\sigma_{[i]}\zeta_{[i]}\Big|\lesssim_\mathbb{P} n^{\frac{1}{\nu}}. \end{align*} Then, using summation by parts, \begin{align*} &\,\sup_{x\in\mathcal{X}}\left|\sum_{i=1}^{n}\mathscr{K}(x, x_{(i)})\eta^{(1)}(\mu_0(x_{(i)})+\mathbf{w}_{[i]}'\boldsymbol{\gamma}_0)\psi^\ddagger(\eta(\mu_0(x_{(i)})+\mathbf{w}_{[i]}'\boldsymbol{\gamma}_0))(\psi^\dagger(\epsilon_{[i]})-\tilde\sigma_{[i]}\zeta_{[i]})\right|\\ =&\,\sup_{x\in\mathcal{X}}\left|\mathscr{K}(x, x_{(n)})S_{n} -\sum_{i=1}^{n-1}S_{i}\left(\mathscr{K}(x,x_{(i+1)}) -\mathscr{K}(x,x_{(i)})\right)\right|\\ \leq&\,\sup_{x\in\mathcal{X}}\max_{1\leq i\leq n}|\mathscr{K}(x,x_i)||S_{n}| +\sup_{x\in\mathcal{X}}\left| \frac{\widehat{\mathbf{b}}_{p,s}^{(v)}(x)'\bar{\mathbf{Q}}^{-1}}{\sqrt{\bar\Omega_{\mu^{(v)}}(x)}} \sum_{i=1}^{n-1}S_{i}\Big(\widehat{\mathbf{b}}_{p,s}(x_{(i+1)}) - \widehat{\mathbf{b}}_{p,s}(x_{(i)})\Big)\right|\\ \leq&\,\sup_{x\in\mathcal{X}}\max_{1\leq i\leq n}|\mathscr{K}(x,x_i)||S_{n}| +\sup_{x\in\mathcal{X}} \Bigg\|\frac{\bar{\mathbf{Q}}^{-1}\widehat{\mathbf{b}}_{p,s}^{(v)}(x)}{\sqrt{\bar\Omega_{\mu^{(v)}}(x)}}\Bigg\|_1 \left\|\sum_{i=1}^{n-1}S_{i}\Big(\widehat{\mathbf{b}}_{p,s}(x_{(i+1)})- \widehat{\mathbf{b}}_{p,s}(x_{(i)})\Big)\right\|_\infty. \end{align*} By Lemmas (ref), (ref) and (ref), $\sup_{x\in\mathcal{X}}\sup_{x_i\in\mathcal{X}}|\mathscr{K}(x,x_i)|\lesssim_\mathbb{P} \sqrt{J}$, and $$\sup_{x\in\mathcal{X}}\Bigg\|\frac{\bar{\mathbf{Q}}^{-1}\widehat{\mathbf{b}}_{p,s}^{(v)}(x)}{\sqrt{\bar\Omega_{\mu^{(v)}}(x)}}\Bigg\|_1\lesssim_\mathbb{P} 1.$$ Then, notice that \begin{equation*} \max_{1\leq l \leq K_{p,s}} \bigg|\sum_{i=1}^{n-1}\Big(\widehat{b}_{p,s,l}(x_{(i+1)})-\widehat{b}_{p,s,l}(x_{(i)})\Big)S_{l}\bigg| \leq \max_{1\leq l \leq K_{p,s}} \sum_{i=1}^{n-1}\Big|\widehat{b}_{p,s,l}(x_{(i+1)})-\widehat{b}_{p,s,l}(x_{(i)})\Big| \max_{1\leq \ell \leq n}\Big|S_{\ell}\Big|. \end{equation*} By construction of the ordering, $\max_{1\leq l \leq K_{p,s}} \sum_{i=1}^{n-1}\Big|\widehat{b}_{p,s,l}(x_{(i+1)})-\widehat{b}_{p,s,l}(x_{(i)})\Big|\lesssim \sqrt{J}$. Under the rate restriction in the theorem, this suffices to show that for any $\xi>0$, $$ \mathbb{P}\bigg(\sup_{x\in\mathcal{X}}\, \Big|\mathbb{G}_n[\mathscr{K}(x,x_i)\eta^{(1)}(\mu_0(x_i)+\mathbf{w}_i'\boldsymbol{\gamma}_0)(\psi(y_i,\eta_i)-\sigma(x_i,\mathbf{w}_i)\zeta_i)]\Big| >\xi a_n^{-1}\Big|\mathscr{F}_{XW\Delta}\bigg)=o_\mathbb{P}(1), $$ where we recover the original ordering. Since $\mathbb{G}_n[\widehat{\mathbf{b}}_{p,s}(x_i)\zeta_i\sigma(x_i,\mathbf{w}_i)\eta_{i,1}]=_{d|\mathscr{F}_{XW\Delta}}\mathbf{N}(0, \bar{\boldsymbol{\Sigma}})$ ($=_{d|\mathscr{F}_{XW}}$ denotes “equal in distribution conditional on $\mathscr{F}_{XW\Delta}$”), the above steps construct the following approximating process: \[ \bar{Z}_{\mu^{(v)},p}(x):=\frac{\widehat{\mathbf{b}}_{p,s}^{(v)}(x)'\bar{\mathbf{Q}}^{-1}}{\sqrt{\bar\Omega_{\mu^{(v)}}(x)}}\bar{\boldsymbol{\Sigma}}^{1/2}\mathbf{N}_{K_{p,s}}. \] Step 3: Suppose that Assumption (ref)(ii) also holds. Note that \[ \begin{split} &\sup_{x\in\mathcal{X}}|\bar{Z}_{\mu^{(v)},p}(x)-Z_{\mu^{(v)},p}(x)|\\ \leq& \sup_{x\in\mathcal{X}}\bigg|\frac{\widehat{\mathbf{b}}^{(v)}(x)'(\bar{\mathbf{Q}}^{-1}-\mathbf{Q}_0^{-1})}{\sqrt{\Omega_{\mu^{(v)}}(x)}}\bar{\boldsymbol{\Sigma}}^{1/2}\mathbf{N}_{K_{p,s}}\bigg| + \sup_{x\in\mathcal{X}}\bigg|\frac{\widehat{\mathbf{b}}^{(v)}(x)'\mathbf{Q}_0^{-1}}{\sqrt{\Omega_{\mu^{(v)}}(x)}}\Big(\bar{\boldsymbol{\Sigma}}^{1/2}-\boldsymbol{\Sigma}_0^{1/2}\Big)\mathbf{N}_{K_{p,s}}\bigg|+\\ &\sup_{x\in\mathcal{X}}\bigg|\frac{\widehat{\mathbf{b}}_{p,0}^{(v)}(x)' (\widehat{\mathbf{T}}_s-\mathbf{T}_s)\mathbf{Q}_0^{-1}}{\sqrt{\Omega_{\mu^{(v)}}(x)}}\boldsymbol{\Sigma}_0^{1/2}\mathbf{N}_{K_{p,s}}\bigg| +\sup_{x\in\mathcal{X}}\bigg|\Big( \frac{1}{\sqrt{\bar\Omega_{\mu^{(v)}}(x)}}- \frac{1}{\sqrt{\Omega_{\mu^{(v)}}(x)}} \Big) \widehat{\mathbf{b}}_{p,0}^{(v)}(x)' \widehat{\mathbf{T}}_s\bar\mathbf{Q}^{-1}\bar\boldsymbol{\Sigma}^{1/2}\mathbf{N}_{K_{p,s}}\bigg|\\ =&\; I+II+III+IV, \end{split} \] where each term on the right-hand side is a mean-zero Gaussian process conditional on $\mathscr{F}_{XW\Delta}$. By Theorem (ref) (see Step 4 of its proof), $\sup_{x\in\mathcal{X}}|\bar\Omega_{\mu^{(v)}}(x)-\Omega_{\mu^{(v)}}(x)|\lesssim_\mathbb{P} J^{1+2v}(\sqrt{J\log n/n}+\mathfrak{r}_{\tt RP})$. By a similar calculation given in Step 1 and the rate condition imposed, the last term is $o_\mathbb{P}(a_n^{-1})$. By Lemmas (ref) and (ref), $\|\bar{\mathbf{Q}}^{-1}-\mathbf{Q}_0^{-1}\|\lesssim_\mathbb{P}\sqrt{J\log J/n}$ and $\|\widehat{\mathbf{T}}_s-\mathbf{T}_s\|\lesssim_\mathbb{P}\sqrt{J\log J/n}$. Also, using the argument in the proof of Lemma (ref) and Theorem X.3.8 of \citet*{Bhatia_2013_book}, $\|\bar{\boldsymbol{\Sigma}}^{1/2}-\boldsymbol{\Sigma}_0^{1/2}\|\lesssim_\mathbb{P}\sqrt{J\log J/n}$. By Gaussian Maximal Inequality \citep*[Corollary 2.2.8]{vandevarrt-Wellner_1996_book}, \[\mathbb{E}\Big[I+II+III\Big|\mathscr{F}_{XW\Delta} \Big]\lesssim_\mathbb{P} \sqrt{\log J} \Big(\|\bar{\boldsymbol{\Sigma}}^{1/2}-\boldsymbol{\Sigma}_0^{1/2}\|+\|\bar{\mathbf{Q}}^{-1}-\mathbf{Q}_0^{-1}\|+\|\widehat{\mathbf{T}}_s-\mathbf{T}_s\|\Big)=o_\mathbb{P}(a_n^{-1}) \] where the last line follows from the imposed rate restriction. Then the proof for part (i) is complete. The results in parts (ii) and (iii) immediately follow by Theorem (ref) and the fact that the leading variance term in the Bahadur representation for $\widehat\vartheta(x,\widehat\mathsf{w})$ or $\widehat\zeta(x,\widehat\mathsf{w})$ differs from that for $\widehat\mu(x)$ or $\widehat\mu^{(1)}(x)$ up to a sign only.

Proof of Theorem (ref)

proofThis conclusion follows from Lemmas (ref), (ref), Theorem (ref) and Gaussian Maximal Inequality as applied in Step 3 in the proof of Theorem (ref).

Proof of Theorem (ref)

proofWe first show that \[ \sup_{u\in\mathbb{R}}\Big|\mathbb{P}\Big(\sup_{x\in\mathcal{X}} |T_{\mu^{(v)},p}(x)|\leq u\Big) - \mathbb{P}\Big(\sup_{x\in\mathcal{X}}|Z_{\mu^{(v)},p}(x)|\leq u\Big)\Big|=o(1). \] By Theorem (ref), there exists a sequence of constants $\xi_n$ such that $\xi_n=o(1)$ and \[ \mathbb{P}\Big(\Big|\sup_{x\in\mathcal{X}}|T_{\mu^{(v)},p}(x)|-\sup_{x\in\mathcal{X}}|Z_{\mu^{(v)},p}(x)|\Big|>\xi_n/a_n\Big)=o(1). \] Then, \begin{align*} \mathbb{P}\Big(\sup_{x\in\mathcal{X}}|T_{\mu^{(v)},p}(x)|\leq u\Big)&\leq \mathbb{P}\Big(\Big\{\sup_{x\in\mathcal{X}}|T_{\mu^{(v)},p}(x)|\leq u\Big\}\cap\Big\{\Big|\sup_{x\in\mathcal{X}}|T_{\mu^{(v)},p}(x)|- \sup_{x\in\mathcal{X}}|Z_{\mu^{(v)},p}(x)|\Big|\leq\xi_n/a_n\Big\}\Big)+o(1)\\ &\leq \mathbb{P}\Big(\sup_{x\in\mathcal{X}}|Z_{\mu^{(v)},p}(x)|\leq u+\xi_n/a_n\Big) +o(1)\\ &\leq \mathbb{P}\Big(\sup_{x\in\mathcal{X}}|Z_{\mu^{(v)},p}(x)|\leq u\Big)+ \sup_{u\in\mathbb{R}}\mathbb{E}\Big[\mathbb{P}\Big(\Big|\sup_{x\in\mathcal{X}}|Z_{\mu^{(v)},p}(x)|-u\Big|\leq \xi_n/a_n\Big|\widehat{\Delta}\Big)\Big]\\ &\leq \mathbb{P}\Big(\sup_{x\in\mathcal{X}}|Z_{\mu^{(v)},p}(x)|\leq u\Big)+ \mathbb{E}\Big[\sup_{u\in\mathbb{R}}\mathbb{P}\Big(\Big|\sup_{x\in\mathcal{X}} |Z_{\mu^{(v)},p}(x)|-u\Big|\leq \xi_n/a_n\Big|\widehat{\Delta}\Big)\Big] +o(1). \end{align*} Apply the Anti-Concentration Inequality conditional on $\widehat\Delta$ Chernozhukov-Chetverikov-Kato_2014b_AoS to the second term: \begin{align*} \sup_{u\in\mathbb{R}}\; \mathbb{P}\Big(\Big|\sup_{x\in\mathcal{X}}|Z_{\mu^{(v)},p}(x)|-u\Big|\leq \xi_n/a_n\Big|\widehat{\Delta}\Big)&\leq 4\xi_na_n^{-1}\mathbb{E}\Big[\sup_{x\in\mathcal{X}}|Z_{\mu^{(v)},p}(x)|\Big|\widehat\Delta\Big]+o(1)\\ &\lesssim_\mathbb{P} \xi_n a_n^{-1}\sqrt{\log J}+o(1)\to 0 \end{align*} where the last step uses Gaussian Maximal Inequality \citep*[see][Corollary 2.2.8]{vandevarrt-Wellner_1996_book}. By Dominated Convergence Theorem, \[\mathbb{E}\Big[ \sup_{u\in\mathbb{R}}\mathbb{P}\Big(\Big|\sup_{x\in\mathcal{X}}|Z_{\mu^{(v)},p}(x)|-u\Big|\leq \xi_n/a_n\Big|\widehat{\Delta}\Big)\Big]=o(1). \] The other side of the inequality follows similarly. By similar argument, using Theorem (ref), we have \[ \sup_{u\in\mathbb{R}}\Big|\mathbb{P}\Big(\sup_{x\in\mathcal{X}}|\widehat{Z}_{\mu^{(v)},p}(x)|\leq u\Big|\mathbf{D},\widehat{\Delta}\Big) - \mathbb{P}\Big(\sup_{x\in\mathcal{X}}|Z_{\mu^{(v)},p}(x)|\leq u\Big|\widehat{\Delta}\Big)\Big|=o_\mathbb{P}(1). \] Then, it remains to show that \begin{equation} \sup_{u\in\mathbb{R}}\Big|\mathbb{P}\Big(\sup_{x\in\mathcal{X}}|Z_{\mu^{(v)},p}(x)|\leq u\Big)- \mathbb{P}\Big(\sup_{x\in\mathcal{X}}|Z_{\mu^{(v)},p}(x)|\leq u|\widehat{\Delta}\Big)\Big|=o_\mathbb{P}(1). \end{equation} We can write \[ Z_{\mu^{(v)},p}(x)=\frac{\widehat{\mathbf{b}}_{p,0}^{(v)}(x)'}{\sqrt{\widehat{\mathbf{b}}_{p,0}^{(v)}(x)'\mathbf{V}_0\widehat{\mathbf{b}}_{p,0}^{(v)}(x)}}\breve{\mathbf{N}}_{K_{p,0}} \] where $\mathbf{V}_0=\mathbf{T}_s'\mathbf{Q}_0^{-1}\boldsymbol{\Sigma}_0\mathbf{Q}_0^{-1}\mathbf{T}_s$ and $\breve{\mathbf{N}}_{K_{p,0}}:=\mathbf{T}_s'\mathbf{Q}_0^{-1}\boldsymbol{\Sigma}_0^{1/2}\mathbf{N}_{K_{p,s}}$ is a $K_{p,0}$-dimensional Gaussian random vector. Importantly, by this construction, $\breve{\mathbf{N}}_{K_{p,0}}$ and $\mathbf{V}_0$ do not depend on $\widehat{\Delta}$ and $x$, and they are only determined by the deterministic partition $\Delta_0$. First consider $v=0$. For any two partitions $\Delta_1, \Delta_2\in\Pi$, for any $x\in\mathcal{X}$, there exists $\check{x}\in\mathcal{X}$ such that \[\mathbf{b}_{p,0}^{(0)}(x;\Delta_1)=\mathbf{b}_{p,0}^{(0)}(\check{x};\Delta_2),\] and vice versa. Therefore, the following two events are equivalent: $\{\omega:\sup_{x\in\mathcal{X}}|Z_p(x;\Delta_1)|\leq u\}= \{\omega:\sup_{x\in\mathcal{X}}|Z_p(x;\Delta_2)|\leq u\}$ for any $u$. Thus, $$\mathbb{E}\Big[\mathbb{P}\Big(\sup_{x\in\mathcal{X}}|Z_{\mu^{(v)},p}(x)|\leq u \Big|\widehat\Delta\Big)\Big]=\mathbb{P}\Big(\sup_{x\in\mathcal{X}}|Z_{\mu^{(v)},p}(x)|\leq u \Big|\widehat\Delta\Big)+o_\mathbb{P}(1).$$ Then for $v=0$, the desired result follows. For $v>0$, simply notice that $\widehat{\mathbf{b}}_{p,0}^{(v)}(x)=\widehat{\mathfrak{T}}_v\widehat{\mathbf{b}}_{p,0}(x)$ for some transformation matrix $\widehat{\mathfrak{T}}_v$. Clearly, $\widehat{\mathfrak{T}}_v$ takes a similar structure as $\widehat{\mathbf{T}}_s$: each row and each column only have a finite number of nonzeros. Each nonzero element is simply $\hat{h}_j^{-v}$ up to some constants. By Lemma (ref), it can be shown that $\|\widehat{\mathfrak{T}}_v-\mathfrak{T}_v\|\lesssim J^v\sqrt{J\log J/n}$ where $\mathfrak{T}_v$ is the population analogue ($\hat{h}_j$ replaced by $h_j$). Repeating the argument given in the proof of Theorems (ref) and (ref), we can replace $\widehat{\mathfrak{T}}_v$ in $Z_{\mu^{(v)},p}(x)$ by $\mathfrak{T}_v$ without affecting the approximation rate. Then the desired result for $T_{\mu^{(v)},p}(x)$ follows by repeating the argument given for $v=0$ above. Finally, the result for $T_{\vartheta,p}(x)$ ($T_{\zeta,p}(x)$) follows by the fact that $Z_{\vartheta,p}(x)$ and $\widehat{Z}_{\vartheta,p}(x)$ ($Z_{\zeta,p}(x)$ and $\widehat{Z}_{\zeta,p}(x)$) differ from $Z_{\mu^{(v)},p}(x)$ and $\widehat{Z}_{\mu^{(v)},p}(x)$ up to a sign only.

Proof of Theorem (ref)

proofWe only consider $\widehat{I}_{\mu^{(v)},p}(x)$. The results in part (ii) and part (iii) follow similarly. Let $\xi_{1,n}=o(1)$, $\xi_{2,n}=o(1)$ and $\xi_{3,n}=o(1)$. Then, \begin{align*} \mathbb{P}\left[\sup_{x\in\mathcal{X}}|T_{\mu^{(v)},p}(x)|\leq \mathfrak{c}_{\mu^{(v)}} \right]&\leq \mathbb{P}\left[\sup_{x\in\mathcal{X}}|\bar{Z}_{\mu^{(v)},p}(x)|\leq \mathfrak{c}_{\mu^{(v)},p}+\xi_{1,n}/a_n \right]+o(1)\\ &\leq \mathbb{P}\left[\sup_{x\in\mathcal{X}}|\bar{Z}_{\mu^{(v)},p}(x)|\leq c^0(1-\alpha+\xi_{3,n})+(\xi_{1,n}+\xi_{2,n})/a_n \right]+o(1)\\ &\leq \mathbb{P}\left[\sup_{x\in\mathcal{X}}|\bar{Z}_{\mu^{(v)},p}(x)|\leq c^0(1-\alpha+\xi_{3,n})\right]+o(1)\to 1-\alpha, \end{align*} where $c^0(1-\alpha+\xi_{3,n})$ denotes the $(1-\alpha+\xi_{3,n})$-quantile of $\sup_{x\in\mathcal{X}}|\bar{Z}_{\mu^{(v)},p}(x)|$ conditional on $\mathscr{F}_{XW\Delta}$ (the $\sigma$-field generated by $\mathbf{X}$, $\mathbf{W}$ and the partition $\widehat\Delta$), the first inequality holds by Theorem (ref), the second by Lemma A.1 of Belloni-Chernozhukov-Chetverikov-Kato_2015_JoE, and the third by Anti-Concentration Inequality in Chernozhukov-Chetverikov-Kato_2014b_AoS. The other side of the bound follows similarly.

Proof of Theorem (ref)

proofWe only consider the proof for part (i). The results in part (ii) and part (iii) follow similarly. Throughout this proof, we let $\xi_{1,n}=o(1)$, $\xi_{2,n}=o(1)$ and $\xi_{3,n}=o(1)$ be sequences of vanishing constants. Moreover, let $A_n$ be a sequence of diverging constants such that $\sqrt{\log J}A_n\lesssim \sqrt{\frac{n}{J^{1+2v}}}$. Note that \[ \sup_{x\in\mathcal{X}} |\dot{T}_{\mu^{(v)},p}(x)|\leq \sup_{x\in\mathcal{X}} \bigg|\frac{\widehat{\mu}(x)- \mu_0^{(v)}(x)} {\sqrt{\widehat\Omega_{\mu^{(v)}}(x)/n}}\bigg|+ \sup_{x\in\mathcal{X}}\bigg| \frac{\mu_0^{(v)}(x)- m^{(v)}(x;\widetilde{\boldsymbol{\theta}})} {\sqrt{\widehat{\Omega}_{\mu^{(v)}}(x)/n}}\bigg|. \] Therefore, under $\dot{\mathsf{H}}_0^{\mu^{(v)}}$, \begin{align*} \mathbb{P}\Big[\sup_{x\in\mathcal{X}}|\dot{T}_{\mu^{(v)},p}(x)|>\mathfrak{c}_{\mu^{(v)}}\Big]& \leq \mathbb{P}\bigg[\sup_{x\in\mathcal{X}}|T_{\mu^{(v)},p}(x)|>\mathfrak{c}_{\mu^{(v)}}- \sup_{x\in\mathcal{X}}\bigg|\frac{\mu_0^{(v)}(x)-m^{(v)}(x;\widetilde{\boldsymbol{\theta}})}{\sqrt{\widehat{\Omega}_{\mu^{(v)}}(x)/n}}\bigg|\bigg] \\ &\leq \mathbb{P}\bigg[\sup_{x\in\mathcal{X}}|\bar{Z}_{\mu^{(v)},p}(x)|>\mathfrak{c}_{\mu^{(v)}}-\xi_{1,n}/a_n - \sup_{x\in\mathcal{X}}\bigg|\frac{\mu_0^{(v)}(x)-m^{(v)}(x;\widetilde{\boldsymbol{\theta}})}{\sqrt{\widehat{\Omega}_{\mu^{(v)}}(x)/n}}\bigg|\bigg] + o(1)\\ &\leq \mathbb{P}\bigg[\sup_{x\in\mathcal{X}}|\bar{Z}_{\mu^{(v)},p}(x)|>c^0(1-\alpha-\xi_{3,n})-(\xi_{1,n}+\xi_{2,n})/a_n - \\ &\qquad \sup_{x\in\mathcal{X}}\bigg|\frac{\mu_0^{(v)}(x)-m^{(v)}(x;\widetilde{\boldsymbol{\theta}})}{\sqrt{\widehat{\Omega}_{\mu^{(v)}}(x)/n}}\bigg|\bigg] + o(1)\\ &\leq \mathbb{P}\Big[\sup_{x\in\mathcal{X}}|\bar{Z}_{\mu^{(v)},p}(x)|>c^0(1-\alpha-\xi_{3,n})\Big]+o(1)\\ &=\alpha+o(1) \end{align*} where $c^0(1-\alpha-\xi_{3,n})$ denotes the $(1-\alpha-\xi_{3,n})$-quantile of $\sup_{x\in\mathcal{X}}|\bar{Z}_{\mu^{(v)},p}(x)|$ conditional on $\mathscr{F}_{XW\Delta}$ (the $\sigma$-field generated by $\mathbf{X}$, $\mathbf{W}$ and $\widehat\Delta$), the second inequality holds by Theorem (ref), the third by Lemma A.1 of Belloni-Chernozhukov-Chetverikov-Kato_2015_JoE, the fourth by the fact that $\sup_{x\in\mathcal{X}}\big|\frac{\mu_0^{(v)}(x)-m^{(v)}(x;\widetilde{\boldsymbol{\theta}})}{\sqrt{\widehat{\Omega}_{\mu^{(v)}}(x)/n}}\big|=o_\mathbb{P}(\frac{1}{\sqrt{\log J}})$ and Anti-Concentration Inequality in Chernozhukov-Chetverikov-Kato_2014b_AoS. The other side of the bound follows similarly. On the other hand, under $\dot{\mathsf{H}}_\text{A}^{\mu^{(v)}}$, \begin{align*} &\mathbb{P}\Big[\sup_{x\in\mathcal{X}}|\dot{T}_{\mu^{(v)},p}(x)|>\mathfrak{c}_{\mu^{(v)}}\Big] \\ =\,&\mathbb{P}\Big[\sup_{x\in\mathcal{X}}\Big|T_{\mu^{(v)},p}(x)+ \frac{\mu_0^{(v)}(x)-m^{(v)}(x;\bar{\boldsymbol{\theta}})}{\sqrt{\widehat{\Omega}_{\mu^{(v)}}(x)/n}}+ \frac{m^{(v)}(x;\bar{\boldsymbol{\theta}})-m^{(v)}(x;\widetilde{\boldsymbol{\theta}})}{\sqrt{\widehat{\Omega}_{\mu^{(v)}}(x)/n}}\Big|>\mathfrak{c}_{\mu^{(v)}}\Big]\\ \geq\,& \mathbb{P}\bigg[\sup_{x\in\mathcal{X}}|T_{\mu^{(v)},p}(x)|< \sup_{x\in\mathcal{X}}\bigg|\frac{\mu_0^{(v)}(x)-m^{(v)}(x;\bar{\boldsymbol{\theta}})} {\sqrt{\widehat{\Omega}_{\mu^{(v)}}(x)/n}} + \frac{m^{(v)}(x;\bar{\boldsymbol{\theta}})-m^{(v)}(x;\widetilde{\boldsymbol{\theta}})}{\sqrt{\widehat{\Omega}_{\mu^{(v)}}(x)/n}}\bigg| -\mathfrak{c}_{\mu^{(v)}}\bigg] \\ \geq\,& \mathbb{P}\Big[\sup_{x\in\mathcal{X}} |\bar{Z}_{\mu^{(v)},p}(x)|\leq \sqrt{\log J}A_n-\xi_{1,n}/a_n\Big]-o(1)\\ \geq\,& 1-o(1). \end{align*} where the fourth line holds by Lemma (ref), Theorem (ref), Theorem (ref), the condition that $J^v\sqrt{J\log J/n}=o(1)$ and the definition of $A_n$, and the last by the Talagrand-Samorodnitsky Concentration Inequality vandevarrt-Wellner_1996_book.

Proof of Theorem (ref)

proofWe only consider the proof for part (i). The results in part (ii) and part (iii) follow similarly. Throughout this proof, the definitions of $A_n$, $\xi_{1,n}, \xi_{2,n}$ and $\xi_{3,n}$ are the same as in the proof of Theorem (ref). Under $\ddot{\mathsf{H}}_0^{\mu^{(v)}}$, \[ \sup_{x\in\mathcal{X}}\ddot{T}_{\mu^{(v)},p}(x) \leq\sup_{x\in\mathcal{X}}T_{\mu^{(v)},p}(x)+ \sup_{x\in\mathcal{X}}\frac{|m^{(v)}(x;\bar{\boldsymbol{\theta}})- m^{(v)}(x;\widetilde{\boldsymbol{\theta}})|} {\sqrt{\widehat{\Omega}_{\mu^{(v)}}(x)/n}}. \] Then, \begin{align*} \mathbb{P}\Big[ \sup_{x\in\mathcal{X}} \ddot{T}_{\mu^{(v)},p}(x)> \mathfrak{c}_{\mu^{(v)}}\Big] &\leq \mathbb{P}\bigg[ \sup_{x\in\mathcal{X}} T_{\mu^{(v)},p}(x)> \mathfrak{c}_{\mu^{(v)}} - \sup_{x\in\mathcal{X}}\frac{|m^{(v)}(x; \bar{\boldsymbol{\theta}})- m^{(v)}(x;\widetilde{\boldsymbol{\theta}})|} {\sqrt{\widehat{\Omega}_{\mu^{(v)}}(x)/n}} \bigg] \\ &\leq \mathbb{P}\Big[ \sup_{x\in\mathcal{X}} \bar{Z}_{\mu^{(v)},p}(x)> \mathfrak{c}_{\mu^{(v)}}-\xi_{1,n}/a_n\Big] + o(1)\\ &\leq \mathbb{P}\Big[ \sup_{x\in\mathcal{X}} \bar{Z}_{\mu^{(v)},p}(x)> c^0(1-\alpha-\xi_{3,n})-(\xi_{1,n}+\xi_{2,n})/a_n\Big] + o(1)\\ &\leq \mathbb{P}\Big[ \sup_{x\in\mathcal{X}} \bar{Z}_{\mu^{(v)},p}(x)> c^0(1-\alpha-\xi_{3,n})\Big] + o(1)\\ &=\alpha+o(1) \end{align*} where $c^0(1-\alpha-\xi_{3,n})$ denotes the $(1-\alpha-\xi_{3,n})$-quantile of $\sup_{x\in\mathcal{X}}\bar{Z}_{\mu^{(v)},p}(x)$ conditional on $\mathscr{F}_{XW\Delta}$ (the $\sigma$-field generated by $\mathbf{X}$, $\mathbf{W}$ and $\widehat\Delta$), the second line holds by Theorem (ref), the third by Lemma A.1 of Belloni-Chernozhukov-Chetverikov-Kato_2015_JoE, the fourth by Anti-Concentration Inequality in Chernozhukov-Chetverikov-Kato_2014b_AoS. On the other hand, under $\ddot{\mathsf{H}}_\text{A}^{\mu^{(v)}}$, \begin{align*} \mathbb{P}\Big[ \sup_{x\in\mathcal{X}} \ddot{T}_{\mu^{(v)},p}(x)> \mathfrak{c}_{\mu^{(v)}}\Big] &=\mathbb{P}\bigg[ \sup_{x\in\mathcal{X}}\Big( T_{\mu^{(v)},p}(x)+ \frac{\mu_0^{(v)}(x)- m^{(v)}(x;\widetilde{\boldsymbol{\theta}})} {\sqrt{\widehat{\Omega}_{\mu^{(v)}}(x)/n}}-\mathfrak{c}_{\mu^{(v)}}\Big)>0\bigg]\\ &\geq \mathbb{P}\Big[ \sup_{x\in\mathcal{X}} |T_{\mu^{(v)},p}(x)|< \sup_{x\in\mathcal{X}}\frac{ \mu_0^{(v)}(x)- m^{(v)}(x;\widetilde{\boldsymbol{\theta}})} {\sqrt{\widehat{\Omega}_{\mu^{(v)}}(x)/n}}-\mathfrak{c}_{\mu^{(v)}}, \\ &\sup_{x\in\mathcal{X}}\frac{\mu_0^{(v)}(x)- m^{(v)}(x;\widetilde{\boldsymbol{\theta}})} {\sqrt{\widehat{\Omega}_{\mu^{(v)}}(x)/n}}>\mathfrak{c}_{\mu^{(v)}}\Big]\\ &\geq \mathbb{P}\Big[ \sup_{x\in\mathcal{X}} |T_{\mu^{(v)},p}(x)|< \sup_{x\in\mathcal{X}}\frac{\mu_0^{(v)}(x)- m^{(v)}(x;\widetilde{\boldsymbol{\theta}})}{\sqrt{\widehat{\Omega}_{\mu^{(v)}}(x)/n}} -\mathfrak{c}_{\mu^{(v)}}\Big]-o(1)\\ &\geq \mathbb{P}\Big[ \sup_{x\in\mathcal{X}} |T_{\mu^{(v)},p}(x)|< \sqrt{\log J}A_n\Big]-o(1)\\ &\geq \mathbb{P}\Big[ \sup_{x\in\mathcal{X}} |\bar{Z}_{\mu^{(v)},p}(x)|<\sqrt{\log J}A_n-\xi_{1,n}/a_n\Big] - o(1) \\ &\geq 1 - o(1) \end{align*} where the third line holds by Lemma (ref), Theorem (ref), Lemma A.1 of Belloni-Chernozhukov-Chetverikov-Kato_2015_JoE, the assumption that $\sup_{x\in\mathcal{X}}|m^{(v)}(x;\widetilde{\boldsymbol{\theta}})- m^{(v)}(x;\bar{\boldsymbol{\theta}})|=o_\mathbb{P}(1)$ and $J^v\sqrt{J\log J/n}=o(1)$, the fourth by definition of $A_n$, and the fifth by Theorem (ref), and the last by Proposition A.2.7 in vandevarrt-Wellner_1996_book.

References

\begingroup \endgroup