Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
53,858 characters · 10 sections · 59 citation commands
Higher-Order Refinements of Small Bandwidth Asymptotics for Density-Weighted Average Derivative Estimators
Keywords: density weighted average derivatives, Edgeworth expansions, small bandwidth asymptotics. \thispagestyle{empty}
\setcounter{page}{1} \pagestyle{plain}
Identification, estimation, and inference in the context of semiparametric models has a long tradition in econometrics Powell_1994_Handbook. Canonical two-step semiparametric estimands are finite dimensional functionals of some other unknown infinite dimensional parameters in the model (e.g., a density or regression function). A leading example of such a finite dimensional estimand is the density weighted average derivative (DWAD) of a regression function. This paper seeks to honor the many contributions of Jim Powell to semiparametric theory in econometrics by juxtaposing the higher-order distributional properties of Powell-Stock-Stoker_1989_ECMA's Powell-Stock-Stoker_1989_ECMA two-step kernel-based DWAD estimator under two alternative large sample approximation regimes: one based on the classical asymptotic linear representation, and the other based on a more general quadratic distributional approximation known as small bandwidth asymptotics.\footnote{Jim Powell's contributions to semiparametric theory are numerous. Honore-Powell_1994_JOE, Powell-Stoker_1996_JoE, Blundell-Powell_2004_RESTUD, AradillasLopez-Honore-Powell_2007_IER, Ahn-Ichimura-Powell-Ruud_2018_JBES, and Graham-Niu-Powell_2023_JOE are some of the most closely connected to the our work: these papers employ U-statistics methods for two-step kernel-based estimators similar to those considered herein. See Powell_2017_JEP for more discussion and references.}
In a landmark contribution, Powell-Stock-Stoker_1989_ECMA proposed a kernel-based DWAD estimator and obtained first-order, asymptotically linear distribution theory employing ideas from the U-statistics literature in statistics to develop valid inference procedures in large samples. This work sparked a wealth of subsequent developments in the econometrics literature: Robinson_1995_ECMA obtained Berry-Esseen bounds, Powell-Stoker_1996_JoE considered mean square error expansions, Nishiyama-Robinson_2000_ECMA,Nishiyama-Robinson_2001_ChBook,Nishiyama-Robinson_2005_ECMA developed Edgeworth expansions, and Newey-Hsieh-Robins_2004_Ecma investigated bias properties, just to mention a few contributions. The two-step semiparametric estimator in this literature employs a preliminary kernel-based estimator of a density function, which requires choosing two main tuning parameters: a bandwidth and a kernel function. The “optimal” choices for these tuning parameters depend on the goal of interest (e.g., point estimation vs. inference), as well as on the features of the underlying data generating process (e.g., smoothness of the unknown density and dimension of the covariates).
Classical first-order distribution theory for kernel-based DWAD estimators has focused on cases where tuning parameter restrictions and model assumptions imply an asymptotic linear representation of the two-step semiparametric point estimator Bickel-Klaassen-Ritov-Wellner_1993_Book,Newey-McFadden_1994_Handbook,Ichimura-Todd_2007_Handbook. That is, the two-step estimator is approximated by a sample average based on an influence function. This approach can be used to construct semiparametrically efficient inference procedures, but requires potentially high smoothness levels of the underlying unknown functions, thereby forcing the use of higher-order kernels or other debiasing techniques. Further, the implied distributional approximation may not be “robust” to tuning parameter choices and/or model features. More specifically, the limiting distribution emerging from the asymptotic linear representation of the centered and scaled point estimator is invariant to the way that the preliminary nonparametric estimators are constructed. At its core, an asymptotic linear approximation assumes away the contribution of additional terms forming the statistic of interest, despite the fact that these terms do contribute to the sampling variability of the two-step semiparametric estimator and, more importantly, do reflect the impact of tuning parameter choices in finite samples.
Cattaneo-Crump-Jansson_2014a_ET proposed an alternative distributional approximation for kernel-based DWAD estimators that allows for, but does not require, asymptotic linearity. The idea is to capture the joint contribution to the sampling distribution of both linear and quadratic terms forming the kernel-based DWAD estimator, because the quadratic term explicitly captures the effect of the choice of bandwidth and kernel function. To operationalize this idea, Cattaneo-Crump-Jansson_2014a_ET introduced an asymptotic experiment where the bandwidth sequence is allowed (but not required) to vanish at a speed that would render the classical asymptotic linear representation invalid because the quadratic term becomes first order even in large samples, which they termed small bandwidth asymptotics. This framework was carefully developed to obtain a distributional approximation that explicitly depends on both linear and quadratic terms, thereby forcing a more careful analysis of how the nonparametric first stage contributes to the sampling distribution of the statistic.
Inference methods based on small bandwidth asymptotics for kernel-based DWAD estimators were found to perform well in simulations Cattaneo-Crump-Jansson_2010_JASA,Cattaneo-Crump-Jansson_2014a_ET,Cattaneo-Crump-Jansson_2014b_ET, but no formal justification for this finite sample success is available in the literature. Methodologically, this alternative distributional approximation leads to a new way of conducting inference (e.g., constructing confidence interval estimators) because the original standard error formula proposed by Powell-Stock-Stoker_1989_ECMA must be modified to make the asymptotic approximation valid across the full range of allowable bandwidths (including the region where asymptotic linearity fails). Theoretically, however, the empirical success of small bandwidth asymptotics could come from two distinct sources: (i) it could deliver a better distributional approximation to the sampling distribution of the point estimator; or (ii) it could deliver a better distributional approximation to the sampling distribution of the studentized t-statistic because the standard error formula is modified.
Employing Edgeworth expansions Bhattacharya-Rao_1976_Book,Hall_1992_Book, this paper shows that the higher-order distributional properties of inference procedures motivated by the small bandwidth asymptotics approximation framework are demonstrably superior to those of procedures motivated by asymptotic linear approximations. We study both standardized and studentized estimators and show that those emerging from the small bandwidth regime offer higher-order corrections, as measured by the second cumulant underlying their Edgeworth expansions. An immediate implication of our results is that the small bandwidth asymptotic framework simultaneously enjoys two advantages: delivering a better distributional approximation (Theorem (ref), standardized t-statistic) and leading to a better standard error construction (Theorem (ref), studentized t-statistic). Therefore, our results have theoretical and practical implications for empirical work in economics, in addition to providing a theory-based explanation for prior simulation findings documenting better numerical performance of inference procedures motivated by small bandwidth asymptotics relative to those motivated by classical asymptotically linear distributional approximations.
The closest antecedent to our work is Nishiyama-Robinson_2000_ECMA,Nishiyama-Robinson_2001_ChBook, who also studied Edgeworth expansions for kernel-based DWAD estimators. Their expansions, however, were motivated by the asymptotic linear approximation of the point estimator, and hence cannot be used to compare and contrast to the distributional approximation emerging from the alternative small bandwidth asymptotic regime. Therefore, from a technical perspective, this paper also offers novel Edgeworth expansions that allow for different standardization and studentization schemes, thereby allowing us to plug-and-play when juxtaposing the two asymptotic approximation frameworks. More specifically, Theorem (ref) below concerns a generic standardized t-statistic and is proven based on Theorem (ref) in the appendix, which may be of independent technical interest due to is generality. In contrast, Theorem (ref) below concerns a more specialized class of studentized t-statistic because establishing valid Edgeworth expansions is considerably harder when dealing with studentization.
The idea of employing more general asymptotic approximation frameworks that do not enforce asymptotic linearity for two-step semiparametric estimators has also featured in other contexts: (i) semi-linear series-based, many covariates, and many instruments estimation Cattaneo-Jansson-Newey_2018_ET,Cattaneo-Jansson-Newey_2018_JASA, (ii) non-linear two-step semiparametric estimation Cattaneo-Crump-Jansson_2013_JASA,Cattaneo-Jansson_2018_ECMA,Cattaneo-Jansson-Ma_2019_RESTUD,Cattaneo-Jansson_2022_ET, and (iii) network estimation Matsushita-Otsu_2021_Biometrika. While our theoretical developments and results focus specifically on the case of kernel-based DWAD estimation, our main conceptual conclusions can be extrapolated to those settings as well. The main takeaway is that employing alternative asymptotic frameworks can deliver improved inference with smaller higher-order distributional approximation errors, thereby offering more robust inference procedures in finite samples. Furthermore, our theoretical and methodological results can also be leveraged to study the higher-order distributional properties of bootstrap-based methods for inference. Although a complete theoretical analysis is beyond the scope of this paper, we provide further discussion about the bootstrap in Section (ref).
The paper continues as follows. Section (ref) introduces the setup and main assumptions. Section (ref) reviews the classical first-order distributional approximation based on asymptotic linearity and the more general small bandwidth distributional approximation, along with their corresponding choices of standard error formulas. Section (ref) presents the main results of our paper. Section (ref) concludes. The appendix is organized in three parts: Appendix (ref) provides a self-contained generic Edgeworth expansion for second-order U-statistics, which may be of independent technical interest, Appendix (ref) gives the proof of Theorem (ref), and Appendix (ref) gives the proof of Theorem (ref).
Suppose $Z_i=(Y_i, X_i')'$, $i = 1,\dots,n$, is a random sample from the distribution of the random vector $Z=(Y, X')'$, where $Y$ is an outcome variable and $X$ takes values on $\mathbb{R}^d$ with Lebesgue density $f$. We consider \[\theta := \mathbb{E}[f(X)\dot{g}(X)], \qquad g(X) := \mathbb{E}[Y|X],\] the DWAD of the regression function $g$, where, for any (differentiable) function $a$, $\dot{a}(x)$ denotes $\partial a(x) / \partial x$, and where existence of $\theta$ is implied by parts (b) and (c)) of the following assumption, which collects the regularity conditions under which our subsequent analysis will proceed.
Under Assumption (ref) and using integration by parts, the DWAD vector can be expressed as
which motivates the celebrated plug-in analog estimator of Powell-Stock-Stoker_1989_ECMA given by
where $\widehat{f_i}(\cdot)$ is a “leave-one-out” kernel density estimator employing a symmetric and differentiable kernel function $K:\mathbb{R}^{d}\rightarrow \mathbb{R}$ and a positive vanishing (bandwidth) sequence $h$.
The estimator $\widehat{\theta}$ can be expressed as a second-order U-statistic with an $n$-varying kernel:
Our analysis of $\widehat{\theta}$ is based on this representation and proceeds under the following assumption about the kernel function.\footnote{In Assumption (ref) (c) and elsewhere, we employ standard multi-index notation: For $a := (a_1,\dots,a_d)' \in \mathbb{Z}_+^d$, we have (i) $[a] := a_1+\dots +a_d$, (ii) $a! := a_1!\dots a_d!$, (iii) $x^a := x_1^{a_1}\dots x_d^{a_d}$ for $x := (x_1,\dots,x_d)' \in \mathbb{R}^d$, and (iv) $\partial^a q(x) / \partial x^a := \partial^{[a]} q(x) / (\partial x_1^{a_1}\dots \partial x_d^{a_d})$ for (sufficiently smooth) $q:\mathbb{R}^d\to\mathbb{R}$.}
Before presenting our main results concerning the higher-order distributional properties of different statistics based on $\widehat{\theta}$, we review conventional and alternative first-order asymptotic distributional approximations, as well as the distinct variance estimation methods emerging from each of those approximation frameworks. Limits are taken as $h\to0$ and $n\to\infty$ unless otherwise noted, $\to_\mathbb{P}$ denotes convergence in probability, and $\rightsquigarrow$ denotes convergence in law.
Under appropriate restrictions on $h$ and $K$, the estimator $\widehat{\theta}$ is asymptotically linear with influence function $\psi$ and asymptotic variance $\Sigma$. More precisely, Powell-Stock-Stoker_1989_ECMA showed that if Assumptions (ref) and (ref) hold and if $nh^{2(P \land S)}\to 0$ and $nh^{d+2}\to \infty $ (where $a \land b$ denotes $\min(a,b)$), then
A proof of (ref) can be based on the $U$-statistic representation in (ref) and its Hoeffding decomposition $\widehat{\theta} = \mathbb{E}[U_{ij}] + \bar{L} + \bar{Q}$, where $\bar{L}$ and $\bar{Q}$ are mean zero random vectors given by
and
respectively: because $\mathbb{E}[U_{ij}] = \theta + O(h^{P \land S})$ and $\mathbb{V}[\bar{Q}]=O(n^{-2} h^{-d-2})$, we have
from which the result (ref) follows upon noting that $\mathbb{V}[L_i - \psi(Z_i)] = O(h^{P \land S})$. Using Edgeworth expansions, Nishiyama-Robinson_2000_ECMA,Nishiyama-Robinson_2001_ChBook studied the quality of the distributional approximation implied by (ref); their result is contained as a special case of our Theorem (ref).
The Hoeffding decomposition and subsequent analysis of each of its terms shows that the estimator admits a bilinear form representation in general, which then is reduced to a sample average approximation by assuming a bandwidth sequence and kernel shape that makes both the misspecification error (smoothing bias) and the variability introduced by $\bar{Q}$ (a “quadratic” term) negligible in large samples. As a result, provided that such tuning parameter choices are feasible, the estimator will be asymptotically linear.
Asymptotic linearity of a semiparametric estimator has several distinct features that may be considered attractive from a theoretical point of view. In particular, it is a necessary condition for semiparametric efficiency, and it leads to a limiting distribution that is invariant to the choice of the first-step nonparametric estimator entering the two-step semiparametric procedure Newey_1994_ECMA. However, insisting on asymptotic linearity may also have its drawbacks because it requires several potentially strong assumptions, and because it leads to a large sample theory that may not accurately represent the finite sample behavior of the statistic. In the case of $\widehat{\theta}$, asymptotic linearity requires $P>2$ unless $d=1$; that is, the use of higher-order kernels or similar debiasing techniques Chernozhukov-etal_2022_ECMA is necessary in order to achieve asymptotic linearity. In addition, asymptotic linearity leads to a limiting experiment which is invariant to the particular choices of smoothing ($K$) and bandwidth ($h$) tuning parameters involved in the construction of the estimator. As a result, large sample distribution theory based on (or implying) asymptotic linearity is silent with respect to the impact that tuning parameter choices may have on the finite sample behavior of the two-step semiparametric statistic.
To address the aforementioned limitations of distribution theory based on asymptotic linearity, Cattaneo-Crump-Jansson_2014a_ET proposed a more general distributional approximation for kernel-based DWAD estimators that accommodates, but does not enforce, asymptotic linearity. The idea is to characterize the joint asymptotic distributional features of both the linear ($\bar{L}$) and quadratic ($\bar{Q}$) terms, and in the process develop a more general first-order asymptotic theory that allows for weaker assumptions than those imposed in the classical asymptotically linear distribution theory. Formally, if Assumptions (ref) and (ref) hold, and if $(nh^{d+2} \land 1) nh^{2(P \land S) }\to 0$ and $n^{2}h^{d}\to \infty$, then
where $\mathbb{V}[\widehat{\theta}] = \mathbb{V}[\bar{L}] + \mathbb{V}[\bar{Q}]$ with
and
This more general distributional approximation was developed explicitly in an attempt to better characterize the finite sample behavior of $\widehat{\theta}$. The result in (ref) shows that the conditions on the bandwidth sequence may be considerably weakened without invalidating the limiting Gaussian distribution, although the asymptotic variance formula changes. Importantly, if $nh^{d+2} $ is bounded then $\widehat{\theta}$ is no longer asymptotically linear and its limiting distribution will cease to be invariant with respect to the underlying preliminary nonparametric estimator. In particular, if $nh^{d+2}\to c >0$ then $\widehat{\theta}$ is root-$n$ consistent, but not asymptotically linear. The bias of the estimator is also controlled in a different way because the bandwidth is allowed to be “smaller” than usual, which may remove the need for higher-order kernels. Interestingly, (ref) allows for the point estimator to not even be consistent for $\theta$, which occurs for sufficiently small bandwidth sequences.
Beyond the aforementioned technical considerations, the result in (ref) can conceptually be interpreted as a more refined first-order distributional approximation for $\widehat{\theta}$, which by relying on a quadratic approximation (i.e., accounting for the contributions of both $\bar{L}$ and $\bar{Q}$) is expected to offer a “better” distributional approximation than approximations relying on asymptotic linearity (i.e., accounting only for the contribution of $\bar{L}$). The idea of standardizing a U-statistic by the joint variance of the linear and quadratic terms underlying its Hoeffding decomposition can be traced back to the original paper of Hoeffding_1948_AMS. Simulation evidence reported in Cattaneo-Crump-Jansson_2010_JASA,Cattaneo-Crump-Jansson_2014a_ET,Cattaneo-Crump-Jansson_2014b_ET corroborated those conceptual interpretations numerically, but no formal justification is available in the literature. Theorem (ref) below will offer the first theoretical result in the literature highlighting specific robustness features of the distributional approximation in (ref) by showing that such approximation has a demonstrably smaller higher-order distributional approximation error.
Motivated by the asymptotic linearity result (ref), Powell-Stock-Stoker_1989_ECMA also proposed the “plug-in” variance estimator
and proved its consistency (i.e., $\widehat{\Sigma}\to_\mathbb{P} \Sigma$) under the same bandwidth sequences required for asymptotic linearity (i.e., assuming $nh^{2(P \land S)}\to 0$ and $nh^{d+2}\to \infty $). Combining this consistency result with (ref), we obtain the following result about a studentized version of $\widehat{\theta}$:
Using Edgeworth expansions, Nishiyama-Robinson_2000_ECMA,Nishiyama-Robinson_2001_ChBook studied the quality of the distributional approximation implied by (ref); their result is contained as a special case of our Theorem (ref).
Complementing the small bandwidth asymptotic representation (ref), Cattaneo-Crump-Jansson_2014a_ET showed that
which implies among other things that the consistency result $\widehat{\Sigma}\to_\mathbb{P} \Sigma$ is valid only if $nh^{d+2}\to \infty$; otherwise, $\widehat{\Sigma}$ is in general asymptotically upwards biased relative to $\mathbb{V}[\widehat{\theta}]$ in (ref). Because $\widehat{\Sigma}$ is asymptotically equivalent to the jackknife variance estimator of $\widehat{\theta}$, Cattaneo-Crump-Jansson_2014b_ET also noted that the asymptotic bias of $\widehat{\Sigma}$ is a consequence of a more generic phenomena underlying jackknife variance estimators studied in Efron-Stein_1981_AOS. See also Matsushita-Otsu_2021_Biometrika for related discussion.
To conduct asymptotically valid inference under the more general small bandwidth asymptotic regime, Cattaneo-Crump-Jansson_2014a_ET proposed several “debiased” variance estimators, including
and showed that $\widehat{\Delta}\to_\mathbb{P} \Delta$ under the same bandwidth sequences required for (ref) to hold (i.e., assuming $nh^{2(P \land S)}\to 0$ and $n^2 h^d \to \infty$). The estimator $\widehat{\Delta}$ is asymptotically equivalent to the debiasing procedure proposed in Efron-Stein_1981_AOS. By design, the result
holds under more general conditions than those required for (ref), suggesting that inference procedures based on $\widehat{V}_\mathtt{SB}$ are more “robust” than procedures based on $\widehat{V}_\mathtt{AL}$.
Conceptually, robustness manifests itself in two distinct ways. First, the underlying Gaussian distributional approximation holds under weaker bandwidth restrictions, a property achieved in part by employing a standardization factor depending explicitly on tuning parameter choices. Second, the new variance estimator $\widehat{V}_\mathtt{SB}$ is obtained from the more general small bandwidth approximation and explicitly accounts for the contribution of terms regarded as higher-order under asymptotic linearity.
While not reproduced here to conserve space, the in-depth Monte Carlo evidence reported in Cattaneo-Crump-Jansson_2010_JASA,Cattaneo-Crump-Jansson_2014a_ET,Cattaneo-Crump-Jansson_2014b_ET also showed that employing inference procedures based on (ref) lead to remarkable improvements in terms of “robustness” to bandwidth choice and other tuning inputs, when compared to classical asymptotically linear inference procedures based on (ref). Theorem (ref) below will show formally that the distributional approximation (ref) has demonstrably smaller higher-order errors than the distributional approximation (ref), thereby providing a theory-based explanation for the empirical success of feasible inference procedures developed under the small bandwidth approximation framework.
Letting $\mathsf{v}\in\mathbb{R}^d$ be a fixed vector and defining $\widehat{\theta}_\mathsf{v}:=\mathsf{v}'\widehat{\theta}$, this section presents Edgeworth expansions for standardized and studentized statistics based on $\widehat{\theta}_\mathsf{v}$. Section (ref) studies standardized statistics of the form $(\widehat{\theta}_\mathsf{v} - \theta_\mathsf{v})/\vartheta_\mathsf{v}$, where $\theta_\mathsf{v}:=\mathsf{v}'\theta$ and $\vartheta_\mathsf{v}$ is an approximate standard deviation of $\widehat{\theta}_\mathsf{v}$; that is, $\vartheta_\mathsf{v}$ is positive, non-random, and such that $(\widehat{\theta}_\mathsf{v} - \theta_\mathsf{v})/\vartheta_\mathsf{v}$ is asymptotically standard normal. The main purpose of studying standardized statistics is to allow us to compare the quality of the distributional approximations (ref) and (ref) based on asymptotic linearity and small bandwidth asymptotics, respectively. Section (ref) then studies studentized statistics of the form $(\widehat{\theta}_\mathsf{v} - \theta_\mathsf{v})/\widehat{\vartheta}_\mathsf{v}$, where $\widehat{\vartheta}_\mathsf{v}^2$ is (random and) equal to either $\mathsf{v}'\widehat{V}_\mathtt{AL}\mathsf{v}$ or $\mathsf{v}'\widehat{V}_\mathtt{SB}\mathsf{v}$. The main purpose of studying studentized statistics is to allow us to investigate the impact of variance estimation on the quality of the distributional approximations (ref) and (ref) based on asymptotic linearity and small bandwidth asymptotics, respectively.
Nishiyama-Robinson_2000_ECMA,Nishiyama-Robinson_2001_ChBook obtained valid Edgeworth expansions for the distribution of the standardized and studentized statistics employing $\vartheta_\mathsf{v}^2 = \mathsf{v}'\Sigma\mathsf{v}/n$ and $\widehat{\vartheta}_\mathsf{v}^2 = \mathsf{v}'\widehat{\Sigma}\mathsf{v}/n$, respectively. Those results were obtained under assumptions implying asymptotic linearity. Although our main interest is in standardization and studentization schemes whose (first-order) validity does not require asymptotic linearity, we retain the assumption of asymptotic linearity to ensure a fair comparison; that is, our results are derived under the same assumptions as those imposed in prior work, in which case all inference procedures are asymptotically valid, and therefore amenable to juxtaposition. While beyond the scope of this paper, allowing for departures from asymptotic linearity is an interesting topic for future research.
Suppressing the dependence on $\mathsf{v}$ and $\vartheta_\mathsf{v}$, let
be the cumulative distribution function (cdf) of $(\widehat{\theta}_\mathsf{v} - \theta_\mathsf{v})/\vartheta_\mathsf{v}$, where $\vartheta_\mathsf{v}$ is positive and non-random. Letting $\Phi$ denote the standard normal cdf, it follows from (ref) that
under assumptions implying in particular that $\omega_\mathsf{v}^2/\vartheta_\mathsf{v}^2 \to 1$, where
with $\sigma_\mathsf{v}^2 := \mathsf{v}'\Sigma\mathsf{v}$ and $\delta_\mathsf{v}^2 := \mathsf{v}'\Delta\mathsf{v}$.
Our first theorem provides a refinement of (ref). To state the theorem, let
and define the following quantities (all of which are finite under the assumptions of Theorem (ref)):
Also, let $\phi$ denote the standard normal probability density function.
The proof of the theorem proceeds by verifying the high-level conditions of more general results presented in Appendix (ref). The general results establish a valid Edgeworth expansion for a generic class of U-statistics with $n$-varying kernels and may be of independent theoretical interest. Theorem (ref) generalizes Nishiyama-Robinson_2000_ECMA by allowing for a generic standardization factors $\vartheta_\mathsf{v}$ instead of their specific choice $\sqrt{\mathsf{v}'\Sigma\mathsf{v}/n} = \sigma_\mathsf{v}/\sqrt{n}$. The latter generalization is important for our purposes, as it enables us to compare the different distributional approximations implied by (ref) and (ref).
As is customary with Edgeworth expansions, the square-bracketed term in the function $G$ is a “correction” term capturing the extent to which the first three cumulants of the statistic differ from those of the standard normal distribution. To be specific, the first and third terms correct for bias and skewness, respectively. None of these correction terms depend on the particular $\vartheta_\mathsf{v}$ used for standardization purposes. In contrast, and as was to be expected, the variance correction term does depend on $\vartheta_\mathsf{v}$, being proportional to $\omega_\mathsf{v}^2/\vartheta_\mathsf{v}^2 - 1$.
The asymptotic linearity result (ref) suggests setting $\vartheta_\mathsf{v}^2=\sigma_\mathsf{v}^2/n$. Doing so, and in agreement with Nishiyama-Robinson_2000_ECMA, we have
the approximation error being $o(r_n)$. In the display, the term on the right hand side involves $\delta_\mathsf{v}^2 = \lim_{n\to\infty}h^{d+2}\mathbb{V}[\mathsf{v}'Q_{ij}\mathsf{v}]$ and is therefore interpretable as a variance correction term intrinsically associated with approximations based on asymptotic linearity, as such approximations ignore the contribution of the “quadratic” terms $Q_{ij}$ to the variability of $\widehat{\theta}$. Unlike (ref), the small bandwidth formulation (ref) explicitly accounts for the presence of “quadratic” terms and the standardization factor $\vartheta_\mathsf{v}^2 = \mathbb{V}[\widehat{\theta}_\mathsf{v}] = \omega_\mathsf{v}^2$ suggested by the small bandwidth formulation is one for which the variance correction term in $G$ vanishes altogether.
Although the details of the results reported here are specific to DWAD estimation, one important qualitative conclusion appears to generalize: the feature that it can be advantageous (in a higher-order sense) to capture the full variability of a statistic when approximating its distribution is known to be shared by certain statistics arising in the context of nonparametric kernel-based density and local polynomial regression inference; for details, see Calonico-Cattaneo-Farrell_2018_JASA,Calonico-Cattaneo-Farrell_2022_Bernoulli.
In isolation, Theorem (ref) is mostly of theoretical interest, the reason being that it is concerned with standardized (as opposed to studentized) estimators. The consequences of employing studentization (i.e., replacing $\vartheta_\mathsf{v}$ with an estimator) will be explored in the next subsection. One important qualitative conclusion of that subsection concerns inference. That conclusion can be anticipated with the help of Theorem (ref). We conclude this subsection by doing so.
For any $\alpha\in(0,1)$, a natural (albeit infeasible) $100(1-\alpha)\%$ two-sided confidence interval for $\theta_\mathsf{v}$ has endpoints given by $\widehat{\theta}_\mathsf{v} \pm c_\alpha \vartheta_\mathsf{v}$, where $c_\alpha := \Phi^{-1}(1-\alpha/2)$ and where $\vartheta_\mathsf{v}^2$ is an approximate variance of $\widehat{\theta}_\mathsf{v}$. Under the assumptions of Theorem (ref), the coverage probability of this interval satisfies
so to the order considered the coverage error is proportional to the term $\omega_\mathsf{v}^2/\vartheta_\mathsf{v}^2 - 1$ discussed previously and our conclusions about this term therefore apply directly. In particular, the coverage error of an infeasible interval using $\vartheta_\mathsf{v}^2 = \mathbb{V}[\widehat{\theta}_\mathsf{v}]$ (as suggested by the small bandwidth asymptotic result (ref)) is $o(r_n)$, while the coverage errors of an infeasible intervals using $\vartheta_\mathsf{v}^2 = \sigma_\mathsf{v}^2/n$ (as suggested by the asymptotic linearity result (ref)) or its pre-asymptotic counterpart $\vartheta_\mathsf{v}^2 = \mathbb{V}[\mathsf{v}'L_i]/n$ are of larger magnitude.
Next, we investigate the role of variance estimation by obtaining Edgeworth expansions for studentized versions of $\widehat{\theta}$. For specificity, and inspired by (ref) and (ref), we compare
and
Studying $\widehat{F}_\mathtt{AL}$, Nishiyama-Robinson_2000_ECMA found that if the assumptions of Theorem (ref) below are satisfied, then
with
where, once again, the three terms in square brackets correct for bias, variance, and skewness, respectively. In light of the results of the previous subsection, one would expect the Edgeworth approximation to $\widehat{F}_\mathtt{SB}$ to be similar to $\widehat{G}_\mathtt{AL}$, the only (possible) difference being the variance correction term. The following result shows that this is indeed the case.
This theorem shows that employing studentization based on small bandwidth asymptotics offers demonstrable improvements in terms of distributional approximations for the resulting feasible t-test: the variance correction term present in $\widehat{G}_\mathtt{AL}$ is absent from $\widehat{G}_\mathtt{SB}$. As in Nishiyama-Robinson_2000_ECMA,Nishiyama-Robinson_2001_ChBook, the result is obtained under the somewhat stronger moment condition $\mathbb{E}[Y^6] < \infty$ than the condition $\mathbb{E}[|Y|^3] < \infty$ of Theorem (ref), the purpose of the strengthened condition being to help control the contribution of the random denominator of the studentized version of $\widehat{\theta}$.
The main practical implication of Theorem (ref) can be illustrated by analyzing the coverage error of $100(1-\alpha)\%$ confidence intervals with endpoints $\widehat{\theta}_\mathsf{v} \pm c_\alpha \widehat{\vartheta}_\mathsf{v}$. Setting $\widehat{\vartheta}_\mathsf{v} = \widehat{\vartheta}_{\mathtt{AL},\mathsf{v}}$ and applying Nishiyama-Robinson_2000_ECMA, we have
whereas setting $\widehat{\vartheta}_\mathsf{v} = \widehat{\vartheta}_{\mathtt{SB},\mathsf{v}}$ and applying Theorem (ref) gives
In other words, confidence intervals based on (ref) are demonstrably superior to those based on (ref) from a higher-order asymptotic point of view. This finding provides a theoretical explanation of the simulation evidence reported in Cattaneo-Crump-Jansson_2014a_ET,Cattaneo-Crump-Jansson_2014b_ET,Cattaneo-Crump-Jansson_2010_JASA, where feasible confidence intervals based on small bandwidth asymptotics were shown to offer better finite sample performance in terms of coverage error than their counterparts based on classical asymptotic linear approximations.
In the previous subsections, we investigated the higher-order performance of large sample distributional approximations under two alternative asymptotic frameworks: asymptotic linearity and small bandwidth asymptotics. We found that the choice of studentization matters in terms of distributional approximation errors, even under conditions guaranteeing that asymptotic linearity holds ($nh^{d+2}\to\infty$), in which case both asymptotic frameworks are first-order valid. As a consequence, our Edgeworth expansions (reported in Theorems (ref) and (ref)) provide alternative validation of the main conclusions obtained by Cattaneo-Crump-Jansson_2010_JASA,Cattaneo-Crump-Jansson_2014a_ET using first-order distributional approximations: in terms of distributional approximation accuracy, the small bandwidth framework justifying (ref) and (ref) dominates the asymptotic linear framework justifying (ref) and (ref).
It is natural to ask whether a similar ranking emerges when employing bootstrap-based inference procedures. Cattaneo-Crump-Jansson_2014b_ET studied the first-order properties of the nonparametric bootstrap under small bandwidth asymptotics for the kernel-based DWAD estimator, and showed that a similar first-order pattern emerges in that case: Under slightly stronger assumptions than Assumptions (ref) and (ref), they showed that if $(nh^{d+2} \land 1) nh^{2(P \land S) }\to 0$ and if $n^2 h^d \to \infty$, then
where $\rightsquigarrow_\mathbb{P}$ denotes weak convergence in probability,
and
with $\widehat{\theta}^*$, $\bar{L}^*$, $\bar{Q}^*$, $\widehat{\Sigma}^*$ and $\widehat{\Delta}^*$ denoting nonparametric bootstrap analogs of $\widehat{\theta}$, $\bar{L}$, $\bar{Q}$, $\widehat{\Sigma}$ and $\widehat{\Delta}$, respectively, and $\mathbb{V}^*[\cdot]$ denoting the variance computed conditional on the original data. It follows from those results that under small bandwidth asymptotics the bootstrap consistently estimates the distribution of a studentized version of $\widehat{\theta}$ when $\widehat{V}_\mathtt{SB}$ is used for studentization purposes, but not when $\widehat{V}_\mathtt{AL}$ is. These conclusions are in perfect agreement with those discussed in Section (ref), the only notable difference in the details being that the variability induced by the bootstrap is larger outside the asymptotic linear regime because $\mathbb{V}^*[\bar{Q}^*]/\mathbb{V}[\bar{Q}]=3+o_\mathbb{P}(1)$. Furthermore, the second term of the bootstrap-based jackknife variance estimator $\widehat{\Sigma}^*$ is asymptotically doubled relative to the second term of the jackknife variance estimator $\widehat{\Sigma}$.
Nishiyama-Robinson_2005_ECMA used Edgeworth expansions to study the properties of inference procedures based on the nonparametric bootstrap under assumptions implying asymptotic linearity. Based on the findings in this paper and those in Cattaneo-Crump-Jansson_2014b_ET, we conjecture that analogous conclusions to those obtained herein will be valid for the case of bootstrap-based inference. While the conceptual parallelism between bootstrap-based inference and the results reported in this paper are clear, formalizing our conjecture requires substantial additional technical work due to the added complications associated with the data resampling, and hence we leave the theoretical analysis for future work.
Employing Edgeworth expansions, we compared the higher-order properties of two first-order distributional approximations and their associated confidence intervals for the kernel-based DWAD estimator of Powell-Stock-Stoker_1989_ECMA. We showed that small bandwidth asymptotics not only give demonstrably better distributional approximations than those implied by asymptotic linearity, but also justifies employing a variance estimator for studentization purposes that improves the distributional approximation. The main takeaway from our results is that in two-step semiparametric settings, and related problems, alternative asymptotic approximations that capture higher-order terms, which are ignored by more traditional asymptotic linearity-based approximations, can deliver better distributional approximations and, by implication, more accurate inference procedures in finite samples. See Cattaneo-Jansson-Newey_2018_ET for related discussion.
While beyond the scope of this paper, it would be of interest to develop analogous Edgeworth expansions for more general linear and non-linear two-step semiparametric procedures employing either Gaussian or resampling approximations under both conventional and alternative asymptotic frameworks Cattaneo-Crump-Jansson_2013_JASA,Cattaneo-Jansson_2018_ECMA,Cattaneo-Jansson-Ma_2019_RESTUD,Cattaneo-Jansson_2022_ET. In particular, for the special case of kernel-based DWAD estimators, which is a linear two-step kernel-based semiparametric estimator, Nishiyama-Robinson_2005_ECMA already obtained Edgeworth expansions for bootstrap-based inference procedures under asymptotic linearity that could be contrasted with those obtained under small bandwidth asymptotics Cattaneo-Crump-Jansson_2014b_ET, after establishing more general Edgeworth expansions accounting for the bootstrap.