Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
81,327 characters · 21 sections · 0 citation commands
Bootstrap Inference in the Presence of Bias
\noindentKeywords:\ Asymptotic bias, bootstrap, incidental parameter bias, model averaging, nonparametric regression, prepivoting.
\def\spacingset#1 \spacingset{1} \thispagestyle{empty} \spacingset{1.1}
Suppose that $\theta$ is a scalar parameter of interest and let $\hat{\theta}_{n}$ denote an estimator for which
where $g(n)\rightarrow\infty$ is the rate of convergence of $\hat{\theta}_{n} $, $\xi_{1}$ is a continuous random variable centered at zero, and $B$ is an asymptotic bias (our theory in fact allows for a more general formulation of the bias). A typical example is $g(n)=n^{1/2}$ and $\xi_{1}\sim N(0,\sigma ^{2})$. Unless $B$ can be consistently estimated, which is often difficult or impossible, classic (first-order) asymptotic inference on $\theta$ based on quantiles of $\xi_{1}$ in ((ref)) is not feasible. Furthermore, the bootstrap, which is well known to deliver asymptotic refinements over first-order asymptotic approximations as well as bias corrections (Hall, 1992; Horowitz, 2001; Cattaneo and Jansson, 2018, 2022; Cattaneo, Jansson, and Ma, 2019), cannot in general be applied to solve the asymptotic bias problem when a consistent estimator of $B$ does not exist. Examples are given below.
Our goal is to justify bootstrap inference based on $T_{n}$ in the context of asymptotically biased estimators and where a consistent estimator of $B$ does not exist. Consider the bootstrap statistic $T_{n}^{\ast}:=g(n)(\hat{\theta }_{n}^{\ast}-\hat{\theta}_{n})$, where $\hat{\theta}_{n}^{\ast}$ is a bootstrap version of $\hat{\theta}_{n}$, such that
where $\hat{B}_{n}$ is the implicit bootstrap bias, and `$\overset{d^{\ast }}{\rightarrow}_{p}$' denotes weak convergence in probability (defined below). When $\hat{B}_{n}-B=o_{p}(1)$, the bootstrap is asymptotically valid in the usual sense that the bootstrap distribution of $T_{n}^{\ast}$ is consistent for the asymptotic distribution of $T_{n}$, i.e., $\sup_{x\in\mathbb{R} }|P^{\ast}(T_{n}^{\ast}\leq x)-P(T_{n}\leq x)|=o_{p}(1)$.
We consider situations where $\hat{B}_{n}-B$ is not asymptotically negligible so the bootstrap fails to replicate the asymptotic bias. For example, this happens when the asymptotic bias term in the bootstrap world includes a random (additive) component, i.e.\
where $\xi_{2}$ is a random variable centered at zero. In this case, the bootstrap distribution is random in the limit and hence cannot mimic the asymptotic distribution given in ((ref)). Moreover, the distribution of the bootstrap p-value, $\hat{p}_{n}:=P^{\ast}(T_{n}^{\ast}\leq T_{n})$, is not asymptotically uniform, and the bootstrap cannot in general deliver hypothesis tests (or confidence intervals) with the desired null rejection probability (or coverage probability).
In this paper, we show that in this non-standard case valid inference can successfully be restored by proper implementation of the bootstrap. This is done by focusing on properties of the bootstrap p-value rather than on the bootstrap as a means of estimating limiting distributions, which is infeasible due to the asymptotic bias. In particular, we show that such implementations lead to bootstrap inferences that are valid in the sense that they provide asymptotically uniformly distributed p-values.
Our inference strategy is based on the fact that, for some bootstrap schemes, the large-sample distribution of the bootstrap p-value, say $H(u)$, $u\in\lbrack0,1]$, although not uniform, does not depend on $B$. That is, we can search for bootstrap algorithms which generate bootstrap p-values that, in large samples, are not affected by unknown bias terms. When this is possible, we can make use of the prepivoting approach of Beran (1987, 1988), which --- as we will show in this paper --- allows to restore bootstrap validity. Specifically, our proposed modified p-value is defined as \[ \tilde{p}_{n}:=\hat{H}_{n}(\hat{p}_{n}), \] where $\hat{H}_{n}(u)$ is any consistent estimator of $H(u)$, uniformly over $u\in\lbrack0,1]$. The (asymptotic) probability integral transform $\hat {p}_{n}\mapsto H(\hat{p}_{n})$, continuity of $H(u)$, and consistency of $\hat{H}_{n}(u)$ then guarantee that $\tilde{p}_{n}$ is asymptotically uniformly distributed. Interestingly, Beran (1987, 1988) proposed this approach to obtain asymptotic refinements for the bootstrap, but did not consider asymptotically biased estimators as we do here.
We propose two approaches to estimating $H$. First, if $H=H_{\gamma}$, where $\gamma$ is a finite-dimensional parameter vector, and a consistent estimator $\hat{\gamma}_{n}$ of $\gamma$ is available, then a `plug-in' approach setting $\hat{H}_{n}=H_{\hat{\gamma}_{n}}$ can deliver asymptotically uniform p-values. Second, if estimation of $\gamma$ is difficult (e.g., when $\gamma$ does not have a closed form expression), we can use a `double bootstrap' scheme (Efron, 1983; Hall, 1986), where estimation of $H$ is achieved by resampling from the bootstrap data originated in the first level.
For both methods, we provide general high-level conditions that imply validity of the proposed approach. Our conditions are not specific to a given bootstrap method; rather, they can in principle be applied to any bootstrap scheme satisfying the proposed sufficient conditions for asymptotic validity.
Our approach is related to recent work by Shao and Politis (2013) and Cavaliere and Georgiev (2020). In particular, a common feature is that the distribution function of the bootstrap statistic, conditional on the original data, is random in the limit. Cavaliere and Georgiev (2020)\ emphasize that randomness of the limiting bootstrap measure does not prevent the bootstrap from delivering an asymptotically uniform p-value (bootstrap `unconditional' validity), and provide results to assess such asymptotic uniformity. Our context is different, since the presence of an asymptotic bias term renders the distribution of the bootstrap p-value non-uniform, even asymptotically. In this respect, our work is related to Shao and Politis (2013), who show that $t$-statistics based on subsampling or block bootstrap methods with bandwidth proportional to sample size may deliver non-uniformly distributed p-values that, however, can be estimated.
To illustrate the practical relevance of our results and to show how to implement them in applied problems, we consider three examples involving estimators that feature an asymptotic bias term. In the first two examples (model averaging and ridge regression), $B$ is not consistently estimable due to the presence of local-to-zero parameters and the standard bootstrap fails. In the third example (nonparametric regression), the bootstrap fails because $B$ depends on the second-order derivative of the conditional mean function, whose estimation requires the use of a different (suboptimal) bandwidth. In these examples, $\xi_{1}$ is normal, but $g(n)$ and $B$ are example-specific. Two additional examples are presented in the supplement. The fourth is a simple location model without the assumption of finite variance, where $\xi_{1}$ is not normal and estimators converge at an unknown rate. The fifth example considers inference for dynamic panel data models, where $B$ is the incidental parameter bias.
The remainder of the paper is organized as follows. In Section (ref) we introduce our three leading examples. Section (ref) contains our general results, which we apply to the three examples in Section (ref). Section (ref) concludes. The supplemental material contains two appendices. Appendix (ref) specializes the general theory to the case of asymptotically Gaussian statistics, and Appendix (ref) contains details and proofs for the three leading examples, as well as two additional examples.
Throughout this paper, the notation $\sim$ indicates equality in distribution. For instance, $Z\sim N(0,1)$ means that $Z$ is distributed as a standard normal random variable. We write `$x:=y$' and `$y=:x$' to mean that $x$ is defined by $y$. The standard Gaussian cumulative distribution function (cdf) is denoted by $\Phi$; $U_{[0,1]}$ is the uniform distribution on $[0,1]$, and $\mathbb{I}_{\{\cdot\}}$ is the indicator function. If $F$ is a cdf, $F^{-1}$ denotes the generalized inverse, i.e.\ the quantile function, $F^{-1} (u):=\inf\{v\in\mathbb{R}:F(v)\geq u\}$, $u\in\mathbb{R}$. Unless specified otherwise, all limits are for $n\rightarrow\infty$. For matrices $a,b,c$ with $n$ rows, we let $S_{ab}:=a^{\prime}b/n$ and $S_{ab.c}:=S_{ab}-S_{ac} S_{cc}^{-1}S_{cb}$, assuming that $S_{cc}$ has full rank.
For a (single level or first-level) bootstrap sequence, say $Y_{n}^{\ast}$, we use $Y_{n}^{\ast}\overset{p^{\ast}}{\rightarrow}_{p}0$, or equivalently $Y_{n}^{\ast}\overset{p^{\ast}}{\rightarrow}0$, in probability, to mean that, for any $\epsilon>0$, $P^{\ast}(|Y_{n}^{\ast}|>\epsilon)\rightarrow_{p}0$, where $P^{\ast}$ denotes the probability measure conditional on the original data $D_{n}$. An equivalent notation is $Y_{n}^{\ast}=o_{p^{\ast}}(1)$ (where we omit the qualification \textquotedblleft in probability\textquotedblright \ for brevity). Similarly, for a double (or second-level) bootstrap sequence, say $Y_{n}^{\ast\ast}$, we write $Y_{n}^{\ast\ast}=o_{p^{\ast\ast}}(1)$ to mean that for all $\epsilon>0$, $P^{\ast\ast}(|Y_{n}^{\ast\ast}|>\epsilon )\overset{p^{\ast}}{\rightarrow}_{p}0$, where $P^{\ast\ast}$ is the probability measure conditional on the first-level bootstrap data $D_{n} ^{\ast}$ and on $D_{n}$.
We use $Y_{n}^{\ast}\overset{d^{\ast}}{\rightarrow}_{p}\xi$, or equivalently $Y_{n}^{\ast}\overset{d^{\ast}}{\rightarrow}\xi$, in probability, to mean that, for all continuity points $u\in\mathbb{R}$ of the cdf of $\xi$, say $G(u):=P(\xi\leq u)$, it holds that $P^{\ast}(Y_{n}^{\ast}\leq u)-G(u)\rightarrow_{p}0$. Similarly, for a double bootstrap sequence $Y_{n}^{\ast\ast}$, we use $Y_{n}^{\ast\ast}\overset{d^{\ast\ast} }{\rightarrow}_{p^{\ast}}\xi$, in probability, to mean that $P^{\ast\ast }(Y_{n}^{\ast\ast}\leq u)-G(u)\overset{p^{\ast}}{\rightarrow}_{p}0$ for all continuity points $u$ of $G$.
In this section we introduce our three leading examples. Example-specific regularity conditions, formally stated results, and additional definitions are given in Appendix (ref). For each of these examples, we argue that ((ref)), ((ref)), and ((ref)) hold, such that the bootstrap p-values $\hat{p}_{n}$ are not uniformly distributed rendering standard bootstrap inference invalid. We then return to each example in Section (ref), where we discuss how to implement our proposed method and prove its validity.
\noindentSetup. We consider inference based on a model averaging estimator obtained as a weighted average of least squares estimates (Hansen, 2007). Assume that data are generated according to the linear model
where $\beta$ is the (scalar) parameter of interest and $\varepsilon$ is an $n$-vector of identically and independently distributed random variables with mean zero and variance $\sigma^{2}$ (henceforth i.i.d.$(0,\sigma^{2})$), conditional on $W:=(x,Z)$.
The researcher fits a set of $M$ models, each of them based on different exclusion restrictions on the $q$-dimensional vector $\delta$. This setup allows for model averaging both explicitly and implicity. The former follows, e.g., Hansen (2007). The latter includes the common practice of robustness checks in applied research, where the significance of a target coefficient is evaluated through an (often informal) assessment of its significance across a set of regressions based on different sets of controls; see Oster (2019) and the references therein. Specifically, letting $R_{m}$ denote a $q\times q_{m}$ selection matrix, the $m^{\text{th}}$ model includes $x$ and $Z_{m}:=ZR_{m}$ as regressors, and the corresponding OLS\ estimator of $\beta$ is $\tilde{\beta}_{m,n}=S_{xx.Z_{m}}^{-1}S_{xy.Z_{m}}$. Given a set of fixed weights $\omega:=(\omega_{1},\dots,\omega_{M})^{\prime}$ such that $\omega _{m}\in\lbrack0,1]$ and $\sum_{m=1}^{M}\omega_{m}=1$, the model averaging estimator is $\tilde{\beta}_{n}:=\sum_{m=1}^{M}\omega_{m}\tilde{\beta}_{m,n}$. Then $T_{n}:=n^{1/2}(\tilde{\beta}_{n}-\beta)$ satisfies $T_{n}-B_{n} \rightarrow_{d}\xi_{1}\sim N(0,v^{2})$, where $v^{2}>0$ and \[ B_{n}:=Q_{n}n^{1/2}\delta,\text{\quad}Q_{n}:=\sum_{m=1}^{M}\omega _{m}S_{xx.Z_{m}}^{-1}S_{xZ.Z_{m}}. \] Thus, the magnitude of the asymptotic bias $B_{n}$ depends on $n^{1/2}\delta$. If $\delta$ is local to zero in the sense that $\delta=cn^{-1/2}$ for some vector $c\in\mathbb{R}^{q}$ (as in, e.g., Hjort and Claeskens, 2003; Liu, 2015; Hounyo and Lahiri, 2023), then $B_{n}\rightarrow_{p}B:=Qc$ with $Q:=\operatorname*{plim}Q_{n}$, so that ((ref)) is satisfied with nonzero $B$ in general. Because $B$ depends on $c$, which is not consistently estimable, we cannot obtain valid inference from a Gaussian distribution based on sample analogues of $B$ and $v^{2}$.
\noindentFixed regressor bootstrap. We generate the bootstrap sample as $y^{\ast}=x\hat{\beta}_{n}+Z\hat{\delta}_{n}+\varepsilon^{\ast}$, where $\varepsilon^{\ast}|D_{n}\sim N(0,\hat{\sigma}_{n}^{2}I_{n})$, $(\hat{\beta }_{n},\hat{\delta}_{n}^{\prime},\hat{\sigma}_{n}^{2})$ is the OLS\ estimator from the full model, and $D_{n}=\{y,W\}$. Similar results can be established for the nonparametric bootstrap where $\varepsilon^{\ast}$ is resampled from the full model residuals. The bootstrap model averaging estimator is given by $\tilde{\beta}_{n}^{\ast}:=\sum_{m=1}^{M}\omega_{m}\tilde{\beta}_{m,n}^{\ast} $, where $\tilde{\beta}_{m,n}^{\ast}:=S_{xx.Z_{m}}^{-1}S_{xy^{\ast}.Z_{m}}$. Letting $T_{n}^{\ast}:=n^{1/2}(\tilde{\beta}_{n}^{\ast}-\hat{\beta}_{n})$, we can show that ((ref)) holds with $\hat{B}_{n}:=Q_{n}n^{1/2}\hat{\delta }_{n}$ such that, as in ((ref)), \[ \hat{B}_{n}-B_{n}=Q_{n}n^{1/2}(\hat{\delta}_{n}-\delta)\overset{d}{\rightarrow }\xi_{2}\sim N(0,v_{22}),\quad v_{22}>0, \] given in particular the asymptotic normality of $n^{1/2}(\hat{\delta} _{n}-\delta)$. Because the bias term in the bootstrap world is random in the limit, the conditional distribution of $T_{n}^{\ast}$ is also random in the limit, and in particular does not mimic the asymptotic distribution of the original statistic $T_{n}$.
\noindentPairs bootstrap. Consider now a pairs (random design) bootstrap sample $\{y_{t}^{\ast},x_{t}^{\ast},z_{t}^{\ast};t=1,\dots,n\}$, based on resampling with replacement from the tuples $\{y_{t},x_{t} ,z_{t};t=1,\dots,n\}$. As is standard, it is useful to recall that the bootstrap data have the representation \[ y^{\ast}=x^{\ast}\hat{\beta}_{n}+Z^{\ast}\hat{\delta}_{n}+\varepsilon^{\ast}, \] where $\varepsilon^{\ast}=(\varepsilon_{1}^{\ast},\ldots,\varepsilon_{n} ^{\ast})^{\prime}$ and $\varepsilon_{t}^{\ast}$ is an i.i.d.\ draw from $\hat{\varepsilon}_{t}=y_{t}-x_{t}\hat{\beta}_{n}-z_{t}^{\prime}\hat{\delta }_{n}$. The pairs bootstrap model averaging estimator is \[ \tilde{\beta}_{n}^{\ast}:=\sum_{m=1}^{M}\omega_{m}\tilde{\beta}_{m,n}^{\ast }\text{ with }\tilde{\beta}_{m,n}^{\ast}:=S_{x^{\ast}x^{\ast}.Z_{m}^{\ast} }^{-1}S_{x^{\ast}y^{\ast}.Z_{m}^{\ast}} \] and $Z_{m}^{\ast}=Z^{\ast}R_{m}$. The pairs bootstrap statistic is then \[ T_{n}^{\ast}:=n^{1/2}(\tilde{\beta}_{n}^{\ast}-\hat{\beta}_{n})=B_{n}^{\ast }+n^{1/2}S_{x^{\ast}x^{\ast}}^{-1}S_{x^{\ast}\varepsilon^{\ast}}, \] where \[ B_{n}^{\ast}:=\sum_{m=1}^{M}\omega_{m}S_{x^{\ast}x^{\ast}.Z_{m}^{\ast}} ^{-1}S_{x^{\ast}Z^{\ast}.Z_{m}^{\ast}}n^{1/2}\hat{\delta}_{n}. \] Therefore, and in contrast with the fixed regressor bootstrap (FRB), the term $B_{n}^{\ast}$ is stochastic under the bootstrap probability measure and replaces the bias term$~\hat{B}_{n}$. This difference is not innocuous because it implies that $T_{n}^{\ast}-\hat{B}_{n}$ no longer replicates the asymptotic distribution of $T_{n}-B_{n}$ and ((ref)) does not hold. However, this does not prevent our method from working, but it will require a different set of conditions which we will give in Section (ref).
\noindentSetup. We consider estimation of a vector of regression parameters through regularization; in particular, by using a ridge estimator. The model is $y_{t}=\theta^{\prime}x_{t}+\varepsilon_{t}$, $t=1,\dots,n$, where $x_{t}$ is a $p\times1$ non-stochastic vector and $\varepsilon_{t}\sim ~$i.i.d.$(0,\sigma^{2})$. Interest is on testing $\mathsf{H}_{0}:g^{\prime }\theta=r$, based on ridge estimation of $\theta$. Specifically, the ridge estimator has closed form expression $\tilde{\theta}_{n}=\tilde{S}_{xx} ^{-1}S_{xy}$, where $\tilde{S}_{xx}:=S_{xx}+n^{-1}c_{n}I_{p}$ and $c_{n}$ is a tuning parameter that controls the degree of shrinkage towards zero. Clearly, $c_{n}=0$ corresponds to the OLS\ estimator, $\hat{\theta}_{n}$. We are interested in the case where the regressors have limited explanatory power, i.e.,\ where $\theta=\delta n^{-1/2}$ is local to zero, which can in fact be taken as a motivation for shrinkage towards zero and hence for ridge estimation. To test $\mathsf{H}_{0}$, we consider the test statistic $T_{n}=n^{1/2}(g^{\prime}\tilde{\theta}_{n}-r)$. If $n^{-1}c_{n}\rightarrow c_{0}\geq0$ (as in, e.g., Fu and Knight, 2000) then, under the null, it holds that $T_{n}-B_{n}\rightarrow_{d}\xi_{1}\sim N(0,v^{2})$, where \[ B_{n}:=-c_{n}n^{-1/2}g^{\prime}\tilde{S}_{xx}^{-1}\theta=-c_{n}n^{-1} g^{\prime}\tilde{S}_{xx}^{-1}\delta\rightarrow B:=-c_{0}g^{\prime} \tilde{\Sigma}_{xx}^{-1}\delta \] with $\tilde{\Sigma}_{xx}:=\Sigma_{xx}+c_{0}I_{p}$ and $\Sigma_{xx}:=\lim S_{xx}$. Hence, for $c_{0}>0$, $\tilde{\theta}_{n}$ is asymptotically biased and the bias term cannot be consistently estimated. Consequently, ((ref)) is satisfied, and inference based on the quantiles of the $N(0,v^{2})$ distribution is invalid unless $c_{0}=0$.
\noindentBootstrap. Consider a pairs (random design) bootstrap sample $\{y_{t}^{\ast},x_{t}^{\ast};t=1,\dots,n\}$ built by i.i.d.\ resampling from the tuples $\{y_{t},x_{t};t=1,\dots,n\}$. The bootstrap analogue of the ridge estimator is $\tilde{\theta}_{n}^{\ast}:=\tilde{S}_{x^{\ast}x^{\ast}} ^{-1}S_{x^{\ast}y^{\ast}}$, where $\tilde{S}_{x^{\ast}x^{\ast}}:=S_{x^{\ast }x^{\ast}}+n^{-1}c_{n}I_{p}$. The bootstrap statistic is $T_{n}^{\ast }:=n^{1/2}g^{\prime}(\tilde{\theta}_{n}^{\ast}-\hat{\theta}_{n})$, which is centered using $\hat{\theta}_{n}$ to guarantee that $\varepsilon_{t}^{\ast}$ and $x_{t}^{\ast}$ are uncorrelated in the bootstrap world. Because we have used a pairs bootstrap, we now have $T_{n}^{\ast}-B_{n}^{\ast} \overset{d}{\rightarrow}_{p^{\ast}}\xi_{1}$ for $B_{n}^{\ast}:=-c_{n} n^{-1/2}g^{\prime}\tilde{S}_{x^{\ast}x^{\ast}}^{-1}\hat{\theta}_{n}$. However, $B_{n}^{\ast}-\hat{B}_{n}=o_{p^{\ast}}(1)$ with $\hat{B}_{n}:=-c_{n} n^{-1/2}g^{\prime}\tilde{S}_{xx}^{-1}\hat{\theta}_{n}$, such that $T_{n} ^{\ast}-\hat{B}_{n}$ still satisfies ((ref)). Then ((ref)) holds with \[ \hat{B}_{n}-B_{n}=-c_{n}n^{-1}g^{\prime}\tilde{S}_{xx}^{-1}n^{1/2}(\hat {\theta}_{n}-\theta)\overset{d}{\rightarrow}\xi_{2}\sim N(0,v_{22}),\quad v_{22}>0, \] so the bootstrap fails to approximate the asymptotic distribution of $T_{n}$ (see also Chatterjee and Lahiri, 2010, 2011).
\noindentSetup. Consider the model
where $\beta(\cdot)$ is a smooth function and $\varepsilon_{t}\sim $ i.i.d.$(0,\sigma^{2})$. For simplicity, we consider a fixed-design model; i.e., $x_{t}=t/n$. The goal is inference on $\beta(x)$ for a fixed $x\in (0,1)$. We apply the standard Nadaraya-Watson (fixed-design) estimator $\hat{\beta}_{h}(x)=(nh)^{-1}\sum_{t=1}^{n}K((x_{t}-x)/h)y_{t}$, where $h=cn^{-1/5}$ for some $c>0$ is the MSE-optimal bandwidth and $K$ is the kernel function. We do not consider the more general local polynomial regression case, although we conjecture that very similar results will hold. We leave that case for future research. The statistic $T_{n}=(nh)^{1/2} (\hat{\beta}_{h}(x)-\beta(x))$ satisfies $T_{n}-B_{n}\rightarrow_{d}\xi _{1}\sim N(0,v^{2})$, where $v^{2}:=\sigma^{2}\int K(u)^{2}du>0$ and
with $k_{t}:=K((x_{t}-x)/h)$. The bias $B_{n}$ satisfies
where $\kappa_{2}:=\int u^{2}K(u)du$ and $\beta^{\prime\prime}(x)\ $denotes the second-order derivative of $\beta(x)$. Thus, ((ref)) is satisfied. Estimating $B$ or $B_{n}$ is challenging because it involves estimating $\beta ^{\prime\prime}(x)$, and although theoretically valid estimators exist, they perform poorly in finite samples. This issue is pointed out by Calonico, Cattaneo, and Titunik (2014) and Calonico, Cattaneo, and Farrell (2018), who propose more accurate bias correction techniques specifically for regression discontinuity designs and nonparametric curve estimation.
\noindentBootstrap. The (parametric) bootstrap sample is generated as $y_{t}^{\ast}=\hat{\beta}_{h}(x_{t})+\varepsilon_{t}^{\ast}$, $t=1,\dots ,n$,\ where $\varepsilon_{t}^{\ast}|D_{n}\sim$ i.i.d.$N(0,\hat{\sigma}_{n} ^{2})$ with $D_{n}=\{y_{t},t=1,\dots,n\}$ and $\hat{\sigma}_{n}^{2}$ denotes a consistent estimator of $\sigma^{2}$; e.g.\ the residual variance. Let $\hat{\beta}_{h}^{\ast}(x)=(nh)^{-1}\sum_{t=1}^{n}k_{t}y_{t}^{\ast}$ and $T_{n}^{\ast}=(nh)^{1/2}(\hat{\beta}_{h}^{\ast}(x)-\hat{\beta}_{h}(x))$. Then ((ref)) is satisfied with \[ \hat{B}_{n}:=(nh)^{1/2}\left( \frac{1}{nh}\sum_{t=1}^{n}k_{t}\hat{\beta} _{h}(x_{t})-\hat{\beta}_{h}(x)\right) . \] Because $h=cn^{-1/5}$, ((ref)) holds with \[ \hat{B}_{n}-B_{n}=(nh)^{1/2}\left( \frac{1}{nh}\sum_{t=1}^{n}k_{t}(\hat {\beta}_{h}(x_{t})-\beta(x_{t}))-(\hat{\beta}_{h}(x)-\beta(x))\right) \overset{d}{\rightarrow}\xi_{2}\sim N(0,v_{22}), \] where $v_{22}>0$, so the bootstrap is invalid. Two possible solutions to this problem are to generate the bootstrap sample as $y_{t}^{\ast} =\hat{\beta}_{g}(x_{t})+\varepsilon_{t}^{\ast}$, where $g$ is an oversmoothing bandwidth satisfying $ng^{5}\rightarrow\infty$ (e.g., H\"{a}rdle and Marron, 1991) or to center the bootstrap statistic at its expected value and add a consistent estimator of $B$ (e.g., H\"{a}rdle and Bowman, 1988; Eubank and Speckman, 1993). Both approaches require selecting two bandwidths, which is not straightforward. An alternative approach suggested by Hall and Horowitz (2013) focuses on an asymptotic theory-based confidence interval and applies the bootstrap to calibrate its coverage probability. However, this requires an additional averaging step across a grid of $x$ (their step 6) to asymptotically eliminate\ $\xi_{2}$, and it results in an asymptotically conservative interval. Finally, a non-bootstrap-based solution is undersmoothing using a bandwidth $h$ satisfying $nh^{5}\rightarrow0$, although of course that is not MSE-optimal and may result in trivial power against certain local alternatives; see Section (ref).
The general framework is as follows. We have a statistic $T_{n}$ defined as a general function of a sample $D_{n}$, for which we would like to compute a valid bootstrap p-value. Usually $T_{n}$ is a test statistic or a (possibly normalized) parameter estimator; for example, $T_{n}=g(n)(\hat{\theta} _{n}-\theta_{0})$. Let $D_{n}^{\ast}$ denote the bootstrap sample, which depends on the original data and on some auxiliary bootstrap variates (which we assume defined jointly with $D_{n}$ on a possibly extended probability space). Let $T_{n}^{\ast}$ denote the bootstrap version of $T_{n}$ computed on $D_{n}^{\ast}$; for example, $T_{n}^{\ast}=g(n)(\hat{\theta}_{n}^{\ast} -\hat{\theta}_{n})$. Let $\hat{L}_{n}(u):=P^{\ast}(T_{n}^{\ast}\leq u)$, $u\in\mathbb{R}$, denote its distribution function, conditional on the original data. The bootstrap p-value is defined as \[ \hat{p}_{n}:=P^{\ast}(T_{n}^{\ast}\leq T_{n})=\hat{L}_{n}(T_{n}). \]
First-order asymptotic validity of $\hat{p}_{n}$ requires that $\hat{p}_{n}$ converges in distribution to a standard uniform distribution; i.e.,\ that $\hat{p}_{n}\rightarrow_{d}U_{[0,1]}$. In this section we focus on a class of statistics $T_{n}$ and $T_{n}^{\ast}$ for which this condition is not necessarily satisfied. The main reason is the presence of an additive `bias'\ term $B_{n}$ that contaminates the distribution of $T_{n}$ and cannot be replicated by the bootstrap distribution of $T_{n}^{\ast}$.
When $B_{n}$ converges to a nonzero constant $B$, Assumption (ref) can be written $T_{n}\rightarrow_{d}B+\xi_{1}$ as in ((ref)). If $T_{n}$ is a normalized version of a (scalar) parameter estimator, i.e.,\ $T_{n} =g(n)(\hat{\theta}_{n}-\theta_{0})$, then we can think of $B$ as the asymptotic bias of\ $\hat{\theta}_{n}$ because $\xi_{1}$ is centered at zero. Although we allow for the possibility that $B_{n}$ does not have a limit (and it may even diverge), we will still refer to $B_{n}$ as a `bias term'. More generally, in Assumption (ref) we cover any statistic $T_{n}$ that is not necessarily Gaussian (even asymptotically) and whose limiting distribution is $G$ only after we subtract the sequence $B_{n}$. The limiting distribution $G$ may depend on a parameter such that $T_{n}-B_{n}$ is not an asymptotic pivot.
Inference based on the asymptotic distribution of $T_{n}$ requires estimating $B_{n}$ and any parameter in $G$. Alternatively, we can use the bootstrap to bypass parameter estimation and directly compute a bootstrap p-value that relies on $T_{n}^{\ast}$ and $T_{n}$ alone; that is, we consider $\hat{p} _{n}:=P^{\ast}(T_{n}^{\ast}\leq T_{n})$. A set of high-level conditions on $T_{n}^{\ast}$ and $T_{n}$ that allow us to derive the asymptotic properties of this p-value are described next.
Assumption (ref)(i) states that $T_{n}^{\ast}-\hat{B}_{n}$ converges in distribution to a random variable $\xi_{1}$ having the same distribution function $G$ as $T_{n}-B_{n}$.\footnote{Note that we write $T_{n}^{\ast} -\hat{B}_{n}\overset{d^{\ast}}{\rightarrow}_{p}\xi_{1}$ to mean that $T_{n}^{\ast}-\hat{B}_{n}$ has (conditionally on $D_{n}$) the same asymptotic distribution function as the random variable $\xi_{1}$. We could alternatively write that $T_{n}^{\ast}-\hat{B}_{n}\overset{d^{\ast}}{\rightarrow}_{p}\xi _{1}^{\ast}$ and $T_{n}-B_{n}\overset{d}{\rightarrow}\xi_{1}$ where $\xi _{1}^{\ast}$ and $\xi_{1}$ are two independent copies of the same distribution, i.e.\ $P(\xi_{1}\leq u)=$ $P(\xi_{1}^{\ast}\leq u)$. We do not make this distinction because we care only about distributional results, but it should be kept in mind.} Thus, $\hat{B}_{n}$ can be thought of as an implicit bootstrap bias that affects the statistic $T_{n}^{\ast}$, in the same way that $B_{n}$ affects the original statistic $T_{n}$. Assumption (ref)(ii) complements Assumption (ref) by requiring the joint convergence of $T_{n}-B_{n}$ and $\hat{B}_{n}-B_{n}$ towards $\xi_{1}$ and $\xi_{2}$, respectively; see also ((ref))--((ref)).
Given Assumption (ref)(i), we could use the bootstrap distribution of $T_{n}^{\ast}-\hat{B}_{n}$ to approximate the distribution of $T_{n}-B_{n}$. Since $B_{n}$ is typically unknown, this result is not very useful for inference unless $\hat{B}_{n}$ is consistent for $B_{n}$. In this case, Assumption (ref) together with Assumption (ref) imply that $\hat{p}_{n}$ is asymptotically distributed as $U_{[0,1]}$. This follows by noting that if $\hat{B}_{n}-B_{n}=o_{p}(1)$, then $\xi_{2}=0$ a.s., implying that $F(u)=G(u)$. Consequently,
where the last distributional equality holds by $F=G$\ and the probability integral transform. However, this result does not hold if $\hat {B}_{n}-B_{n}$ does not converge to zero in probability. Specifically, if $\hat{B}_{n}-B_{n}\rightarrow_{d}\xi_{2}$ (jointly with $T_{n}-B_{n} \rightarrow_{d}\xi_{1}$), then \[ T_{n}-\hat{B}_{n}=(T_{n}-B_{n})-(\hat{B}_{n}-B_{n})\overset{d}{\rightarrow} \xi_{1}-\xi_{2}\sim F^{-1}(U_{[0,1]}) \] under Assumptions (ref) and (ref)(ii). When $\xi_{2}$ is nondegenerate, $F\neq G$, implying that $\hat{p}_{n}=G(T_{n}-\hat{B} _{n})+o_{p}(1)$ is not asymptotically distributed as a standard uniform random variable. This result is summarized in the following theorem.
\noindentProof.\ First notice that $\hat{p}_{n}$ and $G(T_{n}-\hat {B}_{n})$ have the same asymptotic distribution because Assumption (ref)(i) and continuity of $G$ imply that, by Polya's Theorem, \[ |\hat{p}_{n}-G(T_{n}-\hat{B}_{n})|\leq\sup_{u\in\mathbb{R}}|P^{\ast} (T_{n}^{\ast}-\hat{B}_{n}\leq u)-G(u)|\overset{p}{\rightarrow}0. \] Next, by Assumption (ref)(ii), $T_{n}-\hat{B}_{n}\rightarrow_{d} \xi_{1}-\xi_{2}$, such that \[ G(T_{n}-\hat{B}_{n})\overset{d}{\rightarrow}G(\xi_{1}-\xi_{2}) \] by continuity of $G$ and the continuous mapping theorem. Since $\xi_{1} -\xi_{2}$ has continuous cdf $F$, it holds that $\xi_{1}-\xi_{2}\sim F^{-1}(U_{[0,1]})$, which completes the proof.$\hfill\square$
Next, we describe two possible solutions to the invalidity of the standard bootstrap p-value $\hat{p}_{n}$. One relies on the prepivoting approach of Beran (1987, 1988); see Section (ref). The basic idea is that we modify $\hat{p}_{n}$ by applying the mapping $\hat{p}_{n}\mapsto H(\hat{p} _{n})$, where $H(u)$ is the asymptotic cdf of $\hat{p}_{n}$, which makes the modified p-value $H(\hat{p}_{n})$ asymptotically standard uniform. Contrary to Beran (1987, 1988), who proposed prepivoting as a way of providing asymptotic refinements for the bootstrap, here we show how to use prepivoting to solve the invalidity of the standard bootstrap p-value $\hat{p}_{n}$. This result is new in the bootstrap literature. The second approach relies on computing a standard bootstrap p-value based on the modified statistic given by $T_{n}-\hat{B}_{n}$; see Section (ref). Thus, we modify the test statistic rather than modifying the way we compute the bootstrap p-value.
Theorem (ref) implies that \[ P(\hat{p}_{n}\leq u)\rightarrow P(G(F^{-1}(U_{[0,1]}))\leq u)=P(U_{[0,1]}\leq F(G^{-1}(u)))=F(G^{-1}(u))=:H(u) \] uniformly over $u\in\lbrack0,1]$ by Polya's Theorem, given the continuity of $G$ and $F$. Although $H$ is not the uniform distribution, unless $G=F$, it is continuous because $G$ is strictly increasing. Thus, the following corollary to Theorem (ref) holds by the probability integral transform.
Therefore, the mapping of $\hat{p}_{n}$ into $H(\hat{p}_{n})$ transforms $\hat{p}_{n}$ into a new p-value, $H(\hat{p}_{n})$, whose asymptotic distribution is the standard uniform distribution on $[0,1]$. Inference based on $H(\hat{p}_{n})$ is generally infeasible, because we do not observe $H(u)$. However, if we can replace $H(u)$ with a uniformly consistent estimator $\hat{H}_{n}(u)$ then this approach will deliver a feasible modified p-value $\tilde{p}_{n}:=\hat{H}_{n}(\hat{p}_{n})$. Since the limit distribution of $\tilde{p}_{n}$ is the standard uniform distribution, $\tilde{p}_{n}$ is an asymptotically valid p-value. The mapping of $\hat{p}_{n}$ into $\tilde{p} _{n}=\hat{H}_{n}(\hat{p}_{n})$ by the estimated distribution of the former corresponds to what Beran (1987) calls `prepivoting'. In the following sections, we describe two methods of obtaining a consistent estimator of $H(u)$.
Suppose $H(u)=H_{\gamma}(u)$ depends on a finite-dimensional parameter, $\gamma$. In view of Theorem (ref), a simple approach to estimating $H(u)$ is to use \[ \hat{H}_{n}(u)=H_{\hat{\gamma}_{n}}(u), \] where $\hat{\gamma}_{n}$ denotes a consistent estimator of $\gamma$. This leads to a plug-in modified p-value defined as \[ \tilde{p}_{n}=H_{\hat{\gamma}_{n}}(\hat{p}_{n}). \] By consistency of $\hat{\gamma}_{n}$ and under the assumption that $H_{\gamma }$ is continuous in $\gamma$, it follows immediately that \[ \tilde{p}_{n}=H(\hat{p}_{n})+o_{p}(1)\overset{d}{\rightarrow}F(G^{-1} (G(F^{-1}(U_{[0,1]}))))=U_{[0,1]}. \] This result is summarized next.
The plug-in approach relies on a consistent estimator of the asymptotic distribution $H$, but does not require estimating the `bias term' $B_{n}$. When estimating $\gamma$ is simple, this approach is attractive since it does not require any double resampling. Examples are given in Section (ref). However, computation of $\gamma$ is case-specific and may be cumbersome in practice. An automatic approach is to use the bootstrap to estimate $H(u)$, as we describe next.
Following Beran (1987, 1988), we can estimate $H(u)$ with the bootstrap. That is, we let \[ \hat{H}_{n}(u)=P^{\ast}(\hat{p}_{n}^{\ast}\leq u), \] where $\hat{p}_{n}^{\ast}$ is the bootstrap analogue of $\hat{p}_{n}$. Since $\hat{p}_{n}$ is itself a bootstrap p-value, computing $\hat{p}_{n}^{\ast}$ requires a double bootstrap. In particular, let $D_{n}^{\ast\ast}$ denote a further bootstrap sample of size $n$ based on $D_{n}^{\ast}$ and some additional bootstrap variates (defined jointly with $D_{n}$ and $D_{n}^{\ast}$ on a possibly extended probability space), and let $T_{n}^{\ast\ast}$ denote the bootstrap version of $T_{n}^{\ast}$ computed on $D_{n}^{\ast\ast}$. With this notation, the second-level bootstrap p-value is defined as \[ \hat{p}_{n}^{\ast}:=P^{\ast\ast}(T_{n}^{\ast\ast}\leq T_{n}^{\ast}), \] where $P^{\ast\ast}$ denotes the bootstrap probability measure conditional on $D_{n}^{\ast}$ and $D_{n}$ (making $\hat{p}_{n}^{\ast}$ a function of $D_{n}^{\ast}$ and $D_{n}$). This leads to a double bootstrap modified p-value, as given by \[ \tilde{p}_{n}:=\hat{H}_{n}(\hat{p}_{n})=P^{\ast}(\hat{p}_{n}^{\ast}\leq\hat {p}_{n}). \]
In order to show that $\tilde{p}_{n}=\hat{H}_{n}(\hat{p}_{n})\rightarrow _{d}U_{[0,1]}$, we add the following assumption.
Assumption (ref) complements Assumptions (ref) and (ref) by imposing high-level conditions on the second-level bootstrap statistics. Specifically, Assumption (ref)(i) assumes that $T_{n}^{\ast\ast}$ has asymptotic distribution $G$ only after we subtract $\hat{B}_{n}^{\ast}$. This term is the second-level bootstrap analogue of $\hat{B}_{n}$. It depends only on the first-level bootstrap data $D_{n}^{\ast}$ and is not random under $P^{\ast\ast}$. The second part of Assumption (ref) follows from Assumption (ref) in the special case that $\hat{B}_{n}^{\ast}-\hat{B}_{n}=o_{p^{\ast}}(1)$, in probability; i.e., when $\xi_{2}=0$ a.s., implying $F=G$. When $F\neq G$, $\hat{B} _{n}^{\ast}$ is not a consistent estimator of $\hat{B}_{n}$. However, under Assumption (ref), \[ T_{n}^{\ast}-\hat{B}_{n}^{\ast}\,=(T_{n}^{\ast}-\hat{B}_{n})-(\hat{B} _{n}^{\ast}-\hat{B}_{n})\overset{d^{\ast}}{\rightarrow}_{p}\xi_{1}-\xi _{2}=F^{-1}(U_{[0,1]}) \] implying that $T_{n}^{\ast}-\hat{B}_{n}^{\ast}\,$\ mimics the distribution of $T_{n}-\hat{B}_{n}$. This suffices for proving the asymptotic validity of the double bootstrap modified p-value, $\tilde{p}_{n}=\hat{H}_{n}(\hat{p}_{n})$, as proved next.
\noindentProof.\ To prove this result, recall that $\hat{H} _{n}(u)=P^{\ast}(\hat{p}_{n}^{\ast}\leq u)$ and $P(\hat{p}_{n}\leq u)\rightarrow H(u)=F(G^{-1}(u))$ uniformly in $u\in\mathbb{R}$, since $H$ is a continuous distribution function by Assumptions (ref) and (ref). Thus, we have that
where $G(F^{-1}(U_{[0,1]}))$ is a random variable whose distribution function is $H$. Hence, \[ \sup_{u\in\mathbb{R}}|\hat{H}_{n}(u)-H(u)|=o_{p}(1). \] Since $H(\hat{p}_{n})\rightarrow_{d}U_{[0,1]}$, we can conclude that $\tilde{p}_{n}=\hat{H}_{n}(\hat{p}_{n})\rightarrow_{d}U_{[0,1]}$ .$\hfill\square$
Theorem (ref) shows that prepivoting the standard bootstrap p-value $\hat{p}_{n}$ by applying the mapping $\hat{H}_{n}$ transforms it into an asymptotically uniformly distributed random variable. This result holds under Assumptions (ref), (ref), and (ref), independently of whether $G=F$ or not. When $G=F$ then $\hat{p}_{n}\rightarrow_{d}U_{[0,1]}$ (as implied by Theorem (ref)). In this case, the prepivoting approach is not necessary to obtain a first-order asymptotically valid test, although it might help further reducing the size distortion of the test. This corresponds to the setting of Beran (1987, 1988), where prepivoting was proposed as a way of reducing the level distortions of confidence intervals. When $G\neq F$ then $\hat{p}_{n}$ is not asymptotically uniform and a standard bootstrap test based on $\hat{p}_{n}$ is asymptotically invalid, as shown in Theorem (ref). In this case, prepivoting transforms an asymptotically invalid bootstrap p-value into one that is asymptotically valid. This setting was not considered by Beran (1987, 1988) and is new to our paper.
In this section we explicitly consider a testing situation. Suppose we are interested in testing $\mathsf{H}_{0}:\theta=\bar{\theta}$ against $\mathsf{H}_{1}:\theta<\bar{\theta}$. Specifically, defining $T_{n} (\theta):=g(n)(\hat{\theta}_{n}-\theta)$, we consider the test statistic $T_{n}(\bar{\theta})$. The corresponding bootstrap p-value is $\hat{p} _{n}(\bar{\theta})$ with $\hat{p}_{n}(\theta):=P^{\ast}(T_{n}^{\ast}\leq T_{n}(\theta))$. When the null hypothesis is true, i.e., when $\bar{\theta }=\theta_{0}$ with $\theta_{0}$ denoting the true value, we find $T_{n} (\bar{\theta})=T_{n}(\theta_{0})=T_{n}$ and $\hat{p}_{n}(\bar{\theta})=\hat {p}_{n}(\theta_{0})=\hat{p}_{n}$, where $T_{n}$ and $\hat{p}_{n}$ are as defined previously. If Assumptions (ref) and (ref) hold under the null, Theorem (ref) and Corollary (ref) imply that tests based on $H(\hat{p}_{n}(\bar{\theta}))$ have correct asymptotic size, where $H$ continues to denote the asymptotic cdf of $\hat{p}_{n}$.
To analyze power, we consider $\theta_{0}=\bar{\theta}+a_{n}$ for some deterministic sequence $a_{n}$. Then $a_{n}=0$ under the null hypothesis, whereas $a_{n}=a<0$ corresponds to a fixed alternative and $a_{n}=a/g(n)$ for $a<0$ corresponds to a local alternative. Thus, we define $\pi_{n} :=g(n)(\theta_{0}-\bar{\theta})=g(n)a_{n}$ so that $T_{n}(\bar{\theta} )=T_{n}+\pi_{n}$.
\noindentProof.\ As in the proof of Theorem (ref) we have, by Assumption (ref)(i), \[ \hat{p}_{n}(\bar{\theta})=P^{\ast}(T_{n}^{\ast}\leq T_{n}(\bar{\theta }))=P^{\ast}(T_{n}^{\ast}-\hat{B}_{n}\leq T_{n}-\hat{B}_{n}+\pi_{n} )=G(T_{n}-\hat{B}_{n}+\pi_{n})+o_{p}(1). \] If $\pi_{n}\rightarrow\pi$ then $\hat{p}_{n}(\bar{\theta})\rightarrow _{d}G(F^{-1}(U_{[0,1]})+\pi)$ by Assumption (ref)(ii), so that \[ H(\hat{p}_{n}(\bar{\theta}))\overset{d}{\rightarrow}H(G(F^{-1}(U_{[0,1]} )+\pi))=F(F^{-1}(U_{[0,1]})+\pi) \] by definition of $H(u)$. If $\pi_{n}\rightarrow-\infty$ then $\hat{p}_{n} (\bar{\theta})\rightarrow_{p}0$ because $T_{n}-\hat{B}_{n}=O_{p}(1)$ by Assumption (ref)(ii), so that $H(\hat{p}_{n}(\bar{\theta} ))\rightarrow_{p}H(0)=0$ and $P(H(\hat{p}_{n}(\bar{\theta}))\leq \alpha)\rightarrow1$ for any $\alpha>0$.$\hfill\square\medskip$
It follows from Theorem (ref)(ii) that a left-tailed test that rejects for small values of $H(\hat{p}_{n}(\bar{\theta}))$ is consistent. Furthermore, it follows from Theorem (ref)(i) that such a test has non-trivial asymptotic local power against $\pi<0$. Specifically, the asymptotic local power against $\pi$ is given by $P(H(\hat{p}_{n}(\bar{\theta }))\leq\alpha)\rightarrow F(F^{-1}(\alpha)-\pi)$. Interestingly, this only depends on $F$ and not on $G$. As above, to implement the modified p-value, $H(\hat{p}_{n}(\bar{\theta}))$, in practice, we would need a (uniformly) consistent estimator of $H$, i.e., the asymptotic distribution of the bootstrap p-value when the null hypothesis is true.\ This could be either the plug-in or double bootstrap estimators, as discussed in Sections (ref) and (ref).
Note that Assumption (ref) is still assumed to hold in Theorem (ref). That is, the bootstrap statistic $T_{n}^{\ast}$ is assumed to have the same asymptotic behavior under the null and under the alternative. This is commonly the case when the bootstrap algorithm does not impose the null hypothesis when generating the bootstrap data.
The double bootstrap modified p-value $\tilde{p}_{n}$ depends only on the statistic $T_{n}$ and their bootstrap analogues $T_{n}^{\ast}$ and $T_{n} ^{\ast\ast}$. It does not involve computing explicitly $\hat{B}_{n}$ or $\hat{B}_{n}^{\ast}$, but in some applications it can be computationally costly as it requires two levels of resampling. As it turns out, $\tilde {p}_{n}$ is asymptotically equivalent to a single-level bootstrap p-value that is based on bootstrapping the statistic $T_{n}-\hat{B}_{n}$, as we show next.
By definition, the double bootstrap modified p-value is given by $\tilde {p}_{n}:=P^{\ast}(\hat{p}_{n}^{\ast}\leq\hat{p}_{n})$, where \[ \hat{p}_{n}^{\ast}:=P^{\ast\ast}(T_{n}^{\ast\ast}\leq T_{n}^{\ast} )=P^{\ast\ast}(T_{n}^{\ast\ast}-\hat{B}_{n}^{\ast}\leq T_{n}^{\ast}-\hat {B}_{n}^{\ast})=G(T_{n}^{\ast}-\hat{B}_{n}^{\ast})+o_{p^{\ast}}(1), \] in probability, given Assumption (ref). Similarly, under Assumptions (ref) and (ref), \[ \hat{p}_{n}:=P^{\ast}(T_{n}^{\ast}\leq T_{n})=P^{\ast}(T_{n}^{\ast}-\hat {B}_{n}\leq T_{n}-\hat{B}_{n})=G(T_{n}-\hat{B}_{n})+o_{p}(1). \] It follows that
because $G$ is continuous. We summarize this result in the following corollary.
Theorem (ref) shows that $\tilde{p}_{n}\rightarrow _{d}U_{[0,1]}$ and hence is asymptotically valid. In view of this, Corollary (ref) shows that removing $\hat{B}_{n}$ from $T_{n}$ and computing a bootstrap p-value based on the new statistic, $T_{n}-\hat{B}_{n}$, also solves the invalidity problem of the standard bootstrap p-value, $\hat{p}_{n}=P^{\ast}(T_{n}^{\ast}\leq T_{n})$. Note that we do not require $\xi_{2}=0$, i.e.\ $\hat{B}_{n}-B_{n}$ and $\hat{B} _{n}^{\ast}-\hat{B}_{n}$ do not need to converge to zero.
When $\hat{B}_{n}$ and $\hat{B}_{n}^{\ast}$ are easy to compute, e.g.,\ when they are available analytically as functions of $D_{n}$ and $D_{n}^{\ast}$, respectively, Corollary (ref) is useful as it avoids implementing a double bootstrap. When this is not the case, i.e., when deriving $\hat{B}_{n}$ and $\hat{B}_{n}^{\ast}$ explicitly is cumbersome or impossible, we may be able to estimate $\hat{B}_{n}$ from the bootstrap and $\hat{B}_{n}^{\ast}$ from a double bootstrap. Corollary (ref) then shows that the double bootstrap modified p-value $\tilde{p}_{n}$ is a convenient alternative since it depends only on $T_{n}$, $T_{n}^{\ast}$, and $T_{n}^{\ast\ast}$. It is important to note that none of these approaches requires the consistency of $\hat{B}_{n}$ and $\hat{B}_{n}^{\ast}$.
We conclude this section by providing an alternative set of high-level conditions that cover bootstrap methods for which $T_{n}^{\ast}-\hat{B}_{n}$ has a different limiting distribution than $T_{n}-B_{n}$. This may happen, for example, for the pairs bootstrap; see Section (ref) and Remark (ref).
Under Assumption (ref), $T_{n}^{\ast}-\hat{B}_{n}$ does not replicate the distribution of $T_{n}-B_{n}$. This is to be understood in the sense that there does not exist a $P^{\ast}$-measurable term $\hat{B}_{n}$ such that $T_{n}^{\ast}-\hat{B}_{n}$ has the same asymptotic distribution as $T_{n}-B_{n}$.
An important generalization provided by Assumption (ref) compared with Assumption (ref) is to allow for bootstrap methods where the `centering term', say $B_{n}^{\ast}$, depends on the bootstrap data. That is, to allow cases where there is a random (with respect to $P^{\ast}$, i.e., depending on the bootstrap data) term $B_{n}^{\ast}$ such that $T_{n}^{\ast }-B_{n}^{\ast}\overset{d^{\ast}}{\rightarrow}_{p}\xi_{1}$ and hence has the same asymptotic distribution as $T_{n}-B_{n}$. Clearly, this violates Assumption (ref) unless $B_{n}^{\ast}-\hat{B}_{n}\overset{p^{\ast }}{\rightarrow}_{p}0$ (as in the ridge regression in Section (ref)).\ However, letting $\zeta_{1}$ be such that $B_{n}^{\ast}-\hat{B}_{n}\overset{d^{\ast}}{\rightarrow}_{p}\zeta_{1}-\xi_{1} $, then Assumption (ref) covers the former case.
The asymptotic distribution of the bootstrap p-value under Assumption (ref) is given in the following theorem. The proof is identical to that of Theorem (ref), with $G$ replaced by $J$, and hence omitted.
Theorem (ref) implies that now $P(\hat{p}_{n}\leq u)\rightarrow P(J(F^{-1}(U_{[0,1]}))\leq u)=F(J^{-1}(u))=:H(u)$. Clearly, a plug-in approach to estimating this $H(u)$ based on $G$ as described in Section (ref) would be invalid because $G\neq J$ in general. However, it follows straightforwardly by the same arguments as applied in Section (ref) that a plug-in approach based on $J$ will deliver an asymptotically valid plug-in modified p-value.
To implement an asymptotically valid double bootstrap modified p-value we consider the following high-level condition.
Under Assumption (ref), the second-level bootstrap statistic, $T_{n}^{\ast\ast}-\hat{B}_{n}^{\ast}$, replicates the distribution of the first-level statistic, $T_{n}^{\ast}-\hat{B}_{n}$. Thus, the second-level bootstrap p-value is
under Assumption (ref). Hence, the second-level bootstrap p-value has the same asymptotic distribution as the original bootstrap p-value. It follows that the double bootstrap modified p-value, $\tilde{p} _{n}:=\hat{H}_{n}(\hat{p}_{n})=P^{\ast}(\hat{p}_{n}^{\ast}\leq\hat{p}_{n})$, is asymptotically valid, which is stated next. The proof is essentially identical to that of Theorem (ref) and hence omitted.
In this section we revisit our three leading examples from Section (ref), where we argued that standard boostrap inference is invalid due to the presence of bias. In this section\ we show how to apply our general theory in each example. Again, we refer to Appendix (ref) for detailed derivations.
\noindentFixed regressor bootstrap. Extending the arguments in Section (ref), we obtain the following result.
By Lemma (ref), the conditions of Theorem (ref) hold with $G(u)=\Phi(u/v_{11})$ and $F(u)=\Phi(u/v_{d})$, where $v_{d}^{2}=v_{11} +v_{22}-2v_{12}>0$. Then Theorem (ref) implies that the standard bootstrap p-value satisfies $\hat{p}_{n}\rightarrow_{d}\Phi(m\Phi ^{-1}(U_{[0,1]}))$ with $m^{2}:=v_{d}^{2}/v^{2}$. Because $\omega$ is known and $\sigma^{2},\Sigma_{WW}$ are easily estimated, a consistent estimator $\hat{m}_{n}\rightarrow_{p}m$ is available, and the plug-in approach in Corollary (ref) can be implemented by considering the modified p-value, $\tilde{p}_{n}=\Phi(\hat{m}_{n}^{-1}\Phi^{-1}(\hat{p}_{n} ))$. Inspection of the proofs shows that our modified bootstrap approach is asymptotically valid whether $\delta$ is fixed or local-to-zero. In the former case, $B_{n}$ is $O_{p}(n^{1/2})$ rather than $O_{p}(1)$, implying that $B_{n}$ diverges in probability and $\tilde{\beta}_{n}$ is not even consistent for $\beta$. Despite this, the modified bootstrap p-value is asymptotically valid.
Alternatively, we can implement the double bootstrap as in Section (ref). Specifically, let \[ y^{\ast\ast}=x\hat{\beta}_{n}^{\ast}+Z\hat{\delta}_{n}^{\ast}+\varepsilon ^{\ast\ast}, \] where $\varepsilon^{\ast\ast}|\{D_{n},D_{n}^{\ast}\}\sim N(0,\hat{\sigma} _{n}^{\ast2}I_{n})$, $(\hat{\beta}_{n}^{\ast},\hat{\delta}_{n}^{\ast^{\prime} },\hat{\sigma}_{n}^{\ast2})\,$is the OLS estimator obtained from the full model estimated on the first-level bootstrap data, and $D_{n}^{\ast} =\{y^{\ast},W\}$. The double bootstrap statistic is $T_{n}^{\ast\ast} :=n^{1/2}(\tilde{\beta}_{n}^{\ast\ast}-\hat{\beta}_{n}^{\ast})$, where $\tilde{\beta}_{n}^{\ast\ast}:=\sum_{m=1}^{M}\omega_{m}\tilde{\beta} _{m,n}^{\ast\ast}$ with $\tilde{\beta}_{m,n}^{\ast\ast}:=S_{xx.Z_{m}} ^{-1}S_{xy^{\ast\ast}.Z_{m}}$ defined as the double bootstrap OLS estimator from the $m^{\text{th}}$ model. The double bootstrap modified p-value is then $\tilde{p}_{n}=P^{\ast}(\hat{p}_{n}^{\ast}\leq\hat{p}_{n})$ with $\hat{p} _{n}^{\ast}=P^{\ast\ast}(T_{n}^{\ast\ast}\leq T_{n}^{\ast})$.
Lemma (ref) shows that Assumption (ref) is verified in this example. The asymptotic validity of the double bootstrap modified p-value now follows from Lemmas (ref) and (ref) and Theorem (ref).
\noindentPairs bootstrap. For the pairs bootstrap we verify the high-level conditions in Section (ref). To simplify the discussion we consider the case with scalar $z_{t}$ in ((ref)) and where we \textquotedblleft average\textquotedblright \ over only one model ($M=1$), which is the simplest model in which $z_{t}$ is omitted from the regression. That is, we estimate $\beta$ by regression of $y$ on $x$, i.e., $\tilde{\beta}_{n}=S_{xx}^{-1}S_{xy}$. In this special case, $T_{n}-B_{n}\rightarrow_{d}N(0,v^{2})$ with $v^{2}=\sigma^{2}\Sigma_{xx}^{-1}$ and $B_{n}=S_{xx}^{-1}S_{xz}n^{1/2}\delta$.
Notice that, in contrast to the FRB, the asymptotic variance of $T_{n}^{\ast}$ fails to replicate that of $T_{n}$ because of the term $\kappa^{2}>0$. This implies that the methodology developed in Theorem (ref) and its corollaries no longer applies. Instead we can apply the theory of Section (ref). In particular, Lemma (ref) shows that Assumption (ref)(i) holds in this case with $\zeta_{1}\sim N(0,v^{2}+\kappa^{2})$. Lemma (ref) also shows that $\hat{B}_{n}$ is the same for the pairs bootstrap and the FRB, such that Lemma (ref) shows that Assumptions (ref) and (ref)(ii) are verified. This implies that Theorem (ref) holds for this example. Using similar arguments, it can be shown that Assumption (ref) also holds for this example, which would imply that the double bootstrap $p$-values are asymptotically uniformly distributed.
Under local alternatives of the form $\beta_{0}=\bar{\beta}+an^{-1/2}$, where $\bar{\beta}$ is the value under the null (Section (ref)), the asymptotic local power function for the modified p-value is given by $\Phi(\Phi^{-1}(\alpha)-a/v_{d})$; see Theorem (ref). It is not difficult to verify that this is the same power function as that obtained from a test based directly on $\hat{\beta}_{n}$ from the full model ((ref)).
To complete the example in Section (ref), we can proceed as in the previous example.
As in Section (ref), Lemma (ref) and Theorem (ref) imply that the standard bootstrap p-value satisfies $\hat{p}_{n}\rightarrow_{d}\Phi(m\Phi^{-1}(U_{[0,1]}))$, where we now have $m^{2}=(g^{\prime}\tilde{\Sigma}_{xx}^{-1}\Sigma_{xx}\tilde{\Sigma}_{xx} ^{-1}g)^{-1}g^{\prime}\Sigma_{xx}^{-1}g$. Note that this result holds irrespectively of $\theta$ being fixed or local to zero. Thus, the bootstrap is invalid unless $c_{0}=0$ which implies $m=1$. For the plug-in method, a simple consistent estimator of $m$ is given by $\hat{m}_{n}^{2}:=(g^{\prime }\tilde{S}_{xx}^{-1}S_{xx}\tilde{S}_{xx}^{-1}g)^{-1}g^{\prime}S_{xx}^{-1}g$, and inference based on the plug-in modified p-value $\tilde{p}_{n}=\Phi (\hat{m}_{n}^{-1}\Phi^{-1}(\hat{p}_{n}))$ is then asymptotically valid by Corollary (ref).
To implement the double bootstrap method, we can draw the double bootstrap sample $\{y_{t}^{\ast\ast},x_{t}^{\ast\ast};t=1,\dots,n\}$ as i.i.d.\ from $\{y_{t}^{\ast},x_{t}^{\ast};t=1,\dots,n\}$. Accordingly, the second-level bootstrap ridge estimator is $\tilde{\theta}_{n}^{\ast\ast}:=\tilde {S}_{x^{\ast\ast}x^{\ast\ast}}^{-1}S_{x^{\ast\ast}y^{\ast\ast}}$ with associated test statistic $T_{n}^{\ast\ast}:=n^{1/2}g^{\prime}(\tilde{\theta }_{n}^{\ast\ast}-\hat{\theta}_{n}^{\ast})$, which is centered at the first-level bootstrap OLS estimator, $\hat{\theta}_{n}^{\ast}$. It is straightforward to show that, without additional conditions, Assumption (ref) holds.
Validity of the double bootstrap modified p-value $\tilde{p}_{n}=P^{\ast} (\hat{p}_{n}^{\ast}\leq\hat{p}_{n})$ now follows by application of Theorem (ref).
Again, we complete the example in Section (ref) by proceeding as in the previous examples.
As before, Lemma (ref) and Theorem (ref) imply that the standard bootstrap p-value satisfies $\hat{p}_{n}\rightarrow_{d}\Phi (m\Phi^{-1}(U_{[0,1]}))$, where we now have $m^{2}:=4+(\int K^{2} (u)du)^{-1}(\int(\int K(s-u)K(s)ds)^{2}du-4\int K(u)\int K(u-s)K(s)dsdu)$. Thus, in this example, $m$ need not be estimated because it is observed once $K$ is chosen. Therefore, valid inference is feasible with the modified p-value $\tilde{p}_{n}=H(\hat{p}_{n})=\Phi(m^{-1}\Phi^{-1}(\hat{p}_{n}))$; see Corollary (ref).
We can also apply a double bootstrap modification. Let $y_{t}^{\ast\ast} =\hat{\beta}_{h}^{\ast}(x_{t})+\varepsilon_{t}^{\ast\ast}$, $t=1,\dots ,n$,\ where $\varepsilon_{t}^{\ast\ast}|\{D_{n},D_{n}^{\ast}\}\sim ~$i.i.d.$N(0,\hat{\sigma}_{n}^{\ast2})$ with $D_{n}^{\ast}:=\{y_{t}^{\ast },t=1,\dots,n\}$ and $\hat{\sigma}_{n}^{\ast2}$ denoting the residual variance from the first-level bootstrap data. The double bootstrap analogue of $T_{n}$ is $T_{n}^{\ast\ast}:=(nh)^{1/2}(\hat{\beta}_{h}^{\ast\ast}(x)-\hat{\beta} _{h}^{\ast}(x))$, where $\hat{\beta}_{h}^{\ast\ast}(x):=(nh)^{-1}\sum _{t=1}^{n}k_{t}y_{t}^{\ast\ast}$. This can be decomposed as $T_{n}^{\ast\ast }=\xi_{1,n}^{\ast\ast}+\hat{B}_{n}^{\ast}$, where $\hat{B}_{n}^{\ast }:=(nh)^{1/2}(({nh})^{-1}\sum_{t=1}^{n}k_{t}\hat{\beta}_{h}^{\ast}(x_{t} )-\hat{\beta}_{h}^{\ast}(x))$. Unfortunately, although $\xi_{1,n}^{\ast\ast}$ satisfies Assumption (ref)(i), $\hat{B}_{n}^{\ast}$ does not satisfy Assumption (ref)(ii). The reason is that $\hat{B}_{n}^{\ast}-\hat {B}_{n}=\xi_{2,n}^{\ast}+\hat{B}_{2,n}-\hat{B}_{n}$, where $\xi_{2,n}^{\ast}$ satisfies Assumption (ref)(ii), but $\hat{B}_{2,n}:=(nh)^{-1} \sum_{t=1}^{n}k_{t}\hat{B}_{n}(x_{t})$ is a smoothed version of $\hat{B}_{n}$ (evaluated at $x_{t}$) and although $\hat{B}_{2,n}-\hat{B}_{n}$ is mean zero it is not $o_{p}(1)$. However, $\hat{B}_{2,n}-\hat{B}_{n}$ is observed, so this is easily corrected by defining $\bar{T}_{n}^{\ast\ast}:=T_{n}^{\ast\ast }-(\hat{B}_{2,n}-\hat{B}_{n})$. Then we have the following result.
The validity of the double bootstrap modified p-value $\tilde{p}_{n}=P^{\ast }(\hat{p}_{n}^{\ast}\leq\hat{p}_{n})$, where $\hat{p}_{n}^{\ast}:=P^{\ast\ast }(\bar{T}_{n}^{\ast\ast}\leq T_{n}^{\ast})$, follows from Lemma (ref) and Theorem (ref). This in turn implies that confidence intervals based on the double bootstrap are asymptotically valid; see also Remark (ref). We note that Hall and Horowitz (2013) also proposed, without theory, a version of their calibration method based on the double bootstrap. Our double bootstrap-based method for confidence intervals corresponds to their steps 1--5, and where we need a correction they have instead a step 6 in which they average over a grid of $x$.
Finally, under local alternatives of the form $\beta_{0}(x)=\bar{\beta }+an^{-2/5}$, where $\bar{\beta}$ is the value under the null (Section (ref)), the asymptotic local power function for the modified p-value is given by $\Phi(\Phi^{-1}(\alpha)-a/v_{d})$; see Theorem (ref). Alternatively, we could consider a \textquotedblleft bias-free\textquotedblright\ test based on undersmoothing; that is using a bandwidth $h$ satisfying $nh^{5}\rightarrow0$ such that $B_{n}\rightarrow0$ and inference can be based on quantiles of $\xi_{1}\sim N(0,v_{11}^{2})$. In contrast to our procedure, however, such a test has only trivial power against $\bar{\beta}+an^{-2/5}$ because $(nh)^{1/2}an^{-2/5}\rightarrow0$.
In this paper, we have shown that in statistical problems involving bias terms that cannot be estimated, the bootstrap can be modified to provide asymptotically valid inference. Intuitively, the main idea is the following: in some important cases, the bootstrap can be used to `debias' a statistic whose bias is non-negligible, but when doing so additional `noise' is injected. This additional noise does not vanish because the bias cannot be consistently estimated, but it can be handled either by a `plug-in' method or by an additional (i.e., double) bootstrap layer. Specifically, our solution is simple and involves (i) focusing on the bootstrap p-value; (ii) estimating its asymptotic distribution; (iii) mapping the original (invalid)\ p-value into a new (valid)\ p-value using the prepivoting approach. These steps are easy to implement in practice and we provide sufficient conditions for asymptotic validity of the associated tests and confidence intervals.
Our results can be generalized in several directions. For instance, there is a growing literature where inference on a parameter of interest is combined with some auxiliary information in the form of a bound on the bias of the estimator in question. These bounds appear, e.g., in Oster (2019) and Li and M\"{u}ller (2021). It is of interest to investigate how our analysis can be extended in order to incorporate such bounds. Other possible extensions include non-ergodic problems, large-dimensional models, and multivariate estimators or statistics. All these extensions are left for future research.
We thank Federico Bandi, Matias Cattaneo, Christian Gourieroux, Philip Heiler, Michael Jansson, Anders Kock, Damian Kozbur, Marcelo Moreira, David Preinerstorfer, Mikkel S\o lvsten, Luke Taylor, Michael Wolf, and participants at the AiE Conference in Honor of Joon Y.\ Park, 2022 Conference on Econometrics and Business Analytics (CEBA), 2023 Conference on Robust Econometric Methods in Financial Econometrics, 2022 EC$^{2}$ conference, 2$^{\text{nd}}$ `High Voltage Econometrics' workshop, 2023 IAAE Conference, 3$^{\text{rd}}$ Italian Congress of Econometrics and Empirical Economics, 3$^{\text{rd}}$\ Italian Meeting on Probability and Mathematical Statistics, 19$^{\text{th}}$ School of Time Series and Econometrics, Brazilian Statistical Association, 2023 Soci\'{e}t\'{e} Canadienne de Sciences \'{E}conomiques, 2022 Virtual Time Series Seminars, as well as seminar participants at Aarhus University, CREST, FGV - Rio, FGV - S\ {a}o Paulo, Ludwig Maximilian University of Munich, Queen Mary University, Singapore Management University, UFRGS, University of the Balearic Islands, University of Oxford, University of Pittsburgh, University of Victoria, York University, for useful comments and feedback. Cavaliere thanks the the Italian Ministry of University and Research (PRIN 2017 Grant 2017TA7TYC) for financial support. Gon\c{c}alves thanks the Natural Sciences and Engineering Research Council of Canada for financial support (NSERC grant number RGPIN-2021-02663). Nielsen thanks the Danish National Research Foundation for financial support (DNRF Chair grant number DNRF154).