EconBase
← Back to paper

Bootstrap Diagnostic Tests

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

95,543 characters · 20 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
center[center omitted — 1,724 chars of source]

Violation of the assumptions underlying classical (Gaussian) limit theory often yields unreliable statistical inference. This paper shows that the bootstrap can detect such violations by delivering simple and powerful diagnostic tests that (a) induce no pre-testing bias, (b) use the same critical values across applications, and (c) are consistent against deviations from asymptotic normality. The tests compare the conditional distribution of a bootstrap statistic with the Gaussian limit implied by valid specification and assess whether the resulting discrepancy is large enough to indicate failure of the asymptotic Gaussian approximation. The method is computationally straightforward and only requires a sample of i.i.d. draws of the bootstrap statistic. We derive sufficient conditions for the randomness in the data to mix with the randomness in the bootstrap repetitions in a way such that (a), (b) and (c) above hold. We demonstrate the practical relevance and broad applicability of bootstrap diagnostics by considering several scenarios where the asymptotic Gaussian approximation may fail, including weak instruments, non-stationarity, parameters on the boundary of the parameter space, infinite variance data and singular Jacobian in applications of the delta method. An illustration drawn from the empirical macroeconomic literature concludes.

\noindentKeywords:\ Bootstrap inference; Pre-testing bias; Random bootstrap measures.

center[center omitted — 24 chars of source]

introduction

Consider a standardized estimator $T_{n}:=n^{1/2}(\hat{\theta} _{n}-\theta_{0})/\hat{\sigma}_{n}$ based on a data sample $D_{n}$, where $\hat{\sigma}_{n}^{2}$ is an estimator of the asymptotic variance. Classical asymptotic theory is usually based on a set of assumptions guaranteeing that, in large samples, the distribution of $T_{n}$ is well-approximated, to the first order, by some standard distribution, usually the normal one. That is, $T_{n}\overset{d}{\rightarrow}Z$, $Z\sim \mathscr{N}\! (0,1)$, with `$\overset{d}{\rightarrow}$' denoting convergence in distribution. When $\hat{\theta}_{n}$ is an extremum estimator, provided the objective function can be expanded around a (pseudo- ) true value $\theta_{0} $, such assumptions are usually related to (i) existence of moments, (ii) stationarity and ergodicity, (iii) non-singularity of the Hessian (or, for transformations of the original estimator, full rank of the implied Jacobian), (iv) (pseudo-) true parameter in the interior of the parameter space; see, e.g., Newey and McFadden (1994). For extremum estimators based on instrumental variables [IV], (v) assumptions on the strength of the instruments are usually required. Examples of estimators requiring assumptions such as (i)--(v) to hold are, inter alia, (quasi) ML estimators, GMM estimators, nonlinear least squares estimators and minimum distance estimators. With `valid specification' we mean that such assumptions are met.

The detection of invalid specifications is crucial in applications. A key challenge to proper statistical tests of valid specification is that they typically induce a `pre-testing' bias in subsequent inferences. While in some cases the pre-testing bias may be associated with conservative inference under the null hypothesis (see, e.g., de Chaisemartin and D'Haultfoeuille, 2024), tests conditional on the non-rejection of correct specification may be severely oversized, even asymptotically.

In this paper, we show the novel result that the bootstrap delivers, as a by-product, consistent (diagnostic) tests of invalid specification which do not induce any pre-testing bias into subsequent inference procedures when the null of valid specification is not rejected. That is, post-test (conditional) inference is asymptotically exact when conditioning is upon the bootstrap tests not rejecting the null hypothesis of valid specification.

Our approach starts from the observation that -- usually under mild additional requirements -- a bootstrap analog of $T_{n}$, say $T_{n}^{\ast}:=n^{1/2} (\hat{\theta}_{n}^{\ast}-\hat{\theta}_{n})/\hat{\sigma}_{n}$ or $T_{n}^{\ast }:=n^{1/2}(\hat{\theta}_{n}^{\ast}-\hat{\theta}_{n})/\hat{\sigma}_{n}^{\ast}$, will also be asymptotically normal under valid specification; i.e., $T_{n}^{\ast}\overset{d^{\ast}}{\rightarrow}_{p}Z$, with `$\overset{d^{\ast} }{\rightarrow}_{p}$' denoting convergence in distribution conditionally on $D_{n}$; see Appendix (ref) for a formal definition. In contrast, under invalid specification, $T_{n}^{\ast}$ may no longer be asymptotically normal; rather, it usually has a non-Gaussian, random limiting distribution; see Cavaliere and Georgiev (2020) and the references therein. For instance, randomness of the limiting bootstrap measure may arise when (i) the score contributions have infinite variance (Athreya, 1987; Knight, 1989; Cavaliere, Georgiev and Taylor, 2016), (ii) when the data are non-stationary (Basawa, Mallik, McCormick and Taylor, 1991; Cavaliere, Nielsen and Rahbek, 2015), (iii) when the Hessian or the Jacobian are (near-) rank deficient (Datta, 1995; Angelini, Cavaliere and Fanelli, 2022; see also Han and McCloskey, 2019), (iv) when the (pseudo-) true value lies near or on the boundary of the parameter space (Andrews, 2000); (v) for IV\ estimators, when the instruments are weak or irrelevant (see Section (ref)).

This different asymptotic behavior of the bootstrap distribution of $T_{n}^{\ast}$ can be exploited to detect invalid specification. To see why, consider the discrepancy between the cumulative distribution function [cdf] of $T_{n}^{\ast}$ conditional on $D_{n}$, say $\hat{G}_{n}$, and the limiting standard Gaussian cdf $\Phi$, measured as $\hat{d}_{n}:=\parallel\hat{G} _{n}-\Phi\parallel$, where $\parallel\cdot\parallel$ is a user-chosen norm or seminorm on the space of distribution functions. The distance $\hat{d}_{n}$ is expected to shrink to zero at a specific rate when asymptotic normality holds, while converging to a random limit when the assumptions fail to hold. Hence, a simple test of valid specification could assess whether the realized discrepancy $\hat{d}_{n}$ is large enough to reject the null.

As is standard in the bootstrap literature, $\hat{d}_{n}$ can be treated as a known function of the data $D_{n}$, since $\hat{G}_{n}$ can be approximated with any desired precision via the empirical distribution function [edf] $\hat{G}_{n,m}^{\ast}$ of $m$ simulated realizations of $T_{n}^{\ast}$, with $m$ arbitrarily large; that is, for any fixed $n$, as $m\rightarrow\infty$, with probability one, $\hat{G}_{n,m}^{\ast}\rightarrow\hat{G}_{n}$ uniformly, and hence $\hat{d}_{n,m}^{\ast}:=\lVert\hat{G}_{n,m}^{\ast}-\Phi \rVert\rightarrow\hat{d}_{n}$ a.s. if $\parallel\cdot\parallel$ is continuous with respect to uniform convergence.

While it could be tempting to construct diagnostic tests based on $\hat{d} _{n}$ (or a properly normalized version, such as $\sqrt{n}\hat{d}_{n}$), such tests would suffer from at least two important drawbacks. First, their large-$n$ asymptotic properties would depend on the particular bootstrap application of interest, and would often be very difficult to derive, as they require studying higher-order asymptotic expansions of $\hat{G}_{n}$. Second, since $\hat{d}_{n}$ is a function of the data, these tests might give rise to a pre-testing bias.

We show that these drawbacks disappear if inference, rather than being based on $\hat{d}_{n}$, is based on its approximation $\hat{d}_{n,m}^{\ast}$, where $m$ and $n$ diverge jointly under the requirement that $m$ cannot be too large when $n$ is finite. This asymptotic regime differs from the standard sequential bootstrap asymptotics, where $m\rightarrow\infty$ first (such that $\hat{G}_{n,m}^{\ast}-\hat{G}_{n}\approx0$) followed by $n\rightarrow\infty$ (such that $\hat{G}_{n}-\Phi\approx0$); see Andrews and Buchinsky (2000). In particular, we show that when $m$ and $n$ diverge jointly, with $m$ diverging at a proper rate relative to $n$, a test with a known asymptotic distribution and based on the bootstrap statistic $\hat{d}_{n,m}^{\ast}$ can be designed to assess specification invalidity. Moreover, this approach is computationally straightforward, as it just requires to use the set of $m$ bootstrap repetitions to compute $\hat{d}_{n,m}^{\ast}$. Put differently, it is equivalent to the application of distance-based normality tests to the set of $m$ bootstrap repetitions, with different normality tests corresponding to different choices of the employed (semi)norm $\parallel\cdot\parallel$. Finally, it can be performed using the same critical values in a broad range of applications, and\ consistently detects deviations from asymptotic Gaussianity.

The role of the rate condition on $m$ relative to $n$ is to ensure that, under the null hypothesis of valid specification, the test decision becomes, in the limit, stochastically independent of the original data. This fact guarantees that a bootstrap test based on $\hat{d}_{n,m}^{\ast}$ does not induce pre-testing bias in large samples, in contrast to standard pre-tests for specification (in)validity (e.g., tests of finite variance, stationarity tests, pre-tests on the instrument strengths, and so forth). Instead, under a set of relevant alternatives the test is consistent as it exploits, in the limit, the information in the data alone. In summary, according to the validity of either the null or an alternative, the data choose asymptotically between acceptance based on an independent random device or rejection with probability approaching one. Interestingly, in a recent paper, de Chaisemartin and D'Haultf\oe uille (2024)\ show that certain specification tests\footnote{De Chaisemartin and D'Haultf\oe uille define `valid specification'\ as the scenario in which the probability distribution generating $D_{n}$ is such that a target estimator is consistent and asymptotically normal. They consider pre-tests that can detect when such an estimator becomes inconsistent. This definition differs from the one employed here, where `valid specification'\ means that the underlying conditions ensuring the applicability of standard asymptotic inference are satisfied in the estimated model.}, when used as pre-tests, lead to conservative post-test inference under the null of valid specification. We complement their result by showing that the bootstrap delivers tests of specification validity leading to exact post-test inference as the sample size diverges.

This diagnostic potential of the bootstrap has not been developed in the extant literature. Beran (1997) was the first to suggest examining the bootstrap distribution to diagnose bootstrap failure. Davidson (2017) proposes simulation-based diagnostics in order to determine when a given bootstrap procedure works well or not in finite samples. B\aa rdsen and Fanelli (2015) provide prima facie evidence that non-Gaussian bootstrap distributions may be linked to weak identification in DSGE models; see also Angelini, Cavaliere and Fanelli (2022, 2024), and Zhan (2018). Within the problem of statistical reporting in a Bayesian communication framework, Andrews and Shapiro (2025) show that the (Bayesian) bootstrap distribution can serve as a surrogate posterior, and propose comparing it with the Gaussian approximation, using distance measures such as the signed Kolmogorov metric, to assess whether the conventional report adequately conveys uncertainty. A related approach is found in Wang (2025), who proposes detecting violations of asymptotic normality for GMM and extremum estimators by comparing quasi-Bayesian posterior distributions with the Gaussian benchmark. Our contribution extends this literature by demonstrating how distances between bootstrap and Gaussian distributions can be used to construct formal specification tests that avoid pre-testing bias. In terms of econometric theory, a further novelty lies in the asymptotic regime we adopt, where both the sample size $n$ and the number of bootstrap repetitions $m$ pass to infinity simultaneously. This setting is rarely considered in the literature; a notable exception is Andrews and Buchinsky (2000), who exploit this joint asymptotic regime to guide the choice of $m$ in applied work.

To demonstrate the practical relevance and broad applicability of the bootstrap in the detection of specification invalidity, we discuss its use in the five scenarios (i)--(v) mentioned above. As regards case (i) of possible weak instruments, we also present an empirical illustration where we revisit, through the lens of bootstrap diagnostics, K\"{a}nzig's (2021) empirical strategy for identifying the macroeconomic effects of a structural oil supply news shock.

\paragraph*{structure of the paper.}

The paper is organized as follows. In Section (ref) we introduce a running example based on instrumental variable estimation. Section (ref) contains our general results. Section (ref) establishes the key result that the bootstrap procedure induces no pre-test bias in large samples. Additional results and extensions are reported in Section (ref), while Section (ref) illustrates four further applications. An empirical example is presented in Section (ref), and Section (ref) concludes. Notation and definitions used throughout the paper can be found in Appendix (ref). Proofs are collected in Appendices (ref) and (ref).

an example based on instrumental variables

Consider the following linear IV\ regression with one endogenous regressor: \[ y_{i}=\beta x_{i}+u_{i}\text{, }x_{i}=\pi^{\prime}z_{i}+v_{i} \] where the $k\times1$ vector of instruments $z_{i}$ is non-stochastic and the errors $(u_{i},v_{i})^{\prime}$ are i.i.d. with mean zero and variance-covariance matrix $\Sigma$; to simplify, the diagonal elements of $\Sigma$ are set to $\sigma_{u}^{2}=\sigma_{v}^{2}=1$ and the off-diagonal elements to $\rho_{uv}\in(0,1)$. Without loss of generality, we also set $S_{zz}=I_{k}$, using the generic notation $S_{ab}:=n^{-1}\sum_{i=1}^{n} a_{i}b_{i}^{\prime}$. Given a sample of $n$ observations, the 2SLS estimator of $\beta$ is $\hat{\beta}_{n}:=S_{xx.z}^{-1}S_{xy.z}=(S_{xz}S_{zy} )/(S_{xz}S_{zx})$.

Consider further a (Gaussian) parametric bootstrap where the instruments are fixed in the bootstrap world. The bootstrap data are generated as \[ y_{i}^{\ast}=\hat{\beta}_{n}x_{i}^{\ast}+u_{i}^{\ast}\text{, }x_{i}^{\ast }=\hat{\pi}_{n}^{\prime}z_{i}+v_{i}^{\ast} \] where $\hat{\pi}_{n}:=S_{zx}$ is the OLS\ estimator from the (first-stage) regression of $x_{i}$ on $z_{i}$. For simplicity, assume that $\Sigma$ is known, such that $(u_{i}^{\ast},v_{i}^{\ast})^{\prime}$ can be taken as i.i.d. $ \mathscr{N}\! \left( 0,\Sigma\right) $ conditionally on the data.

Let first $\pi\neq0$ be fixed (i.e., the instruments be `strong'), a case that we label a `valid specification'. Then, under standard assumptions for the central limit theorem (CLT), \[ T_{n}:=\sqrt{n}\frac{\hat{\beta}_{n}-\beta}{\omega}=\frac{\sqrt{n}\pi^{\prime }S_{zu}}{(\pi^{\prime}\pi)^{1/2}}+o_{p}(1)\overset{d}{\rightarrow} \mathscr{N}\! (0,1) \] where $\omega^{2}:=(\pi^{\prime}\pi)^{-1}$. Moreover, as $\hat{\pi} _{n}^{\prime}\hat{\pi}_{n}\rightarrow_{p}\pi^{\prime}\pi\neq0$, conditionally on the data,

equation[equation omitted — 308 chars of source]

with $\hat{\omega}_{n}^{2}:=(\hat{\pi}_{n}^{\prime}\hat{\pi}_{n})^{-1}$. Therefore $T_{n}^{\ast}\overset{d^{\ast}}{\rightarrow}_{p} \mathscr{N}\! (0,1)$.

Instead, suppose next that the instruments are weak as in Staiger and Stock (1997), i.e., $\pi=\lambda n^{-1/2}$ for some vector $\lambda$; see also Bound, Jaeger and Baker (1995). This is an instance of an invalid specification. Assuming that \[ \sqrt{n}\binom{S_{zu}}{S_{zv}}=n^{-1/2}\sum_{i=1}^{n}\binom{u_{i}}{v_{i} }\otimes z_{i}\overset{d}{\rightarrow}\zeta:=(\zeta_{u}^{\prime},\zeta _{v}^{\prime})^{\prime} \] with $\zeta$ a zero-mean Gaussian vector with variance-covariance matrix $\Sigma\otimes I_{k}$, we have

equation[equation omitted — 449 chars of source]

Similarly, the bootstrap analog of $\hat{\beta}_{n}-\beta$ can be written as \[ \hat{\beta}_{n}^{\ast}-\hat{\beta}_{n}=\frac{S_{x^{\ast}z}S_{zu^{\ast}} }{S_{x^{\ast}z}S_{zx^{\ast}}}=\frac{\left( \sqrt{n}\hat{\pi}_{n}+\sqrt {n}S_{zv^{\ast}}\right) ^{\prime}\sqrt{n}S_{zu^{\ast}}}{\left( \sqrt{n} \hat{\pi}_{n}+\sqrt{n}S_{zv^{\ast}}\right) ^{\prime}\left( \sqrt{n}\hat{\pi }_{n}+\sqrt{n}S_{zv^{\ast}}\right) } \] where, conditionally on the data, $n^{1/2}(S_{zu^{\ast}}^{\prime},S_{zv^{\ast }}^{\prime})^{\prime}\sim\zeta^{\ast}$, with $\zeta^{\ast}:=(\zeta_{u} ^{\ast\prime},\zeta_{v}^{\ast\prime})^{\prime}$ distributed as $\zeta$. Since, jointly with ((ref)), $\sqrt{n}\hat{\pi} _{n}\rightarrow_{d}\ell:=\lambda+\zeta_{v}$, we find using arguments as in Cavaliere and Georgiev (2020, proof of Theorem 4.1) that $\hat{\beta} _{n}^{\ast}-\hat{\beta}_{n}\overset{d^{\ast}}{\rightarrow}_{w}\xi(\ell ,\zeta^{\ast})|\ell$ and\ that

equation[equation omitted — 319 chars of source]

see Appendix (ref) for the definition of $\overset{d^{\ast} }{\rightarrow}_{w}$. Hence, the bootstrap distribution has a random limit and, as $n\rightarrow\infty$, the bootstrap cdf satisfies $\hat{G}_{n} (x)\rightarrow_{w} \mathscr{G}\! (x):=\mathbb{P(}\xi(\ell,\zeta^{\ast})/\sqrt{\ell^{\prime}\ell}\leq x|\ell)$, $x\in\mathbb{R}$, on $\mathscr{D}_{\mathbb{R}}$, where the limit differs from the Gaussian cdf.

remarkThe limiting randomness of the bootstrap distribution $\hat{G}_{n}$ under weak instruments is illustrated in Figure (ref), where we report the fan chart of $M=1,000$ i.i.d. realizations of $\hat{G}_{n}$ for $k=1$, $n=1,000$ and using a standard Gaussian parametric bootstrap with $z_{i}=1$ and $\rho_{uv}=0.9$. The randomness and non-normality in the bootstrap cdf $\hat{G}_{n}$, quite evident for small values of $\lambda$, ameliorates as $\lambda$ increases and (up to sampling error due to finite $n$) disappears in the strong-instrument case where $\pi\neq0$ is fixed.$\square$
remarkNotice that under weak instruments the bootstrap does not replicate the asymptotic distribution of the original statistic, given by $\xi(\lambda ,\zeta)$ of ((ref)), essentially because in the limiting bootstrap experiment $\lambda$ is replaced by the random vector $\ell$, where $\ell\neq\lambda$ with probability one. This result explains the inconsistency of the bootstrap in the weak IV\ framework, as previously documented in the literature (see, e.g., Davidson and MacKinnon, 2010), in terms of randomness of the limit bootstrap measure.$\square$
remarkThe results in this section do not substantially change if $z_{i}$ is random and resampled along with $\varepsilon_{i}^{\ast}$ and $u_{i}^{\ast}$, and/or if $\{\varepsilon _{i}^{\ast},u_{i}^{\ast}\}_{i=1}^{n}$ are i.i.d. draws from the (centered) residuals $\{\hat{\varepsilon}_{i},\hat{u}_{i}\}_{i=1}^{n}$ as in standard non-parametric bootstrap designs (in the latter case, the random asymptotic distribution in ((ref)) is more involved).$\square$
figure[figure omitted — 318 chars of source]

testing specification validity

Assume that $\theta\in\mathbb{R}$ and consider a bootstrap statistic $T_{n}^{\ast}$ that is asymptotically $ \mathscr{N}\! (0,1)$ under valid specification. $T_{n}^{\ast}$ can be, e.g., of the form $T_{n}^{\ast}:=(\hat{\theta}_{n}^{\ast}-\hat{\theta}_{n})/\operatorname*{se} (\hat{\theta}_{n}^{\ast})$ or $T_{n}^{\ast}:=(\hat{\theta}_{n}^{\ast} -\hat{\theta}_{n})/\operatorname*{se}(\hat{\theta}_{n})$ (non-studentized statistics of the form $T_{n}^{\ast}:=\sqrt{n}(\hat{\theta}_{n}^{\ast} -\hat{\theta}_{n})$ can also be considered). As we shall see in this section, a test of valid specification, which does not induce pre-testing bias, can be obtained by assessing whether the bootstrap replicates the Gaussian asymptotic distribution. As an initial step, we quantify the discrepancy between the bootstrap cdf and the asymptotic Gaussian cdf.

preliminaries

Let $\hat{G}_{n}(\cdot):=\mathbb{P}^{\ast}(T_{n}^{\ast}\leq\cdot)$ be the cdf of $T_{n}^{\ast}$, conditional on the data $D_{n}$. We consider measuring the discrepancy between $\hat{G}_{n}$ and the $ \mathscr{N}\! (0,1)$ cdf $\Phi$ by $\hat{d}_{n}:=\parallel\hat{G}_{n}-\Phi\parallel$ where $\Vert\cdot\Vert:\mathscr{D}_{\mathbb{R}}\rightarrow\lbrack0,\infty)$ is a user-chosen norm or seminorm; alternative discrepancy measures are discussed in Section (ref). For instance, setting $\Vert\cdot\Vert=\Vert\cdot\Vert_{\infty}$, i.e., the sup norm on $\mathscr{D}_{\mathbb{R}}$, delivers the well known Kolmogorov-Smirnov [ KS ] distance

equation[equation omitted — 163 chars of source]

Under the bootstrap consistency hypothesis $T_{n}^{\ast}\overset{d^{\ast} }{\rightarrow}_{p} \mathscr{N}\! (0,1)$, by Polya's theorem,

equation[equation omitted — 115 chars of source]

and similarly, for any (semi)norm $\Vert\cdot\Vert\ $that is continuous on $\mathscr{D}_{\mathbb{R}}$ with the uniform metric;\ examples are the signed KS (Andrews and Shapiro, 2025) and the Cramer-von Mises norms. However, should bootstrap consistency for the Gaussian limit fail, also ((ref)) will fail to hold. In particular, for the IV\ example of Section (ref) as well as for all the examples in Section (ref), it holds that $\hat{G}_{n}\overset {w}{\rightarrow} \mathscr{G}\! $ in $\mathscr{D}_{\mathbb{R}}$, where $ \mathscr{G}\! $ is a (random) cdf such that $ \mathscr{G}\! \neq\Phi$ (a.s.). By the continuous mapping theorem [CMT], provided the transformation $\Vert\cdot\Vert$ is continuous on $\mathscr{D}_{\mathbb{R}} $,\footnote{If $ \mathscr{G}\! $ is sample-path continuous, continuity of $\Vert\cdot\Vert$ on $\mathscr{D}(\mathbb{R})$ equipped with the uniform metric suffices.} $\hat {d}_{n}$ itself has a (possibly) random limit:

equation[equation omitted — 150 chars of source]

where $ \mathscr{Y} >0$ a.s. if $\Vert\cdot\Vert$ is a norm.\footnote{For seminorms, e.g., $\hat{d}_{n}(A):=\sup_{u\in A\subset\mathbb{R}}|\hat{G}_{n}(u)-\Phi(u)|$, also $\mathbb{P}( \mathscr{Y} =0)>0$ may hold; for an example see the parameter-on-the-boundary case of Section (ref), where the positivity of $ \mathscr{Y} $ for $\hat{d}_{n}(A)$ depends on the choice of the set $A$.} The different asymptotic behaviors in ((ref)) and ((ref)) will be exploited to develop bootstrap tests for valid specification.

As is standard, the cdf $\hat{G}_{n}$, as well as the discrepancy $\hat{d} _{n}$, can be approximated using a sample $T_{n:i}^{\ast}$, $i=1,...,m$, of conditionally independent copies of $T_{n}^{\ast}$ (obtained by simulation):

equation[equation omitted — 251 chars of source]

with $m$ user-chosen. Both $\hat{d}_{n,m}^{\ast}$ and $\hat{G}_{n,m}^{\ast}$, the latter being usually employed in applications of the bootstrap to compute p-values and confidence sets, are key to assess specification validity without inducing pre-testing bias.

asymptotic regimes

Consider $\hat{G}_{n,m}^{\ast}$ and $\hat{d}_{n,m}^{\ast}$ of ((ref)). By the Glivenko-Cantelli theorem, for any $n$ and as $m\rightarrow\infty$, $\Vert\hat{G}_{n,m}^{\ast}-\hat{G}_{n} \Vert_{\infty}\overset{a.s.^{\ast}}{\rightarrow}0$ (a.s.) Then, for any $n$, also $\hat{d}_{n,m}^{\ast}\overset{a.s.^{\ast}}{\rightarrow}\hat{d}_{n}$ (a.s.), provided that $\Vert\cdot\Vert$ is continuous on $\mathscr{D}_{\mathbb{R}}$ with the Skorokhod $J_{1}$ or the $\sup$ metric. Hence, for practical purposes $\hat{G}_{n}$ and $\hat{d}_{n}$ can be treated as known and asymptotic inference based on transformations of $\hat{G} _{n}-\Phi$, like $\hat{d}_{n}$, is in principle feasible.

As anticipated in Section (ref), however, it turns out that inference based on such transformations is unattractive in practice, as (i)\ their null asymptotic distribution as $n\rightarrow\infty$ is in general not only very hard to obtain, but also application-specific (see Section (ref) for an example involving $\hat{d}_{n}$); and (ii) a problem of post-diagnostic test bias emerges.

A closer look shows that these issues arise under an implicit sequential asymptotic regime -- which we label as $(m,n\rightarrow\infty)_{\text{seq }}$ using the notation in Phillips and Moon (1999) -- where first $m\rightarrow \infty$ in order for $\hat{d}_{n,m}^{\ast}$ to collapse to $\hat{d}_{n}$, and then $n\rightarrow\infty$. Standard bootstrap asympotic theory, which would discuss $\hat{d}_{n}$ as $n\rightarrow\infty$ rather than the actually computed $\hat{d}_{n,m}^{\ast}$, can be interpreted as employing $(m,n\rightarrow\infty)_{\text{seq }}$ asymptotics.

We now ask whether a different asymptotic regime can be chosen such that (i)\ the asymptotic distributions of test statistics are invariant across a wide range of applications under the null of valid specification, (ii) no post-diagnostic bias is present under the null, and (iii) the tests are consistent against a relevant class of alternatives. Two candidate asymptotic regimes are as follows.

First, consider the sequential regime $(n,m\rightarrow\infty)_{\text{seq }}$, where $n\rightarrow\infty$ followed by $m\rightarrow\infty$. In the case of the $\sup$ norm, $\hat{d}_{n,m}^{\ast}=\Vert\hat{G}_{n,m}^{\ast}-\Phi \Vert_{\infty}$ can be written as $\hat{d}_{n,m}^{\ast}=\phi_{m}(T_{n:1} ^{\ast},...,T_{n:m}^{\ast})$ for a continuous $\phi_{m}\nolinebreak :\nolinebreak\mathbb{R}^{m}\nolinebreak\rightarrow\nolinebreak\mathbb{R}$. Because under the null of valid specification $T_{n}^{\ast}\overset{d^{\ast} }{\rightarrow}_{p}Z\sim \mathscr{N}\! (0,1)$, it holds that \[ \hat{d}_{n,m}^{\ast}=\phi_{m}(T_{n:1}^{\ast},...,T_{n:m}^{\ast})\overset {d^{\ast}}{\rightarrow}_{p}d_{m}^{\ast}:=\phi_{m}(Z_{1},...,Z_{m}) \] as $n\rightarrow\infty$, with the $Z_{i}$'s being independent $ \mathscr{N}\! (0,1)$, $i=1,...,m$. Now, let $m\rightarrow\infty$ after $n\rightarrow\infty$. Under the null, this yields by standard empirical process theory that, with $W$ denoting a standard Brownian bridge, $\sqrt{m}d_{m}^{\ast}\overset {d}{\rightarrow}\Vert W(\Phi)\Vert_{\infty}$, i.e., the Kolmogorov distribution. Therefore, if $(n,m\rightarrow\infty)_{\text{seq}}$, \[ \mathscr{T} _{n,m}^{\ast}:=\sqrt{m}\hat{d}_{n,m}^{\ast}=\sqrt{m}\Vert\hat{G}_{n,m}^{\ast }-\Phi\Vert_{\infty}\overset{d^{\ast}}{\rightarrow}_{p}\Vert W(\Phi )\Vert_{\infty}. \] In contrast, under alternatives such that $\hat{G}_{n}\overset{w}{\rightarrow} \mathscr{G}\! \neq\Phi$ a.s., it holds that $ \mathscr{T} _{n,m}^{\ast}\overset{p^{\ast}}{\rightarrow}_{p}\infty$ as $(n,m\rightarrow \infty)_{\text{seq }}$, yielding consistent tests.

Although the sequential `$m$ after $n$' asymptotic regime leads to tractable derivations, it provides little justification for conducting inference based on $\Vert W(\Phi)\Vert_{\infty}$. Indeed, in practice, one has more control on $m$ (number of bootstrap repetitions) rather than $n$ (sample length). Hence, we also consider an asymptotic regime where $m$ and $n$ diverge jointly, denoted as $(n,m\rightarrow\infty)$. Under the null and under an additional rate condition relating $m$ to $n$, this regime allows to achieve asymptotic distributions that are invariant across applications (e.g., the $\Vert W(\Phi)\Vert_{\infty}$ limit seen previously for the $\sup$ norm), whereas under relevant alternatives the resulting tests are consistent. Focusing on the statistic $ \mathscr{T} _{n,m}^{\ast}:=\sqrt{m}\hat{d}_{n,m}^{\ast}$, the rate condition makes sure that $\hat{G}_{n,m}^{\ast}-\hat{G}_{n}$ is dominant in the derivation of $ \mathscr{T} _{n,m}^{\ast}$'s asymptotic distribution under the null, whereas under alternatives of interest, $\hat{G}_{n}-\Phi$ is dominant in the limit.

Notice that both the $(n,m\rightarrow\infty)_{\text{seq}}$ and $(n,m\rightarrow\infty)$ regimes are intended for justifying inference employing $\hat{G}_{n,m}^{\ast}$ rather than true $\hat{G}_{n}$. Although this choice implies a loss of information for every fixed $n$, as $\hat{G} _{n,m}^{\ast}$ is used before it has converged to $\hat{G}_{n}$, it is the bootstrap randomness in $\hat{G}_{n,m}^{\ast}-\hat{G}_{n}$ that can make the asymptotic distribution of $\hat{d}_{n,m}^{\ast}$ invariant across applications and can eliminate the post-diagnostic bias. Indeed, for both regimes no post-diagnostic test bias is present asymptotically, as the proof of Theorem (ref) below for the joint regime goes through also for the sequential `$m$ after $n$' regime.

the diagnostic test

Redefine $ \mathscr{T} _{n,m}^{\ast}$ introduced above using a generic (semi)norm, i.e., $ \mathscr{T} _{n,m}^{\ast}:=m^{1/2}\hat{d}_{n,m}^{\ast}$, $\hat{d}_{n,m}^{\ast}:=\Vert \hat{G}_{n,m}^{\ast}-\Phi\Vert$ with $\hat{G}_{n,m}^{\ast}(\cdot):=m^{-1} \sum\nolimits_{i=1}^{m}\mathbb{I}_{\mathbb{\{}T_{n:i}^{\ast}\leq\cdot\}}$ and the $T_{n:i}^{\ast}$'s being i.i.d. copies of $T_{n}^{\ast}$ conditionally on the data $D_{n}$. We note the following.

First, $ \mathscr{T} _{n,m}^{\ast}$ depends on the data $D_{n}$ as well as on $m$ auxiliary variates, say $W_{n}^{\ast}$, used to generate the $m$ bootstrap draws $T_{n:i}^{\ast}$, $i=1,\ldots,m$. For instance $W_{n}^{\ast}$, which is defined jointly with $D_{n}$ on a possibly extended probability space, could be thought of as a vector of $m$ i.i.d. $ \mathscr{U}\! _{[0,1]}$ r.v.s, independent of $D_{n}$, such that $T_{n:i}^{\ast}=\hat{G} _{n}^{-1}(W_{n,i}^{\ast})$, $i=1,\ldots,m$, with $\hat{G}_{n}^{-1}$ the generalized inverse of $\hat{G}_{n}$.

Second, $ \mathscr{T} _{n,m}^{\ast}$ can be written as

equation[equation omitted — 188 chars of source]

with $a_{n,m}^{\ast}$ implicitly defined. Here, $\mathcal{Z}_{n,m}^{\ast}$ is a rescaled measure of the distance between the estimator $\hat{G}_{n,m}^{\ast }$ and $\hat{G}_{n}$, while $a_{n,m}^{\ast}$ is related to the distance between $\hat{G}_{n}$ and the Gaussian cdf $\Phi$. In particular, for any $n$, $|a_{n,m}^{\ast}|\leq\sqrt{m}\Vert\hat{G}_{n}-\Phi\Vert$ by Lemma (ref) in Appendix (ref).

We now provide conditions such that, under the valid specification hypothesis, the asymptotic distribution of $ \mathscr{T} _{n,m}^{\ast}$ can be derived and hence used to perform a proper statistical test. In particular, we consider the following assumption.

conditionThe (semi)norm $\Vert\cdot\Vert$ is continuous on $\mathscr{D}_{\mathbb{R}}$. Moreover, $\Vert\hat{G}_{n}-\Phi\Vert =O_{p}(n^{-\alpha})$ for some $\alpha>0$.

Assumption (ref) is a condition on the rate of convergence of the bootstrap cdf to the Gaussian cdf under valid specification. It is satisfied with $\alpha=1/2$ if $\Vert\cdot\Vert\leq C\Vert\cdot\Vert_{\infty}$ for some finite constant $C$ (as is the case of, e.g., the signed KS norm or the Cramer-von Mises norm), and the bootstrap statistic has a one-term Edgeworth expansion of the form $\hat{G}_{n}(x)=\Phi(x)+\hat{q}_{n} (x)n^{-1/2}+o_{p}(n^{-1/2})$ uniformly in $x$, as is usually the case; see, e.g., Hall (1992). For some symmetric statistics, Assumption (ref) can hold with $\alpha=1$. For Gaussian parametric bootstraps, Assumption (ref) may be even satisfied with arbitrary $\alpha>0$; see, e.g., the case in Section (ref).

The following result holds under Assumption (ref).

theoremLet $T_{n}^{\ast}\overset{d^{\ast}}{\rightarrow}_{p} \mathscr{N}\! (0,1)$. Under Assumption (ref), if $(n,m\rightarrow\infty)$ and $m/n^{2\alpha}\rightarrow0$, \begin{description} • (i) $ \mathscr{T} _{n,m}^{\ast}=\mathcal{Z}_{n,m}^{\ast}+o_{p}(1)\overset{d^{\ast}}{\rightarrow }_{p} \mathscr{K}\! :=\Vert W(\Phi)\Vert$, where $W$ is standard Brownian bridge. • (ii) $p_{n,m}^{\ast}:=1-H( \mathscr{T} _{n,m}^{\ast})\overset{w^{\ast}}{\rightarrow}_{p} \mathscr{U}\! _{[0,1]}$ if the cdf $H$ of $ \mathscr{K}\! $ is continuous. \end{description}

Theorem (ref) suggests that a simple diagnostic procedure can be obtained by comparing $ \mathscr{T} _{n,m}^{\ast}$ with critical values from the distribution of $ \mathscr{K}\! $, or through the associated asymptotic p-value $p_{n,m}^{\ast}$. Provided both $m$ and $n$ grow to infinity at a proper relative rate, Theorem (ref) guarantees that the procedure is asymptotically correctly sized.

remarkThe logic of the proof can be easily followed in the case of the signed pointwise discrepancy $\hat{d}_{n}(x):=\hat{G}_{n}(x)-\Phi(x)$ for some fixed $x\in\mathbb{R}$. In this case, $\mathcal{Z}_{n,m}^{\ast}$ and $a_{n,m}^{\ast }$ of ((ref)) are given by $\mathcal{Z}_{n,m}^{\ast}=m^{1/2} (\hat{G}_{n,m}^{\ast}(x)-\hat{G}_{n}(x))$ and $a_{n,m}^{\ast}=m^{1/2}(\hat {G}_{n}(x)-\Phi(x))$. If $\hat{d}_{n}(x)$ satisfies Assumption (ref) for some $\alpha>0$, it holds that $a_{n,m}^{\ast}\rightarrow_{p}0$ as $(n,m\rightarrow\infty)$ with $m/n^{2\alpha}\rightarrow0$. Moreover, with $\xi_{i}^{\ast}:=\mathbb{I}_{\mathbb{\{}T_{n:i}^{\ast}\leq x\}}-\mathbb{E} ^{\ast}[\mathbb{I}_{\mathbb{\{}T_{n:i}^{\ast}\leq x\}}]$, such that $\mathcal{Z}_{n,m}^{\ast}=m^{-1/2}\sum_{i=1}^{m}\xi_{i}^{\ast}$, by a standard Berry-Esseen bound we have that $\mathcal{\tilde{Z}}_{n,m}^{\ast} :=\mathbb{E}^{\ast}[\xi_{i}^{\ast}{}^{2}]^{-1/2}\mathcal{Z}_{n,m}^{\ast}$ satisfies, for some $C<\infty$, \[ \parallel\mathbb{P}^{\ast}\mathbb{(}\mathcal{\tilde{Z}}_{n,m}^{\ast}\leq \cdot)-\Phi(\cdot)\parallel_{\infty}\leq Cm^{-1/2}\mathbb{E}^{\ast}[|\xi _{i}^{\ast}|^{3}]\leq Cm^{-1/2} \] whenever $\mathbb{E}^{\ast}[\xi_{i}^{\ast2}]=\hat{G}_{n}(x)(1-\hat{G} _{n}(x))\neq0$, which occurs with probability approaching one as $n\rightarrow\infty$ since $\mathbb{E}^{\ast}[\xi_{i}^{\ast2}]\rightarrow _{p}v^{2}(x):=\Phi\left( x\right) (1-\Phi(x))$ under the hypothesis that $T_{n}^{\ast}\overset{d^{\ast}}{\rightarrow}_{p} \mathscr{N}\! (0,1)$ as $n\rightarrow\infty$. We conclude that $\mathcal{Z}_{n,m}^{\ast }=v(x)\mathcal{\tilde{Z}}_{n,m}^{\ast}+o_{p}^{\ast}(1)\overset{d^{\ast} }{\rightarrow}_{p} \mathscr{N}\! \left( 0,v(x)\right) \sim W(\Phi(x))$ as $(n,m\rightarrow\infty)$ .$\square$
remarkFor commonly used norms, the distribution of $ \mathscr{K}\! $ is well known; for example, with $\Vert\cdot\Vert=\Vert\cdot\Vert_{\infty}$, $ \mathscr{K}\! $ follows the Kolmogorov distribution which has a continuous cdf. In general, critical values (as well as p-values) can be determined by Monte Carlo simulation with arbitrary accuracy.$\square$

The test based on $ \mathscr{T} _{n,m}^{\ast}$ also has non-trivial asymptotic power against bootstrap inconsistency for the standard Gaussian distribution. This is shown next.

theoremSuppose that $\hat{G}_{n}\overset{w}{\rightarrow} \mathscr{G}\! $ in $ \mathscr{D} _{\mathbb{R}}$ as $n\rightarrow\infty$.\ Suppose further that $\Vert\cdot \Vert$ is continuous on $ \mathscr{D} _{\mathbb{R}}$ and $\Vert \mathscr{G}\! -\Phi\Vert>0$ a.s. Then, for any $c\in\left( 0,\infty\right) $, it holds that $\mathbb{P}^{\ast}( \mathscr{T} _{n,m}^{\ast}\geq c)\overset{p}{\rightarrow}1$ as $(n,m\rightarrow\infty)$.

An inspection of the proof of Theorem (ref) reveals that, as expected, for large $m$ the power of the test is determined by the realized value of $\hat{d}_{n}:=\Vert\hat{G}_{n}-\Phi\Vert$, which is asymptotically distributed as the r.v. $ \mathscr{Y} :=\Vert \mathscr{G}\! -\Phi\Vert>0$ a.s. Such realization depends on the original data $D_{n}$ only, and not on the $m$ bootstrap repetitions used to generate the bootstrap statistic. Larger outcomes of $\hat{d}_{n}$ correspond -- ceteris paribus -- to larger power.

an example based on instrumental variables (cont'd)

When $\pi\neq0$ is fixed, such that the instruments are strong, for the parametric bootstrap described in Section (ref) it holds, without additional assumptions, that $\Vert\hat{G}_{n}-\Phi\Vert_{\infty }=O_{p}(n^{-\alpha})$ for any $\alpha\in(0,\frac{1}{2})$; this is shown in Appendix (ref) by using the machinery of parametric tail estimates. Hence, Assumption (ref) is verified with $\alpha \in(0,\frac{1}{2})$ for the $\sup$ norm and its dominated norms, and Theorem (ref) applies to them. Using uniform Edgeworth expansions, also $\Vert\hat{G}_{n}-\Phi\Vert_{\infty}=O_{p}(n^{-1/2})$ has been shown to hold for the non-parametric i.i.d. bootstrap (where $u_{i}^{\ast}$ and $v_{i} ^{\ast}$ are resampled from the residuals $\hat{u}_{i}:=y_{i}-\hat{\beta} _{n}x_{i}$ and $\hat{v}_{i}:=x_{i}-\hat{\pi}_{n}^{\prime}z_{i}$), under mild regularity conditions on $(u_{i},v_{i})$ (essentially, existence of higher-order moments); see Moreira, Porter and Suarez (2009, Theorem 3).

Under weak instruments, we found previously that the bootstrap cdf of $T_{n}^{\ast}$ satisfies $\hat{G}_{n}(x)\rightarrow_{w} \mathscr{G}\! (x):=\mathbb{P(}\xi(\ell,\zeta^{\ast})/\sqrt{\ell^{\prime}\ell}\leq x$ $|$ $\ell)$ on $\mathscr{D}_{\mathbb{R}}$, which is non-Gaussian with probability one. Thus, Theorem (ref) applies in this case.

table[table omitted — 2,377 chars of source]

We conclude by reporting in Table (ref) the (percentage) empirical rejection probabilities (ERPs) of the bootstrap test based on the sup norm (KS); for comparison, we also consider the test based on the Anderson-Darling norm (AD). The data generating process is as in Remark (ref) with $z_{t}\sim$i.i.d.$ \mathscr{N} (0,1)$ and the bootstrap is non-parametric (see Remark (ref)) with $z_{t}$ fixed in the bootstrap world; a constant is included in estimation. The upper panel corresponds to the strong instrument case ($\pi=1$), while the lower panel considers the weak instrument case with $\pi=\lambda n^{-1/2}$ and $\lambda=2$. To assess how well the asymptotic theory $(n,m\rightarrow\infty,\;m/n\rightarrow0)$ approximates finite-sample behavior, we consider different choices of $m$, namely $m=\lfloor n^{\zeta}\rfloor$ with $\zeta\in\{0.50,0.55,\ldots ,0.70,0.80,0.90,1.00\}$. ERPs are computed using $10,000$ Monte Carlo replications; the nominal level is $5\%$.

In the strong instrument case (upper panel), the ERPs remain relatively close to the nominal level, particularly for lower values of $m$ relative to $n$. They increase as $m$ grows; for KS they never exceed $11\%$, while for AD they can rise up to about $10\%$. Consistently with the theoretical expectation, under the weak instrument scenario (lower panel), the ERPs increase with $m,n$.

post-diagnostics inference

The bootstrap approach developed in the previous section can be used in diagnostic pre-testing of specification validity, with standard inference carried out when the diagnostic procedure does not reject. An important question is whether post-diagnostic inferences are biased by the outcome of the pre-test, in the sense that conditionally on a correct non-rejection by the pre-test, the rejection probabilities of post-diagnostic tests are affected even asymptotically. The answer is no.

The key for this result is that, under valid specification, the bootstrap statistic $ \mathscr{T} _{n,m}^{\ast}$ becomes independent of the original data $D_{n}$ under the joint asymptotic regime $(n,m\rightarrow\infty)$ and the rate condition of Theorem (ref). That is, in the limit $ \mathscr{T} _{n,m}^{\ast}$ only depends on the bootstrap variates used to generate $\hat{G}_{n,m}^{\ast}$ and no longer on the data $D_{n}$, thus eliminating any post-diagnostic bias.

To see why, it suffices to note that the existence of a non-random limit for the conditional law of a bootstrap statistic given the data is an asymptotic independence property. It is this very property that the conditional law of the bootstrap diagnostic statistic $ \mathscr{T} _{n,m}^{\ast}$ enjoys under the conditions of Theorem (ref). The meaning of the implied asymptotic independence is clarified in the next theorem, where $ \mathscr{T} _{n,m}^{\ast}$ can stand for any bootstrap statistic and $(n,m\rightarrow \infty,R)$ means that $(n,m\rightarrow\infty)$ under a rate condition $R$ (such as the one given in Assumption 1).

theoremLet the conditional distribution of a bootstrap statistic $ \mathscr{T} _{n,m}^{\ast}$ given the data $D_{n}$ converge in probability to a nonrandom distribution as $(n,m\rightarrow\infty,R)$. Then, as $(n,m\rightarrow \infty,R)$: \begin{description} • (a) for measurable real functions $f$ and continuous bounded real functions $g$ with matching domains, \[ \sup\nolimits_{\Vert f\Vert_{\infty}\leq1}\left\vert \mathbb{E[}f(D_{n})g( \mathscr{T} _{n,m}^{\ast})]-\mathbb{E[}f(D_{n})]\mathbb{E[}g( \mathscr{T} _{n,m}^{\ast})]\right\vert \rightarrow0\text{;} \] • (b)\ for non-negligible continuity sets $B$ of $ \mathscr{T} _{n,m}^{\ast}$'s limit distribution, \[ \sup\nolimits_{A\in\sigma(D_{n})}|\mathbb{P}(A| \mathscr{T} _{n,m}^{\ast}\in B)-\mathbb{P}(A)|\rightarrow0, \] where $\sigma(D_{n})$ is the $\sigma$-algebra generated by the data $D_{n}$. \end{description}

A non-negligible continuity set is a set with positive probability whose boundary has zero probability. For instance, if the limit distribution of $ \mathscr{T} _{n,m}^{\ast}$ has a continuous cdf $H$, then $B=[0,t]$ is a non-negligible continuity set whenever $t>0$ is such that $H(t)>0$. This gives rise to the following corollary relevant for testing.

corollaryConsider any statistic $\hat{\rho}_{n}\in\mathbb{R}$, measurable with respect to the data $D_{n}$ and with asymptotic cdf $F_{\rho}$ as $n\rightarrow\infty$. Let the conditional distribution of $ \mathscr{T} _{n,m}^{\ast}$ given the data converge in probability to a nonrandom distribution with a continuous cdf $H$ as $(n,m\rightarrow\infty,R)$. Then, for every continuity point $s$ of $F_{\rho}$ any every $t$ with $H(t)>0$, it holds that, as $(n,m\rightarrow\infty,R)$, \[ \sup_{s\in\mathbb{R}}|\mathbb{P}(\hat{\rho}_{n}\leq s| \mathscr{T} _{n,m}^{\ast}\leq t)-F_{\rho}(s)|\rightarrow0. \]

Under the conditions of Theorem (ref), including $T_{n}^{\ast} \overset{d^{\ast}}{\rightarrow}_{p} \mathscr{N}\! (0,1)$, the diagnostic statistic $ \mathscr{T} _{n,m}^{\ast}$ satisfies the assumptions of Corollary (ref) with $H=\Phi$ and $R$ being the rate condition in Assumption (ref). Therefore, in large samples the finite-sample quantiles of $\hat{\rho}_{n}$ conditional on $ \mathscr{T} _{n,m}^{\ast}$ being below (or above) a given critical value\ $t$ approximate well the unconditional quantiles of $\hat{\rho}_{n}$'s asymptotic distribution. Hence, should the diagnostics based on $ \mathscr{T} _{n,m}^{\ast}$ correctly fail to reject specification validity, inference based on $\hat{\rho}_{n}$ is free of pre-testing bias as $n\rightarrow\infty$.

an example based on instrumental variables (cont'd)

We now briefly show Theorem (ref) and Corollary (ref) in action by considering the 2SLS $t$\nobreakdash-test for the null $\beta=\beta_{0}$, denoted $t_{\mathsf{IV} ,n}$, in the setting of the IV\ example of Sections (ref) and (ref), assuming that $z_{t}$ is strong (i.e., a valid specification), so that $t_{\mathsf{IV},n}\rightarrow_{d} \mathscr{N}\! (0,1)$. Indeed, an extensive literature on IV regression documents substantial bias in pre-tests of instrument strength (see, e.g., Andrews, Stock, and Sun, 2019), making it natural to ask whether the bootstrap diagnostic test introduces any such bias.

We compare two different scenarios. In the first one, the researcher runs the $t$\nobreakdash-test without pre-testing valid specification. In the second one, the researcher initially (pre-)tests for valid specification using the bootstrap diagnostic test $ \mathscr{T} _{n,m}^{\ast}$, and then runs the $t$\nobreakdash-test only if the bootstrap test does not reject valid specification. According to Corollary (ref) applied with $\hat{\rho}_{n}=t_{\mathsf{IV} ,n}$ and $F_{\rho}=\Phi$, the unconditional rejection probabilities of the $t$\nobreakdash-test (computed without pre-testing valid specification) and the conditional rejection probabilities (computed conditionally on the bootstrap diagnostic test failing to reject) should coincide as $(n,m\rightarrow\infty)$ at proper relative rates, see Section (ref). These probabilities are estimated in Table (ref) using the same Monte Carlo design as in Section (ref). The unconditional and the conditional ERPs are virtually indistinguishable, thus supporting the results in Corollary (ref).

table[table omitted — 1,502 chars of source]

extensions

diagnostics reporting

An intrinsic feature of any diagnostics based on $ \mathscr{T} _{n,m}^{\ast}$ is that the associated p-value, say $p_{n,m}^{\ast}$, is not $D_{n}$-measurable, as it depends also on the realization of the auxiliary bootstrap variates $W_{n}^{\ast}$ used to generate $ \mathscr{T} _{n,m}^{\ast}$; see Section (ref). This differs from ordinary bootstrap inference, where the bootstrap p-values are usually regarded as measurable with respect to the data $D_{n}$, as is the case where $(m,n\rightarrow\infty)_{\text{seq}}$. Moreover, under the conditions of Theorem (ref), the dependence of $p_{n,m}^{\ast}$ on the bootstrap variates does not vanish as $(n,m\rightarrow\infty)$, in the sense that $p_{n,m}^{\ast(1)}-p_{n,m}^{\ast(2)}$ need not go to zero for two conditionally independent copies $p_{n,m}^{\ast(1)}\ $and $p_{n,m}^{\ast(2)}$ of $p_{n,m}^{\ast}$.

As a consequence, if the bootstrap sample is not given ex ante but, rather, is generated by the practitioner, then for any given data $D_{n}$ different researchers may end up obtaining different p-values $p_{n,m}^{\ast}$. This may create issues in terms of reproducibility, as it is unclear which value of $p_{n,m}^{\ast}$ should be used for decision-making and reporting.\footnote{In fact, in some applications the set of bootstrap draws used to compute $ \mathscr{T} _{n,m}^{\ast}$ is given as an input for the diagnostic procedure. This, for instance, may happen when a replication package includes the set of bootstrap repetitions used to compute a bootstrap statistic but not the code used to generate them. In such cases, unless the set of repetitions are split into subsamples, only one realization of $p_{n,B}^{\ast}$ can be computed.}

A viable way to mitigate this issue is to use the convergence fact $p_{n,m}^{\ast}\overset{d^{\ast}}{\rightarrow}_{p} \mathscr{U}\! _{[0,1]}$ under specification validity, see Theorem (ref). This convergence and Theorem (ref) have the following corollary in terms of $\hat{\pi}_{n,m}(\eta):=\mathbb{P}^{\ast}(p_{n,m}^{\ast}\leq\eta)$, $\eta \in(0,1)$.

corollaryUnder the assumptions of Theorem (ref), $\hat{\pi}_{n,m}(\eta )\rightarrow_{p}\eta$ for any $\eta\in(0,1)$. In contrast, under the assumptions of Theorem (ref), $\hat{\pi}_{n,m}(\eta)\rightarrow_{p}1$.

Hence, rather than reporting a single draw $p_{n,m}^{\ast}$, the diagnostic procedure could additionally include an informal check of whether $\hat{\pi }_{n,m}(\eta)$ substantially exceeds the user-chosen significance level $\eta $. In practice $\hat{\pi}_{n,m}(\eta)$ is not known but, as is standard, it can be approximated with any desired precision by drawing an arbitrarily large number $K$ of i.i.d. (conditionally on the data) realizations of $p_{n,m}^{\ast}$, say $p_{n,m:k}^{\ast}$, $k=1,\ldots K$, and letting $\hat{\pi}_{n,m,K}^{\ast}:=K^{-1}\sum_{k=1}^{K}\mathbb{I}_{\{p_{n,m:k}^{\ast }\leq\eta\}}$.

In principle, the significance of $\hat{\pi}_{n,m,K}^{\ast}(\eta)-\eta$ could be assessed formally by means of a test, using the (pointwise) standard error $(\hat{\pi}_{n,m,K}^{\ast}(\eta)(1-\hat{\pi}_{n,m,K}^{\ast}(\eta))/K)^{1/2}$ or, with the null imposed, $(\eta\left( 1-\eta\right) /K)^{1/2}$. Also a uniform test over $\eta$ could be performed, using the convergence $\sqrt {K}\sup_{\eta\in\lbrack0,1]}|(\hat{\pi}_{n,m,K}^{\ast}(\eta)-\eta )|\overset{w^{\ast}}{\rightarrow}_{p}\sup_{[0,1]}|W|$ as $(n,m,K\rightarrow \infty)$ implied by Proposition (ref) below. Under valid specification, however, such tests would reproduce the problem of outcome dependence on the bootstrap variates employed, here $W_{n,k}^{\ast}$ ($k=1,...,K$), and results would again differ across researchers.

propositionUnder the assumptions of Theorem (ref), $\sqrt{K}(\hat{\pi}_{n,m,K}^{\ast}(\cdot)-(\cdot))\overset{w^{\ast} }{\rightarrow}_{p}W(\cdot)$ on $\mathscr{D}[0,1]$ as $(n,m,K\rightarrow \infty)$, where $W$ denotes a standard Brownian bridge.

alternative discrepancy measures

The discussion so far has focused on the KS or uniform norm, and the norms it dominates. Apart from, e.g., the mentioned signed KS and Cramer-von Mises norms, also the seminorm $\sup_{A}|\cdot|$ for an interval $A$ (e.g., $A:=[1.96,\infty)$), leading to $\hat{d}_{n}(A)=\sup_{u\in A}|\hat{G}_{n}(u)-\Phi(u)|$, is dominated by the KS norm, and similarly $\hat{d}_{n}(x):=|\hat{G}_{n}(x)-\Phi(x)|$ obtained by focusing on a single point in the support (e.g., $x=1.96$). They all satisfy the rate condition in Assumption (ref) whenever the KS norm does so, and are covered by the theory.

In addition, range-based, quantile-based discrepancy measures or measures focusing on specific moments of the bootstrap distributions could be used. For instance, a moment-based discrepancy based on the third and fourth moments can be defined as \[ \hat{d}_{n}:=\Vert v_{n}\Vert_{\Omega}:=v_{n}^{\prime}\Omega v_{n}\text{, }v_{n}:=\int_{\mathbb{R}}(u^{3},u^{4})^{\prime}(d\hat{G}_{n}(u)-d\Phi (u))=\mathbb{E}^{\ast}(T_{n}^{\ast3},T_{n}^{\ast4}-3)^{\prime} \] with $\Omega$ a symmetric p.d. matrix. Such discrepancy measures need not be continuous on $\mathscr{D}_{\mathbb{R}}$ and a strengthening of ((ref)) along the lines discussed, e.g., in Hahn and Liao (2021) may be required for their convergence to zero in probability.

A minimal justification for the use of $\Vert\cdot\Vert_{\Omega}$ in applications is that the ensuing $\hat{d}_{n,m}^{\ast}$ can be written as $\hat{d}_{n,m}^{\ast}=\psi_{m}(T_{n:1}^{\ast},...,T_{n:m}^{\ast})$ for a continuous $\psi_{m}\nolinebreak:\nolinebreak\mathbb{R}^{m}\nolinebreak \rightarrow\nolinebreak\mathbb{R}$. Under the null that $T_{n}^{\ast} \overset{d^{\ast}}{\rightarrow}_{p}Z\sim \mathscr{N}\! (0,1)$, it holds that $\hat{d}_{n,m}^{\ast}\overset{d^{\ast}}{\rightarrow} _{p}d_{m}^{\ast}:=\psi_{m}(Z_{1},...,Z_{m})$ as $n\rightarrow\infty$, with the $Z_{i}$'s independent $ \mathscr{N}\! (0,1)$, $i=1,...,m$. Therefore, for large $n$, tests could be conducted using the finite-$m$ quantiles of $\psi_{m}(Z_{1},...,Z_{m})$, or even its large $m$ asymptotic quantiles. Similar considerations apply to the popular Shapiro-Wilks statistic.

alternative null hypotheses

The discrepancy measures discussed so far compare the bootstrap cdf with the $ \mathscr{N}\! (0,1)$ cdf. Suppose, however, that interest is in assessing the convergence $T_{n}^{\ast}\overset{d^{\ast}}{\rightarrow}_{p} \mathscr{N}\! (0,\sigma^{2})$ for some unspecified $\sigma^{2}>0$. This could be done by redefining the reference statistic as $\tilde{T}_{n}^{\ast}:=T_{n}^{\ast} /\hat{\sigma}_{n}$, $\hat{\sigma}_{n}^{2}:=\mathbb{V}^{\ast}[T_{n}^{\ast}]$, provided $\hat{\sigma}_{n}^{2}$ is consistent\footnote{The condition $\mathbb{E}^{\ast}|T_{n}^{\ast}|^{2+\epsilon}=O_{p}(1)$ for some $\epsilon>0$ suffices.} for $\sigma^{2}$, such that $\tilde{T}_{n}^{\ast}\overset{d^{\ast} }{\rightarrow}_{p} \mathscr{N}\! (0,1)$. Similarly, assessing the convergence $T_{n}^{\ast}\overset{d^{\ast} }{\rightarrow}_{p} \mathscr{N}\! (\mu,\sigma^{2})$ with unknown $\mu$ and $\sigma^{2}>0$ can be done by considering $\check{T}_{n}^{\ast}:=(T_{n}^{\ast}-\hat{\mu}_{n})/\hat{\sigma }_{n},$ $\hat{\mu}_{n}:=\mathbb{E}^{\ast}[T_{n}^{\ast}]$, such that $\check {T}_{n}^{\ast}\overset{d^{\ast}}{\rightarrow}_{p} \mathscr{N}\! (0,1)$ if $(\hat{\mu}_{n},\hat{\sigma}_{n}^{2})$ is consistent for $(\mu,\sigma^{2})$. Here both $\hat{\mu}_{n}$ and $\hat{\sigma}_{n}^{2}$ can be calculated with arbitrary precision by using a sufficiently large number $M\gg m$ of bootstrap repetitions.

Because $\hat{\mu}_{n}\ $and $\hat{\sigma}_{n}^{2}$ are functions of the data (i.e., $D_{n}$-measurable), the theory of Section (ref) can be applied to $\tilde{T}_{n}^{\ast}$ and $\check{T}_{n}^{\ast}$, such that under the conditions of Theorem (ref) their asymptotic distributions are not affected by uncertainty due to the estimation of $\sigma^{2}$ (and $\mu$). Specifically, maintaining the notation $\hat{G}_{n}$ for the bootstrap cdf of $T_{n}^{\ast}$, the rate condition in Assumption (ref) boils down to $||\hat{G}(\cdot\hat{\sigma}_{n})-\Phi\left( \cdot\right) ||=O_{p} (n^{-\alpha})$ and $||\hat{G}(\cdot\hat{\sigma}_{n}+\hat{\mu}_{n})-\Phi\left( \cdot\right) ||=O_{p}(n^{-\alpha})$ respectively for $\tilde{T}_{n}^{\ast}$ and $\check{T}_{n}^{\ast}$. The condition can be checked either directly or by using the estimates

align*[align* omitted — 394 chars of source]

which hold whenever $\hat{\sigma}_{n}^{2}$ and $\hat{\mu}_{n}$ are consistent. Then, if $||\hat{G}(\cdot\sigma)-\Phi\left( \cdot\right) ||=O_{p} (n^{-\alpha})$ and $\hat{\sigma}_{n}^{2}-\sigma^{2}=O_{p}(n^{-1/2})$, it follows that $||\hat{G}(\cdot\hat{\sigma}_{n})-\Phi\left( \cdot\right) ||=O_{p}(n^{-\min\{\alpha,1/2\}})$, and similarly in the case of additional centering.

Finally, statistics such as $ \mathscr{T} _{n,m}^{\ast}:=m^{1/2}\breve{d}_{n,m}^{\ast}$, $\breve{d}_{n,m}^{\ast} :=\Vert\breve{G}_{n,m}^{\ast}-\Phi\Vert$, could also be implemented, with $\breve{G}_{n,m}^{\ast}$ the edf of the standardized bootstrap sample $(T_{n:i}^{\ast}-m_{n,m}^{\ast})/s_{n,m}^{\ast}$, $i=1,\ldots,m$, and $m_{n,m}^{\ast}$ and $s_{n,m}^{\ast}$ the sample mean and standard deviation of $T_{n:1}^{\ast},\ldots,T_{n,m}^{\ast}$, respectively. The resulting tests, which would be similar to Lilliefors' normality test, are not covered by the theory in Section (ref) because $(T_{n:i}^{\ast}-m_{n,m}^{\ast })/s_{n,m}^{\ast}$ are not conditionally i.i.d. given the data. The continuity considerations of Section (ref) carry over, however.

applications

We introduce here some additional applications where the bootstrap can be used to detect specification invalidity. These applications, which are kept deliberately simple, focus on detecting failures of the following assumptions: (i) stationarity, (ii) parameter in the interior of the parameter space, (iii) finite variance, and (iv)\ non-singular Jacobian in applications of the delta method. For each application, we discuss the set up and conditions ensuring that the main theory results from Sections (ref) and (ref) hold.

non-stationarity

Assume that the data are generated by a standard autoregressive recursion: \[ y_{t}=\phi_{0}y_{t-1}+\varepsilon_{t}\text{, }t=1,\ldots,n \] where $\{\varepsilon_{t}\}_{t=1}^{\infty}$ is an i.i.d. sequence r.v.'s with $\mathbb{E[}\varepsilon_{t}]=0$ and $\sigma^{2}:=\mathbb{E[}\varepsilon _{t}^{2}]=1$; $y_{0}$ is assumed to be fixed. The least squares estimator of $\phi_{0}$ is $\hat{\phi}_{n}:=\sum_{t=1}^{n}y_{t}y_{t-1}/\sum_{t=1} ^{n}y_{t-1}^{2}$; under the stability condition $|\phi_{0}|<1$, it holds that \[ T_{n}:=\frac{\hat{\phi}_{n}-\phi_{0}}{\operatorname*{se}(\hat{\phi}_{n} )}\overset{d}{\rightarrow}Z\sim \mathscr{N}\! (0,1) \] where $\operatorname*{se}(\hat{\phi}_{n}):=\hat{\sigma}_{n}(\sum_{t=1} ^{n}y_{t-1}^{2})^{-1/2}$, $\hat{\sigma}_{n}^{2}$ being the sample variance of $\hat{\varepsilon}_{t}:=y_{t}-\hat{\phi}_{n}y_{t-1}$. A simple parametric, recursive bootstrap generates the bootstrap data as follows: \[ y_{t}^{\ast}=\hat{\phi}_{n}y_{t-1}^{\ast}+\varepsilon_{t}^{\ast},\text{ }\varepsilon_{t}^{\ast}\sim\text{i.i.d.} \mathscr{N}\! (0,1)\text{, }t=1,\ldots n, \] initialized at $y_{0}^{\ast}:=y_{0}$, with $\{\varepsilon_{t}^{\ast} \}_{t=1}^{\infty}$ independent of the original data. The bootstrap analog of $T_{n}$ satisfies, as $n\rightarrow\infty$, \[ T_{n}^{\ast}:=\frac{\hat{\phi}_{n}^{\ast}-\hat{\phi}_{n}}{\operatorname*{se} (\hat{\phi}_{n}^{\ast})}\overset{d^{\ast}}{\rightarrow}_{p}Z\sim \mathscr{N}\! (0,1), \] $\hat{\phi}_{n}^{\ast}$ and $\operatorname*{se}(\hat{\phi}_{n}^{\ast})$ being the bootstrap analogs of $\hat{\phi}_{n}$ and $\operatorname*{se}(\hat{\phi }_{n})$, respectively; see Bose (1988).

Suppose now that $\phi_{0}=1$ or, more generally, that $\phi_{0}=\phi _{n}=1+\lambda n^{-1}$, $\lambda\in\mathbb{R}$, such that the data are non-stationary. Then \[ T_{n}\overset{d}{\rightarrow}\xi(\lambda,B):=\Big(\int_{0}^{1}J_{\lambda }(u)^{2}du\Big)^{-1/2}\int_{0}^{1}J_{\lambda}(u)dB(u) \] where $B$ is a standard Brownian motion on $\mathscr{D}{}_{[0,1]}$ and $J_{\lambda}$ the Ornstein-Uhlenbeck process on $\mathscr{D}{}_{[0,1]}$ satisfying the stochastic differential equation $dJ=\lambda J+dB$ (see, e.g., Phillips, 1987, or Andrews and Guggenberger, 2009). An implication is that the bootstrap statistic $T_{n}^{\ast}$ has a non-Gaussian, random limit:

equation[equation omitted — 114 chars of source]

where the random term $\ell\sim\int_{0}^{1}J_{\lambda}(u)dB(u)/\int_{0} ^{1}J_{\lambda}(u)^{2}du$ arises from the weak convergence $n(\hat{\phi} _{n}-\phi_{0})\overset{d}{\rightarrow}\ell$ and $B^{\ast}$ is a standard Brownian motion on $[0,1]$, independent of $\ell$; see Basawa et al. (1991). In terms of cdf's, ((ref)) is equivalent to the weak convergence in $\mathscr{D}{}_{\mathbb{R}}$

equation[equation omitted — 204 chars of source]

Again, $ \mathscr{G}\! \neq\Phi$ a.s. as asymptotic normality of the bootstrap statistic fails.

\paragraph*{validity of the diagnostic procedure.}

Bose (1988) provides sufficient conditions for the nonparametric bootstrap based on least-squares residuals to admit an Edgeworth expansion. These primarily require that $\mathbb{E}[\varepsilon_{t}^{8}]<\infty$ and that the autoregressive characteristic roots (i.e., $1/\phi_{0}$) lie outside the unit circle; see his conditions (A.1)--(A.3). In particular, these conditions imply that $\parallel\hat{G}_{n}-\Phi\parallel_{\infty}=O_{p}(n^{-1/2})$; hence, Assumption (ref) is satisfied with $\alpha=1/2$ for the KS norm and dominated norms. Consequently, by Theorem (ref) validity of the diagnostic procedure for the null hypothesis of stationarity requires that $m/n\rightarrow0$ as $(n,m\rightarrow\infty)$. If, as suggested earlier in this section, the Gaussian parametric bootstrap is used instead, then (A.1) and (A.2) in Bose (1988)\ are automatically satisfied on the bootstrap data, irrespective of the properties of the original $\varepsilon_{t}$'s, provided that $\hat{\phi}_{n}$, which is used to generate the bootstrap data recursively, is consistent.

Under non-stationarity ($\phi_{0}=\phi_{n}=1+\lambda n^{-1}$), it holds that $\parallel\hat{G}_{n}-\Phi\parallel_{\infty}\overset{d}{\rightarrow} \mathscr{Y} :=\parallel \mathscr{G}\! -\Phi\parallel_{\infty}>0$ with $ \mathscr{G}\! $ as in ((ref)), and Theorem (ref) guarantees that non-stationarity can be detected with probability approaching one by tests employing the KS or dominated norms.

parameter on the boundary

As in Andrews (2000), consider an i.i.d. sample $D_{n}:=\{y_{i}\}_{i=1}^{n}$ from a population with $\mathbb{E[}y_{i}]=:\theta_{0}\in\Theta:=[0,\infty)$, $\Theta$ being the parameter space, and $\mathbb{V[}y_{i}^{2}]=1$. The Gaussian quasi-maximum likelihood estimator [QMLE], given by $\hat{\theta} _{n}:=\max\{0,\bar{y}_{n}\}$ with $\bar{y}_{n}:=n^{-1} {\textstyle\sum\nolimits_{i=1}^{n}} y_{i}$, is such that, when $\theta_{0}\in\operatorname*{int}\Theta$,

equation[equation omitted — 216 chars of source]

Consider the Gaussian parametric bootstrap MLE, $\hat{\theta}_{n}^{\ast} :=\max\{0,\bar{y}_{n}^{\ast}\}$, where conditionally on the original data $D_{n}$, $\bar{y}_{n}^{\ast}\sim \mathscr{N}\! (\hat{\theta}_{n},n^{-1/2})$. The bootstrap analog of $T_{n}$ is $T_{n}^{\ast }:=\sqrt{n}(\hat{\theta}_{n}^{\ast}-\hat{\theta}_{n})=\max\{-\sqrt{n} \hat{\theta}_{n},$ $\sqrt{n}(\bar{y}_{n}^{\ast}-\hat{\theta}_{n})\}$ and satisfies \[ T_{n}^{\ast}|D_{n}\sim\max\{-\sqrt{n}\hat{\theta}_{n},Z^{\ast}\}\text{, }Z^{\ast}\sim \mathscr{N}\! (0,1) \] with conditional cdf $\hat{G}_{n}(x)=\Phi(x)\mathbb{I}_{\mathbb{\{}x\geq -\sqrt{n}\hat{\theta}_{n}\}}$. When $\theta_{0}$ is an interior point of $\Theta$, since $\sqrt{n}\hat{\theta}_{n}\rightarrow-\infty$ (a.s.) as $n\rightarrow\infty$, it holds that

equation[equation omitted — 141 chars of source]

and the bootstrap replicates the standard normal distribution.

Suppose now that $\theta_{0}$ is on the boundary ($\theta_{0}=0$) or `close' to it ($\theta_{0}=\lambda n^{-1/2}$, $\lambda\in\lbrack0,\infty)$). In this case, ((ref)) no longer holds; instead \[ T_{n}=\max\{-\sqrt{n}\theta_{0},\sqrt{n}(\bar{y}_{n}-\theta_{0})\}\overset {d}{\rightarrow}\xi(\lambda):=\max\{-\lambda,Z\} \] which differs from a normal random variable. Similarly, and in contrast to ((ref)), the bootstrap cdf satisfies $\hat{G}_{n} (\cdot)=\Phi(\cdot)\mathbb{I}_{\mathbb{\{}\cdot\geq-\sqrt{n}\hat{\theta} _{n}\}}\rightarrow_{w}\Phi(\cdot)\mathbb{I}_{\mathbb{\{}\cdot\geq-\ell\}}$ in $\mathscr{D}{}_{\mathbb{R}}$, which is established using the convergence fact $\sqrt{n}\hat{\theta}_{n}=\sqrt{n}(\hat{\theta}_{n}-\theta_{0})+\sqrt{n} \theta_{0}\overset{d}{\rightarrow}\xi(\lambda)+\lambda=:\ell$. This is equivalent to the weak convergence in distribution \[ T_{n}^{\ast}\overset{d^{\ast}}{\rightarrow}_{w}\xi^{\ast}(\ell)|\ell \] where $\xi^{\ast}(\ell):=\max\{-\ell,Z^{\ast}\}$, $Z^{\ast}\sim \mathscr{N}\! (0,1)$ independent of $\ell$. Since $\ell$ is finite a.s., the conditional cdf of $T_{n}^{\ast}$ is non-Gaussian in the limit, with probability one.

\paragraph*{validity of the diagnostic procedure.}

As seen above, conditionally on the data the parametric bootstrap statistic $T_{n}^{\ast}$ has conditional cdf $\hat{G}_{n}(x)=\Phi(x)\mathbb{I} _{\{x\geq-\sqrt{n}\hat{\theta}_{n}\}}$, and the associated KS \ distance $\hat{d}_{n}:=\parallel\hat{G}_{n}-\Phi\parallel_{\infty}$ satisfies \[ \hat{d}_{n}=\sup_{x\in\mathbb{R}}|\Phi(x)\mathbb{I}_{\{x\geq-\sqrt{n} \hat{\theta}_{n}\}}-\Phi(x)|=\sup_{x\in\mathbb{R}}|\Phi(x)\mathbb{I} _{\{x\leq-\sqrt{n}\hat{\theta}_{n}\}}|=\Phi(-\sqrt{n}\hat{\theta}_{n})\text{.} \] Under the null that $\theta_{0}$ is a fixed interior point ($\theta_{0}>0$), $\hat{d}_{n}\overset{p}{\rightarrow}0$ exponentially fast, and the KS norm satisfies Assumption (ref) for any $\alpha>0$. This implies that any power growth rate of $m$ in terms of $n$ is allowed in Theorem (ref) for tests employing the KS \ distance.

In contrast, when $\theta_{0}$ is on (or near) the boundary, $\hat{d}_{n}$ satisfies \[ \hat{d}_{n}=\Phi(-\sqrt{n}\hat{\theta}_{n})\overset{d}{\rightarrow} \mathscr{Y} =\Phi\left( -\ell\right) , \] with $\ell$ as previously defined. Thus, $\hat{d}_{n,m}^{\ast}$ diverges at the rate of $\sqrt{m}$ as $(n,m\rightarrow\infty)$.

infinite variance

As in Section (ref), assume that $D_{n}:=\{y_{i} \}_{i=1}^{n}$ where the $y_{i}$'s are i.i.d. with $\theta_{0}:=\mathbb{E[} y_{i}]$ and $\sigma^{2}:=\mathbb{V[}y_{i}^{2}]\in(0,\infty)$. This assumption implies that the sample mean $\hat{\theta}_{n}:=n^{-1}\sum_{i=1}^{n}y_{i}$ satisfies $T_{n}:=\sqrt{n}(\hat{\theta}_{n}-\theta_{0})/\sigma\overset {d}{\rightarrow}Z$, $Z\sim \mathscr{N}\! (0,1)$, as $n\rightarrow\infty$. With $\{y_{i}^{\ast}\}_{i=1}^{n}$ denoting a sample of $n$ draws, independent conditionally on $D_{n}$, from the edf of $\{y_{i}\}_{i=1}^{n}$, the i.i.d. bootstrap counterpart of $T_{n}$ is $T_{n}^{\ast}:=\sqrt{n}(\hat{\theta}_{n}^{\ast}-\hat{\theta}_{n})/\hat{\sigma }_{n}$, where $\hat{\theta}_{n}^{\ast}:=n^{-1}\sum_{i=1}^{n}y_{i}^{\ast}$, $\hat{\sigma}_{n}^{2}:=n^{-1}\sum_{i=1}^{n}(y_{i}-\hat{\theta}_{n})^{2}$. As is known, as $n\rightarrow\infty$ (e.g., Singh, 1981), $T_{n}^{\ast} \overset{d^{\ast}}{\rightarrow}_{p}Z$, $Z\sim \mathscr{N}\! (0,1)$; in terms of cdfs, $\hat{G}_{n}(x):=\mathbb{P}^{\ast}(T_{n}^{\ast}\leq x)\overset{p}{\rightarrow}\Phi(x)$.

Consider now the case where $\mathbb{E[}y_{t}^{2}]=\infty$. Specifically, assume that $y_{t}$ is in the domain of attraction of a symmetric stable law with tail index $\nu\in(1,2)$, denoted as $ \mathscr{S}\!\! $\thinspace$(\nu)$. In this case the CLT fails to hold; instead, for some diverging real sequence $\{a_{n}\}$, \[ a_{n}n^{1/2}T_{n}=a_{n}n(\hat{\theta}_{n}-\theta_{0})\overset{d}{\rightarrow} \mathscr{S}\!\! \,(\nu); \] see, e.g., Feller (1971). Note that $ \mathscr{S}\!\! \,\left( \nu\right) $ can be written as $ \mathscr{S}\!\! $\thinspace$(\nu)\sim\sum\nolimits_{k=1}^{\infty}\delta_{k}Z_{k}$, where the $\delta_{k}$'s are i.i.d. Rademacher and $Z_{k}^{1/\nu}:=\sum_{i=1}^{k}E_{i}$ with $\{E_{i}\}_{i=1}^{\infty}$ an i.i.d. sequence of exponential r.v.'s with $\mathbb{E}[E_{i}]=1$; see Lepage, Woodroofe and Zinn (1981). Athreya (1987) and Knight (1989) show that it this case also the i.i.d. bootstrap counterpart of $T_{n}$ is not asymptotically Gaussian; precisely, the bootstrap measure is random in the limit:

equation[equation omitted — 281 chars of source]

where the $\delta_{k}$'s and $Z_{k}$'s are as previously defined. The $M_{k}^{\ast}$'s, which are i.i.d. $ \mathscr{P}\! (1)$ r.v.'s, independent of $\{\delta_{k},Z_{k}\}$, induce the randomness in the limit bootstrap measure; see Theorem 2 in Knight (1989). If a wild bootstrap with Rademacher multipliers is used instead, that is, $\hat{\theta }_{n}^{\ast}:=\hat{\theta}_{n}+n^{-1}\sum_{i=1}^{n}(y_{i}-\hat{\theta} _{n})w_{i}^{\ast}$, with the $w_{i}^{\ast}$'s being i.i.d. Rademacher random variables conditionally on $D_{n}$, then $T_{n}^{\ast}\overset{d^{\ast} }{\rightarrow}_{w}\sum\nolimits_{k=1}^{\infty}\delta_{k}^{\ast}Z_{k} (\sum\nolimits_{k=1}^{\infty}Z_{k}^{2})^{-1/2} \big| \{Z_{k}\}$, see Cavaliere, Georgiev and Taylor (2013, 2016). For both bootstrap schemes, the asymptotic bootstrap measure is random and a.s. non-Gaussian.

\paragraph*{validity of the diagnostic procedure.}

Consider first the case of valid specification where the $y_{t}$'s have finite variance. As seen above, $\hat{G}_{n}(x):=\mathbb{P}^{\ast}(T_{n}^{\ast}\leq x)\overset{p}{\rightarrow}\Phi(x)$. The rate of the previous convergence, required as an input in Assumption (ref), can be established under slightly more than finite second moments.

Specifically, if $\mathbb{E[}|y_{t}|^{2\kappa}]<\infty$ for some $\kappa>1$, then by the Berry-Esseen bound \[ \Vert\hat{G}_{n}-\Phi\Vert_{\infty}\leq\frac{C}{n^{1/2}}\mathbb{E}^{\ast }[|y_{i}^{\ast}-\hat{\theta}_{n}|^{3}]=\frac{C}{n^{1/2}}\frac{1}{n}\sum _{i=1}^{n}|y_{i}-\hat{\theta}_{n}|^{3} \] For $\kappa\geq\frac{3}{2}$ (such that $\mathbb{E}|y_{i}|^{3}<\infty$), the r.h.s. of the previous equation is $O_{a.s.}(n^{-1/2})$, while for $\kappa \in\left( 1,3/2\right) $ it is $o_{a.s.}(n^{-\frac{3(\kappa-1)}{2\kappa}})$ by the Marcinkiewicz-Zigmund strong law. Hence, the KS norm satisfies Assumption (ref) with $\alpha=\min\{\frac{3(\kappa -1)}{2\kappa},\frac{1}{2}\}$ and Theorem (ref) applies. Agnostically, one may select $m$ as growing at a logarithmic rate, e.g., $m=\ln(n)$, which would satisfy the requirement in Theorem (ref) for any $\kappa>1$.

Consider now the case where $\mathbb{E[}y_{i}^{2}]=\infty$. Then, $T_{n} ^{\ast}$ has the random limit given in ((ref)); by Theorem 3 in Knight (1989), its random limiting cdf, $ \mathscr{G}\! $ say, is a.s. sample-path continuous. Hence, ((ref)) implies that $\Vert\hat{G}_{n}-\Phi\Vert_{\infty}\rightarrow_{d} \mathscr{Y} :=\Vert \mathscr{G}\! -\Phi\Vert_{\infty}>0$ a.s.; Theorem (ref) applies and the diagnostic test based on the KS norm rejects with probability approaching one.

near-singular Jacobian

Suppose that an estimator $\hat{\theta}_{n}$ of an unknown parameter $\theta_{0}$ satisfies

equation[equation omitted — 154 chars of source]

and a bootstrap analog $Z_{n}^{\ast}:=\sqrt{n}(\hat{\theta}_{n}^{\ast} -\hat{\theta}_{n})/\hat{\sigma}_{n}$ is available such that $Z_{n}^{\ast }\overset{d^{\ast}}{\rightarrow}_{p}Z^{\ast}\sim \mathscr{N}\! (0,1)$ jointly with ((ref)), with $Z^{\ast}$ independent of $Z$ and $\hat{\sigma}_{n}^{2}\overset{p}{\rightarrow} \tilde{\sigma}^{2}>0$ not necessarily equal to $\sigma^{2}$.

Consider inference on $\tau_{0}:=g(\theta_{0})$, where $g$ is twice continuously differentiable. With $\hat{\tau}_{n}:=g(\hat{\theta}_{n})$, let $T_{n}:=\sqrt{n}(\hat{\tau}_{n}-\tau_{0})$. By a second-order expansion around $\theta_{0}$, \[ T_{n}=\dot{g}Z_{n}+\tfrac{1}{2}g^{\prime\prime}(\bar{\theta}_{n} )n^{-1/2}\sigma^{2}Z_{n}^{2} \] where $\bar{\theta}_{n}$ is on the line segment between $\theta_{0}$ and $\hat{\theta}_{n}$, and $\dot{g}:=g^{\prime}(\theta_{0})$. Provided $\dot {g}\neq0$, it holds that $T_{n}\overset{d}{\rightarrow}\sigma\dot{g}Z$.

Now, consider the bootstrap analog of $T_{n}$. With $\hat{\tau}_{n}^{\ast }:=g(\hat{\theta}_{n}^{\ast})$ and $\hat{\tau}_{n}:=g(\hat{\theta}_{n})$, we have \[ T_{n}^{\ast}:=\sqrt{n}(\hat{\tau}_{n}^{\ast}-\hat{\tau}_{n})=g^{\prime} (\hat{\theta}_{n})\hat{\sigma}_{n}Z_{n}^{\ast}+\tfrac{1}{2}g^{\prime\prime }(\bar{\theta}_{n}^{\ast})n^{-1/2}\hat{\sigma}_{n}^{2}Z_{n}^{\ast2} \] for some $\bar{\theta}_{n}^{\ast}$ between $\hat{\theta}_{n}$ and $\hat {\theta}_{n}^{\ast}$. As for $T_{n}$, if $\dot{g}\neq0$ then $T_{n}^{\ast }\overset{d^{\ast}}{\rightarrow}_{p}\tilde{\sigma}\dot{g}Z$; hence, $T_{n}^{\ast}$ is asymptotically normal.

Suppose instead that $\dot{g}:=g^{\prime}(\theta_{0})=\lambda n^{-1/2}$, $\lambda\in\lbrack0,\infty)$, such that the Jacobian is singular ($\lambda=0$) or nearly singular ($\lambda\in\left( 0,\infty\right) $). Then $T_{n} =o_{p}(1)$ while, provided $\ddot{g}:=g^{\prime\prime}(\theta_{0})\neq0$, $\sqrt{n}T_{n}$ has a chi-square-type limit distribution: \[ \sqrt{n}T_{n}=\lambda\sigma Z_{n}+\tfrac{1}{2}g^{\prime\prime}(\hat{\theta }_{n}^{\ast})\sigma^{2}Z_{n}^{2}\overset{d}{\rightarrow}\lambda\sigma Z+\tfrac{1}{2}\ddot{g}\sigma^{2}Z^{2}. \] Similarly, using the convergence fact \[ \sqrt{n}g^{\prime}(\hat{\theta}_{n})=\sqrt{n}g^{\prime}(\theta_{0} )+g^{\prime\prime}(\theta_{0})\sqrt{n}(\hat{\theta}_{n}-\theta_{0} )+o_{p}(1)\overset{d}{\rightarrow}\ell:=\lambda+\ddot{g}\sigma Z, \] which holds jointly with the convergence of $Z_{n}^{\ast}$, it follows that $T_{n}^{\ast}$ satisfies

equation[equation omitted — 395 chars of source]

where $Z^{\ast}\sim \mathscr{N}\! (0,1)$ is independent of $\ell$. The right hand side of ((ref)) defines a random, non-Gaussian distribution, which we denote by $ \mathscr{G}\! $.

\paragraph*{validity of the diagnostic procedure.}

Assume first that the Jacobian is non-singular; i.e., $\ddot{g}\neq0$. If $Z_{n}^{\ast}$ admits a standard Edgeworth expansion, such that its conditional cdf, say $\hat{F}_{n}$, satisfies $\parallel\hat{F}_{n} -\Phi\parallel_{\infty}=O_{p}(n^{-1/2})$, then also $\hat{G}_{n}$, the conditional cdf of $T_{n}^{\ast}$, satisfies $\parallel\hat{G}_{n}(g^{\prime }(\hat{\theta}_{n})\hat{\sigma}_{n}\cdot)-\Phi(\cdot)\parallel_{\infty} =O_{p}(n^{-1/2})$; see Appendix (ref). As $\hat {G}_{n}(g^{\prime}(\hat{\theta}_{n})\hat{\sigma}_{n}\cdot)\ $is the edf of the bootstrap t-ratio $\tilde{T}_{n}^{\ast}:=T_{n}^{\ast}/(g^{\prime}(\hat{\theta }_{n})\hat{\sigma}_{n})$, which is well-defined with probability approaching one, it follows that Assumption (ref) is verified with $\alpha=1/2$ for $\tilde{T}_{n}^{\ast}$ and the KS norm, and Theorem (ref) applies for diagnostics based on $\tilde{T} _{n}^{\ast}$.

In contrast, when the Jacobian is near zero, $\sqrt{n}g^{\prime}(\hat{\theta }_{n})\overset{d}{\rightarrow}\ell$ and ((ref)) hold jointly. As a result, it is shown in Appendix (ref) that \[ \hat{d}_{n}:=\Vert\hat{G}_{n}(g^{\prime}(\hat{\theta}_{n})\hat{\sigma} _{n}\cdot)-\Phi\Vert_{\infty}\rightarrow1- \mathscr{G}\! (0)\mathbb{I}_{\{\ell\geq0\}}>0\text{ a.s.} \] Further, the conclusions of Theorem (ref) are valid for $\tilde{T} _{n}^{\ast}$-based diagnostics employing the KS norm; see again Appendix (ref).

an empirical illustration

To illustrate the usefulness and potential of our bootstrap diagnostic approach, in this section we consider an example from the empirical macroeconomic framework. We focus on the strategy employed by K\"{a}nzig (2021) to identify a structural oil supply news shock within a structural VAR identified with an external instrument (SVAR-IV), see Stock and Watson (2018). The novelty of K\"{a}nzig's (2021) analysis lies in his construction of an external instrument, $z_{t}$, for the oil supply news shock, $\varepsilon _{\text{oil},t}$, by extending the high-frequency (HF) approach, originally introduced by Gertler and Karadi (2015), outside the monetary policy framework. Specifically, in Kanzig's (2021) baseline model, the instrument $z_{t}$ is the first principal component of six time series which reflect variations in oil price futures with different maturities around OPEC production announcements.

K\"{a}nzig's (2021) model includes six variables ($g=6$): real oil prices, world oil production, world oil inventories, world industrial production, US industrial production, and the US consumer price index (CPI). These variables, observed monthly over the period 1974M1--2017M12 ($n=528$), are collected in the $g\times1$ vector $y_{t}$ and modeled through a VAR system with $p=12$ lags. Interest is in the dynamic responses of $y_{t+h}$ at horizon $h\in\{0,1,\ldots\}$ to an oil supply news shock of magnitude $\varepsilon _{\text{oil},t}:=\mathsf{x}$, measured as

equation[equation omitted — 276 chars of source]

where $\varepsilon_{t}:=(\varepsilon_{\text{oil},t},\varepsilon_{2,t}^{\prime })^{\prime}$, $\varepsilon_{2,t}$ being the vector of latent non-target (not of interest) shocks, $\mathcal{I}_{t-1}$ is the information set at time $t-1$, $\mathcal{C}_{\Pi}$ the\ VAR\ companion matrix, $R$ a selection matrix such that $R^{\prime}R=I_{g}$, and $b$ a $g\times1$ vector of structural parameters capturing the instantaneous response of the variables to the shock (note that $ \textsf{IRF} (0)=b\mathsf{x}$).

The dynamic causal effects in ((ref)) are identified and estimated using $z_{t}$ as instrument for $\varepsilon_{\text{oil},t}$. Under relevance and exogeneity of $z_{t}$, i.e., $\mathbb{E}[z_{t}\varepsilon _{\text{oil},t}]=\phi\neq0$ and $\mathbb{E}[z_{t}\varepsilon_{2,t}^{\prime }]=0$,

equation[equation omitted — 123 chars of source]

here, $b$ is partitioned conformably with the vector of structural shocks. The moment condition ((ref)) is the key ingredient for the estimation of ((ref)). With $u_{t}:=(u_{1,t},u_{2,t} ^{\prime})^{\prime}$ (and $y_{t}:=(y_{1,t},y_{2,t}^{\prime})^{\prime}$) also partitioned conformably, ((ref)) implies the representation:

equation[equation omitted — 127 chars of source]

where $\beta$ captures the (relative) instantaneous responses of the variables in $y_{2,t}$ to a supply news shock of magnitude $\varepsilon_{\text{oil} ,t}=\mathsf{x}=b_{1}^{-1}$, i.e., such that the instantaneous response of $y_{1,t}$ to $\varepsilon_{\text{oil},t}$ equals $1$ (the so-called `unit effect' normalization). The IV estimator of $\beta$ is given by $\hat{\beta }:=S_{\hat{u}_{1}z}^{-1}S_{\hat{u}_{2}z}$, where $\hat{u}_{t}:=(\hat{u} _{1,t},\hat{u}_{2,t}^{\prime})^{\prime}$, $t=1,...,n$, are the VAR residuals. For $\mathsf{x}=b_{1}^{-1}$, the IRFs ((ref))\ are estimated by replacing $\beta$ with $\hat{\beta}$ and the VAR\ companion matrix $\mathcal{C}_{\Pi}$ with the corresponding OLS estimate.

Similarly to the IV regression framework of Section (ref), Montiel Olea, Stock and Watson (2020) show that if $z_{t}$ is a weak instrument, the estimator of $ \textsf{IRF} (h)$ is inconsistent and asymptotically non-Gaussian. As is typical in the literature, K\"{a}nzig (2021) pre-tests the strength of the instrument $z_{t}$ using an F-test from a first-stage regression of $\hat{u}_{1,t}$ on $z_{t}$. The reported `regular' F-statistic in his Table 1 is $22.7$, which is above the homoskedastic threshold of $10.3$ (Stock and Yogo, 2005). In addition, K\"{a}nzig (2021) reports a `robust' F-statistic of $10.6$. Based on this evidence, which K\"{a}nzig (2021) considers supportive of $z_{t}$ being a strong instrument, the IRFs are estimated as described above, and their uncertainty is assessed using $68\%$ and $90\%$ moving block bootstrap confidence intervals (Jentsch and Lunsford, 2019, 2022), based on $10,000$ bootstrap replications. These confidence intervals, shown in K\"{a}nzig's Figure 3, appear broadly in line with the uncertainty commonly encountered by applied macroeconomists.

figure[figure omitted — 539 chars of source]

By exploiting K\"{a}nzig's bootstrap computations, we are able to reassess specification validity of his SVAR-IV model through the lens of bootstrap diagnostics. In this case, assuming correct specification of the VAR\ model, the null hypothesis of valid specification corresponds to the instrument $z_{t}$ being sufficiently strong for the target shock $\varepsilon _{\text{oil},t}$, a condition that ensures that IRFs are estimated consistently and standard asymptotic inference is valid.

Using K\"{a}nzig's (2021) Matlab code, we first estimate the $5$ bootstrap distributions $\hat{G}_{n}^{(j)}(\cdot):=\mathbb{P}^{\ast} (\hat{\beta}_{j}^{\ast}-\hat{\beta}_{j}\leq\cdot)$, $j=1,\ldots,5$, using $10,000$ moving block bootstrap replications of the vector $\hat{\beta}^{\ast }$ (properly standardized, see Section (ref)). The $\hat{G}_{n}^{(j)}$'s, reported in the upper panel of Figure (ref), are substantially non-Gaussian. For each $j=1,\ldots,5$, we run $K=500$ tests based on independent samples of $m\in\{10,20\}$ bootstrap replications; $m$ is deliberately small with respect to $n$. The lower panel of Figure (ref) reports, for significance levels $\eta\in\left( 0,0.1\right) $, the average rejection rates $\hat{\pi}_{n,m,K}^{\ast}(\eta):=K^{-1}\sum_{k=1}^{K}\mathbb{I} _{\mathbb{\{}p_{n,m:k}^{\ast}\leq\eta\}}$, where $p_{n,m:k}^{\ast},$ $k=1,\ldots,K$ are the p-values associated to the $K$ tests; see Section (ref). Note that due to use of the moving block bootstrap, these results are robust to the presence of conditional heteroskedasticity of unknown form in the data (Jentsch and Lunsford, 2022).

It can be observed that for all components of $\beta$, the null hypothesis is strongly rejected at any conventional significance level. Larger values of $m$ reinforce this conclusion. This result\footnote{Results are robust to (i) the choice of the horizon $h$, (ii) the choice of the norm and (iii)\ the standardization of the bootstrap estimator.} likely indicates that the SVAR-IV estimation is based on a weak instrument. Interestingly, neither K\"{a}nzig's (2021) $68\%$ and $90\%$ bootstrap confidence intervals nor the F-tests fully capture this phenomenon, which the bootstrap diagnostic test detects effectively.

concluding remarks

This paper shows that the bootstrap delivers, as a by-product, a versatile diagnostic procedure for detecting invalid specifications, i.e., violations of the assumptions underlying standard asymptotic theory. The procedure, which is based on measures of the discrepancy between the conditional distribution of a bootstrap statistic and the Gaussian distribution, provides a flexible and computationally straightforward alternative to traditional misspecification tests. A key advantage of this approach is that it does not induce pre-testing bias, thereby improving the reliability of post-test statistical inference. That is, under the null of valid specification, inference conditional on not rejecting the bootstrap tests is asymptotically exact.

A further major feature of this procedure is its flexibility. On the one hand, it can be tailored to test the validity of a specific assumption, such as stationarity of the data or relevance of a set of instrumental variables, hence allowing to construct tests specifically designed to have power against alternatives of interest. On the other hand, it can function as a more general diagnostic test for bootstrap consistency; that is, to test whether the distribution of a given bootstrap statistic is close to a theoretical Gaussian limit. This dual capability makes the method broadly applicable, whether testing for specific irregularities or assessing the validity of bootstrap procedures in general.

Our paper also contributes to the literature on bootstrap inference. The usual approach in bootstrap theory is to analyze the asymptotic properties of a bootstrap cdf, say $\hat{G}_{n}$, as the sample size $n$ diverges. In practice, however, $\hat{G}_{n}$ is estimated using the edf $\hat{G} _{n,m}^{\ast}$ of $m$ (conditionally i.i.d.) realization of the bootstrap statistic. Therefore, standard bootstrap theory is implicitly based on letting $m\rightarrow\infty$ first (such that the estimation error $\hat{G} _{n,m}^{\ast}-\hat{G}_{n}$ disappears), followed by $n\rightarrow\infty$. Our paper complements the standard approach by exploring the case where $n,m\rightarrow\infty$ jointly, rather than sequentially. We show that, under suitable conditions $\hat{G}_{n,m}^{\ast}$ has an asymptotic distribution which is known under a wide range of applications and, crucially, becomes asymptotically independent of the original data. This asymptotic independence result is key to avoid that the bootstrap test of valid specification distorts post-test inference.

The results in this paper can be extended in several directions. For instance, we have not discussed the role of the choice of the norm in determining the finite sample behavior and the power properties of the proposed tests. Second, an open question is how to efficiently choose $m$ in practice, such that an optimal balance of size and power is achieved. Third, all the examples considered here assume that the reference statistical model is of low dimension. It is of interest to establish whether our approach could be used to detect validity of the Gaussian asymptotic approximations in high-dimensional cases. Fourth, to what extent the properties of the bootstrap highlighted in this paper are shared by other methods such as the jackknife (see, e.g., Hansen, 2025), permutation tests (Young, 2019), subsampling (Politis, Romano and Wolf, 1999)\ or (quasi) Bayesian methods (as in Wang, 2025)\ are open questions. These are left to future research.