EconBase
← Back to paper

How Reliable are Bootstrap-based Heteroskedasticity Robust Tests?

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

153,215 characters · 19 sections · 69 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

How Reliable are Bootstrap-based Heteroskedasticity Robust Tests?

abstractWe develop theoretical finite-sample results concerning the size of wild bootstrap-based heteroskedasticity robust tests in linear regression models. In particular, these results provide an efficient diagnostic check, which can be used to weed out tests that are unreliable for a given testing problem in the sense that they overreject substantially. This allows us to assess the reliability of a large variety of wild bootstrap-based tests in an extensive numerical study.

Introduction

Testing hypotheses on the parameters in a regression model with potentially heteroskedastic errors is a time-honored problem in econometrics and statistics. As the classical $t$-statistic ($F$-statistic, respectively) is not pivotal, or asymptotically pivotal, in such a case in general, even under Gaussianity of the errors, so-called heteroskedasticity robust (aka heteroskedasticity consistent) modifications of these test statistics have been proposed. These statistics are asymptotically standard normally (chi-square, respectively) distributed under the null. The first generation of such procedures is rooted in the results of E63,E67, see also H77, and has been popularized in econometrics by W80. It soon transpired that tests obtained from these heteroskedasticity robust test statistics by relying on critical values obtained from the respective asymptotic distributions are prone to overrejecting the null hypothesis in finite samples, especially so if the design matrix contains leverage points; see, e.g., MacW85, DavidsonMacKinnon1985, and CheshJewitt1987. One factor contributing to this tendency to overreject is a downward bias present in the covariance matrix estimators used in these test statistics, see CheshJewitt1987. Attempts at remedying the overrejection problem have led to the development of second generation heteroskedasticity robust test statistics (often denoted by HC1 through HC4, with HC0 denoting the first generation test statistic). These statistics use various ways of rescaling the least-squares residuals before computing the covariance matrix estimator employed in the construction of the test statistic; see H77, MacW85, and Crib2004. Simulation studies reported in, e.g., DavidsonMacKinnon1985 and Crib2004 show that these modifications, especially HC3 and HC4, ameliorate the overrejection problem to some extent, but do not eliminate it. Further numerical results are provided in CheshAust_1991, see also Chesh_1989. DavidsonMacKinnon1985 also consider variants of HC0-HC3, denoted by HC0R-HC3R, obtained by using restricted instead of unrestricted least-squares residuals in the computation of the covariance matrix estimators employed by the various test statistics (the restriction alluded to being the restriction defining the null hypothesis). In their simulation experiments, this typically leads to tests that do not overreject (but that may underreject); see also the simulation results in Godfrey2006, who additionally also considers HC4R. Of course, these simulation results do not rule out that the tests based on HC0R-HC4R (relying on critical values suggested by asymptotic theory) may overreject in some situations outside of the scope of the simulation studies; in fact, PP5HC provide numerical proof that also these tests can suffer from considerable overrejection. Note that under the typical assumptions used in the literature all of the modifications of HC0 discussed so far have the same asymptotic distribution under the null, and thus use the same critical value.

An alternative approach is to use bootstrap methods to compute critical values for the test statistics HC0-HC4 or HC0R-HC4R, with the intention to improve upon the critical values derived from the asymptotic null distributions.\footnote{ Another possibility is to use Edgeworth expansions to find better critical values, see Rothenberg1988 for the case of the HC0 test statistic and DavidsonMacKinnon1985 for the HC0R test statistic. Simulation results in MacW85 and DavidsonMacKinnon1985 indicate that this does not work too well in practice. Of course, such expansions could also be worked out for the other versions of the test statistics mentioned, but this does not seem to have been pursued in the literature. Other adjustments are discussed in Imbkoles2016; however, as shown in PP5HC, these adjustments do also not resolve the overrejection problem in general.} Inspired by earlier work in the statistics literature (e.g., WuCFJ1986 and the discussion in beran1986, LiuR1988, Mammen1993), horowitz1997 used the wild bootstrap to obtain critical values for HC0. This was followed up by Flachaire1999, Flachaire2005 and DavidsFlach2008, who considered the test statistics HC0-HC3 as well as HC0R-HC3R, and who stressed the version of the wild bootstrap that imposes the null restriction on the bootstrap data generating process; see also GodfreyOrme2004 and the more recent survey Mackinnon2013. Further simulation studies, some of which also include HC4 and HC4R, can be found in Crib2004, CribariLima2009, Godfrey2006, and PRich2017. Once one has turned to bootstrap methods, one can also think of reverting to the classical (i.e., uncorrected) $t$-statistic ($F$-statistic, respectively) and to apply the bootstrap methods to determine appropriate critical values. This has been considered in Mammen1993; see also Godfrey2006 for some Monte Carlo results. Since the majority of the literature on bootstrap-based heteroskedasticity robust tests favors the wild bootstrap over other bootstrap methods, we shall concentrate on the wild bootstrap in the sequel.

While the before mentioned bootstrap procedures have their merits and overrejection is ameliorated in many of the cases considered in the simulation studies cited, it is unclear whether these observations generalize beyond the situations studied in these simulation experiments. In particular, it is unclear if -- and under which conditions -- these bootstrap procedures (and which of the many variants thereof) are immune to overrejection in finite samples.\footnote{ We are not interested in asymptotic justifications here.}$^{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{,}}$ \footnote{ In the quite special case where the number of restrictions tested equals the number of regression parameters,\ DavidsFlach2008 have a result which implies that certain wild bootstrap-based tests have size equal to the nominal significance level (and hence do not overreject) in finite samples. We note that this result in DavidsFlach2008 is not entirely correct as stated, but needs some amendments and corrections; see also Footnote (ref).} In the present paper we set out to study this question theoretically and numerically. On the theoretical side we show the following finite-sample result: For any test statistic $T$ from a large class of test statistics (including HC0-HC4, HC0R-HC4R, the classical $F$-statistic and a variant thereof that uses restricted residuals) and for any bootstrap method from a large class of wild bootstrap methods (including virtually all wild bootstrap methods considered in the literature) there is a computable number ${\Greekmath 0123} $ (depending only on observables like the design matrix, the restriction to be tested, etc),\ such that the size of the corresponding bootstrap-based test is $1$ for nominal significance levels ${\Greekmath 010B} $ satisfying ${\Greekmath 010B} >{\Greekmath 0123} $.\footnote{ The size is the maximal (i.e., worst-case) null rejection probability, where one maximizes over all possible forms of heteroskedasticity, reflecting agnosticism about the form of the heteroskedasticity; see ((ref)) for a formal definition. It is also assumed that the normalized regression error vector $(\mathbf{u}_{1}/\limfunc{Var}^{1/2}(\mathbf{u}_{1}),\ldots ,\mathbf{u }_{n}/\limfunc{Var}^{1/2}(\mathbf{u}_{n}))^{\prime }$ follows a given (fixed) distribution (e.g., a normal distribution). Note that the theoretical result mentioned in the main text does not depend on the choice of this distribution (as long as it is absolutley continuous), see Section (ref).} That is, for ${\Greekmath 010B} >{\Greekmath 0123} $ the bootstrap-based test fails miserably in that it has null rejection probabilities arbitrarily close to $1$ for some forms of heteroskedasticity. \footnote{ By construction ${\Greekmath 0123} \leq 1$ always holds. If ${\Greekmath 0123} <1$ (which will often be the case) we can then conclude that the bootstrap-based test has size $1$ at least for some values of ${\Greekmath 010B} $.} We note that our results also provide information concerning the infimal coverage probabilities of confidence sets obtained by \textquotedblleft inverting\textquotedblright\ the bootstrap-based tests under consideration. We discuss this in more detail in Remark (ref).

In practice our theoretical finite-sample result can be used as a diagnostic tool to weed out procedures in the following sense: as mentioned before, there is a large menu of heteroskedasticity robust test statistics and wild bootstrap methods available in the literature from which the applied researcher has to choose. As it is unlikely that simulation results in the literature are available that precisely fit the problem the researcher is interested in (i.e., use the same design matrix and the same restriction to be tested), the researcher is typically left with little guidance on which of the many bootstrap-based test procedures to choose for the problem at hand. Based on our theoretical results, the applied researcher can now eliminate procedures that break down in the researchers problem, by computing -- for any initially selected procedure -- the corresponding $ {\Greekmath 0123} $ for the given design matrix and restriction to be tested. If it turns out that ${\Greekmath 0123} <{\Greekmath 010B} $ holds, this procedure should not be used, because these tests have size equal to one according to our theoretical results. Numerical routines for computing ${\Greekmath 0123} $ are provided in the associated R-package wbsd by wbsd.\footnote{ As discussed in Section (ref) and Sections (ref), (ref) of Appendix (ref), computing ${\Greekmath 0123} $ is a nontrivial numerical problem. Supplementing the calculation of ${\Greekmath 0123} $ by numerically evaluating null rejection probabilities for strategically chosen heteroskedasticity structures as discussed in Section (ref) of Appendix (ref) may be advisable.}

The before mentioned theoretical result will typically have practical consequences only in testing problems for which ${\Greekmath 0123} $ is sufficiently small so that standard choices of ${\Greekmath 010B} $ like $0.05$ or $0.1$ satisfy the condition ${\Greekmath 010B} >{\Greekmath 0123} $. We hence investigate this numerically for the test statistics HC0-HC4, HC0R-HC4R, for the classical $F$-statistic, and for a variant thereof that uses restricted residuals, each combined with a large variety of wild bootstrap methods.\footnote{ In the wild bootstrap methods we vary the following elements: (i) centering the bootstrap sample at the unrestricted versus at the restriced least squares estimator, (ii) bootstrapping from unrestricted versus restricted residuals, (iii)\ the distribution of the bootstrap noise, and (iv) various bootstrap multiplicator weights. See Section (ref) for more details.} We now summarize the results of our numerical experiments for $n=10$ (the results for $n=20,30$ being similar): For each combination of test statistic and wild bootstrap method ($960$ combinations) we compute ${\Greekmath 0123} $ for a range of design matrices and null hypotheses (i.e., restrictions to be tested) and report ${\Greekmath 0123} _{\min }$, the smallest of these values of $ {\Greekmath 0123} $.\footnote{ For reasons of numerical stability we actually compute an upper bound for $ {\Greekmath 0123} $, see Section (ref) in Appendix (ref) for more information.}$^{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{,}}$\footnote{ Some of the $960$ combinations actually give rise to one and the same bootstrap-based test. The reasons for nevertheless considering all $960$ combinations are discussed in Sections (ref) and (ref). } We find that for $826$ of the $960$ combinations ${\Greekmath 0123} _{\min }$ is less than $0.05$, and for $936$ combinations ${\Greekmath 0123} _{\min }$ is less than $0.1$. As a consequence, for the bootstrap-based tests corresponding to these $826$ ($936$, respectively) combinations our theoretical results imply that size is equal to $1$ for some design matrices and null hypotheses, if a nominal significance level of $0.05$ ($0.1$, respectively) is being used. Thus these bootstrap-based tests are found not to be reliable in general, in that they suffer from severe overrejection for some design matrices and null hypotheses. Furthermore, for each combination of test statistic/wild bootstrap method we also compute a lower bound for the size of the bootstrap-based test conducted at nominal significance level ${\Greekmath 010B} =0.05$ (as well as at ${\Greekmath 010B} =0.1$).\footnote{ For the size computations we assume the errors to be normally distributed.} We find that for $95$ out of the remaining $134$ combinations ($11$ out of the remaining $24$ combinations, respectively) the (lower bound for the) size exceeds $3{\Greekmath 010B} $ for some of the design matrices and null hypotheses, sometimes by a considerable margin. Thus also these combinations do not lead to reliable bootstrap-based tests. This leaves us with $39$ ($13$, respectively) combinations. Exploiting that some of these combinations left are in fact equivalent to some of the above mentioned unreliable procedures (see Sections (ref) and (ref) for an explanation), allows us to even conclude that in the end only $16$ ($4$, respectively) bootstrap-based heteroskedasticity robust test procedures do not exhibit severe overrejection within the range of our numerical study when $n=10$.

Combining the just-described results with similar findings for the other sample sizes $n=20,30$ leads to the sobering conclusion that none of the bootstrap-based tests considered is reliable for all sample sizes and for ${\Greekmath 010B} =0.05$ as well as ${\Greekmath 010B} =0.1$. That is, for every combination of test statistic and bootstrap method considered, there is a sample size $ n\in \{10,20,30\}$, a significance level ${\Greekmath 010B} \in \{0.05,0.1\}$, a testing problem and a design matrix, such that the size of the corresponding bootstrap-based test equals $1$ by our theoretical results or is numerically found to exceed $3{\Greekmath 010B} $. We must hence conclude that none of the bootstrap-based tests considered is guaranteed to be immune to overrejection, and thus such tests are no reliable panacea for heteroskedasticity robust testing.

If one considers a fixed ${\Greekmath 010B} $, the situation is somewhat more encouraging. While there is no bootstrap-based test that is reliable for all sample sizes for the significance level ${\Greekmath 010B} =0.1$, for ${\Greekmath 010B} =0.05$ there are two bootstrap-based tests that are found not to break down in the above sense for any of the sample sizes considered in the numerical study. Both of these tests use a heteroskedasticity robust test statistic based on a HC3R covariance estimator, a wild bootstrap method based on the Mammen distribution, and impose the null restriction on the bootstrap data generating process. For more details see Section (ref). It is interesting to note that these findings call into question the recommendation in DavidsFlach2008 to base the wild bootstrap on the Rademacher distribution.

Of course, the above are worst-case results in spirit and do not preclude a given bootstrap-based test to be reasonably sized for certain instances of design matrix and null hypothesis. Therefore, in a given application, one could in principle imagine the following strategy: Numerically evaluate the size of the given bootstrap-based test (this will require to commit to a distributional assumption on the errors) and use the test only if the so-evaluated size does not exceed the nominal level ${\Greekmath 010B} $ (by much). \footnote{ Of course, this could also be pursued with non-bootstrap-based tests.} Otherwise, switch to another one of the many other bootstrap-based tests, repeat, and stop upon finding an acceptable test. As mentioned before, a partial shortcut for this strategy could be to compute ${\Greekmath 0123} $ first and to check if ${\Greekmath 010B} >{\Greekmath 0123} $, as we then know from our theoretical results that size must be equal to $1$. Of course, such a strategy would be computationally expensive and moreover would only be a stab into the dark, as there is no guarantee that one would end up with a bootstrap-based test that performs well in the sense of delivering size less than or equal to $ {\Greekmath 010B} $. It seems that a better and more direct strategy is to forgo the bootstrap idea and rather to construct size-controlling critical values for the original test statistics, e.g., for HC0-HC4 or HC0R-HC4R. This is pursued in the companion paper PP5HC. Certainly, this also leads to a computationally intensive method, but one that comes with guaranteed size control.

All test statistics mentioned so far are based on the ordinary least squares estimator. An alternative is to start from a feasible generalized least squares estimator, computed from a (potentially misspecified) model for heteroskedasticity. Again heteroskedasticity robust test procedures can then be developed in a similar manner, see, e.g., Cragg_1983, Cragg_1992, Flachaire_2005b, RomanoWolf2017, Lin_Chou_2018, DiCiccio_Romao_Wolf_2019. While results similar to the ones given in the present paper can probably also be developed for this alternative class of heteroskedasticity robust test procedures, we do not pursue this avenue here.

Framework

Consider the linear regression model

equation[equation omitted — 70 chars of source]

where $X$ is a (real) nonstochastic regressor (design) matrix of dimension $ n\times k$ and where ${\Greekmath 010C} \in \mathbb{R}^{k}$ denotes the unknown regression parameter vector. We always assume $\limfunc{rank}(X)=k$ and $ 1\leq k<n$. We furthermore assume that the $n\times 1$ disturbance vector $ \mathbf{U}=(\mathbf{u}_{1},\ldots ,\mathbf{u}_{n})^{\prime }$ has mean zero and unknown covariance matrix ${\Greekmath 011B} ^{2}\Sigma $, where $\Sigma $ varies in a user-specified (nonempty) set $\mathfrak{C}$ describing the allowed forms of heteroskedasticity, with $\mathfrak{C}$ satisfying $\mathfrak{C} \subseteq \mathfrak{C}_{Het}$, and where $0<{\Greekmath 011B} ^{2}<\infty $ holds ($ {\Greekmath 011B} $ always denoting the positive square root).\footnote{ Since we are concerned with finite-sample results only, the elements of $ \mathbf{Y}$, $X$, and $\mathbf{U}$ (and even the probability space supporting $\mathbf{Y}$ and $\mathbf{U}$) may depend on sample size $n$, but this will not be expressed in the notation. Furthermore, the obvious dependence of$\ \mathfrak{C}$ on $n$ will also not be shown in the notation.} The set $\mathfrak{C}$ will be referred to as the \textquotedblleft heteroskedasticity model\textquotedblright . Here

equation*[equation* omitted — 348 chars of source]

where $\limfunc{diag}({\Greekmath 011C} _{1}^{2},\ldots ,{\Greekmath 011C} _{n}^{2})$ denotes the $ n\times n$ matrix with diagonal elements given by ${\Greekmath 011C} _{i}^{2}$. That is, the errors in the regression model are uncorrelated but can be heteroskedastic. In particular, if $\mathfrak{C}$ is chosen to be $\mathfrak{ C}_{Het}$, one allows for heteroskedasticity of completely unknown form. The normalization condition $\sum_{i=1}^{n}{\Greekmath 011C} _{i}^{2}=1$ is included here only in order to guarantee identifiability of ${\Greekmath 011B} ^{2}$ and $\Sigma $, and could be replaced by any other normalization condition such as, e.g., $ \max {\Greekmath 011C} _{i}^{2}=1$, or ${\Greekmath 011C} _{1}^{2}=1$, without affecting the final results (because any of these normalizations leads to the same overall set of covariance matrices ${\Greekmath 011B} ^{2}\Sigma $ when ${\Greekmath 011B} ^{2}$ varies through the positive real line).

Although of no real significance for the results of this paper as explained in Section (ref), we shall, for ease of exposition, maintain in the sequel that the disturbance vector $\mathbf{U}$\ is normally distributed. The linear model described in ((ref)), together with the just made Gaussianity assumption on $\mathbf{U}$ and with the given heteroskedasticity model $\mathfrak{C}$, then induces a collection of distributions on the Borel-sets of $\mathbb{R}^{n}$, the sample space of $ \mathbf{Y}$. Denoting a Gaussian probability measure with mean ${\Greekmath 0116} \in \mathbb{R}^{n}$ and (possibly singular) covariance matrix $A$ by $P_{{\Greekmath 0116} ,A}$ , the induced collection of distributions is then given by

equation[equation omitted — 206 chars of source]

Since every $\Sigma \in \mathfrak{C}$ is positive definite by assumption, each element of the set in the previous display is absolutely continuous with respect to (w.r.t.) Lebesgue measure on $\mathbb{R}^{n}$.

We shall consider the problem of testing a linear (better: affine) hypothesis on the parameter vector ${\Greekmath 010C} \in \mathbb{R}^{k}$, i.e., the problem of testing the null $R{\Greekmath 010C} =r$ against the alternative $R{\Greekmath 010C} \neq r$, where $R$ is a $q\times k$ matrix always of rank $q\geq 1$ and $r\in \mathbb{R}^{q}$. Set $\mathfrak{M}=\limfunc{span}(X)$. Define the affine space

equation*[equation* omitted — 216 chars of source]

and let

equation*[equation* omitted — 222 chars of source]

Adopting these definitions, the above testing problem can then be written more precisely as

equation[equation omitted — 336 chars of source]

With $\mathfrak{M}_{0}^{lin}$ we shall denote the linear space parallel to $ \mathfrak{M}_{0}$, i.e., $\mathfrak{M}_{0}^{lin}=\mathfrak{M}_{0}-{\Greekmath 0116} _{0}=\left\{ X{\Greekmath 010C} :R{\Greekmath 010C} =0\right\} $ where ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$ . Of course, $\mathfrak{M}_{0}^{lin}$ does not depend on the choice of ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$.

As already mentioned, the assumption of Gaussianity is made for the sake of exposition only and does not really restrict the scope of the results in the paper as is discussed in Section (ref). The assumption of nonstochastic regressors entails little loss of generality either: For example, if $X$ is random and $\mathbf{U}$ is conditionally on $X$ distributed as $N(0,{\Greekmath 011B} ^{2}\Sigma )$, with ${\Greekmath 011B} ^{2}={\Greekmath 011B} ^{2}(X)$ and $\Sigma =\Sigma (X)\in \mathfrak{C}_{Het}$, the results of the paper can be applied after one conditions on $X$ (and a similar statement applies to the generalizations to non-Gaussianity discussed in Section (ref)). See Section (ref) for more discussion. For arguments supporting conditional inference see, e.g., RO1979. Note that such a "strict exogeneity" assumption is quite natural in the situation considered here.

We next collect some further terminology and notation used throughout the paper. A (nonrandomized) test is the indicator function of a Borel-set $W$ in $\mathbb{R}^{n}$, with $W$ called the corresponding rejection region. The size of such a test (rejection region) is -- as usual -- defined as the supremum over all rejection probabilities under the null hypothesis $H_{0}$ given in ((ref)), i.e.,

equation[equation omitted — 200 chars of source]

In slight abuse of terminology, we shall sometimes refer to this quantity as `the size of $W$ over $\mathfrak{C}$' when we want to emphasize the r\^{o}le of $\mathfrak{C}$. Throughout the paper we let $\hat{{\Greekmath 010C}}(y)=\left( X^{\prime }X\right) ^{-1}X^{\prime }y$, where $X$ is the design matrix appearing in ((ref)) and $y\in \mathbb{R}^{n}$. The corresponding ordinary least squares (OLS) residual vector is denoted by $\hat{u}(y)=y-X \hat{{\Greekmath 010C}}(y)$ and its elements are denoted by $\hat{u}_{t}(y)$. The elements of $X$ are denoted by $x_{ti}$, while $x_{t\cdot }$ and $x_{\cdot i} $ denote the $t$-th row and $i$-th column of $X$, respectively. For $ \mathcal{A}$ an affine subspace of $\mathbb{R}^{n}$ satisfying $\mathcal{A} \subseteq \limfunc{span}(X)$ let $\tilde{{\Greekmath 010C}}_{\mathcal{A}}(y)$ denote the restricted least-squares estimator, i.e., $X\tilde{{\Greekmath 010C}}_{\mathcal{A}}(y)$ solves

equation*[equation* omitted — 61 chars of source]

Lebesgue measure on the Borel-sets of $\mathbb{R}^{n}$ will be denoted by $ {\Greekmath 0115} _{\mathbb{R}^{n}}$. The set of real matrices of dimension $l\times m$ is denoted by $\mathbb{R}^{l\times m}$ (all matrices in the paper will be real matrices). The Euclidean norm is denoted by $\left\Vert \cdot \right\Vert $. Let $B^{\prime }$ denote the transpose of a matrix $B\in \mathbb{R}^{l\times m}$ and let $\mathrm{\limfunc{span}}(B)$ denote the subspace in $\mathbb{R}^{l}$ spanned by its columns. For a symmetric and nonnegative definite matrix $B$ we denote the unique symmetric and nonnegative definite square root by $B^{1/2}$. For a linear subspace $ \mathcal{L}$ of $\mathbb{R}^{n}$ we let $\mathcal{L}^{\bot }$ denote its orthogonal complement and we let $\Pi _{\mathcal{L}}$ denote the orthogonal projection onto $\mathcal{L}$. The $j$-th standard basis vector in $\mathbb{R }^{n}$ is written as $e_{j}(n)$. Furthermore, we let $\mathbb{N}$ denote the set of all positive integers. A sum (product, respectively) over an empty index set is to be interpreted as $0$ ($1$, respectively). For a subset $A$ of a topological space we denote by $\limfunc{int}(A)$ the interior of $A$ (w.r.t. the ambient space). Finally, for $\mathcal{A}$ an affine subspace of $\mathbb{R}^{n}$, let $G(\mathcal{A})$ denote the group of all affine transformations $y\mapsto {\Greekmath 010E} (y-a)+a^{\ast }$ where ${\Greekmath 010E} \in \mathbb{R }$, ${\Greekmath 010E} \neq 0$, and $a$ as well as $a^{\ast }$ are elements of $ \mathcal{A}$; for more information see Section 5.1 of PP2016.

Heteroskedasticity robust test statistics using unrestricted residuals

We next introduce two test statistics that will feature prominently. Variants of these statistics using restricted residuals are discussed in Section (ref). For a result pertaining to a more general class of test statistics see Theorem (ref) in Appendix (ref). The test statistic we shall consider first is a standard heteroskedasticity robust test statistic frequently considered in the literature and is given by

equation[equation omitted — 478 chars of source]

where $\hat{\Omega}_{Het}=R\hat{\Psi}_{Het}R^{\prime }$ and where $\hat{\Psi} _{Het}$ is a heteroskedasticity robust estimator as considered in E63,E67, which later on has found its way into the econometrics literature (e.g., W80). It is of the form

equation*[equation* omitted — 213 chars of source]

where the constants $d_{i}>0$ sometimes depend on the design matrix. Typical choices for $d_{i}$ suggested in the literature are $d_{i}=1$, $ d_{i}=n/(n-k) $, $d_{i}=\left( 1-h_{ii}\right) ^{-1}$, or $d_{i}=\left( 1-h_{ii}\right) ^{-2}$, where $h_{ii}$ denotes the $i$-th diagonal element of the projection matrix $X(X^{\prime }X)^{-1}X^{\prime }$, see LE2000 for an overview. Another suggestion is $d_{i}=\left( 1-h_{ii}\right) ^{-{\Greekmath 010E} _{i}}$ for ${\Greekmath 010E} _{i}=\min (nh_{ii}/k,4)$, see Crib2004. For the last three choices of $d_{i}$ just given, we use the convention that we set $d_{i}=1$ in case $h_{ii}=1$. Note that $h_{ii}=1$ implies $\hat{u} _{i}\left( y\right) =0$ for every $y$, and hence it is irrelevant which real value is assigned to $d_{i}$ in case $h_{ii}=1$.\footnote{ In fact, $h_{ii}=1$ is equivalent to $\hat{u}_{i}\left( y\right) =0$ for every $y$, each of which in turn is equivalent to $e_{i}(n)\in $ $\limfunc{ span}(X)$.} The five examples for the weights $d_{i}$ just given correspond to what is often called HC0-HC4 weights in the literature.

In conjunction with the test statistic $T_{Het}$, we shall consider the following mild assumption, which is Assumption 3 in PP2016. As discussed further below, this assumption is in a certain sense unavoidable when using $T_{Het}$. It furthermore also entails that our choice of assigning $T_{Het}\left( y\right) $ the value zero in case $\hat{\Omega} _{Het}\left( y\right) $ is singular has no import on the rejection probabilities of the (non-bootstrap-based) tests obtained from $T_{Het}$ (because of Lemma (ref)(c) below and absolute continuity of the measures $P_{{\Greekmath 0116} ,{\Greekmath 011B} ^{2}\Sigma }$). As will be seen later, our results for the corresponding bootstrap-based tests do also not depend on this choice.

assumptionLet $1\leq i_{1}<\ldots <i_{s}\leq n$ denote all the indices for which $e_{i_{j}}(n)\in \limfunc{span}(X)$ holds where $e_{j}(n)$ denotes the $j$-th standard basis vector in $\mathbb{R}^{n}$. If no such index exists, set $s=0$. Let $X^{\prime }\left( \lnot (i_{1},\ldots i_{s})\right) $ denote the matrix which is obtained from $X^{\prime }$ by deleting all columns with indices $i_{j}$, $1\leq i_{1}<\ldots <i_{s}\leq n$ (if $s=0$ no column is deleted). Then $\limfunc{rank}\left( R(X^{\prime }X)^{-1}X^{\prime }\left( \lnot (i_{1},\ldots i_{s})\right) \right) =q$ holds.

Observe that this assumption only depends on $X$ and $R$ and hence can be checked. Obviously, a simple sufficient condition for Assumption (ref) to hold is that $s=0$ (i.e., that $e_{j}(n)\notin \limfunc{span} (X) $ for all $j$), a generically satisfied condition. Furthermore, we introduce the matrix

eqnarray[eqnarray omitted — 327 chars of source]

The facts collected in the subsequent lemma will be used in the sequel. Parts (a)-(c) have been shown in Lemma 4.1 in PP2016, while Part (d) is taken from Lemma 5.18 of PP3. Part (e) is obvious (observe that $ B(y)$ depends only on $\hat{u}(y)$ and that $\hat{u}({\Greekmath 010D} (y-{\Greekmath 0116} )+{\Greekmath 0116} ^{\bullet })={\Greekmath 010D} \hat{u}(y)$ for every ${\Greekmath 010D} \in \mathbb{R}$, every $ {\Greekmath 0116} \in \limfunc{span}(X)$, and every ${\Greekmath 0116} ^{\bullet }\in \limfunc{span}(X)$ ).

lemma(a) $\hat{\Omega}_{Het}\left( y\right) $ is nonnegative definite for every $y\in \mathbb{R}^{n}$. (b) $\hat{\Omega}_{Het}\left( y\right) $ is singular (zero, respectively) if and only if $\limfunc{rank}\left( B(y)\right) <q$ ($B(y)=0$, respectively). (c) The set $\mathsf{B}$ given by $\left\{ y\in \mathbb{R}^{n}:\limfunc{rank} \left( B(y)\right) <q\right\} $ (or in view of (b) equivalently given by $ \{y\in \mathbb{R}^{n}:\det (\hat{\Omega}_{Het}\left( y\right) )=0\}$) is either a ${\Greekmath 0115} _{\mathbb{R}^{n}}$-null set or the entire sample space $ \mathbb{R}^{n}$. The latter occurs if and only if Assumption (ref) is violated (in which case the test based on $T_{Het}$ becomes trivial, as then $T_{Het}$ is identically zero). (d) Under Assumption (ref), the set $\mathsf{B}$ is a finite union of proper linear subspaces of $\mathbb{R}^{n}$; in case $q=1$, $\mathsf{B}$ is even a proper linear subspace itself.\footnote{ If Assumption (ref) is violated, $\mathsf{B}$ equals $\mathbb{R}^{n}$ by Part (c).} (e) $\mathsf{B}$ is a closed set and contains $\limfunc{span}(X)$. Furthermore, $\mathsf{B}$ is $G(\mathfrak{M})$-invariant and, in particular, $\mathsf{B}+\limfunc{span}(X)=\mathsf{B}$ holds.

In light of Part (c) of the lemma, we see that Assumption (ref) is a natural and unavoidable condition if one wants to obtain a sensible test from $T_{Het}$.\footnote{ If this assumption is violated then $T_{Het}$ is identically zero, an uninteresting trivial case.} Furthermore, note that, if $\mathsf{B}=\limfunc{ span}(X)$ is true, then Assumption (ref) must be satisfied (since $ \limfunc{span}(X)$ is a ${\Greekmath 0115} _{\mathbb{R}^{n}}$-null set due to the maintained assumption $k<n$). As shown in Lemma A.3 in PP3, for any given restriction matrix $R$, the relation $\mathsf{B}=\limfunc{span}(X)$ holds generically in various universes of design matrices. For later use we also mention that under Assumption (ref) the test statistic $T_{Het}$ is continuous at every $y\in \mathbb{R}^{n}\backslash \mathsf{B}$.\footnote{ If Assumption (ref) is violated, then $T_{Het}$ is constant equal to zero, and hence is trivially continuous everywhere.}

Next, we also consider the classical (i.e., uncorrected) F-test statistic, i.e.,

equation[equation omitted — 482 chars of source]

where $\hat{{\Greekmath 011B}}^{2}(y)=\hat{u}\left( y\right) ^{\prime }\hat{u}\left( y\right) /(n-k)\geq 0$ (which vanishes if and only if $y\in \limfunc{span} (X) $). Our choice to set $T_{uc}(y)=0$ for $y\in \limfunc{span}(X)$ has no import on the rejection probabilities of the (non-bootstrap-based) tests obtained from $T_{uc}$, since $\limfunc{span}(X)$ is a ${\Greekmath 0115} _{\mathbb{R} ^{n}}$-null set as a consequence of the maintained assumption that $k<n$ (and since the measures $P_{{\Greekmath 0116} ,{\Greekmath 011B} ^{2}\Sigma }$ are absolutely continuous). It will turn out also not to affect our results for bootstrap-based tests obtained from $T_{uc}$. For reasons of comparability with ((ref)) we have chosen not to normalize the numerator in ((ref)) by $q$, the number of restrictions to be tested, as is often done in the definition of the classical F-test statistic. This also has no import on the results as the bootstrap automatically adapts to scaling. For later use we also mention that the test statistic $T_{uc}$ is continuous at every $y\in \mathbb{R}^{n}\backslash \limfunc{span}(X)$.

remark(i) The test statistics $T_{Het}$ as well as $T_{uc}$ are $G( \mathfrak{M}_{0})$-invariant as is easily seen (with the respective exceptional sets $\mathsf{B}$ and $\limfunc{span}(X)$ being $G(\mathfrak{M})$ -invariant). (ii) Both statistics actually belong to the class of nonsphericity-corrected F-type test statistics in the sense of Section 5.4 in PP2016 (terminology being somewhat unfortunate in the case of $T_{uc}$ as no correction for the non-sphericity is applied in this case). See Remark (ref) in Appendix (ref) for more discussion.

Some intuition for the size one results

The mechanism leading to the size one results put forward formally in the next section is a concentration effect in the distribution generating the data $\mathbf{Y}=(y_{1},\ldots ,y_{n})^{\prime }$, entailing a similar effect in the distribution of $T(\mathbf{Y})$, where we denote by $T$ any of the test statistics considered in the paper. This concentration effect emerges when the data-generating process (DGP) is \textquotedblleft strongly heteroskedastic\textquotedblright . For simplicity, in this section we call a data-generating process strongly heteroskedastic, if a single observation has a (relatively) high variance, whereas all other observations have (relatively) low variance. We shall denote the index corresponding to the highly varying observation by $i^{\ast }$. Denote the expectation of the data vector $\mathbf{Y}$ by ${\Greekmath 0116} _{0}$, where we assume for the discussion in this section that ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$, i.e., that the null hypothesis is satisfied.

We now provide a nonrigorous explanation of the above-mentioned concentration effect and how it leads to the size one results:

enumerate• If the DGP is strongly heteroskedastic, only the single highly varying observation $y_{i^{\ast }}$ will substantially deviate from its expectation $ {\Greekmath 0116} _{0}^{(i^{\ast })}$, whereas all other observations will be very close to their expectations. That is, under such a DGP we approximately have \begin{equation*} \mathbf{Y}\approx {\Greekmath 0116} _{0}+(y_{i^{\ast }}-{\Greekmath 0116} _{0}^{(i^{\ast })})e_{i^{\ast }}(n), \end{equation*} where we recall that $e_{i^{\ast }}(n)$ is the $i^{\ast }$-th $n\times 1$ standard basis vector. That is, essentially, the data are concentrated on a one-dimensional affine subspace of the sample space $\mathbb{R}^{n}$. • Invariance properties of $T$ common to all test statistics used in this paper (and in practice) imply that \begin{equation*} T({\Greekmath 0116} _{0}+ce_{i^{\ast }}(n))=T({\Greekmath 0116} _{0}+e_{i^{\ast }}(n))\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ for every } c\neq 0. \end{equation*} That is, the test statistic under consideration is essentially constant on the one-dimensional affine subspace just obtained in the previous item. • Combining the two previous observations (and ignoring the case where $ y_{i^{\ast }}={\Greekmath 0116} _{0}^{(i^{\ast })}$), suggests that for strongly heteroskedastic DGPs we have \begin{equation*} T(\mathbf{Y})\approx T({\Greekmath 0116} _{0}+e_{i^{\ast }}(n)). \end{equation*} That is, essentially, the distribution of the test statistic collapses at the value $T({\Greekmath 0116} _{0}+e_{i^{\ast }}(n))$.

Now, recall that a wild bootstrap-based test rejects if the test statistic evaluated at the data $T(\mathbf{Y})$ exceeds the bootstrap critical value. This bootstrap critical value is a $1-{\Greekmath 010B} $ quantile of the distribution of the test statistic, but now induced by the distribution that corresponds to a bootstrap scheme $\mathbf{Y}^{\ast }$, say. In general, the distribution of $T(\mathbf{Y}^{\ast })$ depends on two sources of randomness: first, the DGP itself, and second, the randomization mechanism used to generate the bootstrap scheme $\mathbf{Y}^{\ast }$. Making use of the concentration mechanism outlined above, one can, for the class of bootstrap schemes considered, however, show that the dependence on the DGP essentially vanishes for strongly heteroskedastic DGPs. That is, the distribution function of $T(\mathbf{Y}^{\ast })$ approximately equals a distribution function $\digamma _{i^{\ast }}$, say, which depends on $ i^{\ast }$, but on no other aspect of the DGP. The bootstrap critical value is then a $1-{\Greekmath 010B} $ quantile of $\digamma _{i^{\ast }}$. Recalling from 3. that for strongly heteroskedastic DGPs $T(\mathbf{Y})\approx T({\Greekmath 0116} _{0}+e_{i^{\ast }}(n))$, it follows that for such DGPs the event that the wild bootstrap based test rejects the null hypothesis essentially coincides with the event that $T({\Greekmath 0116} _{0}+e_{i^{\ast }}(n))$ exceeds the $1-{\Greekmath 010B} $ quantile of $\digamma _{i^{\ast }}$. Both $T({\Greekmath 0116} _{0}+e_{i^{\ast }}(n))$ and $\digamma _{i^{\ast }}$ are non-random. Hence (recall that we are operating under the null hypothesis) for strongly heteroskedastic DGPs the test will have a rejection probability close to one if $\digamma _{i^{\ast }}(T({\Greekmath 0116} _{0}+e_{i^{\ast }}(n)))>1-{\Greekmath 010B} $.\footnote{ More precisely, if the left hand side limit $\digamma _{i^{\ast }}(T({\Greekmath 0116} _{0}+e_{i^{\ast }}(n))-)$ exceeds $1-{\Greekmath 010B} $. We ignore this technical detail here for the sake of simplicity.} In the above argument $i^{\ast }$ was fixed. Varying $i^{\ast }\in \{1,\ldots ,n\}$, we finally come to the conclusion that the maximal rejection probability under the null will be close to $1$ in case

equation*[equation* omitted — 132 chars of source]

In other words, the bootstrap-based test under consideration will have rejection probabilities close to one for all levels of significance satisfying

equation*[equation* omitted — 145 chars of source]

The quantity to the right is closely related to our constants ${\Greekmath 0123} $. Note that the above reasoning is nonrigorous and, in particular, does not take into consideration some technical subtleties that arise in the just given approximation arguments and that we have tacitly ignored in the preceding discussion. Therefore, the expressions for the constants $ {\Greekmath 0123} $ we arrive at in the theorems in the subsequent section are somewhat more complicated, albeit the underlying intuition is the same.

As transpires from the preceding heuristic discussion, the method for establishing the size one results given in the next section relies on the assumption that the heteroskedasticity model employed is rich enough to approximate extreme cases of strongly heteroskedastic DGPs, namely the ones where all but one observation have zero variance, arbitrarily well. This is certainly so for the leading case of the heteroskedasticity model $\mathfrak{ C}_{Het}$, which describes agnosticism about the form of heteroskedasticity. Therefore, the results in the next section are presented for this case, and a discussion to which other heteroskedasticity models these results generalize is given in Section (ref).

If one maintains a heteroskedasticity model that does not allow one to approximate any of the above mentioned extreme cases of strongly heteroskedastic DGPs (such as, e.g., the heteroskedasticity model $\mathfrak{ C}_{Het}(a)$ which consists of all error-covariance matrices in $\mathfrak{C} _{Het}$ with diagonal elements bounded from below by $a>0$), then the method of proof underlying our size one results no longer is applicable. However, this does not imply that the size of a bootstrap-based test over $ \mathfrak{C}_{Het}(a)$ is about right: Since the rejection probabilities are continuous in the parameters (in particular, in $\Sigma $), the size over $ \mathfrak{C}_{Het}(a)$ will be much larger than the nominal significance level at least for small $a$ (in fact, it will be close to one if $a$ is sufficiently small) in any situation where the size over $\mathfrak{C}_{Het}$ equals one (e.g., in the situations described in the theorems further below). [The actual size over $\mathfrak{C}_{Het}(a)$ depends on the chosen bound $a$ and on the design matrix, the hypothesis to be tested, the test statistic, and also on the bootstrap scheme used.] Furthermore, the bound $a$ has to be decided on prior to the data analysis and is part of modeling the form of heteroskedasticity. It is difficult to see how one would come up with a reasonable bound $a$ in practice: if $a$ is chosen to be very small, this may result in a heteroskedasticity model under which the tests are still severely oversized as just discussed, while choosing $a$ large will typically not be defendable as it presumes considerable knowledge about the admissible forms of heteroskedasticity.

Size one results

In this section we provide sufficient conditions for the size of bootstrap-based heteroskedasticity robust tests to be equal to one when the heteroskedasticity model is $\mathfrak{C}_{Het}$, which is the largest possible heteroskedasticity model and which reflects agnosticism regarding the form of heteroskedasticity. For extensions to other heteroskedasticity models see Section (ref). We next discuss the bootstrap schemes that will be considered and which all are based on the wild bootstrap idea. The first bootstrap scheme is given by

equation[equation omitted — 194 chars of source]

for every $y\in \mathbb{R}^{n}$, where $\mathcal{A}$ will always be an affine subspace of $\mathbb{R}^{n}$ satisfying $\mathfrak{M}_{0}\subseteq \mathcal{A}\subseteq \limfunc{span}(X)$, and where ${\Greekmath 0118} $ is a draw from $ \Xi $, a given (Borel) probability measure on $\mathbb{R}^{n}$. Typical choices in the literature are $\mathcal{A}=\mathfrak{M}_{0}$, i.e., one uses restricted residuals in the wild bootstrap, or $\mathcal{A}=\limfunc{span} (X) $, in which case unrestricted residuals are used.\footnote{ Because all test statistics (and associated exceptional sets) considered below are at least $G(\mathfrak{M}_{0})$-invariant (see Remarks (ref) and (ref)), the bootstrap scheme ((ref)) can be replaced by $y^{\ast \ast }(y,{\Greekmath 0118} )={\Greekmath 0116} _{0}+\limfunc{diag}({\Greekmath 0118} )(y-X\tilde{{\Greekmath 010C}}_{ \mathcal{A}}(y))$ for an arbitrary value ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$ without affecting the bootstrapped test statistic.} In practice only these two choices will typically arise, but the theory given below covers the more general case where $\mathfrak{M}_{0}\subseteq \mathcal{A}\subseteq \limfunc{ span}(X)$ at no extra cost. The measure $\Xi $ may depend on observable quantities like, e.g., $X$, $R$, or $\mathcal{A}$, but not on $y$. For example, $\Xi $ could be the $n$-fold product of Mammen or Rademacher distributions, but other choices (e.g., ones obtained by modifying the aforementioned distributions by weights, or non-discrete distributions) are also covered. See Section (ref) for some examples. For the theoretical results in this section there is no need to specify a particular form of $\Xi $.\footnote{ Suppose $\Xi $ is the empirical distribution of $B$ draws (possibly modified by weights) from an underlying distribution $\Xi _{0}$, which will often be the case if $n$ is large and the ideal bootstrap using $\Xi _{0}$ is infeasible. In this case $\Xi $ is strictly speaking a random probability measure (depending on the particular sample of size $B$ drawn from $\Xi _{0}$ ) and the bootstrap-based tests also depend on this sample. However, working conditionally on this sample, brings us back into the current framework.}

The second bootstrap scheme differs from the first one only insofar as centering is at the unrestricted estimator $X\hat{{\Greekmath 010C}}(y)$ rather than at the restricted estimator $X\tilde{{\Greekmath 010C}}_{\mathfrak{M}_{0}}(y)$. That is, the second bootstrap scheme is given by

equation[equation omitted — 178 chars of source]

for every $y\in \mathbb{R}^{n}$. Note that $y^{\ast }(y,{\Greekmath 0118} )$ as well as $ y^{\maltese }(y,{\Greekmath 0118} )$ depend also on the choice of $\mathcal{A}$, but we shall not show this dependence in the notation.

Bootstrap-based tests derived from $T_{Het}$ and $T_{uc}$

In the subsequent theorems $\Xi $ is always a (Borel) probability measure on $\mathbb{R}^{n}$, and $\mathcal{A}$ is an affine subspace of $\mathbb{R}^{n}$ satisfying $\mathfrak{M}_{0}\subseteq \mathcal{A}\subseteq \limfunc{span}(X)$ . If we use the first bootstrap scheme, i.e., ((ref)), the bootstrapped test statistic corresponding to $T_{Het}$ is given by $ T_{Het}^{\ast }$, where $T_{Het}^{\ast }:\mathbb{R}^{n}\times \mathbb{R} ^{n}\rightarrow \mathbb{R}$ is defined via

equation*[equation* omitted — 109 chars of source]

Furthermore, for every $y\in \mathbb{R}^{n}$ denote the distribution function of the bootstrapped test statistic under $\Xi $ by $F_{Het,y}$, i.e., $F_{Het,y}(t)=\Xi \mathbb{(}T_{Het}^{\ast }(y,{\Greekmath 0118} )\leq t)$ for $t\in \mathbb{R}$. For reasons that are discussed further below, we also need to consider a modification of $T_{Het}$ defined by $T_{Het}^{\blacktriangle }\left( y\right) =T_{Het}\left( y\right) $ if $y\notin \mathsf{B}$ and $ T_{Het}^{\blacktriangle }(y)=\infty $ otherwise. Its bootstrapped version is then given by $T_{Het}^{\blacktriangle ,\ast }(y,{\Greekmath 0118} )=T_{Het}^{\blacktriangle }\left( y^{\ast }(y,{\Greekmath 0118} )\right) $. Similarly as before, for every $y\in \mathbb{R}^{n}$ we denote its distribution function under $\Xi $ by $F_{Het,y}^{\blacktriangle }$, i.e., $F_{Het,y}^{ \blacktriangle }(t)=\Xi \mathbb{(}T_{Het}^{\blacktriangle ,\ast }(y,{\Greekmath 0118} )\leq t)$ for $t\in \mathbb{R\cup \{\infty \}}$.

theoremSuppose Assumption (ref) holds. (a) For every ${\Greekmath 010B} \in (0,1)$, let $f_{Het,1-{\Greekmath 010B} }(y)$ denote a $ (1-{\Greekmath 010B} )$-quantile of $F_{Het,y}$. Define ${\Greekmath 0123} _{Het}=1-\max ({\Greekmath 0123} _{1,Het},{\Greekmath 0123} _{2,Het})$, where \begin{equation} {\Greekmath 0123} _{1,Het}=\max_{\substack{ i=1,\ldots ,n, \\ e_{i}(n)\notin \mathsf{B}}}\Xi \left( \left\{ {\Greekmath 0118} :T_{Het}^{\ast }({\Greekmath 0116} _{0}+e_{i}(n),{\Greekmath 0118} )<T_{Het}({\Greekmath 0116} _{0}+e_{i}(n)),y^{\ast }({\Greekmath 0116} _{0}+e_{i}(n),{\Greekmath 0118} )\notin \mathsf{ B}\right\} \right) \end{equation} and \begin{equation} {\Greekmath 0123} _{2,Het}=\max_{\substack{ i=1,\ldots ,n, \\ e_{i}(n)\in \limfunc{ span}(X),R\hat{{\Greekmath 010C}}(e_{i}(n))\neq 0}}\Xi \left( \left\{ {\Greekmath 0118} :y^{\ast }({\Greekmath 0116} _{0}+e_{i}(n),{\Greekmath 0118} )\notin \mathsf{B}\right\} \right) \end{equation} for some ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$, with the convention that ${\Greekmath 0123} _{1,Het}=0$ (${\Greekmath 0123} _{2,Het}=0$, respectively) if the index set in the maximum operator in ((ref)) (((ref)), respectively) is empty. Then neither ${\Greekmath 0123} _{1,Het}$ nor ${\Greekmath 0123} _{2,Het}$ depend on the choice of ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$. Furthermore, for every ${\Greekmath 010B} \in (0,1)$ such that ${\Greekmath 010B} >{\Greekmath 0123} _{Het}$ holds, we have \begin{equation} \sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }\left( T_{Het}\geq f_{Het,1-{\Greekmath 010B} }\right) \geq \sup_{\Sigma \in \mathfrak{C} _{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }\left( T_{Het}>f_{Het,1-{\Greekmath 010B} }\right) =1 \end{equation} for every ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$ and every $0<{\Greekmath 011B} ^{2}<\infty $ (where the probabilities in ((ref)) are to be interpreted as inner probabilities\footnote{ This allows one to ignore measurability issues regarding $f_{Het,1-{\Greekmath 010B} }$. }). (b) For every ${\Greekmath 010B} \in (0,1)$, let $f_{Het,1-{\Greekmath 010B} }^{\blacktriangle }(y) $ denote a $(1-{\Greekmath 010B} )$-quantile of $F_{Het,y}^{\blacktriangle }$. Then, with ${\Greekmath 0123} _{Het}$ defined in Part (a), for every ${\Greekmath 010B} \in (0,1)$ such that ${\Greekmath 010B} >{\Greekmath 0123} _{Het}$ holds, we have \begin{equation} \sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }\left( T_{Het}\geq f_{Het,1-{\Greekmath 010B} }^{\blacktriangle }\right) \geq \sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }\left( T_{Het}>f_{Het,1-{\Greekmath 010B} }^{\blacktriangle }\right) =1 \end{equation} for every ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$ and every $0<{\Greekmath 011B} ^{2}<\infty $ (where the probabilities in ((ref)) are to be interpreted as inner probabilities).

Part (a) of the preceding theorem implies that for every nominal significance level ${\Greekmath 010B} >{\Greekmath 0123} _{Het}$ the size (over $\mathfrak{C} _{Het}$) of the bootstrap-based test derived from $T_{Het}$ is equal to $1$ and thus is inflated (and this is true whether the bootstrap-based test uses the rejection region $\left\{ y:T_{Het}(y)\geq f_{Het,1-{\Greekmath 010B} }(y)\right\} $ or $\left\{ y:T_{Het}(y)>f_{Het,1-{\Greekmath 010B} }(y)\right\} $).\footnote{ We note that in principle it is conceivable that these two rejection regions have different probabilities under $P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }$.} Note that the lower bound ${\Greekmath 0123} _{Het}$ is observable and can be computed, see Section (ref) for some more detail. As we shall see from the numerical results in Section (ref), the lower bound ${\Greekmath 0123} _{Het} $ can be quite small, the results in the theorem thus covering standard choices for ${\Greekmath 010B} $ such as ${\Greekmath 010B} =0.05$. A consequence of Theorem (ref) thus is, in particular, that there is in general no guarantee for bootstrap-based tests derived from $T_{Het}$ (or from the other statistics considered in the theorems further below), conducted at a nominal significance level ${\Greekmath 010B} $, to be truly level ${\Greekmath 010B} $ tests. Although trivial, we note that Theorem (ref) provides only a sufficient condition for size being equal to one and thus, in case $ {\Greekmath 010B} \leq {\Greekmath 0123} _{Het}$ holds, the size of the bootstrap-based test may nevertheless be much larger than ${\Greekmath 010B} $ (and may perhaps even be equal to $1$).

The significance of Part (b) of the theorem is as follows: Recall from Lemma (ref) that under Assumption (ref) the way $T_{Het}$ is defined on $\mathsf{B}$ is immaterial for the rejection probabilities of (non-bootstrap-based) tests obtained from this test statistic since the set $ \mathsf{B}$ is a Lebesgue null set and since the probability measures in ( (ref)) are all absolutely continuous w.r.t. Lebesgue measure; in particular, the (non-bootstrap-based) tests derived from $T_{Het}$ and $ T_{Het}^{\blacktriangle }$ have the same rejection probabilities. However, when it comes to the bootstrapped test statistics, the situation becomes more complicated as $\Xi $ often will be a discrete measure. That is, it is a priori conceivable that the value we assign to the test statistic on the set $\mathsf{B}$ may have an effect on the bootstrapped test statistic and thus on the $(1-{\Greekmath 010B} )$-quantile computed from it; in particular, it might be that an assignment of a value different from zero on the set $\mathsf{B}$ may lead to a larger $(1-{\Greekmath 010B} )$ -quantile. This then raises the question, whether a bootstrap-based test that uses such a (potentially) larger $(1-{\Greekmath 010B} )$-quantile may have a smaller size than when the quantile $f_{Het,1-{\Greekmath 010B} }$ is being used. Within the context of the theorem, Part (b) answers this in the negative by showing that, even if one defines the bootstrapped test statistic as $ \infty $ on the event where the bootstrap sample $y^{\ast }(y,{\Greekmath 0118} )$ falls into the exceptional set $\mathsf{B}$ and uses a resulting $(1-{\Greekmath 010B} )$ -quantile, the bootstrap-based test again has size $1$ under the same condition on ${\Greekmath 010B} $. [As any other way of defining the bootstrapped test statistic on the event $y^{\ast }(y,{\Greekmath 0118} )\in \mathsf{B}$ obviously leads to $ (1-{\Greekmath 010B} )$-quantiles not larger than an (appropriately chosen) $(1-{\Greekmath 010B} ) $-quantile of $F_{Het,y}^{\blacktriangle }$, Part (b) covers also any such alternative definition of the bootstrapped test statistic.]\footnote{ An alternative approach, which -- if successful -- would make considering Part (b) obsolete, would be to try to show that the set of $y^{\prime }s$ for which $T_{Het}^{\blacktriangle ,\ast }$ and $T_{Het}^{\ast }$ coincide $ \Xi $-a.e., and thus their quantiles coincide, is the complement of a Lebesgue null set. While this alternative approach actually can be shown to work in some special cases, it does not so in general, as can be seen from examples.} For additional discussion see also Remark (ref).

We also stress that the results in the preceding theorem hold for any choice $f_{Het,1-{\Greekmath 010B} }$ ($f_{Het,1-{\Greekmath 010B} }^{\blacktriangle }(y)$, respectively) from the set of $(1-{\Greekmath 010B} )$-quantiles of $F_{Het,y}$ ($ F_{Het,y}^{\blacktriangle }$, respectively).\footnote{Suppose $ 0<{\Greekmath 010E} <1$ and $F$ is a cdf defined on $\mathbb{R}$ ($\mathbb{R\cup \{\infty \}}$, respectively). An element $q\in \mathbb{R}$ ($q\in \mathbb{ R\cup \{\infty \}}$, respectively) is said to be a ${\Greekmath 010E} $-quantile of $F$ iff it satisfies $F(q)\geq {\Greekmath 010E} \geq F(q-)$, where $F(q-)$ denotes the left-hand limit of $F$ at $q$. Note that $q$ need not be unique in general. There is always a smallest and a largest ${\Greekmath 010E} $-quantile among all $ {\Greekmath 010E} $-quantiles. The smallest one is given by $F^{-1}({\Greekmath 010E} $), where $ F^{-1}$ is the "generalized" inverse of $F$. If ${\Greekmath 010E} $ does not belong to the range of $F$, then $F^{-1}({\Greekmath 010E} $) is also the largest ${\Greekmath 010E} $ -quantile. Otherwise, the largest ${\Greekmath 010E} $-quantile is given by $\sup \left\{ x\in \mathbb{R}:F(x)={\Greekmath 010E} \right\} $ ($\sup \left\{ x\in \mathbb{ R\cup \{\infty \}}:F(x)={\Greekmath 010E} \right\} $, respectively), which may or may not coincide with $F^{-1}({\Greekmath 010E} )$.}

Furthermore, we note that the preceding theorem holds with the same lower bound ${\Greekmath 0123} _{Het}$ for a much larger class of error distributions than just Gaussian errors (an assumption we have made only for convenience), see Section (ref). Hence, in this sense the lower bound $ {\Greekmath 0123} _{Het}$ is \textquotedblleft distribution free\textquotedblright .

We next turn to the test statistic $T_{uc}$. Again using the first bootstrap scheme, the bootstrapped test statistic is then given by $T_{uc}^{\ast }$ where $T_{uc}^{\ast }:\mathbb{R}^{n}\times \mathbb{R}^{n}\rightarrow \mathbb{ R}$ is defined via

equation*[equation* omitted — 107 chars of source]

Furthermore, for every $y\in \mathbb{R}^{n}$ denote the distribution function of the bootstrapped test statistic under $\Xi $ by $F_{uc,y}$, i.e., $F_{uc,y}(t)=\Xi \mathbb{(}T_{uc}^{\ast }(y,{\Greekmath 0118} )\leq t)$ for $t\in \mathbb{R}$. As before, we also need to consider the modification of $T_{uc}$ defined by $T_{uc}^{\blacktriangle }\left( y\right) =T_{uc}\left( y\right) $ if $y\notin \limfunc{span}(X)$ and $T_{uc}^{\blacktriangle }(y)=\infty $ otherwise. Its bootstrapped version is then given by $T_{uc}^{\blacktriangle ,\ast }(y,{\Greekmath 0118} )=T_{uc}^{\blacktriangle }\left( y^{\ast }(y,{\Greekmath 0118} )\right) $. Similarly as before, for every $y\in \mathbb{R}^{n}$ we denote its distribution function under $\Xi $ by $F_{uc,y}^{\blacktriangle }$, i.e., $ F_{uc,y}^{\blacktriangle }(t)=\Xi \mathbb{(}T_{uc}^{\blacktriangle ,\ast }(y,{\Greekmath 0118} )\leq t)$ for $t\in \mathbb{R\cup \{\infty \}}$.

theorem(a) For every ${\Greekmath 010B} \in (0,1)$, let $f_{uc,1-{\Greekmath 010B} }(y)$ denote a $(1-{\Greekmath 010B} )$-quantile of $F_{uc,y}$. Define ${\Greekmath 0123} _{uc}=1-\max ({\Greekmath 0123} _{1,uc},{\Greekmath 0123} _{2,uc})$, where \begin{equation} {\Greekmath 0123} _{1,uc}=\max_{\substack{ i=1,\ldots ,n, \\ e_{i}(n)\notin \limfunc{span}(X)}}\Xi \left( \left\{ {\Greekmath 0118} :T_{uc}^{\ast }({\Greekmath 0116} _{0}+e_{i}(n),{\Greekmath 0118} )<T_{uc}({\Greekmath 0116} _{0}+e_{i}(n)),y^{\ast }({\Greekmath 0116} _{0}+e_{i}(n),{\Greekmath 0118} )\notin \limfunc{span}(X)\right\} \right) \end{equation} and \begin{equation} {\Greekmath 0123} _{2,uc}=\max_{\substack{ i=1,\ldots ,n, \\ e_{i}(n)\in \limfunc{ span}(X),R\hat{{\Greekmath 010C}}(e_{i}(n))\neq 0}}\Xi \left( \left\{ {\Greekmath 0118} :y^{\ast }({\Greekmath 0116} _{0}+e_{i}(n),{\Greekmath 0118} )\notin \limfunc{span}(X)\right\} \right) \end{equation} for some ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$, with the convention that ${\Greekmath 0123} _{2,uc}=0$ if the index set in the maximum operator in ((ref)) is empty.\footnote{ Note that the index set in the maximum operator in ((ref)) can not be empty since we have assumed $k<n$.} Then neither ${\Greekmath 0123} _{1,uc}$ nor ${\Greekmath 0123} _{2,uc}$ depend on the choice of ${\Greekmath 0116} _{0}\in \mathfrak{M} _{0} $. Furthermore, for every ${\Greekmath 010B} \in (0,1)$ such that ${\Greekmath 010B} >{\Greekmath 0123} _{uc}$ holds, we have \begin{equation} \sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }\left( T_{uc}\geq f_{uc,1-{\Greekmath 010B} }\right) \geq \sup_{\Sigma \in \mathfrak{C} _{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }\left( T_{uc}>f_{uc,1-{\Greekmath 010B} }\right) =1 \end{equation} for every ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$ and every $0<{\Greekmath 011B} ^{2}<\infty $ (where the probabilities in ((ref)) are to be interpreted as inner probabilities). (b) For every ${\Greekmath 010B} \in (0,1)$, let $f_{uc,1-{\Greekmath 010B} }^{\blacktriangle }(y)$ denote a $(1-{\Greekmath 010B} )$-quantile of $F_{uc,y}^{\blacktriangle }$. Then, with $ {\Greekmath 0123} _{uc}$ defined in Part (a), for every ${\Greekmath 010B} \in (0,1)$ such that ${\Greekmath 010B} >{\Greekmath 0123} _{uc}$ holds, we have \begin{equation} \sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }\left( T_{uc}\geq f_{uc,1-{\Greekmath 010B} }^{\blacktriangle }\right) \geq \sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }\left( T_{uc}>f_{uc,1-{\Greekmath 010B} }^{\blacktriangle }\right) =1 \end{equation} for every ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$ and every $0<{\Greekmath 011B} ^{2}<\infty $ (where the probabilities in ((ref)) are to be interpreted as inner probabilities).

Mutatis mutandis, a discussion similar to the one given subsequently to Theorem (ref) also applies here.

So far we have only considered the bootstrap scheme ((ref)). We now turn to the second bootstrap scheme given by ((ref)). Here the bootstrapped version of $T_{Het}$ is given by

equation*[equation* omitted — 376 chars of source]

if $y^{\maltese }(y,{\Greekmath 0118} )\notin \mathsf{B}$, and by $T_{Het}^{\maltese }(y,{\Greekmath 0118} )=0$ if $y^{\maltese }(y,{\Greekmath 0118} )\in \mathsf{B}$. And the bootstrapped version of $T_{uc}$ is given by

equation*[equation* omitted — 431 chars of source]

if $y^{\maltese }(y,{\Greekmath 0118} )\not\in \limfunc{span}(X)$, and by $ T_{uc}^{\maltese }(y,{\Greekmath 0118} )=0$ if $y^{\maltese }(y,{\Greekmath 0118} )\in \limfunc{span}(X)$ . Furthermore, $T_{Het}^{\blacktriangle ,\maltese }(y,{\Greekmath 0118} )$ and $ T_{uc}^{\blacktriangle ,\maltese }(y,{\Greekmath 0118} )$, are defined in exactly the same way, except that $T_{Het}^{\blacktriangle ,\maltese }(y,{\Greekmath 0118} )=\infty $ if $ y^{\maltese }(y,{\Greekmath 0118} )\in \mathsf{B}$ and that $T_{uc}^{\blacktriangle ,\maltese }(y,{\Greekmath 0118} )=\infty $ if $y^{\maltese }(y,{\Greekmath 0118} )\in \limfunc{span}(X)$.

We will show in the next lemma that $T_{Het}^{\maltese }(y,{\Greekmath 0118} )$ coincides with $T_{Het}^{\ast }(y,{\Greekmath 0118} )$, and that the same is true for $ T_{uc}^{\maltese }(y,{\Greekmath 0118} )$ and $T_{uc}^{\ast }(y,{\Greekmath 0118} )$ (as well as for $ T_{Het}^{\blacktriangle ,\maltese }(y,{\Greekmath 0118} )$ and $T_{Het}^{\blacktriangle ,\ast }(y,{\Greekmath 0118} )$, and $T_{uc}^{\blacktriangle ,\maltese }(y,{\Greekmath 0118} )$ and $ T_{uc}^{\blacktriangle ,\ast }(y,{\Greekmath 0118} )$), provided the same affine space $ \mathcal{A}$ is used in ((ref)) and ((ref)). As a consequence, this -- together with Remark (ref) -- shows that Theorems (ref) and (ref) also apply immediately to the bootstrap-based test when the second bootstrap scheme, i.e., ((ref)), is used (with the same $\mathcal{A}$ and $\Xi $). The lemma is certainly not new and is a variant of a similar result given as Proposition 1 in vG_K_2002.

lemmaWe have $T_{Het}^{\ast }(y,{\Greekmath 0118} )=T_{Het}^{\maltese }(y,{\Greekmath 0118} )$, $ T_{uc}^{\ast }(y,{\Greekmath 0118} )=T_{uc}^{\maltese }(y,{\Greekmath 0118} )$, $T_{Het}^{\blacktriangle ,\ast }(y,{\Greekmath 0118} )=T_{Het}^{\blacktriangle ,\maltese }(y,{\Greekmath 0118} )$, and $ T_{uc}^{\blacktriangle ,\ast }(y,{\Greekmath 0118} )=T_{uc}^{\blacktriangle ,\maltese }(y,{\Greekmath 0118} )$ for every $y\in \mathbb{R}^{n}$ and every ${\Greekmath 0118} \in \mathbb{R}^{n}$ . [Here it is understood that both bootstrap schemes are based on the same affine space $\mathcal{A}$.]
remarkDefine ${\Greekmath 0112} _{Het}$ exactly in the same way as ${\Greekmath 0123} _{Het}$, except that $T_{Het}^{\ast }$ and $y^{\ast }$ are replaced by $ T_{Het}^{\maltese }$ and $y^{\maltese }$. Similarly define ${\Greekmath 0112} _{uc}$. Then ${\Greekmath 0112} _{Het}={\Greekmath 0123} _{Het}$ and ${\Greekmath 0112} _{uc}={\Greekmath 0123} _{uc}$ in view of Lemma (ref) and the fact that $y^{\ast }(y,{\Greekmath 0118} )\notin \mathsf{B} $ iff $y^{\maltese }(y,{\Greekmath 0118} )\notin \mathsf{B}$ and $y^{\ast }(y,{\Greekmath 0118} )\notin \limfunc{span}(X)$ iff $y^{\maltese }(y,{\Greekmath 0118} )\notin \limfunc{span}(X)$ (note that $y^{\ast }(y,{\Greekmath 0118} )-y^{\maltese }(y,{\Greekmath 0118} )\in \limfunc{span}(X)$ and that $\mathsf{B}+\limfunc{span}(X)=\mathsf{B}$).

Bootstrap-based tests derived from $\tilde{T}_{Het}$ and $\tilde{ T}_{uc}$

Based on suggestions in the literature on bootstrapping heteroskedasticity robust tests, we next consider two further test statistics, which are versions of $T_{Het}$ and $T_{uc}$ with the only difference that the covariance matrix estimators used are computed from restricted -- instead of unrestricted -- residuals. We thus define

equation[equation omitted — 498 chars of source]

where $\tilde{\Omega}_{Het}=R\tilde{\Psi}_{Het}R^{\prime }$ and where $ \tilde{\Psi}_{Het}$ is given by

equation*[equation* omitted — 235 chars of source]

where the constants $\tilde{d}_{i}>0$ sometimes depend on the design matrix and on the restriction matrix $R$. Here $\tilde{u}\left( y\right) =y-X\tilde{ {\Greekmath 010C}}_{\mathfrak{M}_{0}}(y)=\Pi _{(\mathfrak{M}_{0}^{lin})^{\bot }}(y-{\Greekmath 0116} _{0})$, where the last expression does not depend on the choice of ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$, and where $\tilde{u}_{t}\left( y\right) $ denotes the $t$-th component of $\tilde{u}\left( y\right) $. Typical choices for $ \tilde{d}_{i}$ are $\tilde{d}_{i}=1$, $\tilde{d}_{i}=n/(n-(k-q))$, $\tilde{d} _{i}=(1-\tilde{h}_{ii})^{-1}$, or $\tilde{d}_{i}=(1-\tilde{h}_{ii})^{-2}$ where $\tilde{h}_{ii}$ denotes the $i$-th diagonal element of the projection matrix $\Pi _{\mathfrak{M}_{0}^{lin}}$, see, e.g., DavidsonMacKinnon1985. Another suggestion is $\tilde{d}_{i}=(1-\tilde{h} _{ii})^{-\tilde{{\Greekmath 010E}}_{i}}$ for $\tilde{{\Greekmath 010E}}_{i}=\min (n\tilde{h} _{ii}/(k-q),4)$ with the convention that $\tilde{{\Greekmath 010E}}_{i}=0$ if $k=q$. \footnote{ Note that in case $k=q$ we have $\tilde{h}_{ii}=0$, and hence $\tilde{d} _{i}=1$ regardless of our convention for $\tilde{{\Greekmath 010E}}_{i}$.} For the last three choices of $\tilde{d}_{i}$ just given we use the convention that we set $\tilde{d}_{i}=1$ in case $\tilde{h}_{ii}=1$. Note that $\tilde{h} _{ii}=1 $ implies $\tilde{u}_{i}\left( y\right) =0$ for every $y$, and hence it is irrelevant which real value is assigned to $\tilde{d}_{i}$ in case $ \tilde{h}_{ii}=1$.\footnote{ In fact, $\tilde{h}_{ii}=1$ is equivalent to $\tilde{u}_{i}\left( y\right) =0 $ for every $y$, each of which in turn is equivalent to $e_{i}(n)\in $ $ \mathfrak{M}_{0}^{lin}$.} The five examples for the weights $\tilde{d}_{i}$ just given correspond to what is often called HC0R-HC4R weights in the literature.\footnote{ In the case $k=q$ the HC0R-HC4R weights all coincide ($\tilde{d}_{i}=1$ for every $i$), and hence result in the same test statistic.}

The subsequent assumption ensures that the set of $y$'s for which $\tilde{ \Omega}_{Het}\left( y\right) $ is singular is a Lebesgue null set, implying that our choice of assigning $\tilde{T}_{Het}\left( y\right) $ the value zero in case $\tilde{\Omega}_{Het}\left( y\right) $ is singular has no import on the rejection probabilities of the (non-bootstrap-based) tests obtained from $\tilde{T}_{Het}$ (as the measures $P_{{\Greekmath 0116} ,{\Greekmath 011B} ^{2}\Sigma } $ are absolutely continuous). As will be seen later, our results for the corresponding bootstrap-based tests do also not depend on this choice. Also, as discussed further below, the assumption is in a certain sense unavoidable when using $\tilde{T}_{Het}$.

assumptionLet $1\leq i_{1}<\ldots <i_{s}\leq n$ denote all the indices for which $e_{i_{j}}(n)\in \mathfrak{M}_{0}^{lin}$ holds where $ e_{j}(n)$ denotes the $j$-th standard basis vector in $\mathbb{R}^{n}$. If no such index exists, set $s=0$. Let $X^{\prime }\left( \lnot (i_{1},\ldots i_{s})\right) $ denote the matrix which is obtained from $X^{\prime }$ by deleting all columns with indices $i_{j}$, $1\leq i_{1}<\ldots <i_{s}\leq n$ (if $s=0$ no column is deleted). Then $\limfunc{rank}\left( R(X^{\prime }X)^{-1}X^{\prime }\left( \lnot (i_{1},\ldots i_{s})\right) \right) =q$ holds.

Observe that this assumption only depends on $X$ and $R$ and hence can be checked. Obviously, a simple sufficient condition for Assumption (ref) to hold is that $s=0$ (i.e., that $e_{j}(n)\notin \mathfrak{M }_{0}^{lin}$ for all $j$), a generically satisfied condition. Furthermore, we introduce the matrix

eqnarray[eqnarray omitted — 407 chars of source]

Note that this matrix does not depend on the choice of ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$. The following lemma collects some important properties of $\tilde{\Omega}_{Het}$ and $\mathsf{\tilde{B}}$ (defined in that lemma). Its proof is given in Appendix (ref).

lemma(a) $\tilde{\Omega}_{Het}\left( y\right) $ is nonnegative definite for every $y\in \mathbb{R}^{n}$. (b) $\tilde{\Omega}_{Het}\left( y\right) $ is singular (zero, respectively) if and only if $\limfunc{rank}(\tilde{B}(y))<q$ ($\tilde{B}(y)=0$, respectively). (c) The set $\mathsf{\tilde{B}}$ given by $\{y\in \mathbb{R}^{n}:\limfunc{ rank}(\tilde{B}(y))<q\}$ (or, in view of (b), equivalently given by $\{y\in \mathbb{R}^{n}:\det (\tilde{\Omega}_{Het}\left( y\right) )=0\}$) is either a ${\Greekmath 0115} _{\mathbb{R}^{n}}$-null set or the entire sample space $\mathbb{R} ^{n}$. The latter occurs if and only if Assumption (ref) is violated (in which case the test based on $\tilde{T}_{Het}$ becomes trivial, as then $\tilde{T}_{Het}$ is identically zero). (d) Suppose Assumption (ref) holds. Then for every ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$ the set $\mathsf{\tilde{B}}-{\Greekmath 0116} _{0}$ is a finite union of proper linear subspaces; in case $q=1$, $\mathsf{\tilde{B}}-{\Greekmath 0116} _{0} $ is even a proper linear subspace itself. \footnote{ Consequently, $\mathsf{\tilde{B}}$ is a finite union of proper affine subspaces, and is a proper affine subspace itself in case $q=1$.}$^{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{,} } $\footnote{ If Assumption (ref) is violated, then $\mathsf{\tilde{B}}-{\Greekmath 0116} _{0}=\mathsf{\tilde{B}}=\mathbb{R}^{n}$ in view of Part (c).} [Note that $ \mathsf{\tilde{B}}-{\Greekmath 0116} _{0}$ does not depend on the choice of ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$. In particular, if $r=0$, i.e., if $\mathfrak{M}_{0}$ is linear, we thus may set ${\Greekmath 0116} _{0}=0$.] (e) $\mathsf{\tilde{B}}$ is a closed set and contains $\mathfrak{M}_{0}$. Also $\mathsf{\tilde{B}}$ is $G(\mathfrak{M}_{0})$-invariant, and in particular $\mathsf{\tilde{B}}+\mathfrak{M}_{0}^{lin}=\mathsf{\tilde{B}}$.

In light of Part (c) of the lemma, we see that Assumption (ref) is a natural and unavoidable condition if one wants to obtain a sensible test from $\tilde{T}_{Het}$.\footnote{ If this assumption is violated then $\tilde{T}_{Het}$ is identically zero, an uninteresting trivial case.} Furthermore, note that if $\mathsf{\tilde{B}} =\mathfrak{M}_{0}$ is true, then Assumption (ref) must be satisfied (since $\mathfrak{M}_{0}$ is a ${\Greekmath 0115} _{\mathbb{R}^{n}}$-null set as $k-q<n$ is always the case). For later use we also mention that under Assumption (ref) the statistic $\tilde{T}_{Het}$ is continuous at every $y\in \mathbb{R}^{n}\backslash \mathsf{\tilde{B}}$.\footnote{ If Assumption (ref) is violated, then $\tilde{T}_{Het}$ is constant equal to zero, and hence trivially continuous everywhere.}

We finally consider in analogy with $T_{uc}$

equation[equation omitted — 498 chars of source]

where $\tilde{{\Greekmath 011B}}^{2}(y)=\tilde{u}\left( y\right) ^{\prime }\tilde{u} \left( y\right) /(n-(k-q))\geq 0$ (which vanishes if and only if $y\in \mathfrak{M}_{0}$). Of course, our choice to set $\tilde{T}_{uc}(y)=0$ for $ y\in \mathfrak{M}_{0}$ has no import on the rejection probabilities of the (non-bootstrap-based) tests obtained from $\tilde{T}_{uc}$, since $\mathfrak{ M}_{0}$ is a ${\Greekmath 0115} _{\mathbb{R}^{n}}$-null set (and since the measures $ P_{{\Greekmath 0116} ,{\Greekmath 011B} ^{2}\Sigma }$ are absolutely continuous). It will turn out also not to affect our results for bootstrap-based tests obtained from $ \tilde{T}_{uc}$. For later use we also mention that $\tilde{T}_{uc}$ is continuous at every $y\in \mathbb{R}^{n}\backslash \mathfrak{M}_{0}$.

remarkThe test statistics $\tilde{T}_{Het}$ as well as $\tilde{ T}_{uc}$ are $G(\mathfrak{M}_{0})$-invariant as is easily seen (with the respective exceptional sets $\mathsf{\tilde{B}}$ and $\mathfrak{M}_{0}$ also being $G(\mathfrak{M}_{0})$-invariant), but typically they are not nonsphericity-corrected F-type tests in the sense of Section 5.4 in PP2016.

In the theorems given in the next two subsections we use the same bootstrap schemes as before (i.e., ((ref)) and ((ref))); in particular, recall that $\Xi $ is a (Borel) probability measure on $\mathbb{R}^{n}$, and that $\mathcal{A}$ is an affine subspace of $\mathbb{R}^{n}$ satisfying $ \mathfrak{M}_{0}\subseteq \mathcal{A}\subseteq \limfunc{span}(X)$.

The first bootstrap scheme

We start with results where the first bootstrap scheme,\ i.e., ((ref) ), is being used. The bootstrapped test statistic corresponding to $\tilde{T} _{Het}$ is then given by $\tilde{T}_{Het}^{\ast }$, where $\tilde{T} _{Het}^{\ast }:\mathbb{R}^{n}\times \mathbb{R}^{n}\rightarrow \mathbb{R}$ is defined via

equation*[equation* omitted — 125 chars of source]

Furthermore, for every $y\in \mathbb{R}^{n}$ denote the distribution function of the bootstrapped test statistic under $\Xi $ by $\tilde{F} _{Het,y}$, i.e., $\tilde{F}_{Het,y}(t)=\Xi \mathbb{(}\tilde{T}_{Het}^{\ast }(y,{\Greekmath 0118} )\leq t)$ for $t\in \mathbb{R}$. For similar reasons as in Section (ref), we also consider the modification of $\tilde{T}_{Het}$ defined by $\tilde{T}_{Het}^{\blacktriangle }\left( y\right) =\tilde{T} _{Het}\left( y\right) $ if $y\notin \mathsf{\tilde{B}}$ and $\tilde{T} _{Het}^{\blacktriangle }(y)=\infty $ otherwise. Its bootstrapped version is then given by $\tilde{T}_{Het}^{\blacktriangle ,\ast }(y,{\Greekmath 0118} )=\tilde{T} _{Het}^{\blacktriangle }\left( y^{\ast }(y,{\Greekmath 0118} )\right) $. For every $y\in \mathbb{R}^{n}$ we denote its distribution function under $\Xi $ by $\tilde{F }_{Het,y}^{\blacktriangle }$, i.e., $\tilde{F}_{Het,y}^{\blacktriangle }(t)=\Xi \mathbb{(}\tilde{T}_{Het}^{\blacktriangle ,\ast }(y,{\Greekmath 0118} )\leq t)$ for $t\in \mathbb{R\cup \{\infty \}}$.

theoremSuppose Assumption (ref) holds. (a) For every ${\Greekmath 010B} \in (0,1)$, let $\tilde{f}_{Het,1-{\Greekmath 010B} }(y)$ denote a $(1-{\Greekmath 010B} )$-quantile of $\tilde{F}_{Het,y}$. Define \begin{equation} \tilde{{\Greekmath 0123}}_{Het}=1-\max_{\substack{ i=1,\ldots ,n, \\ {\Greekmath 0116} _{0}+e_{i}(n)\notin \mathsf{\tilde{B}}}}\Xi \left( \left\{ {\Greekmath 0118} :\tilde{T} _{Het}^{\ast }({\Greekmath 0116} _{0}+e_{i}(n),{\Greekmath 0118} )<\tilde{T}_{Het}({\Greekmath 0116} _{0}+e_{i}(n)),y^{\ast }({\Greekmath 0116} _{0}+e_{i}(n),{\Greekmath 0118} )\notin \mathsf{\tilde{B}} \right\} \right) \end{equation} for some ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$, with the convention that $\tilde{ {\Greekmath 0123}}_{Het}=1$ if the index set in the maximum operator in ((ref)) is empty. Then $\tilde{{\Greekmath 0123}}_{Het}$ does not depend on the choice of ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$. Furthermore, for every $ {\Greekmath 010B} \in (0,1)$ such that ${\Greekmath 010B} >\tilde{{\Greekmath 0123}}_{Het}$ holds, we have \begin{equation} \sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }\left( \tilde{T}_{Het}\geq \tilde{f}_{Het,1-{\Greekmath 010B} }\right) \geq \sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }\left( \tilde{T}_{Het}> \tilde{f}_{Het,1-{\Greekmath 010B} }\right) =1 \end{equation} for every ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$ and every $0<{\Greekmath 011B} ^{2}<\infty $ (where the probabilities in ((ref)) are to be interpreted as inner probabilities). (b) For every ${\Greekmath 010B} \in (0,1)$, let $\tilde{f}_{Het,1-{\Greekmath 010B} }^{\blacktriangle }(y)$ denote a $(1-{\Greekmath 010B} )$-quantile of $\tilde{F} _{Het,y}^{\blacktriangle }$. Then, with $\tilde{{\Greekmath 0123}}_{Het}$ defined in Part (a), for every ${\Greekmath 010B} \in (0,1)$ such that ${\Greekmath 010B} >\tilde{{\Greekmath 0123}} _{Het}$ holds, we have \begin{equation} \sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }\left( \tilde{T}_{Het}\geq \tilde{f}_{Het,1-{\Greekmath 010B} }^{\blacktriangle }\right) \geq \sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }\left( \tilde{T}_{Het}>\tilde{f}_{Het,1-{\Greekmath 010B} }^{\blacktriangle }\right) =1 \end{equation} for every ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$ and every $0<{\Greekmath 011B} ^{2}<\infty $ (where the probabilities in ((ref)) are to be interpreted as inner probabilities).

We next turn to the test statistic $\tilde{T}_{uc}$. Again using the first bootstrap scheme, the bootstrapped test statistic is then given by $\tilde{T} _{uc}^{\ast }$, where $\tilde{T}_{uc}^{\ast }:\mathbb{R}^{n}\times \mathbb{R} ^{n}\rightarrow \mathbb{R}$ is defined via

equation*[equation* omitted — 123 chars of source]

Furthermore, for every $y\in \mathbb{R}^{n}$ denote the distribution function of the bootstrapped test statistic under $\Xi $ by $\tilde{F} _{uc,y} $, i.e., $\tilde{F}_{uc,y}(t)=\Xi \mathbb{(}\tilde{T}_{uc}^{\ast }(y,{\Greekmath 0118} )\leq t)$ for $t\in \mathbb{R}$. We also consider the modification of $\tilde{T}_{uc}$ defined by $\tilde{T}_{uc}^{\blacktriangle }\left( y\right) =\tilde{T}_{uc}\left( y\right) $ if $y\notin \mathfrak{M}_{0}$ and $ \tilde{T}_{uc}^{\blacktriangle }(y)=\infty $ otherwise. Its bootstrapped version is then given by $\tilde{T}_{uc}^{\blacktriangle ,\ast }(y,{\Greekmath 0118} )= \tilde{T}_{uc}^{\blacktriangle }\left( y^{\ast }(y,{\Greekmath 0118} )\right) $. For every $y\in \mathbb{R}^{n}$ we denote its distribution function under $\Xi $ by $ \tilde{F}_{uc,y}^{\blacktriangle }$, i.e., $\tilde{F}_{uc,y}^{\blacktriangle }(t)=\Xi \mathbb{(}\tilde{T}_{uc}^{\blacktriangle ,\ast }(y,{\Greekmath 0118} )\leq t)$ for $t\in \mathbb{R\cup \{\infty \}}$.

theorem(a) For every ${\Greekmath 010B} \in (0,1)$, let $\tilde{f} _{uc,1-{\Greekmath 010B} }(y)$ denote a $(1-{\Greekmath 010B} )$-quantile of $\tilde{F}_{uc,y}$. Define \begin{equation} \tilde{{\Greekmath 0123}}_{uc}=1-\max_{\substack{ i=1,\ldots ,n, \\ {\Greekmath 0116} _{0}+e_{i}(n)\notin \mathfrak{M}_{0}}}\Xi \left( \left\{ {\Greekmath 0118} :\tilde{T} _{uc}^{\ast }({\Greekmath 0116} _{0}+e_{i}(n),{\Greekmath 0118} )<\tilde{T}_{uc}({\Greekmath 0116} _{0}+e_{i}(n)),y^{\ast }({\Greekmath 0116} _{0}+e_{i}(n),{\Greekmath 0118} )\notin \mathfrak{M} _{0}\right\} \right) \end{equation} for some ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$.\footnote{ Note that the index set in the maximum operator in ((ref) ) can not be empty since $k-q\leq k$ and we have assumed $k<n$.} Then $ \tilde{{\Greekmath 0123}}_{uc}$ does not depend on the choice of ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$. Furthermore, for every ${\Greekmath 010B} \in (0,1)$ such that $ {\Greekmath 010B} >\tilde{{\Greekmath 0123}}_{uc}$ holds, we have \begin{equation} \sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }\left( \tilde{T}_{uc}\geq \tilde{f}_{uc,1-{\Greekmath 010B} }\right) \geq \sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }\left( \tilde{T}_{uc}> \tilde{f}_{uc,1-{\Greekmath 010B} }\right) =1 \end{equation} for every ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$ and every $0<{\Greekmath 011B} ^{2}<\infty $ (where the probabilities in ((ref)) are to be interpreted as inner probabilities). (b) For every ${\Greekmath 010B} \in (0,1)$, let $\tilde{f}_{uc,1-{\Greekmath 010B} }^{\blacktriangle }(y)$ denote a $(1-{\Greekmath 010B} )$-quantile of $\tilde{F} _{uc,y}^{\blacktriangle }$. Then, with $\tilde{{\Greekmath 0123}}_{uc}$ defined in Part (a), for every ${\Greekmath 010B} \in (0,1)$ such that ${\Greekmath 010B} >\tilde{{\Greekmath 0123}} _{uc}$ holds, we have \begin{equation} \sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }\left( \tilde{T}_{uc}\geq \tilde{f}_{uc,1-{\Greekmath 010B} }^{\blacktriangle }\right) \geq \sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }\left( \tilde{T}_{uc}>\tilde{f}_{uc,1-{\Greekmath 010B} }^{\blacktriangle }\right) =1 \end{equation} for every ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$ and every $0<{\Greekmath 011B} ^{2}<\infty $ (where the probabilities in ((ref)) are to be interpreted as inner probabilities).

Mutatis mutandis, a discussion similar to the one given subsequently to Theorem (ref) also applies to the preceding two theorems.

The second bootstrap scheme

For the test statistics $\tilde{T}_{Het}$ and $\tilde{T}_{uc}$ an analogon to Lemma (ref) is not available. Hence, we need to provide separate theorems for the case where the second bootstrap scheme, i.e., ((ref) ), is being used. This is done next. With this bootstrap scheme, the bootstrapped test statistic corresponding to $\tilde{T}_{Het}$ is given by

equation*[equation* omitted — 387 chars of source]

if $y^{\maltese }(y,{\Greekmath 0118} )\notin \mathsf{\tilde{B}}$, and by $\tilde{T} _{Het}^{\maltese }(y,{\Greekmath 0118} )=0$ if $y^{\maltese }(y,{\Greekmath 0118} )\in \mathsf{\tilde{B}} $. For every $y\in \mathbb{R}^{n}$ denote the distribution function of the bootstrapped test statistic under $\Xi $ by $\tilde{H}_{Het,y}$, i.e., $ \tilde{H}_{Het,y}(t)=\Xi \mathbb{(}\tilde{T}_{Het}^{\maltese }(y,{\Greekmath 0118} )\leq t) $ for $t\in \mathbb{R}$. Furthermore, $\tilde{T}_{Het}^{\blacktriangle ,\maltese }(y,{\Greekmath 0118} )$ is defined exactly as is $\tilde{T}_{Het}^{\maltese }(y,{\Greekmath 0118} )$, except that $\tilde{T}_{Het}^{\blacktriangle ,\maltese }(y,{\Greekmath 0118} )=\infty $ if $y^{\maltese }(y,{\Greekmath 0118} )\in \mathsf{\tilde{B}}$. For every $y\in \mathbb{R}^{n}$ denote its distribution function under $\Xi $ by $\tilde{H} _{Het,y}^{\blacktriangle }$, i.e., $\tilde{H}_{Het,y}^{\blacktriangle }(t)=\Xi \mathbb{(}\tilde{T}_{Het}^{\blacktriangle ,\maltese }(y,{\Greekmath 0118} )\leq t) $ for $t\in \mathbb{R\cup \{\infty \}}$.

theoremSuppose Assumption (ref) holds. (a) For every ${\Greekmath 010B} \in (0,1)$, let $\tilde{h}_{Het,1-{\Greekmath 010B} }(y)$ denote a $(1-{\Greekmath 010B} )$-quantile of $\tilde{H}_{Het,y}$. Define \begin{equation} \tilde{{\Greekmath 0112}}_{Het}=1-\max_{\substack{ i=1,\ldots ,n, \\ {\Greekmath 0116} _{0}+e_{i}(n)\notin \mathsf{\tilde{B}}}}\Xi \left( \left\{ {\Greekmath 0118} :\tilde{T} _{Het}^{\maltese }({\Greekmath 0116} _{0}+e_{i}(n),{\Greekmath 0118} )<\tilde{T}_{Het}({\Greekmath 0116} _{0}+e_{i}(n)),y^{\maltese }({\Greekmath 0116} _{0}+e_{i}(n),{\Greekmath 0118} )\notin \mathsf{\tilde{B}} \right\} \right) \end{equation} for some ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$, with the convention that $\tilde{ {\Greekmath 0112}}_{Het}=1$ if the index set in the maximum operator in ((ref)) is empty. Then $\tilde{{\Greekmath 0112}}_{Het}$ does not depend on the choice of ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$. Furthermore, for every $ {\Greekmath 010B} \in (0,1)$ such that ${\Greekmath 010B} >\tilde{{\Greekmath 0112}}_{Het}$ holds, we have \begin{equation} \sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }\left( \tilde{T}_{Het}\geq \tilde{h}_{Het,1-{\Greekmath 010B} }\right) \geq \sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }\left( \tilde{T}_{Het}> \tilde{h}_{Het,1-{\Greekmath 010B} }\right) =1 \end{equation} for every ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$ and every $0<{\Greekmath 011B} ^{2}<\infty $ (where the probabilities in ((ref)) are to be interpreted as inner probabilities). (b) For every ${\Greekmath 010B} \in (0,1)$, let $\tilde{h}_{Het,1-{\Greekmath 010B} }^{\blacktriangle }(y)$ denote a $(1-{\Greekmath 010B} )$-quantile of $\tilde{H} _{Het,y}^{\blacktriangle }$. Then, with $\tilde{{\Greekmath 0112}}_{Het}$ defined in Part (a), for every ${\Greekmath 010B} \in (0,1)$ such that ${\Greekmath 010B} >\tilde{{\Greekmath 0112}} _{Het}$ holds, we have \begin{equation} \sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }\left( \tilde{T}_{Het}\geq \tilde{h}_{Het,1-{\Greekmath 010B} }^{\blacktriangle }\right) \geq \sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }\left( \tilde{T}_{Het}>\tilde{h}_{Het,1-{\Greekmath 010B} }^{\blacktriangle }\right) =1 \end{equation} for every ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$ and every $0<{\Greekmath 011B} ^{2}<\infty $ (where the probabilities in ((ref)) are to be interpreted as inner probabilities).

With the bootstrap scheme considered in this subsection, the bootstrapped test statistic corresponding to $\tilde{T}_{uc}$ is given by

equation*[equation* omitted — 442 chars of source]

if $y^{\maltese }(y,{\Greekmath 0118} )\notin \mathfrak{M}_{0}$, and by $\tilde{T} _{uc}^{\maltese }(y,{\Greekmath 0118} )=0$ if $y^{\maltese }(y,{\Greekmath 0118} )\in \mathfrak{M}_{0}$. For every $y\in \mathbb{R}^{n}$ denote the distribution function of the bootstrapped test statistic under $\Xi $ by $\tilde{H}_{uc,y}$, i.e., $ \tilde{H}_{uc,y}(t)=\Xi \mathbb{(}\tilde{T}_{uc}^{\maltese }(y,{\Greekmath 0118} )\leq t)$ for $t\in \mathbb{R}$. Furthermore, $\tilde{T}_{uc}^{\blacktriangle ,\maltese }(y,{\Greekmath 0118} )$ is defined exactly as is $\tilde{T}_{uc}^{\maltese }(y,{\Greekmath 0118} )$, except that $\tilde{T}_{uc}^{\blacktriangle ,\maltese }(y,{\Greekmath 0118} )=\infty $ if $y^{\maltese }(y,{\Greekmath 0118} )\in \mathfrak{M}_{0}$. For every $y\in \mathbb{R}^{n}$ denote its distribution function under $\Xi $ by $\tilde{H} _{uc,y}^{\blacktriangle }$, i.e., $\tilde{H}_{uc,y}^{\blacktriangle }(t)=\Xi \mathbb{(}\tilde{T}_{uc}^{\blacktriangle ,\maltese }(y,{\Greekmath 0118} )\leq t)$ for $ t\in \mathbb{R\cup \{\infty \}}$.

theorem(a) For every ${\Greekmath 010B} \in (0,1)$, let $\tilde{h} _{uc,1-{\Greekmath 010B} }(y)$ denote a $(1-{\Greekmath 010B} )$-quantile of $\tilde{H}_{uc,y}$. Define \begin{equation} \tilde{{\Greekmath 0112}}_{uc}=1-\max_{\substack{ i=1,\ldots ,n, \\ {\Greekmath 0116} _{0}+e_{i}(n)\notin \mathfrak{M}_{0}}}\Xi \left( \left\{ {\Greekmath 0118} :\tilde{T} _{uc}^{\maltese }({\Greekmath 0116} _{0}+e_{i}(n),{\Greekmath 0118} )<\tilde{T}_{uc}({\Greekmath 0116} _{0}+e_{i}(n)),y^{\maltese }({\Greekmath 0116} _{0}+e_{i}(n),{\Greekmath 0118} )\notin \mathfrak{M} _{0}\right\} \right) \end{equation} for some ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$.\footnote{ Note that the index set in the maximum operator in ((ref)) can not be empty since $k-q\leq k$ and we have assumed $k<n$.} Then $\tilde{{\Greekmath 0112}}_{uc}$ does not depend on the choice of $ {\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$. Furthermore, for every ${\Greekmath 010B} \in (0,1)$ such that ${\Greekmath 010B} >\tilde{{\Greekmath 0112}}_{uc}$ holds, we have \begin{equation} \sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }\left( \tilde{T}_{uc}\geq \tilde{h}_{uc,1-{\Greekmath 010B} }\right) \geq \sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }\left( \tilde{T}_{uc}> \tilde{h}_{uc,1-{\Greekmath 010B} }\right) =1 \end{equation} for every ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$ and every $0<{\Greekmath 011B} ^{2}<\infty $ (where the probabilities in ((ref)) are to be interpreted as inner probabilities). (b) For every ${\Greekmath 010B} \in (0,1)$, let $\tilde{h}_{uc,1-{\Greekmath 010B} }^{\blacktriangle }(y)$ denote a $(1-{\Greekmath 010B} )$-quantile of $\tilde{H} _{uc,y}^{\blacktriangle }$. Then, with $\tilde{{\Greekmath 0112}}_{uc}$ defined in Part (a), for every ${\Greekmath 010B} \in (0,1)$ such that ${\Greekmath 010B} >\tilde{{\Greekmath 0112}}_{uc}$ holds, we have \begin{equation} \sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }\left( \tilde{T}_{uc}\geq \tilde{h}_{uc,1-{\Greekmath 010B} }^{\blacktriangle }\right) \geq \sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }\left( \tilde{T}_{uc}>\tilde{h}_{uc,1-{\Greekmath 010B} }^{\blacktriangle }\right) =1 \end{equation} for every ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$ and every $0<{\Greekmath 011B} ^{2}<\infty $ (where the probabilities in ((ref)) are to be interpreted as inner probabilities).

Mutatis mutandis, a discussion similar to the one given subsequently to Theorem (ref) also applies to the preceding two theorems.

Further remarks

remarkAs already noted earlier, the various lower bounds for ${\Greekmath 010B} $ given in the theorems depend only on observable quantities, and can thus be computed numerically. In particular, ${\Greekmath 0123} _{Het}$ in Theorem (ref) depends only on $X$, $R$, $r$, $\mathcal{A}$, $\Xi $, and on the $d_{i}$'s appearing in the definition of the test statistic, while $ {\Greekmath 0123} _{uc}$ in Theorem (ref) depends only on $X$, $R$, $ r$, $\mathcal{A}$, and $\Xi $. Similarly, $\tilde{{\Greekmath 0123}}_{Het}$ in Theorem (ref) and $\tilde{{\Greekmath 0112}}_{Het}$ in Theorem (ref) depend only on $X$, $R$, $r$, $\mathcal{A}$, $\Xi $, and on the $\tilde{d}_{i}$'s appearing in the definition of the test statistic. Finally, $\tilde{{\Greekmath 0123}}_{uc}$ in Theorem (ref) and $ \tilde{{\Greekmath 0112}}_{uc}$ in Theorem (ref) depend only on $X$, $R$ , $r$, $\mathcal{A}$, and $\Xi $.
remark(i) Relations ((ref)) and ((ref)) continue to hold a fortiori if $T_{Het}$ is replaced by $T_{Het}^{\blacktriangle }$, as $T_{Het}^{\blacktriangle }$ is never smaller than $T_{Het}$ (in fact, $T_{Het}^{\blacktriangle }$ and $ T_{Het}$ coincide except on a Lebesgue null set under Assumption (ref)). (ii) Relations ((ref)) and ((ref)) continue to hold a fortiori if $T_{uc}$ is replaced by $T_{uc}^{\blacktriangle }$, as $T_{uc}^{\blacktriangle }$ is never smaller than $T_{uc}$ (in fact, both coincide except on $\limfunc{span} (X)$, a Lebesgue null set). (iii) Relations ((ref)), ((ref)), ((ref)), and ((ref)) continue to hold a fortiori if $ \tilde{T}_{Het}$ is replaced by $\tilde{T}_{Het}^{\blacktriangle }$, as $ \tilde{T}_{Het}^{\blacktriangle }$ is never smaller than $\tilde{T}_{Het}$ (in fact, both coincide except on a Lebesgue null set under Assumption (ref)). (iv) Relations ((ref)), ((ref)), ((ref) ), and ((ref)) continue to hold a fortiori if $\tilde{T}_{uc}$ is replaced by $\tilde{T}_{uc}^{\blacktriangle } $, as $\tilde{T}_{uc}^{\blacktriangle }$ is never smaller than $\tilde{T} _{uc}$ (in fact, both coincide except on $\mathfrak{M}_{0}$, a Lebesgue null set).
remarkIn case $q=1$, inspection of the proof of Theorem (ref) shows that the bound ${\Greekmath 0123} _{Het}$ can be somewhat improved by allowing in the definition of ${\Greekmath 0123} _{2,Het}$ the index $i$ to range over all indices such that $e_{i}(n)\in \mathsf{B}$ and $R\hat{{\Greekmath 010C}}(e_{i}(n))\neq 0$ . [This is so, since in case $q=1$ singularity of $\hat{\Omega}_{Het}\left( y\right) $ is equivalent to $\hat{\Omega}_{Het}\left( y\right) =0$.]
remarkThe test statistics $T_{Het}$ using HC0 and HC1 weights, respectively, differ only by a multiplicative constant, and hence result in the same bootstrap-based test. For design matrices $X$ with $h_{ii}$ not depending on $i$, the same conclusion applies for all weights HC0-HC4. A similar remark applies to $\tilde{T}_{Het}$ (with $\tilde{h}_{ii}$ taking the r\^{o}le of $h_{ii}$).
remarkThe size one results for bootstrap-based tests given in the preceding theorems are easily seen to imply infimal coverage zero results for the corresponding confidence sets for $R{\Greekmath 010C} $ obtained by \textquotedblleft inverting\textquotedblright\ the tests. The computation of such confidence sets is straightforward and leads to ellipsoids in the case where the bootstrap-based test is obtained from $T_{Het}$ or $T_{uc}$ and the bootstrap scheme ((ref)) with $\mathcal{A}=\limfunc{span}(X)$ is used. This is so, since the covariance matrix estimator employed in $T_{Het}$ (or $T_{uc}$, respectively) does not depend on $r$, and since the quantile of the bootstrap distribution is easily seen also not to depend on $r$ in this case. By Lemma (ref) the same is true for the bootstrap-based tests obtained from $T_{Het}$ or $T_{uc}$ and the bootstrap scheme ((ref)) with $\mathcal{A}=\limfunc{span}(X)$. In all other combinations of test statistics and bootstrap schemes the "inversion" is typically more complicated and becomes numerically burdensome, as then the covariance matrix estimator employed in the test statistic and/or the quantile of the bootstrap distribution will typically depend on $r$.

Some special cases

Here we consider the special case where the null hypothesis is simple (i.e., $q=k$). If restricted residuals are used in the bootstrap scheme ((ref)) (i.e., if $\mathcal{A}=\mathfrak{M}_{0}$ holds), the following result shows that in these cases our theorems become vacuous, and thus do not allow us to draw any conclusion about the sizes of the corresponding bootstrap-based tests. [Of course, this by itself does not preclude the possibility that in these cases the size may be equal to one or may substantially exceed ${\Greekmath 010B} $.] The observations made in the theorem below are in line with a result in DavidsFlach2008 implying that -- in the case corresponding to Part (c) of the subsequent theorem -- the bootstrap-based test using the bootstrap scheme ((ref)) with $ \mathcal{A}=\mathfrak{M}_{0}$ indeed has size equal to the nominal significance level ${\Greekmath 010B} $, provided a particular choice of $\Xi $ and particular values of ${\Greekmath 010B} $ are used. [In fact, this result, which is Theorem 1 in DavidsFlach2008, is not entirely correct in the form given, but needs some amendments and corrections, which we shall not provide here.\footnote{A simple counterexample to Theorem 1 in DavidsFlach2008 is provided by a regression model which has a standard basis vector as its only regressor. It is then easy to see that the test statistic is (almost surely) constant and coincides with the bootstrapped test statistic. Consequently, the bootstrap-based test becomes trivial. Its null rejection probabilities are equal to $0$ if the p-value is defined as in DavidsFlach2008. [They are equal to $1$ if an alternative definition of the p-value is used.]}]

theoremSuppose $q=k$ and $\Xi (\left\{ {\Greekmath 0118} \in \mathbb{R} ^{n}:{\Greekmath 0118} _{i}\neq 0\right\} )=1$ for every $i=1,\ldots ,n$. Then: (a) ${\Greekmath 0123} _{Het}=1$ holds in Theorem (ref), if $ \mathcal{A}=\mathfrak{M}_{0}$ is used in the bootstrap scheme.\footnote{ Theorem (ref) maintains Assumption (ref) . If this assumption is violated, then $T_{Het}$ is identically equal to zero, leading to a useless test. If one would formally apply bootstrap scheme ((ref)), the bootstrapped test statistic $T_{Het}^{\ast }$ would then also be identically zero, leading to a bootstrap critical value of zero. The rejection probability is then always equal to zero or always equal to one, depending on whether one uses a strict or weak inequality in the definition of the rejection region.} (b) ${\Greekmath 0123} _{uc}=1$ holds in Theorem (ref), if $\mathcal{ A}=\mathfrak{M}_{0}$ is used in the bootstrap scheme. (c) $\tilde{{\Greekmath 0123}}_{Het}=1$ holds in Theorem (ref), if $ \mathcal{A}=\mathfrak{M}_{0}$ is used in the bootstrap scheme.\footnote{ Theorem (ref) maintains Assumption (ref). If this assumption is violated, then a similar comment as in Footnote (ref) applies.} (d) $\tilde{{\Greekmath 0123}}_{uc}=1$ holds in Theorem (ref), if $ \mathcal{A}=\mathfrak{M}_{0}$ is used in the bootstrap scheme.
remark(i) Part (a) (Part (b), respectively) of the preceding theorem applies to the bootstrap-based test derived from $T_{Het}$ ($T_{uc}$, respectively) when the bootstrap scheme ((ref)) with $\mathcal{A}=\mathfrak{M}_{0}$ is employed. In view of Lemma (ref) and Remark (ref), these results also apply if the bootstrap scheme ((ref)), again with $ \mathcal{A}=\mathfrak{M}_{0}$, is used. However, as can be seen from simple examples, this is not so in the context of Parts (c) and (d) of the preceding theorem (i.e., $\tilde{{\Greekmath 0112}}_{Het}<1$ and $\tilde{{\Greekmath 0112}}_{uc}<1$ can occur in Theorems (ref) and (ref). respectively, even if $q=k$, $\mathcal{A}=\mathfrak{M}_{0}$, and $\Xi $ is as in Theorem (ref)). (ii) If bootstrap schemes ((ref)) or ((ref)) are used for the bootstrap-based tests derived from any of $T_{Het}$, $T_{uc}$, $\tilde{T} _{Het}$, and $\tilde{T}_{uc}$, but now with $\mathfrak{M}_{0}\subsetneqq \mathcal{A}$, simple examples show that ${\Greekmath 0123} _{Het}<1$, ${\Greekmath 0123} _{uc}<1$, $\tilde{{\Greekmath 0123}}_{Het}<1$, $\tilde{{\Greekmath 0123}}_{uc}<1$, $\tilde{ {\Greekmath 0112}}_{Het}<1$, and $\tilde{{\Greekmath 0112}}_{uc}<1$ can occur.

As a consequence of the preceding theorem, the function in the R-package wbsd for computing ${\Greekmath 0123} _{Het}$, etc. first checks if the condition of the theorem are satisfied, and if so, outputs $1$ for $ {\Greekmath 0123} _{Het}$, etc.

Extensions and generalizations

Other covariance models

The results given so far refer to the size of bootstrap-based tests when the covariance model $\mathfrak{C}_{Het}$ is maintained. Inspection of the proofs of the theorems in Sections (ref) and (ref) as well as of Theorem (ref) in Appendix (ref) shows that they also hold with $\mathfrak{C}_{Het}$ replaced by any covariance model $\mathfrak{C} \subseteq \mathfrak{C}_{Het}$ other than $\mathfrak{C}_{Het}$, provided the closure of $\mathfrak{C}$ contains, for $i=1,\ldots ,n$, the matrices $ e_{i}(n)e_{i}(n)^{\prime }$. More generally, if the closure of $\mathfrak{C}$ contains the matrices $e_{i}(n)e_{i}(n)^{\prime }$ only for $i\in I\subseteq \{1,\ldots ,n\}$, then Theorem (ref) continues to hold with $ \mathfrak{C}_{Het}$ replaced by $\mathfrak{C}$, provided the range of the maximum operator in ((ref)) is intersected with $I$, and the other theorems mentioned before continue to hold with $\mathfrak{C}_{Het}$ replaced by $\mathfrak{C}$, provided the ranges of the maximum operators appearing in the definitions of the various quantities ${\Greekmath 0123} _{1,Het}$, ${\Greekmath 0123} _{2,Het}$, ${\Greekmath 0123} _{1,uc}$, ${\Greekmath 0123} _{2,uc}$, $\tilde{ {\Greekmath 0123}}_{Het}$, $\tilde{{\Greekmath 0123}}_{uc}$, $\tilde{{\Greekmath 0112}}_{Het}$, and $ \tilde{{\Greekmath 0112}}_{uc}$ are intersected with $I$.\footnote{ It is here understood that a maximum is interpreted as zero if it extends over an empty range.}

Nongaussian errors

Consider now the regression model as in Section (ref), except for the Gaussianity assumption.

(i) If we assume a (possibly semiparametric) model for the distribution of the errors such that the implied model for the distributions of $\mathbf{Y}$ contains all the Gaussian distributions shown in ((ref)), then the results of the paper continue to hold a fortiori (with the same lower bounds for ${\Greekmath 010B} $), since the size of any test computed w.r.t. such a larger model for the distributions of $\mathbf{Y}$ is certainly not smaller than the size of the same test when computed w.r.t. to the Gaussian model ((ref)).

(ii) Suppose next we assume that the standardized errors ${\Greekmath 011B} ^{-1}\Sigma ^{-1/2}\mathbf{U}$ follow a (fixed) distribution $G$ that does not depend on $({\Greekmath 0116} ,{\Greekmath 011B} ,\Sigma )$, and let $Q_{{\Greekmath 0116} ,{\Greekmath 011B} ^{2}\Sigma ,G}$ denote the implied distribution of $\mathbf{Y}$. If $G$ is absolutely continuous w.r.t. ${\Greekmath 0115} _{\mathbb{R}^{n}}$, then the results of the paper continue to hold with $P_{{\Greekmath 0116} ,{\Greekmath 011B} ^{2}\Sigma }$ replaced by $Q_{{\Greekmath 0116} ,{\Greekmath 011B} ^{2}\Sigma ,G} $ (and with the same lower bounds for ${\Greekmath 010B} $). This is easily seen from an inspection of the proofs.\footnote{The assumption of absolute continuity of $G$ can, in fact, be relaxed to the assumption that none of its one-dimensional marginals has positive mass at zero in case of Theorems (ref), (ref) - (ref), and of the weaker versions of Theorems (ref) and (ref) discussed in Remark (ref) in Appendix (ref). The proofs of Theorems (ref) and (ref) also extend to the situation discussed here under any additional condition that guarantees that the exceptional sets $\mathsf{B}$ and $\limfunc{span} (X) $, respectively, have probability zero under any $Q_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma ,G}$.}

(iii) Suppose we have the same framework as in (ii), except that now $G$ varies in a set $\mathfrak{G}$ (independently of $({\Greekmath 0116} ,{\Greekmath 011B} ,\Sigma )$), i.e., we have a semiparametric model. If at least one member $G\in \mathfrak{ G}$ is absolutely continuous, then the results of the paper continue to hold a fortiori (with the same lower bounds for ${\Greekmath 010B} $) for reasons similar to the ones given in (i).\footnote{ Again, the absolute continuity assumption can be weakened, cf. Footnote (ref).}

(iv) Also note that the lower bounds for ${\Greekmath 010B} $ in all the result do not involve the distribution of $\mathbf{Y}$ (and thus of $\mathbf{U})$, and, in particular, do not involve the Gaussianity assumption. Hence, in this sense the lower bounds are \textquotedblleft distribution free\textquotedblright .

Stochastic regressors

The assumption of nonstochastic regressors can be easily relaxed as follows: Suppose $X$ is random and $\mathbf{U}$ is conditionally on $X$ distributed as $N(0,{\Greekmath 011B} ^{2}\Sigma )$, with ${\Greekmath 011B} ^{2}={\Greekmath 011B} ^{2}(X)>0$ and $ \Sigma =\Sigma (X)\in \mathfrak{C}_{Het}$, where ${\Greekmath 011B} ^{2}(\cdot )$ and $ \Sigma (\cdot )$ vary in given classes of functions. Suppose further that $ {\Greekmath 011B} ^{2}(X)$ and $\Sigma (X)$ vary independently through all of $ (0,\infty )$ and $\mathfrak{C}_{Het}$, respectively, for (almost) every realization of $X$, when the functions ${\Greekmath 011B} ^{2}(\cdot )$ and $\Sigma (\cdot )$ vary in the before mentioned function classes.\footnote{ This is certainly the case if no restrictions on the functions ${\Greekmath 011B} ^{2}(\cdot )$ and $\Sigma (\cdot )$ are imposed beyond ${\Greekmath 011B} ^{2}(\cdot )$ and $\Sigma (\cdot )$, respectively, taking values in $(0,\infty )$ and in the set of diagonal matrices with positive diagnal elements. Another instance where this is seen to be satisfied is the case where ${\Greekmath 011B} ^{2}(X)\Sigma (X)=\limfunc{diag}({\Greekmath 0124} (x_{1\cdot }),\ldots ,{\Greekmath 0124} (x_{n\cdot }))$ with no further restrictions on the function ${\Greekmath 0124} $ (besides positivity) and with at least one regressor being absolutely continuous.} Then the results of the paper obviously apply after one conditions on $X$ provided (almost) all realizations of $X$ satisfy the assumptions of our theorems, which will typically be the case (for brevity we do not provide a formal statement here). And again similar generalizations to non-Gaussianity as discussed in the preceding subsection are possible here.

Numerical results

There is a considerable body of simulation studies investigating finite sample properties of bootstrap-based heteroskedasticity robust tests, see the references mentioned in the Introduction. While these studies provide helpful information, there is -- as always with simulation studies -- an issue to what extent conclusions of such a study generalize. This is particularly so with positive findings (such as, e.g., that a particular bootstrap-based test has null rejection probabilities close to the nominal significance level) as it is less than clear that such a finding allows for generalization beyond the design matrices $X$, the restrictions (given by $R$, $r$), and the forms of heteroskedasticity considered in the simulation study. It is less of an issue with negative results (such as, e.g., that a particular bootstrap-based test has null rejection probabilities much larger than the nominal significance level), since they can be viewed as counterexamples disproving good behavior of the bootstrap-based test in general.

For these reasons we set out to study the worst-case size performance of a variety of bootstrap-based tests.\footnote{ The concept of the size of a test is by itself already a worst-case concept, as it is the supremum of the rejection probabilites over the null hypothesis (where $X$ and the restrictions are being held fixed), cf. ((ref)). The term "worst-case" in "worst-case size performance" here refers to varying $X$ and the restrictions to be tested.} That is, for any given bootstrap-based test in a large class, we try to "break" the test by searching for a design matrix $X$, a restriction (given by $R$, $r$) to be tested, and a form of heteroskedasticity, such that the null rejection probability of the test is substantially larger than the nominal significance level ${\Greekmath 010B} $. Note that this is equivalent to finding $X$, $R$, and $r$ such that the size of the bootstrap-based test computed over the heteroskedasticity model $ \mathfrak{C}_{Het}$ is substantially larger than ${\Greekmath 010B} $. A bootstrap-based test that is "broken" in our study, should probably not be used by practitioners (at least not without first assessing its properties in the particular testing problem put before the practitioner, e.g., by attempting to determining the size of the test in that problem by Monte Carlo methods). A bootstrap-based test that "survives" the "stress test" imposed by our study may perhaps be considered to be a better choice, but note that our numerical results do not provide any guarantee for good performance in the practitioner's testing problem either (and thus again an assessment of its properties in the practitioner's testing problem may be called for).

Our theoretical results obtained in the previous sections play an important r \^{o}le in our study of the size performance of bootstrap-based tests, as these results allow us to deduce abysmal size behavior (i.e., size equal to $ 1$) of a bootstrap-based test (for a given design matrix $X$ and restriction $(R$, $r)$ to be tested) by comparing the (numerically evaluated) ${\Greekmath 0123} $ with the nominal significance level ${\Greekmath 010B} $; in this section the symbol $ {\Greekmath 0123} $ serves as a generic abbreviation for ${\Greekmath 0123} _{Het}$($={\Greekmath 0112} _{Het}$) (cf. Theorem (ref)), ${\Greekmath 0123} _{uc}$($={\Greekmath 0112} _{uc}$) (cf. Theorem (ref)), $\tilde{{\Greekmath 0123}}_{Het}$ (cf. Theorem (ref)), $\tilde{{\Greekmath 0112}}_{Het}$ (cf. Theorem (ref)), $\tilde{{\Greekmath 0123}}_{uc}$ (cf. Theorem (ref)), or $\tilde{{\Greekmath 0112}}_{uc}$ (cf. Theorem (ref)), depending on which of the theorems listed in parentheses (possibly after an appeal to Lemma (ref) and Remark (ref)) applies to the bootstrap-based test under consideration. Recall that these theorems show that if ${\Greekmath 010B} >{\Greekmath 0123} $ holds, then the corresponding bootstrap-based test has size $1$ (over the heteroskedasticity model $ \mathfrak{C}_{Het}$), and thus certainly breaks down (for the given testing problem, i.e., for the given $X$, $R$, $r$). Hence, we shall search for worst-case $X$, $R$, and $r$ that lead to small values of ${\Greekmath 0123} $. \footnote{ Evaluating ${\Greekmath 0123} $ numerically is a nontrivial task as discussed in Section (ref) in Appendix (ref). In order to be on the safe side and to bias our results in favor of the tests (recall that we are after negative results), the ${\Greekmath 0123} $ we shall report will actually be a numerical obtained upper bound for the true ${\Greekmath 0123} $.}

More precisely, we shall consider three settings, where setting refers to sample size $n=10,20$, and $30$, and various scenarios, where scenario refers to a combination of $k$ (number of regressors) and $q$ (number of restrictions to be tested and where $R=(0:I_{q})$, $r=0$ ). In every setting, we roughly do the following: we compute for every bootstrap-based test included in our study the value of ${\Greekmath 0123} $ for a variety of design matrices in a range of scenarios. We then determine the minimal value of ${\Greekmath 0123} $ over all design matrices considered. Then, we compare this minimum with two commonly used levels of significance (${\Greekmath 010B} =0.05$ and ${\Greekmath 010B} =0.1$). In addition to studying the behavior of this minimal value of ${\Greekmath 0123} $, we shall complement this by numerical size lower bound computations.

All computations were carried out in R (R) version 3.6.3 using version 1.0.0 of the R-package wbsd (\textquotedblleft wild bootstrap size diagnostics\textquotedblright ) by wbsd generated with Rtools35. The package wbsd provides computationally efficient routines for determining the quantities ${\Greekmath 0123} _{Het}$($={\Greekmath 0112} _{Het}$ ), ${\Greekmath 0123} _{uc}$($={\Greekmath 0112} _{uc}$), $\tilde{{\Greekmath 0123}}_{Het}$, $\tilde{ {\Greekmath 0112}}_{Het}$, $\tilde{{\Greekmath 0123}}_{uc}$, and $\tilde{{\Greekmath 0112}}_{uc}$, and for obtaining bootstrap p-values in order to obtain the numerical results reported here. The tools for computing ${\Greekmath 0123} $ provided in the R-package wbsd can be used by practitioners as a diagnostic device to check whether a bootstrap-based test is provably unreliable (in that $ {\Greekmath 010B} >{\Greekmath 0123} $, which implies size equal to $1$) in a given testing problem. The package is available on CRAN.

In the following subsections we describe the bootstrap-based tests studied, we explain the computations carried out for each test, and discuss the results obtained. Some of the details are deferred to Appendix (ref).

Description of the bootstrap-based tests studied

The number of bootstrap-based tests we cover in our study is vast: In total we consider the $960$ possible combinations of the $12$ test statistics discussed in Section (ref) and at the beginning of Section (ref) around Equations ((ref)), ((ref)), ((ref)), and ((ref)) with the bootstrap schemes discussed further below ($80$ in total), which are popular special cases of the two general bootstrap schemes discussed in Section (ref).

To be precise, the $12$ test statistics studied are:\footnote{We use here the conventions for $d_{i}$ and $\tilde{d}_{i}$ given below ((ref)) and ((ref)), respectively.}

enumerate• Test statistics based on unrestricted residuals: $T_{uc}$; and $T_{Het}$ with $d_{i}=1$ (HC0); with $d_{i}=n/(n-k)$ (HC1); with $ d_{i}=(1-h_{ii})^{-1}$ (HC2); with $d_{i}=(1-h_{ii})^{-2}$ (HC3); and with $ d_{i}=(1-h_{ii})^{{\Greekmath 010E} _{i}}$ for ${\Greekmath 010E} _{i}=\min (nh_{ii}/k,4)$ (HC4). • Test statistics based on restricted residuals: $\tilde{T} _{uc} $; $\tilde{T}_{Het}$ with $\tilde{d}_{i}=1$ (HC0R); with $\tilde{d} _{i}=n/(n-(k-q))$ (HC1R); with $\tilde{d}_{i}=(1-\tilde{h}_{ii})^{-1}$ (HC2R); with $\tilde{d}_{i}=(1-\tilde{h}_{ii})^{-2}$ (HC3R); and with $ \tilde{d}_{i}=(1-\tilde{h}_{ii})^{\tilde{{\Greekmath 010E}}_{i}}$ for $\tilde{{\Greekmath 010E}} _{i}=\min (n\tilde{h}_{ii}/(k-q),4)$ (HC4R).

The bootstrap schemes we study are $y^{\ast }$ as defined in ((ref)), and $y^{\maltese }$ as defined in ((ref)). Both bootstrap schemes are applied with $\mathcal{A}=\mathfrak{M}_{0}$ as well as with $ \mathcal{A}=\limfunc{span}(X)$. In addition to choosing $\mathcal{A}$, both bootstrap schemes require a concrete choice of $\Xi $. All distributions $ \Xi $ we consider are constructed in the following way: first an auxiliary distribution $\Xi ^{\bullet }$ on $\{-1,1\}^{n}$ has to be chosen. The way we choose this auxiliary distribution depends on the magnitude of $n$. We consider three cases: Setting A ($n=10$), Setting B ($n=20$), and Setting C ( $n=30$).

itemize• In Setting A, we consider (i) $\Xi ^{\bullet }$ equal to the $n$-fold Rademacher distribution\ (i.e., the $n$-fold product of the uniform distribution on $\{-1,1\}$), and (ii) $\Xi ^{\bullet }$ equal to the $n$ -fold Mammen distribution\ (i.e., the $n$-fold product of the distribution on $\{-(\sqrt{5}-1)/2,(\sqrt{5}+1)/2\}$ that assigns mass $(\sqrt{5}+1)/(2 \sqrt{5})$ to $-(\sqrt{5}-1)/2$). • In Settings B and C, we consider $\Xi ^{\bullet }$ equal to an empirical distribution of a sample of size $10n-1$ from the $n$-fold Rademacher distribution and from the $n$-fold Mammen distribution, respectively.\footnote{ The reason for treating Settings B and C differently from Setting A is that for values of $n$ such as $20$ or $30$ an enumeration of all support points of the $n$-fold Rademacher or $n$-fold Mammen distribution is numerically too costly.}

Given an auxiliary distribution $\Xi ^{\bullet }$, the distribution $\Xi $ actually used in the bootstrap scheme depends on a vector of weights $w$, itself typically depending on $X$ or on $X$ and $R$. Given a weights vector $ w$, $\Xi $ is then obtained as the distribution of $\limfunc{diag}(w){\Greekmath 0118} ^{^{\bullet }}$ where ${\Greekmath 0118} ^{^{\bullet }}$ follows the distribution $\Xi ^{\bullet }$. We consider the following choices for the vector of weights $w$ :\footnote{ We use here the same conventions as mentioned in Footnote (ref).}

enumerate• Unrestricted HC0-HC4 weights $w=(w_{1},\ldots ,w_{n})$ : $w_{i}=1$ (HC0), $w_{i}=[n/(n-k)]^{1/2}$ (HC1), $w_{i}=(1-h_{ii})^{-1/2}$ (HC2), $w_{i}=(1-h_{ii})^{-1}$ (HC3), and $w_{i}=(1-h_{ii})^{{\Greekmath 010E} _{i}/2}$ for ${\Greekmath 010E} _{i}=\min (nh_{ii}/k,4)$ (HC4). • Null-restricted HC0R-HC4R weights $w=(\tilde{w}_{1},\ldots , \tilde{w}_{n})$: $\tilde{w}_{i}=1$ (HC0R), $\tilde{w} _{i}=[n/(n-(k-q))]^{1/2}$ (HC1R), $\tilde{w}_{i}=(1-\tilde{h}_{ii})^{-1/2}$ (HC2R), $\tilde{w}_{i}=(1-\tilde{h}_{ii})^{-1}$ (HC3R), and $\tilde{w} _{i}=(1-\tilde{h}_{ii})^{\tilde{{\Greekmath 010E}}_{i}/2}$ for $\tilde{{\Greekmath 010E}}_{i}=\min (n\tilde{h}_{ii}/(k-q),4)$ (HC4R).

In total this gives $960$ possible combinations of test statistics and bootstrap schemes. We emphasize that some of these combinations result in the same bootstrap-based test: (i) For reasons discussed in Lemma (ref), (ii) when changing HC0 weights to HC0R weights in the bootstrap scheme, and (iii) when changing HC0 (HC0R) weights to HC1 (HC1R) weights in the definition of $T_{Het}$ ($\tilde{T}_{Het}$), cf. Remark (ref). Concerning the run-time of the simulations, one could certainly argue that, for the computations, one should keep only one of the combinations that lead to the same bootstrap-based test. However, we have chosen not to, because we can then exploit the ensuing additional computations as a double-check for the methods that \textquotedblleft survive\textquotedblright\ the worst-case analysis (as the design matrices are generated separately for each of the $ 960$ combinations).

Given a test statistic, a bootstrap scheme, and a level of significance $ {\Greekmath 010B} $, the corresponding bootstrap-based test is throughout taken as the test that, observing $y$, rejects the null hypothesis, if the bootstrap p-value computed for $y$ is strictly smaller than ${\Greekmath 010B} $. Here, bootstrap p-value refers to the mass assigned by $\Xi $ to the points ${\Greekmath 0118} $ that give rise to elements in the bootstrap sample at which the test statistic is greater than or equal to the test statistic evaluated at $y$. To be precise, if, e.g., $T_{Het}$ is used as a test statistic, we define the bootstrap p-value as $\Xi ({\Greekmath 0118} :T_{Het}^{\blacktriangle ,\ast }(y,{\Greekmath 0118} )\geq T_{Het}(y)) $ if a bootstrap scheme of the form $y^{\ast }$ is used, and as $ \Xi ({\Greekmath 0118} :T_{Het}^{\blacktriangle ,\maltese }(y,{\Greekmath 0118} )\geq T_{Het}(y))$ if a bootstrap scheme of the form $y^{\maltese }$ is used; for the other test statistics $T_{uc}$, $\tilde{T}_{Het}$, and $\tilde{T}_{uc}$ we proceed similarly. Recall that $T_{Het}^{\blacktriangle }$ coincides with $T_{Het}$, except on the exceptional set $\mathsf{B}$, on which $T_{Het}^{ \blacktriangle }$ is set equal to $\infty $ (and a similar statement applies for the other test statistics). The reason for using $T_{Het}^{ \blacktriangle ,\ast }$ ($T_{Het}^{\blacktriangle ,\maltese }$, respectively) rather than $T_{Het}^{\ast }$ ($T_{Het}^{\maltese }$, respectively) in the definition of the p-value (and similarly for the other test statistics)\ is that this potentially gives a smaller rejection region, thus biasing the result in favor of the test (recall we are after negative results!); cf. the discussion relating to Part (b) of Theorem (ref) given subsequent to this theorem. For the same reason we use $\geq $, and not $>$, in the definition of the p-value. It is easy to see that -- in case a bootstrap scheme of the form $y^{\ast }$ is used -- the bootstrap-based test just defined via p-values can be rewritten as the test that rejects if $T_{Het}(y)>f_{Het,1-{\Greekmath 010B} }^{\blacktriangle ,upper}(y) $, where $f_{Het,1-{\Greekmath 010B} }^{\blacktriangle ,upper}(y)$ is the upper (i.e., largest) $(1-{\Greekmath 010B} )$-quantile of $F_{Het,y}^{ \blacktriangle }$; and a similar statement applies if a bootstrap scheme of the form $y^{\maltese }$ (or one of the other test statistics) is being used. [This also shows that the above defined rejection region is the smallest among all the rejection regions that can appear in the formulations of the theorems in Section (ref).]

Computations carried out in each setting

In each setting ($n=10,20,30$) and for each of the $960$ combinations of test statistics and bootstrap schemes described above we perform a two-step procedure. A detailed description of the computations carried out can be found in Sections (ref) and (ref) in Appendix (ref). Here we only provide a brief summary of the two-step procedure to the extent needed for an understanding of the results presented in Section (ref).

enumerate• The main goal of Step 1 is to find a scenario and a corresponding design matrix leading to a small value of ${\Greekmath 0123} $. Essentially, this is done by randomly generating $n\times k$ design matrices (with first column the intercept, and the remaining coordinates i.i.d. log-(standard) normally distributed) and by computing the corresponding values of ${\Greekmath 0123} $ for the testing problems $R=(0:I_{q})$ and $r=0$. In preparation for Step 2, for a suitably chosen subset of the design matrices generated, we also compute null rejection probabilities for strategically chosen variance parameters (assuming normality). All this is done for every pair $(k,q)$ with $k=2,\ldots ,5$ and $q=1,\ldots ,k-1$. • The goals of Step 2 are twofold: (a) to check the numerical reliability of the computation of ${\Greekmath 0123} $ in Step 1; and (b) to compute lower bounds on the size of the test, if necessary. We do the following for $ {\Greekmath 010B} \in \{0.05,0.1\}$: If ${\Greekmath 0123} _{\min }$, the overall smallest ${\Greekmath 0123} $ identified in Step 1, turns out to be smaller than ${\Greekmath 010B} $, we further check the numerical reliability of ${\Greekmath 0123} _{\min }$ by making use of the null rejection probabilities computed in Step 1. If this numerical check, described in Section (ref) in Appendix (ref), is not passed, we update ${\Greekmath 0123} _{\min }$. Once this check is passed, we distinguish two cases: (i) If ${\Greekmath 0123} _{\min }<{\Greekmath 010B} $ or if the maximum of the null rejection probabilities just referred to (maximized over the strategically chosen variance parameters) exceeds $3{\Greekmath 010B} $, we stop. For these cases we report the value of ${\Greekmath 0123} _{\min }$ together with the maximal rejection probability obtained for the design matrix pertaining to $ {\Greekmath 0123} _{\min }$. (ii) For the exceptional set of tests for which $ {\Greekmath 0123} _{\min }\geq {\Greekmath 010B} $ and the maximum of the null rejection probabilities does not exceed $3{\Greekmath 010B} $ we perform a second search (again sampling as in Step 1) to find design matrices leading to high rejection probabilities under the null. For these cases we report the highest null rejection probability found, and the value of ${\Greekmath 0123} $ corresponding to the design matrix that led to the highest null rejection probability.

Results and discussion

The results of the two-step procedure described in Section (ref) are summarized in Figure (ref) in the form of $6$ plots corresponding to the six combinations of the three settings A, B, and C, and of the two values for ${\Greekmath 010B} $ (${\Greekmath 010B} \in \{0.05,0.1\}$). In each plot, the vertical dashed line intersects the axis at ${\Greekmath 010B} $, the lower (upper, respectively) horizontal dashed line intersects the axis at ${\Greekmath 010B} $ (at $ 3{\Greekmath 010B} $, respectively).

For every combination of the setting and the value of ${\Greekmath 010B} $, the plot is obtained as follows: for every bootstrap-based test procedure (i.e., combination of test statistic and bootstrap scheme), a null rejection probability is plotted against a corresponding ${\Greekmath 0123} $ indicated by a black or red circle. The black circles correspond to test procedures for which Step 2 terminated without starting a second set of searches (which was the case for the vast majority of procedures, see Footnote (ref) in Appendix (ref)). The red circles correspond to the remaining (exceptional) test procedures for which further null rejection probabilities were computed in Step 2.

For all test procedures corresponding to black circles with ${\Greekmath 0123} <{\Greekmath 010B} $, the reliability check applied in Step 1 guarantees that the corresponding null rejection probability found in Step 1 is greater than $ 0.4 $. Note that these null rejection probabilities do not coincide with the sizes of the respective tests, which actually all are equal to $1$ by our theoretical results; they are only lower bounds for the size that are reported for completeness. For all black circles with ${\Greekmath 0123} \geq {\Greekmath 010B} $, the null rejection probabilities are not less than $3{\Greekmath 010B} $, which can be gathered from an inspection of the plots (and which is so by construction of the two-step procedure, see Section (ref) in Appendix (ref)); hence, they are much too large compared to the nominal significance level ${\Greekmath 010B} $. A bootstrap-based test procedure corresponding to a black circle hence \textquotedblleft fails the worst-case check\textquotedblright\ (in the setting and for the ${\Greekmath 010B} $ considered), because a scenario (i.e., $k$ and $q$) and a corresponding design matrix has been found for which the size of the test is $1$, or is exceedingly large, i.e., larger than or equal to $3{\Greekmath 010B} $.

For the test procedures corresponding to red circles extra computations were carried out in Step 2. A test procedure resulting in a red circle is declared to \textquotedblleft fail the worst-case check\textquotedblright\ (in the setting and for the ${\Greekmath 010B} $ considered) if the null rejection probability plotted is not less than $3{\Greekmath 010B} $ (as then a scenario (i.e., $k $ and $q$) and a corresponding design matrix have been found for which the size of the test is exceedingly large, namely larger than or equal to $ 3{\Greekmath 010B} $); otherwise it is declared to \textquotedblleft pass the worst-case check\textquotedblright\ (in the setting and for the ${\Greekmath 010B} $ considered) for the time being. [There are a few instances of red circles for which the null rejection probability plotted is less than $ 3{\Greekmath 010B} $, but ${\Greekmath 0123} <{\Greekmath 010B} $ holds. While our theoretical results then tell us that the size of the test should be equal to $1$, and we hence should classify the test procedure as "failing the worst-case check", we do not do so as we do not want to rely too much on the information provided by $ {\Greekmath 0123} $ in such cases, since no reliability check for computing $ {\Greekmath 0123} $ is included in the computation in Step 2.]

Bootstrap-based test procedures that fail the worst-case check (in at least one setting and for one of the values of ${\Greekmath 010B} $) should thus not be expected to be reliable in general. Therefore, such a test procedure should not be used in practice without first obtaining further guarantees concerning its size properties in the specific problem at hand, e.g., by running additional simulations geared towards the problem at hand.

Figure (ref) shows that in all settings and for both significance levels considered, the vast majority of circles are black (see Footnote (ref) in Appendix (ref)), already leading to the conclusion that most bootstrap-based test procedures fail the worst-case check. Furthermore, most of these test procedures fail in such a way that the corresponding ${\Greekmath 0123} $ is smaller than the significance level ${\Greekmath 010B} $, and thus the test is known to have size equal to $1$ (at least in one of the scenarios and for one of the design matrices considered) as a consequence of our theoretical results. Inspection of Figure (ref) also shows that a good portion of the test procedures corresponding to red circles fail the worst-case check in that the red circles are above or on the $3{\Greekmath 010B} $ -line. Comparing the figures across different settings and values of ${\Greekmath 010B} $, we also see that the number of red circles decreases when passing from $ {\Greekmath 010B} =0.05$ to ${\Greekmath 010B} =0.1$. The figure also shows that the number of red circles increases when increasing $n$. One reason could be that the randomized search for design matrices leading to low values of ${\Greekmath 0123} $ becomes more difficult as $n$ increases. We used the same randomized search algorithm in all three scenarios, which could explain the difference. A conceptual difference between the methods used in Setting A and Settings B and C is that for Setting A ($n=10$) exact computations (concerning $\Xi ^{\bullet }$) are carried out, while for Settings B ($n=20$) and C ($n=30$) approximate computations (based on empirical distributions $\Xi ^{\bullet }$ ) are done. This introduces an additional source of variation in the computations in Settings B and C, which could also be responsible for the increase in the number of red circles.

Important questions now are whether (i) there is a bootstrap-based test procedure left that passes the worst-case check in all settings and for both significance levels considered; (ii) there is a test procedure that passes the worst-case check in all settings considered for a fixed ${\Greekmath 010B} $; (iii) there is a pattern, in the sense that certain combinations of test statistics and bootstrap schemes often pass the worst-case check? To answer such questions, we shall next provide information on the procedures that pass the worst-case check in each setting and for each ${\Greekmath 010B} $ considered.

figure[figure omitted — 123 chars of source]

The bootstrap-based test procedures which (for the time being) pass the worst-case check in Setting A, B, and C, respectively, are summarized in Tables (ref), (ref), and (ref). In each table, the first row contains the test procedures that pass the worst-case check for ${\Greekmath 010B} =0.05$, the second row the ones that pass for ${\Greekmath 010B} =0.1$, and the third row the ones that pass at both nominal levels of significance.

In these tables, to facilitate the exposition, we use the following way of encoding a bootstrap-based test procedure: to each of the $960$ possible procedures (i.e., combinations of test statistics and bootstrap schemes) considered we associate a $7$ digit code: $x=x_{1}{:}x_{2}{:}x_{3}{:}x_{4}{:} x_{5}{:}x_{6}{:}x_{7}$. The encoding of the digits is as follows:

enumerate• indicates which covariance matrix estimator is used in the test statistic (\textquotedblleft -1" stands for the uncorrected estimator based on restricted or unrestricted residuals depending on whether $x_{2}$ is set to F or T; and $0,...,4$ stand for HC0,...,HC4, respectively, or for HC0R,...,HC4R, respectively, depending on whether $x_{2}$ is set to F or T). • indicates whether the covariance matrix estimator used in the test statistic is based on null-restricted residuals (T) or not (F). • indicates the distribution underlying $\Xi ^{\bullet }$ (\textquotedblleft r\textquotedblright\ stands for Rademacher, \textquotedblleft m\textquotedblright\ stands for Mammen). • indicates the weights $w$ used in constructing $\Xi $ from $\Xi ^{\bullet }$ ($0,...,4$ stands for HC0,...,HC4, respectively, or for HC0R,...,HC4R, respectively, depending on whether $x_{5}$ is set to F or T). • indicates whether the weights $w$ are null-restricted (T) or not (F). • indicates whether the bootstrap-scheme was based on null-restricted residuals (T) (i.e., $\mathcal{A}=\mathfrak{M}_{0}$) or not (F) (i.e. $\mathcal{A}=\limfunc{span}(X)$). • indicates whether the bootstrap scheme $y^{\ast }$ (T) or $y^{\maltese }$ (F) was used.

As an example, the code\textquotedblleft -1:T:m:2:F:T:F\textquotedblright\ translates to the bootstrap-based test procedure which uses the test statistic $\tilde{T}_{uc}$ (determined by the first two digits of the code) and the following bootstrap scheme: $\Xi ^{\bullet }$ based on the Mammen distribution, modified by HC2 weights (based on unrestricted residuals), and $y^{\maltese }$ with $\mathcal{A}=\mathfrak{M}_{0}$.

Before interpreting the results, we also need to explain why some test procedures are struck out in Tables (ref), (ref), and (ref) . Recall (e.g., from the discussion in the last but one paragraph of Section (ref)) that some of the $960$ procedures studied are in fact equivalent, meaning that they lead to exactly the same bootstrap-based test. Because the design matrices are generated anew for every of the $960$ procedures, the results found for equivalent procedures can be different. Therefore, it can happen that a procedure passes the worst-case check, but an equivalent version of this procedure does not. Now, a procedure is struck out in a given table, if -- while this procedure passed the worst-case check underlying the table -- an equivalent procedure did not (and thus does not appear in this table). Procedures that are struck out in a table are now no longer considered as having passed the worst-case check (although they appear in that table).

We exploit the following three reasons for equivalence between test procedures: (1) Lemma (ref) shows that a bootstrap-based test using a covariance matrix estimator based on unrestricted residuals does not depend on whether $y^{\ast }$ or $y^{\maltese }$ is used as a bootstrap scheme. Therefore, two procedures with codes $x$ and $x^{\prime }$ are equivalent in case $x_{2}=x_{2}^{\prime }=F$ and $x_{i}=x_{i}^{\prime }$ for $i=1,\ldots ,6 $. (2) Changing the weights vector from HC0 to HC0R (or vice versa) in the construction of $\Xi $ does not change the bootstrap-based test. Therefore, two procedures with codes $x$ and $x^{\prime }$ are equivalent in case $x_{4}=x_{4}^{\prime }=0$, and $x_{i}=x_{i}^{\prime }$ for all $i\neq 5$ . (3) Changing HC0 (HC0R) weights to HC1 (HC1R) weights in the definition of $T_{Het}$ ($\tilde{T}_{Het}$) also does not change the resulting bootstrap-based test, cf. Remark (ref). Therefore, two procedures with codes $x$ and $x^{\prime }$ are equivalent in case $x_{1}\in \{0,1\}$, $ x_{1}^{\prime }\in \{0,1\}$, and $x_{i}=x_{i}^{\prime }$ for all $i=2,\ldots ,7$. Procedures $x$ that are struck out by a slash, i.e., $\cancel{x}$, were eliminated based on reason (1); procedures that are struck out by a backslash, i.e., $\bcancel{x}$, were eliminated based on reason (2); and procedures that are struck out by a horizontal line, i.e., \sout{$x$}, were eliminated based on reason (3). Note that a procedure can be struck out, e.g., based on reasons (1) and (2), and thus is then crossed out, i.e., is marked by $\xcancel{x}$, etc.

table[table omitted — 915 chars of source]
table[table omitted — 1,536 chars of source]
table[table omitted — 1,898 chars of source]

Concerning the questions (i)-(iii) raised above, inspection of Tables (ref), (ref), and (ref) now delivers the following answers:

enumerate• There is no bootstrap-based test procedure that passes the worst-case check in all settings and for both significance levels. • The tests 3:T:m:1:T:T:T and 3:T:m:2:T:T:T pass the worst-case checks in all three settings for the significance level ${\Greekmath 010B} =0.05$. The corresponding rejection probabilities shown in Figure (ref) for the tests 3:T:m:1:T:T:T and 3:T:m:2:T:T:T are $0.077$ and $0.067$ (Setting A), $ 0.033$ and $0.080$ (Setting B), $0.050$ and $0.033$ (Setting C), respectively. For ${\Greekmath 010B} =0.1$ there is no test that passes the worst-case checks in all three settings. • The majority of test procedures appearing in the tables is based on the HC3 or HC3R covariance estimator, and uses a bootstrap scheme based on the Mammen-distribution. In Settings B and C, the tests passing the worst-case check are typically based on $\tilde{T}_{Het}$, i.e., they use restricted residuals in the construction of the covariance estimator.

On the one hand our results issue a distinct warning: overall none of the bootstrap-based tests considered comes with a guarantee that its size is (about) right. On the other hand, when restricting attention only to ${\Greekmath 010B} =0.05$, the tests 3:T:m:1:T:T:T and 3:T:m:2:T:T:T did not break down in our worst-case analysis.\footnote{ However, recall that we have only computed a lower bound for the size.} These two tests are based on $\tilde{T}_{Het}$ using a HC3R covariance estimator, and use the Mammen-distribution in the bootstrap scheme; properties that are common to many of the tests that pass the check (in one of the settings considered). While this obviously does not prove that 3:T:m:1:T:T:T and 3:T:m:2:T:T:T always will have perfect size properties for ${\Greekmath 010B} =0.05$, it shows that in the settings considered (and for ${\Greekmath 010B} =0.05$) they seem to have the best size performance among all bootstrap-based tests considered, and should therefore perhaps be preferred (over the other procedures) by practitioners who insist on applying a bootstrap-based test. However one should keep in mind that, while we have examined a considerable and reasonable range of scenarios and design matrices, the size-behavior of the bootstrap-based tests outside of the range studied can potentially be even worse.

In light of the findings above a better way forward seems to use heteroskedasticity robust test procedures that guarantee size-control as expounded in PP5HC.