EconBase
← Back to paper

Valid Heteroskedasticity Robust Testing

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

190,864 characters · 24 sections · 145 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Valid Heteroskedasticity Robust Testing

abstractTests based on heteroskedasticity robust standard errors are an important technique in econometric practice. Choosing the right critical value, however, is not simple at all: conventional critical values based on asymptotics often lead to severe size distortions; and so do existing adjustments including the bootstrap. To avoid these issues, we suggest to use smallest size-controlling critical values, the generic existence of which we prove in this article for the commonly used test statistics. Furthermore, sufficient and often also necessary conditions for their existence are given that are easy to check. Granted their existence, these critical values are the canonical choice: larger critical values result in unnecessary power loss, whereas smaller critical values lead to over-rejections under the null hypothesis, make spurious discoveries more likely, and thus are invalid. We suggest algorithms to numerically determine the proposed critical values and provide implementations in accompanying software. Finally, we numerically study the behavior of the proposed testing procedures, including their power properties.

Introduction

Testing hypotheses on the parameters in a regression model with potentially heteroskedastic errors is an important problem in econometrics and statistics; see Mackinnon2013 for a recent survey. Since the classical $t$-statistic ($F$-statistic, respectively) is not pivotal, or asymptotically pivotal, in such a case in general, even under Gaussianity of the errors, so-called heteroskedasticity robust (aka heteroskedasticity consistent) modifications of these test statistics have been proposed, which are asymptotically standard normally (chi-square, respectively) distributed under the null. These modifications date back to E63,E67, see also H77, and have later been popularized in econometrics by W80 with great success (see Mackinnon2013). Unfortunately, it turned out that tests obtained from these heteroskedasticity robust test statistics by relying on critical values derived from the respective asymptotic null distributions have a tendency to overreject the null hypothesis in finite samples (and thus are invalid), especially so if the design matrix contains high-leverage points; see, e.g., MacW85, DavidsonMacKinnon1985 , and CheshJewitt1987. One factor contributing to this overrejection tendency is a downward bias in the covariance matrix estimators used in these test statistics, see CheshJewitt1987. In an attempt to reduce the overrejection problem, variants of the before-mentioned heteroskedasticity robust test statistics (often denoted by HC1 through HC4, with HC0 denoting the original proposal) have been considered; see H77 , MacW85, and Crib2004.\footnote{ For a recent contribution geared towards high-dimensional models see Catt.} These variants rescale the least-squares residuals before computing the covariance matrix estimator employed in the construction of the test statistic. According to simulation studies reported in, e.g., DavidsonMacKinnon1985 and Crib2004, these modifications, especially HC3 and HC4, seem to ameliorate the overrejection problem to some extent, but do not eliminate it. Further numerical results are provided in CheshAust_1991, see also Chesh_1989. Numerical results in Section (ref) confirm these observations. Variants of HC0-HC3, denoted by HC0R-HC3R, obtained by using restricted instead of unrestricted least-squares residuals in the computation of the covariance matrix estimators employed by the various test statistics (the restriction alluded to being the restriction defining the null hypothesis) have been introduced in DavidsonMacKinnon1985. In their simulation experiments, this typically leads to tests that do not overreject, but that may substantially underreject; see also the simulation results in Godfrey2006, who additionally also considers HC4R. However, as will be shown in Section (ref), also these tests are in general not immune to (sometimes substantial) overrejection.

Note that, under the typical assumptions used in the literature, all the modifications of HC0 discussed so far have the same asymptotic distribution as HC0, and thus the same critical value as for HC0 (obtained from the asymptotic null distribution) is also used for these modifications in the before mentioned literature. Sometimes small-sample adjustments to the asymptotic critical values are attempted by using the quantiles from a $ t_{d} $-distribution rather than from the asymptotic normal distribution, where the degrees of freedom $d$ are either set to $n-k$ ($n$ and $k$ denoting sample size and number of regressors, respectively), or are obtained through proposals set down by Satterth or BellMcCa; see also Imbkoles2016. While these adjustments can lead to improvements, numerical results presented in Section (ref) show that these adjustments are also not able to solve the overrejection problem in general. An alternative approach is to use bootstrap methods to compute critical values for the test statistics HC0-HC4 or HC0R-HC4R. The relevant literature is reviewed in PPBoot, and it is shown that such methods are again not immune to the overrejection problem in general.\footnote{ Another possibility is to use Edgeworth expansions to find better critical values, see Rothenberg1988 for the case of the HC0 test statistic and DavidsonMacKinnon1985 for the HC0R test statistic. Simulation results in MacW85 and DavidsonMacKinnon1985 indicate that this does not work too well in practice. Of course, such expansions could also be worked out for the other versions of the test statistics mentioned, but this does not seem to have been pursued in the literature.} A referee has pointed out the recent papers by Chuetal2021 and Hansen2021, both of which propose a testing procedure that can be viewed as a parametric bootstrap method. No theoretical justification is given in those papers. In fact, as we show in Appendix (ref), the proposed procedures can be considerably oversized, a feature that can already be seen to some extent in the numerical results given in Chuetal2021 and Hansen2021.

A result by BakiSzek2005 needs to be mentioned here which states that -- in the special case of testing a hypothesis on the location parameter of a heteroskedastic location model with errors that are Gaussian or scale mixtures thereof -- the classical two-sided $t$-test (with the usual critical value) has null rejection probability not exceeding the nominal significance level under any form of heteroskedasticity (for a certain range of significance levels); see IbragMuell2010 for more discussion. IbragMuell2016, extending a result in MickeyBrown1966, provide a related result in the case of the comparison of two heteroskedastic populations; see also Bakirov98. Section (ref) provides some more discussion. We note that all results mentioned in that section are applicable only to testing certain scalar linear contrasts.

Except for the BakiSzek2005 result and the variations discussed in Section (ref), which apply only to quite special situations like, e.g., the heteroskedastic location model, none of the methods discussed so far comes with a theoretical result implying that their associated (finite sample) null rejection probabilities are guaranteed not to exceed the nominal significance level whatever the form of heteroskedasticity may be. \footnote{ In the special case where the number of restrictions tested equals the number of regression parameters,\ DavidsFlach2008 have a result which implies that certain wild bootstrap-based heteroskedasticity robust tests have size equal to the nominal significance level (and hence do not overreject) in finite samples. We note that this result in DavidsFlach2008 is not entirely correct as stated, but needs some amendments and corrections; see PPBoot.} In fact, it transpires from the preceding discussion and the numerical results in Section (ref) that for any of these methods instances of testing problems can be found for which the method in question overrejects substantially. Therefore, it is imperative to be able to find size-controlling critical values for the test statistics considered, i.e., critical values such that the resulting worst-case rejection probability under the null hypothesis does not exceed the nominal significance level. We shall hence pursue in this paper the construction of size-controlling critical values for the test statistics HC0-HC4, HC0R-HC4R, as well as for (two variants of) the classical (i.e., uncorrected) $F$-statistic (including the absolute value of the $t$ -statistic as a special case).

In the present paper we consider classes of test statistics that contain the before mentioned heteroskedasticity robust test statistics as special cases and show under which conditions -- and how -- a critical value can be found such that the resulting test is guaranteed to have size less than or equal to $\alpha $, the prescribed significance level.\footnote{ A less principled attempt at finding a valid test in a given testing problem (i.e., for given design matrix and restriction to be tested) could consist in the practitioner studying the size of a handful of tests (obtained from a few of the above mentioned test statistics in conjunction with a few of the proposed critical values) by means of an extensive Monte Carlo study and in hoping that one of the test procedures emerges from this study as valid for the particular testing problem at hand. Besides being a numerically costly procedure, it does not come with any guarantee of success. } It turns out that the conditions for size controllability are broadly satisfied; in particular, for the commonly used test statistics they are satisfied generically in a sense made precise further below.

We want to emphasize that the existence of size-controlling critical values for heteroskedasticity robust test statistics is not a trivial matter, as it has been shown in PP2016, Section 4, that there are cases where the size of such tests is always one, regardless of the choice of critical value; see also the discussion in Proposition (ref) further below. And even in cases where size control is possible by an appropriate choice of critical value, the standard critical values proposed in the literature (including the small-sample adjustments discussed above) are not guaranteed to deliver size control; in fact, they may fail to do so by a considerable margin (i.e., they are much too small to control size at the desired level) as shown in Section (ref). Our theoretical results also show the existence of a computable "threshold" $C^{\ast }$, say, such that any critical value $C$ satisfying $ C<C^{\ast }$ necessarily leads to a test with size $1$; see Proposition (ref). Since $C^{\ast }$ is not difficult to compute, it can be used as a simple check to weed out unsuitable proposals for critical values.

Apart from avoiding overrejection by construction, the use of smallest size-controlling, rather than conventional, critical values offers also advantages in terms of power in instances where conventional critical values lead to underrejection (i.e., lead to a worst-case rejection probability under the null hypothesis less than the nominal significance level) as is sometimes the case; see Sections (ref) and (ref). In fact, once one has decided on a test statistic to be used for the given null hypothesis, using the smallest size-controlling critical value (provided it exists) is obviously the optimal way to proceed.

We also discuss how the critical values that lead to size control can be determined numerically and provide the R-package hrt (hrt) for their computation. The usefulness of the proposed algorithms and their implementation in the R-package are illustrated numerically on some testing problems in Section (ref). In particular, we compare tests obtained from various of the above mentioned test statistics when used with smallest size-controlling critical values in terms of their power functions. The package hrt also contains a routine for determining the size of a test obtained from a user-supplied critical value. It is important to note that if in a particular application one uses the observed value of the test statistic as the user-supplied critical value in this routine, this routine actually returns a \textquotedblleft valid p-value\textquotedblright\ in the following sense: Checking whether or not this \textquotedblleft p-value\textquotedblright\ is smaller than the prescribed significance level $\alpha $ is equivalent to checking whether or not the observed value of the test statistic is larger than or equal to the smallest size-controlling critical value. Note that the former check avoids the need to actually compute the smallest size-controlling critical value, which is advantageous from a computational point of view. See Section (ref) for more details.

In the paper we work under a Gaussianity assumption. We stress, however, that this assumption is mainly made for convenience of presentation; as shown in Section (ref), this assumption can be relaxed considerably.

While a trivial remark, we would like to note that the size control results given in this paper can easily be translated into results stating that the minimal coverage probability of the associated confidence set obtained by \textquotedblleft inverting\textquotedblright\ the test is not less than the nominal confidence level.

The paper is organized as follows: After introducing notation and the most important test statistics in Sections (ref) and (ref) , Section (ref) provides some intuition for our size-control results which are presented in Sections (ref) and (ref), with some further results relegated to Appendix (ref). Section (ref) discusses ways of relaxing the underlying assumptions. Possible extensions to other classes of test statistics are discussed in Section (ref), while a few comments on power are collected in Section (ref). Section (ref) provides the numerical results including a power study, with some details relegated to Appendix (ref). Section (ref) concludes. Proofs and some technical results can be found in Appendices (ref)-(ref). The algorithms for computing rejection probabilities (including size) and smallest size-controlling critical values are outlined in Section (ref), and are presented in detail in Appendix (ref). Appendix (ref) contains a discussion of Chuetal2021 and Hansen2021.

Framework

Consider the linear regression model

equation[equation omitted — 58 chars of source]

where $X$ is a (real) nonstochastic regressor (design) matrix of dimension $ n\times k$ and where $\beta \in \mathbb{R}^{k}$ denotes the unknown regression parameter vector. We always assume $\limfunc{rank}(X)=k$ and $ 1\leq k<n$. We furthermore assume that the $n\times 1$ disturbance vector $ \mathbf{U}=(\mathbf{u}_{1},\ldots ,\mathbf{u}_{n})^{\prime }$ has mean zero and unknown covariance matrix $\sigma ^{2}\Sigma $, where $\Sigma $ varies in a user-specified (nonempty) set $\mathfrak{C}$ describing the allowed forms of heteroskedasticity, with $\mathfrak{C}$ satisfying $\mathfrak{C} \subseteq \mathfrak{C}_{Het}$, and where $0<\sigma ^{2}<\infty $ holds ($ \sigma $ always denoting the positive square root).\footnote{ Since we are concerned with finite-sample results only, the elements of $ \mathbf{Y}$, $X$, and $\mathbf{U}$ (and even the probability space supporting $\mathbf{Y}$ and $\mathbf{U}$) may depend on sample size $n$, but this will not be expressed in the notation. Furthermore, the obvious dependence of$\ \mathfrak{C}$ on $n$ will also not be shown in the notation.} The set $\mathfrak{C}$ will be referred to as the \textquotedblleft heteroskedasticity model\textquotedblright . Here

equation*[equation* omitted — 176 chars of source]

where $\limfunc{diag}(\tau _{1}^{2},\ldots ,\tau _{n}^{2})$ denotes the $ n\times n$ matrix with diagonal elements given by $\tau _{i}^{2}$. That is, the errors in the regression model are uncorrelated but can be heteroskedastic. In particular, if $\mathfrak{C}$ is chosen to be $\mathfrak{ C}_{Het}$, one allows for heteroskedasticity of completely unknown form. The normalization condition $\sum_{i=1}^{n}\tau _{i}^{2}=1$ is included here only in order to guarantee identifiability of $\sigma ^{2}$ and $\Sigma $, and could be replaced by any other normalization condition such as, e.g., $ \max \tau _{i}^{2}=1$, or $\tau _{1}^{2}=1$, without affecting the final results (because any of these normalizations leads to the same overall set of covariance matrices $\sigma ^{2}\Sigma $ when $\sigma ^{2}$ varies through the positive real line). Although a trivial observation, we stress the fact that all conceivable forms of heteroskedasticity, including parametric ones, can (possibly after normalization) be cast as submodels $ \mathfrak{C}$ of $\mathfrak{C}_{Het}$.

Mainly for ease of exposition, we shall maintain in the sequel that the disturbance vector $\mathbf{U}$\ is normally distributed. This assumption can be substantially relaxed as discussed in Section (ref). The linear model described in ((ref)), together with the just made Gaussianity assumption on $\mathbf{U}$ and with the given heteroskedasticity model $\mathfrak{C}$, then induces a collection of distributions on the Borel-sets of $\mathbb{R}^{n}$, the sample space of $ \mathbf{Y}$. Denoting a Gaussian probability measure with mean $\mu \in \mathbb{R}^{n}$ and (possibly singular) covariance matrix $A$ by $P_{\mu ,A}$ , the induced collection of distributions is then given by

equation[equation omitted — 156 chars of source]

where $\mathrm{\limfunc{span}}(X)$ denotes the column space of $X$. Since every $\Sigma \in \mathfrak{C}$ is positive definite by assumption, each element of the set in the previous display is absolutely continuous with respect to (w.r.t.) Lebesgue measure on $\mathbb{R}^{n}$.

We shall consider the problem of testing a linear (better: affine) hypothesis on the parameter vector $\beta \in \mathbb{R}^{k}$, i.e., the problem of testing the null $R\beta =r$ against the alternative $R\beta \neq r$, where $R$ is a $q\times k$ matrix always of rank $q\geq 1$ and $r\in \mathbb{R}^{q}$. Set $\mathfrak{M}=\limfunc{span}(X)$. Define the affine space

equation*[equation* omitted — 104 chars of source]

and let

equation*[equation* omitted — 110 chars of source]

Adopting these definitions, this testing problem can then be written more precisely as

equation[equation omitted — 226 chars of source]

With $\mathfrak{M}_{0}^{lin}$ we shall denote the linear space parallel to $ \mathfrak{M}_{0}$, i.e., $\mathfrak{M}_{0}^{lin}=\mathfrak{M}_{0}-\mu _{0}=\left\{ X\beta :R\beta =0\right\} $ where $\mu _{0}\in \mathfrak{M}_{0}$ . Of course, $\mathfrak{M}_{0}^{lin}$ does not depend on the choice of $\mu _{0}\in \mathfrak{M}_{0}$.

As already mentioned, the assumption of Gaussianity is made mainly for simplicity of presentation and can be relaxed substantially; see Section (ref). The assumption of nonstochastic regressors entails little loss of generality either, which is important to emphasize: If $X$ is random and $\mathbf{U}$ is conditionally on $X$ distributed as $N(0,\sigma ^{2}\Sigma )$, with $\sigma ^{2}=\sigma ^{2}(X)>0$ and $\Sigma =\Sigma (X)\in \mathfrak{C}_{Het}$, the results of the paper can be applied after one conditions on $X$ (and a similar statement applies to the generalizations to non-Gaussianity discussed in Section (ref) ). See Section (ref) for more discussion and details. For arguments supporting conditional inference see, e.g., RO1979. Note that such a \textquotedblleft strict exogeneity\textquotedblright\ assumption is quite natural in the situation considered here.

We next collect some further terminology and notation used throughout the paper. A (nonrandomized) test is the indicator function of a Borel-set $W$ in $\mathbb{R}^{n}$, with $W$ called the corresponding rejection region. The size of such a test (rejection region) is -- as usual -- defined as the supremum over all rejection probabilities under the null hypothesis $H_{0}$ given in ((ref)), i.e.,

equation*[equation* omitted — 137 chars of source]

In slight abuse of terminology, we shall sometimes refer to this quantity as `the size of $W$ over $\mathfrak{C}$' when we want to emphasize the r\^{o}le of $\mathfrak{C}$. Throughout the paper we let $\hat{\beta}(y)=\left( X^{\prime }X\right) ^{-1}X^{\prime }y$, where $X$ is the design matrix appearing in ((ref)) and $y\in \mathbb{R}^{n}$. The corresponding ordinary least-squares (OLS) residual vector is denoted by $\hat{u}(y)=y-X \hat{\beta}(y)$ and its elements are denoted by $\hat{u}_{t}(y)$. The elements of $X$ are denoted by $x_{ti}$, while $x_{t\cdot }$ and $x_{\cdot i} $ denote the $t$-th row and $i$-th column of $X$, respectively. For $ \mathcal{A}$ an affine subspace of $\mathbb{R}^{n}$ satisfying $\mathcal{A} \subseteq \limfunc{span}(X)$ let $\tilde{\beta}_{\mathcal{A}}(y)$ denote the restricted least-squares estimator, i.e., $X\tilde{\beta}_{\mathcal{A}}(y)$ solves

equation*[equation* omitted — 61 chars of source]

Lebesgue measure on the Borel-sets of $\mathbb{R}^{n}$ will be denoted by $ \lambda _{\mathbb{R}^{n}}$, whereas Lebesgue measure on an arbitrary affine subspace $\mathcal{A}$ of $\mathbb{R}^{n}$ (but viewed as a measure on the Borel-sets of $\mathbb{R}^{n}$) will be denoted by $\lambda _{\mathcal{A}}$, with zero-dimensional Lebesgue measure being interpreted as point mass. The set of real matrices of dimension $l\times m$ is denoted by $\mathbb{R} ^{l\times m}$ (all matrices in the paper will be real matrices) and Lebesgue measure on this set equipped with its Borel $\sigma $-field is denoted by $ \lambda _{\mathbb{R}^{l\times m}}$. Let $B^{\prime }$ denote the transpose of a matrix $B\in \mathbb{R}^{l\times m}$ and let $\mathrm{\limfunc{span}} (B) $ denote the subspace in $\mathbb{R}^{l}$ spanned by its columns. For a symmetric and nonnegative definite matrix $B$ we denote the unique symmetric and nonnegative definite square root by $B^{1/2}$. For a linear subspace $ \mathcal{L}$ of $\mathbb{R}^{n}$ we let $\mathcal{L}^{\bot }$ denote its orthogonal complement and we let $\Pi _{\mathcal{L}}$ denote the orthogonal projection onto $\mathcal{L}$. The Euclidean norm is denoted by $\left\Vert \cdot \right\Vert $, but the same symbol is also used to denote a norm of a matrix. The $j$-th standard basis vector in $\mathbb{R}^{n}$ is written as $ e_{j}(n)$. Furthermore, we let $\mathbb{N}$ denote the set of all positive integers. A sum (product, respectively) over an empty index set is to be interpreted as $0$ ($1$, respectively). Finally, for $\mathcal{A}$ an affine subspace of $\mathbb{R}^{n}$, let $G(\mathcal{A})$ denote the group of all affine transformations $y\mapsto \delta (y-a)+a^{\ast }$ where $\delta \in \mathbb{R}$, $\delta \neq 0$, and $a$ as well as $a^{\ast }$ are elements of $\mathcal{A}$; for more information see Section 5.1 of PP2016.

Heteroskedasticity robust test statistics using unrestricted residuals

We now introduce two test statistics that will feature prominently in the following. Variants thereof that use restricted residuals are discussed in Section (ref). For results pertaining to other classes of test statistics see Section (ref). The test statistic we shall consider first is a standard heteroskedasticity robust test statistic frequently encountered in the literature. It is given by

equation[equation omitted — 334 chars of source]

where $\hat{\Omega}_{Het}=R\hat{\Psi}_{Het}R^{\prime }$ and where $\hat{\Psi} _{Het}$ is a heteroskedasticity robust estimator as considered in E63,E67, which later on has found its way into the econometrics literature (e.g., W80). It is of the form

equation*[equation* omitted — 213 chars of source]

where the constants $d_{i}>0$ sometimes depend on the design matrix. Typical choices for $d_{i}$ suggested in the literature are $d_{i}=1$, $ d_{i}=n/(n-k) $, $d_{i}=\left( 1-h_{ii}\right) ^{-1}$, or $d_{i}=\left( 1-h_{ii}\right) ^{-2}$ where $h_{ii}$ denotes the $i$-th diagonal element of the projection matrix $X(X^{\prime }X)^{-1}X^{\prime }$, see LE2000 for an overview. Another suggestion is $d_{i}=\left( 1-h_{ii}\right) ^{-\delta _{i}}$ for $\delta _{i}=\min (nh_{ii}/k,4)$, see Crib2004. For the last three choices of $d_{i}$ just given, we use the convention that we set $d_{i}=1$ in case $h_{ii}=1$. Note that $h_{ii}=1$ implies $\hat{u} _{i}\left( y\right) =0$ for every $y$, and hence it is irrelevant which real value is assigned to $d_{i}$ in case $h_{ii}=1$.\footnote{ In fact, $h_{ii}=1$ is equivalent to $\hat{u}_{i}\left( y\right) =0$ for every $y$, each of which in turn is equivalent to $e_{i}(n)\in $ $\limfunc{ span}(X)$.} The five examples for the weights $d_{i}$ just given correspond to what is often called HC0-HC4 weights in the literature.

In conjunction with the test statistic $T_{Het}$, we shall consider the following mild assumption, which is Assumption 3 in PP2016. As discussed further below, this assumption is in a certain sense unavoidable when using $T_{Het}$. It furthermore also entails that our choice of assigning $T_{Het}\left( y\right) $ the value zero in case $\hat{\Omega} _{Het}\left( y\right) $ is singular has no import on the probabilistic results of the paper (because of Lemma (ref)(c) below and absolute continuity of the measures $P_{\mu ,\sigma ^{2}\Sigma }$).

assumptionLet $1\leq i_{1}<\ldots <i_{s}\leq n$ denote all the indices for which $e_{i_{j}}(n)\in \limfunc{span}(X)$ holds where $e_{j}(n)$ denotes the $j$-th standard basis vector in $\mathbb{R}^{n}$. If no such index exists, set $s=0$. Let $X^{\prime }\left( \lnot (i_{1},\ldots i_{s})\right) $ denote the matrix which is obtained from $X^{\prime }$ by deleting all columns with indices $i_{j}$, $1\leq i_{1}<\ldots <i_{s}\leq n$ (if $s=0$ no column is deleted). Then $\limfunc{rank}\left( R(X^{\prime }X)^{-1}X^{\prime }\left( \lnot (i_{1},\ldots i_{s})\right) \right) =q$ holds.

Observe that this assumption only depends on $X$ and $R$ and hence can be checked. Obviously, a simple sufficient condition for Assumption (ref) to hold is that $s=0$ (i.e., that $e_{j}(n)\notin \limfunc{span} (X) $ for all $j$), a generically satisfied condition. Furthermore, we introduce the matrix

eqnarray[eqnarray omitted — 327 chars of source]

The facts collected in the subsequent lemma, which is taken from PPBoot (but see also Lemma 4.1 in PP2016 and Lemma 5.18 in PP3), will be used in the sequel.

lemma(a) $\hat{\Omega}_{Het}\left( y\right) $ is nonnegative definite for every $y\in \mathbb{R}^{n}$. (b) $\hat{\Omega}_{Het}\left( y\right) $ is singular (zero, respectively) if and only if $\limfunc{rank}\left( B(y)\right) <q$ ($B(y)=0$, respectively). (c) The set $\mathsf{B}$ given by $\left\{ y\in \mathbb{R}^{n}:\limfunc{rank} \left( B(y)\right) <q\right\} $ (or in view of (b) equivalently given by $ \{y\in \mathbb{R}^{n}:\det (\hat{\Omega}_{Het}\left( y\right) )=0\}$) is either a $\lambda _{\mathbb{R}^{n}}$-null set or the entire sample space $ \mathbb{R}^{n}$. The latter occurs if and only if Assumption (ref) is violated (in which case the test based on $T_{Het}$ becomes trivial, as then $T_{Het}$ is identically zero). (d) Under Assumption (ref), the set $\mathsf{B}$ is a finite union of proper linear subspaces of $\mathbb{R}^{n}$; in case $q=1$, $\mathsf{B}$ is even a proper linear subspace itself.\footnote{ If Assumption (ref) is violated, $\mathsf{B}$ equals $\mathbb{R}^{n}$ by Part (c).} (e) $\mathsf{B}$ is a closed set and contains $\limfunc{span}(X)$. Furthermore, $\mathsf{B}$ is $G(\mathfrak{M})$-invariant and, in particular, $\mathsf{B}+\limfunc{span}(X)=\mathsf{B}$ holds.

In light of Part (c) of the lemma, we see that Assumption (ref) is a natural and unavoidable condition if one wants to obtain a sensible test from $T_{Het}$.\footnote{ If this assumption is violated then $T_{Het}$ is identically zero, an uninteresting trivial case.} Furthermore, note that, if $\mathsf{B}=\limfunc{ span}(X)$ is true, then Assumption (ref) must be satisfied (since $ \limfunc{span}(X)$ is a $\lambda _{\mathbb{R}^{n}}$-null set due to the maintained assumption $k<n$). As shown in Lemma A.3 in PP3, for any given restriction matrix $R$, the relation $\mathsf{B}=\limfunc{span}(X)$ holds generically in various universes of design matrices. For later use we also mention that under Assumption (ref) the test statistic $T_{Het}$ is continuous at every $y\in \mathbb{R}^{n}\backslash \mathsf{B}$.\footnote{ If Assumption (ref) is violated, then $T_{Het}$ is constant equal to zero, and hence is trivially continuous everywhere.}

Next, we also consider the classical (i.e., uncorrected) F-test statistic, i.e.,

equation[equation omitted — 328 chars of source]

where $\hat{\sigma}^{2}(y)=\hat{u}\left( y\right) ^{\prime }\hat{u}\left( y\right) /(n-k)\geq 0$ (which vanishes if and only if $y\in \limfunc{span}(X) $). Our choice to set $T_{uc}(y)=0$ for $y\in \limfunc{span}(X)$ again has no import on the probabilistic results in the paper, since $\limfunc{span}(X) $ is a $\lambda _{\mathbb{R}^{n}}$-null set as a consequence of the maintained assumption that $k<n$ (and since the measures $P_{\mu ,\sigma ^{2}\Sigma }$ are absolutely continuous). For reasons of comparability with ( (ref)) we have chosen not to normalize the numerator in ((ref) ) by $q$, the number of restrictions to be tested, as is often done in the definition of the classical F-test statistic. This also has no import on the results as the factor $1/q$ can be absorbed into the critical value. For later use we also mention that the test statistic $T_{uc}$ is continuous at every $y\in \mathbb{R}^{n}\backslash \limfunc{span}(X)$.

remark(i) The test statistics $T_{Het}$ as well as $T_{uc}$ are $G( \mathfrak{M}_{0})$-invariant as is easily seen (with the respective exceptional sets $\mathsf{B}$ and $\limfunc{span}(X)$ being $G(\mathfrak{M})$ -invariant). (ii) Both statistics actually belong to the class of nonsphericity-corrected F-type test statistics in the sense of Section 5.4 in PP2016 (terminology being somewhat unfortunate in case of $T_{uc}$ as no correction for the non-sphericity is applied in this case). See Remark (ref) in Appendix (ref) for more discussion.
remarkFor later use we note the following: Suppose $(R,r)$ and $( \bar{R},\bar{r})$ are both of dimension $q\times (k+1)$ and have $\limfunc{ rank}(R)=\limfunc{rank}(\bar{R})=q.$ (i) Then $(R,r)$ and $(\bar{R},\bar{r})$ give rise to the same set $\mathfrak{M}_{0}$, and thus to the same testing problem ((ref)), if and only if $(AR,Ar)=(\bar{R},\bar{r})$ holds for a nonsingular $q\times q$ matrix $A$. (ii) The test statistics $ T_{Het}$ and $T_{uc}$ remain the same whether they are computed using $(R,r)$ or $(\bar{R},\bar{r})$ provided $(AR,Ar)=(\bar{R},\bar{r})$ holds for a nonsingular $q\times q$ matrix $A$. [To see this note that the respective exceptional sets $\mathsf{B}$ and $\limfunc{span}(X)$ are the same irrespective of whether $(R,r)$ or $(\bar{R},\bar{r})$ is used, and that $A$ cancels out in the respective quadratic forms appearing in the definitions of the test statistics.]

Some intuition on why conventional critical values can lead to overrejection

We begin the heuristic discussion by considering the testing problem ((ref)) with heteroskedasticity model $\mathfrak{C}=\mathfrak{C} _{Het}$ (i.e., heteroskedasticity of unknown form). Let $T$ stand for any of the test statistics introduced in Section (ref), with rejection occurring whenever $T\geq C$, $C$ a critical value.\footnote{ In case of $T=T_{Het}$ Assumption (ref) is supposed to hold.}$^{ \text{,}}$\footnote{ The discussion similarly applies to the test statistics introduced in Section (ref).}. For simplicity of presentation we assume $r=0$. As discussed in Section (ref), basing the test on the conventional critical value $C_{\chi ^{2}(q),0.05}$ (the 95% quantile of a chi-square distribution with $q$ degrees of freedom) often leads to substantial overrejection, i.e., the size of the test (over $\mathfrak{C} _{Het}$) is substantially larger than the desired value $\alpha =0.05$. One mechanism leading to such overrejection is constituted by a concentration phenomenon discussed at some length in PP2016: In the present situation, the distribution $P_{0,\sigma ^{2}\Sigma }$ \textquotedblleft concentrates\textquotedblright\ on a so-called concentration subspace (given by $\limfunc{span}(e_{i}(n))$) when $\Sigma $ is \textquotedblleft close\textquotedblright\ to one of the singular matrices $ e_{i}(n)e_{i}(n)^{\prime }$.\footnote{ There are also other concentration subspaces in the present situation which we can ignore for the heuristic discussion.} In such a case, depending on the design matrix $X$ and the hypothesis given by $(R,r)$, the concentration space may fall into the rejection region $\{T\geq C_{\chi ^{2}(q),0.05}\}$, leading to a rejection probability close to one, and thus much larger than $ \alpha =0.05$.\footnote{ This is an oversimplified description ignoring some technical details.} Even if the concentration subspace $\limfunc{span}(e_{i}(n))$ is not contained in the rejection region, but is sufficiently close to it, a considerable portion of the mass of $P_{0,\sigma ^{2}\Sigma }$ may nevertheless fall into the rejection region if $\Sigma $ is close to, but not too close to $ e_{i}(n)e_{i}(n)^{\prime }$. This again leads to a relatively large rejection probability. Overrejection will often be especially pronounced if certain high-leverage points are present in the design matrix.\footnote{ We note, however, that there are testing problems (e.g., testing the mean in a heteroskedastic location model using the test statistic $T_{uc}$) for which the text-book critical values obtained under homoskedasticity are actually valid, see BakiSzek2005. The reason is that the \textquotedblleft worst case\textquotedblright\ distribution in this case corresponds to homoskedasticity.}

In order to obtain a test that has size controlled by $\alpha $ (i.e., size $ \leq \alpha $) in situations as just described, the rejection region $ \{T\geq C_{\chi ^{2}(q),0.05}\}$ has to be narrowed down, i.e., $C_{\chi ^{2}(q),0.05}$ has to be replaced by a suitably larger critical value $C$. Whether or not this can successfully be accomplished by a (finite) $C$, is a non-trivial question, the answer depending on whether or not all possible concentration subspaces can be made to fall outside of the rejection region $ \{T\geq C\}$ by an appropriate choice of $C$ larger than $C_{\chi ^{2}(q),0.05}$. Sufficient conditions when this is possible are provided in Theorems (ref) and (ref). Note that, in such a situation, the resulting size-controlling critical values $C$ are then necessarily larger than $C_{\chi ^{2}(q),0.05}$.

In light of the preceding discussion, a natural question is whether or not imposing a heteroskedasticity model more narrow than $\mathfrak{C}_{Het}$ such as, e.g.,

equation*[equation* omitted — 207 chars of source]

where $\tau _{\ast }$, $0<\tau _{\ast }<n^{-1/2}$, is a pre-specified constant set by the user, would mitigate the failure of conventional critical values. Indeed, under the heteroskedasticity model $\mathfrak{C} _{Het,\tau _{\ast }}$ extreme concentration effects leading to rejection probabilities (arbitrarily) close to one cannot occur, and it is possible to prove that size-controlling critical values always exist when $\mathfrak{C} _{Het,\tau _{\ast }}$ is used, see Appendix (ref). Unfortunately, however, this does not imply that conventional critical values such as $C_{\chi ^{2}(q),0.05}$ will work. In fact, the size over $\mathfrak{C} _{Het,\tau _{\ast }}$ of tests using the critical value $C_{\chi ^{2}(q),0.05}$ can still be considerably larger than $\alpha $: To see this, observe that the sets $\mathfrak{C}_{Het,\tau _{\ast }}$are an increasing sequence of sets as $\tau _{\ast }\downarrow 0$, the union of which is $ \mathfrak{C}_{Het}$. Consequently, if $\tau _{\ast }$ is small, the size over $\mathfrak{C}_{Het,\tau _{\ast }}$ will be close to the size over $ \mathfrak{C}_{Het}$, and thus the former will be much larger than $\alpha $ in case the latter is so. As a consequence, also in case of the more narrow heteroskedasticity model $\mathfrak{C}_{Het,\tau _{\ast }}$ size-controlling critical values larger than $C_{\chi ^{2}(q),0.05}$ will have to be used in such a case. Furthermore, the bound $\tau _{\ast }$ has to be decided upon prior to the data analysis and is thus part of modeling the form of heteroskedasticity. It is difficult to see how one would come up with a reasonable value of $\tau _{\ast }$ in practice: If $\tau _{\ast }$ is chosen to be small, this may result in a heteroskedasticity model under which the test based on $C_{\chi ^{2}(q),0.05}$ is still plagued by overrejection as just discussed, while choosing $\tau _{\ast }$ large will typically not be defensible as it presumes considerable knowledge about the admissible forms of heteroskedasticity.

Size control results for $T_{Het}$ and $T_{uc}$ when $\mathfrak{C}= \mathfrak{C}_{Het}$

We introduce the following notation: For a given linear subspace $\mathcal{L} $ of $\mathbb{R}^{n}$ we define the set of indices $I_{0}(\mathcal{L})$ via

equation*[equation* omitted — 93 chars of source]

We set $I_{1}(\mathcal{L})=\left\{ 1,\ldots ,n\right\} \backslash I_{0}( \mathcal{L})$. Clearly, $\func{card}(I_{0}(\mathcal{L}))\leq \dim (\mathcal{L })$ holds. In particular, if $\dim (\mathcal{L})<n$ holds (which, in particular, is so in the leading case $\mathcal{L}=\mathfrak{M}_{0}^{lin}$, since $\dim (\mathfrak{M}_{0}^{lin})=k-q<n$), then $\func{card}(I_{0}( \mathcal{L}))<n$, and thus $\func{card}(I_{1}(\mathcal{L}))\geq 1$.

We have the following size control result for $T_{uc}$ as well as for $ T_{Het}$ over the heteroskedasticity model $\mathfrak{C}_{Het}$ (more precisely, over the null hypothesis $H_{0}$ described in ((ref)) with $\mathfrak{C}=\mathfrak{C}_{Het}$). Note that $\mathfrak{C} _{Het}$ is the largest possible heteroskedasticity model and reflects complete ignorance about the form of heteroskedasticity.

theorem(a) For every $0<\alpha <1$ there exists a real number $C(\alpha )$ such that \begin{equation} \sup_{\mu _{0}\in \mathfrak{M}_{0}}\sup_{0<\sigma ^{2}<\infty }\sup_{\Sigma \in \mathfrak{C}_{Het}}P_{\mu _{0},\sigma ^{2}\Sigma }(T_{uc}\geq C(\alpha ))\leq \alpha \end{equation} holds, provided that \begin{equation} e_{i}(n)\notin \func{span}(X) \ \ for every \ i\in I_{1}(\mathfrak{M} _{0}^{lin}). \end{equation} Furthermore, under condition ((ref)), even equality can be achieved in ((ref)) by a proper choice of $ C(\alpha )$, provided $\alpha \in (0,\alpha ^{\ast }]\cap (0,1)$ holds, where $\alpha ^{\ast }=\sup_{C\in (C^{\ast },\infty )}\sup_{\Sigma \in \mathfrak{C}_{Het}}P_{\mu _{0},\Sigma }(T_{uc}\geq C)$ is positive and where $C^{\ast }=\max \{T_{uc}(\mu _{0}+e_{i}(n)):i\in I_{1}(\mathfrak{M} _{0}^{lin})\}$ for $\mu _{0}\in \mathfrak{M}_{0}$ (with neither $\alpha ^{\ast }$ nor $C^{\ast }$ depending on the choice of $\mu _{0}\in \mathfrak{M }_{0}$). (b) Suppose Assumption (ref) is satisfied.\footnote{ Condition ((ref)) clearly implies that the set $\mathsf{B}$ is a proper subset of $\mathbb{R}^{n}$ (as $\limfunc{card}(I_{1}(\mathfrak{M} _{0}^{lin}))\geq 1$) and thus implies Assumption (ref). Hence, we could have dropped this assumption from the formulation of the theorem. For clarity of presentation we have, however, chosen to explicitly mention Assumption (ref). A similar remark applies to some of the other results given below and will not be repeated.} Then for every $0<\alpha <1$ there exists a real number $C(\alpha )$ such that \begin{equation} \sup_{\mu _{0}\in \mathfrak{M}_{0}}\sup_{0<\sigma ^{2}<\infty }\sup_{\Sigma \in \mathfrak{C}_{Het}}P_{\mu _{0},\sigma ^{2}\Sigma }(T_{Het}\geq C(\alpha ))\leq \alpha \end{equation} holds, provided that \begin{equation} e_{i}(n)\notin \mathsf{B} \ \ for every \ i\in I_{1}(\mathfrak{M} _{0}^{lin}). \end{equation} Furthermore, under condition ((ref)), even equality can be achieved in ((ref)) by a proper choice of $C(\alpha )$, provided $\alpha \in (0,\alpha ^{\ast }]\cap (0,1)$ holds, where now $\alpha ^{\ast }=\sup_{C\in (C^{\ast },\infty )}\sup_{\Sigma \in \mathfrak{C} _{Het}}P_{\mu _{0},\Sigma }(T_{Het}\geq C)$ is positive and where $C^{\ast }=\max \{T_{Het}(\mu _{0}+e_{i}(n)):i\in I_{1}(\mathfrak{M}_{0}^{lin})\}$ for $\mu _{0}\in \mathfrak{M}_{0}$ (with neither $\alpha ^{\ast }$ nor $ C^{\ast }$ depending on the choice of $\mu _{0}\in \mathfrak{M}_{0}$). (c) Under the assumptions of Part (a) (Part (b), respectively) implying existence of a critical value $C(\alpha )$ satisfying ((ref)) (((ref)), respectively), a smallest critical value, denoted by $C_{\Diamond }(\alpha )$, satisfying ( (ref)) (((ref)), respectively) exists for every $0<\alpha <1$. And $C_{\Diamond }(\alpha )$ corresponding to Part (a) (Part (b), respectively) is also the smallest among the critical values leading to equality in ((ref)) (((ref)), respectively) whenever such critical values exist. [Although $C_{\Diamond }(\alpha )$ corresponding to Part (a) and (b), respectively, will typically be different, we use the same symbol.]\footnote{ Cf. also Appendix (ref).}

We see from the theorem that the condition for size control of $T_{Het}$ ($ T_{uc}$, respectively) over $\mathfrak{C}_{Het}$, i.e., condition ((ref)) (((ref)), respectively), only depends on $X$ and $R$; in particular, in case of $T_{Het}$, it does not depend on how the weights $d_{i}$ figuring in the definition of $T_{Het}$ have been chosen (note that the set $\mathsf{B}$ only depends on $X$ and $R$). Moreover, the sufficient conditions for size control are generically satisfied in the universe of all $n\times k$ design matrices $X$ (of rank $k$), see Example (ref) and the attending discussion further below. Furthermore, it is plain that the size-controlling critical values $C(\alpha )$ in Theorem (ref) will depend on the choice of test statistic as well as on the testing problem at hand. More concretely, the size-controlling critical values in Part (b) of the theorem thus depend only on $X$, $R$, and $r$, as well as on the choice of weights $d_{i}$, whereas in Part (a) the dependence is only on $X$, $R$, and $r$. We do not show these dependencies in the notation. In fact, as discussed in Remark (ref) below, it turns out that the size-controlling critical values in both cases actually do not depend on the value of $r$ at all (provided the weights $d_{i}$ are not allowed to depend on $r$ in case of $T_{Het}$). Similarly, it is easy to see that $C^{\ast }$ and $\alpha ^{\ast }$ in Theorem (ref) do not depend on $r$ (under the same provision as before in case of $T_{Het}$).

Another observation is that any critical value delivering size control over $ \mathfrak{C}_{Het}$ also delivers size control over any other heteroskedasticity model $\mathfrak{C}$ since $\mathfrak{C}\subseteq \mathfrak{C}_{Het}$. Of course, for such a $\mathfrak{C}$ even smaller critical values (than needed for $\mathfrak{C}_{Het}$) may already suffice for size control. Also note that sufficient conditions implying size control over $\mathfrak{C}_{Het}$ may be more restrictive than sufficient conditions implying only size control over a smaller heteroskedasticity model $ \mathfrak{C}$. For size control results tailored to such smaller models $ \mathfrak{C}$ see Appendix (ref).

In light of the results of CheshJewitt1987 and Chesh_1989, it is useful to interpret the sufficient conditions for size control, i.e., ( (ref)) and ((ref)), in terms of high-leverage points. First, note that $e_{i}(n)\in \limfunc{span}(X)$ is equivalent to $h_{ii}=1$, which corresponds to the $i$-th observation being an \textquotedblleft extreme high-leverage point\textquotedblright . Hence, ( (ref)) is equivalent to $h_{ii}<1$ for every $i\in \mathcal{I}_{1}(\mathfrak{M}_{0}^{lin})$. In other words, the condition for a size-controlling critical value to exist in Part (a) of Theorem (ref) requires that none of the indices in $\mathcal{I}_{1}( \mathfrak{M}_{0}^{lin})$ corresponds to an extreme high-leverage point. [It is interesting to observe that all indices in $\mathcal{I}_{0}(\mathfrak{M} _{0}^{lin})$ (note that this set may be empty) correspond to extreme high-leverage points.] Hence, for the condition in ((ref) ) not to be satisfied, not only must extreme high-leverage points be present, but the lever needs to be of a particular type depending on the hypothesis given by $(R,r)$ (namely, it must have $i\in \mathcal{I}_{1}( \mathfrak{M}_{0}^{lin})$). Second, note that a sufficient, but not necessary, condition for ((ref)) is $h_{ii}<1$ for $ i=1,\ldots ,n$. Sufficiency is obvious from the preceding discussion. That the condition is not necessary can be seen from Example (ref) further below. Finally, condition ((ref)) implies condition ((ref)) (since $\limfunc{span}(X)\subseteq \mathsf{B}$), and hence implies $h_{ii}<1$ for every $i\in \mathcal{I}_{1}(\mathfrak{M} _{0}^{lin})$. The converse is not always true: even $h_{ii}<1$ for every $ i=1,\ldots ,n$ does not guarantee ((ref)) to be satisfied, see Example (ref) further below. However, generically ((ref)) and ((ref)) coincide (see Lemma A.3 in PP3), in which case the discussion given above for ((ref)) also applies to ((ref)).

remark(Independence of the value of $r$ and implications for confidence sets) (i) As already noted before, the sufficient conditions for size control in both parts of Theorem (ref) only depend on $X$ and $R$. In particular, they do not depend on the value of $r$. (ii) The size of the test based on $T_{uc}$ ($T_{Het}$, respectively) in Theorem (ref) as well as the size-controlling critical values $ C(\alpha )$ (for both test statistics) do also not depend on the value of $r$ (provided the weights $d_{i}$ are not allowed to depend on $r$ in case of $ T_{Het}$). This follows from Lemma 5.15 in PP3 combined with Remark (ref) in Appendix (ref).\footnote{ For this argument we impose Assumption (ref) in case of $T_{Het}$, the case where this assumption is violated being trivial.} This observation is of some importance, as it allows one easily to obtain confidence sets for $R\beta $ by \textquotedblleft inverting\textquotedblright\ the test without the need of recomputing the critical value for every value of $r$.
remark(Some equivalencies) If the respective smallest size-controlling critical values are used (provided they exist), the tests obtained from $T_{Het}$ with the HC0 and the HC1 weights, respectively, are identical, as these two test statistics differ only by a multiplicative constant. The same reasoning applies to the test statistics based on the HC0-HC4 weights, respectively, in case $h_{ii}$ does not depend on $i$.
remark(Positivity of size-controlling critical values) For every $0<\alpha <1$ any $C(\alpha )$ satisfying ((ref)) or ((ref)) is necessarily positive. To see this observe that $\{T_{uc}\geq C\}=\{T_{Het}\geq C\}= \mathbb{R}^{n}$ for $C\leq 0$, since both test statistics are nonnegative everywhere.

The next proposition complements Theorem (ref) and provides a useful lower bound for the size-controlling critical values (other than the trivial bound given in the preceding remark).

proposition\footnote{ It is not difficult to show in the context of Parts (a) and (b) of the proposition that any critical value $C>C^{\ast }$ actually leads to size less than $1$. This follows from a reasoning similar as in Remark 5.4 of PP3.}$^{\text{,}}$\footnote{ If ((ref)) in Part (b) of the proposition does not hold, the conclusion of Part (b) can be shown to continue to hold with $C^{\ast }$ as defined in Theorem (ref)(b), and also with $C^{\ast }$ as defined in Lemma 5.11 of \ PP3 (note that under the assumptions of Part (b) of the proposition both definitions of $C^{\ast }$ actually coincide as shown in the proof of Theorem (ref)). [Recall that under violation of ((ref)) size-controlling critical values may or may not exist.] If Assumption (ref) is not satisfied, then $ T_{Het}\equiv 0$, and the conclusion of Part (b) holds trivially (as $ C^{\ast }=0$ with both definitions). If ((ref)) in Part (a) of the proposition is not satisfied, then no size-controlling critical value exists by Proposition (ref); hence, the conclusion of Part (a) holds trivially, again regardless of which of the two definitions of $C^{\ast }$ is adopted.}(a) Suppose that ((ref)) is satisfied. Then any $C(\alpha )$ satisfying ((ref)) necessarily has to satisfy $C(\alpha )\geq C^{\ast }$, where $C^{\ast }$ is as in Part (a) of Theorem (ref). In fact, for any $C<C^{\ast }$ we have $\sup_{\Sigma \in \mathfrak{C}_{Het}}P_{\mu _{0},\sigma ^{2}\Sigma }(T_{uc}\geq C)=1$ for every $\mu _{0}\in \mathfrak{M}_{0}$ and every $ \sigma ^{2}\in (0,\infty )$. (b) Suppose that Assumption (ref) and ((ref)) are satisfied. Then any $C(\alpha )$ satisfying ((ref)) necessarily has to satisfy $C(\alpha )\geq C^{\ast }$, where $C^{\ast }$ is as in Part (b) of Theorem (ref). In fact, for any $C<C^{\ast }$ we have $\sup_{\Sigma \in \mathfrak{C}_{Het}}P_{\mu _{0},\sigma ^{2}\Sigma }(T_{Het}\geq C)=1$ for every $\mu _{0}\in \mathfrak{M}_{0}$ and every $ \sigma ^{2}\in (0,\infty )$.

The preceding observation is useful in two ways: First, critical values suggested in the literature (such as, e.g., the $(1-\alpha )$-quantile of a chi-square distribution with $q$ degrees of freedom or critical values obtained from a degree of freedom adjustment) can immediately be dismissed if they turn out to be less than $C^{\ast }$, as they then certainly will not guarantee size control.\footnote{ In contrast, if the critical value turns out to be larger than or equal to $ C^{\ast }$, it does not follow that size is less than or equal to $ \alpha $. In fact, substantially oversized tests using a critical value $ C>C^{\ast }$ are certainly possible; see, e.g., Table (ref) and the pertaining discussion.} We use this line of reasoning in the numerical results in Section (ref). Second, if the observed value of the test statistic $T_{Het}$ ($T_{uc}$, respectively) is less than $C^{\ast } $, the decision not to reject the null hypothesis can be taken without further need to compute size-controlling critical values. Note that $C^{\ast }$ as given in Theorem (ref) is quite easy to compute in any given application.

remarkSuppose the assumptions of Part (a) (Part (b), respectively) of Theorem (ref) are satisfied. Then we know from that theorem that the size (over $\mathfrak{C}_{Het})$ of $\{T_{uc}\geq C_{\Diamond }(\alpha )\}$ ($\{T_{Het}\geq C_{\Diamond }(\alpha )\}$, respectively) equals $\alpha $ provided $\alpha \in (0,\alpha ^{\ast }]\cap (0,1)$. If now $\alpha ^{\ast }<\alpha <1$, then the size (over $\mathfrak{C} _{Het})$ of $\{T_{uc}\geq C_{\Diamond }(\alpha )\}$ ($\{T_{Het}\geq C_{\Diamond }(\alpha )\}$, respectively) equals $\alpha ^{\ast }$ (where the $C_{\Diamond }(\alpha )$'s pertaining to Parts (a) and (b) may be different). This follows from $C_{\Diamond }(\alpha )\geq C^{\ast }$ (see Proposition (ref) above) and Remark 5.13(i) in PP3.\footnote{ The assumptions for Part A of Proposition 5.12 in PP3 required in Remark 5.13 of that paper are satisfied under the assumptions of Theorem (ref) as shown in the proof of Theorem (ref) in Appendix (ref). In this proof also the condition $\lambda _{\mathbb{R}^{n}}(T_{uc}=C^{\ast })=0$ ($\lambda _{ \mathbb{R}^{n}}(T_{Het}=C^{\ast })=0$, respectively) required in Remark 5.13 of PP3 is verified.} This argument actually also delivers that $ C_{\Diamond }(\alpha )=C^{\ast }$ must hold in case $\alpha ^{\ast }<\alpha <1$.

We next discuss to what extent the sufficient conditions for size control in Theorem (ref) are also necessary.

proposition(a) If ((ref)) is violated, then $ \sup_{\Sigma \in \mathfrak{C}_{Het}}P_{\mu _{0},\sigma ^{2}\Sigma }(T_{uc}\geq C)=1$ for every choice of critical value $C$, every $\mu _{0}\in \mathfrak{M}_{0}$, and every $\sigma ^{2}\in (0,\infty )$ (implying that size equals $1$ for every $C$). As a consequence, the sufficient condition for size control ((ref)) in Part (a) of Theorem (ref) is also necessary. (b) Suppose Assumption (ref) is satisfied.\footnote{ If this assumption is violated then $T_{Het}$ is identically zero, an uninteresting trivial case.} If ((ref)) is violated, then $\sup_{\Sigma \in \mathfrak{C}_{Het}}P_{\mu _{0},\sigma ^{2}\Sigma }(T_{Het}\geq C)=1$ for every choice of critical value $C$, every $ \mu _{0}\in \mathfrak{M}_{0}$, and every $\sigma ^{2}\in (0,\infty )$ (implying that size equals $1$ for every $C$). [In case $X$ and $R$ are such that $\mathsf{B}=\limfunc{span}(X)$, conditions ((ref)) and ((ref)) coincide; hence the sufficient condition for size control ((ref)) in Part (b) of Theorem (ref) is then also necessary in this case.]
remarkSuppose Assumption (ref) is satisfied. In case $\mathsf{B}\neq \limfunc{span}(X)$ and ((ref)) hold, but ((ref)) is violated, neither Part (b) of Theorem (ref) nor Part (b) of Proposition (ref) apply. We note that there are instances of this situation (see Example (ref)) for which it can be shown by other methods that $T_{Het}$ is size controllable despite failure of ((ref));\footnote{ In this example actually $e_{i}(n)\in \mathsf{B}$ holds for all $i=1,\ldots ,n$.} as a consequence, ((ref)) is not necessary for ((ref)) in general. We conjecture that there are other instances of the situation described here where size control is not possible, but we have not investigated this in any detail. [What can be said in general in this situation is that the size of the rejection region $\left\{ T_{Het}\geq C\right\} $ over $\mathfrak{C}_{Het}$ is certainly equal to $1$ for every $ C<\max \left\{ T_{Het}(\mu _{0}+e_{i}(n)):e_{i}(n)\notin \mathsf{B}\right\} $ , where we use the convention that this maximum is $-\infty $ in case the set over which the maximum is taken is empty. This follows from Lemma 4.1 in PP4 with $\mathbb{K}$ equal to the collection $\{\Pi _{(\mathfrak{M} _{0}^{lin})^{\bot }}e_{i}(n):e_{i}(n)\notin \mathsf{B}\}$.]
remarkLet $T$ stand for either $T_{Het}$ or $T_{uc}$, and suppose that Assumption (ref) is satisfied in case of $T=T_{Het}$: By Remark (ref) in Appendix (ref) and Lemma 5.16 in PP3 the rejection regions $\{y:T(y)\geq C\}$ and $\{y:T(y)>C\}$ differ only by a $\lambda _{\mathbb{R}^{n}}$-null set. Since the measures $P_{\mu ,\sigma ^{2}\Sigma }$ are absolutely continuous w.r.t.$~\lambda _{\mathbb{R}^{n}}$ when $\Sigma $ is nonsingular, $P_{\mu ,\sigma ^{2}\Sigma }(T\geq C)=P_{\mu ,\sigma ^{2}\Sigma }(T>C)$ then follows, and hence the results in this section given for rejection probabilities $P_{\mu ,\sigma ^{2}\Sigma }(T\geq C)$ apply to rejection probabilities $P_{\mu ,\sigma ^{2}\Sigma }(T>C)$ equally well (under the above provision in case of $T=T_{Het}$). A similar remark applies to the results in Appendix (ref).

Some examples

We illustrate Theorem (ref) and Proposition (ref) with a few examples.

example(i) Suppose the design matrix satisfies $e_{i}(n)\notin \func{span}(X)$ for every $1\leq i\leq n$ (which will typically be the case). Then obviously the sufficient condition ((ref) ) is satisfied (in fact, for every choice of $\mathfrak{M}_{0}$, i.e., for every choice of restriction to be tested). And the sufficient condition ((ref)) is also satisfied provided $\mathsf{B}=\func{span}(X)$. (ii) Suppose the design matrix $X$ and the restriction $R$ are such that $ e_{i}(n)\notin \mathsf{B}$ for every $1\leq i\leq n$. Then the sufficient condition ((ref)) is clearly satisfied.

This example shows, in particular, that the sufficient conditions for size control are generically satisfied in the universe of all $n\times k$ design matrices $X$ (of rank $k$). Given the example, this is obvious for $T_{uc}$; and it follows for $T_{Het}$ by additionally noting that, for every given choice of restriction to be tested, the relation $\mathsf{B}=\func{span}(X)$ holds generically in the universe of all $n\times k$ design matrices $X$ (of rank $k$); see Lemma A.3 in PP3. The next example discusses the case where a standard basis vector is among the regressors.

exampleSuppose that $e_{1}(n)$ is the first column of $X$ and that $ e_{i}(n)\notin \func{span}(X)$ for every $2\leq i\leq n$. Suppose further that $R$ is of the form $R=(0,\tilde{R})$, where $\tilde{R}$ has dimension $ q\times (k-1)$. That is, the restriction to be tested does not involve the coefficient of the first regressor. Then it is easy to see that ((ref)) is satisfied and size control for $T_{uc}$ is thus possible. If also $\mathsf{B}=\func{span}(X)$ holds, then the same is true for ((ref)) and $T_{Het}$. [In case $R$ is not as above, but has a nonzero first coordinate, then it is easy to see that $1\in I_{1}( \mathfrak{M}_{0}^{lin})$, and hence ((ref)) is violated. It follows from Proposition (ref) that the rejection region $ \left\{ T_{uc}\geq C\right\} $ indeed has size $1$ for every choice of critical value $C$ when $\mathfrak{C}_{Het}$ is the heteroskedasticity model; and the same is true for $T_{Het}$, provided Assumption (ref) is satisfied.\footnote{ If Assumption (ref) is violated then $T_{Het}$ is identically zero, an uninteresting trivial case.}]

We continue with a few more examples where $X$ has a particular structure.

example(Heteroskedastic location model) Suppose $k=1$, $ x_{t1}=1$ for all $t$, $q=1$, $R=1$, and $r\in \mathbb{R}$. The heteroskedasticity model is given by $\mathfrak{C}_{Het}$. Then the conditions for size control in both parts of Theorem (ref) are satisfied (since it is easy to see that $\mathsf{B}$ coincides with $ \limfunc{span}(X)$ and that Assumption (ref) is satisfied). Note also that in this example $T_{Het}$ and $T_{uc}$ actually coincide in case $ d_{i}=n/(n-1)$ for all $i$, i.e., if the HC1, HC2, or HC4 weights are used, and differ only by a multiplicative constant if the HC0 or HC3 weights are employed; in particular, all these test statistics give rise to one and the same test if the respective smallest size-controlling critical values are used (cf. Remark (ref)).\footnote{ In fact, more is true in the location model: The test statistics $\tilde{T} _{Het}$ using the HC0R-HC4R weights (defined in Section (ref) below) all coincide (cf. Footnote (ref)), and they also coincide with $\tilde{T}_{uc}$ (also defined in Section (ref) below). Perusing the connection between $\tilde{T} _{uc}$ and $T_{uc}$ established in Section (ref), we can then even conclude that all the test statistics $T_{uc}$, $T_{Het}$ with HC0-HC4 weights, $\tilde{T}_{uc}$, and $\tilde{T}_{Het}$ with HC0R-HC4R weights give rise to (essentially) the same test, provided the respective smallest size-controlling critical values are used.} Furthermore, note that the here observed size controllability is in line with results in BakiSzek2005 stating that, for a certain range of significance levels $ \alpha $, the usual critical values obtained from an $F_{1,n-1}$ -distribution actually can be used as size-controlling critical values $ C(\alpha )$ for the test statistic $T_{uc}$ (in fact, these are then the smallest size-controlling critical values $C_{\Diamond }(\alpha )$).

The subsequent example is closely related to the Behrens-Fisher problem, see Remark (ref) in Appendix (ref).

example(Comparing the means of two heteroskedastic groups) Consider the problem of testing the equality of the means of two independent normal populations where the variances of each item may be different, even within a group. In our framework this corresponds to the case $k=2$, $x_{t1}=1$ for $1\leq t\leq n_{1}$, $x_{t1}=0$ for $n_{1}<t\leq n_{1}+n_{2}=n$, $x_{t2}=1-x_{t1}$, and $R=(1,-1)$ with $r=0$. The heteroskedasticity model is then again $\mathfrak{C}_{Het}$. We first assume that $n_{i}\geq 2$ holds for $i=1,2$. Note that in the present context $ T_{uc}$ is nothing else than the square of the two-sample t-statistic that uses a pooled variance estimator, and that $T_{Het}$ is the square of the two-sample t-statistic that uses appropriate variance estimators from each group (the particular form of \ the variance estimator being determined by the choice of $d_{i}$). Now, $e_{i}(n)\notin \func{span}(X)$ for every $1\leq i\leq n$ holds, and hence $T_{uc}$ is size controllable (cf. Example (ref)(i)). This is in line with results in Bakirov98, cf. also Section (ref). Furthermore, it is obvious that Assumption (ref) is satisfied (as $s=0$) and a simple calculation shows that $ B(y)=\hat{u}(y)^{\prime }A$, where $A$ is a diagonal matrix with $ a_{ii}=n_{1}^{-1}$ for $1\leq i\leq n_{1}$ and $a_{ii}=-n_{2}^{-1}$ else. This shows that the set $\mathsf{B}$ coincides with $\func{span}(X)$. Consequently, also $T_{Het}$ is size controllable (again cf. Example (ref)(i)). We also note here that the observed size controllability of $T_{Het}$ is in line with results in IbragMuell2016 stating that for a certain range of significance levels $\alpha $ and group sizes $n_{i}$ the usual critical values obtained from an $F_{1,\min (n_{1},n_{2})-1}$ -distribution actually can be used as size-controlling critical values $ C(\alpha )$ for the test statistic $T_{Het}$ in case $d_{i}$ is set equal to $\left( 1-h_{ii}\right) ^{-1}$; in fact, they are then the smallest size-controlling critical values, cf. the discussion preceding Theorem 1 in IbragMuell2016. In the rather uninteresting case $n_{1}=1$ and $ n_{2}\geq 2$, it is easy to see that Assumption (ref) is satisfied and that the size of both tests equals $1$ for all choices of critical values in view of Proposition (ref), since $e_{1}(n)\in \func{ span}(X)$ and $1\in I_{1}(\mathfrak{M}_{0}^{lin})=\{1,\ldots ,n\}$. The same is true if $n_{1}\geq 2$ and $n_{2}=1$. [The remaining and uninteresting case $n_{1}=n_{2}=1$ falls outside of our framework since we always require $ n>k$.]

The next example is an extension of the previous problem to the case of more than two groups. An interesting phenomenon occurs here: The sufficient conditions for size control of $T_{Het}$ given in Theorem (ref) are violated, but size controllability can nevertheless be established by additional arguments. Hence, this example provides an instance where the conditions in Part (b) of Theorem (ref) are not necessary.

example(Comparing the means of $k$\ heteroskedastic groups) We are given $k$ integers $n_{j}\geq 1$ with $\sum_{j=1}^{k}n_{j}=n$ describing group sizes where $k\geq 3$ holds. The regressors $x_{ti}$ for $ 1\leq i\leq k$ indicate group membership, i.e., they satisfy $x_{ti}=1$ for $ \sum_{j=1}^{i-1}n_{j}<t\leq \sum_{j=1}^{i}n_{j}$ and $x_{ti}=0$ otherwise. The heteroskedasticity model is given by $\mathfrak{C}_{Het}$. We are interested in testing $\beta _{1}=\ldots =\beta _{k}$. We thus may choose the $(k-1)\times k$ restriction matrix $R$ with $j$-th row $(1,0,\ldots 0,-1,\ldots ,0)$ where the entry $-1$ is at position $j+1$. Of course, $ q=k-1 $ and $r=0$ hold. We first consider the case where $n_{j}\geq 2$ for all $j$. Then clearly $k<n$ is satisfied. With regard to $T_{uc}$ we see immediately that $e_{i}(n)\notin \func{span}(X)$ for every $1\leq i\leq n$ follows (since $n_{j}\geq 2$ for all $j$) and thus the sufficient condition ( (ref)) for size control of $T_{uc}$ is satisfied. Turning to $T_{Het}$, it is easy to see that Assumption (ref) is satisfied (since $s=0$ in view of $n_{j}\geq 2$). Furthermore, the $j$-th row of $R(X^{\prime }X)^{-1}X^{\prime }$ is seen to be of the form \begin{equation*} (n_{1}^{-1},\ldots ,n_{1}^{-1},0,\ldots ,0,-n_{j+1}^{-1},\ldots ,-n_{j+1}^{-1},0\ldots ,0), \end{equation*} from which it follows that \begin{equation} R(X^{\prime }X)^{-1}X^{\prime }\limfunc{diag}(d_{1}\hat{u}_{1}^{2}(y),\ldots ,d_{n}\hat{u}_{n}^{2}(y))X(X^{\prime }X)^{-1}R=S_{1}\iota \iota ^{\prime }+ \limfunc{diag}(S_{2},\ldots ,S_{k}), \end{equation} where $\iota $ is the $(k-1)$-dimensional vector with entries all equal to $ 1 $ and where $S_{j}=n_{j}^{-2}\tsum_{t}d_{t}\hat{u}_{t}^{2}(y)=n_{j}^{-2} \tsum_{t}d_{t}(y_{t}-\bar{y}_{(j)})^{2}$ with the summation index $t$ running over all elements in the $j$-th group, and where $\bar{y}_{(j)}$ is the mean in group $j$. From ((ref)) it is not difficult to verify that the set $\mathsf{B}$ is given by \begin{equation*} \mathsf{B}=\tbigcup_{i,j=1,i\neq j}^{k}\left\{ y\in \mathbb{R} ^{n}:S_{i}(y)=S_{j}(y)=0\right\} =\tbigcup_{i,j=1,i\neq j}^{k}\limfunc{span} \left( x_{\cdot i},x_{\cdot j},\left\{ e_{l}(n):x_{li}=x_{lj}=0\right\} \right) . \end{equation*} Note that $\mathsf{B}$ is not a linear space and is strictly larger than $ \limfunc{span}(X)$. The set $\mathfrak{M}_{0}^{lin}$ is given by the span of the vector $e=(1,1,\ldots ,1)^{\prime }$. Hence, $I_{1}(\mathfrak{M} _{0}^{lin})=\left\{ 1,\ldots ,n\right\} $. Since $e_{i}(n)\in \mathsf{B}$ holds for every $i$, we conclude that the sufficient condition ((ref)) for size control of $T_{Het}$ is not satisfied and hence Part (b) of Theorem (ref) does not apply. However, it can be shown by additional arguments, see Proposition (ref) in Appendix\ (ref), that $T_{Het}$ is nevertheless size controllable, i.e., that ((ref)) holds.\footnote{ A smallest size-controlling critical value then also exists in view of Appendix (ref).} Next, in the case where $n_{j}=1$ for some $j$, but not for all $j$, Proposition (ref) shows that the size of the test based on $T_{uc}$ equals $1$ for all choices of critical values, since then for some $i$ the standard basis vector $e_{i}(n)$ is one of the regressors and thus we have $e_{i}(n)\in \func{span}(X)$ and $i\in I_{1}( \mathfrak{M}_{0}^{lin})=\{1,\ldots ,n\}$. For $T_{Het}$ the same is true if $ n_{j}=1$ holds for exactly one $j$ (because of Part (b) of Proposition (ref) and since then Assumption (ref) is satisfied as is easily seen); in case $n_{j}=1$ is true for (at least) two, but not all, values of $j$, $T_{Het}$ is identically zero (as then Assumption (ref) is violated), and thus is size-controllable in a trivial way. [The remaining and uninteresting case $n_{j}=1$ for all $j$ falls outside of our framework since we always require $n>k$.]

We close this section by one more example. Again, the sufficient conditions in Part (b) of Theorem (ref) fail to hold, but additional arguments based on Example (ref) establish size controllability of the test based on $T_{Het}$.

exampleConsider again the situation of Example (ref), except that now $R=I_{2}$, the $2\times 2$ identity matrix, (and again $r=0$). Then $q=k=2$ holds. Consider first the case where $ n_{i}\geq 2$ for $i=1,2$. Condition ((ref)) is then obviously satisfied, and hence $T_{uc}$ is size controllable. We next turn to $T_{Het}$. Since $\mathfrak{M}_{0}^{lin}=\left\{ 0\right\} $ we have $ I_{1}(\mathfrak{M}_{0}^{lin})=\{1,\ldots ,n\}$. Furthermore, simple computations show that Assumption (ref) is satisfied and that \begin{equation*} \mathsf{B}=\limfunc{span}\left( x_{\cdot 1},\left\{ e_{i}(n):i>n_{1}\right\} \right) \cup \limfunc{span}\left( x_{\cdot 2},\left\{ e_{i}(n):i\leq n_{1}\right\} \right) . \end{equation*} Obviously, the sufficient condition ((ref)) for size control of $T_{Het}$ is violated. Nevertheless, $T_{Het}$ is size controllable by the following argument:\footnote{ A smallest size-controlling critical value then also exists in view of Appendix (ref).} Simple computations show that $ T_{Het}(y)=T_{1}(y)+T_{2}(y)$ for $y\notin \mathsf{B}$, where $ T_{1}(y)=n_{1}^{2}\hat{\beta}_{1}^{2}(y)/\sum_{t=1}^{n_{1}}d_{t}\hat{u} _{t}^{2}(y)$ and $T_{2}(y)=n_{2}^{2}\hat{\beta}_{2}^{2}(y)/ \sum_{t=n_{1}+1}^{n}d_{t}\hat{u}_{t}^{2}(y)$. [If the denominator in the formula for $T_{i}(y)$ is zero for some $y\in \mathbb{R}^{n}$, we define $ T_{i}(y)$ as zero.] Since $\mathsf{B}$ is a $\lambda _{\mathbb{R}^{n}}$-null set, $P_{0,\sigma ^{2}\Sigma }(T_{Het}\geq C)\leq P_{0,\sigma ^{2}\Sigma }(T_{1}\geq C/2)+P_{0,\sigma ^{2}\Sigma }(T_{2}\geq C/2)$ for $C>0$. Now, it is easy to see that $P_{0,\sigma ^{2}\Sigma }(T_{i}\geq C/2)$ for $i=1,2$ coincides with the null rejection probability of a test for the mean in a heteroskedastic location model (based on a test statistic of the form ((ref))). However, as shown in Example (ref), such a test is size controllable. [In the case $n_{1}=1$ and $n_{2}\geq 2$ (or vice versa) condition ((ref)) is violated and the rejection region $ \left\{ T_{uc}\geq C\right\} $ has size $1$ for every $C$; furthermore, Assumption (ref) is violated, and hence $T_{Het}$ is identically zero. The case $n_{1}=n_{2}=1$ falls outside of our framework as then $k=n$.]

In Appendix (ref) we discuss yet another example where the sufficient condition of Part (b) of Theorem (ref) fails, but size-controllability can nevertheless be established.

Some variations on BakiSzek2005

(i) As noted in IbragMuell2010, testing a hypothesis regarding a scalar linear contrast in a heteroskedastic (Gaussian) linear regression model more general than a location model can often be converted to a testing problem in a heteroskedastic (Gaussian) location model by suitably dividing the data into subgroups and by considering groupwise least-squares estimators, thus making it amenable to the BakiSzek2005 result mentioned in Section (ref). However, this introduces additional questions such as how to divide up the data. In any case, this approach is limited to testing hypotheses on scalar linear contrasts. It also requires that the linear contrast subject to test is estimable in each subgroup.

(ii) In case the linear contrast subject to test is not estimable in each subgroup, but can be written as the difference of two linear contrasts where the first contrast is estimable in the first $G_{1}$ groups whereas the second contrast is estimable in the last $G_{2}$ groups (where we consider a total of $G_{1}+G_{2}$ groups), IbragMuell2016 point out that the problem can be converted into the problem of comparing two heteroskedastic (Gaussian) populations. Now, for such a two-sample comparison problem Bakirov98 shows for a certain two-sample $t$ -statistic (the square of which is $T_{uc}$, cf. Example (ref) above) how -- in the presence of heteroskedasticity -- size-controlling critical values can be constructed by appropriately transforming quantiles of a $t$-distribution; this result imposes conditions which entail that the nominal significance level $\alpha $ must be quite small (requiring $\alpha $ not to exceed $0.01$ for many group sizes, and often to be considerably smaller). This somewhat limits the applicability of Bakirov's result. Thus IbragMuell2016 go on to consider another two-sample $t$-statistic (the square of which is $T_{Het}$ with $d_{i}=\left( 1-h_{ii}\right) ^{-1}$, cf. Example (ref) above) and -- extending a result in MickeyBrown1966 -- provide a BakiSzek2005-type result, i.e., they show that the $(1-\alpha /2)$-quantile of a $t$-distribution with degrees of freedom equal to the smaller of the two sample sizes minus $1$ provides the smallest size-controlling critical value (for the two-sided test) even under heteroskedasticity.\footnote{ In the balanced case (i.e., if the two samples have the same cardinality) the test statistic considered in Bakirov98 actually coincides with the test statistic in IbragMuell2016.} This result holds under certain conditions on the sample sizes and only for small $\alpha $, but, e.g., allows for the choice $\alpha =0.05$. [We note here that the description of Bakirov98's result in IbragMuell2016 is inaccurate in that a certain transformation of the critical value is being ignored.]

(iii) In the problem of comparing two heteroskedastic (Gaussian) populations based on samples of equal size (\textquotedblleft balanced design\textquotedblright ) one can -- instead of using the two-sample $t$ -test statistics considered in Bakirov98 and IbragMuell2016 -- employ the Bartlett test statistic, which simply is the usual $t$-test statistic computed from the differences between the observations in the two samples.\footnote{ Certainly, there is some arbitrariness in how the observations are being \textquotedblleft paired\textquotedblright .} An advantage of this approach is that the original BakiSzek2005 result is directly applicable, and there is no need to resort to the results described in (ii).

(iv) Another quite special case that can be brought under the realm of the BakiSzek2005 result is a heteroskedastic (Gaussian) regression model with only one regressor that never takes the value zero. Dividing the $t$-th equation in the regression model by $x_{t}$, converts this into a heteroskedastic location problem.

(v) The results in (i)-(iv) immediately also apply if the errors in the regression are distributed as scale mixtures of Gaussians (cf. also Section (ref)).

Results for heteroskedasticity robust test statistics using restricted residuals

In this section we consider two further test statistics which are versions of $T_{Het}$ and $T_{uc}$ with the only difference that the covariance matrix estimators used are based on restricted -- instead of unrestricted -- residuals. The first one of these test statistics has been suggested in the literature, e.g., in DavidsonMacKinnon1985. We thus define

equation[equation omitted — 354 chars of source]

where $\tilde{\Omega}_{Het}=R\tilde{\Psi}_{Het}R^{\prime }$ and where $ \tilde{\Psi}_{Het}$ is given by

equation*[equation* omitted — 235 chars of source]

where the constants $\tilde{d}_{i}>0$ sometimes depend on the design matrix and on the restriction matrix $R$. Here $\tilde{u}\left( y\right) =y-X\tilde{ \beta}_{\mathfrak{M}_{0}}(y)=\Pi _{(\mathfrak{M}_{0}^{lin})^{\bot }}(y-\mu _{0})$, where the last expression does not depend on the choice of $\mu _{0}\in \mathfrak{M}_{0}$, and where $\tilde{u}_{t}\left( y\right) $ denotes the $t$-th component of $\tilde{u}\left( y\right) $. Typical choices for $ \tilde{d}_{i}$ are $\tilde{d}_{i}=1$, $\tilde{d}_{i}=n/(n-(k-q))$, $\tilde{d} _{i}=(1-\tilde{h}_{ii})^{-1}$, or $\tilde{d}_{i}=(1-\tilde{h}_{ii})^{-2}$ where $\tilde{h}_{ii}$ denotes the $i$-th diagonal element of the projection matrix $\Pi _{\mathfrak{M}_{0}^{lin}}$, see, e.g., DavidsonMacKinnon1985. Another suggestion is $\tilde{d}_{i}=(1-\tilde{h} _{ii})^{-\tilde{\delta}_{i}}$ for $\tilde{\delta}_{i}=\min (n\tilde{h} _{ii}/(k-q),4)$ with the convention that $\tilde{\delta}_{i}=0$ if $k=q$. \footnote{ Note that in case $k=q$ we have $\tilde{h}_{ii}=0$, and hence $\tilde{d} _{i}=1$ regardless of our convention for $\tilde{\delta}_{i}$.} For the last three choices of $\tilde{d}_{i}$ just given we use the convention that we set $\tilde{d}_{i}=1$ in case $\tilde{h}_{ii}=1$. Note that $\tilde{h} _{ii}=1 $ implies $\tilde{u}_{i}\left( y\right) =0$ for every $y$, and hence it is irrelevant which real value is assigned to $\tilde{d}_{i}$ in case $ \tilde{h}_{ii}=1$.\footnote{ In fact, $\tilde{h}_{ii}=1$ is equivalent to $\tilde{u}_{i}\left( y\right) =0 $ for every $y$, each of which in turn is equivalent to $e_{i}(n)\in $ $ \mathfrak{M}_{0}^{lin}$.} The five examples for the weights $\tilde{d}_{i}$ just given correspond to what is often called HC0R-HC4R weights in the literature.\footnote{In the case $k=q$ the HC0R-HC4R weights all coincide ($\tilde{d}_{i}=1$ for every $i$), and hence result in the same test statistic.}

The subsequent assumption ensures that the set of $y$'s for which $\tilde{ \Omega}_{Het}\left( y\right) $ is singular is a Lebesgue null set, implying that our choice of assigning $\tilde{T}_{Het}\left( y\right) $ the value zero in case $\tilde{\Omega}_{Het}\left( y\right) $ is singular has no import on the probabilistic results of the paper (as the measures $P_{\mu ,\sigma ^{2}\Sigma }$ are absolutely continuous). Also, as discussed further below, the assumption is in a certain sense unavoidable when using $\tilde{T} _{Het}$.

assumptionLet $1\leq i_{1}<\ldots <i_{s}\leq n$ denote all the indices for which $e_{i_{j}}(n)\in \mathfrak{M}_{0}^{lin}$ holds where $ e_{j}(n)$ denotes the $j$-th standard basis vector in $\mathbb{R}^{n}$. If no such index exists, set $s=0$. Let $X^{\prime }\left( \lnot (i_{1},\ldots i_{s})\right) $ denote the matrix which is obtained from $X^{\prime }$ by deleting all columns with indices $i_{j}$, $1\leq i_{1}<\ldots <i_{s}\leq n$ (if $s=0$ no column is deleted). Then $\limfunc{rank}\left( R(X^{\prime }X)^{-1}X^{\prime }\left( \lnot (i_{1},\ldots i_{s})\right) \right) =q$ holds.

Observe that this assumption only depends on $X$ and $R$ and hence can be checked. Obviously, a simple sufficient condition for Assumption (ref) to hold is that $s=0$ (i.e., that $e_{j}(n)\notin \mathfrak{M }_{0}^{lin}$ for all $j$), a generically satisfied condition. Furthermore, we introduce the matrix

eqnarray[eqnarray omitted — 379 chars of source]

Note that this matrix does not depend on the choice of $\mu _{0}\in \mathfrak{M}_{0}$. The following lemma collects some important properties of $\tilde{\Omega}_{Het}$ and $\mathsf{\tilde{B}}$ (defined in that lemma) and is reproduced from PPBoot for ease of reference.

lemma(a) $\tilde{\Omega}_{Het}\left( y\right) $ is nonnegative definite for every $y\in \mathbb{R}^{n}$. (b) $\tilde{\Omega}_{Het}\left( y\right) $ is singular (zero, respectively) if and only if $\limfunc{rank}(\tilde{B}(y))<q$ ($\tilde{B}(y)=0$, respectively). (c) The set $\mathsf{\tilde{B}}$ given by $\{y\in \mathbb{R}^{n}:\limfunc{ rank}(\tilde{B}(y))<q\}$ (or, in view of (b), equivalently given by $\{y\in \mathbb{R}^{n}:\det (\tilde{\Omega}_{Het}\left( y\right) )=0\}$) is either a $\lambda _{\mathbb{R}^{n}}$-null set or the entire sample space $\mathbb{R} ^{n}$. The latter occurs if and only if Assumption (ref) is violated (in which case the test based on $\tilde{T}_{Het}$ becomes trivial, as then $\tilde{T}_{Het}$ is identically zero). (d) Suppose Assumption (ref) holds. Then for every $\mu _{0}\in \mathfrak{M}_{0}$ the set $\mathsf{\tilde{B}}-\mu _{0}$ is a finite union of proper linear subspaces; in case $q=1$, $\mathsf{\tilde{B}}-\mu _{0} $ is even a proper linear subspace itself.\footnote{ Consequently, $\mathsf{\tilde{B}}$ is a finite union of proper affine subspaces, and is a proper affine subspace itself in case $q=1$.}$^{\text{,} } $\footnote{ If Assumption (ref) is violated, then $\mathsf{\tilde{B}}-\mu _{0}=\mathsf{\tilde{B}}=\mathbb{R}^{n}$ in view of Part (c).} [Note that $ \mathsf{\tilde{B}}-\mu _{0}$ does not depend on the choice of $\mu _{0}\in \mathfrak{M}_{0}$. In particular, if $r=0$, i.e., if $\mathfrak{M}_{0}$ is linear, we thus may set $\mu _{0}=0$.] (e) $\mathsf{\tilde{B}}$ is a closed set and contains $\mathfrak{M}_{0}$. Also $\mathsf{\tilde{B}}$ is $G(\mathfrak{M}_{0})$-invariant, and in particular $\mathsf{\tilde{B}}+\mathfrak{M}_{0}^{lin}=\mathsf{\tilde{B}}$.

In light of Part (c) of the lemma, we see that Assumption (ref) is a natural and unavoidable condition if one wants to obtain a sensible test from $\tilde{T}_{Het}$.\footnote{ If this assumption is violated then $\tilde{T}_{Het}$ is identically zero, an uninteresting trivial case.} Furthermore, note that if $\mathsf{\tilde{B}} =\mathfrak{M}_{0}$ is true, then Assumption (ref) must be satisfied (since $\mathfrak{M}_{0}$ is a $\lambda _{\mathbb{R}^{n}}$-null set as $k-q<n$ is always the case). For later use we also mention that under Assumption (ref) the statistic $\tilde{T}_{Het}$ is continuous at every $y\in \mathbb{R}^{n}\backslash \mathsf{\tilde{B}}$.\footnote{ If Assumption (ref) is violated, then $\tilde{T}_{Het}$ is constant equal to zero, and hence trivially continuous everywhere.}

We finally consider for completeness, and in analogy with $T_{uc}$,

equation[equation omitted — 343 chars of source]

where $\tilde{\sigma}^{2}(y)=\tilde{u}\left( y\right) ^{\prime }\tilde{u} \left( y\right) /(n-(k-q))\geq 0$ (which vanishes if and only if $y\in \mathfrak{M}_{0}$). Of course, our choice to set $\tilde{T}_{uc}(y)=0$ for $ y\in \mathfrak{M}_{0}$ again has no import on the probabilistic results in the paper, since $\mathfrak{M}_{0}$ is a $\lambda _{\mathbb{R}^{n}}$-null set (and since the measures $P_{\mu ,\sigma ^{2}\Sigma }$ are absolutely continuous). For later use we also mention that $\tilde{T}_{uc}$ is continuous at every $y\in \mathbb{R}^{n}\backslash \mathfrak{M}_{0}$. As we shall see in Section (ref), there is a close connection between $\tilde{T}_{uc}$ and $T_{uc}$.

remarkThe test statistics $\tilde{T}_{Het}$ as well as $\tilde{ T}_{uc}$ are $G(\mathfrak{M}_{0})$-invariant as is easily seen (with the respective exceptional sets $\mathsf{\tilde{B}}$ and $\mathfrak{M}_{0}$ also being $G(\mathfrak{M}_{0})$-invariant), but typically they are not nonsphericity-corrected F-type tests in the sense of Section 5.4 in PP2016.
remarkRemark (ref) also applies to $\tilde{T}_{Het}$ and $\tilde{T}_{uc}$. [To see this note that the respective exceptional sets $\mathsf{\tilde{B}}$ and $\mathfrak{M}_{0}$ are the same irrespective of whether $(R,r)$ or $( \bar{R},\bar{r})$ is used, and that $A$ cancels out in the respective quadratic forms appearing in the definitions of the test statistics.]

Size control results for $\tilde{T}_{Het}$ and $\tilde{T}_{uc}$ when $\mathfrak{C}=\mathfrak{C}_{Het}$

Here we discuss size control results for $\tilde{T}_{uc}$ as well as for $ \tilde{T}_{Het}$ over the heteroskedasticity model $\mathfrak{C}_{Het}$ (more precisely, over the null hypothesis $H_{0}$ described in ((ref)) with $\mathfrak{C}=\mathfrak{C}_{Het}$). Some peculiar properties of the test statistics $\tilde{T}_{uc}$ and $\tilde{T}_{Het}$ are then discussed in the following section.

We note that the first statement in Part (a) of the subsequent theorem is actually trivial, since $\tilde{T}_{uc}$ is bounded as shown in the next section (which also provides a discussion when non-trivial size-controlling critical values exist).

theorem(a) For every $0<\alpha <1$ there exists a real number $C(\alpha )$ such that \begin{equation} \sup_{\mu _{0}\in \mathfrak{M}_{0}}\sup_{0<\sigma ^{2}<\infty }\sup_{\Sigma \in \mathfrak{C}_{Het}}P_{\mu _{0},\sigma ^{2}\Sigma }(\tilde{T}_{uc}\geq C(\alpha ))\leq \alpha \end{equation} holds. Furthermore, even equality can be achieved in ((ref)) by a proper choice of $C(\alpha )$, provided $\alpha \in (0,\alpha ^{\ast }]\cap (0,1)$ holds, where $\alpha ^{\ast }=\sup_{C\in (C^{\ast },\infty )}\sup_{\Sigma \in \mathfrak{C} _{Het}}P_{\mu _{0},\Sigma }(\tilde{T}_{uc}\geq C)$ and where $C^{\ast }=\max \{\tilde{T}_{uc}(\mu _{0}+e_{i}(n)):i\in I_{1}(\mathfrak{M}_{0}^{lin})\}$ for $\mu _{0}\in \mathfrak{M}_{0}$ (with neither $\alpha ^{\ast }$ nor $ C^{\ast }$ depending on the choice of $\mu _{0}\in \mathfrak{M}_{0}$). (b) Suppose Assumption (ref) is satisfied.\footnote{ Condition ((ref)) clearly implies that the set $\mathsf{ \tilde{B}}$ is a proper subset of $\mathbb{R}^{n}$ and thus implies Assumption (ref). Hence, we could have dropped this assumption from the formulation of the theorem. A similar remark applies to some of the other results given below and will not be repeated.} Suppose further that $ \tilde{T}_{Het}$ is not constant on $\mathbb{R}^{n}\backslash \mathsf{\tilde{ B}}$.\footnote{The case where $\tilde{T}_{Het}$ is constant on $\mathbb{R}^{n}\backslash \mathsf{\tilde{B}}$ can actually occur under Assumption (ref), see Remark (ref) in Appendix (ref). In such a case $\tilde{T}_{Het}$ is trivially size-controllable (since $\mathsf{\tilde{B}}$ is a $\lambda _{\mathbb{R}^{n}}$-null set under Assumption (ref) and since all probability measures in ((ref)) are absolutely continuous). However, neither a smallest size-controlling critical value exists (when considering rejection regions of the form $\{\tilde{T}_{Het}\geq C\}$) nor can exact size controllability be achieved for $0<\alpha <1$. [If Assumption (ref) is violated, $\tilde{T}_{Het}$ is identically zero and a similar remark applies.]} Then for every $0<\alpha <1$ there exists a real number $C(\alpha )$ such that \begin{equation} \sup_{\mu _{0}\in \mathfrak{M}_{0}}\sup_{0<\sigma ^{2}<\infty }\sup_{\Sigma \in \mathfrak{C}_{Het}}P_{\mu _{0},\sigma ^{2}\Sigma }(\tilde{T}_{Het}\geq C(\alpha ))\leq \alpha \end{equation} holds, provided that for some $\mu _{0}\in \mathfrak{M}_{0}$ (and hence for all $\mu _{0}\in \mathfrak{M}_{0}$) \begin{equation} \mu _{0}+e_{i}(n)\notin \mathsf{\tilde{B}} \ \ for every \ i\in I_{1}( \mathfrak{M}_{0}^{lin}). \end{equation} Furthermore, under condition ((ref)), even equality can be achieved in ((ref)) by a proper choice of $ C(\alpha )$, provided $\alpha \in (0,\alpha ^{\ast }]\cap (0,1)$ holds, where now $\alpha ^{\ast }=\sup_{C\in (C^{\ast },\infty )}\sup_{\Sigma \in \mathfrak{C}_{Het}}P_{\mu _{0},\Sigma }(\tilde{T}_{Het}\geq C)$ and where $ C^{\ast }=\max \{\tilde{T}_{Het}(\mu _{0}+e_{i}(n)):i\in I_{1}(\mathfrak{M} _{0}^{lin})\}$ for $\mu _{0}\in \mathfrak{M}_{0}$ (with neither $\alpha ^{\ast }$ nor $C^{\ast }$ depending on the choice of $\mu _{0}\in \mathfrak{M }_{0}$). (c) Under the assumptions of Part (a) (Part (b), respectively) implying existence of a critical value $C(\alpha )$ satisfying ((ref)) (((ref)), respectively), a smallest critical value, denoted by $C_{\Diamond }(\alpha )$ , satisfying ((ref)) (((ref)), respectively) exists for every $0<\alpha <1$. \footnote{ Note that there are in fact no assumptions for Part (a). We have chosen this formulation for reasons of brevity.} And $C_{\Diamond }(\alpha )$ corresponding to Part (a) (Part (b), respectively) is also the smallest among the critical values leading to equality in ((ref)) (((ref)), respectively) whenever such critical values exist. [Although $C_{\Diamond }(\alpha )$ corresponding to Part (a) and (b), respectively, will typically be different, we use the same symbol.]\footnote{ Cf. also Appendix (ref).}

We see from the theorem that $\tilde{T}_{uc}$ is always size controllable over $\mathfrak{C}_{Het}$, but as discussed in Section (ref) below there is a caveat: Unless ((ref)), i.e., the necessary and sufficient condition for size-controllability of $T_{uc}$, is satisfied, size-controlling $\tilde{T}_{uc}$ leads to trivial tests. We also see that the condition for size control of $\tilde{T}_{Het}$ over $\mathfrak{ C}_{Het}$, i.e., condition ((ref)) is always satisfied in case $\mathsf{\tilde{B}}=\mathfrak{M}_{0}$ (since ((ref)) is then equivalent to $e_{i}(n)\notin \mathfrak{M}_{0}^{lin}$ for every $ i\in I_{1}(\mathfrak{M}_{0}^{lin})$). Furthermore, condition ((ref)) always only depends on $X$ and $R$; in particular, it does not depend on how the weights $\tilde{d}_{i}$ figuring in the definition of $\tilde{T}_{Het}$ have been chosen (note that $\mu _{0}+e_{i}(n)\notin \mathsf{\tilde{B}}$ is equivalent to $e_{i}(n)\notin \mathsf{\tilde{B}}-\mu _{0}$ and that the set $\mathsf{\tilde{B}}-\mu _{0}$ depends only on $X$ and $R$). Furthermore, the size-controlling critical values $C(\alpha )$ in Part (b) of the preceding theorem depend only on $X$, $R$, and $r$, as well as on the choice of weights $\tilde{d}_{i}$, whereas in Part (a) the dependence is only on $X$, $R$, and $r$. We do not show these dependencies in the notation. In fact, as shown in Lemma (ref) in Appendix (ref), it turns out that the size and the size-controlling critical values in both cases actually do not depend on the value of $r$ at all (provided the weights $\tilde{d}_{i}$ are not allowed to depend on $r$ in case of $\tilde{T}_{Het}$). Similarly, it is easy to see that $\alpha ^{\ast }$ and $C^{\ast }$ do not depend on $r$ (under the same provision as before in case of $\tilde{T}_{Het}$).

Similarly as in Section (ref), a critical value delivering size control over $\mathfrak{C}_{Het}$ also delivers size control over any other heteroskedasticity model $\mathfrak{C}$ since $\mathfrak{C} \subseteq \mathfrak{C}_{Het}$. Of course, for such a $\mathfrak{C}$ even smaller critical values (than needed for $\mathfrak{C}_{Het}$) may already suffice for size control. Also note that sufficient conditions implying size control over $\mathfrak{C}_{Het}$ may be more restrictive than sufficient conditions only implying size control over a smaller heteroskedasticity model $\mathfrak{C}$. For size control results tailored to such smaller models $\mathfrak{C}$ see Appendix (ref).

remark(Some equivalencies) If the respective smallest size-controlling critical values are used (provided they exist), the tests obtained from $\tilde{T}_{Het}$ with the HC0R and the HC1R weights, respectively, are identical, as these two test statistics differ only by a multiplicative constant. The same reasoning applies to the test statistics based on the HC0R-HC4R weights, respectively, in case $\tilde{h} _{ii}$ does not depend on $i$.
remark(Positivity of size-controlling critical values) For every $0<\alpha <1$ any $C(\alpha )$ satisfying ((ref)) or ((ref)) is necessarily positive. To see this observe that $\{\tilde{T}_{uc}\geq C\}=\{ \tilde{T}_{Het}\geq C\}=\mathbb{R}^{n}$ for $C\leq 0$, since both test statistics are nonnegative everywhere.

The next proposition complements Theorem (ref) and provides a lower bound for the size-controlling critical values (other than the trivial bound given in the preceding remark). The lower bound is useful for the same reasons as discussed subsequent to Proposition (ref).

proposition\footnote{ It is not difficult to show in the context of Parts (a) and (b) of the proposition that any critical value $C>C^{\ast }$ actually leads to size less than $1$. This follows from a reasoning similar as in Remark 5.4 of PP3.}$^{\text{,}}$\footnote{ If ((ref)) in Part (b) of the proposition does not hold, the conclusion of Part (b) can be shown to continue to hold with $C^{\ast }$ as defined in Theorem (ref)(b), and also with $C^{\ast }$ as defined in Lemma 5.11 of \ PP3 (note that under the assumptions of Part (b) of the proposition both definitions of $C^{\ast }$ actually coincide as shown in the proof of Theorem (ref)). If $ \tilde{T}_{Het}$ is constant on $\mathbb{R}^{n}\backslash \mathsf{\tilde{B}}$ or if Assumption (ref) fails (the latter implying $\tilde{T} _{Het}\equiv 0$), the conclusion of Part (b) also holds as is easily seen (regardless of which of the two definitions of $C^{\ast }$ is adopted).}(a) Any $C(\alpha )$ satisfying ((ref)) necessarily has to satisfy $C(\alpha )\geq C^{\ast }$, where $C^{\ast }$ is as in Part (a) of Theorem (ref). In fact, for any $ C<C^{\ast }$ we have $\sup_{\Sigma \in \mathfrak{C}_{Het}}P_{\mu _{0},\sigma ^{2}\Sigma }(\tilde{T}_{uc}\geq C)=1$ for every $\mu _{0}\in \mathfrak{M} _{0} $ and every $\sigma ^{2}\in (0,\infty )$. (b) Suppose Assumption (ref) and ((ref)) are satisfied, and that $\tilde{T}_{Het}$ is not constant on $\mathbb{R} ^{n}\backslash \mathsf{\tilde{B}}$. Then any $C(\alpha )$ satisfying ((ref)) necessarily has to satisfy $C(\alpha )\geq C^{\ast }$, where $C^{\ast }$ is as in Part (b) of Theorem (ref) . In fact, for any $C<C^{\ast }$ we have $\sup_{\Sigma \in \mathfrak{C} _{Het}}P_{\mu _{0},\sigma ^{2}\Sigma }(\tilde{T}_{Het}\geq C)=1$ for every $ \mu _{0}\in \mathfrak{M}_{0}$ and every $\sigma ^{2}\in (0,\infty )$.
remarkSuppose the assumptions of Part (a) (Part (b), respectively) of Theorem (ref) are satisfied. Then we know from that theorem that the size (over $\mathfrak{C}_{Het})$ of $\{ \tilde{T}_{uc}\geq C_{\Diamond }(\alpha )\}$ ($\{\tilde{T}_{Het}\geq C_{\Diamond }(\alpha )\}$, respectively) equals $\alpha $ provided $\alpha \in (0,\alpha ^{\ast }]\cap (0,1)$. If now $\alpha ^{\ast }<\alpha <1$, then the size (over $\mathfrak{C}_{Het})$ of $\{\tilde{T}_{uc}\geq C_{\Diamond }(\alpha )\}$ ($\{\tilde{T}_{Het}\geq C_{\Diamond }(\alpha )\}$, respectively) equals $\alpha ^{\ast }$ (where the $C_{\Diamond }(\alpha )$'s pertaining to Parts (a) and (b) may be different). This follows from $ C_{\Diamond }(\alpha )\geq C^{\ast }$ (see Proposition (ref) above) and Remark 5.13(i) in PP3).\footnote{ The assumptions for Part A of Proposition 5.12 in PP3 required in Remark 5.13 of that paper are satisfied under the assumptions of Theorem (ref) as shown in the proof of Theorem (ref) in Appendix (ref). In this proof also the condition $\lambda _{\mathbb{R}^{n}}(\tilde{T}_{uc}=C^{\ast })=0$ ($ \lambda _{\mathbb{R}^{n}}(\tilde{T}_{Het}=C^{\ast })=0$, respectively) required in Remark 5.13 of PP3 is verified.} This argument actually also delivers that $C_{\Diamond }(\alpha )=C^{\ast }$ must hold in case $ \alpha ^{\ast }<\alpha <1$.
remarkIn contrast to Section (ref), we have little information on the extent to which the sufficient conditions for size control in Part (b) of Theorem (ref) are also necessary. This is due to the fact that $\tilde{T}_{Het}$ is typically not a nonsphericity-corrected F-type test as noted in Remark (ref). What can be said in general in the context of Part (b) of Theorem (ref) in case ((ref)) is violated, is that the size of the rejection region $\{\tilde{T}_{Het}\geq C\}$ over $ \mathfrak{C}_{Het}$ is certainly equal to $1$ for every $C<\max \{\tilde{T} _{Het}(\mu _{0}+e_{i}(n)):\mu _{0}+e_{i}(n)\notin \mathsf{\tilde{B}}\}$, where $\mu _{0}\in \mathfrak{M}_{0}$ is arbitrary (the maximum being independent of the choice of $\mu _{0}\in \mathfrak{M}_{0}$) and where we use the convention that this maximum is $-\infty $ in case the set over which the maximum is taken is empty. This follows from Lemma 4.1 in PP4 with $\mathbb{K}$ equal to the collection $\{\Pi _{(\mathfrak{M} _{0}^{lin})^{\bot }}e_{i}(n):\mu _{0}+e_{i}(n)\notin \mathsf{\tilde{B}}\}$.
remarkSuppose $q=k$. Then Assumption (ref) is always satisfied (since $\mathfrak{M}_{0}$ being a singleton $\{\mu _{0}\}$ implies $\mathfrak{M}_{0}^{lin}=\{0\}$, and thus $s=0$ in Assumption (ref)). The subsequent claims are proved in Appendix (ref). (i) In case $q=k>1$, it is not difficult to see that then $\mu _{0}+e_{i}(n)\in \mathsf{\tilde{B}}$ for every $i=1,\ldots ,n$ holds, implying that the sufficient condition ((ref)) in Theorem (ref)(b) is violated. [In contrast, in case $q=k=1$, both examples where ((ref)) is satisfied as well as examples where ((ref)) is not satisfied can be found.] (ii) Despite of (i), in case $q=k\geq 1$ the test statistic $\tilde{T}_{Het}$ is always size-controllable over $\mathfrak{C}_{Het}$. This is so since in case $q=k\geq 1$ the statistic $\tilde{T}_{Het}$ is a bounded function. (iii) We also note that in case $q=k\geq 1$ both the case where $\tilde{T} _{Het}$ is constant on $\mathbb{R}^{n}\backslash \mathsf{\tilde{B}}$ as well as the case where $\tilde{T}_{Het}$ is not constant on $\mathbb{R} ^{n}\backslash \mathsf{\tilde{B}}$ can occur. [In the latter case a smallest size-controlling critical value exists in view of Appendix (ref). In the former case no smallest size-controlling critical value exists (when considering rejection regions of the form $\{\tilde{T} _{Het}\geq C\}$).]
remarkLet $\tilde{T}$ stand for either $\tilde{T}_{Het}$ or $\tilde{T}_{uc}$, where in case of $\tilde{T}=\tilde{T}_{Het}$ we suppose that Assumption (ref) is satisfied and that $\tilde{T}_{Het}$ is not constant on $\mathbb{R}^{n}\backslash \mathsf{\tilde{B}}$: By Lemma (ref) in Appendix (ref) the rejection regions $\{y:\tilde{T }(y)\geq C\}$ and $\{y:\tilde{T}(y)>C\}$ differ only by a $\lambda _{\mathbb{ R}^{n}}$-null set. Since the measures $P_{\mu ,\sigma ^{2}\Sigma }$ are absolutely continuous w.r.t.$~\lambda _{\mathbb{R}^{n}}$ when $\Sigma $ is nonsingular, $P_{\mu ,\sigma ^{2}\Sigma }(\tilde{T}\geq C)=P_{\mu ,\sigma ^{2}\Sigma }(\tilde{T}>C)$ then follows, and hence the results in this and the subsequent section given for rejection probabilities $P_{\mu ,\sigma ^{2}\Sigma }(\tilde{T}\geq C)$ apply to rejection probabilities $P_{\mu ,\sigma ^{2}\Sigma }(\tilde{T}>C)$ equally well (under the above provision in case of $T=\tilde{T}_{Het}$). A similar remark applies to the results in Appendix (ref).

Tests obtained from $\tilde{T}_{uc}$ or $\tilde{T}_{Het}$ can be trivial

For the test statistic $T_{uc}$ the rejection regions $\{T_{uc}\geq C\}$, as well as their complements, have positive ($n$-dimensional) Lebesgue measure for every positive real number $C$.\footnote{ The case $C\leq 0$ is uninteresting as the rejection region of $T_{uc}$ (and of all other test statistics considered) then are the entire space $\mathbb{R }^{n}$, since $T_{uc}$ (and the other test statistics considered) take on only nonnegative values.} This follows from Parts 5&6 of Lemma 5.15 in PP2016 together with Remark (ref) in Appendix (ref). As a consequence, all rejection probabilities -- under the null as well as under the alternative -- are positive and less than one regardless of the choice of $C>0$. [This is so because of our Gaussianity assumption and the fact that all $\Sigma \in \mathfrak{C}_{Het}$ are positive definite.] For similar reasons, the same is true for $T_{Het}$ provided Assumption (ref) is satisfied.\footnote{ If Assumption (ref) is not satisfied then $T_{Het}\equiv 0$, and the resulting test (with rejection region $\{T_{Het}\geq C\}$) is trivial as it never rejects for $C>0$, while it always rejects for $C\leq 0$.} The situation is somewhat different for tests derived from $\tilde{T}_{uc}$ or $ \tilde{T}_{Het}$ as we shall discuss next. In the course of this, we also establish a connection between $T_{uc}$ and $\tilde{T}_{uc}$ that is of independent interest. In this section the size of a test always refers to size over $\mathfrak{C}_{Het}$.

The case of $\tilde{T}_{uc}$

First, observe that $\tilde{T}_{uc}(y)\leq n-(k-q)$ holds for every $y\in \mathbb{R}^{n}$ and that this bound is sharp. To see this, note that using standard least-squares theory

equation[equation omitted — 159 chars of source]

for $y\notin \mathfrak{M}_{0}$ and that $\tilde{T}_{uc}(y)=0$ else; the bound is attained precisely for $y\in \limfunc{span}(X)\backslash \mathfrak{M }_{0}$. An immediate consequence of this observation is that any critical value $C\geq (n-(k-q))$ leads to a test with rejection region $\{\tilde{T} _{uc}\geq C\}$ that is either empty (if $C>n-(k-q)$) or is a $\lambda _{ \mathbb{R}^{n}}$-null set, namely $\limfunc{span}(X)\backslash \mathfrak{M} _{0}$ (if $C=n-(k-q)$). Consequently, such a test is trivial in that all rejection probabilities (under the null as well as under the alternative) are zero (because of our Gaussianity assumption and the fact that all $ \Sigma \in \mathfrak{C}_{Het}$ are positive definite). As an aside we note that any $C<n-(k-q)$ leads to a non-trivial test as is easily seen.

Of course, a critical value $C$ satisfying $C\geq n-(k-q)$ is certainly size-controlling, but is useless since it leads to a trivial test as just discussed. We now ask if and when the smallest size-controlling critical value $C_{\Diamond }(\alpha )$, guaranteed to exist by Part (c) of Theorem (ref), leads to a non-trivial test. [This is certainly so if $\alpha ^{\ast }$ in Part (a) of Theorem (ref) is positive, but note that the theorem is silent on this issue.] To obtain insight, we establish a simple, but important, relationship between the test statistics $\tilde{T}_{uc}$ and $T_{uc}$ that is of independent interest also: Note that standard least-squares theory gives

equation*[equation* omitted — 118 chars of source]

for $y\notin \limfunc{span}(X)$, and recall $T_{uc}(y)=0$ for $y\in \limfunc{ span}(X)$. Hence, we obtain

equation[equation omitted — 113 chars of source]

for every $y\notin \limfunc{span}(X)$, where $g:[0,\infty )\rightarrow \lbrack 0,n-(k-q))$ is continuous and strictly increasing with $ \lim_{x\rightarrow \infty }g(x)=(n-(k-q))$. [Since $T_{uc}(y_{m})\rightarrow \infty $ for every sequence $y_{m}\rightarrow y\in \limfunc{span} (X)\backslash \mathfrak{M}_{0}$, the sharpness of the bound $n-(k-q)$ can thus also be read-off from ((ref)).] As a consequence, for every critical value $C>0$, the rejection regions $\{\tilde{T}_{uc}\geq C\}$ and $ \{T_{uc}\geq g^{-1}(C)\}$ differ at most by $\limfunc{span}(X)$, which is a $ \lambda _{\mathbb{R}^{n}}$-null set; in particular, the rejection probabilities (under the null as well as under the alternative) are the same. \footnote{ This is so because of our Gaussianity assumption and the fact that all $ \Sigma \in \mathfrak{C}_{Het}$ are positive definite.} That is, the test statistics $\tilde{T}_{uc}$\ and $T_{uc}$\ give rise to (essentially) the same test, if the critical values chosen are linked by the function $g$\ as above. In particular, as we shall see, this is the case if the respective smallest size-controlling critical values are used for both test statistics (provided both these values exist).

To see what the preceding discussion entails for the existence of non-trivial size-controlling critical values for $\tilde{T}_{uc}$ we distinguish two cases. In the first case we shall see that non-trivial size-controlling critical values do not exist, whereas in the second case they do indeed exist.

Case 1: Condition ((ref)) is violated. Recall from Proposition (ref) that then the size of $\{T_{uc}\geq D\}$ is $1$ for every real $D$ (in particular, implying that $T_{uc}$ is not size controllable). It transpires from the preceding discussion, that hence the size of $\{\tilde{T}_{uc}\geq C\}$ must equal $1$ for every $C$ satisfying $ 0<C<n-(k-q)$ (and a fortiori for $C\leq 0$), because $D:=g^{-1}(C)$ is well-defined and real for $0<C<n-(k-q)$. As a consequence, any size-controlling critical value $C$ for $\tilde{T}_{uc}$ must satisfy $C\geq n-(k-q)$ (with the smallest size-controlling critical value given by $ n-(k-q) $), thus leading to a rejection region that is trivial in that it is empty (if $C>n-(k-q)$) or is a $\lambda _{\mathbb{R}^{n}}$-null set, namely $ \limfunc{span}(X)\backslash \mathfrak{M}_{0}$ (if $C=n-(k-q)$). That is -- while $\tilde{T}_{uc}$ is size-controllable in the present case -- it is so only in a trivial way.\footnote{ The trivial size-controlling critical values $C$ for $\tilde{T}_{uc}$ sort of correspond to using $\infty $ as a \textquotedblleft size-controlling critical value\textquotedblright\ for $T_{uc}$.} [Another way of arriving at the above conclusion is to use Part (a) of Proposition (ref) and to observe that in Part (a) of Theorem (ref) the quantity $C^{\ast }$ equals $n-(k-q)$. To see the latter, note that violation of condition ((ref)) implies existence of an index $i\in I_{1}(\mathfrak{M}_{0}^{lin})$ with $e_{i}(n)\in \limfunc{span} (X)$. In particular, $\hat{u}(\mu _{0}+e_{i}(n))=0$. Since $e_{i}(n)\notin \mathfrak{M}_{0}^{lin}$ must hold in view of $i\in I_{1}(\mathfrak{M} _{0}^{lin})$, and thus $\mu _{0}+e_{i}(n)\notin \mathfrak{M}_{0}$ for every $ \mu _{0}\in \mathfrak{M}_{0}$ must be true, we may use ((ref)) to arrive at $\tilde{T}_{uc}(\mu _{0}+e_{i}(n))=n-(k-q)$ for this $i\in I_{1}( \mathfrak{M}_{0}^{lin})$. This shows $C^{\ast }\geq n-(k-q)$. Equality then follows since $C^{\ast }\leq n-(k-q)$ trivially holds by ((ref)). As a point of interest we also note that $C^{\ast }=n-(k-q)$ implies that $ \alpha ^{\ast }$ in Part (a) of Theorem (ref) satisfies $ \alpha ^{\ast }=0$.]

Case 2: Condition ((ref)) is satisfied. In this case $T_{uc}$ is size controllable according to Theorem (ref). In particular, for any given $\alpha \in (0,1)$ there exists a smallest real number $D_{\Diamond }(\alpha )$ such that the size of $\{T_{uc}\geq D_{\Diamond }(\alpha )\}$ is less than or equal to $\alpha $, with equality holding for $\alpha \in (0,\alpha _{T_{uc}}^{\ast }]\cap (0,1)$ where $ \alpha _{T_{uc}}^{\ast }$ refers to $\alpha ^{\ast }$ appearing in Theorem (ref)(a), and recall from that theorem that $\alpha _{T_{uc}}^{\ast }>0$; and $D_{\Diamond }(\alpha )>0$ by Remark (ref).\footnote{If $\alpha _{T_{uc}}^{\ast }<\alpha <1$ , then the size, in fact, equals $\alpha _{T_{uc}}^{\ast }$; see Remark (ref).} Also note that the rejection region $\{T_{uc}\geq D_{\Diamond }(\alpha )\}$ is not trivial as it has positive $\lambda _{ \mathbb{R}^{n}}$-measure (and the same is true for its complement); see the discussion at the very beginning of Section (ref). Setting $ C_{\Diamond }(\alpha )=g(D_{\Diamond }(\alpha ))$ and using that $\{\tilde{T} _{uc}\geq C_{\Diamond }(\alpha )\}$ and $\{T_{uc}\geq g^{-1}(C_{\Diamond }(\alpha ))\}=\{T_{uc}\geq D_{\Diamond }(\alpha )\}$ differ at most by the $ \lambda _{\mathbb{R}^{n}}$-null set $\limfunc{span}(X)$, we see that (i) $ 0<C_{\Diamond }(\alpha )<n-(k-q)$, (ii) the size of $\{\tilde{T}_{uc}\geq C_{\Diamond }(\alpha )\}$ is less than or equal to $\alpha $, with equality holding for $\alpha \in (0,\alpha _{T_{uc}}^{\ast }]\cap (0,1)$, (iii) $ C_{\Diamond }(\alpha )$ is the smallest size-controlling critical value (recall that $g$ is strictly increasing), and (iv) the rejection region $\{ \tilde{T}_{uc}\geq C_{\Diamond }(\alpha )\}$ is not trivial as it has positive $\lambda _{\mathbb{R}^{n}}$-measure (and the same is true for its complement). In particular, note that $\tilde{T}_{uc}$ and $T_{uc}$ give rise to (essentially) the same test if the respective smallest size-controlling critical values are used. We furthermore note that in the present situation $C_{\tilde{T}_{uc}}^{\ast }=g(C_{T_{uc}}^{\ast })$ and $ \alpha _{\tilde{T}_{uc}}^{\ast }=\alpha _{T_{uc}}^{\ast }$ hold, where $ C_{T_{uc}}^{\ast }$, $\alpha _{T_{uc}}^{\ast }$ correspond to $C^{\ast }$, $ \alpha ^{\ast }$ in Part (a) of Theorem (ref), whereas $C_{ \tilde{T}_{uc}}^{\ast }$, $\alpha _{\tilde{T}_{uc}}^{\ast }$ correspond to $ C^{\ast }$, $\alpha ^{\ast }$ in Part (a) of Theorem (ref).\footnote{ If $\alpha _{\tilde{T}_{uc}}^{\ast }<\alpha <1$, then the size of $\{\tilde{T }_{uc}\geq C_{\Diamond }(\alpha )\}$ is, in fact, equal to $\alpha _{T_{uc}}^{\ast }=\alpha _{\tilde{T}_{uc}}^{\ast }$; cf. Footnote (ref) and Remark (ref).} In particular, $ \alpha _{\tilde{T}_{uc}}^{\ast }>0$ and $0\leq C_{\tilde{T}_{uc}}^{\ast }<n-(k-q)$ follow. These claims can be seen as follows: Under condition ((ref)) we have $\mu _{0}+e_{i}(n)\notin \limfunc{span}(X)$ for every $i\in I_{1}(\mathfrak{M}_{0}^{lin})$ and every $\mu _{0}\in \mathfrak{M}_{0}$. Consequently, $\tilde{T}_{uc}(\mu _{0}+e_{i}(n))=g(T_{uc}(\mu _{0}+e_{i}(n)))$, which proves $C_{\tilde{T} _{uc}}^{\ast }=g(C_{T_{uc}}^{\ast })$ in view of strict monotonicity of $g$. The relation $\alpha _{\tilde{T}_{uc}}^{\ast }=\alpha _{T_{uc}}^{\ast }$ then follows from the definitions of $\alpha _{\tilde{T}_{uc}}^{\ast }$ and $ \alpha _{T_{uc}}^{\ast }$ using that $\{\tilde{T}_{uc}\geq C\}$ and $ \{T_{uc}\geq g^{-1}(C)\}$ differ at most by the $\lambda _{\mathbb{R}^{n}}$ -null set $\limfunc{span}(X)$ for every $C>0$. Positivity of $\alpha _{ \tilde{T}_{uc}}^{\ast }$ now follows from positivity of $\alpha _{T_{uc}}^{\ast }$ discussed before, and $C_{\tilde{T}_{uc}}^{\ast }<n-(k-q)$ follows since $C_{\tilde{T}_{uc}}^{\ast }=g(C_{T_{uc}}^{\ast })$ and $ C_{T_{uc}}^{\ast }<\infty $. [Another way of proving $\alpha _{\tilde{T} _{uc}}^{\ast }>0$ and $0\leq C_{\tilde{T}_{uc}}^{\ast }<n-(k-q)$ without using relationship ((ref)), is to first establish $C_{\tilde{T} _{uc}}^{\ast }<n-(k-q)$ (from observing that $\hat{u}(\mu _{0}+e_{i}(n))\neq 0$ (as $\mu _{0}+e_{i}(n)\notin \limfunc{span}(X)$) for every $i\in I_{1}( \mathfrak{M}_{0}^{lin})$, which implies $\tilde{T}_{uc}(\mu _{0}+e_{i}(n))<n-(k-q)$ for every such $i$ in view of ((ref))) and then to proceed analogously as in the proof of Theorem (ref) below.]

While $\tilde{T}_{uc}$ is always size-controllable, whereas $T_{uc}$ is not, this does not represent any real advantage of $\tilde{T}_{uc}$ over $T_{uc}$ , as we have seen that $\tilde{T}_{uc}$ admits only trivial size-controlling critical values in the case where $T_{uc}$ is not size-controllable. Even more importantly, and already noted above, these test statistics give rise to (essentially) the same test if for both test statistics the respective smallest size-controlling critical values are used (provided they both exist).

The case of $\tilde{T}_{Het}$

For $\tilde{T}_{Het}$ we find that, not infrequently, it is also a bounded function, although we have no proof that this is always so. We illustrate the problems that can arise here first by an example. See also Remark (ref).

exampleConsider the $n\times 2$ design matrix $X$ where the first column represents an intercept, the second column is $x:=(1,-1,0,\ldots ,0)^{\prime }$, and $n\geq 3$. Let $R=(0,1)$, $r=0$, hence $q=1$. Obviously, the first column of $X$ spans $\mathfrak{M}_{0}^{lin}$. Since $ e_{i}(n)\notin \mathfrak{M}_{0}^{lin}$ for every $i=1,\ldots ,n$, Assumption (ref) holds. Furthermore, $\tilde{h}_{ii}=n^{-1}$. Thus $ \tilde{d}_{i}=\tilde{d}_{1}$ holds for every $i=1,\ldots ,n$ and for every of the five choices HC0R-HC4R. Note that $\tilde{d}_{1}^{-1}=1$ (HC0R), $ \tilde{d}_{1}^{-1}=1-n^{-1}$ (HC1R), $\tilde{d}_{1}^{-1}=1-n^{-1}$ (HC2R), $ \tilde{d}_{1}^{-1}=(1-n^{-1})^{2}$ (HC3R), and $\tilde{d}_{1}^{-1}=1-n^{-1}$ (HC4R), and hence $0<\tilde{d}_{1}^{-1}\leq 1$ for all five choices. Straightforward computations now show that $\tilde{\Omega}_{Het}(y)=\tilde{d} _{1}\left[ \left( y_{1}-\bar{y}\right) ^{2}+\left( y_{2}-\bar{y}\right) ^{2} \right] /4$ and \begin{equation} \tilde{T}_{Het}(y)=\tilde{d}_{1}^{-1}\left( y_{1}-y_{2}\right) ^{2}/\left[ \left( y_{1}-\bar{y}\right) ^{2}+\left( y_{2}-\bar{y}\right) ^{2}\right] \end{equation} whenever the numerator is positive, and $\tilde{T}_{Het}(y)=0$ otherwise. Here $\bar{y}$ denotes the arithmetic mean of the observations $y_{i}$. [For later use we also note that the set $\mathsf{\tilde{B}}$ is given by $\{y\in \mathbb{R}^{n}:y_{1}=y_{2}=\bar{y}\}$, and that the size control condition ( (ref)) is satisfied, since $e_{i}(n)\notin \mathsf{\tilde{ B}}$ for every $i=1,\ldots ,n$ (also note that $\mu _{0}$ can be chosen to be zero because of $r=0$). Furthermore, $\tilde{T}_{Het}$ is not constant on $\mathbb{R}^{n}\backslash \mathsf{\tilde{B}}$, since $\tilde{T} _{Het}(e_{1}(n))=\tilde{T}_{Het}(e_{2}(n))=\tilde{d} _{1}^{-1}n^{2}/[(n-1)^{2}+1]$ and $\tilde{T}_{Het}(e_{i}(n))=0$ for $i\geq 3$ (note $n\geq 3$) and since $e_{i}(n)\notin \mathsf{\tilde{B}}$ for every $i$ .] It is now evident from ((ref)) that $\tilde{T}_{Het}(y)\leq 2 \tilde{d}_{1}^{-1}$ for every $y\in \mathbb{R}^{n}$ and that this bound is attained whenever $y_{1}+y_{2}=2\bar{y}$ and $y_{1}\neq y_{2}$ (e.g., for $ y=x$). It follows that any critical value $C\geq 2\tilde{d}_{1}^{-1}$ leads to a test with rejection region that is empty if $C>2\tilde{d}_{1}^{-1}$, and is a Lebesgue null-set if $C=2\tilde{d}_{1}^{-1}$ (the latter following from Lemma (ref)(d) in Appendix (ref) together with some of the observations just noted after ((ref))); thus in both cases all the rejection probabilities are zero under the null as well as under the alternative (given our Gaussianity assumption and the fact that all $\Sigma \in \mathfrak{C}_{Het}$ are positive definite); in particular, these tests have zero power. Since $\tilde{d}_{1}^{-1}\leq 1$, this eliminates all critical values $C\geq 2$ from practical use. In particular, this eliminates the commonly used choice where $C$ is the $95\%$-quantile of a chi-square distribution with $1$ degree of freedom, which is approximately equal to $ 3.8415$.

In the preceding example any critical value $C\geq 2\tilde{d}_{1}^{-1}$ is trivially a size-controlling critical value for the given significance level $\alpha $ ($0<\alpha <1$), but it is \textquotedblleft too large\textquotedblright\ and leads to a trivial test. Certainly, one would prefer to use the smallest size-controlling critical value $C_{\Diamond }(\alpha )$ instead (which in the preceding example exists by Theorem (ref) and by what has been shown in the example) and one would hope that the resulting test is not trivial. As we shall show, this is indeed the case. To this end we first give a general result that, in particular, is applicable to the preceding example. Recall that $C_{\Diamond }(\alpha )$ is positive (Remark (ref)), and that Theorem (ref) is silent on whether $\alpha ^{\ast }>0$ or not.

theoremSuppose Assumption (ref) and ( (ref)) are satisfied, and that $\tilde{T}_{Het}$ is not constant on $\mathbb{R}^{n}\backslash \mathsf{\tilde{B}}$. Let $\alpha $ satisfy $0<\alpha <1$, and let $C^{\ast }$ and $\alpha ^{\ast }$\ be as defined in Part (b) of Theorem (ref). If $C^{\ast }<\sup_{y\in \mathbb{R}^{n}}\tilde{T}_{Het}(y)$ holds, then we have $\alpha ^{\ast }>0$, and the rejection region $\{\tilde{T}_{Het}\geq C_{\Diamond }(\alpha )\}$ is not a $\lambda _{\mathbb{R}^{n}}$-null set, where $ C_{\Diamond }(\alpha )$ is the smallest size-controlling critical value as in Part (c) of Theorem (ref).
remark(i) The preceding theorem clearly implies that -- under its assumptions -- the rejection probabilities associated with the rejection region $\{\tilde{T}_{Het}\geq C_{\Diamond }(\alpha )\}$ are positive under the null as well as under the alternative (in view of our Gaussianity assumption and the fact that all $\Sigma \in \mathfrak{C}_{Het}$ are positive definite). [While we already know from Theorem (ref)(b) and Remark (ref) that the rejection region $\{\tilde{T}_{Het}\geq C_{\Diamond }(\alpha )\}$ has size equal to $\alpha $ in case $\alpha \in (0,\alpha ^{\ast }]\cap (0,1)$, and has size equal to $\alpha ^{\ast }$ if $\alpha ^{\ast }<\alpha <1$, this by itself does not allow one to conclude that the rejection region has positive $\lambda _{\mathbb{R}^{n}}$-measure as the case $\alpha ^{\ast }=0$ is not ruled out by Theorem (ref)(b) and Remark (ref).] (ii) Suppose $C^{\ast }=\sup_{y\in \mathbb{R}^{n}}\tilde{T}_{Het}(y)$, but that the other assumptions of Theorem (ref) hold. \footnote{ We have not investigated whether this case can actually occur for $\tilde{T} _{Het}$. Recall that for $\tilde{T}_{uc}$ this case indeed can occur, see Case 1 in Section (ref).} Then the rejection region $\{ \tilde{T}_{Het}\geq C_{\Diamond }(\alpha )\}$ is a $\lambda _{\mathbb{R} ^{n}} $-null set; thus also the smallest (and hence any) size-controlling critical value leads to a trivial test. To prove the claim, note that by Proposition (ref) we have $C_{\Diamond }(\alpha )\geq C^{\ast }$ , implying that the rejection regions are either empty or coincide with the sets $\{\tilde{T}_{Het}=C^{\ast }\}$, respectively. In the latter case apply Part (d) of Lemma (ref) in Appendix (ref). We also point out that in the present case $\alpha ^{\ast }=0$ must hold since the rejection regions appearing in the definition of $\alpha ^{\ast }$ are all empty (because of $C>C^{\ast }=\sup_{y\in \mathbb{R}^{n}}\tilde{T}_{Het}(y)$ in the definition of $\alpha ^{\ast }$). (iii) If Assumption (ref) holds, but $\tilde{T}_{Het}$ is constant on $\mathbb{R}^{n}\backslash \mathsf{\tilde{B}}$, any rejection region of the form $\{\tilde{T}_{Het}\geq C\}$ is trivial in that the rejection region or its complement is a $\lambda _{\mathbb{R}^{n}}$-null set. [This case can actually occur, see Remark (ref) in Appendix (ref).] If Assumption (ref) is violated, $ \tilde{T}_{Het}$ is identically zero and a similar comment applies.
exampleWe continue the discussion of Example (ref). As noted prior to Theorem (ref), any critical value $C\geq 2\tilde{d} _{1}^{-1}$ is size-controlling in a trivial way, but leads to trivial rejection regions. We now show that the smallest size-controlling critical value $C_{\Diamond }(\alpha )$ indeed leads to a non-trivial test (which, in particular, has positive rejection probabilities in view of our Gaussianity assumption and the fact that all $\Sigma \in \mathfrak{C}_{Het}$ are positive definite). For this it suffices to verify the assumptions of Theorem (ref). The first three assumptions have already been verified above. From the calculations in Example (ref) it is now easy to see that $C^{\ast }=\tilde{d} _{1}^{-1}n^{2}/[(n-1)^{2}+1]$, which is smaller than $2\tilde{d} _{1}^{-1}=\sup_{y\in \mathbb{R}^{n}}\tilde{T}_{Het}(y)$. This completes the proof of the assertion. From Remark (ref)(i) we furthermore see that the rejection region $\{\tilde{T}_{Het}\geq C_{\Diamond }(\alpha )\} $ has size equal to $\alpha $ if $\alpha \in (0,\alpha ^{\ast }]\cap (0,1)$, and has size equal to $\alpha ^{\ast }$ if $\alpha ^{\ast }<\alpha <1 $. Finally we note that size-controlling critical values that do not lead to trivial tests must lie in the interval $[\tilde{d} _{1}^{-1}n^{2}/[(n-1)^{2}+1],2\tilde{d}_{1}^{-1})$ which is quite narrow as it is contained in the interval $[\tilde{d}_{1}^{-1},2\tilde{d}_{1}^{-1})$.

While the situation in Example (ref) is somewhat particular, the example may perhaps contribute to a better understanding of the Monte Carlo findings in DavidsonMacKinnon1985 and Godfrey2006, namely that the tests, obtained from $\tilde{T}_{Het}$ (employing HC0R-HC4R weights) in conjunction with conventional critical values such as the $95\%$-quantile of a chi-square distribution with appropriate degrees of freedom, can suffer from severe underrejection under the null.

remarkAnother class of examples where $\tilde{T}_{Het}$ is bounded is the case $q=k$ discussed in Remark (ref). Recall from that remark that in case $q=k>1$ condition ((ref)) is, however, never satisfied and thus Theorem (ref) is then not applicable. We have not further investigated non-triviality of tests based on $\tilde{T}_{Het}$ in case $q=k$ beyond the observations made in Remark (ref)(iii) that constancy of $\tilde{T}_{Het}$ on $\mathbb{R} ^{n}\backslash \mathsf{\tilde{B}}$ is possible in case $q=k\geq 1$ and thus then Remark (ref)(iii) applies.

Generalizations

Generalizations beyond Gaussianity

(i) All results in the preceding sections (as well as the extensions described in Appendix (ref)) referring to properties under the null hypothesis carry over as they stand to the situation where the error term $ \mathbf{U}$ in ((ref)) is elliptically symmetric distributed and has no atom at zero, i.e., $\mathbf{U}$ is distributed as $\sigma \Sigma ^{1/2} \mathbf{z}$ where $\mathbf{z}$ has a spherically symmetric distribution on $ \mathbb{R}^{n}$ that has no atom at zero.\footnote{ Note that all results in the preceding sections (as well as the extensions in Appendix (ref)), except for a few comments in Section (ref), are results referring to properties under the null hypothesis, } This is so since -- under this distributional model -- the null rejection probabilities of any $G(\mathfrak{M}_{0})$-invariant rejection region coincide with the corresponding null rejection probabilities under the Gaussian model (i.e., where $\mathbf{z}$ is standard Gaussian); see the discussion in Section 5.5 of PP2016 and Appendix E.1 of PP3. \footnote{ Note that all rejection regions considered in the preceding sections are $G( \mathfrak{M}_{0})$-invariant, because the test statistics considered are so.} This implies, in particular, not only that the sufficient conditions for size controllability under the above elliptically symmetric distributed model as well as under the Gaussian model are the same, but that also the numerical values of the size-controlling critical values coincide. As a consequence, the algorithms for computing the size-controlling critical values in the Gaussian case (used in Section (ref) and described in Section (ref) and Appendix (ref)) can be used in the above elliptically symmetric distributed case without any change whatsoever. The same is actually true if $\mathbf{z}$ has a distribution in a certain class larger than the class of spherical symmetric distributions with no atom at zero, see Appendix E.1 of PP3.

(ii) Furthermore, as discussed in detail in Appendix E.2 of PP3, the sufficient conditions for size controllability that we have derived under Gaussianity also imply size controllability for many more forms of distribution of $\mathbf{z}$ than those mentioned in (i); however, the corresponding size-controlling critical values may then differ from the size-controlling critical values that apply under Gaussianity.

(iii) Similarly as in Section 5.5 of PP2016, the negative results given in the preceding sections (as well as the ones described in Appendix (ref)) such as, e.g., size $1$ results, extend in a trivial way beyond the Gaussian model as long as the maintained assumptions on the feasible error distributions are weak enough to ensure that the implied (possibly semiparametric) model, i.e., set of distributions for $\mathbf{Y}$, contains the set given in ((ref)), but possibly contains also other distributions.

(iv) A further generalization beyond Gaussianity in the important special case where $\mathfrak{C}=\mathfrak{C}_{Het}$ is as follows: Suppose $\mathbf{ U}$ is distributed as $\sigma \Sigma ^{1/2}\limfunc{diag}(\mathbf{r)z}$ where $\mathbf{z}$ is standard normally distributed on $\mathbb{R}^{n}$ and where the $n$-dimensional random vector $\mathbf{r}$ is independent of $ \mathbf{z}$ with distribution $\rho $, where $\rho $ is a distribution on $ (0,\infty )^{n}$. [This includes the case where the elements of $\limfunc{ diag}(\mathbf{r)z}$ form an i.i.d. sample from a scale mixture of normals.] Let $Q_{\mu ,\sigma ^{2}\Sigma ,\rho }$ denote the implied distribution for $ \mathbf{Y}$ given by ((ref)) where $\mu =X\beta $. Consider now instead of ((ref)) the (semiparametric) model given by all distributions $Q_{\mu ,\sigma ^{2}\Sigma ,\rho }$ where $\mu \in \mathrm{\limfunc{span}}(X)$, $ 0<\sigma ^{2}<\infty $, $\Sigma \in \mathfrak{C}$, and $\rho $ is an arbitrary distribution on $(0,\infty )^{n}$. Then the sufficient conditions for size controllability derived under Gaussianity in earlier sections (and in Appendix (ref)) also imply size controllability in this larger model. In fact, the size-controlling critical values that apply under Gaussianity deliver also size control under this more general model. This follows from the following reasoning: Let $W$ be a Borel set in $\mathbb{R} ^{n}$ such that $P_{\mu _{0},\sigma ^{2}\Sigma }(W)\leq \alpha $ for every $ \mu _{0}\in \mathfrak{M}_{0}$, every $0<\sigma ^{2}<\infty $, and every $ \Sigma \in \mathfrak{C}_{Het}$. Then for every such $\mu _{0}$, $\sigma ^{2}$ , $\Sigma $, and every distribution $\rho $ on $(0,\infty )^{n}$ we have

eqnarray*[eqnarray* omitted — 448 chars of source]

where $\Sigma _{\mathbf{r}}^{1/2}:=\Sigma ^{1/2}\limfunc{diag}(\mathbf{r)/} s_{\mathbf{r}}$ with $s_{\mathbf{r}}$ denoting the positive square root of the sum of the diagonal elements of $(\Sigma ^{1/2}\limfunc{diag}(\mathbf{r)) }^{2}=\Sigma \limfunc{diag}^{2}(\mathbf{r)}$ and where $\sigma _{\mathbf{r} }=\sigma s_{\mathbf{r}}$. Here we have used that $P_{\mu ,\sigma _{\mathbf{r} }^{2}\Sigma _{\mathbf{r}}}(W)\leq \alpha $ by assumption since $\Sigma _{ \mathbf{r}}=\Sigma \limfunc{diag}^{2}(\mathbf{r)/}s_{\mathbf{r}}^{2}\in \mathfrak{C}_{Het}$ and $0<\sigma _{\mathbf{r}}<\infty $ hold for every realization of $\mathbf{r}$. In the above $\Pr $ denotes the probability measure governing $(\mathbf{r},\mathbf{z)}$ and $\mathbb{E}$ the corresponding expectation operator. \ As a consequence, the smallest size-controlling critical value under Gaussianity is also the smallest size-controlling critical value under the semiparametric model considered here, as the latter model contains the Gaussian model as a submodel. [In the special case where $\limfunc{diag}(\mathbf{r)}$ is a (random) multiple of the identity matrix $I_{n}$, the assumption $\mathfrak{C}=\mathfrak{C}_{Het}$ is superfluous as then $\Sigma _{\mathbf{r}}=\Sigma $, which by assumption belongs to the given $\mathfrak{C}$. In this case $\mathbf{U}$ satisfies the assumptions in (i), and hence (iv) adds little new, except that -- in contrast to (i) -- the reasoning works without use of $G(\mathfrak{M}_{0})$ -invariance.]

(v) It is apparent from the reasoning in (iv) that Gaussianity of $\mathbf{z} $ can be replaced by any other distributional assumption for which size controllability has already been established. E.g., one can in (iv) choose $ \mathbf{z}$ to have a spherically symmetric distribution without an atom at zero or to have a distribution in the more general class mentioned in (i) (note that all relevant rejection regions discussed in earlier sections are $ G(\mathfrak{M}_{0})$-invariant and thus (i) applies). In a similar vein, one can combine the results in Appendix E.2 of PP3 discussed in (ii) above with the reasoning outlined in (iv). We abstain from presenting details.

Generalizations to stochastic regressors

The assumption of nonstochastic regressors can be easily relaxed as follows: Suppose $X$ is random and $\mathbf{U}$ is conditionally on $X$ distributed as $N(0,\sigma ^{2}\Sigma )$, with $\sigma ^{2}=\sigma ^{2}(X)>0$ and $ \Sigma =\Sigma (X)\in \mathfrak{C}_{Het}$ where $\sigma ^{2}(\cdot )$ and $ \Sigma (\cdot )$ may vary in given classes of functions. The size control results such as Theorems (ref) and (ref) can then obviously be applied after one conditions on $X$ provided almost all realizations of $X$ satisfy the assumptions of those theorems, which will typically be the case (for brevity we do not provide a formal statement here).\footnote{ An appropriately modified statement applies to the size control results in Appendix (ref).} The resulting conditional size control statements then immediately imply that the so-obtained conditional size-controlling critical values $C=C(\alpha ,X)$ also control size unconditionally. Size $1$ results such as, e.g., Propositions (ref), (ref), or (ref) also extend to conditional size $1$ results in a similar manner provided $\sigma ^{2}(X)$ and $\Sigma (X)$ vary independently through all of $(0,\infty )$ and $\mathfrak{C}_{Het}$, respectively, for (almost) every realization of $X$, when the functions $\sigma ^{2}(\cdot )$ and $ \Sigma (\cdot )$ vary in the before mentioned function classes.\footnote{ See Footnote 40 in PPBoot for a discussion of sufficient conditions.} Generalizations to non-Gaussianity similarly as discussed in Section (ref) are also possible in the present context.

Results for other classes of tests

The results in Sections (ref) and (ref) (and in Appendix (ref)) have been obtained with the help of a general theory developed in Section 5 of PP2016, Section 5 of PP3, and Section 3.1 of PP4 that covers a very broad class of test statistics (and actually allows also for correlated errors). We note that, like in Section (ref), Gaussianity is again not essential for a good portion of this general theory, see Section 5.5 of PP2016 as well as Appendix E of PP3.\footnote{ Also arguments like in (iv) and (v) of Section (ref) can be applied to try to obtain generalizations.} We next discuss a few further situations that can also be handled by the general theory just mentioned but we refrain from spelling out the details:\footnote{ Applying some of the main results of this general theory (e.g., Corollary 5.6 or Proposition 5.12 of PP3) will require one to determine the set $\mathbb{J}(\mathcal{L},\mathfrak{C})$ defined in Appendix (ref). For the important cases $\mathfrak{C}=\mathfrak{C}_{Het}$ and $\mathfrak{C}= \mathfrak{C}_{(n_{1},\ldots ,n_{m})}$ (defined in Appendix (ref)), this is already accomplished in Propositions (ref) and (ref) in Appendix (ref) below.}

(i) The test statistic considered is an OLS-based test statistic like $ T_{Het}$, but where $\hat{\Omega}_{Het}$ is now replaced by an appropriate estimator derived from a given (possibly misspecified) parametric heteroskedasticity model described by a parameter vector $\theta $.

(ii) The test statistic is a Wald-type test statistic based on a (feasible) generalized least-squares estimator together with an appropriate covariance matrix estimator based on a given (possibly misspecified) parametric model. [This includes the (quasi-)maximum likelihood estimator (provided $\theta $ is unrelated to $\beta $).] Alternatively, the test statistic is the (quasi-)likelihood ratio or (quasi-)score test statistic based on this parametric model.

(iii) The test statistic is a Wald-type test statistic as in (ii), except that the covariance matrix estimator is now nonparametric (in the spirit of heteroskedasticity robust testing) as described in RomanoWolf2017. See also Cragg_1983, Cragg_1992, Flachaire_2005b, Wooldridge2010, Wooldridge2012, RomanoWolf2017, Lin_Chou_2018 , DiCiccio_Romao_Wolf_2019.

Some comments on power

Under our maintained assumptions, heteroskedasticity robust tests based on $ T_{Het}$ or $T_{uc}$ (using an arbitrary critical value $C$, including size-controlling ones) have positive power everywhere in the alternative (cf. the discussion at the beginning of Section (ref)). These tests can furthermore be shown to have power that goes to one as one moves away from the null hypothesis along sequences $(\mu _{l},\sigma _{l}^{2},\Sigma _{l})$ where $\mu _{l}$ moves further and further away from $ \mathfrak{M}_{0}$ (the affine space of means described by the restrictions $ R\beta =r$) in an orthogonal direction as $l\rightarrow \infty $, where $ \sigma _{l}^{2}$ converges to some finite and positive $\sigma ^{2}$, and $ \Sigma _{l}$ converges to a positive definite matrix. Despite of what has just been said, these tests can have, in fact not infrequently will have, infimal power equal to zero if $\mathfrak{C}$ is sufficiently rich, e.g., if $\mathfrak{C}=\mathfrak{C}_{Het}$; cf. Theorem 4.2 in PP2016, Lemma 5.11 in PP3, and Theorem 4.2 in PP4. [This does not contradict the before mentioned result as for this result sequences $\Sigma _{l}$ that converge to a singular matrix as $l\rightarrow \infty $ were ruled out.]

For tests based on $\tilde{T}_{Het}$ or $\tilde{T}_{uc}$ the situation is somewhat different. As shown in Section (ref), tests based on $ \tilde{T}_{Het}$ or $\tilde{T}_{uc}$ can be trivial for some choices of critical values $C$ (and then will have power zero everywhere in the alternative). However, if $C$ is chosen to be the smallest size-controlling critical value (provided it exists), the resulting tests obtained form $ \tilde{T}_{Het}$ or $\tilde{T}_{uc}$ will typically have positive power (under appropriate assumptions). In particular, then the test based on $ \tilde{T}_{uc}$ has the same power function as the test based on $T_{uc}$ that uses its smallest size-controlling critical value, provided the latter exists, see Section (ref). We have not further investigated the power properties of the tests based on $\tilde{T}_{Het}$ in any more detail on a theoretical level. The numerical results in Section (ref) seem to suggest that for these tests power may not go to one along sequences $(\mu _{l},\sigma _{l}^{2},\Sigma _{l})$ as mentioned above: in fact, power does not rise above the significance level $\alpha $ in some examples (on the range of alternatives considered). This feature makes tests based on $\tilde{T}_{Het}$ rather undesirable.

Computing the size and smallest size-controlling critical values

Consider a testing problem as in Equation ((ref)) with $ \mathfrak{C}=\mathfrak{C}_{Het}$ and let $T$ be one of the test statistics considered in the present article (e.g., $T_{Het}$ with some choice for the weights $d_{i}$). Suppose we want to numerically determine the size of the test with rejection region $\left\{ T\geq C\right\} $ for some user-supplied critical value $C$, i.e., we want to determine

equation[equation omitted — 169 chars of source]

Now, for all test statistics $T$ considered in the present article, this can be simplified to

equation[equation omitted — 93 chars of source]

where, subject to $\mu _{0}\in \mathfrak{M}_{0}$, $\mu _{0}$ can be chosen as desired. This is due to invariance properties of $T$, cf. Remarks (ref) and (ref). The quantity in ((ref)) can now be approximated numerically by any maximization algorithm where the probabilities are evaluated by Monte-Carlo methods or by the algorithm described in davies in case $q=1$, cf. Appendix (ref). \footnote{ Alternative to davies other algorithms like Imhof's algorithm, etc. can be used, some of which are also implementented in the R-package CompQuadForm (Duchesne).}

Suppose next that we want to numerically determine the smallest size-controlling critical value $C_{\Diamond }(\alpha )\in \mathbb{R}$ ($ 0<\alpha <1$) when using the test statistic $T$. [We assume here that the user knows that the smallest size-controlling critical value indeed exists, e.g., because the user has checked that the sufficient conditions developed in the present article hold, or because of other reasoning as, e.g., used in Example (ref).] Then, in view of ((ref)) and ((ref)), we need to compute $C_{\Diamond }(\alpha )$ as the smallest real number $C$ for which

equation[equation omitted — 115 chars of source]

holds. The quantity to the left in ((ref)) is non-increasing in the critical value $C$. Hence, to determine the smallest size-controlling critical value $C_{\Diamond }(\alpha )$, any line-search algorithm (in combination with an algorithm to determine the sizes as described before) can be used to compute $C_{\Diamond }(\alpha )$. We stress that it is of foremost importance to know that the testing problem at hand actually allows for size control before one attempts to numerically determine $C_{\Diamond }(\alpha )$. Hence, the theoretical results of the present article are of paramount importance also for the algorithmic aspect of the problem.

The specific algorithms we use to determine size and size-controlling critical values in our numerical studies are based on the above observations and are described in detail in Appendix (ref). They are made available in the R-package hrt (hrt) for the convenience of the user. The numerical procedures we use are heuristic in nature. Questions of efficacy of these algorithms or about theoretical guarantees are certainly important, but are beyond the scope of the present article.

Determining smallest size-controlling values numerically is important, e.g., if one wants to compare their magnitude with that of standard critical values in some special cases, as we do inter alia in the next section, or if one wants to obtain a confidence interval. However, a user who has observed the data and only wants to decide whether or not to reject the null hypothesis at significance level $\alpha $ ($0<\alpha <1$) when using $T$ combined with the smallest size-controlling critical value $C_{\Diamond }(\alpha )$, can actually perform this test without needing to compute $ C_{\Diamond }(\alpha )$: Let $y_{obs}$ be the observed data. Define the \textquotedblleft maximal p-value" as

eqnarray[eqnarray omitted — 334 chars of source]

where the second equality in the display follows from the invariance properties mentioned before (and $\mu _{0}\in \mathfrak{M}_{0}$ can be chosen as desired). It is now not difficult to see that $p(y_{obs})\leq \alpha $ is equivalent to $T(y_{obs})\geq C_{\Diamond }(\alpha )$. That is, rejecting if and only if $p(y_{obs})\leq \alpha $ leads to exactly the same test as rejecting if and only if $T(y_{obs})\geq C_{\Diamond }(\alpha )$, with the former description having the advantage that the more costly computation of $C_{\Diamond }(\alpha )$ can be avoided. What needs to be computed is ((ref)), which, however, is nothing else than the size of the test when using the \textquotedblleft critical value\textquotedblright\ $ T(y_{obs})$. Hence, $p(y_{obs})$ can be determined by any algorithm that determines the size ((ref)) for the user-supplied \textquotedblleft critical value\textquotedblright\ $C=T(y_{obs})$. In particular, the routine \textquotedblleft size\textquotedblright\ provided in the R-package hrt (hrt) can be used for this purpose. Note that checking whether $p(y_{obs})\leq \alpha $ avoids the line-search part (as outlined following ((ref))), and is thus computationally more efficient than first determining $C_{\Diamond }(\alpha )$ (as outlined above) and then checking whether $T(y_{obs})\geq C_{\Diamond }(\alpha )$.

Finally, we note that if (contrary to what we assume in this section) no size-controlling critical value exists for a given significance level $ \alpha \in (0,1)$, then the maximal p-value in ((ref)) is larger than $ \alpha $ for every possible observed value $y_{obs}$, and the corresponding test thus never rejects and thus is uninformative. Hence, while the explicit computation of a smallest size-controlling critical value can be avoided for performing a single test, knowing its existence is important as then the resulting test is guaranteed to be informative (non-trivial) if $T_{Het}$ or $T_{uc}$ is being used; and the same is true for $\tilde{T}_{Het}$ and $\tilde{T}_{uc}$ under the conditions discussed in Section (ref).

We also note that in view of the discussion in Section (ref) the algorithms for computing null rejection probabilities, size, and smallest size-controlling critical values discussed in this section and Appendix (ref) remain valid for elliptically symmetric distributed data without any need for modification. With regard to computing size and smallest size-controlling critical values, the same is also true for the semiparametric model described in (iv) of Section (ref).

Numerical results

In this section we pursue two goals:

enumerate• In Subsection (ref) we show numerically that any of the usual heteroskedasticity robust tests can suffer from overrejection of the null hypothesis (sometimes by a large margin) when they are based on conventional critical values. While this adds to similar evidence already present in the literature for the HC0-HC4 based tests (see Section (ref)), this seems to be a new observation for the HC0R-HC4R based tests. In any case, this drives home the point that none of these heteroskedasticity robust tests based on conventional critical values comes with a guarantee that size is controlled by the nominal significance level $ \alpha $. Consequently, instead of using conventional critical values, this strongly suggests to use (smallest) size-controlling critical values as investigated in this paper. • In Subsection (ref) we then numerically compute smallest size-controlling critical values and study the power behavior of tests based on such size-controlling critical values in some examples.

In this section (and in the attending Appendices (ref) and (ref)) we shall often refer to $T_{Het}$ as HC0-HC4 when we want to stress that the weights $d_{i}$ being used are the HC0-HC4 weights, respectively, see Section (ref). Similarly, we shall refer to $\tilde{T}_{Het}$ as HC0R-HC4R when the HC0R-HC4R weights are used, see Section (ref). For reasons of uniformity of notation, we shall then often denote $T_{uc}$ as UC and $\tilde{T}_{uc}$ as UCR. Furthermore, throughout this section we consider the heteroskedastic Gaussian linear model with $\mathfrak{C}=\mathfrak{C}_{Het}$ as introduced in Section (ref); in particular, the notion of size in the present section (and the attending appendices) always refers to this model.

The algorithms for computing rejection probabilities, the size of a test, and size-controlling critical values used in the before-mentioned numerical computations are described in Section (ref) and Appendix (ref). Implementations are available as an R-package hrt ( hrt).

Tests based on conventional critical values

We consider the important case $q=1$, and first illustrate numerically that none of the test statistics UC, HC0-HC4, UCR, and HC0R-HC4R combined with the critical value $C_{\chi ^{2},0.05}\approx 3.8415$ results in a test that is guaranteed to have size less than or equal to $\alpha =0.05$. This is achieved by providing instances of design matrices $X$ and of hypotheses, described by $(R,r)$, such that the respective test has size larger than the nominal significance level $\alpha =0.05$, often by a large margin. Here $ C_{\chi ^{2},0.05}$ denotes the $95\%$-quantile of a chi-square-distribution with $1$ degree of freedom. [This critical value has a justification for use with HC0-HC4 or HC0R-HC4R via asymptotic considerations, but, in general, there is no such justification for use with UC or UCR, which we nevertheless include here for completeness.\footnote{ Of course, in the special case of homoskedasticity, the before mentioned justification also applies to UC and UCR.}] That is, in the instances we exhibit, this conventional critical value turns out to be too small. We next show similar results for other suggestions of critical values, e.g., for \textquotedblleft degree-of-freedom\textquotedblright\ adjustments to the conventional chi-square based critical value such as the Bell-McCaffrey adjustment (BellMcCa, Imbkoles2016). It is important to note here that in all the instances mentioned our conditions for size-controllability are satisfied, showing that size-controlling critical values can actually be found; hence, the overrejection problems mentioned before are not intrinsic problems, but only reflect the fact that conventional critical values can be a bad choice and do not guarantee size control. [In the present context it is worth recalling that for the test statistics HC0R-HC4R we have already shown in Example (ref) in Section (ref) that other situations can be found in which conventional critical values such as, e.g., $C_{\chi ^{2},0.05}$ are too large, as the resulting tests reject with probability zero only (under the null as well as under the alternative), rendering these tests useless.]

To uncover instances where the conventional critical value $C_{\chi ^{2},0.05}$ is too small, we make use of the following observation: In case a given test statistic from the above list (together with a given design matrix $X$ and hypothesis described by $(R,r)$) is such that the lower bound $C^{\ast }$ on size-controlling critical values obtained in Propositions (ref) ((ref), respectively) exceeds $C_{\chi ^{2},0.05}$, we are done, as we then know that the critical value $C_{\chi ^{2},0.05}$ leads to a test that has size $1$. [As noted subsequent to Theorems (ref) and (ref), the value of $r$ actually plays no role here, and we may set it to zero.]

Since the lower bounds $C^{\ast }$ for size-controlling critical values in Propositions (ref) ((ref), respectively) depend on the given test statistic, on $X$ and on $R$, we may -- for any given choice of test statistic and any given $R$ -- numerically search for particularly \textquotedblleft hostile\textquotedblright\ design matrices, i.e., for design matrices for which the lower bound is large, to see whether matrices $ X$ exist for which the lower bound exceeds $C_{\chi ^{2},0.05}$. We only do this for $k=2$, $R=(0,1)$, $r=0$, and $n=25$, and restrict ourselves to matrices $X$ with first column representing an intercept. The concrete search used is detailed in Appendix (ref), see Algorithm (ref) in particular. Table (ref) provides, for every test statistic considered, the lower bound $C^{\ast }$ corresponding to the most \textquotedblleft hostile\textquotedblright\ design matrix found by the search. [As the searches are run separately for each test statistic, the resulting \textquotedblleft hostile\textquotedblright\ design matrices will typically differ across the runs.]\footnote{ Since, for example, HC0 is a multiple of HC1, where the factor is $ n/(n-k)=1.09$, we know that the \textquotedblleft hostile\textquotedblright\ design matrix obtained from the search for HC1 leads to a $C^{\ast }$-value of $1.09\ast 1711.19=1865.20$ for HC0, larger than the value 95.56 obtained from the search for HC0, cf. Table (ref). We could have reported this larger value, but decided to present the raw results from our searches as this is sufficient for our purposes. We also note that our search procedure detailed in Appendix (ref) does not seriously attempt to optimize the $C^{\ast }$-value (for every one of the test statistics considered) over the set of all feasible $X$, but is only a crude search for finding a matrix resulting in a $C^{\ast }$-value sufficiently large for our purposes.}

table[table omitted — 714 chars of source]

In combination with the theoretical results from Propositions (ref) and (ref), Table (ref) shows that for some design matrices $X$ the critical value $C_{\chi ^{2},0.05}\approx 3.8415$ results in a test with size equal to $1$ when combined with UC, HC0-HC2, and also with UCR. [This is so despite the fact that, for any of the twelve test statistics considered, the sufficient conditions for size-control in the pertaining theorems in Sections (ref) and (ref) are satisfied for all relevant $X$ matrices encountered in the numerical procedure (as we have checked), and hence it is known that size-controlling critical values exist in all these situations!] Table (ref) is not informative about the size of the remaining seven tests, since the corresponding entries in that table are all less than $C_{\chi ^{2},0.05}$. To obtain insight into the sizes of the remaining seven tests we do the following: for each of the tests we numerically compute the size for various instances of design matrices (the ones that give rise to Table (ref)) and report the largest one of these sizes (\textquotedblleft worst case\textquotedblright\ sizes) in Table (ref).\footnote{ Of course, considering additional design matrices $X$ would potentially lead to even larger sizes.} We actually do this for all twelve tests considered. The algorithm used in the size computation is the implementation of Algorithm (ref) in the R-package hrt (hrt), cf. the description in Appendices (ref) and (ref). Table (ref) now clearly shows that for every test statistic considered an instance can be found, in which the size of the test (when using the critical value $C_{\chi ^{2},0.05}$) clearly exceeds the nominal significance level $\alpha =.05$. The lowest value in that table is attained by HC4R, but a size of $0.10$ is still twice the nominal significance level $ \alpha $.

We note that the numbers shown in Table (ref) actually only represent numerically determined lower bounds for the actual sizes, as their computation involves (for any given $X$) a numerical search procedure (over the set $\mathfrak{C}_{Het}$) for the worst-case null rejection probability; that is, the numbers shown in Table (ref) correspond to the null rejection probability computed from a \textquotedblleft bad\textquotedblright\ covariance matrix $\Sigma $, but potentially not for the \textquotedblleft worst\textquotedblright\ possible one. [In this process, for any given $\Sigma \in \mathfrak{C}_{Het}$, we have to numerically compute the null rejection probability, which can be done quite accurately in case $q=1$ by algorithms like the Davies algorithm, see Appendix (ref) as well as Appendix (ref).] In particular, the entries in the $0.98$-$0.99$ range in Table (ref) are numerically determined lower bounds for the size, which, in fact, we know to be equal to $1$ in light of Table (ref). [We could have used this knowledge to replace the entries in question in Table (ref) by $1$, but we decided otherwise in order to showcase the concrete outcome of the numerical algorithm that has been run. Of course, one could also improve this outcome by using a higher accuracy parameter in the optimization procedures involved.]

Sometimes -- without much theoretical justification in general -- it is suggested in the literature to replace $C_{\chi ^{2},0.05}$ by the $95\%$ -quantile of an $F_{1,n-k}$-distribution, which is approximately $4.28$ in the situation considered here ($n-k=23$). Obviously, from Table (ref) we see that the conclusions regarding UC, HC0-HC2, and UCR remain the same when this critical value is used. Repeating the exercise that has led to Table (ref), but with $C_{\chi ^{2},0.05}$ replaced by the $95\%$-quantile of an $F_{1,n-k}$-distribution, gives Table (ref), leading essentially to the same conclusions.

table[table omitted — 317 chars of source]

\textquotedblleft Degree-of-freedom\textquotedblright\ adjustments to the conventional chi-square based critical value such as the Bell-McCaffrey adjustment (BellMcCa) have been discussed in the literature. In particular, Imbkoles2016 suggested to use this adjustment with the HC2 statistic. We have repeated the above exercise that has led to the entry for HC2 in Table (ref), but with $C_{\chi ^{2},0.05}$ replaced by the Bell-McCaffrey adjustment. For the computation of the Bell-McCaffrey adjustment we relied on the R-package dfadjust (dfadjust). For the resulting test, the largest size that was found in our computations was $0.24$, which is more than four times the nominal significance level. It transpires that this adjustment does also not come with a size-guarantee.

We conclude here by stressing that the negative findings in this subsection were obtained in a very simple model with only two regressors and where only one of the parameters is subject to test. For more complex models and test problems the size distortions may even be worse.

Power comparison of tests based on size-controlling critical values

A power comparison of two tests, both conducted at a given nominal significance level $\alpha $, makes sense only if both tests actually are level $\alpha $ tests, i.e., if both tests have a size not exceeding the given $\alpha $. For this reason, we now compare the tests obtained from the statistics UC, HC0-HC4, UCR, HC0R-HC4R only when respective smallest size-controlling critical values are used. Our theoretical results concerning the existence of size-controlling critical values, together with the algorithms for their computation in Appendix (ref), allow for such a comparison in terms of power. In all cases considered in this section $q=1$ will hold.

Throughout, in addition to the power functions of the before-mentioned tests, we also show as a benchmark the power function of the infeasible (i.e., oracle) GLS-based $F$-test conducted at the $5\%$-significance level, that makes use of knowledge of $\Sigma $. For given $\Sigma \in \mathfrak{C}_{Het}$, the distribution of this infeasible GLS-based $F$-test statistic is (under $P_{X\beta ,\sigma ^{2}\Sigma }$ with $\beta \in \mathbb{ R}^{k}$, $\sigma ^{2}\in (0,\infty )$) a noncentral $F_{1,n-k}$-distribution with noncentrality parameter $\delta ^{2}$, where

equation*[equation* omitted — 97 chars of source]

Since the power functions of all the tests considered in our study depend on the parameters $\beta $, $\sigma ^{2}$, and $\Sigma $ only through $(R\beta -r)/\sigma $ and $\Sigma $ (because of $G(\mathfrak{M}_{0})$-invariance and Proposition 5.4 in PP2016), and thus depend only on $\delta $ and $ \Sigma $, we shall -- for given $\Sigma $ -- present all these power functions as a function of $\delta $. We show only results for $\delta \geq 0 $, as the power functions in fact depend on $\delta $ only through $ \left\vert \delta \right\vert $ (for given $\Sigma $); see Proposition 5.4 in PP2016.

Comparing the means of two heteroskedastic groups

As a practically relevant example, we here compare the power of tests based on size-controlling critical values in the context of Example (ref). That is, we treat the problem of comparing the means of two heteroskedastic groups (e.g., a treatment and a control group), the null hypothesis being that the difference of expected outcomes in each group is zero. We consider the case where $n=30$ and $\alpha =0.05$. Furthermore, we vary the size $n_{1}$ of the first group ($n_{1}\in \{3,9,15\}$), corresponding to a \textquotedblleft strongly unbalanced\textquotedblright , \textquotedblleft moderately unbalanced\textquotedblright , and \textquotedblleft balanced\textquotedblright\ design, respectively. We compute the power for a number of covariance matrices $\Sigma _{a}$ given as follows: For $a=1,5,9$ define

equation*[equation* omitted — 178 chars of source]

where the first $n_{1}$ (and last $n-n_{1}$, respectively) diagonal entries of each $\Sigma _{a}$ are constant. That is, we look at power functions evaluated at covariance matrices under which the subjects in the same group actually have the same variances. [For brevity we do not report power functions for covariance matrices not sharing this property.] For the balanced design, we note that $\Sigma _{1}$ and $\Sigma _{9}$ lead to the same power of each test (but we report all results for completeness), and that $\Sigma _{5}$ corresponds to homoskedasticity.

The critical values are chosen in each case as the smallest critical value guaranteeing size control over $\mathfrak{C}_{Het}$ (implying, of course, that the corresponding tests can have null rejection probabilities smaller than $\alpha $ for the covariance matrices $\Sigma _{a}$ considered). The existence of said critical values follows from our theory and is discussed in detail in Example (ref) for the test statistics UC and HC0-HC4; in particular, all assumptions of Theorems (ref) are satisfied. For UCR the existence is guaranteed by Part (a) of Theorem (ref). With regard to the test statistics HC0R-HC4R, note that Assumption (ref) is satisfied since $e_{i}(n)\notin \mathfrak{M}_{0}^{lin}=\mathfrak{M}_{0}=\limfunc{span}((1,\ldots ,1)^{\prime })$ for every $i=1,\ldots ,n$ as $n=30>k=2$. This also shows that the sufficient condition for size control ((ref)) is satisfied as $\mathsf{\tilde{B}}=\mathfrak{M}_{0}$ is easily verified and since one may set $\mu _{0}=0$. We have verified the non-constancy assumption on the test statistics HC0R-HC4R in Theorem (ref) numerically. As a consequence, all assumptions of Part (b) of Theorem (ref) are satisfied.

We note that some of the test statistics differ from each other only by a known multiplicative constant and hence are equivalent in the sense that they give rise to the same test when the respective smallest size-controlling critical value is employed, see Remarks (ref) and (ref): In the unbalanced case ($n_{1}\in \{3,9\}$),\ HC0 and HC1 are equivalent in this sense, as are HC0R-HC4R (the latter is so since $\tilde{h}_{ii}=1/n$ which does not depend on $i$). In the balanced case ($n=15$), UC and HC0-HC4 are all equivalent, and the same is true for UCR and HC0R-HC4R as is not difficult to see. Furthermore, in the balanced as well as in the unbalanced case, the rejection regions of the tests based on UC and UCR coincide essentially (i.e., up to a $\lambda _{\mathbb{R}^{n}}$ -null set) as a consequence of the relationship established in Section (ref). In particular, it follows that in the balanced case, all tests considered (essentially) coincide. We nevertheless compute the power functions for each of the tests separately without making use of the noted equivalencies; this provides a double-check of our numerical results.\footnote{ The equivalencies mentioned in this paragraph for the two-group-comparison problem analogously hold for general $n$, $n_{1}$, and $n_{2}$ as is easily seen.}

Numerically the critical values were determined through the implementation of Algorithms (ref) and (ref) in the R-package hrt ( hrt) version 1.0.0, and the power functions were computed with the implementation of the algorithm by davies in the R-package CompQuadForm (Duchesne) version 1.4.3; see Appendices (ref) and (ref) for more details. For the sake of illustration, we also report the critical values obtained for every test considered and every balancedness condition in Table (ref).

table[table omitted — 617 chars of source]

In relation to Table (ref) we note that the equivalences discussed before predict, e.g., that the ratio between the entries in the column labeled HC0 and the corresponding entries in the column labeled HC1 should be equal to $n/(n-2)=30/28\approx 1.0714$. The ratios computed from the table are $1.0414$, $1.0761$, and $1.0721$ (for $n_{1}=3,9,15)$, which is in pretty good agreement (especially if one converts the critical values shown in the table to critical values for the corresponding \textquotedblleft $t$ -test\textquotedblright\ versions by computing their square roots). The agreement between theoretical and observed ratios for the HC0R-HC4R columns is similar. In the balanced case one can also use the additional equivalences mentioned before and one again finds very good agreement. Similarly, the critical values for UC and UCR in Table (ref) are in excellent agreement with their theoretical relationship found in Section (ref). The reason for the small discrepancies observed lies in the fact that the algorithm underlying the computations for Table (ref) makes use of a random search algorithm. Concerning Table (ref), we also mention that, in the example considered here and for the test statistic HC2, IbragMuell2016 prove in their Theorem 1 (see also the discussion preceding that theorem) that the smallest size-controlling critical values are given by\ $18.51$ ($n_{1}=3$), $5.32$ ($n_{1}=9$ ), and $4.60$ ($n_{1}=15$), respectively. The numerically determined critical values in Table (ref) are reasonably close to these values (after conversion of the critical values to corresponding \textquotedblleft $ t$-test\textquotedblright\ critical values the maximal difference is about $ 0.1$). Of course, the accuracy of our algorithm could be increased by using more stringent accuracy parameters in the optimization routines underlying the computation of the critical value, but this would come with a longer runtime.

From Table (ref) it is clear that for the tests based on unrestricted residuals the smallest size-controlling critical values obtained are always larger, sometimes considerably, than $C_{\chi ^{2},0.05}\approx 3.8415$, again showing that the latter critical value is not effecting size control. For the tests based on restricted residuals the smallest size-controlling critical values sometimes fall below $C_{\chi ^{2},0.05}$ in the strongly unbalanced case (which is not completely surprising in view of Section (ref)); while in this case $C_{\chi ^{2},0.05}$ effects size-control, using the smaller size-controlling critical values given in Table (ref) can only be advantageous in terms of power.

That being said, we emphasize a trivial, but important point, namely that comparing the magnitudes\ of size-controlling critical values relating to different test statistics is not very meaningful and, in particular, not a valid way of comparing the quality of the resulting tests. That is, while it may be tempting to infer from Table (ref) that the HC0 test should be considerably more conservative than the HC4 test, or that the UC test should be considerably more conservative than the UCR test, such a conclusion would be false and not warranted at all (in particular, recall that UC and UCR in fact result in (essentially) the same test if the critical values from Table (ref) are being used). While this would be correct if the critical values were all meant to be used with the same test statistic (which they are not), critical values belonging to different test statistics can certainly not be compared in such a way. Instead, one has to compare the corresponding power functions, which is what we shall do next.

The power functions are shown in Figure (ref) (\textquotedblleft strongly unbalanced\textquotedblright , $n_{1}=3$), Figure (ref) (\textquotedblleft moderately unbalanced\textquotedblright , $n_{1}=9$), and Figure (ref) (\textquotedblleft balanced\textquotedblright , $ n_{1}=15$), where only the first two figures are shown in the main text, and the last figure (in which the power functions of all the feasible tests lie \textquotedblleft on top of each other\textquotedblright ) is available in Appendix (ref). Readers are referred to the online version for colored figures.

figure[figure omitted — 595 chars of source]
figure[figure omitted — 595 chars of source]

The power functions illustrate that the testing problem is getting easier, (i.e., power gets closer to the oracle benchmark), for more balanced design, which has intuitive appeal. Except for the strongly unbalanced case ($ n_{1}=3 $), the power loss of the tests based on HC0-HC4 and HC0R-HC4R relative to the oracle benchmark is surprisingly small (see Figure (ref) as well as Figure (ref) in Appendix (ref)). In the unbalanced cases ($n_{1}\in \{3,9\}$) the HC0-HC4-based tests behave all very similarly, with the power functions of the HC0- and HC1-based test being virtually indistinguishable (as they should in view of the before discussed equivalence). The UC-based test shows markedly worse power performance. Similarly, the HC0R-HC4R-based tests have virtually indistinguishable power functions (as they should because of the before discussed equivalence). The UCR-based test again is inferior (and its power function coincides with the one of UC as mentioned before). There appears also to be little difference between basing the test statistics on unrestricted or restricted residuals in this example. In the balanced case we know that all the feasible tests have exactly the same power function in view of our earlier discussion. This is visible in Figure (ref) in Appendix (ref). Also the different forms of heteroskedasticity considered seem not to have much effect on the power functions (when expressed as a function of $\delta $), except for UC and UCR in the unbalanced cases.

Hence, within the scenario considered in this section, perhaps the most important conclusion concerning the choice of a test statistic appears to be to avoid UC and UCR. Everything apart from that, i.e., whether one uses unrestricted or restricted residuals to construct the test or which specific heteroskedasticity-correction one decides to use, seems to be a comparably irrelevant part of the problem once the right (i.e., smallest size-controlling) critical value is used. We shall see in the next subsection that this conclusion very much depends on the scenario considered here and does not generalize beyond, illustrating the danger of drawing conclusions from a limited numerical study.

A high-leverage design matrix

In this section, we consider testing $\beta _{2}=0$ in a model with intercept and a single regressor $x=(10,\cos (2),\cos (3),\ldots ,\cos (n))^{\prime }$. Obviously, the regressor has a dominant first coordinate, leading to diagonal elements $h_{ii}$ of $X(X^{\prime }X)^{-1}X^{\prime }$ such that the ratio of largest to smallest $h_{ii}$ is roughly $26$ ($\max h_{ii}\simeq 0.879$, $\min h_{ii}\simeq 0.033$). Hence, the design matrix $X$ provides (on purpose) an extreme case, which leads to quite interesting results. We consider again the case $n=30$ and $\alpha =0.05$, but now show power functions for $\Sigma _{a}^{\ast }$, $a=0,\ldots ,4$, where

equation*[equation* omitted — 146 chars of source]

Note that $\Sigma _{0}^{\ast }=n^{-1}I_{n}$ and that increasing $a$ from $0$ to $4$ leads to covariance matrices that approach the degenerate matrix $ e_{1}(n)e_{1}(n)^{\prime }$. All conditions in Theorems (ref) and (ref) are seen to be satisfied in this example: As no vector $e_{i}(n)$ belongs to $\limfunc{span}(X)$ (and thus also not to $ \mathfrak{M}_{0}^{lin}$), Assumptions (ref) and (ref) as well as the sufficient condition for size control ((ref)) are obviously satisfied. The size control conditions ( (ref)) and ((ref)) have been checked numerically, as has been the condition that none of the test statistics HC0R-HC4R is constant on $\mathbb{R}^{n}\backslash \mathsf{\tilde{B}}$.

As in the preceding subsection, the critical values for each test statistic are again chosen as the smallest critical value guaranteeing size control over $\mathfrak{C}_{Het}$ and they are presented in Table (ref) below. [Existence follows from our theory since all assumptions are satisfied as noted before.] For their computation the same algorithms were used as in Section (ref), with a similar statement applying to the numerical routines used for computing the power functions. Note that the critical values for the test statistics UC, HC0-HC3 are large, reflecting the high-leverage in the design matrix; an exception is HC4, the reason being that some of the HC4-weights are considerably larger than the weights for HC0-HC3. Similarly as in the preceding subsection, the tests based on HC0 and HC1 coincide (since HC0 and HC1 differ only by a multiplicative constant and since smallest size-controlling critical values are being used), and the same is true for the tests based on HC0R-HC4R, see Remarks (ref) and (ref). It is easily checked that the ratios of the respective critical values provided in Table (ref) are in good agreement with the theoretical ratios predicted by theory. Furthermore, the tests based on UC and UCR coincide (see Section (ref)), and the critical values for UC and UCR in Table (ref) are in excellent agreement with their theoretical relationship found in Section (ref).

Table (ref) shows that in this example the smallest size-controlling critical values are -- except in one case -- always larger, sometimes considerably larger, than $C_{\chi ^{2},0.05}\approx 3.8415$, once more showing that the latter critical value is not effecting size control in general. In the exceptional case, namely when the HC4 test statistic is used, $C_{\chi ^{2},0.05}$ is considerably larger than the smallest size-controlling critical value, which is $1.12$; while in this case $ C_{\chi ^{2},0.05}$ effects size-control, using the smaller size-controlling critical value $1.12$ can only be advantageous in terms of power.

table[table omitted — 371 chars of source]

The power functions, when the size-controlling critical values from Table (ref) are being used, are shown in Figure (ref). Readers are referred to the online version for a colored figure. Again, as predicted by theory, the power functions of the tests based on HC0 and HC1 shown in Figure (ref) coincide, as do the power functions of the tests based on HC0R-HC4R; the same is true for the power functions of the tests based on UC and UCR. The figure furthermore shows that in the setting considered here, there is now a marked difference between tests based on HC0-HC4 and on HC0R-HC4R, respectively: the power of the tests based on HC0R-HC4R is nowhere greater than $\alpha $, their power function being even non-monotonic, whereas the tests based on HC0-HC4 have increasing power as a function of $\delta $. In contrast to the example considered in the preceding subsection, the power functions of the tests based on HC0-HC4 and on UC are now all markedly different and typically intersect, an exception being the case of $\Sigma _{4}^{\ast }$ where the test based on UC offers the highest power for that covariance matrix. Overall, however, there is no clear ranking between the tests using unrestricted residuals in the example considered here, although we note that the test based on UC (or, equivalently, on UCR) performs very badly in the case of $\Sigma _{0}^{\ast } $. This is not surprising as $\Sigma _{0}^{\ast }$ corresponds to homoskedasticity and the critical value used here is much larger than the classical critical value one would use given knowledge of this homoskedasticity. Furthermore, and in contrast to the results in the preceding subsection, the different forms of heteroskedasticity considered have a noticeable effect on the power functions. The main takeaway is that tests based on HC0R-HC4R (and probably on UC and UCR) should rather be avoided.

figure[figure omitted — 653 chars of source]

Conclusion

The usual heteroskedasticity robust test statistics such as $T_{Het}$ (using HC0-HC4 weights) or $\tilde{T}_{Het}$ (using HC0R-HC4R weights), used in conjunction with conventional critical values obtained from the asymptotic null distribution, are often plagued by overrejection under the null. This has been clearly documented in the literature for $T_{Het}$, and is shown numerically for $\tilde{T}_{Het}$ (as well as for $T_{Het}$) in Section (ref) above. Not surprisingly, similar observations apply to the \textquotedblleft uncorrected\textquotedblright\ test statistics $T_{uc}$ and $\tilde{T}_{uc}$. We show theoretically that all these test statistics can be size-controlled under quite weak conditions by an appropriate choice of critical values.

From the above discussion and the numerical results in Section (ref) it transpires that smallest size-controlling critical values rather than conventional critical values should be used in order to avoid the risk of overrejection. For the computation of smallest size-controlling critical values we provide algorithms which have been implemented in the R-package hrt (hrt) and thus are readily available for the user.

An additional advantage from using smallest size-controlling critical values over conventional critical values is that this typically leads to improved power in instances where conventional critical values lead to underrejection (i.e., lead to worst-case rejection probability under the null less than the nominal significance level) as is sometimes the case; see Sections (ref) and (ref).

If smallest size-controlling critical values are adopted (as they should), the numerical results in Section (ref) suggest that the test statistic $\tilde{T}_{Het}$ (with the usual weights HC0R-HC4R) should be avoided, as the resulting tests may have very poor power properties (see the example in Section (ref)). The test statistic $T_{Het}$ seems to perform better in terms of power, with no clear ranking emerging with regards to the weights HC0-HC4 being used. The \textquotedblleft uncorrected\textquotedblright\ test statistics $T_{uc}$ and $\tilde{T}_{uc}$ appear to be inferior to $T_{Het}$ in terms of power in almost all of the numerical examples considered. We also point out that -- when using smallest size-controlling critical values -- the tests based on $T_{Het}$ employing the HC0 and the HC1 weights, respectively, in fact coincide; and the same holds for tests based on $\tilde{T}_{Het}$ employing the HC0R and the HC1R weights, respectively. Also, the tests based on $T_{uc}$ and $\tilde{T}_{uc}$ then (essentially) coincide. See Remarks (ref), (ref), and Section (ref) as well as the pertaining discussion in Section (ref) for more information, including additional equivalencies when the design matrix $X$ and the restriction $R$ have certain special properties.