Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
190,864 characters · 24 sections · 145 citation commands
Valid Heteroskedasticity Robust Testing
Testing hypotheses on the parameters in a regression model with potentially heteroskedastic errors is an important problem in econometrics and statistics; see Mackinnon2013 for a recent survey. Since the classical $t$-statistic ($F$-statistic, respectively) is not pivotal, or asymptotically pivotal, in such a case in general, even under Gaussianity of the errors, so-called heteroskedasticity robust (aka heteroskedasticity consistent) modifications of these test statistics have been proposed, which are asymptotically standard normally (chi-square, respectively) distributed under the null. These modifications date back to E63,E67, see also H77, and have later been popularized in econometrics by W80 with great success (see Mackinnon2013). Unfortunately, it turned out that tests obtained from these heteroskedasticity robust test statistics by relying on critical values derived from the respective asymptotic null distributions have a tendency to overreject the null hypothesis in finite samples (and thus are invalid), especially so if the design matrix contains high-leverage points; see, e.g., MacW85, DavidsonMacKinnon1985 , and CheshJewitt1987. One factor contributing to this overrejection tendency is a downward bias in the covariance matrix estimators used in these test statistics, see CheshJewitt1987. In an attempt to reduce the overrejection problem, variants of the before-mentioned heteroskedasticity robust test statistics (often denoted by HC1 through HC4, with HC0 denoting the original proposal) have been considered; see H77 , MacW85, and Crib2004.\footnote{ For a recent contribution geared towards high-dimensional models see Catt.} These variants rescale the least-squares residuals before computing the covariance matrix estimator employed in the construction of the test statistic. According to simulation studies reported in, e.g., DavidsonMacKinnon1985 and Crib2004, these modifications, especially HC3 and HC4, seem to ameliorate the overrejection problem to some extent, but do not eliminate it. Further numerical results are provided in CheshAust_1991, see also Chesh_1989. Numerical results in Section (ref) confirm these observations. Variants of HC0-HC3, denoted by HC0R-HC3R, obtained by using restricted instead of unrestricted least-squares residuals in the computation of the covariance matrix estimators employed by the various test statistics (the restriction alluded to being the restriction defining the null hypothesis) have been introduced in DavidsonMacKinnon1985. In their simulation experiments, this typically leads to tests that do not overreject, but that may substantially underreject; see also the simulation results in Godfrey2006, who additionally also considers HC4R. However, as will be shown in Section (ref), also these tests are in general not immune to (sometimes substantial) overrejection.
Note that, under the typical assumptions used in the literature, all the modifications of HC0 discussed so far have the same asymptotic distribution as HC0, and thus the same critical value as for HC0 (obtained from the asymptotic null distribution) is also used for these modifications in the before mentioned literature. Sometimes small-sample adjustments to the asymptotic critical values are attempted by using the quantiles from a $ t_{d} $-distribution rather than from the asymptotic normal distribution, where the degrees of freedom $d$ are either set to $n-k$ ($n$ and $k$ denoting sample size and number of regressors, respectively), or are obtained through proposals set down by Satterth or BellMcCa; see also Imbkoles2016. While these adjustments can lead to improvements, numerical results presented in Section (ref) show that these adjustments are also not able to solve the overrejection problem in general. An alternative approach is to use bootstrap methods to compute critical values for the test statistics HC0-HC4 or HC0R-HC4R. The relevant literature is reviewed in PPBoot, and it is shown that such methods are again not immune to the overrejection problem in general.\footnote{ Another possibility is to use Edgeworth expansions to find better critical values, see Rothenberg1988 for the case of the HC0 test statistic and DavidsonMacKinnon1985 for the HC0R test statistic. Simulation results in MacW85 and DavidsonMacKinnon1985 indicate that this does not work too well in practice. Of course, such expansions could also be worked out for the other versions of the test statistics mentioned, but this does not seem to have been pursued in the literature.} A referee has pointed out the recent papers by Chuetal2021 and Hansen2021, both of which propose a testing procedure that can be viewed as a parametric bootstrap method. No theoretical justification is given in those papers. In fact, as we show in Appendix (ref), the proposed procedures can be considerably oversized, a feature that can already be seen to some extent in the numerical results given in Chuetal2021 and Hansen2021.
A result by BakiSzek2005 needs to be mentioned here which states that -- in the special case of testing a hypothesis on the location parameter of a heteroskedastic location model with errors that are Gaussian or scale mixtures thereof -- the classical two-sided $t$-test (with the usual critical value) has null rejection probability not exceeding the nominal significance level under any form of heteroskedasticity (for a certain range of significance levels); see IbragMuell2010 for more discussion. IbragMuell2016, extending a result in MickeyBrown1966, provide a related result in the case of the comparison of two heteroskedastic populations; see also Bakirov98. Section (ref) provides some more discussion. We note that all results mentioned in that section are applicable only to testing certain scalar linear contrasts.
Except for the BakiSzek2005 result and the variations discussed in Section (ref), which apply only to quite special situations like, e.g., the heteroskedastic location model, none of the methods discussed so far comes with a theoretical result implying that their associated (finite sample) null rejection probabilities are guaranteed not to exceed the nominal significance level whatever the form of heteroskedasticity may be. \footnote{ In the special case where the number of restrictions tested equals the number of regression parameters,\ DavidsFlach2008 have a result which implies that certain wild bootstrap-based heteroskedasticity robust tests have size equal to the nominal significance level (and hence do not overreject) in finite samples. We note that this result in DavidsFlach2008 is not entirely correct as stated, but needs some amendments and corrections; see PPBoot.} In fact, it transpires from the preceding discussion and the numerical results in Section (ref) that for any of these methods instances of testing problems can be found for which the method in question overrejects substantially. Therefore, it is imperative to be able to find size-controlling critical values for the test statistics considered, i.e., critical values such that the resulting worst-case rejection probability under the null hypothesis does not exceed the nominal significance level. We shall hence pursue in this paper the construction of size-controlling critical values for the test statistics HC0-HC4, HC0R-HC4R, as well as for (two variants of) the classical (i.e., uncorrected) $F$-statistic (including the absolute value of the $t$ -statistic as a special case).
In the present paper we consider classes of test statistics that contain the before mentioned heteroskedasticity robust test statistics as special cases and show under which conditions -- and how -- a critical value can be found such that the resulting test is guaranteed to have size less than or equal to $\alpha $, the prescribed significance level.\footnote{ A less principled attempt at finding a valid test in a given testing problem (i.e., for given design matrix and restriction to be tested) could consist in the practitioner studying the size of a handful of tests (obtained from a few of the above mentioned test statistics in conjunction with a few of the proposed critical values) by means of an extensive Monte Carlo study and in hoping that one of the test procedures emerges from this study as valid for the particular testing problem at hand. Besides being a numerically costly procedure, it does not come with any guarantee of success. } It turns out that the conditions for size controllability are broadly satisfied; in particular, for the commonly used test statistics they are satisfied generically in a sense made precise further below.
We want to emphasize that the existence of size-controlling critical values for heteroskedasticity robust test statistics is not a trivial matter, as it has been shown in PP2016, Section 4, that there are cases where the size of such tests is always one, regardless of the choice of critical value; see also the discussion in Proposition (ref) further below. And even in cases where size control is possible by an appropriate choice of critical value, the standard critical values proposed in the literature (including the small-sample adjustments discussed above) are not guaranteed to deliver size control; in fact, they may fail to do so by a considerable margin (i.e., they are much too small to control size at the desired level) as shown in Section (ref). Our theoretical results also show the existence of a computable "threshold" $C^{\ast }$, say, such that any critical value $C$ satisfying $ C<C^{\ast }$ necessarily leads to a test with size $1$; see Proposition (ref). Since $C^{\ast }$ is not difficult to compute, it can be used as a simple check to weed out unsuitable proposals for critical values.
Apart from avoiding overrejection by construction, the use of smallest size-controlling, rather than conventional, critical values offers also advantages in terms of power in instances where conventional critical values lead to underrejection (i.e., lead to a worst-case rejection probability under the null hypothesis less than the nominal significance level) as is sometimes the case; see Sections (ref) and (ref). In fact, once one has decided on a test statistic to be used for the given null hypothesis, using the smallest size-controlling critical value (provided it exists) is obviously the optimal way to proceed.
We also discuss how the critical values that lead to size control can be determined numerically and provide the R-package hrt (hrt) for their computation. The usefulness of the proposed algorithms and their implementation in the R-package are illustrated numerically on some testing problems in Section (ref). In particular, we compare tests obtained from various of the above mentioned test statistics when used with smallest size-controlling critical values in terms of their power functions. The package hrt also contains a routine for determining the size of a test obtained from a user-supplied critical value. It is important to note that if in a particular application one uses the observed value of the test statistic as the user-supplied critical value in this routine, this routine actually returns a \textquotedblleft valid p-value\textquotedblright\ in the following sense: Checking whether or not this \textquotedblleft p-value\textquotedblright\ is smaller than the prescribed significance level $\alpha $ is equivalent to checking whether or not the observed value of the test statistic is larger than or equal to the smallest size-controlling critical value. Note that the former check avoids the need to actually compute the smallest size-controlling critical value, which is advantageous from a computational point of view. See Section (ref) for more details.
In the paper we work under a Gaussianity assumption. We stress, however, that this assumption is mainly made for convenience of presentation; as shown in Section (ref), this assumption can be relaxed considerably.
While a trivial remark, we would like to note that the size control results given in this paper can easily be translated into results stating that the minimal coverage probability of the associated confidence set obtained by \textquotedblleft inverting\textquotedblright\ the test is not less than the nominal confidence level.
The paper is organized as follows: After introducing notation and the most important test statistics in Sections (ref) and (ref) , Section (ref) provides some intuition for our size-control results which are presented in Sections (ref) and (ref), with some further results relegated to Appendix (ref). Section (ref) discusses ways of relaxing the underlying assumptions. Possible extensions to other classes of test statistics are discussed in Section (ref), while a few comments on power are collected in Section (ref). Section (ref) provides the numerical results including a power study, with some details relegated to Appendix (ref). Section (ref) concludes. Proofs and some technical results can be found in Appendices (ref)-(ref). The algorithms for computing rejection probabilities (including size) and smallest size-controlling critical values are outlined in Section (ref), and are presented in detail in Appendix (ref). Appendix (ref) contains a discussion of Chuetal2021 and Hansen2021.
Consider the linear regression model
where $X$ is a (real) nonstochastic regressor (design) matrix of dimension $ n\times k$ and where $\beta \in \mathbb{R}^{k}$ denotes the unknown regression parameter vector. We always assume $\limfunc{rank}(X)=k$ and $ 1\leq k<n$. We furthermore assume that the $n\times 1$ disturbance vector $ \mathbf{U}=(\mathbf{u}_{1},\ldots ,\mathbf{u}_{n})^{\prime }$ has mean zero and unknown covariance matrix $\sigma ^{2}\Sigma $, where $\Sigma $ varies in a user-specified (nonempty) set $\mathfrak{C}$ describing the allowed forms of heteroskedasticity, with $\mathfrak{C}$ satisfying $\mathfrak{C} \subseteq \mathfrak{C}_{Het}$, and where $0<\sigma ^{2}<\infty $ holds ($ \sigma $ always denoting the positive square root).\footnote{ Since we are concerned with finite-sample results only, the elements of $ \mathbf{Y}$, $X$, and $\mathbf{U}$ (and even the probability space supporting $\mathbf{Y}$ and $\mathbf{U}$) may depend on sample size $n$, but this will not be expressed in the notation. Furthermore, the obvious dependence of$\ \mathfrak{C}$ on $n$ will also not be shown in the notation.} The set $\mathfrak{C}$ will be referred to as the \textquotedblleft heteroskedasticity model\textquotedblright . Here
where $\limfunc{diag}(\tau _{1}^{2},\ldots ,\tau _{n}^{2})$ denotes the $ n\times n$ matrix with diagonal elements given by $\tau _{i}^{2}$. That is, the errors in the regression model are uncorrelated but can be heteroskedastic. In particular, if $\mathfrak{C}$ is chosen to be $\mathfrak{ C}_{Het}$, one allows for heteroskedasticity of completely unknown form. The normalization condition $\sum_{i=1}^{n}\tau _{i}^{2}=1$ is included here only in order to guarantee identifiability of $\sigma ^{2}$ and $\Sigma $, and could be replaced by any other normalization condition such as, e.g., $ \max \tau _{i}^{2}=1$, or $\tau _{1}^{2}=1$, without affecting the final results (because any of these normalizations leads to the same overall set of covariance matrices $\sigma ^{2}\Sigma $ when $\sigma ^{2}$ varies through the positive real line). Although a trivial observation, we stress the fact that all conceivable forms of heteroskedasticity, including parametric ones, can (possibly after normalization) be cast as submodels $ \mathfrak{C}$ of $\mathfrak{C}_{Het}$.
Mainly for ease of exposition, we shall maintain in the sequel that the disturbance vector $\mathbf{U}$\ is normally distributed. This assumption can be substantially relaxed as discussed in Section (ref). The linear model described in ((ref)), together with the just made Gaussianity assumption on $\mathbf{U}$ and with the given heteroskedasticity model $\mathfrak{C}$, then induces a collection of distributions on the Borel-sets of $\mathbb{R}^{n}$, the sample space of $ \mathbf{Y}$. Denoting a Gaussian probability measure with mean $\mu \in \mathbb{R}^{n}$ and (possibly singular) covariance matrix $A$ by $P_{\mu ,A}$ , the induced collection of distributions is then given by
where $\mathrm{\limfunc{span}}(X)$ denotes the column space of $X$. Since every $\Sigma \in \mathfrak{C}$ is positive definite by assumption, each element of the set in the previous display is absolutely continuous with respect to (w.r.t.) Lebesgue measure on $\mathbb{R}^{n}$.
We shall consider the problem of testing a linear (better: affine) hypothesis on the parameter vector $\beta \in \mathbb{R}^{k}$, i.e., the problem of testing the null $R\beta =r$ against the alternative $R\beta \neq r$, where $R$ is a $q\times k$ matrix always of rank $q\geq 1$ and $r\in \mathbb{R}^{q}$. Set $\mathfrak{M}=\limfunc{span}(X)$. Define the affine space
and let
Adopting these definitions, this testing problem can then be written more precisely as
With $\mathfrak{M}_{0}^{lin}$ we shall denote the linear space parallel to $ \mathfrak{M}_{0}$, i.e., $\mathfrak{M}_{0}^{lin}=\mathfrak{M}_{0}-\mu _{0}=\left\{ X\beta :R\beta =0\right\} $ where $\mu _{0}\in \mathfrak{M}_{0}$ . Of course, $\mathfrak{M}_{0}^{lin}$ does not depend on the choice of $\mu _{0}\in \mathfrak{M}_{0}$.
As already mentioned, the assumption of Gaussianity is made mainly for simplicity of presentation and can be relaxed substantially; see Section (ref). The assumption of nonstochastic regressors entails little loss of generality either, which is important to emphasize: If $X$ is random and $\mathbf{U}$ is conditionally on $X$ distributed as $N(0,\sigma ^{2}\Sigma )$, with $\sigma ^{2}=\sigma ^{2}(X)>0$ and $\Sigma =\Sigma (X)\in \mathfrak{C}_{Het}$, the results of the paper can be applied after one conditions on $X$ (and a similar statement applies to the generalizations to non-Gaussianity discussed in Section (ref) ). See Section (ref) for more discussion and details. For arguments supporting conditional inference see, e.g., RO1979. Note that such a \textquotedblleft strict exogeneity\textquotedblright\ assumption is quite natural in the situation considered here.
We next collect some further terminology and notation used throughout the paper. A (nonrandomized) test is the indicator function of a Borel-set $W$ in $\mathbb{R}^{n}$, with $W$ called the corresponding rejection region. The size of such a test (rejection region) is -- as usual -- defined as the supremum over all rejection probabilities under the null hypothesis $H_{0}$ given in ((ref)), i.e.,
In slight abuse of terminology, we shall sometimes refer to this quantity as `the size of $W$ over $\mathfrak{C}$' when we want to emphasize the r\^{o}le of $\mathfrak{C}$. Throughout the paper we let $\hat{\beta}(y)=\left( X^{\prime }X\right) ^{-1}X^{\prime }y$, where $X$ is the design matrix appearing in ((ref)) and $y\in \mathbb{R}^{n}$. The corresponding ordinary least-squares (OLS) residual vector is denoted by $\hat{u}(y)=y-X \hat{\beta}(y)$ and its elements are denoted by $\hat{u}_{t}(y)$. The elements of $X$ are denoted by $x_{ti}$, while $x_{t\cdot }$ and $x_{\cdot i} $ denote the $t$-th row and $i$-th column of $X$, respectively. For $ \mathcal{A}$ an affine subspace of $\mathbb{R}^{n}$ satisfying $\mathcal{A} \subseteq \limfunc{span}(X)$ let $\tilde{\beta}_{\mathcal{A}}(y)$ denote the restricted least-squares estimator, i.e., $X\tilde{\beta}_{\mathcal{A}}(y)$ solves
Lebesgue measure on the Borel-sets of $\mathbb{R}^{n}$ will be denoted by $ \lambda _{\mathbb{R}^{n}}$, whereas Lebesgue measure on an arbitrary affine subspace $\mathcal{A}$ of $\mathbb{R}^{n}$ (but viewed as a measure on the Borel-sets of $\mathbb{R}^{n}$) will be denoted by $\lambda _{\mathcal{A}}$, with zero-dimensional Lebesgue measure being interpreted as point mass. The set of real matrices of dimension $l\times m$ is denoted by $\mathbb{R} ^{l\times m}$ (all matrices in the paper will be real matrices) and Lebesgue measure on this set equipped with its Borel $\sigma $-field is denoted by $ \lambda _{\mathbb{R}^{l\times m}}$. Let $B^{\prime }$ denote the transpose of a matrix $B\in \mathbb{R}^{l\times m}$ and let $\mathrm{\limfunc{span}} (B) $ denote the subspace in $\mathbb{R}^{l}$ spanned by its columns. For a symmetric and nonnegative definite matrix $B$ we denote the unique symmetric and nonnegative definite square root by $B^{1/2}$. For a linear subspace $ \mathcal{L}$ of $\mathbb{R}^{n}$ we let $\mathcal{L}^{\bot }$ denote its orthogonal complement and we let $\Pi _{\mathcal{L}}$ denote the orthogonal projection onto $\mathcal{L}$. The Euclidean norm is denoted by $\left\Vert \cdot \right\Vert $, but the same symbol is also used to denote a norm of a matrix. The $j$-th standard basis vector in $\mathbb{R}^{n}$ is written as $ e_{j}(n)$. Furthermore, we let $\mathbb{N}$ denote the set of all positive integers. A sum (product, respectively) over an empty index set is to be interpreted as $0$ ($1$, respectively). Finally, for $\mathcal{A}$ an affine subspace of $\mathbb{R}^{n}$, let $G(\mathcal{A})$ denote the group of all affine transformations $y\mapsto \delta (y-a)+a^{\ast }$ where $\delta \in \mathbb{R}$, $\delta \neq 0$, and $a$ as well as $a^{\ast }$ are elements of $\mathcal{A}$; for more information see Section 5.1 of PP2016.
We now introduce two test statistics that will feature prominently in the following. Variants thereof that use restricted residuals are discussed in Section (ref). For results pertaining to other classes of test statistics see Section (ref). The test statistic we shall consider first is a standard heteroskedasticity robust test statistic frequently encountered in the literature. It is given by
where $\hat{\Omega}_{Het}=R\hat{\Psi}_{Het}R^{\prime }$ and where $\hat{\Psi} _{Het}$ is a heteroskedasticity robust estimator as considered in E63,E67, which later on has found its way into the econometrics literature (e.g., W80). It is of the form
where the constants $d_{i}>0$ sometimes depend on the design matrix. Typical choices for $d_{i}$ suggested in the literature are $d_{i}=1$, $ d_{i}=n/(n-k) $, $d_{i}=\left( 1-h_{ii}\right) ^{-1}$, or $d_{i}=\left( 1-h_{ii}\right) ^{-2}$ where $h_{ii}$ denotes the $i$-th diagonal element of the projection matrix $X(X^{\prime }X)^{-1}X^{\prime }$, see LE2000 for an overview. Another suggestion is $d_{i}=\left( 1-h_{ii}\right) ^{-\delta _{i}}$ for $\delta _{i}=\min (nh_{ii}/k,4)$, see Crib2004. For the last three choices of $d_{i}$ just given, we use the convention that we set $d_{i}=1$ in case $h_{ii}=1$. Note that $h_{ii}=1$ implies $\hat{u} _{i}\left( y\right) =0$ for every $y$, and hence it is irrelevant which real value is assigned to $d_{i}$ in case $h_{ii}=1$.\footnote{ In fact, $h_{ii}=1$ is equivalent to $\hat{u}_{i}\left( y\right) =0$ for every $y$, each of which in turn is equivalent to $e_{i}(n)\in $ $\limfunc{ span}(X)$.} The five examples for the weights $d_{i}$ just given correspond to what is often called HC0-HC4 weights in the literature.
In conjunction with the test statistic $T_{Het}$, we shall consider the following mild assumption, which is Assumption 3 in PP2016. As discussed further below, this assumption is in a certain sense unavoidable when using $T_{Het}$. It furthermore also entails that our choice of assigning $T_{Het}\left( y\right) $ the value zero in case $\hat{\Omega} _{Het}\left( y\right) $ is singular has no import on the probabilistic results of the paper (because of Lemma (ref)(c) below and absolute continuity of the measures $P_{\mu ,\sigma ^{2}\Sigma }$).
Observe that this assumption only depends on $X$ and $R$ and hence can be checked. Obviously, a simple sufficient condition for Assumption (ref) to hold is that $s=0$ (i.e., that $e_{j}(n)\notin \limfunc{span} (X) $ for all $j$), a generically satisfied condition. Furthermore, we introduce the matrix
The facts collected in the subsequent lemma, which is taken from PPBoot (but see also Lemma 4.1 in PP2016 and Lemma 5.18 in PP3), will be used in the sequel.
In light of Part (c) of the lemma, we see that Assumption (ref) is a natural and unavoidable condition if one wants to obtain a sensible test from $T_{Het}$.\footnote{ If this assumption is violated then $T_{Het}$ is identically zero, an uninteresting trivial case.} Furthermore, note that, if $\mathsf{B}=\limfunc{ span}(X)$ is true, then Assumption (ref) must be satisfied (since $ \limfunc{span}(X)$ is a $\lambda _{\mathbb{R}^{n}}$-null set due to the maintained assumption $k<n$). As shown in Lemma A.3 in PP3, for any given restriction matrix $R$, the relation $\mathsf{B}=\limfunc{span}(X)$ holds generically in various universes of design matrices. For later use we also mention that under Assumption (ref) the test statistic $T_{Het}$ is continuous at every $y\in \mathbb{R}^{n}\backslash \mathsf{B}$.\footnote{ If Assumption (ref) is violated, then $T_{Het}$ is constant equal to zero, and hence is trivially continuous everywhere.}
Next, we also consider the classical (i.e., uncorrected) F-test statistic, i.e.,
where $\hat{\sigma}^{2}(y)=\hat{u}\left( y\right) ^{\prime }\hat{u}\left( y\right) /(n-k)\geq 0$ (which vanishes if and only if $y\in \limfunc{span}(X) $). Our choice to set $T_{uc}(y)=0$ for $y\in \limfunc{span}(X)$ again has no import on the probabilistic results in the paper, since $\limfunc{span}(X) $ is a $\lambda _{\mathbb{R}^{n}}$-null set as a consequence of the maintained assumption that $k<n$ (and since the measures $P_{\mu ,\sigma ^{2}\Sigma }$ are absolutely continuous). For reasons of comparability with ( (ref)) we have chosen not to normalize the numerator in ((ref) ) by $q$, the number of restrictions to be tested, as is often done in the definition of the classical F-test statistic. This also has no import on the results as the factor $1/q$ can be absorbed into the critical value. For later use we also mention that the test statistic $T_{uc}$ is continuous at every $y\in \mathbb{R}^{n}\backslash \limfunc{span}(X)$.
We begin the heuristic discussion by considering the testing problem ((ref)) with heteroskedasticity model $\mathfrak{C}=\mathfrak{C} _{Het}$ (i.e., heteroskedasticity of unknown form). Let $T$ stand for any of the test statistics introduced in Section (ref), with rejection occurring whenever $T\geq C$, $C$ a critical value.\footnote{ In case of $T=T_{Het}$ Assumption (ref) is supposed to hold.}$^{ \text{,}}$\footnote{ The discussion similarly applies to the test statistics introduced in Section (ref).}. For simplicity of presentation we assume $r=0$. As discussed in Section (ref), basing the test on the conventional critical value $C_{\chi ^{2}(q),0.05}$ (the 95% quantile of a chi-square distribution with $q$ degrees of freedom) often leads to substantial overrejection, i.e., the size of the test (over $\mathfrak{C} _{Het}$) is substantially larger than the desired value $\alpha =0.05$. One mechanism leading to such overrejection is constituted by a concentration phenomenon discussed at some length in PP2016: In the present situation, the distribution $P_{0,\sigma ^{2}\Sigma }$ \textquotedblleft concentrates\textquotedblright\ on a so-called concentration subspace (given by $\limfunc{span}(e_{i}(n))$) when $\Sigma $ is \textquotedblleft close\textquotedblright\ to one of the singular matrices $ e_{i}(n)e_{i}(n)^{\prime }$.\footnote{ There are also other concentration subspaces in the present situation which we can ignore for the heuristic discussion.} In such a case, depending on the design matrix $X$ and the hypothesis given by $(R,r)$, the concentration space may fall into the rejection region $\{T\geq C_{\chi ^{2}(q),0.05}\}$, leading to a rejection probability close to one, and thus much larger than $ \alpha =0.05$.\footnote{ This is an oversimplified description ignoring some technical details.} Even if the concentration subspace $\limfunc{span}(e_{i}(n))$ is not contained in the rejection region, but is sufficiently close to it, a considerable portion of the mass of $P_{0,\sigma ^{2}\Sigma }$ may nevertheless fall into the rejection region if $\Sigma $ is close to, but not too close to $ e_{i}(n)e_{i}(n)^{\prime }$. This again leads to a relatively large rejection probability. Overrejection will often be especially pronounced if certain high-leverage points are present in the design matrix.\footnote{ We note, however, that there are testing problems (e.g., testing the mean in a heteroskedastic location model using the test statistic $T_{uc}$) for which the text-book critical values obtained under homoskedasticity are actually valid, see BakiSzek2005. The reason is that the \textquotedblleft worst case\textquotedblright\ distribution in this case corresponds to homoskedasticity.}
In order to obtain a test that has size controlled by $\alpha $ (i.e., size $ \leq \alpha $) in situations as just described, the rejection region $ \{T\geq C_{\chi ^{2}(q),0.05}\}$ has to be narrowed down, i.e., $C_{\chi ^{2}(q),0.05}$ has to be replaced by a suitably larger critical value $C$. Whether or not this can successfully be accomplished by a (finite) $C$, is a non-trivial question, the answer depending on whether or not all possible concentration subspaces can be made to fall outside of the rejection region $ \{T\geq C\}$ by an appropriate choice of $C$ larger than $C_{\chi ^{2}(q),0.05}$. Sufficient conditions when this is possible are provided in Theorems (ref) and (ref). Note that, in such a situation, the resulting size-controlling critical values $C$ are then necessarily larger than $C_{\chi ^{2}(q),0.05}$.
In light of the preceding discussion, a natural question is whether or not imposing a heteroskedasticity model more narrow than $\mathfrak{C}_{Het}$ such as, e.g.,
where $\tau _{\ast }$, $0<\tau _{\ast }<n^{-1/2}$, is a pre-specified constant set by the user, would mitigate the failure of conventional critical values. Indeed, under the heteroskedasticity model $\mathfrak{C} _{Het,\tau _{\ast }}$ extreme concentration effects leading to rejection probabilities (arbitrarily) close to one cannot occur, and it is possible to prove that size-controlling critical values always exist when $\mathfrak{C} _{Het,\tau _{\ast }}$ is used, see Appendix (ref). Unfortunately, however, this does not imply that conventional critical values such as $C_{\chi ^{2}(q),0.05}$ will work. In fact, the size over $\mathfrak{C} _{Het,\tau _{\ast }}$ of tests using the critical value $C_{\chi ^{2}(q),0.05}$ can still be considerably larger than $\alpha $: To see this, observe that the sets $\mathfrak{C}_{Het,\tau _{\ast }}$are an increasing sequence of sets as $\tau _{\ast }\downarrow 0$, the union of which is $ \mathfrak{C}_{Het}$. Consequently, if $\tau _{\ast }$ is small, the size over $\mathfrak{C}_{Het,\tau _{\ast }}$ will be close to the size over $ \mathfrak{C}_{Het}$, and thus the former will be much larger than $\alpha $ in case the latter is so. As a consequence, also in case of the more narrow heteroskedasticity model $\mathfrak{C}_{Het,\tau _{\ast }}$ size-controlling critical values larger than $C_{\chi ^{2}(q),0.05}$ will have to be used in such a case. Furthermore, the bound $\tau _{\ast }$ has to be decided upon prior to the data analysis and is thus part of modeling the form of heteroskedasticity. It is difficult to see how one would come up with a reasonable value of $\tau _{\ast }$ in practice: If $\tau _{\ast }$ is chosen to be small, this may result in a heteroskedasticity model under which the test based on $C_{\chi ^{2}(q),0.05}$ is still plagued by overrejection as just discussed, while choosing $\tau _{\ast }$ large will typically not be defensible as it presumes considerable knowledge about the admissible forms of heteroskedasticity.
We introduce the following notation: For a given linear subspace $\mathcal{L} $ of $\mathbb{R}^{n}$ we define the set of indices $I_{0}(\mathcal{L})$ via
We set $I_{1}(\mathcal{L})=\left\{ 1,\ldots ,n\right\} \backslash I_{0}( \mathcal{L})$. Clearly, $\func{card}(I_{0}(\mathcal{L}))\leq \dim (\mathcal{L })$ holds. In particular, if $\dim (\mathcal{L})<n$ holds (which, in particular, is so in the leading case $\mathcal{L}=\mathfrak{M}_{0}^{lin}$, since $\dim (\mathfrak{M}_{0}^{lin})=k-q<n$), then $\func{card}(I_{0}( \mathcal{L}))<n$, and thus $\func{card}(I_{1}(\mathcal{L}))\geq 1$.
We have the following size control result for $T_{uc}$ as well as for $ T_{Het}$ over the heteroskedasticity model $\mathfrak{C}_{Het}$ (more precisely, over the null hypothesis $H_{0}$ described in ((ref)) with $\mathfrak{C}=\mathfrak{C}_{Het}$). Note that $\mathfrak{C} _{Het}$ is the largest possible heteroskedasticity model and reflects complete ignorance about the form of heteroskedasticity.
We see from the theorem that the condition for size control of $T_{Het}$ ($ T_{uc}$, respectively) over $\mathfrak{C}_{Het}$, i.e., condition ((ref)) (((ref)), respectively), only depends on $X$ and $R$; in particular, in case of $T_{Het}$, it does not depend on how the weights $d_{i}$ figuring in the definition of $T_{Het}$ have been chosen (note that the set $\mathsf{B}$ only depends on $X$ and $R$). Moreover, the sufficient conditions for size control are generically satisfied in the universe of all $n\times k$ design matrices $X$ (of rank $k$), see Example (ref) and the attending discussion further below. Furthermore, it is plain that the size-controlling critical values $C(\alpha )$ in Theorem (ref) will depend on the choice of test statistic as well as on the testing problem at hand. More concretely, the size-controlling critical values in Part (b) of the theorem thus depend only on $X$, $R$, and $r$, as well as on the choice of weights $d_{i}$, whereas in Part (a) the dependence is only on $X$, $R$, and $r$. We do not show these dependencies in the notation. In fact, as discussed in Remark (ref) below, it turns out that the size-controlling critical values in both cases actually do not depend on the value of $r$ at all (provided the weights $d_{i}$ are not allowed to depend on $r$ in case of $T_{Het}$). Similarly, it is easy to see that $C^{\ast }$ and $\alpha ^{\ast }$ in Theorem (ref) do not depend on $r$ (under the same provision as before in case of $T_{Het}$).
Another observation is that any critical value delivering size control over $ \mathfrak{C}_{Het}$ also delivers size control over any other heteroskedasticity model $\mathfrak{C}$ since $\mathfrak{C}\subseteq \mathfrak{C}_{Het}$. Of course, for such a $\mathfrak{C}$ even smaller critical values (than needed for $\mathfrak{C}_{Het}$) may already suffice for size control. Also note that sufficient conditions implying size control over $\mathfrak{C}_{Het}$ may be more restrictive than sufficient conditions implying only size control over a smaller heteroskedasticity model $ \mathfrak{C}$. For size control results tailored to such smaller models $ \mathfrak{C}$ see Appendix (ref).
In light of the results of CheshJewitt1987 and Chesh_1989, it is useful to interpret the sufficient conditions for size control, i.e., ( (ref)) and ((ref)), in terms of high-leverage points. First, note that $e_{i}(n)\in \limfunc{span}(X)$ is equivalent to $h_{ii}=1$, which corresponds to the $i$-th observation being an \textquotedblleft extreme high-leverage point\textquotedblright . Hence, ( (ref)) is equivalent to $h_{ii}<1$ for every $i\in \mathcal{I}_{1}(\mathfrak{M}_{0}^{lin})$. In other words, the condition for a size-controlling critical value to exist in Part (a) of Theorem (ref) requires that none of the indices in $\mathcal{I}_{1}( \mathfrak{M}_{0}^{lin})$ corresponds to an extreme high-leverage point. [It is interesting to observe that all indices in $\mathcal{I}_{0}(\mathfrak{M} _{0}^{lin})$ (note that this set may be empty) correspond to extreme high-leverage points.] Hence, for the condition in ((ref) ) not to be satisfied, not only must extreme high-leverage points be present, but the lever needs to be of a particular type depending on the hypothesis given by $(R,r)$ (namely, it must have $i\in \mathcal{I}_{1}( \mathfrak{M}_{0}^{lin})$). Second, note that a sufficient, but not necessary, condition for ((ref)) is $h_{ii}<1$ for $ i=1,\ldots ,n$. Sufficiency is obvious from the preceding discussion. That the condition is not necessary can be seen from Example (ref) further below. Finally, condition ((ref)) implies condition ((ref)) (since $\limfunc{span}(X)\subseteq \mathsf{B}$), and hence implies $h_{ii}<1$ for every $i\in \mathcal{I}_{1}(\mathfrak{M} _{0}^{lin})$. The converse is not always true: even $h_{ii}<1$ for every $ i=1,\ldots ,n$ does not guarantee ((ref)) to be satisfied, see Example (ref) further below. However, generically ((ref)) and ((ref)) coincide (see Lemma A.3 in PP3), in which case the discussion given above for ((ref)) also applies to ((ref)).
The next proposition complements Theorem (ref) and provides a useful lower bound for the size-controlling critical values (other than the trivial bound given in the preceding remark).
The preceding observation is useful in two ways: First, critical values suggested in the literature (such as, e.g., the $(1-\alpha )$-quantile of a chi-square distribution with $q$ degrees of freedom or critical values obtained from a degree of freedom adjustment) can immediately be dismissed if they turn out to be less than $C^{\ast }$, as they then certainly will not guarantee size control.\footnote{ In contrast, if the critical value turns out to be larger than or equal to $ C^{\ast }$, it does not follow that size is less than or equal to $ \alpha $. In fact, substantially oversized tests using a critical value $ C>C^{\ast }$ are certainly possible; see, e.g., Table (ref) and the pertaining discussion.} We use this line of reasoning in the numerical results in Section (ref). Second, if the observed value of the test statistic $T_{Het}$ ($T_{uc}$, respectively) is less than $C^{\ast } $, the decision not to reject the null hypothesis can be taken without further need to compute size-controlling critical values. Note that $C^{\ast }$ as given in Theorem (ref) is quite easy to compute in any given application.
We next discuss to what extent the sufficient conditions for size control in Theorem (ref) are also necessary.
We illustrate Theorem (ref) and Proposition (ref) with a few examples.
This example shows, in particular, that the sufficient conditions for size control are generically satisfied in the universe of all $n\times k$ design matrices $X$ (of rank $k$). Given the example, this is obvious for $T_{uc}$; and it follows for $T_{Het}$ by additionally noting that, for every given choice of restriction to be tested, the relation $\mathsf{B}=\func{span}(X)$ holds generically in the universe of all $n\times k$ design matrices $X$ (of rank $k$); see Lemma A.3 in PP3. The next example discusses the case where a standard basis vector is among the regressors.
We continue with a few more examples where $X$ has a particular structure.
The subsequent example is closely related to the Behrens-Fisher problem, see Remark (ref) in Appendix (ref).
The next example is an extension of the previous problem to the case of more than two groups. An interesting phenomenon occurs here: The sufficient conditions for size control of $T_{Het}$ given in Theorem (ref) are violated, but size controllability can nevertheless be established by additional arguments. Hence, this example provides an instance where the conditions in Part (b) of Theorem (ref) are not necessary.
We close this section by one more example. Again, the sufficient conditions in Part (b) of Theorem (ref) fail to hold, but additional arguments based on Example (ref) establish size controllability of the test based on $T_{Het}$.
In Appendix (ref) we discuss yet another example where the sufficient condition of Part (b) of Theorem (ref) fails, but size-controllability can nevertheless be established.
(i) As noted in IbragMuell2010, testing a hypothesis regarding a scalar linear contrast in a heteroskedastic (Gaussian) linear regression model more general than a location model can often be converted to a testing problem in a heteroskedastic (Gaussian) location model by suitably dividing the data into subgroups and by considering groupwise least-squares estimators, thus making it amenable to the BakiSzek2005 result mentioned in Section (ref). However, this introduces additional questions such as how to divide up the data. In any case, this approach is limited to testing hypotheses on scalar linear contrasts. It also requires that the linear contrast subject to test is estimable in each subgroup.
(ii) In case the linear contrast subject to test is not estimable in each subgroup, but can be written as the difference of two linear contrasts where the first contrast is estimable in the first $G_{1}$ groups whereas the second contrast is estimable in the last $G_{2}$ groups (where we consider a total of $G_{1}+G_{2}$ groups), IbragMuell2016 point out that the problem can be converted into the problem of comparing two heteroskedastic (Gaussian) populations. Now, for such a two-sample comparison problem Bakirov98 shows for a certain two-sample $t$ -statistic (the square of which is $T_{uc}$, cf. Example (ref) above) how -- in the presence of heteroskedasticity -- size-controlling critical values can be constructed by appropriately transforming quantiles of a $t$-distribution; this result imposes conditions which entail that the nominal significance level $\alpha $ must be quite small (requiring $\alpha $ not to exceed $0.01$ for many group sizes, and often to be considerably smaller). This somewhat limits the applicability of Bakirov's result. Thus IbragMuell2016 go on to consider another two-sample $t$-statistic (the square of which is $T_{Het}$ with $d_{i}=\left( 1-h_{ii}\right) ^{-1}$, cf. Example (ref) above) and -- extending a result in MickeyBrown1966 -- provide a BakiSzek2005-type result, i.e., they show that the $(1-\alpha /2)$-quantile of a $t$-distribution with degrees of freedom equal to the smaller of the two sample sizes minus $1$ provides the smallest size-controlling critical value (for the two-sided test) even under heteroskedasticity.\footnote{ In the balanced case (i.e., if the two samples have the same cardinality) the test statistic considered in Bakirov98 actually coincides with the test statistic in IbragMuell2016.} This result holds under certain conditions on the sample sizes and only for small $\alpha $, but, e.g., allows for the choice $\alpha =0.05$. [We note here that the description of Bakirov98's result in IbragMuell2016 is inaccurate in that a certain transformation of the critical value is being ignored.]
(iii) In the problem of comparing two heteroskedastic (Gaussian) populations based on samples of equal size (\textquotedblleft balanced design\textquotedblright ) one can -- instead of using the two-sample $t$ -test statistics considered in Bakirov98 and IbragMuell2016 -- employ the Bartlett test statistic, which simply is the usual $t$-test statistic computed from the differences between the observations in the two samples.\footnote{ Certainly, there is some arbitrariness in how the observations are being \textquotedblleft paired\textquotedblright .} An advantage of this approach is that the original BakiSzek2005 result is directly applicable, and there is no need to resort to the results described in (ii).
(iv) Another quite special case that can be brought under the realm of the BakiSzek2005 result is a heteroskedastic (Gaussian) regression model with only one regressor that never takes the value zero. Dividing the $t$-th equation in the regression model by $x_{t}$, converts this into a heteroskedastic location problem.
(v) The results in (i)-(iv) immediately also apply if the errors in the regression are distributed as scale mixtures of Gaussians (cf. also Section (ref)).
In this section we consider two further test statistics which are versions of $T_{Het}$ and $T_{uc}$ with the only difference that the covariance matrix estimators used are based on restricted -- instead of unrestricted -- residuals. The first one of these test statistics has been suggested in the literature, e.g., in DavidsonMacKinnon1985. We thus define
where $\tilde{\Omega}_{Het}=R\tilde{\Psi}_{Het}R^{\prime }$ and where $ \tilde{\Psi}_{Het}$ is given by
where the constants $\tilde{d}_{i}>0$ sometimes depend on the design matrix and on the restriction matrix $R$. Here $\tilde{u}\left( y\right) =y-X\tilde{ \beta}_{\mathfrak{M}_{0}}(y)=\Pi _{(\mathfrak{M}_{0}^{lin})^{\bot }}(y-\mu _{0})$, where the last expression does not depend on the choice of $\mu _{0}\in \mathfrak{M}_{0}$, and where $\tilde{u}_{t}\left( y\right) $ denotes the $t$-th component of $\tilde{u}\left( y\right) $. Typical choices for $ \tilde{d}_{i}$ are $\tilde{d}_{i}=1$, $\tilde{d}_{i}=n/(n-(k-q))$, $\tilde{d} _{i}=(1-\tilde{h}_{ii})^{-1}$, or $\tilde{d}_{i}=(1-\tilde{h}_{ii})^{-2}$ where $\tilde{h}_{ii}$ denotes the $i$-th diagonal element of the projection matrix $\Pi _{\mathfrak{M}_{0}^{lin}}$, see, e.g., DavidsonMacKinnon1985. Another suggestion is $\tilde{d}_{i}=(1-\tilde{h} _{ii})^{-\tilde{\delta}_{i}}$ for $\tilde{\delta}_{i}=\min (n\tilde{h} _{ii}/(k-q),4)$ with the convention that $\tilde{\delta}_{i}=0$ if $k=q$. \footnote{ Note that in case $k=q$ we have $\tilde{h}_{ii}=0$, and hence $\tilde{d} _{i}=1$ regardless of our convention for $\tilde{\delta}_{i}$.} For the last three choices of $\tilde{d}_{i}$ just given we use the convention that we set $\tilde{d}_{i}=1$ in case $\tilde{h}_{ii}=1$. Note that $\tilde{h} _{ii}=1 $ implies $\tilde{u}_{i}\left( y\right) =0$ for every $y$, and hence it is irrelevant which real value is assigned to $\tilde{d}_{i}$ in case $ \tilde{h}_{ii}=1$.\footnote{ In fact, $\tilde{h}_{ii}=1$ is equivalent to $\tilde{u}_{i}\left( y\right) =0 $ for every $y$, each of which in turn is equivalent to $e_{i}(n)\in $ $ \mathfrak{M}_{0}^{lin}$.} The five examples for the weights $\tilde{d}_{i}$ just given correspond to what is often called HC0R-HC4R weights in the literature.\footnote{In the case $k=q$ the HC0R-HC4R weights all coincide ($\tilde{d}_{i}=1$ for every $i$), and hence result in the same test statistic.}
The subsequent assumption ensures that the set of $y$'s for which $\tilde{ \Omega}_{Het}\left( y\right) $ is singular is a Lebesgue null set, implying that our choice of assigning $\tilde{T}_{Het}\left( y\right) $ the value zero in case $\tilde{\Omega}_{Het}\left( y\right) $ is singular has no import on the probabilistic results of the paper (as the measures $P_{\mu ,\sigma ^{2}\Sigma }$ are absolutely continuous). Also, as discussed further below, the assumption is in a certain sense unavoidable when using $\tilde{T} _{Het}$.
Observe that this assumption only depends on $X$ and $R$ and hence can be checked. Obviously, a simple sufficient condition for Assumption (ref) to hold is that $s=0$ (i.e., that $e_{j}(n)\notin \mathfrak{M }_{0}^{lin}$ for all $j$), a generically satisfied condition. Furthermore, we introduce the matrix
Note that this matrix does not depend on the choice of $\mu _{0}\in \mathfrak{M}_{0}$. The following lemma collects some important properties of $\tilde{\Omega}_{Het}$ and $\mathsf{\tilde{B}}$ (defined in that lemma) and is reproduced from PPBoot for ease of reference.
In light of Part (c) of the lemma, we see that Assumption (ref) is a natural and unavoidable condition if one wants to obtain a sensible test from $\tilde{T}_{Het}$.\footnote{ If this assumption is violated then $\tilde{T}_{Het}$ is identically zero, an uninteresting trivial case.} Furthermore, note that if $\mathsf{\tilde{B}} =\mathfrak{M}_{0}$ is true, then Assumption (ref) must be satisfied (since $\mathfrak{M}_{0}$ is a $\lambda _{\mathbb{R}^{n}}$-null set as $k-q<n$ is always the case). For later use we also mention that under Assumption (ref) the statistic $\tilde{T}_{Het}$ is continuous at every $y\in \mathbb{R}^{n}\backslash \mathsf{\tilde{B}}$.\footnote{ If Assumption (ref) is violated, then $\tilde{T}_{Het}$ is constant equal to zero, and hence trivially continuous everywhere.}
We finally consider for completeness, and in analogy with $T_{uc}$,
where $\tilde{\sigma}^{2}(y)=\tilde{u}\left( y\right) ^{\prime }\tilde{u} \left( y\right) /(n-(k-q))\geq 0$ (which vanishes if and only if $y\in \mathfrak{M}_{0}$). Of course, our choice to set $\tilde{T}_{uc}(y)=0$ for $ y\in \mathfrak{M}_{0}$ again has no import on the probabilistic results in the paper, since $\mathfrak{M}_{0}$ is a $\lambda _{\mathbb{R}^{n}}$-null set (and since the measures $P_{\mu ,\sigma ^{2}\Sigma }$ are absolutely continuous). For later use we also mention that $\tilde{T}_{uc}$ is continuous at every $y\in \mathbb{R}^{n}\backslash \mathfrak{M}_{0}$. As we shall see in Section (ref), there is a close connection between $\tilde{T}_{uc}$ and $T_{uc}$.
Here we discuss size control results for $\tilde{T}_{uc}$ as well as for $ \tilde{T}_{Het}$ over the heteroskedasticity model $\mathfrak{C}_{Het}$ (more precisely, over the null hypothesis $H_{0}$ described in ((ref)) with $\mathfrak{C}=\mathfrak{C}_{Het}$). Some peculiar properties of the test statistics $\tilde{T}_{uc}$ and $\tilde{T}_{Het}$ are then discussed in the following section.
We note that the first statement in Part (a) of the subsequent theorem is actually trivial, since $\tilde{T}_{uc}$ is bounded as shown in the next section (which also provides a discussion when non-trivial size-controlling critical values exist).
We see from the theorem that $\tilde{T}_{uc}$ is always size controllable over $\mathfrak{C}_{Het}$, but as discussed in Section (ref) below there is a caveat: Unless ((ref)), i.e., the necessary and sufficient condition for size-controllability of $T_{uc}$, is satisfied, size-controlling $\tilde{T}_{uc}$ leads to trivial tests. We also see that the condition for size control of $\tilde{T}_{Het}$ over $\mathfrak{ C}_{Het}$, i.e., condition ((ref)) is always satisfied in case $\mathsf{\tilde{B}}=\mathfrak{M}_{0}$ (since ((ref)) is then equivalent to $e_{i}(n)\notin \mathfrak{M}_{0}^{lin}$ for every $ i\in I_{1}(\mathfrak{M}_{0}^{lin})$). Furthermore, condition ((ref)) always only depends on $X$ and $R$; in particular, it does not depend on how the weights $\tilde{d}_{i}$ figuring in the definition of $\tilde{T}_{Het}$ have been chosen (note that $\mu _{0}+e_{i}(n)\notin \mathsf{\tilde{B}}$ is equivalent to $e_{i}(n)\notin \mathsf{\tilde{B}}-\mu _{0}$ and that the set $\mathsf{\tilde{B}}-\mu _{0}$ depends only on $X$ and $R$). Furthermore, the size-controlling critical values $C(\alpha )$ in Part (b) of the preceding theorem depend only on $X$, $R$, and $r$, as well as on the choice of weights $\tilde{d}_{i}$, whereas in Part (a) the dependence is only on $X$, $R$, and $r$. We do not show these dependencies in the notation. In fact, as shown in Lemma (ref) in Appendix (ref), it turns out that the size and the size-controlling critical values in both cases actually do not depend on the value of $r$ at all (provided the weights $\tilde{d}_{i}$ are not allowed to depend on $r$ in case of $\tilde{T}_{Het}$). Similarly, it is easy to see that $\alpha ^{\ast }$ and $C^{\ast }$ do not depend on $r$ (under the same provision as before in case of $\tilde{T}_{Het}$).
Similarly as in Section (ref), a critical value delivering size control over $\mathfrak{C}_{Het}$ also delivers size control over any other heteroskedasticity model $\mathfrak{C}$ since $\mathfrak{C} \subseteq \mathfrak{C}_{Het}$. Of course, for such a $\mathfrak{C}$ even smaller critical values (than needed for $\mathfrak{C}_{Het}$) may already suffice for size control. Also note that sufficient conditions implying size control over $\mathfrak{C}_{Het}$ may be more restrictive than sufficient conditions only implying size control over a smaller heteroskedasticity model $\mathfrak{C}$. For size control results tailored to such smaller models $\mathfrak{C}$ see Appendix (ref).
The next proposition complements Theorem (ref) and provides a lower bound for the size-controlling critical values (other than the trivial bound given in the preceding remark). The lower bound is useful for the same reasons as discussed subsequent to Proposition (ref).
For the test statistic $T_{uc}$ the rejection regions $\{T_{uc}\geq C\}$, as well as their complements, have positive ($n$-dimensional) Lebesgue measure for every positive real number $C$.\footnote{ The case $C\leq 0$ is uninteresting as the rejection region of $T_{uc}$ (and of all other test statistics considered) then are the entire space $\mathbb{R }^{n}$, since $T_{uc}$ (and the other test statistics considered) take on only nonnegative values.} This follows from Parts 5&6 of Lemma 5.15 in PP2016 together with Remark (ref) in Appendix (ref). As a consequence, all rejection probabilities -- under the null as well as under the alternative -- are positive and less than one regardless of the choice of $C>0$. [This is so because of our Gaussianity assumption and the fact that all $\Sigma \in \mathfrak{C}_{Het}$ are positive definite.] For similar reasons, the same is true for $T_{Het}$ provided Assumption (ref) is satisfied.\footnote{ If Assumption (ref) is not satisfied then $T_{Het}\equiv 0$, and the resulting test (with rejection region $\{T_{Het}\geq C\}$) is trivial as it never rejects for $C>0$, while it always rejects for $C\leq 0$.} The situation is somewhat different for tests derived from $\tilde{T}_{uc}$ or $ \tilde{T}_{Het}$ as we shall discuss next. In the course of this, we also establish a connection between $T_{uc}$ and $\tilde{T}_{uc}$ that is of independent interest. In this section the size of a test always refers to size over $\mathfrak{C}_{Het}$.
First, observe that $\tilde{T}_{uc}(y)\leq n-(k-q)$ holds for every $y\in \mathbb{R}^{n}$ and that this bound is sharp. To see this, note that using standard least-squares theory
for $y\notin \mathfrak{M}_{0}$ and that $\tilde{T}_{uc}(y)=0$ else; the bound is attained precisely for $y\in \limfunc{span}(X)\backslash \mathfrak{M }_{0}$. An immediate consequence of this observation is that any critical value $C\geq (n-(k-q))$ leads to a test with rejection region $\{\tilde{T} _{uc}\geq C\}$ that is either empty (if $C>n-(k-q)$) or is a $\lambda _{ \mathbb{R}^{n}}$-null set, namely $\limfunc{span}(X)\backslash \mathfrak{M} _{0}$ (if $C=n-(k-q)$). Consequently, such a test is trivial in that all rejection probabilities (under the null as well as under the alternative) are zero (because of our Gaussianity assumption and the fact that all $ \Sigma \in \mathfrak{C}_{Het}$ are positive definite). As an aside we note that any $C<n-(k-q)$ leads to a non-trivial test as is easily seen.
Of course, a critical value $C$ satisfying $C\geq n-(k-q)$ is certainly size-controlling, but is useless since it leads to a trivial test as just discussed. We now ask if and when the smallest size-controlling critical value $C_{\Diamond }(\alpha )$, guaranteed to exist by Part (c) of Theorem (ref), leads to a non-trivial test. [This is certainly so if $\alpha ^{\ast }$ in Part (a) of Theorem (ref) is positive, but note that the theorem is silent on this issue.] To obtain insight, we establish a simple, but important, relationship between the test statistics $\tilde{T}_{uc}$ and $T_{uc}$ that is of independent interest also: Note that standard least-squares theory gives
for $y\notin \limfunc{span}(X)$, and recall $T_{uc}(y)=0$ for $y\in \limfunc{ span}(X)$. Hence, we obtain
for every $y\notin \limfunc{span}(X)$, where $g:[0,\infty )\rightarrow \lbrack 0,n-(k-q))$ is continuous and strictly increasing with $ \lim_{x\rightarrow \infty }g(x)=(n-(k-q))$. [Since $T_{uc}(y_{m})\rightarrow \infty $ for every sequence $y_{m}\rightarrow y\in \limfunc{span} (X)\backslash \mathfrak{M}_{0}$, the sharpness of the bound $n-(k-q)$ can thus also be read-off from ((ref)).] As a consequence, for every critical value $C>0$, the rejection regions $\{\tilde{T}_{uc}\geq C\}$ and $ \{T_{uc}\geq g^{-1}(C)\}$ differ at most by $\limfunc{span}(X)$, which is a $ \lambda _{\mathbb{R}^{n}}$-null set; in particular, the rejection probabilities (under the null as well as under the alternative) are the same. \footnote{ This is so because of our Gaussianity assumption and the fact that all $ \Sigma \in \mathfrak{C}_{Het}$ are positive definite.} That is, the test statistics $\tilde{T}_{uc}$\ and $T_{uc}$\ give rise to (essentially) the same test, if the critical values chosen are linked by the function $g$\ as above. In particular, as we shall see, this is the case if the respective smallest size-controlling critical values are used for both test statistics (provided both these values exist).
To see what the preceding discussion entails for the existence of non-trivial size-controlling critical values for $\tilde{T}_{uc}$ we distinguish two cases. In the first case we shall see that non-trivial size-controlling critical values do not exist, whereas in the second case they do indeed exist.
Case 1: Condition ((ref)) is violated. Recall from Proposition (ref) that then the size of $\{T_{uc}\geq D\}$ is $1$ for every real $D$ (in particular, implying that $T_{uc}$ is not size controllable). It transpires from the preceding discussion, that hence the size of $\{\tilde{T}_{uc}\geq C\}$ must equal $1$ for every $C$ satisfying $ 0<C<n-(k-q)$ (and a fortiori for $C\leq 0$), because $D:=g^{-1}(C)$ is well-defined and real for $0<C<n-(k-q)$. As a consequence, any size-controlling critical value $C$ for $\tilde{T}_{uc}$ must satisfy $C\geq n-(k-q)$ (with the smallest size-controlling critical value given by $ n-(k-q) $), thus leading to a rejection region that is trivial in that it is empty (if $C>n-(k-q)$) or is a $\lambda _{\mathbb{R}^{n}}$-null set, namely $ \limfunc{span}(X)\backslash \mathfrak{M}_{0}$ (if $C=n-(k-q)$). That is -- while $\tilde{T}_{uc}$ is size-controllable in the present case -- it is so only in a trivial way.\footnote{ The trivial size-controlling critical values $C$ for $\tilde{T}_{uc}$ sort of correspond to using $\infty $ as a \textquotedblleft size-controlling critical value\textquotedblright\ for $T_{uc}$.} [Another way of arriving at the above conclusion is to use Part (a) of Proposition (ref) and to observe that in Part (a) of Theorem (ref) the quantity $C^{\ast }$ equals $n-(k-q)$. To see the latter, note that violation of condition ((ref)) implies existence of an index $i\in I_{1}(\mathfrak{M}_{0}^{lin})$ with $e_{i}(n)\in \limfunc{span} (X)$. In particular, $\hat{u}(\mu _{0}+e_{i}(n))=0$. Since $e_{i}(n)\notin \mathfrak{M}_{0}^{lin}$ must hold in view of $i\in I_{1}(\mathfrak{M} _{0}^{lin})$, and thus $\mu _{0}+e_{i}(n)\notin \mathfrak{M}_{0}$ for every $ \mu _{0}\in \mathfrak{M}_{0}$ must be true, we may use ((ref)) to arrive at $\tilde{T}_{uc}(\mu _{0}+e_{i}(n))=n-(k-q)$ for this $i\in I_{1}( \mathfrak{M}_{0}^{lin})$. This shows $C^{\ast }\geq n-(k-q)$. Equality then follows since $C^{\ast }\leq n-(k-q)$ trivially holds by ((ref)). As a point of interest we also note that $C^{\ast }=n-(k-q)$ implies that $ \alpha ^{\ast }$ in Part (a) of Theorem (ref) satisfies $ \alpha ^{\ast }=0$.]
Case 2: Condition ((ref)) is satisfied. In this case $T_{uc}$ is size controllable according to Theorem (ref). In particular, for any given $\alpha \in (0,1)$ there exists a smallest real number $D_{\Diamond }(\alpha )$ such that the size of $\{T_{uc}\geq D_{\Diamond }(\alpha )\}$ is less than or equal to $\alpha $, with equality holding for $\alpha \in (0,\alpha _{T_{uc}}^{\ast }]\cap (0,1)$ where $ \alpha _{T_{uc}}^{\ast }$ refers to $\alpha ^{\ast }$ appearing in Theorem (ref)(a), and recall from that theorem that $\alpha _{T_{uc}}^{\ast }>0$; and $D_{\Diamond }(\alpha )>0$ by Remark (ref).\footnote{If $\alpha _{T_{uc}}^{\ast }<\alpha <1$ , then the size, in fact, equals $\alpha _{T_{uc}}^{\ast }$; see Remark (ref).} Also note that the rejection region $\{T_{uc}\geq D_{\Diamond }(\alpha )\}$ is not trivial as it has positive $\lambda _{ \mathbb{R}^{n}}$-measure (and the same is true for its complement); see the discussion at the very beginning of Section (ref). Setting $ C_{\Diamond }(\alpha )=g(D_{\Diamond }(\alpha ))$ and using that $\{\tilde{T} _{uc}\geq C_{\Diamond }(\alpha )\}$ and $\{T_{uc}\geq g^{-1}(C_{\Diamond }(\alpha ))\}=\{T_{uc}\geq D_{\Diamond }(\alpha )\}$ differ at most by the $ \lambda _{\mathbb{R}^{n}}$-null set $\limfunc{span}(X)$, we see that (i) $ 0<C_{\Diamond }(\alpha )<n-(k-q)$, (ii) the size of $\{\tilde{T}_{uc}\geq C_{\Diamond }(\alpha )\}$ is less than or equal to $\alpha $, with equality holding for $\alpha \in (0,\alpha _{T_{uc}}^{\ast }]\cap (0,1)$, (iii) $ C_{\Diamond }(\alpha )$ is the smallest size-controlling critical value (recall that $g$ is strictly increasing), and (iv) the rejection region $\{ \tilde{T}_{uc}\geq C_{\Diamond }(\alpha )\}$ is not trivial as it has positive $\lambda _{\mathbb{R}^{n}}$-measure (and the same is true for its complement). In particular, note that $\tilde{T}_{uc}$ and $T_{uc}$ give rise to (essentially) the same test if the respective smallest size-controlling critical values are used. We furthermore note that in the present situation $C_{\tilde{T}_{uc}}^{\ast }=g(C_{T_{uc}}^{\ast })$ and $ \alpha _{\tilde{T}_{uc}}^{\ast }=\alpha _{T_{uc}}^{\ast }$ hold, where $ C_{T_{uc}}^{\ast }$, $\alpha _{T_{uc}}^{\ast }$ correspond to $C^{\ast }$, $ \alpha ^{\ast }$ in Part (a) of Theorem (ref), whereas $C_{ \tilde{T}_{uc}}^{\ast }$, $\alpha _{\tilde{T}_{uc}}^{\ast }$ correspond to $ C^{\ast }$, $\alpha ^{\ast }$ in Part (a) of Theorem (ref).\footnote{ If $\alpha _{\tilde{T}_{uc}}^{\ast }<\alpha <1$, then the size of $\{\tilde{T }_{uc}\geq C_{\Diamond }(\alpha )\}$ is, in fact, equal to $\alpha _{T_{uc}}^{\ast }=\alpha _{\tilde{T}_{uc}}^{\ast }$; cf. Footnote (ref) and Remark (ref).} In particular, $ \alpha _{\tilde{T}_{uc}}^{\ast }>0$ and $0\leq C_{\tilde{T}_{uc}}^{\ast }<n-(k-q)$ follow. These claims can be seen as follows: Under condition ((ref)) we have $\mu _{0}+e_{i}(n)\notin \limfunc{span}(X)$ for every $i\in I_{1}(\mathfrak{M}_{0}^{lin})$ and every $\mu _{0}\in \mathfrak{M}_{0}$. Consequently, $\tilde{T}_{uc}(\mu _{0}+e_{i}(n))=g(T_{uc}(\mu _{0}+e_{i}(n)))$, which proves $C_{\tilde{T} _{uc}}^{\ast }=g(C_{T_{uc}}^{\ast })$ in view of strict monotonicity of $g$. The relation $\alpha _{\tilde{T}_{uc}}^{\ast }=\alpha _{T_{uc}}^{\ast }$ then follows from the definitions of $\alpha _{\tilde{T}_{uc}}^{\ast }$ and $ \alpha _{T_{uc}}^{\ast }$ using that $\{\tilde{T}_{uc}\geq C\}$ and $ \{T_{uc}\geq g^{-1}(C)\}$ differ at most by the $\lambda _{\mathbb{R}^{n}}$ -null set $\limfunc{span}(X)$ for every $C>0$. Positivity of $\alpha _{ \tilde{T}_{uc}}^{\ast }$ now follows from positivity of $\alpha _{T_{uc}}^{\ast }$ discussed before, and $C_{\tilde{T}_{uc}}^{\ast }<n-(k-q)$ follows since $C_{\tilde{T}_{uc}}^{\ast }=g(C_{T_{uc}}^{\ast })$ and $ C_{T_{uc}}^{\ast }<\infty $. [Another way of proving $\alpha _{\tilde{T} _{uc}}^{\ast }>0$ and $0\leq C_{\tilde{T}_{uc}}^{\ast }<n-(k-q)$ without using relationship ((ref)), is to first establish $C_{\tilde{T} _{uc}}^{\ast }<n-(k-q)$ (from observing that $\hat{u}(\mu _{0}+e_{i}(n))\neq 0$ (as $\mu _{0}+e_{i}(n)\notin \limfunc{span}(X)$) for every $i\in I_{1}( \mathfrak{M}_{0}^{lin})$, which implies $\tilde{T}_{uc}(\mu _{0}+e_{i}(n))<n-(k-q)$ for every such $i$ in view of ((ref))) and then to proceed analogously as in the proof of Theorem (ref) below.]
While $\tilde{T}_{uc}$ is always size-controllable, whereas $T_{uc}$ is not, this does not represent any real advantage of $\tilde{T}_{uc}$ over $T_{uc}$ , as we have seen that $\tilde{T}_{uc}$ admits only trivial size-controlling critical values in the case where $T_{uc}$ is not size-controllable. Even more importantly, and already noted above, these test statistics give rise to (essentially) the same test if for both test statistics the respective smallest size-controlling critical values are used (provided they both exist).
For $\tilde{T}_{Het}$ we find that, not infrequently, it is also a bounded function, although we have no proof that this is always so. We illustrate the problems that can arise here first by an example. See also Remark (ref).
In the preceding example any critical value $C\geq 2\tilde{d}_{1}^{-1}$ is trivially a size-controlling critical value for the given significance level $\alpha $ ($0<\alpha <1$), but it is \textquotedblleft too large\textquotedblright\ and leads to a trivial test. Certainly, one would prefer to use the smallest size-controlling critical value $C_{\Diamond }(\alpha )$ instead (which in the preceding example exists by Theorem (ref) and by what has been shown in the example) and one would hope that the resulting test is not trivial. As we shall show, this is indeed the case. To this end we first give a general result that, in particular, is applicable to the preceding example. Recall that $C_{\Diamond }(\alpha )$ is positive (Remark (ref)), and that Theorem (ref) is silent on whether $\alpha ^{\ast }>0$ or not.
While the situation in Example (ref) is somewhat particular, the example may perhaps contribute to a better understanding of the Monte Carlo findings in DavidsonMacKinnon1985 and Godfrey2006, namely that the tests, obtained from $\tilde{T}_{Het}$ (employing HC0R-HC4R weights) in conjunction with conventional critical values such as the $95\%$-quantile of a chi-square distribution with appropriate degrees of freedom, can suffer from severe underrejection under the null.
(i) All results in the preceding sections (as well as the extensions described in Appendix (ref)) referring to properties under the null hypothesis carry over as they stand to the situation where the error term $ \mathbf{U}$ in ((ref)) is elliptically symmetric distributed and has no atom at zero, i.e., $\mathbf{U}$ is distributed as $\sigma \Sigma ^{1/2} \mathbf{z}$ where $\mathbf{z}$ has a spherically symmetric distribution on $ \mathbb{R}^{n}$ that has no atom at zero.\footnote{ Note that all results in the preceding sections (as well as the extensions in Appendix (ref)), except for a few comments in Section (ref), are results referring to properties under the null hypothesis, } This is so since -- under this distributional model -- the null rejection probabilities of any $G(\mathfrak{M}_{0})$-invariant rejection region coincide with the corresponding null rejection probabilities under the Gaussian model (i.e., where $\mathbf{z}$ is standard Gaussian); see the discussion in Section 5.5 of PP2016 and Appendix E.1 of PP3. \footnote{ Note that all rejection regions considered in the preceding sections are $G( \mathfrak{M}_{0})$-invariant, because the test statistics considered are so.} This implies, in particular, not only that the sufficient conditions for size controllability under the above elliptically symmetric distributed model as well as under the Gaussian model are the same, but that also the numerical values of the size-controlling critical values coincide. As a consequence, the algorithms for computing the size-controlling critical values in the Gaussian case (used in Section (ref) and described in Section (ref) and Appendix (ref)) can be used in the above elliptically symmetric distributed case without any change whatsoever. The same is actually true if $\mathbf{z}$ has a distribution in a certain class larger than the class of spherical symmetric distributions with no atom at zero, see Appendix E.1 of PP3.
(ii) Furthermore, as discussed in detail in Appendix E.2 of PP3, the sufficient conditions for size controllability that we have derived under Gaussianity also imply size controllability for many more forms of distribution of $\mathbf{z}$ than those mentioned in (i); however, the corresponding size-controlling critical values may then differ from the size-controlling critical values that apply under Gaussianity.
(iii) Similarly as in Section 5.5 of PP2016, the negative results given in the preceding sections (as well as the ones described in Appendix (ref)) such as, e.g., size $1$ results, extend in a trivial way beyond the Gaussian model as long as the maintained assumptions on the feasible error distributions are weak enough to ensure that the implied (possibly semiparametric) model, i.e., set of distributions for $\mathbf{Y}$, contains the set given in ((ref)), but possibly contains also other distributions.
(iv) A further generalization beyond Gaussianity in the important special case where $\mathfrak{C}=\mathfrak{C}_{Het}$ is as follows: Suppose $\mathbf{ U}$ is distributed as $\sigma \Sigma ^{1/2}\limfunc{diag}(\mathbf{r)z}$ where $\mathbf{z}$ is standard normally distributed on $\mathbb{R}^{n}$ and where the $n$-dimensional random vector $\mathbf{r}$ is independent of $ \mathbf{z}$ with distribution $\rho $, where $\rho $ is a distribution on $ (0,\infty )^{n}$. [This includes the case where the elements of $\limfunc{ diag}(\mathbf{r)z}$ form an i.i.d. sample from a scale mixture of normals.] Let $Q_{\mu ,\sigma ^{2}\Sigma ,\rho }$ denote the implied distribution for $ \mathbf{Y}$ given by ((ref)) where $\mu =X\beta $. Consider now instead of ((ref)) the (semiparametric) model given by all distributions $Q_{\mu ,\sigma ^{2}\Sigma ,\rho }$ where $\mu \in \mathrm{\limfunc{span}}(X)$, $ 0<\sigma ^{2}<\infty $, $\Sigma \in \mathfrak{C}$, and $\rho $ is an arbitrary distribution on $(0,\infty )^{n}$. Then the sufficient conditions for size controllability derived under Gaussianity in earlier sections (and in Appendix (ref)) also imply size controllability in this larger model. In fact, the size-controlling critical values that apply under Gaussianity deliver also size control under this more general model. This follows from the following reasoning: Let $W$ be a Borel set in $\mathbb{R} ^{n}$ such that $P_{\mu _{0},\sigma ^{2}\Sigma }(W)\leq \alpha $ for every $ \mu _{0}\in \mathfrak{M}_{0}$, every $0<\sigma ^{2}<\infty $, and every $ \Sigma \in \mathfrak{C}_{Het}$. Then for every such $\mu _{0}$, $\sigma ^{2}$ , $\Sigma $, and every distribution $\rho $ on $(0,\infty )^{n}$ we have
where $\Sigma _{\mathbf{r}}^{1/2}:=\Sigma ^{1/2}\limfunc{diag}(\mathbf{r)/} s_{\mathbf{r}}$ with $s_{\mathbf{r}}$ denoting the positive square root of the sum of the diagonal elements of $(\Sigma ^{1/2}\limfunc{diag}(\mathbf{r)) }^{2}=\Sigma \limfunc{diag}^{2}(\mathbf{r)}$ and where $\sigma _{\mathbf{r} }=\sigma s_{\mathbf{r}}$. Here we have used that $P_{\mu ,\sigma _{\mathbf{r} }^{2}\Sigma _{\mathbf{r}}}(W)\leq \alpha $ by assumption since $\Sigma _{ \mathbf{r}}=\Sigma \limfunc{diag}^{2}(\mathbf{r)/}s_{\mathbf{r}}^{2}\in \mathfrak{C}_{Het}$ and $0<\sigma _{\mathbf{r}}<\infty $ hold for every realization of $\mathbf{r}$. In the above $\Pr $ denotes the probability measure governing $(\mathbf{r},\mathbf{z)}$ and $\mathbb{E}$ the corresponding expectation operator. \ As a consequence, the smallest size-controlling critical value under Gaussianity is also the smallest size-controlling critical value under the semiparametric model considered here, as the latter model contains the Gaussian model as a submodel. [In the special case where $\limfunc{diag}(\mathbf{r)}$ is a (random) multiple of the identity matrix $I_{n}$, the assumption $\mathfrak{C}=\mathfrak{C}_{Het}$ is superfluous as then $\Sigma _{\mathbf{r}}=\Sigma $, which by assumption belongs to the given $\mathfrak{C}$. In this case $\mathbf{U}$ satisfies the assumptions in (i), and hence (iv) adds little new, except that -- in contrast to (i) -- the reasoning works without use of $G(\mathfrak{M}_{0})$ -invariance.]
(v) It is apparent from the reasoning in (iv) that Gaussianity of $\mathbf{z} $ can be replaced by any other distributional assumption for which size controllability has already been established. E.g., one can in (iv) choose $ \mathbf{z}$ to have a spherically symmetric distribution without an atom at zero or to have a distribution in the more general class mentioned in (i) (note that all relevant rejection regions discussed in earlier sections are $ G(\mathfrak{M}_{0})$-invariant and thus (i) applies). In a similar vein, one can combine the results in Appendix E.2 of PP3 discussed in (ii) above with the reasoning outlined in (iv). We abstain from presenting details.
The assumption of nonstochastic regressors can be easily relaxed as follows: Suppose $X$ is random and $\mathbf{U}$ is conditionally on $X$ distributed as $N(0,\sigma ^{2}\Sigma )$, with $\sigma ^{2}=\sigma ^{2}(X)>0$ and $ \Sigma =\Sigma (X)\in \mathfrak{C}_{Het}$ where $\sigma ^{2}(\cdot )$ and $ \Sigma (\cdot )$ may vary in given classes of functions. The size control results such as Theorems (ref) and (ref) can then obviously be applied after one conditions on $X$ provided almost all realizations of $X$ satisfy the assumptions of those theorems, which will typically be the case (for brevity we do not provide a formal statement here).\footnote{ An appropriately modified statement applies to the size control results in Appendix (ref).} The resulting conditional size control statements then immediately imply that the so-obtained conditional size-controlling critical values $C=C(\alpha ,X)$ also control size unconditionally. Size $1$ results such as, e.g., Propositions (ref), (ref), or (ref) also extend to conditional size $1$ results in a similar manner provided $\sigma ^{2}(X)$ and $\Sigma (X)$ vary independently through all of $(0,\infty )$ and $\mathfrak{C}_{Het}$, respectively, for (almost) every realization of $X$, when the functions $\sigma ^{2}(\cdot )$ and $ \Sigma (\cdot )$ vary in the before mentioned function classes.\footnote{ See Footnote 40 in PPBoot for a discussion of sufficient conditions.} Generalizations to non-Gaussianity similarly as discussed in Section (ref) are also possible in the present context.
The results in Sections (ref) and (ref) (and in Appendix (ref)) have been obtained with the help of a general theory developed in Section 5 of PP2016, Section 5 of PP3, and Section 3.1 of PP4 that covers a very broad class of test statistics (and actually allows also for correlated errors). We note that, like in Section (ref), Gaussianity is again not essential for a good portion of this general theory, see Section 5.5 of PP2016 as well as Appendix E of PP3.\footnote{ Also arguments like in (iv) and (v) of Section (ref) can be applied to try to obtain generalizations.} We next discuss a few further situations that can also be handled by the general theory just mentioned but we refrain from spelling out the details:\footnote{ Applying some of the main results of this general theory (e.g., Corollary 5.6 or Proposition 5.12 of PP3) will require one to determine the set $\mathbb{J}(\mathcal{L},\mathfrak{C})$ defined in Appendix (ref). For the important cases $\mathfrak{C}=\mathfrak{C}_{Het}$ and $\mathfrak{C}= \mathfrak{C}_{(n_{1},\ldots ,n_{m})}$ (defined in Appendix (ref)), this is already accomplished in Propositions (ref) and (ref) in Appendix (ref) below.}
(i) The test statistic considered is an OLS-based test statistic like $ T_{Het}$, but where $\hat{\Omega}_{Het}$ is now replaced by an appropriate estimator derived from a given (possibly misspecified) parametric heteroskedasticity model described by a parameter vector $\theta $.
(ii) The test statistic is a Wald-type test statistic based on a (feasible) generalized least-squares estimator together with an appropriate covariance matrix estimator based on a given (possibly misspecified) parametric model. [This includes the (quasi-)maximum likelihood estimator (provided $\theta $ is unrelated to $\beta $).] Alternatively, the test statistic is the (quasi-)likelihood ratio or (quasi-)score test statistic based on this parametric model.
(iii) The test statistic is a Wald-type test statistic as in (ii), except that the covariance matrix estimator is now nonparametric (in the spirit of heteroskedasticity robust testing) as described in RomanoWolf2017. See also Cragg_1983, Cragg_1992, Flachaire_2005b, Wooldridge2010, Wooldridge2012, RomanoWolf2017, Lin_Chou_2018 , DiCiccio_Romao_Wolf_2019.
Under our maintained assumptions, heteroskedasticity robust tests based on $ T_{Het}$ or $T_{uc}$ (using an arbitrary critical value $C$, including size-controlling ones) have positive power everywhere in the alternative (cf. the discussion at the beginning of Section (ref)). These tests can furthermore be shown to have power that goes to one as one moves away from the null hypothesis along sequences $(\mu _{l},\sigma _{l}^{2},\Sigma _{l})$ where $\mu _{l}$ moves further and further away from $ \mathfrak{M}_{0}$ (the affine space of means described by the restrictions $ R\beta =r$) in an orthogonal direction as $l\rightarrow \infty $, where $ \sigma _{l}^{2}$ converges to some finite and positive $\sigma ^{2}$, and $ \Sigma _{l}$ converges to a positive definite matrix. Despite of what has just been said, these tests can have, in fact not infrequently will have, infimal power equal to zero if $\mathfrak{C}$ is sufficiently rich, e.g., if $\mathfrak{C}=\mathfrak{C}_{Het}$; cf. Theorem 4.2 in PP2016, Lemma 5.11 in PP3, and Theorem 4.2 in PP4. [This does not contradict the before mentioned result as for this result sequences $\Sigma _{l}$ that converge to a singular matrix as $l\rightarrow \infty $ were ruled out.]
For tests based on $\tilde{T}_{Het}$ or $\tilde{T}_{uc}$ the situation is somewhat different. As shown in Section (ref), tests based on $ \tilde{T}_{Het}$ or $\tilde{T}_{uc}$ can be trivial for some choices of critical values $C$ (and then will have power zero everywhere in the alternative). However, if $C$ is chosen to be the smallest size-controlling critical value (provided it exists), the resulting tests obtained form $ \tilde{T}_{Het}$ or $\tilde{T}_{uc}$ will typically have positive power (under appropriate assumptions). In particular, then the test based on $ \tilde{T}_{uc}$ has the same power function as the test based on $T_{uc}$ that uses its smallest size-controlling critical value, provided the latter exists, see Section (ref). We have not further investigated the power properties of the tests based on $\tilde{T}_{Het}$ in any more detail on a theoretical level. The numerical results in Section (ref) seem to suggest that for these tests power may not go to one along sequences $(\mu _{l},\sigma _{l}^{2},\Sigma _{l})$ as mentioned above: in fact, power does not rise above the significance level $\alpha $ in some examples (on the range of alternatives considered). This feature makes tests based on $\tilde{T}_{Het}$ rather undesirable.
Consider a testing problem as in Equation ((ref)) with $ \mathfrak{C}=\mathfrak{C}_{Het}$ and let $T$ be one of the test statistics considered in the present article (e.g., $T_{Het}$ with some choice for the weights $d_{i}$). Suppose we want to numerically determine the size of the test with rejection region $\left\{ T\geq C\right\} $ for some user-supplied critical value $C$, i.e., we want to determine
Now, for all test statistics $T$ considered in the present article, this can be simplified to
where, subject to $\mu _{0}\in \mathfrak{M}_{0}$, $\mu _{0}$ can be chosen as desired. This is due to invariance properties of $T$, cf. Remarks (ref) and (ref). The quantity in ((ref)) can now be approximated numerically by any maximization algorithm where the probabilities are evaluated by Monte-Carlo methods or by the algorithm described in davies in case $q=1$, cf. Appendix (ref). \footnote{ Alternative to davies other algorithms like Imhof's algorithm, etc. can be used, some of which are also implementented in the R-package CompQuadForm (Duchesne).}
Suppose next that we want to numerically determine the smallest size-controlling critical value $C_{\Diamond }(\alpha )\in \mathbb{R}$ ($ 0<\alpha <1$) when using the test statistic $T$. [We assume here that the user knows that the smallest size-controlling critical value indeed exists, e.g., because the user has checked that the sufficient conditions developed in the present article hold, or because of other reasoning as, e.g., used in Example (ref).] Then, in view of ((ref)) and ((ref)), we need to compute $C_{\Diamond }(\alpha )$ as the smallest real number $C$ for which
holds. The quantity to the left in ((ref)) is non-increasing in the critical value $C$. Hence, to determine the smallest size-controlling critical value $C_{\Diamond }(\alpha )$, any line-search algorithm (in combination with an algorithm to determine the sizes as described before) can be used to compute $C_{\Diamond }(\alpha )$. We stress that it is of foremost importance to know that the testing problem at hand actually allows for size control before one attempts to numerically determine $C_{\Diamond }(\alpha )$. Hence, the theoretical results of the present article are of paramount importance also for the algorithmic aspect of the problem.
The specific algorithms we use to determine size and size-controlling critical values in our numerical studies are based on the above observations and are described in detail in Appendix (ref). They are made available in the R-package hrt (hrt) for the convenience of the user. The numerical procedures we use are heuristic in nature. Questions of efficacy of these algorithms or about theoretical guarantees are certainly important, but are beyond the scope of the present article.
Determining smallest size-controlling values numerically is important, e.g., if one wants to compare their magnitude with that of standard critical values in some special cases, as we do inter alia in the next section, or if one wants to obtain a confidence interval. However, a user who has observed the data and only wants to decide whether or not to reject the null hypothesis at significance level $\alpha $ ($0<\alpha <1$) when using $T$ combined with the smallest size-controlling critical value $C_{\Diamond }(\alpha )$, can actually perform this test without needing to compute $ C_{\Diamond }(\alpha )$: Let $y_{obs}$ be the observed data. Define the \textquotedblleft maximal p-value" as
where the second equality in the display follows from the invariance properties mentioned before (and $\mu _{0}\in \mathfrak{M}_{0}$ can be chosen as desired). It is now not difficult to see that $p(y_{obs})\leq \alpha $ is equivalent to $T(y_{obs})\geq C_{\Diamond }(\alpha )$. That is, rejecting if and only if $p(y_{obs})\leq \alpha $ leads to exactly the same test as rejecting if and only if $T(y_{obs})\geq C_{\Diamond }(\alpha )$, with the former description having the advantage that the more costly computation of $C_{\Diamond }(\alpha )$ can be avoided. What needs to be computed is ((ref)), which, however, is nothing else than the size of the test when using the \textquotedblleft critical value\textquotedblright\ $ T(y_{obs})$. Hence, $p(y_{obs})$ can be determined by any algorithm that determines the size ((ref)) for the user-supplied \textquotedblleft critical value\textquotedblright\ $C=T(y_{obs})$. In particular, the routine \textquotedblleft size\textquotedblright\ provided in the R-package hrt (hrt) can be used for this purpose. Note that checking whether $p(y_{obs})\leq \alpha $ avoids the line-search part (as outlined following ((ref))), and is thus computationally more efficient than first determining $C_{\Diamond }(\alpha )$ (as outlined above) and then checking whether $T(y_{obs})\geq C_{\Diamond }(\alpha )$.
Finally, we note that if (contrary to what we assume in this section) no size-controlling critical value exists for a given significance level $ \alpha \in (0,1)$, then the maximal p-value in ((ref)) is larger than $ \alpha $ for every possible observed value $y_{obs}$, and the corresponding test thus never rejects and thus is uninformative. Hence, while the explicit computation of a smallest size-controlling critical value can be avoided for performing a single test, knowing its existence is important as then the resulting test is guaranteed to be informative (non-trivial) if $T_{Het}$ or $T_{uc}$ is being used; and the same is true for $\tilde{T}_{Het}$ and $\tilde{T}_{uc}$ under the conditions discussed in Section (ref).
We also note that in view of the discussion in Section (ref) the algorithms for computing null rejection probabilities, size, and smallest size-controlling critical values discussed in this section and Appendix (ref) remain valid for elliptically symmetric distributed data without any need for modification. With regard to computing size and smallest size-controlling critical values, the same is also true for the semiparametric model described in (iv) of Section (ref).
In this section we pursue two goals:
In this section (and in the attending Appendices (ref) and (ref)) we shall often refer to $T_{Het}$ as HC0-HC4 when we want to stress that the weights $d_{i}$ being used are the HC0-HC4 weights, respectively, see Section (ref). Similarly, we shall refer to $\tilde{T}_{Het}$ as HC0R-HC4R when the HC0R-HC4R weights are used, see Section (ref). For reasons of uniformity of notation, we shall then often denote $T_{uc}$ as UC and $\tilde{T}_{uc}$ as UCR. Furthermore, throughout this section we consider the heteroskedastic Gaussian linear model with $\mathfrak{C}=\mathfrak{C}_{Het}$ as introduced in Section (ref); in particular, the notion of size in the present section (and the attending appendices) always refers to this model.
The algorithms for computing rejection probabilities, the size of a test, and size-controlling critical values used in the before-mentioned numerical computations are described in Section (ref) and Appendix (ref). Implementations are available as an R-package hrt ( hrt).
We consider the important case $q=1$, and first illustrate numerically that none of the test statistics UC, HC0-HC4, UCR, and HC0R-HC4R combined with the critical value $C_{\chi ^{2},0.05}\approx 3.8415$ results in a test that is guaranteed to have size less than or equal to $\alpha =0.05$. This is achieved by providing instances of design matrices $X$ and of hypotheses, described by $(R,r)$, such that the respective test has size larger than the nominal significance level $\alpha =0.05$, often by a large margin. Here $ C_{\chi ^{2},0.05}$ denotes the $95\%$-quantile of a chi-square-distribution with $1$ degree of freedom. [This critical value has a justification for use with HC0-HC4 or HC0R-HC4R via asymptotic considerations, but, in general, there is no such justification for use with UC or UCR, which we nevertheless include here for completeness.\footnote{ Of course, in the special case of homoskedasticity, the before mentioned justification also applies to UC and UCR.}] That is, in the instances we exhibit, this conventional critical value turns out to be too small. We next show similar results for other suggestions of critical values, e.g., for \textquotedblleft degree-of-freedom\textquotedblright\ adjustments to the conventional chi-square based critical value such as the Bell-McCaffrey adjustment (BellMcCa, Imbkoles2016). It is important to note here that in all the instances mentioned our conditions for size-controllability are satisfied, showing that size-controlling critical values can actually be found; hence, the overrejection problems mentioned before are not intrinsic problems, but only reflect the fact that conventional critical values can be a bad choice and do not guarantee size control. [In the present context it is worth recalling that for the test statistics HC0R-HC4R we have already shown in Example (ref) in Section (ref) that other situations can be found in which conventional critical values such as, e.g., $C_{\chi ^{2},0.05}$ are too large, as the resulting tests reject with probability zero only (under the null as well as under the alternative), rendering these tests useless.]
To uncover instances where the conventional critical value $C_{\chi ^{2},0.05}$ is too small, we make use of the following observation: In case a given test statistic from the above list (together with a given design matrix $X$ and hypothesis described by $(R,r)$) is such that the lower bound $C^{\ast }$ on size-controlling critical values obtained in Propositions (ref) ((ref), respectively) exceeds $C_{\chi ^{2},0.05}$, we are done, as we then know that the critical value $C_{\chi ^{2},0.05}$ leads to a test that has size $1$. [As noted subsequent to Theorems (ref) and (ref), the value of $r$ actually plays no role here, and we may set it to zero.]
Since the lower bounds $C^{\ast }$ for size-controlling critical values in Propositions (ref) ((ref), respectively) depend on the given test statistic, on $X$ and on $R$, we may -- for any given choice of test statistic and any given $R$ -- numerically search for particularly \textquotedblleft hostile\textquotedblright\ design matrices, i.e., for design matrices for which the lower bound is large, to see whether matrices $ X$ exist for which the lower bound exceeds $C_{\chi ^{2},0.05}$. We only do this for $k=2$, $R=(0,1)$, $r=0$, and $n=25$, and restrict ourselves to matrices $X$ with first column representing an intercept. The concrete search used is detailed in Appendix (ref), see Algorithm (ref) in particular. Table (ref) provides, for every test statistic considered, the lower bound $C^{\ast }$ corresponding to the most \textquotedblleft hostile\textquotedblright\ design matrix found by the search. [As the searches are run separately for each test statistic, the resulting \textquotedblleft hostile\textquotedblright\ design matrices will typically differ across the runs.]\footnote{ Since, for example, HC0 is a multiple of HC1, where the factor is $ n/(n-k)=1.09$, we know that the \textquotedblleft hostile\textquotedblright\ design matrix obtained from the search for HC1 leads to a $C^{\ast }$-value of $1.09\ast 1711.19=1865.20$ for HC0, larger than the value 95.56 obtained from the search for HC0, cf. Table (ref). We could have reported this larger value, but decided to present the raw results from our searches as this is sufficient for our purposes. We also note that our search procedure detailed in Appendix (ref) does not seriously attempt to optimize the $C^{\ast }$-value (for every one of the test statistics considered) over the set of all feasible $X$, but is only a crude search for finding a matrix resulting in a $C^{\ast }$-value sufficiently large for our purposes.}
In combination with the theoretical results from Propositions (ref) and (ref), Table (ref) shows that for some design matrices $X$ the critical value $C_{\chi ^{2},0.05}\approx 3.8415$ results in a test with size equal to $1$ when combined with UC, HC0-HC2, and also with UCR. [This is so despite the fact that, for any of the twelve test statistics considered, the sufficient conditions for size-control in the pertaining theorems in Sections (ref) and (ref) are satisfied for all relevant $X$ matrices encountered in the numerical procedure (as we have checked), and hence it is known that size-controlling critical values exist in all these situations!] Table (ref) is not informative about the size of the remaining seven tests, since the corresponding entries in that table are all less than $C_{\chi ^{2},0.05}$. To obtain insight into the sizes of the remaining seven tests we do the following: for each of the tests we numerically compute the size for various instances of design matrices (the ones that give rise to Table (ref)) and report the largest one of these sizes (\textquotedblleft worst case\textquotedblright\ sizes) in Table (ref).\footnote{ Of course, considering additional design matrices $X$ would potentially lead to even larger sizes.} We actually do this for all twelve tests considered. The algorithm used in the size computation is the implementation of Algorithm (ref) in the R-package hrt (hrt), cf. the description in Appendices (ref) and (ref). Table (ref) now clearly shows that for every test statistic considered an instance can be found, in which the size of the test (when using the critical value $C_{\chi ^{2},0.05}$) clearly exceeds the nominal significance level $\alpha =.05$. The lowest value in that table is attained by HC4R, but a size of $0.10$ is still twice the nominal significance level $ \alpha $.
We note that the numbers shown in Table (ref) actually only represent numerically determined lower bounds for the actual sizes, as their computation involves (for any given $X$) a numerical search procedure (over the set $\mathfrak{C}_{Het}$) for the worst-case null rejection probability; that is, the numbers shown in Table (ref) correspond to the null rejection probability computed from a \textquotedblleft bad\textquotedblright\ covariance matrix $\Sigma $, but potentially not for the \textquotedblleft worst\textquotedblright\ possible one. [In this process, for any given $\Sigma \in \mathfrak{C}_{Het}$, we have to numerically compute the null rejection probability, which can be done quite accurately in case $q=1$ by algorithms like the Davies algorithm, see Appendix (ref) as well as Appendix (ref).] In particular, the entries in the $0.98$-$0.99$ range in Table (ref) are numerically determined lower bounds for the size, which, in fact, we know to be equal to $1$ in light of Table (ref). [We could have used this knowledge to replace the entries in question in Table (ref) by $1$, but we decided otherwise in order to showcase the concrete outcome of the numerical algorithm that has been run. Of course, one could also improve this outcome by using a higher accuracy parameter in the optimization procedures involved.]
Sometimes -- without much theoretical justification in general -- it is suggested in the literature to replace $C_{\chi ^{2},0.05}$ by the $95\%$ -quantile of an $F_{1,n-k}$-distribution, which is approximately $4.28$ in the situation considered here ($n-k=23$). Obviously, from Table (ref) we see that the conclusions regarding UC, HC0-HC2, and UCR remain the same when this critical value is used. Repeating the exercise that has led to Table (ref), but with $C_{\chi ^{2},0.05}$ replaced by the $95\%$-quantile of an $F_{1,n-k}$-distribution, gives Table (ref), leading essentially to the same conclusions.
\textquotedblleft Degree-of-freedom\textquotedblright\ adjustments to the conventional chi-square based critical value such as the Bell-McCaffrey adjustment (BellMcCa) have been discussed in the literature. In particular, Imbkoles2016 suggested to use this adjustment with the HC2 statistic. We have repeated the above exercise that has led to the entry for HC2 in Table (ref), but with $C_{\chi ^{2},0.05}$ replaced by the Bell-McCaffrey adjustment. For the computation of the Bell-McCaffrey adjustment we relied on the R-package dfadjust (dfadjust). For the resulting test, the largest size that was found in our computations was $0.24$, which is more than four times the nominal significance level. It transpires that this adjustment does also not come with a size-guarantee.
We conclude here by stressing that the negative findings in this subsection were obtained in a very simple model with only two regressors and where only one of the parameters is subject to test. For more complex models and test problems the size distortions may even be worse.
A power comparison of two tests, both conducted at a given nominal significance level $\alpha $, makes sense only if both tests actually are level $\alpha $ tests, i.e., if both tests have a size not exceeding the given $\alpha $. For this reason, we now compare the tests obtained from the statistics UC, HC0-HC4, UCR, HC0R-HC4R only when respective smallest size-controlling critical values are used. Our theoretical results concerning the existence of size-controlling critical values, together with the algorithms for their computation in Appendix (ref), allow for such a comparison in terms of power. In all cases considered in this section $q=1$ will hold.
Throughout, in addition to the power functions of the before-mentioned tests, we also show as a benchmark the power function of the infeasible (i.e., oracle) GLS-based $F$-test conducted at the $5\%$-significance level, that makes use of knowledge of $\Sigma $. For given $\Sigma \in \mathfrak{C}_{Het}$, the distribution of this infeasible GLS-based $F$-test statistic is (under $P_{X\beta ,\sigma ^{2}\Sigma }$ with $\beta \in \mathbb{ R}^{k}$, $\sigma ^{2}\in (0,\infty )$) a noncentral $F_{1,n-k}$-distribution with noncentrality parameter $\delta ^{2}$, where
Since the power functions of all the tests considered in our study depend on the parameters $\beta $, $\sigma ^{2}$, and $\Sigma $ only through $(R\beta -r)/\sigma $ and $\Sigma $ (because of $G(\mathfrak{M}_{0})$-invariance and Proposition 5.4 in PP2016), and thus depend only on $\delta $ and $ \Sigma $, we shall -- for given $\Sigma $ -- present all these power functions as a function of $\delta $. We show only results for $\delta \geq 0 $, as the power functions in fact depend on $\delta $ only through $ \left\vert \delta \right\vert $ (for given $\Sigma $); see Proposition 5.4 in PP2016.
As a practically relevant example, we here compare the power of tests based on size-controlling critical values in the context of Example (ref). That is, we treat the problem of comparing the means of two heteroskedastic groups (e.g., a treatment and a control group), the null hypothesis being that the difference of expected outcomes in each group is zero. We consider the case where $n=30$ and $\alpha =0.05$. Furthermore, we vary the size $n_{1}$ of the first group ($n_{1}\in \{3,9,15\}$), corresponding to a \textquotedblleft strongly unbalanced\textquotedblright , \textquotedblleft moderately unbalanced\textquotedblright , and \textquotedblleft balanced\textquotedblright\ design, respectively. We compute the power for a number of covariance matrices $\Sigma _{a}$ given as follows: For $a=1,5,9$ define
where the first $n_{1}$ (and last $n-n_{1}$, respectively) diagonal entries of each $\Sigma _{a}$ are constant. That is, we look at power functions evaluated at covariance matrices under which the subjects in the same group actually have the same variances. [For brevity we do not report power functions for covariance matrices not sharing this property.] For the balanced design, we note that $\Sigma _{1}$ and $\Sigma _{9}$ lead to the same power of each test (but we report all results for completeness), and that $\Sigma _{5}$ corresponds to homoskedasticity.
The critical values are chosen in each case as the smallest critical value guaranteeing size control over $\mathfrak{C}_{Het}$ (implying, of course, that the corresponding tests can have null rejection probabilities smaller than $\alpha $ for the covariance matrices $\Sigma _{a}$ considered). The existence of said critical values follows from our theory and is discussed in detail in Example (ref) for the test statistics UC and HC0-HC4; in particular, all assumptions of Theorems (ref) are satisfied. For UCR the existence is guaranteed by Part (a) of Theorem (ref). With regard to the test statistics HC0R-HC4R, note that Assumption (ref) is satisfied since $e_{i}(n)\notin \mathfrak{M}_{0}^{lin}=\mathfrak{M}_{0}=\limfunc{span}((1,\ldots ,1)^{\prime })$ for every $i=1,\ldots ,n$ as $n=30>k=2$. This also shows that the sufficient condition for size control ((ref)) is satisfied as $\mathsf{\tilde{B}}=\mathfrak{M}_{0}$ is easily verified and since one may set $\mu _{0}=0$. We have verified the non-constancy assumption on the test statistics HC0R-HC4R in Theorem (ref) numerically. As a consequence, all assumptions of Part (b) of Theorem (ref) are satisfied.
We note that some of the test statistics differ from each other only by a known multiplicative constant and hence are equivalent in the sense that they give rise to the same test when the respective smallest size-controlling critical value is employed, see Remarks (ref) and (ref): In the unbalanced case ($n_{1}\in \{3,9\}$),\ HC0 and HC1 are equivalent in this sense, as are HC0R-HC4R (the latter is so since $\tilde{h}_{ii}=1/n$ which does not depend on $i$). In the balanced case ($n=15$), UC and HC0-HC4 are all equivalent, and the same is true for UCR and HC0R-HC4R as is not difficult to see. Furthermore, in the balanced as well as in the unbalanced case, the rejection regions of the tests based on UC and UCR coincide essentially (i.e., up to a $\lambda _{\mathbb{R}^{n}}$ -null set) as a consequence of the relationship established in Section (ref). In particular, it follows that in the balanced case, all tests considered (essentially) coincide. We nevertheless compute the power functions for each of the tests separately without making use of the noted equivalencies; this provides a double-check of our numerical results.\footnote{ The equivalencies mentioned in this paragraph for the two-group-comparison problem analogously hold for general $n$, $n_{1}$, and $n_{2}$ as is easily seen.}
Numerically the critical values were determined through the implementation of Algorithms (ref) and (ref) in the R-package hrt ( hrt) version 1.0.0, and the power functions were computed with the implementation of the algorithm by davies in the R-package CompQuadForm (Duchesne) version 1.4.3; see Appendices (ref) and (ref) for more details. For the sake of illustration, we also report the critical values obtained for every test considered and every balancedness condition in Table (ref).
In relation to Table (ref) we note that the equivalences discussed before predict, e.g., that the ratio between the entries in the column labeled HC0 and the corresponding entries in the column labeled HC1 should be equal to $n/(n-2)=30/28\approx 1.0714$. The ratios computed from the table are $1.0414$, $1.0761$, and $1.0721$ (for $n_{1}=3,9,15)$, which is in pretty good agreement (especially if one converts the critical values shown in the table to critical values for the corresponding \textquotedblleft $t$ -test\textquotedblright\ versions by computing their square roots). The agreement between theoretical and observed ratios for the HC0R-HC4R columns is similar. In the balanced case one can also use the additional equivalences mentioned before and one again finds very good agreement. Similarly, the critical values for UC and UCR in Table (ref) are in excellent agreement with their theoretical relationship found in Section (ref). The reason for the small discrepancies observed lies in the fact that the algorithm underlying the computations for Table (ref) makes use of a random search algorithm. Concerning Table (ref), we also mention that, in the example considered here and for the test statistic HC2, IbragMuell2016 prove in their Theorem 1 (see also the discussion preceding that theorem) that the smallest size-controlling critical values are given by\ $18.51$ ($n_{1}=3$), $5.32$ ($n_{1}=9$ ), and $4.60$ ($n_{1}=15$), respectively. The numerically determined critical values in Table (ref) are reasonably close to these values (after conversion of the critical values to corresponding \textquotedblleft $ t$-test\textquotedblright\ critical values the maximal difference is about $ 0.1$). Of course, the accuracy of our algorithm could be increased by using more stringent accuracy parameters in the optimization routines underlying the computation of the critical value, but this would come with a longer runtime.
From Table (ref) it is clear that for the tests based on unrestricted residuals the smallest size-controlling critical values obtained are always larger, sometimes considerably, than $C_{\chi ^{2},0.05}\approx 3.8415$, again showing that the latter critical value is not effecting size control. For the tests based on restricted residuals the smallest size-controlling critical values sometimes fall below $C_{\chi ^{2},0.05}$ in the strongly unbalanced case (which is not completely surprising in view of Section (ref)); while in this case $C_{\chi ^{2},0.05}$ effects size-control, using the smaller size-controlling critical values given in Table (ref) can only be advantageous in terms of power.
That being said, we emphasize a trivial, but important point, namely that comparing the magnitudes\ of size-controlling critical values relating to different test statistics is not very meaningful and, in particular, not a valid way of comparing the quality of the resulting tests. That is, while it may be tempting to infer from Table (ref) that the HC0 test should be considerably more conservative than the HC4 test, or that the UC test should be considerably more conservative than the UCR test, such a conclusion would be false and not warranted at all (in particular, recall that UC and UCR in fact result in (essentially) the same test if the critical values from Table (ref) are being used). While this would be correct if the critical values were all meant to be used with the same test statistic (which they are not), critical values belonging to different test statistics can certainly not be compared in such a way. Instead, one has to compare the corresponding power functions, which is what we shall do next.
The power functions are shown in Figure (ref) (\textquotedblleft strongly unbalanced\textquotedblright , $n_{1}=3$), Figure (ref) (\textquotedblleft moderately unbalanced\textquotedblright , $n_{1}=9$), and Figure (ref) (\textquotedblleft balanced\textquotedblright , $ n_{1}=15$), where only the first two figures are shown in the main text, and the last figure (in which the power functions of all the feasible tests lie \textquotedblleft on top of each other\textquotedblright ) is available in Appendix (ref). Readers are referred to the online version for colored figures.
The power functions illustrate that the testing problem is getting easier, (i.e., power gets closer to the oracle benchmark), for more balanced design, which has intuitive appeal. Except for the strongly unbalanced case ($ n_{1}=3 $), the power loss of the tests based on HC0-HC4 and HC0R-HC4R relative to the oracle benchmark is surprisingly small (see Figure (ref) as well as Figure (ref) in Appendix (ref)). In the unbalanced cases ($n_{1}\in \{3,9\}$) the HC0-HC4-based tests behave all very similarly, with the power functions of the HC0- and HC1-based test being virtually indistinguishable (as they should in view of the before discussed equivalence). The UC-based test shows markedly worse power performance. Similarly, the HC0R-HC4R-based tests have virtually indistinguishable power functions (as they should because of the before discussed equivalence). The UCR-based test again is inferior (and its power function coincides with the one of UC as mentioned before). There appears also to be little difference between basing the test statistics on unrestricted or restricted residuals in this example. In the balanced case we know that all the feasible tests have exactly the same power function in view of our earlier discussion. This is visible in Figure (ref) in Appendix (ref). Also the different forms of heteroskedasticity considered seem not to have much effect on the power functions (when expressed as a function of $\delta $), except for UC and UCR in the unbalanced cases.
Hence, within the scenario considered in this section, perhaps the most important conclusion concerning the choice of a test statistic appears to be to avoid UC and UCR. Everything apart from that, i.e., whether one uses unrestricted or restricted residuals to construct the test or which specific heteroskedasticity-correction one decides to use, seems to be a comparably irrelevant part of the problem once the right (i.e., smallest size-controlling) critical value is used. We shall see in the next subsection that this conclusion very much depends on the scenario considered here and does not generalize beyond, illustrating the danger of drawing conclusions from a limited numerical study.
In this section, we consider testing $\beta _{2}=0$ in a model with intercept and a single regressor $x=(10,\cos (2),\cos (3),\ldots ,\cos (n))^{\prime }$. Obviously, the regressor has a dominant first coordinate, leading to diagonal elements $h_{ii}$ of $X(X^{\prime }X)^{-1}X^{\prime }$ such that the ratio of largest to smallest $h_{ii}$ is roughly $26$ ($\max h_{ii}\simeq 0.879$, $\min h_{ii}\simeq 0.033$). Hence, the design matrix $X$ provides (on purpose) an extreme case, which leads to quite interesting results. We consider again the case $n=30$ and $\alpha =0.05$, but now show power functions for $\Sigma _{a}^{\ast }$, $a=0,\ldots ,4$, where
Note that $\Sigma _{0}^{\ast }=n^{-1}I_{n}$ and that increasing $a$ from $0$ to $4$ leads to covariance matrices that approach the degenerate matrix $ e_{1}(n)e_{1}(n)^{\prime }$. All conditions in Theorems (ref) and (ref) are seen to be satisfied in this example: As no vector $e_{i}(n)$ belongs to $\limfunc{span}(X)$ (and thus also not to $ \mathfrak{M}_{0}^{lin}$), Assumptions (ref) and (ref) as well as the sufficient condition for size control ((ref)) are obviously satisfied. The size control conditions ( (ref)) and ((ref)) have been checked numerically, as has been the condition that none of the test statistics HC0R-HC4R is constant on $\mathbb{R}^{n}\backslash \mathsf{\tilde{B}}$.
As in the preceding subsection, the critical values for each test statistic are again chosen as the smallest critical value guaranteeing size control over $\mathfrak{C}_{Het}$ and they are presented in Table (ref) below. [Existence follows from our theory since all assumptions are satisfied as noted before.] For their computation the same algorithms were used as in Section (ref), with a similar statement applying to the numerical routines used for computing the power functions. Note that the critical values for the test statistics UC, HC0-HC3 are large, reflecting the high-leverage in the design matrix; an exception is HC4, the reason being that some of the HC4-weights are considerably larger than the weights for HC0-HC3. Similarly as in the preceding subsection, the tests based on HC0 and HC1 coincide (since HC0 and HC1 differ only by a multiplicative constant and since smallest size-controlling critical values are being used), and the same is true for the tests based on HC0R-HC4R, see Remarks (ref) and (ref). It is easily checked that the ratios of the respective critical values provided in Table (ref) are in good agreement with the theoretical ratios predicted by theory. Furthermore, the tests based on UC and UCR coincide (see Section (ref)), and the critical values for UC and UCR in Table (ref) are in excellent agreement with their theoretical relationship found in Section (ref).
Table (ref) shows that in this example the smallest size-controlling critical values are -- except in one case -- always larger, sometimes considerably larger, than $C_{\chi ^{2},0.05}\approx 3.8415$, once more showing that the latter critical value is not effecting size control in general. In the exceptional case, namely when the HC4 test statistic is used, $C_{\chi ^{2},0.05}$ is considerably larger than the smallest size-controlling critical value, which is $1.12$; while in this case $ C_{\chi ^{2},0.05}$ effects size-control, using the smaller size-controlling critical value $1.12$ can only be advantageous in terms of power.
The power functions, when the size-controlling critical values from Table (ref) are being used, are shown in Figure (ref). Readers are referred to the online version for a colored figure. Again, as predicted by theory, the power functions of the tests based on HC0 and HC1 shown in Figure (ref) coincide, as do the power functions of the tests based on HC0R-HC4R; the same is true for the power functions of the tests based on UC and UCR. The figure furthermore shows that in the setting considered here, there is now a marked difference between tests based on HC0-HC4 and on HC0R-HC4R, respectively: the power of the tests based on HC0R-HC4R is nowhere greater than $\alpha $, their power function being even non-monotonic, whereas the tests based on HC0-HC4 have increasing power as a function of $\delta $. In contrast to the example considered in the preceding subsection, the power functions of the tests based on HC0-HC4 and on UC are now all markedly different and typically intersect, an exception being the case of $\Sigma _{4}^{\ast }$ where the test based on UC offers the highest power for that covariance matrix. Overall, however, there is no clear ranking between the tests using unrestricted residuals in the example considered here, although we note that the test based on UC (or, equivalently, on UCR) performs very badly in the case of $\Sigma _{0}^{\ast } $. This is not surprising as $\Sigma _{0}^{\ast }$ corresponds to homoskedasticity and the critical value used here is much larger than the classical critical value one would use given knowledge of this homoskedasticity. Furthermore, and in contrast to the results in the preceding subsection, the different forms of heteroskedasticity considered have a noticeable effect on the power functions. The main takeaway is that tests based on HC0R-HC4R (and probably on UC and UCR) should rather be avoided.
The usual heteroskedasticity robust test statistics such as $T_{Het}$ (using HC0-HC4 weights) or $\tilde{T}_{Het}$ (using HC0R-HC4R weights), used in conjunction with conventional critical values obtained from the asymptotic null distribution, are often plagued by overrejection under the null. This has been clearly documented in the literature for $T_{Het}$, and is shown numerically for $\tilde{T}_{Het}$ (as well as for $T_{Het}$) in Section (ref) above. Not surprisingly, similar observations apply to the \textquotedblleft uncorrected\textquotedblright\ test statistics $T_{uc}$ and $\tilde{T}_{uc}$. We show theoretically that all these test statistics can be size-controlled under quite weak conditions by an appropriate choice of critical values.
From the above discussion and the numerical results in Section (ref) it transpires that smallest size-controlling critical values rather than conventional critical values should be used in order to avoid the risk of overrejection. For the computation of smallest size-controlling critical values we provide algorithms which have been implemented in the R-package hrt (hrt) and thus are readily available for the user.
An additional advantage from using smallest size-controlling critical values over conventional critical values is that this typically leads to improved power in instances where conventional critical values lead to underrejection (i.e., lead to worst-case rejection probability under the null less than the nominal significance level) as is sometimes the case; see Sections (ref) and (ref).
If smallest size-controlling critical values are adopted (as they should), the numerical results in Section (ref) suggest that the test statistic $\tilde{T}_{Het}$ (with the usual weights HC0R-HC4R) should be avoided, as the resulting tests may have very poor power properties (see the example in Section (ref)). The test statistic $T_{Het}$ seems to perform better in terms of power, with no clear ranking emerging with regards to the weights HC0-HC4 being used. The \textquotedblleft uncorrected\textquotedblright\ test statistics $T_{uc}$ and $\tilde{T}_{uc}$ appear to be inferior to $T_{Het}$ in terms of power in almost all of the numerical examples considered. We also point out that -- when using smallest size-controlling critical values -- the tests based on $T_{Het}$ employing the HC0 and the HC1 weights, respectively, in fact coincide; and the same holds for tests based on $\tilde{T}_{Het}$ employing the HC0R and the HC1R weights, respectively. Also, the tests based on $T_{uc}$ and $\tilde{T}_{uc}$ then (essentially) coincide. See Remarks (ref), (ref), and Section (ref) as well as the pertaining discussion in Section (ref) for more information, including additional equivalencies when the design matrix $X$ and the restriction $R$ have certain special properties.