Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
37,129 characters · 5 sections · 68 citation commands
A Necessary and Sufficient Condition for Size Controllability of Heteroskedasticity Robust Test Statistics
\and
}
Tests and confidence intervals based on so-called heteroskedasticity robust standard errors date back to E63, E67 and constitute, at least since W80, a major component of the applied econometrician's toolbox. Although these early methods come with well-understood large sample properties, when based on critical values derived from asymptotic theory their finite sample properties often deviate substantially from what asymptotic theory suggests: tests may substantially overreject under the null and corresponding confidence intervals may undercover. Strong leverage points have been identified early on as one major reason for these deviations, see, e.g., MacW85, DavidsonMacKinnon1985, and CheshJewitt1987. This has led to various developments trying to attenuate such drawbacks:
Although these developments sometimes lead to improvements, they come with no general finite sample guarantees concerning the size of the tests or the coverage of related confidence intervals, cf. the discussion in PPBoot,PP21 for detailed accounts.
Motivated by this lack of finite sample guarantees, PP21 studied the question under which conditions heteroskedasticity robust test statistics as well as the standard (uncorrected) F-test statistic can actually be paired with appropriate (finite) critical values, so that one obtains tests that have their (finite sample) size controlled by the prescribed significance value ${\Greekmath 010B} $ (i.e., have size $\leq {\Greekmath 010B} $) even though one is completely agnostic about the form of heteroskedasticity.\footnote{ The null-hypothesis to be tested is given by a set of affine restrictions.} Under appropriate assumptions on the errors, allowing for Gaussian as well as substantial non-Gaussian behavior, they have shown that the standard (uncorrected) F-test statistic can be size-controlled (in finite samples) by using an appropriately chosen (finite) critical value if and only if the following simple condition holds:
see (8) in PP21 for a formal statement of this condition.
Under a generally stronger condition than ((ref)) (see (10) in PP21), it was furthermore shown that large classes of heteroskedasticity robust test statistics (e.g., HC0-HC4) can be size-controlled by appropriate (finite) critical values. That condition, however, although satisfied for many testing problems (and even often identical to ((ref)), cf. Theorem 3.9 and Lemma A.3 in PP3), is not necessary in general, as shown in examples given in PP21 ; e.g., their Example 5.5 or Example C.1 in their Appendix C.\footnote{ Appendices to PP21 are published in the Supplementary Material available at the publisher's website of that article.} These examples consider the case of testing linear contrasts in the expected outcomes of subjects belonging to two or more groups, scenarios that are practically relevant. Further examples are provided in Examples A.1-A.4 in Appendix (ref) further below.\footnote{ Example 5.5 in PP21 concerns simultaneously testing multiple retrictions, while Example C.1 in Appendix C of PP21 as well as Examples A.1-A.4 in Appendix (ref) of the present article concern the case of testing a single restriction.}
For the important case of testing problems involving only a single restriction (i.e., the case $q=1$ in the notation of PP21), we show in the present article that the condition in ((ref)) is then in fact necessary and sufficient also for size controllability of the above mentioned classes of heteroskedasticity robust test statistics, including HC0-HC4.
Here we recall the most relevant notions from Sections 2 and 3 of PP21 , to which we refer the reader for further information and discussion. We consider the linear regression model
where $X$ is a (real) nonstochastic regressor (design) matrix of dimension $ n\times k$ and where ${\Greekmath 010C} \in \mathbb{R}^{k}$ denotes the unknown regression parameter vector. Throughout, we assume $\limfunc{rank}(X)=k$ and $1\leq k<n$. We furthermore assume that the $n\times 1$ disturbance vector $ \mathbf{U}=(\mathbf{u}_{1},\ldots ,\mathbf{u}_{n})^{\prime }$ ($^{\prime }$ denoting transposition) has mean zero and unknown covariance matrix ${\Greekmath 011B} ^{2}\Sigma $ ($0<{\Greekmath 011B} <\infty $), where $\Sigma $ varies in the \textquotedblleft heteroskedasticity model\textquotedblright\ given by
and where $\limfunc{diag}({\Greekmath 011C} _{1}^{2},\ldots ,{\Greekmath 011C} _{n}^{2})$ denotes the diagonal $n\times n$ matrix with diagonal elements given by ${\Greekmath 011C} _{i}^{2}$. That is, the disturbances are uncorrelated but can be heteroskedastic of arbitrary form. [In Appendix (ref) we shall also consider another heteroskedasticity model.]\footnote{ Since we are concerned with finite-sample results only, the elements of $ \mathbf{Y}$, $X$, and $\mathbf{U}$ (and even the probability space supporting $\mathbf{Y}$ and $\mathbf{U}$) may depend on sample size $n$, but this will not be expressed in the notation. Furthermore, the obvious dependence of $\mathfrak{C}_{Het}$ on $n$ will also not be shown in the notation, and the same applies to the heteroskedasticity model defined in Appendix (ref).}
For ease of exposition, we shall maintain in the sequel that the disturbance vector $\mathbf{U}$ is normally distributed. Generalizations to classes of non-normal disturbances can be obtained following the arguments in Section 7.1 of PP21, see Remark 2.2 further below. Denoting a Gaussian probability measure with mean ${\Greekmath 0116} \in \mathbb{R}^{n}$ and (possibly singular) covariance matrix $A$ by $P_{{\Greekmath 0116} ,A}$ , the collection of distributions on $\mathbb{R}^{n}$ (the sample space of $ \mathbf{Y}$) induced by the linear model just described together with the Gaussianity assumption is then given by
where $\mathrm{\limfunc{span}}(X)$ denotes the column space of $X$.\footnote{ Since every $\Sigma \in \mathfrak{C}_{Het}$ is positive definite, the measure $P_{{\Greekmath 0116} ,{\Greekmath 011B} ^{2}\Sigma }$ is absolutely continuous with respect to Lebesgue measure on $\mathbb{R}^{n}$.}
We focus on testing the null $R{\Greekmath 010C} =r$ against the alternative $R{\Greekmath 010C} \neq r$, where $R\neq 0$ is a $1\times k$ vector and $r\in \mathbb{R}$. That is, throughout this paper we focus on testing a single restriction, whereas the theory developed in PP21 allows for simultaneously testing multiple restrictions (that is, we here consider only the special case corresponding to $q=1$ in PP21). Set $\mathfrak{M}=\limfunc{span} (X)$, define the affine space
and let
Adopting these definitions, the testing problem we consider can be written more precisely as
We also write $\mathfrak{M}_{0}^{lin}=\mathfrak{M}_{0}-{\Greekmath 0116} _{0}=\left\{ X{\Greekmath 010C} :R{\Greekmath 010C} =0\right\} $ where ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$. Of course, $\mathfrak{M}_{0}^{lin}$ does not depend on the choice of ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$. Furthermore, if $\mathcal{L}$ is a linear subspace of $ \mathbb{R}^{n}$, $\Pi _{\mathcal{L}}$ denotes the orthogonal projection onto $\mathcal{L}$, while $\mathcal{L}^{\bot }$ denotes the orthogonal complement of $\mathcal{L}$ in $\mathbb{R}^{n}$.
The assumption of nonstochastic regressors made above entails little loss of generality, and results for models with stochastic regressors can be obtained from the ones derived in the present paper by the same arguments as the ones given in Section 7.2 of PP21.
We consider the same test statistics as in Section 3 of PP21. Simplified to the setting of testing a single restriction considered in the present article, they are given by
where $\hat{{\Greekmath 010C}}(y)=\left( X^{\prime }X\right) ^{-1}X^{\prime }y$ and where $\hat{\Omega}_{Het}(y)=R\hat{\Psi}_{Het}(y)R^{\prime }$. Here
with $\hat{u}(y)=\left( \hat{u}_{1}(y),\ldots ,\hat{u}_{n}(y)\right) ^{\prime }=y-X\hat{{\Greekmath 010C}}(y)$. The constants $d_{i}>0$ sometimes depend on the design matrix; see PP21 for examples of the weights $d_{i}$, including HC0-HC4 weights. We also recall the following assumption from the latter reference, again specialized to the setting of testing only a single restriction (i.e., to the case $q=1$ in the notation of PP21).
This assumption can be checked in any particular application as it only depends on the observable quantities $R$ and $X$; and a sufficient condition for Assumption (ref) obviously is $s=0$. Assumption (ref) is unavoidable if one wants to obtain a sensible test from the statistic $ T_{Het}$, see Section 3 of PP21 for more discussion. We note that $ e_{j}(n)\in \limfunc{span}(X)$ is equivalent to $h_{jj}=1$, where $h_{jj}$ denotes the $j$-th diagonal element of the `hat matrix' $H=X(X^{\prime }X)^{-1}X^{\prime }$.\footnote{ This follows from $h_{jj}=e_{j}(n)^{\prime }He_{j}(n)=(He_{j}(n))^{\prime }He_{j}(n)$ and the fact that $H$ represents the orthogonal projection onto $ \limfunc{span}(X)$.}
As in PP21, we introduce
Define (recall that $R$ is a nonzero row vector in this article)
It is now easy to see that $\limfunc{span}(X)\subseteq \mathsf{B}$ and that $ \mathsf{B}$ is a linear space (cf. also Lemma 3.1 in PP21). Simple examples can be constructed to show that $\limfunc{span}(X)\neq \mathsf{B}$, in general; cf. Example C.1 in Appendix C of PP21 as well as Examples A.1-A.4 in Appendix (ref) further below.
To summarize the main size controllability statements from PP21 for the above class of test statistics, we first have to recall the following notation: For a given linear subspace $\mathcal{L}$ of $\mathbb{R}^{n}$ we define the set of indices $I_{0}(\mathcal{L})$ via
We set $I_{1}(\mathcal{L})=\left\{ 1,\ldots ,n\right\} \backslash I_{0}( \mathcal{L})$. Clearly, $\func{card}(I_{0}(\mathcal{L}))\leq \dim (\mathcal{L })$ holds. And $I_{1}(\mathcal{L})$ is nonempty provided $\dim (\mathcal{L} )<n$; in particular, $I_{1}(\mathfrak{M}_{0}^{lin})$ is always nonempty since $\dim (\mathfrak{M}_{0}^{lin})=k-1<n-1$. The results in PP21 concerning size controllability of tests for ((ref)) based on $T_{Het}$ can now be summarized as follows; some intuition for why size control cannot always be achieved is provided further below as well as in Section 4 in PP21:
To obtain some intuition for Theorem (ref), recall that the diagonal elements of $\Sigma \in \mathfrak{C}_{Het}$ are positive and sum up to one (by definition). Now, for a matrix $\Sigma $ with $i$-th diagonal entry close to $1$, all other diagonal entries must therefore be close to $0$ , so that $\Sigma \approx e_{i}(n)e_{i}(n)^{\prime }$ then holds. Note that if $\Sigma \approx e_{i}(n)e_{i}(n)^{\prime }$, the distribution $P_{{\Greekmath 0116} ,{\Greekmath 011B} ^{2}\Sigma }$ of the data is strongly \textquotedblleft concentrated\textquotedblright\ around the one-dimensional space ${\Greekmath 0116} + \limfunc{span}(e_{i}(n))$. From an intuitive point of view, whether a given test statistic admits a size-controlling critical value or not, should therefore depend on the \textquotedblleft behavior\textquotedblright\ of the test statistic for values on or close to the spaces ${\Greekmath 0116} _{0}+\limfunc{span} (e_{i}(n))$ with ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$. It turns out that this is intimately related to ((ref)) and ((ref)). See Section 4 in PP21 for more discussion.
Most importantly, the above theorem shows that, given Assumption (ref), the condition in ((ref)) is sufficient for the existence of a (finite) size-controlling critical value $C({\Greekmath 010B} )$ satisfying ((ref)), while the weaker condition ((ref)) is necessary. Furthermore, in case the design matrix $ X$ and the vector $R$ are such that $\mathsf{B}=\limfunc{span}(X)$, and hence the condition in ((ref)) coincides with that in ((ref)), the condition ((ref)) is also necessary. However, $\mathsf{B}=\limfunc{span}(X)$ is not always true (see Example C.1 in Appendix C of PP21 or the examples in Appendix (ref) further below), although the equality holds generically (cf. Theorem 3.9 and Lemma A.3 in PP3). We now show in the subsequent theorem that in the situation considered in this article, namely testing only a single restriction, the condition in ((ref)) in Theorem (ref) can actually always be replaced by that in ((ref)). Before we present that theorem, we discuss an equivalent formulation of condition ((ref)) that is expressed in terms of certain diagonal elements of the `hat matrix' $H $, see ((ref)) below.\footnote{ An informal verbal description of ((ref)) is given in ( (ref))\ in the Introduction.}
Remark 2.1: (i) Condition ((ref)) is equivalent to "$h_{ii}<1$ for every $i\in I_{1}(\mathfrak{M}_{0}^{lin})$".\footnote{ Note that $h_{ii}=1$ always holds if $i\in I_{0}(\mathfrak{M}_{0}^{lin})$.}
(ii) Condition ((ref)) can also equivalently be written as
see Remark B.1(iii) in Appendix (ref) further below.\footnote{ Comparing ((ref)) and ((ref)) could lead one to conjecture equivalence of the conditions $i\in I_{1}(\mathfrak{M}_{0}^{lin})$ and $R(X^{\prime }X)^{-1}x_{i\cdot }^{\prime }\neq 0$. This is incorrect in general, see Example A.1 in Appendix A. However, $R(X^{\prime }X)^{-1}x_{i\cdot }^{\prime }\neq 0$ implies $i\in I_{1}(\mathfrak{M} _{0}^{lin})$, see Part 3 of Lemma (ref).} And this in turn is now equivalent to
The last form of the condition may be more appealing to some readers. We issue a warning here, however, namely that the condition ((ref) ) is, in general, stronger than the condition "$e_{i}(n)\notin \mathsf{B}$ for every $i$ satisfying $R(X^{\prime }X)^{-1}x_{i\cdot }^{\prime }\neq 0$", see Remark B.1(iv) in Appendix (ref).
We now present the announced theorem.
The main take-away of Theorem (ref) is that, given Assumption (ref) holds, the condition in ((ref)) (or equivalently ((ref))) is necessary and sufficient for the existence of a (smallest) finite size-controlling critical value when one is testing only a single restriction.\footnote{ By contraposition, the design matrices $X$ and restrictions $R$ for which size control fails are precisely characterized by failure of\ ((ref)) (or equivalently ((ref))). One example is when $X$ contains the dummy $e_{i}(n)$ as its first column, say, and $R=(1,0,\ldots ,0)$ (or, more generally, $R$ has a non-zero first entry). Another example arises when the first two columns of $X$ are given by $(1,\ldots ,1)^{\prime }$ and $(1,-1,\ldots ,-1)^{\prime }$, and the first two entries of $R$ are both equal to $1$.} The condition "$e_{i}(n)\notin \func{span}(X)$ for every $i=1,\ldots ,n$" (which is tantamount to "$h_{ii}<1$ for every $i=1,\ldots ,n $") implies ((ref)), and thus is sufficient for size-controllability of $T_{Het}$ (but not necessary, see, e.g., Example A.2). Note that the conditions in ((ref)), ((ref)), as well as ((ref)) do not depend on the weights used in the construction of the covariance matrix estimator or on $r$. They only depend on $X$ and $R$. This and more (e.g., how the conditions relate to high-leverage points) is discussed subsequent to Theorem 5.1 (and in Remarks 5.2-5.4, 5.6, and 5.9) in PP21 to which we refer the reader for a detailed account. As a point of interest we also note that condition ( (ref)) given above is exactly the same as condition (8) in PP21 (with $q=1$); in that reference, the latter condition is shown to be necessary and sufficient for size control of the standard (uncorrected) F-test statistic (regardless of whether $q=1$ or not).
We also note here that Theorem (ref) disproves -- for the special case of testing a single restriction -- a conjecture in Remark 5.8 of PP21, namely that there would exist cases where Assumption (ref) holds, ((ref)) is satisfied, ((ref)) does not hold, and size control by a (finite) critical value is not possible.
To see why the refinement of Theorem (ref) provided in Theorem (ref) can matter in practice, it is enough to consider the textbook example of a matrix $X$ with two columns, the first indicating membership to the treatment group and the second indicating membership to the control group (a special case of Example C.1 in Appendix C of PP21). Assume that the first $n_{1}\geq 2$ observations belong to the treatment group and the remaining $n_{2}\geq 2$ observations belong to the control group. Assume further that one wants to test whether ${\Greekmath 010C} _{1}$, the expected outcome of the treatment group, equals a given value (e.g., because one wants to obtain a confidence interval through test inversion). Example C.1 in Appendix C of PP21 shows that in this case
In particular, $e_{i}(n)\in \mathsf{B}$ if and only if $i>n_{1}$, so that ( (ref)) is not satisfied, while ((ref)) holds, and size-controlling critical values hence exist by Theorem (ref) (and can be used for constructing confidence intervals). Further examples are provided in Appendix (ref) below.
We next explain the key observation underlying the proof of Theorem (ref): To this end, define the (possibly empty) set of indices
where $x_{i\cdot }$ denotes the $i$-th row of $X$, and define (the span of the empty set will throughout be interpreted as $\{0\}$) the space
the inclusion holding because $\mathsf{B}$ is a linear space as noted earlier (recall that $R$ is $1\times k$ dimensional in this article). \footnote{ We note that $\mathcal{I}_{\#}$ is a proper subset of $\{1,\ldots ,n\}$ since $R\neq 0$.} Recall that under Assumption (ref) the test statistic $T_{Het}$ as well as $\mathsf{B}$ are invariant with respect to (w.r.t.) the group $G(\mathfrak{M}_{0})$ (i.e., the group of transformations $y\mapsto {\Greekmath 010E} (y-{\Greekmath 0116} _{0})+{\Greekmath 0116} _{0}^{\ast }$ with ${\Greekmath 010E} \in \mathbb{R}$ nonzero and ${\Greekmath 0116} _{0}$ and ${\Greekmath 0116} _{0}^{\ast }$ in $\mathfrak{M}_{0}$), see Remark C.1 in Appendix C of PP21.\footnote{ The invariance holds trivially if Assumption (ref) is violated.} The results in PP21 are based on this invariance property. The crucial observation exploited in the proof of Theorem (ref) now is that, in the special case of testing a single restriction considered in this article, the test statistic $T_{Het}$ as well as $\mathsf{B}$ are invariant, not only w.r.t. $G(\mathfrak{M}_{0})$, but also w.r.t. addition of elements of $ \mathcal{V}_{\#}$. This additional invariance property involving $\mathcal{V} _{\#}$, paired with a careful application of the general theory for size-controlling critical values in PP3, then allows us to deduce the refined statement in Theorem (ref). It turns out fortunate that the general theory in PP3 explicitly allows one to incorporate additional invariance properties beyond $G(\mathfrak{M}_{0})$. For details and proofs the reader is referred to Appendices (ref) and (ref).
Finally, we remark that Theorem (ref) is deduced from Theorem (ref) in Appendix (ref), which is a more general statement that also allows for heteroskedasticity models other than $\mathfrak{C} _{Het} $ (and which are defined in ((ref)) below).
Remark 2.2: (Extensions to non-Gaussian errors) (i) All the theorems in this article continue to hold as they stand, if the disturbance vector $\mathbf{U}$ follows an elliptically symmetric distribution that has no atom at the origin; more precisely, $\mathbf{U}$ is assumed to be distributed as ${\Greekmath 011B} \Sigma ^{1/2}\mathbf{z}$, where $\mathbf{z}$ has a spherically symmetric distribution on $\mathbb{R}^{n}$ that has no atom at the origin, and where ${\Greekmath 011B} $ and $\Sigma $ are as in Section (ref). This is so, since the size under Gaussianity is the same as the size under the elliptical symmetry assumption. In particular, the smallest size-controlling critical values under the elliptical symmetry assumption coincide with the smallest size-controlling critical values under Gaussianity, and thus can be computed from the algorithms relying on Gaussianity described in PP21. See Appendix E.1 of PP3 and Section 7.1(i) of PP21 for more details. The same is actually true for a wider class of distribution for $\mathbf{U}$, namely where $\mathbf{z}$ has a distribution in the class $Z_{ua}$ defined in Appendix E.1 of PP3.
(ii) All the theorems in this article except for Theorem (ref) in Appendix (ref) (i.e., all theorems using the heteroskedasticity model $\mathfrak{C}_{Het}$) continue to hold as they stand, if it is assumed that the disturbance vector $\mathbf{U}$ follows a distribution from the semiparametric model defined in Section 7.1(iv) in PP21 (a model that contains inter alia all distributions corresponding to i.i.d. samples of scale-mixtures of normals). Again, this is so since the size under Gaussianity is the same as the size under this semiparametric model. In particular, the smallest size-controlling critical values under this semiparametric model coincide with the smallest size-controlling critical values under Gaussianity, and thus can be computed from the algorithms relying on Gaussianity described in PP21. See Section 7.1(iv) in PP21 and note that the Gaussian model is a submodel of the semiparametric model considered there.
(iii) Furthermore, as discussed in detail in Appendix E.2 of PP3, any condition sufficient for size controllability under Gaussianity of the disturbance vector $\mathbf{U}$ also implies size controllability for large classes of distributions for $\mathbf{U}$ that satisfy appropriate domination conditions; however, the corresponding size-controlling critical values may then differ from the size-controlling critical values that apply under Gaussianity.
In the case of testing a single restriction, we have shown that the sufficient condition for size controllability of heteroskedasticity robust test statistics in PP21 can be replaced by a weaker sufficient condition that is also necessary. This allows one -- in the case of testing a single restriction -- to resolve the question of existence of (finite) size-controlling critical values in all cases, including those that remain inconclusive under the results in PP21.
We finally remark that the algorithms designed to compute size-controlling critical values as discussed in Section 10 and Appendix E of PP21 can be used as they stand also in situations where (a single restriction is tested and) size controllability has been verified through checking condition ((ref)) (or equivalently ((ref))) and appealing to Theorem (ref), but where ((ref)) does not hold. This is so since the discussion of the before mentioned algorithms in PP21 only requires existence of a (finite) size-controlling critical value, but does not depend on the way this existence is verified.