EconBase
← Back to paper

A Necessary and Sufficient Condition for Size Controllability of Heteroskedasticity Robust Test Statistics

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

37,099 characters

A Necessary and Sufficient Condition for Size Controllability of Heteroskedasticity Robust Test Statistics



\title{A Necessary and Sufficient Condition for Size Controllability of
Heteroskedasticity Robust Test Statistics\thanks{
We thank Mikkel S\o lvsten for helpful discussions and for suggesting to
re-express condition (\ref{non-incl_Het_uncorr}) as condition (\ref{sol}) in
Remark 2.1(ii). We are also grateful to two referees and a Co-Editor for
helpful comments.}}
\author{
\begin{tabular}{c}
Benedikt M. P\"{o}tscher\thanks{
Corresponding author.} \\
{\footnotesize University of Vienna} \\
{\footnotesize Department of Statistics} \\
{\footnotesize A-1090 Vienna, Oskar-Morgenstern Platz 1} \\
{\footnotesize [email removed]}
\end{tabular}
\and
\begin{tabular}{c}
David Preinerstorfer \\
{\footnotesize WU Vienna University of Economics and Business} \\
{\footnotesize Institute for Statistics and Mathematics} \\
{\footnotesize A-1020 Vienna, Welthandelsplatz 1} \\
{\footnotesize [email removed]}
\end{tabular}
}
\date{First version: June 2024\\
Second version: August 2024\\
Third version: November 2024\\
Fourth version: September 2025\\
This version: April 2026}
\maketitle

\begin{abstract}
We revisit size controllability results in \cite{PP21} concerning
heteroskedasticity robust test statistics in regression models. For the
special, but important, case of testing a single restriction (e.g., a zero
restriction on a single coefficient), we povide a necessary \emph{and}
sufficient condition for size controllability, whereas the condition in \cite
{PP21} is, in general, only sufficient (even in the case of testing a single
restriction).
\end{abstract}

\section{Introduction\label{Intro}}

Tests and confidence intervals based on so-called heteroskedasticity robust
standard errors date back to \cite{E63, E67} and constitute, at least since
\cite{W80}, a major component of the applied econometrician's toolbox.
Although these early methods come with well-understood large sample
properties, when based on critical values derived from asymptotic theory
their finite sample properties often deviate substantially from what
asymptotic theory suggests: tests may substantially overreject under the
null and corresponding confidence intervals may undercover. Strong leverage
points have been identified early on as one major reason for these
deviations, see, e.g., \cite{MacW85}, \cite{DavidsonMacKinnon1985}, and \cite
{CheshJewitt1987}. This has led to various developments trying to attenuate
such drawbacks:

\begin{enumerate}
\item modifications of the covariance matrix estimators in \cite{E63, E67}
and \cite{W80} led to tests based on what are now frequently called HC1-HC4
covariance matrix estimators (see, e.g., \cite{LE2000}, and \cite{Crib2004}
for an overview of the relevant literature), with HC0 denoting the original
proposal;

\item some authors investigated degree-of-freedom corrections to obtain
modified critical values (e.g., \cite{Satterth} or \cite{BellMcCa}, see also
\cite{Imbkoles2016});

\item wild bootstrap methods were investigated (for an overview of the
relevant literature see \cite{PPBoot}) and, more recently, parametric
bootstrap methods were studied in \cite{Chuetal2021} and \cite{Hansen2021}.
\end{enumerate}

Although these developments sometimes lead to improvements, they come with
no general finite sample guarantees concerning the size of the tests or the
coverage of related confidence intervals, cf.~the discussion in \cite
{PPBoot,PP21} for detailed accounts.

Motivated by this lack of finite sample guarantees, \cite{PP21} studied the
question under which conditions heteroskedasticity robust test statistics as
well as the standard (uncorrected) F-test statistic can actually be paired
with appropriate (finite) critical values, so that one obtains tests that
have their (finite sample) size controlled by the prescribed significance
value ${\Greekmath 010B} $ (i.e., have size $\leq {\Greekmath 010B} $) even though one is
completely agnostic about the form of heteroskedasticity.\footnote{
The null-hypothesis to be tested is given by a set of affine restrictions.}
Under appropriate assumptions on the errors, allowing for Gaussian as well
as substantial non-Gaussian behavior, they have shown that the standard
(uncorrected) F-test statistic can be size-controlled (in finite samples) by
using an appropriately chosen (finite) critical value if and only if the
following simple condition holds:
\begin{eqnarray}
&&\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{\emph{no standard basis vector that lies in the column span of the
design matrix}}  \notag \\
&&\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{\emph{is \textquotedblleft involved\textquotedblright\ in the affine
restrictions to be tested,}}  \label{q:cond}
\end{eqnarray}
see (8) in \cite{PP21} for a formal statement of this condition.

Under a generally \emph{stronger} condition than (\ref{q:cond}) (see (10) in
\cite{PP21}), it was furthermore shown that large classes of
heteroskedasticity robust test statistics (e.g., HC0-HC4) can be
size-controlled by appropriate (finite) critical values. That condition,
however, although satisfied for many testing problems (and even often
identical to (\ref{q:cond}), cf.~Theorem 3.9 and Lemma A.3 in \cite{PP3}),
is \emph{not} necessary in general, as shown in examples given in \cite{PP21}
; e.g., their Example 5.5 or Example C.1 in their Appendix C.\footnote{
Appendices to \cite{PP21} are published in the Supplementary Material
available at the publisher's website of that article.} These examples
consider the case of testing linear contrasts in the expected outcomes of
subjects belonging to two or more groups, scenarios that are practically
relevant. Further examples are provided in Examples A.1-A.4 in Appendix \ref
{app:proofs} further below.\footnote{
Example 5.5 in \cite{PP21} concerns simultaneously testing multiple
retrictions, while Example C.1 in Appendix C of \cite{PP21} as well as
Examples A.1-A.4 in Appendix \ref{app:proofs} of the present article concern
the case of testing a single restriction.}

For the important case of testing problems involving only a \emph{single
restriction }(i.e., the case $q=1$ in the notation of \cite{PP21}), we show
in the present article that the condition in (\ref{q:cond}) is then in fact
necessary and sufficient also for size controllability of the above
mentioned classes of heteroskedasticity robust test statistics, including
HC0-HC4.

\section{Results on size controllability\label{frame}}

\subsection{Framework\label{frm}}

Here we recall the most relevant notions from Sections 2 and 3 of \cite{PP21}
, to which we refer the reader for further information and discussion. We
consider the linear regression model
\begin{equation}
\mathbf{Y}=X{\Greekmath 010C} +\mathbf{U},  \label{lm}
\end{equation}
where $X$ is a (real) nonstochastic regressor (design) matrix of dimension $
n\times k$ and where ${\Greekmath 010C} \in \mathbb{R}^{k}$ denotes the unknown
regression parameter vector. Throughout, we assume $\limfunc{rank}(X)=k$ and
$1\leq k<n$. We furthermore assume that the $n\times 1$ disturbance vector $
\mathbf{U}=(\mathbf{u}_{1},\ldots ,\mathbf{u}_{n})^{\prime }$ ($^{\prime }$
denoting transposition) has mean zero and unknown covariance matrix ${\Greekmath 011B}
^{2}\Sigma $ ($0<{\Greekmath 011B} <\infty $), where $\Sigma $ varies in the
\textquotedblleft heteroskedasticity model\textquotedblright\ given by
\begin{equation*}
\mathfrak{C}_{Het}=\left\{ \limfunc{diag}({\Greekmath 011C} _{1}^{2},\ldots ,{\Greekmath 011C}
_{n}^{2}):{\Greekmath 011C} _{i}^{2}>0\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ for all }i\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, }\sum_{i=1}^{n}{\Greekmath 011C}
_{i}^{2}=1\right\} ,
\end{equation*}
and where $\limfunc{diag}({\Greekmath 011C} _{1}^{2},\ldots ,{\Greekmath 011C} _{n}^{2})$ denotes the
diagonal $n\times n$ matrix with diagonal elements given by ${\Greekmath 011C} _{i}^{2}$.
That is, the disturbances are uncorrelated but can be heteroskedastic of
arbitrary form. [In Appendix \ref{app:B} we shall also consider another
heteroskedasticity model.]\footnote{
Since we are concerned with finite-sample results only, the elements of $
\mathbf{Y}$, $X$, and $\mathbf{U}$ (and even the probability space
supporting $\mathbf{Y}$ and $\mathbf{U}$) may depend on sample size $n$, but
this will not be expressed in the notation. Furthermore, the obvious
dependence of $\mathfrak{C}_{Het}$ on $n$ will also not be shown in the
notation, and the same applies to the heteroskedasticity model defined in
Appendix \ref{app:B}.}

\emph{For ease of exposition, we shall maintain in the sequel that the
disturbance vector }$\mathbf{U}$ \emph{is normally distributed.
Generalizations to classes of non-normal disturbances can be obtained
following the arguments in Section 7.1 of \cite{PP21}, see Remark 2.2
further below.} Denoting a Gaussian probability measure with mean ${\Greekmath 0116} \in
\mathbb{R}^{n}$ and (possibly singular) covariance matrix $A$ by $P_{{\Greekmath 0116} ,A}$
, the collection of distributions on $\mathbb{R}^{n}$ (the sample space of $
\mathbf{Y}$) induced by the linear model just described together with the
Gaussianity assumption is then given by
\begin{equation*}
\left\{ P_{{\Greekmath 0116} ,{\Greekmath 011B} ^{2}\Sigma }:{\Greekmath 0116} \in \mathrm{\limfunc{span}}
(X),0<{\Greekmath 011B} ^{2}<\infty ,\Sigma \in \mathfrak{C}_{Het}\right\} ,
\end{equation*}
where $\mathrm{\limfunc{span}}(X)$ denotes the column space of $X$.\footnote{
Since every $\Sigma \in \mathfrak{C}_{Het}$ is positive definite, the
measure $P_{{\Greekmath 0116} ,{\Greekmath 011B} ^{2}\Sigma }$ is absolutely continuous with respect
to Lebesgue measure on $\mathbb{R}^{n}$.}

We focus on testing the null $R{\Greekmath 010C} =r$ against the alternative $R{\Greekmath 010C}
\neq r$, where $R\neq 0$ is a $1\times k$ vector and $r\in \mathbb{R}$. That
is, \emph{throughout this paper we focus on testing a single restriction,
whereas the theory developed in \cite{PP21} allows for simultaneously
testing multiple restrictions} (that is, we here consider only the special
case corresponding to $q=1$ in \cite{PP21}). Set $\mathfrak{M}=\limfunc{span}
(X)$, define the affine space
\begin{equation*}
\mathfrak{M}_{0}=\left\{ {\Greekmath 0116} \in \mathfrak{M}:{\Greekmath 0116} =X{\Greekmath 010C} \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ and }R{\Greekmath 010C}
=r\right\} ,
\end{equation*}
and let
\begin{equation*}
\mathfrak{M}_{1}=\left\{ {\Greekmath 0116} \in \mathfrak{M}:{\Greekmath 0116} =X{\Greekmath 010C} \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ and }R{\Greekmath 010C}
\neq r\right\} .
\end{equation*}
Adopting these definitions, the testing problem we consider can be written
more precisely as
\begin{equation}
H_{0}:{\Greekmath 0116} \in \mathfrak{M}_{0},\ 0<{\Greekmath 011B} ^{2}<\infty ,\ \Sigma \in
\mathfrak{C}_{Het}\quad \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ vs. }\quad H_{1}:{\Greekmath 0116} \in \mathfrak{M}_{1},\
0<{\Greekmath 011B} ^{2}<\infty ,\ \Sigma \in \mathfrak{C}_{Het}.
\label{testing problem}
\end{equation}
We also write~$\mathfrak{M}_{0}^{lin}=\mathfrak{M}_{0}-{\Greekmath 0116} _{0}=\left\{
X{\Greekmath 010C} :R{\Greekmath 010C} =0\right\} $ where ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$. Of course,
$\mathfrak{M}_{0}^{lin}$ does not depend on the choice of ${\Greekmath 0116} _{0}\in
\mathfrak{M}_{0}$. Furthermore, if $\mathcal{L}$ is a linear subspace of $
\mathbb{R}^{n}$, $\Pi _{\mathcal{L}}$ denotes the orthogonal projection onto
$\mathcal{L}$, while $\mathcal{L}^{\bot }$ denotes the orthogonal complement
of $\mathcal{L}$ in $\mathbb{R}^{n}$.

The assumption of nonstochastic regressors made above entails little loss of
generality, and results for models with stochastic regressors can be
obtained from the ones derived in the present paper by the same arguments as
the ones given in Section 7.2 of \cite{PP21}.

\subsection{Test statistics, size controllability, and a new result\label
{Sec_2.2}}

We consider the same test statistics as in Section 3 of \cite{PP21}.
Simplified to the setting of testing a \emph{single} restriction considered
in the present article, they are given by
\begin{equation}
T_{Het}\left( y\right) =\left\{
\begin{array}{cc}
(R\hat{{\Greekmath 010C}}\left( y\right) -r)^{2}/\hat{\Omega}_{Het}(y) & \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{if }\hat{
\Omega}_{Het}\left( y\right) \neq 0, \\
0 & \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{if }\hat{\Omega}_{Het}\left( y\right) =0,
\end{array}
\right.  \label{T_het}
\end{equation}
where $\hat{{\Greekmath 010C}}(y)=\left( X^{\prime }X\right) ^{-1}X^{\prime }y$ and
where $\hat{\Omega}_{Het}(y)=R\hat{\Psi}_{Het}(y)R^{\prime }$. Here
\begin{equation*}
\hat{\Psi}_{Het}\left( y\right) =(X^{\prime }X)^{-1}X^{\prime }\limfunc{diag}
\left( d_{1}\hat{u}_{1}^{2}\left( y\right) ,\ldots ,d_{n}\hat{u}
_{n}^{2}\left( y\right) \right) X(X^{\prime }X)^{-1},
\end{equation*}
with $\hat{u}(y)=\left( \hat{u}_{1}(y),\ldots ,\hat{u}_{n}(y)\right)
^{\prime }=y-X\hat{{\Greekmath 010C}}(y)$. The constants $d_{i}>0$ sometimes depend on
the design matrix; see \cite{PP21} for examples of the weights $d_{i}$,
including HC0-HC4 weights. We also recall the following assumption from the
latter reference, again specialized to the setting of testing only a \emph{
single} restriction (i.e., to the case $q=1$ in the notation of \cite{PP21}).

\begin{assumption}
\label{R_and_X}Let $1\leq i_{1}<\ldots <i_{s}\leq n$ denote all the indices
for which $e_{i_{j}}(n)\in \limfunc{span}(X)$ holds where $e_{j}(n)$ denotes
the $j$-th standard basis vector in $\mathbb{R}^{n}$. If no such index
exists, set $s=0$. Let $X^{\prime }\left( \lnot (i_{1},\ldots i_{s})\right) $
denote the matrix which is obtained from $X^{\prime }$ by deleting all
columns with indices $i_{j}$, $1\leq i_{1}<\ldots <i_{s}\leq n$ (if $s=0$,
no column is deleted). Then $R(X^{\prime }X)^{-1}X^{\prime }\left( \lnot
(i_{1},\ldots i_{s})\right) \neq 0$ holds.
\end{assumption}

This assumption can be checked in any particular application as it only
depends on the observable quantities $R$ and $X$; and a sufficient condition
for Assumption \ref{R_and_X} obviously is $s=0$. Assumption \ref{R_and_X} is
unavoidable if one wants to obtain a sensible test from the statistic $
T_{Het}$, see Section 3 of \cite{PP21} for more discussion. We note that $
e_{j}(n)\in \limfunc{span}(X)$ is equivalent to $h_{jj}=1$, where $h_{jj}$
denotes the $j$-th diagonal element of the `hat matrix' $H=X(X^{\prime
}X)^{-1}X^{\prime }$.\footnote{
This follows from $h_{jj}=e_{j}(n)^{\prime }He_{j}(n)=(He_{j}(n))^{\prime
}He_{j}(n)$ and the fact that $H$ represents the orthogonal projection onto $
\limfunc{span}(X)$.}

As in \cite{PP21}, we introduce
\begin{equation*}
B(y)=R(X^{\prime }X)^{-1}X^{\prime }\limfunc{diag}\left( \hat{u}
_{1}(y),\ldots ,\hat{u}_{n}(y)\right) .
\end{equation*}
Define (recall that $R$ is a nonzero row vector in this article)
\begin{equation*}
\mathsf{B}=\left\{ y\in \mathbb{R}^{n}:\limfunc{rank}(B(y))<1\right\}
=\left\{ y\in \mathbb{R}^{n}:B(y)=0\right\} .
\end{equation*}
It is now easy to see that $\limfunc{span}(X)\subseteq \mathsf{B}$ and that $
\mathsf{B}$ is a linear space (cf.~also Lemma 3.1 in \cite{PP21}). Simple
examples can be constructed to show that $\limfunc{span}(X)\neq \mathsf{B}$,
in general; cf.~Example C.1 in Appendix C of \cite{PP21} as well as Examples
A.1-A.4 in Appendix \ref{app:proofs} further below.

To summarize the main size controllability statements from \cite{PP21} for
the above class of test statistics, we first have to recall the following
notation: For a given linear subspace $\mathcal{L}$ of $\mathbb{R}^{n}$ we
define the set of indices $I_{0}(\mathcal{L})$ via
\begin{equation}
I_{0}(\mathcal{L})=\left\{ i:1\leq i\leq n,e_{i}(n)\in \mathcal{L}\right\} .
\label{eqn:I0def}
\end{equation}
We set $I_{1}(\mathcal{L})=\left\{ 1,\ldots ,n\right\} \backslash I_{0}(
\mathcal{L})$. Clearly, $\func{card}(I_{0}(\mathcal{L}))\leq \dim (\mathcal{L
})$ holds. And $I_{1}(\mathcal{L})$ is nonempty provided $\dim (\mathcal{L}
)<n$; in particular, $I_{1}(\mathfrak{M}_{0}^{lin})$ is always nonempty
since $\dim (\mathfrak{M}_{0}^{lin})=k-1<n-1$. The results in \cite{PP21}
concerning size controllability of tests for (\ref{testing problem}) based
on $T_{Het}$ can now be summarized as follows; some intuition for why size
control cannot always be achieved is provided further below as well as in
Section 4 in \cite{PP21}:

\begin{theorem}[Theorem 5.1(b,c) and Propositions 5.5(b) and 5.7(b) in
\protect\cite{PP21} for the case $q=1$]
\label{thm:pp21}\footnote{
The corresponding results in \cite{PP21} for $q\geq 1$ take exactly the same
form, but with the definitions of the relevant quantities adapted to that
more general setting.} Suppose that Assumption \ref{R_and_X} is satisfied.
Then the following statements hold:

\begin{enumerate}
\item For every $0<{\Greekmath 010B} <1$ there exists a real number $C({\Greekmath 010B} )$ such
that
\begin{equation}
\sup_{{\Greekmath 0116} _{0}\in \mathfrak{M}_{0}}\sup_{0<{\Greekmath 011B} ^{2}<\infty }\sup_{\Sigma
\in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }(T_{Het}\geq C({\Greekmath 010B}
))\leq {\Greekmath 010B}  \label{size-control_Het}
\end{equation}
holds, provided that
\begin{equation}
e_{i}(n)\notin \mathsf{B}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ \ \ for every \ }i\in I_{1}(\mathfrak{M}
_{0}^{lin}).  \label{non-incl_Het}
\end{equation}
Furthermore, under condition (\ref{non-incl_Het}), even equality can be
achieved in (\ref{size-control_Het}) by a proper choice of $C({\Greekmath 010B} )$,
provided ${\Greekmath 010B} \in (0,{\Greekmath 010B} ^{\ast }]\cap (0,1)$ holds, where
\begin{equation}
{\Greekmath 010B} ^{\ast }=\sup_{C\in (C^{\ast },\infty )}\sup_{\Sigma \in \mathfrak{C}
_{Het}}P_{{\Greekmath 0116} _{0},\Sigma }(T_{Het}\geq C)  \label{alpha}
\end{equation}
is positive and where
\begin{equation}
C^{\ast }=\max \{T_{Het}({\Greekmath 0116} _{0}+e_{i}(n)):i\in I_{1}(\mathfrak{M}
_{0}^{lin})\}  \label{Cstar}
\end{equation}
for ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$ (with neither ${\Greekmath 010B} ^{\ast }$ nor $
C^{\ast }$ depending on the choice of ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$).

\item Suppose (\ref{non-incl_Het}) is satisfied. Then a smallest critical
value, denoted by $C_{\Diamond }({\Greekmath 010B} )$, satisfying (\ref
{size-control_Het}) exists for every $0<{\Greekmath 010B} <1$. And $C_{\Diamond
}({\Greekmath 010B} )$ is also the smallest among the critical values leading to
equality in (\ref{size-control_Het}) whenever such critical values exist.

\item Suppose (\ref{non-incl_Het}) is satisfied. Then any $C({\Greekmath 010B} )$
satisfying (\ref{size-control_Het}) necessarily has to satisfy $C({\Greekmath 010B}
)\geq C^{\ast }$. In fact, for any $C<C^{\ast }$ we have $\sup_{\Sigma \in
\mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B} ^{2}\Sigma }(T_{Het}\geq C)=1$ for
every ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$ and every ${\Greekmath 011B} ^{2}\in (0,\infty )$.

\item If the condition
\begin{equation}
e_{i}(n)\notin \func{span}(X)\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ \ \ for every \ }i\in I_{1}(\mathfrak{M}
_{0}^{lin})  \label{non-incl_Het_uncorr}
\end{equation}
is violated, then $\sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B}
^{2}\Sigma }(T_{Het}\geq C)=1$ for \emph{every} choice of critical value $C$
, every ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$, and every ${\Greekmath 011B} ^{2}\in (0,\infty
) $ (implying that size equals $1$ for every $C$).\footnote{\label{FN}It is
understood here that critical values are less than infinity.}
\end{enumerate}
\end{theorem}

To obtain some intuition for Theorem \ref{thm:pp21}, recall that the
diagonal elements of $\Sigma \in \mathfrak{C}_{Het}$ are positive and sum up
to one (by definition). Now, for a matrix $\Sigma $ with $i$-th diagonal
entry close to $1$, all other diagonal entries must therefore be close to $0$
, so that $\Sigma \approx e_{i}(n)e_{i}(n)^{\prime }$ then holds. Note that
if $\Sigma \approx e_{i}(n)e_{i}(n)^{\prime }$, the distribution $P_{{\Greekmath 0116}
,{\Greekmath 011B} ^{2}\Sigma }$ of the data is strongly \textquotedblleft
concentrated\textquotedblright\ around the one-dimensional space ${\Greekmath 0116} +
\limfunc{span}(e_{i}(n))$. From an intuitive point of view, whether a given
test statistic admits a size-controlling critical value or not, should
therefore depend on the \textquotedblleft behavior\textquotedblright\ of the
test statistic for values on or close to the spaces ${\Greekmath 0116} _{0}+\limfunc{span}
(e_{i}(n))$ with ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$. It turns out that this is
intimately related to (\ref{non-incl_Het}) and (\ref{non-incl_Het_uncorr}).
See Section 4 in \cite{PP21} for more discussion.

Most importantly, the above theorem shows that, given Assumption~\ref
{R_and_X}, the condition in (\ref{non-incl_Het}) is sufficient for the
existence of a (finite) size-controlling critical value $C({\Greekmath 010B} )$
satisfying (\ref{size-control_Het}), while the weaker condition (\ref
{non-incl_Het_uncorr}) is necessary. Furthermore, in case the design matrix $
X$ and the vector $R$ are such that $\mathsf{B}=\limfunc{span}(X)$, and
hence the condition in (\ref{non-incl_Het}) coincides with that in (\ref
{non-incl_Het_uncorr}), the condition (\ref{non-incl_Het}) is also
necessary. However, $\mathsf{B}=\limfunc{span}(X)$ is not always true (see
Example C.1 in Appendix C of \cite{PP21} or the examples in Appendix \ref
{app:proofs} further below), although the equality holds generically
(cf.~Theorem 3.9 and Lemma A.3 in \cite{PP3}). We now show in the subsequent
theorem that in the situation considered in this article, namely testing
only a single restriction, the condition in (\ref{non-incl_Het}) in Theorem
\ref{thm:pp21} can actually always be replaced by that in (\ref
{non-incl_Het_uncorr}). Before we present that theorem, we discuss an
equivalent formulation of condition (\ref{non-incl_Het_uncorr}) that is
expressed in terms of \emph{certain} diagonal elements of the `hat matrix' $H
$, see (\ref{sol}) below.\footnote{
An informal verbal description of (\ref{non-incl_Het_uncorr}) is given in (
\ref{q:cond})\ in the Introduction.}

\textbf{Remark 2.1:} (i) Condition (\ref{non-incl_Het_uncorr}) is equivalent
to "$h_{ii}<1$ for every $i\in I_{1}(\mathfrak{M}_{0}^{lin})$".\footnote{
Note that $h_{ii}=1$ always holds if $i\in I_{0}(\mathfrak{M}_{0}^{lin})$.}

(ii) Condition (\ref{non-incl_Het_uncorr}) can also equivalently be written
as
\begin{equation}
e_{i}(n)\notin \func{span}(X)\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ for every }i\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ satisfying }
R(X^{\prime }X)^{-1}x_{i\cdot }^{\prime }\neq 0,  \label{sol0}
\end{equation}
see Remark B.1(iii) in Appendix \ref{app:B} further below.\footnote{
Comparing (\ref{non-incl_Het_uncorr}) and (\ref{sol0}) could lead one to
conjecture equivalence of the conditions $i\in I_{1}(\mathfrak{M}_{0}^{lin})$
and $R(X^{\prime }X)^{-1}x_{i\cdot }^{\prime }\neq 0$. This is incorrect in
general, see Example A.1 in Appendix A. However, $R(X^{\prime
}X)^{-1}x_{i\cdot }^{\prime }\neq 0$ implies $i\in I_{1}(\mathfrak{M}
_{0}^{lin})$, see Part 3 of Lemma \ref{gensol}.} And this in turn is now
equivalent to
\begin{equation}
h_{ii}<1\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ for every }i\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ satisfying }R(X^{\prime }X)^{-1}x_{i\cdot
}^{\prime }\neq 0.  \label{sol}
\end{equation}
The last form of the condition may be more appealing to some readers. We
issue a warning here, however, namely that the condition (\ref{non-incl_Het}
) is, in general, stronger than the condition "$e_{i}(n)\notin \mathsf{B}$
for every $i$ satisfying $R(X^{\prime }X)^{-1}x_{i\cdot }^{\prime }\neq 0$",
see Remark B.1(iv) in Appendix \ref{app:B}.

We now present the announced theorem.

\begin{theorem}
\label{thm:pp24}\footnote{
Following the suggestion of some readers, we mention the equivalent
condition (\ref{sol}) explicitly in this theorem, although the equivalence
with (\ref{non-incl_Het_uncorr}) has already been noted in Remark 2.1.}
Suppose that Assumption \ref{R_and_X} is satisfied. Then the following
statements hold:

\begin{enumerate}
\item For every $0<{\Greekmath 010B} <1$ there exists a real number $C({\Greekmath 010B} )$ such
that (\ref{size-control_Het}) holds, provided that (\ref{non-incl_Het_uncorr}
) (or equivalently (\ref{sol})) holds. Furthermore, under condition (\ref
{non-incl_Het_uncorr}) (or equivalently (\ref{sol})), even equality can be
achieved in (\ref{size-control_Het}) by a proper choice of $C({\Greekmath 010B} )$,
provided ${\Greekmath 010B} \in (0,{\Greekmath 010B} ^{\ast }]\cap (0,1)$ holds, where ${\Greekmath 010B}
^{\ast }$ given by (\ref{alpha}) is positive and where $C^{\ast }$is given
by (\ref{Cstar}) for ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$ (with neither ${\Greekmath 010B}
^{\ast }$ nor $C^{\ast }$ depending on the choice of ${\Greekmath 0116} _{0}\in \mathfrak{M
}_{0}$).

\item Suppose (\ref{non-incl_Het_uncorr}) (or equivalently (\ref{sol})) is
satisfied. Then a smallest critical value, denoted by $C_{\Diamond }({\Greekmath 010B}
) $, satisfying (\ref{size-control_Het}) exists for every $0<{\Greekmath 010B} <1$. And
$C_{\Diamond }({\Greekmath 010B} )$ is also the smallest among the critical values
leading to equality in (\ref{size-control_Het}) whenever such critical
values exist.

\item Suppose (\ref{non-incl_Het_uncorr}) (or equivalently (\ref{sol})) is
satisfied. Then any $C({\Greekmath 010B} )$ satisfying (\ref{size-control_Het})
necessarily has to satisfy $C({\Greekmath 010B} )\geq C^{\ast }$. In fact, for any $
C<C^{\ast }$ we have $\sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B}
^{2}\Sigma }(T_{Het}\geq C)=1$ for every ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$ and
every ${\Greekmath 011B} ^{2}\in (0,\infty )$.

\item If (\ref{non-incl_Het_uncorr}) (or equivalently (\ref{sol})) is
violated, then $\sup_{\Sigma \in \mathfrak{C}_{Het}}P_{{\Greekmath 0116} _{0},{\Greekmath 011B}
^{2}\Sigma }(T_{Het}\geq C)=1$ for \emph{every} choice of critical value $C$
, every ${\Greekmath 0116} _{0}\in \mathfrak{M}_{0}$, and every ${\Greekmath 011B} ^{2}\in (0,\infty
) $ (implying that size equals $1$ for every $C$).\footnote{
Cf. Footnote \ref{FN}.}
\end{enumerate}
\end{theorem}

The main take-away of Theorem \ref{thm:pp24} is that, given Assumption \ref
{R_and_X} holds, the condition in (\ref{non-incl_Het_uncorr}) (or
equivalently (\ref{sol})) is necessary and sufficient for the existence of a
(smallest) finite size-controlling critical value when one is testing only a
single restriction.\footnote{
By contraposition, the design matrices $X$ and restrictions $R$ for which
size control fails are precisely characterized by failure of\ (\ref
{non-incl_Het_uncorr}) (or equivalently (\ref{sol})). One example is when $X$
contains the dummy $e_{i}(n)$ as its first column, say, and $R=(1,0,\ldots
,0)$ (or, more generally, $R$ has a non-zero first entry). Another example
arises when the first two columns of $X$ are given by $(1,\ldots ,1)^{\prime
}$ and $(1,-1,\ldots ,-1)^{\prime }$, and the first two entries of $R$ are
both equal to $1$.} The condition "$e_{i}(n)\notin \func{span}(X)$ for every
$i=1,\ldots ,n$" (which is tantamount to "$h_{ii}<1$ for every $i=1,\ldots ,n
$") implies (\ref{non-incl_Het_uncorr}), and thus is sufficient for
size-controllability of $T_{Het}$ (but not necessary, see, e.g., Example
A.2). Note that the conditions in (\ref{non-incl_Het}), (\ref
{non-incl_Het_uncorr}), as well as (\ref{sol}) do not depend on the weights
used in the construction of the covariance matrix estimator or on $r$. They
only depend on $X$ and $R$. This and more (e.g., how the conditions relate
to high-leverage points) is discussed subsequent to Theorem 5.1 (and in
Remarks 5.2-5.4, 5.6, and 5.9) in \cite{PP21} to which we refer the reader
for a detailed account. As a point of interest we also note that condition (
\ref{non-incl_Het_uncorr}) given above is exactly the same as condition (8)
in \cite{PP21} (with $q=1$); in that reference, the latter condition is
shown to be necessary and sufficient for size control of the standard
(uncorrected) F-test statistic (regardless of whether $q=1$ or not).

We also note here that Theorem \ref{thm:pp24} disproves -- for the special
case of testing a single restriction -- a conjecture in Remark 5.8 of \cite
{PP21}, namely that there would exist cases where Assumption \ref{R_and_X}
holds, (\ref{non-incl_Het_uncorr}) is satisfied, (\ref{non-incl_Het}) does
not hold, and size control by a (finite) critical value is not possible.

To see why the refinement of Theorem \ref{thm:pp21} provided in Theorem \ref
{thm:pp24} can matter in practice, it is enough to consider the textbook
example of a matrix $X$ with two columns, the first indicating membership to
the treatment group and the second indicating membership to the control
group (a special case of Example C.1 in Appendix C of \cite{PP21}). Assume
that the first $n_{1}\geq 2$ observations belong to the treatment group and
the remaining $n_{2}\geq 2$ observations belong to the control group. Assume
further that one wants to test whether ${\Greekmath 010C} _{1}$, the expected outcome of
the treatment group, equals a given value (e.g., because one wants to obtain
a confidence interval through test inversion). Example C.1 in Appendix C of
\cite{PP21} shows that in this case
\begin{equation*}
\mathcal{I}_{1}(\mathfrak{M}_{0}^{lin})=\{1,\ldots ,n\}\quad \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ and }
\quad \mathsf{B}=\{y\in \mathbb{R}^{n}:y_{1}=\ldots =y_{n_{1}}\}\neq \func{
span}(X).
\end{equation*}
In particular, $e_{i}(n)\in \mathsf{B}$ if and only if $i>n_{1}$, so that (
\ref{non-incl_Het}) is not satisfied, while (\ref{non-incl_Het_uncorr})
holds, and size-controlling critical values hence exist by Theorem \ref
{thm:pp24} (and can be used for constructing confidence intervals). Further
examples are provided in Appendix \ref{app:proofs} below.

We next explain the key observation underlying the proof of Theorem \ref
{thm:pp24}: To this end, define the (possibly empty) set of indices
\begin{equation*}
\mathcal{I}_{\#}=\left\{ i:1\leq i\leq n,~R(X^{\prime }X)^{-1}x_{i\cdot
}^{\prime }=0\right\} ,
\end{equation*}
where $x_{i\cdot }$ denotes the $i$-th row of $X$, and define (the span of
the empty set will throughout be interpreted as $\{0\}$) the space
\begin{equation}
\mathcal{V}_{\#}=\limfunc{span}\left( \{e_{i}(n):i\in \mathcal{I}
_{\#},~e_{i}(n)\in \mathsf{B}\}\right) \subseteq \mathsf{B},
\label{eqn:Vdef}
\end{equation}
the inclusion holding because $\mathsf{B}$ is a linear space as noted
earlier (recall that $R$ is $1\times k$ dimensional in this article).
\footnote{
We note that $\mathcal{I}_{\#}$ is a proper subset of $\{1,\ldots ,n\}$
since $R\neq 0$.} Recall that under Assumption \ref{R_and_X} the test
statistic $T_{Het}$ as well as $\mathsf{B}$ are invariant with respect to
(w.r.t.) the group $G(\mathfrak{M}_{0})$ (i.e., the group of transformations
$y\mapsto {\Greekmath 010E} (y-{\Greekmath 0116} _{0})+{\Greekmath 0116} _{0}^{\ast }$ with ${\Greekmath 010E} \in \mathbb{R}$
nonzero and ${\Greekmath 0116} _{0}$ and ${\Greekmath 0116} _{0}^{\ast }$ in $\mathfrak{M}_{0}$), see
Remark C.1 in Appendix C of \cite{PP21}.\footnote{
The invariance holds trivially if Assumption \ref{R_and_X} is violated.} The
results in \cite{PP21} are based on this invariance property. The crucial
observation exploited in the proof of Theorem \ref{thm:pp24} now is that, in
the special case of testing a single restriction considered in this article,
the test statistic $T_{Het}$ as well as $\mathsf{B}$ are invariant, not only
w.r.t.~$G(\mathfrak{M}_{0})$, but also w.r.t.~addition of elements of $
\mathcal{V}_{\#}$. This additional invariance property involving $\mathcal{V}
_{\#}$, paired with a careful application of the general theory for
size-controlling critical values in \cite{PP3}, then allows us to deduce the
refined statement in Theorem \ref{thm:pp24}. It turns out fortunate that the
general theory in \cite{PP3} explicitly allows one to incorporate additional
invariance properties beyond $G(\mathfrak{M}_{0})$. For details and proofs
the reader is referred to Appendices \ref{app:proofs} and \ref{app:B}.

Finally, we remark that Theorem \ref{thm:pp24} is deduced from Theorem \ref
{thm:moregen} in Appendix \ref{app:B}, which is a more general statement
that also allows for heteroskedasticity models other than $\mathfrak{C}
_{Het} $ (and which are defined in (\ref{eqn:covmodmg}) below).

\bigskip

\textbf{Remark 2.2:} \emph{(Extensions to non-Gaussian errors) }(i) All the
theorems in this article continue to hold as they stand, if the disturbance
vector $\mathbf{U}$ follows an elliptically symmetric distribution that has
no atom at the origin; more precisely, $\mathbf{U}$ is assumed to be
distributed as ${\Greekmath 011B} \Sigma ^{1/2}\mathbf{z}$, where $\mathbf{z}$ has a
spherically symmetric distribution on $\mathbb{R}^{n}$ that has no atom at
the origin, and where ${\Greekmath 011B} $ and $\Sigma $ are as in Section \ref{frm}.
This is so, since the size under Gaussianity is the same as the size under
the elliptical symmetry assumption. In particular, the smallest
size-controlling critical values under the elliptical symmetry assumption
coincide with the smallest size-controlling critical values under
Gaussianity, and thus can be computed from the algorithms relying on
Gaussianity described in \cite{PP21}. See Appendix E.1 of \cite{PP3} and
Section 7.1(i) of \cite{PP21} for more details. The same is actually true
for a wider class of distribution for $\mathbf{U}$, namely where $\mathbf{z}$
has a distribution in the class $Z_{ua}$ defined in Appendix E.1 of \cite
{PP3}.

(ii) All the theorems in this article except for Theorem \ref{thm:moregen}
in Appendix \ref{app:B} (i.e., all theorems using the heteroskedasticity
model $\mathfrak{C}_{Het}$) continue to hold as they stand, if it is assumed
that the disturbance vector $\mathbf{U}$ follows a distribution from the
semiparametric model defined in Section 7.1(iv) in \cite{PP21} (a model that
contains inter alia all distributions corresponding to i.i.d. samples of
scale-mixtures of normals). Again, this is so since the size under
Gaussianity is the same as the size under this semiparametric model. In
particular, the smallest size-controlling critical values under this
semiparametric model coincide with the smallest size-controlling critical
values under Gaussianity, and thus can be computed from the algorithms
relying on Gaussianity described in \cite{PP21}. See Section 7.1(iv) in \cite
{PP21} and note that the Gaussian model is a submodel of the semiparametric
model considered there.

(iii) Furthermore, as discussed in detail in Appendix E.2 of \cite{PP3}, any
condition sufficient for size controllability under Gaussianity of the
disturbance vector $\mathbf{U}$ also implies size controllability for large
classes of distributions for $\mathbf{U}$ that satisfy appropriate
domination conditions; however, the corresponding size-controlling critical
values may then differ from the size-controlling critical values that apply
under Gaussianity.

\section{Conclusion}

In the case of testing a \emph{single} restriction, we have shown that the
sufficient condition for size controllability of heteroskedasticity robust
test statistics in \cite{PP21} can be replaced by a weaker sufficient
condition that is also necessary. This allows one -- in the case of testing
a single restriction -- to resolve the question of existence of (finite)
size-controlling critical values in all cases, including those that remain
inconclusive under the results in \cite{PP21}.

We finally remark that the algorithms designed to compute size-controlling
critical values as discussed in Section 10 and Appendix E of \cite{PP21} can
be used as they stand also in situations where (a single restriction is
tested and) size controllability has been verified through checking
condition (\ref{non-incl_Het_uncorr}) (or equivalently (\ref{sol})) and
appealing to Theorem~\ref{thm:pp24}, but where (\ref{non-incl_Het}) does not
hold. This is so since the discussion of the before mentioned algorithms in
\cite{PP21} only requires existence of a (finite) size-controlling critical
value, but does not depend on the way this existence is verified.