EconBase
← Back to paper

A Modern Gauss-Markov Theorem? Really?

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

58,570 characters · 6 sections · 122 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Modern Gauss-Markov Theorem? Really?

abstractWe show that the theorems in Hansen_1.5 (the version accepted by Econometrica), except for one, are not new as they coincide with classical theorems like the good old Gauss-Markov or Aitken Theorem, respectively; the exceptional theorem is incorrect. Hansen_2 corrects this theorem. As a result, all theorems in the latter version coincide with the above mentioned classical theorems. Furthermore, we also show that the theorems in Hansen_3 (the version published in Econometrica) either coincide with the classical theorems just mentioned, or contain extra assumptions that are alien to the Gauss-Markov or Aitken Theorem.

Introduction

Hansen_1.5, Hansen_2, Hansen_3 contain several assertions from which he claims it would follow that the linearity condition can be dropped from the Gauss-Markov Theorem or from the Aitken Theorem. We show that this conclusion is unwarranted, as his assertions on which this conclusion rests turn out to be only (intransparent) reformulations of the classical Gauss-Markov or the classical Aitken Theorem, into which he has reintroduced linearity through the backdoor, or contain extra assumptions alien to the Gauss-Markov or Aitken Theorem.

The present paper is mainly pedagogical in nature. In particular, the results will not come as a surprise to anyone well-versed in the theory of linear models and familiar with basic concepts of statistical decision theory, but -- given the confusion introduced by Hansen_1.5, Hansen_2, Hansen_3 -- the paper will benefit the econometrics community.

One important upshot of the present paper is that one should not follow Hansen's plea to drop the linearity condition in teaching the Gauss-Markov Theorem or the Aitken Theorem. Depending on which formulation of the Gauss-Markov Theorem one starts with (Theorem (ref) or (ref) given below), dropping linearity from the formulation of that theorem at best leads to a result equivalent to the usual Gauss-Markov Theorem, and at worst leads to an incorrect result. The same goes for the Aitken Theorem. Unfortunately, in heeding his own advice Hansen has included an incorrect formulation of the Gauss-Markov Theorem in the August 2021 version of his forthcoming text-book (Theorem 4.4. in HABook) available on his webpage for an extended time period.

Hansen_1.5 is the version accepted by Econometrica and which has been available on Econometrica's webpage of forthcoming papers. Hansen_2 is an updated version that corrects an incorrect result in Hansen_1.5 (but otherwise is identical to the latter paper), and is available from Hansen's webpage. Hansen_3 refers to the version finally published in Econometrica, which contains several nontrivial changes relative to Hansen_1.5, Hansen_2 introduced into the paper at the proof-reading stage. Because Hansen_1.5, Hansen_2 have been widely circulated and discussed, and because Hansen_3 has been published in Econometrica, there is a need to discuss all three versions. We shall start by first discussing Hansen_1.5, Hansen_2 in Sections (ref) and (ref), a discussion that has considerable bearings also on Hansen_3. We then move on to discuss the changes introduced into Hansen_3 at the proof-reading stage and their ramifications in Section (ref). Section (ref) discusses the situation when one restricts attention to independent identically distributed errors.

After the first version of this paper had been circulated, Stephen Portnoy sent us a paper of his (Portnoy) that has a result somewhat similar to our Theorem (ref) with a different proof. For a discussion see Section (ref).

The Framework

As in Hansen_1.5, Hansen_2 we consider throughout the paper the linear regression model

equation[equation omitted — 42 chars of source]

where $Y$ is of dimension $n\times 1$ and $X$ is a (non-random) $n\times k$ design matrix with full column rank $k$ satisfying $1\leq k<n$.\footnote{ We make the assumption $k<n$ in order to use exactly the same framework as in Hansen's papers.} It is assumed that

equation[equation omitted — 36 chars of source]

and

equation[equation omitted — 62 chars of source]

where $\sigma ^{2}$, $0<\sigma ^{2}<\infty $, is unknown and $\Sigma $ is a known symmetric and positive definite $n\times n$ matrix ($Ee^{\prime }e<\infty $). This model implies a distribution $F$ for $Y$, which, for the given $X$, depends on $\beta $ and the distribution of $e$, in particular on $\sigma ^{2}$ and $\Sigma $. Now define $\mathbf{F}_{2}(\Sigma )$ as the class of all such distributions $F$ when $\beta $ varies through $\mathbb{R} ^{k}$ and the distribution of $e$ varies through all distributions compatible with ((ref)) and ((ref)) for the given $\Sigma $ (and arbitrary $\sigma ^{2}$, $0<\sigma ^{2}<\infty $). We furthermore introduce the set $\mathbf{F}_{2}$ as the larger class where we also vary $\Sigma $ through the set of all symmetric and positive definite $n\times n$ matrices. In other words,

equation*[equation* omitted — 74 chars of source]

where the union is taken over all symmetric and positive definite $n\times n$ matrices.\footnote{ Note that $\mathbf{F}_{2}(\Sigma _{1})\cap \mathbf{F}_{2}(\Sigma _{2})=\emptyset $ iff $\Sigma _{1}$ and $\Sigma _{2}$ are not proportional. And $\mathbf{F}_{2}(\Sigma _{1})=\mathbf{F}_{2}(\Sigma _{2})$ iff $\Sigma _{1}$ and $\Sigma _{2}$ are proportional.} [Of course, $\mathbf{F} _{2}(\Sigma )$ as well as $\mathbf{F}_{2}$ also depend on the given $X$, but this dependence is not shown in the notation.] The set $\mathbf{F}_{2}^{0}$ defined in Hansen_1.5, Hansen_2 is nothing else than $\mathbf{F} _{2}(I_{n})$, where $I_{n}$ denotes the $n\times n$ identity matrix. \footnote{ Note that in Hansen_3 the symbol $\mathbf{F}_{2}^{0}$ is used to denote a different set of distributions; see Section (ref) below.} In the following $E_{F}$ ($Var_{F}$, respectively) will denote the expectation (variance-covariance matrix, respectively) taken under the distribution $F$. A word on notation: Given $F\in \mathbf{F}_{2}$, there is a unique $\beta $, denoted by $\beta (F)$, and a unique $\sigma ^{2}\Sigma $ , denoted by $(\sigma ^{2}\Sigma )(F)$, compatible with the distribution $F$.

remark(Ambiguity in the definition in Hansen_1.5, Hansen_2)\ Hansen_1.5, Hansen_2 also define a set $\mathbf{F}_{2}$, unfortunately somewhat ambiguously: Taking the first sentence mentioning his set $\mathbf{F }_{2}$ literally, his set would coincide with our $\mathbf{F}_{2}(\Sigma )$. The two sentences following that sentence, however, intimate that his set $ \mathbf{F}_{2}$ was intended to coincide with our set $\mathbf{F}_{2}$. This is confirmed by an inspection of his proofs; furthermore, if one would interpret his set $\mathbf{F}_{2}$ to mean our $\mathbf{F}_{2}(\Sigma )$, then the relation $\mathbf{F}_{2}^{0}\subset \mathbf{F}_{2}$ given below (4) in Hansen_1.5, Hansen_2 (which in our notation would become $\mathbf{F }_{2}(I_{n})\subset \mathbf{F}_{2}(\Sigma )$) could not hold (except for $ \Sigma $ proportional to $I_{n}$). In the following we hence interpret Hansen's set $\mathbf{F}_{2}$ to coincide with our definition of $\mathbf{F} _{2}$. In a remark further below we discuss what happens if one would adopt the interpretation of Hansen's $\mathbf{F}_{2}$ as coinciding with our $ \mathbf{F}_{2}(\Sigma )$.

The Gauss-Markov Case

To focus the discussion, we first treat the situation of a regression model with homoskedastic and uncorrelated errors, i.e., we assume that in ((ref)) we have

equation[equation omitted — 44 chars of source]

Let $\hat{\beta}_{OLS}=(X^{\prime }X)^{-1}X^{\prime }Y$ denote the ordinary least-squares estimator. The classical Gauss-Markov Theorem then reads as follows. Recall that a linear estimator is of the form $AY$, were $A$ is a (nonrandom) $k\times n$ matrix. Also recall that $\mathbf{F}_{2}^{0}=\mathbf{ F}_{2}(I_{n})$.

theoremIf $\hat{\beta}$ is a linear estimator that is unbiased under all $ F\in \mathbf{F}_{2}^{0}$ (meaning that $E_{F}\hat{\beta}=\beta (F)$ for every $F\in \mathbf{F}_{2}^{0}$), then \begin{equation*} Var_{F}(\hat{\beta})\succeq Var_{F}(\hat{\beta}_{OLS}) \end{equation*} for every $F\in \mathbf{F}_{2}^{0}$. [Here $\succeq $ denotes Loewner order.]

The theorem can equivalently be stated in the following more unusual form, which is the form chosen by Hansen (see Theorem 1 in Hansen_1.5, Hansen_2).\footnote{As formulated in Hansen_1.5, Hansen_2 , his\ Theorem 1 has $\sigma ^{2}(X^{\prime }X)^{-1}$ instead of $Var_{F}( \hat{\beta}_{OLS})$ on the r.h.s. of the inequality. Taken literally this leaves $\sigma ^{2}$ unspecified. To obtain a mathematically well-defined statement $\sigma ^{2}$ needs to be interpreted as $\sigma ^{2}(F)$, the variance of the data under $F$, the distribution under which the variance-covariance matrices of the estimators are computed.}

theoremIf $\hat{\beta}$ is a linear estimator that is unbiased under all $F\in \mathbf{F}_{2}$ (meaning that $E_{F}\hat{\beta}=\beta (F)$ for every $F\in \mathbf{F}_{2}$), then \begin{equation} Var_{F}(\hat{\beta})\succeq Var_{F}(\hat{\beta}_{OLS}) \end{equation} for every $F\in \mathbf{F}_{2}^{0}$.

In the latter theorem the unbiasedness is requested to hold over the larger class $\mathbf{F}_{2}$ of distributions rather than only over $ \mathbf{F}_{2}^{0}$. Of course, this is immaterial here and the two theorems are equivalent, because the estimators are required to be linear in both theorems and thus their expectations depend only on the first moment of $Y$ and not on the second moments at all.\footnote{For linear estimators $\hat{\beta}$ the condition $E_{F}\hat{\beta}=\beta (F)$ for every $F\in \mathbf{F}_{2}$ is, in fact also equivalent to $E_{F}\hat{ \beta}=\beta (F)$ for every $F\in \mathbf{G}$, whenever $\mathbf{G\subseteq } $ $\mathbf{F}_{2}$ holds and $\{\beta (F):F\in \mathbf{G}\}$ contains a basis of $\mathbb{R}^{k}$. This is obvious since any of these unbiasedness conditions are equivalent to $AX=I_{k}$, where $A$ is the matrix representing the linear estimator $\hat{\beta}$, i.e., $\hat{\beta}=AY$. [For $\mathbf{G=F}_{2}(I_{n})$ we obtain the equivalence noted above in the main text; a similar equivalence is obtained for $\mathbf{G=F}_{2}(\Sigma )$ .] Furthermore, if $\mathbf{G}$ is chosen to correspond to all distributions in $\mathbf{F}_{2}(I_{n})$ such that $e/\sigma $ follows a given distribution (\textquotedblleft parametric linear regression model\textquotedblright ), the before noted equivalence applies, and we thus can obtain a version of the Gauss-Markov Theorem for the parametric linear regression model. A similar remark applies to the Aitken Theorem.} While the difference in the unbiasedness conditions is immaterial in the preceding theorems, it is worth pointing out that the unbiasedness condition as given in Theorem (ref) requires that an estimator is not only unbiased in the underlying model with uncorrelated and homoskedastic errors one is studying, but also requires unbiasedness under correlated and/or heteroskedastic errors (i.e., under structures that are `outside' of the model that is being considered). Why one would want to impose such a requirement when the underlying model has uncorrelated and homoskedastic errors is at least debatable. However, we stress once more that in the context of the preceding two theorems this does not matter due to the assumed linearity of the estimators.

We next discuss what happens if one eliminates the linearity condition in the two equivalent theorems. Dropping the linearity conditions leads to the following assertions, which will turn out to be no longer equivalent to each other:

Assertion 1: If $\hat{\beta}$ is an estimator (i.e., a Borel-measurable function of $Y$) that is unbiased under all $F\in \mathbf{F} _{2}^{0}$ (meaning that $E_{F}\hat{\beta}=\beta (F)$ for every $F\in \mathbf{ F}_{2}^{0}$), then

equation[equation omitted — 89 chars of source]

for every $F\in \mathbf{F}_{2}^{0}$.

Assertion 2: If $\hat{\beta}$ is an estimator that is unbiased under all $F\in \mathbf{F}_{2}$ (meaning that $E_{F}\hat{\beta}=\beta (F)$ for every $F\in \mathbf{F}_{2}$), then

equation*[equation* omitted — 71 chars of source]

for every $F\in \mathbf{F}_{2}^{0}$.

Before discussing Assertions 1 and 2, we need to make a remark on the interpretation of inequalities like ((ref)).

remark(i) In Theorems (ref) and (ref) the objects $ Var_{F}(\hat{\beta})$ as well as $Var_{F}(\hat{\beta}_{OLS})$ are well-defined as real matrices because all estimators considered are linear, and hence $E_{F}(\parallel \hat{\beta}\parallel ^{2})<\infty $, $ E_{F}(\parallel \hat{\beta}_{OLS}\parallel ^{2})<\infty $ holds for every $ F\in \mathbf{F}_{2}^{0}$ where $\parallel .\parallel $ denotes the Euclidean norm. In contrast, in Assertions 1 and 2 estimators $\hat{\beta}$ with $ E_{F}(\parallel \hat{\beta}\parallel ^{2})=\infty $ for some $F\in \mathbf{F} _{2}^{0}$ are permissible. [Note that $E_{F}(\parallel \hat{\beta}\parallel ^{2})=\infty $ for some $F\in \mathbf{F}_{2}^{0}$ and $E_{F}(\parallel \hat{ \beta}\parallel ^{2})<\infty $ for some other $F\in \mathbf{F}_{2}^{0}$ may occur.] This necessitates some discussion how Assertions 1 and 2 are then to be read. In the first version of this paper we unfortunately had glossed over this issue, but an explicit discussion is warranted. For the subsequent discussion note that in both assertions $E_{F}(\hat{\beta})$ is well-defined and finite for every $F\in \mathbf{F}_{2}^{0}$ as a consequence of the respective unbiasedness assumption (and because $\mathbf{F}_{2}^{0}\subseteq \mathbf{F}_{2}$). (ii) In the scalar case (i.e., $k=1$), there is no problem as the object $Var_{F}(\hat{\beta})$ is well-defined for every $F\in \mathbf{F} _{2}^{0}$ as an element of the extended real line, regardless of whether $E_{F}(\parallel \hat{\beta}\parallel ^{2})<\infty $ or not. Hence, inequality ((ref)) always makes sense in case $k=1$. (iii) For general $k$, in case the estimator $\hat{\beta}$ satisfies $ E_{F}(\parallel \hat{\beta}\parallel ^{2})<\infty $ for a given $F\in \mathbf{F}_{2}^{0}$, the object $Var_{F}(\hat{\beta})$ is well-defined as a real matrix. Note that the inequality ((ref)) can then equivalently be expressed as $Var_{F}(c^{\prime }\hat{\beta})\geq Var_{F}(c^{\prime }\hat{\beta}_{OLS})$ for every $c\in \mathbb{R}^{k}$. (iv) In the case $k>1$, the object $Var_{F}(\hat{\beta})$ is not well-defined if $E_{F}(\parallel \hat{\beta}\parallel ^{2})=\infty $ ($F\in \mathbf{F}_{2}^{0}$), and hence it is not immediately clear how ((ref)) should then be understood. However, the inequalities $ Var_{F}(c^{\prime }\hat{\beta})\geq Var_{F}(c^{\prime }\hat{\beta}_{OLS})$ for every $c\in \mathbb{R}^{k}$ still make sense in view of (ii) above. We hence may and will interpret ((ref)) (with $F\in \mathbf{F} _{2}^{0} $) as a symbolic shorthand notation for $Var_{F}(c^{\prime } \hat{\beta})\geq Var_{F}(c^{\prime }\hat{\beta}_{OLS})$ for every $c\in \mathbb{R}^{k}$ (which works both in the case $E_{F}(\parallel \hat{\beta} \parallel ^{2})<\infty $ and in the case $E_{F}(\parallel \hat{\beta} \parallel ^{2})=\infty $). We have chosen to write inequality ((ref)) as given (abusing notation), rather than the more conventional and more precise $Var_{F}(c^{\prime }\hat{\beta})\geq Var_{F}(c^{\prime } \hat{\beta}_{OLS})$ for every $c\in \mathbb{R}^{k}$, in order for our discussion to be easily comparable with the presentation in Hansen's papers; his papers are silent on this issue. The same convention applies mutatis mutandis to similar statements such as, e.g., Assertions 3 and 4, etc. (v) The above discussion would become moot, if one would introduce the extra assumption $E_{F}(\parallel \hat{\beta}\parallel ^{2})<\infty $ for every $F\in \mathbf{F}_{2}^{0}$ into Assertions 1 and 2. However, such an additional assumption, which has little justification, would (potentially) narrow down the class of estimators competing with $\hat{\beta}_{OLS}$. As we shall see later on, such an extra assumption actually would have no effect on Assertion 2 (and thus on the corresponding theorems in Hansen's papers) at all in view of our Theorem (ref). The effect it would have on Assertion 1 (and some other results) is discussed in Appendix (ref).

We now turn to discussing Assertions 1 and 2. Not unexpectedly, Assertion 1 is incorrect in general.\footnote{ I.e., there exist design matrices $X$ such that the assertion is false.} This is known. For the benefit of the reader we provide some counterexamples and attending discussion in Appendix (ref). In particular, we see that in the classical Gauss-Markov Theorem as it is usually formulated (i.e., in Theorem (ref)) one can not eliminate the linearity condition in general!

Concerning Assertion 2, note that it coincides with Theorem 5 in Hansen_1.5, Hansen_2 (his `modern Gauss-Markov Theorem').\footnote{ The same caveat as expressed in Footnote (ref) also applies to the formulation of Theorem 5 in Hansen_1.5, Hansen_2.} Obvious questions now are (i) whether Assertion 2 (i.e., Theorem 5 in Hansen_1.5, Hansen_2) is correct, and (ii) if so, what is the reason for Assertion 2 to be correct while Assertion 1 is incorrect in general although in both assertions the linearity condition has been dropped. The answer to the latter question lies in the fact that Assertion 2 is requiring a stricter unbiasedness condition, namely unbiasedness over $\mathbf{F}_{2}$ rather than only unbiasedness over $\mathbf{F}_{2}^{0}$. While the two unbiasedness conditions effectively coincide for linear estimators as discussed before, this is no longer the case once we leave the realm of linear estimators. Hence, the (potential) correctness of Assertion 2 (i.e., of Theorem 5 in Hansen_1.5, Hansen_2) must crucially rest on imposing the stricter unbiasedness condition, which not only requires unbiasedness under the model considered (regression with homoskedastic and uncorrelated errors), but oddly also under structures `outside' of the maintained model (namely under heteroskedastic and/or correlated errors). Note that the class of competitors to $\hat{\beta}_{OLS}$ figuring in Assertion 1 is, in general, larger than the class of competitors appearing in Assertion 2. Nevertheless, Hansen_1.5, Hansen_2 (and also Hansen_3) are quiet on the use of this stricter unbiasedness condition.

Having understood what distinguishes Assertion 2 (i.e., Theorem 5 in Hansen_1.5, Hansen_2) from Assertion 1, the question remains whether the former is indeed correct, and if so, what its scope is, i.e., how much larger than the class of linear (unbiased) estimators the class of estimators covered by Assertion 2 (i.e., by Theorem 5 in Hansen_1.5, Hansen_2) is. We answer this now: As we shall show in the subsequent theorem, the only estimators $\hat{\beta}$ satisfying the unbiasedness condition of Assertion 2 (i.e., of Theorem 5 in Hansen_1.5, Hansen_2) are linear estimators! In other words, Theorem 5 in Hansen_1.5, Hansen_2 (i.e., his `modern Gauss-Markov Theorem') is nothing else than the good old(fashioned) Gauss-Markov Theorem (i.e., Theorem (ref) above), just stated in a somewhat unusual and intransparent way! \footnote{ Recall from before that for linear estimators the unbiasedness conditions in Theorems (ref) and (ref) are equivalent.} [While the word `linear' does not appear in the formulation of Theorem 5 in Hansen_1.5, Hansen_2, linearity of the estimators is introduced indirectly through a backdoor provided by the stricter unbiasedness condition.] While Theorem 5 in Hansen_1.5, Hansen_2 thus turns out to be correct, it is certainly not new!\footnote{ We have not checked whether the proofs in Hansen_1.5, Hansen_2 are correct or not.} Theorem 6 in Hansen_2 is a special case of his Theorem 5 for the location model, and thus is also not new; in contrast, Theorem 6 in Hansen_1.5 is a special case of Assertion 1. Example (ref) in Appendix (ref) shows that this theorem is false. What has been said so far also serves as a reminder that one has to be careful with statements such as \textquotedblleft best unbiased equals best linear unbiased\textquotedblright . While this statement is incorrect in the context of Assertion 1 in general, it is trivially correct in the context of Assertion 2 (i.e., of Theorem 5 in Hansen_1.5, Hansen_2) as a consequence of the subsequent Theorem (ref).

An upshot of the preceding discussion is that -- despite a plea to the contrary in Hansen_1.5, Hansen_2, Hansen_3 -- one should not drop `linearity' from the pedagogy of the Gauss-Markov Theorem. There is nothing to gain and a lot to lose: It will lead to an incorrect assertion, if one starts from the usual formulation of the classical Gauss-Markov Theorem (i.e., from Theorem (ref));\ otherwise (i.e., if one starts from Theorem (ref)), it will lead to a correct, but rather intransparent, assertion that is in fact equivalent to the classical Gauss-Markov Theorem. Unfortunately, Hansen has fallen victim to his own advice as the Gauss-Markov Theorem (Theorem 4.4) given in the August 2021 version of his forthcoming text-book HABook is incorrect in general (as it coincides with Assertion 1).

We now provide the theorem alluded to above. After the first version of this paper had been circulated, we learned about Portnoy, which establishes a related result using different arguments than the ones we use; for more discussion see Remark (ref) further below.

theoremIf $\hat{\beta}$ is an estimator (i.e., a Borel-measurable function of $Y$) that is unbiased under all $F\in \mathbf{F} _{2}$ (meaning that $E_{F}\hat{\beta}=\beta (F)$ for every $F\in \mathbf{F} _{2}$), then $\hat{\beta}$ is a linear estimator (i.e., $\hat{\beta}=AY$ for some $k\times n$ matrix $A$).\footnote{ By unbiasedness, such an $A$ must then also satisfy $AX=I_{k}$.}

We give a first "proof" based on Theorem 4.3 in koopmann (also reported as Theorem 2.1 in Gnotetal), but see the discussion immediately following this "proof" for a caveat.\footnote{ Curiously, the result by koopmann in question is actually mentioned in Section 1 of Hansen_1.5, Hansen_2, Hansen_3.}

A first "proof": The unbiasedness assumption of the theorem obviously translates into

equation[equation omitted — 106 chars of source]

for every symmetric and positive definite $\Sigma $ of dimension $ n\times n$; specializing to the case $\Sigma =I_{n}$, we, in particular, obtain\footnote{ Instead of $I_{n}$ we could have chosen any other symmetric and positive definite $n\times n$ matrix $\Sigma _{0}$ instead.}

equation[equation omitted — 105 chars of source]

Condition ((ref)), together with Theorem 4.3 in koopmann (see also Theorem 2.1 in Gnotetal\footnote{ Note that $X^{-}$ in that reference runs through all possible $g$-inverses of $X$.}$^{\text{,}}$\footnote{Gnotetal assume $\sigma ^{2}>0$ whereas koopmann allows also $\sigma ^{2}=0$. However, both theorems are equivalent as unbiasedness under every $F\in \mathbf{F} _{2}(I_{n})$ also implies unbiasedness under the point distributions at $ X\beta $ (i.e., the distributions corresponding to $\sigma ^{2}=0$). This is easily seen by considering those distributions in $\mathbf{F}_{2}(I_{n})$ that correspond to $X\beta +e$ with the components of $e$ being independent identically distributed according to $\varepsilon _{m}(\delta _{-1}+\delta _{1})/2+(1-\varepsilon _{m})\delta _{0}$. Here $\varepsilon _{m}$, $ 0<\varepsilon _{m}<1$, converges to zero for $m\rightarrow \infty $ and $ \delta _{x}$ denotes point mass at $x\in \mathbb{R}$. A similar argument applies in the case of $\mathbf{F}_{2}(\Sigma )$.}), implies that $\hat{\beta }$ is of the form

equation[equation omitted — 114 chars of source]

where $A^{0}$ satisfies $A^{0}X=I_{k}$ and $H_{i}^{0}$ are matrices satisfying $\limfunc{tr}(H_{i}^{0})=0$ and $X^{\prime }H_{i}^{0}X=0$ for $ i=1,\ldots ,k$. It is easy to see that we may without loss of generality assume that the matrices $H_{i}^{0}$ are symmetric (otherwise replace $ H_{i}^{0}$ by $(H_{i}^{0}+H_{i}^{0\prime })/2$). Inserting ((ref)) into ((ref)) yields

equation*[equation* omitted — 169 chars of source]

and this has to hold for every symmetric and positive definite $ \Sigma $. Standard calculations involving the trace operator and division by $\sigma ^{2}$ now give

equation[equation omitted — 176 chars of source]

For every $j=1,\ldots ,n$, choose now a sequence of symmetric and positive definite matrices $\Sigma _{m}^{(j)}$ (each of dimension $n\times n)$ that converges to $e_{j}(n)e_{j}(n)^{\prime }$ as $m\rightarrow \infty $, where $ e_{j}(n)$ denotes the $j$-th standard basis vector in $\mathbb{R}^{n}$ (such sequences obviously exist). Plugging this sequence into ((ref)), letting $m$ go to infinity, and exploiting properties of the trace-operator, we obtain

equation*[equation* omitted — 151 chars of source]

In other words, all the diagonal elements of $H_{i}^{0}$ are zero for every $ i=1,\ldots ,k$. Next, for every $j,l=1,\ldots ,n$, $j\neq l$, choose a sequence of symmetric and positive definite matrices $\Sigma _{m}^{\{j,l\}}$ (each of dimension $n\times n)$ that converges to $ (e_{j}(n)+e_{l}(n))(e_{j}(n)+e_{l}(n))^{\prime }$ as $m\rightarrow \infty $ (such sequences obviously exist). Then exactly the same argument as before delivers

equation*[equation* omitted — 189 chars of source]

Recall that the matrices $H_{i}^{0}$ are symmetric. Together with the already established fact that the diagonal elements are all zero, we obtain that also all the off-diagonal elements in any of the matrices $H_{i}^{0}$ are zero; i.e., $H_{i}^{0}=0$ for every $i=1,\ldots ,k$. This completes the proof.\footnote{ A slightly different version of the first "proof" can be obtained as follows. Theorem 4.3 in koopmann (together with Footnote (ref) ) shows for every given (fixed) $\Sigma $ that any $\hat{\beta}$ satisfying ( (ref)) is of the form $AY+(Y^{\prime }H_{1}Y,\ldots ,Y^{\prime }H_{k}Y)^{\prime }$ where $AX=I_{k}$, the $H_{i}$'s satisfy $\limfunc{tr} (H_{i}\Sigma )=0$, and $X^{\prime }H_{i}X=0$ for $i=1,\ldots ,k$. Again it is easy to see that we may assume the matrices $H_{i}$ to be symmetric. Note that the matrices $A$ and $H_{i}$ flowing from Theorem 4.3 in koopmann in principle could depend on $\Sigma $. The following argument shows that this is, however, not the case (after symmetrization of the $H_{i}$'s) in the present situation: If $\hat{\beta}$ had two distinct linear-quadratic representations with symmetric $H_{i}$'s, then the difference of these two representations would be a vector of multivariate polynomials (at least one of which is nontrivial) that would have to vanish everywhere, which is impossible since the zero-set of a nontrivial multivariate polynomial is a Lebesgue null-set. Given now the independence (from $\Sigma $) of the matrices $H_{i}$, one can then exploit the before mentioned relations $ \limfunc{tr}(H_{i}\Sigma )=0$ in the same way as is done following ((ref)) in the main text.} $\blacksquare $

Theorem 4.3 in koopmann is proved by reducing it to Theorem 3.1 (via Theorems 3.2, 4.1, and 4.2) in the same reference. Unfortunately, a full proof of Theorem 3.1 is not provided in koopmann, only a very rough outline is given. Thus the status of Theorem 4.3 in koopmann is not entirely clear. For this reason we next give a direct proof of our Theorem (ref) which does not rely on any result in koopmann. \footnote{ Alternatively, one could try to provide a complete proof of the result in koopmann. We have not pursued this, but have chosen the route via a direct proof of our Theorem (ref).}

A direct proof: It suffices to establish $\hat{\beta}(y+z)=\hat{ \beta}(y)+\hat{\beta}(z)$ as well as $\hat{\beta}(cz)=c\hat{\beta}(z)$ for every $y$ and $z$ in $\mathbb{R}^{n}$ and every $c\in \mathbb{R}$. For every $m\in \mathbb{N}$ with $m\geq 2$, every $V=(v_{1},\ldots ,v_{m})\in \mathbb{R }^{n\times m}$ and $\alpha \in (0,1)^{m}$ such that $\sum_{i=1}^{m}\alpha _{i}=1$, define a probability measure (distribution) via

equation*[equation* omitted — 76 chars of source]

where $\delta _{z}$ denotes unit point mass at $z\in \mathbb{R}^{n}$. The expectation of $\mu _{V,\alpha }$ equals $V\alpha $, and its variance-covariance matrix equals $V\limfunc{diag}(\alpha )V^{\prime }-(V\alpha )(V\alpha )^{\prime }$. Denote the expectation operator w.r.t. $ \mu _{V,\alpha }$ by $E_{V,\alpha }$. Note that in case $V\alpha =0$ and $ \limfunc{rank}(V)=n$ the measure $\mu _{V,\alpha }$ has expectation zero and a positive definite variance-covariance matrix; thus, $\mu _{V,\alpha }$ corresponds to an $F\in \mathbf{F}_{2}$ which has $\beta (F)=0$. From the unbiasedness assumption imposed on $\hat{\beta}$ we obtain that

equation[equation omitted — 167 chars of source]

Step 1: Fix $z\in \mathbb{R}^{n}$ and define $\alpha ^{(1)}=2^{-1}(n^{-1},\ldots ,n^{-1})^{\prime }\in \mathbb{R}^{2n}$, $\alpha ^{(2)}=2^{-1}((n+1)^{-1},\ldots ,(n+1)^{-1})^{\prime }\in \mathbb{R} ^{2(n+1)} $, $V_{1}=(I_{n},-I_{n})$ and $V_{2}=(I_{n},-I_{n},z,-z)$. Clearly $V_{1}\alpha ^{(1)}=V_{2}\alpha ^{(2)}=0$ and $\limfunc{rank}(V_{1})= \limfunc{rank}(V_{2})=n$. Furthermore,

equation[equation omitted — 143 chars of source]

which implies

equation*[equation* omitted — 154 chars of source]

Applying ((ref)) to $E_{V_{2},\alpha ^{(2)}}(\hat{\beta})$ and $ E_{V_{1},\alpha ^{(1)}}(\hat{\beta})$ now yields $0=\hat{\beta}(z)+\hat{\beta }(-z)$, i.e., we have shown that

equation[equation omitted — 102 chars of source]

in particular $\hat{\beta}(0)=0$ follows.

Step 2: Let $y$ and $z$ be elements of $\mathbb{R}^{n}$. Define the matrix

equation*[equation* omitted — 78 chars of source]

where $e_{i}(n)$ denotes the $i$-th standard basis vector in $\mathbb{R}^{n}$ , and set

equation*[equation* omitted — 134 chars of source]

Then, we obtain $V\alpha =0$ and $\limfunc{rank}(V)=n$. Using ((ref)) and ((ref)) it follows that

equation*[equation* omitted — 101 chars of source]

which by ((ref)) is equivalent to

equation[equation omitted — 113 chars of source]

Using ((ref)) with $y$ replaced by $y+z$ and $z$ replaced by $0$ yields

equation*[equation* omitted — 99 chars of source]

Since $\hat{\beta}(0)=0$ as shown before, we obtain

equation[equation omitted — 142 chars of source]

That is, we have shown that $\hat{\beta}$ is additive, i.e., is a group homomorphism between the additive groups $\mathbb{R}^{n}$ and $\mathbb{R} ^{k} $. By assumption it is also Borel-measurable. It then follows by a result due to Banach and Pettis (e.g., Theorem 2.2 in rosendal_2009) that $\hat{\beta}$ is also continuous. Homogeneity of $\hat{\beta}$ now follows from a standard argument, dating back to Cauchy, so that $\hat{\beta} $ is in fact linear. We give the details for the convenience of the reader: Relation ((ref)) (which contains ((ref)) as a special case) implies $\hat{\beta}(lz)=l\hat{\beta}(z)$ for every integer $l$. Replacing $ z $ by $z/l $ ($l\neq 0$) in the latter relation gives $\hat{\beta}(z)/l= \hat{\beta}(z/l) $ for integer $l\neq 0$. It immediately follows that $\hat{ \beta}(pz/q)=(p/q)\hat{\beta}(z)$ for every pair of integers $p$ and $q$ ($ q\neq 0$). Let $c\in \mathbb{R}$ be arbitrary. Choose a sequence of rational numbers $c_{s}$ that converges to $c$. Then by continuity of $\hat{\beta}$

equation*[equation* omitted — 221 chars of source]

This concludes the proof. $\blacksquare $

remarkInspection of the direct proof above shows that it does not make use of the full force of the unbiasedness condition ($E_{F}\hat{\beta} =\beta (F)$ for every $F\in \mathbf{F}_{2}$), but only exploits unbiasedness for certain strategically chosen discrete distributions $F$, each with finite support and satisfying $\beta (F)=0$.
remark(i) Portnoy uses a somewhat weaker unbiasedness condition than the one used in our Theorem (ref) (but see Remark (ref)), and then establishes only Lebesgue almost everywhere linearity of the estimators rather than linearity. This is an important distinction for the following reason: The results in Hansen_1.5, Hansen_2, Hansen_3 allow also for discrete distributions. For such distributions positive probability mass can fall into the exceptional Lebesgue null set, showing that any attempt to enforce linearity by appropriately redefining the estimator on the exceptional null set will in general not preserve the statistical properties of the estimator. In particular, the claim in Comment (a) in Section 3 of Portnoy that his result \textquotedblleft implies Hansen's result\textquotedblright\ is not warranted. Furthermore, at several instances in the discussion in Portnoy linearity is incorrectly claimed although only linearity Lebesgue almost everywhere is actually established in his paper. (ii) Portnoy emphasizes in his introduction as well as in Comment (b) in his Section 3 that his result allows for distributions that have no finite second moment. The following comment seems to be in order: The proof in Portnoy relies on requiring unbiasedness over a certain class $ \mathbf{P}$, say, of distributions which have compact support, and thus have finite moments of all orders. Trivially, then Portnoy's result holds a fortiori if one requires unbiasedness to hold over a larger class $\mathbf{P} ^{\ast }\mathbf{\supseteq P}$ of distributions, where $\mathbf{P}^{\ast }$ may contain also distributions that only have a finite first moment, but no finite second moment. The direct proof of our Theorem 3.4 effectively relies only on unbiasedness over a family of discrete distributions, each having finite support (cf. Remark (ref)). Again, then our linearity result trivially holds a fortiori if unbiasedness is required over any class of distributions containing the before mentioned family of discrete distributions. Of course, such a class may then also contain distributions that only have a finite first, but no finite second moment.
remark(Ambiguity in the definition in Hansen_1.5, Hansen_2 continued) If Hansen's $\mathbf{F}_{2}$ would be interpreted as coinciding with our $\mathbf{F}_{2}(\Sigma )$ (here with $\Sigma =I_{n}$ because of ((ref))) then the formulations of Theorems (ref) and (ref) as well as the formulations of Assertions 1 and 2 would coincide. In particular, with such an interpretation of Hansen's $\mathbf{F}_{2}$ his Theorem 5 in Hansen_1.5, Hansen_2 would be false.

The Aitken Case

In this section we drop the assumption ((ref)), i.e., $\Sigma $ in ((ref)) need not be the identity matrix. We make a preparatory remark: Similarly to observations made in Section (ref) (see Footnote (ref)), the rendition of Aitken's Theorem (for linear estimators) as given in Theorem 3 in Hansen_1.5, Hansen_2 needs some interpretation to convert it into a mathematically well-defined statement: The product $\sigma ^{2}\Sigma $, on which the r.h.s. of the inequality in that theorem depends (note that $\sigma ^{2}$ and $\Sigma $ enter the expression only via the product), is unspecified, and needs to be interpreted as $(\sigma ^{2}\Sigma )(F)$, the variance-covariance matrix of the data under the relevant $F$ w.r.t. which the variance-covariances in this inequality are taken. The same comment applies to Theorem 4 in Hansen_1.5, Hansen_2.

Aitken's Theorem as usually given in the literature reads as follows. Let $ \hat{\beta}_{GLS}=\hat{\beta}_{GLS}(\Sigma )=(X^{\prime }\Sigma ^{-1}X)^{-1}X^{\prime }\Sigma ^{-1}Y$ denote the generalized least-squares estimator using the known matrix $\Sigma $. Linear estimators are of the form $\hat{\beta}=AY$ where $A$ is a (nonrandom) $k\times n$ matrix.

theoremLet $\Sigma $ be an arbitrary known symmetric and positive definite $n\times n$ matrix. If $\hat{\beta}$ is a linear estimator that is unbiased under all $F\in \mathbf{F}_{2}(\Sigma )$ (meaning that $E_{F}\hat{ \beta}=\beta (F)$ for every $F\in \mathbf{F}_{2}(\Sigma )$), then \begin{equation*} Var_{F}(\hat{\beta})\succeq Var_{F}(\hat{\beta}_{GLS}) \end{equation*} for every $F\in \mathbf{F}_{2}(\Sigma )$.

Similar as in Section (ref), due to linearity of the estimators, an equivalent version of the theorem is obtained if the unbiasedness requirement is extended to all of $\mathbf{F}_{2}$.\footnote{ Cf. Footnote (ref).} This is precisely what happens in Theorem 3 in Hansen_1.5, Hansen_2, his rendition of the Aitken Theorem (for linear estimators). Note that the subsequent theorem is obviously equivalent to Theorem 3 in Hansen_1.5, Hansen_2 and perhaps is more transparent. [To see the equivalence, note that the all-quantor over $\Sigma $ in Theorem (ref) can be "absorbed" by replacing $\mathbf{F}_{2}(\Sigma )$ in that theorem with $\mathbf{F}_{2}$, provided the quantity $\sigma ^{2}\Sigma $ appearing in the expression $Var_{F}(\hat{\beta}_{GLS})=\sigma ^{2}(X^{\prime }\Sigma ^{-1}X)^{-1}=(X^{\prime }(\sigma ^{2}\Sigma )^{-1}X)^{-1}$ in ((ref)) below is understood as $(\sigma ^{2}\Sigma )(F)$ , as is necessary anyway for Theorem 3 in Hansen_1.5, Hansen_2 to formally make sense as noted earlier.]

theoremLet $\Sigma $ be an arbitrary known symmetric and positive definite $n\times n$ matrix. If $\hat{\beta}$ is a linear estimator that is unbiased under all $F\in \mathbf{F}_{2}$ (meaning that $E_{F}\hat{ \beta}=\beta (F)$ for every $F\in \mathbf{F}_{2}$), then \begin{equation} Var_{F}(\hat{\beta})\succeq Var_{F}(\hat{\beta}_{GLS}) \end{equation} for every $F\in \mathbf{F}_{2}(\Sigma )$.

Dropping linearity in both theorems now leads to two assertions.

Assertion 3: Let $\Sigma $ be an arbitrary known symmetric and positive definite $n\times n$ matrix. If $\hat{\beta}$ is an estimator that is unbiased under all $F\in \mathbf{F}_{2}(\Sigma )$ (meaning that $E_{F} \hat{\beta}=\beta (F)$ for every $F\in \mathbf{F}_{2}(\Sigma )$), then

equation*[equation* omitted — 71 chars of source]

for every $F\in \mathbf{F}_{2}(\Sigma )$.

Assertion 4: Let $\Sigma $ be an arbitrary known symmetric and positive definite $n\times n$ matrix. If $\hat{\beta}$ is an estimator that is unbiased under all $F\in \mathbf{F}_{2}$ (meaning that $E_{F}\hat{\beta} =\beta (F)$ for every $F\in \mathbf{F}_{2}$), then

equation*[equation* omitted — 71 chars of source]

for every $F\in \mathbf{F}_{2}(\Sigma )$.

Assertion 3 is again incorrect in general for reasons similar to the ones given for Assertion 1 in the previous section, cf. Appendix (ref). Assertion 4 is equivalent to Theorem 4 in Hansen_1.5, Hansen_2 (to which we shall refer as his `modern Aitken Theorem'); this is seen in the same way as the equivalence of Theorem (ref) above with Theorem 3 in Hansen_1.5, Hansen_2. Assertion 4 is indeed correct, but again not new, as the class of estimators figuring in Assertion 4 consists only of linear estimators as a consequence of Theorem (ref) above.\footnote{ Adding the extra condition $E_{F}(\parallel \hat{\beta}\parallel ^{2})<\infty $ for every $F\in \mathbf{F}_{2}(\Sigma )$ would have no effect on Assertion 4 in view of our Theorem (ref). The effect this extra condition would have on Assertion 3 is discussed in Appendix (ref).} Furthermore, a comment like Remark (ref) also applies here. We conclude this section by noting that the rendition of Aitken's Theorem in the text-book HABook (Theorem 4.5) is ambiguously formulated, making it difficult to decide whether it coincides with the (incorrect) Assertion 3 or with Assertion 4, which is (trivially) correct.

The Results in Hansen_3

In Hansen_3 the same model given by ((ref)), ((ref)), and ((ref)) as in Hansen_1.5, Hansen_2 is considered and $\mathbf{ F}_{2}$ is defined in the same manner.\footnote{ The assumption in Hansen_1.5, Hansen_2 that $\Sigma $ is known and positive definite and that $\sigma ^{2}$ is positive has been dropped in Hansen_3. Nevertheless positive definiteness of $\Sigma $ as well as $ \sigma ^{2}>0$ are frequently used in Hansen_3 (e.g., inverses of $ \Sigma $ are taken; the proof of Theorem 4 makes use of both properties). We hence will continue to assume $\sigma ^{2}>0$ and positive definiteness of $ \Sigma $ in our discussion. We furthermore note that the ambiguity in the definition of $\mathbf{F}_{2}$ in Hansen_1.5, Hansen_2 is now being avoided in Hansen_3 as $\Sigma $ is no longer assumed to be known. Of course, then $\sigma ^{2}$ and $\Sigma $ are no longer identifiable.} A set $ \mathbf{F}_{2}^{\ast }$ representing the subset of $\mathbf{F}_{2}$ corresponding to independent errors $e_{1},\ldots ,e_{n}$ is also defined; here $e_{i}$ denotes the $i$-th component of the error vector $e$. Furthermore, the subset of $\mathbf{F}_{2}^{\ast }$ corresponding to independent homoskedastic errors is denoted by $\mathbf{F}_{2}^{0}$ in Hansen_3. It should be noted that this set is not the same as the set $\mathbf{F}_{2}^{0}$ in Hansen_1.5, Hansen_2. To avoid any confusion we shall in the following write $\mathbf{F}_{2}^{0,new}$ for the set denoted by $\mathbf{F}_{2}^{0}$ in Hansen_3.

We start with a discussion of the treatment of Aitken's Theorem in Hansen_3. Hansen first gives a rendition of the classical Aitken Theorem (Theorem 3 in Hansen_3) which is identical to Theorem 3 in Hansen_1.5, Hansen_2, and thus to Theorem (ref) in the preceding section. He proceeds to provide his `modern Aitken Theorem' (Theorem 4 in Hansen_3), which is identical to the corresponding Theorem 4 in Hansen_1.5, Hansen_2 and which in turn is equivalent to Assertion 4 as just discussed in Section (ref) above.\footnote{ The same caveat regarding the formulation of Hansen's theorems as discussed in Section (ref) applies here.} Consequently, the discussion given in Section (ref) above applies. In particular, the estimators figuring in the `modern Aitken Theorem' in Hansen_3 are all automatically linear by our Theorem (ref), and hence the `modern Aitken Theorem' in Hansen_3 is not new, but reduces to the classical Aitken Theorem.

Hansen then goes on to provide a further result (Theorem 5 in Hansen_3 ) which can equivalently be stated as the following assertion (the equivalence is seen in the same way as the equivalence between Theorem 3 in Hansen_1.5, Hansen_2 and Theorem (ref) in Section (ref) above).\footnote{ The same caveat regarding the formulation of Hansen's theorems discussed in Section (ref) applies also to Theorem 5 in Hansen_3.} For $\Sigma $ a diagonal $n\times n$ matrix with positive diagonal elements, define $ \mathbf{F}_{2}^{\ast }(\Sigma )=\{F\in \mathbf{F}_{2}^{\ast }:Var_{F}(e)\propto \Sigma \}$, where $\propto $ denotes proportionality. Of course, then $\mathbf{F}_{2}^{\ast }=\tbigcup \{\mathbf{F}_{2}^{\ast }(\Sigma ):\Sigma $ diagonal with positive diagonal elements$\}$ and $ \mathbf{F}_{2}^{\ast }(I_{n})=\mathbf{F}_{2}^{0,new}$. Recall that $\hat{ \beta}_{GLS}=\hat{\beta}_{GLS}(\Sigma )=(X^{\prime }\Sigma ^{-1}X)^{-1}X^{\prime }\Sigma ^{-1}Y$.

Assertion 5: Let $\Sigma $ be an arbitrary known diagonal $ n\times n$ matrix with positive diagonal elements. If $\hat{\beta}$ is an estimator that is unbiased under all $F\in \mathbf{F}_{2}^{\ast }$ (meaning that $E_{F}\hat{\beta}=\beta (F)$ for every $F\in \mathbf{F}_{2}^{\ast }$), then

equation*[equation* omitted — 71 chars of source]

for every $F\in \mathbf{F}_{2}^{\ast }(\Sigma )$.

This assertion is very much different from an Aitken Theorem:\footnote{ Assertion 5 (equivalently, Theorem 5 in Hansen_3) seems to be correct. However, we have not checked the correctness of the proofs in Section 6 of Hansen_3 in any detail.} (i) The errors in the model corresponding to $F\in \mathbf{F}_{2}^{\ast }(\Sigma )$ (i.e., distributions $F$ for which the variance inequality has to hold) need to be independent, an assumption alien to Aitken's Theorem (even if $\Sigma $ is diagonal) as this theorem relies only on first and second moment assumptions (as opposed to an independence assumption). (ii) The unbiasedness assumption is -- like in results discussed earlier -- required to hold under the wider class of distributions $\mathbf{F}_{2}^{\ast }$, and not only under the distributions $F$ describing the data generating mechanism (i.e., $F\in \mathbf{F}_{2}^{\ast }(\Sigma )$).\footnote{Requiring unbiasedness over the wider class $\mathbf{F}_{2}^{\ast }$ is crucial here: If in Assertion 5 unbiasedness is only required to hold for $F\in \mathbf{F} _{2}^{\ast }(\Sigma )$ rather than for $F\in \mathbf{F}_{2}^{\ast }$, the resulting statement is incorrect in general. This follows for $\Sigma =I_{n}$ from Example (ref) in Appendix (ref). [Note that the estimator constructed in this example is unbiased even under every $F\in \mathbf{F} _{2}(I_{n})=\mathbf{F}_{2}^{0}$, and that the offending distribution constructed in this example belongs to $\mathbf{F}_{2}^{\ast }(I_{n})$, in fact even corresponds to independent identically distributed errors with finite second moments.] Another counterexample is provided by Example (ref) in Appendix (ref), which covers the location case. Similar examples can easily be constructed for any diagonal $\Sigma $ with positive diagonal elements by a transformation argument.} And (iii) an Aitken Theorem should allow for general $\Sigma $. While in the context of Assertions 2 and 4 no nonlinear unbiased estimator exists, in the context of Assertion 5 nonlinear unbiased estimators indeed exist (at least for some matrices $X$), cf. Remark (ref)\ below.

We next turn to the treatment of the Gauss-Markov Theorem in Hansen_3 : He starts with Theorem 1 (for linear estimators), which despite given the label Gauss-Markov, is not the Gauss-Markov Theorem, but a much weaker result relying on the unnecessarily restrictive assumption that the errors in the model are independent (and homoskedastic).\footnote{ A similar caveat as in Footnote (ref) also applies to Theorems 1, 6, and 7 in Hansen_3.} Such an independence assumption is superfluous in the classical Gauss-Markov Theorem. [Note that as long as only linear estimators are considered, requiring unbiasedness for all $F\in \mathbf{F}_{2}^{\ast }$ , as is done in Theorem 1 of Hansen_3, is identical to requiring unbiasedness for all $F\in \mathbf{F}_{2}^{0,new}$, or for all $F\in \mathbf{ F}_{2}^{0}=\mathbf{F}_{2}(I_{n})$ for that matter; cf. the discussion following Theorem (ref) in Section (ref).] The `modern Gauss-Markov Theorem' (Theorem 6 in Hansen_3) is now the special case of Assertion 5 for $\Sigma =I_{n}$; note that this result is just Theorem 1 in Hansen_3 with the linearity requirement dropped.\footnote{ The `modern Gauss-Markov Theorem' as given in Theorem 5 of \ Hansen_1.5, Hansen_2 is no longer presented, but of course is an immediate consequence of Theorem 4 in Hansen_3. Also Theorem 6 of Hansen_1.5, Hansen_2 is no longer given.} For reasons (i) and (ii) discussed in the preceding paragraph in connection with Assertion 5, this result can not legitimately be called a (modern) Gauss-Markov Theorem. Finally, Theorem 7 in Hansen_3 just specializes Theorem 6 in the same reference to the location model, hence the same remarks apply.

To sum up, Theorems 4-7 in Hansen_3 are either an intransparent restatement of the classical Aitken Theorem introducing linearity of the estimators through the backdoor (Theorem 4 in Hansen_3), or are results modelled on the Gauss-Markov or Aitken Theorem but employing substantial extra conditions such as independence assumptions, etc. (Theorems 5-7 in Hansen_3). [The significance and scope of the latter results is unclear for the reasons discussed before.] As a consequence, the advertisements regarding dropping of the linearity assumption made in Hansen_3 are by no means justified. In particular, the claim made in the abstract and repeated at the end of Section 3 of Hansen_3, that his theorems would show that the label "linear estimator" can be dropped from the pedagogy of the Gauss-Markov Theorem, is without any base. We thus repeat our warning against dropping the linearity assumption from the Gauss-Markov or Aitken Theorem.

remarkIn the discussion following Theorem 5 in Hansen_3, the author gives an example of a nonlinear estimator that is unbiased under all $F\in \mathbf{F}_{2}^{\ast }$. The object$\ \tilde{\beta}$ given there, however, is not well-defined as it is the sum of two components that are of different dimension (unless $k=1$). This can be rectified by redefining $ \tilde{\beta}$ as $\hat{\beta}_{OLS}+Y_{i}(Y_{j}-X_{j}^{\prime }\hat{\beta} _{-i})a$ ($i\neq j$) for any chosen $k\times 1$ vector $a\neq 0$ (here $ Y_{j} $ and $X_{j}^{\prime }$ denote the $j$-th row of $Y$ and $X$, respectively).\footnote{ There is an implicit assumption here, namely that the design matrix continues to have full column rank even after the $i$-th row is deleted.} This object is indeed unbiased under all $F\in \mathbf{F}_{2}^{\ast }$ (but, in general, not under all $F\in \mathbf{F}_{2}$). The claim in Hansen_3 that this is a nonlinear estimator, however, is not generally true for any design matrix $X$. For example, if $n=k+1$, then any leave-one-out residual is zero, and hence $\tilde{\beta}=\hat{\beta}_{OLS}$ is linear (for any choice of $i$ and $j$). Another example where $\tilde{ \beta}=\hat{\beta}_{OLS}$ is when $k=1$, the regressor is the first standard basis vector, $i\neq 1$, and $j=1$. Fortunately, there are examples of design matrices for which $\tilde{\beta}$ is indeed truly nonlinear.

Independent Identically Distributed Errors

We round-off the discussion by briefly considering in this section what happens if we add the condition

equation[equation omitted — 68 chars of source]

to the model. Let $\boldsymbol{F}_{2}^{iid}$ be the subset of $\boldsymbol{F} _{2}^{0}$ corresponding to distributions $F$ that result from ((ref)), ((ref)), ((ref)), and ((ref)). In particular, we ask what is the status of the following assertion which is analogous to Assertion 1.

Assertion 6: If $\hat{\beta}$ is an estimator that is unbiased under all $F\in \mathbf{F}_{2}^{iid}$ (meaning that $E_{F}\hat{\beta}=\beta (F)$ for every $F\in \mathbf{F}_{2}^{iid}$), then

equation*[equation* omitted — 71 chars of source]

for every $F\in \mathbf{F}_{2}^{iid}$.

Note that Assertion 6 differs from Assertion 1 in two respects: (i) the set of competitors to $\hat{\beta}_{OLS}$, i.e., the set of unbiased estimators in Assertion 6 is potentially larger than the corresponding set in Assertion 1, and (ii) the set of distributions $F$ for which the variance inequality has to hold has gotten smaller compared to Assertion 1. Hence, the truth-status of Assertion 1 does not inform us about the corresponding status of Assertion 6.

Fortunately, Example (ref) in Appendix (ref) comes to the rescue and shows that Assertion 6 is incorrect in general (meaning that a design matrix can be found such that it is false). This is so since the nonlinear estimator constructed in that example is a fortiori unbiased under $\mathbf{F }_{2}^{iid}$, and since the offending $F$ found in that example in fact belongs to $\mathbf{F}_{2}^{iid}$. However, in the special case of the location model Assertion 6 is actually true. This follows directly from Theorem 5 in Halmos.\footnote{Halmos allows for $\sigma ^{2}=0$ . However, this is immaterial as a consequence of the discussion in Footnote (ref).} [Recall that, in contrast, Assertion 1 is false in the case of a location model; cf. Example (ref) in Appendix (ref).\footnote{ It is perhaps interesting to note that the assertion one obtains from Assertion 6 by replacing $\mathbf{F}_{2}^{iid}$ by $\mathbf{F}_{2}^{\ast }(I_{n})$ at every occurrence in Assertion 6 is also incorrect in general, and even in the location case; see the discussion in Footnote (ref).} ]

For results in the location case pertaining to classes of absolutely continuous distributions (without or with symmetry restrictions) see Example 4.2 in Section 2.4 of LehCas and the discussion following this example.

A nice result is due to KS: Suppose we restrict to i.i.d. errors in our regression model, but where now the distribution of the errors, $G$ say, is known (and has finite second moments). Suppose also that $n\geq 2k+1$ and that the design matrix has no rows of zeroes. Then, if $\hat{\beta}_{OLS}$ is best unbiased in this model, the distribution $G$ must be Gaussian. [KS actually prove a more general result.] A related result for the location model with independent (not necessarily identically distributed) errors is given in Theorem 7.4.1 of KLR. For more results in that direction see Sections 7.4-7.9 in the same reference.

There is probably more in the mathematical statistics literature we are not aware of, but this is what a quick search has turned up.