Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
What Estimators Are Unbiased For Linear Models?
abstractThe recent thought-provoking paper by hansen2022modern proved that the Gauss-Markov theorem continues to hold without the requirement that competing estimators are linear in the vector of outcomes. Despite the elegant proof, it was shown by the authors and other researchers that the main result in the earlier version of Hansen's paper does not extend the classic Gauss-Markov theorem because no nonlinear unbiased estimator exists under his conditions. To address the issue, hansen2022modern added statements in the latest version with new conditions under which nonlinear unbiased estimators exist.
Motivated by the lively discussion, we study a fundamental problem: what estimators are unbiased for a given class of linear models? We first review a line of highly relevant work dating back to the 1960s, which, unfortunately, have not drawn enough attention. Then, we introduce notation that allows us to restate and unify results from earlier work and hansen2022modern. The new framework also allows us to highlight differences among previous conclusions. Lastly, we establish new representation theorems for unbiased estimators under different restrictions on the linear model, allowing the coefficients and covariance matrix to take only a finite number of values, the higher moments of the estimator and the dependent variable to exist, and the error distribution to be discrete, absolutely continuous, or dominated by another probability measure. Our results substantially generalize the claims of parallel commentaries on hansen2022modern and a remarkable result by koopmann1982parameterschatzung.
Introduction
The celebrated Gauss-Markov theorem plays a central role in econometrics theory for fixed-design linear models. It states that the ordinary least squares (OLS) estimator is the best linear unbiased estimator (BLUE) when the variance-covariance has a scalar form. The result was later generalized by aitkin1935least, who proved that the generalized least squares (GLS) estimator is BLUE under a linear model with the covariance matrix known up to a multiplicative constant. Notably, both efficiency results are proved within the class of linear estimators, raising the possibility that there might exist unbiased estimators nonlinear in the response variable that outperform the OLS/GLS estimators.
To the best of our knowledge, the problem of finding unbiased, nonlinear estimators with smaller variances was first studied by Theodore W. Anderson in the early 1960s anderson1962least, though the work seems to have received little attention. Anderson found that the OLS estimator does not have the minimal variance among all estimators that are unbiased under homoskedastic linear models with independent errors and an essentially unrestricted design matrix. In particular, he pointed out that linear-plus-quadratic (LPQ) estimators can outperform the OLS estimator for certain error distributions.
Very recently, hansen2022modern proved that, among all estimators that are unbiased under any linear model with a finite covariance matrix, the OLS estimator achieves the minimal variance under a scalar covariance matrix. In fact, Hansen proved the more general result that the GLS estimator achieves the minimal variance when the covariance matrix is proportional to the weight matrix (Theorem 4 therein). In this sense, the OLS/GLS estimator is the best unbiased estimator (BUE), a stronger property than BLUE. Hansen's proof relies on a clever combination of the distribution-tilting technique, which is often applied in semiparametric statistics to study asymptotic efficiency bickel1993efficient, and the Cram\'{e}r-Rao lower bound, which is used to study finite-sample efficiency (among unbiased estimators). These two results are not contradictory because anderson1962least examines a larger class of candidate estimators which are only required to be unbiased under a scalar covariance matrix while hansen2022modern restricts the candidates to be unbiased without constraints on the covariance matrix.
The broader model class considered by hansen2022modern only restricts the first moment of the dependent variable, casting doubt on existence of nonlinear unbiased estimators under such weak restrictions. In fact, since the acceptance of hansen2022modern, there has been a lively online discussion on this question led by the current authors\footnote{See e.g., \url{https://twitter.com/jmwooldridge/status/1492990971218440197?s=21} and
\url{https://twitter.com/lihua_lei_stat/status/1493291015129550849?s=21}} and investigated by other researchers potscher2022modern, portnoy2022linear. They unanimously conclude, with different proof strategies, that only linear estimators are unbiased under every linear model with finite covariance matrix. This, unfortunately, implies that BUE and BLUE are equivalent and hence Hansen's new results are restatements of the classical Gauss-Markov theorem.
In an attempt to address the issue, hansen2022modern added new results in the latest version (Theorems 5 and 6). The new results imply that, among all estimators that are unbiased under linear models with independent (rather than uncorrelated) errors, the OLS estimator is BUE under homoskedasticity and independence. This class of unbiased estimators surely includes nonlinear ones. Apparently, this is not the only way to enrich the set of unbiased estimators. In the 1980s, koopmann1982parameterschatzung studied the class of linear models with a covariance matrix known up to a multiplicative constant and proved that all unbiased estimators under this class are LPQ estimators with certain restrictions on the coefficients. Building upon this remarkable representation theorem, gnot1992nonlinear proved that the OLS/GLS estimator is no longer BUE among these LPQ estimators under the same class of linear models. Again, this result does not contradict hansen2022modern because they consider non-nested classes of unbiased estimators and different model classes for which the BUE is proved.
The differences among the above results are subtle --- any statement that the OLS/GLS estimator is or is not BUE hinges on (a) the set of competing estimators and (b) the class of linear models under examination. To address the nuances, we introduce a framework that allows restatement of all existing results anderson1962least, gnot1992nonlinear, hansen2022modern without ambiguity. This can help clear up the confusion that appears to be prevalent in the community and pave the way for further technical discussion.
After that, unlike the other commentaries, we will mainly focus on the characterization of unbiased estimators under different model classes, instead of conditions under which the OLS/GLS estimator is (not) BUE. Our main contributions are new representation theorems for unbiased estimators under various classes of linear models. Our results (Theorem (ref) and Theorem (ref)) substantially generalize potscher2022modern, portnoy2022linear, and koopmann1982parameterschatzung by allowing the coefficients and covariance matrix to take only a finite number of values, the higher moments of the estimator and the dependent variable to exist, and the error distribution to be discrete, absolutely continuous, or dominated by another probability measure. They are based on (a slightly stronger version of) a generic result in ruschendorf1987unbiased, which unifies and generalizes a line of work on unbiasedness and completeness halmos1946theory, fraser1954completeness, hoeffding1977more, fisher1982unbiased, koopmann1982parameterschatzung. Roughly speaking, it states that the class of unbiased estimators of zero under a class of distributions defined by a set of moment conditions are linear combinations of these moments. We believe the result can be applied to study other model classes and provide a user-friendly proof in Appendix (ref) for readers who are less familiar with functional analysis in abstract Banach spaces. Finally, we apply the representation theorem to detect a new case where the OLS/GLS estimator is BUE. It is analogous to Theorem 5 of hansen2022modern but closer to the spirit of classical Gauss-Markov theorem that only constrains the moments.
Notation and literature review without ambiguity
Families of linear models
We start by formally defining several families of linear models that have been investigated in the literature and will be studied in this paper. We require an intricate notation in order to address the nuances that have caused considerable confusion.
Let $X\in \mathbb{R}^{n\times k}$ denote the design matrix and $Y\in \mathbb{R}^{n}$ denote the vector of the dependent variable. Throughout the paper we treat $X$ as fixed and assume the measurable space of $Y$ is equipped with the standard Borel $\sigma$-algebra on $R^{n}$. For any $r\ge 1$ and coefficient vector $\beta\in \mathbb{R}^{k}$, we denote by $\mathbf{F}_r(X; \beta)$ the class of distributions of $Y\in \mathbb{R}^{n}$ with mean $X\beta$ and finite $r$-th moment,
equation[equation omitted — 137 chars of source]
where $\|\cdot\|$ denotes the Euclidean norm in $\mathbb{R}^{n}$. The second condition $\mathbb{E}_{F}[\|Y\|^r] < \infty$ becomes redundant when $r = 1$. The standard (fixed-design) linear model can be characterized as
equation[equation omitted — 109 chars of source]
often with $r = 2$ and sometimes with $r = 1$ jensen1979linear. It is clear that $\mathbf{F}_s(X)\subset \mathbf{F}_r(X)$ if $s \ge r$.
In the econometrics literature, the covariance matrix is often assumed to be known or estimable. We further define the stratum $\mathbf{F}_r(X; \beta, \Lambda)$ of $\mathbf{F}_r(X; \beta)$ by fixing the covariance matrix $\Lambda$, i.e.,
equation[equation omitted — 189 chars of source]
Note that $\mathbf{F}_r(X; \beta, \Lambda) = \mathbf{F}_2(X; \beta, \Lambda)$ for any $r \in [1, 2]$ and, for any $r\ge 2$,
\[\bigcup_{\Lambda\succeq \textbf{0}_{k\times k}}\mathbf{F}_r(X; \beta, \Lambda) = \mathbf{F}_r(X; \beta),\]
where $\textbf{0}_{k\times k}$ denotes the $k\times k$ zero matrix and $A \succeq B$ (resp. $\succ$) iff $A - B$ is positive semidefinite (resp. positive definite). Using this notation, the linear model considered in Aitken's theorem can be expressed as $\mathbf{F}_2^\Sigma(X)$ where
equation[equation omitted — 182 chars of source]
In particular, the sets $\mathbf{F}_2^0$ and $\mathbf{F}_2$ defined in hansen2022modern are equivalent to $\mathbf{F}_2^{I_{n}}(X)$, where $I_{n}$ denotes the $n\times n$ identity matrix, and $\bigcup_{\Sigma\succeq \textbf{0}_{k\times k}}\mathbf{F}_2^\Sigma(X) = \mathbf{F}_2(X)$ in our notation; see Remark 2.1 of potscher2022modern for a similar clarification.
All above classes only restrict the moments of $Y$. Sometimes one is willing to impose further independence assumptions on the error terms $Y - X\beta$. Specifically, let
equation[equation omitted — 207 chars of source]
Similarly, we can further restrict the errors to be identically distributed:
equation[equation omitted — 200 chars of source]
Analogous to (ref) and (ref), we can define
equation[equation omitted — 359 chars of source]
Note that $\mathbf{F}_{2, \mathrm{indep}}^{I_{n}}(X)$ (resp. $\mathbf{F}_{2, \mathrm{iid}}^{I_{n}}(X)$) gives all linear models with homoskedastic and independent errors (resp. with i.i.d. errors). All have been previously studied in the literature anderson1962least, gnot1992nonlinear, hansen2022modern and will be examined in later sections.
Sometimes we further restrict the above model classes by intersecting them with a dominating class:
equation[equation omitted — 119 chars of source]
or a subset of $\mathcal{P}(\mu)$ that only includes distributions with bounded Radon-Nikodym derivatives:
equation[equation omitted — 141 chars of source]
Note that the set of absolutely continuous distributions in $\mathbb{R}^{n}$, denoted by $\mathcal{P}_{\mathrm{cont}}$, is given by $\mathcal{P}(\mu)$ with $\mu$ chosen as the Lebesgue measure\footnote{$\mathcal{P}(\mu)$ would include mixed distributions if $\mu$ is mixed.}. Furthermore, as with potscher2022modern and portnoy2022linear, we also consider the class of discrete distributions\footnote{All results in the paper continue to hold if we redefine $\mathcal{P}_{\mathrm{disc}}$ to include discrete distributions with a countable support.}:
equation[equation omitted — 152 chars of source]
Unbiased estimators for fixed-design linear models
Unbiased estimators exist only if $\beta$ is identifiable. A sufficient and necessary condition, which will be assumed throughout, is
\[\mathrm{rank}(X) = k \le n,\]
in which case the coefficient $\beta$ can be identified as
\[\beta(F) = (X' X)^{-1}X'\mathbb{E}_{F}[Y].\]
where $'$ denotes the transpose. For notational convenience, we will simply write $\beta(F)$ as $\beta$, though it should be kept in mind that $\beta$ is a functional of the distribution of $Y$.
For any $r\ge 1$ and family of linear models $\mathbf{F}(X)$ with a full-rank $X$, denote by $\mathbf{U}_{r}(\mathbf{F}(X))$ as the class of estimators (i.e., measurable functions of $Y$) that are unbiased for $\beta$ with finite $r$-th moments under every distribution in $\mathbf{F}(X)$, i.e.,
equation[equation omitted — 206 chars of source]
where $\mathcal{L}^{r}(F)$ denotes the set of measurable functions that have finite $r$-th moments under $F$. The recent work hansen2022modern, potscher2022modern, portnoy2022linear studied the special case $r = 1$. Clearly, $\mathbf{U}_r(\cdot)$ satisfies the following monotonicity property, which will be repeatly invoked.
propositionFor any $r\ge 1$, $\mathbf{U}_{r}(\mathbf{F}(X))\subset \mathbf{U}_{r}(\tilde{\mathbf{F}}(X))$ if $\tilde{\mathbf{F}}(X) \subset \mathbf{F}(X)$.
In the following, we will consider the class of linear estimators
equation[equation omitted — 123 chars of source]
and the class of linear unbiased estimators (LUE):
equation[equation omitted — 137 chars of source]
It is well-known that all LUE estimators are unbiased.
propositionFor any $r\ge 1$, let $\mathbf{F}_r(X)$ be defined in (ref). Then,
\[\mathbf{H}_{\mathrm{LUE}}(X)\subset \mathbf{U}_r(\mathbf{F}_r(X)).\]
proofFor any $F\in \mathbf{F}(X)$ and $A\in \mathbb{R}^{n\times k}$ satisfying $A'X = I_{k}$,
\[\mathbb{E}_{F}[A'Y] = A'X \beta = \beta.\]
By Rosenthal's inequality rosenthal1970subspaces, \[F\in \mathbf{F}_{r}(X)\Longrightarrow \mathbb{E}_{F}[\|Y\|^r]<\infty \Longrightarrow \mathbb{E}_{F}[\|A'Y\|^r] < \infty\Longrightarrow A'Y \in \mathcal{L}^{r}(F).\]
Thus, $u(y) = A'y\in \mathbf{U}_r(\mathbf{F}_r(X))$.
Most recent discussions about hansen2022modern are essentially about when $\mathbf{U}_1(\mathbf{F}_1(X))\subset \mathbf{H}_{\mathrm{lin}}^{n, k}$. The following result shows that whenever it is true, the LUE estimators represent all unbiased estimators.
propositionFor any $r\ge 1$ and $\mathbf{F}(X)\subset \mathbf{F}_r(X)$ such that $\mathrm{span}\{\beta: F\in \mathbf{F}(X)\} = \mathbb{R}^{k}$,
\[\mathbf{U}_r(\mathbf{F}(X))\subset \mathbf{H}_{\mathrm{lin}}^{n, k}\Longrightarrow \mathbf{U}_r(\mathbf{F}(X)) = \mathbf{H}_{\mathrm{LUE}}(X).\]
proofFor any $A\in \mathbb{R}^{n\times k}$, $u(y) = A'y \in \mathbf{U}_r(\mathbf{F}(X))$ iff for any $F\in \mathbf{F}(X)$, $u\in \mathcal{L}^{r}(F)$ and
\[\mathbb{E}_{F}[A'Y] = A'X \beta = \beta.\]
By Rosenthal's inequality rosenthal1970subspaces, $A'Y \in \mathcal{L}^{r}(F)$ for any $A$. Since $\{\beta: F\in \mathbf{F}(X)\}$ spans $\mathbb{R}^{k}$, we must have $I_{k} - A'X = \textbf{0}_{k\times k}$ and hence $u\in \mathbf{H}_{\mathrm{LUE}}(X)$. The proof is then completed by Proposition (ref).
The main result of potscher2022modern \footnote{The authors posted a proof on social media that coincides with the first proof by potscher2022modern two weeks before; see \url{https://twitter.com/lihua_lei_stat/status/1493291015129550849?s=21}.} can be paraphrased as follows.
theorem[potscher2022modern, Theorem 3.4; see also portnoy2022linear]
\[\mathbf{U}_1(\mathbf{F}_2(X)\cap \mathcal{P}_{\mathrm{disc}})\subset \mathbf{H}_{\mathrm{lin}}^{n, k}.\]
Together with Proposition (ref), Theorem (ref) implies that only LUE estimators can be unbiased under linear models without further constraints on the covariance matrix, even if the model only involves discrete distributions with finite second moments, i.e.,
\[\mathbf{U}_1(\mathbf{F}_2(X)\cap \mathcal{P}_{\mathrm{disc}}) = \mathbf{H}_{\mathrm{LUE}}(X);\]
see also the footnote 7 of potscher2022modern.
While only linear estimators can be unbiased under $\mathbf{F}_2(X)$, the conclusion fails for $\mathbf{F}_{2, \mathrm{iid}}^{I_{k}}(X)$.
proposition[hansen2022modern, remark above Theorem 5]
\[\mathbf{H}_{\mathrm{LUE}}(X)\subsetneq \mathbf{U}_1\left(\mathbf{F}_{2,\mathrm{iid}}^{I_{k}}(X)\right).\]
hansen2022modern provides a simple example that adds a mean-zero term, $Y_i$ multiplied by the leave-$i$-th-observation-out residual, onto the OLS estimator.
Theorem (ref) also fails if further assumptions are imposed on the covariance matrix. Consider the class of LPQ estimators:
equation[equation omitted — 221 chars of source]
where $\mathbb{R}_{\mathrm{sym}}^{n\times n}$ is the set of all symmetric matrices in $\mathbb{R}^{n\times n}$. This class was first studied by anderson1962least who showed that, in the model class $\mathbf{F}_{2, \mathrm{iid}}^{I_{k}}$, an LPQ estimator with appropriately chosen $B_j$'s can have a smaller variance for estimating a specific linear contrast of $\beta$ than the best LUE estimator, unless $n = k$ or the errors are Gaussian brudno1957dispersion, kagan1969.
Subsequently, koopmann1982parameterschatzung proved a remarkable representation theorem that all unbiased estimators under $\mathbf{F}_2^\Sigma(X)$ are given by a subset of $\mathbf{H}_{\mathrm{LPQ}}^{n, k}$: \footnote{koopmann1982parameterschatzung does not require the symmetry of $B_j$, though he commented in a remark that $B_j$ can be assumed to be symmetric without loss of generality because $y'B_j y = y'(B_j + B_j')y / 2$. }
align[align omitted — 318 chars of source]
theorem[koopmann1982parameterschatzung, Theorem 4.1]
For any given $\Sigma\succeq \textbf{0}_{k\times k}$,
\[\mathbf{U}_1(\mathbf{F}_2^\Sigma(X)) = \mathbf{H}_{\mathrm{Km}}^{\Sigma}(X),\]
where $\mathbf{F}_2^\Sigma(X)$ be defined in (ref).
koopmann1982parameterschatzung only presented a proof sketch for this result in his book (Section 3 and 4, Chapter I). It is relatively easy to prove that $\mathbf{H}_{\mathrm{Km}}^{\Sigma}(X)\subset \mathbf{U}_1(\mathbf{F}_2^\Sigma(X))$ via simple linear algebra. The other direction is much more challenging and we will generalize Theorem (ref) in Section (ref) with a self-contained proof.
proof[Proof of $\mathbf{H}_{\mathrm{Km}}^{\Sigma}(X)\subset \mathbf{U}_1(\mathbf{F}_2^\Sigma(X))$:]
For any $u\in \mathbf{H}_{\mathrm{Km}}^{\Sigma}(X)$ with parameters $A, B_1, \ldots, B_k$ and $F\in \mathbf{U}_1(\mathbf{F}_2^\Sigma(X))$, $u\in \mathcal{L}^2(F)$ and
\[\mathbb{E}_{F}[u(Y)] = \mathbb{E}_{F}[A'Y] + (\mathbb{E}_{F}[Y'B_1 Y], \ldots, \mathbb{E}_{F}[Y'B_k Y]).\]
Let $\epsilon = Y - X\beta$. By definition of $\mathbf{U}_1(\mathbf{F}_2^\Sigma(X))$, we have $\mathbb{E}_{F}[\epsilon] = \textbf{0}_{k}$, where $\textbf{0}_{k}$ denotes the $k$-dimensional zero vector, and $\mathbb{E}_{F}[\epsilon\epsilon'] = \sigma^2\Sigma$ for some $\sigma^2 > 0$. Then,
\[\mathbb{E}_{F}[A'Y] = A'X\beta = \beta.\]
For each $j$,
\begin{align*}
\mathbb{E}_{F}[Y'B_j Y]
&= \mathbb{E}_{F}[\beta'X'B_j X\beta] + 2\mathbb{E}_{F}[\beta' XB_j \epsilon] + \mathbb{E}_{F}[\epsilon' B_j \epsilon]\\
& \stackrel{(i)}{=}\mathbb{E}_{F}[\epsilon' B_j \epsilon] = \mathbb{E}_{F}[\mathrm{tr}(\epsilon'B_j \epsilon)]\\
& \stackrel{(ii)}{=} \mathbb{E}_{F}[\mathrm{tr}(B_j \epsilon\epsilon')] = \sigma^2\mathbb{E}_{F}[\mathrm{tr}(B_j\Sigma)] \\
& \stackrel{(iii)}{=} 0,
\end{align*}
where (i) invokes the restriction that $X'B_j X = \textbf{0}_{k\times k}$ and $\mathbb{E}_{F}[\epsilon] = \textbf{0}_{k}$, (ii) uses the fact that $\mathrm{tr}(AB) = \mathrm{tr}(BA)$ and $\mathbb{E}_{F}[\epsilon\epsilon'] = \sigma^2\Sigma$, and (iii) applies the condition that $\mathbb{E}_{F}[\mathrm{tr}(B_j \Sigma)] = 0$.
Revisiting B(L)UE for fixed-design linear models
definitionWe say an estimator $u^{*}(Y)$ is BUE in a class of unbiased estimators $\mathbf{H}(X)$ with respect to a class of distributions $\mathbf{F}(X)\in \mathbf{F}_2(X)$ if
\[\mathrm{Var}_{F}[u(Y)]\succeq \mathrm{Var}_{F}[u^{*}(Y)], \quad \text{for any }u\in \mathbf{H}(X) \text{ and any }F\in \mathbf{F}(X).\]
When $\mathbf{H}(X) = \mathbf{H}_{\mathrm{LUE}}(X)$, we say $h$ is BLUE with respect to $\mathbf{F}(X)$.
Define the
the GLS estimator associated with a covariance matrix $\Sigma$,
equation[equation omitted — 111 chars of source]
where $A^{-}$ denotes the generalized inverse of $A$. Note that the OLS estimator $\hat{\beta}_{\mathrm{OLS}}(Y) = (X'X)^{-1}X'Y$ is equivalent to $\hat{\beta}_{\mathrm{GLS}}^{I_k}(Y)$.
The confusion often arises when $\mathbf{H}(X)$ and $\mathbf{F}(X)$ are not explicitly defined. We summarize the existing results on BUE and non-BUE of the OLS and GLS estimators separately in the following two theorems. By paraphrasing them on an equal footing, we hope to reveal the nuances.
theorem\begin{enumerate}[(a)]
• (Gauss-Markov theorem) $\hat{\beta}_{\mathrm{OLS}}(Y)$ is BLUE with respect to $\mathbf{F}_2(X)$.
• $\hat{\beta}_{\mathrm{OLS}}(Y)$ is BUE in $\mathbf{H}(X)$ with respect to $\mathbf{F}(X)$ if one of the following conditions hold:
\begin{enumerate}[(1)]
• (hansen2022modern, Theorem 2; implied by Cram\'{e}r-Rao theorem)
\[\mathbf{H}(X) = \mathbf{U}_1(\mathbf{F}(X)), \quad \mathbf{F}(X) = \mathbf{F}_2(X) \cap \{\mathrm{Normal}(X\beta, \sigma^2 I_{n})\text{ for some }\sigma^2 > 0\}.\]
• (anderson1962least, Proposition 2)
\[n = k = \mathrm{rank}(X), \quad \mathbf{H}(X) = \mathbf{U}_1(\mathbf{F}_{2,\mathrm{iid}}^{I_{n}}(X)), \quad \mathbf{F}(X) = \mathbf{F}_{2}(X).\]
In fact, anderson1962least proves this result by showing that $\mathbf{H}(X) = \{\hat{\beta}_{\mathrm{OLS}}(Y)\}$.
• (hansen2022modern, Theorem 6)
\[\mathbf{H}(X) = \mathbf{U}_1\left(\bigcup_{\Sigma \text{ diagonal}}\mathbf{F}_{2,\mathrm{indep}}^{\Sigma}(X)\right), \quad \mathbf{F}(X) = \mathbf{F}_{2,\mathrm{indep}}^{I_{n}}(X).\]
• (fraser1954completeness, together with Rao-Blackwell Theorem)
\[X = (1, 1, \ldots, 1)', \quad \mathbf{H}(X) = \mathbf{U}_1(\mathbf{F}_{2, \mathrm{iid}}^{I_{n}}(X)), \quad \mathbf{F}(X) = \mathbf{F}_{2, \mathrm{iid}}^{I_{n}}(X).\]
It is not implied by Theorem 7 of hansen2022modern, itself a special case of his Theorem 6.
\end{enumerate}
• (anderson1962least, Proposition 5) $\hat{\beta}_{\mathrm{OLS}}(Y)$ is not BUE in $\mathbf{H}(X)$ with respect to $\mathbf{F}(X)$ if
\[\min_{i = 1,\ldots,n}\mathrm{rank}(X_{-i}) = k < n, \quad \mathbf{H}(X) = \mathbf{U}_1(\mathbf{F}_{2,\mathrm{indep}}^{I_{n}}(X)), \quad \mathbf{F}(X) = \mathbf{F}_{2,\mathrm{indep}}^{I_{n}}(X),\]
where $X_{-i}$ is the design matrix with the $i$-th row removed.
\end{enumerate}
theorem\begin{enumerate}[(a)]
• (aitkin1935least) For a given $\Sigma \succ \textbf{0}_{k\times k}$, $\hat{\beta}_{\mathrm{GLS}}^{\Sigma}(Y)$ is BLUE with respect to $\mathbf{F}_2^{\Sigma}(X)$.
• For a given $\Sigma \succ \textbf{0}_{k\times k}$, $\hat{\beta}_{\mathrm{GLS}}^{\Sigma}(Y)$ is BUE in $\mathbf{H}(X)$ with respect to $\mathbf{F}(X)$ if any of the following conditions hold:
\begin{enumerate}[(1)]
• (hansen2022modern, Theorem 4)
\[\mathbf{H}(X) = \mathbf{U}_1(\mathbf{F}_2(X)), \quad \mathbf{F}(X) = \mathbf{F}_2^{\Sigma}(X).\]
However, potscher2022modern and portnoy2022linear showed that $\mathbf{U}_1(\mathbf{F}_2(X)) \subset \mathbf{H}_{\mathrm{lin}}^{n, k}$. Thus, this is alien to (a).
• (hansen2022modern, Theorem 5) $\Sigma$ is diagonal,
\[\mathbf{H}(X) = \mathbf{U}_1\left(\bigcup_{\Sigma \text{ diagonal}}\mathbf{F}_{2,\mathrm{indep}}^{\Sigma}(X)\right), \quad \mathbf{F}(X) = \mathbf{F}_{2, \mathrm{indep}}^{\Sigma}(X).\]
\end{enumerate}
• (gnot1992nonlinear) $\hat{\beta}_{\mathrm{GLS}}^{\Sigma}(Y)$ is not BUE in $\mathbf{H}(X)$ with respect to $\mathbf{F}(X)$ if
\[\mathbf{H}(X) = \mathbf{U}_1(\mathbf{F}_2^{\Sigma}(X)), \quad \mathbf{F}(X) = \mathbf{F}_2^{\Sigma}(X).\]
\end{enumerate}
Contrasting Theorem (ref) (b) (3) with Theorem (ref) (c) or Theorem (ref) (b) (1) with Theorem (ref) (c), we can see that the choice of $\mathbf{H}(X)$ is crucial --- while the OLS/GLS can be BUE in a smaller class of estimators, they may no longer be optimal in a larger class.
Unbiased estimators for fixed-design linear models
A master theorem
For any class of fixed-design linear models $\mathbf{F}(X)\subset \mathbf{F}_r(X)$, let
equation[equation omitted — 199 chars of source]
Since $\hat{\beta}_{\mathrm{OLS}}$ is linear, by Rosenthal's inequality rosenthal1970subspaces,
\[\mathbb{E}_{F}[\|Y\|^r] < \infty\Longrightarrow \hat{\beta}_{\mathrm{OLS}}\in \mathcal{L}^{r}(F).\]
Thus,
\[u\in \mathbf{U}_r(\mathbf{F}(X)) \Longleftrightarrow u - \hat{\beta}_{\mathrm{OLS}} \in \{(u_1, \ldots, u_{k})': u_j\in \mathbf{U}^0_r(F(X))\} \triangleq \mathbf{U}^0_{r}(\mathbf{F}(X))^{\otimes k},\]
where $\otimes k$ denotes the $k$-th order tensor product, and thus,
equation[equation omitted — 145 chars of source]
As a consequence, to find all unbiased estimators of $\beta$, it suffices to characterize all unbiased estimators of $0$.
It is not hard to see that both $\mathbf{F}_{r}(X; \beta)$ and $\mathbf{F}_{r}(X; \beta, \Lambda)$ can be formulated as
equation[equation omitted — 139 chars of source]
where
equation[equation omitted — 462 chars of source]
and $x_i'$ denotes the $i$-th row of $X$.
Intuitively, if we treat $F$ as an element in the “dual space” of real-valued measurable functions, though this is incorrect in the rigorous sense, $\mathbf{F}(X)$ is the orthogonal complement of $\mathbf{G}$ in the “dual space” with respect to the bilinear form $(u, F)\mapsto \mathbb{E}_{F}[u(Y)]$. Similarly, $\mathbf{U}^0_r(\mathbf{F}(X))$ can be regarded as the orthogonal complement of $\mathbf{F}(X)$. Then we should expect, in an unrigorous sense, that $\mathbf{U}^0_{r}(\mathbf{F}(X)) = (\mathbf{G}^{\perp})^{\perp}$, which is the closure of the span of $\mathbf{G}$ in an appropriately defined normed vector space. This intuitive argument, together with (ref), suggests that all unbiased estimators under $\mathbf{F}_r(X; \beta)$ are linear and those under $\mathbf{F}_r(X; \beta, \Lambda)$ are LPQ, hence implying the results by potscher2022modern, portnoy2022linear, and koopmann1982parameterschatzung.
We formalize the above heuristics in the following theorem. Below, for a given class of distributions $\mathbf{F}$, we write $\mathcal{L}^{r}(\mathbf{F})$ for $\bigcap_{F\in \mathbf{F}}\mathcal{L}^{r}(F)$. Moreover, for any subsets of functions $\mathcal{G}$ and $\mathcal{H}$, we say $\mathcal{G} \subset \mathcal{H}$ almost surely under $F$ if for each $g\in \mathcal{G}$ there exists $h\in \mathcal{H}$ such that $g - h = 0$ almost everywhere under $F$, and say $\mathcal{G} = \mathcal{H}$ if $\mathcal{G}\subset \mathcal{H}$ and $\mathcal{H}\subset \mathcal{G}$ almost surely under $F$.
theorem[slight generalization of Theorem 1 of ruschendorf1987unbiased]
Let $\mathbf{G}$ be a set of real-valued measurable functions on $\mathbb{R}^{n}$ (equipped with the Borel $\sigma$-field) and
\[\mathbf{F}_{r, \mathbf{G}} = \left\{F: \mathbb{E}_{F}[g(Y)] = 0, g(Y)\in \mathcal{L}^{r}(F), \,\, \text{for all }g\in \mathbf{G}\right\}.\]
Then for any $F\in \mathbf{F}_{r, \mathbf{G}}$,
\[\mathbf{U}^0_{r}(\mathbf{F}_{r, \mathbf{G}}\cap \mathcal{P}(F)) \subset \mathrm{cl}_{F}(\mathrm{span}(\mathbf{G}))\cap \mathcal{L}^{r}(\mathbf{F}_{r, \mathbf{G}}\cap \mathcal{P}(F)) \subset \mathbf{U}^0_{r}(\mathbf{F}_{r, \mathbf{G}}\cap \mathcal{P}_{B}(F))\text{ almost surely under }F,\]
where $\mathcal{P}(F)$ and $\mathcal{P}_{B}(F)$ are the dominating classes defined in (ref) and (ref), respectively, and $\mathrm{cl}_{F}(\cdot)$ denotes the closure in $\mathcal{L}^{1}(F)$.
When $\mathbf{G}$ only contains a finite number of functions, as in (ref), it is well-known that $\mathrm{span}(\mathbf{G})$ is closed (see Proposition (ref)). Then Theorem (ref) implies that
equation[equation omitted — 138 chars of source]
On the other hand, for any $F \in \mathbf{F}_{r, \mathbf{G}}$,
\[\mathbf{G}\subset \mathcal{L}^{r}(F)\Longrightarrow \mathrm{span}(\mathbf{G})\subset \mathcal{L}^{r}(F), \quad F\in \mathbf{F}_{r, \mathbf{G}}\Longrightarrow F\in \mathbf{F}_{r, \mathrm{span}(\mathbf{G})}.\]
As a consequence,
equation[equation omitted — 191 chars of source]
where the second step follows from Proposition (ref). (ref) and (ref) imply the following simplified version of Theorem (ref).
theoremIf $|\mathbf{G}| < \infty$, then for any $r\ge 1$ and $F\in \mathbf{F}_{r, \mathbf{G}}$,
\[\mathbf{U}^0_r(\mathbf{F}_{r, \mathbf{G}}\cap \mathcal{P}(F)) = \mathrm{span}(\mathbf{G}) \text{ almost surely under }F.\]
Theorem (ref) generalizes Theorem 1A and 2A of hoeffding1977more and Theorem 3.1 of koopmann1982parameterschatzung. It will be our key tool to study unbiased estimators for fixed-design linear models.
Unbiased estimators without second moment constraints
Much of the recent discussion on hansen2022modern is about $\mathbf{F}_{2}(X)$, which solely imposes restrictions on the first moment of $Y$. In this section, we will substantially generalize potscher2022modern and portnoy2022linear, showing that $\mathbf{U}_{r}(\mathbf{F}(X)) = \mathbf{H}_{\mathrm{GLS}}(X)$ even when $r > 1$ and $\mathbf{F}(X)$ is much smaller than $\mathbf{F}_{r}(X)$. We start with a surprising result that all unbiased estimators must be linear even if the domain of $\beta$ has only two distinct values with the same direction and the error distribution is restricted into a dominating class.
lemmaFor any $r\ge 1$, $\beta^{*}\in \mathbb{R}^{k}$ and $c_1 < c_2\in \mathbb{R}$, let
\[\mathbf{F}(X) = \bigcup_{\beta\in \{c_j\beta^{*}: j=1,2\}}\mathbf{F}_{r}(X; \beta).\]
Further, let $F$ be any probability measure in $\mathbf{F}(X)$ such that
\[\mathbf{F}_{r}(X; \beta)\cap \mathcal{P}(F) \neq \emptyset, \quad \text{for all }\beta\in \{c_j\beta^{*}: j=1,2\}.\]
Then
\[\mathbf{U}^0_{r}(\mathbf{F}(X)\cap \mathcal{P}(F))\subset \mathbf{H}_{\mathrm{lin}}^{n, 1}\,\, \text{almost surely under }F.\]
proofThroughout this proof, we will write $=$ and $\subset$ for equality and subsets almost surely under $F$. By Proposition (ref),
\[\mathbf{U}^0_{r}(\mathbf{F}(X)\cap \mathcal{P}(F))\subset \bigcap_{\beta\in \{c_j\beta^{*}: j=1,2\}}\mathbf{U}^0_{r}(\mathbf{F}_{r}(X; \beta)\cap \mathcal{P}(F)).\]
Without loss of generality, assume that $F\in \mathbf{F}_{r}(X; c_1\beta^{*})$. Taking any $u\in \mathbf{U}^0_{r}(\mathbf{F}(X)\cap \mathcal{P}(F))$, we must have $u\in \mathbf{U}^0_{r}(\mathbf{F}_{r}(X; c_1\beta^{*})\cap \mathcal{P}(F))$. By (ref) and Theorem (ref), there exists $a_{i}\in \mathbb{R}$ such that
\[u(y) =\sum_{i=1}^{n}a_{i}(y_{i} - c_1 x_{i}'\beta^{*}).\]
Since $u\in \mathbf{U}^0_{r}(\mathbf{F}_r(X; c_2\beta^{*})\cap \mathcal{P}(F))$ and $\mathbf{F}_{r}(X; c_2\beta^{*})\cap \mathcal{P}(F)\neq \emptyset$, for any $\tilde{F}\in \mathbf{F}_{r}(X; c_2\beta^{*})\cap \mathcal{P}(F)$,
\[0 = \mathbb{E}_{\tilde{F}}[u(Y)] = (c_{2} - c_{1})\sum_{i=1}^{n}a_{i}x_{i}'\beta^{*} \Longrightarrow \sum_{i=1}^{n}a_{i}x_{i}'\beta^{*} = 0.\]
Then $u(y)$ can be simplified into $u(y) = \sum_{i=1}^{n}a_{i}y_{i}$. Therefore,
\[\mathbf{U}^0_{r}(\mathbf{F}(X)\cap \mathcal{P}(F))\subset \mathbf{H}_{\mathrm{lin}}^{n, 1}.\]
With Lemma (ref), we can generalize Theorem (ref) by restricting the values of $\beta$ to a finite set that spans $\mathbb{R}^{k}$ and allowing the error distribution to be absolutely continuous with respect to the Lebesgue measure or any probability measure.
theoremSuppose $\mathcal{B}\subset \mathbb{R}^{k}$ includes at least $2k$ vectors
$\{c_{jm}\beta_{m}^{*}: j = 1, 2, m = 1, \ldots, k\}$ such that $c_{1m} < c_{2m}\in \mathbb{R}$ and
\[ \mathrm{span}(\mathcal{B}) = \mathbb{R}^{k}.\]
For any $r\ge 1$, let
\[\mathbf{F}(X) = \bigcup_{\beta\in \mathcal{B}}\mathbf{F}_{r}(X; \beta).\]
\begin{enumerate}[(a)]
• For any probability measure $F\in \mathbf{F}(X)$ such that $\mathbf{F}_r(X; \beta)\cap \mathcal{P}(F) \neq \emptyset$ for any $\beta\in \mathcal{B}$,
\[\mathbf{U}_r(\mathbf{F}(X)\cap \mathcal{P}(F)) = \mathbf{H}_{\mathrm{GLS}}(X) \,\,\text{almost surely under }F.\]
• Let $\mu$ be the Lebesgue measure on $\mathbb{R}^{k}$. Then
\[\mathbf{U}_r(\mathbf{F}(X)\cap \mathcal{P}_{\mathrm{cont}}) = \mathbf{H}_{\mathrm{GLS}}(X) \,\,\text{almost surely under }\mu.\]
• Let $\mathcal{P}_{\mathrm{disc}}$ be defined in (ref). Then
\[\mathbf{U}_r(\mathbf{F}(X)\cap \mathcal{P}_{\mathrm{disc}}) = \mathbf{H}_{\mathrm{GLS}}(X).\]
\end{enumerate}
proofBy Proposition (ref) and Proposition (ref), for any $r\ge 1$,
\[\mathbf{H}_{\mathrm{GLS}}(X)\subset \mathbf{U}_{r}(\mathbf{F}_r(X))\subset \mathbf{U}_{r}(\mathbf{F}(X)\cap \mathcal{P})\]
for any subset $\mathcal{P}$. Therefore, it remains to prove $\mathbf{U}_{r}(\mathbf{F}(X)\cap \mathcal{P})\subset \mathbf{H}_{\mathrm{GLS}}(X)$ with $\mathcal{B} = \{c_{jm}\beta_{m}^{*}: j=1,2, m=1,\ldots,k\}$ in each case.
\begin{enumerate}[(a)]
• Throughout this proof, we will write $=$ and $\subset$ for equality and subsets almost surely under $F$. Without loss of generality, assume that $F\in \mathbf{F}_{r}(X; c_{11}\beta_1^{*})$. By Proposition (ref) and Lemma (ref) with $\beta^{*} = \beta_1^{*}, (c_1, c_2) = (c_{11}, c_{21})$, we have that
\[\mathbf{U}^0_r(\mathbf{F}(X)\cap \mathcal{P}(F))\subset \mathbf{H}_{\mathrm{lin}}^{n, 1}.\]
Thus, for any $u\in \mathbf{U}^0_r(\mathbf{F}(X)\cap \mathcal{P}(F))$, $ u(y) = a'y$ for some $a\in \mathbb{R}^{n}$.
For any $j = 1,2$ and $m = 1, \ldots, k$, choose any distribution $\tilde{F}$ from $\mathbf{F}_r(X; c_{jm}\beta_{m}^{*})\cap \mathcal{P}(F)$, which is non-empty. Then
\[0 = \mathbb{E}_{\tilde{F}}[u(Y)] = a'X(c_{jm}\beta_{m}^{*}).\]
Since $\mathrm{span}(\mathcal{B}) = \mathbb{R}^{k}$, we must have $X'a = \textbf{0}_{k}$. Thus,
\[\mathbf{U}^0_r(\mathbf{F}(X)\cap \mathcal{P}(F))\subset \{u(y) = a'y: X'a = \textbf{0}_{k}\}.\]
By (ref),
\begin{align*}
\mathbf{U}_r(\mathbf{F}(X)\cap \mathcal{P}(F)) &= \left\{u(y) = \hat{\beta}_{\mathrm{OLS}}(y) + \tilde{A}'y: \tilde{A}'X = 0_{k\times k}\right\}\\
& = \left\{u(y) = ((X'X)^{-1}X' + \tilde{A}')y: \tilde{A}'X = 0_{k\times k}\right\}\\
& \subset \mathbf{H}_{\mathrm{GLS}}(X).
\end{align*}
• Let $F \sim N(X(c_{11}\beta_1^{*}), I_{k})$, the multivariate normal distribution with mean $X(c_{11}\beta_1^{*})$ and covariance matrix $I_{k}$. Clearly, $F\in \mathbf{F}(X)$ and, for any $\beta\in \mathcal{B}$, $j=1,2$, and $m = 1,\ldots, k$,
\[N(X(c_{jm}\beta_{m}^{*}), I_{k})\in \mathbf{F}_{r}(X; c_{jm}\beta_{m}^{*})\cap \mathcal{P}(F).\]
Then, the part (a) entails that
\[\mathbf{U}_r(\mathbf{F}(X)\cap \mathcal{P}(F))\subset \mathbf{H}_{\mathrm{GLS}}(X)\text{ almost surely under }F.\]
Since $F$ has positive density with respect to $\mu$ everywhere, $\mathcal{P}(F) = \mathcal{P}(\mu)$ and a function is almost surely zero under $F$ iff it is so under $\mu$. Therefore,
\[\mathbf{U}_r(\mathbf{F}(X)\cap \mathcal{P}_{\mathrm{cont}}) = \mathbf{U}_r(\mathbf{F}(X)\cap \mathcal{P}(\mu))\subset \mathbf{H}_{\mathrm{GLS}}(X)\text{ almost surely under }\mu.\]
• The proof is quite technical; see Appendix (ref).
\end{enumerate}
Unbiased estimators with second moment constraints
Next, we will consider the case where the covariance matrix is constrained. Perhaps surprisingly, even if $\mathrm{Cov}(Y)$ can only take two distinct values that differ by a multiplicative constant and $\beta$ can only take three distinct values pointing to the same direction, any unbiased estimator must be LPQ.
lemmaFor any $r\ge 2$, $\beta^{*}\in \mathbb{R}^{k}$, $c_1 < c_2 < c_3\in \mathbb{R}$, $\Sigma \succ \textbf{0}_{k\times k}$, and $0 < \sigma_{1}^2 < \sigma_{2}^2$, let
\[\mathbf{F}(X) = \bigcup_{\beta\in \{c_j\beta^{*}: j=1,2,3\}}\bigcup_{\Lambda \in \{\sigma_{j}^2 \Sigma: j = 1,2\}}\mathbf{F}_{r}(X; \beta, \Lambda).\]
Further, let $F$ be any probability measure in $\mathbf{F}(X)$ such that
\[\mathbf{F}_{r}(X; \beta, \Lambda)\cap \mathcal{P}(F) \neq \emptyset, \quad \text{for all }\beta\in \{c_j\beta^{*}: j=1,2,3\}, \,\Lambda\in \{\sigma_{j}^2 \Sigma: j = 1,2\}.\]
Then
\[\mathbf{U}^0_r(\mathbf{F}(X)\cap \mathcal{P}(F))\subset \mathbf{H}_{\mathrm{LPQ}}^{n, 1}\,\, \text{almost surely under }F.\]
proofThroughout this proof, we will write $=$ and $\subset$ for equality and subsets almost surely under $F$. By Proposition (ref),
\[\mathbf{U}^0_{r}(\mathbf{F}(X))\subset \bigcap_{\beta\in \{c_j\beta^{*}: j=1,2,3\}}\bigcap_{\Lambda \in \{\sigma_{j}^2 \Sigma: j = 1,2\}}\mathbf{U}^0_{r}(\mathbf{F}_{r}(X; \beta, \Lambda)).\]
Without loss of generality, assume that $F\in \mathbf{F}_{r}(X; c_1\beta^{*}, \sigma_1^2 \Sigma)$. Taking any $u\in \mathbf{U}^0_{r}(\mathbf{F}(X))$, we must have $u\in \mathbf{U}^0_{r}(\mathbf{F}_{r}(X; c_1\beta^{*}, \sigma_{1}^2 \Sigma))$. By (ref) and Theorem (ref), there exists $a_{i}, B_{i\ell}\in \mathbb{R}$ such that
\[u(y) =\sum_{i=1}^{n}a_{i}(y_{i} - c_1 x_{i}'\beta^{*}) + \sum_{i,\ell=1}^{n}B_{i\ell} \left[(y_{i} - c_1 x_{i}'\beta^{*})(y_{\ell} - c_1 x_{\ell}'\beta^{*}) - \sigma_{1}^2\Sigma_{i\ell}\right].\]
We can enforce $B_{i\ell} = B_{\ell i}$ by replacing both with $(B_{i\ell} + B_{\ell i})/2$. Since $u\in \mathbf{U}^0_{r}(\mathbf{F}(X))$, for any $\tilde{F} \in \mathbf{F}_{r}(X; c_1\beta^{*}, \sigma_{2}^2 \Sigma)\cap \mathcal{P}(F)$,
\[0 = \mathbb{E}_{\tilde{F}}[u(Y)] = (\sigma_{2}^2 - \sigma_{1}^2)\sum_{j,\ell=1}^{k} B_{j\ell}\Sigma_{j\ell}\Longrightarrow \sum_{j,\ell=1}^{k} B_{j\ell}\Sigma_{j\ell} = 0.\]
Thus, $u$ can be simplified into
\[u(y) = \sum_{i=1}^{n}a_{i}(y_{i} - c_1 x_{i}'\beta^{*}) + \sum_{i,\ell=1}^{n}B_{i\ell} (y_{i} - c_1 x_{i}'\beta^{*})(y_{\ell} - c_1 x_{\ell}'\beta^{*}).\]
Again, since $u\in \mathbf{U}^0_{r}(\mathbf{F}(X))$, for any $j=2,3$, $u\in \mathbf{U}^0_r(\mathbf{F}_{r}(X; c_j\beta^{*}, \sigma_{1}^2 \Sigma)\cap \mathcal{P}(F))$. Choose $F_j\in \mathbf{F}_{r}(X; c_j\beta^{*}, \sigma_{1}^2 \Sigma)\cap \mathcal{P}(F)$, which is non-empty. Then
\[0 = \mathbb{E}_{F_j}[u(Y)] = (c_{j} - c_{1})\sum_{i=1}^{n}a_{i}x_{i}'\beta^{*} + (c_{j} - c_{1})^2\sum_{i,\ell=1}^{k}B_{i\ell}(x_{i}'\beta^{*})(x_{\ell}'\beta^{*})\]
\[\Longrightarrow 0 = \sum_{i=1}^{n}a_{i}x_{i}'\beta^{*} + (c_{j} - c_{1})\sum_{i,\ell=1}^{n}B_{i\ell}(x_{i}'\beta^{*})(x_{\ell}'\beta^{*}).\]
Since $c_{2} - c_{1} \neq c_{3} - c_{1}$, we must have
\[\sum_{i=1}^{n}a_{i}x_{i}'\beta^{*} = \sum_{i,\ell=1}^{n}B_{i\ell}(x_{i}'\beta^{*})(x_{\ell}'\beta^{*}) = 0.\]
Then $u(y)$ can be further simplified into
\begin{align*}
u(y) = \sum_{i=1}^{n}a_{i}y_{i} + \sum_{i,\ell=1}^{n}B_{i\ell}y_{i}y_{\ell} - c_{1}\sum_{i,\ell=1}^{n}B_{i \ell} \left\{(x_i \beta^{*})y_{\ell} + (x_\ell'\beta^{*})y_{i}\right\}.
\end{align*}
Therefore,
\[\mathbf{U}^0_{r}(\mathbf{F}(X))\subset \mathbf{H}_{\mathrm{LPQ}}^{n, 1}.\]
We can then generalize Koopmann's representation theorem (Theorem (ref)). Throughout the rest of the paper, we will define the vectorization of any symmetric matrix $\Omega\in \mathbb{R}_{\mathrm{sym}}^{k\times k}$ as
equation[equation omitted — 312 chars of source]
By definition, for any $\Omega_1, \Omega_2\in \mathbb{R}_{\mathrm{sym}}^{k\times k}$, $\mathrm{tr}(\Omega_1\Omega_2) = \mathrm{vec}^{*}(\Omega_1)'\mathrm{vec}^{*}(\Omega_2)$.
theoremSuppose $\mathcal{B}\subset \mathbb{R}^{k}$ includes
$\{c_{jm}\beta_{m}^{*}: j = 1, 2, 3, m = 1, \ldots, M\}$ such that $c_{1m} < c_{2m} < c_{3m}\in \mathbb{R}$ and
\[\mathrm{span}(\{\beta_1^{*}, \ldots, \beta_{M}^{*}\}) = \mathbb{R}^{k}, \,\, \mathrm{span}(\{\beta_1^{*}\beta_1^{*'}, \ldots, \beta_{M}^{*}\beta_{M}^{*'}\}) = \mathbb{R}_{\mathrm{sym}}^{k\times k}.\]
For any $r\ge 1$, $\Sigma \succ \textbf{0}_{k\times k}$, and $0 < \sigma_{1}^2 < \sigma_{2}^2$, let
\[\mathbf{F}(X) = \bigcup_{\beta\in \mathcal{B}}\bigcup_{\sigma^2 \in \{\sigma_{1}^2, \sigma_{2}^2\}}\mathbf{F}_{2r}(X; \beta, \sigma^2\Sigma).\]
\begin{enumerate}[(a)]
• For any probability measure $F\in \mathbf{F}(X)$ such that $\mathbf{F}_r(X; \beta, \sigma^2\Sigma)\cap \mathcal{P}(F) \neq \emptyset$ for any $\beta\in \mathcal{B}$ and $\sigma^2\in \{\sigma_1^2, \sigma_2^2\}$,
\[\mathbf{U}_r(\mathbf{F}(X)\cap \mathcal{P}(F)) = \mathbf{H}_{\mathrm{Km}}^{\Sigma}(X) \,\,\text{almost surely under }F.\]
• Let $\mu$ be the Lebesgue measure on $\mathbb{R}^{k}$. Then
\[\mathbf{U}_r(\mathbf{F}(X)\cap \mathcal{P}_{\mathrm{cont}}) = \mathbf{H}_{\mathrm{Km}}^{\Sigma}(X) \,\,\text{almost surely under }\mu.\]
• Let $\mathcal{P}_{\mathrm{disc}}$ be defined in (ref). Then
\[\mathbf{U}_r(\mathbf{F}(X)\cap \mathcal{P}_{\mathrm{disc}}) = \mathbf{H}_{\mathrm{Km}}^{\Sigma}(X).\]
\end{enumerate}
proofAs shown in the partial proof of Theorem (ref), for any $r\ge 1$, by Proposition (ref),
\[\mathbf{H}_{\mathrm{Km}}^{\Sigma}(X)\subset \mathbf{U}_{1}(\mathbf{F}_{2}(X))\subset \mathbf{U}_{1}(\mathbf{F}_{2r}(X)).\]
By Rosenthal's inequality rosenthal1970subspaces, for any $u\in \mathbf{H}_{\mathrm{LPQ}}^{n, k}$ and $F\in \mathbf{F}_{2r}(X)$,
\[\mathbb{E}_{F}[\|Y\|^r] < \infty\Longrightarrow u\in \mathcal{L}^{r}(\mathbf{F}_{2r}(X)).\]
Thus,
\[\mathbf{H}_{\mathrm{Km}}^{\Sigma}(X)\subset \mathbf{U}_{1}(\mathbf{F}_{2r}(X))\cap \mathcal{L}^{r}(\mathbf{F}_{2r}(X)) = \mathbf{U}_{r}(\mathbf{F}_{2r}(X))\]
By Proposition (ref) again, for any subset $\mathcal{P}$,
\[\mathbf{H}_{\mathrm{Km}}^{\Sigma}(X)\subset \mathbf{U}_{r}(\mathbf{F}(X)\cap \mathcal{P}).\]
Therefore, it remains to prove $\mathbf{U}_{r}(\mathbf{F}(X)\cap \mathcal{P})\subset \mathbf{H}_{\mathrm{Km}}^{\Sigma}(X)$ with $\mathcal{B} = \{c_{jm}\beta_{m}^{*}: j=1,2,3, m=1,\ldots,k\}$ in each case.
\begin{enumerate}[(a)]
• Throughout this proof, we will write $=$ and $\subset$ for equality and subsets almost surely under $F$. Without loss of generality, assume that $F\in \mathbf{F}_{r}(X; c_{11}\beta_1^{*}, \sigma_{1}^2\Sigma)$. By Proposition (ref) and Lemma (ref) with $\beta^{*} = \beta_1^{*}, (c_1, c_2, c_3) = (c_{11}, c_{21}, c_{31})$, we have that
\[\mathbf{U}^0_r(\mathbf{F}(X)\cap \mathcal{P}(F))\subset \mathbf{H}_{\mathrm{LPQ}}^{n, 1}.\]
Then, for any $u\in \mathbf{U}^0_r(\mathbf{F}(X)\cap \mathcal{P}(F))$, there exists $a\in \mathbb{R}^{n}$ and $B\in \mathbb{R}_{\mathrm{sym}}^{n\times n}$ such that
\[u(y) = a'y + y'B y.\]
For any $\beta\in \mathcal{B}, \sigma^2\in \{\sigma_1^2, \sigma_2^2\}$, take $\tilde{F} \in \mathbf{F}_r(X; \beta, \Lambda)\cap \mathcal{P}(F)$, which is non-empty. Then
\begin{align*}
0 & = \mathbb{E}_{\tilde{F}}[u(Y)] = a'\mathbb{E}_{\tilde{F}}[Y] + \mathbb{E}_{\tilde{F}}[\mathrm{tr}(Y'B Y)] \\
& = a'\mathbb{E}_{\tilde{F}}[Y] + \beta'X'B X\beta + 2\mathbb{E}_{\tilde{F}}[(Y - X\beta)']B X\beta + \mathbb{E}_{\tilde{F}}[(Y - X\beta)'B(Y - X\beta)]\\
& = a'\mathbb{E}_{\tilde{F}}[Y] + \beta'X'B X\beta + \mathbb{E}_{\tilde{F}}[(Y - X\beta)'B(Y - X\beta)]\\
& = a'\mathbb{E}_{\tilde{F}}[Y] + \beta'X'B X\beta + \mathbb{E}_{\tilde{F}}[\mathrm{tr}(B(Y - X\beta)(Y - X\beta)')]\\
& = a'X\beta + \beta'X'B X\beta + \mathrm{tr}(B \Sigma)\sigma^2.
\end{align*}
Fixing $\beta$ and setting $\sigma^2$ to be $\sigma_{1}^2$ and $\sigma_{2}^{2}$, we obtain that
\begin{equation}
\mathrm{tr}(B\Sigma)\sigma_1^2 = \mathrm{tr}(B\Sigma)\sigma_2^2\Longrightarrow \mathrm{tr}(B\Sigma) = 0.
\end{equation}
Thus, for any $\beta\in \mathcal{B}$,
\[a'X\beta + \beta'X'B X\beta = 0.\]
For any $m = 1, \ldots, M$, setting $\beta = c_{jm}\beta_{m}^{*}$ yields that, for $c\in \{c_{1m}, c_{2m}, c_{3m}\}$,
\[(a'X\beta_{m}^{*})c + (\beta_{m}^{*'}X'BX \beta_{m}^{*})c^2 = 0.\]
This quadratic equation has three distinct solutions iff
\[a'X\beta_{m}^{*} = 0, \quad \beta_{m}^{*'}X'BX \beta_{m}^{*} = 0.\]
Since $\mathrm{span}(\{\beta_{1}^{*}, \ldots, \beta_{m}^{*}\}) = \mathbb{R}^{k}$, we must have
\begin{equation}
X'a = 0_{k}.
\end{equation}
Similarly,
\[\beta_{m}^{*'}X'BX \beta_{m}^{*} = 0 \Longleftrightarrow \mathrm{tr}\left[(X'BX) (\beta_{m}^{*}\beta_{m}^{*'})\right] = 0\Longleftrightarrow \mathrm{vec}^{*}(X'BX)'\mathrm{vec}^{*}(\beta_{m}^{*}\beta_{m}^{*'}) = 0,\]
Since $\mathrm{span}(\{\beta_{m}^{*}\beta_{m}^{*'}: m = 1,\ldots, M\}) = \mathbb{R}_{\mathrm{sym}}^{k\times k}$,
\[\mathrm{span}(\{\mathrm{vec}^{*}(\beta_{m}^{*}\beta_{m}^{*'}): m = 1,\ldots, M\}) = \mathbb{R}^{k(k-1)/2}.\]
Thus,
\begin{equation}
\mathrm{vec}^{*}(X'B X) = 0\Longrightarrow X'BX = 0_{k\times k}.
\end{equation}
Piecing (ref), (ref), and (ref) together, we conclude that
\[\mathbf{U}_r(\mathbf{F}(X)\cap \mathcal{P}(F)) = \hat{\beta}_{\mathrm{OLS}} + \mathbf{U}^0_{r}(\mathbf{F}(X))^{\otimes k} \subset \mathbf{H}_{\mathrm{Km}}^{\Sigma}(X).\]
• Same as the proof of Theorem (ref) (b) except that $F$ is chosen as $N(X(c_{11}\beta_1^{*}), \sigma_1^2 \Sigma)$.
• The proof is quite technical; see Appendix (ref).
\end{enumerate}
As mentioned earlier, koopmann1982parameterschatzung only presented a proof sketch for Theorem (ref), which is a special case of Theorem (ref) with $\mathbf{F}(X) = \mathbf{F}_{2}^\Sigma(X)$. Therefore, we confirm the remarkable result by koopmann1982parameterschatzung and hence address the concerns raised by potscher2022modern.
When $r\ge 1$ and no constraint is imposed on $\beta$, i.e., $\mathcal{B} = \mathbb{R}^{k}$, Theorem (ref) and Theorem (ref) show that the constraint on the covariance matrix matters because
\[\mathbf{H}_{\mathrm{GLS}}(X) = \mathbf{U}_r(\mathbf{F}_{2r}(X)) = \mathbf{U}_r\left( \bigcup_{\Sigma \succeq \textbf{0}_{k\times k}}\mathbf{F}_{2r}^\Sigma(X)\right), \quad \mathbf{H}_{\mathrm{Km}}^\Sigma(X) = \mathbf{U}_{r}\left(\mathbf{F}_{2r}^\Sigma(X)\right),\]
where $\mathbf{F}_{2r}^{\Sigma}(X)$ is defined in (ref). The following result shows that the quadratic estimators would be excluded even if the domain of $\beta$ and $\Sigma$ are both finite.
theoremLet $\mathcal{B}\subset \mathbb{R}^{k}$ satisfies the condition in Theorem (ref) and $\mathcal{M}\subset \mathbb{R}_{\mathrm{sym}}^{n\times n}$ includes at least $n(n-1)$ matrices $\{\sigma_{jm}^2\Sigma_m: j=1,2, m = 1,\ldots, n(n-1)/2\}$ such that $0 < \sigma_{1m}^2 < \sigma_{2m}^2$ and
\[\mathrm{span}\left( \mathcal{M}\right) = \mathbb{R}_{\mathrm{sym}}^{n\times n}.\]
Given $r\ge 1$, Let
\[\mathbf{F}(X) = \bigcup_{\beta\in \mathcal{B}}\bigcup_{\Lambda\in \mathcal{M}}\mathbf{F}_{2r}(X; \beta, \Lambda),\]
Then,
\[\mathbf{U}_{r}\left( \mathbf{F}(X) \cap \mathcal{P}_{\mathrm{disc}}\right) = \mathbf{H}_{\mathrm{GLS}}(X).\]
The same conclusion holds if $\mathcal{P}_{\mathrm{disc}}$ is replaced by $\mathcal{P}_{\mathrm{cont}}$ and $=$ denotes the almost sure equality.
proofLet $u\in \mathbf{U}_{r}(\mathbf{F}(X)\cap \mathcal{P}_{\mathrm{disc}})$. By Proposition (ref), for any $m$,
\[u \in \mathbf{U}_{r}(\mathbf{F}(X)\cap \mathcal{P}_{\mathrm{disc}}) \subset \mathbf{U}_{r}\left(\bigcup_{\beta\in \mathcal{B}}\bigcup_{\sigma^2\in \{\sigma_{1m}^2, \sigma_{2m}^2\}}\mathbf{F}_{2r}(X; \beta, \sigma^2\Sigma_{m})\cap \mathcal{P}_{\mathrm{disc}}\right).\]
Since $\mathcal{B}$ satisfies the condition of Theorem (ref), we have that
\[\mathbf{U}\left(\bigcup_{\beta\in \mathcal{B}}\bigcup_{\sigma^2\in \{\sigma_{1m}^2, \sigma_{2m}^2\}}\mathbf{F}_{2r}(X; \beta, \sigma^2\Sigma_{m})\cap \mathcal{P}_{\mathrm{disc}}\right) = \mathbf{H}_{\mathrm{Km}}^{\Sigma_{m}}(X).\]
Then, there exists $A\in \mathbb{R}^{n\times k}$ and $B_1, \ldots, B_k \in \mathbb{R}_{\mathrm{sym}}^{n\times n}$ such that
\[u(y) = A'y + (y'B_1 y, \ldots, y'B_k y)\]
where
\begin{equation}
A'X = I_k, \quad \mathrm{tr}(B_j \Sigma_m) = 0, \,\, j = 1, \ldots, k.
\end{equation}
For each $j$,
\[\mathrm{vec}^{*}(B_j)'\mathrm{vec}^{*}(\Sigma_{m}) = 0, \,\, m = 1, \ldots, n(n-1)/2.\]
Since $\{\Sigma_{m}\}$ spans $\mathbb{R}_{\mathrm{sym}}^{n\times n}$, $\{\mathrm{vec}^{*}(\Sigma_{m})\}$ spans $\mathbb{R}^{n(n-1)/2}$. Therefore,
\[\mathrm{vec}^{*}(B_j) = 0\Longrightarrow B_j = \textbf{0}_{n\times n}.\]
Therefore,
\[u(y) = A'y\in \mathbf{H}_{\mathrm{GLS}}(X).\]
This implies $\mathbf{U}_{r}(\mathbf{F}(X)\cap \mathcal{P}_{\mathrm{disc}}) \subset \mathbf{H}_{\mathrm{GLS}}(X)$. By Proposition (ref) and Proposition (ref), we know that $\mathbf{H}_{\mathrm{GLS}}(X)\subset \mathbf{U}_{r}(\mathbf{F}_r(X)) \subset \mathbf{U}_{r}(\mathbf{F}(X)\cap \mathcal{P}_{\mathrm{disc}})$. Thus, the result for $\mathcal{P}_{\mathrm{disc}}$ is proved. The result for $\mathcal{P}_{\mathrm{cont}}$ can be proved using the same argument.
A new nontrivial case where OLS/GLS is BUE
With the powerful representation theorem for $\mathbf{U}_{r}(\mathbf{F}_{2r}^\Sigma(X)\cap \mathcal{P})$ discussed in Theorem (ref), we can show that the OLS/GLS estimator is BUE in a sufficiently rich class that does not rule out quadratic estimators with respect to a subset of $\mathbf{F}_{2r}^\Sigma(X)$.
theoremAssume $n \ge \max\{k + 2, 4\}$. Let $\mathbf{G}_3^\Sigma(X)$ includes all distributions of $Y$ such that
\[\mathbb{E}_{F}[\epsilon_{i_1}\epsilon_{i_2}\epsilon_{i_3}] = 0, \,\, \text{ for any }(i_1, i_2, i_3)\in \{1, \ldots, n\}^3\setminus \{(i, i, i): i \le n\}, \,\, \text{where }\epsilon = \Sigma^{-1/2}(Y - X\beta).\]
Given any $r\ge 1$ and $\Sigma \succ \textbf{0}_{k\times k}$, let $VDV'$ be the eigen-decomposition of $\Sigma$ and
\[\mathbf{F}(X) = \bigcup_{\Sigma\in \mathcal{C}_{\Sigma}}\mathbf{F}_{2r}^\Sigma(X),\]
where
\[\mathcal{C}_{\Sigma} = \{V\tilde{D}V': \tilde{D}\in \mathbb{R}^{n\times n} \text{ is diagonal}\}.\]
Then
\[\mathbf{U}_{r}(\mathbf{F}(X)\cap \mathcal{P})\supsetneq \mathbf{H}_{\mathrm{GLS}}(X),\]
and $\hat{\beta}_{\mathrm{GLS}}^{\Sigma}(Y)$ is BUE in $\mathbf{U}_{r}(\mathbf{F}(X)\cap \mathcal{P})$ with respect to $\mathbf{F}_{2r}^{\Sigma}(X)\cap \mathbf{G}_3^{\Sigma}(X)$ for $\mathcal{P}\in \{\mathcal{P}_{\mathrm{cont}}, \mathcal{P}_{\mathrm{disc}}\}$.
proofWithout loss of generality, we assume that $\Sigma = I_{k}$; otherwise, we replace $Y$ by $\Sigma^{-1/2}Y = VD^{-1/2}Y$ and $X$ by $\Sigma^{-1/2}X = VD^{-1/2}X$, in which case $\hat{\beta}_{\mathrm{GLS}}^{\Sigma}(Y)$ reduces to $\hat{\beta}_{\mathrm{OLS}}(VD^{-1/2}Y)$ and $\mathcal{C}_{\Sigma}$ reduces to $\mathcal{C}_{I_k}$. For any $u\in \mathbf{U}_r(\mathbf{F}(X))$, using the same argument for (ref), we can show that $u(y) = A'y + (y'B_1 y, \ldots, y'B_k y)$ with
\begin{equation}
A'X = I_{k}, \,\, X'B_j X = 0_{k\times k}, \,\, \mathrm{tr}(B_j \Sigma) = 0, j= 1,\ldots, k, \Sigma\in \mathcal{C}_{I_k}.
\end{equation}
Since $\mathcal{C}_{I_k}$ includes all diagonal matrices, we must have
\[B_{j, ii} = 0, \,\, i = 1, \ldots, n, j = 1, \ldots, k;\]
that is, all diagonal elements of $B_{j}$ must be zero. For each $j$, this is equivalent to $n$ linear constraints on $\mathrm{vec}^{*}(B_j)$, i.e.,
\begin{equation}
\mathrm{vec}^{*}(e_i e_i')'\mathrm{vec}^{*}(B_j) = 0, \quad i = 1, \ldots, n
\end{equation}
where $e_i$ is the $i$-th canonical basis in $\mathbb{R}^{n}$. The other constraint $X' B_j X = 0$ can be formulated as $k(k-1)/2$ linear constraints on $B_j$, i.e.,
\begin{equation}
\tilde{e}_{p}'X'B_j X \tilde{e}_{q} = 0\Longleftrightarrow \mathrm{vec}^{*}(Xe_{q}e_{p}'X')'\mathrm{vec}^{*}(B_j) = 0, \quad 1\le p\le q \le n,
\end{equation}
where $\tilde{e}_{p}$ is the $p$-th canonical basis in $\mathbb{R}^{k}$. Note that (ref) and (ref) give all constraints on $B_j$, which correspond to a homogeneous linear system on $\mathrm{vec}^{*}(B_j)\in \mathbb{R}^{n(n-1)/2}$ with $n + k(k - 1)/2$ equations. Since $k \le n - 2$ and $n\ge 4$,
\[n + \frac{k(k-1)}{2}\le n + \frac{(n - 2)(n - 3)}{2} = \frac{n^2 - 3n + 6}{2} < \frac{n(n - 1)}{2}.\]
Therefore, a non-zero solution always exists. For any solution $B$ and $A\in \mathbb{R}^{n\times k}$ with $A'X = I_{k}$, consider the estimator
\[u_{B}(Y) = A'Y + (Y' B Y, Y' B Y, \ldots, Y' B Y).\]
Then, for any $F\in \mathbf{F}(X)$, then
\begin{align}
\mathbb{E}_{F}[Y' B Y] &= \beta' X' B X \beta + 2\beta'X' B\mathbb{E}_{F}[Y - X\beta] + \mathbb{E}_{F}[(Y - X\beta)'B(Y - X\beta)]\nonumber\\
& = \mathbb{E}_{F}[(Y - X\beta)'B(Y - X\beta)] = \mathbb{E}_{F}[\mathrm{tr}(B(Y - X\beta)(Y - X\beta)')] \nonumber\\
& = \mathbb{E}_{F}[B\Sigma] = 0,
\end{align}
where the last equality is due to that all diagonal elements of $B$ are zero and $\Sigma\in \mathcal{C}_{I_k}$ is diagonal. Therefore,
\[\mathbb{E}_{F}[u_{B}(Y)] = \mathbb{E}_{F}[A'Y] = A'X\beta = \beta\Longrightarrow u_{B} \in \mathbf{U}_1(\mathbf{F}(X)\cap \mathcal{P}).\]
Since $\mathbf{F}(X)\subset \mathbf{F}_{2r}(X)$, by Rosenthal's inequality rosenthal1970subspaces, $u_{B}\in \mathcal{L}^{r}(\mathbf{F}(X))$. Therefore,
\[\mathbf{U}_r(\mathbf{F}(X)\cap \mathcal{P})\supset \mathbf{H}_{\mathrm{GLS}}(X)\cup \{u_{B}\}\supsetneq \mathbf{H}_{\mathrm{GLS}}(X)\].
Next, we prove that $\hat{\beta}_{\mathrm{OLS}}(Y)$ is BUE in $\mathbf{U}_r(\mathbf{F}(X))$ with respect to $\mathbf{F}_{2r}^{I_k}(X)\cap \mathbf{G}_{3}^{I_{k}}(X)$. For any $u\in \mathbf{U}_r(\mathbf{F}(X))$, it suffices to show that,
\begin{equation}
\mathrm{Cov}_{F}(\hat{\beta}_{\mathrm{OLS}}(Y), u(Y) - \hat{\beta}_{\mathrm{OLS}}(Y)) = 0_{k\times k}, \quad for every F\in \mathbf{F}_{2r}^{I_{k}}(X)\cap \mathbf{G}_3^{I_{k}}(X).
\end{equation}
By (ref), we can write $u(y)$ as
\[u(y) = A'y + (y' B_1 y, \ldots, y' B_k y)',\]
where $A'X = I_{k}$, $X'B_{j}X = \textbf{0}_{k\times k}$, and $B_{j, ii} = 0$ for all $i = 1,\ldots, n$. Let $\epsilon = Y - X\beta$. Recall that we assume that $\Sigma = I_{k}$ without loss of generality. By definition of $\mathbf{F}_{2r}^{I_{k}}(X)$ in (ref), for any $F\in \mathbf{F}(X)$, $\mathrm{Cov}_{F}[\epsilon] = \sigma^2 I_{k}$ for some $\sigma^2 > 0$, and, analogous to (ref),
\[u(Y) - \mathbb{E}_{F}[u(Y)] = A'\epsilon + 2(B_1 X\beta, B_2 X\beta, \ldots, B_k X\beta)'\epsilon + (\epsilon'B_1 \epsilon, \ldots, \epsilon'B_k \epsilon)'.\]
Then for any $j = 1,\ldots, k$,
\begin{align*}
&\mathrm{Cov}_{F}(\hat{\beta}_{\mathrm{OLS}}(Y), u(Y) - \hat{\beta}_{\mathrm{OLS}}(Y))\\
&= \mathrm{Cov}_{F}((X'X)^{-1}X'\epsilon, (A' - (X'X)^{-1}X')\epsilon) + 2\mathrm{Cov}_{F}((X'X)^{-1}X'\epsilon, (B_1 X\beta, B_2 X\beta, \ldots, B_k X\beta)'\epsilon)\\
& \quad + \mathrm{Cov}_{F}((X'X)^{-1}X'\epsilon, (\epsilon' B_1 \epsilon, \ldots, \epsilon' B_k \epsilon)').
\end{align*}
Since $A'X = I_{k}$,
\begin{align}
&\mathrm{Cov}_{F}((X'X)^{-1}X'\epsilon, (A' - (X'X)^{-1}X')\epsilon)\nonumber\\
& = \mathbb{E}_{F}[(X'X)^{-1}X' \epsilon \epsilon' (A - X(X'X)^{-1})]\nonumber\\
& = \sigma^2 (X'X)^{-1}X'(A - X(X'X)^{-1}) = 0_{k\times k}.
\end{align}
Since $X' B_j X = 0$,
\begin{align}
& \mathrm{Cov}_{F}((X'X)^{-1}X'\epsilon, (B_1 X\beta, B_2 X\beta, \ldots, B_k X\beta)'\epsilon)\nonumber\\
& = \mathbb{E}_{F}[(X'X)^{-1}X'\epsilon\epsilon' (B_1 X\beta, B_2 X\beta, \ldots, B_k X\beta)]\nonumber\\
& = \sigma^2 \mathbb{E}_{F}[(X'X)^{-1}X'(B_1 X\beta, B_2 X\beta, \ldots, B_k X\beta)]\nonumber\\
& = \sigma^2 \mathbb{E}_{F}[((X'X)^{-1}(X'B_1 X)\beta, (X'X)^{-1}(X'B_2 X)\beta, \ldots, (X'X)^{-1}(X'B_k X)\beta)] = 0_{k\times k}.
\end{align}
Consider any $B$ with zero diagonal elements. For any $i = 1, \ldots, n$, let $z_{i}'$ be the $i$-th row of $(X'X)^{-1}X'$. Then
\[\mathrm{Cov}_{F}(z_{i}'\epsilon, \epsilon' B \epsilon) = \sum_{m\le n}\sum_{a\neq b \le n}z_{im}B_{ab}\mathrm{Cov}_{F}(\epsilon_{m}, \epsilon_{a}\epsilon_{b}).\]
Since $F\in \mathbf{G}_3^{I_{k}}(X)$, for any $m, a, b$ with $a\neq b$,
\[\mathbb{E}_{F}[\epsilon_m\epsilon_a\epsilon_b] = 0.\]
Since $\mathbb{E}_{F}[\epsilon_{m}] = 0$,
\[\mathrm{Cov}_{F}(\epsilon_{m}, \epsilon_{a}\epsilon_{b}) = \mathbb{E}_{F}[\epsilon_m\epsilon_a\epsilon_b] - \mathbb{E}_{F}[\epsilon_{m}]\mathbb{E}_{F}[\epsilon_{a}\epsilon_{b}] = 0.\]
Thus,
\begin{equation}
\mathrm{Cov}_{F}((X'X)^{-1}X'\epsilon, (\epsilon' B_1 \epsilon, \ldots, \epsilon' B_k \epsilon)') = 0_{k\times k}.
\end{equation}
Putting (ref) - (ref) together, we prove our goal (ref) and thus $\hat{\beta}_{\mathrm{OLS}}(Y)$ is BUE.
When $\Sigma$ is diagonal, Theorem (ref) is analogous to Theorem 5 of hansen2022modern (see also Theorem (ref) (b) (2) for a restatement in our notation), though we impose restrictions on the third moments instead of independence. It is worth emphasizing that neither result implies the other.