EconBase
← Back to paper

The Yule-Frisch-Waugh-Lovell Theorem for Linear Instrumental Variables Estimation

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

80,188 characters · 25 sections · 140 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

The Yule-Frisch-Waugh-Lovell Theorem for Linear Instrumental Variables Estimation

abstractIn this paper, I discuss three aspects of the Frisch-Waugh-Lovell theorem. First, I show that the theorem holds for linear instrumental variables estimation of a multiple regression model that is either exactly or overidentified. I show that with linear instrumental variables estimation: (a) coefficients on endogenous variables are identical in full and partial (or residualized) regressions; (b) residual vectors are identical for full and partial regressions; and (c) estimated covariance matrices of the coefficient vectors from full and partial regressions are equal (up to a degree of freedom correction) if the estimator of the error vector is a function only of the residual vectors and does not use any information about the covariate matrix other than its dimensions. While estimation of the full model uses the full set of instrumental variables, estimation of the partial model uses the residualized version of the same set of instrumental variables, with residualization carried out with respect to the set of exogenous variables. Second, I show that: (a) the theorem applies in large samples to the K-class of estimators, including the limited information maximum likelihood (LIML) estimator, and (b) the theorem does not apply in general to linear GMM estimators, but it does apply to the two step optimal linear GMM estimator. Third, I trace the historical and analytical development of the theorem and suggest that it be renamed as the Yule-Frisch-Waugh-Lovell (YFWL) theorem to recognize the pioneering contribution of the statistician G. Udny Yule in its development. \\ JEL Codes: C26.\\ Keywords: multiple regression; partial regression; Frisch-Waugh-Lovell theorem; instrumental variables estimation; k-class estimators; linear GMM estimators.

\doublespacing

Introduction

The Frisch-Waugh-Lovell (FWL) theorem is a remarkable result about linear regression models estimated with the method of least squares. The theorem shows that coefficients of variables in a multiple regression are exactly equal to corresponding coefficients in partial regressions that use residualized versions of the dependent and independent variables. While it has been primarily used in econometrics krishnakumar-2006,tielens-vanhove-2017, it has found wider applications in a variety of disciplines, including statistics arendacka-puntanen-2015,cinelli-hazlett-2020,gross-moeller-2023, electrical engineering monsurro-trifiletti-2017, computer science ahrens-etal-2021, and genetics & molecular biology reeves-etal-2012, to name just a few. It is now included in many graduate-level textbooks in econometrics, including davidson-mackinnon, hayashi, davidson-mackinnon-2004, angrist-pischke-2009, and greene. The theorem is widely used for estimating linear regression models with large number of dummy variables by partialiing them out, e.g. in panel data sets with large number of fixed effects mccaffrey-etal-2012, gaure-2013, Correia2017:HDFE.

The first contribution of this paper is to show that the FWL theorem can be derived for generalized linear instrumental variables estimation. In particular, consider a linear regression model with a mix of endogenous and exogenous variables. Suppose a researcher has a set of valid instrumental variables, where the number of instrumental variables is weakly larger than the number of endogenous variables---thus I allow for both exactly and over-identified models. Now consider two models: a full model and a partial model. In the full model, the researcher estimates all the parameters of the regression function using instrumental variables estimation (with the full set of instrumental variables). In the partial model, the researcher first residualizes the outcome variable, the endogenous regressors and the full set of instrumental variables with the set of exogenous variables, and then estimates a linear regression of the residualized outcome variable on the residualizes set of endogenous variables using instrumental variables estimation with the residualized instrumental variables.

In such a context, I show that: (a) the coefficient vector on the endogenous variables are numerically identical for instrumental variables estimation of the full and the partial model; (b) the vector of residuals are numerically identical for instrumental variables estimation of the full and partial model; and (c) estimated covariance matrices of the coefficient vector on the endogenous variables from the full and partial models are identical (up to a degree of freedom adjustment) if the estimator of the vector of regression errors are only functions of the regression residuals (and the only information used from the regressor matrices are their dimensions). Hence, this is the FWL theorem for instrumental variables estimation. This implies, in particular, that the argument about computational efficiency that led to widespread use of the FWL in least squares estimation of linear models with large number of dummy variables mccaffrey-etal-2012, gaure-2013, Correia2017:HDFE, can also be applied to instrumental variables estimation of such models.

The second contribution of the paper is to investigate if the FWL theorem applies to two closely related estimators: the limited information maximum likelihood (LIML) estimator and the linear generalized method of moments (GMM) estimator. I show that the FWL theorem applies in large samples to all K-class estimators (the LIML estimator being an important member of this class) for which the parameter $K$ that defines the estimator, tends to $1$. I also show that the theorem does not apply to linear GMM estimators (for overidentified models) in general, but that it does apply to the two step optimal linear GMM (2SGMM) estimator. The reason for this has to do with whether the estimator in question has the same structure as the 2SLS estimator, including, most importantly, in its expression a projection matrix.\footnote{Any square matrix which is symmetric and idempotent is known as a projection matrix strang-2006.} I show that it is possible to express all K-class estimators, where $K \to 1$ with sample size, in terms of an approximate projection matrix. In a similar manner, I also show that it is possible to express the two step optimal GMM estimator in terms of a suitably defined projection matrix. That is why the FWL theorem applies, approximately, to all consistent K-class estimators, including the LIML estimator, and equally well applies to the 2SGMM estimator. I give an empirical example to illustrate this point for both LIML and the optimal 2SGMM estimators.

The third contribution of the paper is to offer a historical narrative about the analytical development of the FWL theorem. Through this analytical narrative, I wish to highlight two points. The first point relates to the contribution of the statistician G. Udny Yule in the development of the FWL theorem. In the econometrics literature, the theorem is understood to have originated in a $ 1933 $ paper by Ragnar Frisch and Frederick V. Waugh in the first volume of Econometrica frisch-waugh-1933, which was later generalized by Michael C. Lovell lovell-1963. In fact, the result was proved more than two and a half decade ago by Yule in a $ 1907 $ paper udny-1907. This seems to be well known in statistics agresti-2015 and should be recognized in econometrics as well. To recognize Yule's pioneering contribution to the development of this important result, I suggest that we refer to it as the Yule-Frisch-Waugh-Lovell (YFWL) theorem, rather than the currently used Frisch-Waugh-Lovell (FWL) theorem.

The second point is to trace out the analytical development of the theorem through the decades highlighting the contribution at each stage of its development. In udny-1907, who proved the result using basic algebra, the partial regressions are always bivariate regressions of residual vectors. In frisch-waugh-1933, where the proof relied on some basic properties of determinants and Cramer's rule for finding solutions of linear systems of equations, the partial regressions are themselves multiple regressions but the residuals are computed with bivariate regressions. Thus, udny-1907 allows multiple variables in the conditioning set, but conceives of the partial regressions as bivariate regressions. On the other hand, frisch-waugh-1933 allows only one variable (the linear time trend) in the conditioning set, but allows the partial regressions to be multiple regressions.

Thirty years later, lovell-1963 generalizes the theorem significantly. He allows multiple variables in the conditioning set and allows the partial regressions to themselves be multiple regressions.\footnote{lovell-1963 points out that tintner-1957 had extended the Frisch-Waugh result to the case of polynomial time trends.} In terms of methodology, lovell-1963 introduces the use of projection matrices from linear algebra, making the proof compact and elegant. Current presentations of the theorem closely follow Lovell's exposition davidson-mackinnon-2004,greene. Finally, lovell-1963 shows that the vector of residuals from the full and partial regressions are identical, a result that neither udny-1907 nor frisch-waugh-1933 had investigated. This finding led Lovell to comment briefly on degrees of freedom adjustment necessary for statistical inference.

Till the 1960s, the YFWL theorem was used in the context of linear regression models estimated with least squares. While udny-1907 and frisch-waugh-1933 had considered ordinary least squares (OLS) estimation, lovell-1963 had pointed out that the theorem could be extended easily to generalized least squares estimation. Could the same set of results be derived for instrumental variables estimation? giles-1984 provided an answer in the affirmative for an exactly identified model (where the number of instrumental variables are exactly equal to the number of endogenous variables in the model). giles-1984 showed that both the coefficient vector and the residuals would be identical from instrumental variables estimations of the full and partial models.

The issue of statistical inference had not yet been investigated thoroughly, i.e. would researchers be able to conduct inference about the subset of coefficients in a multiple regression using the standard errors (or the covariance matrix) estimated from the partial regression? It is true that lovell-1963 had addressed this issue briefly by pointing out that estimated standard errors from the multiple and partial regressions are equal up to a degree of freedom adjustment when the true error is homoskedastic. But he left it at that. It is only in ding-2021 that we find a systematic treatment of the issue. ding-2021 has demonstrated that, in the context of least squares estimation, estimates of the covariance matrices from the multiple and partial regressions are equal up to a degree of freedom adjustment in the case of homoskedastic errors and, quite surprisingly, are exactly equal for some variants of heteroskedasticity consistent covariance matrices (HCCM), heteroskedasticity and autocorrelation consistent (HAC) covariance matrices, and for some variants of clustered robust covariance matrices.

In fact, the analysis in ding-2021 highlights a general principle: if the estimate of error covariance matrix is a function only of the residuals and does not depend on the matrix of covariates beyond its dimensions, then estimated covariance matrices from multiple and partial regressions are equal, either exactly or with a degree of freedom correction. This principle can be leveraged to offer a computational solution for cases where ding-2021's results do not hold, i.e. estimates and other information from partial regressions are sufficient to compute correct covariance matrices of coefficient vectors in the multiple regression. If only a subset of coefficients is of interest then estimation and proper statistical inference can be conducted by working with the partial regression alone.

The current paper is most closely related to giles-1984 and ding-2021. I extend giles-1984 and and ding-2021 in several ways. First, while giles-1984 had demonstrated the YFWL theorem for instrumental variables estimation of an exactly identified model, I extend the result to over-identified models. I offer a general proof that nests both cases of instrumental variables estimation, and of course least squares estimation (because, in this case, the regressors become their own instruments). Second, giles-1984 had not investigated properties of the estimated covariance matrices from full and partial models. I draw on ding-2021 to do so, and in so doing I also extend the analysis in ding-2021 from ordinary least squares estimation to generalized linear instrumental variables estimation. Third, neither giles-1984 nor ding-2021 had investigated whether the YFWL theorem can be applied to the K-class of estimators, including the LIML estimator, or linear GMM estimators. I do so and thereby extend the coverage of the YFWL theorem significantly.

I would like to end this introductory section by noting a limitation of the YFWL for generalized instrumental variables estimation. The key results of the theorem, i.e. equality of the coefficient vectors and residual vectors from full and partial models can only be derived when then conditioning set (the set of covariates used for residualizing) does not contain any endogenous variables. The reason has to do with the necessity of projection matrices in deriving key results of the YFWL theorem. If endogenous variables appear in the conditioning set, the relevant projection matrices are no longer available for the analysis. But this negative result is less restrictive than appears at first sight, as I discuss in section (ref).

The rest of the paper is organized as follows: in section (ref), I introduce the set-up and pose the main question; in section (ref), I present the YFWL for generalized instrumental variables estimation; in section (ref), I discuss why the YFWL theorem applies, in large samples, to all consistent K-class estimators (including the LIML estimator) but does not apply to linear GMM estimators (including the two step GMM estimator); in section (ref), I present a narrative account of the analytical development of the theorem through the decades; I conclude in section (ref) by highlighting the usefulness of, and highlighting a puzzle about the intellectual history of, the YFWL theorem. Proofs are collected together in the appendix.

The set-up and the question

Consider a linear regression of an outcome variable, $Y$, on a set of $k$ covariates,

equation[equation omitted — 63 chars of source]

where $Y$ is a $N \times 1$ vector of the outcome variable, $W$ is a $N \times k$ matrix of covariates, $\beta$ is a $k \times 1$ of parameters, and $\varepsilon$ is a $N \times 1$ vector of the stochastic error term.

Suppose on the basis of a specific question under investigation, it is possible for the researcher to partition the set of regressors into two groups, $W_1$ and $W_2$, where the former is a $N \times k_1$ matrix and the latter is a $N \times k_2$ matrix, with $k=k_1+k_2$, i.e. \[ W = \left[ W_1 : W_2\right]. \] The model in ((ref)) can then be written as

equation[equation omitted — 84 chars of source]

where $\beta_1$ and $\beta_2$ are $k_1 \times 1$ and $k_2 \times 1$ vectors of parameters, so that $\beta' = \left(\beta_1', \beta_2' \right) $. I would like to ensure that the model satisfies the identification condition greene for both sets of parameters, $\beta_1$ and $\beta_2$, using the following assumption.

assumptionThe matrices $W_1$ and $W_2$ are both of full column rank.

Suppose the researcher is only interested in the second group of regressors, $W_2$, i.e. she is only interested in estimating and conducting inference about $\beta_2$. But, of course, she does not want to completely throw away the information in $W_1$ because the variables in this subset do impact $Y$ and are likely to also be correlated to the variables in $W_2$. Hence, while she wants to condition her analysis on $W_1$, she is not interested in their partial effects on the outcome (e.g. fixed effects in a panel data regression model).

Since the researcher is only interested in the parameters on $W_2$, she could consider the following partial regression model,

equation[equation omitted — 113 chars of source]

where $\widetilde{Y}$ and $\widetilde{W}_2$ are residualized (or partialled out) versions of the outcome variable and the variables of interest in ((ref)), respectively, where residualization has been carried out with respect to $W_1$ (the exogenous covariates in the model), i.e.

equation[equation omitted — 100 chars of source]

where

equation[equation omitted — 101 chars of source]

are the hat and residual maker matrices, respectively, for the set of exogenous regressors, $W_1$.

The matrices $P_{W_1}$ and $M_{W_1}$ play important roles in the whole analysis. They are both symmetric and idempotent. They are referred to, respectively, as the `hat matrix' and the `residual maker matrix' because $P_{W_1}$ projects any vector in $\mathbb{R}^N$ onto the column space of $W_1$, and $M_{W_1}$ projects the same vector onto the orthogonal complement of the column space of $W_1$. Thus, $P_{W_1}$ generates the vector of predicted values and $M_{W_2}$ generates the vector of least squares residuals for a regression of any vector in $\mathbb{R}^N$ on $W_1$.

The researcher faces an additional challenge: while the conditioning variables in the model, $W_1$, are exogenous, i.e. $\mathrm{E} \left(\varepsilon | W_1 \right) = 0$, the variables of interest, $W_2$, are endogenous, i.e. $\mathrm{E} \left(\varepsilon | W_2 \right) \neq 0$. Since the model has endogenous regressors, the parameters cannot be estimated consistently by the method of least squares. Instead, the researcher turns to instrumental variables estimation using a set of $k_3$ instrumental variables, represented by the $n \times k_3$ matrix, $Z_2$, with $k_3 \geq k_2$.

assumptionThe set of instruments, $Z_2$, satisfy the following conditions greene: \begin{enumerate} • Exogeneity: The instrumental variables are uncorrelated with the error term in the full model, i.e. $\mathrm{Cov} \left( \varepsilon, Z_2\right) = 0$; • Relevance: The instrumental variables are correlated with the regressors. In particular, $Z'_2 W_1$ has rank $ k_1 $, and $ Z'_2 W_2 $ has rank $ k_2 $.\footnote{Note that both the exogeneity and relevance conditions can also be stated in asymptotic terms using probability limits.} \end{enumerate}

Given her interest and the set-up, can the researcher avoid estimating the full model in ((ref))? Can she work with the partial model ((ref)) and yet consistently estimate and conduct proper statistical inference on $\beta_2$ in ((ref))? The YFWL theorem provides an answer in the affirmative.

The YFWL theorem for linear instrumental variables estimation

Relationship between estimated coefficients

The full set of instrumental variables is given by \[ Z = \left[ W_1 : Z_2\right], \] so that the number of instrumental variables, $k_1+k_3$, can be larger than the number of endogenous variables, $k_1+k_2$, i.e. we are possibly dealing with an overidentified model. Hence, I will consider the 2SLS estimator, also known as the generalized instrumental variables estimator davidson-mackinnon-2004, for the full regression model ((ref)).\footnote{The 2SLS estimator uses, from among all possible linear combination of the set of instruments, the one that minimizes the asymptotic covariance matrix of the instrumental variables estimator brundy-jorgenson-1971.} The unique 2SLS estimator for coefficients in the full regression model ((ref)) is given by greene,

equation[equation omitted — 246 chars of source]

where $P_Z = Z \left( Z'Z\right)^{-1}Z' $ projects on to the column space of $Z$. Hence,

equation[equation omitted — 136 chars of source]

Turning to the partial model ((ref)), we estimate it with the instrumental variables formed by the residualized set of instrumental variables, $\widetilde{Z}_2=M_{W_1}Z_2$, which is a $N \times k_3$ matrix. Thus, the 2SLS estimator for the partial regression model ((ref)) is given by

equation[equation omitted — 208 chars of source]

where $P_{\widetilde{Z}_2} = \widetilde{Z}_2 \left( \widetilde{Z}_2' \widetilde{Z}_2\right)^{-1} \widetilde{Z}_2' $ projects on to the column space of $\widetilde{Z}_2$.

theoremFor instrumental variables estimation of the coefficients in the full regression model ((ref)) and the partial regression model ((ref)), we have \begin{enumerate} • the estimated coefficient vectors $b_2$ in ((ref)) and $\tilde{b}_2$ in ((ref)) are identical, and • the residual vectors from the full model ((ref)) and from the partial model ((ref)) are identical. \end{enumerate}

The results in theorem (ref) applies to the following special cases: (a) the instrumental variables (IV) estimator for an exactly identified model, when the number of instruments is exactly equal to the number of endogenous variables, i.e. $k_3=k_2$; (b) the OLS estimator, when both sets of regressors in ((ref)), $W_1$ and $W_2$, are exogenous. The latter case was proved by udny-1907, frisch-waugh-1933, lovell-1963; the former case was proved by giles-1984. Theorem (ref) generalizes both.

Relationship between estimated covariance matrices of coefficients

Estimated covariance matrices are equal up to a degree of freedom correction

Using the result in theorem 9.4 in greene, we see that the estimator of the asymptotic covariance matrix of $b'=\left( b'_1, b'_2\right) $, the instrumental variables estimator of $\left( \beta'_1, \beta'_2\right) $ in ((ref)) from the full regression model in ((ref)), is given by

align[align omitted — 545 chars of source]

where $W = \left[ W_1 : W_2\right]$, $Z = \left[ W_1 : Z_2\right]$, $P_Z = Z\left(Z'Z \right)^{-1} Z'$ and $\widehat{\Omega}_f$ (subscript `f' for identifying the full regression) is an estimate of the covariance matrix of the error term, $\varepsilon$, in the full regression model ((ref)). Hence, the estimated variance of $b_2$ is given by

equation[equation omitted — 286 chars of source]

Using a similar argument, we see that the estimator of the covariance matrix of $\tilde{b}_2$, the instrumental variables estimator of $\tilde{\beta}_2$ in the partial regression model ((ref)), is given by

equation[equation omitted — 460 chars of source]

where $\widehat{\Omega}_p$ (subscript `p' for identifying the partial regression) is an estimate of the covariance matrix of the error term, $\tilde{\varepsilon}$, in the partial regression model ((ref)).

theoremIf the estimators of the error vectors in the full and partial regression models are equal, i.e. $\widehat{\Omega}_f=\widehat{\Omega}_p$ (after making degrees of freedom corrections, if necessary), then $\widehat{\mathrm{Var}} \left( b_2\right)$ in ((ref)) and $\widehat{\mathrm{Var}} \left( \tilde{b}_2\right)$ in ((ref)) are equal (up to a degrees of freedom correction, if necessary).

Let me summarize what we get from theorem (ref): (a) if the estimate of the error covariance matrix depends only on the vector of regression residuals, then ((ref)) and ((ref)) are exactly equal; (b) if some degree of freedom correction is applied to generate the estimate of the error covariance matrix, then ((ref)) and ((ref)) are equal up to the relevant degree of freedom correction; (c) if the estimate of the error covariance matrix depends, in addition to the residual vector, on the elements of the matrix of regressors, then ((ref)) and ((ref)) are not, in general, equal even after making degree of freedom adjustments. In particular, theorem (ref) in this paper extends ding-2021 to linear instrumental variables estimation and we have the following results:

enumerate• In models with homoskedastic errors, the covariance matrices from the multiple and partial regressions are equal up to a degree of freedom adjustment , i.e. $(N-k_2)\widehat{\mathrm{Var}} \left( \tilde{b}_2\right) = (N-k) \widehat{\mathrm{Var}} \left( b_2\right)$. This is because $\widehat{\Omega}_p = 1/(N-k_2) \text{diag} \left[ \tilde{u}_i^2\right] $ and $\widehat{\Omega}_f = 1/(N-k) \text{diag} \left[ u_i^2\right] $, where `diag' denotes a $N \times N$ diagonal matrix and $k=k_1+k_2$. • In models with HC0 version of HCCM mackinnon-white-1985, the covariance matrices from the multiple and partial regressions are exactly equal because $\widehat{\Omega}_p = \text{diag} \left[ \tilde{u}_i^2\right] $ and $\widehat{\Omega}_f = \text{diag} \left[ u_i^2\right] $. • In models with HC1 version of HCCM mackinnon-white-1985, the covariance matrices from the multiple and partial regressions are equal up to a degree of freedom adjustment because $\widehat{\Omega}_p = N/(N-k_2)\text{diag} \left[ \tilde{u}_i^2\right] $ and $\widehat{\Omega}_f = N/(N-k) \text{diag} \left[ u_i^2\right] $. Hence, $(N-k_2)\widehat{\mathrm{Var}} \left( \tilde{b}_2\right) = (N-k) \widehat{\mathrm{Var}} \left( b_2\right)$. • In models using heteroskedasticity and autocorrelation (HAC) consistent covariance matrices newey-west-1987, the covariance matrices from the multiple and partial regressions are exactly equal because $\widehat{\Omega}_p = \left( \omega_{|i-j|} \tilde{u}_i \tilde{u}'_j\right)_{1 \leq i, j, \leq N} $ and $\widehat{\Omega}_f = \left( \omega_{|i-j|} \tilde{u}_i \tilde{u}'_j\right)_{1 \leq i, j, \leq N} $, as long as the same weights are used in both regressions. • In models using the cluster robust estimate of the variance matrix (CRVE) mackinnon-etal-2023, the covariance matrices from the multiple and partial regressions are exactly equal if no degree of freedom correction is used or if $G/(G-1)$ is used as the degree of freedom correction, where $G$ is the number of clusters cameron-miller-2015. If $G(N-1)/(G-1)(N-k)$ is used as the degree of freedom correction, then we have what mackinnon-etal-2023 call the $ CV_1 $ feasible CRVE and for this CRVE, the following holds: $(N-k_2)\widehat{\mathrm{Var}} \left( \tilde{b}_2\right) = (N-k) \widehat{\mathrm{Var}} \left( b_2\right)$.

In all these cases, researchers can use estimated covariance matrices from the partial regression ((ref)), with the relevant degree of freedom adjustment if necessary, to conduct proper statistical inference on the parameters from the multiple regression ((ref)).

The cases that were left out

There are some cases where the estimate of the error covariance matrix depends, in addition to the residual vector, on elements of the covariate and/or instrumental variables matrix. In these cases, ((ref)) and ((ref)) will not be equal, even after making degrees of freedom corrections. Some common cases where this happens are: (a) models using HC2, HC3, HC4, or HC5 forms of HCCM;\footnote{HC2 and HC3 were introduced by mackinnon-white-1985; HC4 was introduced by cribari-neto-2004; HC5 was introduced by cribari-neto-etal-2007. For a discussion of all the variants of HCCM, see mackinnon-2013.} (b) some variants of the HAC covariance matrix and the $CV_2$ and $CV_3$ forms of feasible CRVEs mackinnon-etal-2023. In these cases, it is not possible to use standard errors from the partial regression, with or without degrees of freedom correction, for inference about the coefficient vector in the full regression model. But, it is still possible to compute the correct covariance matrix for the coefficient vector in the full regression by using information that becomes available while estimating the partial regression. Thus, estimation of the full regression model can still be avoided.\footnote{See appendix (ref).}

Limitation of the YFWL theorem for instrumental variables estimation

In the YFWL for instrumental variables estimation, I have assumed that the set of variables that are of interest are endogenous and the set of variables used for conditioning are exogenous. Can we switch the role of the endogenous and exogenous variables and still get the YFWL theorem? I would like to show that the answer is in the negative but then also argue that this negative result is not too restrictive.

Let us return to the full regression model ((ref)) and specify that the set of variables of interest, $W_2$, is exogenous, and the set of variables used for conditioning (or residualizing), $W_1$, is endogenous. Suppose, now the researcher has a set of $k_3$ valid instrumental variables, $Z_1$, satisfying the exogeneity and relevance conditions, where $k_3 \geq k_1$ . The full set of instrumental variables is $Z = \left[Z_1 : W_2 \right] $, and the instrumental variables estimator of the parameter vector ($\beta'_1, \beta'_2$) is now given by ($b'_1, b'_2$), where,

equation[equation omitted — 215 chars of source]

so that, using results for the inverse of partitioned matrices greene, we have,

equation[equation omitted — 187 chars of source]

As can be seen from the above expression, we do not get any projection matrices. Neither can the inverse on the left be opened up because $W_2$ is not a square matrix.

Now let us turn to the partial regression, which, in this case will be

equation[equation omitted — 103 chars of source]

where $\widetilde{Y} = M_{W_1}Y$ and $\widetilde{W}_2 = M_{W_1} W_2$. Since $\widetilde{W}_2$ is orthogonal to $W_1$ (the endogenous variables), it is itself exogenous to the error of the full regression model. Hence, ((ref)) can be estimated by the method of ordinary least squares, giving us

equation[equation omitted — 148 chars of source]

If, instead instrumental variables estimation is conducted with the residualized instrumental variables, $\widetilde{Z}_1 = M_{W_1} Z_1 = Z_1$ (because $Z_1$ is orthogonal to $W_1$), we will have

equation[equation omitted — 118 chars of source]

In general, neither ((ref)) nor ((ref)) is equal to ((ref)). Hence, the neat results of the YFWL theorem cannot be derived in this case, i.e. for instrumental variables estimation when the set of conditioning variables are endogenous. This negative result is less restrictive than might appear, as I discuss below by way of summarizing the YFWL for instrumental variables estimation.

Summary

Let me summarize the facts about the YFWL theorem and highlight when it is applicable in terms of properties of the regressors and how they can be partitioned into two subsets. In the case of least squares estimation, ordinary or generalized, the YFWL theorem is applicable without any restrictions, i.e. the set of regressors can be partitioned into two subsets any which way. For instrumental variables estimation, there are some restrictions that need to be kept in mind. These restrictions arise because two different distinctions are at play that might not coincide: exogenous variables versus endogenous variables; and variables of interest versus those not of direct interest (and hence relegated to the conditioning set).

Let $S_W$, $S_{X}$ and $S_{I}$ denote, respectively, the full set of covariates, the subset of exogenous variables and the subset of variables of interest. Let $S_{N}=S_W \setminus S_{X}$ denote the set of endogenous variables, and let $S_{D}=S_W \setminus S_{I}$ denote the set of conditioning variables (variables not of direct interest but necessary for conditioning).

enumerate• When $S_I = S_N$, then the YFWL theorem is applicable for instrumental variables estimation (using the residualized set of instrumental variables) as shown in theorem (ref) and theorem (ref). • If $S_I \supset S_N$, then the YFWL theorem is applicable for instrumental variables estimation. Since the conditioning set does not contain any endogenous variables, the results of theorem (ref) and theorem (ref) apply. The only change that is required is to expand $W_1$ (the set of endogenous variables) to include the exogenous variables that are of interest and to then use these variables as their own instrumental variables. • If $S_I \subset S_N$, then the YFWL theorem can be made applicable for instrumental variables estimation in the following way: take the set of variables of interest as being equal to the set of endogenous variables (or, equivalently, take the conditioning set to be the set of all the exogenous variables only) and apply the result of case 1 (so that results of theorem (ref) and theorem (ref) apply); finally, extract coefficients of interest and their covariance matrix. • If $S_I \supseteq S_X$, then the the YFWL theorem for instrumental variables estimation cannot be applied, as demonstrated in section (ref) (for the case when $S_I = S_X$ or equivalently $S_D = S_N$). Note that this is the only case when the YFWL theorem is not applicable for instrumental variables estimation cannot be applied. Hence, the negative result is less restrictive than it might have appeared at first sight. After all, it only restricts one out of four possibilities

Does the FWL theorem apply to K-class and linear GMM estimators?

The LIML estimator (a member of the K-class estimators) and the linear GMM estimator are closely related to the 2SLS estimator (another member of the K-class of estimators) discussed above. Hence, it is natural to ask whether the FWL theorem, which I have derived above for the 2SLS estimator, applies to the K-class and linear GMM estimators.\footnote{The popular ivreg2 suite of functions in STATA seems to allow for partialling for all these estimators. The paper that explains the technical details of ivreg2 claims in an informal argument, by relying on the FWL theorem, that: “The invariance of the estimation results to partialling-out applies to one- and two-step estimators such as OLS, IV, LIML, and two-step GMM, but not to CUE or to GMM iterated more than two steps.” baum-etal-2007. The paper does not offer any formal proofs of these claims.} I will demonstrate that: (a) the YFWL theorem applies, in large samples, to all K-class estimators for which $K \to 1$, and (b) the YFWL theorem applies to the 2SGMM estimator but does not apply to general linear GMM estimators. The clue to arriving at this answer is offered by trying to understand the common structure of the 2SLS, K-class and linear GMM estimators, and then using a crucial step from the proof of theorem (ref).\footnote{For a discussion of the K-class of estimators, see davidson-mackinnon.}

The theory

To get started, let us recall from ((ref)) that the 2SLS estimator of $b_2$ in the full model ((ref)) is given by

equation[equation omitted — 104 chars of source]

and the 2SLS estimator of $\tilde{b}_2$ in the partial model ((ref)) is given by

equation[equation omitted — 154 chars of source]

In the proof of theorem (ref), I showed that ((ref)) and ((ref)) were equal. What is important is to note that the proof of theorem (ref) relies crucially on $P_Z$ in ((ref)) being a projection matrix, i.e. a symmetric and idempotent $n \times n$ matrix strang-2006. This is important because when $P_Z$ is a projection matrix, we can decompose it into two orthogonal matrices, $P_{W_1}$ and $P_{\widetilde{Z}_2}$ rao-etal-2008. It is because of this decomposition that I could show the crucial equality in theorem (ref). This means that if I can express any estimator in the same form as the 2SLS estimator in ((ref)), then I can replicate the argument that established theorem (ref). This will then allow me to claim that the YFWL theorem applies to that estimator. This will be my general strategy to investigate whether the YFWL theorem applies to the K-class and linear GMM estimators.

K-class of Estimators

General case. The K-class of estimators for $b$ in the full model ((ref)) is given by davidson-mackinnon:

equation[equation omitted — 115 chars of source]

where $K$ is a scalar and $M_Z = I - Z\left(Z'Z \right)Z'$ is the residual maker matrix with respect to the set of instrumental variables. When $K=0$, we get the OLS estimator; when $K=1$, we get the 2SLS estimator; when $K=\hat{\kappa}$, where $\hat{\kappa}$ is the smallest eigenvalue of $\left(U'M_{Z}U \right)^{-1/2} U'M_{Z}U \left(U'M_{Z}U \right)^{-1/2}$, where $U=\left[ Y : W_2\right] $, we get the LIML estimator.\footnote{Note that $\left(W'M_{Z}W \right)^{-1/2}$ denotes the inverse of the square root factor of the positive definite matrix $U'M_{Z}U$.}

With regard to the K-class of estimators in ((ref)), note that, in general, $\left(I - K M_Z \right)$ is symmetric, because $M_Z$ is symmetric, but not idempotent. Hence, $\left(I - K M_Z \right)$ is not a projection matrix for a general value of $K$. This implies that, for a general value of $ K $, the K-class of estimators does not have the same structure as the 2SLS estimator in ((ref)). Thus, the arguments that establish theorem (ref) and (ref) will not work. Hence, my conjecture is that for a general value of $K$, the YFWL does not apply to the K-class of estimators.

Special case. But there is an important subset of K-class estimators where it does. To identify this subset, note that for all K-class estimators, where $K$ tends to $1$ at an appropriately fast rate, $\left(I - K M_Z \right)$ is approximately idempotent (in large samples) because \[ \left(I - KM_Z \right)\left(I - KM_Z \right) = I - KM_Z + \left( K^2 - K\right)M_Z \approx I - M_Z \] where I have used the idempotency of $M_Z$ and the fact that $K \approx 1$ in large samples.\footnote{K-class estimators, where $K$ tends to $1$ at a rate faster than $n^{-1/2}$, are consistent davidson-mackinnon.} Thus, for K-class estimators where $K$ tends to $1$---which includes 2SLS, LIML, Fuller, but not OLS---$\left(I - KM_Z \right)$ behaves approximately like the projection matrix $P_Z=\left(I - M_Z \right)$ in large samples. Thus, these K-class estimators have the same structure as the 2SLS estimator in ((ref)) and the arguments of theorem (ref) and (ref) will apply. That is why the YFWL theorem can be applied approximately in large samples for the subset of K-class estimators for which $K \to 1$ with sample size, including the LIML estimator.\footnote{An alternative argument might be as follows: since the LIML is the limit of an iterative process involving OLS estimators, and since the YFWL theorem applies to OLS estimators, the YFWL theorem is also applicable to the LIML estimator. This argument is incomplete unless one shows that the set of estimators that form the sequence in the iterative process belong to a closed set. If it is an open set, then the limit (LIML) will not necessarily inherit the properties of the estimators that are sequences in the iterative process.}

Linear GMM Estimators

General case. The linear GMM estimator for $b$ in the full model ((ref)) is given by cameron-trivedi-2005:

equation[equation omitted — 82 chars of source]

where $W_N$ is some symmetric, positive definite $k_3 \times k_3$ weighting matrix (which is necessarily of full rank). Note that, in general, while $ZW_NZ'$ is symmetric, it is not idempotent.\footnote{As a counterexample, consider $W_N=I$ (the identity matrix of the correct size), which is symmetric and of full rank. Hence, $ZW_NZ'=ZZ'$. Now note that, in general, $ZZ'ZZ'$ is not equal to $ZZ'$. Thus, $ZW_NZ'$ is not idempotent. Hence, it cannot be a projection matrix. For purposes of illustrating the weighting matrix, cameron-trivedi-2005 also use $W_N=I$.} This implies that in the linear GMM estimator in ((ref)), $ZW_NZ'$ is not a projection matrix, even in large samples and not even approximately. Hence, the arguments that establishes equality of the coefficient vectors from the full and partial regressions in theorem (ref) will not work. Hence, my conjecture is that the FWL theorem cannot be applied to general linear GMM estimators.\footnote{This negative result for linear GMM estimators pertain to overidentified models. For exactly identified models, the linear GMM estimator reduces to the IV estimator, for which we can apply the FWL theorem, as I have shown in theorem (ref) and theorem (ref).}

Special case. But there is a special case of the linear GMM estimator where the YFWL theorem does apply. This is the 2SGMM estimator that uses the specific weighting matrix, $W_N$, to minimize the asymptotic variance of the estimator. For $b$ in the full model ((ref)) the 2SGMM arises with the choice of $W_N = N\left( Z'DZ\right)^{-1} $, where $D$ is a $N \times N$ diagonal matrix with the principal diagonal populated by $\widehat{u}_i^2$, where $\widehat{u}$ is the $N \times 1$ vector of residuals of the full model estimated by some consistent estimator, e.g. 2SLS cameron-trivedi-2005. Thus, the two step linear GMM (2SGMM) estimator for $b$ is given by

equation[equation omitted — 124 chars of source]

where the $N$ in the definition of $W_N$ has canceled out.

Let $\widehat{Z} = D^{1/2}Z$, where $D^{1/2}$ is the $N \times N$ matrix with the principal diagonal populated by $\widehat{u}$ (residuals from the full model). Now define $P_{\widehat{Z}} = D^{1/2}\widehat{Z} \left( \widehat{Z}' D \widehat{Z} \right)^{-1} \widehat{Z}' D^{1/2} = D^{1/2} \widehat{Z} \left( \widehat{Z}' D^{1/2} D^{1/2} \widehat{Z} \right)^{-1} \widehat{Z}' D^{1/2}$. It is easy to see that $P_{\widehat{Z}}'=P_{\widehat{Z}}$ and \[ P_{\widehat{Z}} P_{\widehat{Z}} = D^{1/2}\widehat{Z} \left( \widehat{Z}' D \widehat{Z} \right)^{-1} \widehat{Z}' D^{1/2} D^{1/2}\widehat{Z} \left( \widehat{Z}' D \widehat{Z} \right)^{-1} \widehat{Z}' D^{1/2} = D^{1/2}\widehat{Z} \left( \widehat{Z}' D \widehat{Z} \right)^{-1} \widehat{Z}' D^{1/2} = P_{\widehat{Z}}. \] Thus, $P_{\widehat{Z}}$ is a symmetric and idempotent matrix. Hence, $P_{\widehat{Z}}$ is a projection matrix strang-2006. Let $\widehat{Y} = D^{-1/2}Y$, $\widehat{W} = D^{-1/2}W$, where $D^{-1/2}$ is the $N \times N$ matrix with the principal diagonal populated by the reciprocal of $\widehat{u}$. The two step linear GMM estimator for $b$ in ((ref)) can be written as

equation[equation omitted — 167 chars of source]

Thus, the two step linear GMM estimator in ((ref)) has the same structure as the 2SLS estimator in ((ref)), including, crucially, the role played by a projection matrix in both. Hence, the arguments of theorem (ref) and (ref) can be applied to ((ref)). This implies that the YFWL applies to the 2SGMM estimator.

An empirical example

As an empirical example, I return to David Card's $ 1995 $ study of returns to schooling using geographical proximity as instrumental variable card-1995. The model is

equation[equation omitted — 96 chars of source]

where $y$ is earnings (wage in cents per hour in $1976$), $SCHOOLING$ is years of schooling, and $X$ is a set of exogenous controls like experience (and its square), race (whether black), whether lived in the South, whether lived in an SMSA, and regional dummy variables.

The model ((ref)) is estimated with OLS, 2SLS, LIML, Fuller's K-class estimator, IGMM (a GMM estimator where the identity matrix of the correct dimension is used as the weighting matrix) and 2SGMM. Other than OLS, estimation treats $SCHOOLING$ as endogenous and uses two instrumental variables to deal with endogeneity: (a) proximity to $2$ year college; (b) proximity to $4$ year college.\footnote{I use ivmodel in R to get the data and implement the K-class of estimators. I write my own code for the IGMM and 2SGMM estimators. For the IGMM estimator, I use an identity matrix of the correct size as the weighting matrix. For the 2SGMM estimator, I implement the following steps: In the first step, I estimate the model by 2SLS. Using the residuals from the first stage, I construct the optimal weighting matrix. In the second step, I use the optimal weighting matrix (from the first step) and solve the linear system of equations whose solution appears as the 2SGMM estimator in ((ref)). R code for this example is available here: \url{https://drive.google.com/file/d/1ug0SZVwuW5vuxBHZbsVQWlaNXoxxNBMW/view?usp=sharing}.} The source of the data is the National Longitudinal Survey of Young Men (NLSYM) and the sample size is $3101$.\footnote{“The NLSYM began in 1966 with 5525 men age 14--24 and continued with follow-up surveys through 1981.” card-1995.}

Two sets of results are presented in Table (ref). Row 1 reports results of estimating the full model; row 2 presents estimates of the partial model, where the latter model partials out the full set of exogenous variables from the outcome variable (log earnings), from the endogenous variable (years of schooling) and from the two instrumental variables (proximity to $2$ year and $4$ year college). From Table (ref), we see that the parameter estimates from the full and partial models are identical for the K-class of estimators (OLS, 2SLS, Fuller and LIML), and also for the 2SGMM estimator.\footnote{In a previous version of this paper, an error in my R code had given rise to different 2SGMM estimates for the full and partial model. I would like to thank Kit Baum for pointing out that my 2SGMM results were possibly incorrect.} But the estimates are different between the full and partial model for the IGMM estimator: $0.13753$ for the full model and $0.16274$ for the partial model. This confirms the theoretical argument presented in the previous sub-section as to why the FWL applies to the common K-class of estimators and to the 2SGMM estimator, but not to a general linear GMM estimator.

table[table omitted — 1,173 chars of source]

Historical evolution of the YFWL theorem

Pioneering contribution of G. Udny Yule

Novel notation

udny-1907 considers the relationship among $n$ random variables $X_1$, $ X_2$, $\ldots$, and $X_n$ using a sample which has $N$ observations on each of the variables. In particular, suppose this relationship is captured by a linear regression of the first variable, $X_1$, on the other variables, $X_2, \ldots, X_n$, and a constant. This linear regression is what Yule called a multiple regression. With reference to ((ref)), therefore udny-1907 uses: $k=n$, $X_1=Y$, $X_2$ the second column of $W$, $X_3$ the third column of $W$, and so on.

Subtracting the sample mean from all variables, the multiple regression equation can be equivalently expressed in `deviation from mean' form in which the constant drops out. Let $x_1, x_2, \ldots, x_n$ denote the same random variables, but now expressed as deviations from their respective sample means. Introducing a novel notation to express the coefficients and the residual, udny-1907 writes the sample analogue of the multiple regression function as follows:

equation[equation omitted — 213 chars of source]

In ((ref)), the subscripts of the coefficient on each regressor, as also the residual, is divided into the primary and secondary subscripts. The primary subscripts come before the period; the secondary subscripts come after the period. For the coefficients of regressors, the primary subscripts identify the dependent variable and that particular regressor, in that order; and the secondary subscripts identify the other regressors in the model. For the residual in ((ref)), $x_{1.234 \ldots n}$, the primary subscript identifies the dependent variable and the secondary subscripts identify all the regressors.

It is useful to note three features of the subscripts. First, the order of the primary subscripts is important but the order among the secondary ones are not because the order in which the regressors are arranged in immaterial for estimation and inference. For instance, $b_{12.34 \ldots n}$ denotes the coefficient on $x_2$ for a regression of $x_1$ on $x_2, x_3, \ldots, x_n$. Similarly, $b_{1j.234 \ldots n}$ denotes the coefficient on $x_j$ for a regression of $x_1$ on $x_2, x_3, \ldots, x_n$, and note that the secondary subscripts excludes $ j $. Second, elements of the subscripts cannot be repeated because it is not meaningful to include the dependent variable as a regressor or repeat a regressor itself. Third, since the coefficients of the multiple regression function are related to partial correlations, the notation was also used for denoting relevant partial correlations too.\footnote{See, for instance, equation (2) in udny-1907.}

The theorem

Now consider, with reference to the multiple regression ((ref)), what Yule called partial regressions. These refer to regressions with partialled out variables. For instance, with reference to ((ref)), the partial regression of $x_1$ on $x_2$ would be constructed as follows: (a) first run a regression of $x_1$ on all the variables other than $x_2$, i.e. $x_3, x_4, \ldots, x_n$, and collect the residuals; (b) run a regression of $ x_2 $ on all the other regressors, i.e. $x_3, x_4, \ldots, x_n$, and collect the residuals; (c) run a regression of the residuals from the first step (partialled out outcome variable) on the residuals from the second step (partialled out first regressor).

udny-1907 showed that, if parameters of the multiple regression and the partial regression(s) were estimated by the method of least squares then the corresponding coefficients would be numerically the same. For instance, considering the first regressor, $x_2$, in ((ref)), he demonstrated that

equation[equation omitted — 164 chars of source]

where the summation runs over all observations in the sample.\footnote{For a proof see appendix (ref).}

To see why ((ref)) gives the claimed result, recall that $x_{1.34 \ldots n}$ is the residual from a regression of $ x_1 $ on $ x_3, \ldots, x_n $ and $x_{2.34 \ldots n}$ is the residual from the regression of $ x_2 $ on $ x_3, \ldots, x_n $. By construction, both have zero mean in the sample. Recall, further, that in a regression of a zero-mean random $y$ on another zero-mean random variable $z$, the coefficient on the latter is $\sum y z /\sum z^2$. Hence, the left hand side of ((ref)) is the coefficient in a regression of $x_{1.34 \ldots n}$ on $x_{2.34 \ldots n}$. The right hand side is, of course, the coefficient on $x_2$ in ((ref)). Hence, the result.

The above argument, of course, can be applied to any of the regressors in ((ref)) and thus udny-1907 proved the following general result: the coefficient on any regressor in ((ref)) is the same as the coefficient in a bivariate regression of the residualized dependent variable (the vector of residual from a regression of the dependent variable on the other regressors) and the residualized regressor (the vector of residuals from a regression of the regressor of interest on the other regressors).

With reference to the question posed in section (ref), udny-1907, therefore, showed that a researcher could work with partial regressions if she were only interested in a subset of the coefficients in the multiple regression. She would arrive at the same estimate of the parameters by estimating partial regressions as she would if she had estimated the full model. In particular, when $W_1$ contained only one variable (other than the column of $1$s), udny-1907 provided the following answer: (a) run a regression of $y$ (demeaned $Y$) on the variables in $w_2$ (demeaned variables in $W_2$), and collect the residuals; (b) run a regression of $w_1$ (demeaned variable in $W_1$, i.e. excluding the constant) on the variables in $w_2$, and collect the residuals; (c) regress the first set of residuals on the second to get the desired coefficient.

There are two things to note about Yule's answer. First, he did not provide an answer to the question posed in section (ref) when $W_1$ contained more than one random variable (excluding the constant). Of course, partial regressions could be estimated for each independent variable, but they had to be separately estimated as bivariate regressions with all the other variables used for the partial regressions. Second, udny-1907 did not investigate the relationship among variances of the parameters estimated from multiple and partial regressions. Is the estimated standard error of a coefficient also identical from the full and partial regressions? Yule did not pose, and hence did not provide any answers to, this question.

Application to a regression with a linear time trend

We can immediately apply the result from Yule's theorem given above to a regression with a linear time trend. With ((ref)), let $x_n$ denote the demeaned linear time trend variable. Applying ((ref)), we will have the following result: the coefficient on any of the regressors, $x_2, x_3, x_{n-1}$ in ((ref)) is the same as the coefficient in a bivariate regression of residualized $x_1$ on the relevant regressor. This is of course the exact same result that was presented $26$ later in frisch-waugh-1933.

Frisch and Waugh's result

Notation and result

frisch-waugh-1933 study the relationship among $n+1$ variables, one of which is a linear time trend. In reference to ((ref)), therefore frisch-waugh-1933 use: $k=n$, $X_0=Y$, $X_1$ the second column of $W$, $X_2$ the third column of $W$, and so on. They use the same convention of considering variables in the form of deviations from their respective sample means.\footnote{“Consider now the $ n $ variables $x_0, \ldots, x_{n-1}$ and let time be an $ (n+1)$th variable $x_n$. Let all the variables be measured from their means so that $\sum x_i=0 \quad (i = 0, \ldots, n)$ where $\sum$ denotes a summation over all the observations.” frisch-waugh-1933.} Using the exact same notation as used by udny-1907 and postulating an a priori true linear relationship among these variables, they write,

equation[equation omitted — 134 chars of source]

and consider estimating the parameters of ((ref)) in two ways. First, they consider the multiple regression of $x_0$ on $x_1, \ldots, x_n$:

equation[equation omitted — 122 chars of source]

Second, they consider the multiple regression of $x'_0$ on $x'_1, \ldots, x'_{n-1}$ (note that $x_n$ has been excluded from the set of regressors),

equation[equation omitted — 136 chars of source]

where the primed variables are the corresponding time-demeaned original variables, i.e., for $j = 0, 1, \ldots, n-1$, $x'_j$ is the residual in the regression of $x_j$ on $x_n$.\footnote{Note that ((ref)), ((ref)) and ((ref)) are just the predicted regressions functions. These equations exclude the residuals.} Using the basic rules of determinants and Cramer's rule for solving linear equation systems, frisch-waugh-1933 demonstrated that the coefficients denoted by $b$ in ((ref)) are numerically equal to the corresponding coefficients denoted by $b'$ in ((ref)).\footnote{For a proof see appendix (ref). In his presentation of the Frisch-Waugh theorem, chipman-1998 argues as if Frisch and Waugh had used projection matrices in their proof. That is not correct. frisch-waugh-1933 did not use projection matrices in their proof.}

With regard to the question posed in section (ref), Frisch and Waugh provide the same answer as Yule: coefficients are the same whether they are estimated from the multiple or from partial regressions. There are both similarities and differences between udny-1907 and frisch-waugh-1933. First, whereas in udny-1907, only one variable could be included in $W_1$ (the subset of covariates that was of interest to the researcher), in frisch-waugh-1933 only one random variable could be included in $W_2$ (the subset of covariates that was not of interest to the researcher). Second, much like udny-1907 before them, frisch-waugh-1933 did not investigate the relationship among estimated variances of the parameters of multiple and partial regressions. The question that is relevant for statistical inference, i.e. standard errors, had still not been posed.

Lovell extends the OLS analysis

lovell-1963 extended the reach of the theorem significantly and addressed both questions that had been left unanswered by udny-1907 and frisch-waugh-1933. On the one hand, lovell-1963 partitioned the set of regressors, $W$, into two subsets without any restrictions on the number of variables in each subset; on the other, he laid the groundwork for thinking about the estimated covariance matrices of the coefficient vectors.

In the context of OLS estimation, lovell-1963 demonstrated two important results: (a) the coefficient vectors are numerically the same whether they are estimated from the multiple or the partial regressions, and (b) the vector of residuals from the multiple and partial regressions are numerically the same.\footnote{I omit the proofs because they are just special cases of theorem (ref) in this paper.} The first result completed the YFWL so far as the estimate of the coefficient is concerned because the partitioning of the set of regressors was completely general; the second result laid the groundwork for comparing estimated variances of the coefficient vectors from multiple and partial regressions.\footnote{It is straightforward to extend the YFWL from ordinary least squares to generalized least squares estimation, as Lovell had noted when commenting on autocorrelated errors lovell-1963. Other scholars have worked on variations of Lovell's results; see, for instance, fiebig-bartels-1996, krishnakumar-2006.}

D. Giles extends the theorem to IV estimation

So far, the YFWL theorem was always posed in the context of least squares, ordinary or generalized, estimation. To the best of my knowledge, giles-1984 was the first scholar to extend the theorem to instrumental variables estimation of an exactly identified model, i.e. in the setting where the number of instruments was exactly equal to the number of endogenous variables.\footnote{“In the following discussion the necessary inverse matrices are assumed to exist, and we consider only the case of equal numbers of regressors and instruments as this includes such common estimators as two-stage least squares and other members of the k-class.” giles-1984. The last part of the sentence uses non-standard terminology. In current terminology, the two-stage least squares estimator refers to the case where the number of instruments is larger than the number of endogenous variables.} giles-1984 demonstrated that the YFWL theorem holds for the IV estimator (instrumental variables estimator when the model is exactly identified) under $13$ combinations of partialling. He shows that the coefficient vectors from the full and partial regression models are identical and that their asymptotic covariance matrices are the same. giles-1984 did not comment on the estimators of asymptotic covariance matrices of the coefficient vector and, therefore, did not discuss the need for degrees of freedom corrections.

P. Ding considers standard errors for OLS estimation

For more than $100$ years, the evolution of the theorem had primarily focused on comparing the coefficient vectors from multiple and partial regressions. The YFWL theorem had established that they are numerically the same. This was a very useful result but still did not answer the question that would be relevant if a researcher were interested in conducting statistical inference about the coefficient vector. To conduct inference about $b_2$ using results of estimating $\tilde{b}_2$, we also need to be able to compare their estimated covariance matrices. While lovell-1963 had commented briefly on this issue, to the best of my knowledge, ding-2021 provided the first systematic treatment. ding-2021 demonstrated that, in the context of ordinary least squares estimation, covariance matrices of coefficient vectors from partial and full regressions would, in many contexts, be exactly equal or equal up to a degrees of freedom correction. I have extended these sets of results to instrumental variables estimation and have commented on the different cases in section (ref) above.

Discussion and conclusion

By way of concluding this paper, I would like to offer two sets of comments. My first set of comments is about the substance and usefulness of the YFWL theorem, and the second set is about a puzzle in the history of the YFWL theorem.

There are two reasons that make the YFWL theorem so useful, the first computational and the second conceptual. Let me start with the computational usefulness. For a linear regression model estimated by the method of least squares, the YFWL theorem shows that if a researchers is interested in estimating only a subset of the parameters of the model, she need not estimate the full model. She can partition the set of regressors into two subsets, those that are of interest and those that are not of direct interest (but needs to be used for conditioning). She can regress the outcome variable on the conditioning set to create the residualized outcome variable. She can regress each of the variables of interest on the conditioning set to create corresponding residualized covariates. Finally, she can regress the residualized outcome variable on the residualized covariates of interest to get the estimated parameter of interest.

If, in addition, the variables of interest are endogenous and she has a set of valid instrumental variables, she can carry out the same procedure as above for instrumental variable estimation. In addition to creating the residualized outcome variable and the residualized covariates of interest (the endogenous variables), she must create residualized instrumental variables with respect to the exogenous variables in the model. Now, if she estimates an instrumental variables regression of the residualized outcome variable on the residualized covariates of interest using the residualized instrumental variables, then she will get consistent estimates of the parameters of interest.

In large samples, the same procedure can be justified for the subset of K-class estimators (LIML being an important member of this class) for which $K \to 1$ with sample size because the YFWL theorem applies approximately to this class of estimators in large samples. It is important to note that the same procedure cannot be used for linear GMM estimators in general because the YFWL does not apply to such estimators. But the YFWL does apply to the 2SGMM estimator in linear models. This means that the YFWL theorem applies to OLS estimators, linear instrumental variables estimators, all K-class estimators for which $K \to 1$ with sample size and the 2SGMM estimator. Thus the results in this paper significantly increases the coverage of the YFWL theorem for linear models.

For statistical inference, she can work with a partial regression in many, but not all, cases. If the errors are homoskedastic, then the estimated covariance matrix from the partial regression can be used for inference once a degree of freedom adjustment is made. If the errors are heteroskedastic, as is often the case in cross sectional data sets, and the researcher wishes to use the HC0 or HC1 variant of HCCM, then she can use the estimated covariance matrix from the partial regression without any adjustment in the case of HC0 and with a degree of freedom adjustment in the case of HC1. If the researcher is using a time series data set and wishes to use a HAC covariance matrix, then she can use the estimated covariance matrix from the partial regression as is. If the researcher is using a panel data set and wishes to use the standard cluster robust covariance matrix, then she can use the estimated covariance matrix from the partial regression without any adjustment; if a degree of freedom correction is used to compute the cluster robust covariance matrix, then she will need to apply a corresponding degree of freedom adjustment.

If, on the other hand, the researcher wishes to use the HC2, HC3, HC4 or HC5 variants of heteroskedasticity consistent covariance matrix, or if she wishes to use other variants of the cluster robust covariance matrix or HAC covariance matrices, then she cannot use above results, i.e. the covariance matrices from multiple and partial regressions are no longer equal, even after degrees of freedom adjustments. Instead, the researcher can use the four-step computational method proposed in appendix (ref) if she wishes to use information from the partial regression to conduct proper statistical inference about parameters in the multiple regression.

While the question of computation is certainly important, the YFWL theorem is perhaps even more relevant because of its indispensability for an intuitive understanding of multiple regression and partial correlations. In non-experimental settings, whenever a researcher has to analyze the relationship among three of more variables, she has to address a question that G. Udny Yule, a pioneer of regression analysis, raised more than a century ago.

quote[F]or instance, it might be found that changes in pauperism were highly correlated (positively) with changes in the out-relief ratio, and also with changes in the proportion of old; and the question might arise how far the first correlation was due merely to a tendency to give out-relief more freely to the old than the young, i.e. to a correlation between changes in out-relief and changes in the proportion of the old. yule-1911.

In other words, how would a researcher distinguish between the true, or partial, correlation between changes in pauperism and out-relief ratio from the multiple (or total) correlation that was being observed between these two variables merely because both were correlated with a third variable, the proportion of the old.

To address this question in the most general setting, yule-1911 goes on to explain, the researcher must posit a linear relationship between the $n$ variables of interest, $X_1, X_2, X_3, \ldots, X_n$. Such a relationship, Yule, notes

quote... will be of the form \[ X_1 = a + b_2 X_2 + b_3 X_3 + \cdots + b_n X_n. \] If in such a generalized regression or characteristic equation we find a sensible positive value for any one coefficient such as $b_2$, we know that there must be a positive correlation between $X_1$ and $X_2$ that cannot be accounted for by mere correlations of $X_1$ and $X_2$ with $X_3, X_4$, or $X_n$ for the effects of changes in these variables are allowed for in the remaining terms on the right. The magnitude of $b_2$ gives, in fact, the mean change in $X_1$ associated with a unit change in $X_2$ when all the remaining variables are kept constant. yule-1911.

But how do we know that the “magnitude of $b_2$ gives, in fact, the mean change in $X_1$ associated with a unit change in $X_2$ when all the remaining variables are kept constant?” We know this precisely because of the YFWL theorem, which has rigorously established that $b_2$ is the correlation coefficient (normalized by the ratio of their standard deviations) between residualized version of $X_1$ and $X_2$. By residualizing both $X_1$ and $X_2$ with respect to $X_3, X_4, \ldots, X_n$, we have removed from the picture the possible influence these other variables might have had on $X_1$ and $X_2$. Thus, we are now able to measure the correlation between $X_1$ and $X_2$ after removing the possible influence of the other variables out of the way, i.e. holding “all the remaining variables ... constant.” This essential conceptual meaning of `holding variables constant' in a multiple regression rests completely on the YFWL theorem.\footnote{Yule was a leading figure behind the development of regression analysis. The substantive issue he seems to have analyzed in many studies is the effect of the out-relief ratio (proportion of poor provided monetary, food or other assistance but not required to be in an institution, compared to those who were required to enter an institution like a poor house to receive support) on pauperism. In yule-1899, he demonstrated that out-relief is positive correlated with, perhaps causes and increase in, pauperism. Of course, in substantive terms, Yule's regression analysis suffers from simultaneity bias. After all pauperism might also be a cause of administrative measures that change the out-relief ratio.}

My second comment relates to a puzzle in the intellectual history of the YFWL theorem. The puzzle arises from noting that udny-1907's result is essentially the same as the result proved by frisch-waugh-1933, as I have demonstrated. What is puzzling, therefore, is that frisch-waugh-1933 do not cite udny-1907. This omission seems puzzling to me given two fact. First, frisch-waugh-1933 refer to a previous paper by one of the authors: frisch-1929. It is also to be noted that Frisch had published a related work in 1931 frisch-mudgett-1931. What is interesting is that in both these papers, there is explicit reference to the partial correlation formula introduced by Yule.\footnote{For instance, see frisch-1929 and frisch-mudgett-1931} Recall that Yule had introduced a novel notation for representing partial correlations and coefficients in a multiple regression equation in his 1907 paper that I have discussed. The notation was then introduced to a wider readership in statistics with his 1911 book yule-1911. This book was extremely popular and ran several editions, later ones with Kendall yule-kendall-1948. Second, frisch-waugh-1933 use the exact same notation to represent the regression function that Yule had introduced in his 1907 paper. as I have highlighted.

These two facts suggest that Ragnar Frisch was familiar with Yule's work. In fact, Yule's work on partial correlations and multiple regression was the standard approach that was widely used and accepted by statisticians in the early part of the twentieth century griffin-1931. Therefore, it is a puzzle of intellectual history as to why frisch-waugh-1933 did not cite Yule's $ 1907 $ result about the equality of multiple and partial regression coefficients which was, in substantive terms, exactly what they went on to prove in their paper. Of course, frisch-waugh-1933 used a different method of proof from udny-1907, as I have demonstrated in this paper. But in substantive terms, frisch-waugh-1933 proved the same result that Yule had established more than two decades ago. No matter what the reason for the omission in history, it seems that now is the right time to acknowledge Yule's pioneering contribution---both in terms of computation and in terms of interpretation---by attaching his name to a theorem that is so widely used in econometrics.