Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
109,463 characters · 21 sections · 115 citation commands
Multicollinearity, or correlation among regressors, has posed a problem for statistical analysis since the introduction of Ordinary Least Squares (OLS) in legendre. Inflated standard errors depress statistical significance, and negative coefficient covariance renders models sensitive to small changes in regressor selection or functional form. A growing area of concern has been the emergence of p-hacking, the selection of regressors or functional forms to inflate statistical significance leamer, gelman. At the same time, rapid improvements in computing power and the emergence of machine-learning have increased the number of closely related regressors that can be analyzed in a single model baiardi. The statistical instability induced by even modest levels of multicollinearity in the data exacerbates its influence on reported findings, as strategic or arbitrary inclusion of even a single regressor can radically vary the conclusions. All this raises the stakes for discovering an alternative to OLS that will be both rationally interpretable and robust in the face of multicollinearity.
Various statistical treatments for multicollinearity have been developed, each coming at some cost in terms of either coefficient bias or interpretability. Step-wise methods hocking and discretionary elimination of correlated regressors introduce omitted-variable bias. The advantage of retaining all regressors via ridge regression hoerl comes at the cost of systematic downward bias.
Orthogonalization is a popular machine-learning solution for computation, forecasting, and signal processing wherever directly interpretable coefficients are not required despois. laplace introduced his orthogonalization approach to compute legendre’s newly popularized (legendre) least-squares regression, free of gauss's (gauss) computationally more expensive normal equations stigler. Laplace's method was independently discovered by gram, schmidt, and iwasawa, and coined the Modified Gram-Schmidt Process by wong farebrother.\footnote{langou provides an English translation of Laplace's original laplace manuscript.} The Gram-Schmidt is a special case of upper-triangular matrix decomposition, referred to by francis as QR decomposition, which preserves all information in the original data set, a so-called lossless process. A long list of orthogonal models have followed Laplace, notably pearson's (pearson) lossless Principal Component Analysis (PCA) for covariate-importance ranking and golub's (golub) Singular Value Decomposition (SVD) eigenvalue-based method for data compression.
Research interest in multicollinearity has waned in recent decades due in part to a growing consensus that growing data sizes will ameliorate multicollinearity's impact greene and partly from the unreliability of existing diagnostic approaches like pairwise correlation and the Variance Inflation Factor (VIF) kalnins23. In addition, multicollinearity is commonly perceived to be a statistical rather than economic problem enikolopov. Unfortunately, statistical confidence reported for the incorrect sign (Type I error), so-called sign-switching, increases with sample size kalnins18. We will explore how multicollinearity's statistical symptoms may arise from omitting readily available economic information heretofore overlooked -- the temporal order of the regressors.
This paper makes four contributions. First, we derive the finite sample properties of the Gram-Schmidt least-squares model. We show how the model preserves all regressor information and that its coefficients are unbiased and stable, with standard errors lower than in OLS. Second, we explore conditions under which the model returns average treatment effects when treatment responses are heterogeneous. We show that for late treatments - those applied after all other personal characteristics have been determined - the estimate is identical to OLS, and the average treatment effect bias can be estimated and corrected with existing approaches sloczynski. For early treatments, applied before other characteristics and associations are formed, the model returns an unbiased estimate of the average total treatment effect on the treated (ATTT). Third, we extend the model to groups of simultaneous regressors common in economic data sets, expanding the model's use beyond strictly temporally distinct regressors.\footnote{Here, we refer to regressors that are simultaneous to one another rather than to the regressand or error term as are usually considered in simultaneous and endogenous models.} Finally, we apply our model to sloczynski's (sloczynski) decomposition of angrist09's (angrist09) analysis of job program effectiveness, and lubotsky's (lubotsky) assessment of contributors to child reading-and-comprehension outcomes. Both studies controlled for several collinear and temporally-ordered individual characteristics. Our analysis extends earlier findings to the bohren decomposition, which breaks total discrimination into its direct and systemic elements. Our model expands this decomposition, dividing systemic discrimination further into such channel-specific mechanisms as education, income, and household formation.
Our interpretability result is motivated by earlier research in interdependent dynamic systems, first explored in the 1920s by economists interested in describing and forecasting simultaneous supply and demand relations. There, system identification was a primary focus. Early strategies included the instantaneous equilibrium condition, lagged endogenous variables, causal chains, and instrumental variables. schultz introduced the cobweb model, extended by wright25, wright34 to path analysis and instrumentation. Wold generalized these causal chain and dynamic system approaches to the Linear Systems of Equations Model (LSEM) (wold51, wold60), showing in a time-series framework that model coefficients were asymptotically consistent (wold63). goldberger explored systems with unobservable regressors and early latent factor models. alwin first calculated indirect effects and their confidence intervals by using a product-of-coefficients method.
rubin refocused causal identification away from LSEM's lagged variables and instantaneous equilibria toward the random control trial (RCT) strategies inspired by neyman, the so-called Rubin causal model holland, or Potential Outcomes approach imbens20. imbens94 and card extended these strategies greatly, particularly the use of instrumental variables, difference-in-difference, and shape restrictions, frequently considering models with potential endogeneity from omitted variables. The resulting Causal Inference (CI) framework is used widely across academic disciplines.
Parallel to these efforts, pearl00 expanded LSEM to the probabilistic, graph-centric Causal Mediation (CM) framework, of growing popularity in the biological, computer, and social sciences outside economics. pearl00 showed a regressor's total effect to be a partial derivative of an independent variable with respect to the regressor's path through the CM system hunermund, reformulating wright34's (wright34) method of path coefficients and extending wold63's chain principle (wold63) beyond a time series setting. pearl00 recast wold63's conditions in terms of a probabilistic Directed Acyclic Graph (DAG) or Bayesian Network. One-way causality is achieved either when regressors are temporally separated from one another by the natural occurrence of the data or by experimental design, referred to as Sequential Ignorability. pearl01 shows identification in this system is equivalent to the Gauss-Markov assumptions under the additional conditions of independent residual vectors and freedom from heterogeneity induced by unobservable regressors.\footnote{We show later how independent residuals and non-heterogeneity from unobservable regressors follow directly from the Gauss-Markov zero-conditional-mean condition and linearity assumptions, respectively.}
Simultaneous regressors remain unidentified in both the LSEM and CM frameworks.\footnote{wold63 shows that the reduced-form parameters of an interdependent (simultaneous) system are not identified.} We utilize this graphical representation to motivate the linkage between temporal order and direct and total effects and provide properties sufficient for equivalence between the two frameworks. That DAGs follow from the LSEM framework is well documented, and we will not provide further equivalence conditions here.
Temporal order is by no means the only path to separability among regressors. It may be achieved either partially or fully by experimental design imbens18 or awareness of the underlying economic or other known causal mechanisms borusyak. Temporal order is readily observable from data-collection frequency and timing. The passage of time precludes feedback (simultaneity) among regressors in a way that other data-collection or generating processes sometimes lack. In certain cases regressors may be simultaneously determined when measured over a given time scale, in monthly average expenditures and prices for instance, but well-ordered over smaller time scales such as daily transactions with posted prices.\footnote{lewis-beck provide a helpful discussion of the source of simultaneity posed by strotz and developed by fisher and johnston. They suggest that, by nature, regressors tend to arise sporadically rather than all-at-once. Even when feedback occurs between two particular regressors, it is usually by way of a succession of stimuli and responses. In brief, regressor simultaneity arises in naturally occurring data whenever its collection intervals span regressor creation and feedback response activities.} Finally, temporal order removes multicollinearity but does not preclude endogeneity resulting from unobserved regressors, a phenomenon explored by imbens94. The relationship between temporal and unobserved regressors is being explored in a companion paper.
Linearity will be an important motivating assumption throughout this paper. It is restrictive and no longer explicitly imposed by many relevant CM and CI models hunermund, pearl01. Despite this, linearity provides a tractable approach to temporal order and an accessible alternative to OLS, which remains a persistent tool for applied research imbens15 with growing interest from machine-learning CM approaches kumor.
This paper then proceeds as follows. The next section introduces the Gram-Schmidt process for a simplified data set of two recursive regressors. Section 3 reviews the LSEM and recursive DAG frameworks and provides conditions under which Gram-Schmidt coefficients are equivalent in expected value to a recursive LSEM representation. Section 4 derives finite-sample estimation properties for our proposed model. Section 5 relaxes homogeneity and provides the conditions for decomposing the Gram-Schmidt coefficients into average total treatment effects for both early and late-occurring treatments. Section 6 extends the method to a mixed data set containing simultaneous as well as recursive regressors and derives the extended estimation properties. Section 7 illustrates the new model by replicating and expanding the work of angrist09 and lubotsky, who consider the effects of collinear and temporally-ordered family and individual characteristics on earnings and childhood reading and comprehension. Section 9 concludes.
The Gram-Schmidt process begins with a data set $X$ consisting of $K$ real-valued variables arranged in random order. The second variable is regressed on the first and replaced with the resulting residual; then, the regression coefficient is saved. This process is repeated, each variable regressed successively on all prior and then replaced by the resulting residual, saving the coefficients. The outcome is an upper-triangular matrix of estimation coefficients $C$ and a new variable matrix $U$ constituting the orthogonal basis of the original data set, essentially preserving all variable information but in orthogonal form. The final step is to divide each residual by its standard deviation, creating the regressor set's orthonormal basis. The latter step will be excluded below, although it is sometimes useful, and such coefficients can be interpreted in terms of standard deviation units.
The process can be represented by a set of line-by-line OLS regression equations, with each residual replaced by the prior equation's residual. The simple recursive system represented in Figure 1(a) can be represented as follows: \newline {0pt} {0pt}
where $c_{xd}$, $c_{xy}$, and $c_{dy}$ are estimated coefficients, $x$, $d$, and $y$ are data vectors, and $u_x$, $u_y$, and $u_d$ the residuals. The coefficients can be specified as the inner-product ratios $c_{xd} = x'd / x'x$, where $'$ is the transpose operator.\footnote{For comparability with OLS, the algorithm is presented here in the inner-product form suggested by longley, under which Laplace's computational accuracy remains identical to that of the modified Gram-Schmidt. See farebrother.}
In matrix form, the system including $y$ can be written as $X=U(C+I)$, where the $(K+1) \times N$ matrix $X = [x \; d \; y]$ is decomposed into three components: (i) the $(K+1) \times N$ matrix $U$, containing orthogonal residuals vectors $[u_x \; u_d \; u_y]$; (ii) the identity matrix $I$; and (iii) the upper-triangular $K \times K$ matrix {10pt} {0pt}
This procedure produces a convenient orthogonal data set $U$. But how can coefficient matrix $C$ be interpreted? To answer this, we next explore the LSEM and recursive DAG models.
DAGs are one approach to formulating a structural system of equations, deriving the reduced form, checking identification, and recovering the parameters. Three types of parameters are recovered in such systems: (i) direct effects, where a regressor acts, as is the case with OLS, directly on the dependent variable, holding other regressors constant; (ii) indirect effects, when a first regressor acts indirectly on the dependent variable by influencing a second regressor; and (iii) total effects, the sum of direct and indirect effects.
To illustrate, consider the linear structural equation representation for the recursive example in Figure 1(a), where $x$ and $d$ are centered and ordered regressors; $y$ is the independent variable; $\beta_{ij}$, $i,j=x,d,y$, are the unobservable population parameters; and $\upsilon_i$ is the residual:
Here, variables $x$ and $d$ are recursive because they occur one at a time in causal order by way of a temporal or experimentally designed separation. Specifically, $x$ is determined before $d$ and then affects $d$ directly through parameter $\beta_{xd}$. Regressor $x$ is also determined before y, acting upon $y$ directly through $\beta_{xy}$. Finally, $x$ exerts an indirect effect on $y$ by way of its influence on $d$, in turn influencing $y$ through $\beta_{dy}$. Together, the total effect of $x$ on $y$ is the sum of its direct and indirect effects, i.e., $(\beta_{xy} + \beta_{xd}\beta_{dy})$. By the chain rule, this total effect is also the partial derivative of $y$ with respect to $x$.
We can now assemble the total effects of the reduced form system, where residuals $\upsilon_i$, $i = x, d, y$ now serve as regressors:
Note that parameter $\beta_{dy}$ is both the direct and total effect of $d$ on $y$, given that $d$ is the last regressor to be determined and influences no other regressors. This can also be shown by the Frisch-Waugh-Lovell decomposition theorem frisch, lovell, first described by yule, that $\beta_{dy}$ is recoverable from the regression of the independent variable $y$ on residual $\upsilon_d$ and regressor $d$.
The source of multicollinearity in this recursive system is that each regressor influences those determined later. For instance, when regressors are standardized and no spurious or incidental multicollinearity is present, the covariance between $x$ and $d$ is exactly the direct-effect parameter $\beta_{xd}$.
For convenience, the entire reduced-form system, including the dependent variable, can also be represented as a matrix decomposition $X = U(A+I)$, with upper-triangular parameter matrix:
Matrix $A$ also represents the matrix of first-order partial derivatives because $\tfrac{\partial y} {\partial x} = \beta_{xy} + \beta_{xd} \beta_{dy}$. The parameter matrix $A$ appears similar to the Gram-Schmidt coefficient matrix $C$, which we now formalize.
Under the following assumptions, an equivalence between Gram-Schmidt coefficients $c_{ij} \in C$ in ((ref)), $i,j=1,...,K$, and the reduced-form LSEM parameters $a_{ij} \in A$ in ((ref)) can be shown:
We will denote the sample residual vectors as $u_i \subseteq U$, $i=1,...,K$.
Theorem 1 -- Equivalence. Under assumptions (a) - (c), there exists a unique Gram-Schmidt coefficient matrix $C\subseteq \mathcal{R}^{[K \times K]}$ equivalent in expectation to the recursive LSEM reduced-form matrix $A\subseteq \mathcal{R}^{[K \times K]}$
Proof. See online \hyperref[appn]{Appendix}.
In the next theorem, we will compare the finite sample properties of Gram-Schmidt coefficients $c_{ij}$ in equations ((ref)) - ((ref)) with those estimated by OLS $b_{ij}$. Define: (i) $X_{-j}$ as the full regressor set $X$ excluding regressor $x_j$; (ii) $X_{+k}$ as the regressor set including the additional variable $x_k$; (iii) $X_{<j}$ as the regressor set including ordered regressors $x_i$ up to, but not including, $x_j$, $i=1,...,j-1$; and (iv) $R_{jX_{<j}}^2$ as the coefficient of determination of the regression of $x_j$ on the set of remaining regressors $X_{<j}$. We will assume the regressor order is $i < k < j$ throughout, in which $x_j$ will be the $j^\text{th}$ equation's dependent variable. We make two additional assumptions:\footnote{These complete the classic assumptions of the Gauss-Markov theorem, with homogeneity imposed by the scalar-value parameters in the structural equations ((ref)) - ((ref)).}
Theorem 2 -- Properties. Under assumptions (a) - (e), the Gram-Schmidt system $X=U(C+I)$ obtains the following estimation properties:
Proof. See online \hyperref[appn]{Appendix}. Confidence intervals of the total effects and model inference are obtained directly from the terminal Gram-Schmidt regression, eliminating the need to reconstruct indirect effects by way of either product-of-coefficients alwin or simulation methods.\footnote{imai10a propose what they refer to as a nonparametric indirect-effect estimator, calculated by forecasting the independent variable with and without the regressor interaction of interest. Their model relies on parametric estimators of the structural equations. They show their estimator is consistent when all underlying estimators are also consistent. When regressors are present, model errors lack an analytic asymptotic distribution, so confidence intervals must be simulated imai10b.}
To explore the causal interpretation of the Gram-Schmidt coefficients, we relax the homogeneity assumption implicit in the structural model ((ref)) - ((ref)), allowing individuals to exhibit heterogeneous responses to the treatment. We will examine two cases: the classic (late) treatment case in which a treatment $d$ is administered to a group of individuals with predetermined characteristics $x$, and an alternative (early) treatment case in which the treatment is one such characteristic $x$ or a treatment administered before the formation of such characteristics. The early and late cases are respectively represented by treatment variables $x$ and $d$ in Figure 1(a).
In the late-treatment case $d$, we will show that both the convex combination and causal interpretations of the Gram-Schmidt are identical to those of OLS, introduced by angrist98 and extended by sloczynski. In the early-treatment case $x$ by contrast, the Gram-Schmidt approach provides the unbiased average total treatment effect (ATTE). The average total treatment effect on the treated (ATTT) is identical to that in the untreated (ATTU). This makes intuitive sense because early-treatment total effect $c_x$ includes $x$'s direct effect $\beta_{xy}$ as well as its intermediate influence on the formation of later personal characteristics $\beta_{xd} \beta_{dy}$ - which themselves may include associations with various programs and institutions. In effect, if a control group member were to be moved to the early treatment group, such as being black, all regressor characteristics would follow as if they had been a member of the original treatment group.
Two existence and uniqueness assumptions are required for the convex combination interpretation:
Here $p(z)$ is the probability of treatment, defined as the projection $\hat{z}$ from a regression of the treatment variable $z$ on the predetermined characteristics $m$, as shown in equation ((ref)) for variables $d$ on $x$.
It is also useful to define the projection of $y$ on the probability of treatment $p(z)$, given a treatment group $j=0,1$:
{0pt} {0pt}
The average partial linear effect (APLE) of the treatment on group $j$ is then
A causal interpretation also requires ignorability of the mean and linear probability of a binary treatment variable:
We can now state the interpretation of the late-treatment Gram-Schmidt (LTGS).
Lemma 1 -- Late-treatment Gram-Schmidt is identical to OLS.
Under assumptions (a) - (h) and a heterogeneous treatment response:
Here, $\omega_j$ are the convex, variance-weighted treatment proportions defined in sloczynski.
Proof. (A) The late-treatment Gram-Schmidt coefficient is identical to OLS by the Frisch-Waugh-Lovell decomposition theorem frisch, lovell, allowing us to apply the result from sloczynski's Theorem 1 and Corollary 1.
Lemma 2 -- Early-treatment Gram-Schmidt (ETGS) is the ATTE.
Under assumptions (a) - (h) and a heterogeneous treatment response:
Proof. See online \hyperref[appn]{Appendix}.
Causal identification is also challenged by endogeneity from unobserved regressors, as explored by imbens94 and many others. Gram-Schmidt properties under omitted variables, included-irrelevant variables, and late-treatment variables all have interesting implications for unobserved regressors.
The Gram-Schmidt transformation is equivalent to the recursive DAG and LSEM when every regressor is recursive, that is, strictly separated by time or when the experimental design is such that no feedback can occur between regressors. However, the latter is an unlikely condition in practice because regressor simultaneity is present in many naturally occurring -- and even some experimentally designed -- data sets, especially when data collection intervals are wide. In the present section, we extend the Gram-Schmidt process to allow a block of simultaneous regressors to replace a single recursive regressor. This preserves temporal ordering but avoids endogeneity bias because the simultaneous regressors within a block are not regressed on one another. Instead, we construct a block-upper-triangular system of coefficients, the final column representing the total effects on the dependent variable from the relevant regressor. Coefficients thus remain consistent with the dependent variable's partial derivatives with respect to the relevant regressor's path-specific effect described by wright34 and pearl00, excluding feedback effects. By omitting only the mutual regressions of the simultaneous regressors, we allow the recovered total effects to exclude only direct feedback effects among simultaneous regressors. Indirect effects of such feedback that may progress through the system are preserved in the total effects of earlier-occurring regressors.
It is worthwhile first to explore the identification problem raised by these simultaneous regressors. It is well known that simultaneous equations, such as supply and demand systems, are not directly identifiable, so parameters cannot be immediately recovered. The same is true for simultaneous regressors in structural and DAG models. Consider then a simplified system with two simultaneous regressors $s$ and $x$: {10pt} {0pt}
This system is unidentified because neither the simultaneous coefficients nor the residual vectors are observable.
Partial derivatives can, however, be obtained by way of the Implicit Function Theorem:
Unfortunately, the partial derivatives here are also functions of the unrecoverable parameters $\beta_{sx}$ and $\beta_{xs}$. We will refer to these unrecoverable parameters as feedback effects. Direct effects $\beta_{iy}$ are however fully recoverable because OLS remains unbiased, $E[b_{iy}] = \beta_{iy}$. We will also find a way to recover some feedback information from earlier-occurring regressors.
Multicollinearity is assured by the simultaneity of $x$ and $s$, illustrated by the off-diagonal elements of the variance-covariance matrix $\Sigma_s$ of the standardized regressors, shown here without the influence of additional spurious or incidental multicollinearity:
In the next section, we explore the information that can be recovered when a system contains both simultaneous and recursive regressors.
Consider next a system containing both an earlier- and later-determined simultaneous regressor block as illustrated in Figures 1(b) and 1(c), respectively. We will motivate the problem for the earlier-determined simultaneous block case, Fig 1(b), although the properties of the extended method shown here apply to both cases.
The structural equations assumed in Fig 1(b) are:
As in our simultaneity example in subsection 5.1 above, $s$ and $x$ here are simultaneous. But $d$ has replaced $y$ as the third regressor in the system, so dependent variable $y$ is now a function of three regressors. We know direct feedback effects are unrecoverable, while direct effects $\beta_{id}$ and $\beta_{iy}$ are recoverable, though OLS will suffer from steep multicollinearity. This can be seen in the covariance between $s$ and $d$, illustrated in the following equation for standardized regressors with no additional spurious multicollinearity:
As explored next, it will be possible to recover more information than direct effects alone and continue to remove the recursive portion of the multicollinearity.
Our first step is to derive the reduced-form equations in terms of the recursive residual $\upsilon_d$ -- along with simultaneous regressors $x$ and $s$ -- rather than in terms of the residuals as was the case for $x$ and $d$ in the original Gram-Schmidt model ({(ref)})-((ref)):
This system can be specified as $X = U(A+I)$, where $U = [s \: x \: \upsilon_d \upsilon_y]$ and parameter matrix $A$ -- extending ((ref)) -- is now the block-upper-triangular partial derivatives matrix:
The extended method is then estimated as:
Block 1 consists of the early simultaneous regressors $x$ and $s$, as in Figure 1(b), while blocks 2 and 3 contain only one dependent variable, similar to the recursive case in Figure 1(a). Block 1 feedback coefficients $b_{sx}$ and $b_{xs}$ are omitted, so $x$ and $s$ are not regressed on one another. However, each dependent variable in blocks 2 and 3 is regressed on all regressors in the blocks previous to it so the indirect effects $b_{sd}$ and $b_{xd}$ are recovered.
Residual regressor matrix $U$ is no longer completely orthogonal. Rather, our extended method removes all covariance between any two blocks, in the present case between the simultaneous and recursive blocks, such that (i) the off-diagonal elements of the mixed variance-covariance partitioned matrix $\Sigma_{sr}$ are zero, for instance $Cov[s,u_d] = 0$; but (ii) the simultaneous covariance matrix $\Sigma_s$ from equation ((ref)) remains in the upper left: \[ \Sigma_{sr} = \left[
\right], \] where $\makebox{\text{\large\bfseries 0}}$ is a conforming ($2 \times 1$) zeros vector.
The extended method works similarly when a simultaneous regressor block follows a recursive regressor, as in Figure 1(c):
Block 2 now contains the two simultaneous regressors $s$ and $d$ in Figure 1(c) that will not be regressed on one another, although each will be regressed on $x$ in block 1. Covariances between pairs of blocks are again removed in the mixed variance-covariance partitioned matrix $\Sigma_{rs}$. For instance $Cov[u_s,u_d] = 0$ and the simultaneous covariance matrix $\Sigma_s$ shifts to the lower right: \[ \Sigma_{rs} = \left[
\right]. \]
We can now state the properties of our extended Gram-Schmidt least-squares (GSLS) method, which achieves the Theorem 2 properties across blocks rather than across individual regressors. Multicollinearity is eliminated from the recursive regressors, and recursive total effects are recovered. Multicollinearity among simultaneous regressors in a given block will be reduced but not eliminated. Feedback effects recovered, but the direct and intermediate effects are included in the simultaneous regressors' total effects.
To show this, we revise the orthogonality assumption (a) in Theorems 1 and 2. Let $X \subseteq \mathcal{R}^{[N \times K]}$ be a full-rank matrix of recursively ordered regressor blocks $m,n,h=1,...M$, each block $m$ containing one or more regressors with coefficient $c_{i(m)j(n)}$, and block $m$ determined prior to block $n$, $m<n$. The true underlying model in assumption (d) is now represented by block structural equations ((ref))-((ref)). Throughout, we assume the regressor order to be $i < k < j$ and the block order to be $m < h < n$, with $x_{j(n)}$ the $j^\text{th}$ equation's dependent variable.
With this revision in mind, we can summarize the properties of the GSLS estimator.
Theorem 3 -- Properties. Under our revised assumptions (a) - (e) above, the GSLS system $X=U(C+I)$ obtains the following estimation properties:
Proof. See online \hyperref[appn]{Appendix}.
We now compare the GSLS estimator to OLS using two empirical examples. Readers may replicate the analysis here with data and software packages available for R and stata cross.
First, we extend sloczynski's replication of lalond86's (lalond86) and angrist09's (angrist09) analyses (hereafter SLA) of the National Supported Work (NSW) training program's impact on future earnings. Our late-treatment effect matches the impact of the program reported by SLA. We also look at the early-total treatment effect of being black on both inclusion in the work program and future earnings. GSLS extends the discrimination-decomposition framework of bohren to a regression context by decomposing the total-effect estimate into a direct and a systemic discrimination component. We will find that black individuals were included in the program at higher rates and experienced lower earnings than non-black individuals, a phenomenon linked to both direct and systemic discrimination.
Prior studies controlled for several strongly interrelated individual characteristics drawn from the Current Population Survey (CPS) and the Panel Study of Income Dynamics (PSID). NSW participants were randomly assigned a training program and a control group. Specific job assignments were however assigned locally with candidate response rates varying across demographic groups. This resulted in an over-representation of black and economically disadvantaged individuals in the treatment group sloczynski. To test the extent to which black individuals were treated differently, we aggregate Hispanic and white participants to create a binary black and non-black treatment variable. This will serve as an early treatment because individuals were black before other study characteristics were formed and inclusion in the job program was assigned.
Table (ref) summarizes the NSW data for the 185 job program participants and the 15,992 control group individuals. Several regressor means and standard deviations differ materially between the treatment and control groups, suggesting job assignment completions were not random. In the Causal decomposition subsection, we test for the influence of these sample weightings.
Table (ref) compares, for selected regressors, the OLS direct effects and GSLS total effects and standard errors. In both interpretation and expected value the GSLS coefficients differ from OLS. The latter remains unbiased, $E[b_{xy}] = \beta_{xy}$, in the presence of multicollinearity, and its estimates of structural equation ((ref)) represent direct effects. Coefficients are interpreted in the familiar way, namely as an increase or decrease in $y$ resulting from a unit increase in $x$, ceteris paribus, assuming all other regressors are fixed or independent. In linear models, this is the partial derivative of $y$ with respect to $x$, ignoring the regressor interrelatedness that induces multicollinearity.
GSLS estimates provide the total effect, namely the total unit increase (decrease) in $y$ induced by a one-unit increase in $x$, including any indirect effects on $y$ induced or caused by changes in the remaining $x$-dependent regressors. The latter, by the chain rule, is the partial derivative of $y$ with respect to $x$ given all structural relationships in equations ((ref)) - ((ref)), what wright34 refers to as the path coefficient or pearl01 as path-switching (pearl01) or the path-specific effect (pearl00). GSLS standard errors are lower by the removal of multicollinearity.
The job program's estimated direct effect in the Table (ref) OLS model is identical to its total effect in the GSLS model because, as predicted by Theorem 2, this program is the final regressor in the model. The program's coefficient suggests program participation reduced annual earnings by \$3.47 thousand, slightly larger than the \$3.44 thousand found by sloczynski, the difference attributable to our aggregation of white and Hispanic participants.
To understand the rise in GSLS Black coefficient magnitude, we look to the discrimination decomposition framework of bohren. They define the total expected discrimination function $\Delta(y^0)$ as the sum of direct and systemic discrimination: {10pt} {0pt}
where $b$ and $w$ are groups and $y^0$ is the unobservable initial qualification. The direct and systemic components here can be expressed in terms of expected values: {10pt} {0pt}
Here, $A()$ is the action function and in our case the earnings function; $S_i$ is the set of signals for individual $i$, in our case the individual's race and personal characteristics; $G_i$ is the set of groups ${w,b}$; and $Y^0_i$ is the set of initial qualifications, which we assume to be equal across all participants in the program since, for the early-treatment, it represents their potential for earnings before they are born.
In words, direct discrimination in ((ref)) is the difference in action A, earnings, received by individual $i$, were they to belong to group $w$ rather than group $b$, holding the individual's other signals (background characteristics) $S_i$ and initial qualification $y^0$ constant at the group-$w$ level. This matches the interpretation of OLS results shown in Table (ref). The direct discrimination estimate suggests earnings are now lower by a significant \$2.23 thousand dollars per year.
Systemic discrimination in ((ref)) is the expected difference between the earnings of an individual from group $b$ but with background characteristics typical of group $w$, and the same individual's score given group $b$ background characteristics. It represents the cumulative effects of differences in access to such resources as education and employment, which bohren categorizes as technological sources. Our total discrimination estimate of \$3.74 thousand less per year is significantly greater than the direct effect magnitude, suggesting systemic discrimination is present.
The top panel of Table (ref) reports the regression coefficient and its decomposition into the heterogeneous treatment effect on the treated, untreated, and average participants. The probability of being treated $P(d=1)$ and the variance-adjusted statistical weight $\omega_1$ are also listed along with the treatment and control sample sizes and standard errors.
For comparison, the bottom panel of Table (ref) provides rebalanced estimates from covariate matching along with rebalanced treatment and control sample sizes. Matching methods rebalance data by selecting a second, untreated individual with characteristics similar to those of each treated individual in the sample. This method selects the pair minimizing the Mahalanobis distance between regressor means and variances in the two groups dehejia.\footnote{As desired in the rebalanced model, matched standardized differences (differences in means) are close to zero; and the matched variance ratios are near unity, suggesting that rebalancing reduces the differences between the probability of moments of race, education, and enrollment in the job program despite significant differences between them in the original data. A complete list of rebalanced probability moments is available for both studies in the replication code repository.}
In the GSLS late-treatment job program model in Column 3, the heterogeneous response estimates match closely to those in sloczynski, a result of the small magnitude differences resulting from aggregating white and Hispanic participants. The regression coefficient is not statistically different from the heterogeneous ATTT estimate, consistent with the potentially lower regression bias following from the lower probability of treatment $P(d-1)=$ 0.01 and a statistical weight $\omega_1$ close to unity. The smaller ATTE magnitude follows from the larger ATTU estimate along with the convexity restriction from Lemma 1. Finally, the magnitude of the matching ATTE is also larger than the matching ATTT, though not subject to the Lemma 1 restriction.
The total effect of being black is provided in the center column of Table (ref). Illustrating Lemma 2, the ATTE, ATTT, and ATTU and their standard errors are identical to one another. This can be seen here from the larger total-effect coefficient magnitude, which includes the systemic impacts of being black on such later-formed personal characteristics as grade level, degree completion, and inclusion in the jobs program. Results from rebalancing are similar to the GSLS regression coefficient but with slightly higher standard errors.
The decomposition of the direct effect of being black - in the leftmost data column of Table (ref) - follows a pattern similar to the decomposition in Column 3 but with a much smaller and statistically nonsignificant ATT. This is again partly due to the low probability of treatment $P(d=1)=$ 0.08 and a greater ATU. The magnitude of the matching ATT is also smaller than the matching ATE but by less than the heterogeneity estimate since matching is not subject to the convexity restriction.
Finally, the center column of Table (ref) provides the heterogeneous decomposition of the GSLS early-treatment effect of being black. As suggested by Lemma 2, the GSLS coefficient and heterogeneity ATTE estimate match, and the ATTE is equal to the ATTT. Covariate matching estimates are both within a standard deviation of the regression coefficient.
Our second illustration examines data from the bls National Longitudinal Survey of Youth (NLSY) used by several researchers korenman, blau, lubotsky to explore the role of parental income on child reading comprehension test scores while controlling for several strongly interrelated individual and household characteristics. All three studies cite multicollinearity as a leading motivation of their model design, regressor selection, and interpretation of results.
The NLSY began as a survey of 6,283 women and 6,403 men aged 14 to 21 in 1979 and includes a wide range of characteristics, including employment, income, drug use, marriage, education, cognitive assessments, and childbirth. The sample was rebalanced in 1984 to exclude women serving in the military and again in 1990 and 1991 to correct for the survey's original over-weighting of low-income white women. We will test for any remaining sample selection bias at the end of this section. Sampling frequency in the survey was annual through 1994 and biennial thereafter. A second survey, the Child Supplement, was initiated in 1986 to study the children of women included in the 1979 NLSY cohort, biennially recording cognitive achievement and behavior.
Childhood cognitive development is drawn from the Peabody Reading Comprehension Test. From 1986 to 2014, we include all complete observations of children aged 6-14 who were attempting the test for the first time. We exclude children of mothers dropped from the survey during the 1984, 1990, and 1991 revisions. The final sample includes, from the original NLSY cohort, 6,550 children born to 3,181 mothers.
Table (ref) summarizes the NLSY data by child, broken out into above- and below-median family income levels with 3,275 observations in each category. Several regressor means and standard deviations differ materially between the high- and low-income groups, suggesting the 1984, 1990, and 1991 NLSY rebalancing efforts were insufficient or did not target income specifically. In the Causal decomposition subsection, we test for the impacts of any remaining influence of sample selection bias.
Consistent with all three prior studies of these data, we include annual fixed effects and recursive regressors recorded on the date of test: child's gender and age; logs of family size and income, and mother's race, age, education, and performance on the 1980 Armed Forces Qualification Test (AFQT), a general knowledge exam administered to all participants in the original NLSY cohort. We also control for spousal presence in the household and spousal age and education to compare results with lubotsky. Three simultaneous blocks are specified: (i) annual fixed effects; (ii) mother's race, consisting of a nonwhite indicator variable for black or Hispanic (white the omitted variable); and (iii) spousal characteristics, including spouse's residence in the home. All factors are specified in the natural temporal order in which they are assumed to have been determined. For instance, the child's race is defined in the Child Supplement to be the mother's race originally reported in the NLSY. Race was predetermined at the time of the mother's conception, so precedes the mother's age. In turn, both the mother's race and age were determined before her highest grade level was achieved or her AFQT score was recorded. Spousal characteristics, child age and gender, and family size and income are all recorded when the child attempts the reading and comprehension test for the first time, so they may influence the child's test score, but cannot be influenced by it.
Table (ref) compares, for selected regressors, the OLS direct effects, GSLS total effects, and standard errors. The 2.31 family income coefficient implies a one-percent income rise boosts relative reading comprehension by 2.31 percentile points and is comparable to the range reported by blau and lubotsky.
Consistent with blau, the OLS estimate suggests no significant direct influence of maternal race on child test scores, indicated in the left column of Table (ref).\footnote{Race estimates were not reported in lubotsky.} This would be expected if test administration and scoring mechanisms did not differ systematically between racial groups. The GSLS model shows race's total effect reduced reading scores by 10.83 percentile points among nonwhite children, relative to white children, significant at the 99.9% confidence level. The GSLS nonwhite coefficient standard error is lower than OLS by 19%, illustrating the impacts of removing multicollinearity.\footnote{For a given GSLS regressor, the standard error reduction falls below that indicated by VIF, as the former identifies and removes each regressor's contribution to total model multicollinearity, while the VIF attributes all model multicollinearity to each regressor successively.}
Systemic discrimination is represented by the negative and significant GSLS total effect for race, significantly greater in magnitude than the direct effect. We can use GSLS to further decompose the systemic effect into its channel-specific mechanisms. Table (ref) provides the intermediate-stage GSLS regression results for selected regressors. Race's total effect on education can be seen in the second estimates column showing the regression of maternal school-grade completion on maternal race and age. Here, the total race effect on education was 1.11 fewer grades completed on average by nonwhite mothers than white mothers. In turn, each additional grade completed by the mother raised the child's test score by 1.80 percentile points, as indicated in the last column of Table (ref). Thus, we have a 2.0 percentile point (1.11 $\times$ 1.80) differential between nonwhite and white children attributable to maternal education. This indirect education channel constituted approximately 18% of the 10.8-percentile total race effect.
Table (ref) tests for heterogeneity and sample weighting bias. The GSLS late-treatment income effect in Column 3 shows an ATTE estimate of 1.91, lower than the regression coefficient of 2.31. This downward bias results from the higher probability of treatment $P(d=1)=$ 0.50 and a similar variance-adjusted statistical weight $\omega_1$. The high ATTT is offset by a nonsignificant ATTU close to zero.
The OLS early-treatment nonwhite effect estimates are nonsignificant for every heterogeneous response and covariate matching estimate, suggesting there is no evidence of direct discrimination in test administrations.
The GSLS early-treatment coefficient in the center column again matches the heterogeneous responses exactly. The rebalanced sample matching estimates are - though slightly greater in magnitude - not statistically significantly different from the regression coefficient.
Multicollinearity has posed a statistical challenge for over 200 years. We have shown that when coefficients are linearly separable through temporal order or experimental design, the Gram-Schmidt process removes multicollinearity and produces economically interpretable estimates, equivalent in expected value to total effects from a recursive Linear System of Equations Model or Directed Acyclic Graph. Total effects and statistical inference follow directly from the regression, eliminating the need for ex-post simulation or product-of-coefficients methods. The coefficients are unbiased estimates of the partial derivatives, invariant to the omission of later-determined regressors, with standard errors lower than in OLS. For early treatments, Gram-Schmidt returns the average total treatment effect on the treated even when the treatment response is heterogeneous.
We have exploited these properties to extend the Gram-Schmidt approach to recover causal effects from a previously unidentified mixed system of recursive and simultaneous regressors. The extended approach removes multicollinearity between recursive and simultaneous regressors, allowing unbiased estimates of the total effects that are free from multicollinearity.
Using a mix of temporally ordered and simultaneous individual characteristics, we have illustrated extended Gram-Schmidt least squares by applying it to previous studies using the National Supported Work Program and National Longitudinal Survey of Youth. GSLS total effects tended to be greater in magnitude than OLS direct effects, especially among early-determined or inter-related regressors. The approach expanded our interpretation of discrimination impacts reported in earlier studies and lowered standard errors by removing multicollinearity.