Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
69,409 characters · 17 sections · 67 citation commands
Machine learning the first stage in 2SLS: Practical guidance from bias decomposition and simulation
\onehalfspacing
Machine learning (ML) methods appear in diverse empirical econometric applications. Despite this enthusiasm, the literature has little to say about the practical implications of combining ML methods with two-stage least squares (2SLS). In this paper we provide practical guidance on the potential risks and benefits of inserting standard ML methods into the first stage of 2SLS---drawing insights from a decomposition of potential sources of bias and a series of Monte Carlo simulations. The decomposition shows that ML-in-2SLS overlaps with the canonical forbidden regression but also reveals additional potential sources of bias. Our simulation results illustrate this potential is legitimate: some non-linear ML approaches generate larger second-stage bias than the original endogenous regression.
The motivation behind integrating machine learning in two-stage least squares is clear: to the extent incorporating “better" first-stage predictions is possible, researchers can obtain more precise second-stage estimates.\footnote{The integration of 2SLS and ML has already appeared in applied work across many fields. Early examples have used linear models in the fist stage---predictors that typically perform well in our decomposition and simulation. For example, estimating labor market impacts of imprisonment MuellerSmith2015, the effects of racial-composition shocks during the Great Migration Derenoncourt2019, the effect of expropriation on growth Chen2020, the “true” size of China's GDP growth Chen2019, the inter-generational transmission of health Bevis2020, and the heterogeneous impacts of family size and parental labor supply Biewen2020.} Because most ML methods are built explicitly for prediction---they typically outperform ordinary-least squares (OLS) at this task---using ML for first-stage predictions seems quite natural. Mullainathan2017 likewise highlight the reasoning that might lead a practitioner to inject ML into the 2SLS framework: “Machine learning... revolves around prediction" and “belongs in the part of the toolbox marked $\hat{y}$ rather than in the more familiar $\hat{\beta}$ compartment." The authors immediately recognize that “the first stage of a linear instrumental variables regression is effectively prediction."\footnote{Similar intuition is offered in Belloni2011, Belloni2012, Belloni2013, Chernozhukov2015, Chernozhukov2018, Singh2019, Angrist2020, Singh2020, Chen2020-IV inclusive of new artificial intelligence (AI) methods Hartford2017,Bennett2020, Liu2020. In a recent working paper, Chen2020-IV also recognizes this motivation, suggesting that the traditional OLS-based implementation of 2SLS “leaves on the table some variation provided by the instruments that may improve precision of estimates.” If one is willing to accept the fairly strong assumption that any function (nonlinear or linear) of valid instruments is a valid instrument, Chen2020-IV provides an interesting solution to some of the challenges involved with including ML methods in 2SLS. We do not make this assumption.} The risks of adopting out-of-the-box ML methods for 2SLS-type applications are less clear.
In this paper, we document the risks ML-in-2SLS poses for practitioners of applied econometrics. How does curating and generating first-stage predictions with ML affect the downstream, second-stage causal estimates of two-stage least squares? \footnote{With valid instruments, applying OLS in the first stage of 2SLS produces predictions $(\hat{x})$ that are a linear combination of the exogenous instruments. Thus, $\hat{x}$ is itself exogenous in the traditional 2SLS procedure. Predictions produced by nonlinear functions are not guaranteed to be orthogonal to their residuals, generating additional bias/inconsistency in second-stage estimates.} Using a simple decomposition, we discuss several phenomena that can bias ML-based 2SLS away from its target parameters. Some of these phenomena are implications of the forbidden regression Angrist2001,Angrist2009,Wooldridge2010, which na{\"i}ve implementations of ML in 2SLS are likely to lead to---injecting the predictions of a nonlinear estimator into the first stage of 2SLS.\footnote{Another flavor of the forbidden regression involves applying different specifications of controls in the first and second stages. Most out-of-the-box ML methods do not offer a method to ensure that second-stage controls are used for prediction in the ML-based first stage (and in the correct functional form). There are ad hoc solutions to this problem---writing custom functions that implement the ML algorithm plus a linear specification of the controls/fixed effects, or residualizing (i.e., Frisch-Waugh-Lovell). For an example, see the \href{https://github.com/lrberge/fixest}{fixest} package in \texttt{R} and its \texttt{feNmlm()} function, which is written to efficiently estimate maximum likelihood models with multiple fixed-effect (\textit{i.e.}, large factor variables). This issue is particularly important for situations where conditioning on controls/fixed effects is integral to the instruments' exogeneity. Again, ML methods will, in this way, expose researchers to potential pitfalls.}
Other issues are less common to “traditional" econometrics but become key, we argue, to understanding ML-based results. These include:
In our simulations, most ML-rooted solutions that use common ML procedures in the first stage of 2SLS fail to improve upon standard 2SLS (i.e., using OLS in the first stage) and generate more bias. Two linear estimators are the exception: post-Lasso selection and principal component analysis (PCA). Post-Lasso and PCA perform at least as well as standard OLS-based 2SLS. Perhaps more importantly, we show that highly nonlinear tree-based methods (e.g., random forests and boosted trees) can amplify bias, providing parameter estimates farther from truth than na\"ive OLS regressions that ignore endogeneity. Given sufficient training time, na\"ive implementations of neural networks in 2SLS can reproduce the original OLS bias, with little to no advantage over traditional approaches to recovering exogenous identifying variation through 2SLS.\footnote{This ignores the practical as well. With our resources, the simulation for neural network frequently took several hours to complete which is considerably longer than the time it takes to run a traditional 2SLS.}
In Section (ref) we formalize the theoretical settings and define the estimators. In Section (ref) we introduce two data-generating processes. Because practitioners' use cases differ, we compare two general cases: one simple (and dense) case and a second, more complex and sparse case. In Section (ref) we present the empirical results for the discussed estimators and DGPs. In Section (ref) we also build a theoretical decomposition that explain how and why ML-based 2SLS procedures might increase second-stage bias.
While this paper focuses on a practical, ad hoc approach of inserting ML into the first stage of 2SLS, where ML methods might be thought of as complimenting traditional approaches to causal estimation, a separate strand of the literature offers tailor-made ML-based estimators in 2SLS-like structures. If one is willing to depart from a traditional 2SLS structure and accept different (typically stronger) identifying assumptions, these methods potentially capture more of the available “first-stage” variation without suffering from the issues we highlight in this paper. In Section (ref) we discuss such a solution to the problems inherent to the ad hoc ML-based 2SLS approach. In Section (ref) we offer concluding remarks.
Ultimately we conclude that while ML methods offer many promises for a range of applications, most out-of-the-box ML methods are not well suited for the first stage of two-stage least squares. Moreover, applying the wrong ML method in the first stage can actually generate more bias in parameter estimates than entirely ignoring endogeneity.
Applied researchers commonly apply 2SLS to estimate the causal effect of some $x$ on some $y$ in a setting where the exogeneity of $x$ cannot reasonably be assumed. In other words, where
there is concern over the potential for non-zero covariance between the variable of interest $x$ and the disturbance $u$ when estimating the parameter $\beta_1$.
Let $\mathbf{z}$ denote a vector of instrumental variables. We express the first stage of a 2SLS estimates $x$ as a function of these instruments:
In its traditional OLS-based implementation, $f(\mathbf{z})$ is linear in $\mathbf{z}$.
Defining the predictions from ((ref)) as $\hat{x} = f(\mathbf{z})$, the second stage of the 2SLS procedure then regresses the outcome variable $y$ on $\hat{x}$,
to achieve an estimate for $\beta_1$ in ((ref)). We define $\hat{\gamma}_1$ as this estimate of $\beta_1$. If the instruments are valid (i.e., predictive of $x$ and uncorrelated with $u$) and $\hat{x}$ results from an OLS regression, then $\hat{x}$ will also be exogenous.\footnote{We assume homogeneous treatment effects, which removes the requirement of monotonicity.} The second stage of OLS-implemented 2SLS then generates consistent estimates of $\beta_1$, interpreted as the causal effect of $x$ on $y$.
So why introduce ML? Applications of 2SLS identify the effect of $x$ on $y$ by extracting only a fraction of the “good” (exogenous) variation in $x$. The hope for ML-based 2SLS methods is that researchers can extract more of the good variation in $x$---a more flexible fit of the exogenous variation---while still omitting the bad variation. This desire has likely increased following lee2020, who argue that many traditional evaluations of instrumental variables considerably overestimate their significance.
In the analysis below we examine three classes of 2SLS-motivated estimators:
Class 1: “Traditional" two-stage regression methods: This set of estimators covers the standard two-stage regression estimators in an econometrician's toolbox: two-stage least squares, (unbiased) split-sample IV Angrist1995, the Fuller implementation of limited-information maximum likelihood (LIML) Anderson1949,Fuller1977, and jackknife IV (JIVE) Angrist1999. These methods overlap in three important ways: they (i) employ a two-stage approach (ii) whose first stage creates a linear combination of the instruments (iii) with no formal variable selection.
Class 2: Machine-curated variable selections in standard 2SLS: This second class augments the standard OLS-based version of 2SLS with variable selection/synthesis. Specifically, these methods feature an additional procedure, prior to the first stage, that downselects or combines $\mathbf{z}$ into a more parsimonious set of variables. The elements of this more parsimonious expression of $\mathbf{z}$ then appear in the first stage. The rest of the 2SLS process proceeds as usual (i.e., OLS). Importantly, while these models feature variable selection or synthesis, they also preserve linearity in both stages. Because these estimates result from linear combinations of $\mathbf{z}$, the original exclusion restriction of $\mathbf{z}$ passes through to the selected/synthesized instruments.
Our first machine-curated method is the post-Lasso procedure of Belloni2012, which first estimates the linear relationship between $x$ and $\mathbf{z}$ (a linearized version of (ref)) using penalized regression. This penalized regression minimizes the sum of squared error (SSE) plus a penalty proportional to the sum of the coefficients' magnitudes. That is, $\lambda \times \| \bm{\gamma} \|$, where $\bm{\gamma}$ is the vector of coefficients on the (standardized) instruments and $\lambda$ is a the shrinkage parameter chosen by the researcher (typically via cross validation). Because each instrument's coefficient-based penalty changes discontinuously when moving away from $\gamma_i=0$, Lasso can be used to select a set of stronger instruments (whose coefficients are non-zero). Post-Lasso selects the instruments whose coefficients are non-zero and then estimates standard, OLS-based 2SLS using those selected instruments.\footnote{Angrist2020 notes that this methodology may suffer from potentially unseen pre-test bias. Because our model comes from relatively strong instruments, as with the intuition of Zhao2020, we do not estimate de-biased Lasso models. We therefore allow post-Lasso to serve as a representation of both.}
Principal-component analysis (PCA) offers an alternative route to simplifying $\mathbf{z}$ by selecting $\mathbf{z}$'s first $k$ principal components Pearson1901. Thus, as the second machine-curated method we consider, principal-component analysis (PCA) applied to 2SLS (as in Ng2009 and Winkelried2011). This approach passes a set of principal components into the first stage of standard OLS-based 2SLS. While PCA may reduce the first stage's interpretability, this approach can drastically reduce the number of first-stage instruments while retaining considerable explanatory power.
Class 3: ML-based first stages in 2SLS: Our final class of estimators retains the general two-step framework of 2SLS but replaces the first stage with a variety of cross-validated ML algorithms. We evaluate a meaningful subset of machine-learning methods suitable for regression, including random forest Ho1995,Breiman2001, boosted trees Breiman1997,Mason1999,Friedman2001,Friedman2002, neural networks Turing1948,Mcculloch1943,Farley1954, and Lasso Tibshirani1996,Santosa1986.\footnote{For our purposes, the contributions of Srivastava2014 (dropout), Ioffe2015 (batch normalization), and Kingma2017 (stochastic optimization) are particularly relevant.}\footnote{For a nice review of ML methods in applied economics, including Lasso, tree-based methods, and neural networks, please see Storm2019. For broader and more in-depth coverage (from the authors of many of the methods), see James2013 and Hastie2009.} Notably, most of these algorithms offer considerable flexibility (e.g., nonlinearity in $\mathbf{z}$) and variable selection (to varying degrees). This class offers considerable insights into the merits of off-the-shelf ML methods for machine-assisted 2SLS.
In order to examine the performance of ML in the predictive stage of 2SLS---in absolute terms and relative to “traditional” options---we employ two general data-generating processes (DGPs). For reasons described below we refer to the two DGPs as the low-complexity case and the high-complexity case. In practice, the researcher rarely knows the extent to which her case is complex, particularly in terms of extent of nonlinearity or the efficient number of instruments. While “complexity" is certainly subjective, our intention is to bookend the settings for which an applied researcher might apply ML-based 2SLS.
In this case, we aim to depict the performance of various estimators when the DGP is simple (few instruments and non-sparse) and closely matches the ideal scenario for OLS-based 2SLS: an endogenous regressor that is a linear combination of a relatively small set of strongly predictive, exogenous instruments. This case is applicable to researchers seeking to estimate the causal effect of a variable of interest $x_1$ on outcome $y$,
but facing the challenge (e.g., omitted variables, simultaneity) that $x_1$ is endogenous and $\operatorname{E}\expectarg*{\varepsilon_y \big| x_1}\neq0$ prevents OLS from cleanly identifying $\beta_1$ in ((ref)). Importantly, the causal effect $\beta_1$ is common across all individuals, which ensures differences across estimators are not due to the estimators recovering different local average treatment effects (LATEs).
In this low-complexity scenario, ML-based 2SLS methods are overkill: neither variable selection nor nonlinearity are necessary. In fact, our results demonstrate that ML methods can increase bias relative to 2SLS, even relative to endogenous OLS.
Formally, to model a scenario with a single endogenous regressor $(x_1)$ and a small set of valid (and {\it individually} strong) instruments, we define the DGP as
drawing special attention to the inclusion of $\varepsilon_c$ as the disturbance common to both $x_1$ (the variable of interest) and $x_2$ (the omitted variable). This common error follows a standard normal distribution; $\eta$ is distributed uniformly between $-1$ and $1$.
We assume that a set of valid instruments $\mathbf{z}$ exists such that $\operatorname{E}\expectarg*{\varepsilon_y | \mathbf{z}} = 0$ and $\operatorname{E}\expectarg*{x_1 | \mathbf{z}} \neq 0$ (we focus on the case where $|\mathbf{z}|=7$). We also anticipate that the researcher has no beliefs or insights about the functional form of $g_x(\cdot)$, as is often the case in practice. In the true DGP for this case, $g_{x}(\mathbf{z}) = \sum_{i=1}^{7} z_i$. That is, $g_{x}(\cdot)$ is linear.
In particular, we draw the instruments $\mathbf{z}$ from a multivariate normal distribution centered at zero (i.e., $\operatorname{E}\expectarg*{\mathbf{z}} = \mathbf{0}$) with variance-covariance matrix $\Sigma_\mathbf{z}$ where $\widehat{\text{Cov}}(z_i,\,z_j) = 0.6^{|h-k|}$ (and thus $\text{Var}(z_i) = 1$ for each $i$). By implication, $x_1 \sim N(0,\, \text{Grand Sum}(\Sigma_\mathbf{z}) + 1)$.
In full, then, the data represents the following system of equations:
Notably, the specification of the instruments in this DGP produces a very strong first-stage with a relatively large concentration parameter Belloni2012. Put simply, the concentration parameter $\mu^2$ describes the extent to which the weak-instrument problem may arise within a given DGP. A higher value of $\mu^2$ implies that 2SLS, without variable selection, will converge to the true $\beta_1$ at relatively small sample sizes.\footnote{In this “low-complexit” case, $\mu^2 \approx n\times 20.71$, which exceeds the values in Belloni2012. We discuss $\mu^2$'s role further in the high-complexity case section below.} Consequently, the low-complexity case allows us to test how machine-curated first stages perform when there is little to be gained from variable selection/synthesis.
As our high-complexity case, we follow Belloni2012's DPG with two extensions. This DGP allows the researcher to tailor instruments' strengths with many instruments. Following Belloni2012, the DGP in our high-complexity case results from
where
As before, we imagine the researcher's interest in this case focuses on identifying $\beta_1$. However, unlike the earlier DGP, the high-complexity case produces sets of relevant and exogenous instruments that vary in their correlation and individual strength (i.e., $\pi_i$).
In defining the “exponential” design of the first-stage coefficient vector $\bm{\pi}$, we follow Belloni2012: $\bm{\pi}$ captures a “beta pattern" $\widetilde{\bm{\pi}} = (0.7^{0},\, 0.7^{1},\, 0.7^{2},\, \ \dots,\, 0.7^{99})$ that is then multiplied by a constant $C$, i.e., $\bm{\pi} = C \times \widetilde{\bm{\pi}}.$ The constant $C$ implies a value for the concentration parameter, $\mu^2 = \frac{n \bm{\pi}'\Sigma_z \bm{\pi}}{\sigma^2_v}$.\footnote{For a proof of this statement, see Belloni2012.} In panels (ref)--(ref) (Figure (ref)) we illustrate the three beta patterns that we adopt in the “high-complexity" DGP, generating three subcases of this DGP. As described above, the concentration parameter is useful for determining the behavior of IV estimators. Because we are less interested in the case of weak instruments, we use $\mu^2 = 180$, which creates a strong---though fairly sparse---set of instruments as outlined in Belloni2012.\footnote{It is important to select this value thoughtfully. Choosing a $\mu^2$ that is too small will simulate a weak-instruments problem. Choosing a $\mu^2$ that is too large will yield a scenario in which all instruments are “overpoweringly” valid, which reduces the effectiveness of selection or dimension-reduction techniques. See Hansen2008 for additional discussion of $\mu^2$.}
Belloni2012 arrange the coefficients $\bm{\pi}$ in descending order (i.e., $\pi_1 > \pi_2 > \cdots > \pi_{100}$). However, the definition of $\Sigma_z$ implies that “proximate" instruments are more correlated than “distant” instruments (i.e., $\text{Cor}(z_i,\,z_{i+1}) > \text{Cor}(z_i,\,z_{i+k})$ for $k>1$). Thus, the DGP of Belloni2012 ensures the strongest instruments correlate with each other. While this feature may be desirable in many contexts, we remain agnostic with regard to whether the strongest instruments are most correlated with each other or with other instruments. However, this agnosticism requires that we consider three sub-cases, each arising from alternative orderings of the coefficients in $\bm{\pi}$ and the covariance of the respective variables:
Finally, we define $\sigma^2_v = \bm{\pi}' \Sigma_{z} \bm{\pi}$ (which forces that $\text{Var}(x_1) = 1$) and $\sigma_y = 1$. In panels (ref)-(ref) of Figure (ref) we illustrate the cross-instrument correlations implied by $\Sigma_z$: in Panel (ref) we show a correlation matrix among the 100 instruments, and in Panel (ref) we highlight the correlation of $z_{1}$ and $z_{50}$ to each of the other 100 instruments. Instruments are strongly correlated with their neighbors and weakly correlated with non-neighbors, which limits the information accessible from any single instrument.
We now discuss the results of our simulations. In every simulation, we include an “oracle model” that extracts the entirety of the exogeneous component of $x_1$ (perfectly removing endogeneity) and a simple OLS model (where we entirely ignore endogeneity). While one might expect the oracle and plain OLS models to bookend the biases in 2SLS-related models, our simulations demonstrate that they do not. That is, inserting machine learning into the first stage can lead to outcomes that are even worse than ignoring endogeneity.
In each case, we are interested in the performances of the estimators in terms of their biases and the precision of estimates. Recall that these estimators include three broad classes: (i) traditional methods (OLS-based 2SLS, split-sample IV, LIML, and jackknife IV), (ii) machine-curated 2SLS (variable-selection or -curation via post-Lasso and PCA), and (iii) 2SLS applications with ML-powered predictions in their first stages (i.e., replacing first-stage OLS with either Lasso, boosted trees, random forests, or neural networks).
In Figure (ref) we depict the distributions of point estimates ($\hat{\beta}_1$) for a given method in the givezn DGPs. Panel (ref) illustrates the low-complexity case; panels (ref)--(ref) presents the results for our high-complexity cases. Table (ref) summarizes each of these method-by-DGP combinations with the mean and standard error from each. The target parameter $\beta_1$ equals 1 throughout the simulations (indicated with a thin dashed line). Each distribution results from 1,000 iterations of the simulation.
To those with use-cases that resemble our “low-complexity case,” the simulation results have a clear takeaway: PCA-based 2SLS and post-Lasso perform well and offer very safe choices.\footnote{LIML also performs well, but with slightly larger variance. The Jackknife IV estimator yields very high variance in this low-complexity DGP, as do Neural Networks.} Important for the practitioner: All four nonlinear ML-in-the-first-stage methods (i.e., Lasso, boosted trees, neural networks, and random forests) perform poorly in terms of bias and variance. In fact, random-forest-based 2SLS generates {\it more bias} in $\hat{\beta}_1$ than the OLS estimator that entirely ignores endogeneity---it is possible for an ML-based 2SLS estimator to {\it amplify} bias relative to plain OLS. We discuss the source of this bias amplification in the next section.
In the three high-complexity cases in Table (ref) (columns $B$--$D$) and in panels (ref)--(ref) of Figure (ref), LIML and Jackknife IV generate very little bias in their estimates of $\beta_1$, outperforming 2SLS. Across all three DGPs, 2SLS produces mean estimates roughly 2.3--5.8 percent larger than the true parameter, while the centers of LIML's and JIVE's distributions are within 0.4 percent of the true parameter. Injecting random forests into the first stage, on average, produces more biased estimates than na\"ive (endogenous) OLS, generating coefficient estimates that are 32--56 percent larger than the true estimates. Again, one can worsen endogeneity issues by using ML-based 2SLS estimators.
To diagnose the sources of bias from different methods, we show one can decompose the wedge between $\beta_1$ and $\hat\beta^{\text{2SLS}}_1$ into three components,
where $f$ is non-decreasing with respect to each of its arguments, $\hat{x}$ is the first-stage-based prediction of $x$ from some set of valid instruments, $e$ denotes the resulting first-stage residuals $(x - \hat{x})$, and $u$ represents the population disturbance from regressing $y$ on $x$, i.e., $u = y - (\beta_0 + \beta_1 x)$. Below we derive and elaborate upon ((ref)).
Each component of the wedge offers insights into how first-stage methods differentially produce biases---this delivers helpful intuition regarding the pitfalls that may arise in 2SLS applications that include ML-based first stages.
To see the component parts of the bias drawing $\hat\beta^{\text{2SLS}}_1$ away from $\beta_1$, suppose again that the parameter of interest is $\beta_1$, the causal effect of $x$ on $y$ in
Suppose also that $x$ is endogenous, i.e., $\text{Cov}(x,\,u)\neq 0$. The 2SLS estimate of $\beta_1$ comes from estimating
where again $\hat{x}$ is the first-stage-based prediction of $x$ from some set of valid instruments $\mathbf{z} = z_1,\, z_2,\, \ldots,\, z_p$.
Because we estimate the second stage in ((ref)) via OLS, the estimate for $\beta_1$ can be written
where $\widehat{\text{Cov}}(\cdot)$ and $\widehat{\text{Var}}(\cdot)$ refer to the sample-based covariance and variance.
Using ((ref)) and ((ref)), we can rewrite $w$ as
where $e$ is the first-stage residual, the difference between $x$ and $\hat{x}$.
Substituting ((ref)) for $w$, we can decompose the covariance in ((ref)) into two components:
If the first-stage predictions $(\hat{x})$ come from OLS, then $\widehat{\text{Cov}}(\hat{x},e)$ is mechanically zero. The second term, $\widehat{\text{Cov}}(\hat{x},u)$, is typically small when $\hat{x}$ comes from a linear combination of valid instruments.
Finally, substituting ((ref)) into ((ref)) yields a helpful expression for the 2SLS estimate for $\beta_1$,
Again, first-stage OLS guarantees that $\widehat{\text{Cov}}(\hat{x},e)$ is zero and, with valid instruments, that $\widehat{\text{Cov}}(\hat{x},u)$ is small. Whether $\widehat{\text{Var}}(\hat{x})$ is “small” is typically of little consequence with OLS (as $\beta_1 \widehat{\text{Cov}}(\hat{x},e) + \widehat{\text{Cov}}(\hat{x},u)$ is typically small). However, all three points can generate important issues when we mix ML methods into the first stage of 2SLS. With ML methods, nothing guarantees that $\widehat{\text{Cov}}(\hat{x},e)$ is zero or that $\widehat{\text{Cov}}(\hat{x},u)$ is small. Moreover, many ML methods are constructed to {\it reduce} the variance of predictions, which further amplifies bias. This variance-reduction aspect is particularly relevant for nonlinear methods.
For the term $\beta_1 \mathrm{Cov}(\hat{x},e)$ to differ from zero and generate bias, $\beta_1 \neq 0$ and $\mathrm{Cov}(\hat{x},e)\neq 0$. We assume that the population-regression coefficient $\beta_1$ differs from zero. Consequently, the term $\beta_1 \mathrm{Cov}(\hat{x},e)$ only generates bias when $\mathrm{Cov}(\hat{x},e) \neq 0$; $\beta_1$ scales the bias and affects its direction.
By construction, OLS produces predictions that are orthogonal to their residuals, i.e., $\widehat{\text{Cov}}(\hat{x},e) = 0.$ This first term is therefore irrelevant when the first stage uses OLS. However, when practitioners adopt other methods in the first stage (e.g., non-linear methods) nothing guarantees first-stage predictions are uncorrelated with their residuals. Indeed, this component relates to many researchers' definitions of the forbidden regression. This part of the bias results from using estimators whose predictions correlate with their residuals (rather than resulting from a violation of the exclusion restriction). While nonlinear methods can generate $\widehat{\text{Cov}}(\hat{x},e)=0$, many do not (as illustrated by the column $a$ of Table (ref)).
In addition, because $\widehat{\text{Cov}}(\hat{x},e)$ typically drops out of OLS regression, OLS-based empirical intuition does not help here. One implication from this non-OLS intuition of $\widehat{\text{Cov}}(\hat{x},e)$ is that the bias generated by this component is proportional to the size of the target parameter $\beta_1$. Where treatment effects are larger, the bias transmitted through this component is also larger.
To understand why some methods produce larger values of $\widehat{\text{Cov}}(\hat{x},e)$ than other methods, first decompose this covariance into $\widehat{\text{Cov}}(\hat{x},x)$ and $\widehat{\text{Var}}(\hat{x})$:
While $\widehat{\text{Cov}}(\hat{x},e)$ is not generally signable, it is bounded between $-\widehat{\text{Var}}(\hat{x})$ and $\widehat{\text{Cov}}(\hat{x},x)$.\footnote{We assume predictions, $\hat{x}$, will have non-negative covariance with the true values, $x$.} Further, we can sign $\beta_1\widehat{\text{Cov}}(\hat{x},e)$ in five subcases:\footnote{We assume $e$, $x$, and $\hat{x}$ have variation and that the predictions $\hat{x}$ positively correlate with the true values $x$.}
where $\sigma_x$ refers to the standard deviation of $x$ ($\sigma_{\hat{x}}$ and $\sigma_{e}$ are defined similarly).
As ((ref)) reveals, the sign of $\beta_1 \widehat{\text{Cov}}(\hat{x},e)$ depends on two quantities: (i) the sign of $\beta_1$, and (ii) the sign of $\widehat{\text{Corr}}(\hat{x},x) \, \sigma_{x} - \sigma_{\hat{x}}$. It is difficult to generalize the sign of $\widehat{\text{Cov}}(\hat{x},e)$ without further assumptions. While one may be tempted to assume $\sigma_x > \sigma_{\hat{x}}$, this assumption is not sufficient for signing $\widehat{\text{Cov}}(\hat{x},e)$, as it still depends upon the magnitude of $\widehat{\text{Corr}}(\hat{x},x)$.\footnote{Further, this assumption is equivalent to making an assumption on $\widehat{\text{Cov}}(\hat{x},e)$, which means one is essentially assuming the result. That is, $\widehat{\text{Var}}(x) = \widehat{\text{Var}}(\hat{x}) + \widehat{\text{Var}}(e) + 2~\widehat{\text{Cov}}(\hat{x}, e)$. That said, in every iteration of our simulations, $\widehat{\text{Var}}(x) > \widehat{\text{Var}}(\hat{x})$.} The knife-edge case where $a=0$ appears unlikely except in cases where either $\beta_1=0$ or where $\widehat{\text{Cov}}(\hat{x},e)$ is mechanically zero (e.g., OLS).
Across the twelve models that we consider in Table (ref), only the non-OLS models produce $\widehat{\text{Cov}}(\hat{x},e)\neq 0$---this is unsurprising. Lasso, neural nets, boosted trees, and random forests all produce positive covariance between $\hat{x}$ and $e$. In other words, in all of our DGPs, the term $\widehat{\text{Cov}}(\hat{x},e)$ biases $\hat{\beta}$ upward (positively) whenever it is non-zero.\footnote{This upward bias is partly due to the true parameter $\beta_1$ being positive.} Random forest models generate the largest covariance between $\hat{x}$ and $e$ (and consequently the largest $\widehat{\text{Cov}}(\hat{x},e)$) in each of the DGPs. Depending upon the DGP, Lasso, neural nets, and boosted trees generate the second-highest covariance. Because our shallow subcase of neural nets approximates OLS, its covariance between $\hat{x}$ and $e$ is approximately zero.
One way to ensure that $\widehat{\text{Cov}}(\hat{x},e)=0$ for a nonlinear model is to linearize its output---this is achievable, for example, by using the ML-based prediction $\hat{x}(\mathbf{z})$ as an instrument for $x$, rather than plugging it into the second stage Angrist2001,Chen2020-IV. While this approach forces $\widehat{\text{Cov}}(\hat{x},e)=0$, it requires strengthening assumptions. We discuss this possibility further in Section (ref).
More broadly, the component of bias due to covariance between first-stage predictions $(\hat{x})$ and their residuals $(e)$ (i.e., the $\widehat{\text{Cov}}(\hat{x},e)$ term) accounts for the vast majority of the bias for Lasso and substantial amounts of the bias in random forests, boosted trees, and neural nets (the exact portion of the bias differs across DGPs and iterations). While $\widehat{\text{Cov}}(\hat{x},e)$ does not account for all of the bias, the non-zero covariance between first-stage predictions and residuals is an important (potentially large) component of the bias of ML-based 2SLS models.
Unlike $\widehat{\text{Cov}}(\hat{x},e)$, the second component of the wedge between $\hat\beta_1^{\text{2SLS}}$ and $\beta_1$ can be non-zero for both OLS-based methods and non-OLS models. However, methods that use non-linear predictions of $x$ in the first stage (i.e., ML-assisted 2SLS) require special care to reduce $\widehat{\text{Cov}}(\hat{x},u)$ and produce low-bias estimates of $\beta_1$.
This second term, $\widehat{\text{Cov}}(\hat{x},u)$, is effectively the exclusion restriction, and any 2SLS-inspired estimator can reduce bias in $\hat{\beta_1}$ by ensuring $\widehat{\text{Cov}}(\hat{x},u)$ is approximately zero. Assuming the instruments $\mathbf{z}$ are valid, an arbitrary prediction algorithm can maintain $\widehat{\text{Cov}}(\hat{x},u)\approx 0$ through either of three conditions:
Put simply, sufficiently flexible learning algorithms can recover endogenous variation in $x$---even when using linearly valid instruments. Many ML training methods explicitly incentivize and enable algorithms to do this.
Column $b$ of Table (ref) highlights the tendency of flexible first-stage models (e.g., tree methods and neural nets) to recover endogeneity. As the flexibility of algorithms increases, $\widehat{\text{Cov}}(\hat{x},u)$ tends to increase as well (across all DGPs). This covariance and its associated bias are particularly large for tree-based methods (especially random forests) and neural nets with multiple hidden layers. Notably, in panels (ref)--(ref) of Figure (ref), the densities of unrestricted and narrow neural networks are bimodal. As Appendix Figure (ref) illustrates, the bimodality results from whether the neural network (i) “chooses” zero hidden layers (the less biased mode) or (ii) goes deeper (learning the endogenous error and generating more bias).\footnote{This result highlights the importance of allowing neural networks to choose no hidden layers.}\footnote{Another related concern familiar to the ML literature is overfit. Overfit models tend to produce larger values of $\widehat{\text{Cov}}(\hat{x},u)$ than models that have been cross-validated. Though cross-validation is best/standard practice for machine-learning methods in prediction problems, here it retains importance by preventing the algorithms from overfitting the target variable $x$ in the first stage (even when out-of-sample performance is no longer the goal). We use five-fold cross-validation (CV) to tune the hyperparameters for Lasso-, tree-, and neural-net-based methods. Our neural-net cross-validation departs from standard five-fold CV. In Appendix Section (ref) we detail our cross-validation process for training neural net. One might further avoid overfit by applying holdout-style methods---only generating predictions for observation $i$ when $i$ is not in the training set. JIVE, split-sample IV, and Chen2020-IV all feature this additional safeguard. We do not employ these holdout-based methods because our goal in this paper is to simulate the results of a researcher using off-the-shelf ML tools in the first stage of 2SLS.} This covariance between predictions $\hat{x}$ and the unobserved disturbance $u$ accounts for a substantial amount of the bias in nonlinear methods, which demonstrates that the previously discussed first component $\widehat{\text{Cov}}(\hat{x},e)$ is not the only issue facing these models.
While the first two bias components enter additively, the third component scales their sum. Any method that reduces the variance of the first-stage predictions (i.e., reduces $\widehat{\text{Var}}(\hat{x})$) mechanically inflates the bias produced by $\beta_1\widehat{\text{Cov}}(\hat{x},e) + \widehat{\text{Cov}}(\hat{x},u)$.
In the case of properly specified, OLS-based 2SLS, the variance of the predictions hardly affects bias in $\beta_1$, since $\widehat{\text{Cov}}(\hat{x},e)=0$ and $\widehat{\text{Cov}}(\hat{x},u)\approx 0$. However, most ML algorithms reduce the variance of their predictions while trading between out-of-sample bias and variance. This tradeoff between bias and variance happens {\it outside of a 2SLS framework}. Consequently, when practitioners insert variance-reducing ML methods into 2SLS, the variance reduction actually amplifies bias in the second-stage estimates.
Taking these insights to the results in Panel B of Table (ref), notice that variance reduction can cause methods to perform poorly. For example, Lasso-assisted 2SLS produces the lowest variance $\hat{x}$ in two of the three high-complexity cases. (In cases 2 and 3, Lasso-based 2SLS has the highest $1/\widehat{\text{Var}}{\hat{x}}$-based amplifier of the bias.) This high degree of bias amplification generates notable bias in Lasso relative to many other methods. (This is also evident in Figure (ref)). So while Lasso's $\widehat{\text{Cov}}(\hat{x},e)$ and $\widehat{\text{Cov}}(\hat{x},u)$ are less than or equal to those of many other methods, the amplification produced by variance-reduction in $\hat{x}$ ultimately causes Lasso to have substantial bias. Notably, post-Lasso-based 2SLS produces less bias, partly due to the fact that it includes less variance reduction.
Worse yet, tree-based methods substantially reduce variance and produce relatively large $\widehat{\text{Cov}}(\hat{x},e)$ and $\widehat{\text{Cov}}(\hat{x},u)$ components, resulting in very large bias in their parameter estimates (even larger than na\"ive OLS).
In this paper, we examine the implications of plugging off-the-shelf ML methods into a 2SLS framework. In many cases, injecting ML into the first stage of 2SLS generates substantial bias.
While there are many approaches to combining instrumental-variable intuition and machine learning, they relax the traditional 2SLS structure and require different (generally stronger) identifying assumptions and/or tailor-made ML algorithms.\footnote{For example, MLSS Chen2020-IV, DeepIV Hartford2017, DeepGMM Bennett2020, KIV Singh2019, Adversarial Estimation of Riesz Representers chernozhukov2020, Neural Estimation of SEM liao2020, and Non-Parametric IV kilbertus2020class.} Among the current options, the closest in spirit to our question of “What are the implications of inserting ML into 2SLS?" is the “machine learning split-sample" (MLSS) estimator proposed by Chen2020-IV.\footnote{Angrist2020 also applies split-sample methods to several ML algorithms (i.e., post-Lasso, random forest), both in the first stage of 2SLS and while synthesizing instruments in a stage that precedes the first stage.}
With two fairly simple expansions of the traditional 2SLS framework, MLSS mitigates many biases generated by na\"ively plugging ML methods into the first stage. However, as with other more ML-forward methods, the solution is not without the cost of substantially strengthening the exclusion restriction. Specifically, Chen2020-IV proposes augmenting 2SLS with two simple techniques: restrict ML-based predictions to be explicitly out of sample (using split-sample methods) and use the ML-generated predictions as a “synthetic” instrument that then enters linearly in the first stage.
The idea for out-of-sample (split-sample) ML predictions follows the lead of Jackknife IV and Split Sample IV. By introducing out-of-sample methods to the ML-prediction exercise, Chen2020-IV aims to prevent the ML algorithm from fitting the first-stage errors, and shutting down the bias generated by $\widehat{\text{Cov}}(\hat{x},u)$. One potential drawback, however, is that this out-of-sample step likely increases variability (as seen in the JIVE results of Figure (ref)).
The second component of Chen2020-IV involves a zero\textsuperscript{th} stage (i.e., before the first stage), in which the practitioner trains an ML algorithm to predict $x$ using the instruments $\mathbf{z}$.\footnote{Note that this zero\textsuperscript{th} stage is identical to first stages that na\"ively insert ML methods into 2SLS.} The predictions from this zero\textsuperscript{th} stage are then used as the instrument within a traditional 2SLS framework. The benefit is that the resulting linear first stage (linearizing the results of a potentially forbidden regression) guarantees that $\widehat{\text{Cov}}(\hat{x},e)=0$ and shuts down one avenue through which bias enters.
Importantly, this zero\textsuperscript{th} stage of the MLSS approach requires that no learnable function of instruments meaningfully predicts the structural disturbance $u$. The function spaces of ML algorithms can cover all possible functions of the instruments (e.g., most neural networks are universal approximators), which requires strengthening the exclusion restriction from the assumption of “no correlation” to the actual underlying assumption of conditional mean-independence between the instruments and disturbance. For example, this strengthening includes ML-learned step functions, threshold indicators, kinks, many-way interactions---any learnable function $f(z)$ that is informative for $x$. This potentially infinite-dimensional set of exclusion restrictions is presumably more difficult to justify than the typical identifying assumption assumed in 2SLS applications.
Overall, solutions exist for the bias components that we identify in this paper, but each solution comes at a cost. Some solutions are fairly cheap (e.g., out-of-sample prediction methods require sufficient overlap and can increase noise). Other solutions require more from the practitioner, e.g., strengthening a linear exclusion restriction to an infinite-dimensional function space. Researchers may be comfortable invoking these stronger identifying assumptions in some settings. However, more work is needed in exploring and bounding the relative benefits and costs of these procedures in practice.\footnote{For example, future work could continue to explore the bias from instruments that nonlinearly violate an expanded exclusion restriction (potentially discoverable by nonlinear ML algorithms) but satisfy the traditional linear exclusion restriction. Future work could also consider how ML-based 2SLS approaches estimate potentially different LATEs.}
Our results show that na\"ively inserting machine-learning algorithms into the first stage of a 2SLS structure often produces bias. While some channels of this bias will be familiar to an economics audience, others may be less familiar, and arise from the intersection of “classical" econometric tools with ML methods. In terms of minimum bias, we find that the most-successful cases of integrating machine learning into 2SLS are largely restricted to instrument selection or modification (e.g., post-Lasso) before entering a traditional 2SLS framework. However, inserting highly nonlinear algorithms (e.g., random forests, boosted trees, and relatively deep neural nets) into the first stage of 2SLS drives estimates of causal parameters to be {\it more} biased than much simpler and more-easily interpreted alternatives. Importantly, such methods can even yield more bias than a na\"ive, endogenous OLS regression.
We mechanically decompose the bias within ML-augmented 2SLS estimates into three terms:
Overall, we find that many of the methods pioneered in the 1990s and 2000s (e.g., split-sample IV, LIML, and JIVE) reduce bias resulting from weak or over-identified IV for a linear first stage.\footnote{Angrist2020 comes to a similar conclusion. Focusing more on heterogeneous treatment effects and ML-assisted covariate selection, Angrist2020 does not decompose the bias associated with ML methods applied to the instruments in the first-stage prediction problem (Section (ref) and summarized above). Others have also found that JIVE and LIML are the best choice in the weak-many-instruments case---Hansen2014 explores a regularized JIVE, and Carrasco2015 develops a regularized LIML approach.} At the same time, we find that less-traditional methods that mindfully incorporate machine learning techniques (e.g., PCA-synthesized instruments and post-Lasso) generate improvements over 2SLS across a variety of DGPs.\footnote{See Ackerberg2009 for a discussion of tools for addressing over-identified instrumental variables estimation and Andrews2019 for a discussion of weak instruments in linear IV regression.} Importantly, we show that carelessly injecting ML into the first stage can also produce substantial bias, possibly worsening bias over na\"ive OLS.
The fundamental problem comes from recognizing the first stage of 2SLS solely as a prediction problem without recognizing that it is also part of a larger estimation {\it system}.
ML methods can produce good predictions. However, they were neither developed nor optimized for two-stage procedures that generate minimally biased causal estimates. For example, variance reduction in prediction typically improves out-of-sample prediction accuracy but can actually inflate parameter bias in the second stage of 2SLS. Solutions have been proposed to these and other issues from ML-based 2SLS. However, if one approaches these problems without the caution that we suggest (e.g., blindly relying upon ML to learn the variation in the first stage) bias often results. Unfortunately, the solutions to these ML-in-2SLS problems can often require the practitioner to strengthen the exclusion restriction where complexity is highest (e.g., in high-dimensional data). Because ML methods are typically adopted in complex data environments that are not well understood, employing these methods without safeguards can entice researchers into making heroic assumptions.
Given the performance of existing off-the-shelf estimators, na\"ively injecting ML into 2SLS appears to produce little gain relative to its costs/biases. More broadly, applying ML methods to 2SLS requires the practitioner to face issues present in both the prediction and causal-inference settings---and in addition, less-familiar issues arising from their interaction.
\newgeometry{margin=1in}
\subfile{main-results-table}
\subfile{bias-decomp-table}
\singlespacing