EconBase
← Back to paper

Empirical Decomposition of the IV-OLS Gap with Heterogeneous and Nonlinear Effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

94,843 characters · 0 sections · 108 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Empirical Decomposition of the IV--OLS Gap with Heterogeneous and Nonlinear Effects

abstractThis study proposes an econometric framework to interpret and empirically decompose the difference between IV and OLS estimates given by a linear regression model when the true causal effects of the treatment are nonlinear in treatment levels and heterogeneous across covariates. I show that the IV--OLS coefficient gap consists of three estimable components: the difference in weights on the covariates, the difference in weights on the treatment levels, and the difference in identified marginal effects that arises from endogeneity bias. Applications of this framework to return-to-schooling estimates demonstrate the empirical relevance of this distinction in properly interpreting the IV--OLS gap. JEL Classification: C21, C26, I26

\part[Main Paper]

refsection\section{Introduction} Instrumental variables (IV) regression is the most common approach for estimating the causal effect of a potentially endogenous regressor. A standard empirical approach specifies the following linear model: \begin{equation} Y=\beta X+W'\gamma+\varepsilon, \end{equation} where $Y$ is the outcome of interest, $X$ is a scalar (multivalued) treatment, and $W$ is a vector of covariates. A standard econometric textbook takes the linear regression equation ((ref)) as the true causal relationship and interprets the gap between the IV and ordinary least squares (OLS) coefficient estimates of $X$ as a consequence of endogeneity bias associated with omitted variables, selection, or measurement error. If these interpretations fail to provide a plausible explanation for the IV--OLS coefficient gap, empirical researchers often consider the possibility that it instead arises because of how the IV and OLS coefficients place different weights on different treatment margins or groups of individuals. This interpretation treats the regression equation ((ref)) as a linear projection model rather than a causal model, allowing for the treatment effects to be heterogeneous and nonlinear in the true causal relationship. For example, card1995earnings,card1999causal,card2001estimating suggests that the positive IV--OLS gaps in many return-to-schooling studies could be explained by higher returns among credit-constrained individuals, who are more likely to be affected by cost-related instruments. In addition, researchers sometimes perform an OLS regression restricted to a sample of likely complier groups or complying treatment margins to examine the robustness of the IV--OLS coefficient gap to the weight difference. As a prominent example, angrist1991does compute an OLS estimate of the return to schooling by restricting their sample to individuals with 9--12 years of schooling, who are expected to be most influenced by quarter-of-birth instruments. However, despite the widely recognized importance of the weight difference interpretation, relatively few studies attempt to formally quantify how much it actually matters for the IV--OLS coefficient gap.\footnote{Notable exceptions include kling2001interpreting, lochner2001effect, mogstad2010linearity, loken2012linear, and lochner2015estimating.} In this study, I propose an econometric framework to quantify the sources of the IV--OLS coefficient gap. I begin my analysis by considering the following causal model: \begin{equation} Y=g(X,W)+U, \end{equation} with a valid instrument $Z$ that is uncorrelated with an unobservable $U$ conditional on covariates $W$. With no restriction on the structural function $g$, the equation allows for the treatment effects to be heterogeneous across covariates $W$ and nonlinear in treatment levels $X$ in any manner.\footnote{The separability restriction on the equation, which is relaxed in Section (ref), rules out unobserved heterogeneity in the treatment effects.} I demonstrate that the OLS and IV estimates based on the linear regression equation ((ref)), when the true model is ((ref)), represent different weighted averages of the marginal effects of the treatment they identify. Using the weighted-average representations, I then decompose the IV--OLS coefficient gap into three estimable components. These are: (i) the covariate weight difference, being the difference in how the IV and OLS coefficients place weights on the covariates $W$; (ii) the treatment-level weight difference, as the difference in how they place weights on treatment levels $X$; and (iii) the endogeneity bias (or marginal effect difference), being the difference between the IV- and OLS-identified marginal effects originating from the correlation between the treatment $X$ and the unobservable $U$. For instance, in the return-to-schooling context, (i) arises from heterogeneous returns across observed personal backgrounds with different responses to the instrument, which motivates the conjecture of card1995earnings,card1999causal,card2001estimating, (ii) follows from nonlinear returns across schooling levels with different sensitivity to the instrument, which motivates the robustness check of angrist1991does, and (iii) corresponds to endogeneity bias associated with omitted unobserved ability. To implement the decomposition, I propose a two-step approach to estimate the “IV-weighted OLS” coefficients that serve as the intermediate points between the IV and OLS coefficients.\footnote{The concept of an IV-weighted OLS coefficient originates in \citet*{mogstad2010linearity}, where they use the model $Y=g(X)+U$ and account for the treatment-level weight difference.} The first step predicts the conditional mean of the outcome $Y$ given $(X,W$), and the second step uses the prediction to construct a dependent variable and performs a quasi-IV regression. Depending on the functional form restriction in the first step, the estimated OLS coefficients in the second step have the IV weights on the covariates or on both the covariates and treatment levels. Comparing the estimated IV-weighted OLS coefficients with the IV and OLS coefficients reveals the contributions of the weight difference components (i) and (ii) and the endogeneity bias component (iii). Standard statistical packages can compute the IV-weighted OLS estimates and perform statistical inference based on them.\footnote{The Stata package that implements the decomposition, ivolsdec, is available from the Boston College Statistical Software Components (SSC) archive. Type “ssc install ivolsdec” in the Stata command window to install it.} In an extension, I consider a class of identification strategies that use an instrument $Z$ deterministic in the covariates $W$, as in difference-in-differences (DID) and regression discontinuity (RD) designs.\footnote{Throughout this paper, the term “DID” refers to an identification strategy that exploits DID variation in the instrument using a two-way fixed effects regression, in which the treatment and the instrument can be nonbinary. Examples include acemoglu2000large, duflo2001schooling, black2005apple, and a fuzzy DID setup considered in de2018fuzzy. The term “RD” refers to a fuzzy RD design, which is usually implemented as an IV regression.} I show that the weighted-average interpretation and the decomposition approach can also be applied to these setups, with only a small modification of the weight function. It should be noted that my decomposition framework cannot fully isolate endogeneity bias in the presence of unobserved heterogeneity in the treatment effects. In an extension that relaxes the separability restriction on the equation ((ref)), I demonstrate that the marginal effect difference component (iii) captures not only endogeneity bias but also the unobservable-driven discrepancy between the IV-identified and the average marginal effects.\footnote{This corresponds to a classic impossibility result in the setting with a binary treatment and no covariates, where endogeneity bias cannot generally be separated from the difference between the local and population average treatment effects.} While the weight difference components (i) and (ii) remain informative about the implications of observed heterogeneity and nonlinearity, it can be misleading to attribute the marginal effect difference component (iii) entirely to endogeneity bias. Using my framework, I examine the return-to-schooling estimates using several common IV strategies. The first example employs geographic variation in college costs cameron2004estimation,carneiro2011estimating. The second exploits a discontinuity in the minimum school-leaving age across cohorts oreopoulos2006estimating. The third and final example uses DID variation in compulsory schooling laws across cohorts and regions acemoglu2000large. In these empirical examples, the weight difference components are found to be as important as the endogeneity bias component in explaining the IV--OLS coefficient gap. The direction or extent of endogeneity bias implied by the estimated IV--OLS gap differs entirely by taking into consideration how the two coefficients place weights on the different observed personal backgrounds and schooling margins. \subsection*{Related Literature and Roadmap} This paper advances the literature on the interpretation and decomposition of linear regression coefficients by exploring a general and empirically relevant setting in which the treatment effects are nonlinear in treatment levels and heterogeneous across covariates. The local average treatment effect (LATE) interpretation proposed by imbens1994identification in a binary treatment context originates the idea that the IV coefficient is a weighted average of the marginal causal effects of the treatment. angrist1995two and angrist2000interpretation extend the basic insights of the LATE interpretation to a multivalued treatment case. yitzhaki1996using and angrist1999empirical suggest the analogous weighted-average interpretation of the OLS coefficient.\footnote{The OLS interpretation is also explored by angrist1998estimating, aronow2016does, and Sloczynski2020OLS in a binary treatment case with a focus on observed heterogeneity.} Much of the focus of the literature has been on a univariate model with no or fixed covariates, which makes it difficult to immediately apply these results to empirical settings. While drawing on these existing results, my framework synthesizes them into an empirically relevant format and provides an estimable decomposition of the IV--OLS coefficient gap. Motivated by the weighted-average interpretation developed in the literature, mogstad2010linearity, \citet*{loken2012linear}, and \citet*{lochner2015estimating} propose the empirical decomposition of the IV--OLS coefficient gap into weight difference and endogeneity bias components. They consider a model in which the treatment effects are nonlinear in treatment levels but homogeneous across covariates. My framework generalizes these previous works by allowing the treatment effects to be heterogeneous across covariates. This is an empirically meaningful generalization, as my empirical applications demonstrate the relevance of the covariate weight difference in interpreting the IV--OLS coefficient gap. While the linear IV regression is often advocated for its transparency angrist2010credibility, it is also criticized for its lack of a clear connection to an economic parameter of interest heckman2010comparing.\footnote{For the linear OLS regression, Sloczynski2020OLS shows that the OLS coefficient of a binary treatment with covariates generally does not represent the average treatment effect (ATE), the average treatment effect on treated (ATT), or untreated (ATU).} A related strand of the literature based on the latter view aims to develop alternatives to the linear IV regression, including the policy-relevant treatment effect proposed by heckman2001policy and carneiro2010evaluating. Nevertheless, many empirical researchers use the linear regression for its simplicity. My framework aims to provide useful diagnostics for empirical researchers who use the linear regression while recognizing its potential limitations. The rest of the paper proceeds as follows. Section 2 presents the interpretation of the linear IV and OLS coefficients and proposes the decomposition of the IV--OLS coefficient gap. Section 3 extends these econometric results by exploring settings with alternative assumptions. Section 4 proposes the estimators for the IV-weighted OLS coefficients for empirically performing the decomposition. Section 5 presents the results from the empirical applications, and Section 6 concludes. Online Appendix presents the proofs of the theorems, explores additional econometric results, and provides details for the empirical applications. \section{Econometric Framework} \subsection{Setup and Assumptions} I consider a random draw of $(Y,X,W,Z)$, where $Y$ is a scalar outcome variable, $X$ is a scalar treatment variable, $W$ is a vector of covariates, and $Z$ is a scalar instrument. Multi-instrument two-stage least squares can fit into this setup by regarding the projection of the treatment $X$ onto the instrument vector and covariates $W$ (i.e., the first-stage predicted value) as a synthetic scalar instrument.\footnote{If a regression of $X$ on a vector instrument $(Z_{1},\ldots,Z_{M})$ controlling for $W$ yields a first-stage coefficient $(\pi_{1},\ldots,\pi_{M})$, I regard $Z=\sum_{m=1}^{M}\pi_{m}Z_{m}$ as a synthetic scalar instrument. As my setting allows for negative IV weights, a partial monotonicity condition for each individual instrument $Z_{m}$ as in mogstad2019causal is not required as long as the treatment $X$ and the synthetic instrument $Z$ are correlated conditional on $W$.} Throughout this paper, I assume the existence of the first and second moments of any random variable. I use $L_{w}(R)=w'E(WW')^{-1}E(WR)$ to denote a linear projection of a random variable $R$ onto $W$ evaluated at $W=w$ (i.e., a predicted value from a linear regression of $R$ on $W$).\footnote{The covariate vector $W$ includes unity as one of its elements.} I define $\widetilde{R}=R-L_{W}(R)$ to be a residual from the linear projection. Let $m(x,w)=E[Y|X=x,W=w]$ be the conditional mean function of $Y$ given $(X,W)$. \global\long To assess the impact of the treatment $X$ on the outcome $Y$, a standard approach specifies a linear regression model: \begin{equation} Y=\beta X+W'\gamma+\varepsilon,\,\,E(\varepsilon W)=0. \end{equation} Additional moment conditions $E(\varepsilon Z)=0$ and $E(\varepsilon X)=0$, respectively pin down the linear IV coefficient $\beta_{IV}$ and OLS coefficient $\beta_{OLS}$ as \begin{eqnarray} \beta_{IV} & = & { { E(\widetilde{Y}\widetilde{Z})/E(\widetilde{X}\widetilde{Z})}},\\ \beta_{OLS} & = & { E(\widetilde{Y}\widetilde{X})/E(\widetilde{X}^{2})}. \end{eqnarray} Note that I treat the equation ((ref)) merely as a statistical model to characterize $\beta_{IV}$ and $\beta_{OLS}$, which does not impose any assumption on the underlying causal relationship. To consider the causal interpretation of the IV and OLS coefficients, I define $Y(x)$ to be the potential outcome associated with the treatment level $x$, which produces the observed outcome as $Y=Y(X)$. The derivative $Y'(x)$ is the marginal effect of the treatment in a causal sense. I make the following assumptions. \begin{assumptionx} (Separability) The potential outcome is given by $Y(x)=g(x,W)+U$, $E(U|W)=0$. \end{assumptionx} \begin{assumptionx} (Continuous Treatment) The treatment $X$ is continuously distributed on support $(\underline{x},\overline{x})$, with $-\infty\le\underline{x}<\overline{x}\le\infty$. \end{assumptionx} \begin{assumptionx} \textup{(Regularity Conditions on Derivatives)} Let $V_{a}^{b}(f,w)=\int_{\min\{a,b\}}^{\max\{a,b\}}|\frac{\partial}{\partial x}f(x,w)|dx$ be the total variation of a function $f(x,w)$ differentiable in $x$ between points $a$ and $b$. \textup{(i) } $g(x,w)$ is differentiable in $x$ and $E\left[V_{x_{0}}^{X}(g,W)^{2}\right]<\infty$ for some $x_{0}\in(\underline{x},\overline{x})$; \textup{(ii)} $m(x,w)$ is differentiable in $x$ and $E\left[V_{x_{0}}^{X}(m,W)^{2}\right]<\infty$ for some $x_{0}\in(\underline{x},\overline{x})$. \end{assumptionx} \begin{assumptionx} The instrument $Z$ satisfies the conditions: \textup{(i) (Exogeneity)} $E(U\widetilde{Z})=0$; \textup{(ii) (Relevance)} $E(\widetilde{X}\widetilde{Z})\ne0$. \end{assumptionx} \begin{assumptionx} The treatment residual has positive variance, i.e., $E(\widetilde{X}^{2})>0$. \end{assumptionx} \begin{assumptionx} \textup{(Linearity of Conditional Means)} \textup{(i)} $E(Z|W)$ is linear in $W$; \textup{(ii)} $E(X|W)$ is linear in $W$. \end{assumptionx} Given Assumption (ref), the marginal causal effect of the treatment is $Y'(x)=\frac{\partial}{\partial x}g(x,W)$, which can be nonlinear in treatment levels $x$ and heterogeneous across covariates $W$. However, it rules out unobserved heterogeneity in the effect.\footnote{The literature on the nonparametric IV approach typically makes a similar separability assumption. See, for example, newey2003instrumental, blundell2007semi, and horowitz2011applied.} Section (ref) relaxes this assumption. I make Assumption (ref) to focus on a continuous treatment case, which is merely for expositional convenience.\footnote{As indicated in the assumption, the support $(\underline{x},\overline{x})$ can be unbounded. Any integral expression of $x$ in this paper is taken over the support $(\underline{x},\overline{x})$, which is kept implicit to simplify the exposition.} All econometric results can be applied to a discrete treatment case by extending $g(x,w)$ to nonsupport points of the treatment.\footnote{For example, define $g(x,w)$ at nonsupport points by a linear interpolation without loss of generality. Then, the derivatives and integrals of $x$ in all econometric results can be replaced by the differences and summations of $x$.} Assumption (ref) concerns the derivatives of the structural function $g$ and the conditional mean function $m$. Assumption (ref) is a set of standard IV assumptions that require the instrument to be exogenous and relevant after controlling for covariates. Assumption (ref) is a standard OLS assumption. Assumption (ref) is required for the exact weighted-average interpretation of the regression coefficients. A similar linearity assumption appears in, for example, angrist1999empirical, lochner2015estimating, and sloczynkski2020when. This assumption allows the analysis to abstract away from any omitted variable bias associated with unaccounted nonlinear effects of covariates, which is not as fundamental as endogeneity bias associated with unobservables. While this assumption mechanically holds in a saturated model in which a covariate vector $W$ consists of indicators for disjoint groups, it can be restrictive with continuous covariates. An empirical researcher may then want to choose elements of the vector $W$ based on a series approximation and thereby flexibly account for the nonlinear and interaction effects of the underlying observables.\footnote{Appendix (ref) relaxes Assumption (ref) to explore what happens when a good linear approximation is not feasible because of data limitations.} Assumption (ref) does not hold by construction for identification strategies based on DID or RD designs. Section (ref) considers these cases. \subsection{Weighted-Average Interpretation} I start by presenting a key theorem for interpreting the IV and OLS coefficients. The OLS part of the theorem is shown by angrist1999empirical in a discrete treatment setting. I present it for completeness and for reinterpretation in my setting. \begin{thm} The IV and OLS coefficients have a weighted-average interpretation as below. \begin{enumerate} • With Assumptions (ref), (ref), (ref)--(i), (ref), and (ref)--(i), \begin{align*} \beta_{IV} & =\int\int\frac{\partial}{\partial x}g(x,w)\omega_{Z}(x,w)dF_{W}(w)dx; \end{align*} • angrist1999empirical With Assumptions (ref), (ref)--(ii), (ref), and (ref)--(ii), \begin{align*} \beta_{OLS} & =\int\int\frac{\partial}{\partial x}m(x,w)\omega_{X}(x,w)dF_{W}(w)dx. \end{align*} \end{enumerate} The weight function is given by $\omega_{R}(x,w)=E[\widetilde{\mathbbm{1}}_{X\ge x}\widetilde{R}|W=w]/E(\widetilde{X}\widetilde{R})$ for $R=Z,X$, which satisfies $\int\int\omega_{R}(x,w)dF_{W}(w)dx=1$. \end{thm} Theorem (ref) implies that the IV and OLS coefficients are expressed as weighted averages of the marginal effects they identify. While the IV coefficient identifies a weighted average of the causal effects $\frac{\partial}{\partial x}g(x,w)$, the OLS coefficient identifies a weighted average of the slopes of the conditional mean function $\frac{\partial}{\partial x}m(x,w)$. The relationship between the OLS-identified and IV-identified marginal effects, $\frac{\partial}{\partial x}m(x,w)$ and $\frac{\partial}{\partial x}g(x,w)$, can be expressed as \begin{equation} \frac{\partial}{\partial x}m(x,w)=\frac{\partial}{\partial x}g(x,w)+\frac{\partial}{\partial x}E[U|X=x,W=w]. \end{equation} Therefore, the difference between the two marginal effects arises from endogeneity bias, i.e., the correlation between the treatment $X$ and the unobservable $U$. While endogeneity bias makes the IV and OLS coefficients differ, an important implication from Theorem (ref) is that the difference in the weight functions, $\omega_{Z}$ and $\omega_{X}$, also gives rise to the IV--OLS coefficient gap. To explore what makes the IV weight $\omega_{Z}$ and the OLS weight $\omega_{X}$ differ, let $\overline{\omega}_{R}(w)=\int\omega_{R}(x,w)dx$ be the marginal weight on covariates $W=w$ for $R=Z,X$. The marginal IV and OLS weights on $W=w$ are given by: \begin{align} \overline{\omega}_{Z}(w) & =Cov(X,Z|W=w)/E(\widetilde{X}\widetilde{Z}),\\ \overline{\omega}_{X}(w) & =Var(X|W=w)/E(\widetilde{X}^{2}). \end{align} The IV weight $\overline{\omega}_{Z}(w)$ is proportional to the conditional covariance $Cov(X,Z|W=w)$, which is the product of the regression coefficient of $X$ on $Z$ given $W=w$ and the conditional variance $Var(Z|W=w)$. This means that covariates $W$ with greater sensitivity of the treatment to the instrument or larger variation in the instrument are weighted more. In contrast, the OLS weight $\overline{\omega}_{X}(w)$ is proportional to the conditional variance $Var(X|W=w)$. This implies that covariates $W$ with larger variation in the treatment are weighted more. Similarly, let $\overline{\omega}_{R}(x)=\int\omega_{R}(x,w)dF_{W}(w)$ be the marginal weight on the treatment level $X=x$ for $R=Z,X$. The marginal IV and OLS weights on $X=x$ are given by: \begin{align} \overline{\omega}_{Z}(x) & =E(\widetilde{\mathbbm{1}}_{X\ge x}\widetilde{Z})/E(\widetilde{X}\widetilde{Z}),\\ \overline{\omega}_{X}(x) & =E(\widetilde{\mathbbm{1}}_{X\ge x}\widetilde{X})/E(\widetilde{X}^{2}). \end{align} The IV weight $\overline{\omega}_{Z}(x)$ is proportional to a regression coefficient of $\mathbbm{1}_{X\ge x}$ on the instrument $Z$, controlling for $W$. This implies that the more the instrument $Z$ influences the treatment $X$ at $x$, the more weighted the treatment level $x$ is. For example, if a compulsory schooling instrument $Z$ increases years of schooling $X$ through primary and secondary education but does not influence college education, the IV weight is expected to be positive with $x\le12$ and zero with $x>12$.\footnote{If $X$ is discrete, the weight on $x$ represents a change in treatment levels between $x-1$ and $x$.} In general, the IV weight $\overline{\omega}_{Z}(x)$ is not guaranteed to be positive because the instrument can have positive effects on the treatment at some margins while having negative effects at others. Interpreting the OLS weight expression is less straightforward, but Appendix (ref) shows \begin{equation} \overline{\omega}_{X}(x)\propto\int\int\int(x_{1}-x_{2})\mathbbm{1}_{x_{2}<x\le x_{1}}dF_{X|W}(x_{1}|w)dF_{X|W}(x_{2}|w)dF_{W}(w). \end{equation} This implies that the OLS weight is proportional to the sum of differences between the pairs of conditionally independent observations $X_{1},X_{2}\overset{i.i.d.}{\sim}F_{X|W}(\cdot|w)$ with $X_{1}\ge x>X_{2}$. Therefore, the treatment level $x$ is weighted more if the treatment $X$ is densely distributed both above and below $x$.\footnote{yitzhaki1996using derives the OLS weight function in a simpler case with no covariates and shows that the weight function is $\overline{\omega}_{X}(x)\propto(b-x)(x-a)$ if $X$ is uniformly distributed over $[a,b${]}. The weights are zero at both ends of the support despite flat density because no pair of observations can sandwich the endpoints.} As is evident from ((ref)), the OLS weight $\overline{\omega}_{X}(x)$ on each treatment level $x$ is nonnegative. Theorem (ref) can be considered as a generalization of lochner2015estimating, allowing for heterogeneity in the marginal effects across covariates. In fact, restricting $g(x,w)$ to be additively separable in $x$ and $w$ yields a weighted-average expression comparable to theirs. If $g(x,w)$ is additively separable, $\frac{\partial}{\partial x}g(x,w)$ depends only on $x$. This yields $\beta_{IV}=\int\frac{\partial}{\partial x}g(x,\cdot)\overline{\omega}_{Z}(x)dx$, which matches the weighted-average expression provided by Proposition 1 in lochner2015estimating. \subsection{Related Work} Theorem (ref) closely relates to many existing results in the literature. To provide further intuition for the weighted-average interpretation, I explore the relationship between my results and those in the literature. Denote the IV and OLS coefficients conditional on $W=w$ as \begin{eqnarray} b_{IV}(w) & = & Cov(Y,Z|W=w)/Cov(X,Z|W=w),\\ b_{OLS}(w) & = & Cov(Y,X|W=w)/Var(X|W=w). \end{eqnarray} The following result from lochner2015estimating shows that the IV and OLS coefficients, $\beta_{IV}$ and $\beta_{OLS}$, can be viewed as weighted averages of the covariate-specific coefficients, $b_{IV}(w)$ and $b_{OLS}(w)$. \begin{thm} lochner2015estimating \begin{enumerate} • With Assumptions (ref)--(ii) and (ref)--(i), \begin{align*} \beta_{IV} & =\int b_{IV}(w)\overline{\omega}_{Z}(w)dF_{W}(w); \end{align*} • With Assumptions (ref) and (ref)--(ii), \begin{align*} \beta_{OLS} & =\int b_{OLS}(w)\overline{\omega}_{X}(w)dF_{W}(w). \end{align*} \end{enumerate} \end{thm} Note that this theorem does not rely on assumptions about the causal structure (Assumptions (ref) and (ref)--(i)) because $b_{IV}(w)$ is defined in ((ref)) merely as a ratio of two conditional covariances. The IV and OLS weights on covariates $W=w$ match the marginal weights $\overline{\omega}_{Z}(w)$ and $\overline{\omega}_{X}(w)$ defined in ((ref)) and ((ref)). This weighted-average interpretation is explored intensively in a binary treatment case, in which the treatment effect is linear in treatment levels $x\in\{0,1\}$ by construction and it is possible to focus only on heterogeneity angrist1998estimating,aronow2016does,Sloczynski2020OLS,sloczynkski2020when. The weighted-average interpretation of the covariate-specific coefficients, $b_{IV}(w)$ and $b_{OLS}(w)$, can be derived by applying Theorem (ref) conditional on $W=w$. This result originates in yitzhaki1996using and schechtman2004gini. \begin{thm} Let $\omega_{R}(x|w)=\omega_{R}(x,w)/\overline{\omega}_{R}(w)$ be the conditional weight on the treatment level $x$ given $W=w$. \begin{enumerate} • schechtman2004gini With Assumptions (ref), (ref), (ref)--(i), $Cov(U,Z|W=w)=0$, and $Cov(X,Z|W=w)\ne0$, \begin{align*} b_{IV}(w) & =\int\frac{\partial}{\partial x}g(x,w)\omega_{Z}(x|w)dx; \end{align*} • yitzhaki1996using With Assumptions (ref), (ref)--(ii), and $Var(X|W=w)>0$, \begin{align*} b_{OLS}(w) & =\int\frac{\partial}{\partial x}m(x,w)\omega_{X}(x|w)dx. \end{align*} \end{enumerate} \end{thm} The weighted-average expression for the covariate-specific IV coefficient $b_{IV}(w)$ could also be derived as a special case of the results from angrist1995two, angrist2000interpretation, and heckman2006understanding, which consider a more general setting that allows for unobserved heterogeneity in the treatment effects.\footnote{angrist1995two allow for covariates in a special case with a “saturated” first stage, where $Z=E[X|Z_{1},\ldots,Z_{M},W]$ is generated from an instrument vector $(Z_{1},\ldots,Z_{M})$ and a covariate vector $W$ that both consist of indicators for disjoint groups.} Synthesizing these existing results, Theorem (ref) can be divided into two components: the linear IV and OLS coefficients, $\beta_{IV}$ and $\beta_{OLS}$, are weighted averages of the covariate-specific coefficients, $b_{IV}(w)$ and $b_{OLS}(w)$, with different weights on covariates (Theorem (ref)); and the covariate-specific coefficients are weighted averages of the identified marginal effects, $\frac{\partial}{\partial x}g(x,w)$ and $\frac{\partial}{\partial x}m(x,w)$, with different weights on treatment levels (Theorem (ref)). \subsection{Decomposing the IV--OLS Coefficient Gap} The weighted-average interpretation of the IV and OLS coefficients in Theorems (ref)--(ref) motivates the decomposition of the IV--OLS coefficient gap $\beta_{IV}-\beta_{OLS}$. I decompose the gap into the following three components. \begin{eqnarray} \Delta_{CW} & = & \int b_{OLS}(w)\left(\overline{\omega}_{Z}(w)-\overline{\omega}_{X}(w)\right)dF_{W}(w),\\ \Delta_{TW} & = & \int\int\frac{\partial}{\partial x}m(x,w)\left(\omega_{Z}(x|w)-\omega_{X}(x|w)\right)\overline{\omega}_{Z}(w)dF_{W}(w)dx,\\ \Delta_{ME} & = & \int\int\left(\frac{\partial}{\partial x}g(x,w)-\frac{\partial}{\partial x}m(x,w)\right)\omega_{Z}(x,w)dF_{W}(w)dx. \end{eqnarray} The first component, $\Delta_{CW}$, which I call “the covariate weight difference,” corresponds to how differently the IV and OLS coefficients place weights on the covariates. The second component, $\Delta_{TW}$, which I call “the treatment-level weight difference,” captures how differently the IV and OLS coefficients place weights on the treatment levels, conditional on the covariates. The third component, $\Delta_{ME}$, which I refer to as “the endogeneity bias” or “the marginal effect difference,” captures the difference between the IV- and OLS-identified marginal effects, which arises from the endogeneity of the treatment as in ((ref)). The sum of the three components is the IV--OLS gap, i.e., $\beta_{IV}-\beta_{OLS}=\Delta_{CW}+\Delta_{TW}+\Delta_{ME}$. This decomposition departs from the OLS coefficient $\beta_{OLS}$ and arrives at the IV coefficient $\beta_{IV}$ by first changing the weights and then the marginal effects. This order follows an idea of the “IV-weighted OLS” approach adopted by mogstad2010linearity and lochner2015estimating. The advantage of this approach is that the decomposition is always feasible. The decomposition requires knowledge of the following IV-weighted OLS coefficients as intermediate points: \begin{eqnarray} \beta_{C} & = & \int b_{OLS}(w)\overline{\omega}_{Z}(w)dF_{W}(w),\\ \beta_{CT} & = & \int\int\frac{\partial}{\partial x}m(x,w)\omega_{Z}(x,w)dF_{W}(w)dx. \end{eqnarray} The first coefficient, $\beta_{C}$, is the OLS coefficient with the IV weight on covariates, while the second, $\beta_{CT}$, is the OLS coefficient with the IV weight on both covariates and treatment levels. By construction, $\Delta_{CW}=\beta_{C}-\beta_{OLS}$, $\Delta_{TW}=\beta_{CT}-\beta_{C}$, and $\Delta_{ME}=\beta_{IV}-\beta_{CT}$. These coefficients are always well-defined because $\omega_{Z}(x,w)\ne0$ implies $\omega_{X}(x,w)>0$. This is not a unique order in which the IV--OLS gap can be decomposed, as is the case with Blinder--Oaxaca type decomposition methods. For example, one can instead account for the marginal effect difference first, then the weight differences. However, this alternative order requires knowledge of the “OLS-weighted IV” coefficient, i.e., $\int\int\frac{\partial}{\partial x}g(x,w)\omega_{X}(x,w)dF_{W}(x)dx$, as an intermediate point. This coefficient is not always identified because the instrument may have no variation or no impact on the treatment at some $(x,w)$, even with $\omega_{X}(x,w)>0$.\footnote{loken2012linear propose a mixture of the IV-weighted OLS and the OLS-weighted IV to make decomposition results independent of whether the marginal effect or the weight difference is accounted for first. This approach has the same identification issue as the OLS-weighted IV approach.} This makes the OLS-weighted IV approach less practical despite its potential theoretical appeal.\footnote{In all empirical examples in Section (ref), the OLS-weighted IV coefficient cannot be identified because the instrument influences only a subset of treatment margins.} \section{Extensions} I extend my econometric framework in several directions. Section (ref) considers identification strategies based on DID and RD designs, in which the instrument $Z$ is deterministic in the covariates $W$. Section (ref) relaxes Assumption (ref) and allows for unobserved heterogeneity in the treatment effects. Appendix (ref) explores additional extensions that consider a setting without Assumption (ref), a setting with an invalid instrument, and a setting with DID or RD designs in the presence of unobserved heterogeneity. \subsection{Identification Based on DID or RD Designs} Assumption (ref) does not hold by construction for two important identification strategies: DID and RD designs. I explore the weighted-average interpretation in these cases with an alternative set of assumptions. With a DID-based identification strategy, each observation belongs to a particular group $g\in\{1,\ldots,G\}$ and period $t\in\{1,\ldots,T\}$, and the instrument $Z$ is constant within each $(g,t)$. The regression equation ((ref)) can be written as \[ Y=\beta X+\sum_{g=1}^{G}\gamma_{g}d_{g}+\sum_{t=1}^{T-1}\delta_{t}D_{t}+\varepsilon, \] where $d_{g}$ indicates membership to a group $g$ and $D_{t}$ indicates membership to a period $t$. By construction, the instrument $Z$ is a deterministic and nonlinear function of the covariate vector $W=(d_{1},\ldots,d_{G},D_{1},\ldots,D_{T-1})$. This does not satisfy the requirement by Assumption (ref) that $E(Z|W)$ should be linear in $W$. An empirical example from acemoglu2000large in Section (ref) fits into this setting. With an RD-based identification strategy using a running variable $C$ with a cutoff $c$, the instrument is $Z=\mathbbm{1}_{C\ge c}$. The regression equation ((ref)) can be written as \[ Y=\beta X+\sum_{k=1}^{K}\gamma_{k}p_{k}\left(C\right)+\varepsilon, \] where $(p_{1},\ldots,p_{K})$ is a set of basis functions with $p_{k}(c)=0$.\footnote{Given that an RD with local polynomials can be interpreted as a kernel-weighted version of an RD with global polynomials, my description focuses on a global polynomial case. The IV weight function should be multiplied by a kernel weight in a local polynomial case.} The instrument $Z$ is a deterministic and nonlinear function of the covariate vector $W=\left(p_{1}(C),\ldots,p_{K}(C)\right)$. An empirical example from oreopoulos2006estimating in Section (ref) fits into this setting. The definition of the linear IV and OLS coefficients follows ((ref)) and ((ref)). I rule out a “sharp” DID or RD setup with $X=Z$, in which the IV and OLS coefficients are identical by construction. Instead of Assumption (ref), I make the following assumption to represent DID- and RD-based identification strategies. \begin{assumptionx} \textup{(Linear Structural Function)} $g(x,w)$ is linear in $w$ for any $x\in(\underline{x},\overline{x})$. \end{assumptionx} In a DID, this assumption corresponds to a parallel trend assumption, indicating that the group membership $d_{g}$ and the time membership $D_{t}$ additively affect a potential outcome. In an RD, this assumption implies that the relationship between a potential outcome and a running variable $C$ is well approximated by a linear combination of the basis functions $\left(p_{1}(C),\ldots,p_{K}(C)\right)$ around $C=c$, which also implies continuity at $C=c$. In these settings, Assumption (ref)--(i) follows from $Var(Z|W)=0$, as $E(U\widetilde{Z})=E(E(U|W)\widetilde{Z})=0$. With Assumption (ref) replaced by Assumption (ref), the weighted-average interpretation of the IV coefficient remains valid with a modified weight function expression. \begin{thm} With Assumptions (ref), (ref), (ref)--(i), (ref), and (ref), the IV coefficient $\beta_{IV}$ is given by \begin{align*} \beta_{IV} & =\int\int\frac{\partial}{\partial x}g(x,w)\omega_{Z}^{*}(x,w)dF_{W}(w)dx, \end{align*} where $\omega_{Z}^{*}(x,w)=L_{w}(\mathbbm{1}_{X\ge x}\widetilde{Z})/E(\widetilde{X}\widetilde{Z})$ satisfies $\int\int\omega_{Z}^{*}(x,w)dF_{W}(w)dx=1$. \end{thm} Note that there are infinitely many weight functions other than $\omega_{Z}^{*}$ that can make this weighted-average expression. In particular, $\omega_{Z}^{*}(x,w)+h(x,w)-L_{w}\left(h(x,W)\right)$ can also be a weight function, where $h(x,w)$ is any nonlinear function. This property is a mere artifact of Assumption (ref), which makes the marginal effect $\frac{\partial}{\partial x}g(x,w)$ linear in $w$ and orthogonal to any linear projection residual. Therefore, it is most reasonable to use the linearly projected function $\omega_{Z}^{*}$ and omit the redundant variation in weights.\footnote{Chaisemartin2020 suggest the weight function expression for a DID with a binary treatment, which is an important special case of Theorem (ref). Appendix (ref) discusses the relationship between Theorem (ref) and their results.} Given the weighted-average interpretation provided by Theorem (ref), it is possible to define the weight difference and the marginal effect difference components comparably to ((ref)--(ref)), using $\omega_{Z}^{*}$ instead of $\omega_{Z}$ as the IV weight function.\footnote{It still requires Assumption (ref)--(ii) for the OLS coefficient to have an exact weighted-average interpretation. In this situation, a researcher may want to choose a more flexible specification in performing the OLS regression, e.g., controlling for group--time effects instead of additive group and time effects.} Note that the marginal weights on covariate $W=w$ and treatment level $X=x$ are given by \begin{align*} \overline{\omega}_{Z}^{*}(w) & =\int\omega_{Z}^{*}(x,w)dx=L_{w}(X\widetilde{Z})/E(X\widetilde{Z}),\\ \overline{\omega}_{Z}^{*}(x) & =\int\omega_{Z}^{*}(x,w)dF_{W}(w)=E(\mathbbm{1}_{X\ge x}\widetilde{Z})/E(X\widetilde{Z}). \end{align*} While the treatment-level weight $\overline{\omega}_{Z}^{*}(x)$ is identical to $\overline{\omega}_{Z}(x)$ defined in ((ref)), the covariate weight $\overline{\omega}_{Z}^{*}(w)$ has a different expression from $\overline{\omega}_{Z}(w)$ defined in ((ref)). \subsection{Unobserved Heterogeneity} Assumption (ref) rules out unobserved heterogeneity in the marginal effects of the treatment. This is a common but strong assumption. Because the distinction between observables and unobservables arises merely from data availability, it is natural to consider heterogeneity in both observed and unobserved dimensions in the econometric model. Allowing for unobserved heterogeneity in the marginal effects $Y'(x)$, I define the average marginal effect (AME) as $\tau(x,w)=E\left[Y'(x)|W=w\right]$. With Assumption (ref), the AME reduces to $\tau(x,w)=\frac{\partial}{\partial x}g(x,w)$. Removing Assumptions (ref), (ref)--(i), and (ref)--(i), I make the following set of assumptions about the potential outcome process $Y(x)$. \begin{assumptionx} The potential outcome process $Y(x)$ satisfies the following conditions: \begin{enumerate} • \textup{(Conditional Independence)} $Cov\left(Y(x),Z|W\right)=0$ for any $x\in(\underline{x},\overline{x})$; • \textup{(Differentiability)} $Y'(x)$ exists for any $x\in(\underline{x},\overline{x})$ and the total variation $V_{x_{0}}^{X}\left(Y(\cdot)\right)=\int_{\min\{x_{0},X\}}^{\max\{x_{0},X\}}\left|Y'(x)\right|dx$ satisfies $E\left[V_{x_{0}}^{X}\left(Y(\cdot)\right)^{2}\right]<\infty$ for some $x_{0}\in(\underline{x},\overline{x})$. \end{enumerate} \end{assumptionx} Assumption (ref)--(i) replaces (ref)--(i), and this is a standard exogeneity assumption.\footnote{This setting implicitly rules out any direct causal effects of the instrument on the potential outcome. Some studies explicitly specify the potential outcome $Y(x,z)$ to be a function of the potential instrument assignment $z$ and then assume $Y(x,z')=Y(x,z)$ for any $z\ne z'$.} Assumption (ref)--(ii) replaces (ref)--(i). Unobserved heterogeneity in the treatment effects arises when $Var\left(Y'(x)|W=w\right)>0$. To rule out negative weights, some studies in the IV literature specify the potential treatment process and assume it to be monotonic in the instrument. On the other hand, I do not impose the monotonicity condition and maintain Assumption (ref)--(ii) (i.e., $E(\widetilde{X}\widetilde{Z})\ne0$), thereby allowing for negative weights.\footnote{Appendix (ref) considers the weighted-average interpretation under the monotonicity condition and explores its relationship with the LATE interpretation in angrist1995two and the marginal treatment effect interpretation in heckman2006understanding.} The following theorem extends the weighted-average interpretation of the IV coefficient provided by Theorem (ref). \begin{thm} With Assumptions (ref), (ref), (ref)--(ii), and (ref)--(i), the IV coefficient $\beta_{IV}$ is given by \begin{align*} \beta_{IV} & =\int\int\tau_{IV}(x,w)\omega_{Z}(x,w)dF_{W}(w)dx, \end{align*} \textup{where }the IV-identified marginal effect $\tau_{IV}(x,w)$ at each $(x,w)$ is given by \[ \tau_{IV}(x,w)=E\left[Y'(x)\lambda\left(Y'(x)|x,w\right)|W=w\right] \] with $\lambda(t|x,w)=\frac{Cov\left(\mathbbm{1}_{X\ge x},Z|Y'(x)=t,W=w\right)}{Cov\left(\mathbbm{1}_{X\ge x},Z|W=w\right)}$ and $E\left[\lambda(Y'(x)|x,w)|W=w\right]=1$. \end{thm} Theorem (ref) implies that the IV coefficient is a weighted average of the causal effects $Y'(x)$. Although treatment levels $x$ and covariates $w$ are weighted exactly in the same manner as in Theorem (ref), unobserved heterogeneity influences how the effects $Y'(x)$ are weighted. In particular, the difference between the IV-identified marginal effect $\tau_{IV}(x,w)$ and the AME $\tau(x,w)$ can be written as \[ \tau_{IV}(x,w)-\tau(x,w)=Cov\left(Y'(x),\lambda\left(Y'(x)|x,w\right)|W=w\right). \] This difference arises from unobservable-driven covariance between the treatment effects $Y'(x)$ and the treatment responses to the instrument. For example, suppose some unobservables positively influence $Y'(x)$ and make the treatment more responsive to the instrument. As the weight function $\lambda(t|x,w)$ represents how strongly the instrument $Z$ induces a transition of the treatment $X$ from below $x$ to above $x$ conditional on $Y'(x)=t$ and $W=w$, $\tau_{IV}(x,w)>\tau(x,w)$ results from a greater emphasis on a higher $Y'(x)$. This difference corresponds to the discrepancy between the LATE and ATE in a binary treatment setting with no covariates, although the difference is specific to each treatment level $x$ and covariate value $w$ in this case. Given the weighted-average interpretation, the decomposition in ((ref)--(ref)) remains valid by replacing the marginal effect difference component in ((ref)) with \begin{equation} \Delta_{ME}=\int\int\left(\tau_{IV}(x,w)-\frac{\partial}{\partial x}m(x,w)\right)\omega_{Z}(x,w)dF_{W}(w)dx. \end{equation} However, this component no longer represents endogeneity bias alone. In particular, ((ref)) can be further decomposed as \begin{align*} \Delta_{ME} & =\int\int\left(\tau(x,w)-\frac{\partial}{\partial x}m(x,w)\right)\omega_{Z}(x,w)dF_{W}(w)dx\\ & +\int\int\left(\tau_{IV}(x,w)-\tau(x,w)\right)\omega_{Z}(x,w)dF_{W}(w)dx. \end{align*} The first term captures endogeneity bias through the difference between the AME $\tau(x,w)$ and the slope of the conditional mean function $\frac{\partial}{\partial x}m(x,w)$. The second term represents the unobservable-driven weight difference discussed above. In a nonbinary treatment setting, it is well known that the AME $\tau(x,w)$ or analogous causal objects cannot be identified without any restriction on unobserved heterogeneity or causal structure.\footnote{Examples of restrictions that may enable this identification are: one-dimensional unobservable in the first-stage relationship imbens2009identification; one-dimensional unobservable in the potential outcome chernozhukov2007instrumental; and the linear random coefficients model masten2016identification. One could estimate one of these models and recover the AME to further decompose $\Delta_{ME}$ into endogeneity bias and the unobservable-driven weight difference. Exploring the decomposition under these restrictions is beyond the scope of this paper.} Thus, endogeneity bias and the unobservable-driven weight difference cannot be identified separately in general. Nevertheless, it remains helpful to isolate the covariate and treatment-level weight difference components using the decomposition, rather than observing only the raw IV--OLS gap. \section{Estimation and Inference} \subsection{IV-Weighted OLS Estimators} Performing the decomposition proposed in Section (ref) requires estimators of the IV-weighted OLS coefficients defined in ((ref)--(ref)), which serve as intermediate points between the IV and OLS estimates. To define the estimators, let $(Y_{i},X_{i},Z_{i},W_{i})_{i=1}^{N}$ be an i.i.d. random sample that satisfies the set of assumptions in Section (ref). The most natural estimators for the IV-weighted OLS coefficients are their direct data counterparts: \begin{align} \widehat{\beta}_{C} & =\frac{1}{N}\sum_{i=1}^{N}\int\widehat{b}_{OLS}(W_{i})\widehat{\omega}_{Z}(x,W_{i})dx,\\ \widehat{\beta}_{CT} & =\frac{1}{N}\sum_{i=1}^{N}\int\frac{\partial}{\partial x}\widehat{m}(x,W_{i})\widehat{\omega}_{Z}(x,W_{i})dx, \end{align} using some estimators $(\widehat{b}_{OLS},\widehat{m},\widehat{\omega}_{Z})$ for $(b_{OLS},m,\omega_{Z}$). While there are various choices for these estimators $(\widehat{b}_{OLS},\widehat{m},\widehat{\omega}_{Z})$, I focus on the most practical setting. I assume that $\widehat{m}$ is a consistent estimator for $m$ estimated by minimizing $\frac{1}{N}\sum_{i=1}^{N}\left\{ Y_{i}-\widehat{m}(X_{i},W_{i})\right\} ^{2}$, and that $\widehat{m}(x,w)$ is given in a series form \begin{equation} \widehat{m}(x,w)=\sum_{k=1}^{K_{N}}\widehat{\alpha}_{k}(w)p_{k}(x),\,\,\,\widehat{\alpha}_{k}(w)=\sum_{\ell=1}^{L_{N}^{(k)}}\widehat{\theta}_{k\ell}q_{k\ell}(w), \end{equation} where $\{p_{k}(x),q_{k1}(w),\ldots,q_{kL_{N}^{(k)}}(w)\}_{k=1}^{K_{N}}$ is a set of basis functions chosen by the researcher. The numbers of the basis functions $K_{N}$ and $(L_{N}^{(k)})_{k=1}^{K_{N}}$ are constant in a parametric approach, while they may increase with the sample size $N$ with a nonparametric sieve estimation adopted. While $\widehat{b}_{OLS}$ can be any consistent estimator for $b_{OLS}$, for concreteness I assume that $\widehat{b}_{OLS}$ is also estimated by the series form ((ref)) with $K_{N}=2$ and $\left(p_{1}(x),p_{2}(x)\right)=\left(1,x\right)$, in which $\widehat{\alpha}_{2}(w)$ corresponds to $\widehat{b}_{OLS}(w)$. In addition, as the estimator for $\omega_{Z}(x,W_{i})$, I consider its sample analogue \[ \widehat{\omega}_{Z}(x,W_{i})=\frac{\widetilde{\mathbbm{1}}_{X_{i}\ge x}\widetilde{Z}_{i}}{\frac{1}{N}\sum_{j=1}^{N}\widetilde{\mathbbm{1}}_{X_{j}\ge x}\widetilde{Z}_{j}}. \] Then, ((ref)--(ref)) can be rewritten as \begin{align} \widehat{\beta}_{C} & =\frac{\sum_{i=1}^{N}\widehat{b}_{OLS}(W_{i})\widetilde{X}_{i}\widetilde{Z}_{i}}{\sum_{i=1}^{N}\widetilde{X}_{i}\widetilde{Z}_{i}},\\ \widehat{\beta}_{CT} & =\frac{\sum_{i=1}^{N}\left(\sum_{k=1}^{K_{N}}\widehat{\alpha}_{k}(W_{i})\widetilde{P}_{ik}\right)\widetilde{Z}_{i}}{\sum_{i=1}^{N}\widetilde{X}_{i}\widetilde{Z}_{i}}, \end{align} where $P_{ik}=p_{k}(X_{i})$. Therefore, the following two-step procedure gives the IV-weighted OLS estimates.\begin{description}[leftmargin=0em] • Estimate $\widehat{b}_{OLS}(w)$ and $\left\{ \widehat{\alpha}_{k}(w)\right\} _{k=1}^{K_{N}}$ using the least-squares method with the series specification ((ref)). • Regress $Y_{2i}=\sum_{k=1}^{K_{N}}\widehat{\alpha}_{k}(W_{i})\widetilde{P}_{ik}$ on $(X_{i},W_{i})$ instrumenting $X_{i}$ with $Z_{i}$, which yields $\widehat{\beta}_{CT}$ as the coefficient of $X_{i}$. To estimate $\widehat{\beta}_{C}$, let $Y_{2i}=\widehat{b}_{OLS}(W_{i})\widetilde{X}_{i}$ and perform the same regression. \end{description} Note that this two-step approach naturally generalizes the one proposed by lochner2015estimating by allowing for a more general functional form of $\widehat{m}(x,w)$. In fact, $\widehat{\beta}_{CT}$ given in Step 2 is identical to the reweighted OLS estimator in lochner2015estimating if $\widehat{m}(x,w)$ is specified in Step 1 to be additively separable in $x$ and $w$. \subsubsection*{DID- or RD-type Instrument Case} The setting considered in Section (ref) requires a small modification of the estimators due to a difference in the weight functions. Suppose that the weight function $\omega_{Z}^{*}(x,W_{i})$ defined in Theorem (ref) is estimated by its sample analogue \[ \widehat{\omega}_{Z}^{*}(x,W_{i})=\frac{\widehat{L}_{W_{i}}\left(\mathbbm{1}_{X_{i}\ge x}\widetilde{Z}_{i}\right)}{\frac{1}{N}\sum_{j=1}^{N}\mathbbm{1}_{X_{j}\ge x}\widetilde{Z}_{j}}, \] where $\widehat{L}_{w}\left(R_{i}\right)=w'\left(\sum_{j=1}^{N}W_{j}W_{j}'\right)^{-1}\left(\sum_{j=1}^{N}W_{j}R_{j}\right)$ is the linear projection estimate. Using $\widehat{\omega}_{Z}^{*}$ instead of $\widehat{\omega}_{Z}$ in deriving ((ref)--(ref)) yields \begin{align} \widehat{\beta}_{C} & =\frac{\sum_{i=1}^{N}\widehat{L}_{W_{i}}\left(\widehat{b}_{OLS}(W_{i})\right)X_{i}\widetilde{Z}_{i}}{\frac{1}{N}\sum_{i=1}^{N}\widetilde{X}_{i}\widetilde{Z}_{i}},\\ \widehat{\beta}_{CT} & =\frac{\sum_{i=1}^{N}\left(\sum_{k=1}^{K_{N}}\widehat{L}_{W_{i}}\left(\widehat{\alpha}_{k}(W_{i})\right)P_{ik}\right)\widetilde{Z}_{i}}{\frac{1}{N}\sum_{i=1}^{N}\widetilde{X}_{i}\widetilde{Z}_{i}}. \end{align} These expressions differ from ((ref)--(ref)) in two ways. First, $\widehat{b}_{OLS}(W_{i})$ and $\widehat{\alpha}_{k}(W_{i})$ are linearly projected onto $W_{i}$. In practice, $\widehat{b}_{OLS}(W_{i})$ and $\widehat{\alpha}_{k}(W_{i})$ may be specified to be linear at the outset, rather than estimating them from a flexible nonlinear specification first and then linearly predicting them. Second, $X_{i}$ and $P_{ik}$ instead of $\widetilde{X}_{i}$ and $\widetilde{P}_{ik}$ enter the numerators. These differences slightly change Step 2 in estimating the IV-weighted OLS estimates as follows.\begin{description}[leftmargin=0em] • Regress $Y_{2i}=\sum_{k=1}^{K_{N}}\widehat{L}_{W_{i}}\left(\widehat{\alpha}_{k}(W_{i})\right)P_{ik}$ on $(X_{i},W_{i})$ instrumenting $X_{i}$ with $Z_{i}$, which yields $\widehat{\beta}_{CT}$ as the coefficient of $X_{i}$. To estimate $\widehat{\beta}_{C}$, let $Y_{2i}=\widehat{L}_{W_{i}}\left(\widehat{b}_{OLS}(W_{i})\right)X_{i}$ and perform the same regression.\end{description} \subsection{Asymptotic Properties} Asymptotic properties of the IV-weighted OLS estimator can be derived using the standard econometric results for a two-step estimator. The following discussion focuses on the case in which the first step is parametric, since ackerberg2012practical illustrate that a semiparametric two-step estimator that uses a series approximation in the first step can be treated as if it were a parametric estimator for the purpose of standard error computation. Using the standard formula for a parametric two-step estimator newey1994large, Appendix (ref) derives the asymptotic equivalence: \begin{equation} \sqrt{N}\left(\widehat{\beta}_{CT}-\beta_{CT}\right)\underset{p}{\to}\frac{\frac{1}{\sqrt{N}}\sum_{i=1}^{N}\left(v_{1i}\widehat{\widetilde{Z}_{i}}+v_{2i}\widetilde{Z}_{i}\right)}{E\left[\widetilde{X}_{i}\widetilde{Z}_{i}\right]}, \end{equation} where $v_{1i}=Y_{i}-\sum_{k=1}^{K}\alpha_{k}(W_{i})P_{ik}$ is a residual from the first step, $v_{2i}=\widetilde{Y}_{2i}-\beta_{CT}\widetilde{X}_{i}$ is a residual from the second step, and $\widehat{\widetilde{Z}_{i}}$ is a predicted value of $\widetilde{Z}_{i}$ given by the first step using $\widetilde{Z}_{i}$ instead of $Y_{i}$ as the dependent variable.\footnote{Another correction term appears in the case with DID- or RD-based instruments if $\widehat{b}_{OLS}(W_{i})$ and $\widehat{\alpha}_{k}(W_{i})$ are not specified to be linear, as presented in Appendix (ref).} Most statistical packages can estimate the standard error of $\widehat{\beta}_{CT}$ by estimating the standard error of the right-hand side of ((ref)) under a certain distributional assumption (heteroscedasticity-robust, clustered, etc.). For example, in an i.i.d. heteroscedastic case, the asymptotic variance of $\sqrt{N}\left(\widehat{\beta}_{CT}-\beta_{CT}\right)$ is given by $V_{\beta_{CT}}=E\left[\widetilde{X}_{i}\widetilde{Z}_{i}\right]^{-1}E\left[(v_{1i}\widehat{\widetilde{Z}_{i}}+v_{2i}\widetilde{Z}_{i})^{2}\right]$ and the standard error of $\widehat{\beta}_{CT}$ can be estimated using the sample analogue of $\sqrt{V_{\beta_{CT}}/N}$. Estimating the standard error of $\widehat{\beta}_{C}$ can follow the same procedure, as it is a special case with $K=2$. \subsection{Testing the Treatment Endogeneity} Testing the significance of the marginal effect difference component $\Delta_{ME}=\beta_{IV}-\beta_{CT}$ serves as a generalized Durbin--Wu--Hausman (DWH) test that is robust to the nonlinearity and observed heterogeneity of the treatment effects. This test further extends the generalized DWH test proposed in lochner2015estimating by allowing for observed heterogeneity of the effects. Since $\widehat{\beta}_{IV}$ and $\widehat{\beta}_{CT}$ have a common denominator $\frac{1}{N}\sum_{i=1}^{N}\widetilde{X}_{i}\widetilde{Z}_{i}$, a relevant test statistic is \begin{equation} \widehat{T}=\frac{\frac{1}{N}\sum_{i=1}^{N}d_{i}\widetilde{Z}_{i}}{\widehat{S}}, \end{equation} where $d_{i}=\widetilde{Y}_{i}-\widetilde{Y}_{2i}$ is the difference in (residualized) dependent variables between two regressions that yield $\widehat{\beta}_{IV}$ and $\widehat{\beta}_{CT}$. $\widehat{S}$ is the standard error of the numerator, $\frac{1}{N}\sum_{i=1}^{N}d_{i}\widetilde{Z}_{i}$. For example, in an i.i.d. heteroscedastic case, \begin{equation} N\widehat{S}^{2}=\frac{1}{N}\sum_{i=1}^{N}\left(d_{i}\widetilde{Z}_{i}-\frac{1}{N}\sum_{j=1}^{N}d_{j}\widetilde{Z}_{j}-v_{1i}\widehat{\widetilde{Z}_{i}}\right)^{2}. \end{equation} Under a more general distributional assumption, the right hand side of ((ref)) should be replaced by the estimated asymptotic variance of $\frac{1}{\sqrt{N}}\sum_{i=1}^{N}\left(d_{i}\widetilde{Z}_{i}-\frac{1}{N}\sum_{j=1}^{N}d_{j}\widetilde{Z}_{j}-v_{1i}\widehat{\widetilde{Z}_{i}}\right)$. Appendix (ref) shows that $\widehat{T}$ converges in distribution to $N(0,1)$ under the null hypothesis $\Delta_{ME}=0$ and diverges under the alternative hypothesis $\Delta_{ME}\ne0$.\footnote{As described in footnote (ref), in a setting with multiple instruments $(Z_{i1},\ldots,Z_{iM})$, the scalar instrument $Z_{i}$ is generated as $Z_{i}=\sum_{m=1}^{M}\pi_{m}Z_{im}$. In this case, instead of performing the $t$-test with the synthetic $Z_{i}$, one could separately compute $\frac{1}{N}\sum_{i=1}^{N}d_{i}\widetilde{Z}_{im}$ for each $m=1,\ldots,M$ and perform a chi-squared test to improve efficiency.} Three limitations of this test are worth noting. First, it is a valid test of endogeneity only when the setup in Section (ref) describes the true model. Most importantly, the marginal effect difference $\Delta_{ME}$ cannot be attributed to endogeneity bias alone in the presence of unobserved heterogeneity, as discussed in Section (ref).\footnote{In a binary treatment context, one might consider the frameworks such as donald2014testing and mogstad2018using, which can test endogeneity bias even with unobserved heterogeneity under several additional conditions.} Second, this test cannot detect endogeneity bias if the difference between the IV-identified and OLS-identified marginal effects $\frac{\partial}{\partial x}g(x,w)-\frac{\partial}{\partial x}m(x,w)$ in some regions of $(x,w)$ exactly cancels out the difference in other regions. While the direction of endogeneity bias is expected to be unambiguous in many economic contexts, one might consider the nonparametric test proposed by blundell2007non in the contexts in which the direction of endogeneity may differ across $(x,w)$. Finally, this test shares the fundamental limitations with the standard DWH test that the instrument must be valid and that it has little power to detect endogeneity bias if the instrument is not sufficiently strong.\footnote{The asymptotic property of the test itself would not be influenced by the weak instrument problem, since the test statistic in ((ref)) does not depend on the first stage coefficient, $\frac{1}{N}\sum_{i=1}^{N}\widetilde{X}_{i}\widetilde{Z}_{i}$. staiger1997instrumental show that one version of the DWH test Durbin1954 is robust to the weak instrument problem but two others Wu1973,hausman1978specification are not.} \section{Applications to Return-to-Schooling Estimates} This section describes the empirical applications of my framework to return-to-schooling estimates with three different identification strategies.\footnote{I use the term “returns to schooling” to refer to the causal effect of schooling on log wages, even though it sometimes denotes the internal rate of return associated with schooling in its most narrow sense.} First, I use geographic variation in college costs as an instrument as in cameron2004estimation and carneiro2011estimating and estimate the returns to schooling in the National Longitudinal Survey of Youth 1979 (NLSY79). Next, I exploit a discontinuity in the minimum school-leaving age in the United Kingdom as in oreopoulos2006estimating, using the British General Household Survey (GHS). Finally, I exploit DID variation in compulsory schooling laws across cohorts and states in the United States as in \citet*{acemoglu2000large} and estimate the returns to schooling in the 1960--1980 U.S. Censuses. Appendix (ref) provides additional details for these analyses. \subsection{College Cost Instrument with Geographic Variation} This analysis uses the civilian sample from the NLSY79. The original sample consists of 5,579 males and 5,827 females born in 1957--64. After dropping persons with missing information about their Armed Forces Qualification Test (AFQT) score or county of residence at age 14, persons who have not completed 8th grade by age 22, and persons with no wage information in any year, my sample consists of 4,719 males and 4,986 females. For each person, I use observations between ages 25 and 54. The outcome $Y$ is the log hourly wage at the current or most recent job and the treatment $X$ is years of schooling top-coded at 18 years. Because minority and economically disadvantaged households are sampled at higher rates, persons are weighted by the sampling weights throughout my analysis. Within each person, I use equal weights for the person--year observations.\footnote{As a result, the weight on a person--year observation is the sampling weight for the person divided by the number of years for which the person is observed.} It is important to weight observations appropriately to recover the population OLS and IV coefficients because my framework does not assume the linear structural equation and admits a weighted-average interpretation of the coefficients. As instruments, I use measures of direct and opportunity costs of college attendance as in cameron2004estimation and carneiro2011estimating.\footnote{While I follow their identification strategies in estimating the causal effects of schooling, I do not follow their empirical specifications in my analysis. Their original analyses do not exactly fit into my framework because years of schooling is not the only endogenous variable in cameron2004estimation and the schooling measure is binary in carneiro2011estimating.} In particular, I use the presence of a public four-year college and tuition rate of the nearest in-state public four-year college in the county of residence at age 14 to represent the direct cost of attendance.\footnote{In earlier work, card1995using and kane1995labor use college proximity as the schooling instrument associated with the direct costs of college attendance. They do not use opportunity cost measures in their analyses.} While the tuition rate captures a pecuniary cost of attendance, college proximity captures both pecuniary (due to reduced costs of room and board) and nonpecuniary costs of attendance. I use local earnings and unemployment rate in the county of residence at age 14 in the year in which the person turns age 17 (1974 for the oldest cohort and 1981 for the youngest) to represent the opportunity cost of attendance.\footnote{Local earnings are measured at the county level, while the unemployment rate is at the state level.} For the vector of covariates $W$, I use the AFQT percentile, age, an indicator for female, black and Hispanic dummies, indicators for parental education, the number of siblings, and cohort dummies.\footnote{The specification allows the AFQT scores to have different slopes across tertiles and the ages to have different slopes across the 25--34, 35--44, and 45--54 age groups.}In addition, I include urban status, Census division dummies, and the average local earnings and unemployment rate during 1974--81 of the county of residence at age 14 in the covariate vector $W$. Using the average local labor market conditions as control variables ensures that variation in the corresponding instruments is driven by a temporary shock to the local labor market. Otherwise, the instruments may capture a permanent difference in local labor market conditions, which can be directly associated with potential earnings. Controlling for local labor market conditions is also important to ensure the plausibility of the college proximity instrument, as emphasized by cameron2004estimation. Appendix (ref) provides further details of the sample and presents the first-stage regression results. Table (ref) reports the decomposition results. The first three columns of the table report the linear OLS coefficient, the linear IV coefficient, and the IV--OLS coefficient gap. The next three columns report the estimates of the covariate weight difference, the treatment-level weight difference, and the marginal effect difference. Here, I discuss the first row of the table, which reports the results from the NLSY79. The second and third rows are discussed in Sections (ref) and (ref). In the first row, the point estimates of the OLS and IV coefficients are nearly identical, with the OLS coefficient of 0.065 and the IV coefficient of 0.062. An empirical researcher may be tempted to conclude from this result that there is no evidence of ability bias in these data. However, the decomposition using the IV-weighted OLS coefficients indicates that the IV coefficient would be well below the OLS coefficient if they had the same weights on the covariates and treatment levels. In fact, the covariate weight difference is estimated to be 0.011 and the treatment-level weight difference is estimated to be 0.018. With the weight difference components accounted for, the marginal effect difference is --0.032, indicating that the IV-identified returns to schooling are lower than the OLS-identified returns. This result is rather consistent with the ability bias story in terms of point estimates, even though the generalized DWH test fails to reject at the 5% level given the large standard error. I investigate the mechanisms underlying these results by examining the patterns of the IV and OLS weights. Table (ref) presents the total IV and OLS weights on each group of covariates. The first three columns of the table present the population share, the total IV weight, and the total OLS weight on each group. Each set of weights sums to one across the whole sample by construction. The last column reports the OLS schooling coefficient from a linear regression performed separately for each group, to illustrate the difference in OLS-identified schooling effects across groups. While the OLS weights are close to the population shares, the IV weights are concentrated on persons with advantaged backgrounds in terms of AFQT score, parental education, and race/ethnicity. Schooling coefficients from the separately performed OLS regressions indicate that more-advantaged groups tend to have higher schooling effects. These data patterns result in the positive contribution of the covariate weight difference to the IV--OLS coefficient gap. The pattern of IV weights is consistent with the empirical observation by cameron2004estimation that persons with advantaged backgrounds tend to be more sensitive to local college availability in the NLSY79. However, one may expect that persons with disadvantaged backgrounds should be weighted more because they are expected to be more sensitive to college cost instruments given their financial constraints.\footnote{card1995using and kling2001interpreting find that persons from less advantaged backgrounds are more sensitive to the presence of a local college in the sample from the National Longitudinal Survey of Young Men, which is based on older cohorts than the NLSY79.} Several factors can explain low IV weights on persons with disadvantaged backgrounds. While cameron2004estimation suggest that increased funding on federal student aid programs in the 1970s is one such factor, low college attendance and graduation rates among persons from disadvantaged backgrounds can also be important.\footnote{The share of persons with one or more and four or more years of college education are 19% and 4%, respectively, among the bottom third of AFQT scores. Among persons in the top third of AFQT scores, 80% have one or more years of college education and 56% have four or more years of college education.} A low college attendance rate implies that the majority would be never-takers instead of compliers.\footnote{If the college enrollment decision is explained by a logit or probit model, the share of compliers is the largest among the group with a college attendance rate of 50%.} A low college graduation rate implies that their completed years of schooling would not be strongly affected, even if the instruments affect their college attendance decisions. Panel (a) of Figure (ref) illustrates the total IV and OLS weights on each treatment margin. The IV weights are mostly placed on college education margins, with the weights on high school margins close to zero. This weight pattern is consistent with the expectation that college cost instruments affect years of schooling through college attendance decisions. Higher IV weights on college education margins result in the positive contribution of the weight difference to the IV--OLS coefficient gap because marginal effects of years in college are much higher than marginal effects of years in high school in this sample. In fact, the linear OLS coefficient is 0.031 in the subsample with 12 or fewer years of schooling and 0.071 in the subsample with 12 or more years of schooling. The overall result from this decomposition exercise that the IV coefficient is inflated due to the weight difference appears to match the discount rate bias argument developed by lang1993ability and card1995earnings.\footnote{Building on the canonical model of becker1967human, they consider a model in which individuals invest in education as long as the marginal return to an additional year of schooling exceeds its marginal cost. The model predicts that the marginal return at the chosen schooling level would be higher for credit-constrained individuals with higher discount rates. Given this prediction, they argue that the IV coefficient could exceed the population average return if credit-constrained individuals are more sensitive to the instruments and thus weighted more.} However, the estimated weight patterns indicate that the underlying mechanism is distinct in two ways. First, the IV coefficient places more weight on advantaged rather than disadvantaged groups in the population. Higher IV weights on advantaged groups give rise to the higher IV coefficient because of the higher marginal returns among advantaged groups. This is the opposite mechanism to the discount rate bias argument, even though it shifts the IV coefficient in the same direction. Second, higher marginal returns to college education give rise to the higher IV coefficient in this sample, given the concentration of IV weights on college years relative to high school years. The discount rate bias argument does not consider this possibility, as it assumes nonincreasing marginal returns to schooling. \subsection{RD-Based Compulsory Schooling Instrument} As in oreopoulos2006estimating, I now exploit the 1947 compulsory schooling reform in the United Kingdom to construct an RD instrument. The U.K. government raised the minimum school-leaving age in Great Britain from 14 to 15 years in 1947. The share of people leaving school at age 14 or earlier then fell from 56% in the 1932 birth cohort turning 14 one year before the reform to 9% in the 1934 birth cohort turning 14 one year after the reform. The empirical strategy in this analysis closely follows oreopoulos2006estimating.\footnote{My regression results slightly differ from the originally published results in oreopoulos2006estimating due to the data correction oreopoulos2008estimating and a top-coding treatment of schooling described below.} I use the sample of persons younger than 65 years old from the British General Household Surveys in 1984--98 who turned 14 in 1935--65. I exclude persons with missing data on earnings or education and persons leaving school before age 10. The outcome $Y$ is log annual earnings. The treatment $X$ is years of schooling, which is given by the age when they left full-time education minus five years, with the top-coding at 20 years.\footnote{Approximately 2.5% of persons report having left full-time education after age 25, and their schooling levels are all treated as 20 years. oreopoulos2006estimating does not make this top-coding treatment. See Appendix (ref) for the analysis without top-coding, which reaches the same conclusion regarding the relevance of the weight difference components.} The instrument $Z$ is an indicator for 1933 or later birth cohorts, who turned 14 in 1947 or later. The covariates $W$ are the fourth-order polynomials of birth cohort and age.\footnote{gelman2019high recommend against the use of high-order polynomials in RD designs; nevertheless, I follow the original specification in oreopoulos2006estimating. Using two separate quadratic polynomials for pre- and post-reform cohorts instead of global fourth-order polynomials slightly pushes up the IV estimate. However, it does not affect the finding that the weight difference components are important.} The second row of Table (ref) reports the decomposition of the IV--OLS coefficient gap in this empirical application. As the OLS estimate lies above the IV estimate by 0.021, a researcher who presumes the linear causal model may immediately interpret it as the result of ability bias. However, adjusting for the weight difference suggests otherwise. In fact, the decomposition result indicates that the marginal returns identified by the IV coefficient exceed those identified by the OLS coefficient by 0.023, after accounting for the covariate weight difference and the treatment-level weight difference components. The empirical result is no longer consistent with the ability bias story as a point estimate, although both standard and generalized DWH tests fail to reject at the 5% level due to the imprecise IV estimate. Table (ref) presents the IV and OLS weights on the birth cohorts. The IV weights are concentrated on the birth cohorts turning age 14 in 1941--50. The linear regression restricted to these cohorts yields smaller OLS schooling coefficients than those for the younger cohorts, who receive most of the OLS weights. This explains the negative contribution of the covariate weight difference to the IV--OLS coefficient gap. Panel (b) of Figure (ref) shows that the IV weights are almost exclusively placed on the 10th year of schooling, which is exactly what the 1947 reform mandated. This weight pattern pushes down the IV coefficient because the 10th year of schooling has a smaller marginal return than the other schooling margins. In fact, the linear regression restricted to individuals with 9 or 10 years of schooling yields an OLS schooling coefficient of 0.037, which is much smaller than the full-sample OLS coefficient. This analysis demonstrates that researchers should exercise caution in extending the intuition of imbens1994identification to an empirical setting with covariates and a multivalued treatment. Observing a sharp decline in the dropout rate at age 14 after the reform, oreopoulos2006estimating argues that the IV and OLS coefficients in this setting are expected to identify the treatment effects for the comparable population, providing an analogy to the comparison between the LATE and ATE. In principle, however, the LATE interpretation in imbens1994identification draws on a model with a binary treatment and no covariates. It is still plausible that the LATE and ATE in this setting match \emph{conditional on} the schooling level and the birth cohort, judging by the extensive response to the reform. Nevertheless, the estimated patterns of the weights suggest that the IV and OLS coefficients identify the effects for entirely different birth cohorts at completely distinct schooling margins. \subsection{Compulsory Schooling Instrument with DID Variation} This analysis exploits DID variation in compulsory schooling laws across cohorts and regions, following acemoglu2000large.\footnote{My specification follows one of their main specifications in estimating private returns to schooling that uses the 1960--80 Census data with the child labor laws instrument and no state-of-residence controls acemoglu2000large. While they adjust some variables in the 1960--80 data to incorporate the 1950 Census data in their alternative specifications, I do not make these adjustments as I focus only on the 1960--80 Censuses. See Appendix (ref) for details.} The analysis sample consists of 40--49-year-old white males born in the United States from the 1960--80 Censuses. I use log weekly earnings as the outcome $Y$, and limit my analysis to persons with positive earnings and working for at least one week in the previous year. I use years of schooling as the treatment $X$. I include in a vector of covariates $W$ indicators for the state of birth and indicators for the year of birth. As instruments, I use the status of compulsory schooling laws (CSL) at age 14 in the state of birth. As in acemoglu2000large, I construct the CSL instruments based on the required years of schooling associated with child labor laws. Appendix (ref) provides further details of the sample, first-stage regression results, and decomposition results using other common compulsory schooling instruments. The third row of Table (ref) presents the estimated OLS and IV schooling coefficients and the decomposition of the IV--OLS coefficient gap. While the IV estimate is slightly above the OLS estimate with a coefficient gap of 0.017, the decomposition result indicates that the gap is primarily associated with the weight difference components. Although the IV--OLS gap is not statistically distinguishable from zero even without accounting for the weight difference components, this result reshapes the quantitative implication from the gap. In particular, acemoglu2000large attribute the small positive IV--OLS gap to modest external returns to schooling. Accounting for the weight difference components, this result indicates even less important externality than their original interpretation. Table (ref) presents the total IV and OLS weights on each group of covariates. The IV weights attached to some birth cohorts or birth states are negative. In fact, the IV weights aggregate to negative values among persons in the 1930--39 birth cohorts and among persons born in the Midwest or West. Moreover, I find that 54% of covariate-specific IV weights are negative and sum to --5.93. In a usual IV setting, negative weights imply the presence of both compliers and defiers. In a DID setting, however, the weights are not guaranteed to be positive even with the perfect compliance $X=Z$, as pointed out by Chaisemartin2020 in the binary treatment case.\footnote{In fact, the first stage regression indicates the positive relationship between the CSL requirements and schooling levels, even in the subsample of the 1930--39 birth cohorts or among persons born in the Midwest or West.} While the IV-weighted OLS coefficient $\beta_{C}=\int b_{OLS}(w)\overline{\omega}_{Z}(w)dF_{W}(w)$ is a weighted average of the covariate-specific OLS coefficients $b_{OLS}(w)$, negative weights can push $\beta_{C}$ out of the support of $b_{OLS}(w)$. In fact, the estimates of $b_{OLS}(w)$ are no greater than $0.076$, despite $\beta_{C}$ being estimated to be $0.079$. This result suggests that even a small heterogeneity in treatment effects can give rise to a large contribution of the weight difference to the IV--OLS coefficient gap if the IV strategy relies on DID variation. Panel (c) of Figure (ref) reports the total IV and OLS weights for each schooling level. The IV weights mostly capture the schooling margins up to the 12th year of schooling, which is consistent with the context that primary and secondary education is mandated by the CSL. However, the effect of this treatment-level weight difference on the IV--OLS coefficient gap is mostly obscured by the large contribution of the covariate weight difference. \section{Conclusion} When OLS and IV estimates differ, empirical researchers typically consider two explanations. The first takes the linear regression equation literally and interprets the coefficient gap as endogeneity bias. The second extends the intuition of the LATE interpretation imbens1994identification to a general regression equation and interprets the coefficient gap as the weight difference. My paper enables researchers to proceed a step further and formally quantify the contributions of the weight difference and endogeneity bias components separately. I show that the IV--OLS coefficient gap is explained by differences in the weights on the covariates, the weights on the treatment levels, and the identified marginal effects. The marginal effect difference component captures endogeneity bias in the absence of the unobservable-driven interaction between the treatment effects and treatment responses to the instrument. I propose a simple two-step regression approach to perform the decomposition empirically, which can be implemented in standard statistical packages. I demonstrate the practical value of my framework through its empirical applications to return-to-schooling estimates with compulsory schooling and college cost instruments. The IV--OLS coefficient gaps in these empirical applications are substantially influenced by the weight difference components, and accounting for them leads to different conclusions about the direction or extent of endogeneity bias. \begingroup \setlength\bibitemsep{0.25em} \phantomsection \printbibliography[heading=bibintoc] \endgroup \setcounter{secnumdepth}{0} \phantomsection \section[Tables and Figures] \begin{table}[H] \caption{Decomposition of the IV--OLS Gap in Return-to-Schooling Estimates} \begin{threeparttable} \begin{centering} \begin{tabular}{cccccccccc} & & & & & & & & & \tabularnewline \hline \multirow{2}{*}{IV Strategy} & \multirow{2}{*}{Data} & & \multicolumn{3}{c}{Coefficients} & & \multicolumn{3}{c}{Decomposition}\tabularnewline \cline{4-6} \cline{5-6} \cline{6-6} \cline{8-10} \cline{9-10} \cline{10-10} & & & OLS & IV & IV--OLS & & $\Delta_{CW}$ & $\Delta_{TW}$ & $\Delta_{ME}$\tabularnewline \cline{1-2} \cline{2-2} \cline{4-6} \cline{5-6} \cline{6-6} \cline{8-10} \cline{9-10} \cline{10-10} \noalign{\vskip0.25em} College Cost & \multirow{2}{*}{NLSY79\tnote{1)}} & & 0.065 & 0.062 & --0.004 & & 0.011 & 0.018 & --0.032\tabularnewline Variation & & & (0.003) & (0.087) & (0.087) & & (0.011) & (0.010) & (0.086)\tabularnewline[0.5em] Compulsory & \multirow{2}{*}{British GHS\tnote{2)}} & & 0.084 & 0.062 & --0.021 & & --0.016 & --0.029 & 0.023\tabularnewline Schooling RD & & & (0.002) & (0.083) & (0.082) & & (0.009) & (0.018) & (0.078)\tabularnewline[0.5em] Compulsory & \multirow{2}{*}{U.S. Census\tnote{3)}} & & 0.067 & 0.084 & 0.017 & & 0.011 & 0.003 & 0.003\tabularnewline Schooling DID & & & (0.0004) & (0.022) & (0.022) & & (0.004) & (0.003) & (0.021)\tabularnewline[0.25em] \hline \end{tabular} \end{centering} \begin{tablenotes} • Notes: Standard errors are in parentheses. The first three columns report the OLS estimates, the IV estimates, and their gaps. The next three columns report the estimates of the covariate weight difference, the treatment-level weight difference, and the marginal effect difference components. By construction, these three components sum to the IV--OLS gap. Appendix (ref) describes the empirical specification for estimating the decomposition. • Standard errors are robust to heteroskedasticity and correlation across observations on persons living in the same county at age 14. The instruments are college cost measures as defined in the main text. • Standard errors are robust to heteroskedasticity and correlation across observations on the same birth cohort and survey year. The instrument is an indicator for turning age 14 in 1947 or later. • Standard errors are robust to heteroskedasticity and correlation across observations on the same state and year of birth. The instruments are indicators for compulsory schooling requirements implied by child labor laws (7, 8, and 9 or more years). \end{tablenotes} \end{threeparttable} \end{table} \begin{table}[H] \begin{threeparttable} \caption{The IV and OLS Weights on the Covariate Groups (NLSY79)} \begin{centering} \begin{tabular}{cccccc} & & & & & \tabularnewline \hline \multirow{2}{*}{Variable} & \multirow{2}{*}{Group} & Group & OLS & IV & Subsample\tabularnewline & & share & weight & weight & OLS coef.\tabularnewline \hline \multirow{3}{*}{\begin{tabular}{@c@}AFQT \\ percentile\end{tabular}} & 0--1/3 & 0.32 & 0.25 (0.01) & --0.08 (0.16) & 0.056 (0.006)\tabularnewline & 1/3--2/3 & 0.34 & 0.35 (0.01) & 0.19 (0.13) & 0.062 (0.005)\tabularnewline & 2/3--1 & 0.34 & 0.40 (0.01) & 0.89 (0.20) & 0.074 (0.005)\tabularnewline \hline Highest & Some HS of less & 0.24 & 0.21 (0.01) & --0.12 (0.15) & 0.055 (0.005)\tabularnewline parental & HS graduate & 0.42 & 0.41 (0.01) & 0.66 (0.17) & 0.069 (0.005)\tabularnewline education & Some college & 0.34 & 0.37 (0.01) & 0.46 (0.15) & 0.068 (0.005)\tabularnewline \hline & Black & 0.14 & 0.14 (0.02) & --0.12 (0.11) & 0.069 (0.005)\tabularnewline Race/Ethnicity & Hispanic & 0.06 & 0.06 (0.01) & 0.02 (0.03) & 0.064 (0.006)\tabularnewline & Other & 0.80 & 0.80 (0.02) & 1.10 (0.11) & 0.064 (0.004)\tabularnewline \hline \multirow{2}{*}{Sex} & Male & 0.50 & 0.49 (0.01) & 0.47 (0.17) & 0.057 (0.004)\tabularnewline & Female & 0.50 & 0.51 (0.01) & 0.53 (0.17) & 0.074 (0.005)\tabularnewline \hline Urban residence & Rural & 0.31 & 0.30 (0.03) & 0.53 (0.17) & 0.073 (0.006)\tabularnewline at age 14 & Urban & 0.69 & 0.70 (0.03) & 0.47 (0.17) & 0.062 (0.004)\tabularnewline \hline & Northeast & 0.22 & 0.22 (0.04) & 0.18 (0.14) & 0.061 (0.007)\tabularnewline Region & Midwest & 0.31 & 0.30 (0.04) & 0.51 (0.17) & 0.070 (0.007)\tabularnewline at age 14 & South & 0.32 & 0.32 (0.04) & 0.09 (0.13) & 0.058 (0.005)\tabularnewline & West & 0.15 & 0.16 (0.03) & 0.23 (0.15) & 0.071 (0.005)\tabularnewline \hline \end{tabular} \end{centering} \begin{tablenotes} • Notes: Standard errors are in parentheses and robust to heteroskedasticity and correlation across observations on persons living in the same county at age 14. Each set of weights sums to one across the whole sample. Subsample OLS coefficient of years of schooling is obtained with the same set of control variables as regressions in Table (ref). Appendix (ref) describes the empirical specification for estimating the weights. \end{tablenotes} \end{threeparttable} \end{table} \begin{table}[H] \caption{The IV and OLS Weights on the Covariate Groups (British GHS)} \begin{threeparttable} \begin{centering} \begin{tabular}{ccccc} & & & & \tabularnewline \hline \multirow{2}{*}{Year at 14} & Group & OLS & IV & Subsample\tabularnewline & share & weights & weights & OLS coef.\tabularnewline \hline 1935--40 & \multirow{2}{*}{0.04} & 0.03 & --0.09 & 0.054\tabularnewline & & (0.01) & (0.05) & (0.012)\tabularnewline 1941-45 & \multirow{2}{*}{0.09} & 0.08 & 0.68 & 0.065\tabularnewline & & (0.01) & (0.15) & (0.005)\tabularnewline 1946--50 & \multirow{2}{*}{0.14} & 0.12 & 0.49 & 0.079\tabularnewline & & (0.02) & (0.06) & (0.006)\tabularnewline 1951--55 & \multirow{2}{*}{0.19} & 0.18 & --0.10 & 0.086\tabularnewline & & (0.02) & (0.17) & (0.004)\tabularnewline 1956--60 & \multirow{2}{*}{0.25} & 0.26 & 0.05 & 0.088\tabularnewline & & (0.03) & (0.07) & (0.003)\tabularnewline 1961--65 & \multirow{2}{*}{0.28} & 0.34 & --0.02 & 0.088\tabularnewline & & (0.03) & (0.05) & (0.003)\tabularnewline \hline \end{tabular} \end{centering} \begin{tablenotes} • Notes: Standard errors (in parentheses) are robust to heteroskedasticity and correlation across observations on the same birth cohort and survey year. Each set of weights sums to one across the whole sample. Subsample OLS coefficient of years of schooling is obtained with controlling for quartic terms of ages and birth cohorts. Appendix (ref) describes the empirical specification for estimating the weights. \end{tablenotes} \end{threeparttable} \end{table} \begin{table}[H] \caption{The IV and OLS Weights on the Covariate Groups (U.S. Census)} \begin{threeparttable} \begin{centering} \begin{tabular}{cccccc} & & & & & \tabularnewline \hline \multirow{2}{*}{Variable} & \multirow{2}{*}{Group} & Group & OLS & IV & Subsample\tabularnewline & & share & weights & weights & OLS coef.\tabularnewline \hline & 1910--19 & \multirow{2}{*}{0.32} & 0.33 & 1.56 & 0.063\tabularnewline & & & (0.02) & (0.25) & (0.001)\tabularnewline Year of & 1920--29 & \multirow{2}{*}{0.35} & 0.36 & 1.25 & 0.070\tabularnewline birth & & & (0.02) & (0.28) & (0.001)\tabularnewline & 1930--39 & \multirow{2}{*}{0.33} & 0.31 & --1.81 & 0.067\tabularnewline & & & (0.02) & (0.38) & (0.001)\tabularnewline \hline & Northeast & \multirow{2}{*}{0.29} & 0.26 & 0.24 & 0.069\tabularnewline & & & (0.02) & (0.37) & (0.001)\tabularnewline & Midwest & \multirow{2}{*}{0.33} & 0.27 & --1.32 & 0.066\tabularnewline \multirow{2}{*}{\begin{tabular}{@c@}Region of \\ birth\end{tabular}} & & & (0.02) & (0.34) & (0.001)\tabularnewline & South & \multirow{2}{*}{0.30} & 0.39 & 2.95 & 0.068\tabularnewline & & & (0.02) & (0.41) & (0.001)\tabularnewline & West & \multirow{2}{*}{0.09} & 0.08 & --0.87 & 0.064\tabularnewline & & & (0.01) & (0.15) & (0.001)\tabularnewline \hline \end{tabular} \end{centering} \begin{tablenotes} • Notes: Standard errors (in parentheses) are robust to heteroskedasticity and correlation across observations on the same state and year of birth. Each set of weights sums to one across the whole sample. Subsample OLS coefficient of years of schooling is obtained with controlling for state of birth and year of birth dummies. Appendix (ref) describes the empirical specification for estimating the weights. \end{tablenotes} \end{threeparttable} \end{table} \begin{figure}[H] \caption{The IV and OLS Weights on the Treatment Levels} \begin{threeparttable} \begin{minipage}[t]{0.49\columnwidth} (a) NLSY79 (College Costs IV) \end{minipage} \begin{minipage}[t]{0.49\columnwidth} (b) British GHS (Compulsory Schooling RD) \end{minipage} \begin{minipage}[t]{0.49\columnwidth} (c) U.S. Census (Compulsory Schooling DID) \end{minipage} \begin{minipage}[t]{0.49\columnwidth} \begin{center} \end{center} \end{minipage} \begin{tablenotes} • Notes: The IV weights are presented with standard error bars. The standard error bars for the OLS weights are omitted because they are graphically negligible. Each set of weights sums to one across the whole sample. Appendix (ref) describes the empirical specification for estimating the weights. The weights are zero outside the support of years of education by construction, and thus are not plotted on the graph. \end{tablenotes} \end{threeparttable} \end{figure}

\setcounter{secnumdepth}{3}