EconBase
← Back to paper

Don't (fully) exclude me, it's not necessary! Causal inference with semi-IVs

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

147,398 characters · 26 sections · 132 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

1820 Don't (fully) exclude me, it's not necessary! Causal inference with semi-IVs

abstract{ This paper proposes semi-instrumental variables (semi-IVs) as an alternative to instrumental variables (IVs) to identify the causal effect of a binary (or discrete) endogenous treatment. A semi-IV is a less restrictive form of instrument: it affects the selection into treatment but is excluded only from one, not necessarily both, potential outcomes. Having two continuously distributed semi-IVs, one excluded from the potential outcome under treatment and the other from the potential outcome under control, is sufficient to nonparametrically point identify marginal treatment effect (MTE) and local average treatment effect (LATE) parameters. In practice, semi-IVs provide a solution to the challenge of finding valid IVs because they are often easier to find: many selection-specific shocks, policies, prices, costs, or benefits are valid semi-IVs. As an application, I estimate the returns to working in the manufacturing sector on earnings using sector-specific characteristics as semi-IVs. \\ Keywords: instrumental variable, treatment effect, exclusion restriction, identification, selection, policy evaluation.} \\

\thispagestyle{empty}

\restoregeometry \setstretch{1.25} {6pt} {6pt}

\setcounter{page}{1}

Introduction

Identifying the causal effect of a treatment on an outcome is a central challenge in economics, particularly when the selection of treatment is endogenous. A common approach is to use instrumental variables (IVs), which must satisfy three conditions to be valid: they must affect treatment selection (relevance), be independent of unobserved factors (exogeneity), and have no direct effect on the potential outcomes (exclusion restriction). The IV approach is appealing because it can identify causal effects under relatively mild additional assumptions. For instance, when IVs influence treatment selection uniformly in the same direction for all individuals -- a condition known as monotonicity imbensangrist1994 -- they enable the nonparametric identification of average treatment effects for individuals whose treatment choice is influenced by the IVs. These include marginal treatment effect (MTE) heckmanvytlacil2005 and local average treatment effect (LATE) imbensangrist1994 parameters, which are widely estimated in applied research. However, finding valid IVs that meet these conditions is challenging. In particular, the exclusion restriction is a strong and often controversial assumption, even for commonly used instruments, because it is difficult to find a variable that affects the selection without affecting any potential outcome. \\ Meanwhile, in many settings with a discrete treatment, there are numerous variables that, while not excluded from all potential outcomes, are credibly excluded from some. Think, for instance, of treatment-specific characteristics -- such as costs, prices, benefits, amenities, or policies -- that affect the outcomes of individuals only if they select that specific treatment but not if they select another one. I call such variables semi-instrumental variables (semi-IVs): variables relevant for treatment selection and excluded from some, but not necessarily all, potential outcomes. This raises a critical question: is causal inference possible with semi-IVs alone? Remarkably, yes, this paper shows that semi-IVs can serve as substitutes for IVs in addressing endogeneity, even in general models with heterogenous treatment effects. \\ When the treatment is binary, two continuously distributed semi-IVs -- one excluded from the potential outcome under treatment and the other from the potential outcome under control -- are sufficient to nonparametrically identify MTE and LATE parameters, just as an IV would. Crucially, this identification does not require any additional assumption beyond those in the standard IV framework (monotonicity of the selection, applied to both semi-IVs), nor any restriction on how the semi-IVs affect the potential outcomes from which they are not excluded. The identification is IV-like and stems directly from the observable variations induced by the semi-IVs. By expanding the variation available for causal inference, semi-IVs may offer a solution when valid IVs are difficult to find. Since semi-IVs are often more readily available in applied settings, they significantly broaden researchers' toolkit and open up new possibilities for identifying causal effects. \\ To illustrate this, I present several examples of semi-IVs in occupation and location choice models, with various applications in labor, education, and industrial organization, among others (see Section (ref)). As the main running example, which also corresponds to the empirical application in Section (ref), I focus on the identification of the earnings returns to working in the manufacturing sector heckmansedlacek1985, heckmansedlacek1990. Workers select into sectors based on their unobservable attributes, including their sector-specific skills, leading to endogenous selection. As a result, the comparison of earnings between manufacturing and nonmanufacturing workers is confounded. Ideally, we would like to find an IV to address the selection on unobservables, but finding one that influences the sector choice without also influencing the earnings is particularly challenging. Fortunately, we have access to many sector-specific characteristics, such as sector size (e.g., GDP, number of firms and employees), productivity, and growth, with observable variations across time and markets (e.g., states), which can be leveraged for identification. These are promising semi-IV candidates. Indeed, they are likely to affect workers' sector choices: all else equal, the more attractive a sector's characteristics, the more likely a worker is to select it. Such aggregate (shocks to) sector characteristics are also independent of the workers' idiosyncratic sector-specific unobserved skills and preferences. Unfortunately, they are not valid IVs because a worker's earnings are typically determined as a function of the worker's unobserved skills and the characteristics of the selected sector heckmansedlacek1985. The sector-specific characteristics are included in the corresponding sector-specific potential outcomes. Crucially, however, conditional on the characteristics of the manufacturing sector, (shocks to) the characteristics of the nonmanufacturing sector should not directly affect the potential earnings of manufacturing workers, and vice versa. This partial exclusion is the key property that makes sector-specific characteristics -- and, more generally, many treatment-specific variables -- valid semi-IVs. Once characteristics specific to the treated are controlled for, similar characteristics for the untreated do not affect the potential outcomes under treatment, and vice versa.\footnote{Examples of models where sector-specific characteristics satisfy these semi-IV conditions include heckmansedlacek1985, heckmansedlacek1990 and, more recently, eckardt2024. See Section (ref) for discussion. }

It is important to note what is not required for semi-IVs to be valid. First, the two semi-IVs may be correlated, as is often the case. Their validity only requires that, conditional on the included semi-IV, the other one is relevant and rightfully excluded. In a sense, any effect of the excluded semi-IV on a given potential outcome must be subsumed by the effect of the included one.\footnote{ Thus, semi-IVs can be interdependent or jointly determined in equilibrium, as long as, once they are determined, the potential outcomes depend only on their treatment-specific semi-IV. This conditional exclusion property is particularly natural when semi-IVs are treatment-specific prices; see Section (ref) for discussion. } Second, the exclusion restriction applies to individuals' (latent) potential outcomes, but each semi-IV may still indirectly influence the observable mean outcome of the group from which it is excluded via selection effects. For instance, holding other factors constant, an increase in the size of the manufacturing sector encourages more workers (the compliers) to select into manufacturing, which, in turn, alters the observed average earnings of nonmanufacturing workers solely through changes in the workforce composition. \\ In fact, the observable changes in the average nonmanufacturing earnings and the selection probability, both induced by changes in the excluded manufacturing sector characteristics (as semi-IV), are precisely the variations that identify the mean nonmanufacturing potential earnings for the compliers who select into manufacturing as a result of the change in the semi-IV. Similarly, shifts in the nonmanufacturing sector characteristics identify the mean manufacturing potential earnings for the compliers resulting from these shifts. This mirrors the standard IV nonparametric identification arguments but applied separately to the outcomes of the treated and untreated subpopulations instead of directly to all outcomes. \\ Such a decomposition hints at why semi-IVs are sufficient to identify treatment effects. MTE and LATE parameters compare the potential outcomes under treatment (e.g., manufacturing) and under control (e.g., nonmanufacturing) for specific subpopulations of (marginal) compliers. An IV is excluded from both potential outcomes, so shifting it identifies the mean of both potential outcomes for the same set of compliers at once, and directly yields the treatment effect. While convenient, this is not necessary. Indeed, as previously described, two distinct semi-IVs can separately identify the means of the potential outcomes under treatment and control for specific subpopulations. Then, to reconstruct treatment effects, one needs to align the changes in each semi-IV that induce the same flow of compliers into treatment. This alignment is straightforward under the usual LATE monotonicity assumption. Indeed, as with multiple instruments carneiroheckmanvytlacil2011, monotonicity ensures that the treatment probability summarizes the effect of all semi-IVs on selection into a single measure. At the margin, the marginal complier induced into treatment by an infinitesimal increase in the manufacturing sector size is the same as the one induced into treatment by an infinitesimal decrease in the nonmanufacturing sector size. More generally, any two semi-IV shifts causing the same observable shift in treatment probability induce the same set of compliers. \\ Overall, the primary cost of the semi-IV approach is that researchers must use two semi-IVs, rather than a single IV, to identify similar treatment effects. This entails two distinct costs: a (typically modest) search cost and, in the binary treatment case, a stronger monotonicity requirement.\footnote{Throughout, I require continuously distributed semi-IVs so that complier flows can be aligned across treatment margins. For MTE, this adds no cost: standard IV-based MTE identification also require continuous IVs. For LATE, it is stricter than conventional IV identification, which works with discrete or even binary IVs. This extra cost is mild because most examples of semi-IVs are continuously distributed. Moreover, the support requirement can be relaxed: in fact, LATE can be identified if only one of the two semi-IVs is continuous. With two discrete semi-IVs, LATE may also be obtained by interpolation. } The search cost is usually small because valid semi-IVs often come in treatment-specific pairs (e.g., treatment-specific prices, size). So, researchers generally only need to find a single type of semi-IV that can be adapted to all potential outcomes. A related cost is that, by design, because we use at least two semi-IVs, monotonicity becomes a multi-instrument assumption, making it stronger than in the benchmark single IV case lee2018identifying, mogstadetal2021. That said, since semi-IVs are typically representing the same type of incentive across alternatives (e.g., prices), assuming that all individuals respond in a uniform direction to changes in the semi-IVs (or in their relative comparison) may be more plausible than with unrelated incentives (e.g., college proximity and tuition fees in the standard multi-IV example). Note also that this additional monotonicity cost is specific to the binary treatment case. For discrete treatments with more than two distinct alternatives, standard IV identification also requires multiple instruments and thus a stronger monotonicity condition. On balance, the cost of the semi-IV approach -- especially in discrete-treatment settings -- is often modest relative to its main benefit: relaxing the stringent full exclusion restriction of IVs. Semi-IVs provide a convenient alternative to standard IVs for causal inference. \\ Building on constructive identification arguments, this paper develops both nonparametric and semi-parametric methods for estimating MTEs and LATEs with semi-IVs. In particular, the semi-parametric method adapts the local IV MTE estimation heckmanvytlacil1999, carneiroheckmanvytlacil2011, andresen2018exploring to the semi-IV framework. The estimator is fast, easy-to-implement, and performs well, as illustrated by simulations in Online Appendix (ref). A user-friendly implementation is available in the companion R package semiIVreg semiivreg.\footnote{The package is available for download on \href{https://github.com/cbruneelzupanc/semiIVreg}{Github}. See the associated vignette at \href{https://cbruneelzupanc.github.io/semiIVreg/}{https://cbruneelzupanc.github.io/semiIVreg/} for user guidelines and replicable simulated examples. } \\ To illustrate the method, I estimate the (marginal) earnings returns to working in the manufacturing sector in the US and their evolution from $1999$ to $2018$ using state-level sector-specific characteristics (size) as semi-IVs. The analysis focuses on young white men with only high school education, the group most affected by the decline of the manufacturing sector over the period considered pierce2016surprisingly. I find significant heterogeneity in returns: some workers earn more in manufacturing, others more in nonmanufacturing jobs. Moreover, I find that the returns to working in manufacturing declined over the $20$-year period, closely mirroring the decline in manufacturing employment. This occurs despite an increasing observable earnings gap in favor of manufacturing over nonmanufacturing workers (as estimated by naive OLS), highlighting the importance of controlling for the endogenous change in sector composition. Perhaps most strikingly, I find that the potential earnings in both sectors declined over the period. This decline highlights a marked deterioration in the economic prospects of young, less-educated white males, regardless of their choice of sector. \\ The main focus of this paper is on the effect of a binary treatment. This case provides the most transparent intuition for why semi-IVs can effectively substitute for IVs. However, the semi-IV approach naturally extends to the more general case of a discrete treatment with $J > 2$ distinct alternatives. In this setting, having one alternative-specific semi-IV -- which must be excluded from all but one potential outcome -- for each potential outcome (giving a total of $J$ distinct semi-IVs) is sufficient to nonparametrically identify discrete treatment effects, such as the effect of one treatment relative to the best alternative heckmanurzuavytlacil2006, heckmanvytlacil2007b or margin-specific MTEs mountjoy2022community. Notably, identifying these treatment effects with standard IV methods typically requires $J-1$ special "alternative-specific" IVs, which must affect the latent utility of only one alternative but must be excluded from all potential outcomes, unlike semi-IVs. Thus, semi-IVs provide a natural generalization by conveniently allowing the alternative-specific instruments to also impact their corresponding potential outcomes, hence enlarging the available variation to identify discrete treatment effects. This gain comes at the modest cost of requiring $J$ semi-IVs instead of $J-1$ IVs. Apart from this, and contrary to the binary treatment case, there is no additional cost of using semi-IVs in terms of monotonicity there. In fact, margin-specific MTEs can be also be identified with semi-IVs under the same weaker form of monotonicity imposed on IVs by mountjoy2022community, i.e., unordered partial monotonicity combined with a comparable complier assumption. The formal extension of the discrete treatment identification results with IVs to semi-IVs is provided in Online Appendix (ref). The semi-IV approach does not extend to continuous treatments as it would require infinitely many semi-IVs. \\

Related Literature. Exploiting targeted exclusion restrictions from certain potential outcomes to achieve identification has been used in part of the literature on Roy models heckmansedlacek1985, heckmansedlacek1990, heckmanhonore1990, heckman1990, heckmanvytlacil2007b, bayeretal2011, frenchtaber2011, dhaultfoeuillemaurel2013, mourifie2020sharp. For instance, the sector-specific variables in heckmansedlacek1985, heckmansedlacek1990 and dhaultfoeuillemaurel2013 can be interpreted as semi-IVs, excluded from the potential outcomes of workers selecting into other sectors. Roy models can be parametrically identified using distributional assumptions on the unobserved sector skills, possibly combined with sector-specific variables heckmanhonore1990. \\ However, nonparametric identification results using only sector-specific variables (semi-IVs) are scarce and limited to models with specific selection rules. For Roy models -- where the sector selection depends solely on the difference between the potential earnings, i.e., where workers select the sector giving them the highest potential earnings -- heckmanhonore1990 and frenchtaber2011 show that nonparametric identification of the treatment effects can be obtained with sector-specific variables if these variables have an additively separable effect with respect to the unobservables affecting the outcomes. The identification exploits the Roy model structure, which implies that the semi-IVs affect the first stage solely through their effect on the potential outcomes. Thus, by observing changes in sector choice probabilities induced by variations in semi-IVs, one can infer their effects on the corresponding outcomes and subsequently identify certain treatment effects. However, this reasoning is specific to Roy models.\footnote{Under the Roy model specification, the full model (i.e., the joint distribution of the potential outcomes) can be identified, but its identification relies on identification-at-infinity arguments, requiring sector-specific variables (semi-IVs) to shift the propensity score to zero or one, which is rarely satisfied in practice.} For extended Roy models -- where the selection rule is a function of the differences between the potential earnings and a cost depending only on observable covariates, with no additional unobserved shocks beyond those affecting earnings -- dhaultfoeuillemaurel2013 show that nonparametric identification of the distribution of the treatment effects is also possible if the sector-specific variables have an additively separable effect on the outcomes.\footnote{In both, Roy and extended Roy models, additive separability in the outcome translates to the selection equation, yielding weak separability in the first stage and imposing a stronger form of LATE monotonicity.} Their identification results are more general than those in Roy models but still rely on the extended Roy model structure, in particular, the known link between the residuals in the outcome and the selection equations. \\ In contrast, for the more widely applied generalized Roy models -- where the selection rule is more general and includes additional unobserved shocks (e.g., individual unobserved preferences) that may correlate with those affecting earnings -- nonparametric identification of treatment effects has so far relied on IVs combined with the LATE monotonicity assumption heckmanvytlacil2005, heckmanvytlacil2007b, frenchtaber2011.\footnote{The "LATE model" imbensangrist1994 is, implicitly, a generalized Roy model, in the sense that there is no restriction on the potential outcomes (allowing for heterogenous treatment effects) and no restriction on the selection, apart from monotonicity vytlacil2002. See heckmanvytlacil2007b for a clear-cut distinction between Roy, extended Roy, and generalized Roy models.}\footnote{An exception is eckardt2024, who estimates the effect of training on occupation-specific wages using occupation-specific vacancy rates as instruments to control for the selection, without invoking a Roy/extended Roy structure. Instead, because of the high-dimensional treatment space, the paper adapts lee1983generalized and dahl2002 and imposes parametric assumption on the selection rule. Moreover, its main focus is to estimate homogenous `treatment effects' of training on occupation-specific wages (and not occupation premia). Although not a general nonparametric identification result in a generalized Roy model with semi-IVs, it provides a clear application in which occupation-specific variables help identifying some specific effects. } Even with IVs, the average treatment effect (ATE) for the entire population is generally not identified unless there exists a subpopulation for whom the probability of treatment is zero or one (identification-at-infinity). Fortunately, imbensangrist1994 showed that, even without this hard-to-satisfy identification-at-infinity requirement, an IV can still identify treatment effects of interest for subpopulations of compliers induced into treatment by varying the IV: the LATEs. heckmanvytlacil2005 further extended this result and showed that marginal variations of continuous IVs identify the MTEs for marginal compliers. \\ This paper shows that, under the monotonicity assumption of imbensangrist1994, semi-IVs can also nonparametrically identify treatment effects for subpopulations of compliers induced into treatment by varying the semi-IVs (MTEs and LATEs). The semi-IV identification argument parallels that for an IV, but applied separately to each potential outcome. It does not rely on a specific selection rule (apart from monotonicity), in contrast to existing identification results in Roy and extended Roy models. Thus, semi-IVs provide an alternative to IVs for identifying MTE and LATE parameters in general models with monotonicity imposed on the selection rule. This is particularly advantageous in empirical applications, where we want to avoid structural assumptions and where semi-IVs -- such as occupation or location-specific characteristics -- are often readily available but have so far only been used as control variables, neglecting their identification power. \% as semi-IVs. \\ This paper shows how semi-IVs can substitute for IVs in the LATE framework imbensangrist1994, angristimbensrubin1996, i.e., under the LATE monotonicity assumption. A companion paper bruneelbeyhum2025semi shows that semi-IVs can also substitute for IVs in the IV-quantile regression (IVQR) framework chernozhukovhansen2005 to identify quantile treatment effects (QTE) under the assumption of rank invariance on the potential outcomes instead of LATE monotonicity. The IVQR and LATE frameworks encompass most models used in applied research. Since semi-IVs achieve similar identification results in both frameworks, they can substitute for IVs in many empirical applications.\footnote{This paper focuses on \textit{point} identification. Extensions of IV nonparametric \textit{set} identification results manskipepper2000, chesher2010, chesherrosen2017 to semi-IVs are for future research.} \\ More broadly, this paper contributes to the literature on identifying treatment effects by relaxing key IV assumptions. A more detailed discussion of this literature is in Online Appendix (ref). For a recent survey, see lewbel2019identification or mogstad2024handbook. \\

Outline. Section (ref) adapts the canonical binary treatment IV setup angristimbensrubin1996 to semi-IVs, outlines their properties and illustrates with detailed examples. Section (ref) establishes nonparametric identification of MTEs and LATEs using semi-IVs. Section (ref) describes their nonparametric estimation, and shows how local IV estimation of MTEs can be extended to semi-IVs, offering a simple semi-parametric procedure. Finally, Section (ref) applies this procedure to estimate the returns to working in the manufacturing sector in the US. \\

Framework

First, this section shows how the canonical program evaluation problem of identifying the (possibly heterogenous) effect of an endogenous binary treatment on an outcome using an IV imbensangrist1994, angristimbensrubin1996, heckmanvytlacil2005 naturally extends to semi-IVs. Then, I discuss examples of semi-IVs in a variety of applications.

The semi-IV model

Potential outcomes and endogenous treatment. We want to evaluate the effect of a binary treatment $D \in \mathcal{D}=\{0, 1\}$ on a scalar, real-valued outcome $Y$. Let $Y_d$ be the latent potential outcome under treatment state $d$ $\in \mathcal{D}$. The researcher only observes the realized outcome, $Y$, corresponding to the potential outcome of the selected alternative $D$. That is,

align[align omitted — 58 chars of source]

and the potential outcome corresponding to the non-selected treatment is latent. The challenge faced by economists is that $D$ may be endogenous with respect to the potential outcomes $Y_0$ and $Y_1$, so we cannot simply compare the observable distributions of $Y$ given $D=1$ and $D=0$ in order to identify the causal effect of $D$ on $Y$. It is well known that a valid IV, combined with a monotonicity assumption, allows us to nonparametrically recover some causal effects of interest: the LATE imbensangrist1994 and the MTE heckmanvytlacil1999, heckmanvytlacil2005. I will show that valid semi-IVs allow us to recover these parameters as well. The intuition behind why semi-IVs can substitute for IVs can be gleaned from Equation (ref). The IV approach requires a variable that is fully excluded from $Y$. In the semi-IV approach, it suffices to find two variables (the semi-IVs), one excluded from $Y_1$ to serve as an instrument for the effect of $D$ on $DY = DY_1$, and another one excluded from $Y_0$ to serve as an instrument for the effect of $D$ on $(1-D)Y = (1-D)Y_0$. \\

Semi-IVs. Let $Z_0$ and $Z_1$ be the two (sets of) complementary semi-IVs, excluded from $Y_1$ and $Y_0$, respectively, with supports $\mathcal{Z}_0$ and $\mathcal{Z}_1$. Both semi-IVs are observed regardless of the selected treatment. Let $Z=(Z_0, Z_1) \in \mathcal{Z}$ be the set of all semi-IVs, with support $\mathcal{Z}$. Following imbensangrist1994, denote by $D(z)$ the potential treatment choice if the semi-IVs are exogenously set to $Z=z$. The observed treatment is thus $D=D(Z)$.\footnote{The stable unit treatment value assumption (SUTVA) is implicitly made, meaning each individual's potential outcomes and treatment are unaffected by the others' treatment status or semi-IVs. } The validity of the semi-IVs holds conditional on a set of control variables $X$, with support $\mathcal{X}$, which have no identifying power ($X$ may be neither excluded, exogenous, nor relevant).

assumption[Valid semi-IVs] $(Z_0, Z_1)$ satisfies the following conditions: \begin{enumerate}[label=A$D$(ref).\arabic*] • (Exogeneity and Exclusion) For all $z \in \mathcal{Z}$, (i) $D(z) \perp\!\!\!\perp Z | X$, and (ii) for $d=0,1$, $(Y_d, D(z)) \perp\!\!\!\perp Z_{1-d} | (Z_d, X)$, where $\perp\!\!\!\perp$ denotes conditional independence. • (Relevance) For $d=0,1$, the propensity score function, $P(z_0, z_1, x) = \mathbb{E}[D|Z_0=z_0, Z_1=z_1, X=x]$ is a nontrivial function of $z_d$ for all $(z_{1-d}, x) \in \mathcal{Z}_{1-d} \times \mathcal{X}$. \end{enumerate}

By definition, $(Z_0, Z_1)$ is a set of valid complementary semi-IVs if Assumption (ref) holds. A valid IV, $\tilde{Z}$, would be relevant for the selection into treatment and satisfy the stronger conditional independence assumption $(Y_0, Y_1, D(\tilde{z})) \perp\!\!\!\perp \tilde{Z} | X$ imbensangrist1994. In contrast, each semi-IV is relevant ((ref)), but excluded from only one of the two potential outcomes, not both ((ref)). Hence, we say that an IV is fully excluded while a semi-IV is only partially excluded from the outcome. Conditional on $Z_1$ and $X$, shifting $Z_0$ enables shifts in the selection probability without directly affecting $Y_1$: $Z_0$ acts as an IV for the effect of $D$ on $DY=DY_1$. Yet, $Z_0$ is not excluded from $Y_0$ and may affect it in an unrestricted manner, like any covariate $X$. The reverse holds for $Z_1$. Together, $Z_0$ and $Z_1$ provide an exclusion restriction for each potential outcome, making them complementary. \\ In the remainder of the paper, I refer to the conditional independence Assumption (ref)(ii) as the (partial) exclusion restriction, and call $Z_{1-d}$ excluded from $Y_d$ given $(Z_d, X)$. This terminology is a simplification, as, similar to IVs angristimbensrubin1996, satisfying $Y_d \perp\!\!\!\perp Z_{1-d} | (Z_d, X)$ actually requires two conditions: exclusion and exogeneity. To clarify the distinction between the two, denote by $Y_d(z, x)$ the potential outcome when $D$ is exogenously set to $d$, $Z$ to $z$, and $X$ to $x$. Thus, the potential outcomes defined earlier are $Y_d = Y_d(Z, X)$, and the observed outcome is $Y=Y_D=Y_D(Z,X)$. Then, the exclusion restriction is that, for all $x \in \mathcal{X}$, for all $(z_0, z_1)$, $(z_0, z_1')$, and $(z_0', z_1) \in \mathcal{Z}$, and for each $d=0, 1$,

align*[align* omitted — 73 chars of source]

That is, there is no direct effect on $Y_d$ of exogenously switching $Z_{1-d}$ from $z_{1-d}$ to $z_{1-d}'$, once we already control for $Z_d=z_d$ and $X=x$. Furthermore, exogeneity requires that $Z_{1-d}$ be independent of the potential outcomes and treatment, i.e., $(Y_d(z_d, x), D(z)) \perp\!\!\!\perp Z_{1-d} | (Z_d, X)$.\footnote{A stronger version of exogeneity in (ref) is to assume that $(Y_0(z_0, x), Y_1(z_1, x), D(z)) \perp\!\!\!\perp (Z_0, Z_1) | X$ for all $(z_0, z_1) \in \mathcal{Z}$ and $x \in \mathcal{X}$. This is still very general and only restricts how the semi-IVs affect their corresponding potential outcomes: through a direct effect rather than via their correlation with unobservables affecting the potential outcomes (see Online Appendix (ref)). This assumption enables a clearer causal interpretation of the effect of exogenous changes in semi-IVs on outcomes, though it is stronger than necessary for identification. } Notice that $Y_d(z_d, x)$ is not necessarily independent of $Z_d$ nor $X$. This is because $Z_d$ and $X$ may be correlated with the unobservables affecting $Y_d(z_d, x)$. The distinction between exclusion and exogeneity may be clearer when expressing the potential outcomes as latent variables in a generalized Roy model. The reformulation of the semi-IV model as a generalized Roy model, analogous to the IV framework of heckmanvytlacil2005, is detailed in Online Appendix (ref). \\

Causal graph representation. Figure (ref) provides a schematic causal graph representation of Assumption (ref) conditional on a specific $X=x$ (not displayed). The partial exclusions are visible by the absence of a direct link between $Z_{1-d}$ and $Y_d$ for each $d$.

figure[figure omitted — 1,815 chars of source]

Selection into treatment. As with standard IVs, identifying causal effects requires restricting how the semi-IVs affect treatment selection. I impose the same monotonicity/no-defier condition that imbensangrist1994 impose on IVs on the semi-IVs. Beware that, since it is applied jointly to multiple semi-IVs, the assumption is stronger than with a single IV, as in the multi-IV case mogstadetal2021. See discussion in Section (ref).

assumption[Monotonicity] Conditional on $X$, for all $z, z' \in \mathcal{Z}$, either $D_i(z) \geq D_i(z')$ for all individuals $i$, or $D_i(z) \leq D_i(z')$ for all individuals $i$.

Assumption (ref) states that, conditional on $X$, an exogenous change in the semi-IVs $Z$ from $z$ to $z'$ either weakly increases or weakly decreases selection into treatment for all individuals. This rules out non-uniform effects of $Z$ on treatment choice: if some individuals opt into treatment when $Z$ changes from $z$ to $z'$ ($D_i(z')=1 > D_i(z) = 0$ for some $i$), then no one opts out of the treatment because of this change ($D_i(z') = 0 < D_i(z) = 1$ for no $i$). Conversely, if some opt out, no one opts in. That is, there are only compliers and no defiers. \\

Equivalent selection model with latent resistance to treatment. For IVs, the monotonicity condition is equivalent to the existence of a weakly separable rule for the selection into treatment vytlacil2002. The same holds for semi-IVs: monotonicity is equivalent to assuming that the treatment is determined by a weakly separable rule of the form

align[align omitted — 71 chars of source]

for some function $P$ and a random variable $V$ that is continuously distributed given $X$. \\ Under the selection rule (ref), Assumption (ref) can be rewritten as follows. {\theoremstyle{plain} \newtheorem*{specialassumption}{Assumption (ref)'}

specialassumption(i) $V \perp\!\!\!\perp Z | X$, and (ii) for $d=0,1$, $(Y_d, V) \perp\!\!\!\perp Z_{1-d} | (Z_d, X)$.

} As is standard in the IV literature, since $V$ is continuously distributed given $X=x$ for all $x$, one may normalize $V|X \sim \text{Unif}(0, 1)$ without loss of generality.\footnote{This normalization is standard and innocuous given the model assumptions. Indeed, if the selection rule is $D = \mathds{1}\{ \nu(Z, X) - \Eta \geq 0\}$ where $\Eta$ is a general continuous random variable, the model can always be reparametrized with $P(Z, X) = F_{\Eta|X}(\nu(Z, X))$ and $V=F_{\Eta|X}(\Eta)$, ensuring $V|X \sim \text{Unif}(0, 1)$, which corresponds to Equation (ref) under this normalization.} Given the independence of the semi-IVs and $V$ (Assumption (ref)'), it follows that $V | Z=z, X=x \sim \text{Unif}(0,1)$ for all $(z, x) \in \mathcal{Z}\times\mathcal{X}$. Under this normalization, the function $P(z, x)$ is the propensity score function, i.e., $P(z, x) \equiv \mathbb{E}[D | Z=z, X=x]$, and $P=P(Z,X)$ is the propensity score random variable with support $\mathcal{P}$. \\ The normalized $V$ can be understood as the latent resistance to treatment cornelissen2016late, ranked from the lowest ($V=0$) to the highest ( $V=1$). Individuals with a low $V$ exhibit minimal resistance to treatment, opting in even when the propensity score $P$ is low, meaning they require fewer incentives to select into treatment. Conversely, those with a high $V$ have a strong resistance to treatment and only opt in if $P$ is high. \\

Regularity conditions. Following heckmanvytlacil2005, identifying treatment effects also requires two additional mild regularity conditions: (i) $\mathbb{E}|Y_1|$ and $\mathbb{E}|Y_0|$ are finite (ensuring that the treatment parameters are well defined), and (ii) $0 < \textrm{Pr}(D=1|X=x) < 1$ for all $x \in \mathcal{X}$, ensuring the existence of a treated and an untreated population for every $x$.

Discussion about the model

Discussion. Assumptions (ref) and (ref), along with the regularity conditions, define the semi-IV model. The flexibility of this model is best understood by considering what is not imposed mogstad2018ecta. First, similar to the IV model, it does not restrict the dependence between $D$ and the potential outcomes. This allows for rich forms of observed and unobserved heterogeneity, where the causal effect of $D$ on $Y$ can vary with observed covariates, semi-IVs, and unobservables. Second, except for the monotonicity assumption, the selection rule remains general.\footnote{Attention: when one assumes additive separability of the error term in the $Y_d$ equations frenchtaber2011, dhaultfoeuillemaurel2013, monotonicity is weaker than the (extended) Roy selection rule. However, in general models with nonseparable unobservables in $Y_d$, monotonicity is distinct from -- rather than strictly weaker than -- the (extended) Roy selection rule with semi-IVs. For instance, interactions between semi-IVs and outcome-specific unobservables in $Y_d$ violate monotonicity under a Roy specification. } The model does not specify the reasons behind the individuals' treatment choices, in contrast to Roy models, where $D = \mathds{1}\{ Y_1 > Y_0 \}$, or extended Roy models, where $D = \mathds{1}\{ Y_1 - Y_0 + \nu(Z, X) > 0 \}$ for some function $\nu$ of the observables $Z$ and $X$. This flexibility accommodates various decision-making processes, including partial selection on gains ($Y_1 - Y_0$). Third, the effect of the semi-IVs on the potential outcome from which they are not excluded is unrestricted and possibly heterogenous, varying with factors like latent resistance to treatment, $V$. Finally, the semi-IVs can (and generally will) be correlated. This is not a problem since the exclusion restriction holds conditional on the included semi-IV (and covariates). The only cost of the correlation between the semi-IVs is that it may weaken their relevance by limiting their joint variation.\footnote{Assumption (ref) rules out the degenerate case where $Z_0$ and $Z_1$ are perfectly correlated, as each semi-IV must be relevant for selection into treatment conditional on the other semi-IV (and covariates).} \\

The role of monotonicity. Conditional on $X=x$, if an individual selects $D=1$ when $Z=z$, one can conclude that $V \leq P(z, x)$, using (ref) and the independence of $V$ and $Z$ given $X$. Conversely, $D=0$ when $V > P(z,x)$. Therefore, for any $z' \neq z$, if $P(z,x)=P(z',x)$, then $D_i(z', x) = D_i(z, x)$ for all individuals. The propensity score serves as a sufficient statistic for the effect of the semi-IVs on selection: any two values of the vector of semi-IVs yielding the same propensity score select the same individuals into treatment. \\ This property is key to recover treatment effects, as it allows us to match compliers with respect to different semi-IV margins. Starting from a given environment, if a shift of $Z_0$ and another shift of $Z_1$ induce the same change in observable propensity score, under monotonicity we conclude that they induced the same compliers, i.e., the same set of $V$, into (or out of) the treatment. If each treatment-specific semi-IV represent a completely different type of incentive/disincentive (e.g., tuition fee versus distance from college), this assumption could be arguably strong mogstadetal2021.\footnote{Although monotonicity is stronger with multiple IVs, many empirical paper still use multiple IVs because the practical benefits -- richer first stage variation -- can outweigh the cost carneiroheckmanvytlacil2011.} Since the semi-IVs are typically the same type of incentive but treatment-specific (e.g., treatment-specific "prices"), assuming monotonicity, i.e., that all individuals uniformly value the semi-IVs (or their comparison) is more reasonable mogstad2024policy.\footnote{It is possible to relax the monotonicity assumption. Indeed, to identify mean potential outcomes (marginal treatment responses) separately, a weaker unordered partial monotonicity mogstadetal2021 assumption is sufficient. The issue is that these responses are identified for possibly different sets of compliers across semi-IVs. In order to align these sets of compliers and obtain marginal treatment effects, an additional assumption is needed. Standard monotonicity is a solution, but not the only one. For instance, complier alignment and point identification (of margin-specific marginal treatment effects) can also be achieved using comparable compliers assumptions as in mountjoy2022community. Online Appendix (ref) shows how partial monotonicity combined with comparable compliers identifies margin-specific MTE in the discrete-treatment case with semi-IVs; adapting it to the binary case is straightforward.} \\

Implications of the partial exclusion. Assumption (ref)' (or (ref)) implies that

align[align omitted — 137 chars of source]

Intuitively, conditional on $X$, $Z_d$, and $V$, a shift in the excluded semi-IV $Z_{1-d}$ has no effect on the latent average potential outcome $Y_d$. However, this shift affects the propensity score, thereby altering the composition of the treated population and, as a result, the observed outcomes. Indeed, for the treated population, for all $z=(z_0, z_1) \in \mathcal{Z}$ and $x \in \mathcal{X}$, we have

align*[align* omitted — 238 chars of source]

where the first equality follows from the law of total expectation, the second from $D=1$ being equivalent to $V\leq P(z,x)$, and the third from $(Y_1, V) \perp\!\!\!\perp Z_0|Z_1,X$. Similarly, for the untreated population, we have

align*[align* omitted — 108 chars of source]

This makes it clear that for each $d \in \{0,1\}$ the only effect of $Z_{1-d}$ on the conditional mean of $Y\mathds{1}\{D=d\}$ given $Z_d$ and $X$ is through the propensity score $P(Z,X)$. Since $P(Z,X)$ is identified from data on $D, Z,$ and $X$, we can condition directly on $P$ instead of $Z_{1-d}$, yielding

align[align omitted — 148 chars of source]

for $d=0,1$ and for all $z=(z_0, z_1) \in \mathcal{Z}$ and $x \in \mathcal{X}$. Similar to an IV, $Z_{0}$ (resp. $Z_1$) affects $DY$ (resp. $(1-D)Y$) only through its effect on the selection propensity $P$. The excluded semi-IV ($Z_{1-d}$) can shift the selection probability and alter the composition of the treated and control groups while holding the included semi-IV ($Z_d$) and covariates ($X$) fixed. Therefore, conditional on $Z_d$ and $X$, $Z_{1-d}$ acts as an IV for the effect of $D$ on $Y\mathds{1}\{D=d\}$. \\

Comparison with IVs. Each semi-IV acts as an IV for the observable outcomes of either the treated or the untreated subpopulation. An IV is stronger, excluded from the outcome $Y$, meaning it is excluded from both $DY$ and $(1-D)Y$. Since the conditions for a valid semi-IV are strictly weaker than those for an IV, a valid semi-IV is easier to find in practice. Indeed, as shown in the examples below, there are many selection-specific variables that are invalid IVs because they directly affect their alternative-specific potential outcome; but they are valid semi-IVs because they are credibly excluded from the other potential outcomes. \\ While a single valid semi-IV is strictly easier to find than a valid IV, it cannot serve as a substitute for a valid IV on its own. To replace an IV, at least two complementary semi-IVs are needed, with at least one excluded per potential outcome. Thus, there are two possible approaches to address endogeneity: finding (at least) one IV or two complementary semi-IVs. The full exclusion restriction is often the main obstacle to finding a valid IV. Moreover, as discussed in the examples below, finding two complementary semi-IVs is often not much harder than finding one, since treatment-specific versions of a similar variable can be used as semi-IVs. Therefore, complementary semi-IVs should often be more accessible than an IV. \\ Then, in the binary treatment case, the main cost of the semi-IV approach is that the monotonicity assumption jointly imposed on both semi-IVs is stronger than when imposed on a single IV. Practitioneers therefore face a trade-off between assuming stronger monotonicity on the selection (semi-IV) or full exclusion on the outcome equation (IV). \\ This main trade-off disappears in the discrete treatment case (with $J > 2$ distinct alternatives). There, identification with IVs already requires multiple ($J-1$) IVs and a correspondingly stronger montonicity. I show in Online Appendix (ref) how identification of discrete treatment effects obtained with $J-1$ IVs heckmanurzuavytlacil2006, mountjoy2022community can be obtained under the same assumptions with $J$ semi-IVs instead. For intuition, the main text focuses on the binary case. \\ Another cost of the semi-IVs approach is that matching compliers ideally require at least one continuously distributed semi-IV. This is not an additional cost for MTE identification which also requires continuous IVs. However, this is costly for the LATE which standard IV methods can identify with a single binary instrument while binary semi-IVs generally cannot. Fortunately, as visible in the examples described in the next subsection, most examples of semi-IVs are, in fact, continuously distributed. \\ Finally, and perhaps surprisingly, the possible non-exclusion of the semi-IVs from some potential outcomes has an unexpected upside: if individuals select into treatment partly based on their gains, $Y_1 - Y_0$, a variable affecting the potential outcomes should be relevant for the selection by construction. This should make it easier to find "non-weak" semi-IVs.

Examples and discussion

Example 1 (Occupation choice). Consider the question of identifying the effects of occupation choice ($D$) on earnings ($Y$). Examples include how earnings are impacted by working in the manufacturing sector heckmansedlacek1985, heckmansedlacek1990, agriculture lagakoswaugh2013, or a white-collar or blue-collar occupation keanewolpin1997, as well as by joining a union lee1978unionism, attending college angristkrueger1991, card1995, card2001, carneiroheckmanvytlacil2011, or selecting a specific college major or training arcidiacono2004, arcidiacono2020ex, eckardt2024. To answer these, one needs to address endogeneity due to the self-selection of individuals into occupations based on their unobserved occupation-specific skills. To do so, occupation-specific characteristics such as the occupation-specific wage rate, unemployment rate, task content, firm size, revenue, productivity, or subsidies can serve as valid semi-IVs. More precisely, identification can be achieved by leveraging observable variations of these semi-IVs across time and geographic markets, and the corresponding variations in observed outcomes ($Y$) and occupation choices ($D$). \\ Returns to working in manufacturing on earnings. The main application in this paper (Section (ref)) focuses on the effect of the manufacturing ($D=1$) versus the nonmanufacturing ($D=0$) sector choice on earnings ($Y$). I focus on young (aged $18$ to $30$) white men with only high school education, and all the analysis is conducted conditional on being in this subpopulation ($X$). I use the characteristics of the manufacturing ($Z_1$) and nonmanufacturing ($Z_0$) sector of the market (state of residence) of the workers as semi-IVs. First, these are likely to be relevant variables for sectoral choice because, all else equal, a larger or more active sector is more attractive to new (young) workers. Monotonicity assumes that workers have uniform valuation of the sector-specific characteristics (i.e., not interacted with latent $V$). Second, the aggregate sector-specific characteristics are likely to be exogenous with respect to atomistic workers' unobserved sector-specific skills and latent resistance to treatment ($V$). Exogeneity requires that, aside from differences in the semi-IVs, the various markets represent otherwise comparable environments. To ensure this, I control for time trends and pre-existing differences across markets by including year and market (state) fixed effects in $X$. This addresses concerns about the self-selection of high-skilled workers into more favorable markets and accounts for the global evolution of worker skills. Consequently, the identifying variation consists of deviations of sector-specific characteristics from these fixed effects, which I will occasionally refer to as "shocks to" the variables.\footnote{Instead of fixed effects, one could control for the permanent levels of the semi-IVs in each market in $X$, as done by carneiroheckmanvytlacil2011 for some of their IVs, or use shocks to these variables directly as semi-IVs. } Third, (shocks to) the sector characteristics are not valid IVs because a worker's earnings are determined as a function of the worker's unobserved skills and the characteristics of the selected sector heckmansedlacek1985: the sector-specific semi-IVs are included in the corresponding sector-specific potential earnings, i.e., $Z_d$ influences $Y_d$. However, once the included manufacturing characteristics are controlled for, similar characteristics of the other sectors do not affect the manufacturing potential earnings, and vice versa: $Y_d$ is independent of $Z_{1-d}$ conditional on $Z_d$ and $X$. \\ Note that, in a given market, (the shocks to) the manufacturing and nonmanufacturing characteristics are likely to be correlated. As long as they are not perfectly correlated, so that some markets experience relatively larger shocks to their manufacturing sectors compared to their nonmanufacturing sectors than other markets, this is not a problem for identification. \\ As with standard IVs, I work in partial equilibrium and assume that general equilibrium (GE) effects that could threaten partial exclusion restrictions are negligible. For instance, an increase in manufacturing attractivity might alter worker composition, thereby affecting equilibrium wage rates in both sectors and indirectly influencing nonmanufacturing workers' earnings. Fortunately, and unlike the IV case, GE concerns are partly mitigated with semi-IVs. Indeed, both sectors characteristics can be correlated, or jointly determined in equilibrium, the semi-IVs' validity only requires that, once determined, conditional on $Z_d$, $Z_{1-d}$ is excluded from $Y_d$. This property often holds, even in some GE models heckmansedlacek1985, heckmansedlacek1990 where the effect of $Z_{1-d}$ on $Y_d$ are captured by the "effect" of $Z_{1-d}$ on $Z_d$. Controlling for the appropriate sector characteristics $Z_d$ (e.g., productivity, size) absorbs a large part of the GE concerns. This property is most salient when $Z_0$ and $Z_1$ are the sector-specific wage rates: once we control for the equilibrium wage rate in sector $d$, the wage rate in the other sector has no effect on $Y_d$, which only depends on the workers' unobserved skills and on the sector wage rate/price of skills.\footnote{Note that the GE threats are further mitigated by the fact that I focus on a small subpopulation of young workers, which may be too small to influence the aggregate wage rates.} \\ Which sector-specific characteristics to use? Ideally, one would like to observe sector-specific price of skills (wage rates) across markets. Contrary to other examples below, these are not available in this application. Instead, I use sector-specific size, as measured by sector-specific GDP, as semi-IVs in the application of Section (ref). One could use more characteristics per sector as semi-IVs, but according to the GE model of heckmansedlacek1985, the size should be sufficient, as it determines the sector-specific marginal productivity of labor, and thus the sector-specific wage rates.\footnote{In this example, since I do not use the sector-specific wage rates as semi-IVs, a threat to the semi-IVs' validity could be monopsony power with strategic interactions between sectors, such as manufacturing firms adjusting wages based on the size of the nonmanufacturing sector. However, evidence suggests within-sector strategic responses are small, with staiger2010there estimating an elasticity of only $0.128$ in the case of VA (Veteran Affairs) and non-VA hospitals. Between-sector interactions are likely smaller and negligible, as sectors do not explicitly compete. Given that I focus on large markets (states) with many firms in each sector, individual firms are relatively small and primarily compete within their own sector, reducing the risk of cross-sector strategic interactions. Thus, I assume that firms are price takers without monopsony power. } For more details and a complete example of a sector choice model where the sector-specific sizes (measured by total employee compensation in their case) would be valid semi-IVs, refer to heckmansedlacek1985, heckmansedlacek1990. More recently, eckardt2024 employs a similar model to justify the use of sector-specific vacancy rates as valid semi-IVs. In terms of timing, I use the previous year (one-year lagged) sector sizes as semi-IVs, as I assume it reflects the local market conditions when the workers made their sector choices, i.e., at the beginning of the period. In principle, additional lags or current-year sizes could also be considered. \\

Example 2 (Location/migration choice). Consider the related problem of identifying the marginal earnings returns to endogenous location/migration choices of workers borjas1987, dahl2002, chiquiar2005international, roca2017learning or firms combes2012productivity. Location-specific (variations in) expected earnings kennanwalker2011, welfare benefits borjas1999welfare, abramitzky2009effect, agersnap2020welfare , tax rates akcigit2016inventors, kleven2020taxation, or more generally, any place-based policy neumark2015place or measures of local amenities, may serve as valid semi-IVs. Again, these are (i) relevant for the location choice, (ii) exogenous with respect to the individuals' unobserved characteristics, and (iii) the partial exclusion holds because, conditional on one's location-specific shocks, the shocks to the other locations are excluded from one's potential outcome. One can use temporal variation in these location-specific semi-IVs for identification. \\

Example 3 (Demand estimation). Semi-IVs can be used to estimate the demand for different products. For instance, consider the effect of car type -- gasoline ($D=0$) or diesel ($D=1$) -- on subsequent energy consumption, of gasoline ($Y_0$) or diesel ($Y_1$). Endogeneity arises because individuals with different unobserved preferences (e.g., for higher mileage) select different cars. Estimating demand for gasoline/diesel cars is crucial for designing optimal energy policies, such as taxation or subsidies verboven2002quality, grigolon2018consumer. Local market energy-specific prices serve as semi-IVs: they are relevant for the choice of type of car but do not affect the energy consumption of owners of the other car type (e.g., gas prices do not affect the energy consumption of diesel-car owners).\footnote{To address the potential endogeneity of energy prices with respect to unobserved individual characteristics, one can use supply side shocks to the prices (e.g., cost shifters) or exogenous energy-specific subsidies/tax changes as semi-IVs instead of the prices. Additionally, one can control for market fixed effects. } More broadly, energy-specific (shocks to) prices or subsidies can serve as semi-IVs to estimate the demand for energy based on appliance choice, such as solar panel adoption degrooteverboven2019, feger2022solar or electric versus non-electric home appliance purchases dubin1984. \\

Example 4 (Measuring effectiveness). Semi-IVs can also be used to measure the effectiveness (or value-added) of specific types of institutions or products, net of selection effects. For instance, in education, consider the estimation of the effect of charter schools ($D=1$) abdulkadiroglu2011charter, angristpathakwalters2013, dobbiefryer2020, boarding schools behaghel2017ready, catholic schools altonjietal2005, or private schools angrist2002vouchers compared to public schools ($D=0$), on academic achievements or future labor market outcomes ($Y$). In health, consider the estimation of the effectiveness of different insurance plans (e.g., private vs public) finkelsteinhendrenluttmer2019, abalucketal2021, type of hospitals gowrisankaran1999estimating, geweke2003hospital, or nursing homes einav2022producing on health outcomes. In all these examples, variations in institution-specific prices, subsidies finkelstein2019subsidizing, size, funding, or measures of quality (number of teachers/practitioners per student/patient) can serve as semi-IVs. \\

Randomized experiments with semi-IVs. Even in randomized experiments, a randomly assigned intervention is not necessarily a valid IV. While randomization ensures exogeneity with respect to ex ante unobservables, the intervention itself may not be excluded from the outcome of interest angristimbensrubin1996. For instance, conditional cash transfers, such as vouchers for private schooling in Colombia angrist2002vouchers, likely affect students’ achievements directly through income effects for families who receive them if the student attends a private school. As a result, only the intention-to-treat effect of the policy can be estimated, not the treatment effect of attending private school. However, the voucher assignment has no impact on students attending public schools, who receive no money regardless of their assignment, making it a valid semi-IV ($Z_1$), excluded from outcomes associated with attending a public school ($Y_0$).\footnote{The same kind of reasoning applies to the Vietnam lottery draft angrist1990. Non-veterans who were drafted probably had to behave differently or face the consequence of their non-compliance (e.g., jail for conscientious objectors) compared to non-veterans who were not drafted. Thus, the draft is only excluded from the outcomes (subsequent earnings) of veterans, but not from the outcomes of non-veterans. Identifying the effect of being a veteran would require a complementary semi-IV excluded from the outcome of non-veterans, for example differential local tax rates/benefits for veterans at the time of enrollment. } To identify the (local) treatment effect of private schools, the Colombian government could have designed the experiment differently by assigning two random monetary amounts: one for public schools ($Z_0$) and one for private schools ($Z_1$). Both could be nonzero since $Z_0$ is excluded from private school outcomes ($Y_1$) and $Z_1$ from public school outcomes ($Y_0$). By randomly assigning the monetary amounts, this design would facilitate manipulating treatment probabilities. Moreover, it would allow policymakers to study two effects simultaneously: the treatment effect of private school attendance and the policy/semi-IV effect of income increases on achievement. This approach also addresses ethical concerns by ensuring all families receive some support, not just a select few. \\

Discussion. With the exception of a few education-related examples, credible IVs are rarely available to address endogeneity in the settings and questions discussed above. Identification has so far mainly relied on distributional assumptions, functional form restrictions, structural models, or on assuming selection only on observables. Semi-IVs expand the researchers' toolkit by enabling IV-like identification with readily available variables, allowing them to credibly tackle a broader range of questions, even when credible IVs are hard to find.

Identification

Similar to IVs, semi-IVs identify average treatment effects for compliers induced into treatment by their shifts: the LATE and the MTE. More precisely, each semi-IV separately identifies mean potential outcomes under treatment and control at specific unobserved $V$, which are then combined to obtain treatment effects. In this section, I show the identification of the MTE, followed by the LATE. I also describe the identification of the direct effects of the semi-IVs on the potential outcomes from which they are not excluded.

Marginal treatment effects and responses

Treatment parameters definition

Marginal treatment effect (MTE). The marginal treatment effect (MTE) was introduced by bjorklund1987estimation and generalized to the nonparametric case by heckmanvytlacil1999, heckmanvytlacil2005, heckmanvytlacil2007b. With semi-IVs, the MTE is naturally defined by

align[align omitted — 96 chars of source]

The marginal treatment effect MTE$(v, z_0, z_1, x)$ is the average causal effect of $D$ on $Y$ for individuals with unobserved resistance to treatment $V=v$ and observed characteristics $X=x, Z_0=z_0, Z_1=z_1$. It is called "marginal" treatment effect because it is the effect of treatment for individuals at the margin of taking up treatment when $P=v$. \\ The main difference with the MTE using IVs is that the MTE also depends on the semi-IVs here. Since $Z_0$ may affect $Y_0$, and $Z_1$ may affect $Y_1$, we need to condition on $(Z_0, Z_1)$, just as we condition on $X$. Fortunately, this conditioning does not impede identification. \\

Marginal treatment responses (MTR$_d$). Rather than identifying the MTE directly, we focus on the two marginal treatment responses (MTR) functions mogstad2018ecta, mogstad2018review instead. With semi-IVs, these are defined by

align[align omitted — 212 chars of source]

For each $d$, the MTR$_d$ is the mean potential outcome $Y_d$ for individuals with unobserved resistance to treatment $V=v$ and observed characteristics $Z_d=z_d$ and $X=x$. With the semi-IV exclusion property, $m_d$ only depends on $Z_d$ but not on $Z_{1-d}$. One may refer to $m_1$ and $m_0$ as the mean potential outcomes under treatment and under control, respectively. \\ The MTE is equal to the difference of MTRs,

align[align omitted — 218 chars of source]

The second equality holds because we can remove conditioning on the excluded semi-IV using the mean independence equation (ref), i.e., using the partial exclusion of the semi-IVs.

Identification

For notational convenience, omit $X$ in the notation and proceed conditional on $X=x$. \\ In the data, we observe $(Y, D, Z_0, Z_1)$. Since we observe $(Z_0, Z_1)$ and $D$, we also identify the propensity score $P=P(Z_0, Z_1)=\textrm{Pr}(D=1|Z_0, Z_1)$. Therefore, consider that we observe $Y, D, Z_0, Z_1,$ and $P$. From there, the identification of each MTR with semi-IVs is similar to the identification of MTRs with a single IV mogstad2018review. The only difference is that it is not the same source of variation (not the same semi-IV) that identifies both MTRs at once. Instead, $m_1$ is identified by using $Z_0$ as an instrument for the effect of $D$ on $DY$ given $Z_1$ and $X$, while $m_0$ is identified by using $Z_1$ as an instrument for the effect of $D$ on $(1-D)Y$ given $Z_0$ and $X$. $Z_0$ and $Z_1$ play, in turn, the role of an instrument (for $m_1$ and $m_0$, respectively) and the role of a covariate (for $m_0$ and $m_1$, respectively) in the proof. Once both MTRs have been separately identified, one can recover the MTE using Equation (ref). Note that the monotonicity of the selection implicitly plays a key role: although each MTR is identified by shifting a different semi-IV, any change in a semi-IV can be mapped to its corresponding marginal compliers, characterized by $V$, through the associated propensity score. First, I show the identification of each MTR separately, then that of the MTE. \\

Identification of MTR$_1$. Let us proceed conditional on a specific $Z_1=z_1 \in \mathcal{Z}_1$. To identify the mean potential outcome under treatment ($m_1$), use $Z_0$ as an IV excluded from $Y_1$. $Y_1$ is not observed for the untreated individuals, but recall that $Y=DY + (1-D)Y$ and focus on the observable treated population outcomes, $DY = DY_1$. Any shift in $Z_0$ affects the outcome of interest ($DY$ here) only through its effect on the propensity score, $P=P(Z_0, z_1)$. This enables us to identify the MTR$_1$ by marginally shifting $P$ at fixed $Z_1=z_1$ using $Z_0$. \\ Assume that the semi-IV excluded from $Y_1$, $Z_0$, is continuous and relevant for all $Z_0=z_0 \in \mathcal{Z}_0$, conditional on $Z_1=z_1$, i.e., that $\partial P(z_0, z_1)/\partial z_0 \neq 0$. In this case, $P$ is continuously distributed conditional on $Z_1=z_1$. Using the selection equation (ref), we have that $D=1$ when $V \leq P$. So, for any $P=p$ in the interior of the support of $P$ given $Z_1=z_1$, we observe

align[align omitted — 358 chars of source]

Thus, taking the derivative with respect to the propensity score, we have

align[align omitted — 116 chars of source]

where $\mathbb{E}[ DY | P(z_1, Z_0) = p, Z_1=z_1]$ is observed and so is its derivative if $v$ is in the interior of the support of $P$ given $Z_1=z_1$. Therefore, for any $z_1$ and any $v$ in the interior of the support of $P$ given $Z_1=z_1$, the MTR$_1$ given $Z_1=z_1$ and $V=v$ is identified by (ref). \\ The intuition behind why the derivative gives the MTR can be explained as follows. Fix $Z_1=z_1$. When $P=p$, individuals with $V \leq p$ are treated, and those with $V=p$ are at the margin of the treatment, meaning they are indifferent between being treated or not. Now, consider a marginal increase in $P$ from $p$ to $p'=p+dp$ by shifting $Z_0$ while holding $Z_1=z_1$ fixed. This shift moves the marginal individuals with $V=p$ into treatment. Their expected $Y_1$ is $\mathbb{E}[Y_1 | V=p, Z_1=z_1] = m_1(p, z_1).$ Since these individuals are now treated, the average observable outcome $DY$ changes by the proportion of individuals entering treatment times their average $Y_1$, i.e., $d(\mathbb{E}[DY | Z_1=z_1, P=p]) = dp \times \mathbb{E}[Y_1 | V=p, Z_1=z_1]$. Dividing both sides by $dp$ normalizes this change in $DY$ and yields $m_1(p, z_1)$. Thus, the derivative of the average $DY$ with respect to the propensity score identifies the mean potential outcome under treatment for individuals who are indifferent at $V=p$ (and $Z_1=z_1$). \\ The key to identification is having a source of variation that can shift the propensity score at fixed $Z_1$ without directly affecting $DY$ otherwise. This is precisely the role of $Z_0$, as visible in (ref). Implicitly, the derivative in Equation (ref) is identified by the local IV regression of $YD$ on $D$, instrumented by $Z_0$, at a specific value $\tilde{z}_0$ of $Z_0$ such that $P(\tilde{z}_0, z_1) = v$, i.e.,

align*[align* omitted — 218 chars of source]

In order to marginally shift $P$ by shifting $Z_0$, a relevant and continuously distributed $Z_0$ is needed, similar to the continuous IV requirement for MTE identification with IVs. \\

Identification of MTR$_0$. Similarly, for any $Z_0 = z_0 \in \mathcal{Z}_0$, $m_0(v, z_0)$ is identified for any $v$ in the interior of the support of the propensity score $P$ given $Z_0=z_0$ by\footnote{The minus sign comes from the fact that, $(1-D)=1$ when $V > P$, so $\mathbb{E}[(1-D)Y | Z_0 = z_0, P=p] = \int_p^1 m_0(v, z_0) dv,$ and the derivative with respect to $p$ at $p=v$ is applied to the lower bound of the integral.}

align[align omitted — 122 chars of source]

Holding $Z_0=z_0$ fixed, we shift the propensity score by shifting the other semi-IV, $Z_1$. Again, obtaining marginal shifts in $P$ requires $Z_1$ to be continuously distributed and relevant. \\

Identification of the MTE. Given $Z_1=z_1$, $m_1$ is identified for all $v$ in the interior of the support of the propensity score given $Z_1=z_1$. Given $Z_0=z_0$, $m_0$ is identified for all $v$ in the interior of the support of the propensity score given $Z_0=z_0$. As a consequence, for any $(Z_0, Z_1) = (z_0, z_1) \in \mathcal{Z}$, the MTE is identified by Equation (ref), as

align*[align* omitted — 60 chars of source]

for all $v$ in the interior of both, the support of $P$ given $Z_1=z_1$ and the support of $P$ given $Z_0=z_0$, i.e., the interior of the intersection of these two conditional supports. \\

Link with local IV heckmanvytlacil1999. With an IV, the MTE would be identified directly by the local IV regression of $Y$ on $D$, i.e., using

align[align omitted — 151 chars of source]

The problem with semi-IVs is that it is not possible to shift $P$ while fixing both $Z_0$ and $Z_1$. Instead, we split $Y$ into two subsamples, $YD$ and $Y(1-D)$, and identify $m_1$ and $m_0$ separately on these two subsamples using a different excluded semi-IV ($Z_0$ and $Z_1$, respectively) as an instrument to shift $P$, which is feasible thanks to Property (ref), implied by the partial exclusions. We recover the MTE afterwards by aligning the compliers with the same underlying $V=v$. Thus, the main difference with standard local IV is that the shifts in $P$ used to identify $m_1$ and $m_0$ stem from different underlying semi-IV shifts. \\

figure[figure omitted — 2,778 chars of source]

Visual representation. Figure (ref) illustrates the identification of the MTE (and LATE). It shows two distinct shifts in semi-IVs that separately identify $m_1$ and $m_0$. The black and grey curves represent all the combinations of semi-IVs $(Z_0, Z_1)$ yielding a propensity score of $P=p$ and $P=\bar{p}$, respectively. These are iso-probability curves, which are identified directly from the data. Imagine $\bar{p}=p + dp$, where $dp$ is a marginal change in the propensity score. At a given unobserved resistance to treatment $V=p$, to identify $m_1(p, z_1)$, we use the shift represented by the \textcolor{mtr1color2}{purple arrow}, from $Z=z^1$ to $Z=\bar{z}^1$. This effectively shifts the selection probability from $p$ to $\bar{p}$ by exogenously changing $Z_0$ from $\tilde{z}_0$ to $\tilde{z}_0'$ while holding $Z_1=z_1$ fixed. Similarly, to identify $m_0(p, z_0)$, we use the shift represented by the \textcolor{mtr0color2}{orange arrow}, from $Z=z^0$ to $Z=\bar{z}^0$, which shifts the selection probability from $p$ to $\bar{p}$ by exogenously changing $Z_1$ from $\tilde{z}_1$ to $\tilde{z}_1'$ while holding $Z_0=z_0$ fixed. Since it is not possible to move the probability while holding both $Z_0$ and $Z_1$ fixed, the two shifts identifying the marginal treatment responses will always differ. Note that we generally do not use the point $(z_0, z_1)$ for identification, and it is not necessary that $P(z_0, z_1) = p$ to identify $m_0(p, z_0)$ or $m_1(p, z_1)$.\footnote{One could define the MTE at $Z=z=(z_0, z_1)$, with $P(z)=p$, in an alternative notation as $\widetilde{\text{MTE}}(z)$, equal to MTE$(P(z), z_0, z_1)$ in the general notation. Then for each $d$, one can identify the newly defined MTR, $\tilde{m}_d(z)$ $= m_d(p, z_d)$ by marginally shifting $Z_{1-d}$ while holding $Z_d=z_d$ fixed. These treatment parameters are tied to specific values of the semi-IVs and are thus less general than the main definition. One advantage, however, is that these parameters can be identified while relaxing the monotonicity assumption. Indeed, since we focus on marginal shifts, unordered partial monotonicity combined with a comparable complier assumption mountjoy2022community are sufficient to identify $m_d(z)$ and MTE$(z)$. See Online Appendix (ref). }

Overall, the identification of both MTRs, and of the MTE, at a given $V=p$ and $(Z_0, Z_1)=(z_0, z_1)$, requires that there exists four points as illustrated in Figure (ref): for each $d$, there must be two points, $Z=z^d$ and $\bar{z}^d$, which both have $Z_d=z_d$ and are such that $P(z^d) = p$ and $P(\bar{z}^d)=\bar{p}$ in order to identify $m_d(p, z_d)$. In other words, we need points with $Z_1=z_1$, and other points with $Z_0=z_0$, on both isocurves of probability $p$ and $\bar{p}$, allowing the propensity score to shift from $p$ to $\bar{p}$ while holding either $Z_1=z_1$ or $Z_0=z_0$ fixed. \\ When we focus on the MTRs and MTE with continuous semi-IVs, if $z^1$ with $Z_1=z_1$ exists on the isocurve of probability $p$, the existence of $\bar{z}^1$ is guaranteed by the fact that $V=p$ is in the interior of the support of the propensity score given $Z_1=z_1$ (conversely for $z^0$ and $\bar{z}^0$), hence the previous conditions for the identification of the MTRs. \\ In general, to identify the average treatment effects on a larger set of compliers with $V \in [p, \bar{p}]$ for any general $\bar{p} > p$ (i.e., not only marginal changes in $P$), the requirement can be directly stated in terms of the existence of these four points. This corresponds to the identification of the general LATE$(p, \bar{p}, z_0, z_1)$ that is developed in the next subsection. \\ To contextualize this graph, in the application of earnings returns to manufacturing sector choice, identifying returns for workers with resistance $V=v$ in an environment with manufacturing size $Z_1=z_1$ and nonmanufacturing size $Z_0=z_0$ requires observing two environments with different relative sector sizes but the same propensity to work in manufacturing, $P=v$. The mean manufacturing potential earnings at $V=v$ and $Z_1=z_1$ are identified by marginally decreasing nonmanufacturing size (from $\tilde{z}_0$ to $\tilde{z}_0'$), which increases the propensity score from $p$ to $\bar{p}$ while holding $Z_1=z_1$ fixed. The resulting change in average $DY$ is due solely to the entry of the marginal compliers with $V=v$ into treatment, identifying their mean manufacturing earnings. A converse approach identifies the mean nonmanufacturing potential earnings. The difference between these two mean potential earnings gives the MTE, i.e., the returns to working in manufacturing at $V=v$ given $Z_1=z_1, Z_0=z_0$. Overall, identification relies on comparing four otherwise similar environments that differ in their semi-IV combinations, thus providing the required variation in treatment incentives.

Local average treatment effect (and responses)

Treatment parameters definition

Following heckmanvytlacil2005, define the generalized local average treatment effect (LATE) parameters as a function of the set of compliers they correspond to (in terms of $V$):

adjustwidth{-1cm}{-0.5cm} \begin{align} LATE(v, v', z_0, z_1, x) = \mathbb{E}[Y_1 - Y_0 | v \leq V < v', Z_0=z_0, Z_1=z_1, X=x ]. \end{align}

The $\text{LATE}(v, v', z_0, z_1, x)$ is the average causal effects of $D$ on $Y$ for individuals with unobserved resistance to treatment $V \in [v, v')$ and observed characteristics $X=x, Z_0=z_0, Z_1=z_1$. It can also be interpreted as the mean gains for the individuals (compliers) who would be induced to switch into treatment if the probability of treatment changed from $v$ to $v'$, at fixed $X=x, Z_0=z_0, Z_1=z_1$. This interpretation allows detaching the definition of the LATE parameter from the specific semi-IV shift that identifies it, which is especially convenient with semi-IVs. For a more standard definition and identification of the LATE$(z, z', x)$ tied to specific semi-IV shifts from $Z=z$ to $z'$, à la imbensangrist1994, see Online Appendix (ref). Since imbensangrist1994's LATE can always be mapped to the general LATE (ref), I work with the latter for convenience due to its direct link with the MTE. \\ Following the split of MTE in MTRs, define the local average treatment responses (LATR):

align[align omitted — 126 chars of source]

For each $d$, the LATR$_d$ represents the mean potential outcome $Y_d$ for individuals with $V \in [v, v')$ and covariates $X=x, Z_d=z_d$. Because of the exclusion of the semi-IVs, the LATR$_d$ does not depend on $Z_{1-d}$, and this is what we exploit for identification. \\ Similar to the MTE-MTR link in (ref), the LATE and LATRs are related by

align[align omitted — 132 chars of source]

Thus, the LATE is identified if one can identify both LATR$_d$ for compliers with $V \in [v, v')$.

Identification

Again, for notational convenience, omit $X$ and proceed conditional on $X=x$. \\

Identification using the MTRs and MTE. By definition, note that the LATR and the LATE are directly related to the MTR and MTE by

align[align omitted — 211 chars of source]

Therefore, $\text{LATE}(v, v', z_0, z_1)$ is identified if MTE$(\tilde{v}, z_0, z_1)$ is identified for all $\tilde{v} \in [v, v')$, i.e., if for each $d=0, 1$, MTR$_d(\tilde{v}, z_d)$ is identified for all $\tilde{v} \in [v, v')$. \\

Direct identification (Figure (ref)). The LATE identification via the MTE is not possible if continuously distributed semi-IVs are unavailable to generate continuous variation in both the propensity score given $Z_1=z_1$ and given $Z_0=z_0$ over the entire range $P \in [v, v')$. Fortunately, following the intuition from Figure (ref), the identification of this general LATE can be done directly using only four points (i.e., two semi-IV shifts) without identifying all the intermediary MTRs. The intuition for identifying the MTE extends directly to larger shifts in the propensity score, from $P=v$ to $P=v'$. Using (ref), we have:

align*[align* omitted — 123 chars of source]

So, if there exists $z^1 = (\tilde{z}_0, z_1)$ with $P(z^1) = v$ and $\bar{z}^1 = (\tilde{z}_0', z_1)$ with $P(\bar{z}^1) = v'$, then LATR$_1(v, v', z_1)$ is identified by

align[align omitted — 231 chars of source]

This follows the standard IV interpretation, applied to the outcome $DY$: holding $Z_1=z_1$ fixed, shifts in $Z_0$ affect $DY$ only indirectly through their effects on selection probabilities (see Property (ref)). At fixed $Z_1$, the semi-IV $Z_0$ serves as an IV for the effect of $D$ on $DY$. Thus, observable differences between $\mathbb{E}[DY|Z_1=z_1, Z_0=\tilde{z}_0']$ and $\mathbb{E}[DY|Z_1=z_1, Z_0=\tilde{z}_0]$ stem solely from the effect of shifting $Z_0$ from $\tilde{z}_0$ to $\tilde{z}_0'$ on selection and the corresponding set of compliers, $V \in [v, v')$. This shift corresponds to the \textcolor{mtr1color2}{purple arrow} in Figure (ref). \\ Similarly, $Z_1$ serves as an IV for the effect of $D$ on $(1-D)Y$ given $Z_0$. Thus, if $z^0 = (z_0, \tilde{z}_1)$ with $P(z^0) = v$ and $\bar{z}^0= (z_0, \tilde{z}_1')$ with $P(\bar{z}^0)=v'$ exist, LATR$_0(v, v', z_0)$ is identified by\footnote{The minus sign comes from the same reason as in (ref), see footnote (ref).}

align[align omitted — 148 chars of source]

This identifying semi-IV shift corresponds to the \textcolor{mtr0color2}{orange arrow} in Figure (ref).

Therefore, for any $v, v' \in [0, 1]$, LATE$(v, v', z_0, z_1)$ is identified by Equation (ref), i.e., by

align*[align* omitted — 100 chars of source]

provided that (at least) four semi-IVs combinations exist: (i) $z^1 = (\tilde{z}_0, z_1)$ with $P(z^1) = v$ and $\bar{z}^1 = (\tilde{z}_0', z_1)$ with $P(\bar{z}_1) = v'$, ensuring that LATR$_1(v, v', z_1)$ is identified by (ref), and (ii) $z^0 = (z_0, \tilde{z}_1)$ with $P(z^0) = v$ and $\bar{z}^0= (z_0, \tilde{z}_1')$ with $P(\bar{z})=v'$, ensuring that LATR$_0(v, v', z_0)$ is identified by (ref). Again, the monotonicity assumption plays a key role: it ensures that, although LATR$_1$ and LATR$_0$ are identified from two different semi-IV shifts, both shifts induce comparable sets of compliers, with $V \in [v, v')$, because they result in the same observed change in the propensity score, from $v$ to $v'$. Therefore, one can take the difference between these two LATRs to identify the LATE for $V \in [v, v')$. \\

Discussion. LATE parameters for compliers with $V \in [v, v')$ at $Z_1=z_1, Z_0=z_0$ can be identified if, for both $d$, there exist two semi-IV combinations with $Z_d=z_d$, one yielding $P=v$ and the other $P=v'$ (see Figure (ref)). With at least one continuously distributed semi-IV, a broad set of LATE is likely identifiable. With two discrete semi-IVs, exact complier alignment on the same propensity score values is unlikely. However, provided that the discrete supports are large enough, one can interpolate the LATRs in between the support points, and recover (or at least, bound) some LATE parameters accordingly. In the unfortunate case where the researcher only has two binary semi-IVs ($Z_d \in \{0, 1\}$), identifying any LATE requires specific conditions that the effect of the semi-IVs on the selection probabilities compensate each other exactly.\footnote{More precisely, if $P(0, 0) = P(1, 1)=p$, then LATE$(p, P(1,0), 0, 1)$ and LATE$(p, P(0,1), 1, 0)$ are identified, while if $P(0,1)=P(1,0)=p$ then LATE$(p, P(0,0), 1, 1)$ and LATE$(p, P(1,1), 0, 0)$ are identified. } This is very unlikely to hold, meaning that, contrary to binary IVs, the LATE can generally not be identified with two binary semi-IVs. \\ Finally, note that the environment with $Z=(z_0, z_1)$ is generally not used to identify LATE$(v, v', z_0, z_1)$. Contrary to the standard LATE identification of imbensangrist1994 (see Online Appendix (ref)), the identification of the parameters is detached from the specific value of the semi-IVs. However, in the special case where $P(z_0, z_1)=v$ (resp. $v'$), then identification is simplified because it only requires two additional points belonging to the isocurve of probability $v'$ (resp. $v$): one with $Z_1=z_1$ and another one with $Z_0=z_0$.

Other treatment effects

Direct effect of the semi-IVs (targeted-policy evaluation). In some applications, especially when the semi-IVs are alternative-specific policies, one may be interested in identifying the direct effect of these semi-IVs on their respective potential outcomes, net of the selection/compositional changes. Online Appendix (ref) describes how to identify these effects. \\

Policy relevant treatment effects (PRTE). With MTRs identified with semi-IVs, any parameter expressed as a weighted function of the MTRs is also identified. The generalized LATE is already a specific (with unit weights) policy-relevant treatment effect (PRTE). More broadly, mogstad2018ecta (Table 1) provide a list of parameters expressed as weighted functions of the MTRs, along with their corresponding weights with IVs. Many of these parameters and weights can be adapted to identify their counterparts using semi-IVs. \\ Similarly, when the semi-IVs have limited support -- which restricts the observable propensity score and the range of unobserved $V$ for which MTRs are identified -- treatment effects can be extrapolated using approaches analogous to those in mogstad2018ecta for IVs.

Estimation

We observe a sample of $\{Y_i, D_i, Z_{0i}, Z_{1i}, X_i\}_{i=1}^N$. Each $Z_d$ may include multiple valid, continuously distributed semi-IVs. Building on constructive identification arguments, we estimate the MTR$_d$ functions separately to obtain the MTE (and LATE).\footnote{This section focuses on estimating the MTE and MTR but can be adapted to the estimation of the LATE and LATR with only discrete semi-IVs using a similar control function approach for $P$.} The procedure involves estimating three key objects: the propensity score $P(Z_0, Z_1, X)$ and the two conditional expectations $\mathbb{E}[DY | P, Z_1, X]$ and $\mathbb{E}[(1-D)Y|P, Z_0, X]$, which are used to estimate $m_1$ and $m_0$ as in Equations (ref) and (ref). The MTE is then obtained as their difference (Equation (ref)) on the support where both MTRs are identified. \\ I propose three estimation approaches. First, a fully nonparametric approach. Second, a more practical semi-parametric approach, which is the semi-IV counterpart to the typical MTE estimation with IVs carneiroheckmanvytlacil2011, andresen2018exploring. Finally, I present a revisited $2$SLS for homogenous treatment effect models with semi-IVs.

Nonparametric estimation

I estimate the three objects of interest using local linear regressions, similar to mountjoy2022community, who studies a discrete treatment problem with multiple IVs. For practicality, assume that the functions of interest are globally linear in $X$ to avoid the curse of dimensionality. \\ First, estimate the propensity score $P$ via local linear regression of $D$ on $Z_0, Z_1$ and $X$, using a bivariate kernel on $(Z_0, Z_1)$ for every values of $(z_0, z_1) \in \mathcal{Z}$.\footnote{Assume $Z_0$ and $Z_1$ each contain a single semi-IV; otherwise, estimation quickly becomes untractable. } \\ Next, use the fitted $\hat{P}_i$ as a generated covariate in the second stage to estimate the conditional expectations of the outcome for $d=0$ and $d=1$ by

adjustwidth{-2cm}{-2cm} \begin{align*} &\begin{pmatrix} \hat{\beta}^{Y_d}_{Z_d}(z_d, p) \\ \hat{\beta}^{Y_d}_{P}(z_d, p) \\ \hat{\beta}^{Y_d}_{X}(z_d, p) \end{pmatrix} = \underset{\beta_{Z_d}, \beta_{P}, \beta_X}{argmin} \sum_{i=1}^N K\left( \frac{Z_{di}-z_d}{h_d}, \frac{\hat{P}_{i} - p}{h_P} \right) \left(Y_i\mathds{1}\{D_i=d\} - \beta_{Z_d} Z_{di} - \beta_{P} \hat{P}_{i} - X_i' \beta_X \right)^2. \nonumber \end{align*}

For all values of $z_d$ and $p$, $X$ includes a constant, and $K(\cdot)$ is a bivariate kernel with bandwidths $h_d$ for each $Z_d$ and $h_P$ for the fitted $\hat{P}$. \\ Local linear specifications are attractive because, or any observable $(z_d, p)$, the local coefficient on $P$ directly estimates MTR$_d$:

align*[align* omitted — 175 chars of source]

Then, given any $(z_0, z_1, p)$ for which both MTRs are identified and estimated, the MTE is estimated as their difference, using the empirical counterpart of Equation (ref):

align*[align* omitted — 79 chars of source]

For inference, nonparametric bootstrap is used to obtain proper standard errors accounting for the estimation of $\hat{P}$ and its use as a generated regressor in the second stage.

Semi-parametric estimation

I show how the local IV estimation heckmanvytlacil1999, particularly the approach estimating both MTRs separately andresen2018exploring with IVs, adapts to semi-IVs. \\

The semi-parametric model. The treatment is still determined by the weakly separable selection rule (ref). The potential outcomes are given by

align[align omitted — 116 chars of source]

where $U_d$ is a $d$-specific unobserved shocks affecting the potential outcomes, and $\mu_d$ is a $d$-specific parametric function, that is specified as linear for simplicity.\footnote{Linearity can be relaxed as long as $\mu_d$ follows a known parametric form. Since the dimensions of $X$ and $Z_d$ are unspecified, even a simple linear model can be made more flexible by adding interactions between $X$ and $Z_d$ or incorporating flexible functions of $Z_d$ (e.g., splines or polynomials).} The joint distribution of $(U_0, U_1, V)$ is unrestricted, allowing for endogeneity between $Y_d$ and $D$. Under this model, the partial exclusion of the semi-IVs can be written as $(U_d, V) \perp\!\!\!\perp Z_{1-d} | (Z_d, X)$ for each $d=0,1$. Property (ref) can be expressed in terms of $U_d$ by

align[align omitted — 128 chars of source]

Separability. In addition, following the empirical literature estimating MTE carneiroheckmanvytlacil2011, maestas2013does, eisenhauer2015generalized, brinch2017beyond, cornelissen2018benefits, and existing estimation packages brave2014estimating, andresen2018exploring, I impose the following additional separability assumption to simplify the estimation.\footnote{A stronger version of this separability assumption imposing that $U_d \perp\!\!\!\perp (Z_0, Z_1, X)$ for $d=0, 1$ is sometimes assumed instead. But it is stronger than necessary, especially for $X$, because Assumption (ref) does not restrict the dependence between $X$ and $V$. Since $Z_d \perp\!\!\!\perp V$, the distinction is less important for the semi-IVs. See the Section $6.2$ of mogstad2018review for more discussions. }

assumption[Separability] For each $d \in \{0,1\}$, $\mathbb{E}[U_d | V, Z_0, Z_1, X] = \mathbb{E}[U_d | V].$

Separability is a stronger assumption than the implication (ref) derived from the exclusion restriction in the general model. Importantly, separability does not imply that the potential outcome $Y_d$ is independent of $Z_d$ and $X$. Rather, it restricts their dependence to be fully captured by the conditional mean function, $\mu_d(X, Z_d)$. Combined with the additive separability in (ref), it implies that the effect of $X$, and more importantly, of $Z_d$ on $Y_d$ are homogenous with respect to $U_d$, and thus, with respect to $V$. While this restriction is strong, it is commonly assumed (with respect to $X$) in the empirical literature estimating MTE. Since $Z_d$ behaves like a covariate $X$ for $Y_d$, it is natural to extend the separability assumption to $Z_d$. Treatment effects remain generally heterogenous under separability. \\ Even under this stronger separability assumption, the exclusion of $Z_{1-d}$ from $Y_d$ still holds only conditional on $Z_d$ (and $X$), i.e.,

align*[align* omitted — 82 chars of source]

Thus if one does not control for $Z_d$, one cannot exclude $Z_{1-d}$ from the average $Y_d$, since $Z_{1-d}$ may indirectly affect the average $Y_d$ through its correlation with $Z_d$ in $\mu_d(X, Z_d)$. \\

MTRs. Under Assumption (ref) and model (ref), the MTRs are

align[align omitted — 230 chars of source]

In other words, the mean potential outcomes, $m_d$, can be decomposed into two components: the effect of the observables, $\mu_d(x, z_d)$, and the mean unobservable $U_d$ given $V=v$, $k_d(v)$, which captures selection effects. While more restrictive than the general model, this structure has practical advantages. The direct effects of semi-IVs (and covariates) on $Y_d$ are homogenous and more easily interpretable as they do not vary with the unobserved resistance to treatment $V$. They are also easier to identify.\footnote{For any $z_d, z_d', x$, one can identify $\mu_d(x, z_d') - \mu_d(x, z_d)$ by controlling for $P$, for any $p \in \mathcal{P}$. Indeed,

align*[align* omitted — 129 chars of source]

because $\mathbb{E}[U_d | D=d, Z_d, X, P] = \mathbb{E}[U_d | D=d, P]$ is independent of $Z_d$ and $X$. So the direct effect of changing $Z_d$ from $z_d$ to $z_d'$ (at $X=x$) is identified if one can shift $Z_d$ while holding $P$ fixed (for any $p$), which is feasible if $Z_{1-d}$ is relevant. With the linear specification, identification is even simpler: it suffices that there exist two $z_d \neq z_d'$ that yield the same $P=p$, for any $p \in \mathcal{P}$. In this case, $\delta_d$ is identified by:

align*[align* omitted — 142 chars of source]

} Moreover, since $k_d(v)$ is independent of $(X, Z_d)$, the full support of the propensity score (given $D=d$) can be used for identification, rather than only its support conditional on specific $Z_d=z_d, X=x$ (and $D=d$). \\

Partially linear model. Under separability, the observed average outcomes in the treated and untreated subsamples of the semi-parametric model (ref) follow a partially linear model:

align[align omitted — 162 chars of source]

where $\mathbb{E}[U_d | D=d, P, Z_d, X] = \mathbb{E}[U_d | D=d, P] =: \kappa_d(P)$, a nonparametric function of $P$ under separability.\footnote{The conditioning on $P=p$ remains because $D=1$ is equivalent to $V \leq p$ and $D=0$ to $V > p$.} Thus, average outcomes in each subsample follow a partially linear model with a linear component in covariates and semi-IVs, and an additive nonparametric term in $P$. \\ To estimate the MTRs (ref), it suffices to consistently estimate the partially linear model (ref), which directly provides an estimate for $\beta_d$ and $\delta_d$, and allows recovering $k_d(v)$ as a deterministic function of $\kappa_d(v)$. Indeed, $\kappa_1(p) = \mathbb{E}[U_1 | V \leq p] = \int_0^p \mathbb{E}[U_1|V=v]/p\ dv $ and $\kappa_0(p) = \mathbb{E}[U_0 | V > p] = \int_p^1\mathbb{E}[U_0 | V=v]/(1-p) dv$. Differentiating these gives

align*[align* omitted — 146 chars of source]

Estimation. Following the literature heckmanetal1998, heckmanurzuavytlacil2006, I apply nonparametric robinson1988root's double residual regression to estimate the partially linear models separately for treated and untreated subsamples.\footnote{Alternative estimation. robinson1988root's double residual regression is the most popular procedure to nonparametrically control for $P$, but alternatives exist. For instance, $\kappa_d(p)$ can be specified as flexible polynomials or splines, following a nonparametric sieve approach. In this case, after estimating $P$ in a first stage, the model can be estimated by running the regression of $Y$ on $X, Z_d,$ and a flexible function of $P$ in each subsample with $D=d$. If the chosen specification for $\kappa_d(p)$ is sufficiently flexible, this yields consistent estimates of $\beta_d$, $\delta_d$, and $\kappa_d(p)$ for values of $p$ in the interior of the support of $P$ given $D=d$.} The estimation can be decomposed into three steps. In the first stage, the propensity score is estimated. Then, for each subsample of observations with $D=d$ separately, we estimate the effects of the semi-IVs and covariates on $Y$ (thus, on $Y_d$ in the subsample) by controlling for $P$ nonparametrically. More precisely, nonparametrically regress $Y, X, Z_d$ on $P$, and then run the linear regression of the residuals of $Y$ on the residuals of $X$ and $Z_d$ to estimate $\hat{\beta}_d$ and $\hat{\delta}_d$. Finally, we estimate the remaining selection effects, $k_d(v)$. To do so, nonparametrically regress $Y$ minus the effects of the covariates and semi-IVs, i.e., $Y-\hat{\beta}_d X + \hat{\delta}_d Z_d$, on $P$. The detailed estimation procedure with semi-IVs is described in Online Appendix (ref). \\ These three steps parallel those of MTE estimation with IVs andresen2018exploring.\footnote{For more details, see also carneiroheckmanvytlacil2011, Appendix A.2.1. of heckmanetal1998, or the online supplement of heckmanurzuavytlacil2006.} The key difference is the presence of the direct effects of the semi-IVs, $\delta_d$. Identifying these requires being able to fix the propensity score (i.e., to control for $P$) while shifting $Z_d$. This is achievable by also shifting the other excluded semi-IV, $Z_{1-d}$, to offset the effect of $Z_d$ on $P$. Note that covariates are already incorporated in the standard MTE estimation with IVs andresen2018exploring. Thus, the effects of the semi-IVs are naturally incorporated as additional coefficients estimated in the second stage. Therefore, MTE estimation with semi-IVs is similar but also requires the relevance of the semi-IVs to provide identifying excluded variation. \\

Estimated MTR and MTE. For each $d$, once $\beta_d, \delta_d$ and $k_d(v)$ are estimated, the MTR$_d$ function is obtained by plugging these into (ref) for any $v$ in the interior of the support of $\widehat{P}$ given $D=d$, and for any $(z_d, x)$ in the joint support of $Z_d$ and $X$ given $D=d$, i.e.,

align*[align* omitted — 105 chars of source]

Similarly, the MTE is estimated for any $(z_0, z_1, x) \in \mathcal{Z}\times\mathcal{X}$ and for any $v$ in the common support of $\widehat{P}$, i.e., at the intersection of the supports of $\widehat{P}$ given $D=1$ and given $D=0$:\footnote{In practice, there are very few observations at the extreme ends of the propensity score support, leading to poor estimation of the MTRs in those regions. To address this, it is common to trim the supports further to focus on well-estimated interior points brinch2017beyond. For example, one might restrict $P$ in each subsample to values between the $1\%$ and $99\%$ percentiles of its conditional distribution given $D=d$ and define the common support as the intersection of these trimmed supports.}

align*[align* omitted — 287 chars of source]

Under separability, the MTE can be decomposed into two components: one driven by observables, including the semi-IVs, and another driven by the selection. Note that the MTRs and MTE still depend on the covariates and semi-IVs. Typically, we report MTE/MTR curves at average covariates and semi-IV values (i.e., at $\bar{x}, \bar{z}_0, \bar{z}_1$) across all identified values of $v$ (see Figure (ref) for example). Once the MTRs and MTE are estimated, they can be used to estimate more complex parameters such as the LATR and LATE, or other PRTEs. \\

Inference. As in its IV counterpart, raw analytical standard errors from the final-stage local polynomial estimation of the MTR and MTE are biased, as they ignore that the propensity score and the effect of the semi-IVs and covariates were estimated in previous stages. To account for this, nonparametrically bootstrapping the standard errors is recommended. \\

Performances. The semi-parametric estimator performs well with reasonable sample sizes, as illustrated using Monte Carlo simulations in Online Appendix (ref). \\

Practical implementation. The MTE and MTR estimation can be performed using the semiivreg() command from the semiIVreg package semiivreg. The syntax of the function is designed to mimic the ivreg() command for IV estimation.

lstlisting[lstlisting omitted — 75 chars of source]

Two-Stage Least Squares with semi-IVs

Consider a simpler model with homogenous treatment effects:

align[align omitted — 120 chars of source]

where $\mathbb{E}[U] = 0$ and $\mathbb{E}[U | Z_0, Z_1, X] = 0$ serves as the counterpart to the separability Assumption (ref) in this model. The treatment is still determined by Equation (ref). This is a model with homogenous treatment effects with respect to the unobservables, meaning that MTE$(v, z_0, z_1, x)$ is independent of $v$. However, the average treatment effect remains heterogenous with respect to observables. Indeed

align*[align* omitted — 153 chars of source]

Even when covariates $X$ have the same effects on both potential outcomes ($\beta_1=\beta_0$), the ATE still depends on the semi-IVs. Thus, we typically report specific ATEs, such as the ATE at the mean, ATE$(\bar{z}_0, \bar{z}_1, \bar{x})$. \\ Standard IV-GMM estimation does not directly apply to Model (ref) since the number of parameters exceeds the number of instruments. However, the model can be estimated using a modified two-stage least squares (2SLS) approach. In the first stage, estimate $P(Z_0, Z_1, X)$. Then run a 2SLS of Equation (ref) instrumented by the optimal instruments, $\widehat{P}, \widehat{P}Z_1, (1-\widehat{P})Z_0$ chamberlain1987asymptotic. This yields consistent estimates of all parameters $\alpha, \delta, \delta_0, \delta_1, \beta_0$, and $\beta_1$.\footnote{If the true underlying model is a model with heterogenous treatment effects, e.g., Model (ref), the question of whether the estimates from this semi-IV-2SLS can be interpreted as convex combinations of positively weighted semi-IV LATE (like standard IV 2SLS) remains open.}

Application: returns to working in manufacturing

Context: the manufacturing decline

figure[figure omitted — 1,648 chars of source]

The manufacturing sector's employment share has steadily declined in the US over the past decades baily2014us, pierce2016surprisingly, fort2018new. As shown in Figure (ref), from $1999$ to $2018$, its share fell from $13.9\%$ to $8.8\%$ of total US employment: a $36.7\%$ relative decline over $20$ years. Even among young white men with only a high school education, a group more likely to work in manufacturing, the sector's employment share dropped from $23\%$ to $18\%$ in the same period, a $21.7\%$ relative decline. \\ This decline in manufacturing coincides with the rise of the nonmanufacturing sector, particularly services lee2006intersectoral, buera2012rise, autor2013growth. As Figure (ref) shows, US manufacturing GDP (in base-1999 dollars) barely changed from $1,462$ billion in $1998$ to $1,433$ billion in $2017$ (a $2\%$ decline), while nonmanufacturing grew by $52.5\%$, from $7,800$ to $11,895$ billion over the same period. \\

figure[figure omitted — 752 chars of source]

A naive look at earnings by sector (Figure (ref)) suggests that (i) manufacturing workers earn $13.5\%$ more on average than their nonmanufacturing counterparts throughout the period and, perhaps surprisingly, (ii) this gap increased from $13.2\%$ in $1999$ to $17.2\%$ in $2018$. \\ However, this naive comparison is likely biased due to selection on unobservables: even after controlling for demographics (age, education, race), manufacturing workers systematically differ from those in other sectors heckmansedlacek1985. In principle, an IV could address the endogeneity of sector choice, but finding one that affects sector choice without also affecting subsequent earnings is difficult. Fortunately, measures of local sector-specific strength can serve as valid semi-IVs instead, allowing for unbiased estimates of the returns to manufacturing (net of selection) and their evolution over time.\footnote{I remain agnostic about the reasons behind the evolution of returns to working in manufacturing. The literature highlights two candidate explanations: (i) technological change, which affects the task content of manufacturing jobs autor2003skill, goos2014explaining, and (ii) trade, with rising import competition autor2013china. autor2015untangling propose methods to disentangle these two explanations.}

Model specification and semi-IVs

I estimate the effect of working in manufacturing ($D=1$) on the logarithm of the yearly earnings of workers ($Y$), measured in thousands of real $1999$ US dollars.\footnote{The manufacturing sector corresponds to the NAICS supersectors "31-33". The nonmanufacturing sector includes all other industries, private or public, except military jobs. } I focus on young white male workers aged 18 to 30 with only a high school diploma and no college education, a demographic more likely to work in manufacturing and thus more affected by the sector's decline. To address the endogeneity of sector choice, I use the one-year lagged state-level manufacturing GDP ($Z_1$) and nonmanufacturing GDP ($Z_0$), measured in millions of $1999$ US dollars, as semi-IVs. These should be valid as discussed in Example 1 of Section (ref). \\ I estimate the following semi-parametric model for the potential earnings:

align[align omitted — 197 chars of source]

which corresponds to Model (ref) with covariates $X$ including age, age$^2$, and state and year fixed effects. Sector-specific ($d$-specific) effects are estimated for all covariates, including fixed effects. The sector-specific state and year fixed effects control for intrinsic differences in state-level sectoral markets and national trends, capturing sector-specific costs, technological changes, and trade shocks. The sector-specific time fixed effects are particularly interesting as they track the evolution of earnings in both sectors and, consequently, the returns to working in manufacturing over time. With these fixed effects, the semi-IVs' identifying variations do not come from their levels but from deviations relative to sector-specific permanent state levels and the global time trend.\footnote{The semi-IV effects $\delta_d$ are assumed to be homogenous across time and place, allowing identification using local variations in the semi-IVs over time. Identification would not be possible with interacted year-state fixed effects or if the semi-IVs had state- and year-specific effects, such as $\delta_{d, state, year}$.} I assume time fixed effects in the outcome and selection equations absorb all temporal variation in $U_d$ and $V$. I therefore omit the $t$-index and treat these shocks as time-invariant, implying time-invariant $k_d(v)$ in the MTRs/MTE. This simplification permits pooled estimation across periods rather than separate period-specific models that rely only on within-period, cross-sectional (location-specific) variation. Since $Y_d$ are not directly observed with only data on $(Y, D, Z, X)$, Model (ref) cannot be estimated directly. Instead, it is estimated following the semi-parametric procedure in Section (ref).

Data

The individual-level data on earnings, employment, and sector choice for young white men come from the ACS $1\%$ yearly survey waves (2000-2019), effectively covering earnings from 1999 to 2018.\footnote{As the name suggests, each survey wave includes a random, representative sample of $1\%$ of the US population. However, from 2000 to 2004, the coverage was smaller, with the 2000-2004 waves containing only $0.13\%$, $0.43\%$, $0.38\%$, $0.42\%$, and $0.42\%$ of the population, respectively. } I include all US states except Hawaii, Alaska, and Washington D.C., as these are outliers in manufacturing employment, with some years in the sample having no observed young white male manufacturing workers there. State-level sector-specific GDP data come from the Bureau of Economic Analysis (BEA). All monetary variables are adjusted to real 1999 US dollars for consistency and comparability.

Results

table[table omitted — 1,423 chars of source]

Propensity score estimation ($1^{\text{st}}$ stage)

Propensity score. I estimate the propensity score using a probit regression of $D$ on log($Z_0$), log($Z_1$), age, age$^2$, and state and year fixed effects. The results are reported in Column $1$ of Table (ref). Both semi-IVs are relevant for sector choice $D$. As expected, all else equal, a larger (lagged) manufacturing GDP increases the probability that young workers select manufacturing jobs, while a larger nonmanufacturing GDP reduces it. Several factors may explain this. For instance, a larger sector is more attractive to young workers. Additionally, higher past GDP also allows the sector to recruit more in subsequent periods, leading to a passthrough effect from lagged GDP to future hires. \\ Figure (ref) provides a more detailed view of the distribution of $P$ conditional on $Z_1$. It is visible that $Z_0$ is relevant because, given any fixed $Z_1 = z_1$, $P$ still exhibits substantial variation. This is the variation used for identification (while further controlling for covariates). This shows that, even nonparametrically, we could identify large sets of MTE and LATE parameters at any fixed $Z_1=z_1$. \\

figure[figure omitted — 1,242 chars of source]

Common Support. Figure (ref) reports the distribution of the estimated propensity score in both the manufacturing (treated) and nonmanufacturing (untreated) subsamples. Overall, the propensity score ranges from $3.19\%$ to $42.12\%$, but the lowest and highest values do not appear in both subsamples. Since the estimation is noisier at the tails, I further trim the data by dropping observations with estimated propensity score below $7.9\%$ or above $33.9\%$. This corresponds to the common support after trimming the $2.5\%$ lowest and highest values of $P$ in each subsample.\footnote{All results are robust to less trimming, e.g., $1\%$, or even $0\%$, i.e., restricting to the common support.} As a result, MTE and MTRs are only estimated for unobserved resistance to treatment $V \in \mathcal{P} = [0.079, 0.339]$.\footnote{Following the applied literature with IVs carneiroheckmanvytlacil2011, the common support could be expanded by including additional semi-IVs or modeling interacted effects of the semi-IVs and covariates on $D$. Including more covariates would also widen the support of $P$, but this would rely on the separability assumption. Instead, we prefer to use variation induced by the semi-IVs for identification. }

figure[figure omitted — 3,295 chars of source]

Effects on earnings ($2^{\text{nd}}$ stage)

Given the semi-parametric model (ref), the MTR take the form (ref) as in Model (ref), i.e.,

align*[align* omitted — 202 chars of source]

First, I discuss the estimated effects of the semi-IVs (Table (ref)), then the estimated MTRs and MTE evaluated at average covariates and semi-IVs (Figure (ref) and Table (ref)). \\

Direct effects of the semi-IVs. Table (ref) (Columns $2$-$4$) reports the estimated effects of the semi-IVs on $Y_0$, $Y_1$, and $Y_1-Y_0$, respectively. Each semi-IV significantly affects earnings in its corresponding sector. A $1\%$ increase in nonmanufacturing GDP raises the nonmanufacturing workers' earnings by $0.438\%$ in the next period. The pass-through from lagged manufacturing GDP to manufacturing workers' earnings is significantly smaller: a $1\%$ increase raises next-period manufacturing earnings by $0.146\%$. Note that the ratio of the effects of the semi-IVs on their respective potential outcomes is close to the ratio of their effects on the first stage. This suggests selection on gains, where the semi-IVs primarily influence the first stage through their effect on $Y_d$. Thus, selection is mainly driven by the comparison of the semi-IVs' effects rather than by each semi-IV individually. \\

Heterogenous returns to working in manufacturing. Figure (ref) reports the estimated MTR$_d$ (left panel) and MTE (right panel) curves at $V=v$ for an average individual with $X=\bar{x}, Z_0=\bar{z}_0, Z_1=\bar{z}_1$.\footnote{Due to the partially linear model, the MTR and MTE at the average are equivalent to the average MTR and MTE at $V=v$ in the sample. For the fixed effects, the "average fixed effects" already correspond to the sample's average distribution of state and year. } These can only be estimated within the common support, i.e., for $V \in [0.079, 0.339]$. Over the entire common support, the estimated average earnings return to working in manufacturing (LATE) is $7.5\%$ and not significantly different from zero (Table (ref), first row). However, I find significant heterogeneity in the average potential outcomes (MTR) and treatment effects (MTE) across $V$ (Figure (ref)). Individuals with low resistance to treatment (small $V$) have, on average, high manufacturing earnings and low nonmanufacturing earnings. This is consistent with selection on gains: these individuals are the first to select into manufacturing because they benefit the most from it. For them, the returns to working in manufacturing are high: they earn more than twice as much in the manufacturing sector as in the nonmanufacturing sector (MTE $>1$). As $V$ increases, average manufacturing earnings decrease while nonmanufacturing earnings increase. Around $V=20\%$, the returns to working in manufacturing become negative. \\ This heterogeneity is also reflected in the LATR/LATE estimates (Table (ref)), obtained by integrating the MTR/MTE. Among individuals with relatively low resistance to treatment ($V \in [10\%, 15\%]$), average log earnings are $2.918$ in manufacturing (LATR$_1$) versus $2.184$ in nonmanufacturing (LATR$_0$). This results in a large positive LATE of $0.734$, meaning that they earn, on average, $108\%$ more in manufacturing. Conversely, workers with relatively high $V$ ($V \in [25\%, 30\%]$) earn about $34.8\%$ less in manufacturing (LATE$=-0.428$). \\

comment\begin{table}[!t] \captionsetup{justification=centering} \caption{LATE of compliers induced out of manufacturing due to the evolution of the two sectors between $1999$ and $2018$, by state} \begin{adjustbox}{max width=1\textwidth, center} \resizebox{1.3\textwidth}{!}{ \begin{tabular}{@{\extracolsep{5pt}} cccccccccc} \\[-1.8ex]\hline \hline \\[-1.8ex] & \multicolumn{4}{c}{semi-IVs: evolution from $1999$ to $2018$} & \multicolumn{2}{c}{$P(Z)$ change} & \multicolumn{3}{c}{$Y_0 - Y_1 | V \in [p^s_{1999}, p^s_{2018}]$} \\ \cline{2-5} \cline{6-7} \cline{8-10} \\[-1.8ex] State & log($w_{0, 1999}^s)$ & $\Delta w_0^s$ & log($w_{1, 1999}^s)$ & $\Delta w_1^s$ & $p_{1999}^s$ & $p_{2018}^s$ & \color{mtr0}{LATR$_0$} & \color{mtr1}{LATR$_1$} & \color{mte}{LATE} \\ \\[-1.8ex] \hline \\[-1.8ex] \\ AR & $10.829$ & $0.361$ & $9.532$ & $$-$0.152$ & $0.309$ & $0.249$ & $2.657$ & $2.455$ & $0.202$ \\ PA & $12.640$ & $0.385$ & $11.200$ & $$-$0.213$ & $0.302$ & $0.237$ & $2.749$ & $2.530$ & $0.219$ \\ MN & $11.878$ & $0.376$ & $10.223$ & $0.127$ & $0.286$ & $0.240$ & $2.760$ & $2.559$ & $0.201$ \\ IL & $12.826$ & $0.294$ & $11.216$ & $$-$0.103$ & $0.269$ & $0.224$ & $2.703$ & $2.557$ & $0.146$ \\ SC & $11.318$ & $0.460$ & $10.073$ & $$-$0.051$ & $0.265$ & $0.204$ & $2.655$ & $2.600$ & $0.054$ \\ GA & $12.338$ & $0.445$ & $10.663$ & $$-$0.058$ & $0.239$ & $0.182$ & $2.629$ & $2.582$ & $0.048$ \\ NC & $12.141$ & $0.501$ & $11.078$ & $$-$0.038$ & $0.232$ & $0.172$ & $2.542$ & $2.543$ & $$-$0.001$ \\ WV & $10.461$ & $0.275$ & $8.734$ & $$-$0.193$ & $0.203$ & $0.163$ & $2.473$ & $2.576$ & $$-$0.103$ \\ NY & $13.456$ & $0.415$ & $11.056$ & $$-$0.294$ & $0.182$ & $0.129$ & $2.428$ & $2.662$ & $$-$0.234$ \\ TX & $13.224$ & $0.590$ & $11.602$ & $0.212$ & $0.175$ & $0.127$ & $2.410$ & $2.720$ & $$-$0.310$ \\ MA & $12.281$ & $0.421$ & $10.532$ & $$-$0.171$ & $0.171$ & $0.124$ & $2.489$ & $2.798$ & $$-$0.309$ \\ CA & $13.818$ & $0.506$ & $12.025$ & $0.182$ & $0.163$ & $0.123$ & $2.374$ & $2.701$ & $$-$0.328$ \\ NJ & $12.543$ & $0.277$ & $10.717$ & $$-$0.351$ & $0.131$ & $0.097$ & $2.269$ & $2.823$ & $$-$0.554$ \\ MD & $11.976$ & $0.483$ & $9.565$ & $0.003$ & $0.128$ & $0.090$ & $2.311$ & $2.894$ & $$-$0.583$ \\ \\ \hline \\[-1.8ex] \end{tabular} } \end{adjustbox} \captionsetup{font=footnotesize, justification=justified} \caption*{Notes: States sorted by decreasing baseline $p_{1999}^s$. The $\Delta w_d^s$ columns represents the change of log$(w_{d, 2018}^s) - $log$(w_{d, 1999}^s)$, i.e., the growth of (lagged) GDP of sector $d$ between $1999$ and $2018$. The propensity score columns present the estimated propensity score for $25$ years old workers in state $s$ with semi-IVs $w_{d, year}^s$ in both $1999$ and $2018$. These are estimated using the first stage probit (Table (ref)). The last three columns represent the (reversed) treatment effects $Y_0 - Y_1$ on individuals induced into $D=0$ because of the evolution of the semi-IVs (and of time). The LATE is computed according to Equation (ref). } \end{table} LATE on the `compliers' from the nonmanufacturing growth. Over the period, the relative growth of the nonmanufacturing sector induced a flow of `compliers' to switch \textit{out} of the manufacturing sector. The question is to know whether this switch was beneficial or not for these compliers. \\ For each state, indexed by $s$, we compute the effect of the observed evolution of the each sector $d$, from $w^s_{d, 1999}$ to $w^s_{d, 2018}$ between $1999$ and $2018$, for individuals aged $25$. First, this evolution induced a change of the propensity to work in manufacturing sector from: $P(w^s_{0, 1999}, w^s_{1, 1999}, \text{Year}=1999, \text{State}=s, \text{Age}=25) = p^s_{1999}$ to $P(w^s_{0, 2018}, w^s_{1, 2018}, \text{Year}=2018, \text{State}=s, \text{Age}=25) = p^s_{2018}$. Typically, in every state, the nonmanufacturing sector grew significantly more than the manufacturing sector. Following the first stage estimation (Table (ref)), it leads to $p^s_{1999} > p^s_{1999}$ for all states (in line with the observed employment decline in Figure (ref)).\footnote{The estimated $2018$ year fixed effect in the first stage is negligible, equal to $-0.014$ and not significant. Therefore, most of the evolution of the propensity score is driven by the relative change of the semi-IVs here. } Thus, the `compliers' are individuals who are induced \textit{out} of the manufacturing sector, and for them, the treatment effect is now defined as $Y_0 - Y_1$ instead of $Y_1 - Y_0$. Let us compute the LATE as: \begin{align} &\text{LATE}(p^s_{1999}, p^s_{2018}, w_{0, 2018}^s, w_{1, 2018}^s, x_{2018}^s) \nonumber \\ &= \mathbb{E}[Y_0 - Y_1 | p^s_{1999} \leq V < p^s_{2018}, W_0=w^s_{0, 2018}, W_1 =w^s_{1, 2018}, X=x_{2018}^s], \end{align} where $x_{2018}^s = (\text{Year}=2018, \text{State}=s, \text{Age}=25)$. The values of the semi-IVs in $1999$ only pins down the value of the baseline propensity score, $p^s_{1999}$. \\ These state by state LATE are reported in Table (ref). We observe a large heterogeneity, ranging from $+20.2\%$ in Arizona to $-58.3\%$ in Maryland. Most of the effect is driven by the initial proportion of manufacturing workers. Looking at Figure (ref), one can see that the effect will depend on who (in terms of $V$) is induced into $D=0$. At high $V$, the switch from $D=1$ to $D=0$ (i.e., minus the MTE) may be positive while for small $V$ it is negative. \\ Overall, a relative increase in $W_0$ has a conflicting effect on the sector earnings: it raises the earnings of all nonmanufacturing workers, but also increases the probability of workers selecting into nonmanufacturing. This shift attracts workers with lower 'nonmanufacturing skills' into the nonmanufacturing sector, ultimately lowering its average earnings. For high $V$ though, these newly shifted workers were even worse off in the manufacturing sector, so they benefit from the shift into nonmanufacturing.
figure[figure omitted — 2,131 chars of source]

Evolution of the returns to working in manufacturing over time

Figure (ref) shows the evolution of returns to working in manufacturing over time, relative to the $1999$ baseline. This evolution accounts for time fixed effects, as well as the changes in the semi-IVs over time and their resulting effects on earnings. As shown in Figure (ref), the returns to working in manufacturing declined until $2009$, down to $-20\%$, and have slightly rebounded since then, although they remain about $10\%$ lower than in $1999$.\footnote{I interpret these estimates as changes in the MTE, though under the semi-parametric specification, they also correspond to changes in the global average treatment effects. } I also report the evolution of the manufacturing employment probability in Figure (ref) and observe that it closely follows the evolution of the MTE. This suggests selection on earnings gains: workers primarily choose to enter manufacturing or not based on their expected earnings in each sector. The manufacturing employment share remains lower than in $1999$, largely because returns to working in manufacturing have not fully recovered. This pattern was undetectable in the naive (OLS) estimation of the returns to working manufacturing (Figure (ref)), which misleadingly suggests higher returns in $2018$ than in $1999$. Using the semi-IV approach to correct for sector choice endogeneity provides a clearer picture and helps explain the persistent decline in manufacturing employment. \\ Perhaps surprisingly, the evolution of both MTRs (Figure (ref)) reveals a harsher reality for young, uneducated white men: while their manufacturing earnings ($Y_1$) have declined, their nonmanufacturing earnings ($Y_0$) have also fallen, albeit less sharply, since 1999. This highlights a marked deterioration in the economic prospects for this population, regardless of their sector choice. \\

A practical remark. As highlighted by this last result, the MTRs are often informative on their own. In general, they provide valuable information about the selection process, about the dependence between potential outcomes and resistance to treatment. In this application, we observe a standard scenario: $m_1$ is high when $V$ is low and decreases with $V$, while the reverse holds for $m_0$. However, many other scenarios could yield the same observed MTE. For instance, the MTE could be entirely driven by $m_1$ while $m_0$ would remain flat. Reporting only the MTE thus results in a loss of information. Since MTRs can also be estimated with IVs at no additional cost, I recommend always reporting them alongside the MTE.

Conclusion

This paper shows that semi-IVs provide a viable alternative to IVs for causal inference, particularly for estimating LATE and MTE. This novel approach is especially useful since semi-IVs are often more readily available, as illustrated by the examples and the returns to working in manufacturing application. Expanding the set of valid identifying variations broadens the empirical researchers' toolkit. This expansion should enable researchers to address important economic questions for which finding a credible IV was previously infeasible.

{

spacing{1} {4pt} {0pt} \scriptsize

}