Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
95,885 characters · 20 sections · 89 citation commands
Decomposing Identification Gains and Evaluating Instrument Identification Power for Partially Identified Average Treatment Effects
\affil[1]{Amsterdam School of Economics, University of Amsterdam, The Netherlands} \affil[2]{Department of Econometrics and Business Statistics, Monash University, Australia}
The average treatment effect (ATE) is an important policy relevant measure in causal analysis \citep*{heckman2006understanding,imbens2004nonparametric}, but its identification and estimation in empirical research has long been contentious when the treatment is endogenous and instrumental variables (IVs) are used as the identification strategy. This paper takes an empirical causal analyst’s perspective and examines and illustrates the roles of IVs and other factors in the identification and estimation of the ATE within a partially identified modelling framework. Although there have been significant theoretical developments in the econometric literature on the understanding of conventional IV estimands and treatment effect bounds under broader assumptions \citep*[see e.g.][]{manski1990nonparametric,balke1997bounds,heckman1999local,heckman2001instrumental,heckman2005structural,manski2000monotone,chernozhukov2007estimation,chesher2010instrumental}, the exact role of IVs and the associated estimation of various causal effects have remained not well understood in applied economic studies. By synthesising the existing econometric literature on IVs and ATE bounds, and with the help of some novel analyses, in this paper we aim to facilitate a better understanding of the complex role of IVs in ATE identification in practical applications.
In empirical causal analyses, it is common for researchers to estimate the causal effect of an endogenous treatment by a conventional IV estimator as the “identification strategy” nunn2014us. In models with homogeneous treatment responses, using any one of the valid IVs can lead to point identification of the ATE, and conventional IV estimators correctly estimate the ATE. However, as shown by \citet*{heckman2006understanding}, in heterogeneous treatment effect models, different IVs identify different local treatment effects imbens1994identification, and conventional IV estimates are no longer robust to the choice of alternative IVs. In fact, as explained by \citet*{heckman2006understanding}, the classical IV estimand is not ATE but may be a quantity with no easily interpretable meaning regardless of IV strength.
Evidence against homogeneous treatment effects abounds, with estimates based on different sets of valid IVs often producing different treatment effect estimates in practice carneiro2003understanding,basu2007use,angrist2010extrapolate. Once heterogeneous treatment response is allowed, the ATE is often not point identified. Thus, the analysis in \citet*{heckman2006understanding} presents a convincing argument that if the ATE is of primary interest, the identified set for the ATE based on a partially identified model should be preferred to an analysis based on a conventional IV estimation approach. However, there have only been limited applications of partially identified ATE analysis in empirical studies. Consequently, understanding the role of IVs in this setting is important for promoting heterogeneous treatment models for empirical causal analysis.
In heterogeneous treatment effects models, heckman2001instrumental demonstrate that the property of “identification at infinity” heckman1990varieties, namely, the availability of IVs that produce propensity scores of zero and one in the limit, leads to point identification of the ATE. However, this condition is rarely satisfied in practice, especially when IVs have limited variation. When IAI fails, inference on the ATE can still be carried out by constructing an identified set for the ATE. Therefore, in a partial identification framework, the impact of IVs can be studied by their influence on the ATE bounds.
We focus on models with binary outcome and binary endogenous treatment. Such models have been widely used in empirical studies since the pioneering work of heckman1978dummy. See neal1997effects, thornton2008demand and ashraf2014household for examples that use fully parametric bivariate probit models, and see aakvik2005estimating, bhattacharya2008treatment,bhattacharya2012treatment and kreider2012identifying,kreider2016identifying for examples that rely on nonparametric models. The role played by IVs has been a topic of discussion in the literatures, including the notion of “identification by functional form” \citep*[see e.g.][]{maddala1986limited, wilde2000identification, freedman2010endogeneity, mourifie2014note, han2017identification, li2019bivariate}.\footnote{See li2019bivariate for a summary on the topic, including a sufficient condition regarding the support of exogenous regressors for models such as the bivariate probit to achieve ATE point identification without any IVs. However, once the restrictive parametric assumptions fail to hold, IVs become necessary.}
The important role of IVs has been noted for partially identified ATE in heterogeneous treatment effect models manski1990nonparametric, heckman2001instrumental,chesher2005nonparametric,chesher2010instrumental, shaikh2011partial,li2019bivariate. heckman2001instrumental show that for their model with threshold crossing for the treatment, it is the width between the minimum and maximum propensity scores reached by the available instruments that determines the ATE bound width. chesher2010instrumental also points out that the support and the strength of the IVs are important in determining the ATE bounds, whilst li2018bounds present some simulation results on bound width and IV strength. However, the mechanism through which the IV strength translates to identification gains in partially identified models and whether/how other factors also play a part have not been laid bare in a manner that can be readily understood by practitioners.
In this paper, we examine the role of IVs, as well as their interplay with the degree of endogeneity and exogenous covariates, in the identification of the ATE. Following the partial identification literature,\footnote{For example, see Kitagawa (2009) and Swanson et al (2018) among others.} we use the ATE sign identification and the reduction in the size of the ATE identified set as a measure for identification gains. Focusing on the bivariate joint threshold crossing model and the ATE bounds proposed by shaikh2011partial (henceforth referred to as the SV model and SV bounds), we disentangle the various factors determining the ATE identification, and provide useful insights for the practitioners into the different sources and natures of identification gains.
To this end, our first contribution is to highlight and demonstrate for the case of SV bounds how IVs achieve identification gains via their attained minimum and maximum conditional propensity scores. The implication to empirical researchers is that, unlike in homogeneous treatment effect models, in heterogeneous models, omitting relevant IVs, or misclassifying continuous IVs as binary ones, could result in a loss of identification power and wider ATE bounds.
Second, we show that, unlike the case of heckman2001instrumental, the SV bounds for a binary outcome are additionally impacted by the sign and degree of treatment endogeneity. Interestingly, we find that the endogeneity drives the SV bounds asymmetrically. Specifically, the same propensity score extremes could offer much greater identification power when the ATE and the endogeneity direction are of the opposite signs, relative to the case when they are of the same sign. Thus, it is the interactions of IVs with other features of the model that determine the level of the IV identification power for the ATE. Similar asymmetric influence of treatment endogeneity is also noted in nonlinear parametric models freedman2010endogeneity,frazier.
Our third contribution is to propose a novel decomposition of identification gains into components driven by: (i) the existence of valid IVs that identifies the sign of the ATE; (ii) IV strength that determines the size of the outer set of the ATE identified set; and (iii) the variation of the exogenous covariates that further refines the outer set. The last component is the key driver for achieving the ATE sharp identified set in mourifie2015sharp and the ATE point identification in vytlacil2007dummy. Based on the decomposition, we further propose a measure for IV identification power (hereafter $IIP$), which captures the critical fact that the IV identification information pertaining to the ATE varies with the endogeneity degree.
The decomposition and $IIP$ analysis allow us to shine a light on the internal workings of the ATE partial identification mechanism and thereby characterize the structure of identification gains. This analysis allows us to offer useful insights on the extent of identification gains achievable by each individual factor, and the contribution of each IV when multiple IVs are used. Our analysis includes graphical illustrations of bound reduction anatomy, as well as Monte Carlo results of finite sample performance. We also apply our methods to an empirical study of women’s labour force participation angrist1998children. Two IVs are used in the study: a dummy for the first two children being same-sex siblings (“Samesex”), and a dummy for the second birth being a twin (“Twins”). Our analysis shows that the identification power of Twins is about 1.4 times the identification power of Samesex if they are used separately. If both IVs are used, the total IV identification power is only marginally larger than that if only using the Twins.
Together with the theoretical decomposition analysis, we believe this paper offers useful insights for empirical causal researchers who wish to understand the complex impacts of IVs on ATE partial identification. Furthermore, we offer some practical examples where our analysis can be used for policy relevant instrument design and selection. Our paper also sheds light on instrument relevancy. Our $IIP$ measure is related to existing approaches in the generalized methods of moment (GMM) literature that seek to determine instrument “relevancy”. The ability of our approach to rank sets of IVs by their identification gains, in conjunction with our Monte Carlo simulation results lead us to document, we believe for the first time, an important feature of bivariate triangular models: while in the population, adding irrelevant IVs cannot tighten the ATE bounds, in finite-samples, using such IVs could lead to a loss in IV identification power and wider bounds, when the variation of the covariates is small. We liken this phenomena to the well-known problem of irrelevant moment conditions in GMM breusch1999redundancy,hall2003consistent,hall2005generalized,hall2007information and leave a more rigorous study of this topic for future research.
The rest of this paper is organized as follows. In Section (ref) we present the SV model setup and the SV bounds. In Section (ref) we establish three key factors that affect the ATE bounds. Section (ref) introduces our decomposition of identification gains and the index of $IIP$. A comprehensive numerical analysis and graphical presentation are given in Section (ref). Finite sample evaluation and implications for empirical causal practice are presented in Section (ref), and an empirical example is given in Section (ref). The paper closes in Section (ref) with some summary remarks. All proofs are relegated to Appendix.
Suppose we observe a binary outcome $Y$ and a binary treatment $D$. Consider the joint threshold crossing (JTC) model studied in shaikh2011partial:
where $X$ denotes a vector of exogenous covariates, $Z$ represents a vector of instruments that can be discrete, continuous or mixed, $\nu_1$ and $\nu_2$ are unknown functions, and $\varepsilon_1$ and $\varepsilon_2$ are unobservable error terms. Let $Y_d=1[\nu_1(d,X)>\varepsilon_1]$ denote the potential outcome for $D=d$ with $d=0,1$. The JTC model allows for flexible forms of heterogeneous treatment effects due to the nonseparable error structure heckman2006understanding, and is often used in treatment evaluation studies bhattacharya2008treatment,bhattacharya2012treatment,kreider2012identifying.\footnote{vytlacil2002independence shows that the threshold crossing condition is equivalent to the monotonicity assumption. In the JTC models, it means that all individuals with the same observable characteristics will respond to the treatment and instrument in the same direction. bhattacharya2012treatment demonstrate that the ATE SV bounds under the JTC model (ref) still hold under a rank similarity condition, a weaker property that allows heterogeneity in the sign of the ATE$(x)$.}
We are interested in the most commonly studied treatment effect, the conditional ATE, defined as $$\text{ATE}(x)=\mathbb{E}[Y_1|X=x]-\mathbb{E}[Y_0|X=x].$$ For notational simplicity, for any generic random variables $A$ and $B$, henceforth we will use $\text{Pr}[A|b]$ to represent $\text{Pr}[A|B=b]$, unless otherwise stated. The support of $A$ is denoted as $\Omega_A$ and the support of $A$ conditional on $B=b$ is given by $\Omega_{A|b}$.
Assumption (ref) (a) and (e) ensure that the instrument $Z$ is independent of the error terms and relevant to the treatment. Conditions (b), (c) and (d) are imposed for analytical simplicity. Denote $P=P(X,Z)=\text{Pr}[D=1|X,Z]$ with support $\Omega_P$. In model (ref), $Z$ affects the outcome $Y$ only through the propensity score $P$, which is called index sufficiency.
Proposition (ref) summarizes the key results in shaikh2011partial and is presented for reference. Denote $\underline{p}:=\inf\{p\in\Omega_{P}\}$ and $\overline{p}:=\sup\{p\in\Omega_{P}\}$.
There are several implications of Proposition (ref). First, the SV bounds in (ref) and (ref) consist of two layers of intersection evaluations. The first layer is to intersect all possible values of the conditional propensity score $P$ given $X=x$, or equivalently, of the IVs. The second layer of intersections are taken over values of covariates whose variation can compensate (to some degree) for the variation in the treatment variable in the outcome equation. Thus, both the IVs and the covariates contribute to the ATE partial identification using the SV bounds.
In particular, the SV bounds are informative in the sense that the sign of the ATE$(x)$ is recovered, if $Z$ has nonzero prediction power for the treatment, meaning that there exist at least two different values of $P(x,z),P(x,z')\in\Omega_{P|x}$. It indicates that the IVs' contribution to the SV bounds is achieved in two steps: first, using any nonzero IV variation to identify the ATE sign; and second, using the first layer intersections to further shrink the SV bounds. In addition, when the instrument is binary, the ATE$(x)$ sign is identified by the sign of the “intention-to-treat” parameter $\text{Pr}[Y=1|x,P(x,z)]-\text{Pr}[Y=1|x,P(x,z')]$.\footnote{machado2013instrumental study how the intention-to-treat parameter can be used for identification and inference of the ATE sign under different assumptions, including the JTC model. ura2018heterogeneous and tommasi2020bounding also find the intention-to-treat parameter identifies the ATE sign in heterogeneous treatment effect models.} For the contribution of covariates, vytlacil2007dummy show that it is possible to achieve the ATE point identification via the SV bounds if $X$ contains a continuous element or the exclusion restriction holds in both equations. shaikh2011partial Remark 2.1 also discussed conditions for point identification.
Third, by imposing the support condition $\Omega_{X,P}=\Omega_X\times\Omega_P$, the SV bounds are sharp and can be simplified to be functions of the two extreme values of the propensity score $P$.\footnote{The condition $\Omega_{X,P}=\Omega_X\times\Omega_P$ is saying that for any $x,x'\in\Omega_X$, there exist possible realizations $z,z'$ of $Z$ such that $\text{Pr}[D=1|x,z]=\text{Pr}[D=1|x',z']$. The same support condition is also used in mourifie2014partially. If $X$ and $Z$ are both binary, then the point identification assumption in vytlacil2007dummy is equivalent to $\Omega_{X,P}=\Omega_X\times\Omega_P$.} Using the expressions of the sharp SV bounds, we can confirm that the IAI condition guarantees the ATE point identification in the JTC model: if $\underline{p}=0,\overline{p}=1$ holds, then $L^{SV}(x)=U^{SV}(x)=$Pr$(Y=1|x,\overline{p}=1)-$Pr$(Y=1|x,\underline{p}=0)$. Note that the condition $\Omega_{X,P}=\Omega_X\times\Omega_P$ might fail to hold in practice, especially when the variation in $Z$ is limited. When it does not hold, mourifie2015sharp provides an ATE sharp identified set by further exploiting the information on other individuals with different covariate and propensity score values.\footnote{han2020sharp provide a linear programming method to compute the sharp bounds with binary scalar instrument. While, this paper considers more general settings where the instrument(s) can be discrete, continuous or mixed.}
Given the SV bounds, the width of the SV bounds can be defined as $$\omega^{SV}(x)=U^{SV}(x)-L^{SV}(x)\,.$$ Throughout the paper, we use the sign of the ATE and the reduction in the width of the identification set to measure identification gains and power.
We first study how three key factors drive the SV bounds: the conditional propensity score, the direction and degree of endogeneity, and the variation of covariates.
As discussed in the introduction, the propensity score is the vehicle that carries the identification information in the IVs. We start from the conditional propensity score (CPS), $P(x,Z)=$Pr$[D=1|X=x,Z]$, and examine its features that determine the ATE bounds. Define the two extremes of the CPS as $$\underline{p}(x):=\inf_{z\in\Omega_{Z|x}}\{p\in\Omega_{P|x,z}\}\;\text{ and }\;\overline{p}(x):=\sup_{z\in\Omega_{Z|x}}\{p\in\Omega_{P|x,z}\}.$$
In this subsection, we assume the support condition $\Omega_{X,P}=\Omega_X\times\Omega_P$ holds. shaikh2011partial show that the SV bounds are sharp and are functions of $\underline{p}$ and $\overline{p}$. In addition, the support of the CPS $P(x,Z)$ reduces to $\Omega_P$ for $\forall x\in\Omega_X$. Therefore, the two extreme values of the CPS become to $\underline{p}(x)=\underline{p}$ and $\overline{p}(x)=\overline{p}$, where $\underline{p}=\inf\{p\in\Omega_{P}\}$ and $\overline{p}=\sup\{p\in\Omega_{P}\}$. The proposition below shows that the extreme values of the CPS, i.e. $\underline{p}$ and $\overline{p}$, are crucial determinants of the lower and upper bounds and the width of the ATE sharp identified set.
Complementing the results in shaikh2011partial, Proposition (ref) fills in the blanks by explicitly showing that when the IAI fails to hold, the contribution of IVs to the ATE bounds are delivered by the magnitude of $\underline{p}$ and $\overline{p}$ (while holding anything else fixed). In other words, the two extreme values of the CPS define the IV strength in the JTC model. More importantly, we find a monotone relationship between $\underline{p}$ and $\overline{p}$ and the ATE bounds. We emphasize that this monotone result is not found in heckman1999local,heckman2001instrumental, chesher2010instrumental and shaikh2011partial.\footnote{heckman1999local,heckman2001instrumental show that in the treatment threshold crossing model, the ATE bound width is linearly related to the two extreme values of CPS, while they do not give the similar result for the lower and upper ATE bounds. chesher2010instrumental briefly discusses the importance of IV strength and support in determining the extent of ATE identified set for models with threshold crossing in outcome. shaikh2011partial only show that the bounds are functions of $\underline{p}$ and $\overline{p}$ in the JTC model.}
The implications of Proposition (ref) are significant. For homogeneous treatment response models such as linear regression models, any valid and relevant IVs would lead to the same identification gains because the ATE is point identified (e.g. by Wald estimand or 2SLS), therefore missing IVs or misclassifying continuous IVs into binary ones play no role in identifying the ATE heckman2006understanding. However, in heterogeneous treatment effect models, the fact that the CPS affects the ATE bounds via the distance of its extreme values to zero and one demonstrates that missing and misspecified IVs will result in narrower span of CPS, which may lead to nontrivial loss in identification power.
The support condition $\Omega_{X,P}=\Omega_X\times\Omega_P$ may fail when the instruments have limited variation. mourifie2015sharp establishes the ATE sharp identified set via exploiting the variation of covariates without using the support condition. We are able to establish and verify the relationship between the CPS and a tractable characterization of ATE bounds, which is an outer set of mourifie2015sharp's sharp identified set and also an outer set of the SV bounds.\footnote{The analysis using the outer set is practically useful because empirical studies often construct confidence region based on an outer set chesher2020econometric. Since an outer set always includes the sharp identified set, the analysis is conservative yet valid as long as the model is correctly specified molinari2020microeconometrics,kedagni2020discordant. Moreover, discussions in Section (ref) imply that further improvement in the outer set towards the sharp identified set of mourifie2015sharp can be attributed to the information of covariates.} The explicit expressions of the outer set, denoted by $[\underline{L}^{SV}(x),\overline{U}^{SV}(x)]$ can be found in (ref) and (ref); see the proof of Proposition (ref). The width of the outer set, denoted by $\overline{\omega}(x)=\overline{U}^{SV}(x)-\underline{L}^{SV}(x)$, is given in the proposition below.
We can see that the width of the outer set is monotone in the extreme values of CPS, i.e. $\underline{p}(x)$ and $\overline{p}(x)$. Thus, it is no doubt that the extreme values of CPS are also crucial determinants of the ATE sharp bounds, even though the relation may not be monotone. In addition, Proposition (ref) also confirms that the IAI leads to $\overline{\omega}(x)=0$ and ATE point identification.
Next, we illustrate that, for given IV strength (i.e., $\overline{p}(x)$ and $\underline{p}(x)$), how the treatment endogeneity can enhance or hinder the IV identification power and SV bounds. We introduce a family of bivariate single parameter copulae that specifies the joint distribution of $(\varepsilon_1,\varepsilon_2)$, while we do not require the copula nor the marginal distributions to be known. Denote a copula as $C(\cdot,\cdot;\rho):(0,1)^2\mapsto(0,1)$, where $\rho\in\Omega_\rho$ is a scalar dependence parameter that fully describes the joint dependence between $\varepsilon_1$ and $\varepsilon_2$, and their dependence increases as $\rho$ increases.\footnote{In the special case of a normal bivariate probit model $\rho\in[-1,1]$ represents the correlation between the error terms.} For any given copula, the sign and magnitude of $\rho$ can be understood as the direction and degree of endogeneity.
We also impose additional dependence structure, the concordance ordering, on the copula $C(\cdot,\cdot;\rho)$. Following joe1997multivariate, for $\rho_1\neq\rho_2$ and $u_1,u_2\in(0,1)^2$, we say that the copula $C(\cdot,\cdot;\rho)$ satisfies the concordant ordering with respect to $\rho$, denoted as $C(u_1,u_2;\rho_1)\prec_cC(u_1,u_2;\rho_2)$, if
The concordant ordering is a stochastic dominance restriction and is embodied in many well-known copulae, including the normal copula. Similar stochastic dominance conditions are employed in, e.g., han2017identification and han2019estimation, to derive identification and estimation results for the parametric bivariate probit model and its generalizations. Denote the cumulative distribution function of any random variable $A$ as $F_{A}$.
Assumption (ref) does not require $C(\cdot,\cdot;\rho)$ nor $F_{\varepsilon_1}$ and $F_{\varepsilon_2}$ to be known, and is used to establish the impacts of the dependence parameter $\rho$ on the SV bounds.\footnote{The same general conclusions are expected to follow in cases where the parametric copula is relaxed to nonparametric copulae with similar ordering conditions.}
Proposition (ref) implies that the outer set of the SV bounds is impacted by the direction and the degree of endogeneity. Intuitively, this is because the outer set is constructed using the joint probabilities of the outcome and the treatment. In particular, the effect of endogeneity's direction is asymmetric: given a positive ATE, negative dependence between the two error terms helps narrow down the ATE bound width, while the opposite holds for a negative ATE. Thus, even for IVs that have the same $\underline{p}(x)$ and $\overline{p}(x)$, the information contained in the IVs can be correspondingly scaled via the leverage induced by the direction and degree of endogeneity.
In this section, we summarize some key results of covariates in the ATE identification literature. As we have seen from the construction of the SV bounds, covariates contribute to partially identifying the ATE. It is therefore a useful insight for empirical researchers that having a rich set of covariates can help shrink the ATE bounds, for given IV strength. Here, the impact of covariates can be considered as an extension of the popular propensity score matching method in conventional empirical studies. In fact, even without the IAI condition, in certain situations, the covariate variability can ensure the point identification of the ATE. In other cases, where the covariate variability is minimal, no further tightening can be achieved beyond the outer set. Define
Proposition (ref) is a special case of Theorem 4.1 of vytlacil2007dummy when the outcome is a binary variable, and it is also discussed by shaikh2011partial Remark 2.2. It states that when the variation in $X$ exactly compensates for the variation in $D$, which is satisfied by those values $x\in \mathcal{X}^0\cap \mathcal{X}^1$, the point identification of the ATE can be obtained. Intuitively, the identification is achieved by finding a shift in covariates that offsets a shift in the treatment to make the outcome and the probability of getting treated unchanged. mourifie2015sharp also uses similar idea to tighten the SV bounds.
Proposition (ref) provides two conditions of covariate, under which the SV bounds fail to utilize any information in $X$ to compensate the impact of the variation in $D$ on the outcome, and therefore the SV bounds coincide with its outer set. Particularly, condition (i) summarizes the results described in Section 3.1.2 in chiburis2010semiparametric, when none of other propensity score values match exactly to $P(x,z)$. The condition in (ii) holds, if either $X$ is a void variable, or if $X$ has no impacts on outcome.
The results above and the construction of SV bounds both indicate that for any given IVs strength, if the variation of covariates can be enriched, then SV bounds shrink. Thus, for the given IV strength, any further improvement of the outer set towards SV bounds (or, towards mourifie2015sharp's bounds) can be attributed to additional identification information in covariates, whose variations compensate for the treatment's variation to some degree.
In this section, we present a novel decomposition of the identification gains of the SV bounds into components attributable to IVs and exogenous covariates. In addition, we introduce the IV identification power index, $IIP$.
We first introduce the benchmark ATE bounds of manski1990nonparametric which use no IVs and are often referred to as “the worst case scenario” tamer2010partial,chiburis2010semiparametric,bhattacharya2012treatment. The Manski bounds are used as a benchmark, because if IVs are all irrelevant, the ATE SV bounds collapse to the ATE Manski bounds.\footnote{See Remark 2.1 of shaikh2011partial and chiburis2010semiparametric Corollary 1. In particular, Corollary 1 of chiburis2010semiparametric shows that if no IVs are relevant, although the outcome threshold crossing condition tightens the identified set of $\nu_1(0,x)$ and $\nu_1(1,x)$ relative to those under Manski's framework, it fails to improve the ATE bounds.} Denote $L^M(X)$ and $U^M(x)$ as the lower and upper Manski bound of ATE$(x)$, respectively,
It is apparent that the width of the Manski bounds, defined as $\omega^M(x)=U^M(x)-L^M(x)$, are one, for any $x\in\Omega_X$, with the lower bound and upper bound falling on either side of zero. Our decomposition of identification gains is inspired by the results in Section (ref). We start with the Manski bounds, and divide the Manski bounds into four components, which represent incremental identification gains made by the SV bounds over the benchmark Manski bounds, due to different determinants.
It is easy to see that $C_1(x)+C_2(x)+C_3(x)+C_4(x)=\omega^M(x)=1$. If $\nu_2(X,Z)|X$ is degenerate and the IVs have no explanatory power for the treatment, then $C_1(x)=C_2(x)=C_3(x)=0$ and the SV bounds reduce to Manski bounds. In addition, $C_1(x)$ to $C_4(x)$ can always be identified and estimated from the data. In practice, once the model has been estimated (parametrically or non-parametrically), the estimates can be used to construct the decomposition.
Based on the decomposition, we can then construct a quantitative measurement of IV identification power in the partial identification setting. For $\forall x\in\Omega_X$, define the IV identification power $IIP(x)$ as
where $\overline{\omega}(x)$ is the width ATE outer set defined in Proposition (ref) and $IIP(x)\in[0,1]$. Setting $IIP(x)=0$ when $\nu_2(X,Z)|X=x$ is degenerate is equivalent to setting $\overline{\omega}(x)=\omega^M(x)=1$.\footnote{The definition allows $IIP(x)$ to be discontinuous at $\Omega_{P|x}=p_x$ for some constant $p_x\in[0,1]$, i.e. when $\Omega_{P|x}$ is a singleton.} $IIP(x)$ represents the identification gains that are due to the IVs alone and it can be viewed as an index of the IV identification power. The overall IV identification power can be obtained by $\mathbb{E}_X[IIP(X)]$. The following proposition formalizes some important properties of $IIP(x)$.
Proposition (ref) indicates that values of $IIP(x)$ can be compared, across different sets of IVs, or across different values of $x$ given the same set of IVs, since they are standardized relative to the same benchmark.\footnote{$IIP(x)$ or $\mathbb{E}_X[IIP(X)]$ can also be compared across various studies if necessary.} For example, $IIP(x)=0.4$ can be interpreted as that the Manski bounds can be reduced by 40% by using instruments alone.\footnote{Theoretically, the value of $IIP(x)$ should lie in $[0,1]$ and the width of Manski bounds is always one. Then $IIP(x)$ can be interpreted as the percentage points of the identification gains brought by the IVs. In finite sample settings where the estimated Manski bound width may no longer be exact one, then the sample explanation can be obtained by computing the ratio $\hat{IIP}(x)/\hat\omega^M(x)$ using their estimates.} In addition, the values of $IIP(x)$ at its end points are intuitively interpretable: $IIP(x)=0$ indicates that IVs are completely irrelevant and SV bounds reduce to Manski bounds; and $IIP(x)=1$ when the IVs are able to perfectly predict the treatment status and ATE$(x)$ is point identified.
$IIP(x)$ is a meaningful measure of IV usefulness for improving the ATE partial identification, because it incorporates the impacts of the direction and degree of endogeneity on the ATE bounds. Using $IIP(x)$ instead of the CPS extreme values emphasizes the fact that the latter one cannot provide a full picture of the IV identification power. For example, suppose a set of IVs has a small CPS support and the IVs seem to be “weak”. However, a large $IIP(x)$ can be achieved if the magnitude of treatment endogeneity is large and it has an opposite sign with that of the ATE. See for example, the numerical examples in Figure (ref) of Section (ref), when ATE$(x)>0$, $\rho=-0.8$ and $\lambda$ close to zero,\footnote{In the numerical example, $\lambda$ determines the CPS extremes. Closer to zero $\lambda$ means narrower span of the CPS support.} we have $IIP(x)>$50%. Thus, the identification power of the IV set can actually be very strong. Conversely, if the treatment endogeneity is of the same sign as that of the ATE, a reasonable CPS range may produce barely satisfactory $IIP(x)$ and identification gains. See for example, the numerical examples when ATE$(x)>0$, $\rho=0.8$ and $\lambda$ close to 0.5, we have $IIP(x)<$50%.
In this section we illustrate numerically and graphically the results on how each determinant affects the SV bounds and the decomposition of identification gains.\footnote{To make the paper manageable, we focus on the SV bounds. Our numerical analysis may be of broader interest as a prototype for evaluation of IV identification power that could be conducted in other frameworks under weaker assumptions, such as those in chesher2010instrumental, heckman2001instrumental, manski2000monotone and mourifie2015sharp.} Consider a data generating process (DGP) with a linear additive latent structure, which is similar to that studied in li2019bivariate:
where $X\perp Z$, $X\sim\mathbb{N}(0,1)$ and $Z\in\{-1,1\}$ with $\text{Pr}(Z=1)=1/2$. In addition, $(X,Z)'\perp(\varepsilon_1,\varepsilon_2)$ where $(\varepsilon_1,\varepsilon_2)$ is zero mean bivariate normal with unit variances and correlation $\rho$. We set $\alpha=1$ and $\pi=0$ across all parameter settings. Given this specification, there is a monotonic one-to-one mapping from the coefficient of the IV, $\gamma$, to the extreme values of the CPS.
We study the SV bounds under different parameter settings displayed in Table (ref). We capture the extreme values of the CPS by $\gamma$, the direction and degree of endogeneity by $\rho$, and the impact of exogenous covariate by $\beta$. We compute the SV bounds and the Manski bounds, and implement the identification gains decomposition under the true DGP. In what follows we present the outcomes at $x=\mathbb{E}[X]$.\footnote{Our numerical experiments have been conducted at various quantile points of $X$, but space considerations prevent us from listing all the results. We present the outcomes at $x=E[X]$ as these are representative.}
In Figure (ref), we plot in the first row the SV bounds for the ATE$(x)$, and in the second row the bound width. The SV bounds reduce to the Manski bounds when the IVs are irrelevant with $\gamma=0$ (the separate lines in the graphs at $\gamma=0$). When $\gamma$ moves away from zero, the SV bound width has a significant drop. In addition, since ATE$(x)$ is positive ($\alpha>0$), the SV bound width increases as $\rho$ increases. Moreover, comparison of the plots for different values of $\beta$ reveals that larger $\beta$ produces significantly narrower bound width. When $\beta=0.45$, point identification of the ATE$(x)$ is achieved for most of the $(\gamma,\rho)$ pairs because the condition in Proposition (ref) is satisfied.
Figure (ref) displays the decompositions of identification gains for $\gamma\in\{1,2\}$, $\rho\in\{-0.8,-0.5,0.5,0.8\}$ and $\beta\in\{0.05,0.25,0.45\}$. For the positive ATE$(x)$ case, the contribution of IV validity, $C_1(x)$, is determined by the Manski lower bound, and decreases as $\rho$ increases,\footnote{Conversely the numerical results not reported here show that when the ATE$(x)$ is negative $C_1(x)$ increases as $\rho$ increases.} while $C_1(x)$ is invariant to $\beta$. By way of contrast, the component $C_2(x)$ also does not change by $\beta$, but it increases significantly as the magnitude of $\gamma$ increases. The component of identification gains due to the exogenous covariates, $C_3(x)$, also contributes significantly to the bounds. When $\beta$ is relatively large (e.g. $\beta=0.45$), point identification is virtually achieved.
Figure (ref) depicts the index $IIP(x)$ as a function of $(\gamma,\rho)$. The plot confirms that, firstly, $IIP(x)$ is bigger when IVs are stronger ($|\gamma|$ higher). In addition, for a given IV strength, higher $IIP(x)$ can be achieved if $\rho$ has an opposite sign from the ATE$(x)$ and is of high magnitude ($|\rho|$). If $\rho$ is of the same sign as the ATE$(x)$, then the lower the degree of endogeneity the better the identification power.\footnote{Due to space limitation, numerical results not displayed here show that under the DGP (ref), our results regarding how the IV strength and its interaction with the direction and magnitude of endogeneity improve the benchmark Manski bounds also hold in other common ATE bounding analyses, including heckman2001instrumental and chesher2010instrumental.}
In this section, we first present finite samples results on the decomposition and $IIP(x)$. We also illustrate how $IIP(x)$ can be used to compare identification power of alternative IV sets, as there are empirical situations where selection among alternative IVs may be of interest. Second, we discuss practical implications of our results for applied researchers.
We first examine situations when irrelevant, incomplete, or mis-specified IVs are used. We design DGPs with the same set of IVs but with different directions and degrees of endogeneity and different covariant variations, so that the same IV strength can exert different identification powers. Consider i.i.d. samples drawn from DGP in (ref) with two IVs:
where $Z_1$, $Z_2$ and $X$ are mutually independent, and also independent of $(\varepsilon_1,\varepsilon_2)$, $Z_1\sim Bernoulli(1/2)$ and $Z_2\in\{-3,-2,-1,0,1,2,3\}$ with probabilities $(0.1,0.1,0.2,0.2,0.2,0.1,0.1)$. Set $\alpha=1$, $\beta=1$, $\pi=-1$, $(\gamma_1,\gamma_2)=(0.5,0.2)$, and $(\varepsilon_1,\varepsilon_2)$ is jointly normal with mean zero, variance one and correlation $\rho\in\{0.5,0.8\}$. Consider two cases of covariate variability: $X\sim\mathbb{N}(0,1)$ and $X\sim Bernoulli(1/2)$, as in Table (ref). We focus on ATE$(x)$ with $x=0$.
We introduce two "pseudo" IVs: $\widetilde{Z}_2=1[Z_2>0]$, a misspecified binary IV that only partially reflects $Z_2$, and an irrelevant IV $Z_3\in\{0,1\}$ such that $\text{Pr}[Z_3=1]=2/3$, and $Z_3\perp(\varepsilon_1,\varepsilon_2,Z_1,Z_2,X)$. To evaluate the finite sample performance of $IIP(x)$ and selection of IVs, consider five alternative sets of IVs: (1) only $Z_1$; (2) only $Z_2$; (3) $Z_1$ and misspecified $\widetilde{Z}_2$; (4) $Z_1$ and $Z_2$; and (5) $(Z_1,Z_2)$ and irrelevant $Z_3$.
Table (ref) presents the CPS extremes and $IIP(x)$ of the five IV sets under the true DGP, which are the same for case 1 and 2 because the covariate variability does not impact them. We can see that the $(\underline{p}(x),\overline{p}(x))$ is the widest in (4) when both $Z_1$ and $Z_2$ are used. Adding an irrelevant $Z_3$ does not change the $(\underline{p}(x),\overline{p}(x))$, so theoretically (5) has the same IV strength as (4). The $(\underline{p}(x),\overline{p}(x))$ shrinks when only one valid IV is used as in (1) or (2). As expected, when a valid IV is incorrectly specified as a proxy dummy $\widetilde{Z}_2$ in (3), the $(\underline{p}(x),\overline{p}(x))$ is narrower than that of the best set in (4), but wider than that in (1) with $Z_1$ alone. Interestingly, comparing IV set (3) with (2), set (2) with only one valid IV actually results in wider $(\underline{p}(x),\overline{p}(x))$ than that for the two IVs in set (3) with $Z_2$ misspecified.
Whilst the CPS range indicates the IV strength, it is the $IIP(x)$ that captures the identification power of each IV set, measuring the reduction of SV bound width relative to the benchmark Manski bound width due to the contribution of IVs. As seen from the two $IIP(x)$ columns in Table (ref), the same IV strength can achieve bigger identification gains for $\rho=0.5$ than that with $\rho=0.8$. This is consistent with the results in Section (ref): as $\rho$ and ATE$(x)$ are both positive in this case, the lower absolute value of $\rho$, the higher the $IIP(x)$ is. For example for IV set (4), the Manski bound width can be reduced by $59.4\%$ by the two IVs when $\rho=0.8$, and it increases to $62.5\%$ if $\rho=0.5$. For given $\rho$, the equally most powerful IV sets are (4) and (5), and the least powerful set is (1).
We next present the finite sample estimation of the Manski and SV bounds, and conduct the decomposition analysis. Sample size is $n=500,5000,10000$ and replicate $M=1000$ times. Tables (ref) to (ref) present the sample average (over $M$ replications) of the estimated bounds, $C_{1}(x)$ to $C_{4}(x)$ and $IIP(x)$ of the five IV sets at $x=0$. We use the “half-median-unbiased estimator” (HMUE) of the intersecting bounds proposed by \citet*{chernozhukov2013intersection} (hereafter CLR) to estimate the benchmark Manski bounds and the SV bounds. In particular, we employ maximum likelihood estimation (MLE) to estimate the bounding functions and to select the critical values for bias correction according to the simulation-based methodology of CLR.\footnote{We report the HMUE of the Manski bounds, for comparison purpose. Other estimation methods for Manski bounds are also available, see e.g. imbens2004confidence. See Appendix (ref) for more details about the CLR method.}
Tables (ref) and (ref) present results under two different covariate distributions. The first row in each table lists the ATE bounds and decomposition components under the true DGP. We can see that in case 1 (Table (ref)), where $X$ possesses sufficient variation, the true SV bounds point identify the ATE$(x)$ for both $\rho=0.5$ and $\rho=0.8$. In case 2 (Table (ref)), the true SV bounds fail to point identify the ATE$(x)$ due to the limited variation in $X$.
Next, we focus on the left part of each table, which displays the HMUEs of the ATE bounds, and the Hausdorff distance between the true and estimated bounds, at $x=0$.\footnote{Simulation results at different values $x$ display similar patterns to those at $x=0$, therefore are not reported due to the space limitation. Hausdorff distance has been employed to study convergence properties when a set is the parameter of interest, see e.g. manski2002inference.} For all four tables, we can see that the estimated Manski bounds are the same across all five IV sets, always include zero, and have a width a little over one. The estimated SV bounds identify the sign of ATE$(x)$ for all five IV sets. Moreover, the IV sets with greater identification power lead to narrower estimated SV bounds and also improve the estimation accuracy in most of the scenarios. More precisely, the Hausdorff distance of the estimated SV bounds to the true bounds decreases as the IV identification power increases.
Moving to the right part of each of table, first, we note that for each given IV set, all the estimated $C_1(x)$ to $C_4(x)$ and $IIP(x)$ converges to their true values as sample size $n$ increases, indicating that the estimated identification gain is more accurate for larger sample size.\footnote{Because $C_1(x)$ to $C_4(x)$ are functions of $L^M(x)$, $U^M(x)$, $\overline{\omega}(x)$ and $\omega^{SV}(x)$, the estimates of $C_1(x)$ to $C_4(x)$ are computed using the HMUE of the bounds or their widths. We compute $\overline{\omega}(x)$ as the width of the estimated bounds (by HMUE of CLR) $[\underline{L}^{SV}(x),\overline{U}^{SV}(x)]$ in (ref) if ATE$(x)>0$ is identified, or (ref) if otherwise.} We also note that the estimated $C_1(x)$ is the same for different IV sets. This result is intuitive because the identification gains brought by the IV validity should not vary with the IV strength. Comparison of Tables (ref) and (ref) or Tables (ref) and (ref) reveals that the impacts of endogeneity degree on IV identification power can be captured by the estimated $IIP(x)$. Importantly, the true ranking of $IIP(x)$ as in Table (ref) can be correctly revealed by finite sample estimates of $IIP(x)$.
\FloatBarrier
It is interesting to analyze the effect of adding an additional but completely irrelevant IV on the finite sample performance of ATE partial identification, by comparing the results obtained using IV sets (4) and (5). Adding $Z_3$ to $(Z_1,Z_2)$ actually produces a small decrease of the estimated $IIP(x)$, on average, for almost all different DGP designs considered in this section. The Cramer-Von Mises test and the Kolmogorov–Smirnov test confirm that the average values of the estimates of $IIP(x)$ under scenario (4) are significantly different from those obtained under scenario (5), when sample size is $N=500$ and $N=5000$ for both endogeneity degrees and for both case 1 and case 2. While when sample size is sufficiently large $N=10000$, the estimates of $IIP(x)$ under scenario (4) and (5) are no longer significantly different, except for case 2 with $\rho=0.8$. Particularly, from Table (ref) we can see that when $X$ is a binary variable, on average, the estimated SV bounds using $(Z_1,Z_2)$ are narrower than those estimated by the IV set including the irrelevant IV, especially for small sample size. Analyzing the results across the replications in case 2, we find that about 78% (for both endogeneity degrees) of the replications give narrower estimated SV bounds with IV set $(Z_1,Z_2)$ than those with $(Z_1,Z_2,Z_3)$, for sample size $N=500$; and this rate becomes to 53% ($\rho=0.5$) and 64% ($\rho=0.8$) for sufficiently large sample size $N=10000$. This suggests that in finite sample, the loss of information (efficiency) that arises from using irrelevant IV could lead to wider ATE bounds, especially when the covariate possesses limited variation.
On the other hand, the IV irrelevancy cannot always be detected by simply comparing the estimated SV bound width under different IV sets. That is, adding an irrelevant IV in (5) could further shrink the SV bound width when the covariate $X$ is continuous, although the improvement happens at the third decimal and the degree of the improvement decreases as sample size increases. \footnote{In case 1, the shrinkage of the estimated SV bounds using the irrelevant $Z_3$ is due to the finite sample estimation error. In particular, because the estimates of the coefficient of the irrelevant $Z_3$ will be nonzero with probability one, it results in more matched pairs of $(x,z)$ and $(x',z')$ such that $|$Pr$[D=1|x,z]-$Pr$[D=1|x',z']|<c$ (see Appendix (ref)) especially when covariate is continuous. For case 1 in Table (ref), we find that when sample size is $N=500$, (i) there are 22% ($\rho=0.5$) and 17% ($\rho=0.8$) of the 1000 replications where at least one (either lower or upper) estimated SV bound using $(Z_1,Z_2)$ is closer to its true value, compared to that obtained by using the irrelevant IV; and (ii) 12% of the replications yield wider estimated SV bounds when using the irrelevant IV, for both endogeneity degrees.}
Our analysis has several useful implications for policy design, choice of IVs for cost-effective identification targets, and assessing IV identification power in ATE bounding analysis.
First, as greater span of IV-driven propensity scores tightens the ATE bounds, increasing such span by creating IVs with stronger strength can be a consideration in designing policy instruments. For example, random allocation of treatment incentives or vouchers is often used as IV in studies with imperfect compliance. If there is a choice between two alternative IVs for encouraging treatment uptake with similar costs, the more segmented instrument design with both no incentives and high incentives to even a small proportion of participants, which leads to wider $(\underline{p}(x),\overline{p}(x))$, is preferable, compared to the less segmented ones with narrower span of the CPS. In addition, if an IV is continuous, it may not be a good practice to discretize it for simplicity purposes.
Second, for a specific analysis, sometimes the lower ATE bound being above a certain threshold, for example, zero, is more important for making the policy decision, whilst at other times, narrower ATE bounds may be more desirable. A priori knowledge, from economic theory or pilot experiments, on whether the direction of endogeneity is of the same sign as the ATE can be informative on the choice of identification target. When the ATE and the direction of endogeneity have opposite signs, the relatively narrow ATE bounds can be the target, because it is not too demanding of the IV strength and is sufficient to make policy recommendations. In contrast, if they appear to have the same sign, it is possible that increasing IV strength may only have limited scope for reducing ATE bound width. In this case, the researcher may instead want to focus on if the ATE lower bound exceeds a meaningful threshold. The attempt to create or collect stronger IVs in future studies may not be worthwhile, because stronger IVs would not help too much in improving the bounds.
In addition, for a given treatment variable, when there are multiple outcomes of interest, the preferable instrument design and identification target can be different, because the sign of the ATE and the direction of the treatment endogeneity may change with the choice of outcome.
Lastly, our simulation results show that in finite sample, using irrelevant IVs can result in a loss of efficiency and wider ATE bounds, especially when the covariate has limited variation. Our proposed $IIP$ index could be useful as a measure for detecting irrelevancy. Moreover, we reinforce a-fortiori the warning that simply adding extra IVs without assessing their relevance is unlikely to be a good practical strategy.
In this section, we apply our novel decomposition and IV evaluation method to study the effects of childbearing on women's labor supply. The dataset analyzed here is from the 1980 Census Public Use Micro Samples (PUMS), available at DVN/4W9GW2_2009. We follow the data construction in angrist1998children, where the sample consists of married women aged 21-35 with two or more children. The dateset contains 254,652 observations. The outcome is an indicator of being paid for work in the year prior to the census. The treatment is an indicator of having more than two children.
Following angrist1998children, we use as continuous regressors woman's age, woman's age at first birth, and ages of the first two children (quarters), and binary regressors for first child being a boy, second child being a boy, black, hispanic, and other race, as well as the interactions of the above mentioned continuous and indicator variables. For computational simplicity, we reduce dimension of covariates by utilizing the estimated propensity score conditional on $X$, $X_P:=\widehat{Pr}[D=1|X]$ as a covariate, which is estimated via a probit model and $X$ includes all of the regressors mentioned above. Three sets of IVs are considered: (1) the binary indicator that the first two children are the same sex (“Samesex”), (2) the binary indicator that the second birth is a twin (“Twins”), and (3) both indicators (“Both=\{Samesex,Twins\}”). To provide a comparison of SV bounds with other commonly used ATE bounding analyses, we also compute the ATE bounds in heckman2001instrumental (HV bounds) and chesher2010instrumental (Chesher bounds). To be consistent with our previous numerical analyses in Section (ref), we use the method of CLR to compute all the four bounds of interest.
Table (ref) reports the weighted average of the HMUE and the CLR two-sided confidence interval (CI) at 90%, 95% and 99% significant level of the four bounds of ATE$(X_P)$, with weights given by the estimated kernel density of $X_P$. Panels (a), (b) and (c) display the results using IV Samesex, Twins and Both, respectively. The weighted average of the Manski bounds estimates in all three panels are essentially identical, since the Manski bounds do not depend on IVs. In all panels, the HV bounds make an improvement over the benchmark Manski bounds, with the HV bound width using Twins being narrower than that using Samesex, and the HV bound width using Both being the narrowest. The Chesher bounds using \emph{Samesex} fail to identify the sign of the ATE$(X_P)$, as it is a union of both negative and positive intervals. When the IV \emph{Twins} or \emph{Both} is used instead, the weighted average of 95% CI of the Chesher bounds is $[-0.356,-0.007]$ (using \emph{Twins}) or $[-0.336,-0.022]$ (using \emph{Both}), revealing negative effects of having a third child on women's labor force participation.
For the SV bounds, the results using the IV Twins or Both dramatically outperform those using Samesex. The weighted average of 95% CIs using Samesex, Twins and Both are $[-0.548,-0.023]$, $[-0.276,-0.020]$ and $[-0.263,-0.038]$, respectively. The SV bounds estimates confirm the negative effect of a third child on women's labor force participation.\footnote{The two-stage least square (2SLS) estimates of angrist1998children give an ATE estimate of -0.123 with 95% CI of $[-0.178,-0.068]$ using IV \emph{Samesex}, and an estimate of -0.087 with 95% CI of $[-0.120,-0.054]$ using IV \emph{Twins}. As would be expected, the 95% two-sided CIs of all four bounds cover the 2SLS estimates and their associated 95% CIs for both IVs.} To summarize the results above, we can see that for ATE bounds in which the IV plays a key role in extracting identifying information, i.e. HV, Chesher and SV bounds, the IV \emph{Both} gives us the narrowest bounds.
\FloatBarrier
The ranking of IV identification power of the three IVs revealed by the discussion above is confirmed by the results in Table (ref). The results based on the 95% CI show that, on average, the identification power of Twins (66.6%) is significantly larger than that of Samesex (47.0%). We can see that whenever Twins$=1$ the treatment $D=1$ (i.e. perfect prediction), whereas this is not the case for Samesex. It is this feature, of course, that explains the superior performance when the HV, Chesher and SV bounds are evaluated using Twins compared to Samesex. When both \emph{Samesex} and \emph{Twins} are used, the identification power of \emph{Both} (70.1%) exceeds that of either \emph{Samesex} or \emph{Twins}. Closer inspection on the 95% CIs reveals that any valid IV contributes at least 44.7% of the SV bounds' improvement relative to the benchmark Manski bounds. Moreover, the identification power of \emph{Twins} is 66.6%/47.0%$\approx$1.4 times the identification power of \emph{Samesex} if used individually. In addition, given the contribution of IV validity $C_1$, the extra contribution of \emph{Twins} ($C_2=21.9\%$) via its IV strength to the identification power of \emph{Both} dominates the that of \emph{Samesex} ($C_2=2.3\%$). One remark is that, for the HV and Chesher bounds, IVs with higher $IIP$ clearly lead to narrower bounds.
To explore the heterogeneity of the treatment effects, Figure (ref) graphs the four bounds of interest against $X_P$. We can see that when the more powerful of the three IVs are employed, namely Twins or Both, the HV bounds narrow down the possible range of the ATE$(X_P)$ relative to the benchmark Manski bounds, especially for individuals with a small probability of having a third child. In addition, they can even identify the negative effect for individuals with $X_P$ close to zero. Similar properties are exhibited by the Chesher bounds. The 95% CI of SV bounds indicate that for women who are less likely to have more than two children, it is more probable that there will be a negative effect on their labor force participation once they have a third child, roughly in the region of -5% to -20%. For individuals who are more likely to have more than two children, the effect of having a third child is still negative but with larger possible range, roughly from -5% to -40% when $X_P$ is about 0.6, and roughly from 0% to -30% when $X_P$ is close to one.
To check the heterogeneity of the IV identification power, Figure (ref) displays the decompositions plotted against $X_P$. It depicts the fractions of estimated identification gains $\hat C_1(X_P)$ to $\hat C_4(X_P)$ over the estimated Manski bounds width $\hat\omega^M(X_P)$. It is obvious that the IV identification power of Twins and Both are significantly larger than that of Samesex, across all possible values of $X_P$. Furthermore, the contribution of the covariate appears to be amplified when Twins is involved in deriving the bounds, leading to a further reduction in the width of the unexplained part relative to the benchmark.
The ATE is a commonly used causal effect measure. Although there have been significant econometric developments, the choice among alternative modelling assumptions and IVs for estimating ATE can be an uninformed one for an empirical researcher without a good understanding of the role of IVs in studying the ATE. This paper aims to illustrate the mechanism through which IVs, and their interactions with treatment endogeneity and exogenous covariates achieve ATE identification gains in a partially identified heterogeneous treatment effect model. Our analysis is carried out by decomposing the identification gains of the ATE bounds in shaikh2011partial against the benchmark ATE bounds without IVs in manski1990nonparametric. We propose a measure to encapsulate IV identification power, $IIP$, and graphically illustrate the decomposition under alternative scenarios. Our simulation results show that the $IIP$ works well in finite samples as a tool for measuring IV identification power, selecting IVs, and detecting irrelevant ones.
We believe the analysis in this paper brings useful insights to empirical researchers who hope to understand IVs in partially identified models. As an example of these insights, we demonstrate that for the same IV strength, having a high degree of treatment endogeneity that has the opposite sign from that of the ATE can produce greater IV identification power and narrower ATE bounds, while the converse is true if the two have the same sign. In other words, the IV identification power in a partially identified model for binary dependent variables can be very different from the conventional IV strength measure that is only based on the treatment equation. We also discuss practical implications for policy instrument designs after obtaining data from preliminary trials. The analysis in our paper sheds new light on and offers a potential criterion for IV selection in high-dimensional settings within partially identified models. It also raises new questions as to what constitutes an adequate definition of weak IVs and testing for IV validity in conjunction with ATE bounding analyses. Explorations of these issues are left for future research.
{\setstretch{1.1} }