Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
172,936 characters · 22 sections · 140 citation commands
A Computational Approach to Identification of Treatment Effects for Policy Evaluation
For counterfactual policy evaluation, it is important to ensure that treatment parameters are relevant to the policies in question. This is especially challenging in the presence of unobserved heterogeneity. This challenge is well featured in the definition of the local average treatment effect (LATE). The LATE has been one of the most popular treatment parameters used by empirical researchers since it was introduced by imbens1994identification. It induces a straightforward linear estimation method that requires only a binary instrumental variable (IV), and yet, allows for unrestricted treatment heterogeneity. The unfortunate feature of the LATE is that, as the name suggests, the parameter is intrinsically local, recovering the average treatment effect (ATE) for a specific subgroup of population called compliers. This feature leads to two major challenges in making the LATE a reliable parameter for counterfactual policy evaluation. First, the subpopulation for which the effect is measured (e.g., via randomized experiments) may not be the population of policy interest. Second, the definition of the subpopulation depends on the IV chosen, rendering the parameter even more difficult to extrapolate to targeted environments.
Dealing with the lack of external validity of the LATE has been an important theme in the literature. One approach in theoretical work (angrist2010extrapolate,bertanha2019external) and empirical research (dehejia2019local,muralidharan2019disrupting) has been to show the similarity between complier and non-complier groups based on observables. This approach, however, cannot attend to possible unobservable discrepancies between these groups. heckman2005structural unify well-known treatment parameters by expressing them as weighted averages of what they define as the marginal treatment effect (MTE). This MTE framework has a great potential for extrapolation because a class of treatment parameters that are policy-relevant can also be generated as weighted averages of the MTE.\footnote{See heckman2010building for elaboration of this point.} The only obstacle is that the MTE is identified via a method called local IV (heckman1999local), which requires the continuous variation of the IV that is sometime large depending on the target parameter. This in turn reflects the intrinsic difficulty of extrapolation when available exogenous variation is only discrete. Acknowledging this nature of the challenge, previous studies in the literature have proposed imposing shape restrictions on the MTE, which is a function of the treatment-selection unobservable, while allowing for binary instruments in the framework of heckman2005structural. brinch2017beyond introduce shape restrictions (e.g., linearity) on the MTE functions in an attempt to identify the LATE extrapolated to different subpopulations or to test for its external validity. In interesting recent work, mogstad2018using propose a general partial identification framework where bounds on various policy-relevant treatment parameters can be obtained from a set of “IV-like estimands” that are directly identified from the data and routinely obtained in empirical work. kowalski2021reconciling applies an approach similar to these studies to extrapolate the results from one health insurance experiment to an external setting.
This paper continues this pursuit and investigates the possibility of extrapolating local treatment parameters to different policy settings in the MTE framework when IVs are possibly only binary. We propose a computational approach to calculate sharp nonparametric bounds on various extrapolated treatment parameters for discrete and continuous outcomes. We use IVs that satisfy the statistical independence assumption conditional on covariates. The parameters are defined as weighted averages of the MTE. Examples include the ATE, the treatment effect on the treated, the LATE for subgroups induced by new policies, and the policy-relevant treatment effect (PRTE). We also show how to place in this procedure restrictions from a large menu of identifying assumptions beyond the shape restrictions considered in earlier work.
In this paper, we make four main contributions. First, we propose a novel framework for systematically calculating bounds on policy-relevant treatment parameters. We introduce the distribution of the latent state of the outcome-generating process conditional on the treatment-selection unobservable. This latent conditional distribution is the key ingredient for our analysis, as both the target parameter and the distribution of the observables can be written as linear functionals of it. Because the latent distribution is a fundamental quantity in the data-generating process, it is convenient to impose identifying assumptions. Having the latent distribution as a decision variable, we can formulate infinite-dimensional linear programming (LP) that produces bounds on a targeted treatment parameter. Our approach is reminiscent of balke1997bounds and can be viewed as its generalization to the MTE framework. balke1997bounds characterize bounds on the ATE using a binary outcome, treatment and instrument by introducing a LP approach with the latent response vector as the decision variable. The main distinction of our approach is that the latent distribution is conditioned on the selection unobservable, which makes the program infinite-dimensional, but is important for our extrapolation purpose. We also allow for both discrete and continuous $Y$. To make it feasible to solve the resulting infinite-dimensional program, we use a sieve-like approximation of the program and produce a finite-dimensional LP. We also develop a method to rescale the LP to resolve computational issues that arise with a large sieve dimension.
The use of approximation to construct an LP is similar to mogstad2018using's approach. However, the approach we take differs from theirs in the following way. While they use the MTE function (more precisely, each term in the MTE) as the main ingredient to relate IV-like estimands to target parameters, we use the latent distribution as our main building block to relate the full distribution of the data to target parameters. The main consequence of this difference is that we can exhaust the identifying power of statistical independence of IVs, while their approach can exploit mean independence. The two approaches are complementary. For example, when IVs are generated from randomized experiments, one can comfortably assume full independence, in which case our approach can be applied to enjoy the tighter bounds than those under mean independence. This can be useful when the external validity of experimental results is in question, making the extrapolation of the LATE desirable.
Second, we introduce identifying assumptions that have not been used in the context of the MTE framework or the LATE extrapolation. They include assumptions that there exist exogenous variables other than IVs. One of the main messages we hope to deliver in this paper is that, given the challenge of extrapolation, additional exogenous variation can be useful to conduct informative policy evaluation. We propose two types of exogenous variables that have been used in the literature in the context of identifying the ATE: SV11, mourifie2015sharp, han2017identification, vuong2017counterfactual, and han2019estimation use the first type (entering the outcome and selection equations), and VY07, liu2020two, and balat2020multiple use the second type (only entering the outcome equation). For example, the existence of the second type can be plausible when the agent has imperfect foresight when making the treatment selection decision. We utilize these variables in the context of the MTE framework. Moreover, while the existing papers on the ATE make use of these variables in combination with rank similarity, rank invariance, or additive separability, we show that they independently have identifying power for treatment parameters, including the ATE. In general, it may not be always easy to find such exogenous variables. But when the researcher does find it, it can be a more reliable source of identification than assumptions on counterfactual quantities, as the identifying power comes from the data rather than the researcher's prior.
We also propose identifying assumptions that restrict treatment effect heterogeneity. In particular, we propose a range of uniformity assumptions that relate to rank similarity or rank invariance (chernozhukov2005iv) and monotone treatment response in MP00, including a novel identifying assumption, called rank dominance. The direction of endogeneity can also be incorporated in this MTE framework. This assumption is sometimes imposed in empirical work to characterize selection bias and has been shown to have identifying power for the ATE (MP00).
Third, we show that our approach yields straightforward proof of the sharpness of the resulting bounds, no matter whether the outcome is discrete or continuous and whether additional identifying assumptions are imposed or not. This feature stems from the use of the latent conditional distribution in the linear programming and the convexity of the feasible set in the program. When the MTE itself is the target parameter, we distinguish between the notions of pointwise and uniform sharpness and argue why uniform sharpness is often difficult to achieve.
Fourth, as an application, we study the effects of insurance on medical service utilization by considering various counterfactual policies related to insurance coverage. The LATE for compliers and the bounds on the LATE for always-takers and never-takers reveal that possessing private insurance tend to have the largest effect on medical visits for never-takers, i.e., those who face higher insurance cost. This provides a policy implication that lowering the cost of private insurance is important, because the high cost might hinder people with most need from receiving adequate medical services.
The linear programming approach to partial identification of treatment effects was pioneered by balke1997bounds and recently gained attention in the literature; see, e.g., Chi10, mogstad2018using, machado2018instrumental, kamat2017identification, gunsilius2019bounds, han2019optimal, russell2021sharp, han2023quantile.\footnote{There are also studies that use linear programming for partial identification of parameters that are not necessarily treatment effects; see e.g., honore2006bounds, honore2006bounds2, freyberger2015identification, torgovitsky2019partial, and gu2022partial.} As these papers suggest, there are many settings, including ours, where analytical derivation of bounds is cumbersome or nearly impossible due to the complexity of the problems. Also, the computational approach can streamline the sensitivity analysis of a researcher without needing to analytically derive bounds and prove their sharpness whenever changing the set of identifying assumptions.
As concurrent work to ours, marx2020sharp also considers partial identification of policy-relevant treatment parameters in the MTE framework. In his paper, sharp analytical bounds are derived for treatment parameters for the subset of compliers, and the identifying power of rank similarity and covariates is explored for general treatment parameters. The current paper is similar to his in that we also fully exhaust statistical independence (rather than mean independence) and produce sharp bounds. However, our approach differs in a few important ways. First, we provide a computational framework that enables the systematic calculation of bounds. Second, with the computational approach, we produce bounds for a range of treatment parameters under various identifying assumptions that have not been previously explored in this context. Also, the computational approach makes it convenient to conduct sensitivity analyses with different sets of assumptions.
This paper proceeds as follows. The next section introduces the main observables, maintained assumptions, and parameters of interest. Section (ref) defines the latent conditional probability and formulates the infinite-dimensional LP, and Section (ref) introduces sieve approximation to the program. Section (ref) establishes the connection between this paper and mogstad2018using. Section (ref) introduces additional identifying assumptions and shows how they can easily be incorporated in the LP. So far, the analysis is given with discrete $Y$, which is extended to the case with continuous $Y$ in Section (ref). Section (ref) provides numerical illustrations, and Section (ref) contains an empirical application. In the Appendix, Section (ref) lists other examples of target parameters. Section (ref) discusses (i) rescaling of the LP, (ii) the pointwise and uniform sharpness for the MTE bounds, (iii) the extension with continuous covariates, and (iv) estimation and inference. All proofs are contained in Section (ref). Additional numerical results can be found in Section (ref).
Assume that we observe a discrete or continuous outcome $Y\in\mathcal{Y}\subseteq\mathbb{R}$, binary treatment $D\in\{0,1\}$, and possibly discrete instrument $Z\in\mathcal{Z}\subseteq\mathbb{R}$. The leading case is binary $Z$, which is common especially in randomized experiments. We may additionally observe an exogenous variable $W\in\mathcal{W}\subseteq\mathbb{R}$ and possibly endogenous covariates $X\in\mathcal{X}\subseteq\mathbb{R}^{d_{X}}$.
Let $Y(d)$ be the counterfactual outcome given $d$ and $Y(d,w)$ be the extended counterfactual outcome given $(d,w)$, which are consistent with the observed outcome: $Y=\sum_{d\in\{0,1\}}1\{D=d\}Y(d)=\sum_{d\in\{0,1\},w\in\mathcal{W}}1\{D=d,W=w\}Y(d,w)$.
When there is no $W$ this assumption and all below are understood as $W$ being degenerate. Assumption EX imposes the exclusion restriction and conditional statistical independence for $Z$ (and $W$). One of the contributions of this paper is to propose a framework that can make use of full independence instead of mean independence (mogstad2018using) and show the identifying power of the former relative to the latter. This feature arises regardless of the existence of $W$. We formally discuss the identifying power of full independence relative to mean independence in Section (ref). Also, note that Assumption EX imposes marginal independence rather than joint independence of $\{Y(d,w),U\}_{(d,w)\in\{0,1\}\times\mathcal{W}}\perp(Z,W)|X$.
We consider two different scenarios related to $W$: (a) $W$ directly affects $Y$ but not $D$ and (b) $W$ directly affects both $Y$ and $D$. Accordingly, we maintain the following assumptions.
We introduce $W$ as an additional exogenous variable researchers may be equipped with in addition to the instrument $Z$. Given the challenge of extrapolation with minimal variation in $Z$, it would be important to search for additional exogenous variables. In the case of (a), such variables can be motivated by exogenous shocks that agents cannot fully anticipate at the time of making treatment choices. For example, let $Y$ be the earning and $D$ be the college attendance. In this example, $W$ can be a randomized job training program that directly affects $Y$ but whose lottery outcome cannot be foreseen when making the college decision. As another example, when $Y$ is the health outcome and $D$ is getting an insurance, $W$ can be random health or policy shocks that cannot be fully anticipated when making the insurance decision. In the context of generalized Roy models, (a) is consistent with agents' limited information when comparing the potential outcomes of the two treatment states. We show the identifying power of $W$ even with its minimal variation. This is the first paper that formally uses this type of variable for identification in the MTE framework.\footnote{Relatedly, eisenhauer2015generalized allows variables of type (a) in the context of generalized Roy models. However, they consider these variables only as a feature of agent's limited information but not as a source of identification. Specifically, their identification of the MTE and cost parameters only relies on the exogenous variables that are known to the agent at the time of selection into treatment (i.e., using their notation, they rely on $Z$ to identify the MTE and $X_{I}$ to identify the cost parameters); see pp. 430-431 of their paper.} Note that the requirement of reverse exclusion of $W$ in (a) can be tested from the data by inspecting whether $\Pr[D=1|Z,X,W]=\Pr[D=1|Z,X]$.
Assumption SEL imposes a selection model for $D$, which is important in motivating and interpreting marginal treatment effects later. This assumption is also equivalent to imbens1994identification's monotonicity assumption (vytlacil2002independence). We introduce the standard normalization that $U\sim Unif[0,1]$ conditional on $X=x$.\footnote{Note that for any index function $g(z,x)$ and an unobservable $\varepsilon$ with any distribution, the selection model satisfies $D=1\{\varepsilon\le g(Z,X)\}=1\{F_{\varepsilon|X}(\varepsilon|X)\le F_{\varepsilon|X}(g(Z,X)|X)\}=1\{U\le P(Z,X)\}$, since $P(z,x)=\Pr[\varepsilon\le g(z,x)|X=x]=\Pr[U\le F_{\varepsilon|X}(g(z,x)|x)|X=x]=F_{\varepsilon|X}(g(z,x)|x)$ and $F_{\varepsilon|X}(\varepsilon|X)=U$ is uniformly distributed conditional on $X$.}
In Assumption SEL, Case (a) is where $W$ is a reversely excluded exogenous variable, which we call reverse IV. This type of exogenous variables was considered by VY07, balat2020multiple, and liu2020two. In Case (b), we show that a reverse IV is not necessary, and $W$ can be present in the selection equation. This type of exogenous variables was considered by SV11, mourifie2015sharp, han2017identification, vuong2017counterfactual, and han2019estimation. In both scenarios, however, we show that we can use $W$ for identification without necessarily invoking rank similarity, rank invariance, or additive separability in contrast to the above studies. We show this is possible due to the computational approach we take.�Below, we combine the existence of $W$ with assumptions that are related to rank similarity. Another distinct feature of our approach in comparison to the prior studies is that we consider a broad class of the generalized LATEs as our target parameter, including the ATE considered in those studies. For notational simplicity, we focus on Case (a) henceforth; it is straightforward to draw analogous results for Case (b).
We aim to establish sharp bounds on various treatment parameters. Following heckman2005structural, we express treatment parameters as integral equations of the MTE. The MTE is defined in our setting as
where $Y(d)=Y(d,W)$. Similar to mogstad2018using, it is convenient to introduce the marginal treatment response (MTR) function
where $W$ does not appear as a conditioning variable due to Assumption EX. Now, we define the target parameter $\tau$ to be the difference of the weighted averages of the MTRs:
where
by using $F_{U|X}(u|x)=u$, and $\omega_{d}(u,z,x)$ is a known weight specific to the parameter of interest. This definition agrees with the insight of heckman2005structural. The target parameter includes a wide range of policy-relevant treatment parameters. We list a few examples of the target parameter here; other examples can be found in Table (ref) in the Appendix.
To define a broader class of parameters beyond these examples, the weights $\omega_{0}$ and $\omega_{1}$ can be set asymmetrically. All the parameters we consider in this paper can be defined conditional on $X$ and $W$, although we omit them for succinctness.
Typically, a binary instrument is not sufficient in producing informative bounds on the target parameters. This is because a binary instrument has no extrapolative power for general non-compliers, e.g., always-takers and never-takers, but only identifies the effect for compliers. Prior studies have tried to overcome this challenge by imposing shape restrictions on the MTE (cornelissen2016late, brinch2017beyond, mogstad2018using, kowalski2021reconciling), although these restrictions are not always empirically justified. Evidently, it would be useful to provide empirical researchers with a larger variety of assumptions so that it is easier to find justifiable assumptions that suit their specific examples.
The existence of additional exogenous variables embodied in Assumptions SEL and EX may be appealing as it can be warranted by data with less arbitrariness. We accompany Assumptions SEL and EX with an assumption that $W$ and $Z$ are relevant variables, which make the role of these variables more explicit.
Assumption R(i) is a relevance condition for $W$ in determining $Y$. R(ii) is the standard relevance assumption for the instrument and the positivity assumption. We later show that under Assumptions SEL, EX and R, the variation of $W$ (in addition to $Z$) is a useful source for extrapolation and narrowing the bounds on target parameters.
Our goal is to provide a systematic framework to calculate bounds on the target parameters, which is easy to incorporate various identifying assumptions. To this end and as a crucial first step of our analysis, we define a state variable that determines a specific mapping of
for discrete $y\in\mathcal{Y}=\{y_{1},...,y_{L}\}$. We discuss the extension with continuously distributed $Y$ in Section (ref). Given that $d$ is binary and assuming $w$ is also binary, there are $L^{4}$ possible states or maps from $(d,w)$ onto $y$. Using the extended counterfactual outcome $Y(d,w)$, define a latent vector $\epsilon$ as
and its realized value as $e\equiv(y(0,0),y(0,1),y(1,0),y(1,1))\in\mathcal{E}=\mathcal{Y}^{4}$. Then, each value of $e$ represents each possible state. Table (ref) lists all 16 maps in a leading case of binary $y$ with $\mathcal{Y}=\{0,1\}$.
Now, as a key component of our LP, we define the probability mass function of $\epsilon$ conditional on $(U,X)$: for $e\in\mathcal{E}$,
with $\sum_{e\in\mathcal{E}}q(e|u,x)=1$ for any $(u,x)$. The quantity $q(e|u,x)$ captures the joint distribution of $(\epsilon,U)$ and thus reflects endogenous treatment selection. It is shown below that this latent conditional probability is a building block for various treatment parameters and thus serves as the decision variable in the LP. The introduction of $q(e|u,x)$ distinguishes our approach from those in balke1997bounds and mogstad2018using. Since the probability is conditional on continuously distributed $U$, the simple finite-dimensional linear programming approach of balke1997bounds is no longer applicable. Instead, we use an approximation method similar to mogstad2018using. However, mogstad2018using uses the MTR function as a building block for treatment parameters and introduces the “IV-like” estimands as a means of funneling the information from the data. Unlike in mogstad2018using, $q(e|u,x)$ can be directly related to the distribution of data. This allows us to (i) fully exhaust the full independence assumption, (ii) facilitate proving sharpness and (iii) incorporating a large menu of additional identifying assumptions.
In the remaining section and the next section, we focus on binary $Y$ for simplicity; the extension to general discrete $Y$ is straightforward and is discussed in Section (ref); continuous $Y$ is considered in Section (ref). Also, we will mostly focus on binary $W$ and discrete $X$ for expositional simplicity. Section (ref) in the Appendix extends the framework to incorporate continuously distributed $X$; it is also straightforward to extend to allow for general discrete variable $W$. By (ref), note that
Therefore, the MTR can be expressed as
Combining (ref) and (ref), we have $\tau_{d}(z,w,x)=\sum_{e:y(d,w)=1}\int q(e|u,x)\omega_{d}(u,z,x)du$, and thus the target parameter $\tau=E[\tau_{1}(Z,W,X)]-E[\tau_{0}(Z,W,X)]$ in (ref) can be written as
for some $q$ that satisfies the properties of probability.
The goal of this paper is to (at least partially) infer the target parameter $\tau$ based on the data, i.e., the distribution of $(Y,D,Z,W,X)$. The key insight is that there are observationally equivalent $q(e|u,x)$'s that are consistent with the data, which in turn produces observationally equivalent $\tau$'s that define the identified set.
Let $p(y,d|z,w,x)\equiv\Pr[Y=y,D=d|Z=z,W=w,X=x]$ be the observed conditional probability. This data distribution imposes restrictions on $q(e|u,x)$. For instance, for $D=1$,
by Assumption EX, but
where the second equality is by $\Pr[Y(d,w)=y|U=u,X=x]=\sum_{e:y(d,w)=y}q(e|u,x)$.
To define the identified set for $\tau$, we introduce some simplifying notation. Let $q(u)\equiv\{q(e|u,x)\}_{e\in\mathcal{\mathcal{E}},x\in\mathcal{X}}$ and \[ \mathcal{Q}\equiv\{q(\cdot):\sum_{e\in\mathcal{E}}q(e|u,x)=1\text{ and }q(e|u,x)\ge0\text{ }\forall(e,u,x)\} \] be the class of $q(u)$, and let \[ p\equiv\{p(1,d|z,w,x)\}_{(d,z,w,x)\in\{0,1\}\times\mathcal{Z}\times\mathcal{W}\times\mathcal{X}}. \] Also, let $R_{\omega}:\mathcal{Q}\rightarrow\mathbb{R}$ and $R_{0}:\mathcal{Q}\rightarrow\mathbb{R}^{d_{p}}$ (with $d_{p}$ being the dimension of $p$) denote the linear operators of $q(\cdot)$ that satisfy
where $\mathcal{U}_{z,x}^{d}$ denotes the intervals $\mathcal{U}_{z,x}^{1}\equiv[0,P(z,x)]$ and $\mathcal{U}_{z,x}^{0}\equiv(P(z,x),1]$. Then, we can characterize the baseline identified set for $\tau$ where we only impose modeling primitives. Later, we show how to characterize the identified set with additional assumptions introduced in Section (ref).
In what follows, we formulate the infinite-dimensional LP ($\infty$-LP) that characterizes $\mathcal{T}^{*}$. This program conceptualizes sharp bounds on $\tau$ from the data and the maintained assumptions (Assumptions SEL and EX). The upper and lower bounds on $\tau$ are defined as
subject to
Observe that the set of constraints (ref) does not include
This is because we know a priori that they are redundant in the sense that they do not further restrict the feasible set, namely, the set of $q(e|u,x)$'s that satisfy all the constraints ($q\in\mathcal{Q}$ and (ref)).
The result of this theorem is immediate due to the convexity of the feasible set $\{q:q\in\mathcal{Q}\}\cap\{q:R_{0}q=p\}$ in the LP and the linearity of $R_{\omega}q$ in $q$, which implies that $[\underline{\tau},\overline{\tau}]$ is convex.
Although conceptually useful, the LP (ref)--(ref) is not feasible in practice because $\mathcal{Q}$ is an infinite-dimensional space. In this section, we approximate (ref)--(ref) with a finite-dimensional LP via a sieve approximation of the conditional probability $q(e|u,x)$. We use Bernstein polynomials as the sieve basis. Bernstein polynomials are useful in imposing restrictions on the original function (joy2000bernstein; chen2011sensitivity; chen2017approximation) and therefore have been introduced in the context of linear programming (mogstad2018using; mogstad2021causal; masten2021salvaging). Bernstein approximation based on those polynomials possesses the property of being “shape-preserving,” which effectively prevent undesired distortions from the original function's shape properties (goodman1989shape; carnicer1993shape; goodman1999convexity).
Consider the following sieve approximation of $q(e|u,x)$ using Bernstein polynomials of order $K$
where $b_{k}(u)\equiv b_{k,K}(u)\equiv\tbinom{K}{k}u^{k}(1-u)^{K-k}$ is a univariate Bernstein basis, $\theta_{k}^{e,x}\equiv\theta_{k,K}^{e,x}\equiv q(e|k/K,x)$ is its coefficient, and $K$ is finite. By using this approximation, we implicitly impose smoothness in $q(e|\cdot,x)$. Note that discrete $x$ can index $\theta$, because $q(e|u,x)$ is a saturated function of $x$. By the definition of the Bernstein coefficient, for any $(e,x)$, it satisfies $q(e|u,x)\ge0$ for all $u$ if and only if $\theta_{k}^{e,x}\ge0$ for all $k$. Also, $\sum_{e\in\mathcal{E}}q(e|u,x)=1$ for all $(u,x)$ is approximately equivalent to $\sum_{e\in\mathcal{E}}\theta_{k}^{e,x}=1$ for all $(k,x)$. To see this, first, $\sum_{e\in\mathcal{E}}q(e|u,x)=1$ for all $(u,x)$ implies $\sum_{e\in\mathcal{E}}\theta_{k}^{e,x}=\sum_{e\in\mathcal{E}}q(e|k/K,x)=1$ for all $(k,x)$. Conversely, when $\sum_{e\in\mathcal{E}}\theta_{k}^{e,x}=1$ for all $(k,x)$,
by the binomial theorem (coolidge1949story). Motivated by this approximation, we formally define the following sieve space for $\mathcal{Q}$:
Let $\mathcal{K}\equiv\{1,...,K\}$ and $p(z,w,x)\equiv\Pr[Z=z,W=w,X=x]$. For $q\in\mathcal{Q}_{K}$, by (ref) and (ref), the target parameter $\tau=E[\tau_{1}(Z,W,X)]-E[\tau_{0}(Z,W,X)]$ can be expressed with
where $\gamma_{k}^{d}(w,x)\equiv\sum_{z\in\{0,1\}}p(z,w,x)\int b_{k}(u)\omega_{d}(u,z,x)du$. Also, for $q\in\mathcal{Q}_{K}$ and $D=1$, by (ref), we have
where $\delta_{k}^{d}(z,x)\equiv\int_{\mathcal{U}_{z,x}^{d}}b_{k}(u)du$.
From (ref) and (ref), we can expect that a finite-dimensional LP can be obtained with respect to $\theta_{k}^{e,x}$. Let $\theta\equiv\{\theta_{k}^{e,x}\}_{(e,k,x)\in\mathcal{E}\times\mathcal{K}\times\mathcal{X}}$ and let
Then, we can formulate the following finite-dimensional LP that corresponds to the $\infty$-LP in (ref)--(ref):
subject to
One of the advantages of LP is that it is computationally very easy to solve using standard algorithms, such as the simplex algorithm. Assuming binary $W$ and $X$ and setting $K=50$, we have $\dim(\theta)=1632$, and it takes less than 3 seconds to calculate $\overline{\tau}_{K}$ and $\underline{\tau}_{K}$ with moderate computing power. The increase in the support of $\mathcal{W}$ (and thus the number of maps (ref)) only linearly increases the computation time.
The important remaining question is how to choose $K$ in practice. We discuss this issue in Section (ref). Also, when $K$ is too large, coefficients on $\theta_{k}^{e,x}$ in (ref) tend to take values with incomparable orders of magnitude, which may cause common optimization algorithms to arbitrarily drop small values. To address this, we propose a rescaling method in Section (ref). Finally, extending Proposition 4 in mogstad2018using, we may exactly calculate $\overline{\tau}$ and $\underline{\tau}$ (i.e., $\overline{\tau}=\overline{\tau}_{K}$ and $\underline{\tau}=\underline{\tau}_{K}$) under the assumptions that (i) the weight function $\omega_{d}(u,z,x)$ is piece-wise constant in $u$ and (ii) the constant spline that provides the best mean squared error approximation of $q(e|u,x)$ satisfies all the maintained assumptions (possibly including the identifying assumptions introduced later) that $q(e|u,x)$ itself satisfies; see mogstad2018using for details.
The framework proposed in this paper allows us to systematically calculate sharp bounds of treatment parameters using the (conditional) statistical independence of $Z$ and $W$ (Assumption EX). This is possible because the full distribution from the data enters the LP. This is in contrast to the approach in mogstad2018using who utilize the first moment information of the data and exploit mean independence of $Z$ in their bound analysis. The full independence assumption is common in the nonparametric treatment effect literature and can be justified by, for example, IVs generated by randomized experiments. This section formally show the identifying power of full independence and, more importantly, how our framework allows us to make use of it to produce sharp bounds. To this end, we compare our approach with mogstad2018using's and show when the identified sets coincide and when they do not. We first focus on a simple case where $Z$ is the only exogenous variable. Then, we discuss how $W$ can add further identifying power.
To help the reader, we restate the maintained assumption in mogstad2018using in terms of our notation:
Assumption I.2 imposes mean independence of the instrument $Z$. On the other hand Assumption EX (after suppressing $W$ for clear comparison) implies $Y(d)\perp Z|X,U$, namely, statistical independence of $Z$. Therefore, we can expect that the latter has stronger identifying power. Below, we show that the LP constructed with $q$ being the choice variable instead of the MTR function $m_{d}$ enables us to exploit this extra identifying power.
In mogstad2018using the IV-like estimands are the channel through which the data enters to restrict the set of $m_{d}$. They show (their Proposition 3) that if the IV-like estimands are carefully chosen, then the resulting set of $m_{d}$ is equivalent to the set of $m_{d}$ that are consistent with
Therefore, assuming such IV-like estimands are chosen, we can define their identified set as \[ \mathcal{M}_{id}=\Big\{ m=(m_{0},m_{1}),m_{0},m_{1}\in L^{2}:m_{0},m_{1}\text{ satisfies equation (\ref{eq: infoset0}) and (\ref{eq: infoset1}) a.s.}\Big\}. \] Note that this set is the result of funneling the information from data via the conditional means.
When $Y$ is binary, the mean and full independence assumptions are equivalent. Therefore we expect to find no difference in the resulting identified sets of treatment parameters between the two methods. We show this by comparing the identified set of the MTR functions $\mathcal{M}_{id}$ used in mogstad2018using and the set of MTR functions derived from the feasible set in the LP proposed in the current paper (denoted as $\mathcal{M}_{f}$ below). More importantly, as soon as $Y$ departs from binarity, we show that $\mathcal{M}_{f}$ is a proper subset $\mathcal{M}_{id}$.
For comparison, define the feasible set $\mathcal{Q}_{f}$ of our proposed LP as \[ \mathcal{Q}_{f}=\Big\{ q\in L^{2}([0,1]):q\in\mathcal{Q}\text{ and satisfies equation }(\ref{eq:constr})\Big\}, \] where $w$ is suppressed from (ref). To establish the connection with $\mathcal{M}_{id}$, we construct the set of MTR functions based on $\mathcal{Q}_{f}$. For this purpose, we drop $W$ from the setting and redefine $\epsilon\equiv(Y(0),Y(1))$ and $e\equiv(y(0),y(1))$:
\[ \mathcal{M}_{f}=\Big\{ m=(m_{0},m_{1}):m_{d}(u,x)=\sum_{y\in\mathcal{Y}}y\sum_{e:y(d)=y}q(e|u,x),d=\{0,1\},q\in\mathcal{Q}_{f}\Big\}. \] Then the following holds:
Theorem (ref) demonstrates that, when $Y$ is binary, the constraint used in our LP approach captures the same information as the constraint defined by the IV-like estimands that are appropriately selected. Consistent with Proposition 3 in mogstad2018using, $\mathcal{M}_{f}$ and $\mathcal{M}_{id}$ are sharp in this case.
When $Y$ is non-binary and we continue to impose statistical independence, the above equivalence breaks down. That is, unlike $\mathcal{M}_{id}$, $\mathcal{Q}_{f}$ exploits the full data distribution and the statistical independence assumption. We formally show this with $\mathcal{Y}=\{y_{1},...,y_{L}\}$. We accordingly redefine $p$ and $R_{0}$ in (ref) (i.e., $R_{0}q=p$) in the definition of $\mathcal{Q}_{f}$ as
without excluding redundant constraints and suppressing $w$ for simplicity. We introduce a technical assumption.
When $Y$ takes more than two values, $\mathcal{Q}_{f}$ continues to exploit the full independence assumption and as shown in Theorem (ref), $\mathcal{M}_{f}$ is sharp. Note $\mathcal{M}_{id}$ is a strict outer set of $\mathcal{M}_{f}$ because the former only exploits the first moment implication of full independence. This intuition also suggests that the case of continuous $Y$ would deliver a similar result as in Theorem (ref). Theorems (ref) and (ref) corroborate the simulation results in Section (ref). The simulation results for continuous $Y$ can be found in Section (ref).
Finally, consider the case where $W$ is also present and satisfies Assumptions EX and SEL. After proper modification of their assumptions (i.e., $U\perp(Z,W)|X$ and $E[Y(d)|Z,W,X,U]=E[Y(d)|X,U]$), mogstad2018using's framework may be able to use the exogenous variation of $W$. This can be done by using the moments $E[Y|D=d,Z,W,X]$ ($d=0,1$) as inputs in (ref) and (ref).\footnote{Alternatively, the variation of $W$ can be reflected via IV-like estimands as inputs. In this case, the identification of $\tau_{LATE}(w,w')$ is useful. We show its identifiability in Lemma (ref) below.} However, by the same argument as the one above for $Z$, the full independence assumption with respect to $W$ will have a stronger identifying power than mean independence. When $Y$ is beyond binary, our framework can exploit this. The two frameworks also differ in how other additional identifying assumptions can be incorporated in the procedures. This point appears as Remark (ref) in the next section.
In addition to Assumptions SEL, EX and R, researchers may be willing to restrict the degree of treatment heterogeneity, the direction of endogeneity, or the shape of the MTR functions. Although not necessary, such restrictions play significant roles in yielding informative bounds, especially given the challenge of extrapolating the LATE with minimal variation of the exogenous variables. Among the proposed assumptions, researchers want to use those that they deem plausible in given applications. Also, we provide only a few examples of assumptions here; we believe our proposed framework may open a venue for other assumptions we have not explored.
We present a range of restrictions on treatment heterogeneity in the order of stringency starting from the strongest.
\begin{asU2}For every $w,w'\in\mathcal{W}$ and $x\in\mathcal{X}$, either $\Pr[Y(1,w)\ge Y(0,w')|X=x]=1$ or $\Pr[Y(1,w)\le Y(0,w')|X=x]=1$.\end{asU2}
The following assumption is weaker than Assumption U$^{*}$.
When $W$ is not available at all, this assumption can be understood with $\text{\ensuremath{\mathcal{W}}}$ being degenerate. In Assumption U$^{*}$, $w$ and $w'$ may be the same or different, i.e., the uniformity is for all combinations of $(w,w')\in\{(0,0),(1,1),(1,0),(0,1)\}$. Therefore, Assumption U$^{*}$ implies Assumption U. Assumptions U and U$^{*}$ posit that individuals present uniformity in the sense that the treatment either weakly increases the outcome for all individuals or decreases it for all individuals. Intuitively, Assumption U$^{*}$ is stronger because the uniformity remains to hold even under an outcome shift via $w\neq w'$; this may hold when the treatment effect is strong across individuals. Assumptions U and U$^{*}$ share insights with the monotone treatment response assumption that is introduced to bound the ATE in manski1997monotone and MP00. Assumptions U and U$^{*}$ are also related to rank invariance considered in the literature (e.g., chernozhukov2005iv, marx2020sharp). To see this, consider binary $Y$ and a structural model $Y=1[s(D,W)\ge V_{D}]$ with $V_{D}\equiv DV_{1}+(1-D)V_{0}$ (suppressing $X$). When rank invariance ($V_{1}=V_{0}\equiv V$) holds, Assumption U$^{*}$ holds (and thus Assumption U) because either $P[s(1,w)<V,s(0,w')\ge V]=0$ or $P[s(1,w)\ge V,s(0,w')<V]=0$. On the other hand, Assumption U$^{*}$ can still hold even if $V_{1}\neq V_{0}$ when, for example, the distribution of $(V_{1},V_{0})$ is concentrated around $45^{\circ}$ line. Therefore the converse is not true.\footnote{On the other hand, Assumption U$^{*}$ and rank similarity ($F_{V_{1}|U}=F_{V_{0}|U}$) are not nested.} It is important to note that Assumptions U and U$^{*}$ still allow treatment heterogeneity in terms of $X$ (and similarly of $W$). For instance, Assumption U allows that $Y(1,w)\ge Y(0,w)$ a.s. for $X=x$ but $Y(1,w)\le Y(0,w)$ a.s. for $X=x'$.\footnote{A related idea of conditional rank preservation appears in han2021identification and also used in marx2020sharp.} Assumption U is also used, for example, in chiburis2010semiparametric although his focus is bounds on the ATE with binary $Y$.
In fact, it is possible to further weaken Assumption U. To motivate this with binary $Y$, we state Assumption U as follows (suppressing $X$): for every $w$, either $P[Y(1,w)=0,Y(0,w)=1]=0$ or $P[Y(1,w)=1,Y(0,w)=0]=0$. In other words, all individuals respond weakly monotonically to the treatment in the sense that there is no (strictly) negative-treatment-response type or positive-treatment-response type. By iterated expectation, Assumption U holds if and only if either $P[Y(1,w)=0,Y(0,w)=1|U]=0$ a.s. or $P[Y(1,w)=1,Y(0,w)=0|U]=0$ a.s. Then this assumption can be relaxed by allowing for the existence of both types in the population and instead assuming the dominance of one of the types over the other: for example, $P[Y(1,w)=1,Y(0,w)=0|U]>P[Y(1,w)=0,Y(0,w)=1|U]$ a.s. For general $Y$, such an assumption is written as follows (ignoring a measure zero set of $U$):
\begin{asU3}For every $w\in\mathcal{W}$ and $x\in\mathcal{X}$, either $P[Y(1,w)\ge Y(0,w)|U=u,X=x]\ge P[Y(1,w)\le Y(0,w)|U=u,X=x]$ for all $u$ or $P[Y(1,w)\ge Y(0,w)|U=u,X=x]\le P[Y(1,w)\le Y(0,w)|U=u,X=x]$ for all $u$.
\end{asU3}
Assumption U$^{0}$ is weaker than Assumption U, because when $P[Y(1,w)\le Y(0,w)|X=x]=0$, Assumption U$^{0}$ trivially holds.\footnote{One can come up with assumptions which strength is between Assumption U$^{0}$ and Assumption U. For example with binary $Y$, one can assume that, in addition to Assumption U$^{0}$, $P[Y(1,w)=1,Y(0,w)=1|X=x]\ge P[Y(1,w)=0,Y(0,w)=1|X=x]$ and $P[Y(1,w)=0,Y(0,w)=0|X=x]\ge P[Y(1,w)=0,Y(0,w)=1|X=x]$ hold. We do not explore these assumptions for succinctness.} Compared to Assumption U, Assumption U$^{0}$ allows further treatment heterogeneity in that positive- and negative-treatment-response types can both present in the population, but we impose a uniform order on the probabilities of effect sign across individuals. Researchers may be more comfortable with this assumption than complete uniformity. This paper is the first to propose to use this restriction in the identification of treatment effects. We call assumptions of this type rank dominance.
We show that the directions in Assumptions U$^{*}$--U$^{0}$ may be learned from the data. Suppress $X$ for simplicity and let $Z\in\{0,1\}$. Let $\tau_{LATE}(w,w')\equiv E[Y(1,w)-Y(0,w')|P(0)\le U\le P(1)]$ be the LATE defined using the extended counterfactual outcome $Y(d,w)$ assuming Assumption SEL(a) and $P(1)>P(0)$; the definition of $\tau_{LATE}(w,w')$ under Assumption SEL(b) is analogous.
The practical implication of this lemma is that the direction of each assumption will be detected by the feasibility of relevant LP upon imposing it; the next subsection has more details. The identification of $\tau_{LATE}(w,w')$ extends the result in imbens1994identification. In particular,
The intuition behind (ii) and (iii) is as follows. The directions of monotonicity in Assumptions U$^{*}$ and U can be identified by the signs of relevant LATEs. On the other hand, the direction of inequality in Assumption U$^{0}$ can only be identified with binary $Y$, because otherwise the proportions of positive and negative treatment effects can be offset by the magnitude of the effects. This limited testability reflects the weak identifying power of Assumption U$^{0}$. The detail of the intuition can be found in the formal proof of Lemma (ref) in the Appendix. Previous work has discussed the role of the rank similarity assumption on determining the sign of the ATE (bhattacharya2008treatment,SV11,han2019optimal), and the result above shows that Assumptions U$^{*}$--U$^{0}$ play a similar role in the LP approach.
In some applications, researchers are relatively confident about the direction of treatment endogeneity. The idea of imposing the direction of the selection bias as an identifying assumption appears in MP00, who introduce monotone treatment selection (MTS), in addition to the monotone treatment response assumption mentioned above.
Monotonicity and concavity are common shape restrictions used in the context of MTR and MTE framework.
Assumption M assumes monotonicity and appears in brinch2017beyond and mogstad2018using; while Assumption C assumes concavity and appears in mogstad2018using. Another shape restriction introduced in the literature is separability: $m_{d}(u,w,x)=m_{1d}(w,x)+m_{2d}(w,u)$ is weakly concave in $u\in[0,1]$.
In this section, we show how identifying assumptions introduced in Section (ref) can be easily translated into assumptions on the mapping defined in (ref). This allows us to incorporate the additional assumptions in the formulation of the LP, so that one does not need to manually derive analytical bounds every time she imposes a new assumption (and prove their sharpness).
Before proceeding, we revisit Assumptions SEL, EX, and R in the context of the LP. First, we formally show that the existence and relevance of $W$ (as well as $Z$) embodied in Assumptions SEL, EX, and R can be a useful source in narrowing the bounds.
Heuristically, the improvement occurs because, with R(i), the constraint matrix (i.e., the matrix multiplied to the vector $\theta$ in (ref)) has greater rank with the variation of $W$ than without. See the proof of the theorem for a formal argument. Note that non-redundant constraints on $\theta$ do not always guarantee an improvement of the bounds in (ref)--(ref), because these constraints may still be non-binding. Nevertheless, non-redundancy is a necessary condition for the improvement.
We now show how to incorporate Assumptions U$^{*}$, U, U$^{0}$, MTS, M, and C as additional equality and inequality restrictions in the LP: Given the LP (ref)--(ref), identifying assumptions can be imposed by appending
where $R_{1}$ and $R_{2}$ are linear operators on $\mathcal{Q}$ that correspond to equality and inequality constraints, respectively, and $a_{1}$ and $a_{2}$ are some vectors in Euclidean space. Then, analogous constraints on $\theta$ can be added to the finite-dimensional LP (ref)--(ref). When an assumption violates the true data-generating process, then the identified set will be empty. This corresponds to the situation where the LP does not have a feasible solution. When we reflect sampling errors, this corresponds to the case where the confidence set is empty.\footnote{In order to verify whether the identified set is empty, we need to check whether the feasible set of $\theta$ is empty. An efficient way to do this is to identify vertices of the feasible polytope, if any. This process is no simpler than the simplex algorithm that we use to solve the LP. Therefore, we recommend that one first solves the LP and check if infeasibility is reported.}
Assumptions U and U$^{*}$ are imposed in the LP by “deactivating” relevant maps. Motivating by Lemma (ref), suppose the researcher knows that $Y(1,w)\ge Y(0,w)$ almost surely for all $w\in\{0,1\}$ under Assumption U. This assumption can be imposed as equality constraints (ref), namely, in the form of $R_{1}q=a_{1}$: Suppressing $x$ for simplicity and recalling $\epsilon\equiv(Y(0,0),Y(0,1),Y(1,0),Y(1,1))$,
respectively, corresponding for $w=1$ and $w=0$ in $Y(1,w)\ge Y(0,w)$. Therefore, the corresponding $\theta_{k}^{e}=0$. Then, the effective dimension of $\theta$ will be reduced in (ref)--(ref) and thus yields narrower bounds. As another example, suppose the researcher knows that the following holds almost surely under Assumption U$^{*}$: $Y(1,1)\ge Y(0,0)$, $Y(1,0)\ge Y(0,1)$, $Y(1,1)\ge Y(0,1)$, and $Y(1,0)\ge Y(0,0)$. These inequalities respectively imply
and (ref)--(ref). Recall the discussion that, in Assumption U (Assumption U$^{*}$), the direction of monotonicity is allowed to be different for different $w$ ($(w,w')$ pairs). This direction will be identified from the data (Lemma (ref)). Specifically, the direction can be automatically determined from the LP by inspecting whether the LP has a feasible solution; when wrong maps are removed, there is no feasible solution. This result holds regardless of the existence of $W$. A similar argument applies to Assumption U$^{0}$ with binary $Y$. Finally, suppose $P[Y(1,w)\ge Y(0,w)|U=u]\ge P[Y(1,w)\le Y(0,w)|U=u]$ for all $w\in\{0,1\}$ and $u$ under Assumption U$^{0}$. Then, we can generate the following inequality restrictions:
Next, consider Assumption MTS. This assumption can be imposed in the form of $R_{2}q\le a_{2}$. To see this, Assumption MTS is equivalent to
for all $(d,w,x)$. As is clear from this expression, Assumption MTS imposes restrictions on the joint distribution of $(\epsilon,U)$.
Finally, consider Assumptions M and C. It is straightforward to incorporate the shape restrictions on the MTR or MTE function. They can be imposed via inequality constraints (ref), namely, in the form of $R_{2}q\le a_{2}$. For implications on the finite-dimensional LP (ref)--(ref), recall that for $q\in\mathcal{Q}_{K}$, the MTR satisfies
According to the property of the Bernstein polynomial, Assumption M implies that $\sum_{e:y(d,w)=1}\theta_{k}^{e,x}$ is weakly increasing in $k$, i.e.,
Assumption C implies that
One can obtain analogous assumptions and their implications in the presence of $W$.
The analogous approach of LP can be applied to the case of continuous outcome variable. We consider the continuous outcome with support $\mathcal{Y}=[0,1]$ without loss of generality.\footnote{Note that $\mathcal{Y}=\mathbb{R}$ is homeomorphic to the open interval $(0,1)$. We use the closure of the latter as $\mathcal{Y}$ for notational convenience later.} As a key component of our LP, we define the following conditional distribution:
where $\epsilon\equiv(Y(0,0),Y(0,1),Y(1,0),Y(1,1))$ and $e\equiv(y(0,0),y(0,1),y(1,0),y(1,1))$ as before, and “$\epsilon\le e$” is understood as an element-wise inequality. First, we show how the data distribution imposes restrictions on $\tilde{q}(e|u,x)$. From the data, we observe \[ \pi(y,d|z,w,x)\equiv\Pr\left[Y\leq y,D=d|Z=z,W=w,X=x\right] \] for all $(y,d,z,w,x)$. Then, for example, consider the case with $d=1$. The conditional distribution can be written as
where the second equality follows by Assumption EX and the inner integral is the shorthand for a multiple integral with respect to the vector $e$.\footnote{It is worth noting that even though we use the joint distribution of $\epsilon\equiv(Y(0,0),Y(0,1),Y(1,0),Y(1,1))$ as the building block, we do not require joint independence between $\epsilon$ and $(Z,W)$ for the derivation. This is because the observed data distribution is a marginal distribution in $Y$, and thus the marginal independence is enough. The same explanation applies to the analogous derivation (ref) in Section (ref). }
Similarly for the target parameters, the MTR function can be expressed as follows. For example, for $D=0$,
We now define the identified set of the target parameters. Let $\tilde{q}(u)\equiv\{\tilde{q}(e|u,x)\}_{e\in\mathcal{E},x\in\mathcal{X}}$ be the vector of $\tilde{q}(e|u,x)$'s. We introduce the class of $\tilde{q}(\cdot)$ to be
\footnotetext{A function $\tilde{q}:[0,1]^{4}\rightarrow[0,1]$ is 4-increasing if, for any 4-box $[a_{1},b_{1}]\times\cdots\times[a_{4},b_{4}]\subseteq[0,1]^{4}$, it satisfies \[ \Delta_{a_{4}}^{b_{4}}\Delta_{a_{3}}^{b_{3}}\Delta_{a_{2}}^{b_{2}}\Delta_{a_{1}}^{b_{1}}\tilde{q}(e)\geq0, \] where $\Delta_{a_{1}}^{b_{1}}\tilde{q}(e)\equiv\tilde{q}(b_{1},e_{2},e_{3},e_{4})-\tilde{q}(a_{1},e_{2},e_{3},e_{4})$, $\Delta_{a_{2}}^{b_{2}}\tilde{q}(e)\equiv\tilde{q}(e_{1},b_{2},e_{3},e_{4})-\tilde{q}(e_{1},a_{2},e_{3},e_{4})$, and so on.}Define the vector of CDFs
and the linear operators $\tilde{R}_{\omega}:\tilde{\mathcal{Q}}\mathcal{\to\mathbb{R}}$ and $\tilde{R}_{0}:\tilde{\mathcal{Q}}\to\mathbb{R}^{d_{\textbf{\ensuremath{\pi}}}}$ (with $d_{\textbf{\ensuremath{\pi}}}$ being the dimension of $\textbf{\ensuremath{\pi}}$) of $\tilde{q}(\cdot)$ that satisfy:
where the expectation is taken over $(W,Z,X)$ and $\text{\ensuremath{\mathcal{U}_{z,x}^{1}=[0,P(z,x)]}}$ and $\mathcal{U}_{z,x}^{0}=(P(z,x),1]$.
Then the $\infty$-LP is formulated as:
subject to
Note that the LP is infinite dimensional not only because of $\tilde{q}$ but also (ref), which consists of a continuum of constraints.
Analogous to Section (ref), we approximate the unknown function $\tilde{q}(\cdot)$ using multivariate Bernstein polynomials: \[ \tilde{q}(e|u,x)\approx\sum_{\boldsymbol{k}=1}^{K}\theta_{\boldsymbol{k}}^{x}b_{\boldsymbol{k}}(e,u), \] where $b_{\boldsymbol{k}}(e,u)\equiv b_{\boldsymbol{k},K}(e,u)$ is a 5-variate Bernstein polynomials with $\boldsymbol{k}\equiv(\boldsymbol{k}_{e},k_{u})$ and $\boldsymbol{k}_{e}\equiv(k_{00},k_{01},k_{10},k_{11})$ and its coefficient $\theta_{\boldsymbol{k}}^{x}\equiv\theta_{\boldsymbol{k},K}^{x}\equiv\tilde{q}(\boldsymbol{k}_{e}/K|k_{u}/K,x)$. Note that “$\sum_{\boldsymbol{k}=1}^{K}$” stands for “$\sum_{k_{00},k_{01},k_{10},k_{11},k_{u}=1}^{K}$.” Then the constraint can be written as a linear combination of the unknown parameters $\left\{ \theta_{\boldsymbol{k}}^{x}\right\} _{(\boldsymbol{k},x)\in\mathcal{K}^{5}\times\mathcal{X}}$. For example,
where $\sigma_{\boldsymbol{k}}^{d}(y,z,x)\equiv\int_{\mathcal{U}_{z,x}^{d}}\int_{\{e:y(1,w)\le y\}}b_{\boldsymbol{k}}(e,u)dedu$. Similarly, the target parameter can be written as, for example,
where $\zeta_{\boldsymbol{k}}^{d}(w,x)\equiv\sum_{z\in\{0,1\}}p(z,w,x)\int\left(\int_{0}^{1}\int_{\{e:y(d,w)\le y\}}b_{\boldsymbol{k}}(e,u)dedy\right)\omega_{d}(u,z,x)du$.
To address the challenge that the constraint (ref) is indexed by continuous $y$, we proceed as follows. Note that, for any measurable function $h:\mathcal{Y}\rightarrow\mathbb{R}$, $E|h(Y)|=0$ if and only if $h(y)=0$ almost everywhere in $\mathcal{Y}$ (beresteanu2011sharp). Therefore, the constraint (with general $d$) can be replaced by: \[ E\left|\sum_{\boldsymbol{k}=1}^{K}\theta_{\boldsymbol{k}}^{x}\sigma_{\boldsymbol{k}}^{d}(Y,z,x)-\pi(Y,d|z,w,x)\right|=0, \] which is now a single constraint given $(d,z,w,x)$. In estimation, the population mean can be replaced with the sample mean; see Section (ref). Now, redefine $\theta\equiv\{\theta_{\boldsymbol{k}}^{x}\}_{(x,\boldsymbol{k})\in\mathcal{X}\times\mathcal{K}^{5}}$ and
where $\boldsymbol{K}_{e}\equiv(K,K,K,K)$.Then, the LP can be formulated as
subject to
This section provides numerical results to illustrate our theoretical framework and to show the role of different identifying assumptions in improving bounds on the target parameters. For target parameters, we consider the ATE and the LATEs for always-takers (LATE-AT), never-takers (LATE-NT), and compliers (LATE-C). We calculate the bounds on them based only on the information from the data and then show how additional assumptions (e.g., the existence of additional exogenous variables, uniformity, and shape restrictions) tighten the bounds.
One important question we want to answer in the exercise is how the current LP approach compares to that in mogstad2018using. This question is theoretically explored in Section (ref), where we showed that our LP approach can capture full independence's stronger identifying power than mean independence. We show that the simulation results are consistent with the theoretical finding.
We generate the observables $(Y,D,Z,X,W)$ from the following data-generating process (DGP). We assume that $W$ is a reverse IV, i.e., we maintain Assumptions EX and SEL(a). We allow covariate $X$ to be endogenous. All the variables are set to be binary with $\Pr\left[Z=1\right]=0.5$, $\Pr\left[X=1\right]=0.6$ and $\Pr\left[W=1\right]=0.4$. The treatment $D$ is determined by $Z$ and $X$ through the threshold crossing model specified in Assumption SEL(a), where the propensity scores $P(z,x)$ are specified as follows: $P(0,0)=0.1$, $P(1,0)=0.4$, $P(0,1)=0.4$, and $P(1,1)=0.7$. The outcome $Y$ is generated from $(D,X,W)$ through $Y=DY_{1}+(1-D)Y_{0}$. For the case of binary $Y$, we generate $Y_{d}$ from
with the MTR functions are defined as
where $b_{k}^{K}$ stands for the $k$-th basis function in the Bernstein approximation of degree $K$. These MTR functions are consistent with Assumptions M and C, i.e., to be weakly monotone and concave in $u$ for all $(d,x,w)\in\left\{ 0,1\right\} ^{3}$. Also, the DGP in (ref) satisfies Assumption U$^{*}$ because $\epsilon$ does not depend on $d=0,1$ and the MTR functions satisfy $m_{1}(u,x,w)>m_{0}(u,x,w)$ for all $(d,x,w)\in\left\{ 0,1\right\} ^{3}$. Therefore, the DGP also satisfies Assumptions U and U$^{0}$. Following the second example in Section (ref), the DGP satisfies the following uniform order for the counterfactual outcomes $Y(d,w)$: $Y(1,1)\ge Y(0,1)\geq Y(1,0)\ge Y(0,0)$ a.s. The case with non-binary $Y$ has DGPs with similar structure, which we omit for succinctness. We generate a sample containing 1,000,000 observations and choose $K=50$. We choose the large sample size to mimic the population. Our choice of $K$ is discussed below. The number of unknown parameters $\theta$ in the linear programming is equal to $\dim(\theta)=\left|\mathcal{E}\right|\times\left|\mathcal{X}\right|\times(K+1)$.
To illustrate the usefulness the current approach in incorporating full independence of the IV, we make comparisons with mogstad2018using. Motivated by the theoretical results in Section (ref), we consider a range of cases for the support $\mathcal{Y}$ of $Y$. Specifically, $Y$ takes values in $\{0,1\}$, $\{0,0.5,1\}$, $\{0,0.25,0.5,0.75,1\}$, and $\{0,0.1,0.2,0.3,0.4,0.5,0.6,0.7,0.8,0.9,1\}$. We additionally consider different supports $\mathcal{Z}$ of $Z$ and allow $Z$ taking values from $\{0,1\}$, $\{0,0.5,1\}$ and $\{0,0.25,0.5,0.75,1\}$. Note that we intentionally fix the endpoints of $\mathcal{Y}$ and $\mathcal{Z}$ to remove the effect of increased variation of the variables. According to Section (ref), it is conceivable that $Y$ departing from binary will deliver narrower bounds in our approach than mogstad2018using's. The gain from $Z$ departing from binary can be more subtle as the endpoints are fixed. For different combinations of $\mathcal{Y}$ and $\mathcal{Z}$, we derive the bounds on the ATE.
Figure (ref) compares bounds that are calculated by using the current approach with those that are replicated by using mogstad2018using's approach. The first subfigure on the top-left shows that the bounds are identical between the two methods, because $Y$ is binary and thus full independence between $Y(d)$ and $Z$ is equivalent to mean independence. This equivalence is maintained no matter how many values $Z$ takes. In the next three subfigures, our bounds are tighter than mogstad2018using's as we expect. They show a pattern that, as $\mathcal{Y}$ contains more values, the improvement from the current approach is more substantial. This is because, as $Y$ takes more values, with greater degree, full independence contains richer structure than mean independence. Again, this pattern remains to hold regardless of the choice of $\mathcal{Z}$. It would also be interesting to investigate (i) the case with continuous $Y$ and (ii) the impact of $\mathcal{Z}$ when we move its endpoints further apart. They are explored in Section (ref) of the Appendix.
Focusing on binary $Y$, Table (ref) contains the bounds on the ATE under different assumptions, and these bounds are illustrated in Figure (ref) and (ref). The true ATE value is $0.15$, depicted as the solid red line in the figure. From (ref), the worst-case bounds on the ATE with no additional assumptions (and without using variation from $W$) are $[-0.25,0.45]$. Since the mappings do not involve $W$, we have $|\mathcal{E}|=4$, and the linear programming is solved with $\dim(\theta)=\left|\mathcal{E}\right|\times\left|\mathcal{X}\right|\times(K+1)=4\times2\times51=408$.
For comparison, we calculate the bounds that incorporate the existence of $W$. We express the target parameters with mappings involving $W$ and use data distribution conditional on $W=0$ and $W=1$ as the constraints. With binary $W$, we have $|\mathcal{E}|=16$, which gives $\dim(\theta)=\left|\mathcal{E}\right|\times\left|\mathcal{X}\right|\times(K+1)=16\times2\times51=1,632$. The resulting bounds are depicted in the dotted greenish-blue line. When the variation from $W$ is used, the bounds on the ATE are $[-0.21,0.42]$, which is narrower than without using $W$. This result is consistent with our theoretical finding presented in Theorem (ref) that $W$ can help tighten the bounds as long as it is a relevant variable. Nonetheless, these worst-case bounds are not that informative, e.g., they do not determine the sign of the ATE.
Next, we impose Assumption U$^{0}$ without $W$ and with $W$.\footnote{Assumption U and U$^{0}$ give the same bounds in our exercise, therefore, we use the weaker assumption and present the results.} Under Assumption U$^{0}$, the bounds on the ATE are tightened as we incorporate extra inequality constraints according to the direction of monotonicity. As mentioned in Section (ref), the direction of monotonicity in Assumption U$^{0}$ is determined by the LPs. We solve the LPs with different directions imposed, then choose the one with a feasible solution. This means that the corresponding direction of monotonicity is consistent with the DGP. Under Assumption U$^{0}$, we obtain a bound $[0.05,0.45]$, which is narrower comparing with the worst-case bound. With $W$, under Assumption U$^{0}$, the bounds become $[0.05,0.42]$. In Figure (ref), these bounds under Assumptions U$^{0}$ without and with $W$ are depicted as violet and green dashed lines, respectively. Both sets of bounds identify the sign of the ATE, consistent with the theoretical discussion. The improvement is mainly on the lower bounds and the upper bounds coincide with the corresponding worst-case upper bound without and with $W$. These improvements come from the ability to identify the sign under the uniformity assumptions.
Next, we impose the shape restrictions (Assumptions M and C). As discussed in Section (ref), these assumptions can be easily incorporated in the linear programming by directly imposing inequality constraints on $\theta$. Under these assumptions (and the existence of $W$), the bounds on the ATE shrink to $[0.12,0.19]$, which is displayed with the pink line in Figure (ref). We find that shape restrictions are powerful assumptions and yield narrower bounds compared to those with uniformity assumptions. They function differently in the linear programming: unlike the uniformity assumption, which maintains the ranking of individuals across counterfactual groups, shape restrictions directly control the MTR functions.
Figure (ref) presents the results under Assumption U$^{0}$, versus under Assumption U$^{*}$ with existence of $W$. Under Assumption U$^{*}$, the bounds become $[0.05,0.38]$. While their lower bounds coincide, Assumption U$^{*}$ yields a lower upper bound compared to Assumption U$^{0}$.
Next, we construct bounds on the generalized LATEs. Again, we focus on binary $Y$. The original definition of the LATE is the ATE for compliers (C). Researchers may also have interests in other local treatment effects. We consider two other parameters---LATEs for always-takers (AT) and never-takers (NT). Figure (ref) and (ref) display the bounds on the LATE-AT, LATE-C, and LATE-NT under different assumptions. This analysis is analogous to that with the ATE. Since the covariate $X$ affects the decision of compliance, to avoid confusion in the definition of the compliance groups, we instead establish bounds on the LATEs conditional on $X$. We draw the conditional MTE functions with solid red lines in both panels as a reference.
The DGP implies constant MTE function, therefore, the LATE-AT, LATE-C and LATE-NT are all equivalent to true ATE, equaling to $0.33,0.23,0.13$ and $0.20,0.11,0.08$, conditional on $X=0$ and $X=1$ respectively. The feature that there exists no defiers in the DGP is known. When there is no defier, the LATE-C is point identified, which has an analytical expression of the two-stage least squares estimand. Therefore, even when we add the tuning parameters\footnote{The tuning parameters are used to prevent infeasibility in the LP due to the sampling error. The details are introduced in Section (ref).}, the estimates remain very close to the true values throughout. And when we do not need tuning parameters to adjust the numerical errors or when the tuning parameters are very small, the linear programming yields point estimates as shown in Figure (ref).
For the LATE-AT and LATE-NT, as before, we first consider the worst-case bounds where the existence of $W$ is ignored versus where $W$ is taken into account. Without $W$, we get the bounds $[-0.60,0.40]$ and $[-0.28,0.73]$ on the LATE-AT and the LATE-NT conditional on $X=0$, and $[-0.58,0.43]$ and $[-0.41,0.60]$ conditional on $X=1$; with $W$, we get the bounds $[-0.38,0.39]$ and $[-0.26,0.67]$ on the LATE-AT and the LATE-NT conditional on $X=0$, and $[-0.51,0.42]$ and $[-0.34,0.56]$ conditional on $X=1$. Incorporating information from $W$ helps improve both the upper and lower bounds. Imposing Assumption U$^{0}$ without $W$ helps to identify the sign by raising up the lower bound to $0$, for the LATEs; when considering the case with $W$, the pattern remains the same, but with a lower upper bound and a slightly improved lower bound above $0$. We then apply M and C with $W$ taking into account. The bounds on the LATE-AT and the LATE-NT turn to $[0.27,0.36]$ and $[0.11,0.19]$ conditional on $X=0$, and $[0.10,0.29]$ and $[0.04,0.11]$ conditional on $X=1$.
From Figure (ref), under the Assumption U$^{*}$, the bounds shrink to $[0,0.37]$ and $[0,0.53]$ conditional on $X=0$, and $[0,0.40]$ and $[0,0.55]$ conditional on $X=1$, comparing with the bounds under Assumption U$^{0}$. The improvement is from complete order of 16 mapping types we have in this environment and is most significant for the never-taker LATE upper bound.
As a tuning parameter in the LP, we need to choose the order of Bernstein polynomials, $K$. In general, $K$ should be chosen based on the sample size and the smoothness of the function to be approximated, in our case, $q(\cdot)$. The choice of the sieve dimension or more generally, regularization parameters, is a difficult question (chen2007large) and developing data-driven procedure is a subject of on-going research in various nonparametric contexts of point identification; see, e.g., chen2018optimal and han2020nonparametric. In this partial identification setup, we propose the following heuristic and conservative approach, which is in spirit consistent with the very motivation of partial identification.
First, we do not want to claim any prior knowledge about the smoothness of $q(\cdot)$ because it is the distribution of a latent variable. Because $K$ determines the dimension of unknown parameter $\theta$ in the linear programming, the width of the bounds tends to increase with $K$. At the same time, the computational burden increases with $K$. One interesting numerical finding is that, when $K$ is sufficiently large, the increase of the width slows down and the bounds become stable. This suggests that we may be able to conservatively choose $K$ that acknowledges our lack of knowledge of the smoothness but, at the same time, produces a reasonable computational task for the linear programming.
To illustrate this point, we first consider the conditional MTE as the target parameter and show how its bounds change as $K$ increases. We consider the MTE because it is a fundamental parameter that generates other target parameters, and hence, it is important to understand the sensitivity of its bounds to $K$. The left panel of Figure (ref) shows the evolution of the bounds on the MTE as $K$ grows. We use a DGP similar to the one described earlier. When $K=5$, the bounds are narrow. Although it may be tempting to choose this value of $K$, this attempt should be avoided as it may be subject to the misspecification of the true smoothness. When $K$ increases beyond $30$, the bounds start to converge and become stable. We choose $K=50$, and this is the choice we made in our previous numerical exercises.
To compare this converging pattern with a known benchmark, in the right panel of Figure (ref), we depict the identified set for the ATE relative to Manski's analytical bounds (manski1990nonparametric). We observe the identified set approaches to Manski's bound as $K$ increases and it almost overlaps with Manski's bounds when $K$ is around 50.\footnote{Note that with large $K$, some LP solvers would ignore coefficients with negligible (e.g., $10^{-13}$) values that cause a large range of magnitude in the coefficient matrix. It may be recommended to simultaneously rescale a column and a row to achieve a smaller range in the coefficients; see Section (ref) for details. We found that when $K=50$, the bounds from the rescaled LP and original LP are very close to each other (e.g., for the ATE bounds without extra assumptions, the difference is up to 0.01).}
As discussed in Section (ref) in the Appendix, it is worth mentioning that the bounds on the MTE are pointwise sharp but not uniformly sharp. The graph for the MTE bounds are drawn by calculating the pointwise sharp bounds on MTE at each point of $u$ (after properly discretizing it) and then connecting them. Therefore, these bounds should not be viewed as uniformly sharp bounds. Nonetheless, this graph is still useful for the purpose of our illustration. Given the current DGP, we find that there are no uniformly sharp bounds for the MTE.
It is widely recognized in the empirical literature that health insurance coverage can be an essential factor for the utilization of medical services (hurd1997medical; dunlop2002gender; finkelstein2012oregon,taubman2014medicaid). Prior studies on this topic typically make use of parametric econometric models for the analysis. In their application, han2019estimation relax this common approach by introducing a semiparametric bivariate probit model to measure the average effect of insurance coverage on patients' medical visits. By applying our theoretical framework of partial identification, we further relax the parametric and semiparametric structures used in these studies. More importantly, we try to understand how much we can learn about the effect of insurance that is utilized through various counterfactual policies by learning the effect of different compliance groups.
We use the 2010 wave of the Medical Expenditure Panel Survey (MEPS) and focus on all the medical visits in January 2010. The sample is restricted to contain individuals aged between 25 and 64 and exclude those who had any kind of federal or state insurance in 2010. The outcome $Y$ is a binary variable indicating whether or not an individual has visited a doctor's office; the treatment $D$ is whether an individual has private insurance. We choose whether a firm has multiple locations as the binary instrument $Z$. This IV reflects the size of the firm, and larger firms are more likely to provide fringe benefits, including health insurance. On the other hand, the number of branches of a firm does not directly affect employee decisions about medical visits. To justify the IV, self-employed individuals are excluded. For potentially endogenous covariates $X$, we include the age being 45 and older, gender, income above median. Lastly, for an exogenous covariate $W$, we use the percentage of workers who are provided with paid sick leave benefits within each industry. Following han2019estimation, we assume $W$ satisfies Assumptions SEL$_{W}$(b) and EX$_{W}$(b), as $X$ is controlled. The rationale is the following: First, we assume $W$ is exogenous (conditional on covariates) arguably because it is determined by the employer or is in accordance with the local legislation and thus is not correlated with individual employee's preferences. However, due to its nature, it can influence the employee's health-related decisions, such as enrolling in an insurance program ($D$) or utilizing medical services ($Y$).\footnote{Since the relevance of $W$ to $D$ is slightly less plausible than to $Y$, we test whether the propensity score is a not a function of $W$ but cannot reject the null. Therefore, we decided to use $W$ as a common exogenous variable than a reverse IV.} We construct a categorical variable such that $W=0$ for less than median value of the pay sick leave provision, $W=1$ for above the median.
First, as a benchmark, we report that the LATE-C estimate calculated via our linear programming approach is equal to a singleton of 0.05, which is in fact identical to the 2SLS estimate we separately calculate. In what follows, we extrapolate this LATE beyond the complier group to the ATE. The presence of covariates reduces the effective sample size and thus leads to larger sampling errors in estimating the $p$ of the $\infty$-LP (ref)--(ref). This may create inconsistencies in the set of equality constraints (ref), resulting in no feasible solution. This is in fact what happens in this application. To resolve this estimation problem, we introduce a slackness parameter $\kappa$ and modify (ref) so that, with some slackness, it satisfies
where $\hat{R}_{0}$ and $\hat{p}$ are estimates of $R_{0}$ and $p$. A similarly modified constraint can then be followed in the finite-dimensional LP after approximation, as well as by combining (ref)--(ref). The appropriate value of $\kappa$ should depend on the sample size, the dimension of covariates, and the dimension of the unknown parameter $\theta$. To explain the latter, as $K$ increases, the dimension of $\theta$ (i.e., unknowns) increases, while the number of constraints (i.e., simultaneous equations for the unknowns) is fixed. Therefore, as $K$ increases, the chance that the LP does not have a feasible solution would decrease. Based on the method discussed in the previous section, we set $K=50$ in this application.
We calculate worst-case bounds on the ATE, as well as bounds after imposing Assumptions U$^{0}$ and M and after using $W$. Under Assumption U$^{0}$, the data rules out the possibility that $Y(0)>Y(1)$, indicating that individuals with private insurance are more likely to visit a doctor. Assumption M imposes that the MTR function is weakly increasing in $U=u$. Usually, $U$ is interpreted as the latent cost of obtaining treatment. kowalski2021reconciling interpreted $U$ as eligibility in a similar setup for Medicaid insurance. The eligibility for Medicaid is related to income level and age. In our setup, because the treatment is having the private insurance, we interpret the eligibility as the health status, which is reflected in the premium. Interpreting $U$ as a latent cost (e.g., premium) of getting private insurance, Assumption M states that the chance of making a medical visit (with or without insurance) increases for those with higher cost. This is a reasonable assumption given that sicker individuals typically face higher insurance costs and also visit doctors more often. We choose the slackness parameter $\kappa$ to be consistently 0.01 under all assumptions for a comparable comparison.
The bounds on the ATE are shown in Figure (ref). The worst-case bound on the ATE equals $[-0.42,0.36]$. The bounds become $[0.02,0.36]$ under Assumption U$^{0}$ and $[0.07,0.36]$ under Assumption M. It is interesting to note that the identifying power of the uniformity and the shape restriction is similar in this example. When both Assumption U$^{0}$ and Assumption M are imposed, the bounds are further tightened to $[0.08,0.36]$, although not substantially, indicating that the two assumptions are complementary. However, we do not see gain from incorporating $W$ in this case.
Next, we consider the always-taker, complier, and never-taker LATEs. We consider these generalized LATEs conditional on $X=x$. Specifically, we focus on the treatment effects for males above age 45, with income below the median. The results are shown in Table (ref) and depicted in Figure (ref). The LATE-C is analytically calculated via TSLS.\footnote{When the alternative constraint (ref) is used with the slackness parameter, the LATE-C is no longer a singleton.} For the LATE-AT and LATE-NT, Assumption U$^{0}$ identifies the sign of the effects, and Assumption M nearly identifies it. Using the variation in $W$ mostly improves the bounds compared to the ones without it.\footnote{Most of the extra assumptions we impose help to determine the direction of treatment effect, i.e., to raise the lower bound if the treatment effect is positive. Therefore, improvements on LATE-NT are smaller than LATE-AT after imposing extra assumptions, since the evidence of positive treatment effect is relatively strong even with the worst-case bounds of LATE-NT. } From the results we can conclude that, for a range of identifying assumptions, the private insurance tends to have a large effect on medical visits for never-takers, that is, people who face higher insurance cost. For example, for all the cases, the upper bound on the effect (i.e., the most optimistic scenario) is much larger for the never-takers than the always-takers. Under Assumption W or no assumption, the lower bounds (i.e., the most pessimistic scenario) shows a similar pattern. This suggests a policy implication that lowering the cost of private insurance may be important, because high costs may hinder those with the most need from receiving enough medical services.