Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
114,263 characters · 20 sections · 85 citation commands
Policy Relevant Treatment Effects with Multidimensional Unobserved Heterogeneity
Policy relevant treatment effects (PRTEs) have emerged as pivotal target parameters in both empirical and theoretical studies. Instrumental variables (IVs) are commonly used to identify PRTEs with the endogenous treatment choice. In the threshold-crossing model for the treatment with one-dimensional unobserved heterogeneity, heckman/vytlacil:1999,heckman/vytlacil:2001,heckman/vytlacil:2005 study the identification of various PRTEs as the weighted average of the marginal treatment effects. In the same threshold-crossing model, \citet*{mogstad/santos/torgovitsky:2017} propose a linear programming method to construct a bound on them.
The threshold-crossing assumption for the treatment can be restrictive in many empirical scenarios. As highlighted by heckman/vytlacil:2005, it requires “uniformity” of treatment selection across all individuals and restricts the heterogeneity of the treatment selection. mogstad2020policy also demonstrates that this condition becomes particularly strong when multiple IVs provide distinct incentives that lead to heterogeneous treatment selection. In fact, a number of recent papers have proposed specification tests for the threshold-crossing structure together with the instrument validity.\footnote{See huber/mellace:2015,kitagawa_2015,laffers2017note,mourifie2017testing,kedagni2020generalized,SUN2023105523,carr2023testing among others.} Although unobserved heterogeneity plays a crucial role in individual choice models and causal inference studies heckman2001micro,imbens2014instrumental,mogstad2024instrumental, to our knowledge, it is still an open question whether and how the framework of mogstad/santos/torgovitsky:2017 can be extended to bound PRTEs with multidimensional unobserved heterogeneity affecting both treatment and outcome.
This paper aims to address this question when we have a binary endogenous treatment with IVs. We offer a systematic approach to bound various PRTEs without imposing any threshold-crossing structure on the treatment selection process. Our method explicitly models the unobserved heterogeneity by a multidimensional random vector $V$, and we assume that treatment and potential outcomes are mean-independent once conditioned on $V$. This conditional mean-independence assumption then implies that both the target PRTEs and identifiable estimands are weighted averages of bilinear functions of the unknown propensity score and marginal treatment response functions.\footnote{yap2022sensitivity also derives a similar expression for PRTEs when investigating violations of the threshold-crossing structure.}
The unknown propensity score and marginal treatment response functions appear bilinearly in both the target parameter and the identifiable estimands. Consequently, the PRTE bounds can only be formulated as solutions to nonconvex optimization problems with nonlinear (in particular, bilinear) constraints. We propose to use the convex relaxation method of mccormick1976computability, where we replace each bilinear function in the optimization problems with a single function and characterize its feasible region using McCormick Envelopes. The resulting bounds are obtained through linear problems.
The convex relaxation method has several attractive features. First, it provides conservative yet computationally attractive PRTE bounds, as the bounds can be easily computed using well-established linear programming techniques. Second, while the convex relaxation produces an outer set of the PRTE sharp identifed set, we demonstrate that this outer set remains informative. Specifically, we provide sufficient conditions under which the outer set leads to no identification power loss. Third, it offers a unified framework for incorporating linear shape restrictions or combinations of these shape restrictions—some of which may not have been explored in the literature—to further tighten the bounds. This allows researchers to avoid deriving PRTE bounds under different shape restrictions on a case-by-case basis. Lastly, we can conduct inference by directly applying the regularized support function method introduced by gafarov2019inference. This method yields closed-form asymptotic Gaussian estimators for PRTE bounds, so that uniformly valid confidence sets can be constructed using the standard normal critical values, which substantially reduces the computational burden.
It is worthwhile dicussing further the second feature in the above paragraph. For a large class of target parameters, our bounds are simplified to the identified set of mogstad/santos/torgovitsky:2017 when the treatment selection happens to satisfy the threshold-crossing assumption with one-dimensional unobserved heterogeneity. This result is formally stated in Theorem (ref) in Section (ref) and has a practical takeaway. In practice, the researcher may not know whether the treatment equation satisfies the threshold-crossing assumption or not, and our bounds do not require this knowledge. Even if the treatment equation does not satisfy this assumption, our bounds are still valid, whereas the ones of mogstad/santos/torgovitsky:2017 are not (e.g., their bounds can be empty in some cases). In this sense, we extend mogstad/santos/torgovitsky:2017 to a more general setup with multidimensional unobserved heterogeneity and provide a more robust alternative to practitioners.
Numerical analysis and simulation exercises are used to verify the identification results for various PRTEs and asymptotic properties of the inference method. Importantly, numerical results demonstrate that the PRTE bounds of mogstad/santos/torgovitsky:2017 can be empty if their key assumption (i.e., one-dimensional unobserved heterogeneity in the treatment selection process) is violated. Moreover, the conservativeness introduced by the convex relaxation method does not lead to information loss in some cases. For example, without shape restriction, our PRTE outer set for the average treatment effect (ATE) is the same as its sharp bounds of manski1990nonparametric.
The rest of this paper is organized as follows. Section (ref) discusses related literature. Section (ref) introduces our model setup, target parameters, and some commonly used shape restrictions. Section (ref) presents partial identification results using convex relaxation method and investigates conditions under which our convex-relaxation bounds have no identification power loss. Section (ref) outlines the inference procedure in gafarov2019inference for constructing uniformly valid PRTE confidence sets. Section (ref) provides numerical and simulation results. Section (ref) concludes.
A growing body of work has focused on relaxing the threshold-crossing structure with one-dimensional unobserved heterogeneity to accommodate more flexible forms of unobserved heterogeneity. This paper differs from all previous studies because our framework simultaneously incorporates the following features. First, we begin with a general model setup without imposing any threshold-crossing restrictions on the treatment selection, and our target parameters include a large class of PRTEs. Second, we explicitly model the unobserved heterogeneity using a multi-dimensional random vector. Our identification analysis relies on the conditional mean-independence assumption between the potential outcomes and the treatment variable, given the unobserved heterogeneity, which results in a bilinear structure for our target PRTE parameters. Third and most importantly, to overcome the computational challenge posed by the bilinear structure, we propose computationally attractive and informative bounds obtained by using the convex relaxation method. As a foundational result in the causal effect literature, vytlacil:2002 establishes the equivalence between the classical monotonicity condition of imbens/angrist:1994 and the threshold-crossing model with single-dimensional unobserved heterogeneity. In this section of related literature, we cover the relevant papers which relax one of these two conditions.
This paper is closely related to yap2022sensitivity, which also derives a similar bilinear expression for PRTEs when investigating violations of the threshold-crossing structure for the treatment with one-dimensional unobserved heterogeneity. The author uses unobserved compliance types to capture unobserved heterogeneity and a sensitivity parameter to measure the violations of monotonicity. The PRTE bounds are then derived across different values of the sensitivity parameter, and computed by breaking a nonconvex optimization problem into an inner and an outer linear program. This paper differs from yap2022sensitivity for several reasons. First, our general model setup does not rely on a sensitivity parameter about the violations of monotonicity. Second, by explicitly modeling the unobserved heterogeneity, we can easily impose shape restrictions on the unknown but interpretable functions of the propensity score and MTRs, in the same spirit as heckman/vytlacil:1999,heckman/vytlacil:2005 and brinch/mogstad/wiswall:2017. Third, we address the nonconvex problem using convex relaxation, which leads to computationally attractive bounds and easy-to-implement inference.
This paper also relates to studies that explore various relaxations of the threshold-crossing structure for the treatment with one-dimensional unobserved heterogeneity. klein2010heterogeneous delves into a selection model with local violations of the treshold-crossing structure. dinardo/lee:2011 and small2017instrumental impose the stochastic monotonicity on the propensity score across all strata of unobserved heterogeneity. In cases with multiple IVs, mogstad2020policy propose a partial monotonicity condition for each IV separately. In a related vein, goff2020vector introduces a vector monotonicity condition that is stronger than partial monotonicity. Other studies that explore various weaker monotonicity conditions include de2017tolerating, sloczynski2021should, noack2021sensitivity, dahl2023never, van2023limited, and hoff2023identifying.
Another body of work develops less restrictive models capable of accommodating multidimensional unobserved heterogeneity. For instance, lee2018identifying examine the treatment assignment determined by a vector of threshold-crossing rules. gautier2011triangular and gautier2021relaxing model the unobserved heterogeneity using random coefficient models. han2024set propose a generalization of control function approach and use a non-monotone treatment assignment as a motivating example. navjeevan2023identification establish the point identification of target parameters, where they represent the unobserved heterogeneity using potential outcomes and an unobserved type. Notably, such a point identification result requires the target parameter to be an expectation of a function of unobservables that can be further expressed as a linear function of observables, which may not hold for some commonly studied PRTEs.\footnote{One counterexample that does not satisfy this condition is the ATE in a heterogeneous treatment effect model with an endogenous binary treatment. We know that ATE is an expectation of the difference between two potential outcomes. It is generally impossible to express such a difference as a linear map of functions of observables, and the ATE point identification is also not achievable without additional assumptions.}
There are also studies that consider bounding PRTEs while leaving the treatment selection completely unspecified. Some studies focus on treatment effects that have analytical bounds manski1990nonparametric,manski1997monotone,manski2000monotone,huber2017sharp. Others consider treatment effects whose analytical bounds do not exist or are difficult to derive laffers2019bounding,russell2021sharp.
Consider a canonical program evaluation problem with a binary treatment variable $D\in\{0,1\}$ and a scalar outcome variable $Y\in[y^L,y^U]$.\footnote{When there is no restriction on the support of $Y$, we can set $y^L=-\infty$ and $y^U=\infty$.} Define $$ Y=DY_1+(1-D)Y_0, $$ where $Y_1$ and $Y_0$ are the two potential outcomes. We assume there exists a vector of observable variables, $Z=(Z_0,X)\in\mathcal{Z}$, that includes a vector of instruments $Z_0\in\mathcal{Z}_0$ and a vector of covariates $X\in\mathcal{X}$. The instruments can be discrete, continuous, or mixed. We introduce a random vector ${V}\in{\mathcal{V}}$ to represent the multidimensional unobserved heterogeneity that affects both the treatment $D$ and the outcome $Y$, so that the individual treatment effect $Y_1-Y_0$ can be heterogeneous depending on the values of the unobserved heterogeneity $V$. For a generic probability distribution $P$, we let $E_{P}$ be its expectation operator. We use $\mathbb{P}$ to denote the true distribution for $(Y_0,Y_1,D,Z,V)$ and use $\mathbb{E}=E_{\mathbb{P}}$ as its expectation operator. We maintain the following key assumptions throughout the paper.
Assumption (ref) (i) implies that the treatment endogeneity arises from the unobserved heterogeneity $V$. Consequently, given $(X,{V})$, the treatment and instruments become mean-independent of the potential outcomes. Condition (ii) assumes that the unobserved heterogeneity $V$ is independent of $Z$, and it also introduces a standard normalization to the distribution of $V$.\footnote{ If $V$ is independent of $Z_0$ given $X$ and is continuously distributed given $X$, we can always normalize its distribution to a uniform distribution.} This assumption is weaker than the one imposed in mogstad/santos/torgovitsky:2017. We allow $V$ to be multidimensional (including $K=1$ as a special case), whereas mogstad/santos/torgovitsky:2017 assume $V$ to be one-dimensional.\footnote{Note that the researcher may not know the smallest value of $K$ for which Assumption (ref) holds, but may know its upper bound. We can use this upper bound as $K$ in Assumption (ref). For example, if Assumption (ref) holds with $K=1$, then it holds with $K=2$.}
Let $F$ denote the cumulative distribution function. The following lemma shows conditions under which Assumption (ref) holds with $K=2$.
Lemma (ref) shows that if the potential outcomes are independent of $Z$ given $X$, we can always use the two normalized potential outcomes, $(F_{Y_0\mid X}(Y_0),F_{Y_1\mid Y_0,X}(Y_1))$, to represent the multidimensional $V$. Nonetheless, this particular choice of $V$ only provides one example for the unobserved heterogeneity with $K=2$, and there can be other choices of $V$ that also satisfy Assumption (ref). Note that our model setup is general enough to accommodate various treatment selection models in the literature that assume multidimensional unobserved heterogeneity, including those with random coefficients gautier2011triangular, local violations to monotonicity klein2010heterogeneous, stochastic compliance small2017instrumental, and the double hurdle model where treatment is assigned if and only if two criteria are satisfied simultaneously lee2018identifying. This is because all these models mentioned above assume the same independence condition between potential outcomes and instruments (conditional on covariates, if any) as in our Lemma (ref).
Define the marginal treatment response function $m_{\mathbb{P},d}(v,x)$ with $d=0,1$ and the propensity score $m_{\mathbb{P},D}(v,z)$, conditioning on the unobserved heterogeneity $V=v$, as below: $$ m_{\mathbb{P},d}(v,x)={\mathbb{E}}[Y_d\mid V=v,X=x],\text{ and }~ m_{\mathbb{P},D}(v,z)={\mathbb{E}}[D\mid V=v,Z=z]. $$ These three functions, $(m_{\mathbb{P},0},m_{\mathbb{P},1},m_{\mathbb{P},D})$, are the key ingredients for various policy relevant treatment effect (PRTE) parameters. We are going to consider the target parameters of the form
where $\omega_{dd'}^\star$ with $d,d'\in\{0,1\}$ denotes a known or identifiable weight.\footnote{ We consider a general setup that accommodates a large class of PRTE parameters whose known or identifiable weights, $\omega_{dd'}^\star(v,z)$, are allowed to be functions of $v\in\mathcal{V}$, except for Proposition (ref) and Theorem (ref). We assume that the dimension of $V$ is known. For researchers interested only in PRTEs with weights that do not depend on $v$, however, we can remove the assumption of a known dimension of $V$, as we will discuss in more detail in Proposition (ref) in later sections.} Table (ref) below collects some important examples of such PRTE parameters and provides expressions for their weights. For example, if our target parameter is the average treatment effect, $\mathbb{E}[Y_1-Y_0]$, then $\omega_{00}^\star(v,z)=\omega_{01}^\star(v,z)=-1$ and $\omega_{10}^\star(v,z)=\omega_{11}^\star(v,z)=1$.
We can see that $(m_{\mathbb{P},0},m_{\mathbb{P},1},m_{\mathbb{P},D})$ appears in bilinear forms in the expression of the target PRTEs given in (ref). Therefore, for the rest of the paper, define
Then, we can replace the bilinear functions in (ref) using $m_{\mathbb{P},0D}(v,z)$ and $m_{\mathbb{P},1D}(v,z)$. Denote $$m_\mathbb{P}=(m_{\mathbb{P},0},m_{\mathbb{P},1},m_{\mathbb{P},D},m_{\mathbb{P},0D},m_{\mathbb{P},1D})'\in\mathcal{M},$$ where $\mathcal{M}\subseteq (\mathbf{L}^2(\mathcal{V}\times\mathcal{X},[y^L,y^U]))^2\times\mathbf{L}^2(\mathcal{V}\times\mathcal{Z},[0,1])\times(\mathbf{L}^2(\mathcal{V}\times\mathcal{Z},[y^L,y^U]))^2$ is the parameter space for $m_\mathbb{P}$.\footnote{We define $\mathbf{L}^2(\mathcal{A},\mathcal{B})$ to be the set of all the $L^2$ functions from $\mathcal{A}$ to $\mathcal{B}$.} We define $\Gamma^\star$ to be the linear map from $\mathcal{M}$ to $\mathbb{R}$ with
for a generic function $m=(m_{0},m_{1},m_{D},m_{0D},m_{1D})'\in\mathcal{M}$. Given $\Gamma^\star$, we can express the target parameter in (ref) as $\Gamma^\star(m_{\mathbb{P}})$, which is a linear functional of $m_{\mathbb{P}}$. The functional $\Gamma^\star$ is identifiable from the distribution of $Z$ and thus is known to the researchers. In this sense, our partial identification analysis focuses on how we can bound the unknown function $m_{\mathbb{P}}$. Note that under the threshold-crossing structure for the treatment with one-dimensional unobserved heterogeneity, heckman/vytlacil:2005,heckman2007econometric express the target parameters in Table (ref) as linear functionals of the unknown functions $m_{\mathbb{P},0}$ and $m_{\mathbb{P},1}$ only. This is because $m_{\mathbb{P},D}$ is identified by the propensity score $\mathbb{P}(D=1|Z)$, and it is also the main difference between our model setup and those in the existing literature studying PRTEs under the Heckman-Vytlacil framework.
Shape restrictions on the unknown function $m_{\mathbb{P}}$ have often been employed in the literature to tighten the PRTE bounds. In this section, we discuss a few commonly used shape restrictions.
We begin by examining restrictions on the treatment variable and their implied restrictions on $m_{\mathbb{P},D}$.
Condition (ref), introduced by dinardo/lee:2011, is employed by small2017instrumental to study the identification of the Wald estimand. It imposes a monotone restriction on the propensity score at all strata of the unobserved heterogeneity $V$, requiring $m_{\mathbb{P},D}(v,z)=\mathbb{P}(D=1\mid V,Z_0=z_0,X)$ to be weakly increasing in $z_0$.
Condition (ref), introduced by imbens/angrist:1994, is a widely adopted monotonicity condition and is stronger than stochastic monotonicity.\footnote{In dinardo/lee:2011, they refer to Condition (ref) as “deterministic monotonicity” because it asserts that the treatment status of an individual with a given value of the unobserved heterogeneity $V$ is a deterministic function of $Z$.} It requires all individuals to respond to the same shift in the instrument value in the same direction. Specifically, in the case of a binary IV, the classical monotonicity rules out “defiers” and is the key to point identify the local average treatment effect. vytlacil:2002 shows that imposing the classical monotonicity is equivalent to assuming that the treatment variable is determined by a threshold-crossing model with one-dimensional unobserved heterogeneity. In other words, under Condition (ref), there exists a one-dimensional random variable $V$ such that $D=1\{\mathbb{P}(D=1\mid Z)\geq V\}\mbox{ almost surely}$, where $V$ given $Z$ is uniformly distributed over $[0,1]$, and we can set $m_{\mathbb{P},D}(V,Z)=1\{\mathbb{P}(D=1\mid Z)\geq V\}$.
Next, we consider restrictions on the marginal treatment response functions $(m_{\mathbb{P},0},m_{\mathbb{P},1})$ implied by variants of the monotone treatment response condition.
Condition (ref) is the monotone treatment response condition introduced by manski1997monotone and manski2000monotone. It implies $\mathbb{E}[Y_0|Z=z] \leq \mathbb{E}[Y|Z=z] \leq \mathbb{E}[Y_1|Z=z] \text{ for all }z=(z_0,x)\in\mathcal{Z}, $ which further implies restrictions on $(m_{\mathbb{P},0},m_{\mathbb{P},1})$
Conditions (ref) and (ref) only require monotone treatment response at the mean value and are weaker versions of Condition (ref). Specifically, Condition (ref) implies restrictions on $(m_{\mathbb{P},0},m_{\mathbb{P},1})$ that $$m_{\mathbb{P},1}(v,x)-m_{\mathbb{P},0}(v,x)=\mathbb{E}[Y_1-Y_0\mid V=v,X=x]\geq 0,$$ allowing for heterogeneous treatment responses across individuals for some random reasons.
For Condition (ref), since $\mathbb{E}[Y_1-Y_0\mid X=x]=\mathbb{E}\{\mathbb{E}[Y_1-Y_0\mid V,X=x]\mid X=x\}=\mathbb{E}[m_{\mathbb{P},1}(V,x)-m_{\mathbb{P},0}(V,x)]$, it imposes the sign restriction on $(m_{\mathbb{P},0},m_{\mathbb{P},1})$ that $$\int_{{\mathcal{V}}}(m_{\mathbb{P},1}(v,x)-m_{\mathbb{P},0}(v,x))dv\geq0,$$ allowing for heterogeneous treatment responses across individuals with different values of $V$ for any given $X=x$. A similar condition is explored by carlos2019average to study the ATE bounds.
In this section, we discuss shape restrictions which impose joint restrictions on $(m_{\mathbb{P},0},m_{\mathbb{P},1},m_{\mathbb{P},0D},m_{\mathbb{P},1D})$. manski2000monotone consider a monotone treatment selection assumption, which states that the mean value of the potential outcome is weakly increasing in the treatment.
Given Assumption (ref), Condition (ref) implies the following two inequalities\footnote{See detailed proofs in Appendix (ref).}
Another shape restriction is the monotone selection on the gains, which states that the individuals who self-select into the treatment would gain more from being treated compared to those who remain untreated.
Under Assumption (ref), similar arguments used to show the inequalities in (ref) can be applied to show that Condition (ref) implies the following restriction:
In this section, we investigate the identification of the target PRTE $\Gamma^\star(m_{\mathbb{P}})$. Recall that the functional $\Gamma^\star$ is identifiable, and we only need to bound the unknown function $m_{\mathbb{P}}$. To achieve this goal, we first derive constraints that characterize the identified set for $m_{\mathbb{P}}$. Next, we apply a convex relaxation method to obtain computationally feasible PRTE bounds, which we will refer to as the convex-relaxation bounds. Then, we provide sufficient conditions under which the convex-relaxation bounds result in no identification power loss.
We begin by deriving a set of conditional moment equalities that are satisfied by $m_{\mathbb{P}}$. Under Assumption (ref), by the law of iterated expectation and the conditional mean independence between $Y_d$ and $D$ given $(V,Z)$, we have
Similarly, we can show that $\mathbb{E}[Y(1-D)\mid Z]=\mathbb{E}[m_{\mathbb{P},0D}(V,Z)\mid Z]$ and $\mathbb{E}[D\mid Z]=\mathbb{E}[m_{\mathbb{P},D}(V,Z)\mid Z]$. For any generic function $m=(m_{0},m_{1},m_{D},m_{0D},m_{1D})'\in\mathcal{M}$, let us introduce the following restrictions on $m$:
The conditional moment conditions in Eqs. (ref) to (ref) may not be appealing for estimation and inference when $Z$ is continuously distributed. Following mogstad/santos/torgovitsky:2017, we use IV-like functions to transform the conditional moment conditions into unconditional ones. Let $\mathcal{S}$ denote a (user-specified) subset of $\mathbf{L}^2(\mathcal{Z},\mathbb{R})$. Then, we use an IV-like function $s\in\mathcal{S}$ as a weighting function to integrate the conditional moments in Eqs. (ref) to (ref). Namely, Theorem (ref) implies that $m_{\mathbb{P}}$ satisfies the following unconditional moment conditions for every $s\in\mathcal{S}$:\footnote{The IV-like function $s\in\mathcal{S}$ plays the same role as an instrumental variable. For example, in a simple case where $Z_0$ is one-dimensional, if we set $s(z)=\frac{z_0-\mathbb{E}[Z_0]}{\mathbb{C}ov(D,Z_0)}$, then Eqs. (ref) and (ref) together imply that $\mathbb{E}\Big[\int_{{\mathcal{V}}}m_{\mathbb{P},1D}(v,Z)s(Z)dv\Big]+\mathbb{E}\Big[\int_{{\mathcal{V}}}m_{\mathbb{P},0D}(v,Z)s(Z)dv\Big]=\frac{\mathbb{C}ov(Y,Z_0)}{\mathbb{C}ov(D,Z_0)}$, where the right-hand side is the identifiable IV estimand.}
Recall that we denote $m=(m_0,m_1,m_D,m_{0D},m_{1D})'$ as some generic function in the space $\mathcal{M}$. To characterize the constraints on $m$ that is compatible with $m_{\mathbb{P}}$, let us introduce the following notations. For any $s\in\mathcal{S}\subseteq \mathbf{L}^2(\mathcal{Z},\mathbb{R})$ and $m\in\mathcal{M}$, let us define $$ \Gamma_{s}(m)=\mathbb{E}\Big[s(Z)\int_{{\mathcal{V}}}(m_{1D}(v,Z),m_{0D}(v,Z),m_D(v,Z))'dv\Big]. $$ Then, there are two sets of constraints under which $m\in \mathcal{M}$ is compatible with $m_{\mathbb{P}}$. The first set includes the following nonlinear constraints that come from the definition of $m_{\mathbb{P},dD}$:
The second set includes the linear constraints in Eqs. (ref) to (ref) that are imposed by identifiable estimands using observed data:
For a given $\mathcal{S}$, let us define $\mathcal{M}_{\mathcal{S}}$ as a collection of functions in $\mathcal{M}$ that satisfy both the above linear and nonlinear constraints
We can see from Theorem (ref) that $m_\mathbb{P}\in\mathcal{M}_{\mathcal{S}}$. However, for any given $\mathcal{S}$, $\mathcal{M}_{\mathcal{S}}$ is only an outer identified set for all $m\in\mathcal{M}$ that are compatible with $m_{\mathbb{P}}$, as it may not fully utilize all information from the observed data. If $\mathcal{S}$ expands to include more IV-like functions, then more restrictions are imposed on $m$, and $\mathcal{M_S}$ will become tighter. We formalize this result in the following corollary.
Corollary (ref) provides a sufficient condition under which $\mathcal{M}_{\mathcal{S}}$ reduces to the identified set of $m_{\mathbb{P}}$ characterized by the conditional moment constraints in Eqs. (ref) to (ref) and the constraints imposed by the definition in Eq. (ref). Specifically, it requires the choice of the IV-like functions in $\mathcal{S}$ to exhaust all available information in the data, which is not restrictive. For example, if $Z\in\{0,1\}$ is a binary scalar, then $\mathcal{S}=\{1[z=0],1[z=1]\}$ satisfies this condition. If $Z$ is continuous, then $\mathcal{S}=\{1[z\leq z']:~z'\in\mathcal{Z}\}$ needs to consist of indicators for all half-spaces of $\mathcal{Z}$. For any given $\mathcal{M_S}$, we define the identified set for the target PRTE parameter in Eq. (ref) as $$\{\Gamma^\star(m): m\in\mathcal{M_S}\}.$$
In this section, we define the sharp identified set for $m_{\mathbb{P}}$ and investigate sufficient conditions under which $\mathcal{M_S}$ is sharp. The sharp identified set for the target PRTE is then defined as the set of all possible values of $\Gamma^\star(m)$, where $m$ is in the sharp identified set for $m_\mathbb{P}$.
\noindentDefinition of Sharp Identified Set. Let us first introduce some useful notations. For any generic distribution $P$ for $(Y_0,Y_1,D,Z,V)$, we denote $$m_P=(m_{P,0},m_{P,1},m_{P,D},m_{P,0D},m_{P,1D})',$$ where $m_{P,d}$ with $d=0,1$ and $m_{P,D}$ are the marginal treatment response function and the propensity score function induced by $P$, respectively, and we define $m_{P,0D}=m_{P,0}(1-m_{P,D})$ and $m_{P,1D}=m_{P,1}m_{P,D}$. Let $\mathcal{P}$ be the set of distributions $P$ for $(Y_0,Y_1,D,Z,V)$ that satisfy the following conditions: (i) $P_{Y_D,D,Z}=\mathbb{P}_{Y,D,Z}$, (ii) $E_P[Y_d\mid D,Z,{V}]=E_P[Y_d\mid X,{V}]$ and $E_P[{E_P}[Y_d\mid X,{V}]^2]<\infty$ for each $d=0,1$, and (iii) $V$ conditional on $Z$ is uniformly distributed on $\mathcal{V}$ with a known dimension $K$.\footnote{We use $P_{Y_D,D,Z}$ to denote the distribution for $(Y_D,D,Z)$ induced by $P$, and we use $\mathbb{P}_{Y,D,Z}$ to denote the distribution for $(Y,D,Z)$ induced by the true distribution $\mathbb{P}$.} In other words, $\mathcal{P}$ is the set of distributions that satisfy Assumption (ref) and are observationally equivalent to $\mathbb{P}$. Then, the sharp identified set for $m_{\mathbb{P}}$ can be defined by $$\{m_{P}\in\mathcal{M}: P\in\mathcal{P}\}.$$ If $m_\mathbb{P}$ satisfies certain shape restriction of the form $m_\mathbb{P}\in\mathcal{R}$ for a known set $\mathcal{R}\subseteq\mathcal{M}$, such as those discussed in Section (ref), then the sharp identified set for $m_{\mathbb{P}}$ can be defined as $$ \{m_{P}\in\mathcal{R}: P\in\mathcal{P}\}. $$
\noindentSharpness of $\mathcal{M_S}$ with Binary Outcome. Next, we establish the sharpness of the identified set $\mathcal{M_S}$ defined in (ref) in an empirically relevant case with a binary outcome variable.\footnote{It is interesting to explore sharpness in cases with general outcome variables. However, we leave a rigorous exploration for future research.} In the theorem below, we show that if we utilize all available information from observed data by choosing $\mathcal{S}=\mathbf{L}^2(\mathcal{Z},\mathbb{R})$, then $\mathcal{M_S}$ is the sharp identified set for $m_\mathbb{P}$.
When $\mathcal{M_S}$ is the sharp identified set for $m_\mathbb{P}$, the identified set for PRTE based on $\mathcal{M_S}$ is also sharp. Despite the sharpness demonstrated in Theorem (ref), the PRTE identified set can be computationally challenging to obtain, as it requires solving nonconvex optimization problems due to the nonlinear constraints in (ref). In what follows, we address this challenge by introducing a convex relaxation method.
In the objective function $\Gamma^\star(m)$ and the constraints on $m\in\mathcal{M_S}$, everything is linear in $m$ except for the nonlinear (in particular, bilinear) constraints in Eq. (ref). In this section, we propose to use the convex relaxation method proposed by mccormick1976computability to relax the bilinear constraints in Eq. (ref) into linear constraints. This then allows us to bound $\Gamma^\star(m)$ by solving linear optimization problems. Let us first introduce the convex-relaxation presentation of the bilinear constraints on $m_{\mathbb{P},d}$.
In Lemma (ref), we assume the bounds $m^L_d(v,x)$, $m^U_d(v,x)$, $m^L_D(v,z)$, and $m^U_D(v,z)$ are either known or identifiable. If no prior information about these upper and lower bounds is available, we can set $m^L_D(v,z)$ and $m^U_D (v,z)$ to 0 and 1, respectively, since $D$ is a binary variable, and set $m_d^L(v,x)$ and $m_d^U(v,x)$ to $y^L$ and $y^U$, respectively, if they are finite. From Lemma (ref), we can see that after applying the convex relaxation, the equality constraints in Eq. (ref), which are nonlinear in $m_{\mathbb{P}}$, can be relaxed and replaced by the inequality constraints in Eqs. (ref) to (ref), which are all linear in $m_{\mathbb{P}}$. The boundaries of the set for $m_{\mathbb{P},dD}$, characterized by the inequalities in (ref) and (ref), are referred to as the McCormick envelopes.
Under convex relaxation, for any given user-specified $\mathcal{S}$, let us define $$\mathcal{M}^r_\mathcal{S}=\{m\in \mathcal{M}:~m\text{ satisfies Eq. \eqref{cond:Gamma} for any $s\in \mathcal{S}$, and Eqs. \eqref{eq:mccormick_bounds_m} to \eqref{eq:mccormick_relax4_1}}\}.$$ Then, $\mathcal{M}^r_\mathcal{S}$ is the collection of functions $m\in \mathcal{M}$ characterized by linear constraints from observed data and convex relaxation. Thus, it is an outer set for all $m\in\mathcal{M}$ that are compatible with $m_\mathbb{P}$ even if $S$ is a set of sufficiently rich class of IV-like functions.
Based on $\mathcal{M}^r_\mathcal{S}$, we can define an outer set for the target PRTE parameter as below:
Since both $\Gamma^\star(m)$ and the constraints that characterize $\mathcal{M}^r_\mathcal{S}$ are linear (therefore, convex) in $m$, under Assumption (ref), the PRTE outer set in (ref) is equivalent to an interval. Let us denote this interval as $[\underline\beta^\star,\overline\beta^\star]$. Then, we can characterize the two ending points $\underline\beta^\star$ and $\overline\beta^\star$ using the following linear optimization problems:\footnote{Note that we omit the dependence of $\underline\beta^\star$ and $\overline\beta^\star$ on $\mathcal{S}$ for simplicity. }
There are a few advantages of considering the outer set $\{\Gamma^\star(m): m\in\mathcal{M}^r_\mathcal{S}\}$. First, it gives us computationally tractable PRTE bounds because $\underline\beta^\star$ and $\overline\beta^\star$ can be easily and reliably solved by linear programming techniques that are routinely used in many empirical works. Second, shape restrictions that are linear in $m$ can be easily incorporated into the linear programming problems to further shrink the bounds. Third, under some conditions, we can show that $\underline\beta^\star$ and $\overline\beta^\star$ can be computed without knowing the true dimension of the unobserved heterogeneity $V$. To illustrate this, let us denote
Let $\tilde{m}=(\tilde{m}_0,\tilde{m}_1,\tilde{m}_D,\tilde{m}_{0D},\tilde{m}_{1D})'\in \tilde{\mathcal{M}}$, where $ \tilde{\mathcal{M}}=\{\tilde{m}:~\tilde{m}\mbox{ defined in }\eqref{prop_reduce_dim_fun}\mbox{ for some }m\in\mathcal{M}\}$.
Proposition (ref) provides sufficient conditions under which our proposed convex-relaxation bounds for PRTEs can be obtained by solving linear optimization problems over the space $\tilde{\mathcal{M}}$. Given that any $\tilde{m}\in\tilde{\mathcal{M}}$ is degenerate in $V$, the solutions to these linear optimization problems over $\tilde{\mathcal{M}}$ do not rely on knowledge of the true dimension of $V$. This notable result is contingent upon a crucial condition: the weights $\omega_{dd'}^\star(V,Z)$ and the known bounds for $m_{\mathbb{P},d}$ and $m_{\mathbb{P},D}$ must be degenerate in $V$. Specifically, the former, $\omega_{dd'}^\star(V,Z)=\omega_{dd'}^\star(Z)$, holds for many target PRTE estimands listed in Table (ref). The latter holds if we set constant upper and lower bounds for $m_{\mathbb{P},d}$ and $m_{\mathbb{P},D}$.
Despite the advantages of the convex relaxation method mentioned above, its conservativeness may raise concerns. This is because the resulting identified set for $m_{\mathbb{P}}$, $\mathcal{M}^r_\mathcal{S}$, is only an outer set, which is often larger and less informative than its identified set $\mathcal{M_S}$ that does not rely on convex relaxation. This section investigates this concern. We consider two cases, and in each case, we provide sufficient conditions under which our convex-relaxation bounds result in no identification power loss.
We first consider the case where the known or identifiable lower and upper bounds of $m_D(V,Z)$ used in the convex relaxation coincide, implying that $m_D(V,Z)$ is identifiable. If researchers are aware of this and use it properly by setting $m_D^L(V,Z)=m_D^U(V,Z)=m_D(V,Z)$ when applying the convex relaxation method, we can show that the outer set for $\Gamma^\star(m)$ derived from $\mathcal{M}^r_\mathcal{S}$ reduces to its identified set derived from $\mathcal{M_S}$.
The result in Theorem (ref) is straightforward because under the condition $m_D^L(V,Z)=m_D^U(V,Z)$, the inequality constraints in (ref) to (ref) all reduce to the equality constraints in (ref). We can see that $m_D^L(V,Z)=m_D^U(V,Z)$ holds, if $m_{\mathbb{P},D}(V,Z)$ is point identified. This is the case if we impose the shape restriction in Condition (ref) that $D$ is generated by a random coefficient model, where $m_{\mathbb{P},D}(V,Z)=1[\pi_1Z_1+\pi_2Z_2+...+\pi_KZ_K>0]$ is point identified. Another example for the point identified $m_{\mathbb{P},D}(V,Z)$ is when we impose the shape restriction in Condition (ref) that $D$ is generated by a double hurdle model, where $m_{\mathbb{P},D}(V,Z)=1[U_1<Q_1(Z)\text{ and }U_2<Q_2(Z)]$ is point identified. Besides, if we impose the classical monotonicity restriction on $D$, which is equivalent to assuming that $D$ is generated by a threshold-crossing model with a single-dimensional $V\sim \mathrm{Uniform}[0,1]$, then $m_{\mathbb{P},D}(V,Z)=1[\mathbb{P}[D=1\mid Z]\geq V]$ is also point identified.
By combining Theorems (ref) and (ref), we can see that in the case of a binary outcome, if $m_{\mathbb{P},D}(V,Z)$ is point identified, the convex-relaxation bounds for PRTEs are sharp as long as we choose IV-like functions to exhaust all the information from observed data. Such a sharpness result directly relates to the existing literature. For example, heckman/vylacil:2001:book propose the sharp ATE bound assuming bounded outcomes and threshold crossing in the treatment equation. Below, we show that our convex-relaxation bounds for ATE without shape restrictions coincide with the sharp ATE bounds of heckman/vylacil:2001:book.
Next, we consider another case in which the convex-relaxation bounds are simplified to those derived by \citet*{mogstad/santos/torgovitsky:2017} (hereafter, MST) if the true data satisfy the threshold-crossing model for the treatment with one-dimensional unobserved heterogeneity, even though the researcher neither knows nor uses this information. We use the propensity score given $Z$, defined by $p(z)=\mathbb{P}(D=1|Z=z)$. Suppose the treatment variable is generated by a threshold-crossing model in the true data:
where $U$ conditional on $Z$ is uniformly distributed on $[0,1]$, $\mathbb{E}[Y_d\mid D,Z,U]=\mathbb{E}[Y_d\mid X,U]$, and $\mathbb{E}[Y^2_d]<\infty$ for $d\in\{0,1\}$. Here, we use a different notation $U$ to represent the unobserved heterogeneity, in order to emphasize that the researcher may not know whether the true data satisfy the threshold-crossing assumption for the treatment, and instead, a multidimensional $V$ can be assumed when computing our convex-relaxation bounds.
Let us first introduce the MST bounds assuming the threshold-crossing structure in (ref). With notation abuse, we denote the marginal treatment response functions as $\tilde{m}_{\mathbb{P},d}(u,x)=\mathbb{E}[Y_d\mid U=u, X=x]$ for $d=0,1$. The propensity score given $(U,Z)$, i.e., $\mathbb{E}[D|U=u,Z=z]=1\{p(z)\geq u\}$, is point identified for any $(u,z)\in[0,1]\times\mathcal{Z}$. Therefore, different from the expression of our target parameter given in (ref), the MST target parameter depends only on the marginal treatment response functions:
where $\tau^\star_d(u,z)$ for $d=0,1$ denotes some known or identifiable weights. Under the the threshold-crossing structure in (ref), the expression in (ref) is equivalent to that of our target parameter in (ref) if we set
and all PRTE parameters listed in Table (ref) can be equivalently expressed as in (ref), using the weights given in (ref). For any generic functions $(\tilde{m}_0,\tilde{m}_1)$ in some parameter space $\tilde{\mathcal{M}}_{01}\subseteq(\mathbf{L}^2([0,1]\times\mathcal{X},[y^L,y^U]))^2$, define a map from $\tilde{\mathcal{M}}_{01}$ to $\mathbb{R}$ as below: $$\tilde{\Gamma}^\star(\tilde{m}_0,\tilde{m}_1)=\mathbb{E}\left[\int_0^1\tilde{m}_0(u,X)\tau^\star_0(u,Z)du\right]+\mathbb{E}\left[\int_0^1\tilde{m}_1(u,X)\tau^\star_1(u,Z)du\right].$$ In addition, define $\tilde{\mathcal{S}}$ to be a collection of $\tilde{s}(d,z)$, where $\tilde{s}:\{0,1\}\times\mathcal{Z}\mapsto\mathbb{R}$ denotes the identifiable (or known) IV-like function used by MST to construct identified sets for $\tilde{m}_{\mathbb{P},0}$ and $\tilde{m}_{\mathbb{P},1}$. For any generic functions $\tilde{m}_0$ and $\tilde{m}_1$, they consider the following restriction
MST has shown that $\tilde{m}_{\mathbb{P},0}$ and $\tilde{m}_{\mathbb{P},1}$ must satisfy the restriction in (ref) for any $\tilde{s}\in\tilde{\mathcal{S}}$. Define $$\tilde{\mathcal{M}}_{01,\tilde{\mathcal{S}}}=\left\{(\tilde{m}_0,\tilde{m}_1)\in \tilde{\mathcal{M}}_{01}:~\tilde{m}_0\text{ and }\tilde{m}_1\text{ satisfy \eqref{MST_restrictions} for all }\tilde{s}\in \tilde{\mathcal{S}}\right\}.$$ Then, the MST bounds for the target PRTE parameter $\tilde{\Gamma}^\star(\tilde{m}_0,\tilde{m}_1)$ is
Even if the true data generating process satisfies the threshold-crossing assumption in (ref), the researcher may not be aware of it and may instead compute our convex relaxation bounds. The theorem below provides sufficient conditions under which our convex relaxation bounds for $\Gamma^\star(m)$ coincide with the MST bounds for $\tilde{\Gamma}^\star(\tilde{m}_0,\tilde{m}_1)$.
Theorem (ref) provides sufficient conditions under which our convex relaxation method does not result in identification power loss compared to the MST bounds. That is, if the true data generating process satisfies the threshold-crossing assumption in (ref), our convex-relaxation bounds—despite not utilizing this information—are just as tight as the MST bounds for a large class of PRTE parameters whose weights do not depend on $V$. Moreover, this result also holds under a particular set of linear shape restrictions that are imposed on the integrated treatment response functions. In short, our method generalizes MST, and remains applicable and more robust in settings where practitioners are uncertain whether the threshold-crossing structure in (ref) holds.
Theorem (ref) may resemble Theorem 2 of heckman/vylacil:2001:book, but they are substantively different results. In their Theorem 2, if the threshold-crossing assumption in (ref) holds, the manski1990nonparametric IV-mean-independence sharp bounds, which do not impose the threshold crossing structure for the treatment, are equivalent to the Heckman-Vytlacil sharp bounds. In our Theorem (ref), we do not consider the equivalence between the sharp bounds with and without the threshold-crossing assumption in (ref). Instead, we consider the equivalence between our convex-relaxation bounds and the MST bounds.
In this section, we present our computation and inference procedure for PRTE bounds defined by the linear optimization problems in (ref). To solve the optimization problems using linear programming techniques, we first replace the potentially infinite-dimensional parameter space for $m$ with a finite-dimensional space defined using basis functions. Then, we demonstrate that the computationally tractable inference method of gafarov2019inference can be directly applied in our analysis to construct uniformly valid PRTE confidence sets.
In a similar vein to mogstad/santos/torgovitsky:2017, we approximate the unknown function $m$ using a finite number of known basis functions. Let $\mathbf{B}_d:(\mathcal{V,X})\mapsto\mathbb{R}^{K_d}$ with $d=0,1$ be a vector of known basis functions for the unknown function $m_d$. Similarly, let $\mathbf{B}_{D}:(\mathcal{V,Z})\mapsto\mathbb{R}^{K_D}$ be a vector of known basis functions for $m_D$, and $\mathbf{B}_{dD}:(\mathcal{V,Z})\mapsto\mathbb{R}^{K_{dD}}$ with $d=0,1$ be a vector of known basis functions for $m_{dD}$. Then, we can approximate $m=(m_0,m_1,m_D,m_{0D},m_{1D})'$ by $\mathbf{B}'\eta_2$ with a $d_{\eta_2}\times 1$ vector of unknown parameters $\eta_2$ and a matrix of known basis functions $\mathbf{B}$, where $$ \mathbf{B}'\eta_2=
\eta_2. $$ We denote $\eta_1=\Gamma^\star(\mathbf{B}'\eta_2)$ and $\eta=(\eta_1,\eta_2')'$. Define the parameter space for $\eta$ as below $$\Upsilon=\{\eta=(\eta_1,\eta_2')'\in\mathbb{R}^{d_\eta}: \eta_1=\Gamma^\star(\mathbf{B}'\eta_2)\in\mathbb{R} and \mathbf{B}'\eta_2\in\mathcal{M}\}.$$ Then, the dimension of $\eta$ is $d_\eta=d_{\eta_2}+1=K_0+K_1+K_D+K_{0D}+K_{1D}$. Let $e_1=(1,0,...,0)'\in\mathbb{R}^{d_\eta}$. With this approximation, we can modify the optimization problems in (ref) into
$\eta$ subject to
Note that any linear shape restrictions on $m$ (e.g., those discussed in Section (ref)) can be easily transformed into linear constraints on the unknown parameters in $\eta$.\footnote{For a known matrix $\mathbf{C}$, we can use $\mathbf{C}\eta\leq 0$ to describe the linear shape restrictions on $m$. For example, if there are no covariates and we impose the restriction that $\mathbb{E}[m_1(V)-m_0(V)]\geq0$, then we can set $\mathbf{C}=[0,\mathbb{E}[\mathbf{B}'_0],-\mathbb{E}[\mathbf{B}'_1],\mathbf{0}'_{K_D+K_{0D}+K_{1D}}]$, where $\mathbf{0}_l$ denotes a $l\times 1$ vector of zeros.} Thus, we do not consider shape restrictions in this section for notation simplicity.
Depending on the support of $Z$, the above finite-dimensional approximation of the unknown function $m$ may lead to narrower PRTE bounds compared to the nonparametric bounds $[\underline\beta^\star,\overline\beta^\star]$ in (ref). In other words, $[\underline\beta^\star_a,\overline\beta^\star_a]$ may be a proper subset of $[\underline\beta^\star,\overline\beta^\star]$ or even a subset of $\{\Gamma^\star(m):~m\in\mathcal{M_S}\}$. Below, we show that when $Z$ has a finite support, computing the bounds using finite-dimensional constant splines does not alter the nonparametric bounds. Without loss of generality, assume $X\in\{x_1,...,x_{K_X}\}$ and $Z\in\{z_1,...,z_{K_Z}\}$. For any user-specified partition $\mathcal{V}=\bigcup_{k=1}^{K_V}\mathcal{V}_k$ and $d\in\{0,1\}$, let us consider basis functions $\mathbf{B}_d=\{b_{d,kj}\}$ and $\mathbf{B}_D=\mathbf{B}_{dD}=\{b_{D,kl}\}$, where
We refer to any finite-dimensional approximation $\mathbf{B}'\eta_2$ as constant splines if $\mathbf{B}$ are constructed using basis functions in (ref). For any $m\in\mathcal{M}$, one special constant spline that provides the best mean squared error approximation for $m$, denoted by $\Lambda m=(\Lambda m_0,\Lambda m_1,\Lambda m_D,\Lambda m_{0D},\Lambda m_{1D})$, is given by
which corresponds to setting the unknown coefficients in $\eta_2$ to the conditional means of $m_d(V,X)$ given $(V,X)$, and the conditional means of $m_D(V,Z)$ and $m_{dD}(V,Z)$ given $(V,Z)$. The proposition below provides conditions under which the bound $[\underline\beta^\star_a,\overline\beta^\star_a]$, obtained by solving the linear programs defined by (ref) and (ref) over the space of constant splines, is the same as the nonparametric bound $[\underline\beta^\star,\overline\beta^\star]$ defined in (ref). This result extends the exact computational approach demonstrated by Proposition 4 of mogstad/santos/torgovitsky:2017 to more general settings with multidimensional unobserved heterogeneity.
In general, the above finite-dimensional approximation approach can be applied to compute the bounds as long as we approximate $m$ using a linear-in-parameter specification, regardless of the support of $Z$. For example, if $Z$ is a vector of possibly continuous IVs and covariates, we can employ additive linear structures to approximate the unknown functions: $m_d(v,x)=g_d(v)+x'\eta_{d,X}$, $m_D(v,z)=g_D(v)+z'\eta_{D,Z}$, and $m_{dD}(v,z)=g_{dD}(v)+z'\eta_{dD,Z}$ for $d=0,1$, where $\eta_{d,X}$, $\eta_{D,Z}$, and $\eta_{dD,Z}$ are finite-dimensional vectors of unknown parameters. Then, we further approximate the unknown function $g(v)=(g_0(v),g_1(v),g_D(v),g_{0D}(v),g_{1D}(v))$ using finite-dimensional basis functions.\footnote{However, unlike the case with a finite support of $Z$, if $Z$ is continuous, we may need to increase the dimension of the basis functions as the sample size increases, to ensure that the resulting PRTE bounds converge to the nonparametric bounds $[\underline\beta^\star,\overline\beta^\star]$.}
In this section, we denote $\mathcal{S}=\{s_1,...,s_{|\mathcal{S}|}\}$ by assuming its cardinality is finite, i.e., $|\mathcal{S}|<\infty$. For illustrative purposes, we do not consider shape restrictions in the presentation of the inference procedure. Nonetheless, these restrictions can be easily incorporated with additional notations. Below, we demonstrate that the regularized support function method of gafarov2019inference can be directly applied in our analysis to obtain uniformly valid confidence sets for the PRTEs. In this section, we closely follow the notations and assumptions from gafarov2019inference.
Let us denote by $\Upsilon(\mathbb{P})$ the set of $\eta$ that satisfies all the constraints in (ref): $$\Upsilon(\mathbb{P})=\left\{\eta=(\eta_1,\eta_2')'\in\Upsilon:~\eta_1=\Gamma^\star(\mathbf{B}'\eta_2)\text{ and }\mathbf{B}'\eta_2\in \mathcal{M}^r_\mathcal{S}\mbox{ for any given }\mathcal{S}\right\}.$$ Without loss of generality, let us assume that the set $\Upsilon(\mathbb{P})$ is compact. Denote $\mathcal{J}^{eq}$ as a set of $d_{eq}$ equality constraints which include $\eta_1=\Gamma^\star(\mathbf{B}'\eta_2)$ and those given in (ref). Denote $\mathcal{J}^{ineq}$ as the set of $d_{ineq}$ inequality constraints in (ref) to (ref) that describe the convex relaxation. Since $\Upsilon(\mathbb{P})$ only consists of linear constraints in $\eta$, we can find a random matrix $\mathbf{W}\in\mathbb{R}^{(d_{eq}+d_{ineq})\times (d_\eta+1)}$ that depends only on observed variables, such that $\Upsilon(\mathbb{P})$ is equivalently expressed as a set of $\eta$ that satisfies
where $g_j(\mathbf{W},\eta)=\sum_{l=1}^{d_\eta}W_{j,l}*\eta_l-W_{j,d_\eta+1}$ for all $j\in\mathcal{J}^{eq}\cup\mathcal{J}^{ineq}$, and $W_{j,l}$ denotes the $(j,l)$-th element in $\mathbf{W}$. The detailed expression of $\mathbf{W}$ is given in Appendix (ref). Let us rewrite $$\mathbf{W}=
,$$ where the four blocks $\mathbf{W}^{eq}_1\in\mathbb{R}^{d_{eq}\times d_\eta}$, $\mathbf{W}^{eq}_2\in\mathbb{R}^{d_{eq}\times 1}$, $\mathbf{W}^{ineq}_1\in\mathbb{R}^{d_{ineq}\times d_\eta}$, and $\mathbf{W}^{ineq}_2\in\mathbb{R}^{d_{ineq}\times 1}$ are compatible with the dimensions of $\mathcal{J}^{eq}$ and $\mathcal{J}^{ineq}$. Denote $A_{\mathbb{P}}=\mathbb{E}
$ and $\pi_{\mathbb{P}}=\mathbb{E}
$. Then the linear programs in (ref) and (ref) are equivalent to
where $e_j$ is a vector whose $j$-th entry is one and all other entries are zero. The expression of the linear programs in (ref) is the same as that in gafarov2019inference. Therefore, we can directly apply his method of the regularized support function to construct uniformly valid confidence intervals for $\left[\underline{\beta}_a^\star,\overline{\beta}_a^\star\right]$. The idea of the method is to add a small regularization term to the objective functions and consider the following regularized programs which are strictly convex:
where $\|\cdot\|$ denotes the Euclidian norm, and $\mu_n>0$ with $\mu_n\rightarrow0$ as sample size $n\rightarrow\infty$ is a tuning parameter chosen by the researcher.\footnote{Following gafarov2019inference, we use the tuning parameter $\overline{\mu}_n=\sqrt{\frac{\log(\log(n))}{n}}\hat{\overline{\mu}}_1\text{ and }\underline{\mu}_n=\sqrt{\frac{\log(\log(n))}{n}}\hat{\underline{\mu}}_1,$ where $\hat{\overline{\mu}}_1=\frac{(tr\{\hat{var}[\hat{\overline{\lambda}}_n(0)'\mathbf{W}_i]\})^{1/2}}{d_\eta+1}$ and $\hat{\underline{\mu}}_1=\frac{(tr\{\hat{var}[\hat{\underline{\lambda}}_n(0)'\mathbf{W}_i]\})^{1/2}}{d_\eta+1}$, and $\hat{\overline{\lambda}}_n(0)$ and $\hat{\underline{\lambda}}_n(0)$ are the Lagrange multiplier of the nonregularized maximization and minimization problem, respectively.}
gafarov2019inference implements the inference procedure in three steps.
Step 1. Construct the sample analog of the regularized programs. Denote $\mathbf{W}_i=
$ as a sample observation of $\mathbf{W}$ for $i=1,...,n$. Define $\hat{A}_{n}=\frac{1}{n}\sum_{i=1}^{n}
$ and $\hat{\pi}_{n}=\frac{1}{n}\sum_{i=1}^{n}
$. The sample analog of \eqref{regularized_primal} with some sequence $\mu_n$ is given by
Step 2. Estimate the asymptotic variance of the regularized support function estimators, $\hat{\underline{\beta}}_a^\star(\mu_n)$ and $\hat{\overline{\beta}}_a^\star(\mu_n)$, denoted by $ \hat{\underline{\sigma}}^2_n(\mu_n)$ and $\hat{\overline{\sigma}}_n^2(\mu_n)$, respectively. For any given $\mu_n$, denote $\hat{\underline{\eta}}^\star_n(\mu_n)$ and $\hat{\overline{\eta}}^\star_n(\mu_n)$ as the solutions to the regularized programs in (ref). Then, the asymptotic variance can be computed as below:
where $\hat{\underline{\lambda}}_n(\mu_n)$ and $\hat{\overline{\lambda}}_n(\mu_n)$ are the estimated vectors of Lagrange multipliers of the programs in (ref).
Step 3. Compute the bias correction term, the outer bound estimator, and the uniformly valid confidence set. Recall that compared to $\underline{\beta}_a^\star$ and $\overline{\beta}_a^\star$, the regularization term $\mu_n\|\eta\|^2$ introduces an upward and a downward bias in $\hat{\underline{\beta}}_a^\star(\mu_n)$ and $\hat{\overline{\beta}}_a^\star(\mu_n)$, respectively. Below, we introduce two bias correction terms $\mu_n\|\hat{\underline{\eta}}^{out}_n\|^2$ and $\mu_n\|\hat{\overline{\eta}}^{out}_n\|^2$ in Step 3.1, so that the outer bound estimators defined in Step 3.2 are asymptotically unbiased or biased towards the desirable directions.
The following assumptions are employed by gafarov2019inference to show the asymptotic properties of the confidence interval in (ref). We present them here for completeness.
Let $\mathcal{P}^\star$ be a collection of $P$ that satisfies Assumptions (ref) and (ref). The theorem below demonstrates that the $(1-\alpha)$-confidence interval in (ref) covers $\left[\underline{\beta}_a^\star,\overline{\beta}_a^\star\right]$ uniformly across $\mathcal{P}^\star$ with probability at least $1-\alpha$.
In this section, we numerically illustrate the performance of our proposed convex-relaxation bounds and verify the asymptotic properties in finite sample using Monte Carlo simulations. We omit the covariates $X$ to ease the computation.
Consider the data generating process (DGP) for a binary outcome variable: $$Y_1=1[m_1(V)-\varepsilon_1>0]\mbox{ and }Y_0=1[m_0(V)-\varepsilon_0>0],$$ where $(\varepsilon_0,\varepsilon_1)\sim \mathrm{Uniform}([0,1]^2)$. We consider two different dimensions of the unobserved heterogeneity: (i) $V\sim \mathrm{Uniform}[0,1]$; and (ii) $V=(V_1,V_2)\sim \mathrm{Uniform}([0,1]^2)$. The specifications for $(m_0,m_1)$ are summarized in Table (ref).
Consider two mutually independent instruments $Z=(Z_1,Z_2)\in\mathcal{Z}$, where $\mathcal{Z}=\{0,1,2\}\times\{0.5,1\}$ with $\mathbb{P}(Z_1=0)=0.5$ and $ \mathbb{P}(Z_1=1)=0.4$, and $\mathbb{P}(Z_2=0.5)=0.7$. We generate the treatment variable using a random coefficient model:
where $U\sim N(0,var(U))$ with $sd(U)\in\{0.1,0.5,0.9\}$. In addition, $Z$, $(\varepsilon_0,\varepsilon_1)$, $U$ and $V$ are mutually independent. Let
In the numerical illustration and the simulation studies, we consider three target parameters: (i) ATE$=\mathbb{E}[Y_1-Y_0]$, (ii) ATT$=\mathbb{E}[Y_1-Y_0\mid D=1]$, and (iii) the average selection bias ASB$=\mathbb{E}[Y_0\mid D=1]-\mathbb{E}[Y_0\mid D=0]$.
In this section, we numerically illustrate the performance of our convex-relaxation bounds, starting from the cases without any shape restrictions. For all three target parameters, we compare our method with three existing bounding methods:
Specifically, both Manski and HV derive analytical bounds for $\mathbb{E}[Y_0]$ and $\mathbb{E}[Y_1]$, which only depend on the distributions of observed variables, $(Y,D,Z)$. With these bounds, we can compute the Manski and HV bounds for the three target parameters, as all of them can be expressed as some linear combinations of $\mathbb{E}[Y_0]$ and $\mathbb{E}[Y_1]$:
On the other hand, MST and CvR bounds are both computed using linear programming methods. In particular, we calculate MST bounds using all the IV-like functions in $\tilde{\mathcal{S}}=\mathbf{L}^2(\{0,1\}\times\mathcal{Z},\mathbb{R})$.
For CvR bounds, we implement the linear programming computation outlined in Section (ref), and we set $m_d^L(v,x)=m_D^L(v,z)=0$ and $m_d^U(v,x)=m_D^U(v,z)=1$ for the McCormick convex relaxation. We adopt the IV-like functions in $\mathcal{S}=\mathbf{L}^2(\mathcal{Z},\mathbb{R})$, which exhausts all available information from the observed data.\footnote{In the data generating process of this section, the support $\mathcal{Z}$ is finite, so we only need to use the indicator functions for each point in $\mathcal{Z}$ as the IV-like function.} Because $\omega^\star_{dd'}(v,z)$, $m_d^L(v,x)$, $m_d^U(v,x)$, $m_D^L(v,z)$ and $m_D^U(v,z)$ all do not depend on $v$ for all three target parameters, according to Proposition (ref), we do not need to know the true dimension of $V$ and only need to bound the integral of the function $m$ in Eq.(ref), instead of the function $m$ itself. Thus, for both the single-dimensional and two-dimensional $V$ cases, we use $\mathbf{B}_d=1$ for $\int_{\mathcal{V}}m_d(v)dv$ with $d=0,1$, and $\mathbf{B}_D=\mathbf{B}_{dD}=\{1[z=l]\}_{l\in\mathcal{Z}}$ for $\int_{\mathcal{V}}m_D(v,z)dv$ and $\int_{\mathcal{V}}m_{dD}(v,z)dv$.
\noindentBounds without Shape Restrictions. \quad Numerical results of the four bounding methods are summarized in Table (ref). There are a few points worth mentioning. First, we know that the assumptions required by HV bounds, which include the threshold-crossing structure, are stronger than those required by Manski bounds. However, for ATE, its Manski bounds are narrower than or equal to HV bounds across all DGP designs, as the threshold-crossing structure does not hold. This highlights that imposing more structural assumptions does not always lead to narrower bounds. Whereas for ATT and ABS, Manski bounds and HV bounds coincide with each other. Second, due to model misspecification, there are no solutions to the linear programming problems of the MST method, resulting in empty MST bounds for all three target parameters in all DGP designs.\footnote{The empty MST bounds are an indication of the failure of their assumption (i.e., the threshold-crossing structure). However, it is not vice versa, that is, the MST bounds can be non-empty under other DGPs even if the threshold-crossing structure fails.} Third, our CvR bounds for all three target parameters are the same as their Manski bounds across all DGPs designs. These results demonstrate the consequences of model misspecification in the MST method and the informativeness of our CvR bounds when the threshold-crossing structure fails to hold.
\noindentBounds with Shape Restrictions. \quad Next, we evaluate our CvR method under shape restrictions. First, we consider the restrictions in (ref), which are implied by the monotone treatment response (MTR) in Condition (ref). Define
Second, we consider the shape restrictions in (ref), which are implied by the monotone treatment selection (MTS) in Condition (ref). Define
Third, let us consider the shape restriction in (ref), which is implied by the monotone selection on the gains (MSG) in Condition (ref). Define
The last shape restriction is a combination of MTR, MTS, and MSG, defind as $$\mathcal{R}^4=\mathcal{R}^1\cap\mathcal{R}^2\cap\mathcal{R}^3.$$ Note that all these four sets of shape restrictions are satisfied by our DGPs.
The resulting CvR bounds for the three target parameters under these shape restrictions are presented in Tables (ref), (ref), and (ref). Notably, compared to the CvR bounds without restrictions, the restrictions implied by MTR, as in $\mathcal{R}^1$, significantly reduce the CvR bound size for all three target parameters and identify the signs of both ATE and ATT. Furthermore, the restrictions implied by MTS, as in $\mathcal{R}^2$, improve the upper CvR bounds for ATE and ATT and identify the sign of ASB. Compared to MTR and MTS, the restrictions imposed by MSG, as in $\mathcal{R}^3$, yield relatively smaller improvements to the CvR bounds for all three target parameters. Finally, by combining all three shape restrictions, as in $\mathcal{R}^4$, we obtain the tightest CvR bounds, which coincide with the intersection of the bounds under $\mathcal{R}^1$ to $\mathcal{R}^3$.
To check the finite sample performance of the CvR bounds, we conduct Monte Carlo simulations following the gafarov2019inference inference method described in Sections (ref). The DGPs, target parameters, IV-like functions, and known boundaries $m^L_d$, $m^U_d$, $m^L_D$, and $m^U_D$ in the convex relaxation are the same as those used to generate the numerical results summarized in Table (ref). In this section, we do not consider shape restrictions. We present the coverage rates of the 95% confidence interval for ATE, ATT, and ASB in Figures (ref), (ref), and (ref), respectively. These results demonstrate that the asymptotic properties of the inference method, as described in Theorem (ref), hold true in finite samples. Importantly, the confidence intervals exhibit reasonable width and cover the true CvR bounds with a probability of at least 95%. As the sample size increases, the confidence intervals become narrower, and they exclude values outside the true CvR bounds with an increasing probability.
This paper proposes a unified framework to partially identify various policy relevant treatment effects (PRTEs) using a convex relaxation method, when the threshold-crossing structure for the treatment may not hold. Our proposed method robustifies the one of mogstad/santos/torgovitsky:2017 in the following sense: When the true data generating process satisfies the threshold-crossing assumption for the treatment our convex-relaxation bounds for a large class of PRTE parameters are simplified to the bounds of mogstad/santos/torgovitsky:2017, even if we do not impose this condition. This robustness makes our method a more attractive option for practitioners, as it remains applicable regardless of whether the threshold-crossing structure for the treatment holds or not.
We focus on a flexible model that accommodates multidimensional unobserved heterogeneity. We assume that the potential outcomes and the treatment variable are mean-independent given the unobserved heterogeneity. Under this assumption, both the target parameters and the identifiable estimands can be expressed as weighted averages of bilinear functions of the unknown marginal treatment responses and propensity scores given the unobserved heterogeneity. The PRTE bounds are then characterized by nonconvex optimization problems subject to restrictions imposed by the identifiable estimands.
To address the nonconvexity problem, we introduce a convex relaxation method which is attractive for several reasons. First, it provides a computationally feasible identified set for PRTEs. This method replaces each bilinear function with a new function and four linear inequality constraints, converting a bilinear program to a linear program. Second, we show that our proposed convex-relaxation bounds do not lose identification power under some conditions as long as we fully exploit the available information on observable data. Third, linear shape restrictions and their combinations can be easily incorporated to further improve the bounds. At last, we can directly apply the existing inference method of gafarov2019inference to construct uniformly valid confidence intervals for the PRTE bounds. Numerical and simulation analyses demonstrate that our convex-relaxation bounds are informative under violations of the threshold-crossing structure.