EconBase
← Back to paper

On the Identifying Power of Monotonicity for Average Treatment Effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

52,061 characters · 8 sections · 56 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

On the Identifying Power of Generalized Monotonicity for Average Treatment Effects

spacing{1.2} \begin{abstract} In the context of a binary outcome, treatment, and instrument, balke1993nonparametric,balke1997bounds establish that the monotonicity condition of imbens1994identification has no identifying power beyond instrument exogeneity for average potential outcomes and average treatment effects in the sense that adding it to instrument exogeneity does not decrease the identified sets for those parameters whenever those restrictions are consistent with the distribution of the observable data. This paper shows that this phenomenon holds in a broader setting with a multi-valued outcome, treatment, and instrument, under an extension of the monotonicity condition that we refer to as generalized monotonicity. We further show that this phenomenon holds for any restriction on treatment response that is stronger than generalized monotonicity provided that these stronger restrictions do not restrict potential outcomes. Importantly, many models of potential treatments previously considered in the literature imply generalized monotonicity, including the types of monotonicity restrictions considered by kline2016evaluating, kirkeboen2016field, and heckman2018unordered, and the restriction that treatment selection is determined by particular classes of additive random utility models. We show through a series of examples that restrictions on potential treatments can provide identifying power beyond instrument exogeneity for average potential outcomes and average treatment effects when the restrictions imply that the generalized monotonicity condition is violated. In this way, our results shed light on the types of restrictions required for help in identifying average potential outcomes and average treatment effects. \end{abstract}

KEYWORDS: Multi-valued Treatments, Average Treatment Effects, Endogeneity, Instrumental Variables.

JEL classification codes: C31, C35, C36

\thispagestyle{empty} \setcounter{page}{1}

Introduction

In their analysis of a setting with a binary outcome, treatment, and instrument, balke1993nonparametric,balke1997bounds establish that the monotonicity condition of imbens1994identification has no identifying power beyond instrument exogeneity for average potential outcomes and the average treatment effect (ATE). Here, by no identifying power beyond instrument exogeneity, we mean that adding the monotonicity condition of imbens1994identification to instrument exogeneity does not decrease the identified sets for those parameters whenever those restrictions are consistent with the distribution of the observable data.\footnote{For settings with possibly non-binary outcomes, kitagawa2021identification shows this phenomenon continues to hold for any parameter that is a function of the marginal distributions of potential outcomes.} In this way, their results contrast with the analysis of imbens1994identification, who showed that their monotonicity condition and instrument exogeneity permitted identification of the local average treatment effect (LATE). This paper studies the extent to which this phenomenon holds in the broader context of a multi-valued outcome, treatment, and instrument.

We show that a generalization of the monotonicity condition of imbens1994identification to this richer setting also has no identifying power beyond instrument exogeneity for average potential outcomes and ATEs. We hereafter refer to this condition more succinctly as generalized monotonicity. We show further that this result remains true for any restriction that is in fact stronger than generalized monotonicity provided that these stronger restrictions do not restrict potential outcomes in a sense that we will make precise later. This feature of our results is remarkable because one might expect stronger restrictions, possibly with very complicated restrictions on potential treatments that are difficult to fully characterize, to reduce the identified set at least in some instances, but we show that this is not the case. Using this result, our analysis accommodates many examples of restrictions on potential treatments that have been previously considered in the literature. In particular, we show that encouragement designs, the types of monotonicity restrictions considered by kline2016evaluating, kirkeboen2016field, and heckman2018unordered, and certain additive random utility models, including some studied in lee2023treatment, all satisfy generalized monotonicity.

In establishing our results, we derive the identified sets for average potential outcomes under any such restriction and instrument exogeneity while maintaining the assumption that these restrictions are consistent with the distribution of the observable data. Our derivations reveal that the form of the resulting identified sets parallels the form of those derived by balke1993nonparametric,balke1997bounds for a binary outcome, treatment, and instrument. An implication of the form of the identified sets is that average potential outcomes and ATEs are only identified under an identification-at-infinity-type condition when imposing instrument exogeneity and any such restriction. In our analysis, we also derive the identified sets for average potential outcomes and ATEs when imposing instrument exogeneity alone whenever the distribution of the observable data is consistent with instrument exogeneity and generalized monotonicity. As we explain in Example (ref), this consistency is necessarily satisfied, for example, in the context of a multi-arm randomized controlled trial with one-sided non-compliance when defining the instrument to be random assignment to a given treatment arm.

Our results further provide necessary conditions on restrictions on potential treatments to help in identifying average potential outcomes and ATEs. See Theorem 3.3 and the subsequent discussion for details. We illustrate this phenomenon through a series of examples of models that need not satisfy generalized monotonicity and have identifying power for average potential outcomes and ATEs.

Our paper differs from the closely related literature that, in the context of a binary outcome, treatment, and instrument, considers the identifying power of the monotonicity condition of imbens1994identification and instrument exogeneity for the distribution (as opposed to the average) of potential outcomes, or considers the identifying power of these conditions when combined with additional restrictions on potential outcomes. In particular, kamat2019identifying shows that the monotonicity condition of imbens1994identification does have identifying power beyond instrument exogeneity for the (joint) distribution of potential outcomes. machado2019instrumental show that the monotonicity condition of imbens1994identification does have additional identifying power for the ATE beyond instrument exogeneity if one additionally imposes an assumption that requires potential outcomes to vary monotonically with the treatment. Thus, the phenomenon we explore is sensitive to both the choice of parameter and to whether one imposes assumptions on potential outcomes.

The remainder of the paper is organized as follows. Section (ref) introduces our formal setup, notation and assumptions, including our generalized monotonicity condition. Our main identification results are presented in Section (ref). In Section (ref), we provide several examples of restrictions on potential treatments that imply our generalized monotonicity condition, and are thus examples of restrictions that have no identifying power beyond instrument exogeneity for average potential outcomes or ATEs. In contrast, in Section (ref), we provide several examples of restrictions on potential treatments that imply that our generalized monotonicity condition does not hold, and further show that some of these restrictions in fact have identifying power beyond instrument exogeneity for average potential outcomes and ATEs. Proofs of all results can be found in the Appendix.

Setup and Notation

Denote by $Y \in \mathcal Y$ a multi-valued outcome of interest, by $D \in \mathcal D$ a multi-valued endogenous regressor, and by $Z \in \mathcal Z$ a multi-valued instrumental variable.\footnote{Our restriction to a multi-valued $Y$ facilitates exposition, but is not essential. At the expense of slightly more complicated arguments, we can accommodate more generally any real-valued $Y$.} To rule out degenerate cases, we assume throughout that $2 \leq |\mathcal Y| < \infty$, $2 \leq |\mathcal D| < \infty$, and $2 \leq |\mathcal Z| < \infty$. Further denote by $Y_d \in \mathcal Y$ the potential outcome if $D = d \in \mathcal D$ and by $D_z \in \mathcal D$ the potential treatment if $Z = z \in \mathcal Z$. We impose the usual consistency assumption,

equation[equation omitted — 156 chars of source]

Let $P$ denote the distribution of $(Y, D, Z)$ and $Q$ denote the distribution of $((Y_d : d \in \mathcal D),(D_z : z \in \mathcal Z), Z)$. Note that (ref) defines a mapping $T$ through

equation[equation omitted — 94 chars of source]

and therefore $P = QT^{-1}$. In what follows, we will say that a given $Q$ rationalizes a given $P$ if $P = QT^{-1}$.

Below we will require that $Q \in \mathbf Q$, where $\mathbf Q$ is a class of distributions satisfying assumptions that we will specify. Different choices of $\mathbf Q$ represent different assumptions that we impose on the distribution of potential outcomes and potential treatments. In this sense, $\mathbf Q$ may be viewed as a model for potential outcomes and potential treatments.

Given $P$ and a model $\mathbf Q$, the set of $Q \in \mathbf Q$ that can rationalize $P$ is \[ \mathbf Q_0(P, \mathbf Q) = \{Q \in \mathbf Q: P = Q T^{-1}\}~, \] i.e., the pre-image of $P$ under $T$. We say $\mathbf Q$ is consistent with $P$ if and only if $\mathbf Q_0(P, \mathbf Q ) \ne \emptyset$. We will start by considering models $\mathbf Q$ for which every $Q \in \mathbf Q$ satisfies

assumption[Instrument Exogeneity] $((Y_d : d \in \mathcal D),(D_z : z \in \mathcal Z)) \perp \!\!\! \perp Z$ under $Q$.

Our final result on the identifying power of generalized monotonicity will also apply to the weaker exogeneity restriction in richardson2013single that avoids the “cross-world” restrictions of Assumption (ref); see Assumption (ref) and Corollary (ref) below in Section (ref). If $Q$ satisfies Assumption (ref), then

equation[equation omitted — 134 chars of source]

Since the marginal distribution of $Z$ under $P$ and $Q$ are the same, i.e., for all $z \in \mathcal Z$, \[ P\{Z = z\} = Q\{Z = z\}~, \] $P = QT^{-1}$ if and only if (ref) holds. Thus, if all $Q \in \mathbf Q$ satisfies Assumption (ref), then $\mathbf Q_0(P, \mathbf Q)$ can be simplified as

equation[equation omitted — 200 chars of source]

Let $\theta(Q) = ( \mathbb{E}_Q[Y_d] : d \in \mathcal{D} )$ denote the vector of average potential outcomes. For fixed $P$ and $\mathbf Q$, the identified set for $\theta(Q)$ under $P$ relative to $\mathbf{Q}$ is given by

equation*[equation* omitted — 94 chars of source]

$\Theta_0(P, \mathbf Q) $ is nonempty whenever $\mathbf Q_0(P, \mathbf Q)$ is nonempty. By construction, this set is “sharp” in the sense that for any value in the set there exists $Q \in \mathbf Q_0(P, \mathbf Q)$ for which $\theta(Q)$ equals the prescribed value. The identified set for $\theta(Q)$ immediately implies that the identified set for any parameter $\lambda = \lambda(\theta)$ is given by $\lambda(\Theta_0(P, \mathbf Q))$. An important example is $\mathbb{E}_Q[Y_j]- \mathbb{E}_Q[Y_k]$, the ATE for treatment $j$ versus treatment $k$.

In the next section, we consider identification in any model of potential treatments that implies the following restriction on all $Q \in \mathbf{Q}$:

assumption[Generalized Monotonicity] For each $d \in \mathcal{D}$, there exists $z^\ast = z^{\ast}(d, Q) \in \mathcal{Z}$ such that \begin{equation} Q\{D_{z^\ast} \ne d , D_{z^{\prime}}=d for some z^{\prime} \ne z^\ast \}=0 .\end{equation}

In what follows, we refer to Assumption (ref) as generalized monotonicity. It states that under $Q$, for each treatment status $d \in \mathcal D$, there exists a value (possibly depending on $d$ and $Q$) of the instrument $z^\ast \in \mathcal Z$ that maximally encourages all individuals to $d$. Here, by “maximally encourage”, we mean that if an individual does not choose $d$ when $Z = z^\ast$, then they never choose $d$ for any other value of $Z$. Equivalently, if an individual chooses $d$ when $Z$ is equal to any value other than $z^\ast$, then they have to choose $d$ when $Z = z^\ast$. When $\mathcal{D}= \mathcal Z = \{0,1\}$, Assumption (ref) is equivalent to the monotonicity assumption of imbens1994identification.

We emphasize that Assumption (ref) only requires, for each possible value of the treatment, that there exists a value of the instrument that maximally encourages that treatment; it does not require that the value of the instrument is unique. For a given distribution $Q$ and given treatment $d \in \mathcal D$, let $\mathcal{Z}^\ast(d, Q)$ denote the set of $z^\ast$ that satisfy (ref). In this notation, Assumption (ref) can be restated as $\mathcal{Z}^\ast(d, Q) \neq \emptyset$ for each $d \in \mathcal{D}$. In the statement of Assumption (ref), $z^\ast(d, Q)$ is allowed to change across $Q$. The following lemma shows that $\mathcal{Z}^\ast(d, Q)$ is identified from $P$ and is hence the same for all $Q$ that rationalizes $P$ and satisfies Assumptions (ref) and (ref). In what follows, we will therefore write $\mathcal Z^\ast(d)$ and $z^\ast(d)$ whenever the given distribution $Q$ rationalizes $P$ and satisfies Assumptions (ref) and (ref). This result generalizes the corresponding result in imbens1994identification.

lemmaSuppose $Q$ satisfies Assumptions (ref) and (ref) and $P = Q T^{-1}$. Then, $z \in \mathcal Z^\ast(d, Q)$ if and only if \begin{equation} P \{ D = d \mid Z=z \} \ge P \{ D = d \mid Z=z' \} for all z' \in \mathcal Z . \end{equation}

Below we prove the necessity of (ref); sufficiency is established in the appendix. Note that, for any $d \in \mathcal{D}$, $z' \in \mathcal{Z}$, and any $z \in \mathcal Z^\ast(d, Q)$,

align*[align* omitted — 225 chars of source]

where the first and last equalities exploit Assumption (ref), and the third equality uses Assumption (ref).

Main Result

In order to describe our main result, we first introduce some further notation. Denote by $\mathbf Q_E^\ast$ (where $E$ stands for exogeneity) the set of all distributions that satisfy Assumption (ref) and by $\mathbf Q_{E,M}^\ast$ (where $M$ stands for generalized monotonicity) the set of all distributions that satisfy Assumptions (ref) and (ref). We will further require that the model does not restrict potential outcomes in the following sense:

assumption[Unrestricted Potential Outcomes] Let $Q \in \mathbf Q$ and $Q' \in \mathbf Q_E^\ast$. If the distributions of $(D_z: z \in \mathcal Z)$ under $Q$ and $Q'$ are the same, then $Q' \in \mathbf Q$.

In terms of this notation, our main result can be stated as follows:

theoremSuppose $\mathbf Q \subseteq \mathbf Q_{E,M}^\ast$ and $\mathbf Q$ satisfies Assumption (ref). Then, for any $P$ such that $\mathbf Q_0(P, \mathbf Q) \neq \emptyset$, we have $\Theta_0(P, \mathbf Q) = \Theta_0(P, \mathbf Q_{E,M}^\ast) = \Theta_0(P, \mathbf Q_E^\ast)$.

Theorem (ref) describes the sense in which restrictions on potential treatments stronger than generalized monotonicity have no identifying power for average potential outcomes and ATEs provided that these stronger restrictions do not restrict potential outcomes. This result is established through Theorems (ref) and (ref) below. Theorem (ref), developed in Section (ref), characterizes $\Theta_0(P, \mathbf Q)$ for any model $\mathbf Q$ that is stronger than Assumptions (ref) and (ref), i.e., $\mathbf Q \subseteq \mathbf Q_{E,M}^\ast$, and does not restrict potential outcomes in the sense of Assumption (ref). The result shows, in particular, that $\Theta_0(P, \mathbf Q) = \Theta_0(P, \mathbf Q_{E,M}^\ast)$ for any such model $\mathbf Q$ whenever $\mathbf Q_0(P, \mathbf Q) \neq \emptyset$. Remarkably, this result holds even if the model $\mathbf Q$ is strictly more restrictive than Assumptions (ref) and (ref) in the sense that $\mathbf Q \subsetneqq \mathbf Q_{E,M}^\ast$. On the other hand, Theorem (ref) and Corollary (ref), developed in Section (ref), show that $\Theta_0(P, \mathbf Q_{E,M}^\ast) = \Theta_0(P, \mathbf Q_E^\ast)$ whenever $\mathbf Q_0(P, \mathbf Q_{E,M}^\ast) \neq \emptyset$. Together, these results immediately imply Theorem (ref). In fact, Theorem (ref) shows the stronger result that if a submodel of instrument exogeneity and generalized monotonicity is consistent with $P$, then any model sandwiched between this submodel and the model that only assumes mean independence leads to the same identified set for average potential outcomes and ATEs. This observation allows us to establish that generalized monotonicity also has no identifying power for average potential outcomes and ATEs beyond the weaker exogeneity restriction of richardson2013single.

Identified Sets for $\mathbf Q \subseteq \mathbf Q_{E,M}^\ast$

For $d \in \mathcal D$ and $z \in \mathcal Z$, define $\beta_{d|z} = \mathbb{E}_P[ Y \mathbbm{1}\{D=d\} \mid Z=z]$. In addition, define $y^L = \min(\mathcal{Y})$ and $y^U = \max(\mathcal{Y})$. The following theorem derives the identified set for $\theta(Q)$, relative to any model that assumes instrument exogeneity and generalized monotonicity but does not restrict potential outcomes. Note in particular that the assumptions allow for $\mathbf Q \subsetneqq \mathbf Q_{E,M}^\ast$, in which case the model assumes strictly more than instrument exogeneity and generalized monotonicity.

theoremSuppose $\mathbf Q \subseteq \mathbf Q_{E,M}^\ast$ and $\mathbf Q$ satisfies Assumption (ref). Then, for any $P$ such that $\mathbf Q_0(P, \mathbf Q) \neq \emptyset$, \begin{equation} \Theta_0(P, \mathbf Q) = \prod_{d \in \mathcal{D}} \left[ \beta_{d|z^\ast(d)} + y^L (1- \sum_{y \in \mathcal{Y}} p_{yd|z^\ast(d)}) ) , \beta_{d|z^\ast(d)} + y^U (1- \sum_{y \in \mathcal{Y}} p_{yd|z^\ast(d)}) ) \right] . \end{equation}

We now describe some intuition for Theorem (ref). Note that the distribution of the data, $P$, only contains information on the distribution of $Y_d$ for those individuals who would take treatment $d$ for some value of the instrument. On the other hand, it contains no information on the distribution of $Y_d$ for those individuals who would not take that treatment for any value of the instrument. Assumption (ref) implies that individuals would take treatment $d$ at some value of the instrument if and only if $D_{z^{\ast}(d)}=d$, i.e., when maximally encouraged to do so. Assumption (ref) implies $ \beta_{d|z^\ast(d)} = \mathbb{E}_Q[ \mathbbm{1}\{D_{z^{\ast}(d)}=d\} Y_d ] $, and thus captures all the information from $P$ relevant to $\mathbb{E}_Q[Y_d]$, which is the first part of the lower and upper bounds in (ref). In contrast, Assumptions (ref) and (ref) imply that the probability that an individual would not take treatment $d$ for any value of $Z$ is identified from $P$ to be $Q\{D_{z^{\ast}(d)}\ne d\} = 1- \sum_{y \in \mathcal{Y}} p_{yd|z^\ast(d)},$ but $P$ contains no information on the distribution of $Y_d$ for such individuals. Furthermore, that $\mathbf Q$ satisfies Assumption (ref) implies that the model does not restrict the distribution of $Y_d$ for such individuals beyond $y \in \mathcal{Y}$, so that we can set $Y_d$ to be any value between $y^L$ and $y^U$ for these individuals, which constitutes the second part of the upper and lower bounds in (ref).

remarkUnder the instrument exogeneity and monotonicity assumptions of imbens1994identification, balke1993nonparametric,balke1997bounds found the same form of the identified set for $\theta(Q)$ as (ref) when $\mathcal{Y}=\mathcal{D}=\mathcal{Z}=\{0,1\}$. Theorem (ref) therefore generalizes the result of balke1993nonparametric,balke1997bounds to more than two treatment arms and instrument values, to outcomes taking more than two values, and, more surprisingly, to show that the same identified set holds when imposing possibly stronger restrictions on potential treatments than generalized monotonicity.

Theorem (ref) immediately implies the following result on the identified sets for the ATE of treatment $j$ versus $k$:

corollaryUnder the assumptions of Theorem (ref), the identified set for $\mathbb E_Q[Y_j -Y_k]$ is given by: \begin{multline} \biggl[ (\beta_{j|z^\ast(j)}- \beta_{k|z^\ast(k)}) + (y^L -y^U) + y^U \sum_{y \in \mathcal{Y}} p_{yk|z^\ast(k)} - y^L \sum_{y \in \mathcal{Y}} p_{yj|z^\ast(j)} , \\ (\beta_{j|z^\ast(j)}- \beta_{k|z^\ast(k)}) + (y^U-y^L) + y^L \sum_{y \in \mathcal{Y}} p_{yk|z^\ast(k)} - y^U \sum_{y \in \mathcal{Y}} p_{yj|z^\ast(j)} \biggr] . \end{multline}

e

remarkFrom Corollary (ref), the width of the identified set for $\mathbb E_Q[Y_j -Y_k]$ under the assumptions of Theorem (ref) is given by $$(y^U-y^L) \left(P\{D \ne j \mid Z= z^\ast(j)\} + P\{D \ne k \mid Z= z^\ast(k)\} \right)~,$$ which, following Lemma (ref), equals $$(y^U-y^L) \left( \min_{z \in \mathcal{Z}} P\{D \ne j \mid Z= z\} +\min_{z \in \mathcal{Z}} P\{D \ne k \mid Z= z\} \right)~. $$ Therefore, when imposing generalized monotonicity, as well as when imposing any restriction implying generalized monotonicity, the ATE of $j$ versus $k$ is only identified “at infinity” heckman1990varieties,andrews1998semiparametric in the sense that identification requires \begin{equation} \min_{z \in \mathcal{Z}} P\{D \ne j \mid Z= z\}= \min_{z \in \mathcal{Z}} P\{D \ne k \mid Z= z\}=0 . \end{equation} In other words, identification of the the ATE of $j$ versus $k$ under generalized monotonicity or under any restriction implying generalized monotonicity requires that there is some value of the instrument such that everyone takes treatment $j$ at that value of the instrument, and some value of the instrument such that everyone takes treatment $k$ at that value of the instrument. In contrast, by imposing restrictions on potential treatments that imply that generalized monotonicity is violated, the ATE can sometimes be identified without (ref) even when potential outcomes are unrestricted; see Example (ref) in Section (ref) below.

Identifying Power of Generalized Monotonicity

Theorem (ref) above establishes that, for possibly multi-valued $Y$, $D$, and $Z$, and any model $\mathbf Q \subseteq \mathbf Q_{E,M}^\ast$ such that $\mathbf Q$ does not restrict the potential outcomes, if $\mathbf Q_0(P, \mathbf Q) \neq \emptyset$, then $\Theta_0(P, \mathbf Q)$ equals (ref). We now show that $\Theta_0(P, \mathbf Q_{E,M}^\ast) = \Theta_0(P, \mathbf Q_E^\ast)$ as long as $\mathbf Q_0(P, \mathbf Q_{E,M}^\ast) \neq \emptyset$, so that the identified set for $\theta(Q)$ assuming instrument exogeneity and generalized monotonicity coincides with the identified set assuming instrument exogeneity alone, as long as both assumptions are consistent with the distribution of the observed data. Our result therefore generalizes balke1993nonparametric,balke1997bounds, which study the case of binary $Y$, $Z$, and $D$.

In order to do so, we consider a mean independence assumption even weaker than Assumption (ref), and show the identified set under this even weaker assumption is also (ref). A sandwich argument will then lead to our desired result. In particular, we first establish that the identified set for $\theta(Q)$ in (ref) coincides with the identified set under the weaker mean independence assumption considered by robins1989analysis and manski1990nonparametric:

assumption[Mean Independence] $\mathbb{E}_Q[Y_d \mid Z=z ]= \mathbb{E}_Q[Y_d]$ for all $d \in \mathcal{D}$ and $z \in \mathcal{Z}$.

Note Assumption (ref) is weaker than instrument exogeneity in Assumption (ref), and does not imply (ref). Let $\mathbf Q_{\it MI}^\ast$ denote the set of all $Q$ that satisfies Assumption (ref) (where $MI$ stands for mean independence). Following robins1989analysis and manski1990nonparametric, the following lemma derives the identified set for $\theta(Q)$ under mean independence:

lemmaSuppose $\mathbf Q_0(P, \mathbf Q_{\it MI}^\ast) \neq \emptyset$. Then, \begin{equation} \Theta_0(P, \mathbf Q_{\it MI}^\ast) = \prod_{d \in \mathcal{D} } \left[ \max_{z \in \mathcal{Z}} \{ \beta_{d|z} + y^L (1- \sum_{y \in \mathcal{Y}} p_{yd \mid z} ) \} , \min_{z \in \mathcal{Z}} \{ \beta_{d|z} + y^U (1- \sum_{y \in \mathcal{Y}} p_{yd \mid z} ) \} \right]. \end{equation}

The following lemma, which relies on the observation in Lemma (ref), establishes the equivalence between the identified sets in (ref) and (ref) when $\mathbf Q_0(P, \mathbf Q_{E,M}^\ast) \neq \emptyset$.

lemmaSuppose $\mathbf Q \subseteq \mathbf Q_{E, M}^\ast$, $\mathbf Q$ satisfies Assumption (ref), and $P$ is such that $\mathbf Q_0(P, \mathbf Q) \neq \emptyset$. Then, the sets in (ref) and (ref) coincide.

Using Lemma (ref), we are able to establish our desired result, which asserts that, maintaining Assumption (ref), additionally imposing Assumption (ref) either causes the identified set for $\theta(Q)$ to become empty (if those assumptions are not consistent with $P$) or leaves the identified set for $\theta(Q)$ unchanged (if those assumptions are consistent with $P$). In fact, we will establish a stronger result, that if a submodel of instrument exogeneity and generalized monotonicity is consistent with $P$, then any model sandwiched between this submodel and the model that only assumes mean independence leads to the same identified set for $\theta(Q)$.

theoremSuppose $\mathbf Q \subseteq \mathbf Q_{E, M}^\ast$ and $\mathbf Q$ satisfies Assumption (ref). Further suppose $\mathbf Q'$ satisfies $$\mathbf Q \subseteq \mathbf Q' \subseteq \mathbf Q_{\it MI}^\ast~.$$ Then, for any $P$ such that $\mathbf Q_0(P, \mathbf Q) \neq \emptyset$, $$\Theta_0(P, \mathbf Q) = \Theta_0(P, \mathbf Q') = \Theta_0(P, \mathbf Q_{\it MI}^\ast)~.$$
remarkTheorem (ref) implies that in order for a model to have identifying power for average potential outcomes, it has to be the case that the model does not contain a submodel of instrument exogeneity and generalized monotonicity that is consistent with $P$. In other words, the model has to contradict Assumption (ref) or (ref). We illustrate this observation in Example (ref) below.
corollaryFor any $P$ such that $\mathbf Q_0(P, \mathbf Q_{E,M}^\ast) \neq \emptyset$, $\Theta_0(P, \mathbf Q_{E,M}^\ast) = \Theta_0(P, \mathbf Q_E^\ast)$.
remarkAn implication of Corollary (ref) is that the identified set for $\theta(Q)$ under Assumption (ref) alone will be (ref), regardless of whether Assumption (ref) is imposed, as long as the distribution of the data is consistent with Assumptions (ref) and (ref).

Theorem (ref) further implies that under the following weaker exogeneity assumption, generalized monotonicity also has no identifying power for average potential outcomes and ATEs:

assumption[Weak Instrument Exogeneity] Under $Q$, $(Y_d, D_z ) \perp \!\!\! \perp Z$ for all $ d \in \mathcal D, z \in \mathcal Z$.

See richardson2013single for an analysis of alternative exogeneity restrictions, and in particular how Assumption (ref) avoids the “cross-world” restrictions of the stronger joint independence in Assumption (ref). Denote by $\mathbf Q_{\it WE}^\ast$ the set of all distributions that satisfy Assumption (ref) and $\mathbf Q_{\it WE, M}^\ast$ the set of all distributions that satisfy Assumptions (ref) and (ref).

corollaryFor any $P$ such that $\mathbf Q_0(P, \mathbf Q_{\it WE,M}^\ast) \neq \emptyset$, $\Theta_0(P, \mathbf Q_{\it WE,M}^\ast) = \Theta_0(P, \mathbf Q_{\it WE}^\ast)$.

Examples of Models That Satisfy Assumption (ref)

We now consider some restrictions on potential treatments that have been considered previously in the literature. In each case, we show $\mathbf Q \subseteq \mathbf Q_{E,M}^\ast$; in particular, these restrictions satisfy generalized monotonicity. We emphasize that frequently $\mathbf Q \subsetneqq \mathbf Q_{E,M}^\ast$. In what follows, it is implicitly understood that Assumption (ref) is satisfied, so that the model imposes no restriction on potential outcomes. Thus, in each example Theorem (ref) applies and such restrictions do not provide any identifying power for average potential outcomes or ATEs. In Appendix C, we consider three additional examples: an RCT with a “close substitute” as considered in kline2016evaluating; the monotonicity and “irrelevance” assumptions considered in kirkeboen2016field; and an additive random utility model for a binary treatment.

exampleConsider a multi-arm randomized controlled trial (RCT) with noncompliance, where $Z=d$ denotes random assignment to treatment $d$, $D_d=d$ denotes that the subject would comply with assignment if assigned to treatment $d$, and $|\mathcal D|= |\mathcal Z|$. More generally, not necessarily in the context of an RCT, one can interpret $Z=d$ as encouragement to treatment $d$ and interpret $D_d=d$ as the subject would take treatment $d$ if encouraged to do so. In this example, $Q$ satisfies Assumption (ref) because $Z$ is randomly assigned. We may generalize the “no-defier” restriction of angrist1996identification as: for each $d \in \mathcal D$, \[ Q\{D_{d} \ne d , ~D_{d^{\prime}}=d ~ \mbox{for some} ~ d^{\prime} \ne d \}=0~,\] i.e., there is zero probability that a subject would not take treatment $d$ if assigned to (encouraged to take) $d$ but would take $d$ if assigned (encouraged) to some other treatment $d^{\prime} \ne d$. If $Q$ satisfies this generalized no-defier restriction, then Assumption (ref) holds with $z^\ast(d)=d$ for all $d$. This no-defier restriction in particular holds in the context of an RCT with “one-sided non-compliance,” where we assume \[ Q\{D_z \in \{0,z\}\}=1~, \] for all $z \in \mathcal Z$. Here, non-compliance is one-sided because one can fall back to the control group if assigned to $d$ but cannot choose $d$ if assigned to the control group.
remarkAs explained in Example (ref), in a multi-arm RCT with one-sided noncompliance, $P$ will necessarily be consistent with Assumptions (ref) and (ref), i.e., $\mathbf Q(P,\mathbf Q_{E,M}) \neq \emptyset$. Following Remark (ref), Theorem (ref) therefore implies that the identified set for $\theta(Q)$ under Assumption (ref) alone will be (ref) for such a multi-arm RCT. A further implication is that the identified set under Assumption (ref) on $\mathbb{E}[Y_j-Y_k]$ for a multi-arm RCT with one-sided noncompliance depends only on the treatment arms for random assignment to treatments $j$ and $k$.
examplecheng2006bounds consider an RCT with noncompliance where $\mathcal{D}=\mathcal{Z}=\{0,1,2\}$. Assumption (ref) continues to hold because $Z$ is randomly assigned. In such a setting, they develop bounds on average effects within subgroups defined by potential treatments which, following the terminology of frangakis2002principal, they call “principal strata.” While bai2025inference derives the identified sets for those parameters given their assumptions, we now use our analysis to consider instead identification of average potential outcomes and ATEs given their assumptions. Their “Monotonicity I” assumption is equivalent to one-sided noncompliance in the preceding example. Their “Monotonicity II” assumption states that subjects who would comply with assignment to treatment $2$ would also comply with assignment to treatment $1$, so that \[ Q\{ D_1=1 \mid D_2=2\}=1~. \] They argue that such an assumption is plausible in a medical context when treatment $1$ has fewer side effects than treatment $2$, and in their application to treatments for alcohol dependence, in which complying with treatment $1$ (compliance enhancement therapy) requires less effort by subjects than complying with treatment $2$ (cognitive behavioural therapy). Because their Monotonicity I restriction implies our Assumption (ref), so does imposing both their Monotonicity I and II restrictions.
exampleSuppose $Q$ satisfies Assumption (ref). heckman2018unordered define “unordered monotonicity" as the assumption that, for any $d \in \mathcal{D}$, and any $z, z^{\prime} \in \mathcal{Z}$, \begin{equation} Q \{\mathbbm{1}\{D_z=d\} \ge \mathbbm{1}\{D_{z^{\prime}}=d\}\} = 1 or Q \{\mathbbm{1}\{D_z=d\} \le \mathbbm{1}\{D_{z^{\prime}}=d\}\} = 1 . \end{equation} Assumption (ref) holds for any $Q$ that satisfies (ref). To see this, note that Assumption (ref) can be expressed as the requirement that for each $d \in \mathcal D$, there exists $z^\ast(d) \in \mathcal Z$ such that \[ Q \{\mathbbm 1 \{D_{z^\ast(d)} = d\} \geq \mathbbm 1 \{D_z = d\}\} = 1 \text{ for all } z \in \mathcal Z~, \] which is immediately implied by (ref). Note, however, that $\mathcal Z^\ast(d)$ may not be a singleton unless some inequalities in (ref) are strict.
remarkAlthough unordered monotonicity implies Assumption (ref), the converse is generally false. For example, suppose $\mathcal{Z}=\{0,1,2,3\}$ and $\mathcal{D}=\{0,1\}$. Suppose $ \mathbbm{1}\{D_3=1\} \ge \mathbbm{1}\{D_{z}=1\} $ w.p.1 under $Q$ for $z \ne 3$ and $ \mathbbm{1}\{D_0=0\} \ge \mathbbm{1}\{D_{z}=0\} $ w.p.1 under $Q$ for $z \ne 0$, but $Q\{ D_1=1, D_2=0\}$ and $Q\{ D_1=0, D_2=1\}$ are both strictly positive. Then Assumption (ref) holds with $z^\ast(0) = 0$ and $z^\ast(1) = 3$, but unordered monotonicity fails. In particular, $\mathbbm{1} \{D_1 = 1\}$ and $\mathbbm{1} \{D_2 = 1\}$ are not ordered, thus violating (ref). We thus conclude that if $\mathbf Q$ is defined as the set of distributions that satisfy instrument exogeneity and unordered monotonicity, then $\mathbf Q \subsetneqq \mathbf Q_{E,M}^\ast$.
exampleSuppose under $Q$, $(D_z: z \in \mathcal Z)$ is determined by \begin{equation} D_z = \operatorname*{argmax}_{d \in \mathcal{D}} ( g(z, d) + U_{d} ) , \end{equation} for $g: \mathcal Z \times \mathcal D \to \Re$ where $\Re$ is the set of real numbers and a random vector $(U_d: d \in \mathcal D)$, whose distribution is absolutely continuous with respect to the Lebesgue measure on $\Re^{|\mathcal D|}$ and $Z \perp \!\!\! \perp ((U_d: d \in \mathcal D), (Y_d : d \in \mathcal D))$. Hence, $Q$ satisfies Assumption (ref) by construction. Let $\mathbf{Q}$ denote the set of distributions that are consistent with $(D_z: z \in \mathcal Z)$ being determined by (ref) for some $g$ and $(U_d: d \in \mathcal D)$ satisfying these requirements. The model $\mathbf Q$ is called an additive random utility model (ARUM). A sufficient condition for $Q \in \mathbf Q$ to satisfy Assumption (ref) is that for each $d \in \mathcal D$ there exists $z^\ast(d) \in \mathcal Z$ such that \begin{equation} g(z^\ast(d), d) - g(z^\ast(d), d') > g(z, d) - g(z, d') for all d' \neq d and z \neq z^\ast(d) . \end{equation} We refer to the requirement in (ref) as uniform targeting of treatment $d$. The terminology is intended to reflect that there is a value of the instrument that maximizes the gains (in terms of $g$) of choosing treatment $d$ versus any other treatment $d'$ uniformly across these other possible values of the treatment. In this sense, that value of the instrument targets treatment $d$ uniformly. We now argue by contradiction that (ref) implies (ref); hence, if (ref) holds for all $d \in \mathcal D$, then $Q$ satisfies Assumption (ref). To this end, suppose that, with positive probability, $D_{z^\ast(d)} = d' \neq d$ but $D_{z'} = d$ for $z' \neq z^\ast(d)$. Then, \begin{align*} g(z^\ast(d), d') + U_{d'} & \geq g(z^\ast(d), d) + U_d , \\ g(z', d) + U_d & \geq g(z', d') + U_{d'} . \end{align*} These two inequalities imply \[ g(z^\ast(d), d) - g(z^\ast(d), d') \leq g(z', d) - g(z', d')~, \] which violates (ref). A particular example when $|\mathcal Z| \geq |\mathcal D|$ that satisfies (ref) with $z^\ast(d)=d$ after a suitable relabelling is \begin{equation} g(z, d) = \alpha_d +\beta_d \mathbbm 1 \{z = d\} , \end{equation} with $\beta_d > 0$, so that $Z=d$ strictly increases the latent value of treatment $d$ while leaving the values of the remaining options unchanged.
remarkThe ARUM for a binary treatment is equivalent to the Heckman-Vytlacil nonparametric selection model for a binary treatment considered, e.g., in heckman1999local,heckman2005structural, and shown by vytlacil2002independence to be equivalent to the monotonicity and exogeneity assumptions of imbens1994identification. heckman2001instrumental show that the Heckman-Vytlacil nonparametric selection model for a binary treatment results in an identified set for the ATE of the form in (ref) and that the model has no identifying power beyond instrument exogeneity for the ATE. Example C.1 in Appendix C shows that an ARUM for a binary treatment will satisfy Assumption (ref) and therefore the results of this paper nest the results of heckman2001instrumental. Example (ref), on the other hand, extends their results to nonparametric selection models for a multi-valued treatment. For a partial identification analysis of a class of parameters that includes the ATE under a nonparametric selection model for a binary treatment and sometimes imposing additional restrictions, see, e.g., mogstad2018using, han2024computational, and marx2024sharp.
examplelee2023treatment also consider the ARUM defined by (ref) without imposing the uniform targeting of (ref) for each treatment. Instead, they impose an assumption that they refer to as “strict one-to-one targeting”, in which the set of treatments can be partitioned into a set of treatments $\mathcal{D}^\dagger$ that are “targeted” and a set of treatments $\mathcal{D} \setminus \mathcal{D}^\dagger $ that are “not targeted” such that \begin{enumerate} • For $d \in \mathcal{D} \setminus \mathcal{D}^\dagger$, $g(z, d)$ is the same for all $z \in \mathcal Z$; • For $d \in \mathcal{D}^\dagger$, there exists $z^\dagger(d)$ such that $g(z^\dagger(d), d) > g(z', d)$ for all $z' \neq z^\dagger(d)$ and such that $g(z', d)$ takes the same value for all $z' \neq z^\dagger(d)$; additionally, $z^\dagger(d) \neq z^\dagger(d')$ for $d, d^{\prime} \in \mathcal{D}^\dagger$, $d\ne d'$. \end{enumerate} The terminology “one-to-one” stems from the second requirement above. They further impose that there exists a treatment that is known to be non-targeted. This class of ARUMs is equivalent to imposing (ref) for targeted treatments, imposing $g(z, d)=\alpha_d$ for non-targeted treatments, and imposing that there is at least one non-targeted treatment. In such a setting, lee2023treatment analyze the identification of a particular class of average effects within subgroups defined by potential treatments. We now use our analysis to consider instead the identification of average potential outcomes and ATEs given their assumptions. Suppose strict one-to-one targeting holds, and additionally suppose that there is at least one targeted treatment. In Appendix B.1, we argue that Assumption (ref) holds when $|\mathcal{Z}| > |\mathcal{D}^\dagger|$, so that there are more values of the instrument than there are targeted treatments. We further argue that (ref) does not hold for some treatments unless $|\mathcal{D}|=|\mathcal{Z}|=2$. Such models therefore provide another class of ARUMs, distinct from the one with uniform targeting described in Example (ref), for which Assumption (ref) holds. In Example (ref) below, we show, however, that Assumption (ref) does not hold when $|\mathcal D| \geq 3$, $|\mathcal{Z}| = |\mathcal{D}^\dagger|$, and the support of $(U_d: d \in \mathcal D)$ is $\Re^{|\mathcal D|}$.

Examples of Models That Do Not Satisfy Assumption (ref)

We now consider models that do not satisfy generalized monotonicity (Assumption (ref)). For each model, we show that the identified sets for average potential outcomes are not given by (ref). For the first two examples, we further show that they do in fact provide identifying power beyond instrument exogeneity.

exampleSuppose $\mathcal{Y}=\mathcal{D}=\mathcal{Z}=\{0,1\}$. Let $\mathbf{Q}$ denote all distributions $Q$ that satisfy Assumption (ref) and \[ Q\{ D_0 = D_1\} =0~, \] which, in the language of angrist1996identification, is imposing that all individuals are either compliers or defiers. For any $P$ such that $\mathbf{Q}_0(P, \mathbf{Q}) \neq \emptyset$, $\theta(Q)$ is identified relative to $\mathbf Q$, i.e., $\Theta_0(P, \mathbf Q)$ is a singleton. To see this, note that for any $Q\in \mathbf{Q}_0(P, \mathbf{Q})$ and $y \in \mathcal Y$, \begin{align*} Q \{Y_1 = y\} & = Q \{Y_1 = y, D_0 = 0, D_1 = 1\} + Q \{Y_1 = y, D_0 = 1, D_1 = 0\} \\ & = Q \{Y_1 = y, D_1 = 1\} + Q \{Y_1 = y, D_0 = 1\} \\ & = Q \{Y_1 = y, D_1 = 1 \mid Z = 1\} + Q \{Y_1 = y, D_0 = 1 \mid Z = 0\} \\ & = P \{Y = y, D = 1 \mid Z = 1\} + P \{Y = y, D = 1 \mid Z = 0\} , \end{align*} where the first two equalities follows from $Q \{D_0 = D_1\} = 0$, the third equality follows from Assumption (ref), and the final equality follows from $Q \in \mathbf{Q}_0(P, \mathbf{Q})$. A similar argument establishes identification of $Q\{Y_0 = y\}$. In contrast, we show in Appendix B.2 that there exists a $P$ for which $\mathbf Q_0(P,\mathbf Q) \neq \emptyset$ and (ref) is not a singleton. The identified set for $\theta(Q)$ is therefore not given by (ref). We further show in Appendix B.2 that $\Theta_0(P, \mathbf Q_E^\ast)$ is not a singleton for the same $P$. Thus, the model $\mathbf Q$ does have identifying power beyond instrument exogeneity for $\theta(Q)$. Furthermore, recall as discussed in Remark (ref) that if $\mathbf Q$ has identifying power for $\theta(Q)$, then it cannot contain a submodel satisfying instrument exogeneity and generalized monotonicity that is consistent with $P$. In this example, the model $\mathbf{Q}$ contains such a submodel if and only if $P$ satisfies either (i) $P\{D=0\mid Z=0\}=P\{D=1\mid Z=1\}=1$ (in which case all individuals are compliers) or (ii) $P\{D=1\mid Z=0\}=P\{D=0\mid Z=1\}=1$ (in which case all individuals are defiers). In order for $\mathbf Q$ to have identifying power for $\theta(Q)$, it therefore must be the case that $0 < P\{D=0\mid Z=0\}, P\{D=0\mid Z=1\}< 1$, which is indeed satisfied by the counterexample in Appendix B.2. Further note that (ref) is a singleton in this example if and only if $P$ satisfies either (i) or (ii).
exampleConsider an ordered choice model for treatment. Suppose that $|\mathcal{Z}|\ge 3$, and let $\mathbf{Q}$ denote the set of all distributions that satisfy Assumption (ref) and \begin{equation} Q\{D_j \ge D_k\} =1 for all j \ge k . \end{equation} For example, $D$ might represent quantity of some treatment, and $Z$ might represent levels of subsidy for the treatment. The restriction in (ref) is equivalent to the monotonicity assumption considered in angrist1995two. See vytlacil2006ordered for the connection between this restriction and ordered discrete-choice selection models. Without loss of generality, let $\mathcal D = \{0, \ldots, \bar D\}$. In this case, (ref) is satisfied for $d \in \{0,\bar{D}\}$ for all $Q \in \mathbf{Q}_0(P, \mathbf{Q})$; to see this, take $z^\ast(0)=\min\{\mathcal{Z}\}$ and $z^\ast(\bar{D})=\max\{\mathcal{Z}\}$. By a straightforward modification of the arguments underlying Theorem (ref), one can show that the identified sets for $\mathbb{E}_Q[Y_0]$ and $\mathbb{E}_Q[Y_{\bar{D}}]$ are given by (ref) for any $P$ such that $\mathbf{Q}_0(P, \mathbf{Q}) \neq \emptyset$. The ordered monotonicity assumption in (ref) therefore has no identifying power beyond instrument exogeneity for $\mathbb{E}_Q[Y_0]$ and $\mathbb{E}_Q[Y_{\bar{D}}]$. In contrast, (ref) need not hold for $d \in \mathcal{D} \setminus \{0,\bar{D}\}$ and $Q \in \mathbf{Q}_0(P, \mathbf{Q})$. In Appendix B.3, we show there exists a $P$ for which $\mathbf Q_0(P, \mathbf Q) \neq \emptyset$ and $\Theta_0(P, \mathbf Q)$ is not given by (ref). We further show $\Theta_0(P, \mathbf Q) \subsetneqq \Theta_0(P, \mathbf Q_E^\ast)$ for the same $P$. Thus, the ordered monotonicity assumption in (ref) does have identifying power beyond instrument exogeneity for $\mathbb{E}_Q[Y_d]$ for $d \in \mathcal{D} \setminus \{0,\bar{D}\}$.
exampleIn Example (ref), we considered ARUMs satisfying the strict one-to-one targeting assumption of lee2023treatment. As discussed there, if $|\mathcal D| = 2$ or if $|\mathcal{D}| \ge 3$ and $|\mathcal{Z}| > |\mathcal{D}^\dagger|$, then Assumption (ref) holds and the results in Section (ref) are applicable. Now consider the case in which $|\mathcal{D}| \ge 3$ and $|\mathcal{Z}| = |\mathcal{D}^\dagger|$. Denote by $\mathbf Q$ the ARUM model defined by (ref) under the additional assumption that the support of $(U_d: d \in \mathcal D)$ is $\Re^{|\mathcal{D}|}$. Then, as we show in Appendix B.4, while (ref) will hold for targeted treatments, (ref) cannot hold for any non-targeted treatment, and thus Assumption (ref) is violated. We show that, while the identified set for $\mathbb{E}_Q[Y_d]$ is given by (ref) for targeted treatments when $\mathbf Q_0(P, \mathbf Q) \neq \emptyset$, there exists a $P$ for which $\mathbf Q_0(P, \mathbf Q) \neq \emptyset$ and the identified set for $\mathbb{E}_Q[Y_d]$ is not given by (ref) for non-targeted treatments.

Disclosure Statement

The authors report there are no competing interests to declare.