Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Partial Identification in Nonseparable Binary Response Models with Endogenous Regressors We are grateful to James Heckman, Marc Henry, Roger Koenker, and to seminar audiences at Columbia University and Michigan State University for helpful feedback. We also thank Martin Weidner and the organizers of the Chamberlain Seminar, and are grateful to Florian Gunsilius, Sukjin Han, Wayne Gao, and Takuya Ura for their questions and feedback, and to Adam Rosen for his thoughtful discussion. Jiaying Gu acknowledges financial support from the Social Sciences and Humanities Research Council of Canada. All errors are our own.
abstractThis paper considers (partial) identification of a variety of counterfactual parameters in binary response models with possibly endogenous regressors. Our framework allows for nonseparable index functions with multi-dimensional latent variables, and does not require parametric distributional assumptions. We leverage results on hyperplane arrangements and cell enumeration from the literature on computational geometry in order to provide a tractable means of computing the identified set. We demonstrate how various functional form, independence, and monotonicity assumptions can be imposed as constraints in our optimization procedure to tighten the identified set. Finally, we apply our method to study the effects of health insurance on the decision to seek medical treatment.
Keywords: Binary Choice, Counterfactual Probabilities, Endogeneity, Hyperplane Arrangement, Linear Programming, Partial Identification
\thispagestyle{empty}
Introduction
This paper considers partial identification of counterfactual parameters in a class of binary response models of the form:
align[align omitted — 128 chars of source]
where $Y \in \{0,1\}$ is a binary outcome variable, $X \in \mathcal{X}\subset \mathbb{R}^{d_{x}}$ is vector of (possibly endogenous) covariates with finite support, $U \in \mathcal{U}:=\mathbb{R}^{d_{u}}$ is a vector of latent variables, $\theta \in \Theta \subset \mathbb{R}^{d_{\theta}}$ is a vector of structural parameters, and $\tilde{\varphi}_{1}:\mathcal{X}\times \Theta \to \mathbb{R}^{d_{u}}$ and $\tilde{\varphi}_{2}: \mathcal{X}\times \Theta \to \mathbb{R}$ are known functions. Our approach does not require any parametric distributional assumptions on the latent variables. We then consider counterfactuals that can be expressed as a function $\gamma: \mathcal{X} \to \mathcal{X}$ that reassigns each $x \in \mathcal{X}$ to a new value $\gamma(x) \in \mathcal{X}$. Our main focus is on bounding linear functionals of counterfactual probabilities.
In our setting, nonparametric point-identification of the distribution of latent variables occurs only under restrictive conditions, including strong independence assumptions and large support conditions.\footnote{E.g. ichimura1998maximum.} Control function approaches are often used to address the issue of endogenous regressors, but if endogenous regressors are discrete or the mechanism generating the endogenous regressors is poorly understood, then many of these approaches are not applicable. Partial identification arises as a natural alternative to methods for point-identification in the presence of endogenous and discrete covariates. This paper uses an optimization-based approach to bound counterfactual quantities, which allows researchers to easily construct sharp bounds on counterfactual quantities under a variety of different assumptions by simply altering the constraints in our optimization problems.
In the absence of parametric distributional assumptions, our analysis reveals the importance of a special partition of the latent variable space into response types that have identical responses in all possible counterfactuals. We are not the first to emphasize the importance of response types, and our discussion echoes the insights of balke1994counterfactual and heckman2018unordered, among others. Similar to these works, we show that functional form and monotonicity assumptions amount to assigning zero probability to certain response types. Before bounding counterfactual quantities, it is thus necessary to determine which response type are possible/impossible in a given model. In our particular class of models, the latent variable space admits a partition into cells defined by a collection of hyperplanes, each cell corresponding to a unique response type. Using the cell enumeration algorithm of gu2020nonparametric, we show how to enumerate all response types in a time polynomial in the number of input hyperplanes. After enumerating response types, we then demonstrate how various counterfactual quantities can be bounded by solving a sequence of linear programming problems. We also study the case when the index function is linear in parameters, in which case we show how sharp bounds on counterfactual probabilities can be computed without the need to grid over the entire parameter space. Finally, we present a consistency result, we show how to adapt the inference procedure of cho2021simple to our setting, and we apply our method to study the effects of private health insurance on the decision to seek medical treatment.
A Review of the Relevant Literature
Binary response models with endogenous regressors have been studied extensively. Point identified approaches include linear probability models, maximum likelihood estimation (e.g.\! the bivariate probit), control function approaches, or approaches based on special regressors. All of these approaches have well-documented limitations.\footnote{See lewbel2012comparing for a review. }
Nonparametric identification was studied in binary choice and threshold crossing models by matzkin1992nonparametric, and in more general nonseparable models by matzkin2003nonparametric and chernozhukov2005iv, among many others. vytlacil2007dummy studied nonparametric identification of the average treatment effect in a discrete triangular system with a binary endogenous variable under a weak separability assumption in the outcome equation. Since then a number of paper have studied point identification in similar models with discrete endogenous variables (e.g. han2017identification, vuong2017counterfactual, chen2020identification and khan2021informational) and continuous endogenous variables (e.g. imbens2009identification, d2015identification, torgovitsky2015identification). Important precedents to the work presented here from the literature on point identification in random coefficient models include ichimura1998maximum, gautier2013nonparametric and gu2020nonparametric. However, these papers focus almost exclusively on the point-identified case with a linear index function and exogenous covariates with large support.
Many authors have also used partial identification methods to relax the assumptions required for point-identification in these models. In a relevant series of papers, chesher2013instrumental and chesher2014instrumental show how to use random set theory to characterize the identified set of structures in discrete choice models.\footnote{The general formulation of their approach is presented in chesher2017generalized.} Similar to the current paper, they do not provide a model for the endogenous explanatory variables, rendering the discrete choice model incomplete. They construct bounds on structural parameters using a characterization of the sharp set of constraints based on the results of artstein1983distributions.\footnote{See also norberg1992existence and molchanov2017theory Corollary 1.4.11.}
Although our paper focuses primarily on computational issues that arise when bounding counterfactual parameters, we present a comparison of our approach with the approach of chesher2013instrumental and chesher2014instrumental in Appendix (ref).
Our approach to identification is closely related to approaches in galichon2011set, laffers2019identification and torgovitsky2019partial, who demonstrate how to construct sharp bounds on various parameters in models with discrete variables by partitioning the latent variable space and discretizing the latent variables. We also use an identification argument based on partitioning the latent variable space, and we demonstrate how to practically compute the relevant partition using results from the literature on computational geometry. In certain cases, we also show how to avoid griding over the entire parameter space when computing the identified set.
There are a number of other relevant papers in the literature on partial identification in discrete choice models. manski2007partial considers counterfactual choice probabilities in a setting with partial identification, and shows how these counterfactual choice probabilities can be bounded using optimization problems. However, the response type approach in our paper is quite different. We also show how to practically incorporate different assumptions, and we allow for endogenous explanatory variables.\footnote{This latter point differentiates our work from chiong2017counterfactual and allen2019identification.} In other related work, tebaldi2019nonparametric study the problem of computing various counterfactual quantities in a nonparametric discrete choice model with an application to consumer choice of health insurance in California. However, they focus on quasi-linear utility functions and use the particular structure of their setting to resolve the issue of endogeneity by conditioning on a set of covariates.
Computational considerations are also not the main focus of these papers.
This paper is also related to papers that bound treatment effect parameters. Nonparametric bounds were proposed by manski1990nonparametric, and since then contributions have been made by manski1997monotone, manski2000monotone, heckman2001instrumental, bhattacharya2008treatment, manski2009more, chiburis2010semiparametric, shaikh2011partial, bhattacharya2012treatment and mourifie2015sharp, among many others. However, most papers derive closed-form bounds and provide corresponding proofs of sharpness for a fixed target parameter and under a fixed set of assumptions. Similar to other optimization-based approaches, the advantage of our procedure is its flexibility: the researcher can easily modify the model or target parameter and obtain new bounds without the need to derive closed form expressions, or to provide a corresponding proof of sharpness.\footnote{Examples of optimization-based approaches to bounding treatment effect parameters include balke1997bounds, laffers2019bounding, russell2021sharp, mogstad2018using and gunsilius2020path. } Our optimization-based bounds are shown to be sharp under a variety of different assumptions, including flexible functional form assumptions, and any mix of fixed and random coefficients on either exogenous or endogenous covariates. These assumptions are of direct interest to those familiar with structural models of binary outcome variables, and allow us to obtain many previous results as a special case. We hope our approach may also serve as a useful method of sensitivity analysis for those primarily interested in point identified models.
Finally, our paper makes numerous connections to the literature on computational geometry. Computation of our bounds requires the analysis of a partition of the latent space determined by a finite collection of hyperplanes. This turns out to be a well studied subject in combinatorial geometry, and leads us to consideration of the enumeration algorithm proposed by gu2020nonparametric, who build on the work of rada2018new. Our profiling procedure also makes use of the double-description algorithm proposed by Fukuda.
Paper Outline and Notation
The remainder of the paper proceeds as follows. Section (ref) introduces the main theoretical framework and main assumptions. Section (ref) studies practical implementation of the theoretical framework from Section (ref) and introduces our optimization-based bounding procedure for counterfactual probabilities. Section (ref) then demonstrates how to introduce functional form, independence, and monotonicity assumptions into our bounding procedure. Section (ref) discusses estimation and inference, and Section (ref) applies our methodology to study the impact of health insurance on utilization of health care services. Section (ref) concludes. All proofs can be found in Appendix (ref). Appendix (ref) provides some additional discussion of the results presented in the main text. Appendix (ref) compares our procedure to an approach based on Artstein's inequalities, and Appendix (ref) contains supplementary material for our application. \\ \%, and a comparison of our approach to the approach based on Artstein's inequalities from chesher2013instrumental and chesher2014instrumental is presented in Appendix (ref).\\ \\
Notation: The following notation is relevant for both the main text and the appendices. Given a subset $\mathcal{X}$ of Euclidean space, we use $\mathfrak{B}(\mathcal{X})$ to denote the Borel $\sigma-$algebra on $\mathcal{X}$. For two measurable spaces $(\mathcal{X},\mathfrak{B}(\mathcal{X}))$ and $(\mathcal{X}',\mathfrak{B}(\mathcal{X}'))$, the product $\sigma-$algebra on $\mathcal{X}\times \mathcal{X}'$ is denoted by $\mathfrak{B}(\mathcal{X})\otimes \mathfrak{B}(\mathcal{X}')$. Random variables are denoted using capital letters, and if $X: (\Omega,\mathfrak{A})\to (\mathcal{X},\mathfrak{B}(\mathcal{X}))$ is a random variable defined on the probability space $(\Omega,\mathfrak{A},P)$, then we use $P_{X}$ to denote the probability measure induced on $\mathcal{X}$ by $X$; that is, for any $A \in \mathfrak{B}(\mathcal{X})$, $P_{X}(A) := P(X^{-1}(A))$. Furthermore, we interpret $P_{X\mid X'}(X \in A \mid X' =x')$ as a regular conditional probability measure. Finally, $P_{X\mid X'}$ is used as shorthand for the collection $P_{X\mid X'}:=\{P_{X\mid X'}(\,\cdot\, \mid X'=x') : x' \in \mathcal{X}'\}$. We do not explicitly differentiate between scalars and vectors, or random variables and random vectors. To keep the notation clean, we sometimes omit the transpose when combining column vectors; that is, if $v_{1}$ and $v_{2}$ are two column vectors, rather than write $v=(v_{1}^\top,v_{2}^\top)^\top$ we instead write $v=(v_{1},v_{2})$, where it is understood that $v$ is a column vector unless otherwise specified. The cardinality of a set $\mathcal{X}\subset \mathbb{R}^{d}$ is given by $|\mathcal{X}|$.
General Framework: Theoretical Considerations
Main Assumptions and Definitions
We begin by introducing our main assumptions on the binary response models under consideration.
assumptionThere exists a complete probability space $(\Omega, \mathfrak{A},P)$, a random variable $Y:\Omega \to \{0,1\}$, and random vectors $X:\Omega\to \mathcal{X} \subset \mathbb{R}^{d_{x}}$ and $U:\Omega\to \mathcal{U} = \mathbb{R}^{d_{u}}$ satisfying:
\begin{align}
Y= \mathbbm{1}\{\varphi(X,U,\theta_{0})\geq 0 \}\,\, a.s.,
\end{align}
for some function $\varphi(\,\cdot\,,\theta_{0}): \mathcal{X}\times \mathcal{U} \to \mathbb{R}$ parameterized by $\theta_{0} \in \Theta\subset \mathbb{R}^{d_{\theta}}$ with:
\begin{align}
\varphi(x,u,\theta)=\tilde{\varphi}_{1}(x,\theta)^\top u + \tilde{\varphi}_{2}(x,\theta),
\end{align}
where $\tilde{\varphi}_{1}(\,\cdot\,,\theta):\mathcal{X} \to \mathbb{R}^{d_{u}}$ and $\tilde{\varphi}_{2}(\,\cdot\,,\theta): \mathcal{X} \to \mathbb{R}$ are measurable for each $\theta$. Furthermore, $|\mathcal{X}|=:m < \infty$, the spaces $\mathcal{X}$ and $\mathcal{U}$ are equipped with the Borel $\sigma-$algebra, and the distribution of U assigns zero probability to all sets of the form $\{ u \in \mathcal{U} : \varphi(x,u,\theta_{0}) =0\}$.
In Assumption (ref), $U \in \mathcal{U}$ is a vector of latent variables, $\theta \in \Theta$ is a vector of fixed coefficients, and $X\in \mathcal{X}$ is a vector of covariates. From (ref) we restrict the index function to be linear in the latent variables $U \in \mathcal{U}$, although the model in Assumption (ref) still allows for general nonseparability between covariates and latent variables. Importantly, Assumption (ref) imposes that the random vector $X$ has finite support.
In this model the latent variables can also be interpreted as random coefficients. A special case of linearity occurs when the function $\varphi$ is additively separable in a scalar latent variable $U$, which occurs, for instance, when $\tilde{\varphi}_{1}(x,\theta)=1$. A full analysis of this special case using the framework in this paper is taken up in Appendix (ref).
Finally, assuming the distribution of $U$ assigns zero probability to sets of the form $\{u \in \mathcal{U} : \varphi(x,u,\theta)=0\}$ allows for a simplification of the cell enumeration algorithm introduced in the next section. This assumption is implied by absolute continuity of the distribution of latent variables with respect to the Lebesgue measure, which is a standard assumption in this literature.
comment\setcounter{example}{0}
\begin{example}[Program Evaluation]
Consider the problem of program evaluation. We let $Y$ denote a binary observed outcome variable, and we let $X$ denote a treatment variable. For now we take $X$ to be binary, so that it might be interpreted as indicating attendance to the treatment ($X=1$) or control ($X=0$) groups. The observed outcome is related to two jointly unobserved potential outcomes $Y_{0}$ and $Y_{1}$ through the equation:
\begin{align*}
Y = Y_{1} X + Y_{0}(1-X).
\end{align*}
We also suppose that $Z$ represents a vector of observed covariates, including both continuous and discrete variables, related to the potential outcomes and treatment variable. In an observational study, $X$ may be dependent with potential outcomes $Y_{0}$ and $Y_{1}$ in an arbitrary and unknown way. Typical parameters of interest include the average structural functions (ASFs) or the conditional average treatment effect (CATE). The ASFs are given by:
\begin{align}
P(Y_{1} =1 \mid Z = z) = P(Y_{1} =1 \mid X =1, Z = z) p(z) + P(Y_{1} =1 \mid X =0, Z = z) (1-p(z)),\\
P(Y_{0} =1 \mid Z = z) = P(Y_{0} =1 \mid X =1, Z = z) p(z) + P(Y_{0} =1 \mid X =0, Z = z) (1-p(z)),
\end{align}
where $p(z):=P(X = 1 \mid Z =z )$ denotes the propensity score. The CATE is given by:
\begin{align}
\E[Y_{1} - Y_{0} \mid Z = z] = P(Y_{1} =1 \mid Z = z) - P(Y_{0} =1 \mid Z = z).
\end{align}
That is, the CATE is the difference of the ASFs. Both the propensity score and the “factual” probabilities $P(Y_{1} =1 \mid X =1, Z = z)$ and $P(Y_{1} =0 \mid X =0, Z = z)$ can be consistently estimated using nonparametric methods. However, the probabilities $P(Y_{1} =1 \mid X =0, Z = z)$ and $P(Y_{0} =1 \mid X =1, Z = z)$ are inherently counterfactual, and so are not identified without further assumptions. Furthermore, in cases with multivalued or continuous treatment variables $X$, the number of unidentified counterfactual probabilities can be large relative to the identified factual probabilities.
\end{example}
\begin{example}[Simultaneous Discrete Choice]
Consider the problem of simultaneous discrete choice. Here we let $Y$ and $X$ denote two binary variables, and we let $Z$ and $W$ denote two vectors of covariates. A simple simultaneous discrete choice model is given by:
\begin{align}
Y &= \mathbbm{1}\{ \alpha_{0} + X\alpha_{1} + Z\alpha_{2} -U \geq 0\},\\
X &= \mathbbm{1}\{\theta_{0} + Y\theta_{1} + W\theta_{2} - V \geq 0 \},
\end{align}
where $U$ and $V$ represent unobserved random variables, which may be dependent. We will not impose any parametric distributional assumptions on $U$ or $V$, although we will assume that $Z \independent U \mid X$ and $W \independent V \mid Y$. For the sake of this example we will also suppose that the parameters in (ref) and (ref) are all fixed. Examples of such models include empirical models of social interactions, or empirical entry games (c.f. bresnahan1991empirical, tamer2003incomplete, among many others).
Now suppose that the researcher either (i) has no data on the vector $W$, or (ii) believes that equation (ref) may be misspecified. The researcher may then omit equation (ref) and consider only equation (ref) in isolation. Without modelling the equation determining $X$, and allowing $X$ to be an endogenous variable in (ref), we then return to a single equation discrete choice model with endogenous regressors that falls under our framework. In the empirical game example, a possible object of interest in this case includes the conditional average effect of players $X$'s entry decision on the probability of entry of player $Y$; that is:
\begin{align}
P(U \leq \alpha_{0} + \alpha_{1} + Z\alpha_{2} \mid Z=z) - P(U \leq \alpha_{0} + Z\alpha_{2} \mid Z=z).
\end{align}
Decomposing these probabilities yields:
\begin{align}
P(U \leq \alpha_{0} + \alpha_{1} + Z\alpha_{2} \mid Z=z)&= P(Y=1 \mid X=1, Z=z)p(z)\nonumber\\
&\qquad+ P(U \leq \alpha_{0} + \alpha_{1} + Z\alpha_{2} \mid X=0, Z=z) (1-p(z)),
\end{align}
and:
\begin{align}
P(U \leq \alpha_{0} + Z\alpha_{2} \mid Z=z)&= P(U \leq \alpha_{0} + Z\alpha_{2} \mid X=1, Z=z)p(z)\nonumber\\
&\qquad\qquad\qquad+ P(Y=1 \mid X=0, Z=z) (1-p(z)),
\end{align}
where $p(z):=P(X=1 \mid Z=z)$. The probabilities $p(z)$, $P(Y=1 \mid X=1, Z=z)$ and $P(Y=1 \mid X=0, Z=z)$ can all be consistently estimated using nonparametric methods. The remaining probabilities in these expressions are inherently counterfactual, and require consistent estimates of $\alpha_{0}$, $\alpha_{1}$ and $\alpha_{2}$, as well as an estimate of the conditional distribution of $U$ given $(X,Z)$. Obtaining the latter is complicated by the fact that $U$ is not required to belong to a parametric class of distributions, and by the fact that the dependence between $U$ and $X$ is unrestricted.
\end{example}
We assume that the researcher's objective throughout is to obtain a sharp set of constraints defining the identified set of latent variable distributions, and to use these constraints to bound various counterfactual quantities, such as counterfactual conditional probabilities.\footnote{
Similar to previous works, we take the selection relation as a primitive relation on which to construct a definition of the identified set. The close connection between the selection relation from random set theory and the concept of observational equivalence from the work in econometrics on identification has been appreciated in beresteanu2011sharp, beresteanu2012partial, chesher2013instrumental, chesher2014instrumental, and chesher2017generalized, among many others. We continue this work here.} Define the set:
align[align omitted — 145 chars of source]
chesher2017generalized call this set the $U-$level set; intuitively, it delivers all possible values of the latent variables $u$ consistent with the vector $(y,x,\theta)$ given the binary response model in (ref). A measurable selection from the random set $ \mathcal{U}(Y,X,\theta)$ is a random vector $U: \Omega \to \mathcal{U}$ satisfying $U \in \mathcal{U}(Y,X,\theta)$ a.s.\footnote{A general definition of a selection and a random set is provided in Appendix (ref). In Appendix (ref) we prove that $\mathcal{U}(Y,X,\theta)$ is suitably measurable and thus is a random set under our assumptions (see Lemma (ref)). We also prove the existence of a universally measurable selection (see Lemma (ref)).} Importantly, given a distribution of the observable random vectors $(Y,X)$, a structural function $\varphi$ and a fixed coefficient $\theta \in \Theta$, any two measurable selections $U$ and $U'$ from the random set $\mathcal{U}(Y,X,\theta)$ are observationally equivalent in the sense that both latent variable vectors $U$ and $U'$ are consistent with the observed distribution of $Y$ and $X$ for the vector of parameters $\theta \in \Theta$ through the model (ref). This idea can be used to define the identified set.
definition[Identified Set]
Under Assumption (ref), the (joint) identified set $\mathcal{I}_{Y,X}^{*}$ of conditional latent variable distributions $P_{U\mid Y,X}$ and fixed coefficients $\theta$ is the set of all pairs $(P_{U\mid Y,X},\theta)$ satisfying:
\begin{align}
P_{U\mid Y,X}(U \in \mathcal{U}(Y,X,\theta) \mid Y=y, X=x)=1, \,\, P_{Y,X}-a.s.,
\end{align}
and such that the distribution $P_{U} = P_{U\mid Y,X} P_{Y,X}$ assigns zero probability to all sets of the form $\{u \in \mathcal{U} : \varphi(x,u,\theta)=0\}$.
This definition depends on the observed distribution of $(Y,X)$ through the almost-sure relation in (ref). Conditioning the latent variable distribution on the vector $(Y,X)$ is carried throughout the paper, and we show in Section (ref) that it allows us to bound counterfactual parameters that may be relevant to policy analysis. Definition (ref) can also be used to define other related identified sets, including identified sets for conditional latent variable distributions of the form $P_{U\mid X}$ or $P_{U}$.
exampleConsider the simple additively separable threshold crossing model:
\begin{align*}
Y = \mathbbm{1}\{ X \theta_{0} \geq U \},
\end{align*}
where $X \in \mathcal{X} \subset \mathbb{R}$ has finite support, and $U \in \mathbb{R}$. This is a special case of the model we consider, and is explored in detail in Appendix (ref). Here we have:
\begin{align*}
\mathcal{U}(1,x,\theta) &:= \left\{ u \in \mathbb{R} : u \leq x \theta \right\},\\
\mathcal{U}(0,x,\theta) &:= \left\{ u \in \mathbb{R} : u > x \theta \right\}.
\end{align*}
Now suppose that $P(Y=y, X=x)>0$ for all $(y,x) \in \mathcal{Y} \times \mathcal{X}$. From Definition (ref), the (joint) identified set $\mathcal{I}_{Y,X}^{*}$ is the set of all pairs $(P_{U\mid Y,X},\theta)$ satisfying the conditions:
\begin{align}
P_{U\mid Y,X}( U \leq X \theta \mid Y=1, X=x)=1,\qquad \forall x \in \mathcal{X},\\
P_{U\mid Y,X}( U > X \theta\mid Y=0, X=x)=1, \qquad \forall x \in \mathcal{X},
\end{align}
and such that the distribution $P_{U} = P_{U\mid Y,X} P_{Y,X}$ assigns zero probability to the sets $\{ \{u \in \mathbb{R} : u = x\theta\} : x \in \mathcal{X}\}$.
commentWe use Definition (ref) as a primitive starting point with the goal of providing a definition of the identified set for counterfactual choice probabilities, as well as definitions of the identified set for other objects of interest. For example, using this definition we can also provide definitions of the identified set for various projections of the joint identified set. Under Assumption (ref), the identified set of fixed coefficients $\Theta^{*}$ is given by:
\begin{align}
\Theta^{*}:= \left\{ \theta : \exists P_{\theta\mid Y,X,Z} s.t. (P_{\theta\mid Y,X,Z},\theta) \in \mathcal{I}_{Y,X,Z}^{*}\right\}.
\end{align}
Furthermore, under Assumption (ref), the identified set of conditional latent variable distributions $\mathcal{P}_{\theta\mid Y,X,Z}^{*}$ is given by:
\begin{align}
\mathcal{P}_{\theta\mid Y,X,Z}^{*}:=\left\{P_{\theta\mid Y,X,Z} : \exists \theta s.t. (P_{\theta\mid Y,X,Z},\theta) \in \mathcal{I}_{Y,X,Z}^{*} \right\}.
\end{align}
Conditioning on the value of the observed endogenous outcome variable $Y$ may be new to many readers; however, the identified set $\mathcal{P}_{\theta\mid X,Z}^{*}$ for conditional latent variable distributions $P_{\theta \mid X,Z}$ can constructed via integration of distributions $P_{\theta\mid Y,X,Z} \in \mathcal{P}_{\theta\mid Y,X,Z}^{*}$ with respect to the observed conditional choice probabilities. For the sake of comparison with the previous literature, in Appendix (ref) we show how the identified set of (conditional) latent variable distributions are related to conditional choice probabilities in our setup.
In this paper, we consider counterfactuals characterized by the occurrence of an exogenous intervention that modifies at least one of the explanatory variables.
assumption[Counterfactual Domain]
For some collection of functions $\Gamma$ with typical element $\gamma: \mathcal{X} \to \mathcal{X}$, there exists a collection of random variables $\{Y(\,\cdot\,,\gamma): \Omega \to \{0,1\} \mid \gamma \in \Gamma\}$, abbreviated $Y_{\gamma}:=Y(\,\cdot\,,\gamma)$, representing counterfactual choices for each $\gamma$ such that $Y_{\gamma}: \Omega \to \{0,1\}$ is measurable for each $\gamma$, and:
\begin{align}
P_{Y_{\gamma}\mid Y,X,U}\left(Y_{\gamma} = \mathbbm{1}\{\varphi(\gamma(X),U,\theta_{0}) \geq 0 \}\mid Y=y,X=x,U=u\right)=1,
\end{align}
$P_{Y,X,U}-$a.s. for the same $\theta_{0} \in \Theta$ as in Assumption (ref), and for all $\gamma \in \Gamma$.
Assumption (ref) implies that (i) counterfactual response variables indexed by $\gamma \in \Gamma$ exist on the common probability space from Assumption (ref), and (ii) such counterfactual response variables are equal (almost surely) to the values that would arise after an intervention on the system represented by (ref).
Here each counterfactual is represented by a function $\gamma: \mathcal{X} \to \mathcal{X}$, which allows us to consider a general class of counterfactuals. Although each function $\gamma$ is seen as a map from $\mathcal{X}$ to itself, this does not prevent consideration of counterfactuals where $\gamma$ selects values of $x$ that have never been observed in the data. Such cases can be accommodated by simply extending the support $\mathcal{X}$ from Assumption (ref) to include any counterfactual pair $x$ of interest.\footnote{In particular, if $x \in \mathcal{X}$ but $\gamma(x)=x' \notin \mathcal{X}$, then redefine $\mathcal{X} \leftarrow \mathcal{X} \cup \{x'\}$. This approach does not affect anything we present in this paper, since we always require any relation to the observed distribution of $(Y,X)$ to hold only almost-surely. }
comment\begin{remark}
The random variables $Y_{\gamma}$ can be related to potential outcomes from the literature on treatment effects. Interpreting these variables as potential outcomes helps to clarify the invariance assumption made in Assumption (ref). For example, suppose we have only a binary variable $X\in\{0,1\}$ and latent variables $\theta$ (i.e. no variables $Z$ and no fixed coefficients $\theta$). Then the structural function from (ref) can be written as $\varphi(X,\theta)$. In this simple case we have $\gamma(x)\in\{0,1\}$, and we can define the random variables $Y_{0}$ and $Y_{1}$ as:
\begin{align*}
Y_{0}(\omega):=\mathbbm{1}\{\varphi(0,\theta(\omega))\geq 0 \},\\
Y_{1}(\omega):= \mathbbm{1}\{\varphi(1,\theta(\omega))\geq 0 \}.
\end{align*}
The observed choice for an individual indexed by $\omega \in \Omega$ is then given by:
\begin{align*}
Y(\omega):= Y_{0}(\omega)(1-X(\omega))+Y_{1}(\omega)X(\omega).
\end{align*}
That is, the observed choice corresponds to $Y_{0}(\omega)$ if $X(\omega)=0$ and $Y_{1}(\omega)$ if $X(\omega)=1$, where $Y_{0}(\omega)$ and $Y_{1}(\omega)$ are potential outcomes. Extending the analogy beyond the cases of binary $X$ is straightforward, and it will sometimes be useful to refer to this connection to potential outcomes when discussing the interpretations of some of our parameters.
\end{remark}
definition[Identified Set of Counterfactual Conditional Distributions]
Under Assumptions (ref) and (ref), the identified set of counterfactual conditional distributions $\mathcal{P}_{Y_{\gamma}\mid Y,X,U}^{*}$ is the set of all conditional distributions $P_{Y_{\gamma}\mid Y,X,U}$ satisfying:
\begin{align}
P_{Y_{\gamma}\mid Y,X,U}\left(Y_{\gamma} = \mathbbm{1}\{\varphi(\gamma(X),U,\theta) \geq 0 \}\mid Y=y, X=x,U=u\right)=1,
\end{align}
$P_{Y,X,U}-$a.s. for some $(P_{U\mid Y,X},\theta)\in \mathcal{I}_{Y,X}^{*}$.
Note this definition references the identified set $\mathcal{I}_{Y,X}^{*}$ from Definition (ref). This definition can also be used to define other related identified sets, including for counterfactual distributions $P_{Y_{\gamma}\mid Y,X}$ or $P_{Y_{\gamma}\mid X}$, as well as identified sets for average structural functions and average treatment effects.
\setcounter{example}{0}
example[cont'd]
Consider again the simple additively separable threshold crossing model:
\begin{align*}
Y = \mathbbm{1}\{ X \theta_{0} \geq U \},
\end{align*}
where $X \in \mathcal{X} \subset \mathbb{R}$ has finite support, and $U \in \mathbb{R}$. Now fix some $x' \in \mathcal{X}$ and consider the map $\gamma(x) = x'$ for all $x \in \mathcal{X}$. This corresponds to a counterfactual map $\gamma:\mathcal{X}\to \mathcal{X}$ where all values of $X$ are set to the value $X=x'$. The counterfactual outcome variable $Y_{\gamma}$ satisfies:
\begin{align}
Y_{\gamma} = \mathbbm{1}\{x' \theta_{0} \geq U \},
\end{align}
a.s. for the same $\theta_{0} \in \Theta$ as in Assumption (ref). The identified set of counterfactual conditional distributions $\mathcal{P}_{Y_{\gamma}\mid Y,X,U}^{*}$ is the set of all conditional distributions $P_{Y_{\gamma}\mid Y,X,U}$ satisfying:
\begin{align}
P_{Y_{\gamma}\mid Y,X,U}\left(Y_{\gamma} = \mathbbm{1}\{x' \theta \geq U \}\mid Y=y, X=x,U=u\right)=1,
\end{align}
$P_{Y,X,U}-$a.s. for some $(P_{U\mid Y,X},\theta)\in \mathcal{I}_{Y,X}^{*}$. That is, the identified set contains all (and only those) conditional counterfactual distributions $P_{Y_{\gamma}\mid Y,X,U}$ consistent with the support restriction $Y_{\gamma} = \mathbbm{1}\{\varphi(x',U,\theta) \geq 0 \}$ a.s. for some pair $(P_{U\mid Y,X},\theta)$ belonging to the (joint) identified set of conditional latent variable distributions and fixed coefficients.
commentFor example, the identified set of counterfactual conditional choice probabilities $\mathcal{P}_{Y_{\gamma}\mid Y,X,Z}^{*}$ is the set of all conditional distributions $P_{Y_{\gamma}\mid Y,X,Z}$ satisfying:
\begin{align}
&P_{Y_{\gamma}\mid Y,X,Z}\left(Y_{\gamma} = y'\mid Y=y, X=x,Z=z\right)\nonumber\\
&\qquad\qquad\qquad\qquad= \int P_{Y_{\gamma}\mid Y,X,Z,\theta}\left(Y_{\gamma} = y' \mid Y=y,X=x,Z=z,\theta\right)dP_{\theta\mid Y,X,Z},
\end{align}
$(y,x,z)-$a.s. for some triple $(P_{Y_{\gamma}\mid Y,X,Z,\theta},P_{\theta\mid Y,X,Z},\theta)$ satisfying (ref). The conditional choice probability in (ref) then allows us to answer questions of the form: “given the observed values of $(Y,X,Z)$ are $(y,x,z)$, what would be the expected response if we were to set the values of $(X,Z)$ to be $\gamma(x,z)$?” As we mentioned earlier, conditioning on the value of $Y$ may seem unfamiliar to most, but it allows us to answer a new set of policy questions that condition on the current observed response when evaluating the expected counterfactual response.\footnote{This conforms with a counterfactual in the three-level hierarchy of action, prediction and counterfactuals presented in pearl2009causality. See sections 1.4 and 7.2.}
In addition, the identified set for the (conditional) average structural function $\mathcal{P}_{Y_{\gamma}\mid X,Z}^{*}$ will be the set of all probabilities satisfying:
\begin{align}
&P_{Y_{\gamma}\mid X,Z}\left(Y_{\gamma} = y'\mid X=x,Z=z\right)\nonumber\\
&\qquad\qquad\qquad\qquad= \sum_{y \in \{0,1\}} P_{Y_{\gamma}\mid Y,X,Z}\left(Y_{\gamma} = y'\mid Y=y, X=x,Z=z\right) P(Y=y \mid X=x, Z=z),
\end{align}
$(x,z)-$a.s. for some $P_{Y_{\gamma}\mid Y,X,Z} \in \mathcal{P}_{Y_{\gamma}\mid Y,X,Z}^{*}$. Finally, when $\gamma , \gamma' \in \Gamma$ represent two competing policies, the identified set for the (conditional) average treatment effect from moving from policy $\gamma$ to policy $\gamma'$ is given by the set of all values:
\begin{align}
P_{Y_{\gamma'}\mid Y,X,Z}\left(Y_{\gamma'} = y'\mid X=x,Z=z\right) - P_{Y_{\gamma}\mid Y,X,Z}\left(Y_{\gamma} = y'\mid X=x,Z=z\right),
\end{align}
where both terms are average structural functions. We will show how to construct sharp bounds on all of these objects using our framework.
The following result connects the definitions and assumptions in this section.
theoremSuppose that Assumptions (ref) and (ref) hold. Then a counterfactual conditional distribution $P_{Y_{\gamma}\mid Y,X}$ satisfies $P_{Y_{\gamma}\mid Y,X} \in \mathcal{P}_{Y_{\gamma}\mid Y,X}^{*}$ if and only if there exists a pair $(P_{U\mid Y,X},\theta) \in \mathcal{I}_{Y,X}^{*}$ satisfying:
\begin{align}
P_{Y_{\gamma}\mid Y,X}\left(Y_{\gamma} = 1\mid Y=y,X=x\right)= P_{U\mid Y,X}\left( \varphi(\gamma(X),U,\theta) \geq 0 \mid Y=y,X=x\right),
\end{align}
$P_{Y,X}-$a.s.
Theorem (ref) provides the theoretical link between the identified set of counterfactual conditional distributions, and the identified set for the pair $(P_{U\mid Y,X},\theta)$.
While theoretically straightforward, it hides some important practical difficulties. In particular, verifying the existence of a pair $(P_{U\mid Y,X},\theta)$ that satisfies the conditions from Definition (ref) is a nontrivial task. This is at least partly due to the fact that $P_{U\mid Y,X}$ is an infinite dimensional object, even in the case when $X$ has finite support. This infinite dimensional existence problem is exacerbated in practice by the fact that $P_{U\mid Y,X}$ must satisfy a number of constraints to ensure it is consistent with the binary response model through (ref), and to ensure it is a proper conditional probability measure. We consider these practical difficulties in detail in the next section.
General Framework: Practical Considerations
To make progress, define the following vector-valued function:
align[align omitted — 238 chars of source]
and for a fixed binary vector $s \in \{0,1\}^{m}$ define the set:
align[align omitted — 116 chars of source]
The sets from (ref) partition the space $\mathcal{U}$ into at most $2^{m}$ sets, with each set uniquely associated with a binary vector $s \in \{0,1\}^{m}$. The binary vectors $r(u,\theta)$ represent response types (c.f. balke1994counterfactual and heckman2018unordered).\footnote{The collection of sets defining response types appears to be similar to the “minimal relevant partition” in tebaldi2019nonparametric, as well as the partition described in chesher2014instrumental Appendix B. }
In the discrete choice setting, these response types tell us the choices that an individual with type indexed by $(u,\theta)$ would have made had they been assigned an alternate value of $x$. Any two individuals characterized by values of $u$ from the same set $\mathcal{U}(s,\theta)$ make identical choices under every possible counterfactual assignment $\gamma: \mathcal{X} \to \mathcal{X}$, making this a natural grouping of latent types.
After partitioning the space of latent variables using response types, various counterfactual objects of interest can be written as a disjoint union of the sets $\mathcal{U}(s,\theta)$ from (ref) that comprise the partition. For the sake of illustration, consider the binary vectors:
align[align omitted — 69 chars of source]
for $j=1,\ldots,m$. Note that each set $S_{j}$ is comprised of all binary vectors that have a $j^{th}$ entry equal to $1$, and thus contain exactly $2^{m-1}$ elements.\footnote{It is useful to note that the sets $\{S_{j}\}_{j=1}^{m}$ are not disjoint; indeed, it is easy to show that $S_{j} \cap S_{k} \neq \emptyset$ and $S_{j} \cap S_{k}^{c} \neq \emptyset$ for every $j\neq k$.} By definition of the sets $\mathcal{U}(s,\theta)$ and $S_{j}$ we have:
align[align omitted — 114 chars of source]
Furthermore, for $s' \neq s$ we have $\mathcal{U}(s',\theta) \cap \mathcal{U}(s,\theta) = \emptyset$, so that this union is a disjoint union. Thus, we have the following decomposition:
align[align omitted — 203 chars of source]
Such a decomposition holds for any $j=1,\ldots,m$. When the conditioning value $x$ differs from the value $x_{j}$ in the structural function, an application of Theorem (ref) shows that the left hand side of this display represents a counterfactual conditional distribution, illustrating the connection between response types and counterfactual choices.
comment\begin{remark}
Response types have a natural interpretation in the potential outcome framework as a collection of potential outcomes. To illustrate, let us return to the potential outcome interpretation introduced in Remark (ref). In particular, suppose we have only a binary variable $X\in\{0,1\}$ and latent variables $\theta$ (i.e. no variables $Z$ and no fixed coefficients $\theta$). Then the structural function from (ref) can be written as $\varphi(X,\theta)$ and the binary response vector $r(\theta,\theta)$ can be written as $r(\theta)$. As in Remark (ref), let us define the random variables:
\begin{align*}
Y_{0}(\omega):= \mathbbm{1}\{\varphi(0,\theta(\omega))\geq 0 \},\\
Y_{1}(\omega):= \mathbbm{1}\{\varphi(1,\theta(\omega))\geq 0 \}.
\end{align*}
Now consider the four possible binary vectors $s \in \{0,1\}^{2}$:
\begin{align*}
s_{1} = \begin{bmatrix}
1\\
1\\
\end{bmatrix}, &&s_{2} = \begin{bmatrix}
1\\
0\\
\end{bmatrix},&&s_{3} = \begin{bmatrix}
0\\
1\\
\end{bmatrix},&&s_{4} = \begin{bmatrix}
0\\
0\\
\end{bmatrix}.
\end{align*}
In this simple model, we can see that events of the form $\{ \omega : r(\theta(\omega)) = s\}$ can be written as conjunctions of events involving the potential outcomes $Y_{0}$ and $Y_{1}$; in particular, we have:
\begin{align*}
\{ \omega : r(\theta(\omega)) = s_{1}\} &= \{\omega : Y_{0}(\omega) = 1, \, Y_{1}(\omega)=1\},\\
\{ \omega : r(\theta(\omega)) = s_{2}\} &= \{\omega : Y_{0}(\omega) = 1, \, Y_{1}(\omega)=0\},\\
\{ \omega : r(\theta(\omega)) = s_{3}\} &= \{\omega : Y_{0}(\omega) = 0, \, Y_{1}(\omega)=1\},\\
\{ \omega : r(\theta(\omega)) = s_{4}\} &= \{\omega : Y_{0}(\omega) = 0, \, Y_{1}(\omega)=0\}.
\end{align*}
From here it is easy to see that, if probabilities can be assigned to the events above then---given that these events are disjoint---probabilities can also be assigned to any union of these events; the latter would include parameters like counterfactual choice probabilities, the average structural function or the average treatment effect. An individual with a given response type can thus be equivalently viewed as an individual with a fixed vector of potential outcomes.
\end{remark}
The next result shows that, in order to rationalize a given collection of counterfactual conditional distribution, for each $\theta$ it is both necessary and sufficient to construct a probability measure on sets of the form $\mathcal{U}(s,\theta)$ from (ref) that satisfies the constraints of Theorem (ref). In the statement of the result we redefine $\gamma:\mathbb{N}\to \mathbb{N}$ to denote the index of the point in $\{x_{1},,\ldots,x_{m}\}$ assigned under counterfactual $\gamma$, and we set $S_{\gamma(j)} := \{ s \in \{0,1\}^{m} : s_{\gamma(j)} =1 \}$, the analog of $S_{j}$ from (ref).
theoremSuppose Assumptions (ref) and (ref) hold. Fix some $\theta \in \Theta$ and consider the collection of sets:
\begin{align}
\mathcal{A}(\theta):= \{int(\mathcal{U}(s,\theta)) : s \in \{0,1\}^{m}\}.
\end{align}
Then for any collection of counterfactual conditional distributions $P_{Y_{\gamma} \mid Y,X}$, there exists a collection of Borel conditional probability measures $P_{U\mid Y,X}$ satisfying (ref) with $(P_{U\mid Y,X},\theta) \in \mathcal{I}_{Y,X}^{*}$ if and only if there exists a collection $P_{U \mid Y,X}$ of probability measures on the sets in $\mathcal{A}(\theta)$ satisfying:
\begin{align}
\sum_{s \in S_{j}} P_{U\mid Y,X}\left( int(\mathcal{U}(s,\theta)) \mid Y=1,X=x_{j}\right)&=1,\\
\sum_{s \in S_{j}^{c}} P_{U\mid Y,X}\left( int(\mathcal{U}(s,\theta)) \mid Y=0,X=x_{j}\right)&=1,\\
\sum_{s \in S_{\gamma(j)}} P_{U\mid Y,X}\left( int(\mathcal{U}(s,\theta)) \mid Y=y,X=x_{j}\right)&=P_{Y_{\gamma}\mid Y,X}\left( Y_{\gamma}=1 \mid Y=y,X=x_{j}\right),
\end{align}
for all $y \in \{0,1\}$ and $j \in \{1,\ldots,m\}$ assigned positive probability.
Theorem (ref) reduces our infinite dimensional existence problem to a finite dimensional existence problem. For a fixed $\theta \in \Theta$, instead of verifying there exists a Borel probability measure $P_{U \mid Y, X}$ satisfying (ref) and (ref) (an infinite-dimensional object), Theorem (ref) shows it suffices to verify there exists a finite dimensional probability vector with typical element $P_{U\mid Y,X}\left( \text{int}(\mathcal{U}(s,\theta)) \mid Y=y, X=x\right)$ satisfying (ref) - (ref). Note that this result relies crucially on the finiteness of $\mathcal{X}$. There is a close connection between Theorem (ref) and the bounding approach based on Artstein's inequalities (e.g. chesher2017generalized) and optimal transportation (e.g. galichon2011set).\footnote{In Appendix (ref) we show that Theorem (ref) is equivalent to a characterization based on Artstein's inequalities after conditioning on the value of the endogenous variables. This conditioning allows us to obtain a much smaller number of equality constraints when compared to the full set of unconditional constraints arising from Artstein's inequalities. However, Artstein's inequalities can handle continuous instruments. It is also well known that Artstein's inequalities are equivalent to the existence of a certain zero-cost optimal transport problem (see galichon2009test, galichon2011set, and galichon2016optimal). } Related results have also appeared in laffers2019identification and torgovitsky2019partial. The finite number of linear constraints from Theorem (ref) leads naturally to a linear programming approach to bounds on counterfactual quantities.
commentAll of our counterfactual parameters of interest in this paper can be constructed from the basic counterfactual probability of the form:
\begin{align}
P_{Y_{\gamma}\mid Y,X,Z}\left(Y_{\gamma} = 1\mid Y=y,X=x,Z=z\right).
\end{align}
The result thus also implies that the constraints (ref) - (ref) are sufficient to bound any counterfactual parameter that can be written as a linear function of counterfactual probabilities of the form (ref).
At first glance it may be surprising to note that the constraints in Theorem (ref) depend on the observed distribution only through the condition that each constraint must hold for all values of $(y,x,z)$ assigned positive probability by the observed distribution of $(Y,X,Z)$. Beyond these conditions, the observed distribution plays no role in Theorem (ref). While this may appear to be cause for alarm, we remind readers that Assumptions (ref) and (ref) impose minimal structure on the binary response model; here the function $\varphi$ can be extremely flexible, and the variables $X$ and $Z$ can be endogenous. Theorem (ref) shows that in this very flexible environment, the observed conditional choice probabilities do not provide any substantial information on counterfactual choice probabilities. In the next section we will use this idea to formulate an impossibility result that will be useful to motivate the need for additional assumptions in this binary response model. However, we will first take this basic environment as given and show how bounds on counterfactual choice probabilities can be formulated as optimization problems. Our result on the formulation in terms of optimization problems will then be built upon in later sections when we show how to incorporate further assumptions. Finally, although we focus on an optimization-based approach, we believe the partition of the latent variable space in terms of response types may also be useful for researchers interested in studying sufficient conditions for point identification, as in heckman2018unordered. We will not pursue this approach here.
Optimization Formulation
For simplicity, we suppose throughout this subsection that our objective is to bound the counterfactual probability:
align[align omitted — 113 chars of source]
for some $j \in \{1,\ldots,m\}$. All the results in this section are immediately applicable to the case when we wish to bound some linear function of these counterfactual probabilities.
From Theorem (ref) our counterfactual probability can be written as:
align[align omitted — 179 chars of source]
where $\gamma(j)$ is the index in $\{1,\ldots,m\}$ assigned to $j$ under counterfactual $\gamma$. Define the parameter:
align[align omitted — 135 chars of source]
Furthermore, define:
align[align omitted — 147 chars of source]
align[align omitted — 249 chars of source]
The vector of parameters $\pi(\theta)$ represents the variable over which we optimize in our result ahead, and has dimension $d_{\pi} = 2m2^m$. Without loss of generality, we suppose that each $(y,x)$ is assigned positive probability by the observed distribution. From conditions (ref) and (ref) in Theorem (ref), we have the constraints:
align[align omitted — 147 chars of source]
for $j=1,\ldots,m$. Finally, we require the constraints:
align[align omitted — 182 chars of source]
for all $y\in\{0,1\}$ and $j=1,\ldots,m$ and $s \in \{0,1\}^{m}$, and:
align[align omitted — 83 chars of source]
for all $y \in \{0,1\}$ and $j=1,\ldots,m$.
Note that to impose the constraints from (ref), the researcher must first determine which sets $\mathcal{U}(s,\theta)$ have nonempty interior. We return to this point in the next subsection. We are now ready to state one of the main results.
theoremUnder Assumptions (ref) and (ref), the identified set for the counterfactual conditional probability $P_{Y_{\gamma}\mid Y,X}(Y_{\gamma} = 1\mid Y=y, X=x_{j})$ is given by:
\begin{align}
\bigcup_{\theta \in \Theta} [\pi_{\ell b}(y,x_{j},\theta),\pi_{u b}(y,x_{j},\theta)],
\end{align}
where $\pi_{\ell b}(y,x_{j},\theta)$ and $\pi_{u b}(y,x_{j},\theta)$ are determined by the optimization problems:
\begin{align}
\pi_{\ell b}(y,x_{j},\theta) &:= \min_{\pi(\theta)\in \mathbb{R}^{d_{\pi}}} \sum_{s \in S_{\gamma(j)}} \pi(y,x_{j},s,\theta), subject to (ref), (ref), and (ref),\\
\pi_{u b}(y,x_{j},\theta) &:= \max_{\pi(\theta)\in \mathbb{R}^{d_{\pi}}} \sum_{s \in S_{\gamma(j)}} \pi(y,x_{j},s,\theta), subject to (ref), (ref), and (ref).
\end{align}
remarkIf $\Theta^{*}$ is the identified set for $\theta$, then the interval $[\pi_{\ell b}(y,x_{j},\theta),\pi_{u b}(y,x_{j},\theta)]$ is empty for all values of $\theta \in \Theta\setminus \Theta^{*}$. Thus, it is equivalent to take the union in (ref) over $\Theta$ or $\Theta^{*}$. The statement of Theorem (ref) does not assume the identified set $\Theta^{*}$ is known.
In one direction, Theorem (ref) implies that any counterfactual conditional probability of the form (ref) belonging to the identified set can be written as:
align[align omitted — 131 chars of source]
for some $\theta$ and some vector $\pi(\theta)$ satisfying the constraints (ref), (ref), and (ref). In the opposite direction, the Theorem implies that if for some $\theta$ the vector $\pi(\theta)$ satisfies the constraints (ref), (ref), and (ref) then the conditional probability measure on $\mathcal{U}$ represented by $\pi(\theta)$ can be extended to a (not necessarily unique) Borel probability measure on all of $\mathfrak{B}(\mathcal{U})$ that satisfies the conditions of Theorem (ref). This result can be easily modified to bound any linear function of counterfactual conditional probabilities by simply modifying the objective function in Theorem (ref). We make use of this fact in the application section.
After determining which of the sets $\mathcal{U}(s,\theta)$ have nonempty interior, all the constraints in (ref) and (ref) can be written as linear equality/inequality constraints, so that the optimization problems in (ref) and (ref) are linear programming problems. This is very beneficial, since linear programs can be efficiently solved even in cases with thousands of parameters and constraints.
Interestingly, the proofs for Theorems (ref) and (ref) do not require linearity of the index function in $U$ from Assumption (ref), although this assumption is used starting in the next subsection. This implies that Theorem (ref) can be used to bound counterfactual parameters for models of the form:
align[align omitted — 89 chars of source]
without imposing any restrictions on the index function. Without any restrictions, it is always possible to construct a function $\varphi$ such that all regions $\mathcal{U}(s,\theta)$ have nonempty interior, implying that there are $2^m$ response types. In this case, constraint (ref) only imposes that all probabilities are bounded between zero and one, and the rest of Theorem (ref) remains unchanged. We illustrate our procedure using the model (ref) in the application section.
In the general case, elements of $\pi(\theta)$ corresponding to sets $\mathcal{U}(s,\theta)$ with empty interior can be removed from the parameter vector $\pi(\theta)$ without altering the optimal solutions to the linear programs in (ref) and (ref). This allows for further reduction of the dimension of these linear programs.
Although the number of sets $\mathcal{U}(s,\theta)$ appear to grow exponentially in $m$, in the subsections ahead we show that under linearity of the index function in $U$ the number of sets $\mathcal{U}(s,\theta)$ that have nonempty interior grows at a rate that is polynomial in $m$, substantially reducing the computational burden.
commentSome thought reveals that $\theta$ has no effect on the bounding problem in Theorem (ref) other than through its effect on determining which of the sets $\mathcal{U}(s,\theta)$ are nonempty. Indeed, depending on the form of $\varphi$, for a fixed value of $\theta \in \Theta$ and a fixed vector $s \in \{0,1\}^{m}$ there may be no value of $\theta$ satisfying $r(\theta,\theta)=s$. In practice the identified set can be constructed using Theorem (ref) after fixing a particular functional form for $\varphi$, establishing a grid over the parameter space $\Theta$, and then solving the optimization problems (ref) and (ref) for each value of $\theta$ in the grid. This last step can only be completed after determining which of the sets $\mathcal{U}(s,\theta)$ are nonempty at each value of $\theta$ in the grid. The following proposition demonstrates that, in theory, the researcher need only repeat the procedure just described for finitely many values of $\theta$.
\begin{proposition}
Suppose that Assumptions (ref) and (ref) hold. Then there exists a (not necessarily unique) finite subset $\Theta' \subset \Theta$ such that:
\begin{align*}
&\left\{ \overline{\pi} \in \mathbb{R}^{d_{\pi}} : \exists \theta \in \Theta s.t. \pi(\theta) satisfies (ref), (ref), (ref), and \overline{\pi}=\pi(\theta) \right\}\\
&\qquad\qquad\qquad= \left\{ \overline{\pi} \in \mathbb{R}^{d_{\pi}} : \exists \theta \in \Theta' s.t. \pi(\theta) satisfies (ref), (ref), (ref), and \overline{\pi}=\pi(\theta) \right\}.
\end{align*}
\end{proposition}
We will call the points in the set $\Theta'$ the representative points, although it is important to keep in mind that these points are generally not unique. Assuming the representative points can be determined by the researcher, Proposition (ref) immediately implies that the union over $\theta \in \Theta$ in (ref) can be replaced with a union over $\theta \in \Theta'$. That is, the linear programs in (ref) and (ref) need only be solved at the representative points. Proposition (ref) also implies that the identified set for counterfactual choice probabilities in Theorem (ref) will always be a closed (but possibly disconnected) set. Unfortunately, when the researcher cannot determine the representative points, computing the exact (i.e. not approximate) bounds from Theorem (ref) can become computationally prohibitive. Later on we will provide a tractable way to construct $\Theta'$ in the case when $\varphi$ is linear in $\theta$.
\begin{comment}
An Impossibility Result
To this point, the binary response model has been left almost entirely unrestricted. The following corollary to Theorem (ref) states formally the unsurprising result that, when $\varphi$ is completely unrestricted, counterfactual choice probabilities are not identified under our assumptions.
corollarySuppose Assumptions (ref) and (ref) hold. Then there exists a function $\varphi:\mathcal{X} \times \mathcal{Z} \times \Theta \times \mathcal{U} \to \mathbb{R}$ satisfying Assumptions (ref) and (ref) such that the identified set for any counterfactual choice probability of the form (ref) or of the form (ref) with $\gamma(j)\neq j$ is the interval $[0,1]$.
This result is stated as a corollary since it can be proven using Theorem (ref). Intuitively, Theorem (ref) shows that under Assumptions (ref) and (ref) the observed conditional choice probabilities effectively impose no constraints on counterfactual conditional choice probabilities; indeed, this follows by simple inspection of the constraints (ref) and (ref) from the statement of Theorem (ref), as well as the discussion following the statement of Theorem (ref). Without any constraints imposed by the observed conditional choice probabilities, by choosing a sufficiently flexible function $\varphi:\mathcal{X} \times \mathcal{Z} \times \Theta \times \mathcal{U} \to \mathbb{R}$ it is possible to rationalize all counterfactual conditional choice probabilities in the interval $[0,1]$. This result relies on the form of the counterfactual choice probabilities in (ref) and (ref); for other counterfactual parameters---e.g. the average treatment effect (ref)---it is possible to obtain nontrivial bounds under only Assumptions (ref) and (ref).\footnote{It is important to note that this does not occur because Assumptions (ref) or (ref) have imposed any structure, but instead is because the average treatment effect is a weighted average of unidentified counterfactual conditional choice probabilities and identified conditional choice probabilities.}
This impossibility result shows that it is necessary to entertain additional assumptions in order to obtain informative bounds on counterfactual choice probabilities. In the next section, we will explore additional assumptions that can be used to tighten bounds on counterfactual choice probabilities.\footnote{For additional impossibility results in partially identified discrete response models, see the discussion on pages 1396 and 1402 of manski2007partial as well as Corollary 3 in chesher2013instrumental.}
\end{comment}
Hyperplane Arrangements and Cell Enumeration
To use Theorem (ref), we must first determine which set $\mathcal{U}(s,\theta)$ have nonempty interior in order to impose the constraints from (ref). This section explains how this can be done in our setting by using linearity of the index function in the latent variables in combination with the cell enumeration algorithm of gu2020nonparametric.
First we demonstrate that linearity of the index function in the latent variables imposes restrictions on the model by limiting the number of sets $\mathcal{U}(s,\theta)$ that can be assigned positive probability. Constraining sets of the form $\mathcal{U}(s,\theta)$ to be assigned zero probability is called eliminating response types. Response types corresponding to sets $\mathcal{U}(s,\theta)$ that survive elimination are called admissible, and response types corresponding to sets $\mathcal{U}(s,\theta)$ that are eliminated are called inadmissible. The following simple example shows how we can eliminate response types under Assumption (ref).
exampleSuppose we have a variable $X\in\{0.5,1,2\}$ and latent variables $U \in \mathbb{R}^{2}$. Assume there are no fixed coefficients $\theta$. Then the index function from (ref) can be written as $\varphi(X,U)$ and the binary response vector $r(u,\theta)$ can be written as $r(u)$, where:
\begin{align*}
r(u) =
\begin{bmatrix}
\mathbbm{1}\{\varphi(0.5,u) \geq 0 \}\\
\mathbbm{1}\{\varphi(1,u) \geq 0 \}\\
\mathbbm{1}\{\varphi(2,u) \geq 0 \}
\end{bmatrix}.
\end{align*}
Without any additional restrictions on $\varphi$ there are a total of $2^{|\mathcal{X}|}=8$ possible response types.\footnote{For example, if $U=(U_{1},U_{2})$ take $\varphi(X,U)=\sin(U_{1}X+U_{2})$ and fix $U_{2}=0$. Then it is straightforward to find eight values of the parameter $U_{1} \in [-1,1]$ to rationalize each of the $8$ response types.} That is, $r(u) \in \{s_{1}, \ldots,s_{8}\}$, where:
\begin{align*}
s_{1} = \begin{bmatrix}
0\\
0\\
0
\end{bmatrix}, &&s_{2} = \begin{bmatrix}
1\\
0\\
0
\end{bmatrix},&&s_{3} = \begin{bmatrix}
0\\
1\\
0
\end{bmatrix},&&s_{4} = \begin{bmatrix}
1\\
1\\
0
\end{bmatrix},&&s_{5} = \begin{bmatrix}
0\\
0\\
1
\end{bmatrix}, &&s_{6} = \begin{bmatrix}
1\\
0\\
1
\end{bmatrix},&&s_{7} = \begin{bmatrix}
0\\
1\\
1
\end{bmatrix},&&s_{8} = \begin{bmatrix}
1\\
1\\
1
\end{bmatrix}.
\end{align*}
Now suppose instead that the structural function from (ref) can be written as:
\begin{align}
\varphi(X,U) = XU_{1} - U_{2}.
\end{align}
Then the binary response vector $r(u_{1},u_{2})$ is given by:
\begin{align*}
r(u_{1},u_{2}) =
\begin{bmatrix}
\mathbbm{1}\{u_{1} \geq 2u_{2} \}\\
\mathbbm{1}\{u_{1} \geq u_{2} \}\\
\mathbbm{1}\{2u_{1} \geq u_{2} \}
\end{bmatrix}.
\end{align*}
As is illustrated in Figure (ref), now only $6$ response types are admissible. For a distribution of the latent variables to be admissible in this context, it must assign zero probability to the sets:
\begin{align*}
\mathcal{U}(\theta,s_{3}) &= \left\{(u_{1},u_{2}) : r(u_{1},u_{2}) = s_{3} \right\},\\
\mathcal{U}(\theta,s_{6})&=\left\{(u_{1},u_{2}) : r(u_{1},u_{2}) = s_{6} \right\}.
\end{align*}
These additional constraints must be imposed in our optimization problems from Theorem (ref).
figure[figure omitted — 693 chars of source]
In the general case, it can be shown that when $\varphi$ is restricted to be linear in $U$, there is an upper bound on the number of nonempty sets $\mathcal{U}(s,\theta)$ that grows at a rate that is polynomial in $m$ rather than exponential, which is the case when $\varphi$ is unrestricted.
propositionSuppose that Assumption (ref) is satisfied. Then for each $\theta \in \Theta$, there are at most $\sum_{j=0}^{d_{u}} \binom{m}{j}$ admissible response types.
This result is implied by results in the literature on combinatorial geometry. In particular, linearity of the function $\varphi(\,\cdot\,,u)$ means that for each instance of $(x,\theta)$ the function $\varphi(x,u,\theta)$ defines a hyperplane in $d_{u}-$dimensional space. In the case when the vectors defining these hyperplanes are in general position the upper bound in Proposition (ref) is obtained.\footnote{ A collection of $m$ hyperplanes in $d-$dimensional space are considered to be in general position when any collection of $k$ out of the $m$ hyperplanes intersect in a $d-k$ dimensional space for $1 < k \leq d$, and any collection of $k$ out of $m$ hyperplanes has an empty intersection for $k > d$.} This latter result was initially proven by Buck. Straightforward calculation shows that, for $m>d_{u}+1$, the function $\sum_{j=0}^{d_{u}} \binom{m}{j}$ is bounded above by $(e \cdot m/d_{u})^{d_{u}} = O(m^{d_{u}})$, a polynomial in the number of hyperplanes.
remarkWe conjecture that proposition (ref) is a special case of a more general result. In particular, define the class of functions:
\begin{align}
\Phi_{\theta} := \left\{ \phi:\mathcal{X} \to \mathbb{R} : \phi(x):= \phi(x,u,\theta), u \in \mathcal{U}\right\}.
\end{align}
Now define the function:
\begin{align*}
\Pi_{\Phi_{\theta}}(m):=\sup_{x_{1},x_{2},\ldots,x_{m}} \left| \left\{\begin{bmatrix} \mathbbm{1}\{\phi(x_{1}) \geq 0 \} \\ \mathbbm{1}\{\phi(x_{2}) \geq 0 \} \\ \vdots \\ \mathbbm{1}\{\phi(x_{m}) \geq 0 \}\end{bmatrix} : \phi \in \Phi_{\theta} \right\} \right|.
\end{align*}
If $\mathcal{X}=\{x_{1},\ldots,x_{m}\}$, then $\Pi_{\Phi_{\theta}}(m)$ delivers the largest (over all inputs $\mathcal{X}$ with $|\mathcal{X}|=m$) number of admissible response types consistent with the class of index functions $\Phi_{\theta}$. The Shelah-Sauer Lemma (c.f. anthony1999neural Theorem 3.6) then shows that:
\begin{align*}
\Pi_{\Phi_{\theta}}(m) \leq \sum_{i=0}^{v} \binom{m}{i},
\end{align*}
where $v$ is the Vapnik-Chervonenkis (VC) dimension of $\Phi_{\theta}$.\footnote{The definition of the VC dimension varies depending on the reference; for instance, the definition of the VC dimension (index) in van1996weak is $+1$ larger than the definition of the VC dimension in either anthony1999neural or mohri2012foundations. See the beginning of Section 3.3 in anthony1999neural for the definition of the VC dimension used here. } In the special case when each function $\phi(x,u,\theta)$ from $\Phi_{\theta}$ is linear in $u \in \mathbb{R}^{d_{u}}$, the VC-dimension of $\Phi_{\theta}$ is $d_{u}$, suggesting Proposition (ref) is a special case of the Shelah-Sauer Lemma.\footnote{C.f. Theorem 3.4 in anthony1999neural (note Theorem 3.4 in anthony1999neural is applicable to affine functions rather than linear functions, which increases the VC dimension by 1 relative to our case).}
Define the collection:
align[align omitted — 129 chars of source]
To impose linearity in the latent variables we must determine which sets $\mathcal{U}(s,\theta)$ have nonempty interior, and then ensure that any distribution of the latent variables assigns zero probability to these sets. To compute the collection $S_{\varphi}(\theta)$ we propose to use the enumeration algorithm of gu2020nonparametric.
When the index function $\varphi$ is linear in $U$, for each fixed $\theta$ and $s \in \{0,1\}^{m}$ the set $\mathcal{U}(s,\theta)$ is a convex polyhedron formed by the intersection of at most $m$ halfspaces whose boundaries are hyperplanes of the form $\{u \in \mathcal{U} : \varphi(x,u,\theta)=0\}$. The hyperplane arrangement algorithm of gu2020nonparametric accepts $m$ hyperplanes as an input, and outputs the binary vectors $s$ corresponding to the sets $\mathcal{U}(s,\theta)$ that have nonempty interior, as well as a point from each of these sets. avis1996reverse were the first to provide an enumeration algorithm that runs in a time proportional to the maximum number of sets with nonempty interior. Improvements to the enumeration algorithm were made by sleumer1999output and rada2018new. The algorithm of gu2020nonparametric was developed for the problem of nonparametric maximum likelihood in a linear index random coefficient model. It runs in a time proportional to $O(m^{d_{u}})$, which is near-optimal when the hyperplanes are in general position, according to Proposition (ref).
To understand the algorithm, note that for each $s \in \{0, 1\}^m$ and fixed $\theta$, we can verify using a linear program whether there exists a point in the interior of $\mathcal{U}(s,\theta)$. Consider the following problem:
align[align omitted — 239 chars of source]
where $s_{j}$ is the $j^{th}$ element of our fixed binary vector $s$. If $\varepsilon^{*}$ and $u^{*}$ are the optimal values of the program (ref), then a value $\varepsilon^{*}>0$ indicates that $u^{*}$ is an interior point to the polyhedron $\mathcal{U}(s,\theta)$.
Since the linear program (ref) must be solved for each $s \in \{0, 1\}^m$, checking whether each $\mathcal{U}(s,\theta)$ has nonempty interior requires solving $2^{m}$ linear programs, despite the fact that we know the number of nonempty subsets $\mathcal{U}(s,\theta)$ is polynomial in $m$. To address this issue, the algorithm proposed in gu2020nonparametric builds upon the algorithm in rada2018new. The idea is to add one hyperplane at a time. At step $k$ we start with a collection of $k-1$ hyperplanes from the previous steps, as well as all existing response types found up to step $k-1$. We then introduce a new hyperplane into the arrangement, and determine all newly created response types by solving a collection of linear programs. When a new hyperplane is added the only new cells are those that are created when the existing cells are crossed by the last hyperplane. By efficiently locating those crossed cells at each step, all cells in the full arrangement can be enumerated in polynomial time.
In summary, the hyperplane arrangement algorithm is used as a pre-processing step under Assumption (ref) to determine which sets $\mathcal{U}(s,\theta)$ have nonempty interior in a given application. Eliminating the inadmissible sets $\mathcal{U}(s,\theta)$ can also dramatically reduce the dimension of the parameter vector $\pi(\theta)$ in the bounding optimization problems. In particular, under Assumption (ref) we need only consider a parameter vector $\pi(\theta)$ with typical element $\pi(y,x,s,\theta)$ defined for $s$ corresponding to subsets $\mathcal{U}(s,\theta)$ with nonempty interior. The dimension of the revised parameter vector $\pi(\theta)$ constructed in this way is always upper-bounded by a polynomial in $m$ under Assumption (ref).
In the next subsection we show how the assumption of linearity in parameters $\theta \in \Theta$ can be combined with the hyperplane arrangement algorithm to dramatically simplify the bounding procedure suggested by Theorem (ref).
Profiling Under Linearity in the Fixed Coefficients
Constructing bounds on counterfactual probabilities using Theorem (ref) requires evaluating the linear programs (ref) and (ref) at all values of $\theta \in \Theta$ in the parameter space. In practice this procedure is infeasible, and instead the identified set must be constructed using Theorem (ref) by establishing a grid over $\Theta$. The following proposition demonstrates that, theoretically speaking, the researcher need only consider a finite grid.
propositionSuppose that Assumptions (ref) and (ref) hold. Then there exists a (not necessarily unique) finite subset $\Theta' \subset \Theta$ such that:
\begin{align*}
&\left\{ \overline{\pi} \in \mathbb{R}^{d_{\pi}} : \exists \theta \in \Theta s.t. \pi(\theta) satisfies (ref), (ref), (ref), and \overline{\pi}=\pi(\theta) \right\}\\
&\qquad\qquad\qquad= \left\{ \overline{\pi} \in \mathbb{R}^{d_{\pi}} : \exists \theta \in \Theta' s.t. \pi(\theta) satisfies (ref), (ref), (ref), and \overline{\pi}=\pi(\theta) \right\}.
\end{align*}
We call the points in $\Theta'$ the representative points, although these points are generally not unique. If the representative points can be determined by the researcher, Proposition (ref) implies that the union over $\theta \in \Theta$ in (ref) can be replaced with a union over $\theta \in \Theta'$. That is, the linear programs in (ref) and (ref) need only be solved at the representative points.\footnote{See the discussion at the end of Section 3 in torgovitsky2019partial for the same idea. Proposition (ref) also implies that the identified set for counterfactual conditional distributions in Theorem (ref) is always a closed (but possibly disconnected) set.}
In the case when $\varphi$ is also linear in $\theta$, we provide a polynomial-time algorithm for finding a collection of representative points.\footnote{Linearity in $\theta$ is not necessarily restrictive. Since $X$ is discrete, the researcher can construct a saturated specification with an index function $\sum_{x \in \mathcal{X}} \mathbbm{1}\{X=x\}\theta_{x} + U$. Defining $\theta :=(\theta_{x})_{x \in \mathcal{X}}$, this index function is linear in $(\theta,U)$, and so all of our computational results are applicable.} To introduce the approach, recall the set $S_{\varphi}(\theta)$ from (ref). Note that for any two values of $\theta, \theta' \in \Theta$ with $\theta \neq \theta'$, if $S_{\varphi}(\theta) = S_{\varphi}(\theta')$ then the linear programming problems in Theorem (ref) at $\theta$ and $\theta'$ are identical, since they have an identical set of constraints. The points $\theta$ and $\theta'$ are thus equivalent in the sense that we only need to solve the linear programming problems for one of them. Extending this idea, we can define an equivalence class by the set of all $\theta \in \Theta$ delivering the same collection $S_{\varphi}(\theta)$. We then only need to solve the linear programming problems at one value of $\theta$ belonging to each equivalence class. These values of $\theta$ selected from each equivalence class are exactly what we call representative points.
To see how to find the representative points, partition $x = (x_{r},x_{f})$ and suppose the index function is of the form:
align*[align* omitted — 70 chars of source]
where $x_{r}$ has dimension $d_{u}$ and $x_{f}$ has dimension $d_{f}=d_{x} - d_{u}$. Here $x_{r}$ is a subvector of $x$ associated with a random coefficient $u \in \mathcal{U}$, and $x_{f}$ is a subvector of $x$ associated with a fixed coefficient $\theta \in \Theta$. For any binary vector $s \in \{ 0,1\}^{m}$, define:
align[align omitted — 300 chars of source]
These sets form a unique partition of the space $\mathcal{U}\times \Theta$ into cells defined by $m$ hyperplanes of the form:
align[align omitted — 74 chars of source]
To find the representative points, we project the sets $\mathcal{R}(s)$ onto the parameter space $\Theta$, and intersect the projections of each set $\mathcal{R}(s)$ across $s$. This produces a collection of sets on $\Theta$ corresponding exactly to the equivalence classes discussed above. The main challenge of the procedure is to obtain a tractable characterization of the projected sets.
Define the set:
align*[align* omitted — 143 chars of source]
That is, $S_{\varphi}$ collects all vectors $s \in \{0,1\}^{m}$ corresponding to sets $\mathcal{R}(s)$ with nonempty interior. We begin by determining all vectors $s \in S_{\varphi}$ by running the hyperplane arrangement algorithm of gu2020nonparametric on the $m$ hyperplanes of the form (ref).
These $m$ hyperplanes can be combined as $X_{r} u + X_{f} \theta =0$, where $X_{r}$ is an $m \times d_{u}$ matrix and $X_{f}$ is an $m\times d_{f}$ matrix. Now fix any $s \in S_{\varphi}$, and let $D(s) = \text{diag}(2s-1)$ denote the $m\times m$ diagonal matrix with the sign vector $2s-1$ along its main diagonal. Define $X_{r}(s) := D(s) X_{r}$ and $X_{f}(s):= D(s) X_{f}$. Then $\mathcal{R}(s)$ can be rewritten as:
align[align omitted — 83 chars of source]
The projection of $\mathcal{R}(s)$ on $\Theta$ is:
align[align omitted — 154 chars of source]
The objective is now to write $\Theta(s)$ only in terms of linear inequality constraints in $\theta$; in other words, to “eliminate” the latent variables $u$ from the system of inequalities in (ref).\footnote{It is possible to use Fourier-Motzkin elimination for this purpose, which was explored in a similar context in Section 8.2 of chesher2019generalized. In practice we find that Fourier-Motzkin elimination leads to a prohibitively large number of constraints defining (ref), most of which are redundant. This remains true even when the number of inequalities defining the set (ref) is small.}
To this end, consider the set:
align[align omitted — 125 chars of source]
kohler1967projections showed that the projected set (ref) can be rewritten as:
align*[align* omitted — 150 chars of source]
Furthermore, the Minkowski-Weyl Theorem also allows us to re-write the set $\mathcal{C}(s)$ as:
align[align omitted — 142 chars of source]
where $R(s)$ is some matrix.\footnote{For a general convex set defined by $\Lambda = \{\lambda \in \mathbb{R}^d: A\lambda \leq b\}$, the Minkowski-Weyl Theorem states that every vector $\lambda \in \Lambda$ can be written as $\lambda = \lambda_{1}+\lambda_{2}$, where $\lambda_{1} \in \text{conv} \{ v_1, \dots, v_k\}$ and $\lambda_{2} \in \text{cone}\{v_{k+1}, \dots, v_n\}$. Here $v_1, \dots, v_k$ are called vertices of $\Lambda$ and $v_{k+1}, \dots, v_n$ are the extreme rays of $\Lambda$. In the special case of $b = 0$, where all hyperplanes are through the origin, then $\Lambda$ becomes a polyhedral cone and $k=0$, so that $\Lambda = \text{cone}\{v_1, \dots, v_n\}$. This latter case is what is relevant for us, and the columns of the matrix $R(s)$ are the collections of these extreme rays. } That is, every element belonging to the polyhedral cone $\mathcal{C}(s)$ can be written as a nonnegative linear combination of the columns of $R(s)$. The matrix $R(s)$ is called the generating matrix of the polyhedral cone $\mathcal{C}(s)$, and the problem of finding the minimal generating matrix $R(s)$ is called the extreme ray enumeration problem.\footnote{A minimal generating matrix for $\mathcal{C}(s)$ is a generating matrix with the property that no proper submatrix also generates $\mathcal{C}(s)$ (c.f. Fukuda). Note the minimal generating matrix is unique only up to multiplication by a positive scalar.} After obtaining the matrix $R(s)$, we have the following representation of the projection of $\mathcal{R}(s)$:
align[align omitted — 117 chars of source]
The two representations of the cone $\mathcal{C}(s)$ from (ref) and (ref) are called its H-representation and its V-representation, respectively. Converting from one representation of a convex polyhedron to another is called the double description problem in computational geometry. An efficient double description algorithm was proposed by Fukuda, and we use the R implementation in the package Rcdd by Geyer. avis1997good provides a comparison of different algorithms. The projection procedure is then repeated for all $s \in S_{\varphi}$, resulting in the collection $\{ \Theta(s) : s \in S_{\varphi}\}$. To obtain the final representative points, we stack $R(s)^\top X_f(s) \theta = 0$ for all $s \in S_{\varphi}$, and then run the enumeration algorithm of gu2020nonparametric on this final collection of hyperplanes.
This procedure also sheds light on how to construct the identified set for $\theta$. Since our final arrangement involves only hyperplanes through the origin, each representative point is selected from a polyhedral cone. For some of these representative points, the linear programming problems in our bounding procedure from Theorem (ref) may have an empty feasible region. In these case, the representative points---and all points belonging the corresponding cone---lie outside of the identified set $\Theta^{*}$. Thus, the identified set $\Theta^*$ is the union of the polyhedral cones with representative points associated with nonempty feasible regions in our linear programming problems from Theorem (ref).\footnote{This implies that the identified set $\Theta^*$ may not be connected, and for any $\theta \in \Theta^*$, we also have $\lambda \theta \in \Theta^*$ for all $\lambda \geq 0$. An appropriate normalization---for example, fixing $||\theta||=1$---leads to a bounded identified set $\Theta^*$.}
Additional Assumptions
In this section we describe how to impose additional independence and monotonicity assumptions in our framework.
Independence Assumptions
In some cases the researcher may have access to an observed variable that is independent of the latent variables. If such a variable enters as an argument in the index function, it induces variation in the observed conditional probabilities without affecting the distribution of the latent variables. We refer to such variables as exogenous covariates. A similar intuition applies if the variable is independent of the distribution of latent variables, does not enter as an argument in the structural function, but has nontrivial dependence with the variables that do enter the index function. We refer to such variables as instruments. Both exogenous covariates and instruments can be used to produce a smaller identified set for counterfactual parameters. To introduce our independence assumption, we partition $X=(W,Z)$, where $W \in \mathcal{W} \subset \mathbb{R}^{d_{w}}$ and $Z \in \mathcal{Z} \subset \mathbb{R}^{d_{z}}$ and $d_{w}+d_{z}=d_{x}$.
assumption[Independence]
For all $A \in \mathfrak{B}(\mathcal{U})$ we have $P_{U\mid Z}(A\mid Z=z) = P_{U}(A)$, $P_{Z}-$a.s.
The independence assumption constrains the set of admissible latent variable distributions, and links the conditional distributions of $U \mid Z=z$ across values of $z \in \mathcal{Z}$. Although Assumption (ref) posits full independence between $Z$ and the vector of latent variables $U$, the assumption can be easily modified for the case when a subvector of $Z$, say $Z_{1}$, is conditionally independent of $U$ given some other subvector of $Z$, say $Z_{2}$.
Definition (ref) in Appendix (ref) provides the extension of Definition (ref) to the case when Assumption (ref) also holds, and Corollary (ref) in Appendix (ref) provides the analogous extension of Theorem (ref). To extend the linear programming result of Theorem (ref) we must include additional constraints. Without loss of generality we assume all values of $(y,w,z)$ are assigned positive probability. Then the appropriate constraints on the vector $\pi(\theta)$ are given by:
align[align omitted — 275 chars of source]
for $k=1,\ldots,m_{z}-1$, where $m_{z} = |\mathcal{Z}|$. A formal statement of the extension of Theorem (ref) to the case when the constraints (ref) are also imposed is provided by Corollary (ref) in Appendix (ref).
Monotonicity Assumptions
Let $\mathcal{M} \subset \{1,\ldots,m\}\times \{1,\ldots,m\}$ denote any collection of integer tuples $(j,k)$, where $1 \leq j,k \leq m$.
assumption[Monotonicity]
For each $\theta \in \Theta$ and each tuple $(j,k)$ in the set $\mathcal{M}$, we have $\varphi(x_{j},u,\theta) \leq \varphi(x_{k},u,\theta)$.
This monotonicity assumption states that, when comparing two points $x_{j}$ and $x_{k}$, the value of the structural function can be ordered by the researcher. For instance, if the order determined by the researcher's monotonicity assumption for the points $x_{j}$ and $x_{k}$ is $\varphi(x_{j},u,\theta) \leq \varphi(x_{k},u,\theta)$, then the researcher automatically rules out response types with $\mathbbm{1}\{\varphi(x_{j},u,\theta) \geq 0\} > \mathbbm{1}\{\varphi(x_{k},u,\theta)\geq 0\}$.
exampleSuppose again that we have a binary variable $X\in\{0,1\}$ a latent variable $U$, and no fixed coefficients. Then the structural function from (ref) can be written as $\varphi(X,U)$ and the binary response vector $r(u,\theta)$ can be written as $r(u)$, where:
\begin{align*}
r(u) =
\begin{bmatrix}
\mathbbm{1}\{\varphi(0,u) \geq 0 \}\\
\mathbbm{1}\{\varphi(1,u) \geq 0 \}
\end{bmatrix}.
\end{align*}
Note that there are only four response types; that is, $r(u) \in \{s_{1}, s_{2}, s_{3}, s_{4}\}$ where:
\begin{align*}
s_{1} = \begin{bmatrix}
1\\
1\\
\end{bmatrix}, &&s_{2} = \begin{bmatrix}
1\\
0\\
\end{bmatrix},&&s_{3} = \begin{bmatrix}
0\\
1\\
\end{bmatrix},&&s_{4} = \begin{bmatrix}
0\\
0\\
\end{bmatrix}.
\end{align*}
Without any additional restrictions, all response types---and thus all sets of the form $\mathcal{U}(s,\theta)$ for $s \in \{0,1\}^{2}$---can be assigned positive probability by the optimization problems in Theorem (ref). Now suppose we entertain the monotonicity assumption $\varphi(0,u) \leq \varphi(1,u)$. Imposing this constraint rules out $r(u) = s_{2}$, so the set $\mathcal{U}(\theta, s_{2}) = \{u : r(u) = s_{2}\}$ must be assigned probability zero in any solution to the optimization problems in Theorem (ref). \
Similar monotonicity assumptions in triangular systems have been extensively explored by heckman2018unordered. In particular, heckman2018unordered explore how choice theory can be used to impose monotonicity assumptions and to eliminate response types, and many of their insights are applicable here.
Let $S_{M}$ be the collection of vectors $s \in \{0,1\}^{m}$ that respect the monotonicity relations from Assumption (ref). Definition (ref) in Appendix (ref) provides the extension of Definition (ref) to the case when Assumption (ref) is also imposed. The extension of Theorem (ref) to the case when Assumption (ref) is imposed is provided by Corollary (ref) in Appendix (ref). To extend the results of Theorem (ref) we must simply include the set of constraints imposed by Assumption (ref) in our optimization problems. These constraints are provided in Corollary (ref), and can be written in terms of the parameter vector $\pi(\theta)$ as:
align[align omitted — 97 chars of source]
for all $y\in \{0,1\}$ and $j=1,\ldots,m$ occurring with positive probability. Corollary (ref) in Appendix (ref) then shows the extension of Theorem (ref) to the case when Assumption (ref) is imposed using the constraints (ref).
Consistency, Inference, and Bias Correction
The previous section focused on identification and computation issues, and set up the main optimization procedure for bounding counterfactual quantities. In this section we briefly describe estimation and inference in our setting. Our inference procedure is adapted from the procedure of cho2021simple.
In most applications of our procedure, we expected both the constraints and objective function in the linear programming problems (ref) and (ref) to be data dependent, and our results in this section are designed with this case in mind. We make use of this in the application section. Although motivated by the previous sections, the results in this section do not require any of our previous assumptions to hold. In this sense this section is independent of the previous sections, and may be of interest to researchers working with a similar class of problems.
Let $\mathscr{P}$ denote the set of all probability measures on $\mathcal{Y}\times \mathcal{X}$, and let $\Pi:=\{\pi \in \mathbb{R}^{d_{\pi}} : A \pi \leq b\}$ denote a compact and convex polytope defined by some matrix $A$ and some vector $b$. In our context, $\Pi$ represents the parameter space for the conditional probability vector $\pi$ from Theorem (ref). To introduce our results, we convert each of the equality constraints in the linear programs from Theorem (ref) to two equivalent inequality constraints. Then for a fixed $\theta \in \Theta$, the constraints in the programs (ref) and (ref) can be written as a finite number of moment inequality constraints:
align*[align* omitted — 98 chars of source]
where $\mathcal{J}(\theta)$ is an index set that may depend on $\theta$.\footnote{These moment functions should also include parameter space constraints.} Note this formulation also allows the use of data-dependent constraints in the programs (ref) and (ref), so long as these constraints can be written as a collection of moment inequalities. Furthermore, the objective function in the programs (ref) and (ref) can be written as $\E_{P}[\psi(Y_{i},X_{i},\pi,\theta)]$, which also permits the use of a data-dependent objective function. We require the following assumption.
assumptionThe parameter space $(\Theta, \Pi,\mathcal{P})$ satisfies the following: (i) for each $\theta \in \Theta$, the function $\psi(\,\cdot\,,\theta): \mathcal{Y}\times\mathcal{X} \times \Pi \to \mathbb{R}$ is measurable in $(Y,X)$ and linear in $\pi$ with a (possibly data-dependent) Lipschitz constant $C(\theta)$ satisfying $\sup_{\theta \in \Theta} C(\theta)<\infty$ a.s. (ii) For each $\theta \in \Theta$ the set $\mathcal{J}(\theta)$ is finite, and for each $j\in \mathcal{J}(\theta)$ the functions $m_{j}(\,\cdot\,,\theta): \mathcal{Y}\times \mathcal{X} \times \Pi \to \mathbb{R}$ are measurable in $(Y,X)$ and linear in $\pi \in \Pi$. (iii) For every $P \in \mathcal{P}\subset \mathscr{P}$ there exists $(\pi,\theta) \in \Pi\times \Theta$ such that $\E_{P}[m_{j}(Y,X,\pi,\theta)] \leq 0$ for all $j\in \mathcal{J}(\theta)$. (iv) There exists a $B<\infty$ and a value $\delta>0$ such that:
\begin{align*}
\sup_{P \in \mathcal{P}} \int ||(Y,X)||^{2+\delta} \,dP \leq B.
\end{align*}
(v) There exists some constant $C<\infty$ such that, for each $\theta \in \Theta$:
\begin{align*}
\sup_{P \in \mathcal{P}} \max\left\{\E_{P}||\nabla_{\pi} \psi(Y,X,\pi,\theta)||^{2},\max_{j \in \mathcal{J}(\theta)} \E_{P}||\nabla_{\pi} m_{j}(Y,X,\pi,\theta)||^{2} \right\} \leq C;
\end{align*}
(vi) There exists a finite subset $\Theta' \subset \Theta$ such that:
\begin{align*}
&\left\{ \pi \in \Pi : \exists \theta \in \Theta s.t. \E_{P}[m_{j}(Y,X,\pi,\theta)] \leq 0 for $j\in \mathcal{J}(\theta)$ \right\}\\
&\qquad\qquad\qquad\qquad\qquad= \left\{ \pi \in \Pi : \exists \theta \in \Theta' s.t. \E_{P}[m_{j}(Y,X,\pi,\theta)] \leq 0 for $j\in \mathcal{J}(\theta)$ \right\}.
\end{align*}
(vii) For each $\theta \in \Theta'$ and for some $\varepsilon>0$, define the vector-valued class of functions:
\begin{align*}
\mathcal{F}(\theta)&:= \left\{\left(\psi(\,\cdot\,,\pi,\theta),\left( m_{j}(\,\cdot\,,\pi,\theta)\right)_{j \in \mathcal{J}(\theta)} \right) : \pi \in \Pi_{\varepsilon} \right\},\\
\Pi_{\varepsilon} &:= \left\{\pi \in \mathbb{R}^{d_{\pi}} : A \pi \leq b + \varepsilon \right\},
\end{align*}
where $A\pi \leq b$ are the constraints defining $\Pi$. Then there exists an element-wise measurable envelope $F_{\theta}:\mathcal{Y}\times \mathcal{X} \to \mathbb{R}$ for $\mathcal{F}(\theta)$ that is bounded on $\mathcal{Y}\times \mathcal{X}$.
(viii) The researcher has a sample $\{(Y_{i},X_{i})\}_{i=1}^{n}$, with each $(Y_{i},X_{i})$ an i.i.d. draw from some $P \in \mathcal{P}$.
Part (i) restricts the objective function to be Lipschitz continuous in $\pi$ for each $\theta$. This assumption is satisfied, for instance, when bounding counterfactual probabilities or average treatment effects, as in the application section. Part (ii) requires each moment function to be linear in $\pi$, an assumption which is also satisfied for the model and assumptions in this paper. Assumption (iii) is standard in the moment inequalities literature, and constrains $\mathcal{P}$ to be the collection of data generating processes that satisfy the moment conditions. Part (iv) imposes a uniform integrability requirement on the vector $(Y,X)$ needed for the procedure of cho2021simple, and part (v) imposes a uniform integrability condition on the gradients. These are both easily satisfied with discrete random variables $Y$ and $X$. Part (vi) allows us to replace $\Theta$ with a finite subset $\Theta'$ without impacting the bounding problem. Proposition (ref) shows this is the case in our setting. Part (vii) assumes the existence of a bounded envelope function on a slight expansion of $\Pi$. This is also easily verified, for instance, when bounding counterfactual probabilities or average treatment effects. Finally part (viii) assumes that the sample under consideration is i.i.d. from some $P \in \mathcal{P}$ satisfying the other conditions.
Using these assumptions we will prove a consistency result, and demonstrate how to do inference in our class of problems. For any $c \in \mathbb{R}_{+}$ and any $P \in \mathscr{P}$ let us define:
align*[align* omitted — 257 chars of source]
where:
align*[align* omitted — 215 chars of source]
The value functions $\Psi_{\ell b}(\theta,P,0)$ and $\Psi_{u b}(\theta,P,0)$ are analogous to the value functions $\pi_{\ell b}(y,x,\theta)$ and $\pi_{u b}(y,x,\theta)$ from Theorem (ref). Furthermore, let $\Psi^{*}(P):=\Psi^{*}(P,0)$, and denote the empirical measure as $\mathbb{P}_{n} \in \mathscr{P}$. The following theorem shows that a slight (shrinking) enlargement of the set $\Psi^{*}(\P_{n})$ is a consistent estimator for the set $\Psi^{*}(P)$, where consistency is defined using the Hausdorff metric.\footnote{For two sets $A, B \subset \mathbb{R}^{d}$, the Hausdorff metric is defined as:
align*[align* omitted — 123 chars of source]
}
propositionSuppose that Assumption (ref) holds. Then $d_{H}(\Psi^{*}(\P_{n},b_{n}),\Psi^{*}(P)) = o_{P}(1)$, where $b_{n}$ is any positive user-specified sequence satisfying $b_{n} = O(1/\sqrt{\log(n)})$.
Proposition (ref) is related to results and discussions found in molchanov1998limit, manski2002inference, and chernozhukov2007estimation. It is also a special case of a more general consistency result presented in Appendix (ref), and suggests a consistent estimator for the identified set from Theorem (ref). In the application section we take $b_{n} = b/\sqrt{\log(n)}$ for some $b>0$, and call $\Psi^{*}(\P_{n},b_{n})$ the “plug-in” estimate of the identified set.
In Section (ref) we also use the inference method of cho2021simple, designed for uniform inference on value functions in stochastic linear programming problems. In general, this inference problem is highly irregular, but cho2021simple show that regularity can be restored by introducing infinitesimal random perturbations to the constraints and objective function. After perturbing the problem, they prove consistency of a simple and fast nonparametric bootstrap procedure to construct a confidence set. However, a modification of the procedure of cho2021simple is needed to fit our setting to allow for the profiling points $\theta \in \Theta'$.
Let $\xi\sim P_{\xi}$ denote the random perturbation vector from the procedure of cho2021simple. By Proposition (ref), there exists a finite set $\Theta' \subset \Theta$ of representative points satisfying:
align*[align* omitted — 78 chars of source]
To apply the procedure of cho2021simple, we take each $\theta \in \Theta'$, and use the procedure of cho2021simple to construct a set $C_{n}(1-\alpha,\theta)$ satisfying:
align[align omitted — 214 chars of source]
Here, probability is taken with respect to the product measure $\text{Pr}_{P}\times P_{\xi}$ to account for the random perturbations $\xi \sim P_{\xi}$ introduced to restore regularity. We refer to cho2021simple for additional discussion. We then set:
align*[align* omitted — 86 chars of source]
The following result shows that this confidence set is uniformly valid over $\mathcal{P}$. The proof of the result proceeds by verifying the assumptions of cho2021simple for (ref), and then showing that the modified confidence set $CS_{n}(1-\alpha)$ has the correct coverage.
propositionSuppose that Assumption (ref) holds. Then:
\begin{align}
\liminf_{n\to \infty} \inf_{\{(\psi,P): \psi\in \Psi^{*}(P), P \in \mathcal{P}\}} (Pr_{P}\times P_{\xi}) \left(\psi \in CS_{n}(1-\alpha) \right) \geq 1-\alpha.
\end{align}
In our application ahead, we also report bias-corrected estimates of the lower and upper endpoints of the identified set. Convexity of the minimum in combination with Jensen's inequality shows that the sample analog lower bound is biased upward. Similarly, the sample analog upper bound is biased downward. This leads to an identified set that is on average too narrow. In response, chernozhukov2013intersection proposed the use of half-median unbiased estimators. Half-median unbiased estimates $\hat{\Psi}_{\ell b}$ and $\hat{\Psi}_{ub}$ satisfy $\hat{\Psi}_{\ell b} \leq \Psi_{\ell b}(P)$ and $\Psi_{ub}(P)\leq \hat{\Psi}_{u b}$, each holding with probability at least $1/2$. We construct a bias-corrected estimate of $\Psi_{\ell b}(P)$ ($\Psi_{u b}(P)$) using the inference procedure of cho2021simple by setting $\alpha=1/2$ and by taking our estimate to be the endpoint of the $1-\alpha$ lower (upper) confidence set for $\Psi_{\ell b}(P)$ ($\Psi_{u b}(P)$). In our application we report both the plug-in estimates of our bounds based on Theorem (ref) and the half-median unbiased estimates.
Application
In this section we apply our method to study the impact of private health insurance on an individual's decision to visit a doctor. In general, insurance markets are plagued by problems arising from asymmetric information between consumers and insurance providers (c.f. rothschild1978equilibrium). For example, adverse selection occurs in the health insurance market when individuals have more information about their latent health determinants than the providers of health insurance. A robust prediction of the classical theory of asymmetric information is that those who are more likely to purchase insurance are also those who are more likely to experience the insured risk.\footnote{The “insured risk” refers to the event for which insurance was purchased. In our context, it is any event that would typically require a visit to the doctor.} On the other hand, there has been little and mixed empirical evidence of adverse selection in health insurance markets (see cardon2001asymmetric for a discussion).
In this section we compute various counterfactual parameters while remaining agnostic on the exact nature of the latent variables linking health insurance and health care utilization decisions. We take the decision to visit a doctor as our binary outcome variable of interest, and consider the individuals' private health insurance status as an endogenous explanatory variable. This is consistent with the idea that private insurance status may be dependent with individual-specific latent factors---most importantly, unobserved health determinants and attitudes towards risk---that influence an individual's propensity to visit a doctor. We use data from the 2010 wave of the Medical Expenditure Panel Survey (MEPS), which has also been recently analyzed by han2019estimation and acerenza2021testing. We focus on the same sub-sample considered in these papers. In particular, we focus on the month of January 2010, consider only individuals between ages $25$ and $64$, and drop individuals who obtain either federal or state insurance in 2010 and individuals who are self-employed or unemployed. These restrictions leave us with a sample of $7555$ individuals.
In all specifications $W$ is a binary endogenous variable representing an individual's private insurance status, and we consider a binary health status variable ($Z_{1}$) and a binary marital status variable ($Z_{2}$) as regressors.\footnote{The MEPS data includes information on self-reported health status on a scale from $1- 5$, and we consider values less than or equal to $2$ as being “unhealthy.”} Finally, we use the number of employees working for the individual's firm ($Z_{3}$) as an instrument. This variable provides a measure of the size of a firm and has discrete support in the range $[1,500]$, which we further discretize into 11 bins.\footnote{Variable $Z_3$ is supported on the range $[1, 500]$ and is clearly top-coded. We notice that there is bunching of observations at firm sizes in multiples of five, and some regions of the support of $Z_3$ contain very few observations. In order to get reliable estimates of the conditional choice probabilities, we further discretize the firm size into 11 bins. The bins are respectively $[1, 5]$, $(5, 10]$, $(10,20]$, $(20,30]$, $(30,40]$, $(40,50]$, $(50,60]$, $(60,70]$, $(70,100]$, $(100,200]$ and $(200,500]$.} Using firm size as an instrument is consistent with the evidence that larger firms are more likely to provide health insurance benefits, but do not directly influence an individual's decision to visit a doctor.\footnote{From cardon2001asymmetric p.408: “Another observed symptom, consistent with the theoretical predictions, is that the uninsured tend to work for small employers. Large employers can overcome adverse selection by risk pooling.” }
A possible concern with using firm size as an instrument is that risk averse individuals may be more likely to select into a job with a larger firm size. In an attempt to address this issue, we include an alternate independence assumption that assumes the firm size $Z_3$ is conditionally independent of $U$ given $(Z_1, Z_2)$ only when $Z_{3}$ lies within a certain range. The idea is that once we condition on a particular range of firm size, the remaining variation in firm size is independent of $U$ conditional on $(Z_1, Z_2)$. We consider four ranges, given by $(1, 10]$, $(10, 50]$, $(50, 100]$ and $(100,500]$, and impose our conditional independence assumption for each range separately.
The first parameter we consider is the average treatment effect, defined as:
align[align omitted — 403 chars of source]
This parameter provides the average causal effect of obtaining health insurance on the decision to visit a doctor. Second, we consider the counterfactual choice probability:
align*[align* omitted — 159 chars of source]
for $y \in \{0,1\}$. We focus on the parameter $\mu_{ccp}(0)$ for simplicity, which represents the counterfactual choice probability of visiting a doctor when given private health insurance for the set of individuals who have no insurance and who have chosen not to visit a doctor, averaged across health and marital status. We construct our bounds under the following set of assumptions:
enumerate[label=(A\arabic*)]
• Only Assumptions (ref) and (ref).
• (A1) and monotonicity (Assumption (ref)). See below for further details.
• (A1) and independence between $(Z_{1},Z_{2})$ and $U$ (Assumption (ref)).
• (A1), (A2) and (A3) together.
• (A1) and independence between $(Z_{1},Z_{2}, Z_{3})$ and $U$ (Assumption (ref)).
• (A1), (A2) and (A5) together.
• (A1) and $U \independent (Z_{1}, Z_{2},Z_{3}) \mid Z_{3} \in [a,b]$ for various intervals $[a,b]$ (Assumption (ref)). See the discussion above.
• (A1), (A2) and (A7) together.
Note that the general index function takes the form $\varphi(w,z_{1},z_{2},u,\theta)$. When monotonicity is imposed in (A2), we impose:
align*[align* omitted — 80 chars of source]
for each $z_{2} \in \{0,1\}$. This implies that for an unhealthy individual, the propensity to visit a doctor when the person has private insurance is always weakly greater than without insurance, regardless of marital status. Finally we consider three different models for the binary outcome variable $Y$:
align[align omitted — 297 chars of source]
Recall that the extension of our procedure to cover model (ref) was discussed briefly at the end of Section (ref). Indeed, under model (ref) the index function $\varphi$ need not be explicitly specified and it may not satisfy the linearity assumption made under Assumption (ref). This makes model (ref) the most flexible. Models (ref) and (ref) impose linearity of $\varphi$ in the latent variables and in the parameters. Here we distinguish two cases. In the first case, (ref) regards $(U_{1},U_{2})$ as the latent variables in the model. Model (ref) is the same as (ref) except that we have replaced the random slope coefficient $U_{1}$ from (ref) with a fixed coefficient. Model (ref) represents the additively separable linear index model that is commonly used in the empirical literature, except for the fact that we do not assume a parametric distribution for $U$ and do not have a model for the endogenous variable $W$.
The identified sets for $\mu_{ate}$ under assumptions (A1) - (A8) and models (ref) - (ref) are reported in Table (ref). For simplicity, we report the convex hull of the estimated identified set for each specification. Table (ref) also reports our modified plug-in estimator (see Section (ref)) as well as half-median unbiased estimators and $90\%$ confidence sets constructed using the modified procedure of cho2021simple. Due to a confluence of factors---including the dimension of the empirical choice probability vector, the large number of constraints, and the sample size---we find that the bootstrap standard errors are small, resulting in half-median unbiased estimates that are only slightly more narrow than the 90% confidence sets.
Unsurprisingly, the plug-in bounds on $\mu_{ate}$ shrink as the strength of our assumptions increase. The most flexible model is (ref) under assumption (A1), in which case the length of the bound on $\mu_{ate}$ is one.\footnote{Note that in a potential outcome framework with a binary treatment and binary outcome (and no other additional assumptions), the worst-case bound on the average treatment effect always has a length of one (c.f. manski1990nonparametric). } The identified set for $\mu_{ate}$ also always overlaps zero for model (ref). Results in Table (ref) suggest that full independence of $Z_3$ (A5 and A6) produces more informative bounds for $\mu_{ate}$ than our alternate conditional independence assumption (A7 and A8). In fact, our alternate conditional independence assumption does not provide much identifying power (compare the results under Assumptions (A3) and (A7)). On the other hand, full independence of $Z_3$ does induce a noticeable narrowing of the identified set for $\mu_{ate}$ (compare the results under Assumptions (A3) and (A5)). The results for this model are a useful benchmark to compare with cases where we impose linearity on the index function.
table[table omitted — 2,724 chars of source]
Next, we see in Table (ref) that the linear models from (ref) and (ref) narrow the bounds relative to the case of the general index function under some of the assumptions. Unsurprisingly, the smallest interval for $\mu_{ate}$ for model (ref) is obtained under Assumption (A6), in which case the sign of $\mu_{ate}$ is identified. For models (ref) and (ref) we make use of our method for profiling $\theta$, as described in Section (ref). In model (ref) we must profile on $\theta \in \mathbb{R}^2$ and there are 8 representative points.
Interestingly, we find that under Assumptions (A1) - (A4) and (A7) - (A8), the identified set for $\theta$ is the entire Euclidean space $\mathbb{R}^2$. This illustrates that non-trivial bounds on $\mu_{ate}$ are possible even when the structural parameters are unidentified. Figure (ref) in Appendix (ref) shows the intervals computed using the linear programs of the form (ref) and (ref) for each representative point of $\theta$ under our various assumptions.
In the second linear model (ref), all coefficients are fixed. Thus, we need to profile on a parameter vector $\theta\in \mathbb{R}^3$. Our profiling procedure from Section (ref) returns $96$ representative points, each associated with a polyhedral cone in $\mathbb{R}^3$.
Under Assumptions (A1) and (A2), the identified set for $\theta$ is $\mathbb{R}^3$, while for all other assumptions (A3) - (A8) we get a more informative identified set for $\theta$. In Figure (ref) in Appendix (ref) we also show the intervals computed using the linear programs of the form (ref) and (ref) for each representative point of $\theta$ under our various assumptions. Interestingly, the ATE bounds under (A1) for model (M2) and (M3) are the same as those under model (M1), which suggests that the functional form restrictions only become informative when combined with other modelling assumptions. The sign of $\mu_{ate}$ is identified for model (ref) under assumptions (A4) - (A8). The narrowest bounds for $\mu_{ate}$ under model (ref) are obtained under Assumption (A6), where the $90\%$ confidence interval is $[0.12,0.29]$. For comparison, using the same data but a slightly different model, acerenza2021testing also bound the ATE and obtain a $95\%$ confidence interval of $[0.02,0.29]$.\footnote{In addition to using a different model, acerenza2021testing also use the number of employees (without discretization) as their instrument, and they use inference procedure of chernozhukov2013intersection combined with a sample splitting procedure, which is valid under a very different set of assumptions than those presented in the current paper. }
Next we consider the counterfactual choice probability $\mu_{ccp}(0)$. Table (ref) reports the convex hull of the estimated identified set for $\mu_{ccp}(0)$ under various model specifications and under various assumptions. Similar to the bounds for $\mu_{ate}$, the half-median unbiased estimates are only slightly more narrow than the 90% confidence sets. We also see that the bounds on counterfactual choice probabilities tend to be wide and uninformative for most assumptions. Note that under Assumption (A1) to (A4) we always obtain the interval $[0,1]$ for the estimated identified set.
The narrowest bounds are found in model (ref) under Assumptions (A5) and (A6). These bounds allow us to conclude that the probability an individual visits a doctor when provided private health insurance, given that they have no private health insurance and did not visit a doctor increases noticeably, with a magnitude in the interval $[0.37, 0.55]$ and a 90% confidence set of $[0.22, 0.62]$.
table[table omitted — 2,748 chars of source]
For the sake of comparison, we estimate the following bivariate probit:
align*[align* omitted — 170 chars of source]
where $(Z_1, Z_2, Z_3)$ are assumed to be independent from $(\varepsilon_1, \varepsilon_2)$, which are bivariate normal with mean zero, unit variance and correlation $\rho$. This model was estimated with our data using maximum likelihood, and $\mu_{ate}$ was estimated as $0.16$ with a bootstrapped confidence interval of $[0.11, 0.20]$. This value for $\mu_{ate}$ is not in the plug-in bounds under (M3) and (A5) and (A6), but it does lie within all of the 90% confidence sets in Table (ref), and seems to suggest strong evidence of a positive causal effect of health insurance on the decision to visit the doctor.\footnote{han2019estimation obtain a similar result in a model allowing for $\varepsilon_{1}$ and $\varepsilon_{2}$ to have unrestricted marginals, and a flexible dependence structure. However, they consider a different model from us, and the average treatment effect in han2019estimation is different from ours; we consider the average treatment effect averaged over all values of $(w,z)$, while they report the average treatment effect at the average value of their conditioning variables. They also report the average treatment effect at various quantiles of their conditioning variables.} However, the bivariate probit model is highly parameterized, and the results from Table (ref) suggest that under weaker assumptions the sign of $\mu_{ate}$ may not be identified.\footnote{In fact, acerenza2021testing reject the assumptions of the bivariate probit model using the same data but with a slightly different specification.}
The previous literature studying the effects of health insurance on the utilization of health care services is full of mixed results, and Table (ref) suggests that highly parameterized models may give highly significant, but possibly misleading results relative to models that make weaker assumptions.
Conclusion
This paper considers (partial) identification of a variety of counterfactual parameters in binary response models with possibly endogenous regressors. Importantly, our class of models allows for nonseparability of the index function in latent variables, and does not require any parametric distributional assumptions. Our specific partition of the latent variable space is key to our procedure, and we show how to enumerate the sets in this partition using results from the literature on computational geometry and hyperplane arrangements. In doing so, we provide a feasible method of constructing bounds on counterfactual quantities under a variety of different assumptions with multi-dimensional and nonseparable latent variables. We also thoroughly study the special case when the index function is linear in parameters, and show how to compute exact (i.e. not approximate) sharp bounds on counterfactual quantities. We also show how to adapt a recent inference procedure to the setting in this paper in order to construct confidence sets and bias-corrected estimates of the identified set. Finally, we show how to impose independence and monotonicity assumptions, and we present an application of our method to study the effects of private health insurance on the utilization of health care services.
The consideration of multinomial choice models, triangular systems, or general simultaneous discrete choice models (e.g. games, network formation, or models of social interactions) are all natural future extensions of the framework presented here which we intend to pursue. In addition, this paper emphasizes computational issues that arise in models that are partially identified. We believe exploring applications of state-of-the-art algorithms in computer science to problems in econometrics---as we have attempted here---is a fruitful avenue of future research.
appendix\section{Proofs}
\subsection{Proofs of Results in the Main Text}
\begin{proof}[Proof of Theorem (ref)]
Let $\mathcal{P}_{Y_{\gamma}\mid Y,X}^{**}$ denote the set of all conditional distributions $P_{Y_{\gamma}\mid Y,X}$ such that there exists a pair $(P_{U\mid Y,X},\theta) \in \mathcal{I}_{Y,X}^{*}$ satisfying:
\begin{align}
P_{Y_{\gamma}\mid Y,X}\left(Y_{\gamma} = 1\mid Y=y, X=x\right)= P_{U\mid Y,X}\left( \varphi(\gamma(X),U,\theta) \geq 0 \mid Y=y, X=x\right),
\end{align}
$P_{Y,X}-$a.s. To prove the result it suffices to show $\mathcal{P}_{Y_{\gamma}\mid Y, X}^{*}=\mathcal{P}_{Y_{\gamma}\mid Y, X}^{**}$. To do this, we show that $\mathcal{P}_{Y_{\gamma}\mid Y, X}^{*}\subset\mathcal{P}_{Y_{\gamma}\mid Y, X}^{**}$ and $\mathcal{P}_{Y_{\gamma}\mid Y, X}^{**}\subset\mathcal{P}_{Y_{\gamma}\mid Y,X}^{*}$. To this end, begin by fixing an arbitrary $P_{Y_{\gamma}\mid Y, X} \in \mathcal{P}_{Y_{\gamma}\mid Y, X}^{*}$. By Definition (ref) we have:
\begin{align}
P_{Y_{\gamma}\mid Y,X,U}\left(Y_{\gamma} = \mathbbm{1}\{\varphi(\gamma(X),U,\theta) \geq 0 \}\mid Y=y,X=x,U=u\right)=1,
\end{align}
$P_{Y,X,U}-$a.s. for some $(P_{U\mid Y,X},\theta)\in \mathcal{I}_{Y,X}^{*}$. For this pair $(P_{U\mid Y,X},\theta)$ we have:
\begin{align*}
&P_{Y_{\gamma}\mid Y,X,U}\left(Y_{\gamma} = 1\mid Y=y,X=x,U=u\right)\\
&= P_{Y_{\gamma}\mid Y,X,U}\left(Y_{\gamma} = 1,Y_{\gamma} = \mathbbm{1}\{\varphi(\gamma(X),U,\theta) \geq 0 \}\mid Y=y,X=x, U=u\right),
\end{align*}
$P_{Y,X,U}-$a.s., which follows from (ref). Now note:
\begin{align*}
P_{Y_{\gamma}\mid Y,X,U}\left(Y_{\gamma} =1, Y_{\gamma} = \mathbbm{1}\{\varphi(\gamma(X),U,\theta) \geq 0 \}\mid Y=y,X=x,U=u\right)= \mathbbm{1}\{\varphi(\gamma(x),u,\theta) \geq 0 \},
\end{align*}
$P_{Y,X,U}-$a.s. Thus we have:
\begin{align*}
P_{Y_{\gamma}\mid Y,X}\left(Y_{\gamma} = 1\mid Y=y,X=x\right) &= \int P_{Y_{\gamma}\mid Y,X,U}\left(Y_{\gamma} = 1\mid Y=y,=x,U=u\right) \, dP_{U\mid Y,X}\\
&= \int \mathbbm{1}\{\varphi(\gamma(x),u,\theta) \geq 0 \} \, dP_{U\mid Y,X}\\
&= P_{U\mid Y,X}(\varphi(\gamma(X),U,\theta) \geq 0\mid Y=y,X=x),
\end{align*}
$P_{Y,X}-$a.s. This proves $P_{Y_{\gamma}\mid Y,X} \in \mathcal{P}_{Y_{\gamma}\mid Y,X}^{**}$, and since $P_{Y_{\gamma}\mid Y,X} \in \mathcal{P}_{Y_{\gamma}\mid Y,X}^{*}$ was arbitrary we conclude that $\mathcal{P}_{Y_{\gamma}\mid Y,X}^{*} \subset \mathcal{P}_{Y_{\gamma}\mid Y,X}^{**}$.
For the reverse inclusion, fix any arbitrary $P_{Y_{\gamma}\mid Y,X} \in \mathcal{P}_{Y_{\gamma}\mid Y,X}^{**}$. Then by definition there exists a pair $(P_{U\mid Y,X},\theta) \in \mathcal{I}_{Y,X}^{*}$ satisfying:
\begin{align}
P_{Y_{\gamma}\mid Y,X}\left(Y_{\gamma} = 1\mid Y=y,X=x\right)= P_{U\mid Y,X}\left( \varphi(\gamma(X),U,\theta) \geq 0 \mid Y=y, X=x\right),
\end{align}
$P_{Y,X}-$a.s. It suffices to show that for this pair $(P_{U\mid Y,X},\theta)$ there exists $P_{Y_{\gamma}\mid Y,X,U}$ satisfying:
\begin{align}
P_{Y_{\gamma}\mid Y,X,U}\left(Y_{\gamma} = \mathbbm{1}\{\varphi(\gamma(X),U,\theta) \geq 0 \}\mid Y=y,X=x,U=u\right)=1,
\end{align}
$P_{Y,X,U}-$a.s. By the Radon-Nikodym Theorem, the existence of a (version of) $P_{Y_{\gamma}\mid Y,X,U}$ is guaranteed by the fact that $P_{Y_{\gamma},U\mid Y,X}\ll P_{U \mid Y, X}$ for all $(y,x)$ occurring with positive probability. Since all spaces involved are Euclidean, we can choose the version to be an almost surely unique regular conditional distribution (c.f. durrett2010probability Theorem 5.1.9). By construction this $P_{Y_{\gamma}\mid Y,X,U}$ satisfies:
\begin{align*}
&P_{Y_{\gamma},U\mid Y,X}(Y_{\gamma} \in A, U \in B \mid Y=y,X=x)\\
&\qquad\qquad= \int_{B} P_{Y_{\gamma}\mid Y,X,U}(Y_{\gamma} \in A \mid Y=y,X=x,U=u) \, dP_{U \mid Y,X},
\end{align*}
$P_{Y,X}-$a.s. for every $A \subset \{0,1\}$ and $B \in \mathfrak{B}(\mathcal{U})$. Now note that:
\begin{align*}
P_{Y_{\gamma}\mid Y,X,U}(Y_{\gamma}=1, Y_{\gamma}= \mathbbm{1}\{\varphi(\gamma(X),U,\theta) \geq 0 \} \mid Y=y,X=x,U=u) = \mathbbm{1}\{\varphi(\gamma(x),u,\theta) \geq 0 \},\\
P_{Y_{\gamma}\mid Y,X,U}(Y_{\gamma}=0, Y_{\gamma}= \mathbbm{1}\{\varphi(\gamma(X),U,\theta) \geq 0 \} \mid Y=y,X=x,U=u) = \mathbbm{1}\{\varphi(\gamma(x),u,\theta) < 0 \}.
\end{align*}
$P_{Y,X}-$a.s. Thus:
\begin{align*}
&P_{Y_{\gamma}\mid Y,X}\left(Y_{\gamma} = \mathbbm{1}\{\varphi(\gamma(X),U,\theta) \geq 0 \}\mid Y=y,X=x\right)\\
&= \int_{\mathcal{U}} P_{Y_{\gamma}\mid Y,X,U}(Y_{\gamma} = \mathbbm{1}\{\varphi(\gamma(X),U,\theta) \geq 0 \} \mid Y=y,X=x,U=u) \, dP_{U \mid Y,X}\\
&= \int_{\mathcal{U}} P_{Y_{\gamma}\mid Y,X,U}(Y_{\gamma}=1, Y_{\gamma}= \mathbbm{1}\{\varphi(\gamma(X),U,\theta) \geq 0 \} \mid Y=y,X=x,U=u) \, dP_{U \mid Y,X}\\
&\qquad\qquad\qquad+ \int_{\mathcal{U}} P_{Y_{\gamma}\mid Y,X,U}(Y_{\gamma}=0, Y_{\gamma}= \mathbbm{1}\{\varphi(\gamma(X),U,\theta) \geq 0 \} \mid Y=y,X=x,U=u) \, dP_{U \mid Y,X}\\
&= \int_{\mathcal{U}} \mathbbm{1}\{\varphi(\gamma(x),u,\theta) \geq 0 \} \, dP_{U \mid Y,X}+ \int_{\mathcal{U}} \mathbbm{1}\{\varphi(\gamma(x),u,\theta) < 0 \} \, dP_{U \mid Y,X}\\
&= P_{U\mid Y,X}(\varphi(\gamma(x),u,\theta) \geq 0 \mid Y=y, X=x) + P_{U\mid Y,X}(\varphi(\gamma(x),u,\theta) < 0 \mid Y=y, X=x)\\
&=1,
\end{align*}
$P_{Y,X}-$a.s. This proves (ref) and thus shows $P_{Y_{\gamma}\mid Y,X} \in \mathcal{P}_{Y_{\gamma}\mid Y,X}^{*}$. Since $P_{Y_{\gamma}\mid Y,X} \in \mathcal{P}_{Y_{\gamma}\mid Y,X}^{**}$ was arbitrary we can conclude that $\mathcal{P}_{Y_{\gamma}\mid Y,X}^{**} \subset \mathcal{P}_{Y_{\gamma}\mid Y,X}^{*}$. Combining the two inclusions, we have $\mathcal{P}_{Y_{\gamma}\mid Y,X}^{*} =\mathcal{P}_{Y_{\gamma}\mid Y,X}^{**}$. This completes the proof.
\end{proof}
\begin{proof}[Proof of Theorem (ref)]
Let $P_{Y_{\gamma} \mid Y,X}$ be a collection of conditional distributions, and suppose there exists $(P_{U \mid Y,X},\theta) \in \mathcal{I}_{Y,X}^{*}$ satisfying (ref). Since $P_{U \mid Y,X} \ll P_{U}$ for all $(y,x)$ assigned positive probability, and since $P_{U} = P_{U \mid Y,X} P_{Y,X}$ assigns zero probability to sets of the form $\{u \in \mathcal{U} : \varphi(x,u,\theta)=0\}$, (ref) is equivalent to (ref), so we can conclude that $(P_{U \mid Y,X},\theta)$ satisfies (ref). Furthermore, by definition $(P_{U \mid Y,X},\theta) \in \mathcal{I}_{Y,X}^{*}$ implies that:
\begin{align}
P_{U\mid Y,X}(U \in \mathcal{U}(Y,X,\theta) \mid Y=y, X=x)=1, \,\, P_{Y,X}-a.s.,
\end{align}
Again, since $P_{U \mid Y,X} \ll P_{U}$ for all $(y,x)$ assigned positive probability, and since $P_{U}$ assigns zero probability to sets of the form $\{u \in \mathcal{U} : \varphi(x,u,\theta)=0\}$, the previous display is equivalent to conditions (ref) and (ref). This shows that any pair $(P_{U \mid Y,X},\theta) \in \mathcal{I}_{Y,X}^{*}$ satisfying (ref) satisfies (ref) - (ref).
For the reverse, fix any $\theta\in \Theta$ and any collection $P_{U \mid Y,X}$ of conditional probability measures on the sets in $\mathcal{A}(\theta)$ satisfying (ref) - (ref). We show that $P_{U \mid Y,X}$ can be extended to a (not necessarily unique) probability measure $\tilde{P}_{U \mid Y,X}$ on $\mathfrak{B}(\mathcal{U})$ in a manner that ensures $\tilde{P}_{U \mid Y,X}$ satisfies (ref) and such that $(\tilde{P}_{U \mid Y,X},\theta) \in \mathcal{I}_{Y,X}^{*}$. Furthermore, by the definition of an extension, $\tilde{P}_{U \mid Y,X}$ agrees with $P_{U \mid Y,X}$ on all sets in $\mathcal{A}(\theta)$. To construct the extension, for each $s \in \{0,1\}^{m}$ select a single point $u(s,\theta)$ from $\text{int}(\mathcal{U}(s,\theta))$ if $\text{int}(\mathcal{U}(s,\theta))\neq \emptyset$; otherwise choose $u(s,\theta)$ as an arbitrary point from $\mathcal{U}$. For any set $A \subset \mathcal{U}$, define the indicator:
\begin{align*}
\mathbbm{1}(A,\theta,s) = \mathbbm{1}\{ u(s,\theta)\in A\cap int(\mathcal{U}(s,\theta)) \}.
\end{align*}
Now define the function $\mu_{y,x}:\mathfrak{B}(\mathcal{U})\to \mathbb{R}$ as:
\begin{align*}
\mu_{y,x}(B) := \sum_{s \in \{0,1\}^{m}} \mathbbm{1}(B,\theta,s) P_{U\mid Y,X}\left( int(\mathcal{U}(s,\theta))\mid Y=y,X=x\right).
\end{align*}
To verify that this is a proper probability measure on $\mathfrak{B}(\mathcal{U})$, we must show that (i) $\mu_{y,x}(B) \geq \mu_{y,x}(\emptyset) =0$ for every $B \in \mathfrak{B}(\mathcal{U})$, (ii) $\mu_{y,x}(\mathcal{U})=1$, and (iii) for any countable sequence of disjoint sets $\{A_{i}\}_{i=1}^{\infty}$ in $\mathfrak{B}(\mathcal{U})$, we have:
\begin{align*}
\mu_{y,x} \left( \bigcup_{i=1}^{\infty} A_{i} \right) = \sum_{i=1}^{\infty} \mu_{y,x}(A_{i}).
\end{align*}
The first property holds since $\mathbbm{1}(\emptyset,\theta,s)=0$ for all $s$. To verify the second property, note that $\mathbbm{1}(\mathcal{U},\theta,s)=1$ for all $s$ with $\text{int}(\mathcal{U}(s,\theta))\neq \emptyset$, so that:
\begin{align*}
\mu_{y,x}(\mathcal{U}) &= \sum_{s \in \{0,1\}^{m}} \mathbbm{1}(\mathcal{U},\theta,s) P_{U\mid Y,X}\left( int(\mathcal{U}(s,\theta)) \mid Y=y,X=x\right)\\
&= \sum_{s : int(\mathcal{U}(s,\theta)) \neq \emptyset} P_{U\mid Y,X}\left( int(\mathcal{U}(s,\theta)) \mid Y=y,X=x\right)\\
&= 1,
\end{align*}
where the last line holds since $P_{U\mid Y,X}$ is a probability measure on $\mathcal{A}(\theta)$. For the third property, note that for two disjoint Borel sets $A_{1},A_{2} \in \mathfrak{B}(\mathcal{U})$ we have:
\begin{align*}
\mathbbm{1}(A_{1}\cup A_{2},\theta,s) = \mathbbm{1}(A_{1},\theta,s) + \mathbbm{1}(A_{2},\theta,s).
\end{align*}
Inducting on this formula, we conclude that for countable disjoint sets $\{A_{i}\}_{i=1}^{\infty}$ in $\mathfrak{B}(\mathcal{U})$, we have:
\begin{align*}
\mathbbm{1} \left( \bigcup_{i=1}^{\infty} A_{i}, \theta, s \right) = \sum_{i=1}^{\infty} \mathbbm{1}(A_{i},\theta,s),
\end{align*}
Thus we can conclude:
\begin{align*}
\mu_{y,x} \left( \bigcup_{i=1}^{\infty} A_{i} \right) &= \sum_{s \in \{0,1\}^{m}} \mathbbm{1}\left(\bigcup_{i=1}^{\infty} A_{i},\theta,s\right) P_{U\mid Y,X}\left( int(\mathcal{U}(s,\theta)) \mid Y=y,X=x\right)\\
&= \sum_{s \in \{0,1\}^{m}} \sum_{i=1}^{\infty} \mathbbm{1}(A_{i},\theta,s) P_{U\mid Y,X}\left( \text{int}(\mathcal{U}(s,\theta)) \mid Y=y,X=x\right)\\
&= \sum_{i=1}^{\infty} \sum_{s \in \{0,1\}^{m}} \mathbbm{1}(A_{i},\theta,s) P_{U\mid Y,X}\left( \text{int}(\mathcal{U}(s,\theta)) \mid Y=y,X=x\right)\\
&= \sum_{i=1}^{\infty} \mu_{y,x}(A_{i}).
\end{align*}
Thus, our measure satisfies countable additivity. We conclude that $\mu_{y,x}$ is a proper probability measure. Note that the argument above has been completed for a single pair $(y,x)$ indexing the conditioning variables. However, we can repeat the same argument as above for all $(y,x)$ assigned positive probability, and thus can construct a corresponding probability measure $\mu_{y,x}$ satisfying all the conditions described above for each such $(y,x)$.
Now we define $\tilde{P}_{U\mid Y,X} : \mathfrak{B}(\mathcal{U}) \to [0,1]$ by $\tilde{P}_{U\mid Y,X}(B\mid Y=y,X=x) = \mu_{y,x}(B)$ for all $B \in \mathfrak{B}(\mathcal{U})$ and all $(y,x)$ assigned positive probability. By the above, $\tilde{P}_{U\mid Y,X}(\,\cdot\,\mid Y=y,X=x)$ is a proper probability measure on $\mathfrak{B}(\mathcal{U})$ for each $(y,x)$. Also note that for any pair $(1,x)$ assigned positive probability, the pair $(\tilde{P}_{U\mid Y,X},\theta)$ satisfies:
\begin{align}
&\tilde{P}_{U\mid Y,X}(\mathcal{U}(1,x,\theta) \mid Y=1,X=x)\nonumber\\
&= \sum_{s \in S_{j}} \tilde{P}_{U\mid Y,X}(\mathcal{U}(s,\theta)\mid Y=1,X=x)\nonumber\\
&= \sum_{s \in S_{j}} \sum_{s' \in \{0,1\}^{n}} \mathbbm{1}(\mathcal{U}(s,\theta),\theta,s') P_{U\mid Y,X}\left( \text{int}(\mathcal{U}(s,\theta))\mid Y=1,X=x\right)\nonumber\\
&= \sum_{s \in S_{j}} \mathbbm{1}(\mathcal{U}(s,\theta),\theta,s) P_{U\mid Y,X}\left( \text{int}(\mathcal{U}(s,\theta)) \mid Y=1,X=x\right)\nonumber\\
&= 1,
\end{align}
which follows from (ref). Furthermore, for any pair $(0,x)$ assigned positive probability, the pair $(\tilde{P}_{U\mid Y,X},\theta)$ also satisfies:
\begin{align}
&\tilde{P}_{U\mid Y,X}(\mathcal{U}(0,x,\theta) \mid Y=0,X=x) \nonumber\\
&= \sum_{s \in S_{j}^{c}} \tilde{P}_{U\mid Y,X}(\mathcal{U}(s,\theta)\mid Y=0,X=x)\nonumber\\
&= \sum_{s \in S_{j}^{c}} \sum_{s' \in \{0,1\}^{n}} \mathbbm{1}(\mathcal{U}(s,\theta),\theta,s') P_{U\mid Y,X}\left( \text{int}(\mathcal{U}(s,\theta)) \mid Y=0,X=x\right)\nonumber\\
&= \sum_{s \in S_{j}^{c}} \mathbbm{1}(\mathcal{U}(s,\theta),\theta,s) P_{U\mid Y,X}\left( \text{int}(\mathcal{U}(s,\theta)) \mid Y=0,X=x\right)\nonumber\\
&= 1,
\end{align}
which follows from (ref). Conclude that:
\begin{align*}
\tilde{P}_{U\mid Y,X}(U \in \mathcal{U}(Y,X,\theta) \mid Y=y,X=x) &=1, \,\,a.s.
\end{align*}
It is also straightforward to see that $\tilde{P}_{U}:=\tilde{P}_{U\mid Y,X} P_{Y,X}$ assigns zero probability to all sets of the form $\{u \in \mathcal{U} : \varphi(x,u,\theta) = 0\}$, since these sets have empty intersection with $\text{int}(\mathcal{U}(s,\theta))$ for all $s\in \{0,1\}^{m}$. Combining everything, this shows that $(\tilde{P}_{U\mid Y, X},\theta) \in \mathcal{I}_{Y,X}^{*}$. Finally, setting $C:=\{ u \in \mathcal{U} : \varphi(\gamma(x),u,\theta) \geq 0 \}$, it is straightforward to show that:
\begin{align*}
\tilde{P}_{U\mid Y,X}\left( C \mid Y=y, X=x_{j}\right)&= \sum_{s \in \{0,1\}^{m}} \mathbbm{1}(C,\theta,s) P_{U\mid Y,X}\left( \text{int}(\mathcal{U}(s,\theta)) \mid Y=y,X=x_{j}\right)\\
&=\sum_{s \in S_{\gamma(j)}} P_{U\mid Y,X}\left( \text{int}(\mathcal{U}(s,\theta)) \mid Y=y,X=x_{j}\right)\\
&=P_{Y_{\gamma}\mid Y,X}\left( Y_{\gamma}=1 \mid Y=y,X=x_{j}\right),
\end{align*}
for all $(y,x_{j})$ assigned positive probability, which follows from (ref). This is exactly condition (ref). Conclude that $(\tilde{P}_{U\mid Y, X},\theta) \in \mathcal{I}_{Y,X}^{*}$ and that $(\tilde{P}_{U\mid Y,X},\theta)$ satisfies (ref). This completes the proof.
\end{proof}
\begin{proof}[Proof of Theorem (ref)]
Note that the constraints in (ref) are equivalent to the constraints in (ref) and (ref). Furthermore, the objective function in the optimization problems in Theorem (ref) enforce (ref). Thus, using Theorem (ref), a distribution $\pi(\theta)$ is feasible in the optimization problems from Theorem (ref) if and only if there exists a collection of Borel conditional probability measures $P_{U\mid Y,X}$ satisfying (ref) with $(P_{U\mid Y,X},\theta) \in \mathcal{I}_{Y,X}^{*}$. However, by Theorem (ref), there exists a collection of Borel conditional probability measures $P_{U\mid Y,X}$ satisfying (ref) with $(P_{U\mid Y,X},\theta) \in \mathcal{I}_{Y,X}^{*}$ if and only if $P_{Y_{\gamma} \mid Y,X} \in \mathcal{P}_{Y_{\gamma} \mid Y,X}^{*}$, where $P_{Y_{\gamma} \mid Y,X}$ is the (collection of) conditional distribution(s) satisfying (ref).
\end{proof}
\begin{proof}[Proof of Proposition (ref)]
First note that $\theta \in \Theta$ enters the constraints in Theorem (ref) only through the constraints (ref); in particular, only through its determination of which sets $\text{int}(\mathcal{U}(s,\theta))$ are empty versus nonempty. Now define:
\begin{align*}
\mathcal{S}_{\varphi}(\theta):= \{s \in \{0,1\}^{m} : \text{int}(\mathcal{U}(s,\theta)) \neq \emptyset \}.
\end{align*}
Now define an equivalence relation $\sim$ on $\Theta$ as follows: $\theta \sim \theta'$ if and only if $\mathcal{S}_{\varphi}(\theta) = \mathcal{S}_{\varphi}(\theta')$. This equivalence relation will partition $\Theta$ into at most $2^{2^{m}}$ equivalence classes (which is the total number of ways of choosing $k$ vectors from $\{0,1\}^{m}$ for $k=0,1,\ldots,2^{m}$). Furthermore, any two values $\theta$ and $\theta'$ belonging to the same equivalence class will deliver the same values for the linear programs (ref) and (ref) (by construction of the equivalence class). Thus, it is sufficient to consider only one $\theta$ from each equivalence class in Theorem (ref), showing there are at most $2^{2^{m}}$ such $\theta$'s to consider.
\end{proof}
\begin{comment}
\begin{proof}[Proof of Corollary (ref)]
A counterfactual choice probability of the form in (ref) can be written as:
\begin{align*}
P_{Y_{\gamma}\mid Y,X,Z}\left( Y_{\gamma}=1 \mid Y=y,X=x_{j},Z=z_{j}\right) = \sum_{s \in S_{\gamma(j)}} P_{\theta\mid Y,X,Z}\left( \mathcal{U}(s,\theta) \mid Y=y,X=x_{j},Z=z_{j}\right).
\end{align*}
Note that the result is trivial if we consider $(y,x_{j},z_{j})$ assigned zero probability, since in that case Theorem (ref) implies there are no constraints on the counterfactual choice probability above. Thus, assume that $(y,x_{j},z_{j})$ is assigned positive probability. By assumption, $\gamma(j)\neq j$. We now claim that (i) $S_{\gamma(j)} \cap S_{j} \neq \emptyset$, (ii) $S_{\gamma(j)}\cap S_{j}^{c} \neq \emptyset$, (iii) $S_{\gamma(j)}^{c} \cap S_{j} \neq \emptyset$, (iv) $S_{\gamma(j)}^{c}\cap S_{j}^{c} \neq \emptyset$. In particular, any $s \in \{0,1\}^{m}$ with $j^{th}$ entry equal to $1$ and $\gamma(j)^{th}$ entry equal to $1$ belongs to $S_{\gamma(j)} \cap S_{j}$. Denote such a vector by $t_{1} \in \{0,1\}^{m}$. Similarly, any $s \in \{0,1\}^{m}$ with $j^{th}$ entry equal to $0$ and $\gamma(j)^{th}$ entry equal to $1$ belongs to $S_{\gamma(j)} \cap S_{j}^{c}$. Denote such a vector by $t_{2} \in \{0,1\}^{m}$. Continuing in this way, let $t_{3} \in S_{\gamma(j)}^{c} \cap S_{j}$ and $t_{4} \in S_{\gamma(j)}^{c} \cap S_{j}^{c}$. Now fix any $\varphi$ and $\theta$ such that all $2^{m}$ sets $\mathcal{U}(s,\theta)$ are nonempty (such a choice is always possible under Assumptions (ref) and (ref)). For any $\kappa \in [0,1]$ consider the following conditional distribution on sets $A \in \mathcal{A}(\theta)$:
\begin{align*}
P_{\theta\mid Y,X,Z}\left(A \mid Y=y,X=x_{j},Z=z_{j}\right) = \begin{cases}
\kappa, &\text{ if }A= \mathcal{U}(\theta,t_{1}), \text{ and } y=1,\\
\kappa, &\text{ if }A= \mathcal{U}(\theta,t_{2}), \text{ and } y=0,\\
1-\kappa, &\text{ if }A= \mathcal{U}(\theta,t_{3}), \text{ and } y=1,\\
1-\kappa, &\text{ if }A= \mathcal{U}(\theta,t_{4}) \text{ and } y=0,\\
0, &\text{otherwise.}
\end{cases}
\end{align*}
If $y=1$ we have:
\begin{align}
\sum_{s \in S_{j}} P_{\theta\mid Y,X,Z}\left( \mathcal{U}(s,\theta) \mid Y=1,X=x_{j},Z=z_{j}\right)&= \kappa+ (1-\kappa) = 1,
\end{align}
and if $y=0$ we have:
\begin{align}
\sum_{s \in S_{j}^{c}} P_{\theta\mid Y,X,Z}\left( \mathcal{U}(s,\theta) \mid Y=0,X=x_{j},Z=z_{j}\right)&=\kappa +(1-\kappa) = 1.
\end{align}
This shows that constraints (ref) and (ref) are satisfied. Finally, note that for either $y=0$ or $y=1$ we have:
\begin{align}
\sum_{s \in S_{\gamma(j)}} P_{\theta\mid Y,X,Z}\left( \mathcal{U}(s,\theta) \mid Y=y,X=x_{j},Z=z_{j}\right)&=\kappa.
\end{align}
Since this can be completed for any $\kappa \in [0,1]$, conclude that the identified set for the counterfactual choice probability in (ref) is the interval $[0,1]$. Furthermore:
\begin{align*}
&P_{Y_{\gamma}\mid X,Z}\left( Y_{\gamma}=1 \mid X=x_{j},Z=z_{j}\right) \\
&= \sum_{y \in \{0,1\}} P_{Y_{\gamma}\mid Y,X,Z}\left( Y_{\gamma}=1 \mid Y=y,X=x_{j},Z=z_{j}\right) P_{Y \mid X,Z} (Y=y \mid X=x_{j}, Z=z_{j})\\
&= \kappa.
\end{align*}
Again, since this can be completed for any $\kappa \in [0,1]$, conclude that the identified set for counterfactual choice probabilities of the form $P_{Y_{\gamma}\mid X,Z}\left( Y_{\gamma}=1 \mid X=x_{j},Z=z_{j}\right)$ is the interval $[0,1]$.
\end{proof}
\end{comment}
\begin{proof}[Proof of Proposition (ref)]
This follows immediately from the results of Buck.
\end{proof}
\begin{comment}
It remains only to verify that $P_{\theta\mid X,Z}$ constructed in this way satisfies $P_{\theta\mid X,Z} = P_{\theta\mid X}$. To this end, note that the marginal $P_{\theta\mid X}$ for our constructed distribution $P_{\theta\mid X,Z}$ is given by:
\begin{align*}
P_{\theta\mid X}(A\mid X=x_{j}) = \sum_{\{\ell \, :\, x_{\ell}=x_{j}\}} P_{\theta\mid X,Z}(A\mid X=x_{\ell},Z=z_{\ell}) P(Z=z_{\ell} \mid X=x_{\ell}).
\end{align*}
But by definition of $P_{\theta\mid X,Z}(A\mid X=x_{j},Z=z_{j})$ we have:
\begin{align*}
&P_{\theta\mid X}(A\mid X=x_{j})\\
&= \sum_{\{\ell \, :\, x_{\ell}=x_{j}\}} \mu_{\ell}(A) P(Z=z_{\ell} \mid X=x_{\ell})\\
&= \sum_{\{\ell \, :\, x_{\ell}=x_{j}\}} \sum_{s \in \{0,1\}^{n}} \mathbbm{1}(A,\theta,s) p_{j}(s,\theta)\\
&= \sum_{s \in \{0,1\}^{n}} \mathbbm{1}(A,\theta,s)\sum_{\{\ell \, :\, x_{\ell}=x_{j}\}} p_{j}(s,\theta)\\
&= \sum_{s \in \{0,1\}^{n}} \mathbbm{1}(A,\theta,s) q_{j}(s,\theta).
\end{align*}
where we have used the definition of $q_{j}(s,\theta)$. Now note that:\note{Recall the reason why $c-$dependence wont work. We would require the last sum to be less than or equal to $c$ regardless of the set $A$ under consideration, which seems impossible to guarantee.}
\begin{align*}
&\left|P_{\theta\mid X,Z}(A\mid X=x_{j},Z=z_{j}) - P_{\theta\mid X}(A\mid X=x_{j})\right|\\
&\qquad\qquad=\left|\sum_{s \in \{0,1\}^{n}} \mathbbm{1}(A,\theta,s) p_{j}(s,\theta) - \sum_{s \in \{0,1\}^{n}} \mathbbm{1}(A,\theta,s) q_{j}(s,\theta)\right|\\
&\qquad\qquad=\left|\sum_{s \in \{0,1\}^{n}} \mathbbm{1}(A,\theta,s) \left( p_{j}(s,\theta) - q_{j}(s,\theta)\right) \right|\\
&\qquad\qquad=0,
\end{align*}
where the last line follows from the constraint (ref). Therefore, we conclude that $P_{\theta\mid X,Z}$ is such that $(P_{\theta\mid X,Z},\theta) \in \mathcal{I}_{X,Z}^{*}$ and such that $(P_{\theta\mid X,Z},\theta)$ rationalizes the counterfactual choice probability defined in (ref). Since $p(\theta)$ satisfying (ref), (ref), (ref) and (ref) was arbitrary, we conclude that for any $p(\theta)$ satisfying these constraints, there exists $(P_{\theta\mid X,Z},\theta) \in \mathcal{I}_{X,Z}^{*}$ such that $(P_{\theta\mid X,Z},\theta)$ rationalizes the counterfactual choice probability defined using $p(\theta)$ in (ref). Thus, we conclude that $[\ell b_{j}, ub_{j}] \subset \mathcal{P}_{Y_{\gamma}^{\star}\mid X,Z}(x_{j},z_{j})$. This completes the proof.
\end{comment}
\begin{proof}[Proof of Proposition (ref)]
It suffices to verify the assumptions of Theorem (ref); i.e., Assumption (ref).
\begin{enumerate}[label=(\roman*)]
• In our context, the parameter space $\mathcal{T}$ is given by $\Pi\times \Theta$. The set $\Pi$ is a compact and convex polytope by definition.
• Part (ii) of Assumption (ref) is implied by part (i) of Assumption (ref).
• Part (iii) of Assumption (ref) is implied by part (ii) of Assumption (ref).
• Part (iv) of Assumption (ref) is implied by part (iii) of Assumption (ref).
• Fix any $\theta \in \Theta$ and define:
\begin{align*}
\eta_{n}(\theta):= &\bigg\{\max_{j \in \mathcal{J}(\theta)} \sup_{\pi \in \Pi} |\E_{n}[m_{j}(Y_{i},X_{i},\pi,\theta)] - \E_{P}[m_{j}(Y_{i},X_{i},\pi,\theta)]|,\\
&\qquad\qquad\qquad\qquad\qquad\qquad\sup_{\pi \in \Pi} |\E_{n}[\psi(Y_{i},X_{i},\pi,\theta)] - \E_{P}[\psi(Y_{i},X_{i},\pi,\theta)]| \bigg\}.
\end{align*}
Now define the classes of functions:
\begin{align*}
\mathcal{M}_{j}(\theta)&:= \left\{m_{j}(\,\cdot\,,\pi,\theta):\mathcal{Y}\times \mathcal{X} \to \mathbb{R} : \pi \in \Pi \right\},\\
\tilde{\Psi}(\theta)&:= \left\{\psi(\,\cdot\,,\pi,\theta):\mathcal{Y}\times \mathcal{X} \to \mathbb{R} : \pi \in \Pi \right\}.
\end{align*}
By Assumption (ref), $\psi(\,\cdot\,,\theta): \mathcal{Y}\times\mathcal{X} \times \Pi \to \mathbb{R}$ is measurable in $(Y,X)$ and linear in $\pi \in \Pi$, and (for all $j \in \mathcal{J}(\theta)$) the functions $m_{j}(\,\cdot\,,\theta): \mathcal{Y}\times \mathcal{X} \times \Pi\to \mathbb{R}$ are measurable in $(Y,X)$ and linear in $\pi \in \Pi$. These classes are thus all VC-subgraph classes (c.f. Lemma 2.6.15 in van1996weak) with a bounded (and thus uniformly square integrable) envelope function. From here, standard arguments show that these classes are Donsker uniformly over $P \in \mathcal{P}$, and thus $\eta_{n}(\theta)$ is $O_{P}(n^{-1/2})$. This argument shows that part (v) of Assumption (ref) is satisfied with $a_{n}:= \sqrt{n}$.
• From the previous part, any sequence satisfying $b_{n} = O(1/\sqrt{\log(n)})$ also satisfies $b_{n} \geq \eta_{n}(\theta)$ w.p.a. 1.
• Part (vii) of Assumption (ref) is implied by part (vi) of Assumption (ref). Note this assumption also ensures that we can consider at most a finite number of moment inequalities, indexed by by $j \in \cup_{\theta \in \Theta'} \mathcal{J}(\theta)$.
• The last part of Assumption (ref) (i.i.d. data) is implied by part (viii) of Assumption (ref).
\end{enumerate}
\end{proof}
\begin{proof}[Proof of Proposition (ref)]
Fix some $\theta \in \Theta$ and let $C_{n}(1-\alpha,\theta)$ denote the confidence set constructed using the procedure of cho2021simple. We first show:
\begin{align}
\liminf_{n\to \infty} \inf_{\{(\psi,P): \psi\in \Psi^{*}(\theta,P), P \in \mathcal{P}\}} (\text{Pr}_{P}\times P_{\xi}) \left(\psi_{0} \in CS_{n}(1-\alpha,\theta) \right) \geq 1-\alpha.
\end{align}
To do so, it suffices to show that, for our fixed $\theta \in \Theta$, Assumptions 3.1 and 3.2 in cho2021simple are satisfied. By Assumption (ref), $\psi(\,\cdot\,,\theta): \mathcal{Y}\times\mathcal{X} \times \Pi \to \mathbb{R}$ is linear in $\pi$ and the functions $m_{j}(\,\cdot\,,\theta): \mathcal{Y}\times \mathcal{X} \times \Pi \to \mathbb{R}$ are linear in $\pi \in \Pi$. Since $(Y,X)$ has finite support, we can equip $\mathcal{Y}\times\mathcal{X}$ with the discrete topology, in which case every function on $\mathcal{Y}\times \mathcal{X}$ is continuous. This verifies Assumption 3.1 in cho2021simple. Parts (i), (ii) and (iii) of Assumption 3.2 in cho2021simple are implied by the fact that $\Pi$ is a compact and convex polytope, and parts (iii) and (viii) of Assumption (ref) (resp.). Parts (iv), (v) and (vi) of Assumption 3.2 in cho2021simple are then implied by parts (vii), (iv) and (v) of Assumption (ref) (resp.). This verifies Assumption 3.2 in cho2021simple. Now note:
\begin{align*}
&\liminf_{n\to \infty} \inf_{\{(\psi,P): \psi\in \Psi^{*}(P), P \in \mathcal{P}\}} (\text{Pr}_{P}\times P_{\xi}) \left(\psi_{0} \in CS_{n}(1-\alpha) \right)\\
&=\liminf_{n\to \infty} \inf_{\{(\psi,P): \psi\in \Psi^{*}(P,\theta), \theta \in \Theta', P \in \mathcal{P}\}} (\text{Pr}_{P}\times P_{\xi})\left(\psi_{0} \in CS_{n}(1-\alpha) \right)\\
&=\liminf_{n\to \infty} \min_{\theta \in \Theta'} \inf_{\{(\psi,P): \psi\in \Psi^{*}(\theta,P), P \in \mathcal{P}\}} (\text{Pr}_{P}\times P_{\xi}) \left(\psi_{0} \in CS_{n}(1-\alpha) \right)\\
&=\liminf_{n\to \infty} \min_{\theta \in \Theta'} \inf_{\{(\psi,P): \psi\in \Psi^{*}(\theta,P), P \in \mathcal{P}\}} (\text{Pr}_{P}\times P_{\xi}) \left(\psi_{0} \in \bigcup_{\theta \in \Theta'} CS_{n}(1-\alpha,\theta) \right)\\
&\geq\liminf_{n\to \infty} \min_{\theta \in \Theta'} \inf_{\{(\psi,P): \psi\in \Psi^{*}(\theta,P), P \in \mathcal{P}\}} \min_{\theta \in \Theta'} (\text{Pr}_{P}\times P_{\xi})\left(\psi_{0} \in CS_{n}(1-\alpha,\theta) \right)\\
&=\liminf_{n\to \infty} \min_{\theta \in \Theta'} \inf_{\{(\psi,P): \psi\in \Psi^{*}(\theta,P), P \in \mathcal{P}\}} (\text{Pr}_{P}\times P_{\xi}) \left(\psi_{0} \in CS_{n}(1-\alpha,\theta) \right)\\
&=\min_{\theta \in \Theta'} \liminf_{n\to \infty} \inf_{\{(\psi,P): \psi\in \Psi^{*}(\theta,P), P \in \mathcal{P}\}} (\text{Pr}_{P}\times P_{\xi})\left(\psi_{0} \in CS_{n}(1-\alpha,\theta) \right)\\
&\geq 1-\alpha,
\end{align*}
where the second last line follows from continuity of the minimum, and the last line follows from (ref). This completes the proof.
\end{proof}
\subsection{Measurability Results}
\begin{definition}[Weak Measurability, Random Set, Selection]
Let $(\Omega,\mathfrak{A},P)$ be a probability space, let $\mathcal{V}$ be a Polish space, and let $\mathcal{O}_{\mathcal{V}}$ denote the collection of all open sets on $\mathcal{V}$. A multifunction $\mathbb{V}: \Omega \to 2^{\mathcal{V}}$ is called weakly-measurable if for every $A \in \mathcal{O}_{\mathcal{V}}$ we have $\mathbb{V}^{-}(A):=\{ \omega \in \Omega : \mathbb{V}(\omega) \cap A \neq \emptyset \} \in \mathfrak{A}$.
A random set is a weakly measurable multifunction defined on a probability space.
If $\mathbb{V}: \Omega \to 2^{\mathcal{V}}$ is a random set, then a random element $V: \Omega \to \mathcal{V}$ is called a (measurable) selection of $ V$ if $V(\omega) \in \mathbb{V}(\omega)$ for $P-$almost all $\omega \in \Omega$.
\end{definition}
\begin{lemma}
Suppose Assumption (ref) holds. Then for each $\theta \in \Theta$, the map $\mathcal{U}(Y(\,\cdot\,),X(\,\cdot\,),\theta): \Omega \to 2^\mathcal{U}$ is a weakly-measurable multifunction, and thus is a random set.
\end{lemma}
\begin{proof}[Proof of Lemma (ref)]
Fix any open set $A \in \mathcal{O}_{\mathcal{U}}$. We want to show that:
\begin{align*}
\{ \omega \in \Omega : \mathcal{U}(Y(\omega),X(\omega),\theta) \cap A \neq \emptyset \} \in \mathfrak{A}.
\end{align*}
First, define:
\begin{align*}
B(A):= \{ (y,x) \in \mathcal{Y}\times \mathcal{X} : \mathcal{U}(y,x,\theta) \cap A \neq \emptyset \}.
\end{align*}
Since $\mathcal{Y} \times\mathcal{X}$ is finite, and is equipped with the discrete topology and the Borel $\sigma-$algebra, we trivially have $B(A) \in \mathfrak{B}(\mathcal{Y})\otimes \mathfrak{B}(\mathcal{X})$. Since, $(Y,X) : \Omega \to \mathcal{Y}\times \mathcal{X}$ is measurable by assumption, we have $(Y,X)^{-1}(B) := \{ \omega : (Y(\omega),X(\omega)) \in B\} \in \mathfrak{A}$. Thus:
\begin{align*}
\{ \omega \in \Omega : \mathcal{U}(Y(\omega),X(\omega),\theta) \cap A \neq \emptyset \} = (Y,X)^{-1}(B(A)) \in \mathfrak{A},
\end{align*}
as desired.
\end{proof}
Given a $\sigma-$algebra $\mathfrak{F}$ on a space $\mathcal{R}$, the $P$-completion of $\mathfrak{F}$ is the smallest $\sigma-$algebra containing $\mathfrak{F}$ as well as all $P-$null sets of $\mathcal{R}$. The intersection of all $P-$completions of $\mathfrak{F}$ (over all $P$) is called the \textit{universal $\sigma-$algebra}, and functions that are measurable with respect to the universal $\sigma-$algebra are said to be \textit{universally measurable}. The following Lemma shows that the random set $\text{cl } \mathcal{U}(Y,X,\theta)$ admits a universally measurable selection under Assumption (ref).
\begin{lemma}
Suppose Assumption (ref) holds. Then $\text{cl }\mathcal{U}(Y(\omega),X(\omega),\theta)$ admits a universally measurable selection for every $\theta \in \Theta$ ensuring it is nonempty almost surely.
\end{lemma}
\begin{proof}[Proof of Lemma (ref)]
Fix some $\theta \in \Theta$ ensuring $ \mathcal{U}(Y(\omega),X(\omega),\theta)$ is almost surely nonempty. We can then revise $\mathcal{U}(Y(\omega),X(\omega),\theta)$ on any null set to ensure it is nonempty for all $\omega \in \Omega$. By Lemma (ref), $\mathcal{U}(Y(\omega),X(\omega),\theta)$ is weakly-measurable, and by Theorem 18.6 in aliprantis2006infinite this implies that the graph of $\text{cl }\mathcal{U}(Y(\omega),X(\omega),\theta)$ belongs to $\mathfrak{A}\times \mathfrak{B}(\mathcal{U})$; that is, $\text{cl }\mathcal{U}(Y(\omega),X(\omega),\theta)$ is graph-measurable. The result then follows immediately from Theorem 3 of sainte1974extension.
\end{proof}
Taking the closure in Lemma (ref) is a technical detail that does not impact any of the identification results since under Assumption (ref) the paper restricts attention to selections that assign zero probability to the boundary of $\mathcal{U}(Y,X,\theta)$.
\section{Additional Definitions and Results}
\begin{comment}
\subsection{Identified Set of Conditional Latent Variable Distributions}
For the sake of comparison with the previous literature, we now present a result which connects the observed conditional choice probabilities to our definition of the identified set based on the selection relation. The identified set for $P_{\theta \mid X,Z}$ is given by:
\begin{align}
\mathcal{P}_{\theta\mid X,Z}^{*}:=\left\{P_{\theta\mid X,Z} : \exists (P_{\theta\mid Y,X,Z},\theta) \in \mathcal{I}_{Y,X,Z}^{*} \text{ s.t } P_{\theta\mid X,Z} = \int P_{\theta \mid Y, X,Z} \, dP_{Y \mid X,Z} \text{ a.s} \right\}.
\end{align}
We now have the following result.
\begin{theorem}
Suppose Assumption (ref) holds. Then a collection $P_{\theta\mid X,Z}$ satisfies $P_{\theta\mid X,Z} \in \mathcal{P}_{\theta\mid X,Z}^{*}$ if and only if $P_{\theta\mid X,Z}$ satisfies:
\begin{align}
P_{\theta\mid X,Z}(\varphi(x,z,\theta,\theta) \geq 0 \mid X=x,Z=z) = P_{Y\mid X,Z}(Y=1 \mid X=x,Z=z),
\end{align}
$(x,z)-$a.s. for some $\theta \in \Theta$.
\end{theorem}
\begin{proof}[Proof of Theorem (ref)]
Let us define:
\begin{align*}
\mathcal{P}_{\theta\mid X,Z}^{**}= \left\{ P_{\theta\mid X,Z} : \exists \theta \in \Theta \text{ s.t. }P_{\theta\mid X,Z}(\varphi(x,z,\theta,\theta) \geq 0\mid X=x,Z=z) = P_{Y\mid X,Z}(Y=1\mid X=x,Z=z),\,\,(x,z)-a.s. \right\}.
\end{align*}
We want to show that $\mathcal{P}_{\theta\mid X,Z}^{*}= \mathcal{P}_{\theta\mid X,Z}^{**}$, which will be accomplished by showing both $\mathcal{P}_{\theta\mid X,Z}^{*}\subset \mathcal{P}_{\theta\mid X,Z}^{**}$ and $\mathcal{P}_{\theta\mid X,Z}^{**}\subset \mathcal{P}_{\theta\mid X,Z}^{*}$. To show $\mathcal{P}_{\theta\mid X,Z}^{*}\subset \mathcal{P}_{\theta\mid X,Z}^{**}$, fix any $P_{\theta\mid X,Z} \in \mathcal{P}_{\theta\mid X,Z}^{*}$. Then by Definition (ref) and the definition of $\mathcal{P}_{\theta \mid X,Z}^{*}$ above, there exists $\theta: \Omega \to \mathcal{U}$ with $\theta \sim P_{\theta\mid Y,X,Z}$ and an element $\theta \in \Theta$ such that:
\begin{align}
P_{\theta\mid Y,X,Z}(\theta \in \mathcal{U}(Y,X,Z,\theta) \mid Y=y, X=x, Z=z)=1, \,\, (y,x,z)-a.s.,
\end{align}
and:
\begin{align}
P_{\theta\mid X,Z}(\theta \in A \mid X=x,Z=z)&=\int P_{\theta\mid Y,X,Z}(\theta \in A \mid Y=y, X=x, Z=z)\,dP_{Y\mid X,Z},\,\, (x,z)-a.s.,
\end{align}
for every $A \in \mathfrak{B}(\mathcal{U})$. Now define the sets:
\begin{align*}
B_{1}(x,z,\theta)&:=\{\theta : \varphi(x,z,\theta,\theta) \geq 0\},\\
B_{2}(x,z,\theta)&:=\{\theta : \varphi(x,z,\theta,\theta) < 0\}.
\end{align*}
By continuity of $\varphi(x,z,\cdot,\theta)$, we have $B_{1}(x,z,\theta), B_{2}(x,z,\theta) \in \mathfrak{B}(\mathcal{U})$ for each $(x,z,\theta)$. Now for our pair $(\theta,\theta)$ we have:
\begin{align}
&P_{\theta\mid X,Z}(\theta \in B_{1}(x,z,\theta) \mid X=x,Z=z)\nonumber\\
&=\sum_{y \in \{0,1\}} P_{\theta\mid Y,X,Z}(\theta \in B_{1}(x,z,\theta) \mid Y=y, X=x, Z=z)P_{Y\mid X,Z}(Y=y \mid X=x,Z=z)\nonumber\\
&=\sum_{y \in \{0,1\}} P_{\theta\mid Y,X,Z}(\theta \in B_{1}(x,z,\theta)\cap \mathcal{U}(y,x,z,\theta) \mid Y=y, X=x, Z=z)P_{Y\mid X,Z}(Y=y \mid X=x,Z=z)\qquad\\
&=P_{\theta\mid Y,X,Z}(\theta \in \mathcal{U}(1,x,z,\theta) \mid Y=1, X=x, Z=z)P_{Y\mid X,Z}(Y=1\mid X=x,Z=z)\\
&=P_{Y\mid X,Z}(Y=1\mid X=x,Z=z),
\end{align}
$(x,z)-$a.s. Note that (ref) follows from the fact that $P_{\theta\mid Y,X,Z}(\theta \in \mathcal{U}(y,x,z,\theta) \mid Y=y, X=x, Z=z)=1$ a.s. since $P_{\theta\mid Y,X,Z} \in \mathcal{P}_{\theta\mid Y,X,Z}^{*}$ by assumption; (ref) follows from the fact that $B_{1}(x,z,\theta)\cap \mathcal{U}(0,x,z,\theta)=\emptyset$ and $B_{1}(x,z,\theta)= \mathcal{U}(1,x,z,\theta)$; (ref) follows from the fact that $P_{\theta\mid Y,X,Z}(\theta \in \mathcal{U}(1,x,z,\theta) \mid Y=1, X=x, Z=z)=1$ a.s. since $P_{\theta\mid Y,X,Z} \in \mathcal{P}_{\theta\mid Y,X,Z}^{*}$ by assumption. Repeating an identical derivation shows that:
\begin{align*}
P_{\theta\mid X,Z}(\theta \in B_{2}(x,z,\theta) \mid X=x,Z=z) = P_{Y\mid X,Z}(Y=0\mid X=x,Z=z),
\end{align*}
$(x,z)-$a.s. Since $P_{\theta\mid X,Z} \in \mathcal{P}_{\theta\mid X,Z}^{*}$ was arbitrary, this proves that $\mathcal{P}_{\theta\mid X,Z}^{*}\subset \mathcal{P}_{\theta\mid X,Z}^{**}$.
To show $\mathcal{P}_{\theta\mid X,Z}^{**}\subset \mathcal{P}_{\theta\mid X,Z}^{*}$, fix any $P_{\theta\mid X,Z} \in \mathcal{P}_{\theta\mid X,Z}^{**}$. We want to show that $P_{\theta\mid X,Z} \in \mathcal{P}_{\theta\mid X,Z}^{*}$. To do so, we must show that: (i) there exists $P_{\theta\mid Y,X,Z}$ such that:
\begin{align}
&P_{\theta\mid X,Z}(\theta \in A \mid X=x,Z=z)\nonumber\\
&=\int_{\{0,1\}} P_{\theta\mid Y,X,Z}(\theta \in A \mid Y=y, X=x, Z=z)\,dP_{Y\mid X,Z},\,\, (x,z)-a.s.,
\end{align}
for every $A \in \mathfrak{B}(\mathcal{U})$, and (ii) there is a $\theta \in \Theta$ such that:
\begin{align}
P_{\theta\mid Y,X,Z}(\theta \in \mathcal{U}(Y,X,Z,\theta) \mid Y=y, X=x, Z=z)=1, \,\, (y,x,z)-a.s.,
\end{align}
for the same $P_{\theta\mid Y,X,Z}$ from part (i). First note that, by the Radon-Nikodym Theorem, the existence of a (version of) $P_{\theta\mid Y,X,Z}$ is guaranteed by the fact that $P_{\theta,Y\mid X,Z}\ll P_{Y\mid X,Z}$. Since all spaces involved are euclidean, we can choose the version to be an almost surely unique regular conditional distribution (c.f. durrett2010probability Theorem 5.1.9). By construction $P_{\theta \mid Y, X,Z}$ satisfies:
\begin{align*}
&P_{\theta, Y\mid X,Z}(\theta \in A, Y \in B \mid X=x, Z=z)\\
&\qquad\qquad\qquad= \sum_{y \in B} P_{\theta \mid Y, X,Z}(\theta \in A \mid Y=y, X=x, Z=z)P_{Y\mid X,Z}(Y=y \mid X=x,Z=z),
\end{align*}
for every $A \in \mathfrak{B}(\mathcal{U})$ and $B \subset \{0,1\}$. This verifies part (i). It thus remains only to show that any such $P_{\theta\mid Y,X,Z}$ must also satisfy (ref). Since $P_{\theta\mid X,Z} \in \mathcal{P}_{\theta\mid X,Z}^{**}$, there exists a value $\theta \in \Theta$ such that:
\begin{align*}
P_{\theta\mid X,Z}(\varphi(x,z,\theta,\theta) \geq 0 \mid X=x,Z=z) = P_{Y\mid X,Z}(Y=1\mid X=x,Z=z),\,\,(x,z)-a.s.
\end{align*}
For this value of $\theta$, note that:
\begin{align*}
&P_{\theta\mid X,Z}(\varphi(x,z,\theta,\theta) \geq 0\mid X=x,Z=z)\\
&= P_{\theta\mid Y,X,Z}(\varphi(x,z,\theta,\theta) \geq 0\mid Y=1,X=x,Z=z)P(Y=1\mid X=x,Z=z)\\
&\qquad+ P_{\theta\mid Y,X,Z}(\varphi(x,z,\theta,\theta) \geq 0\mid Y=0,X=x,Z=z)P(Y=0\mid X=x,Z=z).
\end{align*}
Furthermore, by assumption we have:
\begin{align*}
P_{\theta\mid X,Z}(\varphi(x,z,\theta,\theta) \geq 0\mid X=x,Z=z) = P(Y=1\mid X=x,Z=z),\,\, (x,z)-a.s.
\end{align*}
Thus:
\begin{align}
&P(Y=1\mid X=x,Z=z)\nonumber\\
&= P_{\theta\mid Y,X,Z}(\varphi(x,z,\theta,\theta) \geq 0\mid Y=1,X=x,Z=z)P(Y=1\mid X=x,Z=z)\nonumber\\
&\qquad+ P_{\theta\mid Y,X,Z}(\varphi(x,z,\theta,\theta) \geq 0\mid Y=0,X=x,Z=z)P(Y=0\mid X=x,Z=z),
\end{align}
$(x,z)-$a.s. Now note by (ref):
\begin{align*}
&P_{\theta\mid Y,X,Z}(\varphi(x,z,\theta,\theta) \geq 0\mid Y=0,X=x,Z=z)P(Y=0\mid X=x,Z=z)\\
&= P_{\theta, Y\mid X,Z}(\varphi(x,z,\theta,\theta) \geq 0, Y=0 \mid X=x,Z=z)\\
&= P_{\theta \mid X,Z}(\varphi(x,z,\theta,\theta) \geq 0, \varphi(x,z,\theta,\theta)<0 \mid X=x,Z=z)\\
&=0.
\end{align*}
Conclude that (ref) is true if and only if:
\begin{align*}
P_{\theta\mid Y,X,Z}(\varphi(x,z,\theta,\theta) \geq 0\mid Y=1,X=x,Z=z)=1,\,\, (x,z)-a.s.
\end{align*}
Similar logic shows:
\begin{align*}
P_{\theta\mid Y,X,Z}(\varphi(x,z,\theta,\theta) < 0\mid Y=0,X=x,Z=z)=1,\,\, (x,z)-a.s.
\end{align*}
Finally, note that by the definition of $ \mathcal{U}(\cdot,\theta): \mathcal{Y} \times \mathcal{X} \times \mathcal{Z} \to \mathcal{U}$ we have:
\begin{align*}
P_{\theta\mid Y,X,Z}(\varphi(x,z,\theta,\theta) \geq 0\mid Y=1,X=x,Z=z) &= P_{\theta\mid Y,X,Z}(\theta \in \mathcal{U}(Y,X,Z,\theta)\mid Y=1,X=x,Z=z),\\
P_{\theta\mid Y,X,Z}(\varphi(x,z,\theta,\theta) < 0\mid Y=0,X=x,Z=z) &= P_{\theta\mid Y,X,Z}(\theta \in \mathcal{U}(Y,X,Z,\theta)\mid Y=0,X=x,Z=z).
\end{align*}
Thus we conclude that $P_{\theta\mid Y,X,Z}$ satisfies (ref). Since $P_{\theta\mid X,Z} \in \mathcal{P}_{\theta\mid X,Z}^{**}$ was arbitrary, we conclude $\mathcal{P}_{\theta\mid X,Z}^{**}\subset \mathcal{P}_{\theta\mid X,Z}^{*}$. Combining everything, we conclude $\mathcal{P}_{\theta\mid X,Z}^{**}= \mathcal{P}_{\theta\mid X,Z}^{*}$. This completes the proof.
\end{proof}
Theorem (ref) says that in order to verify whether a given collection of distributions $P_{\theta\mid X,Z}$ belongs to the identified set $\mathcal{P}_{\theta\mid X,Z}^{*}$, it suffices to find some value of the fixed coefficient $\theta \in \Theta$ such that $P_{\theta\mid X,Z}$ rationalizes the observed conditional choice probabilities via (ref). Note that Assumption (ref) does not impose any assumptions on the dependence between the variables $X$ and $Z$ and the latent variables $\theta$; in other words, this result holds whether $X$ and $Z$ are endogenous, exogenous, or any combination of the two. As discussed in chesher2014instrumental, a binary response model with endogenous regressors is \textit{incomplete} when the mechanism generating the endogenous regressors is left unspecified, as in our environment.\footnote{There competing definitions of incompleteness in the literature, although the definition of an incomplete model used in chesher2014instrumental is equivalent to the definition in tamer2003incomplete and lewbel2007coherency. The definition of an incomplete model discussed here is consistent with these papers.} In the presence of incompleteness there is no longer a unique distribution of the endogenous outcome variables given fixed primitives of the model. chesher2014instrumental propose the use of Artstein's inequalities from random set theory to characterize the distributions of selections from the incomplete binary response model in (ref), and Theorem 3.1 in chesher2014instrumental provides a general characterization of the identified set of latent variable distributions in the case of a linear index function. The key difference between Theorem (ref) above and Theorem 3.1 in chesher2014instrumental is the fact that we condition on the value of the (possibly endogenous) variables $X$ and $Z$. Conditioning on the value of the endogenous variables allows us to construct a simpler set of constraints than those imposed by Artstein's inequalities, which is demonstrated in Appendix (ref). Intuitively, conditioning on a fixed value of any endogenous regressors resolves the issue of model incompleteness. This strategy is not applicable in all environments when the model is incomplete, but appears to be applicable whenever any endogenous regressors in the model are observable. The identified set for the unconditional latent variable distribution (as was considered in chesher2014instrumental) can then be recovered from $\mathcal{P}_{\theta\mid X,Z}^{*}$.
\end{comment}
\subsection{Independence Assumptions}
Under Assumption (ref), we have the following definition of the identified set, which is analogous to both Definitions (ref) and (ref).
\begin{definition}
Under Assumptions (ref) and (ref), the identified set $\mathcal{I}_{Y,W,Z}^{*}$ is the set of all pairs $(P_{U\mid Y,W,Z},\theta)$ such that:
\begin{enumerate}[label=(\roman*)]
• $(P_{U\mid Y,W,Z},\theta)$ satisfies:
\begin{align}
P_{U\mid Y,W,Z}(U \in \mathcal{U}(Y,W,Z,\theta) \mid Y=y, W=w, Z=z)&=1,
\end{align}
$P_{Y,W,Z}-$a.s.
• The distribution $P_{U} = P_{U \mid Y,W,Z} P_{Y,W,Z}$ assigns zero probability to all sets of the form $\{u \in \mathcal{U} : \varphi(w,z,u,\theta) =0\}$.
• For all Borel sets $A \in \mathfrak{B}(\mathcal{U})$ we have $P_{U\mid Z}(A \mid Z=z) = P_{U}(A)$, $P_{Z}-$a.s.
\end{enumerate}
Furthermore, under Assumptions (ref), (ref) and (ref), the identified set of counterfactual conditional distributions $\mathcal{P}_{Y_{\gamma}\mid Y,W,Z,U}^{*}$ is the set of all conditional distributions $P_{Y_{\gamma}\mid Y,W,Z,U}$ satisfying:
\begin{align}
P_{Y_{\gamma}\mid Y, W,Z,U}\left(Y_{\gamma} = \mathbbm{1}\{\varphi(\gamma(W,Z),U,\theta) \geq 0 \}\mid Y=y, W=w,Z=z,U=u\right)=1,
\end{align}
$P_{Y,W,Z,U}-$a.s. for some pair $(P_{U\mid Y,W,Z},\theta)\in \mathcal{I}_{Y,W,Z}^{*}$.
\end{definition}
Here we do not consider the case when both Assumptions (ref) and (ref) hold, but we again note that this definition (and the results to follow) are easily modified to accommodate the case when any combination of these assumptions hold. We now provide the following Corollary whose proof follows almost identically to that of Theorems (ref) and (ref), with the exception being that we require condition (ii) of Definition (ref) to hold.
\begin{corollary}
Under Assumptions (ref), (ref) and (ref), a counterfactual conditional distribution $P_{Y_{\gamma}\mid Y,W,Z}$ satisfies $P_{Y_{\gamma}\mid Y,W,Z} \in \mathcal{P}_{Y_{\gamma}\mid Y,W,Z}^{*}$ if and only if there exists a pair $(P_{U\mid Y,W,Z},\theta) \in \mathcal{I}_{Y,W,Z}^{*}$ (for $\mathcal{I}_{Y,W,Z}^{*}$ from Definition (ref)) satisfying:
\begin{align}
P_{Y_{\gamma}\mid W,Z}\left(Y_{\gamma} = 1\mid Y=y,W=w,Z=z\right)= P_{U\mid Y,W,Z}\left( \varphi(\gamma(W,Z),U,\theta) \geq 0 \mid Y=y,W=w,Z=z\right),\qquad
\end{align}
$P_{Y,W,Z}-$a.s. Furthermore, for any collection of counterfactual conditional distributions $P_{Y_{\gamma} \mid Y,W,Z}$, there exists a collection of Borel conditional probability measures $P_{U \mid Y,W,Z}$ satisfying (ref) with $(P_{U \mid Y,W,Z},\theta) \in \mathcal{I}_{Y,W,Z}^{*}$ (for $\mathcal{I}_{Y,W,Z}^{*}$ from Definition (ref)) if and only if there exists a collection $P_{U \mid Y,W,Z}$ of probability measures on the sets in $\mathcal{A}(\theta)$ from (ref) satisfying:
\begin{align}
\sum_{s \in S_{j}} P_{U\mid Y,W,Z}\left( \text{int}(\mathcal{U}(s,\theta)) \mid Y=1,W=w_{j},Z=z_{j}\right)&=1,\\
\sum_{s \in S_{j}^{c}} P_{U\mid Y,W,Z}\left(\text{int}(\mathcal{U}(s,\theta)) \mid Y=0,W=w_{j},Z=z_{j}\right)&=1,\,\\
\sum_{s \in S_{\gamma(j)}} P_{U\mid Y,W,Z}\left( \text{int}(\mathcal{U}(s,\theta)) \mid Y=y,W=w_{j},Z=z_{j}\right)&=P_{Y_{\gamma}\mid Y,W,Z}\left( Y_{\gamma}=1 \mid Y=y,W=w_{j},Z=z_{j}\right),
\end{align}
for $y \in \{0,1\}$ and $j\in\{1,\ldots,m\}$ assigned positive probability, and:
\begin{align}
&\sum_{y }\sum_{w} P_{U\mid Y,W,Z}\left( \text{int}(\mathcal{U}(s,\theta)) \mid Y=y,W=w,Z=z_{k}\right) P(Y=y,W=w \mid Z=z_{k})\nonumber\\
&\qquad\qquad=\sum_{y }\sum_{w} P_{U\mid Y,W,Z}\left( \text{int}(\mathcal{U}(s,\theta)) \mid Y=y,W=w,Z=z_{k+1}\right)P(Y=y,W=w \mid Z=z_{k+1}),
\end{align}
for all $s \in \{0,1\}^{m}$ and all $k=1,\ldots,m_{z}-1$ assigned positive probability.
\end{corollary}
\begin{proof}[Proof of Corollary (ref)]
The first statement follows a proof identical to the proof of Theorem (ref). For the second statement, the forward direction is identical to the proof of Theorem (ref). The reverse direction is similar to the proof of Theorem (ref), with the exception that we must show that the extended measure on $\mathfrak{B}(\mathcal{U})$ satisfies independence if the intial measure on $\mathcal{A}(\mathcal{U})$ satisfies independence. Let $\tilde{P}_{U\mid Y,W,Z}$ be the extension of $P_{U\mid Y,W,Z}$ from the proof of Theorem (ref). Then for any $A \in \mathfrak{B}(\mathcal{U})$:
\begin{align*}
&\tilde{P}_{U\mid Z}(A \mid Z=z_{k})\\
&=\sum_{y \in \{0,1\}}\sum_{w \in \mathcal{W}}\sum_{s \in \{0,1\}^{m}} \mathbbm{1}(A,\theta,s) P_{U\mid Y,W,Z}\left( \text{int}(\mathcal{U}(s,\theta)) \mid Y=y,W=w,Z=z_{k}\right) P_{Y,W\mid Z}(Y=y,W=w \mid Z=z_{k})\\
&=\sum_{s \in \{0,1\}^{m}} \mathbbm{1}(A,\theta,s)\sum_{y \in \{0,1\}}\sum_{w \in \mathcal{W}} P_{U\mid Y,W,Z}\left(\text{int}(\mathcal{U}(s,\theta)) \mid Y=y,W=w,Z=z_{k}\right) P_{Y,W\mid Z}(Y=y,W=w \mid Z=z_{k})\\
&=\sum_{s \in \{0,1\}^{m}} \mathbbm{1}(A,\theta,s)\sum_{y \in \{0,1\}}\sum_{w \in \mathcal{W}} P_{U\mid Y,W,Z}\left( \text{int}(\mathcal{U}(s,\theta)) \mid Y=y,W=w,Z=z_{k+1}\right) P_{Y,W\mid Z}(Y=y,W=w \mid Z=z_{k+1})\\
&=\sum_{y \in \{0,1\}}\sum_{w \in \mathcal{W}} \sum_{s \in \{0,1\}^{m}} \mathbbm{1}(A,\theta,s) P_{U\mid Y,W,Z}\left( \text{int}(\mathcal{U}(s,\theta)) \mid Y=y,W=w,Z=z_{k+1}\right)P_{Y,W\mid Z}(Y=y,W=w \mid Z=z_{k+1})\\
&=\tilde{P}_{U\mid Z}(A \mid Z=z_{k+1}),
\end{align*}
for all pairs $z_{k}$ and $z_{k+1}$ assigned positive probability, where the third equality follows from (ref). Conclude that $\tilde{P}_{U\mid Z}$ satisfies the second condition in Definition (ref).
\end{proof}
\begin{comment}
\begin{proof}[Proof of Corollary (ref)]
The first statement follows a proof identical to the proof of Theorem (ref). For the second statement, let $P_{Y_{\gamma} \mid Y,X,Z}$ be a collection of conditional choice probabilities, and suppose there exists $(P_{\theta \mid Y,X,Z},\theta) \in \mathcal{I}_{Y,X,Z}^{*}$ satisfying (ref). Note that (ref) is equivalent to (ref), so we can conclude that $(P_{\theta \mid Y,X,Z},\theta)$ satisfies (ref). Furthermore, by definition $(P_{\theta \mid Y,X,Z},\theta) \in \mathcal{I}_{Y,X,Z}^{*}$ implies that:
\begin{align}
P_{\theta\mid Y,X,Z}(\theta \in \mathcal{U}(Y,X,Z,\theta) \mid Y=y, X=x, Z=z)=1, \,\, (y,x,z)-a.s.,
\end{align}
which is equivalent to conditions (ref) and (ref). This shows that any pair $(P_{\theta \mid Y,X,Z},\theta) \in \mathcal{I}_{Y,X,Z}^{*}$ satisfying (ref) also satisfies (ref) - (ref). Finally, note that (ref) is equivalent to:
\begin{align}
P_{\theta\mid Z}\left( \mathcal{U}(s,\theta) \mid Z=z_{k}\right)= P_{\theta\mid Z}\left( \mathcal{U}(s,\theta) \mid Z=z_{k+1}\right),
\end{align}
for all $s \in \{0,1\}^{m}$ and all $k=1,\ldots,m_{z}-1$ such that $z_{k}$ and $z_{k+1}$ are assigned positive probability. Since $(P_{\theta \mid Y,X,Z},\theta) \in \mathcal{I}_{Y,X,Z}^{*}$, for every $A \in \mathfrak{B}(\mathcal{U})$ we have that:
\begin{align*}
P_{\theta \mid Z}(A \mid Z=z)=P_{\theta}(A),
\end{align*}
$z-$a.s. Thus, since $\mathcal{U}(s,\theta) \in \mathfrak{B}(\mathcal{U})$, we have that $(P_{\theta \mid Y,X,Z},\theta) \in \mathcal{I}_{Y,X,Z}^{*}$ trivially satisfies (ref).
For the reverse, fix any $\theta\in \Theta$ and any collection $P_{\theta \mid Y,X,Z}$ of probability measures on the sets in $\mathcal{A}(\theta)$ satisfying (ref) - (ref). We will show that $P_{\theta \mid Y,X,Z}$ can be extended to a (not necessarily unique) probability measure $\tilde{P}_{\theta \mid Y,X,Z}$ on $\mathfrak{B}(\mathcal{U})$ in a manner that ensures $\tilde{P}_{\theta \mid Y,X,Z}$ satisfies (ref) and such that $(\tilde{P}_{\theta \mid Y,X,Z},\theta) \in \mathcal{I}_{Y,X,Z}^{*}$. Furthermore, by the definition of an extension, $\tilde{P}_{\theta \mid Y,X,Z}$ will agree with $P_{\theta \mid Y,X,Z}$ on all sets of the form $\mathcal{A}(\theta)$. Our extension will be similar to the one found in the proof of Theorem (ref).
To construct the extension, note that the sets in $\mathcal{A}(\theta)$ form a disjoint partition of $\mathcal{U}$. Now select a single point $\theta(s,\theta)$ from each set $\mathcal{U}(s,\theta)$ in the collection $\mathcal{A}(\theta)$; if $\mathcal{U}(s,\theta)$ is empty, choose $\theta(s,\theta)$ as an arbitrary point from $\mathcal{U}$. For any set $A \subset \mathcal{U}$, define the indicator:
\begin{align*}
\mathbbm{1}(A,\theta,s) = \mathbbm{1}\{ \theta(s,\theta)\in A\cap \mathcal{U}(s,\theta) \}
\end{align*}
Furthermore, define the function $\mu_{y,x,z}:\mathfrak{B}(\mathcal{U})\to \mathbb{R}$ as:
\begin{align*}
\mu_{y,x,z}(B) := \sum_{s \in \{0,1\}^{m}} \mathbbm{1}(B,\theta,s) P_{\theta\mid Y,X,Z}\left( \mathcal{U}(s,\theta) \mid Y=y,X=x,Z=z\right).
\end{align*}
Now follow an identical argument as in the proof of Theorem (ref) to conclude that this is a proper probability measure on $\mathfrak{B}(\mathcal{U})$ for each $(y,x,z)$ assigned positive probability.
Now take $\tilde{P}_{\theta\mid Y,X,Z} : \mathfrak{B}(\mathcal{U}) \to [0,1]$ by $\tilde{P}_{\theta\mid Y,X,Z}(B\mid Y=y,X=x,Z=z) = \mu_{y,x,z}(B)$ for all $B \in \mathfrak{B}(\mathcal{U})$ and all $(y,x,z)$ assigned positive probability. Furthermore, we define $\tilde{P}_{\theta\mid Z}$ as:
\begin{align*}
\tilde{P}_{\theta\mid Z}(B \mid Z=z):= \sum_{y \in \{0,1\}}\sum_{x \in \mathcal{X}} \tilde{P}_{\theta\mid Y,X,Z}(B\mid Y=y,X=x,Z=z)
\end{align*}
By the discussion above, $\tilde{P}_{\theta\mid Y,X,Z}(\,\cdot\,\mid Y=y,X=x,Z=z)$ is a proper probability measure on $\mathfrak{B}(\mathcal{U})$ for each $(y,x,z)$. Also note by the proof of Theorem (ref) that:
\begin{align*}
\tilde{P}_{\theta\mid Y,X,Z}(\theta \in \mathcal{U}(y,x,z,\theta) \mid Y=0,X=x,Z=z) &=1, \,\,a.s.
\end{align*}
This is the first condition in Definition (ref). Now note that for any $A \in \mathfrak{B}(\mathcal{U})$:
\begin{align*}
&\tilde{P}_{\theta\mid Z}(A \mid Z=z_{k})\\
&= \sum_{y \in \{0,1\}}\sum_{x \in \mathcal{X}}\sum_{s \in \{0,1\}^{m}} \mathbbm{1}(A,\theta,s) P_{\theta\mid Y,X,Z}\left( \mathcal{U}(s,\theta) \mid Y=y,X=x,Z=z_{k}\right) P_{Y,X\mid Z}(Y=y,X=x \mid Z=z_{k})\\
&= \sum_{s \in \{0,1\}^{m}} \mathbbm{1}(A,\theta,s)\sum_{y \in \{0,1\}}\sum_{x \in \mathcal{X}} P_{\theta\mid Y,X,Z}\left( \mathcal{U}(s,\theta) \mid Y=y,X=x,Z=z_{k}\right) P_{Y,X\mid Z}(Y=y,X=x \mid Z=z_{k})\\
&= \sum_{s \in \{0,1\}^{m}} \mathbbm{1}(A,\theta,s)\sum_{y \in \{0,1\}}\sum_{x \in \mathcal{X}} P_{\theta\mid Y,X,Z}\left( \mathcal{U}(s,\theta) \mid Y=y,X=x,Z=z_{k+1}\right) P_{Y,X\mid Z}(Y=y,X=x \mid Z=z_{k+1})\\
&= \sum_{y \in \{0,1\}}\sum_{x \in \mathcal{X}} \sum_{s \in \{0,1\}^{m}} \mathbbm{1}(A,\theta,s) P_{\theta\mid Y,X,Z}\left( \mathcal{U}(s,\theta) \mid Y=y,X=x,Z=z_{k+1}\right)P_{Y,X\mid Z}(Y=y,X=x \mid Z=z_{k+1})\\
&=\tilde{P}_{\theta\mid Z}(A \mid Z=z_{k+1}),
\end{align*}
for all pairs $z_{k}$ and $z_{k+1}$ assigned positive probability, where the third equality follows from (ref). Conclude that $\tilde{P}_{\theta\mid Z}$ satisfies the second condition in Definition (ref). Thus, we have verified that $(\tilde{P}_{\theta\mid Y, X,Z},\theta) \in \mathcal{I}_{Y,X,Z}^{*}$ where $\mathcal{I}_{Y,X,Z}^{*}$ is from Definition (ref). Finally, setting $C:=\{ \theta : \varphi(\gamma(x,z),\theta,\theta) \geq 0 \}$, it is straightforward to show that:
\begin{align*}
\tilde{P}_{\theta\mid Y,X,Z}\left( C \mid Y=y, X=x,Z=z\right)&= \sum_{s \in \{0,1\}^{m}} \mathbbm{1}(C,\theta,s) P_{\theta\mid Y,X,Z}\left( \mathcal{U}(s,\theta) \mid Y=y,X=x,Z=z\right)\\
&=\sum_{s \in S_{\gamma(j)}} P_{\theta\mid Y,X,Z}\left( \mathcal{U}(s,\theta) \mid Y=y,X=x_{j},Z=z_{j}\right)\\
&=P_{Y_{\gamma}\mid Y,X,Z}\left( Y_{\gamma}=1 \mid Y=y,X=x_{j},Z=z_{j}\right),
\end{align*}
for all $(y,x_{j},z_{j})$ assigned positive probability, which follows from (ref). This is exactly condition (ref). Conclude that $(\tilde{P}_{\theta\mid Y, X,Z},\theta) \in \mathcal{I}_{Y,X,Z}^{*}$ and that $(\tilde{P}_{\theta\mid Y,X,Z},\theta)$ satisfies (ref). This completes the proof.
\end{proof}
\end{comment}
Analogous to Theorem (ref), the first part of Corollary (ref) provides the theoretical link between the identified set for counterfactual conditional distributions and the identified set for the pair $(P_{U\mid Y,W,Z},\theta)$ under the additional independence assumption between $U$ and $Z$. Furthermore, analogous to the result in Theorem (ref), the second part of Corollary (ref) reduces an infinite dimensional existence problem to a finite dimensional existence problem. Importantly, the second part of Corollary (ref) builds on Theorem (ref) by demonstrating that Assumption (ref)---which requires $P_{U\mid Z}(A \mid Z=z) = P_{U}(A)$ a.s. for all Borel sets $A$---can be imposed by considering only a finite number of equality constraints on a distribution $P_{U\mid Y,W,Z}$ defined on sets of the form $\mathcal{U}(s,\theta)$.
We have the following Corollary to Theorem (ref):
\begin{corollary}
Under Assumptions (ref), (ref), and (ref), the identified set for the counterfactual conditional probability $P_{Y_{\gamma}\mid Y,W,Z}(Y_{\gamma} = 1\mid Y=y, W=w_{j},Z=z_{j})$ is given by:
\begin{align}
\bigcup_{\theta \in \Theta} [\pi_{\ell b}(y,w_{j},z_{j},\theta),\pi_{u b}(y,w_{j},z_{j},\theta)],
\end{align}
where $\pi_{\ell b}(y,w_{j},z_{j},\theta)$ and $\pi_{u b}(y,w_{j},z_{j},\theta)$ are determined by the optimization problems:
\begin{align}
\pi_{\ell b}(y,w_{j},z_{j},\theta) &:= \min_{\pi(\theta)\in \mathbb{R}^{d_{\pi}}} \sum_{s \in S_{\gamma(j)}} \pi(y,w_{j},z_{j},s,\theta),\text{ s.t. (ref), (ref), (ref), and (ref),}\\
\pi_{u b}(y,w_{j},z_{j},\theta) &:= \max_{\pi(\theta)\in \mathbb{R}^{d_{\pi}}} \sum_{s \in S_{\gamma(j)}} \pi(y,w_{j},z_{j},s,\theta),\text{ s.t. (ref), (ref), (ref), and (ref).}
\end{align}
\end{corollary}
Note that this Corollary is identical to Theorem (ref) with the exception that we have imposed Assumption (ref), and thus have included constraints of the form in (ref). With the exception of these additional constraints, the optimization problems that characterize the bounding problem are the same as before. Again, this result can be easily modified to bound any linear function of counterfactual conditional distributions by simply modifying the objective function in the optimization problems (ref) and (ref).
\begin{comment}
Finally, we present the following corollary to Proposition (ref).
\begin{corollary}
Suppose that Assumptions (ref), (ref) and (ref) hold. Then there exists a (not necessarily unique) finite subset $\Theta' \subset \Theta$ such that:
\begin{align*}
&\left\{ \overline{\pi} \in \mathbb{R}^{d_{\pi}} : \exists \theta \in \Theta \text{ s.t. } \pi(\theta) \text{ satisfies (ref), (ref), (ref), (ref) and }\overline{\pi}=\pi(\theta) \right\}\\
&\qquad\qquad\qquad= \left\{ \overline{\pi} \in \mathbb{R}^{d_{\pi}} : \exists \theta \in \Theta' \text{ s.t. } \pi(\theta) \text{ satisfies (ref), (ref), (ref), (ref) , and }\overline{\pi}=\pi(\theta) \right\}.
\end{align*}
\end{corollary}
The additional independence constraints in (ref) do not affect the proof of Proposition (ref) in any way, and so the proof of this result is identical to the proof of Proposition (ref).
\end{comment}
\subsection{Monotonicity Assumptions}
When we entertain Assumption (ref), we have the following definition of the identified set, which is analogous to both Definitions (ref) and (ref).
\begin{definition}
Under Assumptions (ref) and (ref), the identified set $\mathcal{I}_{Y,X}^{*}$ is the set of all pairs $(P_{U\mid Y,X},\theta)$ such that:
\begin{enumerate}[label=(\roman*)]
• $(P_{U\mid Y,X},\theta)$ satisfies:
\begin{align}
P_{U\mid Y,X}(U \in \mathcal{U}(Y,X,\theta) \mid Y=y, X=x)=1,
\end{align}
$P_{Y,X}-$a.s.
• The distribution $P_{U} = P_{U\mid Y,X} P_{Y,X}$ assigns zero probability to all sets of the form $\{u \in \mathcal{U} : \varphi(x,u,\theta)=0 \}$.
• For all $(j,k) \in \mathcal{M}$ from Assumption (ref), we have:
\begin{align}
P_{U\mid Y,X}(\varphi(x_{j},U,\theta) \leq \varphi(x_{k},U,\theta) \mid Y=y, X=x)=1 \text{ a.s.}
\end{align}
\end{enumerate}
Furthermore, under Assumptions (ref), (ref), and (ref), the identified set of counterfactual conditional distributions $\mathcal{P}_{Y_{\gamma}\mid Y,X,U}^{*}$ is the set of all conditional distributions $P_{Y_{\gamma}\mid Y,X,U}$ satisfying:
\begin{align}
P_{Y_{\gamma}\mid Y,X,U}\left(Y_{\gamma} = \mathbbm{1}\{\varphi(\gamma(X),U,\theta) \geq 0 \}\mid Y=y, X=x,U=u\right)=1,
\end{align}
$P_{Y,X,U}-$a.s. for some pair $(P_{U\mid Y,X},\theta)\in \mathcal{I}_{Y,X}^{*}$.
\end{definition}
Again, this definition and the results to follow are easily modified to accommodate the case when any combination of Assumptions (ref) and (ref) hold. We now provide the following Corollary whose proof follows almost identically to that of Theorems (ref) and (ref), with the exception being that we require condition (ii) of Definition (ref) to hold.
\begin{corollary}
Under Assumptions (ref), (ref), and (ref), a counterfactual conditional distribution $P_{Y_{\gamma}\mid Y,X}$ satisfies $P_{Y_{\gamma}\mid Y,X} \in \mathcal{P}_{Y_{\gamma}\mid Y,X}^{*}$ if and only if there exists a pair $(P_{U\mid Y,X},\theta) \in \mathcal{I}_{Y,X}^{*}$ (for $\mathcal{I}_{Y,X}^{*}$ from Definition (ref)) satisfying:
\begin{align}
P_{Y_{\gamma}\mid Y,X}\left(Y_{\gamma} = 1\mid Y=y,X=x\right)= P_{U\mid Y,X}\left( \varphi(\gamma(X),U,\theta) \geq 0 \mid Y=y,X=x\right),
\end{align}
$P_{Y,X}-$a.s. Furthermore, for any collection of counterfactual conditional distributions $P_{Y_{\gamma} \mid Y,X}$, there exists a collection of Borel conditional probability measures $P_{U \mid Y,X}$ satisfying (ref) with $(P_{U \mid Y,X},\theta) \in \mathcal{I}_{Y,X}^{*}$ (for $\mathcal{I}_{Y,X}^{*}$ from Definition (ref)) if and only if there exists a collection $P_{U \mid Y,X}$ of probability measures on the sets in $\mathcal{A}(\theta)$ from (ref) satisfying:
\begin{align}
\sum_{s \in S_{j}} P_{U\mid Y,X}\left( \text{int}(\mathcal{U}(s,\theta))\mid Y=1,X=x_{j}\right)&=1,\\
\sum_{s \in S_{j}^{c}} P_{U\mid Y,X}\left( \text{int}(\mathcal{U}(s,\theta)) \mid Y=0,X=x_{j}\right)&=1,\\
\sum_{s \in S_{\gamma(j)}} P_{U\mid Y,X}\left( \text{int}(\mathcal{U}(s,\theta)) \mid Y=y,X=x_{j}\right)&=P_{Y_{\gamma}\mid Y,X}\left( Y_{\gamma}=1 \mid Y=y,X=x_{j}\right),
\end{align}
for $y \in \{0,1\}$ and $j\in \{1,\ldots,m\}$ assigned positive probability, and:
\begin{align}
\sum_{s \in S_{M}^{c}} P_{U\mid Y,X}\left(\text{int}(\mathcal{U}(s,\theta))\mid Y=y,X=x\right) = 0, \,\,a.s.
\end{align}
for all $(y,x)$ assigned positive probability, where $S_{M}$ is as defined in Section (ref).
\end{corollary}
The proof of this corollary is identical to the proof of Theorem (ref) and Theorem (ref). Analogous to Theorem (ref), the first part of Corollary (ref) provides the theoretical link between the identified set for counterfactual conditional distributions and the identified set for the pair $(P_{U\mid Y,X},\theta)$ under the additional monotonicity assumption. Analogous to Theorem (ref), the second part of Corollary (ref) reduces an infinite dimensional existence problem to a finite dimensional existence problem amenable to analysis using optimization problems. Building on the intuition provided in Example (ref), the second part of Corollary (ref) demonstrates that monotonicity as in Assumption (ref) can be imposed by considering only a finite number of equality constraints on a distribution $P_{U\mid Y,X}$ defined on sets of the form $\mathcal{U}(s,\theta)$. By definition of the set $S_M$, condition (ref) simply assigns probability zero to all sets $\mathcal{U}(s,\theta)$ that do not satisfy the monotonicity relation from Assumption (ref). This leads to the following result.
\begin{corollary}
Under Assumptions (ref), (ref), and (ref), the identified set for the counterfactual conditional probability $P_{Y_{\gamma}\mid Y,X}(Y_{\gamma} = 1\mid Y=y, X=x_{j})$ is given by:
\begin{align}
\bigcup_{\theta \in \Theta} [\pi_{\ell b}(y,x_{j},\theta),\pi_{u b}(y,x_{j},\theta)],
\end{align}
where $\pi_{\ell b}(y,x_{j},\theta)$ and $\pi_{u b}(y,x_{j},\theta)$ are determined by the optimization problems:
\begin{align}
\pi_{\ell b}(y,x_{j},\theta) &:= \min_{\pi(\theta)\in \mathbb{R}^{d_{\pi}}} \sum_{s \in S_{\gamma(j)}} \pi(y,x_{j},s,\theta),\text{ s.t. (ref), (ref), (ref), and (ref),}\\
\pi_{u b}(y,x_{j},\theta) &:= \max_{\pi(\theta)\in \mathbb{R}^{d_{\pi}}} \sum_{s \in S_{\gamma(j)}} \pi(y,x_{j},s,\theta),\text{ s.t. (ref), (ref), (ref), and (ref).}
\end{align}
\end{corollary}
Note that this Corollary is identical to Theorem (ref) with the exception that we have imposed Assumption (ref), and thus have included constraints of the form (ref). With the exception of these additional constraints, the optimization problems that characterize the bounding problem are the same as before. Finally, alternative counterfactual quantities can be bounded in the same way by simply modifying the objective function in (ref) and (ref).
\begin{comment}
Finally, we present the following corollary to Proposition (ref).
\begin{corollary}
Suppose that Assumptions (ref), (ref) and (ref) hold. Then there exists a (not necessarily unique) finite subset $\Theta' \subset \Theta$ such that:
\begin{align*}
&\left\{ \overline{\pi} \in \mathbb{R}^{d_{\pi}} : \exists \theta \in \Theta \text{ s.t. } \pi(\theta) \text{ satisfies (ref), (ref), (ref), (ref) and }\overline{\pi}=\pi(\theta) \right\}\\
&\qquad\qquad\qquad= \left\{ \overline{\pi} \in \mathbb{R}^{d_{\pi}} : \exists \theta \in \Theta' \text{ s.t. } \pi(\theta) \text{ satisfies (ref), (ref), (ref), (ref), and }\overline{\pi}=\pi(\theta) \right\}.
\end{align*}
\end{corollary}
The additional monotonicity constraints in (ref) are easily accommodated by in the proof of Proposition (ref), and so the proof of this result follows the same proof of Proposition (ref).
\end{comment}
\subsection{Consistency}
In this subsection we present a consistency result for functionals of a partially identified parameter related to results found in molchanov1998limit, manski2002inference, and chernozhukov2007estimation. It is presented in a form that is more general than necessary for the current paper (and more general than Proposition (ref) in Section (ref)), and so it may be of interest in other applications. We consider an environment where the researcher wishes to compute bounds on a functional $\E_{P}[\psi(X_{i},\tau_{1},\tau_{2})]$, where $\psi: \mathcal{X} \times \mathcal{T} \to \mathbb{R}$, $\mathcal{X}\subset \mathbb{R}^{d_{x}}$ denotes the support of the observed random vector $X_{i}$, and $\mathcal{T} = \mathcal{T}_{1}\times \mathcal{T}_{2} \subset \mathbb{R}^{d_{\tau}}$ denotes the parameter space with typical elements $\tau=(\tau_{1},\tau_{2}) \in \mathcal{T}$. The values of $(\tau_{1},\tau_{2})$ are constrained by moment inequalities of the form:
\begin{align*}
\E_{P}[m_{j}(X_{i},\tau_{1},\tau_{2})] \leq 0, \text{ for }j \in \mathcal{J}(\tau_{2}),
\end{align*}
where $\mathcal{J}(\tau_{2})$ is a finite index set that may depend on $\tau_{2}$. Note this does not rule out moment equalities, since each moment equality can be equivalently written as a combination of two moment inequalities. In this environment, the identified set for $(\tau_{01},\tau_{02}) \in \mathcal{T}$ at the true $P$ is given by:
\begin{align*}
\mathcal{T}^{*}(P)&:= \left\{ (\tau_{1},\tau_{2}) \in \mathcal{T} : \E_{P}[m_{j}(X_{i},\tau_{1},\tau_{2})] \leq 0 \text{ for $j\in \mathcal{J}(\tau_{2})$} \right\}.
\end{align*}
In addition, the identified set for $\psi_{0}:= \E_{P}[\psi(X_{i},\tau_{01},\tau_{02})]$ is given by:
\begin{align*}
\Psi^{*}(P)&:= \left\{ \overline{\psi} \in \mathbb{R} : \exists (\tau_{1},\tau_{2}) \in \mathcal{T}^{*}(P) \text{ s.t. } \overline{\psi}= \E_{P}[\psi(X_{i},\tau_{1},\tau_{2})] \right\}.
\end{align*}
Let us define the projection:
\begin{align*}
\mathcal{T}_{1}^{*}(\tau_{2},P)&:= \left\{ \tau_{1} \in \mathcal{T}_{1} : \E_{P}[m_{j}(X_{i},\tau_{1},\tau_{2})] \leq 0 \text{ for $j\in \mathcal{J}(\tau_{2})$} \right\}.
\end{align*}
Under Assumption (ref) ahead, it is straightforward to show that $\Psi^{*}(P)$ can be rewritten as:
\begin{align*}
\Psi^{*}(P) = \bigcup_{\tau_{2} \in \mathcal{T}_{2}} [\Psi_{\ell b}(\tau_{2},P), \Psi_{u b}(\tau_{2},P)],
\end{align*}
where:
\begin{align*}
\Psi_{\ell b}(\tau_{2},P):= \min_{\tau_{1} \in \mathcal{T}_{1}^{*}(\tau_{2},P)} \E_{P}[\psi(X_{i},\tau_{1},\tau_{2})], && \Psi_{u b}(\tau_{2},P):= \max_{\tau_{1} \in \mathcal{T}_{1}^{*}(\tau_{2},P)} \E_{P}[\psi(X_{i},\tau_{1},\tau_{2})].
\end{align*}
We study the consistency properties of the sample analog estimator for this representation of $\Psi^{*}(P)$. In particular, define:
\begin{align*}
\E_{n}[\psi(X_{i},\tau_{1},\tau_{2})]:= \frac{1}{n} \sum_{i=1}^{n} \psi(X_{i},\tau_{1},\tau_{2}), && \E_{n}[m_{j}(X_{i},\tau_{1},\tau_{2})]:= \frac{1}{n} \sum_{i=1}^{n} m_{j}(X_{i},\tau_{1},\tau_{2}), \text{ for }j\in \mathcal{J}(\tau_{2}).
\end{align*}
Then the sample analog estimator of interest is given by:
\begin{align*}
\Psi^{*}(\P_{n}) = \bigcup_{\tau_{2} \in \mathcal{T}_{2}} [\Psi_{\ell b}(\tau_{2},\P_{n}), \Psi_{u b}(\tau_{2},\P_{n})],
\end{align*}
where:
\begin{align*}
\Psi_{\ell b}(\tau_{2},\P_{n}):= \min_{\tau_{1} \in \mathcal{T}_{1}^{*}(\tau_{2},\P_{n})} \E_{n}[\psi(X_{i},\tau_{1},\tau_{2})], && \Psi_{u b}(\tau_{2},\P_{n}):= \max_{\tau_{1} \in \mathcal{T}_{1}^{*}(\tau_{2},\P_{n})} \E_{n}[\psi(X_{i},\tau_{1},\tau_{2})],
\end{align*}
and:
\begin{align*}
\mathcal{T}_{1}^{*}(\tau_{2},\P_{n})&:= \left\{ \tau_{1} \in \mathcal{T}_{1} : \E_{n}[m_{j}(X_{i},\tau_{1},\tau_{2})] \leq 0 \text{ for $j=1,\ldots,J$} \right\}.
\end{align*}
In the following, we define the sequence $\{\eta_{n}(\tau_{2})\}_{n=1}^{\infty}$ as:
\begin{align*}
\eta_{n}(\tau_{2}) := \max\left\{ \max_{j\in \mathcal{J}(\tau_{2})} \sup_{\tau_{1} \in \mathcal{T}_{1}} | \E_{n}[m_{j}(X_{i},\tau_{1},\tau_{2})] - \E_{P}[m_{j}(X_{i},\tau_{1},\tau_{2})]|, \sup_{\tau_{1} \in \mathcal{T}_{1}} \left| \E_{n}[\psi(X_{i},\tau_{1},\tau_{2})] - \E_{P}[\psi(X_{i},\tau_{1},\tau_{2})] \right|\right\}.
\end{align*}
We impose the following assumption.
\begin{assumption}
The parameter space $(\mathcal{T},\mathcal{P})$ satisfies the following: (i) $\mathcal{T}:=\mathcal{T}_{1}\times\mathcal{T}_{2} \subset \mathbb{R}^{d_{\tau}}$, where $\mathcal{T}_{1}$ is compact and convex; (ii) for each $\tau_{2} \in \mathcal{T}_{2}$, the function $\psi(\,\cdot\,,\tau_{2}): \mathcal{X} \times \mathcal{T}_{1} \to \mathbb{R}$ is measurable in $X_{i} \in \mathcal{X}\subset \mathbb{R}^{d_{x}}$ and is Lipschitz continuous in $\tau_{1}$ with a (possibly data-dependent) Lipschitz constant $C(\tau_{2})$ with $\sup_{\tau_{2} \in \mathcal{T}_{2}} C(\tau_{2})<\infty$ a.s.; (iii) for each $\tau_{2} \in \mathcal{T}_{2}$ and $j\in \mathcal{J}(\tau_{2})$, the moment function $m_{j}(\,\cdot\,,\tau_{2}): \mathcal{X} \times \mathcal{T}_{1} \to \mathbb{R}$ is measurable in $X_{i}$ and convex (and thus continuous) in $\tau_{1}$; (iv) for each $P \in \mathcal{P}$ there exists a $(\tau_{1},\tau_{2}) \in \mathcal{T}$ such that $\E_{P}[m_{j}(X_{i},\tau_{1},\tau_{2})] \leq 0, \text{ for }j\in \mathcal{J}(\tau_{2})$; (v) for each fixed $\tau_{2} \in \mathcal{T}_{2}$, we have $\eta_{n}(\tau_{2}) = O_{P}(a_{n}^{-1})$ for some sequence $a_{n} \uparrow \infty$; (vi) for each fixed $\tau_{2} \in \mathcal{T}_{2}$, there exists a sequence $b_{n} \downarrow 0$ satisfying $b_{n} \geq \eta_{n}(\tau_{2})$ w.p.a. 1; (vii) there exists a finite subset $\mathcal{T}_{2}' \subset \mathcal{T}_{2}$ such that:
\begin{align*}
&\left\{ \tau_{1} \in \mathcal{T}_{1} : \exists \tau_{2} \in \mathcal{T}_{2} \text{ s.t. } \E_{P}[m_{j}(X_{i},\tau_{1},\tau_{2})] \leq 0 \text{ for $j\in \mathcal{J}(\tau_{2})$} \right\}\\
&\qquad\qquad\qquad= \left\{ \tau_{1} \in \mathcal{T}_{1} : \exists \tau_{2} \in \mathcal{T}_{2}' \text{ s.t. } \E_{P}[m_{j}(X_{i},\tau_{1},\tau_{2})] \leq 0 \text{ for $j\in \mathcal{J}(\tau_{2})$} \right\}.
\end{align*}
Finally, the sample $\{X_{i}\}_{i=1}^{n}$ is given by $n$ i.i.d. draws from some $P \in \mathcal{P}$.
\end{assumption}
For any $c \in \mathbb{R}$ let us define:
\begin{align*}
\mathcal{T}_{1}^{*}(\tau_{2},P,c)&:= \left\{ \tau_{1} \in \mathcal{T}_{1} : \E_{P}[m_{j}(X_{i},\tau_{1},\tau_{2})] \leq c \text{ for $j\in \mathcal{J}(\tau_{2})$} \right\},
\end{align*}
and:
\begin{align*}
\Psi^{*}(P,c) = \bigcup_{\tau_{2} \in \mathcal{T}_{2}'} [\Psi_{\ell b}(\tau_{2},P,c), \Psi_{u b}(\tau_{2},P,c)],
\end{align*}
where:
\begin{align*}
\Psi_{\ell b}(\tau_{2},P,c):= \min_{\tau_{1} \in \mathcal{T}_{1}^{*}(\tau_{2},P,c)} \E_{P}[\psi(X_{i},\tau_{1},\tau_{2})], && \Psi_{u b}(\tau_{2},\P_{n},c):= \max_{\tau_{1} \in \mathcal{T}_{1}^{*}(\tau_{2},P,c)} \E_{P}[\psi(X_{i},\tau_{1},\tau_{2})].
\end{align*}
Define the sets $\mathcal{T}_{1}^{*}(\tau_{2},P,c)$ and $\Psi^{*}(P,c)$ analogously. The following Theorem then shows that a slight enlargement of the set $\Psi^{*}(\P_{n})$ is a consistent estimator for the set $\Psi^{*}(P)$, where consistency is defined using the Hausdorff metric.
\begin{theorem}
Suppose that Assumption (ref) holds. Then $d_{H}(\Psi^{*}(\P_{n},b_{n}),\Psi^{*}(P)) = o_{P}(1)$, where $b_{n}$ is any sequence satisfying Assumption (ref).
\end{theorem}
\begin{proof}[Proof of Theorem (ref)]
We have:
\begin{align*}
d_{H}(\Psi^{*}(\P_{n},b_{n}),\Psi^{*}(P)) \leq \sum_{\tau_{2} \in \mathcal{T}_{2}'}d_{H}\left([\Psi_{\ell b}(\tau_{2},\P_{n},b_{n}), \Psi_{u b}(\tau_{2},\P_{n},b_{n})], [\Psi_{\ell b}(\tau_{2},P), \Psi_{u b}(\tau_{2},P)]\right).
\end{align*}
Since $\mathcal{T}_{2}'$ is finite, it suffices to show that:
\begin{align*}
d_{H}\left([\Psi_{\ell b}(\tau_{2},\P_{n},b_{n}), \Psi_{u b}(\tau_{2},\P_{n},b_{n})], [\Psi_{\ell b}(\tau_{2},P), \Psi_{u b}(\tau_{2},P)]\right)=o_{P}(1),
\end{align*}
for each $\tau_{2} \in \mathcal{T}_{2}'$. To this end, fix any $\tau_{2} \in \mathcal{T}_{2}'$. To show the previous display, it suffices to show consistency of the upper and lower bounds; i.e. that $|\Psi_{\ell b}(\tau_{2},\P_{n},b_{n}) - \Psi_{\ell b}(\tau_{2},P)| = o_{P}(1)$ and that $|\Psi_{u b}(\tau_{2},\P_{n},b_{n}) - \Psi_{u b}(\tau_{2},P)| = o_{P}(1)$. We focus on the lower bound, since the upper bound proof is symmetric.
First recall that $\psi(X_{i},\tau_{1},\tau_{2})$ is continuous with respect to $\tau_{1}$ for every $\tau_{2}$ by Assumption (ref), and $\mathcal{T}_{1}$ is compact. Thus, we have that $\psi(X_{i},\tau_{1},\tau_{2})$ is uniformly continuous (w.r.t. $\tau_{1}$) on $\mathcal{T}_{1}$. Thus, for every $\varepsilon>0$ there exists a $\delta(\varepsilon)>0$ such that $|\E_{n}[\psi(X_{i},\tau_{1},\tau_{2})] - \E_{n}[\psi(X_{i},\tau_{1}',\tau_{2})]|< \varepsilon$ whenever $||\tau_{1} - \tau_{1}'|| < \delta(\varepsilon)$. Now note that:
\begin{align*}
&|\Psi_{\ell b}(\tau_{2},\P_{n},b_{n}) - \Psi_{\ell b}(\tau_{2},P)|\\
&= \left|\min_{\tau_{1} \in \mathcal{T}_{1}^{*}(\tau_{2},\P_{n},b_{n})} \E_{n}[\psi(X_{i},\tau_{1},\tau_{2})] - \min_{\tau_{1} \in \mathcal{T}_{1}^{*}(\tau_{2},P)} \E_{P}[\psi(X_{i},\tau_{1},\tau_{2})] \right|,\\
&\leq \left|\min_{\tau_{1} \in \mathcal{T}_{1}^{*}(\tau_{2},\P_{n},b_{n})} \E_{n}[\psi(X_{i},\tau_{1},\tau_{2})] - \min_{\tau_{1} \in \mathcal{T}_{1}^{*}(\tau_{2},P)} \E_{n}[\psi(X_{i},\tau_{1},\tau_{2})] \right|\\
&\qquad\qquad\qquad+ \left|\min_{\tau_{1} \in \mathcal{T}_{1}^{*}(\tau_{2},P)} \E_{n}[\psi(X_{i},\tau_{1},\tau_{2})] - \min_{\tau_{1} \in \mathcal{T}_{1}^{*}(\tau_{2},P)} \E_{P}[\psi(X_{i},\tau_{1},\tau_{2})] \right|,\\
&= \left|\max_{\tau_{1} \in \mathcal{T}_{1}^{*}(\tau_{2},P)} - \E_{n}[\psi(X_{i},\tau_{1},\tau_{2})] - \max_{\tau_{1} \in \mathcal{T}_{1}^{*}(\tau_{2},\P_{n},b_{n})} -\E_{n}[\psi(X_{i},\tau_{1},\tau_{2})] \right|\\
&\qquad\qquad\qquad+ \left|\max_{\tau_{1} \in \mathcal{T}_{1}^{*}(\tau_{2},P)} -\E_{P}[\psi(X_{i},\tau_{1},\tau_{2})] -\max_{\tau_{1} \in \mathcal{T}_{1}^{*}(\tau_{2},P)} -\E_{n}[\psi(X_{i},\tau_{1},\tau_{2})] \right|,\\
&\leq \max_{\{\tau_{1},\tau_{1}' \in \mathcal{T}_{1} : ||\tau_{1} - \tau_{1}'|| \leq d_{H}(\mathcal{T}_{1}^{*}(\tau_{2},\P_{n},b_{n}),\mathcal{T}_{1}^{*}(\tau_{2},P))\}}\left| - \E_{n}[\psi(X_{i},\tau_{1},\tau_{2})] - -\E_{n}[\psi(X_{i},\tau_{1}',\tau_{2})] \right|\\
&\qquad\qquad\qquad+ \max_{\tau_{1} \in \mathcal{T}_{1}^{*}(\tau_{2},P)} \left| -\E_{n}[\psi(X_{i},\tau_{1},\tau_{2})] - - \E_{P}[\psi(X_{i},\tau_{1},\tau_{2})] \right|\\
&= \max_{\{\tau_{1},\tau_{1}' \in \mathcal{T}_{1} : ||\tau_{1} - \tau_{1}'|| \leq d_{H}(\mathcal{T}_{1}^{*}(\tau_{2},\P_{n},b_{n}),\mathcal{T}_{1}^{*}(\tau_{2},P))\}}\left| \E_{n}[\psi(X_{i},\tau_{1}',\tau_{2})] - \E_{n}[\psi(X_{i},\tau_{1},\tau_{2})] \right|\\
&\qquad\qquad\qquad+ \max_{\tau_{1} \in \mathcal{T}_{1}^{*}(\tau_{2},P)} \left| \E_{P}[\psi(X_{i},\tau_{1},\tau_{2})] -\E_{n}[\psi(X_{i},\tau_{1},\tau_{2})] \right|\\
&\leq \max_{\{\tau_{1},\tau_{1}' \in \mathcal{T}_{1} : ||\tau_{1} - \tau_{1}'|| \leq d_{H}(\mathcal{T}_{1}^{*}(\tau_{2},\P_{n},b_{n}),\mathcal{T}_{1}^{*}(\tau_{2},P))\}} C(\tau_{2}) \cdot ||\tau_{1} - \tau_{1}'|| + \max_{\tau_{1} \in \mathcal{T}_{1}^{*}(\tau_{2},P)} \left|\E_{P}[\psi(X_{i},\tau_{1},\tau_{2})] - \E_{n}[\psi(X_{i},\tau_{1},\tau_{2})] \right|\\
&\leq C(\tau_{2}) \cdot d_{H}(\mathcal{T}_{1}^{*}(\tau_{2},\P_{n},b_{n}),\mathcal{T}_{1}^{*}(\tau_{2},P))+ \max_{\tau_{1} \in \mathcal{T}_{1}^{*}(\tau_{2},P)} \left|\E_{P}[\psi(X_{i},\tau_{1},\tau_{2})] - \E_{n}[\psi(X_{i},\tau_{1},\tau_{2})] \right|.
\end{align*}
It suffices to show the two terms in the last line of the previous display converge to zero in probability. The second term converges in probability to zero by Assumption (ref)(vi). Furthermore, since $\sup_{\tau_{2} \in \mathcal{T}_{2}} C(\tau_{2})<\infty$ a.s., the first term converges to zero in probability if we can show that:
\begin{align*}
d_{H}(\mathcal{T}_{1}^{*}(\tau_{2},\P_{n},b_{n}),\mathcal{T}_{1}^{*}(\tau_{2},P)) = o_{P}(1).
\end{align*}
The remainder of the proof focuses on proving this latter fact. Note that:
\begin{align*}
d_{H}(\mathcal{T}_{1}^{*}(\tau_{2},\P_{n},b_{n}),\mathcal{T}_{1}^{*}(\tau_{2},P))&= \inf\{\delta>0 : \mathcal{T}_{1}^{*}(\tau_{2},P) \subset \mathcal{T}_{1}^{*}(\tau_{2},\P_{n},b_{n})^\delta, \text{ and } \mathcal{T}_{1}^{*}(\tau_{2},\P_{n},b_{n}) \subset \mathcal{T}_{1}^{*}(\tau_{2},P)^\delta\},
\end{align*}
where:
\begin{align*}
\mathcal{T}_{1}^{*}(\tau_{2},\P_{n},b_{n})^\delta&:= \{\tau_{1} \in \mathcal{T}_{1} : B_{\delta}(\tau_{1})\cap \mathcal{T}_{1}^{*}(\tau_{2},\P_{n},b_{n})\neq \emptyset \},\\
\mathcal{T}_{1}^{*}(\tau_{2},P)^\delta&:= \{\tau_{1} \in \mathcal{T}_{1} : B_{\delta}(\tau_{1})\cap \mathcal{T}_{1}^{*}(\tau_{2},P)\neq \emptyset \},
\end{align*}
where $B_{\delta}(\tau_{1})$ denotes the closed ball of radius $\delta>0$ around $\tau_{1}$. The next part of the proof closely follows the proof of Theorem 2.1 in molchanov1998limit. Define the function:
\begin{align*}
\rho(\varepsilon):= d_{H}(\mathcal{T}_{1}^{*}(\tau_{2},P,\varepsilon),\mathcal{T}_{1}^{*}(\tau_{2},P)).
\end{align*}
Since each of the moment functions are convex (and thus lower semi-continuous) in $\tau_{1}$ for each $\tau_{2}$, each of the sets $\mathcal{T}_{1}^{*}(\tau_{2},P,\varepsilon)$ and $\mathcal{T}_{1}^{*}(\tau_{2},P)$ are closed and $\rho$ is right continuous. Furthermore, $\rho$ is non-increasing for $\varepsilon<0$ and non-decreasing for $\varepsilon>0$. Now by Assumption (ref) we have with high probability:
\begin{align*}
\mathcal{T}_{1}^{*}(\tau_{2},\P_{n},b_{n}) &= \{\tau_{1} \in \mathcal{T}_{1} : \E_{n}[m_{j}(X_{i},\tau_{1},\tau_{2})] \leq b_{n} \text{ for $j\in \mathcal{J}(\tau_{2})$} \}\\
&\subset \{\tau_{1} \in \mathcal{T}_{1} : \E_{P}[m_{j}(X_{i},\tau_{1},\tau_{2})] \leq \eta_{n}(\tau_{2}) + b_{n} \text{ for $j\in \mathcal{J}(\tau_{2})$} \}\\
&\subset \mathcal{T}_{1}^{*}(\tau_{2},P,2b_{n})\\
&\subset \mathcal{T}_{1}^{*}(\tau_{2},P)^{\rho(2b_{n})}.
\end{align*}
Furthermore, by Assumption (ref) we have with high probability for large enough $n$:
\begin{align*}
\mathcal{T}_{1}^{*}(\tau_{2},P) &\subset \mathcal{T}_{1}^{*}(\tau_{2},P,b_{n}-\eta_{n}(\tau_{2}))\\
&\subset \mathcal{T}_{1}^{*}(\tau_{2},\P_{n},b_{n}).
\end{align*}
Conclude that with high probability for large enough $n$:
\begin{align*}
d_{H}(\mathcal{T}_{1}^{*}(\tau_{2},\P_{n},b_{n}),\mathcal{T}_{1}^{*}(\tau_{2},P)) \leq \rho(2b_{n}) \to 0,
\end{align*}
where the last line follows from right-continuity of the function $\rho(\,\cdot\,)$. Since $\tau_{2} \in \mathcal{T}_{2}'$ was arbitrary, this completes the proof.
\end{proof}
\subsection{The Additively Separable Case}
In this subsection we show how our method can be applied to a model that satisfies the following assumption.
\begin{assumption} The index function $\varphi$ satisfying Assumption (ref) is additively separable in $U$; i.e. we have $\varphi(X,U,\theta) = \tilde{\varphi}(X,\theta) - U$ for some function $\tilde{\varphi}$.
\end{assumption}
This is a well-studied special case of the linear model considered in the main text. In particular, much of the discussion in this section expands upon the insights of chesher2013semiparametric. We consider two cases: (i) when the structural function $\varphi$ is linear in the parameter vector $\theta$, and (ii) when the structural function is unknown. To begin, let us consider the following simple example.
\begin{example}
Suppose we have a scalar variable $X$ with support $\mathcal{X} = \{x_{1},\ldots,x_{m}\}$ and latent variables $U \in [-1,1]$. Consider the following additively separable threshold crossing model:
\begin{align*}
Y = \mathbbm{1}\{X\theta \geq U \},
\end{align*}
where $\theta$ is a fixed scalar coefficient. The response types in this setting are characterized by the $m\times 1$ vectors:
\begin{align*}
r(u,\theta) :=
\begin{bmatrix}
\mathbbm{1}\{x_{1}\theta \geq u \}\\
\mathbbm{1}\{x_{2}\theta \geq u \}\\
\vdots\\
\mathbbm{1}\{x_{m}\theta \geq u \}\\
\end{bmatrix}.
\end{align*}
However, the set of possible response types in this setting depends on the sign of the fixed coefficient $\theta$. In particular, when $\theta\geq0$ we have the response types $r(u,\theta) \in \{s_{1}, \ldots, s_{m+1}\}$, where:
\begin{align}
s_{1}:=
\begin{bmatrix}
0\\
0\\
\vdots\\
0\\
0
\end{bmatrix}, && s_{2}:=
\begin{bmatrix}
0\\
0\\
\vdots\\
0\\
1
\end{bmatrix}, &&\ldots, &&
s_{m}:=
\begin{bmatrix}
0\\
1\\
\vdots\\
1\\
1
\end{bmatrix},&&
s_{m+1}:=
\begin{bmatrix}
1\\
1\\
\vdots\\
1\\
1
\end{bmatrix}.
\end{align}
No other response types are possible when $\theta>0$, and so all other response types must be assigned zero probability. Alternatively, when $\theta<0$ we have the response types $r(u,\theta) \in \{s_{1}', \ldots, s_{m+1}'\}$, where:
\begin{align}
s_{1}':=
\begin{bmatrix}
0\\
0\\
\vdots\\
0\\
0\\
\end{bmatrix}, && s_{2}':=
\begin{bmatrix}
1\\
0\\
\vdots\\
0\\
0
\end{bmatrix}, &&\ldots,&&
s_{m}':=
\begin{bmatrix}
1\\
1\\
\vdots\\
1\\
0
\end{bmatrix},&&
s_{m+1}':=
\begin{bmatrix}
1\\
1\\
\vdots\\
1\\
1
\end{bmatrix}.
\end{align}
Again, all other response types must be assigned zero probability by the distribution of $U$.
The reason that these particular response types arise when $\theta\geq0$ and $\theta<0$ is due to the ordering of the support of $X$ induced by the value of the scalar product $X\theta$. In particular, if we suppose $x_{1} \leq x_{2} \leq \ldots \leq x_{m}$, then when $\theta \geq0$ we have the ordering $x_{1}\theta \leq x_{2}\theta\leq \ldots \leq x_{m}\theta$. This means, for example, that it is impossible to find a value of $u \in [-1,1]$ so that:
\begin{align*}
r(u,\theta)= \begin{bmatrix}
\mathbbm{1}\{x_{1}\theta \geq u \}\\
\mathbbm{1}\{x_{2}\theta \geq u \}\\
\mathbbm{1}\{x_{3}\theta \geq u \}\\
\vdots\\
\mathbbm{1}\{x_{m}\theta \geq u \}
\end{bmatrix} = \begin{bmatrix}
0\\
1\\
0\\
\vdots\\
0
\end{bmatrix}.
\end{align*}
This shows that when $\theta \geq 0$ certain response types are not possible, and so must be assigned probability zero by the distribution of $U$. An identical intuition holds in the case when $\theta<0$. In the end, the response types that can be assigned positive probability in this example when $\theta \geq 0$ and $\theta <0$ are exactly the ones corresponding to the vectors in (ref) and (ref), respectively. Figure (ref) provides an illustration in the case when $\mathcal{X}=\{x_{1},x_{2},x_{3}\}$.
\begin{figure}[!t]
\caption{A figure corresponding to Example (ref) illustrating the partition of the latent variable space according to response types in the case when the index function is additively separable in $U$ and when $\mathcal{X}= \{x_{1},x_{2},x_{3}\}$ with $x_{1} \leq x_{2} \leq x_{3}$. As indicated in the example, the feasible response types are those that correspond to a particular ordering of the points in $\mathcal{X}$ induced by the scalar product $X \theta$.}
\end{figure}
\end{example}
This example illustrates the key ideas behind the implementation of our approach when the index function is additively separable in $U$, as in Assumption (ref). In particular, given the function $\tilde{\varphi}$ from Assumption (ref), the key is to determine the values of $\theta$ such that the function $\tilde{\varphi}(\,\cdot\,,\theta): \mathcal{X} \to \mathbb{R}$ induces a unique ordering of the points in the support $\mathcal{X}$. With a scalar $X$ variable, and $\tilde{\varphi}(X,\theta) = X \theta$, Example (ref) shows that only two orderings are possible, corresponding to the case when $\theta \geq 0$ and $\theta <0$. After the order is determined, we can immediately determine the set of response types that must be assigned zero probability by the distribution of $U$, and then impose these restrictions as an additional constraint in the bounding problems (ref) and (ref). In particular, let $S_{\varphi}(\theta)$ denote the set of all binary vectors $s \in \{0,1\}^{m}$ corresponding to sets $\mathcal{U}(s,\theta)$ that can be assigned positive probability under Assumption (ref), and impose the constraint:
\begin{align}
\sum_{s \in S_{\varphi}(\theta)^{c}} \pi(y,x_{j},\theta,s) = 0,
\end{align}
for all $y\in \{0,1\}$ and $j=1,\ldots,m$ occurring with positive probability. Then Theorem (ref) can be extended to accommodate Assumption (ref) by simply adding the constraints (ref) to the optimization problems (ref) and (ref).
Similar to the discussion in the main text, determining the sets $\mathcal{U}(s,\theta)$ that can be assigned positive probability under Assumption (ref) poses an interesting computational problem. Although Example (ref) illustrates a case when there are only two orderings, in general many more orderings may be possible, even when $\tilde{\varphi}$ is linear in $\theta$. Clearly at most $m!$ orderings are possible, but when the index function is linear in $\theta$ it is possible to show that the maximum number of possible orderings is much smaller than $m!$. In particular, consider the function $\tilde{\varphi}(X,\theta) = X\theta$ where $X$ is a vector of dimension $d$. Label the support $\mathcal{X}$ as $\{x_{1},x_{2},\ldots,x_{m} \}$, and let $\Delta_{jk} := x_{j} - x_{k}$ for $1\leq j < k \leq m$. The set $H_{jk}:=\{ \theta \in \mathbb{R}^{d} : \Delta_{jk} \theta = 0\}$ defines a hyperplane through the origin that is normal to the line connecting $x_{j}$ and $x_{k}$ in $\mathbb{R}^{d}$. The set of all such hyperplanes partitions $\mathbb{R}^{d}$ into at most $Q(m,d)$ nonempty cones, where $Q(m,d)$ is defined recursively as:
\begin{equation}
Q(m, d) = Q(m-1, d) + (m-1) Q(m-1,d-1),
\end{equation}
with $Q(m,1) = 2$ for all $m \geq 2$ and $Q(2, d) = 2$ for all $d \geq 1$. Furthermore, each these nonempty cones corresponds exactly to the equivalence class of vectors $\theta$ that induce a unique ordering of the points in $\mathcal{X}$. Thus, the value $Q(m,d)$ serves as an upper bound on the number of orderings of the points in $\mathcal{X}$ that are inducible by the function $\tilde{\varphi}(X,\theta) = X\theta$. The recursive formula from (ref) defining the upper bound $Q(m,d)$ has been independently discovered in different contexts by many authors; the earliest such account appears in Bennett1956, although the formula was independently discovered again in cover1967number. The upper bound $Q(m,d)$ is obtained when the collection of hyperplanes of the form $H_{jk}$ are in general position. Note that $Q(m,1)=2$ corresponds exactly to Example (ref), where it was shown that only two orderings could be induced when $\tilde{\varphi}(X,\theta) = X \theta$ for scalar $X$ and $\theta$. Typically, $Q(m, d) < m!$, although some inspection of the formula shows that we always have $Q(m,d)=m!$ when $d \geq m-1$.
If we could select one value of $\theta$ from each of the cones defined by the collection of hyperplanes of the form $H_{jk}$, we could then determine the permitted orderings of the support points $\mathcal{X}$ by simply evaluating $x_{j} \theta$ for $j=1,\ldots,m$. This would then allow us to determine which sets $\mathcal{U}(s,\theta)$ must be assigned zero probability under Assumption (ref). Note that under Assumption (ref) the latent variable $U$ obtains a value on the hyperplane $H_{jk}$ with probability zero. Thus, it suffices to select one value of $\theta$ from the interior of each of the cones defined by the collection of hyperplanes of the form $H_{jk}$. This can be done using the hyperplane arrangement algorithm described in the main text applied to the hyperplanes of the form $H_{jk}$ for $1\leq j < k \leq m$.
Our method is also applicable to cases when $\tilde{\varphi}(X,\theta)$ may be non-linear in $\theta$. When $\tilde{\varphi}$ is not restricted by the researcher, all orderings of the support points in $\mathcal{X}$ are possible. The researcher must first fix an ordering of the support points in $\mathcal{X}$, determine the admissible response types $S_{\varphi}(\theta)$ for the fixed ordering, and run the linear programs in (ref) and (ref) subject to the constraint (ref). The researcher must then repeat the procedure for all possible orderings of the support points in $\mathcal{X}$. On each iteration of this procedure the researcher obtains an interval with endpoints determined by the values of the linear programs in (ref) and (ref). The closed convex hull of the identified set for the counterfactual probability is then given by the interval whose lower endpoint is the smallest value of the linear program in (ref) obtained across all orderings, and whose upper endpoint is the largest value of the linear program in (ref) obtained across all orderings. There are $m!$ possible orderings for $\tilde{\varphi}(X,\theta)$ unless additional assumptions are imposed, so that considering all possible orderings can quickly become computationally demanding.
\begin{comment}
\subsection{Additional Parameters of Interest}
Before turning to our application, it is useful to note that our procedure can also be used to bound objects that are not of the form (ref), but might still be of policy interest. In particular, partitioning the space $\mathcal{U}$ by consideration of response types opens the door to many new parameters.
To illustrate this point, let $f_{\gamma}: \{0,1\}^{m} \to \mathbb{R}$ be an arbitrary function and consider the set:
\begin{align}
S_{\gamma} := \{ s \in \{0,1\}^{m} : f_{\gamma}(s)=0\}.
\end{align}
By careful construction of the function $f_{\gamma}$, this set can be made to represented the set of binary vectors $s \in \{0,1\}^{m}$ that satisfy some criterion of interest to the researcher.\footnote{Note that the set $S_{j}$ introduced above is a special case of $S_{\gamma}$ where $f_{\gamma}(s) = e_{j}^\top s -1$, with $e_{j}$ equal to the basis vector with $1$ as its $j^{th}$ entry and $0$'s elsewhere.} Then a more general class of functionals of the distribution $P_{\theta\mid Y,X,Z}$ can be written as:
\begin{align}
P_{\theta\mid Y,X,Z} \left( r(\theta,\theta)\in S_{\gamma} \mid Y=y, X=x,Z=z\right)= \sum_{s \in S_{\gamma}} P_{\theta\mid Y,X,Z}\left( \mathcal{U}(s,\theta) \mid Y=y, X=x,Z=z\right).
\end{align}
That is, the left hand side of (ref) represents the probability that an individual has response type characterized by the binary vectors in $S_{\gamma}$. The simplest possible parameter of interest that can be written in this form is the probability that an individual is a given response type, which can be bounded by setting $S_{\gamma} = \{s\}$ for some $s \in \{0,1\}^{m}$. However, the formulation in (ref) allows the researcher to bound arbitrary unions of response types. Bounds on these parameters can then be constructed simply by letting (ref) be the objective function in Theorem (ref) without any modification to the constraints. Note that, by an application of Theorem (ref), the parameter in (ref) nests the parameter (ref) as a special case.
Our formulation of the bounding problem can also be easily amended to accommodate more complicated parameters that are able to answer even more detailed policy questions. In particular, let $T$ denote any arbitrary collection of binary vectors $s \in \{0,1\}^{m}$, and consider a parameter of the form:
\begin{align}
P_{\theta\mid Y,X,Z} \left( r(\theta,\theta)\in S_{\gamma} \mid Y=y, X=x,Z=z, r(\theta,\theta)\in T\right).
\end{align}
This parameter is analogous to the parameter in (ref), although it also conditions on the response types belonging to some collection $T$.
A special case of this general parameter is given by:
\begin{align}
P_{\theta\mid Y,X,Z} \left( r(\theta,\theta)\in S_{\gamma} \mid Y=y, X=x,Z=z, Y_{\gamma(k_{1})}=y_{1}, Y_{\gamma(k_{2})}=y_{2}, \ldots, Y_{\gamma(k_{r})}=y_{r}\right),
\end{align}
where $\{k_{1},\ldots,k_{r}\}$ is any subset of the indices $\{1,\ldots,m\}$. Here the potential outcome interpretation provided in Remarks (ref) and (ref) becomes especially helpful, since a parameter like (ref) can be interpreted as the probability an individual belongs to a specific collection of response types conditional on having a certain profile of potential outcomes. In practice a parameter like the one in (ref) can be bounded using our approach after three small modifications:
\begin{enumerate}[label=(\roman*)]
• Redefine the parameter vector $\pi(\theta)$ from (ref) to instead have typical element:
\begin{align}
\pi(y,x,z,\theta,s)=P_{\theta\mid Y,X,Z}\left( \mathcal{U}(s,\theta) \mid Y=y,X=x,Z=z, r(\theta,\theta)\in T\right).
\end{align}
• Replace the objective function in the optimization problems (ref) and (ref) with an objective function of the form:
\begin{align}
P_{\theta\mid Y,X,Z} \left( r(\theta,\theta)\in S_{\gamma} \mid Y=y, X=x,Z=z, r(\theta,\theta)\in T\right)= \sum_{s \in S_{\gamma}} \pi(y,x,z,\theta,s)
\end{align}
• Replace the “adding up” constraint (ref) with the constraints:
\begin{align}
\sum_{s \in T} \pi(y,x_{j},z_{j},\theta,s) =1, && \sum_{s \in T^{c}} \pi(y,x_{j},z_{j},\theta,s) =0,
\end{align}
for all $y \in \{0,1\}$ and $j=1,\ldots,m$.
\end{enumerate}
After these three steps, running the analogous optimization problems from Theorem (ref) will produce the closed convex hull for a parameter of the form (ref).\todo{Okay, but it seems as though all constraints introduced in this section will also need to be modified.}
\todo[inline]{Additional parameters such as arbitrary regions in $\mathcal{U}$ space. Can this be accomplished by setting $S_{\gamma} = \{ s : A\cap \mathcal{U}(s,\theta)\neq \emptyset\}$?}
\end{comment}
\section{Comparison to Artstein's Inequalities}
Here we briefly discuss the method proposed by chesher2013instrumental and chesher2014instrumental.\footnote{The general version of the approach can be found in chesher2017generalized.} Our objective is to provide an informal comparison, and to illustrate the connections between the two approaches. To ease the comparison, we focus on the identified set for (conditional) latent variable distributions rather than counterfactual probabilities. Suppose Assumption (ref) holds, and consider the correspondence:
\begin{align}
\overline{\mathcal{U}}(y,x,\theta)&:=\text{cl}\left\{ u \in \mathcal{U} : y = \mathbbm{1}\{\varphi(x,u,\theta) \geq 0 \} \right\}.
\end{align}
The set $\overline{\mathcal{U}}(Y,X,\theta)$ is the closure of the set $\mathcal{U}(Y,X,\theta)$ from (ref). Under some conditions, when the distribution of $U$ is absolutely continuous (which is assumed, for example, in both chesher2013instrumental and chesher2014instrumental), these two random sets are equal almost surely.\footnote{In chesher2014instrumental, absolute continuity combined with a linear index function ensures this statement is true. chesher2013instrumental consider a more general class of latent index functions than chesher2014instrumental, and so also impose strict monotonicity in latent variables of the latent index function in order to ensure their analog of the set $\left\{u \in \mathcal{U} : \varphi(x,u,\theta) = 0 \right\}$ is of Lebesgue measure zero for each $(x,\theta)$.} Also define the correspondence:
\begin{align}
H(u,\theta)&:=\left\{(y,x) \in \mathcal{Y} \times \mathcal{X} : y = \mathbbm{1}\{\varphi(x,u,\theta) \geq 0 \} \right\}.
\end{align}
A result due to artstein1983distributions (also see norberg1992existence and molchanov2017theory Corollary 1.4.11), characterizes the set of selections of a random closed set.
\begin{theorem}
Suppose that Assumption (ref) holds. Then for any $\theta \in \Theta$, the random vector $U$ can be realized as a selection of the random closed set $\overline{\mathcal{U}}(Y,X,\theta)$ if and only if:
\begin{align}
P_{U}(U \in K) \leq P_{Y,X}\left(\overline{\mathcal{U}}(Y,X,\theta)\cap K \neq \emptyset\right),
\end{align}
for all compact sets $K \subset \mathcal{U}$. Furthermore, for any $\theta \in \Theta$, the random vector $(Y,X)$ can be realized as a selection of the random closed set $H(U,\theta)$ if and only if:
\begin{align}
P_{Y,X}((Y,X) \in C) \leq P_{U}( H(U,\theta)\cap C \neq \emptyset),
\end{align}
for all compact sets $C \subset \mathcal{Y}\times \mathcal{X}$.
\end{theorem}
\begin{remark}
The random vector $U$ can be realized as a selection of the random closed set $\overline{\mathcal{U}}(Y,X,\theta)$ if and only if there exists a probability space and random elements $\tilde{U}$ and $\tilde{\mathcal{U}}(Y,X,\theta)$ with identical distributions to $U$ and $\overline{\mathcal{U}}(Y,X,\theta)$ such that $\tilde{U} \in \tilde{\mathcal{U}}(Y,X,\theta)$ a.s. Since $\mathcal{U} \subset \mathbb{R}^{d_{\theta}}$ is locally compact and Hausdorff, it is equivalent that (ref) hold for all open sets $G \subset \mathcal{U}$. Note that since $\mathcal{Y}\times \mathcal{X}$ is finite, all subsets are trivially compact with respect to the discrete topology.
\end{remark}
The first part of this result is very similar to Theorem 1 in chesher2013instrumental and Theorem 3.1 in chesher2014instrumental, and the second part of this result is a direct corollary of Theorem 1 in chesher2017generalized. Either (ref) and (ref) can be used to construct the identified set of unconditional latent variable distributions, say $\mathcal{P}_{U}^{*}$; in practice, this is accomplished by first fixing a value of $\theta \in \Theta$, collecting all distributions $P_{U}$ satisfying either (ref) or (ref), and then taking a union (over all $\theta \in \Theta$) of the resulting collections of distributions. A similar result to Theorem (ref) can be stated after conditioning on $(Y,X)$.
The main difficultly with using the characterization of the identified set based on Artstein's inequalities is the number of constraints that must be imposed. The resulting computational bottleneck is thus associated with a lack of short-term computer memory needed to store all of Artstein's inequalities. Most efforts to reduce the computational burden of the approach based on Artstein's inequalities are directed towards reducing the number of constraints implied by Theorem (ref); see the discussions in galichon2011set, chesher2017generalized, russell2021sharp.
At first glance it appears that (ref) leads to a characterization of the identified set of unconditional latent variable distributions that is intractable, given the number of possible compact subsets of $\mathcal{U}$. However, following the discussion in both chesher2013instrumental and chesher2014instrumental, most of the inequalities of the form (ref) are redundant. For instance, when $\varphi$ is linear in $U$ and $\theta$, the set $\overline{\mathcal{U}}(y,x,\theta)$ is a closed halfspace through the origin. In this case, chesher2014instrumental show it suffices to check the inequalities in (ref) for all sets $K$ that can be written as the intersection of halfspaces of the form $\overline{\mathcal{U}}(y,x,\theta)$. In the case with no exogenous variables and a scalar endogenous variable $X$ with $m$ points of support, chesher2014instrumental demonstrate that there are at most $2m(2m+1)/2$ nonredundant inequalities implied by (ref).
On the other hand, since $r:=|\mathcal{Y}\times \mathcal{X}|$ is finite, (ref) gives $2^r-1$ inequalities. Even with small values of $m$ the resulting number of inequalities can also be prohibitively large. We now show that, at the cost of some pre-processing, our approach leads to a simplification of the set of constraints in (ref): the number of constraints in our approach is proportional to $r$ rather than $2^r$.
Consider imposing (ref) conditional on $(Y,X)$, and assume for simplicity that each $(y,x) \in \mathcal{Y}\times\mathcal{X}$ is assigned positive probability. For any $\theta \in \Theta$, the random vector $U$ can be realized as a selection from $\overline{\mathcal{U}}(y,x,\theta)$ if and only if:
\begin{align*}
\mathbbm{1}\{(y,x) \in C\} \leq P_{U \mid Y,X}( H(U,\theta)\cap C \neq \emptyset \mid Y=y,X=x),
\end{align*}
for all compact $C \subset \mathcal{Y}\times\mathcal{X}$. For a fixed value of $(y,x)$, consider the set of all compact $C \subset \mathcal{Y}\times\mathcal{X}$ containing $(y,x)$. For all such $C$ we must have:
\begin{align*}
P_{U \mid Y,X}( H(U,\theta)\cap C \neq \emptyset \mid Y=y,X=x)=1.
\end{align*}
However, this can hold for all compact sets $C \subset \mathcal{Y}\times\mathcal{X}$ containing $(y,x)$ if and only if it holds for the singleton set $\{(y,x)\}$. Conclude that:
\begin{align*}
P_{U\mid Y,X}( H(U,\theta)\cap \{(y,x)\} \neq \emptyset \mid Y=y,X=x)=1.
\end{align*}
Some basic manipulation shows this holds if and only if:
\begin{align*}
P_{U \mid Y,X}(U \in \overline{\mathcal{U}}(Y,X,\theta) \mid Y=y,X=x)=1.
\end{align*}
This derivation can be used to prove the following Lemma.
\begin{lemma}
Suppose that Assumption (ref) holds and that all $(y,x) \in \mathcal{Y}\times\mathcal{X}$ are assigned positive probability. Then Artstein's inequalities:
\begin{align}
P_{Y,X}((Y,X) \in C) \leq P_{U}( H(U,\theta)\cap C \neq \emptyset),
\end{align}
hold for all compact sets $C \subset \mathcal{Y}\times \mathcal{X}$ if and only if:
\begin{align}
P_{U \mid Y,X}(U \in \overline{\mathcal{U}}(Y,X,\theta) \mid Y=y,X=x)=1,
\end{align}
for all $(y,x)$.
\end{lemma}
Notice that (ref) is nearly identical to the condition in Definition (ref) characterizing $\mathcal{P}_{U\mid Y,X}^{*}$, connecting the approach based on Artstein's inequalities to the approach considered in this paper. Setting $r=|\mathcal{Y} \times \mathcal{X}|$, we conclude that the $2^{r}-1$ inequalities implied by (ref) are equivalent to the (at most) $r$ equalities implied by (ref). For the sake of comparison, recall that in the case with no exogenous variables and a scalar endogenous variable $X$ with $m$ points of support, chesher2014instrumental demonstrate that there are at most $2m(2m+1)/2$ nonredundant inequalities implied by (ref) when $\varphi$ is linear in $(U,\theta)$. In contrast, Lemma (ref) implies that with no exogenous variables and a scalar endogenous variable $X$ with $m$ points of support, there are at most $2m$ constraints implied by (ref), regardless of the functional form specified for $\varphi$.
We conclude by noting that our characterization may not always produce less constraints than the approach based on Artstein's inequalities. In particular, the characterization based on Artstein's inequalities easily handles the incorporation of instruments without increasing the number of inequality constraints (c.f. beresteanu2012partial Proposition 2.5), where our approach may require significantly more constraints. However, in many cases we believe our approach can offer substantial simplifications.
\section{Additional Figures}
This section contains two figures illustrating the intervals computed using the linear programs of the form (ref) and (ref) for each profiling point for certain specifications in the application section. Figure (ref) presents the intervals computed for each representative point of $\theta \in \mathbb{R}^2$ when bounding $\mu_{ate}$ for Model (ref) under various assumptions. Figure (ref) shows the analogous intervals for model (ref).
\begin{sidewaysfigure}
\caption{This figure shows the intervals computed using the linear programs of the form (ref) and (ref) for each representative point of $\theta \in \mathbb{R}^2$ when bounding $\mu_{ate}$ for Model (ref) under various assumptions. The active assumptions are given at the top of each illustration. The axes labelled “Profile Points” correspond to various representative points. }
\end{sidewaysfigure}
\begin{sidewaysfigure}
\caption{This figure shows the intervals computed using the linear programs of the form (ref) and (ref) for each representative point of $\theta \in \mathbb{R}^3$ when bounding $\mu_{ate}$ for Model (ref) under various assumptions. The active assumptions are given at the top of each illustration. The axes labelled “Profile Points” correspond to various representative points.}
\end{sidewaysfigure}