Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
129,297 characters · 21 sections · 86 citation commands
Identification in Multiple Treatment Models under Discrete Variation
\thispagestyle{empty}
KEYWORDS: Multiple treatments, discrete instrument, marginal treatment effects, partial identification, Head Start.
JEL classification codes: C14, C31, C36, C61, I21.
\setcounter{page}{1}
In the analysis of treatment effects using instrumental variables, a threshold crossing selection model with a single dimension of unobserved heterogeneity underlies the predominant framework in the literature imbens/angrist:94, heckman/vytlacil:05, vytlacil:02. Designed for the canonical binary treatment setup, it presumes that individuals choose between only two treatments. In many scenarios, however, individuals choose between multiple treatments. Such cases naturally give rise to selection models with multiple dimensions of unobserved heterogeneity heckman/etal:06,heckman/etal:08,lee/salanie:18.
In this paper, we develop a method to learn about treatment effects in multiple treatment setups under a general class of selection models with multidimensional unobserved heterogeneity. Specifically, we consider the class of multidimensional threshold crossing models from lee/salanie:18. As highlighted in Section (ref), this class nests the standard binary treatment model and also covers many empirically relevant settings with multidimensional unobserved heterogeneity, including multinomial, sequential, multi-stage, and double-hurdle choice models.
In this class of models, our objective is to show how to learn about various treatment effects that can be written as weighted averages of the marginal treatment response functions (MTRs) heckman/vytlacil:99,heckman/vytlacil:05,mogstad/etal:18, which capture the mean potential outcomes conditional on the unobserved heterogeneity governing the selection model. As in the standard binary treatment case, MTRs provide a unified framework to express a number of parameters of interest such as the average treatment effect as well as other averages that target policy relevant subgroups.
To analogously learn about such parameters, lee/salanie:18 assume availability of data with continuous variation in the instrument. They show how a sufficient amount of such variation allows nonparametric identification of the selection model and the MTRs over their entire support, and in turn all the parameters of interest---see also heckman/etal:06, heckman/etal:08, mountjoy:22 and tsuda:23 for related arguments in specific models. However, in many empirical scenarios, instruments are only discrete-valued. In such cases, the parameters are usually partially identified as there generally exist multiple admissible values of the MTRs and primitives characterizing the selection model that are consistent with the data.
We show how to compute the identified set for our parameters of interest in these cases. In particular, we build on mogstad/etal:18, who study this identification problem in the special case of a binary treatment, where a single dimension of unobserved heterogeneity underlies selection. In this case, the primitives characterizing the selection model are point identified and the problem is linear in the remaining primitives, namely the MTRs. Imposing a linear parametrization on the MTRs, mogstad/etal:18 show how these features allow computing sharp bounds by solving two linear programs. They also show that with carefully chosen basis functions over a unidimensional partition, such parametrizations can accommodate fully nonparametric MTRs satisfying only shape restrictions. However, two inherent challenges arise when multidimensional unobserved heterogeneity underlies selection: the primitives characterizing the selection model are not point-identified, and the MTRs are no longer unidimensional.
We propose a two-step solution that accounts for these complications. In an inner step, for a given selection primitive, we first exploit the fact that the problem remains linear in the MTRs. Similarly linearly parametrizing the MTRs, we show this implies that we can continue to use linear programs to compute sharp bounds given the selection primitives. Moreover, we introduce novel arguments to show that such a result can also continue to accommodate entirely nonparametric MTRs. Specifically, as expanded in Section (ref), we show how to leverage the rectangular structure underlying our general selection model and construct a new partition that generalizes the unidimensional one from mogstad/etal:18 to the multidimensional case.
In an outer step, we then introduce alternative arguments to account for the fact that the selection model is not point-identified. Specifically, this step takes the unions of the first step bounds across admissible values of the selection primitives. To do so feasibly, we assume that these primitives can be point identified up to a finite dimensional parameter. We demonstrate in several empirically relevant models how this assumption can be satisfied under familiar parametric assumptions on the distribution of the unobserved heterogeneity. In summary, our method translates to solving two linear programs to compute the identified set at the point identified value of the selection model across various values of the finite dimensional parameter, and then taking the union of these sets. We show how this two-step characterization also translates into tractable two-step strategies to estimate the identified set and construct confidence intervals.
Our method offers several advantages over recent approaches to point identification in settings with multiple treatments and discrete instruments. A common method is to impose monotonicity assumptions that restrict selection patterns such that the researcher can nonparametrically point-identify treatment effects for certain response types angrist/imbens:95,bhuller/sigstad:24,goff:25,heckman/pinto:18,kline/walters:16, kirkeboen/etal:16,lee/salanie:23,navjeevan/etaL:23,pinto:19. As we highlight in Section (ref), many such monotonicity assumptions can be captured by selection models in our general class. In such cases, although our method parametrizes the selection model, sufficiently flexible specifications can exactly replicate the nonparametric estimates these approaches produce.\footnote{kline/walters:19 and mogstad/etal:18 similarly note in the binary treatment case that even if the selection model and MTRs are parametrized, one can produce the nonparametric point identified local average treatment effect.} But our method also importantly extends beyond these settings: it allows researchers to learn about effects that are only partially identified and to accommodate richer selection patterns that would otherwise preclude point identification.
A smaller group of papers similarly considers selection models to learn about additional treatment effects, but limits attention to specific models and parameterizations that are point identified by the data hull:20, kline/walters:16,pinto:19. Our method contributes to these papers by nesting such point identification as a special case while allowing more robust parameterizations of the selection model and MTRs that generally imply bounds on the treatment effects of interest. Moreover, our analysis applies to a general class of selection models. Our method therefore provides a blueprint for robustly learning about treatment effects across a wide range of empirical settings.
Our method also has advantages relative to a complementary literature on partial identification in multiple treatment models. Papers in this literature either make no assumptions on selection manski:97,manski/pepper:00,manski/pepper:09 or impose only nonparametric restrictions bai/etal:2025,kamat:21,lee/salanie:23. However, because these approaches do not model the MTRs, they cannot accommodate parameters such as policy-relevant treatment effects or assumptions such as parametric restrictions and separability that are common in empirical work. Our method allows for such parameters and assumptions as well as flexible variants that enable researchers to assess the robustness of their conclusions.
To concretely demonstrate these various benefits, we use our method to revisit kline/walters:16' (kline/walters:16) empirical analysis of the Head Start preschool program using the Head Start Impact Study. Due to oversubscription, Head Start offers were randomized to applicants in the study, which enabled an experimental evaluation of Head Start by comparing outcomes of children with offers to those without. However, what these experimental effects can identify on the effects of Head Start and their policy relevance are clouded by the fact that many children enroll in competing non-Head Start preschools that provide similar services to Head Start and are also publicly subsidized. To account for this complication, kline/walters:16 formulate a multinomial selection model of preschool choice and slot provision. But, in estimating this model, they impose two simplifying assumptions to ensure that both the selection model and MTRs are point-identified: competing preschools do not ration slots, and the MTRs are linear and separable.
We illustrate how our method can be used to gradually weaken these assumptions in a step-by-step manner and analyze the sensitivity of their main empirical conclusions to them. Our bounds reveal that their first conclusion that the positive experimental effects are driven by a positive effect from families who enroll in no preschools absent a Head Start offer is robust to these assumptions. In particular, this conclusion can hold under nonparametric and nonseparable MTRs, and when the selection model allows rationing of competing preschools. In contrast, for their second conclusion that marginally expanding Head Start has a positive, statistically significant effect, we find their assumptions are generally necessary unless one is willing to take the relation between test score and earnings to be towards the higher end of values reported in the literature.
The remainder of the paper is organized as follows. Section (ref) introduces the class of models and treatment effect parameters we analyze. Section (ref) presents our arguments to compute the identified set. Section (ref) discusses how these arguments can be used to estimate the identified set and construct confidence intervals. Section (ref) presents our empirical application. Section (ref) concludes. Proofs of all results are presented in Appendix (ref).
For each individual, we observe an outcome of interest $Y$, their received treatment $D$, their assigned instrument value $Z$, and baseline covariates $X$. We take the list of possible treatments to be given by a discrete set $\mathcal{D}$, and model the support of the instrument and covariates denoted by $\mathcal{Z}$ and $\mathcal{X}$ to be discrete. We assume the outcome and treatment to be generated by the usual potential outcomes structure. In particular, denoting by $D(z)$ the potential treatment had the individual's assigned instrument value been $z \in \mathcal{Z}$, the observed treatment is given by
and, similarly, denoting by $Y(d)$ the potential outcome had the individual's received treatment been $d \in \mathcal{D}$, the observed outcome is given by
Moreover, as usual, we assume that the instrument is statistically independent of the potential variables conditional on the covariates as follows.
\begin{assumption2}{E}{(Exogeneity)} $(\{Y(d) : d \in \mathcal{D}\}, \{D(z) : z \in \mathcal{Z}\}) \perp Z | X~.$ \end{assumption2}
We assume the potential treatments to be related across instrument values through a selection model. Let $U \equiv (U_1, \ldots, U_J)$ denote a vector of individual unobservables governing selection. Following lee/salanie:18, we take potential treatment to be determined by the region in which $U$ falls, where these regions satisfy certain properties in addition to being disjoint to logically ensure that they cannot imply several different treatments.
\begin{assumption2}{SM}{(Selection Model)} Conditional on $X = x \in \mathcal{X}$, let $U$ be continuously distributed with support contained in $[0,1]^J$ with uniform marginal distributions, and let
for $z \in \mathcal{Z}$, where $\{\mathcal{U}^{\text{sm}}_{d,z|x} : d \in \mathcal{D}\}$ denotes a collection of disjoint subsets of $[0,1]^J$ that are members of the $\sigma$-field generated by the sets $\{\{u \in [0,1]^J : u_j \leq g_{j,z|x}\} : j =1,\ldots,J\}$ for some unknown threshold values $\{g_{j,z|x} : j = 1, \ldots, J\}$. \end{assumption2}
As the generated $\sigma$-field is obtained by taking unions, intersections and complements of the sets in $\{\{u \in [0,1]^J : u_j \leq g_{j,z|x}\} : j =1,\ldots,J\}$, Assumption (ref) requires that each of the regions $\mathcal{U}^{\text{sm}}_{d,z|x}$ determining potential treatment can be written as
for some finite $L_{d}$ and $\underline{u}_{j,l,d,z|x},\bar{u}_{j,l,d,z|x} \in \{0,g_{j,z|x},1\}$, i.e. as a finite union of rectangular sets. As usual, the restriction in Assumption (ref) that $U$ has support contained in $[0,1]^J$ with uniform marginal distributions can be viewed as a normalization. To see this, let $\tilde U = (\tilde U_1,\dots,\tilde U_J)$ denote a vector of latent unobservables with continuous marginal distribution functions $\tilde F_{j\mid x}$ given $X=x$. For any thresholds $\underline{\tilde u}_{j,l,d,z\mid x}$ and $\bar{\tilde u}_{j,l,d,z\mid x}$, \[ 1\{\underline{\tilde u}_{j,l,d,z\mid x} \le \tilde U_j \le \bar{\tilde u}_{j,l,d,z\mid x}\} = 1\{\tilde F_{j\mid x}(\underline{\tilde u}_{j,l,d,z\mid x}) \le \tilde F_{j\mid x}(\tilde U_j) \le \tilde F_{j\mid x}(\bar{\tilde u}_{j,l,d,z\mid x})\}. \] Defining $U_j = \tilde F_{j\mid x}(\tilde U_j)$, $\underline{u}_{j,l,d,z\mid x} = \tilde F_{j\mid x}(\underline{\tilde u}_{j,l,d,z\mid x})$, $ \bar u_{j,l,d,z\mid x} = \tilde F_{j\mid x}(\bar{\tilde u}_{j,l,d,z\mid x})$, we obtain a normalized vector $U=(U_1,\dots,U_J)$ with support $[0,1]^J$ and uniform marginal distributions conditional on $X=x$, and selection regions of the form in (ref) with endpoints $\{\underline{u}_{j,l,d,z\mid x},\bar{u}_{j,l,d,z\mid x}\}$.
As highlighted in lee/salanie:18, Assumption (ref) reduces to the standard threshold crossing model heckman/vytlacil:05, imbens/angrist:94 when we have a binary treatment and a single dimension of unobserved heterogeneity given by
for $z \in \mathcal{Z}$ conditional on $X = x \in \mathcal{X}$. But, in cases where $J > 1$, it provides a unified framework to capture a number of additional common, empirically relevant models of selection with multiple as well as binary treatments. Below, we briefly provide examples of several such models. Our first three examples consider multiple treatment models, while our fourth example considers a binary treatment model. For simplicity, when possible, we take the number of treatments in the multiple treatment examples to be solely equal to three.
In our model, we take the primitives to consist of three objects. First, the marginal treatment response (MTR) functions heckman/vytlacil:99,heckman/vytlacil:05,mogstad/etal:18, defined by $m_{d|x}(u) = E[Y(d)\mid U {=} u, X {=} x]$, which capture the mean treatment response for $d \in \mathcal{D}$ conditional on the unobserved heterogeneity $U {=} u$ and covariates $X {=} x$. Second, the distribution of these unobservables conditional on $X {=} x$, denoted by $F_x$. Third, all threshold values determining selection in Assumption (ref), denoted by $g = \{g_{j,z|x} : j = 1,\ldots,J,~z \in \mathcal{Z},~x \in \mathcal{X}\}$. Let $\beta \equiv (m, h)$ summarize these primitives, where $m$ collects the MTRs and $h = (F, g)$ comprises the selection primitives.
Our parameters of interest are functionals of these primitives. In particular, our analysis applies to any parameter that can be written as a weighted sum of integrals of the MTRs over specific regions of the unobserved heterogeneity. Formally, we consider parameters of the form
where $\mathcal{L}$ is a known index set that we sum across in addition to the covariate value $\mathcal{X}$ and treatment values $\mathcal{D}$, $w_{d,l|x}(h)$ determines the weight each integral receives, and $\mathcal{U}^{\theta}_{d,l|x}(g)$ is the integration region specifying which individuals are included in each weighted average. We assume that the weights $w_{d,l|x}(h)$ are known or identified functions of the selection primitives $h$, and the integration regions $\mathcal{U}^{\theta}_{d,l|x}(g)$ can be expressed as a finite union of sets of the form
where the endpoints $\underline{u}_{j,d,l|x}(g)$ and $\bar{u}_{j,d,l|x}(g)$ are known or identified functions of the thresholds $g$. This rectangular structure mirrors that imposed on the selection regions in (ref), which plays a key role in our nonparametric analysis in Section (ref).
As illustrated in Table (ref), a number of commonly studied treatment effect parameters can be written in terms of (ref). These include the average treatment effect (ATE) between two treatments $d',d'' \in \mathcal{D}$ as well as their analogs that condition on $D=d$ to give the average treatment effect on the treated (ATT) or on a response type $D(z) = d_z \in \mathcal{D}$ for $z \in \mathcal{Z}' \subseteq \mathcal{Z}$ to give local average treatment effects (LATEs). Moreover, they also include the class of policy relevant treatment effects heckman/vytlacil:99,heckman/vytlacil:05 that evaluate the effects of altering the thresholds determining selection. To formally define these latter effects, note that the sets determining treatment in (ref) can be more explicitly stated as $\mathcal{U}^{\text{sm}}_{d,z|x} \equiv \mathcal{U}^{\text{sm}}_{d,z|x}(g)$, i.e. as functions of the unknown thresholds. A policy $\delta$ can then be equivalently captured by the thresholds $g + \delta \equiv \{g_{j,z|x} + \delta_{j,z|x} : j = 1, \ldots, J,~z \in \mathcal{Z},~x \in \mathcal{X}\}$, i.e. changing the original thresholds by some known values $\delta \equiv \{\delta_{j,z|x} : j = 1, \ldots, J,~z \in \mathcal{Z},~x \in \mathcal{X}\}$. Let $D^{\delta} = \sum_{d \in \mathcal{D}}d 1\{ U \in \mathcal{U}^{\text{sm}}_{d,Z|X}(g+\delta)\}$ and $Y^{\delta} = Y(D^{\delta})$ denote the individual's treatment and outcome under this policy, respectively. The policy relevant treatment effect (PRTE) between two policies $\delta'$ and $\delta''$ can then be defined by $\text{PRTE}^{\delta',\delta''} = E[Y^{\delta'} - Y^{\delta''}|D^{\delta'} \neq D^{\delta''}]$, i.e. the effect of the counterfactual change in outcomes between the two altered values of thresholds for those who are affected by this change.
The objective of our analysis is to learn about a pre-specified parameter of interest $\theta(\beta)$ given by (ref). As the function $\theta$ is known, what we can learn about the parameter depends on what we know about the primitives $\beta$. To this end, denoting by $\mathbf{M}^{\dagger}$ and $\mathbf{H}^{\dagger}$ the space of all possible $m$ and $h$, let $\mathbf{B} \subseteq \mathbf{M}^{\dagger} \times \mathbf{H}^{\dagger}$ be the admissible space that $\beta$ is restricted to lie in, which is determined by the various assumptions we may impose on it. The data also provide information on $\beta$. The observed decision and outcome through (ref) and (ref) along with the selection model in (ref) and Assumption (ref) provide the following moments
for $d \in \mathcal{D}$, $z \in \mathcal{Z}$ and $x \in \mathcal{X}$. Denoting by $\mathbf{B}^* = \{\beta \in \mathbf{B} : \beta \text{ satisfies } \eqref{eq:D_moments}-\eqref{eq:Y_moments}\}$ the admissible set of primitives satisfying the assumptions and the moments, what we can learn about the parameter of interest can then be formally captured by the identified set defined by
i.e. the image of the admissible set $\mathbf{B}^*$ under the function $\theta$.
We start by showing how to compute the identified set when the moments of the data in the left hand sides of (ref) and (ref) are assumed to be known without uncertainty. The subsequent section shows how this analysis translates into a strategy to estimate the identified set and construct confidence intervals.
From the definition in (ref), observe that computing the identified set requires searching through the space $\mathbf{B}^*$ and taking the image under the function $\theta$. The challenge is how to perform this search in a tractable manner.
To do so, we build on the approach of mogstad/etal:18, who study this identification problem in the special case of a binary treatment when selection is given by (ref). In this case, because a single dimension of unobserved heterogeneity governs selection and is normalized to be uniformly distributed on $[0,1]$, the distribution $F$ is known. Given $F$ is known, $h$ is then point-identified as the threshold function $g$ is directly point-identified by the treatment shares via (ref). This allows simplifying the search problem as it remains to search only across the admissible values of the MTRs. Specifically, for a given $h \in \mathbf{H}^\dagger$, let $\mathbf{M}(h) \equiv \{m \in \mathbf{M}^\dagger : (m,h) \in \mathbf{B}\}$ denote the set of admissible MTRs and let $\mathbf{M}^*(h) = \{m \in \mathbf{M}(h) : (m,h) \text{ satisfies } \eqref{eq:Y_moments} \}$ denote those additionally consistent with the outcome moments. The identified set then simplifies to $\theta(\mathbf{M}^*(h^*),h^*) \equiv \{\theta(m,h^*) : m \in \mathbf{M}^*(h^*) \}$, where $h^*$ denotes the point identified value of $h$. Assuming $\mathbf{M}(h)$ takes $m$ to be linearly parameterized, mogstad/etal:18 exploit the linearity of $\theta$ and $\mathbf{M}^*(h)$ in terms of this parametrization to show that $\theta(\mathbf{M}^*(h^*),h^*) = [\theta_L(h^*),\theta_U(h^*)]$, where
and where these optimization problems correspond to linear programs.
However, in the case where multidimensional unobserved heterogeneity governs selection, an inherent challenge is the fact that $h$ is not generally point identified. This is because, while the marginals continue to be normalized to uniform distributions, the dependence between them is not necessarily known, implying that both $F$ and $g$ may not be known or point identified by the data. We propose a two-step solution to account for this complication. Specifically, let $\mathbf{H}$ denote the set of admissible values of $h$ and $\mathbf{H}^* = \{h \in \mathbf{H} : h \text{ satisfies } \eqref{eq:D_moments}\}$ denote those additionally consistent with the selection moments. The identified set in (ref) can then be stated as
i.e. a union of the identified sets given $h$ across the admissible values of $h$ consistent with the selection moments. This suggests a two-step strategy, where in an inner step, as in the unidimensional case, we can continue to exploit the linearity to compute $\theta(\mathbf{M}^*(h),h)$ given $h \in \mathbf{H}^*$, and then in an outer step, we can aggregate them across $\mathbf{H}^*$ to obtain the final identified set.
In what follows, we show how to operationalize this strategy. In Sections (ref) and (ref), we first show how to generalize the arguments from the unidimensional case to that with multidimensional MTRs to tractably compute $\theta(\mathbf{M}^*(h),h)$ as in (ref) for each $h \in \mathbf{H}$ when $\mathbf{M}(h)$ takes $m$ to be parameterized or remain entirely nonparametric. In Section (ref), we then show how to parametrize $\mathbf{H}^*$ by a finite-dimensional parameter so that we can feasibly take the union across $\mathbf{H}^*$.
As in mogstad/etal:18, we assume $\mathbf{M}(h)$ to linearly parametrize $m$ along with a parameter space determined by a system of linear inequalities as follows.
\begin{assumption2}{PM}{(Parametrizing MTRs)} \sloppy $\mathbf{M}(h) = \bigl\{ m \in \mathbf{M}^{\dagger} : m_{d|x}(u) = \sum_{k=0}^{K} \gamma_{k,d|x} b_{k}(h)(u) \text{ for } d \in \mathcal{D},~x \in \mathcal{X}, \text{ and } \gamma \equiv (\gamma_{k,d|x} : 0 \leq k \leq K,~d \in \mathcal{D},~x \in \mathcal{X}) \in \mathbf{\Gamma}(h) \bigr\}$, where $\{b_{k}(h) : 0 \leq k \leq K \}$ are known functions and $\mathbf{\Gamma}(h) \subseteq \mathbf{R}^{\text{dim}(\gamma)}$ is characterized by a system of linear inequalities. \end{assumption2}
To see how this allows computing $\theta(\mathbf{M}^*(h),h)$ using linear programs as in (ref) for each $h \in \mathbf{H}$, it is useful to rewrite it in terms of the parametrizing variable $\gamma$. As $\theta$ is linear in $m$ and $m$ is linear in $\gamma$, we can substitute in the relation between $m$ and $\gamma$ from Assumption (ref) into $\theta$ to obtain a function $\theta^{\Gamma}$ linear in $\gamma$ such that $\theta(m,h) = \theta^{\Gamma}(\gamma,h)$. Similarly, rewriting the outcome moments in (ref) in terms of $\gamma$
for $d \in \mathcal{D}$, $z \in \mathcal{Z}$, $x \in \mathcal{X}$, we can write $\mathbf{M}^*(h)$ in terms of $\gamma$ by $\mathbf{\Gamma}^*(h) = \left\{ \gamma \in \mathbf{\Gamma}(h) : \gamma \text{ satisfies } \eqref{eq:Y_moments_alpha} \right\}$. The identified set in terms of $\gamma$ then corresponds to $\theta^{\Gamma}(\mathbf{\Gamma}^*(h),h) \equiv \{\theta^{\Gamma}(\gamma,h) : \gamma \in \mathbf{\Gamma}^*(h)\}$. In the following proposition, we show that the linearity implies that $\theta^{\Gamma}(\mathbf{\Gamma}^*(h),h)$ for $h \in \mathbf{H}$ equals an interval with endpoints given by two optimization problems, which correspond to linear programs as $\theta^{\Gamma}$ is a linear function in $\gamma$ and $\mathbf{\Gamma}^*(h)$ is given by a system of linear inequalities.
Assumption (ref) requires the restrictions on $\gamma$ to correspond to a system of linear inequalities. As highlighted in Table (ref), a number of common shape restrictions from the literature correspond to linear restrictions on $m$ and in turn to a system of linear inequalities on $\gamma$ when considered over a finite grid of points of $u$ in $[0,1]^J$. Specifically, it allows assumptions such as boundedness typically required to ensure the bounds on mean effects are finite manski:08, monotonicity in outcomes between two treatments as in the monotone treatment response assumption from manski:97, or the difference in mean outcomes between two treatments to be bounded as in the bounded variation assumption from manski/pepper:18. It also allows imposing monotonicity across different values of the unobserved heterogeneity in a given dimension to capture how the outcome may vary for those more or less likely to select into treatment as in the monotone treatment selection assumption from manski/pepper:00, or separability between the unobserved heterogeneity and the covariates as in brinch/etal:17.
Although Proposition (ref) parameterizes the MTRs via Assumption (ref), it allows the researcher to compute the identified set under fully nonparametric MTRs in special cases. In particular, let $\mathbf{M}_{\text{np}}(h)$ denote a set of nonparametric MTRs. For certain such sets formalized below, we can show $\theta(\mathbf{M}^*_{\text{np}}(h),h) = \theta(\mathbf{M}^*(h),h)$, where $\mathbf{M}(h)$ denotes the subset of $\mathbf{M}_{\text{np}}(h)$ additionally satisfying Assumption (ref) with $\{b_{k}(h) : 0 \leq k \leq K\}$ given by
for a specific choice of partition $\mathbb{U}(h) \equiv \{\mathcal{U}_{0}(h), \mathcal{U}_{1}(h), \ldots, \mathcal{U}_{K}(h)\}$ of $[0,1]^J$, and with
which can be equivalently written as a system of linear inequalities. This mirrors mogstad/etal:18, which shows such a result in the case with a single dimension of unobserved heterogeneity. In this case, the rectangular sets underlying the selection model and parameters in (ref) and (ref) simplify to just intervals. Exploiting this interval structure, they show that a partition consisting of intervals constructed using the end points of those in (ref) and (ref) can generate the nonparametric bounds.
We construct a new partition that shows how to exploit the general rectangular structure of (ref) and (ref), and generalize the unidimensional partition to the case with multidimensional unobserved heterogeneity. Let $\mathbb{U}^{*}(h) = \{\mathcal{U}^{\text{sm}}_{d,z|x}(g) : d \in \mathcal{D},~ z \in \mathcal{Z},~x \in \mathcal{X} \} \cup \{\mathcal{U}^{\theta}_{d,l|x}(g) : d \in \mathcal{D},~ l \in \mathcal{L},~x \in \mathcal{X}\}$ denote the collection of sets underlying the data moments in (ref) and the parameter of interest in (ref). We require that the partition is rich enough to generate each set in this collection so that it can capture the information provided by the data and on the parameter of interest and hence produce the nonparametric bounds. We formalize this property in the following definition.\footnote{We note that our definition shares conceptual similarities to multidimensional partitions considered in discrete choice analysis to obtain nonparametric bounds in alternative identification problems, where one similarly requires the partition to be sufficiently rich so that each element of the partition predicts the same choice behavior chesher/etal:13, gu/etal:22, kamat/norris:22, tebaldi/etal:23.}
\begin{definition2}{P}{(Partition)} Let $\mathbb{U}(h)$ be a finite partition of $[0,1]^J$ such that for each $\mathcal{U}^* \in \mathbb{U}^{*}(h)$, we have that $\bigcup_{\mathcal{U} \in \mathbb{U}} \mathcal{U} = \mathcal{U}^*$ for some $\mathbb{U} \subseteq \mathbb{U}(h)$. \end{definition2}
Figure (ref) graphically illustrates how we can exploit the rectangular structure underlying the selection model and parameters in (ref) and (ref) to construct a partition satisfying Definition (ref) in the context of a simple version of Example (ref). Figure (ref)(a) first presents the sets in $\mathbb{U}^{*}(h)$. Figure (ref)(b) then shows how to exploit the rectangular structure of these sets to construct a partition satisfying Definition (ref). It takes the endpoints of the various rectangles to form disjoint rectangles that partition the entire space.
This idea can also be applied to construct a partition more generally. Specifically, for each $j=1,\ldots,J$, let $u_{(1),j} \leq \ldots \leq u_{(M_j),j}$ denote the ordered values of all the end points in the $j$th dimension of the underlying rectangular sets in (ref) and (ref) that form the sets in $\mathbb{U}^{*}(h)$ together with 0 and 1. Let $\mathbb{U}(h)$ then denote the partition obtained by
i.e. a collection of hyperrectangles constructed by taking the Cartesian product of the set of intervals based on adjacent endpoints in each dimension.
By construction, due to Definition (ref), the partition implies that for each $m \in \mathbf{M}^\dagger$ satisfying (ref), the parametrized version $m' = \sum_{k=0}^{K} \gamma_{k,d|x} b_{k}(h)$, where $b_k(h)$ satisfies (ref) and $\gamma_{k,d|x} = E[m_{d|x}(U) | U \in \mathcal{U}_k(h)]$, also satisfies (ref) and also generates the same parameter of interest as $m$ in the sense that $\theta(m,h) = \theta(m',h)$. In turn, provided that
for $\mathbf{\Gamma}(h)$ in (ref) so that the parametric $m'$ also satisfies the assumptions imposed on $m$, it will generate the same identified set for $\theta$ as $m$. We show this in the following proposition.
Requiring $\mathbf{\Gamma}(h)$ to satisfy (ref) restricts the sets of nonparametric MTRs to which Proposition (ref) applies. However, amongst the linear restrictions in Table (ref), the condition in (ref) fails only when Assumption $\text{M}_j$ is considered. For the remaining restrictions, namely Assumptions B, $\text{M}_{d',d''}$, $\text{BV}_{d',d''}$, and $S$, we can straightforwardly show that for every $m$ satisfying each of these assumptions, $\gamma$ constructed using $m$, i.e. where $\gamma_{k,d|x} = E[m_{d|x}(U) | U \in \mathcal{U}_k(h)]$, satisfies the linear restrictions stated in terms of $\gamma$ given which $\mathbf{\Gamma}(h)$ satisfies (ref)---see Table (ref) for the derivations of these linear restrictions. For such restrictions imposed on the MTRs, nonparametric bounds on the parameter of interest can continue to be computed using Proposition (ref).
Having computed $\theta(\mathbf{M}^*(h))$ given $h \in \mathbf{H}$, we next show how to ensure that we can feasibly take their unions over $h \in \mathbf{H}^*$ and compute the identified set in (ref). To do so, we make the following assumption on $\mathbf{H}$.
\begin{assumption2}{PS}{(Parametrizing Selection Primitives)} $\mathbf{H} = \bigcup_{\lambda \in \mathbf{L}} \mathbf{H}(\lambda)$, where $\mathbf{L} \subseteq \mathbf{R}^{d_{\lambda}}$ is known, and $\mathbf{H}^*(\lambda) = \{h \in \mathbf{H}(\lambda) : h \text{ satisfies \eqref{eq:D_moments}}\}$ is a singleton or an empty set for each $\lambda \in \mathbf{L}$. \end{assumption2}
Recall in the unidimensional case $h$ is point-identified as $F$ is known and $g$ can be identified by the moments in (ref). Assumption (ref) aims to extend this idea to the general case with multidimensional unobserved heterogeneity by assuming $h$ can be parameterized by $\lambda \in \mathbf{L}$ such that an analogous argument applies for each $\lambda \in \mathbf{L}$. In particular, for each $\lambda \in \mathbf{L}$, it requires $\mathbf{H}(\lambda)$ and the data variation in (ref) to be such that there exists a unique value of $h$ consistent with the data so that it is point identified or that no such $h$ exists and the model is misspecified.
For the purposes of computation, Assumption (ref) along with Proposition (ref) allows us to compute the identified set using a two-step strategy performed across $\lambda \in \mathbf{L}$, which we can implement in practice using a grid of values in $\mathbf{L}$. To see this, observe that (ref) can be rewritten as
where $h^*(\lambda)$ is the point-identified value of $h$ when $\mathbf{H}^*(\lambda)$ is non-empty, $\theta_L$ and $\theta_U$ are defined as in (ref), and $\mathbf{L}^* = \{\lambda \in \mathbf{L} : \mathbf{H}^*(\lambda) \neq \emptyset,~\mathbf{\Gamma}^{*}(h^*(\lambda)) \neq \emptyset \}$ is the set of $\lambda$ for which the model is not misspecified. In turn, in an inner step for each $\lambda \in \mathbf{L}$, we can compute $\theta^{\Gamma}(\mathbf{\Gamma}^*(h^*(\lambda)),h^*(\lambda))$ by solving the linear programs in (ref), where, as elaborated in Section (ref), the point-identified value $h^*(\lambda)$ can be computed by solving a minimization problem based on the distance of the parameters to the data in the moments in (ref). In an outer step, we can then take the union of these sets when non-empty across $\lambda \in \mathbf{L}$ to obtain the identified set $\Theta$.
We conclude this section by revisiting the examples in Section (ref) and highlighting how Assumption (ref) can be naturally satisfied in them under common parametric assumptions on $F$.
We next show how the two-step characterization of the identified set in (ref) lends itself to a tractable strategy to estimate the identified set and construct confidence intervals when the moments of the data in (ref) and (ref) can only be estimated using sample data.
To estimate the identified set, a natural starting point is the plug-in estimator that replaces $P_{d|z,x}$ and $E_{d|z,x}$ in (ref) and (ref) with their sample analogs. However, this approach can deliver an empty estimated set in practice, because the moments may not be exactly satisfied even when the true identified set is non-empty. Following mogstad/etal:18, we instead base our estimator on a criterion that measures how well different values of the underlying primitives match the observed moments and retains those that best match the moments.
To describe the estimation procedure, it is useful to first restate the identified set in terms of such criteria. For the purposes of the point-identified value $h^*(\lambda)$, let
subject to $\int_{\mathcal{U}_{d,z|x}^{\text{sm}}(g)} dF_{x} = \pi_{d|z,x}$ for $d \in \mathcal{D}$, $z \in \mathcal{Z}$, $x \in \mathcal{X}$, where $P$ and $\pi$ denote $(P_{d|z,x} : d \in \mathcal{D},~z \in \mathcal{Z},~x \in \mathcal{X})$ and $(\pi_{d|z,x} : d \in \mathcal{D},~z \in \mathcal{Z},~x \in \mathcal{X})$ in vector notation, and $\Omega_{S}$ is a positive definite weight matrix. This criterion measures the squared distance of the model implied moments to the data selection moments in (ref). As $Q_{\text{sel}}(h) = 0$ whenever $h$ matches these moments, $h^*(\lambda)$ for each $\lambda \in \mathbf{L}$ if it exists can then be defined by
For the criterion that captures the condition that $\mathbf{H}^*(\lambda)$ and $\mathbf{\Gamma}^*(h^*(\lambda))$ are non-empty, we first note that this can be equivalently stated as there existing $\gamma \in \mathbf{R}^{\text{dim}(\gamma)}$ such that
where $\mu(\lambda)$ and $C_2(\lambda)$ are estimable quantities that depend on the right hand sides of the moments in (ref) and (ref) evaluated at $h$ equal to $h^*(\lambda)$, and $C_1(\lambda)$, $C_3(\lambda)$ and $c(\lambda)$ are known quantities underlying the linear restrictions $\mathbf{\Gamma}^*(h^*(\lambda))$ and the condition that $h^*(\lambda) \in \mathbf{H}^*(\lambda)$. For brevity, we present the exact forms of these vectors and matrices in Appendix (ref). As in (ref), let
subject to $C_1(\lambda) \bar{\mu} + (C_1(\lambda) C_2(\lambda) + C_3(\lambda)) \gamma \leq c(\lambda)$, where $\Omega_M(\lambda)$ is a positive definite weight matrix, i.e. the minimum possible squared distance of the model implied moments to the data moments in (ref) and (ref) for a given $(\gamma,\lambda)$. Again, as $Q(\gamma,\lambda) = 0$ whenever $\gamma \in \mathbf{\Gamma}^*(h^*(\lambda))$ and $h^*(\lambda) \in \mathbf{H}^*(\lambda)$, $\mathbf{L}^*$ and $\mathbf{\Gamma}^*(h^*(\lambda))$ in (ref) can then be defined by
The estimated identified set $\widehat{\Theta}$ is obtained by substituting in the estimated versions of (ref), (ref) and (ref), where the criterions in their definitions are replaced by their sample analogs. In particular, the estimated version of $h^*(\lambda)$ in (ref) is given by
subject to $\int_{\mathcal{U}_{d,z|x}^{\text{sm}}(g)} dF_{x} = \pi_{d|z,x}$ for $d \in \mathcal{D}$, $z \in \mathcal{Z}$, $x \in \mathcal{X}$, where $\widehat{P}$ is the sample analog estimator of $P$ and $\widehat{\Omega}_{S}$ is a consistent estimator of $\Omega_{S}$. Similarly, the estimated value of $Q(\gamma,\lambda)$ used to obtain estimates of (ref) and (ref) is given by
subject to $C_1(\lambda) \bar{\mu} + (C_1(\lambda) \widehat{C}_2(\lambda) + C_3(\lambda)) \gamma \leq c(\lambda)$, where $\widehat{C}_2(\lambda)$ and $\widehat{\mu}(\lambda)$ are estimates of $C_2(\lambda)$ and $\mu(\lambda)$ obtained by replacing $P_{d|z,x}$, $E_{d|z,x}$ and $h^*(\lambda)$ in the expression provided in Appendix (ref) by their estimated analogs $\widehat{P}_{d|z,x}$, $\widehat{E}_{d|z,x}$ and $\widehat{h}^*(\lambda)$, and $\widehat{\Omega}_M(\lambda)$ is a consistent estimator of $\Omega_M(\lambda)$. However, simply using $\widehat{Q}$ to obtain the estimated versions of (ref) and (ref), and hence the identified set, can result in an empty set as $\min_{\gamma \in \mathbf{\Gamma}(\widehat{h}^*(\lambda))}\widehat{Q}(\gamma,\lambda)$ may not equal 0 for any $\lambda \in \mathbf{L}$. Following mogstad/etal:18, we therefore instead take the estimated versions of (ref) and (ref) to be
where $Q^* = \min_{\lambda \in \mathbf{L}} \min_{\gamma \in \mathbf{\Gamma}(\widehat{h}^*(\lambda))}\widehat{Q}(\gamma,\lambda)$, i.e. the subsets of $\mathbf{L}$ and $\mathbf{\Gamma}(\widehat{h}^*(\lambda))$ that come closest to satisfying the conditions on the estimated criterions up to a tolerance $\kappa$. By construction, these sets, and hence the estimated identified set, are never empty. Here $\kappa$ is a tuning parameter that goes to 0 with sample size at a specific rate and is necessary to ensure that the estimated set is consistent. In cases where $\mathbf{L}$ is a singleton and there exists a unique value of $\gamma$ that ensures $Q(\gamma,h^*(\lambda))$ = 0 so that the model is point identified, we can take $\kappa$ to equal 0.
The above procedure can be summarized by the following algorithm, which can in practice be implemented using a grid of values in $\mathbf{L}$:
In terms of computation, Step 1 is potentially the most expensive step as (ref) can be a nonlinear optimization problem depending on how $h$ enters the problem. For the parametric distributions considered for the examples in Section (ref), these nonlinear problems are similar to those encountered in the estimation of parametric discrete choice problems train:09. In contrast, as in the computations in (ref), the linearity of $\mathbf{\Gamma}$ and $\theta^{\Gamma}$ and the quadraticity of the sample analog $\widehat{Q}$ in (ref) in terms of $\gamma$ given $\lambda \in \mathbf{L}$ implies that Steps 2 and 4 are more structured and correspond to convex quadratic and quadratically-constrained quadratic programs in $(\bar{\mu},\gamma)$, respectively. In cases where we know the model is point identified and $\kappa$ is taken to be 0, Step 4 can be further simplified. In particular, let $\gamma^* = \text{arg}\min_{\gamma \in \mathbf{\Gamma}(\widehat{h}^*(\lambda))}\widehat{Q}(\gamma,\lambda)$ be the unique solution in Step 2 and the unique value in $\widehat{\mathbf{\Gamma}}^*(\widehat{h}^*(\lambda))$. In Step 4, we then simply have $\widehat{\theta}_L(\widehat{h}^*(\lambda)) = \widehat{\theta}_U(\widehat{h}^*(\lambda)) = \theta^{\Gamma}(\gamma^*,\widehat{h}^*(\lambda))$.
Confidence intervals can be constructed by test inversion, i.e. test at level $\alpha \in (0,1)$ the null hypothesis $H_0 : \theta_0 \in \Theta$ and then construct $100(1-\alpha)\%$-confidence intervals by collecting the values of $\theta_0$ that are not rejected. To do so in a computationally tractable manner, we again exploit, as in our estimation procedure, the fact the $h$ is point identified by $h^*(\lambda)$ in (ref) and the linearity of $\theta^{\Gamma}$ and $\mathbf{\Gamma}^*(h^*(\lambda))$ in $\gamma$ to test $H_{0}(\lambda): \theta_0 \in \theta^{\Gamma}(\mathbf{\Gamma}^*(h^*(\lambda)),h^*(\lambda)),~ h^*(\lambda) \in \mathbf{H}^*(\lambda)$. We then show how to aggregate these tests across $\lambda \in \mathbf{L}$.
For each $\lambda \in \mathbf{L}$, we exploit the observation that, similar to (ref), the null hypothesis can be reformulated as
i.e. testing whether $\gamma$ satisfies a linear system of inequalities, where $\mu(\lambda)$, $C_1(\lambda)$, $C_2(\lambda)$, $C_3(\lambda)$ and $c(\lambda)$ are specified as in (ref), but augmented to include the additional condition that $\theta^{\Gamma}(\gamma,h^*(\lambda)) = \theta_0$---see Appendix (ref) for their exact forms. This formulation allows us to use the procedure from cox/etal:24, who show how to test null hypotheses that can be stated as (ref). Similar to the criterion in (ref), it is based on a Quasi-Likelihood Ratio test statistic given by
subject to $C_1(\lambda) \bar{\mu} + (C_1(\lambda) \widehat{C}_2(\lambda) + C_3(\lambda)) \gamma \leq c(\lambda)$, where $\widehat{\mu}(\lambda)$ and $\widehat{C}_2(\lambda)$ denote estimates of $\mu(\lambda)$ and $C_2(\lambda)$, and $\widehat{\Sigma}(\lambda)$ denotes an estimate of the asymptotic variance of $(\widehat{\mu}(\lambda) - \mu(\lambda) + (\widehat{C}_2(\lambda) - C_2(\lambda))\gamma^*)$ with
subject to $C_1(\lambda) \mu(\lambda) + (C_1(\lambda) C_2(\lambda) + C_3(\lambda)) \gamma \leq c(\lambda)$. To compute the critical value, let $\widehat{\bar{\mu}}$ and $\widehat{\gamma}$ denote the minimizers of the problem in (ref), and let $\widehat{K}$ denote the set of row indices corresponding to the binding inequalities in $C_1(\lambda) \widehat{\bar{\mu}} + (C_1(\lambda) \widehat{C}_2(\lambda) + C_3(\lambda)) \widehat{\gamma} \leq c(\lambda)$. Moreover, let
where $[C_1(\lambda),C_3(\lambda)]_{\widehat{K}}$ and $(C_1(\lambda) \widehat{C}_2(\lambda) + C_3(\lambda))_{\widehat{K}}$ denote the submatrices of $[C_1(\lambda),C_3(\lambda)]$ and $(C_1(\lambda) \widehat{C}_2(\lambda) + C_3(\lambda))$ containing only the rows in $\widehat{K}$. The critical value $cv(\lambda,1-\alpha)$ is given by the $(1-\alpha)$-quantile of a chi-squared distribution with $\widehat{r}$ degrees of freedom. The test then equals $\phi(\lambda) = 1\{TS(\lambda) > cv(\lambda,1-\alpha)\}$, i.e. we reject if the test statistic is above the critical value and do not reject otherwise. Provided that $\widehat{\mu}(\lambda)$ is asymptotically normal and under weak assumptions on the consistency of $\widehat{C}_2(\lambda)$, cox/etal:24 show that this test controls size.
To aggregate these tests across $\lambda \in \mathbf{L}$ and test $H_0$, we exploit the principle underlying the so-called recycling test from bugni/etal:15. Let $\mathbf{L}_0 = \{\lambda \in \mathbf{L} : H_0(\lambda) \text{ is true} \}$ denote the set of $\lambda$ where $H_0(\lambda)$ is satisfied under $H_0$ and, similar to (ref)-(ref), let $\widehat{\mathbf{L}}_0 = \{\lambda \in \mathbf{L} : TS(\lambda) \leq \min_{\lambda \in \mathbf{L}}TS(\lambda) + \kappa \}$ denote a consistent estimate of this set, where $\kappa$ is again a tuning parameter necessary to ensure that the estimated set is consistent. The aggregated test then equals
i.e. we reject if the minimum test statistic across $\mathbf{L}$ is greater than the minimum critical value across $\widehat{\mathbf{L}}_0$. In Appendix (ref), we show that if $\phi(\lambda)$ controls asymptotic size for each $\lambda \in \mathbf{L}_0$ and $\widehat{\mathbf{L}}_0$ is a consistent estimate for $\mathbf{L}_0$, then the test in (ref) also controls asymptotic size.
As in our estimation procedure, we can summarize the test for a pre-specified choice of $\alpha \in (0,1)$ by the following algorithm:
For Step 1, we require, as in (ref), replacing $P_{d|z,x}$, $E_{d|z,x}$ and $h^*(\lambda)$ in the expression of $\mu(\lambda)$ and $C_1(\lambda)$ presented in Appendix (ref) by their estimated analogs $\widehat{P}_{d|z,x}$, $\widehat{E}_{d|z,x}$ and $\widehat{h}^*(\lambda)$. As in Step 1 of the estimation algorithm, this step can be computationally intensive as computing $\widehat{h}^*(\lambda)$ in (ref) can be a nonlinear problem. Step 2 is again more structured as the optimization problems in (ref) and (ref) are both convex quadratic problems. Given this algorithm, confidence intervals immediately follow by applying it over a fine grid of values of $\theta_0$ and then collecting the values where we do not reject. In practice, to speed up computation, this part can be easily parallelized across the different values of $\theta_0$.
In this section, we use our method to revisit the analysis of kline/walters:16, who evaluate Head Start (HS) using data from the Head Start Impact Study (HSIS).
HS is a federally funded program in the United States that provides free preschool to three- and four-year-old children from disadvantaged families. The HSIS was a randomized experiment conducted in Fall 2002 to evaluate the program. Due to oversubscription, HS offers were randomized among applicants, enabling experimental evaluation by comparing outcomes of children who received an offer with those who did not puma_etal:10.
However, HS is not the only available preschool option; many alternatives provide similar services and also receive public subsidies. A non-trivial proportion of families in the experiment enrolled in such preschools. As highlighted in KW, this feature clouds the interpretation of the experimental effects as they estimate a combination of HS effects relative to enrolling in such preschools and no preschools. Moreover, it complicates a cost-benefit analysis of HS using these effects as enrollment out of these preschools may introduce fiscal savings and as these preschools may adjust the number of slots offered in response to that of HS.
Motivated by this feature, KW formulate a selection model that governs choice into the different preschool options and how preschools allocate offers. But, in estimating this model using the experimental data, they impose various assumptions such that the model is point-identified. In what follows, we illustrate how our method can be used to analyze the robustness of their conclusions on the returns to HS to these assumptions.
We follow KW in mapping the experiment into the treatment effect setup in Section (ref). Let $Z \in \{0,1\}$ denote an indicator for whether the child was assigned to the treatment group and offered HS, $D \in \mathcal{D}$ and $Y$ denote their preschool choice and outcome of interest in their first year in the experiment, and $X$ denote covariates. As in KW, we take $\mathcal{D} = \{p,c,n\}$, where $p$ denotes HS program preschools, $c$ denotes competing non-HS preschools, and $n$ denotes no preschool; and the outcome of interest to be the average of the Woodcock Johnson III and Peabody Picture and Vocabulary Test scores, with each test score normed to have mean 0 and variance 1 in the control group by cohort. While KW consider a vector of covariates, we take $X$ to be a single binary covariate for simplicity as it is sufficient to present the main ideas. In particular, we take $X \in \{0,1\}$ to be an indicator for whether the child was in the three- or four-year-old cohort. Our analysis sample includes 3,627 children, all three- and four-year-olds in the sample with complete data for all variables mentioned above.
Table (ref) reports several motivating statistics. Panel (a) presents intention-to-treat (ITT) effects on enrollment into HS and outcomes defined by
and the instrumental variable (IV) estimand, defined as the average ratio of the two ITTs by cohort,
ITT$^Y$ reveals that the offer increased test scores, while the ITT$^D$ reveals noncompliance---not all children with an offer enrolled in HS and some without one did. The IV estimand accounts for this by estimating the effect of the offer for those who comply. We calculate a statistically significant effect of 0.230 standard deviations.\footnote{Our IV estimate is similar to the 0.247 standard deviation effect reported in KW. Minor differences reflect our simplified covariate specification and sample restrictions.}
As noted, evaluating the returns to HS using these experimental effects is complicated by the fact that many children without a HS offer enrolled in competing preschools. Panel (b) presents the proportion of children enrolled in the various care settings by offer status. Close to a third of the children without an offer enrolled in competing preschools, and the offer reduced this share. As the estimates below more precisely capture, this indicates that a non-trivial proportion of children induced into HS by the offer came from competing preschools.
We consider the general multinomial selection model outlined in kline/walters:16 that allows for the possibility that competing preschools were also oversubscribed and randomized offers similar to HS. Recall that $Z_p \equiv Z \in \{0,1\}$ denotes the observed indicator for whether an offer to HS was randomly assigned. Moreover, as in KW, we introduce $Z_c$ to denote an unobserved indicator for whether an offer to a competing preschool was randomly assigned. Let $\delta_{p|x} \equiv E[Z_p | X{=}x]$ and $\delta_{c|x} \equiv E[Z_c | X{=}x]$ denote the offer probabilities with which HS and competing preschools are rationed, respectively, conditional on the cohort $x \in \{0,1\}$.
Following offers that are distributed by lottery in a first stage, families make choices in a second stage, where those who do not receive an offer from a preschool can enroll there by exerting additional effort. Let $V_d(z_p,z_c)$ define the family's utility from alternative $d \in \mathcal{D}$ had their offer statuses from HS and competing preschools been set to $(z_p, z_c) \in \{0,1\}^2$. The potential preschool choice given offer statuses equals the utility-maximizing choice given by
and the observed choice is simply the utility-maximizing one given their realized offers, i.e. $D = D(Z_p,Z_c)$. As in KW, we take $V_n(z_p,z_c) = 0$, i.e. normalize the utility of no preschools to 0, and impose the restriction that
where $\tilde{U}_d$ is the alternative-specific unobserved heterogeneity and $\zeta_x > 0$, i.e. an offer to a preschool only increases the utility for that preschool and the effect of an offer on utilities is the same for both types of preschool.\footnote{We find that the estimated $\zeta_x$ is greater than 0 and hence that the assumption $\zeta_x > 0$ is satisfied by the model. However, we note that we explicitly make this relevance-type assumption as we require $\zeta_x \neq 0$ for the selection model to be identified—see Proposition (ref) below.} Moreover, following KW, we assume the distribution on the unobserved heterogeneity conditional on the cohort $x \in \{0,1\}$ to be $\Phi_{\rho_x}$, i.e. a standard bivariate normal distribution with a correlation parameter $\rho_x$ that varies by cohort.\footnote{The restriction on $V_p(z_p,z_c)$ and the distributional assumption on $(\tilde{U}_p,\tilde{U}_c)$ are stated in kline/walters:16. The restriction on $V_c(z_p,z_c)$ follows from the discussion in kline/walters:16, where they note when using their selection model to compute the proportion induced into competing preschools from an offer they assume the utility value from an offer is comparable to that of a HS offer. }
For the purposes of our method, recall that the selection model is required to satisfy Assumption (ref), i.e. it needs to be point-identified up to a finite-dimensional parameter. To this end, the above model corresponds to that in Example (ref) for $D(z_p,z_c)$, where the thresholds $\tilde{g} = (\tilde{g}_{d,z_p,z_c|x} : d \in \{p,c\},~z_p,z_c, x \in \{0,1\})$ satisfy the restrictions in (ref) and the distribution of the unobservables $(\tilde{U}_p,\tilde{U}_c)$ corresponds to $\Phi_{\rho_x}$, which can be captured by taking $\tilde{g}$ and the distribution of unobservables $\tilde{F} = (\tilde{F}_x : x \in \mathcal{X})$ to satisfy
for some values $(\zeta_{p|x},\zeta_{c|x},\zeta_{x},\rho_x) \in \mathbf{R}^2 \times \mathbf{R}_{++} \times (-1,1)$ for each $x \in \mathcal{X}$. However, unlike showing how Assumption (ref) can be satisfied in Example (ref) in Section (ref), a complication here is that one of the instruments, namely $Z_c$, is not observed. Nonetheless, given the additional restrictions that reduce the number of unknowns characterizing the thresholds, we can identify $(\zeta_{p|x},\zeta_{c|x},\zeta_{x},\rho_x)$ for a fixed value of the competing preschool offer probability $\delta_{c|x}$. In particular, instead of (ref)-(ref), note that the selection moments under the restrictions in (ref)-(ref) and accounting for the fact that $Z_c$ is not observed correspond to
\sloppy for $z_p,x \in \{0,1\}$, where $P_{1,c|x} \equiv \delta_{c|x}$ and $P_{0,c|x} \equiv 1 - \delta_{c|x}$, $\sigma_x = (\zeta_{p|x},\zeta_{c|x},\zeta_x,\rho_{x})$, and $H_{c,z_p}$ and $H_{p,z_p}$ capture the right hand side of these moments in terms of the parameters underlying the restrictions in (ref)-(ref). The following proposition states the identification result given a rank condition on $H = (H_{c,0},H_{c,1},H_{p,0},H_{p,1})$.
Given the rank condition, we can ensure Assumption (ref) by taking $\mathbf{H}(\lambda) = \phi(\tilde{\mathbf{H}})$ for $\lambda = (\delta_{c|0},\delta_{c|1}) \in [0,1]^2 = \mathbf{L}$, where $\tilde{\mathbf{H}}$ is the set of all values of $(\tilde{g},\tilde{F})$ that satisfy (ref)-(ref) and $\phi$ denotes the relation as defined in Example (ref) in Section (ref) such that $(g,F) = \phi((\tilde{g},\tilde{F}))$. It follows by Proposition (ref) that $\mathbf{H}^*(\lambda)$ is a singleton for each $\lambda \in \mathbf{L}$ as $(\tilde{g},\tilde{F})$ can be point identified by (ref)-(ref) taking $(\delta_{c|0},\delta_{c|1}) = \lambda$ given which $(g,F)$ can be identified by $\phi(\tilde{g}),\tilde{F}))$. While it is difficult to analytically show that the rank condition holds given the nonlinear structure of the Jacobian, we can verify it numerically in practice using a grid for the values of $(\sigma,\delta_c)$---see Appendix (ref) for details.\footnote{Existing results on analytically showing full rank of the Jacobian in cases with normally distributed errors are present only for the binary choice case, where one exploits related properties to that of the inverse Mills ratio chern/etal:23,chern/etal:24,han/vytlacil:17. Such results, however, do not straightforwardly extend to the multinomial case, and we hence leave it as an extension to future work.} Doing so we find that the Jacobian does indeed satisfy this condition.
In estimating the above model, KW additionally assume $\delta_{c|x} = 0$ for $x \in \mathcal{X}$.\footnote{While they do not explicitly make this assumption, we note that the special case of the model that they estimate in kline/walters:16 that does not introduce $Z_c$ is only logically consistent with the discussion of their general model when taking $\delta_{c|x} = 0$.} As formalized by the above proposition, this ensures that the selection model is point-identified. In the context of the above model, this restriction can be interpreted as having no lottery for competing preschools and hence eliminating the possibility that their slots are rationed in a manner similar to HS. However, as highlighted in kline/walters:16, it is reasonable to believe that such preschools do indeed ration slots as the majority of states do not have universal preschool mandates and preschools face relatively fixed budgets. As our method does not require such an assumption, we use it below to analyze the robustness of their conclusions to it by allowing $\delta_{c|x}$ to vary in $[0,1]$, in which case the selection model is not necessarily point-identified.
We focus on two sets of parameters that underlie the main conclusions in KW. The first set of parameters decompose the returns to a HS offer evaluated by the IV estimand into those arising from children drawn into HS from competing preschools versus no preschool. Under our selection model, as $\zeta_x > 0$ for $x \in \mathcal{X}$, we can decompose the IV estimand in (ref) as follows
where
capture the proportion induced by a HS offer into HS from alternative $d \in \{n,c\}$ and the (local) average effect for this subgroup of children for cohort $x \in \{0,1\}$, respectively, and
capture this proportion averaged across the two cohorts and the effect weighted by the proportion of the subgroup across the two cohorts, respectively. Here the first line in the decomposition, as outlined in kline/walters:16, follows from the arguments of imbens/angrist:94 as $\zeta_x \geq 0$ implies the monotonicity condition that if $D(0,Z_c) \neq D(1,Z_c)$ then $D(1,Z_c) = p$, while the second line averages the expressions across the two cohorts.
A challenge is that while $S_{np}$ in the decomposition is identified by the data through
due to the monotonicity condition that also implies (ref), $LATE_{np}$ and $LATE_{cp}$ are not. KW exploit the point-identified version of the selection model discussed above along with additional assumptions on the Marginal Treatment Response (MTR) functions under this model to identify these effects. We now examine how their conclusions hold under weaker alternatives.
As in Section (ref), we define the MTRs here by $m_{d|x}(u_p,u_c) = E[Y(d)|U_p = u_p, U_c = u_c, X=x]$ for $d \in \mathcal{D}$, $x \in \mathcal{X}$, where $U_p = \tilde{F}_{\tilde{U}_p|X}(\tilde{U}_p)$ and $U_c = \tilde{F}_{\tilde{U}_c|X}(\tilde{U}_c)$ correspond to transformed versions of the unobserved heterogeneity as discussed in Example (ref). Note that we do not condition on $\tilde{F}_{\tilde{U}_p - \tilde{U}_c}(\tilde{U}_p - \tilde{U}_c)$ here as it is deterministically implied by $U_p$ and $U_c$. To ensure point identification, KW assume the MTRs to be linear in $u_p$ and $u_c$ and separable in $x$, i.e.
To see why this implies point identification, observe the outcome moments in (ref) are given by
for $d \in \mathcal{D}$, $z_p,x \in \{0,1\}$, where $\mathcal{U}^{\text{sm}}_{d,z_p,z_c}(g)$ is defined as in Example (ref), and $(g,F)$ are implied by the restrictions on $(\tilde{g},\tilde{F})$ in (ref)-(ref) and the relation between $(g,F)$ and $(\tilde{g},\tilde{F})$ noted in Section (ref). For each $d$, we have four moments in (ref) as $z_p$ and $x$ take two values. Moreover, given that KW assume $\delta_{c|x}$ for $x \in \mathcal{X}$ to be known, Proposition (ref) implies that there exist only four unknowns underlying these moments, namely $(\gamma_{d,1},\gamma_{d,2},\mu_{d|0},\mu_{d|1})$. As these moments are linear in the MTRs and the MTRs are linear in the unknowns, we can invert this linear system to identify them---and thus the LATEs.
To illustrate how our method can assess robustness, Table (ref) reports estimates and 95% confidence intervals for $LATE_{np}$ and $LATE_{cp}$ under KW's specification and a range of alternatives.\footnote{The estimates and confidence intervals are obtained using the algorithms described in Section (ref). In our implementation, we take a grid for $\mathbf{L}$ given by $\{0,0.05,\ldots,0.95,1\}^2$, and $\mathcal{U}_{\text{grid}} = \{0,0.25,0.5,0.75,1\}^2$ when imposing the shape restrictions in the parametric case in Table (ref). For the estimation algorithm, we take $\widehat{\Omega}_{S}$ to equal the identity matrix, and $\widehat{\Omega}_{M}$ to be computed as in $\widehat{\Sigma}^{-1}(\lambda)$ in (ref) using $\mu(\lambda)$, $C_1(\lambda)$, $C_2(\lambda)$, $C_3(\lambda)$ and $c(\lambda)$ from (ref). We take $\kappa = 0.1$ in both algorithms. To compute $\widehat{\Sigma}(\lambda)$, we use the bootstrap, i.e. we take $\widehat{\Sigma}(\lambda) = 1/B\sum_{b=1}^B (\widehat{\mu}_b(\lambda) - \widehat{\mu}(\lambda) + (\widehat{C}_{2,b}(\lambda) - \widehat{C}_{2}(\lambda))\widehat{\gamma}^*)(\widehat{\mu}_b(\lambda) - \widehat{\mu}(\lambda) + (\widehat{C}_{2,b}(\lambda) - \widehat{C}_{2}(\lambda))\widehat{\gamma}^*)'$, where $\widehat{\gamma}^*$ corresponds to the estimated version of (ref) and $\widehat{\mu}_b(\lambda)$ and $\widehat{C}_{2,b}(\lambda)$ are the bootstrap analogs of $\widehat{\mu}(\lambda)$ and $\widehat{C}_{2}(\lambda)$ estimated using the $b$th bootstrap drawn with replacement at the level of the HS centers and where we take $B=200$.} We do so by allowing flexibility across two separate dimensions: the selection model and the MTRs. For each parameter, the upper left estimate reflects our analog to the baseline model in KW, i.e. the specification in (ref) under the assumption that $\delta_{c|x} = 0$. Moving across the columns highlights flexibility in the specification of the MTRs by relaxing the linearity and separability assumptions. Moving down a row relaxes the selection model assumption that there is no rationing for competing preschools, and instead allows them to be rationed by an unknown $\delta_{c|x} \in [0,1]$.
Under our analog of KW’s baseline specification, the estimates for $LATE_{np}$ and $LATE_{cp}$ reveal that HS increases test scores for those induced from no preschool and competing preschools, respectively, by 0.27 and 0.14. Consistent with KW, however, both of these effects are statistically imprecise. KW considers a number of additional covariates and fixed effects to improve precision. We instead consider shape restrictions to improve precision chetverikov/etal:18. To this end, Column (2) imposes a bounded variation assumption as in Table (ref) given by $|m_{p|x}(u_p,u_c) - m_{n|x}(u_p,u_c)| \leq 0.8$, $|m_{c|x}(u_p,u_c) - m_{n|x}(u_p,u_c)| \leq 0.8$, and $|m_{p|x}(u_p,u_c) - m_{c|x}(u_p,u_c)| \leq 0.2 $. We choose the 0.8 value because this is close to the largest effect for a preschool program noted in the literature kline/walters:16. Similarly, we bound the difference in effects between HS and competing preschools to be at most 0.2, since the two options provide similar services and hence are likely to have more similar effects.
Focusing on $LATE_{np}$ first, we find that its estimate under this specification remains substantively unchanged, capturing that the estimated MTRs satisfy the additional restrictions. But, consistent with KW’s results under their more precise specifications, its confidence interval shrinks, resulting in a statistically significant effect. In Columns (3)-(7), we analyze the robustness of this positive effect to relaxing the linearity and separability assumptions in (ref). Column (3) allows for a more flexible polynomial specification by adding interactions of $u_p$ and $u_c$ and their squares to (ref), and Column (4) allows the MTRs to be entirely nonparametric. Columns (5)-(7) further relax these specifications by removing separability. The parameter in these cases is partially identified, but our estimated bounds as well as confidence intervals reveal that neither linearity nor separability are required to conclude that $LATE_{np}$ is positive.
The second row, which allows for rationing of competing preschools ($\delta_{c|x} \in [0,1]$), reveals similar robustness. In this case, the confidence intervals for $LATE_{np}$ in Columns (2)-(7) lie entirely above zero (with lower bounds around 0.04-0.11). Hence the positive effect when competing preschools may be rationed also remains statistically significant even under nonparametric and nonseparable MTRs.
Turning to $LATE_{cp}$, again consistent with KW, the confidence intervals across these specifications cannot rule out a null effect. To better understand this, Table (ref) also presents results for $S_{np}$. Here, as it is directly identified by the data moments in (ref), we obtain estimates by simply taking sample analogs and confidence intervals using a standard $t$-test. The estimate suggests that the null finding arises because the data may simply be worse-powered for this effect; roughly only 30% of those induced into HS come from competing preschools.
To analyze whether stronger restrictions can allow us to reach a more informative conclusion, we consider a final specification in Column (8) that imposes on Column (2) two additional assumptions: (i) a positive selection assumption, similar to the Assumption $M_j$ in Table (ref), that takes $|m_{d|x}(u_p,u_c) - m_{n|x}(u_p,u_c)|$ to be decreasing in $u_p$ and $u_c$ for $d \in \{c,n\}$, i.e. those more likely to enroll in a preschool have a higher effect; and (ii) a positive effect assumption as in $M_{d,n}$ in Table (ref) for $d \in \{p,c\}$, i.e. children have higher test scores under preschool relative to no preschool. Even under these restrictions, the confidence intervals for $LATE_{cp}$ remain unaffected and continue to cover the entire logical region from -0.2 to 0.2 permitted under the bounded variation assumption. This strengthens our takeaway that the data do not rule out the possibility of a null effect.
The second set of parameters we analyze evaluate the policy returns to marginal expansions of HS by marginally changing the probability of a HS offer, $\delta_{p|x}$. Following KW, we first define the benefits and costs of a HS offer. Let $B$ denote the total after-tax lifetime income for all the children, which relates to test scores by
where $\phi_{\text{ben}}$ denotes the pre-specified relation between test scores and future income and $\tau$ the pre-specified tax rate faced by the children of eligible households. The net costs to the government of financing preschool are given by
where $\phi_p$ and $\phi_c$ denote the pre-specified costs of providing HS and competing preschool services, respectively, and $\tau \phi_{\text{ben}} E[Y]$ captures the revenue generated by taxes on the adult earnings of HS-eligible children. Exploiting the fact that $E[Y]$ and $P(D=d)$ can be written in terms of $E[Y(D(z_p,z_c))|X=x]$, $P[D(z_p,z_c)=d|X=x]$ for $d\in\{p,c\}$ and $(\delta_{p|x},\delta_{c|x})$, we can differentiate these expressions with respect to $\delta_{p|x}$ for $x \in \{0,1\}$ to obtain the marginal benefits and costs of expanding access to HS---see Appendix (ref) for details. In doing so, following KW, we either assume that $\delta_{c|x}$ does not adjust to changes in $\delta_{p|x}$, which we refer to as the non-adjusting case kline/walters:16, or that $\delta_{c|x}$ adjusts such that enrollment in competing preschools stays constant, i.e. $d P(D = c|X=x) / d \delta_{p|x} = 0$, which we refer to as the adjusting case kline/walters:16. Performing this exercise, we can obtain expressions for the marginal benefits (MB) and costs (MC) under the non-adjusting case given by
and under the adjusting case given by
where $\Delta_{R,p|x} = E[R(1,Z_c) - R(0,Z_c)|X=x]$ and $\Delta_{R,c|x} = E[R(Z_p,1) - R(Z_p,0)|X=x]$ for random variable $R$, and $d \delta_{c|x} / d \delta_{p|x} = -\Delta_{1\{D=c\},p|x} / \Delta_{1\{D=c\},c|x}$. The marginal benefits and costs are then compared using the marginal value of public funds (MVPF) hendren:16 defined by
Table (ref) reports estimates and 95$\%$ confidence intervals for the MVPF parameter under the non-adjusting case ($\text{MVPF}_{\text{non-adj}}$) and the adjusting case ($\text{MVPF}_{\text{adj}}$), for the specifications considered in Table (ref).\footnote{In unreported results, we also consider an additional MVPF parameter for the adjusting case derived under the assumption that competing preschool offers only draw children from no preschool alternatives and that the effect for these children equals $LATE_{np}$. While this assumption is not logically consistent with the selection model in Section (ref), we consider it for completeness as it is presented in KW. As it produces qualitatively similar to $\text{MVPF}_{\text{adj}}$, we focus on the latter, more logically consistent parameter for brevity.} For the pre-specified values in the expression, following KW, we take $\phi_{\text{ben}} = \kappa_{\text{ben}} E$, where $E$ corresponds to the average present discounted value of lifetime earnings for HS applicants and $\kappa_{\text{ben}}$ is the percentage change in earnings induced by a one standard deviation increase in test scores. We take $\kappa_{\text{ben}} = 0.13$, which corresponds to the estimate found in chetty/etal:11---see kline/walters:16 for an overview of such estimates in the literature. Moreover, following KW, we take $\tau = 0.35$, $E=\$343,392$, $\phi_p = \$8,000$ and $\phi_c = 0.75 \cdot \phi_p = \$6,000$.
Observe that $\text{MVPF}_{\text{non-adj}}$ is directly point identified by moments of the data. This is because $\Delta_{R,p} = E[R|Z_p=1] - E[R|Z_p=0]$, which simply corresponds to the ITT effects on the variable $R$. Similar to the parameter $S_{np}$ in Table (ref), we estimate it by taking sample analogs of the various ITTs, with confidence intervals derived by inverting a standard $t$-test. Consistent with KW, the estimate of $\text{MVPF}_{\text{non-adj}}$ is greater than one and the lower bound of the confidence interval exceeds one, suggesting that, absent adjustment of competing preschools to HS expansion, marginally expanding HS access yields benefits that outweigh the costs.
For $\text{MVPF}_{\text{adj}}$, as in Table (ref), we obtain estimates and confidence intervals following the algorithms in Section (ref), but slightly adapted to account for the piecewise definition of the MVPF parameter in (ref)---see Appendix (ref) for details. Consistent again with KW, we find that under the linear and separable specification on the MTRs in Columns (1)-(2) and when $\delta_{c|x} = 0$, the estimates highlight that $\text{MVPF}_{\text{adj}}$ is also larger than one. Keeping $\delta_{c|x} = 0$, Columns (3), (5), (6), and (8) reveal this finding is robust to alternative parametric specifications, but Columns (4) and (7) show that the lower bound of the identified set falls below one when the MTRs are nonparametric.
When allowing for the possibility that competing preschools are rationed as in HS by allowing $\delta_{c|x} \in [0,1]$, Columns (1)-(7) show that the lower bound on $\text{MVPF}_{\text{adj}}$ falls below one across all specifications. Only under the additional restrictions of positive selection and positive effects in Column (8) does the lower bound of the identified set remain above one. Furthermore, across all specifications, the confidence intervals do not rule out the possibility that $\text{MVPF}_{\text{adj}}$ is smaller than one.
We conclude our analysis by analyzing whether alternative values for $\kappa_{\text{ben}}$, i.e. the relation between test scores and earnings, can potentially result in $\text{MVPF}_{\text{adj}}$ being statistically greater than one. In Figure (ref), we report $p$-values for the null hypothesis that $\text{MVPF}_{\text{adj}} \leq 1$ for alternative values of $\kappa_{\text{ben}}$ under Specifications (2) and (8) from Table (ref). Our results reveal that for $\delta_{c|x} \in [0,1]$ and under additional sufficiently strong restrictions on the MTRs, namely specification (8), we can statistically reject that $\text{MVPF}_{\text{adj}}$ is at most one at a 5$\%$ significance level if we are willing to assume that the percentage increase in earnings induced by a one standard deviation increase in test scores is close to 30$\%$. This value is near the upper bound of estimates found in the literature, comparable to the Perry Preschool Program heckman/etal:10.
In cases with multiple treatments, treatment selection models generally exhibit multidimensional unobserved heterogeneity. In this paper, we develop a marginal treatment effect based method to learn about treatment effects in a general class of such models with discrete-valued instruments. Allowing the selection model to be identified up to a finite-dimensional parameter, we show how a two-step computational program can be used to compute the identified set for the treatment effect parameters when the marginal treatment response functions underlying them remain nonparametric or are additionally parameterized. We demonstrate the benefits of our method by revisiting the empirical analysis of the Head Start program by kline/walters:16.