EconBase
← Back to paper

Identification in Multiple Treatment Models under Discrete Variation

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

129,297 characters · 21 sections · 86 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification in Multiple Treatment Models under Discrete Variation

abstractWe develop a marginal treatment effect based method to learn about causal effects in multiple treatment models with discrete instruments. We allow selection into treatment to be governed by a general class of threshold crossing models that permit multidimensional unobserved heterogeneity. An inherent complication is that the primitives characterizing the selection model are not generally point-identified. Allowing these primitives to be point-identified up to a finite-dimensional parameter, we show how a two-step computational program can be used to obtain sharp bounds for a number of treatment effect parameters when the marginal treatment response functions are allowed to satisfy only nonparametric shape restrictions or are additionally parameterized. We demonstrate the benefits of our method by revisiting kline/walters:16' (kline/walters:16) empirical analysis of the Head Start program. Our approach relaxes their point-identifying assumptions on the selection model and marginal treatment response functions, allowing us to assess the robustness of their conclusions.

\thispagestyle{empty}

KEYWORDS: Multiple treatments, discrete instrument, marginal treatment effects, partial identification, Head Start.

JEL classification codes: C14, C31, C36, C61, I21.

\setcounter{page}{1}

Introduction

In the analysis of treatment effects using instrumental variables, a threshold crossing selection model with a single dimension of unobserved heterogeneity underlies the predominant framework in the literature imbens/angrist:94, heckman/vytlacil:05, vytlacil:02. Designed for the canonical binary treatment setup, it presumes that individuals choose between only two treatments. In many scenarios, however, individuals choose between multiple treatments. Such cases naturally give rise to selection models with multiple dimensions of unobserved heterogeneity heckman/etal:06,heckman/etal:08,lee/salanie:18.

In this paper, we develop a method to learn about treatment effects in multiple treatment setups under a general class of selection models with multidimensional unobserved heterogeneity. Specifically, we consider the class of multidimensional threshold crossing models from lee/salanie:18. As highlighted in Section (ref), this class nests the standard binary treatment model and also covers many empirically relevant settings with multidimensional unobserved heterogeneity, including multinomial, sequential, multi-stage, and double-hurdle choice models.

In this class of models, our objective is to show how to learn about various treatment effects that can be written as weighted averages of the marginal treatment response functions (MTRs) heckman/vytlacil:99,heckman/vytlacil:05,mogstad/etal:18, which capture the mean potential outcomes conditional on the unobserved heterogeneity governing the selection model. As in the standard binary treatment case, MTRs provide a unified framework to express a number of parameters of interest such as the average treatment effect as well as other averages that target policy relevant subgroups.

To analogously learn about such parameters, lee/salanie:18 assume availability of data with continuous variation in the instrument. They show how a sufficient amount of such variation allows nonparametric identification of the selection model and the MTRs over their entire support, and in turn all the parameters of interest---see also heckman/etal:06, heckman/etal:08, mountjoy:22 and tsuda:23 for related arguments in specific models. However, in many empirical scenarios, instruments are only discrete-valued. In such cases, the parameters are usually partially identified as there generally exist multiple admissible values of the MTRs and primitives characterizing the selection model that are consistent with the data.

We show how to compute the identified set for our parameters of interest in these cases. In particular, we build on mogstad/etal:18, who study this identification problem in the special case of a binary treatment, where a single dimension of unobserved heterogeneity underlies selection. In this case, the primitives characterizing the selection model are point identified and the problem is linear in the remaining primitives, namely the MTRs. Imposing a linear parametrization on the MTRs, mogstad/etal:18 show how these features allow computing sharp bounds by solving two linear programs. They also show that with carefully chosen basis functions over a unidimensional partition, such parametrizations can accommodate fully nonparametric MTRs satisfying only shape restrictions. However, two inherent challenges arise when multidimensional unobserved heterogeneity underlies selection: the primitives characterizing the selection model are not point-identified, and the MTRs are no longer unidimensional.

We propose a two-step solution that accounts for these complications. In an inner step, for a given selection primitive, we first exploit the fact that the problem remains linear in the MTRs. Similarly linearly parametrizing the MTRs, we show this implies that we can continue to use linear programs to compute sharp bounds given the selection primitives. Moreover, we introduce novel arguments to show that such a result can also continue to accommodate entirely nonparametric MTRs. Specifically, as expanded in Section (ref), we show how to leverage the rectangular structure underlying our general selection model and construct a new partition that generalizes the unidimensional one from mogstad/etal:18 to the multidimensional case.

In an outer step, we then introduce alternative arguments to account for the fact that the selection model is not point-identified. Specifically, this step takes the unions of the first step bounds across admissible values of the selection primitives. To do so feasibly, we assume that these primitives can be point identified up to a finite dimensional parameter. We demonstrate in several empirically relevant models how this assumption can be satisfied under familiar parametric assumptions on the distribution of the unobserved heterogeneity. In summary, our method translates to solving two linear programs to compute the identified set at the point identified value of the selection model across various values of the finite dimensional parameter, and then taking the union of these sets. We show how this two-step characterization also translates into tractable two-step strategies to estimate the identified set and construct confidence intervals.

Our method offers several advantages over recent approaches to point identification in settings with multiple treatments and discrete instruments. A common method is to impose monotonicity assumptions that restrict selection patterns such that the researcher can nonparametrically point-identify treatment effects for certain response types angrist/imbens:95,bhuller/sigstad:24,goff:25,heckman/pinto:18,kline/walters:16, kirkeboen/etal:16,lee/salanie:23,navjeevan/etaL:23,pinto:19. As we highlight in Section (ref), many such monotonicity assumptions can be captured by selection models in our general class. In such cases, although our method parametrizes the selection model, sufficiently flexible specifications can exactly replicate the nonparametric estimates these approaches produce.\footnote{kline/walters:19 and mogstad/etal:18 similarly note in the binary treatment case that even if the selection model and MTRs are parametrized, one can produce the nonparametric point identified local average treatment effect.} But our method also importantly extends beyond these settings: it allows researchers to learn about effects that are only partially identified and to accommodate richer selection patterns that would otherwise preclude point identification.

A smaller group of papers similarly considers selection models to learn about additional treatment effects, but limits attention to specific models and parameterizations that are point identified by the data hull:20, kline/walters:16,pinto:19. Our method contributes to these papers by nesting such point identification as a special case while allowing more robust parameterizations of the selection model and MTRs that generally imply bounds on the treatment effects of interest. Moreover, our analysis applies to a general class of selection models. Our method therefore provides a blueprint for robustly learning about treatment effects across a wide range of empirical settings.

Our method also has advantages relative to a complementary literature on partial identification in multiple treatment models. Papers in this literature either make no assumptions on selection manski:97,manski/pepper:00,manski/pepper:09 or impose only nonparametric restrictions bai/etal:2025,kamat:21,lee/salanie:23. However, because these approaches do not model the MTRs, they cannot accommodate parameters such as policy-relevant treatment effects or assumptions such as parametric restrictions and separability that are common in empirical work. Our method allows for such parameters and assumptions as well as flexible variants that enable researchers to assess the robustness of their conclusions.

To concretely demonstrate these various benefits, we use our method to revisit kline/walters:16' (kline/walters:16) empirical analysis of the Head Start preschool program using the Head Start Impact Study. Due to oversubscription, Head Start offers were randomized to applicants in the study, which enabled an experimental evaluation of Head Start by comparing outcomes of children with offers to those without. However, what these experimental effects can identify on the effects of Head Start and their policy relevance are clouded by the fact that many children enroll in competing non-Head Start preschools that provide similar services to Head Start and are also publicly subsidized. To account for this complication, kline/walters:16 formulate a multinomial selection model of preschool choice and slot provision. But, in estimating this model, they impose two simplifying assumptions to ensure that both the selection model and MTRs are point-identified: competing preschools do not ration slots, and the MTRs are linear and separable.

We illustrate how our method can be used to gradually weaken these assumptions in a step-by-step manner and analyze the sensitivity of their main empirical conclusions to them. Our bounds reveal that their first conclusion that the positive experimental effects are driven by a positive effect from families who enroll in no preschools absent a Head Start offer is robust to these assumptions. In particular, this conclusion can hold under nonparametric and nonseparable MTRs, and when the selection model allows rationing of competing preschools. In contrast, for their second conclusion that marginally expanding Head Start has a positive, statistically significant effect, we find their assumptions are generally necessary unless one is willing to take the relation between test score and earnings to be towards the higher end of values reported in the literature.

The remainder of the paper is organized as follows. Section (ref) introduces the class of models and treatment effect parameters we analyze. Section (ref) presents our arguments to compute the identified set. Section (ref) discusses how these arguments can be used to estimate the identified set and construct confidence intervals. Section (ref) presents our empirical application. Section (ref) concludes. Proofs of all results are presented in Appendix (ref).

Setup

Observed and Potential Variables

For each individual, we observe an outcome of interest $Y$, their received treatment $D$, their assigned instrument value $Z$, and baseline covariates $X$. We take the list of possible treatments to be given by a discrete set $\mathcal{D}$, and model the support of the instrument and covariates denoted by $\mathcal{Z}$ and $\mathcal{X}$ to be discrete. We assume the outcome and treatment to be generated by the usual potential outcomes structure. In particular, denoting by $D(z)$ the potential treatment had the individual's assigned instrument value been $z \in \mathcal{Z}$, the observed treatment is given by

align[align omitted — 77 chars of source]

and, similarly, denoting by $Y(d)$ the potential outcome had the individual's received treatment been $d \in \mathcal{D}$, the observed outcome is given by

align[align omitted — 77 chars of source]

Moreover, as usual, we assume that the instrument is statistically independent of the potential variables conditional on the covariates as follows.

\begin{assumption2}{E}{(Exogeneity)} $(\{Y(d) : d \in \mathcal{D}\}, \{D(z) : z \in \mathcal{Z}\}) \perp Z | X~.$ \end{assumption2}

Selection Model

We assume the potential treatments to be related across instrument values through a selection model. Let $U \equiv (U_1, \ldots, U_J)$ denote a vector of individual unobservables governing selection. Following lee/salanie:18, we take potential treatment to be determined by the region in which $U$ falls, where these regions satisfy certain properties in addition to being disjoint to logically ensure that they cannot imply several different treatments.

\begin{assumption2}{SM}{(Selection Model)} Conditional on $X = x \in \mathcal{X}$, let $U$ be continuously distributed with support contained in $[0,1]^J$ with uniform marginal distributions, and let

align[align omitted — 105 chars of source]

for $z \in \mathcal{Z}$, where $\{\mathcal{U}^{\text{sm}}_{d,z|x} : d \in \mathcal{D}\}$ denotes a collection of disjoint subsets of $[0,1]^J$ that are members of the $\sigma$-field generated by the sets $\{\{u \in [0,1]^J : u_j \leq g_{j,z|x}\} : j =1,\ldots,J\}$ for some unknown threshold values $\{g_{j,z|x} : j = 1, \ldots, J\}$. \end{assumption2}

As the generated $\sigma$-field is obtained by taking unions, intersections and complements of the sets in $\{\{u \in [0,1]^J : u_j \leq g_{j,z|x}\} : j =1,\ldots,J\}$, Assumption (ref) requires that each of the regions $\mathcal{U}^{\text{sm}}_{d,z|x}$ determining potential treatment can be written as

align[align omitted — 152 chars of source]

for some finite $L_{d}$ and $\underline{u}_{j,l,d,z|x},\bar{u}_{j,l,d,z|x} \in \{0,g_{j,z|x},1\}$, i.e. as a finite union of rectangular sets. As usual, the restriction in Assumption (ref) that $U$ has support contained in $[0,1]^J$ with uniform marginal distributions can be viewed as a normalization. To see this, let $\tilde U = (\tilde U_1,\dots,\tilde U_J)$ denote a vector of latent unobservables with continuous marginal distribution functions $\tilde F_{j\mid x}$ given $X=x$. For any thresholds $\underline{\tilde u}_{j,l,d,z\mid x}$ and $\bar{\tilde u}_{j,l,d,z\mid x}$, \[ 1\{\underline{\tilde u}_{j,l,d,z\mid x} \le \tilde U_j \le \bar{\tilde u}_{j,l,d,z\mid x}\} = 1\{\tilde F_{j\mid x}(\underline{\tilde u}_{j,l,d,z\mid x}) \le \tilde F_{j\mid x}(\tilde U_j) \le \tilde F_{j\mid x}(\bar{\tilde u}_{j,l,d,z\mid x})\}. \] Defining $U_j = \tilde F_{j\mid x}(\tilde U_j)$, $\underline{u}_{j,l,d,z\mid x} = \tilde F_{j\mid x}(\underline{\tilde u}_{j,l,d,z\mid x})$, $ \bar u_{j,l,d,z\mid x} = \tilde F_{j\mid x}(\bar{\tilde u}_{j,l,d,z\mid x})$, we obtain a normalized vector $U=(U_1,\dots,U_J)$ with support $[0,1]^J$ and uniform marginal distributions conditional on $X=x$, and selection regions of the form in (ref) with endpoints $\{\underline{u}_{j,l,d,z\mid x},\bar{u}_{j,l,d,z\mid x}\}$.

As highlighted in lee/salanie:18, Assumption (ref) reduces to the standard threshold crossing model heckman/vytlacil:05, imbens/angrist:94 when we have a binary treatment and a single dimension of unobserved heterogeneity given by

align[align omitted — 63 chars of source]

for $z \in \mathcal{Z}$ conditional on $X = x \in \mathcal{X}$. But, in cases where $J > 1$, it provides a unified framework to capture a number of additional common, empirically relevant models of selection with multiple as well as binary treatments. Below, we briefly provide examples of several such models. Our first three examples consider multiple treatment models, while our fourth example considers a binary treatment model. For simplicity, when possible, we take the number of treatments in the multiple treatment examples to be solely equal to three.

example{(Multinomial Choice)} Let $\mathcal{D} = \{0,1,2\}$, and, conditional on $X = x \in \mathcal{X}$, let \begin{align} D(z) = \operatorname*{arg\,max}_{d \in \mathcal{D}} \tilde{g}_{d,z|x} - \tilde{U}_d \end{align} for $z \in \mathcal{Z}$, where $(\tilde{g}_{1,z|x},\tilde{g}_{2,z|x})$ is unknown, $(\tilde{U}_1,\tilde{U}_2)$ some continuous unobserved variables, and $\tilde{g}_{0,z|x}$ and $\tilde{U}_0$ are normalized to 0. This is the standard additively separable utility model of multinomial choice as studied in heckman/etal:06,heckman/etal:08 and is empirically used to model selection such as that of children into schools abdulkadirouglu/etaL:20, kline/walters:16 or patients into hospitals hull:20. Alternatively, several papers consider monotonicity restrictions on potential treatments that can be equivalently captured by this selection model. For instance, taking $\mathcal{D}$ to denote different fields of study and $\mathcal{Z} = \{0,1,2\}$ with $z \in \{1,2\}$ denoting offer to field $d = z$ and $z=0$ denoting no offer, kirkeboen/etal:16 impose $D(0) \neq D(d) \implies D(d) = d$ for $d \in \{1,2\}$, i.e. if an offer to field $d$ affects decisions then it does so by only encouraging students into that field. As shown in bai/tabord:25, this corresponds to equivalently assuming (ref) with restricting the thresholds to satisfy $\tilde{g}_{1,0|x} = \tilde{g}_{1,2|x}$ and $\tilde{g}_{2,0|x} = \tilde{g}_{2,1|x}$, i.e. the thresholds for the fields are affected only by its offers. In a similar manner, the related monotonicity restriction in kline/walters:16 can also be equivalently captured by (ref). To write (ref) in terms of Assumption (ref), following lee/salanie:18, let $U_j = \tilde{F}_{j|x}(\tilde{U}_j)$ and $g_{j,z|x} = \tilde{F}_{j|x}(\tilde{g}_{j,z|x})$ for $j \in \{1,2\}$, where $\tilde{F}_{j|x}$ denotes the conditional on $X=x$ distribution function of $\tilde{U}_j$, and let $U_3 = \tilde{F}_{12|x}(\tilde{U}_1 - \tilde{U}_2)$ and $g_{3,z|x} = \tilde{F}_{12|x}(\tilde{g}_{1,z|x} - \tilde{g}_{2,z|x})$, where $\tilde{F}_{12|x}$ is the conditional on $X=x$ distribution function of $\tilde{U}_1 - \tilde{U}_2$. Observe that it then follows that (ref) can be equivalently written as \begin{align} D(z) = \begin{cases} 0 & if U_1 > g_{1,z|x}, U_2 > g_{2,z|x} , \\ 1 & if U_1 \leq g_{1,z|x}, U_3 \leq g_{3,z|x} ,\\ 2 & if U_2 \leq g_{2,z|x}, U_3 > g_{3,z|x} , \end{cases} \end{align} \sloppy which can be straightforwardly re-written in terms of Assumption (ref) by taking $\mathcal{U}^{\text{sm}}_{0,z|x} = \{(u_1,u_2,u_3) \in [0,1]^3 : u_1 > g_{1,z|x},~u_2 > g_{2,z|x} \}$, $\mathcal{U}^{\text{sm}}_{1,z|x} = \{(u_1,u_2,u_3) \in [0,1]^3 : u_1 \leq g_{1,z|x},~u_3 \leq g_{3,z|x} \}$, and $\mathcal{U}^{\text{sm}}_{2,z|x} = \{(u_1,u_2,u_3) \in [0,1]^3 : u_2 \leq g_{2,z|x},~u_3 > g_{3,z|x} \}$.
example{(Sequential Choice)} Let $\mathcal{D} = \{0,1,2\}$, and, conditional on $X = x \in \mathcal{X}$, let \begin{align} D(z) = \begin{cases} 0 & if U_1 > g_{1,z|x} , \\ 1 & if U_1 \leq g_{1,z|x} , U_2 > g_{2,z|x} ,\\ 2 & if U_1 \leq g_{1,z|x} , U_2 \leq g_{2,z|x} , \end{cases} \end{align} for $z \in \mathcal{Z}$, where $(g_{1,z|x},g_{2,z|x})$ is unknown. This model imposes a sequential nature to the decisions where individuals first select into treatment 0 or not, and if not, then select into treatments 1 or 2. For instance, such a model has been used empirically to model judicial decisions for defendants, where judges first decide to convict or not and, if convicted, to additionally incarcerate or not arteaga:21,kamat/etal:22. Additionally, a special case of this model when $U_1 = U_2$ corresponds to that studied in heckman/etal:06, and is also used empirically to model various decisions such as years in preschool or sentence length in prison cornelissen/etal:18, rose/shem:21. Moreover, as shown in kamat/etal:22, this special case also corresponds to limiting versions of the monotonicity restrictions in angrist/imbens:95 and heckman/pinto:18. Observe that (ref) can be written in terms of Assumption (ref) by taking $\mathcal{U}^{\text{sm}}_{0,z|x} = \{(u_1,u_2) \in [0,1]^2 : u_1 > g_{1,z|x} \}$, $\mathcal{U}^{\text{sm}}_{1,z|x} = \{(u_1,u_2) \in [0,1]^2 : u_1 \leq g_{1,z|x},~u_2 > g_{2,z|x} \}$, and $\mathcal{U}^{\text{sm}}_{2,z|x} = \{(u_1,u_2) \in [0,1]^2 : u_1 \leq g_{1,z|x},~u_2 \leq g_{2,z|x} \}$.
example{(Multi-stage or Dynamic Choice)} Assumption (ref) also permits more general versions of Examples (ref)-(ref) such as multi-stage or dynamic selection models, which can be represented as decision trees and where a sequence of threshold-crossing rules determine the final treatment. For instance, in the reanalysis of the Moving to Opportunity experiment to evaluate the effects of moving to low-poverty neighborhoods and mental health, navjeevan/etaL:23 take $\mathcal{Z} = \{0,1\}$, where $z$ denotes an indicator for a MTO voucher, and $\mathcal{D} = \{(d_0,d_1) : d_0,d_1 \in \{0,1\} \}$, where $d_0$ denotes an indicator for whether households moved to low-poverty neighborhoods and $d_1$ denotes an indicator for whether the head of the households reported having positive mental health. Following vytlacil:02, the monotonicity restrictions navjeevan/etaL:23 impose can be stated as a selection model given by \begin{align} D(z) = \begin{cases} (0,0) & if U_1 > g_{1,z|x} , U_2 > g_{2,z|x} , \\ (0,1) & if U_1 > g_{1,z|x} , U_2 \leq g_{2,z|x} ,\\ (1,0) & if U_1 \leq g_{1,z|x} , U_3 > g_{3,z|x} , \\ (1,1) & if U_1 \leq g_{1,z|x} , U_3 \leq g_{3,z|x} , \end{cases} \end{align} for $z \in \{0,1\}$, where $(g_{1,z|x},g_{2,z|x},g_{3,z|x})$ is unknown, i.e. a threshold-crossing rule determines whether households relocate in the first stage and another threshold crossing rule then determines their mental health in the second stage, along with the additional restriction that $U_2 = U_3$, i.e. the same index underlies the second stage regardless of the relocation decision, and $g_{2,0|x} = g_{2,1|x}$ and $g_{3,0|x} = g_{3,1|x}$, i.e. the voucher affects only the decision to relocate and not mental health. Observe that (ref) can be written in terms of Assumption (ref) by taking $\mathcal{U}^{\text{sm}}_{(0,0),z|x} = \{(u_1,u_2,u_3) \in [0,1]^3 : u_1 > g_{1,z|x},~u_2 > g_{2,z|x}\}$, $\mathcal{U}^{\text{sm}}_{(0,1),z|x} = \{(u_1,u_2,u_3) \in [0,1]^3 : u_1 > g_{1,z|x},~u_2 \leq g_{2,z|x}\}$, $\mathcal{U}^{\text{sm}}_{(1,0),z|x} = \{(u_1,u_2,u_3) \in [0,1]^3 : u_1 \leq g_{1,z|x},~u_3 > g_{3,z|x}\}$, $\mathcal{U}^{\text{sm}}_{(1,1),z|x} = \{(u_1,u_2,u_3) \in [0,1]^3 : u_1 \leq g_{1,z|x},~u_3 \leq g_{3,z|x}\}$. In a similar manner, we note that versions of the dynamic model studied in heckman/etal:16 can also be captured by Assumption (ref).
example{(Double Hurdle)} Let $\mathcal{D} = \{0,1\}$, and, for each $z \in \mathcal{Z}$ conditional on $X = x \in \mathcal{X}$, let \begin{align} D(z) = 1\{U_1 \leq g_{1,z|x}, U_2 \leq g_{2,z|x}\} , \end{align} where $(g_{1,z|x},g_{2,z|x})$ is unknown, i.e. both unobserved variables have to be below their thresholds for the individual to receive treatment. Such a model can arise in settings where an individual has to make multiple decisions to receive a treatment poirier:80 or participate in a survey dutz/etal:21, and is also related to the selection model that arises under the partial monotonicity assumption of mogstad/etal:21, mogstad/etal:24 under multiple instruments. Observe that (ref) can be written in terms of Assumption (ref) by taking $\mathcal{U}^{\text{sm}}_{0,z|x} = \{(u_1,u_2) \in [0,1]^2 : u_1 > g_{1,z|x} \} \cup \{(u_1,u_2) \in [0,1]^2 : u_2 > g_{2,z|x} \}$, and $\mathcal{U}^{\text{sm}}_{1,z|x} = \{(u_1,u_2) \in [0,1]^2 : u_1 \leq g_{1,z|x},~u_2 \leq g_{2,z|x} \}$.

Parameters of Interest

In our model, we take the primitives to consist of three objects. First, the marginal treatment response (MTR) functions heckman/vytlacil:99,heckman/vytlacil:05,mogstad/etal:18, defined by $m_{d|x}(u) = E[Y(d)\mid U {=} u, X {=} x]$, which capture the mean treatment response for $d \in \mathcal{D}$ conditional on the unobserved heterogeneity $U {=} u$ and covariates $X {=} x$. Second, the distribution of these unobservables conditional on $X {=} x$, denoted by $F_x$. Third, all threshold values determining selection in Assumption (ref), denoted by $g = \{g_{j,z|x} : j = 1,\ldots,J,~z \in \mathcal{Z},~x \in \mathcal{X}\}$. Let $\beta \equiv (m, h)$ summarize these primitives, where $m$ collects the MTRs and $h = (F, g)$ comprises the selection primitives.

Our parameters of interest are functionals of these primitives. In particular, our analysis applies to any parameter that can be written as a weighted sum of integrals of the MTRs over specific regions of the unobserved heterogeneity. Formally, we consider parameters of the form

align[align omitted — 195 chars of source]

where $\mathcal{L}$ is a known index set that we sum across in addition to the covariate value $\mathcal{X}$ and treatment values $\mathcal{D}$, $w_{d,l|x}(h)$ determines the weight each integral receives, and $\mathcal{U}^{\theta}_{d,l|x}(g)$ is the integration region specifying which individuals are included in each weighted average. We assume that the weights $w_{d,l|x}(h)$ are known or identified functions of the selection primitives $h$, and the integration regions $\mathcal{U}^{\theta}_{d,l|x}(g)$ can be expressed as a finite union of sets of the form

align[align omitted — 98 chars of source]

where the endpoints $\underline{u}_{j,d,l|x}(g)$ and $\bar{u}_{j,d,l|x}(g)$ are known or identified functions of the thresholds $g$. This rectangular structure mirrors that imposed on the selection regions in (ref), which plays a key role in our nonparametric analysis in Section (ref).

As illustrated in Table (ref), a number of commonly studied treatment effect parameters can be written in terms of (ref). These include the average treatment effect (ATE) between two treatments $d',d'' \in \mathcal{D}$ as well as their analogs that condition on $D=d$ to give the average treatment effect on the treated (ATT) or on a response type $D(z) = d_z \in \mathcal{D}$ for $z \in \mathcal{Z}' \subseteq \mathcal{Z}$ to give local average treatment effects (LATEs). Moreover, they also include the class of policy relevant treatment effects heckman/vytlacil:99,heckman/vytlacil:05 that evaluate the effects of altering the thresholds determining selection. To formally define these latter effects, note that the sets determining treatment in (ref) can be more explicitly stated as $\mathcal{U}^{\text{sm}}_{d,z|x} \equiv \mathcal{U}^{\text{sm}}_{d,z|x}(g)$, i.e. as functions of the unknown thresholds. A policy $\delta$ can then be equivalently captured by the thresholds $g + \delta \equiv \{g_{j,z|x} + \delta_{j,z|x} : j = 1, \ldots, J,~z \in \mathcal{Z},~x \in \mathcal{X}\}$, i.e. changing the original thresholds by some known values $\delta \equiv \{\delta_{j,z|x} : j = 1, \ldots, J,~z \in \mathcal{Z},~x \in \mathcal{X}\}$. Let $D^{\delta} = \sum_{d \in \mathcal{D}}d 1\{ U \in \mathcal{U}^{\text{sm}}_{d,Z|X}(g+\delta)\}$ and $Y^{\delta} = Y(D^{\delta})$ denote the individual's treatment and outcome under this policy, respectively. The policy relevant treatment effect (PRTE) between two policies $\delta'$ and $\delta''$ can then be defined by $\text{PRTE}^{\delta',\delta''} = E[Y^{\delta'} - Y^{\delta''}|D^{\delta'} \neq D^{\delta''}]$, i.e. the effect of the counterfactual change in outcomes between the two altered values of thresholds for those who are affected by this change.

Identified Set

The objective of our analysis is to learn about a pre-specified parameter of interest $\theta(\beta)$ given by (ref). As the function $\theta$ is known, what we can learn about the parameter depends on what we know about the primitives $\beta$. To this end, denoting by $\mathbf{M}^{\dagger}$ and $\mathbf{H}^{\dagger}$ the space of all possible $m$ and $h$, let $\mathbf{B} \subseteq \mathbf{M}^{\dagger} \times \mathbf{H}^{\dagger}$ be the admissible space that $\beta$ is restricted to lie in, which is determined by the various assumptions we may impose on it. The data also provide information on $\beta$. The observed decision and outcome through (ref) and (ref) along with the selection model in (ref) and Assumption (ref) provide the following moments

align[align omitted — 262 chars of source]

for $d \in \mathcal{D}$, $z \in \mathcal{Z}$ and $x \in \mathcal{X}$. Denoting by $\mathbf{B}^* = \{\beta \in \mathbf{B} : \beta \text{ satisfies } \eqref{eq:D_moments}-\eqref{eq:Y_moments}\}$ the admissible set of primitives satisfying the assumptions and the moments, what we can learn about the parameter of interest can then be formally captured by the identified set defined by

align[align omitted — 171 chars of source]

i.e. the image of the admissible set $\mathbf{B}^*$ under the function $\theta$.

Identification

We start by showing how to compute the identified set when the moments of the data in the left hand sides of (ref) and (ref) are assumed to be known without uncertainty. The subsequent section shows how this analysis translates into a strategy to estimate the identified set and construct confidence intervals.

Two-Step Approach

From the definition in (ref), observe that computing the identified set requires searching through the space $\mathbf{B}^*$ and taking the image under the function $\theta$. The challenge is how to perform this search in a tractable manner.

To do so, we build on the approach of mogstad/etal:18, who study this identification problem in the special case of a binary treatment when selection is given by (ref). In this case, because a single dimension of unobserved heterogeneity governs selection and is normalized to be uniformly distributed on $[0,1]$, the distribution $F$ is known. Given $F$ is known, $h$ is then point-identified as the threshold function $g$ is directly point-identified by the treatment shares via (ref). This allows simplifying the search problem as it remains to search only across the admissible values of the MTRs. Specifically, for a given $h \in \mathbf{H}^\dagger$, let $\mathbf{M}(h) \equiv \{m \in \mathbf{M}^\dagger : (m,h) \in \mathbf{B}\}$ denote the set of admissible MTRs and let $\mathbf{M}^*(h) = \{m \in \mathbf{M}(h) : (m,h) \text{ satisfies } \eqref{eq:Y_moments} \}$ denote those additionally consistent with the outcome moments. The identified set then simplifies to $\theta(\mathbf{M}^*(h^*),h^*) \equiv \{\theta(m,h^*) : m \in \mathbf{M}^*(h^*) \}$, where $h^*$ denotes the point identified value of $h$. Assuming $\mathbf{M}(h)$ takes $m$ to be linearly parameterized, mogstad/etal:18 exploit the linearity of $\theta$ and $\mathbf{M}^*(h)$ in terms of this parametrization to show that $\theta(\mathbf{M}^*(h^*),h^*) = [\theta_L(h^*),\theta_U(h^*)]$, where

align[align omitted — 169 chars of source]

and where these optimization problems correspond to linear programs.

However, in the case where multidimensional unobserved heterogeneity governs selection, an inherent challenge is the fact that $h$ is not generally point identified. This is because, while the marginals continue to be normalized to uniform distributions, the dependence between them is not necessarily known, implying that both $F$ and $g$ may not be known or point identified by the data. We propose a two-step solution to account for this complication. Specifically, let $\mathbf{H}$ denote the set of admissible values of $h$ and $\mathbf{H}^* = \{h \in \mathbf{H} : h \text{ satisfies } \eqref{eq:D_moments}\}$ denote those additionally consistent with the selection moments. The identified set in (ref) can then be stated as

align[align omitted — 101 chars of source]

i.e. a union of the identified sets given $h$ across the admissible values of $h$ consistent with the selection moments. This suggests a two-step strategy, where in an inner step, as in the unidimensional case, we can continue to exploit the linearity to compute $\theta(\mathbf{M}^*(h),h)$ given $h \in \mathbf{H}^*$, and then in an outer step, we can aggregate them across $\mathbf{H}^*$ to obtain the final identified set.

In what follows, we show how to operationalize this strategy. In Sections (ref) and (ref), we first show how to generalize the arguments from the unidimensional case to that with multidimensional MTRs to tractably compute $\theta(\mathbf{M}^*(h),h)$ as in (ref) for each $h \in \mathbf{H}$ when $\mathbf{M}(h)$ takes $m$ to be parameterized or remain entirely nonparametric. In Section (ref), we then show how to parametrize $\mathbf{H}^*$ by a finite-dimensional parameter so that we can feasibly take the union across $\mathbf{H}^*$.

Identified Set under Parametric MTRs

As in mogstad/etal:18, we assume $\mathbf{M}(h)$ to linearly parametrize $m$ along with a parameter space determined by a system of linear inequalities as follows.

\begin{assumption2}{PM}{(Parametrizing MTRs)} \sloppy $\mathbf{M}(h) = \bigl\{ m \in \mathbf{M}^{\dagger} : m_{d|x}(u) = \sum_{k=0}^{K} \gamma_{k,d|x} b_{k}(h)(u) \text{ for } d \in \mathcal{D},~x \in \mathcal{X}, \text{ and } \gamma \equiv (\gamma_{k,d|x} : 0 \leq k \leq K,~d \in \mathcal{D},~x \in \mathcal{X}) \in \mathbf{\Gamma}(h) \bigr\}$, where $\{b_{k}(h) : 0 \leq k \leq K \}$ are known functions and $\mathbf{\Gamma}(h) \subseteq \mathbf{R}^{\text{dim}(\gamma)}$ is characterized by a system of linear inequalities. \end{assumption2}

To see how this allows computing $\theta(\mathbf{M}^*(h),h)$ using linear programs as in (ref) for each $h \in \mathbf{H}$, it is useful to rewrite it in terms of the parametrizing variable $\gamma$. As $\theta$ is linear in $m$ and $m$ is linear in $\gamma$, we can substitute in the relation between $m$ and $\gamma$ from Assumption (ref) into $\theta$ to obtain a function $\theta^{\Gamma}$ linear in $\gamma$ such that $\theta(m,h) = \theta^{\Gamma}(\gamma,h)$. Similarly, rewriting the outcome moments in (ref) in terms of $\gamma$

align[align omitted — 150 chars of source]

for $d \in \mathcal{D}$, $z \in \mathcal{Z}$, $x \in \mathcal{X}$, we can write $\mathbf{M}^*(h)$ in terms of $\gamma$ by $\mathbf{\Gamma}^*(h) = \left\{ \gamma \in \mathbf{\Gamma}(h) : \gamma \text{ satisfies } \eqref{eq:Y_moments_alpha} \right\}$. The identified set in terms of $\gamma$ then corresponds to $\theta^{\Gamma}(\mathbf{\Gamma}^*(h),h) \equiv \{\theta^{\Gamma}(\gamma,h) : \gamma \in \mathbf{\Gamma}^*(h)\}$. In the following proposition, we show that the linearity implies that $\theta^{\Gamma}(\mathbf{\Gamma}^*(h),h)$ for $h \in \mathbf{H}$ equals an interval with endpoints given by two optimization problems, which correspond to linear programs as $\theta^{\Gamma}$ is a linear function in $\gamma$ and $\mathbf{\Gamma}^*(h)$ is given by a system of linear inequalities.

propositionFor $h \in \mathbf{H}$, let $\mathbf{M}(h)$ satisfy Assumption (ref). If $\mathbf{M}^*(h)$ is empty then so is $\theta(\mathbf{M}^*(h),h)$, and if it is not empty, then $\text{closure}(\theta(\mathbf{M}^*(h),h)) = [\theta_L(h),\theta_U(h)]$, where \begin{align} \theta_L(h) = \inf_{\gamma \in \mathbf{\Gamma}^*(h)} \theta^{\Gamma}(\gamma,h) and \theta_U(h) = \sup_{\gamma \in \mathbf{\Gamma}^*(h)} \theta^{\Gamma}(\gamma,h) . \end{align}

Assumption (ref) requires the restrictions on $\gamma$ to correspond to a system of linear inequalities. As highlighted in Table (ref), a number of common shape restrictions from the literature correspond to linear restrictions on $m$ and in turn to a system of linear inequalities on $\gamma$ when considered over a finite grid of points of $u$ in $[0,1]^J$. Specifically, it allows assumptions such as boundedness typically required to ensure the bounds on mean effects are finite manski:08, monotonicity in outcomes between two treatments as in the monotone treatment response assumption from manski:97, or the difference in mean outcomes between two treatments to be bounded as in the bounded variation assumption from manski/pepper:18. It also allows imposing monotonicity across different values of the unobserved heterogeneity in a given dimension to capture how the outcome may vary for those more or less likely to select into treatment as in the monotone treatment selection assumption from manski/pepper:00, or separability between the unobserved heterogeneity and the covariates as in brinch/etal:17.

Identified Set under Nonparametric MTRs

Although Proposition (ref) parameterizes the MTRs via Assumption (ref), it allows the researcher to compute the identified set under fully nonparametric MTRs in special cases. In particular, let $\mathbf{M}_{\text{np}}(h)$ denote a set of nonparametric MTRs. For certain such sets formalized below, we can show $\theta(\mathbf{M}^*_{\text{np}}(h),h) = \theta(\mathbf{M}^*(h),h)$, where $\mathbf{M}(h)$ denotes the subset of $\mathbf{M}_{\text{np}}(h)$ additionally satisfying Assumption (ref) with $\{b_{k}(h) : 0 \leq k \leq K\}$ given by

align[align omitted — 75 chars of source]

for a specific choice of partition $\mathbb{U}(h) \equiv \{\mathcal{U}_{0}(h), \mathcal{U}_{1}(h), \ldots, \mathcal{U}_{K}(h)\}$ of $[0,1]^J$, and with

align[align omitted — 247 chars of source]

which can be equivalently written as a system of linear inequalities. This mirrors mogstad/etal:18, which shows such a result in the case with a single dimension of unobserved heterogeneity. In this case, the rectangular sets underlying the selection model and parameters in (ref) and (ref) simplify to just intervals. Exploiting this interval structure, they show that a partition consisting of intervals constructed using the end points of those in (ref) and (ref) can generate the nonparametric bounds.

We construct a new partition that shows how to exploit the general rectangular structure of (ref) and (ref), and generalize the unidimensional partition to the case with multidimensional unobserved heterogeneity. Let $\mathbb{U}^{*}(h) = \{\mathcal{U}^{\text{sm}}_{d,z|x}(g) : d \in \mathcal{D},~ z \in \mathcal{Z},~x \in \mathcal{X} \} \cup \{\mathcal{U}^{\theta}_{d,l|x}(g) : d \in \mathcal{D},~ l \in \mathcal{L},~x \in \mathcal{X}\}$ denote the collection of sets underlying the data moments in (ref) and the parameter of interest in (ref). We require that the partition is rich enough to generate each set in this collection so that it can capture the information provided by the data and on the parameter of interest and hence produce the nonparametric bounds. We formalize this property in the following definition.\footnote{We note that our definition shares conceptual similarities to multidimensional partitions considered in discrete choice analysis to obtain nonparametric bounds in alternative identification problems, where one similarly requires the partition to be sufficiently rich so that each element of the partition predicts the same choice behavior chesher/etal:13, gu/etal:22, kamat/norris:22, tebaldi/etal:23.}

\begin{definition2}{P}{(Partition)} Let $\mathbb{U}(h)$ be a finite partition of $[0,1]^J$ such that for each $\mathcal{U}^* \in \mathbb{U}^{*}(h)$, we have that $\bigcup_{\mathcal{U} \in \mathbb{U}} \mathcal{U} = \mathcal{U}^*$ for some $\mathbb{U} \subseteq \mathbb{U}(h)$. \end{definition2}

Figure (ref) graphically illustrates how we can exploit the rectangular structure underlying the selection model and parameters in (ref) and (ref) to construct a partition satisfying Definition (ref) in the context of a simple version of Example (ref). Figure (ref)(a) first presents the sets in $\mathbb{U}^{*}(h)$. Figure (ref)(b) then shows how to exploit the rectangular structure of these sets to construct a partition satisfying Definition (ref). It takes the endpoints of the various rectangles to form disjoint rectangles that partition the entire space.

figure[figure omitted — 852 chars of source]

This idea can also be applied to construct a partition more generally. Specifically, for each $j=1,\ldots,J$, let $u_{(1),j} \leq \ldots \leq u_{(M_j),j}$ denote the ordered values of all the end points in the $j$th dimension of the underlying rectangular sets in (ref) and (ref) that form the sets in $\mathbb{U}^{*}(h)$ together with 0 and 1. Let $\mathbb{U}(h)$ then denote the partition obtained by

align[align omitted — 154 chars of source]

i.e. a collection of hyperrectangles constructed by taking the Cartesian product of the set of intervals based on adjacent endpoints in each dimension.

By construction, due to Definition (ref), the partition implies that for each $m \in \mathbf{M}^\dagger$ satisfying (ref), the parametrized version $m' = \sum_{k=0}^{K} \gamma_{k,d|x} b_{k}(h)$, where $b_k(h)$ satisfies (ref) and $\gamma_{k,d|x} = E[m_{d|x}(U) | U \in \mathcal{U}_k(h)]$, also satisfies (ref) and also generates the same parameter of interest as $m$ in the sense that $\theta(m,h) = \theta(m',h)$. In turn, provided that

align[align omitted — 157 chars of source]

for $\mathbf{\Gamma}(h)$ in (ref) so that the parametric $m'$ also satisfies the assumptions imposed on $m$, it will generate the same identified set for $\theta$ as $m$. We show this in the following proposition.

propositionFor $h \in \mathbf{H}$ and $\mathbf{M}_{\text{np}}(h)$, let $\mathbf{M}(h) = \{m \in \mathbf{M}_{\text{np}}(h) : m = \sum_{k=0}^{K} \gamma_{k,d|x} b_{k}(h),~\gamma \equiv (\gamma_{k,d|x} : 0 \leq k \leq K,~d \in \mathcal{D},~x \in \mathcal{X}) \in \mathbf{\Gamma}(h)\}$ with $\{b_{k}(h) : 0 \leq k \leq K\}$ given by (ref), where $\mathbb{U}(h)$ satisfies Definition (ref), and with $\mathbf{\Gamma}(h)$ given by (ref). Suppose $\mathbf{\Gamma}(h)$ satisfies (ref) for every $m \in \mathbf{M}_{\text{np}}(h)$. It then follows that $\theta(\mathbf{M}_{\text{np}}^*(h),h) = \theta(\mathbf{M}^*(h),h)$, where $\theta(\mathbf{M}^*(h),h)$ is characterized in Proposition (ref).

Requiring $\mathbf{\Gamma}(h)$ to satisfy (ref) restricts the sets of nonparametric MTRs to which Proposition (ref) applies. However, amongst the linear restrictions in Table (ref), the condition in (ref) fails only when Assumption $\text{M}_j$ is considered. For the remaining restrictions, namely Assumptions B, $\text{M}_{d',d''}$, $\text{BV}_{d',d''}$, and $S$, we can straightforwardly show that for every $m$ satisfying each of these assumptions, $\gamma$ constructed using $m$, i.e. where $\gamma_{k,d|x} = E[m_{d|x}(U) | U \in \mathcal{U}_k(h)]$, satisfies the linear restrictions stated in terms of $\gamma$ given which $\mathbf{\Gamma}(h)$ satisfies (ref)---see Table (ref) for the derivations of these linear restrictions. For such restrictions imposed on the MTRs, nonparametric bounds on the parameter of interest can continue to be computed using Proposition (ref).

Parametrization of Selection Model Primitives

Having computed $\theta(\mathbf{M}^*(h))$ given $h \in \mathbf{H}$, we next show how to ensure that we can feasibly take their unions over $h \in \mathbf{H}^*$ and compute the identified set in (ref). To do so, we make the following assumption on $\mathbf{H}$.

\begin{assumption2}{PS}{(Parametrizing Selection Primitives)} $\mathbf{H} = \bigcup_{\lambda \in \mathbf{L}} \mathbf{H}(\lambda)$, where $\mathbf{L} \subseteq \mathbf{R}^{d_{\lambda}}$ is known, and $\mathbf{H}^*(\lambda) = \{h \in \mathbf{H}(\lambda) : h \text{ satisfies \eqref{eq:D_moments}}\}$ is a singleton or an empty set for each $\lambda \in \mathbf{L}$. \end{assumption2}

Recall in the unidimensional case $h$ is point-identified as $F$ is known and $g$ can be identified by the moments in (ref). Assumption (ref) aims to extend this idea to the general case with multidimensional unobserved heterogeneity by assuming $h$ can be parameterized by $\lambda \in \mathbf{L}$ such that an analogous argument applies for each $\lambda \in \mathbf{L}$. In particular, for each $\lambda \in \mathbf{L}$, it requires $\mathbf{H}(\lambda)$ and the data variation in (ref) to be such that there exists a unique value of $h$ consistent with the data so that it is point identified or that no such $h$ exists and the model is misspecified.

For the purposes of computation, Assumption (ref) along with Proposition (ref) allows us to compute the identified set using a two-step strategy performed across $\lambda \in \mathbf{L}$, which we can implement in practice using a grid of values in $\mathbf{L}$. To see this, observe that (ref) can be rewritten as

align[align omitted — 133 chars of source]

where $h^*(\lambda)$ is the point-identified value of $h$ when $\mathbf{H}^*(\lambda)$ is non-empty, $\theta_L$ and $\theta_U$ are defined as in (ref), and $\mathbf{L}^* = \{\lambda \in \mathbf{L} : \mathbf{H}^*(\lambda) \neq \emptyset,~\mathbf{\Gamma}^{*}(h^*(\lambda)) \neq \emptyset \}$ is the set of $\lambda$ for which the model is not misspecified. In turn, in an inner step for each $\lambda \in \mathbf{L}$, we can compute $\theta^{\Gamma}(\mathbf{\Gamma}^*(h^*(\lambda)),h^*(\lambda))$ by solving the linear programs in (ref), where, as elaborated in Section (ref), the point-identified value $h^*(\lambda)$ can be computed by solving a minimization problem based on the distance of the parameters to the data in the moments in (ref). In an outer step, we can then take the union of these sets when non-empty across $\lambda \in \mathbf{L}$ to obtain the identified set $\Theta$.

We conclude this section by revisiting the examples in Section (ref) and highlighting how Assumption (ref) can be naturally satisfied in them under common parametric assumptions on $F$.

example[continues=ex:arum] In this model, it is more convenient to work with the representation in (ref). We first show that $\tilde{g} \equiv (\tilde{g}_{d,z|x} : d \in \{1,2\},~z \in \mathcal{Z},~x \in \mathcal{X})$ can be point identified given a known $\tilde{F} \equiv (\tilde{F}_{x} : x \in \mathcal{X})$, where $\tilde{F}_x$ denotes the conditional on $X=x \in \mathcal{X}$ distribution of $(\tilde{U}_1,\tilde{U}_2)$. Observe that the non-redundant moments in (ref), where the moments for one treatment are dropped given that they sum to one, in terms of $\tilde{g}$ and $\tilde{F}$ can be written as \begin{align} P(D=1|Z=z,X=x) &= \int 1\{\tilde{u}_1 < \tilde{g}_{1,z|x}, \tilde{u}_2 - \tilde{u}_1 > \tilde{g}_{2,z|x} - \tilde{g}_{1,z|x}\}d\tilde{F}_x(\tilde{u}_1,\tilde{u}_2) , \\ P(D=2|Z=z,X=x) &= \int 1\{\tilde{u}_2 < \tilde{g}_{2,z|x}, \tilde{u}_1 - \tilde{u}_2 > \tilde{g}_{1,z|x} - \tilde{g}_{2,z|x}\}d\tilde{F}_x(\tilde{u}_1,\tilde{u}_2) , \end{align} for $z \in \mathcal{Z},~x \in \mathcal{X}$. The following proposition shows that if $\tilde{F}_x$ has a strictly positive density for $x \in \mathcal{X}$, then there exists at most a value of $\tilde{g}$ consistent with these moments. \begin{proposition} If $\tilde{F}_x$ is known, and has a strictly positive density on $\mathbf{R}^2$ for each $x \in \mathcal{X}$, then there exists at most a single value of $\tilde{g}$ that satisfies (ref) and (ref). \end{proposition} Let $\tilde{\mathbf{H}}(\lambda)$ denote the set of all values of $(\tilde{g},\tilde{F})$, where $\tilde{F}$ equals a known parametric distribution $\bar{F}(\lambda)$ with a strictly positive density and parametrized by $\lambda \in \mathbf{L}$, and let $\phi$ denote the relation between $(g,F)$ and $(\tilde{g},\tilde{F})$ implied by $U_1 = \tilde{F}_{1|x}(\tilde{U}_1)$, $U_2 = \tilde{F}_{2|x}(\tilde{U}_2)$ and $U_3 = \tilde{F}_{12|x}(\tilde{U}_1-\tilde{U}_2)$ such that $(g,F) = \phi((\tilde{g},\tilde{F}))$. For example, we can take $\bar{F}(\lambda)$ with $\lambda = (\lambda_x : x \in \mathcal{X}) \in \mathbf{L} = [-1,1]^{|\mathcal{X}|}$ such that \begin{align} \bar{F}_x(\lambda)(\tilde{u}_1,\tilde{u}_2) = \Phi_{\lambda_x}(\tilde{u}_1,\tilde{u}_2) , \end{align} where $\Phi_{\lambda_x}$ denotes the distribution function of a bivariate standard normal distribution with correlation coefficient $\lambda_x \in (-1,1)$ as in a probit model---see train:09 for a discussion of various other common distributional parameterizations. To ensure Assumption (ref), we can then take $\mathbf{H}(\lambda) = \phi(\tilde{\mathbf{H}}(\lambda))$ for $\lambda \in \mathbf{L}$. It follows from Proposition (ref) that $\mathbf{H}^*(\lambda)$ is a singleton for each $\lambda \in \mathbf{L}$ as $\tilde{F} = \bar{F}(\lambda)$ is known and $\tilde{g}$ is point identified by (ref)-(ref), given which both $F$ and $g$ are known and point identified by $(g,F) = \phi((\tilde{g},\tilde{F}))$. We note that in cases where there exist restrictions on $\tilde{g}$, such as those discussed in Section (ref), we can impose them through $\tilde{\mathbf{H}}(\lambda)$, and in which case we can potentially also identify some parameters of $\tilde{F}$ and reduce the dimension of $\lambda$---see Section (ref) for an example in the context of our empirical application.
example[continues=ex:sequential] Observe that the non-redundant moments in (ref) in this model can be written as \begin{align} P(D=0|Z=z,X=x) &= 1 - g_{1,z|x} , \\ P(D=1|Z=z,X=x) &= g_{1,z|x} - F_x(g_{1,z|x},g_{2,z|x}) , \end{align} for $z \in \mathcal{Z}$ and $x \in \mathcal{X}$. The following proposition shows that if $F_x$ is strictly increasing in one of the dimensions for $x \in \mathcal{X}$, then there exists a unique value of $g$ consistent with these moments. \begin{proposition} If $F_x(u_1,u_2)$ is known and strictly increasing in $u_2$ for $u_1 > 0$ for each $x \in \mathcal{X}$, and $g_{1,z|x} > 0$ for each $z \in \mathcal{Z}$ and $x \in \mathcal{X}$, then there exists at most a single value of $g$ that satisfies (ref) and (ref). \end{proposition} As $U$ has uniformly distributed marginals, let $\bar{F}(\lambda)$ denote a known parametric distribution for $F$, where $\bar{F}_x(\lambda)$ is a continuous copula that is strictly increasing in the second dimension and parametrized by $\lambda \in \mathbf{L}$. For example, we can take $\bar{F}(\lambda)$ with $\lambda = (\lambda_x : x \in \mathcal{X}) \in \mathbf{L} = [-1,1]^{|\mathcal{X}|}$ such that \begin{align} \bar{F}_x(\lambda)(u_1,u_2) = \Phi_{\lambda_x}(\Phi^{-1}(u_1),\Phi^{-1}(u_2)) , \end{align} where $\Phi$ additionally denotes the distribution function of a standard normal distribution, i.e. a bivariate normal copula with correlation parameter $\lambda_x \in [-1,1]$---see nelsen:07 for examples of various other families of copulas. To ensure Assumption (ref), we can then take $\mathbf{H}(\lambda)$ to equal the set of all values of $(g,F)$, where $F = \bar{F}(\lambda)$ and $g_{1,z|x} > 0$ for $z \in \mathcal{Z}$ and $x \in \mathcal{X}$. It directly follows from Proposition (ref) that $\mathbf{H}^*(\lambda)$ will be a singleton or empty for each $\lambda \in \mathbf{L}$. We note that a restriction such as $U_1 = U_2$, an empirically relevant case as highlighted in Section (ref), can be accommodated by further restricting $\bar{F}(\lambda)$ to satisfy this requirement. For example, in the context of (ref), this amounts to simply restricting $\lambda_x = 1$ for $x \in \mathcal{X}$.
example[continues=ex:dynamic] Observe that the non-redundant moments in (ref) in this model can be written as \begin{align} P(D = (0,1) | Z=z,X=x) &= g_{2,z|x} - F_x(g_{1,z|x},g_{2,z|x},1) \\ P(D = (1,0) | Z=z,X=x) &= g_{1,z|x} - F_x(g_{1,z|x},1,g_{3,z|x}) \\ P(D = (1,1) | Z=z,X=x) &= F_x(g_{1,z|x},1,g_{3,z|x}) \end{align} for $z \in \mathcal{Z}$ and $x \in \mathcal{X}$. The following proposition shows that if $F_x$ is strictly increasing in its third dimension and its derivative with respect to the second dimension is less than one, and $g_{1,z|x} \in (0,1)$ so that both nodes are possible in the second stage, then there exists a unique value of $g$ consistent with these moments. \begin{proposition} \sloppy If $F_x(u_1,1,u_3)$ is known and strictly increasing in $u_3$ for $u_1 > 0$ and $\partial F_x(u_1,u_2,1) / \partial u_2 < 1$ for $u_1 < 1$ for each $x \in \mathcal{X}$, and $g_{1,z|x} \in (0,1)$ for each $z \in \mathcal{Z}$ and $x \in \mathcal{X}$, then there exists at most a single value of $g$ that satisfies (ref)-(ref). \end{proposition} As $U$ has uniformly distributed marginals, let $\bar{F}(\lambda)$ denote a known parametric distribution for $F$, where $\bar{F}_x(\lambda)$ is a continuous copula that is strictly increasing in the third dimension for $u_1 > 0$ and with the property that $\partial F_x(u_1,u_2,1) / \partial u_2 < 1$ for $u_1 < 1$, and parametrized by $\lambda \in \mathbf{L}$. For example, similar to (ref) and (ref), we can take $\bar{F}(\lambda)$ with $\lambda = ((\lambda_{12|x},\lambda_{13|x},\lambda_{23|x}) : x \in \mathcal{X}) \in \mathbf{L} = [-1,1]^{3|\mathcal{X}|}$ such that \begin{align} \bar{F}_x(\lambda)(u_1,u_2,u_3) = \Phi_{(\lambda_{12|x},\lambda_{13|x},\lambda_{23|x})}(\Phi^{-1}(u_1),\Phi^{-1}(u_2),\Phi^{-1}(u_3)) , \end{align} where $\Phi_{(\lambda_{12|x},\lambda_{13|x},\lambda_{23|x})}$ denotes the distribution function of a trivariate standard normal distribution with correlation between $u_k$ and $u_l$ given by $\lambda_{kl} \in (-1,1)$, i.e. a multivariate normal copula. Indeed, it is increasing in the third dimension, and moreover, as \begin{align*} \frac{\partial}{\partial u_2} \Phi_{(\lambda_{12|x},\lambda_{13|x},\lambda_{23|x})}(\Phi^{-1}(u_1),\Phi^{-1}(u_2),\Phi^{-1}(1)) = \Phi\left( \frac{ \Phi^{-1}(u_1) - \lambda_{12|x} \Phi^{-1}(u_2) }{\sqrt{1 - \lambda^2_{12|x}}} \right) , \end{align*} we also have that it satisfies the property that $\partial F_x(u_1,u_2,1) / \partial u_2 < 1$ for $u_1 < 1$. To ensure Assumption (ref), we can then take $\mathbf{H}(\lambda)$ to equal the set of all values of $(g,F)$, where $F = \bar{F}(\lambda)$ and $g_{1,z|x} \in (0,1)$ for $z \in \mathcal{Z}$ and $x \in \mathcal{X}$. It directly follows from Proposition (ref) that $\mathbf{H}^*(\lambda)$ will be a singleton or empty for each $\lambda \in \mathbf{L}$. As in the above examples, we note that additional restrictions on $g$ and $U$ such as those discussed in Section (ref) can be accommodated through $\mathbf{H}(\lambda)$.
example[continues=ex:double] Observe that the non-redundant moments in (ref) in this model can be written as \begin{align} P(D=1|Z=z,X=x) &= F_x(g_{1,z|x},g_{2,z|x}) , \end{align} for $z \in \mathcal{Z}$ and $x \in \mathcal{X}$. Here, unlike Examples (ref) and (ref) above, there are two unknown thresholds, $g_{1,z|x}$ and $g_{2,z|x}$, but only a single moment for each $x \in \mathcal{X}$. In turn, even when $F_x$ is known, it is generally difficult to identify the remaining two unknowns using a single moment. As in lee/salanie:18, it is therefore useful to suppose that $\mathcal{Z} = \mathcal{Z}_1 \times \mathcal{Z}_2$ and $g$ to be restricted such that $g_{1,z|x} \equiv g_{1,z_1|x}$ and $g_{2,z|x} \equiv g_{2,z_2|x}$ for all $z = (z_1,z_2) \in \mathcal{Z}$ and $x \in \mathcal{X}$, i.e. there exist two instruments and an exclusion restriction imposing that each instrument affects only one of the thresholds, and that $g_{1,z_1'|x}$ is known for some $z_1' \in \mathcal{Z}_1$. The following proposition shows that if $F_x$ is assumed to be strictly increasing in both its dimensions then at most a single value of $g$ can be consistent with the moments. \begin{proposition} For each $x \in \mathcal{X}$, let $\mathcal{Z} = \mathcal{Z}_1 \times \mathcal{Z}_2$ be such that $g_{1,z|x} \equiv g_{1,z_1|x} \in (0,1)$ and $g_{2,z|x} \equiv g_{2,z_2|x} \in (0,1)$ for all $z = (z_1,z_2) \in \mathcal{Z}$, and $g_{1,z'_1|x}$ be known for some $z'_{1} \in \mathcal{Z}_1$. If $F_x$ is strictly increasing in both its dimensions when $u_1,u_2 > 0$ and for each $x \in \mathcal{X}$, then there exists at most a single value of $g$ that satisfies (ref). \end{proposition} Unlike Propositions (ref) and (ref) above, Proposition (ref) requires $g_{1,z'_1|x}$ to be known. This captures the fact that $g_{1,z|x}$ and $g_{2,z|x}$ can generally be point identified only up to a constant---see also lee/salanie:18 that show this holds even in the presence of continuous variation in the thresholds. For the purposes of Assumption (ref), for $\lambda_{1} = (\lambda_{1|x} : x \in \mathcal{X}) \in \mathbf{L}_1 = (0,1)^{|\mathcal{X}|}$, let $\mathbf{G}(\lambda_1)$ equal the set of all $g$ such that, for each $x \in \mathcal{X}$, $g_{1,z|x} \equiv g_{1,z_1|x} \in (0,1)$ and $g_{2,z|x} \equiv g_{2,z_2|x} \in (0,1)$ for $z = (z_1,z_2) \in \mathcal{Z}_1 \times \mathcal{Z}_2 \equiv \mathcal{Z}$, and $g_{1,z'_1|x} = \lambda_{1|x}$ for a given $z_1' \in \mathcal{Z}_1$, i.e. thresholds that takes each instrument to affect only one of the thresholds and where $g_{1,z_1'|x}$ is known and equal to $\lambda_{1|x}$ for a given $z_1' \in \mathcal{Z}_1$. Moreover, as $U$ has uniformly distributed marginals, let $\bar{F}(\lambda_2)$ denote a known parametric distribution for $F$, where $\bar{F}_x(\lambda_2)$ is a continuous copula that is strictly increasing in both dimensions and parametrized by $\lambda_2 \in \mathbf{L}_2$. To ensure Assumption (ref), we can then take $\mathbf{H}(\lambda) = \{(g,F) : g \in \mathbf{G}(\lambda_1),~F = \bar{F}(\lambda_2)\}$ for $\lambda = (\lambda_1,\lambda_2) \in \mathbf{L}_1 \times \mathbf{L}_2$. It directly follows from Proposition (ref) that $\mathbf{H}^*(\lambda)$ will be a singleton or empty for each $\lambda \in \mathbf{L}$. In practice, one could take $\bar{F}(\lambda_2)$ and $\mathbf{L}_2$ similar to that in (ref).

Estimation and Statistical Inference

We next show how the two-step characterization of the identified set in (ref) lends itself to a tractable strategy to estimate the identified set and construct confidence intervals when the moments of the data in (ref) and (ref) can only be estimated using sample data.

Estimation

To estimate the identified set, a natural starting point is the plug-in estimator that replaces $P_{d|z,x}$ and $E_{d|z,x}$ in (ref) and (ref) with their sample analogs. However, this approach can deliver an empty estimated set in practice, because the moments may not be exactly satisfied even when the true identified set is non-empty. Following mogstad/etal:18, we instead base our estimator on a criterion that measures how well different values of the underlying primitives match the observed moments and retains those that best match the moments.

To describe the estimation procedure, it is useful to first restate the identified set in terms of such criteria. For the purposes of the point-identified value $h^*(\lambda)$, let

align[align omitted — 91 chars of source]

subject to $\int_{\mathcal{U}_{d,z|x}^{\text{sm}}(g)} dF_{x} = \pi_{d|z,x}$ for $d \in \mathcal{D}$, $z \in \mathcal{Z}$, $x \in \mathcal{X}$, where $P$ and $\pi$ denote $(P_{d|z,x} : d \in \mathcal{D},~z \in \mathcal{Z},~x \in \mathcal{X})$ and $(\pi_{d|z,x} : d \in \mathcal{D},~z \in \mathcal{Z},~x \in \mathcal{X})$ in vector notation, and $\Omega_{S}$ is a positive definite weight matrix. This criterion measures the squared distance of the model implied moments to the data selection moments in (ref). As $Q_{\text{sel}}(h) = 0$ whenever $h$ matches these moments, $h^*(\lambda)$ for each $\lambda \in \mathbf{L}$ if it exists can then be defined by

align[align omitted — 118 chars of source]

For the criterion that captures the condition that $\mathbf{H}^*(\lambda)$ and $\mathbf{\Gamma}^*(h^*(\lambda))$ are non-empty, we first note that this can be equivalently stated as there existing $\gamma \in \mathbf{R}^{\text{dim}(\gamma)}$ such that

align[align omitted — 130 chars of source]

where $\mu(\lambda)$ and $C_2(\lambda)$ are estimable quantities that depend on the right hand sides of the moments in (ref) and (ref) evaluated at $h$ equal to $h^*(\lambda)$, and $C_1(\lambda)$, $C_3(\lambda)$ and $c(\lambda)$ are known quantities underlying the linear restrictions $\mathbf{\Gamma}^*(h^*(\lambda))$ and the condition that $h^*(\lambda) \in \mathbf{H}^*(\lambda)$. For brevity, we present the exact forms of these vectors and matrices in Appendix (ref). As in (ref), let

align[align omitted — 139 chars of source]

subject to $C_1(\lambda) \bar{\mu} + (C_1(\lambda) C_2(\lambda) + C_3(\lambda)) \gamma \leq c(\lambda)$, where $\Omega_M(\lambda)$ is a positive definite weight matrix, i.e. the minimum possible squared distance of the model implied moments to the data moments in (ref) and (ref) for a given $(\gamma,\lambda)$. Again, as $Q(\gamma,\lambda) = 0$ whenever $\gamma \in \mathbf{\Gamma}^*(h^*(\lambda))$ and $h^*(\lambda) \in \mathbf{H}^*(\lambda)$, $\mathbf{L}^*$ and $\mathbf{\Gamma}^*(h^*(\lambda))$ in (ref) can then be defined by

align[align omitted — 285 chars of source]

The estimated identified set $\widehat{\Theta}$ is obtained by substituting in the estimated versions of (ref), (ref) and (ref), where the criterions in their definitions are replaced by their sample analogs. In particular, the estimated version of $h^*(\lambda)$ in (ref) is given by

align[align omitted — 179 chars of source]

subject to $\int_{\mathcal{U}_{d,z|x}^{\text{sm}}(g)} dF_{x} = \pi_{d|z,x}$ for $d \in \mathcal{D}$, $z \in \mathcal{Z}$, $x \in \mathcal{X}$, where $\widehat{P}$ is the sample analog estimator of $P$ and $\widehat{\Omega}_{S}$ is a consistent estimator of $\Omega_{S}$. Similarly, the estimated value of $Q(\gamma,\lambda)$ used to obtain estimates of (ref) and (ref) is given by

align[align omitted — 183 chars of source]

subject to $C_1(\lambda) \bar{\mu} + (C_1(\lambda) \widehat{C}_2(\lambda) + C_3(\lambda)) \gamma \leq c(\lambda)$, where $\widehat{C}_2(\lambda)$ and $\widehat{\mu}(\lambda)$ are estimates of $C_2(\lambda)$ and $\mu(\lambda)$ obtained by replacing $P_{d|z,x}$, $E_{d|z,x}$ and $h^*(\lambda)$ in the expression provided in Appendix (ref) by their estimated analogs $\widehat{P}_{d|z,x}$, $\widehat{E}_{d|z,x}$ and $\widehat{h}^*(\lambda)$, and $\widehat{\Omega}_M(\lambda)$ is a consistent estimator of $\Omega_M(\lambda)$. However, simply using $\widehat{Q}$ to obtain the estimated versions of (ref) and (ref), and hence the identified set, can result in an empty set as $\min_{\gamma \in \mathbf{\Gamma}(\widehat{h}^*(\lambda))}\widehat{Q}(\gamma,\lambda)$ may not equal 0 for any $\lambda \in \mathbf{L}$. Following mogstad/etal:18, we therefore instead take the estimated versions of (ref) and (ref) to be

align[align omitted — 370 chars of source]

where $Q^* = \min_{\lambda \in \mathbf{L}} \min_{\gamma \in \mathbf{\Gamma}(\widehat{h}^*(\lambda))}\widehat{Q}(\gamma,\lambda)$, i.e. the subsets of $\mathbf{L}$ and $\mathbf{\Gamma}(\widehat{h}^*(\lambda))$ that come closest to satisfying the conditions on the estimated criterions up to a tolerance $\kappa$. By construction, these sets, and hence the estimated identified set, are never empty. Here $\kappa$ is a tuning parameter that goes to 0 with sample size at a specific rate and is necessary to ensure that the estimated set is consistent. In cases where $\mathbf{L}$ is a singleton and there exists a unique value of $\gamma$ that ensures $Q(\gamma,h^*(\lambda))$ = 0 so that the model is point identified, we can take $\kappa$ to equal 0.

The above procedure can be summarized by the following algorithm, which can in practice be implemented using a grid of values in $\mathbf{L}$:

tabularx[tabularx omitted — 860 chars of source]

In terms of computation, Step 1 is potentially the most expensive step as (ref) can be a nonlinear optimization problem depending on how $h$ enters the problem. For the parametric distributions considered for the examples in Section (ref), these nonlinear problems are similar to those encountered in the estimation of parametric discrete choice problems train:09. In contrast, as in the computations in (ref), the linearity of $\mathbf{\Gamma}$ and $\theta^{\Gamma}$ and the quadraticity of the sample analog $\widehat{Q}$ in (ref) in terms of $\gamma$ given $\lambda \in \mathbf{L}$ implies that Steps 2 and 4 are more structured and correspond to convex quadratic and quadratically-constrained quadratic programs in $(\bar{\mu},\gamma)$, respectively. In cases where we know the model is point identified and $\kappa$ is taken to be 0, Step 4 can be further simplified. In particular, let $\gamma^* = \text{arg}\min_{\gamma \in \mathbf{\Gamma}(\widehat{h}^*(\lambda))}\widehat{Q}(\gamma,\lambda)$ be the unique solution in Step 2 and the unique value in $\widehat{\mathbf{\Gamma}}^*(\widehat{h}^*(\lambda))$. In Step 4, we then simply have $\widehat{\theta}_L(\widehat{h}^*(\lambda)) = \widehat{\theta}_U(\widehat{h}^*(\lambda)) = \theta^{\Gamma}(\gamma^*,\widehat{h}^*(\lambda))$.

Statistical Inference

Confidence intervals can be constructed by test inversion, i.e. test at level $\alpha \in (0,1)$ the null hypothesis $H_0 : \theta_0 \in \Theta$ and then construct $100(1-\alpha)\%$-confidence intervals by collecting the values of $\theta_0$ that are not rejected. To do so in a computationally tractable manner, we again exploit, as in our estimation procedure, the fact the $h$ is point identified by $h^*(\lambda)$ in (ref) and the linearity of $\theta^{\Gamma}$ and $\mathbf{\Gamma}^*(h^*(\lambda))$ in $\gamma$ to test $H_{0}(\lambda): \theta_0 \in \theta^{\Gamma}(\mathbf{\Gamma}^*(h^*(\lambda)),h^*(\lambda)),~ h^*(\lambda) \in \mathbf{H}^*(\lambda)$. We then show how to aggregate these tests across $\lambda \in \mathbf{L}$.

For each $\lambda \in \mathbf{L}$, we exploit the observation that, similar to (ref), the null hypothesis can be reformulated as

align[align omitted — 198 chars of source]

i.e. testing whether $\gamma$ satisfies a linear system of inequalities, where $\mu(\lambda)$, $C_1(\lambda)$, $C_2(\lambda)$, $C_3(\lambda)$ and $c(\lambda)$ are specified as in (ref), but augmented to include the additional condition that $\theta^{\Gamma}(\gamma,h^*(\lambda)) = \theta_0$---see Appendix (ref) for their exact forms. This formulation allows us to use the procedure from cox/etal:24, who show how to test null hypotheses that can be stated as (ref). Similar to the criterion in (ref), it is based on a Quasi-Likelihood Ratio test statistic given by

align[align omitted — 167 chars of source]

subject to $C_1(\lambda) \bar{\mu} + (C_1(\lambda) \widehat{C}_2(\lambda) + C_3(\lambda)) \gamma \leq c(\lambda)$, where $\widehat{\mu}(\lambda)$ and $\widehat{C}_2(\lambda)$ denote estimates of $\mu(\lambda)$ and $C_2(\lambda)$, and $\widehat{\Sigma}(\lambda)$ denotes an estimate of the asymptotic variance of $(\widehat{\mu}(\lambda) - \mu(\lambda) + (\widehat{C}_2(\lambda) - C_2(\lambda))\gamma^*)$ with

align[align omitted — 81 chars of source]

subject to $C_1(\lambda) \mu(\lambda) + (C_1(\lambda) C_2(\lambda) + C_3(\lambda)) \gamma \leq c(\lambda)$. To compute the critical value, let $\widehat{\bar{\mu}}$ and $\widehat{\gamma}$ denote the minimizers of the problem in (ref), and let $\widehat{K}$ denote the set of row indices corresponding to the binding inequalities in $C_1(\lambda) \widehat{\bar{\mu}} + (C_1(\lambda) \widehat{C}_2(\lambda) + C_3(\lambda)) \widehat{\gamma} \leq c(\lambda)$. Moreover, let

align[align omitted — 182 chars of source]

where $[C_1(\lambda),C_3(\lambda)]_{\widehat{K}}$ and $(C_1(\lambda) \widehat{C}_2(\lambda) + C_3(\lambda))_{\widehat{K}}$ denote the submatrices of $[C_1(\lambda),C_3(\lambda)]$ and $(C_1(\lambda) \widehat{C}_2(\lambda) + C_3(\lambda))$ containing only the rows in $\widehat{K}$. The critical value $cv(\lambda,1-\alpha)$ is given by the $(1-\alpha)$-quantile of a chi-squared distribution with $\widehat{r}$ degrees of freedom. The test then equals $\phi(\lambda) = 1\{TS(\lambda) > cv(\lambda,1-\alpha)\}$, i.e. we reject if the test statistic is above the critical value and do not reject otherwise. Provided that $\widehat{\mu}(\lambda)$ is asymptotically normal and under weak assumptions on the consistency of $\widehat{C}_2(\lambda)$, cox/etal:24 show that this test controls size.

To aggregate these tests across $\lambda \in \mathbf{L}$ and test $H_0$, we exploit the principle underlying the so-called recycling test from bugni/etal:15. Let $\mathbf{L}_0 = \{\lambda \in \mathbf{L} : H_0(\lambda) \text{ is true} \}$ denote the set of $\lambda$ where $H_0(\lambda)$ is satisfied under $H_0$ and, similar to (ref)-(ref), let $\widehat{\mathbf{L}}_0 = \{\lambda \in \mathbf{L} : TS(\lambda) \leq \min_{\lambda \in \mathbf{L}}TS(\lambda) + \kappa \}$ denote a consistent estimate of this set, where $\kappa$ is again a tuning parameter necessary to ensure that the estimated set is consistent. The aggregated test then equals

align[align omitted — 147 chars of source]

i.e. we reject if the minimum test statistic across $\mathbf{L}$ is greater than the minimum critical value across $\widehat{\mathbf{L}}_0$. In Appendix (ref), we show that if $\phi(\lambda)$ controls asymptotic size for each $\lambda \in \mathbf{L}_0$ and $\widehat{\mathbf{L}}_0$ is a consistent estimate for $\mathbf{L}_0$, then the test in (ref) also controls asymptotic size.

As in our estimation procedure, we can summarize the test for a pre-specified choice of $\alpha \in (0,1)$ by the following algorithm:

tabularx[tabularx omitted — 576 chars of source]

For Step 1, we require, as in (ref), replacing $P_{d|z,x}$, $E_{d|z,x}$ and $h^*(\lambda)$ in the expression of $\mu(\lambda)$ and $C_1(\lambda)$ presented in Appendix (ref) by their estimated analogs $\widehat{P}_{d|z,x}$, $\widehat{E}_{d|z,x}$ and $\widehat{h}^*(\lambda)$. As in Step 1 of the estimation algorithm, this step can be computationally intensive as computing $\widehat{h}^*(\lambda)$ in (ref) can be a nonlinear problem. Step 2 is again more structured as the optimization problems in (ref) and (ref) are both convex quadratic problems. Given this algorithm, confidence intervals immediately follow by applying it over a fine grid of values of $\theta_0$ and then collecting the values where we do not reject. In practice, to speed up computation, this part can be easily parallelized across the different values of $\theta_0$.

Estimating the Returns to Head Start

In this section, we use our method to revisit the analysis of kline/walters:16, who evaluate Head Start (HS) using data from the Head Start Impact Study (HSIS).

Background

HS is a federally funded program in the United States that provides free preschool to three- and four-year-old children from disadvantaged families. The HSIS was a randomized experiment conducted in Fall 2002 to evaluate the program. Due to oversubscription, HS offers were randomized among applicants, enabling experimental evaluation by comparing outcomes of children who received an offer with those who did not puma_etal:10.

However, HS is not the only available preschool option; many alternatives provide similar services and also receive public subsidies. A non-trivial proportion of families in the experiment enrolled in such preschools. As highlighted in KW, this feature clouds the interpretation of the experimental effects as they estimate a combination of HS effects relative to enrolling in such preschools and no preschools. Moreover, it complicates a cost-benefit analysis of HS using these effects as enrollment out of these preschools may introduce fiscal savings and as these preschools may adjust the number of slots offered in response to that of HS.

Motivated by this feature, KW formulate a selection model that governs choice into the different preschool options and how preschools allocate offers. But, in estimating this model using the experimental data, they impose various assumptions such that the model is point-identified. In what follows, we illustrate how our method can be used to analyze the robustness of their conclusions on the returns to HS to these assumptions.

Setting and Data

We follow KW in mapping the experiment into the treatment effect setup in Section (ref). Let $Z \in \{0,1\}$ denote an indicator for whether the child was assigned to the treatment group and offered HS, $D \in \mathcal{D}$ and $Y$ denote their preschool choice and outcome of interest in their first year in the experiment, and $X$ denote covariates. As in KW, we take $\mathcal{D} = \{p,c,n\}$, where $p$ denotes HS program preschools, $c$ denotes competing non-HS preschools, and $n$ denotes no preschool; and the outcome of interest to be the average of the Woodcock Johnson III and Peabody Picture and Vocabulary Test scores, with each test score normed to have mean 0 and variance 1 in the control group by cohort. While KW consider a vector of covariates, we take $X$ to be a single binary covariate for simplicity as it is sufficient to present the main ideas. In particular, we take $X \in \{0,1\}$ to be an indicator for whether the child was in the three- or four-year-old cohort. Our analysis sample includes 3,627 children, all three- and four-year-olds in the sample with complete data for all variables mentioned above.

table[table omitted — 1,902 chars of source]

Table (ref) reports several motivating statistics. Panel (a) presents intention-to-treat (ITT) effects on enrollment into HS and outcomes defined by

align*[align* omitted — 217 chars of source]

and the instrumental variable (IV) estimand, defined as the average ratio of the two ITTs by cohort,

align[align omitted — 151 chars of source]

ITT$^Y$ reveals that the offer increased test scores, while the ITT$^D$ reveals noncompliance---not all children with an offer enrolled in HS and some without one did. The IV estimand accounts for this by estimating the effect of the offer for those who comply. We calculate a statistically significant effect of 0.230 standard deviations.\footnote{Our IV estimate is similar to the 0.247 standard deviation effect reported in KW. Minor differences reflect our simplified covariate specification and sample restrictions.}

As noted, evaluating the returns to HS using these experimental effects is complicated by the fact that many children without a HS offer enrolled in competing preschools. Panel (b) presents the proportion of children enrolled in the various care settings by offer status. Close to a third of the children without an offer enrolled in competing preschools, and the offer reduced this share. As the estimates below more precisely capture, this indicates that a non-trivial proportion of children induced into HS by the offer came from competing preschools.

Model of Preschool Choice and Provision

We consider the general multinomial selection model outlined in kline/walters:16 that allows for the possibility that competing preschools were also oversubscribed and randomized offers similar to HS. Recall that $Z_p \equiv Z \in \{0,1\}$ denotes the observed indicator for whether an offer to HS was randomly assigned. Moreover, as in KW, we introduce $Z_c$ to denote an unobserved indicator for whether an offer to a competing preschool was randomly assigned. Let $\delta_{p|x} \equiv E[Z_p | X{=}x]$ and $\delta_{c|x} \equiv E[Z_c | X{=}x]$ denote the offer probabilities with which HS and competing preschools are rationed, respectively, conditional on the cohort $x \in \{0,1\}$.

Following offers that are distributed by lottery in a first stage, families make choices in a second stage, where those who do not receive an offer from a preschool can enroll there by exerting additional effort. Let $V_d(z_p,z_c)$ define the family's utility from alternative $d \in \mathcal{D}$ had their offer statuses from HS and competing preschools been set to $(z_p, z_c) \in \{0,1\}^2$. The potential preschool choice given offer statuses equals the utility-maximizing choice given by

align[align omitted — 86 chars of source]

and the observed choice is simply the utility-maximizing one given their realized offers, i.e. $D = D(Z_p,Z_c)$. As in KW, we take $V_n(z_p,z_c) = 0$, i.e. normalize the utility of no preschools to 0, and impose the restriction that

align[align omitted — 132 chars of source]

where $\tilde{U}_d$ is the alternative-specific unobserved heterogeneity and $\zeta_x > 0$, i.e. an offer to a preschool only increases the utility for that preschool and the effect of an offer on utilities is the same for both types of preschool.\footnote{We find that the estimated $\zeta_x$ is greater than 0 and hence that the assumption $\zeta_x > 0$ is satisfied by the model. However, we note that we explicitly make this relevance-type assumption as we require $\zeta_x \neq 0$ for the selection model to be identified—see Proposition (ref) below.} Moreover, following KW, we assume the distribution on the unobserved heterogeneity conditional on the cohort $x \in \{0,1\}$ to be $\Phi_{\rho_x}$, i.e. a standard bivariate normal distribution with a correlation parameter $\rho_x$ that varies by cohort.\footnote{The restriction on $V_p(z_p,z_c)$ and the distributional assumption on $(\tilde{U}_p,\tilde{U}_c)$ are stated in kline/walters:16. The restriction on $V_c(z_p,z_c)$ follows from the discussion in kline/walters:16, where they note when using their selection model to compute the proportion induced into competing preschools from an offer they assume the utility value from an offer is comparable to that of a HS offer. }

For the purposes of our method, recall that the selection model is required to satisfy Assumption (ref), i.e. it needs to be point-identified up to a finite-dimensional parameter. To this end, the above model corresponds to that in Example (ref) for $D(z_p,z_c)$, where the thresholds $\tilde{g} = (\tilde{g}_{d,z_p,z_c|x} : d \in \{p,c\},~z_p,z_c, x \in \{0,1\})$ satisfy the restrictions in (ref) and the distribution of the unobservables $(\tilde{U}_p,\tilde{U}_c)$ corresponds to $\Phi_{\rho_x}$, which can be captured by taking $\tilde{g}$ and the distribution of unobservables $\tilde{F} = (\tilde{F}_x : x \in \mathcal{X})$ to satisfy

align[align omitted — 382 chars of source]

for some values $(\zeta_{p|x},\zeta_{c|x},\zeta_{x},\rho_x) \in \mathbf{R}^2 \times \mathbf{R}_{++} \times (-1,1)$ for each $x \in \mathcal{X}$. However, unlike showing how Assumption (ref) can be satisfied in Example (ref) in Section (ref), a complication here is that one of the instruments, namely $Z_c$, is not observed. Nonetheless, given the additional restrictions that reduce the number of unknowns characterizing the thresholds, we can identify $(\zeta_{p|x},\zeta_{c|x},\zeta_{x},\rho_x)$ for a fixed value of the competing preschool offer probability $\delta_{c|x}$. In particular, instead of (ref)-(ref), note that the selection moments under the restrictions in (ref)-(ref) and accounting for the fact that $Z_c$ is not observed correspond to

align[align omitted — 599 chars of source]

\sloppy for $z_p,x \in \{0,1\}$, where $P_{1,c|x} \equiv \delta_{c|x}$ and $P_{0,c|x} \equiv 1 - \delta_{c|x}$, $\sigma_x = (\zeta_{p|x},\zeta_{c|x},\zeta_x,\rho_{x})$, and $H_{c,z_p}$ and $H_{p,z_p}$ capture the right hand side of these moments in terms of the parameters underlying the restrictions in (ref)-(ref). The following proposition states the identification result given a rank condition on $H = (H_{c,0},H_{c,1},H_{p,0},H_{p,1})$.

propositionSuppose $J(\sigma,\delta_{c}) = \partial H(\sigma,\delta_{c}) / \partial \sigma$ has full rank for all $(\sigma,\delta_c) \in \mathbf{R}^2 \times \mathbf{R}_{++} \times (-1,1) \times [0,1]$. Then, for each $\delta_{c|x} \in [0,1]$, there exists at most a single value of $\sigma_x$ that satisfies (ref) and (ref).

Given the rank condition, we can ensure Assumption (ref) by taking $\mathbf{H}(\lambda) = \phi(\tilde{\mathbf{H}})$ for $\lambda = (\delta_{c|0},\delta_{c|1}) \in [0,1]^2 = \mathbf{L}$, where $\tilde{\mathbf{H}}$ is the set of all values of $(\tilde{g},\tilde{F})$ that satisfy (ref)-(ref) and $\phi$ denotes the relation as defined in Example (ref) in Section (ref) such that $(g,F) = \phi((\tilde{g},\tilde{F}))$. It follows by Proposition (ref) that $\mathbf{H}^*(\lambda)$ is a singleton for each $\lambda \in \mathbf{L}$ as $(\tilde{g},\tilde{F})$ can be point identified by (ref)-(ref) taking $(\delta_{c|0},\delta_{c|1}) = \lambda$ given which $(g,F)$ can be identified by $\phi(\tilde{g}),\tilde{F}))$. While it is difficult to analytically show that the rank condition holds given the nonlinear structure of the Jacobian, we can verify it numerically in practice using a grid for the values of $(\sigma,\delta_c)$---see Appendix (ref) for details.\footnote{Existing results on analytically showing full rank of the Jacobian in cases with normally distributed errors are present only for the binary choice case, where one exploits related properties to that of the inverse Mills ratio chern/etal:23,chern/etal:24,han/vytlacil:17. Such results, however, do not straightforwardly extend to the multinomial case, and we hence leave it as an extension to future work.} Doing so we find that the Jacobian does indeed satisfy this condition.

In estimating the above model, KW additionally assume $\delta_{c|x} = 0$ for $x \in \mathcal{X}$.\footnote{While they do not explicitly make this assumption, we note that the special case of the model that they estimate in kline/walters:16 that does not introduce $Z_c$ is only logically consistent with the discussion of their general model when taking $\delta_{c|x} = 0$.} As formalized by the above proposition, this ensures that the selection model is point-identified. In the context of the above model, this restriction can be interpreted as having no lottery for competing preschools and hence eliminating the possibility that their slots are rationed in a manner similar to HS. However, as highlighted in kline/walters:16, it is reasonable to believe that such preschools do indeed ration slots as the majority of states do not have universal preschool mandates and preschools face relatively fixed budgets. As our method does not require such an assumption, we use it below to analyze the robustness of their conclusions to it by allowing $\delta_{c|x}$ to vary in $[0,1]$, in which case the selection model is not necessarily point-identified.

Decomposing IV

We focus on two sets of parameters that underlie the main conclusions in KW. The first set of parameters decompose the returns to a HS offer evaluated by the IV estimand into those arising from children drawn into HS from competing preschools versus no preschool. Under our selection model, as $\zeta_x > 0$ for $x \in \mathcal{X}$, we can decompose the IV estimand in (ref) as follows

align[align omitted — 206 chars of source]

where

align[align omitted — 199 chars of source]

capture the proportion induced by a HS offer into HS from alternative $d \in \{n,c\}$ and the (local) average effect for this subgroup of children for cohort $x \in \{0,1\}$, respectively, and

align[align omitted — 216 chars of source]

capture this proportion averaged across the two cohorts and the effect weighted by the proportion of the subgroup across the two cohorts, respectively. Here the first line in the decomposition, as outlined in kline/walters:16, follows from the arguments of imbens/angrist:94 as $\zeta_x \geq 0$ implies the monotonicity condition that if $D(0,Z_c) \neq D(1,Z_c)$ then $D(1,Z_c) = p$, while the second line averages the expressions across the two cohorts.

A challenge is that while $S_{np}$ in the decomposition is identified by the data through

align[align omitted — 152 chars of source]

due to the monotonicity condition that also implies (ref), $LATE_{np}$ and $LATE_{cp}$ are not. KW exploit the point-identified version of the selection model discussed above along with additional assumptions on the Marginal Treatment Response (MTR) functions under this model to identify these effects. We now examine how their conclusions hold under weaker alternatives.

As in Section (ref), we define the MTRs here by $m_{d|x}(u_p,u_c) = E[Y(d)|U_p = u_p, U_c = u_c, X=x]$ for $d \in \mathcal{D}$, $x \in \mathcal{X}$, where $U_p = \tilde{F}_{\tilde{U}_p|X}(\tilde{U}_p)$ and $U_c = \tilde{F}_{\tilde{U}_c|X}(\tilde{U}_c)$ correspond to transformed versions of the unobserved heterogeneity as discussed in Example (ref). Note that we do not condition on $\tilde{F}_{\tilde{U}_p - \tilde{U}_c}(\tilde{U}_p - \tilde{U}_c)$ here as it is deterministically implied by $U_p$ and $U_c$. To ensure point identification, KW assume the MTRs to be linear in $u_p$ and $u_c$ and separable in $x$, i.e.

align[align omitted — 99 chars of source]

To see why this implies point identification, observe the outcome moments in (ref) are given by

align[align omitted — 171 chars of source]

for $d \in \mathcal{D}$, $z_p,x \in \{0,1\}$, where $\mathcal{U}^{\text{sm}}_{d,z_p,z_c}(g)$ is defined as in Example (ref), and $(g,F)$ are implied by the restrictions on $(\tilde{g},\tilde{F})$ in (ref)-(ref) and the relation between $(g,F)$ and $(\tilde{g},\tilde{F})$ noted in Section (ref). For each $d$, we have four moments in (ref) as $z_p$ and $x$ take two values. Moreover, given that KW assume $\delta_{c|x}$ for $x \in \mathcal{X}$ to be known, Proposition (ref) implies that there exist only four unknowns underlying these moments, namely $(\gamma_{d,1},\gamma_{d,2},\mu_{d|0},\mu_{d|1})$. As these moments are linear in the MTRs and the MTRs are linear in the unknowns, we can invert this linear system to identify them---and thus the LATEs.

To illustrate how our method can assess robustness, Table (ref) reports estimates and 95% confidence intervals for $LATE_{np}$ and $LATE_{cp}$ under KW's specification and a range of alternatives.\footnote{The estimates and confidence intervals are obtained using the algorithms described in Section (ref). In our implementation, we take a grid for $\mathbf{L}$ given by $\{0,0.05,\ldots,0.95,1\}^2$, and $\mathcal{U}_{\text{grid}} = \{0,0.25,0.5,0.75,1\}^2$ when imposing the shape restrictions in the parametric case in Table (ref). For the estimation algorithm, we take $\widehat{\Omega}_{S}$ to equal the identity matrix, and $\widehat{\Omega}_{M}$ to be computed as in $\widehat{\Sigma}^{-1}(\lambda)$ in (ref) using $\mu(\lambda)$, $C_1(\lambda)$, $C_2(\lambda)$, $C_3(\lambda)$ and $c(\lambda)$ from (ref). We take $\kappa = 0.1$ in both algorithms. To compute $\widehat{\Sigma}(\lambda)$, we use the bootstrap, i.e. we take $\widehat{\Sigma}(\lambda) = 1/B\sum_{b=1}^B (\widehat{\mu}_b(\lambda) - \widehat{\mu}(\lambda) + (\widehat{C}_{2,b}(\lambda) - \widehat{C}_{2}(\lambda))\widehat{\gamma}^*)(\widehat{\mu}_b(\lambda) - \widehat{\mu}(\lambda) + (\widehat{C}_{2,b}(\lambda) - \widehat{C}_{2}(\lambda))\widehat{\gamma}^*)'$, where $\widehat{\gamma}^*$ corresponds to the estimated version of (ref) and $\widehat{\mu}_b(\lambda)$ and $\widehat{C}_{2,b}(\lambda)$ are the bootstrap analogs of $\widehat{\mu}(\lambda)$ and $\widehat{C}_{2}(\lambda)$ estimated using the $b$th bootstrap drawn with replacement at the level of the HS centers and where we take $B=200$.} We do so by allowing flexibility across two separate dimensions: the selection model and the MTRs. For each parameter, the upper left estimate reflects our analog to the baseline model in KW, i.e. the specification in (ref) under the assumption that $\delta_{c|x} = 0$. Moving across the columns highlights flexibility in the specification of the MTRs by relaxing the linearity and separability assumptions. Moving down a row relaxes the selection model assumption that there is no rationing for competing preschools, and instead allows them to be rationed by an unknown $\delta_{c|x} \in [0,1]$.

table[table omitted — 6,651 chars of source]

Under our analog of KW’s baseline specification, the estimates for $LATE_{np}$ and $LATE_{cp}$ reveal that HS increases test scores for those induced from no preschool and competing preschools, respectively, by 0.27 and 0.14. Consistent with KW, however, both of these effects are statistically imprecise. KW considers a number of additional covariates and fixed effects to improve precision. We instead consider shape restrictions to improve precision chetverikov/etal:18. To this end, Column (2) imposes a bounded variation assumption as in Table (ref) given by $|m_{p|x}(u_p,u_c) - m_{n|x}(u_p,u_c)| \leq 0.8$, $|m_{c|x}(u_p,u_c) - m_{n|x}(u_p,u_c)| \leq 0.8$, and $|m_{p|x}(u_p,u_c) - m_{c|x}(u_p,u_c)| \leq 0.2 $. We choose the 0.8 value because this is close to the largest effect for a preschool program noted in the literature kline/walters:16. Similarly, we bound the difference in effects between HS and competing preschools to be at most 0.2, since the two options provide similar services and hence are likely to have more similar effects.

Focusing on $LATE_{np}$ first, we find that its estimate under this specification remains substantively unchanged, capturing that the estimated MTRs satisfy the additional restrictions. But, consistent with KW’s results under their more precise specifications, its confidence interval shrinks, resulting in a statistically significant effect. In Columns (3)-(7), we analyze the robustness of this positive effect to relaxing the linearity and separability assumptions in (ref). Column (3) allows for a more flexible polynomial specification by adding interactions of $u_p$ and $u_c$ and their squares to (ref), and Column (4) allows the MTRs to be entirely nonparametric. Columns (5)-(7) further relax these specifications by removing separability. The parameter in these cases is partially identified, but our estimated bounds as well as confidence intervals reveal that neither linearity nor separability are required to conclude that $LATE_{np}$ is positive.

The second row, which allows for rationing of competing preschools ($\delta_{c|x} \in [0,1]$), reveals similar robustness. In this case, the confidence intervals for $LATE_{np}$ in Columns (2)-(7) lie entirely above zero (with lower bounds around 0.04-0.11). Hence the positive effect when competing preschools may be rationed also remains statistically significant even under nonparametric and nonseparable MTRs.

Turning to $LATE_{cp}$, again consistent with KW, the confidence intervals across these specifications cannot rule out a null effect. To better understand this, Table (ref) also presents results for $S_{np}$. Here, as it is directly identified by the data moments in (ref), we obtain estimates by simply taking sample analogs and confidence intervals using a standard $t$-test. The estimate suggests that the null finding arises because the data may simply be worse-powered for this effect; roughly only 30% of those induced into HS come from competing preschools.

To analyze whether stronger restrictions can allow us to reach a more informative conclusion, we consider a final specification in Column (8) that imposes on Column (2) two additional assumptions: (i) a positive selection assumption, similar to the Assumption $M_j$ in Table (ref), that takes $|m_{d|x}(u_p,u_c) - m_{n|x}(u_p,u_c)|$ to be decreasing in $u_p$ and $u_c$ for $d \in \{c,n\}$, i.e. those more likely to enroll in a preschool have a higher effect; and (ii) a positive effect assumption as in $M_{d,n}$ in Table (ref) for $d \in \{p,c\}$, i.e. children have higher test scores under preschool relative to no preschool. Even under these restrictions, the confidence intervals for $LATE_{cp}$ remain unaffected and continue to cover the entire logical region from -0.2 to 0.2 permitted under the bounded variation assumption. This strengthens our takeaway that the data do not rule out the possibility of a null effect.

Policy Effects of Expanding Head Start

The second set of parameters we analyze evaluate the policy returns to marginal expansions of HS by marginally changing the probability of a HS offer, $\delta_{p|x}$. Following KW, we first define the benefits and costs of a HS offer. Let $B$ denote the total after-tax lifetime income for all the children, which relates to test scores by

align[align omitted — 71 chars of source]

where $\phi_{\text{ben}}$ denotes the pre-specified relation between test scores and future income and $\tau$ the pre-specified tax rate faced by the children of eligible households. The net costs to the government of financing preschool are given by

align[align omitted — 94 chars of source]

where $\phi_p$ and $\phi_c$ denote the pre-specified costs of providing HS and competing preschool services, respectively, and $\tau \phi_{\text{ben}} E[Y]$ captures the revenue generated by taxes on the adult earnings of HS-eligible children. Exploiting the fact that $E[Y]$ and $P(D=d)$ can be written in terms of $E[Y(D(z_p,z_c))|X=x]$, $P[D(z_p,z_c)=d|X=x]$ for $d\in\{p,c\}$ and $(\delta_{p|x},\delta_{c|x})$, we can differentiate these expressions with respect to $\delta_{p|x}$ for $x \in \{0,1\}$ to obtain the marginal benefits and costs of expanding access to HS---see Appendix (ref) for details. In doing so, following KW, we either assume that $\delta_{c|x}$ does not adjust to changes in $\delta_{p|x}$, which we refer to as the non-adjusting case kline/walters:16, or that $\delta_{c|x}$ adjusts such that enrollment in competing preschools stays constant, i.e. $d P(D = c|X=x) / d \delta_{p|x} = 0$, which we refer to as the adjusting case kline/walters:16. Performing this exercise, we can obtain expressions for the marginal benefits (MB) and costs (MC) under the non-adjusting case given by

align[align omitted — 308 chars of source]

and under the adjusting case given by

align[align omitted — 479 chars of source]

where $\Delta_{R,p|x} = E[R(1,Z_c) - R(0,Z_c)|X=x]$ and $\Delta_{R,c|x} = E[R(Z_p,1) - R(Z_p,0)|X=x]$ for random variable $R$, and $d \delta_{c|x} / d \delta_{p|x} = -\Delta_{1\{D=c\},p|x} / \Delta_{1\{D=c\},c|x}$. The marginal benefits and costs are then compared using the marginal value of public funds (MVPF) hendren:16 defined by

align[align omitted — 198 chars of source]

Table (ref) reports estimates and 95$\%$ confidence intervals for the MVPF parameter under the non-adjusting case ($\text{MVPF}_{\text{non-adj}}$) and the adjusting case ($\text{MVPF}_{\text{adj}}$), for the specifications considered in Table (ref).\footnote{In unreported results, we also consider an additional MVPF parameter for the adjusting case derived under the assumption that competing preschool offers only draw children from no preschool alternatives and that the effect for these children equals $LATE_{np}$. While this assumption is not logically consistent with the selection model in Section (ref), we consider it for completeness as it is presented in KW. As it produces qualitatively similar to $\text{MVPF}_{\text{adj}}$, we focus on the latter, more logically consistent parameter for brevity.} For the pre-specified values in the expression, following KW, we take $\phi_{\text{ben}} = \kappa_{\text{ben}} E$, where $E$ corresponds to the average present discounted value of lifetime earnings for HS applicants and $\kappa_{\text{ben}}$ is the percentage change in earnings induced by a one standard deviation increase in test scores. We take $\kappa_{\text{ben}} = 0.13$, which corresponds to the estimate found in chetty/etal:11---see kline/walters:16 for an overview of such estimates in the literature. Moreover, following KW, we take $\tau = 0.35$, $E=\$343,392$, $\phi_p = \$8,000$ and $\phi_c = 0.75 \cdot \phi_p = \$6,000$.

table[table omitted — 4,706 chars of source]

Observe that $\text{MVPF}_{\text{non-adj}}$ is directly point identified by moments of the data. This is because $\Delta_{R,p} = E[R|Z_p=1] - E[R|Z_p=0]$, which simply corresponds to the ITT effects on the variable $R$. Similar to the parameter $S_{np}$ in Table (ref), we estimate it by taking sample analogs of the various ITTs, with confidence intervals derived by inverting a standard $t$-test. Consistent with KW, the estimate of $\text{MVPF}_{\text{non-adj}}$ is greater than one and the lower bound of the confidence interval exceeds one, suggesting that, absent adjustment of competing preschools to HS expansion, marginally expanding HS access yields benefits that outweigh the costs.

For $\text{MVPF}_{\text{adj}}$, as in Table (ref), we obtain estimates and confidence intervals following the algorithms in Section (ref), but slightly adapted to account for the piecewise definition of the MVPF parameter in (ref)---see Appendix (ref) for details. Consistent again with KW, we find that under the linear and separable specification on the MTRs in Columns (1)-(2) and when $\delta_{c|x} = 0$, the estimates highlight that $\text{MVPF}_{\text{adj}}$ is also larger than one. Keeping $\delta_{c|x} = 0$, Columns (3), (5), (6), and (8) reveal this finding is robust to alternative parametric specifications, but Columns (4) and (7) show that the lower bound of the identified set falls below one when the MTRs are nonparametric.

When allowing for the possibility that competing preschools are rationed as in HS by allowing $\delta_{c|x} \in [0,1]$, Columns (1)-(7) show that the lower bound on $\text{MVPF}_{\text{adj}}$ falls below one across all specifications. Only under the additional restrictions of positive selection and positive effects in Column (8) does the lower bound of the identified set remain above one. Furthermore, across all specifications, the confidence intervals do not rule out the possibility that $\text{MVPF}_{\text{adj}}$ is smaller than one.

We conclude our analysis by analyzing whether alternative values for $\kappa_{\text{ben}}$, i.e. the relation between test scores and earnings, can potentially result in $\text{MVPF}_{\text{adj}}$ being statistically greater than one. In Figure (ref), we report $p$-values for the null hypothesis that $\text{MVPF}_{\text{adj}} \leq 1$ for alternative values of $\kappa_{\text{ben}}$ under Specifications (2) and (8) from Table (ref). Our results reveal that for $\delta_{c|x} \in [0,1]$ and under additional sufficiently strong restrictions on the MTRs, namely specification (8), we can statistically reject that $\text{MVPF}_{\text{adj}}$ is at most one at a 5$\%$ significance level if we are willing to assume that the percentage increase in earnings induced by a one standard deviation increase in test scores is close to 30$\%$. This value is near the upper bound of estimates found in the literature, comparable to the Perry Preschool Program heckman/etal:10.

figure[figure omitted — 243 chars of source]

Conclusion

In cases with multiple treatments, treatment selection models generally exhibit multidimensional unobserved heterogeneity. In this paper, we develop a marginal treatment effect based method to learn about treatment effects in a general class of such models with discrete-valued instruments. Allowing the selection model to be identified up to a finite-dimensional parameter, we show how a two-step computational program can be used to compute the identified set for the treatment effect parameters when the marginal treatment response functions underlying them remain nonparametric or are additionally parameterized. We demonstrate the benefits of our method by revisiting the empirical analysis of the Head Start program by kline/walters:16.