Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
72,366 characters · 17 sections · 49 citation commands
Identifying the Effects of a Program Offer with an Application to Head Start
\thispagestyle{empty}
KEYWORDS: Program offer, discrete choice, unobserved choice sets, instrumental variables, partial identification, randomized experiments, Head Start Impact Study.
JEL classification codes: C14, C25, C31, I21.
\setcounter{page}{1}
A central task of program evaluation is to evaluate policies relevant to a program. For many programs, the baseline policy corresponds to providing individuals with an offer to participate in the program. For example, in a preschool program, children may be provided with an offer to preschools in the program, while in a microfinance program, individuals may be provided with a loan offer. The evaluation task for such programs then translates to evaluating the effects of providing an offer. Indeed, evaluating the number of individuals who may take up the offer and the outcomes they later experience allows one to assess the costs and benefits of providing an offer and, in turn, the returns to the policy.
To analyze such policy relevant questions, the standard approach in the literature is to model how individuals select into the program and exploit the structure of the model to define parameters capturing the policy question of interest heckman/vytlacil:07. As highlighted in Section (ref), the typical model to do so corresponds to individuals choosing by maximizing utility over the various alternatives, where a single dimension of unobserved heterogeneity governs the utility under each alternative. However, as a program offer corresponds to providing the program in an individual's choice set, analyzing the effects of an offer naturally requires a richer model that distinguishes the role of choice sets and preferences in determining utility.
In this paper, I propose a novel treatment selection model that introduces unobserved heterogeneity in both choice sets and preferences to learn about the effects of providing a program offer. By distinguishing the role that choice sets and preferences play, the richer structure provides a building block to define a range of parameters evaluating the effects of a program offer that correspond to comparing participation decisions and outcomes under counterfactual choice sets with and without the program.
Having defined these parameters, I then show how to exploit the model structure to learn about them in the presence of instrumental variables that induce exogenous variation in choice sets. I allow the model to be nonparametric and the dependence between choices sets, preferences and outcomes to be unrestricted. These features imply that the parameters are generally partially identified. Deriving sharp analytical bounds for these parameters, however, can be difficult given the rich model structure and geometry of the parameters. In turn, by leveraging results from charnes/cooper:62, I illustrate how a linear programming procedure can be used to computationally characterize the identified sets. The procedure flexibly applies to the various parameters and allows a range of shape restrictions to be imposed on the baseline model.
As highlighted in Section (ref), the developed tools can be applied to a number of setups, specifically randomized experiments, where exogenous variation in choice sets naturally arises. For example, in the Head Start Impact Study (HSIS), families were randomly assigned to a treatment and control group, where those in the treatment group were more likely to be assigned a Head start preschool offer than those in the control group. Similarly, in a microfinance experiment in angelucci/etal:15, individuals in treatment villages were more likely to receive loan offer than those in control villages; while in the Oregon Health Field Experiment, those in the treatment group where more likely to receive Medicaid than those in the control group.
To empirically illustrate the tools, I apply it to evaluate the effects of providing a Head Start offer using data from the HSIS. Previous analyses have focused on the instrumental variable estimand that evaluates effects for the subgroup of compliers, and their variants that account for where these children would have enrolled in the control group feller/etal:16, kline/walters:16, puma_etal:10. As highlighted in Section (ref), the complier subpopulation corresponds to the policy relevant one when the policy of interest corresponds to the variation induced by the experiment. In turn, they leave open the question on the effects for the subgroup affected in a more general policy providing offers to Head Start.
The application reveals several conclusions. First, providing an offer to Head Start affects a large percentage of children who take up the offer and participate in Head Start, and can significantly improve their test scores. In particular, 82.5$\%$ of children who are provided an offer choose to take it up, and their subsequent tests scores can improve on average between 0.21 and 0.50 standard deviation points. Second, these effects primarily arise from the subgroup of children who do not have an offer to an alternative (non-Head Start) preschool and hence any outside option. Finally, using the estimates to perform a cost-benefit analysis of the policy, the earning benefits associated with the test score gains can potentially be large and outweigh the net costs associated with the take up of the offer. Specifically, when the relation between test scores and earnings is towards the higher end of estimates noted in the literature, the average benefits net of costs per child can range between $\$$12,365 and $\$$37,873.
The developed tools contribute to a large literature that exploit instrumental variables to learn about treatment effects and, specifically, in setups that permit multiple treatments. One strand of the literature focuses on the instrumental variable estimand and shows how it can evaluate treatment effects local to various complier subgroups under different types of monotonicity conditions angrist/imbens:95,heckman/pinto:18,kline/walters:16, kirkeboen/etal:16,pinto:19. Indeed, in cases where the instrument itself corresponds to the policy of interest, the compliers correspond to those affected by the policy change. But, for other more general policies such as the provision of a program offer, one may require learning about treatment effects for various other subgroups of interest.
To this end, the analysis in this paper is more closely related to an alternative strand of the literature that considers choice models to define and learn about various policy relevant parameters that move beyond local effects heckman/etal:06,heckman/etal:08,heckman/vytlacil:07,kline/walters:16,lee/salanie:19. The proposed choice model complements those in these papers that are each tailored for different questions of interest. Specifically, unlike these papers, the model here introduces structure to distinguish the role of choice sets and preferences in selection, which plays an essential role in defining and analyzing the effects of a program offer.
In developing the selection model, the analysis exploits elements from the literature on discrete choice analysis. As highlighted in Section (ref), manski:07 proposed a choice model with heterogeneity in both choice sets and preferences where one only assumed individuals to have a strict preference relationship over the various alternatives. However, while this model allowed for unobserved preferences, it took the choice sets to be observed and exogenously vary. Designed for setups more commonly arising in the program evaluation context, the model in this paper builds on this model to allow for unobserved choice sets and instrumental variable variation in choice sets. In this sense, the model also contributes to a growing literature on discrete choice analysis under unobserved choice sets abaluck/adams:18,barseghyan/etal:18,cattaneo/etal:18,crawford/etal:19, where models with alternative structure and assumptions between choice sets and preferences have been proposed.
The remainder of the paper is organized as follows. Section (ref) introduces the model, defines the average effects of an offer, and illustrates how various empirical examples fall in this setup. Section (ref) shows how to learn about the parameters of interest. Section (ref) applies the tools to the HSIS data. Section (ref) concludes. Proofs and auxiliary results complementing the main analysis are presented in the Supplementary Appendix.
For each individual, we observe an outcome of interest $Y$, their received treatment $D$, and their assigned instrument value $Z$. We take the list of of possible treatments to be given by a discrete set $\mathcal{D}$, and also model the support of the instrument to be given by a discrete set $\mathcal{Z}$. As usual, I assume the observed outcome to be generated by the potential outcomes structure. Denoting by $Y(d)$ the potential outcome had the individual received treatment $d \in \mathcal{D}$, the observed outcome is given by
In addition, I assume the received treatment to be generated by a selection model that introduces heterogeneity in both choice sets and preferences. Specifically, let $C \in \mathcal{C}$ denote the individual's received choice set that corresponds to a non-empty subset of $\mathcal{D}$ to where the individual receives offers from and has the option to participate, and let $U(d) \in \mathbf{R}$ denote the individual's utility from participating in alternative $d \in \mathcal{D}$. The observed treatment is then given by the treatment that maximizes utility across the alternatives present in their realized choice set as follows
Since the utilities do not possess any cardinal value, different monotonic transformations of them will generate observationally equivalent choices. It is therefore useful to directly refer to the underlying preference type that they represent. To this end, assuming individuals have a strict preference relation over the set of alternatives, let $U$ denote the individual's preference type corresponding to a strict preference relationship over $\mathcal{D}$, and let $\mathcal{U}$ denote the set of such types.\footnote{For example, if $\mathcal{D} = \{1,2\}$, the set of preference types can be given by $\mathcal{U} = \{ (1 \succ 2),~(2 \succ 1)\}.$} In addition, let $d(u,c)$ denote the known choice function that corresponds to what preference type $u \in \mathcal{U}$ would choose under a non-empty subset $c \subseteq \mathcal{D}$, i.e. under a given feasible choice set.\footnote{For example, $d((2 \succ 1), \{2,1\}) = 2$ as preference type $(2 \succ 1)$ prefers $2$ to $1$ and hence would choose $2$ when faced with the choice set containing $2$ and $1$.} Using this notation, the utility maximization relationship in (ref) can be alternatively re-written in terms of the preference type and realized choice set by
I also assume the instrument to affect where an individual receives an offer from and hence their choice set, while being excluded from how it affects utilities and outcomes. To this end, using the potential outcomes structure again, let $C(z)$ denote the individual's potential choice set had they been assigned instrument value $z \in \mathcal{Z}$, and let the realized choice set be given by
As illustrated in Section (ref), the relation between $C(z)$ across $z \in \mathcal{Z}$ will allow imposing how the instrument varies the provision of an offer.
Finally, as usual, I take the instrument to be exogenous and statistically independent of the underlying unobserved variables as follows.
\begin{assumption2}{E} (Exogeneity) $(\{Y(d) : d \in \mathcal{D} \},U,\{C(z) : \in \mathcal{Z}\}) \perp Z~.$ \end{assumption2}
A standard choice model in the literature on treatment effect analysis is to take the observed choice to be given by
where
for $z \in \mathcal{Z}$ with $\bar{U}(d,z)$ corresponding to the utility under $d \in \mathcal{D}$ heckman/vytlacil:07. This choice model is indeed more agnostic about the choice process than that in (ref) and can nest it by taking $\bar{U}(d,z) = U(d)$ if $d \in C(z)$ and $\bar{U}(d,z) = - \infty$ if $d \notin C(z)$ for each $d \in \mathcal{D}$ and $z \in \mathcal{Z}$, i.e. by taking utility to be equal to that of enrolling in that alternative if it is in the choice set and to be equal to negative infinity if not as that alternative cannot be feasibly chosen. In other words, the utility here can be interpreted as combining that of choosing an alternative and whether that alternative is in the choice set into a single dimension.
However, by not separating the heterogeneity in preferences and choice sets, this model does not straightforwardly allow analyzing the effects of providing an offer in comparison to the richer structure of the model in (ref). Importantly, this richness allows modeling the concept of an individual receiving an offer to an alternative or not---namely, by taking that alternative to be present in their choice set or not. As we will observe below, this plays an essential role in defining and then learning about a range of effects of providing a program offer by comparing responses under counterfactual choice sets with and without the program.
Explicitly modeling the role of choice sets is a common practice in the literature on discrete choice analysis when one wants to analyze choice under counterfactual choice sets. To this end, it is useful to highlight that the proposed model draws on this conceptual connection and, in particular, to the model in manski:07 with some novel modifications---see also the related analysis of marschak:59. manski:07 similarly introduced for each individual an unobserved preference type and a choice set which together determine observed choice as in (ref). However, designed with alternative empirical setups in mind, the model assumed the choice set $C$ to be observed and statistically independent of preferences, i.e. $C \perp U$. In the program evaluation setup studied in this paper, this amounts to assuming that from where offers were received are observed and independent of the individual's preferences, both of which may generally not be valid. The proposed model in turn shows how to alternatively consider a choice equation with unobserved choice set and instrumental variable variation in choice sets.
Before proceeding to show how we can define the effects of an offer in the proposed model, I illustrate how various empirical examples fit in the model. I focus on various randomized experiments that naturally induce exogenous variation in choice sets. In each experiment, I highlight what the treatment $D$ and instrument $Z$ correspond to and how the experiment varied provision of offers through the relation between $C(z)$ across $z \in \mathcal{Z}$.
I conclude this section by illustrating how to define various effects of an offer in the proposed model. To this end, recall that an offer to an alternative corresponds to an individual receiving that alternative in their choice set. The effects of an offer can therefore be defined by comparing potential responses under counterfactual choice sets with and without the alternative. To define these effects, it is useful to introduce additional notation to refer to the potential decisions and outcomes under a given choice set. Had the individual's choice set been pre-specified to a non-empty subset $c \subseteq \mathcal{D}$, let $D_{c} = d(U,c)$ denote the alternative they would choice from $c$ and $Y_{c} = Y(D_c)$ the subsequent outcome they would then earn.
For each individual, if an offer to an alternative of interest $d \in \mathcal{D}$, typically corresponding to the program, was provided then their resulting choice set would be given by $C_{+d} = C \cup \{d\}$, while if it was not then it would be given by $C_{-d} = C \setminus \{d\}$. The average (treatment) effect (ATE) of providing an offer to $d \in \mathcal{D}$ on choosing this alternative and subsequent outcomes can then be defined by taking the average value of the difference in the respective potential responses under these two choice sets given by
Note that by construction we have $1\{D_{C_{-d}} = d\} = 0$, which captures the fact that if the individual is not provided an offer to an alternative then they naturally cannot choose it. As a result, observe that the former parameter can more simply be rewritten as
i.e. the proportion who choose the alternative when provided an offer. Observe that providing an offer does not affect the subgroup who do not take up the offer, i.e. those with $D_{C_{+d}} \neq d$, as for these children we have $Y_{C_{+d}} = Y_{C_{-d}}$. It is therefore useful to also define the average (treatment) effect of providing an offer conditional on participation (ATOP), i.e. for those who take it up and are hence affected, by
We can also consider the above average effects of an offer conditional on a subgroup of interest. For example, as done in the empirical analysis in Section (ref), we may be interested in evaluating the heterogeneity in the effects of an offer to an alternative based on whether individuals do or do not have an offer to another alternative. In particular, we can define
to capture the proportion who take up the offer and the effects on their subsequent outcomes conditional on having an offer to alternative $d' \in \mathcal{D}$, or
to capture the analogous effects conditional on not having an offer to alternative $d' \in \mathcal{D}$.
In this section, I show how to exploit the structure of the model and variation in choice sets to learn about the parameters of interest. For these purposes, it is useful to first introduce the underlying primitive of the model on which the analysis is based. Let $W = (\{Y(d) : d \in \mathcal{D} \},U,\{C(z) : \in \mathcal{Z}\})$ summarize the latent variables for each individual. The primitive $Q$ then corresponds to the distribution of $W$. To ensure that $Q$ is a finite-dimensional object, I assume that the outcomes take values in a known discrete set $\mathcal{Y} \equiv \{ y_1, \ldots, y_M \}$. This implies that $Q$ is a probability mass function with support contained in a discrete space $\mathcal{W} \equiv \mathcal{Y}^{|\mathcal{D}|} \times \mathcal{U} \times \mathcal{C}^{|\mathcal{Z}|}$, i.e. $Q : \mathcal{W} \to [0,1]$ and $\sum\limits_{w \in \mathcal{W}} Q(w) = 1$. As we will observe below, this allows characterizing what we can learn about the parameter using finite-dimensional computational programs.
The objective is to learn about a pre-specified parameter of interest defined in Section (ref). Each of these parameters can be written in terms of a known function of $Q$. Denoting by $\theta(Q)$ a generic function for the pre-specified parameter, $\theta$ generally has the following form
where $a_{\text{num}} : \mathcal{W} \to \mathbf{R}$ and $a_{\text{den}} : \mathcal{W} \to \mathbf{R}$ are known functions, i.e. as a fraction of linear functions of $Q$. For example, for the parameter (ref) in the context of Example (ref) when $d=p$, we can write it as a linear function, a special case of linear-fractional functions where the denominator takes a value of one by taking $a_{\text{den}}(w) = 1$ for all $w \in \mathcal{W}$, as follows
where $\mathcal{W}^p = \{ w \in \mathcal{W} : d(u,c(1)) = p\}$. As illustrated in Appendix (ref), we can similarly do so for the remaining parameters of interest. Indeed, note that the parameters that do not condition on any event such as that in (ref) can be written as linear parameters, while those that condition on some event can be written more generally as linear-fractional parameters.
Given that the function $\theta$ is known, what we can learn about the parameter depends on the possible values that $Q$ can take. From the observed data and its relation to the latent variables through the structure (ref)-(ref) along with Assumption (ref), we have that
for all $x=(y,d,z) \in \mathcal{Y} \times \mathcal{D} \times \mathcal{Z} \equiv \mathcal{X}$, where $\mathcal{W}_{x}$ is the set of all $w \equiv (\{y(d'):d' \in \mathcal{D}\},u,\{c(z') : z' \in \mathcal{Z}\})$ in $\mathcal{W}$ such that $d(u,c(z))=d$ and $y = y(d)$. For the purposes of identification, the data moments on the right hand side of the above equation are assumed to be known with certainty. I discuss estimation and inference in Section (ref) below. In addition to the data, the assumptions we impose as in Section (ref) on the variation in choice sets also restrict $Q$. In particular, these restrictions can generally be stated as
where $\mathcal{W}_s$ is some known subset of $\mathcal{W}$, i.e. they impose zero probability on the occurrence of some events. For example, observe that (ref) can be written as (ref) by taking $\mathcal{W}_s = \{w \in \mathcal{W} : c(1) \neq c(0) \cup \{p\}\}$. The restrictions in (ref)-(ref) can be similarly stated in this manner. Moreover, as presented in Section (ref) in context of the empirical application, (ref) can also encompass additional such restrictions on the relation between outcomes and preferences.
Denoting by $\mathbf{Q}_{\mathcal{W}}$ the set of all probability mass functions on the sample space $\mathcal{W}$ and by $\mathbf{Q} = \{Q \in \mathbf{Q}_{\mathcal{W}} : Q \text{ satisfies } \eqref{eq:Qdata} \text{ and } \eqref{eq:Qass}\}$ the subset of $\mathbf{Q}_{\mathcal{W}}$ that satisfies the various restrictions imposed by the data and assumptions, what we can learn about the parameter can then be defined by the identified set,
i.e. the image of the space of admissible primitives under $\theta$. The objective of the analysis then translates to show how to compute the identified set.
A standard approach in the literature is to analytically characterize the identified set---see manski:09 for a textbook treatment. While analytic bounds can provide intuition into what is driving identification, doing so in the above setup can be difficult due to the complicated model structure and parameter function $\theta(Q)$. To this end, in the following proposition, I illustrate how a linear programming procedure can be used to instead computationally characterize it---see, for example, balke/pearl:97, demuynck:15, freyberger/horowitz:15, kline/tartari:16, laffers:13, manski:07,manski:14, mogstad/etal:17 and torgovitsky:16,torgovitsky:18 that similarly propose linear programs to compute identified set in related models. In this proposition, I introduce the following additional quantity $\widetilde{\theta}(Q) \equiv \sum\limits_{w \in \mathcal{W}} a_{\text{num}}(w) \cdot Q(w)~.$
Given the structure of the linear-fractional parameters and the linear restrictions imposed on $Q$, Proposition (ref) illustrates that the identified set for each parameter is an interval and that the bounds of this interval can be computed by solving two linear programming problems. To do so, the proof first exploits the fact that the identified set can be written as solutions to so-called linear-fractional programs and then results from charnes/cooper:62 that show how to transform such programs to equivalent linear programs---see also russell:21 who independently makes a related observation in an alternative treatment effect model.
Computing the identified set using Proposition (ref) requires two conditions. First, it requires $\mathbf{Q}$ to be non-empty, i.e. the model is correctly specified, to ensure that the identified set is a non-empty interval. If this is not the case, the linear programs automatically terminate. Second, it requires the denominator of the parameter of interest to be positive for every $Q \in \mathbf{Q}$ to ensure that the parameter is well-defined for all feasible model distributions. This condition can easily be verified in practice. To see how, note that the denominator is also a linear-fractional parameter and more specifically a linear parameter. The above proposition can hence be employed to compute the lower bound for this auxiliary parameter to check if it is strictly positive.
The above characterization of the identified set assumed that the data moments in (ref) were known with certainty. In practice, we can estimate it using the sample analogue of these moments. I conclude this section by illustrating how to compute confidence intervals for the parameter of interest that account for the resulting statistical uncertainty.
Confidence intervals can be constructed by test inversion. In particular, we can test at level $\alpha \in (0,1)$ the following null hypothesis
i.e. the parameter of interest in (ref) can be equal to some value $\theta_0 \in \mathbf{R}$ given the admissible set of distributions $\mathbf{Q}$. Confidence intervals of level $(1-\alpha)$ for the parameter can then be obtained by collecting the values of $\theta_0$ that are not rejected.
To test this null hypothesis, I propose to use a procedure from bugni/etal:17, which they show has attractive theoretical properties that account for the fact that the parameter is partially identified and, as I highlight below, can be computationally attractive in the above setup---see canay/shaikh:17 for an overview of statistical inference procedures in partially identified setups. Denoting the sample size by $N$, the test is based on a test statistic given by
subject to the constraints that $Q(w) \geq 0$ for each $w \in \mathcal{W}$, $\sum\limits_{w \in \mathcal{W}} Q(w) = 1$, (ref), and
where $\hat{P}$ corresponds to the estimated version of $P$ obtained using the empirical distribution of the data, i.e. the test statistic measures the minimum squared distance from zero of the relation between $Q$ and the data in (ref) subject to the various restrictions on $Q$ and that $Q$ can generate $\theta_0$. The test then computes a critical value using the bootstrap. Specifically, for each bootstrap sample $l=1,\ldots,B$, denoting by $\hat{P}_l$ the analogue of $\hat{P}$ computed using the $l$th bootstrap sample, we compute a so-called minimum resampling bootstrap test statistic given by
where
is a so-called discard resampling test statistic, and
subject to the constraints that $Q(w) \geq 0$ for each $w \in \mathcal{W}$, $\sum\limits_{w \in \mathcal{W}} Q(w) = 1$, (ref), and (ref), is a so-called penalize resampling test statistic---see bugni/etal:17 for intuition behind these two test statistics. Here $\kappa$ is a tuning parameter that satisfies $\kappa \to \infty$ and $\kappa / \sqrt{N} \to 0$ as $N \to \infty$---for example, following bugni/etal:17, we can take $\kappa = \sqrt{\text{ln}N}$.
For practical purposes, note that various test statistics in the above procedure can be relatively straightforwardly computed, where that in (ref) can be directly computed and the optimization problems in (ref) and (ref) can be computed using a quadratic program given the quadratic objective and linear constraints. The critical value of the test $c(1-\alpha,\theta_0)$ can then be computed by taking the $(1-\alpha)$-quantile of the distribution of bootstrap test statistics in (ref), i.e. $\{TS^{MR}_l(\theta_0) : 1 \leq l \leq B\}$, given which the test equals $\phi(\theta_0) = 1\{TS(\theta_0) > c(1-\alpha,\theta_0)\}$, i.e. we reject if the test statistic is larger than the critical value.
In this section, I apply the developed tools to the Head Start Impact Study to empirically analyze the average effects of providing an offer to Head Start.
Head Start is the largest early childhood education program in the United States that provides free preschool education to three- and four-year-old children from disadvantaged households. The Head Start Impact Study (HSIS) was a publicly-funded experiment implemented in Fall 2002 to evaluate Head Start. It was conducted by randomly assigning a sample of three- and four-year-old applicants in certain Head Start preschools to a treatment or control group. Those in the treatment group were provided an offer to Head Start by admitting them in the Head Start preschool to which they had applied, while those in the control were not. However, children in the control group could also receive a Head Start offer if they could be admitted to a preschool not part of the study. See puma_etal:10 for further details on the program and experiment.
The objective of the analysis is to exploit the experimental variation induced by the HSIS to evaluate the effects of providing an offer to Head Start. I do so by modeling the setting as in Example (ref) and then evaluating the parameters defined in Section (ref). It is useful to highlight that previous analyses of the HSIS, in contrast, have primarily focused on evaluating parameters based on the so-called intention-to-treat (ITT) and instrumental variable (IV) estimands feller/etal:16, kline/walters:16, puma_etal:10. The ITT estimands can be defined as
i.e. the difference in outcomes and participation in Head Start between the treatment and control groups, and the IV estimand as
i.e. the ratio of the ITT on outcome to that on participation in Head Start. In the following proposition, I illustrate what these estimands allow learning about the effects of a Head Start offer.
Proposition (ref) illustrates that these estimands evaluate the effects of the change in choice sets induced by the experiment. They can therefore be interpreted as the policy effects of providing an offer through assignment relative to not through assignment to the treatment and control groups, respectively. However, while those in the treatment group received a Head Start offer and hence $C(1) = C_{+p}$, note that those in the control group could also potentially receive an offer and hence $C(0)$ does not necessarily equal $C_{-p}$. In turn, they do not necessarily answer what are the policy effects of providing an offer to Head Start relative to not more generally.
Through the defined parameters, the analysis here aims to precisely evaluate such policy effects. As noted in kline/walters:16, it is also useful to highlight that providing an offer can open slots at alternative preschools and result in equilibrium changes on the distribution on who has offers to various preschools. In turn, as the proposed model does not account for such general equilibrium effects, the defined parameters may more appropriately be viewed as capturing the policy effects of providing a Head Start offer taking the slots at other schools fixed. Alternatively, they may be viewed as the effects of a marginal policy providing an offer to a single child as such a policy is more likely to keep such slots fixed.
The analysis uses the data collected by the HSIS. Following kline/walters:16, I use a version of the data from the first year of the experiment that pools the two cohorts of three- and four-year-olds together. Moreover, following feller/etal:16 and kline/walters:16, I take the participation decision to be a categorized version of an administratively coded focal care setting variable that evaluated care setting attendance for the entire year. I take the test score outcome to be a discretized version of that used in kline/walters:16. They use a summary index of cognitive test scores as measured by the average of the Woodcock Johnson III and the Peabody Picture and Vocabulary Test test scores, where each score is standardized to have mean zero and variance one in the control group for each cohort. Given this summary index, I discretize it using the quantiles of its empirical distribution to take ten support points.\footnote{In unreported results, I compute the bounds in Table (ref) using 15 and 20 support points, and find that the numerical results are relatively similar and that the qualitative conclusions remain unchanged.} See Appendix (ref) for further details on how the analysis sample was constructed.
Table (ref) provides some summary statistics for the analysis sample. Panel (a) reports the participation shares across the different types of care settings by treatment and control groups. A large proportion (82.5$\%$) of the treatment group participated in Head Start, which suggests that many prefer Head Start over their alternative options. Comparing to those in Head Start in the control group suggests that many control children may not have had a Head Start offer. However, given that this proportion is not zero, this suggests that at least some proportion (14$\%$) in fact received an offer even when assigned to the control group.
Table (ref) Panel (b) reports estimates for the estimands in (ref)-(ref). As $\text{ITT}^Y$ does not require the test scores to be discrete, I also report results under the raw undiscretized version of the test scores---the results under both versions are relatively similar. Since some children in the control group received a Head Start offer, these estimands, as shown in Proposition (ref), generally evaluate the effect of providing a Head Start offer through assignment to treatment group relative to the control group. The IV estimate reveals such a policy results in a positive effect of 0.249 standard deviation points for the individuals who are affected by it.
Table (ref) reports estimated identified sets and 90$\%$ confidence intervals for the parameters evaluating the average effect of a Head Start offer in (ref)-(ref) when taking $d$ to be equal to $p$, i.e. Head Start. The different columns correspond to different combinations of assumptions, which we motivate and describe below. As in the rest of the analysis below, the identified sets are estimated by applying the linear programming procedure from Proposition (ref) to the empirical distribution of the data, and the confidence intervals are constructed as in Section (ref) where the bootstrap samples are drawn at the level of Head Start centers.
Under the baseline specification in Column (1), $\text{ATE}^D(C_{+p},C_{-p})$ reveals that providing a Head Start offer affects a large percentage of children (85.5$\%$) who take up the offer and participate in Head Start. Note that this parameter is point identified as $C_{+p} = C(1)$ given (ref) and is hence equal to the proportion participating in Head Start in the treatment group from Table (ref) Panel (a). Turning to the effects on test scores, $\text{ATE}^Y(C_{+p},C_{-p})$ and $\text{ATOP}(C_{+p},C_{-p})$ reveal that the upper bound can be potentially large and greater than the IV estimand, highlighting that large test scores are consistent with the data. However, the lower bound is smaller than zero and hence the data alone does not allow us to conclude whether we have positive or negative.
Intuitively, this arises because the baseline specification leaves the relationship between the test scores and other model variables completely unrestricted. To this end, I consider the following additional assumptions restricting the relation that help is reaching more informative conclusions, and which can be written as (ref) as shown in Appendix (ref).
\begin{assumption2}{NC} (Nested Choice) $d(U,\{n,p\}) = p$ if and only if $d(U,\{n,a\}) = a$ . \end{assumption2}
\begin{assumption2}{MTR} (Monotone Treatment Response) $Y(p) \geq Y(n)$ and $Y(p) \geq Y(a)~$. \end{assumption2}
\begin{assumption2}{Roy} For each $d,d' \in \mathcal{D}$, if $Y(d') > Y(d)$ then $d(U,\{d,d'\}) = d'$ . \end{assumption2}
Assumption (ref) imposes parents to prefer Head Start to home care if and only if they also prefer alternative preschools to home care. It is motivated by the idea that parents may have preferences where they prefer to enroll their child in any preschool, Head Start or an alternative one, to home care or vice versa. Assumption (ref), based on manski:97b, imposes test scores under Head Start to be at least as good as those under alternative preschools and home care. It is motivated by the notion that since Head Start is considered to be a high quality care setting and that families in Head Start are primarily disadvantaged, it may be that Head Start is a better option than home care and alternative preschools that these families could afford. Finally, Assumption (ref) imposes parents to prefer enrolling their child in the care setting with the higher test score, i.e. it takes preferences to be based solely on maximizing outcomes in the spirit of the Roy selection model heckman/honore:90, mourifie/etal:15.
While Assumption (ref) is reasonable, Assumptions (ref) and (ref) may seem potentially strong to impose for every child in the population. To this end, I also consider weaker versions of them, where I assume that they hold for only some pre-specified proportion of children. More formally, let the primitive distribution of the model be given by
for each $w \in \mathcal{W}$, where $\lambda \in [0,1]$ corresponds to the known proportion of children for which either Assumptions (ref) or (ref) hold, and $H_1$ and $H_0$ correspond to the underlying distribution for the children for whom these assumptions hold and may not hold, respectively---i.e., the restrictions imposed by these assumptions of the form (ref) are imposed only on $H_1$ and not on $H_0$, which are otherwise identical in term of mass assigned to preferences and choice sets. Indeed, if $\lambda$ equals one, these weaker versions are equivalent to Assumptions (ref) or (ref). However, by taking this proportion to be smaller than one, we can make these assumptions flexibly weaker and analyze sensitivity of the empirical conclusions to the assumptions---similar in spirit to the breakdown point analysis considered in horowitz/manski:95, kline/santos:13 and masten/poirier:17. In Appendix (ref), I illustrate how Proposition (ref) can be straightforwardly modified to account for these weaker versions of Assumptions (ref) or (ref).
The results are presented in Columns (2)-(7) of Table (ref). From Columns (3) and (6), we can observe that under Assumptions (ref) or (ref), the lower bound of $\text{ATOP}(C_{+p},C_{-p})$ significantly increases and allows concluding that the offer can have a large positive benefit to children who take it up. In particular, their test scores on average can increase between 0.206 and 0.497 standard deviation points. Moreover, from Columns (4)-(5) and (7)-(8), these conclusions can continue to hold under more flexible specifications that impose these additional assumptions not on all children but only on some given proportion, namely $80 \%$ and $90 \%$ in this case.
The above results reveal that a Head Start offer affects a large number of children who take up the offer, and can have a positive effect on their subsequent test scores. To better understand these effects and their driving factors, I next analyze heterogeneity based on whether a child had an offer or not to an alternative (non-Head Start) preschool, i.e. the parameters in (ref)-(ref) taking $d' = a$.
Table (ref) reports estimated identified sets and 90$\%$ confidence intervals for these parameters. Observe that parameters evaluating the effects for the group who do not have an alternative preschool offer are not well-defined under all the specifications. This is because the probability of this subgroup, corresponding to the conditioning event of these parameters, can potentially be zero given the data and specification. In turn, since this implies that (ref) is not satisfied, these parameters are not well-defined.\footnote{In this case, the confidence intervals equal the logical bounds under these specifications as the numerator also equals zeros when the denominator does for these parameters and hence all values of $\theta_0$ permitted under the specification can satisfy (ref).} However, when Assumption (ref) is imposed, this is no longer the case as the probability of the conditioning event is bounded away from zero.
The results reveal that providing a Head Start offer affects more children who don't have an alternative preschool offer than those who do. In particular, Column (2) reveals that between 57.5$\%$ and 82.2$\%$ of children with an alternative preschool offer take up the Head Start offer in comparison to between 82.8$\%$ and 100$\%$ of those without an alternative preschool offer. Moreover, under Assumptions (ref) or (ref), the results reveal that providing an offer can potentially have positive test score gains for only those who don't have an alternative preschool offer and not for those who do. These results together highlight that the effects of providing an offer are primarily driven by the subgroup of children who do not have an alternative preschool offer and hence any outside preschool option.
The motivation to study the above effects of a Head Start offer was to capture the potential costs and benefits of a policy providing an offer. I conclude the empirical analysis by using the various effects to perform a cost-benefit analysis of such a policy, similar in spirit to that performed by kline/walters:16 using the local effects. The analysis can be done by attaching an income value to the test score gains to evaluate the benefits as well as one to the take up of Head Start offer to evaluate the costs. To this end, let $\rho$ denote the income value for the test score gains and $\phi_p$ denote the cost of a student participating in Head Start. Moreover, as emphasized in kline/walters:16, it is also important when measuring the costs to account for the potential cost savings associated with the movement of children out of alternative preschools, as these preschools are often publicly funded as well. To this end, let $\phi_a$ analogously denote the cost of participating in an alternative preschool.
Using these values, the average benefits of providing a Head Start offer relative to not can be given by the average effect of test scores times the income value, i.e.
Similarly, the average costs can be given by the proportion who take up the Head Start offer times the cost of participation along with subtracting the costs associated with those who were induced into Head Start from an alternative preschool, i.e.
Taking the difference of the benefit and cost parameters, we can then define the average surplus
which can be used to perform a cost-benefit of a policy providing an offer. Given values of $\rho$, $\phi_p$ and $\phi_a$, these parameters, similar to (ref)-(ref), can be rewritten as (ref). In turn, the linear programming procedure can be continued to be applied to characterize the identified set for these parameters. Following kline/walters:16, I take $\rho = \kappa \cdot E$, where $E$ corresponding to the average present discounted value of lifetime earnings for Head Start applicants is taken to be $\$$34,339 and $\kappa$ corresponding to the relationship between earnings and a one standard deviation increase in test scores is taken to be either 0.1 or 0.3, which is based on the lower and higher end of the values for this relation noted in the literature kline/walters:16. Moreover, following kline/walters:16, I take $\phi_p = \$ 8,000 $ and $\phi_a = 0.75 \cdot \phi_p = \$ 6,000$.
Table (ref) reports estimated identified sets and 90$\%$ confidence intervals for these parameters for a subset of the specifications from Table (ref) that impose restrictions on test scores. The results reveal that the net costs of providing an offer is relatively low in comparison to the absolute value of participating in Head Start. Similar to that highlighted by kline/walters:16, this arises due to relatively high cost savings from children who enroll out of alternative preschools when provided with a Head Start offer. Taking the benefits into account, the upper bounds of the average surplus reveals that benefits can significantly outweigh costs between the $\$$9,727 and $\$$37,873 per child depending on the relation between test scores and earnings. However, whether the lower bound is positive such that the surplus is in fact positive depends on the strength of the assumptions and the relation between test scores and earnings. Specifically, when Assumptions \ref{ass:MTR} or \ref{ass:Roy} are imposed on all children in Column (4) and (7), the surplus can be positive. But, when these assumptions are imposed only on some proportion of the children in the remaining columns, it can be positive only when $\kappa = 0.3$, i.e. the relation between test scores gains and earnings is based on those towards the higher end of that noted in the literature heckman/etal:10.
I propose a novel treatment selection model with unobserved heterogeneity in both choice sets and preferences to evaluate the average effects of a program offer. I show how to exploit the structure of the model to define parameters capturing these effects and then computationally characterize their identified sets under instrumental variable variation in choice sets. I illustrate these tools by analyzing the effects of providing a Head Start offer using data from the Head Start Impact Study.