Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
111,566 characters · 13 sections · 75 citation commands
Nonparametric Treatment Effect Identification in School Choice
\thispagestyle{empty}
\Copy{specificstudy}{A rapidly growing empirical literature studies causal effects of schools using centralized school assignment.\footnote{See, among others, kapor2020heterogeneous,agarwal2018demand,fack2019beyond,calsamiglia2020structural,angrist2020simple,beuermann2018good,abdulkadirouglu2017research,abdulkadiroglu2017impact,abdulkadirouglu2020parents,abdulkadroglu2017regression,marinho2022causal,kirkeboen2016field,angrist2021credible,ketel2023unimport,angrist2022race.} In several centralized assignment systems, two sources of variation determine how schools enroll applicants and drive quasi-experimental identification. For instance, in New York City, studied by abdulkadiroglu2017impact, some schools randomize priority to students, where students with higher lottery numbers are favored over others. Other schools use non-lottery tiebreakers, such as test scores, to distinguish students with otherwise similar characteristics.
These sources of variation provide opportunities for causal inference. For lottery schools, comparing lucky students who receive favorable priorities to those less fortunate estimates a causal effect. For test-score schools, comparing students just above a cutoff to those just below also identifies causal effects. abdulkadiroglu2017impact use both sources of variation to measure school effectiveness in New York City. School districts elsewhere, such as Chicago and Boston, also feature schools with lottery and non-lottery tiebreakers.\footnote{angrist2023methods write “Some centralized assignment schemes, such as those used for Boston and New York City exam schools and New York City screened schools, employ non-lottery tiebreakers such as test scores instead of, or alongside, lottery numbers.” Schools in the Chicago Public School High School Choice program use either a lottery system or a points system for giving priority to students, where the points system is a composite of test scores and other student characteristics. See \url{https://www.cps.edu/gocps/high-school/results/choice/}. }
We refer to the first source of variation as lottery variation and the second as regression-discontinuity (RD) variation. Researchers often estimate aggregate causal quantities, such as contrasts between sets of “treatment” and “control” schools, pooling over lottery and RD variation and over pairs of schools. abdulkadiroglu2017impact do so for New York City, estimating aggregate effects of enrolling in a school receiving an “A” grade on the school district's report card for quality, compared to enrolling in non-Grade-A schools. }
These aggregate estimates provide convenient causal summaries but can mask rich heterogeneity by pooling many different comparisons. These contrasts may involve different pairs of schools, different student characteristics, or different sources of identifying variation. Aggregate causal estimands are often interpreted as simple average treatment effects between treatment and control schools. Because these aggregate effects blend together different conditional average treatment effects, their interpretation should be more nuanced.
In particular, economically, heterogeneity along RD variation versus lottery variation could be meaningful. Consider a lottery school $s_1$ and a test-score school $s_0$. In many centralized algorithms, lottery variation between these schools occurs only for students preferring $s_{1}$ to $s_{0}$, driven by whether they lose the $s_{1}$ lottery. Meanwhile, RD variation exists only for students who prefer $s_0$ to $s_1$---driven by whether these students narrowly qualify or fail the test used by $s_0$. If submitted preferences select positively on gains, the causal contrasts identified by lottery (resp. RD) variation represent subpopulations with more (resp. less) favorable potential outcomes under $s_1$. Different aggregations thus put different weights on subpopulations with different attitudes toward $s_1$.
Motivated by the rich heterogeneity, this paper seeks to reduce causal effects to their smallest unit, which we call atomic treatment effects (aTEs). We define an atomic treatment effect as a conditional average treatment effect between two schools, conditional on all observed student characteristics. Our first contribution is to characterize the set of point-identified aTEs based solely on lottery and RD variation from the assignment mechanism. Since any identified average treatment effect is thus a weighted average over aTEs, this characterizes the set of estimable treatment effect parameters in these settings.
Our characterization yields a clean classification of aTEs as driven by RD or lottery variation. Building on this classification, one concerning observation is that common regression estimators---often derived under homogeneous treatment effect assumptions---implicitly aggregate aTEs using weights that are not chosen by the researcher and can depend on the sample size. Moreover, these implicit weighting schemes can be at odds with the interpretation of these estimates. Starkly, aggregate treatment effects that pool over lottery and RD variation asymptotically put zero weight on the atomic treatment effects identified by RD variation. Our previous observation then also implies that these aggregates only reflect treatment contrasts for test-score schools for students who disprefer these test-score schools. Nevertheless, such estimators are often interpreted as partly reflecting the RD variation on a broader set of students.\footnote{For instance, abdulkadiroglu2017impact employ such an approach, and the title of abdulkadiroglu2017impact suggests that RD-variation is central. To be sure, there is evidence that in finite samples, their estimator puts nontrivial weight on the RD-driven causal comparisons, even though the weight put on these comparisons converges to zero asymptotically---see (ref).}
The reason for this result is simple: Since RD variation can only occur at a cutoff, atomic treatment effects identified by such variation make comparisons for a “thin” set of student characteristics that have population measure zero khan2010irregular---namely, those students with scores at a cutoff. Translated to estimation, this means that the number of students contributing to RD variation is a vanishing fraction compared to the number of students subject to lottery variation. As a result, treatment effect aggregations that weigh each student equally put vanishing weight on the RD-identified aTEs.
Our second set of contributions speaks to these concerns: We provide practical recommendations, a new diagnostic, and theoretical guarantees. First, we introduce a diagnostic that assesses---in finite samples---how much RD-variation contributes to a particular regression estimator. In some cases, in finite samples, the weight on RD variation for these estimators may be non-zero and substantial. This diagnostic assesses whether particular estimators are subjected to the starkest problems arising from their implicit weighting schemes.
\Copy{aggregationlit}{Second, researchers are encouraged to take an explicit stand on the aggregation of aTEs. Our identification result for aTEs helps practitioners reason about the behavior of identified aggregate estimands. We also provide some reasonable default choices: We first aggregate identified aTEs for each ordered pair of schools $s_1, s_0$, among those that prefer $s_1 \succ s_0$, into parameters $\tau_{s_1\succ s_0}$. This aggregation respects identifying variation, and thus lottery-driven aTEs do not dominate. We then propose further aggregation to school-level value-added by explaining the variation in $\tau_{s_1 \succ s_0}$ through the difference in $s_1$ and $s_0$'s value-added, similar to Bradley--Terry models of tournament rankings. In spirit, our call for taking an explicit stand on aggregation echoes cattaneo2016interpreting and bertanha2020regression, who consider aggregation of regression discontinuity treatment effects at multiple cutoffs. Our setting is additionally complex due to the presence of lotteries and of centralized assignment. }
To implement a user-chosen aggregation of aTEs, we also provide asymptotic theory for estimators of lottery- and RD-identified aTEs---which can then be further aggregated according to a user-chosen weighting. As a theoretical contribution, relative to the existing literature abdulkadiroglu2017impact,abdulkadirouglu2017research, our asymptotic theory more accurately reflects the fact that cutoffs have nontrivial randomness in finite samples, and only converge to fixed quantities asymptotically azevedo2016supply. \Copy {introdiff}{Nevertheless, we show that data-dependent cutoffs do not affect the first-order asymptotic behavior of estimators for these treatment effects, using a more refined analysis than in abdulkadiroglu2017impact,abdulkadirouglu2017research. This is a novel contribution to the statistical theory of estimators in settings with school assignment algorithms.}
This paper applies to any setting where school assignments are made using deferred acceptance-like algorithms and researchers seek causal identification based on the centralized assignments beuermann2018good,abdulkadirouglu2017research,marinho2022causal,kirkeboen2016field,angrist2021credible. It contributes to a growing methodological literature on causal inference in centralized school assignment settings singh2023kernel,marinho2022causal,munrojmp,abdulkadiroglu2017impact,abdulkadirouglu2017research,narita2021theory,narita2021algorithm,che2023leveraging,arkhangelsky2025evaluating. This paper is most related to abdulkadiroglu2017impact. We complement their work by investigating the ramifications of heterogeneous treatment effects on their identification strategy and by formally characterizing the set of atomic treatment effects identified by centralized school assignment.\footnote{In particular, abdulkadiroglu2017impact “look forward to a more detailed investigation of the consequences of heterogeneous treatment effects for identification strategies of the sort considered here,” which is precisely the theme of this paper.} Our results also imply that the regression estimator proposed by abdulkadiroglu2017impact estimates a treatment effect that puts vanishing weight on the RD-identified aTEs, though nevertheless it appears the weight on RD-identified aTEs is nontrivial in finite samples in their empirical application.
This paper proceeds as follows. (ref) introduces notation, setup, a class of school choice mechanisms, and a set of treatment effect estimands; in particular, (ref) discusses in detail the motivation for our identification results and some recommendations for aggregating atomic treatment effects. (ref) characterizes those that are identified. (ref) proves asymptotic properties for standard estimators of lottery- or RD-driven estimands. (ref) illustrates our recommendations and results with a Monte Carlo study. (ref) concludes.
Consider a finite set of students $I = \{1,\ldots, N\}$ and schools $S = \{0, 1, \ldots, M\}.$ The schools have capacities $q_1, \ldots, q_M \in \N$. Assume that school $0$ represents an outside option with infinite capacity. Each student $i\in I$ has observed (by the analyst) characteristics $X_i$. $X_i$ contains characteristics that are relevant for the assignment mechanism.\footnote {It may also contain other observed characteristics that the analyst may condition on. Since these additional covariates do not change our results materially, we suppress them and assume for convenience that they are not available.} For each student $i$ and each school $s$, there is a potential outcome $Y_i(s)$, representing the outcomes of the student if assigned to school $s$.\footnote{Consistent with much of the literature, this formulation rules out peer effects or other violations of the stable unit treatment value assumption(SUTVA). } The assignment mechanism takes in $X_1,\ldots, X_N$ and produces matchings $D_i = [D_ {i1},\ldots, D_{iM}]'$, such that $D_{is} = 1$ if and only if $i$ is matched to $s$, and $(D_1,\ldots, D_N)$ satisfies school capacity constraints.\footnote{That is, $\sum_{i} D_{is} \le q_s$ for all $s$. } The analyst observes $Y_i = \sum_{s} D_{is} Y_i(s)$, $D_i$, and $X_i$ for every individual.
To embed this setup in an asymptotic sequence, assume that $(X_i, Y_i(0),\ldots, Y_i (M)) $ are i.i.d. draws from a superpopulation. The set of schools $S$ is fixed, but their capacities in finite samples are generated by $q_s = \lfloor N q_s^* \rfloor$ for some fixed value $q_s^* \in [0,\infty)$.
Following abdulkadiroglu2017impact, we consider the assignment mechanisms that derive matchings from student-proposing deferred acceptance. This setup accommodates the school choice mechanism in New York City, and nests many mechanisms that either use deferred acceptance or can be reduced to deferred acceptance.\footnote{We consider the same mechanisms as abdulkadiroglu2017impact. See footnote 5 in abdulkadiroglu2017impact for a list of school choice settings accommodated by this setup, which includes mechanisms (e.g. immediate acceptance) that can be represented by deferred acceptance under certain transformations of preferences.} Deferred acceptance requires student rankings over schools (termed preferences) and school rankings over students (termed priorities). Students submit preferences, which are included in $X_i$, and school priorities are computed from $X_i$ as follows.
There are two types of schools, lottery schools and test-score schools. School priorities are lexicographic in either $(Q_{it_s}, R_ {it_s})$ or $(Q_{i \ell_s}, U_{i \ell_s})$, for $R$ the set of test scores and $U$ the set of lottery numbers. The quantity $Q_ {is} \in \br{0,\ldots, \bar q_s}, \bar q_s \in \N \cup \br{0},$ represents certain qualifiers; higher values of $Q_{is}$ mean that $i$ is ranked favorably by $s$. This accommodates settings where, for instance, having a sibling at $s$ makes $i$ high-priority; we may represent such students with $Q_{is} = 1$ and others with $Q_{is} = 0$. When two students $ (i,j)$ have the same discrete qualifier $Q_{is} = Q_{js}$, ties are broken by either a test score $R_{it_s}$ or a lottery number $U_ {i\ell_s}$, where $t_s, \ell_s$ indexes which test or lottery number school $s$ uses.
Formally, assume we partition $S$ into lottery and test-score schools. A lottery school $s$ uses a lottery indexed by $\ell_s \in \br{1,\ldots, L}$, $L\in \N$, and a test-score school uses a test-score indexed by $t_s \in \br{1,\ldots, T}$, $T\in \N$. Two schools may use the same lottery or test score for tie-breaking. Assume that the assignment-relevant information $X_i$ takes the form \[ X_i = (\succ_i, R_i, Q_i) = (\succ_i, \underbrace{(R_{i1}, \ldots, R_{iT})}_{ \text{continuously distributed on $[0,1]^T$}}, (Q_{i0},\ldots, Q_{iM})) \] where $\succ_i$ represents the preferences of student $i$, $R_{it}$ represents $i$'s test score on test $t$, and $Q_{is}$ represents $i$'s discrete qualifier at school $s$.
Let $U_{i} = [U_{i1},\ldots, U_{iL}]' \iid F_U$, supported on $[0,1]^{L}$, be the vector of lottery numbers for student $i$, where components $(U_{i\ell_1}, U_{i\ell_2})$ need not be independent. To implement the lexicographic priorities in $Q_{is}$ and $U_{i\ell}$ (or $R_{it}$), each school $s$ computes a priority score for each student $i$ \[ V_ {is} =
\] $V_{is}$ encodes lexicographic priorities in ($Q$, $U$ or $R$), because the integer part of $ (1+\bar q_s)V_ {is}$ is $Q_{is}$ and the fractional part is $R_{it_s}$ or $U_ {i\ell_s}$. The priority ranking for school $s$ is thus \[ i \rhd_s j \iff \bk{(Q_{is}, R_{it_s}) >_ {\text{lex}} (Q_{js}, R_{jt_s}) \text{ or } (Q_{is}, U_{i\ell_s}) >_{\text{lex}} (Q_ {js}, U_{j\ell_s})} \iff V_ {is} > V_ {js}, \numberthis \label{eq:school_priority_v} \] where $i \rhd_s j$ means $i$ is more favorably ranked than $j$ by school $s$.
The matching $D_1,\ldots, D_N$ is then computed by the student-proposing deferred acceptance algorithm with student preferences $\succ_1,\ldots, \succ_N$ and school priorities $\rhd_ {0},\ldots, \rhd_{M}$ GS1962,abdulkadirouglu2003school. The student-proposal deferred acceptance algorithm is well-known; we reproduce it in (ref).
To make the notation concrete, consider the following setup, which is accommodated by these mechanisms.
azevedo2016supply observe that deferred acceptance can be interpreted as computing a vector of cutoffs. Let \[C_{s,N} =
\] be the least-favorable priority score among those matched to $s$, if $s$ is oversubscribed. The matching computed by deferred acceptance is rationalized by each student $i$ being matched to their favorite school among the set of schools for which $V_{is} \ge C_{s,N}$. This cutoff structure generates RD-driven identification.
A key contribution is to clarify which conditional average treatment effects are identified by centralized assignment without restricting treatment effect heterogeneity. To do so, we consider the “smallest” building block of treatment contrasts in this setting, which we call atomic treatment effects. Specifically, for a pair of schools $s_0, s_1 \in S$, define the atomic treatment effect (aTE) as the conditional average treatment effect between this pair of schools for a particular value of observables $X$: \[ \tau_{s_1, s_0} (x) = \E\bk{ Y(s_1) - Y(s_0) \mid X=x}. \] For a school $s$, define the {atomic potential outcome mean} as $ \mu_s(x) = \E[Y(s) \mid X=x]. $
Our identification results characterize for which $(s_1, s_0, x)$ the aTE $\tau_ {s_1, s_0}(x)$ is point-identified under mild assumptions. This exercise reinforces some claims in the literature,\footnote{(ref) shows that the values for which $\mu_s(x)$ is identified under our formulation correspond exactly to those values $(s, x)$ for which the local deferred acceptance propensity score abdulkadiroglu2017impact is positive.} separates the RD-driven and the lottery-driven causal effects, and clarifies that existing estimates of aggregate treatment effects may have poor interpretation. Our results then show how to estimate (certain granular aggregations of) the aTEs, thus enabling users to construct more interpretable estimands. We pause and discuss the motivation and limitation of this exercise.
\Copy{idlitmotiv}{First, knowing which aTEs are identified informs us which aggregate treatment effects are point-identified given only the assignment algorithm. Existing work abdulkadiroglu2017impact,abdulkadirouglu2017research values centralized assignments partly because they provide analogues of quasi-experimental research designs that are popular in other microeconometric settings.\footnote{For instance, abdulkadiroglu2017impact write, “Centralized student assignment opens new opportunities for the measurement of school quality. The research potential of matching markets is enhanced here by marrying the conditional random assignment generated by lottery tie-breaking with RD-style variation at screened schools.”} We establish formal identification and extends it to heterogeneous treatment effects.\footnote{abdulkadiroglu2017impact calculate what they term local deferred acceptance (DA) propensity scores and show in their Corollary 1 that under constant treatment effects, the treatment effect is identified. Our results characterize the set of identified aTEs. These identification results connect to abdulkadiroglu2017impact: We show that $\mu_s (x)$ is identified if and only if the local DA propensity score for school $s$ at characteristics $x$ is positive. See (ref).}} Naturally, the only treatment effect parameters that are point-identified correspond to aggregate treatment effects that place weight only on identified aTEs: i.e., estimands of the form \[ \tau \equiv \int \sum_{s_1 \neq s_0} w(s_1, s_0 \mid x) \cdot \underbrace{\E[Y(s_1)-Y (s_0) \mid X=x]}_{\tau_{s_1, s_0}(x)} dW (x), \numberthis \label{eq:aggregate_TE} \] where $w(s_1, s_0 \mid x)$ and $W(x)$ only put nonzero weight on $\tau_{s_1, s_0} (x)$ that are point-identified. Thus, a practitioner who seeks to only rely on quasi-experimental research designs can inspect the class of identified aTEs and decide the aggregation that is most informative of their substantive economic question.
Second, disaggregating into aTEs helps us understand what practitioners---who may impose constant-treatment-effect assumptions---target under misspecification. For instance, regression estimators for comparing schools---commonly, estimating causal effects of one group of schools $S_1$ to another $S_0 = S \setminus S_1$ by regressing $Y$ on $\sum_ {s\in S_1} D_{is}$ as well as other controls---pool over aTEs between different school pairs and aTEs using different sources of variation. Our analysis disaggregates such estimands. Our subsequent analysis cleanly separates which aTEs are identified through lottery variation and which are through RD variation, collected in the following definition:
\Copy{aggzero}{In disaggregating these estimands, we find that regression estimators---as used by, e.g., abdulkadiroglu2017impact---pool aTEs from both sources in a way that asymptotically puts zero weight on RD-driven aTEs.} This arises because RD-driven estimands use a much smaller set of students than lottery-driven ones, so inverse-variance aggregation leads lottery-driven estimates to dominate. We provide some diagnostics for assessing the severity of this issue in finite samples. This feature also calls for separating out the lottery and RD driven components when aggregating aTEs, for which our estimation results are helpful.
Third, granular aggregations of aTEs can yield more informative, yet still low-dimensional, summaries of school value-added. Each individual aTE is likely too imprecisely estimated to be useful, since it conditions on a single value of $X$. In constructing useful aggregations, one should consider what features plausibly predict heterogeneity. We speculate that, between schools $(s_1, s_0)$, it matters whether $\tau_{s_1, s_0}(X)$ is lottery- or RD-driven, and it also matters whether $X$ contains individuals who report preferring $s_1$ to $s_0$ or vice versa.
Given $s_1$ and $s_0$, for every student with $s_1 \succ s_0$, there is a maximal identified aggregation \[\tau_{s_1 \succ s_0} = \int \tau_ {s_1,s_0}(x) dW_ {s_1\succ s_0} (x),\numberthis \label{eq:maximal}\] where the measure $W_ {s_1\succ s_0}$ is the conditional distribution of $X$ given that $s_1 \succ_i s_0$ and that $\tau_{s_1, s_0} (X)$ is identified.\footnote{Our identification results produce a set $\mathcal I_{s_1, s_0}$ of $X = (\succ, R, Q)$ values such that $x \in \mathcal I_{s_1, s_0}$ if and only if $\tau_ {s_1, s_0}(x)$ is point-identified. $W_{s_1 \succ s_0}$ is then the conditional distribution $(X \mid X \in \mathcal I_{s_1, s_0}, s_1 \succ s_0)$. For RD-driven aTEs, the set $\mathcal I_{s_1, s_0}$ has zero measure under $P$, and thus we take additional care to define the conditional distribution. See (ref) for defining general aggregate treatment effects, for which $\tau_{s_1 \succ s_0}$ is a special case. } $\tau_ {s_1 \succ s_0}$ does not aggregate across identifying variation: It is either RD-driven or lottery-driven depending on whether $s_1$ is a lottery school, per (ref). (ref) provides asymptotic guarantees for estimating (ref).
Practitioners can further summarize estimates $\hat \tau_ {s_1 \succ s_0}$ into school value-added by positing a Bradley--Terry-style firth2005bradley model in which, for instance, the pairwise treatment effects are explained by differences in school value-added, adjusted for reported preference and source of identification \[ \hat\tau_{s_1 \succ s_0} = \alpha_0 + (\mu_ {s_1}^L - \mu_ {s_0}^L)\one (\text{$s_1$ is lottery}) + (\mu_{s_1}^R - \mu_{s_0}^R)\one(\text{$s_1$ is test-score}) + \epsilon_ {s_1, s_0} \numberthis \label{eq:BT_aggregation}. \] Here, $\mu_s^L, \mu_s^R$ represent school value-added among lottery- and RD-based comparisons, respectively, and $\alpha_0$ adjusts for the fact that $s_1$ is revealed to be preferred to $s_0$. These parameters may be estimated via least-squares (normalizing some effects to zero) or further regularized under random effects or hierarchical Bayesian-type assumptions. Under constant effects ($Y(s) = Y(0) + \beta_s$ almost surely), this model is exactly well-specified, and $\mu_ {s_1}^L - \mu_ {s_0}^L = \mu_{s_1}^R - \mu_{s_0}^R$ recovers $\beta_{s_1}-\beta_{s_0}$. Without such an assumption, $\mu_ {s}^L, \mu_s^R$ can be interpreted as least-squares summaries of school causal effects. This cleanly separates identification of causal effects from their aggregation and summary.
\Copy{contlarge}{ \Copy{cont}{This exercise has limitations. First, as we shall require in (ref), identification of RD-driven aTEs relies on continuity of $\mu_s(x)$ in test scores at cutoffs hahn2001identification. marinho2022causal note that this continuity assumption restricts strategic misreporting of preferences. Suppose students know the admission cutoffs and can strategically report preferences based on their test scores or lottery numbers. With strategic reporting, those who report $\succ$ just below a cutoff may have very different true preferences from those who report $\succ$ just above a cutoff. The mix of students can thus change sharply at the cutoff, possibly making conditional expectations discontinuous.
For strategic misreporting to be a concern for continuity, though, students must be able to reliably forecast cutoffs. Prediction errors would smooth over the mix of students above and below a cutoff, potentially restoring continuity.\footnote{ For instance, in many settings (e.g. New York City che2023leveraging and the Chinese college entrance exam chen-kesten-chinese), students submit preferences before learning their test scores, rendering accurate manipulation difficult. In their empirical application, marinho2022causal do show evidence that students manipulate reported preferences: Their Figure 1 shows that students are more likely to rank school $j$ more than the outside option if their running variable is closer to $j$'s cutoff, but this probability does not appear to change discontinuously at the cutoff.} Moreover, such manipulation plausibly leads to discontinuities in the density of various covariates at the cutoff, and can in principle be empirically assessed.
Second, even if continuity holds, the aTEs may still have limited external validity. They are useful for the types of program evaluation exercise in abdulkadirouglu2017research,abdulkadiroglu2017impact, but are perhaps less directly informative of counterfactuals in which school choice mechanisms change or school admissions policies change. The aTEs condition on reported preferences, test scores on existing tests, and are only identified for certain subpopulations. If any of them change in a counterfactual policy environment, additional extrapolative assumptions are needed. We refer readers to marinho2022causal for assumptions and methods that partially identify conditional average treatment effects that condition on true preferences.}
Despite this caveat, the maximal aggregation of aTEs ((ref)) does have a clear policy interpretation: It measures the effects of policies on the margin. Consider a small increase in $s_1$'s capacity and assume such a change does not alter preference submission. Such an increase diverts some students from enrolling in $s_0$ as they now qualify for $s_1$, which they prefer. This aggregate treatment effect ((ref)) between $s_1, s_0$ exactly measures the gain of these students in terms of the outcome $Y$. Thus, for marginal changes, such as increasing the capacity of school $s_1$ by a small amount or allocating marginal resources to $s_1$, such aggregations of aTEs are plausibly informative of the impact of such policies; see arkhangelsky2025evaluating for a similar argument. Indeed, abdulkadiroglu2017impact are motivated by debates in New York City over whether students have adequate access to Grade A schools. This policy question can be viewed as evaluating marginal expansions of Grade A schools; if the impact is large, current access is suboptimally limited.}
So far, we have discussed identification heuristically. Because the school assignments $D_i$ depend jointly on cutoffs $C_N$, themselves computed from all characteristics $ (X_1,\ldots, X_N)$, $(Y_i, D_i, X_i)$ are not i.i.d. conditional on $C_N$. Thus, the usual cross-sectional notion of identification does not apply.\footnote{A standard definition states that a quantity is identified if no two observationally equivalent distributions of observable data and potential outcomes give rise to different values of the quantity. Here, “distributions” refers to the joint distribution of the observable data, school assignments, and potential outcomes for the finite sample of $N$ students. However, vanishing variation of $C_N$ around $c$ would allow for certain quantities to be “identified” per the standard definition, but not consistently estimable. We give a simple example in (ref) to illustrate the conceptual difficulties.} The literature abdulkadiroglu2017impact,marinho2022causal often appeals to large market asymptotics and treats the $C_N$ as nonrandom instead in identification analysis.\footnote{Notable exceptions include agarwal2018demand and munrojmp.} We conclude this section by making this heuristic precise, and we take $C_N$ as random in our estimation section.
When students are drawn from a population and the number of students is large, the random cutoffs $C_{s,N}$ concentrate around some population quantity, satisfying a law of large numbers. Proposition 3 of azevedo2016supply states that if $ \br{(\succ_{is}, V_ {is}) : s \in S}$ are i.i.d. across students, then under mild conditions,\footnote{Precisely speaking, azevedo2016supply define a notion of deferred acceptance matching acting on the continuum economy---which is the distribution of $\br{(\succ_{is}, V_ {is}) : s \in S}$. The additional regularity condition is that this continuum version of deferred acceptance matching admits a unique stable matching. In other words, it is a very mild regularity condition on the distribution of $\br{(\succ_{is}, V_ {is}) : s \in S}$.} the cutoffs $C_N = [C_ {0,N},\ldots, C_{M,N}]$ concentrate to some population counterpart $c$ at the parametric rate: $ \norm{C_N - c}_\infty = O_P(N^{-1/2}). $
We define identification relative to the population cutoff $c$. For a fixed cutoff $c$, we can define $D_i^*(c)$ as the (fictitious) assignment $i$ would receive if the cutoffs were set to $c$. The tuple $(X_i, D_i^*(c), Y_i (0),\ldots, Y_i(M), U_i)$ are then i.i.d., and identification reduces to its definition in standard settings. Since as $C_N \to c$, $D_i^*(C_N) \to D_i^*(c)$ for $N \to \infty$, this notion of identification captures the information content of the data for large markets.
Formally, let $P \in \mathcal P$ be the distribution of student characteristics, potential outcomes, and lottery numbers $(X_i, Y_i(0),\ldots, Y_i(M), U_i)$, where we assume every member of $\mathcal P$ satisfies (ref) azevedo2016supply. Let $c(P)$ be the corresponding large-market cutoffs, defined as the probability limit of $C_{N}$ when data is sampled according to $P$. Define $D_{i}^*(c)$ as the assignment made if the cutoffs were set to $c$: i.e. $D_ {is}^*(c) = 1$ if and only if $s$ is $i$'s favorite school among those with $V_{is} \ge c_s$. Let $Y^*_{i,\text{obs}}(c) = \sum_{s\in S} D_{is}^*(c) Y_i(s) $ denote the observed outcome under $D_i^*$.
We begin with a simplified setting that contains most of the intuition for our formal results.
For this setup, we can write $X_i = (\succ_i, R_i)$ and omit $Q_{is}$. Note that $\tau_{s,s'}(\succ, R)$ is identified if and only if both $\mu_{s}(\succ, R)$ and $\mu_{s'}(\succ, R)$ are. For a given preference $\succ$ and a school $s$, consider the $s$-eligibility set \[ E_{s}(\succ, c) = \br{ r: \P(D^*_{s} (c) = 1 \mid {\succ}, R=r) > 0 }. \numberthis \label{eq:s_elig_def} \] $E_s (\succ, c)$ collects the test scores $r$ such that someone with $X = (\succ, r)$ has positive probability of being matched to school $s$. As a result, $\mu_s(\succ, r)$ is identified if $r \in E_s(\succ, c)$.
Computing these sets for each type yields:
To illustrate the computation, consider students with preference $\succ_A$. For a given value of $R=r$ and a school $s$, we ask whether $r \in E_s(\succ_A, c)$:
Our subsequent results make the above calculation systematic.
If we further assume that $r\mapsto \mu_s(\succ, r)$ is continuous in $r$ for every preference and school $s$, then we can extend identification to the closure of $E_{s_j}(\succ, c)$ in $[0,1]^T$. Continuity assumptions allow us to take sequences and extend identification to their limits: $\mu_s(\succ, r_k) \to \mu_s(\succ, r)$ if $r_k \to r$. Thus, under continuity of the atomic potential outcome means, we can compute regions on which each aTE $\tau_{s,s'}(\succ, R)$ is identified: For a pair of schools $(s, s')$ and preference type in $\br{A,B,C}$, we tabulate the set of $R$ values for which $\tau_{s,s'} (\succ, R)$ is identified. The following table tabulates $\bar E_{s} \cap \bar E_{s'}$ for choices of $s, s'$ and $\succ$:
This calculation shows that the interpretation of aggregate treatment effects is complex in two senses, a complexity often masked by simple aggregations.
First, aggregate treatment effects may mask heterogeneity in terms of which pairs of schools are compared, since different pairs of schools $(s_i, s_j)$ are comparable on different regions of the test score $R$. For instance, we might wish to estimate the treatment effect of being enrolled in an inside option ($s_1,s_2,s_3$) relative to the outside option $s_0$. Naturally, corresponding aggregate treatment effects are some weighted average of the identified aTEs $\tau_{s_1, s_0}, \tau_{s_2,s_0}, \tau_ {s_3, s_0}$. In this case, interpreting these aggregates as general inside-versus-outside effects overstates their generality:
Therefore, in this case, interpreting aggregate treatment effects as an inside-versus-outside option effect drastically overstates the generality of these estimands in the presence of heterogeneous effects.
Second, aggregate treatment effects may mask heterogeneity in terms of which students are compared for a particular pair of schools, since regions of $R$ that admit comparisons for a given pair of schools differ substantially across student types. This is true in this example for comparisons between $(s_2,s_3)$:
In this case, no identified aggregations of aTEs between $s_2$ and $s_3$ take into account students of preference type $B$. Moreover, the maximal identified aggregation of aTEs between $s_2$ and $s_3$ that weighs each student equally is the $s_2$-against-$s_3$ effect among students of type $C$ with high test scores: $\E[Y(s_2) - Y(s_3) \mid {\succ_C}, R\in [2/3,1]]$, since the set of students of type $A$ with $R = 2/3$ is a measure-zero set. Again, interpreting these estimands as blanket causal comparisons between $s_3$ and $s_2$ assumes that treatment effects for other students are similar to those of type $C$ with test scores in $[2/3, 1]$.
We should expect the heterogeneity in both senses to be even more complex in general, since this example only includes 3 out of the 24 possible preferences and only a single type of test score. Given this heterogeneity, aggregations that pool over many schools, preference types, and test scores can obscure the implicit weights on aTEs. To understand these estimates, the next subsection characterizes the eligibility sets $E_{s}$ as well as pairwise treatment contrasts formally.
\Copy{idintro}{ Like the example above, characterizing identification of atomic treatment effects amounts to computing the $s$-eligibility sets ((ref)). Our main result, (ref), computes them in the general setting. The behavior of intersections of $s$-eligibility sets formalizes the distinction between aTEs that are driven by lottery variation and those driven by RD variation. (ref) then shows that any aggregations of aTEs that aggregate over a non-vanishing subpopulation must place zero weight on the RD-driven aTEs.
In the general setting, recall that $X_i = (\succ_i, R_i, Q_i)$, where we let $R_i$ collect the test scores $(R_ {i1},\ldots,R_{iT})$ and $Q_i$ collect the discrete qualifiers $(Q_{i0},\ldots, Q_{iM}).$ Fix $(\succ_i, Q_i)$, we first define the $s$-eligibility sets as the set for $R_i$ on which assignment to $s$ has positive probability.
}
(ref) assumes that atomic mean potential outcomes are continuous in the test scores (and that moments exist). This is a standard assumption in regression discontinuity hahn2001identification, but it does imply restrictions on strategic misreporting of preferences, as we discuss in (ref). In so far as (ref) holds---intuitively, this is plausible when students fail to accurately forecast cutoffs so that student type mixes do not change abruptly at a cutoff---our setting accommodates mechanisms beyond standard deferred acceptance, since these other mechanisms can be represented as deferred acceptance on transformed preferences abdulkadiroglu2017impact.
The next proposition verifies that the $s$-eligibility sets $E_s(\succ, q, c)$ collect values of $r$ such that $\mu_s (\succ, q, r)$ is identified. Continuity ((ref)) extends identification to the closure of the $s$-eligibility sets.
Our main result is the following characterization of all identified atomic treatment effects: We compute the $s$-eligibility sets for both lottery and test-score schools, and find the intersection of closures of $s$-eligibility sets between two schools.
Fix some school $s_0$ and some student with characteristics $x=(\succ, r, q)$. $\mu_{s_0} (x)$ is identified if such a student has positive probability of being assigned to $s_0$. Suppose $s_0$ is a lottery school. Naturally, in order for a student to have positive probability to be assigned $s_0$, we require that
On the other hand, suppose that $s_0$ is a test-score school that uses test $t_0$. Then the corresponding requirements for identification of $\mu_{s_0}(x)$ are:
It is useful for our estimation results to additionally define $r_{s,t}(c)$ as the unique value among $ \br{ \smash{\underline{r_{t}}} (s, q, c_s): q = 1,\ldots, \bar q_s}$ that is in $(0,1)$; if no such value exists, then set $r_ {s,t} (c) = 0$: \[ r_{s,t} (c) = \max\pr{\br{\smash{\underline{r_{t}}}(s, q, c_s): q = 1,\ldots, \bar q_s} \cap [0,1)}. \numberthis \label{eq:test_score_space_cutoff} \] Intuitively, (ref) translate cutoffs in $V_s$-space to cutoffs in $R_{s,t_s}$-space for students of different discrete priority types ($Q_{is}$), as illustrated in (ref).
Finally, the region on which the aTE $\tau_{s_1, s_0}(x)$ is identified is an intersection between regions on which $\mu_{s_j}(x)$ is identified. The latter region turns out to involve hyperrectangles in $R_{i}$. To compactly describe the intersection, it is useful to define intersecting on coordinate $t$: Given $S_1 \subset [0,1]^T$ and $S_2 \subset [0,1]$, let $ \mathrm{Slice}_t(S_1, S_2) = \br{s \in S_1 : s_t \in S_2}$ be the set that takes the intersection of the $t$\th coordinate of $S_1$ with $S_2$. The following figure illustrates slicing an ellipse $S_1 \subset [0,1]^2$ on the first dimension onto an interval $S_2$:
The following theorem formalizes this intuition and computes the regions on which $\tau_ {s_1, s_0}(x)$ is identified.
The first claim of (ref) characterizes the eligibility set for a lottery school $s_0$. The second claim analogously characterizes the eligibility set for a test-score school $s_0$. Both formalize the verbal intuition we have described. The third claim computes the intersection of the closures of the eligibility sets between two schools $s_0, s_1$. Depending on whether the more-preferred school $s_1$ is a lottery school, the intersection takes different shapes. If $s_1$ is a lottery school, then the intersection is either empty or $\bar E_{s_0}$. Otherwise, the intersection is either empty or the slice of $\bar E_ {s_0}$ that equals $s_1$'s cutoff on test $t_1$.
\Copy{idlit}{These identification results closely relate to the literature. They follow abdulkadiroglu2017impact closely and extend their Corollary 1 to settings with heterogeneous treatment effects. In particular, the set of students for whom $R \in \bar E_s(\succ, Q; c)$ is precisely the set of students for whom the local DA propensity score for school $s$ is positive. Separately, when there are only test-score schools and no discrete qualifies ($Q_{is} = 0$), (ref) characterizes the same set---up to a conditionally measure zero set of students---of identified pairwise RD effects as marinho2022causal's Proposition 1. See (ref) for details.
Despite (ref) identifying the same set of effects as shown elsewhere, disaggregating these effects clarifies their interpretation and highlights potential drivers of heterogeneity. In particular, (ref)(3) rationalizes what we preview in (ref) as lottery- or RD-variation. When the more-preferred school is a lottery school, variation between $s_0$ and $s_1$ is driven by the student potentially losing the lottery at $s_1$. On the other hand, when the more-preferred school is a test-score school, variation between $s_0$ and $s_1$ is driven by the RD variation in whether the student just qualifies for $s_1$ or just fails to qualify for $s_1$. Therefore, the two types of aTEs are different in economically meaningful ways: There is no lottery variation between a less-preferred lottery school and a more-preferred test-score school, and no test-score variation between a less-preferred test-score school and a more-preferred lottery school.}
Examining the disaggregated aTEs in turn illustrates certain perils of simple aggregation schemes. (ref)(3) shows that aTEs driven by lottery variation are identifiable for a set of students with positive measure, but aTEs driven by test-score variation ((ref)) are only identifiable for a set of students with zero measure. As a result, when we aggregate treatment effects such that each student receive equal weights (see (ref) for details), we would effectively only aggregate the lottery-driven aTEs. We state this result informally as (ref), which is formally stated and proved as (ref).
(ref) is a concerning observation for interpreting estimates of aggregate treatment effects, as many of these estimands necessarily put zero weight on the aTEs driven by the RD variation. In contrast, our recommended aggregation (ref) avoids this feature.
Translated to estimation, (ref) implies that estimators for aggregate treatment effects put vanishing weight on the aTEs driven by RD variation, since these estimators can only take comparisons local to a cutoff. While in finite samples these weights can still be positive and nontrivial, these implicit weights do depend on the sample size (and on bandwidth tuning parameters); moreover, the larger these weights are (equivalently, the larger the bandwidth), the more susceptible to bias the estimator is. The share of influence from RD-driven aTEs is larger in smaller samples than in larger samples, which makes such aggregate treatment effects potentially difficult to interpret. Indeed, the vanishing weight issue affects popular estimation approaches proposed in the literature.\footnote{\Copy{fnyata}{These issues are not unique to aggregation in the school-choice context. In settings---for instance studied in narita2021algorithm---where the distribution of treatment assignment is a function of $X$, under continuity assumptions, RD-type treatment effects are identified on the boundary of sets of the form $\br{X : \P(D=1\mid X) > 0}$. Similarly, aggregating these effects with effects in the interior may place vanishing weight on these RD-type effects, since the boundary has zero measure.}}
The next subsection returns to our extended example and illustrate that the regression estimator in abdulkadiroglu2017impact puts vanishing weights on the RD-driven aTEs. To be sure, these estimators can be simple to implement and more efficient when treatment effects are homogeneous goldsmith2021estimating,angrist1998estimating; motivated by these advantages of regression, we also develop a simple diagnostic for the weight put on RD-driven aTEs for linear estimators.
The estimation approach proposed by abdulkadiroglu2017impact converges to aggregations of treatment effects that ignore the RD variation in the limit. We illustrate this with our extended example in (ref). First, we make a few additional assumptions on the example.
Roughly speaking, for a region of test score $A \subset [0,1]$, abdulkadiroglu2017impact compute $\P(\tilde D_i = 1 \mid {\succ}, R \in A)$ by counting those $R$'s near a cutoff as having $1/2$ probability of falling on either side. “Near a cutoff” is determined by a bandwidth parameter $h > 0$. To be more concrete, we partition the space of test scores into five regions, with a bandwidth parameter $h > 0$. Regions II and IV are small bands around the cutoffs $c_1, c_2$, and regions $\text{I}, \text{III},\text{V}$ are large regions in between the cutoffs:
The local deferred acceptance propensity scores, as a function of the region that the test score $R$ falls into, are as follows:\footnote{To explain the propensity score calculation, consider a student of type $B$ with test scores in II:
This computation is heuristic with $h > 0$, but as we take the limit $h \to 0$, the probability $\P(\tilde D_i =1 \mid \text{II}, {\succ_B}) \to 0.25$. }
Let $v \in \br{A, B, C} \times \br{\text{I}, \text{II}, \text{III},\text{IV}, \text{V}}$ denote a student type, according to preferences and coarsened test scores. Let $\psi(v)$ denote the corresponding local propensity score collected in the above table. Consider the population regression\footnote{If $\psi(V)$ is exactly equal to $\E[D \mid V]$, then this regression is equivalent to the more familiar control-for-propensity-score regression angrist1998estimating \[ Y = \alpha + \tau \tilde D + \gamma \psi(V) + \epsilon \] by Frisch--Waugh. This latter specification conforms with the specification used for intent-to-treat effects in abdulkadiroglu2017impact (i.e., the reduced form in their IV specification).} \[Y = \tau (\tilde D - \psi(V)) + \epsilon \numberthis \label{eq:ols_spec} \text{ such that } \tau = \frac{\E[(\tilde D - \psi(V)) Y]}{\E[(\tilde D - \psi(V))^2]}. \] We compute in (ref) that
where $\text{Bias}(h) \to 0$ as $h \to 0$.\footnote{We note that because of the regression specification ((ref)), the estimand weighs according to $(1-\psi (v))\psi(v)$. As a result, this aggregation does not weigh each student equally, but nevertheless the conclusion of (ref) continues to hold for such weighting schemes.}
(ref) shows that the implicit aggregation recovers weighted averages of treatment effects (over preference-test score region cells) up to a bias that vanishes as $h \to 0$. However, the weights on regions $\text{II}$ and $\text{IV}$ also vanish as $h \to 0$, since the corresponding $\P(v) \to 0$. Translated to estimation, this implies that the asymptotic unbiasedness of propensity-score estimators requires $h = h_N \to 0$ at appropriate rates, yet the part of the estimator driven by variation from regression discontinuity then becomes asymptotically negligible, as long as there is lottery-driven variation.
To further relate to finite samples, note that under typical assumptions---as confirmed by our estimation results in (ref)---the popular locally linear regression estimator for regression-discontinuity-type variation converges at the rates no faster than $N^{-2/5} \gg N^{-1/2}$, reflecting that identification for the conditional average treatment effect at the cutoff is irregular khan2010irregular. As a result, asymptotically, estimators for the RD aTEs are much noisier than those for the lottery-driven aTEs, and pooling them with inverse-variance-type weights in ((ref)) results in diminishing weight on the RD aTEs.
Motivated by this decomposition, we introduce a simple diagnostic for regression-based procedures that gives upper and lower bounds on the weight put on students who are subject to RD variation. For simplicity, we limit to considering binarized comparisons. That is, there is some treatment $\tilde D_i = \sum_{s \in S_1 \subset S} D_ {is}$ corresponding to a subset $S_1$ of the schools $S$, where a student is considered treated if they are matched to a school in $S_1$ and untreated otherwise.
We consider linear estimators of the form \[ \hat\tau = \sum_{i=1}^n \hat w_i \cdot Y_i, \] where $\hat w_i$ is a function of $X_{1},\ldots, X_N, \tilde D_1,\ldots, \tilde D_N$. Any linear regression estimator with $Y_i$ on the left-hand side and functions of $X_i, \tilde D_i$ on the right-hand side can be written this way. We will also assume the estimator is associated with some chosen bandwidth parameter $h_N$.
Using the characterization in (ref), we label student observations by whether they are subject to lottery variation:
Intuitively, those with $\overline{\mathrm{RD}}_i = 1$ are individuals who are close enough to the cutoff used by $s_1$ such that, on some lottery realizations, they would be assigned to a test-score school $s_1$ if they clear the cutoff, and to some school $s_0$ of the opposite treatment type otherwise. Those with $\underline{\mathrm{RD}}_i = 1$ are those who do so on all lottery realizations.
We can then define \[ \bar p_{\text{RD}} \equiv \frac12 \sum_{i=1}^n \hat w_i (2 \tilde D_i - 1) \bar{\mathrm{RD}}_i \text{ and } \underline p_{\text{RD}} \equiv \frac12 \sum_{i=1}^n \hat w_i (2 \tilde D_i - 1) \underline{\mathrm{RD}}_i \] as the upper and lower bounds for the weight assigned to the RD variation, and these can be implemented by using $(2 \tilde D_i - 1) {\mathrm{RD}}_i$ as left-hand side variables in the regression. Formally, these estimates are interpreted as the change in $\hat\tau$ if every treated student with $\bar{\mathrm{RD}}_i = 1$ (resp. $\underline{\mathrm{RD}}_i = 1$) has their outcome increase by 1/2 unit, and every untreated student with $\bar{ \mathrm{RD}}_i = 1$ has their outcome decrease by 1/2 unit.\footnote{Our assumptions on linear estimators do not rule out estimators that overly extrapolate (i.e. some $\hat w_i$ has the wrong sign: It is negative for $\tilde D_i = 1$ and positive for $\tilde D_i = 0$). As a result, it is possible that one or both of $\bar p_{\text{RD}}$ and $\underline p_{ \text{RD}} $ are negative, or that $\bar p_{\text{RD}} < \underline p_{\text{RD}} $. These unpleasant realizations serve as an additional diagnostic. They cast doubt on the interpretation of the linear estimator $\hat\tau$, as they reveal that a subpopulation chosen solely on the basis of $X_i$ has weights that are wrong-signed. See chen2025potentialweightsimplicitcausal for general diagnostics in regression estimators. }
abdulkadiroglu2017impact study the impact of attending a “Grade A school” (one that receives grade A on the school district's report card for school quality) in New York City versus attending a non-Grade A school. While we do not have access to their data, abdulkadiroglu2017impact (Appendix Figure D.1) report that about 9,000 students out of 32,866 students have local propensity scores exactly equal to $1/2$, under their bandwidth choice $h_N$. This means that these students are solely subject to RD variation between the treatment schools $S_1$ (Grade A schools in New York City) and the control schools. Correspondingly, these individuals would be definitely subjected according to (ref). Thus, this puts a lower bound of about $9,000/32,866\approx 0.3$ on $\underline p_{\text{RD}}$ for their empirical application.\footnote{$9,000/32,866\approx 0.3$ would be the weight on these observations if every observation were weighted equally. Since the specification ((ref)) weighs students proportionally to $(1-\psi)\psi$, those with propensity score at exactly $1/2$ receive the highest weights. Thus, in this case, 0.3 serves as a lower bound.} Assuming that the bias term is sufficiently small, we may conclude that abdulkadiroglu2017impact estimate aggregations of aTEs that put nontrivial weight on the RD-driven aTEs.
In general---though especially when $\bar p_{\text{RD}}, \underline p_{\text{RD}}$ imply unreasonably small weight on the RD-driven aTEs---researchers may wish to unpack heterogeneity further and isolate aTEs that are driven by RD variation, along the lines in (ref). If researchers have in mind weights for an aggregate treatment effect that they prefer, they can also aggregate these finer treatment effects manually. The next section discusses estimation and inference for maximal aggregations atomic treatment effects between two schools ((ref)).
We close this section with two miscellaneous discussions.
Having characterized the identified atomic treatment effects, we advocate that practitioners first estimate more granular aggregations of aTEs and then summarize these aggregations further. Towards this goal, this section provides asymptotic theory for the maximal aggregation between two schools ((ref)).
To that end, consider two schools $s_1, s_0$ and all students who declare $s_1 \succ_i s_0$. We are interested in the estimand \[ \tau_{s_1 \succ s_0}(c) = \E[Y(s_1) - Y(s_0) \mid R \in \bar E_{s_0}(\succ, Q; c) \cap \bar E_{s_1}(\succ, Q; c), s_1 \succ s_0]. \] Constructing estimators for $\tau_{s_1 \succ s_0}$ depends on whether $s_1$ is a lottery school, but they share some common structure.
For a given set of cutoffs $c$, let $J_i(c) = 1$ indicate the set of $X$ values such that
We can thus write \[ \tau_{s_1 \succ s_0}(c) =
\numberthis \] recalling (ref). $\tau_{s_1\succ s_0}(c)$ corresponds to the maximal aggregation of aTEs for $s_1\succ s_0$ when the population cutoffs equal $c$.
To construct analogue estimators for $\tau_{s_1\succ s_0}$, we account for the fact that---depending on the lottery results---students with $J_i(c) = 1$ may be matched to neither $s_1$ nor $s_0$, e.g., if they qualify for some lottery school they prefer. Let $D_ {i1} (U_i; c)$ indicate the event---as a function of the lottery numbers $U_i$---that student $i$ fails to qualify for any lottery school $s \succ_i s_1$ and---if $s_1$ is lottery---additionally qualifies for $s_1$. Let $D_ {i0} (U_i; c)$ indicate the event that student $i$ fails to qualify for any lottery school $s\succ_i s_0$ and---if $s_0$ is lottery---additionally qualifies for $s_0$. Likewise, let $\pi_{ij}(c) = \int D_ {ij} (u; c)\,dF_U(u)$ be the fixed-cutoff expectation of these events.\footnote{These objects are defined formally in (ref).} By construction, for all fixed cutoff $c$, we have that $\E\bk{ J_i(c) \frac{D_{i1}(U_i; c) Y_i}{\pi_{i1}(c)} } = \E\bk{J_i(c) Y_i(s_1)} $ and that $\E\bk{ J_i(c) \frac{D_{i0}(U_i; c) Y_i}{\pi_{i0}(c)} } = \E\bk{J_i(c) Y_i(s_0)}$. Denote $Y_i^{(j)}(c) \equiv \frac{D_{ij}(U_i; c) Y_i}{\pi_{ij}(c)} $.
We thus have two natural estimators for $\tau_{s_1\succ s_0}(c)$, depending on whether $s_1$ is a lottery school. If $s_1$ is a lottery school, then an inverse propensity estimator takes the form \[ \hat\tau_{s_1 \succ s_0} = \pr{\frac{1}{N}\sum_{i=1}^N J_i(C_N)}^{-1}\pr{\frac{1} {N}\sum_ {i=1}^N J_i (C_N) \br{Y_i^ {(1)}(C_N) - Y_i^{(0)}(C_N)}}. \numberthis \label{eq:ipw_lottery} \] (ref) is simply the inverse-weighting estimator among those with $J_i(C_N) = 1$.
\Copy{llr}{On the other hand, if $s_1$ is a test-score school and we shorthand $\rho(c) = r_ {s_1, t_1}(c)$, then (ref) is akin to an RD estimand among those with $J_i(c) = 1$. We consider the analogue of local linear regression for this estimand. For technical reasons, we limit our consideration to a uniform kernel. Specifically, fix some bandwidth $h_N \to 0$, we let \[ \hat\tau_{s_1 \succ s_0} = \hat\beta_+(h_N) - \hat \beta_-(h_N) \numberthis \label{eq:llr_estimand} \] where, for $J(c,h)$ defined in (ref),
The estimator for the left-limit, $\hat \beta_-(h_N)$, is defined analogously. (ref) implements local linear regression among those with $J_i(C_N, h_N) = 1$, which essentially indicates those with $J_i(C_N) = 1$ and have test $R_{it_1}$ within bandwidth $h_N$ of the cutoff $\rho(C_N)$. We focus on local linear regression since it is a standard choice in the regression-discontinuity literature cattaneo2022regression. We speculate that the same analysis likely extends to local polynomial regression with a uniform kernel as well.}
\Copy{estimation}{ Relative to standard setups, estimation is complicated by the fact that the estimators feature the finite-sample cutoffs $C_N$, computed from all the data. Since $C_N$ enters all terms in sample averages, these terms are no longer independent, precluding a standard asymptotic argument. Our asymptotic results account for the effect of a stochastic $C_N$; they can be understood as applications of two-step GMM analyses where we verify certain stochastic equicontinuity conditions in $C_N \to c$. Interestingly, the asymptotic distributions of the estimators do not depend on the asymptotic distribution $\sqrt{N}(C_N - c)$, meaning that, for instance, we do not need to adjust for the fact that $C_N$ is random in calculating standard errors. Our results thus justify procedures in the literature that treat $C_N$ as fixed.\footnote{The sense in which they are valid is nuanced. See the discussion after (ref). }}
Having introduced the estimator, we turn to assumptions. The key assumption is a substantive restriction on school capacities and the population distribution of student characteristics, such that the large-market cutoffs are not in certain knife-edge configurations. This does limit the uniform validity of our asymptotic results.
(ref) rules out populations where the large-market cutoffs from azevedo2016supply are on the boundary of certain sets. The first assumption simply says that the school $s_1$ is not undersubscribed and not strictly easier to qualify for than $s_0$. The second assumption rules out a scenario where everyone with $Q=q$ does not qualify for $s$ regardless of their tiebreakers and everyone with $Q=q+1$ does qualify for $s$ regardless of their tiebreakers. The third assumption says that undersubscribed schools are eventually undersubscribed, for which it suffices to impose that the population capacity of a school is not exactly at the threshold making the school undersubscribed.\footnote{This assumption is stronger than the $O_p(N^{-1/2})$ convergence of the cutoffs that we assume. However, the assumption holds generically for sufficiently large population school capacities. Precisely speaking, suppose some school $s$ with population capacity $q_s$ is undersubscribed in the population, but the probability that it is undersubscribed in samples of size $N$ does not tend to one (and so violates the third assumption). Then we may add a little slack to the school capacity---for any $\epsilon > 0$, making the capacity $q_s + \epsilon$ instead---to guarantee that the third assumption holds. Intuitively, adding $\epsilon$ to the capacity adds $O(\epsilon N)$ seats to the school in finite samples, but the random fluctuation of the market generates variation in student assignments of size $O (\sqrt{N}).$} Lastly, the fourth assumption assumes that the population cutoffs in test-score space are not exactly the same for two schools that uses the same test.\footnote{This is Assumption 2 in abdulkadiroglu2017impact.} Collectively, these assumptions rule out adversarial scenarios where, for instance, the population intersection is empty, $\bar E_ {s_0} ({\succ_i}, Q_i, c) \cap \bar E_{s_1} ({\succ_i}, Q_i, c) = \varnothing$, but the sample intersection is nonempty with probability non-vanishing in $N$, $\P\bk{\bar E_ {s_0} (\succ_i, Q_i, C_N) \cap \bar E_{s_1} (\succ_i, Q_i, C_N) \neq \varnothing} \not\to 0$.
Additionally, we maintain a few technical assumptions, (ref), stated in (ref). These assumptions assert that the distribution of test scores and lotteries are suitably smooth, have suitably smooth conditional means, and have bounded moments.
We have the following result for lottery comparisons.
(ref), proved in (ref), shows that the scaled estimation error $ \sqrt{N} \pr{\hat\tau_{s_1 \succ s_0} - \tau_{s_1\succ s_0}(C_N)}$ is equivalent to a scaled sample mean of i.i.d. random variables ((ref)), which attains a central limit theorem ((ref)). The influence function representation ((ref)) is conducive to deriving joint convergence statements ((ref)).
One subtlety here is that the distribution of $\hat\tau_{s_1 \succ s_0}$ is centered at the random estimand $\tau_{s_1 \succ s_0} (C_N)$, corresponding to the causal effect holding fixed the cutoff at the random value $C_N$, rather than at the population cutoff $c$, $\tau_ {s_1 \succ s_0}(c)$. The difference between the two estimands is of order $O (1/\sqrt{N})$, $ \sqrt{N} (\tau_ {s_1 \succ s_0}(C_N) - \tau_{s_1 \succ s_0} (c)) = O_P(1)$, and cannot be ignored. Controlling this difference is feasible since the asymptotic distribution $\sqrt{N}(C_N - c)$ can be analyzed along the lines in agarwal2018demand, and we leave it to future work. It is reasonable to consider $\tau_{s_1\succ s_0}(C_N)$ as the target of inference, since whether a student chooses between $s_1$ and $s_0$ is ultimately a function of $C_N$ rather than of $c$. Doing so has the additional convenience that the asymptotic distribution does not depend on the limit distribution of $\sqrt{N}(C_N -c)$---and thus standard errors do not need to be adjusted.
Analogously, we have the following result for RD-type comparisons. To facilitate its statement, let $ \check\tau_ {s_1\succ s_0} = \check\beta_+(h_N) - \check \beta_-(h_N) $ where
and similarly for $\check\beta_-$. The local linear regression estimator $\check\tau_ {s_1\succ s_0}$ uses the population cutoff $c$ to construct $J_i$ and $\rho(c)$, and thus its asymptotic properties are well-understood hahn2001identification.
The proof of (ref) is relegated to (ref) ((ref)).
Like (ref), (ref) shows that $ \sqrt{Nh_N} (\hat \tau_{s_1\succ s_0} - \tau_{s_1\succ s_0}(c))$ is asymptotically equivalent to a quantity that is easier to analyze, $\sqrt{Nh_N} (\check\tau_ {s_1\succ s_0} - \tau_{s_1\succ s_0}(c))$. A central limit theorem---standard in the RD literature hahn2001identification---on this quantity is then used to derive the asymptotic distribution. Compared to (ref), we state this theorem by centering at $\tau_{s_1\succ s_0} (c)$. However, since we scale by $\sqrt{Nh_N} \ll \sqrt{N}$, the difference between the estimands is immaterial, as $\sqrt{Nh_N} (\tau_{s_1 \succ s_0}(c) - \tau_{s_1\succ s_0}(C_N)) = o_P (1)$.
\Copy{rdliterature}{(ref) is related to the literature on regression discontinuity with unknown or estimated cutoffs hansen2000sample,porter2015regression. In some settings card2008tipping, the cutoff that determines treatment assignment is unknown and must be estimated. In those settings the cutoff turns out to be estimable at faster-than-$\sqrt{N}$ rates, and its estimation has no effect on subsequent asymptotics. This setting differs from ours since, here, treatment assignment is determined by $C_N$ and not $c$, and $C_N$ is not superconsistent for $c$. Nevertheless, because nonparametric estimators for RD effects converge at slower-than-$\sqrt{N}$ rates, the randomness in $C_N$ also does not affect the asymptotics of the RD estimators. }
We conclude this section by outlining how these asymptotic results are useful to construct inferential statements for further aggregations of aTEs.
We return to the example in (ref) and set up a Monte Carlo study. Suppose the preference types $\br{A, B, C}$ are equally probable. School capacities are $\br{N, 0.25N, 0.25N, 0.25N}$, respectively for $s_0$ through $s_3$. Suppose the test score $R_i \iid \Unif[0,1]$, independently of preferences. Suppose the lottery $U_i \iid \Unif [0,1]$ independently as well. The potential outcomes are heterogeneous. Their conditional means given preference type and test scores are described by (ref). Detailed construction of the simulated data is documented in the code repository (\url{github.com/jiafengkevinchen/school-choice-monte-carlo}).
In this setup, there are seven pairs of schools with nontrivial identified maximal treatment effects. For each pair $s_1 \succ s_0$ and each preference type, (ref) plots the region where $R_i \in \bar E_{s_1}(\succ, C_N) \cap \bar E_{s_0}(\succ, C_N)$. The maximal treatment effect ($\tau_{s_1\succ s_0}$ in (ref)) is then an aggregation of aTEs with $R_i \in \bar E_{s_1}(\succ, C_N) \cap \bar E_{s_0}(\succ, C_N)$. For instance, the effect $\tau_{3\succ 1}$ is an average of effects between type $A$ and type $C$ students with test scores between the two cutoffs.
For one draw of the data with $N=1000$, (ref) shows estimates and standard errors of $\tau_{s_1 \succ s_0}$, using the estimators in (ref). Wald inference based on the analytic standard error appears accurate, and the bootstrap variance is close to the analytic estimated SEs. The estimates are approximately independent: The sampling correlations between different estimates---themselves estimated by the nonparametric bootstrap---are negligible, with all pairwise correlations less than 0.07.
At this sample size, the RD-driven estimates ($\tau_{1\succ *}$ and $\tau_ {2\succ *}$) are not significantly noisier than the lottery estimates. Thus in aggregations they receive nontrivial weight at $N=1000$. To illustrate vanishing weights, for a variety of sample sizes, (ref) in turn shows our diagnostic in (ref) for the regression estimator with large market propensity scores, treating schools 2 and 3 as treatment schools. As expected, the weight put on RD variation depends on and vanishes with the sample size.
To accentuate the concerns regarding aggregation, we choose the potential outcome distributions so that $s_1$ is particularly good and $s_2$ is particularly bad for students of preference type $B$. This makes students of type $B$ have large negative RD-driven treatment effects for the treated schools $2$ and $3$. Correspondingly, in (ref)(b), we observe that the estimated treatment effect becomes less negative as sample size increases, since less of these RD-driven effects is aggregated.
We illustrate alternative methods for aggregation in (ref), which shows estimates of $\mu_s^L, \mu_s^R$ in a specification like (ref) in panels (a)--(b). Each value-added should be interpreted as a comparison of school $s$ to school 0, where (ref) shows how each is computed from the pairwise effects $\tau_{s_1 \succ s_0}$. Generally speaking, the RD-driven effects $\mu_s^R$ are noisier than the lottery-driven effects. Lottery-driven value-added estimates can be different from the test-score driven ones. For instance, for school $3$, the pairwise effect $\tau_ {1 \succ 3}$ is large and positive, and $\tau_{2\succ1}$ is negative, since they include variation in students of type $B$. These effects are aggregated in the value-added $\mu_3^R$, contributing to its large negative value relative to $\mu_3^L$.
In (ref)(c), we fit a similar regression to (ref), but we further aggregate effects by treatment schools (2 and 3) and control schools (0 and 1). This regression specification implicitly defines the RD-driven effect $\tau^R$ as $ \frac{1}{2} (\tau_ {2\succ 1} - \tau_ {1\succ 3})$ and $\tau^L$ as $\frac{1}{2}(\tau_ {3\succ 0} + \tau_ {3\succ1})$. We find that $\tau^R$ is much more negative than $\tau^L$, partly driven by students of preference type $B$. The corresponding estimate from the local DA propensity score regression is about $-0.53$, lying between these two estimates.
Detailed administrative data from school choice settings provide an exciting frontier for causal inference in observational data. Remarkably, school choice markets are engineered roth2002economist to have desirable properties for market participants, and yet they may yield natural-experiment variation that inform program evaluation and policy objectives. Credibility of empirical studies using such variation demands an understanding of the limits of the data---an understanding of what counterfactual queries the data can and cannot answer (absent further assumptions). Our analyses here provide a step towards that understanding.
As a review, we provide a detailed analysis of treatment effect identification in school choice settings. We characterize the identification of aTEs, the building blocks of aggregate treatment effects. We find that pooling over lottery- and RD-driven aTEs leads the former to dominate the latter asymptotically. We provide a simple regression diagnostic for the weight put on RD-driven aTEs, as well as some suggestions for aggregating aTEs. Lastly, we contribute asymptotic theory for estimating aggregations of RD-driven aTEs.
Declaration of generative AI and AI-assisted technologies in the writing process
During the preparation of this work the author(s) used ChatGPT and Google Gemini in order to structure ideas, copy-edit, and produce code to implement procedures. After using this tool/service, the author(s) reviewed and edited the content as needed and take(s) full responsibility for the content of the publication.
\tiny \singlespacing
\onehalfspacing