The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
111,566 characters
Nonparametric Treatment Effect Identification in School Choice
\begin{abstract}
This paper studies nonparametric identification and estimation of causal
effects in centralized school assignment. In many centralized assignment algorithms,
students face both lottery-driven variation and regression discontinuity- (RD) driven
variation. We characterize the full set of identified \emph{atomic treatment effects}
(aTEs), defined as the conditional average treatment effect between a pair of schools
given student characteristics. Atomic treatment effects are the building blocks of more
aggregated treatment contrasts, and common approaches to estimating aTE aggregations can
mask important heterogeneity. In particular, many aggregations of aTEs put zero weight
on aTEs driven by RD variation, and estimators of such aggregations put asymptotically
vanishing weight on the RD-driven aTEs. We provide a diagnostic and recommend new
aggregation schemes. Lastly, we provide estimators and asymptotic results for inference
on these aggregations.
\vspace{3em}
{\slshape \noindent Keywords: School choice,
nonparametric identification, regression discontinuity, heterogeneous
treatment effects
\noindent JEL codes: C1, D47, J01}
\end{abstract}
\thispagestyle{empty}
\maketitle
\section{Introduction}
\Copy{specificstudy}{A rapidly growing empirical literature studies causal
effects of schools using centralized school assignment.\footnote{See, among others,
\citet
{kapor2020heterogeneous,agarwal2018demand,fack2019beyond,calsamiglia2020structural,angrist2020simple,beuermann2018good,abdulkadirouglu2017research,abdulkadiroglu2017impact,abdulkadirouglu2020parents,abdulkadroglu2017regression,marinho2022causal,kirkeboen2016field,angrist2021credible,ketel2023unimport,angrist2022race}.}
In several centralized assignment systems, two sources of variation determine how
schools enroll applicants and drive quasi-experimental identification. For instance, in
New York City, studied by
\citet{abdulkadiroglu2017impact}, some schools randomize priority to students, where
students with higher lottery numbers are favored over others. Other schools use
non-lottery
tiebreakers, such as test scores, to distinguish students with otherwise similar
characteristics.
These sources of variation provide opportunities for causal inference. For lottery
schools, comparing lucky students who receive favorable priorities to those less
fortunate estimates a causal effect. For test-score schools, comparing students just above a cutoff to those just below also identifies causal effects.
\citet{abdulkadiroglu2017impact} use both sources of variation to measure school
effectiveness in New York City. School districts elsewhere, such as Chicago and Boston, also feature schools with lottery and non-lottery
tiebreakers.\footnote{\citet{angrist2023methods} write ``Some centralized assignment
schemes, such as those used for Boston and New York City exam schools and New York City
screened schools, employ non-lottery tiebreakers such as test scores instead of, or
alongside, lottery numbers.'' Schools in the Chicago Public School High School Choice
program use either a lottery system or a points system for giving priority to students,
where the points system is a composite of test scores and other student characteristics.
See \url{https://www.cps.edu/gocps/high-school/results/choice/}. }
We refer to the first source of variation as lottery variation and the second as
regression-discontinuity (RD) variation. Researchers often estimate aggregate causal
quantities, such as contrasts between sets of ``treatment'' and ``control'' schools,
pooling over lottery and RD variation and over pairs of schools. \citet
{abdulkadiroglu2017impact} do so for New York City, estimating aggregate effects
of enrolling in a school receiving an ``A'' grade on the school district's report card
for quality, compared to enrolling in non-Grade-A schools. }
These aggregate estimates provide convenient causal summaries but can mask rich
heterogeneity by pooling many different comparisons. These contrasts may involve
different pairs of schools, different student characteristics, or different sources of
identifying variation. Aggregate causal estimands are often interpreted as simple
average treatment effects between treatment and control schools. Because these aggregate
effects blend together different conditional average treatment effects, their
interpretation should be more nuanced.
In particular, economically, heterogeneity along RD variation versus lottery variation
could be meaningful. Consider a lottery school $s_1$ and a test-score school $s_0$. In many centralized algorithms, lottery variation between these schools occurs only for students preferring $s_{1}$ to $s_{0}$, driven by whether they lose the $s_{1}$ lottery. Meanwhile, RD
variation exists only for students who prefer $s_0$ to $s_1$---driven by whether these
students narrowly qualify or fail the test used by $s_0$. If submitted
preferences select positively on gains, the causal contrasts identified by lottery (resp.
RD) variation represent subpopulations with more (resp. less) favorable potential
outcomes under $s_1$. Different aggregations thus put different weights
on subpopulations with different attitudes toward $s_1$.
Motivated by the rich heterogeneity, this paper seeks to reduce causal effects to their
smallest unit, which we call \emph{atomic treatment effects} (aTEs). We define an atomic
treatment effect as a conditional average treatment effect between two schools,
conditional on all observed student characteristics. Our first contribution is to
characterize the set of point-identified aTEs based solely on lottery and RD variation
from the assignment mechanism. Since any identified average treatment effect is thus a
weighted average over aTEs, this characterizes the set of estimable treatment effect
parameters in these settings.
Our characterization yields a clean classification of aTEs as driven by RD or lottery
variation. Building on this classification, one concerning observation is that common
regression estimators---often derived under homogeneous treatment effect
assumptions---implicitly aggregate aTEs using weights that are not chosen by the researcher and can depend on the sample size. Moreover, these
implicit weighting schemes can be at odds with the interpretation of these
estimates. Starkly, aggregate treatment effects that pool over lottery and RD variation
asymptotically put \emph{zero weight} on the atomic treatment effects identified by RD
variation. Our previous observation then also implies that these aggregates only reflect
treatment contrasts for test-score schools \emph{for students who disprefer these
test-score schools}. Nevertheless, such estimators are often interpreted as partly
reflecting the RD variation on a broader set of students.\footnote{\label
{fn:intro}For instance, \citet{abdulkadiroglu2017impact} employ such an approach, and the
title of \citet{abdulkadiroglu2017impact} suggests that RD-variation is central. To be
sure, there is evidence that in finite samples, their estimator puts nontrivial weight on
the RD-driven causal comparisons, even though the weight put on these comparisons
converges to zero asymptotically---see \cref{sub:diagnostic}.}
The reason for this result is simple: Since RD variation can only occur at a cutoff,
atomic treatment effects identified by such variation make comparisons for a ``thin'' set
of student characteristics that have population measure zero
\citep{khan2010irregular}---namely, those students with scores at a cutoff. Translated to
estimation, this means that the number of students contributing to RD variation is a
vanishing fraction compared to the number of students subject to lottery variation.
As a result, treatment effect aggregations that weigh each student equally put vanishing
weight on the RD-identified aTEs.
Our second set of contributions speaks to these concerns: We provide practical
recommendations, a new diagnostic, and theoretical guarantees. First, we introduce a
diagnostic that assesses---in finite samples---how much RD-variation contributes to a
particular regression estimator. In some cases, in finite samples, the weight on RD
variation for these estimators may be non-zero and substantial. This diagnostic assesses
whether particular estimators are subjected to the starkest problems arising from their
implicit weighting schemes.
\Copy{aggregationlit}{Second, researchers are encouraged to take an explicit stand on the
aggregation of aTEs. Our identification result for aTEs helps practitioners reason about
the behavior of identified aggregate estimands. We also provide some reasonable default
choices: We first aggregate identified aTEs for each ordered pair of schools $s_1, s_0$,
among those that prefer $s_1 \succ s_0$, into parameters $\tau_{s_1\succ s_0}$. This
aggregation respects identifying variation, and thus lottery-driven aTEs do not
dominate. We then propose further aggregation to school-level value-added by explaining
the variation in $\tau_{s_1 \succ s_0}$ through the difference in $s_1$ and $s_0$'s
value-added, similar to Bradley--Terry models of tournament rankings. In spirit, our
call for taking an explicit stand on aggregation echoes \citet
{cattaneo2016interpreting} and \citet{bertanha2020regression}, who consider aggregation
of regression discontinuity treatment effects at multiple cutoffs. Our setting is
additionally complex due to the presence of lotteries and of centralized assignment. }
To implement a user-chosen aggregation of aTEs, we also provide asymptotic theory for
estimators of lottery- and RD-identified aTEs---which can then be further aggregated
according to a user-chosen weighting. As a theoretical contribution, relative to the
existing literature \citep{abdulkadiroglu2017impact,abdulkadirouglu2017research}, our
asymptotic theory more accurately reflects the fact that cutoffs have nontrivial
randomness in finite samples, and only converge to fixed quantities asymptotically
\citep{azevedo2016supply}. \Copy {introdiff}{Nevertheless, we show that data-dependent
cutoffs do not affect the first-order asymptotic behavior of estimators for these
treatment effects, using a more refined analysis than in
\citet {abdulkadiroglu2017impact,abdulkadirouglu2017research}. This is a novel
contribution to the statistical theory of estimators in settings with school assignment
algorithms.}
This paper applies to any setting where school assignments are made using
deferred acceptance-like algorithms and researchers seek causal identification based on
the centralized assignments
\citep{beuermann2018good,abdulkadirouglu2017research,marinho2022causal,kirkeboen2016field,angrist2021credible}.
It contributes to a growing methodological literature on causal inference in centralized
school assignment settings
\citep{singh2023kernel,marinho2022causal,munrojmp,abdulkadiroglu2017impact,abdulkadirouglu2017research,narita2021theory,narita2021algorithm,che2023leveraging,arkhangelsky2025evaluating}.
This paper is most related to \citet{abdulkadiroglu2017impact}. We complement their work
by investigating the ramifications of heterogeneous treatment effects on their
identification strategy and by formally characterizing the set of atomic treatment effects
identified by centralized school assignment.\footnote{In particular, \citet{abdulkadiroglu2017impact}
``look forward to a more detailed investigation of the consequences of heterogeneous
treatment effects for identification strategies of the sort considered here,'' which is
precisely the theme of this paper.} Our results also imply that the regression
estimator proposed by \citet{abdulkadiroglu2017impact} estimates a treatment effect that
puts vanishing weight on the RD-identified aTEs, though nevertheless it appears the weight
on RD-identified aTEs is nontrivial in finite samples in their empirical application.
This paper proceeds as follows. \Cref{sec:model} introduces notation, setup, a class of
school choice mechanisms, and a set of treatment effect estimands; in particular, \cref
{sec:nonstoch} discusses in detail the motivation for our identification results and some
recommendations for aggregating atomic treatment effects. \Cref {sec:identi} characterizes
those that are identified. \Cref{sub:est} proves asymptotic properties for standard
estimators of lottery- or RD-driven estimands. \Cref {sec:monte_carlo} illustrates our
recommendations and results with a Monte Carlo study. \Cref{sec:conclusion} concludes.
\section{School choice mechanisms, estimands, and identification}
\label{sec:model}
Consider a finite set of students $I = \{1,\ldots, N\}$ and schools $S =
\{0, 1, \ldots, M\}.$ The schools have \emph{capacities} $q_1,
\ldots, q_M \in \N$. Assume that school $0$ represents an outside option with infinite
capacity. Each student $i\in I$ has observed (by the analyst) characteristics $X_i$.
$X_i$ contains characteristics that are relevant for the assignment mechanism.\footnote
{It may also contain other observed characteristics that the analyst may condition on.
Since these additional covariates do not change our results materially, we suppress them
and assume for convenience that they are not available.} For each student $i$ and each
school $s$, there is a potential outcome $Y_i(s)$, representing the outcomes of the
student if assigned to school $s$.\footnote{Consistent with much of the literature, this
formulation rules out peer effects or other violations of the stable unit treatment
value assumption(SUTVA). } The assignment mechanism takes in $X_1,\ldots, X_N$ and
produces matchings $D_i = [D_ {i1},\ldots, D_{iM}]'$, such that $D_{is} = 1$ if and only
if $i$ is matched to $s$, and $(D_1,\ldots, D_N)$ satisfies school capacity
constraints.\footnote{That is, $\sum_{i} D_{is} \le q_s$ for all $s$. } The analyst
observes $Y_i =
\sum_{s}
D_{is} Y_i(s)$, $D_i$, and $X_i$ for every individual.
To embed this setup in an asymptotic sequence, assume that $(X_i, Y_i(0),\ldots, Y_i
(M)) $ are i.i.d. draws from a superpopulation. The set of schools $S$ is fixed, but
their capacities in finite samples are generated by $q_s = \lfloor N q_s^*
\rfloor$ for some fixed value $q_s^* \in [0,\infty)$.
\subsection{Assignment mechanism}
\label{sub:assignment_mechanism} Following \citet{abdulkadiroglu2017impact}, we consider
the assignment mechanisms that derive matchings from student-proposing deferred
acceptance. This setup accommodates the school choice
mechanism in New York City, and nests many mechanisms that either
use deferred acceptance or can be reduced to deferred acceptance.\footnote{\label
{fn:othermech}We consider the same mechanisms as \citet{abdulkadiroglu2017impact}. See
footnote 5 in
\citet{abdulkadiroglu2017impact} for a list of school choice settings accommodated by this
setup, which includes mechanisms (e.g. immediate acceptance) that can be represented by
deferred acceptance under certain transformations of preferences.}
Deferred acceptance requires student rankings over schools (termed \emph{preferences}) and
school rankings over students (termed \emph{priorities}). Students submit preferences, which are included in $X_i$, and school priorities are computed from $X_i$ as follows.
There are two types of schools, lottery schools and test-score schools.
School priorities are lexicographic in either $(Q_{it_s}, R_ {it_s})$ or $(Q_{i
\ell_s}, U_{i
\ell_s})$, for $R$ the set of test scores and $U$ the set of lottery numbers. The quantity
$Q_ {is} \in
\br{0,\ldots, \bar q_s}, \bar q_s \in \N \cup \br{0},$ represents certain qualifiers;
higher values of $Q_{is}$ mean that $i$ is ranked favorably by $s$. This accommodates
settings where, for instance, having a sibling at $s$ makes $i$ high-priority; we may
represent such students with $Q_{is} = 1$ and others with $Q_{is} = 0$. When two
students $ (i,j)$ have the same discrete qualifier $Q_{is} = Q_{js}$, ties are broken by
either a test score $R_{it_s}$ or a lottery number $U_ {i\ell_s}$, where $t_s, \ell_s$
indexes which test or lottery number school $s$ uses.
Formally, assume we partition $S$ into lottery and test-score schools. A lottery school
$s$ uses a lottery indexed by $\ell_s \in \br{1,\ldots, L}$, $L\in \N$, and a test-score
school uses a test-score indexed by $t_s \in \br{1,\ldots, T}$, $T\in \N$. Two schools may
use the same lottery or test score for tie-breaking. Assume that the assignment-relevant
information $X_i$ takes the form \[ X_i = (\succ_i, R_i, Q_i) = (\succ_i,
\underbrace{(R_{i1}, \ldots, R_{iT})}_{ \text{continuously distributed on $[0,1]^T$}},
(Q_{i0},\ldots, Q_{iM}))
\]
where $\succ_i$ represents the preferences of student $i$, $R_{it}$ represents $i$'s test
score on test $t$, and $Q_{is}$ represents $i$'s discrete qualifier at school $s$.
Let $U_{i} = [U_{i1},\ldots, U_{iL}]' \iid F_U$, supported on $[0,1]^{L}$, be the vector of
lottery numbers for
student $i$, where components $(U_{i\ell_1}, U_{i\ell_2})$ need not be
independent. To implement the lexicographic
priorities in $Q_{is}$ and $U_{i\ell}$ (or $R_{it}$), each
school $s$
computes a \emph{priority score} for each student $i$ \[ V_ {is} = \begin{cases}
\frac{Q_{is} + R_{it_s}}{\bar q_s + 1}, &\text{ if $s$ is a test-score school}\\
\frac{Q_{is} + U_{i\ell_s}}{\bar q_s + 1}, &\text{ if $s$ is a lottery school}.
\end{cases}
\] $V_{is}$ encodes lexicographic priorities in ($Q$, $U$ or $R$), because the integer
part of $ (1+\bar q_s)V_ {is}$ is $Q_{is}$ and the fractional part is $R_{it_s}$ or $U_
{i\ell_s}$. The priority ranking for school $s$ is thus \[ i \rhd_s j \iff \bk{(Q_{is},
R_{it_s}) >_ {\text{lex}} (Q_{js}, R_{jt_s}) \text{ or } (Q_{is}, U_{i\ell_s})
>_{\text{lex}} (Q_ {js}, U_{j\ell_s})} \iff V_ {is} > V_ {js},
\numberthis \label{eq:school_priority_v}
\]
where $i \rhd_s j$ means $i$ is more favorably ranked than $j$ by school $s$.
The matching $D_1,\ldots, D_N$ is then computed by the student-proposing deferred
acceptance algorithm with student preferences $\succ_1,\ldots, \succ_N$ and school
priorities $\rhd_ {0},\ldots, \rhd_{M}$ \citep{GS1962,abdulkadirouglu2003school}. The
student-proposal deferred acceptance algorithm is well-known; we reproduce it in
\cref{asec:misc}.
To make the notation concrete, consider the following setup, which is accommodated by
these mechanisms.
\begin{exsq}
\label{ex:notation_example}
Suppose there are four schools in addition to the outside option $s_0$. Two ($s_1, s_2$)
are
lottery schools, and two ($s_3, s_4$) are test-score schools. Suppose the two lottery
schools use the same lottery number ($L=1$), and the two test-score schools use the same
test ($T=1$). Suppose school $s_1$ additionally gives priority to students who have a
sibling ($Q_{i1} = 1$) already at the school; those students are lexicographically
preferred to other students, and ties are broken through lottery numbers. Each student is
characterized by \[(\succ_i, R_{i1}, (Q_{i1}, 0, 0, 0))
\] where $\succ_i$ is her reported preferences, $R_{i1}$ is her performance on the only
test in this setup, and $Q_{i1} \in \br{0,1}$ is whether the student has a sibling enrolled at
$s_1$. Each student additionally has the lottery number $U_{i1} \iid \Unif[0,1]$.
Then, under this setup, for two students $i, j$, the school priorities are as follows:
\begin{itemize}
\item School
$s_1$ prefers $i$ to $j$ ($i \rhd_{s_1} j$) iff either $Q_{i1} > Q_{j1}$ or ($Q_{i1} =
Q_
{j1}$ and $U_{i1} > U_{j1}$).
\item School $s_2$ prefers $i$ to $j$ iff $U_{i1} > U_{j1}$.
\item School $s_3, s_4$ prefer $i$ to $j$ iff $R_{i1} > R_{j1}$.
\end{itemize} These school priorities are equivalently represented by \eqref
{eq:school_priority_v}. Assignments are then taken by running student-proposing
deferred acceptance on student-reported preferences $\succ_i$ and the school priorities
$\rhd_s$.
\end{exsq}
\citet{azevedo2016supply} observe that deferred acceptance can be interpreted
as computing a vector of \emph{cutoffs}. Let \[C_{s,N} = \begin{cases}
\min_{i} \br{V_{is} : D_{is}=1}, &\text{ if $s$ is matched to $q_s$ students}\\
0, & \text{ otherwise.}
\end{cases} \] be the
least-favorable priority score among those matched to $s$, if $s$ is oversubscribed. The
matching computed by deferred acceptance is rationalized by each student $i$ being matched
to their favorite school among the set of schools for which $V_{is} \ge C_{s,N}$. This
cutoff structure generates RD-driven identification.
\subsection{Atomic treatment effects and their aggregation in school choice}
\label{sec:nonstoch}
A key contribution is to clarify which conditional average treatment effects are
identified by centralized assignment without restricting treatment effect heterogeneity.
To do so, we consider the ``smallest'' building block of treatment contrasts in this
setting, which we call \emph{atomic treatment effects}. Specifically, for a pair of
schools $s_0, s_1 \in S$, define the
\emph{atomic treatment effect} (aTE) as the conditional average treatment effect between
this pair of schools for a particular value of observables $X$:
\[
\tau_{s_1, s_0} (x) = \E\bk{
Y(s_1) - Y(s_0) \mid X=x}.
\]
For a school $s$, define the {atomic potential outcome mean} as $
\mu_s(x) = \E[Y(s) \mid X=x].
$
Our identification results characterize for which $(s_1, s_0, x)$ the aTE $\tau_
{s_1, s_0}(x)$ is point-identified under mild assumptions. This exercise reinforces some
claims in the literature,\footnote{\Cref{asub:identification_lit} shows that the values for
which $\mu_s(x)$ is identified under our formulation correspond exactly to those values
$(s, x)$ for which the local deferred acceptance propensity score
\citep{abdulkadiroglu2017impact} is positive.} separates the RD-driven and the
lottery-driven causal effects, and clarifies that existing estimates of aggregate
treatment effects may have poor interpretation. Our results then show how to estimate
(certain granular aggregations of) the aTEs, thus enabling users to construct more
interpretable estimands. We pause and discuss the motivation and limitation of this
exercise.
\Copy{idlitmotiv}{First, knowing which aTEs are identified informs us which aggregate
treatment effects are point-identified given only the assignment algorithm.
Existing work \citep{abdulkadiroglu2017impact,abdulkadirouglu2017research} values
centralized assignments partly because they provide analogues of
quasi-experimental research designs that are popular in other microeconometric
settings.\footnote{For instance, \citet{abdulkadiroglu2017impact} write, ``Centralized
student assignment opens new opportunities for the measurement of school quality. The
research potential of matching markets is enhanced here by marrying the conditional
random assignment generated by lottery tie-breaking with RD-style variation at screened
schools.''} We establish formal identification and extends it to heterogeneous treatment effects.\footnote{\citet{abdulkadiroglu2017impact} calculate
what they term local deferred acceptance (DA) propensity scores and show in their
Corollary 1 that under constant treatment effects, the treatment effect is identified.
Our results characterize the set of identified aTEs. These identification results connect
to \citet{abdulkadiroglu2017impact}: We show that $\mu_s (x)$ is identified if and only
if the local DA propensity score for school $s$ at characteristics $x$ is positive. See
\cref{prop:eligibility,asub:identification_lit}.}} Naturally, the \emph{only} treatment
effect parameters that are point-identified correspond to aggregate treatment effects
that place weight only on identified aTEs: i.e., estimands of the form
\[
\tau \equiv \int \sum_{s_1 \neq s_0} w(s_1, s_0 \mid x) \cdot \underbrace{\E[Y(s_1)-Y
(s_0) \mid
X=x]}_{\tau_{s_1, s_0}(x)} dW (x), \numberthis \label{eq:aggregate_TE}
\] where $w(s_1, s_0 \mid x)$ and $W(x)$ only put nonzero weight on $\tau_{s_1, s_0}
(x)$ that are point-identified. Thus, a
practitioner who seeks to only rely on quasi-experimental research designs can inspect
the class of identified aTEs and decide the aggregation that is most informative of
their substantive economic question.
Second, disaggregating into aTEs helps us understand what practitioners---who may impose
constant-treatment-effect assumptions---target under misspecification. For instance,
regression estimators for comparing schools---commonly, estimating causal effects of one
group of schools $S_1$ to another $S_0 = S \setminus S_1$ by regressing $Y$ on $\sum_
{s\in S_1} D_{is}$ as well as other controls---pool over aTEs between different school
pairs and aTEs using different sources of variation. Our analysis disaggregates such
estimands. Our subsequent analysis cleanly separates which aTEs are identified through
lottery variation and which are through RD variation, collected in the following
definition:
\begin{defn}
\label{defn:lottery_vs_rd}
Fix $x = (\succ, q, r)$. We say that an identified aTE $\tau_{s_1, s_0}(x)$ is
\emph{driven by lottery variation} if $\max_{\succ}(s_1, s_0) $ is a lottery school, and
we refer to all other identified aTEs as \emph{driven by RD variation}.
\end{defn}
\Copy{aggzero}{In disaggregating these estimands, we find that regression estimators---as
used by, e.g., \citet{abdulkadiroglu2017impact}---pool aTEs from both sources in a way
that asymptotically puts zero weight on RD-driven aTEs.} This arises because RD-driven
estimands use a much smaller set of students than lottery-driven ones, so
inverse-variance aggregation leads lottery-driven estimates to dominate. We provide some
diagnostics for assessing the severity of this issue in finite samples. This feature
also calls for separating out the lottery and RD driven components when aggregating
aTEs, for which our estimation results are helpful.
Third, granular aggregations of aTEs can yield more informative, yet still
low-dimensional, summaries of school value-added. Each individual aTE is likely too
imprecisely estimated to be useful, since it conditions on a single value of $X$. In
constructing useful aggregations, one should consider what features plausibly predict
heterogeneity. We speculate that, between schools $(s_1, s_0)$, it matters whether
$\tau_{s_1, s_0}(X)$ is lottery- or RD-driven, and it also matters whether $X$ contains
individuals who report preferring $s_1$ to $s_0$ or vice versa.
Given $s_1$ and $s_0$, for every student with $s_1 \succ s_0$, there is a
\emph{maximal} identified aggregation \[\tau_{s_1 \succ s_0} = \int \tau_ {s_1,s_0}(x)
dW_ {s_1\succ s_0} (x),\numberthis \label{eq:maximal}\] where the measure $W_ {s_1\succ
s_0}$ is the conditional distribution of $X$ given that $s_1 \succ_i s_0$ and that
$\tau_{s_1, s_0} (X)$ is identified.\footnote{Our identification results produce a set
$\mathcal I_{s_1, s_0}$ of $X = (\succ, R, Q)$ values such that $x \in \mathcal I_{s_1,
s_0}$ if and only if $\tau_ {s_1, s_0}(x)$ is point-identified. $W_{s_1 \succ s_0}$ is
then the conditional distribution $(X \mid X \in \mathcal I_{s_1, s_0}, s_1 \succ s_0)$.
For RD-driven aTEs, the set $\mathcal I_{s_1, s_0}$ has zero measure under $P$, and thus
we take additional care to \emph{define} the conditional distribution. See
\cref{sub:ATE_def} for defining general aggregate treatment effects, for which $\tau_{s_1
\succ s_0}$ is a special case. } $\tau_ {s_1
\succ s_0}$ does not aggregate across identifying variation: It is either RD-driven or
lottery-driven depending on whether $s_1$ is a lottery school, per
\cref{defn:lottery_vs_rd}. \cref{sub:est} provides asymptotic guarantees
for estimating \cref{eq:maximal}.
Practitioners can further summarize estimates $\hat \tau_ {s_1 \succ
s_0}$ into school value-added by positing a Bradley--Terry-style \citep{firth2005bradley}
model in which, for instance, the pairwise treatment effects are explained by differences
in school value-added, adjusted for reported preference and source of identification \[
\hat\tau_{s_1 \succ s_0} = \alpha_0 +
(\mu_ {s_1}^L - \mu_ {s_0}^L)\one (\text{$s_1$ is
lottery}) +
(\mu_{s_1}^R - \mu_{s_0}^R)\one(\text{$s_1$ is test-score}) + \epsilon_ {s_1, s_0}
\numberthis \label{eq:BT_aggregation}.
\] Here, $\mu_s^L, \mu_s^R$ represent school value-added among lottery- and RD-based
comparisons, respectively, and $\alpha_0$ adjusts for the fact that $s_1$ is revealed to
be preferred to $s_0$. These parameters may be estimated via least-squares (normalizing
some effects to zero) or further regularized under random effects or hierarchical
Bayesian-type assumptions. Under constant effects ($Y(s) = Y(0) +
\beta_s$ almost surely), this model is exactly well-specified, and $\mu_ {s_1}^L - \mu_
{s_0}^L =
\mu_{s_1}^R - \mu_{s_0}^R$ recovers $\beta_{s_1}-\beta_{s_0}$. Without such an
assumption, $\mu_ {s}^L, \mu_s^R$ can be interpreted as least-squares summaries of
school causal effects. This cleanly separates identification of causal effects from
their aggregation and summary.
\Copy{contlarge}{
\Copy{cont}{This exercise has limitations. First, as we shall require in \cref
{as:cts}, identification of RD-driven aTEs relies on continuity of $\mu_s(x)$ in test
scores at cutoffs \citep{hahn2001identification}. \citet{marinho2022causal} note that
this continuity assumption restricts strategic misreporting of preferences. Suppose
students know the admission cutoffs and can strategically report preferences based on
their test scores or lottery numbers. With strategic reporting, those who report
$\succ$ just below a cutoff may have very different true preferences from those who
report $\succ$ just above a cutoff. The mix of students can thus
change sharply at the cutoff, possibly making conditional expectations discontinuous.
For strategic misreporting to be a concern for continuity, though, students must
be able to reliably forecast cutoffs. Prediction errors would smooth over the mix of
students above and below a cutoff, potentially restoring continuity.\footnote{
\label{fn:blm_fn}For
instance, in many settings (e.g. New York City \citep{che2023leveraging} and the Chinese
college entrance exam \citep{chen-kesten-chinese}), students submit preferences before
learning their test scores, rendering accurate manipulation difficult. In their empirical
application, \citet{marinho2022causal} do show evidence that students manipulate reported
preferences: Their Figure 1 shows that students are more likely to rank school $j$ more
than the outside option if their running variable is closer to $j$'s cutoff, but this
probability does not appear to change discontinuously at the cutoff.}
Moreover, such
manipulation plausibly leads to discontinuities in the density of various covariates at
the cutoff, and can in principle be empirically assessed.
Second, even if continuity holds, the aTEs may still have limited external validity.
They are useful for the types of program evaluation exercise in
\citet{abdulkadirouglu2017research,abdulkadiroglu2017impact}, but are perhaps less
directly informative of counterfactuals in which school choice mechanisms change or
school admissions policies change. The aTEs condition on reported preferences, test
scores on existing tests, and are only identified for certain subpopulations. If any of
them change in a counterfactual policy environment, additional extrapolative
assumptions are needed. We refer readers to \citet{marinho2022causal} for assumptions
and methods that partially identify conditional average treatment effects that
condition on true preferences.}
Despite this caveat, the maximal aggregation of aTEs (\cref{eq:maximal}) does have a
clear policy interpretation: It measures the effects of policies on the margin. Consider
a small increase in $s_1$'s capacity and assume such a change does not alter preference
submission. Such an increase diverts some students from enrolling in $s_0$ as they now
qualify for $s_1$, which they prefer. This aggregate treatment effect (\cref{eq:maximal})
between $s_1, s_0$ exactly measures the gain of these students in terms of the outcome
$Y$. Thus, for \emph{marginal} changes, such as increasing the capacity of school $s_1$
by a small amount or allocating marginal resources to $s_1$, such aggregations of aTEs
are plausibly informative of the impact of such policies; see
\citet{arkhangelsky2025evaluating} for a similar argument.
Indeed,
\citet{abdulkadiroglu2017impact} are motivated by debates in New York City over whether students have adequate access to Grade A schools. This policy question can be viewed as evaluating marginal expansions of Grade A schools; if the impact is large, current access is suboptimally limited.}
So far, we have discussed identification heuristically. Because the school assignments
$D_i$ depend jointly on cutoffs $C_N$, themselves computed from all characteristics $
(X_1,\ldots, X_N)$, $(Y_i, D_i, X_i)$ \emph{are not i.i.d.} conditional on $C_N$. Thus,
the usual cross-sectional notion of identification does not apply.\footnote{A standard definition states that a quantity is identified if no
two observationally equivalent distributions of observable data and potential outcomes
give rise to different values of the quantity. Here, ``distributions'' refers to the
\emph{joint distribution} of the observable data, school assignments, and potential
outcomes \emph{for the finite sample of $N$ students}. However, vanishing variation of
$C_N$ around $c$ would allow for certain quantities to be ``identified'' per the standard
definition, but not consistently estimable. We give a simple example in
\cref{asub:id_example} to illustrate the conceptual difficulties.} The literature
\citep{abdulkadiroglu2017impact,marinho2022causal} often appeals to large market
asymptotics and treats the $C_N$ as nonrandom instead in identification analysis.\footnote{Notable exceptions
include
\citet{agarwal2018demand} and \citet{munrojmp}.} We conclude this section by making this
heuristic precise, and we take $C_N$ as random in our estimation section.
\subsection{Identification and random cutoffs}
When students are drawn from a population and the number of students is large, the random
cutoffs $C_{s,N}$ concentrate around some population quantity, satisfying a law of large
numbers. Proposition 3 of \citet{azevedo2016supply} states that if $
\br{(\succ_{is}, V_ {is}) : s \in S}$ are i.i.d. across students, then under mild
conditions,\footnote{Precisely speaking, \citet{azevedo2016supply} define a notion of
deferred acceptance matching acting on the \emph{continuum economy}---which is the
distribution of $\br{(\succ_{is}, V_ {is}) : s \in S}$. The additional regularity
condition is that this continuum version of deferred acceptance matching admits a unique
stable matching. In other words, it is a very mild regularity condition on the
distribution of $\br{(\succ_{is}, V_ {is}) : s \in S}$.} the cutoffs $C_N = [C_
{0,N},\ldots, C_{M,N}]$ concentrate to some population counterpart $c$ at the parametric
rate: $
\norm{C_N - c}_\infty = O_P(N^{-1/2}).
$
\begin{as}
\label{as:limitcutoff}
The population of students is such that the cutoffs $\br{C_{s,N}}$, satisfy
\[\max_{s\in S}\, \abs{C_{s,N} - c_s} = O_P(N^ {-1/2})\] for some fixed $c = (c_s \in
[0,1) :
s\in S)$.
\end{as}
We define identification relative to the population cutoff $c$. For a fixed cutoff $c$, we
can define $D_i^*(c)$ as the (fictitious) assignment $i$ would receive if the cutoffs were
set to $c$. The tuple $(X_i, D_i^*(c), Y_i (0),\ldots, Y_i(M), U_i)$ are then i.i.d., and
identification reduces to its definition in standard settings. Since as
$C_N \to c$, $D_i^*(C_N)
\to D_i^*(c)$ for $N \to \infty$, this notion of identification captures the information
content of the data for large markets.
Formally, let $P \in \mathcal P$ be the distribution of student characteristics, potential
outcomes, and lottery numbers $(X_i, Y_i(0),\ldots, Y_i(M), U_i)$, where we assume every
member of $\mathcal P$ satisfies \cref{as:limitcutoff} \citep{azevedo2016supply}. Let
$c(P)$ be the corresponding large-market cutoffs, defined as the probability limit of
$C_{N}$ when data is sampled according to $P$. Define $D_{i}^*(c)$ as the assignment made
if the cutoffs were set to $c$: i.e. $D_ {is}^*(c) = 1$ if and only if $s$ is $i$'s
favorite school among those with $V_{is} \ge c_s$. Let $Y^*_{i,\text{obs}}(c) = \sum_{s\in
S} D_{is}^*(c) Y_i(s) $ denote the observed outcome under $D_i^*$.
\begin{defn}
Let $ P^*_{\text{obs}}(P) $ denote the distribution of $(D_i^*(c(P)),
Y^*_{i,\text{obs}}(c(P)), X_i, U_i)$. We say a parameter $\tau(P)$ is \emph{identified at
$P$} if there does not exist $\tilde P$ such that $\tau(P) \neq
\tau(\tilde P)$ and $P,
\tilde P$ are observationally equivalent:
$P^*_{\text{obs}}(P) = P^*_{\text{obs}}(\tilde P)$ and $c(P) = c(\tilde P)$.
\end{defn}
\section{Identification of atomic treatment effects}
\label{sec:identi}
\subsection{An extended example}
\label{ex:running_ex} We begin with a simplified setting that contains most of the intuition for our formal results.
\begin{exsq}[Extended example setup]
Suppose there are four schools $s_0, s_1, s_2, s_3$ and three types of students
denoted by $(A, B, C)$. The students have the following preferences and the same
discrete priorities $Q_{is} = 0$:
\begin{center}
\vspace{1em}
\begin{tabular}{ccc}
\toprule & Preferences \\\midrule $A$ & $s_2 \succ s_3 \succ
s_1 \succ s_0$ \\ $B$ & $s_2 \succ s_1 \succ s_3 \succ s_0$ \\ $C$ & $s_3 \succ
s_2 \succ s_1 \succ s_0$ \\
\bottomrule
\end{tabular}
\vspace{1em}
\end{center}
Additionally, consider the following setup, illustrated in \cref{fig:illustration}:
\begin{enumerate}[wide]
\item The schools $s_1, s_2$ are
test-score schools using the same continuously distributed test score $R \in [0,1]$.
\item $s_3$ is a lottery school, and $s_0$ is an undersubscribed
(lottery) school with sufficient capacity. Assume that $s_3$ is oversubscribed---the
probability of qualifying for $s_3$ for any student is not zero or one.\footnote{Note that
if certain students had discrete priority over others, then it may be the case that
they qualify for $s_3$ with probability one, even if $s_3$ is oversubscribed.}
\item Assume the number of students is large enough that cutoffs $c_1, c_2$ can be treated as
fixed.
\item
Since everyone prefers $s_2 \succ s_1$, school 2 has a more stringent cutoff:
$c_2 > c_1$.
\item Since the distribution of $R$ is unspecified, assume without loss that $c_1 =
\frac13, c_2 = \frac23$.
\end{enumerate}
\begin{figure}[bth]
\centering
\begin{tikzpicture}
\draw[thick] (0,0) -- (5,0);
\draw[thick] (0,-0.2) -- (0,0.2);
\draw[thick] (5,-0.2) -- (5,0.2);
\filldraw (5/3,0) circle (1.5pt) node[below=2mm] {\(c_1\)};
\filldraw (10/3,0) circle (1.5pt) node[below=2mm] {\(c_2\)};
\node[below] at (0,-0.5) {$0$};
\node[below] at (5,-0.5) {$1$};
\node[above] at (2.5, 0.3) {$R$};
\node[right=1cm] at (6,.5) {\(s_0\) - lottery, undersubscribed};
\node[right=1cm] at (6,0) {\(s_1\) - test $R$, cutoff $c_1$};
\node[right=1cm] at (6,-.5) {\(s_2\) - test $R$, cutoff $c_2$};
\node[right=1cm] at (6,-1) {\(s_3\) - lottery, oversubscribed};
\end{tikzpicture}
\caption{Illustration of the example setup}
\label{fig:illustration}
\end{figure}
Our Monte Carlo results in \cref{sec:monte_carlo} revisit this example, with uniformly
distributed $R,U$ that determines the cutoffs $C_{1,N}, C_{2,N}$.
\end{exsq}
For this setup, we can write $X_i = (\succ_i, R_i)$ and omit $Q_{is}$. Note
that $\tau_{s,s'}(\succ, R)$ is identified if and only if both $\mu_{s}(\succ, R)$
and $\mu_{s'}(\succ, R)$ are. For a given preference $\succ$ and a school $s$, consider
the
\emph{$s$-eligibility set} \[ E_{s}(\succ, c) = \br{ r:
\P(D^*_{s} (c) = 1 \mid {\succ}, R=r) > 0 }. \numberthis \label{eq:s_elig_def} \] $E_s
(\succ, c)$ collects the test scores $r$ such that someone with $X = (\succ, r)$ has
positive probability of being matched to school $s$. As a result, $\mu_s(\succ, r)$ is
identified if $r \in E_s(\succ, c)$.
Computing these sets for each type yields:
\begin{center}
\vspace{1em}
\begin{tabular}{ccccc}
\toprule Preference type & $E_{s_0}$ & $E_{s_1}$ & $E_{s_2}$ & $E_
{s_3}$ \\
\midrule
$A$ & $[0, \frac13)$ & $[\frac13,\frac23)$ & $[\frac23, 1]$ & $[0, \frac23)$ \\
$B$ & $[0, \frac13)$ & $[\frac13,\frac23)$ & $[\frac23, 1]$ & $[0, \frac13)$ \\
$C$ & $[0, \frac13)$ & $[\frac13,\frac23)$ & $[\frac23, 1]$ & $[0, 1]$
\\
\bottomrule
\end{tabular}
\vspace{1em}
\end{center}
To illustrate the computation, consider students with preference $\succ_A$. For a given
value of $R=r$ and a school $s$, we ask whether $r \in E_s(\succ_A, c)$:
\begin{itemize}
\item When $R \in [0, \frac13)$, students of type $A$ do not qualify for either $s_1$ or
$s_2$, and they are assigned to $s_3$ if they win the lottery at $s_3$. Otherwise, they
are assigned to $s_0$. Thus $R \in E_{s_0}(\succ, c)$ and $R \in E_{s_3}(\succ, c)$, but
$R \not \in E_{s_2}$ and $R\not\in E_{s_1}$.
\item When $R \in [\frac13, \frac23)$, they do not qualify for $s_2$, but
they do qualify for $s_1$. Since they prefer $s_3$ to $s_1$, they are assigned to $s_3$ if
they win the lottery at $s_3$. Otherwise, they are assigned to $s_1$.
Thus $R \in E_{s_1}(\succ, c)$ and $R \in E_{s_3}(\succ, c)$, but $R \not\in E_{s_2}$ and
$R\not\in E_{s_0}$.
\item When $R \in [\frac23, 1]$, they qualify for $s_2$. Since $s_2$ is
their favorite school, they are assigned to $s_2$. Since they have positive probability of
being assigned only to $s_2$, $R$ is only in $E_{s_2}$, and all of $E_{s_0}, E_{s_1}, E_
{s_3}$ exclude $[\frac23, 1]$.
\end{itemize}
Our subsequent results make the above calculation systematic.
If we further assume that $r\mapsto \mu_s(\succ, r)$ is continuous in $r$ for every
preference and school $s$, then we can extend identification to the \emph{closure} of
$E_{s_j}(\succ, c)$ in $[0,1]^T$. Continuity assumptions allow us to take
sequences and extend identification to their limits: $\mu_s(\succ, r_k) \to \mu_s(\succ,
r)$ if $r_k \to r$. Thus, under continuity
of the atomic potential outcome means, we can compute regions on which each aTE
$\tau_{s,s'}(\succ, R)$ is identified: For a pair of schools $(s, s')$ and preference type
in $\br{A,B,C}$, we tabulate the set of $R$ values for which $\tau_{s,s'} (\succ, R)$ is
identified. The following table tabulates $\bar E_{s} \cap \bar E_{s'}$ for choices of
$s, s'$ and $\succ$:
\begin{center}
\vspace{1em}
\begin{tabular}{cccccccc}
\toprule Preference type & $(s_0, s_1)$ & $(s_0, s_2)$ & $(s_0, s_3)$ &
$(s_1, s_2)$ & $(s_1, s_3)$ & $(s_2, s_3)$ \\
\midrule
$A$ & $\br{\frac13}$ & $\varnothing$ & $[0, \frac13]$ & $\br{\frac23}$ & $[\frac13,\frac23]$ &
$\br{\frac23}$\\
$B$ & $\br{\frac13}$ & $\varnothing$ & $[0, \frac13]$ & $\br{\frac23}$ & $\br{\frac13}$ &
$\varnothing$\\
$C$ & $\br{\frac13}$ & $\varnothing$ & $[0, \frac13]$ & $\br{\frac23}$ & $[\frac13,\frac23]$ &
$[\frac23,1]$ \\
\bottomrule
\end{tabular}
\vspace{1em}
\end{center}
\noindent This calculation shows that the interpretation of aggregate treatment
effects is complex in two senses, a complexity often masked by simple aggregations.
First, aggregate treatment effects may mask heterogeneity in terms of which pairs of
schools are compared, since different pairs of schools $(s_i, s_j)$ are comparable on
different regions of the test score $R$. For instance, we might wish to estimate the
treatment effect of being enrolled in an inside option ($s_1,s_2,s_3$) relative to the
outside option $s_0$. Naturally, corresponding aggregate treatment effects are some
weighted average of the identified aTEs $\tau_{s_1, s_0}, \tau_{s_2,s_0}, \tau_
{s_3, s_0}$. In this case, interpreting these aggregates as general inside-versus-outside effects overstates their generality:
\begin{itemize}
\item For one, since $\tau_{s_2, s_0}$ is identified
for no value of $X$, all identified aggregations of the aTEs must exclude comparisons
between $s_2$ and $s_0$.
\item For another, among aggregations over $\tau_{s_1, s_0},
\tau_ {s_2,s_0}, \tau_{s_3, s_0}$, there is a unique maximal aggregation that weighs
each student equally, namely the $s_3$-against-$s_0$ treatment effect for those with
low test scores: $\E\bk{Y(s_3) - Y(s_0)
\mid R \le \frac{1}{3}}.$
\end{itemize} Therefore, in this case, interpreting aggregate treatment effects as an
inside-versus-outside option effect drastically overstates the generality of these
estimands in the presence of heterogeneous effects.
Second, aggregate treatment effects may mask heterogeneity in terms of which students
are compared for a particular pair of schools, since regions of $R$ that admit
comparisons for a given pair of schools differ substantially across student types.
This is true in this example for comparisons between $(s_2,s_3)$: \begin{itemize}
\item Students of type $A$ have identified aTEs between $s_2, s_3$ at a single
point $\br{\frac23}$
\item There are no identified aTEs for students of type $B$ between $s_2, s_3$.
\item Students of type $C$ have identified aTEs for $(s_2, s_3)$ for the
set $[\frac23,1]$, which has positive measure.
\end{itemize} In this case, no identified aggregations of aTEs between $s_2$ and $s_3$
take into account students of preference type $B$. Moreover, the maximal identified
aggregation of aTEs between $s_2$ and $s_3$ that weighs each student equally is the
$s_2$-against-$s_3$ effect among students of type $C$ with high test scores:
$\E[Y(s_2) - Y(s_3) \mid {\succ_C}, R\in [2/3,1]]$, since the set of students of type
$A$ with $R = 2/3$ is a measure-zero set. Again, interpreting these estimands as
blanket causal comparisons between $s_3$ and $s_2$ assumes that treatment effects for
other students are similar to those of type $C$ with test scores in $[2/3, 1]$.
We should expect the heterogeneity in both senses to be even more complex in
general, since this example only includes 3 out of the 24 possible preferences and
only a single type of test score. Given this heterogeneity, aggregations that pool
over many schools, preference types, and test scores can obscure the implicit weights
on aTEs. To understand these estimates, the next subsection characterizes the
eligibility sets $E_{s}$ as well as pairwise treatment contrasts formally.
\subsection{Identification of atomic treatment effects}
\label{sub:identification_of_atomic_treatment_effects}
\Copy{idintro}{
Like the example above, characterizing
identification of atomic treatment effects amounts to computing the $s$-eligibility sets
(\cref{eq:s_elig_def}). Our main result, \cref{prop:eligibility}, computes them in the
general setting. The behavior of intersections of $s$-eligibility sets formalizes the
distinction between aTEs that are driven by lottery variation and those driven by RD
variation. \Cref{cor:agg} then shows that any aggregations of aTEs that aggregate over a
non-vanishing subpopulation must place zero weight on the RD-driven aTEs.
In the general setting, recall that $X_i = (\succ_i, R_i, Q_i)$, where we let $R_i$
collect the test scores $(R_ {i1},\ldots,R_{iT})$ and $Q_i$ collect the discrete
qualifiers $(Q_{i0},\ldots, Q_{iM}).$ Fix $(\succ_i, Q_i)$, we first define the
$s$-eligibility sets as the set for $R_i$ on which assignment to $s$ has positive
probability.
\begin{defn}
Define the $s$-eligibility set at $(\succ, Q)$ and cutoffs $c$ as \[
E_s(\succ, Q, c) = \br{r \in [0,1]^T : \P(D_{is}^*(c) = 1 \mid R=r, \succ, Q) >
0}. \]
\end{defn}
}
\begin{as}
\label{as:cts}
For all $s$ and $(Q, \succ)$ with positive probability, the map $[0,1]^T \to \R$
\[ r \mapsto \E[Y(s)
\mid {\succ}, R=r, Q]\] exists and is continuous on $[0,1]^T$. Moreover, $\E[|Y(s)|
\mid
{\succ}, R=r, Q] < \infty$.
\end{as}
\Cref{as:cts} assumes that atomic mean potential outcomes are continuous in the test
scores (and that moments exist). This is a standard assumption in regression discontinuity
\citep{hahn2001identification}, but it does imply restrictions on strategic misreporting
of preferences, as we discuss in \cref{sec:nonstoch}. In so far as \cref{as:cts}
holds---intuitively, this is plausible when students fail to accurately forecast cutoffs
so that student type mixes do not change abruptly at a cutoff---our setting accommodates
mechanisms beyond standard deferred acceptance, since these other mechanisms can be
represented as deferred acceptance on transformed preferences \citep[as pointed out by] []
{abdulkadiroglu2017impact}.
The next proposition verifies that the $s$-eligibility sets $E_s(\succ, q, c)$ collect
values of $r$ such that $\mu_s (\succ, q, r)$ is identified. Continuity (\cref{as:cts})
extends identification to the closure of the $s$-eligibility sets.
\begin{restatable}{prop}{propid}
\label{prop:id} Consider some value $q$ of $Q_i$ and $\succ$ of $\succ_i$ with positive
probability at $P$. Consider schools $s_0, s_1, s$ and some value $r$ in the support of
$R_i \mid (Q_i=q, {\succ_i}={\succ})$. Under \cref{as:cts}, $\mu_s(\succ, r, q)$ is
identified at $P$ if and only if \[r \in \bar E_s(\succ, q, c(P)),\] where $\bar E_s$ is
the closure of $E_s$ in $[0,1]^T$. The aTE $\tau_{s_1, s_0}(\succ, r , q)$ is identified
at $P$ if and only if $r \in \bar E_{s_0} (\succ, q, c(P)) \cap \bar E_{s_1}(\succ, q, c
(P)).$
\end{restatable}
Our main result is the following characterization of all identified atomic treatment
effects: We compute the $s$-eligibility sets for both lottery and test-score schools, and
find the intersection of closures of $s$-eligibility sets between two schools.
Fix some school $s_0$ and some student with characteristics $x=(\succ, r, q)$. $\mu_{s_0}
(x)$ is identified if such a student has positive probability of being assigned to $s_0$.
Suppose $s_0$ is a lottery school. Naturally, in order for a student to have
positive probability to be assigned $s_0$, we require that
\begin{enumerate}[label=(\roman*),wide]
\item The probability that $s_0$ is the student's favorite lottery school that she
qualifies
for is positive: \[\pi^*_{s_0}(\succ, q, c) \equiv \P\bk{
s_0 = \argmax_{\succ} \br{s \text{ is a lottery school}: V_{is} > c_s} \mid {\succ},
R=r,Q=q
} > 0.\]
\item The student does not prefer any test-score school $s$ that she qualifies for to
$s_0$. This requires that the student's test scores $r_{t}$ are lower than some value
$\smash{\underline{R_{t}}}(s_0, \succ, q; c)$, representing the most lenient cutoff among schools
that use test $t$ that the student prefers to $s_0$.
Formalizing this is complicated by the presence of $Q_{is}$. For a given value
of $q_s$, consider the worst test score for students with $Q_{is} = q_s$ that clears
the cutoff $c_s$---such a test score is 0 if all students with $Q_{is} = q_s$ clears
the cutoff, and 1 if no students do so:
\[
\smash{\underline{r_{t}}}(s, q_s, c_s) = \min\br{\max\br{(1 +\bar q_s) c_s - q_s, 0}, 1}.
\numberthis
\label{eq:test_score_space_cutoff_q}
\]
Correspondingly, let the most lenient cutoff
for test score $t$ among those
preferred to $s_0$ be \[
\smash{\underline{R_{t}}}(s_0, \succ, q; c) = \min_{s_1: s_1 \succ s_0, t_{s_1} = t} \smash{\underline{r_{t}}}(s_1, q_{s_1},
c_
{s_1}).
\numberthis
\label{eq:most_lenient_cutoff}
\]
If the student has $r_t > \smash{\underline{R_{t}}}(s_0, \succ, q; c)$, then they qualify for some school
that they prefer to $s_0$.
\end{enumerate}
\medskip
On the other hand, suppose that $s_0$ is a test-score school that uses test $t_0$. Then
the corresponding
requirements for identification of $\mu_{s_0}(x)$ are:
\begin{enumerate}[label=(\roman*),wide]
\item The student does not \emph{always} win the lottery at some school $s$ preferred
to $s_0$. Let $L(q,c)$ be the set of schools where students with qualifier $q$ win
the lottery with probability 1. This requirement can be written as $s_0 \succ L
(q,c)$.
\item The student does not qualify for any test-score school $s$ preferred to $s_0$.
As before, this can be written as $r_t < \smash{\underline{R_{t}}}(s_0, \succ, q; c)$ for all $t$.
\item The student qualifies for $s_0$: $r_{t_0} \ge \smash{\underline{r_{t_0}}}(s_0, q_{s_0}, c_{s_0})$.
\end{enumerate}
It is useful for our estimation results to additionally define $r_{s,t}(c)$ as the unique
value among $ \br{
\smash{\underline{r_{t}}} (s, q, c_s): q = 1,\ldots,
\bar q_s}$ that is in $(0,1)$; if no such value exists, then set $r_ {s,t} (c) = 0$: \[
r_{s,t} (c) = \max\pr{\br{\smash{\underline{r_{t}}}(s, q, c_s): q = 1,\ldots,
\bar q_s} \cap [0,1)}.
\numberthis
\label{eq:test_score_space_cutoff}
\]
Intuitively, \cref{eq:test_score_space_cutoff_q,eq:test_score_space_cutoff}
translate cutoffs in $V_s$-space to cutoffs in $R_{s,t_s}$-space for students of different
discrete priority types ($Q_{is}$), as illustrated in
\cref{fig:illustration_testscore_cutoff}.
\begin{figure}[tb]
\centering
\begin{tikzpicture}
\definecolor{colorQ0}{RGB}{200,100,100}
\definecolor{colorQ1}{RGB}{100,200,100}
\definecolor{colorQ2}{RGB}{100,100,200}
\draw[thick] (0,-1.5) -- (6,-1.5);
\foreach \x in {2, 4} {
\draw[thick] (\x,-1.6) -- (\x,-1.4);
}
\node[above, color=colorQ0] at (1,-1) {$q_s = 0$};
\node[above, color=colorQ1] at (3,-1) {$q_s = 1$};
\node[above, color=colorQ2] at (5,-1) {$q_s = 2$};
\node[below] at (2,-1.5) {$\frac{1}{3}$};
\node[below] at (4,-1.5) {$\frac{2}{3}$};
\node[below] at (6,-1.5) {1};
\node[below] at (0,-1.5) {0};
\node[above] at (0,-1.2) {$V_s$};
\fill[color=colorQ0] (0,-1.5) rectangle (2,-1.4);
\fill[color=colorQ1] (2,-1.5) rectangle (4,-1.4);
\fill[color=colorQ2] (4,-1.5) rectangle (6,-1.4);
\fill (3,-1.5) circle (2pt);
\node[below] at (3,-1.8) {$c_s$};
\foreach \i/\y/\colorr/\pos/\q in
{1/0/colorQ0/1/0,2/-1.5/colorQ1/0.5/1,3/-3/colorQ2/0/2} {
\draw[thick, color=\colorr, shift={(8, \y)}] (0,0) -- (6,0);
\fill[color=\colorr, shift={(8, \y)}] (\pos*6,0) circle (2.5pt);
\node[below, color=\colorr, shift={(8, \y)}] at (\pos*6,0) {$\underline{r_t}
(s,\q)$};
\foreach \x/\lab in {0/0, 6/1} {
\draw[thick, shift={(8, \y)}] (\x,-0.1) -- (\x,0.1);
\node[above, shift={(8, \y)}] at (\x,0.1) {\lab};
}
\node[above, shift={(8.5, \y-0.1)}] at (6,0.3) {$R_{t_s}$};
}
\end{tikzpicture}
\caption{Illustration of (\cref{eq:test_score_space_cutoff_q}) and
(\cref{eq:test_score_space_cutoff}). The left axis displays the priority scores $V_s$
for a school with three levels of $Q_s$ and a cutoff at $1/2$. The right axes displays
the test scores for students of different levels of $Q_s$. Every student with $q=0$
would not qualify for $s$, and so the corresponding cutoff in test-score space is
$\smash{\underline{r_{t}}} = 1$. Every student with $q=2$ would qualify for the school, and so the
corresponding test-score cutoff is $\smash{\underline{r_{t}}} = 0$. Students with $q=1$ have interior
test-score cutoffs at $\smash{\underline{r_{t}}} = 1/2 = r_{s,t_s}$. }
\label{fig:illustration_testscore_cutoff}
\end{figure}
Finally, the region on which the aTE $\tau_{s_1, s_0}(x)$ is identified is an intersection
between regions on which $\mu_{s_j}(x)$ is identified. The latter region turns out to
involve hyperrectangles in $R_{i}$. To compactly describe the intersection, it is useful
to define intersecting on coordinate $t$: Given $S_1 \subset [0,1]^T$ and $S_2
\subset [0,1]$, let $
\mathrm{Slice}_t(S_1,
S_2) = \br{s \in S_1 : s_t \in S_2}$ be the set that takes the intersection of the
$t$\th{} coordinate of $S_1$ with $S_2$. The following figure illustrates slicing an
ellipse $S_1 \subset [0,1]^2$ on the first dimension onto an interval $S_2$:
\begin{center}
\begin{tikzpicture}
\draw[thick, ->] (0,-0.1) -- (0,5);
\draw[thick, ->] (-0.1,0) -- (5,0);
\foreach \i in {1,...,4}
{
\draw (-0.1,\i) -- (0.1,\i);
\draw (\i,-0.1) -- (\i,0.1);
}
\node[left] at (0,5) {$[0,1]$};
\node[below] at (5,0) {$[0,1]$};
\fill[blue, opacity=0.3] (2.5,2.5) ellipse (2cm and 1.5cm);
\node[blue] at (4,4) {$S_1$};
\draw[thick, red] (1.5,0) -- (3.5,0);
\foreach \i in {1.5, 3.5}
\fill[red] (\i,0) circle (2pt);
\node[below, red] at (2.5,-0.5) {$S_2$};
\begin{scope}
\clip (1.5,0) rectangle (3.5,5);
\fill[green, opacity=0.5] (2.5,2.5) ellipse (2cm and 1.5cm);
\end{scope}
\draw[dashed, red] (1.5,0) -- (1.5,5);
\draw[dashed, red] (3.5,0) -- (3.5,5);
\node[right, color=green] at (5,2.5) {Slice$_1(S_1, S_2)$};
\end{tikzpicture}
\end{center}
The following theorem formalizes this intuition and computes the regions on which $\tau_
{s_1, s_0}(x)$ is identified.
\begin{restatable}{theorem}{propeligibility}
\label{prop:eligibility}
Fix $\succ, q, c$ and suppress their appearances in $\smash{\underline{r_{t}}}, \smash{\underline{R_{t}}}, L,
\pi^*_s$. Then we have:
\begin{enumerate}[wide]
\item For a lottery school $s_0$, \[E_{s_0}(\succ, q, c) = \begin{cases}
\bigtimes_{t=1}^T [0, \smash{\underline{R_{t}}}(s_0))
, &\text{ if $\pi_{s_0}^* > 0$ and $\smash{\underline{R_{t}}}(s_0) > 0$
for all $t$} \\
\varnothing & \text{ otherwise}.
\end{cases}
\]
\item
For a test-score school $s_0$ using test $t_0$, \[ E_{s_0}(\succ, q, c) =
\begin{cases}
\mathrm{Slice}_{t_0}\pr{\bigtimes_{t=1}^T [0, \smash{\underline{R_{t}}}(s_0)), [
\smash{\underline{r_{t_0}}}(s_0), \smash{\underline{R_{t_0}}}(s_0))
} \\ \quad \quad \quad
\text{if $s_0 \succ L(q,c)$, $\smash{\underline{R_{t}}}(s_0) > 0$ for all $t$, and $\smash{\underline{R_{t_0}}}(s_0) >
\smash{\underline{r_{t_0}}}(s_0)$}\\
\varnothing \quad\quad \text{ otherwise}
\end{cases}
\]
\item Consider two schools $s_0, s_1$ where $s_1 \succ s_0$. Assume that neither $E_
{s_0}$ nor $E_{s_1}$ is empty.
If $s_1$ is a lottery school, regardless of whether $s_0$ is test-score or lottery,
then \[
\bar E_{s_0} \cap \bar E_{s_1} = \bar E_{s_0}.
\]
Otherwise, if $s_1$ is a test-score school with test score $t_1$, regardless of
whether $s_0$ is test-score or lottery, \[
\bar E_{s_0} \cap \bar E_{s_1} = \begin{cases}
\mathrm{Slice}_{t_1} \pr{
\bar E_{s_0}, \br{\smash{\underline{r_{t_1}}}(s_1)}
} , &\text{ if $\smash{\underline{r_{t_1}}}(s_1) = \smash{\underline{R_{t_1}}}(s_0) > 0$}\\
\varnothing & \text{ otherwise}.
\end{cases} \numberthis \label{eq:cutoff_variation_intersection}\]
\end{enumerate}
\end{restatable}
The first claim of \cref{prop:eligibility} characterizes the eligibility set for a lottery
school $s_0$. The second claim analogously characterizes the eligibility set for a
test-score school $s_0$. Both formalize the verbal intuition we have described. The third
claim computes the intersection of the closures of the eligibility sets between two
schools $s_0, s_1$. Depending on whether the more-preferred school $s_1$ is a lottery
school, the intersection takes different shapes. If $s_1$ is a lottery school, then the
intersection is either empty or $\bar E_{s_0}$. Otherwise, the intersection is either
empty or the slice of $\bar E_ {s_0}$ that equals $s_1$'s cutoff on test $t_1$.
\Copy{idlit}{These identification results closely relate to the literature. They follow
\citet{abdulkadiroglu2017impact} closely and extend their Corollary 1 to settings with
heterogeneous treatment effects. In particular, the set of students for whom $R \in \bar
E_s(\succ, Q; c)$ is precisely the set of students for whom the local DA propensity score
for school $s$ is positive. Separately, when there are only test-score schools and no
discrete qualifies ($Q_{is} = 0$), \cref{prop:eligibility} characterizes the same set---up
to a conditionally measure zero set of students---of identified pairwise RD effects as
\citet{marinho2022causal}'s Proposition 1. See \cref{asub:identification_lit} for details.
Despite \cref{prop:eligibility} identifying the same set of effects as shown elsewhere,
disaggregating these effects clarifies their interpretation and highlights potential
drivers of heterogeneity. In particular, \cref{prop:eligibility}(3) rationalizes what we
preview in \cref{defn:lottery_vs_rd} as lottery- or RD-variation. When the more-preferred
school is a lottery school, variation between $s_0$ and $s_1$ is driven by the student
potentially \emph{losing} the lottery at $s_1$. On the other hand, when the more-preferred
school is a test-score school, variation between $s_0$ and $s_1$ is driven by the RD
variation in whether the student just qualifies for $s_1$ or just fails to qualify for
$s_1$. Therefore, the two types of aTEs are different in economically meaningful ways:
There is no lottery variation between a less-preferred lottery school and a more-preferred
test-score school, and no test-score variation between a less-preferred test-score school
and a more-preferred lottery school.}
Examining the disaggregated aTEs in turn illustrates certain perils of simple aggregation
schemes. \Cref{prop:eligibility}(3) shows that aTEs driven by lottery variation are
identifiable for a set of students with positive measure, but aTEs driven by test-score
variation (\cref{eq:cutoff_variation_intersection}) are only identifiable for a set of
students with zero measure. As a result, when we aggregate treatment effects such that
each student receive equal weights (see \cref{def:ATE_def} for details), we would
effectively only aggregate the lottery-driven aTEs. We state this result informally as
\cref{cor:agg}, which is formally stated and proved as \cref{cor:aggApp}.
\begin{cor}
\label{cor:agg}
Under suitable assumptions, an identified aggregated treatment effect that weighs each
student equally either puts weight solely on aTEs driven by lottery variation or
aggregates over a set of students of measure zero.
\end{cor}
\Cref{cor:agg} is a concerning observation for interpreting estimates of aggregate
treatment effects, as many of these estimands necessarily put zero weight on the aTEs
driven by the RD variation. In contrast, our recommended aggregation \cref{eq:maximal}
avoids this feature.
Translated to estimation, \cref{cor:agg} implies that estimators for aggregate treatment
effects put \emph{vanishing} weight on the aTEs driven by RD variation, since these
estimators can only take comparisons local to a cutoff. While in finite samples these
weights can still be positive and nontrivial, these implicit weights do depend on the
sample size (and on bandwidth tuning parameters); moreover, the larger these weights are
(equivalently, the larger the bandwidth), the more susceptible to bias the estimator is.
The share of influence from RD-driven aTEs is larger in smaller samples than in larger samples, which makes such aggregate treatment effects potentially difficult to interpret. Indeed, the vanishing weight issue affects popular estimation
approaches proposed in the literature.\footnote{\label{fn:yata}\Copy{fnyata}{These
issues are not
unique to
aggregation in the school-choice context. In settings---for instance studied in
\citet{narita2021algorithm}---where the distribution of treatment assignment is a function
of $X$, under continuity assumptions, RD-type treatment effects are identified on the
boundary of sets of the form $\br{X : \P(D=1\mid X) > 0}$. Similarly, aggregating these
effects with effects in the interior may place vanishing weight on these RD-type
effects, since the boundary has zero measure.}}
The next subsection returns to our extended example and illustrate that the regression
estimator in \citet{abdulkadiroglu2017impact} puts vanishing weights on the RD-driven
aTEs. To be sure, these estimators can be simple to implement and more efficient when
treatment effects are homogeneous \citep{goldsmith2021estimating,angrist1998estimating};
motivated by these advantages of regression, we also develop a simple diagnostic for the
weight put on RD-driven aTEs for linear estimators.
\subsection{Implications for regression estimators and a diagnostic}
\label{sub:prop_scores}
The estimation approach proposed by \citet{abdulkadiroglu2017impact} converges to
aggregations of treatment effects that ignore the RD variation in the limit. We illustrate
this with our extended example in \cref{ex:running_ex}. First, we make a few additional
assumptions on the example.
\begin{exsq}[Additional setup for the extended example]
Here, we additionally assume that the probability of qualifying for $s_3$ for any student
is approximately equal to $1/2$, in order to compute the local DA propensity scores
\citep{abdulkadiroglu2017impact}. Suppose we are interested in the treatment effect of
$s_2, s_3$ relative to $s_1, s_0$. Consider the treatment indicator $\tilde D_i =
\one(\text{$i$ is assigned to either $s_2$ or $s_3$ at cutoff $c$})$. For simplicity, we
remove
heterogeneity of potential outcomes at the school level and assume $Y_T = Y(s_2) = Y(s_3)$
and $Y_C = Y(s_1) = Y(s_0)$.
\end{exsq}
Roughly speaking, for a region of test score $A \subset [0,1]$,
\citet{abdulkadiroglu2017impact} compute $\P(\tilde D_i = 1 \mid {\succ}, R \in A)$ by
counting those $R$'s near a cutoff as having $1/2$ probability of falling on either side.
``Near a cutoff'' is determined by a bandwidth parameter $h > 0$.
To be more concrete, we partition the space of test scores into five regions, with a
bandwidth parameter $h > 0$. Regions \text{II}{} and \text{IV}{} are small bands around the
cutoffs $c_1, c_2$, and regions $\text{I}, \text{III},\text{V}$ are large regions in between the
cutoffs:
\begin{center}
\begin{tikzpicture}
\pgfmathsetmacro{\totalwidth}{5*3}
\pgfmathsetmacro{\third}{\totalwidth/3}
\pgfmathsetmacro{\posA}{0}
\pgfmathsetmacro{\posB}{\third - 0.04*\totalwidth}
\pgfmathsetmacro{\posC}{\third}
\pgfmathsetmacro{\posD}{\third + 0.04*\totalwidth}
\pgfmathsetmacro{\posE}{2*\third - 0.04*\totalwidth}
\pgfmathsetmacro{\posF}{2*\third}
\pgfmathsetmacro{\posG}{2*\third + 0.04*\totalwidth}
\pgfmathsetmacro{\posH}{\totalwidth}
\draw[thick,->] (-0.1,0) -- (\totalwidth+0.1,0);
\fill[blue!20] (\posA, -0.2) rectangle (\posB, 0.2);
\fill[red!20] (\posB, -0.2) rectangle (\posD, 0.2);
\fill[green!20] (\posD, -0.2) rectangle (\posE, 0.2);
\fill[yellow!20] (\posE, -0.2) rectangle (\posG, 0.2);
\fill[orange!20] (\posG, -0.2) rectangle (\posH, 0.2);
\foreach \x/\label in {\posA/0, \posB/{\frac{1}{3}-h}, \posD/{\frac{1}{3}+h}, \posE/{\frac{2}{3}-h}, \posG/{\frac{2}{3}+h}, \posH/1} {
\draw[thick] (\x,0.2) -- (\x,-0.2) node[below] {$\label$};
}
\node at ({(\posA+\posB)/2},0.5) {I};
\node at ({(\posB+\posD)/2},0.5) {II};
\node at ({(\posD+\posE)/2},0.5) {III};
\node at ({(\posE+\posG)/2},0.5) {IV};
\node at ({(\posG+\posH)/2},0.5) {V};
\end{tikzpicture}
\vspace{1em}
\end{center}
The local deferred acceptance propensity scores, as a function of the region that the test score $R$ falls into,
are as follows:\footnote{To explain the propensity score calculation, consider a student of type $B$ with test
scores in \text{II}:
\begin{itemize}[wide]
\item They fail to qualify for $s_2$ when her test score is in \text{II}{}
with probability one.
\item Heuristically, they qualify for $s_1$ with probability 0.5 since her test
score is in \text{II}{} and near the cutoff for $s_1$.
\item $s_1$ is the only control school preferred to $s_3$
\item They qualify for $s_3$ (via lottery) with probability 0.5.
\item Thus their probability of being treated is $\psi = 0 + (1-0.5)
\cdot 0.5 =
0.25$.
\end{itemize}
This computation is heuristic with $h > 0$, but as we take the limit $h \to 0$,
the probability $\P(\tilde D_i =1 \mid \text{II}, {\succ_B}) \to 0.25$. }
\begin{center}
\vspace{1em}
\begin{tabular}{cccccc}
\toprule
& \text{I} & \text{II} & \text{III} & \text{IV} & \text{V} \\ \midrule
$A$ &0.5&0.5&0.5&0.75&1 \\
$B$ &0.5&0.25&0&0.5&1 \\
$C$ &0.5&0.5&0.5&0.75& 1 \\\bottomrule
\end{tabular}
\vspace{1em}
\end{center}
Let $v \in \br{A, B, C} \times \br{\text{I}, \text{II}, \text{III},\text{IV}, \text{V}}$ denote a student
type, according to preferences and coarsened test scores. Let $\psi(v)$ denote the
corresponding local propensity score collected in the above table. Consider the population
regression\footnote{If $\psi(V)$ is exactly equal to $\E[D \mid V]$, then this regression
is equivalent to the more familiar control-for-propensity-score regression
\citep{angrist1998estimating} \[ Y = \alpha + \tau \tilde D + \gamma \psi(V) + \epsilon
\]
by Frisch--Waugh. This latter specification conforms with the specification used for
intent-to-treat effects in \citet{abdulkadiroglu2017impact} (i.e., the reduced form in
their IV specification).} \[Y = \tau (\tilde D - \psi(V)) + \epsilon \numberthis
\label{eq:ols_spec}
\text{ such that }
\tau = \frac{\E[(\tilde D - \psi(V)) Y]}{\E[(\tilde D - \psi(V))^2]}.
\]
We compute in \cref{asub:prop_score_app} that
\begin{align*}
\tau = \sum_{v \in \br{A, B, C} \times \br{\text{I}, \text{II}, \text{III},\text{IV}, \text{V}}}
\frac{\P(v)(1-\psi(v))\psi(v) }{\sum_u \P(u) (1-\psi(u))\psi(u)} \E[Y_T - Y_C \mid
(\succ, R)
\in v] +
\text{Bias}(h).
\numberthis \label{eq:bias_prop_score}
\end{align*}
where $\text{Bias}(h) \to 0$ as $h \to 0$.\footnote{We note that because of the regression
specification (\cref{eq:ols_spec}), the estimand weighs according to $(1-\psi (v))\psi(v)$.
As a result, this aggregation does not weigh each student equally, but nevertheless the
conclusion of \cref{cor:agg} continues to hold for such weighting schemes.}
\Cref{eq:bias_prop_score} shows that the implicit aggregation recovers weighted averages
of treatment effects (over preference-test score region cells) up to a bias that vanishes
as $h \to 0$. However, the weights on regions $\text{II}$ and $\text{IV}$ also vanish as $h \to
0$, since the corresponding $\P(v) \to 0$. Translated to estimation, this
implies that the asymptotic unbiasedness of propensity-score estimators requires $h = h_N
\to 0$ at appropriate rates, yet the part of the estimator driven by variation from
regression discontinuity then becomes asymptotically negligible, as long as there is
lottery-driven variation.
To further relate to finite samples, note that under typical assumptions---as confirmed by
our estimation results in \cref{sub:est}---the popular locally linear regression estimator
for regression-discontinuity-type variation converges at the rates no faster than
$N^{-2/5}
\gg N^{-1/2}$, reflecting that identification for the conditional average treatment effect
at the cutoff is \emph{irregular} \citep{khan2010irregular}. As a result, asymptotically,
estimators for the RD aTEs are much noisier than those for the lottery-driven aTEs, and
pooling them with inverse-variance-type weights in (\cref{eq:bias_prop_score}) results in
diminishing weight on the RD aTEs.
\subsubsection{Diagnostic}
\label{sub:diagnostic}
Motivated by this decomposition, we introduce a simple diagnostic for regression-based
procedures that gives upper and lower bounds on the weight put on students who are
subject to RD variation. For simplicity, we limit to considering binarized comparisons.
That is, there is some treatment $\tilde D_i = \sum_{s \in S_1 \subset S} D_
{is}$ corresponding to a subset $S_1$ of the schools $S$, where a student is considered
treated if they are matched to a school in $S_1$ and untreated otherwise.
We consider linear estimators of the form \[
\hat\tau = \sum_{i=1}^n \hat w_i \cdot Y_i,
\]
where $\hat w_i$ is a function of $X_{1},\ldots, X_N, \tilde D_1,\ldots, \tilde D_N$. Any
linear regression estimator with $Y_i$ on the left-hand side and functions of $X_i, \tilde
D_i$ on the right-hand side can be written this way. We will also assume the estimator is
associated with some chosen bandwidth parameter $h_N$.
Using the characterization in \cref{prop:eligibility}, we label student observations by
whether they are subject to lottery variation:
\begin{defn}
\label{defn:diagnostic}
For some user-chosen bandwidth parameter $h_N$ used implicitly to construct
$\hat\tau$, we will consider an observation $i$ to be \emph{possibly subjected} (denoted
by $\overline{\mathrm{RD}}_i = 1$) to RD variation if there exists $s_1 \succ_i s_0$ such
that:
\begin{itemize}
\item $s_1$ is a test-score school using test score $t_1$
\item Exactly one of $s_1$ and $s_0$ belongs to the treated group $S_1$
\item $\bar E_{s_0}(\succ_i, Q_i, C_N) \cap \bar E_{s_1}(\succ_i, Q_i, C_N) \neq
\varnothing$.
\item $R_i \in \mathrm{Slice}_{t_1}\pr{
\bar E_{s_0}(\succ_i, Q_i, C_N), [\smash{\underline{r_{t_1}}}(s_1, Q_{is_1},
C_{s_1, N}) - h_N, \smash{\underline{r_{t_1}}}(s_1, Q_{is_1},
C_{s_1, N}) + h_N]
}.$
\end{itemize}
We consider observation $i$ to be \emph{definitely subjected} (denoted by $\underline{
\mathrm{RD}}_i = 1$) to RD variation if it is
possibly subjected and there \emph{does not exist} $s_1 \succ_i s_0$ where:
\begin{itemize}
\item $s_1$ is a lottery school
\item Exactly one of $s_1$ and $s_0$ belongs to the treated group $S_1$
\item $R_i \in \bar E_{s_0}(\succ_i, Q_i, C_N) \cap \bar E_{s_1}(\succ_i, Q_i, C_N).$
\end{itemize}
\end{defn}
Intuitively, those with $\overline{\mathrm{RD}}_i = 1$ are individuals who are close
enough to the cutoff used by $s_1$ such that, \emph{on some lottery realizations}, they
would be assigned to a test-score school $s_1$ if they clear the cutoff, and to some
school $s_0$ of the opposite treatment type otherwise. Those with
$\underline{\mathrm{RD}}_i = 1$ are those who do so on \emph{all} lottery realizations.
We can then define \[
\bar p_{\text{RD}} \equiv \frac12 \sum_{i=1}^n \hat w_i (2 \tilde D_i - 1)
\bar{\mathrm{RD}}_i \text{ and } \underline p_{\text{RD}} \equiv \frac12 \sum_{i=1}^n \hat
w_i (2 \tilde D_i - 1)
\underline{\mathrm{RD}}_i
\]
as the upper and lower bounds for the weight assigned to the RD variation, and these can
be implemented by using $(2 \tilde D_i - 1) {\mathrm{RD}}_i$ as left-hand side variables
in the regression. Formally, these estimates are interpreted as the change in $\hat\tau$
if every treated student with $\bar{\mathrm{RD}}_i = 1$ (resp. $\underline{\mathrm{RD}}_i
= 1$) has their outcome increase by 1/2 unit, and every untreated student with $\bar{
\mathrm{RD}}_i = 1$ has their outcome decrease by 1/2 unit.\footnote{Our assumptions on
linear estimators do not rule out estimators that overly extrapolate (i.e. some $\hat w_i$
has the wrong sign: It is negative for $\tilde D_i = 1$ and positive for $\tilde D_i =
0$). As a result, it is possible that one or both of $\bar p_{\text{RD}}$ and $\underline
p_{ \text{RD}} $ are negative, or that $\bar p_{\text{RD}} < \underline p_{\text{RD}} $.
These unpleasant realizations serve as an additional diagnostic. They cast doubt on the
interpretation of the linear estimator $\hat\tau$, as they reveal that a subpopulation
chosen solely on the basis of $X_i$ has weights that are wrong-signed. See \citet{chen2025potentialweightsimplicitcausal} for general diagnostics in regression estimators. }
\citet{abdulkadiroglu2017impact} study the impact of attending a ``Grade A school'' (one
that receives grade A on the school district's report card for school quality) in New York
City versus attending a non-Grade A school. While we do not have access to their data,
\citet{abdulkadiroglu2017impact} (Appendix Figure D.1) report that about 9,000 students
out of 32,866 students have local propensity scores exactly equal to $1/2$, under their
bandwidth choice $h_N$. This means that these students are solely subject to RD
variation between the treatment schools $S_1$ (Grade A schools in New York City) and the
control schools. Correspondingly, these individuals would be \emph{definitely subjected}
according to \cref{defn:diagnostic}. Thus, this puts a lower bound of about
$9,000/32,866\approx 0.3$ on $\underline p_{\text{RD}}$ for their empirical
application.\footnote{$9,000/32,866\approx 0.3$ would be the weight on these observations
if every observation were weighted equally. Since the specification
(\cref{eq:ols_spec}) weighs students proportionally to $(1-\psi)\psi$, those with
propensity score at exactly $1/2$ receive the highest weights. Thus, in this case, 0.3
serves as a lower bound.} Assuming that the bias term is sufficiently small, we may
conclude that \citet{abdulkadiroglu2017impact} estimate aggregations of aTEs that put
nontrivial weight on the RD-driven aTEs.
In general---though especially when $\bar p_{\text{RD}}, \underline p_{\text{RD}}$ imply
unreasonably small weight on the RD-driven aTEs---researchers may wish to unpack
heterogeneity further and isolate aTEs that are driven by RD variation, along the lines
in \cref{sec:nonstoch}. If researchers have in mind weights for an aggregate treatment
effect that they prefer, they can also aggregate these finer treatment effects manually.
The next section discusses estimation and inference for maximal aggregations atomic treatment effects between two schools (\cref{eq:maximal}).
We close this section with two miscellaneous discussions.
\begin{rmksq}[Design-based inference]
In some school choice markets (e.g. Denver Public Schools and Boston Public Schools), all
schools are lottery schools, and so there is no RD-driven variation. In such settings, we
can directly use the lottery variation to conduct design-based estimation and inference.
Design-based approaches have the benefit that we do not need to assume assignment
mechanisms have the cutoff structure in \cref{as:limitcutoff}, nor do we need to consider
any asymptotic notions of identification. \Copy{rubin}{A previous draft of this paper
(\url{https://arxiv.org/abs/2112.03872v2}) contains results on design-based estimation and
connect propensity-score based regression estimators to Horvitz--Thompson estimators.
These results apply to arbitrary school assignment algorithms, not limited to the deferred
acceptance algorithms that we consider. On the other hand, a design-based approach conditions on the set of students observed and considers randomness solely through the lottery; thus it would not capture the uncertainty in generalizing to matchings of future students \citep{rubin1974estimating}. }
\end{rmksq}
\begin{rmksq}[Noncompliance] Our results define the treatment as the school that a student
is matched to by the assignment mechanism. This may not be the school that a student
eventually attends---some students may attend a school that is not in the centralized
assignment system. Atomic treatment effects are thus interpreted only as intent-to-treat
effects. Studying treatment effects of schools that students eventually attend amounts
to analyzing an instrumental variable setting with heterogeneous treatment effects,
multiple treatments, and a multivalued discrete instrument. Such settings---even without
the additional complication of school choice mechanisms---remain an area of active
research
\citep{lee2018identifying,mogstad2020policy,behaghel2013robustness,kirkeboen2016field,kline2016evaluating}.
We leave such analyses to future work.\footnote{\Copy{fnhomo}{In contrast,
\citet{abdulkadiroglu2017impact} are able to consider noncompliance because they assume
no unobserved heterogeneity in treatment effects \citep{mogstad2024instrumental}.}}
\end{rmksq}
\section{Estimation and inference}
\label{sub:est}
Having characterized the identified atomic treatment effects, we advocate that
practitioners first estimate more granular aggregations of aTEs and then summarize these
aggregations further. Towards this goal, this section provides asymptotic theory for the
maximal aggregation between two schools (\cref{eq:maximal}).
To that end, consider two schools $s_1, s_0$ and all students who declare $s_1 \succ_i
s_0$. We are interested in the estimand \[
\tau_{s_1 \succ s_0}(c) = \E[Y(s_1) - Y(s_0) \mid R \in \bar E_{s_0}(\succ, Q; c)
\cap \bar E_{s_1}(\succ, Q; c), s_1 \succ s_0].
\]
Constructing estimators for $\tau_{s_1 \succ s_0}$ depends on whether $s_1$ is a lottery
school, but they share some common structure.
For a given set of cutoffs $c$, let $J_i(c) = 1$ indicate the set of $X$ values such that
\begin{enumerate}
\item $s_1 \succ s_0$
\item $R_i \in \bar E_{s_0}(\succ_i, Q_i; c)
\cap \bar E_{s_1}(\succ_i, Q_i; c)$, if $s_1$ is a lottery school
\item $R_{i,t} \in \br{r_t : r \in \bar E_{s_0}(\succ_i, Q_i; c)
\cap \bar E_{s_1}(\succ_i, Q_i; c)}$ for all $t \neq t_1$, if $s_1$ is a test-score
school that uses test $t_1$.
\end{enumerate}
We can thus write \[
\tau_{s_1 \succ s_0}(c) = \begin{cases}
\E[Y(s_1) - Y(s_0) \mid J_i(c) = 1] , &\text{ if }s_1 \text{ is a lottery
school}\\
\E[Y(s_1) - Y(s_0) \mid J_i(c) = 1, R_{it_1}=r_{s_1,t_1}(c)] &\text{ if }s_1
\text{ is a test-score school.}
\end{cases}
\numberthis
\label{eq:estimands_estimation}
\]
recalling \cref{eq:test_score_space_cutoff}. $\tau_{s_1\succ s_0}(c)$ corresponds to the
maximal aggregation of aTEs for $s_1\succ s_0$ when the population cutoffs equal $c$.
To construct analogue estimators for $\tau_{s_1\succ s_0}$, we account for the fact
that---depending on the lottery results---students with $J_i(c) = 1$ may be matched to
neither $s_1$ nor $s_0$, e.g., if they qualify for some lottery school they prefer. Let
$D_ {i1} (U_i; c)$ indicate the event---as a function of the lottery numbers $U_i$---that
student $i$ fails to qualify for any lottery school $s
\succ_i s_1$ and---if $s_1$ is lottery---additionally qualifies for $s_1$. Let
$D_ {i0} (U_i; c)$ indicate
the event that student $i$ fails to qualify for any lottery school $s\succ_i s_0$ and---if
$s_0$ is lottery---additionally qualifies for $s_0$. Likewise, let $\pi_{ij}(c) = \int D_
{ij} (u;
c)\,dF_U(u)$ be the fixed-cutoff expectation of these events.\footnote{These objects are
defined formally in \eqref{eq:def_d1}.} By construction, for all
fixed cutoff $c$, we have that $\E\bk{ J_i(c)
\frac{D_{i1}(U_i; c) Y_i}{\pi_{i1}(c)} } = \E\bk{J_i(c) Y_i(s_1)}
$ and that $\E\bk{ J_i(c)
\frac{D_{i0}(U_i; c) Y_i}{\pi_{i0}(c)} } = \E\bk{J_i(c) Y_i(s_0)}$.
Denote $Y_i^{(j)}(c) \equiv \frac{D_{ij}(U_i; c) Y_i}{\pi_{ij}(c)} $.
We thus have two natural estimators for $\tau_{s_1\succ s_0}(c)$, depending on whether
$s_1$ is a lottery school. If $s_1$ is a lottery school, then an inverse propensity
estimator takes the form \[
\hat\tau_{s_1 \succ s_0} = \pr{\frac{1}{N}\sum_{i=1}^N J_i(C_N)}^{-1}\pr{\frac{1}
{N}\sum_ {i=1}^N J_i (C_N) \br{Y_i^
{(1)}(C_N) - Y_i^{(0)}(C_N)}}. \numberthis
\label{eq:ipw_lottery}
\]
\Cref{eq:ipw_lottery} is simply the inverse-weighting estimator among those with $J_i(C_N)
= 1$.
\Copy{llr}{On the other hand, if $s_1$ is a test-score school and we shorthand $\rho(c) =
r_ {s_1,
t_1}(c)$, then \eqref{eq:estimands_estimation} is akin to an RD estimand among those with
$J_i(c) = 1$. We consider the analogue of local linear regression for this estimand. For
technical reasons, we limit our consideration to a uniform kernel. Specifically, fix some
bandwidth $h_N \to 0$, we let \[
\hat\tau_{s_1 \succ s_0} = \hat\beta_+(h_N) - \hat \beta_-(h_N)
\numberthis
\label{eq:llr_estimand}
\]
where, for $J(c,h)$ defined in \cref{eq:jdef}, \begin{align*}
\hat \beta_+(h_N) =
\argmin_{b_0} \min_{b_1} \sum_{i: R_{i, t_1} \in [\rho(C_N),
\rho(C_N) + h_N]} J_i(C_N, h_N) \pr{Y_i^{(1)}(C_N) - b_0 - (R_{it_1} - \rho(C_N))
b_1}^2.
\end{align*} The estimator for the left-limit, $\hat \beta_-(h_N)$, is defined
analogously. \Cref{eq:llr_estimand} implements local linear regression among those with
$J_i(C_N, h_N) = 1$, which essentially indicates those with $J_i(C_N) = 1$ and have test
$R_{it_1}$ within bandwidth $h_N$ of the cutoff $\rho(C_N)$. We focus on local linear
regression since it is a standard choice in the regression-discontinuity
literature \citep{cattaneo2022regression}. We speculate that the same analysis likely
extends to local polynomial regression with a uniform kernel as well.}
\Copy{estimation}{ Relative to standard setups, estimation is complicated by the fact that
the estimators feature the finite-sample cutoffs $C_N$, computed from all the data.
Since $C_N$ enters all terms in sample averages, these terms are no longer independent,
precluding a standard asymptotic argument. Our asymptotic results account for the effect
of a stochastic $C_N$; they can be understood as applications of two-step GMM analyses
where we verify certain stochastic equicontinuity conditions in $C_N \to c$.
Interestingly, the asymptotic distributions of the estimators do not depend on the
asymptotic distribution $\sqrt{N}(C_N - c)$, meaning that, for instance, we do not need
to adjust for the fact that $C_N$ is random in calculating standard errors. Our results
thus justify procedures in the literature that treat $C_N$ as fixed.\footnote{The sense
in which they are valid is nuanced. See the discussion after \cref
{thm:lottery_main_estimation}. }}
Having introduced the estimator, we turn to assumptions. The key assumption is a
substantive restriction on school capacities and the population distribution of student
characteristics, such that the large-market cutoffs are not in certain knife-edge
configurations. This does limit the uniform validity of our asymptotic results.
\begin{restatable}[Population cutoffs are
interior]{as}{interior}
\label{as:interior}
The distribution of student observables satisfies
\cref{as:limitcutoff}, where the population cutoffs $c$ satisfy:
\begin{enumerate}[wide]
\item School $s_1$ is neither undersubscribed nor impossible to qualify for: $\rho
(c)
\in
(0,1)$. If $s_1$ uses the same test as $s_0$, then its cutoff is more stringent
$\rho (c) > r_ {s_0, t_1}(c)$.
\item For each school $s$, the cutoff \[
c_s \not\in \br{\frac{1}{\bar q_s + 1}, \ldots, \frac{\bar q_s}{\bar q_s + 1}, 1}.
\]
\item If $c_s =0$, then $s$ is eventually undersubscribed: $\P
\pr{C_
{s,N} = 0} \to 1.$
\item If two schools $s_3, s_4$ use the same test $t$, then their test score
cutoffs are different, unless both are undersubscribed: If $r_{s_3,t}(c) = r_{
s_4,t}(c)$ then $c_{s_3} = c_{s_4} = 0,$ where $r_{s,t}$ is defined in
\cref{eq:test_score_space_cutoff}.
\end{enumerate}
\end{restatable}
\cref{as:interior} rules out populations where the large-market cutoffs from
\citet{azevedo2016supply} are on the boundary of certain sets. The first assumption
simply says that the school $s_1$ is not undersubscribed and not strictly easier to
qualify for than $s_0$. The second assumption rules out a scenario where everyone with
$Q=q$ does not qualify for $s$ regardless of their tiebreakers and everyone with
$Q=q+1$ does qualify for $s$ regardless of their tiebreakers. The third assumption
says that undersubscribed schools are eventually undersubscribed, for which it
suffices to impose that the population capacity of a school is not \emph{exactly} at
the threshold making the school undersubscribed.\footnote{This assumption is stronger
than the $O_p(N^{-1/2})$ convergence of the cutoffs that we assume. However, the
assumption holds generically for sufficiently large population school capacities.
Precisely speaking, suppose some school $s$ with population capacity $q_s$ is
undersubscribed in the population, but the probability that it is undersubscribed in
samples of size $N$ does not tend to one (and so violates the third assumption). Then
we may add a little slack to the school capacity---for any $\epsilon > 0$, making the
capacity $q_s +
\epsilon$ instead---to guarantee that the third assumption holds.
Intuitively, adding $\epsilon$ to the capacity adds $O(\epsilon N)$ seats to the
school in finite samples, but the random fluctuation of the market generates variation
in student assignments of size $O (\sqrt{N}).$} Lastly, the fourth assumption assumes
that the population cutoffs in test-score space are not exactly the same for two
schools that uses the same test.\footnote{This is Assumption 2 in
\citet{abdulkadiroglu2017impact}.} Collectively, these assumptions rule out
adversarial scenarios where, for instance, the population intersection is empty, $\bar
E_ {s_0}
({\succ_i}, Q_i,
c) \cap \bar E_{s_1} ({\succ_i}, Q_i, c) = \varnothing$, but the sample intersection is
nonempty with probability non-vanishing in $N$, $\P\bk{\bar E_ {s_0} (\succ_i, Q_i,
C_N)
\cap \bar
E_{s_1} (\succ_i, Q_i, C_N) \neq \varnothing} \not\to 0$.
Additionally, we maintain a few technical assumptions,
\cref{as:bounded_density,as:ctsdensity,as:second_moment,as:cts_diff,as:moment_strong},
stated in \cref{asec:estimation}. These assumptions assert that the distribution of
test scores and lotteries are suitably smooth, have suitably smooth conditional means,
and have bounded moments.
We have the following result for lottery comparisons.
\begin{restatable}{theorem}{thmlotterymain}
\label{thm:lottery_main_estimation} Suppose $s_1$ is a lottery school. Suppose $s_1, s_0$
are comparable under the limiting cutoffs, i.e., $\E[J_i(c)] > 0$. Then, under
\cref{as:limitcutoff,as:bounded_density,as:moment_strong,as:interior}, for $g(X_i,
Y_i; c) \equiv J_i (C_N) \br{Y_i^{(1)}(C_N) - Y_i^{(0)}(C_N)}$,
\begin{align*} &\sqrt{N}
\pr{\hat\tau_{s_1 \succ s_0} - \tau_{s_1\succ s_0}(C_N)} \\& = \frac{1}{
\sqrt{N}}
\sum_{i=1}^N \br{\frac{g(X_i, Y_i; c) - \E[g(X_i, Y_i; c)]}{\E[J_i(R_i; c)]} -
\frac{\E [g
(X_i, Y_i; c)]}{(\E J_i(R_i; c))^2} (J_i(R_i; c) - \E J_i(R_i; c))} + o_P(1)
\numberthis \label{eq:influence}\\
&\dto \Norm\pr{
0, \frac{\var\pr{g(X_i, Y_i; c) \mid J_i(R_i; c) = 1}}{\E[J_i(R_i; c)]}
}.
\numberthis \label{eq:clt}
\end{align*}
\end{restatable}
\Cref{thm:lottery_main_estimation}, proved in \cref{asub:lottery_ate}, shows that the
scaled estimation error
$ \sqrt{N}
\pr{\hat\tau_{s_1 \succ s_0} - \tau_{s_1\succ s_0}(C_N)}$ is equivalent to a scaled
sample mean of i.i.d. random variables (\cref{eq:influence}), which attains a central
limit theorem (\cref{eq:clt}). The influence function representation (\cref{eq:influence})
is conducive to deriving joint convergence statements (\cref{rmk:joint}).
One subtlety here is that the distribution of $\hat\tau_{s_1 \succ s_0}$ is centered at
the random estimand $\tau_{s_1 \succ s_0} (C_N)$, corresponding to the causal effect
holding fixed the cutoff at the random value $C_N$, rather than at the population cutoff
$c$, $\tau_ {s_1 \succ s_0}(c)$. The difference between the two estimands is of order $O
(1/\sqrt{N})$,
$
\sqrt{N} (\tau_ {s_1
\succ s_0}(C_N) - \tau_{s_1 \succ s_0} (c)) = O_P(1)$, and cannot be ignored. Controlling
this difference is feasible since the asymptotic distribution $\sqrt{N}(C_N - c)$ can be
analyzed along the lines in \citet{agarwal2018demand}, and we leave it to future work.
It is reasonable to consider $\tau_{s_1\succ s_0}(C_N)$ as the target of inference,
since whether a student chooses between $s_1$ and $s_0$ is ultimately a function of
$C_N$ rather than of $c$. Doing so has the additional convenience that the asymptotic
distribution does not depend on the limit distribution of $\sqrt{N}(C_N -c)$---and thus
standard errors do not need to be adjusted.
Analogously, we have the following result for RD-type comparisons. To facilitate its
statement, let
$
\check\tau_
{s_1\succ s_0} = \check\beta_+(h_N) - \check \beta_-(h_N)
$
where
\begin{align*}
\check \beta_+(h_N) \equiv \argmin_{b_0} \min_{b_1} \sum_{i:\one(R_{i, t_1} \in [\rho
(c),
\rho(c) + h_N])} J_i(c) \pr{Y_i^{(1)}(c) - b_0 - (R_{it_1} - \rho(c))
b_1}^2
\end{align*}
and similarly for $\check\beta_-$. The local linear regression estimator $\check\tau_
{s_1\succ s_0}$ uses the population cutoff $c$ to construct $J_i$ and $\rho(c)$, and thus
its asymptotic properties are well-understood \citep{hahn2001identification}.
\begin{theorem}
\label{thm:estimation_main}
Under
\cref{as:limitcutoff,as:interior,as:bounded_density,as:ctsdensity,as:second_moment,as:cts_diff}, assuming $N^{-1/2} = o
(h_N)$ and $h_N = O(N^{-d})$ with $1/5 < d < 1/4$, \begin{align*}
\sqrt{Nh_N} (\hat \tau_{s_1\succ s_0} - \tau_{s_1\succ s_0}(c)) &= \sqrt{Nh_N}
(\check\tau_
{s_1\succ s_0} - \tau_{s_1\succ s_0}(c)) + o_p (1) \\
&\dto \Norm\pr{0,
\frac{4}{\P(J_i(c) = 1) f(\rho(c))} (\sigma_+^2 + \sigma_-^2) }
\end{align*}
for
$\sigma_+^2 = \var(Y_i(s_1) \mid J_{i}(c) = 1, R_{it_1} = \rho(c))$, $
\sigma_-^2 = \var(Y_i(s_0) \mid J_{i}(c) = 1, R_{it_1} = \rho(c))$, and $f(r)$ the
density of $R_{it_1} \mid J_i(c) = 1$.
\end{theorem}
The proof of \cref{thm:estimation_main} is relegated to \cref{asec:estimation}
(\cref{athm:equiv}).
Like \cref{thm:lottery_main_estimation}, \cref{thm:estimation_main} shows that $
\sqrt{Nh_N} (\hat \tau_{s_1\succ s_0} - \tau_{s_1\succ s_0}(c))$ is asymptotically
equivalent to a quantity that is easier to analyze, $\sqrt{Nh_N} (\check\tau_ {s_1\succ
s_0} - \tau_{s_1\succ s_0}(c))$. A central limit theorem---standard in the RD literature
\citep{hahn2001identification}---on this quantity is then used to derive the asymptotic
distribution. Compared to \cref{thm:lottery_main_estimation}, we state this theorem by
centering at $\tau_{s_1\succ s_0} (c)$. However, since we scale by $\sqrt{Nh_N} \ll
\sqrt{N}$, the difference between the estimands is immaterial, as $\sqrt{Nh_N} (\tau_{s_1
\succ s_0}(c) - \tau_{s_1\succ s_0}(C_N)) = o_P (1)$.
\Copy{rdliterature}{\Cref{thm:estimation_main} is related to the literature on regression
discontinuity with unknown or estimated cutoffs \citep
{hansen2000sample,porter2015regression}. In some settings \citep{card2008tipping}, the
cutoff that determines treatment assignment is unknown and must be estimated. In those
settings the cutoff turns out to be estimable at faster-than-$\sqrt{N}$ rates, and its
estimation has no effect on subsequent asymptotics. This setting differs from ours
since, here, treatment assignment is determined by $C_N$ and not $c$, and $C_N$ is not
superconsistent for $c$. Nevertheless, because nonparametric estimators for RD effects
converge at slower-than-$\sqrt{N}$ rates, the randomness in $C_N$ also does not affect
the asymptotics of the RD estimators. }
We conclude this section by outlining how these asymptotic results are useful to construct
inferential statements for further aggregations of aTEs.
\begin{rmksq}[Joint convergence and further aggregation]
\label{rmk:joint}
The representation results in \cref{thm:lottery_main_estimation,thm:estimation_main}
enable us to consider further aggregation, along the lines described in
\cref{sec:nonstoch}. Both results represent the scaled estimation error as the
following \emph{asymptotically linear} representation: For some $r_N \to \infty$ and
$\E\psi_N (X_i, Y_i;
c) = 0,$\footnote{See \cref{eq:influence_fn_rd} for the influence function
representation of the local linear regression estimator.}
\[
r_N (\hat\tau_{s_1 \succ s_0} - \tau_{s_1 \succ s_0}) = \frac{1}{r_N} \sum_{i=1}^N
\psi_N(X_i, Y_i; c) + o_P(1).
\]
The subsequent asymptotic normality is derived by observing that $\frac{1}{r_N} \sum_
{i=1}^N
\psi_N(X_i, Y_i; c) \dto \Norm(0, \sigma^2_\psi)$. Thus, for a vector of
estimators $\hat\tau_1,\ldots,\hat\tau_K$---where each $\tau_k$ is an aggregation
of pairwise aTEs---this representation implies that for some vector-valued mean
zero function $\bm{\psi}(X_i, Y_i; c)$, we have the joint convergence: For $\odot$
the entrywise product, \[
\colvecb{3}{r_{N,1} (\hat\tau_1 - \tau_1)}{\vdots}{r_{N,K} (\hat\tau_K -
\tau_K)} = \colvecb{3}{r_{N,1}^{-1}}{\vdots}{r_{N,K}^{-1}} \odot \sum_{i=1}^N
\bm{\psi}(X_i, Y_i; c) + o_P(1) \dto \Norm(0, \Sigma_{\bm{\psi}}).
\]
The $(j,k)$\th{} entry of $\Sigma_{\bm{\psi}}$ is equal to $
\lim_{N\to\infty} \E\bk{
\frac{N}{r_{N,j}r_{N,k}}\psi_j(X_i, Y_i; c) \psi_k(X_i, Y_i; c)
}, $ which is consistently estimable by an analogue estimator $\frac{1}{r_
{N,j} r_ {N,k}}
\sum_{i=1}^N
\psi_j (X_i, Y_i; C_N) \psi_k(X_i, Y_i; C_N)$. For instance, we show in \cref
{lemma:variance_est,thm:variance_est} that the asymptotic variance of the RD estimators is
consistently estimable. In fact, since both \cref{eq:ipw_lottery} and \cref
{eq:llr_estimand} can be written as regression estimators on a subsample of
units with $J_i(C_N) = 1$, we implement the standard error estimators in \cref
{sec:monte_carlo} by simply reading off the Eicker--Huber--White standard
errors of the corresponding regression.
Motivated by the joint convergence, for $\hat\Sigma_{\bm\psi}$ the estimated
variance, we can approximate the sampling distribution of the estimators as a
joint Gaussian, centered at the corresponding estimands: \[
\colvecb{3}{\hat\tau_1}{\vdots}
{\hat\tau_K} \overset{a}{\sim} \Norm\pr{
\colvecb{3}{\tau_1}{\vdots}{\tau_K}, \Lambda^{-1}_N\hat \Sigma_{\bm\psi} \Lambda^
{-1}_N
}
\quad \Lambda = \diag(r_{N,1},\ldots, r_{N,K}).
\]
This joint convergence facilitates inference for some user-chosen aggregation
$\Gamma
\tau$. For
instance, the Bradley--Terry-style aggregation scheme (\cref{eq:BT_aggregation})
defines some matrix $\Gamma$ that maps pairwise school effects to the vector of
school value added estimates. Given $\Gamma$, we may base inference on the
following approximation to the sampling distribution of $\Gamma\hat\tau$: \[
\Gamma \hat\tau \overset{a}{\sim} \Norm(\Gamma\tau, \Gamma \Lambda^
{-1}_N\hat \Sigma_{\bm\psi} \Lambda^
{-1}_N \Gamma').
\]
Our Monte Carlo results in the next section illustrate these results.
\end{rmksq}
\section{Monte Carlo study}
\label{sec:monte_carlo}
We return to the example in \cref{ex:running_ex} and set up a Monte Carlo study. Suppose
the preference types $\br{A, B, C}$ are equally probable. School capacities are $\br{N,
0.25N, 0.25N, 0.25N}$, respectively for $s_0$ through $s_3$. Suppose the test score $R_i
\iid \Unif[0,1]$, independently of preferences. Suppose the lottery $U_i \iid \Unif
[0,1]$ independently as well. The potential outcomes are heterogeneous. Their conditional
means given preference type and test scores are described by \cref{fig:mean_pos}. Detailed
construction of the simulated data is documented in the code repository
(\url{github.com/jiafengkevinchen/school-choice-monte-carlo}).
\begin{figure}[tb]
\centering
\includegraphics[width=\textwidth]{assets/pos.pdf}
\caption{Mean potential outcomes given $X = (\succ, R)$. The black solid line is the
analytical large market cutoffs, and the dashed line is $C_N$ for a sample of 1000
students.}
\label{fig:mean_pos}
\end{figure}
\begin{figure}[tb]
\centering
\includegraphics[width=\textwidth]{assets/variation.pdf}
\caption{Aggregation region of the maximal treatment effect $\tau_{s_1 \succ s_0}$}
\label{fig:maximal_agg}
\end{figure}
In this setup, there are seven pairs of schools with nontrivial identified maximal
treatment effects. For each pair $s_1 \succ s_0$ and each preference type,
\cref{fig:maximal_agg} plots the region where $R_i \in \bar E_{s_1}(\succ, C_N) \cap \bar
E_{s_0}(\succ, C_N)$. The maximal treatment effect ($\tau_{s_1\succ s_0}$ in
\cref{eq:maximal}) is then an aggregation of aTEs with $R_i \in \bar E_{s_1}(\succ, C_N)
\cap \bar E_{s_0}(\succ, C_N)$. For instance, the effect $\tau_{3\succ 1}$ is an average
of effects between type $A$ and type $C$ students with test scores between the two
cutoffs.
\begin{table}[tb]
\caption{Estimates of $\tau_{s_1 \succ s_0}$}
\label{tab:estimates}
\centering
\begin{tabular}{rrrrrrrr}\toprule
$s_1$ & $s_0$ & Estimate & Analytic SE & Bootstrap SE & True value & Null t-statistic &
Coverage
\\\midrule
1 & 0 & $0.71$ & $0.34$ & $0.40$ & $0.74$ & $-0.08$ & 0.912\\
1 & 3 & $3.30$ & $0.23$ & $0.19$ & $3.34$ & $-0.18$ & 0.958\\
2 & 1 & $-1.79$ & $0.41$ & $0.38$ & $-2.01$ & $0.54$& 0.946 \\
2 & 3 & $0.32$ & $0.40$ & $0.38$ & $0.85$ & $-1.32$ & 0.952\\
3 & 0 & $-0.15$ & $0.10$ & $0.10$ & $-0.29$ & $1.37$& 0.944 \\
3 & 1 & $-0.30$ & $0.16$ & $0.16$ & $-0.33$ & $0.21$& 0.951 \\
3 & 2 & $0.96$ & $0.40$ & $0.41$ & $1.11$ & $-0.36$ & 0.939\\
\bottomrule
\end{tabular}
\begin{proof}[Notes]
All columns except for the last one are based on one draw with $N = 1000$ with 1000
bootstrap samples. We use bandwidth $h_N = 0.3 N^{-0.24} \approx 0.057$. The last
column computes the proportion that the null $t$-statistic is between $[-1.96, 1.96]$
over 1000 draws of the data.
\end{proof}
\end{table}
For one draw of the data with $N=1000$, \cref{tab:estimates} shows estimates and standard
errors of $\tau_{s_1 \succ s_0}$, using the estimators in \cref{sub:est}. Wald
inference based on the analytic standard error appears accurate, and the bootstrap
variance is close to the analytic estimated SEs. The estimates are approximately
independent: The sampling correlations between different estimates---themselves estimated
by the nonparametric bootstrap---are negligible, with all pairwise correlations less than
0.07.
\begin{figure}[tb]
\centering
(a) Bounds on RD variation
\includegraphics[width=\textwidth]{assets/rd_bounds.pdf}
(b) Regression estimates
\includegraphics[width=\textwidth]{assets/reg_estimate.pdf}
\caption{Pooled regression estimate and bounds on RD variation}
\label{fig:rd_bounds}
\begin{proof}[Notes]
We compute large market propensity scores analogously to \cref{sub:prop_scores}, where
we use the same bandwidth choice $h_N = 0.3 N^{-0.24}$ throughout. The treatment
effect is estimated by the coefficient of $\tilde D_i$ in a regression of $Y_i$ on
$\tilde D_i$ and $\hat\psi_i$, where $\tilde D_i$ indicates whether $i$ is assigned to
schools 2 and 3, and $\hat\psi_i$ is the sum of the large market propensity scores for
schools 2 and 3 for student $i$. This plot shows diagnostic values defined in
\cref{sub:diagnostic} for this regression. We also show the proportion of students
possibly and definitely subject to RD variation at this bandwidth sequence.
\end{proof}
\end{figure} At this sample size, the RD-driven estimates ($\tau_{1\succ *}$ and $\tau_
{2\succ *}$) are not significantly noisier than the lottery estimates. Thus in
aggregations they receive nontrivial weight at $N=1000$. To illustrate vanishing
weights, for a variety of sample sizes, \cref{fig:rd_bounds} in turn shows our
diagnostic in \cref{sub:diagnostic} for the regression estimator with large market
propensity scores, treating schools 2 and 3 as treatment schools. As expected, the
weight put on RD variation depends on and vanishes with the sample size.
To accentuate the concerns regarding aggregation, we choose the potential outcome
distributions so that $s_1$ is particularly good and $s_2$ is particularly bad for
students of preference type $B$. This makes students of type $B$ have large negative
RD-driven treatment effects for the treated schools $2$ and $3$. Correspondingly, in
\cref{fig:rd_bounds}(b), we observe that the estimated treatment effect becomes less
negative as sample size increases, since less of these RD-driven effects is aggregated.
\begin{table}[htb]
\caption{Aggregations of pairwise school effects}
\label{tab:agg}
\centering
\begin{subtable}[t]{0.48\linewidth}
\centering
\caption*{(a) Test score–driven value‑added}
\begin{tabular}{lrrrrrr}
\toprule
& $\tau_{1\succ0}$ & $\tau_{1\succ3}$ & $\tau_{2\succ1}$ & $\tau_{2\succ3}$ & Estimate & SE \\ \midrule
$\mu_1^R$ & 1 & 0 & 0 & 0 & 0.71 & 0.42 \\
$\mu_2^R$ & 1 & $-0.33$ & 0.67 & 0.33 & $-1.48$ & 0.51 \\
$\mu_3^R$ & 1 & $-0.67$ & 0.33 & $-0.33$ & $-2.19$ & 0.47 \\
\bottomrule
\end{tabular}
\end{subtable}
\hfill
\begin{subtable}[t]{0.48\linewidth}
\centering
\caption*{(b) Lottery‑driven value‑added}
\begin{tabular}{lrrrrr}
\toprule
& $\tau_{3\succ0}$ & $\tau_{3\succ1}$ & $\tau_{3\succ2}$ & Estimate & SE \\ \midrule
$\mu_1^L$ & 1 & $-1$ & 0 & 0.15 & 0.18 \\
$\mu_2^L$ & 1 & 0 & $-1$ & $-1.12$ & 0.41 \\
$\mu_3^L$ & 1 & 0 & 0 & $-0.15$ & 0.10 \\
\bottomrule
\end{tabular}
\end{subtable}
\vspace{1em}
\caption*{(c) Treatment‑ vs. control‑school aggregation}
\begin{tabular}{lrrrrrrrrr}
\toprule
& $\tau_{1\succ0}$ & $\tau_{1\succ3}$ & $\tau_{2\succ1}$ & $\tau_{2\succ3}$ & $\tau_{3\succ0}$
& $\tau_{3\succ1}$ & $\tau_{3\succ2}$ & Estimate & SE \\ \midrule
$\tau^R$ & 0.00 & $-0.50$ & 0.50 & 0.00 & 0.00 & 0.00 & 0.00 & $-2.55$ & 0.22 \\
$\tau^L$ & 0.00 & 0.00 & 0.00 & 0.00 & 0.50 & 0.50 & 0.00 & $-0.23$ & 0.09 \\
\bottomrule
\end{tabular}
\begin{proof}[Notes]
Panels (a) and (b) show the estimates in the least-squares specification \[\tau_{s_1
\succ s_0} = \one (
\text{$s_1$ is lottery}) (\mu_{s_1}^L - \mu_{s_0}^L) + \one(
\text{$s_1$ is test-score}) (\mu_{s_1}^R - \mu_{s_0}^R) + \epsilon_{s_1, s_0},\] which
is
\cref{eq:BT_aggregation} without the intercept. Because we did not include the
intercept, we normalize $\mu_0^L = \mu_0^R = 0$, making $\mu_s^L, \mu_s^R$ not
comparable to each other.
Panel (c) estimates instead \[
\tau_{s_1 \succ s_0} =J(s_1, s_0) S_
{s_1} \bk{
\one (
\text{$s_1$ is lottery}) \tau^L + \one(
\text{$s_1$ is test-score}) \tau^R} + \epsilon_{s_1, s_0}
\]
where $J(s_1, s_0) = 1$ if $s_1, s_0$ belongs to different treatment groups and zero
otherwise, and $S_{s_1} = 1$ if $s_1$ is a treatment school and $-1$ otherwise. Again,
the treatment group is defined as schools $2$ and $3$. In each panel, the left pane
shows how each parameter aggregates the pairwise effects determined by the
least-squares regression, which is represented by the matrix $\Gamma$ in
\cref{rmk:joint}: For instance, $\mu_2^R = \tau_ {1\succ 0} - \frac{1} {3}
\tau_{1\succ 3} + \frac{2}{3} \tau_{2\succ 1} + \frac{1}{3} \tau_{2\succ 3}$. The
standard errors in the aggregation are computed from $\Gamma \hat\Sigma \Gamma'$ where
$\hat\Sigma$ is the bootstrap estimate of the variance-covariance matrix of the
pairwise estimates $\hat\tau_{s_1\succ s_0}$.
\end{proof}
\end{table}
We illustrate alternative methods for aggregation in \cref{tab:agg}, which shows estimates
of $\mu_s^L, \mu_s^R$ in a specification like \cref{eq:BT_aggregation} in panels (a)--(b).
Each value-added should be interpreted as a comparison of school $s$ to school 0, where
\cref{tab:agg} shows how each is computed from the pairwise effects $\tau_{s_1 \succ
s_0}$. Generally speaking, the RD-driven effects $\mu_s^R$ are noisier than the
lottery-driven effects. Lottery-driven value-added estimates can be different from the
test-score driven ones. For instance, for school $3$, the pairwise effect $\tau_ {1 \succ
3}$ is large and positive, and
$\tau_{2\succ1}$ is negative, since they include variation in students of type $B$.
These effects are aggregated in the value-added $\mu_3^R$, contributing to its large
negative value relative to $\mu_3^L$.
In \cref{tab:agg}(c), we fit a similar regression to \cref{eq:BT_aggregation}, but we
further aggregate effects by treatment schools (2 and 3) and control schools (0 and 1).
This regression specification implicitly defines the RD-driven effect $\tau^R$ as $
\frac{1}{2} (\tau_ {2\succ 1} - \tau_ {1\succ 3})$ and $\tau^L$ as $\frac{1}{2}(\tau_
{3\succ 0}
+ \tau_ {3\succ1})$. We find that $\tau^R$ is much more negative than $\tau^L$, partly
driven by students of preference type $B$. The corresponding estimate from the local DA
propensity score regression is about $-0.53$, lying between these two estimates.
\section{Conclusion}
\label{sec:conclusion}
Detailed administrative data from school choice settings provide an exciting frontier
for causal inference in observational data. Remarkably, school choice markets are
engineered \citep{roth2002economist} to have desirable properties for market
participants, and yet they may yield natural-experiment variation that inform program
evaluation and policy objectives. Credibility of empirical studies using such variation
demands an understanding of the limits of the data---an understanding of what
counterfactual queries the data can and cannot answer (absent further assumptions). Our
analyses here provide a step towards that understanding.
As a review, we provide a detailed analysis of treatment effect identification in school
choice settings. We characterize the identification of aTEs, the building blocks of
aggregate treatment effects. We find that pooling over lottery- and RD-driven aTEs leads
the former to dominate the latter asymptotically. We provide a simple regression
diagnostic for the weight put on RD-driven aTEs, as well as some suggestions for
aggregating aTEs. Lastly, we contribute asymptotic theory for estimating aggregations of
RD-driven aTEs.
\vspace{3em}
\noindent \textbf{Declaration of generative AI and AI-assisted technologies in the writing
process}
During the preparation of this work the author(s) used ChatGPT and Google Gemini in order
to structure ideas, copy-edit, and produce code to implement procedures. After using this
tool/service, the author(s) reviewed and edited the content as needed and take(s) full
responsibility for the content of the publication.
\tiny
\singlespacing
\bibliographystyle{aer}
\bibliography{main.bib}
\onehalfspacing
\newpage
\footnotesize