EconBase
← Back to paper

Causal Effects in Matching Mechanisms with Strategically Reported Preferences

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

139,994 characters · 14 sections · 86 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Causal Effects in Matching Mechanisms with Strategically Reported Preferences

center[center omitted — 31 chars of source]

A growing number of authorities use mechanisms to allocate students to schools in a way that reflects student preferences and school priorities. However, most real-world mechanisms incentivize students to strategically misreport their preferences. Misreporting complicates the identification of causal parameters that depend on true preferences, which are necessary inputs for a broad class of counterfactual analyses. We provide an identification approach robust to misreporting and derive sharp bounds on causal effects of school assignment. Our approach applies to allocation rules characterized by placement scores and cutoffs. We use data from a deferred acceptance mechanism that assigns students to university programs in Chile. Matching theory predicts and empirical evidence shows that students behave strategically in Chile because they face constraints on preference submission and have good prior information about school accessibility. Our bounds are informative enough to reveal significant heterogeneity in graduation success with respect to preferences and school assignment.

Keywords: Matching, Strategic Reporting, Random Sets, Partial Identification, Constrained School Choice, Regression Discontinuity.

JEL Classification: C12, C21, C26.

Introduction

One of the most important decisions students make is their choice of field and institution of education. Identification of the impact of such choices on future outcomes is a critical step in the study of this decision process and of public education policies. The causal effects of schooling may vary widely across economic agents because of heterogeneous skills and preferences. Moreover, individuals' expectations about potential returns may prompt them to choose schools strategically. Heterogeneity and selection make identification of causal effects very challenging, especially when students face a large number of unordered options. On the positive side, a growing number of schools use centralized assignment mechanisms, which produce credible instruments based on the discontinuities generated by the assignment of comparable students to different schools (kirkeboen2016 and abdul_ecta2022).

Many centralized school assignment mechanisms effectively amount to a quasi\--ex\-per\-i\-men\-tal design where two groups of individuals who share similar scores are assigned to different schools based on how their scores relate to admission cutoffs. The assignment in a matching characterized by such cutoffs depends on the student's preferences over feasible schools. Unlike in a typical regression discontinuity (RD) design, in this setting, students on the same side of the cutoff do not necessarily receive the same assignment. For example, individuals with similar scores just above a certain cutoff may all prefer to go to the same “school $j$” but could have very different second-best options if they fall on the other side of the cutoff. In such a context, kirkeboen2016 construct comparable groups of individuals near a cutoff by conditioning on local preferences---that is, by selecting students whose preferences yield identical first- and second-best options if they fall, respectively, above and below that cutoff. Controlling for local preferences that equal a pair of schools, e.g., $(j,k)$, allows the RD to identify the causal effect of a change in the school assignment from $k$ to $j$, averaged over individuals who prefer $j$ over $k$.

Most real-world school assignment mechanisms create incentives for students to misreport their true preferences. agarwal2018 and fack2019 provide thorough discussions with several real-world examples of this. Subsequent empirical work have employed their methods to estimate true preferences or test for strategic behavior, for example, campos_impact_2024, andersson_beyond_2026, and many others. The empirical literature on school choice typically engages in counterfactual policy evaluations and an important class of such counterfactuals requires identification of causal parameters that are conditional on true preferences (Artemov2023, Section V.B). The identification of such causal parameters is the primary goal of our paper and we propose a novel RD strategy that controls for true local preferences. The challenge resides on the fact that true preferences are unobserved but may be partially identified under general forms of misreporting behavior.

This paper derives sharp bounds for causal effects of school assignment on future outcomes in mechanisms with cutoff characterization and strategic student behavior. We devise a two-step identification approach that is robust to strategic reporting of preferences. In the first step, the researcher partially identifies local preferences and constructs local preference sets for each student. We provide several tools for constructing these sets in the context of a student-proposing deferred acceptance (DA) mechanism where constraints on the submitted preferences lead to strategic behavior. Outside such a context, researchers may employ alternative tools to partially identify local preferences, as the second step of our procedure does not require a particular method to be used in the first step. For example, the identification method of agarwal2018 applies to a general class of mechanisms with cutoff representation that includes variants of the DA mechanism, Boston, First Preferences First, Chinese Parallel, etc. Finally, in the second step, the researcher employs an RD identification strategy that controls for local preference sets and partially identifies the causal effects of school assignment. Our bounds can be substantially tightened—and may even collapse to a single point—depending on the data and on how strong the first-stage assumptions are. The more restrictions we impose on strategic behavior, the more we learn about true local preferences, which then leads to narrower bounds.

Strategic behavior among students depends on the characteristics of the assignment mechanism. A mechanism is said to be strategy proof if submitting true preferences is a weakly dominant strategy for all students. For example, dubins1981 demonstrate that the DA mechanism is strategy proof. However, this result breaks down when the mechanism imposes constraints on the preferences that students can submit. In many real-world school assignment mechanisms, the number of schools is too large for students to feasibly rank all schools. The central authorities running these systems may either limit the number of schools that students may rank or impose costs on the basis of the number of schools submitted (see Table 1 Panel B by fack2019 for examples).

Our first-step tools for partial identification of local preferences naturally require assumptions on students' strategic behavior. We motivate our assumptions following the important contributions of haeringer2009. They study a game where students submit constrained preference rankings and a central mechanism allocates the students to schools. One of their important findings is that it is rational for students to submit partial orders of their true preferences in some mechanisms. Specifically, suppose that a mechanism is strategy proof when students are free to rank any number of schools, as under, e.g., the unconstrained DA or Top Trading Cycles (TTC) mechanisms. Then, if the preference rankings are constrained to having at most $K$ schools, a student can do no better than selecting $K$ schools among her acceptable schools and ranking them according to her true preferences.

The key assumption for our first-step tools is that students submit only partial orders of their preferences. This feature is in addition to the cutoff characterization of the matching, which we assume throughout the paper. Cutoff characterization means that a student is matched to her best feasible school, where “best” is defined according to her true preferences and a school is feasible if the student's placement score clears that school's admission cutoff. Our assumption about cutoff characterization is satisfied when the matching outcome is stable azevedo2016. Stability means that each student is matched to an acceptable school and all slots in preferred schools have been filled with people who have better placement scores. In constrained DA and TTC mechanisms, stability occurs in Nash equilibrium of the preference revelation game under appropriate conditions on the placement scores (Theorems 6.3 and 6.4, haeringer2009). Stability occurs in Nash equilibrium without restrictions on the scores in the constrained serial dictatorship (SD) mechanism, which is a particular case of DA. We characterize sharp local preference sets for every individual that are compatible with the observed data and these model assumptions. For an individual drawn at random, this becomes a random set that contains the true local preference with probability one. We show how to shrink these sets by proposing a menu of behavioral assumptions on strategic behavior that researchers may find more or less appropriate to their empirical contexts. The smaller the sets are in the first step, the narrower the bounds will be in the second step, even collapsing to point identification in some cases.

The second step of our approach relies on the local preference sets constructed in the first step with either our method or an alternative method. Given interest in a pair of schools $(j,k)$, we select all individuals whose local preference sets contain $(j,k)$ and whose placement scores are close to the cutoff for admission at school $j$. This subpopulation of individuals contains all individuals whose true local preferences equal $(j,k)$, but also other individuals. The average outcome in the subpopulation equals a weighted average of two averages: first, the average outcome for individuals with true local preferences $(j,k)$, which is interesting for the identification of causal effects, and second, the average outcome for individuals with true local preferences that differ from $(j,k)$. We do not know which individuals have preferences $(j,k)$, but we do characterize sharp bounds on the proportion of such individuals in the subpopulation using the random sets constructed in the first step. Thus, our setting fits the identification problem with corrupted data studied by horowitz1995. This method allows us to derive closed-form bounds on the first out of the two average outcomes above, which then leads to bounds on the average causal effects. This closed-form approach offers some intuition on when we can expect the bounds to be informative about or equal to the actual average causal effect (i.e., point identification). Although practical and intuitive, these closed-form bounds may not be sharp. Thus, building on molinari2020 and using random set theory, we characterize sharp bounds that are numerically computable when outcomes take finitely many values.

Methods combining RD identification with school matching data have been popular among applied and theoretical researchers in economics for at least 15 years kirabo_nber2009,popurquiola2011,bertanha_jmp2014. To the best of our knowledge, our paper is the first to prove RD identification of returns of school assignment in matching mechanisms with strategically reported preferences. We note that our two-step approach differs from the usual control function approach because our first step partially identifies the control variable instead of point-identifying it as in the usual approach. For this reason, we call it control mapping approach. Various other examples in econometrics feature this structure, see han_kaido_2025.

Our paper unifies and complements two branches of the literature. One branch features the methods proposed by agarwal2018 and fack2019 that take into account strategic reporting and identify students' true preferences, but they do not focus on causal effects of school matches on future outcomes. The other branch features the methods of kirkeboen2016 and abdul_ecta2022 that identify the causal effects of different assignments but control for reported instead of true preferences.\footnote{In the same branch, chen2026 provides a comprehensive set of causal parameters that are identified when assignment is based both on lottery- and RD-driven variation.} An RD strategy that controls for reported preferences identifies causal parameters that could be useful for a class of counterfactuals that differs from the one we consider in this paper. It is important to note, however, that in some empirical contexts, agents may have precise ex ante information about ex post matching cutoffs at the time when they choose preferences to report. In some cases, this can even lead to a discontinuous change of submission behavior at the ex post cutoffs. We present evidence of such behavior in the empirical context of Chile (Section (ref)). Such discontinuous behavior may confound identification in an RD that controls for reported preferences; we find that controlling for true preferences makes the RD strategy robust to this possibility.

Our work, and the kind of parameters we target, open avenues for studying identification in various policy settings while allowing for strategic reporting. Recent work in causal inference seeks to identify counterfactual policy effects under interference from market equilibrium or centralized assignment Arkhangel2025,munro_causal_2025. This line of work relies on truthful reporting but acknowledges its limitations and the importance of new developments that allow for strategic reporting.

We apply our two-step identification strategy to matching data from Chile. Chile has a centralized DA mechanism that assigns students to university--major pairs. To have an idea of the figures, in 2010, 88,000 students were constrained to rank at most eight university-major pairs ($K=8$) out of a total of 1,092 options available, in which case the theory predicts lack of truth telling. Prior work documents strategic behavior in Chile larroucaurios2020,larroucaurios2021 and we present two additional pieces of evidence. First, ex post admission cutoffs are highly predictable ex ante for many university-major pairs (Figure (ref)). Second, preference submission behavior changes when students' ex ante placement scores are close to ex post cutoffs, strongly suggesting students act on information on prior cutoffs (Figure (ref)). For some major pairs, the submission behavior even changes discontinuously at the ex post cutoffs (Figure (ref)). The methods proposed by kirkeboen2016 and abdul_ecta2022 are unable to provide the answers we seek with the Chilean data. First, strategic behavior leads their methods to identify a parameter that differs from an average treatment effect conditional on true preferences, which is our interest. Second, even if all agents report the truth, their methods are not directly applicable because of some specific aspects of the Chilean data.\footnote{ The methods of kirkeboen2016 apply to the SD mechanism, which is a particular case of the Chilean DA mechanism. In DA contexts, individual counterfactual sets must be defined in a more general way (see Definitions (ref) and (ref) below). While abdul_ecta2022 address this issue by creating a propensity-score control variable, their technique relies on placement scores with integer priorities plus continuous full-support lottery priorities. This particular structure does not correspond to the Chilean case, where program-specific placement scores are computed as functions of five primitive scores, and sometimes these functions are nonlinear. The Chilean setting requires again the counterfactual sets of schools to be carefully defined such that controlling for local preferences does not violate the continuity assumptions required by RD (see Assumption (ref) and Lemma (ref) below). }

Our analysis of the Chilean data proceeds in two parts. First, we show evidence that outcomes are heterogenous with respect to true local preferences. We compute bounds for average graduation rates if assigned to program $k$ among students with true local preference $(j,k)$. We fix the next-best alternative $k$ to be the Bachillerato de Ingreso Común at the University of Chile---a selective entry program with a large applicant pool that allows students to explore multiple disciplines before settling on a major, and is often viewed as an alternative route into competitive programs. We vary the first-best option $j$ and find non-overlapping bounds in several cases. The results reveal substantial heterogeneity across $j$, indicating that students’ preferences matter for graduation outcomes and are likely correlated with unobserved factors such as effort or ability not captured by test scores. The second part of our analysis focuses on average effects on graduation outcomes from changing the initial assignment from $k$ to $j$ among students with true local preference $(j,k)$. We choose $j$ to be Medicine at PUC Santiago because it is a highly-selective and popular program with a large applicant pool. We fix $j$ and plot bounds for common next-best alternatives $k$. In most cases, our bounds identify a positive effect of assignment into $j$ on graduating from $j$, with magnitudes that vary across alternatives $k$.

More generally, we compute bounds for three hundred college-major pairs $(j,k)$ under assumptions of varying strength on strategic behavior, which provides a valuable sensitivity exercise for researchers to evaluate the empirical content of their assumptions. We highlight three important takeaways from our bounding exercise. First, bounds are generally wide under the weakest of our assumptions, namely, that students simply select any subset with at most $K$ options from their true preference list while preserving the true relative ordering. We call this weak partial order (WPO). Wide WPO bounds essentially say that RD alone applied to matching data with strategic behavior does not reveal much unless researchers are willing to make stronger assumptions on strategic behavior. Second, bounds are generally narrow and even collapse to point identification in some cases under the strongest of our assumptions: strong partial order (SPO). SPO says that agents who rank less than $K$ programs reveal their true preferences. Third, point-estimates produced by assuming truth-telling for all agents fall outside our SPO bounds in many cases. This is relevant because the truth-telling assumption implies the SPO assumption, and we would expect point-estimates inside the bounds. Our finding suggests violation of the truth-telling assumption despite the fact that 80% of students in our sample rank less than $K$ programs. This runs counter to the common intuition that students will report truthfully whenever the $K$ constraint is non-binding for most of them.

The rest of this paper proceeds as follows. Section (ref) lays out the matching model for a continuum population of students and a finite number of schools. Section (ref) examines point identification of average treatment effects when students are truth-tellers. Section (ref) examines partial identification when students strategically report their preferences, with two subsections: Section (ref) provides tools for construction of local preference sets that apply to constrained DA mechanisms, and Section (ref) discusses how to use local preference sets constructed in this or other ways to derive bounds on the average treatment effects. We illustrate our identification approach with the Chilean data in Section (ref). The appendix presents all proofs for the paper plus additional results.

Model

We consider a continuum population of students and a set of $J$ schools, $\mathcal{J} := \{1,\ldots, J\}$, that have capacities $\{ q_1, \ldots, q_J \}$ defined in terms of shares of the student population azevedo2016. Denote by $\Omega$ the set of all students in the universe of interest and use $\omega$ to index an individual student type. The student type consists of three objects. First, $Q(\omega)$ denotes the true (strict) preference relation of student $\omega$ over the set of options $\mathcal{J}^0 := \mathcal{J} \cup\{0\}$, which includes schools $\mathcal{J}$ and an outside option $0$. For example, if $J=2$ and $Q(\omega) = \{1,2,0\}$, then 1 is preferred to 2 (i.e., $1 Q(\omega) 2 $), 1 is preferred to 0 (i.e., $1 Q(\omega) 0 $), and 2 is preferred to 0 (i.e., $2 Q(\omega) 0$). Let $\mathcal{Q}$ be the set of all strict preference relations over $\mathcal{J}^0$ that admit at least one school that is acceptable. A school $j \in \mathcal{J}$ is “acceptable” for student $\omega$ if it is preferred to that student's outside option, i.e., $jQ(\omega)0$. We define $\bar{Q}$ as the weak preference relation induced by $Q$, i.e., $j \bar{Q} k \Leftrightarrow j Q k $ or $j=k$. The second object of the student type is a vector of scores $\boldsymbol{\mathrm{R}}(\omega):=(R_1(\omega), \ldots, R_J(\omega)) \in \mathcal{R} \subseteq \mathbb{R}^J$, where each school $j$ utilizes $R_j$ to rank students for admission. The third and last object, $Y(\omega,d)$, is the potential outcome of student $\omega$ if the student is assigned to option $d \in \mathcal{J}^0$. Each student has a potential outcome function $Y(\omega,\cdot)$ that maps from $\mathcal{J}^0$ to $\mathcal{Y} \subseteq \mathbb{R}$. We call $\boldsymbol{\mathrm{\Gamma}}$ the set of all possible potential outcome functions. The set of all student types is $\Omega := \mathcal{Q} \times \mathcal{R} \times \boldsymbol{\mathrm{\Gamma}}$. In a continuum economy, there is a probability measure $\mathbb P$ over $\Omega$ and the Borel $\sigma$-algebra of the product space $\Omega$. We suppress the argument $\omega$ whenever it is unnecessary for ease of notation, e.g., $Y(d)$ vs. $Y(\omega,d)$ and $Q$ vs. $Q(\omega)$.

A “matching” is described by a measurable function $\mu:\Omega \to \mathcal{J}^0$ that satisfies two conditions: for every $j\in\mathcal{J}$, (i) the mass of students matched to $j$ is less than or equal to the capacity of school $j$, i.e., $\mathbb P \{\omega: \mu(\omega) = j \} \leq q_j$; and (ii) the set of students who weakly prefer option $j \in \mathcal{J}^0$ over their matching, i.e., $\{\omega: j \bar Q(\omega) \mu(\omega) \}$, is an open set.\footnote{ azevedo2016 impose the same condition to rule out multiplicity of stable matchings that differ in a set of types with measure zero.} For every student type $\omega$, $\mu(\omega)$ is either the school $j$ to which the student is matched or zero. When $\mu(\omega)=0$, the student is unmatched and takes an outside option. An important definition for this paper is that of stability.

definition[Stability] The matching $\mu: \Omega \to \mathcal{J}^0$ is a stable matching if three conditions are satisfied for every $\omega \in \Omega$: (i) $\mu(\omega) \bar Q(\omega) 0$ (individual rationality); (ii) for any $j \in \mathcal{J}$, if $j Q(\omega) \mu(\omega)$, then $j$ is full (no waste); and (iii) for any $j \in \mathcal{J}$ that is full, if $\mu(\omega')=j$ and $ j Q(\omega) \mu(\omega)$, then $R_j(\omega') > R_j(\omega )$ (no justified envy).

A mechanism $\varphi$ matches students to schools by mapping the students’ scores and submitted preference lists to schools. Student $\omega$ submits a preference list $P(\omega) \subseteq \mathcal{J}$, which is an ordered list of her acceptable schools. For example, for $J=3$, if $Q(\omega)=\{1,2,0,3\}$, then $P(\omega)=\{1,2\}$, as long as the student submits her true list of acceptable schools. The number of schools in $P$, denoted $|P|$, is at least one because everyone participating in the match has at least one acceptable school. As with $\bar Q$, we also define $\bar P$ as the weak preference relation induced by $P$. A mechanism takes as inputs everyone's submitted preferences (i.e., $P$ which is a mapping from $\Omega$ to an ordered subset of $\mathcal{J}$) and everyone's scores (i.e., a vector-valued function $\boldsymbol{\mathrm{R}}:\Omega \to \mathcal{R}$) and gives rise to a matching function. Formally, $\varphi(P,\boldsymbol{\mathrm{R}}):\Omega \to \mathcal{J}^0$. We say a mechanism $\varphi$ is strategy proof if, for every student, submitting the true ranking of acceptable schools is a weakly dominant strategy---in other words, any student $\omega$'s misreporting of $P$ never leads to a better option and sometimes leads to a worse option, depending on what other students submit. We say a student is a truth-teller if her $P$ equals her true ranking of acceptable schools. Otherwise, we say she is strategic or not a truth-teller.

The ability to characterize a matching allocation on the basis of cutoffs is fundamental for this paper.

definition[Cutoff Characterization] For placement scores $\boldsymbol{\mathrm{S}} : \Omega \to \mathcal{S} \subseteq \mathbb{R}^J$, $\boldsymbol{\mathrm{S}}(\omega):=(S_1(\omega), \ldots, S_J(\omega))$ and admission cutoffs $\boldsymbol{\mathrm{c}} \in \mathcal{S}$, $\boldsymbol{\mathrm{c}}:=(c_1,\ldots, c_J)$, the set of feasible options of a student $\omega$ equals all schools for which her placement scores clear the admission cutoffs plus the outside option: $\{0\} \cup \{j \in \mathcal{J} : S_j(\omega) \geq c_j \}$; student $\omega$'s best feasible option is the option that ranks first according to $Q(\omega)$ among her feasible options. We say the matching $\mu: \Omega \to \mathcal{J}^0$ has cutoff characterization if there exist placement scores $\boldsymbol{\mathrm{S}} : \Omega \to \mathcal{S}$ and admission cutoffs $\boldsymbol{\mathrm{c}} \in \mathcal S$ such that, for every $\omega \in \Omega$, the matching $\mu(\omega)$ equals student $\omega$'s best feasible option according to $Q(\omega)$.

This paper considers mechanisms that produce matching functions with a cutoff characterization according to Definition (ref). Placement scores $\boldsymbol{\mathrm{S}}$ may or may not equal school priority scores $\boldsymbol{\mathrm{R}}$. The definition gives researchers the freedom to construct special placement scores $\boldsymbol{\mathrm{S}}$ if the mechanism that they consider does not admit cutoff characterization by means of priority scores $\boldsymbol{\mathrm{R}}$. The idea behind this definition stems from the logic of the general class of mechanisms of agarwal2018. Moreover, we assume throughout the paper that, for any $j\in \mathcal{J}$, $S_j$ has a distribution that is absolutely continuous with respect to (wrt) the Lebesgue measure and support $\mathcal{S}_j$ that contains a closed interval around $c_j$. The support of $\boldsymbol{\mathrm{S}}$ is $\mathcal{S}$. For any two scores $S_j$ and $S_l$, we assume that either $\mathbb{P}[S_j = S_l]=1$ or $\mathbb{P}[f(S_j) = S_l]<1$ for any measurable function $f$. This says that the only deterministic function relating any two scores may be the identity function.

azevedo2016 demonstrate that, if the matching function is stable, the matching has cutoff characterization with $\boldsymbol{\mathrm{S}} = \boldsymbol{\mathrm{R}}$ and admission cutoffs constructed as follows. For each $j \in \mathcal{J}$, $c_j := \inf \{ S_j(\omega) : \text{ for } \omega \text{ with } \mu(\omega)=j \}$ if some individuals are matched to $j$ or $c_j:=\inf \mathcal{S}_j$ if nobody is matched to $j$. In fact, most empirical work applying RD to matching data in the last 15 years rely on this stability-based method for constructing cutoffs. Many mechanisms produce stable matchings. For example, SD and DA are strategy proof and lead to stable matchings if agents are truth-tellers.

Regarding settings where agents are not truth-tellers, e.g., because they face constraints in the submission of $P$, haeringer2009 and fack2019 show how stability arises in equilibrium with strategic agents. In particular, fack2019 formally test for stability in their empirical context of Parisian schools and do not find evidence against it. One possible threat to stability is agents that make mistakes during submission of their preferences. artemovCheHe2017,Artemov2023 explicitly study stable matching with mistaken agents. They argue that most of the mistakes observed in equilibrium are likely to be payoff-irrelevant in large markets, and thus not hurting stability of the match. Despite the strong support for stability in the literature, we emphasize that Definition (ref) allows for the general cutoff characterization of agarwal2018 that does not require stability.

Parameter of Interest

Having laid out the physical environment of our matching economy, we now motivate the type of causal parameter that concerns this paper. Causal parameters are often used as policy tools for their internal validity—evaluating the impact of historical interventions on outcomes—or for their external validity—assessing the impact of new interventions based on knowledge of historical interventions, that is, constructing counterfactual analyses; see heckman2005scientific for a detailed discussion.

In this paper, we aim to provide strategies for identifying a set of parameters that are not only internally valid but also possess substantial external validity, thereby enabling the evaluation of counterfactual policies. We seek to identify conditional moments of treatment effects $Y(d')-Y(d)$ when the econometrician has access to an infinite amount of data and observes the joint distribution of the following random objects: $P(\omega)$, $\boldsymbol{\mathrm{S}}(\omega)$, $\mu(\omega)$, and $Y(\omega) := Y(\omega, \mu(\omega))$.\footnote{ We abuse the notation and employ the letter $Y$ for both the observed outcome, $Y(\omega)$, and the potential outcome of being assigned to school $d$, $Y(\omega,d)$.}

Section (ref) in the appendix provides a detailed discussion of a class of counterfactual analyses that is advocated by Artemov2023. Analyzing the welfare effects of a policy change in that class relies on identification of the average structural functions (ASF)

equation[equation omitted — 80 chars of source]

for $j \in \{0,\ldots,J\}$ and $q \in \mathcal{Q}$. Nonparametric identification of these functions is extremely challenging. However, this paper shows that combining an RD identification strategy with data from assignment mechanisms that satisfy our assumptions yields set identification of differences of the ASFs in (ref); in other words, our proposed methods set-identify

equation[equation omitted — 83 chars of source]

at finitely many values of $j$, $k$, and $q$ (see Propositions (ref) and (ref)).

The parameters in (ref) are the object of interest of this paper. They carry the simple intuition of the massively popular RD designs while providing useful identifying information for a large class of counterfactual policy evaluations. Section (ref) in the appendix describes how researchers may combine all information identified by RD with smoothness assumptions on the ASFs to construct bounds on a variety of counterfactual policy effects.

Identification with Truthful Reports

In this section, we consider identification of causal effects when all students are truth-tellers, that is, when they submit their true ranking of acceptable schools. We start with truth-telling in order to introduce our approach and notation in a simple behavioral setting before moving to strategic reports in Section (ref).

assumption[Truth-telling] Students submit their true list of acceptable schools.

The identification strategy of this paper resembles a sharp RD design. Our goal is to identify the effects of the school of assignment on future outcomes. Another interesting question is the effect of the school of graduation on future outcomes---to answer it we would require a strategy resembling a fuzzy RD because some students do not graduate from the same school they are assigned to. We defer this identification problem to future work as several issues beyond the scope of this paper (e.g., multiple compliance types with unordered treatments) arise in that case.

Unlike in a standard sharp RD, that $S_j(\omega)$ clears the cutoff $c_j$ does not automatically determine that student $\omega$ is allocated to school $j$. This is the case only when $j$ is the most preferred school among the schools that are feasible to the student, that is, when $j$ is the favorite school in the set of schools for which the student clears the cutoff.

The first step in the RD is to correctly identify the marginal individuals for a given cutoff and a given change in schools. For example, for any individual with score $S_j$ just to the right of $c_j$, we need to determine two things: that the individual is matched to school $j$, and that the individual would have been matched to school $k$ had her score been just to the left of $c_j$. The cutoff representation implies that these two things depend on the counterfactual sets of available schools on either side of the cutoff and on the individual's preferences over these sets.

It is straightforward to obtain counterfactual sets of available schools in the case where all schools rely on the same placement score, that is, $S_j=S_1$ for every $j$. For example, this is the case under the SD mechanism. In this case, the set of feasible schools is all schools with a cutoff below or equal to score $S_1$. Note that everyone just above (or just below) cutoff $c_j$ has exactly the same set of feasible schools. For someone with $S_1 \geq c_j$, the counterfactual scenario has the score crossing to the left of cutoff $c_j$, and school $j$ is dropped from the set of feasible schools. In turn, for someone with $S_1 < c_j$, the counterfactual scenario adds school $j$ to the set of feasible schools. Unlike under SD, agents near and on the same side of a cutoff in DA differ in their sets of feasible schools. It is not immediately obvious which schools appear in their counterfactual sets. In DA, schools use different scores, and these scores may be functions (e.g., weighted averages) of a small set of primitive scores. That is the case in our application with the Chilean data. This makes the joint support of the distribution of scores highly dependent and complicates the counterfactual analysis. Dealing with this complexity is empirically relevant since many real-world higher education assignment mechanisms use DA. See Table 1 Panel B in fack2019 for a list of examples.\footnote{ A common practice in applied work consists of “cleaning” irrelevant schools from the submitted preference lists in cases where $S_j=S_1$ for every $j$. For instance, say an individual submits $P=\{ 1,2,3 \}$ and $c_2 > S_1 > c_1 > c_3$. Given the cutoff characterization and truth-telling, the matching assignment of this individual is school $1$; the counterfactual assignment when $c_2 > c_1 > S_1 > c_3$ is school $3$ even though $2P3$; this is the case because school $2$ has a cutoff higher than the cutoff of school $1$. In this case, the irrelevant school to be cleaned from $P$ is school $2$. The general idea is to remove all schools ranked below $1$ that have cutoffs higher than $c_1$. See the description of this practice by estrada2017. The practice cannot be used to identify counterfactual assignments in cases where different schools use different placement scores, as under, e.g., the DA mechanism. }

Our framework allows for a variety of joint distributions on the vector of scores $\boldsymbol{\mathrm{S}}$ and works with the definition of counterfactual budget sets below.

definition[Counterfactual Budget Sets] Consider a student with a vector of scores $\boldsymbol{\mathrm{S}}$. The budget set for this student is her set of feasible options, \begin{align} B(\boldsymbol{\mathrm{S}}) := & \{ 0 \} \cup \{m \in \mathcal{J} : S_m \geq c_m \}. \notag \end{align} Fix a school $j \in \mathcal{J}$ with cutoff $c_j$. The right-counterfactual budget set for this student at cutoff $c_j$ is $B^+_j(\boldsymbol{\mathrm{S}}) := B(\boldsymbol{\mathrm{S}}) \cup \{m : S_m=S_j \text{ and } c_m=c_j\}$; the left-counterfactual is $B^-_j(\boldsymbol{\mathrm{S}}) := B(\boldsymbol{\mathrm{S}})\setminus \{m : S_m=S_j \text{ and } c_m=c_j \}$, where $C\setminus D$ equals the set $C$ minus the elements of set $D$.

To fix ideas, we consider a simple example throughout the paper in the context of the SD mechanism.

example*[SD Example, Part I] In SD, $S_j=S_1$ for every $j$, and the definition of the budget set above equals $B(\boldsymbol{\mathrm{S}}) = \{ 0 \} \cup \{m : S_1 \geq c_m \}$. Suppose we have four schools with cutoffs $c_1 < c_2 < c_3 < c_4$. Individuals in this economy have five possible budget sets: $\{0\}$, $\{0, 1\}$, $\{0,1,2\}$, $\{0,1,2,3\}$, and $\{0,1,2,3,4\}$. For individuals near cutoff $c_4$, the counterfactual budget sets are $B^-_4(\boldsymbol{\mathrm{S}}) = \{0,1,2,3\}$ and $B^+_4(\boldsymbol{\mathrm{S}}) = \{0,1,2,3,4\}$.\footnote{ In SD or in DA with independent placement scores, the definitions of the counterfactual equal $B^+_j(\boldsymbol{\mathrm{S}}) = B(\boldsymbol{\mathrm{S}}) \cup \{m : c_m=c_j\}$ and $B^-_j(\boldsymbol{\mathrm{S}}) = B(\boldsymbol{\mathrm{S}}) \setminus \{m : c_m=c_j \}$; moreover, if the cutoffs are unique, $B^+_j(\boldsymbol{\mathrm{S}}) = B(\boldsymbol{\mathrm{S}}) \cup \{j\}$ and $B^-_j(\boldsymbol{\mathrm{S}}) = B(\boldsymbol{\mathrm{S}}) \setminus \{j \}$. }

Next, we follow the intuition of kirkeboen2016 and define the concept of local preferences, that is, the first- and second-best choices for a marginal individual at any given cutoff. This will later become the control variable in our RD identification strategy with truthful agents.

definition[Local Preferences] Fix a school $j \in \mathcal{J}$ with cutoff $c_j$. Consider a student $\omega$ with preference $Q(\omega)$ and scores $\boldsymbol{\mathrm{S}}(\omega) $. For any pair of options $(k,l)\in \mathcal{J}^0 \times \mathcal{J}^0$, we say that $(k,l)$ is the local preference of student $\omega$ at cutoff $c_j$ if the favorite feasible option of student $\omega$ shifts from $l$ to $k$ as we exogenously increase $S_j(\omega)$ from being smaller than $c_j$ to being larger than $c_j$. We define the true local preference of this student as the pair $Q_j(\omega) := (k,l)$. Formally, for a set of options $B \subseteq \mathcal{J}^0$, define the best option in $B$ according to $Q$ as $Q(B)$. We have that $Q(B) = m \Leftrightarrow m \in B \text{ and } m \bar Q(\omega) n ~ \forall n \in B.$ Finally, $Q_j(\omega) = (k,l)$ if, and only if, $Q(B^+_j(\boldsymbol{\mathrm{S}}))=k$ and $Q(B^-_j(\boldsymbol{\mathrm{S}}))=l$. The reported local preference $P_j(\omega)$ is defined in a similar fashion. If $B \cap P \neq \emptyset$, $P(B) = m \Leftrightarrow m \in B \text{ and } m \bar P(\omega) n ~ \forall n \in B;$ otherwise, if $B \cap P = \emptyset$, $P(B)=0$. We have that $P_j(\omega) = (k,l)$ if, and only if, $P(B^+_j(\boldsymbol{\mathrm{S}}))=k$ and $P(B^-_j(\boldsymbol{\mathrm{S}}))=l$.
example*[SD Example, Part II] Consider four individuals whose submitted preferences are $P^{(1)}=\{3,4,2,1\}$, $P^{(2)}=\{4,1,2,3\}$, $P^{(3)}=\{4,2,3,1\}$, and $P^{(4)}=\{4,3,1,2\}$. Their corresponding local preferences at cutoff $c_4$ are $P^{(1)}_4=(3,3)$, $P^{(2)}_4=(4,1)$, $P^{(3)}_4=(4,2)$, and $P^{(4)}_4=(4,3)$.

There is no distinction between $P_j$ and $Q_j$ at this stage because of Assumption (ref). Students are truthful when they submit their list of acceptable schools, so $P$ equals the schools listed higher than $0$ in $Q$ and $P_j=Q_j$. Section (ref) considers the case of strategic misreporting. In that case, $P_j$ is observed but $Q_j$ is not. Cutoff characterization implies that an individual with $Q_j=(k,l)$ is matched to school $k$ if $S_j$ is just above $c_j$ or to school $l$ if $S_j$ is just below $c_j$. The same applies for individuals with $P_j=(k,l)$ under Assumption (ref).

The local preference pair $(j,k)$ at a certain cutoff $c_j$ is useful for identification only if there exists a positive fraction of individuals in the data near cutoff $c_j$ with those local preferences. We collect such useful pairs in the set $\mathcal{P}$.

definition[Comparable Pairs] We say $(j,k) \in \mathcal{J} \times \mathcal{J}$, $j \neq k$, is a comparable pair of alternatives if (i) $c_j$ is an interior point of the support $\mathcal{S}_j$ and (ii) $\mathbb{P} [ Q_j =(j,k) | S_j =s]$ is bounded away from zero for $s$ in an open neighborhood of $c_j$. Finally, we define $\mathcal P \subseteq \mathcal{J} \times \mathcal{J}$ as the set of all comparable pairs.\footnote{ We adopt the convention that comparable pairs do not involve the outside option $0$. We do this because it may be hard to interpret the treatment effects of a change from the outside option to a school when the outside option varies across individuals. We do not consider pairs with $j=k$ because the initial school assignment does not change for these individuals. We also exclude pairs $Q_j = (j',k)$ with $j'\neq j$ from $\mathcal{P}$ to avoid redundancy. We may find individuals with $Q_j = (j',k)$ whenever schools $j$ and $j'$ use the same score and have the same cutoff. As the score $S_j = S_{j'} $ crosses the cutoff $c_j = c_{j'}$, access is granted to both schools $j$ and $j'$, and individuals may differ in their preferences for these schools. Individuals who prefer $j'$ will not appear in $\mathcal{P}$ as having $Q_j = (j',k)$, but they may appear in $\mathcal{P}$ with $Q_{j'} = (j',k)$. }

The purpose of defining counterfactual sets and local preferences is to construct a variable for every student and use it as a control variable in the RD. In this section, this variable is $P_j$, which equals to $Q_j$ because of Assumption (ref). When we focus on students with $P_j =(j,k)$, $k\neq j$, the marginal switch in the allocation around cutoff $c_j$ becomes a function of $S_j$. Controlling for $P_j$ ensures that we apply the RD strategy to all individuals whose assignment switches from $k$ to $j$ at the cutoff; this identifies the effect of the change in the assignment as long as the typical RD continuity assumptions are satisfied.

RD identification requires continuity assumptions on the distribution of individual types conditional on the relevant placement score. In our case, we also need to verify continuity after we condition on $P_j=Q_j$. After all, we do not want to condition on a variable that breaks the central argument for identification in RD: that individuals to the right and the left of the cutoff are “similar on average”. Below, we state an assumption on the continuity of types and prove that it implies the kind of smoothness required by RD. Before we do so, we define the following set of events. For every school $j\in \mathcal{J}$, partition the placement scores $\boldsymbol{\mathrm{S}}$ as the score of school $j$ and all other scores: $\boldsymbol{\mathrm{S}} \equiv (S_j,\boldsymbol{\mathrm{S}}_{-j})$. Define $\boldsymbol{\mathrm{A}}_{-j}$ to be the collection of events on $\boldsymbol{\mathrm{S}}_{-j}$ that determine the availability of all non-$j$ schools. There are $2^{J-1}$ such events in $\boldsymbol{\mathrm{A}}_{-j}$. For example, if $J=2$, $\boldsymbol{\mathrm{A}}_{-1}=\{ \{ S_2\geq c_2\}, \{ S_2 < c_2\}\}$; if $J=3$, $\boldsymbol{\mathrm{A}}_{-1}=\{ \{ S_2 \geq c_2, S_3 \geq c_3 \}, \{ S_2 \geq c_2, S_3 < c_3 \}, \{ S_2 < c_2, S_3 \geq c_3 \}, \{ S_2 < c_2, S_3 < c_3 \} \}$; etc.

assumption(Continuity of Types) Consider a school $j$ with cutoff $c_j$ in the interior of the support $\mathcal{S}_j$. Assume the following functions of $s$ are all continuous at $s=c_j$: (i) $\mathbb{P}[\boldsymbol{\mathrm{S}}_{-j} \in A_0, Q =Q_0| S_j = s ]$ for any $A_0 \in \boldsymbol{\mathrm{A}}_{-j}$ and $Q_0 \in \mathcal{Q}$ and (ii) $\mathbb{E}[ ~ g (Y(d)) ~ \mathbb{I}\{ \boldsymbol{\mathrm{S}}_{-j} \in A_0, Q =Q_0 \} ~ | ~ S_j = s ]$ for any $A_0 \in \boldsymbol{\mathrm{A}}_{-j}$, $Q_0 \in \mathcal{Q}$, and $g\in \mathcal{G}$, where $\mathcal{G}$ is a set of measurable functions $g:\mathbb{R} \to \mathbb{R}$ that includes the constant function $g(y)=1$ and the identity function $g(y)=y$.
lemmaSuppose Assumption (ref) holds. Consider a school $j$ with cutoff $c_j$ in the interior of the support $\mathcal{S}_j$, and choose two schools $k,l \in \mathcal{J}^0$ such that $\mathbb{P}[Q_j=(k,l)|S_j=c_j]>0$. Then, for any function $g \in \mathcal{G}$ and any $d \in \mathcal{J}^0$, we have that $ \mathbb{E}[ g (Y(d)) | Q_j=(k,l), S_j = s ]$ and $\mathbb{P}[ Q_j=(k,l) | S_j = s ]$ are continuous functions of $s$ at $s=c_j$.

The proof of this lemma and all other proofs appear in the appendix. Finally, Assumptions (ref) and (ref) give sufficient conditions for identification for comparable pairs of school changes.

propositionSuppose Assumptions (ref)--(ref) hold. For any pair $(j,k) \in \mathcal{P}$, \begin{align*} &\mathbb{E}[g(Y(j)) - g(Y(k)) | Q_j=(j,k), S_j=c_j ] \\ & = \mathbb{E}[g(Y) | P_j =(j,k), S_j=c_j^+] - \mathbb{E}[g(Y) | P_j = (j,k), S_j=c_j^-], \end{align*} where the condition $S_j=c_j^+$ denotes the limit as $S_j \downarrow c_j$ and the condition $S_j=c_j^-$ denotes the limit as $S_j \uparrow c_j$.

Proposition (ref) shows that a standard RD is valid in the truth-telling case as long as we control for $P_j$. The parameter of interest is the average treatment effect on $g(Y)$ from a change in the school of assignment from $j$ to $k$, averaged over individuals at the cutoff $c_j$ and with true local preferences $(j,k)$ ---fomally:

equation[equation omitted — 156 chars of source]

Note that our parameter of interest (ref) does not condition on the full preference profile $Q = q$ as in (ref) but on local preferences $Q_j=(j,k)$. We do so having in mind the data constraints inherent to the local nature of RD estimation. It follows that one value of $Q_j$ maps to multiple values of $Q$ for individuals with $S_j=c_j$ and (ref) equals weighted averages of (ref) over values of $Q$ (see proof of Lemma (ref) in Section (ref) of the appendix).

Identification with Strategic Reports

This section studies identification of causal effects when students are strategic in reporting their rankings of acceptable schools. Strategic reports make $P_j$ generally different from $Q_j$, and $Q_j$ is not observed. In contrast to Proposition (ref), controlling for $P_j$ likely identifies a weighted average of various average treatment effect parameters; however, the weights are unknown because they depend on the unknown mapping $\omega \mapsto Q_j(\omega)$. In addition, there is the possibility of a more extreme form of strategic behavior, that is, agents may change their preference submission discontinuously at the cutoff. We provide empirical evidence of such behavior in Section (ref). In this case, controlling for $P_j$ may break internal validity of the RD, and the strategy of controlling for $P_j$ no longer identifies a weighted average of average treatment effect parameters. Although our methods are also robust to this extreme possibility, we emphasize that it is not the only situation where researchers may find our methods useful. Our methods seek to identify (ref) but controlling for $P_j$ does not identify (ref) under general forms of strategic behavior, regardless if it is discontinuous or not.

We propose a two-step identification approach. In the first step, the researcher characterizes the set of true local preferences $Q_j$ that is compatible with the data and appropriate behavioral assumptions. In the second step, the researcher controls for the constructed local preference sets and partially identifies the parameters in (ref). We discuss the first and second steps in Sections (ref) and (ref), respectively. Note that our two-step approach differs from the usual two-step control function approach in econometrics. The usual approach is to point-identify the control variable in the first step, while our approach involves partially identifying the control variable. Thus, we refer to our two-step procedure as a control mapping approach.

Section (ref) presents several tools for the identification of local preference sets. These tools rely on assumptions known to be appropriate in SD and DA contexts, although we do not rule out their applicability in contexts with other mechanisms; e.g., the TTC mechanism satisfies one of our assumptions, such that some of the tools from Section (ref) are still useful. More generally, researchers may utilize preference identification tools that work under alternative assumptions, for example, the methods of agarwal2018 and fack2019. Either way, the researcher must construct a set of local preferences for each individual in the first step.

Section (ref) describes the second step of our procedure. This step features high-level assumptions imposed on the local preference sets such that the researcher is not restricted to the methods of Section (ref). In particular, depending on the data and the strength of assumptions in the first step, each individual local preference set may collapse to a unit set. This leads to point identification in our second step. For example, this will be the case if researchers utilize the parametric identification tools of agarwal2018 or fack2019 in the first step.

Partial Identification of Local Preferences

This section provides tools for set identification of local preferences using assumptions on agents' behavior and the mechanism. These assumptions are specific to this subsection, and we motivate them with reference to the context of the constrained DA mechanism studied by haeringer2009. haeringer2009 study a game where students submit constrained preference rankings and a mechanism matches students to schools as a function of $P$, $\boldsymbol{\mathrm{R}}$, and schools’ capacities. Although the unconstrained DA mechanism is strategy proof, many real-world implementations of DA restrict the number of schools that students can submit in their rankings. In this case, there is a cap $K<J$ such that $1 \leq |P| \leq K$, and the submitted ranking $P$ is generally different from the list of acceptable schools in $Q$. When implemented in this way, the DA mechanism is not strategy-proof, and there are no dominant strategies. Strategyproofness also breaks down if, instead of facing a cap, students incur an application cost as a function of the number of schools submitted fack2019.

Lemma 4.2 by haeringer2008 shows that, if a mechanism is strategy proof when $K=J$, then, in the game with $K<J$, any constrained ranking of schools is weakly dominated by the same set of schools ranked according to true preferences. This result implies that a student cannot lose and may possibly gain by taking any arbitrary list with less than or equal to $K$ schools, dropping the unacceptable schools, and ranking the acceptable schools according to her true preferences. A further implication is that if a student's number of acceptable schools is less than or equal to $K$, then her dominant strategy is to submit her true list of acceptable schools (Proposition 4.2 of haeringer2009). These implications give rise to a class of undominated strategies according to the following definitions of partial order.

definition(Weak and Strong Partial Order) We say $P$ is a weak partial order of $Q$ if $P$ is any selection of up to $K$ schools among the acceptable schools in $Q$ and that selection of schools is ranked according to $Q$. Formally, (i) $1 \leq |P| \leq K$, $P \subseteq \{d\in Q: d Q 0 \} $; and (ii) for every $d,d' \in P$, $d' P d \Leftrightarrow d' Q d$. We say $P$ is a strong partial order of $Q$ when a third condition holds in addition to (i) and (ii). Namely, (iii) $|P| = \min\{K, | \{d\in Q: d Q 0 \} | \}$. In other terms, if the number of acceptable schools in $Q$ is less than or equal to $K$ and $P$ is a strong partial order of $Q$, then $P$ equals the list of acceptable schools in $Q$; otherwise, if the number of acceptable schools in $Q$ is greater than $K$, $P$ is a subset of $K$ schools among the acceptable schools in $Q$.

Lemma (ref) in Section (ref) of the appendix summarizes the implications of the result on partial orders from haeringer2008,haeringer2009 in terms of our Definition (ref). In short, for a student with true preferences $Q$, any $P$ is weakly dominated by a weak partial order $P^*$ of $Q$ that has the same acceptable schools as $P$; in turn, $P^*$ is weakly dominated by a strong partial order $P^{**}$ of $Q$ that contains the same set of acceptable schools as $P^*$. Every strong partial order is a weak partial order, but the converse is not true. Our definition of a weak partial order strategy is similar to the definition of the dropping strategy from kojima2009incentives.

Assuming that agents always submit a strong partial order implies they reveal their true ordered list of acceptable schools whenever they submit $P$ with fewer schools than the cap $K$. This could be a strong behavioral assumption in some contexts where agents have more than $K$ acceptable schools but have a strong expectation that they will gain admission to a smaller-than-$K$ set of schools. In this case, they may submit $|P|< K$ not because it reflects their full list of acceptable schools but simply because they may not want to incur the costs of ranking all schools up to $K$. In the rest of this subsection, we consider mechanisms that impose a cap $K$ on $P$ and assume that students submit a weak partial order of their true preferences.

assumption[Submission of Weak Partial Order] Students submit a weak partial order of their true preferences.

Assumption (ref) replaces Assumption (ref) to accommodate mechanisms that are not strategy proof. Submitting a weak partial order is rational in DA mechanisms with cap constraints. That is true for any mechanism that becomes strategy proof once we remove the cap constraint, for example, the TTC mechanism.

An assumption maintained throughout this paper is that the cutoff characterization from Definition (ref) applies to the mechanism. This assumption says that $\mu(\omega) = Q(B(\boldsymbol{\mathrm{S}}(\omega))$ for every $\omega \in \Omega$. azevedo2016 show that stability is equivalent to cutoff characterization with $\boldsymbol{\mathrm{S}}=\boldsymbol{\mathrm{R}}$ and cutoffs that equal the minimum score of the admitted students in each school. Thus, it is worth discussing the stability of the constrained DA mechanism. Theorem 6.3 from haeringer2009 demonstrates that any Nash equilibrium in constrained DA where $\boldsymbol{\mathrm{R}}$ satisfies Ergin acyclicity leads to a stable matching in the finite economy.\footnote{ Ergin acyclicity ensures that no student can block a potential improvement for any two other students without affecting her own assignment. See ergin2002 for the formal definition. } Even without Ergin acyclicity, some Nash equilibria still produce stability. SD always satisfies Ergin acyclicity, and thus every Nash equilibrium in SD produces a stable matching. fack2019 also study the constrained DA mechanism. They extend Theorem 6.3 from haeringer2009 to (pure-strategy) Bayesian Nash equilibria in the continuum economy (Proposition A3, Online Appendix A.2.5, fack2019). fack2019 also provide primitive conditions for finite economies where students play partial orders to converge to a continuum economy with a stable equilibrium (Proposition 5, fack2019). They further provide a test for implications of stability and find no empirical or simulation evidence against it. In the context of the constrained TTC mechanism, any Nash equilibrium leads to a stable matching as long as $\boldsymbol{\mathrm{R}}$ satisfies Kesten acyclicity (Theorem 6.4 from haeringer2009). Therefore, constrained SD, DA, and TTC all satisfy the assumption of cutoff characterization as in Definition (ref) with $\boldsymbol{\mathrm{S}}=\boldsymbol{\mathrm{R}}$ under the appropriate conditions.

There is another interesting feature of the cutoff characterization of DA mechanisms. We know that DA produces a stable matching if agents are truth-tellers. In case agents are not truth-tellers, the matching outcome continues to be “stable” if we replace $Q$ with $P$ in the definition of stability.

definition[Stability wrt $P$] We say the matching $\mu: \Omega \to \mathcal{J}^0$ is a stable matching wrt $P$ if three conditions are satisfied for every $\omega \in \Omega$: (i) $\mu(\omega) \bar P(\omega) 0$ (individual rationality); (ii) for any $j \in \mathcal{J}$, if $j P(\omega) \mu(\omega)$, then $j$ is full (no waste); and (iii) for any $j \in \mathcal{J}$ that is full, if $\mu(\omega')=j$ and $ j P(\omega) \mu(\omega)$, then $R_j(\omega') > R_j(\omega )$ (no justified envy), where we adopt the convention that $m P 0$ for every $m \in P$. This is the same as Definition (ref) except that $P$ appears in the place of $Q$.

The DA mechanism, constrained or unconstrained, produces a matching that is stable wrt reported preferences $P$. Stability wrt $P$ leads to a cutoff characterization wrt $P$ according to the work of azevedo2016. This cutoff characterization has scores $\boldsymbol{\mathrm{S}}=\boldsymbol{\mathrm{R}}$ and admission cutoffs that equal the smallest scores of admitted students in each school. In other words, this is the same cutoff characterization from Definition (ref) except that $Q$ is replaced with $P$. Cutoff characterization wrt $P$ is natural in DA but not necessarily in other mechanisms, so we state it in the following assumption.

assumption[Cutoff Characterization wrt $P$] In addition to the maintained assumption of cutoff characterization as in Definition (ref), the matching function $\mu$ satisfies $\mu(\omega) = P(B(\boldsymbol{\mathrm{S}}(\omega))$ for every $\omega \in \Omega$.

Assumption (ref) essentially says that agents are matched to their best feasible options, where best is now defined according to $P$. Assumption (ref) is convenient because it allows us to write a simple expression for the identified set of local preferences in Proposition (ref) below; however, it is not a necessary assumption for the identification of those sets. The convenience comes from the fact that Assumption (ref) implies $\mu = P(B(\boldsymbol{\mathrm{S}})) = Q(B(\boldsymbol{\mathrm{S}}))$, where both $\mu$ and $P$ are observed and $P$ and $Q$ are related via the weak partial order assumption. If we drop Assumption (ref), we have only one equality $\mu = Q(B(\boldsymbol{\mathrm{S}}))$, which leads to larger sets of local preferences in Proposition (ref) below. This is useful to know for settings such as those with the TTC mechanism, which is not stable wrt $P$.

Next, we characterize all possible pairs of local preferences at a cutoff that are compatible with the data and Assumptions (ref) and (ref).

proposition[Identification of Local Preference Sets] Suppose Assumptions (ref) and (ref) hold. Select a school $j$ with cutoff $c_j$. Consider a student with scores $\boldsymbol{\mathrm{S}} \equiv (S_j, \boldsymbol{\mathrm{S}}_{-j})$ and submitted preferences $P$. Call $(a,b) = P_j$. For this student, define $N_j^+ = B^+_j( \boldsymbol{\mathrm{S}} ) \setminus \left\{ P \cup \{ 0 \} \right\}$ and $N_j^- = B^-_j( \boldsymbol{\mathrm{S}} ) \setminus \left\{ P \cup \{ 0 \} \right\}$, respectively, the sets of unlisted feasible schools in the counterfactual budget sets to the right and the left of the cutoff. Then, the $Q_{j}$ of this student belongs to $\boldsymbol{\mathrm{Q}}_j$, where the set $\boldsymbol{\mathrm{Q}}_j$ is defined as follows: \begin{align} \boldsymbol{\mathrm{Q}}_j = \left\{ \begin{array}{ll} \{ (a,b) \}, & if S_j \geq c_j and a=b, \\ \{ (a,b) \} \cup \left(\{ a \} \times N_{j}^-\right), & if S_j \geq c_j and a \neq b, \\ \{ (a,b) \} \cup \left( (N_{j}^+ \setminus N_{j}^-) \times \{ b \} \right) & if S_j < c_j, \end{array} \right. \end{align} where $\left(\{ a \} \times N_{j}^- \right)$ denotes the set formed by the Cartesian product of $a$ and elements in $N_{j}^-$ and $\left(\{ a \} \times N_{j}^- \right) = \emptyset$ if $ N_{j}^- = \emptyset$. Moreover, assume $P$ is a strong partial order of $Q$. Then, $\boldsymbol{\mathrm{Q}}_j$ becomes: \begin{align} \boldsymbol{\mathrm{Q}}_j = \left\{ \begin{array}{ll} \{ (a,b) \}, & if |P|<K, \text{ or if } |P|=K, S_j \geq c_j, \text{ and } a=b, \\ \{ (a,b) \} \cup \left(\{ a \} \times N_{j}^-\right), & \text{ if } |P|=K, S_j \geq c_j, \text{ and } a \neq b, \\ \{ (a,b) \} \cup \left( (N_{j}^+ \setminus N_{j}^-) \times \{ b \} \right) & \text{ if } |P|=K \text{ and } S_j < c_j. \end{array} \right. \end{align} Finally, the characterization in (ref) is sharp if the distribution of $Q$ conditional on $P$ and $\boldsymbol{\mathrm{S}}$ has full support, that is, if every $Q \in \mathcal{Q}$ that satisfies Assumptions (ref) and (ref) is in that support. Likewise, (ref) is sharp if the distribution of $Q$ conditional on $P$ and $\boldsymbol{\mathrm{S}}$ has full support under Assumptions (ref) and (ref) and $P$ being a strong partial order.

We illustrate the proposition in terms of the SD Example.

example*[SD Example, Part III] Suppose the cap constraint is $K=3$ and the four schools are acceptable for everyone. We consider all agents whose $P_4=(4,2)$. For example, if agents submit strong partial orders, they submit either $P=\{4,2,1\}$ or $P=\{4,2,3\}$. The assumption of cutoff characterization wrt $P$ (Assumption (ref)) says that these agents are matched to school $4$ if $S_1\geq c_4$ and to school $2$ otherwise. To keep things simple, consider five different types of true preferences: $Q^{(1)}=\{4,2,3,1,0\}$, $Q^{(2)}=\{4,3,2,1,0\}$, $Q^{(3)}=\{4,1,3,2,0\}$, $Q^{(4)}=\{3,4,2,1,0\}$, and $Q^{(5)}=\{1,3,2,4,0\}$. The weak partial order assumption rules out $Q^{(5)}$ because $2 Q^{(5)} 4$ contradicts $4$ being reported preferred to $2$. The maintained assumption of cutoff characterization (Definition (ref)) further rules out more types of $Q$, depending on whether $S_1 \geq c_4$ or $S_1 < c_4$: \begin{enumerate} • if $S_1\geq c_4$, $Q^{(4)}$ is not possible because the matching assignment is $4$ but the best feasible option according to $Q^{(4)}$ is $3$; in this case, the possible true local preferences are: $Q^{(1)}_4=(4,2)$, $Q^{(2)}_4=(4,3)$, and $Q^{(3)}_4=(4,1)$; for a student who submits $P=\{4,2,1\}$, $\boldsymbol{\mathrm{Q}}_4 = \{(4,2), (4,3)\} $; otherwise, for someone who submits $P=\{4,2,3\}$, $\boldsymbol{\mathrm{Q}}_4 = \{(4,2), (4,1)\} $; • if $S_1 < c_4$, none of $Q^{(2)}$, $Q^{(3)}$, or $Q^{(4)}$ is possible because the matching assignment is $2$ but the best feasible options according to these $Q$s differ from $2$; in this case, the only possible true local preference is $Q^{(1)}_4=(4,2)$, so that $\boldsymbol{\mathrm{Q}}_4 = \{(4,2)\} $. \end{enumerate}

This example illustrates why an RD at $c_4$ that controls for $P_4=(4,2)$ might be problematic. The range of possibilities for true preference types is different between individuals above and below the cutoff. We see types $(4,1)$, $(4,2)$, and $(4,3)$ above the cutoff but only type $(4,2)$ below the cutoff. Suppose in an extreme case there is a nonzero fraction of individuals with types $(4,1)$ or $(4,3)$ right above the cutoff. True preferences arguably affect outcomes and this implies that the distribution of outcomes conditional on $P_4=(4,2)$ and $S_1=s$ changes as $s$ crosses the cutoff $c_4$ even without a treatment effect. This extreme case breaks internal validity of RD identification. Even if the fraction of types $(4,1)$ or $(4,3)$ at the cutoff is zero and RD identification is valid, a continuous but sudden increase in the fraction of such agents on the right of the cutoff may bring severe bias issues and make inference impossible (e.g., bertanha_impossible, in particular, Sections 5 and A.6).

For the issue to arise when we control for $P_4=(4,2)$, there must be at least a fraction of agents with true local preferences $(4,1)$ and $(4,3)$ who change their submission behavior discontinuously as a function of $S_1$ as $S_1$ crosses the value of $c_4$; i.e., they must change $P$ such that $P_4 \neq (4,2)$ below the cutoff and $P_4 = (4,2)$ above the cutoff. Intuitively, this requires these agents to have good ex-ante knowledge about the ex-post value of the cutoff $c_4$. The more knowledge they have about the ex-post cutoff values, the sharper will be their change in behavior around the cutoff, and the closer we are to the extreme case mentioned above. Section (ref) in the appendix presents a numerical example of an economy with agents that maximize expected utility in face of cutoff uncertainty and exhibit such discontinuous behavior at the ex-post cutoffs. An interesting feature of that example is that agents do not need to have perfect knowledge of the cutoffs to exhibit discontinuous behavior. Finally, Section (ref) displays empirical evidence from Chile that supports agents having good a priori knowledge of cutoffs; we also show that submission behavior changes near ex post cutoffs and the change is discontinuous in some cases.

Proposition (ref) identifies all possible values of $Q_j$ for students near a cutoff $c_j$ as a function of their scores and submitted preferences. In some contexts, students may have a large number of feasible but unlisted programs, resulting in sets $\boldsymbol{\mathrm{Q}}_j$ with many possible values. For instance, in the Chilean data, $K=8$ but there are over 1,000 programs; a student may have many feasible options but choose not to list most of them. We now introduce one approach to reducing the size of $\boldsymbol{\mathrm{Q}}_j$ by placing additional restrictions on the expectations students hold when submitting $P$. Other context-specific assumptions can also be employed—for example, in Section (ref), we use an assumption based on preference for fields of study when analyzing post-secondary education in Chile.

agarwal2018 propose a general framework to rationalize strategic reporting as the optimal solution to an expected utility maximization problem. In this framework, agents have private information about their preferences and scores and form beliefs about the distribution of other people's preferences and scores. These beliefs plus knowledge of the mechanism lead the rational agent to derive probabilities of admission to the various schools as a function of the agent's private information and expectations about other agents. The agent then chooses the submission $P$ that maximizes her expected utility.

For our next proposition, we assume that agents are expected utility maximizers where the uncertainty about their match comes from uncertainty about what the admission cutoffs will be after the matching algorithm is run. As such, we assume each agent forms beliefs on admission cutoffs (see Section (ref) in the appendix for a formal definition of the problem of the agent). Uncertainty about cutoffs is key in our continuum economy with cutoff characterization because cutoffs and scores fully characterize the agent's budget set. Given the student's scores, a distribution of possible cutoffs translates into a distribution of possible budget sets. Under Assumption (ref), the student is admitted to the best school according to the submission $P$ among the available schools in the budget set. Therefore, beliefs on cutoffs translate into probabilities of admission to various schools for any given $P$. We make an assumption on the distribution of cutoffs expected by agents that has to do with the concept of uniformly more accessible schools.

definition[Uniformly More Accessible Schools] For a pair of distinct schools $(d,e)$, we say $e$ is uniformly more accessible than $d$ if two conditions are satisfied: first, if access to school $d$ implies access to school $e$, \[ \{\omega: S_d(\omega) \geq c_d \} \subseteq \{\omega: S_e(\omega) \geq c_e \}, \] and second, if replacing option $d$ with option $e$ in any submission $P$ alters the likelihood of admission for at least one school listed in $P$; formally, for any two fixed (i.e., nonrandom) submissions $P$ and $\ti P$ such that $P$ has $d$ but does not have $e$ and $\ti P$ equals $P$ except for $e$ in the place of $d$, there exists $u \in \{0,1, \ldots, |P| \}$ for which \[ \mathbb{P}\left[ P(B(\boldsymbol{\mathrm{S}})) = P^u \right] \neq \mathbb{P}\left[ \ti P(B(\boldsymbol{\mathrm{S}})) = \widetilde{P}^u \right], \] where $P(B)$ denotes the best choice in set $B$ according to $P$ (Definition (ref)) and $P^u$ denotes the school ranked in the $u$-th position in $P$. In short, we say $(d,e) \in UMAS$, where $UMAS\subseteq \mathcal{J} \times \mathcal{J}$ is the set of all such pairs.

Definition (ref) says that $e$ is uniformly more accessible than $d$ if everyone who qualifies for school $d$ also qualifies for school $e$. Schools $d$ and $e$ must also be relevant in the sense of the second condition: there is always a strictly positive fraction of individuals for whom listing $e$ in the place of $d$ changes their best feasible options. In the SD case, a sufficient condition for Assumption (ref) is that $c_d>c_e$ and the cutoffs are distinct interior points in the support of the placement score. Uniformly more accessible schools do not always exist. Whether they do depends on the mechanism in place and the joint distribution of the placement scores. We use this definition to impose a mild restriction on the expectations of agents regarding cutoffs.

assumptionConsider a student with scores $\boldsymbol{\mathrm{s}} \in \mathcal{S}$ who views uncertain cutoffs as random variables $C_1, \ldots, C_J$ before the matching assignment. Let $\widetilde{B} = \{ 0 \} \cup \{j \in \mathcal{J} ~:~ s_j \geq C_j \}$ be the student’s corresponding random budget set. For every pair $(d,e) \in UMAS$, the distribution of cutoffs for this student is such that two conditions are satisfied: first, \[ \{s_d \geq C_d \} \subseteq \{s_e \geq C_e \}, \] and second, for any two fixed (i.e., nonrandom) submissions $P$ and $\ti P$ such that $P$ has $d$ but does not have $e$ and $\ti P$ equals $P$ except for $e$ in the place of $d$, there exists $u \in \{0,1, \ldots, |P| \}$ for which \[ \mathbb{P}\left[ P(\widetilde{B}) = P^u \right] \neq \mathbb{P}\left[ \ti P(\widetilde{B}) = \widetilde{P}^u \right]. \] This is true for every student in the economy.

Assumption (ref) says that students correctly anticipate which schools will be uniformly more accessible after the matching assignment. For example, agents may learn this information by observing past realizations of the matching in the economy. If a school $e$ is well known to be accessible to everyone who has access to school $d$, then it is natural for a student to expect to have access to $e$ if she ever has access to school $d$. Note that the assumption does not pin down the expected probability of admission or the set of schools to which the student will have access in the ex-post economy. It restricts only the expected hierarchy of school access according to $UMAS$. This assumption has implications for the joint distribution of $(P,Q)$.

propositionSuppose Assumptions (ref)--(ref) hold. Consider a student with reported preference ranking $P$. Let $\left( P \times P^c \right)$ be the Cartesian product of listed and unlisted schools, respectively, $P$ and $P^c$. If $(d,e) \in UMAS \cap \left( P \times P^c \right)$, then $d Q e$.

Proposition (ref) says that if an agent lists school $d$ but does not list the uniformly more accessible school $e$, it must be that this agent prefers $d$ over $e$. This result offers a refinement of Proposition (ref) above.

corollaryConsider the setup of Proposition (ref), where $P_j=(a,b)$, and suppose Assumption (ref) holds. Define $A_j^- = N^-_j \setminus \left\{ e: \exists d \in P \text{ with which } (d,e) \in UMAS \text{ and } b \bar{P} d \right\}$ and\\ $A_j^+ = \left(N^+_j \setminus N^-_j \right) \setminus \left\{ e: \exists d \in P \text{ with which } (d,e) \in UMAS \text{ and } a \bar{P} d \right\}$. Then, under weak partial order, \begin{align} \boldsymbol{\mathrm{Q}}_j = \left\{ \begin{array}{ll} \{ (a,b) \}, & if S_j \geq c_j and a=b, \\ \{ (a,b) \} \cup \left(\{ a\} \times A_{j}^-\right), & if S_j \geq c_j and a \neq b, \\ \{ (a,b) \} \cup \left( A_j^+ \times \{ b \} \right) & if S_j < c_j. \end{array} \right. \end{align} Under strong partial order, \begin{align} \boldsymbol{\mathrm{Q}}_j = \left\{ \begin{array}{ll} \{ (a,b) \}, & if |P|<K, \text{ or if } |P|=K, S_j \geq c_j, \text{ and } a=b, \\ \{ (a,b) \} \cup \left(\{ a \} \times A_{j}^-\right), & \text{ if } |P|=K, S_j \geq c_j, \text{ and } a \neq b, \\ \{ (a,b) \} \cup \left( A_j^+ \times \{ b \} \right) & \text{ if } |P|=K \text{ and } S_j < c_j. \end{array} \right. \end{align} These characterizations are sharp as long as (ref)--(ref) are sharp in their respective contexts in Proposition (ref) and imposing Assumption (ref) sets to zero only the following probabilities: $\mathbb{P}\left[ e Q d | P, \boldsymbol{\mathrm{S}} \right]$ for every $(d,e) \in UMAS \cap \left( P \times P^c \right)$.

Corollary (ref) describes how to use Proposition (ref) to potentially reduce the number of elements in the $\boldsymbol{\mathrm{Q}}_j$ constructed in Proposition (ref). The intuition runs as follows. Suppose that a student submits $P$ and $P_j=(a,b)$. If school $e$ is uniformly more accessible than school $d$ and $d$ is listed in $P$ but $e$ is not listed in $P$, then we know the student truly prefers $d$ over $e$. This excludes some possibilities of $Q_j$ in the $\boldsymbol{\mathrm{Q}}_j$ defined by Proposition (ref). For instance, this person cannot have $Q_j=(a,e)$ if $b \bar{P} d$ because that presupposes $e Q b \bar{Q} d $, which contradicts $d Q e$. Likewise, this person cannot have $Q_j=(e,b)$ if $a \bar{P} d$.

example*[SD Example, Part IV] The set of uniformly more accessible schools is $UMAS = \{(2,1), (3,2), (3,1), (4,3), (4,2), (4,1) \}$. Assumption (ref) shrinks the set $\boldsymbol{\mathrm{Q}}_4$ of those agents with $S_1\geq c_4$ and $P=\{4,2,3\}$. Applying the assumption changes $\boldsymbol{\mathrm{Q}}_4 =\{(4,2), (4,1)\}$ to $\boldsymbol{\mathrm{Q}}_4 = \{(4,2)\}$ because $1$ is uniformly more accessible than $2$, $2$ is listed, and $1$ is not listed, so Proposition (ref) implies $2Q1$.

Partial Identification of Causal Effects

In this section, we lay out conditions and derive bounds on average treatment effects. We assume that the researcher has already identified the set of local preferences at a cutoff of interest. This means that the researcher has a set-valued variable $\boldsymbol{\mathrm{Q}}_j$ for all students in the vicinity of a cutoff $j$ corresponding to a comparable pair $(j,k)$ in $\mathcal{P}$. Researchers may construct $\boldsymbol{\mathrm{Q}}_j$ using the methods in Section (ref) if they find it reasonable to rely on at least some of the specific assumptions in that subsection; otherwise, they may use any other method to construct $\boldsymbol{\mathrm{Q}}_j$. There is no restriction on the choice of the method for constructing $\boldsymbol{\mathrm{Q}}_j$ except for a couple of high-level conditions that we assume to hold in this section. We start by defining the conditional support of partially identified true local preferences.

definition[Support of Local Preference Sets] Consider a pair $(j,k) \in \mathcal{P}$ and corresponding cutoff $c_j$. The support of partially identified true local preferences conditional on $S_j=s$ is defined as \[ \boldsymbol{\mathrm{\Lambda}}_{j}(s) = \left\{ B\subseteq \mathcal{J}^0 \times \mathcal{J}^0 : \mathbb{P}\left[ \boldsymbol{\mathrm{Q}}_j = B | S_j=s \right] >0 \right\}. \] The union set of this support is defined as the collection of all unions of sets in $\boldsymbol{\mathrm{\Lambda}}_{j}(s)$, namely, \[ \boldsymbol{\mathrm{\Lambda}}^{\cup}_{j}(s) = \left\{ B^{\cup} \subseteq \mathcal{J}^0 \times \mathcal{J}^0 : \exists B_1, B_2, \ldots \in \boldsymbol{\mathrm{\Lambda}}_{j}(s) \text{ with } B^{\cup} = \cup_{i} B_i \right\}. \]

The set $\boldsymbol{\mathrm{\Lambda}}_{j}(s)$ collects all values of $\boldsymbol{\mathrm{Q}}_j$ that occur with positive probability conditional on $S_j=s$. In the specific context of Section (ref), $\boldsymbol{\mathrm{Q}}_j$ is constructed from the mapping of observables $(P,\boldsymbol{\mathrm{S}})$ to a subset of $\mathcal{J}^0 \times \mathcal{J}^0$, i.e., $\boldsymbol{\mathrm{Q}}_j = \psi_j (P,\boldsymbol{\mathrm{S}})$. For example, Proposition (ref) and Corollary (ref) give examples of such mapping $\psi_j$. A set $B$ of pairs $(a,b) \in \mathcal{J}^0 \times \mathcal{J}^0$ belongs to the support set $\boldsymbol{\mathrm{\Lambda}}_{j}(s)$ if there is a set of values in the support of the conditional distribution of $(P,\boldsymbol{\mathrm{S}})$ given $S_j=s$ such that $\psi_j$ maps those values to the set $B$. The union set $\boldsymbol{\mathrm{\Lambda}}^{\cup}_{j}(s)$ collects all possible unions of support points of $\boldsymbol{\mathrm{Q}}_j$ conditional on $S_j=s$. These definitions are instrumental in the computation of the partially identified distribution of $Q_j$, as explained in Proposition (ref) below.

Sharpness of identification of the distribution of $Q_j$ requires sharpness in the construction of the sets $\boldsymbol{\mathrm{Q}}_j$. Proposition (ref) and Corollary (ref) gave the conditions for sharpness of $\boldsymbol{\mathrm{Q}}_j$ in the context of Section (ref). Outside that context, researchers may construct $\boldsymbol{\mathrm{Q}}_j$ in a different way, so we impose sharpness of $\boldsymbol{\mathrm{Q}}_j$ in the general form of the assumption below.

assumption[Sharp Local Preference Sets] Consider a pair $(j,k) \in \mathcal{P}$ and corresponding cutoff $c_j$. Assume that: (i) the random variable $Q_{j}$ and the random set $\boldsymbol{\mathrm{Q}}_j$ are both measurable maps on the same probability space and $\mathbb{P}\left[ Q_j \in \boldsymbol{\mathrm{Q}}_j ~|~ S_j \right]=1$ with probability 1; and (ii) $\mathrm{supp}\left[ Q_j ~|~ \boldsymbol{\mathrm{Q}}_j, S_j \right] = \boldsymbol{\mathrm{Q}}_j $ with probability $1$, where $\mathrm{supp}\left[ Y ~|~ X \right]$ denotes the support set of the distribution of $Y$ conditional on $X$.

Assumption (ref)(i) says that $\boldsymbol{\mathrm{Q}}_j(\omega)$ of individual $\omega$ contains the true pair of local preferences $Q_j(\omega)$ of that individual (for almost all individuals), which is a minimum requirement for the construction of $\boldsymbol{\mathrm{Q}}_j(\omega)$. This does not say anything about the sharpness of $\boldsymbol{\mathrm{Q}}_j$. For example, $\boldsymbol{\mathrm{Q}}_j = \mathcal{J}^0 \times \mathcal{J}^0$ is completely uninformative and trivially satisfies Assumption (ref)(i). The sharpness requirement is stated in Assumption (ref)(ii). It says that all possibilities of local preferences listed in $\boldsymbol{\mathrm{Q}}_j$ actually occur in the data with positive probability. This rules out unnecessarily large sets $\boldsymbol{\mathrm{Q}}_j$. Assumption (ref)(ii) may be dropped at the cost of lacking sharpness in the identified sets in the rest of this section.

Partial identification of true local preferences and treatment effects occurs at the limit, as $S_j$ approaches $c_j$, and is conditional on $\boldsymbol{\mathrm{Q}}_j$. For this to work, we impose regularity conditions on the distribution of potential outcomes and $\boldsymbol{\mathrm{Q}}_j$ conditional on $S_j$ at the limit $c_j$.

assumption[Distribution of Local Preference Sets] Consider a pair $(j,k) \in \mathcal{P}$ and corresponding cutoff $c_j$. Assume that: (i) there exist a small $\varepsilon>0$ and collections of subsets of $\mathcal{J}^0 \times \mathcal{J}^0$ denoted $\boldsymbol{\mathrm{\Lambda}}_{j}^{+}$ and $\boldsymbol{\mathrm{\Lambda}}_{j}^{-}$ such that $\boldsymbol{\mathrm{\Lambda}}_{j}^{+} = \boldsymbol{\mathrm{\Lambda}}_{j}(c_j+e)$ $\forall e \in [0,\varepsilon)$ and $\boldsymbol{\mathrm{\Lambda}}_{j}^{-} = \boldsymbol{\mathrm{\Lambda}}_{j}(c_j-e)$ $\forall e \in (0,\varepsilon)$; consistent with Definition (ref), we define $\boldsymbol{\mathrm{\Lambda}}_{j}^{\cup+}$ and $\boldsymbol{\mathrm{\Lambda}}_{j}^{\cup-}$ as union sets of $\boldsymbol{\mathrm{\Lambda}}_{j}^{+}$ and $\boldsymbol{\mathrm{\Lambda}}_{j}^{-}$, respectively; (ii) for any $g \in \mathcal{G}$ of Assumption (ref) and any $\underline{\tau}, \overline{\tau} \in \mathbb{R} \cup \{-\infty, +\infty \}$, $\underline{\tau} < \overline{\tau}$, the side limits of the following expectations are well defined: $\mathbb{E}\left[ g(Y) \mathbb{I}\{ \boldsymbol{\mathrm{Q}}_j = A, \underline{\tau} < g(Y) < \overline{\tau} \} ~|~ S_j= c_j^+ \right] \; \; \forall A \in \boldsymbol{\mathrm{\Lambda}}_{j}^{+}$ and $\mathbb{E}\left[ g(Y) \mathbb{I}\{ \boldsymbol{\mathrm{Q}}_j = A, \underline{\tau} < g(Y) < \overline{\tau} \} ~|~ S_j= c_j^- \right] \; \; \forall A \in \boldsymbol{\mathrm{\Lambda}}_{j}^{-}$.

Assumption (ref)(i) concerns the distribution of $\boldsymbol{\mathrm{Q}}_j$ conditional on $S_j$: the support set of $\boldsymbol{\mathrm{Q}}_j$ is constant as $S_j=s$ approaches the cutoff $c_j$ from either side of it. Part (ii) of the assumption concerns the joint distribution of potential outcomes and $\boldsymbol{\mathrm{Q}}_j$ conditional on $S_j$. For example, Assumption (ref)(ii) implies that $\mathbb{P}\left[ \boldsymbol{\mathrm{Q}}_j = A ~|~ S_j= c_j^+ \right]$ and $\mathbb{E}[ Y ~|~ \boldsymbol{\mathrm{Q}}_j = A, Y < \tau, S_j= c_j^+ ]$ are well-defined limits for any $A \in \boldsymbol{\mathrm{\Lambda}}_{j}^{+}$ and $\tau \in \mathbb{R} \cup +\infty$ provided that $\mathbb{P} [ \boldsymbol{\mathrm{Q}}_j = A, Y < \tau ~|~ S_j= c_j^+ ]>0.$ The next result gives inequalities to construct bounds on $\mathbb{P}[Q_j=(a,b) | S_j=c_j]$ for any pair $(a,b)$.

proposition[Sharp Set of Distributions of Local Preferences] Consider a pair $(j,k) \in \mathcal{P}$. Suppose Assumptions (ref), (ref), and (ref) hold. Then, the sharp set of all possible discrete probability distributions of $Q_{j}$ conditional on $S_j=c_j$ is characterized as follows. For every $A \in \boldsymbol{\mathrm{\Lambda}}^{\cup+}_{j} \cup \boldsymbol{\mathrm{\Lambda}}^{\cup-}_{j}$, each probability distribution in that set implies a value for $\mathbb{P}\left[ Q_j \in A | S_j= c_j \right]$ that satisfies one of the three inequalities below: \begin{enumerate}[(i)] • if $A \in \boldsymbol{\mathrm{\Lambda}}^{\cup+}_{j} \cap \boldsymbol{\mathrm{\Lambda}}^{\cup-}_{j}$, \[ \mathbb{P}\left[ Q_j \in A | S_j= c_j \right] \geq \max \left\{~ \mathbb{P}\left[ \boldsymbol{\mathrm{Q}}_j \subseteq A | S_j= c_j^+ \right] ~;~ \mathbb{P}\left[ \boldsymbol{\mathrm{Q}}_j \subseteq A | S_j= c_j^- \right] ~\right\}; \] • if $A \in \boldsymbol{\mathrm{\Lambda}}^{\cup+}_{j} \setminus \boldsymbol{\mathrm{\Lambda}}^{\cup-}_{j}$, \[ \mathbb{P}\left[ Q_j \in A | S_j= c_j \right] \geq \mathbb{P}\left[ \boldsymbol{\mathrm{Q}}_j \subseteq A | S_j= c_j^+ \right]; \text{ or } \] • if $A \in \boldsymbol{\mathrm{\Lambda}}^{\cup-}_{j} \setminus \boldsymbol{\mathrm{\Lambda}}^{\cup+}_{j}$, \[ \mathbb{P}\left[ Q_j \in A | S_j= c_j \right] \geq \mathbb{P}\left[ \boldsymbol{\mathrm{Q}}_j \subseteq A | S_j= c_j^- \right]. \] \end{enumerate}

Proposition (ref) provides a way to construct the sharp partially identified set of all possible distributions of $Q_{j}$ conditional on $S_j=c_j$. A distribution of $Q_{j}$ conditional on $S_j=c_j$ consists of values $p_{a,b} \in [0,1]$ for every $(a,b) \in \mathcal{J}^0 \times \mathcal{J}^0$ such that $\sum_{(a,b) \in \mathcal{J}^0 \times \mathcal{J}^0} ~ p_{a,b} = 1$, where $p_{a,b} = \mathbb{P}\left[ Q_j =(a,b) | S_j= c_j \right].$ The sharp set is constructed by finding all values of $p_{a,b}$ where $\sum_{(a,b) \in A} ~ p_{a,b}$ satisfies the inequalities of Proposition (ref) for every $A \in \boldsymbol{\mathrm{\Lambda}}^{\cup+}_{j} \cup \boldsymbol{\mathrm{\Lambda}}^{\cup-}_{j}$.

example*[SD Example, Part V] Continue to assume that the four schools are acceptable for everyone. Suppose for a moment that all combinations of $(P,Q)$ that satisfy Assumptions (ref)--(ref) exist in the economy, for both $S_1 \geq c_4$ and $S_1<c_4$. Then, the list of all possible $\boldsymbol{\mathrm{Q}}_4$ is as follows: \begin{enumerate} • if $S_1 \geq c_4$, $\{(1,1)\}$, $\{(2,2)\}$, $\{(3,3)\}$, $\{(4,1)\}$, $\{(4,2)\}$, $\{(4,3)\}$, $\{(4,1),(4,2),(4,3)\}$, $\{(4,1),(4,3)\}$, and $\{(4,2),(4,3)\}$; • if $S_1 < c_4$, $\{(1,1)\}$, $\{(2,2)\}$, $\{(3,3)\}$, $\{(4,1)\}$, $\{(4,2)\}$, $\{(4,3)\}$, $\{(1,1),(4,1)\}$, $\{(2,2),(4,2)\}$, and $\{(3,3),(4,3)\}$. \end{enumerate} Let us focus on the case that $S_1\geq c_4$. To keep things simple, suppose three types of $\boldsymbol{\mathrm{Q}}_4$ occur with positive probability conditional on $S_1=s$ for any $s \geq c_4$: $\{(4,2)\}$ with probability $0.1$, $\{(4,3)\}$ with probability $0.3$, and $\{(4,2),(4,3)\}$ with probability $0.6$. It follows that $\boldsymbol{\mathrm{\Lambda}}_{4}(s)=\boldsymbol{\mathrm{\Lambda}}_{4}^+=\{ \{(4,2)\},\{(4,3)\}, \{(4,2),(4,3)\} \} $ and $\boldsymbol{\mathrm{\Lambda}}_{4}^{\cup}(s)= \boldsymbol{\mathrm{\Lambda}}_{4}^{\cup+}= \{ \{(4,2)\},\{(4,3)\}, \{(4,2),(4,3)\} \} $. The lower bounds $\mathbb{P}[\boldsymbol{\mathrm{Q}}_4 \subseteq A | S_1 = c_4^+]$ of Proposition (ref) are as follows: $0.1$ for $A=\{(4,2)\}$; $0.3$ for $A=\{(4,3)\}$; and $1$ for $A=\{(4,2),(4,3)\}$. Thus, $\mathbb{P}[Q_4 = (4,2) | S_1 = c_4^+]$ has lower bound $0.1$, $\mathbb{P}[Q_4 = (4,3) | S_1 = c_4^+]$ has lower bound $0.3$, and the sum of the two equals $1$. When we look at each individual probability, the bounds are $[0.1,0.7]$ on $\mathbb{P}[Q_4 = (4,2) | S_1 = c_4^+]$ and $[0.3,0.9]$ on $\mathbb{P}[Q_4 = (4,3) | S_1 = c_4^+]$.

Recall that the construction of the random set $\boldsymbol{\mathrm{Q}}_j$ depends on assumptions regarding the behavior of agents when they submit $P$. For example, Section (ref) characterizes $\boldsymbol{\mathrm{Q}}_j$ by assuming weak partial order and cutoff characterization wrt $P$. Alternatively, the identification approach of agarwal2018 makes different types of assumptions on agents' expectations and requires data variation in the choice environment. The theoretical credibility of these types of assumptions depends on the mechanism faced by agents; in practice, the assumptions have testable implications for what we should observe in the data. Section (ref) in the appendix provides a testable implication of our assumptions.

Partial identification of the distribution of local preferences allows us to bound the fraction of individuals near cutoff $c_j$ who have $Q_j=(j,k)$. The average outcome near the cutoff is a weighted average of the average outcomes from two different groups: first, individuals with $Q_j=(j,k)$, who interest us for the identification of treatment effects; and second, individuals with $Q_j \neq (j,k)$. The overall average is identified, but the average in each of the groups is not. A strictly positive lower bound on the fraction of individuals in the $Q_j=(j,k)$ group allows us to construct lower and upper bounds on the average outcome for that group.

Start with all individuals above and near cutoff $c_j$ whose $\boldsymbol{\mathrm{Q}}_j$ contain the comparable pair of interest, $(j,k) \in \mathcal{P}$. The fraction of those individuals who have $Q_j=(j,k)$ equals \[ \delta_{j,k}^+ = \frac{ \mathbb{P}\left[ Q_j=(j,k)|S_j=c_j \right] }{ \mathbb{P}\left[ \boldsymbol{\mathrm{Q}}_j \cap \{ (j,k) \} \neq \emptyset |S_j=c_j^+ \right] }, \] where both numerator and denominator are strictly positive by virtue of $(j,k)$ being a comparable pair (Definition (ref)) and of the sharpness of $\boldsymbol{\mathrm{Q}}_j$ (Assumption (ref)). The denominator of $\delta_{j,k}^+$ is identified from the data, and Proposition (ref) bounds the numerator. All we need for identification of treatment effects is a lower bound on $\delta_{j,k}^+$, which comes from a lower bound on its numerator. Let $\underline{p}_{j,k}$ denote the infimum over all probability values for $\mathbb{P}\left[ Q_j=(j,k)|S_j=c_j \right]$ that belong to the partially identified set of Proposition (ref). The sharp lower bound on $\delta_{j,k}^+$ equals \[ \underline{\delta}_{j,k}^+ = \frac{ \underline{p}_{j,k} }{ \mathbb{P}\left[ \boldsymbol{\mathrm{Q}}_j \cap \{ (j,k) \} \neq \emptyset |S_j=c_j^+ \right] }. \] The denominator of $\underline{\delta}_{j,k}^+$ is strictly positive, but $\underline{p}_{j,k}$ may or may not be strictly positive.

The same idea applies for individuals just below the cutoff:

align*[align* omitted — 351 chars of source]
example*[SD Example, Part VI] We have that $\mathbb{P}\left[ \boldsymbol{\mathrm{Q}}_4 \cap \{ (4,2) \} \neq \emptyset |S_1=c_4^+ \right]=0.7$ and the bounds on $\mathbb{P}\left[ Q_4=(4,2)|S_1=c_4 \right]$ are $[0.1,0.7]$. These imply $\underline{p}_{4,2}=0.1$ and $\underline{\delta}_{4,2}^+ = 1/7$.

The following result utilizes the proportions $\underline{\delta}_{j,k}^+$ and $\underline{\delta}_{j,k}^-$ to partially identify average outcomes for individuals with $Q_j=(j,k)$ on either side of the cutoff. Taking differences of these bounds yield bounds for the averages of the treatment effects $Y(j)-Y(k)$.

propositionSuppose Assumptions (ref), (ref), and (ref) hold. Consider a pair $(j,k) \in \mathcal{P}$ such that $\underline{p}_{j,k}>0$, then we have the following bounds on $\mathbb{E}[g(Y(j)) | Q_j=(j,k), S_j=c_j ]$ and $\mathbb{E}[g(Y(k)) | Q_j=(j,k), S_j=c_j ]$: \begin{align*} & \mathbb{E}\left [F_{j,k+}^{-1}(U) \left| \boldsymbol{\mathrm{Q}}_j \cap \{ (j,k) \} \neq \emptyset, U < \delta_{j,k}^+, S_j=c_j^+ \right. \right] \\ & \leq \mathbb{E}\left[g(Y(j)) \left| Q_j=(j,k), S_j=c_j \right. \right] \leq \\ & \mathbb{E}\left [F_{j,k+}^{-1}(U) \left| \boldsymbol{\mathrm{Q}}_j \cap \{ (j,k) \} \neq \emptyset, U > 1-\delta_{j,k}^+, S_j=c_j^+ \right. \right], \end{align*} and \begin{align} \begin{split} & \mathbb{E} \left[ F_{j,k-}^{-1}(U) \left| \boldsymbol{\mathrm{Q}}_j \cap \{ (j,k) \} \neq \emptyset, U < \delta_{j,k}^-, S_j=c_j^- \right. \right] \\ & \leq \mathbb{E}\left[ g(Y(k)) \left| Q_j=(j,k), S_j=c_j \right. \right] \leq \\ & \mathbb{E} \left[ F_{j,k-}^{-1}(U) \left| \boldsymbol{\mathrm{Q}}_j \cap \{ (j,k) \} \neq \emptyset, U > 1 - \delta_{j,k}^-, S_j=c_j^- \right. \right], \end{split} \end{align} where $U \sim \text{Uniform}[0,1]$ (independent of everything else), \\ $F_{j,k+}^{-1}(u):=\inf\left\{y: \mathbb{P}\left[ g(Y) \leq y \left| \boldsymbol{\mathrm{Q}}_j \cap \{(j,k)\} \neq \emptyset, S_j=c_j^+ \right. \right] \geq u \right\}$ and \\ $F_{j,k-}^{-1}(u) :=\inf\left\{y: \mathbb{P}\left[ g(Y) \leq y \left| \boldsymbol{\mathrm{Q}}_j \cap \{(j,k)\} \neq \emptyset, S_j=c_j^- \right. \right] \geq u \right\}$.

The expressions for the bounds simplify significantly in the cases where $Y$ is binary or continuous. For the sake of brevity, we relegate the detailed formulas to the Appendix (Section (ref)). The bounds in Proposition (ref) build on the work by horowitz1995. To see the intuition, consider the case where $g(Y)=Y$ is a continuous random variable, and focus on individuals just above the cutoff. Among all individuals in the subpopulation with $\boldsymbol{\mathrm{Q}}_j \cap \{(j,k) \} \neq \emptyset$ and $S_j=c_j^+$, a fraction ${\delta}^+_{j,k}$ of them has $Q_j=(j,k)$ and $Y=Y(j)$ by the cutoff characterization. We do not know who these individuals are among those in the subpopulation. However, the lowest possible value for $\mathbb{E}[ Y(j) | Q_j=(j,k), S_j=c_j ]$ occurs if all such individuals are located at the lower tail of the distribution of outcomes in the subpopulation. Likewise, the highest possible value for $\mathbb{E}[ Y(j) | Q_j=(j,k), S_j=c_j ]$ occurs if all of that same fraction of individuals are located in the upper tail of the distribution of outcomes. We do not know ${\delta}^+_{j,k}$, but we do know that it is no smaller than $\underline{\delta}^+_{j,k}>0$. The bounds only grow wider as the fraction ${\delta}^+_{j,k}$ decreases, so the bounds evaluated at ${\delta}^+_{j,k} = \underline{\delta}^+_{j,k}$ take into account all possible values for ${\delta}^+_{j,k}$.

Although intuitive and analytically simple, these bounds are not necessarily sharp because $\underline{\delta}^+_{j,k}$ is not exogenously given as considered by horowitz1995; $\underline{\delta}^+_{j,k}$ is constructed from the marginal distribution of $Q_j$ but there could be additional identifying information in the joint distribution of $Q_j$ and $Y$. Providing a complete characterization of the sharp bounds is complex because it involves deriving bounds on the joint distribution of potential outcomes and $Q_j$, which may not be practical when the potential outcomes are continuous. For the sake of simplicity, we relegate the sharp characterization to Section (ref) in the appendix. The bounds of Proposition (ref) contain the sharp bounds of Section (ref) as long as our model assumptions are true. The lack of sharpness may not matter in practice when Proposition (ref) yields tight bounds for a given dataset. However, if our assumptions are not true, the bounds of Proposition (ref) may not contain the sharp bounds of Section (ref). That said, we may obtain tight bounds from the data using Proposition (ref), but this does not mean they contain the true parameter kedagni2020. Therefore, it is advisable to assess the testable implications of Corollary (ref) as a matter of routine.

Assignment to College Programs and Graduation in Chile

In this section, we illustrate our method using data from college applications in Chile. Before estimating bounds for the effect of assignment to a program on graduation outcomes, we document the presence of strategic behavior in this setting.

Institutional Setting and Evidence of Strategic Behavior

\paragraph{Centralized college application system in Chile.} We use publicly available data on Chile's centralized college application and assignment system from 2004 to 2010 and on graduation from 2007 to 2020. The institutional setting has been described in detail by hastingsneilsonzimmerman2013 and larroucaurios2020,larroucaurios2021, among others. College choice in Chile is organized as a semicentralized system---a subset of universities participate in a centralized market in which a clearinghouse collects rank-order lists from applicants and determines assignments using a variant of the DA algorithm. Students can submit rank-order lists of up to eight major--university pairs (“programs”) out of more than 1,000.\footnote{The cap was increased to ten in 2012.} Priorities are program-specific and determined by a weighted average of scores obtained in a national standardized test (the PSU, for prueba de selecci\'on universitaria) and of high-school GPA.\footnote{In 2014, students' relative rank within their high school was added as one of the “primary” scores to be averaged to construct priorities.} Descriptive statistics about the sample of students and programs are shown in Tables (ref) and (ref).

\paragraph{Strategic behavior.} larroucaurios2020,larroucaurios2021 thoroughly document that Chilean college applicants behave strategically. Using a 2014 survey linked to administrative data on applications, they show that listed programs often do not coincide with the truly preferred programs as elicited by the survey. Focusing on medicine programs, they show that as application scores decline, students become more likely to exclude medicine from the top of their submitted list—even when ranking it first in the survey—and application rates drop sharply in the 700–750 range, where most medicine cutoffs lie. This pattern indicates that students forgo medicine as their admission prospects weaken, despite preferring it over other programs.

Expectations about cutoffs play a central role in program choice. In Chile, students can readily infer these from historical cutoff data, and many programs exhibit stable cutoffs over time (see, e.g., kapor2024, Figure 1). Because students observe their priority scores before submitting $P$, cutoff stability implies considerable certainty about which programs are feasible. Figure (ref) shows that more selective programs have not only higher cutoffs but also more \ predictable cutoffs over time. Panel (a) plots the histogram of cutoff values pooled across years and then compares that with two subsets of programs: first, with programs that we select to illustrate our methods in Section (ref) below; and second, with medicine programs. Section (ref) focuses on programs with a sufficiently large number of applications because we need adequate sample sizes to implement our estimation procedures. We refer to them as programs of interest. Programs of interest are more selective than most, and Medicine programs stand out as exceptionally selective. We see in Panel (a) that cutoffs generally increase with selectivity and tend to be more compressed at the top of the distribution. In order to assess how predictable program-specific cutoffs are, we estimate an auto-regressive (AR) model for cutoff value at time $t$ for each program as a function of past cutoff values. Our sample has seven years of data which allows us to estimate an AR with three lags after de-meaning the data. Panel (b) displays the distribution of these program-specific $R^2$'s. That distribution is again compared with the programs of interest of Section (ref) and with medicine programs. We find that selectivity is generally associated with predictability of cutoffs. For most programs of interest, the $R^2$ is larger than 0.70, while 70% of medicine programs have an $R^2$ greater than 0.90. The survey data evidence on strategic behavior regarding Medicine is consistent with agents having strong a priori knowledge about ex post cutoffs for Medicine.

figure[figure omitted — 1,222 chars of source]

Figure (ref) provides additional evidence of strategic behavior arising from a priori knowledge of cutoffs. Consider a program $j$ with cutoff $c_j$. Suppose that students tend to prefer programs of higher quality, consistent with what is found in the literature. If applicants behave strategically, one would expect applications to program $j$ to peak among students with application scores close to $c_j$. If cutoffs tend to remain in the same neighborhood across years, students with application scores much higher than $c_j$ can expect to be admissible to more selective, higher-quality programs than $j$, which they prefer over $j$. Hence, we expect very few of these students to include $j$ on their list. As application scores drop and are closer to $c_j$, students' chances of admission to the most selective programs decrease, and program $j$ becomes one of the most selective (desirable) programs among those for which they still have a high admission probability. Hence, we expect applications to $c_j$ to increase as application scores decrease and draw closer to $c_j$. As application scores decrease below $c_j$, students realize that their probability of admission to program $j$ is lower, and while program $j$ remains a relatively desirable (selective) alternative, we expect these expectations to drive applications down. This application pattern, expected if students behave strategically, is exactly what we observe in the top panel of Figure (ref). Pooling all programs $j$ together, Figure (ref) shows the fraction of students listing program $j$ in their rank-order lists (ROLs or $P$ in terms of our notation), as a function of the distance between their priority score for program $j$ and the cutoff $c_j$.

It may be difficult to disentangle the role of preferences from the role of expectations about admission probabilities when both may enter students' choice of which programs to include in their ROLs (manski_ecta2004; agarwal2018). The pattern observed in Figure (ref) could, alternatively, be consistent with students not behaving strategically but preferring programs that are a good fit in terms of quality, that is, programs in which their skill level would be close to the marginal skill level. If this were the case, application behavior would not change discontinuously as students’ scores cross a program’s admission cutoff. Although such discontinuities are not required for strategic behavior to arise, examining specific program pairs—rather than pooling all programs as in Figure (ref)—reveals clear discontinuities (Figure (ref)). The left panels plot, for three program pairs $(j,k)$, the density of the application score $S_j$ among applicants with $P_j=(j,k)$ and show a significant jump at $S_j=c_j$. The right panels display the probability that $P_j=(j,k)$ among all potential applicants as a function of $S_j$, again exhibiting a sharp discontinuity at the cutoff. Importantly, the programs $j$ in these pairs have highly predictable cutoffs: the $R^2$ from three-lag autoregressive models forecasting the cutoff at time $t$ exceeds 0.98 in all three cases. This discontinuous type of strategic behavior may lead to issues with internal validity of an RD identification strategy that controls for $P_j$ (see Section (ref) for discussion).

figure[figure omitted — 643 chars of source]
figure[figure omitted — 1,309 chars of source]

Results

We are interested in identifying the effects of assignment to a given postsecondary program on college graduation. College returns are typically thought of as tied to college graduation, motivating our focus on graduation-related outcomes (see kirkeboen2016 and altonjiarcidiaconomaurel2016 for further references). In addition, the extent to which students eventually graduate from the program to which they are assigned to (or from programs their assigned program is a pathway to) can be viewed as measure of performance of the assignment mechanism. However, for the average college program, initial enrollment and eventual graduation are far from being perfectly correlated (OECD 2019, larroucaurios2021). We also consider graduation from a top university as another outcome of interest. The choice of this second outcome is in line with the literature on returns to education and college choice, which highlights the role of institution quality in driving returns (see kirkeboen2016 and altonjiarcidiaconomaurel2016 for more references). In our setting, the set of “top” universities consists of Pontificia Universidad Cat\'olica de Chile in Santiago (PUC Santiago; hereafter PUC) and Universidad de Chile (UChile).

We present three sets of results. First, we report estimates of average structural functions, followed by estimates of treatment effects, and then discuss the economic implications our results. In these two exercises, to keep things simple, we focus on a single program $j$ and the popular next-best programs $k$ associated with it (or a single program $k$ and the popular local first-best $j$ associated with it). Finally, we further illustrate the importance of our approach by showing that ignoring strategic behavior in identifying treatment effects can lead to misleading conclusions. In this exercise, we estimate bounds on the treatment effects of interest for a large number of pairs $(j,k)$ and present statistics across these pairs. The bounds we present throughout this section represent estimates---not inference bounds. Inference is beyond the scope of this paper.\footnote{Beyond the fact that the admission cutoffs are estimated, the fact that the $\delta$s used in the construction of bounds are partially identified must be taken into account when confidence bands are derived---a task that we leave for future research.} As noted earlier, the identification proposed in this paper resembles a sharp RD design and is therefore well suited for studying the effects of assignment to a program on graduation-related outcomes. In contrast, estimating the effects of graduation from a program on earnings typically involves a fuzzy design, which we study in separate ongoing work.\footnote{At the time of writing of the present paper, we do not have access to earnings data and therefore cannot provide bounds here on the intent-to-treat (ITT) effects on earnings.}

Average Structural Functions and the Importance of Preferences for Graduation Outcomes

\paragraph{Setup.} We focus on several program pairs $(j,k)$ that share the same $k$: the Bachillerato de Ingreso Común at UChile. This program is chosen for two reasons: (1) its large applicant pool allows for precise estimation; and (2) its design and purpose make it intrinsically interesting. As a popular entry point at a selective university, the Bachillerato lets students explore multiple disciplines before choosing a major; it is often seen as an alternative route into competitive programs for those without high enough scores for direct admission. This makes it particularly compelling to examine whether assignment to this program affects students’ likelihood of graduating from their original local first-best $j$. Accordingly, we examine two graduation outcomes: (1) graduating from one’s local first-best program $j$ when assigned to the next-best alternative, the Bachillerato de Ingreso Común at UChile; and (2) graduating from a top university. Specifically, we show bounds for the average structural function $\mathbb{E}\left[ Y(k) \left| Q_j=(j,k), S_j=c_j \right. \right]$, as in Eq. (ref). We restrict attention to five of the most common local first-best programs $j$ ---Medicine at UChile, Medicine at Universidad de Santiago de Chile, Odontology at UChile, Kinesiology at UChile, Math at PUC.

As a simple procedure to construct a sample local to each of the admission cutoffs $c_j$ of interest, we use a 30-point bandwidth on either side of $c_j$.\footnote{The resulting estimation sample is therefore invariant across outcomes once the pair $(j,k)$ is fixed.} Given our choice of bandwidth, the next step of the exercise is to recover local preferences $Q_j$ for each student at each cutoff $c_j$. In the absence of strategic behavior, these preferences can be directly inferred from application lists ---direct observation of $P$ and placement scores allows to compute cutoffs, budget sets, and finally $P_j$, which equals $Q_j$ under truth-telling. However, under strategic behavior, $Q_j$ is not fully revealed by the observed $P_j$. Section (ref) shows how to construct the set of possible $Q_j$s, that is, the set $\boldsymbol{\mathrm{Q}}_j$, under various assumptions. The weak partial order assumption (WPO) implies that any submitted list is a subset of acceptable programs ordered just as in the student's true preferences. The strong partial order assumption (SPO) adds that students who submit shorter-than-maximum lists reveal their full set of acceptable programs. Under SPO, we fully observe $Q_j$ for students who rank strictly fewer than the maximum number of programs (eight, in the case of our empirical application). For students ranking eight choices, $Q_j$ is observed only if the set $\mathbf{Q}_j$ comes out as a singleton. The UMAS assumption (Assumption (ref), hereafter UMAS) can be used in combination with WPO and SPO to refine $\mathbf{Q}_j$ for the non-singleton cases (Corollary (ref)). In a context with a large number of alternatives and a small cap $K$, like in Chile's college application system, the $\boldsymbol{\mathrm{Q}}_j$ constructed under WPO will likely be large. As discussed in Section (ref), additional context-specific assumptions can be combined with UMAS to reduce the size of $\boldsymbol{\mathrm{Q}}_j$. In the applied literature, fields of study are widely recognized as key drivers of student preferences—and of the heterogeneity in preferences across students—for post-secondary programs altonjiarcidiaconomaurel2016. Building on this insight, we propose the following {\it fields} assumption:

assumptionSuppose the set of all programs is partitioned into fields of study. Let $\mathcal{F}$ be a set of programs in an arbitrary field of study in this partition. Consider any cutoff $c_j$. Let $\omega$ be a student with score $S_j$ near $c_j$. If no program from field $\mathcal{F}$ appears in the reported list $P(\omega)$, then no program from $\mathcal{F}$ can be part of the student's true local preference at $c_j$, i.e., \begin{align*} & \forall \mathcal{F}, P(\omega) \cap \mathcal{F} = \emptyset \Rightarrow Q_j(\omega) =(a,b) : a \notin \mathcal{F} and b \notin \mathcal{F}. \end{align*}

In practice, we partition programs in five fields of study: Health/Medicine; STEM; Economics/Business; Law; and Other. So for instance, if a student's reported preferences include programs in Medicine, STEM, and Economics only, Assumption (ref) excludes programs in Law and Other from their $\boldsymbol{\mathrm{Q}}_j$. For any program $\ell \in \mathcal{J}$, let $\mathcal{F}(\ell)$ denote the set of all programs in the field of study of program $\ell$. Researchers may apply Assumption (ref) to refine the set $\boldsymbol{\mathrm{Q}}_j(\omega)$ of student $\omega$ as follows: \[ \boldsymbol{\mathrm{Q}}_j^*(\omega) = \boldsymbol{\mathrm{Q}}_j(\omega) \setminus \left\{(\ell,\ell') \in \mathcal{J} \times \mathcal{J}^0 : P(\omega) \cap \mathcal{F}(\ell)=\emptyset ~~\text{or}~~ P(\omega) \cap \mathcal{F}(\ell')=\emptyset \right\}, \] where we adopt the convention that $\mathcal{F}(0) = \mathcal{J}$.

Figure (ref) reports estimated bounds on the probability of graduating if assigned to Bachillerato de Ingreso Común at UChile across local first-best programs $j$, for students whose local next best is Bachillerato de Ingreso Comun at UChile. The left panel shows the probability of graduation from the local first-best program; the right panel shows the probability of graduation from a top university.

figure[figure omitted — 1,383 chars of source]

\paragraph{On the importance of preferences for graduation outcomes.} For both outcomes, we find evidence of heterogeneous effects across different local first-best programs ---in a number of cases, bounds do not overlap across $j$. For instance, in the left panel, students whose local first-best is Medicine at UChile are more likely to graduate from that program after assignment to the Bachillerato than those whose first-best is Medicine at Universidad de Santiago de Chile or Math at PUC. In the right panel, students with Medicine at UChile as their first-best are more likely to graduate from a top university when assigned to the Bachillerato than those with Math at PUC as their first-best.

The observed heterogeneity may arise from two sources. First, differences in preferences—potentially correlated with effort and ability—can influence graduation outcomes. Second, variation in academic preparation or ability, proxied by test scores, could also play a role as we are comparing subpopulations with $S_j=c_j$ for different $j$'s in Figure (ref). We show that this second channel is unlikely to explain the differences observed in Figure (ref). If admissions across programs $j$ were based on a single priority score, it would suffice to show that the cutoffs $c_j$ are close or that applicant score distributions are similar around them. However, priority scores vary by program, as each $S_j$ is a weighted sum of six primary components with program-specific weights.\footnote{Primary scores: NEM, math, Spanish, science, history, and the max of science and history. Weights can be zero but always sum to 100%.} Therefore, to assess comparability across cutoffs, we analyze the Euclidean distance between students’ six-dimensional primary score vectors. We show that at any of the cutoffs $c_j$ of interest, students within the bandwidth of the cutoff are, in terms of the Euclidean distance between their vectors of primary scores, on average as close to each other as they are to students around the other cutoffs of interest. Results are shown in Table (ref) and discussed in detail in Appendix (ref).

We therefore interpret the results in Figure (ref) as evidence that students' preferences matter for their graduation outcomes, suggesting that preferences are correlated with effort choices or ability that is not captured by test scores.\footnote{Evidence from heckmankautz2012 shows that achievement tests do not fully capture soft skills or personality traits that influence labor market outcomes (see also borghans2008, almlund2011). Structural models support this, showing that factors beyond test scores affect educational choices. For example, arci2005 find that, even conditional on SAT scores, some individuals have higher college admission chances, financial aid likelihood, labor market earnings without college, and returns to all majors.}

\paragraph{Role of behavioral assumptions.} We next examine how different behavioral assumptions affect the estimated bounds. Table (ref) in the appendix provides descriptive statistics and the estimated $\underline{\delta}_{j,k}^+$ and $\underline{\delta}_{j,k}^-$ for each pair $(j,k)$ and each set of assumptions. The different assumptions are used to construct $\boldsymbol{\mathrm{Q}}_j$ for each individual and therefore determine the conditioning set $\{\boldsymbol{\mathrm{Q}}_j \cap \{(j,k)\} \neq \emptyset\}$ and the values of $\underline{\delta}_{j,k}^+$ and $\underline{\delta}_{j,k}^-$. Under WPO, $\boldsymbol{\mathrm{Q}}_j$ is larger than under SPO, resulting in a larger conditioning set (Table (ref), Columns (1) vs. (3) and (2) vs. (4)). Given a fixed bandwidth around the cutoff $c_j$, this means that (i) the denominator in the equations defining $\underline{\delta}_{j,k}^+$ and $\underline{\delta}_{j,k}^-$ is larger under WPO than under SPO; and (ii) the share of individuals for whom $Q_j$ is found to be a singleton is lower under WPO than it is under SPO, which in turn implies that the numerator in the equations defining $\underline{\delta}_{j,k}^+$ and $\underline{\delta}_{j,k}^-$ is smaller under WPO than under SPO. Overall, this leads to $\underline{\delta}_{j,k}^+$ and $\underline{\delta}_{j,k}^-$ being smaller under WPO than under SPO. The increase in $\underline{\delta}_{j,k}^+$ and $\underline{\delta}_{j,k}^-$ as we impose SPO instead of WPO unambiguously tends to reduce the width of the bounds. The exact extent to which the width of the bounds shrinks as we impose SPO instead of WPO also depends on the joint distribution of the outcome $Y$ and $\boldsymbol{\mathrm{Q}}_j$, however. Subsequently imposing UMAS and the {\it fields} assumption eliminate certain elements from individuals' $\boldsymbol{\mathrm{Q}}_j$s and futher increase $\underline{\delta}_{j,k}^+$ and $\underline{\delta}_{j,k}^-$ (Columns (1) vs. (2) and (3) vs. (4)). \footnote{Note that the impact of the {\it fields} assumption depends on the overlap between fields of $j$ and $k$. Under stability, the first coordinate of $P_j$ and the first coordinate of any element of $\boldsymbol{\mathrm{Q}}_j$ coincide for students just above $c_j$, while the second coordinates of $P_j$ and any element of $\boldsymbol{\mathrm{Q}}_j$ coincide for students just below $c_j$ (Proposition (ref)). This means that, absent the {\it fields} assumption, each student in the set $\{\omega: \boldsymbol{\mathrm{Q}}_j \cap \{(j,k)\} \neq \emptyset\}$ includes a program from either $\mathcal{F}(j)$ or $\mathcal{F}(k)$ in her reported preference. As a consequence, the {\it fields} assumption can only reduce the set $\{\omega: \boldsymbol{\mathrm{Q}}_j \cap \{(j,k)\} \neq \emptyset\}$ for pairs $(j,k)$ for which $\mathcal{F}(j)\neq \mathcal{F}(k)$. }

Treatment Effects

We now illustrate our method for estimating treatment effects, as opposed to average structural functions. Specifically, we estimate the effect on graduation outcomes of being assigned to Medicine at PUC Santiago instead of being assigned to a second-best option. We focus on this program for two main reasons. First, its high selectivity and popularity ensure a sufficiently large sample for precise estimation. Second, its competitiveness and highly predictable cutoff make it likely that applicants adjust their submissions strategically around the cutoff—a type of behavior our approach is designed to guard against. We consider the five most common next-best options to Medicine at PUC: Medicine at UChile, at Universidad de Concepción, and at Universidad de Santiago de Chile; and Science, and Engineering at PUC. As in earlier analyses, we construct our estimation sample using a 30-point bandwidth around the cutoff $c_j$.

Figure (ref) presents the estimated bounds. The left panel shows treatment effects on the probability of graduating from Medicine at PUC Santiago; the right panel shows treatment effects on the probability of graduating from a top university. Starting with the left panel, we find, reassuringly, that assignment to Medicine at PUC increases the probability of graduating from that program. In all five cases, bounds derived under any of the four sets of assumptions exclude zero, identifying a positive treatment effect. Importantly, the magnitude of the effect varies across next-best alternatives, as reflected by non-overlapping bounds. For example, when it comes to graduating from Medicine at PUC, students whose second-best is Science at PUC appear to benefit less from admission to Medicine at PUC than those whose next-best is Medicine at UChile or at Universidad de Santiago de Chile.

The right panel reveals similar patterns. Admission to Medicine at PUC increases the probability of graduating from a top university in many cases, and regardless of whether the next-best option is at a top university itself. For instance, the estimated bounds indicate a positive effect both for students whose next-best program is Science at PUC and for those whose next-best is Medicine at Universidad de Concepción or Universidad de Santiago de Chile. There again, the bounds highlight substantial heterogeneity across next-best alternatives.

As we relax assumptions, the bounds widen and become less informative. This exercise thus helps clarify the identifying power of each assumption, providing researchers with a framework to evaluate the empirical content of different preference assumptions in similar settings.

figure[figure omitted — 1,207 chars of source]

On the Importance of Accounting for Strategic Behavior

Finally, we illustrate our bounding approach more broadly using 302 pairs $(j,k)$ of interest. We show that ignoring strategic behavior may be misleading when interest lies in a average treatment effect conditional on true preferences as in (ref). In what follows, we compare our estimated bounds to point estimates produced by an RD that controls for reported preferences, as the current practice advocates. Current-practice estimates are consistent for \[\mu_{j,p} := \mathbb{E} [Y\mid P_j=(j,k),S_j=c_j^+]-\mathbb{E}[Y\mid P_j=(j,k),S_j=c_j^-].\]

Under the assumption that every student is a truth-teller at the cutoff $c_j$, the parameter $\mu_{j,p}$ coincides with $\mathbb{E} [Y(j) - Y(k) \mid Q_j=(j,k),S_j=c_j]$, which is our parameter of interest (ref). Since the truth-telling assumption implies our behavioral assumptions, our bounds contain $\mu_{j,p}$ under truth-telling. On the contrary, under strategic behavior, $\mu_{j,p}$ does not equal $\mathbb{E} [Y(j) - Y(k) \mid Q_j=(j,k),S_j=c_j]$ in general, and our bounds may or may not contain $\mu_{j,p}$. The interpretation of $\mu_{j,p}$ depends on the type of strategic behavior. First, as mentioned earlier, strategic behavior may be of the discontinuous type to the point of breaking the continuity of $\mathbb{E}[Y(d)\mid P_j=(j,k),S_j=s]$ as a function of $s$ at $s=c_j$, for $d=j$ or $d=k$ (Figure (ref) presents evidence consistent with this behavior). In this case, the RD is not internally valid and $\mu_{j,p}$ does not equal an average treatment effect. Alternatively, strategic behavior may not be of the discontinuous type but still make $P_j \neq Q_j$. In this case, $\mu_{j,p}$ equals a weighted average of various average treatment effect parameters, where the weights are functions of the unknown mapping $\omega \mapsto Q_j(\omega)$.

We look at 302 pairs $(j,k)$ which we select as follows. First, we consider all pairs $(j,k)$ such that program $j$ is available on the platform at any point between 2004 and 2010 and at least one student lists $k$ right after $j$ in his ROL. Second, we count the number of students with score $S_j$ within a 30-point bandwidth of the year-specific $c_j$ who $P_j=(j,k)$.\footnote{ For a number of $j$'s of interest, there exists another program $\ell$ that uses the same placement score as $j$ and whose cutoff $c_{\ell}$ is within a 30-point distance of $c_j$. In these cases, we use a bandwidth smaller than 30 around the cutoff $c_j$ of interest, so the sample of $c_j$ does not include observations from the sample of $c_{\ell}$ that are subject to a different treatment change. } We find 327 pairs $(j,k)$ with a least 50 such students. These pairs involve 137 distinct programs $j$, each of them associated with a number of $k$'s ranging from one to 18. We exclude the 25 pairs $(j,k)$ associated with three of these programs $j$ because no graduates are recorded in the data for these programs. This yields a final set of 302 pairs involving 134 distinct $j$'s.

Table (ref) provides statistics on the estimated bounds that we obtain for our two outcomes of interest, both under the SPO and the SPO+UMAS+fields assumptions. First, we see that SPO bounds identify the sign of our average treatment effect parameter for many pairs; the estimated bounds even collapse to point-estimates for a small share of pairs. Adding the UMAS and fields assumptions make our estimated bounds narrower, and they identify the sign of average effects for 70% of pairs when outcome is graduation from $j$; they collapse to point estimates for 23% of pairs when outcome is graduation from top university. Most importantly, we see that current-practice estimates fall outside of bounds in a significant share of cases, 12% & 15% under SPO and 25% & 24% under SPO+UMAS+fields. This evidence is consistent with lack of truth-telling at the cutoff of these pairs. More broadly, this also suggests caution in contexts like the Chilean case, where most but not all students submit $P$ with $|P|<K$ or very few students are assigned to their last-listed option. It may not be safe to simply assume that everyone is a truth-teller just because the constraint does not bind for most people. In other words, current-practice estimates may be severely biased for $\mathbb{E} [Y(j) - Y(k) \mid Q_j=(j,k),S_j=c_j]$ even if the share of constrained students is small.

table[table omitted — 1,092 chars of source]

Conclusion

This paper provides a novel approach to partially identify the effects of mechanism assignment on future outcomes that is robust to strategic behavior. We illustrate our approach using data from Chile, where a DA mechanism assigns about 80,000 students to more than 1,000 university-major programs every year, and for which, in line with previous literature, we find substantial evidence of strategic behavior. In two high-stakes settings—a selective entry program at the University of Chile and Medicine at PUC Santiago—we find heterogeneous effects of assignment on graduation outcomes across students’ preferred programs and next-best alternatives, consistent with preferences being linked to unobserved traits.

We illustrate our bounding approach for RD-like parameters at the granular university–program level, where students make choices and matching occurs. In many settings, however, researchers or policymakers may seek more aggregate effects—such as STEM vs. non-STEM—due to limited data near multiple cutoffs or broader policy relevance. Aggregation of local effects brings additional challenges that we address in separate ongoing work.