EconBase
← Back to paper

Identification and Estimation in Many-to-one Two-sided Matching without Transfers

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

96,152 characters

Identification and Estimation in Many-to-one Two-sided Matching without Transfers





\title{{Identification and Estimation in Many-to-one Two-sided Matching without Transfers\thanks{\scriptsize This paper subsumes Sun's job market paper ``Identification and Estimation of Many-to-One Matching with an Application to the U.S. College Admissions'' (first draft: 2018) and He and Sinha's previously circulated paper ``Identification and Estimation in Many-to-One Two-Sided Matching without Transfers'' (first draft: 2020). We thank Nikhil Agarwal, Cristina Gualdani, Thierry Magnac, Rosa Matzkin, Ismael Mourifie, Shuyang Sheng, and Xun Tang, for their useful comments. We would also like to thank the seminar and conference participants at 2018 California Econometrics Conference, 2018 Midwest Econometrics Group Conference, 2019 Canadian Econometric Study Group Meetings, 2019 Network Econometrics Junior's conference at Northwestern University, 2019 Seattle-Vancouver Econometrics Conference, Auburn University,  Econometric Society World Congress 2020, EEA-ESEM 2019, ICEF Moscow, Jinan University, Johns Hopkins University, National University of Singapore, Peking University, Rice University, Simon Fraser University, Toulouse School of Economics, UCLA, UC Riverside, University of Florida, University of Toronto, UNSW Sydney, and Western University. Thanks also to Al\'ipio Ferreira Cantisani for excellent research assistance. Sinha acknowledges funding from the French National Research Agency (ANR) under the Investments for the Future (Investissements d'Avenir) program, grant ANR-17-EURE-0010.}}}
\vspace{-2mm}
\author{YingHua He\thanks{\scriptsize Email: [email removed], Rice University, Houston, Texas, USA.} \text{ } \text{ } Shruti Sinha\thanks{\scriptsize Email: [email removed], Toulouse School of Economics, University of Toulouse Capitole, Toulouse, France.} \text{ } \text{ } Xiaoting Sun\thanks{\scriptsize Email: [email removed], Simon Fraser University, Vancouver, Canada.}}

\setstretch{1.05}
\date{\today}


\maketitle

\begin{abstract}
In a setting of many-to-one two-sided matching with non-transferable utilities, e.g., college admissions, we study conditions under which preferences of both sides are identified with data on one single market. Regardless of whether the market is centralized or decentralized, assuming that the observed matching is stable, we show nonparametric identification of preferences of both sides under certain  exclusion restrictions. To take our results to the data, we use Monte Carlo simulations to evaluate different estimators, including the ones that are directly constructed from the identification. We find that a parametric Bayesian approach with a Gibbs sampler works well in realistically sized problems. Finally, we illustrate our methodology in decentralized admissions to public and private schools in Chile and conduct a counterfactual analysis of an affirmative action policy.
\vspace{2mm}\\
{\scshape \bf Keywords}: Many-to-one Two-sided Matching, Non-transferable Utility, Nonparametric Identification, College Admissions, School Choice.
\end{abstract}

\setstretch{1.5}
\newpage

\setstretch{1.3}

\section{Introduction}\label{sec:intro}
    In a many-to-one two-sided matching market, agents are categorized into two sides; everyone on one side has preferences over those on the other side; an agent on only one of the two sides can have  multiple match partners from the other side.
Many real-life markets fit this description, for example,
the medical resident match \citep{roth1984evolution,agarwal_empirical_2015} in the US, school admissions in Chile \citep{gazmuri2017school} and Hungary \citep{aue2020happens}, college admissions in the US, and graduate program admissions in France \citep{he2020application}. Such markets often exclude personalized transfers, even though limited monetary exchanges may exist.  Hence, the literature defines it as matching without transfers or matching with non-transferable utility.

While the literature has extensively studied this type of matching theoretically \citep[see, e.g.,][]{roth1992two,azevedo_supply_2016}, its econometrics is less explored. Our paper aims to make a contribution by answering the following questions:  Are the preferences of both sides identified from data on who matches with whom? If so, how can the preferences be estimated?

To fix ideas, we proceed in the language of college admissions.   We derive a set of sufficient conditions under which both student and college preferences are nonparametrically identified. Our results are obtained from a single market in which there are a continuum of students and a fixed number of colleges. We use that to approximate a single large market. Further, we provide an estimation procedure that is practical even in settings with many agents, allowing for rich observed and unobserved heterogeneity.  Understanding agent preferences is often crucial for policymaking, and one may analyze a wide range of counterfactual policies with estimated preferences. Potentially, our results open a new avenue of research on such matching markets.

The main challenge in identifying student preferences is that each student's {\it actual} choice set is unobservable to the researcher.  For student~$i$ to be able to enroll at college~$c$, college~$c$ needs to accept $i$. The same difficulty exists in the identification of college preferences. Moreover, each student's and each college's choice sets are endogenously determined in equilibrium without  market-clearing prices.

In our continuum setting, we assume that an observed matching is stable.
That is, no college prefers to reject any of its currently matched students to vacate a seat, and no student prefers to leave her current match to become unmatched or matched with a college that is willing to accept her and, if necessary, reject one of its currently matched students.  Stability is often imposed in the study of various matching markets \citep[see, for a survey,][]{chiappori_econometrics_2016} and is satisfied in equilibrium in our setting in certain game-theoretical models \citep{ACH2017,fack_beyond_2019}.

Importantly, there is generically a unique stable matching that is characterized by the colleges' admission cutoffs  \citep{azevedo_supply_2016}.  When college preferences over individual students are represented by utility functions, a college's cutoff is the lowest utility level among its matched students.  Cutoffs further define a student's actual choice set in equilibrium, called {\it feasible set}. A college is in a student's feasible set if the college's utility of being matched with her is higher than its cutoff.  Stability implies that a student is matched with her most-preferred feasible college, similar to a discrete choice problem, except that feasible sets are unobservable and heterogeneous.

A simple equation, called the $i$-$c$ match probability, is the key to understanding our identification result. Specifically, the conditional probability of student~$i$ being matched with college~$c$ is the sum of conditional probabilities of $i$ choosing $c$ from a given feasible set $L$ weighted by the conditional probability of facing $L$:
\vskip-0.8cm
\begin{footnotesize}\begin{align*}
		~& \mathbb{{P}}(\text{student } i \text{ is matched with college } c \mid x_{i}) \\
	= \quad & \sum_{\text{all possible feasible sets, } L}
	\underbrace{{\mathbb{{P}}(L\,\text{{is}}\, i \text{{'s feasible}}\,\text{{set}} \mid x_{i})}}_{\equiv A\; \text{(college preferences)}}\cdot\underbrace{{\mathbb{{P}}( c\,\text{{is}}\,i \text{{'s most-preferred college}}\,\text{{in}}\,L \mid L,x_{i})}}_{\equiv B\; \text{(student preferences)}},
\end{align*}\end{footnotesize}
where $x_{i}$ consists of all observed characteristics of student $i$ (e.g., pair-specific characteristics like distance to colleges). The equation provides a decomposition of the preferences of the two sides: for each given $L$, piece $A$ only depends on the preferences of {\it all} colleges given cutoffs, while $B$ only depends on $i$'s preferences over all colleges.

We then detail a set of exclusion restrictions, among other regularity conditions, such that the excluded variables act as ``demand shifters'' and ``feasible-set shifters'' (or supply shifters). Sufficient variation in these excluded variables identifies the preferences of colleges and students using the $i$-$c$ match probability described above.

Here are some intuitions. For a college $d$, an $(i,d)$-specific demand shifter traces out how $i$'s preferences for $d$ affects the $i$-$c$ match probability. Similarly, an $(i,d)$-specific feasible-set shifter traces out how $d$'s preference for $i$ affects the $i$-$c$ match probability. A non-excluded variable affects the $i$-$c$ match probability through preferences on both sides for all colleges. By taking derivatives of the $i$-$c$ match probability with respect to (w.r.t.) all the variables, excluded and non-excluded, we derive systems of linear equations that link the effects of variations in demand and supply. Hence, the identification problem reduces to setting up  systems of linear equations and ensuring the existence of a unique solution.

The second objective of our paper is to provide practical methods that can be used to analyze real-life markets. We achieve this by deriving theoretical guidelines and showcasing a practical estimation method.

When taken to the data,  the requirement of a large number of excluded variables may be difficult to meet. To address this, we theoretically characterize the tradeoff between exclusion restrictions and the degree of identifiable preference heterogeneity (\cref{deg-of-heter}). The researcher can use this result as a guideline for empirical studies when having insufficient excluded variables.

We also need a practical estimation method to take these identification results to the data. In fact, our identification arguments are constructive, leading to nonparametric and semiparametric estimators. Monte Carlo simulations suggest that estimating the matrices of partial derivatives in the linear systems using the average derivative estimators of \cite{powell_semiparametric_1989} performs well in finite samples only when the curse of dimensionality is not severe.  In a reasonably sized problem, we resort to a parametric Bayesian approach with a Gibbs sampler \citep{rossi2012bayesian}, resembling applications such as \cite{logan_two-sided_2008} for one-to-one two-sided matching and \cite{abdulkadiroglu_welfare_2017} for a one-sided problem. We demonstrate its good performance in Monte Carlo simulations with high dimensionality.

As an empirical application, we consider the decentralized admissions to secondary schools in Chile. To the best of our knowledge, this is {one of the first attempts to estimate the preferences of both sides in a {\it decentralized} market of many-to-one two-sided matching without transfers.} There is no clearinghouse, and students do not submit rank-order lists of schools. By allowing flexible preference heterogeneity in the Bayesian approach, we estimate the preferences of students and schools. We also consider a counterfactual policy in which students from low-income families are prioritized for admissions to all schools.  Segregation in terms of ability and income decreases, albeit slightly. We find that simply giving low-income students access to schools may not significantly change matching outcomes due to student preferences.


\paragraph{Related Literature.} This paper is related to the literature on the identification of matching models; \cref{tab:sumlit} provides an incomplete summary.

\begin{table}[h!]
\caption{Identification Results of Matching Models\label{tab:sumlit}}
\centering\footnotesize
\resizebox{\textwidth}{!}{
\begin{tabular}{p{0.15\textwidth}p{0.47\textwidth}p{0.47\textwidth}}
\toprule
 & Transferable Utility (TU) & Non-Transferable Utility (NTU) \tabularnewline
\midrule
One-to-one & The match surplus is identified \citep[see, e.g.,][]{choo_who_2006,fox_identification_2010,chiappori2017partner,galichon_cupids_2020}. & The match surplus is identified \citep[see, e.g.,][]{dagsvik_aggregation_2000,menzel_large_2015}. \\
[0.2cm]
Many-to-one   &  The utility function and the
distribution of the unobservables of both sides are identified in a homogeneous setting \citep[see, e.g.,][]{diamond_latent_2017}. &  The utility function and the
distribution of the unobservables of both sides are identified in a homogeneous setting \citep[see, e.g.,][]{diamond_latent_2017}. \\
 &   &  {\it Our paper}:  The utility functions of both sides
with heterogeneity and the distribution of the unobservables are identified.\\
[0.2cm]
Many-to-many &  The match surplus  and/or the distribution of the unobservables are identified \citep[see, e.g.,][]{fox_identification_2010,fox_unobserved_2018}. & The match surplus  is identified \citep[see, e.g.,][]{menzel_strategic_2017}.\\
\bottomrule
\end{tabular}}
\end{table}

This literature is split into several strands depending on the preference structures of the agents -- transferable utility (TU) and non-transferable utilities (NTU);\footnote{See \cite{galichon2019costly} for an example of imperfectly transferable utility (ITU) models.} and the maximum number of links an agent is permitted to form across sides -- one-to-one, many-to-one, and many-to-many \citep[see][for a survey]{chiappori_econometrics_2016}.

There is also a close relationship between the one-to-one TU matching model \citep{choo_who_2006,fox_identification_2010,fox_estimating_2018,GRAHAM2011965,sinha_identification_2015,chiappori2017partner,diamond_latent_2017,galichon_cupids_2020,gualdani_identification_2020} and the many-to-one NTU matching model considered here. Market-clearing college cutoffs in our setting play the role of market-clearing shadow prices, although the endogenous cutoffs do not determine how the surplus is split among the agents.


Most of the work on identification within the NTU framework focuses on one-to-one markets \citep{dagsvik_aggregation_2000,menzel_large_2015,uetake2020entry}. Allowing for infinitely many agents on {\it both} sides of a many-to-one matching market, {\cite{agarwal_empirical_2015}} and \cite{diamond_latent_2017} prove identification under a homogeneity restriction on the preferences,\footnote{When studying the medical resident match, \cite{agarwal_empirical_2015} discusses some intuitions of using exclusion restrictions to identify heterogeneous preferences on each side of the market.}  and \cite{ederer2022} shows identification by relaxing the homogeneity assumption but still restricting preferences. \cite{as_2020} provide a recent survey on empirical models of NTU matching.

Many-to-one NTU matching has been empirically studied in the context of secondary school admissions in Hungary \citep{aue2020happens} and graduate program admissions in France \citep{he2020application}. Their data include information on the preferences of both sides that is reported to a centralized mechanism. Therefore, they can independently identify and estimate the preferences of each side, essentially reducing the two-sided matching to two separate one-sided problems.

Centralized many-to-one NTU matching in the context of school choice has been studied extensively, both theoretically since \cite{abdulkadiroglu_school_2003} and empirically \citep[e.g.,][]{abdulkadiroglu_welfare_2017,agarwal_demand_2018,calsamiglia2020structural,fack_beyond_2019,he_gaming_2017,kapor2020heterogeneous}.
In this literature, school preferences are (assumed to be) known because schools rank students according to certain pre-specified rules. The problem then reduces to identifying and estimating student preferences. See  \cite{agarwal2020revealed} for a survey.


Feasible sets in our setting resemble endogenous consideration sets that arise in one-sided decision problems. In this sense, our paper relates to the growing strand of literature that studies the econometrics of decision problems under consideration set formation \citep[e.g.,][]{abaluck_what_2018, barseghyan_heterogeneous_2019, barseghyan2021discrete,cattaneo2020random}. Our contribution here is that we provide a structural two-sided setting where the consideration probabilities in a student's decision problem are entirely determined by the college (supply side) preferences. Along similar lines to ours, \cite{agarwal2022demand} study consumer choice models with latent choice-set constraints. Their identification conditions and ours are non-nested (see Section~\ref{id}).


The remaining paper is organized as follows: \cref{model} describes
the model and data generating process; \cref{id} discusses the identification of preferences of the agents on both sides of the market; \cref{sec:empirical} illustrates an empirical analysis of the match between students and secondary schools in Chile; and \cref{sec:conclusion} concludes.

\section{Model}\label{model}
    For the sake of exposition, our model is set up as a college admissions problem. Consider a single market with a continuum of students and finitely many colleges.  The set of all students is  $\textbf{I}$, with a probability measure $Q$ defined over it,\footnote{Our probability space is $(\mathbf{I},\mathcal{B}(\mathbf{I}),Q)$ where $\mathcal{B}(\mathbf{I})$ is the Borel set of $\mathbf{I}$, $Q:\mathcal{B}(\mathbf{I})\rightarrow[0,1]$, $Q(\mathbf{I})=1$.} and
the set of all colleges is $\mathbf{C}=\{1,2,\dots,C\}$. College $c \in \mathbf{C}$ has a capacity $q_{c}\in (0,1)$.

For $i\in\mathbf{I}$, the utility of being matched with college $c$ is $u_{ic}$; for $c\in\mathbf{C}$,  the utility of being matched with $i$ is $v_{ci}$. To prepare for  our identification results in \cref{id}, we assume additive separability and excluded variables in the utility functions,\footnote{\label{fn:ns_id}See Online Appendix \ref{app:ns_id} for our identification results in a more general nonseparable utility specification.} although the rest of the current section applies to more general models.

The utilities $u_{ic}$ and $v_{ci}$ depend on the vector $(z_{i},y_{i},w_{i})$, which is observable to the researcher, and $\epsilon_{i} =(\epsilon_{i1}\dots,\epsilon_{iC})$ and $\eta_{i} =(\eta_{1i}\dots,\eta_{Ci})$, which are unobservable to the researcher.  Specifically, $z_{i}\in\mathcal{Z}\subseteq\mathbb{R}^{d_{z}}$, $y_{i}=(y_{i1}\dots,y_{iC})\in\mathcal{{Y}}= \mathcal{Y}_1 \times \dots \times \mathcal{Y}_C \subseteq\mathbb{{R}}^{C}$, and $w_{i}=(w_{1i}\dots,w_{Ci})\in\mathcal{{W}}\subseteq\mathbb{{R}}^{C}$. Further, for any $c\in\mathbf{C}$, $\epsilon_{ic}$ and $\eta_{ci}$ are scalar random variables. $(\epsilon_{i},\eta_{i})$ are independent and identically distributed (i.i.d.) draws from a joint distribution $F$. We make no restrictions on this joint distribution, allowing for arbitrary correlations within $(\epsilon_{i},\eta_{i})$.\par

\begin{assm}\label{assm_struc}
Let $u^{c}: \mathcal{Z}\rightarrow\mathbb{R}$, $r^c : \mathcal{Y}_c\rightarrow\mathbb{R}$, and $v^{c}: \mathcal{Z}\rightarrow\mathbb{R}$ be nonparametric functions, such that
\begin{align}\label{NP_specification}
u_{ic}  =  \tau_{ic}+\epsilon_{ic} & \mbox{ and }
v_{ci}  =  \iota_{ci}+\eta_{ci}, \forall c\in\mathbf{C},
\end{align}
where $\tau_{ic} = u^{c}(z_{i})+r^{c}\left(y_{ic}\right)$  and $\iota_{ci} =  v^{c}(z_{i})+w_{ci}$.
\end{assm}

Scalar $y_{ic}$ is a demand shifter and scalar $w_{ci}$ is a supply shifter: $y_{ic}$ enters only $u_{ic}$ and is excluded from all other utility functions, and $w_{ci}$ enters only $v_{ci}$.

We impose scale normalization on each side. For students, there exists a known value $\overline{y}_{c}$ in the interior of its support such that $\frac{\partial r^{c}\left(\overline{y}_{c}\right)}{\partial y_{ic}}=1$.\footnote{If $y_{ic}$ is known to have a negative effect on student preferences, the partial is normalized to $-1$. The same applies to $w_i$. } This holds trivially if $r^{c}\left(y_{ic}\right)=y_{ic}$.
{For colleges, we assume $w_{ci}$ enters $v_{ci}$ linearly with a coefficient normalized to one. As detailed below, this linearity assumption allows us to  vary $w_i$ to construct a sufficient number of equations without increasing the number of unknowns.\footnote{\label{fn:nonlinear} The linearity in $w_{ci}$ can be relaxed. We can allow $w_{ci}$ to enter $v_{ci}$ through a nonlinear function, which is known for some colleges. Namely, assume $v_{ci}  = v^{c}(z_{i})+s^c(w_{ci})+\eta_{ci}$, where $s^c$: $\mathcal{W}_{c}\to\mathbb{R}$ is a nonparametric function such that (i) for each $c\in \overline{\mathbf{C}} \subset \mathbf{C}$, $s^c$ is \textit{unknown} with $\frac{\partial s^{c}\left(\overline{w}_{c}\right)}{\partial w_{ci}}=1$ for some known value $\overline{w}_{c}$; and (ii) for each $c\in \mathbf{C} \setminus \overline{\mathbf{C}}$, $s^c$ is \textit{known}. In this specification, our identification relies on varying $w_{ci}$ for $c\in \mathbf{C} \setminus \overline{\mathbf{C}}$ instead of all $c\in \mathbf{C}$.}  The assumptions on $y_i$ and $w_i$ can be switched, that is, having nonlinearity in $w_i$ and linearity in $y_i$.}

Students can remain unmatched or, equivalently, be matched with an outside option denoted by ``$0$.'' The utility of the outside option is normalized to 0, $u_{i0}=0$ $\forall i\in\mathbf{I}$.
We assume that colleges have responsive preferences.\footnote{For $\varepsilon>0$, let $N_\varepsilon (i)$ and $N_\varepsilon (i')$ be a neighborhood of students around $v_{ci}$ and $v_{ci'}$, respectively, such that $Q(N_\varepsilon (i)) = Q(N_\varepsilon (i'))$. Responsive preferences imply that, for any $\mathbf{I}^c \subset \mathbf{I}$ with $Q(\mathbf{I}^c) \leq q_c - Q(N_\varepsilon (i))$, $N_\varepsilon (i) \subset  \mathbf{I}\setminus \mathbf{I}^c$, and $N_\varepsilon (i') \subset  \mathbf{I}\setminus \mathbf{I}^c$, college~$c$ prefers $\mathbf{I}^c \cup N_\varepsilon (i) $ to  $\mathbf{I}^c \cup N_\varepsilon (i') $ if and only if $v_{ci} > v_{ci'}$. See \cite{roth1992two} for a definition in a case with discrete students.}  This implies that the total utility of a college from being matched with a subset of students (up to its capacity) is increasing in its utility from each student; for example, the total utility is the sum of the utility from each of its matched students. College~$c$ has an acceptability threshold, $T_c$, and finds student~$i$ unacceptable if $v_{ci}<T_c$.




With the data from one such continuum market on $\{(z_i,y_i,w_i)\}_{i}$, $\{q_c\}_{c}$, and who matches with whom, we aim to identify student and college preferences by identifying $\{u^{c}, r^c, v^{c}, T_c\}_{c}$ and $F$, although, as we shall see, $\{T_c\}_{c}$ are not always point identified. Note that $(z_i,y_i,w_i)$ does not include college-specific variables that are constant across students, as these will be absorbed by the college-specific utility functions, $(u^c,r^c,v^c)$.

We use the continuum market to approximate a data generating process in a large finite market as follows: $(z_i, y_i, w_i,\epsilon_{i},\eta_i)$ is an i.i.d.\ draw from its joint distribution, college~$c$'s capacity is a $q_c$-fraction of the total number of students, and $\{u_{ic}, v_{ci},T_c\}_{i,c}$ determines both sides' preferences. This approximation is close to the matching outcomes when we use the equilibrium concept that will be introduced in \cref{sec:stable}.\footnote{\label{fn:approx}In a typical identification argument, the number of observations is taken to infinity while  the ``game'' is kept constant. Our setting contains one single matching game that changes with the market size. Proposition~3 of \cite{azevedo_supply_2016}, Proposition~4 of \cite{fack_beyond_2019}, and Corollary 2 of \cite{ACH2023} imply that, under certain conditions, the equilibrium outcome in the continuum approximates well an equilibrium outcome in a large finite market. Such an approximation is also used in the network literature \citep[e.g.,][]{menzel2022strategic}.}


 \begin{remark}\label{rm:location}
The location normalization in the functions is worth highlighting because the joint distribution of $(\epsilon_i,\eta_i)$, $F$, is fully nonparametric.
	In student preferences, we already impose the normalization, $u_{i0}=0$, but we need another location normalization for each $c$ on either $u^{c}(z_{i})+r^{c}\left(y_{ic}\right)$ or $\epsilon_{ic}$ to separately identify the two. 	Similarly, for each college~$c$'s preferences, we need to location-normalize two of the three model primitives  $(v^c(z_i), \eta_{ci}, T_c)$ to pin down the third.

In \cref{sec:derivative}, we identify the derivatives of ($u^c, r^c, v^c$), so this additional location normalization is not needed. However, it is necessary for identifying $T_c$ and $F$  in \cref{sec:id_dist}, where we location-normalize $u^{c}(z_{i})+r^{c}\left(y_{ic}\right)$, $v^c(z_i)$, and $\eta_{ci}$.
\end{remark}

\subsection{Matching and Stable Matching}\label{sec:stable}

We define a matching function or, simply, a matching, $\mu:\mathbf{{I}}\to\mathbf{{C}}\cup\{0\}$, such that (i) $\mu(i) = c \iff i \in\mu^{-1}(c)$, and (ii) $\forall\,c\in\mathbf{{C}},\,\mu^{-1}(c)\subseteq{\bf {I}},$ where $0 \leq Q(\mu^{-1}(c))\leq q_{c}$.

The following concepts are important for our analysis: individual rationality, blocking pairs, and stability. For notational reasons, we define them in the case with discrete students, corresponding to our empirical application. The precise definitions for a model with a continuum of students can be found in \cite{azevedo_supply_2016}, with measure-zero sets of students appropriately dealt with.

A matching $\mu$ is {\it individually rational} if $u_{i\mu(i)} \geq u_{i0}$ and $v_{\mu(i)i} \geq T_{\mu(i)}$ for all $i \in \mathbf{I}$. A student-college pair $(i,c)\in\mathbf{{I}}\times\mathbf{C}$ \textit{blocks}
a matching $\mu$ if (i) student $i$ strictly prefers college $c$ to her current match $\mu(i)$, $u_{ic}>u_{i\mu(i)}$; and (ii) either college $c$ has excess capacity, $ Q(\mu^{-1}(c))<q_{c}$, or college $c$ prefers $i$ to one of its matched students, $\exists\,i^{\prime}\in\mu^{-1}(c),\,\text{{s.t.}},\,v_{ci}>v_{ci^{\prime}}$. Finally, a matching is \textit{stable}
if it is individually rational and not blocked by any pair $(i,c)\in\mathbf{{I}}\times\mathbf{{C}}$.

We assume that the matching in the data is stable.\footnote{\label{fn:stable}Stability can be achieved in certain equilibrium if students apply to all acceptable colleges and if a stable mechanism, e.g.,  the deferred acceptance \citep{gale_college_1962}, is used to find the matching. Theoretically, provided that students know what criteria colleges use to rank them, stability can still be satisfied in equilibrium, even if students choose not to apply to all acceptable colleges due to application costs \citep{fack_beyond_2019} or if students make certain application mistakes \citep{ACH2017}.  Importantly, achieving stability does not require the market to be centralized, as shown in laboratory experiments \cite[see, e.g.,][]{Pais-decentralized}, and steps in mechanisms such as the deferred acceptance can be implemented in a decentralized fashion \citep{GHK}. }  A stable matching exists and is generically unique \citep{azevedo_supply_2016}.\footnote{We need the regularity condition that the set $\{i\in\mathbf{I}:\mu(i)\,\text{is strictly less preferred than}\,c\}$  is open $\forall\, c\in \mathbf{C}$. This condition implies that a stable matching always allows an extra measure zero set of students into a college when this can be done without compromising stability.}
Moreover, a stable matching is characterized by college cutoffs. College~$c$'s cutoff is determined by its least-preferred matched student when its capacity constraint is binding; otherwise, it coincides with the  acceptability threshold. Let $\delta_{c}$ be college $c$'s cutoff. Then, $\forall\,c\in\mathbf{{C}}$,
\begin{equation}\label{cutoff}
 \delta_{c}=\inf_{j\in\mu^{-1}(c)}v_{cj} \text{ \ if \ } Q(\mu^{-1}(c))=q_c; \; \delta_{c}=T_c  \text{ \ if \ } Q(\mu^{-1}(c))<q_c.
\end{equation}
By definition, $\delta_{c} \geq T_c$. Under the assumption of responsive college preferences, to determine if student~$i$ can be accepted by $c$, we just need to compare $v_{ci}$ and $\delta_{c}$. With non-responsive preferences, how $c$ ranks $i$ and $j$ would depend on who else $c$ accepts, and $\delta_{c}$ alone would not be sufficient to determine if $i$ could have been accepted by $c$.

In a stable matching of the continuum market, $\{u^{c}, r^{c},v^{c},T_c\}_{c\in\mathbf{C}}$, and $F$ imply a unique vector of cutoffs, $\{\delta_c\}_c$. Therefore, $\{\delta_c\}_c$ is merely a shorthand notation for the expression in equation~\eqref{cutoff} rather than additional parameters.\footnote{\label{fn:approx_cutoff}As mentioned earlier, especially, in footnote~\ref{fn:approx}, the continuum approximates a large finite market. The literature cited therein shows that equilibrium cutoffs, hence matching outcomes, in the large market can be close to $\{\delta_c\}_c$ with agents being practically ``cutoff-takers.'' That is, each agent's realized preferences have a negligible effect on cutoffs in large markets.}


\subsection{Two-sided Discrete Choice Problem in a Stable Matching}\label{sec:notation}
College~$c$ is said to be \textit{feasible} for student~$i$ if and only if $v_{ci} \geq \delta_c$.
A student can ``choose'' to match with any of her feasible colleges, but not any infeasible college. We call the set of all feasible colleges of a student her {\it feasible set}.
Let $\mathcal{{L}}$ be the collection of the $2^{C}$ possible feasible sets, $\mathcal{{L}}\equiv\{L:\,0\,\in L,\,L\setminus\{0\}\subseteq\mathbf{{C}}\}$. By construction, the outside option always belongs to every feasible set.  A matching is stable if and only if every student is matched with her most-preferred feasible college \citep{fack_beyond_2019}. A classic issue in two-sided matching is that students' feasible sets are determined endogenously, unobserved by the researcher, and heterogeneous across students.

For any given matching $\mu$, the probability that $L\in\mathcal{{L}}$ is student $i$'s feasible set conditional on $(z_i,w_i)$ is
\begin{align}\label{lambda}
	\mathbb{{P}}\left(\text{{feasible set is }} L | z_i, w_i\right) &=\mathbb{P}\left(v_{ci}\geq\delta_{c}\;\forall\,c\in L;\,v_{di}<\delta_{d}\;\forall\,d\notin L | z_i, w_i\right)\nonumber\\
&\equiv \lambda_{L}(\iota_{1i},\dots,\iota_{Ci}).
\end{align}
If each student's feasible set was observed,  $ \lambda_L$ could be identified from the data, and recovering student and college preferences would follow from standard arguments in the discrete choice literature. However, we do not observe the feasible sets.

Similarly, for students, $\mathbb{{P}}\left(c=\arg\max_{d\in L}u_{id}|L, z_{i}, y_{i}\right)$ is the probability that utility-maximizing  students with observables $(z_{i}, y_{i})$ ``choose'' $c$ from feasible set $L$. Conditional on $L$, this probability only depends on student preferences. We define
\begin{equation}\label{g}
 g_{c,L}(\tau_{ic};\tau_{id},d\ne c) \equiv \mathbb{{P}}\left(c=\arg\max_{d\in L}u_{id}|L, z_{i}, y_{i}\right) , \forall c\in\mathbf{C},
\end{equation}
where the first argument of $g_{c,L}$ is always $\tau_{ic}$. If $d\notin L$, $g_{c,L}$ does not vary with $\tau_{id}$.






\section{Nonparametric Identification}\label{id}

We now turn to nonparametrically identifying the distribution of student and college preferences given a stable matching $\mu$ and covariates $(z_{i},y_{i},w_{i})$ in {\it one} market.

A matching is stable if and only if every student is matched with the most-preferred college in her feasible set. Thus, stability implies, for $c\in\mathbf{{C}}\cup\{0\}$,
\begin{align}
\sigma_{c}(z_{i},y_{i},w_{i}) & \equiv  \mathbb{{P}}(\mu(i)=c|z_{i},y_{i},w_{i}) \notag  \\
&=\sum_{L\in\mathcal{{L}}}\mathbb{{P}}(\text{{feasible set is }} L |z_i, w_i) \cdot \mathbb{{P}}\left(c=\arg\max_{d\in L}u_{id}|L,z_{i}, y_{i}\right)\nonumber\\
&= \sum_{L\in\mathcal{{L}}}\lambda_{L}(\iota_{1i},\dots,\iota_{Ci})\cdot g_{c,L}(\tau_{ic};\tau_{id},d\ne c).\label{ideq2}
\end{align}
Hence, the conditional match probability, $\sigma_{c}(z_{i},y_{i},w_{i})$, which is the fraction of students with $(z_{i},y_{i},w_{i})$ matched with $c$ and known in the population data, is linked to the model through student and college preferences.

Below, we first study the conditions under which the functions $\{u^{c},\,r^{c},\,v^{c}\}_c$ are nonparametrically identified; we then identify the joint distribution of $(\epsilon_{i},\eta_{i})$, $F$. Later in Section~\ref{sec:practical}, we present results that impose fewer requirements on the data.

\subsection{Identifying the Derivatives of the Utility Functions \label{sec:derivative}}

We nonparametrically identify the derivatives of the functions $\{u^{c},r^{c}, v^{c}\}_c$ w.r.t.\ the
observables. With these derivatives identified, the functions are identified up to a constant, provided that $z_{i}$ and $y_i$ have full support. The idea is to use the variation in the excluded variables to trace out how each argument in equation~\eqref{ideq2} affects the conditional match probabilities. Specifically, the excluded variables in student preferences $(y_i)$ only shift demand, while the excluded variables in college preferences $(w_i)$ shift supply, or feasible sets. The effect of other variables $(z_i)$ that enter both demand and supply can be written as a combination of the effects of $y_i$ and $w_i$. This leads to  a  system of linear equations in the derivatives of the conditional match probabilities, whose solution  is our parameters of interest.

\subsubsection{\label{id_eg}A Simple Example with One College}
We describe the intuition for identification in a one-college example, $\mathbf{C}=\{1\}$. Student utility functions are $u_{i1}=u^1\left(z_{i}\right)+r^1\left(y_{i1}\right)+\epsilon_{i1}$ for college~$1$ and $u_{i0}=0$ for the outside option. Here, $z_{i}$ is a scalar and $\frac{\partial r^{1}\left(\overline{y}_{1}\right)}{\partial y_1}=1$ for a known value, $\overline{y}_{1}$.
College~$1$'s utility function is $v_{1i}=v^1(z_{i})+w_{1i}+\eta_{1i}$, and the (unobserved) cutoff is $\delta_1$.

To identify $\frac{\partial u^1}{\partial z_i}$ and $\frac{\partial v^1}{\partial z_i}$, we fix $y_{i1}=\overline{y}_1$ and consider any value $(z,w_1)$ in the interior of its support. Figure~\ref{partition}(a) shows that the space of $(\epsilon_{i1},\eta_{1i})$
is partitioned into four parts based on the feasibility of college~1 and
student~$i$'s preferences (i.e., the acceptability of college~$1$ to $i$). Moreover, $\mu\left(i\right)=1$ if and only if $\epsilon_{i1}>-u^1\left(z\right)-r^1\left(\overline{y}_1\right)$ (college~$1$ is acceptable to $i$) and $\eta_{1i}>\delta_1-v^1\left(z\right)-w_1$ (college $1$ is feasible to $i$).

\pgfplotsset{ticks=none}
\begin{figure}[h!]
	\centering
	\begin{minipage}[t]{.45\linewidth}
		\centering

		\begin{tikzpicture}[scale=0.65]

			\tikzset{every pin edge/.style={draw=black,<-, very near end}}



			\begin{axis}
				[
				ymin=-5,
				xmin=-5,
				ymax=5,
				xmax=5,
				width=10cm,
				axis on top=true,
				axis x line=middle,
				axis y line=middle,
				yticklabels={,,},
				xticklabels={,,}
				]
				\draw[fill=yellow!60,opacity=0.8, draw=none]      (axis cs:0.5,0.5) -- (axis cs:0.5,5) -- (axis cs:5,5) -- (axis cs:5,0.5) -- cycle ;
				\draw[fill=blue!50,opacity=0.8, draw=none]      (axis cs:0.5,0.5) -- (axis cs:5,0.5) -- (axis cs:5,-5) -- (axis cs:-5,-5) --  (axis cs:-5,5) -- (axis cs:0.5,5) --cycle ;


				\addplot+ [mark=none, black, line width = 1, domain=-5:5] {0.5};
				\addplot +[mark=none, black, line width = 1] coordinates {(0.5,-5) (0.5, 5)};



				\node [below, label={[align=left] \scriptsize College 1  feasible \\ \scriptsize College 1  acceptable \\ \footnotesize $\mu(i)=1$}] at (axis cs: 3.3,2.5){};
\node [below, label={[align=left] \scriptsize College 1  feasible \\ \scriptsize College 1 unacceptable \\ \footnotesize $\mu(i)=0$}] at (axis cs: -3,2.5){};
\node [below, label={[align=left] \scriptsize College 1  infeasible \\ \scriptsize College 1 unacceptable \\ \footnotesize $\mu(i)=0$}] at (axis cs: -3,-4){};
\node [below, label={[align=left] \scriptsize College 1  infeasible \\ \scriptsize College 1  acceptable \\ \footnotesize $\mu(i)=0$}] at (axis cs: 3.3,-4){};


				\node [below] at (axis cs: 4.8,0) {\footnotesize $\epsilon_{i1}$};
				\node [left] at (axis cs: 0,4.8) {\footnotesize $\eta_{1i}$};


            \node [below right] at (axis cs: 0.5,0) {\scriptsize $-u^1(z)-r^1(\overline{y}_1)$};
            \node [above left] at (axis cs: 0,0.5) {\scriptsize $\delta_1-v^1(z)-w_1$};
            \addplot[mark=*, mark size=1.2pt] coordinates {(0,0.5)};
            \addplot[mark=*, mark size=1.2pt] coordinates {(0.5,0)} ;
            \node[label={230:{\scriptsize 0}},circle,fill,inner sep=1pt] at (axis cs:0,0) {};



			\end{axis}

		\end{tikzpicture}
		\subcaption[]
		{{\footnotesize Feasible sets, student preferences, \& $\mu(i)$} }
		\label{fig:panela}


		\begin{tikzpicture} [scale=0.65]

			\tikzset{every pin edge/.style={draw=black, <-,very near end}}

		\begin{axis}
			[
			ymin=-5,
			xmin=-5,
			ymax=5,
			xmax=5,
			width=10cm,
			axis on top=true,
			axis x line=middle,
			axis y line=middle,
			yticklabels={,,},
			xticklabels={,,}
			]
			\draw[fill=yellow!60,opacity=0.8, draw=none]      (axis cs:1.5,0.5) -- (axis cs:1.5,5) -- (axis cs:5,5) -- (axis cs:5,0.5) -- cycle ;
			\draw[fill=blue!50,opacity=0.8, draw=none]      (axis cs:1.5,0.5) -- (axis cs:5,0.5) -- (axis cs:5,-5) -- (axis cs:-5,-5) --  (axis cs:-5,5) -- (axis cs:1.5,5) --cycle ;
			\draw[fill=blue!50,opacity=0.8, draw=none, pattern=north east lines, pattern color=orange] (axis cs:0.5,0.5) -- (axis cs:1.5,0.5) -- (axis cs:1.5,5) -- (axis cs:0.5,5)  --cycle ;


			\addplot+ [mark=none, black, line width = 1, domain=1.5:5] {0.5};
			\addplot +[mark=none, black, line width = 1, dotted] coordinates {(0.5,0.5) (0.5, 5)};
			\addplot +[mark=none, black, line width = 1] coordinates {(1.5,0.5) (1.5, 5)};



			\addplot +[mark=none, black, line width = 0.5, dotted] coordinates {(0,0.5) (1.5, 0.5)};
			\addplot +[mark=none, black, line width = 0.5, dotted] coordinates {(1.5,0) (1.5, 0.5)};
			\addplot +[mark=none, black, line width = 0.5, dotted] coordinates {(0.5,0) (0.5, 0.5)};

			\node [below, label={[align=left] \scriptsize College 1  feasible \\ \scriptsize College 1  acceptable \\ \footnotesize $\mu(i)=1$}] at (axis cs: 3.3,2.5){};

			\node [left] at (axis cs: -2,-2) {\footnotesize $\mu(i)=0$};
			\node [below] at (axis cs: 1,3) {\footnotesize $I_1:$};
			\node [below] at (axis cs: 0.96,2.5) {\tiny $\frac{\partial\sigma_{1}}{\partial y_{1}}\Delta y$};
			\node [below] at (axis cs: 4.8,0) {\footnotesize $\epsilon_{i1}$};
			\node [left] at (axis cs: 0,4.8) {\footnotesize $\eta_{1i}$};


		    \node [above left] at (axis cs: 0,0.15) {\scriptsize $\delta_1-v^1(z)-w_1$};
	     	\addplot[mark=*, mark size=1.2pt] coordinates {(0,0.5)};
	    	\addplot[mark=*, mark size=1.2pt] coordinates {(1.5,0)};

			\node[label={[label distance=0.18cm]275:{\scriptsize $-u^1(z)-r^1(\overline{y}_1-\Delta y)$}},inner sep=0pt] at (axis cs:1.2,0) {};
			\addplot[mark=*, mark size=1.2pt] coordinates {(0.5,0)} node[pin={[pin distance=0.8cm]273:{\scriptsize $-u^1(z)-r^1(\overline{y}_1)$}}]{};
			\node[label={230:{\scriptsize 0}},circle,fill,inner sep=1pt] at (axis cs:0,0) {};




			\node[draw, single arrow,
			minimum height=8mm, minimum width=1mm,
			single arrow head extend=1mm,
			anchor=west, rotate=0] at (axis cs: 0.5,4.6) {};
		\end{axis}
		\end{tikzpicture}
		\subcaption[]
		{{\footnotesize Changes in $\mu(i)$  when $y_{i1} \downarrow$ by ${\Delta}y$}}
		\label{fig:panelb}


	\end{minipage}
	\hspace*{\fill}
	\begin{minipage}[t]{.45\linewidth}

		\centering


				\begin{tikzpicture} [scale=0.65]

			\tikzset{every pin edge/.style={draw=black, <-,very near end}}

			\begin{axis}
				[
				ymin=-5,
				xmin=-5,
				ymax=5,
				xmax=5,
				width=10cm,
				axis on top=true,
				axis x line=middle,
				axis y line=middle,
				yticklabels={,,},
				xticklabels={,,}
				]
				\draw[fill=yellow!60,opacity=0.8, draw=none]      (axis cs:0.5,1.5) -- (axis cs:0.5,5) -- (axis cs:5,5) -- (axis cs:5,1.5) -- cycle ;
				\draw[fill=blue!50,opacity=0.8, draw=none]      (axis cs:0.5,1.5) -- (axis cs:5,1.5) -- (axis cs:5,-5) -- (axis cs:-5,-5) --  (axis cs:-5,5) -- (axis cs:0.5,5) --cycle ;
				\draw[fill=blue!50,opacity=0.8, draw=none, pattern=north east lines, pattern color=orange] (axis cs:0.5,0.5) -- (axis cs:5,0.5) -- (axis cs:5,1.5) -- (axis cs:0.5,1.5)  -- (axis cs:0.5,5) -- (axis cs:0.5,5)--cycle ;


				\addplot+ [mark=none, black, line width = 1, domain=0.5:5, dotted] {0.5};
				\addplot +[mark=none, black, line width = 1] coordinates {(0.5,1.5) (0.5, 5)};
				\addplot+ [mark=none, black, line width = 1, domain=0.5:5] {1.5};




				\addplot +[mark=none, black, line width = 0.5, dotted] coordinates {(0,0.5) (0.5, 0.5)};
				\addplot +[mark=none, black, line width = 0.5, dotted] coordinates {(0,1.5) (0.5, 1.5)};
				\addplot +[mark=none, black, line width = 0.5, dotted] coordinates {(0.5,0) (0.5, 1.5)};

				\node [below, label={[align=left] \scriptsize College 1  feasible \\ \scriptsize College 1  acceptable \\ \footnotesize $\mu(i)=1$}] at (axis cs: 3.3,2.5){};

				\node [left] at (axis cs: -2,-2) {\footnotesize $\mu(i)=0$};
				\node [below] at (axis cs: 2.5,1.5) {\footnotesize $I_2:\frac{\partial\sigma_{1}}{\partial w_{1}}\Delta w$};
				\node [below] at (axis cs: 4.8,0) {\footnotesize $\epsilon_{i1}$};
				\node [left] at (axis cs: 0,4.8) {\footnotesize $\eta_{1i}$};


				\node [below right] at (axis cs: 0.5,0) {\scriptsize $-u^1(z)-r^1(\overline{y}_1)$};
				\node [above left] at (axis cs: 0,0.15) {\scriptsize $\delta_1-v^1(z)-w_1$};
				\node [above left] at (axis cs: 0,1.15) {\scriptsize $\delta_1-v^1(z)-(w_1-\Delta w)$};
				\addplot[mark=*, mark size=1.2pt] coordinates {(0,0.5)};
				\addplot[mark=*, mark size=1.2pt] coordinates {(0.5,0)};
				\addplot[mark=*, mark size=1.2pt] coordinates {(0,1.5)};


				\node[label={230:{\scriptsize 0}},circle,fill,inner sep=1pt] at (axis cs:0,0) {};





				\node[draw, single arrow,
				minimum height=7mm, minimum width=1mm,
				single arrow head extend=1mm,
				anchor=west, rotate=90] at (axis cs: 4.6,0.5) {};
			\end{axis}
		\end{tikzpicture}
		\subcaption[]
		{{\footnotesize Changes in $\mu(i)$ when $w_{1i} \downarrow$ by $\Delta w$}}
		\label{fig:panelc}


		\begin{tikzpicture} [scale=0.65]

			\tikzset{every pin edge/.style={draw=black, <-,very near end}}

		\begin{axis}
			[
			ymin=-5,
			xmin=-5,
			ymax=5,
			xmax=5,
			width=10cm,
			axis on top=true,
			axis x line=middle,
			axis y line=middle,
			yticklabels={,,},
			xticklabels={,,}
			]
			\draw[fill=yellow!60,opacity=0.8, draw=none]      (axis cs:1.5,1.5) -- (axis cs:1.5,5) -- (axis cs:5,5) -- (axis cs:5,1.5) -- cycle ;
			\draw[fill=blue!50,opacity=0.8, draw=none]      (axis cs:1.5,1.5) -- (axis cs:5,1.5) -- (axis cs:5,-5) -- (axis cs:-5,-5) --  (axis cs:-5,5) -- (axis cs:1.5,5) --cycle ;
			\draw[fill=blue!50,opacity=0.8, draw=none, pattern=north east lines, pattern color=orange] (axis cs:0.5,0.5) -- (axis cs:5,0.5) -- (axis cs:5,1.5) -- (axis cs:1.5,1.5)  -- (axis cs:1.5,5) -- (axis cs:0.5,5)--cycle ;

			\addplot+ [mark=none, black, line width = 1, domain=0.5:5, dotted] {0.5};
			\addplot +[mark=none, black, line width = 1, dotted] coordinates {(0.5,0.5) (0.5, 5)};
			\addplot +[mark=none, black, line width = 1] coordinates {(1.5,1.5) (1.5, 5)};
			\addplot+ [mark=none, black, line width = 1, domain=1.5:5] {1.5};



			\addplot +[mark=none, black, line width = 0.5, dotted] coordinates {(0,0.5) (0.5, 0.5)};
			\addplot +[mark=none, black, line width = 0.5, dotted] coordinates {(1.5,0) (1.5, 1.5)};
			\addplot +[mark=none, black, line width = 0.5, dotted] coordinates {(0.5,0) (0.5, 0.5)};
			\addplot +[mark=none, black, line width = 0.5, dotted] coordinates {(0,1.5) (1.5, 1.5)};

			\node [below, label={[align=left]	\scriptsize College 1  feasible \\ \scriptsize College 1  acceptable \\ \footnotesize $\mu(i)=1$}] at (axis cs: 3.3,2.5){};
			\node [left] at (axis cs: -2,-2) {\footnotesize $\mu(i)=0$};
			\node [below] at (axis cs: 2.5,1.4) {\footnotesize $I_3: \frac{\partial\sigma_{1}}{\partial z}\Delta z$};
			\node [below] at (axis cs: 4.8,0) {\footnotesize $\epsilon_{i1}$};
			\node [left] at (axis cs: 0,4.8) {\footnotesize $\eta_{1i}$};


			\node[label={[label distance=0.18cm]275:{\scriptsize $-u^1(z-\Delta z)-r^1(\overline{y}_1)$}},inner sep=0pt] at (axis cs:1.2,0) {};
			\node [above left] at (axis cs: 0,0.15) {\scriptsize $\delta_1-v^1(z)-w_1$};
			\node [above left] at (axis cs: 0,1.15) {\scriptsize $\delta_1-v^1(z-\Delta z)-w_1$};
			\addplot[mark=*, mark size=1.2pt] coordinates {(0.5,0)} node[pin={[pin distance=0.8cm]273:{\scriptsize $-u^1(z)-r^1(\overline{y}_1)$}}]{};
			\addplot[mark=*, mark size=1.2pt] coordinates {(0,0.5)};
			\addplot[mark=*, mark size=1.2pt] coordinates {(1.5,0)};
			\addplot[mark=*, mark size=1.2pt] coordinates {(0,1.5)};


			\node[label={230:{\scriptsize 0}},circle,fill,inner sep=1pt] at (axis cs:0,0) {};



			\node[draw, single arrow,
			minimum height=8mm, minimum width=1mm,
			single arrow head extend=1mm,
			anchor=west, rotate=0] at (axis cs: 0.5,4.6) {};

			\node[draw, single arrow,
			minimum height=7mm, minimum width=1mm,
			single arrow head extend=1mm,
			anchor=west, rotate=90] at (axis cs: 4.6,0.5) {};
		\end{axis}
		\end{tikzpicture}
		\subcaption[]
		{{\footnotesize Changes in $\mu(i)$  when $z_i \downarrow$ by $\Delta z$}}
		\label{fig:paneld}

	\end{minipage}
	\caption{Partitioning the Space of Unobservables in the One-College Case}\label{partition}
	\begin{flushleft}
		{\footnotesize\noindent \textit{Notes:} Panel (a) describes the partition of the $(\epsilon_{i1},\eta_{1i})$ space given $(z_{i}, y_{i1}, w_{1i}) = (z,\overline{y}_1,w_1)$ by student $i$'s feasible set and preferences. The other panels show the changes in $\mu(i)$ when $y_{i}$ decreases by  $\Delta y$ and affects only student preferences (panel~b), when $w_{1i}$ decreases by $\Delta w$ and affects only feasible set (panel~c), and when $z_{i}$ decreases by $\Delta z$ (panel~d).}
	\end{flushleft}
\end{figure}

Figures~\ref{partition}(b)--(d) depict how the marginal effect of $z_{i}$ on match probability is linked to the marginal effects of the excluded variables, $y_{i1}$ and $w_{1i}$.
Panel (b) describes the marginal effect of $y_{i1}$. Specifically, decreasing $y_{i1}$ from $\overline{y}_1$ to $\overline{y}_1-\Delta y$ makes college~$1$ less attractive to student~$i$, and the region in which $\mu(i)=1$ shrinks along the horizontal $\epsilon_{i1}$-axis.  The area $I_{1}$ depicts the set of students whose match differs when $y_{i1}$ decreases. The induced change in the match probability, $\sigma_1=\mathbb{P}\left(\mu(i)=1|(z_{i},y_{i1},w_{1i}) = (z,\overline{y}_1,w_1)\right)\equiv\mathbb{P}(\mu(i)=1|z,\overline{y}_1,w_1)$, is the mass that the density of $\left(\epsilon_{i1},\eta_{1i}\right)$ puts on $I_{1}$, or, for  $(z_{i},y_{i1},w_{1i}) = (z,\overline{y}_1,w_1)$,
\begin{equation}
	\frac{\partial\sigma_{1}(z,\overline{y}_{1},w_{1})}{\partial y_{i1}}=\frac{\partial r^1(\overline{y}_{1})}{\partial y_{i1}}\cdot\sum_{L\in\mathcal{{L}}}\lambda_{L}(\iota_{1})\cdot \frac{\partial g_{1,L}(\tau_{1})}{\partial \tau_{i1}}=\sum_{L\in\mathcal{{L}}}\lambda_{L}(\iota_{1})\cdot \frac{\partial g_{1,L}(\tau_{1})}{\partial \tau_{i1}},
	\label{diffy}
\end{equation}
where $\tau_1 \equiv u^1(z)+r^1(\overline{y}_{1})$ and $\iota_{1} \equiv v^1(z)+w_{1}$. By scale normalization, $\frac{\partial r^{1}\left(\overline{y}_{1}\right)}{\partial y_1}=1$.

Panel~(c) shows a similar graph in which decreasing $w_{1i}$ from $w_1$ to $w_1 - \Delta w$ makes college~$1$ less likely to be feasible to student~$i$. Hence, the region $\mu(i) = 1$ shrinks along the vertical $\eta_{1i}$-axis.
The change in the match probability induced by the decrease in $w_{1i}$ is the mass that
the density of $\left(\epsilon_{i1},\eta_{1i}\right)$ puts on the area $I_{2}$, or,
\begin{equation}
\frac{\partial\sigma_1(z,\overline{y}_{1},w_{1})}{\partial w_{1i}}=\sum_{L\in\mathcal{{L}}}\frac{\partial \lambda_{L}(\iota_{1})}{\partial \iota_{1i}}\cdot g_{1,L}(\tau_{1}). \label{diffw}
\end{equation}


In panel (d), the decrease in $z_{i}$ reduces the attractiveness and feasibility of college~$1$ for student~$i$, because $z_{i}$ enters both student and college preferences. Besides, how
$z_{i}$ changes the region of $\mu(i)=1$ relies on the shape of the functions $u^1$ and $v^1$. The change in the match probability
induced by the change in $z_{i}$ corresponds to the area
$I_{3}$, or,
\begin{align}
\frac{{\partial\sigma_1(z,\overline{y}_{1},w_{1})}}{\partial z_{i}} & =\frac{\partial u^1(z)}{\partial z_{i}}\cdot\sum_{L\in\mathcal{{L}}}\lambda_{L}(\iota_{1})\cdot \frac{\partial g_{1,L}(\tau_{1})}{\partial \tau_{i1}}+\frac{\partial v^1(z)}{\partial z_{i}}\cdot\sum_{L\in\mathcal{{L}}}\frac{\partial \lambda_{L}(\iota_{1})}{\partial \iota_{1i}}\cdot g_{1,L}(\tau_{1}). \label{diffz}
\end{align}
Our identification result relies on the changes caused by $z_{i}$, $y_{i1}$, and $w_{1i}$. Plugging equations \eqref{diffy}
and \eqref{diffw} into equation \eqref{diffz}, we have,   for  $(z_{i},y_{i1},w_{1i}) = (z,\overline{y}_1,w_1)$,
\begin{equation}
\frac{{\partial\sigma_1(z,\overline{y}_{1},w_{1})}}{\partial z_{i}}=\frac{\partial u^1(z)}{\partial z_{i}}\cdot\frac{{\partial\sigma_1(z,\overline{y}_{1},w_{1})}}{\partial y_{i1}}+\frac{\partial v^1(z)}{\partial z_{i}}\cdot\frac{{\partial\sigma_1(z,\overline{y}_{1},w_{1})}}{\partial w_{1i}}.\label{eq:indsim}
\end{equation}
This equation reflects the chain rule: the effect of $z_{i}$ on
the  match probability, $\frac{\partial\sigma_1}{\partial z_{i}}$, is
realized through its effects on utilities $u_{i1}$ and $v_{1i}$,
captured by $\frac{\partial u^1}{\partial z_{i}}$ and
$\frac{\partial v^1}{\partial z_{i}}$, and the effects of the utilities on the  match probability, captured by $\frac{\partial {\sigma_1}}{\partial y_{i1}}$
and $\frac{\partial {\sigma_1}}{\partial w_{1i}}$.
In equation~\eqref{eq:indsim}, the derivatives of the match probability can be recovered from the population data, and the two unknowns, $\frac{\partial {u^1}}{\partial z_{i}}$ and $\frac{\partial v^1}{\partial z_{i}}$ are the parameters of interest.

Importantly, when $w_{1i}$ varies, the conditional match probability changes, but the two unknowns remain constant, a consequence of $w_{1i}$ entering $v_{1i}$ in a known way. If, for any $z$, $(\epsilon_{i1},\eta_{1i})$ has enough
variation such that two distinct values of $w_{1i}$ produce two linearly independent equations that have a unique solution, we identify the unknowns. This requirement is formalized as \cref{cond} later, which imposes a mild restriction on the distribution of $(\epsilon_{i1},\eta_{1i})$ as detailed in Online Appendix \ref{app:rank_onecollege}.

We then identify $\frac{\partial r^1}{\partial y_{i1}}$, which is not one when $y_{i1}\neq \overline{y}_1$.  For any value $(z,y_{1},w_{1})$, plugging equations \eqref{diffy}
	and \eqref{diffw} into equation \eqref{diffz}, we rearrange and obtain
\begin{equation}
\left(\frac{\partial\sigma_{1}(z,y_{1},w_{1})}{\partial z_{i}}-\frac{\partial v^{1}(z)}{\partial z_{i}}\cdot\frac{\partial\sigma_{1}(z,y_{1},w_{1})}{\partial w_{1i}}\right)\cdot\frac{\partial r^{1}(y_{1})}{\partial y_{i1}}=\frac{\partial u^{1}(z)}{\partial z_{i}}\cdot\frac{\partial\sigma_{1}(z,y_{1},w_{1})}{\partial y_{i1}}, \label{eq:indsim2}
\end{equation}
where all the terms except for $\frac{\partial r^1}{\partial y_{i1}}$ are either identified or known. For any $y_{1}$, if $\frac{\partial\sigma_{1}}{\partial z_{i}}-\frac{\partial v^{1}}{\partial z_{i}}\cdot\frac{\partial\sigma_{1}}{\partial w_{1i}} \neq 0$ for some value $(z,w_{1})$,
$\frac{\partial r^1}{\partial y_{i1}}$ is identified; otherwise, equation~\eqref{eq:indsim2} implies that $\frac{\partial\sigma_{1}}{\partial y_{i1}}=0$ for all $(z,w_{1})$ and thus $\frac{\partial r^1}{\partial y_{i1}}$ is also identified and equal to zero.

Below, we extend this example to the case with multiple colleges. We derive equation~\eqref{eq:indsim} for each college, in which the marginal effect of $z_i$ on the probability of being matched with each college is the sum of its marginal effects on $(u^c,v^c)$ for all $c$. By varying $\{w_{ci}\}_c$, we form a system of equations in $\{\frac{\partial u^{c}}{\partial z_{i}}, \frac{\partial v^{c}}{\partial z_{i}}\}_c$.
 The identification of $\frac{\partial r^{c}}{\partial y_{ic}}$ is the same as above and relies on a generalized version of equation~\eqref{eq:indsim2} because we can identify $\frac{\partial r^{c}}{\partial y_{ic}}$ for each $c$ separately by holding $\frac{\partial r^{d}}{\partial y_{id}}=1$ for $d\neq c$.

\subsubsection{Formal Identification Results}\label{sec:id_results}

Our nonparametric identification of $\{\frac{\partial u^{c}}{\partial z_{i}}, \frac{\partial r^{c}}{\partial y_{ic}}, \frac{\partial v^{c}}{\partial z_{i}}\}_{c}$ extends \cite{matzkin_constructive_2019} who uses excluded variables to identify nonparametric nonseparable discrete choice models.
\begin{assm} \label{assm_nonpara}
(i) $z_i$, $y_i$, and $w_i$ are continuously distributed;  (ii) for each $c\in\mathbf{C}$,  the functions, $(u^{c}, r^{c}, v^{c})$, are continuously differentiable; and (iii) $F$ is continuously differentiable.
\end{assm}

Part (i) of Assumption \ref{assm_nonpara} requires that all covariates are continuous (but not necessarily have full-support), which is relaxed in \cref{sec:practical}.

\begin{assm}\label{indep} $(\epsilon_{i},\eta_{i})$ is distributed
independently of $(z_i,y_i,w_i)$. \end{assm}

This exogeneity assumption is made for simplicity. One way to relax this assumption is to adopt a control function approach \citep{heckman1985alternative,blundell2004endogeneity, imbens2009identification}. See Online Appendix~\ref{app:cf} for a discussion.

Additionally, we need a condition on the derivatives of match probabilities w.r.t.\ the excluded variables. Let $\Pi_{y}(z_{i},y_{i},w_{i})$ be a $C\times C$ Jacobian matrix of the match probabilities w.r.t.\ the excluded variable $y_{i}$, whose $(c,d)$ element is $\frac{\partial \sigma_{c}(z_{i},y_{i},w_{i})}{\partial y_{id}}$. Similarly, let $\Pi_{w}(z_{i},y_{i},w_{i})$ be a $C\times C$ Jacobian matrix w.r.t.\ the excluded variable $w_{i}$. Fix $y_i=\overline{y},$ where $\overline{y}=(\overline{y}_1,\cdots,\overline{y}_C)$. We then consider a pair of distinct values of $w_{i}$, $\widehat{w}$ and $\widetilde{w}$, and define a $2C\times2C$ matrix evaluated at $\left(z_{i},y_{i}\right)=\left(z,\overline{y}\right)$,
	\[
	\Pi(z,\overline{y},\widehat{w},\widetilde{w})\equiv\begin{pmatrix}\Pi_{y}(z,\overline{y},\widehat{w}) & \Pi_{w}(z,\overline{y},\widehat{w})\\
		\Pi_{y}(z,\overline{y},\widetilde{w}) & \Pi_{w}(z,\overline{y},\widetilde{w})
	\end{pmatrix}.
	\]
We impose the following testable condition on $\Pi(z,\overline{y},\widehat{w},\widetilde{w})$.\footnote{\label{fn:test}For a statistical test of \cref{cond} given a value of $z_i$ and a pair of values of $w_i$, one may use the method proposed by \cite{chen_improved_2019}. Testing $H_{0}$: rank$(\Pi(z,\overline{y},\widehat{w},\widetilde{w}))\leq2C-1$ against $H_{1}$: rank$(\Pi(z,\overline{y},\widehat{w},\widetilde{w}))>2C-1$ is a special case of setup (1) in \citet[p.1788]{chen_improved_2019}. In practice, one needs to find two values of $w_i$ satisfying \cref{cond}, which can be achieved by the following procedure: (i) choose $m$ pairs of $w_i$ values, (ii) for each pair, apply this test and calculate the $p$ values, and (iii) use the Holm–Bonferroni method to control the overall size of this multiple hypothesis testing problem and then determine which null hypothesis, if any, is rejected. The pair of values of $w_i$ associated with any rejected hypothesis satisfies \cref{cond}.}
\begin{cond}\label{cond}
For any $z$ in the interior of $\mathcal{Z} $, there exist two values of $w_i$, $\widehat{w}$ and $\widetilde{w}$, in $w_i$'s support conditional on $(z_{i},y_{i})=(z,\overline{y})$ such that $\Pi(z,\overline{y},\widehat{w},\widetilde{w})$ has rank $2C$.
\end{cond}

Note that, for any value of $z_i$, \cref{cond} only needs {\it two} values of $w_i$ at which $\Pi(z,\overline{y},\widehat{w},\widetilde{w})$ is full-rank.
In other words, \cref{cond} can hold even when there are \textit{infinitely many} values of $w_i$ at which $\Pi(z,\overline{y},\widehat{w},\widetilde{w})$ is \textit{not} full-rank. The condition fails if, for example, for the given value $(z,\overline{y})$, the student's probability of matching with a certain college is always zero for a neighborhood around $\overline{y}$ and all values of $w_i$.\footnote{This may occur if the college is infeasible or unaccaptable to the student with probability one.}

 For a set of sufficient, yet weak and testable, conditions for \cref{cond}, consider $\widehat{w}$ that is large enough such that all colleges are feasible when $w_i=\widehat{w}$. In this case, the upper half of $\Pi(z,\overline{y},\widehat{w},\widetilde{w})$ corresponds to a one-sided discrete choice model, where the feasible-set shifter $w_i$ has no impact on the matching probabilities and thus $\Pi_{w}(z,\overline{y},\widehat{w})=\boldsymbol{0}_{C\times C}$. This makes $\Pi(z,\overline{y},\widehat{w},\widetilde{w})$ a lower triangular block matrix. Then, \cref{cond} holds if and only if both $\Pi_{y}(z,\overline{y},\widehat{w})$ and $\Pi_{w}(z,\overline{y},\widetilde{w})$ are invertible, which can be achieved under the connected substitutes conditions \citep[ Theorem 2]{berry2013connected}.\footnote{The invertibility of $\Pi_{y}(z,\overline{y},\widehat{w})$ holds under the following two assumptions: (i) for each $d\in\mathbf{C}\cup\{0\}$, $\frac{\partial \sigma_{d}(z_i,y_i,w_i)}{\partial y_{ic}}\leq 0$ for all $c\in\mathbf{C}\backslash\{d\}$; (ii) for any nonempty $\overline{\mathbf{C}}\subseteq\mathbf{C}$, there exists $c\in\overline{\mathbf{C}}$ and $d\notin\overline{\mathbf{C}}$ such that $\frac{\partial\sigma_{d}(z,\overline{y},\widehat{w})}{\partial y_{ic}}<0$. The invertibility of $\Pi_{w}(z,\overline{y},\widetilde{w})$ holds under similar assumptions.}\textsuperscript{, }\footnote{\label{fn:as2022}This sufficient condition for \cref{cond} provides us with a clear comparison with a related study, \cite{agarwal2022demand}. Their conditions for the identification of a similar model and our conditions are not nested. Specifically, they assume large support on $w_i$, impose a substitution condition on $\Pi_{y}(z,y,w)$ stronger than the one in \cite{berry2013connected}, and require the substitution condition to hold for \textit{all} but a finite set of $y_i$ values. In contrast, in this sufficient condition for \cref{cond}, with large support of $w_i$, we need a substitution condition on both $\Pi_{y}(z,\overline{y},w)$ and $\Pi_{w}(z,\overline{y},w)$ \`a la \cite{berry2013connected}, but only for $y_i = \overline{y}$.} This suggests that \cref{cond} holds under common logit or probit models. It is worth noting that the large support of $w_i$ above is mainly used to facilitate a clear connection with the connected substitutes conditions, but is  stronger than necessary. In Online Appendix \ref{app:rank_onecollege}, we show that in a one-college case, \cref{cond} holds for all distribution of $\eta_{1i}$ except for the exponential distribution. In Online Appendix \ref{app:rank_para}, we analyze common logit and probit models with two or more colleges. The results suggest that \cref{cond} is non-restrictive and satisfied for a wide range of values of $(\widehat{w},\widetilde{w})$; in fact, the \textit{failure} of \cref{cond} imposes strict restrictions on the supply, or the conditional probabilities of different feasible sets.

\begin{prop} \label{prop_nonpara} Under Assumptions \ref{assm_struc},  \ref{assm_nonpara}, \ref{indep}, and Condition \ref{cond}, for $k= 1,\dots,d_{z}$,
$\{\frac{\partial u^{c}(z)}{\partial z_{i}^{k}},\frac{\partial v^{c}(z)}{\partial z_{i}^{k}},\frac{\partial r^{c}(y_c)}{\partial y_{ic}}\}_{c\in\mathbf{C}}$ are identified for all $\left(z,y\right)$ in the interior of $\mathcal{Z} \times \mathcal{Y}$.
\end{prop}



Our identification arguments proceed in two steps. First, to identify $\frac{\partial u^{c}(z)}{\partial z_{i}^{k}}$ and  $\frac{\partial v^{c}(z)}{\partial z_{i}^{k}}$, we use the variation in the excluded variables $(y_{i},w_{i})$ to derive a system of linear equations that generalizes equation~\eqref{eq:indsim}. Fixing $y_i=\overline{y}$, for any value $(z,w)$ in the interior of $\mathcal{Z}\times\mathcal{W}$, for each college $d \in \mathbf{C}$, and for any component of $z_{i}$, $z_{i}^{k}$,
\begin{equation}
	\frac{\partial\sigma_{d}\left(z,\overline{y},w\right)}{\partial z_{i}^{k}}=\sum_{c\in\mathbf{C}}\frac{\partial\sigma_{d}\left(z,\overline{y},w\right)}{\partial y_{ic}} \cdot \frac{\partial u^{c}(z)}{\partial z_{i}^{k}}+\sum_{c\in\mathbf{C}}\frac{\partial\sigma_{d}\left(z,\overline{y},w\right)}{\partial w_{ci}} \cdot \frac{\partial v^{c}(z)}{\partial z_{i}^{k}},\label{eq:indsimC}
\end{equation}
which gives $C$ linear equations of $2C$ unknowns.
Note that equation~\eqref{eq:indsimC} is for one value of $w_i$. By evaluating equation~\eqref{eq:indsimC} at $d=1,\dots,C$ and two different values of $w_i$, $\widehat{w}$ and $\widetilde{w}$, and stacking them together, we have
\begin{equation}
	\begin{pmatrix}\frac{\partial\boldsymbol{\sigma}\left(z,\overline{y},\widehat{w}\right)}{\partial z_{i}^{k}}\\
		\frac{\partial\boldsymbol{\sigma}\left(z,\overline{y},\widetilde{w}\right)}{\partial z_{i}^{k}}
	\end{pmatrix}=\Pi(z,\overline{y},\widehat{w},\widetilde{w})\cdot\begin{pmatrix}\frac{\partial\boldsymbol{u}\left(z\right)}{\partial z_{i}^{k}}\\
		\frac{\partial\mathbf{v}\left(z\right)}{\partial z_{i}^{k}}
	\end{pmatrix},\label{eq:np_ideq}
\end{equation}
where $\boldsymbol{\sigma} \equiv (\sigma_{1},\dots,\sigma_{C})'$, $\boldsymbol{u} \equiv (u^{1},\dots,u^{C})'$, and $\boldsymbol{v} \equiv (v^{1},\dots,v^{C})'$. Both $\Pi(z,\overline{y},\widehat{w},\widetilde{w})$ and the left-hand side are known from the population data.
Hence, the invertibility of $\Pi(z,\overline{y},\widehat{w},\widetilde{w})$ in \cref{cond} guarantees the existence of a unique solution to this system, leading to the identification of  $\frac{\partial u^{c}(z)}{\partial z_{i}^{k}}$ and  $\frac{\partial v^{c}(z)}{\partial z_{i}^{k}}$.
Importantly, this only requires a pair of distinct values of $w_{i}$ to satisfy \cref{cond}.

Second, we identify $\frac{\partial r^{c}(y_{c})} {\partial y_{ic}}$ for each $c\in\mathbf{C}$ by generalizing equation~\eqref{eq:indsim2} for the one-college example. See the Appendix for detailed proof.

In certain empirical applications, identifying the above derivatives is sufficient, in which case a large-support assumption on excluded variables is not needed. When the functions $(u^{c}, v^{c}, r^{c})$ must be identified, a full-support condition on their arguments, $(z_{i}, y_{i})$, is often imposed. This is the case for our next result of identifying the joint distribution $F$ and cutoffs $\delta_c$. Additionally, a full-support assumption on $w_i$ is also required. Hence, we will assume all observables, $(z_i,y_i,w_i)$, have full support.  \subsection{Identifying the Cutoffs and Joint Distribution of Unobservables}\label{sec:id_dist}
We now formalize the assumptions that are needed for the identification of the cutoffs, $\{\delta_c\}_c$, and the joint distribution of unobservables, $F$.
\begin{assm}\label{assm:F_id}
	For each $c\in\mathbf{C}$,
		(i) the functions $u^{c}+r^{c}$ and $v^{c}$ are identified;\footnote{\label{fn:id_fn}This requires \cref{prop_nonpara}, a full support assumption on $(z_i,y_i)$, and location normalization on $u^{c}+r^{c}$ and $v^{c}$ for each $c$. See \cref{rm:location} for a discussion on location normalization.}
		(ii)  $y_{ic}$ and $w_{ci}$ possess an everywhere positive Lebesgue density conditional
		on $z_{i}$;
		(iii) the range of the function $r^{c}$ is the whole real line, and $\mathcal{W} =  \mathbb{R}^C$; and
		(iv) the $\rho_c$-quantile of the marginal distribution of $\eta_{ci}$ is $0$, i.e., $\text{Quantile}_{\eta_{ci}}(\rho_c)\equiv \inf\{\eta_{c}:  F_{\eta_{ci}}(\eta_{c}) \geq \rho_{c} \} =0,$ for an arbitrary $\rho_c \in (0,1)$.
\end{assm}

Part~(iv) is a location normalization on college preferences, as mentioned in \cref{rm:location}, which can be done college-by-college. Alternatively, one may replace part~(iv) by normalizing cutoffs $\{\delta_{c}\}_c$ to zero.

Before presenting our formal results, we give some intuitions. To identify cutoff~$\delta_{c}$,  under the full-support assumption on student preferences (parts ii and iii of \cref{assm:F_id}), we consider a mass of students to whom all colleges except for $c$ are unacceptable. {The probability of these students matching with $c$ is $1-F_{\eta_{ci}}(\delta_c - v^c(z_{i}) -w_{ci})$. Given the location normalization of $F_{\eta_{ci}}$ (part~iv of  \cref{assm:F_id}), finding the maximum value of $v^c(z_{i})+w_{ci}$ that sets this probability to $1-\rho_c$ identifies $\delta_{c}$.}

To identify $F$, we use the conditional probability of being \textit{unmatched}. Let us illustrate the intuition in the same one-college example as in \cref{id_eg}. For any given value $(z,y_1,w_1)$,
Figure~\ref{cdf} shows the partition of the space of $(\epsilon_{i1},\eta_{1i})$
by matching outcome, $\mu\left(i\right)=0$ or $1$. The conditional probability of $\mu(i)=0$, highlighted in purple in the figure, can be decomposed into three parts, $R_{1}$,
$R_{2}$, and $R_{3}$. Moreover,
\begin{align}
& \mathbb{P}(\mu(i)=0  |(z_{i},y_{i1},w_{1i}) = (z,y_1,w_1)) \notag \\
= & \mathbb{P}(\epsilon_{i1}<-u^1(z)-r^1\left(y_1\right)\text{ or }\eta_{1i}<\delta_1-v^1(z)-w_1  |z,y_1,w_1 )\nonumber \\
= & \mathbb{P}(R_{1}\cup R_{2}  | z,y_1,w_1 )+\mathbb{P}(R_{3}\cup R_{2}  | z,y_1,w_1 )-\mathbb{P}(R_{2}  | z,y_1,w_1 ),\label{decomp}
\end{align}
where $\ensuremath{\mathbb{P}(R_{2} | z,y_1,w_1 })$ is the joint CDF of $(\epsilon_{i1},\eta_{1i})$,  $F\left(-u^{1}(z)-r^{1}\left(y_1\right),\delta_{1}-v^{1}(z)-w_1\right)$, our parameter of interest.

\begin{figure}[h!]
	\begin{centering}
		\begin{tikzpicture}[scale=0.65]

	    \tikzset{every pin edge/.style={draw=black,<-, very near end}}
		\begin{axis}
		[
		ymin=-5,
		xmin=-5,
		ymax=5,
		xmax=5,
		width=10cm,
		axis on top=true,
		axis x line=middle,
		axis y line=middle,
		yticklabels={,,},
		xticklabels={,,}
		]
		\draw[fill=yellow!50,opacity=0.8, draw=none]      (axis cs:0.5,0.5) -- (axis cs:0.5,5) -- (axis cs:5,5) -- (axis cs:5,0.5) -- cycle ;
		\draw[fill=blue!50,opacity=0.8, draw=none]      (axis cs:0.5,0.5) -- (axis cs:5,0.5) -- (axis cs:5,-5) -- (axis cs:-5,-5) --  (axis cs:-5,5) -- (axis cs:0.5,5) --cycle ;
		\draw[fill=blue!50,opacity=0.8, draw=none, pattern=north east lines, pattern color=blue] (axis cs:0.5,0.5) -- (axis cs:0.5,-5) -- (axis cs:-5,-5) -- (axis cs:-5,0.5)  --cycle ;
		\draw[fill=blue!50,opacity=0.3, draw=none, pattern=dots, pattern color=blue] (axis cs:0.5,0.5) -- (axis cs:0.5,5) -- (axis cs:-5,5) -- (axis cs:-5,0.5)  --cycle ;
		\draw[fill=blue!50,opacity=0.3, draw=none, pattern=dots, pattern color=blue] (axis cs:0.5,0.5) -- (axis cs:0.5,-5) -- (axis cs:5,-5) -- (axis cs:5,0.5)  --cycle ;



		\addplot+ [mark=none, black, line width = 0.5, domain=0.5:5] {0.5};
		\addplot +[mark=none, black, line width = 1] coordinates {(0.5,0.5) (0.5, 5)};

		\addplot+ [name path = A, mark=none, darkgray, dotted, line width = 0.5, domain=-5:0.5] {0.5};
		\addplot +[name path = C, mark=none, dotted, darkgray, line width = 1] coordinates {(0.5, -5) (0.5, 0.5)};



		\node [below] at (axis cs: -3,-2.5) {\footnotesize $R_2$};
		\node [above] at (axis cs: 3,2.5) {\footnotesize $\mu(i)=1$};
		\node [left] at (axis cs: -2.5,3) {\footnotesize $R_1$};
		\node [left] at (axis cs: -0.5,-1.5) {\footnotesize $\mu(i)=0$};
		\node [below] at (axis cs: 3.5,-2.5) {\footnotesize $R_3$};
		\node [below] at (axis cs: 4.8,0) {\footnotesize $\epsilon_{i1}$};
		\node [left] at (axis cs: 0,4.8) {\footnotesize $\eta_{1i}$};


		\node[label={230:{\scriptsize 0}},circle,fill,inner sep=1pt] at (axis cs:0,0) {};


		\node [below right] at (axis cs: 0.5,0) {\scriptsize $-u^1(z)-r^1(y_1)$};
		\node [above left] at (axis cs: 0,0.5) {\scriptsize $\delta_1-v^1(z)-w_1$};
		\addplot[mark=*, mark size=1.2pt] coordinates {(0,0.5)};
		\addplot[mark=*, mark size=1.2pt] coordinates {(0.5,0)} ;


		\end{axis}
		\end{tikzpicture}
		\par\end{centering}
	\centering{\caption{\label{cdf}Partitioning the Space of Unobservables in the One-college Case}}
		\begin{flushleft}
		{\footnotesize\noindent \textit{Notes:} This figure shows the partition of the space of the unobservables $(\epsilon_{i1},\eta_{1i})$  by the matching outcome ($\mu(i) = 0$ or $1$) in a one-college setting given  $(z_{i}, y_{i1}, w_{1i}) = (z,y_1,w_1)$.}
	\end{flushleft}
\end{figure}

Further, $\mathbb{P}(R_{1}\cup R_{2} | z,y_1,w_1 )=F_{\epsilon_{i1}}(-u^1(z)-r^1(y_1))$ is the
marginal CDF of $\epsilon_{i1}$. It can be identified by considering the subset of students whose value of $w_{1i}$ is high enough so that the college will be feasible no matter what value $\eta_{1i}$ takes. That is, we can identify
$\mathbb{P}(R_{1}\cup R_{2} |z,y_1,w_1 )$ by ``shutting down'' the effects
of college preferences. Similarly, $\mathbb{P}(R_{3}\cup R_{2} | z,y_1,w_1 )$
can be identified by focusing on the subset of students whose value
of $r^1\left(y_{i1}\right)$ is large enough so that those students find college~1 acceptable no matter what value $\epsilon_{i1}$ takes.  As $\mathbb{P}(\mu(i)=0 | z,y_1,w_1 )$ is known from the population data, once $\mathbb{P}(R_{1}\cup R_{2} | z,y_1,w_1  )$ and $\mathbb{P}(R_{3}\cup R_{2} | z,y_1,w_1 )$ are identified, equation~\eqref{decomp} implies that $\mathbb{P}(R_{2} |z,y_1,w_1  )$ and thus $F$ are identified.


\begin{prop}\label{np_id_F}
Under Assumptions  \ref{assm_struc}, \ref{indep}, and \ref{assm:F_id}, (i) cutoffs $\{\delta_{c}\}_c$ are identified; (ii) the joint distribution of $(\epsilon_{i},\eta_{i})$, $F$, is identified;  and (iii) acceptability threshold $T_c $ for any college~$c$ with vacancies is identified, but $T_c$ for $c$ with a binding capacity constraint is only partially identified, $T_c \leq \delta_{c}$.
\end{prop}


To show part~(ii) in a many-college setting, with the full, large support (parts ii and iii of Assumptions \ref{assm:F_id}), we apply the same argument as in the one-college example to each college~$c$ sequentially, using extreme values of $r^c(y_{ic})$ or $w_{ci}$ to ``shut down'' the effects of students' preference for college $c$ or $c$'s preference for students.

Part~(iii) is a consequence of part~(i) and the definition of cutoffs (equation~\ref{cutoff}). If $c$ reaches its capacity, $T_c$ can be any value below $\delta_{c}$ and result in the same stable matching. We can identify $T_c$ by imposing additional assumptions, e.g., a full-capacity college having the same acceptability threshold as some college with vacancies.




     \subsection{Practical Issues\label{sec:practical}}
When our identification results are taken to the data, there can be practical issues. For example, vector~$z_i$ may include discrete variables such as gender, and the researcher may not have sufficient excluded variables. Below, we address these issues.


\noindent \textbf{Discrete Random Variables.} Our results can be extended to the case where $z_i$ contains discrete variables.
Suppose that $z_{i}=(z_{i}^1,z_{i}^2)$, where $z_{i}^1$ is a vector of discrete variables and $z_{i}^2$ is a vector of continuous variables. Let the support of $z_i^1$ be a finite set of points $\{z^{1,1}, z^{1,2},\dots, z^{1,J}\}$. For $z_i^1 = z^{1,j}$, we define functions $(u^{c,j}+r^c,v^{c,j})$. Conditional on $z_i^1=z^{1,1}$, we apply the results in \cref{sec:derivative,sec:id_dist} to identify $\{u^{c,1}+r^c,v^{c,1},\delta_c\}_{c}$ and $F$, requiring $y_i$ and $w_i$ being continuous, Assumptions~\ref{assm_nonpara}(ii) and (iii), \ref{indep}, \ref{assm:F_id}(ii)-(iv), \cref{cond}, location normalization on $u^{c,1}+r^c$ and $v^{c,1}$ for each $c$, and $(y_i,z_{i}^2)$ having full support. For $j\neq 1$, conditional on $z_i^1=z^{1,j}$, using the conditional match probability of a mass of students for whom $c$ is the only acceptable college, we can use $F$ to identify $v^{c,j}$ for each $c$; similarly, using the conditional match probability of a mass of students whose feasible set is $\{0,c\}$, we can identify the function $u^{c,j}+r^c$.


\noindent \textbf{Insufficient Excluded Variables.} Allowing for college-level heterogeneity implies that we need to recover $3C$ preference parameters, $\{u^c,r^c,v^c\}_c$. As seen in \cref{sec:derivative}, we need $2C$ excluded variables for identification of their derivatives.

The lack of sufficient excluded variables leads to a loss of identification, but not all is lost.  We show that there is a trade-off between the identifiable degree of heterogeneity and the number of excluded variables.  Consider $\mathbf{C}_1, \mathbf{C}_2 \subseteq \mathbf{C}$ with cardinality $\kappa_1$ and $\kappa_2$, respectively. For all $c\in\mathbf{C}_1$, there is a single excluded variable in student preferences, $y_{ic}=y_{i\ast}$; for all $c\in\mathbf{C}_2$, there is a single excluded variable in college preferences, $w_{ci}=w_{i\ast}$.  We have the following identification result.

\begin{prop}\label{deg-of-heter}
Suppose the preference heterogeneity is reduced: $u^c = u^{\ast}$ and $r^c = r^{\ast}$ for all $c\in\mathbf{C}_1$, and $v^c = v^{\ast}$ for all $c\in\mathbf{C}_2$.
For any $c\in\mathbf{{C}}$, the derivatives of $(u^c,r^c,v^c)$ are identified if: (i) Assumptions~\ref{assm_struc}, \ref{assm_nonpara}, and \ref{indep} hold; and
(ii) the following rank conditions hold:
(a) If $C\leq \kappa_1 + \kappa_2 - 2$, for any $z$ in the interior of $\mathcal{Z} $, there exists a value of $w_i$, $\widehat{w}$, in $w_i$'s support conditional on $(z_{i},y_{i})=(z,\overline{y})$ such that $(\Pi_{y}(z,\overline{y},\widehat{w}),\,\Pi_{w}(z,\overline{y},\widehat{w}))$ has column rank at least $2C-\kappa_1-\kappa_2+2$;
or (b) if $C> \kappa_1 + \kappa_2 - 2$, for any $z$ in the interior of $\mathcal{Z} $, there exist two values of $w_i$, $\widehat{w}$ and $\widetilde{w}$, in $w_i$'s support conditional on $(z_{i},y_{i})=(z,\overline{y})$ such that $\Pi(z,\overline{y},\widehat{w},\widetilde{w})$ has column rank at least $2C-\kappa_1-\kappa_2+2$.

\end{prop}

In other words, we need only $C-\kappa_1+1$ excluded variables on the student side and $C-\kappa_2+1$ on the college side to identify the less heterogeneous model.\footnote{Note that the rank condition in  \cref{deg-of-heter} is weaker than Condition \ref{cond}.} Importantly, this still allows for coefficient heterogeneity across colleges in $\mathbf{C}_1$ and $\mathbf{C}_2$. For example, it can be incorporated parametrically by interacting college-specific observables with $z_i$. For $c\in\mathbf{C}_1$, let $u^c(z_i) = \beta p_c z_i$, where $p_c$ is a college-specific observable. This amounts to $z_i$ having a college-specific parameter $\beta_c = \beta p_c$.

Since our identification approach is constructive, it directly implies a nonparameteric estimator. In the Monte Carlo simulations in Online Appendix~\ref{app:mc}, we apply this estimator in a semiparametric setting to avoid the well-known curse of dimensionality; yet the curse remains. Instead, we find that a parametric model based on a Bayesian approach works well. Below, we apply it to a real-life setting.








\section{Secondary School Admissions in Chile}\label{sec:empirical}

Guided by our identification results, we study the admissions to public and private secondary schools (grades 9-12) in Chile in 2007. The market is organized similarly to college admissions in the US. It is decentralized, both sides have preferences, and students do not submit rank-order lists of schools. We use a parametric Bayesian approach for preference estimation and then conduct counterfactual analysis.

\subsection{Institutional Background and Data}\label{sec:background}
Since 2003, secondary education has been compulsory for all Chileans up to 21 years of age. In principle, a public school must accept any student who is willing to enroll; a private school can be subsidized by the government or non-subsidized, but in either case, it can select students based on its preferences.

We focus on a relatively independent market, {\it Market Valparaiso}, that includes five municipalities (Valparaiso, Vi\~na del Mar, Concon, Quilpue, and Villa Alema\~na) as defined in \cite{gazmuri2017school}. Our data includes the municipality of a student's residence and the geographical coordinates of each school, which identifies every agent in the market.  A student is defined to be in Market Valparaiso if in 2008 she resided in a municipality within the boundary of the market. A secondary school is in Market Valparaiso if it is located in and admits students from the market.\footnote{There are 38 secondary schools located in Market Valparaiso that do not have any 9th graders from Market Valparaiso in 2007 and another 15 that have fewer than three 9th graders from Market Valparaiso on average in 2005, 2007, and 2009. We consider these 53 schools as outside options.} In total, there are $9,304$ students and $125$ schools. This reasonably large size makes it plausible that the continuum market in \cref{id} is a good approximation.

We use the SIMCE dataset, provided by \textit{La Agencia de Calidad de la Educaci\~on }(Agency for the Quality of Education), on all 10th graders in 2008 to identify who started secondary school  in 2007.  The SIMCE is Chile's standardized testing program and tracks students' math  and  language  performance. The data includes students' parental income,  parental  schooling,  and  other  characteristics  from  a  parental  questionnaire sent home with students. Most school attributes are calculated from student characteristics of {\it the 10th graders in a school in 2006}, and thus are pre-determined in the 2007 admissions that we study.
As tuition fees are largely fixed within each school and not completely flexible at the student level, we consider the problem as matching without transfers. See Online Appendix~\ref{app:data} for more details on data construction and summary statistics of student characteristics and school attributes.




    \subsection{Empirical Model} \label{sec:EmpModel}
We allow student preferences to be school-type-specific. For student~$i$, the utility of attending school~$c$ of type~$t$ $\in$ \{public, private non-subsidized, private subsidized\} is
\begin{eqnarray}\label{eq:stu_u}
  u_{ict} = \alpha_{ft} \times female_i  + \alpha_{mt} \times male_i + X'_{ic}\beta_t + \zeta_t \epsilon_{ic},
\end{eqnarray}
where $\alpha_{ft}$ is a school-type fixed effect for female students;  $female_i$ is a dummy variable for female;  $\alpha_{mt}$ and $male_i$ are similarly defined; $\epsilon_{ic}$ is  i.i.d.\ standard normal; $\zeta_t$ ($>0$) allows type-specific variances; and $X_{ic}$ are student-school-specific variables, including (i) the distance between $i$'s residence and school~$c$;
(ii) 6 school attributes, and (iii) 3 interactions between school attributes and student characteristics.

Each student has an outside option,  $u_{i0} = \epsilon_{i0}$, with $\epsilon_{i0}$ being standard normal. We impose the usual scale normalization through $\zeta_t=1$ for public school,\footnote{The variance of $\epsilon_{i0}$ is also normalized to be one because there is insufficient variation to estimate the variance of $u_{i0}$. Only 59 students (out of 9,304) choose an outside option (see \cref{tab:sumstu}).} while the location normalization is imposed by setting the deterministic part of $u_{i0}$ to zero.


School preferences are also type-specific, but public schools do not have a utility function because they cannot select students.  For private school~$c$ of type~$t$ (subsidized or non-subsidized), its acceptability threshold is $T_c=0$, and its utility function is
\begin{eqnarray}\label{eq:sch_u}
  v_{cit} = \theta_t + Z'_{ci}\gamma_t + \eta_{ci},
\end{eqnarray}
where $\theta_t$ is a type-specific intercept; $\eta_{ci}$ is i.i.d.\ standard normal; and the vector $Z_{ci}$ includes
(i) 5 student characteristics and (ii) 3 interactions between student characteristics and school attributes.

 In school preferences, the variance of $\eta_{ci}$ being one is the scale normalization and $T_c=0$ is the location normalization.  Allowing the type-specific intercept $\theta_t$ and assuming $T_c=0$  imply that schools of the same type have the same acceptability threshold. Because each type has some schools with vacancies, we can separately identify $\delta_c$ and $T_c$. Otherwise, we would lose its identification (\cref{np_id_F}).


The above specification uses  distance as an \textit{$(i,c)$-pair-specific} excluded variable in student preferences.  In school preferences, math and language scores are \textit{student-specific} excluded variables, implying that there are insufficient excluded variables to estimate school-specific utility functions. Using the results in \cref{deg-of-heter}, we limit preference heterogeneity and assume type-specific utility functions for schools. Our scale and location normalization is guided by \cref{np_id_F} and \cref{rm:location}, while certain normalization is imposed via $F$.

\subsection{Estimation and Results}\label{sec:results}

We use a Bayesian approach with a Gibbs sampler for the estimation, which is first illustrated in Monte Carlo simulations (Online Appendix~\ref{sec:mc_bayes}). We provide the details on the updating of the Markov Chain in Online Appendices~\ref{sec:mc_bayes} and \ref{app:bayes}.

The estimation results are summarized in \cref{tab:est_results}. A caveat is in order when we interpret the results. Because we do not deal with endogeneity issues that may arise due to the correlation between preference shocks and school attributes or student characteristics, the estimates  may not have a causal interpretation.

\begin{table}[h!]
	\centering \footnotesize
	\caption{Estimation Results: Student and School Preferences}   \label{tab:est_results}
	\resizebox{\textwidth}{!}{
		\begin{tabular}{lcccccccc}
			\toprule
			& \multicolumn{2}{c}{} & & \multicolumn{5}{c}{Private Schools}  \\
			\cline{5-9}
			& \multicolumn{2}{c}{Public schools} &&  \multicolumn{2}{c}{Subsidized} &&  \multicolumn{2}{c}{Non-subsidized} \\
			\cmidrule{2-3}     \cmidrule{5-6}     \cmidrule{8-9}
			& \multicolumn{1}{c}{{coef.}}  & \multicolumn{1}{c}{{s.e.}} &&  {coef.} & {s.e.} & &  {coef.} & {s.e.} \\
			\midrule
			\textit{\textbf{Panel A. Student Preferences}} &       &       &       &       &       & & &  \\
			Female & \multicolumn{1}{c}{0.458} & \multicolumn{1}{c}{(0.834)} &       & 10.500 & (0.611) &       & 9.848 & (5.631) \\
			Male  & \multicolumn{1}{c}{0.442} & \multicolumn{1}{c}{(0.834)} &       & 10.387 & (0.610) &       & 9.089 & (5.628) \\
			Distance & \multicolumn{1}{c}{-0.175} & \multicolumn{1}{c}{(0.004)} &       & -0.145 & (0.008) &       & -0.624 & (0.070) \\
			log(tuition) & \multicolumn{1}{c}{0.577} & \multicolumn{1}{c}{(0.119)} &       & -0.385 & (0.102) &       & -15.645 & (1.567) \\
			log(tuition) $\times$ log(income) & \multicolumn{1}{c}{-0.019} & \multicolumn{1}{c}{(0.010)} &       & 0.044 & (0.008) &       & 0.881 & (0.084) \\
			log(median income) & \multicolumn{1}{c}{-0.307} & \multicolumn{1}{c}{(0.072)} &       & -0.720 & (0.067) &       & 2.454 & (0.544) \\
			Teacher experience & \multicolumn{1}{c}{0.003} & \multicolumn{1}{c}{(0.002)} &       & 0.005 & (0.001) &       & 0.020  & (0.015) \\
			Fraction of female students & \multicolumn{1}{c}{-0.010} & \multicolumn{1}{c}{(0.034)} &       & -0.320 & (0.054) &       & -1.428 & (0.773) \\
			Average composite score & \multicolumn{1}{c}{-1.948} & \multicolumn{1}{c}{(0.124)} &       & -2.354 & (0.132) &       & -16.104 & (2.241) \\
			Average composite score $\times$ Composite score & \multicolumn{1}{c}{6.282} & \multicolumn{1}{c}{(0.232)} &       & 6.209 & (0.237) &       & 18.167 & (1.616) \\
			Average mother's education & \multicolumn{1}{c}{0.038} & \multicolumn{1}{c}{(0.023)} &       & -0.283 & (0.024) &       & -1.572 & (0.298) \\
			Average mother's education $\times$ Mother's education & \multicolumn{1}{c}{0.007} & \multicolumn{1}{c}{(0.001)} &       & 0.011 & (0.001) &       & 0.056 & (0.007) \\
			Standard deviation of the utility shock & \multicolumn{2}{c}{Normalized to 1} &       & 1.093 & (0.063) &       & 8.646 & (0.859) \\
			[0.5em]
			\textit{\textbf{Panel B. School Preferences}} &       &&        &       &  &      &       &  \\
			Constant &       &       &       & 3.686 & (0.852) &       & 20.283 & (1.863) \\
			Female &       &       &       & -3.545 & (0.193) &       & -1.438 & (0.274) \\
			Female $\times$  Fraction of female &       &       &       & 6.892 & (0.333) &       & 1.968 & (0.443) \\
			Math score &       &       &       & -3.251 & (0.541) &       & -7.053 & (0.920) \\
			Average math score $\times$ Math score &       &       &       & 7.955 & (0.938) &       & 6.217 & (0.985) \\
			Language score &       &       &       & -0.030 & (0.460) &       & -4.593 & (0.859) \\
			Average language score $\times$ Language score &       &       &       & 1.053 & (0.822) &       & 5.087 & (1.135) \\
			Mother's education &       &       &       & -0.010 & (0.016) &       & -0.136 & (0.047) \\
			log(income) &       &       &       & -0.225 & (0.072) &       & -1.030 & (0.146) \\
			\bottomrule
	\end{tabular}}
	\begin{tabnotes}
		This table presents the posterior mean and standard deviation of each coefficient in student and school utility functions (equations~\ref{eq:stu_u} and \ref{eq:sch_u}). The Bayesian approach goes through a Markov Chain 1.75 million times, and the last 0.75 million iterations are used to calculate these statistics. Our Monte Carlo simulations suggest that the posterior standard deviation provides a useful estimate of the standard deviation of our estimators, albeit with a slight tendency to underestimate. See \cref{tab:mc_bayes} for more details.
	\end{tabnotes}
\end{table}


Panel A shows the estimates of student preferences. Most coefficients are of an expected sign. Interestingly, the coefficient on tuition is positive in the utility function for public schools, and parental income pushes it towards negative. This may reflect the fact that tuition at public schools is generally low (see \cref{tab:sumsch}) and may be correlated with unobserved school quality.
There are also a few coefficients with an unexpected sign in school preferences.  Panel B shows that non-subsidized schools negatively value a student's parental income and mother's education. This may be because, on average, students at those schools often have a high income (above 1.4 million CLP; see \cref{tab:sumstu}) and a highly educated mother (above 17 years).

Recall that each school's acceptability threshold is normalized to zero, so we can use the estimated school preferences to calculate if a student is acceptable to a private school. On average, a subsidized school finds 79\% of the students acceptable, while a non-subsidized school finds 84\% acceptable.  The higher acceptability rate at non-subsidized schools does not imply that students are more often matched with them because their high tuition lowers their desirability to many students, especially those with a low parental income (see Table~\ref{tab:est_results}).

Before conducting counterfactual analysis, we evaluate model fit in Online Appendix~\ref{app:bayes}. It shows that our model fits the data reasonably well when we compare the observed matching with the one predicted based on our model.

\subsection{Counterfactual: Prioritizing Low-income Students}

We consider a counterfactual policy in which students from low-income families are prioritized for admissions to all schools. A student is of low income if her parental income is among the lowest 40\%.\footnote{This resembles a policy adopted in 2008 in Chile as documented by \cite{gazmuri2017school}. It benefits 44\% of elementary school students in 2012 in terms of admission priorities and a tuition waiver. Online Appendix \cref{sumstu_byinc} shows summary statistics of the students by income status.}
Each private school's preferences over students are made lexicographical: low-income students are above others, and within each group of students, a school ranks them as in the current regime; all low-income students are acceptable, while others' acceptability is the same as in the current regime. We do not change anything in student or school preferences beyond the admission priority, which may mitigate the potential bias in our estimation due to aforementioned unaddressed endogeneity issues.


To simulate the counterfactual outcome, we keep 1,500 draws of the utilities in the Markov Chain in the Bayesian estimation.\footnote{{Specifically, there are 15 blocks of 100 draws.  The blocks are equally spaced in the 0.75 million iterations in the Markov chain that are used to calculate the posterior means and standard deviations.}} We run the Gale-Shapley deferred acceptance with each draw and obtain 1,500 sets of counterfactual stable matchings. We then report the average of these counterfactual matchings; recall that we only have one matching outcome under the current regime---the observed one.



\begin{table}[h!]
  \centering \footnotesize
  \caption{Sorting and Student Welfare in the Current and Counterfactual Regimes}  \label{tab:counterfactual}
\resizebox{0.9\textwidth}{!}{
    \begin{tabular}{lccccc}
\toprule
    & \multicolumn{2}{c}{Low-income students} &          & \multicolumn{2}{c}{Non-low-income students}  \\
\cline{2-3} \cline{5-6}
& {Current} & {Counterfactual} &       & {Current} & {Counterfactual} \\
& (1) & (2) &       & (3) & (4) \\
\midrule
\cmidrule{1-3}\cmidrule{5-6}
        Average composite score (same cohort) at matched school  & 0.374 & 0.385 &       & 0.585 & 0.576 \\
        Average parental income (same cohort)  at matched school & 216,389
         & 223,003
          &       & 592,848
           & 587,884
            \\
    [0.1em]
    {\it Fraction enrolled at each school type:} & & & & & \\
    \multicolumn{1}{r}{Public} & 0.679 & 0.602 &       & 0.233 & 0.260 \\
    \multicolumn{1}{r}{Private subsidized} & 0.315 & 0.392 &       & 0.533 & 0.504 \\
    \multicolumn{1}{r}{Private non-subsidized} & 0.002 & 0.002 &       & 0.227 & 0.228 \\
    \multicolumn{1}{r}{Outside option} & 0.004 & 0.003 &       & 0.008 & 0.008 \\
         [0.1em]
    {\it Welfare effects of moving from current to counterfactual:} & & & & & \\
     \multicolumn{1}{r}{Average utility change (reduction in distance, km)} &       \multicolumn{2}{c}{0.332}   &       & \multicolumn{2}{c}{$-$0.225} \\
     \multicolumn{1}{r}{Winners (fraction)} &      \multicolumn{2}{c}{0.120} &       &       \multicolumn{2}{c}{0.000} \\
     \multicolumn{1}{r}{Losers (fraction)} &       \multicolumn{2}{c}{0.000}       &    &   \multicolumn{2}{c}{0.072} \\
     \multicolumn{1}{r}{Indifferent (fraction)} &       \multicolumn{2}{c}{0.881} &       &       \multicolumn{2}{c}{0.927} \\
\bottomrule
\end{tabular}}
    \begin{tabnotes}
The outcome in the current regime is the one observed in the data. To simulate the counterfactual outcome, we keep 1,500 draws of the utilities in the Markov Chain in the Bayesian estimation and obtain 1,500 sets of stable matching outcomes. The statistics for the counterfactual regime are averages across the 1,500 outcomes. The average utility change is measured in terms of willingness to travel to a public school in kilometers.
    \end{tabnotes}
\end{table}


\cref{tab:counterfactual} presents the results. There are several noticeable patterns when we move from the current regime to the counterfactual.
First, low-income students are in schools that have higher-ability and higher-income students in the same cohort, while the opposite is true for non-low-income students.  Second, some low-income students leave public schools for private subsidized schools, while crowding out some other students to public schools. Lastly, the policy benefits low-income students and hurts others. On average, low-income students' welfare gain is equivalent to decreasing travel distance (to a public school) by 0.332 km. This gain is concentrated among 12\% of the low-income students, while others are not affected. Correspondingly, 7.2\% of the non-low-income students are worse off, and none is better off.


These results indicate that low-income students dislike private non-subsidized schools. We explore why it is the case. With the 1,500 draws of student preferences, we examine low-income students' favorite school of each type. We find that low-income students on average value their favorite public school at $2.72$ and private subsidized school at $2.56$, while their favorite non-subsidized school is only valued at $-21.76$, in general unacceptable to them.  We calculate the contribution of different variables to these differences by shutting down their effect in the utility functions. The results show that high tuition at non-subsidized schools is the main contributor. It is worth noting that tuition could be correlated with unobserved school quality. Hence, low-income students may dislike a non-subsidized school because of its high costs and/or their tastes. Additionally, mother's education, both a student's own and a school's average, is also an important factor.  Student ability, measured by their composite score, and distance to each school do not appear to be as important.


In sum, giving low-income students access to schools fails to significantly change matching outcomes due to their own preferences. Low-income students are deterred from private non-subsidized schools by their high tuition. These findings are in line with the preference heterogeneity documented in public school choice \citep[see, e.g.,][]{abdulkadiroglu_welfare_2017,kapor2020heterogeneous},  although tuition  plays no role there.


\section{Concluding Remarks}\label{sec:conclusion}
    We study nonparametric identification of agent preferences in many-to-one two-sided matching without transfers. We derive a set of sufficient conditions for identification and provide guidance for empirical studies. For example, our results clarify the data requirement for the identification of various degrees of preference heterogeneity. To take our results to the data of a reasonably sized market, we propose a Bayesian approach with a Gibbs sampler whose performance is illustrated in Monte Carlo simulations.  Our model encompasses many real-life matching markets, such as college admissions and school choice in many countries, in which our identification results and empirical method can be applied. Hence, this paper opens a new avenue for empirical research.

We illustrate our method in the context of secondary school admissions in Chile. As an example of the usefulness of the estimates, we consider a counterfactual policy in which students from low-income families are prioritized for admissions to all schools. Although the policy benefits low-income students, its effects are small. Such insights are difficult to obtain without estimating the preferences of both sides. In this sense,  our method can help provide an ex-ante evaluation of a range of alternative policies.