Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
96,277 characters · 16 sections · 58 citation commands
Identification and Estimation in Many-to-one Two-sided Matching without Transfers
\setstretch{1.05}
\setstretch{1.5}
\setstretch{1.3}
In a many-to-one two-sided matching market, agents are categorized into two sides; everyone on one side has preferences over those on the other side; an agent on only one of the two sides can have multiple match partners from the other side. Many real-life markets fit this description, for example, the medical resident match roth1984evolution,agarwal_empirical_2015 in the US, school admissions in Chile gazmuri2017school and Hungary aue2020happens, college admissions in the US, and graduate program admissions in France he2020application. Such markets often exclude personalized transfers, even though limited monetary exchanges may exist. Hence, the literature defines it as matching without transfers or matching with non-transferable utility.
While the literature has extensively studied this type of matching theoretically roth1992two,azevedo_supply_2016, its econometrics is less explored. Our paper aims to make a contribution by answering the following questions: Are the preferences of both sides identified from data on who matches with whom? If so, how can the preferences be estimated?
To fix ideas, we proceed in the language of college admissions. We derive a set of sufficient conditions under which both student and college preferences are nonparametrically identified. Our results are obtained from a single market in which there are a continuum of students and a fixed number of colleges. We use that to approximate a single large market. Further, we provide an estimation procedure that is practical even in settings with many agents, allowing for rich observed and unobserved heterogeneity. Understanding agent preferences is often crucial for policymaking, and one may analyze a wide range of counterfactual policies with estimated preferences. Potentially, our results open a new avenue of research on such matching markets.
The main challenge in identifying student preferences is that each student's {\it actual} choice set is unobservable to the researcher. For student $i$ to be able to enroll at college $c$, college $c$ needs to accept $i$. The same difficulty exists in the identification of college preferences. Moreover, each student's and each college's choice sets are endogenously determined in equilibrium without market-clearing prices.
In our continuum setting, we assume that an observed matching is stable. That is, no college prefers to reject any of its currently matched students to vacate a seat, and no student prefers to leave her current match to become unmatched or matched with a college that is willing to accept her and, if necessary, reject one of its currently matched students. Stability is often imposed in the study of various matching markets chiappori_econometrics_2016 and is satisfied in equilibrium in our setting in certain game-theoretical models ACH2017,fack_beyond_2019.
Importantly, there is generically a unique stable matching that is characterized by the colleges' admission cutoffs azevedo_supply_2016. When college preferences over individual students are represented by utility functions, a college's cutoff is the lowest utility level among its matched students. Cutoffs further define a student's actual choice set in equilibrium, called {\it feasible set}. A college is in a student's feasible set if the college's utility of being matched with her is higher than its cutoff. Stability implies that a student is matched with her most-preferred feasible college, similar to a discrete choice problem, except that feasible sets are unobservable and heterogeneous.
A simple equation, called the $i$-$c$ match probability, is the key to understanding our identification result. Specifically, the conditional probability of student $i$ being matched with college $c$ is the sum of conditional probabilities of $i$ choosing $c$ from a given feasible set $L$ weighted by the conditional probability of facing $L$: \vskip-0.8cm
where $x_{i}$ consists of all observed characteristics of student $i$ (e.g., pair-specific characteristics like distance to colleges). The equation provides a decomposition of the preferences of the two sides: for each given $L$, piece $A$ only depends on the preferences of {\it all} colleges given cutoffs, while $B$ only depends on $i$'s preferences over all colleges.
We then detail a set of exclusion restrictions, among other regularity conditions, such that the excluded variables act as “demand shifters” and “feasible-set shifters” (or supply shifters). Sufficient variation in these excluded variables identifies the preferences of colleges and students using the $i$-$c$ match probability described above.
Here are some intuitions. For a college $d$, an $(i,d)$-specific demand shifter traces out how $i$'s preferences for $d$ affects the $i$-$c$ match probability. Similarly, an $(i,d)$-specific feasible-set shifter traces out how $d$'s preference for $i$ affects the $i$-$c$ match probability. A non-excluded variable affects the $i$-$c$ match probability through preferences on both sides for all colleges. By taking derivatives of the $i$-$c$ match probability with respect to (w.r.t.) all the variables, excluded and non-excluded, we derive systems of linear equations that link the effects of variations in demand and supply. Hence, the identification problem reduces to setting up systems of linear equations and ensuring the existence of a unique solution.
The second objective of our paper is to provide practical methods that can be used to analyze real-life markets. We achieve this by deriving theoretical guidelines and showcasing a practical estimation method.
When taken to the data, the requirement of a large number of excluded variables may be difficult to meet. To address this, we theoretically characterize the tradeoff between exclusion restrictions and the degree of identifiable preference heterogeneity ((ref)). The researcher can use this result as a guideline for empirical studies when having insufficient excluded variables.
We also need a practical estimation method to take these identification results to the data. In fact, our identification arguments are constructive, leading to nonparametric and semiparametric estimators. Monte Carlo simulations suggest that estimating the matrices of partial derivatives in the linear systems using the average derivative estimators of powell_semiparametric_1989 performs well in finite samples only when the curse of dimensionality is not severe. In a reasonably sized problem, we resort to a parametric Bayesian approach with a Gibbs sampler rossi2012bayesian, resembling applications such as logan_two-sided_2008 for one-to-one two-sided matching and abdulkadiroglu_welfare_2017 for a one-sided problem. We demonstrate its good performance in Monte Carlo simulations with high dimensionality.
As an empirical application, we consider the decentralized admissions to secondary schools in Chile. To the best of our knowledge, this is {one of the first attempts to estimate the preferences of both sides in a {\it decentralized} market of many-to-one two-sided matching without transfers.} There is no clearinghouse, and students do not submit rank-order lists of schools. By allowing flexible preference heterogeneity in the Bayesian approach, we estimate the preferences of students and schools. We also consider a counterfactual policy in which students from low-income families are prioritized for admissions to all schools. Segregation in terms of ability and income decreases, albeit slightly. We find that simply giving low-income students access to schools may not significantly change matching outcomes due to student preferences.
\paragraph{Related Literature.} This paper is related to the literature on the identification of matching models; (ref) provides an incomplete summary.
This literature is split into several strands depending on the preference structures of the agents -- transferable utility (TU) and non-transferable utilities (NTU);\footnote{See galichon2019costly for an example of imperfectly transferable utility (ITU) models.} and the maximum number of links an agent is permitted to form across sides -- one-to-one, many-to-one, and many-to-many chiappori_econometrics_2016.
There is also a close relationship between the one-to-one TU matching model choo_who_2006,fox_identification_2010,fox_estimating_2018,GRAHAM2011965,sinha_identification_2015,chiappori2017partner,diamond_latent_2017,galichon_cupids_2020,gualdani_identification_2020 and the many-to-one NTU matching model considered here. Market-clearing college cutoffs in our setting play the role of market-clearing shadow prices, although the endogenous cutoffs do not determine how the surplus is split among the agents.
Most of the work on identification within the NTU framework focuses on one-to-one markets dagsvik_aggregation_2000,menzel_large_2015,uetake2020entry. Allowing for infinitely many agents on {\it both} sides of a many-to-one matching market, {agarwal_empirical_2015} and diamond_latent_2017 prove identification under a homogeneity restriction on the preferences,\footnote{When studying the medical resident match, agarwal_empirical_2015 discusses some intuitions of using exclusion restrictions to identify heterogeneous preferences on each side of the market.} and ederer2022 shows identification by relaxing the homogeneity assumption but still restricting preferences. as_2020 provide a recent survey on empirical models of NTU matching.
Many-to-one NTU matching has been empirically studied in the context of secondary school admissions in Hungary aue2020happens and graduate program admissions in France he2020application. Their data include information on the preferences of both sides that is reported to a centralized mechanism. Therefore, they can independently identify and estimate the preferences of each side, essentially reducing the two-sided matching to two separate one-sided problems.
Centralized many-to-one NTU matching in the context of school choice has been studied extensively, both theoretically since abdulkadiroglu_school_2003 and empirically abdulkadiroglu_welfare_2017,agarwal_demand_2018,calsamiglia2020structural,fack_beyond_2019,he_gaming_2017,kapor2020heterogeneous. In this literature, school preferences are (assumed to be) known because schools rank students according to certain pre-specified rules. The problem then reduces to identifying and estimating student preferences. See agarwal2020revealed for a survey.
Feasible sets in our setting resemble endogenous consideration sets that arise in one-sided decision problems. In this sense, our paper relates to the growing strand of literature that studies the econometrics of decision problems under consideration set formation abaluck_what_2018, barseghyan_heterogeneous_2019, barseghyan2021discrete,cattaneo2020random. Our contribution here is that we provide a structural two-sided setting where the consideration probabilities in a student's decision problem are entirely determined by the college (supply side) preferences. Along similar lines to ours, agarwal2022demand study consumer choice models with latent choice-set constraints. Their identification conditions and ours are non-nested (see Section (ref)).
The remaining paper is organized as follows: (ref) describes the model and data generating process; (ref) discusses the identification of preferences of the agents on both sides of the market; (ref) illustrates an empirical analysis of the match between students and secondary schools in Chile; and (ref) concludes.
For the sake of exposition, our model is set up as a college admissions problem. Consider a single market with a continuum of students and finitely many colleges. The set of all students is $\textbf{I}$, with a probability measure $Q$ defined over it,\footnote{Our probability space is $(\mathbf{I},\mathcal{B}(\mathbf{I}),Q)$ where $\mathcal{B}(\mathbf{I})$ is the Borel set of $\mathbf{I}$, $Q:\mathcal{B}(\mathbf{I})\rightarrow[0,1]$, $Q(\mathbf{I})=1$.} and the set of all colleges is $\mathbf{C}=\{1,2,\dots,C\}$. College $c \in \mathbf{C}$ has a capacity $q_{c}\in (0,1)$.
For $i\in\mathbf{I}$, the utility of being matched with college $c$ is $u_{ic}$; for $c\in\mathbf{C}$, the utility of being matched with $i$ is $v_{ci}$. To prepare for our identification results in (ref), we assume additive separability and excluded variables in the utility functions,\footnote{See Online Appendix (ref) for our identification results in a more general nonseparable utility specification.} although the rest of the current section applies to more general models.
The utilities $u_{ic}$ and $v_{ci}$ depend on the vector $(z_{i},y_{i},w_{i})$, which is observable to the researcher, and $\epsilon_{i} =(\epsilon_{i1}\dots,\epsilon_{iC})$ and $\eta_{i} =(\eta_{1i}\dots,\eta_{Ci})$, which are unobservable to the researcher. Specifically, $z_{i}\in\mathcal{Z}\subseteq\mathbb{R}^{d_{z}}$, $y_{i}=(y_{i1}\dots,y_{iC})\in\mathcal{{Y}}= \mathcal{Y}_1 \times \dots \times \mathcal{Y}_C \subseteq\mathbb{{R}}^{C}$, and $w_{i}=(w_{1i}\dots,w_{Ci})\in\mathcal{{W}}\subseteq\mathbb{{R}}^{C}$. Further, for any $c\in\mathbf{C}$, $\epsilon_{ic}$ and $\eta_{ci}$ are scalar random variables. $(\epsilon_{i},\eta_{i})$ are independent and identically distributed (i.i.d.) draws from a joint distribution $F$. We make no restrictions on this joint distribution, allowing for arbitrary correlations within $(\epsilon_{i},\eta_{i})$.
Scalar $y_{ic}$ is a demand shifter and scalar $w_{ci}$ is a supply shifter: $y_{ic}$ enters only $u_{ic}$ and is excluded from all other utility functions, and $w_{ci}$ enters only $v_{ci}$.
We impose scale normalization on each side. For students, there exists a known value $\overline{y}_{c}$ in the interior of its support such that $\frac{\partial r^{c}\left(\overline{y}_{c}\right)}{\partial y_{ic}}=1$.\footnote{If $y_{ic}$ is known to have a negative effect on student preferences, the partial is normalized to $-1$. The same applies to $w_i$. } This holds trivially if $r^{c}\left(y_{ic}\right)=y_{ic}$. {For colleges, we assume $w_{ci}$ enters $v_{ci}$ linearly with a coefficient normalized to one. As detailed below, this linearity assumption allows us to vary $w_i$ to construct a sufficient number of equations without increasing the number of unknowns.\footnote{ The linearity in $w_{ci}$ can be relaxed. We can allow $w_{ci}$ to enter $v_{ci}$ through a nonlinear function, which is known for some colleges. Namely, assume $v_{ci} = v^{c}(z_{i})+s^c(w_{ci})+\eta_{ci}$, where $s^c$: $\mathcal{W}_{c}\to\mathbb{R}$ is a nonparametric function such that (i) for each $c\in \overline{\mathbf{C}} \subset \mathbf{C}$, $s^c$ is unknown with $\frac{\partial s^{c}\left(\overline{w}_{c}\right)}{\partial w_{ci}}=1$ for some known value $\overline{w}_{c}$; and (ii) for each $c\in \mathbf{C} \setminus \overline{\mathbf{C}}$, $s^c$ is known. In this specification, our identification relies on varying $w_{ci}$ for $c\in \mathbf{C} \setminus \overline{\mathbf{C}}$ instead of all $c\in \mathbf{C}$.} The assumptions on $y_i$ and $w_i$ can be switched, that is, having nonlinearity in $w_i$ and linearity in $y_i$.}
Students can remain unmatched or, equivalently, be matched with an outside option denoted by “$0$.” The utility of the outside option is normalized to 0, $u_{i0}=0$ $\forall i\in\mathbf{I}$. We assume that colleges have responsive preferences.\footnote{For $\varepsilon>0$, let $N_\varepsilon (i)$ and $N_\varepsilon (i')$ be a neighborhood of students around $v_{ci}$ and $v_{ci'}$, respectively, such that $Q(N_\varepsilon (i)) = Q(N_\varepsilon (i'))$. Responsive preferences imply that, for any $\mathbf{I}^c \subset \mathbf{I}$ with $Q(\mathbf{I}^c) \leq q_c - Q(N_\varepsilon (i))$, $N_\varepsilon (i) \subset \mathbf{I}\setminus \mathbf{I}^c$, and $N_\varepsilon (i') \subset \mathbf{I}\setminus \mathbf{I}^c$, college $c$ prefers $\mathbf{I}^c \cup N_\varepsilon (i) $ to $\mathbf{I}^c \cup N_\varepsilon (i') $ if and only if $v_{ci} > v_{ci'}$. See roth1992two for a definition in a case with discrete students.} This implies that the total utility of a college from being matched with a subset of students (up to its capacity) is increasing in its utility from each student; for example, the total utility is the sum of the utility from each of its matched students. College $c$ has an acceptability threshold, $T_c$, and finds student $i$ unacceptable if $v_{ci}<T_c$.
With the data from one such continuum market on $\{(z_i,y_i,w_i)\}_{i}$, $\{q_c\}_{c}$, and who matches with whom, we aim to identify student and college preferences by identifying $\{u^{c}, r^c, v^{c}, T_c\}_{c}$ and $F$, although, as we shall see, $\{T_c\}_{c}$ are not always point identified. Note that $(z_i,y_i,w_i)$ does not include college-specific variables that are constant across students, as these will be absorbed by the college-specific utility functions, $(u^c,r^c,v^c)$.
We use the continuum market to approximate a data generating process in a large finite market as follows: $(z_i, y_i, w_i,\epsilon_{i},\eta_i)$ is an i.i.d.\ draw from its joint distribution, college $c$'s capacity is a $q_c$-fraction of the total number of students, and $\{u_{ic}, v_{ci},T_c\}_{i,c}$ determines both sides' preferences. This approximation is close to the matching outcomes when we use the equilibrium concept that will be introduced in (ref).\footnote{In a typical identification argument, the number of observations is taken to infinity while the “game” is kept constant. Our setting contains one single matching game that changes with the market size. Proposition 3 of azevedo_supply_2016, Proposition 4 of fack_beyond_2019, and Corollary 2 of ACH2023 imply that, under certain conditions, the equilibrium outcome in the continuum approximates well an equilibrium outcome in a large finite market. Such an approximation is also used in the network literature menzel2022strategic.}
We define a matching function or, simply, a matching, $\mu:\mathbf{{I}}\to\mathbf{{C}}\cup\{0\}$, such that (i) $\mu(i) = c \iff i \in\mu^{-1}(c)$, and (ii) $\forall\,c\in\mathbf{{C}},\,\mu^{-1}(c)\subseteq{\bf {I}},$ where $0 \leq Q(\mu^{-1}(c))\leq q_{c}$.
The following concepts are important for our analysis: individual rationality, blocking pairs, and stability. For notational reasons, we define them in the case with discrete students, corresponding to our empirical application. The precise definitions for a model with a continuum of students can be found in azevedo_supply_2016, with measure-zero sets of students appropriately dealt with.
A matching $\mu$ is {\it individually rational} if $u_{i\mu(i)} \geq u_{i0}$ and $v_{\mu(i)i} \geq T_{\mu(i)}$ for all $i \in \mathbf{I}$. A student-college pair $(i,c)\in\mathbf{{I}}\times\mathbf{C}$ blocks a matching $\mu$ if (i) student $i$ strictly prefers college $c$ to her current match $\mu(i)$, $u_{ic}>u_{i\mu(i)}$; and (ii) either college $c$ has excess capacity, $ Q(\mu^{-1}(c))<q_{c}$, or college $c$ prefers $i$ to one of its matched students, $\exists\,i^{\prime}\in\mu^{-1}(c),\,\text{{s.t.}},\,v_{ci}>v_{ci^{\prime}}$. Finally, a matching is stable if it is individually rational and not blocked by any pair $(i,c)\in\mathbf{{I}}\times\mathbf{{C}}$.
We assume that the matching in the data is stable.\footnote{Stability can be achieved in certain equilibrium if students apply to all acceptable colleges and if a stable mechanism, e.g., the deferred acceptance gale_college_1962, is used to find the matching. Theoretically, provided that students know what criteria colleges use to rank them, stability can still be satisfied in equilibrium, even if students choose not to apply to all acceptable colleges due to application costs fack_beyond_2019 or if students make certain application mistakes ACH2017. Importantly, achieving stability does not require the market to be centralized, as shown in laboratory experiments Pais-decentralized, and steps in mechanisms such as the deferred acceptance can be implemented in a decentralized fashion GHK. } A stable matching exists and is generically unique azevedo_supply_2016.\footnote{We need the regularity condition that the set $\{i\in\mathbf{I}:\mu(i)\,\text{is strictly less preferred than}\,c\}$ is open $\forall\, c\in \mathbf{C}$. This condition implies that a stable matching always allows an extra measure zero set of students into a college when this can be done without compromising stability.} Moreover, a stable matching is characterized by college cutoffs. College $c$'s cutoff is determined by its least-preferred matched student when its capacity constraint is binding; otherwise, it coincides with the acceptability threshold. Let $\delta_{c}$ be college $c$'s cutoff. Then, $\forall\,c\in\mathbf{{C}}$,
By definition, $\delta_{c} \geq T_c$. Under the assumption of responsive college preferences, to determine if student $i$ can be accepted by $c$, we just need to compare $v_{ci}$ and $\delta_{c}$. With non-responsive preferences, how $c$ ranks $i$ and $j$ would depend on who else $c$ accepts, and $\delta_{c}$ alone would not be sufficient to determine if $i$ could have been accepted by $c$.
In a stable matching of the continuum market, $\{u^{c}, r^{c},v^{c},T_c\}_{c\in\mathbf{C}}$, and $F$ imply a unique vector of cutoffs, $\{\delta_c\}_c$. Therefore, $\{\delta_c\}_c$ is merely a shorthand notation for the expression in equation (ref) rather than additional parameters.\footnote{As mentioned earlier, especially, in footnote (ref), the continuum approximates a large finite market. The literature cited therein shows that equilibrium cutoffs, hence matching outcomes, in the large market can be close to $\{\delta_c\}_c$ with agents being practically “cutoff-takers.” That is, each agent's realized preferences have a negligible effect on cutoffs in large markets.}
College $c$ is said to be feasible for student $i$ if and only if $v_{ci} \geq \delta_c$. A student can “choose” to match with any of her feasible colleges, but not any infeasible college. We call the set of all feasible colleges of a student her {\it feasible set}. Let $\mathcal{{L}}$ be the collection of the $2^{C}$ possible feasible sets, $\mathcal{{L}}\equiv\{L:\,0\,\in L,\,L\setminus\{0\}\subseteq\mathbf{{C}}\}$. By construction, the outside option always belongs to every feasible set. A matching is stable if and only if every student is matched with her most-preferred feasible college fack_beyond_2019. A classic issue in two-sided matching is that students' feasible sets are determined endogenously, unobserved by the researcher, and heterogeneous across students.
For any given matching $\mu$, the probability that $L\in\mathcal{{L}}$ is student $i$'s feasible set conditional on $(z_i,w_i)$ is
If each student's feasible set was observed, $ \lambda_L$ could be identified from the data, and recovering student and college preferences would follow from standard arguments in the discrete choice literature. However, we do not observe the feasible sets.
Similarly, for students, $\mathbb{{P}}\left(c=\arg\max_{d\in L}u_{id}|L, z_{i}, y_{i}\right)$ is the probability that utility-maximizing students with observables $(z_{i}, y_{i})$ “choose” $c$ from feasible set $L$. Conditional on $L$, this probability only depends on student preferences. We define
where the first argument of $g_{c,L}$ is always $\tau_{ic}$. If $d\notin L$, $g_{c,L}$ does not vary with $\tau_{id}$.
We now turn to nonparametrically identifying the distribution of student and college preferences given a stable matching $\mu$ and covariates $(z_{i},y_{i},w_{i})$ in {\it one} market.
A matching is stable if and only if every student is matched with the most-preferred college in her feasible set. Thus, stability implies, for $c\in\mathbf{{C}}\cup\{0\}$,
Hence, the conditional match probability, $\sigma_{c}(z_{i},y_{i},w_{i})$, which is the fraction of students with $(z_{i},y_{i},w_{i})$ matched with $c$ and known in the population data, is linked to the model through student and college preferences.
Below, we first study the conditions under which the functions $\{u^{c},\,r^{c},\,v^{c}\}_c$ are nonparametrically identified; we then identify the joint distribution of $(\epsilon_{i},\eta_{i})$, $F$. Later in Section (ref), we present results that impose fewer requirements on the data.
We nonparametrically identify the derivatives of the functions $\{u^{c},r^{c}, v^{c}\}_c$ w.r.t.\ the observables. With these derivatives identified, the functions are identified up to a constant, provided that $z_{i}$ and $y_i$ have full support. The idea is to use the variation in the excluded variables to trace out how each argument in equation (ref) affects the conditional match probabilities. Specifically, the excluded variables in student preferences $(y_i)$ only shift demand, while the excluded variables in college preferences $(w_i)$ shift supply, or feasible sets. The effect of other variables $(z_i)$ that enter both demand and supply can be written as a combination of the effects of $y_i$ and $w_i$. This leads to a system of linear equations in the derivatives of the conditional match probabilities, whose solution is our parameters of interest.
We describe the intuition for identification in a one-college example, $\mathbf{C}=\{1\}$. Student utility functions are $u_{i1}=u^1\left(z_{i}\right)+r^1\left(y_{i1}\right)+\epsilon_{i1}$ for college $1$ and $u_{i0}=0$ for the outside option. Here, $z_{i}$ is a scalar and $\frac{\partial r^{1}\left(\overline{y}_{1}\right)}{\partial y_1}=1$ for a known value, $\overline{y}_{1}$. College $1$'s utility function is $v_{1i}=v^1(z_{i})+w_{1i}+\eta_{1i}$, and the (unobserved) cutoff is $\delta_1$.
To identify $\frac{\partial u^1}{\partial z_i}$ and $\frac{\partial v^1}{\partial z_i}$, we fix $y_{i1}=\overline{y}_1$ and consider any value $(z,w_1)$ in the interior of its support. Figure (ref)(a) shows that the space of $(\epsilon_{i1},\eta_{1i})$ is partitioned into four parts based on the feasibility of college 1 and student $i$'s preferences (i.e., the acceptability of college $1$ to $i$). Moreover, $\mu\left(i\right)=1$ if and only if $\epsilon_{i1}>-u^1\left(z\right)-r^1\left(\overline{y}_1\right)$ (college $1$ is acceptable to $i$) and $\eta_{1i}>\delta_1-v^1\left(z\right)-w_1$ (college $1$ is feasible to $i$).
\pgfplotsset{ticks=none}
Figures (ref)(b)--(d) depict how the marginal effect of $z_{i}$ on match probability is linked to the marginal effects of the excluded variables, $y_{i1}$ and $w_{1i}$. Panel (b) describes the marginal effect of $y_{i1}$. Specifically, decreasing $y_{i1}$ from $\overline{y}_1$ to $\overline{y}_1-\Delta y$ makes college $1$ less attractive to student $i$, and the region in which $\mu(i)=1$ shrinks along the horizontal $\epsilon_{i1}$-axis. The area $I_{1}$ depicts the set of students whose match differs when $y_{i1}$ decreases. The induced change in the match probability, $\sigma_1=\mathbb{P}\left(\mu(i)=1|(z_{i},y_{i1},w_{1i}) = (z,\overline{y}_1,w_1)\right)\equiv\mathbb{P}(\mu(i)=1|z,\overline{y}_1,w_1)$, is the mass that the density of $\left(\epsilon_{i1},\eta_{1i}\right)$ puts on $I_{1}$, or, for $(z_{i},y_{i1},w_{1i}) = (z,\overline{y}_1,w_1)$,
where $\tau_1 \equiv u^1(z)+r^1(\overline{y}_{1})$ and $\iota_{1} \equiv v^1(z)+w_{1}$. By scale normalization, $\frac{\partial r^{1}\left(\overline{y}_{1}\right)}{\partial y_1}=1$.
Panel (c) shows a similar graph in which decreasing $w_{1i}$ from $w_1$ to $w_1 - \Delta w$ makes college $1$ less likely to be feasible to student $i$. Hence, the region $\mu(i) = 1$ shrinks along the vertical $\eta_{1i}$-axis. The change in the match probability induced by the decrease in $w_{1i}$ is the mass that the density of $\left(\epsilon_{i1},\eta_{1i}\right)$ puts on the area $I_{2}$, or,
In panel (d), the decrease in $z_{i}$ reduces the attractiveness and feasibility of college $1$ for student $i$, because $z_{i}$ enters both student and college preferences. Besides, how $z_{i}$ changes the region of $\mu(i)=1$ relies on the shape of the functions $u^1$ and $v^1$. The change in the match probability induced by the change in $z_{i}$ corresponds to the area $I_{3}$, or,
Our identification result relies on the changes caused by $z_{i}$, $y_{i1}$, and $w_{1i}$. Plugging equations (ref) and (ref) into equation (ref), we have, for $(z_{i},y_{i1},w_{1i}) = (z,\overline{y}_1,w_1)$,
This equation reflects the chain rule: the effect of $z_{i}$ on the match probability, $\frac{\partial\sigma_1}{\partial z_{i}}$, is realized through its effects on utilities $u_{i1}$ and $v_{1i}$, captured by $\frac{\partial u^1}{\partial z_{i}}$ and $\frac{\partial v^1}{\partial z_{i}}$, and the effects of the utilities on the match probability, captured by $\frac{\partial {\sigma_1}}{\partial y_{i1}}$ and $\frac{\partial {\sigma_1}}{\partial w_{1i}}$. In equation (ref), the derivatives of the match probability can be recovered from the population data, and the two unknowns, $\frac{\partial {u^1}}{\partial z_{i}}$ and $\frac{\partial v^1}{\partial z_{i}}$ are the parameters of interest.
Importantly, when $w_{1i}$ varies, the conditional match probability changes, but the two unknowns remain constant, a consequence of $w_{1i}$ entering $v_{1i}$ in a known way. If, for any $z$, $(\epsilon_{i1},\eta_{1i})$ has enough variation such that two distinct values of $w_{1i}$ produce two linearly independent equations that have a unique solution, we identify the unknowns. This requirement is formalized as (ref) later, which imposes a mild restriction on the distribution of $(\epsilon_{i1},\eta_{1i})$ as detailed in Online Appendix (ref).
We then identify $\frac{\partial r^1}{\partial y_{i1}}$, which is not one when $y_{i1}\neq \overline{y}_1$. For any value $(z,y_{1},w_{1})$, plugging equations (ref) and (ref) into equation (ref), we rearrange and obtain
where all the terms except for $\frac{\partial r^1}{\partial y_{i1}}$ are either identified or known. For any $y_{1}$, if $\frac{\partial\sigma_{1}}{\partial z_{i}}-\frac{\partial v^{1}}{\partial z_{i}}\cdot\frac{\partial\sigma_{1}}{\partial w_{1i}} \neq 0$ for some value $(z,w_{1})$, $\frac{\partial r^1}{\partial y_{i1}}$ is identified; otherwise, equation (ref) implies that $\frac{\partial\sigma_{1}}{\partial y_{i1}}=0$ for all $(z,w_{1})$ and thus $\frac{\partial r^1}{\partial y_{i1}}$ is also identified and equal to zero.
Below, we extend this example to the case with multiple colleges. We derive equation (ref) for each college, in which the marginal effect of $z_i$ on the probability of being matched with each college is the sum of its marginal effects on $(u^c,v^c)$ for all $c$. By varying $\{w_{ci}\}_c$, we form a system of equations in $\{\frac{\partial u^{c}}{\partial z_{i}}, \frac{\partial v^{c}}{\partial z_{i}}\}_c$. The identification of $\frac{\partial r^{c}}{\partial y_{ic}}$ is the same as above and relies on a generalized version of equation (ref) because we can identify $\frac{\partial r^{c}}{\partial y_{ic}}$ for each $c$ separately by holding $\frac{\partial r^{d}}{\partial y_{id}}=1$ for $d\neq c$.
Our nonparametric identification of $\{\frac{\partial u^{c}}{\partial z_{i}}, \frac{\partial r^{c}}{\partial y_{ic}}, \frac{\partial v^{c}}{\partial z_{i}}\}_{c}$ extends matzkin_constructive_2019 who uses excluded variables to identify nonparametric nonseparable discrete choice models.
Part (i) of Assumption (ref) requires that all covariates are continuous (but not necessarily have full-support), which is relaxed in (ref).
This exogeneity assumption is made for simplicity. One way to relax this assumption is to adopt a control function approach heckman1985alternative,blundell2004endogeneity, imbens2009identification. See Online Appendix (ref) for a discussion.
Additionally, we need a condition on the derivatives of match probabilities w.r.t.\ the excluded variables. Let $\Pi_{y}(z_{i},y_{i},w_{i})$ be a $C\times C$ Jacobian matrix of the match probabilities w.r.t.\ the excluded variable $y_{i}$, whose $(c,d)$ element is $\frac{\partial \sigma_{c}(z_{i},y_{i},w_{i})}{\partial y_{id}}$. Similarly, let $\Pi_{w}(z_{i},y_{i},w_{i})$ be a $C\times C$ Jacobian matrix w.r.t.\ the excluded variable $w_{i}$. Fix $y_i=\overline{y},$ where $\overline{y}=(\overline{y}_1,\cdots,\overline{y}_C)$. We then consider a pair of distinct values of $w_{i}$, $\widehat{w}$ and $\widetilde{w}$, and define a $2C\times2C$ matrix evaluated at $\left(z_{i},y_{i}\right)=\left(z,\overline{y}\right)$, \[ \Pi(z,\overline{y},\widehat{w},\widetilde{w})\equiv
. \] We impose the following testable condition on $\Pi(z,\overline{y},\widehat{w},\widetilde{w})$.\footnote{For a statistical test of (ref) given a value of $z_i$ and a pair of values of $w_i$, one may use the method proposed by chen_improved_2019. Testing $H_{0}$: rank$(\Pi(z,\overline{y},\widehat{w},\widetilde{w}))\leq2C-1$ against $H_{1}$: rank$(\Pi(z,\overline{y},\widehat{w},\widetilde{w}))>2C-1$ is a special case of setup (1) in chen_improved_2019. In practice, one needs to find two values of $w_i$ satisfying (ref), which can be achieved by the following procedure: (i) choose $m$ pairs of $w_i$ values, (ii) for each pair, apply this test and calculate the $p$ values, and (iii) use the Holm–Bonferroni method to control the overall size of this multiple hypothesis testing problem and then determine which null hypothesis, if any, is rejected. The pair of values of $w_i$ associated with any rejected hypothesis satisfies (ref).}
Note that, for any value of $z_i$, (ref) only needs {\it two} values of $w_i$ at which $\Pi(z,\overline{y},\widehat{w},\widetilde{w})$ is full-rank. In other words, (ref) can hold even when there are infinitely many values of $w_i$ at which $\Pi(z,\overline{y},\widehat{w},\widetilde{w})$ is not full-rank. The condition fails if, for example, for the given value $(z,\overline{y})$, the student's probability of matching with a certain college is always zero for a neighborhood around $\overline{y}$ and all values of $w_i$.\footnote{This may occur if the college is infeasible or unaccaptable to the student with probability one.}
For a set of sufficient, yet weak and testable, conditions for (ref), consider $\widehat{w}$ that is large enough such that all colleges are feasible when $w_i=\widehat{w}$. In this case, the upper half of $\Pi(z,\overline{y},\widehat{w},\widetilde{w})$ corresponds to a one-sided discrete choice model, where the feasible-set shifter $w_i$ has no impact on the matching probabilities and thus $\Pi_{w}(z,\overline{y},\widehat{w})=\boldsymbol{0}_{C\times C}$. This makes $\Pi(z,\overline{y},\widehat{w},\widetilde{w})$ a lower triangular block matrix. Then, (ref) holds if and only if both $\Pi_{y}(z,\overline{y},\widehat{w})$ and $\Pi_{w}(z,\overline{y},\widetilde{w})$ are invertible, which can be achieved under the connected substitutes conditions berry2013connected.\footnote{The invertibility of $\Pi_{y}(z,\overline{y},\widehat{w})$ holds under the following two assumptions: (i) for each $d\in\mathbf{C}\cup\{0\}$, $\frac{\partial \sigma_{d}(z_i,y_i,w_i)}{\partial y_{ic}}\leq 0$ for all $c\in\mathbf{C}\backslash\{d\}$; (ii) for any nonempty $\overline{\mathbf{C}}\subseteq\mathbf{C}$, there exists $c\in\overline{\mathbf{C}}$ and $d\notin\overline{\mathbf{C}}$ such that $\frac{\partial\sigma_{d}(z,\overline{y},\widehat{w})}{\partial y_{ic}}<0$. The invertibility of $\Pi_{w}(z,\overline{y},\widetilde{w})$ holds under similar assumptions.}\textsuperscript{, }\footnote{This sufficient condition for (ref) provides us with a clear comparison with a related study, agarwal2022demand. Their conditions for the identification of a similar model and our conditions are not nested. Specifically, they assume large support on $w_i$, impose a substitution condition on $\Pi_{y}(z,y,w)$ stronger than the one in berry2013connected, and require the substitution condition to hold for all but a finite set of $y_i$ values. In contrast, in this sufficient condition for (ref), with large support of $w_i$, we need a substitution condition on both $\Pi_{y}(z,\overline{y},w)$ and $\Pi_{w}(z,\overline{y},w)$ \`a la berry2013connected, but only for $y_i = \overline{y}$.} This suggests that (ref) holds under common logit or probit models. It is worth noting that the large support of $w_i$ above is mainly used to facilitate a clear connection with the connected substitutes conditions, but is stronger than necessary. In Online Appendix (ref), we show that in a one-college case, (ref) holds for all distribution of $\eta_{1i}$ except for the exponential distribution. In Online Appendix (ref), we analyze common logit and probit models with two or more colleges. The results suggest that (ref) is non-restrictive and satisfied for a wide range of values of $(\widehat{w},\widetilde{w})$; in fact, the failure of (ref) imposes strict restrictions on the supply, or the conditional probabilities of different feasible sets.
Our identification arguments proceed in two steps. First, to identify $\frac{\partial u^{c}(z)}{\partial z_{i}^{k}}$ and $\frac{\partial v^{c}(z)}{\partial z_{i}^{k}}$, we use the variation in the excluded variables $(y_{i},w_{i})$ to derive a system of linear equations that generalizes equation (ref). Fixing $y_i=\overline{y}$, for any value $(z,w)$ in the interior of $\mathcal{Z}\times\mathcal{W}$, for each college $d \in \mathbf{C}$, and for any component of $z_{i}$, $z_{i}^{k}$,
which gives $C$ linear equations of $2C$ unknowns. Note that equation (ref) is for one value of $w_i$. By evaluating equation (ref) at $d=1,\dots,C$ and two different values of $w_i$, $\widehat{w}$ and $\widetilde{w}$, and stacking them together, we have
where $\boldsymbol{\sigma} \equiv (\sigma_{1},\dots,\sigma_{C})'$, $\boldsymbol{u} \equiv (u^{1},\dots,u^{C})'$, and $\boldsymbol{v} \equiv (v^{1},\dots,v^{C})'$. Both $\Pi(z,\overline{y},\widehat{w},\widetilde{w})$ and the left-hand side are known from the population data. Hence, the invertibility of $\Pi(z,\overline{y},\widehat{w},\widetilde{w})$ in (ref) guarantees the existence of a unique solution to this system, leading to the identification of $\frac{\partial u^{c}(z)}{\partial z_{i}^{k}}$ and $\frac{\partial v^{c}(z)}{\partial z_{i}^{k}}$. Importantly, this only requires a pair of distinct values of $w_{i}$ to satisfy (ref).
Second, we identify $\frac{\partial r^{c}(y_{c})} {\partial y_{ic}}$ for each $c\in\mathbf{C}$ by generalizing equation (ref) for the one-college example. See the Appendix for detailed proof.
In certain empirical applications, identifying the above derivatives is sufficient, in which case a large-support assumption on excluded variables is not needed. When the functions $(u^{c}, v^{c}, r^{c})$ must be identified, a full-support condition on their arguments, $(z_{i}, y_{i})$, is often imposed. This is the case for our next result of identifying the joint distribution $F$ and cutoffs $\delta_c$. Additionally, a full-support assumption on $w_i$ is also required. Hence, we will assume all observables, $(z_i,y_i,w_i)$, have full support.
We now formalize the assumptions that are needed for the identification of the cutoffs, $\{\delta_c\}_c$, and the joint distribution of unobservables, $F$.
Part (iv) is a location normalization on college preferences, as mentioned in (ref), which can be done college-by-college. Alternatively, one may replace part (iv) by normalizing cutoffs $\{\delta_{c}\}_c$ to zero.
Before presenting our formal results, we give some intuitions. To identify cutoff $\delta_{c}$, under the full-support assumption on student preferences (parts ii and iii of (ref)), we consider a mass of students to whom all colleges except for $c$ are unacceptable. {The probability of these students matching with $c$ is $1-F_{\eta_{ci}}(\delta_c - v^c(z_{i}) -w_{ci})$. Given the location normalization of $F_{\eta_{ci}}$ (part iv of (ref)), finding the maximum value of $v^c(z_{i})+w_{ci}$ that sets this probability to $1-\rho_c$ identifies $\delta_{c}$.}
To identify $F$, we use the conditional probability of being unmatched. Let us illustrate the intuition in the same one-college example as in (ref). For any given value $(z,y_1,w_1)$, Figure (ref) shows the partition of the space of $(\epsilon_{i1},\eta_{1i})$ by matching outcome, $\mu\left(i\right)=0$ or $1$. The conditional probability of $\mu(i)=0$, highlighted in purple in the figure, can be decomposed into three parts, $R_{1}$, $R_{2}$, and $R_{3}$. Moreover,
where $\ensuremath{\mathbb{P}(R_{2} | z,y_1,w_1 })$ is the joint CDF of $(\epsilon_{i1},\eta_{1i})$, $F\left(-u^{1}(z)-r^{1}\left(y_1\right),\delta_{1}-v^{1}(z)-w_1\right)$, our parameter of interest.
Further, $\mathbb{P}(R_{1}\cup R_{2} | z,y_1,w_1 )=F_{\epsilon_{i1}}(-u^1(z)-r^1(y_1))$ is the marginal CDF of $\epsilon_{i1}$. It can be identified by considering the subset of students whose value of $w_{1i}$ is high enough so that the college will be feasible no matter what value $\eta_{1i}$ takes. That is, we can identify $\mathbb{P}(R_{1}\cup R_{2} |z,y_1,w_1 )$ by “shutting down” the effects of college preferences. Similarly, $\mathbb{P}(R_{3}\cup R_{2} | z,y_1,w_1 )$ can be identified by focusing on the subset of students whose value of $r^1\left(y_{i1}\right)$ is large enough so that those students find college 1 acceptable no matter what value $\epsilon_{i1}$ takes. As $\mathbb{P}(\mu(i)=0 | z,y_1,w_1 )$ is known from the population data, once $\mathbb{P}(R_{1}\cup R_{2} | z,y_1,w_1 )$ and $\mathbb{P}(R_{3}\cup R_{2} | z,y_1,w_1 )$ are identified, equation (ref) implies that $\mathbb{P}(R_{2} |z,y_1,w_1 )$ and thus $F$ are identified.
To show part (ii) in a many-college setting, with the full, large support (parts ii and iii of Assumptions (ref)), we apply the same argument as in the one-college example to each college $c$ sequentially, using extreme values of $r^c(y_{ic})$ or $w_{ci}$ to “shut down” the effects of students' preference for college $c$ or $c$'s preference for students.
Part (iii) is a consequence of part (i) and the definition of cutoffs (equation (ref)). If $c$ reaches its capacity, $T_c$ can be any value below $\delta_{c}$ and result in the same stable matching. We can identify $T_c$ by imposing additional assumptions, e.g., a full-capacity college having the same acceptability threshold as some college with vacancies.
When our identification results are taken to the data, there can be practical issues. For example, vector $z_i$ may include discrete variables such as gender, and the researcher may not have sufficient excluded variables. Below, we address these issues.
Discrete Random Variables. Our results can be extended to the case where $z_i$ contains discrete variables. Suppose that $z_{i}=(z_{i}^1,z_{i}^2)$, where $z_{i}^1$ is a vector of discrete variables and $z_{i}^2$ is a vector of continuous variables. Let the support of $z_i^1$ be a finite set of points $\{z^{1,1}, z^{1,2},\dots, z^{1,J}\}$. For $z_i^1 = z^{1,j}$, we define functions $(u^{c,j}+r^c,v^{c,j})$. Conditional on $z_i^1=z^{1,1}$, we apply the results in (ref) to identify $\{u^{c,1}+r^c,v^{c,1},\delta_c\}_{c}$ and $F$, requiring $y_i$ and $w_i$ being continuous, Assumptions (ref)(ii) and (iii), (ref), (ref)(ii)-(iv), (ref), location normalization on $u^{c,1}+r^c$ and $v^{c,1}$ for each $c$, and $(y_i,z_{i}^2)$ having full support. For $j\neq 1$, conditional on $z_i^1=z^{1,j}$, using the conditional match probability of a mass of students for whom $c$ is the only acceptable college, we can use $F$ to identify $v^{c,j}$ for each $c$; similarly, using the conditional match probability of a mass of students whose feasible set is $\{0,c\}$, we can identify the function $u^{c,j}+r^c$.
Insufficient Excluded Variables. Allowing for college-level heterogeneity implies that we need to recover $3C$ preference parameters, $\{u^c,r^c,v^c\}_c$. As seen in (ref), we need $2C$ excluded variables for identification of their derivatives.
The lack of sufficient excluded variables leads to a loss of identification, but not all is lost. We show that there is a trade-off between the identifiable degree of heterogeneity and the number of excluded variables. Consider $\mathbf{C}_1, \mathbf{C}_2 \subseteq \mathbf{C}$ with cardinality $\kappa_1$ and $\kappa_2$, respectively. For all $c\in\mathbf{C}_1$, there is a single excluded variable in student preferences, $y_{ic}=y_{i\ast}$; for all $c\in\mathbf{C}_2$, there is a single excluded variable in college preferences, $w_{ci}=w_{i\ast}$. We have the following identification result.
In other words, we need only $C-\kappa_1+1$ excluded variables on the student side and $C-\kappa_2+1$ on the college side to identify the less heterogeneous model.\footnote{Note that the rank condition in (ref) is weaker than Condition (ref).} Importantly, this still allows for coefficient heterogeneity across colleges in $\mathbf{C}_1$ and $\mathbf{C}_2$. For example, it can be incorporated parametrically by interacting college-specific observables with $z_i$. For $c\in\mathbf{C}_1$, let $u^c(z_i) = \beta p_c z_i$, where $p_c$ is a college-specific observable. This amounts to $z_i$ having a college-specific parameter $\beta_c = \beta p_c$.
Since our identification approach is constructive, it directly implies a nonparameteric estimator. In the Monte Carlo simulations in Online Appendix (ref), we apply this estimator in a semiparametric setting to avoid the well-known curse of dimensionality; yet the curse remains. Instead, we find that a parametric model based on a Bayesian approach works well. Below, we apply it to a real-life setting.
Guided by our identification results, we study the admissions to public and private secondary schools (grades 9-12) in Chile in 2007. The market is organized similarly to college admissions in the US. It is decentralized, both sides have preferences, and students do not submit rank-order lists of schools. We use a parametric Bayesian approach for preference estimation and then conduct counterfactual analysis.
Since 2003, secondary education has been compulsory for all Chileans up to 21 years of age. In principle, a public school must accept any student who is willing to enroll; a private school can be subsidized by the government or non-subsidized, but in either case, it can select students based on its preferences.
We focus on a relatively independent market, {\it Market Valparaiso}, that includes five municipalities (Valparaiso, Vi\ na del Mar, Concon, Quilpue, and Villa Alema\ na) as defined in gazmuri2017school. Our data includes the municipality of a student's residence and the geographical coordinates of each school, which identifies every agent in the market. A student is defined to be in Market Valparaiso if in 2008 she resided in a municipality within the boundary of the market. A secondary school is in Market Valparaiso if it is located in and admits students from the market.\footnote{There are 38 secondary schools located in Market Valparaiso that do not have any 9th graders from Market Valparaiso in 2007 and another 15 that have fewer than three 9th graders from Market Valparaiso on average in 2005, 2007, and 2009. We consider these 53 schools as outside options.} In total, there are $9,304$ students and $125$ schools. This reasonably large size makes it plausible that the continuum market in (ref) is a good approximation.
We use the SIMCE dataset, provided by La Agencia de Calidad de la Educaci\ on (Agency for the Quality of Education), on all 10th graders in 2008 to identify who started secondary school in 2007. The SIMCE is Chile's standardized testing program and tracks students' math and language performance. The data includes students' parental income, parental schooling, and other characteristics from a parental questionnaire sent home with students. Most school attributes are calculated from student characteristics of {\it the 10th graders in a school in 2006}, and thus are pre-determined in the 2007 admissions that we study. As tuition fees are largely fixed within each school and not completely flexible at the student level, we consider the problem as matching without transfers. See Online Appendix (ref) for more details on data construction and summary statistics of student characteristics and school attributes.
We allow student preferences to be school-type-specific. For student $i$, the utility of attending school $c$ of type $t$ $\in$ \{public, private non-subsidized, private subsidized\} is
where $\alpha_{ft}$ is a school-type fixed effect for female students; $female_i$ is a dummy variable for female; $\alpha_{mt}$ and $male_i$ are similarly defined; $\epsilon_{ic}$ is i.i.d.\ standard normal; $\zeta_t$ ($>0$) allows type-specific variances; and $X_{ic}$ are student-school-specific variables, including (i) the distance between $i$'s residence and school $c$; (ii) 6 school attributes, and (iii) 3 interactions between school attributes and student characteristics.
Each student has an outside option, $u_{i0} = \epsilon_{i0}$, with $\epsilon_{i0}$ being standard normal. We impose the usual scale normalization through $\zeta_t=1$ for public school,\footnote{The variance of $\epsilon_{i0}$ is also normalized to be one because there is insufficient variation to estimate the variance of $u_{i0}$. Only 59 students (out of 9,304) choose an outside option (see (ref)).} while the location normalization is imposed by setting the deterministic part of $u_{i0}$ to zero.
School preferences are also type-specific, but public schools do not have a utility function because they cannot select students. For private school $c$ of type $t$ (subsidized or non-subsidized), its acceptability threshold is $T_c=0$, and its utility function is
where $\theta_t$ is a type-specific intercept; $\eta_{ci}$ is i.i.d.\ standard normal; and the vector $Z_{ci}$ includes (i) 5 student characteristics and (ii) 3 interactions between student characteristics and school attributes.
In school preferences, the variance of $\eta_{ci}$ being one is the scale normalization and $T_c=0$ is the location normalization. Allowing the type-specific intercept $\theta_t$ and assuming $T_c=0$ imply that schools of the same type have the same acceptability threshold. Because each type has some schools with vacancies, we can separately identify $\delta_c$ and $T_c$. Otherwise, we would lose its identification ((ref)).
The above specification uses distance as an $(i,c)$-pair-specific excluded variable in student preferences. In school preferences, math and language scores are student-specific excluded variables, implying that there are insufficient excluded variables to estimate school-specific utility functions. Using the results in (ref), we limit preference heterogeneity and assume type-specific utility functions for schools. Our scale and location normalization is guided by (ref) and (ref), while certain normalization is imposed via $F$.
We use a Bayesian approach with a Gibbs sampler for the estimation, which is first illustrated in Monte Carlo simulations (Online Appendix (ref)). We provide the details on the updating of the Markov Chain in Online Appendices (ref) and (ref).
The estimation results are summarized in (ref). A caveat is in order when we interpret the results. Because we do not deal with endogeneity issues that may arise due to the correlation between preference shocks and school attributes or student characteristics, the estimates may not have a causal interpretation.
Panel A shows the estimates of student preferences. Most coefficients are of an expected sign. Interestingly, the coefficient on tuition is positive in the utility function for public schools, and parental income pushes it towards negative. This may reflect the fact that tuition at public schools is generally low (see (ref)) and may be correlated with unobserved school quality. There are also a few coefficients with an unexpected sign in school preferences. Panel B shows that non-subsidized schools negatively value a student's parental income and mother's education. This may be because, on average, students at those schools often have a high income (above 1.4 million CLP; see (ref)) and a highly educated mother (above 17 years).
Recall that each school's acceptability threshold is normalized to zero, so we can use the estimated school preferences to calculate if a student is acceptable to a private school. On average, a subsidized school finds 79% of the students acceptable, while a non-subsidized school finds 84% acceptable. The higher acceptability rate at non-subsidized schools does not imply that students are more often matched with them because their high tuition lowers their desirability to many students, especially those with a low parental income (see Table (ref)).
Before conducting counterfactual analysis, we evaluate model fit in Online Appendix (ref). It shows that our model fits the data reasonably well when we compare the observed matching with the one predicted based on our model.
We consider a counterfactual policy in which students from low-income families are prioritized for admissions to all schools. A student is of low income if her parental income is among the lowest 40%.\footnote{This resembles a policy adopted in 2008 in Chile as documented by gazmuri2017school. It benefits 44% of elementary school students in 2012 in terms of admission priorities and a tuition waiver. Online Appendix (ref) shows summary statistics of the students by income status.} Each private school's preferences over students are made lexicographical: low-income students are above others, and within each group of students, a school ranks them as in the current regime; all low-income students are acceptable, while others' acceptability is the same as in the current regime. We do not change anything in student or school preferences beyond the admission priority, which may mitigate the potential bias in our estimation due to aforementioned unaddressed endogeneity issues.
To simulate the counterfactual outcome, we keep 1,500 draws of the utilities in the Markov Chain in the Bayesian estimation.\footnote{{Specifically, there are 15 blocks of 100 draws. The blocks are equally spaced in the 0.75 million iterations in the Markov chain that are used to calculate the posterior means and standard deviations.}} We run the Gale-Shapley deferred acceptance with each draw and obtain 1,500 sets of counterfactual stable matchings. We then report the average of these counterfactual matchings; recall that we only have one matching outcome under the current regime---the observed one.
(ref) presents the results. There are several noticeable patterns when we move from the current regime to the counterfactual. First, low-income students are in schools that have higher-ability and higher-income students in the same cohort, while the opposite is true for non-low-income students. Second, some low-income students leave public schools for private subsidized schools, while crowding out some other students to public schools. Lastly, the policy benefits low-income students and hurts others. On average, low-income students' welfare gain is equivalent to decreasing travel distance (to a public school) by 0.332 km. This gain is concentrated among 12% of the low-income students, while others are not affected. Correspondingly, 7.2% of the non-low-income students are worse off, and none is better off.
These results indicate that low-income students dislike private non-subsidized schools. We explore why it is the case. With the 1,500 draws of student preferences, we examine low-income students' favorite school of each type. We find that low-income students on average value their favorite public school at $2.72$ and private subsidized school at $2.56$, while their favorite non-subsidized school is only valued at $-21.76$, in general unacceptable to them. We calculate the contribution of different variables to these differences by shutting down their effect in the utility functions. The results show that high tuition at non-subsidized schools is the main contributor. It is worth noting that tuition could be correlated with unobserved school quality. Hence, low-income students may dislike a non-subsidized school because of its high costs and/or their tastes. Additionally, mother's education, both a student's own and a school's average, is also an important factor. Student ability, measured by their composite score, and distance to each school do not appear to be as important.
In sum, giving low-income students access to schools fails to significantly change matching outcomes due to their own preferences. Low-income students are deterred from private non-subsidized schools by their high tuition. These findings are in line with the preference heterogeneity documented in public school choice abdulkadiroglu_welfare_2017,kapor2020heterogeneous, although tuition plays no role there.
We study nonparametric identification of agent preferences in many-to-one two-sided matching without transfers. We derive a set of sufficient conditions for identification and provide guidance for empirical studies. For example, our results clarify the data requirement for the identification of various degrees of preference heterogeneity. To take our results to the data of a reasonably sized market, we propose a Bayesian approach with a Gibbs sampler whose performance is illustrated in Monte Carlo simulations. Our model encompasses many real-life matching markets, such as college admissions and school choice in many countries, in which our identification results and empirical method can be applied. Hence, this paper opens a new avenue for empirical research.
We illustrate our method in the context of secondary school admissions in Chile. As an example of the usefulness of the estimates, we consider a counterfactual policy in which students from low-income families are prioritized for admissions to all schools. Although the policy benefits low-income students, its effects are small. Such insights are difficult to obtain without estimating the preferences of both sides. In this sense, our method can help provide an ex-ante evaluation of a range of alternative policies.