Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
107,446 characters · 30 sections · 102 citation commands
Role models and revealed gender-specific costs of STEM in an extended Roy model of major choice.
\address{The Pennsylvania State University} \address{Oxford University} \address{University of Toronto, Washington University in St. Louis and NBER} \begingroup \footnote{\scriptsize{The first version is of August 21, 2019. The present version is of July 14, 2023. This research was supported by SSHRC Grant 435-2018-1273 and Leibniz Association Grant SAW-2012-ifo-3. The research was conducted in part, while Isma\^^22el Mourifi\'e was visiting the Becker Friedman Institute at the University of Chicago. He thanks his hosts for their hospitality and support. The authors thank the editor Elie Tamer, the associate editor and a referee for very detailed and insightful reports, that were instrumental to significant improvements in the paper (the usual disclaimer applies of course). The authors also acknowledge the exceptional research assistance of Ivan Sidorov, as well as helpful comments from Roy Allen, Samuel Cohen, Chris Dobronyi, Jonathan Eaton, Jim Heckman, Hiro Kasahara, Thomas Russell, Azeem Shaikh and seminar audiences at Cambridge, Chicago, Notre Dame, Rice, Texas A&M, Vanderbilt and York. Correspondence address: Department of Economics, Max Gluskin House, University of Toronto, 150 St. George St., Toronto, Ontario M5S 3G7, Canada}} \addtocounter{footnote}{-1} \endgroup
{\scriptsize
}
\thispagestyle{empty}
{\scriptsize Keywords: Roy model, partial identification, stochastic monotonicity, sharp bounds, sharp testable implications, education sector choice, role models, women in STEM.
JEL subject classification: C31, C34, I21, J24 }
A large body of evidence supports sorting in the labor market as a dominant mechanism in the explanation of the residual gender wage gap. See Goldin:2021 and references therein. Two salient aspects of sorting in the labor market are the persistent under representation of women in Science, Technology, Engineering and Mathematics (hereafter STEM), which is perceived to be one of the main drivers of the gender wage gap (see daymont1984, zafar2013, and SHB:2019), and the documented influence of role models on sorting choices and eventual outcomes. A survey of the scholarship on both these aspects can be found in KG:2017.
We propose to examine the issue of under representation of women in STEM from the point of view of college major choice, which is a major sorting mechanism for skilled labor, and with the lens of the Roy model. However, the traditional version of the Roy model cannot accommodate stylized facts relating to major and occupational choice. The literature on major and occupational choice, reviewed in AAM:2016, emphasizes the importance of factors beyond expected earnings. This motivates the current theoretical analysis of partial identification of gender-specific non pecuniary drivers of major choice within what is generally known in the literature as the extended Roy model (see HV:99 for details on model genealogy and attribution). The model allows for observed heterogeneity at the individual level, and unobserved heterogeneity at the educational sector level, to influence choices. BKT:2011 and HM:2011 analyze the extended Roy model and provide competing strategies to identify the non pecuniary driver of choice (see the following subsection for details). We propose a utility-based version of the extended Roy model, that allows the implicit non pecuniary costs of each sector to also depend on potential outcomes, unlike BKT:2011 and HM:2011. We make alternative partially identifying assumptions to those in BKT:2011 and HM:2011 and derive sharp bounds on utilities and on the non pecuniary driver of choice. The model allows for multiple interpretations of the non pecuniary component, including job amenities, anticipated lack of support for women in mathematics intensive education and occupations, as well as social conditioning and gender stereotyping of occupations.
We combine an extended Roy model of sorting with shape restrictions induced by the influence of role models. The observed realized earnings variable $Y$ is equal to $Y=Y_1D+Y_0(1-D)$, where $Y_0$ and $Y_1$ are unobserved potential earnings in Sector $0$ (non mathematics intensive college major, hereafter referred to as non-STEM) and Sector $1$ (mathematics intensive college major, hereafter referred to as STEM), and $D$ is the observed random sector selection indicator. Individuals choose Sector $d=0,1,$ when $\mathbb E[u(Y_d,d,W)\vert \mathcal I]>\mathbb E[u(Y_{1-d},1-d,W)\vert \mathcal I]$, where $u$ is the utility function, $\mathcal I$ is the individual's information set at the time of choice, and $W$ is an $\mathcal I$-measurable vector of individual and sector observed variables that influence choice and outcomes. This specification differs in two respects with the extended Roy models in the literature: first we do not impose a tie-breaking mechanism, which is important in case of interval censored observations; second, we do not impose quasi-linearity of utility, and allow varying marginal rates of substitution between consumption and sector amenities. The latter feature is shared with LP:2021.
We formalize the influence of role models on choices and eventual outcomes with the assumption that a subvector $Z$ of the observable variables $W:=(X,Z)$ can only affect potential expected utilities in one direction. Precisely, for all $\nu$ and $x$, the joint probability that $\mathbb E[u(Y_1,1,W)\vert\mathcal I]>\nu$ and $\mathbb E[u(Y_0,0,W)\vert\mathcal I]>\nu$, conditionally on $X=x$ and $Z=z$, is monotone non decreasing in $z$. This shape restriction is inspired by monotone instrumental variables, and related to MP:2000, BGIM:2007 and MHM:2020. We argue that factors that positively impact the formation of cognitive and non cognitive skills, while reducing perceived costs of the STEM sector, are likely to satisfy this shape restriction. Maternal educational attainment and the proportion of female role models (faculty, alumni, invited speakers) are prime candidates (see BGME:2018 and RCM:2014).
The objective of the analysis is to characterize the collection of utility functions that rationalize observed choices. We will refer to this set as the {\em identified set} (also known as sharp identified region in the partial identification literature). We also derive a closed form expression for the minimum perceived cost of mathematics intensive education and activity to rationalize the observed under representation of women in that sector. We interpret this cost as a wage differential that compensates for the real or perceived disamenities of the STEM sector for women. Both the identified set and the closed form lower bound are amenable to inference using existing methods for stochastic monotonicity tests, conditional moment inequalities and intersection bounds.
MHM:2020 document rejection of the Roy model of sorting on labor market outcomes for women in the sample of German graduates in the 2005 and 2009 graduating cohorts of the German DZHW Graduate Survey (see DZHW). We therefore illustrate our theoretical results with an analysis of women's choice of major within the framework of our extended Roy model and confidence regions for the minimum cost of STEM that rationalizes choices. We assume that the proportion of women on the STEM faculty in the individual's region at the time of major choice (as a proxy for the presence of role models and better amenities for women) only affects potential labor market outcomes of women graduates positively. We find significant costs of choosing STEM fields for German women from the former Federal Republic, only when assuming perfect foresight (i.e., perfect knowledge of the potential utilities) when choosing the major. The costs are particularly pronounced for women in the lower quartile of the income distribution and for women whose region had low rates of feminization of the STEM faculty at the time of major choice.
Early results on identification of a parametric version of the extended Roy model are given in Vijverberg:93. Nonparametric identification in extended Roy models is analyzed in HM:2011 and BKT:2011. BKT:2011 restrict attention to non pecuniary costs that are constant across agents, and HM:2011 to non pecuniary costs that depend on observables only, and not on potential outcomes. BKT:2011 identify the pecuniary cost function from a common lower bound on the supports of both potential incomes, or alternatively from the exclusion of a location shifter in the potential outcome equations. HM:2011 combine an additive decomposition of the agent's expected earnings with some smoothness and a continuous covariate. LP:2021 also propose an extended Roy model in structural form. The selection equation in LP:2021 is the same as in our perfect foresight case, so that implicit non pecuniary costs can also depend on potential outcomes. However, identification is achieved from the variation in a utility shifter that is excluded from the potential outcome equations. The latter which may be hard to find and justify in some empirical contexts, such as the one we entertain here.
The inspiration for the shape restriction that drives partial identification goes back at least to MP:2000. More recently, MHM:2020 derive sharp testable implications of stochastic monotonicity constraints in the traditional Roy model. In a related earlier contribution, Chetverikov:2013 shows that regression monotonicity is the testable implication of the monotone treatment response assumption and monotone treatment selection assumption introduced in MP:2000.
The empirical illustration of our proposed methodology fits into a growing applied literature on the causes of under representation of women in university STEM majors. That literature is already sizable, as evidenced by the surveys in AAM:2016 and KG:2017. The most closely related in terms of methodology is Saltiel:2018, which looks at women's college major choice through the lens of a dynamic generalized Roy model inspired by HHV:2018 [HHV:2016,HHV:2018]. As concerns empirical findings, our paper relates more specifically to the literature on the role of female professors and gender specific amenities in shaping major choice and educational outcomes. This includes CR:95, BGME:2018, CPW:2010, CM:2021 on the influence of female professors as role models, and daymont1984, MP:2017, KSW:2018, WZ:2015 [WZ:2015,WZ:2018] and MHM:2020 on the importance of gender specific amenity preferences and gender specific response to non pecuniary factors in the valuation and choice of occupations.
The next section presents the extended Roy model and the main identification results. Section (ref) focuses on the special case of perfect foresight. Section (ref) discusses inference methods. Section (ref) presents a simulation study inspired by the under representation of women in STEM fields, and Section (ref) illustrates the methodology with an analysis of women's major choices in Germany. The last section concludes. Proofs of the main results are collected in the appendix.
We adopt the framework of the potential outcomes model $Y=Y_1D+Y_0(1-D).$ Observed outcome $Y$ has support $\mathcal Y\subseteq[\underline b,\infty)$, (with $\underline b\in\mathbb R\cup\{-\infty\}$ and $\underline b=0$ or $\underline b=-\infty$ in most cases of interest), $D$ is an observed selection indicator, which takes value $1$ if Sector 1 is chosen, and $0$ if Sector $0$ is chosen, and $Y_1$, $Y_0$, are unobserved potential outcomes. In the context of major choice, the outcome of interest will be income in the year following graduation. Sector 1 will consist of all STEM majors and Sector 0, the rest. Decision makers choose their sector of activity based on the realizations of $Y_0$ and $Y_1$, and a vector of observed exogenous characteristics $Z$ with support $\mathcal Z\subseteq\mathbb R^{d_z}$. The vector $Z$ can be a vector $Z=(Z_0,Z_1)$ of sector specific cost shifters\footnote{Unlike potential outcomes $Y_d$, which are only observed in the chosen sector, the sector specific cost shifters are observed in both sectors.}. In the context of women's major choice, for instance, this could be the proportion of women on the faculty in both sectors in the individual's region or prospective university at the time of choice. The shifter $Z$ can also be a non sector specific (possibly vector) cost shifter. In the context of women's major choice, this could be the mother's education attainment. Additional observed exogenous covariates will be omitted from the notation. In the context of major choice in Germany, these include gender, visible minority status and a dummy for residence in the former East Germany, as a less affluent region.
Since $Y,$ $D$ and $Z$ are observed, the distribution $\pi$ of $(Y,D,Z)$ is directly identified from the data. We call $\Pi$ the set of admissible data generating processes. Unless otherwise specified, $\Pi$ is the set of probability distributions on $\mathcal Y\times\{0,1\}\times\mathcal Z$. We summarize the model with the following assumptions.
The original Roy model posits sector selection based only on the comparison of potential outcomes, so that $Y=\max\{Y_0,Y_1\}$. Given our focus on the under representation of women in STEM fields, and rejections of the original Roy model selection rule in MHM:2020, we entertain the possibility that other factors affect sector selection in favor of the non STEM sector. In the context of women's major choices, there might be a gender-specific cost of studying in a STEM field or a gender-specific cost of working in the STEM sector. The former may be the result of the lack of support for female students, or a fear of mathematics carried over from schooling (see XFS:2015 for a survey of the sociological literature on the subject). The latter may be related to family-friendliness of employment outside STEM.
Hence, in our model, the decision by the individual is based on the comparison between expected utilities $\mathbb E[u(Y_d,d,Z)\vert \mathcal I]$ for $d\in\{0,1\}$, where $Y_d$ is the potential outcome (wage) in Sector $d$ and $Z$ is the (possibly vector) value of the non pecuniary factor that may also affect utility. Choices are made using expectations based on the decision maker's information set $\mathcal I$ at the time of decision. We assume throughout this section that the instrument $Z$ is $\mathcal I$-measurable, since it is a vector of variables easily observable at the time of choice. In the context of major choice, the information set involves some knowledge of individual talent for mathematics and non mathematics intensive activities, as well as some anticipation of future labor market conditions and the prices of talent.\footnote{We can allow the information set to contain a vector $W$ of observed exogenous variables to increase the informativeness of our characterization, without changing the analysis: the characterization in Theorem (ref) would then be conditional on $W$.}
Sector specific utility $u(y,d,z)$ can be rationalized by different amenities. Women applicants may prefer Sector $0$ because the proportion of women in the faculty is larger, hence female specific amenities are better provided. Sector specific utility may also be rationalized with reference dependence (Thaler:80 and TK:91) based on gender profiling: social conditioning makes women prefer Sector 0.
The model characterized by Selection Assumption (ref) is an extended Roy model, in the sense that in each sector, and for each discrete socio-economic group, the utility is a deterministic function of the potential outcome and the non pecuniary vector of variables $Z$. The utility may depend on unobservables (hence the dependence on Sector $d$), but there is no unobservable heterogeneity at the individual level, beyond potential outcome $Y_d$.
Without further restrictions, the trivial choice $u(y,d,z)=y$ (traditional Roy model) would rationalize any data generating process under Assumptions (ref) and (ref). Indeed, for any given random vector $(Y,D)$ generating observations, the choice $Y_0=Y_1:=Y$ trivially satisfies Assumptions (ref) and (ref). To restore testability, we introduce a shape restriction on the joint distribution of potential expected utilities $(\mathbb E[u(Y_0,0,Z)\vert \mathcal I],\mathbb E[u(Y_1,1,Z)\vert \mathcal I])$. A traditional approach to restoring testability without parametric restrictions is to allow some observed covariates to affect sector selection only. However, such restrictions are difficult to justify in the context of college major choice. Parental educational attainment is likely to be correlated with unobserved parental cognitive and non cognitive investments in their children (see Card:2001). Distance to college and other instruments designed to affect educational attainment choices are not suitable for major choice. Other variables that are very relevant to a woman's major choice, such as the proportion of women on the faculty, are expected to affect potential outcomes for women as well as their choices. We resort instead to a weaker instrumental notion, where the instrument $Z$ may affect the joint distribution of potential outcomes, but only in one direction, in terms of first order stochastic ordering.
Assumption (ref) holds if the vector $(\mathbb E[u(Y_0,0,Z)\vert \mathcal I],\mathbb E[u(Y_1,1,Z)\vert \mathcal I])$ is stochastically monotone\footnote{See Appendix (ref) for a definition.} with respect to $Z$. The case $\mathbb E[u(Y_d,d,Z)\vert \mathcal I]=\mathbb E[u(Y_d,d,Z)\vert Z]+V_d$, for $d\in\{0,1\},$ with $(V_0,V_1)\perp Z$, coupled with monotonicity in $z$ of $\mathbb E[u(Y_d,d,Z)\vert Z=z],$ for $d\in\{0,1\},$ is a special case of Assumption (ref). Assumption (ref) also holds if the utility $u$ is quasi-linear, increasing in $z$, and the vector of expected potential outcomes $(\mathbb E[Y_0\vert \mathcal I],\mathbb E[Y_1\vert \mathcal I])$ is stochastically monotone with respect to $Z$.
In the context of women's major choice, expected support and amenities for female students in STEM fields are likely to increase with the presence of female faculty and role models in STEM fields and in the educational level of the student's mother. Hence expected utility is likely to increase with the sector selection variables $Z$. Combined with the assumption that more female faculty or a higher educational attainment of the mother cannot hurt a woman's earnings prospects, it yields Assumption (ref).
We characterize the set of all utility functions that can rationalize the data under Assumptions (ref), (ref), and (ref). This will be the content of Theorem (ref). We start with a formal definition of the set of utility functions that rationalize the data under the model assumptions.\footnote{The identified set exhausts all the information contained in the model. As such, it is also known in the literature as {\em sharp identified region}.}
We are interested in characterizing the identified set $\mathcal U(\pi)$ with moment inequalities. Under Assumptions (ref) and (ref), when $D=1$, $\mathbb E[u(Y,D,Z)\vert\mathcal I]=\mathbb E[u(Y_1,1,Z)\vert\mathcal I]\geq \mathbb E[u(Y_0,0,Z)\vert\mathcal I]$, and when $D=0$, $\mathbb E[u(Y,D,Z)\vert\mathcal I]=\mathbb E[u(Y_0,0,Z)\vert\mathcal I]\geq \mathbb E[u(Y_1,1,Z)\vert\mathcal I]$. Hence Assumptions (ref) and (ref) imply that
This relation is the key to deriving testable implications of restrictions on the joint distribution of potential outcomes. This relation also provides a direct proof of the identification of the cost function under the assumptions of HM:2011, as described in Appendix (ref). Equation ((ref)) also implies that $\mathbb E[u(Y,D,Z)\vert Z=z]$ is monotonically non decreasing in $z$, which can be expressed as a collection of moment inequalities involving the infinite dimensional parameter $u$. Conversely, monotonicity of $\mathbb E[u(Y,D,Z)\vert Z=z]$ with respect to $z$ is shown to imply that $u$ can rationalize the data, i.e., that we can construct a vector of potential outcomes $(Y_0,Y_1)$ such that Assumptions (ref), (ref), and (ref) are satisfied for some information set $\mathcal I$. This discussion is formalized in the next theorem (proved in the appendix).
Theorem (ref) shows that the identified set of Definition (ref) is characterized by monotonicity of conditional expected realized utilities with respect to $Z$. A confidence region for the true utility can be obtained by collecting all the utility functions such that the hypothesis of monotonicity of $\mathbb E[u(Y,D,Z)\vert Z=z]$ is not rejected. We provide two applications of Theorem (ref) to operationalize it in our empirical setting. First we propose a direct application to the case, where the utility is parametric (i.e., known up to a finite dimensional parameter vector). Then, we provide closed form expressions for the sharp bounds on the implied gender specific cost of STEM, under additional shape restrictions.
The most direct application of Theorem (ref) is to the case of parametric utility $u(y,d,z;\theta)$ for a finite dimensional parameter $\theta\in\Theta\subseteq\mathbb R^{l_\theta}$. In that case, the identified set $\Theta_I(\pi):=\{\theta\in\Theta:\; u(\cdot,\cdot,\cdot;\theta)\in\mathcal U(\pi)\}$ is characterized by monotonicity of $\mathbb E[u(Y,D,Z;\theta)\vert Z=z]$ with respect to $z$. Hence, a confidence region for the true value $\theta_0$ of the parameter can be obtained by collecting all $\theta\in\Theta$ such that the hypothesis of monotonicity of $\mathbb E[u(Y,D,Z;\theta)\vert Z=z]$ with respect to $z$ is not rejected.
We now seek to characterize a minimal cost function that rationalizes the data under the imperfect foresight Roy model in closed form. We target our model specifically to explain under representation of women in STEM with a specific burden perceived by women in STEM education and STEM professional activities. To that end, we simplify our utility model to allow for an effect of our instrument in the STEM sector only. This is particularly relevant in case the instrument measures the influence of role models in STEM only. We therefore work on a restricted domain to obtain bounds in closed form.
In the rest of this section, we assume that the set $\tilde{\mathcal Z}$ of Definition (ref) is non empty. We start from Theorem (ref), which characterizes cost functions that rationalize the data with the monotonicity in $z$ of the conditional expectation $\mathbb E[u(Y,D,Z)\vert Z=z]$. This is equivalent to monotonicity in $z$ of $\mathbb E[Y-DC(Y,Z)\vert Z=z]$, with $C(y,z):=y-u(y,1,z)$. Since $C$ is non negative on $\mathcal Y\times\tilde{\mathcal Z}$, we seek functions $C$ such that $\mathbb E[Y-DC(Y,Z)\vert Z=z]=\inf\{\mathbb E[Y\vert Z=\tilde z];\,\tilde z\geq z\},$ the lower monotone envelope of $\mathbb E[Y\vert Z=z]$. As shown in the proof of Corollary (ref), this yields the moment equality constraint $\mathbb E[C(Y,Z)\vert D=1,Z=z]=\underline C(z)$, where $\underline C$ is defined in ((ref)) below. The upper bound $\bar C(z)$ is obtained using the worst case bound.
Corollary (ref) yields closed form bounds on the cost function in case of quasi-linear utility $u(y,1,z)=y-C(z)$. Corollary (ref) also yields implications on the testability of the extended Roy model with imperfect foresight. If $\pi$ is such that $\mathbb P(D=1\vert Z=z)>0$ for all $z$ on the support of $Z$, then $\underline C(z)$ is well defined, and the utility function defined by $u(y,0,z):=y$ and $u(y,1,z):=\underline C(z)$ is in $\mathcal U(\pi)$. However, if $\pi$ is such that $\mathbb P(D=1\vert Z=z)=\mathbb P(D=1\vert Z=\tilde z)=0$ and $\mathbb E[Y\vert Z=z]<\mathbb E[Y\vert Z=\tilde z]$ for $z\geq\tilde z$ on the support of $Z$, then Assumptions (ref), (ref) and (ref) are jointly rejected.
The special case, where potential earnings are measurable with respect to the decision maker's information set has important implications and deserves special treatment. The empirical contents of the imperfect and perfect foresight models differ significantly. In the imperfect foresight case, sharp testable implications of the model specification take the form of conditional mean monotonicity constraints. In the perfect foresight case, they take the form of stochastic monotonicity constraints, and the resulting sharp bounds can therefore be considerably tighter. Another significant difference is the richer structural interpretation of the revealed costs of STEM as a compensating differential in the perfect foresight case.
Although the perfect foresight case is nested within the imperfect foresight one described in Section (ref), we state assumptions and results in full for greater clarity. The decision by the individual is based on the comparison between utilities $u(Y_d,d,Z)$ for $d\in\{0,1\}$, where $Y_d$ is the potential outcome (wage) in Sector $d$ and $Z$ is the (possibly vector) value of the non pecuniary factor that may also affect utility.
\begin{assumption+}{(ref)$'$}[Selection] The utility function $(y,d,z)\mapsto u(y,d,z)$ on $\mathcal Y\times\{0,1\}\times\mathcal Z$ is continuous and increasing in its first argument, and satisfies \[ u(Y_d,d,Z)>u(Y_{1-d},1-d,Z)\Rightarrow D=d, \mbox{ for }d\in\{0,1\}. \] \end{assumption+}
In case of quasi-linear utility $u(y,d,z);=y-C(d,z)$, Assumption (ref) is equivalent to $Y_d-C(d,Z)>Y_{1-d}-C(1-d,Z)\Rightarrow D=d$, for $d\in\{0,1\},$ which is equivalent to the traditional extended Roy model selection, under perfect foresight, except for the fact that we don't impose any tie-breaking rule. However, in many applications, the non pecuniary cost may depend on potential outcomes, so that considering non quasi-linear utility comparisons is important. For example, we expect the costs incurred by women studying or working in the mathematics intensive sector to be less pronounced for women with higher mathematics ability, hence decreasing in $Y_1$. In the more general case (beyond quasi-linear utility), Assumption (ref) is equivalent to
with $C(y,z)=y-u^{-1}(u(y,1,z),0,z)$, where $u^{-1}$ denotes the inverse with respect to the first argument. Assumption (ref), therefore, describes a perfect foresight extended Roy model, where the non pecuniary cost may depend on the potential outcome.\footnote{More generally, the cost function $C(y,z)=y-u^{-1}(u(y,1,z),0,z)$ depends on potential outcomes unless the constraint $\partial u(y,1,z)/\partial y=\partial u(u^{-1}(u(y,1,z),0,z),0,z)/\partial y$ holds.}
In the perfect foresight case, the cost function $C$ can be clearly interpreted as a compensating wage differential. Suppose women perceive inferior amenities in the STEM sector. Call $(y,z)\mapsto\tilde C(y,z)$ the compensating differential defined by $u(y,1,z)=u(y-\tilde C(y,z),0,z)$. Then $\tilde C(y,z)=y-u^{-1}(u(y,1,z),0,z)=C(y,z)$. Hence, $C(y,z)$ in Assumption (ref) is a monetary adjustment that makes women, whose talents entitle them to identical (uncompensated) wages in both sectors, indifferent between the two sectors. As defined, it is a willingness to pay for the better amenities of the non-STEM sector, or equivalently, the equivalent variation to a move from non-STEM to STEM.
As in the general case, without further restrictions, the trivial choice $u(y,d,z)=y$ (traditional Roy model) would rationalize any data generating process under Assumptions (ref) and (ref). Indeed, for any given random vector $(Y,D)$ generating observations, the choice $Y_0=Y_1:=Y$ trivially satisfies Assumptions (ref) and (ref). The shape restriction inspired by the influence of role models takes the following form in case of perfect foresight.
\begin{assumption+}{(ref)$'$} The random vector $(u(Y_0,0,Z),u(Y_1,1,Z))$ is such that for each $\nu\in\mathbb R$, the quantity $\mathbb P(u(Y_0,0,Z)\leq \nu,u(Y_1,1,Z)\leq \nu\vert Z=z)$ is monotonically non increasing in $z$. \end{assumption+}
Assumption (ref) allows dependence of potential outcomes on the instrument. Assumption (ref) is weaker than stochastic monotonicity of the vector of potential utilities in $z$, as defined in Appendix (ref). Indeed, stochastic monotonicity implies (but is not equivalent to) monotonicity in $z$ of the quantity $\mathbb P(u(Y_0,0,Z)\leq \nu_0,u(Y_1,1,Z)\leq \nu_1\vert Z=z)$ for all $(\nu_0,\nu_1)\in\mathbb R^2$. Assumption (ref) only involves constraints for a scalar $\nu$ and is hence weaker. Assumption (ref) is in the same spirit, but is not directly comparable to the monotone instrumental variable (MIV) restriction of MP:2000. Unlike MIV, Assumption (ref) places restrictions on the joint distribution of potential outcomes, as opposed to the marginals only, and drives our characterization of the model's empirical content in Theorem (ref).
In the context of women's major choice, we expect support and amenities for female students in STEM fields to increase with the presence of female faculty and role models in STEM fields and in the educational level of the student's mother. Hence we expect utility to increase with the sector selection variables $Z$. Combined with the assumption that more female faculty or a higher educational attainment of the mother cannot hurt a woman's earnings prospects, it yields Assumption (ref).
We characterize the set of all utility functions that can rationalize the data under Assumptions (ref), (ref), and (ref). This will be the content of Theorem (ref). We start with a formal definition of the set of utility functions that rationalize the data under the model assumptions.
\begin{definition+}{(ref)$'$}[Identified Set under perfect foresight] For any $\pi\in\Pi$, we call $\mathcal U^\prime(\pi)$ the collection of functions $u:\mathcal Y\times\{0,1\}\times\mathcal Z\rightarrow\mathbb R$, such that there exists a random vector $(Y_0,Y_1,D,Z)$ where $((1-D)Y_0+DY_1,D,Z)$ has distribution $\pi$ and Assumptions (ref), (ref), and (ref) are satisfied. \end{definition+}
We are interested in characterizing the identified set $\mathcal U^\prime(\pi)$ with moment inequalities. Under Assumptions (ref) and (ref), when $D=1$, $u(Y,D,Z)=u(Y_1,1,Z)\geq u(Y_0,0,Z)$, and when $D=0$, $u(Y,D,Z)=u(Y_0,0,Z)\geq u(Y_1,1,Z)$. Hence we have
Hence, Assumptions (ref), (ref), and (ref) imply that $\mathbb P(u(Y,D,Z)\leq \nu\vert Z=z)$ is monotonically non increasing in $z$, for all $\nu\in\mathbb R$, which is equivalent (by definition) to stochastic monotonicity of $u(Y,D,Z)$ with respect to $Z$, and which can be expressed as a collection of moment inequalities involving the infinite dimensional parameter $u$. Conversely, stochastic monotonicity of $u(Y,D,Z)$ with respect to $Z$ is shown to imply that $u$ can rationalize the data, i.e., that we can construct a vector of potential outcomes $(Y_0,Y_1)$ such that Assumptions (ref), (ref), and (ref) are satisfied. This discussion is formalized in the next theorem (proved in the appendix).
Theorem (ref) shows that the identified set of Definition (ref) is characterized by stochastic monotonicity of $u(Y,D,Z)$ with respect to $Z$, which is equivalent to the collection of moment inequalities in ((ref)). A confidence region for the true utility can be obtained by collecting all the utility functions such that the hypothesis of stochastic monotonicity of $u(Y,D,Z)$ is not rejected. We provide two corollaries to Theorem (ref) to operationalize it in our empirical setting. First we propose a direct application to the case, where the utility is parametric (i.e., known up to a finite dimensional parameter vector). Then, we provide closed form expressions for the sharp bounds on the implied gender specific cost of STEM, under additional shape restrictions.
The most direct application of Theorem (ref) is to the case of parametric utility $u(y,d,z;\theta)$ for a finite dimensional parameter $\theta\in\Theta\subseteq\mathbb R^{l_\theta}$. In that case, the identified set $\Theta_I^\prime(\pi):=\{\theta\in\Theta:\; u(\cdot,\cdot,\cdot;\theta)\in\mathcal U^\prime(\pi)\}$ is characterized by the collection of moment inequalities in ((ref)). Moreover, a confidence region for the true value $\theta_0$ of the parameter can be obtained by collecting all $\theta\in\Theta$ such that the hypothesis of stochastic monotonicity of $u(Y,D,Z;\theta)$ with respect to $Z$ is not rejected.
Assumption (ref) is equivalent to ((ref)) with cost function $C(y,z)=y-u^{-1}(u(y,1,z),0,z)$. We derive closed form bounds on the cost function $C$ under the following shape restrictions on the utility functions.
\begin{definition+}{(ref)$'$} Call $\tilde{\mathcal Z}^\prime$ the subset of $\mathcal Z$ such that, for each $(y,z)\in\mathcal Y\times\tilde{\mathcal Z}^\prime$, (1) $u(y,0,z)\geq u(y,1,z),$ and (2) $u(y,0,z)$ is non increasing in $z$. \end{definition+}
On $\tilde{\mathcal Z}^\prime$, which we assume non empty for the remainder of this section, unobserved amenities are superior in Sector $0$ for all individuals in the socio-economic group of interest. This is empirically relevant in our application, where $z=(z_0,z_1)$ is the vector of proportions of women on the the faculty in each sector, and $u(y,d,z)=u(y,z_d)$. Then, Condition (1) becomes $u(y,z_1)\leq u(y,z_0)$, which is plausible, since STEM departments have far lower proportions of women on the faculty than non STEM ones. As for Condition (2), it is satisfied if the utility function has a satiation point (in $z_d$) below the realized proportions of women on the faculty of non STEM departments.
On $\tilde{\mathcal Z}^\prime$, we have:
where the first inequality and the first equality follow from the definitions of $C$ and $\tilde{\mathcal Z}^\prime$, the second equality follows from Assumption (ref), and the last inequality follows from Assumption (ref) and the fact that $u(y,0,z)$ is increasing in $y$. Hence, for each $y\in\mathcal Y$, we have
where the last inequality uses standard worst-case bounds for the distribution of potential outcome $Y_0$.
Now, the middle term is equal to
which is monotone non increasing in $z\in\tilde{\mathcal Z}^\prime$ for each $y$ by Assumptions (ref), and is right-continuous in $y$. Hence, the random variable $Y-DC(Y,Z)$ is stochastically monotone non decreasing with respect to $Z$. Heuristically, if realized outcomes $Y$ was stochastically monotone non decreasing with respect to $Z$, the function $C=0$ would rationalize the data, and would then be the desired lower bound. In general, the lower bound $\underline C$ for the cost function is such that the distribution of $Y-D\underline C(Y,Z)$ is the monotone lower envelope of the distribution of $Y$, as defined by the first display in ((ref)) below. The second display in ((ref)) defines the upper envelope of the worst case bound in the same way.
Given the stochastic monotonicity of $Y-DC(Y,Z)$ with respect to $Z$, we therefore have
for all $(y,z)\in\mathcal Y\times\tilde{\mathcal Z}^\prime,$ whenever Assumptions (ref), (ref), and (ref) hold.
Lemma (ref) below shows that the envelopes defined in display ((ref)) are themselves cumulative distribution functions, and that they are the bounds of the set of cumulative distribution functions with the desired monotonicity properties.
Equation ((ref)) yields the testable implication for our model that bounds $\underline F(y\vert z)$ and $\bar F(y\vert z)$ cannot cross. Since $\mathbb P(Y-DC(Y,Z)\leq y\vert z) = \mathbb P(Y-C(Y,Z)\leq y, D=1\vert z) + \mathbb P(Y\leq y, D=0\vert z),$ and since $\mathbb P(Y-C(Y,Z)\leq y, D=1\vert z)$ is non decreasing in $y$, ((ref)) also yields the bounds
Finally, we have the following corollary of Theorem (ref), proved in the Appendix, which establishes bounds on the cost functions that correspond to utilities in the identified set of Definition (ref).
The bounds of Corollary (ref) may not be attained under the additional shape restrictions on utilities. We now turn to a characterization of assumptions on the cost function $C(y,z)$ such that the bounds of Corollary (ref) are attained (hence sharp). Since the result below relies on a testable, but high level regularity condition on the data generating process $\pi$, we later complement it with an iterative procedure to tighten the bounds in case the regularity condition fails to hold.
Consider the following restricted set of data generating processes.
The bounds of Corollary (ref) are sharp under the regularity assumption of Definition (ref) and under the shape restrictions on the cost function used to derive ((ref)) and formalized in the following definition.
Sharpness of the bounds is understood here in the following way. Assumptions (ref), (ref), (ref), and the conditions of Definition (ref) imply $C(y,z):=y-u^{-1}(u(y,1,z),0,z)\in\mathcal C(\pi)$. Moreover, Proposition (ref) shows that the bounds of Corollary (ref) belong to $\mathcal C(\pi)$. Finally, the utility functions $\underline u:=y-d\underline C$ and $\bar u:=y-d\bar C$ satisfy Assumptions (ref), (ref), (ref), and the conditions of Definition (ref).
In case the data generating process $\pi$ fails to satisfy the regularity condition of Definition (ref), we propose an iterative procedure to sharpen the bounds. Let $\underline C^{(0)}:= \underline C$, and $\bar C^{(0)}:=\bar C$, as defined in Equation ((ref)) of Corollary (ref). Then, for each $y\in \mathcal{Y}$, $z \in \tilde{\mathcal Z}^\prime$, and $n=1,2, \ldots$, define the sequences:
where $L_-$ is the generalized inverse of $L$ as defined in Corollary (ref). Symmetrically, for the upper bound, define the sequences
where $U^-$ is the generalized inverse of $U$ as defined in Corollary (ref).
Inference based on the sharp iterated bounds of Proposition (ref) poses challenges that are beyond the scope of this paper. In Section (ref), we therefore base inference on the bounds of Proposition (ref), which are only sharp under the additional regularity conditions of Definition (ref), but remain valid without them.
This section proposes guidelines for inference procedures based on the identification results of Sections (ref) and (ref). All proposed inference is based on an i.i.d. sample of outcomes, decisions, covariates and instruments. In Section (ref), we discuss parametric inference based on Theorems (ref) and (ref). Confidence regions for utility parameters are obtained from the inversion of tests of regression and stochastic monotonicity in the imperfect and perfect foresight cases respectively. In Section (ref), we propose pointwise confidence regions for the implied non pecuniary cost of STEM based on the closed form bounds of Corollaries (ref) and (ref) for the imperfect and perfect foresight cases respectively. We propose using the intersection bounds methodology of CLR:2009.
In Theorem (ref), we characterize the set of utility functions that rationalize the data under the extended Roy model of Assumptions (ref), (ref), and (ref). The characterization is regression monotonicity of realized utilities with respect to the instrument. Parametric inference on utilities can therefore be conducted by inverting a test of regression monotonicity. More precisely, assume utility takes the form $u(Y,D,Z;\theta)$ for some parameter vector $\theta\in\Theta$. In Section (ref), we will consider two standard parameterizations of utility, namely a quasi-linear (QL) and a constant elasticity of substitution (CES) form. In both cases, the non pecuniary component in utility is a disamenity from low representation of women among the faculty in STEM fields. The disamenity from belonging to the minority in STEM is assumed to vanish above a threshold $\gamma$ (i.e., $z\geq\gamma$). For $\alpha\in\mathbb R$, $\beta>1$, and $\gamma\in(0,0.5]$, define
Confidence regions for the true value of the parameter can be obtained by collecting all values of $\theta=(\alpha,\beta,\gamma)$ such that the hypothesis of regression monotonicity
is not rejected. Early tests of regression monotonicity are proposed in GSvV:2000 and HH:2000. Chetverikov:2013 and HLS:2016 offer power improvements.
In Theorem (ref), we characterize the set of utility functions that rationalize the data under the extended Roy model of Assumptions (ref), (ref), and (ref). The characterization is stochastic monotonicity of realized utilities with respect to the instrument. Parametric inference on utilities can therefore be conducted by inverting a test of stochastic monotonicity. More precisely, assume utility takes the form $u(Y,D,Z;\theta)$ for some parameter vector $\theta\in\Theta$. For instance, $u(y,z;\theta)$ takes quasi-linear or CES forms in ((ref)) or ((ref)) respectively. Confidence regions for the true value of the parameter can be obtained by collecting all values of $\theta$ such that a hypothesis of stochastic monotonicity
is not rejected. Early tests of stochastic monotonicity, namely LLW:2009, DE:2012 produce limiting rejection rates equal to nominal size under the least favorable DGP. Hence, they can be conservative. More recent stochastic monotonicity tests provide power improvements via pre-estimation of contact sets, as in HLS:2016 and LSW:2018, using directional differentiability of the least concave majorant operator, as in Seo:2018, or adaptivity to smoothness in the conditional cdf, as in CWK:2021.
In what follows, we test both regression monotonicity and stochastic monotonicity using HLS:2016.\footnote{\scriptsize We thank Yu-Chin Hsu, Chu-An Liu and Xiaoxia Shi for sharing their code.} The procedure is valid under continuity of the conditional mean $z\mapsto\mathbb E[u(Y,D,Z;\theta)\vert Z=z]$ and conditional cdf $z\mapsto\mathbb P(u(Y,D,Z;\theta)\leq u\vert Z=z)$ respectively. The sensitivity of inference results to the generalized moment selection procedure is usually the major concern with this type of procedure, see for instance CS:2018\footnote{\scriptsize The generalized moment selection procedure, originally introduced in Hansen:2005, GH:2009 and AS:2010, increases the power of moment inequality tests, while controlling size, by pre-selecting inequalities that are close to binding. In the specific implementation of moment inequality testing in HLS:2016, the threshold according to which moment inequalities are pre-selected depends on the user-chosen quantities $\kappa_n$ and $B_n$.}. We choose the recommended values for the user-chosen parameters governing the generalized moment selection in HLS:2016, namely $B_n=0.85\ln n/\ln\ln n$ and $\kappa_n=0.15\ln n$.
This subsection discusses inference based on the closed form bounds of Corollaries (ref) and (ref) in the imperfect and perfect foresight cases, respectively. The bounds are intersection bounds, and we apply the methodology of Section 4.3 of CLR:2009. Implementation and code are taken from CKLR:2013.
The bounds ((ref)) of Corollary (ref) are intersection bounds, and we propose applying the inference procedure proposed in CLR:2009. We discuss implementation and applicability for the lower bound $\underline C$ in ((ref)). The upper bound is similar. The lower bound $\underline C$ in ((ref)) can be written
where
using the notation of CLR:2009 (not to be confused with parameter $\theta$ in the previous section). The function $\theta(z,\tilde z)$ is estimated with a twice continuously differentiable kernel and a bandwidth sequence that guarantees undersmoothing, hence asymptotically negligible bias (we use the bandwidth recommendation on page 8 of CKLR:2013). Under the latter, Bahadur representation (2) page 1531 of KLX:2010 implies condition NK page 703 of CLR:2009 (as shown in Theorem 8 of Appendix G in the online Supplemental Material). In turn, Condition NK implies validity of the bounds by Theorem 6 page 706 of CLR:2009.
We now propose an inference procedure for the minimal cost function of Corollary (ref). The upper bound can be treated symmetrically. For any given value of $z\in\mathcal Z$, we seek a data driven function $y\mapsto C_n(y,z)$ such that for each $y\in\mathcal Y$,
for some pre-determined level of significance $\alpha$. Define $G(y\vert z,\tilde z):=\mathbb P(Y\leq y\vert \tilde z)-\mathbb P(Y\leq y,D=0\vert z)$. Call $\hat G$ a non parametric estimator for $G$, and define \[ \hat G_-(x\vert z,\tilde z):=\sup\left\{ y\in\mathcal Y \; : \; \hat G(y\vert z,\tilde z)\leq x\right\}. \] Finally, let $\hat F_1$ be a nonparametric estimator for $F_1(y\vert z):=\mathbb P(Y\leq y,D=1\vert z)$. In practice, we use nonparametric estimation procedures in LR:2008. Lemma (ref) (proved in the appendix) shows the applicability of the methodology in CLR:2009 under Condition NK page 703.
Let $s_n(y;z,\tilde z)$ be a standard error for the estimator $\hat G_-(\hat F_1(y\vert z)\vert z,\tilde z)$ and $c^\alpha_n(y;z)$ be the critical value of Definition 3 in CLR:2009. Then, under the assumptions of Theorem 6 (including Condition NK) of CLR:2009, \[ C_n(y,z):=y-\inf_{\tilde z\leq z}\left\{ \hat G_-(\hat F_1(y\vert z)\vert z,\tilde z) + c^\alpha_n(y;z)s_n(y;z,\tilde z)\right\} \] satisfies requirement ((ref)). Condition NK page 703 of CLR:2009 is a complex high-level assumption on the functions $\tilde z\mapsto G_-(F_1(y\vert z)\vert z,\tilde z)$ and $\tilde z\mapsto \hat G_-(\hat F_1(y\vert z)\vert z,\tilde z)$. Developing simpler sufficient conditions on the joint distribution of the vector $(Y,D,Z)$ for Condition NK to hold is beyond the scope of this paper.
The objective of this simulation study is to illustrate and clarify the empirical content of the main assumptions in our model, namely the stochastic monotonicity assumptions (ref) and (ref). We postulate a linear model for potential outcomes.
where $c_d=0.5d+3(1-d)$, $Z$ follows a Beta$(2,5)$ distribution, $\ln (v_0,v_1)$ follows a multivariate normal distribution N$\left(\left(
\right),\left(
\right)\right)$, and $\ln\varepsilon_d=2d+\varepsilon$, where~$\varepsilon$ is standard normal random variable. The choice of means reflects an earnings advantage for sector~1, and we let potential incomes be correlated conditionally on~$Z$. The shifter~$Z$ is observed by the agent and the analyst,~$(v_0,v_1)$ is observed by the agent, but not the analyst, and~$\varepsilon_d$ is observed by neither. We use~$\sigma=0$ to model the case, where agents have perfect foresight, and use~$\sigma=0.8$ otherwise. Since~$Y_d$ is an increasing function of~$Z$ plus noise, Assumptions (ref) and (ref) are satisfied.
We entertain the two parametric models for utility ((ref)) and ((ref)) from Section (ref), with parameter values $\alpha=1$, $\beta=0.2$ in the quasi-linear case, and $2.5$ in the CES case, and $\gamma=1$. The corresponding theoretical cost of choosing Sector 1 is:
Agents choose sector $d$ that maximizes $\mathbb E[u(Y_d,d,Z)\vert Z,v_0,v_1]$ as in Assumption (ref). We consider inference on the non pecuniary cost of Sector 1 under imperfect and perfect foresight respectively.
We first consider the analysis of imperfect foresight agents, assuming imperfect foresight. Nonparametric inference on the costs of Sector 1 under imperfect foresight is based on the bounds in ((ref)). A positive cost is detected when the conditional mean of realized incomes $\mathbb E[Y\vert Z=z]$ is not monotonic in $z$. The non monotonicity of realized incomes $\mathbb E[Y\vert Z=z]$ is visualized in Figure (ref). The empirical content of the model is visualized on Figure (ref). In the latter, the theoretical cost from Equation (ref) is plotted together with the theoretical nonparametric lower bound $\underline C(z)$ from Equation ((ref)). The performance of the inference procedure can then be assessed in Figure (ref), which plots the theoretical lower bound $\underline C(z)$ from Equation ((ref)) together with the lower envelope of the $95\%$ confidence region for each of $100$ samples of size $10,000$ each.
Here we consider the analysis of perfect foresight agents, correctly assuming perfect foresight. Nonparametric inference on the costs of Sector 1 under foresight is based on the bounds in ((ref)). A positive cost is detected when the conditional cdf of realized incomes $\mathbb P[Y\leq y\vert Z=z]$ is not monotonic in $z$. The non monotonicity of realized incomes $\mathbb P[Y\leq y\vert Z=z]$ is visualized in Figure (ref). The empirical content of the model is visualized on Figure (ref). In the latter, the theoretical cost from Equation (ref) is plotted together with the theoretical nonparametric lower bound $\underline C(y,z)$ from Equation ((ref)). The performance of the inference procedure can then be assessed in Figure (ref), which plots the theoretical lower bound $\underline C(y,z)$ from Equation ((ref)) together with the lower envelope of the $95\%$ confidence region for each of $100$ samples of size $10,000$ each. Although there is empirical content, as shown in Figure (ref), the stochastic monotonicity test we use produces overly conservative inference, as can be seen in Figure (ref).
There is a concern that basing inference on a perfect foresight assumption may reveal spurious non pecuniary costs if agents have imperfect foresight. To evaluate this possibility, we analyze inference on the non pecuniary cost of Sector 1 for imperfect foresight quasi-linear utility agents, incorrectly assuming perfect foresight. Figure (ref) shows the theoretical cost from Equation (ref) as a function of $z$, together with the theoretical lower bound $\underline C(y,z)$ from Equation ((ref)). Figure (ref) shows that $\underline C(y,z)$ can be larger than the theoretical cost, which indicates spurious cost detection due to the misspecification of the informational environment.
There is ample evidence that women are severely under-represented in STEM university majors and even more so in STEM fields (see for instance beede2011, zafar2013 and HGHM:2013). Evidence on women's reasons for shunning STEM fields (see for instance the survey in KG:2017) include mathematics gender stereotypes and gender biased amenities, such as family friendliness and work/life balance. Our objective is to document the amplitude of such non pecuniary motivations as revealed in the form of a gender-specific cost of choosing STEM fields. The revealed cost of STEM fields is a function of the rate of feminization of the STEM faculty in the region at the time of choice. We therefore also shed light on the importance of role models in the determination of major choice (KG:2017).
Our empirical analysis relies on surveys of German nationally representative university graduates. The data are collected by the German Centre for Higher Education Research and Science Studies (DZHW) as part of the DZHW Graduate Survey Series. Data and methodology are described in DZHW. The waves we consider include graduates who obtained their highest degree during the academic years 2008-2009. Graduates were interviewed 1 year and 5 years after graduation. At that point, extensive information was collected on their educational experience, employment history, including wages and hours worked, along with detailed socio-economic variables and geographical information about the region where the {\em Abitur} (high school final exam) was completed. We merge the fields of study into two categories. We call STEM the category, which consists of mathematics, physical, life and computer sciences, as well as engineering and related fields. The remaining majors are merged in the non-STEM-degree category. We only consider graduates from institutions in the country of the survey, who are active on their respective country's labor market at the time of the interview.
Our proposed stochastically monotone instrumental variables (SMIV) are the mother's education level and a variable, which we call “feminization of STEM.” The latter is defined as the proportion of women among faculty members by field of study in universities in the individual's region (Land) of residence at the time of choice, i.e., the German Land in which the individual graduated from High School, not the Land where the individual attends university. This variable is calculated for each individual in the sample from data on gender distribution of faculty by field and by Land provided by the Federal Statistical Office of Germany (DESTATIS). The data set provides for each year between 1998 and 2010, the count of faculty members (Scientific and artistic staff “{\em Wissenschaftliches und Künstlerisches Personal}”) by gender in ten fields of study. The variable is computed from the aggregation of mathematics, science and engineering.
The validity of the stochastically monotone instrumental variable rests on a combination of factors. A possible channel is the role model effect: Female students perform better in an environment with more role models. Another channel works through the provision of female specific amenities: Having more women on the faculty is likely to increase the provision of female specific amenities, which will increase women's utility in the sector with a larger proportion of women on the faculty. There is also a selection effect at work, but the latter is more ambiguous. Because of selection, female faculty in STEM may be drawn from a better selected distribution of talent than their male counterparts. Now, more talented women are likely to be more effective role models. However, the selection process we describe also means that if the overall talent distribution is the same in all locations, locations with fewer women on the faculty in STEM would also have a better selected talent distribution. Hence, the effect on role model effectiveness would be the ambiguous result of a trade-off between quality and quantity.
A more serious concern we have (as discussed in our conclusion) is the effect of aggregation into two sectors. The reason it may sometimes lead to violations of monotonicity of potential wages is that within STEM, some sub sectors may have higher wages and lower feminization. This concern is alleviated somewhat by the fact that we do not require wages to be stochastically monotone. The stochastic monotonicity assumption requires that the vector of potential utilities (functions of potential wages and amenities, including feminization) be stochastically monotone with respect to feminization of university faculty. This could be satisfied even in certain cases, where the potential wages themselves are not stochastically monotone.
Several sources of sample selection are of concern. However, given the scarce socio-economic information we have about individuals in the sample, we cannot correct for potential sample selection using matching with external data sets. First, the response rate was $20\%$ for the 2009 cohort with $14\%$ attrition in the second wave. Second, we exclude women who are still enrolled in higher education ($7.8\%$ of the sample), women who are part-time, self-employed ($6.9\%$ of the sample), and women who are unemployed or out of the labor market ($2\%$ of the sample). We keep only graduates who hold a “Bachelor”, “Magister” or “Diplom”, excluding those with “Staatsexamen” and “Lehramt” degrees, which are specific tracks mainly for teachers ($22\%$ of the sample). Finally, Heublein:2014 estimates the drop-out rate for German bachelors degrees to be between $28$ and $30\%$ during the period of interest. The dropout rate is higher for STEM fields, where the estimates range from $30$ to $39\%$. There is some evidence of higher female drop-out rates in other contexts, see Saltiel:2018 and references therein. However, we were unable to locate a reference with estimates of gender-specific drop-out rates in German universities during the period of interest.
The category of individuals we consider consists of women from the former West Germany\footnote{For the sample of women from the former East Germany, the results (omitted here for space constraints) reveal much lower costs of STEM. The difference in behavior between the former East and West may be partly attributable to differences in gender stereotypes, as evidenced in LS:2018}. The variables we consider are average income during the fifth year after graduation, which serves as outcome variable $Y$,\footnote{This choice is motivated by the available data. Lifetime earnings or wage growth opportunities may be important determinants of choice as well, as suggested by the findings in MO:2020. We are grateful to an anonymous referee for bringing the latter reference to our attention.} the choice of major $D$, which takes value $1$ if the chosen sector is STEM and $0$ otherwise, and the feminization of STEM, which serves as our SMIV instrumental variable $Z$.
Table (ref) gives numbers of STEM majors among women from the former West Germany in 2009, as compared to proportions of STEM majors among men of the same category. Figure (ref) shows the distribution of the feminization of STEM variable and the mother's education for cohorts graduating in $2009$. The relation between income, field of study and mother's education or feminization of STEM is investigated with an estimation of the propensity score (probability of choosing STEM) as a function of the feminization of STEM and the mother's education in 2009 in Figure (ref) and quartile regressions of income as a function of the feminization of STEM and the mother's education for 2009 in Figure (ref). Figure (ref) shows a positive relationship between major choice and feminization of STEM (which may or may not be causal).
Note that Figure (ref) suggests a potential violation of stochastic monotonicity of income relative to the feminization of STEM: Stochastic monotonicity of income relative to feminization would imply that the conditional probability of having an income below a given value does not increase with feminization $(P(Y\leq y\vert Z=z ) \searrow z)$. This is clearly not the case in Figure (ref). This implies a violation of the pure Roy model (without non pecuniary cost) as shown in MHM:2020. The violation is confirmed by the more formal test of the pure Roy model entertained in MHM:2020 for this category of individuals\footnote{Note that because of selection, a violation of stochastic monotonicity of realized income relative to feminization does not necessarily imply a violation of stochastic monotonicity of the vector of potential incomes relative to feminization.}.
Figure (ref) shows the lower bound of a one-sided 95% confidence region for the cost function, in the case of perfect foresight. More precisely, the left panel represents a function $C_n(y,z)$ of income and feminization of STEM, such that $\lim_{n\rightarrow\infty}\mathbb P(C_n(y,z)\leq \underline C(y,z))\geq0.95$, where $\underline C$ is the lower bound of ((ref)). The right panel of Figure (ref) shows the same, except that the feminization of STEM is replaced with the mother's education level. We observe that costs tend to be high for low income individuals and remain high for high income individuals, when the rate of feminization of STEM is low. Figure (ref) shows a different visualization the data from Figure (ref). Figure (ref) shows, for each share $c_0$ of total income, the proportion of individuals for whom the upper (resp. lower) bounds of a 95% confidence region for the cost as a share of income is greater than $c_0$. Panels (a) and (b) show this for the feminization of STEM and the mother's education, respectively.
Figure (ref) shows the estimated cost and lower bound of the 95% confidence region for the non pecuniary cost of STEM based on the imperfect foresight closed form expressions in Corollary (ref). The lower bounds include $0$, hence are uninformative. It is expected from the theory that perfect foresight bounds be tighter than the imperfect foresight ones. However, the sharp contrast we observe here between the informativeness of the perfect foresight bounds and the non informativeness of the imperfect foresight ones is a cause for concern. As we have seen in Section (ref), there is a possibility of spurious cost detection as a result of misspecification of the perfect forecast model.
We combine an extended Roy model with a stochastic monotonicity constraint. We are motivated by the need to uncover causes of under representation of women in STEM higher education. On the one hand, it is well documented that expected potential earnings are not the only drivers of major choice, and that non pecuniary, sometimes gender specific factors affect decisions. This motivates the extended Roy model of major choice, where students choose the major that maximizes a utility that depends on both earnings and amenities. On the other hand, there is a sizable literature on the positive effects of role models on choice and outcomes. This motivates the assumption that the vector of expected earnings is stochastically monotone with respect to a variable that measures the influence of role models, namely the mother's education and the proportion of faculty in STEM that are female. We derive testable implications of the model that combines the extended Roy selection with the stochastic monotonicity constraint. We also derive closed form sharp bounds on the implied non pecuniary cost of STEM for women.
Our current methodology cannot disentangle the role of real or perceived gender biased disamenities of STEM fields, in terms of family friendliness and work/life balance, from behavioral and preference biases related to gender stereotypes. An area of concern regarding the validity of our identifying stochastic monotonicity assumption is the aggregation in the STEM category, of areas such as engineering, with extremely low feminization rates and high relative incomes, with areas such as life sciences, with high feminization rates and low relative incomes. Access to more disaggregated data on the proportions of female faculty in mathematics intensive fields (rather than STEM fields as traditionally classified) would considerably alleviate this concern.