Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
101,238 characters · 13 sections · 35 citation commands
Testing the effects of an unobservable factor: Do marriage prospects affect college major choice?
Keywords: Endogenous polychotomous choice, Unobserved factor, Copula, Marriage expectation
JEL Codes: C35, C31, J12
\setcounter{page}{0} \thispagestyle{empty}
Preferences, beliefs, attitudes, expectations, and valuations are critical elements of many economic models and play significant roles in agents' decisions. However, these types of variables are usually not observable in the available data. The econometrics literature has produced various methods (e.g., instrumental variables, the control function approach) to estimate the effects of observable factors on an outcome variable in the presence of unobserved factors. While these estimation methods yield insights in some instances, they cannot be universally applied to understand the impacts of unobservable factors.
This paper focuses on analyzing the effects of an unobservable factor (i.e., a confounder variable) in a scenario involving a polychotomous choice (unordered multinomial choice) and a binary dependent outcome variable. To illustrate this concept, consider the following example related to students' choices of college majors, which we use to outline our empirical framework. In the National Longitudinal Study of Youth (97) Survey dataset (NLSY97), it is evident that college graduates with different majors have considerable differences in marital status during their 30s. This observation raises two questions: (i) does the chosen major impact the marital outcomes of college graduates? and (ii) do students make their major decisions based on the marriage prospects (preferences, expectations) associated with different college majors? While the econometrics literature proposes various approaches to address the first question (which can be considered a treatment effect), there isn't a straightforward method that can satisfactorily answer the second question. In this particular setting, the unobservable marriage prospect factor may influence both marriage outcomes and college major choices. Researchers generally have access to data on college major choices, marital status, and a rich set of control variables related to the socioeconomic status of college graduates. Yet, it remains unclear whether this data can offer insights into whether unobserved marriage prospects indeed influence students' decisions regarding their college majors.
To answer these questions discussed above, we construct an econometric model that integrates the unobservable factor into the structural error components of both the polychotomous choice model and the binary outcome model where, in addition, the binary choice model includes the polychotomous choice variable as an endogenous explanatory variable. In the example mentioned above, the binary outcome variable is marital status, while the polychotomous endogenous explanatory variable is college major. To this end, we extend evans1995finishing and altonji2005evaluation who consider a binary choice model with a binary endogenous explanatory variable to accommodate the case with a polychotomous endogenous explanatory variable. To this end, we use the insight from lee1983generalized and dahl2002mobility, who extend the Heckman type selection model with binary selection to allow for polychotomous selection. The framework we develop allows an econometric procedure to test the significance of a particular unobservable factor, which will is our main methodological contribution. In addition to providing a way to model the causal effect of college major (polychotomous) choice on the marriage decisions (binary outcome), the dependence structure between marginal conditional distributions of students' major choice and the marriage outcome variable allows us to analyze the effect of an unobservable marriage prospect factor. In this endeavor, we employ a copula to characterize the joint distribution of two conditional marginal distributions (i.e., chen2006estimation, arellano2017quantile, callaway2017quantile). As a result, our econometric model provides a framework for analyzing the interplay between major and marriage choices.
While our econometric model leads to a direct identification of the effects of major choices on marriage outcomes (the first question above), identifying the essential dependence structure of the unobservable components between major choice and marital outcomes (the second question above) requires confronting several additional challenges. The primary challenge emanates from the standard normalization requirements of underlying latent utilities within any polychotomous choice model. Specifically, in our econometric model, we normalize them by taking differences between the highest and second-highest latent utilities, and only transformed versions of the unobservable components of the choice model are identified. However, the key dependence of interest is between the untransformed unobservable components of latent utilities and the outcome variable. Because the unobservable marriage prospect factor resides within untransformed unobservable components within our model, we must introduce further structure into our econometric model to test its influence on major choices.
Our testing approach for the dependence structure between the untransformed latent utilities and the outcome variable exploits the ordered nature inherent in the latent utilities of the polychotomous choice model. Building upon siegel1993surprising and rinott1994covariance, which examine the covariance between the ordered and unordered elements of a multivariate normal vector, we extend their insights to recover the dependence between the untransformed unobservable component of latent utilities and the outcome variable, which is an important step for developing our test. In our context, latent utilities of different major options are normalized based on their order, while the potentially correlated unobservable component of the marital outcome equation remains irrelevant to latent utility orders, making it unordered. We employ the estimated covariance matrix obtained by the transformed latent utilities to infer the dependence between the untransformed unobservable component of latent utilities and the outcome variable. The correlation coefficients of the unobservable part of the marital outcome and normalized polychotomous choice model are almost surely zero when there exists no correlation between the unobservable part of the marital outcome and non-normalized polychotomous choices. Consequently, we present an econometric test that examines the correlation vector of the estimated covariance matrix of the unobservable parts of the marital outcome and the normalized polychotomous choice model. Any significant non-zero covariance provides evidence that the unobservable factors that affect the polychotomous choice also affect the binary outcome, or, in other words, that unobservable factors such as marriage preferences/expectations affect the college major choice in addition to marriage itself.
We apply our test using data from the NLYS97 to explore the potential influence of marriage prospects on college major choices. NLYS97 provides comprehensive data on participants' demographic backgrounds, marital status, high school and college-related attributes, graduates' income, and working hours, as well as inquiries regarding participants' expectations. Our test results indicate a significant link between the unobservable factors affecting marital status and major choice. We interpret these unobservable factors as marriage prospects,\footnote{To support our interpretation of the unobservable factors as marriage prospects, we exploit our access to some unique survey questions in the NLSY97 about the self-assessed marriage probabilities for respondents. When we additionally include these responses, our approach no longer detects a relationship between the unobservable factors in the marriage choice and college major choice decisions---this provides a supporting piece of empirical evidence for interpreting the unobservable factors in our main results as marriage prospects.} and, thus, our results indicate that students take into consideration their marriage prospects when they decide on their college major.
The proposed test for the effect of unobservable factors is related to several other papers that consider polychotomous choice models with selection based on unobservables. Motivated by a different application, which is to estimate hospital quality in a model with binary dependent variable (mortality) and non-random selection, geweke2003bayesian consider a binary choice model with a potentially endogenous polychotomous hospital choice variable as one of the explanatory variables. They propose a Bayesian method using a Markov chain Monte Carlo posterior simulator to estimate the model. In contrast, we propose a frequentist approach to deal with a similar model and to provide a direct test for an effect of unobservable factors.\footnote{debtrivedi consider count data models with selectivity, in which they explicitly model the endogeneity of the polychotomous variable as driven by some unobserved factor loadings that are normally distributed and also appear in the equation for the count outcome.} Notably, our test does not impose restrictions on the correlations of the outcome and latent utilities within the choice model, which is commonly used in the literature as a result of the normalization of latent utilities.\footnote{It is worth noting that our paper is also related to the work on estimating treatment effects in models featuring selectivity (e.g., heckman2004using and heckman2006understanding). Our estimation method provides an approach for estimating treatment effects when unobservable factors have an expected effect on the selection into treatment and the outcome. Consequently, our approach contributes to the toolbox of treatment effects estimation techniques in the presence of confounding variables.}
In the labor economics literature, it is common to employ dynamic choice models for the purpose of analyzing sequential decisions. For example, keane_wolphin_97,eckstein1999youths,belzil2002unobserved estimate structural dynamic models of schooling decisions and find significant positive impacts of expected earnings on college participation. arcidiocono_04, arcidiocono_05 focus on the effect of expected earnings on major choices and use structural dynamic models with new techniques. Our approach complements these dynamic choice models in that our approach provides a direct test about whether unobservable factors in the first choice (in our case, a polychotomous choice such as college major) are related to unobservable factors in the second choice (in our case, the binary outcome such as marital status).
Finally, our empirical application contributes novel insights into the determinants of college major choice. Previous studies document the effects of the non-pecuniary factors on college major choice. For instance, in the context of French universities, beffy_12 establish that non-pecuniary considerations significantly guide college major choices. wiswall_zafar_15 study the determinants of college major choice using experimentally generated beliefs and show that heterogeneous tastes are the dominant factors in the college major decision. Moreover, ersoy2022opening illustrate how students alter their college major preferences when exposed to non-earnings-related information within a staggered intervention scenario. The relationship between marriage and college participation decisions (rather than college major choice) is investigated in many theoretical papers (e.g., iyigun_walsh_07, chiappori_iyigun_weiss, gousse2017marriage, and zhang2017marriage, among others), and the effects of marriage prospects on students' college participation decisions are quantified empirically for young women (ge_11). It is noteworthy that our paper introduces the first empirical evidence pertaining to the influence of marriage prospects on college major choices.
The organization of this paper is as follows. Section 2 describes our econometric model for the polychotomous college major choice and the binary marriage choice as well as the marginal and joint model of college major choice and marriage decisions. Section 3 provides details about estimating our model. Section 4 presents the testing procedure for common/correlated unobservable factors (e.g., marriage prospects) affecting both the polychotomous choice and the binary choice. Section 5 presents the data, summary statistics, and preliminary estimation results. In Section 6, we apply the proposed test for the effects of marriage prospects on college major choice and marriage choice using NLSY97 data. Section 7 presents results for females and males separately. Section 8 concludes.
In our empirical application, individuals make selections regarding their college majors, constituting a polychotomous choice, while their marriage decisions correspond to binary outcomes. Consequently, our econometric model encompasses two distinct components: a polychotomous choice and a binary choice. Notably, the binary choice occurs after agents have encountered the consequences of their polychotomous choices. Within the context of the empirical application presented, our central focus pertains to investigating the importance of an unobserved marriage prospect factor in influencing college major choices.
College major choice - Polychotomous choice. College students typically make decisions about their college major either during the initial phases of their undergraduate education or at the point of college admissions. We assume that there are $J$ majors. Given that students weigh various factors when arriving at their major selections, the choice model must account for these considerations. In our model, we include variables that capture expected labor market outcomes and students' personal preferences to account for this complex decision-making process.
We focus exclusively on undergraduate education decisions and omit considerations related to pursuing advanced degrees or discontinuing college studies.\footnote{We do not include duration of education decision in our model compared to beffy_12. This restriction is compatible with our data since less than 5 percent of students have graduate degrees.} Within our model, student $i$ derives $V_{ij}$ utility from graduating with major $j$. We conceptualize this utility as a composite of labor market expectations and other distinctive attributes associated with the chosen major.
Labor market expectations encompass two primary components: (i) earnings and (ii) working hours. The expected labor market earnings vary based on the major $j$ and student characteristics $z_i$. We represent the average expected earnings as $\mathrm{E}[w_{ij} | z_i]$, where $w_{ij}$ is the projected earnings associated with major $j$ for student $i$. Given that our model solely considers college graduates and excludes further educational investments, we omit the opportunity cost of schooling by assuming that educational duration remains consistent across different majors. The second facet of labor market expectations pertains to working hours, a factor with pivotal implications in our model as it captures differences in employment prospects and the adaptability of working arrangements across different majors. We use $\mathrm{E}[h_{ij}|z_i]$ to denote the expected working hours associated with major $j$ for student $i$. In this context, we can encapsulate the attributed value of the expected labor market conditions of the chosen major after we normalize the earnings with working hours.
where $v_{ij}^{l}$ is the expected hourly earnings of student $i$ that chooses major $j$ and depends on students' characteristics.
The remaining part of the utility consists of all other components apart from the labor market conditions linked to major $j$. This encompasses students' inclinations towards various majors, their prospects of marriage, and other preferences exclusive to specific majors. Therefore, we express the non-pecuniary value of major $j$ for individual $i$:
where $z_{i}$ represents student $i$'s characteristics (e.g., gender, school performance variables, race, and regional characteristics), $\beta_{j}$ is the corresponding parameter vector associated with the major $j$. The term $u_{ij}$ accounts for the latent aspect of utility, encapsulating all unobservable non-pecuniary factors, including marriage expectations and leisure preferences. We posit that the latent utility $u$ remains independent of the observed characteristics $z$, which is a standard assumption in the literature on choice modeling.
Students choose a college major that maximizes their expected utility. As a consequence, student $i$ chooses major $j_i^*$ which gives the highest expected utility.
where $V_{ij} = v_{ij}^l+ v_{ij}^n = \frac{\mathrm{E}[w_{ij} | z_i]}{ \mathrm{E}[h_{ij} | z_i]} + z_i \beta_j + u_{ij}$.
Marriage decision - Binary outcome variable. Next, we consider a model for the binary outcome in our framework. In our application's context, the binary outcome variable serves as an indicator denoting whether an individual has ever been married by the age of 30. We make the assumption that marriages exclusively take place after college graduation. We model the marriage choice as a function of observed characteristics, college major, and unobservable factors; in particular,
where $m_i$ is a marriage indicator, which is equal to one if the individual is married at least once and is equal to zero otherwise. In Equation (ref), $x_i$ represents the observed attributes of graduate $i$. These attributes encompass observable individual characteristics (such as age, gender, race, and geographic location) similar to those in the major choice equation. We also include actual earnings and working hours as part of $x_i$---these two variables differ from the expected earnings and hours that were in the student characteristics ($z$) in the college major choice equation. This inclusion allows us to control for the impact of post-graduation dynamics on marriage outcomes.
The parameter $\tau_j$ represents the causal impact of major $j$ on an individual's likelihood of being married—a focal point in numerous econometric models.\footnote{Several models analyze marriage outcomes within dynamic contexts and incorporate variations in the probability of encountering potential marriage partners into their frameworks. (such as ge_11). In our particular scenario, the major dummies encapsulate all these effects.} Identifying and estimating $\tau_j$ provides an indirect method for testing the effect of marriage prospects. This is done by comparing the estimated effects with models that do not account for the unobservable effects of marriage prospects. However, our econometric procedure aims to provide a direct procedure to test the effect of (unobservable) marriage prospects on major choices.
$\epsilon_i$ encompasses the latent components of the marriage outcome equation, including the marriage prospects by design. We assume that $\epsilon$ is independent of individual characteristics $x$. This amounts to assuming that the distribution of unobservables that affect marriage choices (which includes unobserved marriage prospects) is the same across different observable characteristics. However, we refrain from imposing any assumptions regarding the relationship between college major, $j_i$, and the unobservable factors $\epsilon_i$. If students' major choice depends on their marriage prospects, $j_i$ becomes an endogenous variable. Our model allows for this possibility, and testing this hypothesis is a primary objective of our paper.
Given our interest in whether or not unobserved marriage prospects affect college major choice, in this section we consider the models for major choice in Equation (ref) and for marriage outcome in Equation (ref) jointly. Towards this end, for college major choice, we define a polychotomous variable $I_i$ based on the multinomial choice model in Equation (ref). This variable takes values $1$ to $J$, and $I_i=j^*$ if major $j^*$ is chosen. Therefore, $I_i=j^*$ if and only if
\[ \left(\frac{\mathrm{E}[w_{ij^*} | z_i]}{ \mathrm{E}[h_{ij^*} | z_i]} + z_i \beta_{j^*} + u_{ij^*} \right) \geq \underset{j\in\{\{1, \hdots, J\} \setminus j^*\}} \max{y_{ij}} \] where $y_{ij}= \frac{\mathrm{E}[w_{ij} | z_i]}{ \mathrm{E}[h_{ij} | z_i]} + z_i \beta_j + u_{ij}$ is the latent utility of chosen major $j$ for student $i$. Then,
Equation (ref) shows that we can represent a multinomial choice problem with a single inequality, rather than keeping track of $J$ latent utility equations. This simplification further underscores that error terms of multinomial choice models can only be inferred up to a transformation lee1983generalized,dahl2002mobility.
Ultimately, we will be particularly interested in the covariance matrix of the error terms linked to the binary marital outcome and the polychotomous college major choice model because any potential effects stemming from the unobservable marriage prospect factor are contained in the error terms. To advance in this direction, we model the joint distribution of $\epsilon$ and $\xi$ with the eventual goal of recovering the covariance structure of the error terms. Also note that, since the unobservables of the major choice and the marriage outcome equations are not necessarily uncorrelated in our model, estimating $\tau_j$ is not trivial and requires exogenous variation in $z$, the variables that affect major choice.
In order to recover the joint distribution of $\epsilon$ and $\xi$, $F(\epsilon, \xi)$, we first decompose the joint distribution using Sklar's Theorem (sklar1959distribution), which says that there exits a unique copula function, $C(\cdot, \cdot)$ such that $F(\epsilon, \xi)= C(F_1(\epsilon), F_2(\xi))$, where $F_1$ and $F_2$ are the marginal distributions of $\epsilon$ and $\xi$ respectively. Following lee1983generalized, we assume a Gaussian copula. Thus, we have that
where $B$ denotes the bivariate normal distribution. This distribution assumes that the transformed variables are jointly normal with zero means, unit variances, and covariance (or correlation coefficient because of unit variances) $\boldsymbol{\rho}= ( Cov(\epsilon, \xi_1),\hdots,Cov(\epsilon, \xi_J))$. The covariance vector $\boldsymbol{\rho}$ effectively encapsulates the dependence structure inherent within the error terms derived from the transformed latent variables and the binary outcome.
Next, we show that our above model can be related to the observed joint probability of being married and graduating from college major $j^*$. For example, the probability of being unmarried and having major $j^*$ is given by
{
}and, using the same sorts of arguments, the probability of being married and having major $j^*$ is given by
{
} $\mathrm{P}(m_i=0, I_i=j^* | x_i, z_i)$ and $\mathrm{P}(m_i=1, I_i=j^* | x_i, z_i)$ characterize all possible marriage outcomes for all major choices. Equations (ref) and (ref) suggest that we can use these probabilities to estimate all the parameters of the joint distribution of $F(\epsilon,\xi)$, given our model for marital status and college major.
In this section, we describe our estimation procedure, building on the discussion in Section (ref). We proceed in two steps. In the first step, we estimate the polychotomous choice model. In the second step, we simultaneously estimate the parameters from the binary choice model and the copula parameters, given the estimated parameters from the first step.
We first need to specify $F_2\Big(z_i \beta_{j^*} + \frac{\mathrm{E}[w_{ij^*} | z_i]}{ \mathrm{E}[h_{ij^*} | z_i]}\Big)$, where $F_2(\cdot)$ characterizes error components in the polychotomous choice model to estimate the Gaussian copula presented in Equation (ref). In this step, we present a multinomial probit specification for college major and marriage decisions as well as a testing procedure. Multinomial logit model can be used for the ease of computational complexity and the details of them in our setting described in Appendix (ref).
While the latent utilities are normalized with the second highest latent utility of the major choice in our econometric approach, a researcher does not know this choice without additional data on the ranking of major choices. Therefore, we assume that all majors, except the chosen one, have equal chance of being the second choice of the student. If we assume that $u_{ij},\; j=1, \hdots, J$ are distributed with jointly normal, $F_2\left(z_i \beta_{j^*} + \frac{\mathrm{E}[w_{ij^*} | z_i]}{ \mathrm{E}[h_{ij^*} | z_i]} | \theta \right) = \mathrm{P}\left(\xi_{ij^*} < z_i \beta_{j^*} + \frac{\mathrm{E}[w_{ij^*} | z_i]}{ \mathrm{E}[h_{ij^*} | z_i]} \right)$, where $\theta$ includes parameters such as $\beta$, can be represented by using all major options as;
where $\tilde{V}_{ikj^*}= z_i(\beta_k - \beta_{j^*}) + \frac{\mathrm{E}[w_{ij^*} | z_i]}{ \mathrm{E}[h_{ij^*} | z_i]} - \frac{\mathrm{E}[w_{ik} | z_i]}{ \mathrm{E}[h_{ik} | z_i]}$ and $\tilde{u}_{ikj^*}= u_{ik}- u_{ij^*}$. $A_{ij^*} \equiv \{ \tilde{u}_{ij^*} \; s.t. \; \tilde{V}_{ikj^*} + \tilde{u}_{ikj^*} \leq 0 \; \forall k\neq j^*\}$. Since this probability is a $(J-1)$ dimensional integral over error differences $A_{ij^*}$, the subsequent MLE procedure becomes computationally intensive if the the number of majors is high. This matter can be tackled by employing a simulation-based approach to handle the integrals that would otherwise be challenging to compute. However, the utilization of the simulation technique might lead to a nonsmooth objective function owing to the index function. As a result, we use GHK simulator is named after Geweke, Hajivassiliou, and Keane, which is found to be the most reliable method for simulating normal rectangle probability among others (e.g., hajivassiliou1996simulation).\footnote{The details of GHK procedure is presented in Appendix section (ref).}
Next, we need to derive and estimate $\mathrm{P}(m_i=0, I_i=j^* | x_i, z_i)$ and $\mathrm{P}(m_i=1, I_i=j^* | x_i, z_i)$ using a copula function and $F_2\Big(z_i \beta_{j^*} + \frac{\mathrm{E}[w_{ij^*} | z_i]}{ \mathrm{E}[h_{ij^*} | z_i]} \Big)$, which is available after the first step. Recall that we use the Gaussian copula for its practicality in characterizing the dependence of marginal distributions. As a result, using the equality in Equation (ref) and Gaussian copula in Equation (ref), we can write $\mathrm{P}(m_i=1, I_i=j^* | x_i, z_i)$ as the following:\footnote{Details of derivation is available in Appendix (ref).}
where $ V_{ij^*}=z_i \beta_{j^*} + \frac{\mathrm{E}[w_{ij^*} | z_i]}{ \mathrm{E}[h_{ij^*} | z_i]}$.\footnote{Note that we characterize $F_2\left(z_i \beta_{j^*} + \frac{\mathrm{E}[w_{ij^*} | z_i]}{ \mathrm{E}[h_{ij^*} | z_i]} | \theta \right) = \mathrm{P}\left(\xi_{ij^*} < z_i \beta_{j^*} + \frac{\mathrm{E}[w_{ij^*} | z_i]}{ \mathrm{E}[h_{ij^*} | z_i]} \right)$. Therefore, $F_2(V_{ij^*}) = \mathrm{P}\left(\xi_{ij^*} < z_i \beta_{j^*} + \frac{\mathrm{E}[w_{ij^*} | z_i]}{ \mathrm{E}[h_{ij^*} | z_i]} \right)$} Finally, we assume that the marginal distribution of $\epsilon_i$ as a standard normal distribution to follow a standard probit model. Then, $F_1(\cdot)=\Phi(\cdot)$, $\Phi^{-1}(F_1(\epsilon))= \epsilon$. Moreover, we can write the likelihood function based on the derivations above:
where $\theta=\{\beta, \gamma, \mathbf{\tau}, \rho \}$. Therefore, our econometric model can be estimated using maximum likelihood methods. The parameter estimates are the maximizer of the following optimization problem. \[ \hat{\theta}= \underset{\theta} \arg \max \ \sum_{i}^N \log L(\beta, \gamma,\mathbf{\tau}, \rho) \] In practice, this estimation becomes a simulated maximum likelihood estimation (SMLE) if choice probabilities are simulated by Geweke, Hajivassiliou, and Keane (GHK) simulator.
With the proposed estimation method in Section (ref), one can consistently estimate the effects of a particular college major on college graduates' marriage outcome ($\tau_j$). However, our central objective revolves around examining the impact of the unobservable marriage prospect factor on college major choices. In our econometric model, the unobservable marriage prospect factor enters into the error terms of the binary marriage outcome variable and latent utilities of the college major choices. An empirical procedure that tests the effect of the marriage prospects on major choice needs to utilize the covariance structure of these error terms and let us express this covariance as $Cov(\epsilon, u_j)$ for all $j\in 1,\hdots, J$. Since the identification of the polychotomous choice model requires transforming the latent utilities, the true covariance structure between the error terms of the marital status equation and latent utilities of major is not estimated directly, and one can only estimate the covariance between error terms of the marriage equation and transformed latent utilities ($Cov(\epsilon, \xi_j)$ for all $j\in 1,\hdots, J$). Thus, a significant challenge persists: recovering the untransformed covariance structure of the unobserved components within major choice and marital status equations from the covariance structure of the transformed error term covariance matrix originating from the binary marital status and normalized latent utilities of the college major choice models.
As a conventional practice in identifying polychotomous choice models, latent utilities are normalized by computing differences, and we focus on the difference between the highest and the second-highest of latent utilities as presented in Equation (ref). Conditional on observed characteristics of students and majors, the difference between the highest and second-highest latent utilities provides information for the normalized error terms of our econometric model. The choice associated with the highest latent utility is known by the econometrician given that it is presumed to be the agent's chosen alternative, however the choice associated with the second-highest latent utility is unknown. As a result, order statistics of latent utilities turn out to be critical elements for developing our testing procedure for unobservable factors.
Our testing procedure builds upon siegel1993surprising and rinott1994covariance, which study the covariance between variables and their order statistics for multivariate normal variables and presents the following theorem.
Based on the theorem presented, the covariance between a variable and the $r$th order statistic among the remaining variables in a multivariate normal vector is essentially a weighted average. It encompasses a set of covariances involving the selected variable and the other variables, where the weight assigned to each covariance corresponds to the probability of the selected variable being the $r$th order statistic. Consequently, the covariance between an unordered variable and an order statistic can be expressed using a collection of covariances derived from the variable in a multivariate normal vector and the probabilities associated with order statistics.\footnote{Note that, we present a weaker but general version of our testing procedure that does not require multivariate normality assumption in Appendix (ref).}
In our econometric model, our primary focus lies on $Cov(\epsilon,u)$; however, we are constrained to estimating the covariance between $\epsilon$ and the transformed latent utilities ($y_1, \hdots, y_K$), which corresponds to $Cov(\epsilon,\xi)$. As outlined by rinott1994covariance, $\epsilon$ corresponds to an unordered normal variable, while $(y_1, \hdots, y_K)$ corresponds to a vector of ordered normal variables. All these variables are constituents of a multivariate normal vector. It is noteworthy that, in contrast to rinott1994covariance, the vector of latent utilities $(y_1, \hdots, y_K)$ is solely captured through a normalization process within our model. Recalling our definition of $\xi_{ij^*}$ as follows:
In the context of the econometric framework outlined, it is necessary to apply a normality assumption not only to the error terms of the econometric model but also to the latent utilities. This ensures the congruence with the structure presented in Theorem (ref). Consequently, due to the distributional prerequisites, we posit the following assumption.
where $\mu_{\epsilon}, \mu_{y_1}, \hdots, \mu_{y_J}$ denote means of $\epsilon, y_1, \hdots, y_J$. $\rho^*_j= Cov(\epsilon, y_{j})$ is a latent covariance term. $a$ is the covariance between $y_j$ and $y_k$ when $j\neq k$. Note that the ordered nature of the latent utilities is still critical in the normalized setting because our model characterizes the $\xi_{ij}$ term using the second-highest latent utilities.
\textcolor{black}{Although Assumption (ref) may seem limiting, meeting this requirement aligns with the conditions typically encountered in multivariate normal regression or multinomial probit models, following the estimation of parameters $\theta=(\beta, \gamma, \mathbf{\tau}, \rho )$ in Section (ref). These estimates facilitate the attainment of normal latent values under the assumption of normality in error terms, enabling the prediction of latent values.}
Under Assumption (ref), we can express the covariance structure of differences of latent utilities and error term of the marriage outcome equation in the following way:
where $y^{(2)}$ indicates the second order statistic of latent utilities. Since we assume that a college graduate chooses the highest latent utility according to our behavioral model, $y_{j^*} = y^{(1)}$, i.e., the latent utility of the selected major is equal to the highest ordered latent utility, and $u_{j^*}$ is the idiosyncratic part of the latent utility.
To establish covariance between unobservable components of the econometric model presented in Equations (ref) and (ref), we consider the case that the second-highest latent utility can be any major choice except the selected one as represented in the following equation.
where $y_j$ indicates any major $j\in\{1, \hdots, J\}\setminus j^*$, and $\mathrm{P}(y_j=y^{(2)})$ denotes the probability of being the second-highest latent utility for the major $j$. Combining Equations (ref) and (ref) and using the property of covariance (i.e., $Cov(\epsilon, y_j) = Cov(\epsilon, V_j) + Cov(\epsilon, u_j)$) we can write that
In our setting, there is no endogenous variables within $V_j = \frac{\mathrm{E}[w_{ij} | z_i]}{ \mathrm{E}[h_{ij} | z_i]} + z_{i}\beta _{j}$ for $j \in \{1, \hdots, J\}$ and $Cov( \epsilon, V_j) = 0$. Because of the same reason, $Cov(\epsilon, y_j) = \rho^*_j$ becomes equal to $Cov(\epsilon, u_j)$ as presented in the Assumption (ref), which is the primary structural covariance term that we want to make inference about. Therefore, we can express $Cov (\epsilon, \xi_{j^*})$ by rearranging Equations (ref) and (ref) in the following way:
Notably, $Cov (\epsilon, \xi_{j^*})$ is derived from the estimated covariance vector ($\boldsymbol{\rho}= (\rho_1, \hdots, \rho_J) $) obtained through the model detailed in Section (ref) as $ Cov(\Phi^{-1}(F_1(\epsilon)), \Phi^{-1}(F_2(\xi_j))$ for all $j\in\{1, \hdots, J\}$. This vector plays a crucial role in inferring the primary covariance of interest, $Cov(\epsilon, u_j) = \rho^*_j$. In this procedure, the probability of a latent utility being the second-highest is a critical component for making inferences regarding $Cov(\epsilon, u_j)$. When the variance matrix structure is exchangeable, as stipulated in Assumption 1, the probability of being the second-highest latent utility is identical for all choices. Under two conditions, the estimated covariance terms of unobservable components in Equations (ref) and (ref) in our econometric model ($Cov(\epsilon, \xi_j) = \rho_j$ for all $j\in \{1, \hdots, J\}$) become zero: (i) all latent covariance terms (i.e., $Cov(\epsilon, u_j)= \rho^*_j$ for all $j \in \{1, \hdots, J\}$) are equal or (ii) all latent covariance terms are zero (i.e., $Cov(\epsilon, u_j) = \rho^*_j = 0$ for all $j \in \{1, \hdots, J\}$). Hence, we can establish an empirical test for the significance of latent covariance terms ($\rho^*_j = 0$ for all $j \in \{1, \hdots, J\}$) by evaluating whether the estimated covariance terms $\rho_j$ for all $j \in \{1, \hdots, J\}$ are statistically different from zero. If at least one of the $\rho_j$ for $j \in \{1, \hdots, J\}$ significantly deviates from zero, we can almost surely conclude that at least one $Cov(\epsilon, u_j)= \rho^*_j$ is statistically distinct from the others and from zero. This presents a significant correlation relationship between the binary outcome and polychotomous choice due to an unobservable factor.
Formally, the hypothesis for the proposed testing procedure is as follows:
This procedure provides a test to determine the presence of significant correlation patterns among unobservable factors within the binary dependent variable outcome and polychotomous choice models. These correlation patterns correspond to unobservable marriage expectations in the current setting, and, therefore, we can infer that the proposed procedure offers a test for the significance of the effect of marriage expectations on college major choices after controlling the other expected factors.\footnote{Moreover, it's possible to express the probability of a latent utility being the second-highest using the estimated parameters of the polychotomous choice model. If these probabilities exhibit variation across major options given the observed choices, it becomes feasible to further identify each $\rho^*_j$ by solving the system of equations in Equation (ref).}
We examine the relationship between marriage prospects and students' college major choices using data from the National Longitudinal Study of Youth 97 cohort (NLSY97). The NLSY97 is a longitudinal project that tracks the lives of a sample of American youth born between 1980 and 1984. The initial sample consisted of 8,984 respondents who were aged 12 to 17 during the first interview in 1997. This cohort has been surveyed annually from 1997 to 2013 and subsequently biennially.\footnote{The US Bureau of Labor Statistics https://www.nlsinfo.org.} Because the NLSY97 follows participants from an early age, it provides insight into their schooling decisions, including college major choices, as well as subsequent marriage-related information. The dataset also contains a variety of individual-level covariates that we use in our estimation procedure to control for observable characteristics. These covariates include variables like gender, age, race, regional characteristics, school performance metrics, earnings, and school-related attributes.
We focus on survey participants who graduated from higher education institutions. This selection is aligned with the objectives of our study. Initially, there were 8,984 survey participants in the dataset. However, after narrowing down the sample to include only college graduates and those for whom we have observed income and working hours data spanning from 2005 to 2011, the dataset we investigate further is composed of 1,236 observations. Table (ref) presents an overview of the characteristics of male and female participants. Notably, it's crucial to consider the age distribution within our sample. The NLSY97 cohort comprises individuals born between 1980 and 1984, resulting in an age range of 31 to 35 years during the 2015 interview. Consistent with the interview timing, we create a binary marriage outcome variable (Ever Married) to indicate whether or not a person has been married by 2015.
The proportion of female college graduates in our sample amounts to 56%, which is quite close to the population ratio of 58% according to the National Center for Education Statistics' Digest of Education Statistics (2012). However, the NLSY97 dataset exhibits a slightly higher representation of male participants compared to female participants, possibly contributing to the disparities in graduate ratios. An interesting contrast emerges in terms of marital status among graduates, particularly between females and males. In the sample we analyze, around 65% of female college graduates had been married at least once by the year 2015. In contrast, the marriage rate for male college graduates is lower by five percentage points.
Indeed, students' academic performance measures hold a crucial role in the analysis of college major choices. A study by stin_14 sheds light on the discrepancy between students' initial expectations about their final majors upon college entrance and their eventual choices. This disparity is mainly attributed to misperceptions regarding students' overall abilities to perform well. Thus, it becomes imperative to consider and account for these performance measures when investigating college major decisions. Upon examining the observed performance measures, distinct variations between male and female students come into view. Specifically, female students tend to have higher Grade Point Averages (GPAs) during their college years. This trend carries over consistently in the first three terms of college education.\footnote{ It's important to note that college admissions exam scores like SAT and ACT, as well as high school performance indicators, could potentially provide similar control variables. However, these specific variables aren't included in our analysis due to a lack of adequate data availability.} These performance measures not only reflect students' aptitude to gauge their own capabilities but also facilitate the identification of their preferences and strengths, thereby offering guidance for their ultimate college major choices. Thus, our empirical analysis takes into account these performance measures to provide a comprehensive examination of the factors influencing college major decisions.
The 2010 College Course Map serves as the basis for coding students' college majors within the NLSY97 dataset. This methodology involves utilizing information obtained from students' transcripts. In scenarios where a student possesses multiple college transcripts, the major associated with the highest number of credits is selected. The distribution of college major choices for college graduates within our sample is highlighted in Table (ref). Our focus revolves around four primary fields: Business and related studies, Health and related programs, Education programs, and Engineering and Computer Science programs. These fields have been selected due to their immense popularity and the substantial disparities between male and female enrollment rates. It's notable that these chosen college major distributions within our sample closely mirror those present within the broader population of college graduates in the United States. This compatibility ensures that our sample aptly represents the larger demographic, enhancing the robustness and generalizability of our analysis.
A substantial disparity between males and females in terms of college major choice is evident, as highlighted in the lower section of Table (ref). Notably, certain fields tend to be dominated by males, including Business, Engineering, and Computer and Information Technologies. Conversely, fields such as Education and Health and related areas are predominantly female-oriented. For instance, the data indicates that 13% of males select majors within Engineering and Computer Science-related fields, as opposed to a mere 2% of females. Conversely, Education is the major choice for 10% of females, while only 3% of males opt for this field. It's important to recognize that although the precise ratios may differ slightly from those in the general population, the overall pattern of major preferences across genders is consistent.\footnote{For a comprehensive analysis of gender differences across various majors, please refer to bronson.} This pattern of major distribution has also demonstrated stability over time, particularly between 2001 and 2013, as depicted in Figure (ref) in Appendix (ref). This figure illustrates the percentages of chosen college majors among college graduates within the NLSY97 sample across different years, with a clear division between males and females.
The NLSY97 dataset includes a range of questions that researchers can leverage to gather insight into participants' marriage prospects, alongside information on their actual marital status.\footnote{Cohabitation information is also collected, but it is not considered due to its lack of legal basis.} Table (ref) provides a breakdown of marital status and self-stated marriage expectations among college graduates based on their major and gender.\footnote{The survey question regarding marriage expectations asks participants to assess their likelihood of getting married within the next five years, in the years 2000 and 2001.} These stated marriage expectations differ not only between males and females but also among different college majors. Similarly to actual marital status in 2015, it is evident that female college graduates are more inclined to expect marriage within the next five years compared to their male counterparts. For instance, in the year 2000, 47% of female college graduates anticipated marrying within the next five years, while only 39% of male college graduates shared this expectation. Additionally, there are discernible differences in marriage expectations based on college major. Education majors, for instance, exhibit the highest likelihood of expecting marriage within the next five years, with 52% of them anticipating such in 2000. Conversely, Business majors and students majoring in Computer and Engineering programs are less inclined to expect marriage within the same timeframe, with rates of 39% and 45%, respectively. Remarkably, by the year 2015, it becomes apparent that Education majors have indeed realized their marriage expectations at a higher rate (75%) compared to other majors (falling between 62% and 66% for all other categories). The strong correlation between stated marriage expectations and actual marital status, as well as the variation observed across college majors and genders, suggests a potential relationship between marriage prospects and the choice of college major.
Next, we provide some descriptive results with two sets of regressions where the outcomes are (i) college major choice and (ii) whether or not a person has been married by 2015. These descriptive regression outcomes serve as a natural reference point for comparison against the estimates derived from our proposed methodology, which we elaborate on in the subsequent sections. Table (ref) offers an overview of the descriptive estimation outcomes concerning college major choices. To account for monetary motivations in major selection, we incorporate counterfactual normalized hourly earnings predictions that are generated using both sample data and additional exogenous variations.\footnote{The earnings estimation results are available in Table (ref). The computation of counterfactual wages includes degree types, college GPA that uses students' initial GPA as their forecast for final, and average working hours as supplementary exogenous variations. The computation of counterfactual working hours includes attendance for second college or not as additional exogenous variations. Estimation results are presented in Table (ref). Details of expected earnings and annual working hours regressions are available in Appendix (ref).} We normalize the coefficient of expected normalized earnings to 1; therefore, we can interpret results in terms of normalized earnings. Among the observed factors, gender emerges as the most influential and substantial determinant of major choices. Specifically, female students exhibit a statistically significant and quantitatively substantial inclination against selecting Computer Science and Engineering programs. Conversely, they are more inclined towards opting for Education and Health programs. Notably, performance measures play a significant role in influencing major choices. Particularly, within the first three terms of college education, students' GPA has discernible effects on their choice of major. For instance, higher GPA scores in the second semester correspond to a greater tendency to select Computer Science and Engineering programs. Similarly, students with elevated GPAs in the third semester demonstrate an increased likelihood of choosing Education programs.
Table (ref) presents descriptive regression results of marital status determinants under three different covariate sets. It's important to note that, in all three specifications, the constant term is omitted from the regressions in order to present coefficients for all possible college major choices, including other majors. Among the significant findings, a conspicuous difference in marital status emerges between graduates from the Education major and graduates from other majors. Those who majored in Education exhibit higher marriage rates than graduates from other majors, a trend observed across all three specifications. Furthermore, the descriptive analysis also underscores the disparity between females and males and the impact of average earnings on observed marital status.\footnote{Additional descriptive regression results that segment college major choices and marriage outcomes for male and female graduates are provided in Tables (ref) and (ref) in the Appendix. These outcomes elucidate the divergences between male and female students in terms of their major preferences and marital outcomes during their thirties. The disparities identified between male and female students' college major choices and marital statuses underscore the significance of unobservable factors, motivating us to carry out separate analyses for male and female cohorts.}
In this section, we put the developed estimation method into practice to assess the impact of marriage prospects on college major choices using the NLSY97 dataset. In this framework, we allow the unobservable marriage prospect factor to have a simultaneous effect on college major choice and the marital status of college graduates. To examine this, we leverage the estimated vector of correlation coefficients denoted as $\boldsymbol{\rho}$. The essence of our approach lies in testing the statistical correlation between the unobservable elements of college major choice and the observed marital status. Any statistically significant correlation coefficient within the unobservable components would suggest the presence of marriage prospect effects on college major selections.
In order to account for exogenous variation in college major choices, we incorporate college GPA from students' initial three college terms into the college major choice equations. Utilizing performance measures from the early stages of college education offers a robust source of exogenous variations for assessing college major decisions. Importantly, these measures are statistically independent of the marital status of college graduates.\footnote{Referencing the notation introduced in Section (ref), college GPA from the first three college terms constitutes part of $z$ but not of $x$.} Given that these GPAs pertain to the initial three terms, we can reasonably assume their independence from the chosen major's effect. To control for exogenous variations within the marital outcome equation, we incorporate observed average annual earnings and working hours. These variables form components of the marriage outcome equation's regressors ($x$), and they are distinct from the determinants of college major choices ($z$). This approach leverages the diversity present in the dataset, allowing us to avoid confining our identification procedure to nonlinearities.
Table (ref) presents the estimation results of the marriage outcome equation alongside the vector of correlation coefficients derived from our structural estimation approach and Table (ref) presents estimation results of the major choice component. We employ the simulated maximum likelihood estimation (SMLE) with the number of simulations set to 250. Our primary focus is centered on the estimation of $\boldsymbol{\rho}$, as it embodies the correlation between unobservable components across the observed equations. As depicted in Table (ref), $\boldsymbol{\rho}$ yields statistically significant values distinct from zero. These findings strongly suggest a significant correlation between the unobservable factors featured in the observed equations governing marital outcomes and college major choices. This statistical evidence supports the notion of a potential impact of marriage prospects on the decisions related to college major choices.
To systematically compare descriptive and structural estimation results in Tables (ref) and (ref), we present marginal effects in Table (ref) because coefficients from probit regressions do not yield immediately comparable results. Descriptive regressions in Table (ref) (column (3)) indicate that graduating with an Education major increases the chance of getting married by approximately 12 percent compared to graduating with any other major. Moreover, the marginal effects of graduating from Business, Health, Computer Science & Engineering, and other majors are closer to each other. Contrarily, the marginal effects obtained based on our structural estimates present a different story: (i) graduating with an Education major increases the chance of getting married by approximately 19 percent compared to graduating with other majors, (ii) graduating from a Business major decreases the chance of getting married approximately 25 percent, indicating that the probability of getting married varies significantly across these majors. Overall, differences in the marginal effects of reduced form and structural estimation results serve as indicators of the omitted variable bias in an estimation that does not take the correlations between unobservable factors into account.
In Table (ref), we present the marginal effect estimates derived from both the descriptive regressions (top panel) and the proposed structural estimation method (bottom panel) for the major choice. By examining the marginal effects, we can compare the magnitude of the impact of various variables on major choices and present clear interpretation of our empirical analysis. The marginal effect estimates indicate the change in probability of choosing a specific college major compared to choosing majors that are not specifically analyzed in our study (i.e., Others). For example, the marginal effect estimate for female students shows that the probability of choosing an Education field is over 8% higher than choosing other majors.
The disparities in marginal effect estimates underscore the variations in results that the descriptive regressions fail to capture. Notably, there are significant differences in the effects of expected earnings and gender on choice probabilities. For instance, female college students are more likely to choose Education fields and less likely to choose Business and Computer & Engineering fields compared to Health and other fields, with these effects ranging between 7% and 10%. Conversely, the marginal effects obtained from the reduced-form model suggest that female students are more likely to choose the Business field, with the probability of choosing Computer and Engineering fields being more than 48%.
Since our testing procedure is based on the correlation of unobservable elements of the econometric model, the results provide suggestive evidence for the effects of marriage prospects on major choice rather than conclusive findings. In this section, we take a more direct approach by using a proxy variable for marriage prospects to analyze their effects on students' college major choices. We utilize unique survey questions from the NLSY97 regarding the self-assessed probability of getting married in the next five years (i.e., stated marriage expectation variable). Since these questions were answered before students made their college major decisions, they offer valuable variation for a direct test. Therefore, we create a stated marriage expectation measure from these answers and test the effects of marriage prospects directly. We emphasize that in other applications, this sort of exercise is unlikely to be feasible, but the particular setting that we consider provides a unique opportunity to validate our empirical approach.
We implement our testing method after incorporating the stated marriage expectation variable into the existing set of covariates.\footnote{ Descriptive regression results that explore the effects of the stated major expectation measure on major choice are available in Table (ref) in Appendix (ref).} It's important to note that the proposed method takes into account the effects of unobservable factors on both marital outcomes and college major choices. These effects cannot be incorporated into the MNL regression results presented in Table (ref). By including the stated marriage expectation variable in the proposed test, we can explicitly discern the effects of marriage prospects. It's worth mentioning that marriage prospects are part of the unobservable factors discussed in the previous section. The estimation results in Table (ref) provide empirical evidence primarily for the effect of any unobserved factor on college major choices. However, by directly controlling for the stated marriage expectation variable, the proposed method enables us to test the effects of other unobservable factors specifically. This may help us exclude the effects of these other unobserved factors on college major choices.
Table (ref) displays the estimation coefficients and corresponding marginal effects of the marriage outcome equation as well as the vector of correlation coefficients using the proposed testing method, which now includes the stated marriage expectation variable in the covariate set. The absence of significant unobservable factors becomes crucial when comparing the estimation results between Tables (ref) and (ref). As the proposed procedure serves as a test for the presence of unobservable factors, its application helps us assess the significance of other unobservable factors' effects. The results presented in Table (ref) indicate that none of the correlation coefficient parameters achieve statistical significance when the stated marriage expectation variable is explicitly incorporated into our econometric model. This outcome reinforces the idea that the principal catalyst for the correlations among unobservable factors in Table (ref) is most likely the marriage prospects.
Table (ref) presents the marginal effects of selected variables associated with major choice, as obtained through our proposed estimation model. The marginal effects in the choice model can be interpreted as the change in the probability of choosing a specific major, compared to other majors, based on our estimation results. The results indicate that an increase in the expected earnings of a corresponding major significantly increases the probability of choosing that major, having the highest impact on major choice. For example, a \$1000 increase in normalized expected earnings, while holding all other earnings fixed, leads to an over 7% increase in the likelihood of selecting an Education major. This effect can reach over a 30% increase for Computer Science and Engineering majors.
We also observe that female students have different preferences, being 50% less likely to choose Computer Science and Engineering majors compared to male students, after accounting for all other factors. Unsurprisingly, early college performance significantly influences major choice and accounts for a considerable portion of the shifts in choice probabilities. Regarding the effects of marriage expectations on major choice probabilities, we observe effects of relatively smaller magnitude, yet economically important. For example, a 10% increase in the self-assessed probability of getting married in the next five years leads to a 0.4% increase in the likelihood of choosing an Education major.
The significant effects of the female variable in both the structural regression results (Tables (ref) and (ref)) and the descriptive regressions (Tables (ref) and (ref)) indicate notable differences between male and female students in their college major choices and marital status by their 30s. These disparities point to fundamental differences in their decision-making processes, potentially leading to variations in the impact of unobservable factors on their choices. To explore these differences, we conduct separate regressions for male and female participants.
We directly include the stated marriage expectation measure in our econometric framework to present clearer results.\footnote{Estimation results from our proposed estimation method without stated marriage expectation measure are available in Appendix (ref) and suggest the significant impacts of unobservable factors.} Tables (ref) and (ref) provide the estimation results for the marriage outcome with the vector correlation coefficients and the college major choice component parameters, respectively. These results continue to highlight the main insight: a higher stated marriage expectation measure significantly increases the likelihood of marriage for both male and female graduates.
Differences between male and female college graduates exist in multiple dimensions. Firstly, the effects of chosen college majors on marriage outcomes are not significant for male graduates. However, for female graduates, Education and Health majors have significantly higher marriage probabilities in their 30s. Marginal effects obtained from the presented estimations are presented in Table (ref) and provide magnitude of the corresponding variables.
In terms of the effects of marriage prospects on major choice, we present the marginal effects of selected variables associated with major choice, derived from our proposed estimation model, separated by female and male students in Table (ref). The results indicate that the ranking of the impact magnitudes of expected earnings, early GPAs, and self-assessed probability of getting married in the next five years remains consistent between male and female students; however, the magnitudes of these impacts vary. For instance, a \$1000 increase in the normalized expected earnings of Education major graduates, while holding all other earnings fixed, results in an over 6% increase in the likelihood of selecting an Education major for females. This effect can exceed a 47% increase for Computer Science and Engineering majors. For male students, the estimated marginal effect of a \$1000 increase in the normalized expected earnings of Education major graduates does not significantly affect their choice probabilities. Similar to female students, the highest variation in choice probabilities for male students occurs as a result of increased normalized expected earnings for Computer and Engineering degree graduates. A \$1000 increase in normalized expected earnings for Computer and Engineering major graduates increases the probability of selecting these majors by more than 11%.
Early college performance influences major choice, with the impacts on choice probabilities varying between male and female students. The most notable variations are observed in Education and Computer and Engineering majors. For males, a higher GPA in the first term increases the likelihood of selecting a Computer and Engineering majors, whereas for females, it decreases this likelihood. Conversely, a higher first-term GPA increases the probability of choosing an Education major for female students, but it does not significantly impact the choice probabilities for male students.
Regarding the effects of marriage expectations on major choice probabilities, we observe variations in both direction and magnitude between male and female students. For female students, a higher self-assessed probability of getting married in the next five years increases the choice probabilities for Education and Business majors. However, for male students, these effects are not positive or significant. Additionally, for both male and female students, a higher self-assessed probability of getting married in the next five years significantly decreases the likelihood of selecting a Computer & Engineering major. Another difference between male and female students is the magnitude of this effect: for female students, a 10% increase in the self-assessed probability of getting married increases the choice probability of an Education major by almost 2% and decreases the choice probability of a Computer & Engineering major by a little more than 15%. For male students, a 10% increase in this self-assessed probability leads to a non-significant change in the probability of choosing an Education major and a 0.5% decrease in the probability of choosing a Computer & Engineering major.
Finally, comparing the effects of average observed earnings and annual working hours on the observed marital status is crucial for controlling differences among socioeconomic groups and uncovering heterogeneities between males and females. Tables (ref) and (ref) show that annual working hours have opposite effects between females and males. Differences in leisure preferences or the division of labor for housework and childcare could potentially explain the disparity between males and females regarding the effect of annual working hours. If traditional roles in a marriage strongly influence marriage decisions in our dataset, it is expected that males have a higher tendency to work more to support the family, while females have more responsibility for housework and childcare, leaving them less time to work, as indicated by our empirical findings. While the proposed method may not have the power to provide detailed results, it demonstrates its practicality in obtaining empirical results that enable researchers to explore the effects of unobservable factors and make correct interpretations when unobservable factors potentially exert strong effects.
This paper presents a new method for estimating the influence of an unobservable factor that could affect both individuals' choices among multiple options and binary dependent outcome variables. Our econometric model extends a binary choice model with a binary endogenous explanatory variable to accommodate the case with a polychotomous endogenous explanatory variable by utilizing polychotomous selection models and allow us to generate a statistical test to examine covariance structure of error term of binary outcome and polychotomous choice model. Our testing method leverages the ordered nature of latent utilities within the polychotomous choice model, employing a flexible copula-based estimation process. This approach also serves as a complementary tool for dynamic choice models, providing an initial diagnostic method to investigate the impact of confounding factors within econometric models.
We apply our method to investigate the influence of marriage prospects on college major choices. We highlight disparities in the marital status of college graduates and employ an econometric model that simultaneously considers the unobservable marriage prospect factor. Our estimation results reaffirm that, as established in prior research keane_wolphin_97, arcidiocono_04, stin_14, expected earnings, initial academic performance, and gender are pivotal determinants of college major selection. However, our results also reveal that marriage prospects significantly shape college major choices, aligning with findings in the existing literature beffy_12, wiswall_zafar_15 that non-pecuniary factors hold considerable sway in these decisions. Furthermore, given the limited availability of data on certain aspects such as personality traits, life expectations, and educational preferences, our developed method offers researchers a valuable tool to explore other unobservable non-pecuniary factors, extending the possibilities for comprehensive analyses in this field.
The significant influence of marriage prospects on college major selection introduces alternative avenues for designing labor market incentives. From the policy perspective, it's crucial to recognize that students' choices in college majors directly shape skill distributions within the workforce. Any factors related to marriage and labor market conditions that impact individuals' lifestyles can significantly influence career decisions, particularly among women. Despite the increasing college participation and graduation rates among women, their labor market participation rates, in comparison to men, remain considerably lower. In light of these observations, policymakers and employers have the opportunity to create more family-friendly work environments. Such environments can serve as a magnet for women to pursue specific career paths, enhance the labor force participation of college-educated women, and help balance the distribution of skills within the labor market. By addressing these factors, it is possible to foster greater diversity and inclusion in the workforce, thereby contributing to a more equitable and dynamic labor market.