EconBase
← Back to paper

Marital Sorting, Household Inequality and Selection

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

232,181 characters · 22 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Marital Sorting, Household Inequality and Selection

abstractUsing CPS data for 1976 to 2022 we explore how wage inequality has evolved for married couples with both spouses working full time full year, and its impact on household income inequality. We also investigate how marriage sorting patterns have changed over this period. To determine the factors driving income inequality we estimate a model explaining the joint distribution of wages which accounts for the spouses' employment decisions. We find that income inequality has increased for these households and increased assortative matching of wages has exacerbated the inequality resulting from individual wage growth. We find that positive sorting partially reflects the correlation across unobservables influencing both members' of the marriage wages. We decompose the changes in sorting patterns over the 47 years comprising our sample into structural, composition and selection effects and find that the increase in positive sorting primarily reflects the increased skill premia for both observed and unobserved characteristics. \thispagestyle{empty}

Introduction

\addtocounter{page}{-1} While most empirical studies evaluate inequality via an examination of individuals' wages or earnings (see, for example, Katz and Murphy 1992, Murphy and Welch 1992, Juhn, Murphy, and Pierce 1993, Welch 2000, Autor, Katz, and Kearney 2008, Blau and Kahn 2009, Acemoglu and Autor 2011, Autor, Manning, and Smith 2016, Murphy and Topel 2016, and Fern\'{a} ndez-Val et al. 2023a, b), for many individuals their household's income is more determinative of their economic welfare. As households are composed in a variety of ways, it is useful for policy purposes to analyze income inequality for different household types. One type of particular interest is married couples comprising individuals both working full time full year (FTFY). Although the proportion of husbands and wives in the total population aged between 24 and 65 years decreased from 77.7% to 58.2% from 1976 to 2022, this group represents an increasingly larger share of married couples. In 1976 26.2% of married couples in which both the husband and the wife were aged between 24 and 65 years had both members working FTFY and by 2022 this number increased to 51.3%. As the percentage of married males in FTFY employment increased from 82.8 to 84.1 while that for married females went from 31.1 to 56.3, the growth of this group reflects the increasing employment rates of married females.

Examining the patterns and determinants of income inequality across FTFY married couples is interesting from a number of perspectives. First, studies on wage and income inequality generally do not distinguish between married and unmarried individuals. It is useful to examine if inequality among the married shows similar patterns to those of their unmarried counterparts and married individuals with non-working spouses. Second, as household income combines the earnings of both spouses, positive sorting on wages will exacerbate inequality while negative sorting will mitigate it. Finally, the large increase in the participation rates of females, with its implications for selection bias, has been shown to affect females' wages and income inequality (see, for example, Mulligan and Rubinstein 2008, Blau, Kahn, Boboshko, and Comey 2021 and Fern\'{a}ndez-Val et al. 2023b). While selection bias has been largely ignored for males, the existing evidence indicates it is not economically important. However, as increasing married female participation rates increase the sample of males in FTFY married couples, female selection into FTFY employment may have implications for the wage and earnings distributions of married males with working spouses.

There is a vast literature, employing a variety of methodological approaches, on the relationship between marriage, sorting and inequality and it provides mixed evidence on sorting behavior. Kremer (1999) documents that marital sorting on education in the United States, measured by the correlation between spouse's education levels, had declined over the period 1940-1990 and via a calibration exercise finds that marital sorting had a larger effect on intergenerational mobility than inequality. Fern\'{a}ndez and Rogerson (2001) develop and calibrate a dynamic model of \ intergenerational education acquisition, fertility and marital sorting based on the PSID and conclude that an increase in sorting is likely to increase the degree of income inequality. Greenwood, Guner, Kocharkov, and Santos (2015) examine US census data and document an increase in positive assortative matching on education and via the comparison of counterfactual income distributions conclude that the level of income inequality measured by the Gini coefficient has increased. Chiappori, Salani\'{e} and Weiss (2017) provide a model in which increased returns to education results in higher paid couples spending more time with their children and increasing assortative matching on education. Examining US marriages for individuals who are married and born between 1943 and 1972, they find that this occurred for white, but not black, individuals. Eika, Mogstad, and Zafar (2019) examine data for Denmark, Germany, Norway, the United Kingdom, and the United States and show that there is a considerable amount of educational assortative matching although it has changed little since the 1980s. Moreover, while there was an increase in sorting at the bottom of the educational distribution, there is a remarkable decrease of assortative matching at the top of the educational distribution. They conclude that the increases in the Gini coefficient cannot be explained by changes in assortative matching.\ However, they also note that assortative matching on education contributes to the cross-sectional inequality in household income in each country. This supports the earlier work of Breen and Salazar (2011) for the United States that finds a very small impact of educational sorting on changes in income inequality between the late 1970s and early 2000s. Gihleb and Lang (2018) employ a range of statistical measures of assortativeness to the CPS data for 1970 to 2010 and the 2010 American Community Survey to test for increased educational homogany among married couples in the US and matching conclude that there is no evidence of increased assortative matching. However, they do not explore sorting on wages due to the non trivial non participation of wives in market employment. Chiappori, Costa-Dias and Meghir (2020,2023) provide alternative measures of assortativeness to illustrate the difficulty in quantifying how it varies across economies with different marginal distributions of the variable on which it is being evaluated. They conclude that the degree to which educational homogany has changed among married couples in the US depends on the measure employed. The conclusion also varies on where in the educational distribution it is measured. Chiappori, Costa-Dias, Crossman and Meghir (2020) employed related measures in evaluating educational assortativeness in the U.K and conclude that there are no clear patterns. \ However, they also conclude that the changes appear to have only slightly increased income inequality. Our empirical work adds to this literature as we provide a more rigorous and detailed analysis of sorting on wages while accounting for selection although we do so by examining a more homogenous, albeit large, group of workers. However, as our focus is primarily on the role of sorting on inequality we focus on this issue rather than adopting the measures proposed by Gihleb and Lang (2018) and Chiappori and co-authors.

While earnings appear a more appropriate measure than wages for evaluating the welfare implications of inequality, examining different measures may lead to substantially different conclusions. While substantial evidence suggests female wage inequality has risen, Fern\'{a}ndez-Val, Van Vuuren, Vella and Perrachi, hereafter FVVP (2023a), show that annual earnings inequality has decreased for females in the United States due to the shifts in their annual hours of work distribution. This finding is similar in spirit to Cancian and Reed (1998a,1998b) who find household inequality decreased due to the large shifts in the female income distribution resulting from changes in their hours of work. Analyzing changes in annual income is more challenging for married couples as it requires jointly modeling annual hours and wages of both spouses. However, one could begin by restricting attention to those with a relatively homogeneous level of hours while accounting for the accompanying selection. As FTFY individuals generally work a similar numbers of hours, we can then examine income inequality via comparisons of the sum of couples' wages.

We document the changes in the earnings distributions of dual FTFY households and examine if household inequality has been exacerbated by sorting behavior. We do so via an examination of how the wage distributions of husbands and wives and sorting patterns have changed over time. We report how the probability of a male in the $j^{\mathrm{th}}$ decile of the married male wage distribution being married to a female in the $k^{\mathrm{th}}$ decile of the married female wage distribution has evolved over our sample period. We investigate the source of these changes by estimating the conditional joint distribution of husbands' and wives' wages via the bivariate distribution regression methodology of Fern\'{a}ndez-Val, Meier, Van Vuuren and Vella (2023). While this provides insight into how the observed sorting patterns can be explained by observable characteristics and unobservable factors, it does not account for the selection arising from the couple's respective employment decisions. We incorporate these employment decisions by extending the selection model of Chernozhukov, Fern\'{a}ndez-Val, and Luo (2019) (hereafter CFL) based on the Heckman (1974, 1979) selection model to bivariate selection rules. We show point identification of this model under the same exclusion restrictions as in CFL. This analysis relies on a useful result for the multivariate standard normal distribution that might be of independent interest. Lemma (ref) in the appendix characterizes the derivative of this function with respect to each element of the correlation matrix and shows it is positive. While this result is well known in the bivariate case (e.g., Sibuya, 1959) we have not found it for the general multivariate case. Using the model's estimates, we evaluate the role of the different forces generating the observed marital patterns and their implications for the couples' earnings distribution.

Our empirical investigation uncovers a number of notable findings. First, we confirm that wage inequality among couples for which both spouses are working FTFY has increased for the period 1976-2022. Second, we find increasing levels of assortative sorting on wages and that these have contributed to increasing household inequality. Third, there is mixed evidence regarding the role of selection from work decisions on observed sorting patterns. Fourth, the primary factors behind the observed positive sorting patterns are the observed characteristics of the individuals and the correlation between the unobservables driving the wages of each spouse. Finally, we find that positive sorting on wages has increased over the 47 years of our sample and this reflects the increasing market value of observed and unobserved individual characteristics. Consistent with Gihleb and Lang (2018), Eika, Mogstad, and Zafar (2019) and Chiappori, Costa-Dias and Meghir (2020,2023) our results establish that the nature of sorting on observed characteristics, such as education, has not substantially changed. However, we find that the prices of these characteristics have changed to push individuals further up, or down, their wage distributions.

We highlight three important features of our modeling approach treated as exogenous. The first is the individual's decision to marry and the second is their choice of spouse (see, for example, Chiappori, Costa-Dias and Meghir 2019). While each of these is interesting, it is beyond the scope of this paper to model these decisions. We also do not address how household income is allocated across the spouses nor the implications of this allocation for the work or marriage decisions (see, for example, Lise and Seitz 2011, Lise and Yamada 2018 and De Rock, Kovaleva and Potoms 2023). While the failure to address each of these issues represents a shortcoming of our approach, the focus on explaining the sorting pattern of spouses, and its implication for inequality, within the context of endogenous employment decisions remains important.

The following section briefly describes the Current Population Survey data examined and how the sample is selected. It also presents the time series trends in earnings and the wage distributions of FTFY households. Section (ref) describes the observed patterns of marital sorting and Section (ref) provides our econometric model. The measures of sorting are discussed in Section (ref), and estimation is discussed in Section (ref). Our empirical results are presented in Section (ref). Section (ref) concludes.

Data

Data

We employ the Annual Social and Economic Supplement (ASEC) of the Current Population Survey (CPS), or March CPS, for the 47 survey years from 1976 to 2022 which report annual earnings and hours worked for the previous calendar year.\footnote{ The data are taken from the IPUMS-CPS website maintained by the Minnesota Population Center at the University of Minnesota (Flood, King, Ruggles, and Warren 2015).} The 1976 survey is the first for which information on weeks worked and usual hours of work per week last year are available.\footnote{ We refer to the year of the survey and not the calendar year to which it refers.} To avoid issues related to retirement and ongoing educational investment we restrict attention to those aged 24--65 years in the survey year. This produces an overall sample of 2,054,502 males and 2,228,726 females. The annual sample sizes range from a minimum of 30,767 males and 33,924 females in 1976 to a maximum of 55,039 males and 59,622 females in 2001.

Annual hours worked are defined as the product of weeks worked and usual weekly hours of work last year. Those reporting zero hours generally respond not being in the labor force (i.e., they report themselves as doing housework, unable to work, at school, or retired) in the week of the March survey. We define hourly wages as the ratio of reported annual labor earnings in the year before the survey, converted to constant 2021 prices using the consumer price index for all urban consumers, and annual hours worked. Hourly wages are unavailable for those not in the labor force. As annual earnings and hours tend to be poorly measured for the Armed Forces, self-employed, and the unpaid family workers we exclude these groups and focus on civilian dependent employees with positive hourly wages and people out of the labor force last year. This restricted sample comprises 1,783,599 males and 2,097,035 females (respectively 86.8% and 94.1% of the original sample of those aged 24--65). The subsample of civilian dependent employees with positive hourly wages contains 1,540,948 males and 1,465,165 females. Married individuals aged between 24 and 65 years make up 65% of the total sample. While this has decreased from 77% in 1976 to 57% in 2022, they still represent a substantial fraction of the total sample. The percentage of married couples with both spouses working full time increased drastically from 26.2 in 1976 to 51.3 in 2022.

commentThe March CPS differs from the Outgoing Rotation Groups of the CPS, or ORG CPS, which contains information on hourly wages in the survey week for those paid by the hour and on weekly earnings from the primary job during the survey week for those not paid by the hour. Lemieux (2006) and Autor et al.\ (2008) argue that the ORG CPS data are preferable because they provide a point-in-time wage measure and workers paid by the hour (more than half of the U.S.\ workforce) may recall their hourly wages better. However, there is no clear evidence regarding differences in the relative reporting accuracy of hourly wages, weekly earnings and annual earnings. In addition, many workers paid by the hour also work overtime, so their effective hourly wage depends on the importance of overtime work and the wage differential between straight time and overtime. Furthermore, the failure of the March CPS to provide a point-in-time wage measure may be an advantage as it smooths out intra-annual variations in hourly wages.

Descriptive statistics

Figure (ref) presents the time series of various quantiles of annual household labor earnings (in 2021 dollars) for dual FTFY households. Median (Q2) household annual income increased by 29.8% from 97.9 to 127 thousand dollars from 1976 to 2022. However, growth has been more modest at lower quantiles. For example, at the first quartile (Q1), household income increased by 19.0% from 76.5 to 91.0 thousand. Moreover, household income at this quantile was virtually constant from 1976 to 2000. There were increases of 12.5% and 8.9% at the first decile (D1) and the 5th percentile respectively. Increases at higher quantiles have been notably larger. Income at Q3 grew by 49.1% from 122.1 to 182 thousand. Increases at D9 and the 95th percentile were 76.7% and 97.8% respectively. These changes have drastically increased inequality. The Q3/Q1 and D9/D1 ratios are shown in Figure (ref). The former increased from 1.60 to 2.00 and the latter from 2.43 to 3.82.

figure[figure omitted — 8,513 chars of source]
figure[figure omitted — 2,816 chars of source]

FVVP (2023a) show that there is little variation in hours for either males or females working FTFY and the income variation reflects changes in wages. Figure (ref) presents wage growth for married males and females at D1, Q1, Q2, Q3 and D9. For 1976-1996 the real wage for married males at D1 decreased by 17%. It recovers somewhat but in 2022 it remains 11% below its 1976 value. There are large decreases during the financial crisis and a large decrease in the 2010's. The real wage at Q1 for married males shows a similar pattern to that at D1. There is a large decrease in the 1976-1996 period but a recovery during the second half of the sample period despite the two large dips. The 2022 value is 9% lower than the 1976 value. The married male median shows a large decrease at 1996 but virtually no change over the sample period. The overall picture at Q3 is more positive although there are several sharp decreases. However these are offset by large gains during periods of increasing real wages. The 2022 value is 21% higher than the 1976 value. At D9 there is a similar pattern to that at Q3 but with smaller dips and faster increases. This results in a dramatic gain of 42% over the whole sample period.

The time series pattern of the married female wage distribution at the lowest decile in the earlier part of the sample period is similar to that of married males although the decreases are less dramatic. Over the period 1979-1996 it decreases by 8%. For the remaining 26 years it increases by 27% and for the whole period it increases by 19%. The pattern of real wages for married females at Q1 is dissimilar to that at D1. The trend of wage growth appears to be affected by cyclical factors but generally is steadily increasing for the time period examined. This results in an increase of 29%. Growth at Q2 is even more drastic with an increase of 41%. Moreover, with the exception of a dip in the early 1980's the median real wage for this group follows a strong upward trend. \ At Q3 there are periods of substantial gains and these offset the small decreases which are incurred. An overall gain of 60% is achieved. A similar story is observed at D9 but the increases amount to a very substantial 84%.

Given the existence of a marriage premium in mean wages it is interesting to contrast this evidence on married individuals' wage distributions to that of all individuals. Fern\'{a}ndez-Val et al. (2018) provide the corresponding rates of wage changes for all males and all females for the period 1976-2016. Median male wages decline by 13.6% for all working males and the decrease at Q1 is 18.2%. At Q3 there is very modest growth. This clearly suggests that there is a substantial difference between the experience of married and unmarried males. For all females, wage growth at Q1 and Q2 are 17% and 25% respectively and there are strong increases at Q3. Married females have experienced favorable wage growth compared to their unmarried counterparts.

The primary focus of our empirical work below is to model the joint distribution of married couple's real wage rates. To motivate what follows we report how the distribution of the sum of the husband's and wife's wages has evolved over our sample period. This is reported in Figure (ref). We define this as the household hourly wage as it captures the combined market value of an hour of work of each spouse. The household hourly wage at D1 decreases by 12% over the period 1976-1996 but increases over the remainder of the sample period. By 2022 it is 9% higher than in 1976. The household wage at Q1 shows some of the dramatic dips featured in the married males profile but over the sample period there is an increase of 15%. At Q2 it initially resembles the pattern of the male wage with multiple large dips during our sample. However, towards the end of the sample the large increases in wives' wages result in an increase of about 26% percent over the sample period. At Q3 the large shifts in the wive's wage distribution become more important. The household wage shows a dip in the 1980's and a prolonged decrease in the 1990's but over the whole sample period it increases by 43%. The trend at D9 is consistent with the male pattern at this decile combined with those of females at any of the higher quantiles. This produces a steadily rising profile and an increase of almost 70%.

figure[figure omitted — 11,959 chars of source]
figure[figure omitted — 3,804 chars of source]

Consider the implications of these trends for inequality. For males, females and households the Q3/Q1 ratios increased from 1.75, 1.80 and 1.58 in 1976 to 2.33, 2.23 and 1.96 in 2022. The D9/D1 ratios increase from 3.14. 2.90 and 2.39 in 1976 to 5.00, 4.48 and 3.72 in 2022. While we cannot directly infer anything definitive about sorting via the relative magnitudes of these measures we report them to reflect the extent of inequality. However, the rate of growth over the time period is similar for each which appears to suggest positive assortative matching. Fern\'{a}ndez-Val et al. (2018) find that for the period 1976-2016 the D9/D1 ratio for all males increases from 3.6 to 5.4 and that of females increases from 3.7 to 5. This suggests that there is less wage inequality for married individuals. This appears to primarily reflect the relatively higher wages of married individuals at low quantiles.

Marital Sorting

We capture marital sorting patterns by reporting the propensity of the male in the $j^{th}$ decile of the male wage distribution to be married to a female in the $k^{th}$ decile of the female wage distribution. We consider $ j $,$k=1\ldots 10$ producing 100 cells. To overcome small sample sizes we aggregate the data into five year intervals. To contrast the changes in cell sizes over our entire sample period, we compare the beginning and end periods corresponding to 1976-80 and 2018-22. Tables (ref) and (ref) report the contingency tables for these two periods. As random sorting is consistent with each cell containing 0.01 of the data, we divide the sample frequencies by this number. Deviations from 1 are suggestive of sorting behavior. Perfect positive assortative matching implies all couples should appear on the diagonal going from the top left to the bottom right with each element having a value of 10. Perfect negative assortative matching implies the elements on the diagonal going from the bottom left to the top right have value 10. Note that dj denotes that an observation is located between the D(j-1) to D(j) decile.

We acknowledge that a number of factors may generate departures from random sorting. For example, individuals are likely to sort on observed factors such as age, education, race and the individual's region of residence. As these factors are determinants of wages this may spuriously support positive sorting on wages. There may also be wage variations reflecting regional specific cost of living influences. While cost of living components of wages may also partially capture unobserved factors, we do not take a position on whether this necessarily reflects sorting.

For the 1976-80 period there is evidence of non-random sorting. The most striking features of Table (ref) are the cells associated with the extreme values of the joint wage distribution. The two largest observed frequencies correspond to both spouses in the bottom decile (2.19) and both in the top decile (2.39). The next two highest values correspond to the cells immediately adjacent these extremes (1.72 and 1.59). This is consistent with positive sorting. If we define the bottom as (d1+d2) and the top as (d9+d10) we obtain 6.56% and 6.60%. This contrasts with the 4% implied by random sorting. Another interesting feature of the table is revealed by the off-diagonal values. There is an almost monotonically decreasing relationship between distance from the diagonal and the probability of marriage. Finally, the sum of the elements on the diagonal is 13.78, which appears to support positive assortative matching. However, as this ignores behavior away from the diagonal and it is not invariant to the definition of the diagonal we computed the Kendall rank correlation coefficient. Its value of 0.18 supports positive assortative matching.

Table (ref) presents the corresponding values for the 2018-2022 period. As this period is associated with a substantial increase in wage inequality, the level of household inequality is likely to have increased even if the cell frequencies remained at the 1976-80 values. However, it appears that increased marital sorting has exacerbated the individual level inequality. For example, the two extreme cells have increased to 2.94 and 3.25. This represents a particularly large increase in the d1/d1 cell. The frequencies in the bottom 2 and top 2 have increased to 8.13 and 8.48. This reflects a large growth in the frequency in which lower (higher) paid males are married to lower (higher) paid females. The growth in the positive sorting at the bottom is concerning given the manner in which wages, relative to 1976, have fallen for males and only slightly increased for females at these locations of their distributions. The other features of Table (ref) regarding sorting are also generally supported by Table (ref). There is greater evidence that the lowly paid are married to the lowly paid while the highly paid are married to the highly paid. The sum of the diagonal is now 17.24 and the Kendall rank correlation coefficient is 0.27. This suggests increased positive assortative matching relative to the 1976-80 period.

table[table omitted — 1,468 chars of source]
table[table omitted — 1,465 chars of source]

The differences across the two tables are important given their potential implications for inequality. However, the differences might reflect a variety of factors. As employment rates have changed substantially over the sample period, the composition, in terms of observed and unobserved characteristics, of both husbands and wives may have changed and this may have affected the observed patterns of sorting. It is also possible that the prices of these observed and unobserved characteristics have changed. Finally, the nature of marital sorting, as measured by the joint distribution of the couples observed and unobserved characteristics, may have changed.

Many couples share characteristics which are determinants of wages, such as race, geographical location and age, although this does not necessarily reflect sorting on wages. Another important characteristic is education although this might be more reasonably interpreted as indicative of productivity. To examine how these characteristics are allocated across cells, Tables (ref)-(ref) report the average value of some measure of each of these characteristics for the same time periods.

Tables (ref)-(ref) report the ages of wives and husbands. As age is a determinant of wages and spouses generally are close in age, it is possible that the patterns in Tables (ref) and (ref) simply reflect age differences. For the first period the average age for both husbands and wives increases by about 2 years as one goes from the d1/d1 to d10/d10. This seems a remarkably small difference given the large wage discrepancies across these cells. The highest husband age is associated with the highest male decile and the highest ages of wives are for the lowest paid women marrying these men. While this is an interesting result it is beyond the scope of the paper to investigate it further. There are differences across the cells, but it does not appear that age differences are the factors driving the observed sorting. For the later period the average ages of the spouses increase by approximately 2.5 years as one goes from d1/d1 to d10/d10. In percentage terms this is a notable increase compared to the earlier period although it does not appear to be the driving force of increased inequality across the two periods. The oldest males continue to be the highest paid married to the lowest paid women and the oldest wives are those married to these men. The age differences across cells is larger than in the earlier period but are unlikely to explain the observed wage differences.

Tables (ref) to (ref) report the fraction of wives and husbands with university education in the two time periods. Unsurprisingly, educational levels in the higher wage cells are higher than those in the lower. The large increase in individuals obtaining college education is reflected in the substantially higher averages for the later period. We also report the percentage of households in each cell for which both spouses have university education. For the earlier period the product of the two d10 cells is .353(.552*.64) although the corresponding cell in Table (ref) is .467. This difference is supportive of couples sorting on education at high wages. For the later period the corresponding numbers are .868(.938*.926) and .897. Thus while there are more married couples with both spouses university educated, there appears to be more positive sorting on education at high wages in the earlier period. The product of the d1 cells for the first period is .004(.064*.069) while for Table (ref) it is .034. For the later period the corresponding numbers are .012(.107*.115) and .059. This also suggests there is greater positive sorting on education in the earlier period. Tables (ref) to (ref) unsurprisingly indicate that race has some association with an individual's location in each of the wage distribution and that the strength of this relationship has changed drastically over time. Moreover, Tables (ref)- (ref) collectively confirm positive sorting on race. This highlights the necessity to account for these factors in estimating the sorting models below. The final characteristic we consider is the location of residence. This is also likely to affect wage differences as it may capture cost of living differences. Tables (ref) and (ref) suggest that the higher paid are living in metropolitan areas and that the fraction of individuals living in these areas has increased substantially over the sample period. As married individuals generally cohabitate this may also spuriously imply sorting on wages.

Econometric model

Determinants of Marital Sorting

We now focus on estimating the determinants of the marital sorting frequencies in Tables (ref) and (ref). We do so via the bivariate distribution regression (BDR) approach of Fern \'{a}ndez-Val et al. (2023) which employs a Local Gaussian Representation (LGR) of the joint distribution of the wives' and husbands' wages (Chernozhukov, Fern\'{a}ndez-Val, and Luo 2019). We represent this joint distribution as:

equation[equation omitted — 184 chars of source]

where $Y_{w}$ and $Y_{h}$ are the observed wages of wives and husbands, $X$ is a set of observed characteristics, $\Phi _{2}(\cdot ,\cdot ;\rho )$ is the standard bivariate normal CDF with correlation $\rho ,$ and $\mu _{j}(y,x)$ is formally defined below. In the LGR, the marginal conditional CDFs of $Y_{w}$ and $Y_{h}$ are represented by:

equation*[equation* omitted — 104 chars of source]

where $\Phi $ is the standard univariate normal CDF. The parameter $\rho (y_{w},y_{h},x)$ is the local correlation between the unobservables influencing the spouses' wages at $(y_{w},y_{h})$ and captures sorting on unobservables. The unconditional joint distribution of the wives' and husbands' wages can be obtained from the LGR as:

equation*[equation* omitted — 167 chars of source]

where $F_{X}$ is the CDF of $X$, and the corresponding marginals are:

equation*[equation* omitted — 106 chars of source]

Chernozhukov, Fern\'{a}ndez-Val, and Luo (2019) show that the LGR is non-parametric as it does not impose any restrictions on the conditional joint distribution.

The BDR model augments the LGR with two assumptions:

assumption[BDR] (1) $\mu_j(y,x) = P_{j}(x)^{\prime }\beta _{j}(y),$ $j \in \{w,h\},$ where $P_{w}$ and $P_{h}$ are transformations of $x$ and $\beta _{w}(y)$ and $\beta _{h}(y)$ are vectors of coefficients; and (2) $\rho (y_{w},y_{h},x) = \rho(y_{w},y_{h})$.

We allow different specifications for $P_{w}$ and $P_{h}.$ For example, $ P_{w}$ includes the wife's education and age while $P_{h}$ does not. While we assume that the marital sorting parameter does not depend on observed characteristics, we allow it to vary by location in the joint distribution of wages. This model is semiparametric as the parameters $y\mapsto \beta _{j}(y),$ $j\in \{w,h\},$ and $(y_{w},y_{h})\mapsto \rho (y_{w},y_{h})$ are function-valued.

The BDR model describes the joint distribution of wages conditional on the two indices capturing the observed determinants of husbands and wives' wages. Given a random sample of $(Y_{w},Y_{h},X)$, estimation is performed via a series of bivariate probits of the indicators $\mathbf{1}(Y_{w}\leq y_{w})$ and $\mathbf{1}(Y_{h}\leq y_{h})$ on $X$ for multiple values of $ y_{h}$ and $y_{w}$. This is done in two steps. First, estimate $\beta _{j}(y_{j})$ via univariate probit of $\mathbf{1}(Y_{j}\leq y_{j})$ on $X$, for $j\in \{w,h\}$. Second, estimate $\rho (y_{w},y_{h})$ via bivariate probit plugging-in the estimates of $\beta _{w}(y_{w})$ and $\beta _{h}(y_{h})$.

We begin by examining our capacity to explain the variation in Tables (ref) and (ref) via 100 bivariate probits. The entries correspond to estimates of:

multline[multline omitted — 661 chars of source]

where $\underline{y}_{w}$ and $\overline{y}_{w}$ are evaluated at the sample deciles of $Y_{w}$, and $\underline{y}_{h}$ and $\overline{y}_{h}$ at the sample deciles of $Y_{h}$. The predicted probabilities are reported in Tables (ref) for 1976-1980 and Table (ref) for 2018-2022. A comparison of the predicted and empirical probabilities for each of the sample periods indicates that the estimated model reproduces the empirical probabilities across the 100 cells despite the restrictions imposed by Assumption (ref).\footnote{ The predicted probabilities are computed by first using univariate distribution regression models to obtain the quantiles of the marginal distribution of wives and husbands. That is, we estimate the quantiles for wives, $Q_{\tau }^{w};\tau \in \lbrack 0,1]$ solving the empirical analog of $\int \Phi (x\beta (Q_{\tau }^{w}))dF_{X_{w}}(x)=\tau $, where $\beta (\cdot )$ is estimated by univariate distribution regression.The quantiles for husbands are estimated similarly. Based on these quantiles, we can estimate the bivariate distribution regression model as in ((ref)) and estimate the empirical analog of $\int \Phi _{2}(x_{w}\beta (Q_{\tau }^{w}),x_{h}\beta (Q_{\tau }^{h}),\rho )dF_{X_{w},X_{h}}(x_{w},x_{h})$.}

Figure (ref) presents the estimated densities of the 100 estimates of $\rho (y_{w},y_{h})$ for each sample period. The average values for the two periods are $0.18$ and $0.24$ and the figure reveals a clear shift in the distribution to the right over time. Thus, despite the rich nature of $X$, there remains a strong and positive correlation between the unobservables driving wages for spouses and it has increased over time. \footnote{$X$ includes 3 dummy variables for education (high school, some college, college degree or higher), age, age$^{2}$, age interacted with 3 education dummy variables, $age^{2}$ interacted with 3 education dummy variables, a dummy variable for non-white, a dummy variable for Hispanic, 2 dummy variables for metropolitan area (central city, outside central city), and 7 regional dummy variables (middle Atlantic, east north central, west north central, south atlantic, east south central, west south central, mountain, and pacific. base: New England).} This supports the presence of positive and increasing assortative matching on wages. While this may partially reflect common influences on wages, such as local adjustments related to cost of living, it may also capture unobserved ability or other unobserved determinants of productivity.

figure[figure omitted — 2,905 chars of source]

While the estimate of $\rho (y_{w},y_{h})$ is consistent with positive sorting, it is not immediately clear from the coefficient how economically important this parameter is in determining the observed patterns. We employ the estimates from the BDR model to construct counterfactual contingency tables from setting $\rho (y_{w},y_{h})=0$ for all $y_{w},y_{h}$.\footnote{ As the quantiles of the marginal distributions are not affected by setting $ \rho (y_{w},y_{h})$ to zero, we can simply estimate the corresponding distributions in (ref) by setting $\rho (y_{w},y_{h})$ equal to zero.} This corresponds to no correlation or sorting on unobservables while retaining the sorting on the observed characteristics. We implement this counterfactual in the two periods and report the estimates in Tables (ref) and (ref). The predicted probabilities indicate that sorting on unobservables accounts for a large fraction of the observed positive marital sorting. This is remarkable given the positive sorting that occurs on race, age, educational attainments and location of residence highlighted above. Even though it is hard to make any conjectures about the source of the unobserved heterogeneity, Eika, Mogstad, and Zafar (2019) show for Norwegian data that the choice of college major is an important source of educational sorting. This would be captured in $\rho (y_{w},y_{h}).$

table[table omitted — 1,494 chars of source]
table[table omitted — 1,500 chars of source]

Marital Sorting with Employment Selection

The evidence above is based on the subpopulation of working married couples as the BDR model above does not account for selection into employment. As our sample period witnessed drastic increases in the market participation of married females it is possible that both the compositions of the subpopulations of working married females and the working married males with working spouses have changed. We now incorporate the participation decision of each spouse, while allowing for a relationship between these decisions, by extending the BDR model to incorporate endogenous employment decisions. This represents an application of the CFL approach with some provisos we outline below. We begin with some necessary preliminaries.

The sample selection model

Consider the vector of random variables $(D_{w}^{\ast },D_{h}^{\ast },Y_{w}^{\ast },Y_{h}^{\ast })$, where $D_{w}^{\ast }$ and $D_{h}^{\ast }$ are latent variables that determine the employment decision of the wife and husband, and $Y_{w}^{\ast }$ and $Y_{h}^{\ast }$ are the offered wages to the wife and husband. $Z$ is a vector of observed characteristics of which $ X $ is a subset . We make the following assumptions regarding the joint CDF of $(D_{w}^{\ast },D_{h}^{\ast },Y_{w}^{\ast },Y_{h}^{\ast })$ conditional on $Z $:

assumption[LGR, Relevance and Exclusion] (1) For $(d_{w},d_{h},y_{w},y_{h}) \in \mathbb{R}^4,$ \begin{equation*} F_{D_{w}^{\ast },D_{h}^{\ast },Y_{w}^{\ast },Y_{h}^{\ast } \mid Z}(d_{w},d_{h},y_{w},y_{h} \mid z) = \Phi _{4}(\boldsymbol{\mu} (d_{w},d_{h},y_{w},y_{h},z); \boldsymbol{\Sigma} (d_{w},d_{h},y_{w},y_{h},z)), \end{equation*} where $\Phi _{4}(\cdot; \boldsymbol{\Sigma})$ is the standard tetravariate normal CDF with correlation matrix $\boldsymbol{\Sigma}$, \begin{equation*} \boldsymbol{\mu} (d_{w},d_{h},y_{w},y_{h},z)=\left( \begin{array}{c} \mu _{D_{w}^{\ast}}(d_{w},z) \\ \mu _{D_{h}^{\ast}}(d_{h},z) \\ \mu _{Y_{w}^{\ast}}(y_{w},z) \\ \mu _{Y_{h}^{\ast}}(y_{h},z) \end{array} \right) \end{equation*} and \begin{multline*} \boldsymbol{\Sigma} (d_{w},d_{h},y_{w},y_{h},z)= \\ \left( \begin{array}{cccc} 1 & \rho _{D_{w}^{\ast },D_{h}^{\ast }}(d_{w},d_{h},z) & - \rho _{D_{w}^{\ast },Y_{w}^{\ast }}(d_{w},y_{w},z) & - \rho _{D_{w}^{\ast },Y_{h}^{\ast }}(d_{w},y_{h},z) \\ \rho _{D_{w}^{\ast },D_{h}^{\ast }}(d_{w},d_{h},z) & 1 & - \rho _{D_{h}^{\ast },Y_{w}^{\ast }}(d_{h},y_{w},z) & - \rho _{D_{h}^{\ast },Y_{h}^{\ast }}(d_{h},y_{h},z) \\ - \rho _{D_{w}^{\ast },Y_{w}^{\ast }}(d_{w},y_{w},z) & - \rho _{D_{h}^{\ast },Y_{w}^{\ast }}(d_{h},y_{w},z) & 1 & \rho _{Y_{w}^{\ast },Y_{h}^{\ast }}(y_{w},y_{h},z) \\ - \rho _{D_{w}^{\ast },Y_{h}^{\ast }}(d_{w},y_{h},z) & - \rho _{D_{h}^{\ast },Y_{h}^{\ast }}(d_{h},y_{h},z) & \rho _{Y_{w}^{\ast },Y_{h}^{\ast }}(y_{w},y_{h},z) & 1 \end{array} \right) \end{multline*} is non-singular almost everywhere in $z$; (2) $\mu _{D_{w}^{\ast}}(d_{w},z) \neq \mu _{D_{w}^{\ast}}(d_{w},z^{\prime }) $ and $\mu _{D_{h}^{\ast}}(d_{h},z^{\prime \prime }) \neq \mu _{D_{h}^{\ast}}(d_{h},z^{\prime \prime \prime })$ for some $z = (z_1,x)$ , $ z^{\prime } = (z_1^{\prime },x)$, $z^{\prime \prime }= (z_1^{\prime \prime },x)$, and $z^{\prime \prime \prime} = (z_1^{\prime \prime \prime},x)$; and (3) $\mu _{Y_{w}^{\ast }}(y_{w},z) = \mu _{Y_{w}^{\ast }}(y_{w},x)$, $\mu _{Y_{h}^{\ast }}(y_{h},z) = \mu _{Y_{h}^{\ast }}(y_{h},x)$, $ \rho_{D_{w}^{\ast },Y_{w}^{\ast }}(d_{w},y_{w},z) = \rho _{D_{w}^{\ast },Y_{w}^{\ast }}(d_{w},y_{w},x)$, and $\rho _{D_{h}^{\ast},Y_{h}^{\ast }}(d_{h},y_{h},z) = \rho _{D_{h}^{\ast },Y_{h}^{\ast }}(d_{h},y_{h},x)$, for $z = (z_1,x)$.

Assumption (ref)(1) is similar to the LGR of a joint CDF. In contrast to the bivariate case, this representation restricts some features of the joint tetravariate distribution. While the univariate and bivariate marginals remain unrestricted, some restrictions are imposed on the trivariate marginals and joint tetravariate distributions. To highlight this, note that the local dependence between any pair of random variables, as measured by the corresponding component of the matrix $\boldsymbol{\Sigma }(d_{w},d_{h},y_{w},y_{h},z)$, does not depend on the value of the other components. For example, local pairwise independence of all the components, that is $\boldsymbol{\Sigma }(d_{w},d_{h},y_{w},y_{h},z)$ equal to the identity matrix, implies joint local independence of all the components.

Assumption (ref)(3) embodies exclusion restrictions on the marginal distributions of $Y_w^*$ and $Y_h^*$, and the local dependence matrix $ \boldsymbol{\Sigma} (d_{w},d_{h},y_{w},y_{h},z)$. Thus, $Y_w^*$ and $Y_h^*$ are independent of the components of $Z$ not included in $X$. Moreover, these components do not affect the local dependence between all the components in $(D_{w}^{\ast },D_{h}^{\ast },Y_{w}^{\ast },Y_{h}^{\ast })$. Assumption (ref)(2) is a relevance condition of these excluded components of $Z$ on the expectations of $D_w^*$ and $D_h^*$.

The observed variables $(D_{w},D_{h},Y_{w},Y_{h})$ are related to the latent variables as:

equation*[equation* omitted — 127 chars of source]

where $D_{w}$ ($D_{h})$ equals 1 when the wife (husband) is working FTFY and 0 otherwise. Moreover,

equation*[equation* omitted — 165 chars of source]

where $Y_{w}$ and $Y_{h}$ are the FTFY hourly wages of the wife and husband. These are only observed for FTFY working couples.

We show in Appendix (ref) that all the parameters of the LGR of the latent variables are identified from the distribution of the observed variables.

theorem[Identification Under Employment Selection] Under Assumption (ref), $\boldsymbol{\mu} (0,0,y_{w},y_{h},z)$ and $\boldsymbol{\Sigma} (0,0,y_{w},y_{h},z)$ are identified from the joint distribution of $(D_{w},D_{h},Y_{w},Y_{h},Z)$.
comment\subsection{Identification} We define $\nu _{D_{w}^{\ast }}:=\nu _{D_{w}^{\ast }}(1)$, $\nu _{D_{h}^{\ast }}:=\nu _{D_{h}^{\ast }}(1)$ and $\rho _{D_{w}^{\ast },D_{h}^{\ast }}:=\rho _{D_{w}^{\ast },D_{h}^{\ast }}(d_{w},d_{h})$. The parameters $\nu _{D_{w}^{\ast }}(z)$, $\nu _{D_{h}^{\ast }}(z)$ and $\rho _{D_{w}^{\ast },D_{h}^{\ast }}(z)$ are identified through the following system of equations: \begin{equation*} \begin{split} \mathbb{P}(D_{w}& =1;D_{h}=1)=1-\Phi _{2}(\nu _{D_{w}^{\ast }}(z),\nu _{D_{h}^{\ast }}(z),\rho _{D_{w}^{\ast },D_{h}^{\ast }}(z)) \\ \mathbb{P}(D_{w}& =1;D_{h}=0)=\Phi (\nu _{D_{h}^{\ast }}(z))-\Phi _{2}(\nu _{D_{w}^{\ast }}(z),\nu _{D_{h}^{\ast }}(z),\rho _{D_{w}^{\ast },D_{h}^{\ast }}(z)) \\ \mathbb{P}(D_{w}& =0;D_{h}=1)=\Phi (\nu _{D_{w}^{\ast }}(z))-\Phi _{2}(\nu _{D_{w}^{\ast }}(z),\nu _{D_{h}^{\ast }}(z),\rho _{D_{w}^{\ast },D_{h}^{\ast }}(z)) \end{split} . \end{equation*} We have the following three restrictions for the distribution of $Y_{w}$ and $Y_{h}$ for the selected sample: \begin{equation*} \begin{split} \mathbb{P}(D_{w}& =1;D_{h}=1,Y_{w}\leq y_{w},Y_{h}\leq y_{h}) \\ & =\int_{\nu _{D_{w}^{\ast }}(z)}^{\infty }\int_{\nu _{D_{h}^{\ast }}(z)}^{\infty }\int_{-\infty }^{\mu _{Y_{w}^{\ast }}(y_{w},z)}\int_{-\infty }^{\mu _{Y_{h}^{\ast }}(y_{h},z)}d\Phi _{4}(u_{1},u_{2},u_{3},u_{4},\Sigma (y_{1},y_{2},z)) \\ \mathbb{P}(D_{1}& =1;D_{2}=1,Y_{w}\leq y_{1},Y_{h}>y_{2})= \\ & \int_{\nu _{D_{w}^{\ast }}(z)}^{\infty }\int_{\nu _{D_{h}^{\ast }}(z)}^{\infty }\int_{-\infty }^{\mu _{Y_{w}^{\ast }}(y_{w}(z))}\int_{\mu _{Y_{h}^{\ast }}(y_{h}(z))}^{\infty }d\Phi _{4}(u_{1},u_{2},u_{3},u_{4},\Sigma (y_{w},y_{h},z)) \\ \mathbb{P}(D_{1}& =1;D_{2}=1,Y_{w}>y_{w},Y_{h}\leq y_{h})= \\ & \int_{\nu _{D_{w}^{\ast }}(z)}^{\infty }\int_{\nu _{D_{h}^{\ast }}(z)}^{\infty }\int_{\mu _{Y_{w}^{\ast }}(y_{w},z)}^{\infty }\int_{-\infty }^{\mu _{Y_{h}^{\ast }}(y_{h},z)}d\Phi _{4}(u_{1},u_{2},u_{3},u_{4},\Sigma (y_{w},y_{h},z)) \end{split} \end{equation*} where $\Sigma (y_{w},y_{h},z):=\Sigma (0,0,y_{w},y_{h},z)$. The 7 remaining unknowns are: $\mu _{Y_{w}^{\ast }}(y_{w},z)$, $\mu _{Y_{h}^{\ast }}(y_{2},z) $, $\rho _{D_{w}^{\ast },Y_{w}^{\ast }}(y_{w},z)$, $\rho _{D_{w}^{\ast },Y_{h}^{\ast }}(y_{h},z)$, $\rho _{D_{h}^{\ast },Y_{w}^{\ast }}(y_{w},z)$, $ \rho _{D_{h}^{\ast },Y_{h}^{\ast }}(y_{h},z)$ and $\rho _{Y_{w}^{\ast },Y_{h}^{\ast }}(y_{w},y_{h},z)$ and these parameters are not point identified without further exclusion restrictions. Define $X\subset Z$ as the subset of the characteristics $Z$. Define $Z:=(X,\widetilde{Z})$. We make the following assumptions about and excluded random variable $ \widetilde{Z}$: \begin{enumerate} • $0 < \mathbb{P}(D_w=1;D_h=1|X=x) < 1$ and $0 < \mathbb{P}(\widetilde{Z} =z|D_w=1,D_h=1,X=x) < 1$ for $z \in {z_0,z_1,z_2}; z_0 < z_1 < z_2$. • $\mathbb{P}(D_w=1;D_h=1|\widetilde{Z}=z_0, X=x) < \mathbb{P} (D_w=1;D_h=1|\widetilde Z=z_1,X=x) < \mathbb{P}(D_w=1;D_h=1|\widetilde Z=z_2) < 1$. • $\mu_{Y^*_w}(y_h, z) = \mu_{Y^*_w}(y_1,x)$ and $\mu_{Y^*_h}(y_h, z) = \mu_{Y^*_h}(y_2,x)$ for all $y_h \in \mathbb{R}$, $y_w \in \mathbb{R}$ and $ z \in \{z_0,z_1,z_2\}$. • $\rho_{D^*_w, Y^*_w}(y_h,z) = \rho_{D^*_h, Y^*_w}(y_1,x)$, $ \rho_{D^*_w, Y^*_h}(y_2,z) = \rho_{D^*_w Y^*_h}(y_h,x) $, $\rho_{D^*_h, Y^*_w}(y_w,z) = \rho_{D^*_h Y^*_w}(y_1,x)$, $\rho_{D^*_h, Y^*_w}(y_h,z) = \rho_{D^*_h, Y^*_w}(y_w,x)$ and $\rho_{Y^*_w, Y^*_h}(y_w,y_h,z) = \rho_{Y^*_w, Y^*_h}(y_w,y_h,x)$ for all $y_w \in \mathbb{R}$, $y_h \in \mathbb{R}$ and $z \in \{z_0,z_1,z_2\}$. \end{enumerate} The first restriction implies that $Z$ has at least three mass points on the selected sample of both partners working. The second restriction is a combined monotonicity assumption of the exclusion restriction and a restriction that implies relevance. The third and fourth assumptions form the basis of the exclusion restriction implying that the distribution of the wages are not directly affected by the outcome of $Z$. Using these restrictions we write the model as: \begin{equation*} F_{D_{w}^{\ast },D_{h}^{\ast },Y_{w}^{\ast },Y_{h}^{\ast }|Z}(0,0,y_{w},y_{h}|z)=\Phi _{4}(\mu (y_{w},y_{h},z),\Sigma (y_{w},y_{h},z)) \end{equation*} with \begin{equation*} \mu (y_{w},y_{h},z)=\left( \begin{array}{c} \nu _{D_{w}^{\ast }}(z) \\ \nu _{D_{h}^{\ast }}(z) \\ \mu _{Y_{w}^{\ast }}(y_{w},x) \\ \mu _{Y_{h}^{\ast }}(y_{h},x) \end{array} \right) \end{equation*} and \begin{equation*} \Sigma (y_{w},y_{h}|z)=\left( \begin{array}{cccc} 1 & \rho _{D_{w}^{\ast },D_{h}^{\ast }}(z) & \rho _{D_{w}^{\ast },Y_{w}^{\ast }}(y_{w},x) & \rho _{D_{w}^{\ast },Y_{h}^{\ast }}(y_{h},x) \\ \rho _{D_{w}^{\ast },D_{h}^{\ast }}(z) & 1 & \rho _{D_{h}^{\ast },Y_{w}^{\ast }}(y_{w},x) & \rho _{D_{h}^{\ast },Y_{h}^{\ast }}(y_{h},x) \\ \rho _{D_{w}^{\ast },Y_{w}^{\ast }}(y_{w},x) & \rho _{D_{h}^{\ast },Y_{w}^{\ast }}(y_{w},x) & 1 & \rho _{Y_{w}^{\ast },Y_{h}^{\ast }}(y_{w},y_{h},x) \\ \rho _{D_{w}^{\ast },Y_{h}^{\ast }}(y_{h},x) & \rho _{D_{h}^{\ast },Y_{h}^{\ast }}(y_{h},x) & \rho _{Y_{w}^{\ast },Y_{h}^{\ast }}(y_{w},y_{h},x) & 1 \end{array} \right) . \end{equation*} The parameters $\nu _{D_{w}^{\ast }}(z_{j})$, $\nu _{D_{h}^{\ast }}(z_{j})$ and $\rho _{D_{w}^{\ast },D_{h}^{\ast }}(z_{j})$ for $j=0,1,2$ are identified by the following set of equations: \begin{equation*} \begin{split} \mathbb{P}(D_{w}& =1;D_{h}=1|Z=z_{j})=1-\Phi _{2}(\nu _{D_{w}^{\ast }}(z_{j}),\nu _{D_{h}^{\ast }}(z_{j}),\rho _{D_{w}^{\ast },D_{h}^{\ast }}(z)) \\ \mathbb{P}(D_{w}& =1;D_{h}=0|Z=z_{j})=\Phi (\nu _{D_{h}^{\ast }}(z_{j}))-\Phi _{2}(\nu _{D^{\ast }w}(z),\nu _{D_{h}^{\ast }}(z),\rho _{D_{w}^{\ast },D_{h}^{\ast }}(z)) \\ \mathbb{P}(D_{w}& =0;D_{h}=1|Z=z_{j})=\Phi (\nu _{D_{w}^{\ast }}(z_{j}))-\Phi _{2}(\nu _{D_{w}^{\ast }}(z_{j}),\nu _{h}(z_{j}),\rho _{D_{w}^{\ast },D_{h}^{\ast }}(z_{j})) \end{split} \end{equation*} for $j=0,1,2$. Even so we have the following three restrictions for the distribution of $Y_{w}$ and $Y_{h}$ for the selected sample: \begin{equation*} \begin{split} & \mathbb{P}(D_{w}=1;D_{h}=1,Y_{w}\leq y_{w},Y_{h}\leq y_{h}|Z=z_{j})= \\ & \int_{\nu _{D_{w}^{\ast }}(z_{j})}^{\infty }\int_{\nu _{D_{h}^{\ast }}(z_{j})}^{\infty }\int_{-\infty }^{\mu _{Y_{w}^{\ast }}(y_{w},x)}\int_{-\infty }^{\mu _{Y_{h}^{\ast }}(y_{h},x)}d\Phi _{4}(u_{1},u_{2},u_{3},u_{4},\Sigma (y_{w},y_{h},z_{j})) \\ & \mathbb{P}(D_{w}=1;D_{h}=1,Y_{w}\leq y_{w},Y_{h}>y_{h}|Z=z_{j})= \\ & \int_{\nu _{D_{w}^{\ast }}(z_{j})}^{\infty }\int_{\nu _{D_{h}^{\ast }}(z_{j})}^{\infty }\int_{-\infty }^{\mu _{Y_{w}^{\ast }}(y_{w},x)}\int_{\mu _{Y_{h}^{\ast }}(y_{h},x)}^{\infty }d\Phi _{4}(u_{1},u_{2},u_{3},u_{4},\Sigma (y_{w},y_{h},z_{j})) \\ & \mathbb{P}(D_{w}=1;D_{h}=1,Y_{w}>y_{w},Y_{h}\leq y_{h}|Z=z_{j})= \\ & \int_{\nu _{D_{w}^{\ast }}(z_{j})}^{\infty }\int_{\nu _{D_{h}^{\ast }}(z_{j})}^{\infty }\int_{\mu _{Y_{w}^{\ast }}(y_{w},x)}^{\infty }\int_{-\infty }^{\mu _{Y_{h}^{\ast }}(y_{h},x)}d\Phi _{4}(u_{1},u_{2},u_{3},u_{4},\Sigma (y_{w},y_{h},z_{j})) \end{split} \end{equation*} for $j=0,1,2$. This gives a system of 9 equations and 7 unknowns. \subsection{The model} We assume that \begin{equation*} \begin{split} F_{W_w^*, W_h^*, H_w^*, H_h^*|X, Z}(w_w, w_h, 0,0|x,z) = \\ \Phi_4(-\beta^T_w(w_w, w_h) x, -\beta^T_h(w_w, w_h) x, - \pi^T_w z, -\pi^T_h z, \Sigma_4(w_w, w_h)) \end{split} \end{equation*} where \begin{equation*} \begin{split} &\Sigma_4(w_w, w_h,x,z) = \\ &\left( \begin{array}{cccc} 1 & \rho_{10}(w_w, w_h, x, z) & \rho_{20}(w_w, w_h, x, z) & \rho_{30}(w_w, w_h, x, z) \\ \rho_{10}(w_w, w_h,x,z) & 1 & \rho_{21}(w_w, w_h,x,z) & \rho_{31}(w_w, w_h,x,z) \\ \rho_{20}(w_w, w_h,x,z) & \rho_{21}(w_w, w_h,x,z) & 1 & \rho_{32}(x,z) \\ \rho_{30}(w_w, w_h,x,z) & \rho_{31}(w_w, w_h,x,z) & \rho_{32}(x,z) & 1 \end{array} \right) \end{split} \end{equation*} and where $W_w^*$ and $W_h^*$ are the latent wages of the wife and the husband while $H^*_w$ and $H^*_h$ are their corresponding latent working hours. The vector $X$ represents the observed characteristics that determine both the latent wage as well as the latent number of working hours for the wife as well as the husband. The vector $Z$ represents the observed characteristics that only affect the number of working hours. Define \begin{equation*} D_{w} = \mathbf{1}(H^*_w > 0) \end{equation*} as a dummy variable in the case that the wife works and $D_h$ is defined likewise for the husband. Moreover, we define $W_w$ as the observed wage for the case that the wife works, i.e. $W_w = W^*_w$ in case that $D_w =1$ and otherwise it is undefined. These assumption imply that the joint wage distribution of the wife and the husband among families where both partners work equals \begin{equation} \begin{split} \mathbb{P}(W_w \leq w_w; W_h\leq y_w|D_w=1;D_h=1,X=x , Z = z) = \\ \frac{\mathbb{P}(W_w \leq w_w; W_h\leq w_h,D_w=1;D_h=1|Z = z)}{\mathbb{P} (D_w=1;D_h=1_h|X=x, Z = z)} = \\ \frac{\int^{-\beta_w(w_w) x}_{-\infty} \int^{-\beta_h(w_h)x}_{-\infty} \int_{-\gamma_w (x,z)}^{\infty} \int_{-\gamma_h (x,z)}^{\infty} \varphi_4(u_1, u_2, u_3, u_4, \Sigma_4(w_w, w_h, x,z)) du_4 du_3 du_2 du_1} { \int_{-\gamma^T_w (x,z)}^{\infty} \int_{-\gamma^T_h(x,z)}^{\infty} \varphi_2(u_1, u_2, \rho_{32}(x,z)) du_2 du_1} \end{split} \end{equation}

Measures of sorting in the presence of selection

We consider how the expressions for the contingency tables entries change when we account for endogenous sample selection. They are now written:

equation[equation omitted — 431 chars of source]

which corresponds to the ratio of the joint distribution of $(Y_{w},Y_{h})$ to the product of the marginals in the selected population. It is equal to 1 when the wages of the wives and husbands are independent from each other in the selected population.

The values of $s$ are identified from the wage data among couples in which both partners work and does not depend on the identification of the sample selection model. Nevertheless, it is interesting to understand how changes in the sorting measure can be attributed to changes in the model's parameters. This can be conducted via counterfactuals. Note that the numerator of this sorting measure can be written as:

equation[equation omitted — 208 chars of source]

where the integrand equals:

equation*[equation* omitted — 612 chars of source]

with

multline*[multline* omitted — 728 chars of source]

Counterfactuals based on a specific time period

For each time period counterfactuals can be obtained by replacing $ \boldsymbol{\Sigma }(d_{w},d_{h},y_{w},y_{h},x)$ with an alternative positive definite $\widetilde{\boldsymbol{\Sigma }} (d_{w},d_{h},y_{w},y_{h},x)$ reflecting different model parameters. For example, setting $\rho _{D_{w}^{\ast },D_{h}^{\ast }}(0,0,x)$ to zero examines the impact of making the employment decisions for husbands and wives conditionally independent. Setting $\rho _{D_{w}^{\ast },Y_{w}^{\ast }}(0,y_{w},x)$, $\rho _{D_{h}^{\ast },Y_{h}^{\ast }}(0,y_{h},x)$, $\rho _{D_{w}^{\ast },Y_{h}^{\ast }}(0,y_{h},x)$, and $\rho _{D_{h}^{\ast },Y_{w}^{\ast }}(0,y_{w},x)$ to zero eliminates the role of selection while setting $\rho _{Y_{w}^{\ast },Y_{h}^{\ast }}(y_{w},y_{h},x)$ to zero eliminates sorting on the unobservables correlated with wages. Sorting may still arise here on the basis of observed characteristics. Another potential counterfactual distribution integrates ((ref)) over the distribution of $Z$ for the whole population of married couples in our sample rather than the population of FTFY working couples.

Counterfactuals based on different periods of time

To investigate intertemporal changes in sorting we examine counterfactuals over time. We start by rewriting the elements of the contingency matrix in (ref) as:

multline*[multline* omitted — 660 chars of source]

where

equation*[equation* omitted — 128 chars of source]

and

equation*[equation* omitted — 112 chars of source]

We calculate counterfactual joint distributions assuming employment selection is as in year $q$, the wage structure is as in year $r,$ and the composition of the work force is as in year $s$ by:

multline[multline omitted — 186 chars of source]

where $F_{Z}^{s}$ is the distribution of $Z$ in year $s$, and

multline*[multline* omitted — 317 chars of source]

with

equation*[equation* omitted — 240 chars of source]

and\footnote{ A practical problem is that there is no guarantee that $\boldsymbol{\Sigma } ^{q,r}(0,0,y_{w},y_{h},x)$ is positive definite. Nevertheless, we did not encounter this problem in our empirical analysis.}

multline*[multline* omitted — 776 chars of source]

Comparable counterfactual marginal distributions can be computed for the denominator.

Comparisons should be based on wage levels of husbands and wives which incorporate the changes in the wage distributions. Since the real wages of males decreased at the bottom of the distribution while the real wages increased over the whole distribution for females, using fixed wage levels is not appropriate. For example, we are more likely to find a husband earning more than 20 dollars an hour with a wife earning more than 15 dollars an hour in 2020 than in 1976. This, however, does not imply a change in sorting behavior. Rather, it reflects the changes in the likelihoods from the respective marginal distributions. As we use quantiles rather than fixed wages to make these comparisons, we need the counterfactual quantiles of the corresponding marginal wage distributions. For example, the $\tau $-th quantile of the marginal wage distribution of the wives is:

multline[multline omitted — 201 chars of source]

Estimation

We consider a semiparametric BDR model with selection that imposes:

assumption[BDR with Selection] (1) $\mu_{Y^*_j}(y,x) = P_{j}(x)^{\prime }\beta _{j}(y),$ $\mu_{D^*_j}(0,z) = Q_{j}(z)^{\prime }\gamma_j,$ where $P_{j}$ and $Q_{j}$ are transformations of $x$ and $z$, and $\beta _{j}(y)$ and $\gamma _{j}$ are vectors of coefficients, $j \in \{w,h\}$; and (2) $\boldsymbol{ \Sigma} (d_{w},d_{h},y_{w},y_{h},x) = \boldsymbol{\Sigma} (d_{w},d_{h},y_{w},y_{h})$.

Estimation of the local model parameters

Using Assumption (ref) with $P_{j}(x)=x$ and $Q_{j}(z)=z$ for $j\in \{w,h\}$ we estimate the parameters in 2 steps:

enumerate• Bivariate probit to obtain $\gamma _{w}$, $\gamma _{h}$ and $\rho _{D_{w}^{\ast },D_{h}^{\ast }}:=\rho _{D_{w}^{\ast },D_{h}^{\ast }}(0,0)$ using: \begin{equation*} \mathbb{P}(D_{w}=1,D_{h}=1\mid Z=z)=\Phi _{2}(z^{\prime }\gamma _{w},z^{\prime }\gamma _{h};\rho _{D_{w}^{\ast },D_{h}^{\ast }}). \end{equation*} We denote the estimators as $\widehat{\gamma }_{w},$ $\widehat{\gamma }_{h}$ and $\widehat{\rho }_{D_{w}^{\ast },D_{h}^{\ast }}$. • Multivariate probit with sample selection correction to estimate the remaining parameters using: \begin{multline*} \mathbb{P}(Y_{w}\leq y_{w};Y_{h}\leq y_{w}|D_{w}=1;D_{h}=1,X=x,Z=z)\propto \\ \Phi _{4}(z^{\prime }\gamma _{w},z^{\prime }\gamma _{h},x^{\prime }\beta _{w}(y_{w}),x^{\prime }\beta _{h}(y_{h});\boldsymbol{\Sigma } (0,0,y_{w},y_{h})), \end{multline*} which can be estimated by a small adaption of the multivariate probit model after plugging in the first-stage estimators $\widehat{\gamma }_{w},$ $ \widehat{\gamma }_{h}$ and $\widehat{\rho }_{D_{w}^{\ast },D_{h}^{\ast }}$.

The first step is standard. The second step is straightforward, but the calculation of higher-order integrals in the multivariate probit is both computationally intensive and imprecise. The imprecision is especially unfortunate when combined with a numerical optimization method that assumes smoothness of the first and second order derivative of the criterion function. Therefore, we employ the GHK importance sampling simulator of Geweke (1991), Hajivassiliou and McFadden (1990) and Keane (1990) to simulate these probabilities. This importance sampling simulator uses the result that multivariate normal distributions, conditional on realizations of one or more elements of the outcome vector, are also normally distributed but with a lower dimension.

Estimation of the measures of sorting

Estimation of the sorting measures requires estimates of the quantiles of the marginal distributions of wives' and husbands' wage distributions. These are obtained via application of the generalized inverse or rearrangement operator to the plug-in estimator of the distribution. For example:

equation[equation omitted — 230 chars of source]

with:

multline*[multline* omitted — 424 chars of source]

for any $y_{h}$, where $\widehat{\boldsymbol{\Sigma }} ^{q,r}(0,0,y_{w},y_{h}) $ is the plug-in estimator of $\boldsymbol{\Sigma } ^{q,r}(0,0,y_{w},y_{h})$, and the integrals are calculated using the GHK importance sampling method.

Equation (ref) is solved using bisection implying that we need to calculate the term on the right-hand side of that equation for different trial values of the quantile. The counterfactual quantiles for husbands can be estimated in a similar way.

The estimator of the measures of sorting is:

multline*[multline* omitted — 822 chars of source]

where $\widehat{F}_{Y_{w},Y_{h}\mid D_{w},D_{h}}^{q,r,s}$, $\widehat{F} _{Y_{w}\mid D_{w},D_{h}}^{q,r,s}$ and $\widehat{F}_{Y_{h}\mid D_{w},D_{h}}^{q,r,s}$ are plug-in estimators of $F_{Y_{w},Y_{h}\mid D_{w},D_{h}}^{q,r,s}$, $F_{Y_{w}\mid D_{w},D_{h}}^{q,r,s}$ and $F_{Y_{h}\mid D_{w},D_{h}}^{q,r,s}$ in (ref), respectively; and $\underline{y} _{j}$ and $\overline{y}_{j}$ evaluated at $\widehat{Q} _{Y_{j}}^{q,r,s}((i-1)/10)$ and $\widehat{Q}_{Y_{k}}^{q,r,s}(i/10)$, $ i=1,\ldots ,10$ and $j\in \{w,h\}$.

Empirical Results

Parameter Estimates

We estimate the model at each decile for both of the males and females wage distribution for 5 year periods starting from 1976 and ending at 2022. As the total number of years is not divisible by 5 we use two periods of 6 years. These are 1991-1996 and 1997-2002. We estimate one hundred different sets of coefficients for each time period corresponding to the different combinations of deciles for each spouse. As our primary focus is on the role of selection and sorting, we do not report the 100 sets of parameter estimates but Figures (ref)-(ref) present the estimates of the different $\rho ^{\prime }$s at the same quartile of both distributions along with their bootstrapped standard errors. Results for the other quantiles are available from the authors. Note that we employ the 100 sets of parameter estimates in conducting the counterfactuals reported below and the discussion of the $\rho ^{\prime }$s which follows is only to provide some insight into the behavior of these selection and sorting parameters.

figure[figure omitted — 3,684 chars of source]

We discuss our model specification before proceeding to the results. Footnote (ref) listed the variables in $X$ employed in the BDR specification. We continue to use these variable but identification now requires some variables in $Z$ which are excluded from $X.$ We follow Mulligan and Rubinstein (2008) and employ a dummy variable for children at home, family size, and a dummy variable for children at home under 5 years of age in this role. These variables are seen as contentious and we discuss the implications of these choices below.

Consider the correlation between the unobservables in the spouses' employment equations, $\rho _{D_{w}^{\ast },D_{h}^{\ast }},$ recalling it is invariant to the location in the wage distributions. The positive value indicates that the unobservables affecting the spouses' work decisions are positively correlated. The coefficient is small in magnitude but is statistically significantly different from zero. It generally increases over time, except for a dip in the middle of the sample period, and reveals that the unobserved factors which make spouses both work FTFY are increasingly correlated over time. We leave a discussion of the marginal impact of changing this, and other parameters, to the following section.

The estimated selection coefficients for wives ($\rho _{D_{w}^{\ast },Y_{w}^{\ast }}$) and husbands ($\rho _{D_{h}^{\ast },Y_{h}^{\ast }}$) capture the correlation between the unobservables affecting the individual's respective work decision and their own wages. While in conventional selection models they capture the relationship between the selection decision and the mean wage, we evaluate this correlation at different quantiles of the wage distribution. The estimate for wives is particularly interesting and consistent with earlier evidence using similar identification approaches (see, for example, Mulligan and Rubinstein, 2008, and FVVP, 2023b) for FTFY workers. Similar to these earlier papers we find evidence of negative selection changing to positive selection. At each quantile there is negative selection in the earlier period with an estimate around -0.3 which is statistically different from zero. For the latter sample period the estimate has increased. At Q1/Q1 the estimate has increased to zero, while at Q2/Q2 and Q3/Q3 the estimates are approximately 0.2 and 0.3. As discussed in FVVP (2023b), the sign of $\rho _{D_{w}^{\ast },Y_{w}^{\ast }}$ is contentious given its interpretation and its implication for selection. They note that the sign of the selection terms appears to reflect the impact of the variables used to explain participation which are excluded from the wage equation. The change of sign appears to capture the changing impact of these variables on the participation decision. \footnote{ FVVP (2023b) show that this negative coefficient for the earlier period is likely to be due to the exclusion restrictions. The reader is referred to that paper for a detailed discussion of how the exclusion restrictions may be generating the selection related results. One could reproduce the counterfactual that follows using the FVVP (2023b) approach which employs an alternative identification strategy but that is not feasible without making a number of additional assumptions.} While the sign of $\rho _{D_{w}^{\ast },Y_{w}^{\ast }}$ is controversial we explore the impact of changing its value in counterfactual exercises below.

It is not typical to account for selection when estimating male wage equations. However, it is important to do so here as we do not know apriori the impact of selection in this model. Moreover, as we account for the role of male selection on the female wage it is necessary to model the male work decision. The parameter $\rho _{D_{h}^{\ast },Y_{h}^{\ast }}$ is very poorly estimated at each of the quantiles reported. At Q1/Q1 and Q2/Q2 it is negative and large but very imprecisely estimated. At Q3/Q3 it is both negative and positive but generally imprecisely estimated. We attribute this result to our inability to identify this parameter from these data.

The parameters $\rho _{D_{w}^{\ast },Y_{h}^{\ast }}$ and $\rho _{D_{h}^{\ast },Y_{w}^{\ast }}$ also vary by location in the respective wage distributions and capture the correlation between the unobservables affecting the work decision of the individual with the unobservables affecting the wage of their spouse at a certain quantile in the spouses wage distribution. First consider $\rho _{D_{w}^{\ast },Y_{h}^{\ast }}$. Figures (ref)-(ref) suggest that at each quartile the estimate is both negative and positive at each quantile depending on the time period. However, the confidence bands suggest that we cannot reject that the estimate is zero for many of the periods with the exception of some middle years at the first quartile. $\rho _{D_{h}^{\ast },Y_{w}^{\ast }}$ captures how the unobservables driving the husband's participation decision affects the wife's wage. Given the evidence above on male selection it is unsurprising that this parameter is imprecisely estimated although it appears to be generally negative and statistically different from zero.

figure[figure omitted — 14,203 chars of source]
figure[figure omitted — 14,235 chars of source]
figure[figure omitted — 14,182 chars of source]

$\rho _{Y_{w}^{\ast },Y_{h}^{\ast }}$ is an economically important parameter as it captures sorting via the correlation between the unobservables generating the husband's and wife's wages. If individuals are sorting with respect to these unobservables we expect this parameter to be positive. Moreover, as this parameter also captures the prices of these unobservables it will capture similar structural effects associated with the implicit prices of observed characteristics. As noted above, this parameter will capture factors related to influences, such as cost of living or wage premia, which are not captured by the conditioning variables but which are shared by spouses. For example, while the data measure various aspects of the location in which the spouses live we are unable to distinguish between those living in costly urban areas. It will also capture other factors specific to unobserved issues shared by the spouses. For example, it captures that both may work for the same firm and or went to the same college and share the costs and/or benefits associated with their wages. The estimate of this parameter is positive and generally ranges between 0.2 and 0.3. Figure (ref) plots the estimated densities for this parameter for the starting and ending time periods for all 100 models while Figures (ref)-(ref) present the estimates at the various quantiles we consider. There is some evidence that it is increasing at these different quantiles over time but the confidence intervals do not appear to reject that the impact is constant.

With respect to sorting the only clear and consistent evidence is associated with $\rho _{Y_{w}^{\ast },Y_{h}^{\ast }}.$ The evidence from the counterfactuals based on the BDR estimates revealed that setting $\rho _{Y_{w}^{\ast },Y_{h}^{\ast }}$ to zero drastically reduced the presence of positive sorting. However, this may now be offset by allowing for selection even though the estimates of $\rho _{D_{w}^{\ast },D_{h}^{\ast }},$ $\rho _{D_{w}^{\ast },Y_{w}^{\ast }}$, $\rho _{D_{h}^{\ast },Y_{h}^{\ast }}$, $ \rho _{D_{w}^{\ast },Y_{h}^{\ast }}$ and $\rho _{D_{h}^{\ast },Y_{w}^{\ast }} $ do not individually or collectively present a clear story regarding sorting.

The estimates of the parameters related to the husband's selection decision warrant further discussion. As male participation is typically treated as exogenous there is little empirical work on the impact of male selection on wages. Moreover, as methods to analyze the role of selection at different quantiles have only recently been developed there is little existing evidence on how it varies across the wage distribution. However, the parameters are not precisely estimated in this setting and in some instances may be unreasonable in magnitude. This is likely to reflect the form of identification employed. Alternatively, husbands' work decisions may be exogenous. Given the restricted version of the model is a special case of the richer model we decided to proceed with the results with this full model but we also estimated the model treating the husbands' work decisions as exogenous and conducted the corresponding counterfactuals. We comment on those results below although to anticipate the main finding, the treatment of husbands' work decisions as exogenous did not alter the substantive results. We report the estimates of $\rho _{D_{w}^{\ast },Y_{w}^{\ast }}$, $ \rho _{D_{w}^{\ast },Y_{h}^{\ast }}$ and $\rho _{D_{h}^{\ast },Y_{w}^{\ast }} $ from this restricted model in Figures (ref)- (ref).

\pgfplotsset{/pgfplots/group/.cd, horizontal sep=2cm, vertical sep=1.5cm }

figure[figure omitted — 7,139 chars of source]
figure[figure omitted — 7,290 chars of source]
figure[figure omitted — 7,120 chars of source]

Counterfactual Sorting Patterns

The estimated model's capacity to reproduce the frequencies in Tables (ref) and (ref) is reflected in Tables (ref) and (ref) and, although the large number of cells makes it difficult to draw a conclusion by a visual inspection, the estimated cells appear close to the true values. We now examine the role of various model parameters in generating the observed data by changing selected model parameters and examining the predicted allocation across cells. We compare these counterfactuals to Tables (ref) and (ref) as they represent the predictions of our estimated model.

table[table omitted — 1,502 chars of source]
table[table omitted — 1,628 chars of source]
table[table omitted — 1,679 chars of source]
table[table omitted — 1,537 chars of source]

We begin by setting $\rho _{D_{w}^{\ast },D_{h}^{\ast }}$ to zero. Table (ref) presents the results \ for 1976-1980 and Table (ref) those for 2018-2022. They appear similar to Table (ref) and Table (ref) respectively and we conclude that this parameter is unimportant for determining the joint wage distribution.

table[table omitted — 1,502 chars of source]
table[table omitted — 1,636 chars of source]
table[table omitted — 1,680 chars of source]
table[table omitted — 1,543 chars of source]

Tables (ref) and (ref) present the results when $\rho _{D_{w}^{\ast },Y_{w}^{\ast }},$ $\rho _{D_{h}^{\ast },Y_{h}^{\ast }}$, $\rho _{D_{w}^{\ast },Y_{h}^{\ast }}$ and $\rho _{D_{h}^{\ast },Y_{w}^{\ast }}$ are also set to zero. We characterize these parameters as collectively capturing the selection process. For the earlier period the d10/d10 cell increases from 2.29 to 2.94 while the sum in the (d9+d10)/(d9+d10) cells increases from 6.7 to 7.4. This suggests that negative selection bias operating in both labor markets is reducing household inequality. Negative selection is removing the relatively higher paid males and females from FTFY employment. Setting these parameters to zero inserts more higher paid workers into the market. This results in a higher propensity of the higher paid males and females to marry. This is an interesting result although it may reflect the identification strategy. Interestingly the frequencies in Table (ref) provide a somewhat comparable story although the changes are less drastic. Setting the other parameters to zero appears to have little affect on the lower cells. However, while the d10/d10 cell decreases marginally the sum in the (d9+d10)/(d9+d10) cells increases from 7.9 to 8.4. This is similar to the earlier period although the estimated positive selection effects for females at upper quantiles is offsetting the negative selection effects for males throughout the period.

We also examine the counterfactual in which the above parameters are set to zero and we use the $Z^{\prime }$s of the whole sample rather than only those who are working. The results are in Tables (ref) and (ref) and a comparison with Tables (ref) and (ref) indicates they do not affect the observed sorting patterns. Although previous studies (see, for example, Chernozhukov, Fern\'{a}ndez-Val, and Luo 2019) have analyzed the impact of selection on the distribution of wages and have found it to have some impact, our results indicate it would not change the sorting patterns even if it is changing the distribution of both, or either of, the male and female wages.

Finally we set $\rho _{Y_{w}^{\ast },Y_{h}^{\ast }}$ to zero. For the BDR model the corresponding counterfactual produced a substantial reduction in sorting. Table (ref) reports the results for the earlier time period. Given the large number of cells it is useful to focus on the extreme cases as these are the outcomes more closely associated with inequality. First consider the d1/d1 and d10/d10 cells. The former decreases from 2.21 to 1.35 while the latter decreases even more dramatically from 2.89 to 1.36. Expanding the 4 lower and upper combinations to include d2 and d9 respectively produces reductions from 6.6 and 7.4 to 5.1 and 4.96 respectively. It is interesting that the more substantial reductions appear to occur at the cells for the higher wage deciles. Turning to the 2018-2022 time period we see a similar pattern. The d1/d1 cell decreases from 2.88 to 1.54 and d10/d10 decreased from 3.07 to 1.43. Expanding to the four lowest and highest we see reductions of 6.82 to 5.21 and 8.3 to 5.07 respectively. This clearly suggests that the correlation in the unexplained components of wages is largely driving the observed sorting patterns. That is, sorting on unobservables is an important factor in driving inequality.

table[table omitted — 414 chars of source]

To capture the sorting behavior which includes the off diagonals we compute the Kendall rank correlation coefficient for these counterfactuals. These are reported in Table (ref) and confirm the patterns described above. For the 1976-1980 period neither the selection parameters nor the use of the working or total population composition of $Z^{\prime }$s affect the rank correlation coefficient. For each of these experiments the rank correlation coefficient does not differ greatly from its value for the data of .18. However, when we additionally set the value of $\rho _{Y_{w}^{\ast },Y_{h}^{\ast }}$ to zero the value falls to .07. For the 2018-2022 period the results are similar despite the rank correlation coefficient of .27 being 50% higher than the 1976-1980 value. For this period the only counterfactual which generates a different pattern of sorting is that corresponding to $\rho _{Y_{w}^{\ast },Y_{h}^{\ast }}=0.$ For this latter period the reduced rank correlation coefficient is also .07. This is consistent with Tables (ref) and (ref) .

figure[figure omitted — 6,312 chars of source]

We noted that we estimated the model treating the male employment decision as exogenous. Although we do not report the results here, we conducted the corresponding counterfactuals to those for the full model. The general flavor of the simulations were similar although there was stronger evidence of selection affecting the observed sorting for the earlier period. However, the result that setting $\rho _{Y_{w}^{\ast },Y_{h}^{\ast }}$ to zero reduced the observed level of positive sorting was also clearly supported.

Decomposing Changes in the Sorting Patterns

We now focus on the changes in the sorting patterns in the earlier tables and examine how the probability of both spouses being in a specific quantile of their wage distribution changes over our sample period. We decompose the total change into structural, selection and composition effects. Note that unlike the previous section, we now also examine the impact of the changes in the $\beta ^{\prime }$s over time. As we employ a multivariate normal approximation for each point of the four dimensional partition of the data we can use the model parameter estimates to decompose the observed changes into the various components. While the composition effects are defined as the changes in the conditioning variables, the allocation of the other parameters are less straightforward. The changes in the $\beta ^{\prime }$s can be assigned the interpretation of conventional structural effects. We assign $\rho _{D_{w}^{\ast },Y_{w}^{\ast }}$, $\rho _{D_{h}^{\ast },Y_{h}^{\ast }},$ $\rho _{D_{w}^{\ast },Y_{h}^{\ast }}$ and $\rho _{D_{h}^{\ast },Y_{w}^{\ast }}$ to the selection effects. We isolate the impact of $\rho _{Y_{w}^{\ast },Y_{h}^{\ast }}$ parameter as it captures the correlation between the unobservables which impact the spouses' wages. While this is a type of structural effect we separate it from the structural effects operating through the $\beta ^{\prime }$s. Although we can evaluate the sorting behavior observed in 100 cells it appears that the interesting behavior occurs in the top and bottom cells. Accordingly we conduct the decompositions for both members of the couple being located in the bottom and top deciles and quartiles. These plots are presented in Figure (ref). The decompositions represent how the change in the measures can be explained by changes in the various components. If a specific component appears to be close to zero then this indicates that its contribution has remained unchanged and not that its contribution is zero.

We start with the probability of both spouses being located in the bottom decile and quartile. Over the 47 years of our sample the former increases by almost 60 percent and the latter by almost 40 percent. For d1 the large positive change is driven by changes in the structural effect and $\rho _{Y_{w}^{\ast },Y_{h}^{\ast }}.$ Changes in the selection effect are unimportant and the changes in the composition effect are negative. For Q1 the results are essentially the same although the changes in the contribution of $\rho _{Y_{w}^{\ast },Y_{h}^{\ast }}$ are generally unimportant except for a decrease towards the end of the sample. At the top of the distribution there is a similar finding. The probability of both in the top decile and quartile increases by around 60% and 35% respectively. The change in the probability of both spouses being in d10 is again driven by changes in the structural effect and $\rho _{Y_{w}^{\ast },Y_{h}^{\ast }}$ although the change in the structural effect diminishes towards the end of the sample. The change in the probability of both spouses being in Q3 is due to changes in the structural effects. Changes in the composition and selection effects appear unimportant for either d10 or Q3. The results on the composition effects are consistent with the evidence as noted above in Breen and Salazar (2011), Gihleb and Lang (2018), Eika, Mogstad, and Zafar (2019) and Chiappori, Costa-Dias and Meghir (2020,2023). Education is a component of the composition effect and there is no impact at either Q1 or Q3. At D1 and D9 it is less clear as both probabilities decline.

The increasing presence of married females into the FTFY in the 1970's and 1980's had little impact on the sorting process in the marriage market. Overall the selection effects are small. Second, composition effects are generally unimportant and have reduced positive sorting. This reflects that the increased level of acquired education has increased everyone's probability of marrying a relatively highly educated spouse. This is consistent with the findings of others noted above regarding the impact of composition effects on positive sorting as measured by education. The clearest evidence is associated with the structural effects. These are the most important contributors to the total effect\ and their interpretation is also quite clear. Sorting behavior has not greatly changed in terms of the observed characteristics of the spouses. However, the higher skill premia, and the heterogeneity of the skill premia, have resulted in the highly (lowly) paid being even more highly (lowly) paid. This is consistent with empirical evidence on decompositions of changes in wage inequality. The evidence on the role of the $\rho _{Y_{w}^{\ast },Y_{h}^{\ast }}$ provides a similar story. While this parameter captures the impact of unobservables it also reflects how they are priced. Our evidence suggests that the increased observed level of positive assortative sorting reflects the increases in prices of observables and unobservables. Individuals appear to marry the \textquotedblleft same" spouses but the value of their characteristics has changed due to the increased skill premia. It is also possible that $\rho _{Y_{w}^{\ast },Y_{h}^{\ast }}$ partially captures how an individual's wage is influenced by the value of their spouse's characteristics. For example, individuals with highly successful spouses may benefit from exposure to their spouse's network and this may produce some relationship between their wage and, for example, their spouse's education level. Some preliminary investigations of the data suggest that this relationship exists, and that it changes over time, but we do not pursue it here.

Household Income Inequality

We now explore the implication of the model's estimates for household income equality. We assume that all individuals comprising the FTFY couples are working the same number of hours and we employ the household wage, defined as the sum of the spouses' hourly wages, as our measure of family income. We examine how inequality has evolved by examining the D8/D2 ratio under different counterfactuals. We focus on this ratio as the manner in which its components have changed is clear from the tables above. The results are shown in Table (ref). Note that we do not isolate the role of structural effects operating through the $\beta ^{\prime }$s nor the composition effects reflecting the changes in the $X^{\prime }$s over time. As listed at the start of our Introduction, there is a large existing literature which has clearly established the impact of these factors on wage inequality. We repeat the same exercise as in the counterfactual sorting exercises of Section (ref) in which we incrementally set the parameters capturing different features of the model to zero. We also explore the impact of employing the population $X^{\prime }$ s rather than those of the working sample.

table[table omitted — 453 chars of source]

The first row of Table (ref) presents the D8/D2 ratio for our two extreme sample periods. There is a large increase of 29.6%, from 1.79 to 2.32, noting that it is similar to the growth in comparable measures of inequality presented in the discussion of the data. The second row presents the comparable estimates from the predictions of our model and indicates that the level of inequality observed in the data are maintained. To examine the role of sorting we first estimate the corresponding measure in a setting of unconditional random sorting in which husbands are assigned to wives in a random manner. This is shown in the table's bottom row. The results represent the D8/D2 ratio based on 100,000 draws. The estimates are 1.68 and 2.05 and reflect an increase of 22% over the sample period. This large impact reflects the large increase in wage inequality which occurs for both husbands and wives. More interestingly, comparing the top and bottom rows of this table indicates that sorting increases inequality by 7.0% in the earlier period and 13.1% in the latter. This suggests that sorting has substantially increased household inequality.

To investigate the role of the various model parameters, rows 3 to 6 repeat the counterfactuals conducted above. As with the tables above, each of the rows reflects the incremental change in the inequality measure as the parameters associated with some feature of the model are set to zero. For the earlier period selection has reduced inequality although we have highlighted our concerns regarding this result as it is reliant on the form of identification. The results regarding the use of the population $ Z^{\prime }$s rather than the working sample $Z^{\prime }$s confirms our earlier results that this is does not appear to affect the results regarding sorting. Row 6 also highlights the recurring result that positive sorting is operating through the correlation of the spouses' unobservables. For the 2018-2022 period we have similar findings although that related to selection does not hold. However the importance of sorting on unobservables for inequality is again highlighted. Setting $\rho _{Y_{w}^{\ast },Y_{h}^{\ast }}=0$ decreases the D8/D2 ratio for the earlier period by 16% from 2.06 to 1.73 and for the later period by 7% from 2.26 to 2.12.

While there is an important role for $\rho _{Y_{w}^{\ast },Y_{h}^{\ast }}=0,$ the remaining large increase in inequality in the household wage appears due to changes in composition and structural effects. However, the existing evidence, here and in the existing literature, suggests the sorting on observables has not changed for this time period and composition effects are unimportant. This suggests that the increase in wage inequality across households is due to structural effects capturing changes in the skill premia.

Conclusion

We examine the role of marital sorting on household income inequality in a period of substantial and increasing wage inequality. We do so by examining a sample of married couples from the CPS for the years 1976 to 2022 in which each of the spouses works full time/full year. To investigate the determinants of marital sorting and its impact on inequality we estimate a model explaining the spouse's employment decisions and their location in the gender specific wage distributions. This is done via a multivariate distribution regression approach which incorporates selection. We provide a number of important empirical findings.

First, we confirm that wage inequality has increased for both spouses in dual FTFY couples. Changes in inequality for these groups do not precisely correspond to those for all males and females but the trends are similar.

Second, we find clear evidence of positive sorting defined as higher incidences of high (low) wage males marrying high (low) wage females than expected under random sorting. However, we acknowledge that random sorting is not a realistic alternative to positive sorting as couples are frequently similar in terms of age, race and location of residence and each of these are known to determine wages. However, the disproportionate fraction of males in the extremes of the wage distribution marrying females in the corresponding extremes of their distribution is supportive of positive sorting. Moreover, this positive sorting has increased over our time period. Combining this increasing sorting with increasing spouse specific inequality has increased inequality across couples.

Third, a series of counterfactual experiments do not provide any evidence that unobservables related to either of the spouses' work decisions have any implications for the observed sorting behavior. However, there is compelling evidence that the correlation between the unobservables influencing the spouses wages is substantially contributing to the observed sorting behavior. While some of this correlation may be due to factors, such as cost of living premia, which are incurred by both members of the couple it is also possible that this captures factors corresponding to unobserved ability.

Finally, we find that positive sorting associated with the increased share of extreme cells has increased over the 47 years of our sample. A decomposition of the changes in these cells shares over the period examined reveals that their growth is almost entirely due to the prices of observable and unobservable characteristics. This suggests that the increase in observed sorting pattern in not due to different behavior nor the increase of more females in the labor market. Rather as the skill premia has increased this has increased the wages of both members of some couples while couples without these skills appear to be both negatively affected.

\thinspace\ References

description• Autor, D.H., L.F. Katz , and M. Kearney (2008), \textquotedblleft Trends in U.S.\ wage inequality: Revising the revisionists\textquotedblright, Review of Economics and Statistics 90 300--23. • Autor, D.H., A. Manning, and C.L. Smith (2016), \textquotedblleft The contribution of the minimum wage to US wage inequality over three decades: A reassessment.\textquotedblright\ American Economic Journal: Applied Economics 8, 58--99. • \textsc{Blau, F. and L.M. Kahn} (2009), “Inequality and earnings distribution”, In: W. Salverda, B. Nolan, and T.M. Smeeding, “Oxford Handbook on Economic Inequality”, Oxford University Press, 177--203. • \textsc{Blau, F.D., L.M. Kahn, N. Boboshko, and M.L. Comey} (2021), \textquotedblleft The impact of selection into the labor force on the gender wage gap\textquotedblright , working paper, Cornell University. • \textsc{Breen, R. and L. Salazar} (2011), \textquotedblleft Educational assortative mating and earnings inequality in the United States\textquotedblright, \emph{American Journal of Sociology} \textbf{117}, 808-43. • \textsc{Cancian, M. and D. Reed} (1998a), \textquotedblleft Assessing the effects of wives' earnings on family income inequality\textquotedblright , \emph{The Review of Economics and Statistics} \textbf{80}, 73--79. • \textsc{Cancian, M. and D. Reed } (1998b), \textquotedblleft The impact of wives' earnings on income inequality: issues and estimates\textquotedblright , \emph{Demography} \textbf{36}, 173--84. • \textsc{Chiappori, P.A., B. Salani\'e and Y.Weiss}, (2017), “Partner choice, investment in children, and the marital college premium", \emph{ American Economic Review} \textbf{107}, 2109--67. • \textsc{Chiappori, P.A., Costa-Dias, M. and C.Meghir}, (2019), “The marriage market, labor supply, and education choice", \emph{Journal of Political Economy} \textbf{126}, s26-s72. • \textsc{Chiappori, P.A., Costa-Dias, M., Crossman, S. and C.Meghir} (2020), “Changes in assortative matching and inequality in income: Evidence for the UK", \emph{Fiscal Studies} \textbf{41}, 39--63. • \textsc{Chiappori, P.A., Costa-Dias, M. and C.Meghir}, (2020). “Changes in assortative matching: Theory and evidence for the US", NBER Working Paper 26932. • \textsc{Chiappori, P.A., Costa-Dias, M. and C.Meghir,} (2023), “The measuring of assortativeness in marriage", working paper, Yale University. • \textsc{Chernozhukov, V., I. Fern\'{a}ndez-Val, and S. Luo}{ \ } (2019), \textquotedblleft Distribution regression with sample selection, with an application to wage decompositions in the UK", working paper, MIT, Cambridge (MA). • \textsc{De Rock, B., M. Kolvaleva, and T. Potoms} (2023), “A spouse and a house are all we need? Housing demand, labor supply and divorce over the lifecycle", working paper, Free University of Brussels. • \textsc{Eika, L., M. Mogstad, and B. Zafar} (2019), \textquotedblleft Educational assortative mating and household income Inequality\textquotedblright , \emph{Journal of Political Economy} \textbf{ 127}, 2795--2835. • \textsc{Fern\'{a}ndez-Val I., A. van Vuuren, and F. Vella} (2018), \textquotedblleft Decomposing Real Wage Changes in the United States\textquotedblright , \emph{IZA Discussion paper} \textbf{12044}, Bonn. • \textsc{Fern\'{a}ndez-Val I., A. van Vuuren, F. Vella and F. Peracchi} (2023a), \textquotedblleft Hours worked and the U.S distribution of real annual earnings 1976--2019\textquotedblright , \emph{Journal of Applied Econometrics}, forthcoming. • \textsc{Fern\'{a}ndez-Val I., A. van Vuuren, F. Vella and F. Peracchi} (2023b) \textquotedblleft \textquotedblleft Selection and the distribution of female real hourly wages in the US\textquotedblright \textquotedblright , \emph{Quantitative Economics} \textbf{14}, 571-607. • \textsc{Fern\'{a}ndez-Val I., J. Meier, A. van Vuuren and F. Vella} (2023) \textquotedblleft A bivariate distribution regression model for intergenerational mobility\textquotedblright , working paper, Boston University. • \textsc{Fern\'{a}ndez, R. and R. Rogerson} (2001), \textquotedblleft Sorting and long-run inequality\textquotedblright , \emph{Quarterly Journal of Economics} \textbf{116}, 1305--41. • \textsc{Flood, S., M. King, S. Ruggles, and J.R. Warren} (2015), \textquotedblleft Integrated public use microdata series, Current Population Survey: Version 4.0 [Machine-readable database]\textquotedblright, working paper, University of Minnesota. • \textsc{Geweke, J.} (1991), \textquotedblleft Efficient simulation from the multivariate normal and student-t distributions subject to linear constraints\textquotedblright , \emph{Computing Science and Statistics} \textbf{23}, 571--78. • \textsc{Gihleb, R. and K. Lang} (2020), “Educational homogamy and assortative mating have not increased”, \emph{Research in Labor Economics} \textbf{48}, 1--26. • \textsc{Greenwood, J., N. Guner, G. Kocharkov and C. Santos} (2014), \textquotedblleft Marry your like: assortative mating and income inequality\textquotedblright , \emph{American Economic Review} \textbf{104}, 348--54. • \textsc{Hajivassiliou, V.} (1990) \textquotedblleft Smooth simulation estimation of panel data LDV models\textquotedblright , working paper, Yale University. • \textsc{Heckman J.J.} (1974), \textquotedblleft Shadow prices, market wages and labor supply\textquotedblright ,\ \emph{Econometrica} \textbf{42}, 679--94. • \textsc{Heckman J.J.} (1979), \textquotedblleft Sample selection bias as a specification error\textquotedblright , \emph{Econometrica} \textbf{47} , 153--61. • \textsc{Juhn C., K.M. Murphy, and B. Pierce} (1993), \textquotedblleft Wage inequality and the rise in returns to skill\textquotedblright , \emph{ Journal of Political Economy} \textbf{101}, 410--42. • \textsc{Katz L. F., and K. Murphy} (1992), \textquotedblleft Changes in relative wages, 1963--1987: Supply and demand factors\textquotedblright\ \emph{Quarterly Journal of Economics} \textbf{107}, 35--78. • \textsc{Keane, M.} (1994), \textquotedblleft A computationally practical simulation estimator for panel data\textquotedblright , \emph{ Econometrica} \textbf{62}, 95-116. • \textsc{Kremer, M.} (1997), \textquotedblleft How much does sorting increase inequality?\textquotedblright , \emph{Quarterly Journal of Economics } \textbf{112}, 115--39. • \textsc{Lise, J. and S. Seitz} (2011), “Consumption Inequality and Intra-household Allocations", \emph{Review of Economic Studies }\textbf{78}, 328--55. • \textsc{Lise, J. and K. Yamada (2018)},“Household sharing and commitment: Evidence from panel data on individual expenditures and time use", \emph{Review of Economic Studies }\textbf{86}, 2184-2219. • \textsc{Mulligan, C. and Y. Rubinstein} (2008), \textquotedblleft Selection, investment, and women's relative wages over time\textquotedblright, \emph{Quarterly Journal of Economics} \textbf{123}, 1061--1110. • \textsc{Murphy, K.M. and F. Welch} (1992), \textquotedblleft The structure of wages\textquotedblright , \emph{Quarterly Journal of Economics} \textbf{107}, 285--326. • \textsc{Sibuya, M.} (1959), \textquotedblleft Bivariate extreme statistics, I,\textquotedblright \emph{Annals of the Institute of Statistical Mathematics} \textbf{11}, 195--210. • \textsc{Welch F.} (2000), \textquotedblleft Growth in women's relative wages and inequality among men: One phenomenon or two?\textquotedblright\ \emph{American Economic Review Papers & Proceedings} \textbf{90}, 444--49.