EconBase
← Back to paper

Partially Identified Heterogeneous Treatment Effect with Selection: An Application to Gender Gaps

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

102,541 characters · 18 sections · 139 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Partially Identified Heterogeneous Treatment Effect with Selection: An Application to Gender Gaps

\affil[1]{Monash University }

abstractThis paper addresses the sample selection model within the context of the gender gap problem, where even random treatment assignment is affected by selection bias. By offering a robust alternative free from distributional or specification assumptions, we bound the treatment effect under the sample selection model with an exclusion restriction, an assumption whose validity is tested in the literature. This exclusion restriction allows for further segmentation of the population into distinct types based on observed and unobserved characteristics. For each type, we derive the proportions and bound the gender gap accordingly. Notably, trends in type proportions and gender gap bounds reveal an increasing proportion of always-working individuals over time, alongside variations in bounds, including a general decline across time and consistently higher bounds for those in high-potential wage groups. Further analysis, considering additional assumptions, highlights persistent gender gaps for some types, while other types exhibit differing or inconclusive trends. This underscores the necessity of separating individuals by type to understand the heterogeneous nature of the gender gap. Keywords: Sample selection model; Partial identification; Heterogeneous treatment effect; Gender gaps. JEL Classification: C21; J16.

Introduction

In this study, we contribute to the understanding of gender gaps in labor markets through two key avenues. Firstly, we introduce a partial identification method that bounds gender gaps without excessive specification assumptions, maintaining certain literature-based assumptions while relaxing others. Secondly, we move beyond conventional wage and employment status comparisons and delve into subpopulations defined by various traits, including years, covariate groups, and importantly, types of individuals. This approach allows us to uncover heterogeneous treatment effects based on both observable and unobservable characteristics.

A key focus of our method lies in obtaining gender gaps for different types based on these characteristics. For example, we consider types of individuals who may opt to work under certain conditions but not under others, reflecting diverse motivations and social norms. Understanding the gender gap across such types is important as it sheds light on the complexities of individual choices and societal expectations. Even among nonworkers, identifying the gender gap for different types remains important, offering insights into disparities shaped by factors such as family dynamics and personal motivations. By investigating wage gap bounds for different types of individuals under these distinct employment scenarios, influenced by both observable attributes like age and education, and underlying societal norms or motivations, our study provides valuable insights into gender disparities across diverse subpopulations. This comprehensive approach enhances our understanding of treatment effects within varied social contexts and individual circumstances.

Labor force participation and wages serve as vital metrics for assessing the performance of subgroups or types within a population. Labor force participation, representing whether individuals choose to work or not, holds significance as wages are observable only for those who are employed. Consequently, the absence of wage data for non-working individuals is not random but rather stems from their deliberate decisions, highlighting the importance of addressing selection issues in gender gap analyses.

Various approaches exist to tackle the selection problem in gender gap research and related fields. Some methods involve making distributional assumptions and specifying models to estimate parameters in wage equations, followed by a comparison of these parameters between men and women. Others focus on identifying average treatment effects using information from the conditional distribution of selected individuals. Another approach involves considering both working and non-working individuals and employing assumptions to bound the wage distributions for men and women separately. This approach utilizes partial identification techniques, wherein the wage distributions are bounded first, followed by the calculation of differences between these distributions to quantify the gender gap.

We propose another partial identification method within the potential outcome framework to directly assess the effect of gender on wages, bypassing the need to bound wage distributions separately for men and women, whether employed or not. Instead, we bound the Average Treatment Effect (ATE) of gender based on different types, offering a distinct approach from traditional methods reliant on wage equation estimation. Our method does not hinge on specific model specifications or distributional assumptions, unlike methods estimating wage equations, which often require such assumptions. Consequently, rather than estimating the gender gap parameter, our method bounds it, providing a robust check or supplement to existing results. Notably, works by Heckman, such as Heckman1974 and Heckman1979, provide point identification for treatment effects in models with selection. However, their approaches rely on specific models and distribution assumptions. For instance, Heckman1974 and Heckman1979 utilize the relationship between the conditional and unconditional distributions of unobserved error terms to point identify parameters, relying on conditions such as the comparison between wages and reservation wages or the inclusion of selection rules in model equations. Despite their contributions, these methods require specific models and distribution assumptions, whereas our approach offers a more flexible and robust alternative. In gender gap studies, BlauBeller1988 and Mulligan2008 employ the model from Heckman1979 to estimate the parameters in the wage equations.

While our method shares similarities with approaches identifying average treatment effects through conditional distribution information from selected samples, we distinguish ourselves by incorporating conditional distributions on observable covariates and the selection, which integrates the underling unobservable covariates. Unlike methods seeking point identification, our estimation method does not involve estimating distributions, conditional expectations, or copula parameters. Utilizing data from the Current Population Survey (CPS) in the United States spanning 1976 to 2013, sourced from Maasoumi2019, our findings serve as a complementary analysis to theirs. In contrast, Fernandez2021 (following Blundell2007) analyze data from the Family Expenditure Survey (FES) in the United Kingdom, focusing on the effect of education on log hourly wage with a censored outcome variable and a non-binary selection variable, namely, working hours. Their approach centres on identifying functions such as conditional expectations and distributions of observable variables, employing semiparametric methods like logistic and quantile regression. They observe a relatively stable effect of education but note an increasing impact of experience on wages in the 1990s.

Expanding upon their methodology, Fernandez2023 introduce a secondary selection equation into the model to address selection bias based on the distribution of hourly wages among women. Their approach, based on a “control function" methodology with identification from Buchinsky1995, involves semiparametric estimation techniques applied to CPS data spanning 1975 to 2020. They detect positive selection and note evolving selection effects over time for all workers (full-time workers and part-time workers). Throughout the estimation process, they make use of various specification assumptions regarding distributions and apply parametric and distribution regressions.

Instead of imposing functional form assumptions, Huber2014ER identifies the average and quantile treatment effects by adapting inverse probability weighting to the sample selection model. The method relies on the exclusion restriction—using an instrument for selection—to identify the sample selection propensity score. This propensity score is then used as a `control function.' The author applies this method to estimate the returns of high school graduation for females. The instrument, which is the number of young children and its interaction terms, must include a continuous variable for identification purposes, but it is not intended to identify more types in the population.

In contrast, Maasoumi2019 employ a quantile copula function approach, as introduced by ArellanoBook2017 and Arellano2017, to estimate parameters in quantile wage functions under quantile selection models. This approach utilizes the Frank copula. The copula parameter serves to capture the dependence between error terms within both the wage and selection equations. Implementing this method entails three primary steps under point identification. Notably, their approach necessitates an exclusion restriction for the selection part, and they define the gender gap unconventionally through evaluation functions, focusing solely on full-time workers within their dataset and reporting gender gap measures both in conventional metrics, such as the average, and in an unconventional way.

Our method bears resemblance to partial identification techniques employed in previous studies, such as the approach used by Blundell2007, which partially identifies the distributions of log hourly wages for men and women through bounds. Specifically, they focus on bounding the distributions of wages for nonworking individuals based on assumptions about the distribution of workers and nonworkers; the wage distribution of nonworkers is first-order stochastically dominated by the distribution of workers. However, unlike their approach, we treat gender as a treatment and bound the gender gap by types determined by observable and unobservable characteristics. Utilizing data from the Family Expenditure Survey (FES) in the United Kingdom, Blundell2007 finds that under certain assumptions, including stochastic dominance assumptions, the bounds for changes in the quantile treatment effect of gender at the median suggest a decline for individuals without college experience, while the gender gap decline is less apparent for college-educated individuals. Notably, their approach assumes positive selection for men and women, considering differences in the wage distributions between workers and nonworkers without categorizing individuals into specific types. Thus, while both methods aim to bound treatment effects, our approach offers a nuanced understanding of the gender gap by considering diverse types of individuals.

Our method aligns with the partial identification approach under a sample selection model, building on the foundational work of Manski1990, who introduced nonparametric bounds for treatment effects. These bounds initially remain wide without additional assumptions, but Manski1990 proposed methods to narrow them by incorporating specific assumptions and intersections. Within the framework of sample selection models, economists have developed bounds on ATE for various subpopulations. For instance, Zhang2003 established upper and lower bounds for the ATE in the “always observed” subpopulation, where outcome variables are observed irrespective of treatment status. Similarly, Lee2009 derived bounds for this subpopulation under monotonicity assumptions. Zhang2008 extended this work by bounding treatment effects for four subpopulations, including those always observed, only observed when treated, observed when untreated, and never observed individuals. They provided results without relying on monotonicity or stochastic dominance assumptions and contrasted these with results under those assumptions. Additionally, Blanco2013 began with a sample selection model devoid of distributional and monotonicity assumptions and applied nonparametric bounding methods proposed by Horowitz2000. They progressively introduced more assumptions to obtain tighter bounds on the average treatment effects of training programs such as Job Corps on wages, focusing on always-employed individuals and utilizing monotonicity and stochastic dominance assumptions to narrow the bounds further.

Our approach adopts the bounding method introduced in Lee2009 and Zhang2008, employing a nonseparate model framework where gender serves as the “treatment". Specifically, we measure the gender gap as the disparity between a man's log hourly wage and the potential outcome if he were treated as a woman. Unlike Lee2009 and Zhang2008, we incorporate an exclusion restriction, to distinguish between various population segments beyond the conventional distinctions of “always observed,” “only observed when treated,” “observed when untreated,” and “never observed” individuals. This segmentation allows us to delineate types associated with social norms and family roles, defined by both observable and unobservable characteristics. Thus, we are able to bound the gender gap, defined as the difference between the average log hourly wages of women and men, or the ATE for each type, conditional on some observed characteristics. Consequently, the treatment effect becomes heterogeneous across different types, and the introduction of the exclusion restriction enables us to obtain more information from the instrument for selection, bounding the ATE for each value of the instrument. We then intersect these bounds by instrument values, later adjusting them using the CLR method from Chernozhukov2013.

Based on the CPS data from Maasoumi2019, we draw three key insights. Overall, our findings underscore the significance of accounting for diverse individual types when assessing the gender gap. Firstly, for the dominant type of individuals who always work, we observe rapid declines in the upper bound (UB) of the gender gap until the 1990s, followed by a slower decrease thereafter, with occasional rebounds. Conversely, the decline in the lower bound (LB) is notably slower, and sometimes in later years, the LB is below 0, aligning with Maasoumi2019's results and suggesting that we do not rule out the possibility that the gender gap disappears. Additionally, the UB and LB for high-potential wage workers consistently surpass those for their low-potential counterparts, corroborating the findings in Maasoumi2019. Notably, this type always chooses to work under all kinds of conditions with high motivation. The rising proportion of those individuals is mainly due to the fact that more and more women with young children at home choose to work. Also, these results are obtained without assuming any stochastic dominance assumptions, though imposing such assumptions increases the LB, excluding 0 from the bounds while maintaining similar trends.

Secondly, we observe that the bounds for ATE vary significantly across the other types of individuals compared to the relatively stable bounds for those who always work. This variability arises from shifts in the proportions of these types over time. Notably, as the proportion of women opting not to work, particularly those with young children at home, decreases and the proportion of women choosing to work remains stable or increases, the bounds for ATE widen considerably for these types. This increased variability in bounds, reflected in fluctuating LBs and UBs, likely stems from the smaller proportion of individuals in these types. Despite some bounds increasingly excluding 0 in later years, indicating a persistent gender gap, the wide 95% confidence intervals suggest a potentially larger gender gap or no gender gap for those types.

Thirdly, our analysis sheds light on the issue of selection using our method, yielding results consistent with those of Maasoumi2019 regarding women. Specifically, our findings indicate a shift from negative to positive selection for women, aligning with previous research. However, our method reveals that this shift occurs earlier in the timeline, with more years exhibiting positive selection. This observation echoes the conclusions drawn by Fernandez2023, who also focus on all workers (full-time and part-time workers), despite differences in definitions, methodologies, and datasets. The consistency of these findings underscores the significance of addressing selection issues in understanding gender disparities in the labour market.

The structure of the paper is as follows: Section 2 outlines the basic model, including the assumptions, bounds, estimators, and their asymptotic properties for ATE. Section 3 extends this analysis by introducing additional assumptions and corresponding bounds. In Section 4, we present the findings from our empirical investigation into the gender gap. Finally, Section 5 offers our conclusions. Additional tables and graphs are provided in the Appendix, with further empirical results detailed in the Supplementary Appendix.

the Basic Model

In this section, we introduce the basic model, focusing on the impact of the binary treatment variable $D$ on the outcome variable $Y$, while considering unobservable types and observable characteristics within a sample selection model. Unlike the endogeneity of treatment in studies like Imbens1994 and Angrist1996, here the endogenous variable is the selection indicator $S$. $X$ represents a vector of observable individual characteristics.\footnote{Initially, we examine the scenario where covariates ($X$) are excluded. We then add $X$ to the analysis by conditioning the assumptions, the conditional probabilities, and the conditional expectations, for instance, on $X$. The analysis then proceeds along the same lines as when $X$ was excluded.} As discussed in Mellace2011 and Huber2012, we first adopt a structural model:

align*[align* omitted — 69 chars of source]

Utilizing the potential outcome framework and following Lee2009 and Huber2015, the structural model becomes:

align*[align* omitted — 79 chars of source]

In the context of gender gap studies, for instance, in Maasoumi2019, $S$ corresponds to employment status, $D$ represents gender, $Z$ denotes the indicator of not having young children (see below), and $Y$ is log hourly wage rate.

The researcher observes ${D, Z, S, Y}$ with $Z$ taken as an instrumental variable, and by inserting $Z$ within the selection function for $S$ we are able to ascertain different behavior patterns and provide insights into different types of individual.

asuThe pair $(D, Z)$ is independent of the pair $(U, V)$; denoted as $(D, Z) \perp (U, V)$. If we include covariates $X$ in the discussion, then this assumption is generalised to $(D, Z) \perp (U, V)|X$.

Assumption (ref) implies that $(S(d,z), Y^*(d))$ and the treatment are independent, or the potential outcomes are independent of the treatment. Lee2009 and Huber2015 define the types without the exclusion restriction, using $S(d)$ to define the types and separate individuals into four subpopulations. In our work, we define the types via $S(d,z)$. After that, we bound the average treatment effect for each type, i.e., $E(Y^*(1) - Y^*(0)|\text{Type T})$.

The exclusion restriction on $Z$ is testable. Huber2014 propose a testing method under a potential outcome framework, and Maasoumi2019 utilize this method to test the validity of the exact same instrument as we use in the application section without rejecting the hypothesis about the exclusion restriction of the instrument.

asu$S = g(D, Z, V)$, where $V \in (0,1)$ is continuously distributed and $g$ is weakly nonincreasing in $V$. The function $g$ is normalized so that V $\sim U(0,1)$.

Assumption (ref) is a standard assumption found in works such as Chesher2010 and Bartalotti2023. In Chesher2010, this assumption applies to discrete outcome variables, implying that the outcome follows a threshold-crossing model. Similarly, in Bartalotti2023, where the treatment variable is discrete and there exists an instrument for it, this assumption suggests that the model equation for the treatment involves threshold crossing. In our context, the instrument enters solely via the selection component, indicating that it is only this part of the model that operates as a threshold-crossing process and hence that Assumption (ref) is not restrictive. After imposing Assumption (ref) the selection equation of the structural model becomes $S = 1[h(D, Z)-V > 0]$, with threshold-crossing form for $S(d, Z)$, $d\in \{0,1\}$, given by

equation*[equation* omitted — 150 chars of source]

In the first step, we aim to determine the proportions of individuals belonging to each type. Subsequently, we use these proportions to generate bounds for the average treatment effect for each type. In gender gap studies, for instance, this process identifies the proportion of individuals who would work if they were women with young children at home. To find the proportions of a certain type, we identify the proportions of individuals who will work when they are subjected to treatment or remain untreated, influenced by a specific value of the instrument. $P_r(S = 1| D = 1, Z = z)$ represents the proportion of individuals observed ($S = 1$) when they receive treatment ($D = 1$) with $Z = z$, while $P_r(S = 1| D = 0, Z = z)$ indicates the proportion of individuals observed when they are untreated ($D = 0$). With $d\in \{0,1\}$, $$ P_r(S = 1| D = d, Z = z) = P_r(S(d,z) = 1| D = d, Z = z) = P_r(S(d,z) = 1) = P_r(V < h(d, z))\,. $$ The first equality derives from the model, the second from Assumption (ref), and the final one from Assumption (ref). These equations show that when Assumption (ref) and Assumption (ref) hold the probability $P_r(S = 1| D = d, Z = z)$ is $P_r(V < h(d, z))$, i.e., $h(d, z)$. That is, if we identify the conditional probabilities, we identify the thresholds.

Introduction of Different Types

In the selection framework, we build upon definitions proposed by Huber2015 based on potential selection outcomes. For instance, individuals with $S(d) = 1$ in Huber2015 and Lee2009 are termed “always observed," while $(S(1) = 1, S(0) = 0)$ are “compliers," and $(S(1) = 0, S(0) = 1)$ are “defiers." However, in our framework, with $S(d,z)$ instead of $S(d)$, we need new definitions. Our application in Section (ref) will clarify the importance of these distinctions.

In our gender gap study, $Z$ indicates the presence of young children and is exclusively tied to the selection part of the model. Thus, $Z$ serves as an indicator of an individuals' observability (whether they are employed or not employed). Without additional assumptions, this implies that there are 16 distinct types determined by the values of $S$, $D$, and $Z$, as shown in the first four columns of Table (ref).

The identification challenge with $S(d,z)$ resembles that of $S(d)$, as each individual is observed with only one $S(d,z)$ value, which can be either 0 or 1. Without additional assumptions, multiple types may correspond to one observed case for all possible combinations of $D$ and $Z$ with $S = 1$ (e.g., 8 types in the case of $(S = 1, D = 1, Z = 1)$), and there will not be point identification for the proportions of each type. We will need to bound the proportions of each type, as in Huber2015. Here, we introduce the assumptions to define these types and point identify those proportions, as in Lee2009.

table[table omitted — 1,248 chars of source]
asu(Monotonicity w.r.t D) $S(1,z) \geq S(0,z)$ with probability 1.

Assumption (ref) asserts that the probability of $S(1,z)$ being greater than or equal to $S(0,z)$ holds with certainty for each $Z = z$. This assumption helps rule out situations where potential selection status decreases as a result of the increased treatment. Assumption (ref) excludes Types 5, 7, 9, 10, 13, 14, and 15, as detailed in Table (ref). For instance, Types 9, 10, 13, and 14, when $Z = 0$, have potential selection values of $S(1,0)$ less than $S(0,0)$, and Types 5, 7, 13, and 15, when $Z = 1$, have $S(1,1)$ less than $S(0,1)$.

asu(Monotonicity w.r.t Z by treatment status):\\ \refstepcounter{subassumption} (\alph{subassumption}) \ignorespaces $S(0,0) \geq S(0,1)$ with probability 1. \\ \refstepcounter{subassumption} (\alph{subassumption}) \ignorespaces $S(1,0) \leq S(1,1)$ with probability 1.

The direction of the inequalities in Assumptions (ref) and (ref) will vary depending on the specific empirical context. For our example, where $D = 0$ denotes females, Assumption (ref) implies that women will not work when they have young children if they are not working when they do not have young children. Similarly, Assumption (ref) suggests that for males, having young children at home will work if they work when they do not have young children. Under Assumptions (ref) and (ref), certain types (3, 5, 6, 7, 8, 11, and 15) are excluded, as shown in Table (ref).

Assumptions (ref) and (ref) are not testable, but there are corresponding necessary conditions for them. In our empirical application, Assumptions (ref) aligns with situations where the probability of $S = 1$, given $D = 1$ and $Z = z$, is greater than or equal to the probability when $D = 0$ and $Z = z$. That is, $P_r(S = 1|D = 1, Z = z) \geq P_r(S = 1|D = 0, Z = z)$ for $z \in \{ 0, 1 \}$. Assumption (ref) aligns with the data when $P_r(S = 1|D = 0, Z = 0) \geq P_r(S = 1|D = 0, Z = 1)$. Assumption (ref) is consistent with the data when $P_r(S = 1|D = 1, Z = 1) \geq P_r(S = 1|D = 1, Z = 0)$. However, it is important to note that these inequalities are necessary but not sufficient conditions for Assumptions (ref) and (ref) to hold; that is, if the inequalities are satisfied in the empirical application it does not mean that the monotonicity assumptions are valid.

figure[figure omitted — 178 chars of source]
table[table omitted — 1,088 chars of source]
table[table omitted — 536 chars of source]

There are five distinct types defined under Assumptions (ref)-(ref), as shown in the final column of Table (ref) and visually represented in Figure (ref). In Table (ref), we list the five types or strata according to the value of $S(d,z)$. We introduce another kind of notation for each type. Specifically, we use `E' to denote the condition $S(d, z) = 1$ and `N' to denote $S(d, z) = 0$. For example, the notation `EEEE' represents the stratum consisting of individuals always observed under all combinations of $d \in \{0,1\}$ and $z \in \{0,1\}$, characterized by $\{S(1,1) = S(1,0) = S(0,1) = S(0,0) = 1\}$. Based on this table, we construct Table (ref) to illustrate the selection outcomes corresponding to each type under the case $(D = d, Z = z)$ with $d \in {0,1}$ and $z \in {0,1}$. Table (ref) provides insight into the relationship between those observed cases and the five types.

Let us focus on what each type signifies. Women belonging to Type 2 (EENE stratum) will engage in work if they do not have young children (those aged less than 5 years). On the other hand, women of Type 4 (EENN stratum) will refrain from work regardless of whether they have young children or not. Interestingly, if we were to consider males in Type 2 and Type 4, they would be associated with employment. Conversely, women categorized as Type 12 (ENNN stratum) will abstain from work, regardless of whether they have young children in their care. However, males of the same type would only choose to work when they have young children to look after. From this discussion and Table (ref), we see that among those who do not work, there are different types. This classification system provides insight into the behaviour of individuals of different types in various scenarios, shedding light on how various factors influence their decisions about work.

Identification of the Population Proportions

Let us now shift our focus to the proportions of each type in the population conditional on $D$ and $Z$. It is important to remember that $V$ represents the error term within the selection equation, effectively influencing the assignment of types. This is shown in Figure (ref). The unobservable $V$ and the observable characteristics define the types. The relationship between $U$ and $V$ could be quite complex. Under the basic model, we do not assume any assumptions between $U$ and $V$.

Assumption (ref) ensures that $D, Z$ and $V, U$ are independent. Consequently, the proportion of each type conditional on the treatment ($D$) and the instrument ($Z$) is the same as the unconditional proportion of each type in the whole population. In other words, in Table (ref), the proportion of each type within a specific case $(D = d, Z = z)$ when conditioned on treatment and the instrument is the same across all combinations of $D, Z$. The intuition is the same as in Lee2009. For instance, within the cell where $(S = 1, D = 0, Z = 1)$ in Table (ref), Type 1 prevails and all individuals within the cell are of Type 1. Thus, in cases where $(S = 1, D = 0, Z = 1)$ an individuals type is known and the unconditional proportion of Type 1 ($\pi_{T1}$) is defined as $\pi_{T1} = P_r(S =1| D = 0, Z = 1)$.\footnote{$ P_r(S =1| D = 0, Z = 1) = P_r(S(0,1) =1| D = 0, Z = 1) = P_r(S(0,1) =1) = P_r(S(0,1) =1, S(0,0) =1,S(1,1) =1,S(1,0) =1) = P_r(\text{Type 1}) = \pi_{T1}$. The first equality is from the model setting, the second one is from Assumption (ref), the third is from all of the four Assumptions so far, and the fourth one is from the definition of Type 1.}

Moving to the cell characterized by $(S = 1, D = 0, Z = 0)$ in Table (ref), we find individuals belonging to Type 1 or Type 2. Given our knowledge of the unconditional proportion of Type 1, we can determine the proportion of Type 1 within this group as follows: $\frac{\pi_{T1}}{\pi_{T1}+ \pi_{T2}}= \frac{P_r(S =1| D = 0, Z = 1)}{P_r(S =1| D = 0, Z = 0)}$. \footnote{$P_r(S =1| D = 0, Z = 0) = P_r(S(0,0) =1| D = 0, Z = 0) = P_r(S(0,0) =1) = P_r(S(0,0) =1, S(0,1) =1) + P_r(S(0,0) =1, S(0,1) =0) = P_r(\text{Type 1}) + P_r(\text{Type 2}) = \pi_{T1}+ \pi_{T2}$. Hence, $\frac{P_r(S =1| D = 0, Z = 1)}{P_r(S =1| D = 0, Z = 0)} = \frac{\pi_{T1}}{\pi_{T1}+ \pi_{T2}}$ under the four assumptions.} Simultaneously, we can calculate the proportion of Type 2 within the same group using the following expression: $\frac{\pi_{T2}}{\pi_{T1}+ \pi_{T2}} = \frac{P_r(S =1| D = 0, Z = 0) - P_r(S =1| D = 0, Z = 1)}{P_r(S =1| D = 0, Z = 0)}$ \footnote{$P_r(S =1| D = 0, Z = 0) - P_r(S =1| D = 0, Z = 1) = \pi_{T1}+ \pi_{T2} - \pi_{T1} = \pi_{T2}$}. These calculations utilize the unconditional proportion of Type 2 ($\pi_{T2}$), which is determined as the difference between the probability of $S = 1$ when $D = 0$ and $Z = 0$ and the probability when $D = 0$ and $Z = 1$: $\pi_{T2} = P_r(S =1| D = 0, Z = 0) - P_r(S =1| D = 0, Z = 1)$.

Next, in the cell $(S = 1, D = 1, Z = 0)$ we observe the presence of three different types: Types 1, 2, and 4. All three types share the characteristic of having $S(1,0) = 1$. We can ascertain the unconditional proportion of Type 4 ($\pi_{T4}$) as follows: $\pi_{T4} = P_r(S =1| D = 1, Z = 0) - P_r(S =1| D = 0, Z = 0)$ \footnote{ $P_r(S =1| D = 1, Z = 0) - P_r(S =1| D = 0, Z = 0) = P_r(S(1,0) =1) - P_r(S(0,0) =1) = P_r(S(1,0) =1, S(0,0) =1)+ P_r(S(1,0) =1, S(0,0) =0) - P_r(S(1,0) =1, S(0,0) =1) = P_r(S(1,0) =1, S(0,0) =0) = P_r(\text{Type 4}) = \pi_{T4}$}. Consequently, the proportion of Type 4 within this cell can be calculated as: $\frac{\pi_{T4}}{\pi_{T1}+ \pi_{T2}+ \pi_{T4}}$. Meanwhile, the proportion of Type 1 in the same cell is $\frac{\pi_{T1}}{\pi_{T1}+ \pi_{T2}+ \pi_{T4}}$, and that of Type 2 is $\frac{\pi_{T2}}{\pi_{T1}+ \pi_{T2}+ \pi_{T4}}$.

Finally, the probability of Type 12 in the cell characterized by $D = 1, Z = 1, S = 1$ can be determined using the following expression: $\frac{\pi_{T12}}{\pi_{T1}+ \pi_{T2}+ \pi_{T4}+\pi_{T12}}$ with $\pi_{T12} = P_r(S =1| D = 1, Z = 1) - P_r(S =1| D = 1, Z = 0)$ \footnote{ $P_r(S =1| D = 1, Z = 1) - P_r(S =1| D = 1, Z = 0) = P_r(S(1,1) =1| D = 1, Z = 1) - P_r(S(1,0) =1| D = 1, Z = 0) = P_r(S(1,1) =1) - P_r(S(1,0) =1) = P_r(S(1,1) =1, S(1,0) =1) + P_r(S(1,1) =1, S(1,0) =0)- P_r(S(1,0) =1,S(1,1) =1) = P_r(S(1,1) =1, S(1,0) =0) = P_r(\text{Type 12}) = \pi_{T12}$}. The proportion of Type 1 within the same cell is determined as: $\frac{\pi_{T1}}{\pi_{T1}+ \pi_{T2}+ \pi_{T4}+\pi_{T12}}$. The proportion of Type 2 can be calculated as: $\frac{\pi_{T2}}{\pi_{T1}+ \pi_{T2}+ \pi_{T4}+\pi_{T12}}$, and the proportion of Type 4 within this group can be determined as: $\frac{\pi_{T6}}{\pi_{T1}+ \pi_{T2}+ \pi_{T4}+\pi_{T12}}$.

This analysis enables us to discern the proportions of different types within various observed cases. In the context of the gender gap study, it sheds light on the proportions of each type. With data spanning multiple years, it reveals how the types evolve over time within the gender gap study. Moreover, by segmenting the data set into different covariate groups for each year, it demonstrates how the covariate groups and types interact. Thus, it contributes to a deeper understanding of the distribution of these types across different scenarios.

Derivation of Bounds

In this subsection, we discuss the bounds for the ATE for all types within the population: Types 1, 2, 4, 12, and 16. These bounds are established under Assumptions (ref)-(ref), collectively referred to as the basic assumptions. Under these assumptions, we establish the widest possible bounds, representing the most conservative scenario. The ATE for Type T is expressed as $E(Y^*(1) - Y^*(0)|\text{Type T})$. For Type 16, without additional assumptions, the ATE bounds extend across the entire real number line $(-\infty, \infty)$. However, by assuming that potential outcomes for all types are bounded within the interval $[Y^{LB}, Y^{UB}]$, denoted as $Y^*(1)$ and $Y^*(0)$, the ATE bounds narrow down to $[Y^{LB} - Y^{UB}, Y^{UB} - Y^{LB}]$.

Bounds for \texorpdfstring{$E(Y^*(1)|\text{Type T})$}{E(Y(1)|Type T)}

As discussed earlier, we begin by deriving the sharp bounds for potential outcomes when treated by the 2 values of the instrument. The proportions of each type within cells $S = 1, D = 1, Z = 1$ and $S = 1, D = 1, Z = 0$ are calculated as $\frac{\pi_{T.}}{\pi_{T1}+ \pi_{T2}+ \pi_{T4}+\pi_{T12}}$ and $\frac{\pi_{T.}}{\pi_{T1}+ \pi_{T2}+ \pi_{T4}}$, respectively, under Assumptions (ref) to (ref). Next, we follow the same logic as in Proposition 1a in Lee2009 and Section 3 of Huber2015 to find the sharp upper and lower bounds for each value of $Z$. This is standard procedure in sample selection models. Hence, we have two sets of upper and lower bounds for the same parameter.

When $Z = 1$, we have the upper and lower bounds for $E(Y^*(1)|\text{Type T})$. The upper and lower bounds are in the following: $$UB = E[Y|D = 1, S = 1, Y \geq Y_{1 - q_T}, Z = 1],$$ $$LB= E[Y|D = 1, S = 1, Y \leq Y_{q_T}, Z = 1]$$ Here, $q_T$ represents the proportion of Type T in that cell, such as $\frac{\pi_{T1}}{\pi_{T1}+ \pi_{T2}+ \pi_{T4} + \pi_{T12}}$ for Type 1. The bounds for Types 2, 4, and 12 in the same cell are calculated similarly, with the only difference being in the proportions used for each type.\footnote{We show the relations between observable distributions and unobservable distributions for types and proportions in Appendix}

When $Z = 0$, under Assumption (ref), (ref), (ref), and (ref), the upper and lower bounds for $E(Y^*(1)|\text{Type T})$ are computed as: $$UB = E[Y|D = 1, S = 1, Y \geq Y_{1 - q_T}, Z = 0]$$ $$LB=E[Y|D = 1, S = 1, Y \leq Y_{q_T}, Z = 0]$$ These bounds for Types 1, 2, and 4 are based on the proportions of each type in the respective case.

Since $Y^*(1)$ does not depend on Z, the bounds when $Z = 1$ and $Z = 0$ provide more information to identify the same parameter. As a result, for Types 1, 2, and 4, we take the intersection of these two sets of bounds in the second step, further refining the ATE bounds. This is the same idea of intersection as in Chesher2010, Huber2015, and Huber2017.

Bounds for \texorpdfstring{$E(Y^*(0)|\text{Type T})$}{E(Y(0)|Type T)} and ATE

In this subsection, we calculate the average potential outcomes when untreated for Type 1 and Type 2. When $D = 0$ and $Z = 1$, the observed ($S = 1$) type is Type 1. In our empirical application, this means that females with young children (less than 5 years old) who are employed are of Type 1. Recall that in the cell $(S = 1, D = 0, Z = 1)$ of Table (ref), there is only Type 1. Thus, $E(Y|S = 1, D = 0, Z = 1)$ is the average potential outcome when untreated for Type 1, i.e., $E(Y^*(0)|\text{Type 1})$. \footnote{$ E(Y|D = 0, S = 1, Z = 1) = E(Y^*(0)|D = 0, S(0,1) = 1, Z = 1) = E(Y^*(0)|S(0,1) = 1) = E(Y^*(0)|\text{Type 1})$}.

Next, we identify $E(Y^*(0)|\text{Type 2})$. Notice that Type 1 and Type 2 are inside the cell $(S = 1, D = 0, Z = 0)$ of Table (ref). With the information we have, we point identify the average potential outcome when untreated for Type 2. Indeed, we point identify the $E(Y^*(0)|\text{Type 1})$, the proportions of Type 1 and Type 2 in the cell ($\frac{\pi_{T1}}{\pi_{T1}+ \pi_{T2}}$ and $\frac{\pi_{T2}}{\pi_{T1}+ \pi_{T2}}$), and the overall average potential outcome when untreated for Type 1 and Type 2, $E(Y|S = 1, D = 0, Z = 0)$. The point identification for the untreated potential outcome of Type 2 is achieved. This is shown in the following. The proof is in the Appendix. $$E(Y|S = 1, D = 0, Z = 0) = E(Y^*(0)|\text{Type 1}) * \frac{\pi_{T1}}{\pi_{T1}+ \pi_{T2}} + E(Y^*(0)| \text{Type 2})*\frac{\pi_{T2}}{\pi_{T1}+ \pi_{T2}}$$

Notice that individuals of Types 4 and 12, when untreated, are not observed for any value of $Z$. That is, in Table (ref) or Table (ref), $S (0, z)$ is 0 for these two types. Hence, we do not obtain any information from the observed data for $E(Y^*(0)|\text{Type T})$ with $T \in \{4, 12\}$. Without any further assumption, we use theoretical bounds for those values: $E(Y^*(0)|\text{Type T}) \in [Y^{LB}, Y^{UB}]$.

To summarize, the bounds for the ATE for Types 1 and 2 are calculated as follows: $$UB_{T} = \min_{z} E[Y|D = 1, S = 1, Y \geq Y_{1 - q_{T_z}}, Z = z]-E(Y^*(0)|\text{Type T})$$ $$LB_{T}= \max_{z} E[Y|D = 1, S = 1, Y \leq Y_{q_{T_z}}, Z = z]-E(Y^*(0)|\text{Type T})$$

For Type 4, the ATE bounds are: $$UB_{T4} = \min_{z} E[Y|D = 1, S = 1, Y \geq Y_{1 - q_{T4_0}}, Z = z]-Y^{LB}$$ $$LB_{T4}= \max_{z} E[Y|D = 1, S = 1, Y \leq Y_{q_{T4_0}}, Z = z]-Y^{UB}$$

And for Type 12: $$UB_{T12}= E[Y|D = 1, S = 1, Y \geq Y_{1 - q_{T12_0}}, Z = 1]-Y^{LB}$$ $$LB_{T12}= E[Y|D = 1, S = 1, Y \leq Y_{q_{T12_0}}, Z = 1]-Y^{UB}$$

The bounds are sharp under Assumptions (ref) to (ref). That is, they are the tightest bounds that we get under the assumptions. The detailed discussion on the sharpness of the bounds is in Proposition 1a. from Lee2009.

Bounds with Covariates

In the previous discussion, we temporarily excluded the consideration of covariates in our model. Here, we delve into the explicit inclusion of covariates $X$ or individual characteristics. It is straightforward. That is, we apply the discussion, proofs, and results with respect to each value of the covariate in $X$, as in Lee2009. This is a similar method to working with covariates, as in Angrist1996. That is, if education level is one of the variables inside covariates $X$, the discussion is applied with respect to each education level. In the gender gap study, the covariates include many individual characteristics. We will discuss this in Section (ref).

To integrate $X$ into our analysis, we introduce an amended set of assumptions. For example, Assumption (ref) now takes the form: $D$ and $Z$ are independent of $U$ and $V$ given $X$. Similarly, monotonicity assumptions are based on each value of each covariate in $X$. We directly update the assumption versions. The revised version of Assumption (ref) is a conditional independence assumption discussed in Holland1986 and is a direct extension of the confoundedness assumption from Imbens2004Review.

In the gender gap study, we employ this revised set of assumptions, recognizing the significance of covariates in influencing both outcomes and selection. These covariates, along with their polynomials, are also accounted for in Maasoumi2019, where we adopt the same covariates and polynomial specifications.

The bounds for the ATE for Types 1 and 2 are calculated as: $$UB_{T} = \min_{z} E[Y|D = 1, S = 1, Y \geq Y_{1 - q_{T_z}}, Z = z, X = x]-E(Y^*(0)|\text{Type T}, X = x)$$ $$LB_{T}= \max_{z} E[Y|D = 1, S = 1, Y \leq Y_{q_{T_z}}, Z = z, X = x]-E(Y^*(0)|\text{Type T}, X = x)$$

For Type 4, the ATE bounds are: $$UB_{T4} = \min_{z} E[Y|D = 1, S = 1, Y \geq Y_{1 - q_{T4_z}}, Z = z, X = x]-Y^{LB}$$ $$LB_{T4}= \max_{z} E[Y|D = 1, S = 1, Y \leq Y_{q_{T4_z}}, Z = z, X = x]-Y^{UB}$$

And for Type 12: $$UB_{T12} = E[Y|D = 1, S = 1, Y \geq Y_{1 - q_{T12_1}}, Z = 1, X = x]-Y^{LB}$$ $$LB_{T12}= E[Y|D = 1, S = 1, Y \leq Y_{q_{T12_1}}, Z = 1, X = x]-Y^{UB}$$

The bounds for the various types are sharp, following the approach outlined in Lee2009 and Blanco2013. We employ the covariates to tighten these bounds, as they possess predictive power for $Y$. By conditioning on the covariates, we group individuals with similar predicted values of $Y$, thereby potentially reducing the variability in $Y$. Given that our bounding method relies on the values of $Y$, the intervals for treatment effects based on covariate groups are consequently narrower. When examining the overall effect of the ATE across all $X$ groups, we compute the average LB and UB across these groups based on the method proposed by Lee2009 for calculating `total' in Table 5.

Estimation and Inference

In this subsection, we compute the sample versions of the population parameters outlined in the previous subsection, following the methodology presented in Lee2009. For illustrative purposes, we simplify by disregarding the covariates in the bounds. Additionally, we streamline the illustration by only furnishing estimates for Type 1 and Type 2 when $Z = 1$. The estimators for Type 4 and Type 12 follow a similar approach. The estimators for Type 1 when $Z = 1$ are:

align[align omitted — 790 chars of source]

The initial component of $\widehat{UB_{T1}}$ contains the upper bound for $E(Y^*(1)|\text{Type 1})$, while the corresponding segment of $\widehat{LB_{T1}}$ is the lower bound for $E(Y^*(1)|\text{Type 1})$. $\widehat{y_p} $ is the quantile, and $ \widehat{q}$ is the proportion of Type 1 in the case of $(S = 1, D = 1, Z = 1)$ in Table (ref). The numerator of $\widehat{q}$ serves as an estimator for $\pi_{T1}$ and maintains consistency across different values of $Z$. The denominator of $\widehat{q}$, however, changes with the values of $Z$. The ultimate equation furnishes the estimate for $E(Y^*(0)|\text{Type 1})$. When computing the bounds for $Z = 0$, we substitute $Z$ with $(1-Z)$ in the equations above, with the exception of the numerator of $\widehat{q}$ and the final equation.

For Type 2, the estimators are similar except for the last two equations, that is, the estimator for the proportion of Type 2 for each value of $Z$ and the one for $E(Y^*(0)|\text{Type 2})$. For this type,

align[align omitted — 559 chars of source]

Note that the denominator of $\widehat{Y_{0,T2}}$ is $\widehat{\pi_{T2}}$. In our gender gap study, this proportion is a positive number, which is a necessary condition for Assumptions (ref) and (ref) to be valid.

For Types 4 and 12, the estimates $E(Y^*(0)|\text{Type T})$ under the basic assumptions are in the theoretical bounds $[Y^{LB}, Y^{UB}]$. We follow the steps in Huber2015 using the minimum and maximum values from the dataset for each year.

Asymptotic Properties

We derive the asymptotic properties, including consistency and asymptotic normality, for the estimators in Section (ref) by employing the GMM method proposed in Propositions 2 and 3 of Lee2009. This approach is suitable as the estimators represent solutions to just-identified GMM problems, as outlined in Newey1994. Specifically, the parameters of interest are included within moments, resulting in a just-identified case where the number of parameters to estimate equals the number of moments. Using these moments, we construct sample analogs and subsequently derive GMM-type estimators, as detailed in Section (ref). Drawing from the methodologies outlined in Newey1994 and Lee2009, we establish the asymptotic properties for these estimators. Our approach aligns with the assumptions proposed by Lee2009, with additional considerations incorporating the instrument variable as an extra condition for each type.

theorem(Consistency and Asymptotic Normality) $Y(1)$, $Y(0)$ are bounded. With Assumptions (ref) - (ref), and extra assumptions $0 < E[S|D = d, Z = z] < 1$ for $d, z \in \{0, 1\}$, $ \widehat{UB_T} \rightarrow^{p} UB_T$ and $\widehat{LB_T}\rightarrow^{p} LB_T$, as well as\\ $\sqrt{n}(\widehat{UB_T} - UB_T) \rightarrow^d N(0, V^{UB}_T + V^C_T)$ and $\sqrt{n}(\widehat{LB_T} - LB_T) \rightarrow^d N(0, V^{LB}_T + V^C_T)$ with $V^{LB}_T$ and $V^C_T$ defined for each type and $Z$ in the Appendix.

Given that $0 < E[S|D = d, Z = z] < 1$ for $d, z \in \{0, 1\}$ this indicates that the conditional probabilities are constrained to be between 0 and 1. With Assumptions (ref) through (ref), it follows that the proportions of types also fall within the range of 0 to 1. The proof for Theorem (ref) is provided in the appendix. In the appendix, we also provided the covariance of $\widehat{UB}_T - UB_T$ and $\widehat{LB}_T - LB_T$ when $Z = 1$ and $Z = 0$ and the covariance between $V^C_T$ of Type 1 and $V^C_T$ of Type 2. Those covariances are needed to calculate the CLR bounds from Chernozhukov2013. CLR bounds are necessary, because the intersection bounds contain the min and max operators, creating biased UBs and LBs. In Section (ref), we obtain half-median unbiased estimates of UBs and LBs, as well as confidence intervals using the method from Chernozhukov2013.

Further Assumptions and Tighter Bounds

Mean Dominance Assumption

In this section, we introduce further assumptions to tighten the bounds. The utilization of stochastic dominance assumptions in sample selection frameworks has been demonstrated by previous researchers. See, inter alia, Zhang2003, Zhang2008, Blanco2013, Huber2015, and Bartalotti2023. For instance, Bartalotti2023 have employed stochastic assumptions to refine the bounds for individuals always observed. They assume that the distribution of potential outcomes when treated for the always observed subpopulation exhibits first-order stochastic dominance over the distributions of these potential outcomes for compliers, defiers, or those who are never observed. Notably, Bartalotti2023 introduce this assumption when an instrument is used for the treatment, so there will be extra conditions under their setting. Blanco2013 and Huber2017, on the other hand, rely on a mean dominance assumption. We adopt similar assumptions to Blanco2013 and Huber2017 because we are focused on the mean effects for each type.

We apply Assumptions (ref) and (ref) in our gender gap application where the treatment is exogenous and the instrument pertains to the selection process. \footnote{It is important to note that the direction of the inequality is based on the context or the specific application. In a different application, the direction of the inequality may change. This is similar to Assumptions (ref) and (ref).}

asu(Mean Dominance Assumption) \\ \refstepcounter{subassumption} (\alph{subassumption}) \ignorespaces $E(Y^*(0)| \text{Type j}) \geq E(Y^*(0)| \text{Type k})$ with $j \in \{1, 2\}$ and $k \in \{4, 12\}$\\ \refstepcounter{subassumption} (\alph{subassumption}) \ignorespaces $E(Y^*(1)| \text{Type 1}) \geq E(Y^*(1)| \text{Type 2}) \geq E(Y^*(1)| \text{Type 4}) \geq E(Y^*(1)| \text{Type 12})$

In Assumption (ref), it is stated that the mean potential outcomes of Type 1 and Type 2 surpass those of Type 4 and Type 12 when $D = 0$. Applied to the gender gap, this suggests that the average hourly wages of women in Type 1 and Type 2 exceed those in Type 4 and Type 12. Recall that women in Type 4 and Type 12 do not work, while those in Type 1 always work, and those in Type 2 work if there are no young children at home. The idea behind the assumption is that individuals of Type 1 and Type 2 may exhibit capability traits or are more motivated perhaps. Thus, their performance in the labour market is `better' on average. The assumption only affects the bounds for Type 4 and Type 12, leaving the bounds for Type 1 and Type 2 unchanged. It implies that individuals who work tend to fare better in the labour market, suggesting a form of positive selection. We provide further discussion in Section 4.

In Zhang2008 and Blanco2013, the stochastic dominance assumption is also assumed between types, the always working individuals and the other types. Zhang2008 assume the distribution of income for the always employed individuals first-order stochastic dominate that for the compliers. Blanco2013 compare the mean potential outcomes across the two types: always observed type and compliers.

Assumption (ref) extends the logic to the average hourly wages of men in the gender gap. Similarly, it implies that individuals in Type 1 and Type 2 typically achieve higher mean potential outcomes compared to those in Type 4 and Type 12. However, it suggests that for men, there is a ranking of average hourly wages for different types when individuals are engaged in the labour market.

Both of the assumptions are not directly testable; however, we are able to provide some insights into the mean dominance assumption. Specifically, we identify the mean potential outcomes when $D = 0$ for Type 1 and Type 2. Comparing the estimates for $E(Y^*(0)|\text{Type 1})$ and $E(Y^*(0)|\text{Type 2})$, we have certain intuition about whether women who are always observed have higher potential outcome values than those who work intermittently due to childcare responsibilities. It is also important to note that these insights are based solely on the fundamental assumptions outlined in Assumptions (ref)-(ref). With this comparison, we extend this intuition to other types.

Tighter Bounds for Types 1, 2, 4, and 12 under Assumptions (ref) and (ref)

Assuming the mean dominance assumption for $D = 1$ and $D = 0$ scenarios leads to tighter bounds for all types. Under Assumption (ref), the bounds of untreated potential outcomes for Types 4 and 12 are narrowed, as $E(Y^*(0)|\text{Type 1})$ and $E(Y^*(0)|\text{Type 2})$ are point identified.\footnote{The bounds under the basic model and Assumption (ref) are presented in the Appendix.}

Assumption (ref) ensures that in the $(S = 1, Z = 1, D = 1)$ cell, the expected $Y^*(1)$ for Type 1 surpasses the expectation of Types 1, 2, 4, and 12, following the intuition from Huber2015. As a result, the lower bound for the potential outcome of Type 1 when treated and $Z = 1$ is $E(Y|D = 1, S = 1, Z = 1)$. Conversely, for Type 12 in the $(S = 1, Z = 1, D = 1)$ cell, the expected $Y^*(1)$ is lower than the expectation for the cell; that is, the upper bound for the potential outcome of Type 12 when treated is $E(Y|D = 1, S = 1, Z = 1)$.

Similarly, in the $(S = 1, Z = 0, D = 1)$ cell, the expected $Y^*(1)$ for Type 1 outperforms the overall expectation: $E(Y|D = 1, S = 1, Z = 0)$. The expectation for the potential outcome when treated for Type 4 should be lower than the expectation. Whether the expectation for the potential outcome when treated for Type 2 will be lower or higher depends on the proportion of each type.

Overall, under Assumption (ref), the bounds for the ATE for Type 1 are calculated as follows: $$UB_{T1} = \min_{z} E[Y|D = 1, S = 1, Y \geq Y_{1 - q_{T1_z}}, Z = z]-E(Y^*(0)|\text{Type 1})$$ $$LB_{T1}=  \max_{z} \{E(Y|D = 1, S = 1, Z = z)\}-E(Y^*(0)|\text{Type 1})$$

For Type 2, the lower bounds remain the same. However, the upper bounds when $Z = 1$ and $Z = 0$ change under the first inequality sign of Assumption (ref) since the upper bounds of Type 2 should be lower than the upper bounds of Type 1. The updated bounds for the ATE for Types 2 are as follows: $$UB_{T2} = \min_{z} E[Y|D = 1, S = 1, Y \geq Y_{1 - q_{T1_z}}, Z = z]-E(Y^*(0)|\text{Type 2})$$ $$LB_{T2}=  \max_{z} E[Y|D = 1, S = 1, Y \leq Y_{q_{T2_z}}, Z = z]-E(Y^*(0)|\text{Type 2})$$

With Assumptions (ref) and (ref), for Type 4, the bounds on the ATE are as follows: $$UB_{T4} =  \min \{E(Y|D = 1, S = 1, Z = 0), C\}-Y^{LB}$$ $$LB_{T4}=  \max_{z} E[Y|D = 1, S = 1, Y \leq Y_{q_{T4_z}}, Z = z]- \min_{T \in \{1,2\}}  E(Y^*(0)| \text{Type T})$$

The term $C =E[Y|D = 1, S = 1, Y \geq Y_{1 - q_{T_z}}, Z = z]$. i.e., C is the upper bound of $E(Y^*(1)|\text{Type T})$ for Type 1, Type 2, and Type 4.

For Type 12, the bounds on the ATE are as follows: $$UB_{T12} =  \min_z \{E(Y|D = 1, S = 1, Z = z) \}-Y^{LB}$$ $$LB_{T12}=  E[Y|D = 1, S = 1, Y \leq Y_{q_{T12_0}}, Z = 1]- \min_{T \in \{1,2\}}  E(Y^*(0)| \text{Type T})$$

If we only assume that Type 1's expected potential outcome when treated dominates that of Types 4 and 12, then the upper bounds of Types 4 and 12 are as follows: $$UB_{T} = \min_{z} E[Y|D = 1, S = 1, Y \geq Y_{1 - q_{T1_z}}, Z = z]-Y^{LB}.$$

For estimation, we follow the previous estimation strategy for the trimmed conditional expectations and calculate the sample analogue for the conditional expectations $E(Y|D = 1, S = 1, Z = z)$ when $z \in \{0, 1\}$. The asymptotic distribution also follows Theorem (ref).

Empirical Application

In this section, we apply our method to bound the average treatment effect of gender on the log hourly wage\footnote{The hourly wage is the ratio between the previous year's wage and the hours worked last year. For specific details, see Maasoumi2019. The extremely low values of hourly wages are recorded as zeros in the data set. Hence, the missing values of the log of hourly wage variable used in Maasoumi2019 include those that are missing in the hourly wage and also extremely low values.} for each type based on observed and unobserved characteristics. The data is from Maasoumi2019 based on CPS in the United States from 1976 to 2013. For each year, the sample size is different. For instance, in 1976, among those 64,249 records, 34,306 are from females and 29943 are from males. In 2013, there are 61,657 female records, and 57,542 are from males. We focus on not only the full-time workers but also the part-time workers, as Fernandez2023.

Before we construct our bounds, we need to consider the effects of baseline covariates. In Maasoumi2019, those covariates are controls. In our analysis, we follow the method proposed in Lee2009 to bound the ATE conditional on these observable characteristics. Specifically, we use the same group of covariates, their polynomials, and interaction terms as in Maasoumi2019 to construct a proxy for predicted hourly wage for each individual by linear regression estimation. This proxy is calculated for each individual, regardless of their employment status. Based on the proxy, we divide the sample of each year into five groups, i.e., covariate groups. After this division, we calculate the proportion and construct the bounds of ATE for each type in every covariate group. Additionally, we consider the weights of each recorded data point, as in Maasoumi2019 in our calculations for bounds.

In summary, we find that the bounds for Type 1 (EEEE stratum) are narrower than the other types. Also, we see that the LBs and UBs for Type 1 drop rapidly in the early years of the sample (1976–1990s). After that, the bounds sometimes contain 0 and sometimes exclude 0. For Type 2 (EENE stratum), the LBs and UBs of ATE show more and more bounds exclusive to 0. The 95% Confidence Intervals (CIs) are very wide; most of them contain 0. After we employ the mean dominance assumption, the bounds of ATE for Type 1 exclude the 0 through the period. The UBs decreases display the same trend as under the basic model. For Type 4 (EENN stratum), LBs and the LBs of 95% CIs are above 0 in the early years for higher predicted wage groups. For Type 12 (ENNN stratum), like Type 2, the bounds of ATE exclude 0 in the later years, yet the 95% CIs do not.

Results under Basic Model

In this subsection, we present the proportions and the bounds of the ATE for each type under the basic model setting. For the bounds, we show the intersection bounds, CLR bounds, and their corresponding 95% CIs. Intersection bounds refer to the intersection of the bounds of the ATE under $Z = 0$ and $Z = 1$. However, intersection bounds ignore the covariance between the bounds, leading to bias in finite sample settings. Therefore, we apply the method proposed in Chernozhukov2013 to obtain a bias-adjusted version, known as CLR bounds (or bounds under CLR adjustment). Specifically, the CLR bounds here are the half-median unbiased estimates of the lower and upper bounds (Chernozhukov2013).\footnote{It is noteworthy that Xuan Chen generously provided the STATA code from ChenXuan2015, which utilizes the CLR methods from Chernozhukov2013. We express our sincere appreciation for this contribution and have generated an R version of the code based on the method outlined in Chernozhukov2013.} Correspondingly, there are two types of 95% CIs. The first one is from Imbens2004, calculated as $(\widehat{LB} - 1.645 \widehat{SE_{LB}}, \widehat{UB} + 1.645 \widehat{SE_{UB}})$, where the estimators and their Standard Errors are from Theorem (ref). The second one is the CLR 95% CI, from Chernozhukov2013. CLR bounds and CLR 95% CI will be wider than their counterparts: intersection bounds and 95% CI from Imbens2004.

First, we examine the proportions of each type across different years and covariate groups. Figure (ref) illustrates these proportions for each type, with five separate figures provided for each type. Notably, Type 1 (EEEE stratum) consistently exhibits the highest proportions across all years and covariate groups. This is evident from the axis displaying the proportions, where the lowest bar represents a proportion of 0.35 for Type 1, while the highest possible value for other types is 0.3. The figure for Type 1 also shows that there is a big jump from 1976 to 1990; that is, the proportion of always-working individuals (men and women) rises substantially over those years. The magnitude of this increase is most pronounced in covariate group 5 compared to the other four groups. Additionally, when comparing proportions across covariate groups, we notice that as we ascend towards higher predicted income groups, the proportions of consistently working individuals also increase. This is reasonable, as individuals with higher wages tend to engage in employment. Recall that when we calculate the proportions of Type 1 in the whole population, we use the conditional probability $P_r(S =1| D = 0, Z = 1)$. The rise in $P_r(S =1| D = 0, Z = 1)$ shows that more and more mothers with young children choose to work. This is also noted in Blau1998.

However, for the other four types, the trends vary. For Type 2 (EENE stratum) and Type 4 (EENN stratum), there is a noticeable downward trend across the years for most covariate groups. For Type 2, although it is unclear which covariate group dominates the proportion, we observe a significant decrease in the proportions of Type 2 individuals in covariate group 1 over the years, with proportions nearing zero in later years. Recall that the proportion of Type 2 is calculated as the gap between the proportion of women who work and that of women who choose to work with young children at home, that is, the proportion of women who choose not to work with young children at home. This downward trend indicates a decline in the proportion of women who choose not to work with young children at home. Similarly, for Type 4, there is a substantial decrease in proportions for covariate groups 3, 4, and 5 around the 1990s, followed by a slight increase in covariate group 5.

Type 12 (bottom left figure) exhibits higher proportions in later years, similar to Type 1. However, while the proportion of Type 12 (ENNN stratum) individuals increases over time across all covariate groups, it decreases as we progress from covariate group 1 to covariate group 5. Remember, individuals categorized as Type 12 only work if they are male and have young children at home. This disparity across covariate groups indicates a greater prevalence of Type 12 individuals in the lowest predicted income bracket, which is understandable. Lastly, it is worth noting the steady presence of Type 16 in our model, representing individuals who never work. While we report its proportion, we do not provide the bounds of ATE for Type 16.

figure[figure omitted — 151 chars of source]

We delve into the bounds of ATEs for Type 1 and Type 2 across five covariate groups for four years: 1976, 1996, 2011, and 2013 under the basic model, reported in Table (ref), Figure (ref), and Figure (ref) in the Appendix. The choice of 1976 allows us to establish a baseline, while 1996 showcases the pattern in the bounds. For a comprehensive overview of the results spanning 38 years, refer to the tables with complete results in the Supplementary Appendix.

First, let us focus on Type 1. Initially, our attention is drawn to the lower bound (LB) of the intervals for ATEs, as it signifies whether a non-negative gender gap exists. In 1976, the lower bound of CLR on ATE consistently exceeds 0 across all covariate groups, indicating the presence of a gender gap. However, for the first covariate group (the group with predicted lower income levels), the lower bounds are all below 0, implying that 0 is contained within the CIs. Conversely, for all the other four covariate groups, 0 is not included in the CIs. Notably, the trend for Type 1 suggests that the intervals for the gender gap gradually encompass 0 in the later years. Although the trend varies across different covariate groups, the overall direction indicates a gradual inclusion of 0 in the intervals as time progresses. Furthermore, the intervals for the gender gap among individuals in covariate groups 3, 4, and 5 increasingly encompass 0. Notably, covariate groups 3, 4, and 5 represent individuals predicted to have higher income levels. For covariate group 1, the intervals start to include 0 after 1981 and remain inconclusive for most of the sample years. Next, the Upper Bound (UB) for the gender gap in covariate group 1 is consistently the lowest among the five groups. Conversely, covariate group 5 generally exhibits higher UB values. Additionally, UB values decrease gradually over time. This observation aligns with findings from Maasoumi2019, BlauKahn2017, and Goldin2014, indicating that the gender gap has decreased over time, although the decease is not the same across all groups of individuals. For instance, for lower predicted income groups, women do not appear to earn less than men. While in groups with high potential income, the gender gaps are generally larger.

It is essential to note the evolving proportions of individuals in Type 1 over time, ranging from a minimum of 0.356 to a maximum of 0.709. The median proportion stands at 0.596, with the proportion of Type 1 surpassing 0.50 across all five covariate groups after 1985. These relatively high proportions provide us with relatively low standard errors. Thus, there is no big difference between the bounds and CIs. Also, individuals classified as Type 1 always choose to work. In earlier years, these individuals predominantly comprised men who supported the family and women from impoverished backgrounds who married less affluent partners, resulting in substantial gender gaps across the five covariate groups. However, over time, an increasing number of women opt to enter the workforce, leading to a general reduction in the gender gap.

Figure (ref), based on the table on CLR bounds for Type 1 in the Appendix, further supports these findings by incorporating gender gap results from Maasoumi2019 (Table 10 first column: Mean). Notably, the CLR bounds (and the intersection bounds also) contain the mean across five quantiles in Table 10 of Maasoumi2019 after smoothing by using lowess, suggesting alignment without imposing numerous assumptions.

Next, for Type 2, we see that although 1979 marks the first year when the CLR bounds for high potential wage groups no longer encompass 0, as indicated in the Supplementary Appendix, throughout most of the sample period, the 95% CIs for the gender gap remain wide and inclusive of 0. This is primarily due to the comparatively low proportion of Type 2 individuals, particularly in later years. This change in the proportion is shown in Figure (ref). By examining the numbers, the proportion of Type 2 never exceeds 0.100 after 2008. This low proportion contributes to the wide bounds of the ATE, indicating substantial uncertainty, large standard errors, and wide 95% CIs.

The CLR bounds and intersection bounds for Type 2 of potential high-wage groups (covariate groups 4 and 5) exclude 0 for more than 20 years, particularly in 1979, 1981, and subsequent years. Additionally, the LB of the gender gap exceeds 0.5 after 2009 for covariate group 5, suggesting a significant gender gap among individuals predicted to have higher income levels. In other words, men probably perform better than women for those covariate groups in those years. Recall that women of Type 2 will choose not to work when they have young children at home and they will work when they do not, and the covariate groups 4 and 5 contain individuals who are predicted to have higher income levels. This suggests that if women of Type 2, who are predicted to have higher wages, work after their children are older than 5, they will face the gender gap issue. For covariate groups 1, 2, and 3, the intervals consistently include 0 across the years, indicating that women in lower predicted income groups likely do not perform worse than men.

The 95% CIs presented in Table (ref) for Type 2 are wide due to the low proportion, resulting in substantial standard errors. Indeed, with this low proportion, the standard errors of estimates for $E(Y^*(0)|\text{Type 2})$, UB, and LB are large. Consequently, the standard errors of the bounds for ATE are much larger compared to the results for Type 1. As a consequence, the 95% CIs are wide and typically encompass 0. However, there is an exception in 2010. Upon careful examination of tables in the appendix, it is observed that in 2010, the Imbens and Manski 95% CI (the tightest CI) provides bounds with a positive LB. Furthermore, comparing the LBs of CIs across the five covariate groups, it is noted that the LBs for covariate groups 4 and 5 are closer to 0 (less negative). Additionally, upon examining across the years, it is observed that the LBs for covariate groups 5 are gradually moving towards 0 from the negative side.

figure[figure omitted — 239 chars of source]

Next, we compare Type 1 and Type 2 under the basic model. Figure (ref) is drawn to provide a clear depiction of the trend.\footnote{In the tables for Type 2 and covariate group 4, there are 3 big numbers in LB in 1983, 1987, and 2012. This happens because the estimates for $E(Y(0)|Type 2)$ of covariate group 4 in the years 1983, 1987, 2012, and 2013 are below 1. This occurrence is attributed to the very low proportion of Type 2 in the dataset. Across the five covariate groups, such situations only arise in covariate group 4. Therefore, these cases are excluded from our analysis.} This figure presents the bounds of ATEs for Type 1 and Type 2 across years and covariate groups. Based on Table (ref), it is observed that the bounds for the gender gap for Type 1 are much narrower compared to those for Type 2. This discrepancy is reasonable because the proportion of Type 2 is significantly lower than that of Type 1, indicating fewer individuals belong to Type 2. Consequently, the upper bounds of the gender gap for individuals belonging to Type 2 are progressively higher from 1976 to 2013. However, the lower bounds do not decrease notably, especially for covariate group 5. Moreover, the standard errors for Type 2 are higher than those for Type 1, resulting in wider 95% CIs for Type 2 compared to Type 1.

table[table omitted — 4,633 chars of source]

The discussion concerning the ATE for Types 4 and 12 is deferred to a later section. Under the basic model, we rely on theoretical bounds for $E(Y^*(0)| \text{Type j})$ when $j \in \{4, 12\}$. Although the bounds for the ATE for Types 4 and 12 will be considerably wide, they provide valuable information as they are narrower than $[Y^{LB} - Y^{UB}, Y^{UB} - Y^{LB}]$.

Check the Mean Dominance Assumption Using the Data under Basic Model

In this subsection, we compare the estimates of $E(Y^*(0)| \text{Type j})$ when $j \in \{1, 2\}$ to provide a check for the mean dominance assumption for those types. The mean dominance assumption is related to the positive selection assumption. However, the mean dominance assumption compares wages based on types, and positive selection relies on working statuses. For instance, women of Type 2 (EENE stratum) have two possible working statuses: employed and nonemployed. $E(Y^*(0)| \text{Type 1}) > E(Y^*(0)| \text{Type 2})$ implies that wages for Type 1 (EEEE stratum) tend to be higher than those for Type 2 (EENE stratum).

Positive selection, explored in the literature, posits that wages for employed women tend to be higher than those for non-working women. This assumption has been discussed in works such as BlauKahn2006, Blundell2007, Olivetti2008, Mulligan2008, BlauKahn2017, and Maasoumi2019. Blau2023 look into the pattern of selection into the employed status on the gender gap either in the median, mean, or the whole distribution. Their conclusions vary based on methods, including part-time workers or not, the definition of selection, sample data sets, and models. For instance, Mulligan2008 is based on the model and method from Heckman1979, and Maasoumi2019 relies on nonseparate models with the copula method. Blundell2007 assume that for men and women, the wage distribution of nonworkers is first-order stochastically dominated by the distribution of workers. They also provide a weaker version by comparing the median wage of workers and nonworkers. Maasoumi2019 test this positive selection from Blundell2007 using the CPS in the United States. They find that the CDF of the log hourly wage of working women does not always first order stochastic dominate the CDF of not working women. They conclude that for women, the selection changes from negative selection to positive selection. The turning time is the 1990s. For men, Maasoumi2019 conclude that positive selection through the years.

Figure (ref) illustrates the average potential outcomes for women of Type 1 and Type 2 across five covariate groups and years. From the figure, it is evident that for covariate groups 4 and 5, estimates for $E(Y^*(0)| \text{Type 1})$ surpass those of women of Type 2 over the years. Conversely, for covariate groups 1 and 2, the trend is reversed, indicating a notable distinction between Type 1 and Type 2. In covariate group 3, the mean potential outcome of women in Type 1 is initially lower than that of Type 2 for the first two years, but the trend reverses thereafter.

When considering standard errors and 95% confidence intervals, the intervals of the estimates overlap in the early years for covariate groups 4 and 5, while for low potential wage groups, the confidence intervals do not always intersect. Consequently, if we calculate the average across the five covariate groups, the average potential outcomes for women of Type 2 are expected to be higher than those of Type 1 in the early years, as indicated by non-overlapping confidence intervals during the first two years. Regarding selection patterns, negative selection is initially observed, followed by positive selection, as depicted in Table (ref) in Appendix. This is a similar conclusion as in Maasoumi2019 and Mulligan2008. Fernandez2023 analyze the role of selection for CPS data from 1975 to 2020 for all workers using the semiparametric method to estimate functions of distributions based on nonseparate models. They find that there is positive selection, and the selection evolves through time. We consider individuals with full-time jobs and part-time jobs, as in Fernandez2023, and find that the pattern changes in 1979, earlier than Maasoumi2019 observe. Based on the information from Type 1 and Type 2, positive selection predominantly characterizes the pattern, particularly for high-potential outcome groups and later years.

figure[figure omitted — 169 chars of source]

Results under Basic Model with Assumptions (ref) and (ref)

In this subsection, we present the bounds of ATE for all four types under Assumptions (ref)-(ref). With these additional assumptions, we have narrower bounds and 95% CIs compared to those under the basic model alone. We provide results for selected years, with complete outcomes available in the Supplementary Appendix.

Table (ref) reports the results for Type 1 and Type 2 across 5 covariate groups under the basic model and Assumption (ref). Recall that for Types 1 and 2, $E(Y^*(0)| \text{Type j})$ is point identified. Hence, the bounds of ATE for Types 1 and 2 under the basic model and Assumption (ref) are the same as those only under the basic model. With Assumption (ref), the bounds for these two types are narrower than before. We still select the same years to make a comparison with Table (ref).

First, let us focus on Type 1 in Table (ref). We see that the LBs for Type 1 are higher with Assumption (ref). That is, every interval is above 0. The LB and UB are decreasing fast from 1976 to the 1990s. After that, it becomes slower. In some years, the LB and UB even increased. Although there are some differences across the covariate groups, the overall trend stays the same. Among those covariate groups, the highest LBs and UBs happen mostly in covariate group 5. This shows that in the potential high-income group, women who always work probably have relatively poor performance in the labour market compared with men.

Next, Table (ref) shows that for Type 2, men are probably performing better than women in the later years of the sample for people who are predicted to have higher income. The UBs are lower than those under the basic model. However, UB fluctuates. This may be due to the low proportion of individuals who belong to Type 2. For covariate groups 1 and 2, we observed negative UBs for some years. This shows that women probably perform better than men in the group where people are predicted to have lower income levels. However, if we check the 95% confidence intervals we see that both those with and those without adjustment using the CLR method contain 0.

table[table omitted — 4,442 chars of source]
table[table omitted — 4,489 chars of source]

Table (ref) reports the bounds for Type 4 and Type 12 under Assumptions (ref)-(ref). That is, to have the tightest bounds for Type 4 and Type 12, we need both (ref) and (ref). This is because Assumption (ref) provides the upper bounds of $E(Y^*(0)| \text{Type j})$ when $j \in \{4, 12\}$. Assumption (ref), on the other hand, helps shrink the bounds for $E(Y^*(1)| \text{Type j})$ when $j \in \{4, 12\}$. In the Supplementary Appendix, we report the additional results for Type 4 and Type 12 under the basic model with only Assumption (ref). The bounds are wider than the bounds we report and discuss in this subsection. Specifically, the UBs tend to be higher. Just as in Table (ref), we only report the results for several years to show the pattern.

Table (ref) presents the same measures and categories for Type 4 across 5 covariate groups and 4 selected years. In 1976, CLR bounds and CLR 95% CI do not contain 0 when we look at the higher potential income group: covariate group 4. In fact, in the beginning of the sample period (1976--1981), it is quite clear to see that the CLR bound and CLR 95% CI of ATE do not contain 0 for individuals in covariate group 4 if we check the tables in the Supplementary Appendix. That is, we see that in 1976--1981, the gender gap existed in the predicted high-income group (covariate group 4), and the LBs were closer to 0 from the left-hand side (covariate group 5). Recall that women of Type 4 will not work even if they do not have young children at home, and men, on the other hand, will work regardless of whether they have young children or not. In the early years, well-educated women or potential high-wage women may choose to stay at home after they are married. As noted in Neal2004, a high proportion of women in earlier years married to men with high wages and chose to stay at home. Even if they choose to work at that time, their performance in the labour market will probably not be good compared to men. This clear message benefits from the fact that the proportion of Type 4 in that covariate group was the highest at the beginning of the sample period. This is shown in Figure (ref). From 1976 to 1995, 0 was not in the CLR bounds for the covariate group 4 for 12 years (excluding 1983 and 1987). From 1996 to 2008, CLR bounds and intersection bounds sometimes included 0 in covariate groups 3, 4, and 5. After 2009, bounds with or without CLR adjustment do not contain 0 for covariate group 5. These results illustrate that in those periods, women of Type 4 probably had a lower hourly wage than men of the same type. Most of the 95% confidence intervals include 0 in the later years, while only 95% CIs from Imbens2004 do not include 0 after 2010.

Comparing the results for Types 4 and 12, we notice the big difference between the trends of these types. For Type 12, in 1976, the bounds for gender gap included 0 across 4 intervals and 5 covariate groups until 1984 (CLR bounds) or 1982 (intersection bounds). After that, in some years, the lower bounds become positive for covariate group 4. From 2009 onwards, for covariate group 5, the bounds for the ATE and the Imbens2004 95% CIs exclude 0.

Conclusion

This paper addresses a crucial aspect of sample selection within the context of the gender gap problem. In this context, even when treatment assignment is random, the presence of selection bias distorts the treatment's effect on the outcome variable due to unobserved factors. Consequently, bounding the treatment effect without making restrictive distributional or model specification assumptions offers an alternative more robust method. Unlike standard sample selection models, bounding the effect does not necessitate imposing an exclusion restriction. However, employing an exclusion restriction is a common practice in gender gap investigations, with its validity tested in various studies. Utilizing the additional information provided by this exclusion restriction enables the segmentation of the population into distinct subgroups or types, facilitating the derivation of narrower bounds for each type.

Examining the gender gap within the sample selection model and the potential outcome framework allows us to identify the proportions of various types and bound the gender gap for each type accordingly. Notably, we observe distinct trends in the proportion of each type and the corresponding gender gap bounds. For instance, we observe an increasing trend in the proportion of individuals who always work over time, along with varying trends in the bounds for this type. Specifically, we note two key trends. First, the upper and lower bounds of the gender gap tend to shrink over time, albeit at different rates. Prior to the 1990s, the decline was rapid, whereas thereafter, it slowed down. Secondly, we find that the gender gap upper and lower bounds are consistently highest among individuals in the high-potential wage group. With the addition of further assumptions to our basic model, we find that the gender gap bounds for the always-working subgroup show that the gender gap existed in the sample period for this type of individual. However, for other types of individuals, the trend in the gender gap estimates differs or remains ambiguous when examining the 95% CIs. This underscores the significance of disaggregating individuals by type, as it reveals that the gender gap varies across different potential wage levels and types of worker.