Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
70,756 characters · 20 sections · 51 citation commands
Identification of Child Penalties
The transition to parenthood is central to the onset of gender inequality in labor markets goldin2024nobel. An important and growing body of research quantifies the gender gap in the effect of parenthood on labor market outcomes using event‐study designs that normalize estimated effects post treatment kleven2019children. Recent critiques discuss how event-study based estimates may be biased melentyeva2023child, building on the Difference-in-Differences (DID) literature on biases in two‐way fixed‐effects regressions with staggered treatment timing goodman2021difference. While these critiques focus on biases in estimation, less attention has been paid to the identification framework underlying event studies that rely on post‐treatment normalization.
This paper formalizes the identification framework underlying normalized event studies, which I term Normalized Triple Differences (NTD). I reverse-engineer NTD from the validation tests commonly used in applied work. I show that under NTD, the conventional estimator does not identify its target estimand, the gender gap in normalized effects, when the parallel trends assumption from DID is violated. As a solution, I propose targeting the effect of parenthood on the gender earnings ratio and show this estimand is identified under NTD. The analysis focuses on labor market earnings and is illustrated empirically using Israeli administrative data.
I begin by describing the normalized event-study empirical strategy, which proceeds in three steps. First, for each individual, the event is defined as the age at first childbirth, and separate regressions are estimated by gender, regressing the outcome on a set of event time indicators and fixed effects for age and calendar year. Second, predicted outcomes are computed using only the estimated fixed effects, often interpreted as the expected earnings absent treatment. Third, the estimated event-time coefficients are normalized by the mean of the predicted counterfactual outcomes. These normalized estimates capture the gender-specific percentage effects of parenthood. The post-treatment gender gap in these normalized estimates is often labeled the child penalty.
To articulate the identification assumptions underlying the normalized event study approach, I reverse-engineer the assumptions validated by the empirical test: that the pre-childbirth gender gap in normalized estimated effects equals zero kleven2019children,andresen2022causes. I first define the potential outcomes framework and relevant causal and descriptive estimands, then rewrite the empirical validation test in a $2\times2$, as in this simplified setting the assumptions become transparent. Under no anticipation, this test validates that parallel trend violations are equal across genders after normalizing by counterfactual earnings. I term this identification framework Normalized Triple Differences (NTD) and view it as the framework underlying normalized event studies.
I then establish the paper's first main result: in a $2\times2$ where NTD holds, the descriptive gender gap in normalized DID does not identify the causal gender gap in normalized effects when the parallel trends assumption is violated—that is, when counterfactual earnings trajectories differ across treatment and control groups. The direction of bias is systematic: if the control group has steeper counterfactual earnings growth than the treated group, the descriptive gender gap understates the true child penalty. As the normalized event study estimator can be viewed as aggregating multiple $2\times2$s, the result shows it is biased for its target causal estimand even when NTD holds.
Given this result, I examine whether parallel trends is likely to hold in the child penalty context, drawing on insights from canonical models of fertility and human capital ben1967production,becker1990human. In these models, individuals with higher labor-market ability delay childbirth and invest more in human capital. This generates selection on treatment timing: later treatment groups have steeper counterfactual earnings trajectories than earlier groups. When such later treatment groups serve as controls for earlier treatment groups, the parallel trends assumption is violated. As discussed above, under NTD this direction of violation causes the descriptive gender gap to understate the true child penalty. Using Israeli administrative data, I document evidence consistent with this selection mechanism: parents who delay childbirth come from higher-income, more-educated families and score higher on national mathematics exams.
I conclude the discussion on bias of the conventional estimator with a bias-bounding exercise. By imposing an identification assumption on fathers' average treatment effects, I derive a bias-correction formula for mothers' effects under NTD. Using this as a diagnostic tool, I apply a plausible range for fathers' normalized effects (-10% to +10%) to provide suggestive evidence on the magnitude of bias in the conventional estimator. Applying this to the data suggests the bias is heterogeneous across treatment groups. For earlier treatment groups, specifically parents that had their first child at ages 24-26, the exercise suggests substantial bias. For example, for the treatment group 26, bias-corrected estimates are 24–51% more negative five years post-childbirth relative to conventional estimates. This result suggests the conventional estimator understates child penalties in earlier treatment groups, in line with the above discussion. For later-treated parents, the exercise does not provide evidence of bias: starting at age 28, conventional estimates fall within the bounds of the bias-corrected estimates.
Having discussed that parallel trends violations bias the conventional estimator under NTD, I conduct pre-treatment validation tests to assess the parallel trends assumption empirically. Similar tests also validate which treatment groups satisfy the NTD assumptions, informing the alternative identification approach I propose next. I argue tests should be conducted at the $2\times2$ level: since the control group changes with post-treatment event time, each treatment group requires multiple validation tests—one for each control group. This differs from standard practice in the literature, which aggregates across treatment and control groups. Using Israeli data, I document two key patterns. Within gender, pre-treatment DID estimates increase with the gap in treatment timing between treated and control groups. However, pre-treatment NTD estimates are not systematically different from zero for treatment groups aged 26-30. I therefore focus on these treatment groups when estimating post-treatment effects below.
Having established that NTD does not identify the conventional target estimand, I turn to the second main result of the paper. I propose targeting the effect of parenthood on the gender earnings ratio and show it is point identified under NTD. This estimand captures how parenthood affects gender inequality, making it a natural target for the child penalty literature. I develop a corresponding estimator and derive cluster-robust standard errors based on its influence function. Applying both estimators to Israeli data reveals heterogeneity across treatment groups. For earlier-treated parents ($D=26,27$), the two estimators yield similar results. For later-treated groups ($D=29,30$), the new estimator is smaller in magnitude. For example, for treatment group $D=30$ five years post-childbirth, the new estimator is 28% smaller than the conventional estimator. Since the bias-bounding exercise found no evidence of substantial bias for later groups, this difference likely reflects different target estimands rather than bias. The new estimand also enables a direct decomposition of the gender gap: for $D=30$ five years post-childbirth, approximately 34% of observed gender earnings inequality is attributable to parenthood. Finally, I discuss aggregation across treatment groups, noting that differences in treatment distributions can complicate comparisons of aggregates, such as across countries.
This paper is most closely related to the literature that estimates the long-run effect of becoming a parent on gender inequality in the labor market, the so-called child penalty, using administrative data and event‐study designs. bertrand2010dynamics provide an early and highly influential analysis of earnings trajectories among MBAs, documenting the emergence of gender gaps in earnings and labor supply around the first childbirth. angelov2016parenthood extend this approach by using population‐wide data and longer event horizons in a within-couple event study specification, documenting that the effects of parenthood on gender inequality persist for more than a decade after childbirth. kleven2019children further advance the literature by introducing an approach that normalizes event study coefficients to adjust for gender differences unrelated to childbirth. A growing body of subsequent work builds on kleven2019children to study the child penalty kleven2021does,kleven2022geography,andresen2022causes,cortes2023children,kleven2025child,kleven2024family,berniell2021gender,hotz2018parenthood,sieppi2019parenthood.
I contribute to this literature by formalizing the NTD identification framework that the normalized event study approach relies on, establishing both a non-identification result for the gender gap in normalized effects and a new identification result for the effect of parenthood on gender inequality. Beyond the child-penalty context, this new identification framework may be useful in other settings for identifying the effects of a treatment on group inequality, particularly when there are many zeros in the data and hence using logarithms is not plausible. To facilitate replication and application of the proposed estimators, I developed an open-source R package, \href{https://dorleventer.github.io/childpen}{childpen}, which implements the discussed estimators.
This paper also relates to studies that use contemporary DID estimation methods. As examples, in the child penalty context, melentyeva2023child implement a stacked DID design cengiz2019effect,wing2024stacked, lin2025long,fajardo2024there use the estimator of sun2021estimating, and bearth2024beyond adopt the approach of callaway2021difference. While these methods correct estimation bias stemming from staggered treatment timing, they do not address identification bias due to violations of the parallel trends assumption. I discuss such violations when the outcome is labor market earnings, drawing on insights from human capital theory that provide a mechanism by which later-treated control groups have steeper counterfactual earnings trajectories.
This paper also connects to the broader DID methodological literature on parallel trends testing roth2022pretest,ghanem2022selection,roth2023parallel and DID as an aggregate of multiple $2\times2$s callaway2021difference,goodman2021difference. On validation, I argue that pre-trend tests must be performed separately for each treatment–control pair when the control group changes in post-treatment event time, as in not-yet-treated DID designs. On aggregation, I discuss how variation in treatment distribution affects the interpretation of aggregate comparisons. When comparing effects across subpopulations, each subpopulation's estimate is itself an aggregate over multiple treatment groups; if treatment distributions differ, so may the comparisons. In the child penalty context, this is relevant for studies comparing penalties across countries kleven2019child or parent types andresen2022causes.
The rest of the paper is organized as follows. Section (ref) presents the normalized event‐study empirical strategy and the data used for illustrating the theory. Section (ref) presents NTD as the identification framework underlying normalized event studies. Section (ref) establishes that the conventional estimator is biased under NTD. Section (ref) discusses validation tests. Section (ref) presents the new estimator that is unbiased under NTD. Section (ref) concludes.
The child‐penalty literature estimates the gender gap in normalized effects of parenthood on labor‐market outcomes using normalized event‐study designs and administrative earnings data. This section first formalizes normalized event‐studies to clarify the empirical approach which motivates the identification analysis, and then describes the Israeli administrative data used to illustrate the theory throughout the paper.
This section briefly presents the normalized event‐study empirical strategy commonly used in child‐penalty applications kleven2019child,kleven2019children,kleven2022geography,andresen2022causes,kleven2021does,de2022differential.
I briefly present notation needed to formulate the normalized event study approach. Consider a finite population of individuals $i \in \{1, \ldots, n\}$ observed over a finite set of time periods $t \in \{1, \ldots, T\}$. Let $Y_{i,a}$ denote real annual labor‐market earnings of individual $i$ at age $a$, and let $D_i$ denote the age at which individual $i$ has their first child. Define the event time as $E_{i,a} = a - D_i$. Let $G_i \in \{f, m\}$ indicate gender, where $f$ represents female and $m$ male.
The estimation algorithm for the normalized event‐study proceeds in three steps. First, outcomes are regressed on event‐time indicators and fixed effects separately for each gender. The regression model for gender $g \in \{f, m\}$ is
where $1\{\cdot\}$ denotes an indicator function, $\alpha_a^g$ and $\alpha_t^g$ are age and year fixed effects, respectively, and the superscript $g$ indexes gender‐specific coefficients. In the second step, predicted earnings net of event-time coefficients are computed as $\widetilde{Y}_{i,a} = \widehat{\alpha}_{a}^{G_i} + \widehat{\alpha}_{t}^{G_i}$. Finally, the estimated event‐time coefficients $\widehat{\beta}_e^g$ are normalized by the conditional mean of $\widetilde{Y}_{i,a}$ within gender and event time:
where $\mathbb{E}_n[\cdot]$ denotes the sample mean.
Recent work has shown that two-way fixed effects regressions similar to (ref) can produce biased estimates in the presence of multiple treatment groups sun2021estimating,goodman2021difference,de2020two,borusyak2024revisiting. As melentyeva2023child argue, this concern applies to the child penalty setting as well. Our focus is not on biases from the estimation procedure itself, but rather on biases arising from the underlying identification assumptions. In the section below I turn to articulating the identification framework that is rationalized from the above normalized event study empirical strategy.
This section introduces the data used to illustrate the theoretical discussion. Although the arguments developed below are theoretical and generalizable, I illustrate their implications using a specific application to aid with constructing intuition, empirically assess key claims, and document new empirical insights. To that end, I describe the Israeli administrative data used throughout the paper, including its sources, variables, and sample definitions.
The raw dataset covers all Israeli citizens born between 1970 and 2000, matched to their spouses, parents, and children. It was compiled by the Israeli Central Bureau of Statistics (CBS) and integrates data from several administrative sources, including the Population and Immigration Authority’s Civil Registry, the Ministry of Education, the Council for Higher Education, and the Israeli Income Tax Authority.
\noindent1.2.1. Main Variables.
This subsection describes the main variables used in the analysis.
Treatment: Age at birth of first child. Each individual is linked to their biological children. The year of birth of the earliest child is used to define the year of first childbirth. Subtracting the parent’s year of birth yields their age at first birth.
Outcome: Earnings. Annual labor market earnings are observed from 1999 to 2020, based on micro-level tax records. Earnings are coded as zero in years with no reported income. All values are expressed in real 2020 New Israeli Shekels (NIS), using the CBS consumer price index.
The analysis below makes use of several additional variables, defined explicitly in Appendix (ref). These include grandparents' earnings, nationally administrated mathematics test scores called Meitsav, number of credits in high-school subjects, university psychometric entrance test (UPET) scores, years of education and highest education degree. The sample definition also uses ethnicity and religion variables.
\noindent1.2.2. Sample Definitions.
I make the following restrictions on the main sample. First, the analysis focuses on non-Haredi Jews; individuals identified as Arab or Ultra-Orthodox (Haredi) Jews are excluded, due to systematically different fertility and labor market trajectories yakin2021,gould2024child. I further restrict the sample to individuals born between 1975 and 1990. Older cohorts are observed only at advanced ages, while younger cohorts are only partially observed through their late twenties. When adding controls, we limit to birth cohorts 1980 and older, reflecting data availability for education variables.
Furthermore, I drop individual-year observations where the individual is less than 20 years old, corresponding to the typical entry into the labor market after high school completion and mandatory army service.\footnote{The legal minimum working age in Israel is 15. However, most individuals complete high school at 18, followed by mandatory army service—two years for women and three years for men. Some individuals, such as religious women, may be exempt from military service but instead perform national service (e.g., in schools or hospitals).} I also drop parents who had their first child at ages prior 24 or post 40. Births before age 24 would imply pre-trend diagnostics occur before age 20, where few observations exist, while births after age 40 involve very small sample sizes. Additionally, I keep years only in time window where individuals are reported as alive by the Civil Registry. The dataset used in the main analysis, after the above limitations, consists of 13.7 million individual-year observations, made up of 374 thousand mothers and 320 thousand fathers. Further construction details are provided in Appendix (ref).
\noindent1.2.3. Sample Statistics.
Appendix Figure (ref) reports event studies, specifically estimates of (ref), using the Israeli data. Consistent with findings in other countries, the normalized effects for mothers and fathers, $\widehat{\theta}_{\mathrm{ES}}(f,e)$ and $\widehat{\theta}_{\mathrm{ES}}(m,e)$, are very similar before childbirth. After childbirth, a gender gap emerges, with mothers’ normalized effects more negative than fathers’. \footnote{The Israeli results display two patterns that differ somewhat from what is typically observed in other countries. First, the pre-trends for both mothers and fathers are not flat but negative; however, they remain parallel, consistent with the main interpretation of event studies. Second, fathers’ normalized effect $\widehat{\theta}_{\mathrm{ES}}(m,e)$ declines over time after childbirth. Despite these differences, the magnitude and persistence of the gender gap are broadly in line with findings from other settings. Finally, my estimates are similar to other literature on child penalty in Israel, e.g., gould2024child.} While these patterns mirror the main qualitative features found in other studies, interpreting them causally requires caution, as discussed below.
In this section, I formalize the identification framework underlying normalized event studies. I reverse-engineer this framework from the validation test—that pre-childbirth gender gaps in normalized effects equal zero—and term it Normalized Triple Differences (NTD). Under no anticipation, NTD validates that parallel trends violations are equal across genders after normalizing by counterfactual earnings. I begin by defining the potential outcomes framework and the relevant causal and descriptive estimands, and then present the result.
Throughout, for a given gender $g$ I will focus on $2\times2$ comparisons: treatment group $d$, control group $d^\prime$, target age $a$ and pre-treatment age $d-1$. Following melentyeva2023child, I set the control group to be the closest-not-yet-treated treatment group, i.e., $d^\prime=a + 1$. For example, if $a = 30$ then $d^\prime = 31$, if $a = 31$ then $d^\prime = 32$, and so on. I discuss aggregation across treatment groups in Appendix (ref).
Let $W_{i,a} = 1_{\{a \geq D_i\}}$ denote the treatment status of individual $i$ at age $a$, where $D_i$ is the age at first childbirth. This definition implies that child penalties are a staggered adoption design, i.e., $W_{i,a-1}=1\rightarrow W_{i,a}=1$. In this design, each group of parents who experience first birth at a given age $D_i=d$ is treated as a distinct treatment group, untreated before $d$ and treated from age $d$ and onward.
Under the stable unit treatment value assumption (SUTVA) rubin1980randomization, in a staggered adoption design potential outcomes are a function of the timing of treatment callaway2021difference.\footnote{This requires that “age at first birth” satisfies the assumption known as treatment variation irrelevance vanderweele2009concerning, one of the two elements in SUTVA. For example, for a mother who gave birth to her first child at age 25, a counterfactual scenario in which she delays childbirth to age 35 could arise through many distinct causal mechanisms, such as divorce, health issues, or career disruptions. If these different versions of the treatment yield different counterfactual earnings, the potential outcome $Y_{i,a}(d')$ is ill-defined. Since this issue is beyond the scope of the current paper, we abstract from it and assume well-defined potential outcomes throughout.} Formally, let $Y_{i,a}(d)$ be the potential outcome of individual $i$ at age $a$ if the first childbirth occurs at age $d$. Let $Y_{i,a}(\infty)$ denote the potential outcome if $i$ never has a child. Observed outcomes are linked to potential outcomes by the consistency assumption: $Y_{i,a} = Y_{i,a}(D_i)$.
This subsection defines the causal estimands of interest in the child penalty context. I begin with estimands for a single treatment group and a specific gender, and then consider estimands of differences between genders. In Section (ref) I discuss causal estimands which aggregate multiple treatment groups.
Let $APO(g, d, d', a) = \mathbb{E}[ Y_a(d') \mid G = g, D = d]$, denote the average potential outcome (APO) at age $a$ for individuals of gender $g$ who had their first child at age $d$, had they instead had their first child at age $d'$. Next, define the average treatment effect (ATE) for gender $g$, treatment group $d$ at age $a$ as $ATE(g, d, a) = APO(g, d, d, a) - APO(g, d, \infty, a)$, where $APO(g, d, \infty, a)$ denotes the average potential outcome for individuals from treatment group $d$ and gender $g$ in the counterfactual of never having children. To mirror the event study estimator $\widehat{\theta}_{\mathrm{ES}}(g, e)$ in (ref), we define the causal estimand
$\theta(g, d, a)$ captures the proportional earnings loss from childbirth at age $a$, relative to the counterfactual of never giving birth.\footnote{The normalized average treatment effect is similar to estimands studied in the vaccine efficacy literature orenstein1985field, and is also related to target estimands in the excess mortality literature msemburi2023estimates.} This provides a normalized measure of the effect of parenthood on earnings age $a$ for individuals of gender $g$ from treatment group $d$.\footnote{This interpretation aligns with the estimand targeted in child penalty studies, i.e., $\widehat{\theta}_{\mathrm{ES}}(g,e)$ in (ref). For example, kleven2019children describe their child penalty estimator ($P^g_t$ in their notation) as “the year-$t$ effect of children as a percentage of the counterfactual outcome absent children,” where $t$ corresponds to event time $e$ in my notation. Similar quotes can be found in other papers that use the normalized event study strategy.}
The term “child penalty” is frequently used to describe the differential impact of parenthood on labor market outcomes between women and men. Several causal estimands can capture this gender gap. A natural starting point is $ATE(f,d,a) - ATE(m,d,a)$, which reflects the level difference in the impact of parenthood on earnings between females and males. However, since earnings levels may differ by gender even in the absence of children, researchers may prefer a normalized comparison. One such alternative is $\theta(f,d,a) - \theta(m,d,a)$, which captures the gender gap in relative earnings losses from parenthood, that is, the gender gap in normalized effects. Since the normalized event study approach (Section (ref)) compares $\widehat{\theta}_{\mathrm{ES}}(g,e)$ in (ref) across gender, $\theta(f,d,a) - \theta(m,d,a)$ represents the causal estimand implicitly targeted in that empirical strategy.
By a descriptive estimand, I mean a population expectation defined solely in terms of observed outcomes and covariates, without involving potential outcomes beyond the realized outcome $Y$ abadie2020sampling. The following three descriptive estimands—used below to construct validation tests and identify causal estimands—can be thought of as $2\times2$ "p-lim"s of DID estimators for the counterfactual APO, ATE and $\theta$.
In $\delta_{\mathrm{APO}}$, the first term provides the pre-treatment level from the treated group, and the second term adds the trend from the control group. Hence, $\delta_{\mathrm{APO}}$ is how DID constructs the counterfactual APO for the treatment group. $\delta_{\mathrm{ATE}}$ is the conventional DID four expectations and three differences estimand. In our context, it is how DID constructs the effect of parenthood on earnings in levels. $\delta_{\theta}$, which I term "normalized DID", is the ratio of these two, and hence how DID can be used to construct the normalized effect.
I now turn to derive the identification assumption that is implicitly maintained in normalized event studies, as inferred from the empirical validation test commonly used in applied work. Before presenting the result, I introduce the no anticipation and parallel trends identification assumptions.
The no anticipation assumption requires that potential outcomes before childbirth are the same under the observed treatment path and the counterfactual of never giving birth abbring2003nonparametric.\footnote{If outcomes at age $d - 1$ are affected by anticipatory behavior, the no anticipation assumption can instead be imposed at earlier ages (e.g., $d - 2$ or $d - 3$), shifting the baseline period for the treated group. Such concerns may also motivate shifting the closest-not-yet control group may to later ages (e.g., $a+2$ or $a+3$).} Formally,
Next, define the difference in counterfactual earnings trends from age $d - 1$ to $a$ between treatment group $d$ and control group $d'$ as:
The statement $\gamma_{\mathrm{PT}}(g, d, d', a)=0$, i.e., that counterfactual trends are equal across treatment and control, is often called the parallel trends identification assumption in the DID framework. Formally, it can be written as
I now turn to deriving the identification assumptions underlying normalized event studies. As stated above, my approach is to find the assumptions validated by the validation test: whether prior to childbirth the estimated gender gap in normalized effects is zero, where estimates are obtained by gender using (ref). Formally,
The question is what identifying assumptions the test (ref) is validating. To backward-engineer these assumptions, I re-write (ref) in a $2\times2$ comparison and replace estimators with descriptive estimands. The $2\times2$ equivalent of (ref) can be written as follows: For treatment group $d$, later-treated control group $d^\prime$, and pre-treatment age $a < d-1$,
The test in (ref) is similar to (ref) in that it requires that in prior childbirth, the gender gap in normalized DID is zero. The following result shows what restrication on potential outcome if the test (ref) and Assumption (ref) hold in the data.
The proof is presented in Appendix (ref). Proposition (ref) shows an iff statement regarding the restriction on potential outcomes that is implied by the test in (ref) if Assumption (ref) holds. I now state this restriction formally as a new identification assumption:
Assumption (ref) states that the violations of parallel trends, once normalized by the counterfactual APO, are equal across genders. Given Proposition (ref), Assumption (ref) can be interpreted as the implicit identification assumption underlying the normalized event-studies empirical strategy (Section (ref)). Going forward, I refer to the identification framework based on this assumption as the Normalized Triple Differences (NTD) framework.
This section establishes the paper's first main result: under NTD, the conventional estimator does not identify its target causal estimand when the parallel trends assumption is violated. I first derive the bias characterization (Section (ref)), then argue parallel trends violations are likely in the child penalty context due to selection on human capital (Section (ref)), and conclude with a bias-bounding exercise suggesting substantial understatement for early treatment groups (Section (ref)).
The following result characterizes the bias under NTD for post-treatment ages for the descriptive estimand, the gender gap in normalized DID.
Theorem (ref) shows the descriptive gender gap in normalized DID does not identify the gender gap in normalized effects under NTD when the parallel trends assumption (Assumption (ref)) is violated. \footnote{Lemma (ref) in Appendix (ref) shows that $P(d,a)=[ATE(f,d,a)-ATE(m,d,a)]/APO(f,d,\infty,a)$, the $2\times2$ analogue of the estimator that kleven2019children term the child penalty ($P_t$ in their notation), is also not identifiable under NTD when parallel trends in levels fail. I focus instead on the gender gap in normalized effects, $\theta(f,d,a)-\theta(m,d,a)$, since most papers using the normalized event‐study approach in Section (ref) discuss the gender gap in $\widehat{\theta}_{\mathrm{ES}}$ from (ref) as their main result.} I return to the validity of Assumption (ref) in Section (ref). Since this descriptive estimand is the probability limit of the normalized event-study estimator in a single $2\times2$, and event studies aggregate multiple $2\times2$s, Theorem (ref) implies the conventional estimator is biased for its target causal estimand.\footnote{As the bias arises from the normalizing factor, a natural alternative is to use logarithms. Appendix (ref) examines whether a log specification resolves the identification challenges. For DID, the answer is likely negative; combining logs with Triple Differences (TD) may be more promising, though this introduces selection due to zeros in the data.}
A sketch of the proof is as follows. The multiplicative bias arises from normalization by counterfactual earnings, which introduces bias into the denominator that does not cancel when differencing by gender. More formally, recall that $\delta_\theta$ divides $\delta_\mathrm{ATE}$ by $\delta_\mathrm{APO}$ (Section (ref)). By Lemma (ref), $\delta_\mathrm{ATE}$ equals the true ATE plus the parallel trends violation $\gamma_\mathrm{PT}$. Under NTD, the ratio $\gamma_\mathrm{PT}/\delta_\mathrm{APO}$ is equal across genders, so this component cancels when differencing $\delta_\theta$ by gender (see proof in Appendix (ref)). However, bias remains because the denominator $\delta_\mathrm{APO}$ itself contains $\gamma_\mathrm{PT}$. Since $\delta_\mathrm{APO}$ equals the true counterfactual APO plus $\gamma_\mathrm{PT}$, the ratio $ATE/\delta_\mathrm{APO}$ is biased even after gender differencing.
For intuition on the result it is potentially helpful to compare NTD to triple differences (TD). TD makes an additive assumption: that parallel trends violations are equal by gender, $\gamma_\mathrm{PT}(f,d,d^\prime,a)=\gamma_\mathrm{PT}(m,d,d^\prime,a)$. This assumption allows identification of an additive cross-gender estimand, $ATE(f,d,a)-ATE(m,d,a)$. In contrast, NTD makes an assumption on ratios (Assumption (ref)). This does not allow identification for the estimand that takes differences across genders, $\theta(f,d,a)-\theta(m,d,a)$, while identification is restored for an estimand that uses ratios , as shown in Section (ref).
To summarize, even when NTD holds, the conventional normalized event study estimator does not identify its target causal estimand when parallel trends is violated. This motivates examining whether parallel trends is likely to hold in the child penalty context.
I now argue the parallel trends assumption (Assumption (ref)) is unlikely to hold in certain comparisons due to human capital differences across treatment groups. I first present the theoretical argument, then provide supporting empirical evidence.
\noindent3.2.1 Theoretical Argument.
A large body of theoretical and empirical work argues that earnings trajectories reflect underlying differences in human and social capital ben1967production, heckman1976life, heckman2014economics, cunha2007technology, and that the timing of childbirth responds to these same factors becker1990human, de2003inequality, geronimus1992socioeconomic. These two strands combine in life-cycle models of fertility and labor moffitt1984profiles, blackburn1993fertility, adda2017career, francesconi2002joint, keane2010role, jakobsen2022fertility, eckstein2019career.\footnote{adda2017career emphasize that family preferences, rather than initial ability differences, drive occupational sorting differences between early and late fertility mothers. Yet, because fertility preferences shape early-life career and education choices, observed treatment groups will still diverge in their earnings trajectories even in the counterfactual where they do not ultimately have children. Thus, a positive correlation between delayed fertility and human capital investment can emerge through family preferences rather than innate ability.}
Applied to the child penalty context, such models have straightforward implications for parallel trends. First, individuals with higher labor-market ability invest more in human capital and delay childbirth. Second, higher-ability individuals have steeper counterfactual earnings trajectories, particularly early in their careers (ages 25–35). These two facts generate parallel trends violations. For early treatment groups (e.g., $D=25$), the relevant post-treatment ages are the late 20s and early 30s. At these ages, not-yet-treated control groups (e.g., $D=29$–31) consist of higher-ability individuals who, even absent children, would be on steeper earnings trajectories. Consequently, the parallel trends violation is negative: $\gamma_\mathrm{PT}(g,d,d',a) < 0$. In these cases, Theorem (ref) shows the descriptive gender gap will understate child penalties.\footnote{A natural response to selection on treatment timing is to condition on observables. Appendix (ref) discusses identification and estimation of child penalties under DID and TD frameworks that condition on covariates.}
\noindent3.2.2 Empirical Evidence.
Figure (ref) plots several early-life indicators of human and social capital against parents' age at first birth, separately by gender.\footnote{Variable definitions and sample restrictions are detailed in Appendices (ref) and (ref), respectively.} The figure documents that in Israel parents who delay childbirth to around age 30 come from higher-earning and more-educated families, score higher on national mathematics exams, and are more likely to take advanced academic tracks in high school.\footnote{As both gender exhibit selection on treatment, one might consider triple-differencing across genders. Appendix (ref) discuss TD in the child penalty context, and shows how it is related to the counterfactual gender earnings ratio discussed in Section (ref).}
Similar patterns appear in other contexts and when considering human capital later in the life-cycle. melentyeva2023child document a positive correlation between age at first childbirth and grandparents' education in Germany. jensen2024birth find similar results for Denmark using the parents' final educational attainment. Appendix Figure (ref) shows comparable patterns for the USA, while Appendix Figure (ref) documents similar selection for Israel across multiple human-capital outcomes, including both final educational attainment and other related measures.\footnote{These findings relate to the literature on education's causal effect on fertility timing black2008staying,mccrary2011effect.}
Taken together, the evidence indicates selection on treatment in the context of child penalties.\footnote{Prior studies document a link between delayed childbearing and human capital observables buckles2008understanding, but rely on post-treatment variables such as career, earnings or final education. In contrast, I use early-life measures, minimizing concerns about reverse causality.} The selection pattern that emerges from the data shows slopes are largest between ages 20–30, then flatten or decline beyond age 30. Selection patterns for fathers peak slightly later than for mothers, consistent with assortative matching and the average two-year spousal age gap. These patterns support the theoretical prediction that parallel trends violations are likely for comparisons between earlier treatment groups (24-26) and later treatment groups (28-32).
Having discussed that parallel trends violations are likely for certain comparisons, I now turn to quantifying the resulting bias in the conventional estimator.
I now quantify the bias in the conventional estimator using a bias-correction approach that assumes fathers' effects are known. Under a plausible range of fathers' effects, empirical results suggest substantial bias for earlier treatment groups, with the conventional estimator understating child penalties. For later treatment groups, the exercise provides no evidence of bias.
\noindent3.3.1. Theory.
I begin by outlining intuitively how an assumption on fathers' effects enables bias-correction to recover mothers' effects. First, if fathers' normalized effects are known, this identifies their counterfactual APO. Second, this in turn identifies the parallel trends violation $\gamma_\mathrm{PT}$ for fathers. Third, under NTD (Assumption (ref)), the ratio of counterfactual APO to parallel trends violation is equal across genders. Finally, one can use the identified ratio for fathers to adjust the biased DID estimator ($\delta_{\mathrm{APO}}$) of mothers to recover their true counterfactual APO.
Since observed earnings identify APO under realized treatment, an assumption on fathers' normalized effects $\theta(m,d,a)$ is equivalent to an assumption on their counterfactual APO. The following result formalizes this bias-correction approach to identify the conventional causal estimand, the gender gap in normalized effects.
For a sketch of the proof, note that if fathers' counterfactual APO is known, then the parallel trends violation for fathers is identified, which in turn identifies the $\mathrm{Bias}(d,d',a)$ term in Theorem (ref). The result follows by multiplying the expression in Theorem (ref) by the inverse of the bias term.
An immediate limitation is that fathers' counterfactual APO is not known.\footnote{Studies that utilize quasi‐experimental settings, such as those exploiting the random success of in vitro fertilization (IVF) treatments, may provide credible evidence on fathers' effects. Among these, two studies report results for men: lundborg2024there find no significant difference in earnings between successful and unsuccessful IVF treatment couples, while bensnes2023reconciling document, if anything, a small positive effect for men.} I therefore use the bias-correction approach as a diagnostic tool. Since both the conventional and bias-corrected estimators target the same causal estimand, assuming a plausible range for fathers' effects allows us to bound the bias in the conventional estimator. Specifically, if the true effect lies within the assumed range, comparing the conventional estimator to the bias-corrected estimator provides a measure of bias magnitude. Below, I empirically implement this exercise using a range from $-10\%$ to $+10\%$ for fathers' normalized effects.
\noindent3.3.2. Empirical Application.
Figure (ref) implements the bias-bounding exercise described above. The black series shows the conventional estimator: the gender gap in normalized DID by treatment group and target post-treatment age. The colored series show bias-corrected estimates from Proposition (ref) for $\theta(m,d,a)\in\{-0.10,-0.05,0,0.05,0.10\}$. Estimators are constructed using sample analogs of population expectations, and standard errors are constructed using influence functions and clustered at the individual level. See Appendix (ref) for further discussion on estimators and standard errors.
The results suggest the bias is heterogeneous across treatment groups. At five years post-treatment, the conventional estimator for earlier treatment groups ($D=24,25,26$) is less negative than bias-corrected estimates across all assumed values of $\theta(m,d,a)$. To illustrate, consider treatment group $D=26$: the conventional estimate is $-0.104$ (SE $0.008$). Assuming $\theta(m,d,a)=-0.10$ yields a bias-corrected estimate of $-0.129$, while $\theta(m,d,a)=0.10$ yields $-0.157$—that is, $24\%$ and $51\%$ larger in magnitude, respectively. In contrast, for later treatment groups ($D\geq28$), conventional and bias-corrected estimates are similar, providing no evidence of bias.
These differences reflect the bias characterized in Theorem (ref). For earlier treatment groups, bias-corrected estimates are more negative, indicating the conventional estimator attenuates child penalties toward zero. This is consistent with Section (ref): positive selection on human capital implies negative parallel-trends violations for these comparisons, which cause child penalties to be understated.
I now turn to pre-treatment validation tests, which serve two purposes: assessing parallel-trends violations empirically, and determining which treatment groups satisfy NTD for the alternative approach developed in the next section. I argue these tests should be conducted separately for each treatment–control pair, and conduct such tests on the Israeli data.
In DID, the conventional validation approach tests for zero difference in trends in pre-treatment periods, commonly referred to as “pre-trends testing” autor2003outsourcing,roth2022pretest. As discussed in Section (ref), I consider $2\times2$ comparisons with the closest not-yet-treated group serving as the control group. The number of control groups therefore equals the number of post-treatment event times analyzed: estimating child penalties for $e = 0, \ldots, 5$ requires six distinct control groups. Since Assumptions (ref) and (ref) must hold separately for each treatment–control pair, pre-trends testing should be performed at this level as well.\footnote{This point generalizes beyond the closest-not-yet-treated assignment. Whenever the control group consists of not-yet-treated individuals in a staggered adoption design, its composition changes at each post-treatment event time, requiring separate validation tests for each pairing.}
Aggregation poses an additional challenge. Event-study validation tests commonly aggregate across all treatment and control groups (Section (ref)). However, such aggregation does not validate the identification assumptions, which must hold separately for each treatment–control pair. Disaggregated pre-trends tests are therefore essential.
Figure (ref) reports pre-treatment validation tests for treatment–control pairs in the Israeli data. Since I report estimates up to five years post-treatment, each treatment group $d$ requires control groups $d^\prime-d=1,\ldots,6$.\footnote{As discussed in Section (ref), I do not include earnings data prior to age 20. To ensure four pre-treatment years, the youngest treatment group is $D=24$. I also exclude treatment groups past age 40, so the oldest treatment group with five post-treatment years is $D=34$.} I include validation tests for NTD because the next section develops identification results that build on this framework; TD tests are included for completeness.
Three patterns stand out. First, pre-treatment DID estimates (top two rows) increase in magnitude as the treatment–control age gap widens. For example, comparing treatment group $D=25$ to control group $D=26$ (control $+1$ in the legend) shows a small difference with confidence intervals that include zero, while $D=25$ versus $D=30$ diverges sharply. This pattern holds across nearly all treatment groups and for both mothers and fathers. Second, TD and NTD exhibit smaller violations for mid-range groups ($D=26,\ldots,30$). Hence, I focus on these treatment groups when discussing post-treatment results below. Finally, TD and NTD pre-treatment estimates appear smaller in magnitude than DID estimates; however, such differences are difficult to interpret without benchmarking against estimated effects.
The above discussion established that NTD is the identification framework underlying normalized child penalties, and that under NTD the conventional estimator is biased for its target causal estimand. This section presents the second main result of the paper: that the effect of parenthood on the gender earnings ratio is identified under NTD. I first present the identification result, corresponding estimator, and the relationship between the new and conventional estimands. I then present an empirical application of the new estimator. I conclude by discussing two exercises relevant for applied work: decomposing the gender gap and aggregating across treatment groups.
This subsection presents the identification theorem for the new causal estimand and corresponding estimator, then discusses how the new and conventional causal estimands are related.
\noindent5.1.1 Identification.
Let $\rho(d,d^\prime,a)=\tfrac{APO(f,d,d^\prime,a)}{APO(m,d,d^\prime,a)}$ denote the gender earnings ratio for treatment group $d$ at age $a$ under counterfactual treatment timing $d^\prime$. The new target causal estimand is the effect of parenthood on the gender earnings ratio, $\Delta\rho(d,a)=\rho(d,d,a)-\rho(d,\infty,a)$. The next result shows $\Delta\rho(d,a)$ is identified under NTD.
The proof is provided in Appendix (ref). A sketch of the proof is as follows. The gender earnings ratio under the realized treatment $\rho(d,d,a)$ is identified directly from observed data by consistency, while the counterfactual ratio $\rho(d,\infty,a)$ is identified via the gender ratio of $\delta_{\mathrm{APO}}$ under Assumption (ref).
Theorem (ref) has a direct implication for applied work: when NTD holds, researchers should target $\Delta\rho(d,a)$ rather than the gender gap in normalized effects. Importantly, this is not merely a fallback. The estimand $\Delta\rho(d,a)$ quantifies how parenthood affects gender inequality in the labor market, the core question motivating the child penalty literature.
\noindent5.1.2 Estimation.
Estimation follows directly from Theorem (ref): replace population expectations with sample means to obtain estimates for the observed earnings ratio and the $\delta_{\mathrm{APO}}$ gender ratio, then difference to construct the estimator $\widehat{\Delta\rho}(d,a)$. For inference, I derive clustered standard errors based on the estimator's influence function (Appendix (ref)). This approach yields analytical, cluster-robust standard errors. To facilitate application, I developed an R package, \href{https://dorleventer.github.io/childpen}{childpen}.
\noindent5.1.3 Relationship between the Causal Estimands.
The relationship between the new causal estimand, the effect of parenthood on the gender earnings ratio, and the conventional causal estimand, the gender gap in normalized effects, can be expressed as:\footnote{The derivation of this expression is provided in the proof of Theorem (ref).}
(ref) shows the new estimand $\Delta\rho(d,a)$ equals the conventional estimand $\theta(f,d,a) - \theta(m,d,a)$ multiplied by a rescaling factor that depends on the counterfactual gender earnings ratio ($\rho(d,\infty,a)$) and men's normalized effect ($\theta(m,d,a)$). When $\rho(d,\infty,a)$ is smaller than one and male effects are close to zero or positive, the new estimand will be smaller in absolute value than the conventional estimand. While I cannot validate effects for males, I can provide descriptive evidence on the counterfactual gender earnings ratio using pre-treatment data, presented below.
This subsection presents empirical evidence on the new estimand. I first examine observed gender earnings ratios in pre-treatment data, then compare the new and conventional estimators in post-treatment.
\noindent5.2.1 Evidence on $\rho(d,\infty,a)$ in Pre-Treatment.
To study how counterfactual gender inequality $\rho(d,\infty,a)$ evolves across the life cycle and treatment groups, I compute mean earnings at every age $a<d$ for each treatment group $d$ and form the female-to-male observed earnings ratio. Under the no-anticipation assumption (Assumption (ref)), these pre-childbirth earnings identify the counterfactual $APO$s.
Figure (ref) plots the female-to-male observed earnings ratio for each treatment group $d\in[27,38]$. Two patterns emerge. First, within each treatment group the gender earnings ratio before childbirth follows a similar life-cycle pattern: rising up to age 27 and falling from age 28, producing an inverted U-shape. Second, at any given age, parents who delay their first birth generally exhibit higher ratios than earlier-childbearing parents. Importantly, the life-cycle pattern is such that at later ages all groups have a gender earnings ratio smaller than one. Under no anticipation, this implies $\rho(d,\infty,a) < 1$ at these ages.
\noindent5.2.2 Comparing Estimators in Post-Treatment. The new estimand can be visualized as the difference between two gender earnings ratios. Figure (ref) reports post-treatment estimates of the counterfactual gender earnings ratio $\frac{\delta_{\mathrm{APO}}(f,d,d^\prime,a)}{\delta_{\mathrm{APO}}(m,d,d^\prime,a)}$ in red and the observed gender earnings ratio $\frac{\mathbb{E}[Y_a\mid G=f,D=d]}{\mathbb{E}[Y_a\mid G=m,D=d]}$ in blue, separately by treatment group. Under NTD, the red series identifies the gender earnings ratio in the counterfactual absent childbirth ($\rho(d,\infty,a)$), and the blue series identifies the ratio under realized treatment ($\rho(d,d,a)$). The difference between these series estimates the effect of parenthood on the gender earnings ratio, $\Delta\rho(d,a)$.
Figure (ref) compares the conventional (red) and new (blue) estimators for treatment groups $D = 24,...,34$. The conventional estimator estimates $\delta_{\theta}(f,d,d^\prime,a) - \delta_{\theta}(m,d,d^\prime,a)$; the new estimator estimates $\frac{\mathbb{E}[Y_a \mid G=f, D=d]}{\mathbb{E}[Y_a \mid G=m, D=d]} - \frac{\delta_{\mathrm{APO}}(f,d,d^\prime,a)}{\delta_{\mathrm{APO}}(m,d,d^\prime,a)}$. Both replace population expectations with sample analogs. Validation tests suggest Assumption (ref) is most plausible for treatment groups $D=26,...,30$ (Section (ref)). Accordingly, I focus on these groups below.
The figure shows the following patterns. For $D=26,27$, the new estimator starts slightly less negative than the conventional at childbirth but converges and becomes marginally more negative by five years post-treatment. For $D=29,30$, the new estimator is smaller in magnitude than the conventional. For example, at $D=30$ five years post-treatment, the conventional estimator equals $-0.200$ (SE $0.015$) while the new estimator equals $-0.144$ (SE $0.011$), a difference of 28%.
The intuition behind this result for later treatment groups is as follows. The bias-bounding exercise in Section (ref) found no evidence of substantial bias for later treatment groups. The differences observed here likely reflect different target estimands rather than bias. Furthermore, the smaller magnitude of the new estimator is consistent with small effects for fathers and $\rho(d,\infty,a) < 1$ as discussed above theoretically above and empirically in the previous subsection.
This subsection discusses two exercises relevant for applied work: decomposing the gender gap and aggregating across treatment groups.
\noindent5.3.1 Decomposing the Gender Gap.
An important accounting exercise in the child penalty literature is to quantify how much of the overall gender gap is attributable to parenthood kleven2019child,cortes2023children. The identification result in Section (ref) implies that NTD enables this decomposition directly. In Figure (ref), the distance from one to the red series measures the impact of factors other than parenthood on the gender earnings ratio ($1-\rho(d,\infty,a)$), while the distance between the red and blue series measures the direct impact of parenthood ($\Delta\rho(d,a)$).
To illustrate, consider treatment group $D=30$ five years post-childbirth. The estimated counterfactual gender earnings ratio is $0.726$ and the realized gender earnings ratio is $0.582$, implying overall gender inequality of $1 - 0.582 = 0.418$. Of this, parenthood accounts for $0.726 - 0.582 = 0.144$, or approximately $34\%$, while other factors account for the remaining $66\%$.
\noindent5.3.2 Aggregation.
Applied work typically reports estimates aggregated across all treatment groups, as in the normalized event studies (Section (ref)). Appendix (ref) discusses how to aggregate multiple treatment groups, building on the discussion on aggregation in callaway2021difference, for both the conventional and new estimands.
An important caveat arises when comparing aggregated child penalties across strata, such as countries kleven2019child or parent types andresen2022causes: even if single-treatment-group effects are identical across strata, differences in treatment distributions across strata can generate differences in aggregates. For example, countries where parents have children earlier will place more weight on early treatment groups, potentially yielding different aggregate estimates than countries with later childbearing, even if treatment-group-specific effects are identical across countries. Appendix (ref) elaborates on this issue and provides an empirical illustration.
This paper revisits the identification of child penalties in the normalized event‐study framework. I begin by formalizing the underlying identification framework that can be rationalized from validation tests, which I term Normalized Triple Differences (NTD). I show that under NTD, the conventional estimator does not identify its target causal estimand—the gender gap in normalized effects—when the parallel trends assumption is violated. Drawing on human capital theory, I argue such violations are theoretically likely, and an empirical bias-bounding exercise documents substantial bias for early-treated parents. As a solution, I propose targeting the effect of parenthood on the gender earnings ratio and show this estimand is identified under NTD without additional assumptions.
While this paper provides a better understanding of the identification of child penalties, several avenues for future research remain. The analysis relies on both SUTVA and Assumption (ref); exploring how violations of these assumptions affect results is an important next step. Moreover, the discussion focused on earnings as the outcome; extending the framework to employment, hours worked, or firm-level dynamics offers a natural direction for future work. More broadly, the approach developed here can inform other applications with endogenous treatment timing, where selection on treatment and counterfactual inequality affect the validity of identification assumptions.
\printbibliography