Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
171,789 characters · 14 sections · 0 citation commands
Distribution Regression Difference-In-Differences
Keywords: difference in differences, multiple outcomes, distributional and quantile treatment effects
The remarkable popularity of the difference-in-difference (DiD) estimator, inspired by an approach to evaluating the impact of policy interventions on economic outcomes introduced by David Card (see, for example, Card 1990, Card and Krueger, 1994), is one of the most striking features of empirical work on treatment and policy effects. While the methodological innovations in this literature (see Arkhangelsky and Imbens, 2024 for a recent review) include the use of constructed control groups, the staggered timing of treatments, and fuzzy rather than sharp designs, the vast majority of the associated empirical work has estimated the mean effect of the treatment on a single economic outcome. This seems somewhat limited and a fuller evaluation of a policy treatment would be based on an examination of the marginal and joint distributions of all outcomes it potentially influences. This paper provides a simple procedure for estimating distributional treatment effects in the presence of a single treatment when the outcomes of interest are potentially multivariate.
An initial methodological innovation focusing on distributional effects in DiD estimation is the changes-in-changes procedure of Athey and Imbens (2006), which estimates the counterfactual distribution of the treated group in the absence of treatment to compare with its observed distribution in the presence of treatment. Torous et al. (2024) extend the Athey and Imbens approach (2006) to the multivariate outcome setting. Other work has adapted DiD estimation to examine the treatment effects at different quantiles of the outcome via the use of quantile regression. This includes, for example, Callaway and Li (2018, 2019). In contrast, Dube (2019), Goodman-Bacon (2021), and Goodman-Bacon and Schmidt (2020) employ conventional DiD estimation to explore the impact of the treatment at different points of the outcome distribution. Other distributional approaches include Kim and Wooldridge (2023) and Biewen, Fitzenberger, and Rümmele (2022). The former proposes an inverse probability weighting based procedure, while the latter employs a distribution regression (DR) approach. In this paper we also adopt a DR approach to constructing counterfactuals. In contrast to Biewen, Fitzenberger, and Rümmele (2022), who construct the counterfactual distributions via linear probability models, we employ non-linear link functions such as probit or logit models. This has a number of advantages, which we discuss below. In addition, we provide the associated identifying conditions required for this form of implementation of DR-DiD.
While DiD has typically been employed to evaluate the treatment effect on a specified economic outcome, there are many instances in which the treatment may affect multiple outcomes. For example, a change in tax rates on earnings of married couples may affect the hours of work of both husbands and wives. An analysis of such a tax change should include the impact on each of the outcomes. However, a richer analysis would not only examine the impact on the respective marginal hours distributions of husbands and wives but also the joint distribution of hours. Alternatively, while evaluations of minimum wage laws typically evaluate their impact on total employment, they may also affect the joint distribution of part-time and full-time employment. We illustrate how this joint effect can be evaluated via the bivariate distribution regression (BDR) approach of Fern\'andez-Val et al. (2024a). This requires that we first estimate the joint distribution by BDR and then construct the appropriate counterfactual. The treatment effects are then obtained via the appropriate comparisons. A recent alternative to this approach is extending the changes-in-changes procedure to multiple outcomes as is done in Torous et al. (2024).
The following section introduces the model and provides an analysis of the univariate case without covariates. We also extend our analysis to include covariates and contrast our approach with the Athey and Imbens (2006) changes-in-changes procedure. Section (ref) extends our analysis to the multiple outcome case and Section (ref) discusses estimation. Section (ref) provides an empirical illustration via an application of our procedure to the data employed in the Card and Krueger (1994) study of the impact of increasing the minimum wage on employment. Section (ref) concludes.
Consider the standard DiD design with 2 periods, $T \in \{0,1\}$, and 2 groups, $G \in \{0,1\}$ in which a binary treatment, $D \in \{0,1\}$, is administered only to the treatment group with $G = 1$ in the second period $T=1$. Let $Y_0$ and $Y_1$ denote the potential outcomes under the non-treated and treated statuses. The observed outcome is $Y = Y_0 (1-D) + Y_1 D$, which corresponds to $Y_0$ for both groups at $T=0$, $Y_0$ for $G=0$ at $T=1$, and $Y_1$ for $G=1$ at $T=1$. Note that this implicitly imposes a non-anticipation assumption as we do not distinguish between the outcomes of the treated and non-treated state for $G=1$ in period $T=0$.
We are interested in the distributions of the potential outcomes of the treated at $T=1$, that is $F_{Y_1 \,|\, G, T}(y \,|\, 1,1)$ and $F_{Y_0 \,|\, G, T}(y \,|\, 1,1)$. $F_{Y_1 \,|\, G, T}(y \,|\, 1,1)$ is identified from the observed outcome for $G=1$ at $T=1$, $$ F_{Y_1 \,|\, G, T}(y \,|\, 1,1) = F_{Y \,|\, G, T}(y \,|\, 1,1); $$ whereas $F_{Y_0 \,|\, G, T}(y \,|\, 1,1)$ is not identified without further assumptions.
The distribution of $Y_0$ conditional on $G$ and $T$ can be written as:
where $\Lambda$ is an invertible CDF such as the logistic, normal or uniform, and $y \mapsto $ $(\alpha(y), $ $ \beta(y), \gamma(y), \delta(y))$ is a vector of function-valued parameters.
The representation in (ref) does not make any parametric assumption about the underlying distribution of $Y_0 \,|\, G, T$ since the dummy variable representation within the parentheses on the right-hand side is fully saturated. The parameters of the representation are local as they vary with $y$. To understand why (ref) does not impose any restriction, note that $\alpha(y)$, $\beta(y)$, $\gamma(y)$ and $\delta(y)$ can be defined as:\footnote{See also Wooldridge (2023) equations (2.6) and (2.7).} \[
\] We make the following identifying assumptions:
Let $\mathcal{Y}_d(G=g, T=t)$ denote the support of $Y_d \,|\, G=g,T=t$, for $d, g, t \in \{0,1\}$. We also assume:
Assumption (ref) implies that the distribution of the potential outcome $Y_0$ should not change differently in the second period for the treatment group compared to the control group. That is, we allow a difference between the distributions of the potential outcome $Y_0$ between the treatment and control group, but this difference should be identical in both periods. This is a parallel trend type assumption on a transformation of the distribution and can be written as: \[
\] This assumption is sensitive to the link function and imposes restrictions on the distribution $F_{Y_0 \,|\, G, T}$ for some link functions. For example, if $\Lambda$ is the identity link used in the linear probability model as in, for example, Almond et al. (2011), Dube (2019), Cengiz et al. (2019), Goodman-Bacon and Smith (2020), Goodman-Bacon (2021) and Biewen et al. (2022), one needs strong requirements in order to satisfy the parallel trends assumption (Blundell et al., 2004 and Wooldridge, 2023) That is, we need restrictions on the tails of the distribution of $F_{Y_0 \,|\, G, T}(y \,|\, 1,0)$, $F_{Y_0 \,|\, G, T}(y \,|\, 0,1)$ and $F_{Y_0 \,|\, G, T}(y \,|\, 0,0)$ to guarantee that $F_{Y_0 \,|\, G, T}(y \,|\, 1,1)$ is between $0$ and $1$. Thus, it requires that $F_{Y_0 \,|\, G, T}(y \,|\, 1,0) \leq 1 + F_{Y_0 \,|\, G, T}(y \,|\, 0,0) - F_{Y_0 \,|\, G, T}(y \,|\, 0,1)$, which might be restrictive at the top of the distribution, and $F_{Y_0 \,|\, G, T}(y \,|\, 1,0) $$\geq F_{Y_0 \,|\, G, T}(y \,|\, 0,0) - F_{Y_0 \,|\, G, T}(y \,|\, 0,1)$, which might be restrictive at the bottom of the distribution.\footnote{These requirements could be used to develop a specification test for the identity link. Roth and Sant'Anna (2023) proposed a test for the sharp hypothesis that $y \mapsto F_{Y_0 \,|\, G, T}(y \,|\, 1,0) + F_{Y_0 \,|\, G, T}(y \,|\, 0,1) - F_{Y_0 \,|\, G, T}(y \,|\, 0,0)$ be weakly increasing, which can be adapted to our setting. We do not pursue this route as we do not encourage the use of the linear probability model.}$^{,}$\footnote {For example, an increase in 0.2 in probability over time might be realistic for the control group when the initial probability was 0.5. However, if treatment group has a probability of, for example, 0.9, in the first period then it is not possible for the common trends assumption to hold.} Link functions such as the normal or logistic CDFs do not require such restrictions since the transformation expands the range of the distribution to the entire real line.
Assumption (ref) cannot be empirically verified but when we have multiple observations in the pre-treatment period, it is possible to examine whether the “parallel trends” assumption holds pre-treatment. Assumption (ref) is a restriction of the support of the counterfactual outcome of $Y_0$ for the treated group in the treated period.
These two assumptions identify $F_{Y_0 \,|\, G, T}(y \,|\, 1,1)$ since:
under Assumption (ref). The support restrictions in Assumption (ref) ensure that the term inside the squared brackets in (ref) is determined. Note that as $\lim_{x \rightarrow \infty}\Lambda(x) = 1$ and $\lim_{x \rightarrow -\infty}\Lambda(x) = 0$, our assumptions are sufficient but not necessary.
We present this identification result in the following lemma:
In empirical analysis, researchers typically would like to investigate objects that are related to the distributions of the potential outcome variables. One such object is the distributional treatment effect, defined as: \[ \tau(y) := F_{Y_1 \,|\, G,T}(y \,|\, 1,1)(y) - F_{Y_0 \,|\, G,T}(y \,|\, 1,1), \quad y \in \mathbb{R}. \] The distributional treatment effect measures the change in the probability that the outcome is below $y$ as a result of the treatment. Another interesting object is the quantile treatment effect, defined as: \[ \tau^*_q := F^{-1}_{Y_1 \,|\, G,T}(q \,|\, 1,1)(y) - F^{-1}_{Y_0 \,|\, G,T}(q \,|\, 1,1), \quad q \in (0,1). \] The quantile treatment effect measures the difference in the q-th quantile of the outcome variable as a result of the treatment.
Including covariates is appealing as the assumption that $\delta(y) = 0$ may be harder to defend when there are differences in the trend between covariates and/or the composition of the treatment group changes over time in terms of observed characteristics; see also Melly and Santangelo (2015). Covariates are easily incorporated into the identification result by conditioning on them and adding an overlapping support assumption. Specifically, let $X$ be a vector of covariates such that the non-interaction assumption holds conditional on $X$; see Assumption (ref). The distribution of $Y_0$ conditional on $G$, $T$ and $X$ can be written:
where $(y,x) \mapsto (\alpha(y,x), \beta(y,x), \gamma(y,x), \delta(y,x))$ is a vector of unspecified functions.
Let $\mathcal{Y}_d(G=g, T=t; X=x)$ denote the support of $Y_d \,|\, G=g,T=t, X=x$. The identifying assumptions with covariates become:
These two assumptions identify $F_{Y_0 \,|\, G, T, X}(y \,|\, 1,1,x)$ since
under the Assumption (ref). The support restrictions in Assumption (ref) ensure that the term between parentheses in (ref) is determined. Note that as $\lim_{x \rightarrow \infty}\Lambda(x) = 1$ and $\lim_{x \rightarrow -\infty}\Lambda(x) = 0$, our assumptions are sufficient but not necessary.
Let $\mathcal{X}_{11}$ denote the support of $X$ conditional on $G=1$ and $T=1$. The following Lemma states that $F_{Y_0,Z_0 \,|\, G,T,X}$ is identified under the previous assumptions.
We can then identify $F_{Y_0 \,|\, G, T}(y \,|\, 1,1)$ as:
where $F_{X \,|\, G,T}$ is the distribution of $X$ conditional on $G$ and $T$.
As our proposals provide an alternative approach to the changes-in-changes (CiC) procedure of Athey and Imbens (2006), it is useful to contrast their setup and assumptions with ours. CiC assumes that the outcome of an individual without treatment satisfies the relationship $Y_0 = h(U,T)$ for the treatment and control groups, where $U$ is an unobserved and uniformly distributed random variable. It also assumes that $h$ is strictly increasing in the first term and that the distribution of $U$ is independent of time given the treatment outcome, i.e. $U \perp\!\!\!\perp T \,|\, G$. Finally, the support of $U$ for the treated population should be a subset of those of the untreated population. The final assumption implies in terms of the support of the potential outcomes that: \[
\] Their second support restriction is less restrictive than ours but we do not need their first support restriction. The previous assumptions identify the quantile function of $F_{Y_0 \,|\, G, T}(y \,|\, 1,1)$ as: \[
\] where it is assumed that $Y_0$ is continuous with strictly increasing distribution function.
The transformation $\phi$ gives the second period outcome for an individual with an unobserved component $u$ such that $h(u,0) =y$, with $y$ the location at which the distribution function is evaluated (Athey and Imbens, 2006, page 441). Hence, their identification results follow since $\phi$ evaluated in the first period observations of the treatment group is equally distributed as the distribution of the untreated outcome of the treatment group in the second period. Their assumptions imply the transformation $\phi$ that maps quantiles of $Y_0$ from period $0$ to period $1$ is the same for the treatment and control groups. This condition imposes the following restrictions on the coefficients of the representation of the conditional distribution in (ref): $$ \alpha(y) = \alpha(\phi(y)) + \beta(\phi(y)), \quad \gamma(y) = \gamma(\phi(y)) + \delta(\phi(y)) . $$ To see this, note that:
Evaluating (ref) at $g=0$ and applying $F^{-1}_{Y_0 \,|\, G,T}(\cdot \,|\, 0,1)$ to both sides: $$ h(h^{-1}(y,0),1) = F^{-1}_{Y_0 \,|\, G, T}\left( F_{Y_0 \,|\, G, T}(y \,|\, 0,0) \,|\, 0,1\right) =: \phi(y). $$ Replacing $\phi(y)$ back in (ref) and using the representation (ref): $$ \Lambda(\alpha(y) + \gamma(y)g) = \Lambda(\alpha(\phi(y)) + \beta(\phi(y)) + \gamma(\phi(y))g + \delta(\phi(y))g ). $$ The restrictions then follow from equalizing the coefficients in both sides.\footnote{There is only a binding restriction because $\alpha(y) = \alpha(\phi(y)) + \beta(\phi(y))$ holds by definition of $\phi(y)$.} They would complicate estimation in our framework as they involve two different levels of $Y$ and the transformation $\phi$ needs to be estimated.
Roth and Sant'Anna (2023) derive the condition: $$ F_{Y_0 \,|\, G, T}(y \,|\, 1,1) - F_{Y_0 \,|\, G, T}(y \,|\, 1,0) = F_{Y_0 \,|\, G, T}(y \,|\, 0,1) - F_{Y_0 \,|\, G, T}(y \,|\, 0,0), \quad y \in \mathbb{R}, $$ for the parallel trends assumption in expectations: $$ {\mathbb{E}}(Y_0 \,|\, G = 1, T = 1) - {\mathbb{E}}(Y_0 \,|\, G = 1, T = 0) = {\mathbb{E}}(Y_0 \,|\, G = 0, T = 1) - {\mathbb{E}}(Y_0 \,|\, G = 0, T = 0), $$ to be invariant to strictly monotone transformations of $Y_0$. This condition is different from our no-interaction assumption. Indeed, our DR model with no-interaction does not generally satisfy the parallel trends assumption in expectation as: \[
\] depends on $g$ unless $\Lambda$ is the identity map, or $\beta(y)=0$ (no trend) or $\gamma(y)=0$ (random assignment) for $y \in \mathbb{R}$. Roth and Sant'Anna (2023) show that their condition holds if and only if there are no trends, random assignment or a mixture of both. Our model, however, generally satisfies a different invariance property with respect to strictly monotonic transformations that we specify in {subsection} (ref).
The DR model in (ref) with no-interaction is invariant to strictly monotonic transformations in the sense that we specify here. If $Y_0$ follows the DR model and satisfies the no-interaction assumption, then $\tilde Y_0 = h(Y_0)$ also follows the DR model and satisfies the no-interaction assumption for any strictly monotonic transformation $h$. To see this result note that if $h$ is strictly increasing: $$ F_{\tilde Y_0 \,|\, G, T, X}(\tilde y \,|\, g,t,x) = \Lambda(\alpha(h^{-1}(\tilde y)) + \beta(h^{-1}(\tilde y)) t + \gamma(h^{-1}(\tilde y))g ) = \Lambda(\tilde \alpha(\tilde y) + \tilde \beta(\tilde y) t + \tilde \gamma(\tilde y)g ), $$ where $\tilde y \mapsto h^{-1}(\tilde y)$ is the inverse function of $y \mapsto h(y)$, $\tilde \alpha = \alpha \circ h^{-1}$, $\tilde \beta = \beta \circ h^{-1}$ and $\tilde \gamma = \gamma \circ h^{-1}$. A similar argument applies when $h$ is strictly decreasing. Unlike the parallel trends in expectation, the no-interaction or parallel trends in distribution is invariant to strictly monotonic transformations.\footnote{The distributional approach of Kim and Wooldridge (2023) also satisfies this property.}
Some settings may feature multiple outcomes that are potentially affected by the treatment. In these situations, we might be interested not only in how each of the outcomes is affected by the treatment, but also in how the relationship between the outcomes is affected by the treatment. For this, it is necessary to identify the joint distribution of the potential outcomes with and without treatment. We now consider a setting with two outcomes $Y$ and $Z$ and we focus on comparing features of the joint distribution of the potential outcomes with the treatment, $Y_1$ and $Z_1$, and the joint distribution of the potential outcomes without the treatment, $Y_0$ and $Z_0$, for the treated group $G=1$ in the post-treatment period $T=1$. For the sake of illustration we consider two measures of dependence. Namely, Spearman's and Kendall's rank correlation.
Let $F_{Y_d,Z_d \,|\, G,T}$ be the joint distribution of $Y_d$ and $Z_d$ conditional on $G$ and $T$, and $F_{Y_d \,|\, G,T}$ and $F_{Z_d \,|\, G,T}$ be the corresponding marginals. Spearman's rank correlation between $Y_d$ and $Z_d$, $d \in \{0,1\}$, can be expressed:
and Kendall's rank correlation between $Y_d$ and $Z_d$, $d \in \{0,1\}$, can be expressed:
where we assume that $Y_d$ and $Z_d$ are continuous random variables to obtain the expressions on the right hand side.
As in the univariate case, $F_{Y_1,Z_1 \,|\, G,T}(y,z \,|\, 1,1)$ is identified by the joint distribution of the observed outcomes, $F_{Y,Z \,|\, G,T}(y,z \,|\, 1,1)$, whereas $F_{Y_0,Z_0 \,|\, G,T}(y,z \,|\, 1,1)$ is not identified from the data. To analyze identification, we use a variation of the local Gaussian representation (LGR) of a bivariate distribution from Chernozhukov, Fernand\'ez-Val and Luo (2018). Let $\Phi$ denote the Gaussian distribution function and $\Phi_2(\cdot,\cdot;\rho)$ denote the distribution of the bivariate standard normal with parameter $\rho$. Moreover, $\Lambda$ is, again, a strictly increasing cumulative distribution function. As we show in Section (ref), there is a benefit of using the logistic link function in our univariate analysis. Accordingly, we employ this in our empirical analysis for estimating both the univariate and bivariate effects.
The difference between Lemma (ref) and the LGR of Chernozhukov, Fernand\'ez-Val and Luo (2018) is that the marginals are represented by a general link rather than Gaussian links, that is: $$ F_{Y\,|\, X}(y \,|\, x)(y \,|\, x) \equiv \Lambda(\mu_{Y \,|\, X}(y \,|\, x)), \quad F_{Z\,|\, X}(z \,|\, x) \equiv \Lambda(\mu_{Z \,|\, X}(z \,|\, x)). $$ By the LGR, $F_{Y_0,Z_0 \,|\, G,T}$ can be expressed as:
where $\mu_{Y_0 \,|\, G,T}(y \,|\, g,t) = \alpha_Y(y) + \beta_Y(y) t + \gamma_Y(y) g + \delta_Y(y) gt$, $\mu_{Z_0 \,|\, G,T}(y \,|\, g,t) = \alpha_Z(z) + \beta_Z(z) t + \gamma_Z(z) g + \delta_Z(z) gt$, and $\rho_{Y,Z \,|\, G,T}(y,z \,|\, g,t) = \alpha_{Y,Z}(y,z) + \beta_{Y,Z}(y,z) t + \gamma_{Y,Z}(y,z) g + \delta_{Y,Z}(y,z) gt$. In the LGR, the marginals are represented by: $$ F_{Y_0 \,|\, G,T}(y \,|\, g,t) = \Lambda(\alpha_Y(y) + \beta_Y(y) t + \gamma_Y(y) g + \delta_Y(y) gt), $$ and $$ F_{Z_0 \,|\, G,T}(z \,|\, g,t) = \Lambda(\alpha_Z(z) + \beta_Z(z) t + \gamma_Z(z) g + \delta_Z(z) gt). $$ We make the following identifying assumptions with respect to the distribution function in (ref).
Let $\mathcal{YZ}_d(G=g, T=t)$ denote the support of $(Y_d,Z_d) \,|\, G=g,T=t$, for $d, g, t \in {0,1}$. We also assume:
Covariates can be incorporated in a similar fashion as the univariate case. In particular, the LGR of $F_{Y_0,Z_0 \,|\, G,T,X}$ becomes:
where $\mu_{Y_0 \,|\, G,T,X}(y \,|\, g,t,x) = \alpha_Y(y,x) + \beta_Y(y,x) t + \gamma_Y(y,x) g + \delta_Y(y,x) gt$, $\mu_{Z_0 \,|\, G,T,X}(y \,|\, g,t,x) = \alpha_Z(z,x) + \beta_Z(z,x) t + \gamma_Z(z,x) g + \delta_Z(z,x) gt$, and $\rho_{Y,Z \,|\, G,T,X}(y,z \,|\, g,t,x) = \alpha_{Y,Z}(y,z,x) + \beta_{Y,Z}(y,z,x) t + \gamma_{Y,Z}(y,z,x) g + \delta_{Y,Z}(y,z,x) gt$.
Let $\mathcal{YZ}_d(G=g, T=t; X=x)$ denote the support of $(Y_d,Z_d) \,|\, G=g,T=t, X=x$. The identifying assumptions with covariates become:
The marginalized distribution $F_{Y_0,Z_0 \,|\, G,T}(y \,|\, 1,1)$ is then identified by $$ F_{Y_0,Z_0 \,|\, G,T}(y,z \,|\, 1,1) = \int_{\mathcal{X}_{11}} F_{Y_0,Z_0 \,|\, G,T,X}(y,z \,|\, 1,1,x) \mathrm{d} F_{X \,|\, G,T}(x \,|\, 1,1). $$
Assume we have a sample $\{(Y_{i}, X_{i}, G_{i}, T_{i}): 1\leq i \leq N\}$ of $(Y, X, G, T)$. For estimation, we replace the functions $(y,x) \mapsto (\alpha(y,x), \beta(y,x), \gamma(y,x))$ in (ref) by semiparametric linear indexes leading to the DR model for the conditional distribution:
where $p_{\alpha}(x)$, $p_{\beta}(x)$ and $p_{\gamma}(x)$ are vectors including the covariates and their transformations, and $y \mapsto (\alpha(y), \beta(y), \gamma(y))$ is a vector of function-valued parameters.
We implement the DR DiD estimator via the sequence of logit models at each point of the distribution of the outcome variable (Foresi and Peracchi, 1995, Chernozhukov, Fernandez-Val and Melly, 2013). We choose logit because it is the canonical link for binary outcomes allowing for pooled estimation of the distributions of the potential outcomes with and without the treatment (Wooldridge, 2023). Accordingly, we estimate the DR model for the observed outcomes on all observations including those with $D_i = 1$:
where $p_{\alpha}(x)$, $p_{\beta}(x)$, $p_{\gamma}(x)$ and $p_{\theta}(x)$ are vectors including a constant as the first component, and transformations of the covariates, and $\mathcal{Y}$ is a finite grid on $\mathbb{R}$. Let $I_i^y := 1(Y_i \leq y)$ and $\bar I_i^y = 1-I_i^y$.
By the properties of the logistic link, the estimator of $F_{Y_1 \,|\, G, T}(y \,|\, 1,1)$ is identical to the empirical distribution of $Y$ conditional on $G=1$ and $T=1$, $$ \hat F_{Y_1 \,|\, G, T}(y \,|\, 1,1) \equiv \frac{1}{N_{11}} \sum_{i=1}^N G_i T_i \ 1(Y_i \leq y). $$ Note that this estimator is therefore invariant to the specification of $p_{\theta}(x)$. We set $p_{\theta}(x)=1$ to speed up computation. Note that the definition of the estimator for the inverse of the distribution function is appropriate as we rearranged the estimates $y \mapsto \hat F_{Y_d \,|\, G, T}(y \,|\, 1,1)$ on $y \in \mathcal{Y}$, $d \in \{0,1\}$.
Our algorithm has an advantage over the alternative of estimating $p_{\alpha}$, $p_{\beta}$, $p_{\gamma}$ using only those observations for which $D_i = 0$ via direct estimation of (ref). For example, the distributional treatment effect, \[ \int_{\mathcal{X}_{11}} \left[ F_{Y_1|G,T,X}(y|1,1,X) - F_{Y_0|G,T,X}(y|1,1,X) \right] \mathrm{d}F_{X \,|\, G,T}(x \,|\, 1,1), \] equals the average partial effect of $D_i$ for the logit model used for distribution regression. Estimates and standard errors for these effects are routinely reported by many Statitical software packages (Wooldridge, 2023).
Assume we have a sample $\{(Y_{i}, Z_{i}, X_{i}, G_{i}, T_{i}): 1\leq i \leq N\}$ of $(Y, Z, X, G, T)$. For estimation, as in the univariate case, we replace the functions in $\mu_{Y_0 \,|\, G,T,X}$, $\mu_{Z_0 \,|\, G,T,X}$ and $\rho_{Y,Z \,|\, G,T,X}$ by semiparametric generalized linear indexes leading to a bivariate distribution regression (BDR) model:
and
where $p_{\alpha}(x)$, $p_{\beta}(x)$, $p_{\gamma}(x)$, $q_{\alpha}(x)$, $q_{\beta}(x)$, $q_{\gamma}(x)$, $r_{\alpha}(x)$, $r_{\beta}(x)$ and $r_{\gamma}(x)$ are vectors including the covariates and their transformations, and $h(u) = \text{arctanh}(u)$ is the Fisher transformation that enforces $\rho_{Y,Z \,|\, G,T,X}$ to lie in $[-1,1]$.
We estimate all the parameters of $F_{Y_0,Z_0 \,|\, G,T}(y,z \,|\, 1,1)$ using the bivariate distribution regression estimator of Fernandez-Val et al. (2024a). We employ an imputation method that combines the parameter estimates from the sample of the first period for both groups and the sample of the second period for the untreated group, with the sample of the covariates in the second period for the treated group. The distribution $F_{Y_1,Z_1 \,|\, G,T}(y,z \,|\, 1,1)$ is estimated using the empirical distribution of $Y$ and $Z$ in the second period for the treated group. Algorithm (ref) describes the estimation procedure. Let $\mathcal{Y}$ and $\mathcal{Z}$ be finite grids on $\mathbb{R}$, $I_i^y := 1(Y_i \leq y)$, $\bar I_i^y = 1-I_i^y$, $J_i^z := 1(Z_i \leq z)$, $\bar J_i^z = 1-J_i^z$.
Estimators of the functionals of the joint distributions of potential outcomes such as Spearman's and Kendall's rank correlation coefficients can be constructed using the plug-in principle.
The estimators described in Algorithms (ref) and (ref) can be applied to panel and repeated cross-section data. Here we describe a weighted bootstrap algorithm to perform inference on functions of the distributions of potential outcomes designed for panel data. We focus on this case because it is relevant for our empirical application below.
To describe the procedure, we need to introduce an indicator $ID_i$, $i = 1, \ldots, N$, for the units in the panel. For example, if the sample is sorted by unit and time period, $ID = (1,1,2,2,\ldots,n,n)$, where $n=N/2$. The following algorithm describes the weighted bootstrap procedure to construct joint confidence bands for the distributions of the potential outcomes with and without the treatment in the univariate case. Inference for functionals of the distributions and for the bivariate case can be performed using similar algorithms.
We illustrate our approach through a re-examination of data used in Card and Krueger (1994), hereafter CK, investigation of the impact on an increase in the minimum wage on the level of employment. In April 1992 New Jersey increased its minimum wage from the Federal level of 4.25 dollars per hour to 5.05 dollars per hour. CK investigated the impact of this increase on the change in the level of "full-time equivalent employment", measured as the sum of the number of full-time employees plus half the number of part-time employees, in New Jersey fast food restaurants. They use a DiD estimation strategy in which the control group comprises a group of comparable fast food restaurants in the region of Pennsylvania bordering New Jersey. CK concluded that this particular increase in the minimum wage led to a small increase in the level of full-time equivalent employment. This result produced a large and important related literature on the impact of minimum wages on employment.
We employ our procedure using the CK data to investigate the impact of this increase in the minimum wage on the level of full-time, part-time and full-time equivalent employment respectively. We estimate the model using the 409 observations available in CK data set. 80.9 percent of these are observations are treated establishments in New Jersey. We acknowledge that the data set is relatively small and that this is likely to have implications for the level of statistical significance of the results. However, as our objective is to illustrate our approach in a well-known setting, we prefer to work with a data set which has been frequently used and is well understood (see, for a recent example, Torous et al. 2024) rather than providing our own original application. Note that for the models which include covariates, the additional variables are four dummy variables for franchise type and a dummy variable indicating that the establishment is company owned.
The respective actual and counterfactual distributions are reported in Figures (ref)-(ref). These figures are based on the models which include the additional covariates. While there are some differences in the distributions for full-time employment, indicating an increase in employment for establishment sizes above the 1st quartile, the larger differences appear in Figure (ref) which captures the increases in full-time employment. This figure suggests gains at all quantiles. Figure (ref) is suggestive of some small reductions in part-time employment at some quantiles.
The results of the quantile treatment effects for the univariate analyses are reported in Tables (ref) and (ref) noting that those in the former include the covariates while the latter does not. As the results are generally similar, we focus only on Table 1. The DiD estimate of the mean effect on full-time equivalent employment is 2.65 and this is statistically significant at the 10 percent level. An examination of the table reveals that this mean effect is driven by an increase in full-time employment as there is very weak statistical evidence of a small reduction in mean part-time employment. The level of statistical significance probably reflects the small sample size.
Tables (ref) and (ref) also indicate that the effect of the minimum wage increase is different depending on the size of the establishment. For example, at smaller establishments the effect on full-time equivalent employment is negative and this reflects a reduction in these establishments' levels of part-time employment. However, there appears to be growth in full-time equivalent employment at the larger establishment sizes noting that the level of statistical significance is low. The point estimates capturing the changes in part-time employment are either zero or negative. The most striking feature of the table is that the increase in full-time equivalent employment is driven by gains in full-time employment. Moreover, the larger gains in employment are at the upper quantiles noting that the estimates with the higher degrees of statistical significance also occur at these quantiles.
To illustrate the applicability of our methodology to a bivariate analysis we consider the impact of the increase in the minimum wage on the joint distribution of the full-time and part-time employment levels. The counterfactual and actual distributions are shown in Figure (ref). The figure reveals that the joint distribution has changed due to the increase in the minimum wage and that the distribution of the treated population appears to have shifted downward and to the right. This appears to be a similar movement to that reported in Figure 5 of Torous et al. (2024). This suggests the treatment has changed the relationship between full-time and part-time employment. However, it is difficult to identify whether this merely reflecs the changes in the marginal distributions. It is also difficult to interpret the economic implications of the differences shown in this figure. Accordingly in Table 3 we also present estimates of Kendall's tau and Spearman's correlation index to capture the correlation between these two employment levels. Note that Kendall's tau is calculated as: $$ \hat \tau_d = \frac{2}{n_{11}(n_{11}-1)} \sum_{i=1}^N \sum_{j=i+1}^N G_i G_j T_i T_j\text{sgn}(Y_{di} - Y_{dj}) \text{sgn}(Z_{di}-Z_{dj}), \ \ n_{11} = \sum_{i=1}^N G_i T_i, \ \ d \in \{0,1\}. $$ Spearman's correlation index is calculated as:
\[ \hat \rho_{d} = 1 - \frac{6 \sum_{i=1}^{N} G_i T_i D_i^2}{n_{11} (n_{11}^2-1)}, \quad d \in \{0,1\}, \] where $D_i$ is the difference between the ranks of $Y_d$ and $Z_d$ conditional on $G=1$ and $T=1$ for observation $i$.
The Kendall's tau and the Spearman's correlation index for the treated sample in the second period can be calculated from the observed data. For the counterfactual distribution of the treated sample in the second period when not treated, we first sample from the estimated distribution. That is, we sample a value of $Y_0$ using our estimate of its marginal distribution from above. We then sample $Z_0$ from the conditional distribution of $Z_0 \,|\, Y_0$ which can be obtained using our estimates for the bivariate model.
In the absence of treatment the estimates of Kendall's $\tau$ and Spearman's correlation index for these employment levels are -0.0095 and -0.0101 respectively. Following the increase in the minimum wage, this negative relationship becomes stronger, with the corresponding estimate values of -0.1709 and -0.2402. Moreover, despite the relatively small number of observations the difference in the Spearman's correlation is statistically significant at the 10 percent level. These two estimates of the change in the level of correlation both suggest that the increase in the minimum wage has changed the relationship between full-time and part-time employment. The statistically significant stronger negative correlation is consistent with a greater degree of substitutability between part-time and full-time employment in the presence of the higher minimum wage.
We provide a simple distribution regression based estimator to implement the evaluation of treatment effects in a difference-in-difference setting. As our approach provides counterfactual distributions we are able to explore the impact of the treatment at different quantiles of the distribution of the outcome variable. For both the univariate and multivariate cases we provide the identifying assumption and the associated estimation algorithms. A re-examination of the Card and Krueger (1994) study highlights the utility of various aspects of our approach.
Our analysis can easily be extended to the case of multiple time periods and more than two outcomes. We can also extend our distributional regression framework to use time and unit weights as in the synthetic difference-in-difference estimation method of Arkhangelsky et al. (2021). We leave each of these extensions to future research (e.g. Fern\'andez-Val et al., 2024b).