EconBase
← Back to paper

Distribution Regression Difference-In-Differences

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

171,789 characters · 14 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Distribution Regression Difference-In-Differences

abstractWe provide a simple distribution regression estimator for treatment effects in the difference-in-differences (DiD) design. Our procedure is particularly useful when the treatment effect differs across the distribution of the outcome variable. Our proposed estimator easily incorporates covariates and, importantly, can be extended to settings where the treatment potentially affects the joint distribution of multiple outcomes. Our key identifying restriction is that the counterfactual distribution of the treated in the untreated state has no interaction effect between treatment and time. This assumption results in a parallel trend assumption on a transformation of the distribution. We highlight the relationship between our procedure and assumptions with the changes-in-changes approach of Athey and Imbens (2006). We also reexamine the Card and Krueger (1994) study of the impact of minimum wages on employment to illustrate the utility of our approach.

Keywords: difference in differences, multiple outcomes, distributional and quantile treatment effects

Introduction

The remarkable popularity of the difference-in-difference (DiD) estimator, inspired by an approach to evaluating the impact of policy interventions on economic outcomes introduced by David Card (see, for example, Card 1990, Card and Krueger, 1994), is one of the most striking features of empirical work on treatment and policy effects. While the methodological innovations in this literature (see Arkhangelsky and Imbens, 2024 for a recent review) include the use of constructed control groups, the staggered timing of treatments, and fuzzy rather than sharp designs, the vast majority of the associated empirical work has estimated the mean effect of the treatment on a single economic outcome. This seems somewhat limited and a fuller evaluation of a policy treatment would be based on an examination of the marginal and joint distributions of all outcomes it potentially influences. This paper provides a simple procedure for estimating distributional treatment effects in the presence of a single treatment when the outcomes of interest are potentially multivariate.

An initial methodological innovation focusing on distributional effects in DiD estimation is the changes-in-changes procedure of Athey and Imbens (2006), which estimates the counterfactual distribution of the treated group in the absence of treatment to compare with its observed distribution in the presence of treatment. Torous et al. (2024) extend the Athey and Imbens approach (2006) to the multivariate outcome setting. Other work has adapted DiD estimation to examine the treatment effects at different quantiles of the outcome via the use of quantile regression. This includes, for example, Callaway and Li (2018, 2019). In contrast, Dube (2019), Goodman-Bacon (2021), and Goodman-Bacon and Schmidt (2020) employ conventional DiD estimation to explore the impact of the treatment at different points of the outcome distribution. Other distributional approaches include Kim and Wooldridge (2023) and Biewen, Fitzenberger, and Rümmele (2022). The former proposes an inverse probability weighting based procedure, while the latter employs a distribution regression (DR) approach. In this paper we also adopt a DR approach to constructing counterfactuals. In contrast to Biewen, Fitzenberger, and Rümmele (2022), who construct the counterfactual distributions via linear probability models, we employ non-linear link functions such as probit or logit models. This has a number of advantages, which we discuss below. In addition, we provide the associated identifying conditions required for this form of implementation of DR-DiD.

While DiD has typically been employed to evaluate the treatment effect on a specified economic outcome, there are many instances in which the treatment may affect multiple outcomes. For example, a change in tax rates on earnings of married couples may affect the hours of work of both husbands and wives. An analysis of such a tax change should include the impact on each of the outcomes. However, a richer analysis would not only examine the impact on the respective marginal hours distributions of husbands and wives but also the joint distribution of hours. Alternatively, while evaluations of minimum wage laws typically evaluate their impact on total employment, they may also affect the joint distribution of part-time and full-time employment. We illustrate how this joint effect can be evaluated via the bivariate distribution regression (BDR) approach of Fern\'andez-Val et al. (2024a). This requires that we first estimate the joint distribution by BDR and then construct the appropriate counterfactual. The treatment effects are then obtained via the appropriate comparisons. A recent alternative to this approach is extending the changes-in-changes procedure to multiple outcomes as is done in Torous et al. (2024).

The following section introduces the model and provides an analysis of the univariate case without covariates. We also extend our analysis to include covariates and contrast our approach with the Athey and Imbens (2006) changes-in-changes procedure. Section (ref) extends our analysis to the multiple outcome case and Section (ref) discusses estimation. Section (ref) provides an empirical illustration via an application of our procedure to the data employed in the Card and Krueger (1994) study of the impact of increasing the minimum wage on employment. Section (ref) concludes.

Econometric analysis of the univariate case

Consider the standard DiD design with 2 periods, $T \in \{0,1\}$, and 2 groups, $G \in \{0,1\}$ in which a binary treatment, $D \in \{0,1\}$, is administered only to the treatment group with $G = 1$ in the second period $T=1$. Let $Y_0$ and $Y_1$ denote the potential outcomes under the non-treated and treated statuses. The observed outcome is $Y = Y_0 (1-D) + Y_1 D$, which corresponds to $Y_0$ for both groups at $T=0$, $Y_0$ for $G=0$ at $T=1$, and $Y_1$ for $G=1$ at $T=1$. Note that this implicitly imposes a non-anticipation assumption as we do not distinguish between the outcomes of the treated and non-treated state for $G=1$ in period $T=0$.

We are interested in the distributions of the potential outcomes of the treated at $T=1$, that is $F_{Y_1 \,|\, G, T}(y \,|\, 1,1)$ and $F_{Y_0 \,|\, G, T}(y \,|\, 1,1)$. $F_{Y_1 \,|\, G, T}(y \,|\, 1,1)$ is identified from the observed outcome for $G=1$ at $T=1$, $$ F_{Y_1 \,|\, G, T}(y \,|\, 1,1) = F_{Y \,|\, G, T}(y \,|\, 1,1); $$ whereas $F_{Y_0 \,|\, G, T}(y \,|\, 1,1)$ is not identified without further assumptions.

The distribution of $Y_0$ conditional on $G$ and $T$ can be written as:

equation[equation omitted — 146 chars of source]

where $\Lambda$ is an invertible CDF such as the logistic, normal or uniform, and $y \mapsto $ $(\alpha(y), $ $ \beta(y), \gamma(y), \delta(y))$ is a vector of function-valued parameters.

The representation in (ref) does not make any parametric assumption about the underlying distribution of $Y_0 \,|\, G, T$ since the dummy variable representation within the parentheses on the right-hand side is fully saturated. The parameters of the representation are local as they vary with $y$. To understand why (ref) does not impose any restriction, note that $\alpha(y)$, $\beta(y)$, $\gamma(y)$ and $\delta(y)$ can be defined as:\footnote{See also Wooldridge (2023) equations (2.6) and (2.7).} \[

split[split omitted — 665 chars of source]

\] We make the following identifying assumptions:

assumption[No-interaction] $$\delta(y)=0 \text{ for all } y \in \mathbb{R} \text{ in \eqref{eq:dr}}.$$

Let $\mathcal{Y}_d(G=g, T=t)$ denote the support of $Y_d \,|\, G=g,T=t$, for $d, g, t \in \{0,1\}$. We also assume:

assumption[Support] \[ \mathcal{Y}_0(G=1;T=1) \subseteq \mathcal{Y}_0(G=0;T=1) \cup \mathcal{Y}_0(G=1;T=0) \cup \mathcal{Y}_0(G=0;T=0). \]

Assumption (ref) implies that the distribution of the potential outcome $Y_0$ should not change differently in the second period for the treatment group compared to the control group. That is, we allow a difference between the distributions of the potential outcome $Y_0$ between the treatment and control group, but this difference should be identical in both periods. This is a parallel trend type assumption on a transformation of the distribution and can be written as: \[

split[split omitted — 258 chars of source]

\] This assumption is sensitive to the link function and imposes restrictions on the distribution $F_{Y_0 \,|\, G, T}$ for some link functions. For example, if $\Lambda$ is the identity link used in the linear probability model as in, for example, Almond et al. (2011), Dube (2019), Cengiz et al. (2019), Goodman-Bacon and Smith (2020), Goodman-Bacon (2021) and Biewen et al. (2022), one needs strong requirements in order to satisfy the parallel trends assumption (Blundell et al., 2004 and Wooldridge, 2023) That is, we need restrictions on the tails of the distribution of $F_{Y_0 \,|\, G, T}(y \,|\, 1,0)$, $F_{Y_0 \,|\, G, T}(y \,|\, 0,1)$ and $F_{Y_0 \,|\, G, T}(y \,|\, 0,0)$ to guarantee that $F_{Y_0 \,|\, G, T}(y \,|\, 1,1)$ is between $0$ and $1$. Thus, it requires that $F_{Y_0 \,|\, G, T}(y \,|\, 1,0) \leq 1 + F_{Y_0 \,|\, G, T}(y \,|\, 0,0) - F_{Y_0 \,|\, G, T}(y \,|\, 0,1)$, which might be restrictive at the top of the distribution, and $F_{Y_0 \,|\, G, T}(y \,|\, 1,0) $$\geq F_{Y_0 \,|\, G, T}(y \,|\, 0,0) - F_{Y_0 \,|\, G, T}(y \,|\, 0,1)$, which might be restrictive at the bottom of the distribution.\footnote{These requirements could be used to develop a specification test for the identity link. Roth and Sant'Anna (2023) proposed a test for the sharp hypothesis that $y \mapsto F_{Y_0 \,|\, G, T}(y \,|\, 1,0) + F_{Y_0 \,|\, G, T}(y \,|\, 0,1) - F_{Y_0 \,|\, G, T}(y \,|\, 0,0)$ be weakly increasing, which can be adapted to our setting. We do not pursue this route as we do not encourage the use of the linear probability model.}$^{,}$\footnote {For example, an increase in 0.2 in probability over time might be realistic for the control group when the initial probability was 0.5. However, if treatment group has a probability of, for example, 0.9, in the first period then it is not possible for the common trends assumption to hold.} Link functions such as the normal or logistic CDFs do not require such restrictions since the transformation expands the range of the distribution to the entire real line.

Assumption (ref) cannot be empirically verified but when we have multiple observations in the pre-treatment period, it is possible to examine whether the “parallel trends” assumption holds pre-treatment. Assumption (ref) is a restriction of the support of the counterfactual outcome of $Y_0$ for the treated group in the treated period.

These two assumptions identify $F_{Y_0 \,|\, G, T}(y \,|\, 1,1)$ since:

multline[multline omitted — 508 chars of source]

under Assumption (ref). The support restrictions in Assumption (ref) ensure that the term inside the squared brackets in (ref) is determined. Note that as $\lim_{x \rightarrow \infty}\Lambda(x) = 1$ and $\lim_{x \rightarrow -\infty}\Lambda(x) = 0$, our assumptions are sufficient but not necessary.

We present this identification result in the following lemma:

lemma[Identification with Single Outcome] $y \mapsto F_{Y_0 \,|\, G,T}(y \,|\, 1,1)$ is identified on $y \in \mathbb{R}$ under Assumptions (ref) and (ref).
proof[Proof of Lemma (ref)] The results follow from equation (ref).

In empirical analysis, researchers typically would like to investigate objects that are related to the distributions of the potential outcome variables. One such object is the distributional treatment effect, defined as: \[ \tau(y) := F_{Y_1 \,|\, G,T}(y \,|\, 1,1)(y) - F_{Y_0 \,|\, G,T}(y \,|\, 1,1), \quad y \in \mathbb{R}. \] The distributional treatment effect measures the change in the probability that the outcome is below $y$ as a result of the treatment. Another interesting object is the quantile treatment effect, defined as: \[ \tau^*_q := F^{-1}_{Y_1 \,|\, G,T}(q \,|\, 1,1)(y) - F^{-1}_{Y_0 \,|\, G,T}(q \,|\, 1,1), \quad q \in (0,1). \] The quantile treatment effect measures the difference in the q-th quantile of the outcome variable as a result of the treatment.

Inclusion of Covariates

Including covariates is appealing as the assumption that $\delta(y) = 0$ may be harder to defend when there are differences in the trend between covariates and/or the composition of the treatment group changes over time in terms of observed characteristics; see also Melly and Santangelo (2015). Covariates are easily incorporated into the identification result by conditioning on them and adding an overlapping support assumption. Specifically, let $X$ be a vector of covariates such that the non-interaction assumption holds conditional on $X$; see Assumption (ref). The distribution of $Y_0$ conditional on $G$, $T$ and $X$ can be written:

equation[equation omitted — 168 chars of source]

where $(y,x) \mapsto (\alpha(y,x), \beta(y,x), \gamma(y,x), \delta(y,x))$ is a vector of unspecified functions.

Let $\mathcal{Y}_d(G=g, T=t; X=x)$ denote the support of $Y_d \,|\, G=g,T=t, X=x$. The identifying assumptions with covariates become:

assumption[No-interaction with Covariates] $$\delta(y,X)=0 \text{ almost surely for all } y \in \mathbb{R} \text{ in \eqref{eq:dr_with_cov}.}$$
assumption[Support conditions with Covariates] \[ \mathcal{Y}_0(G=1;T=1;X) \subseteq \mathcal{Y}_0(G=0;T=1;X) \cup \mathcal{Y}_0(G=1;T=0;X) \cup \mathcal{Y}_0(G=0;T=0;X), \] almost surely.

These two assumptions identify $F_{Y_0 \,|\, G, T, X}(y \,|\, 1,1,x)$ since

multline[multline omitted — 551 chars of source]

under the Assumption (ref). The support restrictions in Assumption (ref) ensure that the term between parentheses in (ref) is determined. Note that as $\lim_{x \rightarrow \infty}\Lambda(x) = 1$ and $\lim_{x \rightarrow -\infty}\Lambda(x) = 0$, our assumptions are sufficient but not necessary.

Let $\mathcal{X}_{11}$ denote the support of $X$ conditional on $G=1$ and $T=1$. The following Lemma states that $F_{Y_0,Z_0 \,|\, G,T,X}$ is identified under the previous assumptions.

lemma[Identification with Covariates] Under Assumptions (ref) and (ref), $(y,x) \mapsto $ \\ $F_{Y_0 \,|\, G,T,X}(y \,|\, 1,1,x)$ is identified on $(y,x) \in \mathbb{R}\times \mathcal{X}_{11}$.
proof[Proof of Lemma (ref)] The result follows from equations (ref).

We can then identify $F_{Y_0 \,|\, G, T}(y \,|\, 1,1)$ as:

equation[equation omitted — 172 chars of source]

where $F_{X \,|\, G,T}$ is the distribution of $X$ conditional on $G$ and $T$.

Comparison with Changes-In-Changes

As our proposals provide an alternative approach to the changes-in-changes (CiC) procedure of Athey and Imbens (2006), it is useful to contrast their setup and assumptions with ours. CiC assumes that the outcome of an individual without treatment satisfies the relationship $Y_0 = h(U,T)$ for the treatment and control groups, where $U$ is an unobserved and uniformly distributed random variable. It also assumes that $h$ is strictly increasing in the first term and that the distribution of $U$ is independent of time given the treatment outcome, i.e. $U \perp\!\!\!\perp T \,|\, G$. Finally, the support of $U$ for the treated population should be a subset of those of the untreated population. The final assumption implies in terms of the support of the potential outcomes that: \[

split[split omitted — 137 chars of source]

\] Their second support restriction is less restrictive than ours but we do not need their first support restriction. The previous assumptions identify the quantile function of $F_{Y_0 \,|\, G, T}(y \,|\, 1,1)$ as: \[

split[split omitted — 224 chars of source]

\] where it is assumed that $Y_0$ is continuous with strictly increasing distribution function.

The transformation $\phi$ gives the second period outcome for an individual with an unobserved component $u$ such that $h(u,0) =y$, with $y$ the location at which the distribution function is evaluated (Athey and Imbens, 2006, page 441). Hence, their identification results follow since $\phi$ evaluated in the first period observations of the treatment group is equally distributed as the distribution of the untreated outcome of the treatment group in the second period. Their assumptions imply the transformation $\phi$ that maps quantiles of $Y_0$ from period $0$ to period $1$ is the same for the treatment and control groups. This condition imposes the following restrictions on the coefficients of the representation of the conditional distribution in (ref): $$ \alpha(y) = \alpha(\phi(y)) + \beta(\phi(y)), \quad \gamma(y) = \gamma(\phi(y)) + \delta(\phi(y)) . $$ To see this, note that:

equation[equation omitted — 114 chars of source]

Evaluating (ref) at $g=0$ and applying $F^{-1}_{Y_0 \,|\, G,T}(\cdot \,|\, 0,1)$ to both sides: $$ h(h^{-1}(y,0),1) = F^{-1}_{Y_0 \,|\, G, T}\left( F_{Y_0 \,|\, G, T}(y \,|\, 0,0) \,|\, 0,1\right) =: \phi(y). $$ Replacing $\phi(y)$ back in (ref) and using the representation (ref): $$ \Lambda(\alpha(y) + \gamma(y)g) = \Lambda(\alpha(\phi(y)) + \beta(\phi(y)) + \gamma(\phi(y))g + \delta(\phi(y))g ). $$ The restrictions then follow from equalizing the coefficients in both sides.\footnote{There is only a binding restriction because $\alpha(y) = \alpha(\phi(y)) + \beta(\phi(y))$ holds by definition of $\phi(y)$.} They would complicate estimation in our framework as they involve two different levels of $Y$ and the transformation $\phi$ needs to be estimated.

Comparison with Roth and Sant'Anna (2023)

Roth and Sant'Anna (2023) derive the condition: $$ F_{Y_0 \,|\, G, T}(y \,|\, 1,1) - F_{Y_0 \,|\, G, T}(y \,|\, 1,0) = F_{Y_0 \,|\, G, T}(y \,|\, 0,1) - F_{Y_0 \,|\, G, T}(y \,|\, 0,0), \quad y \in \mathbb{R}, $$ for the parallel trends assumption in expectations: $$ {\mathbb{E}}(Y_0 \,|\, G = 1, T = 1) - {\mathbb{E}}(Y_0 \,|\, G = 1, T = 0) = {\mathbb{E}}(Y_0 \,|\, G = 0, T = 1) - {\mathbb{E}}(Y_0 \,|\, G = 0, T = 0), $$ to be invariant to strictly monotone transformations of $Y_0$. This condition is different from our no-interaction assumption. Indeed, our DR model with no-interaction does not generally satisfy the parallel trends assumption in expectation as: \[

split[split omitted — 218 chars of source]

\] depends on $g$ unless $\Lambda$ is the identity map, or $\beta(y)=0$ (no trend) or $\gamma(y)=0$ (random assignment) for $y \in \mathbb{R}$. Roth and Sant'Anna (2023) show that their condition holds if and only if there are no trends, random assignment or a mixture of both. Our model, however, generally satisfies a different invariance property with respect to strictly monotonic transformations that we specify in {subsection} (ref).

Invariance to Strictly Monotonic Transformations

The DR model in (ref) with no-interaction is invariant to strictly monotonic transformations in the sense that we specify here. If $Y_0$ follows the DR model and satisfies the no-interaction assumption, then $\tilde Y_0 = h(Y_0)$ also follows the DR model and satisfies the no-interaction assumption for any strictly monotonic transformation $h$. To see this result note that if $h$ is strictly increasing: $$ F_{\tilde Y_0 \,|\, G, T, X}(\tilde y \,|\, g,t,x) = \Lambda(\alpha(h^{-1}(\tilde y)) + \beta(h^{-1}(\tilde y)) t + \gamma(h^{-1}(\tilde y))g ) = \Lambda(\tilde \alpha(\tilde y) + \tilde \beta(\tilde y) t + \tilde \gamma(\tilde y)g ), $$ where $\tilde y \mapsto h^{-1}(\tilde y)$ is the inverse function of $y \mapsto h(y)$, $\tilde \alpha = \alpha \circ h^{-1}$, $\tilde \beta = \beta \circ h^{-1}$ and $\tilde \gamma = \gamma \circ h^{-1}$. A similar argument applies when $h$ is strictly decreasing. Unlike the parallel trends in expectation, the no-interaction or parallel trends in distribution is invariant to strictly monotonic transformations.\footnote{The distributional approach of Kim and Wooldridge (2023) also satisfies this property.}

Multiple Outcomes

Some settings may feature multiple outcomes that are potentially affected by the treatment. In these situations, we might be interested not only in how each of the outcomes is affected by the treatment, but also in how the relationship between the outcomes is affected by the treatment. For this, it is necessary to identify the joint distribution of the potential outcomes with and without treatment. We now consider a setting with two outcomes $Y$ and $Z$ and we focus on comparing features of the joint distribution of the potential outcomes with the treatment, $Y_1$ and $Z_1$, and the joint distribution of the potential outcomes without the treatment, $Y_0$ and $Z_0$, for the treated group $G=1$ in the post-treatment period $T=1$. For the sake of illustration we consider two measures of dependence. Namely, Spearman's and Kendall's rank correlation.

Let $F_{Y_d,Z_d \,|\, G,T}$ be the joint distribution of $Y_d$ and $Z_d$ conditional on $G$ and $T$, and $F_{Y_d \,|\, G,T}$ and $F_{Z_d \,|\, G,T}$ be the corresponding marginals. Spearman's rank correlation between $Y_d$ and $Z_d$, $d \in \{0,1\}$, can be expressed:

multline*[multline* omitted — 337 chars of source]

and Kendall's rank correlation between $Y_d$ and $Z_d$, $d \in \{0,1\}$, can be expressed:

multline*[multline* omitted — 203 chars of source]

where we assume that $Y_d$ and $Z_d$ are continuous random variables to obtain the expressions on the right hand side.

As in the univariate case, $F_{Y_1,Z_1 \,|\, G,T}(y,z \,|\, 1,1)$ is identified by the joint distribution of the observed outcomes, $F_{Y,Z \,|\, G,T}(y,z \,|\, 1,1)$, whereas $F_{Y_0,Z_0 \,|\, G,T}(y,z \,|\, 1,1)$ is not identified from the data. To analyze identification, we use a variation of the local Gaussian representation (LGR) of a bivariate distribution from Chernozhukov, Fernand\'ez-Val and Luo (2018). Let $\Phi$ denote the Gaussian distribution function and $\Phi_2(\cdot,\cdot;\rho)$ denote the distribution of the bivariate standard normal with parameter $\rho$. Moreover, $\Lambda$ is, again, a strictly increasing cumulative distribution function. As we show in Section (ref), there is a benefit of using the logistic link function in our univariate analysis. Accordingly, we employ this in our empirical analysis for estimating both the univariate and bivariate effects.

lemma[LGR with non-Normal Marginals] The joint distribution of two random variables $Y$ and $Z$ conditional on $X$ can be represented by: $$ F_{Y,Z \,|\, X}(y,z \,|\, x)(y,z \,|\, x) \equiv \Phi_2(\Phi^{-1}(\Lambda(\mu_{Y \,|\, X}(y \,|\, x))), \Phi^{-1}(\Lambda(\mu_{Z \,|\, X}(y \,|\, x))); \rho_{Y,Z \,|\, X}(y,z \,|\, x)), $$ for all $y,z,x$, where $\mu_{Y \,|\, X}(y \,|\, x) = \Lambda^{-1}(F_{Y\,|\, X}(y \,|\, x))$, $\mu_{Z \,|\, X}(y \,|\, x) = \Lambda^{-1}(F_{Z\,|\, X}(z \,|\, x))$, and $\rho_{Y,Z \,|\, X}(y,z \,|\, x))$ is the unique solution in $\rho$ to the equation: $$ F_{Y,Z \,|\, X}(y,z \,|\, x)(y,z \,|\, x) = \Phi_2(\Phi^{-1}(F_{Y\,|\, X}(y \,|\, x)(y,z \,|\, x)), \Phi^{-1}(F_{Z\,|\, X}(z \,|\, x)(y,z \,|\, x)); \rho). $$
proofThe proof is identical to the proof of Lemma 2.1 of Chernozhukov, Fernand\'ez-Val and Luo (2018) using: $$ \Phi^{-1}(\Lambda(\mu_{Y \,|\, X}(y \,|\, x))) = \Phi^{-1}(F_{Y\,|\, X}(y \,|\, x)) $$ and $$ \Phi^{-1}(\Lambda(\mu_{Z \,|\, X}(z \,|\, x))) = \Phi^{-1}(F_{Z\,|\, X}(z \,|\, x)). $$

The difference between Lemma (ref) and the LGR of Chernozhukov, Fernand\'ez-Val and Luo (2018) is that the marginals are represented by a general link rather than Gaussian links, that is: $$ F_{Y\,|\, X}(y \,|\, x)(y \,|\, x) \equiv \Lambda(\mu_{Y \,|\, X}(y \,|\, x)), \quad F_{Z\,|\, X}(z \,|\, x) \equiv \Lambda(\mu_{Z \,|\, X}(z \,|\, x)). $$ By the LGR, $F_{Y_0,Z_0 \,|\, G,T}$ can be expressed as:

multline[multline omitted — 237 chars of source]

where $\mu_{Y_0 \,|\, G,T}(y \,|\, g,t) = \alpha_Y(y) + \beta_Y(y) t + \gamma_Y(y) g + \delta_Y(y) gt$, $\mu_{Z_0 \,|\, G,T}(y \,|\, g,t) = \alpha_Z(z) + \beta_Z(z) t + \gamma_Z(z) g + \delta_Z(z) gt$, and $\rho_{Y,Z \,|\, G,T}(y,z \,|\, g,t) = \alpha_{Y,Z}(y,z) + \beta_{Y,Z}(y,z) t + \gamma_{Y,Z}(y,z) g + \delta_{Y,Z}(y,z) gt$. In the LGR, the marginals are represented by: $$ F_{Y_0 \,|\, G,T}(y \,|\, g,t) = \Lambda(\alpha_Y(y) + \beta_Y(y) t + \gamma_Y(y) g + \delta_Y(y) gt), $$ and $$ F_{Z_0 \,|\, G,T}(z \,|\, g,t) = \Lambda(\alpha_Z(z) + \beta_Z(z) t + \gamma_Z(z) g + \delta_Z(z) gt). $$ We make the following identifying assumptions with respect to the distribution function in (ref).

assumption[Bivariate No-interaction] $$\delta_Y(y) = \delta_Z(z) = \delta_{Y,Z}(y,z) = 0 \text{ for all } (y,z) \in \mathbb{R}^2 \text{ in \eqref{eq:bdr}}. $$

Let $\mathcal{YZ}_d(G=g, T=t)$ denote the support of $(Y_d,Z_d) \,|\, G=g,T=t$, for $d, g, t \in {0,1}$. We also assume:

assumption[Bivariate Support] \[ \mathcal{YZ}_0(G=1;T=1) \subseteq \mathcal{YZ}_0(G=0;T=1) \cup \mathcal{YZ}_0(G=1;T=0) \cup \mathcal{YZ}_0(G=0;T=0). \]
lemma[Identification with Two Outcomes] $(y,z) \mapsto F_{Y_0,Z_0 \,|\, G,T}(y,z \,|\, 1,1)$ is identified on $\mathbb{R}^2$ under Assumptions (ref) and (ref).
proof[Proof of Lemma (ref)] Under the assumptions of the Lemma, $\mu_{Y_0 \,|\, G,T}(y \,|\, g,t) = \alpha_Y(y) + \beta_Y(y) t + \gamma_Y(y) g$, $\mu_{Z_0 \,|\, G,T}(y \,|\, g,t) = \alpha_Z(z) + \beta_Z(z) t + \gamma_Z(z) g$, and $\rho_{Y,Z \,|\, G,T}(y,z \,|\, g,t) = \alpha_{Y,Z}(y,z) + \beta_{Y,Z}(y,z) t + \gamma_{Y,Z}(y,z) g$. The parameters $\alpha_Y(y)$, $\beta_Y(y)$, $\gamma_Y(y)$, $\alpha_Z(z)$, $\beta_Z(z)$, and $\gamma_Z(z)$ are identified from the marginals of $Y$ and $Z$, by Lemma (ref). The parameter $\alpha_{Y,Z}(y,z)$ is identified as the solution in $\alpha$ to: $$ F_{Y,Z \,|\, G,T}(y,z \,|\, 0,0) = \Phi_2(\alpha_Y(y) + \beta_Y(y) t + \gamma_Y(y) g, \alpha_Z(z) + \beta_Z(z) t + \gamma_Z(z) g ; \alpha). $$ This solution exists and is unique because the RHS is strictly increasing in $\alpha$. The parameters $\beta_{Y,Z}(y,z)$ and $\gamma_{Y,Z}(y,z)$ are identified similarly as the solutions in $\beta$ and $\gamma$ of: $$ F_{Y,Z \,|\, G,T}(y,z \,|\, 0,1) = \Phi_2(\alpha_Y(y) + \beta_Y(y) t + \gamma_Y(y) g, \alpha_Z(z) + \beta_Z(z) t + \gamma_Z(z) g ; \alpha_{Y,Z}(y,z) + \beta). $$ and $$ F_{Y,Z \,|\, G,T}(y,z \,|\, 1,0) = \Phi_2(\alpha_Y(y) + \beta_Y(y) t + \gamma_Y(y) g, \alpha_Z(z) + \beta_Z(z) t + \gamma_Z(z) g ; \alpha_{Y,Z}(y,z) + \gamma). $$ Finally, \[ \begin{split} F_{Y_0,Z_0 \,|\, G,T}(y,z \,|\, 1,1) & = \Phi_2(\alpha_Y(y) + \beta_Y(y) + \gamma_Y(y), \alpha_Z(z) \beta_Z(z) + \gamma_Z(z); \alpha_{Y,Z}(y,z) + \\ & \beta_{Y,Z}(y,z) + \gamma_{Y,Z}(y,z)). \end{split} \]

Covariates can be incorporated in a similar fashion as the univariate case. In particular, the LGR of $F_{Y_0,Z_0 \,|\, G,T,X}$ becomes:

multline[multline omitted — 262 chars of source]

where $\mu_{Y_0 \,|\, G,T,X}(y \,|\, g,t,x) = \alpha_Y(y,x) + \beta_Y(y,x) t + \gamma_Y(y,x) g + \delta_Y(y,x) gt$, $\mu_{Z_0 \,|\, G,T,X}(y \,|\, g,t,x) = \alpha_Z(z,x) + \beta_Z(z,x) t + \gamma_Z(z,x) g + \delta_Z(z,x) gt$, and $\rho_{Y,Z \,|\, G,T,X}(y,z \,|\, g,t,x) = \alpha_{Y,Z}(y,z,x) + \beta_{Y,Z}(y,z,x) t + \gamma_{Y,Z}(y,z,x) g + \delta_{Y,Z}(y,z,x) gt$.

Let $\mathcal{YZ}_d(G=g, T=t; X=x)$ denote the support of $(Y_d,Z_d) \,|\, G=g,T=t, X=x$. The identifying assumptions with covariates become:

assumption[Bivariate No-interaction with Covariates] $$\delta_Y(y,X) = \delta_Z(z,X) = \delta_{Y,Z}(y,z,X) = 0 \text{ almost surely for all } (y,z) \in \mathbb{R}^2 \text{ in \eqref{eq:bdr_with_cov}.}$$
assumption[Bivariate Support with Covariates] \begin{multline*} \mathcal{YZ}_0(G=1;T=1;X) \subseteq \\ \mathcal{YZ}_0(G=0;T=1;X) \cup \mathcal{YZ}_0(G=1;T=0;X) \cup \mathcal{YZ}_0(G=0;T=0;X), \end{multline*} almost surely.
lemma[Identification with Two Outcomes and Covariates] Under Assumptions (ref) and (ref), $(y,z) \mapsto F_{Y_0,Z_0 \,|\, G,T,X}(y,z \,|\, 1,1,x)$ is identified on $\mathbb{R}^2 \times \mathcal{X}_{11}$.

The marginalized distribution $F_{Y_0,Z_0 \,|\, G,T}(y \,|\, 1,1)$ is then identified by $$ F_{Y_0,Z_0 \,|\, G,T}(y,z \,|\, 1,1) = \int_{\mathcal{X}_{11}} F_{Y_0,Z_0 \,|\, G,T,X}(y,z \,|\, 1,1,x) \mathrm{d} F_{X \,|\, G,T}(x \,|\, 1,1). $$

Estimation

Univariate Case

Assume we have a sample $\{(Y_{i}, X_{i}, G_{i}, T_{i}): 1\leq i \leq N\}$ of $(Y, X, G, T)$. For estimation, we replace the functions $(y,x) \mapsto (\alpha(y,x), \beta(y,x), \gamma(y,x))$ in (ref) by semiparametric linear indexes leading to the DR model for the conditional distribution:

equation[equation omitted — 184 chars of source]

where $p_{\alpha}(x)$, $p_{\beta}(x)$ and $p_{\gamma}(x)$ are vectors including the covariates and their transformations, and $y \mapsto (\alpha(y), \beta(y), \gamma(y))$ is a vector of function-valued parameters.

We implement the DR DiD estimator via the sequence of logit models at each point of the distribution of the outcome variable (Foresi and Peracchi, 1995, Chernozhukov, Fernandez-Val and Melly, 2013). We choose logit because it is the canonical link for binary outcomes allowing for pooled estimation of the distributions of the potential outcomes with and without the treatment (Wooldridge, 2023). Accordingly, we estimate the DR model for the observed outcomes on all observations including those with $D_i = 1$:

equation[equation omitted — 211 chars of source]

where $p_{\alpha}(x)$, $p_{\beta}(x)$, $p_{\gamma}(x)$ and $p_{\theta}(x)$ are vectors including a constant as the first component, and transformations of the covariates, and $\mathcal{Y}$ is a finite grid on $\mathbb{R}$. Let $I_i^y := 1(Y_i \leq y)$ and $\bar I_i^y = 1-I_i^y$.

algorithm[algorithm omitted — 2,046 chars of source]

By the properties of the logistic link, the estimator of $F_{Y_1 \,|\, G, T}(y \,|\, 1,1)$ is identical to the empirical distribution of $Y$ conditional on $G=1$ and $T=1$, $$ \hat F_{Y_1 \,|\, G, T}(y \,|\, 1,1) \equiv \frac{1}{N_{11}} \sum_{i=1}^N G_i T_i \ 1(Y_i \leq y). $$ Note that this estimator is therefore invariant to the specification of $p_{\theta}(x)$. We set $p_{\theta}(x)=1$ to speed up computation. Note that the definition of the estimator for the inverse of the distribution function is appropriate as we rearranged the estimates $y \mapsto \hat F_{Y_d \,|\, G, T}(y \,|\, 1,1)$ on $y \in \mathcal{Y}$, $d \in \{0,1\}$.

Our algorithm has an advantage over the alternative of estimating $p_{\alpha}$, $p_{\beta}$, $p_{\gamma}$ using only those observations for which $D_i = 0$ via direct estimation of (ref). For example, the distributional treatment effect, \[ \int_{\mathcal{X}_{11}} \left[ F_{Y_1|G,T,X}(y|1,1,X) - F_{Y_0|G,T,X}(y|1,1,X) \right] \mathrm{d}F_{X \,|\, G,T}(x \,|\, 1,1), \] equals the average partial effect of $D_i$ for the logit model used for distribution regression. Estimates and standard errors for these effects are routinely reported by many Statitical software packages (Wooldridge, 2023).

Bivariate Case

Assume we have a sample $\{(Y_{i}, Z_{i}, X_{i}, G_{i}, T_{i}): 1\leq i \leq N\}$ of $(Y, Z, X, G, T)$. For estimation, as in the univariate case, we replace the functions in $\mu_{Y_0 \,|\, G,T,X}$, $\mu_{Z_0 \,|\, G,T,X}$ and $\rho_{Y,Z \,|\, G,T,X}$ by semiparametric generalized linear indexes leading to a bivariate distribution regression (BDR) model:

equation[equation omitted — 155 chars of source]
equation[equation omitted — 155 chars of source]

and

equation[equation omitted — 179 chars of source]

where $p_{\alpha}(x)$, $p_{\beta}(x)$, $p_{\gamma}(x)$, $q_{\alpha}(x)$, $q_{\beta}(x)$, $q_{\gamma}(x)$, $r_{\alpha}(x)$, $r_{\beta}(x)$ and $r_{\gamma}(x)$ are vectors including the covariates and their transformations, and $h(u) = \text{arctanh}(u)$ is the Fisher transformation that enforces $\rho_{Y,Z \,|\, G,T,X}$ to lie in $[-1,1]$.

We estimate all the parameters of $F_{Y_0,Z_0 \,|\, G,T}(y,z \,|\, 1,1)$ using the bivariate distribution regression estimator of Fernandez-Val et al. (2024a). We employ an imputation method that combines the parameter estimates from the sample of the first period for both groups and the sample of the second period for the untreated group, with the sample of the covariates in the second period for the treated group. The distribution $F_{Y_1,Z_1 \,|\, G,T}(y,z \,|\, 1,1)$ is estimated using the empirical distribution of $Y$ and $Z$ in the second period for the treated group. Algorithm (ref) describes the estimation procedure. Let $\mathcal{Y}$ and $\mathcal{Z}$ be finite grids on $\mathbb{R}$, $I_i^y := 1(Y_i \leq y)$, $\bar I_i^y = 1-I_i^y$, $J_i^z := 1(Z_i \leq z)$, $\bar J_i^z = 1-J_i^z$.

algorithm[algorithm omitted — 2,648 chars of source]

Estimators of the functionals of the joint distributions of potential outcomes such as Spearman's and Kendall's rank correlation coefficients can be constructed using the plug-in principle.

Bootstrap Inference

The estimators described in Algorithms (ref) and (ref) can be applied to panel and repeated cross-section data. Here we describe a weighted bootstrap algorithm to perform inference on functions of the distributions of potential outcomes designed for panel data. We focus on this case because it is relevant for our empirical application below.

To describe the procedure, we need to introduce an indicator $ID_i$, $i = 1, \ldots, N$, for the units in the panel. For example, if the sample is sorted by unit and time period, $ID = (1,1,2,2,\ldots,n,n)$, where $n=N/2$. The following algorithm describes the weighted bootstrap procedure to construct joint confidence bands for the distributions of the potential outcomes with and without the treatment in the univariate case. Inference for functionals of the distributions and for the bivariate case can be performed using similar algorithms.

algorithm[algorithm omitted — 3,054 chars of source]
remark[Empirical Bootstrap] Empirical bootstrap can be implemented by drawing the weights in step 1 from a multinomial distribution with values $1,\ldots,n$ and equal probabilities $1/n$.

Empirical application

We illustrate our approach through a re-examination of data used in Card and Krueger (1994), hereafter CK, investigation of the impact on an increase in the minimum wage on the level of employment. In April 1992 New Jersey increased its minimum wage from the Federal level of 4.25 dollars per hour to 5.05 dollars per hour. CK investigated the impact of this increase on the change in the level of "full-time equivalent employment", measured as the sum of the number of full-time employees plus half the number of part-time employees, in New Jersey fast food restaurants. They use a DiD estimation strategy in which the control group comprises a group of comparable fast food restaurants in the region of Pennsylvania bordering New Jersey. CK concluded that this particular increase in the minimum wage led to a small increase in the level of full-time equivalent employment. This result produced a large and important related literature on the impact of minimum wages on employment.

We employ our procedure using the CK data to investigate the impact of this increase in the minimum wage on the level of full-time, part-time and full-time equivalent employment respectively. We estimate the model using the 409 observations available in CK data set. 80.9 percent of these are observations are treated establishments in New Jersey. We acknowledge that the data set is relatively small and that this is likely to have implications for the level of statistical significance of the results. However, as our objective is to illustrate our approach in a well-known setting, we prefer to work with a data set which has been frequently used and is well understood (see, for a recent example, Torous et al. 2024) rather than providing our own original application. Note that for the models which include covariates, the additional variables are four dummy variables for franchise type and a dummy variable indicating that the establishment is company owned.

The respective actual and counterfactual distributions are reported in Figures (ref)-(ref). These figures are based on the models which include the additional covariates. While there are some differences in the distributions for full-time employment, indicating an increase in employment for establishment sizes above the 1st quartile, the larger differences appear in Figure (ref) which captures the increases in full-time employment. This figure suggests gains at all quantiles. Figure (ref) is suggestive of some small reductions in part-time employment at some quantiles.

The results of the quantile treatment effects for the univariate analyses are reported in Tables (ref) and (ref) noting that those in the former include the covariates while the latter does not. As the results are generally similar, we focus only on Table 1. The DiD estimate of the mean effect on full-time equivalent employment is 2.65 and this is statistically significant at the 10 percent level. An examination of the table reveals that this mean effect is driven by an increase in full-time employment as there is very weak statistical evidence of a small reduction in mean part-time employment. The level of statistical significance probably reflects the small sample size.

Tables (ref) and (ref) also indicate that the effect of the minimum wage increase is different depending on the size of the establishment. For example, at smaller establishments the effect on full-time equivalent employment is negative and this reflects a reduction in these establishments' levels of part-time employment. However, there appears to be growth in full-time equivalent employment at the larger establishment sizes noting that the level of statistical significance is low. The point estimates capturing the changes in part-time employment are either zero or negative. The most striking feature of the table is that the increase in full-time equivalent employment is driven by gains in full-time employment. Moreover, the larger gains in employment are at the upper quantiles noting that the estimates with the higher degrees of statistical significance also occur at these quantiles.

To illustrate the applicability of our methodology to a bivariate analysis we consider the impact of the increase in the minimum wage on the joint distribution of the full-time and part-time employment levels. The counterfactual and actual distributions are shown in Figure (ref). The figure reveals that the joint distribution has changed due to the increase in the minimum wage and that the distribution of the treated population appears to have shifted downward and to the right. This appears to be a similar movement to that reported in Figure 5 of Torous et al. (2024). This suggests the treatment has changed the relationship between full-time and part-time employment. However, it is difficult to identify whether this merely reflecs the changes in the marginal distributions. It is also difficult to interpret the economic implications of the differences shown in this figure. Accordingly in Table 3 we also present estimates of Kendall's tau and Spearman's correlation index to capture the correlation between these two employment levels. Note that Kendall's tau is calculated as: $$ \hat \tau_d = \frac{2}{n_{11}(n_{11}-1)} \sum_{i=1}^N \sum_{j=i+1}^N G_i G_j T_i T_j\text{sgn}(Y_{di} - Y_{dj}) \text{sgn}(Z_{di}-Z_{dj}), \ \ n_{11} = \sum_{i=1}^N G_i T_i, \ \ d \in \{0,1\}. $$ Spearman's correlation index is calculated as:

\[ \hat \rho_{d} = 1 - \frac{6 \sum_{i=1}^{N} G_i T_i D_i^2}{n_{11} (n_{11}^2-1)}, \quad d \in \{0,1\}, \] where $D_i$ is the difference between the ranks of $Y_d$ and $Z_d$ conditional on $G=1$ and $T=1$ for observation $i$.

The Kendall's tau and the Spearman's correlation index for the treated sample in the second period can be calculated from the observed data. For the counterfactual distribution of the treated sample in the second period when not treated, we first sample from the estimated distribution. That is, we sample a value of $Y_0$ using our estimate of its marginal distribution from above. We then sample $Z_0$ from the conditional distribution of $Z_0 \,|\, Y_0$ which can be obtained using our estimates for the bivariate model.

In the absence of treatment the estimates of Kendall's $\tau$ and Spearman's correlation index for these employment levels are -0.0095 and -0.0101 respectively. Following the increase in the minimum wage, this negative relationship becomes stronger, with the corresponding estimate values of -0.1709 and -0.2402. Moreover, despite the relatively small number of observations the difference in the Spearman's correlation is statistically significant at the 10 percent level. These two estimates of the change in the level of correlation both suggest that the increase in the minimum wage has changed the relationship between full-time and part-time employment. The statistically significant stronger negative correlation is consistent with a greater degree of substitutability between part-time and full-time employment in the presence of the higher minimum wage.

table[table omitted — 1,287 chars of source]
table[table omitted — 1,340 chars of source]
figure[figure omitted — 25,802 chars of source]
figure[figure omitted — 25,790 chars of source]
figure[figure omitted — 24,025 chars of source]
figure[figure omitted — 28,846 chars of source]
table[table omitted — 923 chars of source]
table[table omitted — 959 chars of source]

Conclusion

We provide a simple distribution regression based estimator to implement the evaluation of treatment effects in a difference-in-difference setting. As our approach provides counterfactual distributions we are able to explore the impact of the treatment at different quantiles of the distribution of the outcome variable. For both the univariate and multivariate cases we provide the identifying assumption and the associated estimation algorithms. A re-examination of the Card and Krueger (1994) study highlights the utility of various aspects of our approach.

Our analysis can easily be extended to the case of multiple time periods and more than two outcomes. We can also extend our distributional regression framework to use time and unit weights as in the synthetic difference-in-difference estimation method of Arkhangelsky et al. (2021). We leave each of these extensions to future research (e.g. Fern\'andez-Val et al., 2024b).

\thinspace\ References

description• Almond, D., H.W. Hoynes, and D.W. Schanzenbach (2011), \textquotedblleft Inside the war on poverty: the impact of food stamps on birth outcomes\textquotedblright, Review of Economics and Statistics 93, 387--403. • Arkhangelsky, D. and G. Imbens (2024), \textquotedblleft Causal Models for longitudinal and panel Data: a survey\textquotedblright, The Econometrics Journal 27, C1-61. • \textsc{Arkhangelsky, D. S. Athey, D.A. Hirshberg, G.W. Imbens, and S. Wager} (2021), \textquotedblleft Synthetic difference-in-differences, \emph{American Economic Review} \textbf{111}, 4088–18. • \textsc{Athey, S. and G.D. Imbens} (2006), \textquotedblleft Identification and inference in nonlinear difference-in-differences models\textquotedblright, \emph{Econometrica} \textbf{74} 431--97. • \textsc{Biewen, M., M. R\"ummele, and B. Fitzenberger} (2022), \textquotedblleft Using distribution regression Difference-in-Differences to evaluate the effects of a minimum wage introduction on the distribution of hourly wages and hours worked, working paper, N\"urnberg. • \textsc{Blundell, R., C. Meghir, M. Costa Dias and J. Van Reenen} (2004), \textquotedblleft Evaluating the employment impact of a mandatory job search program\textquotedblright, \emph{Journal of the European Economic Association} \textbf{2},569--606. • \textsc{Callaway, B. and T. Li} (2019), \textquotedblleft Quantile treatment effects in difference in differences models with panel data\textquotedblright \emph{Quantitative Economics} \textbf{10}, 1579--1618. • \textsc{Card, D.} (1990), \textquotedblleft The Impact of the Mariel Boatlift on the Miami Labor Market\textquotedblright \emph{Industrial and Labor Relations Review,} \textbf{43}, 245–-57. • \textsc{Card, D. and A.B. Krueger} (1994), \textquotedblleft Minimum Wages and Employment: A Case Study of the Fast-Food Industry in New Jersey and Pennsylvania\textquotedblright \emph{American Economic Review} \textbf{84}, 772–-93. • \textsc{Cengiz, D., A. Dube, A. Lindner, and B. Zipperer} (2019), \textquotedblleft The effect of minimum wages on low-wage jobs \textquotedblright, \emph{Quarterly Journal of Economics} \textbf{134}, 1405--54. • \textsc{Chernozhukov, V., I. Fern\'{a}ndez-Val, and Melly B.} (2013), \textquotedblleft Inference on counterfactual distributions\textquotedblright , \emph{Econometrica}, \textbf{81}, 2205--68. • \textsc{Chernozhukov, V., I. Fern\'{a}ndez-Val, and S. Luo}{ \ } (2019), \textquotedblleft Distribution regression with sample selection, with an application to wage decompositions in the UK", working paper, MIT, Cambridge (MA). • \textsc{Chernozhukov, V. I. Fern\'andez-Val, B. Melly, and K. Wüthrich} (2020), \textquotedblleft Generic inference on quantile and quantile effect functions for discrete outcomes\textquotedblright, \emph{Journal of the American Statistical Association} \textbf{115}, 123--37. • \textsc{Dube, A.} (2019), \textquotedblleft Minimum wages and the distribution of family incomes\textquotedblright, \emph{American Economic Journal: Applied Economics} \textbf{11}, 268–-304. • \textsc{Fern\'{a}ndez-Val I., J. Meier, A. van Vuuren and F. Vella} (2024a) \textquotedblleft Bivariate distribution regression with an application to intergenerational mobility\textquotedblright, working paper, Boston University. • \textsc{Fern\'{a}ndez-Val I., J. Meier, A. van Vuuren and F. Vella} (2024b) \textquotedblleft Distributional synthetic difference-in-differences\textquotedblright, working paper, Boston University. • \textsc{Foresi, S. and F. Peracchi} (1995), “The conditional distribution of excess returns: an empirical analysis”, \emph{Journal of the American Statistical Association}, \textbf{90}, 451--66. • \textsc{Goodman-Bacon, A.} (2021), \textquotedblleft The long-run effects of childhood insurance coverage: medicaid implementation, adult health, and labor market outcomes, \emph{American Economic Review} \textbf{111}, 2550-93. • \textsc{Goodman-Bacon, A. and L. Schmidt} (2020), \textquotedblleft Federalizing benefits: The introduction of supplemental security income and the size of the safety net.\textquotedblright, \emph{Journal of Public Economics} \textbf{185}, 104174. • \textsc{Kim, D. and J.M. Wooldridge} (2024), \textquotedblleft Difference-in-differences estimator of quantile treatment effect on the treated \textquotedblright, \emph{Journal of Business and Economic Statistics}, forthcoming. • \textsc{MaCurdy, T.} (2015), \textquotedblleft How effective is the minimum wage at supporting the poor?\textquotedblright, \emph{Journal of Political Economy} \textbf{123}, 497--545. • \textsc{Malesky, E.J., C.V. Nguyen, and A. Trahn} (2014), \textquotedblleft The Impact of recentralization on public services: A difference-in-differences analysis of the abolition of elected councils in Vietnam\textquotedblright, \emph{American Political Science Review} \textbf{108} 144--68. • \textsc{Melly, B. and Santangelo} (2015), \textquotedblleft The changes-in-changes model with covariates\textquotedblright, working paper, Bern University. • \textsc{Roth, J. and P.H.C. Sant'Anna} (2023), \textquotedblleft When Is parallel trends sensitive to functional form?\textquotedblright, \emph{Econometrica} \textbf{91} 737--47. • \textsc{Torous, W., F. Gunsilius, and P. Rigollet} (2024), \textquotedblleft An optimal transport approach to estimating causal effects via nonlinear difference-in-differences, working paper\textquotedblright, University of California, Berkeley. • \textsc{Wooldridge, J.M.} (2023), \textquotedblleft Simple approaches to nonlinear difference-in-differences with panel data\textquotedblright, \emph{Econometric Journal} \textbf{26} C31--66. • \textsc{Williams, O. D. and J.E. Grizzle} (1972), \textquotedblleft Analysis of contingency tables having ordered response categories\textquotedblright, \emph{Journal of the American Statistical Association} \textbf{67}, 55–63.