Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
141,482 characters · 37 sections · 85 citation commands
Evaluating the Impact of Regulatory Policies on Social Welfare in Difference-in-difference Settings
\pagenumbering{gobble}
\setcounter{page}{0}
\pagenumbering{arabic}
Government's regulatory role and its impact on social welfare has been a critical question for economists. These regulatory policies often restrict the budget or choice sets for certain agents in the market by imposing floors or quotas, such as minimum wages, minimum/maximum working time, wage floors for different occupation groups as well as action, reporting and notification thresholds in environmental monitoring. Those types of policies tend to induce behavioral responses that can generate mass points in the outcome of interest. For instance, an important question in the labor economics literature is the effect of an increase or introduction of minimum wages on low-wage jobs or overall employment, see for instance CK1994, NeumarkWascher2008, Cengizetal2019, among many others. The figure below (taken from Cengizetal2019) illustrates that an increase in the minimum wage will shift jobs that were previously paying below the minimum wage $MW$, and then will create “excess jobs” at and slightly above the minimum wage.
This figure also shows the heterogeneous effect of such a policy; it is expected to only affect the wage of low-wage workers and not have an effect on the upper tail of the distribution. In sum, those types of policies have two main features. First, the potential outcomes of interest are likely to exhibit some mass points. Second, the causal effect of the policy is expected to affect only a part of the distribution of the outcomes of interest. As a result, to adequately analyze the impact of these policies, a distributional treatment effect analysis is key, as in Cengizetal2019 for instance; see, also, Almondetal2011,assuncao2022. Furthermore, measuring the impact of such policies on social welfare requires recovering the counterfactual distribution of the outcome of interest.
While these types of policies are widely studied in economics, the existing econometrics methods are not necessarily adequate to recover distributional causal effects in these settings. In the presence of data before and after a new policy, one of the most widely used techniques to assess its impact is the difference-in-differences (DiD) method. Its main drawbacks, however, are two-fold: (1) it does not identify the counterfactual distribution, (2) it is not invariant to monotonic transformations. While there are several methods to identify the counterfactual distribution in difference-in-difference settings AtheyImbens2006,bonhommesauder2011,CallawayLi2019,HavnesMogstad2015, to the best of our knowledge, the distributional DiD and changes-in-changes (CiC) are the only two approaches that are invariant to monotonic transformations.\footnote{The distributional DiD method relies on a parallel trends assumption in the cdfs as opposed to the expectations HavnesMogstad2015,RothSantanna2021. }
RothSantanna2021 show that distributional DiD requires that the distribution of the untreated potential outcome is independent of policy adoption, is stationary across time (within each group), or consists of a mixture of two subpopulations each obeying one of the two restrictions. Such conditions are unlikely to be valid for the policy evaluation questions we are interested in. Indeed, the independence assumption (random assignment) is implausible in our context since the decision to implement a new minimum wage policy is a response to the unsatisfactory features of the pre-policy outcome distribution, such as large wage inequalities, high proportion of workers under poverty, etc. When the policy is not randomly assigned, the validity of the distributional DiD essentially rests on the stationarity assumption, which is restrictive in many practical settings.\footnote{The stationarity assumption can be tested using the control group. RothSantanna2021 provide a sharp specification test of the validity of the distributional DiD assumption in general.}
While the CiC approach introduced in the seminal work by AtheyImbens2006 can accommodate endogenous policy (treatment) assignment as well as time-varying potential outcome distributions, their identification result does not apply to the case where the potential outcomes exhibit some mass points (mixed distributions), as in Figure (ref).\footnote{Mass points are common for a wide range of economic outcomes resulting from censoring bottomcoding1,topcoding1 or bunching bunching3,bunching4,bunching2,bunching5,bunching6,bunching7,bunching8.} In fact, AtheyImbens2006 introduce the CiC approach for either continuous or discrete outcomes that are monotonic (time-varying) functions of a scalar unobservable with a time-invariant distribution across time. In sum, the CiC approach introduced in AtheyImbens2006 should not be applied to evaluate the policies described above.
The current paper provides an alternative, unifying identification result that applies to any type of outcome distribution, is invariant to monotonic transformations, allows for endogeneity of the policy assignment, and does not restrict the evolution of the marginal distribution, nor treatment effect heterogeneity. Our identification result exploits the stability of the dependence (copula) between treatment assignment and the untreated potential outcome across time without imposing restrictions on the structural function that generates the potential outcomes.
Exploiting our copula stability (CS) assumption, we provide a unifying partial identification result for the counterfactual distribution of the treatment group. We then extend our analysis to the case where multiple pre-treatment periods are available. In this case, we show that if copula stability holds for multiple pre-treatment periods, then our multi-period CS bounds exploit the information from the pre-treatment periods to provide tighter bounds.\footnote{We use multi-period CS bounds to refer to CS bounds that use multiple pre-treatment periods.} The presence of multiple pre-treatment periods also allows us to provide a testable restriction of our model assumptions. We demonstrate our theoretical results numerically in Section (ref).
Our CS bounds apply to any type of outcome distribution, whether it is continuous, mixed, or discrete. They shrink to the point-identification result in AtheyImbens2006 for continuous outcomes. Indeed, we show that in this case our copula stability assumption is equivalent to the CiC conditions. For discrete outcomes, we show that our copula stability assumption can be compatible with an underlying production function featuring multi-dimensional unobserved heterogeneity, whereas the CiC bounds for discrete outcomes require a scalar unobservable. For mixed outcomes, we demonstrate that a na\"ive implementation of the CiC approach may lead to a point-estimand that does not coincide with the true counterfactual, whereas our CS bounds will include it.\footnote{We refer to this implementation as na\"ive since AtheyImbens2006 did not provide identification results for mixed outcomes. Nonetheless, an empirical researcher might ignore the mixed-nature of this outcome and implement their point-identification result. }
We also examine the connection between our main identifying assumption and the parallel trends assumption required by DiD. The parallel trends assumption can be equivalently stated as a covariance stability assumption. It is specifically a time invariance assumption on the covariance between treatment assignment and the untreated potential outcome, whereas our assumption maintains the stability of the copula between these two variables. As a result, there are several differences between our copula stability assumption and covariance stability (parallel trends). First, the parallel trends assumption restricts the joint variability of treatment assignment and the untreated potential outcome over time, whereas our copula stability assumption only restricts their dependence structure. Second, while the parallel trends assumption restricts the evolution of the marginal distribution of the untreated potential outcome across time, copula stability does not restrict the evolution of the marginal distribution, nor treatment effect heterogeneity. Last but not least, parallel trends is not invariant to monotonic transformations except under strong conditions on heterogeneity RothSantanna2021. These conditions specifically rule out the existence of a subpopulation that selects into treatment based on unobservables and exhibits changes in its potential outcome distribution. By contrast, our copula stability condition does not rule out such a subpopulation.
Since the motivation behind policies, such as increases in the legal minimum wage, is often to reduce inequality and/or target a specific part of the outcome distribution, we introduce a broad class of social welfare treatment effect parameters that can accommodate the policymaker's objective. While this class includes the average treatment effect on the treated (ATT) as a special case, the ATT corresponds to a social welfare function that is inequality-neutral and gives equal weight to all individuals in the population. As a result, if a policymaker is averse to inequality, then the ATT would be an inadequate causal parameter to judge the policy's effectiveness. In general, the social welfare function adequate to evaluate a specific regulatory policy can be highly context-specific and may depend on the policymaker's preference and/or objective.\footnote{Please see the discussion in Bergeretal2022 which illustrates how the quantitative analysis of the effect of the minimum wage could highly differ depending on the social welfare weights, which are usually unknown to the researcher.} We therefore introduce a broad class of treatment effect parameters that take into account the policy objectives. This broad class specifically includes the class of generalized Gini social welfare functions Mehran1976,Weymark1981. These social welfare functions can take into account measures of inequality by putting higher weight on individuals with lower-ranked outcomes. In addition, we include a class of parameters that can capture the welfare of individuals at the lower tail or a specific interquantile range of the distribution. Bounds on these social welfare treatment effect parameters can be easily computed using our bounds on the counterfactual distribution. We illustrate the usefulness of this broad class of parameters and compare it to the ATT in the context of our empirical application examining the impact of a minimum wage policy (Section (ref)).
We organize the rest of the paper as follows. Section (ref) introduces the analytical framework and presents our main identification results. Section (ref) introduces the class of social welfare treatment effect parameters. Section (ref) provides an empirical illustration examining the impact of minimum wage increases on the wage distribution revisiting Cengizetal2019.
A comparison between our identifying assumption and some of the related approaches in the literature is warranted. bonhommesauder2011 exploit a separable model of the potential outcome to identify the entire counterfactual distribution of the treatment group in a DiD design. By relying on restrictions on the outcome model, it is therefore similar in spirit to the identification approach in AtheyImbens2006. botosarumuris2023 propose identification of counterfactual parameters for a class of semiparametric panel models, whereas our approach can accommodate both repeated cross-sections and panel data and is fully nonparametric. CallawayLi2019 also provide a fully nonparametric identification result exploiting a copula stability restriction on different objects than the ones used in this paper. They require the copula between changes and levels of the untreated potential outcome to be invariant across time for the treatment group, while our copula stability assumption does not restrict the evolution of the marginal distribution of the untreated potential outcome (Remark (ref)). Furthermore, our approach can be applied to repeated cross-sections or panel data and only requires two time periods, whereas CallawayLi2019 require at least three periods of panel data. wooldridge2023simple proposes alternative parallel trends assumptions that are more suitable for binary, fractional and count outcome data. The approach in wooldridge2023simple requires the specification of a parametric transformation model of a linear index for each type of outcome and point-identifies the average treatment effect on the treated, whereas our approach applies to any outcome, is fully nonparametric and partially identifies the counterfactual distribution.
Finally, this paper contributes to a strand in the microeconometrics literature that relies on copula theory. For cross-sectional settings with exogenoeus regressors, rothe2012 provides identification results for partial distributional effects, which hold the copula of the covariates constant, but vary their marginal distributions. mourifie2015 relies on copula theory to provide sharp bounds on the average treatment effect in a binary triangular system. arellanobonhomme2017 propose a method to correct for sample selection in quantile models, where the conditional copula of the error terms in the outcome and selection equations is a key ingredient in their approach.
Following Abadie2005, we consider the following potential outcomes model:\footnote{Note that this model implicitly assumes that there are no anticipatory effects of the treatment, that is, $Y_{00}=Y_{01}=Y_0$.}
where $Y_{t}$ denotes the observed outcome at period $t$ and $Y_{td}$ denotes the potential outcome at period $t\in\{0,1\}$ and treatment status $d\in\{0,1\}$. In the two-group, two-period case, $D$ denotes both group membership and the treatment status in period 1.
We use the following shorthand notation: $p\equiv \mathbb P(D=1)$, $q=1-p$, $\overline{\mathbb R}\equiv\mathbb R \cup \{-\infty, \infty\}$, $Ran H \equiv\{ H(y):y \in \mathbb R\}$, $\overline{\operatorname{Ran}} F\equiv Ran F \cup \{\inf Ran F, \sup Ran F\}$, and $Dom H$ denotes the domain of the function $H$. We consider the following mappings $Q^{\mathbb T,-}_X: [0,1] \rightarrow \mathbb T$, and $Q^{\mathbb T,+}_X: [0,1] \rightarrow \mathbb T$, where $Q^{\mathbb T,-}_X(u)\equiv \inf \{x \in \mathbb T \cup \{\infty\}: F_X(x)\geq u\}$ for all $u \in [0,1]$, $Q^{\mathbb T,+}_X(u)\equiv \sup \{x \in \mathbb T \cup \{-\infty\}: F_X(x)\leq u\}$ for all $u \in [0,1]$. We call $Q^{\mathbb T,+}_X$ and $Q^{\mathbb T,-}_X$ generalized quantile functions whenever $F_X(.)$ is a well-defined cumulative distribution function (cdf). We denote by $\mathbb F$ the space of all well-defined cdfs. $Supp X=\mathbb X$ denotes the support of $X$, and $\mathbb X_{s|d}$ denotes the support of $X_s|D=d$ for $d \in \{0,1\}$. Finally, we define $F_X(x-)\equiv \mathbb{P}(X<x)$.
Our main identification result relies on restrictions imposed on the dependence structure across time. To do so, we rely on copula theory. Copulas are functions that enable us to separate the marginal distributions from the (scale-free) dependence structure of a given multivariate distribution. In our context, we are interested in the subcopula between the untreated potential outcome and group membership across time. Working with copulas in our case will allow us to avoid restricting the type of marginal distribution of the potential outcomes as well as its heterogeneity across time. To fix ideas, let us first provide a formal definition of the (sub)copula.
A copula is a special case of a subcopula where $S_1=S_2=[0,1]$. For a fixed $v \in S_2$, $u \mapsto C(u,v)$ is usually called the horizontal subcopula. The link between the joint distribution and the subcopula has been established by the well-known Sklar (1959) theorem, which provides the following lemma when applied to our context.
To provide intuition for the role of the horizontal subcopula at $q$, it is helpful to divide each side of Equation\ (ref) by $q$, which yields the following for $y\in \overline{\mathbb{R}}$
Now, let us assume that the copula is strictly increasing in its first argument such that its inverse $C_{Y_{t0},D}^{-1}(\cdot~;q)$ is well-defined.\footnote{Note that a horizontal copula is by definition Lipschitz continuous.} We can then show that $C_{Y_{t0},D}^{-1}(\cdot~;q)$ is the main ingredient in the rank mapping between the treatment and control group's untreated potential outcome distribution in period $t$, which we denote by $\Gamma_t(\cdot)$:
where $\Gamma_t(u)\equiv \frac{1}{p}\left( C_{Y_{t0},D}^{-1}\left(uq;q \right)- uq \right)$ for $u\in RanF_{Y_{t0}|D=0}$. The mapping governs the relationship of the rank that a given value $y$ has in the control group's distribution of the untreated potential outcome $F_{Y_{t0}|D=0}$ (factual at each period) onto its rank in the treatment group's distribution of the untreated potential outcome, $F_{Y_{t0}|D=1}$.
Next, we introduce our main assumption.
In the following, we will refer to Assumption (ref) as “copula stabilty” for brevity, but we emphasize that it only requires the stability of the horizontal copula between $Y_{t0}$ and $D$ at $q$, $C_{Y_{t0},D}(u,q)$ for $u\in[0,1]$. There are multiple advantages to our copula stability assumption. First, it is invariant to strictly monotonic transformations. Specifically, for any right-continuous function $g$, that is strictly increasing on $\mathbb Y_{td}$, we have:\footnote{See Embrechts_al2013 (Embrechts_al2013, Proposition 4(2)) for a formal proof.} $$C_{g(Y_{td}),D}(u,q)=C_{Y_{td},D}(u,q) \;\; \forall u \in RanF_{Y_{td}}.$$ Second, it does not impose any restrictions on the variability of the marginal distribution $F_{Y_{t0}}$ across time. Last, but not least, it does not restrict the type of marginal distribution $F_{Y_{t0}}$, whether it is continuous, discrete or mixed.
Assumption (ref) is the key assumption behind our identification approach. It implies that the rank mapping $\Gamma_t(\cdot)$ is stable across periods, i.e. $ \Gamma_1(\cdot)=\Gamma_0(\cdot)\equiv \Gamma(\cdot)$. In the presence of multiple pre-treatment periods, $\Gamma(\cdot)$ can be recovered from each pre-treatment period. As a result, analogous to pre-trend testing in difference-in-differences designs, the time-invariance of $\Gamma(\cdot)$ can also be tested as we demonstrate in Section (ref).
Given the wide use of difference-in-differences, it is also helpful to clarify the relationship between our copula stability assumption and the parallel trends assumption. The parallel trends assumption can be equivalently rewritten as a covariance stability assumption as we show in Appendix (ref),
This equivalence result provides, first, an intuition for why the parallel trends assumption is not invariant to a monotonic transformation since the covariance is not invariant to monotonic transformations. Second, it allows us to observe that the parallel trends assumption jointly restricts the evolution of the marginal distribution of $Y_{t0}$ across time and the dependence between $Y_{t0}$ and $D$. Unlike the parallel trends assumption, our copula stability assumption does not constrain the evolution of the marginal distribution across time, yet it relies only on the stability of the horizontal copula that governs the relationship between $Y_{t0}$ and $D$. As can be seen in the following equation, the two assumptions are non-nested in general:
Indeed, copula stability may hold while $Cov(Y_{10},D)\neq Cov(Y_{00},D)$ because $F_{Y_{10}}\neq F_{Y_{00}}$; and the covariance stability may hold while the copula stability is violated.
In the following, we provide several examples to illustrate the restrictions imposed by our key assumption, and how it compares to some existing assumptions.
The above example demonstrates the copula stability assumption in the context of selection on the gains from the treatment. We next consider selection on untreated potential outcomes. This example shows that copula stability requires comonotonicity between the untreated potential outcomes in the pre- and post-treatment periods.
We next consider selection on time-varying shocks, an example that will not be compatible with copula stability in general.
Next, we proceed to our second identifying assumption, which requires the strict monotonicity of the horizontal copula.
While Assumption (ref) is less critical for our bounding approach, it allows us to simplify the expression of our bounds. It is essentially a restriction on the type of dependence between the potential outcomes and group membership. Many well-known parametric classes of copulas satisfy this assumption, e.g. Frank, Gumbel, Joe, or Gaussian copulas among many others. It excludes, however, extreme types of dependence captured by the Fr\'echet-Hoeffding copula bounds, i.e. $C(u,v)=\min\{u,v\}$ and $C(u,v)=\max\{u+v-1,0\}$. It is worth noting that this assumption is implied by some support conditions on the potential outcome distributions, as we show in the following result.
The main implication of the above lemma is that for continuous potential outcome distributions, we have $Ran F_{Y_{t0}}=[0,1]$, and the strict monotonicity of the copula (Assumption (ref)) is implied by a condition on the support of $Y_{t0}$, $\mathbb Y_{t0|1} \subseteq \mathbb Y_{t0|0}$. That is, the support of the untreated potential outcome of the treatment group is included in the support of the untreated potential outcome of the control group. The support condition imposed in AtheyImbens2006 on the scalar unobservable in the CiC model implies this support condition on the untreated potential outcome.
We next state our main identification result:
Theorem (ref) provides a general (partial) identification result on the counterfactual distribution of the treatment group for any type of potential outcome variables (discrete, continuous, or mixed). Our result neither imposes any restriction on the heterogeneity of potential outcomes within a period nor across periods. We specifically do not impose restrictions on individual treatment effects, $Y_{11}-Y_{10}$, or the evolution of the distribution of the untreated potential outcome across time, $F_{Y_{t0}}, t \in\{0,1\}$. The formal proof is relegated to Appendix (ref). The derived bounds may look involved since we aim to provide a general formulation that covers any type of distribution and want to ensure that our bounds are indeed right-continuous.\footnote{As recognized by AtheyImbens2006, their upper bound in the discrete outcome case may be left-continuous, and therefore may not satisfy the properties of a cdf.} The bounds simplify for some special cases as we will illustrate in Corollary (ref) below.
The intuition behind our (partial) identification result is very simple and can be summarized as follows: In the first period, we identify the joint distribution $\mathbb P(Y_{00}\leq y, D=0)$ and both marginal distributions, $\mathbb P(Y_{00}\leq y)$ and $q$. Using the Sklar result, we can recover the horizontal subcopula $C_{Y_{00},D}(u,q)$ on $RanF_{Y_{0}}$, and thereby the rank mapping $\Gamma(\cdot)$ on $Ran F_{Y_0|D=0}$. Then, since we assume the rank mapping to be stationary across time, we can then carry it over from the pre-treatment period to the post-treatment period to recover the treatment group's distribution of the untreated potential outcome, $F_{Y_{10}|D=1}$, as follows:
The main reason behind the partial identification is that in the first period we recover the subcopula $C_{Y_{00},D}(\cdot,q)$ only on $Ran F_{Y_0}$ ($\Gamma(\cdot)$ only on $Ran F_{Y_0|D=0}$), and we do not know the rank mapping outside this range. We provide a graphical illustration of these functions as well as our bounds in the context of a minimum-wage numerical example in Appendix (ref).
In the case of continuous potential outcomes, $Ran F_{Y_{0}}=[0,1]$, our bounds shrink to a point because the pre-treatment period allows us to recover the entire rank mapping that we carry over to the post-treatment period, as we show in the following corollary of Theorem (ref).
The proof of this corollary is in Appendix (ref). Corollary (ref) recovers the point-identification result obtained in AtheyImbens2006. AtheyImbens2006 provide (partial) identification results for two types of potential outcomes relying on different assumptions for each of the two cases: (i) continuous outcomes that are strictly monotonic in a scalar unobservable, (ii) discrete outcomes that are monotonic in a scalar unobservable. By contrast, Theorem (ref) establishes a unifying identification result for any type of outcome under consideration. In addition to the connection to our identification result, there is a link between the CiC assumptions and our copula stability condition for continuous outcomes. We provide details on this connection and compare the two identification approaches in Section (ref).
In this section, we characterize our bounds in the presence of multiple pre-treatment periods. Suppose we have the following model with $T_0+1$ pre-treatment periods:
We impose the following stability restriction on the horizontal copula at $q$ over multiple pre-treatment periods $t=-T_0,\dots,0$.
The following theorem generalizes Theorem (ref) to the multiple-period case under Assumption (ref). Corollary (ref) then provides testable restrictions of our model assumptions.
We illustrate the arguments in Theorem (ref) and Corollary (ref) in a numerical example motivated by our minimum wage setting in the presence of multiple pre-treatment periods in Section (ref).
Here we illustrate the CS bounds with two pre-treatment periods as well as the testable restrictions in the context of a minimum-wage numerical example. Suppose that both treatment and control groups have a pre-existing minimum wage set at $c_{0}$ in the pre-treatment periods ($t=-1,0$). In the post-treatment period ($t=1$), the minimum wage increases for the treatment group to $c_{1}$. We consider two cases: (i) all model assumptions hold (Figure (ref)), (ii) all assumptions except copula stability hold (Figure (ref)).\footnote{In Appendix (ref), we demonstrate a third case, where copula stability holds, while the strict monotonicity of the horizontal copula is violated. This case demonstrates that we can detect violations of our model assumptions with only one pre-treatment period.}
Figure (ref) demonstrates that when copula stability holds for multiple pre-treatment periods, it can have significant gain in terms of identification as the multi-period CS bounds point-identifies the counterfactual distribution on a larger portion of its support in Panel (c) relative to Panels (a) and (b). Figure (ref)(f) provides our model testable restriction, specifically $\Delta(y)\leq 0$, which holds in this case. Furthermore, Panels (d) and (e) of Figure (ref) present $C_{Y_{t0},D}$ and $\Gamma_t$, respectively, for $t=-1,0$, which are equal on the intersection of their respective ranges.
Next, we demonstrate the case where copula stability only holds for $t\in\{0,1\}$, but not $t\in\{-1,1\}$. In Figure (ref), Panel (a) shows that using the pre-treatment period $t=-1$ only to construct the CS bounds yields bounds that do not include the counterfactual, whereas Panel (b) shows that the counterfactual is included in the CS bounds with pre-treatment period $t=0$ only. When considering the CS bounds using both pre-treatment periods in Figure (ref)(c), we note that the CS lower bound is greater than the CS upper bound, and our model testable restriction is violated as indicated by Figure (ref)(f). Relatedly, Figures (ref)(d) and (ref)(e) demonstrate that the mappings $C_{Y_{t0},D}(\cdot,q)$ and $\Gamma_{t}(\cdot)$, respectively, are not equal for $t=-1,0$, indicating a violation of copula stability.
In this section, we elaborate on the connection between our copula stability assumption and the CiC conditions in AtheyImbens2006. We first show the equivalence between copula stability and the CiC conditions for continuous outcome distributions. Second, while the identification results in AtheyImbens2006 do not account for mixed outcomes, a researcher might still rely on their estimand. Here, we demonstrate that a na\"ive implementation of the CiC approach leads to a point/bound estimand that might not include the true counterfactual, whereas our CS bounds will. Finally, for discrete outcomes, we demonstrate using an analytical example that copula stability can be compatible with multi-dimensional unobserved heterogeneity, whereas the CiC conditions require unobserved heterogeneity to be uni-dimensional.
The following result demonstrates that the CiC conditions for continuous, strictly increasing outcome distributions are equivalent to our copula stability assumption. In Appendix (ref), we demonstrate how this result extends to all continuous outcomes. For other outcome distributions, this equivalence does not hold in general.
The proof of this claim is in Appendix (ref). The main intuition behind it is that for this class of distributions we can write $Y_{t0}=Q_{Y_{t0}}^{\mathbb{R},-}(U_{t0})$, where $U_{t0}=F_{Y_{t0}}(Y_{t0})\sim \mathcal{U}[0,1]$. As a result, the marginal distribution of $U_{t0}$ is stable across time by construction and the stability of the copula between $U_{t0}$ and $D$ is necessary and sufficient for the stability of $U_{t0}|D$, which is the conditional time invariance assumption in AtheyImbens2006. Its equivalence to our copula stability assumption follows from the invariance of the copula under strictly monotonic transformations.
Here, we demonstrate that for mixed outcomes the CiC point/bound estimand may not cover the true counterfactual distribution in the context of the numerical minimum-wage example in Section (ref).
The CiC bounds in the discrete case are defined for any $s \in \mathbb{Y}_{1|0}$ as follows for $t\in\{-1,0\}$,
In the example illustrated in Figure (ref), we have $\mathbb{Y}_{t|0} = \mathbb{R}^+$, and $Y_t|D=0$ has a strictly increasing cdf in $\mathbb{R}^+$. Then the following simplifications hold:
and
where the inequality becomes strict at points of discontinuity.
More importantly, we can see that $F_{t,\text{CiC}}^{\text{LB}}(s)=F_{t,\text{CiC}}^{\text{UB}}(s)$, since $Q_{Y_t|D=0}^{\mathbb{Y}_{t|0},+}(u)=Q_{Y_t|D=0}^{\mathbb{Y}_{t|0},-}(u)$ for $u\in[0,1]$. However, this CiC point estimand is different from the true counterfactual of interest $F_{Y_{10}|D=1}$, as shown in Figure (ref).
Therefore, in this case, our bounds contain the CiC (point/bound) estimands and the true counterfactual \[F_{t,\text{CiC}}^{\text{UB}}(s)\neq F_{Y_{10}|D=1}(s), \text{ where } \{F_{t,\text{CiC}}^{\text{UB}}(s),F_{Y_{10}|D=1}(s)\}\in [F_t^{\text{LB}}(s), F_t^{\text{UB}}(s)].\]
In sum, in this mixed-outcome example, if the researcher ignores the discontinuity and applies the CiC point estimand or applied the CiC bounds for the discrete case, their estimand will not cover the true counterfactual, as shown in Figure (ref).
For the case of discrete outcomes, the following example illustrates that our identifying assumption can be compatible with multi-dimensional unobserved heterogeneity, whereas the CiC conditions require scalar unobserved heterogeneity.
While our assumption accommodates a broader class of binary outcome models than the CiC model assumption, it does not necessarily yield tighter bounds. As illustrated in Figure (ref) in the Online Appendix, both approaches produce the same bounds in the discrete-outcome case.
Building on our unifying, partial identification result for the counterfactual distribution, we provide a class of policy-relevant parameters that quantify the impact of policy on social welfare in the entire population, subpopulations in the lower tail of the distribution or over any interquantile range of the distribution. In general, when a policymaker decides to implement a new policy such as an increase in the legal minimum wage or legal minimum working time, she expects the policy to have a specific social welfare impact. The social welfare function used by the policymaker is not necessarily known to the researcher, however. For instance, the policymaker may consider social welfare functions that put more weight on specific subpopulations, such as lower-income individuals, or considers only social welfare functions with specific properties like social welfare functions that respect the Pigou-Dalton principle of transfers\footnote{The Pigou-Dalton principle states that a transfer of income from a higher-ranked individual to a lower-ranked individual that does not change their ranks is always desirable.} or the rank-dependent social welfare functions introduced by Mehran1976 (Mehran1976).\footnote{See Aabergeetal2013 (Aabergeetal2013) for a detailed discussion.}
As we clarify below, the widely used average treatment effect on the treated (ATT) corresponds to the case where the policymaker is inequality-neutral. If the policymaker is averse to inequality, however, the ATT would not be an adequate causal parameter to measure the impact of the policy or judge its effectiveness.
For this particular reason, we propose a class of parameters of interest that measure the causal effect of a particular policy in terms of a social welfare function,
where $SW_{\omega}(F_X)=\int_{0}^{1}\omega(\tau)Q^{\mathbb R,-}_{X}(\tau)$ denotes the social welfare function associated with a specific distribution $F_X$, and $\omega(\tau) \in [0,1]$ is a weighting function. This social welfare function can be alternatively viewed as a weighted average of the outcomes of individuals $i$ where the weights depend on the rank of $X_i$, $SW_\omega=\int{X_i\omega(Rank(X_i))}di$ KitagawaTetenov2021. Since the social welfare function essentially weights different quantiles of the distribution, the choice of the functional form of the weighting function relates to the inequality aversion of the policymaker and the extent thereof. We next consider several examples of weighting functions and discuss the properties of the social welfare functions they imply.
Before we proceed, it is important to emphasize that, while in many applications where measuring inequality is a concern, the outcome $Y$ is typically income or wages, our framework allows $Y$ to denote other outcomes as well as functions of different outcomes, such as consumption, income and/or human capital. Our $SWTT_{\omega}$ is also a generalization of the quantile treatment effect parameter discussed in Abadieetal2002, Firpo2007, and FrohlichMelly2008.
The class of generalized Gini social welfare functions is the class of rank-dependent, equality-minded social welfare functions which satisfy the Pigou-Dalton principle of transfers and is given by
where $\Lambda(\cdot):[0,1] \mapsto [0,1]$ is a convex, non-increasing, and non-negative function with boundary conditions $\Lambda(0)=1$ and $\Lambda(1)=0$. This class admits the equivalent representation as a weighted sum of quantiles with weighting function $\omega(\tau)=\frac{\partial(1-\Lambda(\tau))}{\partial \tau}$,
As a result, the class of social welfare treatment effect parameters we introduce include this class as a special case. We proceed to present two important special cases of this class of social welfare functions, specifically the utilitarian and Gini social welfare functions.
When $\omega(\tau)=1$, we have $SW_{\omega}(F_X)=\int_{0}^{1}Q^{\mathbb R,-}_{X}(\tau)d\tau$ $=\mathbb E[X].$ This corresponds to the additive welfare function and in this case our proposed parameter boils down to the ATT, i.e. $SWTT_{\omega}=ATT$. The ATT is therefore the appropriate parameter if the policymaker weights subpopulations at different quantiles of the distribution equally.
When $\omega(\tau)=2(1-\tau)$, we have $SW_{\omega}(F_X)=\int_{0}^{1}2(1-\tau)Q^{\mathbb R,-}_{X}(\tau)d\tau$ $=\mathbb E[X]\left(1-I_{Gini}(F_X)\right),$ where $I_{Gini}(F_X)\equiv \frac{\int_{0}^{1}(2\tau-1)Q^{\mathbb R,-}_{X}(\tau)d\tau}{\mathbb E[X]}$ is the widely used Gini inequality index, see Sen1974. $SW_{\omega}(F_X)$ reflects the trade-off between the mean and (in)equality in the distribution $F_X$. The product $\mathbb E[X]I_{Gini}(F_X)$ is a measure of the loss in social welfare due to inequality in the distribution $F_X$. In that case, $SWTT_{\omega}$ captures the impact of the policy using the Gini social welfare function, see BlackorbyDonaldson1978 and Weymark1981. In other words, if the policymaker implements the policy in order to reduce the level of inequality measured by the Gini index, this parameter is the most adequate to judge the impact of this policy.
In many cases, when it is possible to do so, most inequality-averse policymakers like to rank distribution functions consistently with second-degree dominance. For instance, we say $F_{Y_{11}|D=1}$ second-order dominates $F_{Y_{10}|D=1}$ if and only if:
for all $u \in [0,1]$ and holds strictly for some $u$. In this special case, we have $\omega(\tau)=\mathbbm{1}\{\tau \leq u\}$. It is possible, however, that the observed and counterfactual distribution cannot be ranked using this criterion. Furthermore, the policy's objective may be to reduce inequality in a specific part of the distribution. We therefore consider the following quantile-specific Gini social welfare functions.
In the Gini social welfare function discussed above, we assume that the policymaker is interested in the inequality of the whole population. Some policies may be concerned with reducing inequality up to specific quantiles of the distribution, such as minimum-wage policies Dube2019,Cengizetal2019. To quantify the impact of the policy on lower-tail quantiles, we extend the quantile-specific lower-tail Gini social welfare measures introduced in Aabergeetal2013 (Aabergeetal2013) for continuous distributions to any type of distribution in order to accommodate the possibility of discontinuities resulting from censoring or bunching. To do so, we introduce the random variable $X^u=Q_X^{\mathbb{R},-}(V)$, where $V\sim \mathcal{U}[0,u]$ for $u\in(0,1]$.\footnote{For $u\in Ran F_X$, $F_{X^u}(x)=\mathbb P(X\leq x|X\leq Q_X^{\mathbb{R},-}(u))$ for any $x\leq Q_X^{\mathbb{R},-}(u)$, thereby yielding the same truncated random variable introduced in Aabergeetal2013(Aabergeetal2013). For $u\notin Ran F_X$, $X^u$ remains a well-defined random variable.} We relegate the derivations relevant to this section to Appendix (ref).
With this definition of $X^u$, we can show that the lower-tail Gini social welfare function can be decomposed into $\mathbb E[X^u]$ and the Gini coefficient associated with $F_{X^u}$ as follows $$\int_{0}^{1}\frac{2}{u^2}(u-\tau)\mathbbm{1}\{\tau \leq u\}Q^{\mathbb R,-}_{X}(\tau)d\tau=\mathbb E[X^u]\left(1-I_{Gini}(F_{X^u})\right),$$ where $I_{Gini}\left(F_{X^u}\right)\equiv\frac{\int_{0}^{1}(2\tau-u)\mathbbm{1}\{\tau\leq u\}Q^{\mathbb R,-}_{X}(\tau)d\tau}{u^2\mathbb E[X^u]}$ is the lower-tail Gini coefficient at $u$ defined in Aabergeetal2013. Therefore, $SWTT_{\omega}$ with $\omega(\tau)=\frac{2}{u^2}(u-\tau)\mathbbm{1}\{\tau \leq u\}$ yields the following,
and is interpreted as the Quantile-$u$ lower tail Gini social welfare treatment effect on the treated.
Since policies may target other parts of the distribution, such as the upper tail, we can generalize these quantile-specific social welfare treatment effect measures to any range of quantiles $[\underline{u},\overline{u}]$ a researcher may be interested in. Specifically, let $\underline{u}\in[0,1]$, $\overline{u}\in[0,1]$, $\underline{u}<\overline{u}$, $V\sim \mathcal{U}[\underline{u},\overline{u}]$, and $X^{\underline{u},\overline{u}}=Q_X^{\mathbb{R},-}(V)$. A derivation of $F_{X^{\underline{u},\overline{u}}}$ is relegated to Appendix (ref). Now by letting $\omega(\tau)=\frac{2}{(\overline{u}-\underline{u})^2}(\overline{u}-\tau)\mathbbm{1}\{\underline{u}<\tau\leq \overline{u}\}$, we obtain the Gini social welfare function specific to the quantile range $[\underline{u},\overline{u}]$, $$SW_{\omega}(\underline{u},\overline{u})=\int_0^1\frac{2}{(\overline{u}-\underline{u})^2}(\overline{u}-\tau)\mathbbm{1}\{\underline{u}<\tau\leq \overline{u}\}Q_X^{\mathbb{R},-}(\tau)d\tau=\mathbb E[X^{\underline{u},\overline{u}}](1-I_{Gini}(F_{X^{\underline{u},\overline{u}}})),$$ where $\mathbb E[X^{\underline{u},\overline{u}}]\equiv\int_{\underline{u}}^{\overline{u}}Q_X^{\mathbb{R},-}(\tau)d\tau$ and $I_{Gini}\left(F_{X^{\underline{u},\overline{u}}}\right)\equiv \frac{\int_0^1(2\tau-\underline{u}-\overline{u})\mathbbm{1}\{\underline{u}<\tau\leq \overline{u}\}Q_X^{\mathbb{R},-}(\tau)d\tau}{(\overline{u}-\underline{u})^2\mathbb E[X^{\underline{u},\overline{u}}]}$.\footnote{This definition extends the upper tail Gini coefficient to any quantile range $[\underline{u},\overline{u}]$.} The interquantile Gini social welfare treatment effect on the treated over $[\underline{u},\overline{u}]$ is given by
In this section, we illustrate the CS bounds by revisiting the minimum wage study by Cengizetal2019. This application demonstrates the usefulness of the class of policy-relevant parameters we introduce to examine the impact of the minimum wage increase. In particular, the lower-tail quantile social welfare treatment effect estimates allow us to zoom into the lower tail of the distribution, where we expect the minimum wage to have an impact. Overall, our CS bounds document proportionately larger impacts on the Gini social welfare in the lowest part of the distribution, where the minimum wage increase led to increase in the lower-tail mean and Gini social welfare. We also find that the distributional DiD exhibits violations of monotonicity in the lower tail of the distribution and is therefore not suitable for this application.
This empirical illustration highlights two practical advantages of our approach. First, our CS bounds relieve practitioners from having to take a stance on the support of the outcome of interest. Second, our multi-period CS bounds combine information from multiple pre-treatment periods to tighten the bounds on the parameters of interest and to simultaneously test the model assumptions.
Cengizetal2019 examine 138 prominent state-level minimum wage increases between 1979 and 2016 using the individual-level NBER-merged Outgoing Rotation Group Earnings Data of the Current Population Survey. Their goal is to examine the impact of the policy on the wage distribution around the minimum wage, as illustrated in Figure (ref). In order to make the empirical illustration of the multi-period CS bounds succinct, we focus on two pre-treatment periods, 2010 and 2011, and one post-treatment period, 2015, and examine the distributional impact of a nontrivial minimum wage increase of \$0.25 or more.\footnote{Note that starting 2009, the federal minimum has been \$7.25, so a minimum wage increase of \$0.25 or more constitutes an increase of more than 3%. This definition of the treatment variable was also used in the empirical illustration in RothSantanna2021.} For the purpose of this empirical illustration, we focus on the subgroup of states that had a pre-treatment minimum wage of \$8 or higher. We report the results for the remaining states in Appendix (ref).
Table (ref) presents the summary statistics for hourly wage of both treatment and control groups in all three periods we consider. For both subgroups, the summary statistics show that the mean and standard deviation is different across treatment and control groups within the same year as well as within groups before and after the treatment.
In order to estimate the CS bounds on the counterfactual, we rely on Lemma (ref) to re-write the lower bound in a manner that admits straightforward numerical computation, specifically for $y\in\mathbb{Y}_{10|1}$ and for a given pre-treatment period $t$
$F_{Y_{10}|D=1}^{LB,t}(y)$ and $F_{Y_{10}|D=1}^{UB,t}(y)$ are estimated by their sample analogues, $\widehat{F}_{Y_{10}|D=1}^{LB}(y)$ and $\widehat{F}_{Y_{10}|D=1}^{UB}(y)$, respectively, by replacing $F_X$ and $Q_X^{\mathbb{R},-}$ by their empirical counterparts, $\widehat{F}_X$ and $\widehat{Q}_X^{\mathbb{R},-}$, respectively.
The distributional DiD and CiC point estimators of $F_{Y_{10}|D=1}(y)$ are given by
The CS bounds on the counterfactual as well as the observed factual distribution $\widehat{F}_{Y_1|D=1}$ can then be used to obtain the following sample analogues of the lower and upper bounds on the SWTT.\footnote{We compute the integral numerically using a grid with a step size of $0.01$.} For $t\in\{-1,0\}$, we obtain the following CS bounds estimator for the SWTT parameter
Similarly, we compute the multi-period CS bounds on the SWTT parameters.
To compute the SWTT parameters for the distributional DiD and CiC point estimators, we use the following
Figure (ref) presents the observed distribution of the treatment group in 2015, $\widehat{F}_{Y_1|D=1}$, as well as the CS bounds, distributional DiD and CiC point estimators of the counterfactual distribution using 2010 and 2011 as pre-treatment periods. Since the minimum wage is likely to have an impact on the bottom of the distribution, we present those figures for the bottom quartile of the wage distribution where the minimum wage increase is likely to have an impact.\footnote{We relegate the figures of the entire distribution to Figure (ref) in the online appendix.}
First, we examine the CS bounds on the counterfactual distribution using each of the pre-treatment periods separately in Figure (ref)(a) and (ref)(b), respectively. Comparing the observed (factual) distribution with the CS bounds on the counterfactual using each of the pre-treatment periods, we note an obvious change in the censoring point as expected in the context of a minimum wage increase. For instance, in Figure (ref)(b), the CS bounds on the counterfactual distribution exhibit a jump slightly above \$8, whereas the observed (factual) distribution exhibits a jump at about \$9. Furthermore, note that both upper and lower bounds satisfy the properties of a cdf. In addition, since the bounds do not cross, we do not have any detectable violation of the assumptions required for our identification approach. We also plot the sample analogue of the horizontal subcopula $C_{Y_{t0},D}(\cdot,q)$ for 2010 and 2011 to provide a visual check of our copula stability assumption in Figure (ref). This plot is the counterpart of DiD pre-trends plots in our context. While this figure does not provide a formal test of the copula stability assumption, it demonstrates that the copulas governing the dependence between $Y_{t0}$ and $D$ for 2010 and 2011 are fairly similar.
Next, we examine the bottom quartile of the distributional DiD counterfactual estimates using 2010 and 2011 as pre-treatment period in Figure (ref)(c) and (ref)(d), respectively. At first glance, we note violations of the monotonicity property of cdfs in both counterfactual distributions, indicating a violation of the testable implication of the identifying assumption of distributional DiD RothSantanna2021. The magnitude of the monotonocity violation is by far greater for the distributional DiD estimate using the 2010 pre-treatment period; the counterfactual estimate “dips” around the pre-treatment minimum wage of \$8, which is the part of the distribution particularly pertinent for the evaluation of the minimum wage increase.
Finally, we also present the CiC point estimator of the counterfactual using both pre-treatment periods in Figure (ref)(e) and (ref)(f), respectively. As demonstrated in Section (ref), the CiC point estimator coincides with the CS upper bound using the same pre-treatment period. This could translate to the CiC suffering from an upward bias in SWTT estimation as evident from comparing (ref) and (ref).
Next, we quantify the impact of the minimum wage increase on the wage distribution using the ATT and the Gini SWTT both for the overall distribution as well as its lower tail. We report 95% confidence intervals for all SWTT estimators using standard normal critical values and standard errors obtained using nonparametric bootstrap.\footnote{While the formal proof that these confidence intervals provide adequate coverage asymptotically is beyond the scope of the present paper, we have examined their performance in a simulation study mimicking our minimum wage setting which demonstrates that they provide adequate coverage in finite samples.}
Table (ref) presents 95% confidence intervals on the ATT and Gini SWTT using the CS bounds, the distributional DiD and CiC point estimators.
When examining Table (ref), we note that the 95% confidence intervals on the CS bounds for the ATT and Gini SWTT include zero, whether we use 2010 and 2011 as pre-treatment periods separately or use them both in the multi-period CS bounds. This is consistent with the expectation that a minimum wage increase is unlikely to change the mean or inequality of the overall wage distribution. When we consider the 95% confidence intervals using the distributional DiD and CiC point estimators, they suggest no improvement in terms of ATT and Gini SWTT, except using the CiC confidence interval that use the 2011 pre-treatment period. As pointed out in Section (ref), the CiC point estimator of the counterfactual coincides with the CS upper bound. As a result, the corresponding SWTT estimator may be upwardly biased.
In the context of policies such as an increase in the legal minimum wage, the welfare of subpopulations at the lower tail of the wage distribution is an important policy target. Table (ref) provides the lower-tail ATT and Gini social welfare treatment effects, $ATT(u)$ and $Gini~SWTT(u)$ for $u\in\{0.01,0.025,0.05,0.10,0.25\}$, respectively, introduced in Section (ref).
First, we consider the CS bounds using 2010 and 2011 as pre-treatment periods separately as well as the multi-period CS bounds that exploits both pre-treatment periods. Regardless of the pre-treatment year we use, for $u\in\{0.01,0.025,0.05\}$, the 95% confidence intervals on the CS bounds demonstrate statistically significant improvement in terms of lower-tail mean and Gini social welfare. When we consider $u\in\{0.10,0.25\}$, we note that while the CS bounds using the 2011 pre-treatment period demonstrate statistically significant improvements in terms of lower-tail mean and Gini social welfare, the confidence intervals on the CS bounds using the 2010 pre-treatment period are not conclusive on the sign of this impact. Since the multiple-period CS bounds combine the information from both pre-treatment periods, they result in tighter confidence intervals than the CS bounds using 2010 or 2011 by itself for both the lower-tail ATT and Gini SWTT for all quantiles $u$ we consider. These tighter confidence intervals point to improvements both in terms of mean and Gini social welfare up to the lower quartile of the distribution ($u=0.25$). This demonstrates how exploiting the multiple pre-treatment periods can aid to provide tighter bounds that translate to shorter confidence intervals.
Next, we consider the distributional DiD and CiC estimators. The distributional DiD confidence intervals using the 2010 pre-treatment period do not suggest any significant improvement in terms of lower-tail mean and Gini social welfare, whereas the distributional DiD confidence intervals using the 2011 pre-treatment period suggest significant improvements in terms of both lower-tail mean and Gini social welfare for most of the quantiles we consider. When we examine the CiC point estimator, we note that the corresponding confidence intervals suggest significant improvements in terms of mean and Gini social welfare for all of the lower-tail quantiles we consider ($u=0.25$).
The confidence intervals on the lower-tail SWTT parameters demonstrate that the distributional DiD can yield contradictory results that then require an ad-hoc choice by the applied researcher regarding which period to use.\footnote{Since the distributional DiD point estimator of the counterfactual distribution using the 2010 pre-treatment period exhibits monotonicity violations, an applied researcher would likely discard those results and use the distributional DiD estimator using the 2011 pre-treatment period, for which the monotonicity violations are very minor. The selection of the pre-treatment period relies however on a pre-test, which raises the usual post-selection inference concerns. Pre-test bias issues in the context of difference-in-difference designs have been examined in Roth2022.} The confidence intervals based on the CiC point estimator will coincide with the confidence interval on the CS upper bound and may therefore be upwardly biased.
Overall, our empirical application underscores the advantages of the CS bounds in terms of relieving the applied researcher from choosing the pre-treatment period as well as specifying the type of outcome distribution. It also demonstrates how to use the CS bounds on the counterfactual distribution to conduct inference on the SWTT parameters. Finally, The CS bounds on the counterfactual distribution can be used to bound other parameters, such as the parameters examined in Cengizetal2019. We provide these estimates in Appendix (ref).
With the goal of assessing the impact of regulatory policies on social welfare, this paper provides a unifying, partial identification result for the counterfactual distribution of the treatment group in difference-in-difference settings. Exploiting the stability of the dependence (copula) between group membership and the untreated potential outcome across time, our identification result has several advantages: (1) it applies to any outcome distribution, whether continuous, discrete or mixed, (2) it is invariant to monotonic transformations of the outcome, (3) it can allow for nonrandom selection into treatment without restricting the evolution of the marginal distribution of the potential outcomes across time. To quantify the impact of regulatory policies on social welfare, we introduce a broad class of treatment effect parameters. This class includes the ATT as well as the Gini social welfare treatment effect on the treated as a special case. We illustrate the empirical relevance of our results using a minimum wage application revisiting Cengizetal2019.
\setcounter{lemma}{0} \setcounter{claim}{0} \setcounter{example}{0}
Before we proceed to provide a proof of the above lemma, we compare the bounds in Lemma (ref)(1) with those used in AtheyImbens2006, hereinafter AI2006, to bound the counterfactual distribution for discrete outcomes. These bounds are given by the following in our notation,
Now note that the upper bound employed in AI2006 only differs from the upper bound in Lemma (ref)(1) in terms the use of $\mathbb{X}$ instead of $\mathbb{R}$. These two quantiles only differ for $u=0$, since $\{x\in\mathbb{R}:F_X(x)\geq 0\}=\mathbb{R}$, whereas $\{x\in\mathbb{X}:F_X(x)\geq 0\}=\mathbb{X}$. As a result, $Q_X^{\mathbb{R},-}(0)=-\infty$ and $F_X(Q_X^{\mathbb{R},-}(0))=0$, whereas $Q_X^{\mathbb{X},-}(0)=\inf \mathbb{X}$ and $F_X(\inf\mathbb{X})\geq 0$. Therefore, our upper bound is lower than the one used in AI2006 for $u=0$.\footnote{Note that this is inconsequential for their identification result, since they provide bounds on the counterfactual distribution on its support, and set it to zero below the infimum of its support and to one above the supremum of its support.}
The lower bound in Lemma (ref)(1) is starkly different from the lower bound in (ref). As we discuss in Section (ref), the lower bound in (ref) equals the upper bound for several examples with mixed outcomes, due to censoring or bunching, because $Q_X^{\mathbb{X},+}(u)=Q_X^{\mathbb{X},-}(u)$ for $u\in[0,1]$ for some mixed outcome distributions. As a result, the lower bound is not valid in the mixed-outcome case in general. In those cases, the AI2006 bounds would not cover the counterfactual distribution. We demonstrate additional numerical examples in Appendix (ref). By contrast, our lower bound is valid and sharp for any outcome distribution. For discrete outcomes, our bounds collapse to theirs in numerical examples provided in Appendix (ref).
By Sklar's Theorem Nelsen2006, there is a unique subcopula $C_{Y_{10}, D}$ determined on $Ran F_{Y_{10}} \times \{q\}$, such that the following hold:
Using Proposition 1(4) from Embrechts_al2013, we have:
The latter equality holds, because (i) for all $u \in \overline{\operatorname{Ran}} F_{Y_{10}}$ there exists $y \in \overline{\mathbb R}$ such that $y=Q^{\mathbb R,-}_{Y_{10}}(u)$ and (ii) from Proposition 1(4) in Embrechts_al2013 we have $F_{Y_{10}}\left(Q^{\mathbb R,-}_{Y_{10}}(u)\right)=u$ for all $u \in \overline{\operatorname{Ran}} F_{Y_{10}}$. For $u, u' \in \overline{\operatorname{Ran}} F_{Y_{10}}$ such that $u<u'$ we have $Q^{\mathbb R,-}_{Y_{10}}(u) < Q^{\mathbb R,-}_{Y_{10}}(u') \Rightarrow F_{Y_1,D}\left(Q^{\mathbb R,-}_{Y_{10}}(u),0\right) <F_{Y_1,D}\left(Q^{\mathbb R,-}_{Y_{10}}(u'),0\right) \iff C_{Y_{10},D}(u,q) <C_{Y_{10},D}(u',q)$. The first strict inequality holds because by construction $Q^{\mathbb R,-}_{Y_{10}}(u)$ is strictly increasing on $\overline{\operatorname{Ran}} F_{Y_{10}}$. The second holds because $Q^{\mathbb R,-}_{Y_{10}}(\cdot) \in \mathbb Y_{10} \subseteq \mathbb Y_{10|0}$ since $\mathbb Y_{10|1} \subseteq \mathbb Y_{10|0}$.
\qed
The proof follows in three steps. First, we derive the bounds (Section (ref)), then we proceed to show sharpness (Section (ref)). Since the sharpness proof relies on two intermediate lemmata, the last step is then to prove these two lemmata (Section (ref)).
Take a fixed $y \in \mathbb Y_{10|0}$, then the following holds for all $\tilde y< Q^{\mathbb R,+}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(y)\right)$ :
The first line of the inequality trivially holds from Lemma (ref)((ref)) and the fact that $Y_0\leq \tilde{y}$ implies $Y_0 < Q^{\mathbb R,+}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(y)\right)$. The third line holds by Sklar's Theorem Nelsen2006. The fourth line holds under Assumption (ref), and the last line holds under Assumption (ref). Notice that the last line requires $u \mapsto C_{Y_{10},D}(u,q)$ to be strictly increasing only on $\overline{\operatorname{Ran}} F_{Y_{10}}\cup \overline{\operatorname{Ran}} F_{Y_{00}} \subseteq [0,1]$. Now, applying the monotonicity of the function $v-C_{Y_0,D}(v,q)$ on the inequality ((ref)), for all $\tilde y < Q^{\mathbb R,+}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(y)\right)$ we have:
With a slight abuse of notation, we will use $F_{Y_{t0},D}(y,1)\equiv \mathbb{P}(Y_{t0}\leq y, D=1)$. Since $F_{Y_{t0}}(y)=F_{Y_{t0}, D}(y,1)+F_{Y_{t0}, D}(y,0)=F_{Y_{t0}, D}(y,1) + C_{Y_{t0},D}(F_{Y_{t0}}(y),q)$ for $t=0,1$, the latter equality implies the following:
where the second line holds under Assumption (ref). So, to summarize, for any fixed $y \in \mathbb Y_{10|0}$, we have: $$F_{Y_0|D=1}\left(\tilde y\right) \leq F_{Y_{10}|D=1}(y) \leq F_{Y_0|D=1}\left(Q^{\mathbb R,-}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(y)\right)\right), \text{ for all } \tilde y < Q^{\mathbb R,+}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(y)\right).$$ Taking the supremum over $\tilde{y}<Q_{Y_0|D=0}^{\mathbb{R},+}(F_{Y_1|D=0}(y))$ implies that: $$ \sup_{\tilde y < Q^{\mathbb R,+}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(y)\right)}F_{Y_0|D=1}\left(\tilde y\right) \leq F_{Y_{10}|D=1}(y) \leq F_{Y_0|D=1}\left(Q^{\mathbb R,-}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(y)\right)\right),$$ which is equivalent to: $$\underbrace{F_{Y_0|D=1}\left(Q^{\mathbb R,+}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(y)\right)-\right)}_{=F_{Y_0|D=1}\left(\left[Q^{\mathbb R,+}_{Y_0|D=0}\circ F_{Y_{1|D=0}}\right](y)-\right)\equiv F^{LB}(y) } \leq F_{Y_{10}|D=1}(y) \leq \underbrace{F_{Y_0|D=1}\left(Q^{\mathbb R,-}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(y)\right)\right)}_{=\left[F_{Y_0|D=1}\circ Q^{\mathbb R,-}_{Y_0|D=0}\circ F_{Y_1|D=0}\right] (y)\equiv F^{UB}(y) }.$$
We then finally have:
While these above bounds are point-wise sharp for all $y \in \mathbb Y_{10|0},$ they may not be sharp for $y \in \mathbb R \setminus \mathbb Y_{10|0}$. And this is because the upper bound may not be right-continuous in some cases, similarly for the lower bound which may not be right-continuous whenever $\{\tilde y \in \mathbb Y_{0|D=1} \cup \{-\infty\}:F_{Y_0|D=1}(\tilde y)\leq u\}$ is open for some $u\in Ran F_{Y_0|D=1}$.
To clarify this point, let us consider the simple case where $Y_{t0}$, $t \in \{0,1\}$ are all discrete random variables with $\mathbb Y_{10|0}=\{y_0,...,y_K\}$. In this case, $F^{LB}(.)$ is a well-defined cdf, while $F^{UB}(.)$ may not be a right-continuous function. Indeed, the function $u \mapsto Q^{\mathbb Y_{0|0},-}(u)$ is left-continuous and the discontinuities happen at $u\in Ran F_{Y_0|D=0}$. Now, consider that there exists $u_k \in Ran F_{Y_0|D=0} \cap Ran F_{Y_{10}|D=0}$, thus $F^{UB}(.)$ could be left-continuous at $y_k \in \mathbb Y_{10|0}$ such that $F_{Y_{10}|D=0}(y_k)=u_k$. If it is left-continuous and not right-continuous in $y_k$, we have: $\{y \in \overline{\mathbb R}: F^{UB}(y)> F^{UB}(y_k)\}=(y_k,\infty]$. Let us consider $\epsilon>0$ such that $y_k +\epsilon < y_{k+1}$. In such a case, $F_{Y_{10}|D=1}(y_k+\epsilon)=F_{Y_{10}|D=1}(y_k)$, however, by applying naively the bounds to $y_k$ and $y_{k}+\epsilon$ we have:
which implies that the upper bound in ((ref)) is not sharp since $F^{UB}(y_k+\epsilon)> F^{UB}(y_k)$. A valid tighter bound for $F^{LB}(y')$ for $y_{k}<y'<y_{k+1}$ is:
Since extending the bounds in Eq. ((ref)) to the case where $y \notin \mathbb Y_{10|0}$ provides non-sharp bounds, we provide an alternative approach that internalizes the idea that our target function of interest must be right-continuous since it is a cdf. Recall,
then for any fixed $y \in \mathbb R$, we have:
Notice that because $\mathbb Y_{10|1} \subseteq \mathbb Y_{10|0}$, and $F_{Y_{10}|D=1}(\cdot)$ is a right-continuous function, we have the following equality by Lemma (ref)((ref)): $$\lim_{\tilde y \downarrow y}\sup\left\{F_{Y_{10}|D=1}(t): t\leq \tilde y \; \& \; t \in \mathbb Y_{10|0} \cup\{-\infty \} \right\}= F_{Y_{10}|D=1}(y) \text{ for all } y \in \mathbb R.$$ The last inequality therefore becomes:
In the previous subsection (ref), we showed that the bounds are valid. Now, we will show that both bounds are achievable. For the sake of brevity, we will focus only on the upper bound. The main idea is to provide a DGP which is only a function of the observable distributions but verifies the model assumptions and for which $\tilde{F}_{Y_{10}|D=1}(y)$ is equal to the upper bound.
Consider that the unidentified counterfactual distribution is exactly the upper bound: $$\tilde{F}_{Y_{10}\vert D=1}(y)\equiv \lim_{\tilde y \downarrow y}\sup\left\{F^{UB}(t): t\leq \tilde y \; \& \; t \in \mathbb Y_{10|0} \cup\{-\infty \} \right\}.$$ For simplicity, we consider the case where $$\lim_{\tilde y \downarrow y}\sup\left\{F^{UB}(t): t\leq \tilde y \; \& \; t \in \mathbb Y_{10|0} \cup\{-\infty \} \right\}=F^{UB}(y)\equiv F^{UB}_{Y_{10}\vert D=1}(y).$$ We need to define a joint distribution on $(Y_{00},Y_{10},Y_{11}, D)$ such that it is compatible with the data $(Y_0,Y_1,D)$, and Assumptions (ref) and (ref) hold. For any vector $X$, denote $F_{X,D}(x,d)=\mathbb P(X\leq x, D=d)$. Let $F_{Y_{00},Y_{10},Y_{11},D}(y_0,y_{10},y_{11},d)$ be a candidate joint distribution. We define
We construct the proposed distribution using the following rule. For $\tilde{F}_{Y_{00},Y_{10},Y_{11},D}(y_0,y_{10},y_{11},d)$ to be compatible with the data $(Y_0,Y_1,D)$, we must have
The distributions $\tilde{F}_{Y_{11}\vert Y_{00}\leq y_0,Y_{10}\leq y_{10},D=0}(y_{11})$ and $\tilde{F}_{Y_{10}\vert Y_{00}\leq y_0,Y_{11}\leq y_{11},D=1}(y_{10})$ are counterfactual. We set both of them equal to $\tilde{F}_{Y_{10}\vert Y_{00}\leq \infty,Y_{11}\leq \infty,D=1}(y_{10})=F^{UB}_{Y_{10}\vert D=1}(y)$, which is the counterfactual distribution that we consider above.
We now show that $\tilde{F}_{Y_{10}|D=1}(y)$ is a cdf. It is easy to see that $\tilde{F}_{Y_{10}|D=1}(y)$ is nondecreasing since for $y \leq y'$ we have
The limits of the function $\tilde{F}_{Y_{10}|D=1}(y)$ at $-\infty$ and $\infty$ are 0 and 1, respectively. By construction, the function $\tilde{F}_{Y_{10}|D=1}(y)$ is a right-continuous function.
We have
We now need to construct copulas $\tilde{C}_{Y_0,D}(u,q)$, $\tilde{C}_{Y_{10},D}(u,q)$, $\tilde{C}_{Y_{0},Y_{1} \vert D=0}(u_0,u_1)$, and $\tilde{C}_{Y_{0},Y_{10},D}(u_0,u_1,q)$ such that the following holds:
where $\tilde{F}_{Y_{10}}(y_{10}) = p F^{UB}_{Y_{10}|D=1}(y_{10}) + q F_{Y_1 \vert D=0}(y_{10})\equiv F^{UB}_{Y_{10}}(y_{10})$.
Since $\overline{Ran}F_{Y_0}$ is closed, we define {\tiny{
}}
where for any $u \in [0,1]$, $\underline{u}(u)\equiv \sup\{q\in \overline{\operatorname{Ran}} F_{Y_0} \cup \overline{\operatorname{Ran}} \tilde{F}_{Y_{10}}: q \leq u\}$, $\overline{u}(u)\equiv \inf\{q\in \overline{\operatorname{Ran}} F_{Y_0} \cup \overline{\operatorname{Ran}} \tilde{F}_{Y_{10}}: q \geq u\}$, and $ \tilde{Q}^{\mathbb{R},-}_{Y_{10}}(u)\equiv\inf\{y\in\mathbb{R}:\tilde{F}_{Y_{10}}(y)\geq u\}$.
{\tiny{
}} where for $t\in \{0,1\}$ and for any $(u_0,u_1) \in [0,1]^2$, $\underline{u_t}(u)\equiv \sup\{q\in \overline{\operatorname{Ran}} F_{Y_t\vert D=0}: q \leq u\}$, while $\overline{u_t}(u)\equiv \inf\{q\in \overline{\operatorname{Ran}} F_{Y_t \vert D=0}: q \geq u\}$.
We then define for $(u_0,u_1) \in [0,1]^2$
We can verify that $\tilde{C}_{Y_{00},Y_{10},D}(u_0,u_1,q)$ is a well-defined copula. We start by showing that $\tilde{C}_{Y_0,D}(u_0,q)$ is a well-defined subcopula. To do so, we need to introduce two intermediate lemmata:
First, we have $\tilde{C}_{Y_0,D}(1,q)=F_{Y_0,D}(Q_{Y_0}^{\mathbb{R},-}(1),0)=q$. Now let us show that for all $(u,v)\in[0,1]^2$ such that $u < v$, we have $\tilde{C}_{Y_0,D}(u,q)< \tilde{C}_{Y_0,D}(v,q)$. From the definition of $\tilde{C}_{Y_0,D}(u,q)$ and Lemma (ref), it follows that, when $u$ and $v$ belong to the same range, this monotonicity condition holds. We are going to prove it when $u$ and $v$ belong to different ranges. On the one hand, if $u \in Ran\tilde{F}_{Y_{10}}$ and $v \in RanF_{Y_0}$, then from Lemma (ref), we have $\tilde{C}_{Y_0,D}(u,q) < \tilde{C}_{Y_0,D}(v,q)$. On the other hand, if $v \in Ran\tilde{F}_{Y_{10}}$ and $u \in RanF_{Y_0}$, then from Lemma (ref), we have $\tilde{C}_{Y_0,D}(u,q) < \tilde{C}_{Y_0,D}(v,q)$. Since $\tilde{C}_{Y_0,Y_1 \vert D=0}(u_0,u_1)$ is an extended copula of the identified part of the copula of $(Y_0,Y_1) \vert D=0$ through the Sklar theorem, it is a well-defined copula. Any extended copula of this form should work for the proof, as we do not impose any additional restrictions on the true copula of $(Y_0,Y_1) \vert D=0$.
We also need to check that $\tilde{C}_{Y_{00},Y_{10},D}\left(F_{Y_0}(y_0),\tilde{F}_{Y_{10}}(y_{10}),q\right)=F_{Y_0,Y_1,D}(y_0,y_{10},0).$ This latter equality holds by construction of $\tilde{C}_{Y_{00},Y_{10},D}(u_0,u_1,q)$.
When we let $u_0$ go to 1, we obtain
Similarly,
And by construction, we have $\tilde{C}_{Y_{10},D}(u,q)=\tilde{C}_{Y_0,D}(u,q)$ for all $u\in [0,1]$ (Assumption (ref) holds). Furthermore, we have shown above that $\tilde{C}_{Y_0,D}(u,q)$ is strictly increasing in $u$ (Assumption (ref) holds).
By construction, the proposed joint distribution $\tilde{F}_{Y_{00},Y_{10},Y_{11},D}(y_0,y_1,y_2,d)$ is compatible with the data and the proposed copulas $\tilde{C}_{Y_0,D}(u,q)$, and $\tilde{C}_{Y_{10},D}(u,q)$ satisfy Assumptions (ref) and (ref).
The proof is similar for the lower bound on $F_{Y_{10\vert D=1}}(y)$ and any distribution in the identified set of $F_{Y_{10}\vert D=1}(y_{10})$.
To complete the proof, it remains to show the two intermediate lemmata.
First, we start by the following claims:
Proof. We have
Since $F_{Y_0\vert D=1}(h(y)) \in RanF_{Y_0 \vert D=1},$ to obtain the smallest element $v \in RanF_{Y_{0}}$, we need to find the smallest element $s$ on $RanF_{Y_0\vert D=0}$ such that $F_{Y_{1}\vert D=0}(y) \leq s$. From Lemma (ref).((ref)), $s=F_{Y_0\vert D=0}(Q^{\mathbb R,-}_{F_{Y_0\vert D=0}}(F_{Y_{1}\vert D=0}(y)))$. This completes the proof of Claim (ref).\qed
\noindentProof. $u_y\equiv F_{Y_{10}}^{UB}(y) = q F_{Y_{10}\vert D=0}(y) + p F^{UB}_{Y_{10}\vert D=1}(y),$ where $F^{UB}_{Y_{10}\vert D=1}(y)=F_{Y_0\vert D=1}(h(y))$ with $h(y)=Q_{Y_{0\vert 0}}^{\mathbb R,-}(F_{Y_1\vert D=0}(y))$. Then, $u_y \leq q F_{Y_{0}\vert D=0}(h(y)) + p F_{Y_0\vert D=1}(h(y))$, since $F_{Y_{10}\vert D=0}(y) \leq F_{Y_{0}\vert D=0}(h(y))$ by construction. So, $u_y \leq F_{Y_0}(h(y))\equiv v_y \in RanF_{Y_0}$. Now, the following hold:
This completes the proof of Claim (ref).\qed
Now we proceed to complete the proof of the lemma. Take $u \in RanF_{Y_{10}}^{UB}$ and $v\in RanF_{Y_0}$ such that $u < v$. Since $u \in RanF_{Y_{10}}^{UB}$, there exits $y$ such that $u_y=F_{Y_{10}}^{UB}(y)$. Then, from Claim (ref), there exists $v_y= F_{Y_0}(h(y))$ such that $u_y \leq v_y$. From Claim (ref), we have $v_y \leq v$. If $v=v_y$, then we have $F_{Y_{10}}^{UB}(y) < F_{Y_0}(h(y))$, which implies successively
If $v_y < v$, then from Claim (ref) we have $\tilde{C}_{Y_0,D}(u,q)\leq \tilde{C}_{Y_0,D}(v_y,q)$. And since $\tilde{C}_{Y_0,D}(v,q)$ is strictly increasing on $RanF_{Y_0}$ from Lemma (ref), we have $\tilde{C}_{Y_0,D}(v_y,q) < \tilde{C}_{Y_0,D}(v,q)$. Therefore, $\tilde{C}_{Y_0,D}(u,q) < \tilde{C}_{Y_0,D}(v,q)$. \qed
We first start by stating and proving the following claim:
\noindentProof. We have
where $\underline{h}(y)= Q_{Y_0\vert 0}^{\mathbb R,+}(F_{Y_1\vert D=0}(y))$, and the second inequality holds from Lemma (ref). We discuss two cases.
Case 1: $F^{UB}_{Y_{10}\vert D=1}(y)=F^{LB}_{Y_{10}\vert D=1}(y)$
In this case, $u_y=w_y$, we have
Case 2: $F^{LB}_{Y_{10}\vert D=1}(y)< F^{UB}_{Y_{10}\vert D=1}(y)$
In this case, $w_y \notin \overline{\operatorname{Ran}} F_{Y_{10}}^{UB}$. From Lemma (ref), $v_y$ is the highest element of $\overline{\operatorname{Ran}} F_{Y_0}$ such that $w_y \geq v_y$. First, suppose $w_y \notin \overline{\operatorname{Ran}} F_{Y_0}.$ Then $w_y \in (\overline{\operatorname{Ran}} F_{Y_0})^c \cap (\overline{\operatorname{Ran}} F^{UB}_{Y_{10}})^c$. Let $\overline{u}(w_y)\equiv \inf\{q\in \overline{\operatorname{Ran}} F_{Y_0} \cup \overline{\operatorname{Ran}} F^{UB}_{Y_{10}}: q \geq w_y\}$. We have $v_y \leq w_y < \overline{u}(w_y) \leq u_y$, and either $\overline{u}(w_y) \in \overline{\operatorname{Ran}} F_{Y_0}$ or $\overline{u}(w_y) \in \overline{\operatorname{Ran}} F^{UB}_{Y_{10}}$.
If $\overline{u}(w_y) \in \overline{\operatorname{Ran}} F_{Y_0}$, then
Since $0 \leq \frac{w_y-v_y}{\overline{u}(w_y)-v_y} \leq 1$ and $\bigg[F_{Y_0,D}\left(Q_{Y_0}^{\mathbb{R},-}(\overline{u}(w_y)),0\right)-F_{Y_0,D}\left(Q_{Y_0}^{\mathbb{R},-}(v_y),0\right)\bigg] \geq 0$ from Lemma (ref), the following holds:
where the last inequality holds because $Q_{Y_{0}}^{\mathbb{R},-}(u)$ is monotone in $u$. Hence, $$\tilde{C}_{Y_0,D}(v_y,q) \leq \tilde{C}_{Y_0,D}(w_y,q) \leq \tilde{C}_{Y_0,D}(u_y,q).$$
If $\overline{u}(w_y) \in \overline{\operatorname{Ran}} F^{UB}_{Y_{10}}$, then $\overline{u}(w_y) \in \overline{\operatorname{Ran}} F^{UB}_{Y_{10}}=u_y$, and
Since $0 \leq \frac{w_y-v_y}{\overline{u}(w_y)-v_y} \leq 1$ and $\bigg[F_{Y_1,D}(y,0)-F_{Y_0,D}\left(\underline{h}(y)-,0\right)\bigg] \geq 0$ from Lemma (ref), the following holds:
Second, suppose $w_y \in \overline{\operatorname{Ran}} F_{Y_0}.$ Then, from Lemma (ref), we must have $w_y=v_y$, which implies $F_{Y_{1}, D}(y,0)=F_{Y_{0},D}(\underline{h}(y)-,0)$, which in turn implies $\tilde{C}_{Y_0,D}(u_y,q)=\tilde{C}_{Y_0,D}(v_y,q)=\tilde{C}_{Y_0,D}(w_y,q)$.
This completes the proof of Claim (ref). \qed
Now we proceed to complete the proof of the lemma. Take $u \in RanF_{Y_{10}}^{UB}$ and $v\in RanF_{Y_0}$ such that $v < u$. Since $u \in RanF_{Y_{10}}^{UB}$, there exits $y$ such that $u_y=F_{Y_{10}}^{UB}(y)$. From Claim (ref), there exists $w_y \in [0,1]$ and $v_y \in \overline{\operatorname{Ran}} F_{Y_0}$ such that $v\leq v_y < w_y \leq u$.
Case 1: $F^{UB}_{Y_{10}\vert D=1}(y)=F^{LB}_{Y_{10}\vert D=1}(y)$
In this case, $u_y=w_y$, we have
Case 2: $F^{LB}_{Y_{10}\vert D=1}(y)< F^{UB}_{Y_{10}\vert D=1}(y)$
The proof here is very similar to Case 2 in Claim (ref), except the strict inequality $0 < \frac{w_y-v_y}{\overline{u}(w_y)-v_y} < 1$. This strict inequality implies $$\tilde{C}_{Y_0,D}(v,q) \leq \tilde{C}_{Y_0,D}(v_y,q) < \tilde{C}_{Y_0,D}(w_y,q) \leq \tilde{C}_{Y_0,D}(u_y,q).$$ Hence, $\tilde{C}_{Y_0,D}(v,q) < \tilde{C}_{Y_0,D}(u,q)$.
Now we have completed the proof of the two intermediate lemmata and thereby the proof of Theorem (ref).
\qed
The proof of this theorem follows by similar arguments to the proof of Theorem (ref) and is therefore provided in Section (ref) of the online appendix.
(i) $\Longrightarrow$ (ii).
Since the cdf $F_{Y_{t0}}$ is continuous and strictly increasing, we have
By definition, $h_t$ is continuous and strictly increasing as is the quantile function $Q^{\mathbb R,-}_{Y_{t0}}$. Then, the following equalities hold:
where the second equality holds from the invariance principle in Embrechts_al2013 (Embrechts_al2013, Proposition 4(2)). Therefore,
where the second implication follows from $U_{t0} \sim \mathcal U_{[0,1]}$ and $F_D(0)=q$, the third holds from Sklar's theorem, and the fifth follows from $U_{t0} \sim \mathcal U_{[0,1]}$. Hence, we have:
(ii) $\Longrightarrow$ (i). Suppose there exist two strictly increasing functions $h_t(.), t \in \{0,1\}$ and two uniformly distributed random variables over $[0,1]$ $U_{00}$ and $U_{10}$ such that $Y_{t0}=h_t(U_{t0})$ and $U_{00}|D=d \sim U_{10}|D=d$. Then, we have
where the fourth implication holds from Sklar's theorem, the fifth follows from $U_{t0} \sim \mathcal U_{[0,1]}$, the sixth follows by the invariance principle in Embrechts_al2013 (Embrechts_al2013, Proposition 4.(2)), and the last holds from $Y_{t0}=h_t(U_{t0})$. \qed
\setcounter{figure}{0}
\setcounter{table}{0} \setcounter{page}{0} \setcounter{claim}{0} \pagenumbering{gobble}
\startcontents[sections] \printcontents[sections]{l}{1}{\setcounter{tocdepth}{1}} \setcounter{page}{0} \pagenumbering{arabic}