EconBase
← Back to paper

Evaluating the Impact of Regulatory Policies on Social Welfare in Difference-in-Difference Settings

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

141,482 characters · 37 sections · 85 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Evaluating the Impact of Regulatory Policies on Social Welfare in Difference-in-difference Settings

\pagenumbering{gobble}

center[center omitted — 1,652 chars of source]

\setcounter{page}{0}

\pagenumbering{arabic}

Introduction

Government's regulatory role and its impact on social welfare has been a critical question for economists. These regulatory policies often restrict the budget or choice sets for certain agents in the market by imposing floors or quotas, such as minimum wages, minimum/maximum working time, wage floors for different occupation groups as well as action, reporting and notification thresholds in environmental monitoring. Those types of policies tend to induce behavioral responses that can generate mass points in the outcome of interest. For instance, an important question in the labor economics literature is the effect of an increase or introduction of minimum wages on low-wage jobs or overall employment, see for instance CK1994, NeumarkWascher2008, Cengizetal2019, among many others. The figure below (taken from Cengizetal2019) illustrates that an increase in the minimum wage will shift jobs that were previously paying below the minimum wage $MW$, and then will create “excess jobs” at and slightly above the minimum wage.

figure[figure omitted — 668 chars of source]

This figure also shows the heterogeneous effect of such a policy; it is expected to only affect the wage of low-wage workers and not have an effect on the upper tail of the distribution. In sum, those types of policies have two main features. First, the potential outcomes of interest are likely to exhibit some mass points. Second, the causal effect of the policy is expected to affect only a part of the distribution of the outcomes of interest. As a result, to adequately analyze the impact of these policies, a distributional treatment effect analysis is key, as in Cengizetal2019 for instance; see, also, Almondetal2011,assuncao2022. Furthermore, measuring the impact of such policies on social welfare requires recovering the counterfactual distribution of the outcome of interest.

While these types of policies are widely studied in economics, the existing econometrics methods are not necessarily adequate to recover distributional causal effects in these settings. In the presence of data before and after a new policy, one of the most widely used techniques to assess its impact is the difference-in-differences (DiD) method. Its main drawbacks, however, are two-fold: (1) it does not identify the counterfactual distribution, (2) it is not invariant to monotonic transformations. While there are several methods to identify the counterfactual distribution in difference-in-difference settings AtheyImbens2006,bonhommesauder2011,CallawayLi2019,HavnesMogstad2015, to the best of our knowledge, the distributional DiD and changes-in-changes (CiC) are the only two approaches that are invariant to monotonic transformations.\footnote{The distributional DiD method relies on a parallel trends assumption in the cdfs as opposed to the expectations HavnesMogstad2015,RothSantanna2021. }

RothSantanna2021 show that distributional DiD requires that the distribution of the untreated potential outcome is independent of policy adoption, is stationary across time (within each group), or consists of a mixture of two subpopulations each obeying one of the two restrictions. Such conditions are unlikely to be valid for the policy evaluation questions we are interested in. Indeed, the independence assumption (random assignment) is implausible in our context since the decision to implement a new minimum wage policy is a response to the unsatisfactory features of the pre-policy outcome distribution, such as large wage inequalities, high proportion of workers under poverty, etc. When the policy is not randomly assigned, the validity of the distributional DiD essentially rests on the stationarity assumption, which is restrictive in many practical settings.\footnote{The stationarity assumption can be tested using the control group. RothSantanna2021 provide a sharp specification test of the validity of the distributional DiD assumption in general.}

While the CiC approach introduced in the seminal work by AtheyImbens2006 can accommodate endogenous policy (treatment) assignment as well as time-varying potential outcome distributions, their identification result does not apply to the case where the potential outcomes exhibit some mass points (mixed distributions), as in Figure (ref).\footnote{Mass points are common for a wide range of economic outcomes resulting from censoring bottomcoding1,topcoding1 or bunching bunching3,bunching4,bunching2,bunching5,bunching6,bunching7,bunching8.} In fact, AtheyImbens2006 introduce the CiC approach for either continuous or discrete outcomes that are monotonic (time-varying) functions of a scalar unobservable with a time-invariant distribution across time. In sum, the CiC approach introduced in AtheyImbens2006 should not be applied to evaluate the policies described above.

The current paper provides an alternative, unifying identification result that applies to any type of outcome distribution, is invariant to monotonic transformations, allows for endogeneity of the policy assignment, and does not restrict the evolution of the marginal distribution, nor treatment effect heterogeneity. Our identification result exploits the stability of the dependence (copula) between treatment assignment and the untreated potential outcome across time without imposing restrictions on the structural function that generates the potential outcomes.

Exploiting our copula stability (CS) assumption, we provide a unifying partial identification result for the counterfactual distribution of the treatment group. We then extend our analysis to the case where multiple pre-treatment periods are available. In this case, we show that if copula stability holds for multiple pre-treatment periods, then our multi-period CS bounds exploit the information from the pre-treatment periods to provide tighter bounds.\footnote{We use multi-period CS bounds to refer to CS bounds that use multiple pre-treatment periods.} The presence of multiple pre-treatment periods also allows us to provide a testable restriction of our model assumptions. We demonstrate our theoretical results numerically in Section (ref).

Our CS bounds apply to any type of outcome distribution, whether it is continuous, mixed, or discrete. They shrink to the point-identification result in AtheyImbens2006 for continuous outcomes. Indeed, we show that in this case our copula stability assumption is equivalent to the CiC conditions. For discrete outcomes, we show that our copula stability assumption can be compatible with an underlying production function featuring multi-dimensional unobserved heterogeneity, whereas the CiC bounds for discrete outcomes require a scalar unobservable. For mixed outcomes, we demonstrate that a na\"ive implementation of the CiC approach may lead to a point-estimand that does not coincide with the true counterfactual, whereas our CS bounds will include it.\footnote{We refer to this implementation as na\"ive since AtheyImbens2006 did not provide identification results for mixed outcomes. Nonetheless, an empirical researcher might ignore the mixed-nature of this outcome and implement their point-identification result. }

We also examine the connection between our main identifying assumption and the parallel trends assumption required by DiD. The parallel trends assumption can be equivalently stated as a covariance stability assumption. It is specifically a time invariance assumption on the covariance between treatment assignment and the untreated potential outcome, whereas our assumption maintains the stability of the copula between these two variables. As a result, there are several differences between our copula stability assumption and covariance stability (parallel trends). First, the parallel trends assumption restricts the joint variability of treatment assignment and the untreated potential outcome over time, whereas our copula stability assumption only restricts their dependence structure. Second, while the parallel trends assumption restricts the evolution of the marginal distribution of the untreated potential outcome across time, copula stability does not restrict the evolution of the marginal distribution, nor treatment effect heterogeneity. Last but not least, parallel trends is not invariant to monotonic transformations except under strong conditions on heterogeneity RothSantanna2021. These conditions specifically rule out the existence of a subpopulation that selects into treatment based on unobservables and exhibits changes in its potential outcome distribution. By contrast, our copula stability condition does not rule out such a subpopulation.

Since the motivation behind policies, such as increases in the legal minimum wage, is often to reduce inequality and/or target a specific part of the outcome distribution, we introduce a broad class of social welfare treatment effect parameters that can accommodate the policymaker's objective. While this class includes the average treatment effect on the treated (ATT) as a special case, the ATT corresponds to a social welfare function that is inequality-neutral and gives equal weight to all individuals in the population. As a result, if a policymaker is averse to inequality, then the ATT would be an inadequate causal parameter to judge the policy's effectiveness. In general, the social welfare function adequate to evaluate a specific regulatory policy can be highly context-specific and may depend on the policymaker's preference and/or objective.\footnote{Please see the discussion in Bergeretal2022 which illustrates how the quantitative analysis of the effect of the minimum wage could highly differ depending on the social welfare weights, which are usually unknown to the researcher.} We therefore introduce a broad class of treatment effect parameters that take into account the policy objectives. This broad class specifically includes the class of generalized Gini social welfare functions Mehran1976,Weymark1981. These social welfare functions can take into account measures of inequality by putting higher weight on individuals with lower-ranked outcomes. In addition, we include a class of parameters that can capture the welfare of individuals at the lower tail or a specific interquantile range of the distribution. Bounds on these social welfare treatment effect parameters can be easily computed using our bounds on the counterfactual distribution. We illustrate the usefulness of this broad class of parameters and compare it to the ATT in the context of our empirical application examining the impact of a minimum wage policy (Section (ref)).

We organize the rest of the paper as follows. Section (ref) introduces the analytical framework and presents our main identification results. Section (ref) introduces the class of social welfare treatment effect parameters. Section (ref) provides an empirical illustration examining the impact of minimum wage increases on the wage distribution revisiting Cengizetal2019.

Related Literature

A comparison between our identifying assumption and some of the related approaches in the literature is warranted. bonhommesauder2011 exploit a separable model of the potential outcome to identify the entire counterfactual distribution of the treatment group in a DiD design. By relying on restrictions on the outcome model, it is therefore similar in spirit to the identification approach in AtheyImbens2006. botosarumuris2023 propose identification of counterfactual parameters for a class of semiparametric panel models, whereas our approach can accommodate both repeated cross-sections and panel data and is fully nonparametric. CallawayLi2019 also provide a fully nonparametric identification result exploiting a copula stability restriction on different objects than the ones used in this paper. They require the copula between changes and levels of the untreated potential outcome to be invariant across time for the treatment group, while our copula stability assumption does not restrict the evolution of the marginal distribution of the untreated potential outcome (Remark (ref)). Furthermore, our approach can be applied to repeated cross-sections or panel data and only requires two time periods, whereas CallawayLi2019 require at least three periods of panel data. wooldridge2023simple proposes alternative parallel trends assumptions that are more suitable for binary, fractional and count outcome data. The approach in wooldridge2023simple requires the specification of a parametric transformation model of a linear index for each type of outcome and point-identifies the average treatment effect on the treated, whereas our approach applies to any outcome, is fully nonparametric and partially identifies the counterfactual distribution.

Finally, this paper contributes to a strand in the microeconometrics literature that relies on copula theory. For cross-sectional settings with exogenoeus regressors, rothe2012 provides identification results for partial distributional effects, which hold the copula of the covariates constant, but vary their marginal distributions. mourifie2015 relies on copula theory to provide sharp bounds on the average treatment effect in a binary triangular system. arellanobonhomme2017 propose a method to correct for sample selection in quantile models, where the conditional copula of the error terms in the outcome and selection equations is a key ingredient in their approach.

Analytical Framework and Main Identification Results

Following Abadie2005, we consider the following potential outcomes model:\footnote{Note that this model implicitly assumes that there are no anticipatory effects of the treatment, that is, $Y_{00}=Y_{01}=Y_0$.}

eqnarray[eqnarray omitted — 127 chars of source]

where $Y_{t}$ denotes the observed outcome at period $t$ and $Y_{td}$ denotes the potential outcome at period $t\in\{0,1\}$ and treatment status $d\in\{0,1\}$. In the two-group, two-period case, $D$ denotes both group membership and the treatment status in period 1.

We use the following shorthand notation: $p\equiv \mathbb P(D=1)$, $q=1-p$, $\overline{\mathbb R}\equiv\mathbb R \cup \{-\infty, \infty\}$, $Ran H \equiv\{ H(y):y \in \mathbb R\}$, $\overline{\operatorname{Ran}} F\equiv Ran F \cup \{\inf Ran F, \sup Ran F\}$, and $Dom H$ denotes the domain of the function $H$. We consider the following mappings $Q^{\mathbb T,-}_X: [0,1] \rightarrow \mathbb T$, and $Q^{\mathbb T,+}_X: [0,1] \rightarrow \mathbb T$, where $Q^{\mathbb T,-}_X(u)\equiv \inf \{x \in \mathbb T \cup \{\infty\}: F_X(x)\geq u\}$ for all $u \in [0,1]$, $Q^{\mathbb T,+}_X(u)\equiv \sup \{x \in \mathbb T \cup \{-\infty\}: F_X(x)\leq u\}$ for all $u \in [0,1]$. We call $Q^{\mathbb T,+}_X$ and $Q^{\mathbb T,-}_X$ generalized quantile functions whenever $F_X(.)$ is a well-defined cumulative distribution function (cdf). We denote by $\mathbb F$ the space of all well-defined cdfs. $Supp X=\mathbb X$ denotes the support of $X$, and $\mathbb X_{s|d}$ denotes the support of $X_s|D=d$ for $d \in \{0,1\}$. Finally, we define $F_X(x-)\equiv \mathbb{P}(X<x)$.

Identifying Assumptions

Our main identification result relies on restrictions imposed on the dependence structure across time. To do so, we rely on copula theory. Copulas are functions that enable us to separate the marginal distributions from the (scale-free) dependence structure of a given multivariate distribution. In our context, we are interested in the subcopula between the untreated potential outcome and group membership across time. Working with copulas in our case will allow us to avoid restricting the type of marginal distribution of the potential outcomes as well as its heterogeneity across time. To fix ideas, let us first provide a formal definition of the (sub)copula.

definition[Nelsen2006] A two-dimensional subcopula is a function $C$ with the following properties: \begin{enumerate} • $Dom C=S_1\times S_2$, where $S_1$ and $S_2$ are subsets of $[0,1]$ containing $0$ and $1$; • For all $u,u' \in S_1$, and $v,v' \in S_2$ such that $u\leq u'$, and $v\leq v'$, we have: $$ C(u',v') + C(u,v)\geq C(u',v) + C(u,v');$$$C(0,v)=C(u,0)=0$ for all $(u,v) \in S_1\times S_2$, and $C(1,v)=v$, $C(u,1)=u$ for all $(u,v) \in S_1\times S_2$. \end{enumerate}

A copula is a special case of a subcopula where $S_1=S_2=[0,1]$. For a fixed $v \in S_2$, $u \mapsto C(u,v)$ is usually called the horizontal subcopula. The link between the joint distribution and the subcopula has been established by the well-known Sklar (1959) theorem, which provides the following lemma when applied to our context.

lemma[Sklar, 1959] There exists a unique subcopula $C: \overline{\operatorname{Ran}} F_{Y_{t0}}\times \{0,q,1\} \rightarrow [0,1]$ such that \begin{eqnarray} \mathbb P(Y_{t0}\leq y, D=0)=C_{Y_{t0},D}(F_{Y_{t0}}(y),q),\;\;\; for y \in \overline{\mathbb{R}}. \end{eqnarray}

To provide intuition for the role of the horizontal subcopula at $q$, it is helpful to divide each side of Equation\ (ref) by $q$, which yields the following for $y\in \overline{\mathbb{R}}$

eqnarray[eqnarray omitted — 93 chars of source]

Now, let us assume that the copula is strictly increasing in its first argument such that its inverse $C_{Y_{t0},D}^{-1}(\cdot~;q)$ is well-defined.\footnote{Note that a horizontal copula is by definition Lipschitz continuous.} We can then show that $C_{Y_{t0},D}^{-1}(\cdot~;q)$ is the main ingredient in the rank mapping between the treatment and control group's untreated potential outcome distribution in period $t$, which we denote by $\Gamma_t(\cdot)$:

eqnarray[eqnarray omitted — 136 chars of source]

where $\Gamma_t(u)\equiv \frac{1}{p}\left( C_{Y_{t0},D}^{-1}\left(uq;q \right)- uq \right)$ for $u\in RanF_{Y_{t0}|D=0}$. The mapping governs the relationship of the rank that a given value $y$ has in the control group's distribution of the untreated potential outcome $F_{Y_{t0}|D=0}$ (factual at each period) onto its rank in the treatment group's distribution of the untreated potential outcome, $F_{Y_{t0}|D=1}$.

Next, we introduce our main assumption.

assumption[Copula stability] The following condition holds: $C_{Y_{00},D}(u,q)=C_{Y_{10},D}(u,q)$ for all $u \in [0,1]$.

In the following, we will refer to Assumption (ref) as “copula stabilty” for brevity, but we emphasize that it only requires the stability of the horizontal copula between $Y_{t0}$ and $D$ at $q$, $C_{Y_{t0},D}(u,q)$ for $u\in[0,1]$. There are multiple advantages to our copula stability assumption. First, it is invariant to strictly monotonic transformations. Specifically, for any right-continuous function $g$, that is strictly increasing on $\mathbb Y_{td}$, we have:\footnote{See Embrechts_al2013 (Embrechts_al2013, Proposition 4(2)) for a formal proof.} $$C_{g(Y_{td}),D}(u,q)=C_{Y_{td},D}(u,q) \;\; \forall u \in RanF_{Y_{td}}.$$ Second, it does not impose any restrictions on the variability of the marginal distribution $F_{Y_{t0}}$ across time. Last, but not least, it does not restrict the type of marginal distribution $F_{Y_{t0}}$, whether it is continuous, discrete or mixed.

Assumption (ref) is the key assumption behind our identification approach. It implies that the rank mapping $\Gamma_t(\cdot)$ is stable across periods, i.e. $ \Gamma_1(\cdot)=\Gamma_0(\cdot)\equiv \Gamma(\cdot)$. In the presence of multiple pre-treatment periods, $\Gamma(\cdot)$ can be recovered from each pre-treatment period. As a result, analogous to pre-trend testing in difference-in-differences designs, the time-invariance of $\Gamma(\cdot)$ can also be tested as we demonstrate in Section (ref).

Given the wide use of difference-in-differences, it is also helpful to clarify the relationship between our copula stability assumption and the parallel trends assumption. The parallel trends assumption can be equivalently rewritten as a covariance stability assumption as we show in Appendix (ref),

eqnarray[eqnarray omitted — 124 chars of source]

This equivalence result provides, first, an intuition for why the parallel trends assumption is not invariant to a monotonic transformation since the covariance is not invariant to monotonic transformations. Second, it allows us to observe that the parallel trends assumption jointly restricts the evolution of the marginal distribution of $Y_{t0}$ across time and the dependence between $Y_{t0}$ and $D$. Unlike the parallel trends assumption, our copula stability assumption does not constrain the evolution of the marginal distribution across time, yet it relies only on the stability of the horizontal copula that governs the relationship between $Y_{t0}$ and $D$. As can be seen in the following equation, the two assumptions are non-nested in general:

eqnarray*[eqnarray* omitted — 96 chars of source]

Indeed, copula stability may hold while $Cov(Y_{10},D)\neq Cov(Y_{00},D)$ because $F_{Y_{10}}\neq F_{Y_{00}}$; and the covariance stability may hold while the copula stability is violated.

In the following, we provide several examples to illustrate the restrictions imposed by our key assumption, and how it compares to some existing assumptions.

example[Roy selection] Consider the following data generating process (DGP) in which the treatment is received when its gain (treatment effect) is bigger than or equal to a threshold, say 0 for simplicity. This is a simple Roy model where selection into treatment is on the gain. \begin{eqnarray} \left\{ \begin{array}{lcl} Y_0 &=& U_0\\ \\ Y_{1} &=& \eta D+ U_1 \\ \\ D &=& \mathbbm{1}\{\eta\geq 0\} \end{array} \right. \end{eqnarray} where $\left(\begin{array}{c}U_0\\ U_1 \\ \eta \end{array}\right)\sim N(0,\Sigma)$, $\Sigma=\left(\begin{array}{ccc}\sigma_0^2 & \delta \sigma_0 \sigma_1 & \rho_0 \sigma_0\\ \delta \sigma_0 \sigma_1 & \sigma_1^2 & \rho_1 \sigma_1\\ \rho_0 \sigma_0& \rho_1 \sigma_1 & 1\end{array}\right)$, and $\rho_t \neq 0$. In this case, we have the following: \begin{enumerate} • {\bf Copula stability:} $\rho_0=\rho_1 \Leftrightarrow Corr(\eta,Y_{00})=Corr(\eta,Y_{10})$. • {\bf Parallel trends:} $\rho_0 \sigma_0=\rho_1 \sigma_1 \Leftrightarrow Cov(\eta,Y_{00})=Cov(\eta,Y_{10})$. • {\bf Distributional DiD:} $\rho_0=\rho_1$ and $\sigma_0^2=\sigma_1^2$ $\Leftrightarrow$ $Y_{00}|D=d\sim Y_{10}|D=d$, for $d=0,1$. \end{enumerate} As can be seen, the copula stability assumption is equivalent to $\rho_0=\rho_1$, meaning that the correlation between the policy effect $\eta$ and $Y_{t0}$ is stable over time. It does not restrict any moment of the marginal distribution of the potential outcomes $Y_{t0}$. The parallel trends assumption, however, restricts the variances of the potential outcomes $Y_{00}$ and $Y_{10}$, since it is equivalent to $\rho_0 \sigma_0=\rho_1 \sigma_1$. The validity of the distributional DiD in this setting is implausible, since it requires stationarity of $Y_{t0}|D=d$. This could be easily checked using the observed distribution of the control group. Note that while the copula stability condition, result (a), does not rely on the Gaussianity assumption imposed on the marginal distribution, results (b) and (c), which involve the parallel trends and distributional DiD assumptions, are heavily dependent on this distributional assumption. For further details, see Appendices (ref) and (ref).

The above example demonstrates the copula stability assumption in the context of selection on the gains from the treatment. We next consider selection on untreated potential outcomes. This example shows that copula stability requires comonotonicity between the untreated potential outcomes in the pre- and post-treatment periods.

example[Selection on untreated potential outcomes] Consider the following model, where selection into treatment is a function of the pre-treatment outcome, such as in the Ashenfelter dip, \begin{eqnarray*} \left\{ \begin{array}{lcl} Y_0 &=& Y_{00},\\ Y_{1} &=& Y_{11} D + Y_{10} (1-D), \\ D &=& \mathbbm{1}\{Y_{00}> c\}. \end{array} \right. \end{eqnarray*} Assume that $Y_{00}$ has a continuous and strictly increasing cdf. It can be shown that $$ C_{Y_{t0},D}(u,q)=\mathbb P\left(F_{Y_{t0}}(Y_{t0}) \leq u, F_{Y_{00}}(Y_{00}) \leq q\right). $$ By construction, we can see that we always have $C_{Y_{00},D}(u,q)=\min(u,q)$, while $C_{Y_{10},D}(u,q)=\min(u,q)$ if $Y_{00}$ and $Y_{10}$ are comonotone. Thus, in this model with selection on the lagged outcome, the copula stability assumption holds if $Y_{10}=h_1(U)$, $Y_{00}=h_0(U)$ for some non-decreasing functions $h_0$ and $h_1$. Notice that when selection is on lagged outcomes, the parallel trends assumption fails in general (as $\mathbb E[Y_{10}-Y_{00} \vert Y_{00} > c]\neq \mathbb E[Y_{10}-Y_{00} \vert Y_{00} \leq c]$), unless $Y_{10}-Y_{00}$ is independent of $Y_{00}$; that is, $Y_{00}$ follows a martingale process.\footnote{Relatedly, gsw2022 show that for parallel trends to hold under selection on pre-treatment unobservables, a martingale-type restriction on the untreated potential outcome is necessary.} Therefore, in this framework, even when parallel trends fail, copula stability may still hold under a particular mapping between $Y_{10}$ and $Y_{00}$. It is important to note however that the comonotonicity assumption may not be plausible in some applications, and therefore copula stability would fail. Finally, from the above arguments, it is straightforward to show that if selection was on post-treatment (untreated) potential outcomes, $D=\mathbbm{1}\{Y_{10}>c\}$, then copula stability would also require comonotonicity between pre- and post-treatment untreated potential outcomes.

We next consider selection on time-varying shocks, an example that will not be compatible with copula stability in general.

example[Selection on time-varying shocks] Consider a setting where selection into treatment depends only on the post-treatment shock, specifically: \begin{eqnarray*} \left\{ \begin{array}{lcl} Y_{t0} &=& f_t + \lambda + \varepsilon_{t},\\ \\ D &=& \mathbbm{1}\{\varepsilon_{1}>c\}, \end{array} \right. \end{eqnarray*} where $\lambda$ and $\varepsilon_{t}$ are time-invariant and time-varying unobservables, respectively, and $f_t$ captures nonstochastic time trend. Suppose further that $\left(\begin{array}{c}\lambda\\ \varepsilon_{0} \\ \varepsilon_{1} \end{array}\right)\sim N(0,\Sigma)$, $\Sigma=\left(\begin{array}{ccc}\sigma_{\lambda}^2 & 0 & 0\\ 0 & \sigma_{\varepsilon_0}^2 & 0\\ 0& 0 & \sigma_{\varepsilon_1}^2 \end{array}\right)$. In this model, the copula $C_{Y_{t0},D}(u,q)$ is Gaussian, and $Corr(\varepsilon_{1},Y_{00})=0$ while $Corr(\varepsilon_{1},Y_{10})=\frac{\sigma_{\varepsilon_1}}{\sqrt{\sigma^2_{\lambda}+\sigma^2_{\varepsilon_{1}}}}$. Hence, our CS assumption fails to hold. Note, however, that in this DGP, both CiC and PT assumptions fail to hold as well. Again, in this example, we can demonstrate that if selection was on pre-treatment shocks, $\varepsilon_0$, then copula stability would be violated by similar arguments.
example[Firm-specific wage increases] Consider an individual working for a firm $f$. Let $Y_{00}$ and $Y_{10}$ be, respectively, the worker's wage in periods 0 and 1 in the absence of a minimum-wage increase $D$. The worker's wage in period 1 in the absence of the minimum-wage increase would be her wage in period 0 plus any increase that firm $f$ provides to its workers. Suppose there is an increase of $100R_f\%$ in workers' salaries in firm $f$. Then, we can write $Y_{10}=(1+R_f)Y_{00}$. If the wage increase rate $R_f$ is jointly independent of the baseline salary $Y_{00}$ and the policy $D$, i.e., $R_f \perp (Y_{00},D)$, then copula stability holds.\footnote{See proof in Appendix (ref). We allow $D$ to depend on $Y_{00}$ and the treatment effect.} On the other hand, parallel trends as well as distributional parallel trends fail to hold unless $D$ is (mean) independent of $Y_{00}$ (i.e., unless random assignment holds). In practice, the wage increase rate $R_f$ may depend on some firm-level characteristics $X_f$ that can explain $R_f$, $Y_{00}$, and $D$. Then, our horizontal copula stability assumption will hold conditional on firm characteristics $X_f$, i.e., $C_{Y_{10},D\vert X_f}(u,q)=C_{Y_{00},D\vert X_f}(u,q)$ for all $u$, but not unconditionally.

Next, we proceed to our second identifying assumption, which requires the strict monotonicity of the horizontal copula.

assumption[Strictly increasing horizontal copula] The function $u \mapsto C_{Y_{10},D}(u,q)$ is strictly increasing on $[0,1]$.

While Assumption (ref) is less critical for our bounding approach, it allows us to simplify the expression of our bounds. It is essentially a restriction on the type of dependence between the potential outcomes and group membership. Many well-known parametric classes of copulas satisfy this assumption, e.g. Frank, Gumbel, Joe, or Gaussian copulas among many others. It excludes, however, extreme types of dependence captured by the Fr\'echet-Hoeffding copula bounds, i.e. $C(u,v)=\min\{u,v\}$ and $C(u,v)=\max\{u+v-1,0\}$. It is worth noting that this assumption is implied by some support conditions on the potential outcome distributions, as we show in the following result.

lemmaIf $\mathbb Y_{t0|1} \subseteq \mathbb Y_{t0|0}$, then $u \mapsto C_{Y_{t0},D}(u,q)$ is strictly increasing on $Ran F_{Y_{t0}}$ for $t \in \{0,1\}$.

The main implication of the above lemma is that for continuous potential outcome distributions, we have $Ran F_{Y_{t0}}=[0,1]$, and the strict monotonicity of the copula (Assumption (ref)) is implied by a condition on the support of $Y_{t0}$, $\mathbb Y_{t0|1} \subseteq \mathbb Y_{t0|0}$. That is, the support of the untreated potential outcome of the treatment group is included in the support of the untreated potential outcome of the control group. The support condition imposed in AtheyImbens2006 on the scalar unobservable in the CiC model implies this support condition on the untreated potential outcome.

remarkHere, we formally compare our copula stability assumption with the one introduced in CallawayLi2019. To see this, let us define $\Delta Y_{t0}=Y_{t0}-Y_{(t-1)0}$, CallawayLi2019 require $C_{\Delta Y_{t0},Y_{(t-1)0}|D=1}(\cdot,\cdot)=C_{\Delta Y_{(t-1)0},Y_{(t-2)0}|D=1}(\cdot,\cdot)$. As can be seen, their assumption imposes a dependence stability on different objects than ours, and it requires at least three time periods of panel data. In addition, unlike us, their identification results require an additional independence condition between the change in the untreated potential outcome and treatment assignment, $\Delta Y_{t0}\perp D$.

Main Identification Result

We next state our main identification result:

theoremSuppose that $\mathbb Y_{t0|1} \subseteq \mathbb Y_{t0|0}$ for $t \in \{0,1\}$, then under Assumptions (ref) and (ref), the bounds on the unobserved counterfactual distribution $F_{Y_{10}|D=1}(.)$ are: {\begin{eqnarray*} &&\lim_{\tilde y \downarrow y}\sup\left\{F^{LB}(t): t\leq \tilde y \; & \; t \in \mathbb Y_{10|0} \cup\{-\infty \} \right\} \&&\qquad \leq F_{Y_{10}|D=1}(y)\leq \nonumber\&& \lim_{\tilde y \downarrow y}\sup\left\{ F^{UB}(t): t\leq \tilde y \; & \; t \in \mathbb Y_{10|0} \cup\{-\infty \} \right\} \end{eqnarray*} for all $y \in \mathbb R$, where \begin{eqnarray*} F^{LB}(t)&=&F_{Y_0|D=1}\left(Q^{\mathbb R,+}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(t)\right)-\right)\; \\\; F^{UB}(t)&=&F_{Y_0|D=1}\left(Q^{\mathbb R,-}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(t)\right)\right). \end{eqnarray*}} The above bounds are shown to be sharp when $\overline{\operatorname{Ran}} F_{Y_{0}}$ is closed.\footnote{We conjecture that the sharpness statement remains valid without this closure requirement, but it requires a more involved construction of the subcopula that rationalizes the data.}

Theorem (ref) provides a general (partial) identification result on the counterfactual distribution of the treatment group for any type of potential outcome variables (discrete, continuous, or mixed). Our result neither imposes any restriction on the heterogeneity of potential outcomes within a period nor across periods. We specifically do not impose restrictions on individual treatment effects, $Y_{11}-Y_{10}$, or the evolution of the distribution of the untreated potential outcome across time, $F_{Y_{t0}}, t \in\{0,1\}$. The formal proof is relegated to Appendix (ref). The derived bounds may look involved since we aim to provide a general formulation that covers any type of distribution and want to ensure that our bounds are indeed right-continuous.\footnote{As recognized by AtheyImbens2006, their upper bound in the discrete outcome case may be left-continuous, and therefore may not satisfy the properties of a cdf.} The bounds simplify for some special cases as we will illustrate in Corollary (ref) below.

The intuition behind our (partial) identification result is very simple and can be summarized as follows: In the first period, we identify the joint distribution $\mathbb P(Y_{00}\leq y, D=0)$ and both marginal distributions, $\mathbb P(Y_{00}\leq y)$ and $q$. Using the Sklar result, we can recover the horizontal subcopula $C_{Y_{00},D}(u,q)$ on $RanF_{Y_{0}}$, and thereby the rank mapping $\Gamma(\cdot)$ on $Ran F_{Y_0|D=0}$. Then, since we assume the rank mapping to be stationary across time, we can then carry it over from the pre-treatment period to the post-treatment period to recover the treatment group's distribution of the untreated potential outcome, $F_{Y_{10}|D=1}$, as follows:

eqnarray[eqnarray omitted — 132 chars of source]

The main reason behind the partial identification is that in the first period we recover the subcopula $C_{Y_{00},D}(\cdot,q)$ only on $Ran F_{Y_0}$ ($\Gamma(\cdot)$ only on $Ran F_{Y_0|D=0}$), and we do not know the rank mapping outside this range. We provide a graphical illustration of these functions as well as our bounds in the context of a minimum-wage numerical example in Appendix (ref).

In the case of continuous potential outcomes, $Ran F_{Y_{0}}=[0,1]$, our bounds shrink to a point because the pre-treatment period allows us to recover the entire rank mapping that we carry over to the post-treatment period, as we show in the following corollary of Theorem (ref).

corollaryUnder Assumption (ref), whenever $\mathbb Y_{t0|1} \subseteq \mathbb Y_{t0|0}$ for $t \in \{0,1\}$ and the cdfs $F_{Y_{t0}\vert D=d}(.)$, $t, d \in \{0,1\}$ are continuous, we have: $$F_{Y_{10}|D=1}(y)=F_{Y_0|D=1}\left(Q^{\mathbb R,-}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(y)\right)\right)$$ for all $ y \in \mathbb R$.

The proof of this corollary is in Appendix (ref). Corollary (ref) recovers the point-identification result obtained in AtheyImbens2006. AtheyImbens2006 provide (partial) identification results for two types of potential outcomes relying on different assumptions for each of the two cases: (i) continuous outcomes that are strictly monotonic in a scalar unobservable, (ii) discrete outcomes that are monotonic in a scalar unobservable. By contrast, Theorem (ref) establishes a unifying identification result for any type of outcome under consideration. In addition to the connection to our identification result, there is a link between the CiC assumptions and our copula stability condition for continuous outcomes. We provide details on this connection and compare the two identification approaches in Section (ref).

remarkThe bounds in Theorem (ref) may cross, indicating that at least one of our key assumptions does not hold. We present the formal testable implication in the following subsection.

Multiple pre-treatment periods

In this section, we characterize our bounds in the presence of multiple pre-treatment periods. Suppose we have the following model with $T_0+1$ pre-treatment periods:

eqnarray[eqnarray omitted — 148 chars of source]

We impose the following stability restriction on the horizontal copula at $q$ over multiple pre-treatment periods $t=-T_0,\dots,0$.

assumption[Dependence stability over multiple periods] For all $t=-T_0,\dots,0$ and $u \in [0,1]$, $$C_{Y_{t0},D}(u,q)=C_{Y_{10},D}(u,q).$$

The following theorem generalizes Theorem (ref) to the multiple-period case under Assumption (ref). Corollary (ref) then provides testable restrictions of our model assumptions.

theoremSuppose that $\mathbb Y_{t0|1} \subseteq \mathbb Y_{t0|0}$ for $t \in \{-T_0,\dots,0\}$. If Assumptions (ref) and (ref) hold, then the bounds on the unobserved counterfactual distribution $F_{Y_{10}|D=1}(.)$ are: {\begin{eqnarray*} &&\lim_{\tilde y \downarrow y}\sup\left\{ \max_{t \in \{-T_0,\dots,0\}}F_t^{LB}(s): s\leq \tilde y \; & \; s \in \mathbb Y_{10|0} \cup\{-\infty \} \right\} \&& \qquad \leq F_{Y_{10}|D=1}(y) \leq \nonumber\&& \lim_{\tilde y \downarrow y}\sup\left\{ \min_{t \in \{-T_0,\dots,0\}}F_t^{UB}(s): s\leq \tilde y \; & \; s \in \mathbb Y_{10|0} \cup\{-\infty \} \right\}, \end{eqnarray*} for all $y \in \mathbb R$, where for $t\in\{-T_0,\dots,0\}$ \begin{eqnarray*} F_t^{LB}(s)&=&F_{Y_t|D=1}\left(Q^{\mathbb R,+}_{Y_t|D=0}\left(F_{Y_{1|D=0}}(s)\right)-\right)\; \\\; F_t^{UB}(s)&=&F_{Y_t|D=1}\left(Q^{\mathbb R,-}_{Y_t|D=0}\left(F_{Y_{1|D=0}}(s)\right)\right). \end{eqnarray*}}
corollary[Model's Testable Restriction] Suppose that $\mathbb Y_{t0|1} \subseteq \mathbb Y_{t0|0}$ for $t \in \{-T_0,\dots,0\}$, then if Assumptions (ref) and (ref) hold, the following inequalities must be satisfied: \begin{eqnarray*} &&\Delta(y)\leq 0 \quad \forall y \in \mathbb{Y}_{10|0}, where \end{eqnarray*} $\Delta(y)\equiv \max_{t \in \{-T_0,\dots,0\}}F_{Y_t}\left(Q^{\mathbb R,+}_{Y_t|D=0}\left(F_{Y_{1|D=0}}(y)\right)-\right)- \min_{t \in \{-T_0,\dots,0\}}F_{Y_t}\left(Q^{\mathbb R,-}_{Y_t|D=0}\left(F_{Y_{1|D=0}}(y)\right)\right)$.

We illustrate the arguments in Theorem (ref) and Corollary (ref) in a numerical example motivated by our minimum wage setting in the presence of multiple pre-treatment periods in Section (ref).

figure[figure omitted — 2,082 chars of source]
figure[figure omitted — 1,625 chars of source]
remark[Staggered adoption design] Suppose that we observe multiple post-treatment periods, $t=1,\dots,T_1$, where $D\in\{1,2,\dots,T_1,\infty\}$. For $t=1,\dots,T_1$, $D=t$ denotes the group that adopts the treatment in period $t$, and $D=\infty$ denotes the control group that is never-treated. Let $\mathbb{P}(D\leq \tau)=q_{\tau}$ for $\tau\in\{\infty,1,\dots, T_1-1\}$ and $Y_t^{\infty}$ denote the potential outcome in the control state. We can extend our identification approach to this setting under a suitable copula stability assumption, specifically assuming $C_{Y_{t}^{\infty},D}(\cdot,q_{\tau})=C_{Y_{(t-1)}^{\infty},D}(\cdot,q_{\tau})$ for $\tau\in\{\infty,1,\dots,T_1-1\}$ and $t\in\{-T_0+1,\dots,-1,0,1,\dots, T_1\}$.

Numerical Illustration

Here we illustrate the CS bounds with two pre-treatment periods as well as the testable restrictions in the context of a minimum-wage numerical example. Suppose that both treatment and control groups have a pre-existing minimum wage set at $c_{0}$ in the pre-treatment periods ($t=-1,0$). In the post-treatment period ($t=1$), the minimum wage increases for the treatment group to $c_{1}$. We consider two cases: (i) all model assumptions hold (Figure (ref)), (ii) all assumptions except copula stability hold (Figure (ref)).\footnote{In Appendix (ref), we demonstrate a third case, where copula stability holds, while the strict monotonicity of the horizontal copula is violated. This case demonstrates that we can detect violations of our model assumptions with only one pre-treatment period.}

Figure (ref) demonstrates that when copula stability holds for multiple pre-treatment periods, it can have significant gain in terms of identification as the multi-period CS bounds point-identifies the counterfactual distribution on a larger portion of its support in Panel (c) relative to Panels (a) and (b). Figure (ref)(f) provides our model testable restriction, specifically $\Delta(y)\leq 0$, which holds in this case. Furthermore, Panels (d) and (e) of Figure (ref) present $C_{Y_{t0},D}$ and $\Gamma_t$, respectively, for $t=-1,0$, which are equal on the intersection of their respective ranges.

Next, we demonstrate the case where copula stability only holds for $t\in\{0,1\}$, but not $t\in\{-1,1\}$. In Figure (ref), Panel (a) shows that using the pre-treatment period $t=-1$ only to construct the CS bounds yields bounds that do not include the counterfactual, whereas Panel (b) shows that the counterfactual is included in the CS bounds with pre-treatment period $t=0$ only. When considering the CS bounds using both pre-treatment periods in Figure (ref)(c), we note that the CS lower bound is greater than the CS upper bound, and our model testable restriction is violated as indicated by Figure (ref)(f). Relatedly, Figures (ref)(d) and (ref)(e) demonstrate that the mappings $C_{Y_{t0},D}(\cdot,q)$ and $\Gamma_{t}(\cdot)$, respectively, are not equal for $t=-1,0$, indicating a violation of copula stability.

Connection to Changes-in-Changes

In this section, we elaborate on the connection between our copula stability assumption and the CiC conditions in AtheyImbens2006. We first show the equivalence between copula stability and the CiC conditions for continuous outcome distributions. Second, while the identification results in AtheyImbens2006 do not account for mixed outcomes, a researcher might still rely on their estimand. Here, we demonstrate that a na\"ive implementation of the CiC approach leads to a point/bound estimand that might not include the true counterfactual, whereas our CS bounds will. Finally, for discrete outcomes, we demonstrate using an analytical example that copula stability can be compatible with multi-dimensional unobserved heterogeneity, whereas the CiC conditions require unobserved heterogeneity to be uni-dimensional.

Continuous outcomes

The following result demonstrates that the CiC conditions for continuous, strictly increasing outcome distributions are equivalent to our copula stability assumption. In Appendix (ref), we demonstrate how this result extends to all continuous outcomes. For other outcome distributions, this equivalence does not hold in general.

claimAssume the cdfs $F_{Y_{t0}}(.)$ for $t \in \{0,1\}$ are continuous and strictly increasing, then the following two statements are equivalent: \begin{enumerate} • $C_{Y_{00},D}(u,q)=C_{Y_{10},D}(u,q)$ for all $u \in [0,1]$. • There exist two strictly increasing functions $h_t(.), t \in \{0,1\}$ and two uniformly distributed random variables over $[0,1]$, $U_{00}$ and $U_{10}$, such that $Y_{t0}=h_t(U_{t0})$ and $U_{00}|D=d \sim U_{10}|D=d$ for $d \in \{0,1\}$. \end{enumerate}

The proof of this claim is in Appendix (ref). The main intuition behind it is that for this class of distributions we can write $Y_{t0}=Q_{Y_{t0}}^{\mathbb{R},-}(U_{t0})$, where $U_{t0}=F_{Y_{t0}}(Y_{t0})\sim \mathcal{U}[0,1]$. As a result, the marginal distribution of $U_{t0}$ is stable across time by construction and the stability of the copula between $U_{t0}$ and $D$ is necessary and sufficient for the stability of $U_{t0}|D$, which is the conditional time invariance assumption in AtheyImbens2006. Its equivalence to our copula stability assumption follows from the invariance of the copula under strictly monotonic transformations.

Mixed outcomes

Here, we demonstrate that for mixed outcomes the CiC point/bound estimand may not cover the true counterfactual distribution in the context of the numerical minimum-wage example in Section (ref).

The CiC bounds in the discrete case are defined for any $s \in \mathbb{Y}_{1|0}$ as follows for $t\in\{-1,0\}$,

eqnarray*[eqnarray* omitted — 261 chars of source]

In the example illustrated in Figure (ref), we have $\mathbb{Y}_{t|0} = \mathbb{R}^+$, and $Y_t|D=0$ has a strictly increasing cdf in $\mathbb{R}^+$. Then the following simplifications hold:

multline*[multline* omitted — 236 chars of source]

and

multline*[multline* omitted — 336 chars of source]

where the inequality becomes strict at points of discontinuity.

More importantly, we can see that $F_{t,\text{CiC}}^{\text{LB}}(s)=F_{t,\text{CiC}}^{\text{UB}}(s)$, since $Q_{Y_t|D=0}^{\mathbb{Y}_{t|0},+}(u)=Q_{Y_t|D=0}^{\mathbb{Y}_{t|0},-}(u)$ for $u\in[0,1]$. However, this CiC point estimand is different from the true counterfactual of interest $F_{Y_{10}|D=1}$, as shown in Figure (ref).

Therefore, in this case, our bounds contain the CiC (point/bound) estimands and the true counterfactual \[F_{t,\text{CiC}}^{\text{UB}}(s)\neq F_{Y_{10}|D=1}(s), \text{ where } \{F_{t,\text{CiC}}^{\text{UB}}(s),F_{Y_{10}|D=1}(s)\}\in [F_t^{\text{LB}}(s), F_t^{\text{UB}}(s)].\]

In sum, in this mixed-outcome example, if the researcher ignores the discontinuity and applies the CiC point estimand or applied the CiC bounds for the discrete case, their estimand will not cover the true counterfactual, as shown in Figure (ref).

figure[figure omitted — 668 chars of source]

Discrete outcomes

For the case of discrete outcomes, the following example illustrates that our identifying assumption can be compatible with multi-dimensional unobserved heterogeneity, whereas the CiC conditions require scalar unobserved heterogeneity.

example[Binary outcome model with multidimensional unobserved heterogeneity] Consider the following model \begin{eqnarray} Y_t &=&1-\mathbbm{1}\{\eta t D + U_t \leq c_t,\tilde{\eta} t D + \tilde{U}_t \leq \tilde{c}_t\}, t=0,1,\nonumber\\ D&=&\mathbbm{1}\{V>q\}, \end{eqnarray} where $(Y_0,Y_1,D)$ is an observed random vector, $(\eta, \tilde{\eta}, U_0, U_1, \tilde{U}_0,\tilde{U}_1)$ is a latent random vector, and $(c_t, \tilde{c}_t)$ is a constant vector. For simplicity, we normalize $U_t$, $\tilde{U}_t$ and $V$ to be uniformly distributed on $[0,1]$. The untreated potential outcome $Y_{t0}$ is $$Y_{t0} =1-\mathbbm{1}\{U_t \leq c_t,\tilde{U}_t \leq \tilde{c}_t\}.$$ For instance, $D$ could be the student loan forgiveness program, $Y_t$ could be a college attendance decision, $U_t$ and $\tilde{U}_t$ could respectively be father's and mother's wealth in the absence of the program. This model assumes that an individual decides to attend college if at least one of the parents' wealth is above a (parent-specific) threshold, whether they were to receive the loan forgiveness program or not. While the CiC approach does not allow multidimensional unobserved heterogeneity, we show in Appendix (ref) that for the wide class of Archimedean copulas the stability of the dependence structure of the latent variables $(U_t,\tilde{U}_t,V)$ over time implies our copula stability assumption.

While our assumption accommodates a broader class of binary outcome models than the CiC model assumption, it does not necessarily yield tighter bounds. As illustrated in Figure (ref) in the Online Appendix, both approaches produce the same bounds in the discrete-outcome case.

Policy-relevant parameters: Social welfare treatment effect on the treated (SWTT)

Building on our unifying, partial identification result for the counterfactual distribution, we provide a class of policy-relevant parameters that quantify the impact of policy on social welfare in the entire population, subpopulations in the lower tail of the distribution or over any interquantile range of the distribution. In general, when a policymaker decides to implement a new policy such as an increase in the legal minimum wage or legal minimum working time, she expects the policy to have a specific social welfare impact. The social welfare function used by the policymaker is not necessarily known to the researcher, however. For instance, the policymaker may consider social welfare functions that put more weight on specific subpopulations, such as lower-income individuals, or considers only social welfare functions with specific properties like social welfare functions that respect the Pigou-Dalton principle of transfers\footnote{The Pigou-Dalton principle states that a transfer of income from a higher-ranked individual to a lower-ranked individual that does not change their ranks is always desirable.} or the rank-dependent social welfare functions introduced by Mehran1976 (Mehran1976).\footnote{See Aabergeetal2013 (Aabergeetal2013) for a detailed discussion.}

As we clarify below, the widely used average treatment effect on the treated (ATT) corresponds to the case where the policymaker is inequality-neutral. If the policymaker is averse to inequality, however, the ATT would not be an adequate causal parameter to measure the impact of the policy or judge its effectiveness.

For this particular reason, we propose a class of parameters of interest that measure the causal effect of a particular policy in terms of a social welfare function,

eqnarray*[eqnarray* omitted — 223 chars of source]

where $SW_{\omega}(F_X)=\int_{0}^{1}\omega(\tau)Q^{\mathbb R,-}_{X}(\tau)$ denotes the social welfare function associated with a specific distribution $F_X$, and $\omega(\tau) \in [0,1]$ is a weighting function. This social welfare function can be alternatively viewed as a weighted average of the outcomes of individuals $i$ where the weights depend on the rank of $X_i$, $SW_\omega=\int{X_i\omega(Rank(X_i))}di$ KitagawaTetenov2021. Since the social welfare function essentially weights different quantiles of the distribution, the choice of the functional form of the weighting function relates to the inequality aversion of the policymaker and the extent thereof. We next consider several examples of weighting functions and discuss the properties of the social welfare functions they imply.

Before we proceed, it is important to emphasize that, while in many applications where measuring inequality is a concern, the outcome $Y$ is typically income or wages, our framework allows $Y$ to denote other outcomes as well as functions of different outcomes, such as consumption, income and/or human capital. Our $SWTT_{\omega}$ is also a generalization of the quantile treatment effect parameter discussed in Abadieetal2002, Firpo2007, and FrohlichMelly2008.

Generalized Gini social welfare function

The class of generalized Gini social welfare functions is the class of rank-dependent, equality-minded social welfare functions which satisfy the Pigou-Dalton principle of transfers and is given by

eqnarray*[eqnarray* omitted — 58 chars of source]

where $\Lambda(\cdot):[0,1] \mapsto [0,1]$ is a convex, non-increasing, and non-negative function with boundary conditions $\Lambda(0)=1$ and $\Lambda(1)=0$. This class admits the equivalent representation as a weighted sum of quantiles with weighting function $\omega(\tau)=\frac{\partial(1-\Lambda(\tau))}{\partial \tau}$,

eqnarray*[eqnarray* omitted — 103 chars of source]

As a result, the class of social welfare treatment effect parameters we introduce include this class as a special case. We proceed to present two important special cases of this class of social welfare functions, specifically the utilitarian and Gini social welfare functions.

Utilitarian welfare function

When $\omega(\tau)=1$, we have $SW_{\omega}(F_X)=\int_{0}^{1}Q^{\mathbb R,-}_{X}(\tau)d\tau$ $=\mathbb E[X].$ This corresponds to the additive welfare function and in this case our proposed parameter boils down to the ATT, i.e. $SWTT_{\omega}=ATT$. The ATT is therefore the appropriate parameter if the policymaker weights subpopulations at different quantiles of the distribution equally.

Gini social welfare function

When $\omega(\tau)=2(1-\tau)$, we have $SW_{\omega}(F_X)=\int_{0}^{1}2(1-\tau)Q^{\mathbb R,-}_{X}(\tau)d\tau$ $=\mathbb E[X]\left(1-I_{Gini}(F_X)\right),$ where $I_{Gini}(F_X)\equiv \frac{\int_{0}^{1}(2\tau-1)Q^{\mathbb R,-}_{X}(\tau)d\tau}{\mathbb E[X]}$ is the widely used Gini inequality index, see Sen1974. $SW_{\omega}(F_X)$ reflects the trade-off between the mean and (in)equality in the distribution $F_X$. The product $\mathbb E[X]I_{Gini}(F_X)$ is a measure of the loss in social welfare due to inequality in the distribution $F_X$. In that case, $SWTT_{\omega}$ captures the impact of the policy using the Gini social welfare function, see BlackorbyDonaldson1978 and Weymark1981. In other words, if the policymaker implements the policy in order to reduce the level of inequality measured by the Gini index, this parameter is the most adequate to judge the impact of this policy.

Second-order dominance

In many cases, when it is possible to do so, most inequality-averse policymakers like to rank distribution functions consistently with second-degree dominance. For instance, we say $F_{Y_{11}|D=1}$ second-order dominates $F_{Y_{10}|D=1}$ if and only if:

eqnarray*[eqnarray* omitted — 216 chars of source]

for all $u \in [0,1]$ and holds strictly for some $u$. In this special case, we have $\omega(\tau)=\mathbbm{1}\{\tau \leq u\}$. It is possible, however, that the observed and counterfactual distribution cannot be ranked using this criterion. Furthermore, the policy's objective may be to reduce inequality in a specific part of the distribution. We therefore consider the following quantile-specific Gini social welfare functions.

Quantile-specific lower tail Gini social welfare function

In the Gini social welfare function discussed above, we assume that the policymaker is interested in the inequality of the whole population. Some policies may be concerned with reducing inequality up to specific quantiles of the distribution, such as minimum-wage policies Dube2019,Cengizetal2019. To quantify the impact of the policy on lower-tail quantiles, we extend the quantile-specific lower-tail Gini social welfare measures introduced in Aabergeetal2013 (Aabergeetal2013) for continuous distributions to any type of distribution in order to accommodate the possibility of discontinuities resulting from censoring or bunching. To do so, we introduce the random variable $X^u=Q_X^{\mathbb{R},-}(V)$, where $V\sim \mathcal{U}[0,u]$ for $u\in(0,1]$.\footnote{For $u\in Ran F_X$, $F_{X^u}(x)=\mathbb P(X\leq x|X\leq Q_X^{\mathbb{R},-}(u))$ for any $x\leq Q_X^{\mathbb{R},-}(u)$, thereby yielding the same truncated random variable introduced in Aabergeetal2013(Aabergeetal2013). For $u\notin Ran F_X$, $X^u$ remains a well-defined random variable.} We relegate the derivations relevant to this section to Appendix (ref).

With this definition of $X^u$, we can show that the lower-tail Gini social welfare function can be decomposed into $\mathbb E[X^u]$ and the Gini coefficient associated with $F_{X^u}$ as follows $$\int_{0}^{1}\frac{2}{u^2}(u-\tau)\mathbbm{1}\{\tau \leq u\}Q^{\mathbb R,-}_{X}(\tau)d\tau=\mathbb E[X^u]\left(1-I_{Gini}(F_{X^u})\right),$$ where $I_{Gini}\left(F_{X^u}\right)\equiv\frac{\int_{0}^{1}(2\tau-u)\mathbbm{1}\{\tau\leq u\}Q^{\mathbb R,-}_{X}(\tau)d\tau}{u^2\mathbb E[X^u]}$ is the lower-tail Gini coefficient at $u$ defined in Aabergeetal2013. Therefore, $SWTT_{\omega}$ with $\omega(\tau)=\frac{2}{u^2}(u-\tau)\mathbbm{1}\{\tau \leq u\}$ yields the following,

eqnarray*[eqnarray* omitted — 160 chars of source]

and is interpreted as the Quantile-$u$ lower tail Gini social welfare treatment effect on the treated.

Interquantile Gini social welfare function

Since policies may target other parts of the distribution, such as the upper tail, we can generalize these quantile-specific social welfare treatment effect measures to any range of quantiles $[\underline{u},\overline{u}]$ a researcher may be interested in. Specifically, let $\underline{u}\in[0,1]$, $\overline{u}\in[0,1]$, $\underline{u}<\overline{u}$, $V\sim \mathcal{U}[\underline{u},\overline{u}]$, and $X^{\underline{u},\overline{u}}=Q_X^{\mathbb{R},-}(V)$. A derivation of $F_{X^{\underline{u},\overline{u}}}$ is relegated to Appendix (ref). Now by letting $\omega(\tau)=\frac{2}{(\overline{u}-\underline{u})^2}(\overline{u}-\tau)\mathbbm{1}\{\underline{u}<\tau\leq \overline{u}\}$, we obtain the Gini social welfare function specific to the quantile range $[\underline{u},\overline{u}]$, $$SW_{\omega}(\underline{u},\overline{u})=\int_0^1\frac{2}{(\overline{u}-\underline{u})^2}(\overline{u}-\tau)\mathbbm{1}\{\underline{u}<\tau\leq \overline{u}\}Q_X^{\mathbb{R},-}(\tau)d\tau=\mathbb E[X^{\underline{u},\overline{u}}](1-I_{Gini}(F_{X^{\underline{u},\overline{u}}})),$$ where $\mathbb E[X^{\underline{u},\overline{u}}]\equiv\int_{\underline{u}}^{\overline{u}}Q_X^{\mathbb{R},-}(\tau)d\tau$ and $I_{Gini}\left(F_{X^{\underline{u},\overline{u}}}\right)\equiv \frac{\int_0^1(2\tau-\underline{u}-\overline{u})\mathbbm{1}\{\underline{u}<\tau\leq \overline{u}\}Q_X^{\mathbb{R},-}(\tau)d\tau}{(\overline{u}-\underline{u})^2\mathbb E[X^{\underline{u},\overline{u}}]}$.\footnote{This definition extends the upper tail Gini coefficient to any quantile range $[\underline{u},\overline{u}]$.} The interquantile Gini social welfare treatment effect on the treated over $[\underline{u},\overline{u}]$ is given by

eqnarray*[eqnarray* omitted — 364 chars of source]
remarkIt is important to note that when defining interquantile $SWTT_{\omega}(\underline{u},\overline{u})$ when $[\underline{u},\overline{u}]\subset [0,1]$, caution is required in interpreting these parameters, as we may not be comparing the same population unless certain assumptions hold. However, this concern is shared by most of the existing literature on recovering quantile treatment effects, including Abadieetal2002, Firpo2007, FrohlichMelly2008 and CallawayLi2019, among many others. This issue disappears once we assume rank invariance—i.e., that there exists $U \sim \mathcal{U}[0,1]$ such that $Y_{1d|D=1} = Q_{Y_{1d|D=1}}^{\mathbb{R},-}(U), \quad \text{for } d \in \{0,1\}$.

Empirical Illustration

In this section, we illustrate the CS bounds by revisiting the minimum wage study by Cengizetal2019. This application demonstrates the usefulness of the class of policy-relevant parameters we introduce to examine the impact of the minimum wage increase. In particular, the lower-tail quantile social welfare treatment effect estimates allow us to zoom into the lower tail of the distribution, where we expect the minimum wage to have an impact. Overall, our CS bounds document proportionately larger impacts on the Gini social welfare in the lowest part of the distribution, where the minimum wage increase led to increase in the lower-tail mean and Gini social welfare. We also find that the distributional DiD exhibits violations of monotonicity in the lower tail of the distribution and is therefore not suitable for this application.

This empirical illustration highlights two practical advantages of our approach. First, our CS bounds relieve practitioners from having to take a stance on the support of the outcome of interest. Second, our multi-period CS bounds combine information from multiple pre-treatment periods to tighten the bounds on the parameters of interest and to simultaneously test the model assumptions.

Data and Implementation

Cengizetal2019 examine 138 prominent state-level minimum wage increases between 1979 and 2016 using the individual-level NBER-merged Outgoing Rotation Group Earnings Data of the Current Population Survey. Their goal is to examine the impact of the policy on the wage distribution around the minimum wage, as illustrated in Figure (ref). In order to make the empirical illustration of the multi-period CS bounds succinct, we focus on two pre-treatment periods, 2010 and 2011, and one post-treatment period, 2015, and examine the distributional impact of a nontrivial minimum wage increase of \$0.25 or more.\footnote{Note that starting 2009, the federal minimum has been \$7.25, so a minimum wage increase of \$0.25 or more constitutes an increase of more than 3%. This definition of the treatment variable was also used in the empirical illustration in RothSantanna2021.} For the purpose of this empirical illustration, we focus on the subgroup of states that had a pre-treatment minimum wage of \$8 or higher. We report the results for the remaining states in Appendix (ref).

Table (ref) presents the summary statistics for hourly wage of both treatment and control groups in all three periods we consider. For both subgroups, the summary statistics show that the mean and standard deviation is different across treatment and control groups within the same year as well as within groups before and after the treatment.

In order to estimate the CS bounds on the counterfactual, we rely on Lemma (ref) to re-write the lower bound in a manner that admits straightforward numerical computation, specifically for $y\in\mathbb{Y}_{10|1}$ and for a given pre-treatment period $t$

eqnarray[eqnarray omitted — 205 chars of source]

$F_{Y_{10}|D=1}^{LB,t}(y)$ and $F_{Y_{10}|D=1}^{UB,t}(y)$ are estimated by their sample analogues, $\widehat{F}_{Y_{10}|D=1}^{LB}(y)$ and $\widehat{F}_{Y_{10}|D=1}^{UB}(y)$, respectively, by replacing $F_X$ and $Q_X^{\mathbb{R},-}$ by their empirical counterparts, $\widehat{F}_X$ and $\widehat{Q}_X^{\mathbb{R},-}$, respectively.

table[table omitted — 1,009 chars of source]
figure[figure omitted — 1,573 chars of source]
figure[figure omitted — 560 chars of source]

The distributional DiD and CiC point estimators of $F_{Y_{10}|D=1}(y)$ are given by

eqnarray[eqnarray omitted — 298 chars of source]

The CS bounds on the counterfactual as well as the observed factual distribution $\widehat{F}_{Y_1|D=1}$ can then be used to obtain the following sample analogues of the lower and upper bounds on the SWTT.\footnote{We compute the integral numerically using a grid with a step size of $0.01$.} For $t\in\{-1,0\}$, we obtain the following CS bounds estimator for the SWTT parameter

eqnarray[eqnarray omitted — 330 chars of source]

Similarly, we compute the multi-period CS bounds on the SWTT parameters.

To compute the SWTT parameters for the distributional DiD and CiC point estimators, we use the following

eqnarray[eqnarray omitted — 348 chars of source]

Bounds on the counterfactual distribution

Figure (ref) presents the observed distribution of the treatment group in 2015, $\widehat{F}_{Y_1|D=1}$, as well as the CS bounds, distributional DiD and CiC point estimators of the counterfactual distribution using 2010 and 2011 as pre-treatment periods. Since the minimum wage is likely to have an impact on the bottom of the distribution, we present those figures for the bottom quartile of the wage distribution where the minimum wage increase is likely to have an impact.\footnote{We relegate the figures of the entire distribution to Figure (ref) in the online appendix.}

First, we examine the CS bounds on the counterfactual distribution using each of the pre-treatment periods separately in Figure (ref)(a) and (ref)(b), respectively. Comparing the observed (factual) distribution with the CS bounds on the counterfactual using each of the pre-treatment periods, we note an obvious change in the censoring point as expected in the context of a minimum wage increase. For instance, in Figure (ref)(b), the CS bounds on the counterfactual distribution exhibit a jump slightly above \$8, whereas the observed (factual) distribution exhibits a jump at about \$9. Furthermore, note that both upper and lower bounds satisfy the properties of a cdf. In addition, since the bounds do not cross, we do not have any detectable violation of the assumptions required for our identification approach. We also plot the sample analogue of the horizontal subcopula $C_{Y_{t0},D}(\cdot,q)$ for 2010 and 2011 to provide a visual check of our copula stability assumption in Figure (ref). This plot is the counterpart of DiD pre-trends plots in our context. While this figure does not provide a formal test of the copula stability assumption, it demonstrates that the copulas governing the dependence between $Y_{t0}$ and $D$ for 2010 and 2011 are fairly similar.

Next, we examine the bottom quartile of the distributional DiD counterfactual estimates using 2010 and 2011 as pre-treatment period in Figure (ref)(c) and (ref)(d), respectively. At first glance, we note violations of the monotonicity property of cdfs in both counterfactual distributions, indicating a violation of the testable implication of the identifying assumption of distributional DiD RothSantanna2021. The magnitude of the monotonocity violation is by far greater for the distributional DiD estimate using the 2010 pre-treatment period; the counterfactual estimate “dips” around the pre-treatment minimum wage of \$8, which is the part of the distribution particularly pertinent for the evaluation of the minimum wage increase.

Finally, we also present the CiC point estimator of the counterfactual using both pre-treatment periods in Figure (ref)(e) and (ref)(f), respectively. As demonstrated in Section (ref), the CiC point estimator coincides with the CS upper bound using the same pre-treatment period. This could translate to the CiC suffering from an upward bias in SWTT estimation as evident from comparing (ref) and (ref).

Bounds on treatment effects

Next, we quantify the impact of the minimum wage increase on the wage distribution using the ATT and the Gini SWTT both for the overall distribution as well as its lower tail. We report 95% confidence intervals for all SWTT estimators using standard normal critical values and standard errors obtained using nonparametric bootstrap.\footnote{While the formal proof that these confidence intervals provide adequate coverage asymptotically is beyond the scope of the present paper, we have examined their performance in a simulation study mimicking our minimum wage setting which demonstrates that they provide adequate coverage in finite samples.}

Overall social welfare treatment effects

Table (ref) presents 95% confidence intervals on the ATT and Gini SWTT using the CS bounds, the distributional DiD and CiC point estimators.

table[table omitted — 2,294 chars of source]
table[table omitted — 3,669 chars of source]

When examining Table (ref), we note that the 95% confidence intervals on the CS bounds for the ATT and Gini SWTT include zero, whether we use 2010 and 2011 as pre-treatment periods separately or use them both in the multi-period CS bounds. This is consistent with the expectation that a minimum wage increase is unlikely to change the mean or inequality of the overall wage distribution. When we consider the 95% confidence intervals using the distributional DiD and CiC point estimators, they suggest no improvement in terms of ATT and Gini SWTT, except using the CiC confidence interval that use the 2011 pre-treatment period. As pointed out in Section (ref), the CiC point estimator of the counterfactual coincides with the CS upper bound. As a result, the corresponding SWTT estimator may be upwardly biased.

Lower-tail social welfare treatment effects

In the context of policies such as an increase in the legal minimum wage, the welfare of subpopulations at the lower tail of the wage distribution is an important policy target. Table (ref) provides the lower-tail ATT and Gini social welfare treatment effects, $ATT(u)$ and $Gini~SWTT(u)$ for $u\in\{0.01,0.025,0.05,0.10,0.25\}$, respectively, introduced in Section (ref).

First, we consider the CS bounds using 2010 and 2011 as pre-treatment periods separately as well as the multi-period CS bounds that exploits both pre-treatment periods. Regardless of the pre-treatment year we use, for $u\in\{0.01,0.025,0.05\}$, the 95% confidence intervals on the CS bounds demonstrate statistically significant improvement in terms of lower-tail mean and Gini social welfare. When we consider $u\in\{0.10,0.25\}$, we note that while the CS bounds using the 2011 pre-treatment period demonstrate statistically significant improvements in terms of lower-tail mean and Gini social welfare, the confidence intervals on the CS bounds using the 2010 pre-treatment period are not conclusive on the sign of this impact. Since the multiple-period CS bounds combine the information from both pre-treatment periods, they result in tighter confidence intervals than the CS bounds using 2010 or 2011 by itself for both the lower-tail ATT and Gini SWTT for all quantiles $u$ we consider. These tighter confidence intervals point to improvements both in terms of mean and Gini social welfare up to the lower quartile of the distribution ($u=0.25$). This demonstrates how exploiting the multiple pre-treatment periods can aid to provide tighter bounds that translate to shorter confidence intervals.

Next, we consider the distributional DiD and CiC estimators. The distributional DiD confidence intervals using the 2010 pre-treatment period do not suggest any significant improvement in terms of lower-tail mean and Gini social welfare, whereas the distributional DiD confidence intervals using the 2011 pre-treatment period suggest significant improvements in terms of both lower-tail mean and Gini social welfare for most of the quantiles we consider. When we examine the CiC point estimator, we note that the corresponding confidence intervals suggest significant improvements in terms of mean and Gini social welfare for all of the lower-tail quantiles we consider ($u=0.25$).

The confidence intervals on the lower-tail SWTT parameters demonstrate that the distributional DiD can yield contradictory results that then require an ad-hoc choice by the applied researcher regarding which period to use.\footnote{Since the distributional DiD point estimator of the counterfactual distribution using the 2010 pre-treatment period exhibits monotonicity violations, an applied researcher would likely discard those results and use the distributional DiD estimator using the 2011 pre-treatment period, for which the monotonicity violations are very minor. The selection of the pre-treatment period relies however on a pre-test, which raises the usual post-selection inference concerns. Pre-test bias issues in the context of difference-in-difference designs have been examined in Roth2022.} The confidence intervals based on the CiC point estimator will coincide with the confidence interval on the CS upper bound and may therefore be upwardly biased.

Overall, our empirical application underscores the advantages of the CS bounds in terms of relieving the applied researcher from choosing the pre-treatment period as well as specifying the type of outcome distribution. It also demonstrates how to use the CS bounds on the counterfactual distribution to conduct inference on the SWTT parameters. Finally, The CS bounds on the counterfactual distribution can be used to bound other parameters, such as the parameters examined in Cengizetal2019. We provide these estimates in Appendix (ref).

Conclusion

With the goal of assessing the impact of regulatory policies on social welfare, this paper provides a unifying, partial identification result for the counterfactual distribution of the treatment group in difference-in-difference settings. Exploiting the stability of the dependence (copula) between group membership and the untreated potential outcome across time, our identification result has several advantages: (1) it applies to any outcome distribution, whether continuous, discrete or mixed, (2) it is invariant to monotonic transformations of the outcome, (3) it can allow for nonrandom selection into treatment without restricting the evolution of the marginal distribution of the potential outcomes across time. To quantify the impact of regulatory policies on social welfare, we introduce a broad class of treatment effect parameters. This class includes the ATT as well as the Gini social welfare treatment effect on the treated as a special case. We illustrate the empirical relevance of our results using a minimum wage application revisiting Cengizetal2019.

singlespace

\setcounter{lemma}{0} \setcounter{claim}{0} \setcounter{example}{0}

appendix

Proofs of the main results

An Additional Result

lemmaLet $X$ be a random variable, we then have: \begin{enumerate} • The following bounds are pointwise sharp, \begin{eqnarray} F_{X}\left(Q^{\mathbb R,+}_{X}\left(u\right)-\right) &\leq& u \leq F_{X}\left(Q^{\mathbb R,-}_{X}\left(u\right)\right), for all u \in [0,1]. \end{eqnarray} • Let $\mathbb{X}\subseteq \mathbb{Z}$, $\sup\left\{F_{X}\left(t\right): t\leq x \; \& \; t \in \mathbb Z \cup\{-\infty \} \right\}=F_X(x).$ \end{enumerate}

Before we proceed to provide a proof of the above lemma, we compare the bounds in Lemma (ref)(1) with those used in AtheyImbens2006, hereinafter AI2006, to bound the counterfactual distribution for discrete outcomes. These bounds are given by the following in our notation,

eqnarray[eqnarray omitted — 105 chars of source]

Now note that the upper bound employed in AI2006 only differs from the upper bound in Lemma (ref)(1) in terms the use of $\mathbb{X}$ instead of $\mathbb{R}$. These two quantiles only differ for $u=0$, since $\{x\in\mathbb{R}:F_X(x)\geq 0\}=\mathbb{R}$, whereas $\{x\in\mathbb{X}:F_X(x)\geq 0\}=\mathbb{X}$. As a result, $Q_X^{\mathbb{R},-}(0)=-\infty$ and $F_X(Q_X^{\mathbb{R},-}(0))=0$, whereas $Q_X^{\mathbb{X},-}(0)=\inf \mathbb{X}$ and $F_X(\inf\mathbb{X})\geq 0$. Therefore, our upper bound is lower than the one used in AI2006 for $u=0$.\footnote{Note that this is inconsequential for their identification result, since they provide bounds on the counterfactual distribution on its support, and set it to zero below the infimum of its support and to one above the supremum of its support.}

The lower bound in Lemma (ref)(1) is starkly different from the lower bound in (ref). As we discuss in Section (ref), the lower bound in (ref) equals the upper bound for several examples with mixed outcomes, due to censoring or bunching, because $Q_X^{\mathbb{X},+}(u)=Q_X^{\mathbb{X},-}(u)$ for $u\in[0,1]$ for some mixed outcome distributions. As a result, the lower bound is not valid in the mixed-outcome case in general. In those cases, the AI2006 bounds would not cover the counterfactual distribution. We demonstrate additional numerical examples in Appendix (ref). By contrast, our lower bound is valid and sharp for any outcome distribution. For discrete outcomes, our bounds collapse to theirs in numerical examples provided in Appendix (ref).

proof(Lemma (ref))\\ (1) $Q_X^{\mathbb{R},-}(u)\equiv\inf\{x\in\mathbb{R}:F_X(x)\geq u\}$. We know from the properties of a quantile function that $F_X(Q_X^{\mathbb{R},-}(u)) \geq u$. We now show that this inequality is sharp. Suppose that there exists $\tilde{x} \in \mathbb R: F_X(\tilde{x}) \geq u$ and $F_X(\tilde{x}) < F_X(Q_X^{\mathbb{R},-}(u))$. On the one hand, we have $F_X(\tilde{x}) < F_X(Q_X^{\mathbb{R},-}(u)) \Longrightarrow \tilde{x} < Q_X^{\mathbb{R},-}(u),$ since $F_X$ is nondecreasing. On the other hand, $F_X(\tilde{x}) \geq u \Longrightarrow \tilde{x} \in \{x\in \mathbb R: F_X(x) \geq u\}$. Therefore, $\tilde{x} \geq \inf\{x\in \mathbb R: F_X(x) \geq u\}=Q_X^{\mathbb{R},-}(u),$ which contradicts $\tilde{x} < Q_X^{\mathbb{R},-}(u)$. We next show $F_X(Q_X^{\mathbb{R},+}(u)-)\leq u$. For a fixed $u \in [0,1]$, let us define $\Omega=\{y \in \mathbb R: F_X(y) \leq u\}$. We first show this implication: $z < Q_X^{\mathbb R,+}(u) \Longrightarrow F_X(z) \leq u$. By contradiction, suppose that (i) $z < Q_X^{\mathbb R,+}(u)$ and (ii) $F_X(z) > u$. Take $y \in \Omega$, then by (ii) we have $F_X(z) > u \geq F_X(y)$, which implies $F_X(z) > F_X(y)$, which in turn implies $y \leq z$ since $F_X$ is nondecreasing. Therefore, for all $y \in \Omega,$ we have $\ y \leq z$. It follows that $\sup \Omega \leq z$, i.e., $Q_X^{\mathbb R,+}(u) \leq z$. This leads to a contradiction since $z<Q_X^{\mathbb R,+}(u)$ by (i). Hence, we have shown that $z < Q_X^{\mathbb R,+}(u) \Longrightarrow F_X(z) \leq u$. Second, by definition, we have $F_X(Q_X^{\mathbb{R},+}(u)-)\equiv\sup_{z < Q_X^{\mathbb R,+}(u)} F_X(z) \leq \sup_{z < Q_X^{\mathbb R,+}(u)} u=u$, where the inequality holds from the previous implication. Now we proceed to show that $F_X(Q_X^{\mathbb{R},+}(u)-)\leq u$ is sharp. First, let us show that there does not exist any $\tilde{x} \in \mathbb R$ such that (i) $F_X(\tilde{x}) \leq u$ and (ii) $F_X(Q_X^{\mathbb{R},+}(u)-) < F_X(\tilde{x}-)$. By contradiction, suppose there exists such an $\tilde{x}\in\mathbb R$. From (ii), $\sup_{z < Q^{\mathbb R,+}(u)} $ $F_X(z)$ $\equiv F_X(Q_X^{\mathbb{R},+}(u)-) < F_X(\tilde{x}-) \equiv \sup_{z < \tilde{x}} F_X(z)$, we deduce that $\{z < Q^{\mathbb R,+}(u)\} \subset \{z < \tilde{x}\}$. Therefore, $Q^{\mathbb R,+}(u) < \tilde{x}$. From (i), $F_X(\tilde{x}) \leq u$, we have $\tilde{x} \in \{x \in \mathbb R: F_X(x) \leq u\}$. Therefore, $\tilde{x} \leq \sup\{x \in \mathbb R: F_X(x) \leq u\}=Q^{\mathbb R,+}(u)$, which leads to a contradiction. It follows that there does not exist any $\tilde{x} \in \mathbb R$ such that $F_X(\tilde{x}) \leq u$ and $F_X(Q_X^{\mathbb{R},+}(u)-) < F_X(\tilde{x})$. \\ Second, let us show that there does not exist any $\tilde{x} \in \mathbb R$ such that $F_X(\tilde{x}-) \leq u$ and $F_X(Q_X^{\mathbb{R},+}(u)-) < F_X(\tilde{x}-)$. If $F_X(Q_X^{\mathbb{R},+}(u)-) < F_X(\tilde{x}-)$, then from the previous result, we must have $F_X(\tilde{x}) > u$. Hence, we have $F_X(\tilde{x}) > u \geq F_X(\tilde{x}-)$, which implies $\tilde{x}=Q_X^{\mathbb{R},+}(u)$, which in turn contradicts $F_X(Q_X^{\mathbb{R},+}(u)-) < F_X(\tilde{x}-)$.

Proof of Lemma (ref)

By Sklar's Theorem Nelsen2006, there is a unique subcopula $C_{Y_{10}, D}$ determined on $Ran F_{Y_{10}} \times \{q\}$, such that the following hold:

eqnarray[eqnarray omitted — 138 chars of source]

Using Proposition 1(4) from Embrechts_al2013, we have:

eqnarray[eqnarray omitted — 170 chars of source]

The latter equality holds, because (i) for all $u \in \overline{\operatorname{Ran}} F_{Y_{10}}$ there exists $y \in \overline{\mathbb R}$ such that $y=Q^{\mathbb R,-}_{Y_{10}}(u)$ and (ii) from Proposition 1(4) in Embrechts_al2013 we have $F_{Y_{10}}\left(Q^{\mathbb R,-}_{Y_{10}}(u)\right)=u$ for all $u \in \overline{\operatorname{Ran}} F_{Y_{10}}$. For $u, u' \in \overline{\operatorname{Ran}} F_{Y_{10}}$ such that $u<u'$ we have $Q^{\mathbb R,-}_{Y_{10}}(u) < Q^{\mathbb R,-}_{Y_{10}}(u') \Rightarrow F_{Y_1,D}\left(Q^{\mathbb R,-}_{Y_{10}}(u),0\right) <F_{Y_1,D}\left(Q^{\mathbb R,-}_{Y_{10}}(u'),0\right) \iff C_{Y_{10},D}(u,q) <C_{Y_{10},D}(u',q)$. The first strict inequality holds because by construction $Q^{\mathbb R,-}_{Y_{10}}(u)$ is strictly increasing on $\overline{\operatorname{Ran}} F_{Y_{10}}$. The second holds because $Q^{\mathbb R,-}_{Y_{10}}(\cdot) \in \mathbb Y_{10} \subseteq \mathbb Y_{10|0}$ since $\mathbb Y_{10|1} \subseteq \mathbb Y_{10|0}$.

\qed

Proof of Theorem (ref)

The proof follows in three steps. First, we derive the bounds (Section (ref)), then we proceed to show sharpness (Section (ref)). Since the sharpness proof relies on two intermediate lemmata, the last step is then to prove these two lemmata (Section (ref)).

Derivation of the bounds

Take a fixed $y \in \mathbb Y_{10|0}$, then the following holds for all $\tilde y< Q^{\mathbb R,+}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(y)\right)$ :

eqnarray[eqnarray omitted — 907 chars of source]

The first line of the inequality trivially holds from Lemma (ref)((ref)) and the fact that $Y_0\leq \tilde{y}$ implies $Y_0 < Q^{\mathbb R,+}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(y)\right)$. The third line holds by Sklar's Theorem Nelsen2006. The fourth line holds under Assumption (ref), and the last line holds under Assumption (ref). Notice that the last line requires $u \mapsto C_{Y_{10},D}(u,q)$ to be strictly increasing only on $\overline{\operatorname{Ran}} F_{Y_{10}}\cup \overline{\operatorname{Ran}} F_{Y_{00}} \subseteq [0,1]$. Now, applying the monotonicity of the function $v-C_{Y_0,D}(v,q)$ on the inequality ((ref)), for all $\tilde y < Q^{\mathbb R,+}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(y)\right)$ we have:

eqnarray*[eqnarray* omitted — 372 chars of source]

With a slight abuse of notation, we will use $F_{Y_{t0},D}(y,1)\equiv \mathbb{P}(Y_{t0}\leq y, D=1)$. Since $F_{Y_{t0}}(y)=F_{Y_{t0}, D}(y,1)+F_{Y_{t0}, D}(y,0)=F_{Y_{t0}, D}(y,1) + C_{Y_{t0},D}(F_{Y_{t0}}(y),q)$ for $t=0,1$, the latter equality implies the following:

eqnarray*[eqnarray* omitted — 833 chars of source]

where the second line holds under Assumption (ref). So, to summarize, for any fixed $y \in \mathbb Y_{10|0}$, we have: $$F_{Y_0|D=1}\left(\tilde y\right) \leq F_{Y_{10}|D=1}(y) \leq F_{Y_0|D=1}\left(Q^{\mathbb R,-}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(y)\right)\right), \text{ for all } \tilde y < Q^{\mathbb R,+}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(y)\right).$$ Taking the supremum over $\tilde{y}<Q_{Y_0|D=0}^{\mathbb{R},+}(F_{Y_1|D=0}(y))$ implies that: $$ \sup_{\tilde y < Q^{\mathbb R,+}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(y)\right)}F_{Y_0|D=1}\left(\tilde y\right) \leq F_{Y_{10}|D=1}(y) \leq F_{Y_0|D=1}\left(Q^{\mathbb R,-}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(y)\right)\right),$$ which is equivalent to: $$\underbrace{F_{Y_0|D=1}\left(Q^{\mathbb R,+}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(y)\right)-\right)}_{=F_{Y_0|D=1}\left(\left[Q^{\mathbb R,+}_{Y_0|D=0}\circ F_{Y_{1|D=0}}\right](y)-\right)\equiv F^{LB}(y) } \leq F_{Y_{10}|D=1}(y) \leq \underbrace{F_{Y_0|D=1}\left(Q^{\mathbb R,-}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(y)\right)\right)}_{=\left[F_{Y_0|D=1}\circ Q^{\mathbb R,-}_{Y_0|D=0}\circ F_{Y_1|D=0}\right] (y)\equiv F^{UB}(y) }.$$

We then finally have:

eqnarray[eqnarray omitted — 115 chars of source]

While these above bounds are point-wise sharp for all $y \in \mathbb Y_{10|0},$ they may not be sharp for $y \in \mathbb R \setminus \mathbb Y_{10|0}$. And this is because the upper bound may not be right-continuous in some cases, similarly for the lower bound which may not be right-continuous whenever $\{\tilde y \in \mathbb Y_{0|D=1} \cup \{-\infty\}:F_{Y_0|D=1}(\tilde y)\leq u\}$ is open for some $u\in Ran F_{Y_0|D=1}$.

To clarify this point, let us consider the simple case where $Y_{t0}$, $t \in \{0,1\}$ are all discrete random variables with $\mathbb Y_{10|0}=\{y_0,...,y_K\}$. In this case, $F^{LB}(.)$ is a well-defined cdf, while $F^{UB}(.)$ may not be a right-continuous function. Indeed, the function $u \mapsto Q^{\mathbb Y_{0|0},-}(u)$ is left-continuous and the discontinuities happen at $u\in Ran F_{Y_0|D=0}$. Now, consider that there exists $u_k \in Ran F_{Y_0|D=0} \cap Ran F_{Y_{10}|D=0}$, thus $F^{UB}(.)$ could be left-continuous at $y_k \in \mathbb Y_{10|0}$ such that $F_{Y_{10}|D=0}(y_k)=u_k$. If it is left-continuous and not right-continuous in $y_k$, we have: $\{y \in \overline{\mathbb R}: F^{UB}(y)> F^{UB}(y_k)\}=(y_k,\infty]$. Let us consider $\epsilon>0$ such that $y_k +\epsilon < y_{k+1}$. In such a case, $F_{Y_{10}|D=1}(y_k+\epsilon)=F_{Y_{10}|D=1}(y_k)$, however, by applying naively the bounds to $y_k$ and $y_{k}+\epsilon$ we have:

eqnarray[eqnarray omitted — 275 chars of source]

which implies that the upper bound in ((ref)) is not sharp since $F^{UB}(y_k+\epsilon)> F^{UB}(y_k)$. A valid tighter bound for $F^{LB}(y')$ for $y_{k}<y'<y_{k+1}$ is:

eqnarray*[eqnarray* omitted — 99 chars of source]

Since extending the bounds in Eq. ((ref)) to the case where $y \notin \mathbb Y_{10|0}$ provides non-sharp bounds, we provide an alternative approach that internalizes the idea that our target function of interest must be right-continuous since it is a cdf. Recall,

eqnarray[eqnarray omitted — 105 chars of source]

then for any fixed $y \in \mathbb R$, we have:

eqnarray*[eqnarray* omitted — 440 chars of source]

Notice that because $\mathbb Y_{10|1} \subseteq \mathbb Y_{10|0}$, and $F_{Y_{10}|D=1}(\cdot)$ is a right-continuous function, we have the following equality by Lemma (ref)((ref)): $$\lim_{\tilde y \downarrow y}\sup\left\{F_{Y_{10}|D=1}(t): t\leq \tilde y \; \& \; t \in \mathbb Y_{10|0} \cup\{-\infty \} \right\}= F_{Y_{10}|D=1}(y) \text{ for all } y \in \mathbb R.$$ The last inequality therefore becomes:

multline[multline omitted — 319 chars of source]

Sharpness of the bounds

In the previous subsection (ref), we showed that the bounds are valid. Now, we will show that both bounds are achievable. For the sake of brevity, we will focus only on the upper bound. The main idea is to provide a DGP which is only a function of the observable distributions but verifies the model assumptions and for which $\tilde{F}_{Y_{10}|D=1}(y)$ is equal to the upper bound.

Consider that the unidentified counterfactual distribution is exactly the upper bound: $$\tilde{F}_{Y_{10}\vert D=1}(y)\equiv \lim_{\tilde y \downarrow y}\sup\left\{F^{UB}(t): t\leq \tilde y \; \& \; t \in \mathbb Y_{10|0} \cup\{-\infty \} \right\}.$$ For simplicity, we consider the case where $$\lim_{\tilde y \downarrow y}\sup\left\{F^{UB}(t): t\leq \tilde y \; \& \; t \in \mathbb Y_{10|0} \cup\{-\infty \} \right\}=F^{UB}(y)\equiv F^{UB}_{Y_{10}\vert D=1}(y).$$ We need to define a joint distribution on $(Y_{00},Y_{10},Y_{11}, D)$ such that it is compatible with the data $(Y_0,Y_1,D)$, and Assumptions (ref) and (ref) hold. For any vector $X$, denote $F_{X,D}(x,d)=\mathbb P(X\leq x, D=d)$. Let $F_{Y_{00},Y_{10},Y_{11},D}(y_0,y_{10},y_{11},d)$ be a candidate joint distribution. We define

eqnarray*[eqnarray* omitted — 274 chars of source]

We construct the proposed distribution using the following rule. For $\tilde{F}_{Y_{00},Y_{10},Y_{11},D}(y_0,y_{10},y_{11},d)$ to be compatible with the data $(Y_0,Y_1,D)$, we must have

eqnarray*[eqnarray* omitted — 346 chars of source]

The distributions $\tilde{F}_{Y_{11}\vert Y_{00}\leq y_0,Y_{10}\leq y_{10},D=0}(y_{11})$ and $\tilde{F}_{Y_{10}\vert Y_{00}\leq y_0,Y_{11}\leq y_{11},D=1}(y_{10})$ are counterfactual. We set both of them equal to $\tilde{F}_{Y_{10}\vert Y_{00}\leq \infty,Y_{11}\leq \infty,D=1}(y_{10})=F^{UB}_{Y_{10}\vert D=1}(y)$, which is the counterfactual distribution that we consider above.

We now show that $\tilde{F}_{Y_{10}|D=1}(y)$ is a cdf. It is easy to see that $\tilde{F}_{Y_{10}|D=1}(y)$ is nondecreasing since for $y \leq y'$ we have

eqnarray*[eqnarray* omitted — 201 chars of source]

The limits of the function $\tilde{F}_{Y_{10}|D=1}(y)$ at $-\infty$ and $\infty$ are 0 and 1, respectively. By construction, the function $\tilde{F}_{Y_{10}|D=1}(y)$ is a right-continuous function.

We have

eqnarray*[eqnarray* omitted — 245 chars of source]

We now need to construct copulas $\tilde{C}_{Y_0,D}(u,q)$, $\tilde{C}_{Y_{10},D}(u,q)$, $\tilde{C}_{Y_{0},Y_{1} \vert D=0}(u_0,u_1)$, and $\tilde{C}_{Y_{0},Y_{10},D}(u_0,u_1,q)$ such that the following holds:

eqnarray*[eqnarray* omitted — 293 chars of source]

where $\tilde{F}_{Y_{10}}(y_{10}) = p F^{UB}_{Y_{10}|D=1}(y_{10}) + q F_{Y_1 \vert D=0}(y_{10})\equiv F^{UB}_{Y_{10}}(y_{10})$.

Since $\overline{Ran}F_{Y_0}$ is closed, we define {\tiny{

eqnarray*[eqnarray* omitted — 3,182 chars of source]

}}

where for any $u \in [0,1]$, $\underline{u}(u)\equiv \sup\{q\in \overline{\operatorname{Ran}} F_{Y_0} \cup \overline{\operatorname{Ran}} \tilde{F}_{Y_{10}}: q \leq u\}$, $\overline{u}(u)\equiv \inf\{q\in \overline{\operatorname{Ran}} F_{Y_0} \cup \overline{\operatorname{Ran}} \tilde{F}_{Y_{10}}: q \geq u\}$, and $ \tilde{Q}^{\mathbb{R},-}_{Y_{10}}(u)\equiv\inf\{y\in\mathbb{R}:\tilde{F}_{Y_{10}}(y)\geq u\}$.

{\tiny{

eqnarray*[eqnarray* omitted — 1,093 chars of source]

}} where for $t\in \{0,1\}$ and for any $(u_0,u_1) \in [0,1]^2$, $\underline{u_t}(u)\equiv \sup\{q\in \overline{\operatorname{Ran}} F_{Y_t\vert D=0}: q \leq u\}$, while $\overline{u_t}(u)\equiv \inf\{q\in \overline{\operatorname{Ran}} F_{Y_t \vert D=0}: q \geq u\}$.

We then define for $(u_0,u_1) \in [0,1]^2$

eqnarray*[eqnarray* omitted — 181 chars of source]

We can verify that $\tilde{C}_{Y_{00},Y_{10},D}(u_0,u_1,q)$ is a well-defined copula. We start by showing that $\tilde{C}_{Y_0,D}(u_0,q)$ is a well-defined subcopula. To do so, we need to introduce two intermediate lemmata:

lemmaFor any $u \in RanF_{Y_{10}}^{UB}$ and $v\in RanF_{Y_0}$ such that $u < v$, we have $\tilde{C}_{Y_0,D}(u,q) < \tilde{C}_{Y_0,D}(v,q)$.
lemmaSuppose $F_{Y_0}(y-) \in RanF_{Y_0}$ for all $y$. For any $u \in RanF_{Y_{10}}^{UB}$ and $v\in RanF_{Y_0}$ such that $v < u$, we have $\tilde{C}_{Y_0,D}(v,q) < \tilde{C}_{Y_0,D}(u,q)$.

First, we have $\tilde{C}_{Y_0,D}(1,q)=F_{Y_0,D}(Q_{Y_0}^{\mathbb{R},-}(1),0)=q$. Now let us show that for all $(u,v)\in[0,1]^2$ such that $u < v$, we have $\tilde{C}_{Y_0,D}(u,q)< \tilde{C}_{Y_0,D}(v,q)$. From the definition of $\tilde{C}_{Y_0,D}(u,q)$ and Lemma (ref), it follows that, when $u$ and $v$ belong to the same range, this monotonicity condition holds. We are going to prove it when $u$ and $v$ belong to different ranges. On the one hand, if $u \in Ran\tilde{F}_{Y_{10}}$ and $v \in RanF_{Y_0}$, then from Lemma (ref), we have $\tilde{C}_{Y_0,D}(u,q) < \tilde{C}_{Y_0,D}(v,q)$. On the other hand, if $v \in Ran\tilde{F}_{Y_{10}}$ and $u \in RanF_{Y_0}$, then from Lemma (ref), we have $\tilde{C}_{Y_0,D}(u,q) < \tilde{C}_{Y_0,D}(v,q)$. Since $\tilde{C}_{Y_0,Y_1 \vert D=0}(u_0,u_1)$ is an extended copula of the identified part of the copula of $(Y_0,Y_1) \vert D=0$ through the Sklar theorem, it is a well-defined copula. Any extended copula of this form should work for the proof, as we do not impose any additional restrictions on the true copula of $(Y_0,Y_1) \vert D=0$.

We also need to check that $\tilde{C}_{Y_{00},Y_{10},D}\left(F_{Y_0}(y_0),\tilde{F}_{Y_{10}}(y_{10}),q\right)=F_{Y_0,Y_1,D}(y_0,y_{10},0).$ This latter equality holds by construction of $\tilde{C}_{Y_{00},Y_{10},D}(u_0,u_1,q)$.

When we let $u_0$ go to 1, we obtain

eqnarray*[eqnarray* omitted — 283 chars of source]

Similarly,

eqnarray*[eqnarray* omitted — 283 chars of source]

And by construction, we have $\tilde{C}_{Y_{10},D}(u,q)=\tilde{C}_{Y_0,D}(u,q)$ for all $u\in [0,1]$ (Assumption (ref) holds). Furthermore, we have shown above that $\tilde{C}_{Y_0,D}(u,q)$ is strictly increasing in $u$ (Assumption (ref) holds).

By construction, the proposed joint distribution $\tilde{F}_{Y_{00},Y_{10},Y_{11},D}(y_0,y_1,y_2,d)$ is compatible with the data and the proposed copulas $\tilde{C}_{Y_0,D}(u,q)$, and $\tilde{C}_{Y_{10},D}(u,q)$ satisfy Assumptions (ref) and (ref).

The proof is similar for the lower bound on $F_{Y_{10\vert D=1}}(y)$ and any distribution in the identified set of $F_{Y_{10}\vert D=1}(y_{10})$.

To complete the proof, it remains to show the two intermediate lemmata.

Proofs of Intermediate Lemmata

Proof of Lemma (ref)

First, we start by the following claims:

claimFor any $u_y=F_{Y_{10}}^{UB}(y) \in RanF_{Y_{10}}^{UB}$, the smallest $v \in RanF_{Y_{0}}$ such that $u_y \leq v$ is $v_y=F_{Y_0}(h(y))$.

Proof. We have

eqnarray*[eqnarray* omitted — 134 chars of source]

Since $F_{Y_0\vert D=1}(h(y)) \in RanF_{Y_0 \vert D=1},$ to obtain the smallest element $v \in RanF_{Y_{0}}$, we need to find the smallest element $s$ on $RanF_{Y_0\vert D=0}$ such that $F_{Y_{1}\vert D=0}(y) \leq s$. From Lemma (ref).((ref)), $s=F_{Y_0\vert D=0}(Q^{\mathbb R,-}_{F_{Y_0\vert D=0}}(F_{Y_{1}\vert D=0}(y)))$. This completes the proof of Claim (ref).\qed

claimFor any $u \in RanF_{Y_{10}}^{UB}$, there exists $v \in RanF_{Y_{0}}$ such that $u \leq v \Longrightarrow \tilde{C}_{Y_0,D}(u,q)\leq \tilde{C}_{Y_0,D}(v,q).$

\noindentProof. $u_y\equiv F_{Y_{10}}^{UB}(y) = q F_{Y_{10}\vert D=0}(y) + p F^{UB}_{Y_{10}\vert D=1}(y),$ where $F^{UB}_{Y_{10}\vert D=1}(y)=F_{Y_0\vert D=1}(h(y))$ with $h(y)=Q_{Y_{0\vert 0}}^{\mathbb R,-}(F_{Y_1\vert D=0}(y))$. Then, $u_y \leq q F_{Y_{0}\vert D=0}(h(y)) + p F_{Y_0\vert D=1}(h(y))$, since $F_{Y_{10}\vert D=0}(y) \leq F_{Y_{0}\vert D=0}(h(y))$ by construction. So, $u_y \leq F_{Y_0}(h(y))\equiv v_y \in RanF_{Y_0}$. Now, the following hold:

eqnarray*[eqnarray* omitted — 300 chars of source]
eqnarray*[eqnarray* omitted — 189 chars of source]

This completes the proof of Claim (ref).\qed

Now we proceed to complete the proof of the lemma. Take $u \in RanF_{Y_{10}}^{UB}$ and $v\in RanF_{Y_0}$ such that $u < v$. Since $u \in RanF_{Y_{10}}^{UB}$, there exits $y$ such that $u_y=F_{Y_{10}}^{UB}(y)$. Then, from Claim (ref), there exists $v_y= F_{Y_0}(h(y))$ such that $u_y \leq v_y$. From Claim (ref), we have $v_y \leq v$. If $v=v_y$, then we have $F_{Y_{10}}^{UB}(y) < F_{Y_0}(h(y))$, which implies successively

eqnarray*[eqnarray* omitted — 366 chars of source]

If $v_y < v$, then from Claim (ref) we have $\tilde{C}_{Y_0,D}(u,q)\leq \tilde{C}_{Y_0,D}(v_y,q)$. And since $\tilde{C}_{Y_0,D}(v,q)$ is strictly increasing on $RanF_{Y_0}$ from Lemma (ref), we have $\tilde{C}_{Y_0,D}(v_y,q) < \tilde{C}_{Y_0,D}(v,q)$. Therefore, $\tilde{C}_{Y_0,D}(u,q) < \tilde{C}_{Y_0,D}(v,q)$. \qed

Proof of Lemma (ref)

We first start by stating and proving the following claim:

claimSuppose $F_{Y_0}(y-) \in RanF_{Y_0}$ for all $y$. For any $u_y=F_{Y_{10}}^{UB}(y) \in RanF_{Y_{10}}^{UB}$, there exist $w_y \in [0,1]$ and $v_y\in RanF_{Y_0}$ such that $v_y \leq w_y \leq u_y$ and $\tilde{C}_{Y_0,D}(v_y,q) \leq \tilde{C}_{Y_0,D}(w_y,q) \leq \tilde{C}_{Y_0,D}(u_y,q)$.

\noindentProof. We have

eqnarray*[eqnarray* omitted — 359 chars of source]

where $\underline{h}(y)= Q_{Y_0\vert 0}^{\mathbb R,+}(F_{Y_1\vert D=0}(y))$, and the second inequality holds from Lemma (ref). We discuss two cases.

Case 1: $F^{UB}_{Y_{10}\vert D=1}(y)=F^{LB}_{Y_{10}\vert D=1}(y)$

In this case, $u_y=w_y$, we have

eqnarray*[eqnarray* omitted — 554 chars of source]

Case 2: $F^{LB}_{Y_{10}\vert D=1}(y)< F^{UB}_{Y_{10}\vert D=1}(y)$

In this case, $w_y \notin \overline{\operatorname{Ran}} F_{Y_{10}}^{UB}$. From Lemma (ref), $v_y$ is the highest element of $\overline{\operatorname{Ran}} F_{Y_0}$ such that $w_y \geq v_y$. First, suppose $w_y \notin \overline{\operatorname{Ran}} F_{Y_0}.$ Then $w_y \in (\overline{\operatorname{Ran}} F_{Y_0})^c \cap (\overline{\operatorname{Ran}} F^{UB}_{Y_{10}})^c$. Let $\overline{u}(w_y)\equiv \inf\{q\in \overline{\operatorname{Ran}} F_{Y_0} \cup \overline{\operatorname{Ran}} F^{UB}_{Y_{10}}: q \geq w_y\}$. We have $v_y \leq w_y < \overline{u}(w_y) \leq u_y$, and either $\overline{u}(w_y) \in \overline{\operatorname{Ran}} F_{Y_0}$ or $\overline{u}(w_y) \in \overline{\operatorname{Ran}} F^{UB}_{Y_{10}}$.

If $\overline{u}(w_y) \in \overline{\operatorname{Ran}} F_{Y_0}$, then

eqnarray*[eqnarray* omitted — 278 chars of source]

Since $0 \leq \frac{w_y-v_y}{\overline{u}(w_y)-v_y} \leq 1$ and $\bigg[F_{Y_0,D}\left(Q_{Y_0}^{\mathbb{R},-}(\overline{u}(w_y)),0\right)-F_{Y_0,D}\left(Q_{Y_0}^{\mathbb{R},-}(v_y),0\right)\bigg] \geq 0$ from Lemma (ref), the following holds:

eqnarray*[eqnarray* omitted — 337 chars of source]

where the last inequality holds because $Q_{Y_{0}}^{\mathbb{R},-}(u)$ is monotone in $u$. Hence, $$\tilde{C}_{Y_0,D}(v_y,q) \leq \tilde{C}_{Y_0,D}(w_y,q) \leq \tilde{C}_{Y_0,D}(u_y,q).$$

If $\overline{u}(w_y) \in \overline{\operatorname{Ran}} F^{UB}_{Y_{10}}$, then $\overline{u}(w_y) \in \overline{\operatorname{Ran}} F^{UB}_{Y_{10}}=u_y$, and

eqnarray*[eqnarray* omitted — 277 chars of source]
eqnarray*[eqnarray* omitted — 167 chars of source]

Since $0 \leq \frac{w_y-v_y}{\overline{u}(w_y)-v_y} \leq 1$ and $\bigg[F_{Y_1,D}(y,0)-F_{Y_0,D}\left(\underline{h}(y)-,0\right)\bigg] \geq 0$ from Lemma (ref), the following holds:

eqnarray*[eqnarray* omitted — 135 chars of source]

Second, suppose $w_y \in \overline{\operatorname{Ran}} F_{Y_0}.$ Then, from Lemma (ref), we must have $w_y=v_y$, which implies $F_{Y_{1}, D}(y,0)=F_{Y_{0},D}(\underline{h}(y)-,0)$, which in turn implies $\tilde{C}_{Y_0,D}(u_y,q)=\tilde{C}_{Y_0,D}(v_y,q)=\tilde{C}_{Y_0,D}(w_y,q)$.

This completes the proof of Claim (ref). \qed

Now we proceed to complete the proof of the lemma. Take $u \in RanF_{Y_{10}}^{UB}$ and $v\in RanF_{Y_0}$ such that $v < u$. Since $u \in RanF_{Y_{10}}^{UB}$, there exits $y$ such that $u_y=F_{Y_{10}}^{UB}(y)$. From Claim (ref), there exists $w_y \in [0,1]$ and $v_y \in \overline{\operatorname{Ran}} F_{Y_0}$ such that $v\leq v_y < w_y \leq u$.

Case 1: $F^{UB}_{Y_{10}\vert D=1}(y)=F^{LB}_{Y_{10}\vert D=1}(y)$

In this case, $u_y=w_y$, we have

eqnarray*[eqnarray* omitted — 704 chars of source]

Case 2: $F^{LB}_{Y_{10}\vert D=1}(y)< F^{UB}_{Y_{10}\vert D=1}(y)$

The proof here is very similar to Case 2 in Claim (ref), except the strict inequality $0 < \frac{w_y-v_y}{\overline{u}(w_y)-v_y} < 1$. This strict inequality implies $$\tilde{C}_{Y_0,D}(v,q) \leq \tilde{C}_{Y_0,D}(v_y,q) < \tilde{C}_{Y_0,D}(w_y,q) \leq \tilde{C}_{Y_0,D}(u_y,q).$$ Hence, $\tilde{C}_{Y_0,D}(v,q) < \tilde{C}_{Y_0,D}(u,q)$.

Now we have completed the proof of the two intermediate lemmata and thereby the proof of Theorem (ref).

\qed

Proof of Corollary (ref)

proofIn the continuous cdfs case, we have \begin{eqnarray*} F^{LB}(t)&=&F_{Y_0|D=1}\left(Q^{\mathbb R,+}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(t)\right)-\right)\; \\\; &=&\mathbb P\left(Y_0 < Q^{\mathbb R,+}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(t)\right)\vert D=1\right) \; by definition \\\; &=&\mathbb P\left(Y_0 \leq Q^{\mathbb R,+}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(t)\right)\vert D=1\right) \; under continuity\\\; &=&F_{Y_0|D=1}\left(Q^{\mathbb R,+}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(t)\right)\right)\; \\\; F^{UB}(t)&=&F_{Y_0|D=1}\left(Q^{\mathbb R,-}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(t)\right)\right). \end{eqnarray*} We know that $Q^{\mathbb R,-}_{Y_0|D=0}(u) \leq Q^{\mathbb R,+}_{Y_0|D=0}(u)$ since $Q^{\mathbb R,+}_{Y_0|D=0}(u)=\sup\{y \in \mathbb R: F_{Y_0\vert D=0}(y)=u\}$ and $Q^{\mathbb R,-}_{Y_0|D=0}(u)=\inf\{y \in \mathbb R: F_{Y_0\vert D=0}(y)=u\}$ in the continuous cdf case. Since $F_{Y_0\vert D=1}$ is nondecreasing, $F_{Y_0|D=1}\left(Q^{\mathbb R,-}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(t)\right)\right) \leq F_{Y_0|D=1}\left(Q^{\mathbb R,+}_{Y_0|D=0}\left(F_{Y_{1|D=0}}(t)\right)\right)$, that is, $F^{UB}(t)\leq F^{LB}(t)$. We know that under our model assumptions $F^{LB}(t)\leq F^{UB}(t)$, therefore it follows that $F^{LB}(t)=F^{UB}(t)$.

Proof of Theorem (ref)

The proof of this theorem follows by similar arguments to the proof of Theorem (ref) and is therefore provided in Section (ref) of the online appendix.

Proof of Claim (ref)

(i) $\Longrightarrow$ (ii).

Since the cdf $F_{Y_{t0}}$ is continuous and strictly increasing, we have

eqnarray*[eqnarray* omitted — 284 chars of source]

By definition, $h_t$ is continuous and strictly increasing as is the quantile function $Q^{\mathbb R,-}_{Y_{t0}}$. Then, the following equalities hold:

eqnarray*[eqnarray* omitted — 63 chars of source]

where the second equality holds from the invariance principle in Embrechts_al2013 (Embrechts_al2013, Proposition 4(2)). Therefore,

eqnarray*[eqnarray* omitted — 497 chars of source]

where the second implication follows from $U_{t0} \sim \mathcal U_{[0,1]}$ and $F_D(0)=q$, the third holds from Sklar's theorem, and the fifth follows from $U_{t0} \sim \mathcal U_{[0,1]}$. Hence, we have:

eqnarray*[eqnarray* omitted — 493 chars of source]

(ii) $\Longrightarrow$ (i). Suppose there exist two strictly increasing functions $h_t(.), t \in \{0,1\}$ and two uniformly distributed random variables over $[0,1]$ $U_{00}$ and $U_{10}$ such that $Y_{t0}=h_t(U_{t0})$ and $U_{00}|D=d \sim U_{10}|D=d$. Then, we have

eqnarray*[eqnarray* omitted — 660 chars of source]

where the fourth implication holds from Sklar's theorem, the fifth follows from $U_{t0} \sim \mathcal U_{[0,1]}$, the sixth follows by the invariance principle in Embrechts_al2013 (Embrechts_al2013, Proposition 4.(2)), and the last holds from $Y_{t0}=h_t(U_{t0})$. \qed

\setcounter{figure}{0}

\setcounter{table}{0} \setcounter{page}{0} \setcounter{claim}{0} \pagenumbering{gobble}

center[center omitted — 239 chars of source]

\startcontents[sections] \printcontents[sections]{l}{1}{\setcounter{tocdepth}{1}} \setcounter{page}{0} \pagenumbering{arabic}