The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
56,117 characters
Sensitivity Analysis in Unconditional Quantile Effects
\doublespacing
\normalem
\title{Sensitivity Analysis in Unconditional Quantile Effects
\textcolor{red}{}}
\author{Juli\'an Mart\'inez-Iriarte\thanks{
\begin{doublespace}
Email: [email removed]. I am deeply indebted to my advisor Yixiao Sun for his constant support and guidance. I thank Jeffrey Clemens, Dimitris Christelis, Antonio Galvao, Peter Hull, Ying-Ying Lee, Michael Leung, Jessie Li, Xinwei Ma, Augusto Nieto-Barthaburu, Joel Sobel, Pietro Spini, Kaspar Wuthrich, Ying Zhu, and seminar participants at various universities for helpful comments. Two anonymous referees provided excellent comments that greatly improved the paper. All errors remain my own. \end{doublespace}
} \\
Department of Economics,
UC Santa Cruz}
\date{\textbf{\today}}
\maketitle
\thispagestyle{empty}
\vspace{-2em}
\begin{abstract}
This paper proposes a framework to analyze the effects of counterfactual policies on the unconditional quantiles of an outcome variable. For a given counterfactual policy, we obtain identified sets for the effect of both marginal and global changes in the proportion of treated individuals. To conduct a sensitivity analysis, we introduce the quantile breakdown frontier, a curve that $(i)$ indicates whether a sensitivity analysis is possible or not, and $(ii)$ when a sensitivity analysis is possible, quantifies the amount of selection bias consistent with a given conclusion of interest across different quantiles. To illustrate our method, we perform a sensitivity analysis on the effect of unionizing low income workers on the quantiles of the distribution of (log) wages.
\end{abstract}
\textbf{Keywords}: unconditional quantile effects, partial identification, sensitivity analysis.\bigskip
\clearpage
\pagenumbering{arabic}
\section{Introduction}
In this paper we propose a sensitivity analysis on the effect of counterfactual policies that change the proportion of treated individuals. Consider a situation where a policy maker is interested in treating non-treated individuals.
The key identification challenge is that we do not observe the counterfactual outcome
of individuals who switch groups, that is, the \emph{newly} treated individuals. In some cases, however, it is still possible to recover the distribution of the unobserved counterfactual outcome. For example, suppose that treatment status is randomly assigned, and a policy maker increases the proportion of treated individuals by randomly selecting non-treated individuals.\footnote{We assume full compliance in both randomizations.} Although we do not observe the counterfactual outcome of the \emph{newly} treated individuals, we know it is drawn from the same distribution as the \emph{already} treated individuals. Hence, we identified the counterfactual distribution of \emph{newly} treated individuals.
When treatment status is \emph{not} randomly assigned in the first place, the identification strategy previously described breaks down. The reason is that due to the selection bias in the original treatment status, a random selection of individuals from the control group will be drawn from a different distribution. Thus, in the presence of selection bias, identification of the counterfactual distribution requires that the policy maker has enough information to devise a policy such that the (unobservable) distribution of the \emph{newly} treated ``matches'' the distribution of the \emph{already} treated individuals. This is usually infeasible. Even if the policy maker has this information, such as when treatment status is randomly assigned, they might not be interested in a policy that merely selects the \emph{newly} treated individuals at random.
The previous discussion highlights that identification of counterfactual distributions results in either very stringent information requirements, or in policies that might not be interesting. In both cases, the distribution of the \emph{newly} treated individuals is restricted. From the point of view of the policy maker, this can rule out many interesting policies. To see this, consider the following example. A policy maker might like to know if an increase in the unionization rate reduces inequality. If unionized workers are relatively high-skilled, and a policy expands unionization with low-skilled workers, then the distribution of wages conditional on being in the union, is likely to change.
In order to analyze a richer set of counterfactual policies, we drop the restrictions on the distribution of the \emph{newly} treated individuals and provide partial identification results for two effects. The first one is a global effect that compares the quantiles of the observed outcome, to those of the counterfactual outcome, where the proportion of treated individuals has been increased by $\delta$. The second one is a marginal effect where we let $\delta$ go to zero, and analyze its limiting effect on the unconditional quantiles of the outcome.
Another important contribution of this paper is to propose a framework for a sensitivity analysis on certain conclusions of interest. To do this, we quantify the departure from point identification by the vertical distance between the distributions of the \emph{newly} treated individuals and the \emph{already} treated individuals. We introduce a curve called the \emph{quantile breakdown frontier}, which first indicates whether a sensitivity analysis is possible, and second, it quantifies the maximum departure from point identification such that a given set of conclusions holds across different quantiles. Using this curve, we bound the global effects curve using this maximum departure derived from the quantile breakdown frontier. In this way, we obtain an identified region for the global effect curve consistent with the desired conclusions. Estimation of both the quantile breakdown frontier and the bounds on the global effect are based on empirical distribution functions and empirical quantiles, and are $\sqrt n$-consistent.
The departure from point identification is due to the selection bias induced by the counterfactual policy. We call this the \emph{policy selection bias}. The usual selection bias states that treated and non-treated individuals are different in a sense, and that is what explains the selection in the first place. Instead, the policy selection bias is the difference between the distributions of the \emph{newly} treated individuals and the \emph{already} treated individuals. Returning to the unionization example, the policy selection bias arises because the union wages of \emph{newly} unionized workers may not be drawn from distribution of the \emph{already} unionized workers. We do not know the distribution of union wages of newly unionized workers, hence we can only partially identify the global and marginal effects.
The policy selection bias can be non-negligible even if the original selection into treatment is randomly assigned. The reason is that, for the policy selection bias, what matters is who the \emph{newly} treated individuals are. Conversely, if there is selection bias initially, but the distribution of the \emph{newly} treated ``matches'' the distribution of the \emph{already} treated individuals, then there will be no policy selection bias. Thus, the policy selection bias depends on the particular counterfactual policy being analyzed, not whether there is selection bias in the original selection mechanism.
We apply these methods to the study of unions and inequality, which has long been of interest to labor economics. A recent comprehensive review of this extensive literature is provided by \cite{Farber2020}. Using the data in \cite{Firpo2009}, our empirical application considers the effect of expanding unionization on the quantiles of the distribution of (log) wages. Our approach allows us to tackle the question from a different perspective. Using the tools developed in this paper, we can quantify the amount of policy selection bias that is consistent with a policy that increases the unionization rate by unionizing low earnings workers. By looking at the global effect in the $30$\textsuperscript{th} quantile of the distribution of wages we investigate the amount of policy selection bias consistent with unions reducing overall inequality. To this end, we examine the following conclusion: whether the $30$\textsuperscript{th} quantile increases by more than 10\%. We find that this is consistent with moderate values of policy selection bias. The bounds on the global effect for other quantiles reveals that this policy can have a bigger effect on quantiles below the $30$\textsuperscript{th} quantile, without hurting those at the middle and top of the distribution of income.
\textbf{Related Literature} There is an extensive literature devoted to the analysis of counterfactual distributions. A good reference is \cite{Firpo2011}. In this paper, we focus on counterfactual distributions that arise as a result of a counterfactual policy that changes the proportion of treated individuals. The Policy Relevant Treatment Effect (PRTE) of \cite{heckman_prte, Heckman2005}, and the Marginal PRTE (MPRTE) of \cite{Carneiro2010,Carneiro2011} are examples of the aforementioned global and marginal effects. The difference is that they analyze the unconditional mean of the outcome. Identification relies on the a separable threshold model for the selection equation, and the availability of a continuous instrumental variable. In this setting, the proportion of treated individuals is changed by manipulating the instrumental variable. Our analysis does not make any assumptions on the selection equation. We do not require an instrumental variable either.
The marginal effect on the unconditional quantiles of an outcome was first studied by \cite{Firpo2009}. The identification arguments of \cite{Firpo2009} are based on a distributional invariance assumption: the distribution of the outcome for the original treatment group (under the original policy regime) is the same as that for the new treatment group (under the new policy regime), and this also holds for the control groups under the two policy regimes.\footnote{See the proof to Corollary 3 of the working paper version \cite{Firpo2007}.} For the case of an endogenous binary covariate, where distributional invariance might not hold, \cite{yixiao2020} achieve identification by generalizing the Marginal Treatment Effect framework. \cite{Kasy2016} also analyzes counterfactual policies which assign a binary treatment, but focuses on a welfare ranking.
\cite{Rothe2012} provides a general treatment for functionals of the unconditional distribution of the outcome. What we call a global effect, \cite{Rothe2012} refers to as a \emph{Fixed Partial Policy Effect}, and what we call a marginal effect, \cite{Rothe2012} refers to as a \emph{Marginal Partial Distributional Policy}. However, \cite{Rothe2012} imposes different identifying assumptions, namely a form of conditional exogeneity, which also yield a partial identified set. We do not impose such assumptions in order to broaden the types of policies we can analyze.
It is important to highlight that we do not estimate a quantile treatment effect. The quantile treatment effect is the difference between the $\tau$-quantile under treatment and the $\tau$-quantile under control, and depends on the distribution of the covariates. In a recent contribution, \cite{Lieli2020} investigate the changes in this effect when the distribution of the covariates is manipulated. Aside from treatment status, we do not manipulate the distribution of covariates.
Our sensitivity analysis is based on the breakdown analysis of \cite{Kline2013} and \cite{Masten2020}. \cite{Kline2013} perform a sensitivity analysis in a different context: departures from a missing (data) at random assumption. In a manner similar to us, this departure is measured as the Kolmogorov-Smirnov distance between the distribution of observed outcomes and the (unobserved) distribution of missing outcomes. Our quantile breakdown frontier builds on the breakdown frontier introduced by \cite{Masten2020}. However, instead of relaxing two parameters, we relax just one, and plot it against different quantiles. Another recent application of the breakdown analysis is \cite{noack} in the context of LATE.
\textbf{Notation} All the CDFs are denoted by $F$ with a subscript indicating the random variable. So, the CDF of $Y$ is $F_Y(y)$. Conditional CDFs are denoted similarly. For example, the CDF of $Y$ conditional on $D=1$ and $X=x$ is denoted by $F_{Y|D=1,X=x}(y)$. The $\tau$-quantile of $Y$ is denoted by $F^{-1}_Y(\tau)$. Weak convergence is denoted by $\rightsquigarrow$.
\section{Counterfactual Policies and Unconditional Effects}\label{section_uncond_effects}
We will work with the potential outcomes framework. For some unknown functions $h_0$ and $h_1$
\begin{align*}
Y(0) &= h_0(X,U_0),\\
Y(1) &= h_1(X,U_1),
\end{align*}
where $X$ are observed covariates and $U_0$ and $U_1$ consist of unobservables. We do not impose any restriction on the dimension of the unobservables. The observed outcome is thus
\begin{align*}
Y &= D\cdot h_1(X,U_1) + (1-D)\cdot h_0(X,U_0).\notag\\
:&=h(D,X,U),
\end{align*}
for a general nonseparable function $h$, where $D$ is a binary random variable taking values $0$ and $1$, and $U:=(U_0,U_1)'$. The variable $D$ can be interpreted as the treatment status, and $p:=\Pr(D=1)$ is the proportion of treated individuals.
A counterfactual policy is an alternative assignment of individuals to treatment. It is given by a binary random variable $D_\delta$, such that $\Pr(D_\delta=1)=p+\delta$ for a fixed $\delta \in(-p,1-p)$. It is called counterfactual because it may assign $D_\delta=1$ to an individual whose $D=0$. As $\delta$ varies over $(-p,1-p)$, we obtain a collection of (counterfactual) policies which is denoted by $\mathcal D$. When a particular counterfactual policy $D_\delta$ belongs to $\mathcal D$ we write $D_\delta\in\mathcal D$. The counterfactual outcome we would observe for a given $D_\delta\in\mathcal D$ is
\begin{align*}
Y_{D_\delta}&=h(D_\delta,X,U),
\end{align*}
where we implicitly assumes that the potential outcomes are not affected by the manipulation of $D$.
Strictly speaking, the counterfactual outcome $Y_{D_\delta}$ is not well defined until we define $\mathcal D$, the collection of counterfactual policies. We will restrict ourselves to policies that shift a portion of individuals in the control group to the treatment group. We refer to such individuals as \emph{newly treated}. This means that for every individual, $D_\delta-D\geq 0$. This is shown in Figure \ref{figure_introduction}.
\begin{figure}
\centering
\begin{tikzpicture}[
scale=1.5]
\draw [green, thin, fill=green, fill opacity=0.2] plot [smooth] coordinates {(0.5,2) (0,3) (-1.3,3) (-2,2) (-0.5,1)};
\draw [green, thin, fill=cyan!20, fill opacity=0.7] plot [smooth] coordinates {(-0.5,1) (0,0) (1.355,1) (1.5, 1.5) (0.5,2) };
\draw [green, thin, fill=green, fill opacity=0.2] plot [smooth] coordinates {(6.5,2) (6,3) (4.7,3) (4,2) (5.5,1)};
\draw [green, thin, fill=cyan!20, fill opacity=0.7] plot [smooth] coordinates {(5.5,1) (6,0) (7.355,1)};
\draw [green, fill=red, fill opacity=0.2] plot [smooth] coordinates {(7.355,1) (7.5, 1.5) (6.5,2) };
\fill[fill=red, fill opacity=0.2] (6.5,2)--(7.355,1)--(5.5,1);
\draw [->, thick] (0.6,2.3)-- (3.8,2.3);
\node [anchor=south] at (2.2,2.3) {\footnotesize\emph{counterfactual policy}};
\draw [solid, thick] (-0.5,1)-- (0.5,2);
\node [anchor=north] at (-1,2.3) {\footnotesize$D=1$};
\node [anchor=north] at (0.5,1.3) {\footnotesize$D=0$};
\draw [dotted, thick] (5.5,1)-- (6.5,2);
\draw [solid, thick] (5.5,1)-- (7.355,1);
\node [anchor=north] at (5,2.3) {\footnotesize$D=1$};
\node [anchor=north] at (5.05,2) {\footnotesize$D_\delta=1$};
\node [anchor=north] at (6.7,1.7) {\footnotesize$D=0$};
\node [anchor=north] at (6.75,1.4) {\footnotesize$D_\delta=1$};
\node [anchor=north] at (6.25,0.9) {\footnotesize$D=0$};
\node [anchor=north] at (6.3,0.63) {\footnotesize$D_\delta=0$};
\end{tikzpicture}
\caption{\emph{A counterfactual policy where $D_\delta-D\geq 0$.}}
\label{figure_introduction}
\end{figure}
\begin{assumption}[Counterfactual Policies]\label{assumption_policy}
The collection of policies $\mathcal D$ satisfies
\begin{enumerate}
\item $\Pr(D_\delta=1)=p+\delta$ for $\delta\in[0,1-p)$ and $D_\delta\in\mathcal D$;
\item Monotonicity: $D_\delta-D\geq 0$;
\end{enumerate}
\end{assumption}
The monotonicity assumption $D_\delta-D\geq 0$ is mainly for expositional simplicity. We can do without this assumption, but we need to make some minor changes to our approach. However, there is also a practical purpose. In a context where $D$ is union status, and $D=1$ denotes unionized individuals, Assumption \ref{assumption_policy} requires that we increase the unionization rate by unionizing previously nonunionized workers. It would probably be hard to simultaneously unionize and deunionize different workers.
Another way to look at the monotonicity assumption is by inspecting the joint distribution of $D$ and $D_\delta$ it induces:
\begin{equation*}
\centering
\begin{tabular}{l|*{2}{c}r}
& $D_{\delta }=0$ & $D_{\delta }=1$ \\
\hline
$D=0$ & $1-p-\delta$& $\delta$ \\
$D=1$ & $0$ & $p$ \\
\end{tabular}
\end{equation*}
In other words, Assumption \ref{assumption_policy} rules out the presence of \emph{newly untreated} individuals. Also, in the limit, when $\delta=0$, we return to the original distribution of individuals. We will evaluate the effect of a counterfactual policy with two parameters: the global and the marginal effects. Let $F_Y^{-1}(\tau)$ and $F_{Y_{D_\delta}}^{-1}(\tau)$ denote the $\tau$-quantiles of $Y$ and $Y_{D_\delta}$ respectively.
\begin{definition}[Global and Marginal Effects]
For a given collection of policies $\mathcal D$, the unconditional global effect at the $\tau$-quantile of $D_\delta\in\mathcal D$ is
\begin{align*}
G_{\tau, D_\delta}&:=F_{Y_{D_\delta}}^{-1}(\tau)-F_{Y}^{-1}(\tau),
\end{align*}
and the unconditional marginal effect at the $\tau$-quantile is
\begin{equation*}
M_{\tau,\mathcal D}:=\lim_{\delta\to 0}\frac{F_{Y_{D_\delta}}^{-1}(\tau)-F_{Y}^{-1}(\tau)}{\delta}
\end{equation*}
whenever this limit exists.
\end{definition}
The global effect $G_{\tau, D_\delta}$ is the comparison of quantiles of the counterfactual distribution vs. the observed distribution for a fixed policy $D_\delta$. Naturally, for a collection $\mathcal D$, we have a corresponding collection on global effects. The marginal effect $M_{\tau,\mathcal D}$ can be interpreted as an ordinary derivative: for small $\delta$, it provides an approximation to the direction of the change in a given $\tau$-quantile. The main text will focus on the global effect, while the marginal effect is treated in detail in Appendix \ref{app_maginal_effect}.
\begin{remark}[\cite{Firpo2009}]
The marginal effect $M_{\tau,\mathcal D}$ was originally studied by \cite{Firpo2009}. Instead of Assumption \ref{assumption_policy}, \cite{Firpo2009} assume a form of distributional invariance: $F_{Y_{D_\delta}|D_\delta=d}=F_{Y|D=d}$ and obtain point identification. See the proof to Corollary 3 of the working paper version \cite{Firpo2007}. When both $D$ and $D_\delta$ are independent of $U$ and $X$, then distributional invariance will be satisfied. In this particular case, a policy maker can randomize $D_{\delta}$ so that for a given $\delta$, a fraction $p+\delta$ of individuals is randomly assigned to treatment. However, if we allow for $D$ to be endogenous, and if, as is usually the case, the structural form of endogeneity is unknown, then it may be impossible for the policy maker to design a sequence $\mathcal D$, such that for every $D_\delta\in\mathcal D$, $F_{Y_{D_\delta}|D_\delta=d}$ ``matches'' $F_{Y|D=d}$. From the point of view of the policy maker, this is a significant restriction on the types of counterfactual policies they can consider.
\end{remark}
\begin{remark}[Policy Relevant Treatment Effects]
\cite{heckman_prte, Heckman2005} and \cite{Carneiro2010,Carneiro2011} investigate the effect on the unconditional mean of the outcome. Using our notation, the Policy Relevant Treatment Effect (PRTE) of \cite{heckman_prte, Heckman2005} is
\begin{equation*}
\text{\textrm{PRTE}}_{D_\delta }=\frac{E(Y_{D_\delta })-E(Y)}{\delta}
\end{equation*}
and taking the limit $\delta \rightarrow 0$ yields the Marginal PRTE (MPRTE) of \cite{Carneiro2010,Carneiro2011}:
\begin{equation*}
\mathrm{MPRTE_\mathcal D}=\lim_{\delta \rightarrow 0}\text{\textrm{PRTE}}_{\delta }.
\end{equation*}
\cite{yixiao2020} show how to generalize the MPRTE to cover the case of \cite{Firpo2009} as well.
\end{remark}
\begin{remark}[\cite{Rothe2012}]
\cite{Rothe2012} also studies the global and marginal effects but under a different identifying assumption, namely a form of conditional exogeneity. This assumption also yields an identified set. Let the outcome be $Y=h(D,X,U)$. For uniformly distributed random variables $\tilde U_1$ and $\tilde U_2$, the outcome can be represented as $Y=h(Q_D(\tilde U_1),Q_X(\tilde U_2),U)$ where $Q_D$ and $Q_X$ are the quantile functions. Then $Q_D$ is changed to another quantile function $Q_D^*$, generating a counterfactual distribution, which is identified when $\tilde U_1\perp U\| X$ and $D$ is continuous. When $D$ is discrete, $\tilde U_1$ is not uniquely determined, so that a range of possible counterfactual distributions is possible resulting in partial identification.
\end{remark}
The next task is to define \emph{who} are the \emph{newly treated} individuals, that is, how does $D_\delta$ determine who receives treatment among the individuals whose $D=0$? In this paper we will focus on two types of policies: a policy that simply chooses individuals whose $D=0$ at random and assigns them to $D_\delta=1$, and a policy that chooses individuals based on a user-specified criterion. We will refer to these two types of policies as \emph{randomized policy} and \emph{non-randomized policy} respectively.
\begin{example}[Randomized policy]\label{example_rand_policy}
A randomized policy satisfies: for any $\delta\in[0,1-p)$
\begin{align*}
D_\delta =
\left\{
\begin{aligned}
1 &\quad \text{if } D = 1 \\
0 \text{ or } 1 &\quad \text{if } D = 0
\end{aligned}
\right.
\end{align*}
and the \emph{newly treated} are selected at random. Using the conditional independence notation
we write $D_\delta\perp Y(1), Y(0)\|D=0$:
\begin{align}\label{eqn_random_cia}
Pr(D_\delta=1|D=0)=\Pr(D_\delta=1|D=0, Y(1), Y(0))=\frac{\delta}{1-p}.
\end{align}
\end{example}
\begin{example}[Non-randomized policy]\label{example_non_rand_policy}
An example of a non-randomized policy is the following: for any $\delta\in[0,1-p)$
\begin{align}\label{eqn_non_random}
D_\delta =
\left\{
\begin{aligned}
1 &\quad \text{if } D = 1 \\
1 &\quad \text{if } D = 0 \text{ and } Z \le F^{-1}_{Z|D=0}\left( \frac{\delta}{1 - p} \right) \\
0 &\quad \text{otherwise}
\end{aligned}
\right.
\end{align}
for some observable random variable $Z$. In this case, the individuals in the group $\left\{D=0\right\}$ whose $Z$ is less than the $\frac{\delta}{1-p}$-quantile of this group are shifted to $D_\delta=1$. This rule guarantees that, in expectation, a proportion $\delta$ of individuals is shifted.
\end{example}
For a collection of policies $\mathcal D$ that satisfies Assumptions \ref{assumption_policy}, the counterfactual distribution $F_{Y_{D_\delta}}(y)$ can be decomposed as
\begin{align}\label{count_dist_dec}
F_{Y_{D_\delta}}(y) &= pF_{Y|D=1}(y) + (1-p-\delta) F_{Y|D_\delta=0}(y) +\delta F_{Y(1)|D=0,D_\delta=1}(y)
\end{align}
for each $D_\delta$. Here, $F_{Y(1)|D=0,D_\delta=1}(y)$ corresponds to the distribution of the \emph{newly treated}. This distribution cannot be identified from the data because it requires observing $Y(1)$ for a subpopulation for which we only observe their $Y(0)$. Consequently, $F_{Y_{D_\delta}}(y)$ is not identified either. The goal is to bound the quantiles of $F_{Y_{D_\delta}}(y)$. To that end, we make the following regularity assumptions.
\begin{assumption}[Regularity Assumptions]\label{assumption_regularity}
\begin{enumerate}
\item\label{assumption_yd_regular} For $d=0,1$, $F_{Y|D=d,X=x}(y)$ is continuous and strictly increasing for all $y$ such that $0<F_{Y|D=d,X=x}(y)<1$, and for all $x\in\mathcal X$.
\item\label{assumption_yd_regular_2} For every $D_\delta\in\mathcal D$, $F_{Y|D_\delta=0}(y)$ is continuous and strictly increasing for all $y$ such that $0<F_{Y|D_\delta=0}(y)<1$.
\item For every $D_\delta\in\mathcal D$, $F_{Y(1)|D=0,D_\delta=1}(y)$ is continuous and strictly increasing for all $y$ such that $0<F_{Y(1)|D=0,D_\delta=1}(y)<1$.
\item\label{assumption_x_support} For every $D_\delta\in\mathcal D$, the support of $X$ conditional on $D=0$ and $D_\delta=1$ is included in the support of $X$ conditional on $D=1$.
\end{enumerate}
\end{assumption}
The next assumption is our main working assumption to obtain the bounds on the quantiles of $F_{Y_{D_\delta}}(y)$.
\begin{assumption}[KS-distance]\label{assumption_ks_distance}
For a given $D_\delta\in\mathcal D$, there exists a known $c\in[0,1]$ such that
\begin{align}\label{eq:ks_distance}
\sup_{y\in\mathbb R}\left |F_{Y(1)|D=0,D_\delta=1}(y)-\int_{\mathcal X}F_{Y|D=1,X=x}(y)dF_{X|D=0,D_\delta=1}(x)\right |\leq c
\end{align}
\end{assumption}
We refer to the left hand side of \eqref{eq:ks_distance} as the \emph{policy selection bias}. The idea is that in the absence of policy selection bias, we can take $c=0$ and $F_{Y(1)|D=0,D_\delta=1}(y)$ can be recovered by matching \emph{already} unionized individuals with \emph{newly} unionized individuals and integrating against the characteristics of the \emph{newly} unionized individuals, using $F_{X|D=0,D_\delta=1}(x)$. This is akin to a \textit{joint} unconfoundedness assumption: $D\perp Y(1)\| X$ \textit{and} $D_\delta\perp Y(1)\| X$.\footnote{If $D, D_\delta\perp Y(1)\| X$, then $F_{Y|D=1, X=x}(y) = F_{Y(1)|D=0,D_\delta = 1, X=x}(y)$,
so that integrating against $F_{X|D=0,D_\delta=1}(x)$ yields $F_{Y(1)|D=0,D_\delta=1}(y) $.} By allowing $c$ to be non-zero, we are relaxing this particular type of conditional independence assumption (\cite{Masten2018}). Here, Assumption \ref{assumption_regularity}.\ref{assumption_x_support} becomes relevant.
\cite{Kline2013} also use the Kolmogorov-Smirnov distance to bound the quantiles of the outcome to allow for the possibility that data might not be missing at random.
\begin{remark}
For now we take $c$ as known. In the next section it will be the quantity on which we will base the sensitivity analysis.
\end{remark}
\begin{theorem}[Bounds on the Global Effect]\label{thm_bounds_global}
If a policy $D_\delta$ satisfies assumptions \ref{assumption_policy}, \ref{assumption_regularity}, and \ref{assumption_ks_distance}, then for $\tau\in(\delta,1-\delta)$
\begin{align*}
\max\left\{\tilde F_A^{-1}(\tau-\delta) ,F_A^{-1}(\tau-\delta c)\right\} - F_{Y}^{-1}(\tau) \leq G_{\tau, D_\delta}\leq \min\left\{\tilde F_A^{-1}(\tau) ,F_A^{-1}(\tau+\delta c)\right\}-F_{Y}^{-1}(\tau)
\end{align*}
where $\tilde F_A(y):= p F_{Y|D=1}(y) + (1-p-\delta) F_{Y|D_\delta=0}(y),$ and
\begin{align*}
F_A(y)&:=\tilde F_A(y) +\delta \int_{\mathcal X}F_{Y|D=1,X=x}(y)dF_{X|D=0,D_\delta=1}(x).
\end{align*}
For a fixed $\tau\in(\delta,1-\delta)$, these bounds are sharp.
\end{theorem}
\begin{remark}
The bounds are sharp in the sense that for any possible value of the global effect within the bounds, we can construct a corresponding newly treated distribution $F_{Y(1)|D=0,D_\delta=1}(y)$ which delivers precisely that same global effect. This construction depends on $\tau$, implying that, rather than uniform, sharpness is pointwise in $\tau\in(\delta,1-\delta)$.
\end{remark}
\begin{remark}
When $c=0$, $G_{\tau, D_\delta}$ is point identified because the counterfactual distribution in \eqref{count_dist_dec} becomes $F_A(y)$, which is identified. We refer to $F_A(y)$ as the apparent distribution, hence the subscript $A$, because it is the distribution that ``appears'' to be the counterfactual distribution. Naturally, for $c=0$, we have that the lower and upper bound are identical, and equal to $F_A^{-1}(\tau)$.\footnote{\label{footnote}
For $c=0$, the upper bound is $F_A^{-1}(\tau)$ because $F_A^{-1}(\tau)\leq \tilde F_A^{-1}(\tau) $. For the lower bound we need to show that $\max\left\{\tilde F_A^{-1}(\tau-\delta),F_A^{-1}(\tau)\right\}=F_A^{-1}(\tau)$. Note that $\tilde F_A^{-1}(\tau-\delta)$ satisfies:
\begin{align*}
p F_{Y|D=1}(\tilde F_A^{-1}(\tau-\delta)) + (1-p-\delta) F_{Y|D_\delta=0}(\tilde F_A^{-1}(\tau-\delta)) = \tau-\delta
\end{align*}
On the other hand, $F_A^{-1}(\tau)$ when $c=0$ satisfies
\begin{align*}
pF_{Y|D=1}(F_A^{-1}(\tau)) + (1-p-\delta) F_{Y|D_\delta=0}(F_A^{-1}(\tau)) +\delta \int_{\mathcal X}F_{Y|D=1,X=x}(F_A^{-1}(\tau))dF_{X|D=0,D_\delta=1}(x)=\tau
\end{align*}
Comparing these two expression, we obtain
\begin{align*}
p F_{Y|D=1}(\tilde F_A^{-1}(\tau-\delta)) + (1-p-\delta) F_{Y|D_\delta=0}(\tilde F_A^{-1}(\tau-\delta)) \leq
pF_{Y|D=1}(F_A^{-1}(\tau)) + (1-p-\delta) F_{Y|D_\delta=0}(F_A^{-1}(\tau))
\end{align*}
which implies that we must have $\tilde F_A^{-1}(\tau-\delta)\leq F_A^{-1}(\tau)$.
} Thus, for $c=0$, $G_{\tau, D_\delta}= F_A^{-1}(\tau)-F_{Y}^{-1}(\tau)$.
\end{remark}
\begin{remark}
By construction, we have that $\tilde F_A(y)\leq F_A(y)$, which implies, for $\tau\in(\delta,1-\delta)$, that $ F_A^{-1}(\tau)\leq \tilde F_A^{-1}(\tau)$. Therefore, whether $F_A$ or $\tilde F_A$ bounds the global effect depends on $\tau$, $\delta$, and $c$. In other words, the $\min$ and $\max$ are not redundant.
\end{remark}
\begin{remark}\label{remark_tau}
The range of $\tau$ is restricted to $(\delta,1-\delta)$ in order to ensure that $c$ to plays a role in the bounds. For $\tau$ outside this range, the bounds might not depend on $c$, if $c$ is close enough to 1.
\end{remark}
\section{Quantile Breakdown Frontier}
The quantile breakdown frontier is a curve that allows to perform to a sensitivity analysis with respect to $c$, the policy selection bias. Suppose we are interested in a target conclusion $G_{\tau,D_\delta}\geq g_{\tau,L} $ for some $g_{\tau,L}$. If the conclusion does not hold when $c=0$, then there is no point in performing a sensitivity analysis. On the other hand, if the conclusion \textit{does} hold under $c=0$, we would like to know the maximum amount of $c$ such that the conclusion continues to hold. This is the sensitivity analysis.
The quantile breakdown frontier for the global effect tackles both of these issues. The frontier is the map
\begin{align}\label{qbf_def}
\tau\mapsto c_{\tau,L}= \frac{\tau -F_A\left ( F_{Y}^{-1}(\tau)+ g_{\tau,L} \right )}{\delta}
\end{align}
for $\tau\in(\delta,1-\delta)$.\footnote{We can also look at conclusions of the type $G_{\tau,D_\delta}\leq g_{\tau,U}$. In this case, the quantile breakdown frontier is the map
\begin{align*}
\tau\mapsto c_{\tau,U}= \frac{F_A\left ( F_{Y}^{-1}(\tau)+ g_{\tau,U} \right )-\tau}{\delta}.
\end{align*}
}
An explanation of the derivation is in section \ref{app_frontier_der} of the appendix. When $c_{\tau,L}<0$, the desired target conclusion does not hold. If $c_{\tau,L}>0,$ then for $c\leq c_{\tau,L}$ then target conclusion holds under point identification. If $c> c_{\tau,L}>0$, then the target conclusion might not hold. In this sense, when $c_{\tau,L}>0,$ the frontier provides an amount of policy selection bias which is compatible with the conclusion.
The next lemma contains some properties of the quantile breakdown frontier.
\begin{lemma}\label{qbf_global_lemma}
Under the assumptions of Theorem \ref{thm_bounds_global}, the quantile breakdown frontier for $G_{\tau,D_\delta}$ defined in \eqref{qbf_def} satisfies the following: (i) if $\tau\mapsto g_{\tau,L}$ is continuous, then $\tau\mapsto c_{\tau,L}$ is continuous; (ii) if $c_{\tau,L}<0$, then the conclusion does not hold for $c=0$; (iii) if $c_{\tau,L}\geq 0$, then for any $c\geq 0$ such that $c\leq c_{\tau,L}$, the conclusion holds, and for $c> c_{\tau,L}$ the conclusion might not hold; and (iv) if $c_{\tau,L}\geq 0$ and $\tilde F_A^{-1}(\tau-\delta)-F_Y^{-1}(\tau)\geq g_{\tau,L} $, then the conclusion holds for any $c\in[0,1]$.
\end{lemma}
Part $(i)$ of the lemma puts a restriction on the types of conclusion by requiring continuity of the family of target conclusions. This is not essential, but illustrates the fact that the smoothness of the frontier depends, among other things, on the conclusions. For notational simplicity, later on we assume $g_{\tau,L}\equiv g$. Part $(iv)$ addresses a potential ``conservadurism'' in the breakdown analysis. This stems from the fact there is a possibility that the lower bound for the global effect does not depend on $c$. In the notation of Theorem \ref{thm_bounds_global}, this means that we cannot rule out $\max\left\{\tilde F_A^{-1}(\tau-\delta) ,F_A^{-1}(\tau-\delta c_{\tau,L})\right\} = \tilde F_A^{-1}(\tau-\delta).$ This is related to part $(ii)$. Indeed, if the opposite is true, $\tilde F_A^{-1}(\tau-\delta)< F_A^{-1}(\tau-\delta c_{\tau,L})$, then we can modify the language of part $(ii)$
to say that for $c> c_{\tau,L}$ the conclusion \textit{will} not hold.\footnote{In the empirical application we test for $\tilde F_A^{-1}(\tau-\delta)-F_Y^{-1}(\tau) < g_{\tau,L}$.}
\subsection{Some comparative statics}
It is possible to examine the behavior of $c_{\tau,L}$ with respect to $g_{\tau,L}$. Differentiating \eqref{qbf_def}, we get, for a fixed $\tau$:
\begin{align}\label{der_qbf_g}
\frac{\partial c_{\tau,L}}{\partial g_{\tau,L}}= -\frac{f_A\left ( F_{Y}^{-1}(\tau)+ g_{\tau,L} \right )}{\delta}<0.
\end{align}
provided $\delta>0$. This means that if the conclusion is less stringent, \textit{i.e.}, $g_{\tau,L}$ is reduced, then the frontier moves upwards, indicating that more policy selection bias can be allowed before the conclusion breaks down. In this sense, $c_{\tau,L}$ provides a minimum amount of policy selection bias for all $g$ such that $g\leq g_{\tau,L}$.
Now consider the derivative of \eqref{qbf_def} with respect to $\tau$:
\begin{align*}
\frac{\partial c_{\tau,L}}{\partial \tau}= \frac{1}{\delta}-\frac{f_A\left ( F_{Y}^{-1}(\tau)+ g_{\tau,L} \right )}{\delta}\left(\frac{\partial F_{Y}^{-1}(\tau)}{\partial \tau} + \frac{\partial g_{\tau,L}}{\partial \tau} \right)
\end{align*}
It is not possible to sign the slope of the frontier \textit{a priori}. It can depend largely on $\tau$, that is, whether we are the center or the tails of the distribution of $Y$, and whether the outcome is bounded or not, among other reasons.
The derivative of \eqref{qbf_def} with respect to $\delta$ is more complicated due to the fact that $F_A$ depends on $\delta$. A general expression is the following:
\begin{align*}
\frac{\partial c_{\tau,L}}{\partial \delta}= -\frac{c_{\tau,L}}{\delta}-\frac{1}{\delta}\frac{\partial F_A\left ( F_{Y}^{-1}(\tau)+ g_{\tau,L} \right )}{\partial \delta}
\end{align*}
where, following the definition of $F_A ( y)$ given in \eqref{complete_apparent}, we get (heuristically)
\begin{align*}
\frac{\partial F_A ( y)}{\partial \delta} &= \int_{\mathcal X}F_{Y|D=1,X=x}(y)dF_{X|D=0,D_\delta=1}(x)-F_{Y|D_\delta=0}(y)\\& + \delta \int_{\mathcal X}F_{Y|D=1,X=x}(y)\frac{\partial f_{X|D=0,D_\delta=1}(x)}{\partial \delta} dx - \delta \frac{\partial F_{Y|D_\delta=0}(y)}{\partial \delta}.
\end{align*}
\subsection{Bounds derived from the QBF}
Suppose that a policy maker choose a particular quantile $\tau^*$ and a target conclusion $g_{\tau^*,L}$. Then, provided that $c_{\tau^*,L}\in [0,1]$, we can obtain bounds on the global effect other quantiles $\tau\neq \tau^*$. The interpretation is that, as long as the policy selection bias satisfies $c\leq c_{\tau^*,L}$, the global effect will be bounded across quantiles by the bounds of Theorem \ref{thm_bounds_global} evaluated at $c_{\tau^*,L}$. In a sense, this allows us to extend the sensitivity analysis to other quantiles. That is for $\tau\in(\delta,1-\delta)$, we have
\begin{align}\label{eq:bounds_derived}
\max\left\{\tilde F_A^{-1}(\tau-\delta) ,F_A^{-1}(\tau-\delta c_{\tau^*,L})\right\} - F_{Y}^{-1}(\tau) &\leq G_{\tau, D_\delta}\notag\\
&\leq \min\left\{\tilde F_A^{-1}(\tau) ,F_A^{-1}(\tau+\delta c_{\tau^*,L})\right\}-F_{Y}^{-1}(\tau).
\end{align}
\section{Estimation and Inference}
We work in the space $\ell^{\infty}(\delta,1-\delta)$ of bounded real-valued functions defined on $(\delta,1-\delta)$. As usual, we endow this space with the supremum norm: $\|x\|_\infty:=\sup_{t\in(\delta,1-\delta)}|x(t)|$.\footnote{The reason we restrict the space to be $\ell^{\infty}(\delta,1-\delta)$ and not $\ell^{\infty}(0,1)$ is due to the fact that for a given $\delta$, we cannot reach quantiles below $\delta$ or above $1-\delta$. See Remark \ref{remark_tau} above.} In order to simplify notation, and ensure the continuity of the quantile breakdown frontier, we are going to focus on the case where the threshold is constant across $\tau$.
\begin{assumption}[Constant Threshold]\label{uniform_threshold}
For some scalar $g$, the threshold $g_{\tau,L}$ satisfies $g_{\tau,L}=g$ for any $\tau\in(\delta,1-\delta)$.
\end{assumption}
This assumption can be relaxed at the expense of more complicated notation. However, we still require smoothness in the map $\tau\mapsto g_\tau.$ For the case of the quantile breakdown for the sign of the marginal effect, we will set $g=0$.
\begin{assumption}[DGP]\label{dgp}
We observe an i.i.d. sample $\{ Y_i, D_i, X_i \}_{i=1}^n$, where the support of $X$ is finite.
\end{assumption}
\subsection{Global Effect - Estimation}
To estimate $ c_{\tau,L}$ given in \eqref{qbf_def} we need to specify two components: $(i)$ $\delta$, the increase in the treated proportion, and $(ii)$
$D_{i,\delta}$, the counterfactual treatment assignment, as in examples \ref{example_rand_policy} and \ref{example_non_rand_policy}. Once this has been done,
we use sample analogs. The estimator of the quantile breakdown frontier for the global effect is
\begin{align}\label{qbf_estimation}
\hat c_{\tau,L}= \frac{\tau - \hat F_A\left ( \hat F_{Y}^{-1}(\tau)+ g \right )}{\delta},
\end{align}
where $ \hat F_{Y}^{-1}(\tau)$ is the empirical $\tau$-quantile of $Y$: $\hat F_Y^{-1}(\tau) :=\inf \left\{ y:\hat F_Y(y)\geq \tau \right\}$, and $\hat F_A(y)$ is the empirical counterpart of $F_A(y)$ given in Theorem \ref{thm_bounds_global}. Detailed expressions are provided in Appendix \ref{appendix_notation}.
This is similar to a quantile-quantile transformation (see Exercise 4 in Chapter 3.9 in \cite{vandervaart1996}). We base our proof of the asymptotic distribution of $\sqrt n(\hat c_{\tau,L}-c_{\tau,L})$ on the proof of Lemma A.1 in \cite{Beare2019}.\footnote{\cite{Beare2019} also offer some interesting historical context for the result.} We view the map $\tau\mapsto\hat c_{\tau,L}$ as a random element of $\ell^{\infty}(\delta,1-\delta)$. In that case, we denote it simply by $\hat c_{L}.$
The main assumption is the following.
\begin{assumption}[Functional CLT]\label{brownian_bridge}
The following multivariate functional central limit theorem holds
\begin{align*}
\sqrt n\begin{pmatrix}
\hat F_{Y} - F_Y\\
\hat F_{A}- F_{A}
\end{pmatrix} \rightsquigarrow \begin{pmatrix}
\mathbb G_Y\\
\mathbb G_{A}
\end{pmatrix},
\end{align*}
where $\mathbb G_Y$ and $\mathbb G_A$ are Brownian bridges in $\ell^{\infty}(\mathcal Y)$.
\end{assumption}
\begin{remark}
Appendix \ref{app_primitive} lists primitive assumptions in terms of the CDFs that comprise $F_A$ that lead to Assumption \ref{brownian_bridge}.
\end{remark}
The following assumption is needed to establish the Hadamard differentiable of different functions used in the construction of $\theta$.
\begin{assumption}[Conditions for Hadamard Differentiability]\label{ass_had_dif}
\item
\begin{enumerate}
\item \label{ass_f_y_quantiles}For some $\varepsilon>0$, $F_Y$ is continuously differentiable in $[F_Y^{-1}(\delta)-\varepsilon,F_Y^{-1}(1-\delta)+\varepsilon]\subset \mathcal Y$ with strictly positive derivative $f_Y$.
\item \label{ass_f_a_uniform}$y\mapsto F_A(y)$ is differentiable, with uniformly continuous and bounded derivatives.
\end{enumerate}
\end{assumption}
The first item in Assumption \ref{ass_had_dif} concerns the support $\mathcal Y$ and the smoothness of $F_Y$. It is used to guarantee the Hadamard differentiability of the quantile process $\tau\mapsto F_Y^{-1}(\tau)$ for $\tau\in(\delta,1-\delta)$. The second item ensures that $F_A(y)$ has a uniformly continuous and bounded derivative which we denote by $f_A(y)$. It is needed to establish the Hadamard differentiability of the composition map $(F_A,F_Y^{-1})\mapsto F_A\circ (F_Y^{-1}+g)$.\footnote{Section 3.9 in \cite{vandervaart1996} studies the Hadamard differentiability of composition maps.}
\begin{theorem}\label{dist_theta}
Under Assumptions \ref{uniform_threshold}, \ref{dgp}, \ref{brownian_bridge}, and \ref{ass_had_dif}
\begin{align*}
\sqrt n\begin{pmatrix}
\hat c_{L}- c_{L}
\end{pmatrix} \rightsquigarrow
\frac{1}{\delta}\mathbb G_A\circ (F_Y^{-1}+g)-\frac{1}{\delta}f_A\circ (F_Y^{-1}+g) \frac{\mathbb G_Y\circ F_Y^{-1}}{f_Y\circ F_Y^{-1}},
\end{align*}
a tight Gaussian element in $\ell^{\infty}(\delta,1-\delta)$.
\end{theorem}
\subsection{Global Effect - Inference}
We propose a two-step inference procedure to accommodate the results of Lemma \ref{qbf_global_lemma}. First, we test the following null hypothesis: $ H_{1,0}: \tilde F_A^{-1}(\tau-\delta)-F_Y^{-1}(\tau)< g,$
against the alternative $ H_{1,a}: \tilde F_A^{-1}(\tau-\delta)-F_Y^{-1}(\tau)\geq g.$
If the null $H_{1,0}$ is not rejected, we proceed to test: $H_{2,0}: c_{\tau,L} = \bar c$, against the alternative $ H_{2,a}: c_{\tau,L} \neq \bar c,$ for some user-specified $\bar c$.
Let $\alpha_1$ denote the \emph{size} of the first test, and let $\alpha_2$ denote the \emph{conditional size} of the second test. The family-wise error rate (FWER), defined as the probability of making at least one false rejection among all true null hypotheses, is $ \alpha_1 + (1-\alpha_1)\alpha_2$. For example, to keep the FWER at $5\%$, then setting $\alpha_1=0.025$ yields $\alpha_2\le(0.05-\alpha_1)/(1-\alpha_1)\approx 0.02564$.
The first null hypothesis requires estimation of $d_{\tau}:=\tilde F_A^{-1}(\tau-\delta)-F_Y^{-1}(\tau)$. Estimation follows by resorting to the sample analogs: $\hat d_{\tau}=\hat{\tilde F}_A^{-1}(\tau-\delta)-\hat F_Y^{-1}(\tau)$.
Details can be found in Appendix \ref{appendix_d_tau}.
\subsection{Bounds derived from the QBF}
To estimate the bounds derived the from the QBF for the global effect, we use the sample counterpart of \eqref{eq:bounds_derived}. Besides the target conclusion $g$ for the global effect, we need to supply, a quantile level $\tau^*$ of interest. For
$\tau\in(\delta,1-\delta)$, the estimator of the derived bounds is
\begin{align*}
\max\left\{\hat{\tilde F}_A^{-1}(\tau-\delta) ,\hat F_A^{-1}(\tau-\delta \hat {\tilde c}_{\tau^*,L})\right\} - \hat F_{Y}^{-1}(\tau) &\leq G_{\tau, D_\delta}\\
&\leq \min\left\{\hat{\tilde F}_A^{-1}(\tau) ,\hat F_A^{-1}(\tau+\delta \hat {\tilde c}_{\tau^*,L})\right\}- \hat F_{Y}^{-1}(\tau),
\end{align*}
where
\begin{align}
\hat {\tilde c}_{\tau^*,L} = \max \left \{ \min \left \{ \hat c_{\tau^*,L}, 1\right\},0 \right\}.
\end{align}
Here, $\hat c_{\tau^*,L}$ is the estimator given in \eqref{qbf_estimation} evaluated $\tau=\tau^*$. The asymptotic distribution of the bounds cannot be Gaussian due to being the composition of maps which are not Hadamard fully differentiable--only directional.
Hence, by Corollary 3.1 in \cite{Fang2019}, the standard bootstrap will fail. This means that if we attempt to construct confidence intervals in the usual way by resampling, we will not obtain correct asymptotic coverage. Instead, we use the numerical delta method of \cite{Hong2018}. A detailed analysis of the asymptotic distribution of the bounds is in Appendix \ref{distribution_of_bounds_appendix}.
\section{Empirical application: What do unions do?}\label{section_empirical}
There is an extensive literature that studies unions and inequality. A recent contribution by \cite{Farber2020} contains a review of the literature. In our empirical application, in particular, we look at how unions affect the distribution of wages for \emph{all} workers. Unions can have a variety of effects on the distribution of wages. As argued by \cite{Freeman1980}, unions can raise the wages of unionized workers relative to non-unionized workers, possibly through more bargaining power. So, if higher paid workers unionize, the dispersion of wages can increase, but if lower paid workers unionize, the dispersion of wages can decrease. Furthermore, within a given industry, the union can reduce the dispersion of wages by standardizing the wages. This will impact the distribution of wages more or less depending on the size of the industry and the wages it pays.
A key difficulty in identifying the causal effect of unions on wages is that selection into unions is non-random. Hence, any measurement of the union premium--the difference in wages between similar union and nonunion workers--will be biased for the causal effect. Indeed, this has been a long standing concern of labor economists.\footnote{Indeed, the opening words of \cite{Card1996} are: \begin{quotation}
\emph{Despite a large and sophisticated literature there is still substantial disagreement over the extent to which differences in the structure of wages between union and nonunion workers represent an effect of trade unions, rather than a consequence of the nonrandom selection of unionized workers.}
\end{quotation}} With respect to selection into unions, \cite{Card1996} argues that unionized workers with low observed skills, tend to have high unobserved skills. The reverse happens with high skilled unionized workers: they tend to have low unobservable skills. Due to this selection bias, it might be impossible for a policy maker to devise a policy where the \emph{newly} unionized workers are selected in a way such that they are drawn from the distribution of the \emph{already} unionized workers.
Using the techniques developed in this paper, we are going to consider the effect of both globally and marginally expanding union coverage. We will explicitly allow for non-random selection into unions.
This allows for the \emph{newly} unionized and \emph{already} unionized workers to be drawn from different distributions. We do not use any imputation method to impute the union premium of the newly unionized workers.
Following \cite{Freeman1980}, \cite{Card2001} and \cite{Card2004} we consider a two sector economy. Each worker has a well-defined pair of potential (log) wages: $Y_i(1)$ for the unionized sector and $Y_i(0)$ for the nonunionized sector. Under Assumption \ref{assumption_policy}, and for any policy $D_\delta$, we have the following classification of individuals:
\begin{equation*}
\centering
\begin{tabular}{l|*{2}{c}r}
& $D_{\delta }=0$ & $D_{\delta }=1$ \\
\hline
$D=0$ & \emph{nonunionized} & \emph{newly unionized} \\
$D=1$ & - & \emph{unionized} \\
\end{tabular}
\end{equation*}
The relevant unobserved distribution is then $F_{Y(1)|D=0, D_\delta=1}$: the union wages of the newly unionized workers. Following \eqref{eq:ks_distance}, we look at departures of $F_{Y(1)|D=0, D_\delta=1}$ from
\begin{align*}
\int_{\mathcal X}F_{Y|D=1,X=x}(y)dF_{X|D=0,D_\delta=1}(x),
\end{align*}
which is observed. This difference is what we refer to as the policy selection bias.
Using the data in \cite{Firpo2009} we estimate the quantile breakdown frontier for marginal and global effects of different type of policies on the distribution of real log hourly wages. We use the 1983-1985 Outgoing Rotation Group (ORG) Supplement of the Current Population Survey. Our sample consists of 266,956 observations on U.S. males, of which $73.8\%$ are non-unionized, and $26.2\%$ are unionized. See \cite{Lemieux2006} for more details about the data. The covariates are: years of education, age, marital status, dummy for race (nonwhite), years of experience, and union status indicator. Table \ref{tab:wage_cov_diffs} contains sample means by union status, and the difference. All of the differences are significantly different from $0$ at the $5\%$ level, possibly reflecting selection into unions.
\begin{table}[ht]
\centering
\begin{tabular}{lccc}
\hline
Variable & Non-unionized & Unionized & Difference \\
\hline
\hline
Log(wage) & 1.718 & 1.963 & 0.245 \\
& (0.601) & (0.407) & [0.241, 0.249] \\[1ex]
Age (years) & 34.953 & 39.783 & 4.831 \\
& (12.455) & (11.566) & [4.729, 4.933] \\[1ex]
Education (years) & 12.966 & 12.498 & -0.468 \\
& (2.912) & (2.650) & [-0.491, -0.444] \\[1ex]
Experience (years) & 15.987 & 21.285 & 5.298 \\
& (12.668) & (12.252) & [5.192, 5.405] \\[1ex]
Married (share) & 0.611 & 0.742 & 0.131 \\
& (0.488) & (0.438) & [0.127, 0.135] \\[1ex]
Nonwhite (share) & 0.102 & 0.134 & 0.032 \\
& (0.303) & (0.341) & [0.029, 0.035] \\[1ex]
Number of Observations & 197,115 & 69,841 & \\[2ex]
\hline
\hline
\end{tabular}
\caption{Sample means by union status. Standard deviations in parentheses, and $95\%$ CIs for the differences in brackets.}
\label{tab:wage_cov_diffs}
\end{table}
The unionization rate in the dataset is $0.26$. Figure \ref{hump_union_rates} shows the typical hump-shaped pattern of the unionization rates by quantiles of the distribution of wages. For lower quantiles, unionization rates are quite low. They peak in the past the middle of the distribution and then drop at the higher quantiles.
\begin{figure}
\centering
\includegraphics[scale=0.42]{plot_hump.png}
\caption{\emph{Unionization rates by quantiles of the distribution of wages.}}
\label{hump_union_rates}
\end{figure}
Consider a policy that increase in the unionization rate by $10\%$. It consists of unionizing workers whose wages are below the $.10/(1-p)$-quantile $\approx 0.14$-quantile of the wages of the nonunionized sector.\footnote{This guarantees that the unionization rate increases by roughly 10\%. Indeed, the mean of $D_\delta$ is now $0.36$.} In the notation of this paper, we have $D=1$ if a worker is unionized, $D_\delta=1$ if a worker is unionized under the policy, $Y$ is (log) wage, and $\delta=0.1$. That is, $D_\delta$ is given by
\begin{align*}
D_\delta =
\left\{
\begin{aligned}
1 &\quad \text{if } D = 1 \\
1 &\quad \text{if } D = 0 \text{ and } Y \le F^{-1}_{Y|D=0}(0.14) \\
0 &\quad \text{otherwise}
\end{aligned}
\right.
\end{align*}
Figure \ref{empirical_qbf_many_2} shows a family of quantile breakdown frontiers for $g=0, 0.025, 0.05, 0.075, 0.1$, for a grid of $\tau\in (0.15,0.85)$. The top one corresponds to $g=0$. As $g$ increases, the frontiers shift down. This is in agreement with \eqref{der_qbf_g}, which states that, for a fixed $\tau$, the derivative of $c_{\tau,L}$ with respect to $g$ is negative.
Figure \ref{empirical_qbf} takes a closer look at the quantile breakdown frontier for $g=0.05$. Confidence intervals are obtained via 1,000 bootstrap replications. Since the outcome variable is log wages, this means approximately a $5\%$ increase in wages. First, we can see that for lower quantiles, the conclusion will hold as long as $c$ is below the frontier. On the other hand, for upper quantiles, the conclusion does not hold under point identification. Thus there is no point in doing a sensitivity analysis for upper quantiles.
Suppose the policy maker is interested in $\tau^*=.3$. This is indicated by the vertical dashed red line. The $.3$-quantile of the unconditional distribution of log wages is $1.465$. An increase of $0.05$ would bring this the $.3$-quantile to $1.515$ which is close to the $.33$-quantile which is $1.521$. In this case, $\hat c_{\tau^*,L}=\hat c_{.3,L}=0.223$, and $[0.176,0.270]$ is the approximately $97.5\%$ confidence interval. Importantly, $\hat d_{\tau^*} = \hat d_{.3} = 0.0001$ which is less that $g=0.05$, and, as shown in Figure \ref{plot_d_tau}, it is (uniformly) statistically less that $g=0.05$. This means, by parts $(iii)$ and $(iv)$ of Lemma \ref{qbf_global_lemma}, that the conclusion \textit{only} holds for $c\leq 0.223$, and hence the quantile breakdown frontier is sharp at $\tau=.3$. Using equation \eqref{eq:bounds_derived}, we can find the bounds for the global effect for all quantiles which are consistent with this amount of policy selection bias. This is shown in Figure \ref{bounds_qbf}. Note that at $\tau=0.3$, the lower bound is $.05$ by construction. At upper quantiles, the lower bound is slightly negative.
\begin{figure}
\centering
\includegraphics[scale=0.18]{qbf_many.png}
\caption{\emph{Quantile Breakdown Frontiers for $G_{\tau,D_\delta}\geq g$ and $g=0, 0.025, 0.05, 0.075, 0.1.$ Top one is $g=0$.}}
\label{empirical_qbf_many_2}
\end{figure}
\begin{figure}
\centering
\includegraphics[scale=0.18]{d_tau.png}
\caption{\emph{Pre-testing: $97.5\%$ one-sided upper uniform confidence band for $d_\tau.$}}
\label{plot_d_tau}
\end{figure}
\begin{figure}
\centering
\includegraphics[scale=0.18]{qbf_global_effect_uniform.png}
\caption{\emph{Quantile Breakdown Frontier for $G_{\tau,D_\delta}\geq0.05$ with $\approx97.5\%$ uniform confidence bands.}}
\label{empirical_qbf}
\end{figure}
\begin{figure}
\centering
\includegraphics[scale=0.18]{plot_bounds_and_point_identified_red_line.png}
\caption{\emph{Bounds derived from the QBF with $95\%$ uniform confidence bands, and the point identified global effects.}}
\label{bounds_qbf}
\end{figure}
In order to interpret the magnitude of the QBF, we compare it to the observable difference
\begin{align}\label{eq:obs_psb}
c_{obs}:=\sup_{y\in \mathbb R}\left |F_{Y(0)|D=0,D_\delta=1}(y)-\int_{\mathcal X}F_{Y|D=1,X=x}(y)dF_{X|D=0,D_\delta=1}(x)\right |
\end{align}
This is different from the KS-distance in Assumption \ref{assumption_ks_distance}, which is stated in terms of $F_{Y(1)|D=0,D_\delta=1}(y)$, an unobserved distribution. In \eqref{eq:obs_psb}, we are comparing the difference between the newly-treated when they are not treated, versus the treated matched with the newly treated covariates. The empirical counterpart of \eqref{eq:obs_psb} is $\hat c_{obs}=0.243$, and $(0.231 ,0.256)$ is the $95\%$ confidence interval computed via the usual bootstrap. This is similar to the value of the QBF at $\tau=.3$ which is $\hat c_{.3,L}=0.223$. If the treatment does not have a very large effect on the distribution of the outcome, then one could argue that the conclusion of interest might hold. However, if the treatment might have a very strong effect on the distribution, then it might be less likely that the conclusion holds.
\section{Conclusion}
This paper provides a way to perform a sensitivity analysis on the unconditional quantiles effects of policies that manipulate the proportion of treated individuals. To do so, we introduce the quantile breakdown frontier as a tool to examine, across quantiles, whether a sensitivity analysis is possible, and how much policy selection bias is compatible with a given conclusion. Next, we use the information from the curve at a particular quantile to provide bounds on the effect of a policy across quantiles. Our empirical application looks at the effect of increasing the proportion of unionized workers by unionizing lower earners. We find that an increase of $10\%$ for the $30$\textsuperscript{th} quantile is consistent under moderate values of selection bias. Moreover, this implies that the effect on lower quantiles would be even bigger, and that lower quantiles would not be hurt.
\clearpage
\bibliographystyle{aea}
\bibliography{references}
\clearpage