EconBase
← Back to paper

Inference on Counterfactual Distributions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

123,362 characters · 19 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inference on Counterfactual Distributions

abstract{ Counterfactual distributions are important ingredients for policy analysis and decomposition analysis in empirical economics. In this article we develop modeling and inference tools for counterfactual distributions based on regression methods. The counterfactual scenarios that we consider consist of ceteris paribus changes in either the distribution of covariates related to the outcome of interest or the conditional distribution of the outcome given covariates. For either of these scenarios we derive joint functional central limit theorems and bootstrap validity results for regression-based estimators of the status quo and counterfactual outcome distributions. These results allow us to construct simultaneous confidence sets for function-valued effects of the counterfactual changes, including the effects on the entire distribution and quantile functions of the outcome as well as on related functionals. These confidence sets can be used to test functional hypotheses such as no-effect, positive effect, or stochastic dominance. Our theory applies to general counterfactual changes and covers the main regression methods including classical, quantile, duration, and distribution regressions. We illustrate the results with an empirical application to wage decompositions using data for the United States. } { As a part of developing the main results, we introduce distribution regression as a comprehensive and flexible tool for modeling and estimating the entire conditional distribution. We show that distribution regression encompasses the Cox duration regression and represents a useful alternative to quantile regression. We establish functional central limit theorems and bootstrap validity results for the empirical distribution regression process and various related functionals. \newline } { Key Words: Counterfactual distribution, decomposition analysis, policy analysis, quantile regression, distribution regression, duration/transformation regression, Hadamard differentiability of the counterfactual operator, exchangeable bootstrap, unconditional quantile and distribution effects }

\thispagestyle{empty}

\pagestyle{plain}

Introduction

\setcounter{page}{1}

Counterfactual distributions are important ingredients for policy analysis (Stock, 1989, Heckman and Vytlacil, 2007) and decomposition analysis (e.g., Juhn, Murphy, and Pierce, 1993, DiNardo, Fortin, and Lemieux, 1996, Fortin, Lemieux, and Firpo, 2011) in empirical economics. For example, we might be interested in predicting the effect of cleaning up a local hazardous waste site on the marginal distribution of housing prices (Stock, 1991). Or, we might be interested in decomposing differences in wage distributions between men and women into a discrimination effect, arising due to pay differences between men and women with the same characteristics, and a composition effect, arising due to differences in characteristics between men and women (Oaxaca, 1973, and Blinder, 1973). In either example, the key policy or decomposition effects are differences between observed and counterfactual distributions. Using econometric terminology, we can often think of a counterfactual distribution as the result of either a change in the distribution of a set of covariates $X$ that determine the outcome variable of interest $Y$, or as a change in the relationship of the covariates with the outcome, i.e. a change in the conditional distribution of $Y$ given $X$. Counterfactual analysis consists of evaluating the effects of such changes.

The main objective and contribution of this paper is to provide estimation and inference procedures for the entire marginal counterfactual distribution of $Y$ and its functionals based on regression methods. Starting from regression estimators of the conditional distribution of the outcome given covariates and nonparametric estimators of the covariate distribution, we obtain uniformly consistent and asymptotically Gaussian estimators for functionals of the status quo and counterfactual marginal distributions of the outcome. Examples of these functionals include distribution functions, quantile functions, quantile effects, distribution effects, Lorenz curves, and Gini coefficients. We then construct confidence sets that take into account the sampling variation coming from the estimation of the conditional and covariate distributions. These confidence sets are uniform in the sense that they cover the entire functional with pre-specified probability and can be used to test functional hypotheses such as no-effect, positive effect, or stochastic dominance.

Our analysis specifically targets and covers the regression methods for estimating conditional distributions most commonly used in empirical work, including classical, quantile, duration/transformation, and distribution regressions. We consider simple counterfactual scenarios consisting of marginal changes in the values of a given covariate, as well as more elaborate counterfactual scenarios consisting of general changes in the covariate distribution or in the conditional distribution of the outcome given covariates. For example, the changes in the covariate and conditional distributions can correspond to known transformations of these distributions in a population or to the distributions in different populations. This array of alternatives allows us to answer a wide variety of counterfactual questions such as the ones mentioned above.

This paper contains two sets of new theoretical results on counterfactual analysis. First, we establish the validity of the estimation and inference procedures under two high-level conditions. The first condition requires the first stage estimators of the conditional and covariate distributions to satisfy a functional central limit theorem. The second condition requires validity of the bootstrap for estimating the limit laws of the first stage estimators. Under the first condition, we derive functional central limit theorems for the estimators of the counterfactual functionals of interest, taking into account the sampling variation coming from the first stage. Under both conditions, we show that the bootstrap is valid for estimating the limit laws of the estimators of the counterfactual functionals. The key new theoretical ingredient to all these results is the Hadamard differentiability of the counterfactual operator -- that maps the conditional distributions and covariate distributions into the marginal counterfactual distributions -- with respect to its arguments, which we establish in Lemma (ref). Given this key ingredient, the other theoretical results above follow from the functional delta method. A convenient and important feature of these results is that they automatically imply estimation and inference validity of any existing or potential estimation method that obeys the two high-level conditions set forth above.

The second set of results deals with estimation and inference under primitive conditions in two leading regression methods. Specifically, we prove that the high-level conditions -- functional central limit theorem and validity of bootstrap -- hold for estimators of the conditional distribution based on quantile and distribution regression. In the process of proving these results we establish also some auxiliary results, which are of independent interest. In particular, we derive a functional central limit theorem and prove the validity of exchangeable bootstrap for the (entire) empirical coefficient process of distribution regression and related functionals. We also prove the validity of the exchangeable bootstrap for the (entire) empirical coefficient process of quantile regression and related functionals. (These are a consequence of a more general result on functional delta method for Z-processes that we establish in Appendix (ref)).\footnote{ Prior work by Hahn (1995, 1997) showed empirical and weighted bootstrap validity for estimating pointwise laws of quantile regression coefficients; see also Chamberlain and Imbens (2003) and Chen and Pouzo (2009). An important exception is the independent work by Kline and Santos (2013) that established validity of weighted bootstrap for the entire coefficient process of Chamberlain (1994) minimum distance estimator in models with discrete covariates.} Note that the exchangeable bootstrap covers the empirical, weighted, subsampling, and $m$ out of $n$ bootstraps as special cases, which gives much flexibility to the practitioner.

This paper contributes to the previous literature on counterfactual analysis based on regression methods. Stock (1989) introduced integrated kernel regression-based estimators to evaluate the mean effect of policy interventions. Gosling, Machin, and Meghir (2000) and Machado and Mata (2005) proposed quantile regression-based estimators to evaluate distributional effects, but provided no econometric theory for these estimators. We also work with quantile regression-based estimators for evaluating counterfactual effects, though our estimators differ in some details, and we establish the limit laws as well as inference theory for our estimators. We also consider distribution regression-based estimators for evaluating counterfactual effects, and derive the corresponding limit laws as well as inference theory for our estimators. Moreover, our main results are generic and apply to any estimator of the conditional and covariate distributions that satisfy the conditions mentioned above, including classical regression (Juhn, Murphy and Pierce, 1993), flexible duration regression (Donald, Green and Paarsch, 2000), and other potential approaches.

Let us comment on the results for distribution regression separately. Distribution regression, as defined here, consists of the application of a continuum of binary regressions to the data. We introduce distribution regression as a comprehensive tool for modeling and estimating the entire conditional distribution. This partly builds on, but significantly differs from Foresi and Peracchi (1995), that proposed to use several binary regressions as a partial description of the conditional distribution. \footnote{ Foresi and Peracchi (1995) considered a fixed number of binary regressions. In sharp contrast, we consider a continuum of binary regressions, and show that the continuum provides a coherent and flexible model for the entire conditional distribution. Moreover, we derive the limit theory for the continuum of binary regressions, establishing functional central limit theorems and bootstrap functional central limit theorems for the distribution regression coefficient process and other related functionals, including the estimators of the entire conditional distribution and quantile functions.} We show that distribution regression encompasses the Cox (1972) transformation/duration model as a special case, and represents a useful alternative to Koenker and Bassett's (1978) quantile regression.

An alternative approach to counterfactual analysis, which is not covered by our theoretical results, consists in reweighting the observations using the propensity score, in the spirit of Horvitz and Thompson (1952). For instance, DiNardo, Fortin, and Lemieux (1996) applied this idea to estimate counterfactual densities, Firpo (2007) applied to to quantile treatment effects, and Donald and Hsu (2013) applied it to the distribution and quantile functions of potential outcomes. Under correct specification, the regression and the weighting approaches are equally valid. In particular, if we use saturated specifications for the propensity score and conditional distribution, then both approaches lead to numerically identical results. An advantage of the regression approach is that the intermediate step---the estimation of the conditional model---is often of independent economic interest. For example, Buchinsky (1994) applies quantile regression to analyze the determinants of conditional wage distributions. This model nests the classical Mincer wage regression and is useful for decomposing changes in the wage distribution into factors associated with between-group and within-group inequality.

We illustrate our estimation and inference procedures with a decomposition analysis of the evolution of the U.S. wage distribution, motivated by the in influential article by DiNardo, Fortin, and Lemieux (1996). We complement their analysis by employing a wide range of regression methods (instead of reweighting methods), providing standard errors for the estimates of the main effects, and extending the analysis to the entire distribution using simultaneous confidence bands. The use of standard errors allows us to disentangle the economic significance of various effects from the statistical uncertainty, which was previously ignored in most decomposition analyses in economics. We also compare quantile and distribution regression as competing models for the conditional distribution of wages and discuss the different choices that must be made to implement our estimators. Our empirical results highlight the important role of the decline in the real minimum wage and the minor role of de-unionization in explaining the increase in wage inequality during the 80s.

We organize the rest of the paper as follows. Section 2 presents our setting, the counterfactual distributions and effects of interest, and gives conditions under which these effects have a causal interpretation. In Section 3 we describe regression models for the conditional distribution, introduce the distribution regression method and contrast it with the quantile regression method, define our proposed estimation and inference procedures, and outline the main estimation and inference results. Section 4 contains the main theoretical results under simple high-level conditions, which cover a broad array of estimation methods. In Section 5 we verify the previous high-level conditions for the main estimators of the conditional distribution function---quantile and distribution regression---under suitable primitive conditions. In Section 6 we present the empirical application, and in Section 7 we conclude with a summary of the main results and pointing out some possible directions of future research. In the Appendix, we include all the proofs and additional technical results. We give a consistency result for bootstrap confidence bands, a numerical example comparing quantile and distribution regression, and additional empirical results in the online supplemental material (Chernozhukov, Fernandez-Val, and Melly, 2013).

The Setting for Counterfactual Analysis

Counterfactual distributions

In order to motivate the analysis, let us first set up a simple running example. Suppose we would like to analyze the wage differences between men and women. Let $0$ denote the population of men and $1$ the population of women. $Y_{j}$ denotes wages and $X_{j}$ denotes job market-relevant characteristics affecting wages for populations $j=0$ and $j=1$. The conditional distribution functions $F_{Y_{0}|X_{0}}(y|x)$ and $ F_{Y_{1}|X_{1}}(y|x)$ describe the stochastic assignment of wages to workers with characteristics $x$, for men and women, respectively. Let $F_{Y{\langle 0|0\rangle }}$ and $F_{Y{\langle 1|1\rangle }}$ represent the observed distribution function of wages for men and women, and $F_{Y{\langle 0|1\rangle }}$ represent the counterfactual distribution function of wages that would have prevailed for women had they faced the men's wage schedule $ F_{Y_{0}|X_{0}}$:

equation*[equation* omitted — 105 chars of source]

The latter distribution is called counterfactual, since it does not arise as a distribution from any observable population. Rather, this distribution is constructed by integrating the conditional distribution of wages for men with respect to the distribution of characteristics for women. This quantity is well defined if $\mathcal{X}_{0},$ the support of men's characteristics, includes $\mathcal{X}_{1}$, the support of women's characteristics, namely $ \mathcal{X}_{1}\subseteq \mathcal{X}_{0}.$

The difference in the observed wage distributions between men and women can be decomposed in the spirit of Oaxaca (1973) and Blinder (1973) as follows:

equation*[equation* omitted — 191 chars of source]

where the first term in brackets is due to differences in the wage structure and the second term is a composition effect due to differences in characteristics. We can decompose similarly any functional of the observed wage distributions such as the quantile function or Lorenz curve into wage structure and composition effects. These counterfactual effects are well defined econometric parameters and are widely used in empirical analysis, e.g. the first term of the decomposition is a measure of gender discrimination. It is important to note that these effects do not necessarily have a causal interpretation without additional conditions. Section 2.3 provides sufficient conditions for such an interpretation to be valid. Thus, our theory covers both the descriptive decomposition analysis and the causal policy analysis, because the econometric objects -- the counterfactual distributions and their functionals -- are the same in either case.

In what follows we formalize these definitions and treat the more general case with several populations. We suppose that the populations are labeled by $k\in \mathcal{K}$, and that for each population $k$ there is a random $ d_{x}$-vector $X_{k}$ of covariates and a random outcome variable $Y_{k}$. The covariate vector is observable in all populations, but the outcome is only observable in populations $j\in \mathcal{J}\subseteq \mathcal{K}$. Given observability, we can identify the covariate distribution $F_{X_{k}}$ in each population $k\in \mathcal{K},$ and the conditional distribution $ F_{Y_{j}|X_{j}}$ in each population $j\in \mathcal{J}$, as well as the corresponding conditional quantile function $Q_{Y_{j}|X_{j}}$.\footnote{ The inference theory of Section 4 does not rely on observability of $X_k$ and $(X_j,Y_j)$, but it only requires that $F_{X_{k}}$ and $F_{Y_{j}|X_{j}}$ are identified and estimable at parametric rates. In principle, $F_{X_{k}}$ and $F_{Y_{j}|X_{j}}$ can correspond to distributions of latent random variables. For example, $F_{Y_{j}|X_{j}}$ might be the conditional distribution of an outcome that we observe censored due to top coding, or it might be a structural conditional function identified by IV methods in a model with endogeneity. We focus on the case of observable random variables, because it is convenient for the exposition and covers our leading examples in Section 5. We briefly discuss extensions to models with endogeneity in the conclusion.} Thus, we can associate each $F_{X_{k}}$ with label $k$ and each $F_{Y_{j}|X_{j}}$ with label $j$. We denote the support of $X_{k}$ by $ \mathcal{X}_{k}\subseteq \mathbb{R}^{d_{x}}$ and the region of interest for $ Y_{j}$ by $\mathcal{Y}_{j}\subseteq \mathbb{R}$.\footnote{ We shall typically exclude tail regions of $Y_{j}$ in estimation, as in Koenker (2005, p. 148).} We assume for simplicity that the number of populations, $|\mathcal{K}|$, is finite. Further, we define $\mathcal{Y}_{j} \mathcal{X}_{j}=\{(y,x):y\in \mathcal{Y}_{j},x\in \mathcal{X}_{j}\}$, $ \mathcal{Y}\mathcal{X}\mathcal{J}=\{(y,x,j):(y,x)\in \mathcal{Y}_{j}\mathcal{ X}_{j},j\in \mathcal{J}\},$ and generate other index sets by taking Cartesian products, e.g., $\mathcal{J}\mathcal{K}=\{(j,k):j\in \mathcal{J} ,k\in \mathcal{K}\}$.

Our main interest lies in the counterfactual distribution and quantile functions created by combining the conditional distribution in population $j$ with the covariate distribution in population $k$, namely:

align[align omitted — 260 chars of source]

where $F_{Y{\langle j|k\rangle }}^{\leftarrow }$ is the left-inverse function of $F_{Y{\langle j|k\rangle }}$ defined in Appendix A. In the definition ((ref)) we assume the support condition:

equation[equation omitted — 126 chars of source]

which ensures that the integral is well defined. This condition is analogous to the overlap condition in treatment effect models with unconfoundedness (Rosenbaum and Rubin, 1983). In the gender wage gap example, it means that every female worker can be matched with a male worker with the same characteristics. If this condition is not met initially, we need to explicitly trim the supports and define the parameters relative to the common support.\footnote{ Specifically, given initial supports $\mathcal{X}_{j}^{o}$ and $\mathcal{X} _{k}^{o}$ such that $\mathcal{X}_{k}^{o}\not\subseteq \mathcal{X}_{j}^{o}$, we can set $\mathcal{X}_{k}=\mathcal{X}_{j}=(\mathcal{X}_{k}^{o}\cap \mathcal{X}_{j}^{o})$. Then the covariate distributions are redefined over this support. See, e.g., Heckman, Ichimura, Smith, and Todd (1998), and Crump, Hotz, Imbens, and Mitnik (2009) for relevant discussions.}

The counterfactual distribution $F_{Y{\langle j|k\rangle}}$ is the distribution function of the counterfactual outcome $Y{\langle j|k\rangle}$ created by first sampling the covariate $X_{k}$ from the distribution $ F_{X_{k}}$ and then sampling $Y{\langle j|k\rangle}$ from the conditional distribution $F_{Y_{j}|X_{j}}(\cdot|X_{k})$. This mechanism has a strong representation in the form\footnote{ This representation for counterfactuals was suggested by Roger Koenker in the context of quantile regression, as noted in Machado and Mata (2005).}

equation[equation omitted — 178 chars of source]

This representation is useful for connecting counterfactual analysis with various forms of regression methods that provide models for conditional quantiles. In particular, conditional quantile models imply conditional distribution models through the relation:

equation[equation omitted — 134 chars of source]

In what follows, we define a counterfactual effect as the result of a shift from one counterfactual distribution $F_{Y{\langle l|m\rangle}}$ to another $F_{Y{\langle j|k\rangle}},$ for some $j,l \in \mathcal{J}$ and $k,m \in \mathcal{K}$. Thus, we are interested in estimating and performing inference on the distribution and quantile effects

equation*[equation* omitted — 195 chars of source]

as well as other functionals of the counterfactual distributions. For example, Lorenz curves, commonly used to measure inequality, are ratios of partial means to overall means

equation*[equation* omitted — 204 chars of source]

defined for non-negative outcomes only, i.e. $\mathcal{Y}_j \subseteq [0,\infty)$. In general, the counterfactual effects take the form

equation[equation omitted — 146 chars of source]

This includes, as special cases, the previous distribution and quantile effects; Lorenz effects, with $\Delta(y) = L(y,F_{Y{\langle j|k\rangle}}) - L(y, F_{Y{\langle l|m\rangle}})$; Gini coefficients, with $\Delta= 1 - 2 \int_{\mathcal{Y}_j} L(F_{Y{\langle j|k\rangle}},y) dy =: G_{Y{\langle j|k\rangle}} $; and Gini effects, with $\Delta= G_{Y{\langle j|k\rangle}} - G_{Y{\langle l|m\rangle}}$.

Types of counterfactuals effects

Focusing on quantile effects as the leading functional of interest, we can isolate the following special cases of counterfactual effects (CE):

equation*[equation* omitted — 447 chars of source]

In the gender wage gap example mentioned at the beginning of the section, the wage structure effect is a type 1 CE (with $j=1$, $k=1$, and $l=0$), while the composition effect is an example of a type 2 CE (with $j=0$, $k=1$ , and $m=0 $). In the wage decomposition application in Section 6 the populations correspond to time periods, the minimum wage is treated as a feature of the conditional distribution, and the covariates include union status and other worker characteristics. We consider type 1 CE by sequentially changing the minimum wage and the wage structure. We also consider type 2 CE by sequentially changing the components of the covariate distribution. The CE of simultaneously changing the conditional and covariate distributions are also covered by our theoretical results but are less common in applications.

While in the previous examples the populations correspond to different demographic groups or time periods, we can also create populations artificially by transforming status quo populations. This is especially useful when considering type 2 CE. Formally, we can think of $X_{k}$ as being created through a known transformation of $X_{0}$ in population $0$:

equation[equation omitted — 128 chars of source]

This case covers, for example, adding one unit to the first covariate, $ X_{1k}=X_{10}+1,$ holding the rest of the covariates constant. The resulting effect becomes the unconditional quantile regression, which measures the effect of a unit change in a given covariate component on the unconditional quantiles of $Y$.\footnote{ The resulting notion of unconditional quantile regression is related but strictly different from the notion introduced by Firpo, Fortin and Lemieux (2009). The latter notion measures a first order approximation to such an effect, whereas the notion described here measures the exact size of such an effect on the unconditional quantiles. When the change is small, the two notions coincide approximately, but generally they can differ substantially.} For example, this type of counterfactual is useful for estimating the effect of smoking on the marginal distribution of infant birth weights. Another example is a mean preserving redistribution of the first covariate implemented as $X_{1k}=(1-\alpha )E[X_{10}]+\alpha X_{10}.$ These and more general types of transformation defined in ((ref)) are useful for estimating the effect of a change in taxation on the marginal distribution of food expenditure, or the effect of cleaning up a local hazardous waste site on the marginal distribution of housing prices (Stock, 1991).

Even though the previous examples correspond to conceptually different thought experiments, our econometric analysis covers all of them.

When counterfactual effects have a causal interpretation

Under an assumption called conditional exogeneity, selection on observables or unconfoundedness (e.g., Rosenbaum and Rubin, 1983, Heckman and Robb, 1985, and Imbens, 2004), CE can be interpreted as causal effects. In order to explain this assumption and define causal effects, it is convenient to rely upon the potential outcome notation. Let $(Y_{j}^{\ast } : j \in \mathcal{J})$ denote a vector of potential outcome variables for various values of a policy, $j\in \mathcal{J}$, and let $X$ be a vector of control variables or, simply, covariates.\footnote{ We use the term policy in a very broad sense, which could include any program or treatment. The definition of potential outcomes relies implicitly on a notion of manipulability of the policy via some thought experiment. Here there are different views about whether such thought experiment should be implementable or could be a purely mental act, see e.g., Rubin (1978) and Holland (1986) for the former view, and Heckman (1998, 2008) and Bollen and Pearl (2012) for the latter view. Following the treatment effects literature, we exclude general equilibrium effects in the definition of potential outcomes.} Let $J$ denote the random variable describing the realized policy, and $Y:=Y_{J}^{\ast }$ the realized outcome variable. When the policy $J$ is not randomly assigned, it is well known that the distribution of the observed outcome $Y$ conditional on $J=j,$ i.e. the distribution of $Y \mid J =j$, may differ from the distribution of $ Y_{j}^{\ast }$. However, if $J$ is randomly assigned conditional on the control variables $X$---i.e. if the conditional exogeneity assumption holds---then the distributions of $Y\mid X,J=j$ and $Y_{j}^{\ast }\mid X$ agree. In this case the observable conditional distributions have a causal interpretation, and so do the counterfactual distributions generated from these conditionals by integrating out $X$.

To explain this point formally, let $F_{Y_{j}^{\ast }\mid J}(y\mid k)$ denote the distribution of the potential outcome $Y_{j}^{\ast }$ in the population with $J=k\in \mathcal{J}$. The causal effect of exogenously changing the policy from $l$ to $j$ on the distribution of the potential outcome in the population with realized policy $J=k$, is

equation*[equation* omitted — 83 chars of source]

In the notation of the previous sections, the policy $J$ corresponds to an indicator for the population labels $j \in \mathcal{J}$, and the observed outcome and covariates are generated as $Y_j = Y \mid J = j,$ and $X_k = X \mid J = k.$\footnote{ The notation $Y_{j} = Y\mid J=j$ designates that $Y_{j} = Y$ if $J=j$, and $ X_k = X \mid J = k$ designates that $X_k = X$ if $J=k$.} The lemma given below shows that under conditional exogeneity, for any $j,k\in \mathcal{J}$ the counterfactual distribution $F_{Y{\langle j|k\rangle }}\left( y\right) $ exactly corresponds to $F_{Y_{j}^{\ast }\mid J}(y\mid k)$, and hence the causal effect of exogenously changing the policy from $l$ to $j$ in the population with $J=k$ corresponds to the CE of changing the conditional distribution from $l$ to $j,$ i.e.,

equation*[equation* omitted — 170 chars of source]
lemma[Causal interpretation for counterfactual distributions] Suppose that \begin{equation} (Y^*_j:j\in \mathcal{J}) \perp \!\!\!\perp J\mid X, \ a.s., \end{equation} where $\perp \!\!\!\perp $ denotes independence. Under ((ref)) and ( (ref)), \begin{equation*} F_{Y{\langle j|k\rangle }}(\cdot ) = F_{Y^*_j\mid J}(\cdot \mid k),\ \ j,k\in \mathcal{J}. \end{equation*}

The CE of changing the covariate distribution, $F_{Y{\langle j|k\rangle } }(y)-F_{Y{\langle j|m\rangle }}(y)$, also has a causal interpretation as the policy effect of changing exogenously the covariate distribution from $ F_{X_{m}}$ to $F_{X_{k}}$ under the assumption that the policy does not affect the conditional distribution. Such a policy effect arises, for example, in Stock (1991)'s analysis of the impact of cleaning up a hazardous site on housing prices. Here, the distance to the nearest hazardous site is one of the characteristics, $X$, that affect the price of a house, $Y$, and the cleanup changes the distribution of $X$, say, from $F_{X_{m}}$ to $ F_{X_{k}}$. The assumption for causality is that the cleanup does not alter the hedonic pricing function $F_{Y_{m}|X_{m}}(y|x)$, which describes the stochastic assignment of prices $y$ to houses with characteristics $x$. We do not discuss explicitly the potential outcome notation and the formal causal interpretation for this case.

Modeling Choices and Inference Methods for Counterfactual Analysis

In this section we discuss modeling choices, introduce our proposed estimation and inference methods, and outline our results, without submersing into mathematical details. Counterfactual distributions in our framework have the form ((ref)), so we need to model and estimate the conditional distributions $F_{Y_{j}|X_{j}}$ and covariate distributions $F_{X_{k}}$. As leading approaches for modeling and estimating $F_{Y_{j}|X_{j}}$ we shall use semi-parametric quantile and distribution regression methods. As the leading approach to estimating $F_{X_k}$ we shall consider an unrestricted nonparametric method. Note that our proposal of using distribution regressions is new for counterfactual analysis, while our proposal of using quantile regressions builds on earlier work by Machado and Mata (2005), though differs in algorithmic details.

Regression models for conditional distributions

The counterfactual distributions of interest depend on either the underlying conditional distribution, $F_{Y_{j}|X_{j}} $, or the conditional quantile function, $Q_{Y_{j}|X_{j}},$ through the relation ( (ref)). Thus, we can proceed by modeling and estimating either of these conditional functions. There are several principal approaches to carry out these tasks, and our theory covers these approaches as leading special cases. In this section we drop the dependence on the population index $j$ to simplify the notation.

1. Conditional quantile models. Classical regression is one of the principal approaches to modeling and estimating conditional quantiles. The classical location-shift model takes the linear-in-parameters form: $ Y=P(X)^{\prime }\beta +V,\ \ V=Q_{V}(U),$ where $U\sim U(0,1)$ is independent of $X$, $P(X)$ is a vector of transformations of $X$ such as polynomials or B-splines, and $P(X)^{\prime }\beta $ is a location function such as the conditional mean. The additive disturbance $V$ has unknown distribution and quantile functions $F_{V}$ and $Q_{V}$. The conditional quantile function of $Y$ given $X$ is $Q_{Y|X}(u|x)=P(X)^{\prime }\beta +Q_{V}(u),$ and the corresponding conditional distribution is $ F_{Y|X}(y|x)=F_{V}(y-P(X)^{\prime }\beta ).$ This model, used in Juhn, Murphy and Pierce (1993), is parsimonious but restrictive, since no matter how flexible $P(X)$ is, the covariates impact the outcome only through the location. In applications this model as well as its location-scale generalizations are often rejected, so we cannot recommend its use without appropriate specification checks.

A major generalization and alternative to classical regression is quantile regression, which is a rather complete method for modeling and estimating conditional quantile functions (Koenker and Bassett, 1978, Koenker, 2005). \footnote{ Quantile regression is one of most important methods of regression analysis in economics. For applications, including to counterfactual analysis, see, e.g., Buchinsky (1994), Chamberlain (1994), Abadie (1997), Gosling, Machin, and Meghir (2000), Machado and Mata (2005), Angrist, Chernozhukov, and Fern\'andez-Val (2006), and Autor, Katz, and Kearney (2006b).} In this approach, we have the general non-separable representation: $Y = Q_{Y|X}(U | X) = P(X)^{\prime}\beta(U),$ where $U \sim U(0,1)$ is independent of $X$ (Koenker, 2005, p. 59). We can back out the conditional distribution from the conditional quantile function through the integral transform:

equation*[equation* omitted — 106 chars of source]

The main advantage of quantile regression is that it permits covariates to impact the outcome by changing not only the location or scale of the distribution but also its entire shape. Moreover, quantile regression is flexible in that by considering $P(X)$ that is rich enough, one could approximate the true conditional quantile function arbitrarily well, when $Y$ has a smooth conditional density (Koenker, 2005, p. 53).

2. Conditional distribution models. A common approach to model conditional distributions is through the Cox (1972) transformation (duration regression) model: $ F_{Y|X}(y | x) = 1 - \exp(-\exp(t(y) - P(x)^{\prime}\beta)),$ where $ t(\cdot) $ is an unknown monotonic transformation. This conditional distribution corresponds to the following location-shift representation: $ t(Y) = P(X)^{\prime}\beta+ V, $ where $V$ has an extreme value distribution and is independent of $X$. In this model, covariates impact an unknown monotone transformation of the outcome only through the location. The role of covariates is therefore limited in an important way. Note, however, that since $t(\cdot)$ is unknown this model is not a special case of quantile regression.

Instead of restricting attention to the transformation model for the conditional distribution, we advocate modelling $F_{Y|X}(y|x)$ separately at each threshold $y \in \mathcal{Y}$, building upon related, but different, contributions by Foresi and Peracchi (1995) and Han and Hausman (1990). Namely, we propose considering the distribution regression model

equation[equation omitted — 112 chars of source]

where $\Lambda $ is a known link function and $\beta (\cdot )$ is an unknown function-valued parameter. This specification includes the Cox (1972) model as a strict special case, but allows for a much more flexible effect of the covariates. Indeed, to see the inclusion, we set the link function to be the complementary log-log link, $\Lambda (v)=1-\exp (-\exp (v))$, take $P(x)$ to include a constant as the first component, and let $P(x)^{\prime }\beta (y)=t(y)-P(x)^{\prime }\beta $, so that only the first component of $\beta (y)$ varies with the threshold $y$. To see the greater flexibility of ((ref)), we note that ((ref)) allows all components of $\beta (y)$ to vary with $y$.

The fact that distribution regression with a complementary log-log link nests the Cox model leads us to consider this specification as an important reference point. Other useful link functions include the logit, probit, linear, log-log, and Gosset functions (see Koenker and Yoon, 2009, for the latter). We also note that the distribution regression model is flexible in the sense that, for any given link function $\Lambda $, we can approximate the conditional distribution function $F_{Y|X}(y|x)$ arbitrarily well by using a rich enough $P(X)$.\footnote{ Indeed, let $P(X)$ denote the first $p$ components of a basis in $L^{2}( \mathcal{X},P)$. Suppose that $\Lambda ^{-1}(F_{Y|X}(y|X))\in L^{2}(\mathcal{ X},P)$ and $\lambda(t) =\partial \Lambda(t) / \partial t$ is bounded above by $\bar{\lambda}$. Then, there exists $\beta (y)$ depending on $p$, such that $\delta _{p}=E\left[ \Lambda ^{-1}(F_{Y|X}(y|X))-P(X)^{\prime }\beta (y) \right] ^{ 2}\rightarrow 0$ as $p$ grows, so that $E\left[ F_{Y|X}(y|X)-\Lambda \left( P(X)^{\prime }\beta \left( y\right) \right) \right] ^{2}\leq \bar{\lambda}\delta _{p}\rightarrow 0$.} Thus, the choice of the link function is not important for sufficiently rich $P(X)$.

3. Comparison of distribution regression vs. quantile regression. It is important to compare and contrast the quantile regression and distribution regression models. Just like quantile regression generalizes location regression by allowing all the slope coefficients $\beta (u) $ to depend on the quantile index $u$, distribution regression generalizes transformation (duration) regression by allowing all the slope coefficients $ \beta (y)$ to depend on the threshold index $y$. Both models therefore generalize important classical models and are semiparametric because they have infinite-dimensional parameters $\beta (\cdot )$. When the specification of $P(X)$ is saturated, the quantile regression and distribution regression models coincide.\footnote{ For example, when $P(X)$ contains indicators of all points of support of $X$ , if the support of $X$ is finite.} When the specification of $P(X)$ is not saturated, distribution and quantile regression models may differ substantially and are not nested. Accordingly, the model choice cannot be made on the basis of generality.

Both models are flexible in the sense that by allowing for a sufficiently rich $P(X)$, we can approximate the conditional distribution arbitrarily well. However, linear-in-parameters quantile regression is only flexible if $ Y$ has a smooth conditional density, and may provide a poor approximation to the conditional distribution otherwise, e.g. when $Y$ is discrete or has mass points, as it happens in our empirical application. In sharp contrast, distribution regression does not require smoothness of the conditional density, since the approximation is done pointwise in the threshold $y$, and thus handles continuous, discrete, or mixed $Y$ without any special adjustment. Another practical consideration is determined by the functional of interest. For example, we show in Remark 3.1 that the algorithm to compute estimates of the counterfactual distribution involves simpler steps for distribution regression than for quantile regression, whereas this computational advantage does not carry over the counterfactual quantile function. Thus, in practice, we recommend researchers to choose one method over the other on the basis of empirical performance, specification testing, ability to handle complicated data situations, or the functional of interest. In section 6 we explain how these factors influence our decision in a wage decomposition application.

Estimation of counterfactual distributions and their functionals

The estimator of each counterfactual distribution is obtained by the plug-in-rule, namely integrating an estimator of the conditional distribution $\widehat{F}_{Y_{j}|X_{j}}$ with respect to an estimator of the covariate distribution $\widehat{F}_{X_{k}}$,

equation[equation omitted — 217 chars of source]

For counterfactual quantiles and other functionals, we also obtain the estimators via the plug-in rule:

equation[equation omitted — 251 chars of source]

where $\widehat{F}_{Y{\langle j|k\rangle }}^{r}$ denotes the rearrangement of $\widehat{F}_{Y{\langle j|k\rangle }}$ if $\widehat{F}_{Y{\langle j|k\rangle }}$ is not monotone (see Chernozhukov, Fernandez-Val, and Galichon, 2010).\footnote{ If a functional $\phi _{0}$ requires proper distribution functions as inputs, we assume that the rearrangement is applied before applying $\phi _{0}$. Hence formally, to keep notation simple, we interpret the final functional $\phi $ as the composition of the original functional $\phi _{0}$ with the rearrangement.}

Assume that we have samples $\{ (Y_{ki}, X_{ki}): i=1,...,n_{k}\}$ composed of i.i.d. copies of $(Y_{k},X_{k})$ for all populations $k \in \mathcal{K}$, where $Y_{ji}$ is observable only for $j \in\mathcal{J} \subseteq \mathcal{K} $. We estimate the covariate distribution $F_{X_{k}} $ using the empirical distribution function

equation[equation omitted — 148 chars of source]

To estimate the conditional distribution $F_{Y_{j}|X_{j}},$ we develop methods based on the regression models described in Section 3.1. The estimator based on distribution regression (DR) takes the form:

eqnarray[eqnarray omitted — 429 chars of source]

where $d_p = \dim P(X_j)$. The estimator based on quantile regression (QR) takes the form:

eqnarray[eqnarray omitted — 411 chars of source]

for some small constant $\varepsilon > 0$. The trimming by $\varepsilon$ avoids estimation of tail quantiles (Koenker, 2005, p. 148), and is valid under the conditions set forth in Theorem (ref).\footnote{ In our empirical example, we use $\varepsilon=.01$. Tail trimming seems unavoidable in practice, unless we impose stringent tail restrictions on the conditional density or use explicit extrapolation to the tails as in Chernozhukov and Du (2008).}

We provide additional examples of estimators of the conditional distribution function in the working paper version (Chernozhukov, Fernandez-Val and Melly, 2009). Also our conditions in Section 4 allow for various additional estimators of the covariate distribution.

To sum-up, our estimates are computed using the following algorithm:

algorithm[algorithm omitted — 531 chars of source]
remarkIn practice, the quantile regression coefficients can be estimated on a fine mesh $\varepsilon = u_{1} < ... < u_{S} = 1 - \varepsilon$, with meshwidth $ \delta$ such that $\delta\sqrt{n_{j}} \to 0 $. In this case the final counterfactual distribution estimator is computed as: $\widehat F_{Y{\langle j|k\rangle}}(y)= \varepsilon + n_{k}^{-1} \delta\sum_{i=1}^{n_{k}} \sum_{s=1}^{S} 1\{P(X_{ki})^{\prime}\widehat\beta(u_{s}) \leq y \}. $ For distribution regression, the counterfactual distribution estimator takes the computationally convenient form $\widehat F_{Y{\langle j|k\rangle}}(y) = n_{k}^{-1} \sum_{i=1}^{n_{k}} \Lambda( P(X_{ki})^{\prime}\widehat\beta_{j}(y))$ that does not involve inversion nor trimming \qed

Inference

The estimators of the counterfactual effects follow functional central limit theorems under conditions that we will make precise in the next section. For example, the estimators of the counterfactual distributions satisfy

equation*[equation* omitted — 177 chars of source]

where $n$ is a sample size index (say, $n$ denotes the sample size of population $0$) and $\bar Z_{jk}$ are zero-mean Gaussian processes. We characterize the limit processes for our leading examples in Section 5, so that we can perform inference using standard analytical methods. However, for ease of inference, we recommend and prove the validity of a general resampling procedure called the exchangeable bootstrap (e.g., Praestgaard and Wellner, 1993, and van der Vaart and Wellner, 1996). This procedure incorporates many popular forms of resampling as special cases, namely the empirical bootstrap, weighted bootstrap, $m$ out of $n$ bootstrap, and subsampling. It is quite useful for applications to have all of these schemes covered by our theory. For example, in small samples with categorical covariates, we might want to use the weighted bootstrap to gain accuracy and robustness to “small cells", whereas in large samples, where computational tractability can be an important consideration, we might prefer subsampling.

In the rest of this section we briefly describe the exchangeable bootstrap method and its implementation details, leaving a more technical discussion of the method to Sections 4 and 5. Let $(w_{k1},...,w_{kn_{k}}),$ $k\in \mathcal{K},$ be vectors of nonnegative random variables that are independent of data, and satisfy Condition EB in Section 5. For example, $ (w_{k1},...,w_{kn_{k}})$ are multinomial vectors with dimension $n_{k}$ and probabilities $(1/n_{k},\ldots ,1/n_{k})$ in the empirical bootstrap. The exchangeable bootstrap uses the components of $(w_{k1},...,w_{kn_{k}})$ as random sampling weights in the construction of the bootstrap version of the estimators. Thus, the bootstrap version of the estimator of the counterfactual distribution is

equation[equation omitted — 246 chars of source]

The component $\widehat{F}_{X_{k}}^{\ast }$ is a bootstrap version of covariate distribution estimator. For example, if using the estimator of $ F_{X_{k}}$ in ((ref)), set

equation[equation omitted — 193 chars of source]

for $n_{k}^{\ast }=\sum_{i=1}^{n_{k}}w_{ki}$. The component $\widehat{F} _{Y_{j}|X_{j}}^{\ast }$ is a bootstrap version of the conditional distribution estimator. For example, if using DR, set $\widehat{F} _{Y_{j}|X_{j}}^{\ast }(y|x)=\Lambda (P(x)^{\prime }\widehat{\beta } _{j}^{\ast }(y)),\ (y,x)\in \mathcal{Y}_{j}\mathcal{X}_{j}$, $j\in \mathcal{J }$, for

equation*[equation* omitted — 224 chars of source]

If using QR, set $\widehat{F}_{Y_{j}|X_{j}}^{\ast }(y|x)=\varepsilon +\int_{\varepsilon }^{1-\varepsilon }1\{P(x)^{\prime }\widehat{\beta } _{j}^{\ast }(u)\leq y\}du$, $(y,x)\in \mathcal{Y}_{j}\mathcal{X}_{j}$, $j\in \mathcal{J}$, for

equation*[equation* omitted — 176 chars of source]

Bootstrap versions of the estimators of the counterfactual quantiles and other functionals are obtained by monotonizing $\widehat F^{*}_{Y{\langle j|k\rangle}}$ using rearrangement if required and setting

equation[equation omitted — 278 chars of source]

The following algorithm describes how to obtain an exchangeable bootstrap draw of a counterfactual estimator.

algorithm[algorithm omitted — 788 chars of source]

The exchangeable bootstrap distributions are useful to perform asymptotically valid inference on the counterfactual effects of interest. We focus on uniform methods that cover standard pointwise methods for real-valued parameters as special cases, and that also allow us to consider richer functional parameters and hypotheses. For example, an asymptotic simultaneous $(1-\alpha )$-confidence band for the counterfactual distribution $F_{Y{\langle j|k\rangle }}(y)$ over the region $y\in \mathcal{Y }_{j}$ is defined by the end-point functions

equation[equation omitted — 189 chars of source]

such that

equation[equation omitted — 269 chars of source]

Here, $\widehat{\Sigma }_{jk}(y)$ is a uniformly consistent estimator of $ \Sigma_{jk} (y)$, the asymptotic variance function of $\sqrt{n}(\widehat{F} _{Y{\langle j|k\rangle }}(y)-F_{Y{\langle j|k\rangle }}(y))$. In order to achieve the coverage property ((ref)), we set the critical value $\widehat{t}_{1-\alpha}$ as a consistent estimator of the $(1-\alpha )$ -quantile of the Kolmogorov-Smirnov maximal $t$-statistic:

equation*[equation* omitted — 157 chars of source]

The following algorithm describes how to obtain uniform bands using exchangeable bootstrap:

algorithm[algorithm omitted — 1,211 chars of source]

We can obtain similar uniform bands for the counterfactual quantile functions and other functionals replacing $\widehat{F}_{Y{\langle j|k\rangle }}^{\ast }$ by $\widehat{Q}_{Y{\langle j|k\rangle }}^{\ast }$ or $\widehat{ \Delta }^{\ast }$ and adjusting the indexing sets accordingly. If the sample size is large, we can reduce the computational complexity of step (i) of the algorithm by resampling the first order approximation to the estimators of the conditional distribution, by using subsampling, or by simulating the limit process $\bar{Z}_{jk}$ using multiplier methods (Barrett and Donald, 2003).

Our confidence bands can be used to test functional hypotheses about counterfactual effects. For example, it is straightforward to test no-effect, positive effect or stochastic dominance hypotheses by verifying whether the entire null hypothesis falls within the confidence band of the relevant counterfactual functional, e.g., as in Barrett and Donald (2003) and Linton, Song, and Whang (2010).\footnote{ For further references and other approaches, see McFadden (1989), Klecan, McFadden, and McFadden (1991), Anderson (1996), Davidson and Duclos (2000), Abadie (2002), Chernozhukov and Fernandez-Val (2005), Linton, Massoumi, and Whang (2005), Chernozhukov and Hansen (2006), or Maier (2011), among others.}

remark[On Validity of Confidence Bands] Algorithm (ref) uses the rescaled bootstrap interquartile range $\widehat{\Sigma }_{jk}(y)$ as a robust estimator of $\Sigma_{jk} (y)$. Other choices of quantile spreads are also possible. If $\Sigma_{jk} (y)$ is bounded away from zero on the region $ y\in\mathcal{Y}_j,$ uniform consistency of $\widehat{\Sigma }_{jk}(y)$ over $ y\in \mathcal{Y}_j$ and consistency of the confidence bands follow from the consistency of bootstrap for estimating the law of the limit Gaussian process $\bar{Z}_{jk},$ shown in Sections 4 and 5, and Lemma 1 in Chernozhukov and Fernandez-Val (2005). Appendix A of the Supplemental Material provides a formal proof for the sake of completeness. The bootstrap standard deviation is a natural alternative estimator for $\Sigma_{jk}(y),$ but its consistency requires the uniform integrability of $\{\widehat{Z} _{jk}^{\ast }(y)^2 : y \in \mathcal{Y}_j\}$ , which in turn requires additional technical conditions that we do not impose (see Kato, 2011). \qed

Inference Theory for Counterfactual Analysis under General Conditions

This section contains the main theoretical results of the paper. We state the results under simple high-level conditions, which cover a broad array of estimation methods. We verify the high-level conditions for the principal approaches -- quantile and distribution regressions -- in the next section. Throughout this section, $n$ denotes a sample size index and all limits are taken as $n \to \infty$. We refer to Appendix A for additional notation.

Theory under general conditions

We begin by gathering the key modeling conditions introduced in Section 2.

Condition S. \ (a) The condition ((ref)) on the support inclusion holds, so that the counterfactual distributions ((ref)) are well defined. (b) The sample size $n_k$ for the $k$ -th population is nondecreasing in the index $n$ and $n/n_{k} \to s_{k} \in[0, \infty),$ for all $k \in\mathcal{K}$, as $n \to \infty$.

We impose high-level regularity conditions on the following empirical processes:

equation*[equation* omitted — 208 chars of source]

indexed by $( y,x,j,k, f) \in\mathcal{Y}\mathcal{X}\mathcal{J}\mathcal{K} \mathcal{F}$, where $\widehat{F}_{Y_{j}|X_{j}}$ is the estimator of the conditional distribution $F_{Y_{j}|X_{j}}$, $\widehat{F}_{X_{k}} $ is the estimator of the covariate distribution $F_{X_{k}}$, and $\mathcal{F}$ is a function class specified below.\footnote{ Throughout the paper we assume that $\widehat{F}_{Y_{j}|X_{j}}$ takes values in $[0,1]$. This can always be imposed in practice by truncating negative values to $0$ and values above $1$ to $1$.} We require that these empirical processes converge to well-behaved Gaussian processes. In what follows, we consider $\mathcal{Y}_j\mathcal{X}_j$ as a subset of $\overline{\mathbb{R}} ^{1+d_x}$ with topology induced by the standard metric $\rho$ on $\overline{ \mathbb{R}}^{1+d_x}$, where $\overline{\mathbb{R}} = \mathbb{R} \cup \{ + \infty, - \infty\}$. We also let $\lambda_k(f,\tilde{f}) = [\int (f - \tilde f)^2 dF_{X_k}]^{1/2}$ be a metric on $\mathcal{F} $.

Condition D. \ Let $\mathcal{F}$ be a class of measurable functions that includes $\{ F_{Y_{j}|X_{j}}(y| \cdot): y \in\mathcal{Y}_j, j \in\mathcal{J}\} $ as well as the indicators of all the rectangles in $ \overline{\mathbb{R}}^{d_x}$, such that $\mathcal{F}$ is totally bounded under $\lambda_k$. (a) In the metric space $\ell^{\infty}(\mathcal{YXJKF} )^{2}$,

equation*[equation* omitted — 106 chars of source]

as stochastic processes indexed by $(y,x,j,k,f) \in\mathcal{YXJKF}.$ The limit process is a zero-mean tight Gaussian process, where $Z_{j}$ a.s. has uniformly continuous paths with respect to $\rho$, and $G_{{k}}$ a.s. has uniformly continuous paths with respect to the metric $\lambda_k$ on $ \mathcal{F}$. } (b) The map $y \mapsto F_{Y_{j}|X_{j}}(y| \cdot)$ is uniformly continuous with respect to the metric $\lambda_k$ for all $(j,k) \in\mathcal{J}\mathcal{K}$, namely as $\delta \to 0$, $\sup_{ \rho(y,\bar y) \leq \delta} \lambda_k ( F_{Y_{j}|X_{j}}(y| \cdot), F_{Y_{j}|X_{j}}(\bar y| \cdot) ) \to 0$ uniformly in $(j,k) \in\mathcal{J}\mathcal{K}$.

Condition D requires that a uniform central limit theorem holds for the estimators of the conditional and covariate distributions. We verify Condition D for semi-parametric estimators of the conditional distribution function, such as quantile and distribution regression, under i.i.d. sampling assumption. For the case of duration/transformation regression, this condition follows from the results of Andersen and Gill (1982) and Burr and Doss (1993). For the case of classical (location) regression, this condition follows from the results reported in the working paper version (Chernozhukov, Fernandez-Val and Melly, 2009). We expect Condition D to hold in many other applied settings. The requirement $\widehat {G}_{{k}} \rightsquigarrow{G}_{{k}}$ on the estimated measures is weak and is satisfied when $\widehat F_{X_{k}}$ is the empirical measure based on a random sample, as in the previous section. Finally, we note that Condition D does not even impose the i.i.d sampling conditions, only that a functional central limit theorem is satisfied. Thus, Condition D can be expected to hold more generally, which may be relevant for time series applications.

remarkCondition D does not require that the regions $\mathcal{Y}_j$ and $\mathcal{X }_{k}$ are compact subsets of $\mathbb{R}$ and $\mathbb{R}^{d_x}$, but we shall impose compactness when we provide primitive conditions. The requirement $\widehat {G}_{{k}} \rightsquigarrow{G}_{{k}} $ holds not only for empirical measures but also for various smooth empirical measures; in fact, in the latter case the indexing class of functions $\mathcal{F}$ can be much larger than Glivenko-Cantelli or Donsker; see Radulovic and Wegkamp (2003) and Gine and Nickl (2008). \qed
theorem[Uniform limit theory for counterfactual distributions and quantiles] Suppose that Conditions S and D hold. (1) Then, \begin{equation} \sqrt{n} \left( \widehat{F}_{Y\langle j|k\rangle}(y) - F_{Y\langle j|k\rangle}(y) \right) \rightsquigarrow\bar Z_{jk}(y) \end{equation} as a stochastic process indexed by $(y,j,k) \in\mathcal{Y}\mathcal{J} \mathcal{K}$ in the metric space $\ell^{\infty}( \mathcal{Y}\mathcal{J} \mathcal{K})$, where $\bar Z_{jk}$ is a tight zero-mean Gaussian process with a.s. uniformly continuous paths on $\mathcal{Y}_j$, defined by \begin{equation} \bar Z_{jk}(y) := \sqrt{s_{j}} \int Z_{j}(y,x) d F_{X_{k}}(x) + \sqrt{s_{k}} G_{{k}}(F_{Y_{j}|X_{j}}(y| \cdot)). \end{equation} (2) If in addition $F_{Y{\langle j|k\rangle}}$ admits a positive continuous density $f_{Y{\langle j|k\rangle}}$ on an interval $[a,b]$containing an $ \epsilon $-enlargement of the set $\{Q_{Y{\langle j|k\rangle}}(\tau): \tau \in{\mathcal{T}}\}$ in $\mathcal{Y}_j$, where ${\mathcal{T}} \subset (0,1)$, then \begin{equation} \sqrt{n} \left( \widehat{Q}_{Y{\langle j|k\rangle}}(\tau) - Q_{Y{\langle j|k\rangle}}(\tau) \right) \rightsquigarrow- \bar Z_{jk}(Q_{Y{\langle j|k\rangle}}(\tau)) / f_{Y{\langle j|k\rangle}}( Q_{Y_{\langle j|k\rangle}}(\tau)) =: V_{jk}(\tau), \end{equation} as a stochastic process indexed by $(\tau,j,k) \in{\mathcal{T}}\mathcal{J} \mathcal{K}$ in the metric space $\ell^{\infty}( {\mathcal{T}}\mathcal{J} \mathcal{K})$, where $V_{jk}$ is a tight zero mean Gaussian process with a.s. uniformly continuous paths on ${\mathcal{T}}$.

This is the first main and new result of the paper. It shows that if the estimators of the conditional and covariate distributions satisfy a functional central limit theorem, then the estimators of the counterfactual distributions and quantiles also obey a functional central limit theorem. This result forms the basis of all inference results on counterfactual estimators.

As an application of the result above, we derive functional central limit theorems for distribution and quantile effects. Let $\mathcal{Y} \subseteq \mathcal{Y}_j \cap \mathcal{Y}_l,$ ${\mathcal{T}} \subset (0,1),$ and

align*[align* omitted — 411 chars of source]
corollary[Limit theory for quantile and distribution effects] Under the conditions of Theorem (ref), part 1, \begin{equation} \sqrt{n} \left( \widehat\Delta^{DE} (y) - \Delta^{DE} (y) \right) \rightsquigarrow \bar Z_{jk}(y) - \bar Z_{lm}(y) =: S_{t}(y), \end{equation} as a stochastic process indexed by $y \in\mathcal{Y}$ in the space $ \ell^{\infty}( \mathcal{Y} )$, where $S_{t}$ is a tight zero-mean Gaussian process with a.s. uniformly continuous paths on $\mathcal{Y}$. Under conditions of Theorem (ref), part 2, \begin{equation} \sqrt{n} \left( \widehat\Delta^{QE} (\tau) - \Delta^{QE} (\tau) \right) \rightsquigarrow V_{jk}(\tau) - V_{lm}(\tau) =: W_{t}(\tau), \end{equation} as a stochastic process indexed by $\tau \in{\mathcal{T}}$ in the space $ \ell^{\infty}( {\mathcal{T}} )$, where $W_{t}$ is a tight zero-mean Gaussian process with a.s. uniformly continuous paths on $\mathcal{T}$.

The following corollary is another application of the result above. It shows that plug-in estimators of Hadamard-differentiable functionals also satisfy functional central limit theorems. Examples include Lorenz curves and Lorenz effects, as well as real-valued parameters, such as Gini coefficients and Gini effects. Regularity conditions for Hadamard-differentiability of Lorenz and Gini functionals are given in Bhattacharya (2007).

corollary[Limit theory for smooth functionals] Consider the parameter $\theta$ as an element of a parameter space $\mathbb{D}_{\theta} \subset \mathbb{D}= \times_{(jk) \in \mathcal{J} \mathcal{K}} \ell^{\infty}(\mathcal{Y}_j)$, with $\mathbb{D}_{\theta}$ containing the true value $\theta_0 = (F_{Y\langle j|k\rangle} : (j,k) \in \mathcal{J}\mathcal{K}) $. Consider the plug-in estimator $\widehat \theta = (\widehat F_{Y\langle j|k\rangle} : (j,k) \in\mathcal{J} \mathcal{K})$. Suppose $\phi(\theta) $, a functional of interest mapping $\mathbb{D} _{\theta}$ to $\ell^{\infty}( \mathcal{W})$, is Hadamard differentiable in $ \theta$ at $\theta_0$ tangentially to $\times_{(jk) \in \mathcal{J}\mathcal{K }} C(\mathcal{Y}_j)$ with derivative $(\phi^{\prime}_{jk}: (j,k) \in\mathcal{ J}\mathcal{K})$. Let $\Delta = \phi(\theta_0)$ and $\widehat\Delta = \phi(\widehat\theta)$. Then, under the conditions of Theorem (ref), part 1, \begin{equation} \sqrt{n} \left( \widehat{\Delta}(w) - \Delta(w) \right) \rightsquigarrow\sum_{(j,k) \in\mathcal{J}\mathcal{K}} (\phi^{\prime}_{jk} \bar Z_{jk})(w) =: T (w), \end{equation} as a stochastic processes indexed by $w \in\mathcal{W}$ in $\ell^{\infty}( \mathcal{W})$, where $T$ is a tight zero-mean Gaussian process.

Validity of resampling and other simulation methods for counterfactual analysis

As we discussed in Section 3.3, Kolmogorov-Smirnov type procedures offer a convenient and computationally attractive approach for performing inference on function-valued parameters using functional central limit theorems. A complication in our case is that the limit processes in ((ref))--((ref)) are non-pivotal, as their covariance functions depend on unknown, though estimable, nuisance parameters.\footnote{ Similar non-pivotality issues arise in a variety of goodness-of-fit problems studied by Durbin and others, and are referred to as the Durbin problem by Koenker and Xiao (2002).} We deal with this non-pivotality by using resampling and simulation methods. An attractive result shown as part of our theoretical analysis is that the counterfactual operator is Hadamard differentiable with respect to the underlying conditional and covariate distributions. As a result, if bootstrap or any other method consistently estimates the limit laws of the estimators of the conditional and covariate distributions, it also consistently estimates the limit laws of the estimators of the counterfactual distributions and their smooth functionals. This convenient result follows from the functional delta method for bootstrap of Hadamard differentiable functionals.

In order to state the results formally, we follow the notation and definitions in van der Vaart and Wellner (1996). Let $D_{n}$ denote the data vector and $M_{n}$ be the vector of random variables used to generate bootstrap draws or simulation draws given $D_{n}$ (this may depend on the particular resampling or simulation method). Consider the random element $ \mathbb{Z}^{*}_{n} = \mathbb{Z}_{n}(D_{n}, M_{n})$ in a normed space $ \mathbb{D}$. We say that the bootstrap law of $\mathbb{Z}^{*}_{n}$ consistently estimates the law of some tight random element $\mathbb{Z}$ and write $\mathbb{Z}^{*}_{n} \rightsquigarrow_{\mathbb{P}} \mathbb{Z} $ in $\mathbb{D}$ if

equation[equation omitted — 203 chars of source]

where $\text{BL}_{1}(\mathbb{D})$ denotes the space of functions with Lipschitz norm at most 1 and $E_{M_{n}}$ denotes the conditional expectation with respect to $M_{n}$ given the data $D_{n}$; $\rightarrow_{\mathbb{P}}$ denotes convergence in (outer) probability.

Next, consider the processes $\widehat\vartheta(t)$ $=$ $(\widehat F_{Y_{j}|X_{j}}(y|x),$ $\int f d \widehat F_{X_{k}})$ and $\vartheta(t)$ $=$ $(F_{Y_{j}|X_{j}}(y|x), \int f d F_{X_{k}}),$ indexed by $t=(y,x,j,k,f) \in T =\mathcal{Y}\mathcal{X}\mathcal{J}\mathcal{K}\mathcal{F}$, as elements of $ \mathbb{E}_{\vartheta} = \ell^{\infty} (T)^{2}$. Condition D(a) can be restated as $\sqrt{n} (\widehat\vartheta - \vartheta) \rightsquigarrow \mathbb{Z}_{\vartheta}$ in $\mathbb{E}_{\vartheta}$, where $\mathbb{Z} _{\vartheta}$ denotes the limit process in Condition D(a). Let $ \widehat\vartheta^{*}$ be the bootstrap draw of $\widehat\vartheta$. Consider the functional of interest $\phi=\phi(\vartheta)$ in the normed space $\mathbb{E}_{\phi}$, which can be either the counterfactual distribution and quantile functions considered in Theorem (ref) , the distribution or quantile effects considered in Corollary (ref), or any of the functionals considered in Corollary (ref). Denote the plug-in estimator of $\phi$ as $\widehat\phi= \phi(\widehat\vartheta)$ and the corresponding bootstrap draw as $\widehat \phi^{*} = \phi( \widehat\vartheta^{*})$. Let $\mathbb{Z}_{\phi}$ denote the limit law of $ \sqrt{n}(\widehat\phi- \phi)$, as described in Theorem (ref), Corollary (ref), and Corollary (ref).

theorem[Validity of resampling and other simulation methods for counterfactual analysis] Assume that the conditions of Theorem (ref) hold. If $\sqrt{n}(\widehat{\vartheta }^{\ast }-\widehat{\vartheta } )\rightsquigarrow _{\mathbb{P} }\mathbb{Z}_{\vartheta }$ in $\mathbb{E}_{\vartheta } $, then $\sqrt{n}(\widehat{\phi }^{\ast }-\widehat{\phi})\rightsquigarrow _{\mathbb{P} }\mathbb{Z}_{\phi } $ in $\mathbb{E}_{\phi }$. In words, if the exchangeable bootstrap or any other simulation method consistently estimates the law of the limit stochastic process in Condition D, then this method also consistently estimates the laws of the limit stochastic processes ((ref))--((ref)) for the estimators of counterfactual distributions, quantiles, distribution effects, quantile effects, and other functionals.

This is the second main and new result of the paper. Theorem (ref) shows that any resampling method is valid for estimating the limit laws of the estimators of the counterfactual effects, provided this method is valid for estimating the limit laws of the (function-valued) estimators of the conditional and covariate distributions. We verify the latter condition for our principal estimators in Section 5, where we establish the validity of exchangeable bootstrap methods for estimating the laws of function-valued estimators of the conditional distribution based on quantile regression and distribution regression processes. As noted in Remark (ref) , this result also implies the validity of the Kolmogorov-Smirnov type confidence bands for counterfactual effects under non-degeneracy of the variance function of the limit processes for the estimators of these effects; see Appendix A of the supplemental material for details.

Inference Theory for Counterfactual Analysis under Primitive Conditions

We verify that the high-level conditions of the previous section hold for the principal estimators of the conditional distribution functions, and so the various conclusions on inference methods also apply to this case. We also present new results on limit distribution theory for distribution regression processes and exchangeable bootstrap validity for quantile and distribution regression processes, which may be of a substantial independent interest. Throughout this section, we re-label $P(X)$ to $X$ to simplify the notation. This entails no loss of generality when $P(X)$ includes $X$ as one of the components.

Preliminaries on sampling.

We assume there are samples $\{ (Y_{ki}, X_{ki}): i=1,...,n_{k}\}$ composed of i.i.d. copies of $(Y_{k},X_{k})$ for all populations $k \in \mathcal{K}$. The samples are independent across $k \in \mathcal{K}_0 \subseteq \mathcal{K} $. We assume that $Y_{ji}$ is observable only for $j \in\mathcal{J} \subseteq \mathcal{K}_0$. We shall call the case with $\mathcal{K} = \mathcal{K}_0$ the independent samples case. The independent samples case arises, for example, in the wage decomposition application of Section 6. In addition, we may have transformation samples indexed by $k \in \mathcal{K}_t$ created via transformation of some “originating" samples $l \in \mathcal{K}_0$. For example, in the unconditional quantile regression mentioned in Section 2, we create a transformation sample by shifting one of the covariates in the original sample up by a unit.

We need to account for the dependence between the transformation and originating samples in the limit theory for counterfactual estimators. In order to do so formally, we specify the relation of each transformation sample, with index $k \in \mathcal{K}_t$, to an originating sample, with index $l(k) \in \mathcal{K}_0$, as follows: $(Y_{k i}, X_{k i}) = g_{l(k),k} (Y_{l(k) i}, X_{l(k)i}), i = 1,..., n_k$, where $g_{l(k),k}$ is a known measurable transformation map, and $l: \mathcal{K}_{t} \to\mathcal{K}_0 $ is the indexing function that gives the index $l(k)$ of the sample from which the transformation sample $k$ is originated. We also let $\mathcal{K} = \mathcal{K}_t \cup \mathcal{K}_0$. The main requirement on the map $ g_{l(k),k}: \mathbb{R}^{d_x+1} \mapsto \mathbb{R}^{d_x+1} $ is that it preserves the Dudley-Koltchinskii-Pollard's (DKP) sufficient condition for universal Donskerness: given a class $\mathcal{F}$ of suitably measurable and bounded functions mapping a measurable subset of $\mathbb{R}^{d_x+1}$ to $\mathbb{R} $ that obeys Pollard's entropy condition, the class $\mathcal{F} \circ g_{l(k),k} $ continues to contain bounded and suitably measurable functions, and obeys Pollard's entropy condition.\footnote{ The definitions of suitably measurable and Pollard's entropy condition are recalled in Appendix A. Together with boundedness, these are well-known sufficient conditions for a function class to be universal Donsker (Dudley, 1987, Koltchinskii, 1981, and Pollard, 1982).} For example, this holds if $ g_{l(k),k}$ is an affine or a uniformly Lipschits map. The following condition states formally the sampling requirements.

Condition SM. \ The samples $D_{k} = \{(Y_{ki}, X_{ki}): 1 \leq i \leq n_{k} \}$, $k \in \mathcal{K},$ are generated as follows: (a) For each population $k \in \mathcal{K}_0$, $D_{k}$ contains i.i.d. copies of the random vector $(Y_{k}, X_{k})$ that has probability law $P_{k}$, and $ D_k $ are independent across $k \in \mathcal{K}_0$. (b) For each population $ k \in \mathcal{K}_t$, the samples $D_{k}$ are created by transformation, $ D_k = \{g_{l(k),k} (Y_{l(k)i}, X_{l(k)i}): 1 \leq i \leq n_{l(k)} \}$ for $ l(k) \in \mathcal{K}_0$, where the maps $g_{l(k),k} $ preserve the DKP condition.

Lemma (ref) in Appendix D shows the following result under Condition SM: As $n \to\infty$ the empirical processes

equation[equation omitted — 175 chars of source]

converge weakly,

equation[equation omitted — 105 chars of source]

as stochastic processes indexed by $(k,f) \in\mathcal{K}\mathcal{F}$ in $ \ell^{\infty }(\mathcal{K}\mathcal{F})$. The limit processes $\mathbb{G}_{k}$ are tight $P_{k} $-Brownian bridges, which are independent across $k \in \mathcal{K}_0$,\footnote{ A zero-mean Gaussian process ${\mathbb{G}}_{k}$ is a $P_k$-Brownian bridge if its covariance function takes the form $E[{\mathbb{G}}_k(f) {\mathbb{G}} _k(l) ] = \int f l d P_k - \int f d P_k \int l d P_k$, for any $f$ and $l$ in $L^2(F_{X_k}) $; see van der Vaart (1998).} and for $k \in \mathcal{K}_t$ are defined by:

equation[equation omitted — 134 chars of source]

Exchangeable bootstrap

The following condition specifies how we should draw the bootstrap weights to mimic the dependence between the samples in the exchangeable bootstrap version of the estimators of counterfactual functionals described in Section 3.

Condition EB. \ For each $n_{k}$ and $k \in\mathcal{K}_{0}$ , $(w_{k1}, ..., w_{kn_{k}})$ is an exchangeable,\footnote{ A sequence of random variables $X_1, X_2, ..., X_n$ is exchangeable if for any finite permutation $\sigma$ of the indices $1,2, ..., n$ the joint distribution of the permuted sequence $X_{\sigma(1)}, X_{\sigma(2)}, ...,X_{\sigma(n)} $ is the same as the joint distribution of the original sequence.} nonnegative random vector, which is independent of the data $ (D_k)_{k \in \mathcal{K}}$, such that for some $\epsilon> 0$

equation[equation omitted — 263 chars of source]

where $\bar w_{k} = {n_{k}}^{-1} \sum_{i=1}^{n_{k}} w_{ki} $. Moreover, the vectors $(w_{k1}, ..., w_{kn_{k}})$ are independent across $k \in \mathcal{K} _0$. For each $k \in\mathcal{K}_{t}$,

equation[equation omitted — 82 chars of source]

}

remark[Common bootstrap schemes] As pointed out in van der Vaart and Wellner (1996), by appropriately selecting the distribution of the weights, exchangeable bootstrap covers the most common bootstrap schemes as special cases. The empirical bootstrap corresponds to the case where $(w_{k1},...,w_{kn_{k}})$ is a multinomial vector with parameter $n_{k}$ and probabilities $(1/n_{k},...,1/n_{k})$. The weighted bootstrap corresponds to the case where $w_{k1},...,w_{kn_{k}}$ are i.i.d. nonnegative random variables with $E[w_{k1}]=Var[w_{k1}]=1$, e.g. standard exponential. The $m$ out of $n$ bootstrap corresponds to letting $(w_{k1},...,w_{kn_{k}})$ be equal to $\sqrt{n_{k}/m_{k}}$ times multinomial vectors with parameter $ m_{k}$ and probabilities $(1/n_{k},...,1/n_{k})$. The subsampling bootstrap corresponds to letting $(w_{k1},...,w_{kn_{k}})$ be a row in which the number $n_{k}(n_{k}-m_{k})^{-1/2}m_{k}^{-1/2}$ appears $m_{k}$ times and 0 appears $n_{k}-m_{k}$ times ordered at random, independent of the data. \qed

Inference theory for counterfactual estimators based on quantile regression

We proceed to impose the following conditions on $(Y_j,X_j)$ for each $j \in \mathcal{J}$.

Condition QR. \ (a) The conditional quantile function takes the form $Q_{Y_{j}|X_{j}}(u|x)= x^{\prime}\beta_{j}(u)$ for all $u \in \mathcal{U}=[\varepsilon, 1- \varepsilon]$ with $0< \varepsilon< 1/2$, and $ x \in\mathcal{X}_{j}$. (b) The conditional density function $ f_{Y_{j}|X_{j}}(y|x)$ exists, is uniformly continuous in $(y,x)$ in the support of $(Y_j, X_j)$, and is uniformly bounded. (c) The minimal eigenvalue of $J_{j}(u)$ $=$ $E[f_{Y_{j}|X_{j}}( X_{j}^{\prime}\beta_j(u)|X_{j}) X_{j}X_{j}^{\prime}]$ is bounded away from zero uniformly over $u \in{\mathcal{U}}$. (d) $E\|X_{j}\|^{2+\epsilon} < \infty$ for some $\epsilon>0$.

In order to state the next result, let us define

eqnarray*[eqnarray* omitted — 361 chars of source]
theorem[Validity of QR based counterfactual analysis] Suppose that for each $j \in\mathcal{J},$ Conditions S, SM, and QR hold, the region of interest $\mathcal{Y}_{j}\mathcal{X}_{j}$ is a compact subset of $\mathbb{R}^{1+ d_x},$ and $\mathcal{U}_j := \{ u : x^{\prime }\beta_j (u) \in \mathcal{Y}_{j}, \text{ for some } x \in \mathcal{ X}_j\} \subseteq \mathcal{U}$. Then, (1) Condition D holds for the quantile regression estimator ((ref)) of the conditional distribution and the empirical distribution estimator ((ref)) of the covariate distribution. The limit processes are given by \begin{equation*} Z_{j}(y,x) = {\mathbb{G}}_{j}( \ell_{j,y,x}), \ \ G_{{k}}(f) = {\mathbb{G}} _{k}(f), \ \ (j,k) \in \mathcal{J}\mathcal{K}, \end{equation*} where ${\mathbb{G}}_{k}$ are the $P_{k}$-Brownian bridges defined in ((ref)) and ((ref)). In particular, $ \{F_{Y_{j}|X_{j}}(y|\cdot): y \in\mathcal{Y}_{j}\}$ is a universal Donsker class. (2) Exchangeable bootstrap consistently estimates the limit law of these processes under Condition EB. (3) Therefore, all conclusions of Theorems (ref)- (ref) and Corollaries (ref) - (ref) apply. In particular, the limit law for the estimated counterfactual distribution is given by $\bar Z_{jk}(y) := {\mathbb{G}}_{j} ( \kappa_{jk,y}), $ with covariance function $E[\bar Z_{jk}(y) \bar Z_{lm}(\bar y)] = E[ \kappa_{jk,y} \kappa_{ lm,\bar y} ] - E[ \kappa_{jk, y} ] E[ \kappa_{lm,\bar y}]$.

This is the third main and new result of the paper. It derives the joint functional central limit theorem for the quantile regression estimator of the conditional distribution and the empirical distribution function estimator of the covariate distribution. It also shows that exchangeable bootstrap consistently estimates the limit law. Moreover, the result characterizes the limit law of the estimator of the counterfactual distribution in Theorem 4.1, which in turn determines the limit laws of the estimators of the counterfactual quantile functions and other functionals, via Theorem 4.1 and Corollaries 4.1 and 4.2. Note that $\mathcal{U}_j \subseteq \mathcal{U}$ is the condition that permits the use of trimming in ( (ref)), since it says that the conditional distribution of $Y_j$ given $X_j$ on the region of interest $\mathcal{Y}_j\mathcal{X}_j$ is not determined by the tail conditional quantiles.

While proving Theorem (ref), we establish the following corollaries that may be of independent interest.

corollary[Validity of exchangeable bootstrap for QR coefficient process] Let $\{(Y_{ji},X_{ji}): 1\leq i \leq n_{j}\}$ be a sample of i.i.d. copies of the random vector $(Y_{j},X_{j})$ that has probability law $ P_{j}$ and obeys Condition QR. (1) As $n_{j} \to\infty$, the QR coefficient process possesses the following first order approximation and limit law: $ \sqrt{n_{j}}(\widehat\beta_{j}(\cdot) - \beta_{j}(\cdot)) = \widehat{\mathbb{ G}}_j(\psi_{j,\cdot}) + o_{\mathbb{P}}(1) \rightsquigarrow {\mathbb{G}} _{j}(\psi_{j,\cdot} ) $ in $\ell^{\infty}({\mathcal{U}})^{d_x}$, where ${ \mathbb{G}}_j$ is a $P_{j}$- Brownian Bridge. (2) The exchangeable bootstrap law is consistent for the limit law, namely, as $n_{j} \to\infty$, \begin{equation*} \sqrt{n_{j}}(\widehat\beta^{*}_{j}(\cdot) - \widehat\beta_{j}(\cdot)) \rightsquigarrow _{\mathbb{P}} {\mathbb{G}}_{j}(\psi_{j,\cdot} ) in \ell^{\infty}({\mathcal{U}})^{d_x}. \end{equation*}

The result (2) is new and shows that exchangeable bootstrap (which includes empirical bootstrap, weighted bootstrap, $m$ out of $n$ bootstrap, and subsampling) is valid for estimating the limit law of the entire QR coefficient process. Previously, such a result was available only for pointwise cases (e.g. Hahn, 1995, 1997, and Feng, He, and Hu, 2011), and the process result was available only for subsampling (Chernozhukov and Fernandez-Val, 2005, and Chernozhukov and Hansen, 2006).

Let $\widehat Q_{Y_j|X_j}(u|x):= x^{\prime }\widehat \beta_j(u)$ be the QR estimator of the conditional quantile function, and $u \mapsto \widehat Q^r_{Y_j|X_j}(u|x)$ be the non-decreasing rearrangement of $u \mapsto \widehat Q_{Y_j|X_j}(u|x)$ over the region $\mathcal{U}_j$. Let $\widehat F_{Y_j|X_j}(y|x) $ be the QR estimator of the conditional distribution function defined in ((ref)). Also, we use the star superscript to denote the bootstrap versions of all these estimators, and define

equation*[equation* omitted — 83 chars of source]
corollary[Limit law and exchangeable bootstrap for QR-based estimators of conditional distribution and quantile functions] Suppose that the conditions of Theorem (ref) hold. Then, (1) As $n_{j} \to\infty$, in $\ell^{\infty}({\mathcal{U}}_{j}\mathcal{X }_{j})$, $\sqrt{n_{j}}(\widehat Q_{Y_j|X_j}(u|x) - Q_{Y_j|X_j}(u|x) ) = \widehat { \mathbb{G}}_{j}( \bar \ell_{j,u,x}) + o_{\mathbb{P}}(1) \rightsquigarrow{ \mathbb{G}}_{j}( \bar \ell_{j,u,x}),$ and $\sqrt{n_{j}}(\widehat Q^r_{Y_j|X_j}(u|x) - Q_{Y_j|X_j}(u|x) ) = \widehat {\mathbb{G}}_{j}( \bar \ell_{j,u,x}) + o_{\mathbb{P}}(1) \rightsquigarrow{\mathbb{G}}_{j}( \bar \ell_{j,u,x}),$ as stochastic processes indexed by $(u,x) \in {\mathcal{U}} _{j}\mathcal{X}_{j}$. In $\ell^{\infty}({\mathcal{Y}}_{j}\mathcal{X}_{j})$, $ \sqrt{n_{j}}(\widehat F_{Y_j|X_j}(y|x) - F_{Y_j|X_j}(y|x) ) = \widehat { \mathbb{G}}_{j}( \ell_{j,y,x} ) + o_{\mathbb{P}}(1) \rightsquigarrow{\mathbb{G}} _{j}(\ell_{j,y,x}),$ as a stochastic process indexed by $(y,x) \in {\mathcal{ Y}}_{j}\mathcal{X}_{j}$. (2) The exchangeable bootstrap law is consistent for the limit laws, namely, as $n_{j} \to\infty$, in $\ell^{\infty}({ \mathcal{U}}_{j}\mathcal{X}_{j})$, $\sqrt{n_{j}}(\widehat Q^*_{Y_j|X_j}(u|x) - \widehat Q_{Y_j|X_j}(u|x) ) \rightsquigarrow_{\mathbb{P}} {\mathbb{G}}_{j}( \bar \ell_{j,u,x}),$ and $\sqrt{n_{j}}(\widehat Q^{r*}_{Y_j|X_j}(u|x) - \widehat Q^{r}_{Y_j|X_j}(u|x) ) \rightsquigarrow_{\mathbb{P}} {\mathbb{G}}_{j}( \bar \ell_{j,u,x}),$ as stochastic processes indexed by $(u,x) \in {\mathcal{U}} _{j}\mathcal{X}_{j}$. In $\ell^{\infty}({\mathcal{Y}}_{j}\mathcal{X}_{j})$, $ \sqrt{n_{j}}(\widehat F^*_{Y_j|X_j}(y|x) - \widehat F_{Y_j|X_j}(y|x) ) \rightsquigarrow_{\mathbb{P}} {\mathbb{G}}_{j}( \ell_{j,y,x}),$ as a stochastic process indexed by $(y,x) \in {\mathcal{Y}}_{j}\mathcal{X}_{j}$.

Corollary (ref) establishes first order approximations, functional central limit theorems and exchangeable bootstrap validity for QR-based estimators of the conditional distribution and quantile functions. The two estimators of the conditional quantile function -- $\widehat Q _{Y_j|X_j}$ and $\widehat Q^r_{Y_j|X_j}$ -- are asymptotically equivalent. However, $ \widehat Q_{Y_j|X_j}$ is not necessarily monotone, while $\widehat Q^r_{Y_j|X_j}$ is monotone and has better finite sample properties (Chernozhukov, Fernandez-Val, and Galichon, 2009).

Inference Theory for Counterfactual Estimators based on Distribution Regression

We shall impose the following conditions on $(Y_{j},X_{j})$ for each $j \in \mathcal{J}$.

Condition DR. \ (a) The conditional distribution function takes the form $F_{Y_{j}|X_{j}}(y|x)=\Lambda(x^{\prime}\beta_j(y))$ for all $ y \in \mathcal{Y}_j$ and $x \in \mathcal{X}_{j}$, where $\Lambda$ is either the probit or logit link function. (b) The region of interest $\mathcal{Y}_j$ is either a compact interval in $\mathbb{R}$ or a finite subset of $\mathbb{R }$. In the former case, the conditional density function $ f_{Y_{j}|X_{j}}(y|x)$ exists, is uniformly bounded and uniformly continuous in $(y,x)$ in the support of $(Y_j,X_j)$. (c) $E\|X_{j}\|^{2} < \infty$ and the minimum eigenvalue of

equation*[equation* omitted — 192 chars of source]

is bounded away from zero uniformly over $y \in \mathcal{Y}_{j}$, where $ \lambda$ is the derivative of $\Lambda$.}

In order to state the next result, we define

eqnarray*[eqnarray* omitted — 471 chars of source]
theorem[Validity of DR based counterfactual analysis] Suppose that for each $j \in\mathcal{J}$, Conditions S, SM, and DR hold, and the region $\mathcal{Y}_{j}\mathcal{X}_{j}$ is a compact subset of $\mathbb{R}^{1 + d_x}$. Then, (1) Condition D holds for the distribution regression estimator ((ref)) of the conditional distribution and the empirical distribution estimator ((ref)) of the covariate distribution, with limit processes given by \begin{equation*} Z_{j}(y,x) = {\mathbb{G}}_{j}( \ell_{j,y,x}), \ \ G_{{k}}(f) = {\mathbb{G}} _{k}(f), \ \ (j,k) \in \mathcal{J}\mathcal{K}, \end{equation*} where ${\mathbb{G}}_{k}$ are the $P_{k}$-Brownian bridges defined in ((ref)) and ((ref)). In particular, $ \{F_{Y_{j}|X_{j}}(y|\cdot) : y \in\mathcal{Y}_{j}\}$ is a universal Donsker class. (2) Exchangeable bootstrap consistently estimates the limit law of these processes under Condition EB. (c) Therefore, all conclusions of Theorem (ref) and (ref), and of Corollaries (ref) and (ref) apply to this case. In particular, the limit law for the estimated counterfactual distribution is given by $\bar Z_{jk}(y) := { \mathbb{G}}_{j} ( \kappa_{jk,y}), $ with covariance function $E [\bar Z_{jk}(y) \bar Z_{lm}(\bar y)] = E[ \kappa_{jk,y} \kappa_{ lm,\bar y} ] - E[ \kappa_{jk, y} ] E[ \kappa_{lm,\bar y}]$.

This is the fourth main and new result of the paper. It derives the joint functional central limit theorem for the distribution regression estimator of the conditional distribution and the empirical distribution function estimator of the covariate distribution. It also shows that bootstrap consistently estimates the limit law. Moreover, the result characterizes the limit law of the estimator of the counterfactual distribution in Theorem 4.1, which in turn determines the limit laws of the estimators of the counterfactual quantiles and other functionals, via Theorem 4.1 and Corollaries 4.1 and 4.2.

While proving Theorem (ref), we also establish the following corollaries that may be of independent interest.

corollary[Limit law and exchangeable bootstrap for DR coefficient process] Let $\{(Y_{ji},X_{ji}): 1 \leq i \leq n_{j}\}$ be a sample of i.i.d. copies of the random vector $(Y_{j},X_{j})$ that has probability law $ P_{j}$ and obeys Condition DR. (1) As $n_{j} \to\infty$, the DR coefficient process possesses the following first order approximation and limit law: \begin{equation*} \sqrt{n_{j}}(\widehat\beta_{j}(\cdot) - \beta_{j}(\cdot)) = \widehat { \mathbb{G}}_{j}( \psi_{j,\cdot}) + o_{\mathbb{P}}(1) \rightsquigarrow{\mathbb{G}} _{j}(\psi_{j,\cdot}) in \ell^{\infty}({\mathcal{Y}}_{j})^{d_x}, \end{equation*} where ${\mathbb{G}}_j$ is a $P_{j}$- Brownian Bridge. (2) The exchangeable bootstrap law is consistent for the limit law, namely, as $n_{j} \to\infty$, \begin{equation*} \sqrt{n_{j}}(\widehat\beta^{*}_{j}(\cdot) - \widehat\beta_{j}(\cdot)) \rightsquigarrow _{\mathbb{P}} {\mathbb{G}}_{j}(\psi_{j,\cdot} ) in \ell^{\infty }({\mathcal{Y}}_{j})^{d_x}. \end{equation*}

Let $\widehat F_{Y_j|X_j}(y|x):= \Lambda(x^{\prime }\widehat \beta_j(y))$ be the DR estimator of the conditional distribution function, and $y \mapsto \widehat F^r_{Y_j|X_j}(y|x)$ be the non-decreasing rearrangement of $y \mapsto \widehat F_{Y_j|X_j}(y|x)$ over the region $\mathcal{Y}_j$. Let $\widehat Q_{Y_j|X_j}(u|x) = \widehat F^{r\leftarrow}_{Y_j|X_j}(u|x)$ be the DR estimator of the conditional quantile function, obtained by inverting the rearranged estimator of the distribution function over the region $ \mathcal{U}_j$. Here, $\mathcal{U}_j \subset (0,1)$ can be any compact interval of quantile indices such that an $\epsilon$-expansion of the region $\{Q_{Y_j|X_j}(u|x) : u \in \mathcal{U}_j\}$ is contained in $\mathcal{Y}_j$ , for all $x \in \mathcal{X}_j$. Also, we use the star superscript to denote the bootstrap versions of all these estimators, and define

equation*[equation* omitted — 130 chars of source]
corollary[Limit law and exchangeable bootstrap for DR-based estimators of conditional distribution and quantile functions] Suppose that the region of interest $\mathcal{Y}_{j}\mathcal{X }_{j}$ is a compact subset of $\mathbb{R}^{1+d_x}$, $\mathcal{Y}_j$ is an interval, the conditions of Corollary (ref) hold, and $ f_{Y_j|X_j}(y|x)>0$ on $\mathcal{Y}_j\mathcal{X}_j$. Then, (1) As $n_{j} \to\infty$, in $\ell^{\infty}({\mathcal{Y}}_{j}\mathcal{X}_{j})$, $\sqrt{ n_{j}}(\widehat F_{Y_j|X_j}(y|x) - F_{Y_j|X_j}(y|x) ) = \widehat {\mathbb{G}} _{j}( \ell_{j,y,x}) + o_{\mathbb{P}}(1) \rightsquigarrow{\mathbb{G}}_{j}( \ell_{j,y,x}),$ and $\sqrt{n_{j}}(\widehat F^r_{Y_j|X_j}(y|x) - F_{Y_j|X_j}(y|x) ) = \widehat {\mathbb{G}}_{j}( \ell_{j,y,x}) + o_{\mathbb{P}}(1) \rightsquigarrow{\mathbb{G}}_{j}( \ell_{j,y,x}),$ as stochastic processes indexed by $(y,x) \in {\mathcal{Y}}_{j}\mathcal{X}_{j}$. In $\ell^{\infty}({ \mathcal{U}}_{j}\mathcal{X}_{j})$, $\sqrt{n_{j}}(\widehat Q_{Y_j|X_j}(u|x) - Q_{Y_j|X_j}(u|x) ) = \widehat {\mathbb{G}}_{j}( \bar \ell_{j,u,x} ) + o_{\mathbb{P}}(1) \rightsquigarrow{\mathbb{G}}_{j}( \bar \ell_{j,u,x}),$ as a stochastic process indexed by $(u,x) \in {\mathcal{U}}_{j}\mathcal{X}_{j}$. (2) The exchangeable bootstrap law is consistent for the limit laws, namely, as $n_{j} \to\infty$, in $\ell^{\infty}({\mathcal{Y}}_{j}\mathcal{X}_{j})$, $ \sqrt{n_{j}}(\widehat F^*_{Y_j|X_j}(y|x) - \widehat F_{Y_j|X_j}(y|x) ) \rightsquigarrow_{\mathbb{P}} {\mathbb{G}}_{j}( \ell_{j,y,x}),$ and $\sqrt{n_{j}} (\widehat F^{r*}_{Y_j|X_j}(y|x) - \widehat F^{r}_{Y_j|X_j}(y|x) ) \rightsquigarrow_{\mathbb{P}} {\mathbb{G}}_{j}( \ell_{j,y,x}),$ as stochastic processes indexed by $(y,x) \in {\mathcal{Y}}_{j}\mathcal{X}_{j}$. In $ \ell^{\infty}({\mathcal{U}}_{j}\mathcal{X}_{j})$, $\sqrt{n_{j}}(\widehat Q^*_{Y_j|X_j}(u|x) - \widehat Q_{Y_j|X_j}(u|x) ) \rightsquigarrow_{\mathbb{P}} { \mathbb{G}}_{j}( \bar \ell_{j,u,x}),$ as a stochastic process indexed by $ (u,x) \in {\mathcal{U}}_{j}\mathcal{X}_{j}$.

Corollary (ref) establishes first order approximations, functional central limit theorems and exchangeable bootstrap validity for DR-based estimators of the conditional distribution and quantile functions. The two estimators of the conditional distribution function -- $\widehat F_{Y_j|X_j}$ and $\widehat F^r_{Y_j|X_j}$ -- are asymptotically equivalent. However, $ \widehat F_{Y_j|X_j}$ is not necessarily monotone, while $\widehat F^r_{Y_j|X_j}$ is monotone and has better finite sample properties (Chernozhukov, Fernandez-Val, and Galichon, 2009).

The limit distribution and bootstrap consistency results in Corollaries (ref) and (ref) are new. They have already been applied in several studies (Chernozhukov, Fernandez-Val and Kowalski, 2011, Rothe, 2012, and Rothe and Wied, 2012). Note that unlike Theorem (ref) and Corollary (ref), Corollary (ref) does not rely on compactness of the region $\mathcal{Y}_{j}\mathcal{X}_{j}$.

Labor Market Institutions and the Distribution of Wages

In this section we apply our estimation and inference procedures to re-analyze the evolution of the U.S. wage distribution between 1979 and 1988. The first goal here is to compare the methods proposed in Section 3 and to discuss the various choices that practitioners need to make. The second goal is to provide support for the findings of DiNardo, Fortin, and Lemieux (1996, DFL hereafter) with a rigorous econometric analysis. Indeed, we provide confidence intervals for real-valued and function-valued effects of the institutional and labor market factors driving changes in the wage distribution, thereby quantifying their economic and statistical significance. We also provide a variance decomposition of the covariate composition effect into within-group and between-group components.

We use the same dataset and variables as in DFL, extracted from the outgoing rotation groups of the Current Population Surveys (CPS) in 1979 and 1988. The outcome variable of interest is the hourly log-wage in 1979 dollars. The regressors include a union status indicator, nine education dummy variables interacted with experience, a quartic term in experience, two occupation dummy variables, twenty industry dummy variables, and indicators for race, SMSA, marital status, and part-time status. Following DFL we weigh the observations by the product of the CPS sampling weights and the hours worked. We analyze the data only for men for the sake of brevity.\footnote{ Results for women can be found in Appendix C of the supplemental material.}

The major factors suspected to have an important role in the evolution of the wage distribution between 1979 and 1988 are the minimum wage, whose real value declined by 27 percent, the level of unionization, whose level declined from 32 percent to 21 percent in our sample, and the characteristics of the labor force, whose education levels and other characteristics changed substantially during this period. Thus, following DFL, we decompose the total change in the US wage distribution into the sum of four effects: (1) the effect of the change in minimum wage, (2) the effect of de-unionization, (3) the effect of changes in the characteristics of the labor force other than unionization, and (4) the wage structure effect. We stress that this decomposition has a causal interpretation only under additional conditions analogous to the ones laid out in Section 2.3.

We formally define these four effects as differences between appropriately chosen counterfactual distributions. Let $F_{Y{\langle (t,s)|(r,v)\rangle }}$ denote the counterfactual distribution of log-wages $Y$ when the wage structure is as in year $t$, the minimum wage $M$ is at the level observed in year $s$, the union status $U$ is distributed as in year $r$, and the other worker characteristics $C$ are distributed as in year $v$. We use two indexes to refer to the conditional and covariate distributions because we treat the minimum wage as a feature of the conditional distribution and we want to separate union status from the other covariates. Given these counterfactual distributions, we can decompose the observed change in the distribution of wages between 1979 (year 0) and 1988 (year 1) into the sum of the previous four effects: {

equation[equation omitted — 524 chars of source]

} In constructing the decompositions ((ref)), we follow the same sequential order as in DFL.\footnote{ The order of the decomposition matters because it defines the counterfactual distributions and effects of interest. We report some results for the reverse sequential order in Appendix C of the supplemental material. The results are similar under the two alternative sequential orders.}

We next describe how to identify and estimate the various counterfactual distributions appearing in ((ref)). The first counterfactual distribution is $F_{Y{\langle (1,0)|(1,1)\rangle }}$, the distribution of wages that we would observe in 1988 if the real minimum wage was as high as in 1979. Identifying this quantity requires additional assumptions.\footnote{ We cannot identify this quantity from random variation in minimum wage, since the same federal minimum wage applies to all individuals and state level minimum wages varied little across states in the years considered.} Following DFL, the first strategy we employ is to assume the conditional wage density at or below the minimum wage depends only on the value of the minimum wage, and the minimum wage has no employment effects and no spillover effects on wages above its level. Under these conditions, DFL show that

equation[equation omitted — 333 chars of source]

where $F_{Y_{(t,s)}|X_{t}}(y|x)$ denotes the conditional distribution of wages in year $t$ given worker characteristics $X_{t}=(U_{t},C_{t})$ when the level of the minimum wage is as in year $s$, and $m_{s}$ denotes the level of the minimum wage in year $s$. The second strategy we employ completely avoids modeling the conditional wage distribution below the minimal wage by simply censoring the observed wages below the minimum wage to the value of the minimum wage, i.e.

equation[equation omitted — 208 chars of source]

Given either ((ref)) or ((ref)) we identify the counterfactual distribution of wages using the representation:

equation[equation omitted — 116 chars of source]

where $F_{X_{t}}$ is the joint distribution of worker characteristics and union status in year $t$. The other counterfactual marginal distributions we need are

equation[equation omitted — 154 chars of source]

and

equation[equation omitted — 142 chars of source]

All the components of these distributions are identified and we estimate them using the plug-in principle. In particular, we estimate the conditional distribution $F_{U_{0}|C_{0}}(u|c),u\in \{0,1\},$ by logistic regression, and $F_{X_{1}}$, $F_{C_{1}}$ and $F_{X_{0}}$ by the empirical distributions.

From a practical standpoint, the main implementation decision concerns the choice of the estimator of the conditional distributions, $ F_{Y_{(j,j)}|X_{j}}\left( y|x\right) ,$ for $j\in \{0,1\}$. We consider the use of quantile regression, distribution regression, classical regression, and duration/transformation regression. The classical regression and the duration regression models are parsimonious special cases of the first two models. However, these models are not appropriate in this application due to substantial conditional heteroskedasticity in log wages (Lemieux, 2006, and Angrist, Chernozhukov, and Fernandez-Val, 2006). As the additional restrictions that these two models impose are rejected by the data, we focus on the distribution and quantile regression approaches.

Distribution and quantile regressions impose different parametric restrictions on the data generating process. A linear model for the conditional quantile function may not provide a good approximation to the conditional quantiles near the minimum wage, where the conditional quantile function may be highly nonlinear. Indeed, under the assumptions of DFL the conditional wage function has different determinants below and above the minimum wage. In contrast, the distribution regression model may well capture this type of behavior, since it allows the model coefficients to depend directly on the wage levels.

A second characteristic of our data set is the sizeable presence of mass points around the minimum wage and at some other round-dollar amounts. For instance, 20% of the wages take exactly 1 out of 6 values and 50% of the wages take exactly 1 out of 25 values. We compare the distribution and quantile regression estimators in a simulation exercise calibrated to fit many properties of the data set. The results presented in Appendix B of the supplemental material show that quantile regression is more accurate when the dependent variable is continuous but performs worse than distribution regression in the presence of realistic mass points. Based on these simulations and on specification tests that reject the linear quantile regression model, we employ the distribution regression approach to generate the main empirical results.\footnote{ Rothe and Wied (2012) propose new specification tests for conditional distribution models. Applying their tests to a similar dataset, they reject the quantile regression model but not the distribution regression model.} Since most of the problems for quantile regression take place in the region of the minimum wage, we also check the robustness of our results with a censoring approach. We censor wages from below at the value of the minimum wage and then apply censored quantile and distribution regressions to the resulting data.

We present our empirical results in Table 1 and Figures 1--3. In Table 1, we report the estimation and inference results for the decomposition ((ref)) of the changes in various measures of wage dispersion between 1979 and 1988 estimated using logit distribution regressions. Figures 1-3 refine these results by presenting estimates and 95% simultaneous confidence intervals for several major counterfactual effects of interest, including quantile, distribution and Lorenz effects. We construct the simultaneous confidence bands using 100 bootstrap replications and a grid of quantile indices $\left\{ 0.02,0.021,...,0.98\right\} $.

We see in the top panels of Figures 1-3 that the low end of the distribution is significantly lower in 1988 while the upper end is significantly higher in 1988. This pattern reflects the well-known increase in wage inequality during this period. Next we turn to the decomposition of the total change into the sum of the four effects. For this decomposition we focus mostly on quantile functions for comparability with recent studies and to facilitate interpretation.\footnote{ Discreteness of wage data implies that the quantile functions have jumps. To avoid this erratic behavior in the graphical representations of the results, we display smoothed quantile functions. The non-smoothed results are available from the authors. The quantile functions were smoothed using a bandwidth of 0.015 and a Gaussian kernel. The results in Table 1 have not been smoothed.} From Figure 1, we see that the contribution of de-unionization to the total change is quantitatively small and has a U-shaped effect across the quantile indexes. The magnitude and shape of this effect on the marginal quantiles between the first and last decile sharply contrast with the quantitatively large and monotonically decreasing shape of the effect of the union status on the conditional quantile function for this range of indexes (Chamberlain, 1994).\footnote{ We find similar estimates to Chamberlain (1994) for the effect of union status on the conditional quantile function in our CPS data.} This comparison illustrates the difference between conditional and unconditional effects. The unconditional effects depend not only on the conditional effects but also on the characteristics of the workers who switched their unionization status. Obviously, de-unionization cannot affect those who were not unionized at the beginning of the period, which is 70 percent of the workers. In our data, the unionization rate declines from 32 to 21 percent, thus affecting only 11 percent of the workers. Thus, even though the conditional impact of switching from union to non-union status can be quantitatively large, it has a quantitatively small effect on the marginal distribution.

From Figure 1, we also see that the change in the distribution of worker characteristics (other than union status) is responsible for a large part of the increase in wage inequality. The importance of these composition effects has been recently stressed by Lemieux (2006) and Autor, Katz and Kearney (2008). The composition effect, including the de-unionization and worker characteristics effects, is realized through two channels: between-group and within-group inequality. To understand the effect of these channels on wage dispersion it is useful to consider a linear quantile model $Y=X^{\prime }\beta (U)$, where $X$ is independent of $U$. By the law of total variance, we can decompose the variance of $Y$ into:

equation[equation omitted — 140 chars of source]

where between-group inequality corresponds to the first term and within-group inequality corresponds to the second term.\footnote{ See Aaberge, Bjerve, and Doksum (2005) for an analogous decomposition of the pseudo-Lorenz curve. Similar within-between decompositions can also be constructed using distribution regression models.} When we keep the coefficients fixed, a change in the distribution of the covariates increases inequality through the first channel if the variance of the covariates increases and through the second channel if the proportion of high-variance groups increases. In our case, both components increased by about 10% between 1979 and 1988. The increase in the proportion of college graduate from 19% to 23% is an example of the observed composition changes. It raised between-group inequality because highly educated workers earn conditional average wages well above the unconditional average, and within-group inequality because of the higher wage volatility faced by these workers.\footnote{ This is an empirical fact in our data set and not a theoretical fact. Increasing the proportion of educated workers can in principle either increase or decrease either component. To compute these effects, we kept the coefficients constant at their values obtained from estimating $F_{Y\left( 1,0\right) |X}$ and changed the distribution of the covariates from $ F_{X_{0}}$ to $F_{X_{1}}$. See Appendix D of the supplemental material for more details on the computation of the variance decomposition.}

We also include estimates of the wage structure effect, sometimes referred to as the price effect, which captures changes in the conditional distribution of log hourly wages. It represents the difference we would observe if the distribution of worker characteristics and union status, and the minimum wage remained unchanged during this period. This effect has a U-shaped pattern, which is similar to the pattern Autor, Katz and Kearney (2006a) find for the period between 1990 and 2000. They relate this pattern to a bi-polarization of employment into low and high skill jobs. However, they do not find a U-shaped pattern for the period between 1980 and 1990. A possible explanation for the apparent absence of this pattern in their analysis might be that the declining minimum wage masks this phenomenon. In our analysis, once we control for this temporary factor, we do uncover the U-shaped pattern for the price component in the 80s.

In Figure C1 of the supplemental material , we check the robustness of the results with respect to the link function used to implement the DR estimator. The results previously analyzed were obtained with a logistic link function. The differences between the estimates obtained with the logistic, normal, uniform (linear probability model), Cauchy and complementary log-log link functions are so modest that the lines are almost indistinguishable. As we mentioned above, the assumptions about the minimum wage are also delicate, since the mechanism that generates wages strictly below this level is not clear; it could be measurement error, non-coverage, or non-compliance with the law. To check the robustness of the results to the DFL assumptions about the minimum wage and to our semi-parametric model of the conditional distribution, we re-estimate the decomposition using censored linear quantile regression and censored distribution regression with a logit link, censoring the wage data below the minimum wage. For censored quantile regression, we use Powell's (1986) censored quantile regression estimated by Chernozhukov and Hong's (2002) algorithm. For censored distribution regression, we simply censor to zero the distribution regression estimates of the conditional distributions below the minimum wage and recompute the functionals of interest. We find the results in Figure C2 of the Supplemental Material to be very similar for the quantile and distribution regressions, and they are not very sensitive to the censoring.

Overall, our estimates and confidence intervals reinforce the findings of DFL, giving them a rigorous econometric foundation. Even though the sample size is large, the precision of some of the estimates was unclear to us a priori. For instance, only a relatively small proportion of workers are affected by unions. We provide standard errors and confidence intervals, which demonstrate the statistical and economic significance of the results. Moreover, we validate the results with a wide array of estimation methods. The similarity of the estimates may come as a surprise because the estimators make different parametric assumptions. However, in a fully saturated model all the estimators we have applied would give numerically the same results. The similarity of the results can be explained by the flexibility of our parametric model. Finally, we give a variance decomposition of the composition effect that shows that the increase in wage inequality is due to both between-group and within-group inequality components.

Conclusion and directions for future work

This paper develops methods for performing inference about the effect on an outcome of interest of a change in either the distribution of covariates or the relationship of the outcome with these covariates. The validity of the proposed inference procedures in large samples relies only on the applicability of a functional central limit theorem and the consistency of the bootstrap for the estimators of the covariate and conditional distributions. These conditions hold for the empirical distribution function estimator of the covariate distribution and for the most common regression estimators of the conditional distribution, such as classical, quantile, duration/transformation, and distribution regressions. Thus, we offer valid inference procedures for several popular existing estimators and introduce distribution regression to estimate counterfactual distributions.

We focus on functionals of the marginal counterfactual distributions but we do not consider their joint distribution. This joint distribution is required to compute other economically interesting quantities such as the distribution of the counterfactual effects. Abbring and Heckman (2007) discuss various ways to identify the distribution of these effects. The working paper version of this article provides inference procedures under a rank invariance assumption.

We focus on semi-parametric estimators of the conditional distribution due to their dominant role in empirical work (Angrist and Pischke, 2008). We hope to extend the analysis to nonparametric estimators in future work. Fully nonparametric estimators are in principle attractive but their implementation in samples of moderate size might be problematic. Rothe (2010) makes first steps in this direction and highlights some of the difficulties.

As mentioned in Foonote 1, our general results do not require the observability of the outcome of interest. If $F_{Y_{j}|X_{j}}\left( y|x\right) $ is redefined as the conditional distribution of a latent outcome and an estimator $\widehat{F}_{Y_{j}|X_{j}}\left( y|x\right) $ that satisfies Condition D is available, then the inference results in Section 4 apply. An interesting example is given by the policy relevant treatment effects of Heckman and Vytlacil (2005). They consider a class of policies that affect the probability of participation in a program but do not affect directly the structural function of the outcome in a model with endogeneity. For instance, one may be interested in the effect of decreasing college tuition on wages. In their model, the policy relevant treatment effect is the conditional marginal treatment effect integrated over the covariate distribution and the error term in the participation equation. This type of policy effects is outside of the scope of this paper but is certainly worth pursuing in future research.