Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
102,890 characters · 11 sections · 69 citation commands
Generic Inference on Quantile and Quantile Effect Functions for Discrete Outcomes
The quantile function (QF), introduced by galton1874proposed, has become a standard tool for descriptive and inferential analysis due to its straightforward and intuitive interpretation. For instance, quantiles play a crucial role in exploratory data analysis as advocated by Tukey and his co-authors. In a classical book, tukey1977exploratory promoted the use of the five-number summary, box-plot and median smoothers, which are all based on quantiles. doksum1974empirical and erich1975nonparametrics suggested to report the quantile effect (QE) function---the difference between two QFs---to compare the entire distribution of a random variable in two different populations. In randomized control trials and natural experiments, QEs have a causal interpretation, and are usually referred to as quantile treatment effects (QTEs). Quantile regression (QR), introduced by koenker1978regression, extends this concept to non-binary treatments and multivariate models. The monograph of Koenker05 and the handbook edited by koenker2017handbook provide thorough reviews of recent developments of the quantile regression methodology with applications to treatment effects, survival analysis, longitudinal data, time series, and financial data, among others. chernozhukov+13inference develop the use of QR and distribution regression (DR) methods to construct counterfactual distribution and quantile functions, which serve as building blocks in policy and decomposition analysis.
In this paper we propose a generic procedure to obtain confidence bands for QFs and QE functions that are valid for continuous, discrete and mixed discrete-continuous outcomes, including counterfactual or latent outcomes. Plotting these bands allows us to visualize the sampling uncertainty associated with the estimates of these functions. The bands have a straightforward interpretation: they cover the true functions with a pre-specified probability, e.g. 95%, such that any function that lies outside of the band even at a single quantile can be rejected at the corresponding level, e.g. 5%. In addition, they are versatile: the same confidence band can be used for testing different null hypotheses. The researcher does not even need to know the hypothesis that will be considered by the reader. For instance, the hypothesis that a treatment has no effect on an outcome can be rejected if the confidence band for the QE function does not cover the zero line. First-order stochastic dominance implies that some non-negative values are covered at all quantiles. The location-shift hypothesis that the QE function is constant implies that there is at least one value covered by the band at all quantiles.
The proposed method relies on a natural transformation of simultaneous confidence bands for distribution functions (DFs) into simultaneous confidence bands for QFs and QE functions. We invert confidence bands for DFs (DF-bands) into confidence bands for QFs (QF-bands) and impose shape and support restrictions. We then take the Minkowski difference of the QF-bands, viewed as sets, to construct confidence bands for the QE functions (QE-bands). This method is generic and applies to a wide collection of model-based estimators of conditional and marginal DFs of discrete, continuous and mixed continuous-discrete outcomes with and without covariates. The only requirement is the existence of a valid method for obtaining simultaneous DF-bands, which is readily available under general sampling conditions for cross-section, time series and panel data. This includes the classical settings of the empirical DFs as a special case, but is much more general than that. For instance, in our empirical applications, we analyze QFs and QE functions that are obtained by inverting counterfactual DFs constructed from regression models with covariates.\footnote{A counterfactual DF is the DF of a potential outcome that is not directly observable but can be constructed using a model. In Section (ref), the counterfactual DFs are formed by integrating the distribution of an outcome conditional on a vector of covariates in one population with respect to the marginal distribution of the covariates in a different population.}
Our method can be used to construct three types of bands. First, we show how to invert a DF-band into a QF-band. We prove that there is no loss of coverage by the inversion in that the resulting QF-band covers the entire QF with the same probability as the source DF-band covers the entire DF. Second, we iterate the method to construct simultaneous QF-bands for multiple QFs. Here simultaneity not only means that the bands are uniform---in that they cover the whole function---but also that all the functions are covered by the corresponding bands jointly with the prescribed probability. These bands can be used to test any comparison between the QFs such as that differences or ratios of them are constant. By construction, our simultaneous bands are not conservative whenever the source DF-bands are not conservative. Third, for the leading case of differences of QFs, we construct QE-bands as differences of QF-bands. Our QE-bands can be conservative due to the projection implicit when we take differences of the QF-bands. However, as we discuss below in the literature review, we are not aware of any generic method to construct valid QE-bands for discrete outcomes. We also show how to make the the QF-bands and QE-bands more informative by imposing support restrictions when the outcome is discrete. To implement all types of bands, we provide explicit algorithms based on bootstrap.
One important application of our method concerns the estimation of QFs and QE functions from models with covariates, which can be used for causal inference as we show in our empirical examples. When the outcome is continuous, quantile regression is a convenient model to incorporate covariates. In many interesting applications, however, the outcome is not continuously distributed. This is naturally the case with count data, ordinal data, and discrete duration data, but it also concerns test scores that are functions of a finite number of questions, censored variables, and other mixed discrete-continuous variables. Examples include the number of doctor visits in our first application (see Panel A in Figure (ref)), IQ test scores for children in our second application (see Panels B and C in Figure (ref)), and wages that have mass points at round values and at the minimum wage.
For discrete outcomes, however, the existing (uniform) inference methods for QR (e.g., gutenbrunner1992process, koenker2002inference, angrist2006misspecification, qu2015nonparametric, belloni2017series) break down. In addition, the linearity assumption for the conditional quantiles underlying QR is highly implausible in that case. For instance, the linear Poisson regression model does not have linear conditional quantiles.
The classical models for discrete outcomes such as Poisson, Cox proportional hazard or ordered response models are highly parametrized. These models have the advantage of being parsimonious in terms of parameters and easy to interpret. However, they impose strong homogeneity restrictions on the effects of the covariates. For instance, if a covariate increases the average outcome, then it must increase all the quantiles of the outcome distribution. Moreover, Poisson models imply a restrictive single crossing property on the sign of the estimated probability effects winkelmann06. To avoid these limitations we employ the DR model in our applications. williams1972analysis introduced this model to analyze ordered outcomes without assuming proportional odds. They estimated a different binary logistic regression for each category of the outcome instead of assuming that the slope coefficients are the same for all categories. foresiperacchi95 estimated a sequence of logistic regression to obtain the conditional distribution of excess returns at a fixed number of thresholds. chernozhukov+13inference considered a continuum of binary regressions and showed that the continuum provides a coherent and flexible model for the entire conditional distribution. rothe2013testing proposed specification tests for DR. DR is a comprehensive tool for modeling and estimating the entire conditional distribution of any type of outcome (discrete, continuous, or mixed). It allows the covariates to affect differently the outcome at different points of the distribution and encompasses several classical parametric models as special cases (classical linear regression, Cox proportional hazard, Poisson regression).
The cost of this flexibility is that the DR parameters can be hard to interpret because they do not correspond to QEs. To overcome this problem, we propose to report QEs computed as differences between the QFs of counterfactual distributions estimated by DR in conjunction with confidence bands constructed using our projection method. These one-dimensional functions provide an intuitive summary of the effects of the covariates. We argue that this combination of our generic procedure with the DR model provides a comprehensive and practical approach for estimating QFs and QE functions with discrete data that can be utilized for causal inference. While we focus on DR in this paper, we emphasize that our method also combines well with classical parametric models such as Poisson, Cox proportional hazard and ordered response regression models. To the best of out knowledge, there is no inference method available to construct valid QF-bands and QE-bands even in these simple models. In addition, our method works in conjunction with more recent inference approaches for DFs with potentially discrete data. Examples include frandsen2012rdd, donald2014estimation, hsu2015estimation, and belloni2017program.
We apply our approach to two problems, featuring two common types of discrete outcomes. In the first application, we exploit a large-scale randomized control trial in Oregon to estimate the distributional impact of universal insurance coverage on health care utilization measured by the number of doctor visits. Since this outcome is a count, we estimate the conditional DFs using both Poisson and distribution regressions. Poisson regression clearly underestimates the probability of having zero visits as well as that of having a large number of visits. The more flexible DR finds a positive effect, especially at the upper tail of the distribution. This is an interesting empirical finding in its own right; it complements the mean regression analysis results reported in finkelstein+12. In the second application, we reanalyze the racial test score gap of young children. As shown in Figure (ref), test scores are discrete. We find that while there is very little gap at eight months, a large gap arises at seven years. In addition, looking at the whole distribution, we uncover that the observed racial gap is widening in the upper tail of the distribution of test scores. The increase in the gap can be mostly explained by differences in socio-economic factors affecting the development of the child as captured by observed covariates. These results complement and expand the findings of fryer2013testing for the mean racial test score gap, revealing what happens to the entire distribution.
\paragraph{Literature Review} To the best of our knowledge, this is the first paper that provides asymptotically similar simultaneous QF-bands for discrete outcomes. scheffe1945non were the first to consider empirical quantiles for discrete data. They showed that pointwise confidence intervals obtained by inverting pointwise confidence intervals for the DF based on the empirical DF are still valid but conservative when the outcome is discrete.\footnote{scheffe1945non studied the properties of confidence intervals for quantiles based on order statistics. This way of obtaining confidence intervals is equivalent to inverting confidence intervals for the DF. woodruff1952confidence and francisco1991quantile graphically illustrated and formally defined, respectively, the inversion idea. However, their formal results only apply to continuous outcomes.} frydman2008discrete and larocque2008confidence suggested methods to obtain the exact coverage rate of these confidence intervals. In contrast, our confidence bands for the QFs are uniform in the probability index and not conservative, and can be based on more general estimators of the DF.
Another strand of the literature tried to overcome the discreteness in the data by adding a small random noise to the outcome (also called jittering), see for instance machado2005quantiles and the applications in koenker2002inference and chernozhukov+13inference. ma+11asymptotic considered an alternative definition of quantiles based on linearly interpolated DFs. These strategies restore asymptotic Gaussianity of the empirical QFs and QE functions, at the price of changing the estimand. One might argue that this change is not a serious issue when the number of points in the support of the outcome is large, but we find it more transparent to work directly with the observed discrete outcome. Thus, we keep the focus on the original QE function at the price that our QE-bands might be conservative. However, we find that they lead to informative inferences in two empirical applications and in extensive numerical simulations, where they are not conservative in most cases. Since we are not aware of any alternative generic method to construct asymptotically similar QE-bands of discrete outcomes, we believe this is a useful addition to the statistical toolkit.
Our method can also be applied to obtain QF-bands and QE-bands of continuous outcomes, complementing the existing well-established methods for this type of outcomes. For example, kh15 provided a method to construct exact DF-bands and QF-bands based on the empirical distribution of independent and identically distributed data.\footnote{We refer to kh15 for an excellent review on construction of confidence bands from the empirical distribution function.} Our method has the more modest goal of constructing asymptotic confidence bands, but applies more generally to any estimator of the distribution that obeys a functional central limit theorem. This includes Poisson, Cox proportional hazard and DR-based estimators of the distribution from weakly dependent data. chernozhukov+13inference also used DR as the basis for constructing QF-bands and QE-bands of continuous outcomes. Their construction consists of two steps. First, obtain estimators of the QFs and QE functions from estimators of DFs by inversion. Second, construct QF-bands and QE-bands from the limit distributions of the estimators of the QFs and QE functions derived from the limit distribution of the estimators of the distribution functions via delta method. This construction has the advantage of producing asymptotically similar QE-bands, but breaks down for discrete outcomes because the quantile (left-inverse) mapping is not smooth (Hadamard differentiable), which precludes the application of the delta method in the second step. Our construction of the confidence bands consists of similar steps but the order of the steps is different. First, we construct DF-bands using the limit distribution of the estimators of the DFs. Second, we construct the QF-bands by inversion the DF-bands and take differences to construct the QE-bands. This difference in the order of the steps completely avoids the delta method and is the key to apply our method to discrete outcomes. Hence, we see our method as complementary to chernozhukov+13inference.
Finally, we would like to emphasize that none of the existing methods that we are aware of can be applied to construct valid QF-bands and QE-bands in our two empirical applications where the outcomes are discrete and the estimators of the distribution are model-based.
\paragraph{Outline} The rest of the paper is organized as follows. Section (ref) introduces our generic method to construct QF-bands and QE-bands. Section (ref) provides an explicit algorithm based on bootstrappable estimators for the DFs. Section (ref) presents the two empirical applications. Appendix (ref) shows how to improve the finite sample properties of DF-bands by imposing logical monotonicity or range restrictions. Appendix (ref) provides an additional algorithm to construct confidence bands for single QFs. Appendix (ref) reports the results of a simulation study.
This section contains the main theoretical results of the paper. Our only assumption is the availability of simultaneous confidence bands for DFs. Since the seminal work of kolmogoroff1933, various methods to obtain simultaneous confidence bands have been developed.\footnote{The original Kolmogorov bands are actually conservative for discrete random variables, see kolmogoroff1941confidence. Alternative methods, such as those described in Section (ref), are asymptotically exact.} In Section (ref) we provide an algorithm to construct simultaneous DF-bands that can be applied when the estimators of the DFs are known to be bootstrappable, which is often the case.
Let $\mathcal{Y}$ be a closed subinterval in the extended real number line $ \overline{\mathbb{R}} = \mathbb{R} \cup \{-\infty,+\infty\}$. Let $\mathbb{D}$ denote the set of nondecreasing functions, mapping $ \mathcal{Y}$ to $[0,1]$. A function $F$ is nondecreasing if for all $x,y \in \mathcal{Y}$ such that $x \leq y$, one has $ F(x) \leq F(y)$. We will call the elements of the set $\mathbb{D}$ \textquotedblleft distribution functions", albeit some of them need not be proper DFs. In what follows, we let $ F$ denote some target DF. $F$ could be a conditional DF, a marginal DF, or a counterfactual DF.
In many applications the point estimates $\hat{F}$ and confidence bands $ [L^{\prime },U^{\prime }]$ for the target distribution $F$ do not satisfy logical monotonicity or range restrictions, namely they do not take values in the set $\mathbb{D}$. Appendix (ref) shows that given such an ordered triple $L^{\prime }\leq \hat{F}\leq U^{\prime }$, we can always transform it into another ordered triple $L\leq \check{F}\leq U$ that obeys the logical monotonicity and range restrictions. Such a transformation will generally improve the finite sample properties of the point estimates and confidence bands.
Here we discuss the construction of confidence bands for the left-inverse function of $F$, $F^{\leftarrow}$, which we call “quantile function” of $F$.
The following theorem provides a confidence band $I^{\leftarrow }$ for the QF $F^{\leftarrow }$ based on a generic confidence band $I$ for $ F $.
Proof. Here we adopt the convention $\inf\{\emptyset\} = + \infty$ so that the left-inverse function can be defined as $G^{\leftarrow }(a):=\inf \{y\in \mathcal{Y}:G(y)\geq a\}\wedge \sup \{y\in \mathcal{Y}\}$, which avoids distinguishing the two cases of Definition (ref).
It suffices to show that $U^\leftarrow \le F^\leftarrow \le L^\leftarrow$ if and only if $L \le F \le U$. We first show that $F^\leftarrow \le L^\leftarrow$ if and only if $F \ge L$. For the “if” part, note that for any $a\in \lbrack 0,1]$, since $L(y)\leq F(y)$ for each $y\in \mathcal{Y} $ and $F,L \in \mathbb{D}$,
For the “only if” part, we use that for any $G \in \mathbb{D}$, $G^{\to} \circ G^{\leftarrow} = G$, where $G^{\to}$ denotes the right-inverse of $G^{\leftarrow}$ defined by $$ G^{\to} \circ G^{\leftarrow} (y) := \sup\{a \in [0,1] : G^{\leftarrow}(a) \leq y \} \vee 0, $$ where we use the convention $\sup\{\emptyset\} = - \infty$. Then, for any $y\in \mathcal{Y}$, since $F^{\leftarrow}(a)\le L^{\leftarrow}(a)$ for each $a\in [0,1]$ and $a \mapsto F^{\leftarrow}(a)$ and $a \mapsto L^{\leftarrow}(a)$ are nondecreasing,
Analogously, we can conclude that $F^\leftarrow \ge U^\leftarrow$ if and only if $F \le U$. {\tiny \ensuremath{\blacksquare} }
We can narrow $I^{\leftarrow }$ without affecting its coverage by exploiting the support restriction that the quantiles can only take the values of the underlying random variable. This is relevant when the variable of interest is discrete as in the applications presented in Section (ref). Suppose that $T$ is the support of the random variable with DF $F$. Then it makes sense to exploit the support restriction that $F^{\leftarrow }(a)\in T$ by intersecting the confidence band for $F^{\leftarrow }$ with $ T $. Clearly, this will not affect the coverage properties of the bands.
Figure (ref) illustrates the construction of bands using Theorem (ref) and Corollary (ref). The left panel shows a DF $F:[0,10]\mapsto \lbrack 0,1]$ covered by a DF-band $I=[L,U]$. The middle panel shows that the inverse map $F^{\leftarrow }:[0,1]\mapsto \lbrack 0,10]$ is covered by the inverted band $I^{\leftarrow }=[U^{\leftarrow },L^{\leftarrow }]$. The band $I^{\leftarrow }$ is easy to obtain by rotating and flipping $I$, but does not exploit the fact that the support of the variable with distribution $F$ in this example is the set $ T=\{0,1,\ldots ,10\}$. By intersecting $I^{\leftarrow }$ with $T$ we obtain in the right panel the QF-band $\tilde{I}^{\leftarrow }$ which reflects the support restrictions.
The quantile effect (QE) function $a\mapsto \Delta _{j,m}(a)$ is the difference between the QFs of two random variables with DFs $F_{j}$ and $ F_{m}$ and support sets $T_{j}$ and $T_{m}$, i.e.,
Our next goal is to construct simultaneous confidence bands that jointly cover the DFs, $\left( F_{k}\right) _{k\in \mathcal{K}}$, the corresponding QFs, $\left( F_{k}^{\leftarrow}\right) _{k\in \mathcal{K}}$, and the QE functions, $\left( \Delta _{j,m}\right) _{\left( j,m\right) \in \mathcal{K}^{2}}$, where $\mathcal{K}$ is a finite set. For example, $\mathcal{K}=\{0,1\}$ (treated and control outcome distributions) in our first application and $\mathcal{K}=\{W,B,C\}$ (white, black and counterfactual test score distributions) in our second application.
Specifically, suppose we have the DF-bands $\left( I_{k}\right) _{k\in \mathcal{K}}$, which jointly cover the DFs $\left( F_{k}\right) _{k\in \mathcal{K}}$ with probability at least $p$. For example, we can construct these bands using Theorem (ref) in conjunction with the Bonferroni inequality.\footnote{The joint coverage of two confidence bands with marginal coverage probabilities $\tilde{p}$ is at least $p=2\tilde{p}-1$ by Bonferroni inequality.} Alternatively, the generic Algorithm (ref) presented in Section (ref) provides a construction of a joint confidence band that is not conservative. First we construct the QF-bands $\left( I_{k}^{\leftarrow}\right) _{k\in \mathcal{K}}$, which jointly cover the QFs $ \left( F_{k}^{\leftarrow}\right) _{k\in \mathcal{K}}$ with probability at least $p$ by Theorem (ref). Then we convert these bands to confidence bands for $\left(\Delta _{j,m}\right)_{(j,k)\in \mathcal{K}^{2}}$ by taking the pointwise Minkowski difference $\ominus$ of each of the pairs of the two bands, viewed as sets. Recall that the Minkowski difference between two subsets $V$ and $U$ of a vector space is $V\ominus U:=\{v-u:v\in V,u\in U\}$. We note that if $V$ and $U$ are intervals, $[v_{1},v_{2}]$ and $ [u_{1},u_{2}]$, then
In words, the upper-end of the interval for the difference $v-u$ is the difference between the upper-end of the interval for $v$ and the lower-end of the interval for $u$. Symmetrically, the lower-end of the interval for the difference $v-u$ is the difference between the lower-end of the interval for $v$ and the upper-end of the interval for $u$. This greatly simplifies the practical computation of the bands.
Proof. The results follows from the definition of the Minkowski difference and because the event $\cap _{k\in \mathcal{K}}\{F_{k}\in I_k\}$ is equivalent to the event $\cap _{k\in \mathcal{K}}\{F_{k}^{\leftarrow }\in I_{k}^{\leftarrow }\}$ by Theorem (ref), which implies the event $\cap _{\left( j,m\right) \in \mathcal{K}^{2}}\{F_{j}^{\leftarrow }-F_{m}^{\leftarrow }\in I_{\Delta \left( j,m\right) }^{\leftarrow }\}$. {\tiny \ensuremath{\blacksquare} }
Theorem (ref) shows that simultaneous QF-bands can be obtained by inverting simultaneous DF-bands, and QE-bands by taking the Minkowski difference between the two simulataneous QF-bands for the corresponding QFs. As in Theorem (ref), we can narrow the band $I_{\Delta }^{\leftarrow }$ without affecting coverage by imposing support restrictions as demonstrated in Corollary (ref).
Figure (ref) illustrates the construction of QE-bands using Theorem (ref) and Corollary (ref). The left panel shows the bands $I_{0}^{\leftarrow }$ and $I_{1}^{\leftarrow }$ for the QFs $F_{0}^{\leftarrow }$ and $ F_{1}^{\leftarrow }$. The middle panel shows the band $I_{\Delta \left( 1,0\right) }$ for the QE function $\Delta _{1,0}=F_{1}^{\leftarrow }-F_{0}^{\leftarrow }$, obtained by taking the Minkowski difference of $ I_{1}^{\leftarrow }$ and $I_{0}^{\leftarrow }$. The right panel shows the confidence band $\tilde{I}_{\Delta \left( 1,0\right) }$ for the QE function $ \Delta _{1,0}$ resulting from imposing the support restrictions. As the Theorem (ref) proves, the QE function $\Delta _{1,0}$ is covered by the QE-band $I_{\Delta \left( 1,0\right) }$.
In Section (ref) we assumed the existence of simultaneous DF-bands. Here we describe an algorithm that is shown to provide asymptotically valid simultaneous bands for any bootstrappable estimator of the DFs. Many commonly used estimators of the DF are bootstrappable under suitable conditions. For example, chernozhukov+13inference give conditions for bootstrap consistency for the DR-based estimators that we use in the empirical applications. Maximum likelihood estimators, such as the Poisson regression that we use as a benchmark in the first application, are also bootstrappable under weak differentiability conditions, see arcones1992bootstrap. We note that if the data are discrete, these existing results still yield validity of bootstrapping the DF estimator to construct DF-bands, but they do not justify the validity of bootstrapping the QF and QE estimators to construct QF-bands and QE-bands. The reason is that the delta-method breaks down because the left-inverse mapping is no longer (Hadamard) differentiable.
Algorithm (ref) provides simultaneous confidence bands that asymptotically jointly cover the DFs $\left( F_{k}\right) _{k\in \mathcal{K} } $, the corresponding QFs $\left( F^{\leftarrow}_{k}\right) _{k\in \mathcal{ K}}$, and the QE functions $F_{j}^{\leftarrow }-F_{k}^{\leftarrow }$ for all $(j,k)\in \mathcal{K}^{2},$ with probability $p$. In practice, we estimate the DFs on a grid of points. Let $T$ be a finite subset of $\mathcal{Y}$. For the DF of a discrete random variable $Y$ with finite support, we can choose $T$ as the support of $Y$. Otherwise, we can set $T$ as a grid of values covering the region of interest of the support of $Y$.
In step (1) we bootstrap jointly all the estimators of the DFs. In our applications it is important to obtain jointly the bootstrap draws of these estimators because they are not independent. There are multiple ways to obtain the bootstrap draws of $\hat{F}$. A generic resampling procedure is the exchangeable bootstrap pw:93,vdV-W, which recomputes $\hat{F}$ using sampling weights drawn independently from the data. This procedure incorporates many popular bootstrap schemes as special cases by a suitable choice of the distribution of the weights. For example, the empirical bootstrap corresponds to multinomial weights, and the weighted or Bayesian bootstrap corresponds to standard exponential weights. Exchangeable bootstrap can also accommodate dependence or clustering in the data by drawing the same weight for all the observations that belong to the same cluster sc:97,cheng2013cluster. For example, in the application of Section (ref) we draw the same weights for all the individuals of the same household.
In the second step we estimate pointwise standard errors. We use the bootstrap rescaled interquartile range because it is more robust than the bootstrap standard deviation in that it requires weaker conditions for consistency chernozhukov+13inference. In the third step, we compute, for each bootstrap draw, the weighted recentered Kolmogorov-Smirnov maximal $t$-statistic over all distributions $F_{k}\left( y\right) $ with $k\in \mathcal{K}$ and $y\in T$. The maximum over $k\in \mathcal{K}$ ensures joint coverage of all the DFs. Then we take the $p$ -quantile of the bootstrap Kolmogorov-Smirnov statistics. This allows us, in the fourth step, to construct preliminary DF-bands that jointly cover all the DFs $\left( F_{k}\right) _{k\in \mathcal{K}}$ with probability $p$. We improve these bands by imposing the shape restrictions.
In the fifth step we invert the DF-bands to obtain QF-bands, as justified by Theorem (ref). In the last step we obtain the QE-bands by taking Minkowski differences of the QF-bands, as justified by Theorem (ref). If needed, we can impose the support conditions in the last two steps.
The following corollary of Theorem (ref) provides theoretical justification for Algorithm (ref). To state the result, let $ \ell ^{\infty }(\mathcal{Y})$ denote the metric space of bounded functions from $\mathcal{Y}$ to $\mathbb{R}$ equipped with the sup-norm and $|\mathcal{K}|$ denote the cardinality of the set $\mathcal{K}$.
Proof. Lemma SA.1 of chernozhukov+13inference implies that $ \lim_{n\rightarrow \infty }{\mathrm{P}} (\cap _{k\in \mathcal{K}}\{F_{k}\in \lbrack L_{k}^{\prime },U_{k}^{\prime }]\})=p$. The result then follows from Lemma (ref), Theorems (ref) and (ref), and Corollaries (ref) and (ref).{\tiny \ensuremath{\blacksquare} }
Algorithm (ref) provides confidence bands that jointly cover the DFs, the QFs, and the QE functions. If one is only interested in one single QF, say $F_1^{\leftarrow}$, the corresponding QF-band obtained from Algorithm (ref) can be conservative. This is because we compute the maximal $t$-statistic over all distributions $ \left(F_{k}\right)_{k\in \mathcal{K}} $ to ensure joint coverage, which is not required if one is only interested in $F_1^{\leftarrow}$. Appendix (ref) provides a bootstrap algorithm that yields an asymptotically similar QF-band for a single QF.
In this section we apply our approach to two data sets, corresponding to two common types of discrete outcomes.\footnote{The data and code in R R18 for the empirical analysis is available at \href{https://github.com/bmelly/discreteQ}{https://github.com/bmelly/discreteQ}.} In both cases we use the distribution regression model and obtain QE as differences between counterfactual distributions. For this reason, we first introduce the specific methods and then present both empirical illustrations.
In the absence of covariates, the empirical DF is a minimal sufficient statistic for a non-parametric marginal DF. Distribution regression (DR) generalizes this concept to a conditional DF like OLS generalizes the univariate mean to the conditional mean function. The key, simple observation underlying DR is that the conditional distribution of the outcome $Y$ given the covariates $X$ at a point $y$ can be expressed as $F_{Y\mid X}(y\mid x)={\mathrm{E}}[1\{Y\leq y\}\mid X=x]$. Accordingly, we can construct a collection of binary response variables, which record the events that the outcome $Y$ falls bellow a set of thresholds $T$, i.e.,
and use a binary regression model for each variable in this collection. This yields the DR model:
where $\Lambda _{y}(\cdot )$ is a known link function which is allowed to change with the threshold level $y$; $B(x)$ is a vector of transformations of $x$ with good approximating properties such as polynomials, B-splines, and interactions; and $\beta \left( y\right) $ is an unknown vector of parameters. Knowledge of the function $y\mapsto \beta (y)$ implies knowledge of the distribution of $Y$ conditional on $X$. The DR model is flexible in the sense that, for any given link function, we can approximate the conditional DF arbitrarily well by using a rich enough set of transformations of the original covariates $B(x)$. In the extreme case when $ X$ is discrete and $B(x)$ is fully saturated, the estimated conditional distribution is numerically equal to the empirical DF in each cell of $X$ for any monotonic link function. When $B(x)$ is not fully saturated, one can choose a DF such as the normal or logistic as the link function to guarantee that the model probabilities lie between 0 and 1.
DR nests a variety of classical models such as the Normal regression, the Cox proportional hazard, ordered logit, ordered probit, Poisson regression, as well as other generalized linear models. Example (ref) shows the inclusion of the Poisson regression model which we use as a benchmark in our first empirical application. In what follows we set $B(x)=x$ to lighten the notation without loss of generality.
Assume that we have a sample $\{(Y_{i},X_{i}):i=1,...,n\}$ of $(Y,X)$. The DR estimator of the conditional distribution is
where
williams1972analysis introduced DR in the context of ordered outcomes. foresiperacchi95 applied this method to estimate the conditional distribution of excess return evaluated at a finite number of points. chernozhukov+13inference extended williams1972analysis's definition to arbitrary outcomes and established functional central limit theorems and bootstrap validity results for DR as an estimator of the whole conditional distribution. One of the main advantages of DR is that it not only accommodates continuous but also discrete and mixed discrete continuous outcomes very naturally.
We show how to utilize DR for causal inference in two empirical applications. In both applications there are two groups: the treated and control units in the first application, and the black and white children in the second application. We use DR to model and estimate the conditional distribution of the outcome in each group at each value of the covariates, that we denote by $F_{Y_{0}|X_{0}}(y\mid x)$ and $F_{Y_{1}|X_{1}}(y\mid x)$. The difference between these two high-dimensional DFs is, however, difficult to convey. Instead, we integrate these conditional distributions with respect to observed covariate distributions and compare the resulting marginal distributions.
For instance, in the first application, the marginal distribution
where $F_{X}$ is the distribution of $X$ in the entire population including the treated and control units, represents the distribution of a potential outcome. When $k=1$, $F_{{\langle 1\rangle }}$ is the outcome distribution that would be observed if every units were treated, and when $k=0$, $F_{{\langle 0\rangle }}$ is the outcome distribution if every units were not treated. These two distributions are called counterfactual, since they do not arise as distributions from any observable population. They nevertheless have a causal interpretation as distributions of potential outcomes when the treatment is randomized conditionally on the control variables $X$.
Let $\hat{F}_{Y_{k}|X_{k}}$ denote the DR estimator of $F_{Y_{k}|X_{k}}$, $ k\in \{0,1\}$. We estimate $F_{{\langle k\rangle }}$ by the plugging-in rule, namely integrating $\hat{F}_{Y_{k}|X_{k}}$ with respect to the empirical distribution of $X$ for treated and control units. For $k\in \left\{ 0,1\right\} $,
We then report the empirical QE function:
chernozhukov+13inference derived joint functional central limit theorems for $(\hat{F}_{{\langle 0\rangle }},\hat{F}_{{\langle 1\rangle }})$ and established bootstrap validity. We can thus use the algorithms in Section (ref) to construct asymptotically valid simultaneous confidence bands for the counterfactual QFs $(F_{{\langle 1\rangle }}^{\leftarrow },F_{{\langle 0\rangle }}^{\leftarrow })$ and the QE function $\Delta =F_{{\langle 1\rangle }}^{\leftarrow }-F_{{\langle 0\rangle }}^{\leftarrow }$.
Our first application illustrates the construction of confidence bands using data from the Oregon health insurance experiment. In 2008, the state of Oregon initiated a limited expansion of its Medicaid program for uninsured low-income adults by offering insurance coverage to the lottery winners from a waiting list of 90,000 people (see \href{http://www.nber.org/oregon/} {www.nber.org/oregon} for details). This experiment constitutes a unique opportunity to study the impact of insurance by means of a large-scale randomized controlled trial finkelstein+12,baicker+13,baicker+14,taubman+14.
We investigate the impact of insurance coverage on health care utilization as analyzed in finkelstein+12 using a publicly available dataset finkelstein+12data. The data are available via: \href{http://www.nber.org/oregon/4.data.html}{ http://www.nber.org/oregon/4.data.html}. Detailed information about the dataset and descriptive statistics are available in finkelstein+12 and the corresponding online appendix. We focus on one count outcome $Y$: the number of outpatient visits in the last six months, which was elicited via a large mail survey. After excluding individuals with missing information in any of the variables used in the analysis, the resulting sample consists of 23,441 observations. The top histogram in Figure (ref) illustrates the discrete nature of our dependent variable. Almost 40% of the outcomes are zeros, more than 90% of the mass is concentrated between zero and five, but a few people have a greater number of visits.
finkelstein+12 find a positive effect of winning the lottery on the number of outpatient visits.\footnote{They label these effects intention-to-treat (ITT) effects and also report local average treatment effects (LATE) estimated using IV regressions. In this section, we focus on ITT effects.} Their results are based on ordinary least squares (OLS) regressions, where the covariates $X$ include household size, indicators for the survey wave, and interactions of the household size indicators and the survey wave. Although individuals were chosen randomly, these covariates are included as controls because the entire household for any selected individual became eligible to apply for insurance and the fraction of treated individuals varies across survey waves. We complement their findings by looking at the whole distribution of the number outpatient visits. We first estimate the conditional outcome distributions separately for the lottery winners and losers via Poisson regression and DR. For DR, we use the exponentiated incomplete gamma link in (ref) such that DR nests the Poisson regression as an exact special case. As explained in Section (ref), we integrate the conditional outcome distributions with respect to the covariate distribution for both lottery winners and losers to obtain estimates of the counterfactual distributions $ F_{\left\langle 1\right\rangle }$ and $F_{\left\langle 0\right\rangle }$.
The top panel of Figure (ref) displays the DFs $ \hat{F}_{\left\langle 1\right\rangle }$ and $\hat{F}_{\left\langle 0\right\rangle }$ estimated by the Poisson regression and DR. The corresponding QFs $\hat{F}_{\left\langle 1\right\rangle }^{\leftarrow }$ and $F_{\left\langle 0\right\rangle }^{\leftarrow }$ are displayed in both middle panels. Finally, the estimated QE functions, $\hat{F}_{{\langle 1\rangle }}^{\leftarrow }-\hat{F}_{{\langle 0\rangle }}^{\leftarrow }$, are plotted in the bottom panels. In all cases, the figure also shows 95% simultaneous confidence bands, constructed using Algorithm (ref) with $B=1,000$ Bayesian bootstrap draws that take into account the possible clustering of the observations at the household level. Reflecting the discrete nature of our outcome variables, we impose the support restrictions $T_0 = T_1 =\{0,1,\ldots \}$.
A comparison between the Poisson and DR results reveals striking differences. The Poisson model predicts a much lower mass at zero and a much thinner upper tail of the distribution for both groups. Indeed, these differences are statistically significant as the Poisson and DR simultaneous DF-bands and QF-bands do not overlap for a large part of the support. A formal test rejects the equality of these distributions with a p-value below $0.001$. Since the DR model with exponentiated incomplete gamma link nests the Poisson model, we conclude that the Poisson model is rejected by the data. For this reason, we focus the discussion on the DR results.
The QE-band do not fully cover the zero-line and thus we can reject the null hypothesis that winning the lottery has no effect on the number of outpatient visits. We can also reject the hypothesis that $F_{\left\langle 0\right\rangle }$ first-order stochastically dominates $F_{\left\langle 1\right\rangle }$ because the band for $F_{\left\langle 0\right\rangle }^{\leftarrow}$ is strictly below the band for $ F_{\left\langle 1\right\rangle }^{\leftarrow}$ at some probability indexes. However, we cannot reject the opposite hypothesis. In other words, at no quantile index the confidence band contains strictly negative effects while at some probability indexes it contains strictly positive effects.
Health economists distinguish between the treatment effect on the extensive (whether to see a doctor) and intensive (the number of visits given at least one) margins. The first effect is easy to estimate: the probability of not seeing a doctor decreased significantly from 43% to 37% with the treatment. The effect on the intensive margin is more difficult to gauge because we do not observe both potential outcomes for any individual. If we assume that the individuals induced to see a doctor by the insurance coverage are not seriously sick and visit the doctor only once, then the effect on the intensive margin can also be seen in Figure (ref): the effect from 0 to 1 visit represents the effect on the extensive margin and the effect on the rest of the distribution represents the effect on the intensive margin. Both effects are statistically significant. We note in particular that the quantile differences do not vanish at the top of the distribution.
The assumption made to justify this interpretation may be too strong and lead to an overestimation of the effect on the intensive margin. For instance, the doctor may find a serious problem and schedule other visits. Following zhang2003estimation and angrist2006long, we can bound the effect on the intensive margin from below by assuming that patients who see a doctor anyway visit their doctor at least as often as patients who see a doctor only if insured. Under this weaker assumption, the effect on the intensive margin is bounded from below by the QE function obtained by keeping only observations with at least one visit. We also find a positive treatment effect with this method, which reinforces the evidence of a positive effect among the existing users.
As a second application, we reanalyze the racial IQ test score gap examined in fryer2013testing. We use data from the US Collaborative Perinatal Project (CPP). These data contain information on children from 30,002 women who gave birth in 12 medical centers between 1959 and 1965. Our main outcomes of interest are the standardized test scores at the ages of eight months (Bayley Scale of Infant Development) and seven years (both Stanford-Binet and Wechsler Intelligence Test). In addition to the test score measures, the dataset contains a rich set of background characteristics for the children, $X$, including information on age, gender, region, socioeconomic status, home environment, prenatal conditions, and interviewer fixed effects. fryer2013testing provide a comprehensive description of the dataset and extensive descriptive statistics.
A key feature of the test scores is the discrete nature of their distribution. We observe only 76 and 128 different values for the standardized test scores at the ages of eight months and seven years, respectively. The middle and bottom panels of Figure (ref) present the corresponding histograms. Note that each bar corresponds to exactly one value. For instance, at eight months, almost 12% of the observations have exactly the same score and 60% of the observations have one of the most frequent six values. This is a common feature of test scores, which are necessarily discrete because they are based on a finite number of questions.
To gain a better understanding of the causes of the observed black-white test score gap, we provide a distributional decomposition into explained and unexplained parts by observable background characteristics. Let $F_{{\langle W|W\rangle }}$ and $F_{{\langle B|B\rangle }}$ represent the observed test score DFs for white and black children, and $F_{\langle W|B\rangle }$ represents the counterfactual DF of test scores that would have prevailed for white children had they had the distribution of background characteristics of black children, $F_{X_{B}}$, namely,
With this counterfactual test score distribution it is possible to decompose the quantiles of the observed black-white test score gap into
where the first term in brackets corresponds to the composition effect due to differences in observable background characteristics and the second term is the unexplained difference.
We estimate $F_{\langle W\mid W\rangle }$ and $F_{\langle B\mid B\rangle }$ by the empirical test score distributions for white and black children, respectively. We estimate the counterfactual distribution $F_{\langle W\mid B\rangle }$ by the sample analog of (ref) replacing $ F_{Y_{W}|X_{W}}$ by the DR estimator for white children, and $F_{X_B}$ by the empirical distribution of $X$ for black children. We use the logistic link function for the DR, but the results using the linear link function or the normal link function are similar.
Figures (ref) and (ref) report the results for the eight months and seven years outcomes, respectively. The first panels show the observed and counterfactual QFs, $F_{\langle W\mid W\rangle }^{\leftarrow }$, $F_{\langle B\mid B\rangle }^{\leftarrow }$ and $F_{\langle W\mid B\rangle }^{\leftarrow }$. The second panels show the difference between the observed QFs, $F_{\langle W\mid W\rangle }^{\leftarrow }-F_{\langle B\mid B\rangle }^{\leftarrow }$. The third and fourth panels decompose these observed differences into the composition effect ($F_{\langle W\mid W\rangle }^{\leftarrow }-F_{\langle W\mid B\rangle }^{\leftarrow }$) and the unexplained component ($F_{\langle W\mid B\rangle }^{\leftarrow }-F_{\langle B\mid B\rangle }^{\leftarrow }$). The point estimates are shown with their respective 95% simultaneous confidence bands constructed using Algorithm (ref) with $B=1,000$ Bayesian bootstrap draws. The bands impose the restrictions that the supports of the test scores correspond to the observed values in the sample.
For eight-month-old children, we find very small differences between the test score distributions of black and white children. The black-white gap is positive at the lower tail and is mainly due to unobserved characteristics. While these effects are statistically significant, they are so small in magnitude that they should not worry any policy maker. The composition effect is very small, probably simply because there was no difference to explain to begin with.
The results are completely different for seven-year-old children. We find a large and statistically significant positive raw black-white gap. A formal test based on the uniform bands rejects the null hypothesis of a zero or a negative racial test score gap at all quantiles. The estimated QE function is increasing in the probability index ranging from below $0.6$ standard deviation units at the lower tail up to over one standard deviation unit at the upper tail of the distribution. The quantile differences at the tails substantially differ from the mean difference of $ 0.85$ standard deviation units reported in fryer2013testing. In fact, we can formally reject the null hypothesis of a constant raw test score gap across the distribution because we can not draw a horizontal line at any value of the difference of test scores, which is covered by the confidence band of the QE function at all probability indexes.
Our decomposition analysis shows that about two third of this gap can be explained by differences in the distribution of observable characteristics. Nevertheless, the remaining unexplained difference is significant, both in economic and in statistical terms. Looking at the QE function, we can see that there is substantial effect heterogeneneity along the distribution. Interestingly, the increase in the test score gap at the upper quantiles can be fully explained by differences in background characteristics between black and white children. The resulting unexplained difference is maximized in the center of the distribution. Finally, our simultaneous confidence bands allow for testing several interesting hypothesis' about the whole QE function. For instance, we can reject the null hypothesis that the composition effect and the unexplained difference are zero, negative, or constant at all quantiles but we cannot reject that they are positive everywhere.