Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
102,659 characters · 13 sections · 131 citation commands
{ {Abstract:}}{ \ In a unified framework, we provide estimators and confidence bands for a variety of treatment effects when the outcome of interest, typically a duration, is subjected to right censoring. Our methodology accommodates average, distributional, and quantile treatment effects under different identifying assumptions including unconfoundedness, local treatment effects, and nonlinear differences-in-differences. The proposed estimators are easy to implement, have close-form representation, are fully data-driven upon estimation of nuisance parameters, and do not rely on parametric distributional assumptions, shape restrictions, or on restricting the potential treatment effect heterogeneity across different subpopulations. These treatment effects results are obtained as a consequence of more general results on two-step Kaplan-Meier estimators that are of independent interest: we provide conditions for applying $(i)$ uniform law of large numbers, $(ii)$ functional central limit theorems, and $(iii)$ we prove the validity of the ordinary nonparametric bootstrap in a two-step estimation procedure where the outcome of interest may be randomly censored. }
{ Keywords: Kaplan-Meier Integrals; Survival Analysis; Policy Evaluation; Treatment effects; Duration models.} \pagenumbering{arabic}
Assessing whether a policy has any effect on a particular outcome has been one of the main concerns in empirical research. As summarized in Heckman2007 and Imbens2009, the focus of the policy evaluation literature has been mainly confined to situations where the realized outcome of interest is completely observed for the treated and the control groups. However, when the outcome variable is subjected to censoring, such inference procedures may provide misleading conclusions on the effect of the proposed policy. Important empirical examples of such a setting include the evaluation of labor market programs on the length of unemployment, of correctional programs on recidivism of criminal activities, and of clinical therapy on the survival time.
The main objective and contribution of this paper is to provide a unified framework to derive estimation and inference procedures for policy evaluation when the outcome of interest, typically a duration, is subjected to right-censoring. Our methodology accommodates average, distributional, and quantile treatment effects in a variety of identifying assumptions such as selection on observable, cf. Hirano2003, Firpo2007, and Donald2013; access to a binary instrumental variable, cf. Imbens1994, Abadie2002a, Abadie2003 and Frolich2013; and access to repeated observations over time, cf. Athey2006. To the best of our knowledge, this paper is the first to propose such broad policy evaluation tools for right-censored outcomes without relying on parametric assumptions or shape restrictions.
Our policy evaluation results build on the fact that many treatment effect measures commonly used can be written as (smooth) functions of moment equations of the type
where $Y$ is the outcome of interest, $T$ is the treatment status, and $X$ is a vector of covariates; $\varphi _{z,h_{0}}$ is some integrable function, potentially indexed by $z$, and by (infinite dimensional) nuisance parameters $h_{0}$; and $F$ is the joint cumulative distribution function (CDF). Therefore, our policy evaluation problem can be translated into the more general task of estimating moments of the type of ((ref)).
In the presence of right-censored outcomes, the main challenge in estimating ((ref)) is the fact that $Y$ is not always observed. That is, instead of observing a random sample $\left\{ Y_{i},X_{i},T_{i}\right\} _{i=1}^{n}$ of $\left( Y,X,T\right) $ as in the \textquotedblleft complete data\textquotedblright\ setup, one observes $iid$ copies $\left\{ Q_{i},\delta _{i},X_{i},T_{i}\right\} _{i=1}^{n}$ of $(Q,\delta ,X,T)$, where $Q=\min \left( Y,C\right) $, $\delta =1\left\{ Y\leq C\right\} $, and $ C$ is a censoring random variable. Right-censoring is a common feature of duration outcomes, and may arise for different reasons, such as the end of a follow-up, or drop out. Thus, when estimating ((ref)), one must take into account this data limitation. In fact, ignoring the censoring problem or restricting the analysis to uncensored observations leads to biased and inconsistent estimators for ((ref)).
To overcome such problems we propose the following two-step procedure. In the first step, one consistently estimate $h_{0}$ using parametric, semiparametric or nonparametric methods, and denote such generic estimator by $\hat{h}_{n}$. In the second step, one plugs $\hat{h}_{n}$ into ((ref)), and then replace $F$ with $\hat{F}_{n}^{km},$ where $\hat{F} _{n}^{km}$ is a nonparametric multivariate extension of the time-honored Kaplan1958 product-limit estimator that naturally address the censoring issue\footnote{ Following VanNoorden2014, Kaplan1958 is, based on Thomson Reuters' Web of Science as 7 October 2014, the most cited paper in statistics, and the $11^{th}$ most cited paper in all sciences, with 38,600 citations.}. By combining these two steps, we propose to estimate ((ref) ) by
We label the estimator in ((ref)) as the two-step Kaplan-Meier (2SKM) estimator.
The 2SKM estimator inherits many attractive features. First, it is very easy to implement, has a simple close-form representation, is fully data-driven upon estimation of the nuisance parameters $h_{0}$, and does not depend on parametric functional form assumptions on the joint distribution $Y$, $X$ and $T$. This last property is in sharp contrast with Cox1972 proportional hazard models, or Buckley1979 accelerated failure time models, two of the most popular duration models in the literature. Second, in the absence of censoring, ((ref)) reduces to the empirical analogue of ((ref)),
implying that one can interpret our proposal as a natural generalization of standard two-step estimation procedures such as Pakes1989 and Chen2003 to situations in which the outcome is censored.
This article contains two sets of new theoretical results on Kaplan-Meier integrals ((ref)). First, we present a set of sufficient conditions under which the 2SKM estimator is uniformly consistent, and converges weakly to a tight Gaussian process. Furthermore, since the limiting variance function may depend on the data generating process in rather complicated forms, we propose and prove the validity of the ordinary nonparametric bootstrap, which can be used to construct asymptotic valid confidence bands.
The second set of results deals with estimation and inference under primitive conditions in three leading policy evaluation methods. Specifically, we prove that the high-level conditions to establish the functional central limit theorem and validity of bootstrap hold for average, distributional, and quantile treatment effects under the unconfoundedness, local treatment effects, and nonlinear differences-in-differences setups.
This article contributes to the literature on treatment effects with censored data. Contrary to Ham1996, Eberwein1997, Hubbard2000, Anstrom2001, Abbring2003, and Vanderlaan2003, our methodology does not rely on parametric models, separability or proportionality restrictions. In contrast with Frandsen2014, our proposal can easily accommodate covariates, does not rely on the potentially restrictive condition that the censoring variable is always observed, and does not require choosing truncation parameters. Furthermore, it is important to emphasize that, in contrast to all the aforementioned proposals, our main results are generic, can be used under a variety of identification conditions, and apply to any functional of interest that satisfy the relatively weak conditions.
We also contribute to the literature on Kaplan-Meier integrals, cf. Stute1993a, Stute1993, Stute1995a,Stute1996a, Stute1996b, Stute1999 , Wang1999, Akritas2000a, and Sellero2005. The available results in this literature are not directly applicable to our two-step framework in which the integrand is indexed by unknown, possibly infinite-dimensional nuisance parameters that have to be estimated beforehand. Thus, our results for 2SKM estimators complement and extend those available in the literature.
In order to achieve the aforementioned results, one must bear in mind that although we do not restrict the dependence between $Y$, $X$ and $T$, our estimation and inference procedure relies on the maintained assumptions that $(a)$ conditionally on the treatment status $T$, the outcome of interest ($Y$ ) is independent of the censoring variable $\left( C\right) $, and $(b)$ conditionally on $Y$ and $T$, the vector of available covariates $(X)$ does not provide any additional information if censoring will take place. These assumptions are standard in censoring models, and nest the setups considered by, e.g. Powell1986, Honore2002, Hong2003a, Lee2005, Blundell2007, and Frandsen2014. Nonetheless, these maintained assumptions are stronger than assuming that, conditionally on $X$ and $T$, $Y$ is independent of $C$, and may be violated in some applications. Thus, as a form of specification test for 2SKM estimators, it may be desirable to test our maintained assumptions on the censoring mechanism. In the supplemental appendix we show that such a task is feasible, and discuss how one can implement a likelihood ratio type test for the assumptions. Constructing such a nonparametric test is only feasible at the cost of introducing additional smoothness and support restrictions on the underlying data generating process, on top of making use of tuning parameters such as bandwidths.
The rest of the paper is organized as follows. In Section (ref) we motivate the problem at hands by showing that different treatment effects parameters can be written as smooth functions of moment equations of the type of ((ref)). In Section (ref) we discuss the identification and estimation of generic moments of the type of ((ref)) when the outcome of interest is censored. Section (ref) discusses some sufficient conditions to derive (uniform) law of large numbers and (functional) central limit theorems for the proposed 2SKM estimators. We also discuss some regularity conditions for establishing the validity of the ordinary nonparametric bootstrap for censored data. In Section (ref) we use our general results on 2SKM estimators to establish the asymptotic properties of the treatment effect parameters discussed in Section (ref) in the presence of censored outcomes. In Section (ref) we conduct a small scale Monte Carlo exercise to illustrate the finite sample properties of our proposal. Section (ref) concludes with a summary of the main results. A supplemental appendix includes: $\left( i\right) $ the proofs of the results herein; $\left( ii\right) $ a discussion on how one can test the maintained assumptions on the censoring mechanism; and $\left( iii\right) $ the complete set of Monte Carlo results.
In this section, we show that, under different identification scenarios, one can use (smooth) functions of moment equations of the type of ((ref)) to characterize the average, distributional, and quantile treatment effects. We particularly focus on three popular identification setups: $(i)$ selection on observables, $\left( ii\right) $ access to a binary instrumental variable, and $(iii)$ access to repeated observations over time.
We use the following notation. Let $Y_{0}$ and $Y_{1}$ be the potential individual outcomes under the control and treatment group, respectively. Upon inflow, an individual is assigned to a treatment $(T=1)$ or to a control $\left( T=0\right) $ group. The realized outcome of interest is $ Y\equiv $ $TY_{1}+(1-T)Y_{0}$, and $X$ is a $k$-dimensional vector of pre-treatment observable covariates. Let $ \protect\mathpalette{\protect\independentT}{\perp} $ mean \textquotedblleft is independent\textquotedblright , and $\mathcal{Y}$ denote the support of the random variable $Y$.
Example 2.1 (Unconfoundedness setup). One of the most popular identification strategies in policy evaluation is to assume that selection into treatment is solely based on observable characteristics, i.e. $\left( Y_{0},Y_{1}\right) \protect\mathpalette{\protect\independentT}{\perp} T|X$ $a.s\mathbf{.}$. This is the so called unconfoundedness setup. Here, popular parameters of interest are the overall average, distributional, and quantile treatment effects
respectively, where for $t\in \left\{ 0,1\right\} $, $q_{Y_{t}}\left( \tau \right) \equiv \inf \left\{ y:\mathbb{P}\left( Y_{t}\leq y\right) \geq \tau \right\} .$
As shown by Rosenbaum1983, provided that individuals with the same $X$ values have a positive probability of being both at the treatment and the control group, the aforementioned treatment effect parameters are identified by
where $p\left( X\right) \equiv \mathbb{P}\left( T=1|X\right) $ is the propensity score, i.e. the probability of selection into treatment,
and, for $t\in \left\{ 0,1\right\} $, $F_{Y_{t}}^{-1}\left( \tau \right) \equiv \inf \left\{ y:F_{Y_{t}}\left( y\right) \geq \tau \right\} $\footnote{ The average, distributional, and quantile treatment effects on treated subpopulation can also be identified using a similar strategy.}.
Notice that ((ref)) and ((ref)) are simple differences of moment equations of the type of ((ref)), where, in both cases, $p\left( \cdot \right) $ plays the role of the unknown nuisance parameter $h_{0}$, and $ y\in \mathcal{W}\subseteq \mathcal{Y}$ plays the role of $z$ in ((ref)). Although one cannot write the quantile treatment effects ((ref)) as moment equations of the type of ((ref)), its identification follows from the one-to-one relationship between the quantile function $ F_{Y_{t}}^{-1}\left( \tau \right) $ and the CDF $F_{Y_{t}}\left( y\right) $, $t\in \left\{ 0,1\right\} $\footnote{ For estimation and inference purposes, when one is interested in quantile treatment effects, we will impose additional continuity restrictions on the DGP, such that the functional delta method can be applied, see e.g. Chapter 3.9 of VanderVaart1996. We defer discussion of these assumptions to Section (ref).}. Thus, the treatment effect measures ((ref))-((ref)) fit well into our framework.
Example 2.2 (Local treatment effects setup) In many circumstances, the assumption that the selection into treatment is based only on observable characteristics may be unrealistic. Imbens1994 and Angrist1996 point out that when this is the case and a binary instrument ($Z)$ for the selection into treatment is available, one can only nonparametrically identify treatment effect measures for the subpopulation of compliers, that is, individuals who comply with their actual assignment of treatment, and would have complied with the alternative assignment. Such policy evaluation framework is know as the local treatment effect (LTE) setup.
By following similar arguments as Rosenbaum1983, Abadie2003 and Frolich2013 show that, under some regularity conditions to be discussed in Section (ref), the average, distributional and quantile treatment effects for the subpopulation of compliers,
respectively, can be identified by
where, for $t\in \left\{ 0,1\right\} $,
and
$F_{Y_{t}^{c}}^{-1}\left( \tau \right) =\inf \left\{ y:F_{Y_{t}^{c}}\left( y\right) \geq \tau \right\} $, and $e\left( X\right) \equiv \mathbb{P} (Z=1|X) $.
From ((ref)) and ((ref)), one can see that $\mathbb{E}\left[ Y_{t}^{c} \right] $ and $F_{Y_{t}^{c}}\left( y\right) $ are scaled differences of moment equations of the type of ((ref)). Analogously to the unconfoundedness setup, $e\left( \cdot \right) $ plays the role of $h_{0}$, and $y\in \mathcal{W}\subseteq \mathcal{Y},$ and $\tau \in \left( 0,1\right) $ play the role of $z$ in ((ref)), and ((ref)), respectively. Although identification of the aforementioned treatment effects involve $ \kappa _{t}\left( e\right) $, for estimation and inference purpose, we can treat $\kappa _{t}\left( e\right) $ as a known function, cf. Abadie2003 and Frolich2013. Thus, as in Example 2.1, the treatment effect measures ((ref))-((ref)) fit well into our framework.
Example 2.3 (Differences-in-Differences setup) This example is concerned with treatment effects when one has access to repeated observations over time, the so called differences-in-differences (DID) approach, cf. Angrist1999. In its basic form, a control group is not treated at two time periods, whereas a treatment group is treated at the second period. In such a setup, $T=G\cdot I$, $G=\left\{ 0,1\right\} $, $ I=\left\{ 0,1\right\} ,$ where $G$ is equal to 1 for the treatment group and 0 otherwise, and $I$ is a time indicator such that $I=0$ for the pre-treatment period and $I=1$ for the post-treatment period. Covariates $X$ are not available.
In this setup, one is usually interested in estimating the average, distributional, and quantile treatment effects for the treated subpopulation,
In a seminal work, Athey2006 show that, although the classical DID model as in Card1994 may not be adequate to estimate treatment effects beyond the average, a generalization of the DID model, the changes-in-changes (CIC) model, can be used to nonparametrically identify the $ATT$, $DTT\left( y\right) $ and $QTT\left( \tau \right) $. More specifically, Athey2006 show that, under some conditions to be discussed in Section (ref),
where, for $g=\left\{ 0,1\right\} $, $j=\left\{ 0,1\right\} $, $Y_{gj}$ are the realized outcome $Y$ conditional on $G=g$ and $I=j$, and $ F_{Y_{gj}}\left( y\right) =\mathbb{E}\left( 1\left\{ Y_{gj}\leq y\right\} \right) $, and $F_{Y_{gj}}^{-1}\left( \tau \right) =\inf \left\{ y:F_{Y_{gj}}\left( y\right) \geq \tau \right\} $.
Different from previous examples, not all terms in ((ref))-((ref)) are indexed by unknown functions, and when they do, there is more than one nuisance function. That is, ((ref)) is the difference between $\mathbb{E} \left[ Y_{11}\right] ,$ which does not depend on nuisance parameters, and $ \mathbb{E}\left[ F_{Y_{01}}^{-1}(F_{Y_{00}}\left( Y_{10}\right) \right] $, where $F_{Y_{01}}^{-1}$ and $F_{Y_{00}}$ play the role of $h$ here. Moving to ((ref)), $y$ plays the role of $z$, and $F_{Y_{01}}^{-1}$ and $ F_{Y_{01}}$ play the role of $h$. Finally, as in Examples 2.1 and 2.2, ((ref)) is a consequence of ((ref)). Thus, ((ref))-((ref)) fit into our framework.
Let $\left( Y,X,T\right) \in $ $\mathcal{Y\times X\times T\subseteq }$ $ \mathbb{R}\times \mathbb{R}^{k}\times \left\{ 0,1\right\} $, $F\left( y,x,t\right) \equiv \mathbb{P}\left( Y\leq y,X\leq x,T\leq t\right) $, and $ \varphi _{z,h_{0}}\left( Y,X,T\right) $ be a generic known, measurable, real-valued function indexed by $z\in \mathcal{W\subseteq Y\times X\times T}, $ and by potentially infinite dimensional nuisance parameters $h_{0}\in \mathcal{H}$\thinspace , where $\mathcal{H}$ is a Banach space with the supremum norm. Our goal is to make inference about ((ref)), but due to censoring mechanism, instead of always $Y$, one observes $Q=\min \left( Y,C\right) $, together with the non-censoring indicator $\delta =1\left\{ Y\leq C\right\} $. Hence, the available data consist of a random sample $ \left\{ \left( Q_{i},\delta _{i},X_{i},T_{i}\right) \right\} _{i=1}^{n}$ from $\left( Q,\delta ,X,T\right) $, and not $\left\{ \left( Y_{i},,X_{i},T_{i}\right) \right\} _{i=1}^{n}$ from $\left( Y,X,T\right) $. In this section, we discuss how one can identify and estimate ((ref)) with censored outcomes. Throughout the rest of this paper, all random variables are defined on a common probability space $\left( \Omega ,\mathcal{ A},\mathbb{P}\right) $.
We make the following assumption about the censoring mechanism.
Assumption (ref) states that, conditionally on the treatment status, the outcome of interest is independent of the censoring random variable, and that, given the underlying duration\ $Y$ and treatment status $T$, the covariates do not provide any further information whether censoring will take place, that is, $\delta $ and $X$ are conditionally independent given $Y$ and $T$. For instance, a particular case in which Assumption (ref) is satisfied is when $C$ is independent of $\left( Y,X,T\right) $, as assumed by e.g. Honore2002, Lee2005, Blundell2007, and Frandsen2014. It is important to have in mind that Assumption (ref) is more general than this particular case; it does not impose any restriction on how $Y$ and $C$ depends on $T$, and it allows some dependency between $C$ , $T$ and $X.$ Overall, such an assumption is not restrictive when censoring is fixed, or when the data comes from standard follow-up studies.
Next, we discuss the identification of ((ref)) with randomly-censored data when Assumption (ref) is satisfied. Denote $ H_{t}\left( y\right) =\mathbb{P}\left( Q\leq y|T=t\right) $, $G_{t}\left( y\right) =\mathbb{P}\left( C\leq y|T=t\right) $ and $H_{1t}\left( y,x\right) =\mathbb{P}\left( Q\leq y,X\leq x,\delta =1|T=t\right) $. Under Assumption (ref), the joint cumulative hazard function for the subpopulation with $\left\{ T=t\right\} $ is given by\footnote{ To see this, note that the probability that a random individual, taken at random from subpopulation $\left\{ T=t,Y\geq \bar{y}\right\} $, exits the state of interest\ before $\bar{y}+dy$ and have characteristics $\left\{ X\leq x\right\} $ is $\mathbb{P}\left( \bar{y}\leq Y<\bar{y}+dy,X\leq x|Y\geq \bar{y},T=t\right) =\left[ F_{t}\left( \left( \bar{y}+dy\right) -,x\right) -F_{t}\left( \bar{y}-,x\right) \right] /\left[ 1-F_{t}\left( \bar{ y}-,\infty \right) \right] $. The desired result is achived by integration.}
where $F_{t}\left( y,x\right) \equiv $ $\mathbb{P}\left( Y\leq y,X\leq x|T=t\right) $ and for any generic function $J$, $J\left( y-\right) =\lim_{a\uparrow y}J\left( a\right) $, and $J\left\{ y\right\} =J\left( y\right) -J\left( y-\right) $. For $t\in \left\{ 0,1\right\} $, let $\tau _{H_{t}}=\inf \left\{ y:H_{t}\left( y\right) =1\right\} $, $\tau _{F_{t}}=\inf \left\{ y:F_{t}\left( y,\infty \right) =1\right\} $, $\tau _{G_{t}}=\inf \left\{ y:G_{t}\left( y\right) =1\right\} $ be the least upper bound of the support of $H_{t}\left( \cdot \right) ,$ $F_{t}\left( \cdot ,\infty \right) $ and $G_{t}\left( \cdot \right) $, respectively. Let $\tau _{H}=\min \left( \tau _{H_{0}},\tau _{H_{1}}\right) $, and $\tau _{F}$ and $ \tau _{G}$ are defined analogously.
Next proposition shows that, under Assumption (ref) , we can identify $F\left( y,x,t\right) ,$ which is key to establish the identification of ((ref)). In contrast to \textquotedblleft inverse probability of censoring\textquotedblright\ (IPC)\ literature, see e.g. Robins1992, Vanderlaan2003, and references therein, our identification results do not require that $G_{t}\left( \cdot \right) <1$ $ a.s.,t\in \left\{ 0,1\right\} $, nor relies on continuity assumptions on $Y$ and $C$.
From Proposition (ref) one can see that the joint cumulative hazard plays a major role in the identification of $F\left( y,x,t\right) $. Once we establish that $\Lambda \left( y,x|T=t\right) $ can be written in terms of $ \left( Q,\delta ,X,T\right) $, we just need to plug in $\Lambda ^{cens}\left( y,x|T=t\right) $ into ((ref)) to recover $F\left( y,x,t\right) $. Another important implication of Proposition (ref) is that nonparametric identification of $F\left( y,x,t\right) $ over the entire support of $Y$ may not be feasible. This is intuitive since outcomes beyond $\tau _{H}=\min \left( \tau _{F},\tau _{G}\right) $ are never observed for both treatment and control groups. Such restriction is important, because it implies that the general moment condition ((ref)) will be identified only if one of the following conditions holds:
In order to better understand these conditions, notice that Condition (ref) implies that\ $\tau _{H}=\tau _{F}$. It turns out that the support of the censoring random variable being larger than or equal to the support of the outcome of interest is a necessary and sufficient condition for identifying $F\left( y,x,t\right) $ over its entire support. In fact, Condition (ref) can only be dispensed for identification of ((ref) ) if $\varphi _{z,h_{0}}$ satisfies Condition (ref). When $\tau _{H}=\tau _{G}<\tau _{F}$ outcomes beyond $\tau _{G}$ are never observed, and because $\mathbb{P}\left( \tau _{G}<Y\leq \tau _{F}\right) >0$ $a.s.$, identification of ((ref)) can only be attained if $\varphi _{z,h_{0}}\left( Y,X,T\right) =0$ in $\left[ \tau _{G},\tau _{F}\right] $. If neither Condition (ref) nor Condition (ref) is satisfied, one can only nonparametrically point-identify a truncated version of ((ref) ). Hence, identification of ((ref)) depends mainly on two things: the support of $Y$ and $C$, and the type of function $\varphi _{z,h_{0}}$ one is willing to analyze.
Proposition (ref) can also be exploited for estimation purposes. Intuitively, to estimate $F\left( y,x,t\right) $ we need to estimate $ \Lambda ^{cens}\left( y,x|T=t\right) $ and $\mathbb{P}\left( T=t\right) ,$ and plug in these estimators into ((ref)). But notice that $\Lambda ^{cens}\left( y,x|T=t\right) $ only depends on $H_{1t}\left( y,x\right) $ and $H_{t}\left( y\right) $, and both can be estimated by their sample analogues
where, for $t\in \left\{ 0,1\right\} $, $n_{t}=\sum_{i=1}^{n}1\left\{ T_{i}=t\right\} $. Hence, $\Lambda ^{cens}\left( y,x|T=t\right) $ can be estimated by
where $Q_{1:n_{t}}\leq $ $\cdots \leq Q_{n_{t}:n_{t}}$ are the ordered $Q$ -values in the subpopulation with $\left\{ T=t\right\} $, and $X_{\left[ i:n_{t}\right] }$, $\delta _{\left[ i:n_{t}\right] }$ are the concomitants of the $ith$ order statistics in the $t^{th}$ subpopulation, that is, the $X$ and $\delta $ paired with $Q_{i:n_{t}}$. Since $\hat{\Lambda} _{n}^{cens}\left( y,x|T=t\right) $ is purely discrete and that $\mathbb{P} \left( T=t\right) $ can be estimated by $n_{t}/n$, by plugging ((ref)) and $n_{t}/n$ into ((ref)) we have that
Although ((ref)) seems to have a complicated formula, in the next corollary we show that this is not the case, that ((ref)) can be written as a simple data-driven weighted average.
Corollary (ref) is important because it shows that, in practice, one does not need to first estimate $\Lambda ^{cens}\left( y,x|T=t\right) $ to get an estimator for $F\left( y,x,t\right) $. This is automatically achieved by the weights $W_{in_{t}}$. Additionally, in the absence of covariates $(x=\infty )$ and treatments ($n_{0}=n$, and $n_{1}=0)$, ((ref)) reduces to the time-honored Kaplan1958 product limit estimator of $F\left( y,\infty ,\infty \right) ,$
cf. Stute1993a and Stute1993. Thus, we argue that ((ref)) can be viewed as a multivariate extension of the Kaplan1958 product limit estimator, where the treatment status may affect the censoring and the outcome distribution in an arbitrary way.
With $\hat{F}_{n}^{km}\left( y,x,t\right) $ at hands, one can estimate ((ref)) by
where $\hat{h}_{n}$ is a generic first-step estimator for the unknown nuisance parameter $h_{0}$. The estimator in ((ref)) is what we refer as the two-step Kaplan-Meier estimator for ((ref)).
It is clear from ((ref)) that the 2SKM estimator has a close form representation, does not depend on functional form assumptions on the joint distribution $Y$, $X$ and $T$, and is fully data-driven upon estimation of the nuisance parameters $h_{0}$. Furthermore, in the absence of censoring, $ W_{in_{t}}=n^{-1}$ $a.s.$, implying that ((ref)) collapses to
the sample analogue of ((ref)). Hence, one can clearly see that indeed the 2SKM estimator ((ref)) is a natural extension of ((ref)) to the cases in which our outcome of interest is subjected to random right-censoring.
In this section we derive the asymptotic properties of the 2SKM estimator ( (ref)). We adopt the following notation: for a generic set $\mathcal{G}$ , let $l^{\infty }\left( \mathcal{G}\right) $ be the Banach space of all uniformly bounded real functions on $\mathcal{G}$ equipped with the uniform metric $\left\Vert f\right\Vert _{\mathcal{G}}\equiv \sup_{z\in \mathcal{G} }\left\vert f\left( z\right) \right\vert $. Let $\mathcal{W}\mathcal{ \subseteq }$ $\left( -\infty ,\tau _{H}\right) \times \mathbb{R}^{k}$ $ \times \left\{ 0,1\right\} $. We study the weak convergence of ((ref)) and related processes as elements of $l^{\infty }\left( \mathcal{W}\right) $ . Let $\Rightarrow $ denote weak convergence on $\left( l^{\infty }\left( \mathcal{W}\right) ,\mathcal{B}_{\infty }\right) $ in the sense of J. Hoffmann-J$\phi $rgensen, where $\mathcal{B}_{\infty }$ denotes the corresponding Borel $\sigma $-algebra - cf. VanderVaart1996.
For a generic $h\in \mathcal{H},$ $z\in \mathcal{W}$, define
Therefore, $S^{\varphi }\left( z,h_{0}\right) $ and $\hat{S}_{n}^{\varphi }\left( z,\hat{h}_{n}\right) $ are respectively equal to the target function ((ref)) and its 2SKM estimator ((ref)).
In the following, we derive a set of sufficient conditions under which $\hat{ S}_{n}^{\varphi }\left( z,\hat{h}_{n}\right) $ is uniformly consistent, and converges weakly to a tight Gaussian process. Furthermore, we show that one can use the ordinary nonparametric bootstrap to conduct asymptotically valid inference. These results are novel, complementing and extending those available in the literature on Kaplan-Meier integrals, cf. Stute1993a , Stute1993, Stute1995a, Stute1996a, Stute1996b, Stute2004, Stute2000, and Sellero2005.
For the 2SKM estimator in ((ref)) to be uniformly consistent, we state the following sufficient conditions.
Assumptions (ref)-(ref) are standard requirements in two-step estimation procedures, cf. Chen2003, and are not related to the censoring problem. Assumption (ref) requires consistent estimation of the nuisance parameters $h_{0}$. Assumption (ref) is a standard continuity condition, and is weaker than directly imposing a continuity assumption in $\varphi _{z,h}\left( Y,X,T\right) $. Finally, Assumption (ref) put some restrictions on the class of functions $\varphi _{z,h}\left( Y,X,T\right) .$ Now we state the uniform consistency for $\hat{S }_{n}^{\varphi }\left( z,\hat{h}_{n}\right) $.
Theorem (ref) is the first main and new result of the paper. It shows that under some relatively weak regularity conditions our 2SKM estimator satisfies a uniform law of large numbers.
To derive the limiting distribution of $\hat{S}_{n}^{\varphi }\left( z,\hat{h }_{n}\right) $, we impose the following sufficient conditions:
Assumptions (ref)-(ref) are not related to the censoring problem, and are standard in two-step estimation procedures, cf. Chen2003. Assumption (ref) strengthens Assumption (ref) such that the estimator of the nuisance parameter converges at a rate faster than $n^{-1/4}$. Assumption (ref) is a smooth condition for $ S^{\varphi }\left( z,h_{0}\right) $ that strengthens Assumption (ref). Assumption (ref) imposes additional restrictions on $ \varphi _{z,h},$ and it may be verified by using Theorem 3 of Chen2003 , for example. Assumption (ref) is related to the estimation of the nuisance parameter $h_{0}$, and it is a sufficient condition to $\sqrt{n} \left( \Gamma ^{\varphi }\left( z,h_{0}\right) [\hat{h}_{n}-h_{0}]\right) $ converge weakly. It assumes that $\sqrt{n}\left( \Gamma ^{\varphi }\left( z,h_{0}\right) [h-h_{0}]\right) $ is a smooth linear functional of $ [h-h_{0}],$ and that one can use a functional central limit theorem in its linear representation. When $h_{0}$ is consistently estimated by parametric methods, Assumption (ref) will be satisfied under mild integrability and smoothness conditions. When $h_{n}$ is nonparametric and has a closed form expression, under mild conditions, one can use the Riesz representation approach to obtain $\kappa ^{\varphi }\left( z,h_{0}\right) $. Once $\kappa ^{\varphi }\left( z,h_{0}\right) $ is obtained, Assumption (ref)$(ii)$ can be verified using empirical process theory, cf. VanderVaart1996.
It turns out that Assumptions (ref)-(ref) are not sufficient to derive the asymptotic distribution of $\hat{S}_{n}^{\varphi }\left( z, \hat{h}_{n}\right) .$We need some additional conditions due to the censoring problem. Define
where, for a generic $h\in \mathcal{H}$,
and $H_{t}\left( y\right) $ and $H_{1t}\left( y,x\right) $ are defined as before and $H_{0t}\left( y\right) =\mathbb{P}\left( Q\leq y,\delta =0|T=t\right) .$
Assumption (ref) is a modified \textquotedblleft finite second moment\textquotedblright\ condition for censored data. In the absence of censoring, such condition reduces to $\sup_{z}\left\vert \mathbb{E}\left[ \varphi _{z,h_{0}}\left( Y,X,T\right) ^{2}\right] \right\vert <\infty $. Assumption (ref) guarantees that ((ref)) has a finite variance. Assumption (ref) is to control the bias of the $\hat{S} _{n}^{\varphi }\left( z,h_{0}\right) $. Although the bias of $\hat{S} _{n}^{\varphi }\left( z,h_{0}\right) $ converges to $0$, the rate of convergence may be faster than $\sqrt{n}$, and Assumption (ref) guarantees that the bias is of the order $o\left( n^{-1/2}\right) $. This issue has been discussed in detail in Stute1994a. Whenever Condition (ref) is satisfied, Assumptions (ref) and (ref) will be satisfied provided that $\sup_{z}\left\vert \mathbb{E}\left[ \varphi _{z,h_{0}}\left( Y,X,T\right) ^{2}\right] \right\vert <\infty $, which is implied by Assumption (ref). However, this is not necessarily the case for a generic $\varphi _{z,h_{0}}$ when Condition (ref) is not satisfied.
Next theorem presents the weak convergence result for the 2SKM estimator.
Theorem (ref) is the second main and new result of the paper. It shows that under some relatively weak regularity conditions our 2SKM estimator satisfies a functional central limit theorem. This result forms the basis of all inference results on policy evaluation with censored data.
As an application of the result above, we can show that plug-in estimators of Hadamard differentiable functionals also satisfy functional central limit theorems. Examples include quantile curves, as well as Lorenz curves, and Gini coefficients.
From Theorem (ref) we have that the asymptotic covariance function ((ref)) depends on the underlying data generating process and standardization can be complicated. To see this, note that in order to estimate $V^{\varphi }\left( \cdot ,\cdot \right) $, one needs to estimate $\eta ^{\varphi }$ and $\kappa ^{\varphi }$, plug in our estimator for $S^{\varphi }\left( z,h_{0}\right) $, $\hat{S}_{n}^{\varphi }\left( z,\hat{h}_{n}\right) ,$ and then compute the sample second moment of these quantities. But in order to estimate $\eta ^{\varphi }$ one needs to estimate $\gamma _{0t}$, $\gamma _{1t,z,h}^{\varphi }$ and $\gamma _{2t,z,h}^{\varphi },t\in \left\{ 0,1\right\} $. Furthermore, different estimators of $\kappa ^{\varphi }$ could be needed depending on how one chooses to estimate $h_{0}$. It turns out that estimating these nuisance functions can be difficult, and may involve tuning parameters such as bandwidths, cf. SantAnna2016a. To avoid these issues, we follow an alternative route and use the ordinary nonparametric bootstrap to conduct asymptotically valid inference.
In order to compute the bootstrap confidence bands, let $B$ be a large integer. For each $b=1,\dots ,B:$
Then, the $\left( 1-\alpha \right) 100\%$ asymptotic confidence band is calculated as
where $c_{1-\alpha }^{B}$ denotes the empirical $\left( 1-\alpha \right) $ quantile of the simulated sample $\left\{ L^{b,\varphi }\right\} _{b=1}^{B}$ . In practice, the maximum in step 3 is taken over a discretized subset $ \mathcal{W}.$
Next, we establish the asymptotic validity of the aforementioned bootstrap procedure considering the following additional conditions on the nuisance parameters. Here and subsequently, superscript $\ast $ denotes probability or moment computed under the bootstrap distribution conditional on the original data set.
Theorem (ref) is the third main and new result of the paper. It shows that the limiting distribution of the bootstrap estimator is the same as that of Theorem (ref), and hence, our proposed resample scheme is able to mimic the asymptotic distribution of interest. Such a result is very powerful and allows one to use the ordinary nonparametric bootstrap to conduct asymptotically valid inference.
By combining Theorem (ref) with the functional delta method for the bootstrap, cf. Theorem 3.9.11 in VanderVaart1996, we can show the bootstrap validity of plug-in estimators of Hadamard differentiable functionals as well.
In this section we illustrate the general applicability of our 2SKM approach by revisiting the motivating examples of Section (ref). In short, we show that, under relatively weak regularity conditions, one can consistently estimate, and construct asymptotically valid confidence bands for the average, distributional, and quantile treatment effects discussed in Examples 2.1, 2.2 and 2.3 when the outcomes is randomly censored. These results are novel to the literature, and are obtained by verifying the high-level conditions in Theorems (ref)-(ref).
We use the same potential outcome notation as in Section (ref), but due to the censoring mechanisms, instead of observing $Y$, one observes $ Q\equiv TQ_{1}+(1-T)Q_{0},$ where $Q_{0}=\min \left\{ Y_{0},C_{0}\right\} $, $Q_{1}=\min \left\{ Y_{1},C_{1}\right\} $, $C_{0}$ and $C_{1}$ being potential censoring random variables under the control and treatment groups, respectively. In addition to $Q$, one also observes the censoring indicator $ \delta \equiv T\delta _{1}+\left( 1-T\right) \delta _{0}$, where, for $t\in \left\{ 0,1\right\} $, $\delta _{t}=1\left\{ Y_{t}\leq C_{t}\right\} $. It is important to emphasize that, in the following, we can accommodate covariates, allow the treatment status to affect the censoring variable in an arbitrary way, and we do not impose the potentially restrictive condition that censoring variable $C$ is always observed.
We first revisit unconfoundedness setup discussed in Example 2.1. We impose the following conditions.
Assumptions (ref)$(i)$ and $(ii)$ are standard in the literature, cf. Rosenbaum1983, Hirano2003, Ichimura2005, Firpo2007, Donald2013, among others. If censoring is not present, Assumptions (ref)$(i)$ and $(ii)$ suffice to identify our treatment effects of interest. Nonetheless, censoring introduces another source of confounding because the probability of censoring is related to potential outcomes. This additional identification challenge can be overcome under Assumption (ref)$(iii)$, the analogous of Assumptions (ref) in the unconfoundedness context\footnote{ As discussed in Supplemental Appendix, such an assumption is testable as long as one imposes additional smoothness and support restrictions in the DGP.}.
In the absence of censoring, Rosenbaum1983, Hirano2003, Ichimura2005, Firpo2007, Donald2013, among others, have proposed estimators for ((ref))-((ref)), where one first estimate $ p\left( \cdot \right) $ by parametric or nonparametric methods, plugs it into ((ref))-((ref)), and then use the analogy principle to estimate ((ref))-((ref)). As we have seen in Section (ref), although such a procedure is not feasible when $Y$ is subject to censoring mechanisms, one can use the 2SKM procedure to overcome this issue. That is, under Assumption (ref), one can use the 2SKM methodology, and estimate ((ref))-((ref)) by
respectively, where
$\hat{F}_{n,Y_{t}}^{km,-1}\left( \tau \right) $ is the empirical $\tau $ -quantile of the rearrangement of $\hat{F}_{n,Y_{t}}^{km}\left( y\right) $ if $\hat{F}_{n,Y_{t}}^{km}\left( y\right) $ is not monotone, cf. Chernozhukov2010, $t\in \left\{ 0,1\right\} $, and $\hat{p}_{n}\left( \cdot \right) $ is a first-step estimator for the propensity score $p\left( \cdot \right) .$ Here, for $1\leq i\leq n_{1}$, $Q_{i:n_{1}}$ is the $ith$ order statistics in the treated subsample, and $X_{\left[ i:n_{1}\right] }$ is the concomitants of the $ith$ order statistics in the treated subpopulation; $Q_{i:n_{0}}$ and $X_{\left[ i:n_{0}\right] }$ are defined analogously but for the control subsample.
In practice, one can estimate $p\left( \cdot \right) $ by parametric, semi-parametric or nonparametric methods, e.g. Rosenbaum1983, Hahn1998, Hirano2003 and Ichimura2005. Nonetheless, it is important to have in mind that different regularity conditions might be needed depending on the estimation method you use. In the Appendix we discuss these conditions for three popular estimators of $p\left( \cdot \right) $: the parametric estimator (e.g. Logit or Probit specifications), the nonparametric leave-one-out Nadaraya-Watson kernel-based estimator, and the nonparametric Logit Series estimator. We can show that as long as the required regularity (smooth) conditions are met, the 2SKM estimators ((ref))-((ref)) are uniform consistent, converge weakly, and the ordinary nonparametric bootstrap procedure can be used to conduct asymptotically valid inference. These results are summarized in the next proposition.
The results in Proposition (ref) are new to the literature. To the best of our knowledge, the only available results related to Proposition (ref) are Hubbard2000, who, for a fixed $y$, proposes an alternative estimator for the $DTE\left( y\right) $ that relies on a parametric specification for the propensity score, and Anstrom2001 who builds on Hubbard2000 and proposes an estimator for the $ATE$. Nonetheless, it is important to notice that the results in Proposition (ref) go beyond this particular case: it allows one to use nonparametric estimators of the propensity score, and justify the use of the bootstrap to conduct uniform asymptotically valid inference. On one hand, allowing for the propensity score to be estimated by nonparametric methods can be particularly important for two reasons: $(a)$ as shown by Huber2013, misspecification of the propensity score may lead to severe distortion on the policy evaluation parameters of interest; and $(b)$ as shown by Hirano2003 and Chen2008, even when the propensity score is correctly specified, using nonparametric estimates can lead to efficiency gains. On the other hand, since our bootstrapped confidence sets are uniformly valid in the sense that they cover the entire functional of interest with pre-specified probability, they can be used to test functional hypotheses such as no-effect, positive effect, or stochastic dominance, cf. Abadie2002.
This section proposes and derives the asymptotic properties of 2SKM estimators of the local average, distributional and quantile treatment effects described in Example 2.2. To do so, we need to introduce additional notation. Let $Y_{0},Y_{1},Q_{1},Q_{0},\delta _{1},\,\delta _{0},C_{0},C_{1}$ and $X$ be defined as in the unconfoundedness framework. The local treatment effect (LTE) setup presumes the availability of a binary instrumental variable $Z$ for the treatment assignment. Denote $T_{0}$ and $T_{1}$ the values that $T$ would have taken if $Z$ is equal to zero or one, respectively. The realized treatment is $T=ZT_{1}+\left( 1-Z\right) T_{0}.$ Thus, the observed sample consist of $iid$ copies $\left\{ Q_{i},\delta _{i},X_{i},T_{i},Z_{i}\right\} _{i=1}^{n}$ of $(Q,\delta ,X,T,Z)$. Denote $ e\left( X\right) \equiv \mathbb{P}(Z=1|X)$.
In order to identify the LTE for the subpopulation of compliers, we impose the following assumptions.
Assumption (ref)$\left( i\right) $-$\left( iii\right) $ are standard, cf. Abadie2003 and Frolich2013\footnote{ Although standard in the literature, Assumption (ref)$(iii)$ can be relaxed, see DeChaisemartin2014 for details.}. Assumption (ref) $(iv)$ is related to the censoring mechanisms and is the analogous of Assumption (ref) in the LTE context; it solves the additional identification challenge that censoring introduces into the LTE setup. It is important to notice that Assumption (ref) does not restrict how treatment status and instruments affects the censoring variable, which is weaker than the assumptions commonly used in the literature, cf. Frandsen2014.
In the absence of censoring, Abadie2003, Frolich2007, and Frolich2013 propose estimators for ((ref))-((ref)). Although their procedures are not feasible when $Y$ is censored, we know from the discussion in Sections (ref) and (ref) that, under Assumption (ref), we can apply the 2SKM procedure to estimate ((ref))-((ref)) in the present context.
The first step towards estimating ((ref))-((ref)) is to estimate $ e\left( \cdot \right) $. Noticing that the available instrument $Z$ for $T$ is binary, one can treat $e\left( \cdot \right) $ as an \textquotedblleft instrumental propensity score\textquotedblright\ and estimate it using parametric models such as the Logit or Probit specification, or using nonparametric Kernel or Series estimators as described in the Appendix. We denote the estimator of $e\left( \cdot \right) $ by $\hat{e}_{n}\left( \cdot \right) $.
With $\hat{e}_{n}\left( \cdot \right) $ at hands, the next task is to estimate ((ref)) and ((ref)) with censored outcomes. First, there is no (new) challenge into estimating $\kappa _{t}\left( e\right) $ because one can simply use its sample analogue,
Next, we plug in $\hat{\kappa}_{t,n}\left( \hat{e}_{n}\right) $ into ((ref)) and ((ref)), and by using our Kaplan-Meier approach to handle the censoring problem, we estimate ((ref)) and ((ref)) by
where $n_{tz}=\sum_{i=1}^{n}1\left\{ T=t\right\} 1\left\{ Z=z\right\} $, $ z\in \left\{ 0,1\right\} $, and for $1\leq i\leq n_{tz},$ $Q_{1:n_{tz}}\leq $ $\cdots \leq Q_{n_{td}:n_{tz}}$ are the ordered $Q$-values in the subsample with $\left\{ T=t,Z=z\right\} $, $X_{\left[ i:n_{tz}\right] }$ and $\delta _{ \left[ i:n_{tz}\right] }$ are the $X$ and $\delta $ paired with $Q_{i:n_{tz}} $, and
is the Kaplan-Meier weights for the subsample with $\left\{ T=t,Z=z\right\} $ . Once such measures are available, our 2SKM estimators for ((ref))-( (ref)) are given by
where, for $t\in \left\{ 0,1\right\} $, $\mathbb{E}_{n}^{km}\left[ Y_{t}^{c} \right] $ is given by ((ref)), $\hat{F}_{n,Y_{t}^{c}}^{km}\left( y\right) $ is given by ((ref)) and $\hat{F}_{n,Y_{t}^{c}}^{km,-1}\left( \tau \right) =\inf \left( y:\hat{F}_{n,Y_{t}^{c}}^{km,r}\left( y\right) \geq \tau \right) $, where $\hat{F}_{n,Y_{t}^{c}}^{km,r}\left( y\right) $ denotes the rearrangement of $\hat{F}_{n,Y_{t}^{c}}^{km}\left( y\right) $ if $\hat{F} _{n,Y_{t}^{c}}^{km}\left( y\right) $ is not monotone, cf. Chernozhukov2010\footnote{ To construct ((ref))-((ref)), we split the sample into four sub-samples depending on the treatment status $T$ and on the value of the instrument $D$. This is necessary because Assumptions (ref)$\left( iv\right) $-$\left( v\right) $ does not impose any restriction on how $T$ and $D$ affect the censoring probability. If one is willing to strengthen Assumptions (ref)$\left( iv\right) $-$\left( v\right) $ to the case in which these assumptions hold unconditionally on $D$, one would need to split the sample only on treated and control groups, like in the unconfoundedness setup. For the sake of generality, we avoid doing so.}.
Next proposition shows that the 2SKM estimators ((ref))-((ref) ) are uniformly consistent, converge weakly, and one can use the bootstrap to perform asymptotically valid inference. Let
The results in Proposition (ref) are novel to the literature. To the best of our knowledge, the only related results to Proposition (ref) is Frandsen2014, who proposes estimators for the distributional and quantile treatment effects ((ref)) and ((ref)), but in the much simpler setup than ours: Frandsen2014's proposal cannot accommodate covariates, relies on the censoring variable being always observed, and requires appropriate support restrictions that excludes from the analysis some functionals of interest such as the $ATE$. Furthermore, even when Frandsen2014 putative conditions are satisfied, one can show that our 2SKM estimators are more efficient than his, even though the 2SKM estimator does not use the full sample of $C_{i}$ values, cf. Portnoy2010. These features highlights the flexibility and power of our proposal.
In this section we propose 2SKM estimators for ((ref))-((ref)) in the Changes-in-Changes (CIC) setup described in Example 2.3. We make the following assumptions.
Assumptions (ref)$(i)$-$\left( vi\right) $ define the CIC classical setup of Athey2006. Assumption (ref)$(vii)$ is related to the censoring mechanism, and states that conditionally on the group status and on the time period, the outcome of interest is independent of the censoring random variable.
The first step towards estimating ((ref))-((ref)) is to estimate the nuisance functions $F_{Y_{01}}\left( \cdot \right) $ and $ F_{Y_{00}}^{-1}\left( \cdot \right) $. Notice that, in contrast with the unconfoundedness and local treatment effect setups, here the nuisance functions are affect by the censoring problem. Nonetheless, they can be estimated by their Kaplan-Meier analogues
$g\in \left\{ 0,1\right\} $, $j\in \left\{ 0,1\right\} $, where $n_{gj}$ $ =\sum_{i=1}^{n}1\left\{ G_{i}=g\right\} 1\left\{ I_{i}=j\right\} ,$ $ Q_{1:n_{gj}}\leq $ $\cdots \leq Q_{n_{gj}:n_{gj}}$ are the ordered $Q$ -values in the subsample with $\left\{ G=g,I=j\right\} $, $X_{\left[ i:n_{gj} \right] }$ and $\delta _{\left[ i:n_{gj}\right] }$ are the $Q_{i:n_{gj}}$ concomitants, and for $1\leq i\leq n_{gj}$,
is the size of the Kaplan-Meier jump for observation $i$ in the subsample with $\left\{ G=g,I=j\right\} $. Notice that these nonparametric estimators are fully data-driven, and do not require the use of tuning parameters such as bandwidths.
With the first-step estimators at hands, we can use our 2SKM approach to estimate ((ref))-((ref)). More precisely, we propose to estimate ( (ref))-((ref)) by
where
and $\hat{F}_{n,Y_{0}|T=1}^{km,r}\left( y\right) $ denotes the rearrangement of $\hat{F}_{n,Y_{0}|T=1}^{km}\left( y\right) $ if $\hat{F} _{n,Y_{0}|T=1}^{km}\left( y\right) $ is not monotone, cf. Chernozhukov2010.
In the next proposition we show that the 2SKM estimators ((ref))-((ref)) are uniformly consistent, converge weakly, and one can use the ordinary nonparametric bootstrap to perform asymptotically valid inference. Let
These results in Proposition (ref) are new even when censoring is not an issue. First, it generalizes Athey2006 pointwise results to hold uniformly. Second, it proves that one can use the bootstrap to perform inference in the CIC setup. Both of these points are of practical relevance: $(i)$ because our results hold uniformly, one can test for first-or second-order stochastic dominance in the same spirit of Abadie2002; $ \left( ii\right) $ by using bootstrapped confidence intervals to conduct inference on $QTT\left( \cdot \right) $, one completely avoids the need of estimating density functions to construct standard errors, a task that would involve choosing tuning parameters. Proposition (ref) shows that these desirable features naturally carry out to the randomly censored CIC setup.
In this section, we conduct a small scale Monte Carlo exercise in order to study the finite sample properties of our proposed policy evaluation estimators. More precisely, we compare the performance of the two-step Kaplan-Meier (2SKM) estimators proposed here with those based on $(a)$ the \textquotedblleft naive\textquotedblright\ approach that uses inverse probability weighted (IPW) estimators ignoring that the outcome of interested is subjected to censoring (we label such an approach as \textquotedblleft Ignore \textquotedblright ); $\left( b\right) $ the \textquotedblleft naive\textquotedblright\ approach that uses IPW estimators after dropping all censored data (we label such an approach as \textquotedblleft Uncens \textquotedblright ); $\left( c\right) $ the Cox1972, Cox1975 Proportional hazard model for the treated and control groups (we label such an approach as \textquotedblleft Cox\textquotedblright ), in which we exploit the relationship between the conditional hazard rates, the conditional CDF's, and then integrate out the covariate vector to get the unconditional CDF's; and $\left( d\right) $ the Frandsen2014 's proposal (we label such an approach as \textquotedblleft Frandsen\textquotedblright )\footnote{ For detailed description on how to compute the policy evaluation parameters using these competing methods, see the Supplemental Appendix.}. For conciseness, we focus on the unconfoundedness setup.
We consider the following four designs:
where $X,\varepsilon _{0}$ and $\varepsilon _{1}$ are independently distributed as standard normals, and $\varepsilon _{c}$ is independently distributed as exponential with parameter $a_{c}$,\ where $a_{c}$ is chosen such that the percentage of censoring in the sample is approximately equal to 10 or 30 percent. Note that, because $\varepsilon _{c}$ is exponentially distributed, censoring is more concentrated on the upper tail of the distribution, as is typically the case. All designs are adapted from Frandsen2014. Design $1$ is the baseline setup, in which potential outcomes do not depend on covariates, and the treatment effect is homogenous (constant) across the entire distribution. Design $2$ introduces heterogeneity by allowing the policy intervention to affect both the mean and the variance of the potential outcomes, whereas Design $3$ introduces heterogeneity by allowing potential outcomes to depend on covariates $X$. Design $4$ is the most \textquotedblleft heterogeneous\textquotedblright\ design: it combines Designs $2$ and $3$. In all designs, $P\left( T=1|X\right) =\exp (0.5X)/(1+\exp \left( 0.5X\right) )$, and $\mathbb{E} \left( Y_{1}\right) =F_{Y_{1}}^{-1}\left( 0.5\right) =1$, $\mathbb{E}\left( Y_{0}\right) =F_{Y_{0}}^{-1}\left( 0.5\right) =0$, implying that the $ ATE=QTE\left( 0.5\right) =1$. The observed data is $\left\{ Q_{i},\delta _{i},X_{i},T_{i}\right\} _{i=1}^{n}$, where $Q_{i}=\min \left( Y_{i},C_{i}\right) $ and $\delta _{i}=1\left\{ Y_{i}\leq C_{i}\right\} $. Nonetheless, in order to use Frandsen2014 approach, we assume that $ C_{i}$ is observed for both censored and uncensored observations, though we do not need such restrictive condition to compute the 2SKM, the \textquotedblleft naive approaches\textquotedblright , or the Cox based estimators.
The finite sample comparisons are based on bias\footnote{ In the Supplemental Appendix we also compare the root mean square errors of the competing methods.} for $\mathbb{E}\left( Y_{1}\right) $ , $\mathbb{E}\left( Y_{0}\right) ,$ $F_{Y_{1}}^{-1}\left( 0.5\right) $, $ F_{Y_{0}}^{-1}\left( 0.5\right) ,$ $ATE$ and $QTE\left( 0.5\right) $. When censoring is not present, the 2SKM estimators are numerically equivalent to those base on the \textquotedblleft naive approaches\textquotedblright . Thus, we report only the 2SKM, Cox, and Frandsen2014 estimators in these simulation setups. All simulations are based on a thousand Monte Carlo experiments, with a sample size of $n=1,000$ across all scenarios. We estimate $p\left( \cdot \right) $ using Hirano2003 series logit estimator with $1,X,X^{2},X^{3}$ as power functions.
The simulation results are presented in Table (ref). The simulations show that the proposed 2SKM estimators for $\mathbb{E}\left( Y_{1}\right) ,$ $\mathbb{E}\left( Y_{0}\right) ,$ $F_{Y_{1}}^{-1}\left( 0.5\right) $, $F_{Y_{0}}^{-1}\left( 0.5\right) ,$ $ATE$ and $QTE\left( 0.5\right) $ have minimal bias across all DGP's, and outperforms all other methods, specially when covariates play an important role. This is not surprising, since the 2SKM approach is the only appropriate method to estimate all measures of interest in the presence of censoring and covariates, without relying on functional form assumptions. Even when the potential outcomes do not depend on covariates, however, our proposed 2SKM estimators perform nearly as well as Frandsen2014's estimators, even though we make use of less information (we do not use $C_{i}^{\prime }s$ whatsoever). Such a feature stress the flexibility and appeal of our 2SKM estimators.
As expected, the \textquotedblleft naive\textquotedblright\ estimators that ignore the censoring issue, or use only uncensored observations are severely biased. Another feature worth mentioning is that estimators for $\mathbb{E} \left( Y_{1}\right) $, $\mathbb{E}\left( Y_{0}\right) $ and $ATE$ based on the Cox proportional hazard model have close to minimal bias, even though the model is misspecified (the conditional hazards are not proportional in the analyzed DGP's). However, the same is not true for the Cox estimators for the median. As discussed by Portnoy2003, this is due to the fact that the Cox model greatly restricts the behavior of the quantile effects, leading to inconsistent and severely biased estimates when the underlying assumptions of the model are not satisfied, as it is the case here. Finally, notice that Frandsen2014's estimators are unbiased in DGP's $1$ and $2 $, but are severely biased in DGP's $3$ and $4$. This is a simple consequence of Frandsen2014 not being able to accommodate covariates into the analysis, which turns out to be crucial in the last two DGP's.
In summary, our simulations highlights that our proposed 2SKM estimators exhibit very good finite sample properties in all analyzed designs. On the other hand, ignoring the censoring problem, imposing ad hoc functional form restrictions in the distribution of the potential outcomes, or not accommodating covariates into the analysis may lead to spurious conclusions about the policy effectiveness.
In this paper we proposed a class of Kaplan-Meier two-step estimators when the outcome of interest is subjected to right-censoring mechanisms. We provided sufficient conditions for the 2SKM estimator to be uniformly consistent and converge weakly to a tight Gaussian process with mean zero, and covariance function that may depend on the underlying DGP in rather complicated ways. To conduct asymptotically valid inference, we have shown that one can use the ordinary nonparametric bootstrap. We illustrate the relevance and applicability of our general results by proposing new average, distributional, and quantile treatment effects estimators under the unconfoundedness, local treatment effect, and changes-in-changes setups with censored outcomes.
Although we have focused on the aforementioned three policy setups, the proposed 2SKM tools can be applied to other designs such as multi-valued treatments, cf. Cattaneo2010; dynamic treatment effects, cf. Sianesi2004, Fredriksson2008, vandenberg2009, and Vikstrom2014; fuzzy differences-in-differences, cf. DeChaisemartin2015; distributional differences-in-differences, cf. Callaway2015, and Callaway2015a; and also to identify other parameters of interest such as the marginal treatment effects, cf. Heckman2001, Heckman2005; or to conduct Oaxaca-Blinder-type decompositions, cf. Fortin2011 for a review, and Garcia-Suaza2015 for related results with censored outcomes. In short, in this paper we have shown that, by using the 2SKM approach, many policy evaluation tools available for \textquotedblleft complete data\textquotedblright\ can be extended to accommodate randomly censored outcomes.